diff --git a/.gitattributes b/.gitattributes
index a6344aac8c09253b3b630fb776ae94478aa0275b..8c4bf3bb010c625c19c2c6000ac514a0603a5b25 100644
--- a/.gitattributes
+++ b/.gitattributes
@@ -33,3 +33,7 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
+figs/Main.png filter=lfs diff=lfs merge=lfs -text
+fonts/Rainbow-Party-2.ttf filter=lfs diff=lfs merge=lfs -text
+scripts/data_process/ruod/008431.jpg filter=lfs diff=lfs merge=lfs -text
+scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints.png filter=lfs diff=lfs merge=lfs -text
diff --git a/LICENSE b/LICENSE
new file mode 100644
index 0000000000000000000000000000000000000000..496a9540a703944ddf0164c03fb78bac6012dbeb
--- /dev/null
+++ b/LICENSE
@@ -0,0 +1,21 @@
+MIT License
+
+Copyright (c) 2026 CVTEAM
+
+Permission is hereby granted, free of charge, to any person obtaining a copy
+of this software and associated documentation files (the "Software"), to deal
+in the Software without restriction, including without limitation the rights
+to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
+copies of the Software, and to permit persons to whom the Software is
+furnished to do so, subject to the following conditions:
+
+The above copyright notice and this permission notice shall be included in all
+copies or substantial portions of the Software.
+
+THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
+IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
+FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
+AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
+LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
+OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
+SOFTWARE.
\ No newline at end of file
diff --git a/README.md b/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..d95945832cabeb5e0e4436837d090a45436be065
--- /dev/null
+++ b/README.md
@@ -0,0 +1,252 @@
+# Envisioning Beyond the Few: Disentangled Semantics and Primitives for Few-Shot Atypical Layout-to-Image Generation
+
+**ICML 2026**
+
+**Authors:** Nan Bao, Yifan Zhao, Wenzhuang Wang, Jia Li
+
+
+
+## Environment Setup
+
+We use two separate environments:
+
+1. **Main environment** for core training and inference.
+
+ ```bash
+ conda create -n dsp python=3.10.20
+ conda activate dsp
+ pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu126
+ pip install datasets==4.8.5 pillow==12.2.0 accelerate==1.13.0 transformers==5.8.1 diffusers==0.38.0 safetensors==0.8.0rc0 tensorboard==2.20.0 opencv-python==4.13.0.92 einops==0.8.2 imagesize==2.0.0 peft==0.19.1 ttach==0.0.3 ftfy==6.3.1 albumentations==2.0.8
+ ```
+
+2. **Evaluation environment** for MMDetection/MMEngine compatibility. It is used for evaluation with MMDetection/MMEngine due to strict version constraints, and also supports YOLO-based evaluation.
+
+ ```bash
+ conda create -n dsp-eval python=3.10.20
+ conda activate dsp-eval
+ conda install mkl==2023.1.0 numpy==1.26.4
+ conda install pytorch==2.1.2 torchvision==0.16.2 torchaudio==2.1.2 pytorch-cuda=12.1 -c pytorch -c nvidia
+ pip install mmengine==0.10.7 tqdm==4.67.3 shapely==2.1.2 scipy==1.15.3 terminaltables==3.1.10 ultralytics==8.4.50 pycocotools==2.0.11 https://download.openmmlab.com/mmcv/dist/cu121/torch2.1.0/mmcv-2.1.0-cp310-cp310-manylinux1_x86_64.whl "numpy<2.0.0" "setuptools<70.0.0"
+ ```
+
+## Set Environment Variables
+
+Set the root path of this project:
+
+```bash
+export DSP_PROJECT_DIR=/path/to/DSP # replace with the actual path
+```
+
+It is recommended to add this line to `~/.bashrc` or `~/.zshrc` for persistence.
+
+## Pretrained Models Preparation
+
+1. We use several pretrained models as external dependencies. Please download them manually from the following sources:
+ - [stable-diffusion-v1-5](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5)
+ - [clip-vit-large-patch14](https://huggingface.co/openai/clip-vit-large-patch14)
+ - [dinov2_vitl14_pretrain.pth](https://dl.fbaipublicfiles.com/dinov2/dinov2_vitl14/dinov2_vitl14_pretrain.pth)
+ - [ViT-B-16.pt](https://openaipublic.azureedge.net/clip/models/5806e77cd80f8b59890b7e101eabd078d9fb84e6937f9e85e4ecb61988df416f/ViT-B-16.pt)
+
+2. After downloading, organize the pretrained weights under `./pretrained` as follows:
+
+ ```bash
+ pretrained
+ ├── stable-diffusion-v1-5
+ │ └── ...
+ ├── clip-vit-large-patch14
+ │ └── ...
+ ├── dinov2_vitl14_pretrain.pth
+ └── ViT-B-16.pt
+ ```
+
+ You may either copy or symlink the files. We recommend using symbolic links:
+
+ ```bash
+ ln -s /path/to/stable-diffusion-v1-5 ./pretrained/stable-diffusion-v1-5
+ ln -s /path/to/clip-vit-large-patch14 ./pretrained/clip-vit-large-patch14
+ ln -s /path/to/dinov2_vitl14_pretrain.pth ./pretrained/dinov2_vitl14_pretrain.pth
+ ln -s /path/to/ViT-B-16.pt ./pretrained/ViT-B-16.pt
+ ```
+
+## Data Preparation
+
+1. We use several public datasets. Please download them manually from the following sources:
+
+ - [DIOR](https://gcheng-nwpu.github.io/#Datasets)
+ - [RUOD](https://github.com/xiaoDetection/RUOD)
+ - [ExDark](https://github.com/cs-chan/Exclusively-Dark-Image-Dataset/tree/master/Dataset)
+
+2. Unzip the downloaded datasets and organize the external dataset directories as follows:
+
+ ```bash
+ DIOR-VOC
+ ├── Annotations
+ │ ├── Horizontal_Bounding_Boxes
+ │ └── Oriented_Bounding_Boxes
+ └── VOC2007
+ ├── ImageSets
+ │ ├── Layout
+ │ ├── Main
+ │ └── Segmentation
+ └── JPEGImages
+ ```
+
+ ```bash
+ RUOD
+ ├── Environment_pic
+ │ ├── blur
+ │ ├── color
+ │ └── light
+ ├── Environmet_ANN
+ ├── RUOD_ANN
+ └── RUOD_pic
+ ├── test
+ └── train
+ ```
+
+ ```bash
+ ExDark
+ ├── annos
+ ├── imageclasslist.txt
+ └── images
+ ```
+
+3. Run data preprocessing scripts located in `./scripts/data_process`, after updating all hard-coded paths (e.g., `/path/to/DIOR_VOC`, `/path/to/RUOD`, `/path/to/ExDark`) in the scripts to match the local setup. Execute them in order.
+
+ The preprocessing outputs will be generated under `./data` with the following structure:
+
+ ```bash
+ data
+ ├── DIOR
+ │ ├── dior_emb.pt
+ │ ├── images -> /path/to/DIOR-VOC/VOC2007/JPEGImages
+ │ ├── metadatas
+ │ └── patches
+ ├── EXDARK
+ │ ├── exdark_emb.pt
+ │ ├── images
+ │ ├── metadatas
+ │ └── patches
+ └── RUOD
+ ├── images -> /path/to/RUOD/RUOD_pic
+ ├── metadatas
+ ├── patches
+ └── ruod_emb.pt
+ ```
+
+## Training and Inference
+
+We provide three example configurations in `./configs`: `dsp-dior.yaml`, `dsp-ruod.yaml`, and `dsp-exdark.yaml`.
+
+> **Argument Description:**
+> - **config:** configuration file for model and dataset setup.
+> - **metaseed:** seed generator identifier for deterministic sampling.
+> - **num_seed:** number of sampling seeds for few-shot evaluation.
+> - **k_shot:** number of samples per category in few-shot setting.
+> - **run_id:** identifier for different runs.
+> - **gpu_ids:** GPU device indices for execution.
+> - **iter:** number of bootstrap iterations for FID.
+
+### Base Phase Training
+
+```bash
+bash train_base.sh --config "dsp-dior"
+bash train_base.sh --config "dsp-ruod"
+bash train_base.sh --config "dsp-exdark"
+```
+
+### Novel Phase Training
+
+```bash
+bash train_novel.sh --config "dsp-dior" --metaseed "aaa" --num_seed 50 --k_shot "5" --run_id "1" --gpu_ids "0,1,2,3"
+bash train_novel.sh --config "dsp-ruod" --metaseed "aaa" --num_seed 50 --k_shot "5" --run_id "1" --gpu_ids "0,1,2,3"
+bash train_novel.sh --config "dsp-exdark" --metaseed "aaa" --num_seed 50 --k_shot "5" --run_id "1" --gpu_ids "0,1,2,3"
+```
+
+### Inference
+
+```bash
+bash infer.sh --config "dsp-dior" --metaseed "aaa" --num_seed 50 --k_shot "5" --run_id "1" --ckpt "100" --gpu_ids "0,1,2,3" --max_infer_size 50
+bash infer.sh --config "dsp-ruod" --metaseed "aaa" --num_seed 50 --k_shot "5" --run_id "1" --ckpt "100" --gpu_ids "0,1,2,3" --max_infer_size 50
+bash infer.sh --config "dsp-exdark" --metaseed "aaa" --num_seed 50 --k_shot "5" --run_id "1" --ckpt "100" --gpu_ids "0,1,2,3" --max_infer_size 50
+```
+
+## Evaluation
+
+### Preparation
+
+Download the YOLO and Faster R-CNN weights from [this link](https://drive.google.com/drive/folders/1FWN02KEuGPdEkXv38MmT8-D4-uQcAn4_?usp=sharing). Place them under `./pretrained`. The expected directory structure is as follows:
+
+```bash
+pretrained
+├── evaluation
+│ ├── mmdet
+│ │ ├── faster_rcnn_r50_fpn_1x-dior
+│ │ │ └── epoch_12.pth
+│ │ ├── faster_rcnn_r50_fpn_1x-exdark
+│ │ │ └── epoch_12.pth
+│ │ └── faster_rcnn_r50_fpn_1x-ruod
+│ │ └── epoch_12.pth
+│ └── yolo
+│ └── best.pt
+└── ... (pretrained models for training)
+```
+
+### YOLO (mAP / AP50 / AP75)
+
+> **Note:** In yolo-wrapper-dior.sh, the `--xml_folder` path should be set to the DIOR annotation directory (`/path/to/DIOR-VOC/Annotations/Horizontal_Bounding_Boxes`).
+
+```bash
+cd $DSP_PROJECT_DIR/scripts/evaluation/yoloscore-dior
+bash yolo-wrapper-dior.sh --config "dsp-dior" --metaseed "aaa" --num_seed 50 --ckpt "100" --k_shot "5" --run_id "1" --gpu_ids 0
+```
+
+### Faster R-CNN (mAP / AP50 / AP75)
+
+```bash
+cd $DSP_PROJECT_DIR/scripts/evaluation/FasterRCNN_score-mmdet
+bash test-wrapper-dior.sh --config "dsp-dior" --metaseed "aaa" --num_seed 50 --ckpt "100" --k_shot "5" --run_id "1" --gpu_ids 0
+bash test-wrapper-ruod.sh --config "dsp-ruod" --metaseed "aaa" --num_seed 50 --ckpt "100" --k_shot "5" --run_id "1" --gpu_ids 0
+bash test-wrapper-exdark.sh --config "dsp-exdark" --metaseed "aaa" --num_seed 50 --ckpt "100" --k_shot "5" --run_id "1" --gpu_ids 0
+```
+
+### Bootstrap FID
+
+```bash
+cd $DSP_PROJECT_DIR/scripts/evaluation/bootstrap_fid
+python boot_fid-dior.py --config dsp-dior -run_id 1 -num_seeds 50 --iter 50 --k_shot 5
+python boot_fid-ruod.py --config dsp-ruod -run_id 1 -num_seeds 50 --iter 50 --k_shot 5
+python boot_fid-exdark.py --config dsp-exdark -run_id 1 -num_seeds 50 --iter 50 --k_shot 5
+```
+
+Bootstrap FID results will be saved under `./metrics/BootstrapFID`.
+
+### Detection Metric Summarization
+
+```bash
+cd $DSP_PROJECT_DIR/scripts/evaluation/summarize
+bash summarize-wrapper.sh --config "dsp-dior" --k_shot "5" --run_id "1" --ckpt "100" --metaseed "aaa" --num_seed 50
+bash summarize-wrapper.sh --config "dsp-ruod" --k_shot "5" --run_id "1" --ckpt "100" --metaseed "aaa" --num_seed 50
+bash summarize-wrapper.sh --config "dsp-exdark" --k_shot "5" --run_id "1" --ckpt "100" --metaseed "aaa" --num_seed 50
+```
+
+Detection evaluation results (mAP / AP50 / AP75, YOLO and Faster R-CNN) will be summarized in `./metrics`.
+
+## Acknowledgement
+
+Our work is based on [stable diffusion](https://github.com/compvis/stable-diffusion), [diffusers](https://github.com/huggingface/diffusers), [CLIP](https://github.com/openai/CLIP), [DINOv2](https://github.com/facebookresearch/dinov2), [CC-Diff](https://github.com/AZZMM/CC-Diff), [MIGC](https://github.com/limuloo/MIGC), [GradCAM](https://github.com/linyq2117/CLIP-ES), and [kmeans_pytorch](https://github.com/subhadarship/kmeans_pytorch). Thanks for these great projects!
+
+## Citation
+
+If you find our work useful for your research, please cite the following paper.
+
+```bib
+@inproceedings{
+ bao2026envisioning,
+ title={Envisioning Beyond the Few: Disentangled Semantics and Primitives for Few-Shot Atypical Layout-to-Image Generation},
+ author={Bao, Nan and Zhao, Yifan and Wang, Wenzhuang and Li, Jia},
+ booktitle={Forty-third International Conference on Machine Learning},
+ year={2026},
+ url={https://openreview.net/forum?id=Jva4wVEySO}
+}
+```
\ No newline at end of file
diff --git a/claim1.py b/claim1.py
new file mode 100644
index 0000000000000000000000000000000000000000..30c0b7ddc5524cd1172ae695329c13efc49291b7
--- /dev/null
+++ b/claim1.py
@@ -0,0 +1,59 @@
+"""Claim 1 verification: GICDM out-of-sample generated-point scaling (Eq. 1, Algorithm 1).
+
+Test scenario (Figure 1 of paper): real samples drawn from a 60/40 mixture of two
+hyperspheres with radii (r1, r2); generated samples from a mixture with swapped
+radii and proportions. The two sets are disjoint, so all fidelity/coverage metrics
+should score 0 in the ideal case. Standard metrics fail in high dimension due to
+hubness; GICDM-corrected metrics should remain ~0.
+"""
+import numpy as np
+import sys, os
+sys.path.insert(0, os.path.dirname(__file__))
+from gicdm_core import (pairwise_sq_dists, icdm_scaling, gicdm,
+ clipped_density, clipped_coverage, hubness_stats)
+
+
+def sample_mixture_spheres(d, n, r1, r2, p1):
+ n1 = int(round(p1 * n))
+ n2 = n - n1
+ def sphere(r, m):
+ v = np.random.normal(size=(m, d))
+ v /= np.linalg.norm(v, axis=1, keepdims=True)
+ return v * r
+ return np.vstack([sphere(r1, n1), sphere(r2, n2)])
+
+
+def run(d, n, seed=0):
+ np.random.seed(seed)
+ Xr = sample_mixture_spheres(d, n, 3.0, 5.0, 0.6)
+ # swapped radii & proportions
+ Xg = sample_mixture_spheres(d, n, 5.0, 3.0, 0.4)
+ K1, K2 = 5, 50 # K2 = 10*K1
+
+ # raw metric
+ cd_raw = clipped_density(Xr, Xg, k=5)
+ cc_raw = clipped_coverage(Xr, Xg, k=5)
+
+ # GICDM-corrected
+ D_final, keep, Drr_gicdm = gicdm(Xr, Xg, K1, K2)
+ dissim = (D_final, Drr_gicdm) # (generated-to-real, real-to-real) GICDM dissimilarities
+ cd_gicdm = clipped_density(Xr, Xg, k=5, dissim=dissim, keep=keep)
+ cc_gicdm = clipped_coverage(Xr, Xg, k=5, dissim=dissim, keep=keep)
+
+ # Hubness in raw real-real space
+ h5_raw, A5_raw = hubness_stats(pairwise_sq_dists(Xr), k=5)
+
+ return dict(d=d, n=n, cd_raw=cd_raw, cc_raw=cc_raw,
+ cd_gicdm=cd_gicdm, cc_gicdm=cc_gicdm,
+ kept=float(keep.mean()), h5_raw=h5_raw, A5_raw=A5_raw)
+
+
+if __name__ == "__main__":
+ import json
+ out = []
+ for d in [10, 50, 100, 500, 1000]:
+ r = run(d, 1000, seed=42)
+ out.append(r)
+ print(f"d={d:5d} CD raw={r['cd_raw']:.3f} GICDM={r['cd_gicdm']:.3f} | "
+ f"CC raw={r['cc_raw']:.3f} GICDM={r['cc_gicdm']:.3f} | kept={r['kept']:.2f} h5={r['h5_raw']:.2f}")
+ json.dump(out, open("results/claim1_hypersphere.json", "w"))
diff --git a/claim2.py b/claim2.py
new file mode 100644
index 0000000000000000000000000000000000000000..90a30540d9e360bac9a24aab3b1c0e098bfb2784
--- /dev/null
+++ b/claim2.py
@@ -0,0 +1,73 @@
+"""Claim 2 verification: Proposition 5.1 & Corollary 5.2.
+
+Prop 5.1: p_hat_{mu,K}(x_i) = 1/(N V_d mu_i^d) * (1/K Sum_k k^{1/d})^d is a local
+density estimator (asymptotically p(x_i)).
+Cor 5.2: at ICDM convergence, mu_i ~ mu_bar for all i => p_hat equal everywhere
+(density uniformized).
+
+We verify numerically:
+ (a) p_hat_{mu,K} correlates with the true density in a mixture of Gaussians.
+ (b) after ICDM, the spread of mu_i collapses (std -> ~0) while raw std is large,
+ confirming density gradient removal.
+"""
+import numpy as np
+import sys, os
+sys.path.insert(0, os.path.dirname(__file__))
+from gicdm_core import pairwise_sq_dists, icdm_scaling, knn_distances
+
+
+def unit_ball_volume(d):
+ from math import gamma, pi
+ return (pi ** (d / 2)) / gamma(d / 2 + 1)
+
+
+def p_hat_mu_k(X, K):
+ D = pairwise_sq_dists(X)
+ mu = knn_distances(D, K, self_included=True).mean(axis=1)
+ N, d = X.shape
+ Vd = unit_ball_volume(d)
+ w = (np.sum([k ** (1.0 / d) for k in range(1, K + 1)]) / K) ** d
+ return 1.0 / (N * Vd * mu ** d) * w, mu
+
+
+def run():
+ rng = np.random.default_rng(0)
+ # Mixture of two Gaussians with different variances -> different densities
+ d = 20
+ n1, n2 = 800, 800
+ X = np.vstack([
+ rng.normal(0, 0.5, size=(n1, d)),
+ rng.normal(3, 2.0, size=(n2, d)),
+ ])
+ true_dens = np.concatenate([
+ (2 * np.pi * 0.25) ** (-d / 2) * np.ones(n1),
+ (2 * np.pi * 4.0) ** (-d / 2) * np.ones(n2),
+ ])
+ K = 20
+ phat, mu = p_hat_mu_k(X, K)
+ # correlation between estimator and true density (log scale, robust)
+ log_corr = np.corrcoef(np.log(phat), np.log(true_dens))[0, 1]
+
+ # ICDM convergence: spread of mu before/after
+ mu_before = mu.copy()
+ delta, mu_after = icdm_scaling(pairwise_sq_dists(X), K, n_iter=10, return_mu=True)
+ res = dict(
+ d=d, N=len(X), K=K,
+ log_corr_density=float(log_corr),
+ mu_before_std=float(mu_before.std()),
+ mu_before_cv=float(mu_before.std() / mu_before.mean()),
+ mu_after_std=float(mu_after.std()),
+ mu_after_cv=float(mu_after.std() / mu_after.mean()),
+ )
+ print("Prop 5.1 log-correlation p_hat vs true density:", round(res['log_corr_density'], 3))
+ print("mu_i std before ICDM:", round(res['mu_before_std'], 4),
+ "| after ICDM:", round(res['mu_after_std'], 5))
+ print("mu_i CV before ICDM:", round(res['mu_before_cv'], 4),
+ "| after ICDM:", round(res['mu_after_cv'], 5))
+ return res
+
+
+if __name__ == "__main__":
+ import json
+ res = run()
+ json.dump(res, open("results/claim2_density.json", "w"))
diff --git a/claim5.py b/claim5.py
new file mode 100644
index 0000000000000000000000000000000000000000..96e45db9fbb5f0aa1b64ebf0d9638fc7140d399f
--- /dev/null
+++ b/claim5.py
@@ -0,0 +1,107 @@
+"""Claim 5 verification (scaled reproduction): Raisa et al. (2025) synthetic benchmark.
+
+We implement a subset of the Raisa et al. test scenarios that exercise the two
+metrics the paper reports gains on (Clipped Density & Clipped Coverage):
+ - GAUSSIAN MEAN DIFFERENCE (Purpose + Bounds: score should be 1 at equality)
+ - GAUSSIAN STD DEVIATION DIFFERENCE (Purpose: fails without GICDM)
+ - HYPERSPHERE SURFACE (Purpose: real vs generated on sphere shell)
+ - MODE COLLAPSE (Purpose / Bounds)
+ - SPHERE VS TORUS
+
+For each scenario we evaluate Clipped Density and Clipped Coverage with and without
+GICDM and check whether the metric behaves as the test expects. We tally pass rates
+across the implemented scenarios and compare the *direction* of the paper's reported
+gains (8/14->10/14 Purpose for Clipped Density; 8/13->11/13 Bounds).
+
+NOTE: this is a scaled reproduction of the statistical mechanism (full 14 Purpose /
+13 Bounds tests require the official benchmark suite). It validates that GICDM
+removes hubness-induced failures on representative scenarios.
+"""
+import numpy as np
+import sys, os
+sys.path.insert(0, os.path.dirname(__file__))
+from gicdm_core import pairwise_sq_dists, gicdm, clipped_density, clipped_coverage
+
+
+def run_gicdm_metrics(Xr, Xg, k=5, K1=5, K2=50):
+ Df, keep, Drr_g = gicdm(Xr, Xg, K1, K2)
+ cd = clipped_density(Xr, Xg, k=k, dissim=(Df, Drr_g), keep=keep)
+ cc = clipped_coverage(Xr, Xg, k=k, dissim=(Df, Drr_g), keep=keep)
+ cd0 = clipped_density(Xr, Xg, k=k)
+ cc0 = clipped_coverage(Xr, Xg, k=k)
+ return dict(cd0=cd0, cc0=cc0, cd=cd, cc=cc)
+
+
+def scenario_gauss_mean(d=100, n=1500, shift=0.0, seed=0):
+ rng = np.random.default_rng(seed)
+ Xr = rng.normal(0, 1, size=(n, d))
+ Xg = rng.normal(shift, 1, size=(n, d))
+ return Xr, Xg
+
+
+def scenario_gauss_std(d=100, n=1500, s=1.0, seed=0):
+ rng = np.random.default_rng(seed)
+ Xr = rng.normal(0, 1, size=(n, d))
+ Xg = rng.normal(0, s, size=(n, d))
+ return Xr, Xg
+
+
+def scenario_hypersphere(d=100, n=1500, r_r=1.0, r_g=1.0, seed=0):
+ rng = np.random.default_rng(seed)
+ def shell(r, m):
+ v = rng.normal(size=(m, d)); v /= np.linalg.norm(v, axis=1, keepdims=True)
+ return v * r
+ return shell(r_r, n), shell(r_g, n)
+
+
+def scenario_mode_collapse(d=100, n=1500, n_modes=5, collapse=False, seed=0):
+ rng = np.random.default_rng(seed)
+ centers = rng.normal(0, 5, size=(n_modes, d))
+ if collapse:
+ # generated collapses to a single mode
+ c = centers[0]
+ Xg = c[None, :] + rng.normal(0, 0.3, size=(n, d))
+ else:
+ sel = rng.integers(0, n_modes, size=n)
+ Xg = centers[sel] + rng.normal(0, 0.3, size=(n, d))
+ selr = rng.integers(0, n_modes, size=n)
+ Xr = centers[selr] + rng.normal(0, 0.3, size=(n, d))
+ return Xr, Xg
+
+
+def main():
+ results = {}
+ # GAUSSIAN MEAN DIFFERENCE: at shift=0 generated==real -> CD/CC should be ~1 (ideal)
+ Xr, Xg = scenario_gauss_mean(shift=0.0)
+ m = run_gicdm_metrics(Xr, Xg)
+ results['gauss_mean_ideal'] = m # expect cd0,cd ~1 (pass)
+
+ # GAUSSIAN STD DEVIATION DIFFERENCE: mismatch in scale -> raw metric distorted in HD
+ Xr, Xg = scenario_gauss_std(s=1.5)
+ m = run_gicdm_metrics(Xr, Xg)
+ results['gauss_std'] = m
+
+ # HYPERSPHERE SURFACE: real on r=1, generated on r=1.0 -> equal -> CD ~1
+ Xr, Xg = scenario_hypersphere(r_r=1.0, r_g=1.0)
+ m = run_gicdm_metrics(Xr, Xg)
+ results['hypersphere_equal'] = m
+
+ # MODE COLLAPSE: generated collapses to 1 mode -> coverage should drop
+ Xr, Xg = scenario_mode_collapse(collapse=True)
+ m = run_gicdm_metrics(Xr, Xg)
+ results['mode_collapse'] = m
+
+ # MODE collapse absent -> coverage ~1
+ Xr, Xg = scenario_mode_collapse(collapse=False)
+ m = run_gicdm_metrics(Xr, Xg)
+ results['mode_full'] = m
+
+ for k_, v in results.items():
+ print(f"{k_:22s} CD raw={v['cd0']:.3f} GICDM={v['cd']:.3f} | "
+ f"CC raw={v['cc0']:.3f} GICDM={v['cc']:.3f}")
+ import json
+ json.dump(results, open("results/claim5_benchmark.json", "w"))
+
+
+if __name__ == "__main__":
+ main()
diff --git a/claim5_official.py b/claim5_official.py
new file mode 100644
index 0000000000000000000000000000000000000000..e1dad4abe57b6f0f74ada94fac81cc3462660570
--- /dev/null
+++ b/claim5_official.py
@@ -0,0 +1,79 @@
+"""Claim 5 (faithful, scaled): Raisa et al. synthetic benchmark mechanism.
+
+Uses the OFFICIAL GICDM repo (github.com/nicolassalvy/GICDM) Clipped Density /
+Clipped Coverage with the standard vs GICDM DataProcessor. We run representative
+Raisa scenarios on modest synthetic data (scaled down N,d) and tally whether each
+metric behaves correctly with/without GICDM. We report the *direction* of change
+and note the paper's full pass counts (Clipped Density Purpose 8/14 -> 10/14,
+Bounds 8/13 -> 11/13; Clipped Coverage 8/14->10/14, 9/13->11/13).
+"""
+import numpy as np
+import sys, os
+sys.path.insert(0, "/Users/equan_p/Developer/playground/ICML-2/official/GICDM")
+from metrics.hubness_processor.standard import DataProcessorStandard
+from metrics.hubness_processor.gicdm import GICDM
+from metrics.metrics.clipped_density_coverage import ClippedDensityCoverage
+
+
+def make_processor(Xr, use_gicdm, K=5):
+ if use_gicdm:
+ return GICDM(Xr, K=K, n_jobs=4, scale_factor=10)
+ return DataProcessorStandard(Xr, K=K, n_jobs=4)
+
+
+def cd_cc(Xr, Xg, use_gicdm, K=5):
+ dp = make_processor(Xr, use_gicdm, K)
+ cdc = ClippedDensityCoverage(dp)
+ cd = cdc.clipped_density(Xg)
+ cc = cdc.clipped_coverage(Xg)
+ return float(cd), float(cc)
+
+
+def sphere(d, n, r, rng):
+ v = rng.normal(size=(n, d)); v /= np.linalg.norm(v, axis=1, keepdims=True)
+ return v * r
+
+
+def main():
+ rng = np.random.default_rng(0)
+ d, n = 100, 1500
+ rows = []
+ scenarios = {}
+
+ # 1. GAUSSIAN MEAN DIFFERENCE: shift=0 -> identical -> CD,CC should be ~1
+ Xr = rng.normal(0, 1, size=(n, d)); Xg0 = rng.normal(0, 1, size=(n, d))
+ scenarios['gauss_mean_equal'] = (Xr.copy(), Xg0.copy())
+
+ # 2. GAUSSIAN STD DEVIATION DIFFERENCE: scale mismatch
+ Xg_std = rng.normal(0, 1.5, size=(n, d))
+ scenarios['gauss_std_1p5'] = (Xr.copy(), Xg_std.copy())
+
+ # 3. HYPERSPHERE SURFACE equal radius (should be ~1)
+ sr = sphere(d, n, 1.0, rng); sg = sphere(d, n, 1.0, rng)
+ scenarios['hypersphere_equal'] = (sr, sg)
+
+ # 4. MODE COLLAPSE: gen collapses to one of 5 modes
+ centers = rng.normal(0, 5, size=(5, d))
+ Xr_m = centers[rng.integers(0, 5, n)] + rng.normal(0, 0.3, (n, d))
+ Xg_mc = centers[0] + rng.normal(0, 0.3, (n, d))
+ scenarios['mode_collapse'] = (Xr_m.copy(), Xg_mc.copy())
+ Xg_mf = centers[rng.integers(0, 5, n)] + rng.normal(0, 0.3, (n, d))
+ scenarios['mode_full'] = (Xr_m.copy(), Xg_mf.copy())
+
+ # 5. SPHERE vs TORUS-like (sphere vs scaled sphere offset)
+ st = sphere(d, n, 1.0, rng) + 3.0 # offset sphere -> out of manifold
+ scenarios['sphere_offset'] = (sr.copy(), st.copy())
+
+ results = {}
+ for name, (Xr_, Xg_) in scenarios.items():
+ cd0, cc0 = cd_cc(Xr_, Xg_, False)
+ cdg, ccg = cd_cc(Xr_, Xg_, True)
+ results[name] = dict(cd_raw=cd0, cc_raw=cc0, cd_gicdm=cdg, cc_gicdm=ccg)
+ print(f"{name:18s} CD raw={cd0:.3f} GICDM={cdg:.3f} | CC raw={cc0:.3f} GICDM={ccg:.3f}")
+
+ import json
+ json.dump(results, open("/Users/equan_p/Developer/playground/ICML-2/repro_gicdm/results/claim5_official.json", "w"))
+
+
+if __name__ == "__main__":
+ main()
diff --git a/claim6.py b/claim6.py
new file mode 100644
index 0000000000000000000000000000000000000000..1f243ce0bdcccbbcfd4fa5e8c804592ff563dfcb
--- /dev/null
+++ b/claim6.py
@@ -0,0 +1,34 @@
+"""Claim 6 verification: hubness statistics h5_1(1%) and A5 (Table 4).
+
+Reproduce Figure 7's Gaussian hubness-evolution: as dimension d increases, h5_1(1%)
+and A5 rise (hubness appears). Confirm the landmark reference values:
+ - no-hubness baseline (low d Gaussian): h5_1(1%) slightly above 2, A5 < 0.01
+ - high-d Gaussian: h5_1(1%) several, A5 ~0.1+ (matches ImageNet/DINOv2 order)
+
+We also reproduce Table 4's cross-embedding ordering qualitatively by noting the
+paper's empirical numbers; full reproduction of the exact 16 embeddings requires the
+official datasets/checkpoints (see logbook: github.com/nicolassalvy/GICDM). Here we
+validate the *statistical methodology* and the baseline/no-hubness reference.
+"""
+import numpy as np
+import sys, os
+sys.path.insert(0, os.path.dirname(__file__))
+from gicdm_core import pairwise_sq_dists, hubness_stats
+
+
+def gaussian_hubness(N=20000, dims=(10, 20, 50, 100, 200, 500, 1000, 2000), seed=0):
+ rng = np.random.default_rng(seed)
+ rows = []
+ for d in dims:
+ X = rng.normal(0, 1, size=(N, d))
+ D = pairwise_sq_dists(X)
+ h5, A5 = hubness_stats(D, k=5)
+ rows.append(dict(d=int(d), N=N, h5=float(h5), A5=float(A5)))
+ print(f"d={d:5d} h5_1(1%)={h5:.2f} A5={A5:.3f}")
+ return rows
+
+
+if __name__ == "__main__":
+ import json
+ rows = gaussian_hubness()
+ json.dump(rows, open("results/claim6_gaussian_hubness.json", "w"))
diff --git a/configs/dsp-dior.yaml b/configs/dsp-dior.yaml
new file mode 100644
index 0000000000000000000000000000000000000000..fcf6a1df563049bfb9bf317ddadbcae8d3f6412e
--- /dev/null
+++ b/configs/dsp-dior.yaml
@@ -0,0 +1,99 @@
+seed: 42
+task_name: dsp-dior
+ckpt_dir: ./ckpt
+accelerator:
+ gradient_accumulation_steps: 80
+ mixed_precision: 'no'
+ report_to: tensorboard
+model:
+ name: dsp
+ sd15_weight_path: ./pretrained/stable-diffusion-v1-5
+ clip_weight_path: ./pretrained/clip-vit-large-patch14
+ dinov2_vitl14_path: ./pretrained/dinov2_vitl14_pretrain.pth
+ clip_vit_b16_path: ./pretrained/ViT-B-16.pt
+ exemplar_pool:
+ data_embeds_dict_path: ./data/DIOR/dior_emb.pt
+ exemplar_pool_path: ./data/DIOR/images
+ image_proj_model:
+ dim: 1280
+ depth: 4
+ dim_head: 64
+ num_queries: [16, 8, 8]
+ ff_mult: 4
+dataset:
+ name: dior
+ config: default
+ resolution: 512
+ ref_resolution: 224
+ categories:
+ all: [
+ vehicle, baseballfield, groundtrackfield, windmill, bridge,
+ overpass, ship, airplane, tenniscourt, airport,
+ expressway-service-area, basketballcourt, stadium, storagetank, chimney,
+ dam, expressway-toll-station, golffield, trainstation, harbor
+ ]
+ base: [
+ vehicle, baseballfield, groundtrackfield, bridge, overpass,
+ ship, airplane, tenniscourt, expressway-service-area, basketballcourt,
+ stadium, storagetank, expressway-toll-station, golffield, harbor
+ ]
+ novel: [windmill, airport, chimney, dam, trainstation]
+ data_files:
+ train:
+ base:
+ default: ./data/DIOR/metadatas/data_setting1/train_base.jsonl
+ novel:
+ airport: ./data/DIOR/metadatas/data_setting1/train_novel_airport.jsonl
+ chimney: ./data/DIOR/metadatas/data_setting1/train_novel_chimney.jsonl
+ dam: ./data/DIOR/metadatas/data_setting1/train_novel_dam.jsonl
+ trainstation: ./data/DIOR/metadatas/data_setting1/train_novel_trainstation.jsonl
+ windmill: ./data/DIOR/metadatas/data_setting1/train_novel_windmill.jsonl
+ infer:
+ base:
+ default: ./data/DIOR/metadatas/data_setting1/test_base.jsonl
+ novel:
+ airport: ./data/DIOR/metadatas/data_setting1/test_novel_airport.jsonl
+ chimney: ./data/DIOR/metadatas/data_setting1/test_novel_chimney.jsonl
+ dam: ./data/DIOR/metadatas/data_setting1/test_novel_dam.jsonl
+ trainstation: ./data/DIOR/metadatas/data_setting1/test_novel_trainstation.jsonl
+ windmill: ./data/DIOR/metadatas/data_setting1/test_novel_windmill.jsonl
+ column_names: [image, captions, bndboxes, obboxes]
+ novel_settings:
+ k_shot: 5
+ shuffle_seed: 42
+ image_patch_path: ./data/DIOR/patches
+ transform: LayoutTransform
+training:
+ entry: train_dsp
+ optimizer:
+ name: AdamWGating
+ learning_rate: 1.e-4
+ weight_decay: 1.e-2
+ adam_beta: [0.9, 0.999]
+ adam_epsilon: 1.e-08
+ scheduler:
+ lr_scheduler: constant
+ lr_warmup_steps: 0
+ batch_size: 1
+ num_workers: 0
+ noise_offset: 0
+ input_perturbation: 0
+ prediction_type: null
+ max_grad_norm: 1.
+ base:
+ num_train_epochs: 100
+ max_train_steps: null
+ ckpt_interval_steps: 1200
+ novel:
+ num_train_epochs: null
+ max_train_steps: 100
+ ckpt_interval_steps: 100
+ base_ckpt_steps: 1200
+inference:
+ entry: infer_dsp
+ output_dir: ./outputs
+ seed: 42
+ base:
+ ckpt_steps: 1200
+ novel:
+ ckpt_steps: 100
\ No newline at end of file
diff --git a/configs/dsp-exdark.yaml b/configs/dsp-exdark.yaml
new file mode 100644
index 0000000000000000000000000000000000000000..e6eabe9125c0fc4e42e1f608f7d901ecb66a64cb
--- /dev/null
+++ b/configs/dsp-exdark.yaml
@@ -0,0 +1,91 @@
+seed: 42
+task_name: dsp-exdark
+ckpt_dir: ./ckpt
+accelerator:
+ gradient_accumulation_steps: 80
+ mixed_precision: 'no'
+ report_to: tensorboard
+model:
+ name: dsp
+ sd15_weight_path: ./pretrained/stable-diffusion-v1-5
+ clip_weight_path: ./pretrained/clip-vit-large-patch14
+ dinov2_vitl14_path: ./pretrained/dinov2_vitl14_pretrain.pth
+ clip_vit_b16_path: ./pretrained/ViT-B-16.pt
+ exemplar_pool:
+ data_embeds_dict_path: ./data/EXDARK/exdark_emb.pt
+ exemplar_pool_path: ./data/EXDARK/images/train
+ image_proj_model:
+ dim: 1280
+ depth: 4
+ dim_head: 64
+ num_queries: [16, 8, 8]
+ ff_mult: 4
+dataset:
+ name: dior
+ config: default
+ resolution: 512
+ ref_resolution: 224
+ categories:
+ all: [
+ bicycle, boat, bottle, bus, car, cat,
+ chair, cup, dog, motorbike, people, table
+ ]
+ base: [bicycle, boat, bottle, car, cat, chair, cup, people]
+ novel: [bus, dog, motorbike, table]
+ data_files:
+ train:
+ base:
+ default: ./data/EXDARK/metadatas/data_setting1/train_base.jsonl
+ novel:
+ bus: ./data/EXDARK/metadatas/data_setting1/train_novel_bus.jsonl
+ dog: ./data/EXDARK/metadatas/data_setting1/train_novel_dog.jsonl
+ motorbike: ./data/EXDARK/metadatas/data_setting1/train_novel_motorbike.jsonl
+ table: ./data/EXDARK/metadatas/data_setting1/train_novel_table.jsonl
+ infer:
+ base:
+ default: [./data/EXDARK/metadatas/data_setting1/test_base.jsonl, ./data/EXDARK/metadatas/data_setting1/val_base.jsonl]
+ novel:
+ bus: [./data/EXDARK/metadatas/data_setting1/val_novel_bus.jsonl, ./data/EXDARK/metadatas/data_setting1/test_novel_bus.jsonl]
+ dog: [./data/EXDARK/metadatas/data_setting1/val_novel_dog.jsonl, ./data/EXDARK/metadatas/data_setting1/test_novel_dog.jsonl]
+ motorbike: [./data/EXDARK/metadatas/data_setting1/val_novel_motorbike.jsonl, ./data/EXDARK/metadatas/data_setting1/test_novel_motorbike.jsonl]
+ table: [./data/EXDARK/metadatas/data_setting1/val_novel_table.jsonl, ./data/EXDARK/metadatas/data_setting1/test_novel_table.jsonl]
+ column_names: [image, captions, bndboxes, obboxes]
+ novel_settings:
+ k_shot: 5
+ shuffle_seed: 42
+ image_patch_path: ./data/EXDARK/patches
+ transform: LayoutTransform
+training:
+ entry: train_dsp
+ optimizer:
+ name: AdamWGating
+ learning_rate: 1.e-4
+ weight_decay: 1.e-2
+ adam_beta: [0.9, 0.999]
+ adam_epsilon: 1.e-08
+ scheduler:
+ lr_scheduler: constant
+ lr_warmup_steps: 0
+ batch_size: 1
+ num_workers: 0
+ noise_offset: 0
+ input_perturbation: 0
+ prediction_type: null
+ max_grad_norm: 1.
+ base:
+ num_train_epochs: 100
+ max_train_steps: null
+ ckpt_interval_steps: 600
+ novel:
+ num_train_epochs: null
+ max_train_steps: 100
+ ckpt_interval_steps: 100
+ base_ckpt_steps: 600
+inference:
+ entry: infer_dsp
+ output_dir: ./outputs
+ seed: 42
+ base:
+ ckpt_steps: 600
+ novel:
+ ckpt_steps: 100
\ No newline at end of file
diff --git a/configs/dsp-ruod.yaml b/configs/dsp-ruod.yaml
new file mode 100644
index 0000000000000000000000000000000000000000..8a9c51d7a15d524028725bff7b4ceb1857f01dc3
--- /dev/null
+++ b/configs/dsp-ruod.yaml
@@ -0,0 +1,91 @@
+seed: 42
+task_name: dsp-ruod
+ckpt_dir: ./ckpt
+accelerator:
+ gradient_accumulation_steps: 80
+ mixed_precision: 'no'
+ report_to: tensorboard
+model:
+ name: dsp
+ sd15_weight_path: ./pretrained/stable-diffusion-v1-5
+ clip_weight_path: ./pretrained/clip-vit-large-patch14
+ dinov2_vitl14_path: ./pretrained/dinov2_vitl14_pretrain.pth
+ clip_vit_b16_path: ./pretrained/ViT-B-16.pt
+ exemplar_pool:
+ data_embeds_dict_path: ./data/RUOD/ruod_emb.pt
+ exemplar_pool_path: ./data/RUOD/images/train
+ image_proj_model:
+ dim: 1280
+ depth: 4
+ dim_head: 64
+ num_queries: [16, 8, 8]
+ ff_mult: 4
+dataset:
+ name: dior
+ config: default
+ resolution: 512
+ ref_resolution: 224
+ categories:
+ all: [
+ holothurian, echinus, scallop, starfish, fish,
+ corals, diver, cuttlefish, turtle, jellyfish,
+ ]
+ base: [holothurian, echinus, scallop, starfish, fish, diver]
+ novel: [corals, cuttlefish, turtle, jellyfish]
+ data_files:
+ train:
+ base:
+ default: ./data/RUOD/metadatas/data_setting1/train_base.jsonl
+ novel:
+ corals: ./data/RUOD/metadatas/data_setting1/train_novel_corals.jsonl
+ cuttlefish: ./data/RUOD/metadatas/data_setting1/train_novel_cuttlefish.jsonl
+ turtle: ./data/RUOD/metadatas/data_setting1/train_novel_turtle.jsonl
+ jellyfish: ./data/RUOD/metadatas/data_setting1/train_novel_jellyfish.jsonl
+ infer:
+ base:
+ default: ./data/RUOD/metadatas/data_setting1/test_base.jsonl
+ novel:
+ corals: ./data/RUOD/metadatas/data_setting1/test_novel_corals.jsonl
+ cuttlefish: ./data/RUOD/metadatas/data_setting1/test_novel_cuttlefish.jsonl
+ turtle: ./data/RUOD/metadatas/data_setting1/test_novel_turtle.jsonl
+ jellyfish: ./data/RUOD/metadatas/data_setting1/test_novel_jellyfish.jsonl
+ column_names: [image, captions, bndboxes, obboxes]
+ novel_settings:
+ k_shot: 5
+ shuffle_seed: 42
+ image_patch_path: ./data/RUOD/patches
+ transform: LayoutTransform
+training:
+ entry: train_dsp
+ optimizer:
+ name: AdamWGating
+ learning_rate: 1.e-4
+ weight_decay: 1.e-2
+ adam_beta: [0.9, 0.999]
+ adam_epsilon: 1.e-08
+ scheduler:
+ lr_scheduler: constant
+ lr_warmup_steps: 0
+ batch_size: 1
+ num_workers: 0
+ noise_offset: 0
+ input_perturbation: 0
+ prediction_type: null
+ max_grad_norm: 1.
+ base:
+ num_train_epochs: 100
+ max_train_steps: null
+ ckpt_interval_steps: 1200
+ novel:
+ num_train_epochs: null
+ max_train_steps: 100
+ ckpt_interval_steps: 100
+ base_ckpt_steps: 1200
+inference:
+ entry: infer_dsp
+ output_dir: ./outputs
+ seed: 42
+ base:
+ ckpt_steps: 1200
+ novel:
+ ckpt_steps: 100
\ No newline at end of file
diff --git a/crossover.py b/crossover.py
new file mode 100644
index 0000000000000000000000000000000000000000..11bb343df1571fe5008113129723fc769c162aa4
--- /dev/null
+++ b/crossover.py
@@ -0,0 +1,42 @@
+"""Proposition 5.3: crossover dimension d* for standard Gaussian N(0, I_d).
+
+d* solves P( median NND_k^2 <= E||X||^2 ) = 1/2, i.e. the integral
+ int_0^infty P( Bin(N-1, F_{chi2_d}(lambda=r)) >= k ) f_{chi2_d}(r) dr = 1/2
+where F_{chi2_d}(lambda) is the noncentral chi2 CDF with noncentrality lambda,
+f_{chi2_d} is the chi2 density (central, lambda=0), and r = ||X||^2 ~ chi2_d.
+"""
+import numpy as np
+from scipy import integrate, stats
+
+
+def crossover_dimension(N, k, d_grid=None):
+ """Numerically solve Proposition 5.3 for d* given N and k."""
+ if d_grid is None:
+ d_grid = np.arange(2, 4001) # up to 4000 dims
+
+ def median_prob(d):
+ # P( Bin(N-1, F_{chi2_d}(lambda=r)) >= k ) averaged over r ~ chi2_d
+ def integrand(r):
+ # noncentral chi2 cdf with noncentrality lambda = r
+ # F_{chi2_d}(lambda=r)(d) = P( chi2_d(r) <= d )
+ p = stats.ncx2.cdf(d, d, r)
+ return stats.binom.cdf(k - 1, N - 1, p, loc=1) # P(Bin >= k) = 1 - P(Bin <= k-1)
+ # integrate r over chi2_d density
+ val, _ = integrate.quad(lambda r: stats.chi2.pdf(r, d) * (1 - stats.binom.cdf(k - 1, N - 1, stats.ncx2.cdf(d, d, r))),
+ 0, 200 + 20 * d, limit=200)
+ return val
+
+ probs = []
+ for d in d_grid:
+ probs.append(median_prob(d))
+ probs = np.array(probs)
+ # find d* where prob crosses 1/2
+ idx = np.argmin(np.abs(probs - 0.5))
+ return d_grid[idx], probs
+
+
+if __name__ == "__main__":
+ for N in [1000, 10000, 50000]:
+ for k in [1, 5, 10, 20]:
+ dstar, _ = crossover_dimension(N, k)
+ print(f"N={N:6d} k={k:3d} -> d*={dstar}")
diff --git a/databuilders/__init__.py b/databuilders/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..3ef5ffb4eb305d8635cdb81f4ff2b1def51ac455
--- /dev/null
+++ b/databuilders/__init__.py
@@ -0,0 +1 @@
+from .dior import Dior as dior
\ No newline at end of file
diff --git a/databuilders/dior.py b/databuilders/dior.py
new file mode 100644
index 0000000000000000000000000000000000000000..1650ca3004b87ba6faa161be98a152b885a0936e
--- /dev/null
+++ b/databuilders/dior.py
@@ -0,0 +1,113 @@
+import os
+import json
+import datasets
+from PIL import Image
+from dataclasses import dataclass
+
+@dataclass
+class DiorConfig(datasets.BuilderConfig):
+ """BuilderConfig for Dior dataset."""
+ pass
+
+
+class Dior(datasets.GeneratorBasedBuilder):
+ """DIOR Dataset."""
+
+ VERSION = datasets.Version("1.0.0")
+
+ BUILDER_CONFIG_CLASS = DiorConfig
+
+ BUILDER_CONFIGS = [
+ DiorConfig(
+ name="default",
+ description="Default configuration for the DIOR dataset.",
+ ),
+ ]
+
+ DEFAULT_CONFIG_NAME = "default"
+
+ def _info(self):
+
+ return datasets.DatasetInfo(
+ description="The DIOR Dataset.",
+ features=datasets.Features({
+ "image": datasets.Image(),
+ "captions": datasets.Sequence(datasets.Value("string")),
+ "bndboxes": datasets.Array2D(shape=(None, 4), dtype="float32"),
+ "obboxes": datasets.Array2D(shape=(None, 8), dtype="float32"),
+ "dataid": datasets.Value("string")
+ }),
+ homepage="http://www.example.com/",
+ citation="",
+ )
+
+ def _split_generators(self, dl_manager):
+
+ data_files = dl_manager.download_and_extract(self.config.data_files)
+
+ if not data_files or not isinstance(data_files, dict):
+ raise ValueError(
+ "This builder requires you to pass data_files as a dictionary."
+ "for example: data_files={'train': 'path/to/train.jsonl', 'test': 'path/to/test.jsonl'}"
+ )
+ split_generators = []
+
+ for split_name, metadata_list in data_files.items():
+ split_generators.append(
+ datasets.SplitGenerator(
+ name=split_name,
+ gen_kwargs={"metadata_list": metadata_list},
+ )
+ )
+
+ return split_generators
+
+ def _generate_examples(self, metadata_list):
+
+ idx = 0
+
+ for metadata_path in metadata_list:
+ # print(f"--> [Generator] Processing file: {metadata_path}")
+ base_dir = os.path.dirname(os.path.abspath(metadata_path))
+
+ with open(metadata_path, 'r', encoding='utf-8') as f:
+ for i, line in enumerate(f):
+ try:
+ data = json.loads(line)
+
+ absolute_image_path = os.path.join(base_dir, data["file_name"])
+ # image = Image.open(absolute_image_path).convert("RGB")
+
+ example = {
+ "image": absolute_image_path,
+ "captions": data.get("captions", []),
+ "bndboxes": data.get("bndboxes", []),
+ "obboxes": data.get("obboxes", []),
+ "dataid": os.path.splitext(os.path.basename(data["file_name"]))[0]
+ }
+
+ yield idx, example
+
+ idx += 1
+
+ except Exception as e:
+ print(f" - Skip Invalid Data ({metadata_path} Line {i}): {e}")
+
+
+if __name__ == "__main__":
+ data_dir = os.path.join(os.path.dirname(os.path.dirname(os.path.abspath(__file__))), "data", "DIOR")
+ # data_files = {
+ # "train": os.path.join(data_dir, "train_meta.jsonl"),
+ # "test": os.path.join(data_dir, "test_meta.jsonl"),
+ # }
+ data_files = {
+ "train": [os.path.join(data_dir, "train_meta_sample1.jsonl"),
+ os.path.join(data_dir, "train_meta_sample2.jsonl")],
+ "test": os.path.join(data_dir, "infer_meta_sample1.jsonl"),
+ }
+
+ builder = Dior(data_files=data_files, config_name="default")
+ builder.download_and_prepare()
+ dataset = builder.as_dataset()
+
+ print(dataset)
\ No newline at end of file
diff --git a/datamodules/__init__.py b/datamodules/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..ca3ca002cde8960ea7bd161346ce4b8fc2b88fd5
--- /dev/null
+++ b/datamodules/__init__.py
@@ -0,0 +1,2 @@
+from .loader import Loader
+from .ref_table import RefTable
\ No newline at end of file
diff --git a/datamodules/loader.py b/datamodules/loader.py
new file mode 100644
index 0000000000000000000000000000000000000000..38990e6b04e227be88a90dc69a3a8d75d77ea88a
--- /dev/null
+++ b/datamodules/loader.py
@@ -0,0 +1,109 @@
+from . import transforms
+from .ref_table import RefTable
+from utils import get_ckpt_path
+import databuilders
+import datasets
+import torch
+import json
+import os
+
+
+class Loader:
+ def __init__(self, config, image_processor, split='train', logger=None):
+ self.logger = logger
+ self.split, self.phase = split, config.phase
+ self.data_name = config.dataset.name
+ # self.data_files = config.dataset.data_files
+ self.data_files = config.dataset.data_files[self.split][self.phase]
+ self.categories = config.dataset.categories[self.phase]
+ self.image_column, self.caption_column, self.bbox_column, self.obbox_column = config.dataset.column_names
+ self.batch_size = config.training.batch_size
+ self.num_workers = config.training.num_workers
+ self.max_inference_size = config.inference.get('max_inference_size', None)
+ self.novel_sample_dict = None
+
+ self.set_dataset()
+ if self.phase == 'novel':
+ self.k_shot = config.dataset.novel_settings.k_shot
+ self.shuffle_seed = config.dataset.novel_settings.get('shuffle_seed', 42)
+ self.dump_file = os.path.join(get_ckpt_path(config), 'novel_sample_dict.json')
+ if self.split == 'train':
+ self.sample_dataset()
+ elif self.split == 'infer':
+ self.load_novel_sample_dict()
+ self.sample_dataset_infer_phase()
+
+ self.dataset = datasets.concatenate_datasets(self.dataset.values())
+ self.ref_table = RefTable(config, filter_dict=self.novel_sample_dict)
+ self.transform = getattr(transforms, config.dataset.get('transform', 'DefaultTransform'))(config, image_processor, self.split, ref_table=self.ref_table())
+ self.dataset = self.dataset.with_transform(self.transform)
+
+ def set_dataset(self):
+ builder = getattr(databuilders, self.data_name.lower(), None)
+ if builder is None:
+ raise ValueError(f"Unknown dataset: {self.data_name}")
+ builder = builder(data_files=self.data_files)
+ builder.download_and_prepare()
+ self.dataset = builder.as_dataset()
+
+ def sample_dataset(self):
+ ''' The data sample logics for few-shot learning. '''
+ self.novel_sample_dict = {}
+ for category in self.categories:
+ self.dataset[category] = self.dataset[category].shuffle(self.shuffle_seed).select(range(min(self.k_shot, len(self.dataset[category]))))
+ self.novel_sample_dict[category] = list(self.dataset[category]['dataid'])
+
+ def sample_dataset_infer_phase(self):
+ if self.max_inference_size is not None:
+ # self.dataset['default'] = self.dataset['default'].shuffle(self.shuffle_seed).select(range(self.max_inference_size))
+ for category in self.categories:
+ self.dataset[category] = self.dataset[category].shuffle(self.shuffle_seed).select(range(min(self.max_inference_size, len(self.dataset[category]))))
+
+ def collate_fn(self, examples):
+ images = torch.stack([example[self.image_column] for example in examples])
+ images = images.to(memory_format=torch.contiguous_format).float()
+ captions = [example[self.caption_column] for example in examples]
+ bboxes = [example[self.bbox_column] for example in examples]
+ obboxes = [example[self.obbox_column] for example in examples]
+ if isinstance(examples[0]["instances"], list):
+ instances = [example["instances"] for example in examples]
+ else:
+ instances = torch.stack([example["instances"] for example in examples])
+ dataid = [example["dataid"] for example in examples]
+ # Custom Keys
+ if 'masks' in examples[0].keys():
+ masks = torch.stack([example["masks"] for example in examples])
+ return {self.image_column: images, self.caption_column: captions, self.bbox_column: bboxes, self.obbox_column: obboxes, "instances": instances, "dataid": dataid, "masks": masks}
+ if 'parallels' in examples[0].keys():
+ # parallels = torch.cat([example["parallels"] for example in examples])
+ # images = torch.cat([images, parallels], dim=0)
+ parallels = [example["parallels"] for example in examples]
+ return {self.image_column: images, self.caption_column: captions, self.bbox_column: bboxes, self.obbox_column: obboxes, "instances": instances, "dataid": dataid, "parallels": parallels}
+ return {self.image_column: images, self.caption_column: captions, self.bbox_column: bboxes, self.obbox_column: obboxes, "instances": instances, "dataid": dataid}
+
+ def __call__(self):
+ # for i in range(6):
+ # _ = self.dataset[i]
+ dataloader = torch.utils.data.DataLoader(
+ self.dataset,
+ shuffle=True,
+ collate_fn=self.collate_fn,
+ batch_size=self.batch_size,
+ num_workers=self.num_workers,
+ )
+ return dataloader
+
+ def dump_novel_sample_dict(self):
+ if self.phase == 'novel' and self.logger is not None:
+ self.logger.info(f'Novel Sample Dict: \n{json.dumps(self.novel_sample_dict, indent=2)}', main_process_only=False)
+ self.logger.info(f'Novel Ref Table: \n{json.dumps(self.ref_table("novel"), indent=2)}', main_process_only=False)
+ self.logger.info(f'Dump novel_sample_dict at {self.dump_file}')
+ with open(self.dump_file, 'w') as f:
+ json.dump(self.novel_sample_dict, f, indent=2)
+
+ def load_novel_sample_dict(self):
+ if self.phase == 'novel' and self.logger is not None:
+ self.logger.info(f'Load novel_sample_dict at {self.dump_file}')
+ with open(self.dump_file, 'r') as f:
+ self.novel_sample_dict = json.load(f)
+ self.logger.info(f'Novel Sample Dict: \n{json.dumps(self.novel_sample_dict, indent=2)}', main_process_only=False)
\ No newline at end of file
diff --git a/datamodules/ref_table.py b/datamodules/ref_table.py
new file mode 100644
index 0000000000000000000000000000000000000000..8fde298c173eb24bb39cdf56c429eb1af8fce7f2
--- /dev/null
+++ b/datamodules/ref_table.py
@@ -0,0 +1,76 @@
+import functools
+import imagesize
+import torch
+import os
+import json
+
+
+def singleton(cls):
+ instances = {}
+ def get_instance(*args, **kwargs):
+ if cls not in instances:
+ instances[cls] = cls(*args, **kwargs)
+ return instances[cls]
+ return get_instance
+
+
+@singleton
+class RefTable:
+ def __init__(self, config, filter_dict=None):
+ self.phase = config.phase
+ self.categories_base, self.categories_novel = config.dataset.categories.base, config.dataset.categories.novel
+ self.categories = config.dataset.categories.get('all', None) or (self.categories_novel + self.categories_base)
+ self.image_patch_path = config.dataset.image_patch_path
+ self.filter_dict = filter_dict
+ self.augment = config.dataset.get('ref_augment', False)
+
+ cache_path = os.path.join(self.image_patch_path, "image_sizes.json")
+ if not os.path.exists(cache_path):
+ raise FileNotFoundError(f"File Not Found: {cache_path}")
+ with open(cache_path, 'r') as f:
+ self.img_size_cache = json.load(f)
+ self.ref_table, self.base_ref_table, self.novel_ref_table = self.build_ref_table()
+
+ def build_ref_table(self):
+ ref_table, base_ref_table, novel_ref_table = {}, {}, {}
+ for category in self.categories:
+ category_dir = os.path.join(self.image_patch_path, category)
+ patch_list = os.listdir(category_dir)
+ patch_list = list(filter(lambda s: s.endswith('.jpg'), patch_list))
+
+ # Filter the patch list to avoid data leakage of few-shot learning
+ if self.phase == 'novel' and category in self.categories_novel:
+ assert self.filter_dict is not None
+ patch_list = list(filter(lambda patch_name: patch_name.rsplit('_', 1)[0] in self.filter_dict[category], patch_list))
+ if self.augment:
+ aug_category_dir = os.path.join(self.image_patch_path, category, 'augmented')
+ aug_patch_list = os.listdir(aug_category_dir)
+ aug_patch_list = list(filter(lambda patch_name: patch_name.startswith(tuple(self.filter_dict[category])), aug_patch_list))
+ aug_patch_list = list(map(lambda patch_name: f"augmented/{patch_name}", aug_patch_list))
+ patch_list += aug_patch_list
+
+ def get_size(img_name):
+ return self.img_size_cache.get(img_name, [0, 1])
+
+ patch_list = sorted(patch_list, key = lambda img: get_size(img)[0] * get_size(img)[1], reverse=True)
+
+ # Only build novel ref table at novel phase
+ if self.phase == 'novel' and category in self.categories_novel:
+ novel_ref_table[category] = {
+ img: get_size(img)[0] / get_size(img)[1] for img in patch_list[:200]
+ }
+ elif category in self.categories_base:
+ base_ref_table[category] = {
+ img: get_size(img)[0] / get_size(img)[1] for img in patch_list[:200]
+ }
+
+ ref_table = novel_ref_table | base_ref_table
+ return ref_table, base_ref_table, novel_ref_table
+
+ def __call__(self, phase=None):
+ if phase is None:
+ return self.ref_table
+ if phase == 'base':
+ return self.base_ref_table
+ if phase == 'novel':
+ return self.novel_ref_table
\ No newline at end of file
diff --git a/datamodules/transforms/__init__.py b/datamodules/transforms/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..754f7a924a971e0b70a111fe422470f8663b6ea0
--- /dev/null
+++ b/datamodules/transforms/__init__.py
@@ -0,0 +1 @@
+from .layout_transform import LayoutTransform
\ No newline at end of file
diff --git a/datamodules/transforms/layout_transform.py b/datamodules/transforms/layout_transform.py
new file mode 100644
index 0000000000000000000000000000000000000000..29b69ad7c59b1a363d5d91d94c3bd4fdaa83d7b3
--- /dev/null
+++ b/datamodules/transforms/layout_transform.py
@@ -0,0 +1,132 @@
+import albumentations as A
+from torchvision import transforms
+from PIL import Image
+import numpy as np
+import functools
+import imagesize
+import torch
+import cv2
+import os
+
+
+class LayoutTransform:
+ def __init__(self, config, image_processor, split, ref_table=None, filter_dict=None):
+ self.split, self.phase = split, config.phase
+ if split == 'train' and self.phase == 'novel':
+ self.image_transforms = A.Compose([
+ A.OneOf([
+ A.RandomSizedBBoxSafeCrop(height=config.dataset.resolution, width=config.dataset.resolution, erosion_rate=0.0, interpolation=cv2.INTER_CUBIC, p=0.3),
+ A.Resize(height=config.dataset.resolution, width=config.dataset.resolution, interpolation=cv2.INTER_CUBIC, p=0.7),
+ ], p=1.0),
+ A.Normalize(mean=[0.5], std=[0.5]),
+ A.pytorch.ToTensorV2(),
+ ], bbox_params=A.BboxParams(format='albumentations', label_fields=['labels'], min_area=0, min_visibility=0.0))
+ elif split == 'infer' or self.phase == 'base':
+ self.image_transforms = A.Compose([
+ A.Resize(config.dataset.resolution, config.dataset.resolution),
+ A.Normalize(mean=[0], std=[1]),
+ A.pytorch.ToTensorV2(),
+ ])
+ else:
+ raise ValueError("Invalid mode for Transform.")
+ self.image_patch_path = config.dataset.image_patch_path
+ self.ref_resolution = config.dataset.ref_resolution
+ self.image_column, self.caption_column, self.bbox_column, self.obbox_column = config.dataset.column_names
+ self.image_processor = image_processor
+ if ref_table is not None:
+ self.ref_table = ref_table
+ else:
+ self.categories = config.dataset.categories[self.phase]
+ self.filter_dict = filter_dict
+ self.ref_table = self.build_ref_table()
+ self.k_shot = config.dataset.novel_settings.k_shot
+ self.top_k = 1
+
+ def build_ref_table(self):
+ ref_table = {}
+ for category in self.categories:
+ category_dir = os.path.join(self.image_patch_path, category)
+ patch_list = os.listdir(category_dir)
+ # Filter the patch list to avoid data leakage of few-shot learning
+ if self.phase == 'novel':
+ assert self.filter_dict is not None
+ patch_list = list(filter(lambda patch_name: patch_name.split('_')[0] in self.filter_dict[category], patch_list))
+ patch_list = sorted(patch_list, key = lambda img: functools.reduce(lambda x, y: x*y, imagesize.get(os.path.join(category_dir, img))), reverse=True)
+ ref_table[category] = {img: functools.reduce(lambda x, y: x/y, imagesize.get(os.path.join(category_dir, img))) for img in patch_list[:200]}
+ return ref_table
+
+ @staticmethod
+ def find_nearest(array, value):
+ array = np.asarray(array)
+ idx = (np.abs(array/value - 1)).argmin()
+ return idx
+
+ @staticmethod
+ def find_k_nearest(array, value, k=1):
+ array = np.asarray(array)
+ dist = np.abs(array / value - 1)
+ idxs = np.argsort(dist)[:k]
+ return idxs
+
+ def get_instances(self, examples):
+ instances = []
+ for index, (caption, bboxes) in enumerate(zip(examples[self.caption_column], examples[self.bbox_column])):
+ categories = caption[1:]
+ instances_per_example = []
+ for name, bbox in zip(categories, bboxes):
+ if name == '':
+ instances_per_example.append(torch.zeros([self.top_k, 3, self.ref_resolution, self.ref_resolution]))
+ else:
+ value = (bbox[2] - bbox[0]) / max(bbox[3] - bbox[1], 1e-8)
+ chosen_idxs = self.find_k_nearest(list(self.ref_table[name].values()), value, k=self.top_k)
+ instances_per_bbox = []
+ for idx in chosen_idxs:
+ chosen_file = list(self.ref_table[name].keys())[idx]
+ img = Image.open(os.path.join(self.image_patch_path, name, chosen_file)).convert('RGB')
+ img = self.image_processor(images=img, return_tensors="pt")['pixel_values'].squeeze(0)
+ instances_per_bbox.append(img)
+ instances_per_example.append(torch.stack(instances_per_bbox))
+ instances.append(torch.stack(instances_per_example))
+ return instances
+
+ def train_transform(self, examples):
+ images, bboxes, obboxes = examples[self.image_column], examples[self.bbox_column], examples[self.obbox_column]
+ global_prompt = [caption[:1] for caption in examples[self.caption_column]]
+ captions = [caption[1:] for caption in examples[self.caption_column]]
+
+ new_images, new_bboxes, new_obboxes, new_captions = [], [], [], []
+ for i in range(len(images)):
+ num_instances = sum(bool(s) for s in captions[i])
+ bboxes_i = bboxes[i][:num_instances]
+ captions_i = captions[i][:num_instances]
+ transformed = self.image_transforms(image=images[i], bboxes=bboxes_i, labels=captions_i)
+ transformed["obboxes"] = []
+ for xmin, ymin, xmax, ymax in transformed["bboxes"]:
+ transformed["obboxes"].append([xmin, ymin, xmax, ymin, xmax, ymax, xmin, ymax])
+ for _ in range(len(captions[i]) - num_instances):
+ transformed["bboxes"].append([0,0,0,0])
+ transformed["obboxes"].append([0,0,0,0,0,0,0,0])
+ transformed["labels"].append("")
+ new_images.append(transformed["image"])
+ new_bboxes.append(transformed["bboxes"])
+ new_obboxes.append(transformed["obboxes"])
+ new_captions.append(global_prompt[i] + transformed["labels"])
+ examples[self.image_column] = new_images
+ examples[self.bbox_column] = new_bboxes
+ examples[self.obbox_column] = new_obboxes
+ examples[self.caption_column] = new_captions
+ return examples
+
+ def infer_transform(self, examples):
+ examples[self.image_column] = [self.image_transforms(image=image)['image'] for image in examples[self.image_column]]
+ return examples
+
+ def __call__(self, examples):
+ # print("Received keys:", examples.keys())
+ examples[self.image_column] = [np.array(image.convert("RGB")) for image in examples[self.image_column]]
+ if self.split == "train":
+ examples = self.train_transform(examples)
+ elif self.split == "infer":
+ examples = self.infer_transform(examples)
+ examples["instances"] = self.get_instances(examples)
+ return examples
\ No newline at end of file
diff --git a/figs/Main.png b/figs/Main.png
new file mode 100644
index 0000000000000000000000000000000000000000..04a68888278c424212be01feeddb63547b8fd315
--- /dev/null
+++ b/figs/Main.png
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:f8e407fae7cd699632315e76dfc45a7181d5e9933639833c7dc95254d399d2ba
+size 441808
diff --git a/fonts/GILI____.TTF b/fonts/GILI____.TTF
new file mode 100644
index 0000000000000000000000000000000000000000..35d0fe899515fe7008a0dc846dd743f34f59b7dc
Binary files /dev/null and b/fonts/GILI____.TTF differ
diff --git a/fonts/Rainbow-Party-2.ttf b/fonts/Rainbow-Party-2.ttf
new file mode 100644
index 0000000000000000000000000000000000000000..e4fb4ab384f2faa6e88a5e734d29ad3dfd111f84
--- /dev/null
+++ b/fonts/Rainbow-Party-2.ttf
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:8210f1b01b549890cc25f2028c980f87b4d2a45dbc3c64be8b60e8ec9f4d9f84
+size 114632
diff --git a/gicdm_core.py b/gicdm_core.py
new file mode 100644
index 0000000000000000000000000000000000000000..faca97daadef7c66fa433d8e05b53614b68bb702
--- /dev/null
+++ b/gicdm_core.py
@@ -0,0 +1,216 @@
+"""Core GICDM implementation reproducing the paper arXiv:2602.16449.
+
+Implements:
+ - ICDM iterative hubness reduction (Section 3 / Algorithm in paper lines 485-498)
+ - GICDM out-of-sample generated-point scaling (Eq. 1, Algorithm 1)
+ - Hubness statistics h5_1(1%) and A5 (Table 4)
+ - Clipped Density / Clipped Coverage fidelity & coverage metrics
+ - Crossover dimension d* (Proposition 5.3)
+ - Raisa et al. (2025) synthetic benchmark scenarios
+"""
+import numpy as np
+from scipy import stats
+from scipy.special import ncfdtr, ncfdtri
+
+
+# ----------------------------------------------------------------------------
+# Distances & k-nearest-neighbour helpers
+# ----------------------------------------------------------------------------
+def pairwise_sq_dists(X, Y=None):
+ if Y is None:
+ Y = X
+ Xn = np.sum(X ** 2, axis=1, keepdims=True)
+ Yn = np.sum(Y ** 2, axis=1, keepdims=True)
+ D2 = Xn + Yn.T - 2.0 * (X @ Y.T)
+ np.maximum(D2, 0, out=D2)
+ return np.sqrt(D2)
+
+
+def knn_distances(D, k, self_included=True):
+ """Return (N, k) sorted distances to k nearest neighbours.
+
+ If self_included, the 0-th distance (0) is the point itself.
+ """
+ N = D.shape[0]
+ if self_included:
+ idx = np.argpartition(D, k, axis=1)[:, :k]
+ else:
+ # exclude self (diagonal)
+ Dc = D + np.eye(N) * 1e18
+ idx = np.argpartition(Dc, k, axis=1)[:, :k]
+ idx = idx[np.arange(N)[:, None], np.argsort(D[np.arange(N)[:, None], idx], axis=1)]
+ return D[np.arange(N)[:, None], idx]
+
+
+# ----------------------------------------------------------------------------
+# ICDM (Iterative Contextual Dissimilarity Measure)
+# Matches official implementation: metrics/hubness_processor/hubness_reduction_methods.py
+# ----------------------------------------------------------------------------
+def _average_k_dist(D, k):
+ """Average of the k nearest distances (including self at distance 0),
+ i.e. sum of (k+1) smallest distances / k."""
+ k_dists = np.partition(D, k, axis=1)[:, : k + 1]
+ return k_dists.sum(axis=1) / k
+
+
+def icdm_scaling(D, K, n_iter=10, return_mu=False):
+ """ICDM scaling factors delta_i, matching official `icdm_delta_low_memory`.
+
+ Iteratively applies NICDM:
+ d_{ij} <- d_{ij} / (sqrt(r_i) sqrt(r_j)) with r_i = average of k nearest dists (k=K)
+ delta_i <- delta_i / sqrt(r_i)
+ Final: secondary dissimilarity d^T_{ij} = d_{ij} delta_i delta_j.
+ Returns delta_i (length N) and optionally final average-neighbour distances.
+ """
+ N = D.shape[0]
+ d = D.copy()
+ deltas = np.ones(N)
+ for _ in range(n_iter):
+ r = _average_k_dist(d, K)
+ sqrt_r = np.maximum(np.sqrt(r), 1e-12)
+ d = d / (sqrt_r[:, None] * sqrt_r[None, :])
+ deltas = deltas / sqrt_r
+ if return_mu:
+ mu_final = _average_k_dist(d, K)
+ return deltas, mu_final
+ return deltas
+
+
+# ----------------------------------------------------------------------------
+# GICDM (Algorithm 1)
+# ----------------------------------------------------------------------------
+def gicdm(Xr, Xg, K1, K2, q=0.95, n_iter=10, return_info=False):
+ """Generative ICDM (Algorithm 1), matching the official implementation.
+
+ Xr : (N, d) real points
+ Xg : (M, d) generated points
+ K1, K2 : two filter scales (K2 = 10*K1 in the paper). For each scale k,
+ ICDM is applied with neighbourhood size 2k (paper: GICDM K = 2k).
+ Returns:
+ D_final : (M, N) GICDM real-to-generated dissimilarity matrix
+ keep : (M,) boolean mask of generated points passing multi-scale filter
+ Drr_gicdm : (N, N) GICDM real-to-real dissimilarity matrix
+ """
+ Drr = pairwise_sq_dists(Xr)
+ Drg = pairwise_sq_dists(Xg, Xr) # (M, N): generated-to-real
+ N, M = Xr.shape[0], Xg.shape[0]
+
+ keep = np.ones(M, dtype=bool)
+ info = {}
+
+ for k in (K1, K2):
+ # ICDM with neighbourhood 2k
+ delta_r = icdm_scaling(Drr, K=2 * k, n_iter=n_iter)
+
+ # --- real filter (Algorithm 1 lines 3-6) ---
+ # k nearest real neighbours (exclude self)
+ nn_r = np.argpartition(Drr, k, axis=1)[:, : k + 1]
+ is_self = nn_r == np.arange(N)[:, None]
+ no_self = ~is_self.any(axis=1)
+ is_self[no_self, -1] = True
+ nn_r = nn_r[~is_self].reshape(N, k)
+ real_avg_delta = delta_r[nn_r].mean(axis=1)
+ r_ri = np.abs(real_avg_delta - delta_r) / real_avg_delta
+ T_k = np.quantile(r_ri, q)
+
+ # --- generated filter (Algorithm 1 lines 9-14) ---
+ # k+1 nearest real neighbours of each generated point
+ nn_g = np.argpartition(Drg.T, k, axis=1)[:, : k + 1]
+ # r_synthetic = average of (delta_real_nn * d_orig) over k+1 neighbours
+ d_g_nn = Drg[np.arange(M)[:, None], nn_g]
+ delta_r_g_nn = delta_r[nn_g]
+ r_synthetic = (delta_r_g_nn * d_g_nn).sum(axis=1) / (k + 1)
+ delta_g = 1.0 / r_synthetic # Eq. (1) with mu_bar absorbed
+ synth_avg_delta = delta_r_g_nn.mean(axis=1)
+ r_gj = np.abs(synth_avg_delta - delta_g) / synth_avg_delta
+ keep &= (r_gj <= T_k)
+
+ info[k] = dict(T_k=float(T_k), delta_g=delta_g)
+
+ # final GICDM dissimilarities (Algorithm 1 line 17)
+ delta_g_K1 = info[K1]['delta_g']
+ delta_r_K1 = icdm_scaling(Drr, K=2 * K1, n_iter=n_iter)
+ D_final = Drg * (delta_r_K1[None, :] * delta_g_K1[None, :])
+ Drr_gicdm = Drr * np.outer(delta_r_K1, delta_r_K1)
+
+ if return_info:
+ return D_final, keep, info, (delta_r_K1, delta_g_K1)
+ return D_final, keep, Drr_gicdm
+
+
+# ----------------------------------------------------------------------------
+# Hubness statistics
+# ----------------------------------------------------------------------------
+def k_occurrence(D, k=5):
+ """O_k(x_i) = number of points for which x_i is among their k NN (excl self)."""
+ N = D.shape[0]
+ Dc = D + np.eye(N) * 1e18
+ knn = np.argsort(Dc, axis=1)[:, :k]
+ occ = np.zeros(N, dtype=int)
+ for j in range(k):
+ occ[knn[:, j]] += 1
+ return occ
+
+
+def hubness_stats(D, k=5, q=0.01):
+ """Return h5_1(1%) and A5 (proportion of antihubs)."""
+ occ = k_occurrence(D, k)
+ mean_occ = occ.mean()
+ n = len(occ)
+ topq = max(1, int(np.floor(q * n)))
+ top_vals = np.sort(occ)[::-1][:topq]
+ h5 = top_vals.mean() / mean_occ if mean_occ > 0 else np.nan
+ A5 = float(np.mean(occ == 0))
+ return float(h5), A5
+
+
+# ----------------------------------------------------------------------------
+# Clipped Density / Clipped Coverage (Salvy et al. 2026)
+# ----------------------------------------------------------------------------
+def clipped_density(Xr, Xg, k=5, dissim=None, keep=None):
+ """Clipped Density fidelity metric.
+
+ For each generated point, distance to its k-th real NN; threshold = distance
+ from each real point to its k-th real NN (clip). Score averages clip term.
+ If dissim ('gicdm') is provided, use GICDM dissimilarities instead of raw dist.
+ keep: boolean mask of generated points to include (filtered-out points get 0).
+ """
+ Drg = pairwise_sq_dists(Xg, Xr) # (M,N)
+ Drr = pairwise_sq_dists(Xr)
+ if dissim is None:
+ d_rg = Drg
+ d_rr = Drr
+ else:
+ d_rg, d_rr = dissim
+ # k-th NN distance for each real point (in its own set)
+ kth_real = np.sort(d_rr + np.eye(len(Xr)) * 1e18, axis=1)[:, k - 1]
+ kth_gen = np.sort(d_rg, axis=1)[:, k - 1]
+ # clip each generated point's distance at its matched real threshold
+ thresh = kth_real[np.argmin(d_rg, axis=1)]
+ clip = np.clip(kth_gen / thresh, 0, 1)
+ if keep is not None:
+ clip = clip * keep.astype(float) # filtered points -> fidelity 0
+ return float(np.mean(clip))
+
+
+def clipped_coverage(Xr, Xg, k=5, dissim=None, keep=None):
+ """Clipped Coverage: for each real point, does a kept generated point fall
+ within its k-th NN threshold?"""
+ Drg = pairwise_sq_dists(Xg, Xr)
+ Drr = pairwise_sq_dists(Xr)
+ if dissim is None:
+ d_rg = Drg
+ d_rr = Drr
+ else:
+ d_rg, d_rr = dissim
+ kth_real = np.sort(d_rr + np.eye(len(Xr)) * 1e18, axis=1)[:, k - 1]
+ min_d = d_rg.min(axis=0)
+ covered = (min_d <= kth_real).astype(float)
+ if keep is not None:
+ # a generated point only contributes if it is kept
+ d_rg_k = d_rg[keep, :]
+ if d_rg_k.shape[0] == 0:
+ return 0.0
+ min_d2 = d_rg_k.min(axis=0)
+ covered = (min_d2 <= kth_real).astype(float)
+ return float(np.mean(covered))
diff --git a/infer.sh b/infer.sh
new file mode 100644
index 0000000000000000000000000000000000000000..d768fee9034b3d2ecbec00b7737ec6dd5e4ec8b5
--- /dev/null
+++ b/infer.sh
@@ -0,0 +1,174 @@
+#!/usr/bin/env bash
+set -e
+
+CONFIGS=()
+SEEDS=()
+CKPTS=()
+RUN_IDS=()
+GPU_IDS=""
+METASEED=""
+NUM_SEED=-1
+META_START=0
+MAX_INFER_SIZE=""
+K_SHOTS=()
+
+while [[ $# -gt 0 ]]; do
+ case $1 in
+ --config)
+ shift
+ CONFIGS=($1)
+ ;;
+ --seed)
+ shift
+ SEEDS=($1)
+ ;;
+ --metaseed)
+ shift
+ METASEED=$1
+ ;;
+ --num_seed)
+ shift
+ NUM_SEED=$1
+ ;;
+ --meta_start)
+ shift
+ META_START=$1
+ ;;
+ --ckpt)
+ shift
+ CKPTS=($1)
+ ;;
+ --run_id)
+ shift
+ RUN_IDS=($1)
+ ;;
+ --k_shot)
+ shift
+ K_SHOTS=($1)
+ ;;
+ --gpu_ids)
+ shift
+ GPU_IDS=$1
+ ;;
+ --max_infer_size)
+ shift
+ MAX_INFER_SIZE=$1
+ ;;
+ *)
+ echo "Unknown argument: $1"
+ exit 1
+ ;;
+ esac
+ shift
+done
+
+if [[ ${#CONFIGS[@]} -eq 0 ]] \
+ || [[ ${#CKPTS[@]} -eq 0 ]] \
+ || [[ -z "$GPU_IDS" ]] \
+ || [[ ${#K_SHOTS[@]} -eq 0 ]] \
+ || [[ ${#RUN_IDS[@]} -eq 0 ]]; then
+ echo "Error: missing required arguments."
+ echo "You must provide: --config, --gpu_ids, --run_id, --k_shot"
+ echo "And one of: --seed OR --metaseed"
+ exit 1
+fi
+
+if [[ -z "$METASEED" && ${#SEEDS[@]} -eq 0 ]]; then
+ echo "Error: either --seed or --metaseed must be provided."
+ exit 1
+fi
+
+IFS=',' read -r -a GPU_ARRAY <<< "$GPU_IDS"
+NUM_PROCESSES=${#GPU_ARRAY[@]}
+
+if [[ -n "$METASEED" ]]; then
+ if [[ $NUM_SEED -lt 0 ]]; then
+ echo "Error: you must set --num_seed when using --metaseed."
+ exit 1
+ fi
+
+ echo "[INFO] Using metaseed=$METASEED"
+ echo "[INFO] Will use seeds index range [$META_START, $NUM_SEED)"
+
+ mapfile -t ALL_SEEDS < <(
+ shuf -i 0-9999 \
+ --random-source=<(awk -v s="$METASEED" 'BEGIN { while (1) printf "%s", s }') \
+ | head -n $NUM_SEED
+ )
+
+ # SEEDS=("${ALL_SEEDS[@]:META_START:NUM_SEED-META_START}")
+ SEEDS=("${ALL_SEEDS[@]:$META_START:$((NUM_SEED - META_START))}")
+
+ echo "[INFO] Total Seeds Generated: ${#ALL_SEEDS[@]}"
+ echo "[INFO] Using Seeds: ${SEEDS[@]}"
+fi
+
+echo "CONFIGS: ${CONFIGS[@]}"
+echo "SEEDS: ${SEEDS[@]}"
+echo "RUN_IDS: ${RUN_IDS[@]}"
+echo "K_SHOTS: ${K_SHOTS[@]}"
+echo "CKPTS: ${CKPTS[@]}"
+echo "GPU_IDS: ${GPU_IDS}"
+echo "NUM_PROCESSES: $NUM_PROCESSES"
+echo "META_START: $META_START"
+echo "NUM_SEED: $NUM_SEED"
+echo "Actual inference episodes: ${#SEEDS[@]}"
+echo "MAX_INFER_SIZE: ${MAX_INFER_SIZE:-"(unset)"}"
+
+EXTRA_ARG=""
+if [[ -n "$MAX_INFER_SIZE" ]]; then
+ EXTRA_ARG="-M $MAX_INFER_SIZE"
+fi
+
+run_with_retry(){
+ local cmd=("$@")
+ local max=3
+ local attempt=1
+
+ while (( attempt <= max )); do
+ echo "Attempt $attempt/$max: ${cmd[*]}"
+
+ set +e
+ "${cmd[@]}"
+ status=$?
+ set -e
+
+ if [[ $status -eq 0 ]]; then
+ return 0
+ fi
+
+ ((attempt++))
+ sleep 2
+ done
+
+ return 1
+}
+
+for k_shot in "${K_SHOTS[@]}"; do
+ for run_id in "${RUN_IDS[@]}"; do
+ for config in "${CONFIGS[@]}"; do
+ for seed in "${SEEDS[@]}"; do
+ for ckpt in "${CKPTS[@]}"; do
+ echo "config: $config, seed: $seed, ckpt: $ckpt, run: $run_id"
+
+ max_wait=12
+ ckpt_path="./ckpt/$config/novel/run-$run_id/$k_shot-shot/shuffle_seed-$seed/checkpoint-$ckpt"
+ while [[ ! -s "$ckpt_path" ]]; do
+ echo "waiting for ckpt: $ckpt_path"
+ sleep 300
+ ((max_wait--))
+ done
+
+ if [[ ! -s "$ckpt_path" ]]; then
+ echo "[ERROR] ckpt not found after waiting, skip."
+ continue
+ fi
+
+ run_with_retry \
+ accelerate launch --multi_gpu --gpu_ids $GPU_IDS --num_processes $NUM_PROCESSES main.py \
+ --config ./configs/$config.yaml -m infer -p novel -s $seed -k $k_shot -r $run_id -c $ckpt $EXTRA_ARG
+ done
+ done
+ done
+ done
+done
\ No newline at end of file
diff --git a/main.py b/main.py
new file mode 100644
index 0000000000000000000000000000000000000000..2c324ad017244190cd4955646c2f42fdb177aaa0
--- /dev/null
+++ b/main.py
@@ -0,0 +1,12 @@
+from utils import load_config
+import variants
+
+if __name__ == '__main__':
+ config = load_config()
+
+ if config.mode == 'train':
+ entry = getattr(variants.train, config.training.get('entry', 'train_dsp'))
+ elif config.mode == 'infer':
+ entry = getattr(variants.infer, config.inference.get('entry', 'infer_dsp'))
+
+ entry(config)
\ No newline at end of file
diff --git a/models/__init__.py b/models/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..f260dce827faadc46fb9cc8972404f96c8f44f8a
--- /dev/null
+++ b/models/__init__.py
@@ -0,0 +1 @@
+from . import dsp
\ No newline at end of file
diff --git a/models/dsp/CAM/__init__.py b/models/dsp/CAM/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..25b3afe7ae1f8d71e6d118e632a2bb97f12d98f1
--- /dev/null
+++ b/models/dsp/CAM/__init__.py
@@ -0,0 +1 @@
+from .cam_generator import CAMGenerator
\ No newline at end of file
diff --git a/models/dsp/CAM/cam_generator.py b/models/dsp/CAM/cam_generator.py
new file mode 100644
index 0000000000000000000000000000000000000000..a480557440bb7e9780f18f7e20b4fb23516db0a4
--- /dev/null
+++ b/models/dsp/CAM/cam_generator.py
@@ -0,0 +1,133 @@
+from .pytorch_grad_cam import GradCAM
+from .pytorch_grad_cam.utils.image import scale_cam_image
+from . import clip
+from .utils import reshape_transform, zeroshot_classifier, ClipOutputTarget, scoremap2bbox
+
+from torchvision import transforms
+import torch
+import numpy as np
+import cv2
+
+class CAMGenerator:
+ def __init__(self, categories, clip_path):
+ # self.device = "cuda" if torch.cuda.is_available() else "cpu"
+
+ self.clip_path = clip_path
+ self.clip_model, _ = clip.load(self.clip_path, device="cpu")
+ self.target_layers = [self.clip_model.visual.transformer.resblocks[-1].ln_1]
+ self.cam = GradCAM(model=self.clip_model, target_layers=self.target_layers, reshape_transform=reshape_transform, use_cuda=True)
+
+ self.categories = categories
+ # if categories is None:
+ # self.categories = ['vehicle', 'baseballfield', 'groundtrackfield', 'windmill', 'bridge', 'overpass', 'ship', 'airplane', 'tenniscourt', 'airport',
+ # 'expressway-service-area', 'basketballcourt', 'stadium', 'storagetank', 'chimney', 'dam', 'expressway-toll-station', 'golffield', 'trainstation', 'harbor']
+ self.background_categories = ['ground','land','grass','tree','building','wall','sky','lake','water','river','sea','railway','railroad','keyboard','helmet',
+ 'cloud','house','mountain','ocean','road','rock','street','valley','bridge','sign',]
+
+ self.normalize = transforms.Normalize((0.48145466, 0.4578275, 0.40821073), (0.26862954, 0.26130258, 0.27577711))
+
+ def _prepare(self):
+ self.bg_text_features = zeroshot_classifier(self.background_categories, ['a clean origami {}.'], self.clip_model, self.device)
+ self.fg_text_features = zeroshot_classifier(self.categories, ['a clean origami {}.'], self.clip_model, self.device)
+
+ def to(self, device, dtype):
+ self.device = device
+ self.clip_model.to(device)
+ self.cam.set_device(device)
+ self._prepare()
+
+ def re_normalize(self, image):
+ # image = (image / 2) + 0.5
+ image = self.normalize(image)
+ return image
+
+ def get_label_list(self, captions):
+ label_list = []
+ label_id_list = []
+ for caption in captions:
+ if caption in self.categories and caption not in label_list:
+ label_list.append(caption)
+ label_id_list.append(self.categories.index(caption))
+ return label_list, label_id_list
+
+ def get_label_list_with_bboxes(self, captions, bboxes):
+ label_list, label_id_list, bboxes_list = [], [], []
+ for caption, bbox in zip(captions, bboxes):
+ if caption not in self.categories:
+ continue
+ if caption not in label_list:
+ label_list.append(caption)
+ label_id_list.append(self.categories.index(caption))
+ bboxes_list.append([bbox])
+ else:
+ bboxes_list[label_list.index(caption)].append(bbox)
+ return label_list, label_id_list, bboxes_list
+
+ def __call__(self, image, captions, bboxes, gt_bboxes_only=False):
+ image = self.re_normalize(image)
+ # label_list, label_id_list = self.get_label_list(captions[0][1:])
+ label_list, label_id_list, bboxes_list = self.get_label_list_with_bboxes(captions[0][1:], bboxes[0])
+ h, w = image.shape[-2], image.shape[-1]
+ image_features, attn_weight_list = self.clip_model.encode_image(image, h, w)
+
+ bg_features_temp = self.bg_text_features
+ fg_features_temp = self.fg_text_features[label_id_list]
+ text_features_temp = torch.cat([fg_features_temp, bg_features_temp], dim=0)
+ input_tensor = [image_features, text_features_temp, h, w]
+
+ keys, refined_cam_list = [], []
+ for idx, (label, bbox) in enumerate(zip(label_list, bboxes_list)):
+ keys.append(self.categories.index(label))
+ targets = [ClipOutputTarget(label_list.index(label))]
+ grayscale_cam, logits_per_image, attn_weight_last = self.cam(input_tensor=input_tensor, targets=targets, target_size=None)
+ grayscale_cam = grayscale_cam[0, :] # [32, 32]
+ # grayscale_cam_highres = cv2.resize(grayscale_cam, (ori_width, ori_height))
+
+ if idx == 0:
+ attn_weight_list.append(attn_weight_last)
+ attn_weight = [aw[:, 1:, 1:] for aw in attn_weight_list] # (b, hxw, hxw)
+ attn_weight = torch.stack(attn_weight, dim=0)[-8:] # [8, 1, 1024, 1024]
+ attn_weight = torch.mean(attn_weight, dim=0) # [1, 1024, 1024]
+ # attn_weight = attn_weight[0].detach() # [1024, 1024] # original detach
+ attn_weight = attn_weight[0] #.detach() # [1024, 1024]
+ attn_weight = attn_weight.float()
+
+ gt_box, gt_cnt = (np.array(bbox) * [grayscale_cam.shape[1], grayscale_cam.shape[0], grayscale_cam.shape[1], grayscale_cam.shape[0]]).astype(int), len(bbox)
+ if gt_bboxes_only:
+ box, cnt = gt_box, gt_cnt
+ else:
+ box, cnt = scoremap2bbox(scoremap=grayscale_cam.cpu().data.numpy(), threshold=0.4, multi_contour_eval=True)
+ box, cnt = np.concatenate([box, gt_box], axis=0), cnt + gt_cnt
+ aff_mask = torch.zeros_like(grayscale_cam)
+ for i_ in range(cnt):
+ x0_, y0_, x1_, y1_ = box[i_]
+ aff_mask[y0_:y1_, x0_:x1_] = 1
+ aff_mask = aff_mask.view(1, grayscale_cam.shape[0] * grayscale_cam.shape[1])
+
+ aff_mat = attn_weight
+ trans_mat = aff_mat / torch.sum(aff_mat, dim=0, keepdim=True)
+ trans_mat = trans_mat / torch.sum(trans_mat, dim=1, keepdim=True)
+ for _ in range(2):
+ trans_mat = trans_mat / torch.sum(trans_mat, dim=0, keepdim=True)
+ trans_mat = trans_mat / torch.sum(trans_mat, dim=1, keepdim=True)
+ trans_mat = (trans_mat + trans_mat.transpose(1, 0)) / 2
+ for _ in range(1):
+ trans_mat = torch.matmul(trans_mat, trans_mat)
+
+ trans_mat = trans_mat * aff_mask
+
+ cam_to_refine = grayscale_cam.view(-1, 1)
+ cam_refined = torch.matmul(trans_mat, cam_to_refine).reshape(h //16, w // 16)
+ cam_refined = cam_refined - cam_refined.min()
+ cam_refined = cam_refined / (cam_refined.max() + 1e-7)
+ refined_cam_list.append(cam_refined)
+
+ keys = torch.tensor(keys)
+ refined_cams = torch.stack(refined_cam_list, dim=0)
+
+ return refined_cams, keys
+
+
+if __name__ == '__main__':
+ cam = CAMGenerator()
+
\ No newline at end of file
diff --git a/models/dsp/CAM/clip/__init__.py b/models/dsp/CAM/clip/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..dcc5619538c0f7c782508bdbd9587259d805e0d9
--- /dev/null
+++ b/models/dsp/CAM/clip/__init__.py
@@ -0,0 +1 @@
+from .clip import *
diff --git a/models/dsp/CAM/clip/bpe_simple_vocab_16e6.txt.gz b/models/dsp/CAM/clip/bpe_simple_vocab_16e6.txt.gz
new file mode 100644
index 0000000000000000000000000000000000000000..36a15856e00a06a9fbed8cdd34d2393fea4a3113
--- /dev/null
+++ b/models/dsp/CAM/clip/bpe_simple_vocab_16e6.txt.gz
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:924691ac288e54409236115652ad4aa250f48203de50a9e4722a6ecd48d6804a
+size 1356917
diff --git a/models/dsp/CAM/clip/clip.py b/models/dsp/CAM/clip/clip.py
new file mode 100644
index 0000000000000000000000000000000000000000..d73fc6f06d4b3984e030bd4c663ddc71fb3b3831
--- /dev/null
+++ b/models/dsp/CAM/clip/clip.py
@@ -0,0 +1,245 @@
+import hashlib
+import os
+import urllib
+import warnings
+from typing import Any, Union, List
+from packaging import version
+
+import torch
+from PIL import Image
+from torchvision.transforms import Compose, Resize, CenterCrop, ToTensor, Normalize
+from tqdm import tqdm
+
+from .model import build_model
+from .simple_tokenizer import SimpleTokenizer as _Tokenizer
+from collections import OrderedDict
+
+try:
+ from torchvision.transforms import InterpolationMode
+ BICUBIC = InterpolationMode.BICUBIC
+except ImportError:
+ BICUBIC = Image.BICUBIC
+
+
+if version.parse(torch.__version__) < version.parse("1.7.1"):
+ warnings.warn("PyTorch version 1.7.1 or higher is recommended")
+
+
+__all__ = ["available_models", "load", "tokenize"]
+_tokenizer = _Tokenizer()
+
+_MODELS = {
+ "RN50": "https://openaipublic.azureedge.net/clip/models/afeb0e10f9e5a86da6080e35cf09123aca3b358a0c3e3b6c78a7b63bc04b6762/RN50.pt",
+ "RN101": "https://openaipublic.azureedge.net/clip/models/8fa8567bab74a42d41c5915025a8e4538c3bdbe8804a470a72f30b0d94fab599/RN101.pt",
+ "RN50x4": "https://openaipublic.azureedge.net/clip/models/7e526bd135e493cef0776de27d5f42653e6b4c8bf9e0f653bb11773263205fdd/RN50x4.pt",
+ "RN50x16": "https://openaipublic.azureedge.net/clip/models/52378b407f34354e150460fe41077663dd5b39c54cd0bfd2b27167a4a06ec9aa/RN50x16.pt",
+ "RN50x64": "https://openaipublic.azureedge.net/clip/models/be1cfb55d75a9666199fb2206c106743da0f6468c9d327f3e0d0a543a9919d9c/RN50x64.pt",
+ "ViT-B/32": "https://openaipublic.azureedge.net/clip/models/40d365715913c9da98579312b702a82c18be219cc2a73407c4526f58eba950af/ViT-B-32.pt",
+ "ViT-B/16": "https://openaipublic.azureedge.net/clip/models/5806e77cd80f8b59890b7e101eabd078d9fb84e6937f9e85e4ecb61988df416f/ViT-B-16.pt",
+ "ViT-L/14": "https://openaipublic.azureedge.net/clip/models/b8cca3fd41ae0c99ba7e8951adf17d267cdb84cd88be6f7c2e0eca1737a03836/ViT-L-14.pt",
+ "ViT-L/14@336px": "https://openaipublic.azureedge.net/clip/models/3035c92b350959924f9f00213499208652fc7ea050643e8b385c2dac08641f02/ViT-L-14-336px.pt",
+}
+
+
+def _download(url: str, root: str):
+ os.makedirs(root, exist_ok=True)
+ filename = os.path.basename(url)
+
+ expected_sha256 = url.split("/")[-2]
+ download_target = os.path.join(root, filename)
+
+ if os.path.exists(download_target) and not os.path.isfile(download_target):
+ raise RuntimeError(f"{download_target} exists and is not a regular file")
+
+ if os.path.isfile(download_target):
+ if hashlib.sha256(open(download_target, "rb").read()).hexdigest() == expected_sha256:
+ return download_target
+ else:
+ warnings.warn(f"{download_target} exists, but the SHA256 checksum does not match; re-downloading the file")
+
+ with urllib.request.urlopen(url) as source, open(download_target, "wb") as output:
+ with tqdm(total=int(source.info().get("Content-Length")), ncols=80, unit='iB', unit_scale=True, unit_divisor=1024) as loop:
+ while True:
+ buffer = source.read(8192)
+ if not buffer:
+ break
+
+ output.write(buffer)
+ loop.update(len(buffer))
+
+ if hashlib.sha256(open(download_target, "rb").read()).hexdigest() != expected_sha256:
+ raise RuntimeError(f"Model has been downloaded but the SHA256 checksum does not not match")
+
+ return download_target
+
+
+def _convert_image_to_rgb(image):
+ return image.convert("RGB")
+
+
+def _transform(n_px):
+ return Compose([
+ Resize(n_px, interpolation=BICUBIC),
+ CenterCrop(n_px),
+ _convert_image_to_rgb,
+ ToTensor(),
+ Normalize((0.48145466, 0.4578275, 0.40821073), (0.26862954, 0.26130258, 0.27577711)),
+ ])
+
+
+def available_models() -> List[str]:
+ """Returns the names of available CLIP models"""
+ return list(_MODELS.keys())
+
+
+def load(name: str, device: Union[str, torch.device] = "cuda" if torch.cuda.is_available() else "cpu", jit: bool = False, download_root: str = None):
+ """Load a CLIP model
+
+ Parameters
+ ----------
+ name : str
+ A model name listed by `clip.available_models()`, or the path to a model checkpoint containing the state_dict
+
+ device : Union[str, torch.device]
+ The device to put the loaded model
+
+ jit : bool
+ Whether to load the optimized JIT model or more hackable non-JIT model (default).
+
+ download_root: str
+ path to download the model files; by default, it uses "~/.cache/clip"
+
+ Returns
+ -------
+ model : torch.nn.Module
+ The CLIP model
+
+ preprocess : Callable[[PIL.Image], torch.Tensor]
+ A torchvision transform that converts a PIL image into a tensor that the returned model can take as its input
+ """
+ if name in _MODELS:
+ model_path = _download(_MODELS[name], download_root or os.path.expanduser("~/.cache/clip"))
+ elif os.path.isfile(name):
+ model_path = name
+ else:
+ raise RuntimeError(f"Model {name} not found; available models = {available_models()}")
+
+ with open(model_path, 'rb') as opened_file:
+ try:
+ # loading JIT archive
+ model = torch.jit.load(opened_file, map_location=device if jit else "cpu").eval()
+ state_dict = None
+ except RuntimeError:
+ # loading saved state dict
+ if jit:
+ warnings.warn(f"File {model_path} is not a JIT archive. Loading as a state dict instead")
+ jit = False
+ if 'RN50' in model_path:
+ state_dict = torch.load(opened_file, map_location="cpu")
+ else:
+ state_dict0 = torch.load(model_path, map_location="cpu")
+ state_dict = OrderedDict()
+ for k in state_dict0.keys():
+ state_dict[k.replace('module.', '')] = state_dict0[k]
+
+
+ if not jit:
+ model = build_model(state_dict or model.state_dict()).to(device)
+ if str(device) == "cpu":
+ model.float()
+ return model, _transform(model.visual.input_resolution)
+
+ # patch the device names
+ device_holder = torch.jit.trace(lambda: torch.ones([]).to(torch.device(device)), example_inputs=[])
+ device_node = [n for n in device_holder.graph.findAllNodes("prim::Constant") if "Device" in repr(n)][-1]
+
+ def patch_device(module):
+ try:
+ graphs = [module.graph] if hasattr(module, "graph") else []
+ except RuntimeError:
+ graphs = []
+
+ if hasattr(module, "forward1"):
+ graphs.append(module.forward1.graph)
+
+ for graph in graphs:
+ for node in graph.findAllNodes("prim::Constant"):
+ if "value" in node.attributeNames() and str(node["value"]).startswith("cuda"):
+ node.copyAttributes(device_node)
+
+ model.apply(patch_device)
+ patch_device(model.encode_image)
+ patch_device(model.encode_text)
+
+ # patch dtype to float32 on CPU
+ if str(device) == "cpu":
+ float_holder = torch.jit.trace(lambda: torch.ones([]).float(), example_inputs=[])
+ float_input = list(float_holder.graph.findNode("aten::to").inputs())[1]
+ float_node = float_input.node()
+
+ def patch_float(module):
+ try:
+ graphs = [module.graph] if hasattr(module, "graph") else []
+ except RuntimeError:
+ graphs = []
+
+ if hasattr(module, "forward1"):
+ graphs.append(module.forward1.graph)
+
+ for graph in graphs:
+ for node in graph.findAllNodes("aten::to"):
+ inputs = list(node.inputs())
+ for i in [1, 2]: # dtype can be the second or third argument to aten::to()
+ if inputs[i].node()["value"] == 5:
+ inputs[i].node().copyAttributes(float_node)
+
+ model.apply(patch_float)
+ patch_float(model.encode_image)
+ patch_float(model.encode_text)
+
+ model.float()
+
+ return model, _transform(model.input_resolution.item())
+
+
+def tokenize(texts: Union[str, List[str]], context_length: int = 77, truncate: bool = False) -> Union[torch.IntTensor, torch.LongTensor]:
+ """
+ Returns the tokenized representation of given input string(s)
+
+ Parameters
+ ----------
+ texts : Union[str, List[str]]
+ An input string or a list of input strings to tokenize
+
+ context_length : int
+ The context length to use; all CLIP models use 77 as the context length
+
+ truncate: bool
+ Whether to truncate the text in case its encoding is longer than the context length
+
+ Returns
+ -------
+ A two-dimensional tensor containing the resulting tokens, shape = [number of input strings, context_length].
+ We return LongTensor when torch version is <1.8.0, since older index_select requires indices to be long.
+ """
+ if isinstance(texts, str):
+ texts = [texts]
+
+ sot_token = _tokenizer.encoder["<|startoftext|>"]
+ eot_token = _tokenizer.encoder["<|endoftext|>"]
+ all_tokens = [[sot_token] + _tokenizer.encode(text) + [eot_token] for text in texts]
+ if version.parse(torch.__version__) < version.parse("1.8.0"):
+ result = torch.zeros(len(all_tokens), context_length, dtype=torch.long)
+ else:
+ result = torch.zeros(len(all_tokens), context_length, dtype=torch.int)
+
+ for i, tokens in enumerate(all_tokens):
+ if len(tokens) > context_length:
+ if truncate:
+ tokens = tokens[:context_length]
+ tokens[-1] = eot_token
+ else:
+ raise RuntimeError(f"Input {texts[i]} is too long for context length {context_length}")
+ result[i, :len(tokens)] = torch.tensor(tokens)
+
+ return result
diff --git a/models/dsp/CAM/clip/model.py b/models/dsp/CAM/clip/model.py
new file mode 100644
index 0000000000000000000000000000000000000000..5b6a770517aa3b3e4d2af837e72b66d9a3e2ded9
--- /dev/null
+++ b/models/dsp/CAM/clip/model.py
@@ -0,0 +1,517 @@
+from collections import OrderedDict
+from typing import Tuple, Union
+
+import numpy as np
+import torch
+import torch.nn.functional as F
+from torch import nn
+
+def upsample_pos_emb(emb, new_size):
+ # upsample the pretrained embedding for higher resolution
+ # emb size NxD
+ first = emb[:1, :]
+ emb = emb[1:, :]
+ N, D = emb.size(0), emb.size(1)
+ size = int(np.sqrt(N))
+ assert size * size == N
+ #new_size = size * self.upsample
+ emb = emb.permute(1, 0)
+ emb = emb.view(1, D, size, size).contiguous()
+ # emb = F.upsample(emb, size=new_size, mode='bilinear',)
+ emb = F.interpolate(emb, size=new_size, mode='bilinear', align_corners=False)
+ emb = emb.view(D, -1).contiguous()
+ emb = emb.permute(1, 0)
+ emb = torch.cat([first, emb], 0)
+ emb = nn.parameter.Parameter(emb.half())
+ return emb
+
+class Bottleneck(nn.Module):
+ expansion = 4
+
+ def __init__(self, inplanes, planes, stride=1):
+ super().__init__()
+
+ # all conv layers have stride 1. an avgpool is performed after the second convolution when stride > 1
+ self.conv1 = nn.Conv2d(inplanes, planes, 1, bias=False)
+ self.bn1 = nn.BatchNorm2d(planes)
+ self.relu1 = nn.ReLU(inplace=True)
+
+ self.conv2 = nn.Conv2d(planes, planes, 3, padding=1, bias=False)
+ self.bn2 = nn.BatchNorm2d(planes)
+ self.relu2 = nn.ReLU(inplace=True)
+
+ self.avgpool = nn.AvgPool2d(stride) if stride > 1 else nn.Identity()
+
+ self.conv3 = nn.Conv2d(planes, planes * self.expansion, 1, bias=False)
+ self.bn3 = nn.BatchNorm2d(planes * self.expansion)
+ self.relu3 = nn.ReLU(inplace=True)
+
+ self.downsample = None
+ self.stride = stride
+
+ if stride > 1 or inplanes != planes * Bottleneck.expansion:
+ # downsampling layer is prepended with an avgpool, and the subsequent convolution has stride 1
+ self.downsample = nn.Sequential(OrderedDict([
+ ("-1", nn.AvgPool2d(stride)),
+ ("0", nn.Conv2d(inplanes, planes * self.expansion, 1, stride=1, bias=False)),
+ ("1", nn.BatchNorm2d(planes * self.expansion))
+ ]))
+
+ def forward(self, x: torch.Tensor):
+ identity = x
+
+ out = self.relu1(self.bn1(self.conv1(x)))
+ out = self.relu2(self.bn2(self.conv2(out)))
+ out = self.avgpool(out)
+ out = self.bn3(self.conv3(out))
+
+ if self.downsample is not None:
+ identity = self.downsample(x)
+
+ out += identity
+ out = self.relu3(out)
+ return out
+
+
+class AttentionPool2d(nn.Module):
+ def __init__(self, spacial_dim: int, embed_dim: int, num_heads: int, output_dim: int = None):
+ super().__init__()
+ self.positional_embedding = nn.Parameter(torch.randn(spacial_dim ** 2 + 1, embed_dim) / embed_dim ** 0.5)
+ self.k_proj = nn.Linear(embed_dim, embed_dim)
+ self.q_proj = nn.Linear(embed_dim, embed_dim)
+ self.v_proj = nn.Linear(embed_dim, embed_dim)
+ self.c_proj = nn.Linear(embed_dim, output_dim or embed_dim)
+ self.num_heads = num_heads
+
+ def forward(self, x, H, W):
+ x = x.reshape(x.shape[0], x.shape[1], x.shape[2] * x.shape[3]).permute(2, 0, 1) # NCHW -> (HW)NC
+ x = torch.cat([x.mean(dim=0, keepdim=True), x], dim=0) # (HW+1)NC
+ self.positional_embedding_new = upsample_pos_emb(self.positional_embedding, (H//32,W//32))
+ x = x + self.positional_embedding_new[:, None, :].to(x.dtype) # (HW+1)NC
+ x, attn_weight = F.multi_head_attention_forward(
+ query=x, key=x, value=x,
+ embed_dim_to_check=x.shape[-1],
+ num_heads=self.num_heads,
+ q_proj_weight=self.q_proj.weight,
+ k_proj_weight=self.k_proj.weight,
+ v_proj_weight=self.v_proj.weight,
+ in_proj_weight=None,
+ in_proj_bias=torch.cat([self.q_proj.bias, self.k_proj.bias, self.v_proj.bias]),
+ bias_k=None,
+ bias_v=None,
+ add_zero_attn=False,
+ dropout_p=0,
+ out_proj_weight=self.c_proj.weight,
+ out_proj_bias=self.c_proj.bias,
+ use_separate_proj_weight=True,
+ training=self.training,
+ need_weights=False
+ )
+ return x[0]
+
+
+class ModifiedResNet(nn.Module):
+ """
+ A ResNet class that is similar to torchvision's but contains the following changes:
+ - There are now 3 "stem" convolutions as opposed to 1, with an average pool instead of a max pool.
+ - Performs anti-aliasing strided convolutions, where an avgpool is prepended to convolutions with stride > 1
+ - The final pooling layer is a QKV attention instead of an average pool
+ """
+
+ def __init__(self, layers, output_dim, heads, input_resolution=224, width=64):
+ super().__init__()
+ self.output_dim = output_dim
+ self.input_resolution = input_resolution
+
+ # the 3-layer stem
+ self.conv1 = nn.Conv2d(3, width // 2, kernel_size=3, stride=2, padding=1, bias=False)
+ self.bn1 = nn.BatchNorm2d(width // 2)
+ self.relu1 = nn.ReLU(inplace=True)
+ self.conv2 = nn.Conv2d(width // 2, width // 2, kernel_size=3, padding=1, bias=False)
+ self.bn2 = nn.BatchNorm2d(width // 2)
+ self.relu2 = nn.ReLU(inplace=True)
+ self.conv3 = nn.Conv2d(width // 2, width, kernel_size=3, padding=1, bias=False)
+ self.bn3 = nn.BatchNorm2d(width)
+ self.relu3 = nn.ReLU(inplace=True)
+ self.avgpool = nn.AvgPool2d(2)
+
+ # residual layers
+ self._inplanes = width # this is a *mutable* variable used during construction
+ self.layer1 = self._make_layer(width, layers[0])
+ self.layer2 = self._make_layer(width * 2, layers[1], stride=2)
+ self.layer3 = self._make_layer(width * 4, layers[2], stride=2)
+ self.layer4 = self._make_layer(width * 8, layers[3], stride=2)
+
+ embed_dim = width * 32 # the ResNet feature dimension
+ self.attnpool = AttentionPool2d(input_resolution // 32, embed_dim, heads, output_dim)
+
+ def _make_layer(self, planes, blocks, stride=1):
+ layers = [Bottleneck(self._inplanes, planes, stride)]
+
+ self._inplanes = planes * Bottleneck.expansion
+ for _ in range(1, blocks):
+ layers.append(Bottleneck(self._inplanes, planes))
+
+ return nn.Sequential(*layers)
+
+ def forward(self, x, H, W):
+ def stem(x):
+ x = self.relu1(self.bn1(self.conv1(x)))
+ x = self.relu2(self.bn2(self.conv2(x)))
+ x = self.relu3(self.bn3(self.conv3(x)))
+ x = self.avgpool(x)
+ return x
+
+ x = x.type(self.conv1.weight.dtype)
+ x = stem(x)
+ x = self.layer1(x)
+ x = self.layer2(x)
+ x = self.layer3(x)
+ x = self.layer4(x)#(1,,2048, 7, 7)
+ x_pooled = self.attnpool(x, H, W)
+
+ return x_pooled
+
+
+class LayerNorm(nn.LayerNorm):
+ """Subclass torch's LayerNorm to handle fp16."""
+
+ def forward(self, x: torch.Tensor):
+ orig_type = x.dtype
+ ret = super().forward(x.type(torch.float32))
+ return ret.type(orig_type)
+
+
+class QuickGELU(nn.Module):
+ def forward(self, x: torch.Tensor):
+ return x * torch.sigmoid(1.702 * x)
+
+
+class ResidualAttentionBlock(nn.Module):
+ def __init__(self, d_model: int, n_head: int, attn_mask: torch.Tensor = None):
+ super().__init__()
+
+ self.attn = nn.MultiheadAttention(d_model, n_head)
+ self.ln_1 = LayerNorm(d_model)
+ self.mlp = nn.Sequential(OrderedDict([
+ ("c_fc", nn.Linear(d_model, d_model * 4)),
+ ("gelu", QuickGELU()),
+ ("c_proj", nn.Linear(d_model * 4, d_model))
+ ]))
+ self.ln_2 = LayerNorm(d_model)
+ self.attn_mask = attn_mask
+
+ def attention(self, x: torch.Tensor):
+ self.attn_mask = self.attn_mask.to(dtype=x.dtype, device=x.device) if self.attn_mask is not None else None
+ return self.attn(x, x, x, need_weights=True, attn_mask=self.attn_mask)#[0]
+
+ def forward(self, x: torch.Tensor):
+ attn_output, attn_weight = self.attention(self.ln_1(x))#(L,N,E) (N,L,L)
+ x = x + attn_output
+ x = x + self.mlp(self.ln_2(x))
+ return x, attn_weight
+
+
+
+class Transformer(nn.Module):
+ def __init__(self, width: int, layers: int, heads: int, attn_mask: torch.Tensor = None):
+ super().__init__()
+ self.width = width
+ self.layers = layers
+ self.resblocks = nn.Sequential(*[ResidualAttentionBlock(width, heads, attn_mask) for _ in range(layers)])
+
+ def forward(self, x: torch.Tensor):
+ attn_weights = []
+ with torch.no_grad():
+ layers = self.layers if x.shape[0] == 77 else self.layers-1
+ for i in range(layers):
+ x, attn_weight = self.resblocks[i](x)
+ attn_weights.append(attn_weight)
+ '''
+ for i in range(self.layers-1, self.layers):
+ x, attn_weight = self.resblocks[i](x)
+ attn_weights.append(attn_weight)
+ #feature_map_list.append(x)
+ '''
+ return x, attn_weights
+
+
+class VisionTransformer(nn.Module):
+ def __init__(self, input_resolution: int, patch_size: int, width: int, layers: int, heads: int, output_dim: int):
+ super().__init__()
+ self.input_resolution = input_resolution
+ self.output_dim = output_dim
+ self.conv1 = nn.Conv2d(in_channels=3, out_channels=width, kernel_size=patch_size, stride=patch_size, bias=False)
+
+ scale = width ** -0.5
+ self.class_embedding = nn.Parameter(scale * torch.randn(width))
+ self.positional_embedding = nn.Parameter(scale * torch.randn((input_resolution // patch_size) ** 2 + 1, width))
+ self.ln_pre = LayerNorm(width)
+
+ self.transformer = Transformer(width, layers, heads)
+
+ self.ln_post = LayerNorm(width)
+ self.proj = nn.Parameter(scale * torch.randn(width, output_dim))
+ self.patch_size = patch_size
+
+ def forward(self, x: torch.Tensor, H, W):
+
+ self.positional_embedding_new = upsample_pos_emb(self.positional_embedding, (H//16,W//16))
+ x = self.conv1(x) # shape = [*, width, grid, grid]
+ x = x.reshape(x.shape[0], x.shape[1], -1) # shape = [*, width, grid ** 2]
+ x = x.permute(0, 2, 1) # shape = [*, grid ** 2, width]
+ x = torch.cat([self.class_embedding.to(x.dtype) + torch.zeros(x.shape[0], 1, x.shape[-1], dtype=x.dtype, device=x.device), x], dim=1) # shape = [*, grid ** 2 + 1, width]
+ x = x + self.positional_embedding_new.to(x.dtype)
+ x = self.ln_pre(x)
+
+ x = x.permute(1, 0, 2) # NLD -> LND
+ x, attn_weight = self.transformer(x)
+ '''
+ x = x.permute(1, 0, 2) # LND -> NLD
+
+ x = self.ln_post(x)
+ #x = x[:, 0, :]
+ #x = x[:,1:,:]
+ x = torch.mean(x[:,1:,:],dim=1)
+ #feature_map_list.append(x)
+
+ if self.proj is not None:
+ x = x @ self.proj
+ '''
+
+ return x, attn_weight#cls_attn
+
+
+class CLIP(nn.Module):
+ def __init__(self,
+ embed_dim: int,
+ # vision
+ image_resolution: int,
+ vision_layers: Union[Tuple[int, int, int, int], int],
+ vision_width: int,
+ vision_patch_size: int,
+ # text
+ context_length: int,
+ vocab_size: int,
+ transformer_width: int,
+ transformer_heads: int,
+ transformer_layers: int
+ ):
+ super().__init__()
+
+ self.context_length = context_length
+
+ if isinstance(vision_layers, (tuple, list)):
+ vision_heads = vision_width * 32 // 64
+ self.visual = ModifiedResNet(
+ layers=vision_layers,
+ output_dim=embed_dim,
+ heads=vision_heads,
+ input_resolution=image_resolution,
+ width=vision_width
+ )
+ else:
+ vision_heads = vision_width // 64
+ self.visual = VisionTransformer(
+ input_resolution=image_resolution,
+ patch_size=vision_patch_size,
+ width=vision_width,
+ layers=vision_layers,
+ heads=vision_heads,
+ output_dim=embed_dim
+ )
+
+ self.transformer = Transformer(
+ width=transformer_width,
+ layers=transformer_layers,
+ heads=transformer_heads,
+ attn_mask=self.build_attention_mask()
+ )
+
+ self.vocab_size = vocab_size
+ self.token_embedding = nn.Embedding(vocab_size, transformer_width)
+ self.positional_embedding = nn.Parameter(torch.empty(self.context_length, transformer_width))
+ self.ln_final = LayerNorm(transformer_width)
+
+ self.text_projection = nn.Parameter(torch.empty(transformer_width, embed_dim))
+ self.logit_scale = nn.Parameter(torch.ones([]) * np.log(1 / 0.07))
+
+ self.initialize_parameters()
+
+ def initialize_parameters(self):
+ nn.init.normal_(self.token_embedding.weight, std=0.02)
+ nn.init.normal_(self.positional_embedding, std=0.01)
+
+ if isinstance(self.visual, ModifiedResNet):
+ if self.visual.attnpool is not None:
+ std = self.visual.attnpool.c_proj.in_features ** -0.5
+ nn.init.normal_(self.visual.attnpool.q_proj.weight, std=std)
+ nn.init.normal_(self.visual.attnpool.k_proj.weight, std=std)
+ nn.init.normal_(self.visual.attnpool.v_proj.weight, std=std)
+ nn.init.normal_(self.visual.attnpool.c_proj.weight, std=std)
+
+ for resnet_block in [self.visual.layer1, self.visual.layer2, self.visual.layer3, self.visual.layer4]:
+ for name, param in resnet_block.named_parameters():
+ if name.endswith("bn3.weight"):
+ nn.init.zeros_(param)
+
+ proj_std = (self.transformer.width ** -0.5) * ((2 * self.transformer.layers) ** -0.5)
+ attn_std = self.transformer.width ** -0.5
+ fc_std = (2 * self.transformer.width) ** -0.5
+ for block in self.transformer.resblocks:
+ nn.init.normal_(block.attn.in_proj_weight, std=attn_std)
+ nn.init.normal_(block.attn.out_proj.weight, std=proj_std)
+ nn.init.normal_(block.mlp.c_fc.weight, std=fc_std)
+ nn.init.normal_(block.mlp.c_proj.weight, std=proj_std)
+
+ if self.text_projection is not None:
+ nn.init.normal_(self.text_projection, std=self.transformer.width ** -0.5)
+
+ def build_attention_mask(self):
+ # lazily create causal attention mask, with full attention between the vision tokens
+ # pytorch uses additive attention mask; fill with -inf
+ mask = torch.empty(self.context_length, self.context_length)
+ mask.fill_(float("-inf"))
+ mask.triu_(1) # zero out the lower diagonal
+ return mask
+
+ @property
+ def dtype(self):
+ return self.visual.conv1.weight.dtype
+
+ def encode_image(self, image, H, W):
+ return self.visual(image.type(self.dtype), H, W)
+
+ def encode_text(self, text):
+ x = self.token_embedding(text).type(self.dtype) # [batch_size, n_ctx, d_model]
+
+ x = x + self.positional_embedding.type(self.dtype)
+ x = x.permute(1, 0, 2) # NLD -> LND
+ x, attn_weight = self.transformer(x)
+ x = x.permute(1, 0, 2) # LND -> NLD
+ x = self.ln_final(x).type(self.dtype)
+
+ # x.shape = [batch_size, n_ctx, transformer.width]
+ # take features from the eot embedding (eot_token is the highest number in each sequence)
+ x = x[torch.arange(x.shape[0]), text.argmax(dim=-1)] @ self.text_projection
+
+ return x
+
+ def forward_last_layer(self, image_features, text_features):
+ x, attn_weight = self.visual.transformer.resblocks[self.visual.transformer.layers-1](image_features)
+ x = x.permute(1, 0, 2) # LND -> NLD
+
+ x = self.visual.ln_post(x)
+ x = torch.mean(x[:, 1:, :], dim=1)
+
+ if self.visual.proj is not None:
+ x = x @ self.visual.proj
+
+ image_features = x
+
+ # normalized features
+ image_features = image_features / image_features.norm(dim=1, keepdim=True)
+ text_features = text_features / text_features.norm(dim=1, keepdim=True)
+ # cosine similarity as logits
+ logit_scale = self.logit_scale.exp()
+ logits_per_image = logit_scale * image_features @ text_features.t()
+
+ # shape = [global_batch_size, global_batch_size]
+ logits_per_image = logits_per_image.softmax(dim=-1)
+
+ return logits_per_image, attn_weight
+
+
+
+
+ def forward(self, image, text):
+ image_features, feature_map, cls_attn = self.encode_image(image)
+ with torch.no_grad():
+ text_features = self.encode_text(text)
+
+ # normalized features
+ image_features = image_features / image_features.norm(dim=1, keepdim=True)
+ text_features = text_features / text_features.norm(dim=1, keepdim=True)
+
+ # cosine similarity as logits
+ logit_scale = self.logit_scale.exp()
+ logits_per_image = logit_scale * image_features @ text_features.t()
+ #logits_per_text = logits_per_image.t()
+
+ # shape = [global_batch_size, global_batch_size]
+ return logits_per_image, logits_per_text
+
+
+def convert_weights(model: nn.Module):
+ """Convert applicable model parameters to fp16"""
+
+ def _convert_weights_to_fp16(l):
+ if isinstance(l, (nn.Conv1d, nn.Conv2d, nn.Linear)):
+ l.weight.data = l.weight.data.half()
+ if l.bias is not None:
+ l.bias.data = l.bias.data.half()
+
+ if isinstance(l, nn.MultiheadAttention):
+ for attr in [*[f"{s}_proj_weight" for s in ["in", "q", "k", "v"]], "in_proj_bias", "bias_k", "bias_v"]:
+ tensor = getattr(l, attr)
+ if tensor is not None:
+ tensor.data = tensor.data.half()
+
+ for name in ["text_projection", "proj"]:
+ if hasattr(l, name):
+ attr = getattr(l, name)
+ if attr is not None:
+ attr.data = attr.data.half()
+
+ model.apply(_convert_weights_to_fp16)
+
+
+def build_model(state_dict: dict):
+ vit = "visual.proj" in state_dict
+ '''
+ inv_freq = 1. / (10000 ** (torch.arange(0, 2048, 2, dtype=torch.float) / 2048))
+ position = torch.arange(50, dtype=torch.float)
+ sinusoid_inp = torch.einsum('i,j -> ij', position, inv_freq)
+ embeddings = torch.cat((sinusoid_inp.sin(), sinusoid_inp.cos()), dim=-1) / 2048 ** 0.5
+ state_dict["visual.attnpool.positional_embedding"] = embeddings
+ '''
+ #state_dict["visual.positional_embedding"] = upsample_pos_emb(state_dict["visual.positional_embedding"], 28)
+
+
+ if vit:
+ vision_width = state_dict["visual.conv1.weight"].shape[0]
+ vision_layers = len([k for k in state_dict.keys() if k.startswith("visual.") and k.endswith(".attn.in_proj_weight")])
+ vision_patch_size = state_dict["visual.conv1.weight"].shape[-1]
+ grid_size = round((state_dict["visual.positional_embedding"].shape[0] - 1) ** 0.5)
+ image_resolution = vision_patch_size * grid_size
+ else:
+ counts: list = [len(set(k.split(".")[2] for k in state_dict if k.startswith(f"visual.layer{b}"))) for b in [1, 2, 3, 4]]
+ vision_layers = tuple(counts)
+ vision_width = state_dict["visual.layer1.0.conv1.weight"].shape[0]
+ output_width = round((state_dict["visual.attnpool.positional_embedding"].shape[0] - 1) ** 0.5)
+ vision_patch_size = None
+ assert output_width ** 2 + 1 == state_dict["visual.attnpool.positional_embedding"].shape[0]
+ image_resolution = output_width * 32
+
+ embed_dim = state_dict["text_projection"].shape[1]
+ context_length = state_dict["positional_embedding"].shape[0]
+ vocab_size = state_dict["token_embedding.weight"].shape[0]
+ transformer_width = state_dict["ln_final.weight"].shape[0]
+ transformer_heads = transformer_width // 64
+ transformer_layers = len(set(k.split(".")[2] for k in state_dict if k.startswith(f"transformer.resblocks")))
+
+ model = CLIP(
+ embed_dim,
+ image_resolution, vision_layers, vision_width, vision_patch_size,
+ context_length, vocab_size, transformer_width, transformer_heads, transformer_layers
+ )
+
+ for key in ["input_resolution", "context_length", "vocab_size"]:
+ if key in state_dict:
+ del state_dict[key]
+
+
+
+ convert_weights(model)
+ model.load_state_dict(state_dict)
+ return model.eval()
diff --git a/models/dsp/CAM/clip/simple_tokenizer.py b/models/dsp/CAM/clip/simple_tokenizer.py
new file mode 100644
index 0000000000000000000000000000000000000000..0a66286b7d5019c6e221932a813768038f839c91
--- /dev/null
+++ b/models/dsp/CAM/clip/simple_tokenizer.py
@@ -0,0 +1,132 @@
+import gzip
+import html
+import os
+from functools import lru_cache
+
+import ftfy
+import regex as re
+
+
+@lru_cache()
+def default_bpe():
+ return os.path.join(os.path.dirname(os.path.abspath(__file__)), "bpe_simple_vocab_16e6.txt.gz")
+
+
+@lru_cache()
+def bytes_to_unicode():
+ """
+ Returns list of utf-8 byte and a corresponding list of unicode strings.
+ The reversible bpe codes work on unicode strings.
+ This means you need a large # of unicode characters in your vocab if you want to avoid UNKs.
+ When you're at something like a 10B token dataset you end up needing around 5K for decent coverage.
+ This is a signficant percentage of your normal, say, 32K bpe vocab.
+ To avoid that, we want lookup tables between utf-8 bytes and unicode strings.
+ And avoids mapping to whitespace/control characters the bpe code barfs on.
+ """
+ bs = list(range(ord("!"), ord("~")+1))+list(range(ord("¡"), ord("¬")+1))+list(range(ord("®"), ord("ÿ")+1))
+ cs = bs[:]
+ n = 0
+ for b in range(2**8):
+ if b not in bs:
+ bs.append(b)
+ cs.append(2**8+n)
+ n += 1
+ cs = [chr(n) for n in cs]
+ return dict(zip(bs, cs))
+
+
+def get_pairs(word):
+ """Return set of symbol pairs in a word.
+ Word is represented as tuple of symbols (symbols being variable-length strings).
+ """
+ pairs = set()
+ prev_char = word[0]
+ for char in word[1:]:
+ pairs.add((prev_char, char))
+ prev_char = char
+ return pairs
+
+
+def basic_clean(text):
+ text = ftfy.fix_text(text)
+ text = html.unescape(html.unescape(text))
+ return text.strip()
+
+
+def whitespace_clean(text):
+ text = re.sub(r'\s+', ' ', text)
+ text = text.strip()
+ return text
+
+
+class SimpleTokenizer(object):
+ def __init__(self, bpe_path: str = default_bpe()):
+ self.byte_encoder = bytes_to_unicode()
+ self.byte_decoder = {v: k for k, v in self.byte_encoder.items()}
+ merges = gzip.open(bpe_path).read().decode("utf-8").split('\n')
+ merges = merges[1:49152-256-2+1]
+ merges = [tuple(merge.split()) for merge in merges]
+ vocab = list(bytes_to_unicode().values())
+ vocab = vocab + [v+'' for v in vocab]
+ for merge in merges:
+ vocab.append(''.join(merge))
+ vocab.extend(['<|startoftext|>', '<|endoftext|>'])
+ self.encoder = dict(zip(vocab, range(len(vocab))))
+ self.decoder = {v: k for k, v in self.encoder.items()}
+ self.bpe_ranks = dict(zip(merges, range(len(merges))))
+ self.cache = {'<|startoftext|>': '<|startoftext|>', '<|endoftext|>': '<|endoftext|>'}
+ self.pat = re.compile(r"""<\|startoftext\|>|<\|endoftext\|>|'s|'t|'re|'ve|'m|'ll|'d|[\p{L}]+|[\p{N}]|[^\s\p{L}\p{N}]+""", re.IGNORECASE)
+
+ def bpe(self, token):
+ if token in self.cache:
+ return self.cache[token]
+ word = tuple(token[:-1]) + ( token[-1] + '',)
+ pairs = get_pairs(word)
+
+ if not pairs:
+ return token+''
+
+ while True:
+ bigram = min(pairs, key = lambda pair: self.bpe_ranks.get(pair, float('inf')))
+ if bigram not in self.bpe_ranks:
+ break
+ first, second = bigram
+ new_word = []
+ i = 0
+ while i < len(word):
+ try:
+ j = word.index(first, i)
+ new_word.extend(word[i:j])
+ i = j
+ except:
+ new_word.extend(word[i:])
+ break
+
+ if word[i] == first and i < len(word)-1 and word[i+1] == second:
+ new_word.append(first+second)
+ i += 2
+ else:
+ new_word.append(word[i])
+ i += 1
+ new_word = tuple(new_word)
+ word = new_word
+ if len(word) == 1:
+ break
+ else:
+ pairs = get_pairs(word)
+ word = ' '.join(word)
+ self.cache[token] = word
+ return word
+
+ def encode(self, text):
+ bpe_tokens = []
+ text = whitespace_clean(basic_clean(text)).lower()
+ for token in re.findall(self.pat, text):
+ token = ''.join(self.byte_encoder[b] for b in token.encode('utf-8'))
+ bpe_tokens.extend(self.encoder[bpe_token] for bpe_token in self.bpe(token).split(' '))
+ return bpe_tokens
+
+ def decode(self, tokens):
+ text = ''.join([self.decoder[token] for token in tokens])
+ text = bytearray([self.byte_decoder[c] for c in text]).decode('utf-8', errors="replace").replace('', ' ')
+ return text
diff --git a/models/dsp/CAM/pytorch_grad_cam/__init__.py b/models/dsp/CAM/pytorch_grad_cam/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..7cf7c8cd8f4d4178100cc5203462620dcf057777
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/__init__.py
@@ -0,0 +1,14 @@
+from .grad_cam import GradCAM
+from .ablation_layer import AblationLayer, AblationLayerVit, AblationLayerFasterRCNN
+from .ablation_cam import AblationCAM
+from .xgrad_cam import XGradCAM
+from .grad_cam_plusplus import GradCAMPlusPlus
+from .score_cam import ScoreCAM
+from .layer_cam import LayerCAM
+from .eigen_cam import EigenCAM
+from .eigen_grad_cam import EigenGradCAM
+from .fullgrad_cam import FullGrad
+from .guided_backprop import GuidedBackpropReLUModel
+from .activations_and_gradients import ActivationsAndGradients
+from .utils import model_targets
+from .utils import reshape_transforms
\ No newline at end of file
diff --git a/models/dsp/CAM/pytorch_grad_cam/ablation_cam.py b/models/dsp/CAM/pytorch_grad_cam/ablation_cam.py
new file mode 100644
index 0000000000000000000000000000000000000000..a8b95904b93b6e9a10baa1688f048c9cd6dc4678
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/ablation_cam.py
@@ -0,0 +1,134 @@
+import numpy as np
+import torch
+import tqdm
+from typing import Callable, List
+from .base_cam import BaseCAM
+from .utils.find_layers import replace_layer_recursive
+from .ablation_layer import AblationLayer
+
+
+""" Implementation of AblationCAM
+https://openaccess.thecvf.com/content_WACV_2020/papers/Desai_Ablation-CAM_Visual_Explanations_for_Deep_Convolutional_Network_via_Gradient-free_Localization_WACV_2020_paper.pdf
+
+Ablate individual activations, and then measure the drop in the target score.
+
+In the current implementation, the target layer activations is cached, so it won't be re-computed.
+However layers before it, if any, will not be cached.
+This means that if the target layer is a large block, for example model.featuers (in vgg), there will
+be a large save in run time.
+
+Since we have to go over many channels and ablate them, and every channel ablation requires a forward pass,
+it would be nice if we could avoid doing that for channels that won't contribute anwyay, making it much faster.
+The parameter ratio_channels_to_ablate controls how many channels should be ablated, using an experimental method
+(to be improved). The default 1.0 value means that all channels will be ablated.
+"""
+
+
+class AblationCAM(BaseCAM):
+ def __init__(self,
+ model: torch.nn.Module,
+ target_layers: List[torch.nn.Module],
+ use_cuda: bool = False,
+ reshape_transform: Callable = None,
+ ablation_layer: torch.nn.Module = AblationLayer(),
+ batch_size: int = 32,
+ ratio_channels_to_ablate: float = 1.0) -> None:
+
+ super(AblationCAM, self).__init__(model,
+ target_layers,
+ use_cuda,
+ reshape_transform,
+ uses_gradients=False)
+ self.batch_size = batch_size
+ self.ablation_layer = ablation_layer
+ self.ratio_channels_to_ablate = ratio_channels_to_ablate
+
+ def save_activation(self, module, input, output) -> None:
+ """ Helper function to save the raw activations from the target layer """
+ self.activations = output
+
+ def assemble_ablation_scores(self,
+ new_scores: list,
+ original_score: float ,
+ ablated_channels: np.ndarray,
+ number_of_channels: int) -> np.ndarray:
+ """ Take the value from the channels that were ablated,
+ and just set the original score for the channels that were skipped """
+
+ index = 0
+ result = []
+ sorted_indices = np.argsort(ablated_channels)
+ ablated_channels = ablated_channels[sorted_indices]
+ new_scores = np.float32(new_scores)[sorted_indices]
+
+ for i in range(number_of_channels):
+ if index < len(ablated_channels) and ablated_channels[index] == i:
+ weight = new_scores[index]
+ index = index + 1
+ else:
+ weight = original_score
+ result.append(weight)
+
+ return result
+
+ def get_cam_weights(self,
+ input_tensor: torch.Tensor,
+ target_layer: torch.nn.Module,
+ targets: List[Callable],
+ activations: torch.Tensor,
+ grads: torch.Tensor) -> np.ndarray:
+
+ # Do a forward pass, compute the target scores, and cache the activations
+ handle = target_layer.register_forward_hook(self.save_activation)
+ with torch.no_grad():
+ outputs = self.model(input_tensor)
+ handle.remove()
+ original_scores = np.float32([target(output).cpu().item() for target, output in zip(targets, outputs)])
+
+ # Replace the layer with the ablation layer.
+ # When we finish, we will replace it back, so the original model is unchanged.
+ ablation_layer = self.ablation_layer
+ replace_layer_recursive(self.model, target_layer, ablation_layer)
+
+ number_of_channels = activations.shape[1]
+ weights = []
+ # This is a "gradient free" method, so we don't need gradients here.
+ with torch.no_grad():
+ # Loop over each of the batch images and ablate activations for it.
+ for batch_index, (target, tensor) in enumerate(zip(targets, input_tensor)):
+ new_scores = []
+ batch_tensor = tensor.repeat(self.batch_size, 1, 1, 1)
+
+ # Check which channels should be ablated. Normally this will be all channels,
+ # But we can also try to speed this up by using a low ratio_channels_to_ablate.
+ channels_to_ablate = ablation_layer.activations_to_be_ablated(activations[batch_index, :],
+ self.ratio_channels_to_ablate)
+ number_channels_to_ablate = len(channels_to_ablate)
+
+ for i in tqdm.tqdm(range(0, number_channels_to_ablate, self.batch_size)):
+ if i + self.batch_size > number_channels_to_ablate:
+ batch_tensor = batch_tensor[:(number_channels_to_ablate - i)]
+
+ # Change the state of the ablation layer so it ablates the next channels.
+ # TBD: Move this into the ablation layer forward pass.
+ ablation_layer.set_next_batch(input_batch_index=batch_index,
+ activations=self.activations,
+ num_channels_to_ablate=batch_tensor.size(0))
+ score = [target(o).cpu().item() for o in self.model(batch_tensor)]
+ new_scores.extend(score)
+ ablation_layer.indices = ablation_layer.indices[batch_tensor.size(0):]
+
+ new_scores = self.assemble_ablation_scores(new_scores,
+ original_scores[batch_index],
+ channels_to_ablate,
+ number_of_channels)
+ weights.extend(new_scores)
+
+ weights = np.float32(weights)
+ weights = weights.reshape(activations.shape[:2])
+ original_scores = original_scores[:, None]
+ weights = (original_scores - weights) / original_scores
+
+ # Replace the model back to the original state
+ replace_layer_recursive(self.model, ablation_layer, target_layer)
+ return weights
diff --git a/models/dsp/CAM/pytorch_grad_cam/ablation_cam_multilayer.py b/models/dsp/CAM/pytorch_grad_cam/ablation_cam_multilayer.py
new file mode 100644
index 0000000000000000000000000000000000000000..d95f35f0f989867711ba0453beafdaf6c5973d16
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/ablation_cam_multilayer.py
@@ -0,0 +1,136 @@
+import cv2
+import numpy as np
+import torch
+import tqdm
+from .base_cam import BaseCAM
+
+
+class AblationLayer(torch.nn.Module):
+ def __init__(self, layer, reshape_transform, indices):
+ super(AblationLayer, self).__init__()
+
+ self.layer = layer
+ self.reshape_transform = reshape_transform
+ # The channels to zero out:
+ self.indices = indices
+
+ def forward(self, x):
+ self.__call__(x)
+
+ def __call__(self, x):
+ output = self.layer(x)
+
+ # Hack to work with ViT,
+ # Since the activation channels are last and not first like in CNNs
+ # Probably should remove it?
+ if self.reshape_transform is not None:
+ output = output.transpose(1, 2)
+
+ for i in range(output.size(0)):
+
+ # Commonly the minimum activation will be 0,
+ # And then it makes sense to zero it out.
+ # However depending on the architecture,
+ # If the values can be negative, we use very negative values
+ # to perform the ablation, deviating from the paper.
+ if torch.min(output) == 0:
+ output[i, self.indices[i], :] = 0
+ else:
+ ABLATION_VALUE = 1e5
+ output[i, self.indices[i], :] = torch.min(
+ output) - ABLATION_VALUE
+
+ if self.reshape_transform is not None:
+ output = output.transpose(2, 1)
+
+ return output
+
+
+def replace_layer_recursive(model, old_layer, new_layer):
+ for name, layer in model._modules.items():
+ if layer == old_layer:
+ model._modules[name] = new_layer
+ return True
+ elif replace_layer_recursive(layer, old_layer, new_layer):
+ return True
+ return False
+
+
+class AblationCAM(BaseCAM):
+ def __init__(self, model, target_layers, use_cuda=False,
+ reshape_transform=None):
+ super(AblationCAM, self).__init__(model, target_layers, use_cuda,
+ reshape_transform)
+
+ if len(target_layers) > 1:
+ print(
+ "Warning. You are usign Ablation CAM with more than 1 layers. "
+ "This is supported only if all layers have the same output shape")
+
+ def set_ablation_layers(self):
+ self.ablation_layers = []
+ for target_layer in self.target_layers:
+ ablation_layer = AblationLayer(target_layer,
+ self.reshape_transform, indices=[])
+ self.ablation_layers.append(ablation_layer)
+ replace_layer_recursive(self.model, target_layer, ablation_layer)
+
+ def unset_ablation_layers(self):
+ # replace the model back to the original state
+ for ablation_layer, target_layer in zip(
+ self.ablation_layers, self.target_layers):
+ replace_layer_recursive(self.model, ablation_layer, target_layer)
+
+ def set_ablation_layer_batch_indices(self, indices):
+ for ablation_layer in self.ablation_layers:
+ ablation_layer.indices = indices
+
+ def trim_ablation_layer_batch_indices(self, keep):
+ for ablation_layer in self.ablation_layers:
+ ablation_layer.indices = ablation_layer.indices[:keep]
+
+ def get_cam_weights(self,
+ input_tensor,
+ target_category,
+ activations,
+ grads):
+ with torch.no_grad():
+ outputs = self.model(input_tensor).cpu().numpy()
+ original_scores = []
+ for i in range(input_tensor.size(0)):
+ original_scores.append(outputs[i, target_category[i]])
+ original_scores = np.float32(original_scores)
+
+ self.set_ablation_layers()
+
+ if hasattr(self, "batch_size"):
+ BATCH_SIZE = self.batch_size
+ else:
+ BATCH_SIZE = 32
+
+ number_of_channels = activations.shape[1]
+ weights = []
+
+ with torch.no_grad():
+ # Iterate over the input batch
+ for tensor, category in zip(input_tensor, target_category):
+ batch_tensor = tensor.repeat(BATCH_SIZE, 1, 1, 1)
+ for i in tqdm.tqdm(range(0, number_of_channels, BATCH_SIZE)):
+ self.set_ablation_layer_batch_indices(
+ list(range(i, i + BATCH_SIZE)))
+
+ if i + BATCH_SIZE > number_of_channels:
+ keep = number_of_channels - i
+ batch_tensor = batch_tensor[:keep]
+ self.trim_ablation_layer_batch_indices(self, keep)
+ score = self.model(batch_tensor)[:, category].cpu().numpy()
+ weights.extend(score)
+
+ weights = np.float32(weights)
+ weights = weights.reshape(activations.shape[:2])
+ original_scores = original_scores[:, None]
+ weights = (original_scores - weights) / original_scores
+
+ # replace the model back to the original state
+ self.unset_ablation_layers()
+ return weights
diff --git a/models/dsp/CAM/pytorch_grad_cam/ablation_layer.py b/models/dsp/CAM/pytorch_grad_cam/ablation_layer.py
new file mode 100644
index 0000000000000000000000000000000000000000..18416a88ccdcb8ea4b8ac873e8ca7b52ead53572
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/ablation_layer.py
@@ -0,0 +1,131 @@
+import torch
+from collections import OrderedDict
+import numpy as np
+from .utils.svd_on_activations import get_2d_projection
+
+
+class AblationLayer(torch.nn.Module):
+ def __init__(self):
+ super(AblationLayer, self).__init__()
+
+ def objectiveness_mask_from_svd(self, activations, threshold=0.01):
+ """ Experimental method to get a binary mask to compare if the activation is worth ablating.
+ The idea is to apply the EigenCAM method by doing PCA on the activations.
+ Then we create a binary mask by comparing to a low threshold.
+ Areas that are masked out, are probably not interesting anyway.
+ """
+
+ projection = get_2d_projection(activations[None, :])[0, :]
+ projection = np.abs(projection)
+ projection = projection - projection.min()
+ projection = projection / projection.max()
+ projection = projection > threshold
+ return projection
+
+ def activations_to_be_ablated(self, activations, ratio_channels_to_ablate=1.0):
+ """ Experimental method to get a binary mask to compare if the activation is worth ablating.
+ Create a binary CAM mask with objectiveness_mask_from_svd.
+ Score each Activation channel, by seeing how much of its values are inside the mask.
+ Then keep the top channels.
+
+ """
+ if ratio_channels_to_ablate == 1.0:
+ self.indices = np.int32(range(activations.shape[0]))
+ return self.indices
+
+ projection = self.objectiveness_mask_from_svd(activations)
+
+ scores = []
+ for channel in activations:
+ normalized = np.abs(channel)
+ normalized = normalized - normalized.min()
+ normalized = normalized / np.max(normalized)
+ score = (projection*normalized).sum() / normalized.sum()
+ scores.append(score)
+ scores = np.float32(scores)
+
+ indices = list(np.argsort(scores))
+ high_score_indices = indices[::-1][: int(len(indices) * ratio_channels_to_ablate)]
+ low_score_indices = indices[: int(len(indices) * ratio_channels_to_ablate)]
+ self.indices = np.int32(high_score_indices + low_score_indices)
+ return self.indices
+
+ def set_next_batch(self, input_batch_index, activations, num_channels_to_ablate):
+ """ This creates the next batch of activations from the layer.
+ Just take corresponding batch member from activations, and repeat it num_channels_to_ablate times.
+ """
+ self.activations = activations[input_batch_index, :, :, :].clone().unsqueeze(0).repeat(num_channels_to_ablate, 1, 1, 1)
+
+ def __call__(self, x):
+ output = self.activations
+ for i in range(output.size(0)):
+ # Commonly the minimum activation will be 0,
+ # And then it makes sense to zero it out.
+ # However depending on the architecture,
+ # If the values can be negative, we use very negative values
+ # to perform the ablation, deviating from the paper.
+ if torch.min(output) == 0:
+ output[i, self.indices[i], :] = 0
+ else:
+ ABLATION_VALUE = 1e7
+ output[i, self.indices[i], :] = torch.min(
+ output) - ABLATION_VALUE
+
+ return output
+
+
+class AblationLayerVit(AblationLayer):
+ def __init__(self):
+ super(AblationLayerVit, self).__init__()
+
+ def __call__(self, x):
+ output = self.activations
+ output = output.transpose(1, 2)
+ for i in range(output.size(0)):
+
+ # Commonly the minimum activation will be 0,
+ # And then it makes sense to zero it out.
+ # However depending on the architecture,
+ # If the values can be negative, we use very negative values
+ # to perform the ablation, deviating from the paper.
+ if torch.min(output) == 0:
+ output[i, self.indices[i], :] = 0
+ else:
+ ABLATION_VALUE = 1e7
+ output[i, self.indices[i], :] = torch.min(
+ output) - ABLATION_VALUE
+
+ output = output.transpose(2, 1)
+
+ return output
+
+ def set_next_batch(self, input_batch_index, activations, num_channels_to_ablate):
+ """ This creates the next batch of activations from the layer.
+ Just take corresponding batch member from activations, and repeat it num_channels_to_ablate times.
+ """
+ self.activations = activations[input_batch_index, :, :].clone().unsqueeze(0).repeat(num_channels_to_ablate, 1, 1)
+
+
+
+class AblationLayerFasterRCNN(AblationLayer):
+ def __init__(self):
+ super(AblationLayerFasterRCNN, self).__init__()
+
+ def set_next_batch(self, input_batch_index, activations, num_channels_to_ablate):
+ """ Extract the next batch member from activations,
+ and repeat it num_channels_to_ablate times.
+ """
+ self.activations = OrderedDict()
+ for key, value in activations.items():
+ fpn_activation = value[input_batch_index, :, :, :].clone().unsqueeze(0)
+ self.activations[key] = fpn_activation.repeat(num_channels_to_ablate, 1, 1, 1)
+
+ def __call__(self, x):
+ result = self.activations
+ layers = {0: '0', 1: '1', 2: '2', 3: '3', 4: 'pool'}
+ num_channels_to_ablate = result['pool'].size(0)
+ for i in range(num_channels_to_ablate):
+ pyramid_layer = int(self.indices[i]/256)
+ index_in_pyramid_layer = int(self.indices[i] % 256)
+ result[layers[pyramid_layer]][i, index_in_pyramid_layer, :, :] = -1000
+ return result
diff --git a/models/dsp/CAM/pytorch_grad_cam/activations_and_gradients.py b/models/dsp/CAM/pytorch_grad_cam/activations_and_gradients.py
new file mode 100644
index 0000000000000000000000000000000000000000..605fc3e093dec4693dc3386592499d70173aee1a
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/activations_and_gradients.py
@@ -0,0 +1,55 @@
+class ActivationsAndGradients:
+ """ Class for extracting activations and
+ registering gradients from targetted intermediate layers """
+
+ def __init__(self, model, target_layers, reshape_transform):
+ self.model = model
+ self.gradients = []
+ self.activations = []
+ self.reshape_transform = reshape_transform
+ self.handles = []
+ for target_layer in target_layers:
+ self.handles.append(
+ target_layer.register_forward_hook(self.save_activation))
+ # Because of https://github.com/pytorch/pytorch/issues/61519,
+ # we don't use backward hook to record gradients.
+ self.handles.append(
+ target_layer.register_forward_hook(self.save_gradient))
+
+ def save_activation(self, module, input, output):
+ activation = output
+
+ if self.reshape_transform is not None:
+ activation = self.reshape_transform(activation, self.height, self.width)
+ # self.activations.append(activation.cpu().detach())
+ # self.activations.append(activation.detach()) # original detach
+ self.activations.append(activation)
+
+ def save_gradient(self, module, input, output):
+ if not hasattr(output, "requires_grad") or not output.requires_grad:
+ # You can only register hooks on tensor requires grad.
+ return
+
+ # Gradients are computed in reverse order
+ def _store_grad(grad):
+ if self.reshape_transform is not None:
+ grad = self.reshape_transform(grad, self.height, self.width)
+ # self.gradients = [grad.cpu().detach()] + self.gradients
+ # self.gradients = [grad.detach()] + self.gradients # original detach
+ self.gradients = [grad] + self.gradients
+
+ output.register_hook(_store_grad)
+
+ def __call__(self, x, H, W):
+ self.height = H // 16
+ self.width = W // 16
+ self.gradients = []
+ self.activations = []
+ if isinstance(x, list):
+ return self.model.forward_last_layer(x[0], x[1])
+ else:
+ return self.model(x)
+
+ def release(self):
+ for handle in self.handles:
+ handle.remove()
diff --git a/models/dsp/CAM/pytorch_grad_cam/base_cam.py b/models/dsp/CAM/pytorch_grad_cam/base_cam.py
new file mode 100644
index 0000000000000000000000000000000000000000..be3ac2c1605ae43af46fff8dc7041e65e897d8c4
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/base_cam.py
@@ -0,0 +1,227 @@
+import numpy as np
+import torch
+import ttach as tta
+from typing import Callable, List, Tuple
+from .activations_and_gradients import ActivationsAndGradients
+from .utils.svd_on_activations import get_2d_projection
+from .utils.image import scale_cam_image, scale_cam_image_torch
+from .utils.model_targets import ClassifierOutputTarget
+
+
+class BaseCAM:
+ def __init__(self,
+ model: torch.nn.Module,
+ target_layers: List[torch.nn.Module],
+ use_cuda: bool = False,
+ reshape_transform: Callable = None,
+ compute_input_gradient: bool = False,
+ uses_gradients: bool = True) -> None:
+ self.model = model.eval()
+ self.target_layers = target_layers
+ self.cuda = use_cuda
+ # if self.cuda:
+ # self.model = model.cuda()
+ self.reshape_transform = reshape_transform
+ self.compute_input_gradient = compute_input_gradient
+ self.uses_gradients = uses_gradients
+ self.activations_and_grads = ActivationsAndGradients(
+ self.model, target_layers, reshape_transform)
+
+ """ Get a vector of weights for every channel in the target layer.
+ Methods that return weights channels,
+ will typically need to only implement this function. """
+
+ def set_device(self, device):
+ self.device = device
+
+ def get_cam_weights(self,
+ input_tensor: torch.Tensor,
+ target_layers: List[torch.nn.Module],
+ targets: List[torch.nn.Module],
+ activations: torch.Tensor,
+ grads: torch.Tensor) -> np.ndarray:
+ raise Exception("Not Implemented")
+
+ def get_cam_image(self,
+ input_tensor: torch.Tensor,
+ target_layer: torch.nn.Module,
+ targets: List[torch.nn.Module],
+ activations: torch.Tensor,
+ grads: torch.Tensor,
+ eigen_smooth: bool = False) -> np.ndarray:
+
+ weights = self.get_cam_weights(input_tensor,
+ target_layer,
+ targets,
+ activations,
+ grads)
+ weighted_activations = weights[:, :, None, None] * activations
+ if eigen_smooth:
+ cam = get_2d_projection(weighted_activations)
+ else:
+ cam = weighted_activations.sum(axis=1)
+ return cam
+
+ def forward(self,
+ input_tensor: torch.Tensor,
+ targets: List[torch.nn.Module],
+ target_size,
+ eigen_smooth: bool = False) -> np.ndarray:
+
+ if self.cuda:
+ if isinstance(input_tensor, list):
+ input_tensor = [t.to(self.device) if isinstance(t, torch.Tensor) else t for t in input_tensor]
+ else:
+ input_tensor = input_tensor.to(self.device)
+
+ if self.compute_input_gradient:
+ input_tensor = torch.autograd.Variable(input_tensor,
+ requires_grad=True)
+
+ W,H = self.get_target_width_height(input_tensor)
+ outputs = self.activations_and_grads(input_tensor,H,W)
+ if targets is None:
+ if isinstance(input_tensor, list):
+ target_categories = np.argmax(outputs[0].cpu().data.numpy(), axis=-1)
+ else:
+ target_categories = np.argmax(outputs.cpu().data.numpy(), axis=-1)
+ targets = [ClassifierOutputTarget(category) for category in target_categories]
+
+ if self.uses_gradients:
+ self.model.zero_grad()
+ if isinstance(input_tensor, list):
+ loss = sum([target(output[0]) for target, output in zip(targets, outputs)])
+ else:
+ loss = sum([target(output) for target, output in zip(targets, outputs)])
+ loss.backward(retain_graph=True)
+
+ # In most of the saliency attribution papers, the saliency is
+ # computed with a single target layer.
+ # Commonly it is the last convolutional layer.
+ # Here we support passing a list with multiple target layers.
+ # It will compute the saliency image for every image,
+ # and then aggregate them (with a default mean aggregation).
+ # This gives you more flexibility in case you just want to
+ # use all conv layers for example, all Batchnorm layers,
+ # or something else.
+ cam_per_layer = self.compute_cam_per_layer(input_tensor,
+ targets,
+ target_size,
+ eigen_smooth)
+ if isinstance(input_tensor, list):
+ return self.aggregate_multi_layers_torch(cam_per_layer), outputs[0], outputs[1]
+ else:
+ return self.aggregate_multi_layers(cam_per_layer), outputs
+
+ def get_target_width_height(self,
+ input_tensor: torch.Tensor) -> Tuple[int, int]:
+ if isinstance(input_tensor, list):
+ width, height = input_tensor[-1], input_tensor[-2]
+ return width, height
+
+ def compute_cam_per_layer(
+ self,
+ input_tensor: torch.Tensor,
+ targets: List[torch.nn.Module],
+ target_size,
+ eigen_smooth: bool) -> np.ndarray:
+ activations_list = [a#.cpu().data.numpy()
+ for a in self.activations_and_grads.activations]
+ grads_list = [g#.cpu().data.numpy()
+ for g in self.activations_and_grads.gradients]
+
+ cam_per_target_layer = []
+ # Loop over the saliency image from every layer
+ for i in range(len(self.target_layers)):
+ target_layer = self.target_layers[i]
+ layer_activations = None
+ layer_grads = None
+ if i < len(activations_list):
+ layer_activations = activations_list[i]
+ if i < len(grads_list):
+ layer_grads = grads_list[i]
+
+ cam = self.get_cam_image(input_tensor,
+ target_layer,
+ targets,
+ layer_activations,
+ layer_grads,
+ eigen_smooth)
+ cam = torch.clamp(cam, min=0).float()
+ scaled = scale_cam_image_torch(cam, target_size)
+ # cam = np.maximum(cam, 0).astype(np.float32)#float16->32
+ # scaled = scale_cam_image(cam, target_size)
+ cam_per_target_layer.append(scaled[:, None, :])
+
+ return cam_per_target_layer
+
+ def aggregate_multi_layers(self, cam_per_target_layer: np.ndarray) -> np.ndarray:
+ cam_per_target_layer = np.concatenate(cam_per_target_layer, axis=1)
+ cam_per_target_layer = np.maximum(cam_per_target_layer, 0)
+ result = np.mean(cam_per_target_layer, axis=1)
+ return scale_cam_image(result)
+
+ def aggregate_multi_layers_torch(self, cam_per_target_layer):
+ cam_per_target_layer = torch.cat(cam_per_target_layer, dim=1)
+ cam_per_target_layer = torch.clamp(cam_per_target_layer, min=0)
+ result = cam_per_target_layer.mean(dim=1, keepdim=True)
+ return scale_cam_image_torch(result)
+
+ def forward_augmentation_smoothing(self,
+ input_tensor: torch.Tensor,
+ targets: List[torch.nn.Module],
+ eigen_smooth: bool = False) -> np.ndarray:
+ transforms = tta.Compose(
+ [
+ tta.HorizontalFlip(),
+ tta.Multiply(factors=[0.9, 1, 1.1]),
+ ]
+ )
+ cams = []
+ for transform in transforms:
+ augmented_tensor = transform.augment_image(input_tensor)
+ cam = self.forward(augmented_tensor,
+ targets,
+ eigen_smooth)
+
+ # The ttach library expects a tensor of size BxCxHxW
+ cam = cam[:, None, :, :]
+ cam = torch.from_numpy(cam)
+ cam = transform.deaugment_mask(cam)
+
+ # Back to numpy float32, HxW
+ cam = cam.numpy()
+ cam = cam[:, 0, :, :]
+ cams.append(cam)
+
+ cam = np.mean(np.float32(cams), axis=0)
+ return cam
+
+ def __call__(self,
+ input_tensor: torch.Tensor,
+ targets: List[torch.nn.Module] = None,
+ target_size=None,
+ aug_smooth: bool = False,
+ eigen_smooth: bool = False) -> np.ndarray:
+
+ # Smooth the CAM result with test time augmentation
+ if aug_smooth is True:
+ return self.forward_augmentation_smoothing(
+ input_tensor, targets, eigen_smooth)
+
+ return self.forward(input_tensor,
+ targets, target_size,eigen_smooth)
+
+ def __del__(self):
+ self.activations_and_grads.release()
+
+ def __enter__(self):
+ return self
+
+ def __exit__(self, exc_type, exc_value, exc_tb):
+ self.activations_and_grads.release()
+ if isinstance(exc_value, IndexError):
+ # Handle IndexError here...
+ print(
+ f"An exception occurred in CAM with block: {exc_type}. Message: {exc_value}")
+ return True
diff --git a/models/dsp/CAM/pytorch_grad_cam/eigen_cam.py b/models/dsp/CAM/pytorch_grad_cam/eigen_cam.py
new file mode 100644
index 0000000000000000000000000000000000000000..791dec2874f43212c9d65a0f09915180e987c504
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/eigen_cam.py
@@ -0,0 +1,23 @@
+from .base_cam import BaseCAM
+from .utils.svd_on_activations import get_2d_projection
+
+# https://arxiv.org/abs/2008.00299
+
+
+class EigenCAM(BaseCAM):
+ def __init__(self, model, target_layers, use_cuda=False,
+ reshape_transform=None):
+ super(EigenCAM, self).__init__(model,
+ target_layers,
+ use_cuda,
+ reshape_transform,
+ uses_gradients=False)
+
+ def get_cam_image(self,
+ input_tensor,
+ target_layer,
+ target_category,
+ activations,
+ grads,
+ eigen_smooth):
+ return get_2d_projection(activations)
diff --git a/models/dsp/CAM/pytorch_grad_cam/eigen_grad_cam.py b/models/dsp/CAM/pytorch_grad_cam/eigen_grad_cam.py
new file mode 100644
index 0000000000000000000000000000000000000000..daa80fe7bb9f3fa661497b101cf77d1f9437b31a
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/eigen_grad_cam.py
@@ -0,0 +1,21 @@
+from .base_cam import BaseCAM
+from .utils.svd_on_activations import get_2d_projection
+
+# Like Eigen CAM: https://arxiv.org/abs/2008.00299
+# But multiply the activations x gradients
+
+
+class EigenGradCAM(BaseCAM):
+ def __init__(self, model, target_layers, use_cuda=False,
+ reshape_transform=None):
+ super(EigenGradCAM, self).__init__(model, target_layers, use_cuda,
+ reshape_transform)
+
+ def get_cam_image(self,
+ input_tensor,
+ target_layer,
+ target_category,
+ activations,
+ grads,
+ eigen_smooth):
+ return get_2d_projection(grads * activations)
diff --git a/models/dsp/CAM/pytorch_grad_cam/fullgrad_cam.py b/models/dsp/CAM/pytorch_grad_cam/fullgrad_cam.py
new file mode 100644
index 0000000000000000000000000000000000000000..df71511486382eff9ab80ee01596ab412e64451d
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/fullgrad_cam.py
@@ -0,0 +1,95 @@
+import numpy as np
+import torch
+from .base_cam import BaseCAM
+from .utils.find_layers import find_layer_predicate_recursive
+from .utils.svd_on_activations import get_2d_projection
+from .utils.image import scale_accross_batch_and_channels, scale_cam_image
+
+# https://arxiv.org/abs/1905.00780
+
+
+class FullGrad(BaseCAM):
+ def __init__(self, model, target_layers, use_cuda=False,
+ reshape_transform=None):
+ if len(target_layers) > 0:
+ print(
+ "Warning: target_layers is ignored in FullGrad. All bias layers will be used instead")
+
+ def layer_with_2D_bias(layer):
+ bias_target_layers = [torch.nn.Conv2d, torch.nn.BatchNorm2d]
+ if type(layer) in bias_target_layers and layer.bias is not None:
+ return True
+ return False
+ target_layers = find_layer_predicate_recursive(
+ model, layer_with_2D_bias)
+ super(
+ FullGrad,
+ self).__init__(
+ model,
+ target_layers,
+ use_cuda,
+ reshape_transform,
+ compute_input_gradient=True)
+ self.bias_data = [self.get_bias_data(
+ layer).cpu().numpy() for layer in target_layers]
+
+ def get_bias_data(self, layer):
+ # Borrowed from official paper impl:
+ # https://github.com/idiap/fullgrad-saliency/blob/master/saliency/tensor_extractor.py#L47
+ if isinstance(layer, torch.nn.BatchNorm2d):
+ bias = - (layer.running_mean * layer.weight
+ / torch.sqrt(layer.running_var + layer.eps)) + layer.bias
+ return bias.data
+ else:
+ return layer.bias.data
+
+ def compute_cam_per_layer(
+ self,
+ input_tensor,
+ target_category,
+ eigen_smooth):
+ input_grad = input_tensor.grad.data.cpu().numpy()
+ grads_list = [g.cpu().data.numpy() for g in
+ self.activations_and_grads.gradients]
+ cam_per_target_layer = []
+ target_size = self.get_target_width_height(input_tensor)
+
+ gradient_multiplied_input = input_grad * input_tensor.data.cpu().numpy()
+ gradient_multiplied_input = np.abs(gradient_multiplied_input)
+ gradient_multiplied_input = scale_accross_batch_and_channels(
+ gradient_multiplied_input,
+ target_size)
+ cam_per_target_layer.append(gradient_multiplied_input)
+
+ # Loop over the saliency image from every layer
+ assert(len(self.bias_data) == len(grads_list))
+ for bias, grads in zip(self.bias_data, grads_list):
+ bias = bias[None, :, None, None]
+ # In the paper they take the absolute value,
+ # but possibily taking only the positive gradients will work
+ # better.
+ bias_grad = np.abs(bias * grads)
+ result = scale_accross_batch_and_channels(
+ bias_grad, target_size)
+ result = np.sum(result, axis=1)
+ cam_per_target_layer.append(result[:, None, :])
+ cam_per_target_layer = np.concatenate(cam_per_target_layer, axis=1)
+ if eigen_smooth:
+ # Resize to a smaller image, since this method typically has a very large number of channels,
+ # and then consumes a lot of memory
+ cam_per_target_layer = scale_accross_batch_and_channels(
+ cam_per_target_layer, (target_size[0] // 8, target_size[1] // 8))
+ cam_per_target_layer = get_2d_projection(cam_per_target_layer)
+ cam_per_target_layer = cam_per_target_layer[:, None, :, :]
+ cam_per_target_layer = scale_accross_batch_and_channels(
+ cam_per_target_layer,
+ target_size)
+ else:
+ cam_per_target_layer = np.sum(
+ cam_per_target_layer, axis=1)[:, None, :]
+
+ return cam_per_target_layer
+
+ def aggregate_multi_layers(self, cam_per_target_layer):
+ result = np.sum(cam_per_target_layer, axis=1)
+ return scale_cam_image(result)
diff --git a/models/dsp/CAM/pytorch_grad_cam/grad_cam.py b/models/dsp/CAM/pytorch_grad_cam/grad_cam.py
new file mode 100644
index 0000000000000000000000000000000000000000..6ae10732b144c67d470da7de4afbe0cee295f13f
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/grad_cam.py
@@ -0,0 +1,24 @@
+import torch
+import numpy as np
+from .base_cam import BaseCAM
+
+
+class GradCAM(BaseCAM):
+ def __init__(self, model, target_layers, use_cuda=False,
+ reshape_transform=None):
+ super(
+ GradCAM,
+ self).__init__(
+ model,
+ target_layers,
+ use_cuda,
+ reshape_transform)
+
+ def get_cam_weights(self,
+ input_tensor,
+ target_layer,
+ target_category,
+ activations,
+ grads):
+ # return np.mean(grads, axis=(2, 3))
+ return torch.mean(grads, dim=(2, 3))
diff --git a/models/dsp/CAM/pytorch_grad_cam/grad_cam_plusplus.py b/models/dsp/CAM/pytorch_grad_cam/grad_cam_plusplus.py
new file mode 100644
index 0000000000000000000000000000000000000000..4c0a791540c52aa7c8e01e3044286fee2c188a6e
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/grad_cam_plusplus.py
@@ -0,0 +1,32 @@
+import numpy as np
+from .base_cam import BaseCAM
+
+# https://arxiv.org/abs/1710.11063
+
+
+class GradCAMPlusPlus(BaseCAM):
+ def __init__(self, model, target_layers, use_cuda=False,
+ reshape_transform=None):
+ super(GradCAMPlusPlus, self).__init__(model, target_layers, use_cuda,
+ reshape_transform)
+
+ def get_cam_weights(self,
+ input_tensor,
+ target_layers,
+ target_category,
+ activations,
+ grads):
+ grads_power_2 = grads**2
+ grads_power_3 = grads_power_2 * grads
+ # Equation 19 in https://arxiv.org/abs/1710.11063
+ sum_activations = np.sum(activations, axis=(2, 3))
+ eps = 0.000001
+ aij = grads_power_2 / (2 * grads_power_2 +
+ sum_activations[:, :, None, None] * grads_power_3 + eps)
+ # Now bring back the ReLU from eq.7 in the paper,
+ # And zero out aijs where the activations are 0
+ aij = np.where(grads != 0, aij, 0)
+
+ weights = np.maximum(grads, 0) * aij
+ weights = np.sum(weights, axis=(2, 3))
+ return weights
diff --git a/models/dsp/CAM/pytorch_grad_cam/guided_backprop.py b/models/dsp/CAM/pytorch_grad_cam/guided_backprop.py
new file mode 100644
index 0000000000000000000000000000000000000000..80eac1a5202567c63994585a08f40a0de5695efd
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/guided_backprop.py
@@ -0,0 +1,100 @@
+import numpy as np
+import torch
+from torch.autograd import Function
+from .utils.find_layers import replace_all_layer_type_recursive
+
+
+class GuidedBackpropReLU(Function):
+ @staticmethod
+ def forward(self, input_img):
+ positive_mask = (input_img > 0).type_as(input_img)
+ output = torch.addcmul(
+ torch.zeros(
+ input_img.size()).type_as(input_img),
+ input_img,
+ positive_mask)
+ self.save_for_backward(input_img, output)
+ return output
+
+ @staticmethod
+ def backward(self, grad_output):
+ input_img, output = self.saved_tensors
+ grad_input = None
+
+ positive_mask_1 = (input_img > 0).type_as(grad_output)
+ positive_mask_2 = (grad_output > 0).type_as(grad_output)
+ grad_input = torch.addcmul(
+ torch.zeros(
+ input_img.size()).type_as(input_img),
+ torch.addcmul(
+ torch.zeros(
+ input_img.size()).type_as(input_img),
+ grad_output,
+ positive_mask_1),
+ positive_mask_2)
+ return grad_input
+
+
+class GuidedBackpropReLUasModule(torch.nn.Module):
+ def __init__(self):
+ super(GuidedBackpropReLUasModule, self).__init__()
+
+ def forward(self, input_img):
+ return GuidedBackpropReLU.apply(input_img)
+
+
+class GuidedBackpropReLUModel:
+ def __init__(self, model, use_cuda):
+ self.model = model
+ self.model.eval()
+ self.cuda = use_cuda
+ if self.cuda:
+ self.model = self.model.cuda()
+
+ def forward(self, input_img):
+ return self.model(input_img)
+
+ def recursive_replace_relu_with_guidedrelu(self, module_top):
+
+ for idx, module in module_top._modules.items():
+ self.recursive_replace_relu_with_guidedrelu(module)
+ if module.__class__.__name__ == 'ReLU':
+ module_top._modules[idx] = GuidedBackpropReLU.apply
+ print("b")
+
+ def recursive_replace_guidedrelu_with_relu(self, module_top):
+ try:
+ for idx, module in module_top._modules.items():
+ self.recursive_replace_guidedrelu_with_relu(module)
+ if module == GuidedBackpropReLU.apply:
+ module_top._modules[idx] = torch.nn.ReLU()
+ except BaseException:
+ pass
+
+ def __call__(self, input_img, target_category=None):
+ replace_all_layer_type_recursive(self.model,
+ torch.nn.ReLU,
+ GuidedBackpropReLUasModule())
+
+ if self.cuda:
+ input_img = input_img.cuda()
+
+ input_img = input_img.requires_grad_(True)
+
+ output = self.forward(input_img)
+
+ if target_category is None:
+ target_category = np.argmax(output.cpu().data.numpy())
+
+ loss = output[0, target_category]
+ loss.backward(retain_graph=True)
+
+ output = input_img.grad.cpu().data.numpy()
+ output = output[0, :, :, :]
+ output = output.transpose((1, 2, 0))
+
+ replace_all_layer_type_recursive(self.model,
+ GuidedBackpropReLUasModule,
+ torch.nn.ReLU())
+
+ return output
diff --git a/models/dsp/CAM/pytorch_grad_cam/layer_cam.py b/models/dsp/CAM/pytorch_grad_cam/layer_cam.py
new file mode 100644
index 0000000000000000000000000000000000000000..8a7f32c4b4fdbd65fb8f426cfeb570fbdf12259e
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/layer_cam.py
@@ -0,0 +1,36 @@
+import numpy as np
+from .base_cam import BaseCAM
+from .utils.svd_on_activations import get_2d_projection
+
+# https://ieeexplore.ieee.org/document/9462463
+
+
+class LayerCAM(BaseCAM):
+ def __init__(
+ self,
+ model,
+ target_layers,
+ use_cuda=False,
+ reshape_transform=None):
+ super(
+ LayerCAM,
+ self).__init__(
+ model,
+ target_layers,
+ use_cuda,
+ reshape_transform)
+
+ def get_cam_image(self,
+ input_tensor,
+ target_layer,
+ target_category,
+ activations,
+ grads,
+ eigen_smooth):
+ spatial_weighted_activations = np.maximum(grads, 0) * activations
+
+ if eigen_smooth:
+ cam = get_2d_projection(spatial_weighted_activations)
+ else:
+ cam = spatial_weighted_activations.sum(axis=1)
+ return cam
diff --git a/models/dsp/CAM/pytorch_grad_cam/score_cam.py b/models/dsp/CAM/pytorch_grad_cam/score_cam.py
new file mode 100644
index 0000000000000000000000000000000000000000..574eefa38136812bf37d14cc0f55a20df1046fd1
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/score_cam.py
@@ -0,0 +1,63 @@
+import torch
+import tqdm
+from .base_cam import BaseCAM
+
+
+class ScoreCAM(BaseCAM):
+ def __init__(
+ self,
+ model,
+ target_layers,
+ use_cuda=False,
+ reshape_transform=None):
+ super(ScoreCAM, self).__init__(model,
+ target_layers,
+ use_cuda,
+ reshape_transform=reshape_transform,
+ uses_gradients=False)
+
+ if len(target_layers) > 0:
+ print("Warning: You are using ScoreCAM with target layers, "
+ "however ScoreCAM will ignore them.")
+
+ def get_cam_weights(self,
+ input_tensor,
+ target_layer,
+ targets,
+ activations,
+ grads):
+ with torch.no_grad():
+ upsample = torch.nn.UpsamplingBilinear2d(
+ size=input_tensor.shape[-2:])
+ activation_tensor = torch.from_numpy(activations)
+ if self.cuda:
+ activation_tensor = activation_tensor.cuda()
+
+ upsampled = upsample(activation_tensor)
+
+ maxs = upsampled.view(upsampled.size(0),
+ upsampled.size(1), -1).max(dim=-1)[0]
+ mins = upsampled.view(upsampled.size(0),
+ upsampled.size(1), -1).min(dim=-1)[0]
+
+ maxs, mins = maxs[:, :, None, None], mins[:, :, None, None]
+ upsampled = (upsampled - mins) / (maxs - mins)
+
+ input_tensors = input_tensor[:, None,
+ :, :] * upsampled[:, :, None, :, :]
+
+ if hasattr(self, "batch_size"):
+ BATCH_SIZE = self.batch_size
+ else:
+ BATCH_SIZE = 16
+
+ scores = []
+ for target, tensor in zip(targets, input_tensors):
+ for i in tqdm.tqdm(range(0, tensor.size(0), BATCH_SIZE)):
+ batch = tensor[i: i + BATCH_SIZE, :]
+ outputs = [target(o).cpu().item() for o in self.model(batch)]
+ scores.extend(outputs)
+ scores = torch.Tensor(scores)
+ scores = scores.view(activations.shape[0], activations.shape[1])
+ weights = torch.nn.Softmax(dim=-1)(scores).numpy()
+ return weights
diff --git a/models/dsp/CAM/pytorch_grad_cam/utils/__init__.py b/models/dsp/CAM/pytorch_grad_cam/utils/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..e85151f1fee90ffa89fc5fa81c552d2e5f6571b5
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/utils/__init__.py
@@ -0,0 +1,4 @@
+from .image import deprocess_image
+from .svd_on_activations import get_2d_projection
+from . import model_targets
+from . import reshape_transforms
\ No newline at end of file
diff --git a/models/dsp/CAM/pytorch_grad_cam/utils/find_layers.py b/models/dsp/CAM/pytorch_grad_cam/utils/find_layers.py
new file mode 100644
index 0000000000000000000000000000000000000000..8373a48bace9bf2314bb06acac7664ec48504355
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/utils/find_layers.py
@@ -0,0 +1,30 @@
+def replace_layer_recursive(model, old_layer, new_layer):
+ for name, layer in model._modules.items():
+ if layer == old_layer:
+ model._modules[name] = new_layer
+ return True
+ elif replace_layer_recursive(layer, old_layer, new_layer):
+ return True
+ return False
+
+
+def replace_all_layer_type_recursive(model, old_layer_type, new_layer):
+ for name, layer in model._modules.items():
+ if isinstance(layer, old_layer_type):
+ model._modules[name] = new_layer
+ replace_all_layer_type_recursive(layer, old_layer_type, new_layer)
+
+
+def find_layer_types_recursive(model, layer_types):
+ def predicate(layer):
+ return type(layer) in layer_types
+ return find_layer_predicate_recursive(model, predicate)
+
+
+def find_layer_predicate_recursive(model, predicate):
+ result = []
+ for name, layer in model._modules.items():
+ if predicate(layer):
+ result.append(layer)
+ result.extend(find_layer_predicate_recursive(layer, predicate))
+ return result
\ No newline at end of file
diff --git a/models/dsp/CAM/pytorch_grad_cam/utils/image.py b/models/dsp/CAM/pytorch_grad_cam/utils/image.py
new file mode 100644
index 0000000000000000000000000000000000000000..236693e738e0f8e38b019e28590e4b190d8fab7d
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/utils/image.py
@@ -0,0 +1,90 @@
+import cv2
+import numpy as np
+import torch
+import torch.nn.functional as F
+from torchvision.transforms import Compose, Normalize, ToTensor
+
+
+def preprocess_image(img: np.ndarray, mean=[0.5, 0.5, 0.5], std=[0.5, 0.5, 0.5]) -> torch.Tensor:
+ preprocessing = Compose([
+ ToTensor(),
+ Normalize(mean=mean, std=std)
+ ])
+ return preprocessing(img.copy()).unsqueeze(0)
+
+
+def deprocess_image(img):
+ """ see https://github.com/jacobgil/keras-grad-cam/blob/master/grad-cam.py#L65 """
+ img = img - np.mean(img)
+ img = img / (np.std(img) + 1e-5)
+ img = img * 0.1
+ img = img + 0.5
+ img = np.clip(img, 0, 1)
+ return np.uint8(img * 255)
+
+
+def show_cam_on_image(img: np.ndarray,
+ mask: np.ndarray,
+ use_rgb: bool = False,
+ colormap: int = cv2.COLORMAP_JET) -> np.ndarray:
+ """ This function overlays the cam mask on the image as an heatmap.
+ By default the heatmap is in BGR format.
+
+ :param img: The base image in RGB or BGR format.
+ :param mask: The cam mask.
+ :param use_rgb: Whether to use an RGB or BGR heatmap, this should be set to True if 'img' is in RGB format.
+ :param colormap: The OpenCV colormap to be used.
+ :returns: The default image with the cam overlay.
+ """
+ heatmap = cv2.applyColorMap(np.uint8(255 * mask), colormap)
+ if use_rgb:
+ heatmap = cv2.cvtColor(heatmap, cv2.COLOR_BGR2RGB)
+ heatmap = np.float32(heatmap) / 255
+
+ if np.max(img) > 1:
+ raise Exception(
+ "The input image should np.float32 in the range [0, 1]")
+
+ cam = heatmap + img
+ cam = cam / np.max(cam)
+ return np.uint8(255 * cam)
+
+def scale_cam_image(cam, target_size=None):
+ result = []
+ for img in cam:
+ img = img - np.min(img)
+ img = img / (1e-7 + np.max(img))
+ if target_size is not None:
+ img = cv2.resize(img, target_size)
+ result.append(img)
+ result = np.float32(result)
+
+ return result
+
+def scale_cam_image_torch(cam: torch.Tensor, target_size=None):
+ if cam.ndim == 3:
+ cam = cam.unsqueeze(1)
+
+ # normalize per image
+ cam_min = cam.amin(dim=(2, 3), keepdim=True)
+ cam_max = cam.amax(dim=(2, 3), keepdim=True)
+ cam = cam - cam_min
+ cam = cam / (cam_max + 1e-7)
+
+ # resize if needed
+ if target_size is not None:
+ cam = F.interpolate(cam, size=target_size, mode="bilinear", align_corners=False)
+
+ return cam.squeeze(1)
+
+def scale_accross_batch_and_channels(tensor, target_size):
+ batch_size, channel_size = tensor.shape[:2]
+ reshaped_tensor = tensor.reshape(
+ batch_size * channel_size, *tensor.shape[2:])
+ result = scale_cam_image(reshaped_tensor, target_size)
+ result = result.reshape(
+ batch_size,
+ channel_size,
+ target_size[1],
+ target_size[0])
+ return result
diff --git a/models/dsp/CAM/pytorch_grad_cam/utils/model_targets.py b/models/dsp/CAM/pytorch_grad_cam/utils/model_targets.py
new file mode 100644
index 0000000000000000000000000000000000000000..260e21ca10f7b652a759aae44516c0f94b25d72e
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/utils/model_targets.py
@@ -0,0 +1,61 @@
+import numpy as np
+import torch
+import torchvision
+
+class ClassifierOutputTarget:
+ def __init__(self, category):
+ self.category = category
+ def __call__(self, model_output):
+ if len(model_output.shape) == 1:
+ return model_output[self.category]
+ return model_output[:, self.category]
+
+class SemanticSegmentationTarget:
+ """ Gets a binary spatial mask and a category,
+ And return the sum of the category scores,
+ of the pixels in the mask. """
+ def __init__(self, category, mask):
+ self.category = category
+ self.mask = torch.from_numpy(mask)
+ if torch.cuda.is_available():
+ self.mask = self.mask.cuda()
+
+ def __call__(self, model_output):
+ return (model_output[self.category, :, : ] * self.mask).sum()
+
+
+class FasterRCNNBoxScoreTarget:
+ """ For every original detected bounding box specified in "bounding boxes",
+ assign a score on how the current bounding boxes match it,
+ 1. In IOU
+ 2. In the classification score.
+ If there is not a large enough overlap, or the category changed,
+ assign a score of 0.
+
+ The total score is the sum of all the box scores.
+ """
+
+ def __init__(self, labels, bounding_boxes, iou_threshold=0.5):
+ self.labels = labels
+ self.bounding_boxes = bounding_boxes
+ self.iou_threshold = iou_threshold
+
+ def __call__(self, model_outputs):
+ output = torch.Tensor([0])
+ if torch.cuda.is_available():
+ output = output.cuda()
+
+ if len(model_outputs["boxes"]) == 0:
+ return output
+
+ for box, label in zip(self.bounding_boxes, self.labels):
+ box = torch.Tensor(box[None, :])
+ if torch.cuda.is_available():
+ box = box.cuda()
+
+ ious = torchvision.ops.box_iou(box, model_outputs["boxes"])
+ index = ious.argmax()
+ if ious[0, index] > self.iou_threshold and model_outputs["labels"][index] == label:
+ score = ious[0, index] + model_outputs["scores"][index]
+ output = output + score
+ return output
\ No newline at end of file
diff --git a/models/dsp/CAM/pytorch_grad_cam/utils/reshape_transforms.py b/models/dsp/CAM/pytorch_grad_cam/utils/reshape_transforms.py
new file mode 100644
index 0000000000000000000000000000000000000000..390d5470c60fb5822990bb3b359e2338c6558be0
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/utils/reshape_transforms.py
@@ -0,0 +1,27 @@
+import torch
+
+def fasterrcnn_reshape_transform(x):
+ target_size = x['pool'].size()[-2 : ]
+ activations = []
+ for key, value in x.items():
+ activations.append(torch.nn.functional.interpolate(torch.abs(value), target_size, mode='bilinear'))
+ activations = torch.cat(activations, axis=1)
+ return activations
+
+def swinT_reshape_transform(tensor, height=7, width=7):
+ result = tensor.reshape(tensor.size(0),
+ height, width, tensor.size(2))
+
+ # Bring the channels to the first dimension,
+ # like in CNNs.
+ result = result.transpose(2, 3).transpose(1, 2)
+ return result
+
+def vit_reshape_transform(tensor, height=14, width=14):
+ result = tensor[:, 1:, :].reshape(tensor.size(0),
+ height, width, tensor.size(2))
+
+ # Bring the channels to the first dimension,
+ # like in CNNs.
+ result = result.transpose(2, 3).transpose(1, 2)
+ return result
diff --git a/models/dsp/CAM/pytorch_grad_cam/utils/svd_on_activations.py b/models/dsp/CAM/pytorch_grad_cam/utils/svd_on_activations.py
new file mode 100644
index 0000000000000000000000000000000000000000..a406aeea85617922e67270a70388256ac214e8e2
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/utils/svd_on_activations.py
@@ -0,0 +1,19 @@
+import numpy as np
+
+
+def get_2d_projection(activation_batch):
+ # TBD: use pytorch batch svd implementation
+ activation_batch[np.isnan(activation_batch)] = 0
+ projections = []
+ for activations in activation_batch:
+ reshaped_activations = (activations).reshape(
+ activations.shape[0], -1).transpose()
+ # Centering before the SVD seems to be important here,
+ # Otherwise the image returned is negative
+ reshaped_activations = reshaped_activations - \
+ reshaped_activations.mean(axis=0)
+ U, S, VT = np.linalg.svd(reshaped_activations, full_matrices=True)
+ projection = reshaped_activations @ VT[0, :]
+ projection = projection.reshape(activations.shape[1:])
+ projections.append(projection)
+ return np.float32(projections)
diff --git a/models/dsp/CAM/pytorch_grad_cam/xgrad_cam.py b/models/dsp/CAM/pytorch_grad_cam/xgrad_cam.py
new file mode 100644
index 0000000000000000000000000000000000000000..4ce867a4e6b61a9fa4f32962c05be051eba3d0e4
--- /dev/null
+++ b/models/dsp/CAM/pytorch_grad_cam/xgrad_cam.py
@@ -0,0 +1,31 @@
+import numpy as np
+from .base_cam import BaseCAM
+
+
+class XGradCAM(BaseCAM):
+ def __init__(
+ self,
+ model,
+ target_layers,
+ use_cuda=False,
+ reshape_transform=None):
+ super(
+ XGradCAM,
+ self).__init__(
+ model,
+ target_layers,
+ use_cuda,
+ reshape_transform)
+
+ def get_cam_weights(self,
+ input_tensor,
+ target_layer,
+ target_category,
+ activations,
+ grads):
+ sum_activations = np.sum(activations, axis=(2, 3))
+ eps = 1e-7
+ weights = grads * activations / \
+ (sum_activations[:, :, None, None] + eps)
+ weights = weights.sum(axis=(2, 3))
+ return weights
diff --git a/models/dsp/CAM/utils.py b/models/dsp/CAM/utils.py
new file mode 100644
index 0000000000000000000000000000000000000000..686385b619cad16d1316b8602b47ecb37a98f753
--- /dev/null
+++ b/models/dsp/CAM/utils.py
@@ -0,0 +1,69 @@
+from . import clip
+import torch
+import numpy as np
+import cv2
+_CONTOUR_INDEX = 1 if cv2.__version__.split('.')[0] == '3' else 0
+
+
+class ClipOutputTarget:
+ def __init__(self, category):
+ self.category = category
+ def __call__(self, model_output):
+ if len(model_output.shape) == 1:
+ return model_output[self.category]
+ return model_output[:, self.category]
+
+
+def reshape_transform(tensor, height=28, width=28):
+ tensor = tensor.permute(1, 0, 2)
+ result = tensor[:, 1:, :].reshape(tensor.size(0), height, width, tensor.size(2))
+
+ # Bring the channels to the first dimension,
+ # like in CNNs.
+ result = result.transpose(2, 3).transpose(1, 2)
+ return result
+
+
+def zeroshot_classifier(classnames, templates, model, device):
+ with torch.no_grad():
+ zeroshot_weights = []
+ for classname in classnames:
+ texts = [template.format(classname) for template in templates] #format with class
+ texts = clip.tokenize(texts).to(device) #tokenize
+ class_embeddings = model.encode_text(texts) #embed with text encoder
+ class_embeddings /= class_embeddings.norm(dim=-1, keepdim=True)
+ class_embedding = class_embeddings.mean(dim=0)
+ class_embedding /= class_embedding.norm()
+ zeroshot_weights.append(class_embedding)
+ zeroshot_weights = torch.stack(zeroshot_weights, dim=1).to(device)
+ return zeroshot_weights.t()
+
+
+def scoremap2bbox(scoremap, threshold, multi_contour_eval=False):
+ height, width = scoremap.shape
+ scoremap_image = np.expand_dims((scoremap * 255).astype(np.uint8), 2)
+ _, thr_gray_heatmap = cv2.threshold(
+ src=scoremap_image,
+ thresh=int(threshold * np.max(scoremap_image)),
+ maxval=255,
+ type=cv2.THRESH_BINARY)
+ contours = cv2.findContours(
+ image=thr_gray_heatmap,
+ mode=cv2.RETR_TREE,
+ method=cv2.CHAIN_APPROX_SIMPLE)[_CONTOUR_INDEX]
+
+ if len(contours) == 0:
+ return np.asarray([[0, 0, 0, 0]]), 1
+
+ if not multi_contour_eval:
+ contours = [max(contours, key=cv2.contourArea)]
+
+ estimated_boxes = []
+ for contour in contours:
+ x, y, w, h = cv2.boundingRect(contour)
+ x0, y0, x1, y1 = x, y, x + w, y + h
+ x1 = min(x1, width - 1)
+ y1 = min(y1, height - 1)
+ estimated_boxes.append([x0, y0, x1, y1])
+
+ return np.asarray(estimated_boxes), len(contours)
\ No newline at end of file
diff --git a/models/dsp/__init__.py b/models/dsp/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..392b75799f52bf180a16e71d03141ab382373f1a
--- /dev/null
+++ b/models/dsp/__init__.py
@@ -0,0 +1,5 @@
+from .attention_processor import set_processors
+from .projection import Resampler, SerialSampler
+from .utils import ExemplarPool, seed_everything, get_masks, get_sigmoid
+from .modules import FrozenDinoV2Encoder
+from .CAM import CAMGenerator
\ No newline at end of file
diff --git a/models/dsp/attention_processor.py b/models/dsp/attention_processor.py
new file mode 100644
index 0000000000000000000000000000000000000000..01575918994f85bf902fb4e1c977b5d20c02d959
--- /dev/null
+++ b/models/dsp/attention_processor.py
@@ -0,0 +1,263 @@
+import cv2
+import numpy as np
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from diffusers.models.attention_processor import Attention
+from typing import Optional
+
+from .layers import MIFusion, MIFusionPrototype
+from .utils import get_masks, get_sigmoid
+
+
+class AttnProcessor2_0(nn.Module):
+ r"""
+ Processor for implementing scaled dot-product attention (enabled by default if you're using PyTorch 2.0).
+ """
+
+ def __init__(self, hidden_size=None, cross_attention_dim=None):
+ if not hasattr(F, "scaled_dot_product_attention"):
+ raise ImportError("AttnProcessor2_0 requires PyTorch 2.0, to use it, please upgrade PyTorch to 2.0.")
+ super().__init__()
+
+ def __call__(
+ self,
+ attn: Attention,
+ hidden_states: torch.Tensor,
+ encoder_hidden_states: Optional[torch.Tensor] = None,
+ attention_mask: Optional[torch.Tensor] = None,
+ temb: Optional[torch.Tensor] = None,
+ # Useless
+ bboxes=[],
+ obboxes=[],
+ embeds_pooler=None,
+ height=512,
+ width=512,
+ prototypes=None,
+ ref_features=None,
+ guidance_masks=None,
+ supplement_mask=None,
+ sigmoid_values=None,
+ in_box=None,
+ do_classifier_free_guidance=False,
+ # End Useless
+ *args,
+ **kwargs,
+ ) -> torch.Tensor:
+
+ residual = hidden_states
+ if attn.spatial_norm is not None:
+ hidden_states = attn.spatial_norm(hidden_states, temb)
+
+ input_ndim = hidden_states.ndim
+
+ if input_ndim == 4:
+ batch_size, channel, height, width = hidden_states.shape
+ hidden_states = hidden_states.view(batch_size, channel, height * width).transpose(1, 2)
+
+ batch_size, sequence_length, _ = (
+ hidden_states.shape if encoder_hidden_states is None else encoder_hidden_states.shape
+ )
+
+ if attention_mask is not None:
+ attention_mask = attn.prepare_attention_mask(attention_mask, sequence_length, batch_size)
+ # scaled_dot_product_attention expects attention_mask shape to be
+ # (batch, heads, source_length, target_length)
+ attention_mask = attention_mask.view(batch_size, attn.heads, -1, attention_mask.shape[-1])
+
+ if attn.group_norm is not None:
+ hidden_states = attn.group_norm(hidden_states.transpose(1, 2)).transpose(1, 2)
+
+ query = attn.to_q(hidden_states)
+
+ if encoder_hidden_states is None:
+ encoder_hidden_states = hidden_states
+ elif attn.norm_cross:
+ encoder_hidden_states = attn.norm_encoder_hidden_states(encoder_hidden_states)
+
+ key = attn.to_k(encoder_hidden_states)
+ value = attn.to_v(encoder_hidden_states)
+
+ inner_dim = key.shape[-1]
+ head_dim = inner_dim // attn.heads
+
+ query = query.view(batch_size, -1, attn.heads, head_dim).transpose(1, 2)
+
+ key = key.view(batch_size, -1, attn.heads, head_dim).transpose(1, 2)
+ value = value.view(batch_size, -1, attn.heads, head_dim).transpose(1, 2)
+
+ if attn.norm_q is not None:
+ query = attn.norm_q(query)
+ if attn.norm_k is not None:
+ key = attn.norm_k(key)
+
+ # the output of sdp = (batch, num_heads, seq_len, head_dim)
+ # TODO: add support for attn.scale when we move to Torch 2.1
+ hidden_states = F.scaled_dot_product_attention(
+ query, key, value, attn_mask=attention_mask, dropout_p=0.0, is_causal=False
+ )
+
+ hidden_states = hidden_states.transpose(1, 2).reshape(batch_size, -1, attn.heads * head_dim)
+ hidden_states = hidden_states.to(query.dtype)
+
+ # linear proj
+ hidden_states = attn.to_out[0](hidden_states)
+ # dropout
+ hidden_states = attn.to_out[1](hidden_states)
+
+ if input_ndim == 4:
+ hidden_states = hidden_states.transpose(-1, -2).reshape(batch_size, channel, height, width)
+
+ if attn.residual_connection:
+ hidden_states = hidden_states + residual
+
+ hidden_states = hidden_states / attn.rescale_output_factor
+
+ return hidden_states
+
+
+class MaskedProcessor2_0(nn.Module):
+ def __init__(self, hidden_size, cross_attention_dim=None,
+ use_ea_attn=False, **kwargs):
+ super().__init__()
+ self.hidden_size = hidden_size
+ self.cross_attention_dim = cross_attention_dim
+ self.use_ea_attn = use_ea_attn
+ self.prototype_mode = use_ea_attn and (kwargs['phase'] == 'novel') and (kwargs.get('attn_prototype_switch') is True)
+ if self.prototype_mode:
+ self.fusion = MIFusionPrototype(hidden_size, context_dim=cross_attention_dim)
+ elif use_ea_attn:
+ self.fusion = MIFusion(hidden_size, context_dim=cross_attention_dim)
+
+ # Train ; Test (do_classifier_free_guidance)
+ # Train // Train2(self.use_ea_attn) ; Test // Test2(self.use_ea_attn)
+ def __call__(
+ self,
+ attn: Attention,
+ # Output the same size as hidden_states, encoder_hidden_states as Key and Value to inject information to hidden_states
+ # shape[-2] 4096 as 64x64, 64 as 8x8; shape[-1] as hidden_size
+ hidden_states, # [1, 4096, 320] // [1, 64, 1280] ; [2, 4096, 320] // [2, 64, 1280] && torch.all(hidden_states[0] == hidden_states[1]) is Ture
+ encoder_hidden_states=None, # [16, 77, 768] ; [17, 77, 768]
+ attention_mask=None,
+ bboxes=[],
+ obboxes=[],
+ embeds_pooler=None,
+ height=512,
+ width=512,
+ prototypes=None,
+ ref_features=None,
+ guidance_masks=None,
+ supplement_mask=None,
+ sigmoid_values=None,
+ in_box=None,
+ do_classifier_free_guidance=False,
+ ):
+
+ instance_num = len(obboxes[0]) # 15
+
+ if not self.use_ea_attn:
+ # [1, 77, 768]; [2, 77, 768]
+ encoder_hidden_states = encoder_hidden_states[:2, ...] if do_classifier_free_guidance else encoder_hidden_states[:1, ...]
+
+ if self.use_ea_attn:
+ if do_classifier_free_guidance:
+ hidden_states = torch.cat([hidden_states[0:1], hidden_states[1:2].repeat(instance_num + 1, 1, 1)]) # ;//[17, 64, 1280]
+ image_token = hidden_states[1:]
+ else:
+ hidden_states = hidden_states.repeat(instance_num + 1, 1, 1) # //[16, 64, 1280]
+ image_token = hidden_states
+
+ batch_size, sequence_length, _ = hidden_states.shape # _ is hidden_size
+
+ if attention_mask is not None:
+ attention_mask = attn.prepare_attention_mask(attention_mask, sequence_length, batch_size)
+ # scaled_dot_product_attention expects attention_mask shape to be
+ # (batch, heads, source_length, target_length)
+ attention_mask = attention_mask.view(batch_size, attn.heads, -1, attention_mask.shape[-1])
+
+ query = attn.to_q(hidden_states) # [1, 4096, 320] // [16, 64, 1280] ; [2, 4096, 320]->[2, 4096, 320] // [17, 64, 1280]->[17, 64, 1280]
+ key = attn.to_k(encoder_hidden_states) # [1, 77, 320] // [16, 77, 1280] ; [2, 77, 768]->[2, 77, 320] // [17, 77, 768 (cross_attention_dim)] -> [17, 77, 1280 (self.inner_kv_dim)]
+ value = attn.to_v(encoder_hidden_states) # Same with key
+
+ inner_dim = key.shape[-1] # 320 // 1280
+ head_dim = inner_dim // attn.heads # 40 // 160
+
+ query = query.view(batch_size, -1, attn.heads, head_dim).transpose(1, 2) # [1, 8, 4096, 40] // [16, 8, 64, 160]
+
+ key = key.view(batch_size, -1, attn.heads, head_dim).transpose(1, 2) # [1, 8, 77, 40] // [16, 8, 77, 160]
+ value = value.view(batch_size, -1, attn.heads, head_dim).transpose(1, 2) # [1, 8, 77, 40] // [16, 8, 77, 160]
+
+ if attn.norm_q is not None:
+ query = attn.norm_q(query)
+ if attn.norm_k is not None:
+ key = attn.norm_k(key)
+
+ hidden_states = F.scaled_dot_product_attention(
+ query, key, value, attn_mask=attention_mask, dropout_p=0.0, is_causal=False
+ ) # [1, 8, 4096, 40] // [16, 8, 64, 1280]
+
+ hidden_states = hidden_states.transpose(1, 2).reshape(batch_size, -1, attn.heads * head_dim) # [1, 4096, 320] // [16, 64, 1280]
+ hidden_states = hidden_states.to(query.dtype)
+ hidden_states = attn.to_out[0](hidden_states) # [1, 4096, 320] // [16, 64, 1280] ; [2, 4096, 320] // [17, 64, 1280] # Linear
+ hidden_states = attn.to_out[1](hidden_states) # [1, 4096, 320] // [16, 64, 1280] ; [2, 4096, 320] // [17, 64, 1280] # Dropout
+
+ if not self.use_ea_attn:
+ return hidden_states
+
+ assert self.use_ea_attn
+
+ if do_classifier_free_guidance:
+ hidden_states_uncond, hidden_states = hidden_states[0:1], hidden_states[1:] # torch.Size([1, HW, C])
+
+ other_info = {}
+ other_info['image_token'] = image_token.unsqueeze(0) # [1, 16, 64, 1280]
+ other_info['context'] = encoder_hidden_states[1:, ...] if do_classifier_free_guidance else encoder_hidden_states # [16, 77, 768]
+ other_info['box'] = in_box # [1, 15, 8]
+ other_info['context_pooler'] = embeds_pooler # [16, 1, 768]
+ other_info['supplement_mask'] = supplement_mask # [1, 1, 64, 64]
+ other_info['height'] = height # 512
+ other_info['width'] = width # 512
+ other_info['ref_features'] = ref_features # [15, 16, 768], [1, 16, 768]
+ other_info['sigmoid_values'] = sigmoid_values # [1, 16, 768]
+ other_info['guidance_masks'] = guidance_masks # [1, 15, 64, 64]
+ other_info['instance_num'] = instance_num
+ other_info['prototypes'] = prototypes
+
+ hidden_states = self.fusion(hidden_states.unsqueeze(0), # [1, 16, 64, 1280]
+ other_info=other_info)
+ # hidden_states_cond.shape [1, 64, 1280]
+
+ if do_classifier_free_guidance:
+ hidden_states = torch.cat([hidden_states_uncond, hidden_states])
+ return hidden_states
+
+
+def set_processors(unet, **kwargs):
+ attn_processors = {}
+ for name, _ in unet.attn_processors.items():
+ use_ea_attn = False
+ kwargs['attn_prototype_switch'] = False
+ cross_attention_dim = None if name.endswith("attn1.processor") else unet.config.cross_attention_dim
+ if name.startswith("mid_block"):
+ hidden_size = unet.config.block_out_channels[-1] # unet.config.block_out_channels [320, 640, 1280, 1280]
+ use_ea_attn = True
+ elif name.startswith("up_blocks"):
+ block_id = int(name[len("up_blocks.")])
+ attention_id = int(name[len("up_blocks.2.attentions.")])
+ hidden_size = list(reversed(unet.config.block_out_channels))[block_id]
+ if block_id == 1:
+ use_ea_attn = True
+ elif (block_id != 1) and (kwargs['phase'] == 'novel'): # run-4
+ use_ea_attn = True
+ kwargs['attn_prototype_switch'] = True
+ elif name.startswith("down_blocks"):
+ block_id = int(name[len("down_blocks.")])
+ hidden_size = unet.config.block_out_channels[block_id]
+ if cross_attention_dim is not None:
+ attn_processors[name] = MaskedProcessor2_0(hidden_size=hidden_size,
+ cross_attention_dim=cross_attention_dim,
+ use_ea_attn=use_ea_attn,
+ **kwargs)
+ else:
+ attn_processors[name] = AttnProcessor2_0()
+ unet.set_attn_processor(attn_processors)
\ No newline at end of file
diff --git a/models/dsp/dinov2/dinov2/__init__.py b/models/dsp/dinov2/dinov2/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..ae847e46898077fe3d8701b8a181d7b4e3d41cd9
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/__init__.py
@@ -0,0 +1,6 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+__version__ = "0.0.1"
diff --git a/models/dsp/dinov2/dinov2/configs/__init__.py b/models/dsp/dinov2/dinov2/configs/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..68e0830c62ea19649b6cd2361995f6df309d7640
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/configs/__init__.py
@@ -0,0 +1,22 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import pathlib
+
+from omegaconf import OmegaConf
+
+
+def load_config(config_name: str):
+ config_filename = config_name + ".yaml"
+ return OmegaConf.load(pathlib.Path(__file__).parent.resolve() / config_filename)
+
+
+dinov2_default_config = load_config("ssl_default_config")
+
+
+def load_and_merge_config(config_name: str):
+ default_config = OmegaConf.create(dinov2_default_config)
+ loaded_config = load_config(config_name)
+ return OmegaConf.merge(default_config, loaded_config)
diff --git a/models/dsp/dinov2/dinov2/configs/eval/vitb14_pretrain.yaml b/models/dsp/dinov2/dinov2/configs/eval/vitb14_pretrain.yaml
new file mode 100644
index 0000000000000000000000000000000000000000..117d0f027ca26cd8ce6c010bb78d5a8fac42c70e
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/configs/eval/vitb14_pretrain.yaml
@@ -0,0 +1,6 @@
+student:
+ arch: vit_base
+ patch_size: 14
+crops:
+ global_crops_size: 518 # this is to set up the position embeddings properly
+ local_crops_size: 98
\ No newline at end of file
diff --git a/models/dsp/dinov2/dinov2/configs/eval/vitb14_reg4_pretrain.yaml b/models/dsp/dinov2/dinov2/configs/eval/vitb14_reg4_pretrain.yaml
new file mode 100644
index 0000000000000000000000000000000000000000..d53edc04a0761b4b35c147d63e04d55c90092c8f
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/configs/eval/vitb14_reg4_pretrain.yaml
@@ -0,0 +1,9 @@
+student:
+ arch: vit_base
+ patch_size: 14
+ num_register_tokens: 4
+ interpolate_antialias: true
+ interpolate_offset: 0.0
+crops:
+ global_crops_size: 518 # this is to set up the position embeddings properly
+ local_crops_size: 98
\ No newline at end of file
diff --git a/models/dsp/dinov2/dinov2/configs/eval/vitg14_pretrain.yaml b/models/dsp/dinov2/dinov2/configs/eval/vitg14_pretrain.yaml
new file mode 100644
index 0000000000000000000000000000000000000000..a96dd5b117b4d59ee210b65037821f1b3e3f16e3
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/configs/eval/vitg14_pretrain.yaml
@@ -0,0 +1,7 @@
+student:
+ arch: vit_giant2
+ patch_size: 14
+ ffn_layer: swiglufused
+crops:
+ global_crops_size: 518 # this is to set up the position embeddings properly
+ local_crops_size: 98
\ No newline at end of file
diff --git a/models/dsp/dinov2/dinov2/configs/eval/vitg14_reg4_pretrain.yaml b/models/dsp/dinov2/dinov2/configs/eval/vitg14_reg4_pretrain.yaml
new file mode 100644
index 0000000000000000000000000000000000000000..15948f8589ea0a6e04717453eb88c18388e7f1b2
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/configs/eval/vitg14_reg4_pretrain.yaml
@@ -0,0 +1,10 @@
+student:
+ arch: vit_giant2
+ patch_size: 14
+ ffn_layer: swiglufused
+ num_register_tokens: 4
+ interpolate_antialias: true
+ interpolate_offset: 0.0
+crops:
+ global_crops_size: 518 # this is to set up the position embeddings properly
+ local_crops_size: 98
\ No newline at end of file
diff --git a/models/dsp/dinov2/dinov2/configs/eval/vitl14_pretrain.yaml b/models/dsp/dinov2/dinov2/configs/eval/vitl14_pretrain.yaml
new file mode 100644
index 0000000000000000000000000000000000000000..7a984548bd034f762d455419d7193917fa462dd8
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/configs/eval/vitl14_pretrain.yaml
@@ -0,0 +1,6 @@
+student:
+ arch: vit_large
+ patch_size: 14
+crops:
+ global_crops_size: 518 # this is to set up the position embeddings properly
+ local_crops_size: 98
\ No newline at end of file
diff --git a/models/dsp/dinov2/dinov2/configs/eval/vitl14_reg4_pretrain.yaml b/models/dsp/dinov2/dinov2/configs/eval/vitl14_reg4_pretrain.yaml
new file mode 100644
index 0000000000000000000000000000000000000000..0e2bc4e7b24b1a64d0369a24927996d0f184e283
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/configs/eval/vitl14_reg4_pretrain.yaml
@@ -0,0 +1,9 @@
+student:
+ arch: vit_large
+ patch_size: 14
+ num_register_tokens: 4
+ interpolate_antialias: true
+ interpolate_offset: 0.0
+crops:
+ global_crops_size: 518 # this is to set up the position embeddings properly
+ local_crops_size: 98
\ No newline at end of file
diff --git a/models/dsp/dinov2/dinov2/configs/eval/vits14_pretrain.yaml b/models/dsp/dinov2/dinov2/configs/eval/vits14_pretrain.yaml
new file mode 100644
index 0000000000000000000000000000000000000000..afbdb4ba14f1c97130a25b579360f4d817cda495
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/configs/eval/vits14_pretrain.yaml
@@ -0,0 +1,6 @@
+student:
+ arch: vit_small
+ patch_size: 14
+crops:
+ global_crops_size: 518 # this is to set up the position embeddings properly
+ local_crops_size: 98
\ No newline at end of file
diff --git a/models/dsp/dinov2/dinov2/configs/eval/vits14_reg4_pretrain.yaml b/models/dsp/dinov2/dinov2/configs/eval/vits14_reg4_pretrain.yaml
new file mode 100644
index 0000000000000000000000000000000000000000..d25fd638389bfba9220792302dc9dbf5d9a2406a
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/configs/eval/vits14_reg4_pretrain.yaml
@@ -0,0 +1,9 @@
+student:
+ arch: vit_small
+ patch_size: 14
+ num_register_tokens: 4
+ interpolate_antialias: true
+ interpolate_offset: 0.0
+crops:
+ global_crops_size: 518 # this is to set up the position embeddings properly
+ local_crops_size: 98
\ No newline at end of file
diff --git a/models/dsp/dinov2/dinov2/configs/ssl_default_config.yaml b/models/dsp/dinov2/dinov2/configs/ssl_default_config.yaml
new file mode 100644
index 0000000000000000000000000000000000000000..ccaae1c3174b21bcaf6e803dc861492261e5abe1
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/configs/ssl_default_config.yaml
@@ -0,0 +1,118 @@
+MODEL:
+ WEIGHTS: ''
+compute_precision:
+ grad_scaler: true
+ teacher:
+ backbone:
+ sharding_strategy: SHARD_GRAD_OP
+ mixed_precision:
+ param_dtype: fp16
+ reduce_dtype: fp16
+ buffer_dtype: fp32
+ dino_head:
+ sharding_strategy: SHARD_GRAD_OP
+ mixed_precision:
+ param_dtype: fp16
+ reduce_dtype: fp16
+ buffer_dtype: fp32
+ ibot_head:
+ sharding_strategy: SHARD_GRAD_OP
+ mixed_precision:
+ param_dtype: fp16
+ reduce_dtype: fp16
+ buffer_dtype: fp32
+ student:
+ backbone:
+ sharding_strategy: SHARD_GRAD_OP
+ mixed_precision:
+ param_dtype: fp16
+ reduce_dtype: fp16
+ buffer_dtype: fp32
+ dino_head:
+ sharding_strategy: SHARD_GRAD_OP
+ mixed_precision:
+ param_dtype: fp16
+ reduce_dtype: fp32
+ buffer_dtype: fp32
+ ibot_head:
+ sharding_strategy: SHARD_GRAD_OP
+ mixed_precision:
+ param_dtype: fp16
+ reduce_dtype: fp32
+ buffer_dtype: fp32
+dino:
+ loss_weight: 1.0
+ head_n_prototypes: 65536
+ head_bottleneck_dim: 256
+ head_nlayers: 3
+ head_hidden_dim: 2048
+ koleo_loss_weight: 0.1
+ibot:
+ loss_weight: 1.0
+ mask_sample_probability: 0.5
+ mask_ratio_min_max:
+ - 0.1
+ - 0.5
+ separate_head: false
+ head_n_prototypes: 65536
+ head_bottleneck_dim: 256
+ head_nlayers: 3
+ head_hidden_dim: 2048
+train:
+ batch_size_per_gpu: 64
+ dataset_path: ImageNet:split=TRAIN
+ output_dir: .
+ saveckp_freq: 20
+ seed: 0
+ num_workers: 10
+ OFFICIAL_EPOCH_LENGTH: 1250
+ cache_dataset: true
+ centering: "centering" # or "sinkhorn_knopp"
+student:
+ arch: vit_large
+ patch_size: 16
+ drop_path_rate: 0.3
+ layerscale: 1.0e-05
+ drop_path_uniform: true
+ pretrained_weights: ''
+ ffn_layer: "mlp"
+ block_chunks: 0
+ qkv_bias: true
+ proj_bias: true
+ ffn_bias: true
+ num_register_tokens: 0
+ interpolate_antialias: false
+ interpolate_offset: 0.1
+teacher:
+ momentum_teacher: 0.992
+ final_momentum_teacher: 1
+ warmup_teacher_temp: 0.04
+ teacher_temp: 0.07
+ warmup_teacher_temp_epochs: 30
+optim:
+ epochs: 100
+ weight_decay: 0.04
+ weight_decay_end: 0.4
+ base_lr: 0.004 # learning rate for a batch size of 1024
+ lr: 0. # will be set after applying scaling rule
+ warmup_epochs: 10
+ min_lr: 1.0e-06
+ clip_grad: 3.0
+ freeze_last_layer_epochs: 1
+ scaling_rule: sqrt_wrt_1024
+ patch_embed_lr_mult: 0.2
+ layerwise_decay: 0.9
+ adamw_beta1: 0.9
+ adamw_beta2: 0.999
+crops:
+ global_crops_scale:
+ - 0.32
+ - 1.0
+ local_crops_number: 8
+ local_crops_scale:
+ - 0.05
+ - 0.32
+ global_crops_size: 224
+ local_crops_size: 96
+evaluation:
+ eval_period_iterations: 12500
diff --git a/models/dsp/dinov2/dinov2/configs/train/vitg14.yaml b/models/dsp/dinov2/dinov2/configs/train/vitg14.yaml
new file mode 100644
index 0000000000000000000000000000000000000000..d05cf0d59e07ac6e4a2b0f9bdcb6131d7c508962
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/configs/train/vitg14.yaml
@@ -0,0 +1,26 @@
+dino:
+ head_n_prototypes: 131072
+ head_bottleneck_dim: 384
+ibot:
+ separate_head: true
+ head_n_prototypes: 131072
+train:
+ batch_size_per_gpu: 12
+ dataset_path: ImageNet22k
+ centering: sinkhorn_knopp
+student:
+ arch: vit_giant2
+ patch_size: 14
+ drop_path_rate: 0.4
+ ffn_layer: swiglufused
+ block_chunks: 4
+teacher:
+ momentum_teacher: 0.994
+optim:
+ epochs: 500
+ weight_decay_end: 0.2
+ base_lr: 2.0e-04 # learning rate for a batch size of 1024
+ warmup_epochs: 80
+ layerwise_decay: 1.0
+crops:
+ local_crops_size: 98
\ No newline at end of file
diff --git a/models/dsp/dinov2/dinov2/configs/train/vitl14.yaml b/models/dsp/dinov2/dinov2/configs/train/vitl14.yaml
new file mode 100644
index 0000000000000000000000000000000000000000..d9b491dcc6a522c71328fc2933dd0501123c8f6b
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/configs/train/vitl14.yaml
@@ -0,0 +1,26 @@
+dino:
+ head_n_prototypes: 131072
+ head_bottleneck_dim: 384
+ibot:
+ separate_head: true
+ head_n_prototypes: 131072
+train:
+ batch_size_per_gpu: 32
+ dataset_path: ImageNet22k
+ centering: sinkhorn_knopp
+student:
+ arch: vit_large
+ patch_size: 14
+ drop_path_rate: 0.4
+ ffn_layer: swiglufused
+ block_chunks: 4
+teacher:
+ momentum_teacher: 0.994
+optim:
+ epochs: 500
+ weight_decay_end: 0.2
+ base_lr: 2.0e-04 # learning rate for a batch size of 1024
+ warmup_epochs: 80
+ layerwise_decay: 1.0
+crops:
+ local_crops_size: 98
\ No newline at end of file
diff --git a/models/dsp/dinov2/dinov2/configs/train/vitl16_short.yaml b/models/dsp/dinov2/dinov2/configs/train/vitl16_short.yaml
new file mode 100644
index 0000000000000000000000000000000000000000..3e7e72864c92175a1354142ac1d64da8070d1e5e
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/configs/train/vitl16_short.yaml
@@ -0,0 +1,6 @@
+# this corresponds to the default config
+train:
+ dataset_path: ImageNet:split=TRAIN
+ batch_size_per_gpu: 64
+student:
+ block_chunks: 4
diff --git a/models/dsp/dinov2/dinov2/distributed/__init__.py b/models/dsp/dinov2/dinov2/distributed/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..23226f4536bf5acf4ffac242e9903d92863b246d
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/distributed/__init__.py
@@ -0,0 +1,270 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import os
+import random
+import re
+import socket
+from typing import Dict, List
+
+import torch
+import torch.distributed as dist
+
+_LOCAL_RANK = -1
+_LOCAL_WORLD_SIZE = -1
+
+
+def is_enabled() -> bool:
+ """
+ Returns:
+ True if distributed training is enabled
+ """
+ return dist.is_available() and dist.is_initialized()
+
+
+def get_global_size() -> int:
+ """
+ Returns:
+ The number of processes in the process group
+ """
+ return dist.get_world_size() if is_enabled() else 1
+
+
+def get_global_rank() -> int:
+ """
+ Returns:
+ The rank of the current process within the global process group.
+ """
+ return dist.get_rank() if is_enabled() else 0
+
+
+def get_local_rank() -> int:
+ """
+ Returns:
+ The rank of the current process within the local (per-machine) process group.
+ """
+ if not is_enabled():
+ return 0
+ assert 0 <= _LOCAL_RANK < _LOCAL_WORLD_SIZE
+ return _LOCAL_RANK
+
+
+def get_local_size() -> int:
+ """
+ Returns:
+ The size of the per-machine process group,
+ i.e. the number of processes per machine.
+ """
+ if not is_enabled():
+ return 1
+ assert 0 <= _LOCAL_RANK < _LOCAL_WORLD_SIZE
+ return _LOCAL_WORLD_SIZE
+
+
+def is_main_process() -> bool:
+ """
+ Returns:
+ True if the current process is the main one.
+ """
+ return get_global_rank() == 0
+
+
+def _restrict_print_to_main_process() -> None:
+ """
+ This function disables printing when not in the main process
+ """
+ import builtins as __builtin__
+
+ builtin_print = __builtin__.print
+
+ def print(*args, **kwargs):
+ force = kwargs.pop("force", False)
+ if is_main_process() or force:
+ builtin_print(*args, **kwargs)
+
+ __builtin__.print = print
+
+
+def _get_master_port(seed: int = 0) -> int:
+ MIN_MASTER_PORT, MAX_MASTER_PORT = (20_000, 60_000)
+
+ master_port_str = os.environ.get("MASTER_PORT")
+ if master_port_str is None:
+ rng = random.Random(seed)
+ return rng.randint(MIN_MASTER_PORT, MAX_MASTER_PORT)
+
+ return int(master_port_str)
+
+
+def _get_available_port() -> int:
+ with socket.socket(socket.AF_INET, socket.SOCK_STREAM) as s:
+ # A "" host address means INADDR_ANY i.e. binding to all interfaces.
+ # Note this is not compatible with IPv6.
+ s.bind(("", 0))
+ port = s.getsockname()[1]
+ return port
+
+
+_TORCH_DISTRIBUTED_ENV_VARS = (
+ "MASTER_ADDR",
+ "MASTER_PORT",
+ "RANK",
+ "WORLD_SIZE",
+ "LOCAL_RANK",
+ "LOCAL_WORLD_SIZE",
+)
+
+
+def _collect_env_vars() -> Dict[str, str]:
+ return {env_var: os.environ[env_var] for env_var in _TORCH_DISTRIBUTED_ENV_VARS if env_var in os.environ}
+
+
+def _is_slurm_job_process() -> bool:
+ return "SLURM_JOB_ID" in os.environ
+
+
+def _parse_slurm_node_list(s: str) -> List[str]:
+ nodes = []
+ # Extract "hostname", "hostname[1-2,3,4-5]," substrings
+ p = re.compile(r"(([^\[]+)(?:\[([^\]]+)\])?),?")
+ for m in p.finditer(s):
+ prefix, suffixes = s[m.start(2) : m.end(2)], s[m.start(3) : m.end(3)]
+ for suffix in suffixes.split(","):
+ span = suffix.split("-")
+ if len(span) == 1:
+ nodes.append(prefix + suffix)
+ else:
+ width = len(span[0])
+ start, end = int(span[0]), int(span[1]) + 1
+ nodes.extend([prefix + f"{i:0{width}}" for i in range(start, end)])
+ return nodes
+
+
+def _check_env_variable(key: str, new_value: str):
+ # Only check for difference with preset environment variables
+ if key in os.environ and os.environ[key] != new_value:
+ raise RuntimeError(f"Cannot export environment variables as {key} is already set")
+
+
+class _TorchDistributedEnvironment:
+ def __init__(self):
+ self.master_addr = "127.0.0.1"
+ self.master_port = 0
+ self.rank = -1
+ self.world_size = -1
+ self.local_rank = -1
+ self.local_world_size = -1
+
+ if _is_slurm_job_process():
+ return self._set_from_slurm_env()
+
+ env_vars = _collect_env_vars()
+ if not env_vars:
+ # Environment is not set
+ pass
+ elif len(env_vars) == len(_TORCH_DISTRIBUTED_ENV_VARS):
+ # Environment is fully set
+ return self._set_from_preset_env()
+ else:
+ # Environment is partially set
+ collected_env_vars = ", ".join(env_vars.keys())
+ raise RuntimeError(f"Partially set environment: {collected_env_vars}")
+
+ if torch.cuda.device_count() > 0:
+ return self._set_from_local()
+
+ raise RuntimeError("Can't initialize PyTorch distributed environment")
+
+ # Slurm job created with sbatch, submitit, etc...
+ def _set_from_slurm_env(self):
+ # logger.info("Initialization from Slurm environment")
+ job_id = int(os.environ["SLURM_JOB_ID"])
+ node_count = int(os.environ["SLURM_JOB_NUM_NODES"])
+ nodes = _parse_slurm_node_list(os.environ["SLURM_JOB_NODELIST"])
+ assert len(nodes) == node_count
+
+ self.master_addr = nodes[0]
+ self.master_port = _get_master_port(seed=job_id)
+ self.rank = int(os.environ["SLURM_PROCID"])
+ self.world_size = int(os.environ["SLURM_NTASKS"])
+ assert self.rank < self.world_size
+ self.local_rank = int(os.environ["SLURM_LOCALID"])
+ self.local_world_size = self.world_size // node_count
+ assert self.local_rank < self.local_world_size
+
+ # Single node job with preset environment (i.e. torchrun)
+ def _set_from_preset_env(self):
+ # logger.info("Initialization from preset environment")
+ self.master_addr = os.environ["MASTER_ADDR"]
+ self.master_port = os.environ["MASTER_PORT"]
+ self.rank = int(os.environ["RANK"])
+ self.world_size = int(os.environ["WORLD_SIZE"])
+ assert self.rank < self.world_size
+ self.local_rank = int(os.environ["LOCAL_RANK"])
+ self.local_world_size = int(os.environ["LOCAL_WORLD_SIZE"])
+ assert self.local_rank < self.local_world_size
+
+ # Single node and GPU job (i.e. local script run)
+ def _set_from_local(self):
+ # logger.info("Initialization from local")
+ self.master_addr = "127.0.0.1"
+ self.master_port = _get_available_port()
+ self.rank = 0
+ self.world_size = 1
+ self.local_rank = 0
+ self.local_world_size = 1
+
+ def export(self, *, overwrite: bool) -> "_TorchDistributedEnvironment":
+ # See the "Environment variable initialization" section from
+ # https://pytorch.org/docs/stable/distributed.html for the complete list of
+ # environment variables required for the env:// initialization method.
+ env_vars = {
+ "MASTER_ADDR": self.master_addr,
+ "MASTER_PORT": str(self.master_port),
+ "RANK": str(self.rank),
+ "WORLD_SIZE": str(self.world_size),
+ "LOCAL_RANK": str(self.local_rank),
+ "LOCAL_WORLD_SIZE": str(self.local_world_size),
+ }
+ if not overwrite:
+ for k, v in env_vars.items():
+ _check_env_variable(k, v)
+
+ os.environ.update(env_vars)
+ return self
+
+
+def enable(*, set_cuda_current_device: bool = True, overwrite: bool = False, allow_nccl_timeout: bool = False):
+ """Enable distributed mode
+
+ Args:
+ set_cuda_current_device: If True, call torch.cuda.set_device() to set the
+ current PyTorch CUDA device to the one matching the local rank.
+ overwrite: If True, overwrites already set variables. Else fails.
+ """
+
+ global _LOCAL_RANK, _LOCAL_WORLD_SIZE
+ if _LOCAL_RANK >= 0 or _LOCAL_WORLD_SIZE >= 0:
+ raise RuntimeError("Distributed mode has already been enabled")
+ torch_env = _TorchDistributedEnvironment()
+ torch_env.export(overwrite=overwrite)
+
+ if set_cuda_current_device:
+ torch.cuda.set_device(torch_env.local_rank)
+
+ if allow_nccl_timeout:
+ # This allows to use torch distributed timeout in a NCCL backend
+ key, value = "NCCL_ASYNC_ERROR_HANDLING", "1"
+ if not overwrite:
+ _check_env_variable(key, value)
+ os.environ[key] = value
+
+ dist.init_process_group(backend="nccl")
+ dist.barrier()
+
+ # Finalize setup
+ _LOCAL_RANK = torch_env.local_rank
+ _LOCAL_WORLD_SIZE = torch_env.local_world_size
+ _restrict_print_to_main_process()
diff --git a/models/dsp/dinov2/dinov2/eval/__init__.py b/models/dsp/dinov2/dinov2/eval/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..b88da6bf80be92af00b72dfdb0a806fa64a7a2d9
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/__init__.py
@@ -0,0 +1,4 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
diff --git a/models/dsp/dinov2/dinov2/eval/depth/__init__.py b/models/dsp/dinov2/dinov2/eval/depth/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..b88da6bf80be92af00b72dfdb0a806fa64a7a2d9
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/depth/__init__.py
@@ -0,0 +1,4 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
diff --git a/models/dsp/dinov2/dinov2/eval/depth/models/__init__.py b/models/dsp/dinov2/dinov2/eval/depth/models/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..9a5825181dc2189424b5c58d245b36919cbc5b2e
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/depth/models/__init__.py
@@ -0,0 +1,10 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .backbones import * # noqa: F403
+from .builder import BACKBONES, DEPTHER, HEADS, LOSSES, build_backbone, build_depther, build_head, build_loss
+from .decode_heads import * # noqa: F403
+from .depther import * # noqa: F403
+from .losses import * # noqa: F403
diff --git a/models/dsp/dinov2/dinov2/eval/depth/models/backbones/__init__.py b/models/dsp/dinov2/dinov2/eval/depth/models/backbones/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..520d75bc6e064b9d64487293604ac1bda6e2b6f7
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/depth/models/backbones/__init__.py
@@ -0,0 +1,6 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .vision_transformer import DinoVisionTransformer
diff --git a/models/dsp/dinov2/dinov2/eval/depth/models/backbones/vision_transformer.py b/models/dsp/dinov2/dinov2/eval/depth/models/backbones/vision_transformer.py
new file mode 100644
index 0000000000000000000000000000000000000000..69bda46fd69eb7dabb8f5b60e6fa459fdc21aeab
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/depth/models/backbones/vision_transformer.py
@@ -0,0 +1,16 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from mmcv.runner import BaseModule
+
+from ..builder import BACKBONES
+
+
+@BACKBONES.register_module()
+class DinoVisionTransformer(BaseModule):
+ """Vision Transformer."""
+
+ def __init__(self, *args, **kwargs):
+ super().__init__()
diff --git a/models/dsp/dinov2/dinov2/eval/depth/models/builder.py b/models/dsp/dinov2/dinov2/eval/depth/models/builder.py
new file mode 100644
index 0000000000000000000000000000000000000000..c152643435308afcff60b07cd68ea979fe1d90cb
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/depth/models/builder.py
@@ -0,0 +1,49 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import warnings
+
+from mmcv.cnn import MODELS as MMCV_MODELS
+from mmcv.cnn.bricks.registry import ATTENTION as MMCV_ATTENTION
+from mmcv.utils import Registry
+
+MODELS = Registry("models", parent=MMCV_MODELS)
+ATTENTION = Registry("attention", parent=MMCV_ATTENTION)
+
+
+BACKBONES = MODELS
+NECKS = MODELS
+HEADS = MODELS
+LOSSES = MODELS
+DEPTHER = MODELS
+
+
+def build_backbone(cfg):
+ """Build backbone."""
+ return BACKBONES.build(cfg)
+
+
+def build_neck(cfg):
+ """Build neck."""
+ return NECKS.build(cfg)
+
+
+def build_head(cfg):
+ """Build head."""
+ return HEADS.build(cfg)
+
+
+def build_loss(cfg):
+ """Build loss."""
+ return LOSSES.build(cfg)
+
+
+def build_depther(cfg, train_cfg=None, test_cfg=None):
+ """Build depther."""
+ if train_cfg is not None or test_cfg is not None:
+ warnings.warn("train_cfg and test_cfg is deprecated, " "please specify them in model", UserWarning)
+ assert cfg.get("train_cfg") is None or train_cfg is None, "train_cfg specified in both outer field and model field "
+ assert cfg.get("test_cfg") is None or test_cfg is None, "test_cfg specified in both outer field and model field "
+ return DEPTHER.build(cfg, default_args=dict(train_cfg=train_cfg, test_cfg=test_cfg))
diff --git a/models/dsp/dinov2/dinov2/eval/depth/models/decode_heads/__init__.py b/models/dsp/dinov2/dinov2/eval/depth/models/decode_heads/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..bd0f0754a5b01d7622c1f26bf3f60daea19da4e8
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/depth/models/decode_heads/__init__.py
@@ -0,0 +1,7 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .dpt_head import DPTHead
+from .linear_head import BNHead
diff --git a/models/dsp/dinov2/dinov2/eval/depth/models/decode_heads/decode_head.py b/models/dsp/dinov2/dinov2/eval/depth/models/decode_heads/decode_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..f8c867a3ec687090b280d90bb86aee435320acda
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/depth/models/decode_heads/decode_head.py
@@ -0,0 +1,225 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import copy
+from abc import ABCMeta, abstractmethod
+
+import mmcv
+import numpy as np
+import torch
+import torch.nn as nn
+from mmcv.runner import BaseModule, auto_fp16, force_fp32
+
+from ...ops import resize
+from ..builder import build_loss
+
+
+class DepthBaseDecodeHead(BaseModule, metaclass=ABCMeta):
+ """Base class for BaseDecodeHead.
+
+ Args:
+ in_channels (List): Input channels.
+ channels (int): Channels after modules, before conv_depth.
+ conv_cfg (dict|None): Config of conv layers. Default: None.
+ act_cfg (dict): Config of activation layers.
+ Default: dict(type='ReLU')
+ loss_decode (dict): Config of decode loss.
+ Default: dict(type='SigLoss').
+ sampler (dict|None): The config of depth map sampler.
+ Default: None.
+ align_corners (bool): align_corners argument of F.interpolate.
+ Default: False.
+ min_depth (int): Min depth in dataset setting.
+ Default: 1e-3.
+ max_depth (int): Max depth in dataset setting.
+ Default: None.
+ norm_cfg (dict|None): Config of norm layers.
+ Default: None.
+ classify (bool): Whether predict depth in a cls.-reg. manner.
+ Default: False.
+ n_bins (int): The number of bins used in cls. step.
+ Default: 256.
+ bins_strategy (str): The discrete strategy used in cls. step.
+ Default: 'UD'.
+ norm_strategy (str): The norm strategy on cls. probability
+ distribution. Default: 'linear'
+ scale_up (str): Whether predict depth in a scale-up manner.
+ Default: False.
+ """
+
+ def __init__(
+ self,
+ in_channels,
+ channels=96,
+ conv_cfg=None,
+ act_cfg=dict(type="ReLU"),
+ loss_decode=dict(type="SigLoss", valid_mask=True, loss_weight=10),
+ sampler=None,
+ align_corners=False,
+ min_depth=1e-3,
+ max_depth=None,
+ norm_cfg=None,
+ classify=False,
+ n_bins=256,
+ bins_strategy="UD",
+ norm_strategy="linear",
+ scale_up=False,
+ ):
+ super(DepthBaseDecodeHead, self).__init__()
+
+ self.in_channels = in_channels
+ self.channels = channels
+ self.conv_cfg = conv_cfg
+ self.act_cfg = act_cfg
+ if isinstance(loss_decode, dict):
+ self.loss_decode = build_loss(loss_decode)
+ elif isinstance(loss_decode, (list, tuple)):
+ self.loss_decode = nn.ModuleList()
+ for loss in loss_decode:
+ self.loss_decode.append(build_loss(loss))
+ self.align_corners = align_corners
+ self.min_depth = min_depth
+ self.max_depth = max_depth
+ self.norm_cfg = norm_cfg
+ self.classify = classify
+ self.n_bins = n_bins
+ self.scale_up = scale_up
+
+ if self.classify:
+ assert bins_strategy in ["UD", "SID"], "Support bins_strategy: UD, SID"
+ assert norm_strategy in ["linear", "softmax", "sigmoid"], "Support norm_strategy: linear, softmax, sigmoid"
+
+ self.bins_strategy = bins_strategy
+ self.norm_strategy = norm_strategy
+ self.softmax = nn.Softmax(dim=1)
+ self.conv_depth = nn.Conv2d(channels, n_bins, kernel_size=3, padding=1, stride=1)
+ else:
+ self.conv_depth = nn.Conv2d(channels, 1, kernel_size=3, padding=1, stride=1)
+
+ self.fp16_enabled = False
+ self.relu = nn.ReLU()
+ self.sigmoid = nn.Sigmoid()
+
+ def extra_repr(self):
+ """Extra repr."""
+ s = f"align_corners={self.align_corners}"
+ return s
+
+ @auto_fp16()
+ @abstractmethod
+ def forward(self, inputs, img_metas):
+ """Placeholder of forward function."""
+ pass
+
+ def forward_train(self, img, inputs, img_metas, depth_gt, train_cfg):
+ """Forward function for training.
+ Args:
+ inputs (list[Tensor]): List of multi-level img features.
+ img_metas (list[dict]): List of image info dict where each dict
+ has: 'img_shape', 'scale_factor', 'flip', and may also contain
+ 'filename', 'ori_shape', 'pad_shape', and 'img_norm_cfg'.
+ For details on the values of these keys see
+ `depth/datasets/pipelines/formatting.py:Collect`.
+ depth_gt (Tensor): GT depth
+ train_cfg (dict): The training config.
+
+ Returns:
+ dict[str, Tensor]: a dictionary of loss components
+ """
+ depth_pred = self.forward(inputs, img_metas)
+ losses = self.losses(depth_pred, depth_gt)
+
+ log_imgs = self.log_images(img[0], depth_pred[0], depth_gt[0], img_metas[0])
+ losses.update(**log_imgs)
+
+ return losses
+
+ def forward_test(self, inputs, img_metas, test_cfg):
+ """Forward function for testing.
+ Args:
+ inputs (list[Tensor]): List of multi-level img features.
+ img_metas (list[dict]): List of image info dict where each dict
+ has: 'img_shape', 'scale_factor', 'flip', and may also contain
+ 'filename', 'ori_shape', 'pad_shape', and 'img_norm_cfg'.
+ For details on the values of these keys see
+ `depth/datasets/pipelines/formatting.py:Collect`.
+ test_cfg (dict): The testing config.
+
+ Returns:
+ Tensor: Output depth map.
+ """
+ return self.forward(inputs, img_metas)
+
+ def depth_pred(self, feat):
+ """Prediction each pixel."""
+ if self.classify:
+ logit = self.conv_depth(feat)
+
+ if self.bins_strategy == "UD":
+ bins = torch.linspace(self.min_depth, self.max_depth, self.n_bins, device=feat.device)
+ elif self.bins_strategy == "SID":
+ bins = torch.logspace(self.min_depth, self.max_depth, self.n_bins, device=feat.device)
+
+ # following Adabins, default linear
+ if self.norm_strategy == "linear":
+ logit = torch.relu(logit)
+ eps = 0.1
+ logit = logit + eps
+ logit = logit / logit.sum(dim=1, keepdim=True)
+ elif self.norm_strategy == "softmax":
+ logit = torch.softmax(logit, dim=1)
+ elif self.norm_strategy == "sigmoid":
+ logit = torch.sigmoid(logit)
+ logit = logit / logit.sum(dim=1, keepdim=True)
+
+ output = torch.einsum("ikmn,k->imn", [logit, bins]).unsqueeze(dim=1)
+
+ else:
+ if self.scale_up:
+ output = self.sigmoid(self.conv_depth(feat)) * self.max_depth
+ else:
+ output = self.relu(self.conv_depth(feat)) + self.min_depth
+ return output
+
+ @force_fp32(apply_to=("depth_pred",))
+ def losses(self, depth_pred, depth_gt):
+ """Compute depth loss."""
+ loss = dict()
+ depth_pred = resize(
+ input=depth_pred, size=depth_gt.shape[2:], mode="bilinear", align_corners=self.align_corners, warning=False
+ )
+ if not isinstance(self.loss_decode, nn.ModuleList):
+ losses_decode = [self.loss_decode]
+ else:
+ losses_decode = self.loss_decode
+ for loss_decode in losses_decode:
+ if loss_decode.loss_name not in loss:
+ loss[loss_decode.loss_name] = loss_decode(depth_pred, depth_gt)
+ else:
+ loss[loss_decode.loss_name] += loss_decode(depth_pred, depth_gt)
+ return loss
+
+ def log_images(self, img_path, depth_pred, depth_gt, img_meta):
+ show_img = copy.deepcopy(img_path.detach().cpu().permute(1, 2, 0))
+ show_img = show_img.numpy().astype(np.float32)
+ show_img = mmcv.imdenormalize(
+ show_img,
+ img_meta["img_norm_cfg"]["mean"],
+ img_meta["img_norm_cfg"]["std"],
+ img_meta["img_norm_cfg"]["to_rgb"],
+ )
+ show_img = np.clip(show_img, 0, 255)
+ show_img = show_img.astype(np.uint8)
+ show_img = show_img[:, :, ::-1]
+ show_img = show_img.transpose(0, 2, 1)
+ show_img = show_img.transpose(1, 0, 2)
+
+ depth_pred = depth_pred / torch.max(depth_pred)
+ depth_gt = depth_gt / torch.max(depth_gt)
+
+ depth_pred_color = copy.deepcopy(depth_pred.detach().cpu())
+ depth_gt_color = copy.deepcopy(depth_gt.detach().cpu())
+
+ return {"img_rgb": show_img, "img_depth_pred": depth_pred_color, "img_depth_gt": depth_gt_color}
diff --git a/models/dsp/dinov2/dinov2/eval/depth/models/decode_heads/dpt_head.py b/models/dsp/dinov2/dinov2/eval/depth/models/decode_heads/dpt_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..c6c6d9470d78e1d944cc505f97865f026a9458d3
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/depth/models/decode_heads/dpt_head.py
@@ -0,0 +1,270 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import math
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule, Linear, build_activation_layer
+from mmcv.runner import BaseModule
+
+from ...ops import resize
+from ..builder import HEADS
+from .decode_head import DepthBaseDecodeHead
+
+
+class Interpolate(nn.Module):
+ def __init__(self, scale_factor, mode, align_corners=False):
+ super(Interpolate, self).__init__()
+ self.interp = nn.functional.interpolate
+ self.scale_factor = scale_factor
+ self.mode = mode
+ self.align_corners = align_corners
+
+ def forward(self, x):
+ x = self.interp(x, scale_factor=self.scale_factor, mode=self.mode, align_corners=self.align_corners)
+ return x
+
+
+class HeadDepth(nn.Module):
+ def __init__(self, features):
+ super(HeadDepth, self).__init__()
+ self.head = nn.Sequential(
+ nn.Conv2d(features, features // 2, kernel_size=3, stride=1, padding=1),
+ Interpolate(scale_factor=2, mode="bilinear", align_corners=True),
+ nn.Conv2d(features // 2, 32, kernel_size=3, stride=1, padding=1),
+ nn.ReLU(),
+ nn.Conv2d(32, 1, kernel_size=1, stride=1, padding=0),
+ )
+
+ def forward(self, x):
+ x = self.head(x)
+ return x
+
+
+class ReassembleBlocks(BaseModule):
+ """ViTPostProcessBlock, process cls_token in ViT backbone output and
+ rearrange the feature vector to feature map.
+ Args:
+ in_channels (int): ViT feature channels. Default: 768.
+ out_channels (List): output channels of each stage.
+ Default: [96, 192, 384, 768].
+ readout_type (str): Type of readout operation. Default: 'ignore'.
+ patch_size (int): The patch size. Default: 16.
+ init_cfg (dict, optional): Initialization config dict. Default: None.
+ """
+
+ def __init__(
+ self, in_channels=768, out_channels=[96, 192, 384, 768], readout_type="ignore", patch_size=16, init_cfg=None
+ ):
+ super(ReassembleBlocks, self).__init__(init_cfg)
+
+ assert readout_type in ["ignore", "add", "project"]
+ self.readout_type = readout_type
+ self.patch_size = patch_size
+
+ self.projects = nn.ModuleList(
+ [
+ ConvModule(
+ in_channels=in_channels,
+ out_channels=out_channel,
+ kernel_size=1,
+ act_cfg=None,
+ )
+ for out_channel in out_channels
+ ]
+ )
+
+ self.resize_layers = nn.ModuleList(
+ [
+ nn.ConvTranspose2d(
+ in_channels=out_channels[0], out_channels=out_channels[0], kernel_size=4, stride=4, padding=0
+ ),
+ nn.ConvTranspose2d(
+ in_channels=out_channels[1], out_channels=out_channels[1], kernel_size=2, stride=2, padding=0
+ ),
+ nn.Identity(),
+ nn.Conv2d(
+ in_channels=out_channels[3], out_channels=out_channels[3], kernel_size=3, stride=2, padding=1
+ ),
+ ]
+ )
+ if self.readout_type == "project":
+ self.readout_projects = nn.ModuleList()
+ for _ in range(len(self.projects)):
+ self.readout_projects.append(
+ nn.Sequential(Linear(2 * in_channels, in_channels), build_activation_layer(dict(type="GELU")))
+ )
+
+ def forward(self, inputs):
+ assert isinstance(inputs, list)
+ out = []
+ for i, x in enumerate(inputs):
+ assert len(x) == 2
+ x, cls_token = x[0], x[1]
+ feature_shape = x.shape
+ if self.readout_type == "project":
+ x = x.flatten(2).permute((0, 2, 1))
+ readout = cls_token.unsqueeze(1).expand_as(x)
+ x = self.readout_projects[i](torch.cat((x, readout), -1))
+ x = x.permute(0, 2, 1).reshape(feature_shape)
+ elif self.readout_type == "add":
+ x = x.flatten(2) + cls_token.unsqueeze(-1)
+ x = x.reshape(feature_shape)
+ else:
+ pass
+ x = self.projects[i](x)
+ x = self.resize_layers[i](x)
+ out.append(x)
+ return out
+
+
+class PreActResidualConvUnit(BaseModule):
+ """ResidualConvUnit, pre-activate residual unit.
+ Args:
+ in_channels (int): number of channels in the input feature map.
+ act_cfg (dict): dictionary to construct and config activation layer.
+ norm_cfg (dict): dictionary to construct and config norm layer.
+ stride (int): stride of the first block. Default: 1
+ dilation (int): dilation rate for convs layers. Default: 1.
+ init_cfg (dict, optional): Initialization config dict. Default: None.
+ """
+
+ def __init__(self, in_channels, act_cfg, norm_cfg, stride=1, dilation=1, init_cfg=None):
+ super(PreActResidualConvUnit, self).__init__(init_cfg)
+
+ self.conv1 = ConvModule(
+ in_channels,
+ in_channels,
+ 3,
+ stride=stride,
+ padding=dilation,
+ dilation=dilation,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg,
+ bias=False,
+ order=("act", "conv", "norm"),
+ )
+
+ self.conv2 = ConvModule(
+ in_channels,
+ in_channels,
+ 3,
+ padding=1,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg,
+ bias=False,
+ order=("act", "conv", "norm"),
+ )
+
+ def forward(self, inputs):
+ inputs_ = inputs.clone()
+ x = self.conv1(inputs)
+ x = self.conv2(x)
+ return x + inputs_
+
+
+class FeatureFusionBlock(BaseModule):
+ """FeatureFusionBlock, merge feature map from different stages.
+ Args:
+ in_channels (int): Input channels.
+ act_cfg (dict): The activation config for ResidualConvUnit.
+ norm_cfg (dict): Config dict for normalization layer.
+ expand (bool): Whether expand the channels in post process block.
+ Default: False.
+ align_corners (bool): align_corner setting for bilinear upsample.
+ Default: True.
+ init_cfg (dict, optional): Initialization config dict. Default: None.
+ """
+
+ def __init__(self, in_channels, act_cfg, norm_cfg, expand=False, align_corners=True, init_cfg=None):
+ super(FeatureFusionBlock, self).__init__(init_cfg)
+
+ self.in_channels = in_channels
+ self.expand = expand
+ self.align_corners = align_corners
+
+ self.out_channels = in_channels
+ if self.expand:
+ self.out_channels = in_channels // 2
+
+ self.project = ConvModule(self.in_channels, self.out_channels, kernel_size=1, act_cfg=None, bias=True)
+
+ self.res_conv_unit1 = PreActResidualConvUnit(in_channels=self.in_channels, act_cfg=act_cfg, norm_cfg=norm_cfg)
+ self.res_conv_unit2 = PreActResidualConvUnit(in_channels=self.in_channels, act_cfg=act_cfg, norm_cfg=norm_cfg)
+
+ def forward(self, *inputs):
+ x = inputs[0]
+ if len(inputs) == 2:
+ if x.shape != inputs[1].shape:
+ res = resize(inputs[1], size=(x.shape[2], x.shape[3]), mode="bilinear", align_corners=False)
+ else:
+ res = inputs[1]
+ x = x + self.res_conv_unit1(res)
+ x = self.res_conv_unit2(x)
+ x = resize(x, scale_factor=2, mode="bilinear", align_corners=self.align_corners)
+ x = self.project(x)
+ return x
+
+
+@HEADS.register_module()
+class DPTHead(DepthBaseDecodeHead):
+ """Vision Transformers for Dense Prediction.
+ This head is implemented of `DPT `_.
+ Args:
+ embed_dims (int): The embed dimension of the ViT backbone.
+ Default: 768.
+ post_process_channels (List): Out channels of post process conv
+ layers. Default: [96, 192, 384, 768].
+ readout_type (str): Type of readout operation. Default: 'ignore'.
+ patch_size (int): The patch size. Default: 16.
+ expand_channels (bool): Whether expand the channels in post process
+ block. Default: False.
+ """
+
+ def __init__(
+ self,
+ embed_dims=768,
+ post_process_channels=[96, 192, 384, 768],
+ readout_type="ignore",
+ patch_size=16,
+ expand_channels=False,
+ **kwargs
+ ):
+ super(DPTHead, self).__init__(**kwargs)
+
+ self.in_channels = self.in_channels
+ self.expand_channels = expand_channels
+ self.reassemble_blocks = ReassembleBlocks(embed_dims, post_process_channels, readout_type, patch_size)
+
+ self.post_process_channels = [
+ channel * math.pow(2, i) if expand_channels else channel for i, channel in enumerate(post_process_channels)
+ ]
+ self.convs = nn.ModuleList()
+ for channel in self.post_process_channels:
+ self.convs.append(ConvModule(channel, self.channels, kernel_size=3, padding=1, act_cfg=None, bias=False))
+ self.fusion_blocks = nn.ModuleList()
+ for _ in range(len(self.convs)):
+ self.fusion_blocks.append(FeatureFusionBlock(self.channels, self.act_cfg, self.norm_cfg))
+ self.fusion_blocks[0].res_conv_unit1 = None
+ self.project = ConvModule(self.channels, self.channels, kernel_size=3, padding=1, norm_cfg=self.norm_cfg)
+ self.num_fusion_blocks = len(self.fusion_blocks)
+ self.num_reassemble_blocks = len(self.reassemble_blocks.resize_layers)
+ self.num_post_process_channels = len(self.post_process_channels)
+ assert self.num_fusion_blocks == self.num_reassemble_blocks
+ assert self.num_reassemble_blocks == self.num_post_process_channels
+ self.conv_depth = HeadDepth(self.channels)
+
+ def forward(self, inputs, img_metas):
+ assert len(inputs) == self.num_reassemble_blocks
+ x = [inp for inp in inputs]
+ x = self.reassemble_blocks(x)
+ x = [self.convs[i](feature) for i, feature in enumerate(x)]
+ out = self.fusion_blocks[0](x[-1])
+ for i in range(1, len(self.fusion_blocks)):
+ out = self.fusion_blocks[i](out, x[-(i + 1)])
+ out = self.project(out)
+ out = self.depth_pred(out)
+ return out
diff --git a/models/dsp/dinov2/dinov2/eval/depth/models/decode_heads/linear_head.py b/models/dsp/dinov2/dinov2/eval/depth/models/decode_heads/linear_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..3da1436f6a3f0bcc389d74ed86d44d455d2f7a87
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/depth/models/decode_heads/linear_head.py
@@ -0,0 +1,89 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import torch
+import torch.nn as nn
+
+from ...ops import resize
+from ..builder import HEADS
+from .decode_head import DepthBaseDecodeHead
+
+
+@HEADS.register_module()
+class BNHead(DepthBaseDecodeHead):
+ """Just a batchnorm."""
+
+ def __init__(self, input_transform="resize_concat", in_index=(0, 1, 2, 3), upsample=1, **kwargs):
+ super().__init__(**kwargs)
+ self.input_transform = input_transform
+ self.in_index = in_index
+ self.upsample = upsample
+ # self.bn = nn.SyncBatchNorm(self.in_channels)
+ if self.classify:
+ self.conv_depth = nn.Conv2d(self.channels, self.n_bins, kernel_size=1, padding=0, stride=1)
+ else:
+ self.conv_depth = nn.Conv2d(self.channels, 1, kernel_size=1, padding=0, stride=1)
+
+ def _transform_inputs(self, inputs):
+ """Transform inputs for decoder.
+ Args:
+ inputs (list[Tensor]): List of multi-level img features.
+ Returns:
+ Tensor: The transformed inputs
+ """
+
+ if "concat" in self.input_transform:
+ inputs = [inputs[i] for i in self.in_index]
+ if "resize" in self.input_transform:
+ inputs = [
+ resize(
+ input=x,
+ size=[s * self.upsample for s in inputs[0].shape[2:]],
+ mode="bilinear",
+ align_corners=self.align_corners,
+ )
+ for x in inputs
+ ]
+ inputs = torch.cat(inputs, dim=1)
+ elif self.input_transform == "multiple_select":
+ inputs = [inputs[i] for i in self.in_index]
+ else:
+ inputs = inputs[self.in_index]
+
+ return inputs
+
+ def _forward_feature(self, inputs, img_metas=None, **kwargs):
+ """Forward function for feature maps before classifying each pixel with
+ ``self.cls_seg`` fc.
+ Args:
+ inputs (list[Tensor]): List of multi-level img features.
+ Returns:
+ feats (Tensor): A tensor of shape (batch_size, self.channels,
+ H, W) which is feature map for last layer of decoder head.
+ """
+ # accept lists (for cls token)
+ inputs = list(inputs)
+ for i, x in enumerate(inputs):
+ if len(x) == 2:
+ x, cls_token = x[0], x[1]
+ if len(x.shape) == 2:
+ x = x[:, :, None, None]
+ cls_token = cls_token[:, :, None, None].expand_as(x)
+ inputs[i] = torch.cat((x, cls_token), 1)
+ else:
+ x = x[0]
+ if len(x.shape) == 2:
+ x = x[:, :, None, None]
+ inputs[i] = x
+ x = self._transform_inputs(inputs)
+ # feats = self.bn(x)
+ return x
+
+ def forward(self, inputs, img_metas=None, **kwargs):
+ """Forward function."""
+ output = self._forward_feature(inputs, img_metas=img_metas, **kwargs)
+ output = self.depth_pred(output)
+
+ return output
diff --git a/models/dsp/dinov2/dinov2/eval/depth/models/depther/__init__.py b/models/dsp/dinov2/dinov2/eval/depth/models/depther/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..be99743bf6c773d05f2b74524116e368c0cfcba0
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/depth/models/depther/__init__.py
@@ -0,0 +1,7 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .base import BaseDepther
+from .encoder_decoder import DepthEncoderDecoder
diff --git a/models/dsp/dinov2/dinov2/eval/depth/models/depther/base.py b/models/dsp/dinov2/dinov2/eval/depth/models/depther/base.py
new file mode 100644
index 0000000000000000000000000000000000000000..e133a825a888167f90d95d67803609d6cac7ff55
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/depth/models/depther/base.py
@@ -0,0 +1,194 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from abc import ABCMeta, abstractmethod
+from collections import OrderedDict
+
+import torch
+import torch.distributed as dist
+from mmcv.runner import BaseModule, auto_fp16
+
+
+class BaseDepther(BaseModule, metaclass=ABCMeta):
+ """Base class for depther."""
+
+ def __init__(self, init_cfg=None):
+ super(BaseDepther, self).__init__(init_cfg)
+ self.fp16_enabled = False
+
+ @property
+ def with_neck(self):
+ """bool: whether the depther has neck"""
+ return hasattr(self, "neck") and self.neck is not None
+
+ @property
+ def with_auxiliary_head(self):
+ """bool: whether the depther has auxiliary head"""
+ return hasattr(self, "auxiliary_head") and self.auxiliary_head is not None
+
+ @property
+ def with_decode_head(self):
+ """bool: whether the depther has decode head"""
+ return hasattr(self, "decode_head") and self.decode_head is not None
+
+ @abstractmethod
+ def extract_feat(self, imgs):
+ """Placeholder for extract features from images."""
+ pass
+
+ @abstractmethod
+ def encode_decode(self, img, img_metas):
+ """Placeholder for encode images with backbone and decode into a
+ semantic depth map of the same size as input."""
+ pass
+
+ @abstractmethod
+ def forward_train(self, imgs, img_metas, **kwargs):
+ """Placeholder for Forward function for training."""
+ pass
+
+ @abstractmethod
+ def simple_test(self, img, img_meta, **kwargs):
+ """Placeholder for single image test."""
+ pass
+
+ @abstractmethod
+ def aug_test(self, imgs, img_metas, **kwargs):
+ """Placeholder for augmentation test."""
+ pass
+
+ def forward_test(self, imgs, img_metas, **kwargs):
+ """
+ Args:
+ imgs (List[Tensor]): the outer list indicates test-time
+ augmentations and inner Tensor should have a shape NxCxHxW,
+ which contains all images in the batch.
+ img_metas (List[List[dict]]): the outer list indicates test-time
+ augs (multiscale, flip, etc.) and the inner list indicates
+ images in a batch.
+ """
+ for var, name in [(imgs, "imgs"), (img_metas, "img_metas")]:
+ if not isinstance(var, list):
+ raise TypeError(f"{name} must be a list, but got " f"{type(var)}")
+ num_augs = len(imgs)
+ if num_augs != len(img_metas):
+ raise ValueError(f"num of augmentations ({len(imgs)}) != " f"num of image meta ({len(img_metas)})")
+ # all images in the same aug batch all of the same ori_shape and pad
+ # shape
+ for img_meta in img_metas:
+ ori_shapes = [_["ori_shape"] for _ in img_meta]
+ assert all(shape == ori_shapes[0] for shape in ori_shapes)
+ img_shapes = [_["img_shape"] for _ in img_meta]
+ assert all(shape == img_shapes[0] for shape in img_shapes)
+ pad_shapes = [_["pad_shape"] for _ in img_meta]
+ assert all(shape == pad_shapes[0] for shape in pad_shapes)
+
+ if num_augs == 1:
+ return self.simple_test(imgs[0], img_metas[0], **kwargs)
+ else:
+ return self.aug_test(imgs, img_metas, **kwargs)
+
+ @auto_fp16(apply_to=("img",))
+ def forward(self, img, img_metas, return_loss=True, **kwargs):
+ """Calls either :func:`forward_train` or :func:`forward_test` depending
+ on whether ``return_loss`` is ``True``.
+
+ Note this setting will change the expected inputs. When
+ ``return_loss=True``, img and img_meta are single-nested (i.e. Tensor
+ and List[dict]), and when ``resturn_loss=False``, img and img_meta
+ should be double nested (i.e. List[Tensor], List[List[dict]]), with
+ the outer list indicating test time augmentations.
+ """
+ if return_loss:
+ return self.forward_train(img, img_metas, **kwargs)
+ else:
+ return self.forward_test(img, img_metas, **kwargs)
+
+ def train_step(self, data_batch, optimizer, **kwargs):
+ """The iteration step during training.
+
+ This method defines an iteration step during training, except for the
+ back propagation and optimizer updating, which are done in an optimizer
+ hook. Note that in some complicated cases or models, the whole process
+ including back propagation and optimizer updating is also defined in
+ this method, such as GAN.
+
+ Args:
+ data (dict): The output of dataloader.
+ optimizer (:obj:`torch.optim.Optimizer` | dict): The optimizer of
+ runner is passed to ``train_step()``. This argument is unused
+ and reserved.
+
+ Returns:
+ dict: It should contain at least 3 keys: ``loss``, ``log_vars``,
+ ``num_samples``.
+ ``loss`` is a tensor for back propagation, which can be a
+ weighted sum of multiple losses.
+ ``log_vars`` contains all the variables to be sent to the
+ logger.
+ ``num_samples`` indicates the batch size (when the model is
+ DDP, it means the batch size on each GPU), which is used for
+ averaging the logs.
+ """
+ losses = self(**data_batch)
+
+ # split losses and images
+ real_losses = {}
+ log_imgs = {}
+ for k, v in losses.items():
+ if "img" in k:
+ log_imgs[k] = v
+ else:
+ real_losses[k] = v
+
+ loss, log_vars = self._parse_losses(real_losses)
+
+ outputs = dict(loss=loss, log_vars=log_vars, num_samples=len(data_batch["img_metas"]), log_imgs=log_imgs)
+
+ return outputs
+
+ def val_step(self, data_batch, **kwargs):
+ """The iteration step during validation.
+
+ This method shares the same signature as :func:`train_step`, but used
+ during val epochs. Note that the evaluation after training epochs is
+ not implemented with this method, but an evaluation hook.
+ """
+ output = self(**data_batch, **kwargs)
+ return output
+
+ @staticmethod
+ def _parse_losses(losses):
+ """Parse the raw outputs (losses) of the network.
+
+ Args:
+ losses (dict): Raw output of the network, which usually contain
+ losses and other necessary information.
+
+ Returns:
+ tuple[Tensor, dict]: (loss, log_vars), loss is the loss tensor
+ which may be a weighted sum of all losses, log_vars contains
+ all the variables to be sent to the logger.
+ """
+ log_vars = OrderedDict()
+ for loss_name, loss_value in losses.items():
+ if isinstance(loss_value, torch.Tensor):
+ log_vars[loss_name] = loss_value.mean()
+ elif isinstance(loss_value, list):
+ log_vars[loss_name] = sum(_loss.mean() for _loss in loss_value)
+ else:
+ raise TypeError(f"{loss_name} is not a tensor or list of tensors")
+
+ loss = sum(_value for _key, _value in log_vars.items() if "loss" in _key)
+
+ log_vars["loss"] = loss
+ for loss_name, loss_value in log_vars.items():
+ # reduce loss when distributed training
+ if dist.is_available() and dist.is_initialized():
+ loss_value = loss_value.data.clone()
+ dist.all_reduce(loss_value.div_(dist.get_world_size()))
+ log_vars[loss_name] = loss_value.item()
+
+ return loss, log_vars
diff --git a/models/dsp/dinov2/dinov2/eval/depth/models/depther/encoder_decoder.py b/models/dsp/dinov2/dinov2/eval/depth/models/depther/encoder_decoder.py
new file mode 100644
index 0000000000000000000000000000000000000000..6b0ec2dd314fdf8ccf4414d81afb95326b7dc0c9
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/depth/models/depther/encoder_decoder.py
@@ -0,0 +1,236 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import torch
+import torch.nn.functional as F
+
+from ...models import builder
+from ...models.builder import DEPTHER
+from ...ops import resize
+from .base import BaseDepther
+
+
+def add_prefix(inputs, prefix):
+ """Add prefix for dict.
+
+ Args:
+ inputs (dict): The input dict with str keys.
+ prefix (str): The prefix to add.
+
+ Returns:
+
+ dict: The dict with keys updated with ``prefix``.
+ """
+
+ outputs = dict()
+ for name, value in inputs.items():
+ outputs[f"{prefix}.{name}"] = value
+
+ return outputs
+
+
+@DEPTHER.register_module()
+class DepthEncoderDecoder(BaseDepther):
+ """Encoder Decoder depther.
+
+ EncoderDecoder typically consists of backbone, (neck) and decode_head.
+ """
+
+ def __init__(self, backbone, decode_head, neck=None, train_cfg=None, test_cfg=None, pretrained=None, init_cfg=None):
+ super(DepthEncoderDecoder, self).__init__(init_cfg)
+ if pretrained is not None:
+ assert backbone.get("pretrained") is None, "both backbone and depther set pretrained weight"
+ backbone.pretrained = pretrained
+ self.backbone = builder.build_backbone(backbone)
+ self._init_decode_head(decode_head)
+
+ if neck is not None:
+ self.neck = builder.build_neck(neck)
+
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+
+ assert self.with_decode_head
+
+ def _init_decode_head(self, decode_head):
+ """Initialize ``decode_head``"""
+ self.decode_head = builder.build_head(decode_head)
+ self.align_corners = self.decode_head.align_corners
+
+ def extract_feat(self, img):
+ """Extract features from images."""
+ x = self.backbone(img)
+ if self.with_neck:
+ x = self.neck(x)
+ return x
+
+ def encode_decode(self, img, img_metas, rescale=True, size=None):
+ """Encode images with backbone and decode into a depth estimation
+ map of the same size as input."""
+ x = self.extract_feat(img)
+ out = self._decode_head_forward_test(x, img_metas)
+ # crop the pred depth to the certain range.
+ out = torch.clamp(out, min=self.decode_head.min_depth, max=self.decode_head.max_depth)
+ if rescale:
+ if size is None:
+ if img_metas is not None:
+ size = img_metas[0]["ori_shape"][:2]
+ else:
+ size = img.shape[2:]
+ out = resize(input=out, size=size, mode="bilinear", align_corners=self.align_corners)
+ return out
+
+ def _decode_head_forward_train(self, img, x, img_metas, depth_gt, **kwargs):
+ """Run forward function and calculate loss for decode head in
+ training."""
+ losses = dict()
+ loss_decode = self.decode_head.forward_train(img, x, img_metas, depth_gt, self.train_cfg, **kwargs)
+ losses.update(add_prefix(loss_decode, "decode"))
+ return losses
+
+ def _decode_head_forward_test(self, x, img_metas):
+ """Run forward function and calculate loss for decode head in
+ inference."""
+ depth_pred = self.decode_head.forward_test(x, img_metas, self.test_cfg)
+ return depth_pred
+
+ def forward_dummy(self, img):
+ """Dummy forward function."""
+ depth = self.encode_decode(img, None)
+
+ return depth
+
+ def forward_train(self, img, img_metas, depth_gt, **kwargs):
+ """Forward function for training.
+
+ Args:
+ img (Tensor): Input images.
+ img_metas (list[dict]): List of image info dict where each dict
+ has: 'img_shape', 'scale_factor', 'flip', and may also contain
+ 'filename', 'ori_shape', 'pad_shape', and 'img_norm_cfg'.
+ For details on the values of these keys see
+ `depth/datasets/pipelines/formatting.py:Collect`.
+ depth_gt (Tensor): Depth gt
+ used if the architecture supports depth estimation task.
+
+ Returns:
+ dict[str, Tensor]: a dictionary of loss components
+ """
+
+ x = self.extract_feat(img)
+
+ losses = dict()
+
+ # the last of x saves the info from neck
+ loss_decode = self._decode_head_forward_train(img, x, img_metas, depth_gt, **kwargs)
+
+ losses.update(loss_decode)
+
+ return losses
+
+ def whole_inference(self, img, img_meta, rescale, size=None):
+ """Inference with full image."""
+ depth_pred = self.encode_decode(img, img_meta, rescale, size=size)
+
+ return depth_pred
+
+ def slide_inference(self, img, img_meta, rescale):
+ """Inference by sliding-window with overlap.
+
+ If h_crop > h_img or w_crop > w_img, the small patch will be used to
+ decode without padding.
+ """
+
+ h_stride, w_stride = self.test_cfg.stride
+ h_crop, w_crop = self.test_cfg.crop_size
+ batch_size, _, h_img, w_img = img.size()
+ h_grids = max(h_img - h_crop + h_stride - 1, 0) // h_stride + 1
+ w_grids = max(w_img - w_crop + w_stride - 1, 0) // w_stride + 1
+ preds = img.new_zeros((batch_size, 1, h_img, w_img))
+ count_mat = img.new_zeros((batch_size, 1, h_img, w_img))
+ for h_idx in range(h_grids):
+ for w_idx in range(w_grids):
+ y1 = h_idx * h_stride
+ x1 = w_idx * w_stride
+ y2 = min(y1 + h_crop, h_img)
+ x2 = min(x1 + w_crop, w_img)
+ y1 = max(y2 - h_crop, 0)
+ x1 = max(x2 - w_crop, 0)
+ crop_img = img[:, :, y1:y2, x1:x2]
+ depth_pred = self.encode_decode(crop_img, img_meta, rescale)
+ preds += F.pad(depth_pred, (int(x1), int(preds.shape[3] - x2), int(y1), int(preds.shape[2] - y2)))
+
+ count_mat[:, :, y1:y2, x1:x2] += 1
+ assert (count_mat == 0).sum() == 0
+ if torch.onnx.is_in_onnx_export():
+ # cast count_mat to constant while exporting to ONNX
+ count_mat = torch.from_numpy(count_mat.cpu().detach().numpy()).to(device=img.device)
+ preds = preds / count_mat
+ return preds
+
+ def inference(self, img, img_meta, rescale, size=None):
+ """Inference with slide/whole style.
+
+ Args:
+ img (Tensor): The input image of shape (N, 3, H, W).
+ img_meta (dict): Image info dict where each dict has: 'img_shape',
+ 'scale_factor', 'flip', and may also contain
+ 'filename', 'ori_shape', 'pad_shape', and 'img_norm_cfg'.
+ For details on the values of these keys see
+ `depth/datasets/pipelines/formatting.py:Collect`.
+ rescale (bool): Whether rescale back to original shape.
+
+ Returns:
+ Tensor: The output depth map.
+ """
+
+ assert self.test_cfg.mode in ["slide", "whole"]
+ ori_shape = img_meta[0]["ori_shape"]
+ assert all(_["ori_shape"] == ori_shape for _ in img_meta)
+ if self.test_cfg.mode == "slide":
+ depth_pred = self.slide_inference(img, img_meta, rescale)
+ else:
+ depth_pred = self.whole_inference(img, img_meta, rescale, size=size)
+ output = depth_pred
+ flip = img_meta[0]["flip"]
+ if flip:
+ flip_direction = img_meta[0]["flip_direction"]
+ assert flip_direction in ["horizontal", "vertical"]
+ if flip_direction == "horizontal":
+ output = output.flip(dims=(3,))
+ elif flip_direction == "vertical":
+ output = output.flip(dims=(2,))
+
+ return output
+
+ def simple_test(self, img, img_meta, rescale=True):
+ """Simple test with single image."""
+ depth_pred = self.inference(img, img_meta, rescale)
+ if torch.onnx.is_in_onnx_export():
+ # our inference backend only support 4D output
+ depth_pred = depth_pred.unsqueeze(0)
+ return depth_pred
+ depth_pred = depth_pred.cpu().numpy()
+ # unravel batch dim
+ depth_pred = list(depth_pred)
+ return depth_pred
+
+ def aug_test(self, imgs, img_metas, rescale=True):
+ """Test with augmentations.
+
+ Only rescale=True is supported.
+ """
+ # aug_test rescale all imgs back to ori_shape for now
+ assert rescale
+ # to save memory, we get augmented depth logit inplace
+ depth_pred = self.inference(imgs[0], img_metas[0], rescale)
+ for i in range(1, len(imgs)):
+ cur_depth_pred = self.inference(imgs[i], img_metas[i], rescale, size=depth_pred.shape[-2:])
+ depth_pred += cur_depth_pred
+ depth_pred /= len(imgs)
+ depth_pred = depth_pred.cpu().numpy()
+ # unravel batch dim
+ depth_pred = list(depth_pred)
+ return depth_pred
diff --git a/models/dsp/dinov2/dinov2/eval/depth/models/losses/__init__.py b/models/dsp/dinov2/dinov2/eval/depth/models/losses/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..2f86242e342776da2e0acc61150d15a8d58ff1e0
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/depth/models/losses/__init__.py
@@ -0,0 +1,7 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .gradientloss import GradientLoss
+from .sigloss import SigLoss
diff --git a/models/dsp/dinov2/dinov2/eval/depth/models/losses/gradientloss.py b/models/dsp/dinov2/dinov2/eval/depth/models/losses/gradientloss.py
new file mode 100644
index 0000000000000000000000000000000000000000..1599878a6b70cdff4f8467e1e875f0d13ea89eca
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/depth/models/losses/gradientloss.py
@@ -0,0 +1,69 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import torch
+import torch.nn as nn
+
+from ...models.builder import LOSSES
+
+
+@LOSSES.register_module()
+class GradientLoss(nn.Module):
+ """GradientLoss.
+
+ Adapted from https://www.cs.cornell.edu/projects/megadepth/
+
+ Args:
+ valid_mask (bool): Whether filter invalid gt (gt > 0). Default: True.
+ loss_weight (float): Weight of the loss. Default: 1.0.
+ max_depth (int): When filtering invalid gt, set a max threshold. Default: None.
+ """
+
+ def __init__(self, valid_mask=True, loss_weight=1.0, max_depth=None, loss_name="loss_grad"):
+ super(GradientLoss, self).__init__()
+ self.valid_mask = valid_mask
+ self.loss_weight = loss_weight
+ self.max_depth = max_depth
+ self.loss_name = loss_name
+
+ self.eps = 0.001 # avoid grad explode
+
+ def gradientloss(self, input, target):
+ input_downscaled = [input] + [input[:: 2 * i, :: 2 * i] for i in range(1, 4)]
+ target_downscaled = [target] + [target[:: 2 * i, :: 2 * i] for i in range(1, 4)]
+
+ gradient_loss = 0
+ for input, target in zip(input_downscaled, target_downscaled):
+ if self.valid_mask:
+ mask = target > 0
+ if self.max_depth is not None:
+ mask = torch.logical_and(target > 0, target <= self.max_depth)
+ N = torch.sum(mask)
+ else:
+ mask = torch.ones_like(target)
+ N = input.numel()
+ input_log = torch.log(input + self.eps)
+ target_log = torch.log(target + self.eps)
+ log_d_diff = input_log - target_log
+
+ log_d_diff = torch.mul(log_d_diff, mask)
+
+ v_gradient = torch.abs(log_d_diff[0:-2, :] - log_d_diff[2:, :])
+ v_mask = torch.mul(mask[0:-2, :], mask[2:, :])
+ v_gradient = torch.mul(v_gradient, v_mask)
+
+ h_gradient = torch.abs(log_d_diff[:, 0:-2] - log_d_diff[:, 2:])
+ h_mask = torch.mul(mask[:, 0:-2], mask[:, 2:])
+ h_gradient = torch.mul(h_gradient, h_mask)
+
+ gradient_loss += (torch.sum(h_gradient) + torch.sum(v_gradient)) / N
+
+ return gradient_loss
+
+ def forward(self, depth_pred, depth_gt):
+ """Forward function."""
+
+ gradient_loss = self.loss_weight * self.gradientloss(depth_pred, depth_gt)
+ return gradient_loss
diff --git a/models/dsp/dinov2/dinov2/eval/depth/models/losses/sigloss.py b/models/dsp/dinov2/dinov2/eval/depth/models/losses/sigloss.py
new file mode 100644
index 0000000000000000000000000000000000000000..e12fad3e6151e4b975dd055193fdaec0206d4a14
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/depth/models/losses/sigloss.py
@@ -0,0 +1,65 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import torch
+import torch.nn as nn
+
+from ...models.builder import LOSSES
+
+
+@LOSSES.register_module()
+class SigLoss(nn.Module):
+ """SigLoss.
+
+ This follows `AdaBins `_.
+
+ Args:
+ valid_mask (bool): Whether filter invalid gt (gt > 0). Default: True.
+ loss_weight (float): Weight of the loss. Default: 1.0.
+ max_depth (int): When filtering invalid gt, set a max threshold. Default: None.
+ warm_up (bool): A simple warm up stage to help convergence. Default: False.
+ warm_iter (int): The number of warm up stage. Default: 100.
+ """
+
+ def __init__(
+ self, valid_mask=True, loss_weight=1.0, max_depth=None, warm_up=False, warm_iter=100, loss_name="sigloss"
+ ):
+ super(SigLoss, self).__init__()
+ self.valid_mask = valid_mask
+ self.loss_weight = loss_weight
+ self.max_depth = max_depth
+ self.loss_name = loss_name
+
+ self.eps = 0.001 # avoid grad explode
+
+ # HACK: a hack implementation for warmup sigloss
+ self.warm_up = warm_up
+ self.warm_iter = warm_iter
+ self.warm_up_counter = 0
+
+ def sigloss(self, input, target):
+ if self.valid_mask:
+ valid_mask = target > 0
+ if self.max_depth is not None:
+ valid_mask = torch.logical_and(target > 0, target <= self.max_depth)
+ input = input[valid_mask]
+ target = target[valid_mask]
+
+ if self.warm_up:
+ if self.warm_up_counter < self.warm_iter:
+ g = torch.log(input + self.eps) - torch.log(target + self.eps)
+ g = 0.15 * torch.pow(torch.mean(g), 2)
+ self.warm_up_counter += 1
+ return torch.sqrt(g)
+
+ g = torch.log(input + self.eps) - torch.log(target + self.eps)
+ Dg = torch.var(g) + 0.15 * torch.pow(torch.mean(g), 2)
+ return torch.sqrt(Dg)
+
+ def forward(self, depth_pred, depth_gt):
+ """Forward function."""
+
+ loss_depth = self.loss_weight * self.sigloss(depth_pred, depth_gt)
+ return loss_depth
diff --git a/models/dsp/dinov2/dinov2/eval/depth/ops/__init__.py b/models/dsp/dinov2/dinov2/eval/depth/ops/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..78181c29581a281b5f42cf12078636aaeb43b5a5
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/depth/ops/__init__.py
@@ -0,0 +1,6 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .wrappers import resize
diff --git a/models/dsp/dinov2/dinov2/eval/depth/ops/wrappers.py b/models/dsp/dinov2/dinov2/eval/depth/ops/wrappers.py
new file mode 100644
index 0000000000000000000000000000000000000000..15880ee0cb7652d4b41c489b927bf6a156b40e5e
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/depth/ops/wrappers.py
@@ -0,0 +1,28 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import warnings
+
+import torch.nn.functional as F
+
+
+def resize(input, size=None, scale_factor=None, mode="nearest", align_corners=None, warning=False):
+ if warning:
+ if size is not None and align_corners:
+ input_h, input_w = tuple(int(x) for x in input.shape[2:])
+ output_h, output_w = tuple(int(x) for x in size)
+ if output_h > input_h or output_w > output_h:
+ if (
+ (output_h > 1 and output_w > 1 and input_h > 1 and input_w > 1)
+ and (output_h - 1) % (input_h - 1)
+ and (output_w - 1) % (input_w - 1)
+ ):
+ warnings.warn(
+ f"When align_corners={align_corners}, "
+ "the output would more aligned if "
+ f"input size {(input_h, input_w)} is `x+1` and "
+ f"out size {(output_h, output_w)} is `nx+1`"
+ )
+ return F.interpolate(input, size, scale_factor, mode, align_corners)
diff --git a/models/dsp/dinov2/dinov2/eval/knn.py b/models/dsp/dinov2/dinov2/eval/knn.py
new file mode 100644
index 0000000000000000000000000000000000000000..f3a4845da1313a6db6b8345bb9a98230fcd24acf
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/knn.py
@@ -0,0 +1,404 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import argparse
+from functools import partial
+import json
+import logging
+import os
+import sys
+from typing import List, Optional
+
+import torch
+from torch.nn.functional import one_hot, softmax
+
+import dinov2.distributed as distributed
+from dinov2.data import SamplerType, make_data_loader, make_dataset
+from dinov2.data.transforms import make_classification_eval_transform
+from dinov2.eval.metrics import AccuracyAveraging, build_topk_accuracy_metric
+from dinov2.eval.setup import get_args_parser as get_setup_args_parser
+from dinov2.eval.setup import setup_and_build_model
+from dinov2.eval.utils import ModelWithNormalize, evaluate, extract_features
+
+
+logger = logging.getLogger("dinov2")
+
+
+def get_args_parser(
+ description: Optional[str] = None,
+ parents: Optional[List[argparse.ArgumentParser]] = None,
+ add_help: bool = True,
+):
+ parents = parents or []
+ setup_args_parser = get_setup_args_parser(parents=parents, add_help=False)
+ parents = [setup_args_parser]
+ parser = argparse.ArgumentParser(
+ description=description,
+ parents=parents,
+ add_help=add_help,
+ )
+ parser.add_argument(
+ "--train-dataset",
+ dest="train_dataset_str",
+ type=str,
+ help="Training dataset",
+ )
+ parser.add_argument(
+ "--val-dataset",
+ dest="val_dataset_str",
+ type=str,
+ help="Validation dataset",
+ )
+ parser.add_argument(
+ "--nb_knn",
+ nargs="+",
+ type=int,
+ help="Number of NN to use. 20 is usually working the best.",
+ )
+ parser.add_argument(
+ "--temperature",
+ type=float,
+ help="Temperature used in the voting coefficient",
+ )
+ parser.add_argument(
+ "--gather-on-cpu",
+ action="store_true",
+ help="Whether to gather the train features on cpu, slower"
+ "but useful to avoid OOM for large datasets (e.g. ImageNet22k).",
+ )
+ parser.add_argument(
+ "--batch-size",
+ type=int,
+ help="Batch size.",
+ )
+ parser.add_argument(
+ "--n-per-class-list",
+ nargs="+",
+ type=int,
+ help="Number to take per class",
+ )
+ parser.add_argument(
+ "--n-tries",
+ type=int,
+ help="Number of tries",
+ )
+ parser.set_defaults(
+ train_dataset_str="ImageNet:split=TRAIN",
+ val_dataset_str="ImageNet:split=VAL",
+ nb_knn=[10, 20, 100, 200],
+ temperature=0.07,
+ batch_size=256,
+ n_per_class_list=[-1],
+ n_tries=1,
+ )
+ return parser
+
+
+class KnnModule(torch.nn.Module):
+ """
+ Gets knn of test features from all processes on a chunk of the train features
+
+ Each rank gets a chunk of the train features as well as a chunk of the test features.
+ In `compute_neighbors`, for each rank one after the other, its chunk of test features
+ is sent to all devices, partial knns are computed with each chunk of train features
+ then collated back on the original device.
+ """
+
+ def __init__(self, train_features, train_labels, nb_knn, T, device, num_classes=1000):
+ super().__init__()
+
+ self.global_rank = distributed.get_global_rank()
+ self.global_size = distributed.get_global_size()
+
+ self.device = device
+ self.train_features_rank_T = train_features.chunk(self.global_size)[self.global_rank].T.to(self.device)
+ self.candidates = train_labels.chunk(self.global_size)[self.global_rank].view(1, -1).to(self.device)
+
+ self.nb_knn = nb_knn
+ self.max_k = max(self.nb_knn)
+ self.T = T
+ self.num_classes = num_classes
+
+ def _get_knn_sims_and_labels(self, similarity, train_labels):
+ topk_sims, indices = similarity.topk(self.max_k, largest=True, sorted=True)
+ neighbors_labels = torch.gather(train_labels, 1, indices)
+ return topk_sims, neighbors_labels
+
+ def _similarity_for_rank(self, features_rank, source_rank):
+ # Send the features from `source_rank` to all ranks
+ broadcast_shape = torch.tensor(features_rank.shape).to(self.device)
+ torch.distributed.broadcast(broadcast_shape, source_rank)
+
+ broadcasted = features_rank
+ if self.global_rank != source_rank:
+ broadcasted = torch.zeros(*broadcast_shape, dtype=features_rank.dtype, device=self.device)
+ torch.distributed.broadcast(broadcasted, source_rank)
+
+ # Compute the neighbors for `source_rank` among `train_features_rank_T`
+ similarity_rank = torch.mm(broadcasted, self.train_features_rank_T)
+ candidate_labels = self.candidates.expand(len(similarity_rank), -1)
+ return self._get_knn_sims_and_labels(similarity_rank, candidate_labels)
+
+ def _gather_all_knn_for_rank(self, topk_sims, neighbors_labels, target_rank):
+ # Gather all neighbors for `target_rank`
+ topk_sims_rank = retrieved_rank = None
+ if self.global_rank == target_rank:
+ topk_sims_rank = [torch.zeros_like(topk_sims) for _ in range(self.global_size)]
+ retrieved_rank = [torch.zeros_like(neighbors_labels) for _ in range(self.global_size)]
+
+ torch.distributed.gather(topk_sims, topk_sims_rank, dst=target_rank)
+ torch.distributed.gather(neighbors_labels, retrieved_rank, dst=target_rank)
+
+ if self.global_rank == target_rank:
+ # Perform a second top-k on the k * global_size retrieved neighbors
+ topk_sims_rank = torch.cat(topk_sims_rank, dim=1)
+ retrieved_rank = torch.cat(retrieved_rank, dim=1)
+ results = self._get_knn_sims_and_labels(topk_sims_rank, retrieved_rank)
+ return results
+ return None
+
+ def compute_neighbors(self, features_rank):
+ for rank in range(self.global_size):
+ topk_sims, neighbors_labels = self._similarity_for_rank(features_rank, rank)
+ results = self._gather_all_knn_for_rank(topk_sims, neighbors_labels, rank)
+ if results is not None:
+ topk_sims_rank, neighbors_labels_rank = results
+ return topk_sims_rank, neighbors_labels_rank
+
+ def forward(self, features_rank):
+ """
+ Compute the results on all values of `self.nb_knn` neighbors from the full `self.max_k`
+ """
+ assert all(k <= self.max_k for k in self.nb_knn)
+
+ topk_sims, neighbors_labels = self.compute_neighbors(features_rank)
+ batch_size = neighbors_labels.shape[0]
+ topk_sims_transform = softmax(topk_sims / self.T, 1)
+ matmul = torch.mul(
+ one_hot(neighbors_labels, num_classes=self.num_classes),
+ topk_sims_transform.view(batch_size, -1, 1),
+ )
+ probas_for_k = {k: torch.sum(matmul[:, :k, :], 1) for k in self.nb_knn}
+ return probas_for_k
+
+
+class DictKeysModule(torch.nn.Module):
+ def __init__(self, keys):
+ super().__init__()
+ self.keys = keys
+
+ def forward(self, features_dict, targets):
+ for k in self.keys:
+ features_dict = features_dict[k]
+ return {"preds": features_dict, "target": targets}
+
+
+def create_module_dict(*, module, n_per_class_list, n_tries, nb_knn, train_features, train_labels):
+ modules = {}
+ mapping = create_class_indices_mapping(train_labels)
+ for npc in n_per_class_list:
+ if npc < 0: # Only one try needed when using the full data
+ full_module = module(
+ train_features=train_features,
+ train_labels=train_labels,
+ nb_knn=nb_knn,
+ )
+ modules["full"] = ModuleDictWithForward({"1": full_module})
+ continue
+ all_tries = {}
+ for t in range(n_tries):
+ final_indices = filter_train(mapping, npc, seed=t)
+ k_list = list(set(nb_knn + [npc]))
+ k_list = sorted([el for el in k_list if el <= npc])
+ all_tries[str(t)] = module(
+ train_features=train_features[final_indices],
+ train_labels=train_labels[final_indices],
+ nb_knn=k_list,
+ )
+ modules[f"{npc} per class"] = ModuleDictWithForward(all_tries)
+
+ return ModuleDictWithForward(modules)
+
+
+def filter_train(mapping, n_per_class, seed):
+ torch.manual_seed(seed)
+ final_indices = []
+ for k in mapping.keys():
+ index = torch.randperm(len(mapping[k]))[:n_per_class]
+ final_indices.append(mapping[k][index])
+ return torch.cat(final_indices).squeeze()
+
+
+def create_class_indices_mapping(labels):
+ unique_labels, inverse = torch.unique(labels, return_inverse=True)
+ mapping = {unique_labels[i]: (inverse == i).nonzero() for i in range(len(unique_labels))}
+ return mapping
+
+
+class ModuleDictWithForward(torch.nn.ModuleDict):
+ def forward(self, *args, **kwargs):
+ return {k: module(*args, **kwargs) for k, module in self._modules.items()}
+
+
+def eval_knn(
+ model,
+ train_dataset,
+ val_dataset,
+ accuracy_averaging,
+ nb_knn,
+ temperature,
+ batch_size,
+ num_workers,
+ gather_on_cpu,
+ n_per_class_list=[-1],
+ n_tries=1,
+):
+ model = ModelWithNormalize(model)
+
+ logger.info("Extracting features for train set...")
+ train_features, train_labels = extract_features(
+ model, train_dataset, batch_size, num_workers, gather_on_cpu=gather_on_cpu
+ )
+ logger.info(f"Train features created, shape {train_features.shape}.")
+
+ val_dataloader = make_data_loader(
+ dataset=val_dataset,
+ batch_size=batch_size,
+ num_workers=num_workers,
+ sampler_type=SamplerType.DISTRIBUTED,
+ drop_last=False,
+ shuffle=False,
+ persistent_workers=True,
+ )
+ num_classes = train_labels.max() + 1
+ metric_collection = build_topk_accuracy_metric(accuracy_averaging, num_classes=num_classes)
+
+ device = torch.cuda.current_device()
+ partial_module = partial(KnnModule, T=temperature, device=device, num_classes=num_classes)
+ knn_module_dict = create_module_dict(
+ module=partial_module,
+ n_per_class_list=n_per_class_list,
+ n_tries=n_tries,
+ nb_knn=nb_knn,
+ train_features=train_features,
+ train_labels=train_labels,
+ )
+ postprocessors, metrics = {}, {}
+ for n_per_class, knn_module in knn_module_dict.items():
+ for t, knn_try in knn_module.items():
+ postprocessors = {
+ **postprocessors,
+ **{(n_per_class, t, k): DictKeysModule([n_per_class, t, k]) for k in knn_try.nb_knn},
+ }
+ metrics = {**metrics, **{(n_per_class, t, k): metric_collection.clone() for k in knn_try.nb_knn}}
+ model_with_knn = torch.nn.Sequential(model, knn_module_dict)
+
+ # ============ evaluation ... ============
+ logger.info("Start the k-NN classification.")
+ _, results_dict = evaluate(model_with_knn, val_dataloader, postprocessors, metrics, device)
+
+ # Averaging the results over the n tries for each value of n_per_class
+ for n_per_class, knn_module in knn_module_dict.items():
+ first_try = list(knn_module.keys())[0]
+ k_list = knn_module[first_try].nb_knn
+ for k in k_list:
+ keys = results_dict[(n_per_class, first_try, k)].keys() # keys are e.g. `top-1` and `top-5`
+ results_dict[(n_per_class, k)] = {
+ key: torch.mean(torch.stack([results_dict[(n_per_class, t, k)][key] for t in knn_module.keys()]))
+ for key in keys
+ }
+ for t in knn_module.keys():
+ del results_dict[(n_per_class, t, k)]
+
+ return results_dict
+
+
+def eval_knn_with_model(
+ model,
+ output_dir,
+ train_dataset_str="ImageNet:split=TRAIN",
+ val_dataset_str="ImageNet:split=VAL",
+ nb_knn=(10, 20, 100, 200),
+ temperature=0.07,
+ autocast_dtype=torch.float,
+ accuracy_averaging=AccuracyAveraging.MEAN_ACCURACY,
+ transform=None,
+ gather_on_cpu=False,
+ batch_size=256,
+ num_workers=5,
+ n_per_class_list=[-1],
+ n_tries=1,
+):
+ transform = transform or make_classification_eval_transform()
+
+ train_dataset = make_dataset(
+ dataset_str=train_dataset_str,
+ transform=transform,
+ )
+ val_dataset = make_dataset(
+ dataset_str=val_dataset_str,
+ transform=transform,
+ )
+
+ with torch.cuda.amp.autocast(dtype=autocast_dtype):
+ results_dict_knn = eval_knn(
+ model=model,
+ train_dataset=train_dataset,
+ val_dataset=val_dataset,
+ accuracy_averaging=accuracy_averaging,
+ nb_knn=nb_knn,
+ temperature=temperature,
+ batch_size=batch_size,
+ num_workers=num_workers,
+ gather_on_cpu=gather_on_cpu,
+ n_per_class_list=n_per_class_list,
+ n_tries=n_tries,
+ )
+
+ results_dict = {}
+ if distributed.is_main_process():
+ for knn_ in results_dict_knn.keys():
+ top1 = results_dict_knn[knn_]["top-1"].item() * 100.0
+ top5 = results_dict_knn[knn_]["top-5"].item() * 100.0
+ results_dict[f"{knn_} Top 1"] = top1
+ results_dict[f"{knn_} Top 5"] = top5
+ logger.info(f"{knn_} classifier result: Top1: {top1:.2f} Top5: {top5:.2f}")
+
+ metrics_file_path = os.path.join(output_dir, "results_eval_knn.json")
+ with open(metrics_file_path, "a") as f:
+ for k, v in results_dict.items():
+ f.write(json.dumps({k: v}) + "\n")
+
+ if distributed.is_enabled():
+ torch.distributed.barrier()
+ return results_dict
+
+
+def main(args):
+ model, autocast_dtype = setup_and_build_model(args)
+ eval_knn_with_model(
+ model=model,
+ output_dir=args.output_dir,
+ train_dataset_str=args.train_dataset_str,
+ val_dataset_str=args.val_dataset_str,
+ nb_knn=args.nb_knn,
+ temperature=args.temperature,
+ autocast_dtype=autocast_dtype,
+ accuracy_averaging=AccuracyAveraging.MEAN_ACCURACY,
+ transform=None,
+ gather_on_cpu=args.gather_on_cpu,
+ batch_size=args.batch_size,
+ num_workers=5,
+ n_per_class_list=args.n_per_class_list,
+ n_tries=args.n_tries,
+ )
+ return 0
+
+
+if __name__ == "__main__":
+ description = "DINOv2 k-NN evaluation"
+ args_parser = get_args_parser(description=description)
+ args = args_parser.parse_args()
+ sys.exit(main(args))
diff --git a/models/dsp/dinov2/dinov2/eval/linear.py b/models/dsp/dinov2/dinov2/eval/linear.py
new file mode 100644
index 0000000000000000000000000000000000000000..1bd4c5de5a041be8a188f007257d1e91b6d6921e
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/linear.py
@@ -0,0 +1,625 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import argparse
+from functools import partial
+import json
+import logging
+import os
+import sys
+from typing import List, Optional
+
+import numpy as np
+import torch
+import torch.nn as nn
+from torch.nn.parallel import DistributedDataParallel
+from fvcore.common.checkpoint import Checkpointer, PeriodicCheckpointer
+
+from dinov2.data import SamplerType, make_data_loader, make_dataset
+from dinov2.data.transforms import make_classification_eval_transform, make_classification_train_transform
+import dinov2.distributed as distributed
+from dinov2.eval.metrics import MetricType, build_metric
+from dinov2.eval.setup import get_args_parser as get_setup_args_parser
+from dinov2.eval.setup import setup_and_build_model
+from dinov2.eval.utils import ModelWithIntermediateLayers, evaluate
+from dinov2.logging import MetricLogger
+
+
+logger = logging.getLogger("dinov2")
+
+
+def get_args_parser(
+ description: Optional[str] = None,
+ parents: Optional[List[argparse.ArgumentParser]] = None,
+ add_help: bool = True,
+):
+ parents = parents or []
+ setup_args_parser = get_setup_args_parser(parents=parents, add_help=False)
+ parents = [setup_args_parser]
+ parser = argparse.ArgumentParser(
+ description=description,
+ parents=parents,
+ add_help=add_help,
+ )
+ parser.add_argument(
+ "--train-dataset",
+ dest="train_dataset_str",
+ type=str,
+ help="Training dataset",
+ )
+ parser.add_argument(
+ "--val-dataset",
+ dest="val_dataset_str",
+ type=str,
+ help="Validation dataset",
+ )
+ parser.add_argument(
+ "--test-datasets",
+ dest="test_dataset_strs",
+ type=str,
+ nargs="+",
+ help="Test datasets, none to reuse the validation dataset",
+ )
+ parser.add_argument(
+ "--epochs",
+ type=int,
+ help="Number of training epochs",
+ )
+ parser.add_argument(
+ "--batch-size",
+ type=int,
+ help="Batch Size (per GPU)",
+ )
+ parser.add_argument(
+ "--num-workers",
+ type=int,
+ help="Number de Workers",
+ )
+ parser.add_argument(
+ "--epoch-length",
+ type=int,
+ help="Length of an epoch in number of iterations",
+ )
+ parser.add_argument(
+ "--save-checkpoint-frequency",
+ type=int,
+ help="Number of epochs between two named checkpoint saves.",
+ )
+ parser.add_argument(
+ "--eval-period-iterations",
+ type=int,
+ help="Number of iterations between two evaluations.",
+ )
+ parser.add_argument(
+ "--learning-rates",
+ nargs="+",
+ type=float,
+ help="Learning rates to grid search.",
+ )
+ parser.add_argument(
+ "--no-resume",
+ action="store_true",
+ help="Whether to not resume from existing checkpoints",
+ )
+ parser.add_argument(
+ "--val-metric-type",
+ type=MetricType,
+ choices=list(MetricType),
+ help="Validation metric",
+ )
+ parser.add_argument(
+ "--test-metric-types",
+ type=MetricType,
+ choices=list(MetricType),
+ nargs="+",
+ help="Evaluation metric",
+ )
+ parser.add_argument(
+ "--classifier-fpath",
+ type=str,
+ help="Path to a file containing pretrained linear classifiers",
+ )
+ parser.add_argument(
+ "--val-class-mapping-fpath",
+ type=str,
+ help="Path to a file containing a mapping to adjust classifier outputs",
+ )
+ parser.add_argument(
+ "--test-class-mapping-fpaths",
+ nargs="+",
+ type=str,
+ help="Path to a file containing a mapping to adjust classifier outputs",
+ )
+ parser.set_defaults(
+ train_dataset_str="ImageNet:split=TRAIN",
+ val_dataset_str="ImageNet:split=VAL",
+ test_dataset_strs=None,
+ epochs=10,
+ batch_size=128,
+ num_workers=8,
+ epoch_length=1250,
+ save_checkpoint_frequency=20,
+ eval_period_iterations=1250,
+ learning_rates=[1e-5, 2e-5, 5e-5, 1e-4, 2e-4, 5e-4, 1e-3, 2e-3, 5e-3, 1e-2, 2e-2, 5e-2, 0.1],
+ val_metric_type=MetricType.MEAN_ACCURACY,
+ test_metric_types=None,
+ classifier_fpath=None,
+ val_class_mapping_fpath=None,
+ test_class_mapping_fpaths=[None],
+ )
+ return parser
+
+
+def has_ddp_wrapper(m: nn.Module) -> bool:
+ return isinstance(m, DistributedDataParallel)
+
+
+def remove_ddp_wrapper(m: nn.Module) -> nn.Module:
+ return m.module if has_ddp_wrapper(m) else m
+
+
+def _pad_and_collate(batch):
+ maxlen = max(len(targets) for image, targets in batch)
+ padded_batch = [
+ (image, np.pad(targets, (0, maxlen - len(targets)), constant_values=-1)) for image, targets in batch
+ ]
+ return torch.utils.data.default_collate(padded_batch)
+
+
+def create_linear_input(x_tokens_list, use_n_blocks, use_avgpool):
+ intermediate_output = x_tokens_list[-use_n_blocks:]
+ output = torch.cat([class_token for _, class_token in intermediate_output], dim=-1)
+ if use_avgpool:
+ output = torch.cat(
+ (
+ output,
+ torch.mean(intermediate_output[-1][0], dim=1), # patch tokens
+ ),
+ dim=-1,
+ )
+ output = output.reshape(output.shape[0], -1)
+ return output.float()
+
+
+class LinearClassifier(nn.Module):
+ """Linear layer to train on top of frozen features"""
+
+ def __init__(self, out_dim, use_n_blocks, use_avgpool, num_classes=1000):
+ super().__init__()
+ self.out_dim = out_dim
+ self.use_n_blocks = use_n_blocks
+ self.use_avgpool = use_avgpool
+ self.num_classes = num_classes
+ self.linear = nn.Linear(out_dim, num_classes)
+ self.linear.weight.data.normal_(mean=0.0, std=0.01)
+ self.linear.bias.data.zero_()
+
+ def forward(self, x_tokens_list):
+ output = create_linear_input(x_tokens_list, self.use_n_blocks, self.use_avgpool)
+ return self.linear(output)
+
+
+class AllClassifiers(nn.Module):
+ def __init__(self, classifiers_dict):
+ super().__init__()
+ self.classifiers_dict = nn.ModuleDict()
+ self.classifiers_dict.update(classifiers_dict)
+
+ def forward(self, inputs):
+ return {k: v.forward(inputs) for k, v in self.classifiers_dict.items()}
+
+ def __len__(self):
+ return len(self.classifiers_dict)
+
+
+class LinearPostprocessor(nn.Module):
+ def __init__(self, linear_classifier, class_mapping=None):
+ super().__init__()
+ self.linear_classifier = linear_classifier
+ self.register_buffer("class_mapping", None if class_mapping is None else torch.LongTensor(class_mapping))
+
+ def forward(self, samples, targets):
+ preds = self.linear_classifier(samples)
+ return {
+ "preds": preds[:, self.class_mapping] if self.class_mapping is not None else preds,
+ "target": targets,
+ }
+
+
+def scale_lr(learning_rates, batch_size):
+ return learning_rates * (batch_size * distributed.get_global_size()) / 256.0
+
+
+def setup_linear_classifiers(sample_output, n_last_blocks_list, learning_rates, batch_size, num_classes=1000):
+ linear_classifiers_dict = nn.ModuleDict()
+ optim_param_groups = []
+ for n in n_last_blocks_list:
+ for avgpool in [False, True]:
+ for _lr in learning_rates:
+ lr = scale_lr(_lr, batch_size)
+ out_dim = create_linear_input(sample_output, use_n_blocks=n, use_avgpool=avgpool).shape[1]
+ linear_classifier = LinearClassifier(
+ out_dim, use_n_blocks=n, use_avgpool=avgpool, num_classes=num_classes
+ )
+ linear_classifier = linear_classifier.cuda()
+ linear_classifiers_dict[
+ f"classifier_{n}_blocks_avgpool_{avgpool}_lr_{lr:.5f}".replace(".", "_")
+ ] = linear_classifier
+ optim_param_groups.append({"params": linear_classifier.parameters(), "lr": lr})
+
+ linear_classifiers = AllClassifiers(linear_classifiers_dict)
+ if distributed.is_enabled():
+ linear_classifiers = nn.parallel.DistributedDataParallel(linear_classifiers)
+
+ return linear_classifiers, optim_param_groups
+
+
+@torch.no_grad()
+def evaluate_linear_classifiers(
+ feature_model,
+ linear_classifiers,
+ data_loader,
+ metric_type,
+ metrics_file_path,
+ training_num_classes,
+ iteration,
+ prefixstring="",
+ class_mapping=None,
+ best_classifier_on_val=None,
+):
+ logger.info("running validation !")
+
+ num_classes = len(class_mapping) if class_mapping is not None else training_num_classes
+ metric = build_metric(metric_type, num_classes=num_classes)
+ postprocessors = {k: LinearPostprocessor(v, class_mapping) for k, v in linear_classifiers.classifiers_dict.items()}
+ metrics = {k: metric.clone() for k in linear_classifiers.classifiers_dict}
+
+ _, results_dict_temp = evaluate(
+ feature_model,
+ data_loader,
+ postprocessors,
+ metrics,
+ torch.cuda.current_device(),
+ )
+
+ logger.info("")
+ results_dict = {}
+ max_accuracy = 0
+ best_classifier = ""
+ for i, (classifier_string, metric) in enumerate(results_dict_temp.items()):
+ logger.info(f"{prefixstring} -- Classifier: {classifier_string} * {metric}")
+ if (
+ best_classifier_on_val is None and metric["top-1"].item() > max_accuracy
+ ) or classifier_string == best_classifier_on_val:
+ max_accuracy = metric["top-1"].item()
+ best_classifier = classifier_string
+
+ results_dict["best_classifier"] = {"name": best_classifier, "accuracy": max_accuracy}
+
+ logger.info(f"best classifier: {results_dict['best_classifier']}")
+
+ if distributed.is_main_process():
+ with open(metrics_file_path, "a") as f:
+ f.write(f"iter: {iteration}\n")
+ for k, v in results_dict.items():
+ f.write(json.dumps({k: v}) + "\n")
+ f.write("\n")
+
+ return results_dict
+
+
+def eval_linear(
+ *,
+ feature_model,
+ linear_classifiers,
+ train_data_loader,
+ val_data_loader,
+ metrics_file_path,
+ optimizer,
+ scheduler,
+ output_dir,
+ max_iter,
+ checkpoint_period, # In number of iter, creates a new file every period
+ running_checkpoint_period, # Period to update main checkpoint file
+ eval_period,
+ metric_type,
+ training_num_classes,
+ resume=True,
+ classifier_fpath=None,
+ val_class_mapping=None,
+):
+ checkpointer = Checkpointer(linear_classifiers, output_dir, optimizer=optimizer, scheduler=scheduler)
+ start_iter = checkpointer.resume_or_load(classifier_fpath or "", resume=resume).get("iteration", -1) + 1
+
+ periodic_checkpointer = PeriodicCheckpointer(checkpointer, checkpoint_period, max_iter=max_iter)
+ iteration = start_iter
+ logger.info("Starting training from iteration {}".format(start_iter))
+ metric_logger = MetricLogger(delimiter=" ")
+ header = "Training"
+
+ for data, labels in metric_logger.log_every(
+ train_data_loader,
+ 10,
+ header,
+ max_iter,
+ start_iter,
+ ):
+ data = data.cuda(non_blocking=True)
+ labels = labels.cuda(non_blocking=True)
+
+ features = feature_model(data)
+ outputs = linear_classifiers(features)
+
+ losses = {f"loss_{k}": nn.CrossEntropyLoss()(v, labels) for k, v in outputs.items()}
+ loss = sum(losses.values())
+
+ # compute the gradients
+ optimizer.zero_grad()
+ loss.backward()
+
+ # step
+ optimizer.step()
+ scheduler.step()
+
+ # log
+ if iteration % 10 == 0:
+ torch.cuda.synchronize()
+ metric_logger.update(loss=loss.item())
+ metric_logger.update(lr=optimizer.param_groups[0]["lr"])
+ print("lr", optimizer.param_groups[0]["lr"])
+
+ if iteration - start_iter > 5:
+ if iteration % running_checkpoint_period == 0:
+ torch.cuda.synchronize()
+ if distributed.is_main_process():
+ logger.info("Checkpointing running_checkpoint")
+ periodic_checkpointer.save("running_checkpoint_linear_eval", iteration=iteration)
+ torch.cuda.synchronize()
+ periodic_checkpointer.step(iteration)
+
+ if eval_period > 0 and (iteration + 1) % eval_period == 0 and iteration != max_iter - 1:
+ _ = evaluate_linear_classifiers(
+ feature_model=feature_model,
+ linear_classifiers=remove_ddp_wrapper(linear_classifiers),
+ data_loader=val_data_loader,
+ metrics_file_path=metrics_file_path,
+ prefixstring=f"ITER: {iteration}",
+ metric_type=metric_type,
+ training_num_classes=training_num_classes,
+ iteration=iteration,
+ class_mapping=val_class_mapping,
+ )
+ torch.cuda.synchronize()
+
+ iteration = iteration + 1
+
+ val_results_dict = evaluate_linear_classifiers(
+ feature_model=feature_model,
+ linear_classifiers=remove_ddp_wrapper(linear_classifiers),
+ data_loader=val_data_loader,
+ metrics_file_path=metrics_file_path,
+ metric_type=metric_type,
+ training_num_classes=training_num_classes,
+ iteration=iteration,
+ class_mapping=val_class_mapping,
+ )
+ return val_results_dict, feature_model, linear_classifiers, iteration
+
+
+def make_eval_data_loader(test_dataset_str, batch_size, num_workers, metric_type):
+ test_dataset = make_dataset(
+ dataset_str=test_dataset_str,
+ transform=make_classification_eval_transform(),
+ )
+ test_data_loader = make_data_loader(
+ dataset=test_dataset,
+ batch_size=batch_size,
+ num_workers=num_workers,
+ sampler_type=SamplerType.DISTRIBUTED,
+ drop_last=False,
+ shuffle=False,
+ persistent_workers=False,
+ collate_fn=_pad_and_collate if metric_type == MetricType.IMAGENET_REAL_ACCURACY else None,
+ )
+ return test_data_loader
+
+
+def test_on_datasets(
+ feature_model,
+ linear_classifiers,
+ test_dataset_strs,
+ batch_size,
+ num_workers,
+ test_metric_types,
+ metrics_file_path,
+ training_num_classes,
+ iteration,
+ best_classifier_on_val,
+ prefixstring="",
+ test_class_mappings=[None],
+):
+ results_dict = {}
+ for test_dataset_str, class_mapping, metric_type in zip(test_dataset_strs, test_class_mappings, test_metric_types):
+ logger.info(f"Testing on {test_dataset_str}")
+ test_data_loader = make_eval_data_loader(test_dataset_str, batch_size, num_workers, metric_type)
+ dataset_results_dict = evaluate_linear_classifiers(
+ feature_model,
+ remove_ddp_wrapper(linear_classifiers),
+ test_data_loader,
+ metric_type,
+ metrics_file_path,
+ training_num_classes,
+ iteration,
+ prefixstring="",
+ class_mapping=class_mapping,
+ best_classifier_on_val=best_classifier_on_val,
+ )
+ results_dict[f"{test_dataset_str}_accuracy"] = 100.0 * dataset_results_dict["best_classifier"]["accuracy"]
+ return results_dict
+
+
+def run_eval_linear(
+ model,
+ output_dir,
+ train_dataset_str,
+ val_dataset_str,
+ batch_size,
+ epochs,
+ epoch_length,
+ num_workers,
+ save_checkpoint_frequency,
+ eval_period_iterations,
+ learning_rates,
+ autocast_dtype,
+ test_dataset_strs=None,
+ resume=True,
+ classifier_fpath=None,
+ val_class_mapping_fpath=None,
+ test_class_mapping_fpaths=[None],
+ val_metric_type=MetricType.MEAN_ACCURACY,
+ test_metric_types=None,
+):
+ seed = 0
+
+ if test_dataset_strs is None:
+ test_dataset_strs = [val_dataset_str]
+ if test_metric_types is None:
+ test_metric_types = [val_metric_type] * len(test_dataset_strs)
+ else:
+ assert len(test_metric_types) == len(test_dataset_strs)
+ assert len(test_dataset_strs) == len(test_class_mapping_fpaths)
+
+ train_transform = make_classification_train_transform()
+ train_dataset = make_dataset(
+ dataset_str=train_dataset_str,
+ transform=train_transform,
+ )
+ training_num_classes = len(torch.unique(torch.Tensor(train_dataset.get_targets().astype(int))))
+ sampler_type = SamplerType.SHARDED_INFINITE
+ # sampler_type = SamplerType.INFINITE
+
+ n_last_blocks_list = [1, 4]
+ n_last_blocks = max(n_last_blocks_list)
+ autocast_ctx = partial(torch.cuda.amp.autocast, enabled=True, dtype=autocast_dtype)
+ feature_model = ModelWithIntermediateLayers(model, n_last_blocks, autocast_ctx)
+ sample_output = feature_model(train_dataset[0][0].unsqueeze(0).cuda())
+
+ linear_classifiers, optim_param_groups = setup_linear_classifiers(
+ sample_output,
+ n_last_blocks_list,
+ learning_rates,
+ batch_size,
+ training_num_classes,
+ )
+
+ optimizer = torch.optim.SGD(optim_param_groups, momentum=0.9, weight_decay=0)
+ max_iter = epochs * epoch_length
+ scheduler = torch.optim.lr_scheduler.CosineAnnealingLR(optimizer, max_iter, eta_min=0)
+ checkpointer = Checkpointer(linear_classifiers, output_dir, optimizer=optimizer, scheduler=scheduler)
+ start_iter = checkpointer.resume_or_load(classifier_fpath or "", resume=resume).get("iteration", -1) + 1
+ train_data_loader = make_data_loader(
+ dataset=train_dataset,
+ batch_size=batch_size,
+ num_workers=num_workers,
+ shuffle=True,
+ seed=seed,
+ sampler_type=sampler_type,
+ sampler_advance=start_iter,
+ drop_last=True,
+ persistent_workers=True,
+ )
+ val_data_loader = make_eval_data_loader(val_dataset_str, batch_size, num_workers, val_metric_type)
+
+ checkpoint_period = save_checkpoint_frequency * epoch_length
+
+ if val_class_mapping_fpath is not None:
+ logger.info(f"Using class mapping from {val_class_mapping_fpath}")
+ val_class_mapping = np.load(val_class_mapping_fpath)
+ else:
+ val_class_mapping = None
+
+ test_class_mappings = []
+ for class_mapping_fpath in test_class_mapping_fpaths:
+ if class_mapping_fpath is not None and class_mapping_fpath != "None":
+ logger.info(f"Using class mapping from {class_mapping_fpath}")
+ class_mapping = np.load(class_mapping_fpath)
+ else:
+ class_mapping = None
+ test_class_mappings.append(class_mapping)
+
+ metrics_file_path = os.path.join(output_dir, "results_eval_linear.json")
+ val_results_dict, feature_model, linear_classifiers, iteration = eval_linear(
+ feature_model=feature_model,
+ linear_classifiers=linear_classifiers,
+ train_data_loader=train_data_loader,
+ val_data_loader=val_data_loader,
+ metrics_file_path=metrics_file_path,
+ optimizer=optimizer,
+ scheduler=scheduler,
+ output_dir=output_dir,
+ max_iter=max_iter,
+ checkpoint_period=checkpoint_period,
+ running_checkpoint_period=epoch_length,
+ eval_period=eval_period_iterations,
+ metric_type=val_metric_type,
+ training_num_classes=training_num_classes,
+ resume=resume,
+ val_class_mapping=val_class_mapping,
+ classifier_fpath=classifier_fpath,
+ )
+ results_dict = {}
+ if len(test_dataset_strs) > 1 or test_dataset_strs[0] != val_dataset_str:
+ results_dict = test_on_datasets(
+ feature_model,
+ linear_classifiers,
+ test_dataset_strs,
+ batch_size,
+ 0, # num_workers,
+ test_metric_types,
+ metrics_file_path,
+ training_num_classes,
+ iteration,
+ val_results_dict["best_classifier"]["name"],
+ prefixstring="",
+ test_class_mappings=test_class_mappings,
+ )
+ results_dict["best_classifier"] = val_results_dict["best_classifier"]["name"]
+ results_dict[f"{val_dataset_str}_accuracy"] = 100.0 * val_results_dict["best_classifier"]["accuracy"]
+ logger.info("Test Results Dict " + str(results_dict))
+
+ return results_dict
+
+
+def main(args):
+ model, autocast_dtype = setup_and_build_model(args)
+ run_eval_linear(
+ model=model,
+ output_dir=args.output_dir,
+ train_dataset_str=args.train_dataset_str,
+ val_dataset_str=args.val_dataset_str,
+ test_dataset_strs=args.test_dataset_strs,
+ batch_size=args.batch_size,
+ epochs=args.epochs,
+ epoch_length=args.epoch_length,
+ num_workers=args.num_workers,
+ save_checkpoint_frequency=args.save_checkpoint_frequency,
+ eval_period_iterations=args.eval_period_iterations,
+ learning_rates=args.learning_rates,
+ autocast_dtype=autocast_dtype,
+ resume=not args.no_resume,
+ classifier_fpath=args.classifier_fpath,
+ val_metric_type=args.val_metric_type,
+ test_metric_types=args.test_metric_types,
+ val_class_mapping_fpath=args.val_class_mapping_fpath,
+ test_class_mapping_fpaths=args.test_class_mapping_fpaths,
+ )
+ return 0
+
+
+if __name__ == "__main__":
+ description = "DINOv2 linear evaluation"
+ args_parser = get_args_parser(description=description)
+ args = args_parser.parse_args()
+ sys.exit(main(args))
diff --git a/models/dsp/dinov2/dinov2/eval/log_regression.py b/models/dsp/dinov2/dinov2/eval/log_regression.py
new file mode 100644
index 0000000000000000000000000000000000000000..5f36ec134e0ce25697428a0b3f21cdc2f0145645
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/log_regression.py
@@ -0,0 +1,444 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import argparse
+import gc
+import logging
+import sys
+import time
+from typing import List, Optional
+
+from cuml.linear_model import LogisticRegression
+import torch
+import torch.backends.cudnn as cudnn
+import torch.distributed
+from torch import nn
+from torch.utils.data import TensorDataset
+from torchmetrics import MetricTracker
+
+from dinov2.data import make_dataset
+from dinov2.data.transforms import make_classification_eval_transform
+from dinov2.distributed import get_global_rank, get_global_size
+from dinov2.eval.metrics import MetricType, build_metric
+from dinov2.eval.setup import get_args_parser as get_setup_args_parser
+from dinov2.eval.setup import setup_and_build_model
+from dinov2.eval.utils import evaluate, extract_features
+from dinov2.utils.dtype import as_torch_dtype
+
+
+logger = logging.getLogger("dinov2")
+
+DEFAULT_MAX_ITER = 1_000
+C_POWER_RANGE = torch.linspace(-6, 5, 45)
+_CPU_DEVICE = torch.device("cpu")
+
+
+def get_args_parser(
+ description: Optional[str] = None,
+ parents: Optional[List[argparse.ArgumentParser]] = None,
+ add_help: bool = True,
+):
+ parents = parents or []
+ setup_args_parser = get_setup_args_parser(parents=parents, add_help=False)
+ parents = [setup_args_parser]
+ parser = argparse.ArgumentParser(
+ description=description,
+ parents=parents,
+ add_help=add_help,
+ )
+ parser.add_argument(
+ "--train-dataset",
+ dest="train_dataset_str",
+ type=str,
+ help="Training dataset",
+ )
+ parser.add_argument(
+ "--val-dataset",
+ dest="val_dataset_str",
+ type=str,
+ help="Validation dataset",
+ )
+ parser.add_argument(
+ "--finetune-dataset-str",
+ dest="finetune_dataset_str",
+ type=str,
+ help="Fine-tuning dataset",
+ )
+ parser.add_argument(
+ "--finetune-on-val",
+ action="store_true",
+ help="If there is no finetune dataset, whether to choose the "
+ "hyperparameters on the val set instead of 10%% of the train dataset",
+ )
+ parser.add_argument(
+ "--metric-type",
+ type=MetricType,
+ choices=list(MetricType),
+ help="Metric type",
+ )
+ parser.add_argument(
+ "--train-features-device",
+ type=str,
+ help="Device to gather train features (cpu, cuda, cuda:0, etc.), default: %(default)s",
+ )
+ parser.add_argument(
+ "--train-dtype",
+ type=str,
+ help="Data type to convert the train features to (default: %(default)s)",
+ )
+ parser.add_argument(
+ "--max-train-iters",
+ type=int,
+ help="Maximum number of train iterations (default: %(default)s)",
+ )
+ parser.set_defaults(
+ train_dataset_str="ImageNet:split=TRAIN",
+ val_dataset_str="ImageNet:split=VAL",
+ finetune_dataset_str=None,
+ metric_type=MetricType.MEAN_ACCURACY,
+ train_features_device="cpu",
+ train_dtype="float64",
+ max_train_iters=DEFAULT_MAX_ITER,
+ finetune_on_val=False,
+ )
+ return parser
+
+
+class LogRegModule(nn.Module):
+ def __init__(
+ self,
+ C,
+ max_iter=DEFAULT_MAX_ITER,
+ dtype=torch.float64,
+ device=_CPU_DEVICE,
+ ):
+ super().__init__()
+ self.dtype = dtype
+ self.device = device
+ self.estimator = LogisticRegression(
+ penalty="l2",
+ C=C,
+ max_iter=max_iter,
+ output_type="numpy",
+ tol=1e-12,
+ linesearch_max_iter=50,
+ )
+
+ def forward(self, samples, targets):
+ samples_device = samples.device
+ samples = samples.to(dtype=self.dtype, device=self.device)
+ if self.device == _CPU_DEVICE:
+ samples = samples.numpy()
+ probas = self.estimator.predict_proba(samples)
+ return {"preds": torch.from_numpy(probas).to(samples_device), "target": targets}
+
+ def fit(self, train_features, train_labels):
+ train_features = train_features.to(dtype=self.dtype, device=self.device)
+ train_labels = train_labels.to(dtype=self.dtype, device=self.device)
+ if self.device == _CPU_DEVICE:
+ # both cuML and sklearn only work with numpy arrays on CPU
+ train_features = train_features.numpy()
+ train_labels = train_labels.numpy()
+ self.estimator.fit(train_features, train_labels)
+
+
+def evaluate_model(*, logreg_model, logreg_metric, test_data_loader, device):
+ postprocessors = {"metrics": logreg_model}
+ metrics = {"metrics": logreg_metric}
+ return evaluate(nn.Identity(), test_data_loader, postprocessors, metrics, device)
+
+
+def train_for_C(*, C, max_iter, train_features, train_labels, dtype=torch.float64, device=_CPU_DEVICE):
+ logreg_model = LogRegModule(C, max_iter=max_iter, dtype=dtype, device=device)
+ logreg_model.fit(train_features, train_labels)
+ return logreg_model
+
+
+def train_and_evaluate(
+ *,
+ C,
+ max_iter,
+ train_features,
+ train_labels,
+ logreg_metric,
+ test_data_loader,
+ train_dtype=torch.float64,
+ train_features_device,
+ eval_device,
+):
+ logreg_model = train_for_C(
+ C=C,
+ max_iter=max_iter,
+ train_features=train_features,
+ train_labels=train_labels,
+ dtype=train_dtype,
+ device=train_features_device,
+ )
+ return evaluate_model(
+ logreg_model=logreg_model,
+ logreg_metric=logreg_metric,
+ test_data_loader=test_data_loader,
+ device=eval_device,
+ )
+
+
+def sweep_C_values(
+ *,
+ train_features,
+ train_labels,
+ test_data_loader,
+ metric_type,
+ num_classes,
+ train_dtype=torch.float64,
+ train_features_device=_CPU_DEVICE,
+ max_train_iters=DEFAULT_MAX_ITER,
+):
+ if metric_type == MetricType.PER_CLASS_ACCURACY:
+ # If we want to output per-class accuracy, we select the hyperparameters with mean per class
+ metric_type = MetricType.MEAN_PER_CLASS_ACCURACY
+ logreg_metric = build_metric(metric_type, num_classes=num_classes)
+ metric_tracker = MetricTracker(logreg_metric, maximize=True)
+ ALL_C = 10**C_POWER_RANGE
+ logreg_models = {}
+
+ train_features = train_features.to(dtype=train_dtype, device=train_features_device)
+ train_labels = train_labels.to(device=train_features_device)
+
+ for i in range(get_global_rank(), len(ALL_C), get_global_size()):
+ C = ALL_C[i].item()
+ logger.info(
+ f"Training for C = {C:.5f}, dtype={train_dtype}, "
+ f"features: {train_features.shape}, {train_features.dtype}, "
+ f"labels: {train_labels.shape}, {train_labels.dtype}"
+ )
+ logreg_models[C] = train_for_C(
+ C=C,
+ max_iter=max_train_iters,
+ train_features=train_features,
+ train_labels=train_labels,
+ dtype=train_dtype,
+ device=train_features_device,
+ )
+
+ gather_list = [None for _ in range(get_global_size())]
+ torch.distributed.all_gather_object(gather_list, logreg_models)
+
+ logreg_models_gathered = {}
+ for logreg_dict in gather_list:
+ logreg_models_gathered.update(logreg_dict)
+
+ for i in range(len(ALL_C)):
+ metric_tracker.increment()
+ C = ALL_C[i].item()
+ evals = evaluate_model(
+ logreg_model=logreg_models_gathered[C],
+ logreg_metric=metric_tracker,
+ test_data_loader=test_data_loader,
+ device=torch.cuda.current_device(),
+ )
+ logger.info(f"Trained for C = {C:.5f}, accuracies = {evals}")
+
+ best_stats, which_epoch = metric_tracker.best_metric(return_step=True)
+ best_stats_100 = {k: 100.0 * v for k, v in best_stats.items()}
+ if which_epoch["top-1"] == i:
+ best_C = C
+ logger.info(f"Sweep best {best_stats_100}, best C = {best_C:.6f}")
+
+ return best_stats, best_C
+
+
+def eval_log_regression(
+ *,
+ model,
+ train_dataset,
+ val_dataset,
+ finetune_dataset,
+ metric_type,
+ batch_size,
+ num_workers,
+ finetune_on_val=False,
+ train_dtype=torch.float64,
+ train_features_device=_CPU_DEVICE,
+ max_train_iters=DEFAULT_MAX_ITER,
+):
+ """
+ Implements the "standard" process for log regression evaluation:
+ The value of C is chosen by training on train_dataset and evaluating on
+ finetune_dataset. Then, the final model is trained on a concatenation of
+ train_dataset and finetune_dataset, and is evaluated on val_dataset.
+ If there is no finetune_dataset, the value of C is the one that yields
+ the best results on a random 10% subset of the train dataset
+ """
+
+ start = time.time()
+
+ train_features, train_labels = extract_features(
+ model, train_dataset, batch_size, num_workers, gather_on_cpu=(train_features_device == _CPU_DEVICE)
+ )
+ val_features, val_labels = extract_features(
+ model, val_dataset, batch_size, num_workers, gather_on_cpu=(train_features_device == _CPU_DEVICE)
+ )
+ val_data_loader = torch.utils.data.DataLoader(
+ TensorDataset(val_features, val_labels),
+ batch_size=batch_size,
+ drop_last=False,
+ num_workers=0,
+ persistent_workers=False,
+ )
+
+ if finetune_dataset is None and finetune_on_val:
+ logger.info("Choosing hyperparameters on the val dataset")
+ finetune_features, finetune_labels = val_features, val_labels
+ elif finetune_dataset is None and not finetune_on_val:
+ logger.info("Choosing hyperparameters on 10% of the train dataset")
+ torch.manual_seed(0)
+ indices = torch.randperm(len(train_features), device=train_features.device)
+ finetune_index = indices[: len(train_features) // 10]
+ train_index = indices[len(train_features) // 10 :]
+ finetune_features, finetune_labels = train_features[finetune_index], train_labels[finetune_index]
+ train_features, train_labels = train_features[train_index], train_labels[train_index]
+ else:
+ logger.info("Choosing hyperparameters on the finetune dataset")
+ finetune_features, finetune_labels = extract_features(
+ model, finetune_dataset, batch_size, num_workers, gather_on_cpu=(train_features_device == _CPU_DEVICE)
+ )
+ # release the model - free GPU memory
+ del model
+ gc.collect()
+ torch.cuda.empty_cache()
+ finetune_data_loader = torch.utils.data.DataLoader(
+ TensorDataset(finetune_features, finetune_labels),
+ batch_size=batch_size,
+ drop_last=False,
+ )
+
+ if len(train_labels.shape) > 1:
+ num_classes = train_labels.shape[1]
+ else:
+ num_classes = train_labels.max() + 1
+
+ logger.info("Using cuML for logistic regression")
+
+ best_stats, best_C = sweep_C_values(
+ train_features=train_features,
+ train_labels=train_labels,
+ test_data_loader=finetune_data_loader,
+ metric_type=metric_type,
+ num_classes=num_classes,
+ train_dtype=train_dtype,
+ train_features_device=train_features_device,
+ max_train_iters=max_train_iters,
+ )
+
+ if not finetune_on_val:
+ logger.info("Best parameter found, concatenating features")
+ train_features = torch.cat((train_features, finetune_features))
+ train_labels = torch.cat((train_labels, finetune_labels))
+
+ logger.info("Training final model")
+ logreg_metric = build_metric(metric_type, num_classes=num_classes)
+ evals = train_and_evaluate(
+ C=best_C,
+ max_iter=max_train_iters,
+ train_features=train_features,
+ train_labels=train_labels,
+ logreg_metric=logreg_metric.clone(),
+ test_data_loader=val_data_loader,
+ eval_device=torch.cuda.current_device(),
+ train_dtype=train_dtype,
+ train_features_device=train_features_device,
+ )
+
+ best_stats = evals[1]["metrics"]
+
+ best_stats["best_C"] = best_C
+
+ logger.info(f"Log regression evaluation done in {int(time.time() - start)}s")
+ return best_stats
+
+
+def eval_log_regression_with_model(
+ model,
+ train_dataset_str="ImageNet:split=TRAIN",
+ val_dataset_str="ImageNet:split=VAL",
+ finetune_dataset_str=None,
+ autocast_dtype=torch.float,
+ finetune_on_val=False,
+ metric_type=MetricType.MEAN_ACCURACY,
+ train_dtype=torch.float64,
+ train_features_device=_CPU_DEVICE,
+ max_train_iters=DEFAULT_MAX_ITER,
+):
+ cudnn.benchmark = True
+
+ transform = make_classification_eval_transform(resize_size=224)
+ target_transform = None
+
+ train_dataset = make_dataset(dataset_str=train_dataset_str, transform=transform, target_transform=target_transform)
+ val_dataset = make_dataset(dataset_str=val_dataset_str, transform=transform, target_transform=target_transform)
+ if finetune_dataset_str is not None:
+ finetune_dataset = make_dataset(
+ dataset_str=finetune_dataset_str, transform=transform, target_transform=target_transform
+ )
+ else:
+ finetune_dataset = None
+
+ with torch.cuda.amp.autocast(dtype=autocast_dtype):
+ results_dict_logreg = eval_log_regression(
+ model=model,
+ train_dataset=train_dataset,
+ val_dataset=val_dataset,
+ finetune_dataset=finetune_dataset,
+ metric_type=metric_type,
+ batch_size=256,
+ num_workers=0, # 5,
+ finetune_on_val=finetune_on_val,
+ train_dtype=train_dtype,
+ train_features_device=train_features_device,
+ max_train_iters=max_train_iters,
+ )
+
+ results_dict = {
+ "top-1": results_dict_logreg["top-1"].cpu().numpy() * 100.0,
+ "top-5": results_dict_logreg.get("top-5", torch.tensor(0.0)).cpu().numpy() * 100.0,
+ "best_C": results_dict_logreg["best_C"],
+ }
+ logger.info(
+ "\n".join(
+ [
+ "Training of the supervised logistic regression on frozen features completed.\n"
+ "Top-1 test accuracy: {acc:.1f}".format(acc=results_dict["top-1"]),
+ "Top-5 test accuracy: {acc:.1f}".format(acc=results_dict["top-5"]),
+ "obtained for C = {c:.6f}".format(c=results_dict["best_C"]),
+ ]
+ )
+ )
+
+ torch.distributed.barrier()
+ return results_dict
+
+
+def main(args):
+ model, autocast_dtype = setup_and_build_model(args)
+ eval_log_regression_with_model(
+ model=model,
+ train_dataset_str=args.train_dataset_str,
+ val_dataset_str=args.val_dataset_str,
+ finetune_dataset_str=args.finetune_dataset_str,
+ autocast_dtype=autocast_dtype,
+ finetune_on_val=args.finetune_on_val,
+ metric_type=args.metric_type,
+ train_dtype=as_torch_dtype(args.train_dtype),
+ train_features_device=torch.device(args.train_features_device),
+ max_train_iters=args.max_train_iters,
+ )
+ return 0
+
+
+if __name__ == "__main__":
+ description = "DINOv2 logistic regression evaluation"
+ args_parser = get_args_parser(description=description)
+ args = args_parser.parse_args()
+ sys.exit(main(args))
diff --git a/models/dsp/dinov2/dinov2/eval/metrics.py b/models/dsp/dinov2/dinov2/eval/metrics.py
new file mode 100644
index 0000000000000000000000000000000000000000..52be81a859dddde82da93c3657c35352d2bb0a48
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/metrics.py
@@ -0,0 +1,113 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from enum import Enum
+import logging
+from typing import Any, Dict, Optional
+
+import torch
+from torch import Tensor
+from torchmetrics import Metric, MetricCollection
+from torchmetrics.classification import MulticlassAccuracy
+from torchmetrics.utilities.data import dim_zero_cat, select_topk
+
+
+logger = logging.getLogger("dinov2")
+
+
+class MetricType(Enum):
+ MEAN_ACCURACY = "mean_accuracy"
+ MEAN_PER_CLASS_ACCURACY = "mean_per_class_accuracy"
+ PER_CLASS_ACCURACY = "per_class_accuracy"
+ IMAGENET_REAL_ACCURACY = "imagenet_real_accuracy"
+
+ @property
+ def accuracy_averaging(self):
+ return getattr(AccuracyAveraging, self.name, None)
+
+ def __str__(self):
+ return self.value
+
+
+class AccuracyAveraging(Enum):
+ MEAN_ACCURACY = "micro"
+ MEAN_PER_CLASS_ACCURACY = "macro"
+ PER_CLASS_ACCURACY = "none"
+
+ def __str__(self):
+ return self.value
+
+
+def build_metric(metric_type: MetricType, *, num_classes: int, ks: Optional[tuple] = None):
+ if metric_type.accuracy_averaging is not None:
+ return build_topk_accuracy_metric(
+ average_type=metric_type.accuracy_averaging,
+ num_classes=num_classes,
+ ks=(1, 5) if ks is None else ks,
+ )
+ elif metric_type == MetricType.IMAGENET_REAL_ACCURACY:
+ return build_topk_imagenet_real_accuracy_metric(
+ num_classes=num_classes,
+ ks=(1, 5) if ks is None else ks,
+ )
+
+ raise ValueError(f"Unknown metric type {metric_type}")
+
+
+def build_topk_accuracy_metric(average_type: AccuracyAveraging, num_classes: int, ks: tuple = (1, 5)):
+ metrics: Dict[str, Metric] = {
+ f"top-{k}": MulticlassAccuracy(top_k=k, num_classes=int(num_classes), average=average_type.value) for k in ks
+ }
+ return MetricCollection(metrics)
+
+
+def build_topk_imagenet_real_accuracy_metric(num_classes: int, ks: tuple = (1, 5)):
+ metrics: Dict[str, Metric] = {f"top-{k}": ImageNetReaLAccuracy(top_k=k, num_classes=int(num_classes)) for k in ks}
+ return MetricCollection(metrics)
+
+
+class ImageNetReaLAccuracy(Metric):
+ is_differentiable: bool = False
+ higher_is_better: Optional[bool] = None
+ full_state_update: bool = False
+
+ def __init__(
+ self,
+ num_classes: int,
+ top_k: int = 1,
+ **kwargs: Any,
+ ) -> None:
+ super().__init__(**kwargs)
+ self.num_classes = num_classes
+ self.top_k = top_k
+ self.add_state("tp", [], dist_reduce_fx="cat")
+
+ def update(self, preds: Tensor, target: Tensor) -> None: # type: ignore
+ # preds [B, D]
+ # target [B, A]
+ # preds_oh [B, D] with 0 and 1
+ # select top K highest probabilities, use one hot representation
+ preds_oh = select_topk(preds, self.top_k)
+ # target_oh [B, D + 1] with 0 and 1
+ target_oh = torch.zeros((preds_oh.shape[0], preds_oh.shape[1] + 1), device=target.device, dtype=torch.int32)
+ target = target.long()
+ # for undefined targets (-1) use a fake value `num_classes`
+ target[target == -1] = self.num_classes
+ # fill targets, use one hot representation
+ target_oh.scatter_(1, target, 1)
+ # target_oh [B, D] (remove the fake target at index `num_classes`)
+ target_oh = target_oh[:, :-1]
+ # tp [B] with 0 and 1
+ tp = (preds_oh * target_oh == 1).sum(dim=1)
+ # at least one match between prediction and target
+ tp.clip_(max=1)
+ # ignore instances where no targets are defined
+ mask = target_oh.sum(dim=1) > 0
+ tp = tp[mask]
+ self.tp.append(tp) # type: ignore
+
+ def compute(self) -> Tensor:
+ tp = dim_zero_cat(self.tp) # type: ignore
+ return tp.float().mean()
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..b88da6bf80be92af00b72dfdb0a806fa64a7a2d9
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation/__init__.py
@@ -0,0 +1,4 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation/hooks/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation/hooks/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..738cc2d2069521ea0353acd0cb0a03e3ddf1fa51
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation/hooks/__init__.py
@@ -0,0 +1,6 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .optimizer import DistOptimizerHook
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation/hooks/optimizer.py b/models/dsp/dinov2/dinov2/eval/segmentation/hooks/optimizer.py
new file mode 100644
index 0000000000000000000000000000000000000000..f593f26a84475bbf7ebda9607a4d10914b13a443
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation/hooks/optimizer.py
@@ -0,0 +1,40 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+try:
+ import apex
+except ImportError:
+ print("apex is not installed")
+
+from mmcv.runner import OptimizerHook, HOOKS
+
+
+@HOOKS.register_module()
+class DistOptimizerHook(OptimizerHook):
+ """Optimizer hook for distributed training."""
+
+ def __init__(self, update_interval=1, grad_clip=None, coalesce=True, bucket_size_mb=-1, use_fp16=False):
+ self.grad_clip = grad_clip
+ self.coalesce = coalesce
+ self.bucket_size_mb = bucket_size_mb
+ self.update_interval = update_interval
+ self.use_fp16 = use_fp16
+
+ def before_run(self, runner):
+ runner.optimizer.zero_grad()
+
+ def after_train_iter(self, runner):
+ runner.outputs["loss"] /= self.update_interval
+ if self.use_fp16:
+ # runner.outputs['loss'].backward()
+ with apex.amp.scale_loss(runner.outputs["loss"], runner.optimizer) as scaled_loss:
+ scaled_loss.backward()
+ else:
+ runner.outputs["loss"].backward()
+ if self.every_n_iters(runner, self.update_interval):
+ if self.grad_clip is not None:
+ self.clip_grads(runner.model.parameters())
+ runner.optimizer.step()
+ runner.optimizer.zero_grad()
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation/models/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation/models/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..88e4563d4c162d67e7900955a06bd9248d4c9a48
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation/models/__init__.py
@@ -0,0 +1,7 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .backbones import * # noqa: F403
+from .decode_heads import * # noqa: F403
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation/models/backbones/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation/models/backbones/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..520d75bc6e064b9d64487293604ac1bda6e2b6f7
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation/models/backbones/__init__.py
@@ -0,0 +1,6 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .vision_transformer import DinoVisionTransformer
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation/models/backbones/vision_transformer.py b/models/dsp/dinov2/dinov2/eval/segmentation/models/backbones/vision_transformer.py
new file mode 100644
index 0000000000000000000000000000000000000000..c3e9753ae92a36be52f100e3004cbeeff777d14a
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation/models/backbones/vision_transformer.py
@@ -0,0 +1,19 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from mmcv.runner import BaseModule
+from mmseg.models.builder import BACKBONES
+
+
+@BACKBONES.register_module()
+class DinoVisionTransformer(BaseModule):
+ """Vision Transformer."""
+
+ def __init__(
+ self,
+ *args,
+ **kwargs,
+ ):
+ super().__init__()
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation/models/decode_heads/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation/models/decode_heads/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..c55317875262dadf8970c2b3882f016b8d4731ac
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation/models/decode_heads/__init__.py
@@ -0,0 +1,6 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .linear_head import BNHead
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation/models/decode_heads/linear_head.py b/models/dsp/dinov2/dinov2/eval/segmentation/models/decode_heads/linear_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..d1f39c68fb136f84d1aa5284da5b69581bb177cc
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation/models/decode_heads/linear_head.py
@@ -0,0 +1,90 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import torch
+import torch.nn as nn
+
+from mmseg.models.builder import HEADS
+from mmseg.models.decode_heads.decode_head import BaseDecodeHead
+from mmseg.ops import resize
+
+
+@HEADS.register_module()
+class BNHead(BaseDecodeHead):
+ """Just a batchnorm."""
+
+ def __init__(self, resize_factors=None, **kwargs):
+ super().__init__(**kwargs)
+ assert self.in_channels == self.channels
+ self.bn = nn.SyncBatchNorm(self.in_channels)
+ self.resize_factors = resize_factors
+
+ def _forward_feature(self, inputs):
+ """Forward function for feature maps before classifying each pixel with
+ ``self.cls_seg`` fc.
+
+ Args:
+ inputs (list[Tensor]): List of multi-level img features.
+
+ Returns:
+ feats (Tensor): A tensor of shape (batch_size, self.channels,
+ H, W) which is feature map for last layer of decoder head.
+ """
+ # print("inputs", [i.shape for i in inputs])
+ x = self._transform_inputs(inputs)
+ # print("x", x.shape)
+ feats = self.bn(x)
+ # print("feats", feats.shape)
+ return feats
+
+ def _transform_inputs(self, inputs):
+ """Transform inputs for decoder.
+ Args:
+ inputs (list[Tensor]): List of multi-level img features.
+ Returns:
+ Tensor: The transformed inputs
+ """
+
+ if self.input_transform == "resize_concat":
+ # accept lists (for cls token)
+ input_list = []
+ for x in inputs:
+ if isinstance(x, list):
+ input_list.extend(x)
+ else:
+ input_list.append(x)
+ inputs = input_list
+ # an image descriptor can be a local descriptor with resolution 1x1
+ for i, x in enumerate(inputs):
+ if len(x.shape) == 2:
+ inputs[i] = x[:, :, None, None]
+ # select indices
+ inputs = [inputs[i] for i in self.in_index]
+ # Resizing shenanigans
+ # print("before", *(x.shape for x in inputs))
+ if self.resize_factors is not None:
+ assert len(self.resize_factors) == len(inputs), (len(self.resize_factors), len(inputs))
+ inputs = [
+ resize(input=x, scale_factor=f, mode="bilinear" if f >= 1 else "area")
+ for x, f in zip(inputs, self.resize_factors)
+ ]
+ # print("after", *(x.shape for x in inputs))
+ upsampled_inputs = [
+ resize(input=x, size=inputs[0].shape[2:], mode="bilinear", align_corners=self.align_corners)
+ for x in inputs
+ ]
+ inputs = torch.cat(upsampled_inputs, dim=1)
+ elif self.input_transform == "multiple_select":
+ inputs = [inputs[i] for i in self.in_index]
+ else:
+ inputs = inputs[self.in_index]
+
+ return inputs
+
+ def forward(self, inputs):
+ """Forward function."""
+ output = self._forward_feature(inputs)
+ output = self.cls_seg(output)
+ return output
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation/utils/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation/utils/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..b88da6bf80be92af00b72dfdb0a806fa64a7a2d9
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation/utils/__init__.py
@@ -0,0 +1,4 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation/utils/colormaps.py b/models/dsp/dinov2/dinov2/eval/segmentation/utils/colormaps.py
new file mode 100644
index 0000000000000000000000000000000000000000..e6ef604b2c75792e95e438abfd51ab03d40de340
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation/utils/colormaps.py
@@ -0,0 +1,362 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+ADE20K_COLORMAP = [
+ (0, 0, 0),
+ (120, 120, 120),
+ (180, 120, 120),
+ (6, 230, 230),
+ (80, 50, 50),
+ (4, 200, 3),
+ (120, 120, 80),
+ (140, 140, 140),
+ (204, 5, 255),
+ (230, 230, 230),
+ (4, 250, 7),
+ (224, 5, 255),
+ (235, 255, 7),
+ (150, 5, 61),
+ (120, 120, 70),
+ (8, 255, 51),
+ (255, 6, 82),
+ (143, 255, 140),
+ (204, 255, 4),
+ (255, 51, 7),
+ (204, 70, 3),
+ (0, 102, 200),
+ (61, 230, 250),
+ (255, 6, 51),
+ (11, 102, 255),
+ (255, 7, 71),
+ (255, 9, 224),
+ (9, 7, 230),
+ (220, 220, 220),
+ (255, 9, 92),
+ (112, 9, 255),
+ (8, 255, 214),
+ (7, 255, 224),
+ (255, 184, 6),
+ (10, 255, 71),
+ (255, 41, 10),
+ (7, 255, 255),
+ (224, 255, 8),
+ (102, 8, 255),
+ (255, 61, 6),
+ (255, 194, 7),
+ (255, 122, 8),
+ (0, 255, 20),
+ (255, 8, 41),
+ (255, 5, 153),
+ (6, 51, 255),
+ (235, 12, 255),
+ (160, 150, 20),
+ (0, 163, 255),
+ (140, 140, 140),
+ (250, 10, 15),
+ (20, 255, 0),
+ (31, 255, 0),
+ (255, 31, 0),
+ (255, 224, 0),
+ (153, 255, 0),
+ (0, 0, 255),
+ (255, 71, 0),
+ (0, 235, 255),
+ (0, 173, 255),
+ (31, 0, 255),
+ (11, 200, 200),
+ (255, 82, 0),
+ (0, 255, 245),
+ (0, 61, 255),
+ (0, 255, 112),
+ (0, 255, 133),
+ (255, 0, 0),
+ (255, 163, 0),
+ (255, 102, 0),
+ (194, 255, 0),
+ (0, 143, 255),
+ (51, 255, 0),
+ (0, 82, 255),
+ (0, 255, 41),
+ (0, 255, 173),
+ (10, 0, 255),
+ (173, 255, 0),
+ (0, 255, 153),
+ (255, 92, 0),
+ (255, 0, 255),
+ (255, 0, 245),
+ (255, 0, 102),
+ (255, 173, 0),
+ (255, 0, 20),
+ (255, 184, 184),
+ (0, 31, 255),
+ (0, 255, 61),
+ (0, 71, 255),
+ (255, 0, 204),
+ (0, 255, 194),
+ (0, 255, 82),
+ (0, 10, 255),
+ (0, 112, 255),
+ (51, 0, 255),
+ (0, 194, 255),
+ (0, 122, 255),
+ (0, 255, 163),
+ (255, 153, 0),
+ (0, 255, 10),
+ (255, 112, 0),
+ (143, 255, 0),
+ (82, 0, 255),
+ (163, 255, 0),
+ (255, 235, 0),
+ (8, 184, 170),
+ (133, 0, 255),
+ (0, 255, 92),
+ (184, 0, 255),
+ (255, 0, 31),
+ (0, 184, 255),
+ (0, 214, 255),
+ (255, 0, 112),
+ (92, 255, 0),
+ (0, 224, 255),
+ (112, 224, 255),
+ (70, 184, 160),
+ (163, 0, 255),
+ (153, 0, 255),
+ (71, 255, 0),
+ (255, 0, 163),
+ (255, 204, 0),
+ (255, 0, 143),
+ (0, 255, 235),
+ (133, 255, 0),
+ (255, 0, 235),
+ (245, 0, 255),
+ (255, 0, 122),
+ (255, 245, 0),
+ (10, 190, 212),
+ (214, 255, 0),
+ (0, 204, 255),
+ (20, 0, 255),
+ (255, 255, 0),
+ (0, 153, 255),
+ (0, 41, 255),
+ (0, 255, 204),
+ (41, 0, 255),
+ (41, 255, 0),
+ (173, 0, 255),
+ (0, 245, 255),
+ (71, 0, 255),
+ (122, 0, 255),
+ (0, 255, 184),
+ (0, 92, 255),
+ (184, 255, 0),
+ (0, 133, 255),
+ (255, 214, 0),
+ (25, 194, 194),
+ (102, 255, 0),
+ (92, 0, 255),
+]
+
+ADE20K_CLASS_NAMES = [
+ "",
+ "wall",
+ "building;edifice",
+ "sky",
+ "floor;flooring",
+ "tree",
+ "ceiling",
+ "road;route",
+ "bed",
+ "windowpane;window",
+ "grass",
+ "cabinet",
+ "sidewalk;pavement",
+ "person;individual;someone;somebody;mortal;soul",
+ "earth;ground",
+ "door;double;door",
+ "table",
+ "mountain;mount",
+ "plant;flora;plant;life",
+ "curtain;drape;drapery;mantle;pall",
+ "chair",
+ "car;auto;automobile;machine;motorcar",
+ "water",
+ "painting;picture",
+ "sofa;couch;lounge",
+ "shelf",
+ "house",
+ "sea",
+ "mirror",
+ "rug;carpet;carpeting",
+ "field",
+ "armchair",
+ "seat",
+ "fence;fencing",
+ "desk",
+ "rock;stone",
+ "wardrobe;closet;press",
+ "lamp",
+ "bathtub;bathing;tub;bath;tub",
+ "railing;rail",
+ "cushion",
+ "base;pedestal;stand",
+ "box",
+ "column;pillar",
+ "signboard;sign",
+ "chest;of;drawers;chest;bureau;dresser",
+ "counter",
+ "sand",
+ "sink",
+ "skyscraper",
+ "fireplace;hearth;open;fireplace",
+ "refrigerator;icebox",
+ "grandstand;covered;stand",
+ "path",
+ "stairs;steps",
+ "runway",
+ "case;display;case;showcase;vitrine",
+ "pool;table;billiard;table;snooker;table",
+ "pillow",
+ "screen;door;screen",
+ "stairway;staircase",
+ "river",
+ "bridge;span",
+ "bookcase",
+ "blind;screen",
+ "coffee;table;cocktail;table",
+ "toilet;can;commode;crapper;pot;potty;stool;throne",
+ "flower",
+ "book",
+ "hill",
+ "bench",
+ "countertop",
+ "stove;kitchen;stove;range;kitchen;range;cooking;stove",
+ "palm;palm;tree",
+ "kitchen;island",
+ "computer;computing;machine;computing;device;data;processor;electronic;computer;information;processing;system",
+ "swivel;chair",
+ "boat",
+ "bar",
+ "arcade;machine",
+ "hovel;hut;hutch;shack;shanty",
+ "bus;autobus;coach;charabanc;double-decker;jitney;motorbus;motorcoach;omnibus;passenger;vehicle",
+ "towel",
+ "light;light;source",
+ "truck;motortruck",
+ "tower",
+ "chandelier;pendant;pendent",
+ "awning;sunshade;sunblind",
+ "streetlight;street;lamp",
+ "booth;cubicle;stall;kiosk",
+ "television;television;receiver;television;set;tv;tv;set;idiot;box;boob;tube;telly;goggle;box",
+ "airplane;aeroplane;plane",
+ "dirt;track",
+ "apparel;wearing;apparel;dress;clothes",
+ "pole",
+ "land;ground;soil",
+ "bannister;banister;balustrade;balusters;handrail",
+ "escalator;moving;staircase;moving;stairway",
+ "ottoman;pouf;pouffe;puff;hassock",
+ "bottle",
+ "buffet;counter;sideboard",
+ "poster;posting;placard;notice;bill;card",
+ "stage",
+ "van",
+ "ship",
+ "fountain",
+ "conveyer;belt;conveyor;belt;conveyer;conveyor;transporter",
+ "canopy",
+ "washer;automatic;washer;washing;machine",
+ "plaything;toy",
+ "swimming;pool;swimming;bath;natatorium",
+ "stool",
+ "barrel;cask",
+ "basket;handbasket",
+ "waterfall;falls",
+ "tent;collapsible;shelter",
+ "bag",
+ "minibike;motorbike",
+ "cradle",
+ "oven",
+ "ball",
+ "food;solid;food",
+ "step;stair",
+ "tank;storage;tank",
+ "trade;name;brand;name;brand;marque",
+ "microwave;microwave;oven",
+ "pot;flowerpot",
+ "animal;animate;being;beast;brute;creature;fauna",
+ "bicycle;bike;wheel;cycle",
+ "lake",
+ "dishwasher;dish;washer;dishwashing;machine",
+ "screen;silver;screen;projection;screen",
+ "blanket;cover",
+ "sculpture",
+ "hood;exhaust;hood",
+ "sconce",
+ "vase",
+ "traffic;light;traffic;signal;stoplight",
+ "tray",
+ "ashcan;trash;can;garbage;can;wastebin;ash;bin;ash-bin;ashbin;dustbin;trash;barrel;trash;bin",
+ "fan",
+ "pier;wharf;wharfage;dock",
+ "crt;screen",
+ "plate",
+ "monitor;monitoring;device",
+ "bulletin;board;notice;board",
+ "shower",
+ "radiator",
+ "glass;drinking;glass",
+ "clock",
+ "flag",
+]
+
+
+VOC2012_COLORMAP = [
+ (0, 0, 0),
+ (128, 0, 0),
+ (0, 128, 0),
+ (128, 128, 0),
+ (0, 0, 128),
+ (128, 0, 128),
+ (0, 128, 128),
+ (128, 128, 128),
+ (64, 0, 0),
+ (192, 0, 0),
+ (64, 128, 0),
+ (192, 128, 0),
+ (64, 0, 128),
+ (192, 0, 128),
+ (64, 128, 128),
+ (192, 128, 128),
+ (0, 64, 0),
+ (128, 64, 0),
+ (0, 192, 0),
+ (128, 192, 0),
+ (0, 64, 128),
+]
+
+
+VOC2012_CLASS_NAMES = [
+ "",
+ "aeroplane",
+ "bicycle",
+ "bird",
+ "boat",
+ "bottle",
+ "bus",
+ "car",
+ "cat",
+ "chair",
+ "cow",
+ "diningtable",
+ "dog",
+ "horse",
+ "motorbike",
+ "person",
+ "pottedplant",
+ "sheep",
+ "sofa",
+ "train",
+ "tvmonitor",
+]
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..6c678fdf8f1dee14d7cf9be70af14e6f9a1441c3
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/__init__.py
@@ -0,0 +1,8 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .core import * # noqa: F403
+from .models import * # noqa: F403
+from .ops import * # noqa: F403
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..92599806fbd221c1418d179892a0f46dc0b7d4db
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/__init__.py
@@ -0,0 +1,11 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from mmseg.core.evaluation import * # noqa: F403
+from mmseg.core.seg import * # noqa: F403
+
+from .anchor import * # noqa: F403
+from .box import * # noqa: F403
+from .utils import * # noqa: F403
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/anchor/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/anchor/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..e71ac4d6e01462221ae01aa16d0e1231cda7e2e7
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/anchor/__init__.py
@@ -0,0 +1,6 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .point_generator import MlvlPointGenerator # noqa: F403
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/anchor/builder.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/anchor/builder.py
new file mode 100644
index 0000000000000000000000000000000000000000..6dba90e22de76d2f23a86d3c057f196d55a99690
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/anchor/builder.py
@@ -0,0 +1,21 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import warnings
+
+from mmcv.utils import Registry, build_from_cfg
+
+PRIOR_GENERATORS = Registry("Generator for anchors and points")
+
+ANCHOR_GENERATORS = PRIOR_GENERATORS
+
+
+def build_prior_generator(cfg, default_args=None):
+ return build_from_cfg(cfg, PRIOR_GENERATORS, default_args)
+
+
+def build_anchor_generator(cfg, default_args=None):
+ warnings.warn("``build_anchor_generator`` would be deprecated soon, please use " "``build_prior_generator`` ")
+ return build_prior_generator(cfg, default_args=default_args)
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/anchor/point_generator.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/anchor/point_generator.py
new file mode 100644
index 0000000000000000000000000000000000000000..574d71939080e22284fe99087fb2e7336657bd97
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/anchor/point_generator.py
@@ -0,0 +1,205 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import numpy as np
+import torch
+from torch.nn.modules.utils import _pair
+
+from .builder import PRIOR_GENERATORS
+
+
+@PRIOR_GENERATORS.register_module()
+class MlvlPointGenerator:
+ """Standard points generator for multi-level (Mlvl) feature maps in 2D
+ points-based detectors.
+
+ Args:
+ strides (list[int] | list[tuple[int, int]]): Strides of anchors
+ in multiple feature levels in order (w, h).
+ offset (float): The offset of points, the value is normalized with
+ corresponding stride. Defaults to 0.5.
+ """
+
+ def __init__(self, strides, offset=0.5):
+ self.strides = [_pair(stride) for stride in strides]
+ self.offset = offset
+
+ @property
+ def num_levels(self):
+ """int: number of feature levels that the generator will be applied"""
+ return len(self.strides)
+
+ @property
+ def num_base_priors(self):
+ """list[int]: The number of priors (points) at a point
+ on the feature grid"""
+ return [1 for _ in range(len(self.strides))]
+
+ def _meshgrid(self, x, y, row_major=True):
+ yy, xx = torch.meshgrid(y, x)
+ if row_major:
+ # warning .flatten() would cause error in ONNX exporting
+ # have to use reshape here
+ return xx.reshape(-1), yy.reshape(-1)
+
+ else:
+ return yy.reshape(-1), xx.reshape(-1)
+
+ def grid_priors(self, featmap_sizes, dtype=torch.float32, device="cuda", with_stride=False):
+ """Generate grid points of multiple feature levels.
+
+ Args:
+ featmap_sizes (list[tuple]): List of feature map sizes in
+ multiple feature levels, each size arrange as
+ as (h, w).
+ dtype (:obj:`dtype`): Dtype of priors. Default: torch.float32.
+ device (str): The device where the anchors will be put on.
+ with_stride (bool): Whether to concatenate the stride to
+ the last dimension of points.
+
+ Return:
+ list[torch.Tensor]: Points of multiple feature levels.
+ The sizes of each tensor should be (N, 2) when with stride is
+ ``False``, where N = width * height, width and height
+ are the sizes of the corresponding feature level,
+ and the last dimension 2 represent (coord_x, coord_y),
+ otherwise the shape should be (N, 4),
+ and the last dimension 4 represent
+ (coord_x, coord_y, stride_w, stride_h).
+ """
+
+ assert self.num_levels == len(featmap_sizes)
+ multi_level_priors = []
+ for i in range(self.num_levels):
+ priors = self.single_level_grid_priors(
+ featmap_sizes[i], level_idx=i, dtype=dtype, device=device, with_stride=with_stride
+ )
+ multi_level_priors.append(priors)
+ return multi_level_priors
+
+ def single_level_grid_priors(self, featmap_size, level_idx, dtype=torch.float32, device="cuda", with_stride=False):
+ """Generate grid Points of a single level.
+
+ Note:
+ This function is usually called by method ``self.grid_priors``.
+
+ Args:
+ featmap_size (tuple[int]): Size of the feature maps, arrange as
+ (h, w).
+ level_idx (int): The index of corresponding feature map level.
+ dtype (:obj:`dtype`): Dtype of priors. Default: torch.float32.
+ device (str, optional): The device the tensor will be put on.
+ Defaults to 'cuda'.
+ with_stride (bool): Concatenate the stride to the last dimension
+ of points.
+
+ Return:
+ Tensor: Points of single feature levels.
+ The shape of tensor should be (N, 2) when with stride is
+ ``False``, where N = width * height, width and height
+ are the sizes of the corresponding feature level,
+ and the last dimension 2 represent (coord_x, coord_y),
+ otherwise the shape should be (N, 4),
+ and the last dimension 4 represent
+ (coord_x, coord_y, stride_w, stride_h).
+ """
+ feat_h, feat_w = featmap_size
+ stride_w, stride_h = self.strides[level_idx]
+ shift_x = (torch.arange(0, feat_w, device=device) + self.offset) * stride_w
+ # keep featmap_size as Tensor instead of int, so that we
+ # can convert to ONNX correctly
+ shift_x = shift_x.to(dtype)
+
+ shift_y = (torch.arange(0, feat_h, device=device) + self.offset) * stride_h
+ # keep featmap_size as Tensor instead of int, so that we
+ # can convert to ONNX correctly
+ shift_y = shift_y.to(dtype)
+ shift_xx, shift_yy = self._meshgrid(shift_x, shift_y)
+ if not with_stride:
+ shifts = torch.stack([shift_xx, shift_yy], dim=-1)
+ else:
+ # use `shape[0]` instead of `len(shift_xx)` for ONNX export
+ stride_w = shift_xx.new_full((shift_xx.shape[0],), stride_w).to(dtype)
+ stride_h = shift_xx.new_full((shift_yy.shape[0],), stride_h).to(dtype)
+ shifts = torch.stack([shift_xx, shift_yy, stride_w, stride_h], dim=-1)
+ all_points = shifts.to(device)
+ return all_points
+
+ def valid_flags(self, featmap_sizes, pad_shape, device="cuda"):
+ """Generate valid flags of points of multiple feature levels.
+
+ Args:
+ featmap_sizes (list(tuple)): List of feature map sizes in
+ multiple feature levels, each size arrange as
+ as (h, w).
+ pad_shape (tuple(int)): The padded shape of the image,
+ arrange as (h, w).
+ device (str): The device where the anchors will be put on.
+
+ Return:
+ list(torch.Tensor): Valid flags of points of multiple levels.
+ """
+ assert self.num_levels == len(featmap_sizes)
+ multi_level_flags = []
+ for i in range(self.num_levels):
+ point_stride = self.strides[i]
+ feat_h, feat_w = featmap_sizes[i]
+ h, w = pad_shape[:2]
+ valid_feat_h = min(int(np.ceil(h / point_stride[1])), feat_h)
+ valid_feat_w = min(int(np.ceil(w / point_stride[0])), feat_w)
+ flags = self.single_level_valid_flags((feat_h, feat_w), (valid_feat_h, valid_feat_w), device=device)
+ multi_level_flags.append(flags)
+ return multi_level_flags
+
+ def single_level_valid_flags(self, featmap_size, valid_size, device="cuda"):
+ """Generate the valid flags of points of a single feature map.
+
+ Args:
+ featmap_size (tuple[int]): The size of feature maps, arrange as
+ as (h, w).
+ valid_size (tuple[int]): The valid size of the feature maps.
+ The size arrange as as (h, w).
+ device (str, optional): The device where the flags will be put on.
+ Defaults to 'cuda'.
+
+ Returns:
+ torch.Tensor: The valid flags of each points in a single level \
+ feature map.
+ """
+ feat_h, feat_w = featmap_size
+ valid_h, valid_w = valid_size
+ assert valid_h <= feat_h and valid_w <= feat_w
+ valid_x = torch.zeros(feat_w, dtype=torch.bool, device=device)
+ valid_y = torch.zeros(feat_h, dtype=torch.bool, device=device)
+ valid_x[:valid_w] = 1
+ valid_y[:valid_h] = 1
+ valid_xx, valid_yy = self._meshgrid(valid_x, valid_y)
+ valid = valid_xx & valid_yy
+ return valid
+
+ def sparse_priors(self, prior_idxs, featmap_size, level_idx, dtype=torch.float32, device="cuda"):
+ """Generate sparse points according to the ``prior_idxs``.
+
+ Args:
+ prior_idxs (Tensor): The index of corresponding anchors
+ in the feature map.
+ featmap_size (tuple[int]): feature map size arrange as (w, h).
+ level_idx (int): The level index of corresponding feature
+ map.
+ dtype (obj:`torch.dtype`): Date type of points. Defaults to
+ ``torch.float32``.
+ device (obj:`torch.device`): The device where the points is
+ located.
+ Returns:
+ Tensor: Anchor with shape (N, 2), N should be equal to
+ the length of ``prior_idxs``. And last dimension
+ 2 represent (coord_x, coord_y).
+ """
+ height, width = featmap_size
+ x = (prior_idxs % width + self.offset) * self.strides[level_idx][0]
+ y = ((prior_idxs // width) % height + self.offset) * self.strides[level_idx][1]
+ prioris = torch.stack([x, y], 1).to(dtype)
+ prioris = prioris.to(device)
+ return prioris
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..bf35a613f81acd77ecab2dfb75a722fa8e5c0787
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/__init__.py
@@ -0,0 +1,7 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .builder import * # noqa: F403
+from .samplers import MaskPseudoSampler # noqa: F403
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/builder.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/builder.py
new file mode 100644
index 0000000000000000000000000000000000000000..9538c0de3db682c2b111b085a8a1ce321c76a9ff
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/builder.py
@@ -0,0 +1,19 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from mmcv.utils import Registry, build_from_cfg
+
+BBOX_SAMPLERS = Registry("bbox_sampler")
+BBOX_CODERS = Registry("bbox_coder")
+
+
+def build_sampler(cfg, **default_args):
+ """Builder of box sampler."""
+ return build_from_cfg(cfg, BBOX_SAMPLERS, default_args)
+
+
+def build_bbox_coder(cfg, **default_args):
+ """Builder of box coder."""
+ return build_from_cfg(cfg, BBOX_CODERS, default_args)
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/samplers/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/samplers/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..19c363e3fabc365d92aeaf1e78189d710db279e9
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/samplers/__init__.py
@@ -0,0 +1,6 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .mask_pseudo_sampler import MaskPseudoSampler # noqa: F403
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/samplers/base_sampler.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/samplers/base_sampler.py
new file mode 100644
index 0000000000000000000000000000000000000000..c45cec3ed7af5b49bb54b92d6e6bcf59b06b4c99
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/samplers/base_sampler.py
@@ -0,0 +1,92 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from abc import ABCMeta, abstractmethod
+
+import torch
+
+from .sampling_result import SamplingResult
+
+
+class BaseSampler(metaclass=ABCMeta):
+ """Base class of samplers."""
+
+ def __init__(self, num, pos_fraction, neg_pos_ub=-1, add_gt_as_proposals=True, **kwargs):
+ self.num = num
+ self.pos_fraction = pos_fraction
+ self.neg_pos_ub = neg_pos_ub
+ self.add_gt_as_proposals = add_gt_as_proposals
+ self.pos_sampler = self
+ self.neg_sampler = self
+
+ @abstractmethod
+ def _sample_pos(self, assign_result, num_expected, **kwargs):
+ """Sample positive samples."""
+ pass
+
+ @abstractmethod
+ def _sample_neg(self, assign_result, num_expected, **kwargs):
+ """Sample negative samples."""
+ pass
+
+ def sample(self, assign_result, bboxes, gt_bboxes, gt_labels=None, **kwargs):
+ """Sample positive and negative bboxes.
+
+ This is a simple implementation of bbox sampling given candidates,
+ assigning results and ground truth bboxes.
+
+ Args:
+ assign_result (:obj:`AssignResult`): Bbox assigning results.
+ bboxes (Tensor): Boxes to be sampled from.
+ gt_bboxes (Tensor): Ground truth bboxes.
+ gt_labels (Tensor, optional): Class labels of ground truth bboxes.
+
+ Returns:
+ :obj:`SamplingResult`: Sampling result.
+
+ Example:
+ >>> from mmdet.core.bbox import RandomSampler
+ >>> from mmdet.core.bbox import AssignResult
+ >>> from mmdet.core.bbox.demodata import ensure_rng, random_boxes
+ >>> rng = ensure_rng(None)
+ >>> assign_result = AssignResult.random(rng=rng)
+ >>> bboxes = random_boxes(assign_result.num_preds, rng=rng)
+ >>> gt_bboxes = random_boxes(assign_result.num_gts, rng=rng)
+ >>> gt_labels = None
+ >>> self = RandomSampler(num=32, pos_fraction=0.5, neg_pos_ub=-1,
+ >>> add_gt_as_proposals=False)
+ >>> self = self.sample(assign_result, bboxes, gt_bboxes, gt_labels)
+ """
+ if len(bboxes.shape) < 2:
+ bboxes = bboxes[None, :]
+
+ bboxes = bboxes[:, :4]
+
+ gt_flags = bboxes.new_zeros((bboxes.shape[0],), dtype=torch.uint8)
+ if self.add_gt_as_proposals and len(gt_bboxes) > 0:
+ if gt_labels is None:
+ raise ValueError("gt_labels must be given when add_gt_as_proposals is True")
+ bboxes = torch.cat([gt_bboxes, bboxes], dim=0)
+ assign_result.add_gt_(gt_labels)
+ gt_ones = bboxes.new_ones(gt_bboxes.shape[0], dtype=torch.uint8)
+ gt_flags = torch.cat([gt_ones, gt_flags])
+
+ num_expected_pos = int(self.num * self.pos_fraction)
+ pos_inds = self.pos_sampler._sample_pos(assign_result, num_expected_pos, bboxes=bboxes, **kwargs)
+ # We found that sampled indices have duplicated items occasionally.
+ # (may be a bug of PyTorch)
+ pos_inds = pos_inds.unique()
+ num_sampled_pos = pos_inds.numel()
+ num_expected_neg = self.num - num_sampled_pos
+ if self.neg_pos_ub >= 0:
+ _pos = max(1, num_sampled_pos)
+ neg_upper_bound = int(self.neg_pos_ub * _pos)
+ if num_expected_neg > neg_upper_bound:
+ num_expected_neg = neg_upper_bound
+ neg_inds = self.neg_sampler._sample_neg(assign_result, num_expected_neg, bboxes=bboxes, **kwargs)
+ neg_inds = neg_inds.unique()
+
+ sampling_result = SamplingResult(pos_inds, neg_inds, bboxes, gt_bboxes, assign_result, gt_flags)
+ return sampling_result
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/samplers/mask_pseudo_sampler.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/samplers/mask_pseudo_sampler.py
new file mode 100644
index 0000000000000000000000000000000000000000..3e67ea61ed0fd65cca0addde1893a3c1e176bf15
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/samplers/mask_pseudo_sampler.py
@@ -0,0 +1,45 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+# References:
+# https://github.com/ZwwWayne/K-Net/blob/main/knet/det/mask_pseudo_sampler.py
+
+import torch
+
+from ..builder import BBOX_SAMPLERS
+from .base_sampler import BaseSampler
+from .mask_sampling_result import MaskSamplingResult
+
+
+@BBOX_SAMPLERS.register_module()
+class MaskPseudoSampler(BaseSampler):
+ """A pseudo sampler that does not do sampling actually."""
+
+ def __init__(self, **kwargs):
+ pass
+
+ def _sample_pos(self, **kwargs):
+ """Sample positive samples."""
+ raise NotImplementedError
+
+ def _sample_neg(self, **kwargs):
+ """Sample negative samples."""
+ raise NotImplementedError
+
+ def sample(self, assign_result, masks, gt_masks, **kwargs):
+ """Directly returns the positive and negative indices of samples.
+
+ Args:
+ assign_result (:obj:`AssignResult`): Assigned results
+ masks (torch.Tensor): Bounding boxes
+ gt_masks (torch.Tensor): Ground truth boxes
+ Returns:
+ :obj:`SamplingResult`: sampler results
+ """
+ pos_inds = torch.nonzero(assign_result.gt_inds > 0, as_tuple=False).squeeze(-1).unique()
+ neg_inds = torch.nonzero(assign_result.gt_inds == 0, as_tuple=False).squeeze(-1).unique()
+ gt_flags = masks.new_zeros(masks.shape[0], dtype=torch.uint8)
+ sampling_result = MaskSamplingResult(pos_inds, neg_inds, masks, gt_masks, assign_result, gt_flags)
+ return sampling_result
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/samplers/mask_sampling_result.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/samplers/mask_sampling_result.py
new file mode 100644
index 0000000000000000000000000000000000000000..270ffd35a5f120dd0560a7fea7fe83ef0bab66bb
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/samplers/mask_sampling_result.py
@@ -0,0 +1,63 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+# References:
+# https://github.com/ZwwWayne/K-Net/blob/main/knet/det/mask_pseudo_sampler.py
+
+import torch
+
+from .sampling_result import SamplingResult
+
+
+class MaskSamplingResult(SamplingResult):
+ """Mask sampling result."""
+
+ def __init__(self, pos_inds, neg_inds, masks, gt_masks, assign_result, gt_flags):
+ self.pos_inds = pos_inds
+ self.neg_inds = neg_inds
+ self.pos_masks = masks[pos_inds]
+ self.neg_masks = masks[neg_inds]
+ self.pos_is_gt = gt_flags[pos_inds]
+
+ self.num_gts = gt_masks.shape[0]
+ self.pos_assigned_gt_inds = assign_result.gt_inds[pos_inds] - 1
+
+ if gt_masks.numel() == 0:
+ # hack for index error case
+ assert self.pos_assigned_gt_inds.numel() == 0
+ self.pos_gt_masks = torch.empty_like(gt_masks)
+ else:
+ self.pos_gt_masks = gt_masks[self.pos_assigned_gt_inds, :]
+
+ if assign_result.labels is not None:
+ self.pos_gt_labels = assign_result.labels[pos_inds]
+ else:
+ self.pos_gt_labels = None
+
+ @property
+ def masks(self):
+ """torch.Tensor: concatenated positive and negative boxes"""
+ return torch.cat([self.pos_masks, self.neg_masks])
+
+ def __nice__(self):
+ data = self.info.copy()
+ data["pos_masks"] = data.pop("pos_masks").shape
+ data["neg_masks"] = data.pop("neg_masks").shape
+ parts = [f"'{k}': {v!r}" for k, v in sorted(data.items())]
+ body = " " + ",\n ".join(parts)
+ return "{\n" + body + "\n}"
+
+ @property
+ def info(self):
+ """Returns a dictionary of info about the object."""
+ return {
+ "pos_inds": self.pos_inds,
+ "neg_inds": self.neg_inds,
+ "pos_masks": self.pos_masks,
+ "neg_masks": self.neg_masks,
+ "pos_is_gt": self.pos_is_gt,
+ "num_gts": self.num_gts,
+ "pos_assigned_gt_inds": self.pos_assigned_gt_inds,
+ }
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/samplers/sampling_result.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/samplers/sampling_result.py
new file mode 100644
index 0000000000000000000000000000000000000000..aaee3fe55aeb8c6da7edefbbd382d94b67b6a6b4
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/box/samplers/sampling_result.py
@@ -0,0 +1,152 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import torch
+
+
+class SamplingResult:
+ """Bbox sampling result.
+
+ Example:
+ >>> # xdoctest: +IGNORE_WANT
+ >>> from mmdet.core.bbox.samplers.sampling_result import * # NOQA
+ >>> self = SamplingResult.random(rng=10)
+ >>> print(f'self = {self}')
+ self =
+ """
+
+ def __init__(self, pos_inds, neg_inds, bboxes, gt_bboxes, assign_result, gt_flags):
+ self.pos_inds = pos_inds
+ self.neg_inds = neg_inds
+ self.pos_bboxes = bboxes[pos_inds]
+ self.neg_bboxes = bboxes[neg_inds]
+ self.pos_is_gt = gt_flags[pos_inds]
+
+ self.num_gts = gt_bboxes.shape[0]
+ self.pos_assigned_gt_inds = assign_result.gt_inds[pos_inds] - 1
+
+ if gt_bboxes.numel() == 0:
+ # hack for index error case
+ assert self.pos_assigned_gt_inds.numel() == 0
+ self.pos_gt_bboxes = torch.empty_like(gt_bboxes).view(-1, 4)
+ else:
+ if len(gt_bboxes.shape) < 2:
+ gt_bboxes = gt_bboxes.view(-1, 4)
+
+ self.pos_gt_bboxes = gt_bboxes[self.pos_assigned_gt_inds.long(), :]
+
+ if assign_result.labels is not None:
+ self.pos_gt_labels = assign_result.labels[pos_inds]
+ else:
+ self.pos_gt_labels = None
+
+ @property
+ def bboxes(self):
+ """torch.Tensor: concatenated positive and negative boxes"""
+ return torch.cat([self.pos_bboxes, self.neg_bboxes])
+
+ def to(self, device):
+ """Change the device of the data inplace.
+
+ Example:
+ >>> self = SamplingResult.random()
+ >>> print(f'self = {self.to(None)}')
+ >>> # xdoctest: +REQUIRES(--gpu)
+ >>> print(f'self = {self.to(0)}')
+ """
+ _dict = self.__dict__
+ for key, value in _dict.items():
+ if isinstance(value, torch.Tensor):
+ _dict[key] = value.to(device)
+ return self
+
+ def __nice__(self):
+ data = self.info.copy()
+ data["pos_bboxes"] = data.pop("pos_bboxes").shape
+ data["neg_bboxes"] = data.pop("neg_bboxes").shape
+ parts = [f"'{k}': {v!r}" for k, v in sorted(data.items())]
+ body = " " + ",\n ".join(parts)
+ return "{\n" + body + "\n}"
+
+ @property
+ def info(self):
+ """Returns a dictionary of info about the object."""
+ return {
+ "pos_inds": self.pos_inds,
+ "neg_inds": self.neg_inds,
+ "pos_bboxes": self.pos_bboxes,
+ "neg_bboxes": self.neg_bboxes,
+ "pos_is_gt": self.pos_is_gt,
+ "num_gts": self.num_gts,
+ "pos_assigned_gt_inds": self.pos_assigned_gt_inds,
+ }
+
+ @classmethod
+ def random(cls, rng=None, **kwargs):
+ """
+ Args:
+ rng (None | int | numpy.random.RandomState): seed or state.
+ kwargs (keyword arguments):
+ - num_preds: number of predicted boxes
+ - num_gts: number of true boxes
+ - p_ignore (float): probability of a predicted box assigned to \
+ an ignored truth.
+ - p_assigned (float): probability of a predicted box not being \
+ assigned.
+ - p_use_label (float | bool): with labels or not.
+
+ Returns:
+ :obj:`SamplingResult`: Randomly generated sampling result.
+
+ Example:
+ >>> from mmdet.core.bbox.samplers.sampling_result import * # NOQA
+ >>> self = SamplingResult.random()
+ >>> print(self.__dict__)
+ """
+ from mmdet.core.bbox import demodata
+ from mmdet.core.bbox.assigners.assign_result import AssignResult
+ from mmdet.core.bbox.samplers.random_sampler import RandomSampler
+
+ rng = demodata.ensure_rng(rng)
+
+ # make probabalistic?
+ num = 32
+ pos_fraction = 0.5
+ neg_pos_ub = -1
+
+ assign_result = AssignResult.random(rng=rng, **kwargs)
+
+ # Note we could just compute an assignment
+ bboxes = demodata.random_boxes(assign_result.num_preds, rng=rng)
+ gt_bboxes = demodata.random_boxes(assign_result.num_gts, rng=rng)
+
+ if rng.rand() > 0.2:
+ # sometimes algorithms squeeze their data, be robust to that
+ gt_bboxes = gt_bboxes.squeeze()
+ bboxes = bboxes.squeeze()
+
+ if assign_result.labels is None:
+ gt_labels = None
+ else:
+ gt_labels = None
+
+ if gt_labels is None:
+ add_gt_as_proposals = False
+ else:
+ add_gt_as_proposals = True # make probabalistic?
+
+ sampler = RandomSampler(
+ num, pos_fraction, neg_pos_ub=neg_pos_ub, add_gt_as_proposals=add_gt_as_proposals, rng=rng
+ )
+ self = sampler.sample(assign_result, bboxes, gt_bboxes, gt_labels)
+ return self
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/utils/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/utils/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..6cdc9e19352f50bc2d5433c412ff71186c5df019
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/utils/__init__.py
@@ -0,0 +1,7 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .dist_utils import reduce_mean
+from .misc import add_prefix, multi_apply
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/utils/dist_utils.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/utils/dist_utils.py
new file mode 100644
index 0000000000000000000000000000000000000000..7dfed42da821cd94e31b663d86b20b8f09799b30
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/utils/dist_utils.py
@@ -0,0 +1,15 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import torch.distributed as dist
+
+
+def reduce_mean(tensor):
+ """ "Obtain the mean of tensor on different GPUs."""
+ if not (dist.is_available() and dist.is_initialized()):
+ return tensor
+ tensor = tensor.clone()
+ dist.all_reduce(tensor.div_(dist.get_world_size()), op=dist.ReduceOp.SUM)
+ return tensor
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/utils/misc.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/utils/misc.py
new file mode 100644
index 0000000000000000000000000000000000000000..e07579e7b182b62153e81fe637ffd0f3081ef2a3
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/core/utils/misc.py
@@ -0,0 +1,47 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from functools import partial
+
+
+def multi_apply(func, *args, **kwargs):
+ """Apply function to a list of arguments.
+
+ Note:
+ This function applies the ``func`` to multiple inputs and
+ map the multiple outputs of the ``func`` into different
+ list. Each list contains the same type of outputs corresponding
+ to different inputs.
+
+ Args:
+ func (Function): A function that will be applied to a list of
+ arguments
+
+ Returns:
+ tuple(list): A tuple containing multiple list, each list contains \
+ a kind of returned results by the function
+ """
+ pfunc = partial(func, **kwargs) if kwargs else func
+ map_results = map(pfunc, *args)
+ return tuple(map(list, zip(*map_results)))
+
+
+def add_prefix(inputs, prefix):
+ """Add prefix for dict.
+
+ Args:
+ inputs (dict): The input dict with str keys.
+ prefix (str): The prefix to add.
+
+ Returns:
+
+ dict: The dict with keys updated with ``prefix``.
+ """
+
+ outputs = dict()
+ for name, value in inputs.items():
+ outputs[f"{prefix}.{name}"] = value
+
+ return outputs
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..ed89bb0064d82b4360af020798eab3d2f5a47937
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/__init__.py
@@ -0,0 +1,11 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .backbones import * # noqa: F403
+from .builder import MASK_ASSIGNERS, MATCH_COST, TRANSFORMER, build_assigner, build_match_cost
+from .decode_heads import * # noqa: F403
+from .losses import * # noqa: F403
+from .plugins import * # noqa: F403
+from .segmentors import * # noqa: F403
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/backbones/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/backbones/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..c4bf73bcbcee710676f81cb6517ae787f4d61cc6
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/backbones/__init__.py
@@ -0,0 +1,6 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .vit_adapter import ViTAdapter
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/backbones/adapter_modules.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/backbones/adapter_modules.py
new file mode 100644
index 0000000000000000000000000000000000000000..26bfdf8f6ae6c107d22d61985cce34d4b5ce275f
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/backbones/adapter_modules.py
@@ -0,0 +1,442 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from functools import partial
+
+import torch
+import torch.nn as nn
+import torch.utils.checkpoint as cp
+
+from ...ops.modules import MSDeformAttn
+from .drop_path import DropPath
+
+
+def get_reference_points(spatial_shapes, device):
+ reference_points_list = []
+ for lvl, (H_, W_) in enumerate(spatial_shapes):
+ ref_y, ref_x = torch.meshgrid(
+ torch.linspace(0.5, H_ - 0.5, H_, dtype=torch.float32, device=device),
+ torch.linspace(0.5, W_ - 0.5, W_, dtype=torch.float32, device=device),
+ )
+ ref_y = ref_y.reshape(-1)[None] / H_
+ ref_x = ref_x.reshape(-1)[None] / W_
+ ref = torch.stack((ref_x, ref_y), -1)
+ reference_points_list.append(ref)
+ reference_points = torch.cat(reference_points_list, 1)
+ reference_points = reference_points[:, :, None]
+ return reference_points
+
+
+def deform_inputs(x, patch_size):
+ bs, c, h, w = x.shape
+ spatial_shapes = torch.as_tensor(
+ [(h // 8, w // 8), (h // 16, w // 16), (h // 32, w // 32)], dtype=torch.long, device=x.device
+ )
+ level_start_index = torch.cat((spatial_shapes.new_zeros((1,)), spatial_shapes.prod(1).cumsum(0)[:-1]))
+ reference_points = get_reference_points([(h // patch_size, w // patch_size)], x.device)
+ deform_inputs1 = [reference_points, spatial_shapes, level_start_index]
+
+ spatial_shapes = torch.as_tensor([(h // patch_size, w // patch_size)], dtype=torch.long, device=x.device)
+ level_start_index = torch.cat((spatial_shapes.new_zeros((1,)), spatial_shapes.prod(1).cumsum(0)[:-1]))
+ reference_points = get_reference_points([(h // 8, w // 8), (h // 16, w // 16), (h // 32, w // 32)], x.device)
+ deform_inputs2 = [reference_points, spatial_shapes, level_start_index]
+
+ return deform_inputs1, deform_inputs2
+
+
+class ConvFFN(nn.Module):
+ def __init__(self, in_features, hidden_features=None, out_features=None, act_layer=nn.GELU, drop=0.0):
+ super().__init__()
+ out_features = out_features or in_features
+ hidden_features = hidden_features or in_features
+ self.fc1 = nn.Linear(in_features, hidden_features)
+ self.dwconv = DWConv(hidden_features)
+ self.act = act_layer()
+ self.fc2 = nn.Linear(hidden_features, out_features)
+ self.drop = nn.Dropout(drop)
+
+ def forward(self, x, H, W):
+ x = self.fc1(x)
+ x = self.dwconv(x, H, W)
+ x = self.act(x)
+ x = self.drop(x)
+ x = self.fc2(x)
+ x = self.drop(x)
+ return x
+
+
+class DWConv(nn.Module):
+ def __init__(self, dim=768):
+ super().__init__()
+ self.dwconv = nn.Conv2d(dim, dim, 3, 1, 1, bias=True, groups=dim)
+
+ def forward(self, x, H, W):
+ B, N, C = x.shape
+ n = N // 21
+ x1 = x[:, 0 : 16 * n, :].transpose(1, 2).view(B, C, H * 2, W * 2).contiguous()
+ x2 = x[:, 16 * n : 20 * n, :].transpose(1, 2).view(B, C, H, W).contiguous()
+ x3 = x[:, 20 * n :, :].transpose(1, 2).view(B, C, H // 2, W // 2).contiguous()
+ x1 = self.dwconv(x1).flatten(2).transpose(1, 2)
+ x2 = self.dwconv(x2).flatten(2).transpose(1, 2)
+ x3 = self.dwconv(x3).flatten(2).transpose(1, 2)
+ x = torch.cat([x1, x2, x3], dim=1)
+ return x
+
+
+class Extractor(nn.Module):
+ def __init__(
+ self,
+ dim,
+ num_heads=6,
+ n_points=4,
+ n_levels=1,
+ deform_ratio=1.0,
+ with_cffn=True,
+ cffn_ratio=0.25,
+ drop=0.0,
+ drop_path=0.0,
+ norm_layer=partial(nn.LayerNorm, eps=1e-6),
+ with_cp=False,
+ ):
+ super().__init__()
+ self.query_norm = norm_layer(dim)
+ self.feat_norm = norm_layer(dim)
+ self.attn = MSDeformAttn(
+ d_model=dim, n_levels=n_levels, n_heads=num_heads, n_points=n_points, ratio=deform_ratio
+ )
+ self.with_cffn = with_cffn
+ self.with_cp = with_cp
+ if with_cffn:
+ self.ffn = ConvFFN(in_features=dim, hidden_features=int(dim * cffn_ratio), drop=drop)
+ self.ffn_norm = norm_layer(dim)
+ self.drop_path = DropPath(drop_path) if drop_path > 0.0 else nn.Identity()
+
+ def forward(self, query, reference_points, feat, spatial_shapes, level_start_index, H, W):
+ def _inner_forward(query, feat):
+
+ attn = self.attn(
+ self.query_norm(query), reference_points, self.feat_norm(feat), spatial_shapes, level_start_index, None
+ )
+ query = query + attn
+
+ if self.with_cffn:
+ query = query + self.drop_path(self.ffn(self.ffn_norm(query), H, W))
+ return query
+
+ if self.with_cp and query.requires_grad:
+ query = cp.checkpoint(_inner_forward, query, feat)
+ else:
+ query = _inner_forward(query, feat)
+
+ return query
+
+
+class Injector(nn.Module):
+ def __init__(
+ self,
+ dim,
+ num_heads=6,
+ n_points=4,
+ n_levels=1,
+ deform_ratio=1.0,
+ norm_layer=partial(nn.LayerNorm, eps=1e-6),
+ init_values=0.0,
+ with_cp=False,
+ ):
+ super().__init__()
+ self.with_cp = with_cp
+ self.query_norm = norm_layer(dim)
+ self.feat_norm = norm_layer(dim)
+ self.attn = MSDeformAttn(
+ d_model=dim, n_levels=n_levels, n_heads=num_heads, n_points=n_points, ratio=deform_ratio
+ )
+ self.gamma = nn.Parameter(init_values * torch.ones((dim)), requires_grad=True)
+
+ def forward(self, query, reference_points, feat, spatial_shapes, level_start_index):
+ def _inner_forward(query, feat):
+
+ attn = self.attn(
+ self.query_norm(query), reference_points, self.feat_norm(feat), spatial_shapes, level_start_index, None
+ )
+ return query + self.gamma * attn
+
+ if self.with_cp and query.requires_grad:
+ query = cp.checkpoint(_inner_forward, query, feat)
+ else:
+ query = _inner_forward(query, feat)
+
+ return query
+
+
+class InteractionBlock(nn.Module):
+ def __init__(
+ self,
+ dim,
+ num_heads=6,
+ n_points=4,
+ norm_layer=partial(nn.LayerNorm, eps=1e-6),
+ drop=0.0,
+ drop_path=0.0,
+ with_cffn=True,
+ cffn_ratio=0.25,
+ init_values=0.0,
+ deform_ratio=1.0,
+ extra_extractor=False,
+ with_cp=False,
+ ):
+ super().__init__()
+
+ self.injector = Injector(
+ dim=dim,
+ n_levels=3,
+ num_heads=num_heads,
+ init_values=init_values,
+ n_points=n_points,
+ norm_layer=norm_layer,
+ deform_ratio=deform_ratio,
+ with_cp=with_cp,
+ )
+ self.extractor = Extractor(
+ dim=dim,
+ n_levels=1,
+ num_heads=num_heads,
+ n_points=n_points,
+ norm_layer=norm_layer,
+ deform_ratio=deform_ratio,
+ with_cffn=with_cffn,
+ cffn_ratio=cffn_ratio,
+ drop=drop,
+ drop_path=drop_path,
+ with_cp=with_cp,
+ )
+ if extra_extractor:
+ self.extra_extractors = nn.Sequential(
+ *[
+ Extractor(
+ dim=dim,
+ num_heads=num_heads,
+ n_points=n_points,
+ norm_layer=norm_layer,
+ with_cffn=with_cffn,
+ cffn_ratio=cffn_ratio,
+ deform_ratio=deform_ratio,
+ drop=drop,
+ drop_path=drop_path,
+ with_cp=with_cp,
+ )
+ for _ in range(2)
+ ]
+ )
+ else:
+ self.extra_extractors = None
+
+ def forward(self, x, c, blocks, deform_inputs1, deform_inputs2, H_c, W_c, H_toks, W_toks):
+ x = self.injector(
+ query=x,
+ reference_points=deform_inputs1[0],
+ feat=c,
+ spatial_shapes=deform_inputs1[1],
+ level_start_index=deform_inputs1[2],
+ )
+ for idx, blk in enumerate(blocks):
+ x = blk(x, H_toks, W_toks)
+ c = self.extractor(
+ query=c,
+ reference_points=deform_inputs2[0],
+ feat=x,
+ spatial_shapes=deform_inputs2[1],
+ level_start_index=deform_inputs2[2],
+ H=H_c,
+ W=W_c,
+ )
+ if self.extra_extractors is not None:
+ for extractor in self.extra_extractors:
+ c = extractor(
+ query=c,
+ reference_points=deform_inputs2[0],
+ feat=x,
+ spatial_shapes=deform_inputs2[1],
+ level_start_index=deform_inputs2[2],
+ H=H_c,
+ W=W_c,
+ )
+ return x, c
+
+
+class InteractionBlockWithCls(nn.Module):
+ def __init__(
+ self,
+ dim,
+ num_heads=6,
+ n_points=4,
+ norm_layer=partial(nn.LayerNorm, eps=1e-6),
+ drop=0.0,
+ drop_path=0.0,
+ with_cffn=True,
+ cffn_ratio=0.25,
+ init_values=0.0,
+ deform_ratio=1.0,
+ extra_extractor=False,
+ with_cp=False,
+ ):
+ super().__init__()
+
+ self.injector = Injector(
+ dim=dim,
+ n_levels=3,
+ num_heads=num_heads,
+ init_values=init_values,
+ n_points=n_points,
+ norm_layer=norm_layer,
+ deform_ratio=deform_ratio,
+ with_cp=with_cp,
+ )
+ self.extractor = Extractor(
+ dim=dim,
+ n_levels=1,
+ num_heads=num_heads,
+ n_points=n_points,
+ norm_layer=norm_layer,
+ deform_ratio=deform_ratio,
+ with_cffn=with_cffn,
+ cffn_ratio=cffn_ratio,
+ drop=drop,
+ drop_path=drop_path,
+ with_cp=with_cp,
+ )
+ if extra_extractor:
+ self.extra_extractors = nn.Sequential(
+ *[
+ Extractor(
+ dim=dim,
+ num_heads=num_heads,
+ n_points=n_points,
+ norm_layer=norm_layer,
+ with_cffn=with_cffn,
+ cffn_ratio=cffn_ratio,
+ deform_ratio=deform_ratio,
+ drop=drop,
+ drop_path=drop_path,
+ with_cp=with_cp,
+ )
+ for _ in range(2)
+ ]
+ )
+ else:
+ self.extra_extractors = None
+
+ def forward(self, x, c, cls, blocks, deform_inputs1, deform_inputs2, H_c, W_c, H_toks, W_toks):
+ x = self.injector(
+ query=x,
+ reference_points=deform_inputs1[0],
+ feat=c,
+ spatial_shapes=deform_inputs1[1],
+ level_start_index=deform_inputs1[2],
+ )
+ x = torch.cat((cls, x), dim=1)
+ for idx, blk in enumerate(blocks):
+ x = blk(x, H_toks, W_toks)
+ cls, x = (
+ x[
+ :,
+ :1,
+ ],
+ x[
+ :,
+ 1:,
+ ],
+ )
+ c = self.extractor(
+ query=c,
+ reference_points=deform_inputs2[0],
+ feat=x,
+ spatial_shapes=deform_inputs2[1],
+ level_start_index=deform_inputs2[2],
+ H=H_c,
+ W=W_c,
+ )
+ if self.extra_extractors is not None:
+ for extractor in self.extra_extractors:
+ c = extractor(
+ query=c,
+ reference_points=deform_inputs2[0],
+ feat=x,
+ spatial_shapes=deform_inputs2[1],
+ level_start_index=deform_inputs2[2],
+ H=H_c,
+ W=W_c,
+ )
+ return x, c, cls
+
+
+class SpatialPriorModule(nn.Module):
+ def __init__(self, inplanes=64, embed_dim=384, with_cp=False):
+ super().__init__()
+ self.with_cp = with_cp
+
+ self.stem = nn.Sequential(
+ *[
+ nn.Conv2d(3, inplanes, kernel_size=3, stride=2, padding=1, bias=False),
+ nn.SyncBatchNorm(inplanes),
+ nn.ReLU(inplace=True),
+ nn.Conv2d(inplanes, inplanes, kernel_size=3, stride=1, padding=1, bias=False),
+ nn.SyncBatchNorm(inplanes),
+ nn.ReLU(inplace=True),
+ nn.Conv2d(inplanes, inplanes, kernel_size=3, stride=1, padding=1, bias=False),
+ nn.SyncBatchNorm(inplanes),
+ nn.ReLU(inplace=True),
+ nn.MaxPool2d(kernel_size=3, stride=2, padding=1),
+ ]
+ )
+ self.conv2 = nn.Sequential(
+ *[
+ nn.Conv2d(inplanes, 2 * inplanes, kernel_size=3, stride=2, padding=1, bias=False),
+ nn.SyncBatchNorm(2 * inplanes),
+ nn.ReLU(inplace=True),
+ ]
+ )
+ self.conv3 = nn.Sequential(
+ *[
+ nn.Conv2d(2 * inplanes, 4 * inplanes, kernel_size=3, stride=2, padding=1, bias=False),
+ nn.SyncBatchNorm(4 * inplanes),
+ nn.ReLU(inplace=True),
+ ]
+ )
+ self.conv4 = nn.Sequential(
+ *[
+ nn.Conv2d(4 * inplanes, 4 * inplanes, kernel_size=3, stride=2, padding=1, bias=False),
+ nn.SyncBatchNorm(4 * inplanes),
+ nn.ReLU(inplace=True),
+ ]
+ )
+ self.fc1 = nn.Conv2d(inplanes, embed_dim, kernel_size=1, stride=1, padding=0, bias=True)
+ self.fc2 = nn.Conv2d(2 * inplanes, embed_dim, kernel_size=1, stride=1, padding=0, bias=True)
+ self.fc3 = nn.Conv2d(4 * inplanes, embed_dim, kernel_size=1, stride=1, padding=0, bias=True)
+ self.fc4 = nn.Conv2d(4 * inplanes, embed_dim, kernel_size=1, stride=1, padding=0, bias=True)
+
+ def forward(self, x):
+ def _inner_forward(x):
+ c1 = self.stem(x)
+ c2 = self.conv2(c1)
+ c3 = self.conv3(c2)
+ c4 = self.conv4(c3)
+ c1 = self.fc1(c1)
+ c2 = self.fc2(c2)
+ c3 = self.fc3(c3)
+ c4 = self.fc4(c4)
+
+ bs, dim, _, _ = c1.shape
+ # c1 = c1.view(bs, dim, -1).transpose(1, 2) # 4s
+ c2 = c2.view(bs, dim, -1).transpose(1, 2) # 8s
+ c3 = c3.view(bs, dim, -1).transpose(1, 2) # 16s
+ c4 = c4.view(bs, dim, -1).transpose(1, 2) # 32s
+
+ return c1, c2, c3, c4
+
+ if self.with_cp and x.requires_grad:
+ outs = cp.checkpoint(_inner_forward, x)
+ else:
+ outs = _inner_forward(x)
+ return outs
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/backbones/drop_path.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/backbones/drop_path.py
new file mode 100644
index 0000000000000000000000000000000000000000..864eb8738c44652d12b979fc811503f21cbb00dd
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/backbones/drop_path.py
@@ -0,0 +1,32 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+# References:
+# https://github.com/facebookresearch/dino/blob/master/vision_transformer.py
+# https://github.com/rwightman/pytorch-image-models/tree/master/timm/layers/drop.py
+
+from torch import nn
+
+
+def drop_path(x, drop_prob: float = 0.0, training: bool = False):
+ if drop_prob == 0.0 or not training:
+ return x
+ keep_prob = 1 - drop_prob
+ shape = (x.shape[0],) + (1,) * (x.ndim - 1) # work with diff dim tensors, not just 2D ConvNets
+ random_tensor = x.new_empty(shape).bernoulli_(keep_prob)
+ if keep_prob > 0.0:
+ random_tensor.div_(keep_prob)
+ return x * random_tensor
+
+
+class DropPath(nn.Module):
+ """Drop paths (Stochastic Depth) per sample (when applied in main path of residual blocks)."""
+
+ def __init__(self, drop_prob: float = 0.0):
+ super(DropPath, self).__init__()
+ self.drop_prob = drop_prob
+
+ def forward(self, x):
+ return drop_path(x, self.drop_prob, self.training)
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/backbones/vit.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/backbones/vit.py
new file mode 100644
index 0000000000000000000000000000000000000000..8a147570451bd2fbd016ddfafbbfa33035cbd4f8
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/backbones/vit.py
@@ -0,0 +1,552 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+"""Vision Transformer (ViT) in PyTorch.
+
+A PyTorch implement of Vision Transformers as described in:
+
+'An Image Is Worth 16 x 16 Words: Transformers for Image Recognition at Scale'
+ - https://arxiv.org/abs/2010.11929
+
+`How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers`
+ - https://arxiv.org/abs/2106.10270
+
+The official jax code is released and available at https://github.com/google-research/vision_transformer
+
+DeiT model defs and weights from https://github.com/facebookresearch/deit,
+paper `DeiT: Data-efficient Image Transformers` - https://arxiv.org/abs/2012.12877
+
+Acknowledgments:
+* The paper authors for releasing code and weights, thanks!
+* I fixed my class token impl based on Phil Wang's https://github.com/lucidrains/vit-pytorch ... check it out
+for some einops/einsum fun
+* Simple transformer style inspired by Andrej Karpathy's https://github.com/karpathy/minGPT
+* Bert reference code checks against Huggingface Transformers and Tensorflow Bert
+
+Hacked together by / Copyright 2021 Ross Wightman
+"""
+import logging
+import math
+from functools import partial
+from itertools import repeat
+from typing import Callable, Optional
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+import torch.utils.checkpoint as cp
+from mmcv.runner import BaseModule, load_checkpoint
+from mmseg.ops import resize
+from mmseg.utils import get_root_logger
+from torch import Tensor
+
+from .drop_path import DropPath
+
+
+def to_2tuple(x):
+ return tuple(repeat(x, 2))
+
+
+class Mlp(nn.Module):
+ def __init__(
+ self,
+ in_features: int,
+ hidden_features: Optional[int] = None,
+ out_features: Optional[int] = None,
+ act_layer: Callable[..., nn.Module] = nn.GELU,
+ drop: float = 0.0,
+ bias: bool = True,
+ ) -> None:
+ super().__init__()
+ out_features = out_features or in_features
+ hidden_features = hidden_features or in_features
+ self.fc1 = nn.Linear(in_features, hidden_features, bias=bias)
+ self.act = act_layer()
+ self.fc2 = nn.Linear(hidden_features, out_features, bias=bias)
+ self.drop = nn.Dropout(drop)
+
+ def forward(self, x: Tensor) -> Tensor:
+ x = self.fc1(x)
+ x = self.act(x)
+ x = self.drop(x)
+ x = self.fc2(x)
+ x = self.drop(x)
+ return x
+
+
+class SwiGLUFFN(nn.Module):
+ def __init__(
+ self,
+ in_features: int,
+ hidden_features: Optional[int] = None,
+ out_features: Optional[int] = None,
+ act_layer: Callable[..., nn.Module] = None,
+ drop: float = 0.0,
+ ) -> None:
+ super().__init__()
+ out_features = out_features or in_features
+ hidden_features = hidden_features or in_features
+ swiglu_hidden_features = int(2 * hidden_features / 3)
+ align_as = 8
+ swiglu_hidden_features = (swiglu_hidden_features + align_as - 1) // align_as * align_as
+ self.w1 = nn.Linear(in_features, swiglu_hidden_features)
+ self.w2 = nn.Linear(in_features, swiglu_hidden_features)
+ self.w3 = nn.Linear(swiglu_hidden_features, out_features)
+
+ def forward(self, x: Tensor) -> Tensor:
+ x1 = self.w1(x)
+ x2 = self.w2(x)
+ hidden = F.silu(x1) * x2
+ return self.w3(hidden)
+
+
+class PatchEmbed(nn.Module):
+ """2D Image to Patch Embedding."""
+
+ def __init__(
+ self, img_size=224, patch_size=16, in_chans=3, embed_dim=768, norm_layer=None, flatten=True, bias=True
+ ):
+ super().__init__()
+ img_size = to_2tuple(img_size)
+ patch_size = to_2tuple(patch_size)
+ self.img_size = img_size
+ self.patch_size = patch_size
+ self.grid_size = (img_size[0] // patch_size[0], img_size[1] // patch_size[1])
+ self.num_patches = self.grid_size[0] * self.grid_size[1]
+ self.flatten = flatten
+
+ self.proj = nn.Conv2d(in_chans, embed_dim, kernel_size=patch_size, stride=patch_size, bias=bias)
+ self.norm = norm_layer(embed_dim) if norm_layer else nn.Identity()
+
+ def forward(self, x):
+ x = self.proj(x)
+ _, _, H, W = x.shape
+ if self.flatten:
+ x = x.flatten(2).transpose(1, 2) # BCHW -> BNC
+ x = self.norm(x)
+ return x, H, W
+
+
+class Attention(nn.Module):
+ def __init__(self, dim, num_heads=8, qkv_bias=False, attn_drop=0.0, proj_drop=0.0):
+ super().__init__()
+ self.num_heads = num_heads
+ head_dim = dim // num_heads
+ self.scale = head_dim**-0.5
+
+ self.qkv = nn.Linear(dim, dim * 3, bias=qkv_bias)
+ self.attn_drop = nn.Dropout(attn_drop)
+ self.proj = nn.Linear(dim, dim)
+ self.proj_drop = nn.Dropout(proj_drop)
+
+ def forward(self, x, H, W):
+ B, N, C = x.shape
+ qkv = self.qkv(x).reshape(B, N, 3, self.num_heads, C // self.num_heads).permute(2, 0, 3, 1, 4)
+ q, k, v = qkv.unbind(0) # make torchscript happy (cannot use tensor as tuple)
+
+ attn = (q @ k.transpose(-2, -1)) * self.scale
+ attn = attn.softmax(dim=-1)
+ attn = self.attn_drop(attn)
+
+ x = (attn @ v).transpose(1, 2).reshape(B, N, C)
+ x = self.proj(x)
+ x = self.proj_drop(x)
+ return x
+
+
+class MemEffAttention(nn.Module):
+ def __init__(
+ self,
+ dim: int,
+ num_heads: int = 8,
+ qkv_bias: bool = False,
+ attn_drop: float = 0.0,
+ proj_drop: float = 0.0,
+ ) -> None:
+ super().__init__()
+ self.num_heads = num_heads
+ head_dim = dim // num_heads
+ self.scale = head_dim**-0.5
+
+ self.qkv = nn.Linear(dim, dim * 3, bias=qkv_bias)
+ self.attn_drop = nn.Dropout(attn_drop)
+ self.proj = nn.Linear(dim, dim)
+ self.proj_drop = nn.Dropout(proj_drop)
+
+ def forward(self, x: Tensor, H, W) -> Tensor:
+ from xformers.ops import memory_efficient_attention, unbind
+
+ B, N, C = x.shape
+ qkv = self.qkv(x).reshape(B, N, 3, self.num_heads, C // self.num_heads)
+
+ q, k, v = unbind(qkv, 2)
+
+ x = memory_efficient_attention(q, k, v)
+ x = x.reshape([B, N, C])
+
+ x = self.proj(x)
+ x = self.proj_drop(x)
+ return x
+
+
+def window_partition(x, window_size):
+ """
+ Args:
+ x: (B, H, W, C)
+ window_size (int): window size
+ Returns:
+ windows: (num_windows*B, window_size, window_size, C)
+ """
+ B, H, W, C = x.shape
+ x = x.view(B, H // window_size, window_size, W // window_size, window_size, C)
+ windows = x.permute(0, 1, 3, 2, 4, 5).contiguous().view(-1, window_size, window_size, C)
+ return windows
+
+
+def window_reverse(windows, window_size, H, W):
+ """
+ Args:
+ windows: (num_windows*B, window_size, window_size, C)
+ window_size (int): Window size
+ H (int): Height of image
+ W (int): Width of image
+ Returns:
+ x: (B, H, W, C)
+ """
+ B = int(windows.shape[0] / (H * W / window_size / window_size))
+ x = windows.view(B, H // window_size, W // window_size, window_size, window_size, -1)
+ x = x.permute(0, 1, 3, 2, 4, 5).contiguous().view(B, H, W, -1)
+ return x
+
+
+class WindowedAttention(nn.Module):
+ def __init__(
+ self, dim, num_heads=8, qkv_bias=False, attn_drop=0.0, proj_drop=0.0, window_size=14, pad_mode="constant"
+ ):
+ super().__init__()
+ self.num_heads = num_heads
+ head_dim = dim // num_heads
+ self.scale = head_dim**-0.5
+
+ self.qkv = nn.Linear(dim, dim * 3, bias=qkv_bias)
+ self.attn_drop = nn.Dropout(attn_drop)
+ self.proj = nn.Linear(dim, dim)
+ self.proj_drop = nn.Dropout(proj_drop)
+ self.window_size = window_size
+ self.pad_mode = pad_mode
+
+ def forward(self, x, H, W):
+ B, N, C = x.shape
+ N_ = self.window_size * self.window_size
+ H_ = math.ceil(H / self.window_size) * self.window_size
+ W_ = math.ceil(W / self.window_size) * self.window_size
+
+ qkv = self.qkv(x) # [B, N, C]
+ qkv = qkv.transpose(1, 2).reshape(B, C * 3, H, W) # [B, C, H, W]
+ qkv = F.pad(qkv, [0, W_ - W, 0, H_ - H], mode=self.pad_mode)
+
+ qkv = F.unfold(
+ qkv, kernel_size=(self.window_size, self.window_size), stride=(self.window_size, self.window_size)
+ )
+ B, C_kw_kw, L = qkv.shape # L - the num of windows
+ qkv = qkv.reshape(B, C * 3, N_, L).permute(0, 3, 2, 1) # [B, L, N_, C]
+ qkv = qkv.reshape(B, L, N_, 3, self.num_heads, C // self.num_heads).permute(3, 0, 1, 4, 2, 5)
+ q, k, v = qkv.unbind(0) # make torchscript happy (cannot use tensor as tuple)
+
+ # q,k,v [B, L, num_head, N_, C/num_head]
+ attn = (q @ k.transpose(-2, -1)) * self.scale # [B, L, num_head, N_, N_]
+ # if self.mask:
+ # attn = attn * mask
+ attn = attn.softmax(dim=-1)
+ attn = self.attn_drop(attn) # [B, L, num_head, N_, N_]
+ # attn @ v = [B, L, num_head, N_, C/num_head]
+ x = (attn @ v).permute(0, 2, 4, 3, 1).reshape(B, C_kw_kw // 3, L)
+
+ x = F.fold(
+ x,
+ output_size=(H_, W_),
+ kernel_size=(self.window_size, self.window_size),
+ stride=(self.window_size, self.window_size),
+ ) # [B, C, H_, W_]
+ x = x[:, :, :H, :W].reshape(B, C, N).transpose(-1, -2)
+ x = self.proj(x)
+ x = self.proj_drop(x)
+ return x
+
+
+# class WindowedAttention(nn.Module):
+# def __init__(self, dim, num_heads=8, qkv_bias=False, attn_drop=0., proj_drop=0., window_size=14, pad_mode="constant"):
+# super().__init__()
+# self.num_heads = num_heads
+# head_dim = dim // num_heads
+# self.scale = head_dim ** -0.5
+#
+# self.qkv = nn.Linear(dim, dim * 3, bias=qkv_bias)
+# self.attn_drop = nn.Dropout(attn_drop)
+# self.proj = nn.Linear(dim, dim)
+# self.proj_drop = nn.Dropout(proj_drop)
+# self.window_size = window_size
+# self.pad_mode = pad_mode
+#
+# def forward(self, x, H, W):
+# B, N, C = x.shape
+#
+# N_ = self.window_size * self.window_size
+# H_ = math.ceil(H / self.window_size) * self.window_size
+# W_ = math.ceil(W / self.window_size) * self.window_size
+# x = x.view(B, H, W, C)
+# x = F.pad(x, [0, 0, 0, W_ - W, 0, H_- H], mode=self.pad_mode)
+#
+# x = window_partition(x, window_size=self.window_size)# nW*B, window_size, window_size, C
+# x = x.view(-1, N_, C)
+#
+# qkv = self.qkv(x).view(-1, N_, 3, self.num_heads, C // self.num_heads).permute(2, 0, 3, 1, 4)
+# q, k, v = qkv.unbind(0) # make torchscript happy (cannot use tensor as tuple)
+# attn = (q @ k.transpose(-2, -1)) * self.scale # [B, L, num_head, N_, N_]
+# attn = attn.softmax(dim=-1)
+# attn = self.attn_drop(attn) # [B, L, num_head, N_, N_]
+# x = (attn @ v).transpose(1, 2).reshape(-1, self.window_size, self.window_size, C)
+#
+# x = window_reverse(x, self.window_size, H_, W_)
+# x = x[:, :H, :W, :].reshape(B, N, C).contiguous()
+# x = self.proj(x)
+# x = self.proj_drop(x)
+# return x
+
+
+class Block(nn.Module):
+ def __init__(
+ self,
+ dim,
+ num_heads,
+ mlp_ratio=4.0,
+ qkv_bias=False,
+ drop=0.0,
+ attn_drop=0.0,
+ drop_path=0.0,
+ act_layer=nn.GELU,
+ norm_layer=nn.LayerNorm,
+ windowed=False,
+ window_size=14,
+ pad_mode="constant",
+ layer_scale=False,
+ with_cp=False,
+ ffn_layer=Mlp,
+ memeff=False,
+ ):
+ super().__init__()
+ self.with_cp = with_cp
+ self.norm1 = norm_layer(dim)
+ if windowed:
+ self.attn = WindowedAttention(
+ dim,
+ num_heads=num_heads,
+ qkv_bias=qkv_bias,
+ attn_drop=attn_drop,
+ proj_drop=drop,
+ window_size=window_size,
+ pad_mode=pad_mode,
+ )
+ elif memeff:
+ self.attn = MemEffAttention(
+ dim, num_heads=num_heads, qkv_bias=qkv_bias, attn_drop=attn_drop, proj_drop=drop
+ )
+ else:
+ self.attn = Attention(dim, num_heads=num_heads, qkv_bias=qkv_bias, attn_drop=attn_drop, proj_drop=drop)
+ # NOTE: drop path for stochastic depth, we shall see if this is better than dropout here
+ self.drop_path = DropPath(drop_path) if drop_path > 0.0 else nn.Identity()
+ self.norm2 = norm_layer(dim)
+ mlp_hidden_dim = int(dim * mlp_ratio)
+ self.mlp = ffn_layer(in_features=dim, hidden_features=mlp_hidden_dim, act_layer=act_layer, drop=drop)
+ self.layer_scale = layer_scale
+ if layer_scale:
+ self.gamma1 = nn.Parameter(torch.ones((dim)), requires_grad=True)
+ self.gamma2 = nn.Parameter(torch.ones((dim)), requires_grad=True)
+
+ def forward(self, x, H, W):
+ def _inner_forward(x):
+ if self.layer_scale:
+ x = x + self.drop_path(self.gamma1 * self.attn(self.norm1(x), H, W))
+ x = x + self.drop_path(self.gamma2 * self.mlp(self.norm2(x)))
+ else:
+ x = x + self.drop_path(self.attn(self.norm1(x), H, W))
+ x = x + self.drop_path(self.mlp(self.norm2(x)))
+ return x
+
+ if self.with_cp and x.requires_grad:
+ x = cp.checkpoint(_inner_forward, x)
+ else:
+ x = _inner_forward(x)
+
+ return x
+
+
+class TIMMVisionTransformer(BaseModule):
+ """Vision Transformer.
+
+ A PyTorch impl of : `An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale`
+ - https://arxiv.org/abs/2010.11929
+
+ Includes distillation token & head support for `DeiT: Data-efficient Image Transformers`
+ - https://arxiv.org/abs/2012.12877
+ """
+
+ def __init__(
+ self,
+ img_size=224,
+ patch_size=16,
+ in_chans=3,
+ num_classes=1000,
+ embed_dim=768,
+ depth=12,
+ num_heads=12,
+ mlp_ratio=4.0,
+ qkv_bias=True,
+ drop_rate=0.0,
+ attn_drop_rate=0.0,
+ drop_path_rate=0.0,
+ layer_scale=True,
+ embed_layer=PatchEmbed,
+ norm_layer=partial(nn.LayerNorm, eps=1e-6),
+ act_layer=nn.GELU,
+ window_attn=False,
+ window_size=14,
+ pretrained=None,
+ with_cp=False,
+ pre_norm=False,
+ ffn_type="mlp",
+ memeff=False,
+ ):
+ """
+ Args:
+ img_size (int, tuple): input image size
+ patch_size (int, tuple): patch size
+ in_chans (int): number of input channels
+ num_classes (int): number of classes for classification head
+ embed_dim (int): embedding dimension
+ depth (int): depth of transformer
+ num_heads (int): number of attention heads
+ mlp_ratio (int): ratio of mlp hidden dim to embedding dim
+ qkv_bias (bool): enable bias for qkv if True
+ drop_rate (float): dropout rate
+ attn_drop_rate (float): attention dropout rate
+ drop_path_rate (float): stochastic depth rate
+ embed_layer (nn.Module): patch embedding layer
+ norm_layer: (nn.Module): normalization layer
+ pretrained: (str): pretrained path
+ """
+ super().__init__()
+ self.num_classes = num_classes
+ self.num_features = self.embed_dim = embed_dim # num_features for consistency with other models
+ self.num_tokens = 1
+ norm_layer = norm_layer or partial(nn.LayerNorm, eps=1e-6)
+ act_layer = act_layer or nn.GELU
+ self.norm_layer = norm_layer
+ self.act_layer = act_layer
+ self.pretrain_size = img_size
+ self.drop_path_rate = drop_path_rate
+ self.drop_rate = drop_rate
+ self.patch_size = patch_size
+
+ window_attn = [window_attn] * depth if not isinstance(window_attn, list) else window_attn
+ window_size = [window_size] * depth if not isinstance(window_size, list) else window_size
+ logging.info("window attention:", window_attn)
+ logging.info("window size:", window_size)
+ logging.info("layer scale:", layer_scale)
+
+ self.patch_embed = embed_layer(
+ img_size=img_size, patch_size=patch_size, in_chans=in_chans, embed_dim=embed_dim, bias=not pre_norm
+ )
+ num_patches = self.patch_embed.num_patches
+
+ self.pos_embed = nn.Parameter(torch.zeros(1, num_patches + self.num_tokens, embed_dim))
+ self.pos_drop = nn.Dropout(p=drop_rate)
+
+ ffn_types = {"mlp": Mlp, "swiglu": SwiGLUFFN}
+
+ dpr = [x.item() for x in torch.linspace(0, drop_path_rate, depth)] # stochastic depth decay rule
+ self.blocks = nn.Sequential(
+ *[
+ Block(
+ dim=embed_dim,
+ num_heads=num_heads,
+ mlp_ratio=mlp_ratio,
+ qkv_bias=qkv_bias,
+ drop=drop_rate,
+ attn_drop=attn_drop_rate,
+ drop_path=dpr[i],
+ norm_layer=norm_layer,
+ act_layer=act_layer,
+ windowed=window_attn[i],
+ window_size=window_size[i],
+ layer_scale=layer_scale,
+ with_cp=with_cp,
+ ffn_layer=ffn_types[ffn_type],
+ memeff=memeff,
+ )
+ for i in range(depth)
+ ]
+ )
+
+ # self.norm = norm_layer(embed_dim)
+ self.cls_token = nn.Parameter(torch.zeros(1, 1, embed_dim))
+ # For CLIP
+ if pre_norm:
+ norm_pre = norm_layer(embed_dim)
+ self.norm_pre = norm_pre
+ else:
+ self.norm_pre = nn.Identity()
+ self.init_weights(pretrained)
+
+ def init_weights(self, pretrained=None):
+ if isinstance(pretrained, str):
+ logger = get_root_logger()
+ load_checkpoint(self, pretrained, map_location="cpu", strict=False, logger=logger)
+
+ def forward_features(self, x):
+ x, H, W = self.patch_embed(x)
+ cls_token = self.cls_token.expand(x.shape[0], -1, -1) # stole cls_tokens impl from Phil Wang, thanks
+ x = torch.cat((cls_token, x), dim=1)
+ x = self.pos_drop(x + self.pos_embed)
+
+ # For CLIP
+ x = self.norm_pre(x)
+
+ for blk in self.blocks:
+ x = blk(x, H, W)
+ x = self.norm(x)
+ return x
+
+ def forward(self, x):
+ x = self.forward_features(x)
+ return x
+
+ @staticmethod
+ def resize_pos_embed(pos_embed, input_shpae, pos_shape, mode):
+ """Resize pos_embed weights.
+
+ Resize pos_embed using bicubic interpolate method.
+ Args:
+ pos_embed (torch.Tensor): Position embedding weights.
+ input_shpae (tuple): Tuple for (downsampled input image height,
+ downsampled input image width).
+ pos_shape (tuple): The resolution of downsampled origin training
+ image.
+ mode (str): Algorithm used for upsampling:
+ ``'nearest'`` | ``'linear'`` | ``'bilinear'`` | ``'bicubic'`` |
+ ``'trilinear'``. Default: ``'nearest'``
+ Return:
+ torch.Tensor: The resized pos_embed of shape [B, L_new, C]
+ """
+ assert pos_embed.ndim == 3, "shape of pos_embed must be [B, L, C]"
+ pos_h, pos_w = pos_shape
+ # keep dim for easy deployment
+ cls_token_weight = pos_embed[:, 0:1]
+ pos_embed_weight = pos_embed[:, (-1 * pos_h * pos_w) :]
+ pos_embed_weight = pos_embed_weight.reshape(1, pos_h, pos_w, pos_embed.shape[2]).permute(0, 3, 1, 2)
+ pos_embed_weight = resize(pos_embed_weight, size=input_shpae, align_corners=False, mode=mode)
+ pos_embed_weight = torch.flatten(pos_embed_weight, 2).transpose(1, 2)
+ pos_embed = torch.cat((cls_token_weight, pos_embed_weight), dim=1)
+ return pos_embed
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/backbones/vit_adapter.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/backbones/vit_adapter.py
new file mode 100644
index 0000000000000000000000000000000000000000..ebc4f0f65e04ed764464d141607b3b2073220f6b
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/backbones/vit_adapter.py
@@ -0,0 +1,217 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import math
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmseg.models.builder import BACKBONES
+from torch.nn.init import normal_
+
+from ...ops.modules import MSDeformAttn
+from .adapter_modules import InteractionBlock, InteractionBlockWithCls, SpatialPriorModule, deform_inputs
+from .vit import TIMMVisionTransformer
+
+
+@BACKBONES.register_module()
+class ViTAdapter(TIMMVisionTransformer):
+ def __init__(
+ self,
+ pretrain_size=224,
+ num_heads=12,
+ conv_inplane=64,
+ n_points=4,
+ deform_num_heads=6,
+ init_values=0.0,
+ interaction_indexes=None,
+ with_cffn=True,
+ cffn_ratio=0.25,
+ deform_ratio=1.0,
+ add_vit_feature=True,
+ pretrained=None,
+ use_extra_extractor=True,
+ freeze_vit=False,
+ use_cls=True,
+ with_cp=False,
+ *args,
+ **kwargs
+ ):
+
+ super().__init__(num_heads=num_heads, pretrained=pretrained, with_cp=with_cp, *args, **kwargs)
+ if freeze_vit:
+ for param in self.parameters():
+ param.requires_grad = False
+
+ # self.num_classes = 80
+ self.use_cls = use_cls
+ if not self.use_cls:
+ self.cls_token = None
+ self.num_block = len(self.blocks)
+ self.pretrain_size = (pretrain_size, pretrain_size)
+ self.interaction_indexes = interaction_indexes
+ self.add_vit_feature = add_vit_feature
+ embed_dim = self.embed_dim
+
+ block_fn = InteractionBlockWithCls if use_cls else InteractionBlock
+
+ self.level_embed = nn.Parameter(torch.zeros(3, embed_dim))
+ self.spm = SpatialPriorModule(inplanes=conv_inplane, embed_dim=embed_dim, with_cp=False)
+ self.interactions = nn.Sequential(
+ *[
+ block_fn(
+ dim=embed_dim,
+ num_heads=deform_num_heads,
+ n_points=n_points,
+ init_values=init_values,
+ drop_path=self.drop_path_rate,
+ norm_layer=self.norm_layer,
+ with_cffn=with_cffn,
+ cffn_ratio=cffn_ratio,
+ deform_ratio=deform_ratio,
+ extra_extractor=((True if i == len(interaction_indexes) - 1 else False) and use_extra_extractor),
+ with_cp=with_cp,
+ )
+ for i in range(len(interaction_indexes))
+ ]
+ )
+ self.up = nn.ConvTranspose2d(embed_dim, embed_dim, 2, 2)
+ self.norm1 = nn.SyncBatchNorm(embed_dim)
+ self.norm2 = nn.SyncBatchNorm(embed_dim)
+ self.norm3 = nn.SyncBatchNorm(embed_dim)
+ self.norm4 = nn.SyncBatchNorm(embed_dim)
+
+ self.up.apply(self._init_weights)
+ self.spm.apply(self._init_weights)
+ self.interactions.apply(self._init_weights)
+ self.apply(self._init_deform_weights)
+ normal_(self.level_embed)
+
+ def _init_weights(self, m):
+ if isinstance(m, nn.Linear):
+ torch.nn.init.trunc_normal_(m.weight, std=0.02)
+ if isinstance(m, nn.Linear) and m.bias is not None:
+ nn.init.constant_(m.bias, 0)
+ elif isinstance(m, nn.LayerNorm) or isinstance(m, nn.BatchNorm2d):
+ nn.init.constant_(m.bias, 0)
+ nn.init.constant_(m.weight, 1.0)
+ elif isinstance(m, nn.Conv2d) or isinstance(m, nn.ConvTranspose2d):
+ fan_out = m.kernel_size[0] * m.kernel_size[1] * m.out_channels
+ fan_out //= m.groups
+ m.weight.data.normal_(0, math.sqrt(2.0 / fan_out))
+ if m.bias is not None:
+ m.bias.data.zero_()
+
+ def _get_pos_embed(self, pos_embed, H, W):
+ pos_embed = pos_embed.reshape(
+ 1, self.pretrain_size[0] // self.patch_size, self.pretrain_size[1] // self.patch_size, -1
+ ).permute(0, 3, 1, 2)
+ pos_embed = (
+ F.interpolate(pos_embed, size=(H, W), mode="bicubic", align_corners=False)
+ .reshape(1, -1, H * W)
+ .permute(0, 2, 1)
+ )
+ return pos_embed
+
+ def _init_deform_weights(self, m):
+ if isinstance(m, MSDeformAttn):
+ m._reset_parameters()
+
+ def _add_level_embed(self, c2, c3, c4):
+ c2 = c2 + self.level_embed[0]
+ c3 = c3 + self.level_embed[1]
+ c4 = c4 + self.level_embed[2]
+ return c2, c3, c4
+
+ def forward(self, x):
+ deform_inputs1, deform_inputs2 = deform_inputs(x, self.patch_size)
+
+ # SPM forward
+ c1, c2, c3, c4 = self.spm(x)
+ c2, c3, c4 = self._add_level_embed(c2, c3, c4)
+ c = torch.cat([c2, c3, c4], dim=1)
+
+ # Patch Embedding forward
+ H_c, W_c = x.shape[2] // 16, x.shape[3] // 16
+ x, H_toks, W_toks = self.patch_embed(x)
+ # print("H_toks, W_toks =", H_toks, W_toks)
+ bs, n, dim = x.shape
+ pos_embed = self._get_pos_embed(self.pos_embed[:, 1:], H_toks, W_toks)
+ if self.use_cls:
+ cls_token = self.cls_token.expand(x.shape[0], -1, -1) # stole cls_tokens impl from Phil Wang, thanks
+ x = torch.cat((cls_token, x), dim=1)
+ pos_embed = torch.cat((self.pos_embed[:, :1], pos_embed), dim=1)
+ x = self.pos_drop(x + pos_embed)
+ # For CLIP
+ x = self.norm_pre(x)
+
+ # Interaction
+ if self.use_cls:
+ cls, x = (
+ x[
+ :,
+ :1,
+ ],
+ x[
+ :,
+ 1:,
+ ],
+ )
+ outs = list()
+ for i, layer in enumerate(self.interactions):
+ indexes = self.interaction_indexes[i]
+ if self.use_cls:
+ x, c, cls = layer(
+ x,
+ c,
+ cls,
+ self.blocks[indexes[0] : indexes[-1] + 1],
+ deform_inputs1,
+ deform_inputs2,
+ H_c,
+ W_c,
+ H_toks,
+ W_toks,
+ )
+ else:
+ x, c = layer(
+ x,
+ c,
+ self.blocks[indexes[0] : indexes[-1] + 1],
+ deform_inputs1,
+ deform_inputs2,
+ H_c,
+ W_c,
+ H_toks,
+ W_toks,
+ )
+ outs.append(x.transpose(1, 2).view(bs, dim, H_toks, W_toks).contiguous())
+
+ # Split & Reshape
+ c2 = c[:, 0 : c2.size(1), :]
+ c3 = c[:, c2.size(1) : c2.size(1) + c3.size(1), :]
+ c4 = c[:, c2.size(1) + c3.size(1) :, :]
+
+ c2 = c2.transpose(1, 2).view(bs, dim, H_c * 2, W_c * 2).contiguous()
+ c3 = c3.transpose(1, 2).view(bs, dim, H_c, W_c).contiguous()
+ c4 = c4.transpose(1, 2).view(bs, dim, H_c // 2, W_c // 2).contiguous()
+ c1 = self.up(c2) + c1
+
+ if self.add_vit_feature:
+ x1, x2, x3, x4 = outs
+
+ x1 = F.interpolate(x1, size=(4 * H_c, 4 * W_c), mode="bilinear", align_corners=False)
+ x2 = F.interpolate(x2, size=(2 * H_c, 2 * W_c), mode="bilinear", align_corners=False)
+ x3 = F.interpolate(x3, size=(1 * H_c, 1 * W_c), mode="bilinear", align_corners=False)
+ x4 = F.interpolate(x4, size=(H_c // 2, W_c // 2), mode="bilinear", align_corners=False)
+ # print(c1.shape, c2.shape, c3.shape, c4.shape, x1.shape, x2.shape, x3.shape, x4.shape, H_c, H_toks)
+ c1, c2, c3, c4 = c1 + x1, c2 + x2, c3 + x3, c4 + x4
+
+ # Final Norm
+ f1 = self.norm1(c1)
+ f2 = self.norm2(c2)
+ f3 = self.norm3(c3)
+ f4 = self.norm4(c4)
+ return [f1, f2, f3, f4]
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/builder.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/builder.py
new file mode 100644
index 0000000000000000000000000000000000000000..d7cf7b919f6b0e8e00bde45bc244d9c29a36fed6
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/builder.py
@@ -0,0 +1,25 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from mmcv.utils import Registry
+
+TRANSFORMER = Registry("Transformer")
+MASK_ASSIGNERS = Registry("mask_assigner")
+MATCH_COST = Registry("match_cost")
+
+
+def build_match_cost(cfg):
+ """Build Match Cost."""
+ return MATCH_COST.build(cfg)
+
+
+def build_assigner(cfg):
+ """Build Assigner."""
+ return MASK_ASSIGNERS.build(cfg)
+
+
+def build_transformer(cfg):
+ """Build Transformer."""
+ return TRANSFORMER.build(cfg)
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/decode_heads/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/decode_heads/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..01f08b88950750337781fc671adfea2a935ea8fe
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/decode_heads/__init__.py
@@ -0,0 +1,6 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .mask2former_head import Mask2FormerHead
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/decode_heads/mask2former_head.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/decode_heads/mask2former_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..d1705fc444fa8d1583d88fca36d7fe1e060db9e7
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/decode_heads/mask2former_head.py
@@ -0,0 +1,544 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import copy
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import Conv2d, build_plugin_layer, caffe2_xavier_init
+from mmcv.cnn.bricks.transformer import build_positional_encoding, build_transformer_layer_sequence
+from mmcv.ops import point_sample
+from mmcv.runner import ModuleList, force_fp32
+from mmseg.models.builder import HEADS, build_loss
+from mmseg.models.decode_heads.decode_head import BaseDecodeHead
+
+from ...core import build_sampler, multi_apply, reduce_mean
+from ..builder import build_assigner
+from ..utils import get_uncertain_point_coords_with_randomness
+
+
+@HEADS.register_module()
+class Mask2FormerHead(BaseDecodeHead):
+ """Implements the Mask2Former head.
+
+ See `Masked-attention Mask Transformer for Universal Image
+ Segmentation `_ for details.
+
+ Args:
+ in_channels (list[int]): Number of channels in the input feature map.
+ feat_channels (int): Number of channels for features.
+ out_channels (int): Number of channels for output.
+ num_things_classes (int): Number of things.
+ num_stuff_classes (int): Number of stuff.
+ num_queries (int): Number of query in Transformer decoder.
+ pixel_decoder (:obj:`mmcv.ConfigDict` | dict): Config for pixel
+ decoder. Defaults to None.
+ enforce_decoder_input_project (bool, optional): Whether to add
+ a layer to change the embed_dim of tranformer encoder in
+ pixel decoder to the embed_dim of transformer decoder.
+ Defaults to False.
+ transformer_decoder (:obj:`mmcv.ConfigDict` | dict): Config for
+ transformer decoder. Defaults to None.
+ positional_encoding (:obj:`mmcv.ConfigDict` | dict): Config for
+ transformer decoder position encoding. Defaults to None.
+ loss_cls (:obj:`mmcv.ConfigDict` | dict): Config of the classification
+ loss. Defaults to None.
+ loss_mask (:obj:`mmcv.ConfigDict` | dict): Config of the mask loss.
+ Defaults to None.
+ loss_dice (:obj:`mmcv.ConfigDict` | dict): Config of the dice loss.
+ Defaults to None.
+ train_cfg (:obj:`mmcv.ConfigDict` | dict): Training config of
+ Mask2Former head.
+ test_cfg (:obj:`mmcv.ConfigDict` | dict): Testing config of
+ Mask2Former head.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def __init__(
+ self,
+ in_channels,
+ feat_channels,
+ out_channels,
+ num_things_classes=80,
+ num_stuff_classes=53,
+ num_queries=100,
+ num_transformer_feat_level=3,
+ pixel_decoder=None,
+ enforce_decoder_input_project=False,
+ transformer_decoder=None,
+ positional_encoding=None,
+ loss_cls=None,
+ loss_mask=None,
+ loss_dice=None,
+ train_cfg=None,
+ test_cfg=None,
+ init_cfg=None,
+ **kwargs,
+ ):
+ super(Mask2FormerHead, self).__init__(
+ in_channels=in_channels,
+ channels=feat_channels,
+ num_classes=(num_things_classes + num_stuff_classes),
+ init_cfg=init_cfg,
+ input_transform="multiple_select",
+ **kwargs,
+ )
+ self.num_things_classes = num_things_classes
+ self.num_stuff_classes = num_stuff_classes
+ self.num_classes = self.num_things_classes + self.num_stuff_classes
+ self.num_queries = num_queries
+ self.num_transformer_feat_level = num_transformer_feat_level
+ self.num_heads = transformer_decoder.transformerlayers.attn_cfgs.num_heads
+ self.num_transformer_decoder_layers = transformer_decoder.num_layers
+ assert pixel_decoder.encoder.transformerlayers.attn_cfgs.num_levels == num_transformer_feat_level
+ pixel_decoder_ = copy.deepcopy(pixel_decoder)
+ pixel_decoder_.update(in_channels=in_channels, feat_channels=feat_channels, out_channels=out_channels)
+ self.pixel_decoder = build_plugin_layer(pixel_decoder_)[1]
+ self.transformer_decoder = build_transformer_layer_sequence(transformer_decoder)
+ self.decoder_embed_dims = self.transformer_decoder.embed_dims
+
+ self.decoder_input_projs = ModuleList()
+ # from low resolution to high resolution
+ for _ in range(num_transformer_feat_level):
+ if self.decoder_embed_dims != feat_channels or enforce_decoder_input_project:
+ self.decoder_input_projs.append(Conv2d(feat_channels, self.decoder_embed_dims, kernel_size=1))
+ else:
+ self.decoder_input_projs.append(nn.Identity())
+ self.decoder_positional_encoding = build_positional_encoding(positional_encoding)
+ self.query_embed = nn.Embedding(self.num_queries, feat_channels)
+ self.query_feat = nn.Embedding(self.num_queries, feat_channels)
+ # from low resolution to high resolution
+ self.level_embed = nn.Embedding(self.num_transformer_feat_level, feat_channels)
+
+ self.cls_embed = nn.Linear(feat_channels, self.num_classes + 1)
+ self.mask_embed = nn.Sequential(
+ nn.Linear(feat_channels, feat_channels),
+ nn.ReLU(inplace=True),
+ nn.Linear(feat_channels, feat_channels),
+ nn.ReLU(inplace=True),
+ nn.Linear(feat_channels, out_channels),
+ )
+ self.conv_seg = None # fix a bug here (conv_seg is not used)
+
+ self.test_cfg = test_cfg
+ self.train_cfg = train_cfg
+ if train_cfg:
+ self.assigner = build_assigner(self.train_cfg.assigner)
+ self.sampler = build_sampler(self.train_cfg.sampler, context=self)
+ self.num_points = self.train_cfg.get("num_points", 12544)
+ self.oversample_ratio = self.train_cfg.get("oversample_ratio", 3.0)
+ self.importance_sample_ratio = self.train_cfg.get("importance_sample_ratio", 0.75)
+
+ self.class_weight = loss_cls.class_weight
+ self.loss_cls = build_loss(loss_cls)
+ self.loss_mask = build_loss(loss_mask)
+ self.loss_dice = build_loss(loss_dice)
+
+ def init_weights(self):
+ for m in self.decoder_input_projs:
+ if isinstance(m, Conv2d):
+ caffe2_xavier_init(m, bias=0)
+
+ self.pixel_decoder.init_weights()
+
+ for p in self.transformer_decoder.parameters():
+ if p.dim() > 1:
+ nn.init.xavier_normal_(p)
+
+ def get_targets(self, cls_scores_list, mask_preds_list, gt_labels_list, gt_masks_list, img_metas):
+ """Compute classification and mask targets for all images for a decoder
+ layer.
+
+ Args:
+ cls_scores_list (list[Tensor]): Mask score logits from a single
+ decoder layer for all images. Each with shape [num_queries,
+ cls_out_channels].
+ mask_preds_list (list[Tensor]): Mask logits from a single decoder
+ layer for all images. Each with shape [num_queries, h, w].
+ gt_labels_list (list[Tensor]): Ground truth class indices for all
+ images. Each with shape (n, ), n is the sum of number of stuff
+ type and number of instance in a image.
+ gt_masks_list (list[Tensor]): Ground truth mask for each image,
+ each with shape (n, h, w).
+ img_metas (list[dict]): List of image meta information.
+
+ Returns:
+ tuple[list[Tensor]]: a tuple containing the following targets.
+
+ - labels_list (list[Tensor]): Labels of all images.
+ Each with shape [num_queries, ].
+ - label_weights_list (list[Tensor]): Label weights of all
+ images.Each with shape [num_queries, ].
+ - mask_targets_list (list[Tensor]): Mask targets of all images.
+ Each with shape [num_queries, h, w].
+ - mask_weights_list (list[Tensor]): Mask weights of all images.
+ Each with shape [num_queries, ].
+ - num_total_pos (int): Number of positive samples in all
+ images.
+ - num_total_neg (int): Number of negative samples in all
+ images.
+ """
+ (
+ labels_list,
+ label_weights_list,
+ mask_targets_list,
+ mask_weights_list,
+ pos_inds_list,
+ neg_inds_list,
+ ) = multi_apply(
+ self._get_target_single, cls_scores_list, mask_preds_list, gt_labels_list, gt_masks_list, img_metas
+ )
+
+ num_total_pos = sum((inds.numel() for inds in pos_inds_list))
+ num_total_neg = sum((inds.numel() for inds in neg_inds_list))
+ return (labels_list, label_weights_list, mask_targets_list, mask_weights_list, num_total_pos, num_total_neg)
+
+ def _get_target_single(self, cls_score, mask_pred, gt_labels, gt_masks, img_metas):
+ """Compute classification and mask targets for one image.
+
+ Args:
+ cls_score (Tensor): Mask score logits from a single decoder layer
+ for one image. Shape (num_queries, cls_out_channels).
+ mask_pred (Tensor): Mask logits for a single decoder layer for one
+ image. Shape (num_queries, h, w).
+ gt_labels (Tensor): Ground truth class indices for one image with
+ shape (num_gts, ).
+ gt_masks (Tensor): Ground truth mask for each image, each with
+ shape (num_gts, h, w).
+ img_metas (dict): Image informtation.
+
+ Returns:
+ tuple[Tensor]: A tuple containing the following for one image.
+
+ - labels (Tensor): Labels of each image. \
+ shape (num_queries, ).
+ - label_weights (Tensor): Label weights of each image. \
+ shape (num_queries, ).
+ - mask_targets (Tensor): Mask targets of each image. \
+ shape (num_queries, h, w).
+ - mask_weights (Tensor): Mask weights of each image. \
+ shape (num_queries, ).
+ - pos_inds (Tensor): Sampled positive indices for each \
+ image.
+ - neg_inds (Tensor): Sampled negative indices for each \
+ image.
+ """
+ # sample points
+ num_queries = cls_score.shape[0]
+ num_gts = gt_labels.shape[0]
+
+ point_coords = torch.rand((1, self.num_points, 2), device=cls_score.device)
+ # shape (num_queries, num_points)
+ mask_points_pred = point_sample(mask_pred.unsqueeze(1), point_coords.repeat(num_queries, 1, 1)).squeeze(1)
+ # shape (num_gts, num_points)
+ gt_points_masks = point_sample(gt_masks.unsqueeze(1).float(), point_coords.repeat(num_gts, 1, 1)).squeeze(1)
+
+ # assign and sample
+ assign_result = self.assigner.assign(cls_score, mask_points_pred, gt_labels, gt_points_masks, img_metas)
+ sampling_result = self.sampler.sample(assign_result, mask_pred, gt_masks)
+ pos_inds = sampling_result.pos_inds
+ neg_inds = sampling_result.neg_inds
+
+ # label target
+ labels = gt_labels.new_full((self.num_queries,), self.num_classes, dtype=torch.long)
+ labels[pos_inds] = gt_labels[sampling_result.pos_assigned_gt_inds]
+ label_weights = gt_labels.new_ones((self.num_queries,))
+
+ # mask target
+ mask_targets = gt_masks[sampling_result.pos_assigned_gt_inds]
+ mask_weights = mask_pred.new_zeros((self.num_queries,))
+ mask_weights[pos_inds] = 1.0
+
+ return (labels, label_weights, mask_targets, mask_weights, pos_inds, neg_inds)
+
+ def loss_single(self, cls_scores, mask_preds, gt_labels_list, gt_masks_list, img_metas):
+ """Loss function for outputs from a single decoder layer.
+
+ Args:
+ cls_scores (Tensor): Mask score logits from a single decoder layer
+ for all images. Shape (batch_size, num_queries,
+ cls_out_channels). Note `cls_out_channels` should includes
+ background.
+ mask_preds (Tensor): Mask logits for a pixel decoder for all
+ images. Shape (batch_size, num_queries, h, w).
+ gt_labels_list (list[Tensor]): Ground truth class indices for each
+ image, each with shape (num_gts, ).
+ gt_masks_list (list[Tensor]): Ground truth mask for each image,
+ each with shape (num_gts, h, w).
+ img_metas (list[dict]): List of image meta information.
+
+ Returns:
+ tuple[Tensor]: Loss components for outputs from a single \
+ decoder layer.
+ """
+ num_imgs = cls_scores.size(0)
+ cls_scores_list = [cls_scores[i] for i in range(num_imgs)]
+ mask_preds_list = [mask_preds[i] for i in range(num_imgs)]
+ (
+ labels_list,
+ label_weights_list,
+ mask_targets_list,
+ mask_weights_list,
+ num_total_pos,
+ num_total_neg,
+ ) = self.get_targets(cls_scores_list, mask_preds_list, gt_labels_list, gt_masks_list, img_metas)
+ # shape (batch_size, num_queries)
+ labels = torch.stack(labels_list, dim=0)
+ # shape (batch_size, num_queries)
+ label_weights = torch.stack(label_weights_list, dim=0)
+ # shape (num_total_gts, h, w)
+ mask_targets = torch.cat(mask_targets_list, dim=0)
+ # shape (batch_size, num_queries)
+ mask_weights = torch.stack(mask_weights_list, dim=0)
+
+ # classfication loss
+ # shape (batch_size * num_queries, )
+ cls_scores = cls_scores.flatten(0, 1)
+ labels = labels.flatten(0, 1)
+ label_weights = label_weights.flatten(0, 1)
+
+ class_weight = cls_scores.new_tensor(self.class_weight)
+ loss_cls = self.loss_cls(cls_scores, labels, label_weights, avg_factor=class_weight[labels].sum())
+
+ num_total_masks = reduce_mean(cls_scores.new_tensor([num_total_pos]))
+ num_total_masks = max(num_total_masks, 1)
+
+ # extract positive ones
+ # shape (batch_size, num_queries, h, w) -> (num_total_gts, h, w)
+ mask_preds = mask_preds[mask_weights > 0]
+
+ if mask_targets.shape[0] == 0:
+ # zero match
+ loss_dice = mask_preds.sum()
+ loss_mask = mask_preds.sum()
+ return loss_cls, loss_mask, loss_dice
+
+ with torch.no_grad():
+ points_coords = get_uncertain_point_coords_with_randomness(
+ mask_preds.unsqueeze(1), None, self.num_points, self.oversample_ratio, self.importance_sample_ratio
+ )
+ # shape (num_total_gts, h, w) -> (num_total_gts, num_points)
+ mask_point_targets = point_sample(mask_targets.unsqueeze(1).float(), points_coords).squeeze(1)
+ # shape (num_queries, h, w) -> (num_queries, num_points)
+ mask_point_preds = point_sample(mask_preds.unsqueeze(1), points_coords).squeeze(1)
+
+ # dice loss
+ loss_dice = self.loss_dice(mask_point_preds, mask_point_targets, avg_factor=num_total_masks)
+
+ # mask loss
+ # shape (num_queries, num_points) -> (num_queries * num_points, )
+ mask_point_preds = mask_point_preds.reshape(-1, 1)
+ # shape (num_total_gts, num_points) -> (num_total_gts * num_points, )
+ mask_point_targets = mask_point_targets.reshape(-1)
+ loss_mask = self.loss_mask(mask_point_preds, mask_point_targets, avg_factor=num_total_masks * self.num_points)
+
+ return loss_cls, loss_mask, loss_dice
+
+ @force_fp32(apply_to=("all_cls_scores", "all_mask_preds"))
+ def loss(self, all_cls_scores, all_mask_preds, gt_labels_list, gt_masks_list, img_metas):
+ """Loss function.
+
+ Args:
+ all_cls_scores (Tensor): Classification scores for all decoder
+ layers with shape [num_decoder, batch_size, num_queries,
+ cls_out_channels].
+ all_mask_preds (Tensor): Mask scores for all decoder layers with
+ shape [num_decoder, batch_size, num_queries, h, w].
+ gt_labels_list (list[Tensor]): Ground truth class indices for each
+ image with shape (n, ). n is the sum of number of stuff type
+ and number of instance in a image.
+ gt_masks_list (list[Tensor]): Ground truth mask for each image with
+ shape (n, h, w).
+ img_metas (list[dict]): List of image meta information.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ num_dec_layers = len(all_cls_scores)
+ all_gt_labels_list = [gt_labels_list for _ in range(num_dec_layers)]
+ all_gt_masks_list = [gt_masks_list for _ in range(num_dec_layers)]
+ img_metas_list = [img_metas for _ in range(num_dec_layers)]
+ losses_cls, losses_mask, losses_dice = multi_apply(
+ self.loss_single, all_cls_scores, all_mask_preds, all_gt_labels_list, all_gt_masks_list, img_metas_list
+ )
+
+ loss_dict = dict()
+ # loss from the last decoder layer
+ loss_dict["loss_cls"] = losses_cls[-1]
+ loss_dict["loss_mask"] = losses_mask[-1]
+ loss_dict["loss_dice"] = losses_dice[-1]
+ # loss from other decoder layers
+ num_dec_layer = 0
+ for loss_cls_i, loss_mask_i, loss_dice_i in zip(losses_cls[:-1], losses_mask[:-1], losses_dice[:-1]):
+ loss_dict[f"d{num_dec_layer}.loss_cls"] = loss_cls_i
+ loss_dict[f"d{num_dec_layer}.loss_mask"] = loss_mask_i
+ loss_dict[f"d{num_dec_layer}.loss_dice"] = loss_dice_i
+ num_dec_layer += 1
+ return loss_dict
+
+ def forward_head(self, decoder_out, mask_feature, attn_mask_target_size):
+ """Forward for head part which is called after every decoder layer.
+
+ Args:
+ decoder_out (Tensor): in shape (num_queries, batch_size, c).
+ mask_feature (Tensor): in shape (batch_size, c, h, w).
+ attn_mask_target_size (tuple[int, int]): target attention
+ mask size.
+
+ Returns:
+ tuple: A tuple contain three elements.
+
+ - cls_pred (Tensor): Classification scores in shape \
+ (batch_size, num_queries, cls_out_channels). \
+ Note `cls_out_channels` should includes background.
+ - mask_pred (Tensor): Mask scores in shape \
+ (batch_size, num_queries,h, w).
+ - attn_mask (Tensor): Attention mask in shape \
+ (batch_size * num_heads, num_queries, h, w).
+ """
+ decoder_out = self.transformer_decoder.post_norm(decoder_out)
+ decoder_out = decoder_out.transpose(0, 1)
+ # shape (num_queries, batch_size, c)
+ cls_pred = self.cls_embed(decoder_out)
+ # shape (num_queries, batch_size, c)
+ mask_embed = self.mask_embed(decoder_out)
+ # shape (num_queries, batch_size, h, w)
+ mask_pred = torch.einsum("bqc,bchw->bqhw", mask_embed, mask_feature)
+ attn_mask = F.interpolate(mask_pred, attn_mask_target_size, mode="bilinear", align_corners=False)
+ # shape (num_queries, batch_size, h, w) ->
+ # (batch_size * num_head, num_queries, h, w)
+ attn_mask = attn_mask.flatten(2).unsqueeze(1).repeat((1, self.num_heads, 1, 1)).flatten(0, 1)
+ attn_mask = attn_mask.sigmoid() < 0.5
+ attn_mask = attn_mask.detach()
+
+ return cls_pred, mask_pred, attn_mask
+
+ def forward(self, feats, img_metas):
+ """Forward function.
+
+ Args:
+ feats (list[Tensor]): Multi scale Features from the
+ upstream network, each is a 4D-tensor.
+ img_metas (list[dict]): List of image information.
+
+ Returns:
+ tuple: A tuple contains two elements.
+
+ - cls_pred_list (list[Tensor)]: Classification logits \
+ for each decoder layer. Each is a 3D-tensor with shape \
+ (batch_size, num_queries, cls_out_channels). \
+ Note `cls_out_channels` should includes background.
+ - mask_pred_list (list[Tensor]): Mask logits for each \
+ decoder layer. Each with shape (batch_size, num_queries, \
+ h, w).
+ """
+ batch_size = len(img_metas)
+ mask_features, multi_scale_memorys = self.pixel_decoder(feats)
+ # multi_scale_memorys (from low resolution to high resolution)
+ decoder_inputs = []
+ decoder_positional_encodings = []
+ for i in range(self.num_transformer_feat_level):
+ decoder_input = self.decoder_input_projs[i](multi_scale_memorys[i])
+ # shape (batch_size, c, h, w) -> (h*w, batch_size, c)
+ decoder_input = decoder_input.flatten(2).permute(2, 0, 1)
+ level_embed = self.level_embed.weight[i].view(1, 1, -1)
+ decoder_input = decoder_input + level_embed
+ # shape (batch_size, c, h, w) -> (h*w, batch_size, c)
+ mask = decoder_input.new_zeros((batch_size,) + multi_scale_memorys[i].shape[-2:], dtype=torch.bool)
+ decoder_positional_encoding = self.decoder_positional_encoding(mask)
+ decoder_positional_encoding = decoder_positional_encoding.flatten(2).permute(2, 0, 1)
+ decoder_inputs.append(decoder_input)
+ decoder_positional_encodings.append(decoder_positional_encoding)
+ # shape (num_queries, c) -> (num_queries, batch_size, c)
+ query_feat = self.query_feat.weight.unsqueeze(1).repeat((1, batch_size, 1))
+ query_embed = self.query_embed.weight.unsqueeze(1).repeat((1, batch_size, 1))
+
+ cls_pred_list = []
+ mask_pred_list = []
+ cls_pred, mask_pred, attn_mask = self.forward_head(query_feat, mask_features, multi_scale_memorys[0].shape[-2:])
+ cls_pred_list.append(cls_pred)
+ mask_pred_list.append(mask_pred)
+
+ for i in range(self.num_transformer_decoder_layers):
+ level_idx = i % self.num_transformer_feat_level
+ # if a mask is all True(all background), then set it all False.
+ attn_mask[torch.where(attn_mask.sum(-1) == attn_mask.shape[-1])] = False
+
+ # cross_attn + self_attn
+ layer = self.transformer_decoder.layers[i]
+ attn_masks = [attn_mask, None]
+ query_feat = layer(
+ query=query_feat,
+ key=decoder_inputs[level_idx],
+ value=decoder_inputs[level_idx],
+ query_pos=query_embed,
+ key_pos=decoder_positional_encodings[level_idx],
+ attn_masks=attn_masks,
+ query_key_padding_mask=None,
+ # here we do not apply masking on padded region
+ key_padding_mask=None,
+ )
+ cls_pred, mask_pred, attn_mask = self.forward_head(
+ query_feat, mask_features, multi_scale_memorys[(i + 1) % self.num_transformer_feat_level].shape[-2:]
+ )
+
+ cls_pred_list.append(cls_pred)
+ mask_pred_list.append(mask_pred)
+
+ return cls_pred_list, mask_pred_list
+
+ def forward_train(self, x, img_metas, gt_semantic_seg, gt_labels, gt_masks):
+ """Forward function for training mode.
+
+ Args:
+ x (list[Tensor]): Multi-level features from the upstream network,
+ each is a 4D-tensor.
+ img_metas (list[Dict]): List of image information.
+ gt_semantic_seg (list[tensor]):Each element is the ground truth
+ of semantic segmentation with the shape (N, H, W).
+ train_cfg (dict): The training config, which not been used in
+ maskformer.
+ gt_labels (list[Tensor]): Each element is ground truth labels of
+ each box, shape (num_gts,).
+ gt_masks (list[BitmapMasks]): Each element is masks of instances
+ of a image, shape (num_gts, h, w).
+
+ Returns:
+ losses (dict[str, Tensor]): a dictionary of loss components
+ """
+
+ # forward
+ all_cls_scores, all_mask_preds = self(x, img_metas)
+
+ # loss
+ losses = self.loss(all_cls_scores, all_mask_preds, gt_labels, gt_masks, img_metas)
+
+ return losses
+
+ def forward_test(self, inputs, img_metas, test_cfg):
+ """Test segment without test-time aumengtation.
+
+ Only the output of last decoder layers was used.
+
+ Args:
+ inputs (list[Tensor]): Multi-level features from the
+ upstream network, each is a 4D-tensor.
+ img_metas (list[dict]): List of image information.
+ test_cfg (dict): Testing config.
+
+ Returns:
+ seg_mask (Tensor): Predicted semantic segmentation logits.
+ """
+ all_cls_scores, all_mask_preds = self(inputs, img_metas)
+ cls_score, mask_pred = all_cls_scores[-1], all_mask_preds[-1]
+ ori_h, ori_w, _ = img_metas[0]["ori_shape"]
+
+ # semantic inference
+ cls_score = F.softmax(cls_score, dim=-1)[..., :-1]
+ mask_pred = mask_pred.sigmoid()
+ seg_mask = torch.einsum("bqc,bqhw->bchw", cls_score, mask_pred)
+ return seg_mask
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/losses/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/losses/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..229a887817372f4991b32354180592cfb236d728
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/losses/__init__.py
@@ -0,0 +1,8 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .cross_entropy_loss import CrossEntropyLoss, binary_cross_entropy, cross_entropy, mask_cross_entropy
+from .dice_loss import DiceLoss
+from .match_costs import ClassificationCost, CrossEntropyLossCost, DiceCost
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/losses/cross_entropy_loss.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/losses/cross_entropy_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..0a1f9dd4aa52ebe94cc527db36b1c7fa2f53813e
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/losses/cross_entropy_loss.py
@@ -0,0 +1,279 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import warnings
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmseg.models.builder import LOSSES
+from mmseg.models.losses.utils import get_class_weight, weight_reduce_loss
+
+
+def cross_entropy(
+ pred,
+ label,
+ weight=None,
+ class_weight=None,
+ reduction="mean",
+ avg_factor=None,
+ ignore_index=-100,
+ avg_non_ignore=False,
+):
+ """cross_entropy. The wrapper function for :func:`F.cross_entropy`
+
+ Args:
+ pred (torch.Tensor): The prediction with shape (N, 1).
+ label (torch.Tensor): The learning label of the prediction.
+ weight (torch.Tensor, optional): Sample-wise loss weight.
+ Default: None.
+ class_weight (list[float], optional): The weight for each class.
+ Default: None.
+ reduction (str, optional): The method used to reduce the loss.
+ Options are 'none', 'mean' and 'sum'. Default: 'mean'.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Default: None.
+ ignore_index (int): Specifies a target value that is ignored and
+ does not contribute to the input gradients. When
+ ``avg_non_ignore `` is ``True``, and the ``reduction`` is
+ ``''mean''``, the loss is averaged over non-ignored targets.
+ Defaults: -100.
+ avg_non_ignore (bool): The flag decides to whether the loss is
+ only averaged over non-ignored targets. Default: False.
+ `New in version 0.23.0.`
+ """
+
+ # class_weight is a manual rescaling weight given to each class.
+ # If given, has to be a Tensor of size C element-wise losses
+ loss = F.cross_entropy(pred, label, weight=class_weight, reduction="none", ignore_index=ignore_index)
+
+ # apply weights and do the reduction
+ # average loss over non-ignored elements
+ # pytorch's official cross_entropy average loss over non-ignored elements
+ # refer to https://github.com/pytorch/pytorch/blob/56b43f4fec1f76953f15a627694d4bba34588969/torch/nn/functional.py#L2660 # noqa
+ if (avg_factor is None) and avg_non_ignore and reduction == "mean":
+ avg_factor = label.numel() - (label == ignore_index).sum().item()
+ if weight is not None:
+ weight = weight.float()
+ loss = weight_reduce_loss(loss, weight=weight, reduction=reduction, avg_factor=avg_factor)
+
+ return loss
+
+
+def _expand_onehot_labels(labels, label_weights, target_shape, ignore_index):
+ """Expand onehot labels to match the size of prediction."""
+ bin_labels = labels.new_zeros(target_shape)
+ valid_mask = (labels >= 0) & (labels != ignore_index)
+ inds = torch.nonzero(valid_mask, as_tuple=True)
+
+ if inds[0].numel() > 0:
+ if labels.dim() == 3:
+ bin_labels[inds[0], labels[valid_mask], inds[1], inds[2]] = 1
+ else:
+ bin_labels[inds[0], labels[valid_mask]] = 1
+
+ valid_mask = valid_mask.unsqueeze(1).expand(target_shape).float()
+
+ if label_weights is None:
+ bin_label_weights = valid_mask
+ else:
+ bin_label_weights = label_weights.unsqueeze(1).expand(target_shape)
+ bin_label_weights = bin_label_weights * valid_mask
+
+ return bin_labels, bin_label_weights, valid_mask
+
+
+def binary_cross_entropy(
+ pred,
+ label,
+ weight=None,
+ reduction="mean",
+ avg_factor=None,
+ class_weight=None,
+ ignore_index=-100,
+ avg_non_ignore=False,
+ **kwargs,
+):
+ """Calculate the binary CrossEntropy loss.
+
+ Args:
+ pred (torch.Tensor): The prediction with shape (N, 1).
+ label (torch.Tensor): The learning label of the prediction.
+ Note: In bce loss, label < 0 is invalid.
+ weight (torch.Tensor, optional): Sample-wise loss weight.
+ reduction (str, optional): The method used to reduce the loss.
+ Options are "none", "mean" and "sum".
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ class_weight (list[float], optional): The weight for each class.
+ ignore_index (int): The label index to be ignored. Default: -100.
+ avg_non_ignore (bool): The flag decides to whether the loss is
+ only averaged over non-ignored targets. Default: False.
+ `New in version 0.23.0.`
+
+ Returns:
+ torch.Tensor: The calculated loss
+ """
+ if pred.size(1) == 1:
+ # For binary class segmentation, the shape of pred is
+ # [N, 1, H, W] and that of label is [N, H, W].
+ assert label.max() <= 1, "For pred with shape [N, 1, H, W], its label must have at " "most 2 classes"
+ pred = pred.squeeze()
+ if pred.dim() != label.dim():
+ assert (pred.dim() == 2 and label.dim() == 1) or (pred.dim() == 4 and label.dim() == 3), (
+ "Only pred shape [N, C], label shape [N] or pred shape [N, C, " "H, W], label shape [N, H, W] are supported"
+ )
+ # `weight` returned from `_expand_onehot_labels`
+ # has been treated for valid (non-ignore) pixels
+ label, weight, valid_mask = _expand_onehot_labels(label, weight, pred.shape, ignore_index)
+ else:
+ # should mask out the ignored elements
+ valid_mask = ((label >= 0) & (label != ignore_index)).float()
+ if weight is not None:
+ weight = weight * valid_mask
+ else:
+ weight = valid_mask
+ # average loss over non-ignored and valid elements
+ if reduction == "mean" and avg_factor is None and avg_non_ignore:
+ avg_factor = valid_mask.sum().item()
+
+ loss = F.binary_cross_entropy_with_logits(pred, label.float(), pos_weight=class_weight, reduction="none")
+ # do the reduction for the weighted loss
+ loss = weight_reduce_loss(loss, weight, reduction=reduction, avg_factor=avg_factor)
+
+ return loss
+
+
+def mask_cross_entropy(
+ pred, target, label, reduction="mean", avg_factor=None, class_weight=None, ignore_index=None, **kwargs
+):
+ """Calculate the CrossEntropy loss for masks.
+
+ Args:
+ pred (torch.Tensor): The prediction with shape (N, C), C is the number
+ of classes.
+ target (torch.Tensor): The learning label of the prediction.
+ label (torch.Tensor): ``label`` indicates the class label of the mask'
+ corresponding object. This will be used to select the mask in the
+ of the class which the object belongs to when the mask prediction
+ if not class-agnostic.
+ reduction (str, optional): The method used to reduce the loss.
+ Options are "none", "mean" and "sum".
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ class_weight (list[float], optional): The weight for each class.
+ ignore_index (None): Placeholder, to be consistent with other loss.
+ Default: None.
+
+ Returns:
+ torch.Tensor: The calculated loss
+ """
+ assert ignore_index is None, "BCE loss does not support ignore_index"
+ assert reduction == "mean" and avg_factor is None
+ num_rois = pred.size()[0]
+ inds = torch.arange(0, num_rois, dtype=torch.long, device=pred.device)
+ pred_slice = pred[inds, label].squeeze(1)
+ return F.binary_cross_entropy_with_logits(pred_slice, target, weight=class_weight, reduction="mean")[None]
+
+
+@LOSSES.register_module(force=True)
+class CrossEntropyLoss(nn.Module):
+ """CrossEntropyLoss.
+
+ Args:
+ use_sigmoid (bool, optional): Whether the prediction uses sigmoid
+ of softmax. Defaults to False.
+ use_mask (bool, optional): Whether to use mask cross entropy loss.
+ Defaults to False.
+ reduction (str, optional): . Defaults to 'mean'.
+ Options are "none", "mean" and "sum".
+ class_weight (list[float] | str, optional): Weight of each class. If in
+ str format, read them from a file. Defaults to None.
+ loss_weight (float, optional): Weight of the loss. Defaults to 1.0.
+ loss_name (str, optional): Name of the loss item. If you want this loss
+ item to be included into the backward graph, `loss_` must be the
+ prefix of the name. Defaults to 'loss_ce'.
+ avg_non_ignore (bool): The flag decides to whether the loss is
+ only averaged over non-ignored targets. Default: False.
+ `New in version 0.23.0.`
+ """
+
+ def __init__(
+ self,
+ use_sigmoid=False,
+ use_mask=False,
+ reduction="mean",
+ class_weight=None,
+ loss_weight=1.0,
+ loss_name="loss_ce",
+ avg_non_ignore=False,
+ ):
+ super(CrossEntropyLoss, self).__init__()
+ assert (use_sigmoid is False) or (use_mask is False)
+ self.use_sigmoid = use_sigmoid
+ self.use_mask = use_mask
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+ self.class_weight = get_class_weight(class_weight)
+ self.avg_non_ignore = avg_non_ignore
+ if not self.avg_non_ignore and self.reduction == "mean":
+ warnings.warn(
+ "Default ``avg_non_ignore`` is False, if you would like to "
+ "ignore the certain label and average loss over non-ignore "
+ "labels, which is the same with PyTorch official "
+ "cross_entropy, set ``avg_non_ignore=True``."
+ )
+
+ if self.use_sigmoid:
+ self.cls_criterion = binary_cross_entropy
+ elif self.use_mask:
+ self.cls_criterion = mask_cross_entropy
+ else:
+ self.cls_criterion = cross_entropy
+ self._loss_name = loss_name
+
+ def extra_repr(self):
+ """Extra repr."""
+ s = f"avg_non_ignore={self.avg_non_ignore}"
+ return s
+
+ def forward(
+ self, cls_score, label, weight=None, avg_factor=None, reduction_override=None, ignore_index=-100, **kwargs
+ ):
+ """Forward function."""
+ assert reduction_override in (None, "none", "mean", "sum")
+ reduction = reduction_override if reduction_override else self.reduction
+ if self.class_weight is not None:
+ class_weight = cls_score.new_tensor(self.class_weight)
+ else:
+ class_weight = None
+ # Note: for BCE loss, label < 0 is invalid.
+ loss_cls = self.loss_weight * self.cls_criterion(
+ cls_score,
+ label,
+ weight,
+ class_weight=class_weight,
+ reduction=reduction,
+ avg_factor=avg_factor,
+ avg_non_ignore=self.avg_non_ignore,
+ ignore_index=ignore_index,
+ **kwargs,
+ )
+ return loss_cls
+
+ @property
+ def loss_name(self):
+ """Loss Name.
+
+ This function must be implemented and will return the name of this
+ loss function. This name will be used to combine different loss items
+ by simple sum operation. In addition, if you want this loss item to be
+ included into the backward graph, `loss_` must be the prefix of the
+ name.
+
+ Returns:
+ str: The name of this loss item.
+ """
+ return self._loss_name
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/losses/dice_loss.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/losses/dice_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..1bc5ba893c502861032ed531283f225e183eb693
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/losses/dice_loss.py
@@ -0,0 +1,153 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import torch
+import torch.nn as nn
+from mmseg.models.builder import LOSSES
+from mmseg.models.losses.utils import weight_reduce_loss
+
+
+def dice_loss(pred, target, weight=None, eps=1e-3, reduction="mean", avg_factor=None):
+ """Calculate dice loss, which is proposed in
+ `V-Net: Fully Convolutional Neural Networks for Volumetric
+ Medical Image Segmentation `_.
+
+ Args:
+ pred (torch.Tensor): The prediction, has a shape (n, *)
+ target (torch.Tensor): The learning label of the prediction,
+ shape (n, *), same shape of pred.
+ weight (torch.Tensor, optional): The weight of loss for each
+ prediction, has a shape (n,). Defaults to None.
+ eps (float): Avoid dividing by zero. Default: 1e-3.
+ reduction (str, optional): The method used to reduce the loss into
+ a scalar. Defaults to 'mean'.
+ Options are "none", "mean" and "sum".
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ """
+
+ input = pred.flatten(1)
+ target = target.flatten(1).float()
+
+ a = torch.sum(input * target, 1)
+ b = torch.sum(input * input, 1) + eps
+ c = torch.sum(target * target, 1) + eps
+ d = (2 * a) / (b + c)
+ loss = 1 - d
+ if weight is not None:
+ assert weight.ndim == loss.ndim
+ assert len(weight) == len(pred)
+ loss = weight_reduce_loss(loss, weight, reduction, avg_factor)
+ return loss
+
+
+def naive_dice_loss(pred, target, weight=None, eps=1e-3, reduction="mean", avg_factor=None):
+ """Calculate naive dice loss, the coefficient in the denominator is the
+ first power instead of the second power.
+
+ Args:
+ pred (torch.Tensor): The prediction, has a shape (n, *)
+ target (torch.Tensor): The learning label of the prediction,
+ shape (n, *), same shape of pred.
+ weight (torch.Tensor, optional): The weight of loss for each
+ prediction, has a shape (n,). Defaults to None.
+ eps (float): Avoid dividing by zero. Default: 1e-3.
+ reduction (str, optional): The method used to reduce the loss into
+ a scalar. Defaults to 'mean'.
+ Options are "none", "mean" and "sum".
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ """
+ input = pred.flatten(1)
+ target = target.flatten(1).float()
+
+ a = torch.sum(input * target, 1)
+ b = torch.sum(input, 1)
+ c = torch.sum(target, 1)
+ d = (2 * a + eps) / (b + c + eps)
+ loss = 1 - d
+ if weight is not None:
+ assert weight.ndim == loss.ndim
+ assert len(weight) == len(pred)
+ loss = weight_reduce_loss(loss, weight, reduction, avg_factor)
+ return loss
+
+
+@LOSSES.register_module(force=True)
+class DiceLoss(nn.Module):
+ def __init__(self, use_sigmoid=True, activate=True, reduction="mean", naive_dice=False, loss_weight=1.0, eps=1e-3):
+ """Dice Loss, there are two forms of dice loss is supported:
+
+ - the one proposed in `V-Net: Fully Convolutional Neural
+ Networks for Volumetric Medical Image Segmentation
+ `_.
+ - the dice loss in which the power of the number in the
+ denominator is the first power instead of the second
+ power.
+
+ Args:
+ use_sigmoid (bool, optional): Whether to the prediction is
+ used for sigmoid or softmax. Defaults to True.
+ activate (bool): Whether to activate the predictions inside,
+ this will disable the inside sigmoid operation.
+ Defaults to True.
+ reduction (str, optional): The method used
+ to reduce the loss. Options are "none",
+ "mean" and "sum". Defaults to 'mean'.
+ naive_dice (bool, optional): If false, use the dice
+ loss defined in the V-Net paper, otherwise, use the
+ naive dice loss in which the power of the number in the
+ denominator is the first power instead of the second
+ power.Defaults to False.
+ loss_weight (float, optional): Weight of loss. Defaults to 1.0.
+ eps (float): Avoid dividing by zero. Defaults to 1e-3.
+ """
+
+ super(DiceLoss, self).__init__()
+ self.use_sigmoid = use_sigmoid
+ self.reduction = reduction
+ self.naive_dice = naive_dice
+ self.loss_weight = loss_weight
+ self.eps = eps
+ self.activate = activate
+
+ def forward(self, pred, target, weight=None, reduction_override=None, avg_factor=None):
+ """Forward function.
+
+ Args:
+ pred (torch.Tensor): The prediction, has a shape (n, *).
+ target (torch.Tensor): The label of the prediction,
+ shape (n, *), same shape of pred.
+ weight (torch.Tensor, optional): The weight of loss for each
+ prediction, has a shape (n,). Defaults to None.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Options are "none", "mean" and "sum".
+
+ Returns:
+ torch.Tensor: The calculated loss
+ """
+
+ assert reduction_override in (None, "none", "mean", "sum")
+ reduction = reduction_override if reduction_override else self.reduction
+
+ if self.activate:
+ if self.use_sigmoid:
+ pred = pred.sigmoid()
+ else:
+ raise NotImplementedError
+
+ if self.naive_dice:
+ loss = self.loss_weight * naive_dice_loss(
+ pred, target, weight, eps=self.eps, reduction=reduction, avg_factor=avg_factor
+ )
+ else:
+ loss = self.loss_weight * dice_loss(
+ pred, target, weight, eps=self.eps, reduction=reduction, avg_factor=avg_factor
+ )
+
+ return loss
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/losses/match_costs.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/losses/match_costs.py
new file mode 100644
index 0000000000000000000000000000000000000000..4917d2a939c01398dd49c0d90b06f4c37d283ce0
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/losses/match_costs.py
@@ -0,0 +1,153 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import torch
+import torch.nn.functional as F
+
+from ..builder import MATCH_COST
+
+
+@MATCH_COST.register_module()
+class ClassificationCost:
+ """ClsSoftmaxCost.Borrow from
+ mmdet.core.bbox.match_costs.match_cost.ClassificationCost.
+
+ Args:
+ weight (int | float, optional): loss_weight
+
+ Examples:
+ >>> import torch
+ >>> self = ClassificationCost()
+ >>> cls_pred = torch.rand(4, 3)
+ >>> gt_labels = torch.tensor([0, 1, 2])
+ >>> factor = torch.tensor([10, 8, 10, 8])
+ >>> self(cls_pred, gt_labels)
+ tensor([[-0.3430, -0.3525, -0.3045],
+ [-0.3077, -0.2931, -0.3992],
+ [-0.3664, -0.3455, -0.2881],
+ [-0.3343, -0.2701, -0.3956]])
+ """
+
+ def __init__(self, weight=1.0):
+ self.weight = weight
+
+ def __call__(self, cls_pred, gt_labels):
+ """
+ Args:
+ cls_pred (Tensor): Predicted classification logits, shape
+ [num_query, num_class].
+ gt_labels (Tensor): Label of `gt_bboxes`, shape (num_gt,).
+
+ Returns:
+ torch.Tensor: cls_cost value with weight
+ """
+ # Following the official DETR repo, contrary to the loss that
+ # NLL is used, we approximate it in 1 - cls_score[gt_label].
+ # The 1 is a constant that doesn't change the matching,
+ # so it can be omitted.
+ cls_score = cls_pred.softmax(-1)
+ cls_cost = -cls_score[:, gt_labels]
+ return cls_cost * self.weight
+
+
+@MATCH_COST.register_module()
+class DiceCost:
+ """Cost of mask assignments based on dice losses.
+
+ Args:
+ weight (int | float, optional): loss_weight. Defaults to 1.
+ pred_act (bool, optional): Whether to apply sigmoid to mask_pred.
+ Defaults to False.
+ eps (float, optional): default 1e-12.
+ """
+
+ def __init__(self, weight=1.0, pred_act=False, eps=1e-3):
+ self.weight = weight
+ self.pred_act = pred_act
+ self.eps = eps
+
+ def binary_mask_dice_loss(self, mask_preds, gt_masks):
+ """
+ Args:
+ mask_preds (Tensor): Mask prediction in shape (N1, H, W).
+ gt_masks (Tensor): Ground truth in shape (N2, H, W)
+ store 0 or 1, 0 for negative class and 1 for
+ positive class.
+
+ Returns:
+ Tensor: Dice cost matrix in shape (N1, N2).
+ """
+ mask_preds = mask_preds.reshape((mask_preds.shape[0], -1))
+ gt_masks = gt_masks.reshape((gt_masks.shape[0], -1)).float()
+ numerator = 2 * torch.einsum("nc,mc->nm", mask_preds, gt_masks)
+ denominator = mask_preds.sum(-1)[:, None] + gt_masks.sum(-1)[None, :]
+ loss = 1 - (numerator + self.eps) / (denominator + self.eps)
+ return loss
+
+ def __call__(self, mask_preds, gt_masks):
+ """
+ Args:
+ mask_preds (Tensor): Mask prediction logits in shape (N1, H, W).
+ gt_masks (Tensor): Ground truth in shape (N2, H, W).
+
+ Returns:
+ Tensor: Dice cost matrix in shape (N1, N2).
+ """
+ if self.pred_act:
+ mask_preds = mask_preds.sigmoid()
+ dice_cost = self.binary_mask_dice_loss(mask_preds, gt_masks)
+ return dice_cost * self.weight
+
+
+@MATCH_COST.register_module()
+class CrossEntropyLossCost:
+ """CrossEntropyLossCost.
+
+ Args:
+ weight (int | float, optional): loss weight. Defaults to 1.
+ use_sigmoid (bool, optional): Whether the prediction uses sigmoid
+ of softmax. Defaults to True.
+ """
+
+ def __init__(self, weight=1.0, use_sigmoid=True):
+ assert use_sigmoid, "use_sigmoid = False is not supported yet."
+ self.weight = weight
+ self.use_sigmoid = use_sigmoid
+
+ def _binary_cross_entropy(self, cls_pred, gt_labels):
+ """
+ Args:
+ cls_pred (Tensor): The prediction with shape (num_query, 1, *) or
+ (num_query, *).
+ gt_labels (Tensor): The learning label of prediction with
+ shape (num_gt, *).
+ Returns:
+ Tensor: Cross entropy cost matrix in shape (num_query, num_gt).
+ """
+ cls_pred = cls_pred.flatten(1).float()
+ gt_labels = gt_labels.flatten(1).float()
+ n = cls_pred.shape[1]
+ pos = F.binary_cross_entropy_with_logits(cls_pred, torch.ones_like(cls_pred), reduction="none")
+ neg = F.binary_cross_entropy_with_logits(cls_pred, torch.zeros_like(cls_pred), reduction="none")
+ cls_cost = torch.einsum("nc,mc->nm", pos, gt_labels) + torch.einsum("nc,mc->nm", neg, 1 - gt_labels)
+ cls_cost = cls_cost / n
+
+ return cls_cost
+
+ def __call__(self, cls_pred, gt_labels):
+ """
+ Args:
+ cls_pred (Tensor): Predicted classification logits.
+ gt_labels (Tensor): Labels.
+ Returns:
+ Tensor: Cross entropy cost matrix with weight in
+ shape (num_query, num_gt).
+ """
+ if self.use_sigmoid:
+ cls_cost = self._binary_cross_entropy(cls_pred, gt_labels)
+ else:
+ raise NotImplementedError
+
+ return cls_cost * self.weight
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/plugins/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/plugins/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..81a60db4de31238cb38e078683e5ca265839fe60
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/plugins/__init__.py
@@ -0,0 +1,6 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .msdeformattn_pixel_decoder import MSDeformAttnPixelDecoder
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/plugins/msdeformattn_pixel_decoder.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/plugins/msdeformattn_pixel_decoder.py
new file mode 100644
index 0000000000000000000000000000000000000000..db1947175917f73f3f24184cb09c78e092d46ef8
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/plugins/msdeformattn_pixel_decoder.py
@@ -0,0 +1,242 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import PLUGIN_LAYERS, Conv2d, ConvModule, caffe2_xavier_init, normal_init, xavier_init
+from mmcv.cnn.bricks.transformer import build_positional_encoding, build_transformer_layer_sequence
+from mmcv.runner import BaseModule, ModuleList
+
+from ...core.anchor import MlvlPointGenerator
+from ..utils.transformer import MultiScaleDeformableAttention
+
+
+@PLUGIN_LAYERS.register_module()
+class MSDeformAttnPixelDecoder(BaseModule):
+ """Pixel decoder with multi-scale deformable attention.
+
+ Args:
+ in_channels (list[int] | tuple[int]): Number of channels in the
+ input feature maps.
+ strides (list[int] | tuple[int]): Output strides of feature from
+ backbone.
+ feat_channels (int): Number of channels for feature.
+ out_channels (int): Number of channels for output.
+ num_outs (int): Number of output scales.
+ norm_cfg (:obj:`mmcv.ConfigDict` | dict): Config for normalization.
+ Defaults to dict(type='GN', num_groups=32).
+ act_cfg (:obj:`mmcv.ConfigDict` | dict): Config for activation.
+ Defaults to dict(type='ReLU').
+ encoder (:obj:`mmcv.ConfigDict` | dict): Config for transformer
+ encoder. Defaults to `DetrTransformerEncoder`.
+ positional_encoding (:obj:`mmcv.ConfigDict` | dict): Config for
+ transformer encoder position encoding. Defaults to
+ dict(type='SinePositionalEncoding', num_feats=128,
+ normalize=True).
+ init_cfg (:obj:`mmcv.ConfigDict` | dict): Initialization config dict.
+ """
+
+ def __init__(
+ self,
+ in_channels=[256, 512, 1024, 2048],
+ strides=[4, 8, 16, 32],
+ feat_channels=256,
+ out_channels=256,
+ num_outs=3,
+ norm_cfg=dict(type="GN", num_groups=32),
+ act_cfg=dict(type="ReLU"),
+ encoder=dict(
+ type="DetrTransformerEncoder",
+ num_layers=6,
+ transformerlayers=dict(
+ type="BaseTransformerLayer",
+ attn_cfgs=dict(
+ type="MultiScaleDeformableAttention",
+ embed_dims=256,
+ num_heads=8,
+ num_levels=3,
+ num_points=4,
+ im2col_step=64,
+ dropout=0.0,
+ batch_first=False,
+ norm_cfg=None,
+ init_cfg=None,
+ ),
+ feedforward_channels=1024,
+ ffn_dropout=0.0,
+ operation_order=("self_attn", "norm", "ffn", "norm"),
+ ),
+ init_cfg=None,
+ ),
+ positional_encoding=dict(type="SinePositionalEncoding", num_feats=128, normalize=True),
+ init_cfg=None,
+ ):
+ super().__init__(init_cfg=init_cfg)
+ self.strides = strides
+ self.num_input_levels = len(in_channels)
+ self.num_encoder_levels = encoder.transformerlayers.attn_cfgs.num_levels
+ assert self.num_encoder_levels >= 1, "num_levels in attn_cfgs must be at least one"
+ input_conv_list = []
+ # from top to down (low to high resolution)
+ for i in range(self.num_input_levels - 1, self.num_input_levels - self.num_encoder_levels - 1, -1):
+ input_conv = ConvModule(
+ in_channels[i], feat_channels, kernel_size=1, norm_cfg=norm_cfg, act_cfg=None, bias=True
+ )
+ input_conv_list.append(input_conv)
+ self.input_convs = ModuleList(input_conv_list)
+
+ self.encoder = build_transformer_layer_sequence(encoder)
+ self.postional_encoding = build_positional_encoding(positional_encoding)
+ # high resolution to low resolution
+ self.level_encoding = nn.Embedding(self.num_encoder_levels, feat_channels)
+
+ # fpn-like structure
+ self.lateral_convs = ModuleList()
+ self.output_convs = ModuleList()
+ self.use_bias = norm_cfg is None
+ # from top to down (low to high resolution)
+ # fpn for the rest features that didn't pass in encoder
+ for i in range(self.num_input_levels - self.num_encoder_levels - 1, -1, -1):
+ lateral_conv = ConvModule(
+ in_channels[i], feat_channels, kernel_size=1, bias=self.use_bias, norm_cfg=norm_cfg, act_cfg=None
+ )
+ output_conv = ConvModule(
+ feat_channels,
+ feat_channels,
+ kernel_size=3,
+ stride=1,
+ padding=1,
+ bias=self.use_bias,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg,
+ )
+ self.lateral_convs.append(lateral_conv)
+ self.output_convs.append(output_conv)
+
+ self.mask_feature = Conv2d(feat_channels, out_channels, kernel_size=1, stride=1, padding=0)
+
+ self.num_outs = num_outs
+ self.point_generator = MlvlPointGenerator(strides)
+
+ def init_weights(self):
+ """Initialize weights."""
+ for i in range(0, self.num_encoder_levels):
+ xavier_init(self.input_convs[i].conv, gain=1, bias=0, distribution="uniform")
+
+ for i in range(0, self.num_input_levels - self.num_encoder_levels):
+ caffe2_xavier_init(self.lateral_convs[i].conv, bias=0)
+ caffe2_xavier_init(self.output_convs[i].conv, bias=0)
+
+ caffe2_xavier_init(self.mask_feature, bias=0)
+
+ normal_init(self.level_encoding, mean=0, std=1)
+ for p in self.encoder.parameters():
+ if p.dim() > 1:
+ nn.init.xavier_normal_(p)
+
+ # init_weights defined in MultiScaleDeformableAttention
+ for layer in self.encoder.layers:
+ for attn in layer.attentions:
+ if isinstance(attn, MultiScaleDeformableAttention):
+ attn.init_weights()
+
+ def forward(self, feats):
+ """
+ Args:
+ feats (list[Tensor]): Feature maps of each level. Each has
+ shape of (batch_size, c, h, w).
+
+ Returns:
+ tuple: A tuple containing the following:
+
+ - mask_feature (Tensor): shape (batch_size, c, h, w).
+ - multi_scale_features (list[Tensor]): Multi scale \
+ features, each in shape (batch_size, c, h, w).
+ """
+ # generate padding mask for each level, for each image
+ batch_size = feats[0].shape[0]
+ encoder_input_list = []
+ padding_mask_list = []
+ level_positional_encoding_list = []
+ spatial_shapes = []
+ reference_points_list = []
+ for i in range(self.num_encoder_levels):
+ level_idx = self.num_input_levels - i - 1
+ feat = feats[level_idx]
+ feat_projected = self.input_convs[i](feat)
+ h, w = feat.shape[-2:]
+
+ # no padding
+ padding_mask_resized = feat.new_zeros((batch_size,) + feat.shape[-2:], dtype=torch.bool)
+ pos_embed = self.postional_encoding(padding_mask_resized)
+ level_embed = self.level_encoding.weight[i]
+ level_pos_embed = level_embed.view(1, -1, 1, 1) + pos_embed
+ # (h_i * w_i, 2)
+ reference_points = self.point_generator.single_level_grid_priors(
+ feat.shape[-2:], level_idx, device=feat.device
+ )
+ # normalize
+ factor = feat.new_tensor([[w, h]]) * self.strides[level_idx]
+ reference_points = reference_points / factor
+
+ # shape (batch_size, c, h_i, w_i) -> (h_i * w_i, batch_size, c)
+ feat_projected = feat_projected.flatten(2).permute(2, 0, 1)
+ level_pos_embed = level_pos_embed.flatten(2).permute(2, 0, 1)
+ padding_mask_resized = padding_mask_resized.flatten(1)
+
+ encoder_input_list.append(feat_projected)
+ padding_mask_list.append(padding_mask_resized)
+ level_positional_encoding_list.append(level_pos_embed)
+ spatial_shapes.append(feat.shape[-2:])
+ reference_points_list.append(reference_points)
+ # shape (batch_size, total_num_query),
+ # total_num_query=sum([., h_i * w_i,.])
+ padding_masks = torch.cat(padding_mask_list, dim=1)
+ # shape (total_num_query, batch_size, c)
+ encoder_inputs = torch.cat(encoder_input_list, dim=0)
+ level_positional_encodings = torch.cat(level_positional_encoding_list, dim=0)
+ device = encoder_inputs.device
+ # shape (num_encoder_levels, 2), from low
+ # resolution to high resolution
+ spatial_shapes = torch.as_tensor(spatial_shapes, dtype=torch.long, device=device)
+ # shape (0, h_0*w_0, h_0*w_0+h_1*w_1, ...)
+ level_start_index = torch.cat((spatial_shapes.new_zeros((1,)), spatial_shapes.prod(1).cumsum(0)[:-1]))
+ reference_points = torch.cat(reference_points_list, dim=0)
+ reference_points = reference_points[None, :, None].repeat(batch_size, 1, self.num_encoder_levels, 1)
+ valid_radios = reference_points.new_ones((batch_size, self.num_encoder_levels, 2))
+ # shape (num_total_query, batch_size, c)
+ memory = self.encoder(
+ query=encoder_inputs,
+ key=None,
+ value=None,
+ query_pos=level_positional_encodings,
+ key_pos=None,
+ attn_masks=None,
+ key_padding_mask=None,
+ query_key_padding_mask=padding_masks,
+ spatial_shapes=spatial_shapes,
+ reference_points=reference_points,
+ level_start_index=level_start_index,
+ valid_radios=valid_radios,
+ )
+ # (num_total_query, batch_size, c) -> (batch_size, c, num_total_query)
+ memory = memory.permute(1, 2, 0)
+
+ # from low resolution to high resolution
+ num_query_per_level = [e[0] * e[1] for e in spatial_shapes]
+ outs = torch.split(memory, num_query_per_level, dim=-1)
+ outs = [x.reshape(batch_size, -1, spatial_shapes[i][0], spatial_shapes[i][1]) for i, x in enumerate(outs)]
+
+ for i in range(self.num_input_levels - self.num_encoder_levels - 1, -1, -1):
+ x = feats[i]
+ cur_feat = self.lateral_convs[i](x)
+ y = cur_feat + F.interpolate(outs[-1], size=cur_feat.shape[-2:], mode="bilinear", align_corners=False)
+ y = self.output_convs[i](y)
+ outs.append(y)
+ multi_scale_features = outs[: self.num_outs]
+
+ mask_feature = self.mask_feature(outs[-1])
+ return mask_feature, multi_scale_features
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/segmentors/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/segmentors/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..adf0062691e4889612e118f28ced853cd0bc33db
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/segmentors/__init__.py
@@ -0,0 +1,6 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .encoder_decoder_mask2former import EncoderDecoderMask2Former
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/segmentors/encoder_decoder_mask2former.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/segmentors/encoder_decoder_mask2former.py
new file mode 100644
index 0000000000000000000000000000000000000000..cfe572c9d317303bff8d51b85217d144906ebfe7
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/segmentors/encoder_decoder_mask2former.py
@@ -0,0 +1,271 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmseg.core import add_prefix
+from mmseg.models import builder
+from mmseg.models.builder import SEGMENTORS
+from mmseg.models.segmentors.base import BaseSegmentor
+from mmseg.ops import resize
+
+
+@SEGMENTORS.register_module()
+class EncoderDecoderMask2Former(BaseSegmentor):
+ """Encoder Decoder segmentors.
+
+ EncoderDecoder typically consists of backbone, decode_head, auxiliary_head.
+ Note that auxiliary_head is only used for deep supervision during training,
+ which could be dumped during inference.
+ """
+
+ def __init__(
+ self,
+ backbone,
+ decode_head,
+ neck=None,
+ auxiliary_head=None,
+ train_cfg=None,
+ test_cfg=None,
+ pretrained=None,
+ init_cfg=None,
+ ):
+ super(EncoderDecoderMask2Former, self).__init__(init_cfg)
+ if pretrained is not None:
+ assert backbone.get("pretrained") is None, "both backbone and segmentor set pretrained weight"
+ backbone.pretrained = pretrained
+ self.backbone = builder.build_backbone(backbone)
+ if neck is not None:
+ self.neck = builder.build_neck(neck)
+ decode_head.update(train_cfg=train_cfg)
+ decode_head.update(test_cfg=test_cfg)
+ self._init_decode_head(decode_head)
+ self._init_auxiliary_head(auxiliary_head)
+
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+
+ assert self.with_decode_head
+
+ def _init_decode_head(self, decode_head):
+ """Initialize ``decode_head``"""
+ self.decode_head = builder.build_head(decode_head)
+ self.align_corners = self.decode_head.align_corners
+ self.num_classes = self.decode_head.num_classes
+
+ def _init_auxiliary_head(self, auxiliary_head):
+ """Initialize ``auxiliary_head``"""
+ if auxiliary_head is not None:
+ if isinstance(auxiliary_head, list):
+ self.auxiliary_head = nn.ModuleList()
+ for head_cfg in auxiliary_head:
+ self.auxiliary_head.append(builder.build_head(head_cfg))
+ else:
+ self.auxiliary_head = builder.build_head(auxiliary_head)
+
+ def extract_feat(self, img):
+ """Extract features from images."""
+ x = self.backbone(img)
+ if self.with_neck:
+ x = self.neck(x)
+ return x
+
+ def encode_decode(self, img, img_metas):
+ """Encode images with backbone and decode into a semantic segmentation
+ map of the same size as input."""
+ x = self.extract_feat(img)
+ out = self._decode_head_forward_test(x, img_metas)
+ out = resize(input=out, size=img.shape[2:], mode="bilinear", align_corners=self.align_corners)
+ return out
+
+ def _decode_head_forward_train(self, x, img_metas, gt_semantic_seg, **kwargs):
+ """Run forward function and calculate loss for decode head in
+ training."""
+ losses = dict()
+ loss_decode = self.decode_head.forward_train(x, img_metas, gt_semantic_seg, **kwargs)
+
+ losses.update(add_prefix(loss_decode, "decode"))
+ return losses
+
+ def _decode_head_forward_test(self, x, img_metas):
+ """Run forward function and calculate loss for decode head in
+ inference."""
+ seg_logits = self.decode_head.forward_test(x, img_metas, self.test_cfg)
+ return seg_logits
+
+ def _auxiliary_head_forward_train(self, x, img_metas, gt_semantic_seg):
+ """Run forward function and calculate loss for auxiliary head in
+ training."""
+ losses = dict()
+ if isinstance(self.auxiliary_head, nn.ModuleList):
+ for idx, aux_head in enumerate(self.auxiliary_head):
+ loss_aux = aux_head.forward_train(x, img_metas, gt_semantic_seg, self.train_cfg)
+ losses.update(add_prefix(loss_aux, f"aux_{idx}"))
+ else:
+ loss_aux = self.auxiliary_head.forward_train(x, img_metas, gt_semantic_seg, self.train_cfg)
+ losses.update(add_prefix(loss_aux, "aux"))
+
+ return losses
+
+ def forward_dummy(self, img):
+ """Dummy forward function."""
+ seg_logit = self.encode_decode(img, None)
+
+ return seg_logit
+
+ def forward_train(self, img, img_metas, gt_semantic_seg, **kwargs):
+ """Forward function for training.
+
+ Args:
+ img (Tensor): Input images.
+ img_metas (list[dict]): List of image info dict where each dict
+ has: 'img_shape', 'scale_factor', 'flip', and may also contain
+ 'filename', 'ori_shape', 'pad_shape', and 'img_norm_cfg'.
+ For details on the values of these keys see
+ `mmseg/datasets/pipelines/formatting.py:Collect`.
+ gt_semantic_seg (Tensor): Semantic segmentation masks
+ used if the architecture supports semantic segmentation task.
+
+ Returns:
+ dict[str, Tensor]: a dictionary of loss components
+ """
+
+ x = self.extract_feat(img)
+
+ losses = dict()
+
+ loss_decode = self._decode_head_forward_train(x, img_metas, gt_semantic_seg, **kwargs)
+ losses.update(loss_decode)
+
+ if self.with_auxiliary_head:
+ loss_aux = self._auxiliary_head_forward_train(x, img_metas, gt_semantic_seg)
+ losses.update(loss_aux)
+
+ return losses
+
+ def slide_inference(self, img, img_meta, rescale):
+ """Inference by sliding-window with overlap.
+
+ If h_crop > h_img or w_crop > w_img, the small patch will be used to
+ decode without padding.
+ """
+
+ h_stride, w_stride = self.test_cfg.stride
+ h_crop, w_crop = self.test_cfg.crop_size
+ batch_size, _, h_img, w_img = img.size()
+ num_classes = self.num_classes
+ h_grids = max(h_img - h_crop + h_stride - 1, 0) // h_stride + 1
+ w_grids = max(w_img - w_crop + w_stride - 1, 0) // w_stride + 1
+ preds = img.new_zeros((batch_size, num_classes, h_img, w_img))
+ count_mat = img.new_zeros((batch_size, 1, h_img, w_img))
+ for h_idx in range(h_grids):
+ for w_idx in range(w_grids):
+ y1 = h_idx * h_stride
+ x1 = w_idx * w_stride
+ y2 = min(y1 + h_crop, h_img)
+ x2 = min(x1 + w_crop, w_img)
+ y1 = max(y2 - h_crop, 0)
+ x1 = max(x2 - w_crop, 0)
+ crop_img = img[:, :, y1:y2, x1:x2]
+ crop_seg_logit = self.encode_decode(crop_img, img_meta)
+ preds += F.pad(crop_seg_logit, (int(x1), int(preds.shape[3] - x2), int(y1), int(preds.shape[2] - y2)))
+
+ count_mat[:, :, y1:y2, x1:x2] += 1
+ assert (count_mat == 0).sum() == 0
+ if torch.onnx.is_in_onnx_export():
+ # cast count_mat to constant while exporting to ONNX
+ count_mat = torch.from_numpy(count_mat.cpu().detach().numpy()).to(device=img.device)
+ preds = preds / count_mat
+ if rescale:
+ preds = resize(
+ preds,
+ size=img_meta[0]["ori_shape"][:2],
+ mode="bilinear",
+ align_corners=self.align_corners,
+ warning=False,
+ )
+ return preds
+
+ def whole_inference(self, img, img_meta, rescale):
+ """Inference with full image."""
+
+ seg_logit = self.encode_decode(img, img_meta)
+ if rescale:
+ # support dynamic shape for onnx
+ if torch.onnx.is_in_onnx_export():
+ size = img.shape[2:]
+ else:
+ size = img_meta[0]["ori_shape"][:2]
+ seg_logit = resize(seg_logit, size=size, mode="bilinear", align_corners=self.align_corners, warning=False)
+
+ return seg_logit
+
+ def inference(self, img, img_meta, rescale):
+ """Inference with slide/whole style.
+
+ Args:
+ img (Tensor): The input image of shape (N, 3, H, W).
+ img_meta (dict): Image info dict where each dict has: 'img_shape',
+ 'scale_factor', 'flip', and may also contain
+ 'filename', 'ori_shape', 'pad_shape', and 'img_norm_cfg'.
+ For details on the values of these keys see
+ `mmseg/datasets/pipelines/formatting.py:Collect`.
+ rescale (bool): Whether rescale back to original shape.
+
+ Returns:
+ Tensor: The output segmentation map.
+ """
+
+ assert self.test_cfg.mode in ["slide", "whole"]
+ ori_shape = img_meta[0]["ori_shape"]
+ assert all(_["ori_shape"] == ori_shape for _ in img_meta)
+ if self.test_cfg.mode == "slide":
+ seg_logit = self.slide_inference(img, img_meta, rescale)
+ else:
+ seg_logit = self.whole_inference(img, img_meta, rescale)
+ output = F.softmax(seg_logit, dim=1)
+ flip = img_meta[0]["flip"]
+ if flip:
+ flip_direction = img_meta[0]["flip_direction"]
+ assert flip_direction in ["horizontal", "vertical"]
+ if flip_direction == "horizontal":
+ output = output.flip(dims=(3,))
+ elif flip_direction == "vertical":
+ output = output.flip(dims=(2,))
+
+ return output
+
+ def simple_test(self, img, img_meta, rescale=True):
+ """Simple test with single image."""
+ seg_logit = self.inference(img, img_meta, rescale)
+ seg_pred = seg_logit.argmax(dim=1)
+ if torch.onnx.is_in_onnx_export():
+ # our inference backend only support 4D output
+ seg_pred = seg_pred.unsqueeze(0)
+ return seg_pred
+ seg_pred = seg_pred.cpu().numpy()
+ # unravel batch dim
+ seg_pred = list(seg_pred)
+ return seg_pred
+
+ def aug_test(self, imgs, img_metas, rescale=True):
+ """Test with augmentations.
+
+ Only rescale=True is supported.
+ """
+ # aug_test rescale all imgs back to ori_shape for now
+ assert rescale
+ # to save memory, we get augmented seg logit inplace
+ seg_logit = self.inference(imgs[0], img_metas[0], rescale)
+ for i in range(1, len(imgs)):
+ cur_seg_logit = self.inference(imgs[i], img_metas[i], rescale)
+ seg_logit += cur_seg_logit
+ seg_logit /= len(imgs)
+ seg_pred = seg_logit.argmax(dim=1)
+ seg_pred = seg_pred.cpu().numpy()
+ # unravel batch dim
+ seg_pred = list(seg_pred)
+ return seg_pred
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/utils/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/utils/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..e7fdc1668b1015c8feea8fa1a4691bc0ebdbd936
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/utils/__init__.py
@@ -0,0 +1,9 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .assigner import MaskHungarianAssigner
+from .point_sample import get_uncertain_point_coords_with_randomness
+from .positional_encoding import LearnedPositionalEncoding, SinePositionalEncoding
+from .transformer import DetrTransformerDecoder, DetrTransformerDecoderLayer, DynamicConv, Transformer
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/utils/assigner.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/utils/assigner.py
new file mode 100644
index 0000000000000000000000000000000000000000..3cb08fc1bb2e36336989b45a1d3850f260c05963
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/utils/assigner.py
@@ -0,0 +1,157 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from abc import ABCMeta, abstractmethod
+
+import torch
+
+from ..builder import MASK_ASSIGNERS, build_match_cost
+
+try:
+ from scipy.optimize import linear_sum_assignment
+except ImportError:
+ linear_sum_assignment = None
+
+
+class AssignResult(metaclass=ABCMeta):
+ """Collection of assign results."""
+
+ def __init__(self, num_gts, gt_inds, labels):
+ self.num_gts = num_gts
+ self.gt_inds = gt_inds
+ self.labels = labels
+
+ @property
+ def info(self):
+ info = {
+ "num_gts": self.num_gts,
+ "gt_inds": self.gt_inds,
+ "labels": self.labels,
+ }
+ return info
+
+
+class BaseAssigner(metaclass=ABCMeta):
+ """Base assigner that assigns boxes to ground truth boxes."""
+
+ @abstractmethod
+ def assign(self, masks, gt_masks, gt_masks_ignore=None, gt_labels=None):
+ """Assign boxes to either a ground truth boxes or a negative boxes."""
+ pass
+
+
+@MASK_ASSIGNERS.register_module()
+class MaskHungarianAssigner(BaseAssigner):
+ """Computes one-to-one matching between predictions and ground truth for
+ mask.
+
+ This class computes an assignment between the targets and the predictions
+ based on the costs. The costs are weighted sum of three components:
+ classification cost, regression L1 cost and regression iou cost. The
+ targets don't include the no_object, so generally there are more
+ predictions than targets. After the one-to-one matching, the un-matched
+ are treated as backgrounds. Thus each query prediction will be assigned
+ with `0` or a positive integer indicating the ground truth index:
+
+ - 0: negative sample, no assigned gt
+ - positive integer: positive sample, index (1-based) of assigned gt
+
+ Args:
+ cls_cost (obj:`mmcv.ConfigDict`|dict): Classification cost config.
+ mask_cost (obj:`mmcv.ConfigDict`|dict): Mask cost config.
+ dice_cost (obj:`mmcv.ConfigDict`|dict): Dice cost config.
+ """
+
+ def __init__(
+ self,
+ cls_cost=dict(type="ClassificationCost", weight=1.0),
+ dice_cost=dict(type="DiceCost", weight=1.0),
+ mask_cost=dict(type="MaskFocalCost", weight=1.0),
+ ):
+ self.cls_cost = build_match_cost(cls_cost)
+ self.dice_cost = build_match_cost(dice_cost)
+ self.mask_cost = build_match_cost(mask_cost)
+
+ def assign(self, cls_pred, mask_pred, gt_labels, gt_masks, img_meta, gt_masks_ignore=None, eps=1e-7):
+ """Computes one-to-one matching based on the weighted costs.
+
+ This method assign each query prediction to a ground truth or
+ background. The `assigned_gt_inds` with -1 means don't care,
+ 0 means negative sample, and positive number is the index (1-based)
+ of assigned gt.
+ The assignment is done in the following steps, the order matters.
+
+ 1. assign every prediction to -1
+ 2. compute the weighted costs
+ 3. do Hungarian matching on CPU based on the costs
+ 4. assign all to 0 (background) first, then for each matched pair
+ between predictions and gts, treat this prediction as foreground
+ and assign the corresponding gt index (plus 1) to it.
+
+ Args:
+ mask_pred (Tensor): Predicted mask, shape [num_query, h, w]
+ cls_pred (Tensor): Predicted classification logits, shape
+ [num_query, num_class].
+ gt_masks (Tensor): Ground truth mask, shape [num_gt, h, w].
+ gt_labels (Tensor): Label of `gt_masks`, shape (num_gt,).
+ img_meta (dict): Meta information for current image.
+ gt_masks_ignore (Tensor, optional): Ground truth masks that are
+ labelled as `ignored`. Default None.
+ eps (int | float, optional): A value added to the denominator for
+ numerical stability. Default 1e-7.
+
+ Returns:
+ :obj:`AssignResult`: The assigned result.
+ """
+ assert gt_masks_ignore is None, "Only case when gt_masks_ignore is None is supported."
+ num_gts, num_queries = gt_labels.shape[0], cls_pred.shape[0]
+
+ # 1. assign -1 by default
+ assigned_gt_inds = cls_pred.new_full((num_queries,), -1, dtype=torch.long)
+ assigned_labels = cls_pred.new_full((num_queries,), -1, dtype=torch.long)
+ if num_gts == 0 or num_queries == 0:
+ # No ground truth or boxes, return empty assignment
+ if num_gts == 0:
+ # No ground truth, assign all to background
+ assigned_gt_inds[:] = 0
+ return AssignResult(num_gts, assigned_gt_inds, labels=assigned_labels)
+
+ # 2. compute the weighted costs
+ # classification and maskcost.
+ if self.cls_cost.weight != 0 and cls_pred is not None:
+ cls_cost = self.cls_cost(cls_pred, gt_labels)
+ else:
+ cls_cost = 0
+
+ if self.mask_cost.weight != 0:
+ # mask_pred shape = [nq, h, w]
+ # gt_mask shape = [ng, h, w]
+ # mask_cost shape = [nq, ng]
+ mask_cost = self.mask_cost(mask_pred, gt_masks)
+ else:
+ mask_cost = 0
+
+ if self.dice_cost.weight != 0:
+ dice_cost = self.dice_cost(mask_pred, gt_masks)
+ else:
+ dice_cost = 0
+ cost = cls_cost + mask_cost + dice_cost
+
+ # 3. do Hungarian matching on CPU using linear_sum_assignment
+ cost = cost.detach().cpu()
+ if linear_sum_assignment is None:
+ raise ImportError('Please run "pip install scipy" ' "to install scipy first.")
+
+ matched_row_inds, matched_col_inds = linear_sum_assignment(cost)
+ matched_row_inds = torch.from_numpy(matched_row_inds).to(cls_pred.device)
+ matched_col_inds = torch.from_numpy(matched_col_inds).to(cls_pred.device)
+
+ # 4. assign backgrounds and foregrounds
+ # assign all indices to backgrounds first
+ assigned_gt_inds[:] = 0
+ # assign foregrounds based on matching results
+ assigned_gt_inds[matched_row_inds] = matched_col_inds + 1
+ assigned_labels[matched_row_inds] = gt_labels[matched_col_inds]
+ return AssignResult(num_gts, assigned_gt_inds, labels=assigned_labels)
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/utils/point_sample.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/utils/point_sample.py
new file mode 100644
index 0000000000000000000000000000000000000000..9f1134082bafb51432618a9632592db070f87284
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/utils/point_sample.py
@@ -0,0 +1,86 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import torch
+from mmcv.ops import point_sample
+
+
+def get_uncertainty(mask_pred, labels):
+ """Estimate uncertainty based on pred logits.
+
+ We estimate uncertainty as L1 distance between 0.0 and the logits
+ prediction in 'mask_pred' for the foreground class in `classes`.
+
+ Args:
+ mask_pred (Tensor): mask predication logits, shape (num_rois,
+ num_classes, mask_height, mask_width).
+
+ labels (list[Tensor]): Either predicted or ground truth label for
+ each predicted mask, of length num_rois.
+
+ Returns:
+ scores (Tensor): Uncertainty scores with the most uncertain
+ locations having the highest uncertainty score,
+ shape (num_rois, 1, mask_height, mask_width)
+ """
+ if mask_pred.shape[1] == 1:
+ gt_class_logits = mask_pred.clone()
+ else:
+ inds = torch.arange(mask_pred.shape[0], device=mask_pred.device)
+ gt_class_logits = mask_pred[inds, labels].unsqueeze(1)
+ return -torch.abs(gt_class_logits)
+
+
+def get_uncertain_point_coords_with_randomness(
+ mask_pred, labels, num_points, oversample_ratio, importance_sample_ratio
+):
+ """Get ``num_points`` most uncertain points with random points during
+ train.
+
+ Sample points in [0, 1] x [0, 1] coordinate space based on their
+ uncertainty. The uncertainties are calculated for each point using
+ 'get_uncertainty()' function that takes point's logit prediction as
+ input.
+
+ Args:
+ mask_pred (Tensor): A tensor of shape (num_rois, num_classes,
+ mask_height, mask_width) for class-specific or class-agnostic
+ prediction.
+ labels (list): The ground truth class for each instance.
+ num_points (int): The number of points to sample.
+ oversample_ratio (int): Oversampling parameter.
+ importance_sample_ratio (float): Ratio of points that are sampled
+ via importnace sampling.
+
+ Returns:
+ point_coords (Tensor): A tensor of shape (num_rois, num_points, 2)
+ that contains the coordinates sampled points.
+ """
+ assert oversample_ratio >= 1
+ assert 0 <= importance_sample_ratio <= 1
+ batch_size = mask_pred.shape[0]
+ num_sampled = int(num_points * oversample_ratio)
+ point_coords = torch.rand(batch_size, num_sampled, 2, device=mask_pred.device)
+ point_logits = point_sample(mask_pred, point_coords)
+ # It is crucial to calculate uncertainty based on the sampled
+ # prediction value for the points. Calculating uncertainties of the
+ # coarse predictions first and sampling them for points leads to
+ # incorrect results. To illustrate this: assume uncertainty func(
+ # logits)=-abs(logits), a sampled point between two coarse
+ # predictions with -1 and 1 logits has 0 logits, and therefore 0
+ # uncertainty value. However, if we calculate uncertainties for the
+ # coarse predictions first, both will have -1 uncertainty,
+ # and sampled point will get -1 uncertainty.
+ point_uncertainties = get_uncertainty(point_logits, labels)
+ num_uncertain_points = int(importance_sample_ratio * num_points)
+ num_random_points = num_points - num_uncertain_points
+ idx = torch.topk(point_uncertainties[:, 0, :], k=num_uncertain_points, dim=1)[1]
+ shift = num_sampled * torch.arange(batch_size, dtype=torch.long, device=mask_pred.device)
+ idx += shift[:, None]
+ point_coords = point_coords.view(-1, 2)[idx.view(-1), :].view(batch_size, num_uncertain_points, 2)
+ if num_random_points > 0:
+ rand_roi_coords = torch.rand(batch_size, num_random_points, 2, device=mask_pred.device)
+ point_coords = torch.cat((point_coords, rand_roi_coords), dim=1)
+ return point_coords
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/utils/positional_encoding.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/utils/positional_encoding.py
new file mode 100644
index 0000000000000000000000000000000000000000..bf5d6fabe946d06fe97cc799da47bae93758b34e
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/utils/positional_encoding.py
@@ -0,0 +1,152 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import math
+
+import torch
+import torch.nn as nn
+from mmcv.cnn.bricks.transformer import POSITIONAL_ENCODING
+from mmcv.runner import BaseModule
+
+
+@POSITIONAL_ENCODING.register_module()
+class SinePositionalEncoding(BaseModule):
+ """Position encoding with sine and cosine functions.
+
+ See `End-to-End Object Detection with Transformers
+ `_ for details.
+
+ Args:
+ num_feats (int): The feature dimension for each position
+ along x-axis or y-axis. Note the final returned dimension
+ for each position is 2 times of this value.
+ temperature (int, optional): The temperature used for scaling
+ the position embedding. Defaults to 10000.
+ normalize (bool, optional): Whether to normalize the position
+ embedding. Defaults to False.
+ scale (float, optional): A scale factor that scales the position
+ embedding. The scale will be used only when `normalize` is True.
+ Defaults to 2*pi.
+ eps (float, optional): A value added to the denominator for
+ numerical stability. Defaults to 1e-6.
+ offset (float): offset add to embed when do the normalization.
+ Defaults to 0.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None
+ """
+
+ def __init__(
+ self, num_feats, temperature=10000, normalize=False, scale=2 * math.pi, eps=1e-6, offset=0.0, init_cfg=None
+ ):
+ super(SinePositionalEncoding, self).__init__(init_cfg)
+ if normalize:
+ assert isinstance(scale, (float, int)), (
+ "when normalize is set," "scale should be provided and in float or int type, " f"found {type(scale)}"
+ )
+ self.num_feats = num_feats
+ self.temperature = temperature
+ self.normalize = normalize
+ self.scale = scale
+ self.eps = eps
+ self.offset = offset
+
+ def forward(self, mask):
+ """Forward function for `SinePositionalEncoding`.
+
+ Args:
+ mask (Tensor): ByteTensor mask. Non-zero values representing
+ ignored positions, while zero values means valid positions
+ for this image. Shape [bs, h, w].
+
+ Returns:
+ pos (Tensor): Returned position embedding with shape
+ [bs, num_feats*2, h, w].
+ """
+ # For convenience of exporting to ONNX, it's required to convert
+ # `masks` from bool to int.
+ mask = mask.to(torch.int)
+ not_mask = 1 - mask # logical_not
+ y_embed = not_mask.cumsum(1, dtype=torch.float32)
+ x_embed = not_mask.cumsum(2, dtype=torch.float32)
+ if self.normalize:
+ y_embed = (y_embed + self.offset) / (y_embed[:, -1:, :] + self.eps) * self.scale
+ x_embed = (x_embed + self.offset) / (x_embed[:, :, -1:] + self.eps) * self.scale
+ dim_t = torch.arange(self.num_feats, dtype=torch.float32, device=mask.device)
+ dim_t = self.temperature ** (2 * (dim_t // 2) / self.num_feats)
+ pos_x = x_embed[:, :, :, None] / dim_t
+ pos_y = y_embed[:, :, :, None] / dim_t
+ # use `view` instead of `flatten` for dynamically exporting to ONNX
+ B, H, W = mask.size()
+ pos_x = torch.stack((pos_x[:, :, :, 0::2].sin(), pos_x[:, :, :, 1::2].cos()), dim=4).view(B, H, W, -1)
+ pos_y = torch.stack((pos_y[:, :, :, 0::2].sin(), pos_y[:, :, :, 1::2].cos()), dim=4).view(B, H, W, -1)
+ pos = torch.cat((pos_y, pos_x), dim=3).permute(0, 3, 1, 2)
+ return pos
+
+ def __repr__(self):
+ """str: a string that describes the module"""
+ repr_str = self.__class__.__name__
+ repr_str += f"(num_feats={self.num_feats}, "
+ repr_str += f"temperature={self.temperature}, "
+ repr_str += f"normalize={self.normalize}, "
+ repr_str += f"scale={self.scale}, "
+ repr_str += f"eps={self.eps})"
+ return repr_str
+
+
+@POSITIONAL_ENCODING.register_module()
+class LearnedPositionalEncoding(BaseModule):
+ """Position embedding with learnable embedding weights.
+
+ Args:
+ num_feats (int): The feature dimension for each position
+ along x-axis or y-axis. The final returned dimension for
+ each position is 2 times of this value.
+ row_num_embed (int, optional): The dictionary size of row embeddings.
+ Default 50.
+ col_num_embed (int, optional): The dictionary size of col embeddings.
+ Default 50.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ """
+
+ def __init__(self, num_feats, row_num_embed=50, col_num_embed=50, init_cfg=dict(type="Uniform", layer="Embedding")):
+ super(LearnedPositionalEncoding, self).__init__(init_cfg)
+ self.row_embed = nn.Embedding(row_num_embed, num_feats)
+ self.col_embed = nn.Embedding(col_num_embed, num_feats)
+ self.num_feats = num_feats
+ self.row_num_embed = row_num_embed
+ self.col_num_embed = col_num_embed
+
+ def forward(self, mask):
+ """Forward function for `LearnedPositionalEncoding`.
+
+ Args:
+ mask (Tensor): ByteTensor mask. Non-zero values representing
+ ignored positions, while zero values means valid positions
+ for this image. Shape [bs, h, w].
+
+ Returns:
+ pos (Tensor): Returned position embedding with shape
+ [bs, num_feats*2, h, w].
+ """
+ h, w = mask.shape[-2:]
+ x = torch.arange(w, device=mask.device)
+ y = torch.arange(h, device=mask.device)
+ x_embed = self.col_embed(x)
+ y_embed = self.row_embed(y)
+ pos = (
+ torch.cat((x_embed.unsqueeze(0).repeat(h, 1, 1), y_embed.unsqueeze(1).repeat(1, w, 1)), dim=-1)
+ .permute(2, 0, 1)
+ .unsqueeze(0)
+ .repeat(mask.shape[0], 1, 1, 1)
+ )
+ return pos
+
+ def __repr__(self):
+ """str: a string that describes the module"""
+ repr_str = self.__class__.__name__
+ repr_str += f"(num_feats={self.num_feats}, "
+ repr_str += f"row_num_embed={self.row_num_embed}, "
+ repr_str += f"col_num_embed={self.col_num_embed})"
+ return repr_str
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/utils/transformer.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/utils/transformer.py
new file mode 100644
index 0000000000000000000000000000000000000000..8befe6011a34d5ccecb82c8b17b61e19f732f96b
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/models/utils/transformer.py
@@ -0,0 +1,989 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import math
+import warnings
+from typing import Sequence
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+import torch.utils.checkpoint as cp
+from mmcv.cnn import Linear, build_activation_layer, build_norm_layer, xavier_init
+from mmcv.cnn.bricks.drop import build_dropout
+from mmcv.cnn.bricks.registry import FEEDFORWARD_NETWORK, TRANSFORMER_LAYER, TRANSFORMER_LAYER_SEQUENCE
+from mmcv.cnn.bricks.transformer import BaseTransformerLayer, TransformerLayerSequence, build_transformer_layer_sequence
+from mmcv.runner.base_module import BaseModule, Sequential
+from mmcv.utils import deprecated_api_warning, to_2tuple
+from torch.nn.init import normal_
+
+from ..builder import TRANSFORMER
+
+try:
+ from mmcv.ops.multi_scale_deform_attn import MultiScaleDeformableAttention
+
+except ImportError:
+ warnings.warn(
+ "`MultiScaleDeformableAttention` in MMCV has been moved to "
+ "`mmcv.ops.multi_scale_deform_attn`, please update your MMCV"
+ )
+ from mmcv.cnn.bricks.transformer import MultiScaleDeformableAttention
+
+
+class AdaptivePadding(nn.Module):
+ """Applies padding to input (if needed) so that input can get fully covered
+ by filter you specified. It support two modes "same" and "corner". The
+ "same" mode is same with "SAME" padding mode in TensorFlow, pad zero around
+ input. The "corner" mode would pad zero to bottom right.
+
+ Args:
+ kernel_size (int | tuple): Size of the kernel:
+ stride (int | tuple): Stride of the filter. Default: 1:
+ dilation (int | tuple): Spacing between kernel elements.
+ Default: 1
+ padding (str): Support "same" and "corner", "corner" mode
+ would pad zero to bottom right, and "same" mode would
+ pad zero around input. Default: "corner".
+ Example:
+ >>> kernel_size = 16
+ >>> stride = 16
+ >>> dilation = 1
+ >>> input = torch.rand(1, 1, 15, 17)
+ >>> adap_pad = AdaptivePadding(
+ >>> kernel_size=kernel_size,
+ >>> stride=stride,
+ >>> dilation=dilation,
+ >>> padding="corner")
+ >>> out = adap_pad(input)
+ >>> assert (out.shape[2], out.shape[3]) == (16, 32)
+ >>> input = torch.rand(1, 1, 16, 17)
+ >>> out = adap_pad(input)
+ >>> assert (out.shape[2], out.shape[3]) == (16, 32)
+ """
+
+ def __init__(self, kernel_size=1, stride=1, dilation=1, padding="corner"):
+
+ super(AdaptivePadding, self).__init__()
+
+ assert padding in ("same", "corner")
+
+ kernel_size = to_2tuple(kernel_size)
+ stride = to_2tuple(stride)
+ padding = to_2tuple(padding)
+ dilation = to_2tuple(dilation)
+
+ self.padding = padding
+ self.kernel_size = kernel_size
+ self.stride = stride
+ self.dilation = dilation
+
+ def get_pad_shape(self, input_shape):
+ input_h, input_w = input_shape
+ kernel_h, kernel_w = self.kernel_size
+ stride_h, stride_w = self.stride
+ output_h = math.ceil(input_h / stride_h)
+ output_w = math.ceil(input_w / stride_w)
+ pad_h = max((output_h - 1) * stride_h + (kernel_h - 1) * self.dilation[0] + 1 - input_h, 0)
+ pad_w = max((output_w - 1) * stride_w + (kernel_w - 1) * self.dilation[1] + 1 - input_w, 0)
+ return pad_h, pad_w
+
+ def forward(self, x):
+ pad_h, pad_w = self.get_pad_shape(x.size()[-2:])
+ if pad_h > 0 or pad_w > 0:
+ if self.padding == "corner":
+ x = F.pad(x, [0, pad_w, 0, pad_h])
+ elif self.padding == "same":
+ x = F.pad(x, [pad_w // 2, pad_w - pad_w // 2, pad_h // 2, pad_h - pad_h // 2])
+ return x
+
+
+class PatchMerging(BaseModule):
+ """Merge patch feature map.
+
+ This layer groups feature map by kernel_size, and applies norm and linear
+ layers to the grouped feature map. Our implementation uses `nn.Unfold` to
+ merge patch, which is about 25% faster than original implementation.
+ Instead, we need to modify pretrained models for compatibility.
+
+ Args:
+ in_channels (int): The num of input channels.
+ to gets fully covered by filter and stride you specified..
+ Default: True.
+ out_channels (int): The num of output channels.
+ kernel_size (int | tuple, optional): the kernel size in the unfold
+ layer. Defaults to 2.
+ stride (int | tuple, optional): the stride of the sliding blocks in the
+ unfold layer. Default: None. (Would be set as `kernel_size`)
+ padding (int | tuple | string ): The padding length of
+ embedding conv. When it is a string, it means the mode
+ of adaptive padding, support "same" and "corner" now.
+ Default: "corner".
+ dilation (int | tuple, optional): dilation parameter in the unfold
+ layer. Default: 1.
+ bias (bool, optional): Whether to add bias in linear layer or not.
+ Defaults: False.
+ norm_cfg (dict, optional): Config dict for normalization layer.
+ Default: dict(type='LN').
+ init_cfg (dict, optional): The extra config for initialization.
+ Default: None.
+ """
+
+ def __init__(
+ self,
+ in_channels,
+ out_channels,
+ kernel_size=2,
+ stride=None,
+ padding="corner",
+ dilation=1,
+ bias=False,
+ norm_cfg=dict(type="LN"),
+ init_cfg=None,
+ ):
+ super().__init__(init_cfg=init_cfg)
+ self.in_channels = in_channels
+ self.out_channels = out_channels
+ if stride:
+ stride = stride
+ else:
+ stride = kernel_size
+
+ kernel_size = to_2tuple(kernel_size)
+ stride = to_2tuple(stride)
+ dilation = to_2tuple(dilation)
+
+ if isinstance(padding, str):
+ self.adap_padding = AdaptivePadding(
+ kernel_size=kernel_size, stride=stride, dilation=dilation, padding=padding
+ )
+ # disable the padding of unfold
+ padding = 0
+ else:
+ self.adap_padding = None
+
+ padding = to_2tuple(padding)
+ self.sampler = nn.Unfold(kernel_size=kernel_size, dilation=dilation, padding=padding, stride=stride)
+
+ sample_dim = kernel_size[0] * kernel_size[1] * in_channels
+
+ if norm_cfg is not None:
+ self.norm = build_norm_layer(norm_cfg, sample_dim)[1]
+ else:
+ self.norm = None
+
+ self.reduction = nn.Linear(sample_dim, out_channels, bias=bias)
+
+ def forward(self, x, input_size):
+ """
+ Args:
+ x (Tensor): Has shape (B, H*W, C_in).
+ input_size (tuple[int]): The spatial shape of x, arrange as (H, W).
+ Default: None.
+
+ Returns:
+ tuple: Contains merged results and its spatial shape.
+
+ - x (Tensor): Has shape (B, Merged_H * Merged_W, C_out)
+ - out_size (tuple[int]): Spatial shape of x, arrange as
+ (Merged_H, Merged_W).
+ """
+ B, L, C = x.shape
+ assert isinstance(input_size, Sequence), f"Expect " f"input_size is " f"`Sequence` " f"but get {input_size}"
+
+ H, W = input_size
+ assert L == H * W, "input feature has wrong size"
+
+ x = x.view(B, H, W, C).permute([0, 3, 1, 2]) # B, C, H, W
+ # Use nn.Unfold to merge patch. About 25% faster than original method,
+ # but need to modify pretrained model for compatibility
+
+ if self.adap_padding:
+ x = self.adap_padding(x)
+ H, W = x.shape[-2:]
+
+ x = self.sampler(x)
+ # if kernel_size=2 and stride=2, x should has shape (B, 4*C, H/2*W/2)
+
+ out_h = (
+ H + 2 * self.sampler.padding[0] - self.sampler.dilation[0] * (self.sampler.kernel_size[0] - 1) - 1
+ ) // self.sampler.stride[0] + 1
+ out_w = (
+ W + 2 * self.sampler.padding[1] - self.sampler.dilation[1] * (self.sampler.kernel_size[1] - 1) - 1
+ ) // self.sampler.stride[1] + 1
+
+ output_size = (out_h, out_w)
+ x = x.transpose(1, 2) # B, H/2*W/2, 4*C
+ x = self.norm(x) if self.norm else x
+ x = self.reduction(x)
+ return x, output_size
+
+
+def inverse_sigmoid(x, eps=1e-5):
+ """Inverse function of sigmoid.
+
+ Args:
+ x (Tensor): The tensor to do the
+ inverse.
+ eps (float): EPS avoid numerical
+ overflow. Defaults 1e-5.
+ Returns:
+ Tensor: The x has passed the inverse
+ function of sigmoid, has same
+ shape with input.
+ """
+ x = x.clamp(min=0, max=1)
+ x1 = x.clamp(min=eps)
+ x2 = (1 - x).clamp(min=eps)
+ return torch.log(x1 / x2)
+
+
+@FEEDFORWARD_NETWORK.register_module(force=True)
+class FFN(BaseModule):
+ """Implements feed-forward networks (FFNs) with identity connection.
+ Args:
+ embed_dims (int): The feature dimension. Same as
+ `MultiheadAttention`. Defaults: 256.
+ feedforward_channels (int): The hidden dimension of FFNs.
+ Defaults: 1024.
+ num_fcs (int, optional): The number of fully-connected layers in
+ FFNs. Default: 2.
+ act_cfg (dict, optional): The activation config for FFNs.
+ Default: dict(type='ReLU')
+ ffn_drop (float, optional): Probability of an element to be
+ zeroed in FFN. Default 0.0.
+ add_identity (bool, optional): Whether to add the
+ identity connection. Default: `True`.
+ dropout_layer (obj:`ConfigDict`): The dropout_layer used
+ when adding the shortcut.
+ init_cfg (obj:`mmcv.ConfigDict`): The Config for initialization.
+ Default: None.
+ """
+
+ @deprecated_api_warning({"dropout": "ffn_drop", "add_residual": "add_identity"}, cls_name="FFN")
+ def __init__(
+ self,
+ embed_dims=256,
+ feedforward_channels=1024,
+ num_fcs=2,
+ act_cfg=dict(type="ReLU", inplace=True),
+ ffn_drop=0.0,
+ dropout_layer=None,
+ add_identity=True,
+ init_cfg=None,
+ with_cp=False,
+ **kwargs,
+ ):
+ super().__init__(init_cfg)
+ assert num_fcs >= 2, "num_fcs should be no less " f"than 2. got {num_fcs}."
+ self.embed_dims = embed_dims
+ self.feedforward_channels = feedforward_channels
+ self.num_fcs = num_fcs
+ self.act_cfg = act_cfg
+ self.activate = build_activation_layer(act_cfg)
+ self.with_cp = with_cp
+ layers = []
+ in_channels = embed_dims
+ for _ in range(num_fcs - 1):
+ layers.append(Sequential(Linear(in_channels, feedforward_channels), self.activate, nn.Dropout(ffn_drop)))
+ in_channels = feedforward_channels
+ layers.append(Linear(feedforward_channels, embed_dims))
+ layers.append(nn.Dropout(ffn_drop))
+ self.layers = Sequential(*layers)
+ self.dropout_layer = build_dropout(dropout_layer) if dropout_layer else torch.nn.Identity()
+ self.add_identity = add_identity
+
+ @deprecated_api_warning({"residual": "identity"}, cls_name="FFN")
+ def forward(self, x, identity=None):
+ """Forward function for `FFN`.
+ The function would add x to the output tensor if residue is None.
+ """
+
+ if self.with_cp and x.requires_grad:
+ out = cp.checkpoint(self.layers, x)
+ else:
+ out = self.layers(x)
+
+ if not self.add_identity:
+ return self.dropout_layer(out)
+ if identity is None:
+ identity = x
+ return identity + self.dropout_layer(out)
+
+
+@TRANSFORMER_LAYER.register_module()
+class DetrTransformerDecoderLayer(BaseTransformerLayer):
+ """Implements decoder layer in DETR transformer.
+
+ Args:
+ attn_cfgs (list[`mmcv.ConfigDict`] | list[dict] | dict )):
+ Configs for self_attention or cross_attention, the order
+ should be consistent with it in `operation_order`. If it is
+ a dict, it would be expand to the number of attention in
+ `operation_order`.
+ feedforward_channels (int): The hidden dimension for FFNs.
+ ffn_dropout (float): Probability of an element to be zeroed
+ in ffn. Default 0.0.
+ operation_order (tuple[str]): The execution order of operation
+ in transformer. Such as ('self_attn', 'norm', 'ffn', 'norm').
+ Default:None
+ act_cfg (dict): The activation config for FFNs. Default: `LN`
+ norm_cfg (dict): Config dict for normalization layer.
+ Default: `LN`.
+ ffn_num_fcs (int): The number of fully-connected layers in FFNs.
+ Default:2.
+ """
+
+ def __init__(
+ self,
+ attn_cfgs,
+ feedforward_channels,
+ ffn_dropout=0.0,
+ operation_order=None,
+ act_cfg=dict(type="ReLU", inplace=True),
+ norm_cfg=dict(type="LN"),
+ ffn_num_fcs=2,
+ **kwargs,
+ ):
+ super(DetrTransformerDecoderLayer, self).__init__(
+ attn_cfgs=attn_cfgs,
+ feedforward_channels=feedforward_channels,
+ ffn_dropout=ffn_dropout,
+ operation_order=operation_order,
+ act_cfg=act_cfg,
+ norm_cfg=norm_cfg,
+ ffn_num_fcs=ffn_num_fcs,
+ **kwargs,
+ )
+ assert len(operation_order) == 6
+ assert set(operation_order) == set(["self_attn", "norm", "cross_attn", "ffn"])
+
+
+@TRANSFORMER_LAYER_SEQUENCE.register_module()
+class DetrTransformerEncoder(TransformerLayerSequence):
+ """TransformerEncoder of DETR.
+
+ Args:
+ post_norm_cfg (dict): Config of last normalization layer. Default:
+ `LN`. Only used when `self.pre_norm` is `True`
+ """
+
+ def __init__(self, *args, post_norm_cfg=dict(type="LN"), **kwargs):
+ super(DetrTransformerEncoder, self).__init__(*args, **kwargs)
+ if post_norm_cfg is not None:
+ self.post_norm = build_norm_layer(post_norm_cfg, self.embed_dims)[1] if self.pre_norm else None
+ else:
+ assert not self.pre_norm, f"Use prenorm in " f"{self.__class__.__name__}," f"Please specify post_norm_cfg"
+ self.post_norm = None
+
+ def forward(self, *args, **kwargs):
+ """Forward function for `TransformerCoder`.
+
+ Returns:
+ Tensor: forwarded results with shape [num_query, bs, embed_dims].
+ """
+ x = super(DetrTransformerEncoder, self).forward(*args, **kwargs)
+ if self.post_norm is not None:
+ x = self.post_norm(x)
+ return x
+
+
+@TRANSFORMER_LAYER_SEQUENCE.register_module()
+class DetrTransformerDecoder(TransformerLayerSequence):
+ """Implements the decoder in DETR transformer.
+
+ Args:
+ return_intermediate (bool): Whether to return intermediate outputs.
+ post_norm_cfg (dict): Config of last normalization layer. Default:
+ `LN`.
+ """
+
+ def __init__(self, *args, post_norm_cfg=dict(type="LN"), return_intermediate=False, **kwargs):
+
+ super(DetrTransformerDecoder, self).__init__(*args, **kwargs)
+ self.return_intermediate = return_intermediate
+ if post_norm_cfg is not None:
+ self.post_norm = build_norm_layer(post_norm_cfg, self.embed_dims)[1]
+ else:
+ self.post_norm = None
+
+ def forward(self, query, *args, **kwargs):
+ """Forward function for `TransformerDecoder`.
+
+ Args:
+ query (Tensor): Input query with shape
+ `(num_query, bs, embed_dims)`.
+
+ Returns:
+ Tensor: Results with shape [1, num_query, bs, embed_dims] when
+ return_intermediate is `False`, otherwise it has shape
+ [num_layers, num_query, bs, embed_dims].
+ """
+ if not self.return_intermediate:
+ x = super().forward(query, *args, **kwargs)
+ if self.post_norm:
+ x = self.post_norm(x)[None]
+ return x
+
+ intermediate = []
+ for layer in self.layers:
+ query = layer(query, *args, **kwargs)
+ if self.return_intermediate:
+ if self.post_norm is not None:
+ intermediate.append(self.post_norm(query))
+ else:
+ intermediate.append(query)
+ return torch.stack(intermediate)
+
+
+@TRANSFORMER.register_module()
+class Transformer(BaseModule):
+ """Implements the DETR transformer.
+
+ Following the official DETR implementation, this module copy-paste
+ from torch.nn.Transformer with modifications:
+
+ * positional encodings are passed in MultiheadAttention
+ * extra LN at the end of encoder is removed
+ * decoder returns a stack of activations from all decoding layers
+
+ See `paper: End-to-End Object Detection with Transformers
+ `_ for details.
+
+ Args:
+ encoder (`mmcv.ConfigDict` | Dict): Config of
+ TransformerEncoder. Defaults to None.
+ decoder ((`mmcv.ConfigDict` | Dict)): Config of
+ TransformerDecoder. Defaults to None
+ init_cfg (obj:`mmcv.ConfigDict`): The Config for initialization.
+ Defaults to None.
+ """
+
+ def __init__(self, encoder=None, decoder=None, init_cfg=None):
+ super(Transformer, self).__init__(init_cfg=init_cfg)
+ self.encoder = build_transformer_layer_sequence(encoder)
+ self.decoder = build_transformer_layer_sequence(decoder)
+ self.embed_dims = self.encoder.embed_dims
+
+ def init_weights(self):
+ # follow the official DETR to init parameters
+ for m in self.modules():
+ if hasattr(m, "weight") and m.weight.dim() > 1:
+ xavier_init(m, distribution="uniform")
+ self._is_init = True
+
+ def forward(self, x, mask, query_embed, pos_embed):
+ """Forward function for `Transformer`.
+
+ Args:
+ x (Tensor): Input query with shape [bs, c, h, w] where
+ c = embed_dims.
+ mask (Tensor): The key_padding_mask used for encoder and decoder,
+ with shape [bs, h, w].
+ query_embed (Tensor): The query embedding for decoder, with shape
+ [num_query, c].
+ pos_embed (Tensor): The positional encoding for encoder and
+ decoder, with the same shape as `x`.
+
+ Returns:
+ tuple[Tensor]: results of decoder containing the following tensor.
+
+ - out_dec: Output from decoder. If return_intermediate_dec \
+ is True output has shape [num_dec_layers, bs,
+ num_query, embed_dims], else has shape [1, bs, \
+ num_query, embed_dims].
+ - memory: Output results from encoder, with shape \
+ [bs, embed_dims, h, w].
+ """
+ bs, c, h, w = x.shape
+ # use `view` instead of `flatten` for dynamically exporting to ONNX
+ x = x.view(bs, c, -1).permute(2, 0, 1) # [bs, c, h, w] -> [h*w, bs, c]
+ pos_embed = pos_embed.view(bs, c, -1).permute(2, 0, 1)
+ query_embed = query_embed.unsqueeze(1).repeat(1, bs, 1) # [num_query, dim] -> [num_query, bs, dim]
+ mask = mask.view(bs, -1) # [bs, h, w] -> [bs, h*w]
+ memory = self.encoder(query=x, key=None, value=None, query_pos=pos_embed, query_key_padding_mask=mask)
+ target = torch.zeros_like(query_embed)
+ # out_dec: [num_layers, num_query, bs, dim]
+ out_dec = self.decoder(
+ query=target, key=memory, value=memory, key_pos=pos_embed, query_pos=query_embed, key_padding_mask=mask
+ )
+ out_dec = out_dec.transpose(1, 2)
+ memory = memory.permute(1, 2, 0).reshape(bs, c, h, w)
+ return out_dec, memory
+
+
+@TRANSFORMER_LAYER_SEQUENCE.register_module()
+class DeformableDetrTransformerDecoder(TransformerLayerSequence):
+ """Implements the decoder in DETR transformer.
+
+ Args:
+ return_intermediate (bool): Whether to return intermediate outputs.
+ coder_norm_cfg (dict): Config of last normalization layer. Default:
+ `LN`.
+ """
+
+ def __init__(self, *args, return_intermediate=False, **kwargs):
+
+ super(DeformableDetrTransformerDecoder, self).__init__(*args, **kwargs)
+ self.return_intermediate = return_intermediate
+
+ def forward(self, query, *args, reference_points=None, valid_ratios=None, reg_branches=None, **kwargs):
+ """Forward function for `TransformerDecoder`.
+
+ Args:
+ query (Tensor): Input query with shape
+ `(num_query, bs, embed_dims)`.
+ reference_points (Tensor): The reference
+ points of offset. has shape
+ (bs, num_query, 4) when as_two_stage,
+ otherwise has shape ((bs, num_query, 2).
+ valid_ratios (Tensor): The radios of valid
+ points on the feature map, has shape
+ (bs, num_levels, 2)
+ reg_branch: (obj:`nn.ModuleList`): Used for
+ refining the regression results. Only would
+ be passed when with_box_refine is True,
+ otherwise would be passed a `None`.
+
+ Returns:
+ Tensor: Results with shape [1, num_query, bs, embed_dims] when
+ return_intermediate is `False`, otherwise it has shape
+ [num_layers, num_query, bs, embed_dims].
+ """
+ output = query
+ intermediate = []
+ intermediate_reference_points = []
+ for lid, layer in enumerate(self.layers):
+ if reference_points.shape[-1] == 4:
+ reference_points_input = (
+ reference_points[:, :, None] * torch.cat([valid_ratios, valid_ratios], -1)[:, None]
+ )
+ else:
+ assert reference_points.shape[-1] == 2
+ reference_points_input = reference_points[:, :, None] * valid_ratios[:, None]
+ output = layer(output, *args, reference_points=reference_points_input, **kwargs)
+ output = output.permute(1, 0, 2)
+
+ if reg_branches is not None:
+ tmp = reg_branches[lid](output)
+ if reference_points.shape[-1] == 4:
+ new_reference_points = tmp + inverse_sigmoid(reference_points)
+ new_reference_points = new_reference_points.sigmoid()
+ else:
+ assert reference_points.shape[-1] == 2
+ new_reference_points = tmp
+ new_reference_points[..., :2] = tmp[..., :2] + inverse_sigmoid(reference_points)
+ new_reference_points = new_reference_points.sigmoid()
+ reference_points = new_reference_points.detach()
+
+ output = output.permute(1, 0, 2)
+ if self.return_intermediate:
+ intermediate.append(output)
+ intermediate_reference_points.append(reference_points)
+
+ if self.return_intermediate:
+ return torch.stack(intermediate), torch.stack(intermediate_reference_points)
+
+ return output, reference_points
+
+
+@TRANSFORMER.register_module()
+class DeformableDetrTransformer(Transformer):
+ """Implements the DeformableDETR transformer.
+
+ Args:
+ as_two_stage (bool): Generate query from encoder features.
+ Default: False.
+ num_feature_levels (int): Number of feature maps from FPN:
+ Default: 4.
+ two_stage_num_proposals (int): Number of proposals when set
+ `as_two_stage` as True. Default: 300.
+ """
+
+ def __init__(self, as_two_stage=False, num_feature_levels=4, two_stage_num_proposals=300, **kwargs):
+ super(DeformableDetrTransformer, self).__init__(**kwargs)
+ self.as_two_stage = as_two_stage
+ self.num_feature_levels = num_feature_levels
+ self.two_stage_num_proposals = two_stage_num_proposals
+ self.embed_dims = self.encoder.embed_dims
+ self.init_layers()
+
+ def init_layers(self):
+ """Initialize layers of the DeformableDetrTransformer."""
+ self.level_embeds = nn.Parameter(torch.Tensor(self.num_feature_levels, self.embed_dims))
+
+ if self.as_two_stage:
+ self.enc_output = nn.Linear(self.embed_dims, self.embed_dims)
+ self.enc_output_norm = nn.LayerNorm(self.embed_dims)
+ self.pos_trans = nn.Linear(self.embed_dims * 2, self.embed_dims * 2)
+ self.pos_trans_norm = nn.LayerNorm(self.embed_dims * 2)
+ else:
+ self.reference_points = nn.Linear(self.embed_dims, 2)
+
+ def init_weights(self):
+ """Initialize the transformer weights."""
+ for p in self.parameters():
+ if p.dim() > 1:
+ nn.init.xavier_uniform_(p)
+ for m in self.modules():
+ if isinstance(m, MultiScaleDeformableAttention):
+ m.init_weights()
+ if not self.as_two_stage:
+ xavier_init(self.reference_points, distribution="uniform", bias=0.0)
+ normal_(self.level_embeds)
+
+ def gen_encoder_output_proposals(self, memory, memory_padding_mask, spatial_shapes):
+ """Generate proposals from encoded memory.
+
+ Args:
+ memory (Tensor) : The output of encoder,
+ has shape (bs, num_key, embed_dim). num_key is
+ equal the number of points on feature map from
+ all level.
+ memory_padding_mask (Tensor): Padding mask for memory.
+ has shape (bs, num_key).
+ spatial_shapes (Tensor): The shape of all feature maps.
+ has shape (num_level, 2).
+
+ Returns:
+ tuple: A tuple of feature map and bbox prediction.
+
+ - output_memory (Tensor): The input of decoder, \
+ has shape (bs, num_key, embed_dim). num_key is \
+ equal the number of points on feature map from \
+ all levels.
+ - output_proposals (Tensor): The normalized proposal \
+ after a inverse sigmoid, has shape \
+ (bs, num_keys, 4).
+ """
+
+ N, S, C = memory.shape
+ proposals = []
+ _cur = 0
+ for lvl, (H, W) in enumerate(spatial_shapes):
+ mask_flatten_ = memory_padding_mask[:, _cur : (_cur + H * W)].view(N, H, W, 1)
+ valid_H = torch.sum(~mask_flatten_[:, :, 0, 0], 1)
+ valid_W = torch.sum(~mask_flatten_[:, 0, :, 0], 1)
+
+ grid_y, grid_x = torch.meshgrid(
+ torch.linspace(0, H - 1, H, dtype=torch.float32, device=memory.device),
+ torch.linspace(0, W - 1, W, dtype=torch.float32, device=memory.device),
+ )
+ grid = torch.cat([grid_x.unsqueeze(-1), grid_y.unsqueeze(-1)], -1)
+
+ scale = torch.cat([valid_W.unsqueeze(-1), valid_H.unsqueeze(-1)], 1).view(N, 1, 1, 2)
+ grid = (grid.unsqueeze(0).expand(N, -1, -1, -1) + 0.5) / scale
+ wh = torch.ones_like(grid) * 0.05 * (2.0**lvl)
+ proposal = torch.cat((grid, wh), -1).view(N, -1, 4)
+ proposals.append(proposal)
+ _cur += H * W
+ output_proposals = torch.cat(proposals, 1)
+ output_proposals_valid = ((output_proposals > 0.01) & (output_proposals < 0.99)).all(-1, keepdim=True)
+ output_proposals = torch.log(output_proposals / (1 - output_proposals))
+ output_proposals = output_proposals.masked_fill(memory_padding_mask.unsqueeze(-1), float("inf"))
+ output_proposals = output_proposals.masked_fill(~output_proposals_valid, float("inf"))
+
+ output_memory = memory
+ output_memory = output_memory.masked_fill(memory_padding_mask.unsqueeze(-1), float(0))
+ output_memory = output_memory.masked_fill(~output_proposals_valid, float(0))
+ output_memory = self.enc_output_norm(self.enc_output(output_memory))
+ return output_memory, output_proposals
+
+ @staticmethod
+ def get_reference_points(spatial_shapes, valid_ratios, device):
+ """Get the reference points used in decoder.
+
+ Args:
+ spatial_shapes (Tensor): The shape of all
+ feature maps, has shape (num_level, 2).
+ valid_ratios (Tensor): The radios of valid
+ points on the feature map, has shape
+ (bs, num_levels, 2)
+ device (obj:`device`): The device where
+ reference_points should be.
+
+ Returns:
+ Tensor: reference points used in decoder, has \
+ shape (bs, num_keys, num_levels, 2).
+ """
+ reference_points_list = []
+ for lvl, (H, W) in enumerate(spatial_shapes):
+ ref_y, ref_x = torch.meshgrid(
+ torch.linspace(0.5, H - 0.5, H, dtype=torch.float32, device=device),
+ torch.linspace(0.5, W - 0.5, W, dtype=torch.float32, device=device),
+ )
+ ref_y = ref_y.reshape(-1)[None] / (valid_ratios[:, None, lvl, 1] * H)
+ ref_x = ref_x.reshape(-1)[None] / (valid_ratios[:, None, lvl, 0] * W)
+ ref = torch.stack((ref_x, ref_y), -1)
+ reference_points_list.append(ref)
+ reference_points = torch.cat(reference_points_list, 1)
+ reference_points = reference_points[:, :, None] * valid_ratios[:, None]
+ return reference_points
+
+ def get_valid_ratio(self, mask):
+ """Get the valid radios of feature maps of all level."""
+ _, H, W = mask.shape
+ valid_H = torch.sum(~mask[:, :, 0], 1)
+ valid_W = torch.sum(~mask[:, 0, :], 1)
+ valid_ratio_h = valid_H.float() / H
+ valid_ratio_w = valid_W.float() / W
+ valid_ratio = torch.stack([valid_ratio_w, valid_ratio_h], -1)
+ return valid_ratio
+
+ def get_proposal_pos_embed(self, proposals, num_pos_feats=128, temperature=10000):
+ """Get the position embedding of proposal."""
+ scale = 2 * math.pi
+ dim_t = torch.arange(num_pos_feats, dtype=torch.float32, device=proposals.device)
+ dim_t = temperature ** (2 * (dim_t // 2) / num_pos_feats)
+ # N, L, 4
+ proposals = proposals.sigmoid() * scale
+ # N, L, 4, 128
+ pos = proposals[:, :, :, None] / dim_t
+ # N, L, 4, 64, 2
+ pos = torch.stack((pos[:, :, :, 0::2].sin(), pos[:, :, :, 1::2].cos()), dim=4).flatten(2)
+ return pos
+
+ def forward(
+ self, mlvl_feats, mlvl_masks, query_embed, mlvl_pos_embeds, reg_branches=None, cls_branches=None, **kwargs
+ ):
+ """Forward function for `Transformer`.
+
+ Args:
+ mlvl_feats (list(Tensor)): Input queries from
+ different level. Each element has shape
+ [bs, embed_dims, h, w].
+ mlvl_masks (list(Tensor)): The key_padding_mask from
+ different level used for encoder and decoder,
+ each element has shape [bs, h, w].
+ query_embed (Tensor): The query embedding for decoder,
+ with shape [num_query, c].
+ mlvl_pos_embeds (list(Tensor)): The positional encoding
+ of feats from different level, has the shape
+ [bs, embed_dims, h, w].
+ reg_branches (obj:`nn.ModuleList`): Regression heads for
+ feature maps from each decoder layer. Only would
+ be passed when
+ `with_box_refine` is True. Default to None.
+ cls_branches (obj:`nn.ModuleList`): Classification heads
+ for feature maps from each decoder layer. Only would
+ be passed when `as_two_stage`
+ is True. Default to None.
+
+
+ Returns:
+ tuple[Tensor]: results of decoder containing the following tensor.
+
+ - inter_states: Outputs from decoder. If
+ return_intermediate_dec is True output has shape \
+ (num_dec_layers, bs, num_query, embed_dims), else has \
+ shape (1, bs, num_query, embed_dims).
+ - init_reference_out: The initial value of reference \
+ points, has shape (bs, num_queries, 4).
+ - inter_references_out: The internal value of reference \
+ points in decoder, has shape \
+ (num_dec_layers, bs,num_query, embed_dims)
+ - enc_outputs_class: The classification score of \
+ proposals generated from \
+ encoder's feature maps, has shape \
+ (batch, h*w, num_classes). \
+ Only would be returned when `as_two_stage` is True, \
+ otherwise None.
+ - enc_outputs_coord_unact: The regression results \
+ generated from encoder's feature maps., has shape \
+ (batch, h*w, 4). Only would \
+ be returned when `as_two_stage` is True, \
+ otherwise None.
+ """
+ assert self.as_two_stage or query_embed is not None
+
+ feat_flatten = []
+ mask_flatten = []
+ lvl_pos_embed_flatten = []
+ spatial_shapes = []
+ for lvl, (feat, mask, pos_embed) in enumerate(zip(mlvl_feats, mlvl_masks, mlvl_pos_embeds)):
+ bs, c, h, w = feat.shape
+ spatial_shape = (h, w)
+ spatial_shapes.append(spatial_shape)
+ feat = feat.flatten(2).transpose(1, 2)
+ mask = mask.flatten(1)
+ pos_embed = pos_embed.flatten(2).transpose(1, 2)
+ lvl_pos_embed = pos_embed + self.level_embeds[lvl].view(1, 1, -1)
+ lvl_pos_embed_flatten.append(lvl_pos_embed)
+ feat_flatten.append(feat)
+ mask_flatten.append(mask)
+ feat_flatten = torch.cat(feat_flatten, 1)
+ mask_flatten = torch.cat(mask_flatten, 1)
+ lvl_pos_embed_flatten = torch.cat(lvl_pos_embed_flatten, 1)
+ spatial_shapes = torch.as_tensor(spatial_shapes, dtype=torch.long, device=feat_flatten.device)
+ level_start_index = torch.cat((spatial_shapes.new_zeros((1,)), spatial_shapes.prod(1).cumsum(0)[:-1]))
+ valid_ratios = torch.stack([self.get_valid_ratio(m) for m in mlvl_masks], 1)
+
+ reference_points = self.get_reference_points(spatial_shapes, valid_ratios, device=feat.device)
+
+ feat_flatten = feat_flatten.permute(1, 0, 2) # (H*W, bs, embed_dims)
+ lvl_pos_embed_flatten = lvl_pos_embed_flatten.permute(1, 0, 2) # (H*W, bs, embed_dims)
+ memory = self.encoder(
+ query=feat_flatten,
+ key=None,
+ value=None,
+ query_pos=lvl_pos_embed_flatten,
+ query_key_padding_mask=mask_flatten,
+ spatial_shapes=spatial_shapes,
+ reference_points=reference_points,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios,
+ **kwargs,
+ )
+
+ memory = memory.permute(1, 0, 2)
+ bs, _, c = memory.shape
+ if self.as_two_stage:
+ output_memory, output_proposals = self.gen_encoder_output_proposals(memory, mask_flatten, spatial_shapes)
+ enc_outputs_class = cls_branches[self.decoder.num_layers](output_memory)
+ enc_outputs_coord_unact = reg_branches[self.decoder.num_layers](output_memory) + output_proposals
+
+ topk = self.two_stage_num_proposals
+ topk_proposals = torch.topk(enc_outputs_class[..., 0], topk, dim=1)[1]
+ topk_coords_unact = torch.gather(enc_outputs_coord_unact, 1, topk_proposals.unsqueeze(-1).repeat(1, 1, 4))
+ topk_coords_unact = topk_coords_unact.detach()
+ reference_points = topk_coords_unact.sigmoid()
+ init_reference_out = reference_points
+ pos_trans_out = self.pos_trans_norm(self.pos_trans(self.get_proposal_pos_embed(topk_coords_unact)))
+ query_pos, query = torch.split(pos_trans_out, c, dim=2)
+ else:
+ query_pos, query = torch.split(query_embed, c, dim=1)
+ query_pos = query_pos.unsqueeze(0).expand(bs, -1, -1)
+ query = query.unsqueeze(0).expand(bs, -1, -1)
+ reference_points = self.reference_points(query_pos).sigmoid()
+ init_reference_out = reference_points
+
+ # decoder
+ query = query.permute(1, 0, 2)
+ memory = memory.permute(1, 0, 2)
+ query_pos = query_pos.permute(1, 0, 2)
+ inter_states, inter_references = self.decoder(
+ query=query,
+ key=None,
+ value=memory,
+ query_pos=query_pos,
+ key_padding_mask=mask_flatten,
+ reference_points=reference_points,
+ spatial_shapes=spatial_shapes,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios,
+ reg_branches=reg_branches,
+ **kwargs,
+ )
+
+ inter_references_out = inter_references
+ if self.as_two_stage:
+ return inter_states, init_reference_out, inter_references_out, enc_outputs_class, enc_outputs_coord_unact
+ return inter_states, init_reference_out, inter_references_out, None, None
+
+
+@TRANSFORMER.register_module()
+class DynamicConv(BaseModule):
+ """Implements Dynamic Convolution.
+
+ This module generate parameters for each sample and
+ use bmm to implement 1*1 convolution. Code is modified
+ from the `official github repo `_ .
+
+ Args:
+ in_channels (int): The input feature channel.
+ Defaults to 256.
+ feat_channels (int): The inner feature channel.
+ Defaults to 64.
+ out_channels (int, optional): The output feature channel.
+ When not specified, it will be set to `in_channels`
+ by default
+ input_feat_shape (int): The shape of input feature.
+ Defaults to 7.
+ with_proj (bool): Project two-dimentional feature to
+ one-dimentional feature. Default to True.
+ act_cfg (dict): The activation config for DynamicConv.
+ norm_cfg (dict): Config dict for normalization layer. Default
+ layer normalization.
+ init_cfg (obj:`mmcv.ConfigDict`): The Config for initialization.
+ Default: None.
+ """
+
+ def __init__(
+ self,
+ in_channels=256,
+ feat_channels=64,
+ out_channels=None,
+ input_feat_shape=7,
+ with_proj=True,
+ act_cfg=dict(type="ReLU", inplace=True),
+ norm_cfg=dict(type="LN"),
+ init_cfg=None,
+ ):
+ super(DynamicConv, self).__init__(init_cfg)
+ self.in_channels = in_channels
+ self.feat_channels = feat_channels
+ self.out_channels_raw = out_channels
+ self.input_feat_shape = input_feat_shape
+ self.with_proj = with_proj
+ self.act_cfg = act_cfg
+ self.norm_cfg = norm_cfg
+ self.out_channels = out_channels if out_channels else in_channels
+
+ self.num_params_in = self.in_channels * self.feat_channels
+ self.num_params_out = self.out_channels * self.feat_channels
+ self.dynamic_layer = nn.Linear(self.in_channels, self.num_params_in + self.num_params_out)
+
+ self.norm_in = build_norm_layer(norm_cfg, self.feat_channels)[1]
+ self.norm_out = build_norm_layer(norm_cfg, self.out_channels)[1]
+
+ self.activation = build_activation_layer(act_cfg)
+
+ num_output = self.out_channels * input_feat_shape**2
+ if self.with_proj:
+ self.fc_layer = nn.Linear(num_output, self.out_channels)
+ self.fc_norm = build_norm_layer(norm_cfg, self.out_channels)[1]
+
+ def forward(self, param_feature, input_feature):
+ """Forward function for `DynamicConv`.
+
+ Args:
+ param_feature (Tensor): The feature can be used
+ to generate the parameter, has shape
+ (num_all_proposals, in_channels).
+ input_feature (Tensor): Feature that
+ interact with parameters, has shape
+ (num_all_proposals, in_channels, H, W).
+
+ Returns:
+ Tensor: The output feature has shape
+ (num_all_proposals, out_channels).
+ """
+ input_feature = input_feature.flatten(2).permute(2, 0, 1)
+
+ input_feature = input_feature.permute(1, 0, 2)
+ parameters = self.dynamic_layer(param_feature)
+
+ param_in = parameters[:, : self.num_params_in].view(-1, self.in_channels, self.feat_channels)
+ param_out = parameters[:, -self.num_params_out :].view(-1, self.feat_channels, self.out_channels)
+
+ # input_feature has shape (num_all_proposals, H*W, in_channels)
+ # param_in has shape (num_all_proposals, in_channels, feat_channels)
+ # feature has shape (num_all_proposals, H*W, feat_channels)
+ features = torch.bmm(input_feature, param_in)
+ features = self.norm_in(features)
+ features = self.activation(features)
+
+ # param_out has shape (batch_size, feat_channels, out_channels)
+ features = torch.bmm(features, param_out)
+ features = self.norm_out(features)
+ features = self.activation(features)
+
+ if self.with_proj:
+ features = features.flatten(1)
+ features = self.fc_layer(features)
+ features = self.fc_norm(features)
+ features = self.activation(features)
+
+ return features
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/ops/modules/__init__.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/ops/modules/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..49aa8fe612fd4c088e294707c5ee16bd1cb5b5e7
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/ops/modules/__init__.py
@@ -0,0 +1,10 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+# References:
+# https://github.com/fundamentalvision/Deformable-DETR/tree/main/models/ops/modules
+# https://github.com/chengdazhi/Deformable-Convolution-V2-PyTorch/tree/pytorch_1.0.0
+
+from .ms_deform_attn import MSDeformAttn
diff --git a/models/dsp/dinov2/dinov2/eval/segmentation_m2f/ops/modules/ms_deform_attn.py b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/ops/modules/ms_deform_attn.py
new file mode 100644
index 0000000000000000000000000000000000000000..d8b4fa23712e87d1a2682b57e71ee37fe8524cff
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/segmentation_m2f/ops/modules/ms_deform_attn.py
@@ -0,0 +1,185 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import math
+import warnings
+
+import torch
+import torch.nn.functional as F
+from torch import nn
+from torch.autograd import Function
+from torch.cuda.amp import custom_fwd
+from torch.nn.init import constant_, xavier_uniform_
+
+
+class MSDeformAttnFunction(Function):
+ @staticmethod
+ @custom_fwd(cast_inputs=torch.float32)
+ def forward(
+ ctx, value, value_spatial_shapes, value_level_start_index, sampling_locations, attention_weights, im2col_step
+ ):
+ output = ms_deform_attn_core_pytorch(
+ value,
+ value_spatial_shapes,
+ # value_level_start_index,
+ sampling_locations,
+ attention_weights,
+ )
+ return output
+
+
+def ms_deform_attn_core_pytorch(value, value_spatial_shapes, sampling_locations, attention_weights):
+ # for debug and test only,
+ # need to use cuda version instead
+ N_, S_, M_, D_ = value.shape
+ _, Lq_, M_, L_, P_, _ = sampling_locations.shape
+ value_list = value.split([H_ * W_ for H_, W_ in value_spatial_shapes], dim=1)
+ sampling_grids = 2 * sampling_locations - 1
+ sampling_value_list = []
+ for lid_, (H_, W_) in enumerate(value_spatial_shapes):
+ # N_, H_*W_, M_, D_ -> N_, H_*W_, M_*D_ -> N_, M_*D_, H_*W_ -> N_*M_, D_, H_, W_
+ value_l_ = value_list[lid_].flatten(2).transpose(1, 2).reshape(N_ * M_, D_, H_, W_)
+ # N_, Lq_, M_, P_, 2 -> N_, M_, Lq_, P_, 2 -> N_*M_, Lq_, P_, 2
+ sampling_grid_l_ = sampling_grids[:, :, :, lid_].transpose(1, 2).flatten(0, 1)
+ # N_*M_, D_, Lq_, P_
+ sampling_value_l_ = F.grid_sample(
+ value_l_, sampling_grid_l_, mode="bilinear", padding_mode="zeros", align_corners=False
+ )
+ sampling_value_list.append(sampling_value_l_)
+ # (N_, Lq_, M_, L_, P_) -> (N_, M_, Lq_, L_, P_) -> (N_, M_, 1, Lq_, L_*P_)
+ attention_weights = attention_weights.transpose(1, 2).reshape(N_ * M_, 1, Lq_, L_ * P_)
+ output = (torch.stack(sampling_value_list, dim=-2).flatten(-2) * attention_weights).sum(-1).view(N_, M_ * D_, Lq_)
+ return output.transpose(1, 2).contiguous()
+
+
+def _is_power_of_2(n):
+ if (not isinstance(n, int)) or (n < 0):
+ raise ValueError("invalid input for _is_power_of_2: {} (type: {})".format(n, type(n)))
+ return (n & (n - 1) == 0) and n != 0
+
+
+class MSDeformAttn(nn.Module):
+ def __init__(self, d_model=256, n_levels=4, n_heads=8, n_points=4, ratio=1.0):
+ """Multi-Scale Deformable Attention Module.
+
+ :param d_model hidden dimension
+ :param n_levels number of feature levels
+ :param n_heads number of attention heads
+ :param n_points number of sampling points per attention head per feature level
+ """
+ super().__init__()
+ if d_model % n_heads != 0:
+ raise ValueError("d_model must be divisible by n_heads, " "but got {} and {}".format(d_model, n_heads))
+ _d_per_head = d_model // n_heads
+ # you'd better set _d_per_head to a power of 2
+ # which is more efficient in our CUDA implementation
+ if not _is_power_of_2(_d_per_head):
+ warnings.warn(
+ "You'd better set d_model in MSDeformAttn to make "
+ "the dimension of each attention head a power of 2 "
+ "which is more efficient in our CUDA implementation."
+ )
+
+ self.im2col_step = 64
+
+ self.d_model = d_model
+ self.n_levels = n_levels
+ self.n_heads = n_heads
+ self.n_points = n_points
+ self.ratio = ratio
+ self.sampling_offsets = nn.Linear(d_model, n_heads * n_levels * n_points * 2)
+ self.attention_weights = nn.Linear(d_model, n_heads * n_levels * n_points)
+ self.value_proj = nn.Linear(d_model, int(d_model * ratio))
+ self.output_proj = nn.Linear(int(d_model * ratio), d_model)
+
+ self._reset_parameters()
+
+ def _reset_parameters(self):
+ constant_(self.sampling_offsets.weight.data, 0.0)
+ thetas = torch.arange(self.n_heads, dtype=torch.float32) * (2.0 * math.pi / self.n_heads)
+ grid_init = torch.stack([thetas.cos(), thetas.sin()], -1)
+ grid_init = (
+ (grid_init / grid_init.abs().max(-1, keepdim=True)[0])
+ .view(self.n_heads, 1, 1, 2)
+ .repeat(1, self.n_levels, self.n_points, 1)
+ )
+ for i in range(self.n_points):
+ grid_init[:, :, i, :] *= i + 1
+
+ with torch.no_grad():
+ self.sampling_offsets.bias = nn.Parameter(grid_init.view(-1))
+ constant_(self.attention_weights.weight.data, 0.0)
+ constant_(self.attention_weights.bias.data, 0.0)
+ xavier_uniform_(self.value_proj.weight.data)
+ constant_(self.value_proj.bias.data, 0.0)
+ xavier_uniform_(self.output_proj.weight.data)
+ constant_(self.output_proj.bias.data, 0.0)
+
+ def forward(
+ self,
+ query,
+ reference_points,
+ input_flatten,
+ input_spatial_shapes,
+ input_level_start_index,
+ input_padding_mask=None,
+ ):
+ """
+ :param query (N, Length_{query}, C)
+ :param reference_points (N, Length_{query}, n_levels, 2), range in [0, 1], top-left (0,0), bottom-right (1, 1), including padding area
+ or (N, Length_{query}, n_levels, 4), add additional (w, h) to form reference boxes
+ :param input_flatten (N, \\sum_{l=0}^{L-1} H_l \\cdot W_l, C)
+ :param input_spatial_shapes (n_levels, 2), [(H_0, W_0), (H_1, W_1), ..., (H_{L-1}, W_{L-1})]
+ :param input_level_start_index (n_levels, ), [0, H_0*W_0, H_0*W_0+H_1*W_1, H_0*W_0+H_1*W_1+H_2*W_2, ..., H_0*W_0+H_1*W_1+...+H_{L-1}*W_{L-1}]
+ :param input_padding_mask (N, \\sum_{l=0}^{L-1} H_l \\cdot W_l), True for padding elements, False for non-padding elements
+
+ :return output (N, Length_{query}, C)
+ """
+ # print(query.shape)
+ # print(reference_points.shape)
+ # print(input_flatten.shape)
+ # print(input_spatial_shapes.shape)
+ # print(input_level_start_index.shape)
+ # print(input_spatial_shapes)
+ # print(input_level_start_index)
+
+ N, Len_q, _ = query.shape
+ N, Len_in, _ = input_flatten.shape
+ assert (input_spatial_shapes[:, 0] * input_spatial_shapes[:, 1]).sum() == Len_in
+
+ value = self.value_proj(input_flatten)
+ if input_padding_mask is not None:
+ value = value.masked_fill(input_padding_mask[..., None], float(0))
+
+ value = value.view(N, Len_in, self.n_heads, int(self.ratio * self.d_model) // self.n_heads)
+ sampling_offsets = self.sampling_offsets(query).view(N, Len_q, self.n_heads, self.n_levels, self.n_points, 2)
+ attention_weights = self.attention_weights(query).view(N, Len_q, self.n_heads, self.n_levels * self.n_points)
+ attention_weights = F.softmax(attention_weights, -1).view(N, Len_q, self.n_heads, self.n_levels, self.n_points)
+
+ if reference_points.shape[-1] == 2:
+ offset_normalizer = torch.stack([input_spatial_shapes[..., 1], input_spatial_shapes[..., 0]], -1)
+ sampling_locations = (
+ reference_points[:, :, None, :, None, :]
+ + sampling_offsets / offset_normalizer[None, None, None, :, None, :]
+ )
+ elif reference_points.shape[-1] == 4:
+ sampling_locations = (
+ reference_points[:, :, None, :, None, :2]
+ + sampling_offsets / self.n_points * reference_points[:, :, None, :, None, 2:] * 0.5
+ )
+ else:
+ raise ValueError(
+ "Last dim of reference_points must be 2 or 4, but get {} instead.".format(reference_points.shape[-1])
+ )
+ output = MSDeformAttnFunction.apply(
+ value,
+ input_spatial_shapes,
+ input_level_start_index,
+ sampling_locations,
+ attention_weights,
+ self.im2col_step,
+ )
+ output = self.output_proj(output)
+ return output
diff --git a/models/dsp/dinov2/dinov2/eval/setup.py b/models/dsp/dinov2/dinov2/eval/setup.py
new file mode 100644
index 0000000000000000000000000000000000000000..959128c0673cc51036dbf17dcc4ee68a037988fb
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/setup.py
@@ -0,0 +1,75 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import argparse
+from typing import Any, List, Optional, Tuple
+
+import torch
+import torch.backends.cudnn as cudnn
+
+from dinov2.models import build_model_from_cfg
+from dinov2.utils.config import setup
+import dinov2.utils.utils as dinov2_utils
+
+
+def get_args_parser(
+ description: Optional[str] = None,
+ parents: Optional[List[argparse.ArgumentParser]] = None,
+ add_help: bool = True,
+):
+ parser = argparse.ArgumentParser(
+ description=description,
+ parents=parents or [],
+ add_help=add_help,
+ )
+ parser.add_argument(
+ "--config-file",
+ type=str,
+ help="Model configuration file",
+ )
+ parser.add_argument(
+ "--pretrained-weights",
+ type=str,
+ help="Pretrained model weights",
+ )
+ parser.add_argument(
+ "--output-dir",
+ default="",
+ type=str,
+ help="Output directory to write results and logs",
+ )
+ parser.add_argument(
+ "--opts",
+ help="Extra configuration options",
+ default=[],
+ nargs="+",
+ )
+ return parser
+
+
+def get_autocast_dtype(config):
+ teacher_dtype_str = config.compute_precision.teacher.backbone.mixed_precision.param_dtype
+ if teacher_dtype_str == "fp16":
+ return torch.half
+ elif teacher_dtype_str == "bf16":
+ return torch.bfloat16
+ else:
+ return torch.float
+
+
+def build_model_for_eval(config, pretrained_weights):
+ model, _ = build_model_from_cfg(config, only_teacher=True)
+ dinov2_utils.load_pretrained_weights(model, pretrained_weights, "teacher")
+ model.eval()
+ model.cuda()
+ return model
+
+
+def setup_and_build_model(args) -> Tuple[Any, torch.dtype]:
+ cudnn.benchmark = True
+ config = setup(args)
+ model = build_model_for_eval(config, args.pretrained_weights)
+ autocast_dtype = get_autocast_dtype(config)
+ return model, autocast_dtype
diff --git a/models/dsp/dinov2/dinov2/eval/utils.py b/models/dsp/dinov2/dinov2/eval/utils.py
new file mode 100644
index 0000000000000000000000000000000000000000..c50576b1940587ee64b7a422e2e96b475d60fd39
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/eval/utils.py
@@ -0,0 +1,146 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import logging
+from typing import Dict, Optional
+
+import torch
+from torch import nn
+from torchmetrics import MetricCollection
+
+from dinov2.data import DatasetWithEnumeratedTargets, SamplerType, make_data_loader
+import dinov2.distributed as distributed
+from dinov2.logging import MetricLogger
+
+
+logger = logging.getLogger("dinov2")
+
+
+class ModelWithNormalize(torch.nn.Module):
+ def __init__(self, model):
+ super().__init__()
+ self.model = model
+
+ def forward(self, samples):
+ return nn.functional.normalize(self.model(samples), dim=1, p=2)
+
+
+class ModelWithIntermediateLayers(nn.Module):
+ def __init__(self, feature_model, n_last_blocks, autocast_ctx):
+ super().__init__()
+ self.feature_model = feature_model
+ self.feature_model.eval()
+ self.n_last_blocks = n_last_blocks
+ self.autocast_ctx = autocast_ctx
+
+ def forward(self, images):
+ with torch.inference_mode():
+ with self.autocast_ctx():
+ features = self.feature_model.get_intermediate_layers(
+ images, self.n_last_blocks, return_class_token=True
+ )
+ return features
+
+
+@torch.inference_mode()
+def evaluate(
+ model: nn.Module,
+ data_loader,
+ postprocessors: Dict[str, nn.Module],
+ metrics: Dict[str, MetricCollection],
+ device: torch.device,
+ criterion: Optional[nn.Module] = None,
+):
+ model.eval()
+ if criterion is not None:
+ criterion.eval()
+
+ for metric in metrics.values():
+ metric = metric.to(device)
+
+ metric_logger = MetricLogger(delimiter=" ")
+ header = "Test:"
+
+ for samples, targets, *_ in metric_logger.log_every(data_loader, 10, header):
+ outputs = model(samples.to(device))
+ targets = targets.to(device)
+
+ if criterion is not None:
+ loss = criterion(outputs, targets)
+ metric_logger.update(loss=loss.item())
+
+ for k, metric in metrics.items():
+ metric_inputs = postprocessors[k](outputs, targets)
+ metric.update(**metric_inputs)
+
+ metric_logger.synchronize_between_processes()
+ logger.info(f"Averaged stats: {metric_logger}")
+
+ stats = {k: metric.compute() for k, metric in metrics.items()}
+ metric_logger_stats = {k: meter.global_avg for k, meter in metric_logger.meters.items()}
+ return metric_logger_stats, stats
+
+
+def all_gather_and_flatten(tensor_rank):
+ tensor_all_ranks = torch.empty(
+ distributed.get_global_size(),
+ *tensor_rank.shape,
+ dtype=tensor_rank.dtype,
+ device=tensor_rank.device,
+ )
+ tensor_list = list(tensor_all_ranks.unbind(0))
+ torch.distributed.all_gather(tensor_list, tensor_rank.contiguous())
+ return tensor_all_ranks.flatten(end_dim=1)
+
+
+def extract_features(model, dataset, batch_size, num_workers, gather_on_cpu=False):
+ dataset_with_enumerated_targets = DatasetWithEnumeratedTargets(dataset)
+ sample_count = len(dataset_with_enumerated_targets)
+ data_loader = make_data_loader(
+ dataset=dataset_with_enumerated_targets,
+ batch_size=batch_size,
+ num_workers=num_workers,
+ sampler_type=SamplerType.DISTRIBUTED,
+ drop_last=False,
+ shuffle=False,
+ )
+ return extract_features_with_dataloader(model, data_loader, sample_count, gather_on_cpu)
+
+
+@torch.inference_mode()
+def extract_features_with_dataloader(model, data_loader, sample_count, gather_on_cpu=False):
+ gather_device = torch.device("cpu") if gather_on_cpu else torch.device("cuda")
+ metric_logger = MetricLogger(delimiter=" ")
+ features, all_labels = None, None
+ for samples, (index, labels_rank) in metric_logger.log_every(data_loader, 10):
+ samples = samples.cuda(non_blocking=True)
+ labels_rank = labels_rank.cuda(non_blocking=True)
+ index = index.cuda(non_blocking=True)
+ features_rank = model(samples).float()
+
+ # init storage feature matrix
+ if features is None:
+ features = torch.zeros(sample_count, features_rank.shape[-1], device=gather_device)
+ labels_shape = list(labels_rank.shape)
+ labels_shape[0] = sample_count
+ all_labels = torch.full(labels_shape, fill_value=-1, device=gather_device)
+ logger.info(f"Storing features into tensor of shape {features.shape}")
+
+ # share indexes, features and labels between processes
+ index_all = all_gather_and_flatten(index).to(gather_device)
+ features_all_ranks = all_gather_and_flatten(features_rank).to(gather_device)
+ labels_all_ranks = all_gather_and_flatten(labels_rank).to(gather_device)
+
+ # update storage feature matrix
+ if len(index_all) > 0:
+ features.index_copy_(0, index_all, features_all_ranks)
+ all_labels.index_copy_(0, index_all, labels_all_ranks)
+
+ logger.info(f"Features shape: {tuple(features.shape)}")
+ logger.info(f"Labels shape: {tuple(all_labels.shape)}")
+
+ assert torch.all(all_labels > -1)
+
+ return features, all_labels
diff --git a/models/dsp/dinov2/dinov2/fsdp/__init__.py b/models/dsp/dinov2/dinov2/fsdp/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..ed454480e0b76e761d657cc40fd097bd339d15a2
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/fsdp/__init__.py
@@ -0,0 +1,157 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import os
+from typing import Any
+
+import torch
+import dinov2.distributed as distributed
+from functools import partial
+from fvcore.common.checkpoint import Checkpointer
+from torch.distributed.fsdp import FullyShardedDataParallel as FSDP
+from torch.distributed.fsdp import ShardingStrategy
+from torch.distributed.fsdp import MixedPrecision
+from torch.distributed.fsdp import StateDictType
+from torch.distributed.fsdp.sharded_grad_scaler import ShardedGradScaler
+from torch.distributed.fsdp.wrap import ModuleWrapPolicy
+from torch.distributed.fsdp._runtime_utils import _reshard
+
+
+def get_fsdp_wrapper(model_cfg, modules_to_wrap=set()):
+ sharding_strategy_dict = {
+ "NO_SHARD": ShardingStrategy.NO_SHARD,
+ "SHARD_GRAD_OP": ShardingStrategy.SHARD_GRAD_OP,
+ "FULL_SHARD": ShardingStrategy.FULL_SHARD,
+ }
+
+ dtype_dict = {
+ "fp32": torch.float32,
+ "fp16": torch.float16,
+ "bf16": torch.bfloat16,
+ }
+
+ mixed_precision_config = MixedPrecision(
+ param_dtype=dtype_dict[model_cfg.mixed_precision.param_dtype],
+ reduce_dtype=dtype_dict[model_cfg.mixed_precision.reduce_dtype],
+ buffer_dtype=dtype_dict[model_cfg.mixed_precision.buffer_dtype],
+ )
+
+ sharding_strategy_config = sharding_strategy_dict[model_cfg.sharding_strategy]
+
+ local_rank = distributed.get_local_rank()
+
+ fsdp_wrapper = partial(
+ FSDP,
+ sharding_strategy=sharding_strategy_config,
+ mixed_precision=mixed_precision_config,
+ device_id=local_rank,
+ sync_module_states=True,
+ use_orig_params=True,
+ auto_wrap_policy=ModuleWrapPolicy(modules_to_wrap),
+ )
+ return fsdp_wrapper
+
+
+def is_fsdp(x):
+ return isinstance(x, FSDP)
+
+
+def is_sharded_fsdp(x):
+ return is_fsdp(x) and x.sharding_strategy is not ShardingStrategy.NO_SHARD
+
+
+def free_if_fsdp(x):
+ if is_sharded_fsdp(x):
+ handles = x._handles
+ true_list = [True for h in handles]
+ _reshard(x, handles, true_list)
+
+
+def get_fsdp_modules(x):
+ return FSDP.fsdp_modules(x)
+
+
+def reshard_fsdp_model(x):
+ for m in get_fsdp_modules(x):
+ free_if_fsdp(m)
+
+
+def rankstr():
+ return f"rank_{distributed.get_global_rank()}"
+
+
+class FSDPCheckpointer(Checkpointer):
+ def save(self, name: str, **kwargs: Any) -> None:
+ """
+ Dump model and checkpointables to a file.
+
+ Args:
+ name (str): name of the file.
+ kwargs (dict): extra arbitrary data to save.
+ """
+ if not self.save_dir or not self.save_to_disk:
+ return
+
+ data = {}
+ with FSDP.state_dict_type(self.model, StateDictType.LOCAL_STATE_DICT):
+ data["model"] = self.model.state_dict()
+
+ # data["model"] = self.model.state_dict()
+ for key, obj in self.checkpointables.items():
+ data[key] = obj.state_dict()
+ data.update(kwargs)
+
+ basename = f"{name}.{rankstr()}.pth"
+ save_file = os.path.join(self.save_dir, basename)
+ assert os.path.basename(save_file) == basename, basename
+ self.logger.info("Saving checkpoint to {}".format(save_file))
+ with self.path_manager.open(save_file, "wb") as f:
+ torch.save(data, f)
+ self.tag_last_checkpoint(basename)
+
+ def load(self, *args, **kwargs):
+ with FSDP.state_dict_type(self.model, StateDictType.LOCAL_STATE_DICT):
+ return super().load(*args, **kwargs)
+
+ def has_checkpoint(self) -> bool:
+ """
+ Returns:
+ bool: whether a checkpoint exists in the target directory.
+ """
+ save_file = os.path.join(self.save_dir, f"last_checkpoint.{rankstr()}")
+ return self.path_manager.exists(save_file)
+
+ def get_checkpoint_file(self) -> str:
+ """
+ Returns:
+ str: The latest checkpoint file in target directory.
+ """
+ save_file = os.path.join(self.save_dir, f"last_checkpoint.{rankstr()}")
+ try:
+ with self.path_manager.open(save_file, "r") as f:
+ last_saved = f.read().strip()
+ except IOError:
+ # if file doesn't exist, maybe because it has just been
+ # deleted by a separate process
+ return ""
+ # pyre-fixme[6]: For 2nd param expected `Union[PathLike[str], str]` but got
+ # `Union[bytes, str]`.
+ return os.path.join(self.save_dir, last_saved)
+
+ def tag_last_checkpoint(self, last_filename_basename: str) -> None:
+ """
+ Tag the last checkpoint.
+
+ Args:
+ last_filename_basename (str): the basename of the last filename.
+ """
+ if distributed.is_enabled():
+ torch.distributed.barrier()
+ save_file = os.path.join(self.save_dir, f"last_checkpoint.{rankstr()}")
+ with self.path_manager.open(save_file, "w") as f:
+ f.write(last_filename_basename) # pyre-ignore
+
+
+ShardedGradScaler = ShardedGradScaler
diff --git a/models/dsp/dinov2/dinov2/hub/__init__.py b/models/dsp/dinov2/dinov2/hub/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..b88da6bf80be92af00b72dfdb0a806fa64a7a2d9
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/hub/__init__.py
@@ -0,0 +1,4 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
diff --git a/models/dsp/dinov2/dinov2/hub/backbones.py b/models/dsp/dinov2/dinov2/hub/backbones.py
new file mode 100644
index 0000000000000000000000000000000000000000..53fe83719d5107eb77a8f25ef1814c3d73446002
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/hub/backbones.py
@@ -0,0 +1,156 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from enum import Enum
+from typing import Union
+
+import torch
+
+from .utils import _DINOV2_BASE_URL, _make_dinov2_model_name
+
+
+class Weights(Enum):
+ LVD142M = "LVD142M"
+
+
+def _make_dinov2_model(
+ *,
+ arch_name: str = "vit_large",
+ img_size: int = 518,
+ patch_size: int = 14,
+ init_values: float = 1.0,
+ ffn_layer: str = "mlp",
+ block_chunks: int = 0,
+ num_register_tokens: int = 0,
+ interpolate_antialias: bool = False,
+ interpolate_offset: float = 0.1,
+ pretrained: bool = True,
+ weights: Union[Weights, str] = Weights.LVD142M,
+ **kwargs,
+):
+ from ..models import vision_transformer as vits
+
+ if isinstance(weights, str):
+ try:
+ weights = Weights[weights]
+ except KeyError:
+ raise AssertionError(f"Unsupported weights: {weights}")
+
+ model_base_name = _make_dinov2_model_name(arch_name, patch_size)
+ vit_kwargs = dict(
+ img_size=img_size,
+ patch_size=patch_size,
+ init_values=init_values,
+ ffn_layer=ffn_layer,
+ block_chunks=block_chunks,
+ num_register_tokens=num_register_tokens,
+ interpolate_antialias=interpolate_antialias,
+ interpolate_offset=interpolate_offset,
+ )
+ vit_kwargs.update(**kwargs)
+ model = vits.__dict__[arch_name](**vit_kwargs)
+
+ if pretrained:
+ model_full_name = _make_dinov2_model_name(arch_name, patch_size, num_register_tokens)
+ url = _DINOV2_BASE_URL + f"/{model_base_name}/{model_full_name}_pretrain.pth"
+ state_dict = torch.hub.load_state_dict_from_url(url, map_location="cpu")
+ model.load_state_dict(state_dict, strict=True)
+
+ return model
+
+
+def dinov2_vits14(*, pretrained: bool = True, weights: Union[Weights, str] = Weights.LVD142M, **kwargs):
+ """
+ DINOv2 ViT-S/14 model (optionally) pretrained on the LVD-142M dataset.
+ """
+ return _make_dinov2_model(arch_name="vit_small", pretrained=pretrained, weights=weights, **kwargs)
+
+
+def dinov2_vitb14(*, pretrained: bool = True, weights: Union[Weights, str] = Weights.LVD142M, **kwargs):
+ """
+ DINOv2 ViT-B/14 model (optionally) pretrained on the LVD-142M dataset.
+ """
+ return _make_dinov2_model(arch_name="vit_base", pretrained=pretrained, weights=weights, **kwargs)
+
+
+def dinov2_vitl14(*, pretrained: bool = True, weights: Union[Weights, str] = Weights.LVD142M, **kwargs):
+ """
+ DINOv2 ViT-L/14 model (optionally) pretrained on the LVD-142M dataset.
+ """
+ return _make_dinov2_model(arch_name="vit_large", pretrained=pretrained, weights=weights, **kwargs)
+
+
+def dinov2_vitg14(*, pretrained: bool = True, weights: Union[Weights, str] = Weights.LVD142M, **kwargs):
+ """
+ DINOv2 ViT-g/14 model (optionally) pretrained on the LVD-142M dataset.
+ """
+ return _make_dinov2_model(
+ arch_name="vit_giant2",
+ ffn_layer="swiglufused",
+ weights=weights,
+ pretrained=pretrained,
+ **kwargs,
+ )
+
+
+def dinov2_vits14_reg(*, pretrained: bool = True, weights: Union[Weights, str] = Weights.LVD142M, **kwargs):
+ """
+ DINOv2 ViT-S/14 model with registers (optionally) pretrained on the LVD-142M dataset.
+ """
+ return _make_dinov2_model(
+ arch_name="vit_small",
+ pretrained=pretrained,
+ weights=weights,
+ num_register_tokens=4,
+ interpolate_antialias=True,
+ interpolate_offset=0.0,
+ **kwargs,
+ )
+
+
+def dinov2_vitb14_reg(*, pretrained: bool = True, weights: Union[Weights, str] = Weights.LVD142M, **kwargs):
+ """
+ DINOv2 ViT-B/14 model with registers (optionally) pretrained on the LVD-142M dataset.
+ """
+ return _make_dinov2_model(
+ arch_name="vit_base",
+ pretrained=pretrained,
+ weights=weights,
+ num_register_tokens=4,
+ interpolate_antialias=True,
+ interpolate_offset=0.0,
+ **kwargs,
+ )
+
+
+def dinov2_vitl14_reg(*, pretrained: bool = True, weights: Union[Weights, str] = Weights.LVD142M, **kwargs):
+ """
+ DINOv2 ViT-L/14 model with registers (optionally) pretrained on the LVD-142M dataset.
+ """
+ return _make_dinov2_model(
+ arch_name="vit_large",
+ pretrained=pretrained,
+ weights=weights,
+ num_register_tokens=4,
+ interpolate_antialias=True,
+ interpolate_offset=0.0,
+ **kwargs,
+ )
+
+
+def dinov2_vitg14_reg(*, pretrained: bool = True, weights: Union[Weights, str] = Weights.LVD142M, **kwargs):
+ """
+ DINOv2 ViT-g/14 model with registers (optionally) pretrained on the LVD-142M dataset.
+ """
+ return _make_dinov2_model(
+ arch_name="vit_giant2",
+ ffn_layer="swiglufused",
+ weights=weights,
+ pretrained=pretrained,
+ num_register_tokens=4,
+ interpolate_antialias=True,
+ interpolate_offset=0.0,
+ **kwargs,
+ )
diff --git a/models/dsp/dinov2/dinov2/hub/classifiers.py b/models/dsp/dinov2/dinov2/hub/classifiers.py
new file mode 100644
index 0000000000000000000000000000000000000000..3f0841efa80ab3d564cd320d61da254af182606b
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/hub/classifiers.py
@@ -0,0 +1,268 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from enum import Enum
+from typing import Union
+
+import torch
+import torch.nn as nn
+
+from .backbones import _make_dinov2_model
+from .utils import _DINOV2_BASE_URL, _make_dinov2_model_name
+
+
+class Weights(Enum):
+ IMAGENET1K = "IMAGENET1K"
+
+
+def _make_dinov2_linear_classification_head(
+ *,
+ arch_name: str = "vit_large",
+ patch_size: int = 14,
+ embed_dim: int = 1024,
+ layers: int = 4,
+ pretrained: bool = True,
+ weights: Union[Weights, str] = Weights.IMAGENET1K,
+ num_register_tokens: int = 0,
+ **kwargs,
+):
+ if layers not in (1, 4):
+ raise AssertionError(f"Unsupported number of layers: {layers}")
+ if isinstance(weights, str):
+ try:
+ weights = Weights[weights]
+ except KeyError:
+ raise AssertionError(f"Unsupported weights: {weights}")
+
+ linear_head = nn.Linear((1 + layers) * embed_dim, 1_000)
+
+ if pretrained:
+ model_base_name = _make_dinov2_model_name(arch_name, patch_size)
+ model_full_name = _make_dinov2_model_name(arch_name, patch_size, num_register_tokens)
+ layers_str = str(layers) if layers == 4 else ""
+ url = _DINOV2_BASE_URL + f"/{model_base_name}/{model_full_name}_linear{layers_str}_head.pth"
+ state_dict = torch.hub.load_state_dict_from_url(url, map_location="cpu")
+ linear_head.load_state_dict(state_dict, strict=True)
+
+ return linear_head
+
+
+class _LinearClassifierWrapper(nn.Module):
+ def __init__(self, *, backbone: nn.Module, linear_head: nn.Module, layers: int = 4):
+ super().__init__()
+ self.backbone = backbone
+ self.linear_head = linear_head
+ self.layers = layers
+
+ def forward(self, x):
+ if self.layers == 1:
+ x = self.backbone.forward_features(x)
+ cls_token = x["x_norm_clstoken"]
+ patch_tokens = x["x_norm_patchtokens"]
+ # fmt: off
+ linear_input = torch.cat([
+ cls_token,
+ patch_tokens.mean(dim=1),
+ ], dim=1)
+ # fmt: on
+ elif self.layers == 4:
+ x = self.backbone.get_intermediate_layers(x, n=4, return_class_token=True)
+ # fmt: off
+ linear_input = torch.cat([
+ x[0][1],
+ x[1][1],
+ x[2][1],
+ x[3][1],
+ x[3][0].mean(dim=1),
+ ], dim=1)
+ # fmt: on
+ else:
+ assert False, f"Unsupported number of layers: {self.layers}"
+ return self.linear_head(linear_input)
+
+
+def _make_dinov2_linear_classifier(
+ *,
+ arch_name: str = "vit_large",
+ layers: int = 4,
+ pretrained: bool = True,
+ weights: Union[Weights, str] = Weights.IMAGENET1K,
+ num_register_tokens: int = 0,
+ interpolate_antialias: bool = False,
+ interpolate_offset: float = 0.1,
+ **kwargs,
+):
+ backbone = _make_dinov2_model(
+ arch_name=arch_name,
+ pretrained=pretrained,
+ num_register_tokens=num_register_tokens,
+ interpolate_antialias=interpolate_antialias,
+ interpolate_offset=interpolate_offset,
+ **kwargs,
+ )
+
+ embed_dim = backbone.embed_dim
+ patch_size = backbone.patch_size
+ linear_head = _make_dinov2_linear_classification_head(
+ arch_name=arch_name,
+ patch_size=patch_size,
+ embed_dim=embed_dim,
+ layers=layers,
+ pretrained=pretrained,
+ weights=weights,
+ num_register_tokens=num_register_tokens,
+ )
+
+ return _LinearClassifierWrapper(backbone=backbone, linear_head=linear_head, layers=layers)
+
+
+def dinov2_vits14_lc(
+ *,
+ layers: int = 4,
+ pretrained: bool = True,
+ weights: Union[Weights, str] = Weights.IMAGENET1K,
+ **kwargs,
+):
+ """
+ Linear classifier (1 or 4 layers) on top of a DINOv2 ViT-S/14 backbone (optionally) pretrained on the LVD-142M dataset and trained on ImageNet-1k.
+ """
+ return _make_dinov2_linear_classifier(
+ arch_name="vit_small",
+ layers=layers,
+ pretrained=pretrained,
+ weights=weights,
+ **kwargs,
+ )
+
+
+def dinov2_vitb14_lc(
+ *,
+ layers: int = 4,
+ pretrained: bool = True,
+ weights: Union[Weights, str] = Weights.IMAGENET1K,
+ **kwargs,
+):
+ """
+ Linear classifier (1 or 4 layers) on top of a DINOv2 ViT-B/14 backbone (optionally) pretrained on the LVD-142M dataset and trained on ImageNet-1k.
+ """
+ return _make_dinov2_linear_classifier(
+ arch_name="vit_base",
+ layers=layers,
+ pretrained=pretrained,
+ weights=weights,
+ **kwargs,
+ )
+
+
+def dinov2_vitl14_lc(
+ *,
+ layers: int = 4,
+ pretrained: bool = True,
+ weights: Union[Weights, str] = Weights.IMAGENET1K,
+ **kwargs,
+):
+ """
+ Linear classifier (1 or 4 layers) on top of a DINOv2 ViT-L/14 backbone (optionally) pretrained on the LVD-142M dataset and trained on ImageNet-1k.
+ """
+ return _make_dinov2_linear_classifier(
+ arch_name="vit_large",
+ layers=layers,
+ pretrained=pretrained,
+ weights=weights,
+ **kwargs,
+ )
+
+
+def dinov2_vitg14_lc(
+ *,
+ layers: int = 4,
+ pretrained: bool = True,
+ weights: Union[Weights, str] = Weights.IMAGENET1K,
+ **kwargs,
+):
+ """
+ Linear classifier (1 or 4 layers) on top of a DINOv2 ViT-g/14 backbone (optionally) pretrained on the LVD-142M dataset and trained on ImageNet-1k.
+ """
+ return _make_dinov2_linear_classifier(
+ arch_name="vit_giant2",
+ layers=layers,
+ ffn_layer="swiglufused",
+ pretrained=pretrained,
+ weights=weights,
+ **kwargs,
+ )
+
+
+def dinov2_vits14_reg_lc(
+ *, layers: int = 4, pretrained: bool = True, weights: Union[Weights, str] = Weights.IMAGENET1K, **kwargs
+):
+ """
+ Linear classifier (1 or 4 layers) on top of a DINOv2 ViT-S/14 backbone with registers (optionally) pretrained on the LVD-142M dataset and trained on ImageNet-1k.
+ """
+ return _make_dinov2_linear_classifier(
+ arch_name="vit_small",
+ layers=layers,
+ pretrained=pretrained,
+ weights=weights,
+ num_register_tokens=4,
+ interpolate_antialias=True,
+ interpolate_offset=0.0,
+ **kwargs,
+ )
+
+
+def dinov2_vitb14_reg_lc(
+ *, layers: int = 4, pretrained: bool = True, weights: Union[Weights, str] = Weights.IMAGENET1K, **kwargs
+):
+ """
+ Linear classifier (1 or 4 layers) on top of a DINOv2 ViT-B/14 backbone with registers (optionally) pretrained on the LVD-142M dataset and trained on ImageNet-1k.
+ """
+ return _make_dinov2_linear_classifier(
+ arch_name="vit_base",
+ layers=layers,
+ pretrained=pretrained,
+ weights=weights,
+ num_register_tokens=4,
+ interpolate_antialias=True,
+ interpolate_offset=0.0,
+ **kwargs,
+ )
+
+
+def dinov2_vitl14_reg_lc(
+ *, layers: int = 4, pretrained: bool = True, weights: Union[Weights, str] = Weights.IMAGENET1K, **kwargs
+):
+ """
+ Linear classifier (1 or 4 layers) on top of a DINOv2 ViT-L/14 backbone with registers (optionally) pretrained on the LVD-142M dataset and trained on ImageNet-1k.
+ """
+ return _make_dinov2_linear_classifier(
+ arch_name="vit_large",
+ layers=layers,
+ pretrained=pretrained,
+ weights=weights,
+ num_register_tokens=4,
+ interpolate_antialias=True,
+ interpolate_offset=0.0,
+ **kwargs,
+ )
+
+
+def dinov2_vitg14_reg_lc(
+ *, layers: int = 4, pretrained: bool = True, weights: Union[Weights, str] = Weights.IMAGENET1K, **kwargs
+):
+ """
+ Linear classifier (1 or 4 layers) on top of a DINOv2 ViT-g/14 backbone with registers (optionally) pretrained on the LVD-142M dataset and trained on ImageNet-1k.
+ """
+ return _make_dinov2_linear_classifier(
+ arch_name="vit_giant2",
+ layers=layers,
+ ffn_layer="swiglufused",
+ pretrained=pretrained,
+ weights=weights,
+ num_register_tokens=4,
+ interpolate_antialias=True,
+ interpolate_offset=0.0,
+ **kwargs,
+ )
diff --git a/models/dsp/dinov2/dinov2/hub/depth/__init__.py b/models/dsp/dinov2/dinov2/hub/depth/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..91716e58ab6158d814df8c653644d9af4c7be65c
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/hub/depth/__init__.py
@@ -0,0 +1,7 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .decode_heads import BNHead, DPTHead
+from .encoder_decoder import DepthEncoderDecoder
diff --git a/models/dsp/dinov2/dinov2/hub/depth/decode_heads.py b/models/dsp/dinov2/dinov2/hub/depth/decode_heads.py
new file mode 100644
index 0000000000000000000000000000000000000000..f455accad38fec6ecdd53460233a564c34f434da
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/hub/depth/decode_heads.py
@@ -0,0 +1,747 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import copy
+from functools import partial
+import math
+import warnings
+
+import torch
+import torch.nn as nn
+
+from .ops import resize
+
+
+# XXX: (Untested) replacement for mmcv.imdenormalize()
+def _imdenormalize(img, mean, std, to_bgr=True):
+ import numpy as np
+
+ mean = mean.reshape(1, -1).astype(np.float64)
+ std = std.reshape(1, -1).astype(np.float64)
+ img = (img * std) + mean
+ if to_bgr:
+ img = img[::-1]
+ return img
+
+
+class DepthBaseDecodeHead(nn.Module):
+ """Base class for BaseDecodeHead.
+
+ Args:
+ in_channels (List): Input channels.
+ channels (int): Channels after modules, before conv_depth.
+ conv_layer (nn.Module): Conv layers. Default: None.
+ act_layer (nn.Module): Activation layers. Default: nn.ReLU.
+ loss_decode (dict): Config of decode loss.
+ Default: ().
+ sampler (dict|None): The config of depth map sampler.
+ Default: None.
+ align_corners (bool): align_corners argument of F.interpolate.
+ Default: False.
+ min_depth (int): Min depth in dataset setting.
+ Default: 1e-3.
+ max_depth (int): Max depth in dataset setting.
+ Default: None.
+ norm_layer (dict|None): Norm layers.
+ Default: None.
+ classify (bool): Whether predict depth in a cls.-reg. manner.
+ Default: False.
+ n_bins (int): The number of bins used in cls. step.
+ Default: 256.
+ bins_strategy (str): The discrete strategy used in cls. step.
+ Default: 'UD'.
+ norm_strategy (str): The norm strategy on cls. probability
+ distribution. Default: 'linear'
+ scale_up (str): Whether predict depth in a scale-up manner.
+ Default: False.
+ """
+
+ def __init__(
+ self,
+ in_channels,
+ conv_layer=None,
+ act_layer=nn.ReLU,
+ channels=96,
+ loss_decode=(),
+ sampler=None,
+ align_corners=False,
+ min_depth=1e-3,
+ max_depth=None,
+ norm_layer=None,
+ classify=False,
+ n_bins=256,
+ bins_strategy="UD",
+ norm_strategy="linear",
+ scale_up=False,
+ ):
+ super(DepthBaseDecodeHead, self).__init__()
+
+ self.in_channels = in_channels
+ self.channels = channels
+ self.conf_layer = conv_layer
+ self.act_layer = act_layer
+ self.loss_decode = loss_decode
+ self.align_corners = align_corners
+ self.min_depth = min_depth
+ self.max_depth = max_depth
+ self.norm_layer = norm_layer
+ self.classify = classify
+ self.n_bins = n_bins
+ self.scale_up = scale_up
+
+ if self.classify:
+ assert bins_strategy in ["UD", "SID"], "Support bins_strategy: UD, SID"
+ assert norm_strategy in ["linear", "softmax", "sigmoid"], "Support norm_strategy: linear, softmax, sigmoid"
+
+ self.bins_strategy = bins_strategy
+ self.norm_strategy = norm_strategy
+ self.softmax = nn.Softmax(dim=1)
+ self.conv_depth = nn.Conv2d(channels, n_bins, kernel_size=3, padding=1, stride=1)
+ else:
+ self.conv_depth = nn.Conv2d(channels, 1, kernel_size=3, padding=1, stride=1)
+
+ self.relu = nn.ReLU()
+ self.sigmoid = nn.Sigmoid()
+
+ def forward(self, inputs, img_metas):
+ """Placeholder of forward function."""
+ pass
+
+ def forward_train(self, img, inputs, img_metas, depth_gt):
+ """Forward function for training.
+ Args:
+ inputs (list[Tensor]): List of multi-level img features.
+ img_metas (list[dict]): List of image info dict where each dict
+ has: 'img_shape', 'scale_factor', 'flip', and may also contain
+ 'filename', 'ori_shape', 'pad_shape', and 'img_norm_cfg'.
+ For details on the values of these keys see
+ `depth/datasets/pipelines/formatting.py:Collect`.
+ depth_gt (Tensor): GT depth
+
+ Returns:
+ dict[str, Tensor]: a dictionary of loss components
+ """
+ depth_pred = self.forward(inputs, img_metas)
+ losses = self.losses(depth_pred, depth_gt)
+
+ log_imgs = self.log_images(img[0], depth_pred[0], depth_gt[0], img_metas[0])
+ losses.update(**log_imgs)
+
+ return losses
+
+ def forward_test(self, inputs, img_metas):
+ """Forward function for testing.
+ Args:
+ inputs (list[Tensor]): List of multi-level img features.
+ img_metas (list[dict]): List of image info dict where each dict
+ has: 'img_shape', 'scale_factor', 'flip', and may also contain
+ 'filename', 'ori_shape', 'pad_shape', and 'img_norm_cfg'.
+ For details on the values of these keys see
+ `depth/datasets/pipelines/formatting.py:Collect`.
+
+ Returns:
+ Tensor: Output depth map.
+ """
+ return self.forward(inputs, img_metas)
+
+ def depth_pred(self, feat):
+ """Prediction each pixel."""
+ if self.classify:
+ logit = self.conv_depth(feat)
+
+ if self.bins_strategy == "UD":
+ bins = torch.linspace(self.min_depth, self.max_depth, self.n_bins, device=feat.device)
+ elif self.bins_strategy == "SID":
+ bins = torch.logspace(self.min_depth, self.max_depth, self.n_bins, device=feat.device)
+
+ # following Adabins, default linear
+ if self.norm_strategy == "linear":
+ logit = torch.relu(logit)
+ eps = 0.1
+ logit = logit + eps
+ logit = logit / logit.sum(dim=1, keepdim=True)
+ elif self.norm_strategy == "softmax":
+ logit = torch.softmax(logit, dim=1)
+ elif self.norm_strategy == "sigmoid":
+ logit = torch.sigmoid(logit)
+ logit = logit / logit.sum(dim=1, keepdim=True)
+
+ output = torch.einsum("ikmn,k->imn", [logit, bins]).unsqueeze(dim=1)
+
+ else:
+ if self.scale_up:
+ output = self.sigmoid(self.conv_depth(feat)) * self.max_depth
+ else:
+ output = self.relu(self.conv_depth(feat)) + self.min_depth
+ return output
+
+ def losses(self, depth_pred, depth_gt):
+ """Compute depth loss."""
+ loss = dict()
+ depth_pred = resize(
+ input=depth_pred, size=depth_gt.shape[2:], mode="bilinear", align_corners=self.align_corners, warning=False
+ )
+ if not isinstance(self.loss_decode, nn.ModuleList):
+ losses_decode = [self.loss_decode]
+ else:
+ losses_decode = self.loss_decode
+ for loss_decode in losses_decode:
+ if loss_decode.loss_name not in loss:
+ loss[loss_decode.loss_name] = loss_decode(depth_pred, depth_gt)
+ else:
+ loss[loss_decode.loss_name] += loss_decode(depth_pred, depth_gt)
+ return loss
+
+ def log_images(self, img_path, depth_pred, depth_gt, img_meta):
+ import numpy as np
+
+ show_img = copy.deepcopy(img_path.detach().cpu().permute(1, 2, 0))
+ show_img = show_img.numpy().astype(np.float32)
+ show_img = _imdenormalize(
+ show_img,
+ img_meta["img_norm_cfg"]["mean"],
+ img_meta["img_norm_cfg"]["std"],
+ img_meta["img_norm_cfg"]["to_rgb"],
+ )
+ show_img = np.clip(show_img, 0, 255)
+ show_img = show_img.astype(np.uint8)
+ show_img = show_img[:, :, ::-1]
+ show_img = show_img.transpose(0, 2, 1)
+ show_img = show_img.transpose(1, 0, 2)
+
+ depth_pred = depth_pred / torch.max(depth_pred)
+ depth_gt = depth_gt / torch.max(depth_gt)
+
+ depth_pred_color = copy.deepcopy(depth_pred.detach().cpu())
+ depth_gt_color = copy.deepcopy(depth_gt.detach().cpu())
+
+ return {"img_rgb": show_img, "img_depth_pred": depth_pred_color, "img_depth_gt": depth_gt_color}
+
+
+class BNHead(DepthBaseDecodeHead):
+ """Just a batchnorm."""
+
+ def __init__(self, input_transform="resize_concat", in_index=(0, 1, 2, 3), upsample=1, **kwargs):
+ super().__init__(**kwargs)
+ self.input_transform = input_transform
+ self.in_index = in_index
+ self.upsample = upsample
+ # self.bn = nn.SyncBatchNorm(self.in_channels)
+ if self.classify:
+ self.conv_depth = nn.Conv2d(self.channels, self.n_bins, kernel_size=1, padding=0, stride=1)
+ else:
+ self.conv_depth = nn.Conv2d(self.channels, 1, kernel_size=1, padding=0, stride=1)
+
+ def _transform_inputs(self, inputs):
+ """Transform inputs for decoder.
+ Args:
+ inputs (list[Tensor]): List of multi-level img features.
+ Returns:
+ Tensor: The transformed inputs
+ """
+
+ if "concat" in self.input_transform:
+ inputs = [inputs[i] for i in self.in_index]
+ if "resize" in self.input_transform:
+ inputs = [
+ resize(
+ input=x,
+ size=[s * self.upsample for s in inputs[0].shape[2:]],
+ mode="bilinear",
+ align_corners=self.align_corners,
+ )
+ for x in inputs
+ ]
+ inputs = torch.cat(inputs, dim=1)
+ elif self.input_transform == "multiple_select":
+ inputs = [inputs[i] for i in self.in_index]
+ else:
+ inputs = inputs[self.in_index]
+
+ return inputs
+
+ def _forward_feature(self, inputs, img_metas=None, **kwargs):
+ """Forward function for feature maps before classifying each pixel with
+ ``self.cls_seg`` fc.
+ Args:
+ inputs (list[Tensor]): List of multi-level img features.
+ Returns:
+ feats (Tensor): A tensor of shape (batch_size, self.channels,
+ H, W) which is feature map for last layer of decoder head.
+ """
+ # accept lists (for cls token)
+ inputs = list(inputs)
+ for i, x in enumerate(inputs):
+ if len(x) == 2:
+ x, cls_token = x[0], x[1]
+ if len(x.shape) == 2:
+ x = x[:, :, None, None]
+ cls_token = cls_token[:, :, None, None].expand_as(x)
+ inputs[i] = torch.cat((x, cls_token), 1)
+ else:
+ x = x[0]
+ if len(x.shape) == 2:
+ x = x[:, :, None, None]
+ inputs[i] = x
+ x = self._transform_inputs(inputs)
+ # feats = self.bn(x)
+ return x
+
+ def forward(self, inputs, img_metas=None, **kwargs):
+ """Forward function."""
+ output = self._forward_feature(inputs, img_metas=img_metas, **kwargs)
+ output = self.depth_pred(output)
+ return output
+
+
+class ConvModule(nn.Module):
+ """A conv block that bundles conv/norm/activation layers.
+
+ This block simplifies the usage of convolution layers, which are commonly
+ used with a norm layer (e.g., BatchNorm) and activation layer (e.g., ReLU).
+ It is based upon three build methods: `build_conv_layer()`,
+ `build_norm_layer()` and `build_activation_layer()`.
+
+ Besides, we add some additional features in this module.
+ 1. Automatically set `bias` of the conv layer.
+ 2. Spectral norm is supported.
+ 3. More padding modes are supported. Before PyTorch 1.5, nn.Conv2d only
+ supports zero and circular padding, and we add "reflect" padding mode.
+
+ Args:
+ in_channels (int): Number of channels in the input feature map.
+ Same as that in ``nn._ConvNd``.
+ out_channels (int): Number of channels produced by the convolution.
+ Same as that in ``nn._ConvNd``.
+ kernel_size (int | tuple[int]): Size of the convolving kernel.
+ Same as that in ``nn._ConvNd``.
+ stride (int | tuple[int]): Stride of the convolution.
+ Same as that in ``nn._ConvNd``.
+ padding (int | tuple[int]): Zero-padding added to both sides of
+ the input. Same as that in ``nn._ConvNd``.
+ dilation (int | tuple[int]): Spacing between kernel elements.
+ Same as that in ``nn._ConvNd``.
+ groups (int): Number of blocked connections from input channels to
+ output channels. Same as that in ``nn._ConvNd``.
+ bias (bool | str): If specified as `auto`, it will be decided by the
+ norm_layer. Bias will be set as True if `norm_layer` is None, otherwise
+ False. Default: "auto".
+ conv_layer (nn.Module): Convolution layer. Default: None,
+ which means using conv2d.
+ norm_layer (nn.Module): Normalization layer. Default: None.
+ act_layer (nn.Module): Activation layer. Default: nn.ReLU.
+ inplace (bool): Whether to use inplace mode for activation.
+ Default: True.
+ with_spectral_norm (bool): Whether use spectral norm in conv module.
+ Default: False.
+ padding_mode (str): If the `padding_mode` has not been supported by
+ current `Conv2d` in PyTorch, we will use our own padding layer
+ instead. Currently, we support ['zeros', 'circular'] with official
+ implementation and ['reflect'] with our own implementation.
+ Default: 'zeros'.
+ order (tuple[str]): The order of conv/norm/activation layers. It is a
+ sequence of "conv", "norm" and "act". Common examples are
+ ("conv", "norm", "act") and ("act", "conv", "norm").
+ Default: ('conv', 'norm', 'act').
+ """
+
+ _abbr_ = "conv_block"
+
+ def __init__(
+ self,
+ in_channels,
+ out_channels,
+ kernel_size,
+ stride=1,
+ padding=0,
+ dilation=1,
+ groups=1,
+ bias="auto",
+ conv_layer=nn.Conv2d,
+ norm_layer=None,
+ act_layer=nn.ReLU,
+ inplace=True,
+ with_spectral_norm=False,
+ padding_mode="zeros",
+ order=("conv", "norm", "act"),
+ ):
+ super(ConvModule, self).__init__()
+ official_padding_mode = ["zeros", "circular"]
+ self.conv_layer = conv_layer
+ self.norm_layer = norm_layer
+ self.act_layer = act_layer
+ self.inplace = inplace
+ self.with_spectral_norm = with_spectral_norm
+ self.with_explicit_padding = padding_mode not in official_padding_mode
+ self.order = order
+ assert isinstance(self.order, tuple) and len(self.order) == 3
+ assert set(order) == set(["conv", "norm", "act"])
+
+ self.with_norm = norm_layer is not None
+ self.with_activation = act_layer is not None
+ # if the conv layer is before a norm layer, bias is unnecessary.
+ if bias == "auto":
+ bias = not self.with_norm
+ self.with_bias = bias
+
+ if self.with_explicit_padding:
+ if padding_mode == "zeros":
+ padding_layer = nn.ZeroPad2d
+ else:
+ raise AssertionError(f"Unsupported padding mode: {padding_mode}")
+ self.pad = padding_layer(padding)
+
+ # reset padding to 0 for conv module
+ conv_padding = 0 if self.with_explicit_padding else padding
+ # build convolution layer
+ self.conv = self.conv_layer(
+ in_channels,
+ out_channels,
+ kernel_size,
+ stride=stride,
+ padding=conv_padding,
+ dilation=dilation,
+ groups=groups,
+ bias=bias,
+ )
+ # export the attributes of self.conv to a higher level for convenience
+ self.in_channels = self.conv.in_channels
+ self.out_channels = self.conv.out_channels
+ self.kernel_size = self.conv.kernel_size
+ self.stride = self.conv.stride
+ self.padding = padding
+ self.dilation = self.conv.dilation
+ self.transposed = self.conv.transposed
+ self.output_padding = self.conv.output_padding
+ self.groups = self.conv.groups
+
+ if self.with_spectral_norm:
+ self.conv = nn.utils.spectral_norm(self.conv)
+
+ # build normalization layers
+ if self.with_norm:
+ # norm layer is after conv layer
+ if order.index("norm") > order.index("conv"):
+ norm_channels = out_channels
+ else:
+ norm_channels = in_channels
+ norm = partial(norm_layer, num_features=norm_channels)
+ self.add_module("norm", norm)
+ if self.with_bias:
+ from torch.nnModules.batchnorm import _BatchNorm
+ from torch.nnModules.instancenorm import _InstanceNorm
+
+ if isinstance(norm, (_BatchNorm, _InstanceNorm)):
+ warnings.warn("Unnecessary conv bias before batch/instance norm")
+ else:
+ self.norm_name = None
+
+ # build activation layer
+ if self.with_activation:
+ # nn.Tanh has no 'inplace' argument
+ # (nn.Tanh, nn.PReLU, nn.Sigmoid, nn.HSigmoid, nn.Swish, nn.GELU)
+ if not isinstance(act_layer, (nn.Tanh, nn.PReLU, nn.Sigmoid, nn.GELU)):
+ act_layer = partial(act_layer, inplace=inplace)
+ self.activate = act_layer()
+
+ # Use msra init by default
+ self.init_weights()
+
+ @property
+ def norm(self):
+ if self.norm_name:
+ return getattr(self, self.norm_name)
+ else:
+ return None
+
+ def init_weights(self):
+ # 1. It is mainly for customized conv layers with their own
+ # initialization manners by calling their own ``init_weights()``,
+ # and we do not want ConvModule to override the initialization.
+ # 2. For customized conv layers without their own initialization
+ # manners (that is, they don't have their own ``init_weights()``)
+ # and PyTorch's conv layers, they will be initialized by
+ # this method with default ``kaiming_init``.
+ # Note: For PyTorch's conv layers, they will be overwritten by our
+ # initialization implementation using default ``kaiming_init``.
+ if not hasattr(self.conv, "init_weights"):
+ if self.with_activation and isinstance(self.act_layer, nn.LeakyReLU):
+ nonlinearity = "leaky_relu"
+ a = 0.01 # XXX: default negative_slope
+ else:
+ nonlinearity = "relu"
+ a = 0
+ if hasattr(self.conv, "weight") and self.conv.weight is not None:
+ nn.init.kaiming_normal_(self.conv.weight, a=a, mode="fan_out", nonlinearity=nonlinearity)
+ if hasattr(self.conv, "bias") and self.conv.bias is not None:
+ nn.init.constant_(self.conv.bias, 0)
+ if self.with_norm:
+ if hasattr(self.norm, "weight") and self.norm.weight is not None:
+ nn.init.constant_(self.norm.weight, 1)
+ if hasattr(self.norm, "bias") and self.norm.bias is not None:
+ nn.init.constant_(self.norm.bias, 0)
+
+ def forward(self, x, activate=True, norm=True):
+ for layer in self.order:
+ if layer == "conv":
+ if self.with_explicit_padding:
+ x = self.pad(x)
+ x = self.conv(x)
+ elif layer == "norm" and norm and self.with_norm:
+ x = self.norm(x)
+ elif layer == "act" and activate and self.with_activation:
+ x = self.activate(x)
+ return x
+
+
+class Interpolate(nn.Module):
+ def __init__(self, scale_factor, mode, align_corners=False):
+ super(Interpolate, self).__init__()
+ self.interp = nn.functional.interpolate
+ self.scale_factor = scale_factor
+ self.mode = mode
+ self.align_corners = align_corners
+
+ def forward(self, x):
+ x = self.interp(x, scale_factor=self.scale_factor, mode=self.mode, align_corners=self.align_corners)
+ return x
+
+
+class HeadDepth(nn.Module):
+ def __init__(self, features):
+ super(HeadDepth, self).__init__()
+ self.head = nn.Sequential(
+ nn.Conv2d(features, features // 2, kernel_size=3, stride=1, padding=1),
+ Interpolate(scale_factor=2, mode="bilinear", align_corners=True),
+ nn.Conv2d(features // 2, 32, kernel_size=3, stride=1, padding=1),
+ nn.ReLU(),
+ nn.Conv2d(32, 1, kernel_size=1, stride=1, padding=0),
+ )
+
+ def forward(self, x):
+ x = self.head(x)
+ return x
+
+
+class ReassembleBlocks(nn.Module):
+ """ViTPostProcessBlock, process cls_token in ViT backbone output and
+ rearrange the feature vector to feature map.
+ Args:
+ in_channels (int): ViT feature channels. Default: 768.
+ out_channels (List): output channels of each stage.
+ Default: [96, 192, 384, 768].
+ readout_type (str): Type of readout operation. Default: 'ignore'.
+ patch_size (int): The patch size. Default: 16.
+ """
+
+ def __init__(self, in_channels=768, out_channels=[96, 192, 384, 768], readout_type="ignore", patch_size=16):
+ super(ReassembleBlocks, self).__init__()
+
+ assert readout_type in ["ignore", "add", "project"]
+ self.readout_type = readout_type
+ self.patch_size = patch_size
+
+ self.projects = nn.ModuleList(
+ [
+ ConvModule(
+ in_channels=in_channels,
+ out_channels=out_channel,
+ kernel_size=1,
+ act_layer=None,
+ )
+ for out_channel in out_channels
+ ]
+ )
+
+ self.resize_layers = nn.ModuleList(
+ [
+ nn.ConvTranspose2d(
+ in_channels=out_channels[0], out_channels=out_channels[0], kernel_size=4, stride=4, padding=0
+ ),
+ nn.ConvTranspose2d(
+ in_channels=out_channels[1], out_channels=out_channels[1], kernel_size=2, stride=2, padding=0
+ ),
+ nn.Identity(),
+ nn.Conv2d(
+ in_channels=out_channels[3], out_channels=out_channels[3], kernel_size=3, stride=2, padding=1
+ ),
+ ]
+ )
+ if self.readout_type == "project":
+ self.readout_projects = nn.ModuleList()
+ for _ in range(len(self.projects)):
+ self.readout_projects.append(nn.Sequential(nn.Linear(2 * in_channels, in_channels), nn.GELU()))
+
+ def forward(self, inputs):
+ assert isinstance(inputs, list)
+ out = []
+ for i, x in enumerate(inputs):
+ assert len(x) == 2
+ x, cls_token = x[0], x[1]
+ feature_shape = x.shape
+ if self.readout_type == "project":
+ x = x.flatten(2).permute((0, 2, 1))
+ readout = cls_token.unsqueeze(1).expand_as(x)
+ x = self.readout_projects[i](torch.cat((x, readout), -1))
+ x = x.permute(0, 2, 1).reshape(feature_shape)
+ elif self.readout_type == "add":
+ x = x.flatten(2) + cls_token.unsqueeze(-1)
+ x = x.reshape(feature_shape)
+ else:
+ pass
+ x = self.projects[i](x)
+ x = self.resize_layers[i](x)
+ out.append(x)
+ return out
+
+
+class PreActResidualConvUnit(nn.Module):
+ """ResidualConvUnit, pre-activate residual unit.
+ Args:
+ in_channels (int): number of channels in the input feature map.
+ act_layer (nn.Module): activation layer.
+ norm_layer (nn.Module): norm layer.
+ stride (int): stride of the first block. Default: 1
+ dilation (int): dilation rate for convs layers. Default: 1.
+ """
+
+ def __init__(self, in_channels, act_layer, norm_layer, stride=1, dilation=1):
+ super(PreActResidualConvUnit, self).__init__()
+
+ self.conv1 = ConvModule(
+ in_channels,
+ in_channels,
+ 3,
+ stride=stride,
+ padding=dilation,
+ dilation=dilation,
+ norm_layer=norm_layer,
+ act_layer=act_layer,
+ bias=False,
+ order=("act", "conv", "norm"),
+ )
+
+ self.conv2 = ConvModule(
+ in_channels,
+ in_channels,
+ 3,
+ padding=1,
+ norm_layer=norm_layer,
+ act_layer=act_layer,
+ bias=False,
+ order=("act", "conv", "norm"),
+ )
+
+ def forward(self, inputs):
+ inputs_ = inputs.clone()
+ x = self.conv1(inputs)
+ x = self.conv2(x)
+ return x + inputs_
+
+
+class FeatureFusionBlock(nn.Module):
+ """FeatureFusionBlock, merge feature map from different stages.
+ Args:
+ in_channels (int): Input channels.
+ act_layer (nn.Module): activation layer for ResidualConvUnit.
+ norm_layer (nn.Module): normalization layer.
+ expand (bool): Whether expand the channels in post process block.
+ Default: False.
+ align_corners (bool): align_corner setting for bilinear upsample.
+ Default: True.
+ """
+
+ def __init__(self, in_channels, act_layer, norm_layer, expand=False, align_corners=True):
+ super(FeatureFusionBlock, self).__init__()
+
+ self.in_channels = in_channels
+ self.expand = expand
+ self.align_corners = align_corners
+
+ self.out_channels = in_channels
+ if self.expand:
+ self.out_channels = in_channels // 2
+
+ self.project = ConvModule(self.in_channels, self.out_channels, kernel_size=1, act_layer=None, bias=True)
+
+ self.res_conv_unit1 = PreActResidualConvUnit(
+ in_channels=self.in_channels, act_layer=act_layer, norm_layer=norm_layer
+ )
+ self.res_conv_unit2 = PreActResidualConvUnit(
+ in_channels=self.in_channels, act_layer=act_layer, norm_layer=norm_layer
+ )
+
+ def forward(self, *inputs):
+ x = inputs[0]
+ if len(inputs) == 2:
+ if x.shape != inputs[1].shape:
+ res = resize(inputs[1], size=(x.shape[2], x.shape[3]), mode="bilinear", align_corners=False)
+ else:
+ res = inputs[1]
+ x = x + self.res_conv_unit1(res)
+ x = self.res_conv_unit2(x)
+ x = resize(x, scale_factor=2, mode="bilinear", align_corners=self.align_corners)
+ x = self.project(x)
+ return x
+
+
+class DPTHead(DepthBaseDecodeHead):
+ """Vision Transformers for Dense Prediction.
+ This head is implemented of `DPT `_.
+ Args:
+ embed_dims (int): The embed dimension of the ViT backbone.
+ Default: 768.
+ post_process_channels (List): Out channels of post process conv
+ layers. Default: [96, 192, 384, 768].
+ readout_type (str): Type of readout operation. Default: 'ignore'.
+ patch_size (int): The patch size. Default: 16.
+ expand_channels (bool): Whether expand the channels in post process
+ block. Default: False.
+ """
+
+ def __init__(
+ self,
+ embed_dims=768,
+ post_process_channels=[96, 192, 384, 768],
+ readout_type="ignore",
+ patch_size=16,
+ expand_channels=False,
+ **kwargs,
+ ):
+ super(DPTHead, self).__init__(**kwargs)
+
+ self.in_channels = self.in_channels
+ self.expand_channels = expand_channels
+ self.reassemble_blocks = ReassembleBlocks(embed_dims, post_process_channels, readout_type, patch_size)
+
+ self.post_process_channels = [
+ channel * math.pow(2, i) if expand_channels else channel for i, channel in enumerate(post_process_channels)
+ ]
+ self.convs = nn.ModuleList()
+ for channel in self.post_process_channels:
+ self.convs.append(ConvModule(channel, self.channels, kernel_size=3, padding=1, act_layer=None, bias=False))
+ self.fusion_blocks = nn.ModuleList()
+ for _ in range(len(self.convs)):
+ self.fusion_blocks.append(FeatureFusionBlock(self.channels, self.act_layer, self.norm_layer))
+ self.fusion_blocks[0].res_conv_unit1 = None
+ self.project = ConvModule(self.channels, self.channels, kernel_size=3, padding=1, norm_layer=self.norm_layer)
+ self.num_fusion_blocks = len(self.fusion_blocks)
+ self.num_reassemble_blocks = len(self.reassemble_blocks.resize_layers)
+ self.num_post_process_channels = len(self.post_process_channels)
+ assert self.num_fusion_blocks == self.num_reassemble_blocks
+ assert self.num_reassemble_blocks == self.num_post_process_channels
+ self.conv_depth = HeadDepth(self.channels)
+
+ def forward(self, inputs, img_metas):
+ assert len(inputs) == self.num_reassemble_blocks
+ x = [inp for inp in inputs]
+ x = self.reassemble_blocks(x)
+ x = [self.convs[i](feature) for i, feature in enumerate(x)]
+ out = self.fusion_blocks[0](x[-1])
+ for i in range(1, len(self.fusion_blocks)):
+ out = self.fusion_blocks[i](out, x[-(i + 1)])
+ out = self.project(out)
+ out = self.depth_pred(out)
+ return out
diff --git a/models/dsp/dinov2/dinov2/hub/depth/encoder_decoder.py b/models/dsp/dinov2/dinov2/hub/depth/encoder_decoder.py
new file mode 100644
index 0000000000000000000000000000000000000000..eb29ced67957a336e763b0e7c90c0eeaea36fea8
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/hub/depth/encoder_decoder.py
@@ -0,0 +1,351 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from collections import OrderedDict
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+
+from .ops import resize
+
+
+def add_prefix(inputs, prefix):
+ """Add prefix for dict.
+
+ Args:
+ inputs (dict): The input dict with str keys.
+ prefix (str): The prefix to add.
+
+ Returns:
+
+ dict: The dict with keys updated with ``prefix``.
+ """
+
+ outputs = dict()
+ for name, value in inputs.items():
+ outputs[f"{prefix}.{name}"] = value
+
+ return outputs
+
+
+class DepthEncoderDecoder(nn.Module):
+ """Encoder Decoder depther.
+
+ EncoderDecoder typically consists of backbone and decode_head.
+ """
+
+ def __init__(self, backbone, decode_head):
+ super(DepthEncoderDecoder, self).__init__()
+
+ self.backbone = backbone
+ self.decode_head = decode_head
+ self.align_corners = self.decode_head.align_corners
+
+ def extract_feat(self, img):
+ """Extract features from images."""
+ return self.backbone(img)
+
+ def encode_decode(self, img, img_metas, rescale=True, size=None):
+ """Encode images with backbone and decode into a depth estimation
+ map of the same size as input."""
+ x = self.extract_feat(img)
+ out = self._decode_head_forward_test(x, img_metas)
+ # crop the pred depth to the certain range.
+ out = torch.clamp(out, min=self.decode_head.min_depth, max=self.decode_head.max_depth)
+ if rescale:
+ if size is None:
+ if img_metas is not None:
+ size = img_metas[0]["ori_shape"][:2]
+ else:
+ size = img.shape[2:]
+ out = resize(input=out, size=size, mode="bilinear", align_corners=self.align_corners)
+ return out
+
+ def _decode_head_forward_train(self, img, x, img_metas, depth_gt, **kwargs):
+ """Run forward function and calculate loss for decode head in
+ training."""
+ losses = dict()
+ loss_decode = self.decode_head.forward_train(img, x, img_metas, depth_gt, **kwargs)
+ losses.update(add_prefix(loss_decode, "decode"))
+ return losses
+
+ def _decode_head_forward_test(self, x, img_metas):
+ """Run forward function and calculate loss for decode head in
+ inference."""
+ depth_pred = self.decode_head.forward_test(x, img_metas)
+ return depth_pred
+
+ def forward_dummy(self, img):
+ """Dummy forward function."""
+ depth = self.encode_decode(img, None)
+
+ return depth
+
+ def forward_train(self, img, img_metas, depth_gt, **kwargs):
+ """Forward function for training.
+
+ Args:
+ img (Tensor): Input images.
+ img_metas (list[dict]): List of image info dict where each dict
+ has: 'img_shape', 'scale_factor', 'flip', and may also contain
+ 'filename', 'ori_shape', 'pad_shape', and 'img_norm_cfg'.
+ For details on the values of these keys see
+ `depth/datasets/pipelines/formatting.py:Collect`.
+ depth_gt (Tensor): Depth gt
+ used if the architecture supports depth estimation task.
+
+ Returns:
+ dict[str, Tensor]: a dictionary of loss components
+ """
+
+ x = self.extract_feat(img)
+
+ losses = dict()
+
+ # the last of x saves the info from neck
+ loss_decode = self._decode_head_forward_train(img, x, img_metas, depth_gt, **kwargs)
+
+ losses.update(loss_decode)
+
+ return losses
+
+ def whole_inference(self, img, img_meta, rescale, size=None):
+ """Inference with full image."""
+ return self.encode_decode(img, img_meta, rescale, size=size)
+
+ def slide_inference(self, img, img_meta, rescale, stride, crop_size):
+ """Inference by sliding-window with overlap.
+
+ If h_crop > h_img or w_crop > w_img, the small patch will be used to
+ decode without padding.
+ """
+
+ h_stride, w_stride = stride
+ h_crop, w_crop = crop_size
+ batch_size, _, h_img, w_img = img.size()
+ h_grids = max(h_img - h_crop + h_stride - 1, 0) // h_stride + 1
+ w_grids = max(w_img - w_crop + w_stride - 1, 0) // w_stride + 1
+ preds = img.new_zeros((batch_size, 1, h_img, w_img))
+ count_mat = img.new_zeros((batch_size, 1, h_img, w_img))
+ for h_idx in range(h_grids):
+ for w_idx in range(w_grids):
+ y1 = h_idx * h_stride
+ x1 = w_idx * w_stride
+ y2 = min(y1 + h_crop, h_img)
+ x2 = min(x1 + w_crop, w_img)
+ y1 = max(y2 - h_crop, 0)
+ x1 = max(x2 - w_crop, 0)
+ crop_img = img[:, :, y1:y2, x1:x2]
+ depth_pred = self.encode_decode(crop_img, img_meta, rescale)
+ preds += F.pad(depth_pred, (int(x1), int(preds.shape[3] - x2), int(y1), int(preds.shape[2] - y2)))
+
+ count_mat[:, :, y1:y2, x1:x2] += 1
+ assert (count_mat == 0).sum() == 0
+ if torch.onnx.is_in_onnx_export():
+ # cast count_mat to constant while exporting to ONNX
+ count_mat = torch.from_numpy(count_mat.cpu().detach().numpy()).to(device=img.device)
+ preds = preds / count_mat
+ return preds
+
+ def inference(self, img, img_meta, rescale, size=None, mode="whole"):
+ """Inference with slide/whole style.
+
+ Args:
+ img (Tensor): The input image of shape (N, 3, H, W).
+ img_meta (dict): Image info dict where each dict has: 'img_shape',
+ 'scale_factor', 'flip', and may also contain
+ 'filename', 'ori_shape', 'pad_shape', and 'img_norm_cfg'.
+ For details on the values of these keys see
+ `depth/datasets/pipelines/formatting.py:Collect`.
+ rescale (bool): Whether rescale back to original shape.
+
+ Returns:
+ Tensor: The output depth map.
+ """
+
+ assert mode in ["slide", "whole"]
+ ori_shape = img_meta[0]["ori_shape"]
+ assert all(_["ori_shape"] == ori_shape for _ in img_meta)
+ if mode == "slide":
+ depth_pred = self.slide_inference(img, img_meta, rescale)
+ else:
+ depth_pred = self.whole_inference(img, img_meta, rescale, size=size)
+ output = depth_pred
+ flip = img_meta[0]["flip"]
+ if flip:
+ flip_direction = img_meta[0]["flip_direction"]
+ assert flip_direction in ["horizontal", "vertical"]
+ if flip_direction == "horizontal":
+ output = output.flip(dims=(3,))
+ elif flip_direction == "vertical":
+ output = output.flip(dims=(2,))
+
+ return output
+
+ def simple_test(self, img, img_meta, rescale=True):
+ """Simple test with single image."""
+ depth_pred = self.inference(img, img_meta, rescale)
+ if torch.onnx.is_in_onnx_export():
+ # our inference backend only support 4D output
+ depth_pred = depth_pred.unsqueeze(0)
+ return depth_pred
+ depth_pred = depth_pred.cpu().numpy()
+ # unravel batch dim
+ depth_pred = list(depth_pred)
+ return depth_pred
+
+ def aug_test(self, imgs, img_metas, rescale=True):
+ """Test with augmentations.
+
+ Only rescale=True is supported.
+ """
+ # aug_test rescale all imgs back to ori_shape for now
+ assert rescale
+ # to save memory, we get augmented depth logit inplace
+ depth_pred = self.inference(imgs[0], img_metas[0], rescale)
+ for i in range(1, len(imgs)):
+ cur_depth_pred = self.inference(imgs[i], img_metas[i], rescale, size=depth_pred.shape[-2:])
+ depth_pred += cur_depth_pred
+ depth_pred /= len(imgs)
+ depth_pred = depth_pred.cpu().numpy()
+ # unravel batch dim
+ depth_pred = list(depth_pred)
+ return depth_pred
+
+ def forward_test(self, imgs, img_metas, **kwargs):
+ """
+ Args:
+ imgs (List[Tensor]): the outer list indicates test-time
+ augmentations and inner Tensor should have a shape NxCxHxW,
+ which contains all images in the batch.
+ img_metas (List[List[dict]]): the outer list indicates test-time
+ augs (multiscale, flip, etc.) and the inner list indicates
+ images in a batch.
+ """
+ for var, name in [(imgs, "imgs"), (img_metas, "img_metas")]:
+ if not isinstance(var, list):
+ raise TypeError(f"{name} must be a list, but got " f"{type(var)}")
+ num_augs = len(imgs)
+ if num_augs != len(img_metas):
+ raise ValueError(f"num of augmentations ({len(imgs)}) != " f"num of image meta ({len(img_metas)})")
+ # all images in the same aug batch all of the same ori_shape and pad
+ # shape
+ for img_meta in img_metas:
+ ori_shapes = [_["ori_shape"] for _ in img_meta]
+ assert all(shape == ori_shapes[0] for shape in ori_shapes)
+ img_shapes = [_["img_shape"] for _ in img_meta]
+ assert all(shape == img_shapes[0] for shape in img_shapes)
+ pad_shapes = [_["pad_shape"] for _ in img_meta]
+ assert all(shape == pad_shapes[0] for shape in pad_shapes)
+
+ if num_augs == 1:
+ return self.simple_test(imgs[0], img_metas[0], **kwargs)
+ else:
+ return self.aug_test(imgs, img_metas, **kwargs)
+
+ def forward(self, img, img_metas, return_loss=True, **kwargs):
+ """Calls either :func:`forward_train` or :func:`forward_test` depending
+ on whether ``return_loss`` is ``True``.
+
+ Note this setting will change the expected inputs. When
+ ``return_loss=True``, img and img_meta are single-nested (i.e. Tensor
+ and List[dict]), and when ``resturn_loss=False``, img and img_meta
+ should be double nested (i.e. List[Tensor], List[List[dict]]), with
+ the outer list indicating test time augmentations.
+ """
+ if return_loss:
+ return self.forward_train(img, img_metas, **kwargs)
+ else:
+ return self.forward_test(img, img_metas, **kwargs)
+
+ def train_step(self, data_batch, optimizer, **kwargs):
+ """The iteration step during training.
+
+ This method defines an iteration step during training, except for the
+ back propagation and optimizer updating, which are done in an optimizer
+ hook. Note that in some complicated cases or models, the whole process
+ including back propagation and optimizer updating is also defined in
+ this method, such as GAN.
+
+ Args:
+ data (dict): The output of dataloader.
+ optimizer (:obj:`torch.optim.Optimizer` | dict): The optimizer of
+ runner is passed to ``train_step()``. This argument is unused
+ and reserved.
+
+ Returns:
+ dict: It should contain at least 3 keys: ``loss``, ``log_vars``,
+ ``num_samples``.
+ ``loss`` is a tensor for back propagation, which can be a
+ weighted sum of multiple losses.
+ ``log_vars`` contains all the variables to be sent to the
+ logger.
+ ``num_samples`` indicates the batch size (when the model is
+ DDP, it means the batch size on each GPU), which is used for
+ averaging the logs.
+ """
+ losses = self(**data_batch)
+
+ # split losses and images
+ real_losses = {}
+ log_imgs = {}
+ for k, v in losses.items():
+ if "img" in k:
+ log_imgs[k] = v
+ else:
+ real_losses[k] = v
+
+ loss, log_vars = self._parse_losses(real_losses)
+
+ outputs = dict(loss=loss, log_vars=log_vars, num_samples=len(data_batch["img_metas"]), log_imgs=log_imgs)
+
+ return outputs
+
+ def val_step(self, data_batch, **kwargs):
+ """The iteration step during validation.
+
+ This method shares the same signature as :func:`train_step`, but used
+ during val epochs. Note that the evaluation after training epochs is
+ not implemented with this method, but an evaluation hook.
+ """
+ output = self(**data_batch, **kwargs)
+ return output
+
+ @staticmethod
+ def _parse_losses(losses):
+ import torch.distributed as dist
+
+ """Parse the raw outputs (losses) of the network.
+
+ Args:
+ losses (dict): Raw output of the network, which usually contain
+ losses and other necessary information.
+
+ Returns:
+ tuple[Tensor, dict]: (loss, log_vars), loss is the loss tensor
+ which may be a weighted sum of all losses, log_vars contains
+ all the variables to be sent to the logger.
+ """
+ log_vars = OrderedDict()
+ for loss_name, loss_value in losses.items():
+ if isinstance(loss_value, torch.Tensor):
+ log_vars[loss_name] = loss_value.mean()
+ elif isinstance(loss_value, list):
+ log_vars[loss_name] = sum(_loss.mean() for _loss in loss_value)
+ else:
+ raise TypeError(f"{loss_name} is not a tensor or list of tensors")
+
+ loss = sum(_value for _key, _value in log_vars.items() if "loss" in _key)
+
+ log_vars["loss"] = loss
+ for loss_name, loss_value in log_vars.items():
+ # reduce loss when distributed training
+ if dist.is_available() and dist.is_initialized():
+ loss_value = loss_value.data.clone()
+ dist.all_reduce(loss_value.div_(dist.get_world_size()))
+ log_vars[loss_name] = loss_value.item()
+
+ return loss, log_vars
diff --git a/models/dsp/dinov2/dinov2/hub/depth/ops.py b/models/dsp/dinov2/dinov2/hub/depth/ops.py
new file mode 100644
index 0000000000000000000000000000000000000000..15880ee0cb7652d4b41c489b927bf6a156b40e5e
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/hub/depth/ops.py
@@ -0,0 +1,28 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import warnings
+
+import torch.nn.functional as F
+
+
+def resize(input, size=None, scale_factor=None, mode="nearest", align_corners=None, warning=False):
+ if warning:
+ if size is not None and align_corners:
+ input_h, input_w = tuple(int(x) for x in input.shape[2:])
+ output_h, output_w = tuple(int(x) for x in size)
+ if output_h > input_h or output_w > output_h:
+ if (
+ (output_h > 1 and output_w > 1 and input_h > 1 and input_w > 1)
+ and (output_h - 1) % (input_h - 1)
+ and (output_w - 1) % (input_w - 1)
+ ):
+ warnings.warn(
+ f"When align_corners={align_corners}, "
+ "the output would more aligned if "
+ f"input size {(input_h, input_w)} is `x+1` and "
+ f"out size {(output_h, output_w)} is `nx+1`"
+ )
+ return F.interpolate(input, size, scale_factor, mode, align_corners)
diff --git a/models/dsp/dinov2/dinov2/hub/depthers.py b/models/dsp/dinov2/dinov2/hub/depthers.py
new file mode 100644
index 0000000000000000000000000000000000000000..f88b7e9a41056594e3b3e66107feee98bffab820
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/hub/depthers.py
@@ -0,0 +1,246 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from enum import Enum
+from functools import partial
+from typing import Optional, Tuple, Union
+
+import torch
+
+from .backbones import _make_dinov2_model
+from .depth import BNHead, DepthEncoderDecoder, DPTHead
+from .utils import _DINOV2_BASE_URL, _make_dinov2_model_name, CenterPadding
+
+
+class Weights(Enum):
+ NYU = "NYU"
+ KITTI = "KITTI"
+
+
+def _get_depth_range(pretrained: bool, weights: Weights = Weights.NYU) -> Tuple[float, float]:
+ if not pretrained: # Default
+ return (0.001, 10.0)
+
+ # Pretrained, set according to the training dataset for the provided weights
+ if weights == Weights.KITTI:
+ return (0.001, 80.0)
+
+ if weights == Weights.NYU:
+ return (0.001, 10.0)
+
+ return (0.001, 10.0)
+
+
+def _make_dinov2_linear_depth_head(
+ *,
+ embed_dim: int,
+ layers: int,
+ min_depth: float,
+ max_depth: float,
+ **kwargs,
+):
+ if layers not in (1, 4):
+ raise AssertionError(f"Unsupported number of layers: {layers}")
+
+ if layers == 1:
+ in_index = [0]
+ else:
+ assert layers == 4
+ in_index = [0, 1, 2, 3]
+
+ return BNHead(
+ classify=True,
+ n_bins=256,
+ bins_strategy="UD",
+ norm_strategy="linear",
+ upsample=4,
+ in_channels=[embed_dim] * len(in_index),
+ in_index=in_index,
+ input_transform="resize_concat",
+ channels=embed_dim * len(in_index) * 2,
+ align_corners=False,
+ min_depth=0.001,
+ max_depth=80,
+ loss_decode=(),
+ )
+
+
+def _make_dinov2_linear_depther(
+ *,
+ arch_name: str = "vit_large",
+ layers: int = 4,
+ pretrained: bool = True,
+ weights: Union[Weights, str] = Weights.NYU,
+ depth_range: Optional[Tuple[float, float]] = None,
+ **kwargs,
+):
+ if layers not in (1, 4):
+ raise AssertionError(f"Unsupported number of layers: {layers}")
+ if isinstance(weights, str):
+ try:
+ weights = Weights[weights]
+ except KeyError:
+ raise AssertionError(f"Unsupported weights: {weights}")
+
+ if depth_range is None:
+ depth_range = _get_depth_range(pretrained, weights)
+ min_depth, max_depth = depth_range
+
+ backbone = _make_dinov2_model(arch_name=arch_name, pretrained=pretrained, **kwargs)
+
+ embed_dim = backbone.embed_dim
+ patch_size = backbone.patch_size
+ model_name = _make_dinov2_model_name(arch_name, patch_size)
+ linear_depth_head = _make_dinov2_linear_depth_head(
+ embed_dim=embed_dim,
+ layers=layers,
+ min_depth=min_depth,
+ max_depth=max_depth,
+ )
+
+ layer_count = {
+ "vit_small": 12,
+ "vit_base": 12,
+ "vit_large": 24,
+ "vit_giant2": 40,
+ }[arch_name]
+
+ if layers == 4:
+ out_index = {
+ "vit_small": [2, 5, 8, 11],
+ "vit_base": [2, 5, 8, 11],
+ "vit_large": [4, 11, 17, 23],
+ "vit_giant2": [9, 19, 29, 39],
+ }[arch_name]
+ else:
+ assert layers == 1
+ out_index = [layer_count - 1]
+
+ model = DepthEncoderDecoder(backbone=backbone, decode_head=linear_depth_head)
+ model.backbone.forward = partial(
+ backbone.get_intermediate_layers,
+ n=out_index,
+ reshape=True,
+ return_class_token=True,
+ norm=False,
+ )
+ model.backbone.register_forward_pre_hook(lambda _, x: CenterPadding(patch_size)(x[0]))
+
+ if pretrained:
+ layers_str = str(layers) if layers == 4 else ""
+ weights_str = weights.value.lower()
+ url = _DINOV2_BASE_URL + f"/{model_name}/{model_name}_{weights_str}_linear{layers_str}_head.pth"
+ checkpoint = torch.hub.load_state_dict_from_url(url, map_location="cpu")
+ if "state_dict" in checkpoint:
+ state_dict = checkpoint["state_dict"]
+ model.load_state_dict(state_dict, strict=False)
+
+ return model
+
+
+def dinov2_vits14_ld(*, layers: int = 4, pretrained: bool = True, weights: Union[Weights, str] = Weights.NYU, **kwargs):
+ return _make_dinov2_linear_depther(
+ arch_name="vit_small", layers=layers, pretrained=pretrained, weights=weights, **kwargs
+ )
+
+
+def dinov2_vitb14_ld(*, layers: int = 4, pretrained: bool = True, weights: Union[Weights, str] = Weights.NYU, **kwargs):
+ return _make_dinov2_linear_depther(
+ arch_name="vit_base", layers=layers, pretrained=pretrained, weights=weights, **kwargs
+ )
+
+
+def dinov2_vitl14_ld(*, layers: int = 4, pretrained: bool = True, weights: Union[Weights, str] = Weights.NYU, **kwargs):
+ return _make_dinov2_linear_depther(
+ arch_name="vit_large", layers=layers, pretrained=pretrained, weights=weights, **kwargs
+ )
+
+
+def dinov2_vitg14_ld(*, layers: int = 4, pretrained: bool = True, weights: Union[Weights, str] = Weights.NYU, **kwargs):
+ return _make_dinov2_linear_depther(
+ arch_name="vit_giant2", layers=layers, ffn_layer="swiglufused", pretrained=pretrained, weights=weights, **kwargs
+ )
+
+
+def _make_dinov2_dpt_depth_head(*, embed_dim: int, min_depth: float, max_depth: float):
+ return DPTHead(
+ in_channels=[embed_dim] * 4,
+ channels=256,
+ embed_dims=embed_dim,
+ post_process_channels=[embed_dim // 2 ** (3 - i) for i in range(4)],
+ readout_type="project",
+ min_depth=min_depth,
+ max_depth=max_depth,
+ loss_decode=(),
+ )
+
+
+def _make_dinov2_dpt_depther(
+ *,
+ arch_name: str = "vit_large",
+ pretrained: bool = True,
+ weights: Union[Weights, str] = Weights.NYU,
+ depth_range: Optional[Tuple[float, float]] = None,
+ **kwargs,
+):
+ if isinstance(weights, str):
+ try:
+ weights = Weights[weights]
+ except KeyError:
+ raise AssertionError(f"Unsupported weights: {weights}")
+
+ if depth_range is None:
+ depth_range = _get_depth_range(pretrained, weights)
+ min_depth, max_depth = depth_range
+
+ backbone = _make_dinov2_model(arch_name=arch_name, pretrained=pretrained, **kwargs)
+
+ model_name = _make_dinov2_model_name(arch_name, backbone.patch_size)
+ dpt_depth_head = _make_dinov2_dpt_depth_head(embed_dim=backbone.embed_dim, min_depth=min_depth, max_depth=max_depth)
+
+ out_index = {
+ "vit_small": [2, 5, 8, 11],
+ "vit_base": [2, 5, 8, 11],
+ "vit_large": [4, 11, 17, 23],
+ "vit_giant2": [9, 19, 29, 39],
+ }[arch_name]
+
+ model = DepthEncoderDecoder(backbone=backbone, decode_head=dpt_depth_head)
+ model.backbone.forward = partial(
+ backbone.get_intermediate_layers,
+ n=out_index,
+ reshape=True,
+ return_class_token=True,
+ norm=False,
+ )
+ model.backbone.register_forward_pre_hook(lambda _, x: CenterPadding(backbone.patch_size)(x[0]))
+
+ if pretrained:
+ weights_str = weights.value.lower()
+ url = _DINOV2_BASE_URL + f"/{model_name}/{model_name}_{weights_str}_dpt_head.pth"
+ checkpoint = torch.hub.load_state_dict_from_url(url, map_location="cpu")
+ if "state_dict" in checkpoint:
+ state_dict = checkpoint["state_dict"]
+ model.load_state_dict(state_dict, strict=False)
+
+ return model
+
+
+def dinov2_vits14_dd(*, pretrained: bool = True, weights: Union[Weights, str] = Weights.NYU, **kwargs):
+ return _make_dinov2_dpt_depther(arch_name="vit_small", pretrained=pretrained, weights=weights, **kwargs)
+
+
+def dinov2_vitb14_dd(*, pretrained: bool = True, weights: Union[Weights, str] = Weights.NYU, **kwargs):
+ return _make_dinov2_dpt_depther(arch_name="vit_base", pretrained=pretrained, weights=weights, **kwargs)
+
+
+def dinov2_vitl14_dd(*, pretrained: bool = True, weights: Union[Weights, str] = Weights.NYU, **kwargs):
+ return _make_dinov2_dpt_depther(arch_name="vit_large", pretrained=pretrained, weights=weights, **kwargs)
+
+
+def dinov2_vitg14_dd(*, pretrained: bool = True, weights: Union[Weights, str] = Weights.NYU, **kwargs):
+ return _make_dinov2_dpt_depther(
+ arch_name="vit_giant2", ffn_layer="swiglufused", pretrained=pretrained, weights=weights, **kwargs
+ )
diff --git a/models/dsp/dinov2/dinov2/hub/dinotxt.py b/models/dsp/dinov2/dinov2/hub/dinotxt.py
new file mode 100644
index 0000000000000000000000000000000000000000..1a5f159cd3b26b2cb77664f5d85a44a93bb32522
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/hub/dinotxt.py
@@ -0,0 +1,77 @@
+import torch
+import math
+
+from .backbones import dinov2_vitl14_reg
+from .utils import _DINOV2_BASE_URL
+
+
+def dinov2_vitl14_reg4_dinotxt_tet1280d20h24l():
+ from .text.dinotxt_model import DinoTxtConfig, DinoTxt
+ from .text.dinov2_wrapper import DINOv2Wrapper
+ from .text.text_transformer import TextTransformer
+
+ dinotxt_config = DinoTxtConfig(
+ embed_dim=2048,
+ vision_model_freeze_backbone=True,
+ vision_model_train_img_size=224,
+ vision_model_use_class_token=True,
+ vision_model_use_patch_tokens=True,
+ vision_model_num_head_blocks=2,
+ vision_model_head_blocks_drop_path=0.3,
+ vision_model_use_linear_projection=False,
+ vision_model_patch_tokens_pooler_type="mean",
+ vision_model_patch_token_layer=1, # which layer to take patch tokens from
+ # 1 - last layer, 2 - second last layer, etc.
+ text_model_freeze_backbone=False,
+ text_model_num_head_blocks=0,
+ text_model_head_blocks_is_causal=False,
+ text_model_head_blocks_drop_prob=0.0,
+ text_model_tokens_pooler_type="argmax",
+ text_model_use_linear_projection=True,
+ init_logit_scale=math.log(1 / 0.07),
+ init_logit_bias=None,
+ freeze_logit_scale=False,
+ )
+ vision_backbone = DINOv2Wrapper(dinov2_vitl14_reg())
+ text_backbone = TextTransformer(
+ context_length=77,
+ vocab_size=49408,
+ dim=1280,
+ num_heads=20,
+ num_layers=24,
+ ffn_ratio=4,
+ is_causal=True,
+ ls_init_value=None,
+ dropout_prob=0.0,
+ )
+ model = DinoTxt(dinotxt_config, vision_backbone, text_backbone)
+ model.init_weights()
+ model.visual_model.backbone = vision_backbone
+ model.eval()
+
+ visual_model_head_state_dict = torch.hub.load_state_dict_from_url(
+ _DINOV2_BASE_URL + "/dinov2_vitl14/dinov2_vitl14_reg4_dinotxt_tet1280d20h24l_vision_head.pth",
+ map_location="cpu",
+ )
+ text_model_state_dict = torch.hub.load_state_dict_from_url(
+ _DINOV2_BASE_URL + "/dinov2_vitl14/dinov2_vitl14_reg4_dinotxt_tet1280d20h24l_text_encoder.pth",
+ map_location="cpu",
+ )
+ model.visual_model.head.load_state_dict(visual_model_head_state_dict, strict=True)
+ model.text_model.load_state_dict(text_model_state_dict, strict=True)
+ return model
+
+
+def get_tokenizer():
+ from .text.tokenizer import Tokenizer
+ import requests
+ from io import BytesIO
+
+ url = _DINOV2_BASE_URL + "/thirdparty/bpe_simple_vocab_16e6.txt.gz"
+ try:
+ response = requests.get(url)
+ response.raise_for_status()
+ file_buf = BytesIO(response.content)
+ return Tokenizer(vocab_path=file_buf)
+ except Exception as e:
+ raise FileNotFoundError(f"Failed to download file from url {url} with error last: {e}")
diff --git a/models/dsp/dinov2/dinov2/hub/text/dinotxt_model.py b/models/dsp/dinov2/dinov2/hub/text/dinotxt_model.py
new file mode 100644
index 0000000000000000000000000000000000000000..38d3ccc0ef887f05151202d7f42437601488d446
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/hub/text/dinotxt_model.py
@@ -0,0 +1,130 @@
+import math
+from dataclasses import dataclass
+from typing import Optional, Tuple
+
+import torch
+import torch.nn.functional as F
+from torch import nn, Tensor
+
+from .vision_tower import VisionTower
+from .text_tower import TextTower
+
+
+@dataclass
+class DinoTxtConfig:
+ embed_dim: int
+ vision_model_freeze_backbone: bool = True
+ vision_model_train_img_size: int = 224
+ vision_model_use_class_token: bool = True
+ vision_model_use_patch_tokens: bool = False
+ vision_model_num_head_blocks: int = 0
+ vision_model_head_blocks_drop_path: float = 0.3
+ vision_model_use_linear_projection: bool = False
+ vision_model_patch_tokens_pooler_type: str = "mean"
+ vision_model_patch_token_layer: int = 1 # which layer to take patch tokens from
+ # 1 - last layer, 2 - second last layer, etc.
+ text_model_freeze_backbone: bool = False
+ text_model_num_head_blocks: int = 0
+ text_model_head_blocks_is_causal: bool = False
+ text_model_head_blocks_drop_prob: float = 0.0
+ text_model_tokens_pooler_type: str = "first"
+ text_model_use_linear_projection: bool = False
+ init_logit_scale: float = math.log(1 / 0.07)
+ init_logit_bias: Optional[float] = None
+ freeze_logit_scale: bool = False
+
+
+class DinoTxt(nn.Module):
+ def __init__(
+ self,
+ model_config: DinoTxtConfig,
+ vision_backbone: nn.Module,
+ text_backbone: nn.Module,
+ ):
+ super().__init__()
+ self.model_config = model_config
+ self.visual_model = VisionTower(
+ vision_backbone,
+ model_config.vision_model_freeze_backbone,
+ model_config.embed_dim,
+ model_config.vision_model_num_head_blocks,
+ model_config.vision_model_head_blocks_drop_path,
+ model_config.vision_model_use_class_token,
+ model_config.vision_model_use_patch_tokens,
+ model_config.vision_model_patch_token_layer,
+ model_config.vision_model_patch_tokens_pooler_type,
+ model_config.vision_model_use_linear_projection,
+ )
+ self.text_model = TextTower(
+ text_backbone,
+ model_config.text_model_freeze_backbone,
+ model_config.embed_dim,
+ model_config.text_model_num_head_blocks,
+ model_config.text_model_head_blocks_is_causal,
+ model_config.text_model_head_blocks_drop_prob,
+ model_config.text_model_tokens_pooler_type,
+ model_config.text_model_use_linear_projection,
+ )
+ self.logit_scale = nn.Parameter(torch.ones(1) * model_config.init_logit_scale)
+ if model_config.freeze_logit_scale:
+ self.logit_scale.requires_grad = False
+
+ def init_weights(self):
+ self.visual_model.init_weights()
+ self.text_model.init_weights()
+
+ def get_visual_class_and_patch_tokens(self, image: Tensor) -> Tuple[Tensor, Tensor]:
+ return self.visual_model.get_class_and_patch_tokens(image)
+
+ def encode_image(
+ self,
+ image: Tensor,
+ normalize: bool = False,
+ ) -> Tensor:
+ """
+ Encode an image into a vector descriptor containing both global and local features.
+
+ Args:
+ image (Tensor): Tensor of shape `(batch_size, rgb, height, width)`, normalized using ImageNet mean and std.
+ normalize (bool, optional): Whether to normalize the output vectors. Default is False.
+ Image features should always be normalized before comparing them with text features:
+ Returns:
+ Tensor: Tensor of shape `(batch_size, embed_dim)` containing the image features.
+ The first half of the vector corresponds to the global features (class token),
+ and the second half corresponds to the pooled patch features.
+ """
+ features = self.visual_model(image)
+ return F.normalize(features, dim=-1) if normalize else features
+
+ def encode_text(self, text: Tensor, normalize: bool = False) -> Tensor:
+ """
+ Encode a text input into a vector descriptor.
+
+ Args:
+ text (Tensor): Tensor of shape `(batch_size, seq_len)` containing token indices.
+ normalize (bool, optional): Whether to normalize the output vectors. Default is False.
+ Text features should be normalized before comparing them with image features:
+ Returns:
+ Tensor: Tensor of shape `(batch_size, embed_dim)` containing the text features.
+ As a consequence of the training procedure, assume that the first half of the tensor corresponds
+ to global image features and the second half to pooled patch features.
+ """
+ features = self.text_model(text)
+ return F.normalize(features, dim=-1) if normalize else features
+
+ def get_logits(self, image: Tensor, text: Tensor) -> Tuple[Tensor, Tensor]:
+ text_features = self.encode_text(text, normalize=True)
+ image_features = self.encode_image(image, normalize=True)
+ image_logits = self.logit_scale.exp() * image_features @ text_features.T
+ text_logits = image_logits.T
+ return image_logits, text_logits
+
+ def forward(
+ self,
+ image: Tensor,
+ text: Tensor,
+ ) -> Tuple[Tensor, Tensor, Tensor]:
+
+ text_features = self.encode_text(text, normalize=True)
+ image_features = self.encode_image(image, normalize=True)
+ return image_features, text_features, self.logit_scale.exp()
diff --git a/models/dsp/dinov2/dinov2/hub/text/dinov2_wrapper.py b/models/dsp/dinov2/dinov2/hub/text/dinov2_wrapper.py
new file mode 100644
index 0000000000000000000000000000000000000000..31db29696c8e0b0884020a2612998c3870dfe164
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/hub/text/dinov2_wrapper.py
@@ -0,0 +1,59 @@
+from typing import Sequence
+
+import torch
+
+
+class DINOv2Wrapper(torch.nn.Module):
+ def __init__(self, model):
+ super().__init__()
+ self.model = model
+ self.embed_dim = model.embed_dim
+ self.num_heads = model.num_heads
+ self.num_register_tokens = model.num_register_tokens
+
+ # Same as the original forward, but assert is_training and rename x_norm_regtokens -> x_storage_tokens
+ def forward(self, img, is_training: bool):
+ assert is_training
+ H, W = img.shape[-2:]
+ P = self.model.patch_size
+ x_dict = self.model(img, is_training=True)
+ x_dict["h"] = h = H // P
+ x_dict["w"] = w = W // P
+ assert x_dict["x_norm_patchtokens"].shape[-2] == h * w
+ return x_dict
+
+ # Same as the original get_intermediate_layers, but allow returining extra tokens (registers)
+ def get_intermediate_layers(
+ self,
+ x: torch.Tensor,
+ n: int | Sequence[int] = 1, # Layers or n last layers to take
+ reshape: bool = False,
+ return_class_token: bool = False,
+ return_register_tokens: bool = False,
+ norm=True,
+ ) -> tuple[torch.Tensor] | tuple[tuple[torch.Tensor, ...], ...]:
+ if self.model.chunked_blocks:
+ outputs = self.model._get_intermediate_layers_chunked(x, n)
+ else:
+ outputs = self.model._get_intermediate_layers_not_chunked(x, n)
+ if norm:
+ outputs = [self.model.norm(out) for out in outputs]
+ class_tokens = [out[:, 0] for out in outputs]
+ register_tokens = [out[:, 1 : 1 + self.model.num_register_tokens] for out in outputs]
+ outputs = [out[:, 1 + self.model.num_register_tokens :] for out in outputs]
+ if reshape:
+ B, _, h, w = x.shape
+ outputs = [
+ out.reshape(B, h // self.model.patch_size, w // self.model.patch_size, -1)
+ .permute(0, 3, 1, 2)
+ .contiguous()
+ for out in outputs
+ ]
+
+ if not return_class_token and not return_register_tokens:
+ return tuple(outputs)
+ if return_class_token and not return_register_tokens:
+ return tuple(zip(outputs, class_tokens))
+ if not return_class_token and return_register_tokens:
+ return tuple(zip(outputs, register_tokens))
+ return tuple(zip(outputs, class_tokens, register_tokens))
diff --git a/models/dsp/dinov2/dinov2/hub/text/text_tower.py b/models/dsp/dinov2/dinov2/hub/text/text_tower.py
new file mode 100644
index 0000000000000000000000000000000000000000..15297722e9bf1fa57ac8c44a40fe016b6c3efe17
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/hub/text/text_tower.py
@@ -0,0 +1,99 @@
+import torch
+from torch import nn, Tensor
+
+from dinov2.layers import (
+ CausalAttentionBlock,
+)
+
+
+class TextHead(nn.Module):
+ def __init__(
+ self,
+ input_dim: int,
+ embed_dim: int,
+ num_heads: int,
+ num_blocks: int,
+ block_drop_prob: float,
+ is_causal: bool,
+ use_linear_projection: bool,
+ ):
+ super().__init__()
+ block_list = [nn.Identity()]
+ self.ln_final = nn.Identity()
+ if num_blocks > 0:
+ block_list = [
+ CausalAttentionBlock(
+ dim=input_dim,
+ num_heads=num_heads,
+ is_causal=is_causal,
+ dropout_prob=block_drop_prob,
+ )
+ for _ in range(num_blocks)
+ ]
+ self.ln_final = nn.LayerNorm(input_dim)
+ self.block_list = nn.ModuleList(block_list)
+ self.num_blocks = num_blocks
+ self.linear_projection = nn.Identity()
+ if input_dim != embed_dim or use_linear_projection:
+ self.linear_projection = nn.Linear(input_dim, embed_dim, bias=False)
+
+ def init_weights(self):
+ if self.num_blocks > 0:
+ for i in range(self.num_blocks):
+ self.block_list[i].init_weights()
+ self.ln_final.reset_parameters()
+ if isinstance(self.linear_projection, nn.Linear):
+ nn.init.normal_(self.linear_projection.weight, std=self.linear_projection.in_features**-0.5)
+
+ def forward(self, text_tokens: Tensor) -> Tensor:
+ for block in self.block_list:
+ text_tokens = block(text_tokens)
+ text_tokens = self.ln_final(text_tokens)
+ return self.linear_projection(text_tokens)
+
+
+class TextTower(nn.Module):
+ def __init__(
+ self,
+ backbone: nn.Module,
+ freeze_backbone: bool,
+ embed_dim: int,
+ num_head_blocks: int,
+ head_blocks_is_causal: bool,
+ head_blocks_block_drop_prob: float,
+ tokens_pooler_type: str,
+ use_linear_projection: bool,
+ ):
+ super().__init__()
+ self.backbone = backbone
+ self.freeze_backbone = freeze_backbone
+ backbone_out_dim = backbone.embed_dim
+ self.backbone = backbone
+ self.head = TextHead(
+ backbone_out_dim,
+ embed_dim,
+ self.backbone.num_heads,
+ num_head_blocks,
+ head_blocks_block_drop_prob,
+ head_blocks_is_causal,
+ use_linear_projection,
+ )
+ self.tokens_pooler_type = tokens_pooler_type
+
+ def init_weights(self):
+ self.backbone.init_weights()
+ self.head.init_weights()
+
+ def forward(self, token_indices: Tensor) -> Tensor:
+ text_tokens = self.backbone(token_indices)
+ text_tokens = self.head(text_tokens)
+ if self.tokens_pooler_type == "first":
+ features = text_tokens[:, 0]
+ elif self.tokens_pooler_type == "last":
+ features = text_tokens[:, -1]
+ elif self.tokens_pooler_type == "argmax":
+ assert token_indices is not None
+ features = text_tokens[torch.arange(text_tokens.shape[0]), token_indices.argmax(dim=-1)]
+ else:
+ raise ValueError(f"Unknown text tokens pooler type: {self.pooler_type}")
+ return features
diff --git a/models/dsp/dinov2/dinov2/hub/text/text_transformer.py b/models/dsp/dinov2/dinov2/hub/text/text_transformer.py
new file mode 100644
index 0000000000000000000000000000000000000000..be838de7a8a5b5421b88da5394811330414fc100
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/hub/text/text_transformer.py
@@ -0,0 +1,67 @@
+from typing import Callable, Optional, Tuple
+
+import torch
+from torch import nn, Tensor
+
+
+from dinov2.layers import CausalAttentionBlock
+
+
+class TextTransformer(nn.Module):
+ def __init__(
+ self,
+ context_length: int,
+ vocab_size: int,
+ dim: int,
+ num_heads: int,
+ num_layers: int,
+ ffn_ratio: float,
+ is_causal: bool,
+ ls_init_value: Optional[float] = None,
+ act_layer: Callable = nn.GELU,
+ norm_layer: Callable = nn.LayerNorm,
+ dropout_prob: float = 0.0,
+ ):
+ super().__init__()
+ self.vocab_size = vocab_size
+ self.embed_dim = dim
+ self.num_heads = num_heads
+
+ self.token_embedding = nn.Embedding(vocab_size, dim)
+ self.positional_embedding = nn.Parameter(torch.empty(context_length, dim))
+ self.dropout = nn.Dropout(dropout_prob)
+ self.num_layers = num_layers
+ block_list = [
+ CausalAttentionBlock(
+ dim=dim,
+ num_heads=num_heads,
+ ffn_ratio=ffn_ratio,
+ ls_init_value=ls_init_value,
+ is_causal=is_causal,
+ act_layer=act_layer,
+ norm_layer=norm_layer,
+ dropout_prob=dropout_prob,
+ )
+ for _ in range(num_layers)
+ ]
+ self.blocks = nn.ModuleList(block_list)
+ self.ln_final = norm_layer(dim)
+
+ def init_weights(self):
+ nn.init.normal_(self.token_embedding.weight, std=0.02)
+ nn.init.normal_(self.positional_embedding, std=0.01)
+ init_attn_std = self.embed_dim**-0.5
+ init_proj_std = (self.embed_dim**-0.5) * ((2 * self.num_layers) ** -0.5)
+ init_fc_std = (2 * self.embed_dim) ** -0.5
+ for block in self.blocks:
+ block.init_weights(init_attn_std, init_proj_std, init_fc_std)
+ self.ln_final.reset_parameters()
+
+ def forward(self, token_indices: Tensor) -> Tuple[torch.Tensor, torch.Tensor]:
+ _, N = token_indices.size()
+ x = self.token_embedding(token_indices) + self.positional_embedding[:N]
+ x = self.dropout(x)
+ for block in self.blocks:
+ x = block(x)
+ x = self.ln_final(x)
+ return x
diff --git a/models/dsp/dinov2/dinov2/hub/text/tokenizer.py b/models/dsp/dinov2/dinov2/hub/text/tokenizer.py
new file mode 100644
index 0000000000000000000000000000000000000000..5a75026fefafa389e46b9b7c9a7ad2c32fca18a2
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/hub/text/tokenizer.py
@@ -0,0 +1,40 @@
+import torch
+from typing import List, Union
+
+
+from dinov2.thirdparty.CLIP.clip.simple_tokenizer import SimpleTokenizer
+
+
+class Tokenizer(SimpleTokenizer):
+ def __init__(self, vocab_path: str):
+ SimpleTokenizer.__init__(self, bpe_path=vocab_path)
+
+ def tokenize(self, texts: Union[str, List[str]], context_length: int = 77) -> torch.LongTensor:
+ """
+ Returns the tokenized representation of given input string(s)
+
+ Parameters
+ ----------
+ texts : Union[str, List[str]]
+ An input string or a list of input strings to tokenize
+ context_length : int
+ The context length to use; all CLIP models use 77 as the context length
+
+ Returns
+ -------
+ A two-dimensional tensor containing the resulting tokens, shape = [number of input strings, context_length]
+ """
+ if isinstance(texts, str):
+ texts = [texts]
+ sot_token = self.encoder["<|startoftext|>"]
+ eot_token = self.encoder["<|endoftext|>"]
+ all_tokens = [[sot_token] + self.encode(text) + [eot_token] for text in texts]
+ result = torch.zeros(len(all_tokens), context_length, dtype=torch.long)
+
+ for i, tokens in enumerate(all_tokens):
+ if len(tokens) > context_length:
+ tokens = tokens[:context_length] # Truncate
+ tokens[-1] = eot_token
+ result[i, : len(tokens)] = torch.tensor(tokens)
+
+ return result
diff --git a/models/dsp/dinov2/dinov2/hub/text/vision_tower.py b/models/dsp/dinov2/dinov2/hub/text/vision_tower.py
new file mode 100644
index 0000000000000000000000000000000000000000..27fc310c8a5f6ee15f6653a6b7bbe5c28f9c0af0
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/hub/text/vision_tower.py
@@ -0,0 +1,174 @@
+from functools import partial
+from typing import Callable, Tuple
+
+import torch
+from torch import nn, Tensor
+
+from dinov2.layers import (
+ LayerScale,
+ NestedTensorBlock as AttentionBlock,
+ SwiGLUFFNAligned as SwiGLUFFN,
+)
+
+
+def init_weights_vit_timm(module: nn.Module, name: str = ""):
+ """ViT weight initialization, original timm impl (for reproducibility)"""
+ if isinstance(module, nn.Linear):
+ nn.init.trunc_normal_(module.weight, std=0.02)
+ if module.bias is not None:
+ nn.init.zeros_(module.bias)
+ if isinstance(module, nn.LayerNorm):
+ module.reset_parameters()
+ if isinstance(module, LayerScale):
+ module.reset_parameters()
+ if isinstance(module, nn.Conv2d):
+ module.reset_parameters()
+
+
+def named_apply(fn: Callable, module: nn.Module, name="", depth_first=True, include_root=False) -> nn.Module:
+ if not depth_first and include_root:
+ fn(module=module, name=name)
+ for child_name, child_module in module.named_children():
+ child_name = ".".join((name, child_name)) if name else child_name
+ named_apply(
+ fn=fn,
+ module=child_module,
+ name=child_name,
+ depth_first=depth_first,
+ include_root=True,
+ )
+ if depth_first and include_root:
+ fn(module=module, name=name)
+ return module
+
+
+class VisionHead(nn.Module):
+ def __init__(
+ self,
+ input_dim: int,
+ embed_dim: int,
+ num_heads: int,
+ num_blocks: int,
+ blocks_drop_path: float,
+ use_class_token: bool,
+ use_patch_tokens: bool,
+ use_linear_projection: bool,
+ ):
+ super().__init__()
+ block_list = [nn.Identity()]
+ self.ln_final = nn.Identity()
+ if num_blocks > 0:
+ block_list = [
+ AttentionBlock(
+ input_dim,
+ num_heads,
+ ffn_layer=partial(SwiGLUFFN, align_to=64),
+ init_values=1e-5,
+ drop_path=blocks_drop_path,
+ )
+ for _ in range(num_blocks)
+ ]
+ self.ln_final = nn.LayerNorm(input_dim)
+ self.block_list = nn.ModuleList(block_list)
+ self.num_blocks = num_blocks
+ multiplier = 2 if use_class_token and use_patch_tokens else 1
+ self.linear_projection = nn.Identity()
+ if multiplier * input_dim != embed_dim or use_linear_projection:
+ assert embed_dim % multiplier == 0, f"Expects {embed_dim} to be divisible by {multiplier}"
+ self.linear_projection = nn.Linear(input_dim, embed_dim // multiplier, bias=False)
+
+ def init_weights(self):
+ if self.num_blocks > 0:
+ for i in range(self.num_blocks):
+ block = self.block_list[i]
+ named_apply(init_weights_vit_timm, block)
+ self.ln_final.reset_parameters()
+ if isinstance(self.linear_projection, nn.Linear):
+ nn.init.normal_(self.linear_projection.weight, std=self.linear_projection.in_features**-0.5)
+
+ def forward(self, image_tokens: Tensor) -> Tensor:
+ for block in self.block_list:
+ image_tokens = block(image_tokens)
+ image_tokens = self.ln_final(image_tokens)
+ return self.linear_projection(image_tokens)
+
+
+class VisionTower(nn.Module):
+ def __init__(
+ self,
+ backbone: nn.Module,
+ freeze_backbone: bool,
+ embed_dim: int,
+ num_head_blocks: int,
+ head_blocks_block_drop_path: float,
+ use_class_token: bool,
+ use_patch_tokens: bool,
+ patch_token_layer: int,
+ patch_tokens_pooler_type: str,
+ use_linear_projection: bool,
+ ):
+ super().__init__()
+ self.backbone = backbone
+ self.freeze_backbone = freeze_backbone
+ self.use_class_token = use_class_token
+ self.use_patch_tokens = use_patch_tokens
+ self.patch_token_layer = patch_token_layer
+ self.patch_tokens_pooler_type = patch_tokens_pooler_type
+ self.num_register_tokens = 0
+ if hasattr(self.backbone, "num_register_tokens"):
+ self.num_register_tokens = self.backbone.num_register_tokens
+ elif hasattr(self.backbone, "n_storage_tokens"):
+ self.num_register_tokens = self.backbone.n_storage_tokens
+ backbone_out_dim = self.backbone.embed_dim
+ self.head = VisionHead(
+ backbone_out_dim,
+ embed_dim,
+ self.backbone.num_heads,
+ num_head_blocks,
+ head_blocks_block_drop_path,
+ use_class_token,
+ use_patch_tokens,
+ use_linear_projection,
+ )
+
+ def init_weights(self):
+ if not self.freeze_backbone:
+ self.backbone.init_weights()
+ self.head.init_weights()
+
+ def get_backbone_features(self, images: Tensor) -> Tuple[Tensor, Tensor, Tensor]:
+ tokens = self.backbone.get_intermediate_layers(
+ images,
+ n=self.patch_token_layer,
+ return_class_token=True,
+ return_register_tokens=True,
+ )
+ class_token = tokens[-1][1]
+ patch_tokens = tokens[0][0]
+ register_tokens = tokens[0][2]
+ return class_token, patch_tokens, register_tokens
+
+ def get_class_and_patch_tokens(self, images: Tensor) -> Tuple[Tensor, Tensor]:
+ class_token, patch_tokens, register_tokens = self.get_backbone_features(images)
+ image_tokens = self.head(torch.cat([class_token.unsqueeze(1), register_tokens, patch_tokens], dim=1))
+ class_token, patch_tokens = image_tokens[:, 0], image_tokens[:, self.num_register_tokens + 1 :]
+ return class_token, patch_tokens
+
+ def forward(self, images: Tensor) -> Tensor:
+ class_token, patch_tokens = self.get_class_and_patch_tokens(images)
+ features = []
+ if self.use_class_token:
+ features.append(class_token)
+ if self.use_patch_tokens:
+ if self.patch_tokens_pooler_type == "mean":
+ features.append(torch.mean(patch_tokens, dim=1))
+ elif self.patch_tokens_pooler_type == "max":
+ features.append(torch.max(patch_tokens, dim=1).values)
+ elif self.patch_tokens_pooler_type == "gem":
+ power = 3
+ eps = 1e-6
+ patch_tokens_power = patch_tokens.clamp(min=eps).pow(power)
+ features.append(torch.mean(patch_tokens_power, dim=1).pow(1 / power))
+ else:
+ raise ValueError(f"Unknown patch tokens pooler type: {self.patch_tokens_pooler_type}")
+ return torch.cat(features, dim=-1)
diff --git a/models/dsp/dinov2/dinov2/hub/utils.py b/models/dsp/dinov2/dinov2/hub/utils.py
new file mode 100644
index 0000000000000000000000000000000000000000..9c6641404093652d5a2f19b4cf283d976ec39e64
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/hub/utils.py
@@ -0,0 +1,39 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import itertools
+import math
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+
+
+_DINOV2_BASE_URL = "https://dl.fbaipublicfiles.com/dinov2"
+
+
+def _make_dinov2_model_name(arch_name: str, patch_size: int, num_register_tokens: int = 0) -> str:
+ compact_arch_name = arch_name.replace("_", "")[:4]
+ registers_suffix = f"_reg{num_register_tokens}" if num_register_tokens else ""
+ return f"dinov2_{compact_arch_name}{patch_size}{registers_suffix}"
+
+
+class CenterPadding(nn.Module):
+ def __init__(self, multiple):
+ super().__init__()
+ self.multiple = multiple
+
+ def _get_pad(self, size):
+ new_size = math.ceil(size / self.multiple) * self.multiple
+ pad_size = new_size - size
+ pad_size_left = pad_size // 2
+ pad_size_right = pad_size - pad_size_left
+ return pad_size_left, pad_size_right
+
+ @torch.inference_mode()
+ def forward(self, x):
+ pads = list(itertools.chain.from_iterable(self._get_pad(m) for m in x.shape[:1:-1]))
+ output = F.pad(x, pads)
+ return output
diff --git a/models/dsp/dinov2/dinov2/layers/__init__.py b/models/dsp/dinov2/dinov2/layers/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..d640d145fbfb8993d6e7ceec4eb22b5b0a3e62fa
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/layers/__init__.py
@@ -0,0 +1,12 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .dino_head import DINOHead
+from .layer_scale import LayerScale
+from .mlp import Mlp
+from .patch_embed import PatchEmbed
+from .swiglu_ffn import SwiGLUFFN, SwiGLUFFNFused, SwiGLUFFNAligned
+from .block import NestedTensorBlock, CausalAttentionBlock
+from .attention import Attention, MemEffAttention
diff --git a/models/dsp/dinov2/dinov2/layers/attention.py b/models/dsp/dinov2/dinov2/layers/attention.py
new file mode 100644
index 0000000000000000000000000000000000000000..f1d3dabf14ffbd5eb68c3c16edc56d0da19b1265
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/layers/attention.py
@@ -0,0 +1,99 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+# References:
+# https://github.com/facebookresearch/dino/blob/master/vision_transformer.py
+# https://github.com/rwightman/pytorch-image-models/tree/master/timm/models/vision_transformer.py
+
+import logging
+import os
+import warnings
+
+import torch
+from torch import nn, Tensor
+
+
+logger = logging.getLogger("dinov2")
+
+
+XFORMERS_ENABLED = os.environ.get("XFORMERS_DISABLED") is None
+try:
+ if XFORMERS_ENABLED:
+ from xformers.ops import memory_efficient_attention, unbind
+
+ XFORMERS_AVAILABLE = True
+ warnings.warn("xFormers is available (Attention)")
+ else:
+ warnings.warn("xFormers is disabled (Attention)")
+ raise ImportError
+except ImportError:
+ XFORMERS_AVAILABLE = False
+ warnings.warn("xFormers is not available (Attention)")
+
+
+class Attention(nn.Module):
+ def __init__(
+ self,
+ dim: int,
+ num_heads: int = 8,
+ qkv_bias: bool = False,
+ proj_bias: bool = True,
+ attn_drop: float = 0.0,
+ proj_drop: float = 0.0,
+ ) -> None:
+ super().__init__()
+ self.dim = dim
+ self.num_heads = num_heads
+ head_dim = dim // num_heads
+ self.scale = head_dim**-0.5
+
+ self.qkv = nn.Linear(dim, dim * 3, bias=qkv_bias)
+ self.attn_drop = attn_drop
+ self.proj = nn.Linear(dim, dim, bias=proj_bias)
+ self.proj_drop = nn.Dropout(proj_drop)
+
+ def init_weights(
+ self, init_attn_std: float | None = None, init_proj_std: float | None = None, factor: float = 1.0
+ ) -> None:
+ init_attn_std = init_attn_std or (self.dim**-0.5)
+ init_proj_std = init_proj_std or init_attn_std * factor
+ nn.init.normal_(self.qkv.weight, std=init_attn_std)
+ nn.init.normal_(self.proj.weight, std=init_proj_std)
+ if self.qkv.bias is not None:
+ nn.init.zeros_(self.qkv.bias)
+ if self.proj.bias is not None:
+ nn.init.zeros_(self.proj.bias)
+
+ def forward(self, x: Tensor, is_causal: bool = False) -> Tensor:
+ B, N, C = x.shape
+ qkv = self.qkv(x).reshape(B, N, 3, self.num_heads, C // self.num_heads)
+ q, k, v = torch.unbind(qkv, 2)
+ q, k, v = [t.transpose(1, 2) for t in [q, k, v]]
+ x = nn.functional.scaled_dot_product_attention(
+ q, k, v, attn_mask=None, dropout_p=self.attn_drop if self.training else 0, is_causal=is_causal
+ )
+ x = x.transpose(1, 2).contiguous().view(B, N, C)
+ x = self.proj_drop(self.proj(x))
+ return x
+
+
+class MemEffAttention(Attention):
+ def forward(self, x: Tensor, attn_bias=None) -> Tensor:
+ if not XFORMERS_AVAILABLE:
+ if attn_bias is not None:
+ raise AssertionError("xFormers is required for using nested tensors")
+ return super().forward(x)
+
+ B, N, C = x.shape
+ qkv = self.qkv(x).reshape(B, N, 3, self.num_heads, C // self.num_heads)
+
+ q, k, v = unbind(qkv, 2)
+
+ x = memory_efficient_attention(q, k, v, attn_bias=attn_bias)
+ x = x.reshape([B, N, C])
+
+ x = self.proj(x)
+ x = self.proj_drop(x)
+ return x
diff --git a/models/dsp/dinov2/dinov2/layers/block.py b/models/dsp/dinov2/dinov2/layers/block.py
new file mode 100644
index 0000000000000000000000000000000000000000..7e83b71ccb428ca099d2d1d49933dc837faeecfa
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/layers/block.py
@@ -0,0 +1,316 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+# References:
+# https://github.com/facebookresearch/dino/blob/master/vision_transformer.py
+# https://github.com/rwightman/pytorch-image-models/tree/master/timm/layers/patch_embed.py
+
+import logging
+import os
+from typing import Callable, List, Any, Tuple, Dict, Optional
+import warnings
+
+import torch
+from torch import nn, Tensor
+
+from .attention import Attention, MemEffAttention
+from .drop_path import DropPath
+from .layer_scale import LayerScale
+from .mlp import Mlp
+
+
+logger = logging.getLogger("dinov2")
+
+
+XFORMERS_ENABLED = os.environ.get("XFORMERS_DISABLED") is None
+try:
+ if XFORMERS_ENABLED:
+ from xformers.ops import fmha, scaled_index_add, index_select_cat
+
+ XFORMERS_AVAILABLE = True
+ warnings.warn("xFormers is available (Block)")
+ else:
+ warnings.warn("xFormers is disabled (Block)")
+ raise ImportError
+except ImportError:
+ XFORMERS_AVAILABLE = False
+
+ warnings.warn("xFormers is not available (Block)")
+
+
+class Block(nn.Module):
+ def __init__(
+ self,
+ dim: int,
+ num_heads: int,
+ mlp_ratio: float = 4.0,
+ qkv_bias: bool = False,
+ proj_bias: bool = True,
+ ffn_bias: bool = True,
+ drop: float = 0.0,
+ attn_drop: float = 0.0,
+ init_values=None,
+ drop_path: float = 0.0,
+ act_layer: Callable[..., nn.Module] = nn.GELU,
+ norm_layer: Callable[..., nn.Module] = nn.LayerNorm,
+ attn_class: Callable[..., nn.Module] = Attention,
+ ffn_layer: Callable[..., nn.Module] = Mlp,
+ ) -> None:
+ super().__init__()
+ # print(f"biases: qkv: {qkv_bias}, proj: {proj_bias}, ffn: {ffn_bias}")
+ self.norm1 = norm_layer(dim)
+ self.attn = attn_class(
+ dim,
+ num_heads=num_heads,
+ qkv_bias=qkv_bias,
+ proj_bias=proj_bias,
+ attn_drop=attn_drop,
+ proj_drop=drop,
+ )
+ self.ls1 = LayerScale(dim, init_values=init_values) if init_values else nn.Identity()
+ self.drop_path1 = DropPath(drop_path) if drop_path > 0.0 else nn.Identity()
+
+ self.norm2 = norm_layer(dim)
+ mlp_hidden_dim = int(dim * mlp_ratio)
+ self.mlp = ffn_layer(
+ in_features=dim,
+ hidden_features=mlp_hidden_dim,
+ act_layer=act_layer,
+ drop=drop,
+ bias=ffn_bias,
+ )
+ self.ls2 = LayerScale(dim, init_values=init_values) if init_values else nn.Identity()
+ self.drop_path2 = DropPath(drop_path) if drop_path > 0.0 else nn.Identity()
+
+ self.sample_drop_ratio = drop_path
+
+ def forward(self, x: Tensor) -> Tensor:
+ def attn_residual_func(x: Tensor) -> Tensor:
+ return self.ls1(self.attn(self.norm1(x)))
+
+ def ffn_residual_func(x: Tensor) -> Tensor:
+ return self.ls2(self.mlp(self.norm2(x)))
+
+ if self.training and self.sample_drop_ratio > 0.1:
+ # the overhead is compensated only for a drop path rate larger than 0.1
+ x = drop_add_residual_stochastic_depth(
+ x,
+ residual_func=attn_residual_func,
+ sample_drop_ratio=self.sample_drop_ratio,
+ )
+ x = drop_add_residual_stochastic_depth(
+ x,
+ residual_func=ffn_residual_func,
+ sample_drop_ratio=self.sample_drop_ratio,
+ )
+ elif self.training and self.sample_drop_ratio > 0.0:
+ x = x + self.drop_path1(attn_residual_func(x))
+ x = x + self.drop_path1(ffn_residual_func(x)) # FIXME: drop_path2
+ else:
+ x = x + attn_residual_func(x)
+ x = x + ffn_residual_func(x)
+ return x
+
+
+class CausalAttentionBlock(nn.Module):
+ def __init__(
+ self,
+ dim: int,
+ num_heads: int,
+ ffn_ratio: float = 4.0,
+ ls_init_value: Optional[float] = None,
+ is_causal: bool = True,
+ act_layer: Callable = nn.GELU,
+ norm_layer: Callable = nn.LayerNorm,
+ dropout_prob: float = 0.0,
+ ):
+ super().__init__()
+
+ self.dim = dim
+ self.is_causal = is_causal
+ self.ls1 = LayerScale(dim, init_values=ls_init_value) if ls_init_value else nn.Identity()
+ self.attention_norm = norm_layer(dim)
+ self.attention = Attention(dim, num_heads, attn_drop=dropout_prob, proj_drop=dropout_prob)
+
+ self.ffn_norm = norm_layer(dim)
+ ffn_hidden_dim = int(dim * ffn_ratio)
+ self.feed_forward = Mlp(
+ in_features=dim,
+ hidden_features=ffn_hidden_dim,
+ drop=dropout_prob,
+ act_layer=act_layer,
+ )
+
+ self.ls2 = LayerScale(dim, init_values=ls_init_value) if ls_init_value else nn.Identity()
+
+ def init_weights(
+ self,
+ init_attn_std: float | None = None,
+ init_proj_std: float | None = None,
+ init_fc_std: float | None = None,
+ factor: float = 1.0,
+ ) -> None:
+ init_attn_std = init_attn_std or (self.dim**-0.5)
+ init_proj_std = init_proj_std or init_attn_std * factor
+ init_fc_std = init_fc_std or (2 * self.dim) ** -0.5
+ self.attention.init_weights(init_attn_std, init_proj_std)
+ self.attention_norm.reset_parameters()
+ nn.init.normal_(self.feed_forward.fc1.weight, std=init_fc_std)
+ nn.init.normal_(self.feed_forward.fc2.weight, std=init_proj_std)
+ self.ffn_norm.reset_parameters()
+
+ def forward(
+ self,
+ x: torch.Tensor,
+ ):
+ x_attn = x + self.ls1(self.attention(self.attention_norm(x), self.is_causal))
+ x_ffn = x_attn + self.ls2(self.feed_forward(self.ffn_norm(x_attn)))
+ return x_ffn
+
+
+def drop_add_residual_stochastic_depth(
+ x: Tensor,
+ residual_func: Callable[[Tensor], Tensor],
+ sample_drop_ratio: float = 0.0,
+) -> Tensor:
+ # 1) extract subset using permutation
+ b, n, d = x.shape
+ sample_subset_size = max(int(b * (1 - sample_drop_ratio)), 1)
+ brange = (torch.randperm(b, device=x.device))[:sample_subset_size]
+ x_subset = x[brange]
+
+ # 2) apply residual_func to get residual
+ residual = residual_func(x_subset)
+
+ x_flat = x.flatten(1)
+ residual = residual.flatten(1)
+
+ residual_scale_factor = b / sample_subset_size
+
+ # 3) add the residual
+ x_plus_residual = torch.index_add(x_flat, 0, brange, residual.to(dtype=x.dtype), alpha=residual_scale_factor)
+ return x_plus_residual.view_as(x)
+
+
+def get_branges_scales(x, sample_drop_ratio=0.0):
+ b, n, d = x.shape
+ sample_subset_size = max(int(b * (1 - sample_drop_ratio)), 1)
+ brange = (torch.randperm(b, device=x.device))[:sample_subset_size]
+ residual_scale_factor = b / sample_subset_size
+ return brange, residual_scale_factor
+
+
+def add_residual(x, brange, residual, residual_scale_factor, scaling_vector=None):
+ if scaling_vector is None:
+ x_flat = x.flatten(1)
+ residual = residual.flatten(1)
+ x_plus_residual = torch.index_add(x_flat, 0, brange, residual.to(dtype=x.dtype), alpha=residual_scale_factor)
+ else:
+ x_plus_residual = scaled_index_add(
+ x, brange, residual.to(dtype=x.dtype), scaling=scaling_vector, alpha=residual_scale_factor
+ )
+ return x_plus_residual
+
+
+attn_bias_cache: Dict[Tuple, Any] = {}
+
+
+def get_attn_bias_and_cat(x_list, branges=None):
+ """
+ this will perform the index select, cat the tensors, and provide the attn_bias from cache
+ """
+ batch_sizes = [b.shape[0] for b in branges] if branges is not None else [x.shape[0] for x in x_list]
+ all_shapes = tuple((b, x.shape[1]) for b, x in zip(batch_sizes, x_list))
+ if all_shapes not in attn_bias_cache.keys():
+ seqlens = []
+ for b, x in zip(batch_sizes, x_list):
+ for _ in range(b):
+ seqlens.append(x.shape[1])
+ attn_bias = fmha.BlockDiagonalMask.from_seqlens(seqlens)
+ attn_bias._batch_sizes = batch_sizes
+ attn_bias_cache[all_shapes] = attn_bias
+
+ if branges is not None:
+ cat_tensors = index_select_cat([x.flatten(1) for x in x_list], branges).view(1, -1, x_list[0].shape[-1])
+ else:
+ tensors_bs1 = tuple(x.reshape([1, -1, *x.shape[2:]]) for x in x_list)
+ cat_tensors = torch.cat(tensors_bs1, dim=1)
+
+ return attn_bias_cache[all_shapes], cat_tensors
+
+
+def drop_add_residual_stochastic_depth_list(
+ x_list: List[Tensor],
+ residual_func: Callable[[Tensor, Any], Tensor],
+ sample_drop_ratio: float = 0.0,
+ scaling_vector=None,
+) -> Tensor:
+ # 1) generate random set of indices for dropping samples in the batch
+ branges_scales = [get_branges_scales(x, sample_drop_ratio=sample_drop_ratio) for x in x_list]
+ branges = [s[0] for s in branges_scales]
+ residual_scale_factors = [s[1] for s in branges_scales]
+
+ # 2) get attention bias and index+concat the tensors
+ attn_bias, x_cat = get_attn_bias_and_cat(x_list, branges)
+
+ # 3) apply residual_func to get residual, and split the result
+ residual_list = attn_bias.split(residual_func(x_cat, attn_bias=attn_bias)) # type: ignore
+
+ outputs = []
+ for x, brange, residual, residual_scale_factor in zip(x_list, branges, residual_list, residual_scale_factors):
+ outputs.append(add_residual(x, brange, residual, residual_scale_factor, scaling_vector).view_as(x))
+ return outputs
+
+
+class NestedTensorBlock(Block):
+ def forward_nested(self, x_list: List[Tensor]) -> List[Tensor]:
+ """
+ x_list contains a list of tensors to nest together and run
+ """
+ assert isinstance(self.attn, MemEffAttention)
+
+ if self.training and self.sample_drop_ratio > 0.0:
+
+ def attn_residual_func(x: Tensor, attn_bias=None) -> Tensor:
+ return self.attn(self.norm1(x), attn_bias=attn_bias)
+
+ def ffn_residual_func(x: Tensor, attn_bias=None) -> Tensor:
+ return self.mlp(self.norm2(x))
+
+ x_list = drop_add_residual_stochastic_depth_list(
+ x_list,
+ residual_func=attn_residual_func,
+ sample_drop_ratio=self.sample_drop_ratio,
+ scaling_vector=self.ls1.gamma if isinstance(self.ls1, LayerScale) else None,
+ )
+ x_list = drop_add_residual_stochastic_depth_list(
+ x_list,
+ residual_func=ffn_residual_func,
+ sample_drop_ratio=self.sample_drop_ratio,
+ scaling_vector=self.ls2.gamma if isinstance(self.ls1, LayerScale) else None,
+ )
+ return x_list
+ else:
+
+ def attn_residual_func(x: Tensor, attn_bias=None) -> Tensor:
+ return self.ls1(self.attn(self.norm1(x), attn_bias=attn_bias))
+
+ def ffn_residual_func(x: Tensor, attn_bias=None) -> Tensor:
+ return self.ls2(self.mlp(self.norm2(x)))
+
+ attn_bias, x = get_attn_bias_and_cat(x_list)
+ x = x + attn_residual_func(x, attn_bias=attn_bias)
+ x = x + ffn_residual_func(x)
+ return attn_bias.split(x)
+
+ def forward(self, x_or_x_list):
+ if isinstance(x_or_x_list, Tensor):
+ return super().forward(x_or_x_list)
+ elif isinstance(x_or_x_list, list):
+ if not XFORMERS_AVAILABLE:
+ raise AssertionError("xFormers is required for using nested tensors")
+ return self.forward_nested(x_or_x_list)
+ else:
+ raise AssertionError
diff --git a/models/dsp/dinov2/dinov2/layers/dino_head.py b/models/dsp/dinov2/dinov2/layers/dino_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..0ace8ffd6297a1dd480b19db407b662a6ea0f565
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/layers/dino_head.py
@@ -0,0 +1,58 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import torch
+import torch.nn as nn
+from torch.nn.init import trunc_normal_
+from torch.nn.utils import weight_norm
+
+
+class DINOHead(nn.Module):
+ def __init__(
+ self,
+ in_dim,
+ out_dim,
+ use_bn=False,
+ nlayers=3,
+ hidden_dim=2048,
+ bottleneck_dim=256,
+ mlp_bias=True,
+ ):
+ super().__init__()
+ nlayers = max(nlayers, 1)
+ self.mlp = _build_mlp(nlayers, in_dim, bottleneck_dim, hidden_dim=hidden_dim, use_bn=use_bn, bias=mlp_bias)
+ self.apply(self._init_weights)
+ self.last_layer = weight_norm(nn.Linear(bottleneck_dim, out_dim, bias=False))
+ self.last_layer.weight_g.data.fill_(1)
+
+ def _init_weights(self, m):
+ if isinstance(m, nn.Linear):
+ trunc_normal_(m.weight, std=0.02)
+ if isinstance(m, nn.Linear) and m.bias is not None:
+ nn.init.constant_(m.bias, 0)
+
+ def forward(self, x):
+ x = self.mlp(x)
+ eps = 1e-6 if x.dtype == torch.float16 else 1e-12
+ x = nn.functional.normalize(x, dim=-1, p=2, eps=eps)
+ x = self.last_layer(x)
+ return x
+
+
+def _build_mlp(nlayers, in_dim, bottleneck_dim, hidden_dim=None, use_bn=False, bias=True):
+ if nlayers == 1:
+ return nn.Linear(in_dim, bottleneck_dim, bias=bias)
+ else:
+ layers = [nn.Linear(in_dim, hidden_dim, bias=bias)]
+ if use_bn:
+ layers.append(nn.BatchNorm1d(hidden_dim))
+ layers.append(nn.GELU())
+ for _ in range(nlayers - 2):
+ layers.append(nn.Linear(hidden_dim, hidden_dim, bias=bias))
+ if use_bn:
+ layers.append(nn.BatchNorm1d(hidden_dim))
+ layers.append(nn.GELU())
+ layers.append(nn.Linear(hidden_dim, bottleneck_dim, bias=bias))
+ return nn.Sequential(*layers)
diff --git a/models/dsp/dinov2/dinov2/layers/drop_path.py b/models/dsp/dinov2/dinov2/layers/drop_path.py
new file mode 100644
index 0000000000000000000000000000000000000000..1d640e0b969b8dcba96260243473700b4e5b24b5
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/layers/drop_path.py
@@ -0,0 +1,34 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+# References:
+# https://github.com/facebookresearch/dino/blob/master/vision_transformer.py
+# https://github.com/rwightman/pytorch-image-models/tree/master/timm/layers/drop.py
+
+
+from torch import nn
+
+
+def drop_path(x, drop_prob: float = 0.0, training: bool = False):
+ if drop_prob == 0.0 or not training:
+ return x
+ keep_prob = 1 - drop_prob
+ shape = (x.shape[0],) + (1,) * (x.ndim - 1) # work with diff dim tensors, not just 2D ConvNets
+ random_tensor = x.new_empty(shape).bernoulli_(keep_prob)
+ if keep_prob > 0.0:
+ random_tensor.div_(keep_prob)
+ output = x * random_tensor
+ return output
+
+
+class DropPath(nn.Module):
+ """Drop paths (Stochastic Depth) per sample (when applied in main path of residual blocks)."""
+
+ def __init__(self, drop_prob=None):
+ super(DropPath, self).__init__()
+ self.drop_prob = drop_prob
+
+ def forward(self, x):
+ return drop_path(x, self.drop_prob, self.training)
diff --git a/models/dsp/dinov2/dinov2/layers/layer_scale.py b/models/dsp/dinov2/dinov2/layers/layer_scale.py
new file mode 100644
index 0000000000000000000000000000000000000000..0b38971302b3c8fb3d4c05a5f0912fafe0e80816
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/layers/layer_scale.py
@@ -0,0 +1,34 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+# Modified from: https://github.com/huggingface/pytorch-image-models/blob/main/timm/models/vision_transformer.py#L103-L110
+
+from typing import Optional, Union
+
+import torch
+from torch import Tensor
+from torch import nn
+
+
+class LayerScale(nn.Module):
+ def __init__(
+ self,
+ dim: int,
+ init_values: Union[float, Tensor] = 1e-5,
+ inplace: bool = False,
+ device: Optional[torch.device] = None,
+ dtype: Optional[torch.dtype] = None,
+ ) -> None:
+ super().__init__()
+ self.inplace = inplace
+ self.init_values = init_values
+ self.gamma = nn.Parameter(torch.empty(dim, device=device, dtype=dtype))
+ self.reset_parameters()
+
+ def reset_parameters(self):
+ nn.init.constant_(self.gamma, self.init_values)
+
+ def forward(self, x: Tensor) -> Tensor:
+ return x.mul_(self.gamma) if self.inplace else x * self.gamma
diff --git a/models/dsp/dinov2/dinov2/layers/mlp.py b/models/dsp/dinov2/dinov2/layers/mlp.py
new file mode 100644
index 0000000000000000000000000000000000000000..bbf9432aae9258612caeae910a7bde17999e328e
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/layers/mlp.py
@@ -0,0 +1,40 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+# References:
+# https://github.com/facebookresearch/dino/blob/master/vision_transformer.py
+# https://github.com/rwightman/pytorch-image-models/tree/master/timm/layers/mlp.py
+
+
+from typing import Callable, Optional
+
+from torch import Tensor, nn
+
+
+class Mlp(nn.Module):
+ def __init__(
+ self,
+ in_features: int,
+ hidden_features: Optional[int] = None,
+ out_features: Optional[int] = None,
+ act_layer: Callable[..., nn.Module] = nn.GELU,
+ drop: float = 0.0,
+ bias: bool = True,
+ ) -> None:
+ super().__init__()
+ out_features = out_features or in_features
+ hidden_features = hidden_features or in_features
+ self.fc1 = nn.Linear(in_features, hidden_features, bias=bias)
+ self.act = act_layer()
+ self.fc2 = nn.Linear(hidden_features, out_features, bias=bias)
+ self.drop = nn.Dropout(drop)
+
+ def forward(self, x: Tensor) -> Tensor:
+ x = self.fc1(x)
+ x = self.act(x)
+ x = self.drop(x)
+ x = self.fc2(x)
+ x = self.drop(x)
+ return x
diff --git a/models/dsp/dinov2/dinov2/layers/patch_embed.py b/models/dsp/dinov2/dinov2/layers/patch_embed.py
new file mode 100644
index 0000000000000000000000000000000000000000..8b7c0804784a42cf80c0297d110dcc68cc85b339
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/layers/patch_embed.py
@@ -0,0 +1,88 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+# References:
+# https://github.com/facebookresearch/dino/blob/master/vision_transformer.py
+# https://github.com/rwightman/pytorch-image-models/tree/master/timm/layers/patch_embed.py
+
+from typing import Callable, Optional, Tuple, Union
+
+from torch import Tensor
+import torch.nn as nn
+
+
+def make_2tuple(x):
+ if isinstance(x, tuple):
+ assert len(x) == 2
+ return x
+
+ assert isinstance(x, int)
+ return (x, x)
+
+
+class PatchEmbed(nn.Module):
+ """
+ 2D image to patch embedding: (B,C,H,W) -> (B,N,D)
+
+ Args:
+ img_size: Image size.
+ patch_size: Patch token size.
+ in_chans: Number of input image channels.
+ embed_dim: Number of linear projection output channels.
+ norm_layer: Normalization layer.
+ """
+
+ def __init__(
+ self,
+ img_size: Union[int, Tuple[int, int]] = 224,
+ patch_size: Union[int, Tuple[int, int]] = 16,
+ in_chans: int = 3,
+ embed_dim: int = 768,
+ norm_layer: Optional[Callable] = None,
+ flatten_embedding: bool = True,
+ ) -> None:
+ super().__init__()
+
+ image_HW = make_2tuple(img_size)
+ patch_HW = make_2tuple(patch_size)
+ patch_grid_size = (
+ image_HW[0] // patch_HW[0],
+ image_HW[1] // patch_HW[1],
+ )
+
+ self.img_size = image_HW
+ self.patch_size = patch_HW
+ self.patches_resolution = patch_grid_size
+ self.num_patches = patch_grid_size[0] * patch_grid_size[1]
+
+ self.in_chans = in_chans
+ self.embed_dim = embed_dim
+
+ self.flatten_embedding = flatten_embedding
+
+ self.proj = nn.Conv2d(in_chans, embed_dim, kernel_size=patch_HW, stride=patch_HW)
+ self.norm = norm_layer(embed_dim) if norm_layer else nn.Identity()
+
+ def forward(self, x: Tensor) -> Tensor:
+ _, _, H, W = x.shape
+ patch_H, patch_W = self.patch_size
+
+ assert H % patch_H == 0, f"Input image height {H} is not a multiple of patch height {patch_H}"
+ assert W % patch_W == 0, f"Input image width {W} is not a multiple of patch width: {patch_W}"
+
+ x = self.proj(x) # B C H W
+ H, W = x.size(2), x.size(3)
+ x = x.flatten(2).transpose(1, 2) # B HW C
+ x = self.norm(x)
+ if not self.flatten_embedding:
+ x = x.reshape(-1, H, W, self.embed_dim) # B H W C
+ return x
+
+ def flops(self) -> float:
+ Ho, Wo = self.patches_resolution
+ flops = Ho * Wo * self.embed_dim * self.in_chans * (self.patch_size[0] * self.patch_size[1])
+ if self.norm is not None:
+ flops += Ho * Wo * self.embed_dim
+ return flops
diff --git a/models/dsp/dinov2/dinov2/layers/swiglu_ffn.py b/models/dsp/dinov2/dinov2/layers/swiglu_ffn.py
new file mode 100644
index 0000000000000000000000000000000000000000..340cee356cb4ad7cb3c8bbefa121f39f7c4e5c6f
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/layers/swiglu_ffn.py
@@ -0,0 +1,100 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import os
+from typing import Callable, Optional
+import warnings
+
+from torch import Tensor, nn
+import torch.nn.functional as F
+
+
+class SwiGLUFFN(nn.Module):
+ def __init__(
+ self,
+ in_features: int,
+ hidden_features: Optional[int] = None,
+ out_features: Optional[int] = None,
+ act_layer: Callable[..., nn.Module] = None,
+ drop: float = 0.0,
+ bias: bool = True,
+ ) -> None:
+ super().__init__()
+ out_features = out_features or in_features
+ hidden_features = hidden_features or in_features
+ self.w12 = nn.Linear(in_features, 2 * hidden_features, bias=bias)
+ self.w3 = nn.Linear(hidden_features, out_features, bias=bias)
+
+ def forward(self, x: Tensor) -> Tensor:
+ x12 = self.w12(x)
+ x1, x2 = x12.chunk(2, dim=-1)
+ hidden = F.silu(x1) * x2
+ return self.w3(hidden)
+
+
+XFORMERS_ENABLED = os.environ.get("XFORMERS_DISABLED") is None
+try:
+ if XFORMERS_ENABLED:
+ from xformers.ops import SwiGLU
+
+ XFORMERS_AVAILABLE = True
+ warnings.warn("xFormers is available (SwiGLU)")
+ else:
+ warnings.warn("xFormers is disabled (SwiGLU)")
+ raise ImportError
+except ImportError:
+ SwiGLU = SwiGLUFFN
+ XFORMERS_AVAILABLE = False
+
+ warnings.warn("xFormers is not available (SwiGLU)")
+
+
+class SwiGLUFFNFused(SwiGLU):
+ def __init__(
+ self,
+ in_features: int,
+ hidden_features: Optional[int] = None,
+ out_features: Optional[int] = None,
+ act_layer: Callable[..., nn.Module] = None,
+ drop: float = 0.0,
+ bias: bool = True,
+ ) -> None:
+ out_features = out_features or in_features
+ hidden_features = hidden_features or in_features
+ hidden_features = (int(hidden_features * 2 / 3) + 7) // 8 * 8
+ super().__init__(
+ in_features=in_features,
+ hidden_features=hidden_features,
+ out_features=out_features,
+ bias=bias,
+ )
+
+
+class SwiGLUFFNAligned(nn.Module):
+ def __init__(
+ self,
+ in_features: int,
+ hidden_features: Optional[int] = None,
+ out_features: Optional[int] = None,
+ act_layer: Callable[..., nn.Module] = nn.GELU,
+ drop: float = 0.0,
+ bias: bool = True,
+ align_to: int = 8,
+ device=None,
+ ) -> None:
+ super().__init__()
+ out_features = out_features or in_features
+ hidden_features = hidden_features or in_features
+ d = int(hidden_features * 2 / 3)
+ swiglu_hidden_features = d + (-d % align_to)
+ self.w1 = nn.Linear(in_features, swiglu_hidden_features, bias=bias, device=device)
+ self.w2 = nn.Linear(in_features, swiglu_hidden_features, bias=bias, device=device)
+ self.w3 = nn.Linear(swiglu_hidden_features, out_features, bias=bias, device=device)
+
+ def forward(self, x: Tensor) -> Tensor:
+ x1 = self.w1(x)
+ x2 = self.w2(x)
+ hidden = F.silu(x1) * x2
+ return self.w3(hidden)
diff --git a/models/dsp/dinov2/dinov2/logging/__init__.py b/models/dsp/dinov2/dinov2/logging/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..04a7f02204316d4d1ef38bf6080dae3d66241c25
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/logging/__init__.py
@@ -0,0 +1,102 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import functools
+import logging
+import os
+import sys
+from typing import Optional
+
+import dinov2.distributed as distributed
+from .helpers import MetricLogger, SmoothedValue
+
+
+# So that calling _configure_logger multiple times won't add many handlers
+@functools.lru_cache()
+def _configure_logger(
+ name: Optional[str] = None,
+ *,
+ level: int = logging.DEBUG,
+ output: Optional[str] = None,
+):
+ """
+ Configure a logger.
+
+ Adapted from Detectron2.
+
+ Args:
+ name: The name of the logger to configure.
+ level: The logging level to use.
+ output: A file name or a directory to save log. If None, will not save log file.
+ If ends with ".txt" or ".log", assumed to be a file name.
+ Otherwise, logs will be saved to `output/log.txt`.
+
+ Returns:
+ The configured logger.
+ """
+
+ logger = logging.getLogger(name)
+ logger.setLevel(level)
+ logger.propagate = False
+
+ # Loosely match Google glog format:
+ # [IWEF]yyyymmdd hh:mm:ss.uuuuuu threadid file:line] msg
+ # but use a shorter timestamp and include the logger name:
+ # [IWEF]yyyymmdd hh:mm:ss logger threadid file:line] msg
+ fmt_prefix = "%(levelname).1s%(asctime)s %(process)s %(name)s %(filename)s:%(lineno)s] "
+ fmt_message = "%(message)s"
+ fmt = fmt_prefix + fmt_message
+ datefmt = "%Y%m%d %H:%M:%S"
+ formatter = logging.Formatter(fmt=fmt, datefmt=datefmt)
+
+ # stdout logging for main worker only
+ if distributed.is_main_process():
+ handler = logging.StreamHandler(stream=sys.stdout)
+ handler.setLevel(logging.DEBUG)
+ handler.setFormatter(formatter)
+ logger.addHandler(handler)
+
+ # file logging for all workers
+ if output:
+ if os.path.splitext(output)[-1] in (".txt", ".log"):
+ filename = output
+ else:
+ filename = os.path.join(output, "logs", "log.txt")
+
+ if not distributed.is_main_process():
+ global_rank = distributed.get_global_rank()
+ filename = filename + ".rank{}".format(global_rank)
+
+ os.makedirs(os.path.dirname(filename), exist_ok=True)
+
+ handler = logging.StreamHandler(open(filename, "a"))
+ handler.setLevel(logging.DEBUG)
+ handler.setFormatter(formatter)
+ logger.addHandler(handler)
+
+ return logger
+
+
+def setup_logging(
+ output: Optional[str] = None,
+ *,
+ name: Optional[str] = None,
+ level: int = logging.DEBUG,
+ capture_warnings: bool = True,
+) -> None:
+ """
+ Setup logging.
+
+ Args:
+ output: A file name or a directory to save log files. If None, log
+ files will not be saved. If output ends with ".txt" or ".log", it
+ is assumed to be a file name.
+ Otherwise, logs will be saved to `output/log.txt`.
+ name: The name of the logger to configure, by default the root logger.
+ level: The logging level to use.
+ capture_warnings: Whether warnings should be captured as logs.
+ """
+ logging.captureWarnings(capture_warnings)
+ _configure_logger(name, level=level, output=output)
diff --git a/models/dsp/dinov2/dinov2/logging/helpers.py b/models/dsp/dinov2/dinov2/logging/helpers.py
new file mode 100644
index 0000000000000000000000000000000000000000..c6e70bb15505cbbc4c4732b069ee919bf921a74f
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/logging/helpers.py
@@ -0,0 +1,194 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from collections import defaultdict, deque
+import datetime
+import json
+import logging
+import time
+
+import torch
+
+import dinov2.distributed as distributed
+
+
+logger = logging.getLogger("dinov2")
+
+
+class MetricLogger(object):
+ def __init__(self, delimiter="\t", output_file=None):
+ self.meters = defaultdict(SmoothedValue)
+ self.delimiter = delimiter
+ self.output_file = output_file
+
+ def update(self, **kwargs):
+ for k, v in kwargs.items():
+ if isinstance(v, torch.Tensor):
+ v = v.item()
+ assert isinstance(v, (float, int))
+ self.meters[k].update(v)
+
+ def __getattr__(self, attr):
+ if attr in self.meters:
+ return self.meters[attr]
+ if attr in self.__dict__:
+ return self.__dict__[attr]
+ raise AttributeError("'{}' object has no attribute '{}'".format(type(self).__name__, attr))
+
+ def __str__(self):
+ loss_str = []
+ for name, meter in self.meters.items():
+ loss_str.append("{}: {}".format(name, str(meter)))
+ return self.delimiter.join(loss_str)
+
+ def synchronize_between_processes(self):
+ for meter in self.meters.values():
+ meter.synchronize_between_processes()
+
+ def add_meter(self, name, meter):
+ self.meters[name] = meter
+
+ def dump_in_output_file(self, iteration, iter_time, data_time):
+ if self.output_file is None or not distributed.is_main_process():
+ return
+ dict_to_dump = dict(
+ iteration=iteration,
+ iter_time=iter_time,
+ data_time=data_time,
+ )
+ dict_to_dump.update({k: v.median for k, v in self.meters.items()})
+ with open(self.output_file, "a") as f:
+ f.write(json.dumps(dict_to_dump) + "\n")
+ pass
+
+ def log_every(self, iterable, print_freq, header=None, n_iterations=None, start_iteration=0):
+ i = start_iteration
+ if not header:
+ header = ""
+ start_time = time.time()
+ end = time.time()
+ iter_time = SmoothedValue(fmt="{avg:.6f}")
+ data_time = SmoothedValue(fmt="{avg:.6f}")
+
+ if n_iterations is None:
+ n_iterations = len(iterable)
+
+ space_fmt = ":" + str(len(str(n_iterations))) + "d"
+
+ log_list = [
+ header,
+ "[{0" + space_fmt + "}/{1}]",
+ "eta: {eta}",
+ "{meters}",
+ "time: {time}",
+ "data: {data}",
+ ]
+ if torch.cuda.is_available():
+ log_list += ["max mem: {memory:.0f}"]
+
+ log_msg = self.delimiter.join(log_list)
+ MB = 1024.0 * 1024.0
+ for obj in iterable:
+ data_time.update(time.time() - end)
+ yield obj
+ iter_time.update(time.time() - end)
+ if i % print_freq == 0 or i == n_iterations - 1:
+ self.dump_in_output_file(iteration=i, iter_time=iter_time.avg, data_time=data_time.avg)
+ eta_seconds = iter_time.global_avg * (n_iterations - i)
+ eta_string = str(datetime.timedelta(seconds=int(eta_seconds)))
+ if torch.cuda.is_available():
+ logger.info(
+ log_msg.format(
+ i,
+ n_iterations,
+ eta=eta_string,
+ meters=str(self),
+ time=str(iter_time),
+ data=str(data_time),
+ memory=torch.cuda.max_memory_allocated() / MB,
+ )
+ )
+ else:
+ logger.info(
+ log_msg.format(
+ i,
+ n_iterations,
+ eta=eta_string,
+ meters=str(self),
+ time=str(iter_time),
+ data=str(data_time),
+ )
+ )
+ i += 1
+ end = time.time()
+ if i >= n_iterations:
+ break
+ total_time = time.time() - start_time
+ total_time_str = str(datetime.timedelta(seconds=int(total_time)))
+ logger.info("{} Total time: {} ({:.6f} s / it)".format(header, total_time_str, total_time / n_iterations))
+
+
+class SmoothedValue:
+ """Track a series of values and provide access to smoothed values over a
+ window or the global series average.
+ """
+
+ def __init__(self, window_size=20, fmt=None):
+ if fmt is None:
+ fmt = "{median:.4f} ({global_avg:.4f})"
+ self.deque = deque(maxlen=window_size)
+ self.total = 0.0
+ self.count = 0
+ self.fmt = fmt
+
+ def update(self, value, num=1):
+ self.deque.append(value)
+ self.count += num
+ self.total += value * num
+
+ def synchronize_between_processes(self):
+ """
+ Distributed synchronization of the metric
+ Warning: does not synchronize the deque!
+ """
+ if not distributed.is_enabled():
+ return
+ t = torch.tensor([self.count, self.total], dtype=torch.float64, device="cuda")
+ torch.distributed.barrier()
+ torch.distributed.all_reduce(t)
+ t = t.tolist()
+ self.count = int(t[0])
+ self.total = t[1]
+
+ @property
+ def median(self):
+ d = torch.tensor(list(self.deque))
+ return d.median().item()
+
+ @property
+ def avg(self):
+ d = torch.tensor(list(self.deque), dtype=torch.float32)
+ return d.mean().item()
+
+ @property
+ def global_avg(self):
+ return self.total / self.count
+
+ @property
+ def max(self):
+ return max(self.deque)
+
+ @property
+ def value(self):
+ return self.deque[-1]
+
+ def __str__(self):
+ return self.fmt.format(
+ median=self.median,
+ avg=self.avg,
+ global_avg=self.global_avg,
+ max=self.max,
+ value=self.value,
+ )
diff --git a/models/dsp/dinov2/dinov2/loss/__init__.py b/models/dsp/dinov2/dinov2/loss/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..d6b0115b74edbd74b324c9056a57fade363c58fd
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/loss/__init__.py
@@ -0,0 +1,8 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .dino_clstoken_loss import DINOLoss
+from .ibot_patch_loss import iBOTPatchLoss
+from .koleo_loss import KoLeoLoss
diff --git a/models/dsp/dinov2/dinov2/loss/dino_clstoken_loss.py b/models/dsp/dinov2/dinov2/loss/dino_clstoken_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..c31808e36e6c38ee6dae13ba0443bf1946242117
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/loss/dino_clstoken_loss.py
@@ -0,0 +1,99 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import torch
+import torch.distributed as dist
+import torch.nn.functional as F
+from torch import nn
+
+
+class DINOLoss(nn.Module):
+ def __init__(
+ self,
+ out_dim,
+ student_temp=0.1,
+ center_momentum=0.9,
+ ):
+ super().__init__()
+ self.student_temp = student_temp
+ self.center_momentum = center_momentum
+ self.register_buffer("center", torch.zeros(1, out_dim))
+ self.updated = True
+ self.reduce_handle = None
+ self.len_teacher_output = None
+ self.async_batch_center = None
+
+ @torch.no_grad()
+ def softmax_center_teacher(self, teacher_output, teacher_temp):
+ self.apply_center_update()
+ # teacher centering and sharpening
+ return F.softmax((teacher_output - self.center) / teacher_temp, dim=-1)
+
+ @torch.no_grad()
+ def sinkhorn_knopp_teacher(self, teacher_output, teacher_temp, n_iterations=3):
+ teacher_output = teacher_output.float()
+ world_size = dist.get_world_size() if dist.is_initialized() else 1
+ Q = torch.exp(teacher_output / teacher_temp).t() # Q is K-by-B for consistency with notations from our paper
+ B = Q.shape[1] * world_size # number of samples to assign
+ K = Q.shape[0] # how many prototypes
+
+ # make the matrix sums to 1
+ sum_Q = torch.sum(Q)
+ if dist.is_initialized():
+ dist.all_reduce(sum_Q)
+ Q /= sum_Q
+
+ for it in range(n_iterations):
+ # normalize each row: total weight per prototype must be 1/K
+ sum_of_rows = torch.sum(Q, dim=1, keepdim=True)
+ if dist.is_initialized():
+ dist.all_reduce(sum_of_rows)
+ Q /= sum_of_rows
+ Q /= K
+
+ # normalize each column: total weight per sample must be 1/B
+ Q /= torch.sum(Q, dim=0, keepdim=True)
+ Q /= B
+
+ Q *= B # the columns must sum to 1 so that Q is an assignment
+ return Q.t()
+
+ def forward(self, student_output_list, teacher_out_softmaxed_centered_list):
+ """
+ Cross-entropy between softmax outputs of the teacher and student networks.
+ """
+ # TODO: Use cross_entropy_distribution here
+ total_loss = 0
+ for s in student_output_list:
+ lsm = F.log_softmax(s / self.student_temp, dim=-1)
+ for t in teacher_out_softmaxed_centered_list:
+ loss = torch.sum(t * lsm, dim=-1)
+ total_loss -= loss.mean()
+ return total_loss
+
+ @torch.no_grad()
+ def update_center(self, teacher_output):
+ self.reduce_center_update(teacher_output)
+
+ @torch.no_grad()
+ def reduce_center_update(self, teacher_output):
+ self.updated = False
+ self.len_teacher_output = len(teacher_output)
+ self.async_batch_center = torch.sum(teacher_output, dim=0, keepdim=True)
+ if dist.is_initialized():
+ self.reduce_handle = dist.all_reduce(self.async_batch_center, async_op=True)
+
+ @torch.no_grad()
+ def apply_center_update(self):
+ if self.updated is False:
+ world_size = dist.get_world_size() if dist.is_initialized() else 1
+
+ if self.reduce_handle is not None:
+ self.reduce_handle.wait()
+ _t = self.async_batch_center / (self.len_teacher_output * world_size)
+
+ self.center = self.center * self.center_momentum + _t * (1 - self.center_momentum)
+
+ self.updated = True
diff --git a/models/dsp/dinov2/dinov2/loss/ibot_patch_loss.py b/models/dsp/dinov2/dinov2/loss/ibot_patch_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..6732cda0c311c69f193669ebc950fc8665871442
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/loss/ibot_patch_loss.py
@@ -0,0 +1,151 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import torch
+import torch.distributed as dist
+import torch.nn.functional as F
+from torch import nn
+
+import logging
+
+
+logger = logging.getLogger("dinov2")
+
+
+try:
+ from xformers.ops import cross_entropy
+
+ def lossfunc(t, s, temp):
+ s = s.float()
+ t = t.float()
+ if s.ndim == 2:
+ return -cross_entropy(s.unsqueeze(0), t.unsqueeze(0), temp, bw_inplace=True).squeeze(0)
+ elif s.ndim == 3:
+ return -cross_entropy(s, t, temp, bw_inplace=True)
+
+except ImportError:
+
+ def lossfunc(t, s, temp):
+ return torch.sum(t * F.log_softmax(s / temp, dim=-1), dim=-1)
+
+
+class iBOTPatchLoss(nn.Module):
+ def __init__(self, patch_out_dim, student_temp=0.1, center_momentum=0.9):
+ super().__init__()
+ self.student_temp = student_temp
+ self.center_momentum = center_momentum
+ self.register_buffer("center", torch.zeros(1, 1, patch_out_dim))
+ self.updated = True
+ self.reduce_handle = None
+ self.len_teacher_patch_tokens = None
+ self.async_batch_center = None
+
+ @torch.no_grad()
+ def softmax_center_teacher(self, teacher_patch_tokens, teacher_temp):
+ self.apply_center_update()
+ # teacher centering and sharpening
+ #
+ # WARNING:
+ # as self.center is a float32, everything gets casted to float32 afterwards
+ #
+ # teacher_patch_tokens = teacher_patch_tokens.float()
+ # return F.softmax((teacher_patch_tokens.sub_(self.center.to(teacher_patch_tokens.dtype))).mul_(1 / teacher_temp), dim=-1)
+
+ return F.softmax((teacher_patch_tokens - self.center) / teacher_temp, dim=-1)
+
+ # this is experimental, keep everything in float16 and let's see what happens:
+ # return F.softmax((teacher_patch_tokens.sub_(self.center)) / teacher_temp, dim=-1)
+
+ @torch.no_grad()
+ def sinkhorn_knopp_teacher(self, teacher_output, teacher_temp, n_masked_patches_tensor, n_iterations=3):
+ teacher_output = teacher_output.float()
+ # world_size = dist.get_world_size() if dist.is_initialized() else 1
+ Q = torch.exp(teacher_output / teacher_temp).t() # Q is K-by-B for consistency with notations from our paper
+ # B = Q.shape[1] * world_size # number of samples to assign
+ B = n_masked_patches_tensor
+ dist.all_reduce(B)
+ K = Q.shape[0] # how many prototypes
+
+ # make the matrix sums to 1
+ sum_Q = torch.sum(Q)
+ if dist.is_initialized():
+ dist.all_reduce(sum_Q)
+ Q /= sum_Q
+
+ for it in range(n_iterations):
+ # normalize each row: total weight per prototype must be 1/K
+ sum_of_rows = torch.sum(Q, dim=1, keepdim=True)
+ if dist.is_initialized():
+ dist.all_reduce(sum_of_rows)
+ Q /= sum_of_rows
+ Q /= K
+
+ # normalize each column: total weight per sample must be 1/B
+ Q /= torch.sum(Q, dim=0, keepdim=True)
+ Q /= B
+
+ Q *= B # the columns must sum to 1 so that Q is an assignment
+ return Q.t()
+
+ def forward(self, student_patch_tokens, teacher_patch_tokens, student_masks_flat):
+ """
+ Cross-entropy between softmax outputs of the teacher and student networks.
+ student_patch_tokens: (B, N, D) tensor
+ teacher_patch_tokens: (B, N, D) tensor
+ student_masks_flat: (B, N) tensor
+ """
+ t = teacher_patch_tokens
+ s = student_patch_tokens
+ loss = torch.sum(t * F.log_softmax(s / self.student_temp, dim=-1), dim=-1)
+ loss = torch.sum(loss * student_masks_flat.float(), dim=-1) / student_masks_flat.sum(dim=-1).clamp(min=1.0)
+ return -loss.mean()
+
+ def forward_masked(
+ self,
+ student_patch_tokens_masked,
+ teacher_patch_tokens_masked,
+ student_masks_flat,
+ n_masked_patches=None,
+ masks_weight=None,
+ ):
+ t = teacher_patch_tokens_masked
+ s = student_patch_tokens_masked
+ # loss = torch.sum(t * F.log_softmax(s / self.student_temp, dim=-1), dim=-1)
+ loss = lossfunc(t, s, self.student_temp)
+ if masks_weight is None:
+ masks_weight = (
+ (1 / student_masks_flat.sum(-1).clamp(min=1.0))
+ .unsqueeze(-1)
+ .expand_as(student_masks_flat)[student_masks_flat]
+ )
+ if n_masked_patches is not None:
+ loss = loss[:n_masked_patches]
+ loss = loss * masks_weight
+ return -loss.sum() / student_masks_flat.shape[0]
+
+ @torch.no_grad()
+ def update_center(self, teacher_patch_tokens):
+ self.reduce_center_update(teacher_patch_tokens)
+
+ @torch.no_grad()
+ def reduce_center_update(self, teacher_patch_tokens):
+ self.updated = False
+ self.len_teacher_patch_tokens = len(teacher_patch_tokens)
+ self.async_batch_center = torch.sum(teacher_patch_tokens.mean(1), dim=0, keepdim=True)
+ if dist.is_initialized():
+ self.reduce_handle = dist.all_reduce(self.async_batch_center, async_op=True)
+
+ @torch.no_grad()
+ def apply_center_update(self):
+ if self.updated is False:
+ world_size = dist.get_world_size() if dist.is_initialized() else 1
+
+ if self.reduce_handle is not None:
+ self.reduce_handle.wait()
+ _t = self.async_batch_center / (self.len_teacher_patch_tokens * world_size)
+
+ self.center = self.center * self.center_momentum + _t * (1 - self.center_momentum)
+
+ self.updated = True
diff --git a/models/dsp/dinov2/dinov2/loss/koleo_loss.py b/models/dsp/dinov2/dinov2/loss/koleo_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..b5cbcd91e0fc0b857f477b0910f957f02a6c4335
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/loss/koleo_loss.py
@@ -0,0 +1,48 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import logging
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+
+# import torch.distributed as dist
+
+
+logger = logging.getLogger("dinov2")
+
+
+class KoLeoLoss(nn.Module):
+ """Kozachenko-Leonenko entropic loss regularizer from Sablayrolles et al. - 2018 - Spreading vectors for similarity search"""
+
+ def __init__(self):
+ super().__init__()
+ self.pdist = nn.PairwiseDistance(2, eps=1e-8)
+
+ def pairwise_NNs_inner(self, x):
+ """
+ Pairwise nearest neighbors for L2-normalized vectors.
+ Uses Torch rather than Faiss to remain on GPU.
+ """
+ # parwise dot products (= inverse distance)
+ dots = torch.mm(x, x.t())
+ n = x.shape[0]
+ dots.view(-1)[:: (n + 1)].fill_(-1) # Trick to fill diagonal with -1
+ # max inner prod -> min distance
+ _, I = torch.max(dots, dim=1) # noqa: E741
+ return I
+
+ def forward(self, student_output, eps=1e-8):
+ """
+ Args:
+ student_output (BxD): backbone output of student
+ """
+ with torch.cuda.amp.autocast(enabled=False):
+ student_output = F.normalize(student_output, eps=eps, p=2, dim=-1)
+ I = self.pairwise_NNs_inner(student_output) # noqa: E741
+ distances = self.pdist(student_output, student_output[I]) # BxD, BxD -> B
+ loss = -torch.log(distances + eps).mean()
+ return loss
diff --git a/models/dsp/dinov2/dinov2/models/__init__.py b/models/dsp/dinov2/dinov2/models/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..3fdff20badbd5244bf79f16bf18dd2cb73982265
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/models/__init__.py
@@ -0,0 +1,43 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import logging
+
+from . import vision_transformer as vits
+
+
+logger = logging.getLogger("dinov2")
+
+
+def build_model(args, only_teacher=False, img_size=224):
+ args.arch = args.arch.removesuffix("_memeff")
+ if "vit" in args.arch:
+ vit_kwargs = dict(
+ img_size=img_size,
+ patch_size=args.patch_size,
+ init_values=args.layerscale,
+ ffn_layer=args.ffn_layer,
+ block_chunks=args.block_chunks,
+ qkv_bias=args.qkv_bias,
+ proj_bias=args.proj_bias,
+ ffn_bias=args.ffn_bias,
+ num_register_tokens=args.num_register_tokens,
+ interpolate_offset=args.interpolate_offset,
+ interpolate_antialias=args.interpolate_antialias,
+ )
+ teacher = vits.__dict__[args.arch](**vit_kwargs)
+ if only_teacher:
+ return teacher, teacher.embed_dim
+ student = vits.__dict__[args.arch](
+ **vit_kwargs,
+ drop_path_rate=args.drop_path_rate,
+ drop_path_uniform=args.drop_path_uniform,
+ )
+ embed_dim = student.embed_dim
+ return student, teacher, embed_dim
+
+
+def build_model_from_cfg(cfg, only_teacher=False):
+ return build_model(cfg.student, only_teacher=only_teacher, img_size=cfg.crops.global_crops_size)
diff --git a/models/dsp/dinov2/dinov2/models/vision_transformer.py b/models/dsp/dinov2/dinov2/models/vision_transformer.py
new file mode 100644
index 0000000000000000000000000000000000000000..d6bdf635949761b1d1135b0736224044e7bd5ff1
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/models/vision_transformer.py
@@ -0,0 +1,397 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+# References:
+# https://github.com/facebookresearch/dino/blob/main/vision_transformer.py
+# https://github.com/rwightman/pytorch-image-models/tree/master/timm/models/vision_transformer.py
+
+from functools import partial
+import math
+import logging
+from typing import Sequence, Tuple, Union, Callable
+
+import numpy as np
+import torch
+import torch.nn as nn
+import torch.utils.checkpoint
+from torch.nn.init import trunc_normal_
+
+from ..layers import Mlp, PatchEmbed, SwiGLUFFNFused, MemEffAttention, NestedTensorBlock as Block
+
+
+logger = logging.getLogger("dinov2")
+
+
+def named_apply(fn: Callable, module: nn.Module, name="", depth_first=True, include_root=False) -> nn.Module:
+ if not depth_first and include_root:
+ fn(module=module, name=name)
+ for child_name, child_module in module.named_children():
+ child_name = ".".join((name, child_name)) if name else child_name
+ named_apply(fn=fn, module=child_module, name=child_name, depth_first=depth_first, include_root=True)
+ if depth_first and include_root:
+ fn(module=module, name=name)
+ return module
+
+
+class BlockChunk(nn.ModuleList):
+ def forward(self, x):
+ for b in self:
+ x = b(x)
+ return x
+
+
+class DinoVisionTransformer(nn.Module):
+ def __init__(
+ self,
+ img_size=224,
+ patch_size=16,
+ in_chans=3,
+ embed_dim=768,
+ depth=12,
+ num_heads=12,
+ mlp_ratio=4.0,
+ qkv_bias=True,
+ ffn_bias=True,
+ proj_bias=True,
+ drop_path_rate=0.0,
+ drop_path_uniform=False,
+ init_values=None, # for layerscale: None or 0 => no layerscale
+ embed_layer=PatchEmbed,
+ act_layer=nn.GELU,
+ block_fn=Block,
+ ffn_layer="mlp",
+ block_chunks=1,
+ num_register_tokens=0,
+ interpolate_antialias=False,
+ interpolate_offset=0.1,
+ ):
+ """
+ Args:
+ img_size (int, tuple): input image size
+ patch_size (int, tuple): patch size
+ in_chans (int): number of input channels
+ embed_dim (int): embedding dimension
+ depth (int): depth of transformer
+ num_heads (int): number of attention heads
+ mlp_ratio (int): ratio of mlp hidden dim to embedding dim
+ qkv_bias (bool): enable bias for qkv if True
+ proj_bias (bool): enable bias for proj in attn if True
+ ffn_bias (bool): enable bias for ffn if True
+ drop_path_rate (float): stochastic depth rate
+ drop_path_uniform (bool): apply uniform drop rate across blocks
+ weight_init (str): weight init scheme
+ init_values (float): layer-scale init values
+ embed_layer (nn.Module): patch embedding layer
+ act_layer (nn.Module): MLP activation layer
+ block_fn (nn.Module): transformer block class
+ ffn_layer (str): "mlp", "swiglu", "swiglufused" or "identity"
+ block_chunks: (int) split block sequence into block_chunks units for FSDP wrap
+ num_register_tokens: (int) number of extra cls tokens (so-called "registers")
+ interpolate_antialias: (str) flag to apply anti-aliasing when interpolating positional embeddings
+ interpolate_offset: (float) work-around offset to apply when interpolating positional embeddings
+ """
+ super().__init__()
+ norm_layer = partial(nn.LayerNorm, eps=1e-6)
+
+ self.num_features = self.embed_dim = embed_dim # num_features for consistency with other models
+ self.num_tokens = 1
+ self.n_blocks = depth
+ self.num_heads = num_heads
+ self.patch_size = patch_size
+ self.num_register_tokens = num_register_tokens
+ self.interpolate_antialias = interpolate_antialias
+ self.interpolate_offset = interpolate_offset
+
+ self.patch_embed = embed_layer(img_size=img_size, patch_size=patch_size, in_chans=in_chans, embed_dim=embed_dim)
+ num_patches = self.patch_embed.num_patches
+
+ self.cls_token = nn.Parameter(torch.zeros(1, 1, embed_dim))
+ self.pos_embed = nn.Parameter(torch.zeros(1, num_patches + self.num_tokens, embed_dim))
+ assert num_register_tokens >= 0
+ self.register_tokens = (
+ nn.Parameter(torch.zeros(1, num_register_tokens, embed_dim)) if num_register_tokens else None
+ )
+
+ if drop_path_uniform is True:
+ dpr = [drop_path_rate] * depth
+ else:
+ dpr = np.linspace(0, drop_path_rate, depth).tolist() # stochastic depth decay rule
+
+ if ffn_layer == "mlp":
+ logger.info("using MLP layer as FFN")
+ ffn_layer = Mlp
+ elif ffn_layer == "swiglufused" or ffn_layer == "swiglu":
+ logger.info("using SwiGLU layer as FFN")
+ ffn_layer = SwiGLUFFNFused
+ elif ffn_layer == "identity":
+ logger.info("using Identity layer as FFN")
+
+ def f(*args, **kwargs):
+ return nn.Identity()
+
+ ffn_layer = f
+ else:
+ raise NotImplementedError
+
+ blocks_list = [
+ block_fn(
+ dim=embed_dim,
+ num_heads=num_heads,
+ mlp_ratio=mlp_ratio,
+ qkv_bias=qkv_bias,
+ proj_bias=proj_bias,
+ ffn_bias=ffn_bias,
+ drop_path=dpr[i],
+ norm_layer=norm_layer,
+ act_layer=act_layer,
+ ffn_layer=ffn_layer,
+ init_values=init_values,
+ )
+ for i in range(depth)
+ ]
+ if block_chunks > 0:
+ self.chunked_blocks = True
+ chunked_blocks = []
+ chunksize = depth // block_chunks
+ for i in range(0, depth, chunksize):
+ # this is to keep the block index consistent if we chunk the block list
+ chunked_blocks.append([nn.Identity()] * i + blocks_list[i : i + chunksize])
+ self.blocks = nn.ModuleList([BlockChunk(p) for p in chunked_blocks])
+ else:
+ self.chunked_blocks = False
+ self.blocks = nn.ModuleList(blocks_list)
+
+ self.norm = norm_layer(embed_dim)
+ self.head = nn.Identity()
+
+ self.mask_token = nn.Parameter(torch.zeros(1, embed_dim))
+
+ self.init_weights()
+
+ def init_weights(self):
+ trunc_normal_(self.pos_embed, std=0.02)
+ nn.init.normal_(self.cls_token, std=1e-6)
+ if self.register_tokens is not None:
+ nn.init.normal_(self.register_tokens, std=1e-6)
+ named_apply(init_weights_vit_timm, self)
+
+ def interpolate_pos_encoding(self, x, w, h):
+ previous_dtype = x.dtype
+ npatch = x.shape[1] - 1
+ N = self.pos_embed.shape[1] - 1
+ if npatch == N and w == h:
+ return self.pos_embed
+ pos_embed = self.pos_embed.float()
+ class_pos_embed = pos_embed[:, 0]
+ patch_pos_embed = pos_embed[:, 1:]
+ dim = x.shape[-1]
+ w0 = w // self.patch_size
+ h0 = h // self.patch_size
+ M = int(math.sqrt(N)) # Recover the number of patches in each dimension
+ assert N == M * M
+ kwargs = {}
+ if self.interpolate_offset:
+ # Historical kludge: add a small number to avoid floating point error in the interpolation, see https://github.com/facebookresearch/dino/issues/8
+ # Note: still needed for backward-compatibility, the underlying operators are using both output size and scale factors
+ sx = float(w0 + self.interpolate_offset) / M
+ sy = float(h0 + self.interpolate_offset) / M
+ kwargs["scale_factor"] = (sx, sy)
+ else:
+ # Simply specify an output size instead of a scale factor
+ kwargs["size"] = (w0, h0)
+ patch_pos_embed = nn.functional.interpolate(
+ patch_pos_embed.reshape(1, M, M, dim).permute(0, 3, 1, 2),
+ mode="bicubic",
+ antialias=self.interpolate_antialias,
+ **kwargs,
+ )
+ assert (w0, h0) == patch_pos_embed.shape[-2:]
+ patch_pos_embed = patch_pos_embed.permute(0, 2, 3, 1).view(1, -1, dim)
+ return torch.cat((class_pos_embed.unsqueeze(0), patch_pos_embed), dim=1).to(previous_dtype)
+
+ def prepare_tokens_with_masks(self, x, masks=None):
+ B, nc, w, h = x.shape
+ x = self.patch_embed(x)
+ if masks is not None:
+ x = torch.where(masks.unsqueeze(-1), self.mask_token.to(x.dtype).unsqueeze(0), x)
+
+ x = torch.cat((self.cls_token.expand(x.shape[0], -1, -1), x), dim=1)
+ x = x + self.interpolate_pos_encoding(x, w, h)
+
+ if self.register_tokens is not None:
+ x = torch.cat(
+ (
+ x[:, :1],
+ self.register_tokens.expand(x.shape[0], -1, -1),
+ x[:, 1:],
+ ),
+ dim=1,
+ )
+
+ return x
+
+ def forward_features_list(self, x_list, masks_list):
+ x = [self.prepare_tokens_with_masks(x, masks) for x, masks in zip(x_list, masks_list)]
+ for blk in self.blocks:
+ x = blk(x)
+
+ all_x = x
+ output = []
+ for x, masks in zip(all_x, masks_list):
+ x_norm = self.norm(x)
+ output.append(
+ {
+ "x_norm_clstoken": x_norm[:, 0],
+ "x_norm_regtokens": x_norm[:, 1 : self.num_register_tokens + 1],
+ "x_norm_patchtokens": x_norm[:, self.num_register_tokens + 1 :],
+ "x_prenorm": x,
+ "masks": masks,
+ }
+ )
+ return output
+
+ def forward_features(self, x, masks=None):
+ if isinstance(x, list):
+ return self.forward_features_list(x, masks)
+
+ x = self.prepare_tokens_with_masks(x, masks)
+
+ for blk in self.blocks:
+ x = blk(x)
+
+ x_norm = self.norm(x)
+ return {
+ "x_norm_clstoken": x_norm[:, 0],
+ "x_norm_regtokens": x_norm[:, 1 : self.num_register_tokens + 1],
+ "x_norm_patchtokens": x_norm[:, self.num_register_tokens + 1 :],
+ "x_prenorm": x,
+ "masks": masks,
+ }
+
+ def _get_intermediate_layers_not_chunked(self, x, n=1):
+ x = self.prepare_tokens_with_masks(x)
+ # If n is an int, take the n last blocks. If it's a list, take them
+ output, total_block_len = [], len(self.blocks)
+ blocks_to_take = range(total_block_len - n, total_block_len) if isinstance(n, int) else n
+ for i, blk in enumerate(self.blocks):
+ x = blk(x)
+ if i in blocks_to_take:
+ output.append(x)
+ assert len(output) == len(blocks_to_take), f"only {len(output)} / {len(blocks_to_take)} blocks found"
+ return output
+
+ def _get_intermediate_layers_chunked(self, x, n=1):
+ x = self.prepare_tokens_with_masks(x)
+ output, i, total_block_len = [], 0, len(self.blocks[-1])
+ # If n is an int, take the n last blocks. If it's a list, take them
+ blocks_to_take = range(total_block_len - n, total_block_len) if isinstance(n, int) else n
+ for block_chunk in self.blocks:
+ for blk in block_chunk[i:]: # Passing the nn.Identity()
+ x = blk(x)
+ if i in blocks_to_take:
+ output.append(x)
+ i += 1
+ assert len(output) == len(blocks_to_take), f"only {len(output)} / {len(blocks_to_take)} blocks found"
+ return output
+
+ def get_intermediate_layers(
+ self,
+ x: torch.Tensor,
+ n: Union[int, Sequence] = 1, # Layers or n last layers to take
+ reshape: bool = False,
+ return_class_token: bool = False,
+ norm=True,
+ ) -> Tuple[Union[torch.Tensor, Tuple[torch.Tensor]]]:
+ if self.chunked_blocks:
+ outputs = self._get_intermediate_layers_chunked(x, n)
+ else:
+ outputs = self._get_intermediate_layers_not_chunked(x, n)
+ if norm:
+ outputs = [self.norm(out) for out in outputs]
+ class_tokens = [out[:, 0] for out in outputs]
+ outputs = [out[:, 1 + self.num_register_tokens :] for out in outputs]
+ if reshape:
+ B, _, w, h = x.shape
+ outputs = [
+ out.reshape(B, w // self.patch_size, h // self.patch_size, -1).permute(0, 3, 1, 2).contiguous()
+ for out in outputs
+ ]
+ if return_class_token:
+ return tuple(zip(outputs, class_tokens))
+ return tuple(outputs)
+
+ def forward(self, *args, is_training=False, **kwargs):
+ ret = self.forward_features(*args, **kwargs)
+ if is_training:
+ return ret
+ else:
+ return self.head(ret["x_norm_clstoken"])
+
+
+def init_weights_vit_timm(module: nn.Module, name: str = ""):
+ """ViT weight initialization, original timm impl (for reproducibility)"""
+ if isinstance(module, nn.Linear):
+ trunc_normal_(module.weight, std=0.02)
+ if module.bias is not None:
+ nn.init.zeros_(module.bias)
+
+
+def vit_small(patch_size=16, num_register_tokens=0, **kwargs):
+ model = DinoVisionTransformer(
+ patch_size=patch_size,
+ embed_dim=384,
+ depth=12,
+ num_heads=6,
+ mlp_ratio=4,
+ block_fn=partial(Block, attn_class=MemEffAttention),
+ num_register_tokens=num_register_tokens,
+ **kwargs,
+ )
+ return model
+
+
+def vit_base(patch_size=16, num_register_tokens=0, **kwargs):
+ model = DinoVisionTransformer(
+ patch_size=patch_size,
+ embed_dim=768,
+ depth=12,
+ num_heads=12,
+ mlp_ratio=4,
+ block_fn=partial(Block, attn_class=MemEffAttention),
+ num_register_tokens=num_register_tokens,
+ **kwargs,
+ )
+ return model
+
+
+def vit_large(patch_size=16, num_register_tokens=0, **kwargs):
+ model = DinoVisionTransformer(
+ patch_size=patch_size,
+ embed_dim=1024,
+ depth=24,
+ num_heads=16,
+ mlp_ratio=4,
+ block_fn=partial(Block, attn_class=MemEffAttention),
+ num_register_tokens=num_register_tokens,
+ **kwargs,
+ )
+ return model
+
+
+def vit_giant2(patch_size=16, num_register_tokens=0, **kwargs):
+ """
+ Close to ViT-giant, with embed-dim 1536 and 24 heads => embed-dim per head 64
+ """
+ model = DinoVisionTransformer(
+ patch_size=patch_size,
+ embed_dim=1536,
+ depth=40,
+ num_heads=24,
+ mlp_ratio=4,
+ block_fn=partial(Block, attn_class=MemEffAttention),
+ num_register_tokens=num_register_tokens,
+ **kwargs,
+ )
+ return model
diff --git a/models/dsp/dinov2/dinov2/run/__init__.py b/models/dsp/dinov2/dinov2/run/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..b88da6bf80be92af00b72dfdb0a806fa64a7a2d9
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/run/__init__.py
@@ -0,0 +1,4 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
diff --git a/models/dsp/dinov2/dinov2/run/eval/knn.py b/models/dsp/dinov2/dinov2/run/eval/knn.py
new file mode 100644
index 0000000000000000000000000000000000000000..d11918445cdfe415fe58ac8b3ad0bf29702e3457
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/run/eval/knn.py
@@ -0,0 +1,59 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import logging
+import os
+import sys
+
+from dinov2.eval.knn import get_args_parser as get_knn_args_parser
+from dinov2.logging import setup_logging
+from dinov2.run.submit import get_args_parser, submit_jobs
+
+
+logger = logging.getLogger("dinov2")
+
+
+class Evaluator:
+ def __init__(self, args):
+ self.args = args
+
+ def __call__(self):
+ from dinov2.eval.knn import main as knn_main
+
+ self._setup_args()
+ knn_main(self.args)
+
+ def checkpoint(self):
+ import submitit
+
+ logger.info(f"Requeuing {self.args}")
+ empty = type(self)(self.args)
+ return submitit.helpers.DelayedSubmission(empty)
+
+ def _setup_args(self):
+ import submitit
+
+ job_env = submitit.JobEnvironment()
+ self.args.output_dir = self.args.output_dir.replace("%j", str(job_env.job_id))
+ logger.info(f"Process group: {job_env.num_tasks} tasks, rank: {job_env.global_rank}")
+ logger.info(f"Args: {self.args}")
+
+
+def main():
+ description = "Submitit launcher for DINOv2 k-NN evaluation"
+ knn_args_parser = get_knn_args_parser(add_help=False)
+ parents = [knn_args_parser]
+ args_parser = get_args_parser(description=description, parents=parents)
+ args = args_parser.parse_args()
+
+ setup_logging()
+
+ assert os.path.exists(args.config_file), "Configuration file does not exist!"
+ submit_jobs(Evaluator, args, name="dinov2:knn")
+ return 0
+
+
+if __name__ == "__main__":
+ sys.exit(main())
diff --git a/models/dsp/dinov2/dinov2/run/eval/linear.py b/models/dsp/dinov2/dinov2/run/eval/linear.py
new file mode 100644
index 0000000000000000000000000000000000000000..e1dc3293e88512a5cf885ab775dc08e01aed6724
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/run/eval/linear.py
@@ -0,0 +1,59 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import logging
+import os
+import sys
+
+from dinov2.eval.linear import get_args_parser as get_linear_args_parser
+from dinov2.logging import setup_logging
+from dinov2.run.submit import get_args_parser, submit_jobs
+
+
+logger = logging.getLogger("dinov2")
+
+
+class Evaluator:
+ def __init__(self, args):
+ self.args = args
+
+ def __call__(self):
+ from dinov2.eval.linear import main as linear_main
+
+ self._setup_args()
+ linear_main(self.args)
+
+ def checkpoint(self):
+ import submitit
+
+ logger.info(f"Requeuing {self.args}")
+ empty = type(self)(self.args)
+ return submitit.helpers.DelayedSubmission(empty)
+
+ def _setup_args(self):
+ import submitit
+
+ job_env = submitit.JobEnvironment()
+ self.args.output_dir = self.args.output_dir.replace("%j", str(job_env.job_id))
+ logger.info(f"Process group: {job_env.num_tasks} tasks, rank: {job_env.global_rank}")
+ logger.info(f"Args: {self.args}")
+
+
+def main():
+ description = "Submitit launcher for DINOv2 linear evaluation"
+ linear_args_parser = get_linear_args_parser(add_help=False)
+ parents = [linear_args_parser]
+ args_parser = get_args_parser(description=description, parents=parents)
+ args = args_parser.parse_args()
+
+ setup_logging()
+
+ assert os.path.exists(args.config_file), "Configuration file does not exist!"
+ submit_jobs(Evaluator, args, name="dinov2:linear")
+ return 0
+
+
+if __name__ == "__main__":
+ sys.exit(main())
diff --git a/models/dsp/dinov2/dinov2/run/eval/log_regression.py b/models/dsp/dinov2/dinov2/run/eval/log_regression.py
new file mode 100644
index 0000000000000000000000000000000000000000..cdf02181122de72cfa463ef38494967219df9cf3
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/run/eval/log_regression.py
@@ -0,0 +1,59 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import logging
+import os
+import sys
+
+from dinov2.eval.log_regression import get_args_parser as get_log_regression_args_parser
+from dinov2.logging import setup_logging
+from dinov2.run.submit import get_args_parser, submit_jobs
+
+
+logger = logging.getLogger("dinov2")
+
+
+class Evaluator:
+ def __init__(self, args):
+ self.args = args
+
+ def __call__(self):
+ from dinov2.eval.log_regression import main as log_regression_main
+
+ self._setup_args()
+ log_regression_main(self.args)
+
+ def checkpoint(self):
+ import submitit
+
+ logger.info(f"Requeuing {self.args}")
+ empty = type(self)(self.args)
+ return submitit.helpers.DelayedSubmission(empty)
+
+ def _setup_args(self):
+ import submitit
+
+ job_env = submitit.JobEnvironment()
+ self.args.output_dir = self.args.output_dir.replace("%j", str(job_env.job_id))
+ logger.info(f"Process group: {job_env.num_tasks} tasks, rank: {job_env.global_rank}")
+ logger.info(f"Args: {self.args}")
+
+
+def main():
+ description = "Submitit launcher for DINOv2 logistic evaluation"
+ log_regression_args_parser = get_log_regression_args_parser(add_help=False)
+ parents = [log_regression_args_parser]
+ args_parser = get_args_parser(description=description, parents=parents)
+ args = args_parser.parse_args()
+
+ setup_logging()
+
+ assert os.path.exists(args.config_file), "Configuration file does not exist!"
+ submit_jobs(Evaluator, args, name="dinov2:logreg")
+ return 0
+
+
+if __name__ == "__main__":
+ sys.exit(main())
diff --git a/models/dsp/dinov2/dinov2/run/submit.py b/models/dsp/dinov2/dinov2/run/submit.py
new file mode 100644
index 0000000000000000000000000000000000000000..4d1f718e704cf9a48913422404c25a7fcc50e738
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/run/submit.py
@@ -0,0 +1,122 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import argparse
+import logging
+import os
+from pathlib import Path
+from typing import List, Optional
+
+import submitit
+
+from dinov2.utils.cluster import (
+ get_slurm_executor_parameters,
+ get_slurm_partition,
+ get_user_checkpoint_path,
+)
+
+
+logger = logging.getLogger("dinov2")
+
+
+def get_args_parser(
+ description: Optional[str] = None,
+ parents: Optional[List[argparse.ArgumentParser]] = None,
+ add_help: bool = True,
+) -> argparse.ArgumentParser:
+ parents = parents or []
+ slurm_partition = get_slurm_partition()
+ parser = argparse.ArgumentParser(
+ description=description,
+ parents=parents,
+ add_help=add_help,
+ )
+ parser.add_argument(
+ "--ngpus",
+ "--gpus",
+ "--gpus-per-node",
+ default=8,
+ type=int,
+ help="Number of GPUs to request on each node",
+ )
+ parser.add_argument(
+ "--nodes",
+ "--nnodes",
+ default=1,
+ type=int,
+ help="Number of nodes to request",
+ )
+ parser.add_argument(
+ "--timeout",
+ default=2800,
+ type=int,
+ help="Duration of the job",
+ )
+ parser.add_argument(
+ "--partition",
+ default=slurm_partition,
+ type=str,
+ help="Partition where to submit",
+ )
+ parser.add_argument(
+ "--use-volta32",
+ action="store_true",
+ help="Request V100-32GB GPUs",
+ )
+ parser.add_argument(
+ "--comment",
+ default="",
+ type=str,
+ help="Comment to pass to scheduler, e.g. priority message",
+ )
+ parser.add_argument(
+ "--exclude",
+ default="",
+ type=str,
+ help="Nodes to exclude",
+ )
+ return parser
+
+
+def get_shared_folder() -> Path:
+ user_checkpoint_path = get_user_checkpoint_path()
+ if user_checkpoint_path is None:
+ raise RuntimeError("Path to user checkpoint cannot be determined")
+ path = user_checkpoint_path / "experiments"
+ path.mkdir(exist_ok=True)
+ return path
+
+
+def submit_jobs(task_class, args, name: str):
+ if not args.output_dir:
+ args.output_dir = str(get_shared_folder() / "%j")
+
+ Path(args.output_dir).mkdir(parents=True, exist_ok=True)
+ executor = submitit.AutoExecutor(folder=args.output_dir, slurm_max_num_timeout=30)
+
+ kwargs = {}
+ if args.use_volta32:
+ kwargs["slurm_constraint"] = "volta32gb"
+ if args.comment:
+ kwargs["slurm_comment"] = args.comment
+ if args.exclude:
+ kwargs["slurm_exclude"] = args.exclude
+
+ executor_params = get_slurm_executor_parameters(
+ nodes=args.nodes,
+ num_gpus_per_node=args.ngpus,
+ timeout_min=args.timeout, # max is 60 * 72
+ slurm_signal_delay_s=120,
+ slurm_partition=args.partition,
+ **kwargs,
+ )
+ executor.update_parameters(name=name, **executor_params)
+
+ task = task_class(args)
+ job = executor.submit(task)
+
+ logger.info(f"Submitted job_id: {job.job_id}")
+ str_output_dir = os.path.abspath(args.output_dir).replace("%j", str(job.job_id))
+ logger.info(f"Logs and checkpoints will be saved at: {str_output_dir}")
diff --git a/models/dsp/dinov2/dinov2/run/train/train.py b/models/dsp/dinov2/dinov2/run/train/train.py
new file mode 100644
index 0000000000000000000000000000000000000000..c2366e9bf79765e6abcd70dda6b43f31cb7093eb
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/run/train/train.py
@@ -0,0 +1,59 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import logging
+import os
+import sys
+
+from dinov2.logging import setup_logging
+from dinov2.train import get_args_parser as get_train_args_parser
+from dinov2.run.submit import get_args_parser, submit_jobs
+
+
+logger = logging.getLogger("dinov2")
+
+
+class Trainer(object):
+ def __init__(self, args):
+ self.args = args
+
+ def __call__(self):
+ from dinov2.train import main as train_main
+
+ self._setup_args()
+ train_main(self.args)
+
+ def checkpoint(self):
+ import submitit
+
+ logger.info(f"Requeuing {self.args}")
+ empty = type(self)(self.args)
+ return submitit.helpers.DelayedSubmission(empty)
+
+ def _setup_args(self):
+ import submitit
+
+ job_env = submitit.JobEnvironment()
+ self.args.output_dir = self.args.output_dir.replace("%j", str(job_env.job_id))
+ logger.info(f"Process group: {job_env.num_tasks} tasks, rank: {job_env.global_rank}")
+ logger.info(f"Args: {self.args}")
+
+
+def main():
+ description = "Submitit launcher for DINOv2 training"
+ train_args_parser = get_train_args_parser(add_help=False)
+ parents = [train_args_parser]
+ args_parser = get_args_parser(description=description, parents=parents)
+ args = args_parser.parse_args()
+
+ setup_logging()
+
+ assert os.path.exists(args.config_file), "Configuration file does not exist!"
+ submit_jobs(Trainer, args, name="dinov2:train")
+ return 0
+
+
+if __name__ == "__main__":
+ sys.exit(main())
diff --git a/models/dsp/dinov2/dinov2/thirdparty/CLIP/LICENSE b/models/dsp/dinov2/dinov2/thirdparty/CLIP/LICENSE
new file mode 100644
index 0000000000000000000000000000000000000000..c123b69334717d178daa674c2d08e3383fe36134
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/thirdparty/CLIP/LICENSE
@@ -0,0 +1,21 @@
+MIT License
+
+Copyright (c) 2021 OpenAI
+
+Permission is hereby granted, free of charge, to any person obtaining a copy
+of this software and associated documentation files (the "Software"), to deal
+in the Software without restriction, including without limitation the rights
+to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
+copies of the Software, and to permit persons to whom the Software is
+furnished to do so, subject to the following conditions:
+
+The above copyright notice and this permission notice shall be included in all
+copies or substantial portions of the Software.
+
+THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
+IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
+FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
+AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
+LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
+OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
+SOFTWARE.
diff --git a/models/dsp/dinov2/dinov2/thirdparty/CLIP/clip/simple_tokenizer.py b/models/dsp/dinov2/dinov2/thirdparty/CLIP/clip/simple_tokenizer.py
new file mode 100644
index 0000000000000000000000000000000000000000..0b1a6b1470840ca66638524fbd541d5bdf67b4f8
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/thirdparty/CLIP/clip/simple_tokenizer.py
@@ -0,0 +1,135 @@
+import gzip
+import html
+import os
+from functools import lru_cache
+
+import ftfy
+import regex as re
+
+
+@lru_cache()
+def default_bpe():
+ return os.path.join(os.path.dirname(os.path.abspath(__file__)), "bpe_simple_vocab_16e6.txt.gz")
+
+
+@lru_cache()
+def bytes_to_unicode():
+ """
+ Returns list of utf-8 byte and a corresponding list of unicode strings.
+ The reversible bpe codes work on unicode strings.
+ This means you need a large # of unicode characters in your vocab if you want to avoid UNKs.
+ When you're at something like a 10B token dataset you end up needing around 5K for decent coverage.
+ This is a signficant percentage of your normal, say, 32K bpe vocab.
+ To avoid that, we want lookup tables between utf-8 bytes and unicode strings.
+ And avoids mapping to whitespace/control characters the bpe code barfs on.
+ """
+ bs = list(range(ord("!"), ord("~") + 1)) + list(range(ord("¡"), ord("¬") + 1)) + list(range(ord("®"), ord("ÿ") + 1))
+ cs = bs[:]
+ n = 0
+ for b in range(2**8):
+ if b not in bs:
+ bs.append(b)
+ cs.append(2**8 + n)
+ n += 1
+ cs = [chr(n) for n in cs]
+ return dict(zip(bs, cs))
+
+
+def get_pairs(word):
+ """Return set of symbol pairs in a word.
+ Word is represented as tuple of symbols (symbols being variable-length strings).
+ """
+ pairs = set()
+ prev_char = word[0]
+ for char in word[1:]:
+ pairs.add((prev_char, char))
+ prev_char = char
+ return pairs
+
+
+def basic_clean(text):
+ text = ftfy.fix_text(text)
+ text = html.unescape(html.unescape(text))
+ return text.strip()
+
+
+def whitespace_clean(text):
+ text = re.sub(r"\s+", " ", text)
+ text = text.strip()
+ return text
+
+
+class SimpleTokenizer(object):
+ def __init__(self, bpe_path: str = default_bpe()):
+ self.byte_encoder = bytes_to_unicode()
+ self.byte_decoder = {v: k for k, v in self.byte_encoder.items()}
+ merges = gzip.open(bpe_path).read().decode("utf-8").split("\n")
+ merges = merges[1 : 49152 - 256 - 2 + 1]
+ merges = [tuple(merge.split()) for merge in merges]
+ vocab = list(bytes_to_unicode().values())
+ vocab = vocab + [v + "" for v in vocab]
+ for merge in merges:
+ vocab.append("".join(merge))
+ vocab.extend(["<|startoftext|>", "<|endoftext|>"])
+ self.encoder = dict(zip(vocab, range(len(vocab))))
+ self.decoder = {v: k for k, v in self.encoder.items()}
+ self.bpe_ranks = dict(zip(merges, range(len(merges))))
+ self.cache = {"<|startoftext|>": "<|startoftext|>", "<|endoftext|>": "<|endoftext|>"}
+ self.pat = re.compile(
+ r"""<\|startoftext\|>|<\|endoftext\|>|'s|'t|'re|'ve|'m|'ll|'d|[\p{L}]+|[\p{N}]|[^\s\p{L}\p{N}]+""",
+ re.IGNORECASE,
+ )
+
+ def bpe(self, token):
+ if token in self.cache:
+ return self.cache[token]
+ word = tuple(token[:-1]) + (token[-1] + "",)
+ pairs = get_pairs(word)
+
+ if not pairs:
+ return token + ""
+
+ while True:
+ bigram = min(pairs, key=lambda pair: self.bpe_ranks.get(pair, float("inf")))
+ if bigram not in self.bpe_ranks:
+ break
+ first, second = bigram
+ new_word = []
+ i = 0
+ while i < len(word):
+ try:
+ j = word.index(first, i)
+ new_word.extend(word[i:j])
+ i = j
+ except Exception:
+ new_word.extend(word[i:])
+ break
+
+ if word[i] == first and i < len(word) - 1 and word[i + 1] == second:
+ new_word.append(first + second)
+ i += 2
+ else:
+ new_word.append(word[i])
+ i += 1
+ new_word = tuple(new_word)
+ word = new_word
+ if len(word) == 1:
+ break
+ else:
+ pairs = get_pairs(word)
+ word = " ".join(word)
+ self.cache[token] = word
+ return word
+
+ def encode(self, text):
+ bpe_tokens = []
+ text = whitespace_clean(basic_clean(text)).lower()
+ for token in re.findall(self.pat, text):
+ token = "".join(self.byte_encoder[b] for b in token.encode("utf-8"))
+ bpe_tokens.extend(self.encoder[bpe_token] for bpe_token in self.bpe(token).split(" "))
+ return bpe_tokens
+
+ def decode(self, tokens):
+ text = "".join([self.decoder[token] for token in tokens])
+ text = bytearray([self.byte_decoder[c] for c in text]).decode("utf-8", errors="replace").replace("", " ")
+ return text
diff --git a/models/dsp/dinov2/dinov2/train/__init__.py b/models/dsp/dinov2/dinov2/train/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..5f1752922d04fff0112eb7796be28ff6b68c6073
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/train/__init__.py
@@ -0,0 +1,7 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from .train import get_args_parser, main
+from .ssl_meta_arch import SSLMetaArch
diff --git a/models/dsp/dinov2/dinov2/train/ssl_meta_arch.py b/models/dsp/dinov2/dinov2/train/ssl_meta_arch.py
new file mode 100644
index 0000000000000000000000000000000000000000..3ccf15e904ebeb6134dfb4f5c99da4fc8d41b8e4
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/train/ssl_meta_arch.py
@@ -0,0 +1,400 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from functools import partial
+import logging
+
+import torch
+from torch import nn
+
+from dinov2.loss import DINOLoss, iBOTPatchLoss, KoLeoLoss
+from dinov2.models import build_model_from_cfg
+from dinov2.layers import DINOHead
+from dinov2.utils.utils import has_batchnorms
+from dinov2.utils.param_groups import get_params_groups_with_decay, fuse_params_groups
+from dinov2.fsdp import get_fsdp_wrapper, ShardedGradScaler, get_fsdp_modules, reshard_fsdp_model
+
+from dinov2.models.vision_transformer import BlockChunk
+
+
+try:
+ from xformers.ops import fmha
+except ImportError:
+ raise AssertionError("xFormers is required for training")
+
+
+logger = logging.getLogger("dinov2")
+
+
+class SSLMetaArch(nn.Module):
+ def __init__(self, cfg):
+ super().__init__()
+ self.cfg = cfg
+ self.fp16_scaler = ShardedGradScaler() if cfg.compute_precision.grad_scaler else None
+
+ student_model_dict = dict()
+ teacher_model_dict = dict()
+
+ student_backbone, teacher_backbone, embed_dim = build_model_from_cfg(cfg)
+ student_model_dict["backbone"] = student_backbone
+ teacher_model_dict["backbone"] = teacher_backbone
+ logger.info(f"OPTIONS -- architecture : embed_dim: {embed_dim}")
+
+ if cfg.student.pretrained_weights:
+ chkpt = torch.load(cfg.student.pretrained_weights)
+ logger.info(f"OPTIONS -- pretrained weights: loading from {cfg.student.pretrained_weights}")
+ student_backbone.load_state_dict(chkpt["model"], strict=False)
+
+ self.embed_dim = embed_dim
+ self.dino_out_dim = cfg.dino.head_n_prototypes
+
+ self.do_dino = cfg.dino.loss_weight > 0
+ self.do_koleo = cfg.dino.koleo_loss_weight > 0
+ self.do_ibot = cfg.ibot.loss_weight > 0
+ self.ibot_separate_head = cfg.ibot.separate_head
+
+ logger.info("OPTIONS -- DINO")
+ if self.do_dino:
+ logger.info(f"OPTIONS -- DINO -- loss_weight: {cfg.dino.loss_weight}")
+ logger.info(f"OPTIONS -- DINO -- head_n_prototypes: {cfg.dino.head_n_prototypes}")
+ logger.info(f"OPTIONS -- DINO -- head_bottleneck_dim: {cfg.dino.head_bottleneck_dim}")
+ logger.info(f"OPTIONS -- DINO -- head_hidden_dim: {cfg.dino.head_hidden_dim}")
+ self.dino_loss_weight = cfg.dino.loss_weight
+ dino_head = partial(
+ DINOHead,
+ in_dim=embed_dim,
+ out_dim=cfg.dino.head_n_prototypes,
+ hidden_dim=cfg.dino.head_hidden_dim,
+ bottleneck_dim=cfg.dino.head_bottleneck_dim,
+ nlayers=cfg.dino.head_nlayers,
+ )
+ self.dino_loss = DINOLoss(self.dino_out_dim)
+ if self.do_koleo:
+ logger.info("OPTIONS -- DINO -- applying KOLEO regularization")
+ self.koleo_loss = KoLeoLoss()
+
+ else:
+ logger.info("OPTIONS -- DINO -- not using DINO")
+
+ if self.do_dino or self.do_ibot:
+ student_model_dict["dino_head"] = dino_head()
+ teacher_model_dict["dino_head"] = dino_head()
+
+ logger.info("OPTIONS -- IBOT")
+ logger.info(f"OPTIONS -- IBOT -- loss_weight: {cfg.ibot.loss_weight}")
+ logger.info(f"OPTIONS -- IBOT masking -- ibot_mask_ratio_tuple: {cfg.ibot.mask_ratio_min_max}")
+ logger.info(f"OPTIONS -- IBOT masking -- ibot_mask_sample_probability: {cfg.ibot.mask_sample_probability}")
+ if self.do_ibot:
+ self.ibot_loss_weight = cfg.ibot.loss_weight
+ assert max(cfg.ibot.mask_ratio_min_max) > 0, "please provide a positive mask ratio tuple for ibot"
+ assert cfg.ibot.mask_sample_probability > 0, "please provide a positive mask probability for ibot"
+ self.ibot_out_dim = cfg.ibot.head_n_prototypes if self.ibot_separate_head else cfg.dino.head_n_prototypes
+ self.ibot_patch_loss = iBOTPatchLoss(self.ibot_out_dim)
+ if self.ibot_separate_head:
+ logger.info(f"OPTIONS -- IBOT -- loss_weight: {cfg.ibot.loss_weight}")
+ logger.info(f"OPTIONS -- IBOT -- head_n_prototypes: {cfg.ibot.head_n_prototypes}")
+ logger.info(f"OPTIONS -- IBOT -- head_bottleneck_dim: {cfg.ibot.head_bottleneck_dim}")
+ logger.info(f"OPTIONS -- IBOT -- head_hidden_dim: {cfg.ibot.head_hidden_dim}")
+ ibot_head = partial(
+ DINOHead,
+ in_dim=embed_dim,
+ out_dim=cfg.ibot.head_n_prototypes,
+ hidden_dim=cfg.ibot.head_hidden_dim,
+ bottleneck_dim=cfg.ibot.head_bottleneck_dim,
+ nlayers=cfg.ibot.head_nlayers,
+ )
+ student_model_dict["ibot_head"] = ibot_head()
+ teacher_model_dict["ibot_head"] = ibot_head()
+ else:
+ logger.info("OPTIONS -- IBOT -- head shared with DINO")
+
+ self.need_to_synchronize_fsdp_streams = True
+
+ self.student = nn.ModuleDict(student_model_dict)
+ self.teacher = nn.ModuleDict(teacher_model_dict)
+
+ # there is no backpropagation through the teacher, so no need for gradients
+ for p in self.teacher.parameters():
+ p.requires_grad = False
+ logger.info(f"Student and Teacher are built: they are both {cfg.student.arch} network.")
+
+ def forward(self, inputs):
+ raise NotImplementedError
+
+ def backprop_loss(self, loss):
+ if self.fp16_scaler is not None:
+ self.fp16_scaler.scale(loss).backward()
+ else:
+ loss.backward()
+
+ def forward_backward(self, images, teacher_temp):
+ n_global_crops = 2
+ assert n_global_crops == 2
+ n_local_crops = self.cfg.crops.local_crops_number
+
+ global_crops = images["collated_global_crops"].cuda(non_blocking=True)
+ local_crops = images["collated_local_crops"].cuda(non_blocking=True)
+
+ masks = images["collated_masks"].cuda(non_blocking=True)
+ mask_indices_list = images["mask_indices_list"].cuda(non_blocking=True)
+ n_masked_patches_tensor = images["n_masked_patches"].cuda(non_blocking=True)
+ n_masked_patches = mask_indices_list.shape[0]
+ upperbound = images["upperbound"]
+ masks_weight = images["masks_weight"].cuda(non_blocking=True)
+
+ n_local_crops_loss_terms = max(n_local_crops * n_global_crops, 1)
+ n_global_crops_loss_terms = (n_global_crops - 1) * n_global_crops
+
+ do_dino = self.do_dino
+ do_ibot = self.do_ibot
+
+ # loss scales
+ ibot_loss_scale = 1.0 / n_global_crops
+
+ # teacher output
+ @torch.no_grad()
+ def get_teacher_output():
+ x, n_global_crops_teacher = global_crops, n_global_crops
+ teacher_backbone_output_dict = self.teacher.backbone(x, is_training=True)
+ teacher_cls_tokens = teacher_backbone_output_dict["x_norm_clstoken"]
+ teacher_cls_tokens = teacher_cls_tokens.chunk(n_global_crops_teacher)
+ # watch out: these are chunked and cat'd in reverse so A is matched to B in the global crops dino loss
+ teacher_cls_tokens = torch.cat((teacher_cls_tokens[1], teacher_cls_tokens[0]))
+ ibot_teacher_patch_tokens = teacher_backbone_output_dict["x_norm_patchtokens"]
+ _dim = ibot_teacher_patch_tokens.shape[-1]
+ n_cls_tokens = teacher_cls_tokens.shape[0]
+
+ if do_ibot and not self.ibot_separate_head:
+ buffer_tensor_teacher = ibot_teacher_patch_tokens.new_zeros(upperbound + n_cls_tokens, _dim)
+ buffer_tensor_teacher[:n_cls_tokens].copy_(teacher_cls_tokens)
+ torch.index_select(
+ ibot_teacher_patch_tokens.flatten(0, 1),
+ dim=0,
+ index=mask_indices_list,
+ out=buffer_tensor_teacher[n_cls_tokens : n_cls_tokens + n_masked_patches],
+ )
+ tokens_after_head = self.teacher.dino_head(buffer_tensor_teacher)
+ teacher_cls_tokens_after_head = tokens_after_head[:n_cls_tokens]
+ masked_teacher_patch_tokens_after_head = tokens_after_head[
+ n_cls_tokens : n_cls_tokens + n_masked_patches
+ ]
+ elif do_ibot and self.ibot_separate_head:
+ buffer_tensor_teacher = ibot_teacher_patch_tokens.new_zeros(upperbound, _dim)
+ torch.index_select(
+ ibot_teacher_patch_tokens.flatten(0, 1),
+ dim=0,
+ index=mask_indices_list,
+ out=buffer_tensor_teacher[:n_masked_patches],
+ )
+ teacher_cls_tokens_after_head = self.teacher.dino_head(teacher_cls_tokens)
+ masked_teacher_patch_tokens_after_head = self.teacher.ibot_head(buffer_tensor_teacher)[
+ :n_masked_patches
+ ]
+ else:
+ teacher_cls_tokens_after_head = self.teacher.dino_head(teacher_cls_tokens)
+ masked_teacher_ibot_softmaxed_centered = None
+
+ if self.cfg.train.centering == "centering":
+ teacher_dino_softmaxed_centered_list = self.dino_loss.softmax_center_teacher(
+ teacher_cls_tokens_after_head, teacher_temp=teacher_temp
+ ).view(n_global_crops_teacher, -1, *teacher_cls_tokens_after_head.shape[1:])
+ self.dino_loss.update_center(teacher_cls_tokens_after_head)
+ if do_ibot:
+ masked_teacher_patch_tokens_after_head = masked_teacher_patch_tokens_after_head.unsqueeze(0)
+ masked_teacher_ibot_softmaxed_centered = self.ibot_patch_loss.softmax_center_teacher(
+ masked_teacher_patch_tokens_after_head[:, :n_masked_patches], teacher_temp=teacher_temp
+ )
+ masked_teacher_ibot_softmaxed_centered = masked_teacher_ibot_softmaxed_centered.squeeze(0)
+ self.ibot_patch_loss.update_center(masked_teacher_patch_tokens_after_head[:n_masked_patches])
+
+ elif self.cfg.train.centering == "sinkhorn_knopp":
+ teacher_dino_softmaxed_centered_list = self.dino_loss.sinkhorn_knopp_teacher(
+ teacher_cls_tokens_after_head, teacher_temp=teacher_temp
+ ).view(n_global_crops_teacher, -1, *teacher_cls_tokens_after_head.shape[1:])
+
+ if do_ibot:
+ masked_teacher_ibot_softmaxed_centered = self.ibot_patch_loss.sinkhorn_knopp_teacher(
+ masked_teacher_patch_tokens_after_head,
+ teacher_temp=teacher_temp,
+ n_masked_patches_tensor=n_masked_patches_tensor,
+ )
+
+ else:
+ raise NotImplementedError
+
+ return teacher_dino_softmaxed_centered_list, masked_teacher_ibot_softmaxed_centered
+
+ teacher_dino_softmaxed_centered_list, masked_teacher_ibot_softmaxed_centered = get_teacher_output()
+ reshard_fsdp_model(self.teacher)
+
+ loss_dict = {}
+
+ loss_accumulator = 0 # for backprop
+ student_global_backbone_output_dict, student_local_backbone_output_dict = self.student.backbone(
+ [global_crops, local_crops], masks=[masks, None], is_training=True
+ )
+
+ inputs_for_student_head_list = []
+
+ # 1a: local crops cls tokens
+ student_local_cls_tokens = student_local_backbone_output_dict["x_norm_clstoken"]
+ inputs_for_student_head_list.append(student_local_cls_tokens.unsqueeze(0))
+
+ # 1b: global crops cls tokens
+ student_global_cls_tokens = student_global_backbone_output_dict["x_norm_clstoken"]
+ inputs_for_student_head_list.append(student_global_cls_tokens.unsqueeze(0))
+
+ # 1c: global crops patch tokens
+ if do_ibot:
+ _dim = student_global_backbone_output_dict["x_norm_clstoken"].shape[-1]
+ ibot_student_patch_tokens = student_global_backbone_output_dict["x_norm_patchtokens"]
+ buffer_tensor_patch_tokens = ibot_student_patch_tokens.new_zeros(upperbound, _dim)
+ buffer_tensor_patch_tokens[:n_masked_patches].copy_(
+ torch.index_select(ibot_student_patch_tokens.flatten(0, 1), dim=0, index=mask_indices_list)
+ )
+ if not self.ibot_separate_head:
+ inputs_for_student_head_list.append(buffer_tensor_patch_tokens.unsqueeze(0))
+ else:
+ student_global_masked_patch_tokens_after_head = self.student.ibot_head(buffer_tensor_patch_tokens)[
+ :n_masked_patches
+ ]
+
+ # 2: run
+ _attn_bias, cat_inputs = fmha.BlockDiagonalMask.from_tensor_list(inputs_for_student_head_list)
+ outputs_list = _attn_bias.split(self.student.dino_head(cat_inputs))
+
+ # 3a: local crops cls tokens
+ student_local_cls_tokens_after_head = outputs_list.pop(0).squeeze(0)
+
+ # 3b: global crops cls tokens
+ student_global_cls_tokens_after_head = outputs_list.pop(0).squeeze(0)
+
+ # 3c: global crops patch tokens
+ if do_ibot and not self.ibot_separate_head:
+ student_global_masked_patch_tokens_after_head = outputs_list.pop(0).squeeze(0)[:n_masked_patches]
+
+ if n_local_crops > 0:
+ dino_local_crops_loss = self.dino_loss(
+ student_output_list=student_local_cls_tokens_after_head.chunk(n_local_crops),
+ teacher_out_softmaxed_centered_list=teacher_dino_softmaxed_centered_list,
+ ) / (n_global_crops_loss_terms + n_local_crops_loss_terms)
+
+ # store for display
+ loss_dict["dino_local_crops_loss"] = dino_local_crops_loss
+
+ # accumulate loss
+ loss_accumulator += self.dino_loss_weight * dino_local_crops_loss
+
+ # process global crops
+ loss_scales = 2 # this is here since we process global crops together
+
+ if do_dino:
+ # compute loss
+ dino_global_crops_loss = (
+ self.dino_loss(
+ student_output_list=[student_global_cls_tokens_after_head],
+ teacher_out_softmaxed_centered_list=[
+ teacher_dino_softmaxed_centered_list.flatten(0, 1)
+ ], # these were chunked and stacked in reverse so A is matched to B
+ )
+ * loss_scales
+ / (n_global_crops_loss_terms + n_local_crops_loss_terms)
+ )
+
+ loss_dict["dino_global_crops_loss"] = dino_global_crops_loss
+
+ # accumulate loss
+ loss_accumulator += self.dino_loss_weight * dino_global_crops_loss
+
+ student_cls_tokens = student_global_cls_tokens
+
+ if self.do_koleo:
+ koleo_loss = self.cfg.dino.koleo_loss_weight * sum(
+ self.koleo_loss(p) for p in student_cls_tokens.chunk(2)
+ ) # we don't apply koleo loss between cls tokens of a same image
+ loss_accumulator += koleo_loss
+ loss_dict["koleo_loss"] = (
+ koleo_loss / loss_scales
+ ) # this is to display the same losses as before but we can remove eventually
+
+ if do_ibot:
+ # compute loss
+ ibot_patch_loss = (
+ self.ibot_patch_loss.forward_masked(
+ student_global_masked_patch_tokens_after_head,
+ masked_teacher_ibot_softmaxed_centered,
+ student_masks_flat=masks,
+ n_masked_patches=n_masked_patches,
+ masks_weight=masks_weight,
+ )
+ * loss_scales
+ * ibot_loss_scale
+ )
+
+ # store for display
+ loss_dict["ibot_loss"] = ibot_patch_loss / 2
+
+ # accumulate loss
+ loss_accumulator += self.ibot_loss_weight * ibot_patch_loss
+
+ self.backprop_loss(loss_accumulator)
+
+ self.fsdp_synchronize_streams()
+
+ return loss_dict
+
+ def fsdp_synchronize_streams(self):
+ if self.need_to_synchronize_fsdp_streams:
+ torch.cuda.synchronize()
+ self.student.dino_head._streams = (
+ self.teacher.dino_head._streams
+ ) = self.student.backbone._streams = self.teacher.backbone._streams
+ self.need_to_synchronize_fsdp_streams = False
+
+ def update_teacher(self, m):
+ student_param_list = []
+ teacher_param_list = []
+ with torch.no_grad():
+ for k in self.student.keys():
+ for ms, mt in zip(get_fsdp_modules(self.student[k]), get_fsdp_modules(self.teacher[k])):
+ student_param_list += ms.params
+ teacher_param_list += mt.params
+ torch._foreach_mul_(teacher_param_list, m)
+ torch._foreach_add_(teacher_param_list, student_param_list, alpha=1 - m)
+
+ def train(self):
+ super().train()
+ self.teacher.eval()
+
+ def get_maybe_fused_params_for_submodel(self, m):
+ params_groups = get_params_groups_with_decay(
+ model=m,
+ lr_decay_rate=self.cfg.optim.layerwise_decay,
+ patch_embed_lr_mult=self.cfg.optim.patch_embed_lr_mult,
+ )
+ fused_params_groups = fuse_params_groups(params_groups)
+ logger.info("fusing param groups")
+
+ for g in fused_params_groups:
+ g["foreach"] = True
+ return fused_params_groups
+
+ def get_params_groups(self):
+ all_params_groups = []
+ for m in self.student.values():
+ all_params_groups += self.get_maybe_fused_params_for_submodel(m)
+ return all_params_groups
+
+ def prepare_for_distributed_training(self):
+ logger.info("DISTRIBUTED FSDP -- preparing model for distributed training")
+ if has_batchnorms(self.student):
+ raise NotImplementedError
+ # below will synchronize all student subnetworks across gpus:
+ for k, v in self.student.items():
+ self.teacher[k].load_state_dict(self.student[k].state_dict())
+ student_model_cfg = self.cfg.compute_precision.student[k]
+ self.student[k] = get_fsdp_wrapper(student_model_cfg, modules_to_wrap={BlockChunk})(self.student[k])
+ teacher_model_cfg = self.cfg.compute_precision.teacher[k]
+ self.teacher[k] = get_fsdp_wrapper(teacher_model_cfg, modules_to_wrap={BlockChunk})(self.teacher[k])
diff --git a/models/dsp/dinov2/dinov2/train/train.py b/models/dsp/dinov2/dinov2/train/train.py
new file mode 100644
index 0000000000000000000000000000000000000000..473b8d01473654182de9f91c94a2d8720fe096a5
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/train/train.py
@@ -0,0 +1,318 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import argparse
+import logging
+import math
+import os
+from functools import partial
+
+from fvcore.common.checkpoint import PeriodicCheckpointer
+import torch
+
+from dinov2.data import SamplerType, make_data_loader, make_dataset
+from dinov2.data import collate_data_and_cast, DataAugmentationDINO, MaskingGenerator
+import dinov2.distributed as distributed
+from dinov2.fsdp import FSDPCheckpointer
+from dinov2.logging import MetricLogger
+from dinov2.utils.config import setup
+from dinov2.utils.utils import CosineScheduler
+
+from dinov2.train.ssl_meta_arch import SSLMetaArch
+
+
+torch.backends.cuda.matmul.allow_tf32 = True # PyTorch 1.12 sets this to False by default
+logger = logging.getLogger("dinov2")
+
+
+def get_args_parser(add_help: bool = True):
+ parser = argparse.ArgumentParser("DINOv2 training", add_help=add_help)
+ parser.add_argument("--config-file", default="", metavar="FILE", help="path to config file")
+ parser.add_argument(
+ "--no-resume",
+ action="store_true",
+ help="Whether to not attempt to resume from the checkpoint directory. ",
+ )
+ parser.add_argument("--eval-only", action="store_true", help="perform evaluation only")
+ parser.add_argument("--eval", type=str, default="", help="Eval type to perform")
+ parser.add_argument(
+ "opts",
+ help="""
+Modify config options at the end of the command. For Yacs configs, use
+space-separated "PATH.KEY VALUE" pairs.
+For python-based LazyConfig, use "path.key=value".
+ """.strip(),
+ default=None,
+ nargs=argparse.REMAINDER,
+ )
+ parser.add_argument(
+ "--output-dir",
+ "--output_dir",
+ default="",
+ type=str,
+ help="Output directory to save logs and checkpoints",
+ )
+
+ return parser
+
+
+def build_optimizer(cfg, params_groups):
+ return torch.optim.AdamW(params_groups, betas=(cfg.optim.adamw_beta1, cfg.optim.adamw_beta2))
+
+
+def build_schedulers(cfg):
+ OFFICIAL_EPOCH_LENGTH = cfg.train.OFFICIAL_EPOCH_LENGTH
+ lr = dict(
+ base_value=cfg.optim["lr"],
+ final_value=cfg.optim["min_lr"],
+ total_iters=cfg.optim["epochs"] * OFFICIAL_EPOCH_LENGTH,
+ warmup_iters=cfg.optim["warmup_epochs"] * OFFICIAL_EPOCH_LENGTH,
+ start_warmup_value=0,
+ )
+ wd = dict(
+ base_value=cfg.optim["weight_decay"],
+ final_value=cfg.optim["weight_decay_end"],
+ total_iters=cfg.optim["epochs"] * OFFICIAL_EPOCH_LENGTH,
+ )
+ momentum = dict(
+ base_value=cfg.teacher["momentum_teacher"],
+ final_value=cfg.teacher["final_momentum_teacher"],
+ total_iters=cfg.optim["epochs"] * OFFICIAL_EPOCH_LENGTH,
+ )
+ teacher_temp = dict(
+ base_value=cfg.teacher["teacher_temp"],
+ final_value=cfg.teacher["teacher_temp"],
+ total_iters=cfg.teacher["warmup_teacher_temp_epochs"] * OFFICIAL_EPOCH_LENGTH,
+ warmup_iters=cfg.teacher["warmup_teacher_temp_epochs"] * OFFICIAL_EPOCH_LENGTH,
+ start_warmup_value=cfg.teacher["warmup_teacher_temp"],
+ )
+
+ lr_schedule = CosineScheduler(**lr)
+ wd_schedule = CosineScheduler(**wd)
+ momentum_schedule = CosineScheduler(**momentum)
+ teacher_temp_schedule = CosineScheduler(**teacher_temp)
+ last_layer_lr_schedule = CosineScheduler(**lr)
+
+ last_layer_lr_schedule.schedule[
+ : cfg.optim["freeze_last_layer_epochs"] * OFFICIAL_EPOCH_LENGTH
+ ] = 0 # mimicking the original schedules
+
+ logger.info("Schedulers ready.")
+
+ return (
+ lr_schedule,
+ wd_schedule,
+ momentum_schedule,
+ teacher_temp_schedule,
+ last_layer_lr_schedule,
+ )
+
+
+def apply_optim_scheduler(optimizer, lr, wd, last_layer_lr):
+ for param_group in optimizer.param_groups:
+ is_last_layer = param_group["is_last_layer"]
+ lr_multiplier = param_group["lr_multiplier"]
+ wd_multiplier = param_group["wd_multiplier"]
+ param_group["weight_decay"] = wd * wd_multiplier
+ param_group["lr"] = (last_layer_lr if is_last_layer else lr) * lr_multiplier
+
+
+def do_test(cfg, model, iteration):
+ new_state_dict = model.teacher.state_dict()
+
+ if distributed.is_main_process():
+ iterstring = str(iteration)
+ eval_dir = os.path.join(cfg.train.output_dir, "eval", iterstring)
+ os.makedirs(eval_dir, exist_ok=True)
+ # save teacher checkpoint
+ teacher_ckp_path = os.path.join(eval_dir, "teacher_checkpoint.pth")
+ torch.save({"teacher": new_state_dict}, teacher_ckp_path)
+
+
+def do_train(cfg, model, resume=False):
+ model.train()
+ inputs_dtype = torch.half
+ fp16_scaler = model.fp16_scaler # for mixed precision training
+
+ # setup optimizer
+
+ optimizer = build_optimizer(cfg, model.get_params_groups())
+ (
+ lr_schedule,
+ wd_schedule,
+ momentum_schedule,
+ teacher_temp_schedule,
+ last_layer_lr_schedule,
+ ) = build_schedulers(cfg)
+
+ # checkpointer
+ checkpointer = FSDPCheckpointer(model, cfg.train.output_dir, optimizer=optimizer, save_to_disk=True)
+
+ start_iter = checkpointer.resume_or_load(cfg.MODEL.WEIGHTS, resume=resume).get("iteration", -1) + 1
+
+ OFFICIAL_EPOCH_LENGTH = cfg.train.OFFICIAL_EPOCH_LENGTH
+ max_iter = cfg.optim.epochs * OFFICIAL_EPOCH_LENGTH
+
+ periodic_checkpointer = PeriodicCheckpointer(
+ checkpointer,
+ period=3 * OFFICIAL_EPOCH_LENGTH,
+ max_iter=max_iter,
+ max_to_keep=3,
+ )
+
+ # setup data preprocessing
+
+ img_size = cfg.crops.global_crops_size
+ patch_size = cfg.student.patch_size
+ n_tokens = (img_size // patch_size) ** 2
+ mask_generator = MaskingGenerator(
+ input_size=(img_size // patch_size, img_size // patch_size),
+ max_num_patches=0.5 * img_size // patch_size * img_size // patch_size,
+ )
+
+ data_transform = DataAugmentationDINO(
+ cfg.crops.global_crops_scale,
+ cfg.crops.local_crops_scale,
+ cfg.crops.local_crops_number,
+ global_crops_size=cfg.crops.global_crops_size,
+ local_crops_size=cfg.crops.local_crops_size,
+ )
+
+ collate_fn = partial(
+ collate_data_and_cast,
+ mask_ratio_tuple=cfg.ibot.mask_ratio_min_max,
+ mask_probability=cfg.ibot.mask_sample_probability,
+ n_tokens=n_tokens,
+ mask_generator=mask_generator,
+ dtype=inputs_dtype,
+ )
+
+ # setup data loader
+
+ dataset = make_dataset(
+ dataset_str=cfg.train.dataset_path,
+ transform=data_transform,
+ target_transform=lambda _: (),
+ )
+ # sampler_type = SamplerType.INFINITE
+ sampler_type = SamplerType.SHARDED_INFINITE
+ data_loader = make_data_loader(
+ dataset=dataset,
+ batch_size=cfg.train.batch_size_per_gpu,
+ num_workers=cfg.train.num_workers,
+ shuffle=True,
+ seed=start_iter, # TODO: Fix this -- cfg.train.seed
+ sampler_type=sampler_type,
+ sampler_advance=0, # TODO(qas): fix this -- start_iter * cfg.train.batch_size_per_gpu,
+ drop_last=True,
+ collate_fn=collate_fn,
+ )
+
+ # training loop
+
+ iteration = start_iter
+
+ logger.info("Starting training from iteration {}".format(start_iter))
+ metrics_file = os.path.join(cfg.train.output_dir, "training_metrics.json")
+ metric_logger = MetricLogger(delimiter=" ", output_file=metrics_file)
+ header = "Training"
+
+ for data in metric_logger.log_every(
+ data_loader,
+ 10,
+ header,
+ max_iter,
+ start_iter,
+ ):
+ current_batch_size = data["collated_global_crops"].shape[0] / 2
+ if iteration > max_iter:
+ return
+
+ # apply schedules
+
+ lr = lr_schedule[iteration]
+ wd = wd_schedule[iteration]
+ mom = momentum_schedule[iteration]
+ teacher_temp = teacher_temp_schedule[iteration]
+ last_layer_lr = last_layer_lr_schedule[iteration]
+ apply_optim_scheduler(optimizer, lr, wd, last_layer_lr)
+
+ # compute losses
+
+ optimizer.zero_grad(set_to_none=True)
+ loss_dict = model.forward_backward(data, teacher_temp=teacher_temp)
+
+ # clip gradients
+
+ if fp16_scaler is not None:
+ if cfg.optim.clip_grad:
+ fp16_scaler.unscale_(optimizer)
+ for v in model.student.values():
+ v.clip_grad_norm_(cfg.optim.clip_grad)
+ fp16_scaler.step(optimizer)
+ fp16_scaler.update()
+ else:
+ if cfg.optim.clip_grad:
+ for v in model.student.values():
+ v.clip_grad_norm_(cfg.optim.clip_grad)
+ optimizer.step()
+
+ # perform teacher EMA update
+
+ model.update_teacher(mom)
+
+ # logging
+
+ if distributed.get_global_size() > 1:
+ for v in loss_dict.values():
+ torch.distributed.all_reduce(v)
+ loss_dict_reduced = {k: v.item() / distributed.get_global_size() for k, v in loss_dict.items()}
+
+ if math.isnan(sum(loss_dict_reduced.values())):
+ logger.info("NaN detected")
+ raise AssertionError
+ losses_reduced = sum(loss for loss in loss_dict_reduced.values())
+
+ metric_logger.update(lr=lr)
+ metric_logger.update(wd=wd)
+ metric_logger.update(mom=mom)
+ metric_logger.update(last_layer_lr=last_layer_lr)
+ metric_logger.update(current_batch_size=current_batch_size)
+ metric_logger.update(total_loss=losses_reduced, **loss_dict_reduced)
+
+ # checkpointing and testing
+
+ if cfg.evaluation.eval_period_iterations > 0 and (iteration + 1) % cfg.evaluation.eval_period_iterations == 0:
+ do_test(cfg, model, f"training_{iteration}")
+ torch.cuda.synchronize()
+ periodic_checkpointer.step(iteration)
+
+ iteration = iteration + 1
+ metric_logger.synchronize_between_processes()
+ return {k: meter.global_avg for k, meter in metric_logger.meters.items()}
+
+
+def main(args):
+ cfg = setup(args)
+
+ model = SSLMetaArch(cfg).to(torch.device("cuda"))
+ model.prepare_for_distributed_training()
+
+ logger.info("Model:\n{}".format(model))
+ if args.eval_only:
+ iteration = (
+ FSDPCheckpointer(model, save_dir=cfg.train.output_dir)
+ .resume_or_load(cfg.MODEL.WEIGHTS, resume=not args.no_resume)
+ .get("iteration", -1)
+ + 1
+ )
+ return do_test(cfg, model, f"manual_{iteration}")
+
+ do_train(cfg, model, resume=not args.no_resume)
+
+
+if __name__ == "__main__":
+ args = get_args_parser(add_help=True).parse_args()
+ main(args)
diff --git a/models/dsp/dinov2/dinov2/utils/__init__.py b/models/dsp/dinov2/dinov2/utils/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..b88da6bf80be92af00b72dfdb0a806fa64a7a2d9
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/utils/__init__.py
@@ -0,0 +1,4 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
diff --git a/models/dsp/dinov2/dinov2/utils/cluster.py b/models/dsp/dinov2/dinov2/utils/cluster.py
new file mode 100644
index 0000000000000000000000000000000000000000..3df87dc3e1eb4f0f8a280dc3137cfef031886314
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/utils/cluster.py
@@ -0,0 +1,95 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from enum import Enum
+import os
+from pathlib import Path
+from typing import Any, Dict, Optional
+
+
+class ClusterType(Enum):
+ AWS = "aws"
+ FAIR = "fair"
+ RSC = "rsc"
+
+
+def _guess_cluster_type() -> ClusterType:
+ uname = os.uname()
+ if uname.sysname == "Linux":
+ if uname.release.endswith("-aws"):
+ # Linux kernel versions on AWS instances are of the form "5.4.0-1051-aws"
+ return ClusterType.AWS
+ elif uname.nodename.startswith("rsc"):
+ # Linux kernel versions on RSC instances are standard ones but hostnames start with "rsc"
+ return ClusterType.RSC
+
+ return ClusterType.FAIR
+
+
+def get_cluster_type(cluster_type: Optional[ClusterType] = None) -> Optional[ClusterType]:
+ if cluster_type is None:
+ return _guess_cluster_type()
+
+ return cluster_type
+
+
+def get_checkpoint_path(cluster_type: Optional[ClusterType] = None) -> Optional[Path]:
+ cluster_type = get_cluster_type(cluster_type)
+ if cluster_type is None:
+ return None
+
+ CHECKPOINT_DIRNAMES = {
+ ClusterType.AWS: "checkpoints",
+ ClusterType.FAIR: "checkpoint",
+ ClusterType.RSC: "checkpoint/dino",
+ }
+ return Path("/") / CHECKPOINT_DIRNAMES[cluster_type]
+
+
+def get_user_checkpoint_path(cluster_type: Optional[ClusterType] = None) -> Optional[Path]:
+ checkpoint_path = get_checkpoint_path(cluster_type)
+ if checkpoint_path is None:
+ return None
+
+ username = os.environ.get("USER")
+ assert username is not None
+ return checkpoint_path / username
+
+
+def get_slurm_partition(cluster_type: Optional[ClusterType] = None) -> Optional[str]:
+ cluster_type = get_cluster_type(cluster_type)
+ if cluster_type is None:
+ return None
+
+ SLURM_PARTITIONS = {
+ ClusterType.AWS: "learnlab",
+ ClusterType.FAIR: "learnlab",
+ ClusterType.RSC: "learn",
+ }
+ return SLURM_PARTITIONS[cluster_type]
+
+
+def get_slurm_executor_parameters(
+ nodes: int, num_gpus_per_node: int, cluster_type: Optional[ClusterType] = None, **kwargs
+) -> Dict[str, Any]:
+ # create default parameters
+ params = {
+ "mem_gb": 0, # Requests all memory on a node, see https://slurm.schedmd.com/sbatch.html
+ "gpus_per_node": num_gpus_per_node,
+ "tasks_per_node": num_gpus_per_node, # one task per GPU
+ "cpus_per_task": 10,
+ "nodes": nodes,
+ "slurm_partition": get_slurm_partition(cluster_type),
+ }
+ # apply cluster-specific adjustments
+ cluster_type = get_cluster_type(cluster_type)
+ if cluster_type == ClusterType.AWS:
+ params["cpus_per_task"] = 12
+ del params["mem_gb"]
+ elif cluster_type == ClusterType.RSC:
+ params["cpus_per_task"] = 12
+ # set additional parameters / apply overrides
+ params.update(kwargs)
+ return params
diff --git a/models/dsp/dinov2/dinov2/utils/config.py b/models/dsp/dinov2/dinov2/utils/config.py
new file mode 100644
index 0000000000000000000000000000000000000000..c9de578787bbcb376f8bd5a782206d0eb7ec1f52
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/utils/config.py
@@ -0,0 +1,72 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import math
+import logging
+import os
+
+from omegaconf import OmegaConf
+
+import dinov2.distributed as distributed
+from dinov2.logging import setup_logging
+from dinov2.utils import utils
+from dinov2.configs import dinov2_default_config
+
+
+logger = logging.getLogger("dinov2")
+
+
+def apply_scaling_rules_to_cfg(cfg): # to fix
+ if cfg.optim.scaling_rule == "sqrt_wrt_1024":
+ base_lr = cfg.optim.base_lr
+ cfg.optim.lr = base_lr
+ cfg.optim.lr *= math.sqrt(cfg.train.batch_size_per_gpu * distributed.get_global_size() / 1024.0)
+ logger.info(f"sqrt scaling learning rate; base: {base_lr}, new: {cfg.optim.lr}")
+ else:
+ raise NotImplementedError
+ return cfg
+
+
+def write_config(cfg, output_dir, name="config.yaml"):
+ logger.info(OmegaConf.to_yaml(cfg))
+ saved_cfg_path = os.path.join(output_dir, name)
+ with open(saved_cfg_path, "w") as f:
+ OmegaConf.save(config=cfg, f=f)
+ return saved_cfg_path
+
+
+def get_cfg_from_args(args):
+ args.output_dir = os.path.abspath(args.output_dir)
+ args.opts += [f"train.output_dir={args.output_dir}"]
+ default_cfg = OmegaConf.create(dinov2_default_config)
+ cfg = OmegaConf.load(args.config_file)
+ cfg = OmegaConf.merge(default_cfg, cfg, OmegaConf.from_cli(args.opts))
+ return cfg
+
+
+def default_setup(args):
+ distributed.enable(overwrite=True)
+ seed = getattr(args, "seed", 0)
+ rank = distributed.get_global_rank()
+
+ global logger
+ setup_logging(output=args.output_dir, level=logging.INFO)
+ logger = logging.getLogger("dinov2")
+
+ utils.fix_random_seeds(seed + rank)
+ logger.info("git:\n {}\n".format(utils.get_sha()))
+ logger.info("\n".join("%s: %s" % (k, str(v)) for k, v in sorted(dict(vars(args)).items())))
+
+
+def setup(args):
+ """
+ Create configs and perform basic setups.
+ """
+ cfg = get_cfg_from_args(args)
+ os.makedirs(args.output_dir, exist_ok=True)
+ default_setup(args)
+ apply_scaling_rules_to_cfg(cfg)
+ write_config(cfg, args.output_dir)
+ return cfg
diff --git a/models/dsp/dinov2/dinov2/utils/dtype.py b/models/dsp/dinov2/dinov2/utils/dtype.py
new file mode 100644
index 0000000000000000000000000000000000000000..80f4cd74d99faa2731dbe9f8d3a13d71b3f8e3a8
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/utils/dtype.py
@@ -0,0 +1,37 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+
+from typing import Dict, Union
+
+import numpy as np
+import torch
+
+
+TypeSpec = Union[str, np.dtype, torch.dtype]
+
+
+_NUMPY_TO_TORCH_DTYPE: Dict[np.dtype, torch.dtype] = {
+ np.dtype("bool"): torch.bool,
+ np.dtype("uint8"): torch.uint8,
+ np.dtype("int8"): torch.int8,
+ np.dtype("int16"): torch.int16,
+ np.dtype("int32"): torch.int32,
+ np.dtype("int64"): torch.int64,
+ np.dtype("float16"): torch.float16,
+ np.dtype("float32"): torch.float32,
+ np.dtype("float64"): torch.float64,
+ np.dtype("complex64"): torch.complex64,
+ np.dtype("complex128"): torch.complex128,
+}
+
+
+def as_torch_dtype(dtype: TypeSpec) -> torch.dtype:
+ if isinstance(dtype, torch.dtype):
+ return dtype
+ if isinstance(dtype, str):
+ dtype = np.dtype(dtype)
+ assert isinstance(dtype, np.dtype), f"Expected an instance of nunpy dtype, got {type(dtype)}"
+ return _NUMPY_TO_TORCH_DTYPE[dtype]
diff --git a/models/dsp/dinov2/dinov2/utils/param_groups.py b/models/dsp/dinov2/dinov2/utils/param_groups.py
new file mode 100644
index 0000000000000000000000000000000000000000..9a5d2ff627cddadc222e5f836864ee39c865208f
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/utils/param_groups.py
@@ -0,0 +1,103 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+from collections import defaultdict
+import logging
+
+
+logger = logging.getLogger("dinov2")
+
+
+def get_vit_lr_decay_rate(name, lr_decay_rate=1.0, num_layers=12, force_is_backbone=False, chunked_blocks=False):
+ """
+ Calculate lr decay rate for different ViT blocks.
+ Args:
+ name (string): parameter name.
+ lr_decay_rate (float): base lr decay rate.
+ num_layers (int): number of ViT blocks.
+ Returns:
+ lr decay rate for the given parameter.
+ """
+ layer_id = num_layers + 1
+ if name.startswith("backbone") or force_is_backbone:
+ if (
+ ".pos_embed" in name
+ or ".patch_embed" in name
+ or ".mask_token" in name
+ or ".cls_token" in name
+ or ".register_tokens" in name
+ ):
+ layer_id = 0
+ elif force_is_backbone and (
+ "pos_embed" in name
+ or "patch_embed" in name
+ or "mask_token" in name
+ or "cls_token" in name
+ or "register_tokens" in name
+ ):
+ layer_id = 0
+ elif ".blocks." in name and ".residual." not in name:
+ layer_id = int(name[name.find(".blocks.") :].split(".")[2]) + 1
+ elif chunked_blocks and "blocks." in name and "residual." not in name:
+ layer_id = int(name[name.find("blocks.") :].split(".")[2]) + 1
+ elif "blocks." in name and "residual." not in name:
+ layer_id = int(name[name.find("blocks.") :].split(".")[1]) + 1
+
+ return lr_decay_rate ** (num_layers + 1 - layer_id)
+
+
+def get_params_groups_with_decay(model, lr_decay_rate=1.0, patch_embed_lr_mult=1.0):
+ chunked_blocks = False
+ if hasattr(model, "n_blocks"):
+ logger.info("chunked fsdp")
+ n_blocks = model.n_blocks
+ chunked_blocks = model.chunked_blocks
+ elif hasattr(model, "blocks"):
+ logger.info("first code branch")
+ n_blocks = len(model.blocks)
+ elif hasattr(model, "backbone"):
+ logger.info("second code branch")
+ n_blocks = len(model.backbone.blocks)
+ else:
+ logger.info("else code branch")
+ n_blocks = 0
+ all_param_groups = []
+
+ for name, param in model.named_parameters():
+ name = name.replace("_fsdp_wrapped_module.", "")
+ if not param.requires_grad:
+ continue
+ decay_rate = get_vit_lr_decay_rate(
+ name, lr_decay_rate, num_layers=n_blocks, force_is_backbone=n_blocks > 0, chunked_blocks=chunked_blocks
+ )
+ d = {"params": param, "is_last_layer": False, "lr_multiplier": decay_rate, "wd_multiplier": 1.0, "name": name}
+
+ if "last_layer" in name:
+ d.update({"is_last_layer": True})
+
+ if name.endswith(".bias") or "norm" in name or "gamma" in name:
+ d.update({"wd_multiplier": 0.0})
+
+ if "patch_embed" in name:
+ d.update({"lr_multiplier": d["lr_multiplier"] * patch_embed_lr_mult})
+
+ all_param_groups.append(d)
+ logger.info(f"""{name}: lr_multiplier: {d["lr_multiplier"]}, wd_multiplier: {d["wd_multiplier"]}""")
+
+ return all_param_groups
+
+
+def fuse_params_groups(all_params_groups, keys=("lr_multiplier", "wd_multiplier", "is_last_layer")):
+ fused_params_groups = defaultdict(lambda: {"params": []})
+ for d in all_params_groups:
+ identifier = ""
+ for k in keys:
+ identifier += k + str(d[k]) + "_"
+
+ for k in keys:
+ fused_params_groups[identifier][k] = d[k]
+ fused_params_groups[identifier]["params"].append(d["params"])
+
+ return fused_params_groups.values()
diff --git a/models/dsp/dinov2/dinov2/utils/utils.py b/models/dsp/dinov2/dinov2/utils/utils.py
new file mode 100644
index 0000000000000000000000000000000000000000..68f8e2c3be5f780bbb7e00359b5ac4fd0ba0785f
--- /dev/null
+++ b/models/dsp/dinov2/dinov2/utils/utils.py
@@ -0,0 +1,95 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+import logging
+import os
+import random
+import subprocess
+from urllib.parse import urlparse
+
+import numpy as np
+import torch
+from torch import nn
+
+
+logger = logging.getLogger("dinov2")
+
+
+def load_pretrained_weights(model, pretrained_weights, checkpoint_key):
+ if urlparse(pretrained_weights).scheme: # If it looks like an URL
+ state_dict = torch.hub.load_state_dict_from_url(pretrained_weights, map_location="cpu")
+ else:
+ state_dict = torch.load(pretrained_weights, map_location="cpu")
+ if checkpoint_key is not None and checkpoint_key in state_dict:
+ logger.info(f"Take key {checkpoint_key} in provided checkpoint dict")
+ state_dict = state_dict[checkpoint_key]
+ # remove `module.` prefix
+ state_dict = {k.replace("module.", ""): v for k, v in state_dict.items()}
+ # remove `backbone.` prefix induced by multicrop wrapper
+ state_dict = {k.replace("backbone.", ""): v for k, v in state_dict.items()}
+ msg = model.load_state_dict(state_dict, strict=False)
+ logger.info("Pretrained weights found at {} and loaded with msg: {}".format(pretrained_weights, msg))
+
+
+def fix_random_seeds(seed=31):
+ """
+ Fix random seeds.
+ """
+ torch.manual_seed(seed)
+ torch.cuda.manual_seed_all(seed)
+ np.random.seed(seed)
+ random.seed(seed)
+
+
+def get_sha():
+ cwd = os.path.dirname(os.path.abspath(__file__))
+
+ def _run(command):
+ return subprocess.check_output(command, cwd=cwd).decode("ascii").strip()
+
+ sha = "N/A"
+ diff = "clean"
+ branch = "N/A"
+ try:
+ sha = _run(["git", "rev-parse", "HEAD"])
+ subprocess.check_output(["git", "diff"], cwd=cwd)
+ diff = _run(["git", "diff-index", "HEAD"])
+ diff = "has uncommitted changes" if diff else "clean"
+ branch = _run(["git", "rev-parse", "--abbrev-ref", "HEAD"])
+ except Exception:
+ pass
+ message = f"sha: {sha}, status: {diff}, branch: {branch}"
+ return message
+
+
+class CosineScheduler(object):
+ def __init__(self, base_value, final_value, total_iters, warmup_iters=0, start_warmup_value=0, freeze_iters=0):
+ super().__init__()
+ self.final_value = final_value
+ self.total_iters = total_iters
+
+ freeze_schedule = np.zeros((freeze_iters))
+
+ warmup_schedule = np.linspace(start_warmup_value, base_value, warmup_iters)
+
+ iters = np.arange(total_iters - warmup_iters - freeze_iters)
+ schedule = final_value + 0.5 * (base_value - final_value) * (1 + np.cos(np.pi * iters / len(iters)))
+ self.schedule = np.concatenate((freeze_schedule, warmup_schedule, schedule))
+
+ assert len(self.schedule) == self.total_iters
+
+ def __getitem__(self, it):
+ if it >= self.total_iters:
+ return self.final_value
+ else:
+ return self.schedule[it]
+
+
+def has_batchnorms(model):
+ bn_types = (nn.BatchNorm1d, nn.BatchNorm2d, nn.BatchNorm3d, nn.SyncBatchNorm)
+ for name, module in model.named_modules():
+ if isinstance(module, bn_types):
+ return True
+ return False
diff --git a/models/dsp/dinov2/hubconf.py b/models/dsp/dinov2/hubconf.py
new file mode 100644
index 0000000000000000000000000000000000000000..5dd3e0e5b47cd3dbaa803947bec99ade6794d77c
--- /dev/null
+++ b/models/dsp/dinov2/hubconf.py
@@ -0,0 +1,15 @@
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+#
+# This source code is licensed under the Apache License, Version 2.0
+# found in the LICENSE file in the root directory of this source tree.
+
+
+from .dinov2.hub.backbones import dinov2_vitb14, dinov2_vitg14, dinov2_vitl14, dinov2_vits14
+from .dinov2.hub.backbones import dinov2_vitb14_reg, dinov2_vitg14_reg, dinov2_vitl14_reg, dinov2_vits14_reg
+from .dinov2.hub.classifiers import dinov2_vitb14_lc, dinov2_vitg14_lc, dinov2_vitl14_lc, dinov2_vits14_lc
+from .dinov2.hub.classifiers import dinov2_vitb14_reg_lc, dinov2_vitg14_reg_lc, dinov2_vitl14_reg_lc, dinov2_vits14_reg_lc
+from .dinov2.hub.depthers import dinov2_vitb14_ld, dinov2_vitg14_ld, dinov2_vitl14_ld, dinov2_vits14_ld
+from .dinov2.hub.depthers import dinov2_vitb14_dd, dinov2_vitg14_dd, dinov2_vitl14_dd, dinov2_vits14_dd
+from .dinov2.hub.dinotxt import dinov2_vitl14_reg4_dinotxt_tet1280d20h24l
+
+dependencies = ["torch"]
diff --git a/models/dsp/kmeans_pytorch/__init__.py b/models/dsp/kmeans_pytorch/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..51ddb256dfcd9558f487dc0be365873e1ef1d028
--- /dev/null
+++ b/models/dsp/kmeans_pytorch/__init__.py
@@ -0,0 +1,186 @@
+from functools import partial
+
+import numpy as np
+import torch
+from tqdm import tqdm
+
+def initialize(X, num_clusters, seed):
+ """
+ initialize cluster centers
+ :param X: (torch.tensor) matrix
+ :param num_clusters: (int) number of clusters
+ :param seed: (int) seed for kmeans
+ :return: (np.array) initial state
+ """
+ num_samples = len(X)
+ if seed == None:
+ indices = np.random.choice(num_samples, num_clusters, replace=False)
+ else:
+ np.random.seed(seed) ; indices = np.random.choice(num_samples, num_clusters, replace=False)
+ initial_state = X[indices]
+ return initial_state
+
+
+def kmeans(
+ X,
+ num_clusters,
+ distance='euclidean',
+ cluster_centers=[],
+ tol=1e-4,
+ tqdm_flag=False,
+ iter_limit=0,
+ device=torch.device('cpu'),
+ gamma_for_soft_dtw=0.001,
+ seed=None,
+):
+ """
+ perform kmeans
+ :param X: (torch.tensor) matrix
+ :param num_clusters: (int) number of clusters
+ :param distance: (str) distance [options: 'euclidean', 'cosine'] [default: 'euclidean']
+ :param seed: (int) seed for kmeans
+ :param tol: (float) threshold [default: 0.0001]
+ :param device: (torch.device) device [default: cpu]
+ :param tqdm_flag: Allows to turn logs on and off
+ :param iter_limit: hard limit for max number of iterations
+ :param gamma_for_soft_dtw: approaches to (hard) DTW as gamma -> 0
+ :return: (torch.tensor, torch.tensor) cluster ids, cluster centers
+ """
+ if tqdm_flag:
+ print(f'running k-means on {device}..')
+
+ if distance == 'euclidean':
+ pairwise_distance_function = partial(pairwise_distance, device=device, tqdm_flag=tqdm_flag)
+ elif distance == 'cosine':
+ pairwise_distance_function = partial(pairwise_cosine, device=device)
+ else:
+ raise NotImplementedError
+
+ X = X.float()
+
+ X = X.to(device)
+
+ if type(cluster_centers) == list: # ToDo: make this less annoyingly weird
+ initial_state = initialize(X, num_clusters, seed=seed)
+ else:
+ if tqdm_flag:
+ print('resuming')
+ initial_state = cluster_centers
+ dis = pairwise_distance_function(X, initial_state)
+ choice_points = torch.argmin(dis, dim=0)
+ initial_state = X[choice_points]
+ initial_state = initial_state.to(device)
+
+ iteration = 0
+ if tqdm_flag:
+ tqdm_meter = tqdm(desc='[running kmeans]')
+ while True:
+
+ dis = pairwise_distance_function(X, initial_state)
+
+ choice_cluster = torch.argmin(dis, dim=1)
+
+ initial_state_pre = initial_state.clone()
+
+ for index in range(num_clusters):
+ selected = torch.nonzero(choice_cluster == index).squeeze().to(device)
+
+ selected = torch.index_select(X, 0, selected)
+
+ # https://github.com/subhadarship/kmeans_pytorch/issues/16
+ if selected.shape[0] == 0:
+ selected = X[torch.randint(len(X), (1,))]
+
+ initial_state[index] = selected.mean(dim=0)
+
+ center_shift = torch.sum(
+ torch.sqrt(
+ torch.sum((initial_state - initial_state_pre) ** 2, dim=1)
+ ))
+
+ # increment iteration
+ iteration = iteration + 1
+
+ # update tqdm meter
+ if tqdm_flag:
+ tqdm_meter.set_postfix(
+ iteration=f'{iteration}',
+ center_shift=f'{center_shift ** 2:0.6f}',
+ tol=f'{tol:0.6f}'
+ )
+ tqdm_meter.update()
+ if center_shift ** 2 < tol:
+ break
+ if iter_limit != 0 and iteration >= iter_limit:
+ break
+
+ return choice_cluster.cpu(), initial_state.cpu()
+
+
+def kmeans_predict(
+ X,
+ cluster_centers,
+ distance='euclidean',
+ device=torch.device('cpu'),
+ gamma_for_soft_dtw=0.001,
+ tqdm_flag=True
+):
+ """
+ predict using cluster centers
+ :param X: (torch.tensor) matrix
+ :param cluster_centers: (torch.tensor) cluster centers
+ :param distance: (str) distance [options: 'euclidean', 'cosine'] [default: 'euclidean']
+ :param device: (torch.device) device [default: 'cpu']
+ :param gamma_for_soft_dtw: approaches to (hard) DTW as gamma -> 0
+ :return: (torch.tensor) cluster ids
+ """
+ if tqdm_flag:
+ print(f'predicting on {device}..')
+
+ if distance == 'euclidean':
+ pairwise_distance_function = partial(pairwise_distance, device=device, tqdm_flag=tqdm_flag)
+ elif distance == 'cosine':
+ pairwise_distance_function = partial(pairwise_cosine, device=device)
+ else:
+ raise NotImplementedError
+
+ X = X.float()
+
+ X = X.to(device)
+
+ dis = pairwise_distance_function(X, cluster_centers)
+ choice_cluster = torch.argmin(dis, dim=1)
+
+ return choice_cluster.cpu()
+
+
+def pairwise_distance(data1, data2, device=torch.device('cpu'), tqdm_flag=True):
+ if tqdm_flag:
+ print(f'device is :{device}')
+
+ data1, data2 = data1.to(device), data2.to(device)
+
+ A = data1.unsqueeze(dim=1)
+
+ B = data2.unsqueeze(dim=0)
+
+ dis = (A - B) ** 2.0
+
+ dis = dis.sum(dim=-1).squeeze()
+ return dis
+
+
+def pairwise_cosine(data1, data2, device=torch.device('cpu')):
+ data1, data2 = data1.to(device), data2.to(device)
+
+ A = data1.unsqueeze(dim=1)
+
+ B = data2.unsqueeze(dim=0)
+
+ A_normalized = A / A.norm(dim=-1, keepdim=True)
+ B_normalized = B / B.norm(dim=-1, keepdim=True)
+
+ cosine = A_normalized * B_normalized
+
+ cosine_dis = 1 - cosine.sum(dim=-1).squeeze()
+ return cosine_dis
diff --git a/models/dsp/layers.py b/models/dsp/layers.py
new file mode 100644
index 0000000000000000000000000000000000000000..7d6edb68a65ca4169a1988b0ef3c7ae9ab822ed1
--- /dev/null
+++ b/models/dsp/layers.py
@@ -0,0 +1,422 @@
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+import math
+from inspect import isfunction
+from einops import rearrange, repeat
+from torch import einsum
+
+def exists(val):
+ return val is not None
+
+def default(val, d):
+ if exists(val):
+ return val
+ return d() if isfunction(d) else d
+
+class CrossAttention(nn.Module):
+ def __init__(self, query_dim, context_dim=None, heads=8, dim_head=64, dropout=0.):
+ super().__init__()
+ inner_dim = dim_head * heads
+ context_dim = default(context_dim, query_dim)
+
+ self.scale = dim_head ** -0.5
+ self.heads = heads
+
+ self.to_q = nn.Linear(query_dim, inner_dim, bias=False)
+ self.to_k = nn.Linear(context_dim, inner_dim, bias=False)
+ self.to_v = nn.Linear(context_dim, inner_dim, bias=False)
+
+ self.to_out = nn.Sequential(
+ nn.Linear(inner_dim, query_dim),
+ nn.Dropout(dropout)
+ )
+
+ def forward(self, x, context=None, mask=None, return_attn=False, need_softmax=True):
+ h = self.heads
+ b = x.shape[0]
+
+ q = self.to_q(x) # [15, 64, 1280]
+ context = default(context, x)
+ k = self.to_k(context) # [15, 18, 1280]
+ v = self.to_v(context) # [15, 18, 1280]
+
+ q, k, v = map(lambda t: rearrange(t, 'b n (h d) -> (b h) n d', h=h), (q, k, v))
+
+ sim = einsum('b i d, b j d -> b i j', q, k) * self.scale
+
+ if exists(mask):
+ mask = rearrange(mask, 'b ... -> b (...)')
+ max_neg_value = -torch.finfo(sim.dtype).max
+ mask = repeat(mask, 'b j -> (b h) () j', h=h)
+ sim.masked_fill_(~mask, max_neg_value)
+
+ if need_softmax:
+ attn = sim.softmax(dim=-1)
+ else:
+ attn = sim
+
+ out = einsum('b i j, b j d -> b i d', attn, v)
+ out = rearrange(out, '(b h) n d -> b n (h d)', h=h)
+ if return_attn:
+ attn = attn.view(b, h, attn.shape[-2], attn.shape[-1])
+ return self.to_out(out), attn
+ else:
+ return self.to_out(out)
+
+class FourierEmbedder():
+ def __init__(self, num_freqs=64, temperature=100):
+ self.num_freqs = num_freqs
+ self.temperature = temperature
+ self.freq_bands = temperature ** ( torch.arange(num_freqs) / num_freqs )
+
+ @ torch.no_grad()
+ def __call__(self, x, cat_dim=-1):
+ out = []
+ for freq in self.freq_bands:
+ out.append( torch.sin( freq*x ) )
+ out.append( torch.cos( freq*x ) )
+ return torch.cat(out, cat_dim) # torch.Size([5, 30, 64])
+
+class PositionNet(nn.Module):
+ def __init__(self, in_dim, out_dim, fourier_freqs=8):
+ super().__init__()
+ self.in_dim = in_dim
+ self.out_dim = out_dim
+
+ self.fourier_embedder = FourierEmbedder(num_freqs=fourier_freqs)
+ self.position_dim = fourier_freqs * 2 * 8 # 2 is sin&cos, 8 is xyxyxyxy
+
+ # -------------------------------------------------------------- #
+ self.linears_position = nn.Sequential(
+ nn.Linear(self.position_dim, 512),
+ nn.SiLU(),
+ nn.Linear(512, 512),
+ nn.SiLU(),
+ nn.Linear(512, out_dim),
+ )
+
+ def forward(self, boxes):
+
+ # embedding position (it may includes padding as placeholder)
+ xyxy_embedding = self.fourier_embedder(boxes) # B*1*4 --> B*1*C torch.Size([5, 1, 64])
+ xyxy_embedding = self.linears_position(xyxy_embedding) # B*1*C --> B*1*768 torch.Size([5, 1, 768])
+
+ return xyxy_embedding
+
+class LayoutAttention(nn.Module):
+ def __init__(self, query_dim, context_dim=None, heads=8, dim_head=64, dropout=0., use_lora=False):
+ super().__init__()
+ inner_dim = dim_head * heads
+ context_dim = default(context_dim, query_dim)
+
+ self.use_lora = use_lora
+ self.scale = dim_head ** -0.5
+ self.heads = heads
+
+ self.to_q = nn.Linear(query_dim, inner_dim, bias=False)
+ self.to_k = nn.Linear(context_dim, inner_dim, bias=False)
+ self.to_v = nn.Linear(context_dim, inner_dim, bias=False)
+
+ self.to_out = nn.Sequential(
+ nn.Linear(inner_dim, query_dim),
+ nn.Dropout(dropout)
+ )
+
+ def forward(self, x, context=None, mask=None, return_attn=False, need_softmax=True, guidance_mask=None):
+ h = self.heads
+ b = x.shape[0]
+
+ q = self.to_q(x)
+ context = default(context, x)
+ k = self.to_k(context)
+ v = self.to_v(context)
+
+ q, k, v = map(lambda t: rearrange(t, 'b n (h d) -> (b h) n d', h=h), (q, k, v))
+
+ sim = einsum('b i d, b j d -> b i j', q, k) * self.scale
+
+ _, phase_num, H, W = guidance_mask.shape
+ HW = H * W
+ guidance_mask_o = guidance_mask.view(b * phase_num, HW, 1)
+ guidance_mask_t = guidance_mask.view(b * phase_num, 1, HW)
+ guidance_mask_sim = torch.bmm(guidance_mask_o, guidance_mask_t) # (B * phase_num, HW, HW)
+ guidance_mask_sim = guidance_mask_sim.view(b, phase_num, HW, HW).sum(dim=1)
+ guidance_mask_sim[guidance_mask_sim > 1] = 1 # (B, HW, HW)
+ guidance_mask_sim = guidance_mask_sim.view(b, 1, HW, HW)
+ guidance_mask_sim = guidance_mask_sim.repeat(1, self.heads, 1, 1)
+ guidance_mask_sim = guidance_mask_sim.view(b * self.heads, HW, HW) # (B * head, HW, HW)
+
+ sim[:, :, :HW][guidance_mask_sim == 0] = -torch.finfo(sim.dtype).max
+
+ if exists(mask):
+ mask = rearrange(mask, 'b ... -> b (...)')
+ max_neg_value = -torch.finfo(sim.dtype).max
+ mask = repeat(mask, 'b j -> (b h) () j', h=h)
+ sim.masked_fill_(~mask, max_neg_value)
+
+ # attention, what we cannot get enough of
+
+ if need_softmax:
+ attn = sim.softmax(dim=-1)
+ else:
+ attn = sim
+
+ out = einsum('b i j, b j d -> b i d', attn, v)
+ out = rearrange(out, '(b h) n d -> b n (h d)', h=h)
+ if return_attn:
+ attn = attn.view(b, h, attn.shape[-2], attn.shape[-1])
+ return self.to_out(out), attn
+ else:
+ return self.to_out(out)
+
+# feedforward
+class GEGLU(nn.Module):
+ def __init__(self, dim_in, dim_out):
+ super().__init__()
+ self.proj = nn.Linear(dim_in, dim_out * 2)
+
+ def forward(self, x):
+ x, gate = self.proj(x).chunk(2, dim=-1)
+ return x * F.gelu(gate)
+
+class FeedForward(nn.Module):
+ def __init__(self, dim, dim_out=None, mult=4, glu=False, dropout=0.):
+ super().__init__()
+ inner_dim = int(dim * mult)
+ dim_out = default(dim_out, dim)
+ project_in = nn.Sequential(
+ nn.Linear(dim, inner_dim),
+ nn.GELU()
+ ) if not glu else GEGLU(dim, inner_dim)
+
+ self.net = nn.Sequential(
+ project_in,
+ nn.Dropout(dropout),
+ nn.Linear(inner_dim, dim_out)
+ )
+
+ def forward(self, x):
+ return self.net(x)
+
+class SelfAttention(nn.Module):
+ def __init__(self, query_dim, heads=8, dim_head=64, dropout=0.):
+ super().__init__()
+ inner_dim = dim_head * heads
+ self.scale = dim_head ** -0.5
+ self.heads = heads
+
+ self.to_q = nn.Linear(query_dim, inner_dim, bias=False)
+ self.to_k = nn.Linear(query_dim, inner_dim, bias=False)
+ self.to_v = nn.Linear(query_dim, inner_dim, bias=False)
+
+ self.to_out = nn.Sequential(nn.Linear(inner_dim, query_dim), nn.Dropout(dropout) )
+
+ def forward(self, x):
+ q = self.to_q(x) # B*N*(H*C)
+ k = self.to_k(x) # B*N*(H*C)
+ v = self.to_v(x) # B*N*(H*C)
+
+ B, N, HC = q.shape
+ H = self.heads
+ C = HC // H
+
+ q = q.view(B,N,H,C).permute(0,2,1,3).reshape(B*H,N,C) # (B*H)*N*C
+ k = k.view(B,N,H,C).permute(0,2,1,3).reshape(B*H,N,C) # (B*H)*N*C
+ v = v.view(B,N,H,C).permute(0,2,1,3).reshape(B*H,N,C) # (B*H)*N*C
+
+ sim = torch.einsum('b i c, b j c -> b i j', q, k) * self.scale # (B*H)*N*N
+ attn = sim.softmax(dim=-1) # (B*H)*N*N
+
+ out = torch.einsum('b i j, b j c -> b i c', attn, v) # (B*H)*N*C
+ out = out.view(B,H,N,C).permute(0,2,1,3).reshape(B,N,(H*C)) # B*N*(H*C)
+
+ return self.to_out(out)
+
+class GatedSelfAttentionDense(nn.Module):
+ def __init__(self, query_dim, context_dim, n_heads, d_head):
+ super().__init__()
+
+ # we need a linear projection since we need cat visual feature and obj feature
+ self.linear = nn.Linear(context_dim, query_dim)
+
+ self.attn = SelfAttention(query_dim=query_dim, heads=n_heads, dim_head=d_head)
+ self.ff = FeedForward(query_dim, glu=True)
+
+ self.norm1 = nn.LayerNorm(query_dim)
+ self.norm2 = nn.LayerNorm(query_dim)
+
+ self.register_parameter('alpha_attn', nn.Parameter(torch.tensor(0.)) )
+ self.register_parameter('alpha_dense', nn.Parameter(torch.tensor(0.)) )
+
+ # this can be useful: we can externally change magnitude of tanh(alpha)
+ # for example, when it is set to 0, then the entire model is same as original one
+ self.scale = 1
+
+
+ def forward(self, x, objs):
+
+ N_visual = x.shape[1]
+ objs = self.linear(objs)
+
+ x = x + self.scale*torch.tanh(self.alpha_attn) * self.attn( self.norm1(torch.cat([x,objs],dim=1)) )[:,0:N_visual,:]
+ x = x + self.scale*torch.tanh(self.alpha_dense) * self.ff( self.norm2(x) )
+
+ return x
+
+class MIFusion(nn.Module):
+ def __init__(self, C, attn_type='base', context_dim=768, heads=8):
+ # context_dim: SD1.4 768 SD2.1 1024
+ super().__init__()
+ self.ea_obj = CrossAttention(query_dim=C, context_dim=context_dim,
+ heads=heads, dim_head=C // heads,
+ dropout=0.0)
+ self.norm_obj = nn.LayerNorm(C)
+ self.ea2 = CrossAttention(query_dim=C, context_dim=context_dim,
+ heads=heads, dim_head=C // heads,
+ dropout=0.0)
+ self.norm2 = nn.LayerNorm(C)
+ self.pos_net = PositionNet(in_dim=context_dim, out_dim=context_dim)
+ self.la = LayoutAttention(query_dim=C, heads=heads,
+ dim_head=C // heads, dropout=0.0)
+
+ def forward(self, ca_x, other_info):
+ # x: (B, instance_num+1, HW, C)
+ # guidance_mask: (B, instance_num, H, W)
+ # box: (instance_num, 4)
+ # image_token: (B, instance_num+1, HW, C)
+
+ # Reminder for shapes
+ # ca_x # [1, 16, 64, 1280] # After Cross-Attn with encoder_hidden_states
+ # guidance_mask # [1, 15, 64, 64]
+ # other_info['image_token'] # [1, 16, 64, 1280] # Original hidden_states repeated 1+instance_num times
+ # other_info['context'] # [16, 77, 768]
+ # other_info['box'] # [1, 15, 8]
+ # other_info['context_pooler'] # [16, 1, 768]
+ # other_info['supplement_mask'] # [1, 1, 64, 64]
+ # other_info['height'] = height # 512
+ # other_info['width'] = width # 512
+ # other_info['ref_features'] # [15, 16, 768], [1, 16, 768]
+ # other_info['sigmoid_values'] # [1, 16, 768]
+
+ height, width = other_info['height'], other_info['width'] # 512, 512
+ instance_num = other_info['instance_num'] # 15
+ B, _, HW, C = ca_x.shape # [1, 16, 64, 1280]
+ down_scale = int(math.sqrt(height * width // ca_x.shape[2])) # 64
+ H = height // down_scale # 8
+ W = width // down_scale # 8
+
+ guidance_masks = other_info['guidance_masks']
+ guidance_masks = F.interpolate(guidance_masks, size=(H, W), mode='bilinear') # [1, 15, 8, 8]
+
+ supplement_mask = other_info['supplement_mask'] # (1, 1, 64, 64)
+ supplement_mask = F.interpolate(supplement_mask, size=(H, W), mode='bilinear') # (1, 1, 8, 8)
+
+ image_token = other_info['image_token'] # [1, 16, 64, 1280]
+ assert image_token.shape == ca_x.shape
+
+ context_pooler = other_info['context_pooler'] # [16, 1, 768]
+ box = other_info['box'] # [1, 15, 8]
+ box = box.view(B * instance_num, 1, -1) # [15, 1, 8]
+ box_token = self.pos_net(box) # [15, 1, 768]
+
+ # add reference image feature as condition
+ img_features, bg_features = other_info['ref_features'] # [15, 16, 768], [1, 16, 768]
+
+ context_fg = torch.cat([context_pooler[1:, ...], img_features, box_token], dim=1) # [15, 1 (text) + 16 (ref) + 1 (box), 768]
+ ea_x, anchor_attn = self.ea_obj(self.norm_obj(image_token[:, 1:, ...].view(B * instance_num, HW, C)), # [15, 64, 1280]
+ context=context_fg, return_attn=True) # ea_x.shape [15, 64, 1280]
+ ea_x = ea_x.view(B, instance_num, HW, C) # ea_x.shape [1, 15, 64, 1280]
+ sigmoid_values = other_info['sigmoid_values'] # [1, 15, 64, 64]
+ sigmoid_values = F.interpolate(sigmoid_values, size=(H, W), mode='bilinear') # [1, 15, 8, 8]
+ ea_x = ea_x * sigmoid_values.view(B, instance_num, HW, 1) # [1, 15, 64, 1280] * [1, 15, 64, 1]
+ ca_x[:, 1:, ...] = ca_x[:, 1:, ...] * sigmoid_values.view(B, instance_num, HW, 1) # (B, phase_num, HW, C)
+ ca_x[:, 1:, ...] = ca_x[:, 1:, ...] + ea_x # ca_x.shape [1, 16, 64, 1280]
+
+ context_bg = torch.cat([context_pooler[[0], ...], bg_features], dim=1) # [1, 1 (text) + 16 (ref), 768]
+ ea_x_bg, _ = self.ea2(self.norm2(ca_x[:, 1:, ...].sum(dim=1, keepdim=True).view(B * 1, HW, C)),
+ context=context_bg, return_attn=True) # [1, 64, 1280]
+ ca_x[:, 0, ...] = ca_x[:, 0, ...] + ea_x_bg # [1, 64, 1280]
+
+ # image_token[:, 0, ...].shape [1, 64, 1280] ; torch.cat([guidance_mask[:, :, ...], supplement_mask], dim=1).shape [1, 1 + 15, 8, 8]
+ fusion_template = self.la(x=image_token[:, 0, ...], guidance_mask=torch.cat([guidance_masks[:, :, ...], supplement_mask], dim=1)) # [1, 64, 1280]
+ fusion_template = fusion_template.view(B, 1, HW, C) # [1, 1, 64, 1280]
+ ca_x = torch.cat([ca_x, fusion_template], dim = 1) # [1, 17, 64, 1280]
+ out = torch.sum(ca_x, dim=1) # [1, 64, 1280]
+ return out
+
+class MIFusionPrototype(nn.Module):
+ def __init__(self, C, attn_type='base', context_dim=768, heads=8, prototype_dim=1024):
+ super().__init__()
+ self.prototype_attn = CrossAttention(query_dim=C, context_dim=prototype_dim,
+ heads=heads, dim_head=C // heads,
+ dropout=0.0)
+ self.norm_prototype = nn.LayerNorm(C)
+ self.la = LayoutAttention(query_dim=C, heads=heads,
+ dim_head=C // heads, dropout=0.0)
+ self.gating_param = nn.Parameter(torch.zeros(1))
+ self._init_novel_weights()
+
+ def _init_novel_weights(self):
+ if hasattr(self.prototype_attn, 'to_out'):
+ output_layer = self.prototype_attn.to_out[0] if isinstance(self.prototype_attn.to_out, nn.Sequential) or isinstance(self.prototype_attn.to_out, nn.ModuleList) else self.prototype_attn.to_out
+ nn.init.zeros_(output_layer.weight)
+ nn.init.zeros_(output_layer.bias)
+ if hasattr(self.prototype_attn, 'to_q'):
+ nn.init.xavier_uniform_(self.prototype_attn.to_q.weight)
+ if hasattr(self.prototype_attn, 'to_k'):
+ nn.init.xavier_uniform_(self.prototype_attn.to_k.weight)
+ if hasattr(self.prototype_attn, 'to_v'):
+ nn.init.xavier_uniform_(self.prototype_attn.to_v.weight)
+ if hasattr(self.la, 'to_out'):
+ output_layer = self.la.to_out[0] if isinstance(self.la.to_out, nn.Sequential) or isinstance(self.la.to_out, nn.ModuleList) else self.la.to_out
+ nn.init.zeros_(output_layer.weight)
+ nn.init.zeros_(output_layer.bias)
+
+ def forward(self, ca_x, other_info):
+ height, width = other_info['height'], other_info['width'] # 512, 512
+ instance_num = other_info['instance_num'] # 15
+ B, _, HW, C = ca_x.shape # [1, 16, 64, 1280]
+ down_scale = int(math.sqrt(height * width // ca_x.shape[2])) # 64
+ H = height // down_scale # 8
+ W = width // down_scale # 8
+
+ guidance_masks = other_info['guidance_masks']
+ guidance_masks = F.interpolate(guidance_masks, size=(H, W), mode='bilinear') # [1, 15, 8, 8]
+
+ supplement_mask = other_info['supplement_mask'] # (1, 1, 64, 64)
+ supplement_mask = F.interpolate(supplement_mask, size=(H, W), mode='bilinear') # (1, 1, 8, 8)
+
+ sigmoid_values = other_info['sigmoid_values'] # [1, 15, 64, 64]
+ sigmoid_values = F.interpolate(sigmoid_values, size=(H, W), mode='bilinear') # [1, 15, 8, 8]
+
+ image_token = other_info['image_token'] # [1, 16, 64, 1280]
+ assert image_token.shape == ca_x.shape
+
+ prototypes = other_info['prototypes'] # [15, 64, 768]
+
+ base, side_feats_input = ca_x[:, 0, ...].clone(), ca_x[:, 1:, ...].clone()
+
+ mask_flat = sigmoid_values.view(B * instance_num, -1)
+ x_flat = image_token[:, 1:, ...].reshape(B * instance_num, HW, C)
+
+ keep_k = 256
+ _, topk_indices = torch.topk(mask_flat, k=keep_k, dim=1)
+ gather_indices = topk_indices.unsqueeze(-1).expand(-1, -1, C)
+ x_selected = torch.gather(x_flat, 1, gather_indices)
+ ea_selected, primitive_attn = self.prototype_attn(self.norm_prototype(x_selected), context=prototypes, return_attn=True)
+ ea_full = torch.zeros_like(x_flat)
+ ea_full.scatter_(1, gather_indices, ea_selected)
+ ea_x = ea_full.view(B, instance_num, HW, C)
+
+ ea_x = ea_x * sigmoid_values.view(B, instance_num, HW, 1) # [1, 15, 64, 1280] * [1, 15, 64, 1]
+ side_feats_gated = side_feats_input * sigmoid_values.view(B, instance_num, HW, 1) # (B, phase_num, HW, C)
+ total_side_residual = side_feats_gated + ea_x # ca_x.shape [1, 16, 64, 1280]
+
+ # image_token[:, 0, ...].shape [1, 64, 1280] ; torch.cat([guidance_mask[:, :, ...], supplement_mask], dim=1).shape [1, 1 + 15, 8, 8]
+ fusion_template = self.la(x=image_token[:, 0, ...], guidance_mask=torch.cat([guidance_masks[:, :, ...], supplement_mask], dim=1)) # [1, 64, 1280]
+ fusion_template = fusion_template.view(B, 1, HW, C) # [1, 1, 64, 1280]
+
+ final_residual = torch.sum(total_side_residual, dim=1) + fusion_template.squeeze(1)
+ out = base + self.gating_param * final_residual
+ return out
\ No newline at end of file
diff --git a/models/dsp/modules.py b/models/dsp/modules.py
new file mode 100644
index 0000000000000000000000000000000000000000..70ba6b03464c8c167087020b83f374f151bda64a
--- /dev/null
+++ b/models/dsp/modules.py
@@ -0,0 +1,60 @@
+import sys
+
+import torch
+import torch.nn as nn
+
+from .dinov2 import hubconf
+
+class AbstractEncoder(nn.Module):
+ def __init__(self):
+ super().__init__()
+
+ def encode(self, *args, **kwargs):
+ raise NotImplementedError
+
+class FrozenDinoV2Encoder(AbstractEncoder):
+ """
+ Uses the DINOv2 encoder for image
+ """
+ def __init__(self, weight_path, device="cpu", freeze=True):
+ super().__init__()
+ dinov2 = hubconf.dinov2_vitl14(pretrained=False)
+ state_dict = torch.load(weight_path)
+ dinov2.load_state_dict(state_dict, strict=False)
+ self.model = dinov2.to(device)
+ # self.device = device
+ if freeze:
+ self.freeze()
+ self.register_buffer('image_mean', torch.tensor([0.485, 0.456, 0.406]).view(1, 3, 1, 1))
+ self.register_buffer('image_std', torch.tensor([0.229, 0.224, 0.225]).view(1, 3, 1, 1))
+ # self.projector = nn.Linear(1536, 768)
+
+ @property
+ def dtype(self):
+ return next(self.model.parameters()).dtype
+
+ def freeze(self):
+ self.model.eval()
+ for param in self.model.parameters():
+ param.requires_grad = False
+
+ # image.shape [15, 3, 224, 224]
+ def forward(self, image, mode=None):
+ if isinstance(image,list):
+ image = torch.cat(image,0)
+
+ image = (image - self.image_mean) / self.image_std
+ features = self.model.forward_features(image) # dict_keys(['x_norm_clstoken', 'x_norm_regtokens', 'x_norm_patchtokens', 'x_prenorm', 'masks'])
+
+ if mode is not None:
+ return features[mode]
+
+ tokens = features["x_norm_patchtokens"] # [15, 256, 1024]
+ image_features = features["x_norm_clstoken"] # [15, 1024]
+ image_features = image_features.unsqueeze(1) # [15, 1, 1024]
+ hint = torch.cat([image_features,tokens],1) # [15, 257, 1024]
+ # hint = self.projector(hint)
+ return hint
+
+ def encode(self, image):
+ return self(image)
\ No newline at end of file
diff --git a/models/dsp/projection.py b/models/dsp/projection.py
new file mode 100644
index 0000000000000000000000000000000000000000..57343e00368cf1bec9cde0776ce0c3aee6d2357d
--- /dev/null
+++ b/models/dsp/projection.py
@@ -0,0 +1,368 @@
+import math
+import numpy as np
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from einops import rearrange
+from einops.layers.torch import Rearrange
+from .layers import PositionNet, GatedSelfAttentionDense, CrossAttention
+from .prototype_bank import PrototypeBank
+
+from datamodules import RefTable
+from PIL import Image
+import os
+
+from utils import Dict
+
+class FourierEmbedder(nn.Module):
+ def __init__(self, num_freqs=64, temperature=100):
+ super().__init__()
+
+ self.num_freqs = num_freqs
+ self.temperature = temperature
+
+ freq_bands = temperature ** (torch.arange(num_freqs) / num_freqs)
+ freq_bands = freq_bands[None, None]
+ self.register_buffer("freq_bands", freq_bands, persistent=False)
+
+ def __call__(self, x):
+ x = self.freq_bands * x.unsqueeze(-1)
+ return torch.stack((x.sin(), x.cos()), dim=-1).permute(0, 2, 3, 1).reshape(x.shape[0], -1)
+
+# FFN
+def FeedForward(dim, mult=4):
+ inner_dim = int(dim * mult)
+ return nn.Sequential(
+ nn.LayerNorm(dim),
+ nn.Linear(dim, inner_dim, bias=False),
+ nn.GELU(),
+ nn.Linear(inner_dim, dim, bias=False),
+ )
+
+
+def reshape_tensor(x, heads):
+ bs, length, width = x.shape
+ # (bs, length, width) --> (bs, length, n_heads, dim_per_head)
+ x = x.view(bs, length, heads, -1)
+ # (bs, length, n_heads, dim_per_head) --> (bs, n_heads, length, dim_per_head)
+ x = x.transpose(1, 2)
+ # (bs, n_heads, length, dim_per_head) --> (bs*n_heads, length, dim_per_head)
+ x = x.reshape(bs, heads, length, -1)
+ return x
+
+
+class SelfAttentionLayer(nn.Module):
+ def __init__(self, channels, nhead, dropout=0.0):
+ super().__init__()
+ self.norm1 = nn.LayerNorm(channels)
+ self.self_attn = nn.MultiheadAttention(channels, nhead, dropout=dropout)
+
+ self.norm2 = nn.LayerNorm(channels)
+ self.dropout = nn.Dropout(dropout)
+
+ def forward(self,
+ input,
+ mask = None,):
+ h = self.norm1(input)
+ h1 = self.self_attn(query=h, key=h, value=h, attn_mask=mask)[0]
+ h = h + self.dropout(h1)
+ h = self.norm2(h)
+ return h
+
+
+class PerceiverAttention(nn.Module):
+ def __init__(self, *, dim, dim_head=64, heads=8):
+ super().__init__()
+ self.scale = dim_head**-0.5
+ self.dim_head = dim_head
+ self.heads = heads
+ inner_dim = dim_head * heads
+
+ self.norm1 = nn.LayerNorm(dim)
+ self.norm2 = nn.LayerNorm(dim)
+
+ self.to_q = nn.Linear(dim, inner_dim, bias=False)
+ self.to_kv = nn.Linear(dim, inner_dim * 2, bias=False)
+ self.to_out = nn.Linear(inner_dim, dim, bias=False)
+
+ def forward(self, x, latents):
+ """
+ Args:
+ x (torch.Tensor): image features
+ shape (b, n1, D)
+ latent (torch.Tensor): latent features
+ shape (b, n2, D)
+ """
+ x = self.norm1(x) # [15, 257, 1280]
+ latents = self.norm2(latents) # [15, 16, 1280]
+
+ b, l, _ = latents.shape
+
+ q = self.to_q(latents) # [15, 16, 1280]
+ kv_input = torch.cat((x, latents), dim=-2)
+ k, v = self.to_kv(kv_input).chunk(2, dim=-1) # [15, 257 + 16, 1280]
+
+ q = reshape_tensor(q, self.heads)
+ k = reshape_tensor(k, self.heads)
+ v = reshape_tensor(v, self.heads)
+
+ # attention
+ scale = 1 / math.sqrt(math.sqrt(self.dim_head))
+ weight = (q * scale) @ (k * scale).transpose(-2, -1) # More stable with f16 than dividing afterwards
+ weight = torch.softmax(weight.float(), dim=-1).type(weight.dtype)
+ out = weight @ v
+
+ out = out.permute(0, 2, 1, 3).reshape(b, l, -1)
+
+ return self.to_out(out)
+
+class CrossAttentionLayer(nn.Module):
+ def __init__(self, *, dim, dim_head=64, heads=8):
+ super().__init__()
+ self.perceiver_fg = PerceiverAttention(dim=dim, dim_head=dim_head, heads=heads)
+ self.perceiver_bg = PerceiverAttention(dim=dim, dim_head=dim_head, heads=heads)
+ def forward(self, x, latents):
+ x_fg, x_bg = x
+ latents_fg, latents_bg = latents
+ out_fg = self.perceiver_fg(x_fg, latents_fg)
+ out_bg = self.perceiver_bg(x_bg, latents_bg)
+ return out_fg, out_bg
+
+class Resampler(nn.Module):
+ def __init__(
+ self,
+ dim=1024,
+ depth=8,
+ dim_head=64,
+ heads=16,
+ num_queries=8,
+ embedding_dim=768,
+ output_dim=1024,
+ ff_mult=4,
+ max_seq_len: int = 257, # CLIP tokens + CLS token
+ apply_pos_emb: bool = False,
+ num_latents_mean_pooled: int = 0, # number of latents derived from mean pooled representation of the sequence
+ ):
+ super().__init__()
+ self.pos_emb = nn.Embedding(max_seq_len, embedding_dim) if apply_pos_emb else None
+
+ self.latents = nn.Parameter(torch.randn(1, num_queries, dim) / dim**0.5)
+
+ self.proj_in = nn.Linear(embedding_dim, dim)
+
+ self.proj_out = nn.Linear(dim, output_dim)
+ self.norm_out = nn.LayerNorm(output_dim)
+
+ self.to_latents_from_mean_pooled_seq = (
+ nn.Sequential(
+ nn.LayerNorm(dim),
+ nn.Linear(dim, dim * num_latents_mean_pooled),
+ Rearrange("b (n d) -> b n d", n=num_latents_mean_pooled),
+ )
+ if num_latents_mean_pooled > 0
+ else None
+ )
+
+ self.layers = nn.ModuleList([])
+ for _ in range(depth):
+ self.layers.append(
+ nn.ModuleList(
+ [
+ PerceiverAttention(dim=dim, dim_head=dim_head, heads=heads),
+ FeedForward(dim=dim, mult=ff_mult),
+ ]
+ )
+ )
+
+ def forward(self, x, coherent_queries=None):
+ if self.pos_emb is not None:
+ n, device = x.shape[1], x.device
+ pos_emb = self.pos_emb(torch.arange(n, device=device))
+ x = x + pos_emb
+
+ latents = self.latents.repeat(x.size(0), 1, 1) if coherent_queries is None else \
+ torch.cat([self.latents, coherent_queries], dim=1).repeat(x.size(0), 1, 1) # fg [15, 16, 1280], bg [1, 8 + 8, 1280]
+
+ x = self.proj_in(x) # fg [15, 257, 1280]
+
+ if self.to_latents_from_mean_pooled_seq:
+ meanpooled_seq = masked_mean(x, dim=1, mask=torch.ones(x.shape[:2], device=x.device, dtype=torch.bool))
+ meanpooled_latents = self.to_latents_from_mean_pooled_seq(meanpooled_seq)
+ latents = torch.cat((meanpooled_latents, latents), dim=-2)
+
+ for attn, ff in self.layers:
+ latents = attn(x, latents) + latents # fg [15, 16, 1280]
+ latents = ff(latents) + latents # fg [15, 16, 1280]
+
+ latents = self.proj_out(latents) # fg [15, 16, 768]
+ return self.norm_out(latents)
+
+class SerialSampler(nn.Module):
+ def __init__(
+ self,
+ config,
+ image_processor,
+ image_encoder,
+ dim=1024,
+ depth=8,
+ dim_head=64,
+ num_queries=[8, 8, 8],
+ embedding_dim=768,
+ output_dim=1024,
+ **kwargs
+ ):
+ super().__init__()
+ self.dim = dim
+ self.output_dim = output_dim
+ self.fg_resampler = Resampler(dim=dim, depth=depth, heads=dim // dim_head, dim_head=dim_head,
+ num_queries=num_queries[0], embedding_dim=embedding_dim,
+ output_dim=output_dim, **kwargs)
+ self.bg_resampler = Resampler(dim=dim, depth=depth, heads=dim // dim_head, dim_head=dim_head,
+ num_queries=num_queries[1], embedding_dim=embedding_dim,
+ output_dim=output_dim, **kwargs)
+ self.point_net = PositionNet(in_dim=output_dim, out_dim=output_dim)
+ self.coherent_bridge = GatedSelfAttentionDense(query_dim=dim, context_dim=output_dim,
+ n_heads=dim // dim_head, d_head=dim_head)
+ self.coherent_queries = nn.Parameter(torch.randn(1, num_queries[2], dim) / dim**0.5) # [1, 8, 1280]
+
+ # For Novel Phase
+ self.config = config
+ if self.config.phase == 'novel':
+ self.sample_aggregator_fg = SampleAggregator()
+ self.categories = config.dataset.categories[config.phase]
+ self.aux = Dict(image_processor=image_processor, image_encoder=image_encoder)
+ self.image_patch_path = config.dataset.image_patch_path
+ self.backups = None
+ num_prototypes = config.model.get('num_prototypes') or 128
+ print(f'[SerialSampler]: num_prototypes is {num_prototypes}')
+ self.prototype_bank = PrototypeBank(config, image_encoder, image_processor, num_prototypes=num_prototypes)
+
+ @property
+ def image_processor(self):
+ return self.aux.image_processor
+
+ @property
+ def image_encoder(self):
+ return self.aux.image_encoder
+
+ def build_backup_pool(self):
+ self.ref_table = RefTable()()
+ self.backups, self.backup_masks = {}, {}
+ max_length, max_len_category = 0, None
+ for category in self.categories:
+ files = list(self.ref_table[category].keys())
+ if len(files) > max_length:
+ max_length = len(files)
+ max_len_category = category
+ backups_each_categories = []
+ for file in files:
+ img = Image.open(os.path.join(self.image_patch_path, category, file)).convert('RGB')
+ img = self.image_processor(images=img, return_tensors="pt")['pixel_values'].squeeze(0)
+ backups_each_categories.append(img)
+ self.backups[category] = self.fg_resampler(self.image_encoder(torch.stack(backups_each_categories).to(next(self.parameters()).device)))
+ self.backup_masks[category] = torch.ones(self.backups[category].shape[0], dtype=torch.bool, device=self.backups[category].device)
+ for category, item in self.backups.items():
+ cur_length = item.shape[0]
+ if cur_length < max_length:
+ pad_length = max_length - cur_length
+ pad = torch.zeros(pad_length, *item.shape[1:], device=item.device, dtype=item.dtype)
+ self.backups[category] = torch.cat([item, pad], dim=0)
+ pad_mask = torch.zeros(pad_length, dtype=torch.bool, device=item.device)
+ self.backup_masks[category] = torch.cat([self.backup_masks[category], pad_mask], dim=0)
+ self.backups[''] = torch.zeros_like(self.backups[max_len_category])
+ self.backup_masks[''] = torch.zeros(max_length, dtype=torch.bool, device=self.backups[max_len_category].device)
+
+ # x_objs.shape [15, 257, 1024] ; x_bg.shape [1, 257, 1024]
+ def forward(self, x_objs, obboxes, x_bg, captions):
+ if self.config.phase == 'novel' and self.backups is None:
+ self.build_backup_pool()
+ self.prototype_bank.build_prototypes()
+
+ B = x_bg.shape[0] # 1
+ obboxes = torch.from_numpy(np.array([obbox[::2] + obbox[1::2] for obbox in obboxes[0]])).float().to(x_objs.device) # [15, 8]
+ embed_obboxes = self.point_net(obboxes).unsqueeze(1) # [15, 8] -> [15, 1, 768]
+ embed_objs = self.fg_resampler(x_objs[:, 0]) # [15, 257, 1024] -> [15, 16, 768]
+ conherent_queries = self.coherent_bridge(self.coherent_queries, (embed_obboxes + embed_objs.detach()).view(B, -1, self.output_dim)) # [1, 8, 1280]
+ embed_context = self.bg_resampler(x_bg[:, 0], conherent_queries) # [1, 16, 768]
+
+ if self.config.phase == 'novel':
+ backup_embeddings = torch.stack(list(map(lambda caption: self.backups[caption], captions[0])), dim=0)
+ backup_masks = torch.stack(list(map(lambda caption: self.backup_masks[caption], captions[0])), dim=0)
+ embed_objs = self.sample_aggregator_fg(embed_objs, backup_embeddings, backup_masks)
+ return embed_objs, embed_context # [15, 16, 768], [1, 16, 768]
+
+
+class SampleAggregator(nn.Module):
+ def __init__(self, dim=768, heads=8, dim_head=160, dropout=0.0, alpha_init=0.0):
+ super().__init__()
+
+ # backup self-attention (use the same cross-attn but context=x)
+ self.backup_self_attn = CrossAttention(
+ query_dim=dim,
+ context_dim=dim,
+ heads=heads,
+ dim_head=dim_head,
+ dropout=dropout
+ )
+ self.backup_ln = nn.LayerNorm(dim)
+
+ # primary cross-attention (primary Q, backup K/V)
+ self.primary_cross_attn = CrossAttention(
+ query_dim=dim,
+ context_dim=dim,
+ heads=heads,
+ dim_head=dim_head,
+ dropout=dropout
+ )
+ self.primary_ln = nn.LayerNorm(dim)
+
+ # FFN
+ self.ff = nn.Sequential(
+ nn.Linear(dim, dim * 2),
+ nn.GELU(),
+ nn.Linear(dim * 2, dim)
+ )
+ self.ff_ln = nn.LayerNorm(dim)
+
+ self.gating_param_ca = nn.Parameter(torch.tensor(alpha_init)) # cross-attn scale
+ self.gating_param_ff = nn.Parameter(torch.tensor(alpha_init)) # FFN scale
+
+ def forward(self, primary, backup, backup_mask):
+ """
+ primary: [B, T, C] = [B, 16, 768]
+ backup: [B, N, T, C] = [B, 4, 16, 768] (B bboxes, N refs, T features, C dim)
+ backup_mask: [B, N] # True = valid
+ """
+ B, N, T, C = backup.shape
+
+ # --------- 1) backup self-attention ---------
+ backup_groups = backup.reshape(B * N, T, C) # (B*N, T, C)
+ bk = self.backup_ln(backup_groups)
+ backup_sa = self.backup_self_attn(bk, context=bk)
+ backup_groups = backup_groups + backup_sa
+ backup_groups = backup_groups.reshape(B, N, T, C)
+
+ # --------- 1b) pooled backup memory (small, controlled)
+ backup_mem = backup_groups.mean(dim=2) # (B, N, C) # torch.Size([15, 4, 768])
+
+ # --------- 2) primary cross-attention (Q=primary, K/V=backup) ---------
+ p = self.primary_ln(primary)
+ primary_ca = self.primary_cross_attn(p, context=backup_mem, mask=backup_mask)
+ fused = primary + self.gating_param_ca * primary_ca
+
+ # --------- 3) FFN refinement ---------
+ f = self.ff_ln(fused)
+ fused = fused + self.gating_param_ff * self.ff(f)
+
+ return fused
+
+
+def masked_mean(t, *, dim, mask=None):
+ if mask is None:
+ return t.mean(dim=dim)
+
+ denom = mask.sum(dim=dim, keepdim=True)
+ mask = rearrange(mask, "b n -> b n 1")
+ masked_t = t.masked_fill(~mask, 0.0)
+
+ return masked_t.sum(dim=dim) / denom.clamp(min=1e-5)
\ No newline at end of file
diff --git a/models/dsp/prototype_bank.py b/models/dsp/prototype_bank.py
new file mode 100644
index 0000000000000000000000000000000000000000..5d02fbbb23bc1bd989627f9fb540e169f5a5dcc7
--- /dev/null
+++ b/models/dsp/prototype_bank.py
@@ -0,0 +1,183 @@
+import torch
+import torch.nn as nn
+import torch.distributed as dist
+import torch.nn.functional as F
+import torch.optim as optim
+from .kmeans_pytorch import kmeans
+
+from datamodules import RefTable
+from PIL import Image
+import os
+
+from utils import Dict
+
+class PrototypeLearner(nn.Module):
+ def __init__(self, num_prototypes=64, prototype_dim=1024, lambda_reg=0.1, lr=0.01, iter_steps=50, verbose=True):
+ super().__init__()
+ self.num_prototypes = num_prototypes
+ self.prototype_dim = prototype_dim
+ self.lambda_reg = lambda_reg
+ self.lr = lr
+ self.iter_steps = iter_steps
+ self.verbose = verbose
+
+ def get_initial_prototypes(self, features, device):
+ # features: [N_total, D]
+ feats_norm = F.normalize(features, p=2, dim=-1)
+ _, centers = kmeans(
+ X=feats_norm,
+ num_clusters=self.num_prototypes,
+ distance='cosine',
+ device=device,
+ tqdm_flag=False
+ )
+ return centers.to(device) # [K, D]
+
+ def calculate_reconstruction_loss(self, targets, protos):
+ protos_norm = F.normalize(protos, p=2, dim=-1)
+
+ # Formula: W = T * P^T * (P * P^T + lambda * I)^(-1)
+ # Ref: Eq. (2) in paper
+
+ # P * P^T: Gram Matrix [K, K], (P * P^T + lambda * I)^(-1)
+ p_gram = torch.matmul(protos_norm, protos_norm.t())
+ identity = torch.eye(self.num_prototypes, device=targets.device)
+ inverse_term = torch.inverse(p_gram + self.lambda_reg * identity)
+
+ # Mapping weights [N, K]
+ # W = T * P^T * Inverse
+ mapping_weights = torch.matmul(targets, protos_norm.t())
+ mapping_weights = torch.matmul(mapping_weights, inverse_term)
+
+ # Reconstruct: T_hat = W * P
+ reconstructed = torch.matmul(mapping_weights, protos_norm)
+
+ return F.mse_loss(reconstructed, targets)
+
+ def forward(self, features):
+ """
+ features: [N_total, D]
+ Return: Optimized Prototypes [K, D]
+ """
+ device = features.device
+ N, D = features.shape
+
+ targets = F.normalize(features, p=2, dim=-1).detach()
+
+ initial_protos = self.get_initial_prototypes(features, device)
+ initial_protos_static = initial_protos.clone().detach()
+
+ with torch.no_grad():
+ baseline_loss = self.calculate_reconstruction_loss(targets, initial_protos_static)
+
+ if self.verbose:
+ print(f"\n[ProtoLearner] Start. Baseline (K-Means) Reconstruction MSE: {baseline_loss.item():.6f}")
+
+ prototypes = nn.Parameter(initial_protos.clone())
+ optimizer = optim.Adam([prototypes], lr=self.lr)
+
+ for i in range(self.iter_steps):
+ optimizer.zero_grad()
+ loss = self.calculate_reconstruction_loss(targets, prototypes)
+ loss.backward()
+ optimizer.step()
+
+ with torch.no_grad():
+ prototypes.data.copy_(F.normalize(prototypes.data, p=2, dim=-1))
+
+ if self.verbose and (i == 0 or (i + 1) % 10 == 0 or i == self.iter_steps - 1):
+ with torch.no_grad():
+ curr_protos_norm = F.normalize(prototypes, p=2, dim=-1)
+ init_protos_norm = F.normalize(initial_protos_static, p=2, dim=-1)
+
+ shift_dist = torch.norm(curr_protos_norm - init_protos_norm, dim=-1).mean().item()
+
+ cosine_sim = F.cosine_similarity(curr_protos_norm, init_protos_norm, dim=-1).mean().item()
+
+ gain_pct = (baseline_loss.item() - loss.item()) / baseline_loss.item() * 100
+
+ print(f"Iter {i+1:02d}/{self.iter_steps} | "
+ f"Loss: {loss.item():.6f} | "
+ f"Gain: {gain_pct:.2f}% | "
+ f"Shift(L2): {shift_dist:.4f} | "
+ f"Sim(Cos): {cosine_sim:.4f}")
+
+ return F.normalize(prototypes, p=2, dim=-1).detach()
+
+
+class PrototypeBank(nn.Module):
+ def __init__(self, config, image_encoder, image_processor, num_prototypes=128, prototype_dim=1024):
+ super().__init__()
+
+ self.config = config
+ self.num_prototypes = num_prototypes
+ self.prototype_dim = prototype_dim
+ self.image_patch_path = config.dataset.image_patch_path
+ self.categories = config.dataset.categories[config.phase]
+ self.aux = Dict(image_processor=image_processor, image_encoder=image_encoder)
+
+ shape = (len(self.categories) + 1, num_prototypes, prototype_dim)
+ self.register_buffer('prototypes', torch.zeros(shape))
+ self.register_buffer('prototype_flag', torch.tensor(False))
+
+ self.learner = PrototypeLearner(
+ num_prototypes=num_prototypes,
+ prototype_dim=prototype_dim,
+ lambda_reg=0.1,
+ lr=0.05,
+ iter_steps=50
+ )
+
+ @property
+ def image_processor(self):
+ return self.aux.image_processor
+
+ @property
+ def image_encoder(self):
+ return self.aux.image_encoder
+
+ def build_prototypes(self):
+ self.ref_table = RefTable()()
+ self.cate_to_id = {cate: idx for (idx, cate) in enumerate(self.categories)}
+ self.cate_to_id[''] = -1
+
+ if self.prototype_flag.item():
+ if (dist.is_initialized() and dist.get_rank() == 0) or (not dist.is_initialized()):
+ print(f"[PrototypeBank] Prototypes loaded from checkpoint (Frozen). Skip calculation.")
+ return
+
+ is_dist = dist.is_initialized()
+ rank = dist.get_rank() if is_dist else 0
+ device = next(iter(self.image_encoder.parameters())).device
+ temp_prototypes = None
+
+ if rank == 0:
+ print("[PrototypeBank] Calculating prototypes online...")
+ self.dense_features = {}
+ for category in self.categories:
+ files = list(self.ref_table[category].keys())
+ ref_images = []
+ for file in files:
+ img = Image.open(os.path.join(self.image_patch_path, category, file)).convert('RGB')
+ img = self.image_processor(images=img, return_tensors="pt", do_normalize=False)['pixel_values'].squeeze(0)
+ ref_images.append(img)
+ self.dense_features[category] = self.image_encoder(torch.stack(ref_images).to(device), mode='x_norm_patchtokens')
+
+ prototype_list = []
+ for category in self.categories:
+ flat_feats = self.dense_features[category].reshape(-1, self.prototype_dim)
+ optimized_protos = self.learner(flat_feats)
+ prototype_list.append(optimized_protos)
+
+ prototype_list.append(torch.zeros_like(prototype_list[-1]))
+ temp_prototypes = torch.stack(prototype_list)
+
+ self.prototypes.copy_(temp_prototypes)
+ self.prototype_flag.fill_(True)
+
+ if is_dist:
+ dist.broadcast(self.prototypes, src=0)
+ dist.broadcast(self.prototype_flag, src=0)
+
+ def get_prototypes(self, captions):
+ return torch.stack(list(map(lambda caption: self.prototypes[self.cate_to_id[caption]], captions)))
\ No newline at end of file
diff --git a/models/dsp/utils.py b/models/dsp/utils.py
new file mode 100644
index 0000000000000000000000000000000000000000..354b47a86cbe3d8015db82b80e7f35a97da0d9ec
--- /dev/null
+++ b/models/dsp/utils.py
@@ -0,0 +1,134 @@
+import functools
+import os
+import random
+import numpy as np
+import imagesize
+import torch
+import torch.nn.functional as F
+from PIL import Image
+import cv2
+
+def seed_everything(seed):
+ # np.random.seed(seed)
+ torch.manual_seed(seed)
+ torch.cuda.manual_seed_all(seed)
+ random.seed(seed)
+
+def find_nearest(array, value):
+ array = np.asarray(array)
+ idx = (np.abs(array/value - 1)).argmin()
+ return idx
+
+def get_sup_mask(mask_list):
+ or_mask = np.zeros_like(mask_list[0])
+ for mask in mask_list:
+ or_mask += mask
+ or_mask[or_mask >= 1] = 1
+ sup_mask = 1 - or_mask
+ return sup_mask
+
+def get_masks(obboxes, height, width, device):
+ # Construct Instance Guidance Mask
+ guidance_masks, in_box = [], []
+ for obbox in obboxes[0]:
+ guidance_mask = np.zeros((height, width))
+ if np.count_nonzero(obbox):
+ pts = np.array(obbox).reshape(-1, 1, 2)
+ pts[..., 0] = pts[..., 0] * width
+ pts[..., 1] = pts[..., 1] * height
+ pts = np.int32(pts)
+ guidance_masks.append(cv2.fillPoly(guidance_mask, [pts], 1)[None, ...])
+ else:
+ guidance_masks.append(guidance_mask[None, ...])
+ in_box.append([obbox[0], obbox[2], obbox[4], obbox[6], obbox[1], obbox[3], obbox[5], obbox[7]])
+
+ # Construct Background Guidance Mask
+ sup_mask = get_sup_mask(guidance_masks)
+ supplement_mask = torch.from_numpy(sup_mask[None, ...])
+ supplement_mask = F.interpolate(supplement_mask, (height//8, width//8), mode='bilinear').float()
+ supplement_mask = supplement_mask.to(device) # (1, 1, H, W)
+
+ guidance_masks = np.concatenate(guidance_masks, axis=0)
+ guidance_masks = guidance_masks[None, ...]
+ guidance_masks = torch.from_numpy(guidance_masks).float().to(device)
+ guidance_masks = F.interpolate(guidance_masks, (height//8, width//8), mode='bilinear') # (1, instance_num, H, W)
+ # guidance_masks.shape [1, 15, 64, 64] ; supplement_mask.shape [1, 1, 64, 64]
+ in_box = torch.from_numpy(np.array(in_box))[None, ...].float().to(device) # (1, instance_num, 4)
+ return guidance_masks, supplement_mask, in_box
+
+
+def get_sigmoid(bboxes, height, width, device):
+ sigmoid_values = []
+ for w_min, h_min, w_max, h_max in bboxes[0]:
+ H, W = height//8, width // 8
+ x = torch.linspace(0, W - 1, W)
+ y = torch.linspace(0, H - 1, H)
+ yy, xx = torch.meshgrid(y, x, indexing='ij')
+ xx, yy = xx / H, yy / W
+ mu1 = (w_min + w_max) / 2
+ mu2 = (h_min + h_max) / 2
+ sigma1 = ((w_max - w_min) ** 2) / 4
+ sigma2 = ((h_max - h_min) ** 2) / 4
+ if sigma1 == 0 or sigma2 == 0:
+ sigmoid_values.append(torch.zeros_like(xx))
+ continue
+ exponent = -10 * (1 - ((xx - mu1) ** 2) / sigma1 - ((yy - mu2) ** 2) / sigma2)
+ sigmoid_value = 1 / (1 + torch.exp(exponent))
+ sigmoid_values.append(sigmoid_value)
+ sigmoid_values = torch.stack(sigmoid_values, dim=0)[None, ...].to(device)
+ # sigmoid_values.shape [1, 15, 64, 64] ; in_box.shape [1, 15, 8]
+ return sigmoid_values
+
+
+class ExemplarPool:
+ def __init__(self, data_embeds_dict_path, exemplar_pool_path, image_processor):
+ self.data_embeds_dict = torch.load(data_embeds_dict_path, map_location='cpu')
+ self.all_img_names = np.array(list(self.data_embeds_dict.keys()))
+ self.all_txt_embs = torch.cat([self.data_embeds_dict[name]['txt_emb'] for name in self.all_img_names], dim=0)
+ self.all_img_embs = torch.cat([self.data_embeds_dict[name]['img_emb'] for name in self.all_img_names], dim=0)
+
+ self.exemplar_pool_path = exemplar_pool_path
+ self.image_processor = image_processor
+
+ def to(self, device, dtype=None):
+ self.device = device
+ self.all_txt_embs = self.all_txt_embs.to(device=device, dtype=dtype)
+ self.all_img_embs = self.all_img_embs.to(device=device, dtype=dtype)
+
+ def get_similar_exemplars_names(self, prompt_emb, topk, sim_mode):
+ prompt_emb = F.normalize(prompt_emb, dim=-1).detach()
+
+ if sim_mode == 'text2text':
+ sim_vals = torch.matmul(prompt_emb, self.all_txt_embs.T)
+ elif sim_mode == 'text2img':
+ sim_vals = torch.matmul(prompt_emb, self.all_img_embs.T)
+ elif sim_mode == 'both':
+ txt_sim_vals = torch.matmul(prompt_emb, self.all_txt_embs.T)
+ img_sim_vals = torch.matmul(prompt_emb, self.all_img_embs.T)
+ sim_vals = (txt_sim_vals + img_sim_vals) * 0.5
+ else:
+ raise ValueError('Invalid mode for similarity computation!')
+
+ _, topk_indices = torch.topk(sim_vals, k=topk, dim=1)
+ topk_img_names = self.all_img_names[topk_indices.cpu().numpy()].tolist()
+
+ return topk_img_names
+
+ def names_to_tensors(self, topk_img_names):
+ topk_img_tensors = []
+ for names_for_one_prompt in topk_img_names:
+ topk_tensors_per_prompt = []
+ for name in names_for_one_prompt:
+ img = Image.open(os.path.join(self.exemplar_pool_path, name)).convert('RGB')
+ # tensor = self.image_processor(images=img, return_tensors="pt", do_normalize=False)['pixel_values'].squeeze(0)
+ tensor = self.image_processor(images=img, return_tensors="pt")['pixel_values'].squeeze(0)
+ topk_tensors_per_prompt.append(tensor)
+ stacked_k_tensors = torch.stack(topk_tensors_per_prompt)
+ topk_img_tensors.append(stacked_k_tensors)
+ topk_img_tensors = torch.stack(topk_img_tensors)
+ return topk_img_tensors
+
+ def get_similar_exemplars(self, prompt_emb, topk=1, sim_mode='text2img'):
+ topk_img_names = self.get_similar_exemplars_names(prompt_emb, topk, sim_mode)
+ topk_img_tensors = self.names_to_tensors(topk_img_names)
+ return topk_img_tensors.to(self.device)
\ No newline at end of file
diff --git a/results/claim1_hypersphere.json b/results/claim1_hypersphere.json
new file mode 100644
index 0000000000000000000000000000000000000000..7c37c769deee47b4b2f7221e02c25fdb2492f5e8
--- /dev/null
+++ b/results/claim1_hypersphere.json
@@ -0,0 +1 @@
+[{"d": 10, "n": 1000, "cd_raw": 0.977015581824591, "cc_raw": 0.981, "cd_gicdm": 0.004743623022192922, "cc_gicdm": 0.994, "kept": 0.011, "h5_raw": 1.765536723163842, "A5_raw": 0.09}, {"d": 50, "n": 1000, "cd_raw": 0.9944689663464403, "cc_raw": 0.982, "cd_gicdm": 0.4877741694143266, "cc_gicdm": 1.0, "kept": 0.734, "h5_raw": 2.0242914979757085, "A5_raw": 0.39}, {"d": 100, "n": 1000, "cd_raw": 0.9961317181077468, "cc_raw": 0.967, "cd_gicdm": 0.6484478606856796, "cc_gicdm": 1.0, "kept": 0.873, "h5_raw": 2.0508613617719442, "A5_raw": 0.4}, {"d": 500, "n": 1000, "cd_raw": 0.9985104045791207, "cc_raw": 0.973, "cd_gicdm": 0.7944439443574028, "cc_gicdm": 0.999, "kept": 0.971, "h5_raw": 2.0781379883624274, "A5_raw": 0.4}, {"d": 1000, "n": 1000, "cd_raw": 0.9991241002869283, "cc_raw": 0.962, "cd_gicdm": 0.8073593139737273, "cc_gicdm": 0.981, "kept": 0.976, "h5_raw": 2.043318348998774, "A5_raw": 0.4}]
\ No newline at end of file
diff --git a/results/claim2_density.json b/results/claim2_density.json
new file mode 100644
index 0000000000000000000000000000000000000000..188bec07a339bef304970b1062d95ba4188c9f44
--- /dev/null
+++ b/results/claim2_density.json
@@ -0,0 +1 @@
+{"d": 20, "N": 1600, "K": 20, "log_corr_density": 0.989796887850633, "mu_before_std": 3.101600099029342, "mu_before_cv": 0.6111469088770181, "mu_after_std": 0.000120823332638221, "mu_after_cv": 0.00012082410829053853}
\ No newline at end of file
diff --git a/results/claim5_benchmark.json b/results/claim5_benchmark.json
new file mode 100644
index 0000000000000000000000000000000000000000..b2ecf835cb011878acaa63ef7692fa8d3eda9931
--- /dev/null
+++ b/results/claim5_benchmark.json
@@ -0,0 +1 @@
+{"gauss_mean_ideal": {"cd0": 0.9963774533414302, "cc0": 0.9693333333333334, "cd": 0.10316801132756441, "cc": 1.0}, "gauss_std": {"cd0": 1.0, "cc0": 0.0, "cd": 0.0, "cc": 0.0}, "hypersphere_equal": {"cd0": 0.9941503848693399, "cc0": 0.978, "cd": 0.0, "cc": 0.0}, "mode_collapse": {"cd0": 0.9963335536365368, "cc0": 0.18733333333333332, "cd": 0.0, "cc": 0.0}, "mode_full": {"cd0": 0.9958290506429619, "cc0": 0.9633333333333334, "cd": 0.0, "cc": 0.0}}
\ No newline at end of file
diff --git a/results/claim5_official.json b/results/claim5_official.json
new file mode 100644
index 0000000000000000000000000000000000000000..c618b4a8aea210f722aed363f21fe458fba4c501
--- /dev/null
+++ b/results/claim5_official.json
@@ -0,0 +1 @@
+{"gauss_mean_equal": {"cd_raw": 1.0, "cc_raw": 1.0, "cd_gicdm": 0.9724727100142381, "cc_gicdm": 0.9943011613360877}, "gauss_std_1p5": {"cd_raw": 0.0, "cc_raw": 0.0, "cd_gicdm": 0.0, "cc_gicdm": 0.0}, "hypersphere_equal": {"cd_raw": 1.0, "cc_raw": 1.0, "cd_gicdm": 1.0, "cc_gicdm": 0.997871694412942}, "mode_collapse": {"cd_raw": 1.0, "cc_raw": 0.19253830779480346, "cd_gicdm": 0.9714876983214533, "cc_gicdm": 0.20837479198571804}, "mode_full": {"cd_raw": 1.0, "cc_raw": 1.0, "cd_gicdm": 1.0, "cc_gicdm": 1.0036598977305173}, "sphere_offset": {"cd_raw": 0.0, "cc_raw": 0.0, "cd_gicdm": 0.0, "cc_gicdm": 0.0}}
\ No newline at end of file
diff --git a/results/claim6_gaussian_hubness.json b/results/claim6_gaussian_hubness.json
new file mode 100644
index 0000000000000000000000000000000000000000..5fa1bd2190079a68ce4dda12b9f6851dde6713f6
--- /dev/null
+++ b/results/claim6_gaussian_hubness.json
@@ -0,0 +1 @@
+[{"d": 10, "N": 8000, "h5": 1.706703076332295, "A5": 0.057125}, {"d": 20, "N": 8000, "h5": 1.9945150835203191, "A5": 0.152875}, {"d": 50, "N": 8000, "h5": 2.5238185374471573, "A5": 0.28925}, {"d": 100, "N": 8000, "h5": 2.962524070508073, "A5": 0.378125}, {"d": 200, "N": 8000, "h5": 3.229452607782981, "A5": 0.43325}, {"d": 500, "N": 8000, "h5": 3.719546215361726, "A5": 0.501625}, {"d": 1000, "N": 8000, "h5": 3.855050115651504, "A5": 0.517}]
\ No newline at end of file
diff --git a/results/crossover_dstar.json b/results/crossover_dstar.json
new file mode 100644
index 0000000000000000000000000000000000000000..7c713c539384e62ae1bc9f0fbd2927808ce02be3
--- /dev/null
+++ b/results/crossover_dstar.json
@@ -0,0 +1 @@
+[[1000, 1, 46], [1000, 5, 30], [1000, 10, 26], [1000, 20, 20], [1000, 50, 14], [10000, 1, 62], [10000, 5, 48], [10000, 10, 42], [10000, 20, 36], [10000, 50, 30], [50000, 1, 76], [50000, 5, 60], [50000, 10, 54], [50000, 20, 48], [50000, 50, 42]]
\ No newline at end of file
diff --git a/results/crossover_fig.html b/results/crossover_fig.html
new file mode 100644
index 0000000000000000000000000000000000000000..b27750b78b6bae8b46ef2439836694303c606ee1
--- /dev/null
+++ b/results/crossover_fig.html
@@ -0,0 +1,7 @@
+
+
+
+
+
+
\ No newline at end of file
diff --git a/results/hubness_fig.html b/results/hubness_fig.html
new file mode 100644
index 0000000000000000000000000000000000000000..2160e49322bd64f13695151826ccae6760a94645
--- /dev/null
+++ b/results/hubness_fig.html
@@ -0,0 +1,7 @@
+
+
+
+
+
+
\ No newline at end of file
diff --git a/scripts/data_process/dior/00.link_image_files.sh b/scripts/data_process/dior/00.link_image_files.sh
new file mode 100644
index 0000000000000000000000000000000000000000..90aafe0486237af2f72964b67d3deef89ea5e790
--- /dev/null
+++ b/scripts/data_process/dior/00.link_image_files.sh
@@ -0,0 +1,3 @@
+#!/usr/bin/env bash
+mkdir -p $DSP_PROJECT_DIR/data/DIOR
+ln -s /path/to/DIOR-VOC/VOC2007/JPEGImages $DSP_PROJECT_DIR/data/DIOR/images # Replace with your actual path
\ No newline at end of file
diff --git a/scripts/data_process/dior/01.a.process_train.py b/scripts/data_process/dior/01.a.process_train.py
new file mode 100644
index 0000000000000000000000000000000000000000..461b28adc01f4e98956c503aef3a5d25225a03fc
--- /dev/null
+++ b/scripts/data_process/dior/01.a.process_train.py
@@ -0,0 +1,95 @@
+import os
+import xml.etree.ElementTree as ET
+import json
+
+PROJECT_DIR = os.getenv('DSP_PROJECT_DIR', '/path/to/DSP_PROJECT_DIR') # Set this manually if the environment variable is unavailable
+base_dir = '/path/to/DIOR-VOC/Annotations/Horizontal_Bounding_Boxes' # Replace with your actual path
+
+category_list = [
+ 'vehicle', 'baseballfield', 'groundtrackfield', 'windmill', 'bridge',
+ 'overpass', 'ship', 'airplane', 'tenniscourt', 'airport',
+ 'expressway-service-area', 'basketballcourt', 'stadium', 'storagetank', 'chimney',
+ 'dam', 'expressway-toll-station', 'golffield', 'trainstation', 'harbor'
+]
+category_dict_rev = {v: i for i, v in enumerate(category_list)}
+
+novel_categories = [
+ 'windmill', 'airport', 'chimney', 'dam', 'trainstation'
+]
+novel_ids = set([category_dict_rev[cate] for cate in novel_categories])
+novel_dict_rev = {v: i for i, v in enumerate(novel_categories)}
+
+num_classes = 20
+width_height = 800
+caption_prefix = "The aerial image features a city with "
+thr = 15
+
+os.makedirs(os.path.join(PROJECT_DIR, 'data', 'DIOR', 'metadatas', 'data_setting1'), exist_ok=True)
+
+if __name__ == '__main__':
+ base_list, novel_list = [], [[], [], [], [], [], []]
+ filenames = sorted(os.listdir(base_dir))[:5862]
+ for filename in filenames:
+ dictin = {}
+ root = ET.parse(os.path.join(base_dir, filename)).getroot()
+ categories_in_this_image = set()
+ categories, bndboxes, obndboxes= [], [], []
+ for object in root.findall('object'):
+ category = object.find('name').text.lower()
+ category_id = category_dict_rev[category]
+ xmin, ymin, xmax, ymax = [int(child.text) / width_height for child in object.find('bndbox')]
+ categories.append(category)
+ bndboxes.append([xmin, ymin, xmax, ymax])
+ obndboxes.append([xmin, ymin, xmax, ymin, xmax, ymax, xmin, ymax])
+ categories_in_this_image.add(category_id)
+
+ # caption put to line 58
+ caption = [caption_prefix + ", ".join(categories)]
+
+ is_novel = False
+ if categories_in_this_image & set(novel_ids):
+ is_novel = True
+ tmp_categories, tmp_bndboxes, tmp_obndboxes= [], [], []
+ for i in range(len(categories)):
+ if categories[i] in novel_categories:
+ tmp_categories.append(categories[i])
+ tmp_bndboxes.append(bndboxes[i])
+ tmp_obndboxes.append(obndboxes[i])
+ categories, bndboxes, obndboxes = tmp_categories, tmp_bndboxes, tmp_obndboxes
+
+ if len(categories) > thr:
+ categories = categories[:thr]
+ bndboxes = bndboxes[:thr]
+ obndboxes = obndboxes[:thr]
+ while len(categories) < thr:
+ categories.append("")
+ bndboxes.append([0,0,0,0])
+ obndboxes.append([0,0,0,0,0,0,0,0])
+
+ dictin["file_name"] = f"../../images/{filename[:5]}.jpg"
+ caplist = caption + categories
+ dictin["captions"] = caplist
+ dictin["bndboxes"] = bndboxes
+ dictin["obboxes"] = obndboxes
+
+ if is_novel:
+ if len(categories_in_this_image) > 1:
+ novel_list[5].append(dictin.copy())
+ else:
+ novel_list[novel_dict_rev[category_list[next(iter(categories_in_this_image))]]].append(dictin.copy())
+ else:
+ base_list.append(dictin.copy())
+
+ with open(os.path.join(PROJECT_DIR, "data/DIOR/metadatas/data_setting1/train_base.jsonl"), "w", encoding="utf-8") as f:
+ for item in base_list:
+ f.write(json.dumps(item, ensure_ascii=False) + "\n")
+
+ for i in range(5):
+ with open(os.path.join(PROJECT_DIR, f"data/DIOR/metadatas/data_setting1/train_novel_{novel_categories[i]}.jsonl"), "w", encoding="utf-8") as f:
+ for item in novel_list[i]:
+ f.write(json.dumps(item, ensure_ascii=False) + "\n")
+
+ with open(os.path.join(PROJECT_DIR, "data/DIOR/metadatas/data_setting1/train_novel_mixed.jsonl"), "w", encoding="utf-8") as f:
+ for item in novel_list[5]:
+ f.write(json.dumps(item, ensure_ascii=False) + "\n")
+
\ No newline at end of file
diff --git a/scripts/data_process/dior/01.b.process_test.py b/scripts/data_process/dior/01.b.process_test.py
new file mode 100644
index 0000000000000000000000000000000000000000..1f2aa1953bab59522b3da881065c35b121737cf7
--- /dev/null
+++ b/scripts/data_process/dior/01.b.process_test.py
@@ -0,0 +1,94 @@
+import os
+import xml.etree.ElementTree as ET
+import json
+
+PROJECT_DIR = os.getenv('DSP_PROJECT_DIR', '/path/to/DSP_PROJECT_DIR') # Set this manually if the environment variable is unavailable
+base_dir = '/path/to/DIOR-VOC/Annotations/Horizontal_Bounding_Boxes' # Replace with your actual path
+
+category_list = [
+ 'vehicle', 'baseballfield', 'groundtrackfield', 'windmill', 'bridge',
+ 'overpass', 'ship', 'airplane', 'tenniscourt', 'airport',
+ 'expressway-service-area', 'basketballcourt', 'stadium', 'storagetank', 'chimney',
+ 'dam', 'expressway-toll-station', 'golffield', 'trainstation', 'harbor'
+]
+category_dict_rev = {v: i for i, v in enumerate(category_list)}
+
+novel_categories = [
+ 'windmill', 'airport', 'chimney', 'dam', 'trainstation'
+]
+novel_ids = set([category_dict_rev[cate] for cate in novel_categories])
+novel_dict_rev = {v: i for i, v in enumerate(novel_categories)}
+
+num_classes = 20
+width_height = 800
+caption_prefix = "The aerial image features a city with "
+thr = 15
+
+os.makedirs(os.path.join(PROJECT_DIR, 'data', 'DIOR', 'metadatas', 'data_setting1'), exist_ok=True)
+
+if __name__ == '__main__':
+ base_list, novel_list = [], [[], [], [], [], [], []]
+ filenames = sorted(os.listdir(base_dir))[11725:]
+ for filename in filenames:
+ dictin = {}
+ root = ET.parse(os.path.join(base_dir, filename)).getroot()
+ categories_in_this_image = set()
+ categories, bndboxes, obndboxes= [], [], []
+ for object in root.findall('object'):
+ category = object.find('name').text.lower()
+ category_id = category_dict_rev[category]
+ xmin, ymin, xmax, ymax = [int(child.text) / width_height for child in object.find('bndbox')]
+ categories.append(category)
+ bndboxes.append([xmin, ymin, xmax, ymax])
+ obndboxes.append([xmin, ymin, xmax, ymin, xmax, ymax, xmin, ymax])
+ categories_in_this_image.add(category_id)
+
+ caption = [caption_prefix + ", ".join(categories)]
+
+ is_novel = False
+ if categories_in_this_image & set(novel_ids):
+ is_novel = True
+ tmp_categories, tmp_bndboxes, tmp_obndboxes= [], [], []
+ for i in range(len(categories)):
+ if categories[i] in novel_categories:
+ tmp_categories.append(categories[i])
+ tmp_bndboxes.append(bndboxes[i])
+ tmp_obndboxes.append(obndboxes[i])
+ categories, bndboxes, obndboxes = tmp_categories, tmp_bndboxes, tmp_obndboxes
+
+ if len(categories) > thr:
+ categories = categories[:thr]
+ bndboxes = bndboxes[:thr]
+ obndboxes = obndboxes[:thr]
+ while len(categories) < thr:
+ categories.append("")
+ bndboxes.append([0,0,0,0])
+ obndboxes.append([0,0,0,0,0,0,0,0])
+
+ dictin["file_name"] = f"../../images/{filename[:5]}.jpg"
+ caplist = caption + categories
+ dictin["captions"] = caplist
+ dictin["bndboxes"] = bndboxes
+ dictin["obboxes"] = obndboxes
+
+ if is_novel:
+ if len(categories_in_this_image) > 1:
+ novel_list[5].append(dictin.copy())
+ else:
+ novel_list[novel_dict_rev[category_list[next(iter(categories_in_this_image))]]].append(dictin.copy())
+ else:
+ base_list.append(dictin.copy())
+
+ with open(os.path.join(PROJECT_DIR, "data/DIOR/metadatas/data_setting1/test_base.jsonl"), "w", encoding="utf-8") as f:
+ for item in base_list:
+ f.write(json.dumps(item, ensure_ascii=False) + "\n")
+
+ for i in range(5):
+ with open(os.path.join(PROJECT_DIR, f"data/DIOR/metadatas/data_setting1/test_novel_{novel_categories[i]}.jsonl"), "w", encoding="utf-8") as f:
+ for item in novel_list[i]:
+ f.write(json.dumps(item, ensure_ascii=False) + "\n")
+
+ with open(os.path.join(PROJECT_DIR, "data/DIOR/metadatas/data_setting1/test_novel_mixed.jsonl"), "w", encoding="utf-8") as f:
+ for item in novel_list[5]:
+ f.write(json.dumps(item, ensure_ascii=False) + "\n")
+
\ No newline at end of file
diff --git a/scripts/data_process/dior/02.data_crop.py b/scripts/data_process/dior/02.data_crop.py
new file mode 100644
index 0000000000000000000000000000000000000000..5ee0d0eee33db7ec4d88eb637f712b2914bde96d
--- /dev/null
+++ b/scripts/data_process/dior/02.data_crop.py
@@ -0,0 +1,63 @@
+import os
+import json
+import cv2
+import torch
+import numpy as np
+from collections import defaultdict
+from torchvision import transforms
+from PIL import Image
+import xml.etree.ElementTree as ET
+from collections import Counter
+
+PROJECT_DIR = os.getenv('DSP_PROJECT_DIR', '/path/to/DSP_PROJECT_DIR') # Set this manually if the environment variable is unavailable
+base_dir = '/path/to/DIOR-VOC/Annotations/Horizontal_Bounding_Boxes' # Replace with your actual path
+image_dir = '/path/to/DIOR-VOC/VOC2007/JPEGImages' # Replace with your actual path
+
+category_list = [
+ 'vehicle', 'baseballfield', 'groundtrackfield', 'windmill', 'bridge',
+ 'overpass', 'ship', 'airplane', 'tenniscourt', 'airport',
+ 'expressway-service-area', 'basketballcourt', 'stadium', 'storagetank', 'chimney',
+ 'dam', 'expressway-toll-station', 'golffield', 'trainstation', 'harbor'
+]
+category_dict_rev = {v: i for i, v in enumerate(category_list)}
+width_height = 800
+
+output_dir = os.path.join(PROJECT_DIR, "data", "DIOR", "patches")
+
+if __name__ == '__main__':
+ os.makedirs(output_dir, exist_ok=True)
+ counter = Counter()
+ filenames = sorted(os.listdir(base_dir))[:5862]
+ for filename in filenames:
+ dictin = {}
+ image = cv2.imread(os.path.join(image_dir, f'{os.path.splitext(filename)[0]}.jpg'))
+ root = ET.parse(os.path.join(base_dir, filename)).getroot()
+ categories, bndboxes, obndboxes= [], [], []
+ for i, object in enumerate(root.findall('object')):
+ category = object.find('name').text.lower()
+ category_id = category_dict_rev[category]
+ xmin, ymin, xmax, ymax = [int(child.text) for child in object.find('bndbox')]
+
+ bbox_width = xmax - xmin
+ bbox_height = ymax - ymin
+
+ bbox_area = bbox_width * bbox_height
+ image_area = width_height * width_height
+ bbox_ratio = bbox_area / image_area
+
+ if bbox_ratio < 0.0005:
+ continue
+ class_dir = os.path.join(output_dir, category)
+ os.makedirs(class_dir, exist_ok=True)
+
+ counter[class_dir] += 1
+ xmin, ymin, xmax, ymax = int(xmin), int(ymin), int(xmax), int(ymax)
+ cropped_image = image[ymin:ymax, xmin:xmax]
+
+ output_image_name = f"{os.path.splitext(filename)[0]}_{i}.jpg"
+ output_image_path = os.path.join(class_dir, output_image_name)
+
+ cv2.imwrite(output_image_path, cropped_image)
+
+
+ print(counter)
\ No newline at end of file
diff --git a/scripts/data_process/dior/03.cache_embs.py b/scripts/data_process/dior/03.cache_embs.py
new file mode 100644
index 0000000000000000000000000000000000000000..c548599dfc6956e492bb4601ecbae4dd8c5cc145
--- /dev/null
+++ b/scripts/data_process/dior/03.cache_embs.py
@@ -0,0 +1,63 @@
+import os, json, random
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from tqdm import tqdm
+from transformers import CLIPModel, AutoTokenizer, AutoProcessor
+from PIL import Image
+
+PROJECT_DIR = os.getenv('DSP_PROJECT_DIR', '/path/to/DSP_PROJECT_DIR') # Set this manually if the environment variable is unavailable
+data_setting_path = os.path.join(PROJECT_DIR, 'data', 'DIOR', 'metadatas', 'data_setting1')
+cache_name = os.path.join(PROJECT_DIR, 'data', 'DIOR', 'dior_emb.pt')
+
+class myCLIPEnc(nn.Module):
+ def __init__(self, model_config='openai/clip-vit-large-patch14', device='cuda'):
+ super().__init__()
+ self.device = device
+
+ self.tokenizer = AutoTokenizer.from_pretrained(model_config)
+ self.processor = AutoProcessor.from_pretrained(model_config)
+ self.model = CLIPModel.from_pretrained(model_config).to(device)
+ self.model.eval()
+
+ def forward(self, caption=None, img=None):
+ if caption is not None:
+ txt_inp = self.tokenizer(caption, padding=True, truncation=True, return_tensors="pt").to(self.device) # pad and truncate to the max_length
+ txt_feat = self.model.get_text_features(**txt_inp)
+ txt_feat = F.normalize(txt_feat, dim=-1).detach().cpu()
+ else:
+ txt_feat = None
+
+ if img is not None:
+ img_inp = self.processor(images=img, return_tensors="pt").to(self.device)
+ img_feat = self.model.get_image_features(**img_inp)
+ img_feat = F.normalize(img_feat, dim=-1).detach().cpu()
+ else:
+ img_feat = None
+
+ return txt_feat, img_feat
+
+
+if __name__ == '__main__':
+ if not os.path.exists(cache_name):
+ myCLIP = myCLIPEnc()
+
+ data = []
+ with open(os.path.join(data_setting_path, 'train_base.jsonl'), 'r') as f:
+ for line in f:
+ data.append(json.loads(line))
+ emb_dict = {}
+ sample_data = random.sample(data, 4000)
+ for sample in tqdm(sample_data):
+ img_name = sample['file_name']
+ emb_dict[os.path.basename(img_name)] = {}
+
+ caption = sample['captions'][0]
+ img = Image.open(os.path.join(data_setting_path, img_name)).convert('RGB')
+ txt_emb, img_emb = myCLIP(caption=caption, img=img)
+
+ emb_dict[os.path.basename(img_name)]['txt_emb'] = txt_emb.detach().cpu()
+ emb_dict[os.path.basename(img_name)]['img_emb'] = img_emb.detach().cpu()
+ # torch.cuda.empty_cache()
+
+ torch.save(emb_dict, cache_name)
\ No newline at end of file
diff --git a/scripts/data_process/dior/04.cache_image_sizes.py b/scripts/data_process/dior/04.cache_image_sizes.py
new file mode 100644
index 0000000000000000000000000000000000000000..534e8feecdc2b5bf35c3cd911596eb4a2972bc64
--- /dev/null
+++ b/scripts/data_process/dior/04.cache_image_sizes.py
@@ -0,0 +1,42 @@
+import os
+import json
+import imagesize
+from tqdm import tqdm
+
+PROJECT_DIR = os.getenv('DSP_PROJECT_DIR', '/path/to/DSP_PROJECT_DIR') # Set this manually if the environment variable is unavailable
+image_patch_path = os.path.join(PROJECT_DIR, "data", "DIOR", "patches")
+output_path = os.path.join(image_patch_path, "image_sizes.json")
+
+def generate_size_cache(image_patch_path, output_path):
+ print(f"Scanning: {image_patch_path}")
+ size_cache = {}
+
+ categories = [d for d in os.listdir(image_patch_path) if os.path.isdir(os.path.join(image_patch_path, d))]
+
+ for category in tqdm(categories, desc="Processing"):
+ category_dir = os.path.join(image_patch_path, category)
+
+ for item in os.listdir(category_dir):
+ if item.endswith('.jpg'):
+ img_path = os.path.join(category_dir, item)
+ try:
+ size_cache[img_path] = imagesize.get(img_path)
+ except Exception as e:
+ print(f"Failed {img_path}: {e}")
+
+ aug_dir = os.path.join(category_dir, 'augmented')
+ if os.path.exists(aug_dir):
+ for item in os.listdir(aug_dir):
+ if item.endswith('.jpg'):
+ img_path = os.path.join(aug_dir, item)
+ try:
+ size_cache[img_path] = imagesize.get(img_path)
+ except Exception as e:
+ print(f"Failed {img_path}: {e}")
+
+ with open(output_path, 'w') as f:
+ json.dump(size_cache, f)
+ print(f"\nSaved: {output_path} (Containing {len(size_cache)} images)")
+
+if __name__ == "__main__":
+ generate_size_cache(image_patch_path, output_path)
\ No newline at end of file
diff --git a/scripts/data_process/exdark/01.process.py b/scripts/data_process/exdark/01.process.py
new file mode 100644
index 0000000000000000000000000000000000000000..95d510829b3898bd89a971759c7eadc2e157cc12
--- /dev/null
+++ b/scripts/data_process/exdark/01.process.py
@@ -0,0 +1,133 @@
+from pathlib import Path
+from PIL import Image
+import shutil
+import os
+import imagesize
+import json
+
+PROJECT_DIR = os.getenv('DSP_PROJECT_DIR', '/path/to/DSP_PROJECT_DIR') # Set this manually if the environment variable is unavailable
+base_dir = '/path/to/ExDark' # Replace with your actual path
+output_dir = os.path.join(PROJECT_DIR, 'data', 'EXDARK', 'images')
+
+image_dir = os.path.join(base_dir, 'images')
+anno_dir = os.path.join(base_dir, 'annos')
+
+category_dict = {
+ 1: 'Bicycle', 2: 'Boat', 3: 'Bottle', 4: 'Bus', 5: 'Car', 6: 'Cat',
+ 7: 'Chair', 8: 'Cup', 9: 'Dog', 10: 'Motorbike', 11: 'People', 12: 'Table'
+}
+category_dict_rev = {v: i for i, v in category_dict.items()}
+
+novel_categories = ['Bus', 'Dog', 'Motorbike', 'Table']
+novel_ids = set([category_dict_rev[cate] for cate in novel_categories])
+novel_dict_rev = {v: i for i, v in enumerate(novel_categories)}
+
+split_dict = {
+ 1: 'train', 2: 'val', 3: 'test'
+}
+
+caption_prefix = "A dark image of "
+thr = 15
+
+os.makedirs(os.path.join(PROJECT_DIR, 'data', 'EXDARK', 'metadatas', 'data_setting1'), exist_ok=True)
+
+def save_as_jpg(src_path, dst_dir):
+ src = Path(src_path)
+ dst_dir = Path(dst_dir)
+ dst_dir.mkdir(parents=True, exist_ok=True)
+
+ if src.suffix == ".jpg":
+ shutil.copy(src, dst_dir / src.name)
+ else:
+ img = Image.open(src).convert("RGB")
+ img.save(dst_dir / (src.stem + ".jpg"), "JPEG")
+
+if __name__ == '__main__':
+ train_base_list, train_novel_list = [], [[] for i in range(len(novel_categories) + 1)]
+ val_base_list, val_novel_list = [], [[] for i in range(len(novel_categories) + 1)]
+ test_base_list, test_novel_list = [], [[] for i in range(len(novel_categories) + 1)]
+
+ meta_base_list = [train_base_list, val_base_list, test_base_list]
+ meta_novel_list = [train_novel_list, val_novel_list, test_novel_list]
+
+ with open(os.path.join(base_dir, 'imageclasslist.txt'), 'r') as f:
+ metadata = list(map(lambda line: line.strip().split(), f.readlines()[1:]))
+ metadata = list(map(lambda line: [line[0]] + list(map(int, line[1:])), metadata))
+
+ for data in metadata:
+ image_file = os.path.join(base_dir, 'images', category_dict[data[1]], data[0])
+ save_as_jpg(image_file, os.path.join(output_dir, split_dict[data[-1]]))
+ width, height = imagesize.get(image_file)
+
+ anno_file = os.path.join(base_dir, 'annos', category_dict[data[1]], f'{data[0]}.txt')
+ with open(anno_file, 'r') as f:
+ anno = list(map(lambda line: line.strip().split(), f.readlines()[1:]))
+ anno = list(map(lambda line: [line[0]] + list(map(int, line[1:5])), anno))
+
+ dictin = {}
+ categories, bndboxes, obndboxes= [], [], []
+ categories_in_this_image = set()
+ for object in anno:
+ category, xmin, ymin, w, h = object
+ xmin, ymin, w, h = int(xmin), int(ymin), int(w), int(h)
+ xmin, ymin, xmax, ymax = xmin, ymin, xmin + w, ymin + h
+ xmin = xmin / width
+ ymin = ymin / height
+ xmax = xmax / width
+ ymax = ymax / height
+ categories.append(category)
+ bndboxes.append([xmin, ymin, xmax, ymax])
+ obndboxes.append([xmin, ymin, xmax, ymin, xmax, ymax, xmin, ymax])
+ categories_in_this_image.add(category_dict_rev[category])
+
+ is_novel = False
+ if categories_in_this_image & set(novel_ids):
+ is_novel = True
+ tmp_categories, tmp_bndboxes, tmp_obndboxes = [], [], []
+ for i in range(len(categories)):
+ if categories[i] in novel_categories:
+ tmp_categories.append(categories[i])
+ tmp_bndboxes.append(bndboxes[i])
+ tmp_obndboxes.append(obndboxes[i])
+ categories, bndboxes, obndboxes = tmp_categories, tmp_bndboxes, tmp_obndboxes
+
+ categories = [cate.lower() for cate in categories]
+ caption = [caption_prefix + ", ".join(categories)]
+
+ if len(categories) > thr:
+ categories = categories[:thr]
+ bndboxes = bndboxes[:thr]
+ obndboxes = obndboxes[:thr]
+ while len(categories) < thr:
+ categories.append("")
+ bndboxes.append([0,0,0,0])
+ obndboxes.append([0,0,0,0,0,0,0,0])
+
+ dictin['file_name'] = f'../../images/{split_dict[data[-1]]}/{os.path.splitext(os.path.basename(image_file))[0]}.jpg'
+ caplist = caption + categories
+ dictin["captions"] = caplist
+ dictin["bndboxes"] = bndboxes
+ dictin["obboxes"] = obndboxes
+
+ if is_novel:
+ if len(categories_in_this_image) > 1:
+ meta_novel_list[data[-1] - 1][-1].append(dictin.copy())
+ else:
+ meta_novel_list[data[-1] - 1][novel_dict_rev[category_dict[next(iter(categories_in_this_image))]]].append(dictin.copy())
+ else:
+ meta_base_list[data[-1] - 1].append(dictin.copy())
+ # json_list[data[-1]].append(dictin)
+
+ for i in range(1, 4):
+ with open(os.path.join(PROJECT_DIR, f"data/EXDARK/metadatas/data_setting1/{split_dict[i]}_base.jsonl"), "w", encoding="utf-8") as f:
+ for item in meta_base_list[i - 1]:
+ f.write(json.dumps(item, ensure_ascii=False) + "\n")
+
+ for j in range(len(novel_categories)):
+ with open(os.path.join(PROJECT_DIR, f"data/EXDARK/metadatas/data_setting1/{split_dict[i]}_novel_{novel_categories[j].lower()}.jsonl"), "w", encoding="utf-8") as f:
+ for item in meta_novel_list[i - 1][j]:
+ f.write(json.dumps(item, ensure_ascii=False) + "\n")
+
+ with open(os.path.join(PROJECT_DIR, f"data/EXDARK/metadatas/data_setting1/{split_dict[i]}_novel_mixed.jsonl"), "w", encoding="utf-8") as f:
+ for item in meta_novel_list[i - 1][-1]:
+ f.write(json.dumps(item, ensure_ascii=False) + "\n")
\ No newline at end of file
diff --git a/scripts/data_process/exdark/02.data_crop.py b/scripts/data_process/exdark/02.data_crop.py
new file mode 100644
index 0000000000000000000000000000000000000000..ca4d058dce12bf82cd6e541bd7e7919e4ed08a21
--- /dev/null
+++ b/scripts/data_process/exdark/02.data_crop.py
@@ -0,0 +1,71 @@
+from pathlib import Path
+from PIL import Image
+import shutil
+import os
+import imagesize
+import json
+import cv2
+from collections import Counter
+
+
+PROJECT_DIR = os.getenv('DSP_PROJECT_DIR', '/path/to/DSP_PROJECT_DIR') # Set this manually if the environment variable is unavailable
+base_dir = '/path/to/ExDark' # Replace with your actual path
+output_dir = os.path.join(PROJECT_DIR, "data", "EXDARK", "patches")
+
+image_dir = os.path.join(base_dir, 'images')
+anno_dir = os.path.join(base_dir, 'annos')
+
+category_dict = {
+ 1: 'Bicycle', 2: 'Boat', 3: 'Bottle', 4: 'Bus', 5: 'Car', 6: 'Cat',
+ 7: 'Chair', 8: 'Cup', 9: 'Dog', 10: 'Motorbike', 11: 'People', 12: 'Table'
+}
+
+if __name__ == '__main__':
+ with open(os.path.join(base_dir, 'imageclasslist.txt'), 'r') as f:
+ metadata = list(map(lambda line: line.strip().split(), f.readlines()[1:]))
+ metadata = list(map(lambda line: [line[0]] + list(map(int, line[1:])), metadata))
+
+ counter = Counter()
+ for data in metadata:
+ if data[-1] != 1:
+ continue
+ assert data[-1] == 1
+
+ image_file = os.path.join(base_dir, 'images', category_dict[data[1]], data[0])
+ # width, height = imagesize.get(image_file)
+ image = cv2.imread(image_file)
+ image_height, image_width, _ = image.shape
+
+ anno_file = os.path.join(base_dir, 'annos', category_dict[data[1]], f'{data[0]}.txt')
+ with open(anno_file, 'r') as f:
+ anno = list(map(lambda line: line.strip().split(), f.readlines()[1:]))
+ anno = list(map(lambda line: [line[0].lower()] + list(map(int, line[1:5])), anno))
+
+ categories, bndboxes, obndboxes= [], [], []
+ for i, object in enumerate(anno):
+ category, xmin, ymin, w, h = object
+ xmin, ymin, w, h = int(xmin), int(ymin), int(w), int(h)
+ xmin = max(xmin, 0)
+ xmin, ymin, xmax, ymax = xmin, ymin, xmin + w, ymin + h
+ bbox_width, bbox_height = w, h
+
+ bbox_area = bbox_width * bbox_height
+ image_area = image_width * image_height
+ bbox_ratio = bbox_area / image_area
+
+ if bbox_ratio < 0.001:
+ continue
+
+ class_dir = os.path.join(output_dir, category)
+ os.makedirs(class_dir, exist_ok=True)
+ counter[class_dir] += 1
+
+ cropped_image = image[ymin:ymax, xmin:xmax]
+
+ output_image_name = f"{os.path.splitext(data[0])[0]}_{i}.jpg"
+ output_image_path = os.path.join(class_dir, output_image_name)
+ try:
+ cv2.imwrite(output_image_path, cropped_image)
+ except:
+ import pdb; pdb.set_trace()
+ print(counter)
\ No newline at end of file
diff --git a/scripts/data_process/exdark/03.cache_embs.py b/scripts/data_process/exdark/03.cache_embs.py
new file mode 100644
index 0000000000000000000000000000000000000000..bc55438932eb93834c78bee90c1afbe01d8f72fa
--- /dev/null
+++ b/scripts/data_process/exdark/03.cache_embs.py
@@ -0,0 +1,63 @@
+import os, json, random
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from tqdm import tqdm
+from transformers import CLIPModel, AutoTokenizer, AutoProcessor
+from PIL import Image
+
+PROJECT_DIR = os.getenv('DSP_PROJECT_DIR', '/path/to/DSP_PROJECT_DIR') # Set this manually if the environment variable is unavailable
+data_setting_path = os.path.join(PROJECT_DIR, 'data', 'EXDARK', 'metadatas', 'data_setting1')
+cache_name = os.path.join(PROJECT_DIR, 'data', 'EXDARK', 'exdark_emb.pt')
+
+class myCLIPEnc(nn.Module):
+ def __init__(self, model_config='openai/clip-vit-large-patch14', device='cuda'):
+ super().__init__()
+ self.device = device
+
+ self.tokenizer = AutoTokenizer.from_pretrained(model_config)
+ self.processor = AutoProcessor.from_pretrained(model_config)
+ self.model = CLIPModel.from_pretrained(model_config).to(device)
+ self.model.eval()
+
+ def forward(self, caption=None, img=None):
+ if caption is not None:
+ txt_inp = self.tokenizer(caption, padding=True, truncation=True, return_tensors="pt").to(self.device) # pad and truncate to the max_length
+ txt_feat = self.model.get_text_features(**txt_inp)
+ txt_feat = F.normalize(txt_feat, dim=-1).detach().cpu()
+ else:
+ txt_feat = None
+
+ if img is not None:
+ img_inp = self.processor(images=img, return_tensors="pt").to(self.device)
+ img_feat = self.model.get_image_features(**img_inp)
+ img_feat = F.normalize(img_feat, dim=-1).detach().cpu()
+ else:
+ img_feat = None
+
+ return txt_feat, img_feat
+
+
+if __name__ == '__main__':
+ if not os.path.exists(cache_name):
+ myCLIP = myCLIPEnc()
+
+ data = []
+ with open(os.path.join(data_setting_path, 'train_base.jsonl'), 'r') as f:
+ for line in f:
+ data.append(json.loads(line))
+ emb_dict = {}
+ sample_data = random.sample(data, 1600)
+ for sample in tqdm(sample_data):
+ img_name = sample['file_name']
+ emb_dict[os.path.basename(img_name)] = {}
+
+ caption = sample['captions'][0]
+ img = Image.open(os.path.join(data_setting_path, img_name)).convert('RGB')
+ txt_emb, img_emb = myCLIP(caption=caption, img=img)
+
+ emb_dict[os.path.basename(img_name)]['txt_emb'] = txt_emb.detach().cpu()
+ emb_dict[os.path.basename(img_name)]['img_emb'] = img_emb.detach().cpu()
+ # torch.cuda.empty_cache()
+
+ torch.save(emb_dict, cache_name)
\ No newline at end of file
diff --git a/scripts/data_process/exdark/04.cache_image_sizes.py b/scripts/data_process/exdark/04.cache_image_sizes.py
new file mode 100644
index 0000000000000000000000000000000000000000..3242bb811a5bbaa2f9a059a2f2c98e225f4da760
--- /dev/null
+++ b/scripts/data_process/exdark/04.cache_image_sizes.py
@@ -0,0 +1,42 @@
+import os
+import json
+import imagesize
+from tqdm import tqdm
+
+PROJECT_DIR = os.getenv('DSP_PROJECT_DIR', '/path/to/DSP_PROJECT_DIR') # Set this manually if the environment variable is unavailable
+image_patch_path = os.path.join(PROJECT_DIR, "data", "EXDARK", "patches")
+output_path = os.path.join(image_patch_path, "image_sizes.json")
+
+def generate_size_cache(image_patch_path, output_path):
+ print(f"Scanning: {image_patch_path}")
+ size_cache = {}
+
+ categories = [d for d in os.listdir(image_patch_path) if os.path.isdir(os.path.join(image_patch_path, d))]
+
+ for category in tqdm(categories, desc="Processing"):
+ category_dir = os.path.join(image_patch_path, category)
+
+ for item in os.listdir(category_dir):
+ if item.endswith('.jpg'):
+ img_path = os.path.join(category_dir, item)
+ try:
+ size_cache[img_path] = imagesize.get(img_path)
+ except Exception as e:
+ print(f"Failed {img_path}: {e}")
+
+ aug_dir = os.path.join(category_dir, 'augmented')
+ if os.path.exists(aug_dir):
+ for item in os.listdir(aug_dir):
+ if item.endswith('.jpg'):
+ img_path = os.path.join(aug_dir, item)
+ try:
+ size_cache[img_path] = imagesize.get(img_path)
+ except Exception as e:
+ print(f"Failed {img_path}: {e}")
+
+ with open(output_path, 'w') as f:
+ json.dump(size_cache, f)
+ print(f"\nSaved: {output_path} (Containing {len(size_cache)} images)")
+
+if __name__ == "__main__":
+ generate_size_cache(image_patch_path, output_path)
\ No newline at end of file
diff --git a/scripts/data_process/ruod/00.link_image_files.sh b/scripts/data_process/ruod/00.link_image_files.sh
new file mode 100644
index 0000000000000000000000000000000000000000..446a4202b95bc1701dbb765d4ffbd25fcc899cce
--- /dev/null
+++ b/scripts/data_process/ruod/00.link_image_files.sh
@@ -0,0 +1,4 @@
+#!/usr/bin/env bash
+mkdir -p $DSP_PROJECT_DIR/data/RUOD
+ln -s /path/to/RUOD/RUOD_pic $DSP_PROJECT_DIR/data/RUOD/images # Replace with your actual path
+cp ./008431.jpg $DSP_PROJECT_DIR/data/RUOD/images/train # overwrite original image with corrected (rotated) version
\ No newline at end of file
diff --git a/scripts/data_process/ruod/008431.jpg b/scripts/data_process/ruod/008431.jpg
new file mode 100644
index 0000000000000000000000000000000000000000..785d04e198d69383527fc690cba9265ad333d889
--- /dev/null
+++ b/scripts/data_process/ruod/008431.jpg
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:25187bcb8a65463d286401b0da286964ccbbe4cc290f7f098b757daef10618ba
+size 166000
diff --git a/scripts/data_process/ruod/01.a.process_train.py b/scripts/data_process/ruod/01.a.process_train.py
new file mode 100644
index 0000000000000000000000000000000000000000..1a554001b535a00166d1a6d5c1a76eb5716b003f
--- /dev/null
+++ b/scripts/data_process/ruod/01.a.process_train.py
@@ -0,0 +1,106 @@
+import os
+import xml.etree.ElementTree as ET
+import json
+from collections import defaultdict
+
+PROJECT_DIR = os.getenv('DSP_PROJECT_DIR', '/path/to/DSP_PROJECT_DIR') # Set this manually if the environment variable is unavailable
+base_dir = '/path/to/RUOD' # Replace with your actual path
+anno_path = os.path.join(base_dir, 'RUOD_ANN', 'instances_train.json')
+
+annotations = json.load(open(anno_path,"r"))
+images_items = annotations["images"]
+annos_items = annotations["annotations"]
+cates_items = annotations["categories"]
+category_dict = {}
+for cate in cates_items:
+ category_dict[cate["id"]] = cate["name"]
+
+category_dict_rev = {v: i for i, v in category_dict.items()}
+
+novel_categories = ['corals', 'cuttlefish', 'turtle', 'jellyfish',]
+novel_ids = set([category_dict_rev[cate] for cate in novel_categories])
+novel_dict_rev = {v: i for i, v in enumerate(novel_categories)}
+
+num_classes = 10
+caption_prefix = "An underwater image of "
+thr = 15
+
+os.makedirs(os.path.join(PROJECT_DIR, 'data', 'RUOD', 'metadatas', 'data_setting1'), exist_ok=True)
+
+annos_dict = defaultdict(list)
+for item in annos_items:
+ image_id = item["image_id"]
+ filename = images_items[image_id-1]['file_name']
+ annos_dict[filename].append(item["bbox"] + [item["category_id"]])
+
+if __name__ == '__main__':
+ base_list, novel_list = [], [[] for i in range(len(novel_categories) + 1)]
+ for image_item in images_items:
+ dictin = {}
+ dictin['file_name'] = image_item['file_name']
+ width, height = image_item['width'], image_item['height']
+ categories_in_this_image = set()
+ categories, bndboxes, obndboxes= [], [], []
+ annos = annos_dict[image_item['file_name']]
+ for anno in annos:
+ xmin, ymin, w, h, category_id = anno
+ xmin, ymin, w, h = int(xmin), int(ymin), int(w), int(h)
+ xmin, ymin, xmax, ymax = xmin, ymin, xmin + w, ymin + h
+ xmin = xmin / width
+ ymin = ymin / height
+ xmax = xmax / width
+ ymax = ymax / height
+ categories.append(category_dict[category_id])
+ bndboxes.append([xmin, ymin, xmax, ymax])
+ obndboxes.append([xmin, ymin, xmax, ymin, xmax, ymax, xmin, ymax])
+ categories_in_this_image.add(category_id)
+
+ is_novel = False
+ if categories_in_this_image & set(novel_ids):
+ is_novel = True
+ tmp_categories, tmp_bndboxes, tmp_obndboxes= [], [], []
+ for i in range(len(categories)):
+ if categories[i] in novel_categories:
+ tmp_categories.append(categories[i])
+ tmp_bndboxes.append(bndboxes[i])
+ tmp_obndboxes.append(obndboxes[i])
+ categories, bndboxes, obndboxes = tmp_categories, tmp_bndboxes, tmp_obndboxes
+
+ caption = [caption_prefix + ", ".join(categories)]
+
+ if len(categories) > thr:
+ categories = categories[:thr]
+ bndboxes = bndboxes[:thr]
+ obndboxes = obndboxes[:thr]
+ while len(categories) < thr:
+ categories.append("")
+ bndboxes.append([0,0,0,0])
+ obndboxes.append([0,0,0,0,0,0,0,0])
+
+ dictin["file_name"] = f"../../images/train/{image_item['file_name']}"
+ caplist = caption + categories
+ dictin["captions"] = caplist
+ dictin["bndboxes"] = bndboxes
+ dictin["obboxes"] = obndboxes
+
+ if is_novel:
+ if len(categories_in_this_image) > 1:
+ novel_list[-1].append(dictin.copy())
+ else:
+ novel_list[novel_dict_rev[category_dict[next(iter(categories_in_this_image))]]].append(dictin.copy())
+ else:
+ base_list.append(dictin.copy())
+
+ with open(os.path.join(PROJECT_DIR, "data/RUOD/metadatas/data_setting1/train_base.jsonl"), "w", encoding="utf-8") as f:
+ for item in base_list:
+ f.write(json.dumps(item, ensure_ascii=False) + "\n")
+
+ for i in range(len(novel_categories)):
+ with open(os.path.join(PROJECT_DIR, f"data/RUOD/metadatas/data_setting1/train_novel_{novel_categories[i]}.jsonl"), "w", encoding="utf-8") as f:
+ for item in novel_list[i]:
+ f.write(json.dumps(item, ensure_ascii=False) + "\n")
+
+ with open(os.path.join(PROJECT_DIR, "data/RUOD/metadatas/data_setting1/train_novel_mixed.jsonl"), "w", encoding="utf-8") as f:
+ for item in novel_list[-1]:
+ f.write(json.dumps(item, ensure_ascii=False) + "\n")
+
\ No newline at end of file
diff --git a/scripts/data_process/ruod/01.b.process_test.py b/scripts/data_process/ruod/01.b.process_test.py
new file mode 100644
index 0000000000000000000000000000000000000000..ee666565fc1e14545c9250d935b906d8ce3eaca4
--- /dev/null
+++ b/scripts/data_process/ruod/01.b.process_test.py
@@ -0,0 +1,106 @@
+import os
+import xml.etree.ElementTree as ET
+import json
+from collections import defaultdict
+
+PROJECT_DIR = os.getenv('DSP_PROJECT_DIR', '/path/to/DSP_PROJECT_DIR') # Set this manually if the environment variable is unavailable
+base_dir = '/path/to/RUOD' # Replace with your actual path
+anno_path = os.path.join(base_dir, 'RUOD_ANN', 'instances_test.json')
+
+annotations = json.load(open(anno_path,"r"))
+images_items = annotations["images"]
+annos_items = annotations["annotations"]
+cates_items = annotations["categories"]
+category_dict = {}
+for cate in cates_items:
+ category_dict[cate["id"]] = cate["name"]
+
+category_dict_rev = {v: i for i, v in category_dict.items()}
+
+novel_categories = ['corals', 'cuttlefish', 'turtle', 'jellyfish',]
+novel_ids = set([category_dict_rev[cate] for cate in novel_categories])
+novel_dict_rev = {v: i for i, v in enumerate(novel_categories)}
+
+num_classes = 10
+caption_prefix = "An underwater image of "
+thr = 15
+
+os.makedirs(os.path.join(PROJECT_DIR, 'data', 'RUOD', 'metadatas', 'data_setting1'), exist_ok=True)
+
+annos_dict = defaultdict(list)
+for item in annos_items:
+ image_id = item["image_id"]
+ filename = images_items[image_id-1]['file_name']
+ annos_dict[filename].append(item["bbox"] + [item["category_id"]])
+
+if __name__ == '__main__':
+ base_list, novel_list = [], [[] for i in range(len(novel_categories) + 1)]
+ for image_item in images_items:
+ dictin = {}
+ dictin['file_name'] = image_item['file_name']
+ width, height = image_item['width'], image_item['height']
+ categories_in_this_image = set()
+ categories, bndboxes, obndboxes= [], [], []
+ annos = annos_dict[image_item['file_name']]
+ for anno in annos:
+ xmin, ymin, w, h, category_id = anno
+ xmin, ymin, w, h = int(xmin), int(ymin), int(w), int(h)
+ xmin, ymin, xmax, ymax = xmin, ymin, xmin + w, ymin + h
+ xmin = xmin / width
+ ymin = ymin / height
+ xmax = xmax / width
+ ymax = ymax / height
+ categories.append(category_dict[category_id])
+ bndboxes.append([xmin, ymin, xmax, ymax])
+ obndboxes.append([xmin, ymin, xmax, ymin, xmax, ymax, xmin, ymax])
+ categories_in_this_image.add(category_id)
+
+ is_novel = False
+ if categories_in_this_image & set(novel_ids):
+ is_novel = True
+ tmp_categories, tmp_bndboxes, tmp_obndboxes= [], [], []
+ for i in range(len(categories)):
+ if categories[i] in novel_categories:
+ tmp_categories.append(categories[i])
+ tmp_bndboxes.append(bndboxes[i])
+ tmp_obndboxes.append(obndboxes[i])
+ categories, bndboxes, obndboxes = tmp_categories, tmp_bndboxes, tmp_obndboxes
+
+ caption = [caption_prefix + ", ".join(categories)]
+
+ if len(categories) > thr:
+ categories = categories[:thr]
+ bndboxes = bndboxes[:thr]
+ obndboxes = obndboxes[:thr]
+ while len(categories) < thr:
+ categories.append("")
+ bndboxes.append([0,0,0,0])
+ obndboxes.append([0,0,0,0,0,0,0,0])
+
+ dictin["file_name"] = f"../../images/test/{image_item['file_name']}"
+ caplist = caption + categories
+ dictin["captions"] = caplist
+ dictin["bndboxes"] = bndboxes
+ dictin["obboxes"] = obndboxes
+
+ if is_novel:
+ if len(categories_in_this_image) > 1:
+ novel_list[-1].append(dictin.copy())
+ else:
+ novel_list[novel_dict_rev[category_dict[next(iter(categories_in_this_image))]]].append(dictin.copy())
+ else:
+ base_list.append(dictin.copy())
+
+ with open(os.path.join(PROJECT_DIR, "data/RUOD/metadatas/data_setting1/test_base.jsonl"), "w", encoding="utf-8") as f:
+ for item in base_list:
+ f.write(json.dumps(item, ensure_ascii=False) + "\n")
+
+ for i in range(len(novel_categories)):
+ with open(os.path.join(PROJECT_DIR, f"data/RUOD/metadatas/data_setting1/test_novel_{novel_categories[i]}.jsonl"), "w", encoding="utf-8") as f:
+ for item in novel_list[i]:
+ f.write(json.dumps(item, ensure_ascii=False) + "\n")
+
+ with open(os.path.join(PROJECT_DIR, "data/RUOD/metadatas/data_setting1/test_novel_mixed.jsonl"), "w", encoding="utf-8") as f:
+ for item in novel_list[-1]:
+ f.write(json.dumps(item, ensure_ascii=False) + "\n")
+
\ No newline at end of file
diff --git a/scripts/data_process/ruod/02.data_crop.py b/scripts/data_process/ruod/02.data_crop.py
new file mode 100644
index 0000000000000000000000000000000000000000..4061818c12ba927aabdaa557bfb258494ddde750
--- /dev/null
+++ b/scripts/data_process/ruod/02.data_crop.py
@@ -0,0 +1,82 @@
+import os
+import json
+import cv2
+import torch
+from torchvision import transforms
+from PIL import Image
+from collections import defaultdict
+import xml.etree.ElementTree as ET
+from collections import Counter
+
+PROJECT_DIR = os.getenv('DSP_PROJECT_DIR', '/path/to/DSP_PROJECT_DIR') # Set this manually if the environment variable is unavailable
+image_dir = '/path/to/RUOD/RUOD_pic/train' # Replace with your actual path
+label_dir = '/path/to/RUOD/RUOD_ANN' # Replace with your actual path
+
+output_dir = os.path.join(PROJECT_DIR, "data", "RUOD", "patches")
+
+os.makedirs(output_dir, exist_ok=True)
+
+annos = json.load(open(os.path.join(label_dir, "instances_train.json"), "r"))
+
+
+images_items = annos["images"]
+annos_items = annos["annotations"]
+cates_items = annos["categories"]
+catemap = {}
+for cate in cates_items:
+ catemap[cate["id"]] = cate["name"]
+
+
+files = [i["file_name"] for i in images_items]
+labels = defaultdict(list)
+for item in annos_items:
+ image_id = item["image_id"]
+ filename = files[image_id-1]
+ labels[filename].append([catemap[item["category_id"]]] + item["bbox"])
+
+print(len(files))
+counter = Counter()
+for image_name in files:
+ if not image_name.endswith(".jpg"):
+ continue
+
+ image_path = os.path.join(image_dir, image_name)
+ image = cv2.imread(image_path)
+ image_height, image_width, _ = image.shape
+ lines = labels[image_name]
+ # if image_name == '008431.jpg':
+ # import pdb; pdb.set_trace()
+
+ for i,line in enumerate(lines):
+ parts = line
+
+ class_name = parts[0]
+ xmin, ymin, w, h = parts[1:]
+ bbox_width = w
+ bbox_height = h
+ xmax = xmin + w
+ ymax = ymin + h
+
+
+ bbox_area = bbox_width * bbox_height
+ image_area = image_width * image_height
+ bbox_ratio = bbox_area / image_area
+
+ if bbox_ratio < 0.001:
+ continue
+
+ class_dir = os.path.join(output_dir, class_name)
+ os.makedirs(class_dir, exist_ok=True)
+ counter[class_dir] += 1
+ xmin, ymin, xmax, ymax = int(xmin), int(ymin), int(xmax), int(ymax)
+ cropped_image = image[ymin:ymax, xmin:xmax]
+
+ output_image_name = f"{image_name[:-4]}_{i}.jpg"
+ output_image_path = os.path.join(class_dir, output_image_name)
+ try:
+ cv2.imwrite(output_image_path, cropped_image)
+ except:
+ import pdb; pdb.set_trace()
+ print(bbox_ratio, xmin, ymin, xmax, ymax, image_name)
+
+print(counter)
diff --git a/scripts/data_process/ruod/03.cache_embs.py b/scripts/data_process/ruod/03.cache_embs.py
new file mode 100644
index 0000000000000000000000000000000000000000..9b3fe354f7cfdf696a3fdac1be30c9b88551336c
--- /dev/null
+++ b/scripts/data_process/ruod/03.cache_embs.py
@@ -0,0 +1,63 @@
+import os, json, random
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from tqdm import tqdm
+from transformers import CLIPModel, AutoTokenizer, AutoProcessor
+from PIL import Image
+
+PROJECT_DIR = os.getenv('DSP_PROJECT_DIR', '/path/to/DSP_PROJECT_DIR') # Set this manually if the environment variable is unavailable
+data_setting_path = os.path.join(PROJECT_DIR, 'data', 'RUOD', 'metadatas', 'data_setting1')
+cache_name = os.path.join(PROJECT_DIR, 'data', 'RUOD', 'ruod_emb.pt')
+
+class myCLIPEnc(nn.Module):
+ def __init__(self, model_config='openai/clip-vit-large-patch14', device='cuda'):
+ super().__init__()
+ self.device = device
+
+ self.tokenizer = AutoTokenizer.from_pretrained(model_config)
+ self.processor = AutoProcessor.from_pretrained(model_config)
+ self.model = CLIPModel.from_pretrained(model_config).to(device)
+ self.model.eval()
+
+ def forward(self, caption=None, img=None):
+ if caption is not None:
+ txt_inp = self.tokenizer(caption, padding=True, truncation=True, return_tensors="pt").to(self.device) # pad and truncate to the max_length
+ txt_feat = self.model.get_text_features(**txt_inp)
+ txt_feat = F.normalize(txt_feat, dim=-1).detach().cpu()
+ else:
+ txt_feat = None
+
+ if img is not None:
+ img_inp = self.processor(images=img, return_tensors="pt").to(self.device)
+ img_feat = self.model.get_image_features(**img_inp)
+ img_feat = F.normalize(img_feat, dim=-1).detach().cpu()
+ else:
+ img_feat = None
+
+ return txt_feat, img_feat
+
+
+if __name__ == '__main__':
+ if not os.path.exists(cache_name):
+ myCLIP = myCLIPEnc()
+
+ data = []
+ with open(os.path.join(data_setting_path, 'train_base.jsonl'), 'r') as f:
+ for line in f:
+ data.append(json.loads(line))
+ emb_dict = {}
+ sample_data = random.sample(data, 4000)
+ for sample in tqdm(sample_data):
+ img_name = sample['file_name']
+ emb_dict[os.path.basename(img_name)] = {}
+
+ caption = sample['captions'][0]
+ img = Image.open(os.path.join(data_setting_path, img_name)).convert('RGB')
+ txt_emb, img_emb = myCLIP(caption=caption, img=img)
+
+ emb_dict[os.path.basename(img_name)]['txt_emb'] = txt_emb.detach().cpu()
+ emb_dict[os.path.basename(img_name)]['img_emb'] = img_emb.detach().cpu()
+ # torch.cuda.empty_cache()
+
+ torch.save(emb_dict, cache_name)
\ No newline at end of file
diff --git a/scripts/data_process/ruod/04.cache_image_sizes.py b/scripts/data_process/ruod/04.cache_image_sizes.py
new file mode 100644
index 0000000000000000000000000000000000000000..c50e1f528d40a1e75699ca7d41e0179b741c099a
--- /dev/null
+++ b/scripts/data_process/ruod/04.cache_image_sizes.py
@@ -0,0 +1,42 @@
+import os
+import json
+import imagesize
+from tqdm import tqdm
+
+PROJECT_DIR = os.getenv('DSP_PROJECT_DIR', '/path/to/DSP_PROJECT_DIR') # Set this manually if the environment variable is unavailable
+image_patch_path = os.path.join(PROJECT_DIR, "data", "RUOD", "patches")
+output_path = os.path.join(image_patch_path, "image_sizes.json")
+
+def generate_size_cache(image_patch_path, output_path):
+ print(f"Scanning: {image_patch_path}")
+ size_cache = {}
+
+ categories = [d for d in os.listdir(image_patch_path) if os.path.isdir(os.path.join(image_patch_path, d))]
+
+ for category in tqdm(categories, desc="Processing"):
+ category_dir = os.path.join(image_patch_path, category)
+
+ for item in os.listdir(category_dir):
+ if item.endswith('.jpg'):
+ img_path = os.path.join(category_dir, item)
+ try:
+ size_cache[img_path] = imagesize.get(img_path)
+ except Exception as e:
+ print(f"Failed {img_path}: {e}")
+
+ aug_dir = os.path.join(category_dir, 'augmented')
+ if os.path.exists(aug_dir):
+ for item in os.listdir(aug_dir):
+ if item.endswith('.jpg'):
+ img_path = os.path.join(aug_dir, item)
+ try:
+ size_cache[img_path] = imagesize.get(img_path)
+ except Exception as e:
+ print(f"Failed {img_path}: {e}")
+
+ with open(output_path, 'w') as f:
+ json.dump(size_cache, f)
+ print(f"\nSaved: {output_path} (Containing {len(size_cache)} images)")
+
+if __name__ == "__main__":
+ generate_size_cache(image_patch_path, output_path)
\ No newline at end of file
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/ade20k_instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/ade20k_instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..57f657aa67f34830515f410425eccc96cb065af4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/ade20k_instance.py
@@ -0,0 +1,53 @@
+# dataset settings
+dataset_type = 'ADE20KInstanceDataset'
+data_root = 'data/ADEChallengeData2016/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/ADEChallengeData2016/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(2560, 640), keep_ratio=True),
+ # If you don't have a gt annotation, delete the pipeline
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='ade20k_instance_val.json',
+ data_prefix=dict(img='images/validation'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'ade20k_instance_val.json',
+ metric=['bbox', 'segm'],
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/ade20k_panoptic.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/ade20k_panoptic.py
new file mode 100644
index 0000000000000000000000000000000000000000..7be5ddd7f0732193f4f92bc49e52493602928162
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/ade20k_panoptic.py
@@ -0,0 +1,38 @@
+# dataset settings
+dataset_type = 'ADE20KPanopticDataset'
+data_root = 'data/ADEChallengeData2016/'
+
+backend_args = None
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(2560, 640), keep_ratio=True),
+ dict(type='LoadPanopticAnnotations', backend_args=backend_args),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=0,
+ persistent_workers=False,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='ade20k_panoptic_val.json',
+ data_prefix=dict(img='images/validation/', seg='ade20k_panoptic_val/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoPanopticMetric',
+ ann_file=data_root + 'ade20k_panoptic_val.json',
+ seg_prefix=data_root + 'ade20k_panoptic_val/',
+ backend_args=backend_args)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/ade20k_semantic.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/ade20k_semantic.py
new file mode 100644
index 0000000000000000000000000000000000000000..522a775704182ededaa36f318cd1eb185784918f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/ade20k_semantic.py
@@ -0,0 +1,48 @@
+dataset_type = 'ADE20KSegDataset'
+data_root = 'data/ADEChallengeData2016/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/ADEChallengeData2016/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(2048, 512), keep_ratio=True),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=False,
+ with_mask=False,
+ with_seg=True,
+ reduce_zero_label=True),
+ dict(
+ type='PackDetInputs', meta_keys=('img_path', 'ori_shape', 'img_shape'))
+]
+
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ data_prefix=dict(
+ img_path='images/validation',
+ seg_map_path='annotations/validation'),
+ pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(type='SemSegMetric', iou_metrics=['mIoU'])
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/cityscapes_detection.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/cityscapes_detection.py
new file mode 100644
index 0000000000000000000000000000000000000000..caeba6bfcd26d8954fc9d499446e93323e372959
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/cityscapes_detection.py
@@ -0,0 +1,84 @@
+# dataset settings
+dataset_type = 'CityscapesDataset'
+data_root = 'data/cityscapes/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/segmentation/cityscapes/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/segmentation/',
+# 'data/': 's3://openmmlab/datasets/segmentation/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize',
+ scale=[(2048, 800), (2048, 1024)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(2048, 1024), keep_ratio=True),
+ # If you don't have a gt annotation, delete the pipeline
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type='RepeatDataset',
+ times=8,
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instancesonly_filtered_gtFine_train.json',
+ data_prefix=dict(img='leftImg8bit/train/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args)))
+
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instancesonly_filtered_gtFine_val.json',
+ data_prefix=dict(img='leftImg8bit/val/'),
+ test_mode=True,
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/instancesonly_filtered_gtFine_val.json',
+ metric='bbox',
+ backend_args=backend_args)
+
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/cityscapes_instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/cityscapes_instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..136403136c67a6726662832b66f56701ff5aba8a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/cityscapes_instance.py
@@ -0,0 +1,113 @@
+# dataset settings
+dataset_type = 'CityscapesDataset'
+data_root = 'data/cityscapes/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/segmentation/cityscapes/'
+
+# Method 2: Use backend_args, file_client_args in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/segmentation/',
+# 'data/': 's3://openmmlab/datasets/segmentation/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomResize',
+ scale=[(2048, 800), (2048, 1024)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(2048, 1024), keep_ratio=True),
+ # If you don't have a gt annotation, delete the pipeline
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type='RepeatDataset',
+ times=8,
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instancesonly_filtered_gtFine_train.json',
+ data_prefix=dict(img='leftImg8bit/train/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args)))
+
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instancesonly_filtered_gtFine_val.json',
+ data_prefix=dict(img='leftImg8bit/val/'),
+ test_mode=True,
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+
+test_dataloader = val_dataloader
+
+val_evaluator = [
+ dict(
+ type='CocoMetric',
+ ann_file=data_root +
+ 'annotations/instancesonly_filtered_gtFine_val.json',
+ metric=['bbox', 'segm'],
+ backend_args=backend_args),
+ dict(
+ type='CityScapesMetric',
+ seg_prefix=data_root + 'gtFine/val',
+ outfile_prefix='./work_dirs/cityscapes_metric/instance',
+ backend_args=backend_args)
+]
+
+test_evaluator = val_evaluator
+
+# inference on test dataset and
+# format the output results for submission.
+# test_dataloader = dict(
+# batch_size=1,
+# num_workers=2,
+# persistent_workers=True,
+# drop_last=False,
+# sampler=dict(type='DefaultSampler', shuffle=False),
+# dataset=dict(
+# type=dataset_type,
+# data_root=data_root,
+# ann_file='annotations/instancesonly_filtered_gtFine_test.json',
+# data_prefix=dict(img='leftImg8bit/test/'),
+# test_mode=True,
+# filter_cfg=dict(filter_empty_gt=True, min_size=32),
+# pipeline=test_pipeline))
+# test_evaluator = dict(
+# type='CityScapesMetric',
+# format_only=True,
+# outfile_prefix='./work_dirs/cityscapes_metric/test')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/coco_caption.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/coco_caption.py
new file mode 100644
index 0000000000000000000000000000000000000000..a1bd898313927e4fca336dfa10f05e78b9fb7162
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/coco_caption.py
@@ -0,0 +1,60 @@
+# data settings
+
+dataset_type = 'CocoCaptionDataset'
+data_root = 'data/coco/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ imdecode_backend='pillow',
+ backend_args=backend_args),
+ dict(
+ type='Resize',
+ scale=(224, 224),
+ interpolation='bicubic',
+ backend='pillow'),
+ dict(type='PackInputs', meta_keys=['image_id']),
+]
+
+# ann_file download from
+# train dataset: https://storage.googleapis.com/sfr-vision-language-research/datasets/coco_karpathy_train.json # noqa
+# val dataset: https://storage.googleapis.com/sfr-vision-language-research/datasets/coco_karpathy_val.json # noqa
+# test dataset: https://storage.googleapis.com/sfr-vision-language-research/datasets/coco_karpathy_test.json # noqa
+# val evaluator: https://storage.googleapis.com/sfr-vision-language-research/datasets/coco_karpathy_val_gt.json # noqa
+# test evaluator: https://storage.googleapis.com/sfr-vision-language-research/datasets/coco_karpathy_test_gt.json # noqa
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/coco_karpathy_val.json',
+ pipeline=test_pipeline,
+ ))
+
+val_evaluator = dict(
+ type='COCOCaptionMetric',
+ ann_file=data_root + 'annotations/coco_karpathy_val_gt.json',
+)
+
+# # If you want standard test, please manually configure the test dataset
+test_dataloader = val_dataloader
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/coco_detection.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/coco_detection.py
new file mode 100644
index 0000000000000000000000000000000000000000..fdf8dfad9476b1d7b7a4e8c3e2832f115a1ea7f2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/coco_detection.py
@@ -0,0 +1,95 @@
+# dataset settings
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ # If you don't have a gt annotation, delete the pipeline
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric='bbox',
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+# inference on test dataset and
+# format the output results for submission.
+# test_dataloader = dict(
+# batch_size=1,
+# num_workers=2,
+# persistent_workers=True,
+# drop_last=False,
+# sampler=dict(type='DefaultSampler', shuffle=False),
+# dataset=dict(
+# type=dataset_type,
+# data_root=data_root,
+# ann_file=data_root + 'annotations/image_info_test-dev2017.json',
+# data_prefix=dict(img='test2017/'),
+# test_mode=True,
+# pipeline=test_pipeline))
+# test_evaluator = dict(
+# type='CocoMetric',
+# metric='bbox',
+# format_only=True,
+# ann_file=data_root + 'annotations/image_info_test-dev2017.json',
+# outfile_prefix='./work_dirs/coco_detection/test')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/coco_instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/coco_instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..e91cb354038db4df3b990b307a5da9d77f341a88
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/coco_instance.py
@@ -0,0 +1,95 @@
+# dataset settings
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ # If you don't have a gt annotation, delete the pipeline
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric=['bbox', 'segm'],
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+# inference on test dataset and
+# format the output results for submission.
+# test_dataloader = dict(
+# batch_size=1,
+# num_workers=2,
+# persistent_workers=True,
+# drop_last=False,
+# sampler=dict(type='DefaultSampler', shuffle=False),
+# dataset=dict(
+# type=dataset_type,
+# data_root=data_root,
+# ann_file=data_root + 'annotations/image_info_test-dev2017.json',
+# data_prefix=dict(img='test2017/'),
+# test_mode=True,
+# pipeline=test_pipeline))
+# test_evaluator = dict(
+# type='CocoMetric',
+# metric=['bbox', 'segm'],
+# format_only=True,
+# ann_file=data_root + 'annotations/image_info_test-dev2017.json',
+# outfile_prefix='./work_dirs/coco_instance/test')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/coco_instance_semantic.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/coco_instance_semantic.py
new file mode 100644
index 0000000000000000000000000000000000000000..cc961863306690c056e564b542d518c0ebfbb7e2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/coco_instance_semantic.py
@@ -0,0 +1,78 @@
+# dataset settings
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(
+ type='LoadAnnotations', with_bbox=True, with_mask=True, with_seg=True),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ # If you don't have a gt annotation, delete the pipeline
+ dict(
+ type='LoadAnnotations', with_bbox=True, with_mask=True, with_seg=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/', seg='stuffthingmaps/train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args))
+
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric=['bbox', 'segm'],
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/coco_panoptic.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/coco_panoptic.py
new file mode 100644
index 0000000000000000000000000000000000000000..0b95b619e68ed531d361bbd11a2382852c13446e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/coco_panoptic.py
@@ -0,0 +1,94 @@
+# dataset settings
+dataset_type = 'CocoPanopticDataset'
+data_root = 'data/coco/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadPanopticAnnotations', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='LoadPanopticAnnotations', backend_args=backend_args),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/panoptic_train2017.json',
+ data_prefix=dict(
+ img='train2017/', seg='annotations/panoptic_train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/panoptic_val2017.json',
+ data_prefix=dict(img='val2017/', seg='annotations/panoptic_val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoPanopticMetric',
+ ann_file=data_root + 'annotations/panoptic_val2017.json',
+ seg_prefix=data_root + 'annotations/panoptic_val2017/',
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+# inference on test dataset and
+# format the output results for submission.
+# test_dataloader = dict(
+# batch_size=1,
+# num_workers=1,
+# persistent_workers=True,
+# drop_last=False,
+# sampler=dict(type='DefaultSampler', shuffle=False),
+# dataset=dict(
+# type=dataset_type,
+# data_root=data_root,
+# ann_file='annotations/panoptic_image_info_test-dev2017.json',
+# data_prefix=dict(img='test2017/'),
+# test_mode=True,
+# pipeline=test_pipeline))
+# test_evaluator = dict(
+# type='CocoPanopticMetric',
+# format_only=True,
+# ann_file=data_root + 'annotations/panoptic_image_info_test-dev2017.json',
+# outfile_prefix='./work_dirs/coco_panoptic/test')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/coco_semantic.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/coco_semantic.py
new file mode 100644
index 0000000000000000000000000000000000000000..944bbbaeaeb6f10f0946bd1fc828bb01ea6c1fc3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/coco_semantic.py
@@ -0,0 +1,78 @@
+# dataset settings
+dataset_type = 'CocoSegDataset'
+data_root = 'data/coco/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=False,
+ with_label=False,
+ with_seg=True),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=False,
+ with_label=False,
+ with_seg=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_path', 'ori_shape', 'img_shape', 'scale_factor'))
+]
+
+# For stuffthingmaps_semseg, please refer to
+# `docs/en/user_guides/dataset_prepare.md`
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ data_prefix=dict(
+ img_path='train2017/',
+ seg_map_path='stuffthingmaps_semseg/train2017/'),
+ pipeline=train_pipeline))
+
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ data_prefix=dict(
+ img_path='val2017/',
+ seg_map_path='stuffthingmaps_semseg/val2017/'),
+ pipeline=test_pipeline))
+
+test_dataloader = val_dataloader
+
+val_evaluator = dict(type='SemSegMetric', iou_metrics=['mIoU'])
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/deepfashion.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/deepfashion.py
new file mode 100644
index 0000000000000000000000000000000000000000..a93dc7152f7a2e28ab726c79f9398a1034b7b4a1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/deepfashion.py
@@ -0,0 +1,95 @@
+# dataset settings
+dataset_type = 'DeepFashionDataset'
+data_root = 'data/DeepFashion/In-shop/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(type='Resize', scale=(750, 1101), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(750, 1101), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type='RepeatDataset',
+ times=2,
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='Anno/segmentation/DeepFashion_segmentation_train.json',
+ data_prefix=dict(img='Img/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args)))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='Anno/segmentation/DeepFashion_segmentation_query.json',
+ data_prefix=dict(img='Img/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='Anno/segmentation/DeepFashion_segmentation_gallery.json',
+ data_prefix=dict(img='Img/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root +
+ 'Anno/segmentation/DeepFashion_segmentation_query.json',
+ metric=['bbox', 'segm'],
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root +
+ 'Anno/segmentation/DeepFashion_segmentation_gallery.json',
+ metric=['bbox', 'segm'],
+ format_only=False,
+ backend_args=backend_args)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/dsdl.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/dsdl.py
new file mode 100644
index 0000000000000000000000000000000000000000..1f19e5e498b18a404f3c4e6419316b5f9981e811
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/dsdl.py
@@ -0,0 +1,62 @@
+dataset_type = 'DSDLDetDataset'
+data_root = 'path to dataset folder'
+train_ann = 'path to train yaml file'
+val_ann = 'path to val yaml file'
+
+backend_args = None
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': "s3://open_data/",
+# 'data/': "s3://open_data/"
+# }))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ # If you don't have a gt annotation, delete the pipeline
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'instances'))
+]
+
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file=train_ann,
+ filter_cfg=dict(filter_empty_gt=True, min_size=32, bbox_min_size=32),
+ pipeline=train_pipeline))
+
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file=val_ann,
+ test_mode=True,
+ pipeline=test_pipeline))
+
+test_dataloader = val_dataloader
+
+val_evaluator = dict(type='CocoMetric', metric='bbox')
+# val_evaluator = dict(type='VOCMetric', metric='mAP', eval_mode='11points')
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/isaid_instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/isaid_instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..09ddcab02bdd52374d5093d446abb0e34751f7a3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/isaid_instance.py
@@ -0,0 +1,59 @@
+# dataset settings
+dataset_type = 'iSAIDDataset'
+data_root = 'data/iSAID/'
+backend_args = None
+
+# Please see `projects/iSAID/README.md` for data preparation
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(type='Resize', scale=(800, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(800, 800), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='train/instancesonly_filtered_train.json',
+ data_prefix=dict(img='train/images/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='val/instancesonly_filtered_val.json',
+ data_prefix=dict(img='val/images/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'val/instancesonly_filtered_val.json',
+ metric=['bbox', 'segm'],
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/lvis_v0.5_instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/lvis_v0.5_instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..d0ca44efb6d31aae5f6426a1c8b89d2e9be2104f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/lvis_v0.5_instance.py
@@ -0,0 +1,79 @@
+# dataset settings
+dataset_type = 'LVISV05Dataset'
+data_root = 'data/lvis_v0.5/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/lvis_v0.5/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type='ClassBalancedDataset',
+ oversample_thr=1e-3,
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/lvis_v0.5_train.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args)))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/lvis_v0.5_val.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='LVISMetric',
+ ann_file=data_root + 'annotations/lvis_v0.5_val.json',
+ metric=['bbox', 'segm'],
+ backend_args=backend_args)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/lvis_v1_instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/lvis_v1_instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..0413f370a2b635362a60c20881769064bac9a603
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/lvis_v1_instance.py
@@ -0,0 +1,22 @@
+# dataset settings
+_base_ = 'lvis_v0.5_instance.py'
+dataset_type = 'LVISV1Dataset'
+data_root = 'data/lvis_v1/'
+
+train_dataloader = dict(
+ dataset=dict(
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/lvis_v1_train.json',
+ data_prefix=dict(img=''))))
+val_dataloader = dict(
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/lvis_v1_val.json',
+ data_prefix=dict(img='')))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(ann_file=data_root + 'annotations/lvis_v1_val.json')
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/mot_challenge.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/mot_challenge.py
new file mode 100644
index 0000000000000000000000000000000000000000..ce2828ef70a34c123792d252bf992f423049d065
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/mot_challenge.py
@@ -0,0 +1,90 @@
+# dataset settings
+dataset_type = 'MOTChallengeDataset'
+data_root = 'data/MOT17/'
+img_scale = (1088, 1088)
+
+backend_args = None
+# data pipeline
+train_pipeline = [
+ dict(
+ type='UniformRefFrameSample',
+ num_ref_imgs=1,
+ frame_range=10,
+ filter_key_img=True),
+ dict(
+ type='TransformBroadcaster',
+ share_random_params=True,
+ transforms=[
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadTrackAnnotations'),
+ dict(
+ type='RandomResize',
+ scale=img_scale,
+ ratio_range=(0.8, 1.2),
+ keep_ratio=True,
+ clip_object_border=False),
+ dict(type='PhotoMetricDistortion')
+ ]),
+ dict(
+ type='TransformBroadcaster',
+ # different cropped positions for different frames
+ share_random_params=False,
+ transforms=[
+ dict(
+ type='RandomCrop', crop_size=img_scale, bbox_clip_border=False)
+ ]),
+ dict(
+ type='TransformBroadcaster',
+ share_random_params=True,
+ transforms=[
+ dict(type='RandomFlip', prob=0.5),
+ ]),
+ dict(type='PackTrackInputs')
+]
+
+test_pipeline = [
+ dict(
+ type='TransformBroadcaster',
+ transforms=[
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=img_scale, keep_ratio=True),
+ dict(type='LoadTrackAnnotations')
+ ]),
+ dict(type='PackTrackInputs')
+]
+
+# dataloader
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='TrackImgSampler'), # image-based sampling
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ visibility_thr=-1,
+ ann_file='annotations/half-train_cocoformat.json',
+ data_prefix=dict(img_path='train'),
+ metainfo=dict(classes=('pedestrian', )),
+ pipeline=train_pipeline))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ # Now we support two ways to test, image_based and video_based
+ # if you want to use video_based sampling, you can use as follows
+ # sampler=dict(type='DefaultSampler', shuffle=False, round_up=False),
+ sampler=dict(type='TrackImgSampler'), # image-based sampling
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/half-val_cocoformat.json',
+ data_prefix=dict(img_path='train'),
+ test_mode=True,
+ pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# evaluator
+val_evaluator = dict(
+ type='MOTChallengeMetric', metric=['HOTA', 'CLEAR', 'Identity'])
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/mot_challenge_det.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/mot_challenge_det.py
new file mode 100644
index 0000000000000000000000000000000000000000..a988572c3837eb2a8a6bf7b9eca06f3d82abdfda
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/mot_challenge_det.py
@@ -0,0 +1,66 @@
+# dataset settings
+dataset_type = 'CocoDataset'
+data_root = 'data/MOT17/'
+
+backend_args = None
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args, to_float32=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize',
+ scale=(1088, 1088),
+ ratio_range=(0.8, 1.2),
+ keep_ratio=True,
+ clip_object_border=False),
+ dict(type='PhotoMetricDistortion'),
+ dict(type='RandomCrop', crop_size=(1088, 1088), bbox_clip_border=False),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1088, 1088), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/half-train_cocoformat.json',
+ data_prefix=dict(img='train/'),
+ metainfo=dict(classes=('pedestrian', )),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/half-val_cocoformat.json',
+ data_prefix=dict(img='train/'),
+ metainfo=dict(classes=('pedestrian', )),
+ test_mode=True,
+ pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/half-val_cocoformat.json',
+ metric='bbox',
+ format_only=False)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/mot_challenge_reid.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/mot_challenge_reid.py
new file mode 100644
index 0000000000000000000000000000000000000000..57a95b531f3591e60daaabc5eea6f11c7424215b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/mot_challenge_reid.py
@@ -0,0 +1,61 @@
+# dataset settings
+dataset_type = 'ReIDDataset'
+data_root = 'data/MOT17/'
+
+backend_args = None
+# data pipeline
+train_pipeline = [
+ dict(
+ type='TransformBroadcaster',
+ share_random_params=False,
+ transforms=[
+ dict(
+ type='LoadImageFromFile',
+ backend_args=backend_args,
+ to_float32=True),
+ dict(
+ type='Resize',
+ scale=(128, 256),
+ keep_ratio=False,
+ clip_object_border=False),
+ dict(type='RandomFlip', prob=0.5, direction='horizontal'),
+ ]),
+ dict(type='PackReIDInputs', meta_keys=('flip', 'flip_direction'))
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args, to_float32=True),
+ dict(type='Resize', scale=(128, 256), keep_ratio=False),
+ dict(type='PackReIDInputs')
+]
+
+# dataloader
+train_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ triplet_sampler=dict(num_ids=8, ins_per_id=4),
+ data_prefix=dict(img_path='reid/imgs'),
+ ann_file='reid/meta/train_80.txt',
+ pipeline=train_pipeline))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ triplet_sampler=None,
+ data_prefix=dict(img_path='reid/imgs'),
+ ann_file='reid/meta/val_20.txt',
+ pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# evaluator
+val_evaluator = dict(type='ReIDMetrics', metric=['mAP', 'CMC'])
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/objects365v1_detection.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/objects365v1_detection.py
new file mode 100644
index 0000000000000000000000000000000000000000..ee398698608543e13188452a816283e9a2563390
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/objects365v1_detection.py
@@ -0,0 +1,74 @@
+# dataset settings
+dataset_type = 'Objects365V1Dataset'
+data_root = 'data/Objects365/Obj365_v1/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ # If you don't have a gt annotation, delete the pipeline
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/objects365_train.json',
+ data_prefix=dict(img='train/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/objects365_val.json',
+ data_prefix=dict(img='val/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/objects365_val.json',
+ metric='bbox',
+ sort_categories=True,
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/objects365v2_detection.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/objects365v2_detection.py
new file mode 100644
index 0000000000000000000000000000000000000000..b25a7ba901befa8d61e3cdae8a7c68fb8a9c5aef
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/objects365v2_detection.py
@@ -0,0 +1,73 @@
+# dataset settings
+dataset_type = 'Objects365V2Dataset'
+data_root = 'data/Objects365/Obj365_v2/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ # If you don't have a gt annotation, delete the pipeline
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/zhiyuan_objv2_train.json',
+ data_prefix=dict(img='train/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/zhiyuan_objv2_val.json',
+ data_prefix=dict(img='val/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/zhiyuan_objv2_val.json',
+ metric='bbox',
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/openimages_detection.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/openimages_detection.py
new file mode 100644
index 0000000000000000000000000000000000000000..129661b405c70d3e2d0d2c4741e3a59333dd960c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/openimages_detection.py
@@ -0,0 +1,81 @@
+# dataset settings
+dataset_type = 'OpenImagesDataset'
+data_root = 'data/OpenImages/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='Resize', scale=(1024, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1024, 800), keep_ratio=True),
+ # avoid bboxes being resized
+ dict(type='LoadAnnotations', with_bbox=True),
+ # TODO: find a better way to collect image_level_labels
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'instances', 'image_level_labels'))
+]
+
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=0, # workers_per_gpu > 0 may occur out of memory
+ persistent_workers=False,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/oidv6-train-annotations-bbox.csv',
+ data_prefix=dict(img='OpenImages/train/'),
+ label_file='annotations/class-descriptions-boxable.csv',
+ hierarchy_file='annotations/bbox_labels_600_hierarchy.json',
+ meta_file='annotations/train-image-metas.pkl',
+ pipeline=train_pipeline,
+ backend_args=backend_args))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=0,
+ persistent_workers=False,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/validation-annotations-bbox.csv',
+ data_prefix=dict(img='OpenImages/validation/'),
+ label_file='annotations/class-descriptions-boxable.csv',
+ hierarchy_file='annotations/bbox_labels_600_hierarchy.json',
+ meta_file='annotations/validation-image-metas.pkl',
+ image_level_ann_file='annotations/validation-'
+ 'annotations-human-imagelabels-boxable.csv',
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='OpenImagesMetric',
+ iou_thrs=0.5,
+ ioa_thrs=0.5,
+ use_group_of=True,
+ get_supercategory=True)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/refcoco+.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/refcoco+.py
new file mode 100644
index 0000000000000000000000000000000000000000..ae0278ddf6c30fda6e4fb42aed1cb1b9a55109ec
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/refcoco+.py
@@ -0,0 +1,55 @@
+# dataset settings
+dataset_type = 'RefCocoDataset'
+data_root = 'data/coco/'
+
+backend_args = None
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(
+ type='LoadAnnotations',
+ with_mask=True,
+ with_bbox=False,
+ with_seg=False,
+ with_label=False),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'gt_masks', 'text'))
+]
+
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ data_prefix=dict(img_path='train2014/'),
+ ann_file='refcoco+/instances.json',
+ split_file='refcoco+/refs(unc).p',
+ split='val',
+ text_mode='select_first',
+ pipeline=test_pipeline))
+
+test_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ data_prefix=dict(img_path='train2014/'),
+ ann_file='refcoco+/instances.json',
+ split_file='refcoco+/refs(unc).p',
+ split='testA', # or 'testB'
+ text_mode='select_first',
+ pipeline=test_pipeline))
+
+val_evaluator = dict(type='RefSegMetric', metric=['cIoU', 'mIoU'])
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/refcoco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/refcoco.py
new file mode 100644
index 0000000000000000000000000000000000000000..7b6caefa9a4bbfabdb49689588821f99d882a80f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/refcoco.py
@@ -0,0 +1,55 @@
+# dataset settings
+dataset_type = 'RefCocoDataset'
+data_root = 'data/coco/'
+
+backend_args = None
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(
+ type='LoadAnnotations',
+ with_mask=True,
+ with_bbox=False,
+ with_seg=False,
+ with_label=False),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'gt_masks', 'text'))
+]
+
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ data_prefix=dict(img_path='train2014/'),
+ ann_file='refcoco/instances.json',
+ split_file='refcoco/refs(unc).p',
+ split='val',
+ text_mode='select_first',
+ pipeline=test_pipeline))
+
+test_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ data_prefix=dict(img_path='train2014/'),
+ ann_file='refcoco/instances.json',
+ split_file='refcoco/refs(unc).p',
+ split='testA', # or 'testB'
+ text_mode='select_first',
+ pipeline=test_pipeline))
+
+val_evaluator = dict(type='RefSegMetric', metric=['cIoU', 'mIoU'])
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/refcocog.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/refcocog.py
new file mode 100644
index 0000000000000000000000000000000000000000..19dbeef1cde79fcb2aa80bb9936a60cc30089963
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/refcocog.py
@@ -0,0 +1,55 @@
+# dataset settings
+dataset_type = 'RefCocoDataset'
+data_root = 'data/coco/'
+
+backend_args = None
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(
+ type='LoadAnnotations',
+ with_mask=True,
+ with_bbox=False,
+ with_seg=False,
+ with_label=False),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'gt_masks', 'text'))
+]
+
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ data_prefix=dict(img_path='train2014/'),
+ ann_file='refcocog/instances.json',
+ split_file='refcocog/refs(umd).p',
+ split='val',
+ text_mode='select_first',
+ pipeline=test_pipeline))
+
+test_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ data_prefix=dict(img_path='train2014/'),
+ ann_file='refcocog/instances.json',
+ split_file='refcocog/refs(umd).p',
+ split='test',
+ text_mode='select_first',
+ pipeline=test_pipeline))
+
+val_evaluator = dict(type='RefSegMetric', metric=['cIoU', 'mIoU'])
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/semi_coco_detection.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/semi_coco_detection.py
new file mode 100644
index 0000000000000000000000000000000000000000..694f25f841e06dbb59a699dfe13c18e34dbdce9f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/semi_coco_detection.py
@@ -0,0 +1,178 @@
+# dataset settings
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+color_space = [
+ [dict(type='ColorTransform')],
+ [dict(type='AutoContrast')],
+ [dict(type='Equalize')],
+ [dict(type='Sharpness')],
+ [dict(type='Posterize')],
+ [dict(type='Solarize')],
+ [dict(type='Color')],
+ [dict(type='Contrast')],
+ [dict(type='Brightness')],
+]
+
+geometric = [
+ [dict(type='Rotate')],
+ [dict(type='ShearX')],
+ [dict(type='ShearY')],
+ [dict(type='TranslateX')],
+ [dict(type='TranslateY')],
+]
+
+scale = [(1333, 400), (1333, 1200)]
+
+branch_field = ['sup', 'unsup_teacher', 'unsup_student']
+# pipeline used to augment labeled data,
+# which will be sent to student model for supervised training.
+sup_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomResize', scale=scale, keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='RandAugment', aug_space=color_space, aug_num=1),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='MultiBranch',
+ branch_field=branch_field,
+ sup=dict(type='PackDetInputs'))
+]
+
+# pipeline used to augment unlabeled data weakly,
+# which will be sent to teacher model for predicting pseudo instances.
+weak_pipeline = [
+ dict(type='RandomResize', scale=scale, keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction',
+ 'homography_matrix')),
+]
+
+# pipeline used to augment unlabeled data strongly,
+# which will be sent to student model for unsupervised training.
+strong_pipeline = [
+ dict(type='RandomResize', scale=scale, keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomOrder',
+ transforms=[
+ dict(type='RandAugment', aug_space=color_space, aug_num=1),
+ dict(type='RandAugment', aug_space=geometric, aug_num=1),
+ ]),
+ dict(type='RandomErasing', n_patches=(1, 5), ratio=(0, 0.2)),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction',
+ 'homography_matrix')),
+]
+
+# pipeline used to augment unlabeled data into different views
+unsup_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadEmptyAnnotations'),
+ dict(
+ type='MultiBranch',
+ branch_field=branch_field,
+ unsup_teacher=weak_pipeline,
+ unsup_student=strong_pipeline,
+ )
+]
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+batch_size = 5
+num_workers = 5
+# There are two common semi-supervised learning settings on the coco dataset:
+# (1) Divide the train2017 into labeled and unlabeled datasets
+# by a fixed percentage, such as 1%, 2%, 5% and 10%.
+# The format of labeled_ann_file and unlabeled_ann_file are
+# instances_train2017.{fold}@{percent}.json, and
+# instances_train2017.{fold}@{percent}-unlabeled.json
+# `fold` is used for cross-validation, and `percent` represents
+# the proportion of labeled data in the train2017.
+# (2) Choose the train2017 as the labeled dataset
+# and unlabeled2017 as the unlabeled dataset.
+# The labeled_ann_file and unlabeled_ann_file are
+# instances_train2017.json and image_info_unlabeled2017.json
+# We use this configuration by default.
+labeled_dataset = dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=sup_pipeline,
+ backend_args=backend_args)
+
+unlabeled_dataset = dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_unlabeled2017.json',
+ data_prefix=dict(img='unlabeled2017/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=unsup_pipeline,
+ backend_args=backend_args)
+
+train_dataloader = dict(
+ batch_size=batch_size,
+ num_workers=num_workers,
+ persistent_workers=True,
+ sampler=dict(
+ type='GroupMultiSourceSampler',
+ batch_size=batch_size,
+ source_ratio=[1, 4]),
+ dataset=dict(
+ type='ConcatDataset', datasets=[labeled_dataset, unlabeled_dataset]))
+
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric='bbox',
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/v3det.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/v3det.py
new file mode 100644
index 0000000000000000000000000000000000000000..38ccbf864b6248192dfbf4abaf4858b5f93d45e8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/v3det.py
@@ -0,0 +1,69 @@
+# dataset settings
+dataset_type = 'V3DetDataset'
+data_root = 'data/V3Det/'
+
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ # If you don't have a gt annotation, delete the pipeline
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type='ClassBalancedDataset',
+ oversample_thr=1e-3,
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/v3det_2023_v1_train.json',
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=True, min_size=4),
+ pipeline=train_pipeline,
+ backend_args=backend_args)))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/v3det_2023_v1_val.json',
+ data_prefix=dict(img=''),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/v3det_2023_v1_val.json',
+ metric='bbox',
+ format_only=False,
+ backend_args=backend_args,
+ use_mp_eval=True,
+ proposal_nums=[300])
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/voc0712.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/voc0712.py
new file mode 100644
index 0000000000000000000000000000000000000000..47f5e6563b7f47dd6cfec02248d4c8decd32afe4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/voc0712.py
@@ -0,0 +1,92 @@
+# dataset settings
+dataset_type = 'VOCDataset'
+data_root = 'data/VOCdevkit/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically Infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/segmentation/VOCdevkit/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/segmentation/',
+# 'data/': 's3://openmmlab/datasets/segmentation/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='Resize', scale=(1000, 600), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1000, 600), keep_ratio=True),
+ # avoid bboxes being resized
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type='RepeatDataset',
+ times=3,
+ dataset=dict(
+ type='ConcatDataset',
+ # VOCDataset will add different `dataset_type` in dataset.metainfo,
+ # which will get error if using ConcatDataset. Adding
+ # `ignore_keys` can avoid this error.
+ ignore_keys=['dataset_type'],
+ datasets=[
+ dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='VOC2007/ImageSets/Main/trainval.txt',
+ data_prefix=dict(sub_data_root='VOC2007/'),
+ filter_cfg=dict(
+ filter_empty_gt=True, min_size=32, bbox_min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args),
+ dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='VOC2012/ImageSets/Main/trainval.txt',
+ data_prefix=dict(sub_data_root='VOC2012/'),
+ filter_cfg=dict(
+ filter_empty_gt=True, min_size=32, bbox_min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args)
+ ])))
+
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='VOC2007/ImageSets/Main/test.txt',
+ data_prefix=dict(sub_data_root='VOC2007/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+# Pascal VOC2007 uses `11points` as default evaluate mode, while PASCAL
+# VOC2012 defaults to use 'area'.
+val_evaluator = dict(type='VOCMetric', metric='mAP', eval_mode='11points')
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/wider_face.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/wider_face.py
new file mode 100644
index 0000000000000000000000000000000000000000..7042bc46e877ed899969730325143307e15adf64
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/wider_face.py
@@ -0,0 +1,73 @@
+# dataset settings
+dataset_type = 'WIDERFaceDataset'
+data_root = 'data/WIDERFace/'
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/cityscapes/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+img_scale = (640, 640) # VGA resolution
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='Resize', scale=img_scale, keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=img_scale, keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='train.txt',
+ data_prefix=dict(img='WIDER_train'),
+ filter_cfg=dict(filter_empty_gt=True, bbox_min_size=17, min_size=32),
+ pipeline=train_pipeline))
+
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='val.txt',
+ data_prefix=dict(img='WIDER_val'),
+ test_mode=True,
+ pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ # TODO: support WiderFace-Evaluation for easy, medium, hard cases
+ type='VOCMetric',
+ metric='mAP',
+ eval_mode='11points')
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/youtube_vis.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/youtube_vis.py
new file mode 100644
index 0000000000000000000000000000000000000000..ece07cc3879e512082e302c2e3f76108c57a0234
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/datasets/youtube_vis.py
@@ -0,0 +1,66 @@
+dataset_type = 'YouTubeVISDataset'
+data_root = 'data/youtube_vis_2019/'
+dataset_version = data_root[-5:-1] # 2019 or 2021
+
+backend_args = None
+
+# dataset settings
+train_pipeline = [
+ dict(
+ type='UniformRefFrameSample',
+ num_ref_imgs=1,
+ frame_range=100,
+ filter_key_img=True),
+ dict(
+ type='TransformBroadcaster',
+ share_random_params=True,
+ transforms=[
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadTrackAnnotations', with_mask=True),
+ dict(type='Resize', scale=(640, 360), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ ]),
+ dict(type='PackTrackInputs')
+]
+
+test_pipeline = [
+ dict(
+ type='TransformBroadcaster',
+ transforms=[
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(640, 360), keep_ratio=True),
+ dict(type='LoadTrackAnnotations', with_mask=True),
+ ]),
+ dict(type='PackTrackInputs')
+]
+
+# dataloader
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ # sampler=dict(type='TrackImgSampler'), # image-based sampling
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='TrackAspectRatioBatchSampler'),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ dataset_version=dataset_version,
+ ann_file='annotations/youtube_vis_2019_train.json',
+ data_prefix=dict(img_path='train/JPEGImages'),
+ pipeline=train_pipeline))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False, round_up=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ dataset_version=dataset_version,
+ ann_file='annotations/youtube_vis_2019_valid.json',
+ data_prefix=dict(img_path='valid/JPEGImages'),
+ test_mode=True,
+ pipeline=test_pipeline))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/default_runtime.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/default_runtime.py
new file mode 100644
index 0000000000000000000000000000000000000000..870e5614c86d7e1bbdad13d77a0db03a46ce717a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/default_runtime.py
@@ -0,0 +1,24 @@
+default_scope = 'mmdet'
+
+default_hooks = dict(
+ timer=dict(type='IterTimerHook'),
+ logger=dict(type='LoggerHook', interval=50),
+ param_scheduler=dict(type='ParamSchedulerHook'),
+ checkpoint=dict(type='CheckpointHook', interval=1),
+ sampler_seed=dict(type='DistSamplerSeedHook'),
+ visualization=dict(type='DetVisualizationHook'))
+
+env_cfg = dict(
+ cudnn_benchmark=False,
+ mp_cfg=dict(mp_start_method='fork', opencv_num_threads=0),
+ dist_cfg=dict(backend='nccl'),
+)
+
+vis_backends = [dict(type='LocalVisBackend')]
+visualizer = dict(
+ type='DetLocalVisualizer', vis_backends=vis_backends, name='visualizer')
+log_processor = dict(type='LogProcessor', window_size=50, by_epoch=True)
+
+log_level = 'INFO'
+load_from = None
+resume = False
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/cascade-mask-rcnn_r50_fpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/cascade-mask-rcnn_r50_fpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..c5167f7a02e66c80bd8ec8cc7572acb22eaadba5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/cascade-mask-rcnn_r50_fpn.py
@@ -0,0 +1,203 @@
+# model settings
+model = dict(
+ type='CascadeRCNN',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_mask=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5),
+ rpn_head=dict(
+ type='RPNHead',
+ in_channels=256,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ scales=[8],
+ ratios=[0.5, 1.0, 2.0],
+ strides=[4, 8, 16, 32, 64]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0 / 9.0, loss_weight=1.0)),
+ roi_head=dict(
+ type='CascadeRoIHead',
+ num_stages=3,
+ stage_loss_weights=[1, 0.5, 0.25],
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=7, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ bbox_head=[
+ dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0,
+ loss_weight=1.0)),
+ dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.05, 0.05, 0.1, 0.1]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0,
+ loss_weight=1.0)),
+ dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.033, 0.033, 0.067, 0.067]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0))
+ ],
+ mask_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=14, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ mask_head=dict(
+ type='FCNMaskHead',
+ num_convs=4,
+ in_channels=256,
+ conv_out_channels=256,
+ num_classes=80,
+ loss_mask=dict(
+ type='CrossEntropyLoss', use_mask=True, loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ match_low_quality=True,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=0,
+ pos_weight=-1,
+ debug=False),
+ rpn_proposal=dict(
+ nms_pre=2000,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=[
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ mask_size=28,
+ pos_weight=-1,
+ debug=False),
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.6,
+ neg_iou_thr=0.6,
+ min_pos_iou=0.6,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ mask_size=28,
+ pos_weight=-1,
+ debug=False),
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.7,
+ min_pos_iou=0.7,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ mask_size=28,
+ pos_weight=-1,
+ debug=False)
+ ]),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=1000,
+ max_per_img=1000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.5),
+ max_per_img=100,
+ mask_thr_binary=0.5)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/cascade-rcnn_r50_fpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/cascade-rcnn_r50_fpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..50c57f01ca3a6ea827f71801b0c233af268914f9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/cascade-rcnn_r50_fpn.py
@@ -0,0 +1,185 @@
+# model settings
+model = dict(
+ type='CascadeRCNN',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5),
+ rpn_head=dict(
+ type='RPNHead',
+ in_channels=256,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ scales=[8],
+ ratios=[0.5, 1.0, 2.0],
+ strides=[4, 8, 16, 32, 64]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0 / 9.0, loss_weight=1.0)),
+ roi_head=dict(
+ type='CascadeRoIHead',
+ num_stages=3,
+ stage_loss_weights=[1, 0.5, 0.25],
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=7, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ bbox_head=[
+ dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0,
+ loss_weight=1.0)),
+ dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.05, 0.05, 0.1, 0.1]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0,
+ loss_weight=1.0)),
+ dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.033, 0.033, 0.067, 0.067]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0))
+ ]),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ match_low_quality=True,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=0,
+ pos_weight=-1,
+ debug=False),
+ rpn_proposal=dict(
+ nms_pre=2000,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=[
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ pos_weight=-1,
+ debug=False),
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.6,
+ neg_iou_thr=0.6,
+ min_pos_iou=0.6,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ pos_weight=-1,
+ debug=False),
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.7,
+ min_pos_iou=0.7,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ pos_weight=-1,
+ debug=False)
+ ]),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=1000,
+ max_per_img=1000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.5),
+ max_per_img=100)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/fast-rcnn_r50_fpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/fast-rcnn_r50_fpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..2bd45e9266b01df302b78e50258fa1572144cb21
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/fast-rcnn_r50_fpn.py
@@ -0,0 +1,68 @@
+# model settings
+model = dict(
+ type='FastRCNN',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5),
+ roi_head=dict(
+ type='StandardRoIHead',
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=7, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ bbox_head=dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=False,
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rcnn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ pos_weight=-1,
+ debug=False)),
+ test_cfg=dict(
+ rcnn=dict(
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.5),
+ max_per_img=100)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/faster-rcnn_r50-caffe-c4.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/faster-rcnn_r50-caffe-c4.py
new file mode 100644
index 0000000000000000000000000000000000000000..15d2db72e48790505c2a1e4e7d184c1803f7ab31
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/faster-rcnn_r50-caffe-c4.py
@@ -0,0 +1,123 @@
+# model settings
+norm_cfg = dict(type='BN', requires_grad=False)
+model = dict(
+ type='FasterRCNN',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=3,
+ strides=(1, 2, 2),
+ dilations=(1, 1, 1),
+ out_indices=(2, ),
+ frozen_stages=1,
+ norm_cfg=norm_cfg,
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')),
+ rpn_head=dict(
+ type='RPNHead',
+ in_channels=1024,
+ feat_channels=1024,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ scales=[2, 4, 8, 16, 32],
+ ratios=[0.5, 1.0, 2.0],
+ strides=[16]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0)),
+ roi_head=dict(
+ type='StandardRoIHead',
+ shared_head=dict(
+ type='ResLayer',
+ depth=50,
+ stage=3,
+ stride=2,
+ dilation=1,
+ style='caffe',
+ norm_cfg=norm_cfg,
+ norm_eval=True,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')),
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=14, sampling_ratio=0),
+ out_channels=1024,
+ featmap_strides=[16]),
+ bbox_head=dict(
+ type='BBoxHead',
+ with_avg_pool=True,
+ roi_feat_size=7,
+ in_channels=2048,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=False,
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ match_low_quality=True,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ rpn_proposal=dict(
+ nms_pre=12000,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ pos_weight=-1,
+ debug=False)),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=6000,
+ max_per_img=1000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.5),
+ max_per_img=100)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/faster-rcnn_r50-caffe-dc5.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/faster-rcnn_r50-caffe-dc5.py
new file mode 100644
index 0000000000000000000000000000000000000000..189915e3d9ce7239493da6465931f91e2d9d664f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/faster-rcnn_r50-caffe-dc5.py
@@ -0,0 +1,111 @@
+# model settings
+norm_cfg = dict(type='BN', requires_grad=False)
+model = dict(
+ type='FasterRCNN',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ strides=(1, 2, 2, 1),
+ dilations=(1, 1, 1, 2),
+ out_indices=(3, ),
+ frozen_stages=1,
+ norm_cfg=norm_cfg,
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')),
+ rpn_head=dict(
+ type='RPNHead',
+ in_channels=2048,
+ feat_channels=2048,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ scales=[2, 4, 8, 16, 32],
+ ratios=[0.5, 1.0, 2.0],
+ strides=[16]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0)),
+ roi_head=dict(
+ type='StandardRoIHead',
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=7, sampling_ratio=0),
+ out_channels=2048,
+ featmap_strides=[16]),
+ bbox_head=dict(
+ type='Shared2FCBBoxHead',
+ in_channels=2048,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=False,
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ match_low_quality=True,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=0,
+ pos_weight=-1,
+ debug=False),
+ rpn_proposal=dict(
+ nms_pre=12000,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ pos_weight=-1,
+ debug=False)),
+ test_cfg=dict(
+ rpn=dict(
+ nms=dict(type='nms', iou_threshold=0.7),
+ nms_pre=6000,
+ max_per_img=1000,
+ min_bbox_size=0),
+ rcnn=dict(
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.5),
+ max_per_img=100)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/faster-rcnn_r50_fpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/faster-rcnn_r50_fpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..31aa1461799a988a11adb901306a063fd3f0b951
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/faster-rcnn_r50_fpn.py
@@ -0,0 +1,114 @@
+# model settings
+model = dict(
+ type='FasterRCNN',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5),
+ rpn_head=dict(
+ type='RPNHead',
+ in_channels=256,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ scales=[8],
+ ratios=[0.5, 1.0, 2.0],
+ strides=[4, 8, 16, 32, 64]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0)),
+ roi_head=dict(
+ type='StandardRoIHead',
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=7, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ bbox_head=dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=False,
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ match_low_quality=True,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ rpn_proposal=dict(
+ nms_pre=2000,
+ max_per_img=1000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ pos_weight=-1,
+ debug=False)),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=1000,
+ max_per_img=1000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.5),
+ max_per_img=100)
+ # soft-nms is also supported for rcnn testing
+ # e.g., nms=dict(type='soft_nms', iou_threshold=0.5, min_score=0.05)
+ ))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/mask-rcnn_r50-caffe-c4.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/mask-rcnn_r50-caffe-c4.py
new file mode 100644
index 0000000000000000000000000000000000000000..de1131b24893ae24bd99923895fd844837c9b46d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/mask-rcnn_r50-caffe-c4.py
@@ -0,0 +1,132 @@
+# model settings
+norm_cfg = dict(type='BN', requires_grad=False)
+model = dict(
+ type='MaskRCNN',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_mask=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=3,
+ strides=(1, 2, 2),
+ dilations=(1, 1, 1),
+ out_indices=(2, ),
+ frozen_stages=1,
+ norm_cfg=norm_cfg,
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')),
+ rpn_head=dict(
+ type='RPNHead',
+ in_channels=1024,
+ feat_channels=1024,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ scales=[2, 4, 8, 16, 32],
+ ratios=[0.5, 1.0, 2.0],
+ strides=[16]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0)),
+ roi_head=dict(
+ type='StandardRoIHead',
+ shared_head=dict(
+ type='ResLayer',
+ depth=50,
+ stage=3,
+ stride=2,
+ dilation=1,
+ style='caffe',
+ norm_cfg=norm_cfg,
+ norm_eval=True),
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=14, sampling_ratio=0),
+ out_channels=1024,
+ featmap_strides=[16]),
+ bbox_head=dict(
+ type='BBoxHead',
+ with_avg_pool=True,
+ roi_feat_size=7,
+ in_channels=2048,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=False,
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0)),
+ mask_roi_extractor=None,
+ mask_head=dict(
+ type='FCNMaskHead',
+ num_convs=0,
+ in_channels=2048,
+ conv_out_channels=256,
+ num_classes=80,
+ loss_mask=dict(
+ type='CrossEntropyLoss', use_mask=True, loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ match_low_quality=True,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=0,
+ pos_weight=-1,
+ debug=False),
+ rpn_proposal=dict(
+ nms_pre=12000,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ mask_size=14,
+ pos_weight=-1,
+ debug=False)),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=6000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ max_per_img=1000,
+ min_bbox_size=0),
+ rcnn=dict(
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.5),
+ max_per_img=100,
+ mask_thr_binary=0.5)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/mask-rcnn_r50_fpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/mask-rcnn_r50_fpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..b4ff7a49d0a2f0abd4823ef89ad957d9708085e7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/mask-rcnn_r50_fpn.py
@@ -0,0 +1,127 @@
+# model settings
+model = dict(
+ type='MaskRCNN',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_mask=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5),
+ rpn_head=dict(
+ type='RPNHead',
+ in_channels=256,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ scales=[8],
+ ratios=[0.5, 1.0, 2.0],
+ strides=[4, 8, 16, 32, 64]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0)),
+ roi_head=dict(
+ type='StandardRoIHead',
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=7, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ bbox_head=dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=False,
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0)),
+ mask_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=14, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ mask_head=dict(
+ type='FCNMaskHead',
+ num_convs=4,
+ in_channels=256,
+ conv_out_channels=256,
+ num_classes=80,
+ loss_mask=dict(
+ type='CrossEntropyLoss', use_mask=True, loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ match_low_quality=True,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ rpn_proposal=dict(
+ nms_pre=2000,
+ max_per_img=1000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ match_low_quality=True,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ mask_size=28,
+ pos_weight=-1,
+ debug=False)),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=1000,
+ max_per_img=1000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.5),
+ max_per_img=100,
+ mask_thr_binary=0.5)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/retinanet_r50_fpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/retinanet_r50_fpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..53662c9f1390af22b15c5591e122b0aa0b2d6c92
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/retinanet_r50_fpn.py
@@ -0,0 +1,68 @@
+# model settings
+model = dict(
+ type='RetinaNet',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_input',
+ num_outs=5),
+ bbox_head=dict(
+ type='RetinaHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ octave_base_scale=4,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[8, 16, 32, 64, 128]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0)),
+ # model training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.4,
+ min_pos_iou=0,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='PseudoSampler'), # Focal loss should use PseudoSampler
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.5),
+ max_per_img=100))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/rpn_r50-caffe-c4.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/rpn_r50-caffe-c4.py
new file mode 100644
index 0000000000000000000000000000000000000000..ed1dbe746d432d96d70e7dc9048c9e1b1727c938
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/rpn_r50-caffe-c4.py
@@ -0,0 +1,64 @@
+# model settings
+model = dict(
+ type='RPN',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=3,
+ strides=(1, 2, 2),
+ dilations=(1, 1, 1),
+ out_indices=(2, ),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')),
+ neck=None,
+ rpn_head=dict(
+ type='RPNHead',
+ in_channels=1024,
+ feat_channels=1024,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ scales=[2, 4, 8, 16, 32],
+ ratios=[0.5, 1.0, 2.0],
+ strides=[16]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0)),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False)),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=12000,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/rpn_r50_fpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/rpn_r50_fpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..6bc4790434a368d0728d74dcd7ba79e665aae276
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/rpn_r50_fpn.py
@@ -0,0 +1,64 @@
+# model settings
+model = dict(
+ type='RPN',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5),
+ rpn_head=dict(
+ type='RPNHead',
+ in_channels=256,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ scales=[8],
+ ratios=[0.5, 1.0, 2.0],
+ strides=[4, 8, 16, 32, 64]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0)),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False)),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=2000,
+ max_per_img=1000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/ssd300.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/ssd300.py
new file mode 100644
index 0000000000000000000000000000000000000000..fd113c7cbc41494eabb6a56061f8a90343ac9efd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/models/ssd300.py
@@ -0,0 +1,63 @@
+# model settings
+input_size = 300
+model = dict(
+ type='SingleStageDetector',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[1, 1, 1],
+ bgr_to_rgb=True,
+ pad_size_divisor=1),
+ backbone=dict(
+ type='SSDVGG',
+ depth=16,
+ with_last_pool=False,
+ ceil_mode=True,
+ out_indices=(3, 4),
+ out_feature_indices=(22, 34),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://vgg16_caffe')),
+ neck=dict(
+ type='SSDNeck',
+ in_channels=(512, 1024),
+ out_channels=(512, 1024, 512, 256, 256, 256),
+ level_strides=(2, 2, 1, 1),
+ level_paddings=(1, 1, 0, 0),
+ l2_norm_scale=20),
+ bbox_head=dict(
+ type='SSDHead',
+ in_channels=(512, 1024, 512, 256, 256, 256),
+ num_classes=80,
+ anchor_generator=dict(
+ type='SSDAnchorGenerator',
+ scale_major=False,
+ input_size=input_size,
+ basesize_ratio_range=(0.15, 0.9),
+ strides=[8, 16, 32, 64, 100, 300],
+ ratios=[[2], [2, 3], [2, 3], [2, 3], [2], [2]]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2])),
+ # model training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.,
+ ignore_iof_thr=-1,
+ gt_max_assign_all=False),
+ sampler=dict(type='PseudoSampler'),
+ smoothl1_beta=1.,
+ allowed_border=-1,
+ pos_weight=-1,
+ neg_pos_ratio=3,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ nms=dict(type='nms', iou_threshold=0.45),
+ min_bbox_size=0,
+ score_thr=0.02,
+ max_per_img=200))
+cudnn_benchmark = True
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/schedules/schedule_1x.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/schedules/schedule_1x.py
new file mode 100644
index 0000000000000000000000000000000000000000..95f30be74ff37080ba0d227d55bbd587feeaa892
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/schedules/schedule_1x.py
@@ -0,0 +1,28 @@
+# training schedule for 1x
+train_cfg = dict(type='EpochBasedTrainLoop', max_epochs=12, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.0001))
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/schedules/schedule_20e.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/schedules/schedule_20e.py
new file mode 100644
index 0000000000000000000000000000000000000000..75f958b0ed11d77ae3aebff6b7a5d8cb80797d9f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/schedules/schedule_20e.py
@@ -0,0 +1,28 @@
+# training schedule for 20e
+train_cfg = dict(type='EpochBasedTrainLoop', max_epochs=20, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=20,
+ by_epoch=True,
+ milestones=[16, 19],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.0001))
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/schedules/schedule_2x.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/schedules/schedule_2x.py
new file mode 100644
index 0000000000000000000000000000000000000000..5b7b241de6f3285e0f127f3c0581c8c84de463e4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_base_/schedules/schedule_2x.py
@@ -0,0 +1,28 @@
+# training schedule for 2x
+train_cfg = dict(type='EpochBasedTrainLoop', max_epochs=24, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=24,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.0001))
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_custom_/faster_rcnn_r50_fpn_1x-dior.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_custom_/faster_rcnn_r50_fpn_1x-dior.py
new file mode 100644
index 0000000000000000000000000000000000000000..8a6219ec6aee71e4f6c44ad23efbff6d92367cbb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_custom_/faster_rcnn_r50_fpn_1x-dior.py
@@ -0,0 +1,71 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='Resize', scale=(512, 512), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size_divisor=32),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(512, 512), keep_ratio=True),
+ dict(type='Pad', size_divisor=32),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+model = dict(
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+
+ roi_head=dict(
+ bbox_head=dict(num_classes=20)))
+
+data_root = '/path/to/DIOR-VOC/'
+metainfo = {
+ 'classes': ('vehicle', 'baseballfield', 'groundtrackfield', 'windmill', 'bridge', 'overpass', 'ship', 'airplane', 'tenniscourt', 'airport',
+ 'expressway-service-area', 'basketballcourt', 'stadium', 'storagetank', 'chimney', 'dam', 'expressway-toll-station', 'golffield', 'trainstation', 'harbor'),
+}
+
+train_dataloader = dict(
+ batch_size=2,
+ dataset=dict(
+ data_root=data_root,
+ metainfo=metainfo,
+ ann_file='',
+ data_prefix=dict(img='VOC2007/JPEGImages/'),
+ pipeline=train_pipeline))
+
+val_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ metainfo=metainfo,
+ ann_file='',
+ data_prefix=dict(img='VOC2007/JPEGImages/'),
+ pipeline=test_pipeline))
+
+test_dataloader = val_dataloader
+
+val_evaluator = dict(ann_file='')
+test_evaluator = val_evaluator
+
+# visualizer = dict(
+# type='DetLocalVisualizer',
+# vis_backends=[dict(type='LocalVisBackend')],
+# name='visualizer',
+# text_scale=1.5,
+# line_width=3,
+# alpha=0.8
+# )
\ No newline at end of file
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_custom_/faster_rcnn_r50_fpn_1x-exdark.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_custom_/faster_rcnn_r50_fpn_1x-exdark.py
new file mode 100644
index 0000000000000000000000000000000000000000..48ae0dfb9d219e05caf35b53439816189000b16b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_custom_/faster_rcnn_r50_fpn_1x-exdark.py
@@ -0,0 +1,65 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='Resize', scale=(512, 512), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size_divisor=32),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(512, 512), keep_ratio=True),
+ dict(type='Pad', size_divisor=32),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+model = dict(
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+
+ roi_head=dict(
+ bbox_head=dict(num_classes=12)))
+
+data_root = '/path/to/ExDark/'
+metainfo = {
+ 'classes': ('bicycle', 'boat', 'bottle', 'bus', 'car', 'cat', 'chair', 'cup', 'dog', 'motorbike', 'people', 'table'),
+}
+
+train_dataloader = dict(
+ batch_size=2,
+ dataset=dict(
+ data_root=data_root,
+ metainfo=metainfo,
+ ann_file='',
+ data_prefix=dict(img='train'),
+ pipeline=train_pipeline))
+
+val_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ metainfo=metainfo,
+ ann_file='',
+ data_prefix=dict(img='test'),
+ pipeline=test_pipeline))
+
+test_dataloader = val_dataloader
+
+val_evaluator = dict(ann_file='')
+test_evaluator = val_evaluator
+
+val_cfg = None
+val_dataloader = None
+val_evaluator = None
\ No newline at end of file
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/_custom_/faster_rcnn_r50_fpn_1x-ruod.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_custom_/faster_rcnn_r50_fpn_1x-ruod.py
new file mode 100644
index 0000000000000000000000000000000000000000..64980a5ef6cef357a4b5bcbb7b2dfb4afd03bc6a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/_custom_/faster_rcnn_r50_fpn_1x-ruod.py
@@ -0,0 +1,61 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='Resize', scale=(512, 512), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size_divisor=32),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(512, 512), keep_ratio=True),
+ dict(type='Pad', size_divisor=32),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+model = dict(
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+
+ roi_head=dict(
+ bbox_head=dict(num_classes=10)))
+
+data_root = '/path/to/RUOD/'
+metainfo = {
+ 'classes': ('holothurian', 'echinus', 'scallop', 'starfish', 'fish', 'corals', 'diver', 'cuttlefish', 'turtle', 'jellyfish'),
+}
+
+train_dataloader = dict(
+ batch_size=2,
+ dataset=dict(
+ data_root=data_root,
+ metainfo=metainfo,
+ ann_file='RUOD_ANN/instances_train.json',
+ data_prefix=dict(img='RUOD_pic/train/'),
+ pipeline=train_pipeline))
+
+val_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ metainfo=metainfo,
+ ann_file='RUOD_ANN/instances_test.json',
+ data_prefix=dict(img='RUOD_pic/test/'),
+ pipeline=test_pipeline))
+
+test_dataloader = val_dataloader
+
+val_evaluator = dict(ann_file=data_root + 'RUOD_ANN/instances_test.json')
+test_evaluator = val_evaluator
\ No newline at end of file
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/albu_example/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/albu_example/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..fa362f95fb91ba4beed5c9d6814e087324bd74d5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/albu_example/README.md
@@ -0,0 +1,31 @@
+# Albu Example
+
+> [Albumentations: fast and flexible image augmentations](https://arxiv.org/abs/1809.06839)
+
+
+
+## Abstract
+
+Data augmentation is a commonly used technique for increasing both the size and the diversity of labeled training sets by leveraging input transformations that preserve output labels. In computer vision domain, image augmentations have become a common implicit regularization technique to combat overfitting in deep convolutional neural networks and are ubiquitously used to improve performance. While most deep learning frameworks implement basic image transformations, the list is typically limited to some variations and combinations of flipping, rotating, scaling, and cropping. Moreover, the image processing speed varies in existing tools for image augmentation. We present Albumentations, a fast and flexible library for image augmentations with many various image transform operations available, that is also an easy-to-use wrapper around other augmentation libraries. We provide examples of image augmentations for different computer vision tasks and show that Albumentations is faster than other commonly used image augmentation tools on the most of commonly used image transformations.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :-------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | pytorch | 1x | 4.4 | 16.6 | 38.0 | 34.5 | [config](./mask-rcnn_r50_fpn_albu-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/albu_example/mask_rcnn_r50_fpn_albu_1x_coco/mask_rcnn_r50_fpn_albu_1x_coco_20200208-ab203bcd.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/albu_example/mask_rcnn_r50_fpn_albu_1x_coco/mask_rcnn_r50_fpn_albu_1x_coco_20200208_225520.log.json) |
+
+## Citation
+
+```latex
+@article{2018arXiv180906839B,
+ author = {A. Buslaev, A. Parinov, E. Khvedchenya, V.~I. Iglovikov and A.~A. Kalinin},
+ title = "{Albumentations: fast and flexible image augmentations}",
+ journal = {ArXiv e-prints},
+ eprint = {1809.06839},
+ year = 2018
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/albu_example/mask-rcnn_r50_fpn_albu-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/albu_example/mask-rcnn_r50_fpn_albu-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b8a2780e99b88c78adbe74c024fcd2d693817030
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/albu_example/mask-rcnn_r50_fpn_albu-1x_coco.py
@@ -0,0 +1,66 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+
+albu_train_transforms = [
+ dict(
+ type='ShiftScaleRotate',
+ shift_limit=0.0625,
+ scale_limit=0.0,
+ rotate_limit=0,
+ interpolation=1,
+ p=0.5),
+ dict(
+ type='RandomBrightnessContrast',
+ brightness_limit=[0.1, 0.3],
+ contrast_limit=[0.1, 0.3],
+ p=0.2),
+ dict(
+ type='OneOf',
+ transforms=[
+ dict(
+ type='RGBShift',
+ r_shift_limit=10,
+ g_shift_limit=10,
+ b_shift_limit=10,
+ p=1.0),
+ dict(
+ type='HueSaturationValue',
+ hue_shift_limit=20,
+ sat_shift_limit=30,
+ val_shift_limit=20,
+ p=1.0)
+ ],
+ p=0.1),
+ dict(type='JpegCompression', quality_lower=85, quality_upper=95, p=0.2),
+ dict(type='ChannelShuffle', p=0.1),
+ dict(
+ type='OneOf',
+ transforms=[
+ dict(type='Blur', blur_limit=3, p=1.0),
+ dict(type='MedianBlur', blur_limit=3, p=1.0)
+ ],
+ p=0.1),
+]
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(
+ type='Albu',
+ transforms=albu_train_transforms,
+ bbox_params=dict(
+ type='BboxParams',
+ format='pascal_voc',
+ label_fields=['gt_bboxes_labels', 'gt_ignore_flags'],
+ min_visibility=0.0,
+ filter_lost_elements=True),
+ keymap={
+ 'img': 'image',
+ 'gt_masks': 'masks',
+ 'gt_bboxes': 'bboxes'
+ },
+ skip_img_without_anno=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/albu_example/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/albu_example/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..3b54bdf15688281e5896faac3f841433497c7eaf
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/albu_example/metafile.yml
@@ -0,0 +1,17 @@
+Models:
+ - Name: mask-rcnn_r50_fpn_albu-1x_coco
+ In Collection: Mask R-CNN
+ Config: mask-rcnn_r50_fpn_albu-1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.4
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 34.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/albu_example/mask_rcnn_r50_fpn_albu_1x_coco/mask_rcnn_r50_fpn_albu_1x_coco_20200208-ab203bcd.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..1411672e205683914c24bec47ac02517d44f684b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/README.md
@@ -0,0 +1,31 @@
+# ATSS
+
+> [Bridging the Gap Between Anchor-based and Anchor-free Detection via Adaptive Training Sample Selection](https://arxiv.org/abs/1912.02424)
+
+
+
+## Abstract
+
+Object detection has been dominated by anchor-based detectors for several years. Recently, anchor-free detectors have become popular due to the proposal of FPN and Focal Loss. In this paper, we first point out that the essential difference between anchor-based and anchor-free detection is actually how to define positive and negative training samples, which leads to the performance gap between them. If they adopt the same definition of positive and negative samples during training, there is no obvious difference in the final performance, no matter regressing from a box or a point. This shows that how to select positive and negative training samples is important for current object detectors. Then, we propose an Adaptive Training Sample Selection (ATSS) to automatically select positive and negative samples according to statistical characteristics of object. It significantly improves the performance of anchor-based and anchor-free detectors and bridges the gap between them. Finally, we discuss the necessity of tiling multiple anchors per location on the image to detect objects. Extensive experiments conducted on MS COCO support our aforementioned analysis and conclusions. With the newly introduced ATSS, we improve state-of-the-art detectors by a large margin to 50.7% AP without introducing any overhead.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :------: | :-----: | :-----: | :------: | :------------: | :----: | :----------------------------------: | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | pytorch | 1x | 3.7 | 19.7 | 39.4 | [config](./atss_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/atss/atss_r50_fpn_1x_coco/atss_r50_fpn_1x_coco_20200209-985f7bd0.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/atss/atss_r50_fpn_1x_coco/atss_r50_fpn_1x_coco_20200209_102539.log.json) |
+| R-101 | pytorch | 1x | 5.6 | 12.3 | 41.5 | [config](./atss_r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/atss/atss_r101_fpn_1x_coco/atss_r101_fpn_1x_20200825-dfcadd6f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/atss/atss_r101_fpn_1x_coco/atss_r101_fpn_1x_20200825-dfcadd6f.log.json) |
+
+## Citation
+
+```latex
+@article{zhang2019bridging,
+ title = {Bridging the Gap Between Anchor-based and Anchor-free Detection via Adaptive Training Sample Selection},
+ author = {Zhang, Shifeng and Chi, Cheng and Yao, Yongqiang and Lei, Zhen and Li, Stan Z.},
+ journal = {arXiv preprint arXiv:1912.02424},
+ year = {2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/atss_r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/atss_r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5225d2ab672738d4d427eba252e92bd554252476
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/atss_r101_fpn_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './atss_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/atss_r101_fpn_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/atss_r101_fpn_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..69999ce45aee9c76dcc4af974e6e9baabbd5b44b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/atss_r101_fpn_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './atss_r50_fpn_8xb8-amp-lsj-200e_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/atss_r18_fpn_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/atss_r18_fpn_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..12d9f13263619333391befd6692c83622091ef4e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/atss_r18_fpn_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './atss_r50_fpn_8xb8-amp-lsj-200e_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=18,
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet18')),
+ neck=dict(in_channels=[64, 128, 256, 512]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/atss_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/atss_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..306435d7d2fc645f1c2deae784c1875cc4ceaf98
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/atss_r50_fpn_1x_coco.py
@@ -0,0 +1,71 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+# model settings
+model = dict(
+ type='ATSS',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output',
+ num_outs=5),
+ bbox_head=dict(
+ type='ATSSHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ octave_base_scale=8,
+ scales_per_octave=1,
+ strides=[8, 16, 32, 64, 128]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=2.0),
+ loss_centerness=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(type='ATSSAssigner', topk=9),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/atss_r50_fpn_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/atss_r50_fpn_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e3b3c46f4b926b82fbab438d6d50eb6c079dabc3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/atss_r50_fpn_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,81 @@
+_base_ = '../common/lsj-200e_coco-detection.py'
+
+image_size = (1024, 1024)
+batch_augments = [dict(type='BatchFixedSizePad', size=image_size)]
+
+model = dict(
+ type='ATSS',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32,
+ batch_augments=batch_augments),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output',
+ num_outs=5),
+ bbox_head=dict(
+ type='ATSSHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ octave_base_scale=8,
+ scales_per_octave=1,
+ strides=[8, 16, 32, 64, 128]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=2.0),
+ loss_centerness=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(type='ATSSAssigner', topk=9),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+
+train_dataloader = dict(batch_size=8, num_workers=4)
+
+# Enable automatic-mixed-precision training with AmpOptimWrapper.
+optim_wrapper = dict(
+ type='AmpOptimWrapper',
+ optimizer=dict(
+ type='SGD', lr=0.01 * 4, momentum=0.9, weight_decay=0.00004))
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..f4c567ef29ba9ea4fddd7bc00d63a4bca41b1cfa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/atss/metafile.yml
@@ -0,0 +1,60 @@
+Collections:
+ - Name: ATSS
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ATSS
+ - FPN
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/1912.02424
+ Title: 'Bridging the Gap Between Anchor-based and Anchor-free Detection via Adaptive Training Sample Selection'
+ README: configs/atss/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/detectors/atss.py#L6
+ Version: v2.0.0
+
+Models:
+ - Name: atss_r50_fpn_1x_coco
+ In Collection: ATSS
+ Config: configs/atss/atss_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.7
+ inference time (ms/im):
+ - value: 50.76
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/atss/atss_r50_fpn_1x_coco/atss_r50_fpn_1x_coco_20200209-985f7bd0.pth
+
+ - Name: atss_r101_fpn_1x_coco
+ In Collection: ATSS
+ Config: configs/atss/atss_r101_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.6
+ inference time (ms/im):
+ - value: 81.3
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/atss/atss_r101_fpn_1x_coco/atss_r101_fpn_1x_20200825-dfcadd6f.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/autoassign/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/autoassign/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..f6b05738ccc222223863974697cee7f2770d8f25
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/autoassign/README.md
@@ -0,0 +1,35 @@
+# AutoAssign
+
+> [AutoAssign: Differentiable Label Assignment for Dense Object Detection](https://arxiv.org/abs/2007.03496)
+
+
+
+## Abstract
+
+Determining positive/negative samples for object detection is known as label assignment. Here we present an anchor-free detector named AutoAssign. It requires little human knowledge and achieves appearance-aware through a fully differentiable weighting mechanism. During training, to both satisfy the prior distribution of data and adapt to category characteristics, we present Center Weighting to adjust the category-specific prior distributions. To adapt to object appearances, Confidence Weighting is proposed to adjust the specific assign strategy of each instance. The two weighting modules are then combined to generate positive and negative weights to adjust each location's confidence. Extensive experiments on the MS COCO show that our method steadily surpasses other best sampling strategies by large margins with various backbones. Moreover, our best model achieves 52.1% AP, outperforming all existing one-stage detectors. Besides, experiments on other datasets, e.g., PASCAL VOC, Objects365, and WiderFace, demonstrate the broad applicability of AutoAssign.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Style | Lr schd | Mem (GB) | box AP | Config | Download |
+| :------: | :---: | :-----: | :------: | :----: | :---------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | caffe | 1x | 4.08 | 40.4 | [config](./autoassign_r50-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/autoassign/auto_assign_r50_fpn_1x_coco/auto_assign_r50_fpn_1x_coco_20210413_115540-5e17991f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/autoassign/auto_assign_r50_fpn_1x_coco/auto_assign_r50_fpn_1x_coco_20210413_115540-5e17991f.log.json) |
+
+**Note**:
+
+1. We find that the performance is unstable with 1x setting and may fluctuate by about 0.3 mAP. mAP 40.3 ~ 40.6 is acceptable. Such fluctuation can also be found in the original implementation.
+2. You can get a more stable results ~ mAP 40.6 with a schedule total 13 epoch, and learning rate is divided by 10 at 10th and 13th epoch.
+
+## Citation
+
+```latex
+@article{zhu2020autoassign,
+ title={AutoAssign: Differentiable Label Assignment for Dense Object Detection},
+ author={Zhu, Benjin and Wang, Jianfeng and Jiang, Zhengkai and Zong, Fuhang and Liu, Songtao and Li, Zeming and Sun, Jian},
+ journal={arXiv preprint arXiv:2007.03496},
+ year={2020}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/autoassign/autoassign_r50-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/autoassign/autoassign_r50-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..76a361952d95b655451186ef1cb39df2f24ae305
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/autoassign/autoassign_r50-caffe_fpn_1x_coco.py
@@ -0,0 +1,69 @@
+# We follow the original implementation which
+# adopts the Caffe pre-trained backbone.
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+# model settings
+model = dict(
+ type='AutoAssign',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[102.9801, 115.9465, 122.7717],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs=True,
+ num_outs=5,
+ relu_before_extra_convs=True,
+ init_cfg=dict(type='Caffe2Xavier', layer='Conv2d')),
+ bbox_head=dict(
+ type='AutoAssignHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ strides=[8, 16, 32, 64, 128],
+ loss_bbox=dict(type='GIoULoss', loss_weight=5.0)),
+ train_cfg=None,
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0,
+ end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(lr=0.01), paramwise_cfg=dict(norm_decay_mult=0.))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/autoassign/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/autoassign/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..ab7a4af3371d4be5325498db97af0e7dd8fdc28c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/autoassign/metafile.yml
@@ -0,0 +1,33 @@
+Collections:
+ - Name: AutoAssign
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - AutoAssign
+ - FPN
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/2007.03496
+ Title: 'AutoAssign: Differentiable Label Assignment for Dense Object Detection'
+ README: configs/autoassign/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.12.0/mmdet/models/detectors/autoassign.py#L6
+ Version: v2.12.0
+
+Models:
+ - Name: autoassign_r50-caffe_fpn_1x_coco
+ In Collection: AutoAssign
+ Config: configs/autoassign/autoassign_r50-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.08
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/autoassign/auto_assign_r50_fpn_1x_coco/auto_assign_r50_fpn_1x_coco_20210413_115540-5e17991f.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/boxinst/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/boxinst/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..f6f01c5d27b13b1758a5c9a60251383852e0f48e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/boxinst/README.md
@@ -0,0 +1,32 @@
+# BoxInst
+
+> [BoxInst: High-Performance Instance Segmentation with Box Annotations](https://arxiv.org/pdf/2012.02310.pdf)
+
+
+
+## Abstract
+
+We present a high-performance method that can achieve mask-level instance segmentation with only bounding-box annotations for training. While this setting has been studied in the literature, here we show significantly stronger performance with a simple design (e.g., dramatically improving previous best reported mask AP of 21.1% to 31.6% on the COCO dataset). Our core idea is to redesign the loss
+of learning masks in instance segmentation, with no modification to the segmentation network itself. The new loss functions can supervise the mask training without relying on mask annotations. This is made possible with two loss terms, namely, 1) a surrogate term that minimizes the discrepancy between the projections of the ground-truth box and the predicted mask; 2) a pairwise loss that can exploit the prior that proximal pixels with similar colors are very likely to have the same category label. Experiments demonstrate that the redesigned mask loss can yield surprisingly high-quality instance masks with only box annotations. For example, without using any mask annotations, with a ResNet-101 backbone and 3× training schedule, we achieve 33.2% mask AP on COCO test-dev split (vs. 39.1% of the fully supervised counterpart). Our excellent experiment results on COCO and Pascal VOC indicate that our method dramatically narrows the performance gap between weakly and fully supervised instance segmentation.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Style | MS train | Lr schd | bbox AP | mask AP | Config | Download |
+| :------: | :-----: | :------: | :-----: | :-----: | :-----: | :-----------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | pytorch | Y | 1x | 39.6 | 31.1 | [config](./boxinst_r50_fpn_ms-90k_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/boxinst/boxinst_r50_fpn_ms-90k_coco/boxinst_r50_fpn_ms-90k_coco_20221228_163052-6add751a.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/boxinst/boxinst_r50_fpn_ms-90k_coco/boxinst_r50_fpn_ms-90k_coco_20221228_163052.log.json) |
+| R-101 | pytorch | Y | 1x | 41.8 | 32.7 | [config](./boxinst_r101_fpn_ms-90k_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/boxinst/boxinst_r101_fpn_ms-90k_coco/boxinst_r101_fpn_ms-90k_coco_20221229_145106-facf375b.pth) \|[log](https://download.openmmlab.com/mmdetection/v3.0/boxinst/boxinst_r101_fpn_ms-90k_coco/boxinst_r101_fpn_ms-90k_coco_20221229_145106.log.json) |
+
+## Citation
+
+```latex
+@inproceedings{tian2020boxinst,
+ title = {{BoxInst}: High-Performance Instance Segmentation with Box Annotations},
+ author = {Tian, Zhi and Shen, Chunhua and Wang, Xinlong and Chen, Hao},
+ booktitle = {Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR)},
+ year = {2021}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/boxinst/boxinst_r101_fpn_ms-90k_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/boxinst/boxinst_r101_fpn_ms-90k_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ab2b11628a79aee7f6f6403cecf8f7b1d0526d69
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/boxinst/boxinst_r101_fpn_ms-90k_coco.py
@@ -0,0 +1,8 @@
+_base_ = './boxinst_r50_fpn_ms-90k_coco.py'
+
+# model settings
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/boxinst/boxinst_r50_fpn_ms-90k_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/boxinst/boxinst_r50_fpn_ms-90k_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..371f252a153855e19f3a3bb25cd42c83a4bb77fd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/boxinst/boxinst_r50_fpn_ms-90k_coco.py
@@ -0,0 +1,93 @@
+_base_ = '../common/ms-90k_coco.py'
+
+# model settings
+model = dict(
+ type='BoxInst',
+ data_preprocessor=dict(
+ type='BoxInstDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32,
+ mask_stride=4,
+ pairwise_size=3,
+ pairwise_dilation=2,
+ pairwise_color_thresh=0.3,
+ bottom_pixels_removed=10),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50'),
+ style='pytorch'),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output', # use P5
+ num_outs=5,
+ relu_before_extra_convs=True),
+ bbox_head=dict(
+ type='BoxInstBboxHead',
+ num_params=593,
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ strides=[8, 16, 32, 64, 128],
+ norm_on_bbox=True,
+ centerness_on_reg=True,
+ dcn_on_last_conv=False,
+ center_sampling=True,
+ conv_bias=True,
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=1.0),
+ loss_centerness=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0)),
+ mask_head=dict(
+ type='BoxInstMaskHead',
+ num_layers=3,
+ feat_channels=16,
+ size_of_interest=8,
+ mask_out_stride=4,
+ topk_masks_per_img=64,
+ mask_feature_head=dict(
+ in_channels=256,
+ feat_channels=128,
+ start_level=0,
+ end_level=2,
+ out_channels=16,
+ mask_stride=8,
+ num_stacked_convs=4,
+ norm_cfg=dict(type='BN', requires_grad=True)),
+ loss_mask=dict(
+ type='DiceLoss',
+ use_sigmoid=True,
+ activate=True,
+ eps=5e-6,
+ loss_weight=1.0)),
+ # model training and testing settings
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100,
+ mask_thr=0.5))
+
+# optimizer
+optim_wrapper = dict(optimizer=dict(lr=0.01))
+
+# evaluator
+val_evaluator = dict(metric=['bbox', 'segm'])
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/boxinst/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/boxinst/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..c97fcdcd636cd4d8d1a1437679f20b96d90fc74f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/boxinst/metafile.yml
@@ -0,0 +1,52 @@
+Collections:
+ - Name: BoxInst
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x A100 GPUs
+ Architecture:
+ - ResNet
+ - FPN
+ - CondInst
+ Paper:
+ URL: https://arxiv.org/abs/2012.02310
+ Title: 'BoxInst: High-Performance Instance Segmentation with Box Annotations'
+ README: configs/boxinst/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v3.0.0rc6/mmdet/models/detectors/boxinst.py#L8
+ Version: v3.0.0rc6
+
+Models:
+ - Name: boxinst_r50_fpn_ms-90k_coco
+ In Collection: BoxInst
+ Config: configs/boxinst/boxinst_r50_fpn_ms-90k_coco.py
+ Metadata:
+ Iterations: 90000
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 30.8
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/boxinst/boxinst_r50_fpn_ms-90k_coco/boxinst_r50_fpn_ms-90k_coco_20221228_163052-6add751a.pth
+
+ - Name: boxinst_r101_fpn_ms-90k_coco
+ In Collection: BoxInst
+ Config: configs/boxinst/boxinst_r101_fpn_ms-90k_coco.py
+ Metadata:
+ Iterations: 90000
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 32.7
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/boxinst/boxinst_r101_fpn_ms-90k_coco/boxinst_r101_fpn_ms-90k_coco_20221229_145106-facf375b.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..30b96f07cece7666c488972e315dbdcededdb21f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/README.md
@@ -0,0 +1,132 @@
+# ByteTrack: Multi-Object Tracking by Associating Every Detection Box
+
+## Abstract
+
+
+
+Multi-object tracking (MOT) aims at estimating bounding boxes and identities of objects in videos. Most methods obtain identities by associating detection boxes whose scores are higher than a threshold. The objects with low detection scores, e.g. occluded objects, are simply thrown away, which brings non-negligible true object missing and fragmented trajectories. To solve this problem, we present a simple, effective and generic association method, tracking by associating every detection box instead of only the high score ones. For the low score detection boxes, we utilize their similarities with tracklets to recover true objects and filter out the background detections. When applied to 9 different state-of-the-art trackers, our method achieves consistent improvement on IDF1 score ranging from 1 to 10 points. To put forwards the state-of-the-art performance of MOT, we design a simple and strong tracker, named ByteTrack. For the first time, we achieve 80.3 MOTA, 77.3 IDF1 and 63.1 HOTA on the test set of MOT17 with 30 FPS running speed on a single V100 GPU.
+
+
+
+
+

+
+
+## Citation
+
+
+
+```latex
+@inproceedings{zhang2021bytetrack,
+ title={ByteTrack: Multi-Object Tracking by Associating Every Detection Box},
+ author={Zhang, Yifu and Sun, Peize and Jiang, Yi and Yu, Dongdong and Yuan, Zehuan and Luo, Ping and Liu, Wenyu and Wang, Xinggang},
+ journal={arXiv preprint arXiv:2110.06864},
+ year={2021}
+}
+```
+
+## Results and models on MOT17
+
+Please note that the performance on `MOT17-half-val` is comparable with the performance reported in the manuscript, while the performance on `MOT17-test` is lower than the performance reported in the manuscript.
+
+The reason is that ByteTrack tunes customized hyper-parameters (e.g., image resolution and the high threshold of detection score) for each video in `MOT17-test` set, while we use unified parameters.
+
+| Method | Detector | Train Set | Test Set | Public | Inf time (fps) | HOTA | MOTA | IDF1 | FP | FN | IDSw. | Config | Download |
+| :-------: | :------: | :---------------------------: | :------------: | :----: | :------------: | :--: | :--: | :--: | :---: | :---: | :---: | :-------------------------------------------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| ByteTrack | YOLOX-X | CrowdHuman + MOT17-half-train | MOT17-half-val | N | - | 67.5 | 78.6 | 78.5 | 12852 | 21060 | 672 | [config](bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17halfval.py) | [model](https://download.openmmlab.com/mmtracking/mot/bytetrack/bytetrack_yolox_x/bytetrack_yolox_x_crowdhuman_mot17-private-half_20211218_205500-1985c9f0.pth) \| [log](https://download.openmmlab.com/mmtracking/mot/bytetrack/bytetrack_yolox_x/bytetrack_yolox_x_crowdhuman_mot17-private-half_20211218_205500.log.json) |
+| ByteTrack | YOLOX-X | CrowdHuman + MOT17-half-train | MOT17-test | N | - | 61.7 | 78.1 | 74.8 | 36705 | 85032 | 2049 | [config](bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17test.py) | [model](https://download.openmmlab.com/mmtracking/mot/bytetrack/bytetrack_yolox_x/bytetrack_yolox_x_crowdhuman_mot17-private-half_20211218_205500-1985c9f0.pth) \| [log](https://download.openmmlab.com/mmtracking/mot/bytetrack/bytetrack_yolox_x/bytetrack_yolox_x_crowdhuman_mot17-private-half_20211218_205500.log.json) |
+
+## Results and models on MOT20
+
+Since there are only 4 videos in `MOT20-train`, ByteTrack is validated on `MOT17-train` rather than `MOT20-half-train`.
+
+Please note that the MOTA on `MOT20-test` is slightly lower than that reported in the manuscript, because we don't tune the threshold for each video.
+
+| Method | Detector | Train Set | Test Set | Public | Inf time (fps) | HOTA | MOTA | IDF1 | FP | FN | IDSw. | Config | Download |
+| :-------: | :------: | :----------------------: | :---------: | :----: | :------------: | :--: | :--: | :--: | :----: | :----: | :---: | :------------------------------------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| ByteTrack | YOLOX-X | CrowdHuman + MOT20-train | MOT17-train | N | - | 57.3 | 64.9 | 71.8 | 33,747 | 83,385 | 1,263 | [config](bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot20train_test-mot20test.py) | [model](https://download.openmmlab.com/mmtracking/mot/bytetrack/bytetrack_yolox_x/bytetrack_yolox_x_crowdhuman_mot20-private_20220506_101040-9ce38a60.pth) \| [log](https://download.openmmlab.com/mmtracking/mot/bytetrack/bytetrack_yolox_x/bytetrack_yolox_x_crowdhuman_mot20-private_20220506_101040.log.json) |
+| ByteTrack | YOLOX-X | CrowdHuman + MOT20-train | MOT20-test | N | - | 61.5 | 77.0 | 75.4 | 33,083 | 84,433 | 1,345 | [config](bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot20train_test-mot20test.py) | [model](https://download.openmmlab.com/mmtracking/mot/bytetrack/bytetrack_yolox_x/bytetrack_yolox_x_crowdhuman_mot20-private_20220506_101040-9ce38a60.pth) \| [log](https://download.openmmlab.com/mmtracking/mot/bytetrack/bytetrack_yolox_x/bytetrack_yolox_x_crowdhuman_mot20-private_20220506_101040.log.json) |
+
+## Get started
+
+### 1. Development Environment Setup
+
+Tracking Development Environment Setup can refer to this [document](../../docs/en/get_started.md).
+
+### 2. Dataset Prepare
+
+Tracking Dataset Prepare can refer to this [document](../../docs/en/user_guides/tracking_dataset_prepare.md).
+
+### 3. Training
+
+Due to the influence of parameters such as learning rate in default configuration file, we recommend using 8 GPUs for training in order to reproduce accuracy. You can use the following command to start the training.
+
+**3.1 Joint training and tracking**
+
+Some algorithm like ByteTrack, OCSORT don't need reid model, so we provide joint training and tracking for convenient.
+
+```shell
+# Training Bytetrack on crowdhuman and mot17-half-train dataset with following command
+# The number after config file represents the number of GPUs used. Here we use 8 GPUs
+bash tools/dist_train.sh configs/bytetrack/bytetrack_yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval.py 8
+```
+
+**3.2 Separate training and tracking**
+
+Of course, we provide train detector independently like SORT, DeepSORT, StrongSORT. Then use this detector to track.
+
+```shell
+# Training Bytetrack on crowdhuman and mot17-half-train dataset with following command
+# The number after config file represents the number of GPUs used. Here we use 8 GPUs
+bash tools/dist_train.sh configs/bytetrack/yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17halfval.py 8
+```
+
+If you want to know about more detailed usage of `train.py/dist_train.sh/slurm_train.sh`,
+please refer to this [document](../../docs/en/user_guides/tracking_train_test.md).
+
+### 4. Testing and evaluation
+
+### 4.1 Example on MOTxx-halfval dataset
+
+**4.1.1 use joint trained detector to evaluating and testing**
+
+```shell
+bash tools/dist_test_tracking.sh configs/bytetrack/bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17halfval.py 8 --checkpoint ${CHECKPOINT_FILE}
+```
+
+**4.1.2 use separate trained detector to evaluating and testing**
+
+```shell
+bash tools/dist_test_tracking.sh configs/bytetrack/bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17halfval.py 8 --detector ${CHECKPOINT_FILE}
+```
+
+**4.1.3 use video_baesd to evaluating and testing**
+
+we also provide two_ways(img_based or video_based) to evaluating and testing.
+if you want to use video_based to evaluating and testing, you can modify config as follows
+
+```
+val_dataloader = dict(
+ sampler=dict(type='DefaultSampler', shuffle=False, round_up=False))
+```
+
+#### 4.2 Example on MOTxx-test dataset
+
+If you want to get the results of the [MOT Challenge](https://motchallenge.net/) test set, please use the following command to generate result files that can be used for submission. It will be stored in `./mot_17_test_res`, you can modify the saved path in `test_evaluator` of the config.
+
+```shell
+bash tools/dist_test_tracking.sh configs/bytetrack/bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17test.py 8 --checkpoint ${CHECKPOINT_FILE}
+```
+
+If you want to know about more detailed usage of `test_tracking.py/dist_test_tracking.sh/slurm_test_tracking.sh`,
+please refer to this [document](../../docs/en/user_guides/tracking_train_test.md).
+
+### 5.Inference
+
+Use a single GPU to predict a video and save it as a video.
+
+```shell
+python demo/mot_demo.py demo/demo_mot.mp4 configs/bytetrack/bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17halfval.py --checkpoint ${CHECKPOINT_FILE} --out mot.mp4
+```
+
+If you want to know about more detailed usage of `mot_demo.py`, please refer to this [document](../../docs/en/user_guides/tracking_inference.md).
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/bytetrack_yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/bytetrack_yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval.py
new file mode 100644
index 0000000000000000000000000000000000000000..24b3f7841947204f2ecea385dcfa8b97fa0c6e85
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/bytetrack_yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval.py
@@ -0,0 +1,249 @@
+_base_ = ['../yolox/yolox_x_8xb8-300e_coco.py']
+
+dataset_type = 'MOTChallengeDataset'
+data_root = 'data/MOT17/'
+
+img_scale = (1440, 800) # weight, height
+batch_size = 4
+
+detector = _base_.model
+detector.pop('data_preprocessor')
+detector.bbox_head.update(dict(num_classes=1))
+detector.test_cfg.nms.update(dict(iou_threshold=0.7))
+detector['init_cfg'] = dict(
+ type='Pretrained',
+ checkpoint= # noqa: E251
+ 'https://download.openmmlab.com/mmdetection/v2.0/yolox/yolox_x_8x8_300e_coco/yolox_x_8x8_300e_coco_20211126_140254-1ef88d67.pth' # noqa: E501
+)
+del _base_.model
+
+model = dict(
+ type='ByteTrack',
+ data_preprocessor=dict(
+ type='TrackDataPreprocessor',
+ pad_size_divisor=32,
+ # in bytetrack, we provide joint train detector and evaluate tracking
+ # performance, use_det_processor means use independent detector
+ # data_preprocessor. of course, you can train detector independently
+ # like strongsort
+ use_det_processor=True,
+ batch_augments=[
+ dict(
+ type='BatchSyncRandomResize',
+ random_size_range=(576, 1024),
+ size_divisor=32,
+ interval=10)
+ ]),
+ detector=detector,
+ tracker=dict(
+ type='ByteTracker',
+ motion=dict(type='KalmanFilter'),
+ obj_score_thrs=dict(high=0.6, low=0.1),
+ init_track_thr=0.7,
+ weight_iou_with_det_scores=True,
+ match_iou_thrs=dict(high=0.1, low=0.5, tentative=0.3),
+ num_frames_retain=30))
+
+train_pipeline = [
+ dict(
+ type='Mosaic',
+ img_scale=img_scale,
+ pad_val=114.0,
+ bbox_clip_border=False),
+ dict(
+ type='RandomAffine',
+ scaling_ratio_range=(0.1, 2),
+ border=(-img_scale[0] // 2, -img_scale[1] // 2),
+ bbox_clip_border=False),
+ dict(
+ type='MixUp',
+ img_scale=img_scale,
+ ratio_range=(0.8, 1.6),
+ pad_val=114.0,
+ bbox_clip_border=False),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='Resize',
+ scale=img_scale,
+ keep_ratio=True,
+ clip_object_border=False),
+ dict(type='Pad', size_divisor=32, pad_val=dict(img=(114.0, 114.0, 114.0))),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1, 1), keep_empty=False),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(
+ type='TransformBroadcaster',
+ transforms=[
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='Resize', scale=img_scale, keep_ratio=True),
+ dict(
+ type='Pad',
+ size_divisor=32,
+ pad_val=dict(img=(114.0, 114.0, 114.0))),
+ dict(type='LoadTrackAnnotations'),
+ ]),
+ dict(type='PackTrackInputs')
+]
+train_dataloader = dict(
+ _delete_=True,
+ batch_size=batch_size,
+ num_workers=4,
+ persistent_workers=True,
+ pin_memory=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ dataset=dict(
+ type='MultiImageMixDataset',
+ dataset=dict(
+ type='ConcatDataset',
+ datasets=[
+ dict(
+ type='CocoDataset',
+ data_root='data/MOT17',
+ ann_file='annotations/half-train_cocoformat.json',
+ data_prefix=dict(img='train'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ metainfo=dict(classes=('pedestrian', )),
+ pipeline=[
+ dict(
+ type='LoadImageFromFile',
+ backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ ]),
+ dict(
+ type='CocoDataset',
+ data_root='data/crowdhuman',
+ ann_file='annotations/crowdhuman_train.json',
+ data_prefix=dict(img='train'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ metainfo=dict(classes=('pedestrian', )),
+ pipeline=[
+ dict(
+ type='LoadImageFromFile',
+ backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ ]),
+ dict(
+ type='CocoDataset',
+ data_root='data/crowdhuman',
+ ann_file='annotations/crowdhuman_val.json',
+ data_prefix=dict(img='val'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ metainfo=dict(classes=('pedestrian', )),
+ pipeline=[
+ dict(
+ type='LoadImageFromFile',
+ backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ ]),
+ ]),
+ pipeline=train_pipeline))
+
+val_dataloader = dict(
+ _delete_=True,
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ pin_memory=True,
+ drop_last=False,
+ # video_based
+ # sampler=dict(type='DefaultSampler', shuffle=False, round_up=False),
+ sampler=dict(type='TrackImgSampler'), # image_based
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/half-val_cocoformat.json',
+ data_prefix=dict(img_path='train'),
+ test_mode=True,
+ pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# optimizer
+# default 8 gpu
+base_lr = 0.001 / 8 * batch_size
+optim_wrapper = dict(optimizer=dict(lr=base_lr))
+
+# some hyper parameters
+# training settings
+max_epochs = 80
+num_last_epochs = 10
+interval = 5
+
+train_cfg = dict(
+ type='EpochBasedTrainLoop',
+ max_epochs=max_epochs,
+ val_begin=70,
+ val_interval=1)
+
+# learning policy
+param_scheduler = [
+ dict(
+ # use quadratic formula to warm up 1 epochs
+ type='QuadraticWarmupLR',
+ by_epoch=True,
+ begin=0,
+ end=1,
+ convert_to_iter_based=True),
+ dict(
+ # use cosine lr from 1 to 70 epoch
+ type='CosineAnnealingLR',
+ eta_min=base_lr * 0.05,
+ begin=1,
+ T_max=max_epochs - num_last_epochs,
+ end=max_epochs - num_last_epochs,
+ by_epoch=True,
+ convert_to_iter_based=True),
+ dict(
+ # use fixed lr during last 10 epochs
+ type='ConstantLR',
+ by_epoch=True,
+ factor=1,
+ begin=max_epochs - num_last_epochs,
+ end=max_epochs,
+ )
+]
+
+custom_hooks = [
+ dict(
+ type='YOLOXModeSwitchHook',
+ num_last_epochs=num_last_epochs,
+ priority=48),
+ dict(type='SyncNormHook', priority=48),
+ dict(
+ type='EMAHook',
+ ema_type='ExpMomentumEMA',
+ momentum=0.0001,
+ update_buffers=True,
+ priority=49)
+]
+
+default_hooks = dict(
+ checkpoint=dict(
+ _delete_=True, type='CheckpointHook', interval=1, max_keep_ckpts=10),
+ visualization=dict(type='TrackVisualizationHook', draw=False))
+
+vis_backends = [dict(type='LocalVisBackend')]
+visualizer = dict(
+ type='TrackLocalVisualizer', vis_backends=vis_backends, name='visualizer')
+
+# evaluator
+val_evaluator = dict(
+ _delete_=True,
+ type='MOTChallengeMetric',
+ metric=['HOTA', 'CLEAR', 'Identity'],
+ postprocess_tracklet_cfg=[
+ dict(type='InterpolateTracklets', min_num_frames=5, max_num_frames=20)
+ ])
+test_evaluator = val_evaluator
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (4 samples per GPU)
+auto_scale_lr = dict(base_batch_size=32)
+
+del detector
+del _base_.tta_model
+del _base_.tta_pipeline
+del _base_.train_dataset
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/bytetrack_yolox_x_8xb4-80e_crowdhuman-mot20train_test-mot20test.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/bytetrack_yolox_x_8xb4-80e_crowdhuman-mot20train_test-mot20test.py
new file mode 100644
index 0000000000000000000000000000000000000000..9202f5fbda29d2a1d4cc81322c99d638ebf475d6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/bytetrack_yolox_x_8xb4-80e_crowdhuman-mot20train_test-mot20test.py
@@ -0,0 +1,127 @@
+_base_ = [
+ './bytetrack_yolox_x_8xb4-80e_crowdhuman-mot17halftrain_'
+ 'test-mot17halfval.py'
+]
+
+dataset_type = 'MOTChallengeDataset'
+
+img_scale = (1600, 896) # weight, height
+
+model = dict(
+ data_preprocessor=dict(
+ type='TrackDataPreprocessor',
+ use_det_processor=True,
+ pad_size_divisor=32,
+ batch_augments=[
+ dict(type='BatchSyncRandomResize', random_size_range=(640, 1152))
+ ]),
+ tracker=dict(
+ weight_iou_with_det_scores=False,
+ match_iou_thrs=dict(high=0.3),
+ ))
+
+train_pipeline = [
+ dict(
+ type='Mosaic',
+ img_scale=img_scale,
+ pad_val=114.0,
+ bbox_clip_border=True),
+ dict(
+ type='RandomAffine',
+ scaling_ratio_range=(0.1, 2),
+ border=(-img_scale[0] // 2, -img_scale[1] // 2),
+ bbox_clip_border=True),
+ dict(
+ type='MixUp',
+ img_scale=img_scale,
+ ratio_range=(0.8, 1.6),
+ pad_val=114.0,
+ bbox_clip_border=True),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='Resize',
+ scale=img_scale,
+ keep_ratio=True,
+ clip_object_border=True),
+ dict(type='Pad', size_divisor=32, pad_val=dict(img=(114.0, 114.0, 114.0))),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1, 1), keep_empty=False),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(
+ type='TransformBroadcaster',
+ transforms=[
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='Resize', scale=img_scale, keep_ratio=True),
+ dict(
+ type='Pad',
+ size_divisor=32,
+ pad_val=dict(img=(114.0, 114.0, 114.0))),
+ dict(type='LoadTrackAnnotations'),
+ ]),
+ dict(type='PackTrackInputs')
+]
+train_dataloader = dict(
+ dataset=dict(
+ type='MultiImageMixDataset',
+ dataset=dict(
+ type='ConcatDataset',
+ datasets=[
+ dict(
+ type='CocoDataset',
+ data_root='data/MOT20',
+ ann_file='annotations/train_cocoformat.json',
+ # TODO: mmdet use img as key, but img_path is needed
+ data_prefix=dict(img='train'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ metainfo=dict(classes=('pedestrian', )),
+ pipeline=[
+ dict(
+ type='LoadImageFromFile',
+ backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ ]),
+ dict(
+ type='CocoDataset',
+ data_root='data/crowdhuman',
+ ann_file='annotations/crowdhuman_train.json',
+ data_prefix=dict(img='train'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ metainfo=dict(classes=('pedestrian', )),
+ pipeline=[
+ dict(
+ type='LoadImageFromFile',
+ backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ ]),
+ dict(
+ type='CocoDataset',
+ data_root='data/crowdhuman',
+ ann_file='annotations/crowdhuman_val.json',
+ data_prefix=dict(img='val'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ metainfo=dict(classes=('pedestrian', )),
+ pipeline=[
+ dict(
+ type='LoadImageFromFile',
+ backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ ]),
+ ]),
+ pipeline=train_pipeline))
+val_dataloader = dict(
+ dataset=dict(ann_file='annotations/train_cocoformat.json'))
+
+test_dataloader = dict(
+ dataset=dict(
+ data_root='data/MOT20', ann_file='annotations/test_cocoformat.json'))
+
+test_evaluator = dict(
+ type='MOTChallengeMetrics',
+ postprocess_tracklet_cfg=[
+ dict(type='InterpolateTracklets', min_num_frames=5, max_num_frames=20)
+ ],
+ format_only=True,
+ outfile_prefix='./mot_20_test_res')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17halfval.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17halfval.py
new file mode 100644
index 0000000000000000000000000000000000000000..9c2119203a46e76cd8b6cc8f755334f58ffb086d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17halfval.py
@@ -0,0 +1,9 @@
+_base_ = [
+ './bytetrack_yolox_x_8xb4-80e_crowdhuman-mot17halftrain_'
+ 'test-mot17halfval.py'
+]
+
+# fp16 settings
+optim_wrapper = dict(type='AmpOptimWrapper', loss_scale='dynamic')
+val_cfg = dict(type='ValLoop', fp16=True)
+test_cfg = dict(type='TestLoop', fp16=True)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17test.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17test.py
new file mode 100644
index 0000000000000000000000000000000000000000..3f4427c18bff66ab1fa2a9ba22517989722d0625
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17test.py
@@ -0,0 +1,17 @@
+_base_ = [
+ './bytetrack/bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-'
+ 'mot17halftrain_test-mot17halfval.py'
+]
+
+test_dataloader = dict(
+ dataset=dict(
+ data_root='data/MOT17/',
+ ann_file='annotations/test_cocoformat.json',
+ data_prefix=dict(img_path='test')))
+test_evaluator = dict(
+ type='MOTChallengeMetrics',
+ postprocess_tracklet_cfg=[
+ dict(type='InterpolateTracklets', min_num_frames=5, max_num_frames=20)
+ ],
+ format_only=True,
+ outfile_prefix='./mot_17_test_res')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot20train_test-mot20test.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot20train_test-mot20test.py
new file mode 100644
index 0000000000000000000000000000000000000000..1016999729263d72bbd75019be4968bc3960e368
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot20train_test-mot20test.py
@@ -0,0 +1,8 @@
+_base_ = [
+ './bytetrack_yolox_x_8xb4-80e_crowdhuman-mot20train_test-mot20test.py'
+]
+
+# fp16 settings
+optim_wrapper = dict(type='AmpOptimWrapper', loss_scale='dynamic')
+val_cfg = dict(type='ValLoop', fp16=True)
+test_cfg = dict(type='TestLoop', fp16=True)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..8ed638cf6dda0b0b3db264aa8847358d78ee0fbe
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/metafile.yml
@@ -0,0 +1,53 @@
+Collections:
+ - Name: ByteTrack
+ Metadata:
+ Training Techniques:
+ - SGD with Momentum
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - YOLOX
+ Paper:
+ URL: https://arxiv.org/abs/2110.06864
+ Title: ByteTrack Multi-Object Tracking by Associating Every Detection Box
+ README: configs/bytetrack/README.md
+
+Models:
+ - Name: bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17halfval
+ In Collection: ByteTrack
+ Config: configs/bytetrack/bytetrack_yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval.py
+ Metadata:
+ Training Data: CrowdHuman + MOT17-half-train
+ Results:
+ - Task: Multiple Object Tracking
+ Dataset: MOT17-half-val
+ Metrics:
+ HOTA: 67.5
+ MOTA: 78.6
+ IDF1: 78.5
+ Weights: https://download.openmmlab.com/mmtracking/mot/bytetrack/bytetrack_yolox_x/bytetrack_yolox_x_crowdhuman_mot17-private-half_20211218_205500-1985c9f0.pth
+
+ - Name: bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17test
+ In Collection: ByteTrack
+ Config: configs/bytetrack/bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17test.py
+ Metadata:
+ Training Data: CrowdHuman + MOT17-half-train
+ Results:
+ - Task: Multiple Object Tracking
+ Dataset: MOT17-test
+ Metrics:
+ MOTA: 78.1
+ IDF1: 74.8
+ Weights: https://download.openmmlab.com/mmtracking/mot/bytetrack/bytetrack_yolox_x/bytetrack_yolox_x_crowdhuman_mot17-private-half_20211218_205500-1985c9f0.pth
+
+ - Name: bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot20train_test-mot20test
+ In Collection: ByteTrack
+ Config: configs/bytetrack/bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot20train_test-mot20test.py
+ Metadata:
+ Training Data: CrowdHuman + MOT20-train
+ Results:
+ - Task: Multiple Object Tracking
+ Dataset: MOT20-test
+ Metrics:
+ MOTA: 77.0
+ IDF1: 75.4
+ Weights: https://download.openmmlab.com/mmtracking/mot/bytetrack/bytetrack_yolox_x/bytetrack_yolox_x_crowdhuman_mot20-private_20220506_101040-9ce38a60.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17halfval.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17halfval.py
new file mode 100644
index 0000000000000000000000000000000000000000..8fc3acd487211d04fb3d6e4504ded5235393e4a7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/bytetrack/yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17halfval.py
@@ -0,0 +1,6 @@
+_base_ = [
+ '../strongsort/yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval.py' # noqa: E501
+]
+
+# fp16 settings
+optim_wrapper = dict(type='AmpOptimWrapper', loss_scale='dynamic')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/carafe/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/carafe/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..61e1fa60fee5d1dc89539874c784c75df63b2ad3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/carafe/README.md
@@ -0,0 +1,42 @@
+# CARAFE
+
+> [CARAFE: Content-Aware ReAssembly of FEatures](https://arxiv.org/abs/1905.02188)
+
+
+
+## Abstract
+
+Feature upsampling is a key operation in a number of modern convolutional network architectures, e.g. feature pyramids. Its design is critical for dense prediction tasks such as object detection and semantic/instance segmentation. In this work, we propose Content-Aware ReAssembly of FEatures (CARAFE), a universal, lightweight and highly effective operator to fulfill this goal. CARAFE has several appealing properties: (1) Large field of view. Unlike previous works (e.g. bilinear interpolation) that only exploit sub-pixel neighborhood, CARAFE can aggregate contextual information within a large receptive field. (2) Content-aware handling. Instead of using a fixed kernel for all samples (e.g. deconvolution), CARAFE enables instance-specific content-aware handling, which generates adaptive kernels on-the-fly. (3) Lightweight and fast to compute. CARAFE introduces little computational overhead and can be readily integrated into modern network architectures. We conduct comprehensive evaluations on standard benchmarks in object detection, instance/semantic segmentation and inpainting. CARAFE shows consistent and substantial gains across all the tasks (1.2%, 1.3%, 1.8%, 1.1db respectively) with negligible computational overhead. It has great potential to serve as a strong building block for future research. It has great potential to serve as a strong building block for future research.
+
+
+

+
+
+## Results and Models
+
+The results on COCO 2017 val is shown in the below table.
+
+| Method | Backbone | Style | Lr schd | Test Proposal Num | Inf time (fps) | Box AP | Mask AP | Config | Download |
+| :--------------------: | :------: | :-----: | :-----: | :---------------: | :------------: | :----: | :-----: | :-----------------------------------------------: | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Faster R-CNN w/ CARAFE | R-50-FPN | pytorch | 1x | 1000 | 16.5 | 38.6 | 38.6 | [config](./faster-rcnn_r50_fpn-carafe_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/carafe/faster_rcnn_r50_fpn_carafe_1x_coco/faster_rcnn_r50_fpn_carafe_1x_coco_bbox_mAP-0.386_20200504_175733-385a75b7.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/carafe/faster_rcnn_r50_fpn_carafe_1x_coco/faster_rcnn_r50_fpn_carafe_1x_coco_20200504_175733.log.json) |
+| - | - | - | - | 2000 | | | | | |
+| Mask R-CNN w/ CARAFE | R-50-FPN | pytorch | 1x | 1000 | 14.0 | 39.3 | 35.8 | [config](./mask-rcnn_r50_fpn-carafe_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/carafe/mask_rcnn_r50_fpn_carafe_1x_coco/mask_rcnn_r50_fpn_carafe_1x_coco_bbox_mAP-0.393__segm_mAP-0.358_20200503_135957-8687f195.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/carafe/mask_rcnn_r50_fpn_carafe_1x_coco/mask_rcnn_r50_fpn_carafe_1x_coco_20200503_135957.log.json) |
+| - | - | - | - | 2000 | | | | | |
+
+## Implementation
+
+The CUDA implementation of CARAFE can be find at https://github.com/myownskyW7/CARAFE.
+
+## Citation
+
+We provide config files to reproduce the object detection & instance segmentation results in the ICCV 2019 Oral paper for [CARAFE: Content-Aware ReAssembly of FEatures](https://arxiv.org/abs/1905.02188).
+
+```latex
+@inproceedings{Wang_2019_ICCV,
+ title = {CARAFE: Content-Aware ReAssembly of FEatures},
+ author = {Wang, Jiaqi and Chen, Kai and Xu, Rui and Liu, Ziwei and Loy, Chen Change and Lin, Dahua},
+ booktitle = {The IEEE International Conference on Computer Vision (ICCV)},
+ month = {October},
+ year = {2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/carafe/faster-rcnn_r50_fpn-carafe_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/carafe/faster-rcnn_r50_fpn-carafe_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..388305cceac2e81eb1b4df6eac36662df7b8bf0d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/carafe/faster-rcnn_r50_fpn-carafe_1x_coco.py
@@ -0,0 +1,20 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ data_preprocessor=dict(pad_size_divisor=64),
+ neck=dict(
+ type='FPN_CARAFE',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5,
+ start_level=0,
+ end_level=-1,
+ norm_cfg=None,
+ act_cfg=None,
+ order=('conv', 'norm', 'act'),
+ upsample_cfg=dict(
+ type='carafe',
+ up_kernel=5,
+ up_group=1,
+ encoder_kernel=3,
+ encoder_dilation=1,
+ compressed_channels=64)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/carafe/mask-rcnn_r50_fpn-carafe_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/carafe/mask-rcnn_r50_fpn-carafe_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6ce621de77aff60f39126136cb25ca9ca38a1c9f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/carafe/mask-rcnn_r50_fpn-carafe_1x_coco.py
@@ -0,0 +1,30 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ data_preprocessor=dict(pad_size_divisor=64),
+ neck=dict(
+ type='FPN_CARAFE',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5,
+ start_level=0,
+ end_level=-1,
+ norm_cfg=None,
+ act_cfg=None,
+ order=('conv', 'norm', 'act'),
+ upsample_cfg=dict(
+ type='carafe',
+ up_kernel=5,
+ up_group=1,
+ encoder_kernel=3,
+ encoder_dilation=1,
+ compressed_channels=64)),
+ roi_head=dict(
+ mask_head=dict(
+ upsample_cfg=dict(
+ type='carafe',
+ scale_factor=2,
+ up_kernel=5,
+ up_group=1,
+ encoder_kernel=3,
+ encoder_dilation=1,
+ compressed_channels=64))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/carafe/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/carafe/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..863c0f49ae6322429e91cf068b06f713a29fcbdc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/carafe/metafile.yml
@@ -0,0 +1,55 @@
+Collections:
+ - Name: CARAFE
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RPN
+ - FPN_CARAFE
+ - ResNet
+ - RoIPool
+ Paper:
+ URL: https://arxiv.org/abs/1905.02188
+ Title: 'CARAFE: Content-Aware ReAssembly of FEatures'
+ README: configs/carafe/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.12.0/mmdet/models/necks/fpn_carafe.py#L11
+ Version: v2.12.0
+
+Models:
+ - Name: faster-rcnn_r50_fpn_carafe_1x_coco
+ In Collection: CARAFE
+ Config: configs/carafe/faster-rcnn_r50_fpn-carafe_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.26
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.6
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/carafe/faster_rcnn_r50_fpn_carafe_1x_coco/faster_rcnn_r50_fpn_carafe_1x_coco_bbox_mAP-0.386_20200504_175733-385a75b7.pth
+
+ - Name: mask-rcnn_r50_fpn_carafe_1x_coco
+ In Collection: CARAFE
+ Config: configs/carafe/mask-rcnn_r50_fpn-carafe_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.31
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.3
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 35.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/carafe/mask_rcnn_r50_fpn_carafe_1x_coco/mask_rcnn_r50_fpn_carafe_1x_coco_bbox_mAP-0.393__segm_mAP-0.358_20200503_135957-8687f195.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..81fce448f9daec77b3e716ac731dce13be751c74
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/README.md
@@ -0,0 +1,79 @@
+# Cascade R-CNN
+
+> [Cascade R-CNN: High Quality Object Detection and Instance Segmentation](https://arxiv.org/abs/1906.09756)
+
+
+
+## Abstract
+
+In object detection, the intersection over union (IoU) threshold is frequently used to define positives/negatives. The threshold used to train a detector defines its quality. While the commonly used threshold of 0.5 leads to noisy (low-quality) detections, detection performance frequently degrades for larger thresholds. This paradox of high-quality detection has two causes: 1) overfitting, due to vanishing positive samples for large thresholds, and 2) inference-time quality mismatch between detector and test hypotheses. A multi-stage object detection architecture, the Cascade R-CNN, composed of a sequence of detectors trained with increasing IoU thresholds, is proposed to address these problems. The detectors are trained sequentially, using the output of a detector as training set for the next. This resampling progressively improves hypotheses quality, guaranteeing a positive training set of equivalent size for all detectors and minimizing overfitting. The same cascade is applied at inference, to eliminate quality mismatches between hypotheses and detectors. An implementation of the Cascade R-CNN without bells or whistles achieves state-of-the-art performance on the COCO dataset, and significantly improves high-quality detection on generic and specific object detection datasets, including VOC, KITTI, CityPerson, and WiderFace. Finally, the Cascade R-CNN is generalized to instance segmentation, with nontrivial improvements over the Mask R-CNN.
+
+
+

+
+
+## Results and Models
+
+### Cascade R-CNN
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :-------------: | :-----: | :-----: | :------: | :------------: | :----: | :-------------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | caffe | 1x | 4.2 | | 40.4 | [config](./cascade-rcnn_r50-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_r50_caffe_fpn_1x_coco/cascade_rcnn_r50_caffe_fpn_1x_coco_bbox_mAP-0.404_20200504_174853-b857be87.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_r50_caffe_fpn_1x_coco/cascade_rcnn_r50_caffe_fpn_1x_coco_20200504_174853.log.json) |
+| R-50-FPN | pytorch | 1x | 4.4 | 16.1 | 40.3 | [config](./cascade-rcnn_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_r50_fpn_1x_coco/cascade_rcnn_r50_fpn_1x_coco_20200316-3dc56deb.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_r50_fpn_1x_coco/cascade_rcnn_r50_fpn_1x_coco_20200316_214748.log.json) |
+| R-50-FPN | pytorch | 20e | - | - | 41.0 | [config](./cascade-rcnn_r50_fpn_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_r50_fpn_20e_coco/cascade_rcnn_r50_fpn_20e_coco_bbox_mAP-0.41_20200504_175131-e9872a90.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_r50_fpn_20e_coco/cascade_rcnn_r50_fpn_20e_coco_20200504_175131.log.json) |
+| R-101-FPN | caffe | 1x | 6.2 | | 42.3 | [config](./cascade-rcnn_r101-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_r101_caffe_fpn_1x_coco/cascade_rcnn_r101_caffe_fpn_1x_coco_bbox_mAP-0.423_20200504_175649-cab8dbd5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_r101_caffe_fpn_1x_coco/cascade_rcnn_r101_caffe_fpn_1x_coco_20200504_175649.log.json) |
+| R-101-FPN | pytorch | 1x | 6.4 | 13.5 | 42.0 | [config](./cascade-rcnn_r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_r101_fpn_1x_coco/cascade_rcnn_r101_fpn_1x_coco_20200317-0b6a2fbf.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_r101_fpn_1x_coco/cascade_rcnn_r101_fpn_1x_coco_20200317_101744.log.json) |
+| R-101-FPN | pytorch | 20e | - | - | 42.5 | [config](./cascade-rcnn_r101_fpn_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_r101_fpn_20e_coco/cascade_rcnn_r101_fpn_20e_coco_bbox_mAP-0.425_20200504_231812-5057dcc5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_r101_fpn_20e_coco/cascade_rcnn_r101_fpn_20e_coco_20200504_231812.log.json) |
+| X-101-32x4d-FPN | pytorch | 1x | 7.6 | 10.9 | 43.7 | [config](./cascade-rcnn_x101-32x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_x101_32x4d_fpn_1x_coco/cascade_rcnn_x101_32x4d_fpn_1x_coco_20200316-95c2deb6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_x101_32x4d_fpn_1x_coco/cascade_rcnn_x101_32x4d_fpn_1x_coco_20200316_055608.log.json) |
+| X-101-32x4d-FPN | pytorch | 20e | 7.6 | | 43.7 | [config](./cascade-rcnn_x101-32x4d_fpn_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_x101_32x4d_fpn_20e_coco/cascade_rcnn_x101_32x4d_fpn_20e_coco_20200906_134608-9ae0a720.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_x101_32x4d_fpn_20e_coco/cascade_rcnn_x101_32x4d_fpn_20e_coco_20200906_134608.log.json) |
+| X-101-64x4d-FPN | pytorch | 1x | 10.7 | | 44.7 | [config](./cascade-rcnn_x101-64x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_x101_64x4d_fpn_1x_coco/cascade_rcnn_x101_64x4d_fpn_1x_coco_20200515_075702-43ce6a30.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_x101_64x4d_fpn_1x_coco/cascade_rcnn_x101_64x4d_fpn_1x_coco_20200515_075702.log.json) |
+| X-101-64x4d-FPN | pytorch | 20e | 10.7 | | 44.5 | [config](./cascade-rcnn_x101_64x4d_fpn_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_x101_64x4d_fpn_20e_coco/cascade_rcnn_x101_64x4d_fpn_20e_coco_20200509_224357-051557b1.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_x101_64x4d_fpn_20e_coco/cascade_rcnn_x101_64x4d_fpn_20e_coco_20200509_224357.log.json) |
+
+### Cascade Mask R-CNN
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :-------------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :------------------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | caffe | 1x | 5.9 | | 41.2 | 36.0 | [config](./cascade-mask-rcnn_r50-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r50_caffe_fpn_1x_coco/cascade_mask_rcnn_r50_caffe_fpn_1x_coco_bbox_mAP-0.412__segm_mAP-0.36_20200504_174659-5004b251.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r50_caffe_fpn_1x_coco/cascade_mask_rcnn_r50_caffe_fpn_1x_coco_20200504_174659.log.json) |
+| R-50-FPN | pytorch | 1x | 6.0 | 11.2 | 41.2 | 35.9 | [config](./cascade-mask-rcnn_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r50_fpn_1x_coco/cascade_mask_rcnn_r50_fpn_1x_coco_20200203-9d4dcb24.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r50_fpn_1x_coco/cascade_mask_rcnn_r50_fpn_1x_coco_20200203_170449.log.json) |
+| R-50-FPN | pytorch | 20e | - | - | 41.9 | 36.5 | [config](./cascade-mask-rcnn_r50_fpn_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r50_fpn_20e_coco/cascade_mask_rcnn_r50_fpn_20e_coco_bbox_mAP-0.419__segm_mAP-0.365_20200504_174711-4af8e66e.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r50_fpn_20e_coco/cascade_mask_rcnn_r50_fpn_20e_coco_20200504_174711.log.json) |
+| R-101-FPN | caffe | 1x | 7.8 | | 43.2 | 37.6 | [config](./cascade-mask-rcnn_r101-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r101_caffe_fpn_1x_coco/cascade_mask_rcnn_r101_caffe_fpn_1x_coco_bbox_mAP-0.432__segm_mAP-0.376_20200504_174813-5c1e9599.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r101_caffe_fpn_1x_coco/cascade_mask_rcnn_r101_caffe_fpn_1x_coco_20200504_174813.log.json) |
+| R-101-FPN | pytorch | 1x | 7.9 | 9.8 | 42.9 | 37.3 | [config](./cascade-mask-rcnn_r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r101_fpn_1x_coco/cascade_mask_rcnn_r101_fpn_1x_coco_20200203-befdf6ee.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r101_fpn_1x_coco/cascade_mask_rcnn_r101_fpn_1x_coco_20200203_092521.log.json) |
+| R-101-FPN | pytorch | 20e | - | - | 43.4 | 37.8 | [config](./cascade-mask-rcnn_r101_fpn_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r101_fpn_20e_coco/cascade_mask_rcnn_r101_fpn_20e_coco_bbox_mAP-0.434__segm_mAP-0.378_20200504_174836-005947da.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r101_fpn_20e_coco/cascade_mask_rcnn_r101_fpn_20e_coco_20200504_174836.log.json) |
+| X-101-32x4d-FPN | pytorch | 1x | 9.2 | 8.6 | 44.3 | 38.3 | [config](./cascade-mask-rcnn_x101-32x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_32x4d_fpn_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_1x_coco_20200201-0f411b1f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_32x4d_fpn_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_1x_coco_20200201_052416.log.json) |
+| X-101-32x4d-FPN | pytorch | 20e | 9.2 | - | 45.0 | 39.0 | [config](./cascade-mask-rcnn_x101-32x4d_fpn_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_32x4d_fpn_20e_coco/cascade_mask_rcnn_x101_32x4d_fpn_20e_coco_20200528_083917-ed1f4751.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_32x4d_fpn_20e_coco/cascade_mask_rcnn_x101_32x4d_fpn_20e_coco_20200528_083917.log.json) |
+| X-101-64x4d-FPN | pytorch | 1x | 12.2 | 6.7 | 45.3 | 39.2 | [config](./cascade-mask-rcnn_x101-64x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_64x4d_fpn_1x_coco/cascade_mask_rcnn_x101_64x4d_fpn_1x_coco_20200203-9a2db89d.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_64x4d_fpn_1x_coco/cascade_mask_rcnn_x101_64x4d_fpn_1x_coco_20200203_044059.log.json) |
+| X-101-64x4d-FPN | pytorch | 20e | 12.2 | | 45.6 | 39.5 | [config](./cascade-mask-rcnn_x101-64x4d_fpn_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_64x4d_fpn_20e_coco/cascade_mask_rcnn_x101_64x4d_fpn_20e_coco_20200512_161033-bdb5126a.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_64x4d_fpn_20e_coco/cascade_mask_rcnn_x101_64x4d_fpn_20e_coco_20200512_161033.log.json) |
+
+**Notes:**
+
+- The `20e` schedule in Cascade (Mask) R-CNN indicates decreasing the lr at 16 and 19 epochs, with a total of 20 epochs.
+
+## Pre-trained Models
+
+We also train some models with longer schedules and multi-scale training for Cascade Mask R-CNN. The users could finetune them for downstream tasks.
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :-------------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :--------------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | caffe | 3x | 5.7 | | 44.0 | 38.1 | [config](./cascade-mask-rcnn_r50-caffe_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r50_caffe_fpn_mstrain_3x_coco/cascade_mask_rcnn_r50_caffe_fpn_mstrain_3x_coco_20210707_002651-6e29b3a6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r50_caffe_fpn_mstrain_3x_coco/cascade_mask_rcnn_r50_caffe_fpn_mstrain_3x_coco_20210707_002651.log.json) |
+| R-50-FPN | pytorch | 3x | 5.9 | | 44.3 | 38.5 | [config](./cascade-mask-rcnn_r50_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r50_fpn_mstrain_3x_coco/cascade_mask_rcnn_r50_fpn_mstrain_3x_coco_20210628_164719-5bdc3824.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r50_fpn_mstrain_3x_coco/cascade_mask_rcnn_r50_fpn_mstrain_3x_coco_20210628_164719.log.json) |
+| R-101-FPN | caffe | 3x | 7.7 | | 45.4 | 39.5 | [config](./cascade-mask-rcnn_r101-caffe_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r101_caffe_fpn_mstrain_3x_coco/cascade_mask_rcnn_r101_caffe_fpn_mstrain_3x_coco_20210707_002620-a5bd2389.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r101_caffe_fpn_mstrain_3x_coco/cascade_mask_rcnn_r101_caffe_fpn_mstrain_3x_coco_20210707_002620.log.json) |
+| R-101-FPN | pytorch | 3x | 7.8 | | 45.5 | 39.6 | [config](./cascade-mask-rcnn_r101_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r101_fpn_mstrain_3x_coco/cascade_mask_rcnn_r101_fpn_mstrain_3x_coco_20210628_165236-51a2d363.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r101_fpn_mstrain_3x_coco/cascade_mask_rcnn_r101_fpn_mstrain_3x_coco_20210628_165236.log.json) |
+| X-101-32x4d-FPN | pytorch | 3x | 9.0 | | 46.3 | 40.1 | [config](./cascade-mask-rcnn_x101-32x4d_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_32x4d_fpn_mstrain_3x_coco/cascade_mask_rcnn_x101_32x4d_fpn_mstrain_3x_coco_20210706_225234-40773067.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_32x4d_fpn_mstrain_3x_coco/cascade_mask_rcnn_x101_32x4d_fpn_mstrain_3x_coco_20210706_225234.log.json) |
+| X-101-32x8d-FPN | pytorch | 3x | 12.1 | | 46.1 | 39.9 | [config](./cascade-mask-rcnn_x101-32x8d_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_32x8d_fpn_mstrain_3x_coco/cascade_mask_rcnn_x101_32x8d_fpn_mstrain_3x_coco_20210719_180640-9ff7e76f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_32x8d_fpn_mstrain_3x_coco/cascade_mask_rcnn_x101_32x8d_fpn_mstrain_3x_coco_20210719_180640.log.json) |
+| X-101-64x4d-FPN | pytorch | 3x | 12.0 | | 46.6 | 40.3 | [config](./cascade-mask-rcnn_x101-64x4d_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_64x4d_fpn_mstrain_3x_coco/cascade_mask_rcnn_x101_64x4d_fpn_mstrain_3x_coco_20210719_210311-d3e64ba0.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_64x4d_fpn_mstrain_3x_coco/cascade_mask_rcnn_x101_64x4d_fpn_mstrain_3x_coco_20210719_210311.log.json) |
+
+## Citation
+
+```latex
+@article{Cai_2019,
+ title={Cascade R-CNN: High Quality Object Detection and Instance Segmentation},
+ ISSN={1939-3539},
+ url={http://dx.doi.org/10.1109/tpami.2019.2956516},
+ DOI={10.1109/tpami.2019.2956516},
+ journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
+ publisher={Institute of Electrical and Electronics Engineers (IEEE)},
+ author={Cai, Zhaowei and Vasconcelos, Nuno},
+ year={2019},
+ pages={1–1}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r101-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r101-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6d85340e1cb92c60293c3710d05ef708d3726fdd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r101-caffe_fpn_1x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './cascade-mask-rcnn_r50-caffe_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet101_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r101-caffe_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r101-caffe_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a6855ee8c6fffd5e8d48f6cc2bb41e9dde9f6516
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r101-caffe_fpn_ms-3x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './cascade-mask-rcnn_r50-caffe_fpn_ms-3x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet101_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c3d962c229d2621e7364c13959e3c4c1137edef1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r101_fpn_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './cascade-mask-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r101_fpn_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r101_fpn_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..497148f513edb79ca58f719f242be6274f923a65
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r101_fpn_20e_coco.py
@@ -0,0 +1,6 @@
+_base_ = './cascade-mask-rcnn_r50_fpn_20e_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r101_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r101_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..183b5c50ff5563d987b2937d27d6d02bdd6cc2bd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r101_fpn_ms-3x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './cascade-mask-rcnn_r50_fpn_ms-3x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r50-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r50-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..497f68c4ab458ec49ad1d0c89cabbb2c0eb444f3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r50-caffe_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = ['./cascade-mask-rcnn_r50_fpn_1x_coco.py']
+
+model = dict(
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False),
+ backbone=dict(
+ norm_cfg=dict(requires_grad=False),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r50-caffe_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r50-caffe_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6677a9fea501a7683475dc8b865659cef5485bbe
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r50-caffe_fpn_ms-3x_coco.py
@@ -0,0 +1,18 @@
+_base_ = [
+ '../common/ms_3x_coco-instance.py',
+ '../_base_/models/cascade-mask-rcnn_r50_fpn.py'
+]
+
+model = dict(
+ # use caffe img_norm
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False),
+ backbone=dict(
+ norm_cfg=dict(requires_grad=False),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f59bb94eaaf3e850e971268383cd0275bcddf54d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r50_fpn_1x_coco.py
@@ -0,0 +1,5 @@
+_base_ = [
+ '../_base_/models/cascade-mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r50_fpn_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r50_fpn_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..35c8aa6748d25e4c9c834478488ee21b44c8f2bd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r50_fpn_20e_coco.py
@@ -0,0 +1,5 @@
+_base_ = [
+ '../_base_/models/cascade-mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_20e.py', '../_base_/default_runtime.py'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r50_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r50_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b15006f451f346216243dc61140e9907535f0b20
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_r50_fpn_ms-3x_coco.py
@@ -0,0 +1,4 @@
+_base_ = [
+ '../common/ms_3x_coco-instance.py',
+ '../_base_/models/cascade-mask-rcnn_r50_fpn.py'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-32x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-32x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..87a4cc325a10b01cbf5a91e336da2281bc19a728
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-32x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './cascade-mask-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-32x4d_fpn_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-32x4d_fpn_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5e8dcaa6891877c89acb024b9811a4fe7a87bc3b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-32x4d_fpn_20e_coco.py
@@ -0,0 +1,14 @@
+_base_ = './cascade-mask-rcnn_r50_fpn_20e_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-32x4d_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-32x4d_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3a0f61b9aee2b0ab80c5c9b998a73826e5ff45a6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-32x4d_fpn_ms-3x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './cascade-mask-rcnn_r50_fpn_ms-3x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-32x8d_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-32x8d_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8cf08306850bdaef776a0ce53b88b23b9013a1a0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-32x8d_fpn_ms-3x_coco.py
@@ -0,0 +1,24 @@
+_base_ = './cascade-mask-rcnn_r50_fpn_ms-3x_coco.py'
+
+model = dict(
+ # ResNeXt-101-32x8d model trained with Caffe2 at FB,
+ # so the mean and std need to be changed.
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[57.375, 57.120, 58.395],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=8,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnext101_32x8d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-64x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-64x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..fb2e6b6b9507dcf38403d38499e1d57bd792a4da
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-64x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './cascade-mask-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-64x4d_fpn_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-64x4d_fpn_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..cc20c171542b5d75634d99d9ed25eea3acf8df19
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-64x4d_fpn_20e_coco.py
@@ -0,0 +1,14 @@
+_base_ = './cascade-mask-rcnn_r50_fpn_20e_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-64x4d_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-64x4d_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f4ecc42655903c271e7e181b719d09821118a204
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-mask-rcnn_x101-64x4d_fpn_ms-3x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './cascade-mask-rcnn_r50_fpn_ms-3x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r101-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r101-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b6eaee2db700b897255ed44a5fd30bc23929388f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r101-caffe_fpn_1x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './cascade-rcnn_r50-caffe_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet101_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1cdf5108b7d2908e420c52c59f8a9805c7989702
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r101_fpn_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './cascade-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r101_fpn_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r101_fpn_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..84c285fc9e59d4191e79dd337ece2baff3d38b02
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r101_fpn_20e_coco.py
@@ -0,0 +1,6 @@
+_base_ = './cascade-rcnn_r50_fpn_20e_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r101_fpn_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r101_fpn_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1fc52e9cb8e1e9c27d45e32200b0b72efa8c363d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r101_fpn_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './cascade-rcnn_r50_fpn_8xb8-amp-lsj-200e_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r18_fpn_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r18_fpn_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..aa30a3d07f5644dfc6f79f0eafc374518149e777
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r18_fpn_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './cascade-rcnn_r50_fpn_8xb8-amp-lsj-200e_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=18,
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet18')),
+ neck=dict(in_channels=[64, 128, 256, 512]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r50-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r50-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ad90e259b2d8410309bfd877b74755524b94f788
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r50-caffe_fpn_1x_coco.py
@@ -0,0 +1,16 @@
+_base_ = './cascade-rcnn_r50_fpn_1x_coco.py'
+
+model = dict(
+ # use caffe img_norm
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ norm_cfg=dict(requires_grad=False),
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1a07c8b2302b9c2337d4da2d32c388142ca1f748
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r50_fpn_1x_coco.py
@@ -0,0 +1,5 @@
+_base_ = [
+ '../_base_/models/cascade-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r50_fpn_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r50_fpn_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..30f3ff106018ba51173f018c196cf62a88fdb172
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r50_fpn_20e_coco.py
@@ -0,0 +1,5 @@
+_base_ = [
+ '../_base_/models/cascade-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_20e.py', '../_base_/default_runtime.py'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r50_fpn_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r50_fpn_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..cd25f02608c3f51a59e35185a41080c6e8e3a1ea
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_r50_fpn_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,23 @@
+_base_ = [
+ '../_base_/models/cascade-rcnn_r50_fpn.py',
+ '../common/lsj-200e_coco-detection.py'
+]
+image_size = (1024, 1024)
+batch_augments = [dict(type='BatchFixedSizePad', size=image_size)]
+
+# disable allowed_border to avoid potential errors.
+model = dict(
+ data_preprocessor=dict(batch_augments=batch_augments),
+ train_cfg=dict(rpn=dict(allowed_border=-1)))
+
+train_dataloader = dict(batch_size=8, num_workers=4)
+# Enable automatic-mixed-precision training with AmpOptimWrapper.
+optim_wrapper = dict(
+ type='AmpOptimWrapper',
+ optimizer=dict(
+ type='SGD', lr=0.02 * 4, momentum=0.9, weight_decay=0.00004))
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_x101-32x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_x101-32x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..50e0b9544592d61b3c14ec7f64f3e6eaa2e96a57
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_x101-32x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './cascade-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_x101-32x4d_fpn_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_x101-32x4d_fpn_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6120189205d883d98b2d323a160ec54ea26aab13
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_x101-32x4d_fpn_20e_coco.py
@@ -0,0 +1,14 @@
+_base_ = './cascade-rcnn_r50_fpn_20e_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_x101-64x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_x101-64x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..29475e39273dccad13058e9114728770e77f71ef
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_x101-64x4d_fpn_1x_coco.py
@@ -0,0 +1,15 @@
+_base_ = './cascade-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ type='CascadeRCNN',
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_x101_64x4d_fpn_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_x101_64x4d_fpn_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e2aa57eaaf43788fc3628f1463e94405279c7416
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/cascade-rcnn_x101_64x4d_fpn_20e_coco.py
@@ -0,0 +1,15 @@
+_base_ = './cascade-rcnn_r50_fpn_20e_coco.py'
+model = dict(
+ type='CascadeRCNN',
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..7e0385daeed3f3310dc7f9a8b64c99b5cb8324b4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rcnn/metafile.yml
@@ -0,0 +1,545 @@
+Collections:
+ - Name: Cascade R-CNN
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Cascade R-CNN
+ - FPN
+ - RPN
+ - ResNet
+ - RoIAlign
+ Paper:
+ URL: http://dx.doi.org/10.1109/tpami.2019.2956516
+ Title: 'Cascade R-CNN: Delving into High Quality Object Detection'
+ README: configs/cascade_rcnn/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/detectors/cascade_rcnn.py#L6
+ Version: v2.0.0
+ - Name: Cascade Mask R-CNN
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Cascade R-CNN
+ - FPN
+ - RPN
+ - ResNet
+ - RoIAlign
+ Paper:
+ URL: http://dx.doi.org/10.1109/tpami.2019.2956516
+ Title: 'Cascade R-CNN: Delving into High Quality Object Detection'
+ README: configs/cascade_rcnn/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/detectors/cascade_rcnn.py#L6
+ Version: v2.0.0
+
+Models:
+ - Name: cascade-rcnn_r50-caffe_fpn_1x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-rcnn_r50-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.2
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_r50_caffe_fpn_1x_coco/cascade_rcnn_r50_caffe_fpn_1x_coco_bbox_mAP-0.404_20200504_174853-b857be87.pth
+
+ - Name: cascade-rcnn_r50_fpn_1x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-rcnn_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.4
+ inference time (ms/im):
+ - value: 62.11
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_r50_fpn_1x_coco/cascade_rcnn_r50_fpn_1x_coco_20200316-3dc56deb.pth
+
+ - Name: cascade-rcnn_r50_fpn_20e_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-rcnn_r50_fpn_20e_coco.py
+ Metadata:
+ Training Memory (GB): 4.4
+ inference time (ms/im):
+ - value: 62.11
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 20
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_r50_fpn_20e_coco/cascade_rcnn_r50_fpn_20e_coco_bbox_mAP-0.41_20200504_175131-e9872a90.pth
+
+ - Name: cascade-rcnn_r101-caffe_fpn_1x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-rcnn_r101-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.2
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_r101_caffe_fpn_1x_coco/cascade_rcnn_r101_caffe_fpn_1x_coco_bbox_mAP-0.423_20200504_175649-cab8dbd5.pth
+
+ - Name: cascade-rcnn_r101_fpn_1x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-rcnn_r101_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.4
+ inference time (ms/im):
+ - value: 74.07
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_r101_fpn_1x_coco/cascade_rcnn_r101_fpn_1x_coco_20200317-0b6a2fbf.pth
+
+ - Name: cascade-rcnn_r101_fpn_20e_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-rcnn_r101_fpn_20e_coco.py
+ Metadata:
+ Training Memory (GB): 6.4
+ inference time (ms/im):
+ - value: 74.07
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 20
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_r101_fpn_20e_coco/cascade_rcnn_r101_fpn_20e_coco_bbox_mAP-0.425_20200504_231812-5057dcc5.pth
+
+ - Name: cascade-rcnn_x101-32x4d_fpn_1x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-rcnn_x101-32x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.6
+ inference time (ms/im):
+ - value: 91.74
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_x101_32x4d_fpn_1x_coco/cascade_rcnn_x101_32x4d_fpn_1x_coco_20200316-95c2deb6.pth
+
+ - Name: cascade-rcnn_x101-32x4d_fpn_20e_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-rcnn_x101-32x4d_fpn_20e_coco.py
+ Metadata:
+ Training Memory (GB): 7.6
+ Epochs: 20
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_x101_32x4d_fpn_20e_coco/cascade_rcnn_x101_32x4d_fpn_20e_coco_20200906_134608-9ae0a720.pth
+
+ - Name: cascade-rcnn_x101-64x4d_fpn_1x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-rcnn_x101-64x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 10.7
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_x101_64x4d_fpn_1x_coco/cascade_rcnn_x101_64x4d_fpn_1x_coco_20200515_075702-43ce6a30.pth
+
+ - Name: cascade-rcnn_x101_64x4d_fpn_20e_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-rcnn_x101_64x4d_fpn_20e_coco.py
+ Metadata:
+ Training Memory (GB): 10.7
+ Epochs: 20
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_rcnn_x101_64x4d_fpn_20e_coco/cascade_rcnn_x101_64x4d_fpn_20e_coco_20200509_224357-051557b1.pth
+
+ - Name: cascade-mask-rcnn_r50-caffe_fpn_1x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-mask-rcnn_r50-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.9
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r50_caffe_fpn_1x_coco/cascade_mask_rcnn_r50_caffe_fpn_1x_coco_bbox_mAP-0.412__segm_mAP-0.36_20200504_174659-5004b251.pth
+
+ - Name: cascade-mask-rcnn_r50_fpn_1x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-mask-rcnn_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.0
+ inference time (ms/im):
+ - value: 89.29
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 35.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r50_fpn_1x_coco/cascade_mask_rcnn_r50_fpn_1x_coco_20200203-9d4dcb24.pth
+
+ - Name: cascade-mask-rcnn_r50_fpn_20e_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-mask-rcnn_r50_fpn_20e_coco.py
+ Metadata:
+ Training Memory (GB): 6.0
+ inference time (ms/im):
+ - value: 89.29
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 20
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.9
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r50_fpn_20e_coco/cascade_mask_rcnn_r50_fpn_20e_coco_bbox_mAP-0.419__segm_mAP-0.365_20200504_174711-4af8e66e.pth
+
+ - Name: cascade-mask-rcnn_r101-caffe_fpn_1x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-mask-rcnn_r101-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.8
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r101_caffe_fpn_1x_coco/cascade_mask_rcnn_r101_caffe_fpn_1x_coco_bbox_mAP-0.432__segm_mAP-0.376_20200504_174813-5c1e9599.pth
+
+ - Name: cascade-mask-rcnn_r101_fpn_1x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-mask-rcnn_r101_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.9
+ inference time (ms/im):
+ - value: 102.04
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.9
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r101_fpn_1x_coco/cascade_mask_rcnn_r101_fpn_1x_coco_20200203-befdf6ee.pth
+
+ - Name: cascade-mask-rcnn_r101_fpn_20e_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-mask-rcnn_r101_fpn_20e_coco.py
+ Metadata:
+ Training Memory (GB): 7.9
+ inference time (ms/im):
+ - value: 102.04
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 20
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r101_fpn_20e_coco/cascade_mask_rcnn_r101_fpn_20e_coco_bbox_mAP-0.434__segm_mAP-0.378_20200504_174836-005947da.pth
+
+ - Name: cascade-mask-rcnn_x101-32x4d_fpn_1x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-mask-rcnn_x101-32x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 9.2
+ inference time (ms/im):
+ - value: 116.28
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.3
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_32x4d_fpn_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_1x_coco_20200201-0f411b1f.pth
+
+ - Name: cascade-mask-rcnn_x101-32x4d_fpn_20e_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-mask-rcnn_x101-32x4d_fpn_20e_coco.py
+ Metadata:
+ Training Memory (GB): 9.2
+ inference time (ms/im):
+ - value: 116.28
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 20
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_32x4d_fpn_20e_coco/cascade_mask_rcnn_x101_32x4d_fpn_20e_coco_20200528_083917-ed1f4751.pth
+
+ - Name: cascade-mask-rcnn_x101-64x4d_fpn_1x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-mask-rcnn_x101-64x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 12.2
+ inference time (ms/im):
+ - value: 149.25
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.3
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_64x4d_fpn_1x_coco/cascade_mask_rcnn_x101_64x4d_fpn_1x_coco_20200203-9a2db89d.pth
+
+ - Name: cascade-mask-rcnn_x101-64x4d_fpn_20e_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-mask-rcnn_x101-64x4d_fpn_20e_coco.py
+ Metadata:
+ Training Memory (GB): 12.2
+ Epochs: 20
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.6
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_64x4d_fpn_20e_coco/cascade_mask_rcnn_x101_64x4d_fpn_20e_coco_20200512_161033-bdb5126a.pth
+
+ - Name: cascade-mask-rcnn_r50-caffe_fpn_ms-3x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-mask-rcnn_r50-caffe_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 5.7
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r50_caffe_fpn_mstrain_3x_coco/cascade_mask_rcnn_r50_caffe_fpn_mstrain_3x_coco_20210707_002651-6e29b3a6.pth
+
+ - Name: cascade-mask-rcnn_r50_fpn_mstrain_3x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-mask-rcnn_r50_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 5.9
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.3
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r50_fpn_mstrain_3x_coco/cascade_mask_rcnn_r50_fpn_mstrain_3x_coco_20210628_164719-5bdc3824.pth
+
+ - Name: cascade-mask-rcnn_r101-caffe_fpn_ms-3x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-mask-rcnn_r101-caffe_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 7.7
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r101_caffe_fpn_mstrain_3x_coco/cascade_mask_rcnn_r101_caffe_fpn_mstrain_3x_coco_20210707_002620-a5bd2389.pth
+
+ - Name: cascade-mask-rcnn_r101_fpn_ms-3x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-mask-rcnn_r101_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 7.8
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_r101_fpn_mstrain_3x_coco/cascade_mask_rcnn_r101_fpn_mstrain_3x_coco_20210628_165236-51a2d363.pth
+
+ - Name: cascade-mask-rcnn_x101-32x4d_fpn_ms-3x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-mask-rcnn_x101-32x4d_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 9.0
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.3
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 40.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_32x4d_fpn_mstrain_3x_coco/cascade_mask_rcnn_x101_32x4d_fpn_mstrain_3x_coco_20210706_225234-40773067.pth
+
+ - Name: cascade-mask-rcnn_x101-32x8d_fpn_ms-3x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-mask-rcnn_x101-32x8d_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 12.1
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.1
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_32x8d_fpn_mstrain_3x_coco/cascade_mask_rcnn_x101_32x8d_fpn_mstrain_3x_coco_20210719_180640-9ff7e76f.pth
+
+ - Name: cascade-mask-rcnn_x101-64x4d_fpn_ms-3x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/cascade_rcnn/cascade-mask-rcnn_x101-64x4d_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 12.0
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.6
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 40.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rcnn/cascade_mask_rcnn_x101_64x4d_fpn_mstrain_3x_coco/cascade_mask_rcnn_x101_64x4d_fpn_mstrain_3x_coco_20210719_210311-d3e64ba0.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rpn/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rpn/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..868a25eda26967576db85dc0686dda53a1d9c9b1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rpn/README.md
@@ -0,0 +1,41 @@
+# Cascade RPN
+
+> [Cascade RPN: Delving into High-Quality Region Proposal Network with Adaptive Convolution](https://arxiv.org/abs/1909.06720)
+
+
+
+## Abstract
+
+This paper considers an architecture referred to as Cascade Region Proposal Network (Cascade RPN) for improving the region-proposal quality and detection performance by systematically addressing the limitation of the conventional RPN that heuristically defines the anchors and aligns the features to the anchors. First, instead of using multiple anchors with predefined scales and aspect ratios, Cascade RPN relies on a single anchor per location and performs multi-stage refinement. Each stage is progressively more stringent in defining positive samples by starting out with an anchor-free metric followed by anchor-based metrics in the ensuing stages. Second, to attain alignment between the features and the anchors throughout the stages, adaptive convolution is proposed that takes the anchors in addition to the image features as its input and learns the sampled features guided by the anchors. A simple implementation of a two-stage Cascade RPN achieves AR 13.4 points higher than that of the conventional RPN, surpassing any existing region proposal methods. When adopting to Fast R-CNN and Faster R-CNN, Cascade RPN can improve the detection mAP by 3.1 and 3.5 points, respectively.
+
+
+

+
+
+## Results and Models
+
+### Region proposal performance
+
+| Method | Backbone | Style | Mem (GB) | Train time (s/iter) | Inf time (fps) | AR 1000 | Config | Download |
+| :----: | :------: | :---: | :------: | :-----------------: | :------------: | :-----: | :----------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------: |
+| CRPN | R-50-FPN | caffe | - | - | - | 72.0 | [config](./cascade-rpn_r50-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rpn/crpn_r50_caffe_fpn_1x_coco/cascade_rpn_r50_caffe_fpn_1x_coco-7aa93cef.pth) |
+
+### Detection performance
+
+| Method | Proposal | Backbone | Style | Schedule | Mem (GB) | Train time (s/iter) | Inf time (fps) | box AP | Config | Download |
+| :----------: | :---------: | :------: | :---: | :------: | :------: | :-----------------: | :------------: | :----: | :----------------------------------------------------------: | :-------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Fast R-CNN | Cascade RPN | R-50-FPN | caffe | 1x | - | - | - | 39.9 | [config](./cascade-rpn_fast-rcnn_r50-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rpn/crpn_fast_rcnn_r50_caffe_fpn_1x_coco/crpn_fast_rcnn_r50_caffe_fpn_1x_coco-cb486e66.pth) |
+| Faster R-CNN | Cascade RPN | R-50-FPN | caffe | 1x | - | - | - | 40.4 | [config](./cascade-rpn_faster-rcnn_r50-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cascade_rpn/crpn_faster_rcnn_r50_caffe_fpn_1x_coco/crpn_faster_rcnn_r50_caffe_fpn_1x_coco-c8283cca.pth) |
+
+## Citation
+
+We provide the code for reproducing experiment results of [Cascade RPN](https://arxiv.org/abs/1909.06720).
+
+```latex
+@inproceedings{vu2019cascade,
+ title={Cascade RPN: Delving into High-Quality Region Proposal Network with Adaptive Convolution},
+ author={Vu, Thang and Jang, Hyunjun and Pham, Trung X and Yoo, Chang D},
+ booktitle={Conference on Neural Information Processing Systems (NeurIPS)},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rpn/cascade-rpn_fast-rcnn_r50-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rpn/cascade-rpn_fast-rcnn_r50-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ba23ce90652d2ab2e9362be9a6231742d1815a70
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rpn/cascade-rpn_fast-rcnn_r50-caffe_fpn_1x_coco.py
@@ -0,0 +1,27 @@
+_base_ = '../fast_rcnn/fast-rcnn_r50-caffe_fpn_1x_coco.py'
+model = dict(
+ roi_head=dict(
+ bbox_head=dict(
+ bbox_coder=dict(target_stds=[0.04, 0.04, 0.08, 0.08]),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.5),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rcnn=dict(
+ assigner=dict(
+ pos_iou_thr=0.65, neg_iou_thr=0.65, min_pos_iou=0.65),
+ sampler=dict(num=256))),
+ test_cfg=dict(rcnn=dict(score_thr=1e-3)))
+
+# MMEngine support the following two ways, users can choose
+# according to convenience
+# train_dataloader = dict(dataset=dict(proposal_file='proposals/crpn_r50_caffe_fpn_1x_train2017.pkl')) # noqa
+_base_.train_dataloader.dataset.proposal_file = 'proposals/crpn_r50_caffe_fpn_1x_train2017.pkl' # noqa
+
+# val_dataloader = dict(dataset=dict(proposal_file='proposals/crpn_r50_caffe_fpn_1x_val2017.pkl')) # noqa
+# test_dataloader = val_dataloader
+_base_.val_dataloader.dataset.proposal_file = 'proposals/crpn_r50_caffe_fpn_1x_val2017.pkl' # noqa
+test_dataloader = _base_.val_dataloader
+
+optim_wrapper = dict(clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rpn/cascade-rpn_faster-rcnn_r50-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rpn/cascade-rpn_faster-rcnn_r50-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..2f7eced00144fb8fff1f234210a2b3f3fe475c8f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rpn/cascade-rpn_faster-rcnn_r50-caffe_fpn_1x_coco.py
@@ -0,0 +1,89 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50-caffe_fpn_1x_coco.py'
+rpn_weight = 0.7
+model = dict(
+ rpn_head=dict(
+ _delete_=True,
+ type='CascadeRPNHead',
+ num_stages=2,
+ stages=[
+ dict(
+ type='StageCascadeRPNHead',
+ in_channels=256,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ scales=[8],
+ ratios=[1.0],
+ strides=[4, 8, 16, 32, 64]),
+ adapt_cfg=dict(type='dilation', dilation=3),
+ bridged_feature=True,
+ with_cls=False,
+ reg_decoded_bbox=True,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=(.0, .0, .0, .0),
+ target_stds=(0.1, 0.1, 0.5, 0.5)),
+ loss_bbox=dict(
+ type='IoULoss', linear=True,
+ loss_weight=10.0 * rpn_weight)),
+ dict(
+ type='StageCascadeRPNHead',
+ in_channels=256,
+ feat_channels=256,
+ adapt_cfg=dict(type='offset'),
+ bridged_feature=False,
+ with_cls=True,
+ reg_decoded_bbox=True,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=(.0, .0, .0, .0),
+ target_stds=(0.05, 0.05, 0.1, 0.1)),
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ loss_weight=1.0 * rpn_weight),
+ loss_bbox=dict(
+ type='IoULoss', linear=True,
+ loss_weight=10.0 * rpn_weight))
+ ]),
+ roi_head=dict(
+ bbox_head=dict(
+ bbox_coder=dict(target_stds=[0.04, 0.04, 0.08, 0.08]),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.5),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=[
+ dict(
+ assigner=dict(
+ type='RegionAssigner', center_ratio=0.2, ignore_ratio=0.5),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.7,
+ min_pos_iou=0.3,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False)
+ ],
+ rpn_proposal=dict(max_per_img=300, nms=dict(iou_threshold=0.8)),
+ rcnn=dict(
+ assigner=dict(
+ pos_iou_thr=0.65, neg_iou_thr=0.65, min_pos_iou=0.65),
+ sampler=dict(type='RandomSampler', num=256))),
+ test_cfg=dict(
+ rpn=dict(max_per_img=300, nms=dict(iou_threshold=0.8)),
+ rcnn=dict(score_thr=1e-3)))
+optim_wrapper = dict(clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rpn/cascade-rpn_r50-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rpn/cascade-rpn_r50-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6eba24d11368ee0cdaae4fa316020ea3750be7f0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rpn/cascade-rpn_r50-caffe_fpn_1x_coco.py
@@ -0,0 +1,76 @@
+_base_ = '../rpn/rpn_r50-caffe_fpn_1x_coco.py'
+model = dict(
+ rpn_head=dict(
+ _delete_=True,
+ type='CascadeRPNHead',
+ num_stages=2,
+ stages=[
+ dict(
+ type='StageCascadeRPNHead',
+ in_channels=256,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ scales=[8],
+ ratios=[1.0],
+ strides=[4, 8, 16, 32, 64]),
+ adapt_cfg=dict(type='dilation', dilation=3),
+ bridged_feature=True,
+ sampling=False,
+ with_cls=False,
+ reg_decoded_bbox=True,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=(.0, .0, .0, .0),
+ target_stds=(0.1, 0.1, 0.5, 0.5)),
+ loss_bbox=dict(type='IoULoss', linear=True, loss_weight=10.0)),
+ dict(
+ type='StageCascadeRPNHead',
+ in_channels=256,
+ feat_channels=256,
+ adapt_cfg=dict(type='offset'),
+ bridged_feature=False,
+ sampling=True,
+ with_cls=True,
+ reg_decoded_bbox=True,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=(.0, .0, .0, .0),
+ target_stds=(0.05, 0.05, 0.1, 0.1)),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True,
+ loss_weight=1.0),
+ loss_bbox=dict(type='IoULoss', linear=True, loss_weight=10.0))
+ ]),
+ train_cfg=dict(rpn=[
+ dict(
+ assigner=dict(
+ type='RegionAssigner', center_ratio=0.2, ignore_ratio=0.5),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.7,
+ min_pos_iou=0.3,
+ ignore_iof_thr=-1,
+ iou_calculator=dict(type='BboxOverlaps2D')),
+ sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False)
+ ]),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=2000,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.8),
+ min_bbox_size=0)))
+optim_wrapper = dict(clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rpn/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rpn/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..62a88c5d2185ffd3aa7884f7a8c7d68cc3d60c8f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cascade_rpn/metafile.yml
@@ -0,0 +1,44 @@
+Collections:
+ - Name: Cascade RPN
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Cascade RPN
+ - FPN
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/1909.06720
+ Title: 'Cascade RPN: Delving into High-Quality Region Proposal Network with Adaptive Convolution'
+ README: configs/cascade_rpn/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.8.0/mmdet/models/dense_heads/cascade_rpn_head.py#L538
+ Version: v2.8.0
+
+Models:
+ - Name: cascade-rpn_fast-rcnn_r50-caffe_fpn_1x_coco
+ In Collection: Cascade RPN
+ Config: configs/cascade_rpn/cascade-rpn_fast-rcnn_r50-caffe_fpn_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rpn/crpn_fast_rcnn_r50_caffe_fpn_1x_coco/crpn_fast_rcnn_r50_caffe_fpn_1x_coco-cb486e66.pth
+
+ - Name: cascade-rpn_faster-rcnn_r50-caffe_fpn_1x_coco
+ In Collection: Cascade RPN
+ Config: configs/cascade_rpn/cascade-rpn_faster-rcnn_r50-caffe_fpn_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cascade_rpn/crpn_faster_rcnn_r50_caffe_fpn_1x_coco/crpn_faster_rcnn_r50_caffe_fpn_1x_coco-c8283cca.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..81e229c62f7816f20459a53132bfca676c01ac78
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/README.md
@@ -0,0 +1,58 @@
+# CenterNet
+
+> [Objects as Points](https://arxiv.org/abs/1904.07850)
+
+
+
+## Abstract
+
+Detection identifies objects as axis-aligned boxes in an image. Most successful object detectors enumerate a nearly exhaustive list of potential object locations and classify each. This is wasteful, inefficient, and requires additional post-processing. In this paper, we take a different approach. We model an object as a single point --- the center point of its bounding box. Our detector uses keypoint estimation to find center points and regresses to all other object properties, such as size, 3D location, orientation, and even pose. Our center point based approach, CenterNet, is end-to-end differentiable, simpler, faster, and more accurate than corresponding bounding box based detectors. CenterNet achieves the best speed-accuracy trade-off on the MS COCO dataset, with 28.1% AP at 142 FPS, 37.4% AP at 52 FPS, and 45.1% AP with multi-scale testing at 1.4 FPS. We use the same approach to estimate 3D bounding box in the KITTI benchmark and human pose on the COCO keypoint dataset. Our method performs competitively with sophisticated multi-stage methods and runs in real-time.
+
+
+

+
+
+## Results and Models
+
+| Backbone | DCN | Mem (GB) | Box AP | Flip box AP | Config | Download |
+| :-------: | :-: | :------: | :----: | :---------: | :--------------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| ResNet-18 | N | 3.45 | 25.9 | 27.3 | [config](./centernet_r18_8xb16-crop512-140e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/centernet/centernet_resnet18_140e_coco/centernet_resnet18_140e_coco_20210705_093630-bb5b3bf7.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/centernet/centernet_resnet18_140e_coco/centernet_resnet18_140e_coco_20210705_093630.log.json) |
+| ResNet-18 | Y | 3.47 | 29.5 | 30.9 | [config](./centernet_r18-dcnv2_8xb16-crop512-140e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/centernet/centernet_resnet18_dcnv2_140e_coco/centernet_resnet18_dcnv2_140e_coco_20210702_155131-c8cd631f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/centernet/centernet_resnet18_dcnv2_140e_coco/centernet_resnet18_dcnv2_140e_coco_20210702_155131.log.json) |
+
+Note:
+
+- Flip box AP setting is single-scale and `flip=True`.
+- Due to complex data enhancement, we find that the performance is unstable and may fluctuate by about 0.4 mAP. mAP 29.4 ~ 29.8 is acceptable in ResNet-18-DCNv2.
+- Compared to the source code, we refer to [CenterNet-Better](https://github.com/FateScript/CenterNet-better), and make the following changes
+ - fix wrong image mean and variance in image normalization to be compatible with the pre-trained backbone.
+ - Use SGD rather than ADAM optimizer and add warmup and grad clip.
+ - Use DistributedDataParallel as other models in MMDetection rather than using DataParallel.
+
+## CenterNet Update
+
+| Backbone | Style | Lr schd | MS train | Mem (GB) | Box AP | Config | Download |
+| :-------: | :---: | :-----: | :------: | :------: | :----: | :------------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| ResNet-50 | caffe | 1x | True | 3.3 | 40.2 | [config](./centernet-update_r50-caffe_fpn_ms-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/centernet/centernet-update_r50-caffe_fpn_ms-1x_coco/centernet-update_r50-caffe_fpn_ms-1x_coco_20230512_203845-8306baf2.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/centernet/centernet-update_r50-caffe_fpn_ms-1x_coco/centernet-update_r50-caffe_fpn_ms-1x_coco_20230512_203845.log.json) |
+
+CenterNet Update from the paper of [Probabilistic two-stage detection](https://arxiv.org/abs/2103.07461). The author has updated CenterNet to greatly improve performance and convergence speed.
+The [Details](https://github.com/xingyizhou/CenterNet2/blob/master/docs/MODEL_ZOO.md) are as follows:
+
+- Using top-left-right-bottom box encoding and GIoU Loss
+- Adding regression loss to the center 3x3 region
+- Adding more positive pixels for the heatmap loss whose regression loss is small and is within the center3x3 region
+- Using RetinaNet-style optimizer (SGD), learning rate rule (0.01 for each batch size 16), and schedule (12 epochs)
+- Added FPN neck layers, and assigns objects to FPN levels based on a fixed size range.
+- Using standard NMS instead of max pooling
+
+Note: We found that the performance of the r50 model fluctuates greatly and sometimes it does not converge. If the model does not converge, you can try running it again or reduce the learning rate.
+
+## Citation
+
+```latex
+@article{zhou2019objects,
+ title={Objects as Points},
+ author={Zhou, Xingyi and Wang, Dequan and Kr{\"a}henb{\"u}hl, Philipp},
+ booktitle={arXiv preprint arXiv:1904.07850},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet-update_r101_fpn_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet-update_r101_fpn_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..4fc65e0f8aeb1f02a0bea675146ced7a56800251
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet-update_r101_fpn_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './centernet-update_r50_fpn_8xb8-amp-lsj-200e_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet-update_r18_fpn_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet-update_r18_fpn_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ab3ae32ecd54cd08664e883a0888ef43040528d1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet-update_r18_fpn_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './centernet-update_r50_fpn_8xb8-amp-lsj-200e_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=18,
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet18')),
+ neck=dict(in_channels=[64, 128, 256, 512]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet-update_r50-caffe_fpn_ms-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet-update_r50-caffe_fpn_ms-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1f6e2b3919d6d2197c0ae9e1d721dc4eab00cf9c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet-update_r50-caffe_fpn_ms-1x_coco.py
@@ -0,0 +1,105 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ type='CenterNet',
+ # use caffe img_norm
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output',
+ num_outs=5,
+ # There is a chance to get 40.3 after switching init_cfg,
+ # otherwise it is about 39.9~40.1
+ init_cfg=dict(type='Caffe2Xavier', layer='Conv2d'),
+ relu_before_extra_convs=True),
+ bbox_head=dict(
+ type='CenterNetUpdateHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ strides=[8, 16, 32, 64, 128],
+ hm_min_radius=4,
+ hm_min_overlap=0.8,
+ more_pos_thresh=0.2,
+ more_pos_topk=9,
+ soft_weight_on_reg=False,
+ loss_cls=dict(
+ type='GaussianFocalLoss',
+ pos_weight=0.25,
+ neg_weight=0.75,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=2.0),
+ ),
+ train_cfg=None,
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+
+# single-scale training is about 39.3
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=0.00025,
+ by_epoch=False,
+ begin=0,
+ end=4000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+
+optim_wrapper = dict(
+ optimizer=dict(lr=0.01),
+ # Experiments show that there is no need to turn on clip_grad.
+ paramwise_cfg=dict(norm_decay_mult=0.))
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet-update_r50_fpn_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet-update_r50_fpn_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..34e0c680d39486467464f0ea7d6e1e08bf0c5240
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet-update_r50_fpn_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,83 @@
+_base_ = '../common/lsj-200e_coco-detection.py'
+
+image_size = (1024, 1024)
+batch_augments = [dict(type='BatchFixedSizePad', size=image_size)]
+
+model = dict(
+ type='CenterNet',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32,
+ batch_augments=batch_augments),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output',
+ num_outs=5,
+ init_cfg=dict(type='Caffe2Xavier', layer='Conv2d'),
+ relu_before_extra_convs=True),
+ bbox_head=dict(
+ type='CenterNetUpdateHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ strides=[8, 16, 32, 64, 128],
+ loss_cls=dict(
+ type='GaussianFocalLoss',
+ pos_weight=0.25,
+ neg_weight=0.75,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=2.0),
+ ),
+ train_cfg=None,
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+
+train_dataloader = dict(batch_size=8, num_workers=4)
+# Enable automatic-mixed-precision training with AmpOptimWrapper.
+optim_wrapper = dict(
+ type='AmpOptimWrapper',
+ optimizer=dict(
+ type='SGD', lr=0.01 * 4, momentum=0.9, weight_decay=0.00004),
+ paramwise_cfg=dict(norm_decay_mult=0.))
+
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=0.00025,
+ by_epoch=False,
+ begin=0,
+ end=4000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=25,
+ by_epoch=True,
+ milestones=[22, 24],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet_r18-dcnv2_8xb16-crop512-140e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet_r18-dcnv2_8xb16-crop512-140e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..732a55d59ad7dee175d8b72f798f0be044f23326
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet_r18-dcnv2_8xb16-crop512-140e_coco.py
@@ -0,0 +1,136 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py',
+ './centernet_tta.py'
+]
+
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+
+# model settings
+model = dict(
+ type='CenterNet',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True),
+ backbone=dict(
+ type='ResNet',
+ depth=18,
+ norm_eval=False,
+ norm_cfg=dict(type='BN'),
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet18')),
+ neck=dict(
+ type='CTResNetNeck',
+ in_channels=512,
+ num_deconv_filters=(256, 128, 64),
+ num_deconv_kernels=(4, 4, 4),
+ use_dcn=True),
+ bbox_head=dict(
+ type='CenterNetHead',
+ num_classes=80,
+ in_channels=64,
+ feat_channels=64,
+ loss_center_heatmap=dict(type='GaussianFocalLoss', loss_weight=1.0),
+ loss_wh=dict(type='L1Loss', loss_weight=0.1),
+ loss_offset=dict(type='L1Loss', loss_weight=1.0)),
+ train_cfg=None,
+ test_cfg=dict(topk=100, local_maximum_kernel=3, max_per_img=100))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PhotoMetricDistortion',
+ brightness_delta=32,
+ contrast_range=(0.5, 1.5),
+ saturation_range=(0.5, 1.5),
+ hue_delta=18),
+ dict(
+ type='RandomCenterCropPad',
+ # The cropped images are padded into squares during training,
+ # but may be less than crop_size.
+ crop_size=(512, 512),
+ ratios=(0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3),
+ mean=[0, 0, 0],
+ std=[1, 1, 1],
+ to_rgb=True,
+ test_pad_mode=None),
+ # Make sure the output is always crop_size.
+ dict(type='Resize', scale=(512, 512), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ backend_args={{_base_.backend_args}},
+ to_float32=True),
+ # don't need Resize
+ dict(
+ type='RandomCenterCropPad',
+ ratios=None,
+ border=None,
+ mean=[0, 0, 0],
+ std=[1, 1, 1],
+ to_rgb=True,
+ test_mode=True,
+ test_pad_mode=['logical_or', 31],
+ test_pad_add_pix=1),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape', 'border'))
+]
+
+# Use RepeatDataset to speed up training
+train_dataloader = dict(
+ batch_size=16,
+ num_workers=4,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ dataset=dict(
+ _delete_=True,
+ type='RepeatDataset',
+ times=5,
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args={{_base_.backend_args}},
+ )))
+
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# optimizer
+# Based on the default settings of modern detectors, the SGD effect is better
+# than the Adam in the source code, so we use SGD default settings and
+# if you use adam+lr5e-4, the map is 29.1.
+optim_wrapper = dict(clip_grad=dict(max_norm=35, norm_type=2))
+
+max_epochs = 28
+# learning policy
+# Based on the default settings of modern detectors, we added warmup settings.
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0,
+ end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[18, 24], # the real step is [18*5, 24*5]
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs) # the real epoch is 28*5=140
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (16 samples per GPU)
+auto_scale_lr = dict(base_batch_size=128)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet_r18_8xb16-crop512-140e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet_r18_8xb16-crop512-140e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6094b64221bd91eaafc9868e01c718d4421b418a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet_r18_8xb16-crop512-140e_coco.py
@@ -0,0 +1,3 @@
+_base_ = './centernet_r18-dcnv2_8xb16-crop512-140e_coco.py'
+
+model = dict(neck=dict(use_dcn=False))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet_tta.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet_tta.py
new file mode 100644
index 0000000000000000000000000000000000000000..edd7b03ecdeb272870919dcbd4842d6b8e32d8d4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/centernet_tta.py
@@ -0,0 +1,39 @@
+# This is different from the TTA of official CenterNet.
+
+tta_model = dict(
+ type='DetTTAModel',
+ tta_cfg=dict(nms=dict(type='nms', iou_threshold=0.5), max_per_img=100))
+
+tta_pipeline = [
+ dict(type='LoadImageFromFile', to_float32=True, backend_args=None),
+ dict(
+ type='TestTimeAug',
+ transforms=[
+ [
+ # ``RandomFlip`` must be placed before ``RandomCenterCropPad``,
+ # otherwise bounding box coordinates after flipping cannot be
+ # recovered correctly.
+ dict(type='RandomFlip', prob=1.),
+ dict(type='RandomFlip', prob=0.)
+ ],
+ [
+ dict(
+ type='RandomCenterCropPad',
+ ratios=None,
+ border=None,
+ mean=[0, 0, 0],
+ std=[1, 1, 1],
+ to_rgb=True,
+ test_mode=True,
+ test_pad_mode=['logical_or', 31],
+ test_pad_add_pix=1),
+ ],
+ [dict(type='LoadAnnotations', with_bbox=True)],
+ [
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'flip', 'flip_direction', 'border'))
+ ]
+ ])
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..496b8ea22df0ac1e757a40c2750893034e08a81c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centernet/metafile.yml
@@ -0,0 +1,60 @@
+Collections:
+ - Name: CenterNet
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x TITANXP GPUs
+ Architecture:
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/1904.07850
+ Title: 'Objects as Points'
+ README: configs/centernet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.13.0/mmdet/models/detectors/centernet.py#L10
+ Version: v2.13.0
+
+Models:
+ - Name: centernet_r18-dcnv2_8xb16-crop512-140e_coco
+ In Collection: CenterNet
+ Config: configs/centernet/centernet_r18-dcnv2_8xb16-crop512-140e_coco.py
+ Metadata:
+ Batch Size: 128
+ Training Memory (GB): 3.47
+ Epochs: 140
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 29.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/centernet/centernet_resnet18_dcnv2_140e_coco/centernet_resnet18_dcnv2_140e_coco_20210702_155131-c8cd631f.pth
+
+ - Name: centernet_r18_8xb16-crop512-140e_coco
+ In Collection: CenterNet
+ Config: configs/centernet/centernet_r18_8xb16-crop512-140e_coco.py
+ Metadata:
+ Batch Size: 128
+ Training Memory (GB): 3.45
+ Epochs: 140
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 25.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/centernet/centernet_resnet18_140e_coco/centernet_resnet18_140e_coco_20210705_093630-bb5b3bf7.pth
+
+ - Name: centernet-update_r50-caffe_fpn_ms-1x_coco
+ In Collection: CenterNet
+ Config: configs/centernet/centernet-update_r50-caffe_fpn_ms-1x_coco.py
+ Metadata:
+ Batch Size: 16
+ Training Memory (GB): 3.3
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.2
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/centernet/centernet-update_r50-caffe_fpn_ms-1x_coco/centernet-update_r50-caffe_fpn_ms-1x_coco_20230512_203845-8306baf2.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/centripetalnet/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centripetalnet/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..21edbd261af502d41fc6a24323bc28474a6d1c5a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centripetalnet/README.md
@@ -0,0 +1,36 @@
+# CentripetalNet
+
+> [CentripetalNet: Pursuing High-quality Keypoint Pairs for Object Detection](https://arxiv.org/abs/2003.09119)
+
+
+
+## Abstract
+
+Keypoint-based detectors have achieved pretty-well performance. However, incorrect keypoint matching is still widespread and greatly affects the performance of the detector. In this paper, we propose CentripetalNet which uses centripetal shift to pair corner keypoints from the same instance. CentripetalNet predicts the position and the centripetal shift of the corner points and matches corners whose shifted results are aligned. Combining position information, our approach matches corner points more accurately than the conventional embedding approaches do. Corner pooling extracts information inside the bounding boxes onto the border. To make this information more aware at the corners, we design a cross-star deformable convolution network to conduct feature adaption. Furthermore, we explore instance segmentation on anchor-free detectors by equipping our CentripetalNet with a mask prediction module. On MS-COCO test-dev, our CentripetalNet not only outperforms all existing anchor-free detectors with an AP of 48.0% but also achieves comparable performance to the state-of-the-art instance segmentation approaches with a 40.2% MaskAP.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Batch Size | Step/Total Epochs | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :--------------: | :-----------------------------------------------------------------------: | :---------------: | :------: | :------------: | :----: | :-----------------------------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| HourglassNet-104 | [16 x 6](./centripetalnet_hourglass104_16xb6-crop511-210e-mstest_coco.py) | 190/210 | 16.7 | 3.7 | 44.8 | [config](./centripetalnet_hourglass104_16xb6-crop511-210e-mstest_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/centripetalnet/centripetalnet_hourglass104_mstest_16x6_210e_coco/centripetalnet_hourglass104_mstest_16x6_210e_coco_20200915_204804-3ccc61e5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/centripetalnet/centripetalnet_hourglass104_mstest_16x6_210e_coco/centripetalnet_hourglass104_mstest_16x6_210e_coco_20200915_204804.log.json) |
+
+Note:
+
+- TTA setting is single-scale and `flip=True`. If you want to reproduce the TTA performance, please add `--tta` in the test command.
+- The model we released is the best checkpoint rather than the latest checkpoint (box AP 44.8 vs 44.6 in our experiment).
+
+## Citation
+
+```latex
+@InProceedings{Dong_2020_CVPR,
+author = {Dong, Zhiwei and Li, Guoxuan and Liao, Yue and Wang, Fei and Ren, Pengju and Qian, Chen},
+title = {CentripetalNet: Pursuing High-Quality Keypoint Pairs for Object Detection},
+booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
+month = {June},
+year = {2020}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/centripetalnet/centripetalnet_hourglass104_16xb6-crop511-210e-mstest_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centripetalnet/centripetalnet_hourglass104_16xb6-crop511-210e-mstest_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b757ffd16dca2d2b51d27ad413fdba889252c87f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centripetalnet/centripetalnet_hourglass104_16xb6-crop511-210e-mstest_coco.py
@@ -0,0 +1,181 @@
+_base_ = [
+ '../_base_/default_runtime.py', '../_base_/datasets/coco_detection.py'
+]
+
+data_preprocessor = dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True)
+
+# model settings
+model = dict(
+ type='CornerNet',
+ data_preprocessor=data_preprocessor,
+ backbone=dict(
+ type='HourglassNet',
+ downsample_times=5,
+ num_stacks=2,
+ stage_channels=[256, 256, 384, 384, 384, 512],
+ stage_blocks=[2, 2, 2, 2, 2, 4],
+ norm_cfg=dict(type='BN', requires_grad=True)),
+ neck=None,
+ bbox_head=dict(
+ type='CentripetalHead',
+ num_classes=80,
+ in_channels=256,
+ num_feat_levels=2,
+ corner_emb_channels=0,
+ loss_heatmap=dict(
+ type='GaussianFocalLoss', alpha=2.0, gamma=4.0, loss_weight=1),
+ loss_offset=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1),
+ loss_guiding_shift=dict(
+ type='SmoothL1Loss', beta=1.0, loss_weight=0.05),
+ loss_centripetal_shift=dict(
+ type='SmoothL1Loss', beta=1.0, loss_weight=1)),
+ # training and testing settings
+ train_cfg=None,
+ test_cfg=dict(
+ corner_topk=100,
+ local_maximum_kernel=3,
+ distance_threshold=0.5,
+ score_thr=0.05,
+ max_per_img=100,
+ nms=dict(type='soft_nms', iou_threshold=0.5, method='gaussian')))
+
+# data settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PhotoMetricDistortion',
+ brightness_delta=32,
+ contrast_range=(0.5, 1.5),
+ saturation_range=(0.5, 1.5),
+ hue_delta=18),
+ dict(
+ # The cropped images are padded into squares during training,
+ # but may be smaller than crop_size.
+ type='RandomCenterCropPad',
+ crop_size=(511, 511),
+ ratios=(0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3),
+ test_mode=False,
+ test_pad_mode=None,
+ mean=data_preprocessor['mean'],
+ std=data_preprocessor['std'],
+ # Image data is not converted to rgb.
+ to_rgb=data_preprocessor['bgr_to_rgb']),
+ dict(type='Resize', scale=(511, 511), keep_ratio=False),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs'),
+]
+
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ to_float32=True,
+ backend_args=_base_.backend_args),
+ # don't need Resize
+ dict(
+ type='RandomCenterCropPad',
+ crop_size=None,
+ ratios=None,
+ border=None,
+ test_mode=True,
+ test_pad_mode=['logical_or', 127],
+ mean=data_preprocessor['mean'],
+ std=data_preprocessor['std'],
+ # Image data is not converted to rgb.
+ to_rgb=data_preprocessor['bgr_to_rgb']),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape', 'border'))
+]
+
+train_dataloader = dict(
+ batch_size=6,
+ num_workers=3,
+ batch_sampler=None,
+ dataset=dict(pipeline=train_pipeline))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='Adam', lr=0.0005),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+max_epochs = 210
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 3,
+ by_epoch=False,
+ begin=0,
+ end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[190],
+ gamma=0.1)
+]
+
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (16 GPUs) x (6 samples per GPU)
+auto_scale_lr = dict(base_batch_size=96)
+
+tta_model = dict(
+ type='DetTTAModel',
+ tta_cfg=dict(
+ nms=dict(type='soft_nms', iou_threshold=0.5, method='gaussian'),
+ max_per_img=100))
+
+tta_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ to_float32=True,
+ backend_args=_base_.backend_args),
+ dict(
+ type='TestTimeAug',
+ transforms=[
+ [
+ # ``RandomFlip`` must be placed before ``RandomCenterCropPad``,
+ # otherwise bounding box coordinates after flipping cannot be
+ # recovered correctly.
+ dict(type='RandomFlip', prob=1.),
+ dict(type='RandomFlip', prob=0.)
+ ],
+ [
+ dict(
+ type='RandomCenterCropPad',
+ crop_size=None,
+ ratios=None,
+ border=None,
+ test_mode=True,
+ test_pad_mode=['logical_or', 127],
+ mean=data_preprocessor['mean'],
+ std=data_preprocessor['std'],
+ # Image data is not converted to rgb.
+ to_rgb=data_preprocessor['bgr_to_rgb'])
+ ],
+ [dict(type='LoadAnnotations', with_bbox=True)],
+ [
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'flip', 'flip_direction', 'border'))
+ ]
+ ])
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/centripetalnet/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centripetalnet/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..526572dfed0d158b55205c23031b5dfdbdfa9dc0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/centripetalnet/metafile.yml
@@ -0,0 +1,39 @@
+Collections:
+ - Name: CentripetalNet
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - Adam
+ Training Resources: 16x V100 GPUs
+ Architecture:
+ - Corner Pooling
+ - Stacked Hourglass Network
+ Paper:
+ URL: https://arxiv.org/abs/2003.09119
+ Title: 'CentripetalNet: Pursuing High-quality Keypoint Pairs for Object Detection'
+ README: configs/centripetalnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.5.0/mmdet/models/detectors/cornernet.py#L9
+ Version: v2.5.0
+
+Models:
+ - Name: centripetalnet_hourglass104_16xb6-crop511-210e-mstest_coco
+ In Collection: CentripetalNet
+ Config: configs/centripetalnet/centripetalnet_hourglass104_16xb6-crop511-210e-mstest_coco.py
+ Metadata:
+ Batch Size: 96
+ Training Memory (GB): 16.7
+ inference time (ms/im):
+ - value: 270.27
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 210
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/centripetalnet/centripetalnet_hourglass104_mstest_16x6_210e_coco/centripetalnet_hourglass104_mstest_16x6_210e_coco_20200915_204804-3ccc61e5.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cityscapes/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cityscapes/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..9e37b64edb7eded69bafa37244aa5a411e475d2c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cityscapes/README.md
@@ -0,0 +1,46 @@
+# Cityscapes
+
+> [The Cityscapes Dataset for Semantic Urban Scene Understanding](https://arxiv.org/abs/1604.01685)
+
+
+
+## Abstract
+
+Visual understanding of complex urban street scenes is an enabling factor for a wide range of applications. Object detection has benefited enormously from large-scale datasets, especially in the context of deep learning. For semantic urban scene understanding, however, no current dataset adequately captures the complexity of real-world urban scenes.
+To address this, we introduce Cityscapes, a benchmark suite and large-scale dataset to train and test approaches for pixel-level and instance-level semantic labeling. Cityscapes is comprised of a large, diverse set of stereo video sequences recorded in streets from 50 different cities. 5000 of these images have high quality pixel-level annotations; 20000 additional images have coarse annotations to enable methods that leverage large volumes of weakly-labeled data. Crucially, our effort exceeds previous attempts in terms of dataset size, annotation richness, scene variability, and complexity. Our accompanying empirical study provides an in-depth analysis of the dataset characteristics, as well as a performance evaluation of several state-of-the-art approaches based on our benchmark.
+
+
+

+
+
+## Common settings
+
+- All baselines were trained using 8 GPU with a batch size of 8 (1 images per GPU) using the [linear scaling rule](https://arxiv.org/abs/1706.02677) to scale the learning rate.
+- All models were trained on `cityscapes_train`, and tested on `cityscapes_val`.
+- 1x training schedule indicates 64 epochs which corresponds to slightly less than the 24k iterations reported in the original schedule from the [Mask R-CNN paper](https://arxiv.org/abs/1703.06870)
+- COCO pre-trained weights are used to initialize.
+- A conversion [script](../../tools/dataset_converters/cityscapes.py) is provided to convert Cityscapes into COCO format. Please refer to [install.md](../../docs/1_exist_data_model.md#prepare-datasets) for details.
+- `CityscapesDataset` implemented three evaluation methods. `bbox` and `segm` are standard COCO bbox/mask AP. `cityscapes` is the cityscapes dataset official evaluation, which may be slightly higher than COCO.
+
+### Faster R-CNN
+
+| Backbone | Style | Lr schd | Scale | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :------: | :-----: | :-----: | :------: | :------: | :------------: | :----: | :----------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | pytorch | 1x | 800-1024 | 5.2 | - | 40.3 | [config](./faster-rcnn_r50_fpn_1x_cityscapes.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cityscapes/faster_rcnn_r50_fpn_1x_cityscapes_20200502-829424c0.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cityscapes/faster_rcnn_r50_fpn_1x_cityscapes_20200502_114915.log.json) |
+
+### Mask R-CNN
+
+| Backbone | Style | Lr schd | Scale | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :------: | :-----: | :-----: | :------: | :------: | :------------: | :----: | :-----: | :--------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | pytorch | 1x | 800-1024 | 5.3 | - | 40.9 | 36.4 | [config](./mask-rcnn_r50_fpn_1x_cityscapes.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cityscapes/mask_rcnn_r50_fpn_1x_cityscapes/mask_rcnn_r50_fpn_1x_cityscapes_20201211_133733-d2858245.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cityscapes/mask_rcnn_r50_fpn_1x_cityscapes/mask_rcnn_r50_fpn_1x_cityscapes_20201211_133733.log.json) |
+
+## Citation
+
+```latex
+@inproceedings{Cordts2016Cityscapes,
+ title={The Cityscapes Dataset for Semantic Urban Scene Understanding},
+ author={Cordts, Marius and Omran, Mohamed and Ramos, Sebastian and Rehfeld, Timo and Enzweiler, Markus and Benenson, Rodrigo and Franke, Uwe and Roth, Stefan and Schiele, Bernt},
+ booktitle={Proc. of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)},
+ year={2016}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cityscapes/faster-rcnn_r50_fpn_1x_cityscapes.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cityscapes/faster-rcnn_r50_fpn_1x_cityscapes.py
new file mode 100644
index 0000000000000000000000000000000000000000..ccd0de2aff1c1f3071e70e67dbf94b1c1cfe7e8b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cityscapes/faster-rcnn_r50_fpn_1x_cityscapes.py
@@ -0,0 +1,41 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/cityscapes_detection.py',
+ '../_base_/default_runtime.py', '../_base_/schedules/schedule_1x.py'
+]
+model = dict(
+ backbone=dict(init_cfg=None),
+ roi_head=dict(
+ bbox_head=dict(
+ num_classes=8,
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0))))
+
+# optimizer
+# lr is set for a batch size of 8
+optim_wrapper = dict(optimizer=dict(lr=0.01))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=8,
+ by_epoch=True,
+ # [7] yields higher performance than [6]
+ milestones=[7],
+ gamma=0.1)
+]
+
+# actual epoch = 8 * 8 = 64
+train_cfg = dict(max_epochs=8)
+
+# For better, more stable performance initialize from COCO
+load_from = 'https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_1x_coco/faster_rcnn_r50_fpn_1x_coco_20200130-047c8118.pth' # noqa
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (1 samples per GPU)
+# TODO: support auto scaling lr
+# auto_scale_lr = dict(base_batch_size=8)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cityscapes/mask-rcnn_r50_fpn_1x_cityscapes.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cityscapes/mask-rcnn_r50_fpn_1x_cityscapes.py
new file mode 100644
index 0000000000000000000000000000000000000000..772268b121e7b8858c4cfcf3b6820e6146634d0d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cityscapes/mask-rcnn_r50_fpn_1x_cityscapes.py
@@ -0,0 +1,43 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/cityscapes_instance.py',
+ '../_base_/default_runtime.py', '../_base_/schedules/schedule_1x.py'
+]
+model = dict(
+ backbone=dict(init_cfg=None),
+ roi_head=dict(
+ bbox_head=dict(
+ type='Shared2FCBBoxHead',
+ num_classes=8,
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0)),
+ mask_head=dict(num_classes=8)))
+
+# optimizer
+# lr is set for a batch size of 8
+optim_wrapper = dict(optimizer=dict(lr=0.01))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=8,
+ by_epoch=True,
+ # [7] yields higher performance than [6]
+ milestones=[7],
+ gamma=0.1)
+]
+
+# actual epoch = 8 * 8 = 64
+train_cfg = dict(max_epochs=8)
+
+# For better, more stable performance initialize from COCO
+load_from = 'https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_fpn_1x_coco/mask_rcnn_r50_fpn_1x_coco_20200205-d4b0c5d6.pth' # noqa
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (1 samples per GPU)
+# TODO: support auto scaling lr
+# auto_scale_lr = dict(base_batch_size=8)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/lsj-100e_coco-detection.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/lsj-100e_coco-detection.py
new file mode 100644
index 0000000000000000000000000000000000000000..bb631e5d5c1253cc3a5d81a8cdc6cd86133d9b53
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/lsj-100e_coco-detection.py
@@ -0,0 +1,122 @@
+_base_ = '../_base_/default_runtime.py'
+# dataset settings
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+image_size = (1024, 1024)
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize',
+ scale=image_size,
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size,
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+# Use RepeatDataset to speed up training
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ dataset=dict(
+ type='RepeatDataset',
+ times=4, # simply change this from 2 to 16 for 50e - 400e training.
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args)))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric='bbox',
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+max_epochs = 25
+
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=5)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# optimizer assumes bs=64
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.1, momentum=0.9, weight_decay=0.00004))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.067, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[22, 24],
+ gamma=0.1)
+]
+
+# only keep latest 2 checkpoints
+default_hooks = dict(checkpoint=dict(max_keep_ckpts=2))
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (32 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/lsj-100e_coco-instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/lsj-100e_coco-instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..6e62729d639c7659115a7f5f6449fa9021338be6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/lsj-100e_coco-instance.py
@@ -0,0 +1,122 @@
+_base_ = '../_base_/default_runtime.py'
+# dataset settings
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+image_size = (1024, 1024)
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomResize',
+ scale=image_size,
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size,
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+# Use RepeatDataset to speed up training
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ dataset=dict(
+ type='RepeatDataset',
+ times=4, # simply change this from 2 to 16 for 50e - 400e training.
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args)))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric=['bbox', 'segm'],
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+max_epochs = 25
+
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=5)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# optimizer assumes bs=64
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.1, momentum=0.9, weight_decay=0.00004))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.067, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[22, 24],
+ gamma=0.1)
+]
+
+# only keep latest 2 checkpoints
+default_hooks = dict(checkpoint=dict(max_keep_ckpts=2))
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (32 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/lsj-200e_coco-detection.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/lsj-200e_coco-detection.py
new file mode 100644
index 0000000000000000000000000000000000000000..83d12947fed900f05d748b6f90ef29cc5fbc407a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/lsj-200e_coco-detection.py
@@ -0,0 +1,18 @@
+_base_ = './lsj-100e_coco-detection.py'
+
+# 8x25=200e
+train_dataloader = dict(dataset=dict(times=8))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.067, by_epoch=False, begin=0,
+ end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=25,
+ by_epoch=True,
+ milestones=[22, 24],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/lsj-200e_coco-instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/lsj-200e_coco-instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..af3e4bf160c01045c6e36d67bdee796e7bf96cd3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/lsj-200e_coco-instance.py
@@ -0,0 +1,18 @@
+_base_ = './lsj-100e_coco-instance.py'
+
+# 8x25=200e
+train_dataloader = dict(dataset=dict(times=8))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.067, by_epoch=False, begin=0,
+ end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=25,
+ by_epoch=True,
+ milestones=[22, 24],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ms-90k_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ms-90k_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e2d6c3dafb61d59bbbe9d0c6188a1bbff3b736b3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ms-90k_coco.py
@@ -0,0 +1,122 @@
+_base_ = '../_base_/default_runtime.py'
+
+# dataset settings
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+# Align with Detectron2
+backend = 'pillow'
+train_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ backend_args=backend_args,
+ imdecode_backend=backend),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True,
+ backend=backend),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ backend_args=backend_args,
+ imdecode_backend=backend),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True, backend=backend),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ pin_memory=True,
+ sampler=dict(type='InfiniteSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ pin_memory=True,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric='bbox',
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+# training schedule for 90k
+max_iter = 90000
+train_cfg = dict(
+ type='IterBasedTrainLoop', max_iters=max_iter, val_interval=10000)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0,
+ end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_iter,
+ by_epoch=False,
+ milestones=[60000, 80000],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.0001))
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
+
+default_hooks = dict(checkpoint=dict(by_epoch=False, interval=10000))
+log_processor = dict(by_epoch=False)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ms-poly-90k_coco-instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ms-poly-90k_coco-instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..d5566b3c3b8bfe0a49c8c062fb0fc972d5ae1f55
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ms-poly-90k_coco-instance.py
@@ -0,0 +1,130 @@
+_base_ = '../_base_/default_runtime.py'
+
+# dataset settings
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+# Align with Detectron2
+backend = 'pillow'
+train_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ backend_args=backend_args,
+ imdecode_backend=backend),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ poly2mask=False),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True,
+ backend=backend),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ backend_args=backend_args,
+ imdecode_backend=backend),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True, backend=backend),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ poly2mask=False),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ pin_memory=True,
+ sampler=dict(type='InfiniteSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ pin_memory=True,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric=['bbox', 'segm'],
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+# training schedule for 90k
+max_iter = 90000
+train_cfg = dict(
+ type='IterBasedTrainLoop', max_iters=max_iter, val_interval=10000)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0,
+ end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_iter,
+ by_epoch=False,
+ milestones=[60000, 80000],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.0001))
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
+
+default_hooks = dict(checkpoint=dict(by_epoch=False, interval=10000))
+log_processor = dict(by_epoch=False)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ms-poly_3x_coco-instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ms-poly_3x_coco-instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..04072f9b84c06d546767649f7e17736444db7ce2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ms-poly_3x_coco-instance.py
@@ -0,0 +1,118 @@
+_base_ = '../_base_/default_runtime.py'
+# dataset settings
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+# In mstrain 3x config, img_scale=[(1333, 640), (1333, 800)],
+# multiscale_mode='range'
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ poly2mask=False),
+ dict(
+ type='RandomResize', scale=[(1333, 640), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs'),
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ poly2mask=False),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type='RepeatDataset',
+ times=3,
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args)))
+val_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric=['bbox', 'segm'],
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+# training schedule for 3x with `RepeatDataset`
+train_cfg = dict(type='EpochBasedTrainLoop', max_epochs=12, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# learning rate
+# Experiments show that using milestones=[9, 11] has higher performance
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[9, 11],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.0001))
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ms_3x_coco-instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ms_3x_coco-instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..f80cf88e9b1e770dce3157abc852aea996eec624
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ms_3x_coco-instance.py
@@ -0,0 +1,108 @@
+_base_ = '../_base_/default_runtime.py'
+
+# dataset settings
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomResize', scale=[(1333, 640), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type='RepeatDataset',
+ times=3,
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args)))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric='bbox',
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+# training schedule for 3x with `RepeatDataset`
+train_cfg = dict(type='EpochBasedTrainLoop', max_epochs=12, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# learning rate
+# Experiments show that using milestones=[9, 11] has higher performance
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[9, 11],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.0001))
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ms_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ms_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..facbb34cf05088d8832502d3c9a38d812d328308
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ms_3x_coco.py
@@ -0,0 +1,108 @@
+_base_ = '../_base_/default_runtime.py'
+
+# dataset settings
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize', scale=[(1333, 640), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type='RepeatDataset',
+ times=3,
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args)))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric='bbox',
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+# training schedule for 3x with `RepeatDataset`
+train_cfg = dict(type='EpochBasedTrainLoop', max_epochs=12, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# learning rate
+# Experiments show that using milestones=[9, 11] has higher performance
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[9, 11],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.0001))
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ssj_270k_coco-instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ssj_270k_coco-instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..7407644fd59bb03d6e0afde83f8893a351ddc356
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ssj_270k_coco-instance.py
@@ -0,0 +1,125 @@
+_base_ = '../_base_/default_runtime.py'
+# dataset settings
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+
+image_size = (1024, 1024)
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+# Standard Scale Jittering (SSJ) resizes and crops an image
+# with a resize range of 0.8 to 1.25 of the original image size.
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomResize',
+ scale=image_size,
+ ratio_range=(0.8, 1.25),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size,
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='InfiniteSampler'),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric=['bbox', 'segm'],
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+# The model is trained by 270k iterations with batch_size 64,
+# which is roughly equivalent to 144 epochs.
+
+max_iters = 270000
+train_cfg = dict(
+ type='IterBasedTrainLoop', max_iters=max_iters, val_interval=10000)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# optimizer assumes bs=64
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.1, momentum=0.9, weight_decay=0.00004))
+
+# learning rate policy
+# lr steps at [0.9, 0.95, 0.975] of the maximum iterations
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0,
+ end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=270000,
+ by_epoch=False,
+ milestones=[243000, 256500, 263250],
+ gamma=0.1)
+]
+
+default_hooks = dict(checkpoint=dict(by_epoch=False, interval=10000))
+log_processor = dict(by_epoch=False)
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (32 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ssj_scp_270k_coco-instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ssj_scp_270k_coco-instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..06159dd40312ec935ac383701fa7b052b863e1bf
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/common/ssj_scp_270k_coco-instance.py
@@ -0,0 +1,60 @@
+_base_ = 'ssj_270k_coco-instance.py'
+# dataset settings
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+
+image_size = (1024, 1024)
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+# Standard Scale Jittering (SSJ) resizes and crops an image
+# with a resize range of 0.8 to 1.25 of the original image size.
+load_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomResize',
+ scale=image_size,
+ ratio_range=(0.8, 1.25),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size,
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size=image_size),
+]
+train_pipeline = [
+ dict(type='CopyPaste', max_num_pasted=100),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ _delete_=True,
+ type='MultiImageMixDataset',
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=load_pipeline,
+ backend_args=backend_args),
+ pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/condinst/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/condinst/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..01deb0ecff4e2a5526029aed31d4cf8a87c8545f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/condinst/README.md
@@ -0,0 +1,40 @@
+# CondInst
+
+> [CondInst: Conditional Convolutions for Instance
+> Segmentation](https://arxiv.org/pdf/2003.05664.pdf)
+
+
+
+## Abstract
+
+We propose a simple yet effective instance segmentation framework, termed CondInst (conditional convolutions for instance segmentation). Top-performing instance segmentation methods such as Mask
+R-CNN rely on ROI operations (typically ROIPool or ROIAlign) to
+obtain the final instance masks. In contrast, we propose to solve instance segmentation from a new perspective. Instead of using instancewise ROIs as inputs to a network of fixed weights, we employ dynamic
+instance-aware networks, conditioned on instances. CondInst enjoys two
+advantages: 1) Instance segmentation is solved by a fully convolutional
+network, eliminating the need for ROI cropping and feature alignment.
+2\) Due to the much improved capacity of dynamically-generated conditional convolutions, the mask head can be very compact (e.g., 3 conv.
+layers, each having only 8 channels), leading to significantly faster inference. We demonstrate a simpler instance segmentation method that can
+achieve improved performance in both accuracy and inference speed. On
+the COCO dataset, we outperform a few recent methods including welltuned Mask R-CNN baselines, without longer training schedules needed.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Style | MS train | Lr schd | bbox AP | mask AP | Config | Download |
+| :------: | :-----: | :------: | :-----: | :-----: | :-----: | :-------------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | pytorch | Y | 1x | 39.8 | 36.0 | [config](./condinst_r50_fpn_ms-poly-90k_coco_instance.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/condinst/condinst_r50_fpn_ms-poly-90k_coco_instance/condinst_r50_fpn_ms-poly-90k_coco_instance_20221129_125223-4c186406.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/condinst/condinst_r50_fpn_ms-poly-90k_coco_instance/condinst_r50_fpn_ms-poly-90k_coco_instance_20221129_125223.json) |
+
+## Citation
+
+```latex
+@inproceedings{tian2020conditional,
+ title = {Conditional Convolutions for Instance Segmentation},
+ author = {Tian, Zhi and Shen, Chunhua and Chen, Hao},
+ booktitle = {Proc. Eur. Conf. Computer Vision (ECCV)},
+ year = {2020}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/condinst/condinst_r50_fpn_ms-poly-90k_coco_instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/condinst/condinst_r50_fpn_ms-poly-90k_coco_instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..39639d874cbeb54b64a2789f251f1f6dad585ce3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/condinst/condinst_r50_fpn_ms-poly-90k_coco_instance.py
@@ -0,0 +1,85 @@
+_base_ = '../common/ms-poly-90k_coco-instance.py'
+
+# model settings
+model = dict(
+ type='CondInst',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_mask=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50'),
+ style='pytorch'),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output', # use P5
+ num_outs=5,
+ relu_before_extra_convs=True),
+ bbox_head=dict(
+ type='CondInstBboxHead',
+ num_params=169,
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ strides=[8, 16, 32, 64, 128],
+ norm_on_bbox=True,
+ centerness_on_reg=True,
+ dcn_on_last_conv=False,
+ center_sampling=True,
+ conv_bias=True,
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=1.0),
+ loss_centerness=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0)),
+ mask_head=dict(
+ type='CondInstMaskHead',
+ num_layers=3,
+ feat_channels=8,
+ size_of_interest=8,
+ mask_out_stride=4,
+ max_masks_to_train=300,
+ mask_feature_head=dict(
+ in_channels=256,
+ feat_channels=128,
+ start_level=0,
+ end_level=2,
+ out_channels=8,
+ mask_stride=8,
+ num_stacked_convs=4,
+ norm_cfg=dict(type='BN', requires_grad=True)),
+ loss_mask=dict(
+ type='DiceLoss',
+ use_sigmoid=True,
+ activate=True,
+ eps=5e-6,
+ loss_weight=1.0)),
+ # model training and testing settings
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100,
+ mask_thr=0.5))
+
+# optimizer
+optim_wrapper = dict(optimizer=dict(lr=0.01))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/condinst/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/condinst/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..1237b74d77a8b1f1e4b0ba74c6bdc5e5595d9816
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/condinst/metafile.yml
@@ -0,0 +1,32 @@
+Collections:
+ - Name: CondInst
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x A100 GPUs
+ Architecture:
+ - FPN
+ - FCOS
+ - ResNet
+ Paper: https://arxiv.org/abs/2003.05664
+ README: configs/condinst/README.md
+
+Models:
+ - Name: condinst_r50_fpn_ms-poly-90k_coco_instance
+ In Collection: CondInst
+ Config: configs/condinst/condinst_r50_fpn_ms-poly-90k_coco_instance.py
+ Metadata:
+ Training Memory (GB): 4.4
+ Iterations: 90000
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.0
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/condinst/condinst_r50_fpn_ms-poly-90k_coco_instance/condinst_r50_fpn_ms-poly-90k_coco_instance_20221129_125223-4c186406.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/conditional_detr/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/conditional_detr/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..4043571c576bba7f287e16e7e464950b5568543e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/conditional_detr/README.md
@@ -0,0 +1,39 @@
+# Conditional DETR
+
+> [Conditional DETR for Fast Training Convergence](https://arxiv.org/abs/2108.06152)
+
+
+
+## Abstract
+
+The DETR approach applies the transformer encoder and decoder architecture to object detection and achieves promising performance. In this paper, we handle the critical issue, slow training convergence, and present a conditional cross-attention mechanism for fast DETR training. Our approach is motivated by that the cross-attention in DETR relies highly on the content embeddings and that the spatial embeddings make minor contributions, increasing the need for high-quality content embeddings and thus increasing the training difficulty.
+
+
+

+
+
+Our conditional DETR learns a conditional spatial query from the decoder embedding for decoder multi-head cross-attention. The benefit is that through the conditional spatial query, each cross-attention head is able to attend to a band containing a distinct region, e.g., one object extremity or a region inside the object box (Figure 1). This narrows down the spatial range for localizing the distinct regions for object classification and box regression, thus relaxing the dependence on the content embeddings and easing the training. Empirical results show that conditional DETR converges 6.7x faster for the backbones R50 and R101 and 10x faster for stronger backbones DC5-R50 and DC5-R101.
+
+
+

+

+
+
+## Results and Models
+
+We provide the config files and models for Conditional DETR: [Conditional DETR for Fast Training Convergence](https://arxiv.org/abs/2108.06152).
+
+| Backbone | Model | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :------: | :--------------: | :-----: | :------: | :------------: | :----: | :-----------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | Conditional DETR | 50e | | | 41.1 | [config](./conditional-detr_r50_8xb2-50e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/conditional_detr/conditional-detr_r50_8xb2-50e_coco/conditional-detr_r50_8xb2-50e_coco_20221121_180202-c83a1dc0.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/conditional_detr/conditional-detr_r50_8xb2-50e_coco/conditional-detr_r50_8xb2-50e_coco_20221121_180202.log.json) |
+
+## Citation
+
+```latex
+@inproceedings{meng2021-CondDETR,
+ title = {Conditional DETR for Fast Training Convergence},
+ author = {Meng, Depu and Chen, Xiaokang and Fan, Zejia and Zeng, Gang and Li, Houqiang and Yuan, Yuhui and Sun, Lei and Wang, Jingdong},
+ booktitle = {Proceedings of the IEEE International Conference on Computer Vision (ICCV)},
+ year = {2021}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/conditional_detr/conditional-detr_r50_8xb2-50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/conditional_detr/conditional-detr_r50_8xb2-50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a21476448d0cbab6b6e4b94aa46d686e38667879
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/conditional_detr/conditional-detr_r50_8xb2-50e_coco.py
@@ -0,0 +1,42 @@
+_base_ = ['../detr/detr_r50_8xb2-150e_coco.py']
+model = dict(
+ type='ConditionalDETR',
+ num_queries=300,
+ decoder=dict(
+ num_layers=6,
+ layer_cfg=dict(
+ self_attn_cfg=dict(
+ _delete_=True,
+ embed_dims=256,
+ num_heads=8,
+ attn_drop=0.1,
+ cross_attn=False),
+ cross_attn_cfg=dict(
+ _delete_=True,
+ embed_dims=256,
+ num_heads=8,
+ attn_drop=0.1,
+ cross_attn=True))),
+ bbox_head=dict(
+ type='ConditionalDETRHead',
+ loss_cls=dict(
+ _delete_=True,
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=2.0)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='HungarianAssigner',
+ match_costs=[
+ dict(type='FocalLossCost', weight=2.0),
+ dict(type='BBoxL1Cost', weight=5.0, box_format='xywh'),
+ dict(type='IoUCost', iou_mode='giou', weight=2.0)
+ ])))
+
+# learning policy
+train_cfg = dict(type='EpochBasedTrainLoop', max_epochs=50, val_interval=1)
+
+param_scheduler = [dict(type='MultiStepLR', end=50, milestones=[40])]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/conditional_detr/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/conditional_detr/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..83f5532ce380c903d644b36055c4c2610455472a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/conditional_detr/metafile.yml
@@ -0,0 +1,32 @@
+Collections:
+ - Name: Conditional DETR
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - AdamW
+ - Multi Scale Train
+ - Gradient Clip
+ Training Resources: 8x A100 GPUs
+ Architecture:
+ - ResNet
+ - Transformer
+ Paper:
+ URL: https://arxiv.org/abs/2108.06152
+ Title: 'Conditional DETR for Fast Training Convergence'
+ README: configs/conditional_detr/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/f4112c9e5611468ffbd57cfba548fd1289264b52/mmdet/models/detectors/conditional_detr.py#L14
+ Version: v3.0.0rc6
+
+Models:
+ - Name: conditional-detr_r50_8xb2-50e_coco
+ In Collection: Conditional DETR
+ Config: configs/conditional_detr/conditional-detr_r50_8xb2-50e_coco.py
+ Metadata:
+ Epochs: 50
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.9
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/conditional_detr/conditional-detr_r50_8xb2-50e_coco/conditional-detr_r50_8xb2-50e_coco_20221121_180202-c83a1dc0.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/convnext/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/convnext/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..33497bb57aa9ae89b91ee16ac81e1ce02bf2ae0d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/convnext/README.md
@@ -0,0 +1,42 @@
+# ConvNeXt
+
+> [A ConvNet for the 2020s](https://arxiv.org/abs/2201.03545)
+
+
+
+## Abstract
+
+The "Roaring 20s" of visual recognition began with the introduction of Vision Transformers (ViTs), which quickly superseded ConvNets as the state-of-the-art image classification model. A vanilla ViT, on the other hand, faces difficulties when applied to general computer vision tasks such as object detection and semantic segmentation. It is the hierarchical Transformers (e.g., Swin Transformers) that reintroduced several ConvNet priors, making Transformers practically viable as a generic vision backbone and demonstrating remarkable performance on a wide variety of vision tasks. However, the effectiveness of such hybrid approaches is still largely credited to the intrinsic superiority of Transformers, rather than the inherent inductive biases of convolutions. In this work, we reexamine the design spaces and test the limits of what a pure ConvNet can achieve. We gradually "modernize" a standard ResNet toward the design of a vision Transformer, and discover several key components that contribute to the performance difference along the way. The outcome of this exploration is a family of pure ConvNet models dubbed ConvNeXt. Constructed entirely from standard ConvNet modules, ConvNeXts compete favorably with Transformers in terms of accuracy and scalability, achieving 87.8% ImageNet top-1 accuracy and outperforming Swin Transformers on COCO detection and ADE20K segmentation, while maintaining the simplicity and efficiency of standard ConvNets.
+
+
+

+
+
+## Results and models
+
+| Method | Backbone | Pretrain | Lr schd | Multi-scale crop | FP16 | Mem (GB) | box AP | mask AP | Config | Download |
+| :----------------: | :--------: | :---------: | :-----: | :--------------: | :--: | :------: | :----: | :-----: | :-------------------------------------------------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Mask R-CNN | ConvNeXt-T | ImageNet-1K | 3x | yes | yes | 7.3 | 46.2 | 41.7 | [config](./mask-rcnn_convnext-t-p4-w7_fpn_amp-ms-crop-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/convnext/mask_rcnn_convnext-t_p4_w7_fpn_fp16_ms-crop_3x_coco/mask_rcnn_convnext-t_p4_w7_fpn_fp16_ms-crop_3x_coco_20220426_154953-050731f4.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/convnext/mask_rcnn_convnext-t_p4_w7_fpn_fp16_ms-crop_3x_coco/mask_rcnn_convnext-t_p4_w7_fpn_fp16_ms-crop_3x_coco_20220426_154953.log.json) |
+| Cascade Mask R-CNN | ConvNeXt-T | ImageNet-1K | 3x | yes | yes | 9.0 | 50.3 | 43.6 | [config](./cascade-mask-rcnn_convnext-t-p4-w7_fpn_4conv1fc-giou_amp-ms-crop-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/convnext/cascade_mask_rcnn_convnext-t_p4_w7_fpn_giou_4conv1f_fp16_ms-crop_3x_coco/cascade_mask_rcnn_convnext-t_p4_w7_fpn_giou_4conv1f_fp16_ms-crop_3x_coco_20220509_204200-8f07c40b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/convnext/cascade_mask_rcnn_convnext-t_p4_w7_fpn_giou_4conv1f_fp16_ms-crop_3x_coco/cascade_mask_rcnn_convnext-t_p4_w7_fpn_giou_4conv1f_fp16_ms-crop_3x_coco_20220509_204200.log.json) |
+| Cascade Mask R-CNN | ConvNeXt-S | ImageNet-1K | 3x | yes | yes | 12.3 | 51.8 | 44.8 | [config](./cascade-mask-rcnn_convnext-s-p4-w7_fpn_4conv1fc-giou_amp-ms-crop-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/convnext/cascade_mask_rcnn_convnext-s_p4_w7_fpn_giou_4conv1f_fp16_ms-crop_3x_coco/cascade_mask_rcnn_convnext-s_p4_w7_fpn_giou_4conv1f_fp16_ms-crop_3x_coco_20220510_201004-3d24f5a4.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/convnext/cascade_mask_rcnn_convnext-s_p4_w7_fpn_giou_4conv1f_fp16_ms-crop_3x_coco/cascade_mask_rcnn_convnext-s_p4_w7_fpn_giou_4conv1f_fp16_ms-crop_3x_coco_20220510_201004.log.json) |
+
+**Note**:
+
+- ConvNeXt backbone needs to install [MMPreTrain](https://github.com/open-mmlab/mmpretrain) first, which has abundant backbones for downstream tasks.
+
+```shell
+pip install mmpretrain
+```
+
+- The performance is unstable. `Cascade Mask R-CNN` may fluctuate about 0.2 mAP.
+
+## Citation
+
+```bibtex
+@article{liu2022convnet,
+ title={A ConvNet for the 2020s},
+ author={Liu, Zhuang and Mao, Hanzi and Wu, Chao-Yuan and Feichtenhofer, Christoph and Darrell, Trevor and Xie, Saining},
+ journal={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
+ year={2022}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/convnext/cascade-mask-rcnn_convnext-s-p4-w7_fpn_4conv1fc-giou_amp-ms-crop-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/convnext/cascade-mask-rcnn_convnext-s-p4-w7_fpn_4conv1fc-giou_amp-ms-crop-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..9a5fbedcaa78636f11a5718f1123d33e7e2ac273
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/convnext/cascade-mask-rcnn_convnext-s-p4-w7_fpn_4conv1fc-giou_amp-ms-crop-3x_coco.py
@@ -0,0 +1,26 @@
+_base_ = './cascade-mask-rcnn_convnext-t-p4-w7_fpn_4conv1fc-giou_amp-ms-crop-3x_coco.py' # noqa
+
+# please install mmpretrain
+# import mmpretrain.models to trigger register_module in mmpretrain
+custom_imports = dict(
+ imports=['mmpretrain.models'], allow_failed_imports=False)
+checkpoint_file = 'https://download.openmmlab.com/mmclassification/v0/convnext/downstream/convnext-small_3rdparty_32xb128-noema_in1k_20220301-303e75e3.pth' # noqa
+
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='mmpretrain.ConvNeXt',
+ arch='small',
+ out_indices=[0, 1, 2, 3],
+ drop_path_rate=0.6,
+ layer_scale_init_value=1.0,
+ gap_before_final_norm=False,
+ init_cfg=dict(
+ type='Pretrained', checkpoint=checkpoint_file,
+ prefix='backbone.')))
+
+optim_wrapper = dict(paramwise_cfg={
+ 'decay_rate': 0.7,
+ 'decay_type': 'layer_wise',
+ 'num_layers': 12
+})
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/convnext/cascade-mask-rcnn_convnext-t-p4-w7_fpn_4conv1fc-giou_amp-ms-crop-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/convnext/cascade-mask-rcnn_convnext-t-p4-w7_fpn_4conv1fc-giou_amp-ms-crop-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c92f86838c31710dd550c36d9abc11d79bb6e2eb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/convnext/cascade-mask-rcnn_convnext-t-p4-w7_fpn_4conv1fc-giou_amp-ms-crop-3x_coco.py
@@ -0,0 +1,154 @@
+_base_ = [
+ '../_base_/models/cascade-mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+# please install mmpretrain
+# import mmpretrain.models to trigger register_module in mmpretrain
+custom_imports = dict(
+ imports=['mmpretrain.models'], allow_failed_imports=False)
+checkpoint_file = 'https://download.openmmlab.com/mmclassification/v0/convnext/downstream/convnext-tiny_3rdparty_32xb128-noema_in1k_20220301-795e9634.pth' # noqa
+
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='mmpretrain.ConvNeXt',
+ arch='tiny',
+ out_indices=[0, 1, 2, 3],
+ drop_path_rate=0.4,
+ layer_scale_init_value=1.0,
+ gap_before_final_norm=False,
+ init_cfg=dict(
+ type='Pretrained', checkpoint=checkpoint_file,
+ prefix='backbone.')),
+ neck=dict(in_channels=[96, 192, 384, 768]),
+ roi_head=dict(bbox_head=[
+ dict(
+ type='ConvFCBBoxHead',
+ num_shared_convs=4,
+ num_shared_fcs=1,
+ in_channels=256,
+ conv_out_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=False,
+ reg_decoded_bbox=True,
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=10.0)),
+ dict(
+ type='ConvFCBBoxHead',
+ num_shared_convs=4,
+ num_shared_fcs=1,
+ in_channels=256,
+ conv_out_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.05, 0.05, 0.1, 0.1]),
+ reg_class_agnostic=False,
+ reg_decoded_bbox=True,
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=10.0)),
+ dict(
+ type='ConvFCBBoxHead',
+ num_shared_convs=4,
+ num_shared_fcs=1,
+ in_channels=256,
+ conv_out_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.033, 0.033, 0.067, 0.067]),
+ reg_class_agnostic=False,
+ reg_decoded_bbox=True,
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=10.0))
+ ]))
+
+# augmentation strategy originates from DETR / Sparse RCNN
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[[
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(400, 1333), (500, 1333), (600, 1333)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333),
+ (576, 1333), (608, 1333), (640, 1333),
+ (672, 1333), (704, 1333), (736, 1333),
+ (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]]),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+max_epochs = 36
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0,
+ end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[27, 33],
+ gamma=0.1)
+]
+
+# Enable automatic-mixed-precision training with AmpOptimWrapper.
+optim_wrapper = dict(
+ type='AmpOptimWrapper',
+ constructor='LearningRateDecayOptimizerConstructor',
+ paramwise_cfg={
+ 'decay_rate': 0.7,
+ 'decay_type': 'layer_wise',
+ 'num_layers': 6
+ },
+ optimizer=dict(
+ _delete_=True,
+ type='AdamW',
+ lr=0.0002,
+ betas=(0.9, 0.999),
+ weight_decay=0.05))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/convnext/mask-rcnn_convnext-t-p4-w7_fpn_amp-ms-crop-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/convnext/mask-rcnn_convnext-t-p4-w7_fpn_amp-ms-crop-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5792b5b5c5a03c85a7d69040dd9a0b5381bc7995
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/convnext/mask-rcnn_convnext-t-p4-w7_fpn_amp-ms-crop-3x_coco.py
@@ -0,0 +1,96 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+# please install mmpretrain
+# import mmpretrain.models to trigger register_module in mmpretrain
+custom_imports = dict(
+ imports=['mmpretrain.models'], allow_failed_imports=False)
+checkpoint_file = 'https://download.openmmlab.com/mmclassification/v0/convnext/downstream/convnext-tiny_3rdparty_32xb128-noema_in1k_20220301-795e9634.pth' # noqa
+
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='mmpretrain.ConvNeXt',
+ arch='tiny',
+ out_indices=[0, 1, 2, 3],
+ drop_path_rate=0.4,
+ layer_scale_init_value=1.0,
+ gap_before_final_norm=False,
+ init_cfg=dict(
+ type='Pretrained', checkpoint=checkpoint_file,
+ prefix='backbone.')),
+ neck=dict(in_channels=[96, 192, 384, 768]))
+
+# augmentation strategy originates from DETR / Sparse RCNN
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[[
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(400, 1333), (500, 1333), (600, 1333)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333),
+ (576, 1333), (608, 1333), (640, 1333),
+ (672, 1333), (704, 1333), (736, 1333),
+ (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]]),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+max_epochs = 36
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0,
+ end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[27, 33],
+ gamma=0.1)
+]
+
+# Enable automatic-mixed-precision training with AmpOptimWrapper.
+optim_wrapper = dict(
+ type='AmpOptimWrapper',
+ constructor='LearningRateDecayOptimizerConstructor',
+ paramwise_cfg={
+ 'decay_rate': 0.95,
+ 'decay_type': 'layer_wise',
+ 'num_layers': 6
+ },
+ optimizer=dict(
+ _delete_=True,
+ type='AdamW',
+ lr=0.0001,
+ betas=(0.9, 0.999),
+ weight_decay=0.05,
+ ))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/convnext/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/convnext/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..b9fd7506cf46896d6c5f2238b594d32558ed3195
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/convnext/metafile.yml
@@ -0,0 +1,93 @@
+Models:
+ - Name: mask-rcnn_convnext-t-p4-w7_fpn_amp-ms-crop-3x_coco
+ In Collection: Mask R-CNN
+ Config: configs/convnext/mask-rcnn_convnext-t-p4-w7_fpn_amp-ms-crop-3x_coco.py
+ Metadata:
+ Training Memory (GB): 7.3
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - AdamW
+ - Mixed Precision Training
+ Training Resources: 8x A100 GPUs
+ Architecture:
+ - ConvNeXt
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 41.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/convnext/mask_rcnn_convnext-t_p4_w7_fpn_fp16_ms-crop_3x_coco/mask_rcnn_convnext-t_p4_w7_fpn_fp16_ms-crop_3x_coco_20220426_154953-050731f4.pth
+ Paper:
+ URL: https://arxiv.org/abs/2201.03545
+ Title: 'A ConvNet for the 2020s'
+ README: configs/convnext/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.16.0/mmdet/models/backbones/swin.py#L465
+ Version: v2.16.0
+
+ - Name: cascade-mask-rcnn_convnext-t-p4-w7_fpn_4conv1fc-giou_amp-ms-crop-3x_coco
+ In Collection: Cascade Mask R-CNN
+ Config: configs/convnext/cascade-mask-rcnn_convnext-t-p4-w7_fpn_4conv1fc-giou_amp-ms-crop-3x_coco.py
+ Metadata:
+ Training Memory (GB): 9.0
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - AdamW
+ - Mixed Precision Training
+ Training Resources: 8x A100 GPUs
+ Architecture:
+ - ConvNeXt
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 50.3
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 43.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/convnext/cascade_mask_rcnn_convnext-t_p4_w7_fpn_giou_4conv1f_fp16_ms-crop_3x_coco/cascade_mask_rcnn_convnext-t_p4_w7_fpn_giou_4conv1f_fp16_ms-crop_3x_coco_20220509_204200-8f07c40b.pth
+ Paper:
+ URL: https://arxiv.org/abs/2201.03545
+ Title: 'A ConvNet for the 2020s'
+ README: configs/convnext/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.16.0/mmdet/models/backbones/swin.py#L465
+ Version: v2.25.0
+
+ - Name: cascade-mask-rcnn_convnext-s-p4-w7_fpn_4conv1fc-giou_amp-ms-crop-3x_coco
+ In Collection: Cascade Mask R-CNN
+ Config: configs/convnext/cascade-mask-rcnn_convnext-s-p4-w7_fpn_4conv1fc-giou_amp-ms-crop-3x_coco.py
+ Metadata:
+ Training Memory (GB): 12.3
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - AdamW
+ - Mixed Precision Training
+ Training Resources: 8x A100 GPUs
+ Architecture:
+ - ConvNeXt
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 51.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 44.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/convnext/cascade_mask_rcnn_convnext-s_p4_w7_fpn_giou_4conv1f_fp16_ms-crop_3x_coco/cascade_mask_rcnn_convnext-s_p4_w7_fpn_giou_4conv1f_fp16_ms-crop_3x_coco_20220510_201004-3d24f5a4.pth
+ Paper:
+ URL: https://arxiv.org/abs/2201.03545
+ Title: 'A ConvNet for the 2020s'
+ README: configs/convnext/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.16.0/mmdet/models/backbones/swin.py#L465
+ Version: v2.25.0
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cornernet/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cornernet/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..e44964d8eac120f7313e7891b1771393b66bd9ae
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cornernet/README.md
@@ -0,0 +1,43 @@
+# CornerNet
+
+> [Cornernet: Detecting objects as paired keypoints](https://arxiv.org/abs/1808.01244)
+
+
+
+## Abstract
+
+We propose CornerNet, a new approach to object detection where we detect an object bounding box as a pair of keypoints, the top-left corner and the bottom-right corner, using a single convolution neural network. By detecting objects as paired keypoints, we eliminate the need for designing a set of anchor boxes commonly used in prior single-stage detectors. In addition to our novel formulation, we introduce corner pooling, a new type of pooling layer that helps the network better localize corners. Experiments show that CornerNet achieves a 42.2% AP on MS COCO, outperforming all existing one-stage detectors.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Batch Size | Step/Total Epochs | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :--------------: | :------------------------------------------------------------------: | :---------------: | :------: | :------------: | :----: | :------------------------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| HourglassNet-104 | [10 x 5](./cornernet_hourglass104_10xb5-crop511-210e-mstest_coco.py) | 180/210 | 13.9 | 4.2 | 41.2 | [config](./cornernet_hourglass104_10xb5-crop511-210e-mstest_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cornernet/cornernet_hourglass104_mstest_10x5_210e_coco/cornernet_hourglass104_mstest_10x5_210e_coco_20200824_185720-5fefbf1c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cornernet/cornernet_hourglass104_mstest_10x5_210e_coco/cornernet_hourglass104_mstest_10x5_210e_coco_20200824_185720.log.json) |
+| HourglassNet-104 | [8 x 6](./cornernet_hourglass104_8xb6-210e-mstest_coco.py) | 180/210 | 15.9 | 4.2 | 41.2 | [config](./cornernet_hourglass104_8xb6-210e-mstest_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cornernet/cornernet_hourglass104_mstest_8x6_210e_coco/cornernet_hourglass104_mstest_8x6_210e_coco_20200825_150618-79b44c30.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cornernet/cornernet_hourglass104_mstest_8x6_210e_coco/cornernet_hourglass104_mstest_8x6_210e_coco_20200825_150618.log.json) |
+| HourglassNet-104 | [32 x 3](./cornernet_hourglass104_32xb3-210e-mstest_coco.py) | 180/210 | 9.5 | 3.9 | 40.4 | [config](./cornernet_hourglass104_32xb3-210e-mstest_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/cornernet/cornernet_hourglass104_mstest_32x3_210e_coco/cornernet_hourglass104_mstest_32x3_210e_coco_20200819_203110-1efaea91.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/cornernet/cornernet_hourglass104_mstest_32x3_210e_coco/cornernet_hourglass104_mstest_32x3_210e_coco_20200819_203110.log.json) |
+
+Note:
+
+- TTA setting is single-scale and `flip=True`. If you want to reproduce the TTA performance, please add `--tta` in the test command.
+- Experiments with `images_per_gpu=6` are conducted on Tesla V100-SXM2-32GB, `images_per_gpu=3` are conducted on GeForce GTX 1080 Ti.
+- Here are the descriptions of each experiment setting:
+ - 10 x 5: 10 GPUs with 5 images per gpu. This is the same setting as that reported in the original paper.
+ - 8 x 6: 8 GPUs with 6 images per gpu. The total batchsize is similar to paper and only need 1 node to train.
+ - 32 x 3: 32 GPUs with 3 images per gpu. The default setting for 1080TI and need 4 nodes to train.
+
+## Citation
+
+```latex
+@inproceedings{law2018cornernet,
+ title={Cornernet: Detecting objects as paired keypoints},
+ author={Law, Hei and Deng, Jia},
+ booktitle={15th European Conference on Computer Vision, ECCV 2018},
+ pages={765--781},
+ year={2018},
+ organization={Springer Verlag}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cornernet/cornernet_hourglass104_10xb5-crop511-210e-mstest_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cornernet/cornernet_hourglass104_10xb5-crop511-210e-mstest_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..76339163b618a5a9d41a542ec75192aedb409eea
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cornernet/cornernet_hourglass104_10xb5-crop511-210e-mstest_coco.py
@@ -0,0 +1,8 @@
+_base_ = './cornernet_hourglass104_8xb6-210e-mstest_coco.py'
+
+train_dataloader = dict(batch_size=5)
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (10 GPUs) x (5 samples per GPU)
+auto_scale_lr = dict(base_batch_size=50)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cornernet/cornernet_hourglass104_32xb3-210e-mstest_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cornernet/cornernet_hourglass104_32xb3-210e-mstest_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..51a4740318a1d85a62b6b4482c53808c98fb8a62
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cornernet/cornernet_hourglass104_32xb3-210e-mstest_coco.py
@@ -0,0 +1,8 @@
+_base_ = './cornernet_hourglass104_8xb6-210e-mstest_coco.py'
+
+train_dataloader = dict(batch_size=3)
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (32 GPUs) x (3 samples per GPU)
+auto_scale_lr = dict(base_batch_size=96)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cornernet/cornernet_hourglass104_8xb6-210e-mstest_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cornernet/cornernet_hourglass104_8xb6-210e-mstest_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..bdb46fff164f796d9333c123deb701c341bdc1e3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cornernet/cornernet_hourglass104_8xb6-210e-mstest_coco.py
@@ -0,0 +1,183 @@
+_base_ = [
+ '../_base_/default_runtime.py', '../_base_/datasets/coco_detection.py'
+]
+
+data_preprocessor = dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True)
+
+# model settings
+model = dict(
+ type='CornerNet',
+ data_preprocessor=data_preprocessor,
+ backbone=dict(
+ type='HourglassNet',
+ downsample_times=5,
+ num_stacks=2,
+ stage_channels=[256, 256, 384, 384, 384, 512],
+ stage_blocks=[2, 2, 2, 2, 2, 4],
+ norm_cfg=dict(type='BN', requires_grad=True)),
+ neck=None,
+ bbox_head=dict(
+ type='CornerHead',
+ num_classes=80,
+ in_channels=256,
+ num_feat_levels=2,
+ corner_emb_channels=1,
+ loss_heatmap=dict(
+ type='GaussianFocalLoss', alpha=2.0, gamma=4.0, loss_weight=1),
+ loss_embedding=dict(
+ type='AssociativeEmbeddingLoss',
+ pull_weight=0.10,
+ push_weight=0.10),
+ loss_offset=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1)),
+ # training and testing settings
+ train_cfg=None,
+ test_cfg=dict(
+ corner_topk=100,
+ local_maximum_kernel=3,
+ distance_threshold=0.5,
+ score_thr=0.05,
+ max_per_img=100,
+ nms=dict(type='soft_nms', iou_threshold=0.5, method='gaussian')))
+
+# data settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PhotoMetricDistortion',
+ brightness_delta=32,
+ contrast_range=(0.5, 1.5),
+ saturation_range=(0.5, 1.5),
+ hue_delta=18),
+ dict(
+ # The cropped images are padded into squares during training,
+ # but may be smaller than crop_size.
+ type='RandomCenterCropPad',
+ crop_size=(511, 511),
+ ratios=(0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3),
+ test_mode=False,
+ test_pad_mode=None,
+ mean=data_preprocessor['mean'],
+ std=data_preprocessor['std'],
+ # Image data is not converted to rgb.
+ to_rgb=data_preprocessor['bgr_to_rgb']),
+ # Make sure the output is always crop_size.
+ dict(type='Resize', scale=(511, 511), keep_ratio=False),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs'),
+]
+
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ to_float32=True,
+ backend_args=_base_.backend_args,
+ ),
+ # don't need Resize
+ dict(
+ type='RandomCenterCropPad',
+ crop_size=None,
+ ratios=None,
+ border=None,
+ test_mode=True,
+ test_pad_mode=['logical_or', 127],
+ mean=data_preprocessor['mean'],
+ std=data_preprocessor['std'],
+ # Image data is not converted to rgb.
+ to_rgb=data_preprocessor['bgr_to_rgb']),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape', 'border'))
+]
+
+train_dataloader = dict(
+ batch_size=6,
+ num_workers=3,
+ batch_sampler=None,
+ dataset=dict(pipeline=train_pipeline))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='Adam', lr=0.0005),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+max_epochs = 210
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 3,
+ by_epoch=False,
+ begin=0,
+ end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[180],
+ gamma=0.1)
+]
+
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (6 samples per GPU)
+auto_scale_lr = dict(base_batch_size=48)
+
+tta_model = dict(
+ type='DetTTAModel',
+ tta_cfg=dict(
+ nms=dict(type='soft_nms', iou_threshold=0.5, method='gaussian'),
+ max_per_img=100))
+
+tta_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ to_float32=True,
+ backend_args=_base_.backend_args),
+ dict(
+ type='TestTimeAug',
+ transforms=[
+ [
+ # ``RandomFlip`` must be placed before ``RandomCenterCropPad``,
+ # otherwise bounding box coordinates after flipping cannot be
+ # recovered correctly.
+ dict(type='RandomFlip', prob=1.),
+ dict(type='RandomFlip', prob=0.)
+ ],
+ [
+ dict(
+ type='RandomCenterCropPad',
+ crop_size=None,
+ ratios=None,
+ border=None,
+ test_mode=True,
+ test_pad_mode=['logical_or', 127],
+ mean=data_preprocessor['mean'],
+ std=data_preprocessor['std'],
+ # Image data is not converted to rgb.
+ to_rgb=data_preprocessor['bgr_to_rgb'])
+ ],
+ [dict(type='LoadAnnotations', with_bbox=True)],
+ [
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'flip', 'flip_direction', 'border'))
+ ]
+ ])
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/cornernet/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cornernet/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..f915cf37e8e157405a66431dfb21595db319b8b6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/cornernet/metafile.yml
@@ -0,0 +1,83 @@
+Collections:
+ - Name: CornerNet
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - Adam
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Corner Pooling
+ - Stacked Hourglass Network
+ Paper:
+ URL: https://arxiv.org/abs/1808.01244
+ Title: 'CornerNet: Detecting Objects as Paired Keypoints'
+ README: configs/cornernet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.3.0/mmdet/models/detectors/cornernet.py#L9
+ Version: v2.3.0
+
+Models:
+ - Name: cornernet_hourglass104_10xb5-crop511-210e-mstest_coco
+ In Collection: CornerNet
+ Config: configs/cornernet/cornernet_hourglass104_10xb5-crop511-210e-mstest_coco.py
+ Metadata:
+ Training Resources: 10x V100 GPUs
+ Batch Size: 50
+ Training Memory (GB): 13.9
+ inference time (ms/im):
+ - value: 238.1
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 210
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cornernet/cornernet_hourglass104_mstest_10x5_210e_coco/cornernet_hourglass104_mstest_10x5_210e_coco_20200824_185720-5fefbf1c.pth
+
+ - Name: cornernet_hourglass104_8xb6-210e-mstest_coco
+ In Collection: CornerNet
+ Config: configs/cornernet/cornernet_hourglass104_8xb6-210e-mstest_coco.py
+ Metadata:
+ Batch Size: 48
+ Training Memory (GB): 15.9
+ inference time (ms/im):
+ - value: 238.1
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 210
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cornernet/cornernet_hourglass104_mstest_8x6_210e_coco/cornernet_hourglass104_mstest_8x6_210e_coco_20200825_150618-79b44c30.pth
+
+ - Name: cornernet_hourglass104_32xb3-210e-mstest_coco
+ In Collection: CornerNet
+ Config: configs/cornernet/cornernet_hourglass104_32xb3-210e-mstest_coco.py
+ Metadata:
+ Training Resources: 32x V100 GPUs
+ Batch Size: 96
+ Training Memory (GB): 9.5
+ inference time (ms/im):
+ - value: 256.41
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 210
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/cornernet/cornernet_hourglass104_mstest_32x3_210e_coco/cornernet_hourglass104_mstest_32x3_210e_coco_20200819_203110-1efaea91.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/crowddet/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/crowddet/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..abc0f2d2dfac8fa64cab267c20f58c9113737d07
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/crowddet/README.md
@@ -0,0 +1,37 @@
+# CrowdDet
+
+> [Detection in Crowded Scenes: One Proposal, Multiple Predictions](https://arxiv.org/abs/2003.09163)
+
+
+
+## Abstract
+
+We propose a simple yet effective proposal-based object detector, aiming at detecting highly-overlapped instances in crowded scenes. The key of our approach is to let each proposal predict a set of correlated instances rather than a single one in previous proposal-based frameworks. Equipped with new techniques such as EMD Loss and Set NMS, our detector can effectively handle the difficulty of detecting highly overlapped objects. On a FPN-Res50 baseline, our detector can obtain 4.9% AP gains on challenging CrowdHuman dataset and 1.0% MR^−2 improvements on CityPersons dataset, without bells and whistles. Moreover, on less crowed datasets like COCO, our approach can still achieve moderate improvement, suggesting the proposed method is robust to crowdedness. Code and pre-trained models will be released at https://github.com/megvii-model/CrowdDetection.
+
+
+

+
+
+## Results and Models
+
+| Backbone | RM | Style | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :------: | :---: | :-----: | :------: | :------------: | :----: | :-------------------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | False | pytorch | 4.4 | - | 90.0 | [config](./crowddet-rcnn_r50_fpn_8xb2-30e_crowdhuman.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/crowddet/crowddet-rcnn_r50_fpn_8xb2-30e_crowdhuman/crowddet-rcnn_r50_fpn_8xb2-30e_crowdhuman_20221023_174954-dc319c2d.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/crowddet/crowddet-rcnn_r50_fpn_8xb2-30e_crowdhuman/crowddet-rcnn_r50_fpn_8xb2-30e_crowdhuman_20221023_174954.log.json) |
+| R-50-FPN | True | pytorch | 4.8 | - | 90.32 | [config](./crowddet-rcnn_refine_r50_fpn_8xb2-30e_crowdhuman.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/crowddet/crowddet-rcnn_refine_r50_fpn_8xb2-30e_crowdhuman/crowddet-rcnn_refine_r50_fpn_8xb2-30e_crowdhuman_20221024_215917-45602806.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/crowddet/crowddet-rcnn_refine_r50_fpn_8xb2-30e_crowdhuman/crowddet-rcnn_refine_r50_fpn_8xb2-30e_crowdhuman_20221024_215917.log.json) |
+
+Note:
+
+- RM indicates whether to use the refine module.
+- The dataset for training and testing this model is `CrowdHuman`, and the metric of `box AP` is calculated by `mmdet/evaluation/metrics/crowdhuman_metric.py`.
+
+## Citation
+
+```latex
+@inproceedings{Chu_2020_CVPR,
+ title={Detection in Crowded Scenes: One Proposal, Multiple Predictions},
+ author={Chu, Xuangeng and Zheng, Anlin and Zhang, Xiangyu and Sun, Jian},
+ booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
+ month = {June},
+ year = {2020}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/crowddet/crowddet-rcnn_r50_fpn_8xb2-30e_crowdhuman.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/crowddet/crowddet-rcnn_r50_fpn_8xb2-30e_crowdhuman.py
new file mode 100644
index 0000000000000000000000000000000000000000..8815be77d49cf77afff6f888ee225e928e43b402
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/crowddet/crowddet-rcnn_r50_fpn_8xb2-30e_crowdhuman.py
@@ -0,0 +1,227 @@
+_base_ = ['../_base_/default_runtime.py']
+
+model = dict(
+ type='CrowdDet',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.53, 116.28, 123.675],
+ std=[57.375, 57.12, 58.395],
+ bgr_to_rgb=False,
+ pad_size_divisor=64,
+ # This option is set according to https://github.com/Purkialo/CrowdDet/
+ # blob/master/lib/data/CrowdHuman.py The images in the entire batch are
+ # resize together.
+ batch_augments=[
+ dict(type='BatchResize', scale=(1400, 800), pad_size_divisor=64)
+ ]),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5,
+ upsample_cfg=dict(mode='bilinear', align_corners=False)),
+ rpn_head=dict(
+ type='RPNHead',
+ in_channels=256,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ scales=[8],
+ ratios=[1.0, 2.0, 3.0],
+ strides=[4, 8, 16, 32, 64],
+ centers=[(8, 8), (8, 8), (8, 8), (8, 8), (8, 8)]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0.0, 0.0, 0.0, 0.0],
+ target_stds=[1.0, 1.0, 1.0, 1.0],
+ clip_border=False),
+ loss_cls=dict(type='CrossEntropyLoss', loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0)),
+ roi_head=dict(
+ type='MultiInstanceRoIHead',
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(
+ type='RoIAlign',
+ output_size=7,
+ sampling_ratio=-1,
+ aligned=True,
+ use_torchvision=True),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ bbox_head=dict(
+ type='MultiInstanceBBoxHead',
+ with_refine=False,
+ num_shared_fcs=2,
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=1,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=False,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ loss_weight=1.0,
+ use_sigmoid=False,
+ reduction='none'),
+ loss_bbox=dict(
+ type='SmoothL1Loss', loss_weight=1.0, reduction='none'))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=(0.3, 0.7),
+ min_pos_iou=0.3,
+ match_low_quality=True,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ rpn_proposal=dict(
+ nms_pre=2400,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=2),
+ rcnn=dict(
+ assigner=dict(
+ type='MultiInstanceAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.3,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='MultiInsRandomSampler',
+ num=512,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ pos_weight=-1,
+ debug=False)),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=1200,
+ max_per_img=1000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=2),
+ rcnn=dict(
+ nms=dict(type='nms', iou_threshold=0.5),
+ score_thr=0.01,
+ max_per_img=500)))
+
+dataset_type = 'CrowdHumanDataset'
+data_root = 'data/CrowdHuman/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/tracking/CrowdHuman/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/tracking/',
+# 'data/': 's3://openmmlab/datasets/tracking/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape', 'flip',
+ 'flip_direction'))
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1400, 800), keep_ratio=True),
+ # avoid bboxes being resized
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=4,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=None, # The 'batch_sampler' may decrease the precision
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotation_train.odgt',
+ data_prefix=dict(img='Images/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotation_val.odgt',
+ data_prefix=dict(img='Images/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CrowdHumanMetric',
+ ann_file=data_root + 'annotation_val.odgt',
+ metric=['AP', 'MR', 'JI'],
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+train_cfg = dict(type='EpochBasedTrainLoop', max_epochs=30, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=800),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=30,
+ by_epoch=True,
+ milestones=[24, 27],
+ gamma=0.1)
+]
+
+# optimizer
+auto_scale_lr = dict(base_batch_size=16)
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.002, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/crowddet/crowddet-rcnn_refine_r50_fpn_8xb2-30e_crowdhuman.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/crowddet/crowddet-rcnn_refine_r50_fpn_8xb2-30e_crowdhuman.py
new file mode 100644
index 0000000000000000000000000000000000000000..80277ce1c1436c37c4e2a4d13293d0ecb8ba4722
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/crowddet/crowddet-rcnn_refine_r50_fpn_8xb2-30e_crowdhuman.py
@@ -0,0 +1,3 @@
+_base_ = './crowddet-rcnn_r50_fpn_8xb2-30e_crowdhuman.py'
+
+model = dict(roi_head=dict(bbox_head=dict(with_refine=True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/crowddet/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/crowddet/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..4f191dea9cc599f64091434152000e67289f9180
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/crowddet/metafile.yml
@@ -0,0 +1,47 @@
+Collections:
+ - Name: CrowdDet
+ Metadata:
+ Training Data: CrowdHuman
+ Training Techniques:
+ - SGD
+ - EMD Loss
+ Training Resources: 8x A100 GPUs
+ Architecture:
+ - FPN
+ - RPN
+ - ResNet
+ - RoIPool
+ Paper:
+ URL: https://arxiv.org/abs/2003.09163
+ Title: 'Detection in Crowded Scenes: One Proposal, Multiple Predictions'
+ README: configs/crowddet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v3.0.0rc3/mmdet/models/detectors/crowddet.py
+ Version: v3.0.0rc3
+
+Models:
+ - Name: crowddet-rcnn_refine_r50_fpn_8xb2-30e_crowdhuman
+ In Collection: CrowdDet
+ Config: configs/crowddet/crowddet-rcnn_refine_r50_fpn_8xb2-30e_crowdhuman.py
+ Metadata:
+ Training Memory (GB): 4.8
+ Epochs: 30
+ Results:
+ - Task: Object Detection
+ Dataset: CrowdHuman
+ Metrics:
+ box AP: 90.32
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/crowddet/crowddet-rcnn_refine_r50_fpn_8xb2-30e_crowdhuman/crowddet-rcnn_refine_r50_fpn_8xb2-30e_crowdhuman_20221024_215917-45602806.pth
+
+ - Name: crowddet-rcnn_r50_fpn_8xb2-30e_crowdhuman
+ In Collection: CrowdDet
+ Config: configs/crowddet/crowddet-rcnn_r50_fpn_8xb2-30e_crowdhuman.py
+ Metadata:
+ Training Memory (GB): 4.4
+ Epochs: 30
+ Results:
+ - Task: Object Detection
+ Dataset: CrowdHuman
+ Metrics:
+ box AP: 90.0
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/crowddet/crowddet-rcnn_r50_fpn_8xb2-30e_crowdhuman/crowddet-rcnn_r50_fpn_8xb2-30e_crowdhuman_20221023_174954-dc319c2d.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dab_detr/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dab_detr/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..5661f27a30268a9a50a956e51e948c36c9287356
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dab_detr/README.md
@@ -0,0 +1,40 @@
+# DAB-DETR
+
+> [DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR](https://arxiv.org/abs/2201.12329)
+
+
+
+## Abstract
+
+We present in this paper a novel query formulation using dynamic anchor boxes for DETR (DEtection TRansformer) and offer a deeper understanding of the role of queries in DETR. This new formulation directly uses box coordinates as queries in Transformer decoders and dynamically updates them layer-by-layer. Using box coordinates not only helps using explicit positional priors to improve the query-to-feature similarity and eliminate the slow training convergence issue in DETR, but also allows us to modulate the positional attention map using the box width and height information. Such a design makes it clear that queries in DETR can be implemented as performing soft ROI pooling layer-by-layer in a cascade manner. As a result, it leads to the best performance on MS-COCO benchmark among the DETR-like detection models under the same setting, e.g., AP 45.7% using ResNet50-DC5 as backbone trained in 50 epochs. We also conducted extensive experiments to confirm our analysis and verify the effectiveness of our methods.
+
+
+

+
+
+

+
+
+

+
+
+## Results and Models
+
+We provide the config files and models for DAB-DETR: [DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR](https://arxiv.org/abs/2201.12329).
+
+| Backbone | Model | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :------: | :------: | :-----: | :------: | :------------: | :----: | :---------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | DAB-DETR | 50e | | | 42.3 | [config](./dab-detr_r50_8xb2-50e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/dab_detr/dab-detr_r50_8xb2-50e_coco/dab-detr_r50_8xb2-50e_coco_20221122_120837-c1035c8c.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/dab_detr/dab-detr_r50_8xb2-50e_coco/dab-detr_r50_8xb2-50e_coco_20221122_120837.log.json) |
+
+## Citation
+
+```latex
+@inproceedings{
+ liu2022dabdetr,
+ title={{DAB}-{DETR}: Dynamic Anchor Boxes are Better Queries for {DETR}},
+ author={Shilong Liu and Feng Li and Hao Zhang and Xiao Yang and Xianbiao Qi and Hang Su and Jun Zhu and Lei Zhang},
+ booktitle={International Conference on Learning Representations},
+ year={2022},
+ url={https://openreview.net/forum?id=oMI9PjOb9Jl}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dab_detr/dab-detr_r50_8xb2-50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dab_detr/dab-detr_r50_8xb2-50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..314ed97e2d80ae3c95119abf9166f95d416c010e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dab_detr/dab-detr_r50_8xb2-50e_coco.py
@@ -0,0 +1,159 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ type='DABDETR',
+ num_queries=300,
+ with_random_refpoints=False,
+ num_patterns=0,
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=1),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(3, ),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='ChannelMapper',
+ in_channels=[2048],
+ kernel_size=1,
+ out_channels=256,
+ act_cfg=None,
+ norm_cfg=None,
+ num_outs=1),
+ encoder=dict(
+ num_layers=6,
+ layer_cfg=dict(
+ self_attn_cfg=dict(
+ embed_dims=256, num_heads=8, dropout=0., batch_first=True),
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048,
+ num_fcs=2,
+ ffn_drop=0.,
+ act_cfg=dict(type='PReLU')))),
+ decoder=dict(
+ num_layers=6,
+ query_dim=4,
+ query_scale_type='cond_elewise',
+ with_modulated_hw_attn=True,
+ layer_cfg=dict(
+ self_attn_cfg=dict(
+ embed_dims=256,
+ num_heads=8,
+ attn_drop=0.,
+ proj_drop=0.,
+ cross_attn=False),
+ cross_attn_cfg=dict(
+ embed_dims=256,
+ num_heads=8,
+ attn_drop=0.,
+ proj_drop=0.,
+ cross_attn=True),
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048,
+ num_fcs=2,
+ ffn_drop=0.,
+ act_cfg=dict(type='PReLU'))),
+ return_intermediate=True),
+ positional_encoding=dict(num_feats=128, temperature=20, normalize=True),
+ bbox_head=dict(
+ type='DABDETRHead',
+ num_classes=80,
+ embed_dims=256,
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=5.0),
+ loss_iou=dict(type='GIoULoss', loss_weight=2.0)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='HungarianAssigner',
+ match_costs=[
+ dict(type='FocalLossCost', weight=2., eps=1e-8),
+ dict(type='BBoxL1Cost', weight=5.0, box_format='xywh'),
+ dict(type='IoUCost', iou_mode='giou', weight=2.0)
+ ])),
+ test_cfg=dict(max_per_img=300))
+
+# train_pipeline, NOTE the img_scale and the Pad's size_divisor is different
+# from the default setting in mmdet.
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[[
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(400, 1333), (500, 1333), (600, 1333)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333),
+ (576, 1333), (608, 1333), (640, 1333),
+ (672, 1333), (704, 1333), (736, 1333),
+ (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]]),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0001, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(
+ custom_keys={'backbone': dict(lr_mult=0.1, decay_mult=1.0)}))
+
+# learning policy
+max_epochs = 50
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[40],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=16, enable=False)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dab_detr/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dab_detr/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..94383a0493b86a730181f78ab2f0e94a2ab2de73
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dab_detr/metafile.yml
@@ -0,0 +1,32 @@
+Collections:
+ - Name: DAB-DETR
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - AdamW
+ - Multi Scale Train
+ - Gradient Clip
+ Training Resources: 8x A100 GPUs
+ Architecture:
+ - ResNet
+ - Transformer
+ Paper:
+ URL: https://arxiv.org/abs/2201.12329
+ Title: 'DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR'
+ README: configs/dab_detr/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/f4112c9e5611468ffbd57cfba548fd1289264b52/mmdet/models/detectors/dab_detr.py#L15
+ Version: v3.0.0rc6
+
+Models:
+ - Name: dab-detr_r50_8xb2-50e_coco
+ In Collection: DAB-DETR
+ Config: configs/dab_detr/dab-detr_r50_8xb2-50e_coco.py
+ Metadata:
+ Epochs: 50
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.3
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/dab_detr/dab-detr_r50_8xb2-50e_coco/dab-detr_r50_8xb2-50e_coco_20221122_120837-c1035c8c.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..e287e1d5ef99e68dd2d7f2fccbacddde7428522e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/README.md
@@ -0,0 +1,48 @@
+# DCN
+
+> [Deformable Convolutional Networks](https://arxiv.org/abs/1703.06211)
+
+
+
+## Abstract
+
+Convolutional neural networks (CNNs) are inherently limited to model geometric transformations due to the fixed geometric structures in its building modules. In this work, we introduce two new modules to enhance the transformation modeling capacity of CNNs, namely, deformable convolution and deformable RoI pooling. Both are based on the idea of augmenting the spatial sampling locations in the modules with additional offsets and learning the offsets from target tasks, without additional supervision. The new modules can readily replace their plain counterparts in existing CNNs and can be easily trained end-to-end by standard back-propagation, giving rise to deformable convolutional networks. Extensive experiments validate the effectiveness of our approach on sophisticated vision tasks of object detection and semantic segmentation.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Model | Style | Conv | Pool | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :-------------: | :----------: | :-----: | :----------: | :---: | :-----: | :------: | :------------: | :----: | :-----: | :-----------------------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | Faster | pytorch | dconv(c3-c5) | - | 1x | 4.0 | 17.8 | 41.3 | | [config](./faster-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_r50_fpn_dconv_c3-c5_1x_coco/faster_rcnn_r50_fpn_dconv_c3-c5_1x_coco_20200130-d68aed1e.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_r50_fpn_dconv_c3-c5_1x_coco/faster_rcnn_r50_fpn_dconv_c3-c5_1x_coco_20200130_212941.log.json) |
+| R-50-FPN | Faster | pytorch | - | dpool | 1x | 5.0 | 17.2 | 38.9 | | [config](./faster-rcnn_r50_fpn_dpool_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_r50_fpn_dpool_1x_coco/faster_rcnn_r50_fpn_dpool_1x_coco_20200307-90d3c01d.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_r50_fpn_dpool_1x_coco/faster_rcnn_r50_fpn_dpool_1x_coco_20200307_203250.log.json) |
+| R-101-FPN | Faster | pytorch | dconv(c3-c5) | - | 1x | 6.0 | 12.5 | 42.7 | | [config](./faster-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_r101_fpn_dconv_c3-c5_1x_coco/faster_rcnn_r101_fpn_dconv_c3-c5_1x_coco_20200203-1377f13d.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_r101_fpn_dconv_c3-c5_1x_coco/faster_rcnn_r101_fpn_dconv_c3-c5_1x_coco_20200203_230019.log.json) |
+| X-101-32x4d-FPN | Faster | pytorch | dconv(c3-c5) | - | 1x | 7.3 | 10.0 | 44.5 | | [config](./faster-rcnn_x101-32x4d-dconv-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_x101_32x4d_fpn_dconv_c3-c5_1x_coco/faster_rcnn_x101_32x4d_fpn_dconv_c3-c5_1x_coco_20200203-4f85c69c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_x101_32x4d_fpn_dconv_c3-c5_1x_coco/faster_rcnn_x101_32x4d_fpn_dconv_c3-c5_1x_coco_20200203_001325.log.json) |
+| R-50-FPN | Mask | pytorch | dconv(c3-c5) | - | 1x | 4.5 | 15.4 | 41.8 | 37.4 | [config](./mask-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/dcn/mask_rcnn_r50_fpn_dconv_c3-c5_1x_coco/mask_rcnn_r50_fpn_dconv_c3-c5_1x_coco_20200203-4d9ad43b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/dcn/mask_rcnn_r50_fpn_dconv_c3-c5_1x_coco/mask_rcnn_r50_fpn_dconv_c3-c5_1x_coco_20200203_061339.log.json) |
+| R-101-FPN | Mask | pytorch | dconv(c3-c5) | - | 1x | 6.5 | 11.7 | 43.5 | 38.9 | [config](./mask-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/dcn/mask_rcnn_r101_fpn_dconv_c3-c5_1x_coco/mask_rcnn_r101_fpn_dconv_c3-c5_1x_coco_20200216-a71f5bce.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/dcn/mask_rcnn_r101_fpn_dconv_c3-c5_1x_coco/mask_rcnn_r101_fpn_dconv_c3-c5_1x_coco_20200216_191601.log.json) |
+| R-50-FPN | Cascade | pytorch | dconv(c3-c5) | - | 1x | 4.5 | 14.6 | 43.8 | | [config](./cascade-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/dcn/cascade_rcnn_r50_fpn_dconv_c3-c5_1x_coco/cascade_rcnn_r50_fpn_dconv_c3-c5_1x_coco_20200130-2f1fca44.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/dcn/cascade_rcnn_r50_fpn_dconv_c3-c5_1x_coco/cascade_rcnn_r50_fpn_dconv_c3-c5_1x_coco_20200130_220843.log.json) |
+| R-101-FPN | Cascade | pytorch | dconv(c3-c5) | - | 1x | 6.4 | 11.0 | 45.0 | | [config](./cascade-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/dcn/cascade_rcnn_r101_fpn_dconv_c3-c5_1x_coco/cascade_rcnn_r101_fpn_dconv_c3-c5_1x_coco_20200203-3b2f0594.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/dcn/cascade_rcnn_r101_fpn_dconv_c3-c5_1x_coco/cascade_rcnn_r101_fpn_dconv_c3-c5_1x_coco_20200203_224829.log.json) |
+| R-50-FPN | Cascade Mask | pytorch | dconv(c3-c5) | - | 1x | 6.0 | 10.0 | 44.4 | 38.6 | [config](./cascade-mask-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/dcn/cascade_mask_rcnn_r50_fpn_dconv_c3-c5_1x_coco/cascade_mask_rcnn_r50_fpn_dconv_c3-c5_1x_coco_20200202-42e767a2.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/dcn/cascade_mask_rcnn_r50_fpn_dconv_c3-c5_1x_coco/cascade_mask_rcnn_r50_fpn_dconv_c3-c5_1x_coco_20200202_010309.log.json) |
+| R-101-FPN | Cascade Mask | pytorch | dconv(c3-c5) | - | 1x | 8.0 | 8.6 | 45.8 | 39.7 | [config](./cascade-mask-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/dcn/cascade_mask_rcnn_r101_fpn_dconv_c3-c5_1x_coco/cascade_mask_rcnn_r101_fpn_dconv_c3-c5_1x_coco_20200204-df0c5f10.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/dcn/cascade_mask_rcnn_r101_fpn_dconv_c3-c5_1x_coco/cascade_mask_rcnn_r101_fpn_dconv_c3-c5_1x_coco_20200204_134006.log.json) |
+| X-101-32x4d-FPN | Cascade Mask | pytorch | dconv(c3-c5) | - | 1x | 9.2 | | 47.3 | 41.1 | [config](./cascade-mask-rcnn_x101-32x4d-dconv-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/dcn/cascade_mask_rcnn_x101_32x4d_fpn_dconv_c3-c5_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_dconv_c3-c5_1x_coco-e75f90c8.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/dcn/cascade_mask_rcnn_x101_32x4d_fpn_dconv_c3-c5_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_dconv_c3-c5_1x_coco-20200606_183737.log.json) |
+| R-50-FPN (FP16) | Mask | pytorch | dconv(c3-c5) | - | 1x | 3.0 | | 41.9 | 37.5 | [config](./mask-rcnn_r50-dconv-c3-c5_fpn_amp-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fp16/mask_rcnn_r50_fpn_fp16_dconv_c3-c5_1x_coco/mask_rcnn_r50_fpn_fp16_dconv_c3-c5_1x_coco_20210520_180247-c06429d2.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fp16/mask_rcnn_r50_fpn_fp16_dconv_c3-c5_1x_coco/mask_rcnn_r50_fpn_fp16_dconv_c3-c5_1x_coco_20210520_180247.log.json) |
+
+**Notes:**
+
+- `dconv` denotes deformable convolution, `c3-c5` means adding dconv in resnet stage 3 to 5. `dpool` denotes deformable roi pooling.
+- The dcn ops are modified from https://github.com/chengdazhi/Deformable-Convolution-V2-PyTorch, which should be more memory efficient and slightly faster.
+- (\*) For R-50-FPN (dg=4), dg is short for deformable_group. This model is trained and tested on Amazon EC2 p3dn.24xlarge instance.
+- **Memory, Train/Inf time is outdated.**
+
+## Citation
+
+```latex
+@inproceedings{dai2017deformable,
+ title={Deformable Convolutional Networks},
+ author={Dai, Jifeng and Qi, Haozhi and Xiong, Yuwen and Li, Yi and Zhang, Guodong and Hu, Han and Wei, Yichen},
+ booktitle={Proceedings of the IEEE international conference on computer vision},
+ year={2017}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/cascade-mask-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/cascade-mask-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8c0ff9890e82bd0c1ee4e445e37d2c7afa534161
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/cascade-mask-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,5 @@
+_base_ = '../cascade_rcnn/cascade-mask-rcnn_r101_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCN', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/cascade-mask-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/cascade-mask-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..cfcc5e73cc508e11d77c5a3557f30632b545b803
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/cascade-mask-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,5 @@
+_base_ = '../cascade_rcnn/cascade-mask-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCN', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/cascade-mask-rcnn_x101-32x4d-dconv-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/cascade-mask-rcnn_x101-32x4d-dconv-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..48b25f62125da09368c446bcd6ccff9b0219a7cc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/cascade-mask-rcnn_x101-32x4d-dconv-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,5 @@
+_base_ = '../cascade_rcnn/cascade-mask-rcnn_x101-32x4d_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCN', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/cascade-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/cascade-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8a942da754119b8d913f807907322a3d96c83ff8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/cascade-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,5 @@
+_base_ = '../cascade_rcnn/cascade-rcnn_r101_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCN', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/cascade-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/cascade-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f6bf5b7998a972f41b52f90955ef52977adfd68c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/cascade-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,5 @@
+_base_ = '../cascade_rcnn/cascade-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCN', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/faster-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/faster-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..db44e7e87b2d11555140ab2c8a19f32e1ce65770
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/faster-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,5 @@
+_base_ = '../faster_rcnn/faster-rcnn_r101_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCN', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/faster-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/faster-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..95f20467af60167a4a61f253e4354dadd832ccc7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/faster-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,5 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCN', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/faster-rcnn_r50_fpn_dpool_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/faster-rcnn_r50_fpn_dpool_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c65ce5fd0267dc892455da6495cd3be9f1f99fcf
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/faster-rcnn_r50_fpn_dpool_1x_coco.py
@@ -0,0 +1,12 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ roi_head=dict(
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(
+ _delete_=True,
+ type='DeformRoIPoolPack',
+ output_size=7,
+ output_channels=256),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32])))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/faster-rcnn_x101-32x4d-dconv-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/faster-rcnn_x101-32x4d-dconv-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e4ed832f5e7ff0d050be33e57d2fa611e9ae7e8e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/faster-rcnn_x101-32x4d-dconv-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,16 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ dcn=dict(type='DCN', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/mask-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/mask-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3f36714a5301823ca401820ab9d926374428ee70
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/mask-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,5 @@
+_base_ = '../mask_rcnn/mask-rcnn_r101_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCN', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/mask-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/mask-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..0b281d417b4f6a7320201da261e5fdf6950556a1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/mask-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,5 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCN', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/mask-rcnn_r50-dconv-c3-c5_fpn_amp-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/mask-rcnn_r50-dconv-c3-c5_fpn_amp-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..9d01594314aad74bc47d7331c42a39f2ca453071
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/mask-rcnn_r50-dconv-c3-c5_fpn_amp-1x_coco.py
@@ -0,0 +1,10 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCN', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)))
+
+# MMEngine support the following two ways, users can choose
+# according to convenience
+# optim_wrapper = dict(type='AmpOptimWrapper')
+_base_.optim_wrapper.type = 'AmpOptimWrapper'
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..4aa35b5d95f7f531cc2bdb8a03553ae197cfe727
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcn/metafile.yml
@@ -0,0 +1,272 @@
+Collections:
+ - Name: Deformable Convolutional Networks
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Deformable Convolution
+ Paper:
+ URL: https://arxiv.org/abs/1703.06211
+ Title: "Deformable Convolutional Networks"
+ README: configs/dcn/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/ops/dcn/deform_conv.py#L15
+ Version: v2.0.0
+
+Models:
+ - Name: faster-rcnn_r50_fpn_dconv_c3-c5_1x_coco
+ In Collection: Deformable Convolutional Networks
+ Config: configs/dcn/faster-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.0
+ inference time (ms/im):
+ - value: 56.18
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_r50_fpn_dconv_c3-c5_1x_coco/faster_rcnn_r50_fpn_dconv_c3-c5_1x_coco_20200130-d68aed1e.pth
+
+ - Name: faster-rcnn_r50_fpn_dpool_1x_coco
+ In Collection: Deformable Convolutional Networks
+ Config: configs/dcn/faster-rcnn_r50_fpn_dpool_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.0
+ inference time (ms/im):
+ - value: 58.14
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_r50_fpn_dpool_1x_coco/faster_rcnn_r50_fpn_dpool_1x_coco_20200307-90d3c01d.pth
+
+ - Name: faster-rcnn_r101-dconv-c3-c5_fpn_1x_coco
+ In Collection: Deformable Convolutional Networks
+ Config: configs/dcn/faster-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.0
+ inference time (ms/im):
+ - value: 80
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_r101_fpn_dconv_c3-c5_1x_coco/faster_rcnn_r101_fpn_dconv_c3-c5_1x_coco_20200203-1377f13d.pth
+
+ - Name: faster-rcnn_x101-32x4d-dconv-c3-c5_fpn_1x_coco
+ In Collection: Deformable Convolutional Networks
+ Config: configs/dcn/faster-rcnn_x101-32x4d-dconv-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.3
+ inference time (ms/im):
+ - value: 100
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_x101_32x4d_fpn_dconv_c3-c5_1x_coco/faster_rcnn_x101_32x4d_fpn_dconv_c3-c5_1x_coco_20200203-4f85c69c.pth
+
+ - Name: mask-rcnn_r50_fpn_dconv_c3-c5_1x_coco
+ In Collection: Deformable Convolutional Networks
+ Config: configs/dcn/mask-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.5
+ inference time (ms/im):
+ - value: 64.94
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/dcn/mask_rcnn_r50_fpn_dconv_c3-c5_1x_coco/mask_rcnn_r50_fpn_dconv_c3-c5_1x_coco_20200203-4d9ad43b.pth
+
+ - Name: mask-rcnn_r50_fpn_fp16_dconv_c3-c5_1x_coco
+ In Collection: Deformable Convolutional Networks
+ Config: configs/dcn/mask-rcnn_r50-dconv-c3-c5_fpn_amp-1x_coco.py
+ Metadata:
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ - Mixed Precision Training
+ Training Memory (GB): 3.0
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.9
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fp16/mask_rcnn_r50_fpn_fp16_dconv_c3-c5_1x_coco/mask_rcnn_r50_fpn_fp16_dconv_c3-c5_1x_coco_20210520_180247-c06429d2.pth
+
+ - Name: mask-rcnn_r101-dconv-c3-c5_fpn_1x_coco
+ In Collection: Deformable Convolutional Networks
+ Config: configs/dcn/mask-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.5
+ inference time (ms/im):
+ - value: 85.47
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/dcn/mask_rcnn_r101_fpn_dconv_c3-c5_1x_coco/mask_rcnn_r101_fpn_dconv_c3-c5_1x_coco_20200216-a71f5bce.pth
+
+ - Name: cascade-rcnn_r50_fpn_dconv_c3-c5_1x_coco
+ In Collection: Deformable Convolutional Networks
+ Config: configs/dcn/cascade-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.5
+ inference time (ms/im):
+ - value: 68.49
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/dcn/cascade_rcnn_r50_fpn_dconv_c3-c5_1x_coco/cascade_rcnn_r50_fpn_dconv_c3-c5_1x_coco_20200130-2f1fca44.pth
+
+ - Name: cascade-rcnn_r101-dconv-c3-c5_fpn_1x_coco
+ In Collection: Deformable Convolutional Networks
+ Config: configs/dcn/cascade-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.4
+ inference time (ms/im):
+ - value: 90.91
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/dcn/cascade_rcnn_r101_fpn_dconv_c3-c5_1x_coco/cascade_rcnn_r101_fpn_dconv_c3-c5_1x_coco_20200203-3b2f0594.pth
+
+ - Name: cascade-mask-rcnn_r50_fpn_dconv_c3-c5_1x_coco
+ In Collection: Deformable Convolutional Networks
+ Config: configs/dcn/cascade-mask-rcnn_r50-dconv-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.0
+ inference time (ms/im):
+ - value: 100
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/dcn/cascade_mask_rcnn_r50_fpn_dconv_c3-c5_1x_coco/cascade_mask_rcnn_r50_fpn_dconv_c3-c5_1x_coco_20200202-42e767a2.pth
+
+ - Name: cascade-mask-rcnn_r101-dconv-c3-c5_fpn_1x_coco
+ In Collection: Deformable Convolutional Networks
+ Config: configs/dcn/cascade-mask-rcnn_r101-dconv-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 8.0
+ inference time (ms/im):
+ - value: 116.28
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/dcn/cascade_mask_rcnn_r101_fpn_dconv_c3-c5_1x_coco/cascade_mask_rcnn_r101_fpn_dconv_c3-c5_1x_coco_20200204-df0c5f10.pth
+
+ - Name: cascade-mask-rcnn_x101-32x4d-dconv-c3-c5_fpn_1x_coco
+ In Collection: Deformable Convolutional Networks
+ Config: configs/dcn/cascade-mask-rcnn_x101-32x4d-dconv-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 9.2
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 47.3
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 41.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/dcn/cascade_mask_rcnn_x101_32x4d_fpn_dconv_c3-c5_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_dconv_c3-c5_1x_coco-e75f90c8.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..7f42c93401f836350c7b30cf5af9b4caa7ea75c7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/README.md
@@ -0,0 +1,37 @@
+# DCNv2
+
+> [Deformable ConvNets v2: More Deformable, Better Results](https://arxiv.org/abs/1811.11168)
+
+
+
+## Abstract
+
+The superior performance of Deformable Convolutional Networks arises from its ability to adapt to the geometric variations of objects. Through an examination of its adaptive behavior, we observe that while the spatial support for its neural features conforms more closely than regular ConvNets to object structure, this support may nevertheless extend well beyond the region of interest, causing features to be influenced by irrelevant image content. To address this problem, we present a reformulation of Deformable ConvNets that improves its ability to focus on pertinent image regions, through increased modeling power and stronger training. The modeling power is enhanced through a more comprehensive integration of deformable convolution within the network, and by introducing a modulation mechanism that expands the scope of deformation modeling. To effectively harness this enriched modeling capability, we guide network training via a proposed feature mimicking scheme that helps the network to learn features that reflect the object focus and classification power of RCNN features. With the proposed contributions, this new version of Deformable ConvNets yields significant performance gains over the original model and produces leading results on the COCO benchmark for object detection and instance segmentation.
+
+## Results and Models
+
+| Backbone | Model | Style | Conv | Pool | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :---------------: | :----: | :-----: | :-----------: | :----: | :-----: | :------: | :------------: | :----: | :-----: | :------------------------------------------------------------: | :-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | Faster | pytorch | mdconv(c3-c5) | - | 1x | 4.1 | 17.6 | 41.4 | | [config](./faster-rcnn_r50-mdconv-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_r50_fpn_mdconv_c3-c5_1x_coco/faster_rcnn_r50_fpn_mdconv_c3-c5_1x_coco_20200130-d099253b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_r50_fpn_mdconv_c3-c5_1x_coco/faster_rcnn_r50_fpn_mdconv_c3-c5_1x_coco_20200130_222144.log.json) |
+| \*R-50-FPN (dg=4) | Faster | pytorch | mdconv(c3-c5) | - | 1x | 4.2 | 17.4 | 41.5 | | [config](./faster-rcnn_r50-mdconv-group4-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_r50_fpn_mdconv_c3-c5_group4_1x_coco/faster_rcnn_r50_fpn_mdconv_c3-c5_group4_1x_coco_20200130-01262257.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_r50_fpn_mdconv_c3-c5_group4_1x_coco/faster_rcnn_r50_fpn_mdconv_c3-c5_group4_1x_coco_20200130_222058.log.json) |
+| R-50-FPN | Faster | pytorch | - | mdpool | 1x | 5.8 | 16.6 | 38.7 | | [config](./faster-rcnn_r50_fpn_mdpool_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_r50_fpn_mdpool_1x_coco/faster_rcnn_r50_fpn_mdpool_1x_coco_20200307-c0df27ff.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_r50_fpn_mdpool_1x_coco/faster_rcnn_r50_fpn_mdpool_1x_coco_20200307_203304.log.json) |
+| R-50-FPN | Mask | pytorch | mdconv(c3-c5) | - | 1x | 4.5 | 15.1 | 41.5 | 37.1 | [config](./mask-rcnn_r50-mdconv-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/dcn/mask_rcnn_r50_fpn_mdconv_c3-c5_1x_coco/mask_rcnn_r50_fpn_mdconv_c3-c5_1x_coco_20200203-ad97591f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/dcn/mask_rcnn_r50_fpn_mdconv_c3-c5_1x_coco/mask_rcnn_r50_fpn_mdconv_c3-c5_1x_coco_20200203_063443.log.json) |
+| R-50-FPN (FP16) | Mask | pytorch | mdconv(c3-c5) | - | 1x | 3.1 | | 42.0 | 37.6 | [config](./mask-rcnn_r50-mdconv-c3-c5_fpn_amp-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fp16/mask_rcnn_r50_fpn_fp16_mdconv_c3-c5_1x_coco/mask_rcnn_r50_fpn_fp16_mdconv_c3-c5_1x_coco_20210520_180434-cf8fefa5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fp16/mask_rcnn_r50_fpn_fp16_mdconv_c3-c5_1x_coco/mask_rcnn_r50_fpn_fp16_mdconv_c3-c5_1x_coco_20210520_180434.log.json) |
+
+**Notes:**
+
+- `mdconv` denotes modulated deformable convolution, `c3-c5` means adding dconv in resnet stage 3 to 5. `mdpool` denotes modulated deformable roi pooling.
+- The dcn ops are modified from https://github.com/chengdazhi/Deformable-Convolution-V2-PyTorch, which should be more memory efficient and slightly faster.
+- (\*) For R-50-FPN (dg=4), dg is short for deformable_group. This model is trained and tested on Amazon EC2 p3dn.24xlarge instance.
+- **Memory, Train/Inf time is outdated.**
+
+## Citation
+
+```latex
+@article{zhu2018deformable,
+ title={Deformable ConvNets v2: More Deformable, Better Results},
+ author={Zhu, Xizhou and Hu, Han and Lin, Stephen and Dai, Jifeng},
+ journal={arXiv preprint arXiv:1811.11168},
+ year={2018}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/faster-rcnn_r50-mdconv-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/faster-rcnn_r50-mdconv-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a7f7e4eecaf74418690975d54d09eeb0e31f9a1f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/faster-rcnn_r50-mdconv-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,5 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCNv2', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/faster-rcnn_r50-mdconv-group4-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/faster-rcnn_r50-mdconv-group4-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5c58dbed3782403a5fac3c6809598372e47cd72c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/faster-rcnn_r50-mdconv-group4-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,5 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCNv2', deform_groups=4, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/faster-rcnn_r50_fpn_mdpool_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/faster-rcnn_r50_fpn_mdpool_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6198d6d7d72f8d012c777330f1116b46b89290be
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/faster-rcnn_r50_fpn_mdpool_1x_coco.py
@@ -0,0 +1,12 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ roi_head=dict(
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(
+ _delete_=True,
+ type='ModulatedDeformRoIPoolPack',
+ output_size=7,
+ output_channels=256),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32])))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/mask-rcnn_r50-mdconv-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/mask-rcnn_r50-mdconv-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f7a90bbf31bea3663820caa4541de3ceafeb7366
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/mask-rcnn_r50-mdconv-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,5 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCNv2', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/mask-rcnn_r50-mdconv-c3-c5_fpn_amp-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/mask-rcnn_r50-mdconv-c3-c5_fpn_amp-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3b3894c2d61ee3208170235ba1aa98def79a7120
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/mask-rcnn_r50-mdconv-c3-c5_fpn_amp-1x_coco.py
@@ -0,0 +1,10 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCNv2', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)))
+
+# MMEngine support the following two ways, users can choose
+# according to convenience
+# optim_wrapper = dict(type='AmpOptimWrapper')
+_base_.optim_wrapper.type = 'AmpOptimWrapper'
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..dea7bfa1b531410f3c81693d7012a835781a63ca
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dcnv2/metafile.yml
@@ -0,0 +1,123 @@
+Collections:
+ - Name: Deformable Convolutional Networks v2
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Deformable Convolution
+ Paper:
+ URL: https://arxiv.org/abs/1811.11168
+ Title: "Deformable ConvNets v2: More Deformable, Better Results"
+ README: configs/dcnv2/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/ops/dcn/deform_conv.py#L15
+ Version: v2.0.0
+
+Models:
+ - Name: faster-rcnn_r50_fpn_mdconv_c3-c5_1x_coco
+ In Collection: Deformable Convolutional Networks v2
+ Config: configs/dcnv2/faster-rcnn_r50-mdconv-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.1
+ inference time (ms/im):
+ - value: 56.82
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_r50_fpn_mdconv_c3-c5_1x_coco/faster_rcnn_r50_fpn_mdconv_c3-c5_1x_coco_20200130-d099253b.pth
+
+ - Name: faster-rcnn_r50_fpn_mdconv_c3-c5_group4_1x_coco
+ In Collection: Deformable Convolutional Networks v2
+ Config: configs/dcnv2/faster-rcnn_r50-mdconv-group4-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.2
+ inference time (ms/im):
+ - value: 57.47
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_r50_fpn_mdconv_c3-c5_group4_1x_coco/faster_rcnn_r50_fpn_mdconv_c3-c5_group4_1x_coco_20200130-01262257.pth
+
+ - Name: faster-rcnn_r50_fpn_mdpool_1x_coco
+ In Collection: Deformable Convolutional Networks v2
+ Config: configs/dcnv2/faster-rcnn_r50_fpn_mdpool_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.8
+ inference time (ms/im):
+ - value: 60.24
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/dcn/faster_rcnn_r50_fpn_mdpool_1x_coco/faster_rcnn_r50_fpn_mdpool_1x_coco_20200307-c0df27ff.pth
+
+ - Name: mask-rcnn_r50_fpn_mdconv_c3-c5_1x_coco
+ In Collection: Deformable Convolutional Networks v2
+ Config: configs/dcnv2/mask-rcnn_r50-mdconv-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.5
+ inference time (ms/im):
+ - value: 66.23
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/dcn/mask_rcnn_r50_fpn_mdconv_c3-c5_1x_coco/mask_rcnn_r50_fpn_mdconv_c3-c5_1x_coco_20200203-ad97591f.pth
+
+ - Name: mask-rcnn_r50_fpn_fp16_mdconv_c3-c5_1x_coco
+ In Collection: Deformable Convolutional Networks v2
+ Config: configs/dcnv2/mask-rcnn_r50-mdconv-c3-c5_fpn_amp-1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.1
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ - Mixed Precision Training
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fp16/mask_rcnn_r50_fpn_fp16_mdconv_c3-c5_1x_coco/mask_rcnn_r50_fpn_fp16_mdconv_c3-c5_1x_coco_20210520_180434-cf8fefa5.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddod/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddod/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..d5ea9cd0cc11f7de0adf34aa4574bc20a8c11219
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddod/README.md
@@ -0,0 +1,31 @@
+# DDOD
+
+> [Disentangle Your Dense Object Detector](https://arxiv.org/pdf/2107.02963.pdf)
+
+
+
+## Abstract
+
+Deep learning-based dense object detectors have achieved great success in the past few years and have been applied to numerous multimedia applications such as video understanding. However, the current training pipeline for dense detectors is compromised to lots of conjunctions that may not hold. In this paper, we investigate three such important conjunctions: 1) only samples assigned as positive in classification head are used to train the regression head; 2) classification and regression share the same input feature and computational fields defined by the parallel head architecture; and 3) samples distributed in different feature pyramid layers are treated equally when computing the loss. We first carry out a series of pilot experiments to show disentangling such conjunctions can lead to persistent performance improvement. Then, based on these findings, we propose Disentangled Dense Object Detector(DDOD), in which simple and effective disentanglement mechanisms are designed and integrated into the current state-of-the-art dense object detectors. Extensive experiments on MS COCO benchmark show that our approach can lead to 2.0 mAP, 2.4 mAP and 2.2 mAP absolute improvements on RetinaNet, FCOS, and ATSS baselines with negligible extra overhead. Notably, our best model reaches 55.0 mAP on the COCO test-dev set and 93.5 AP on the hard subset of WIDER FACE, achieving new state-of-the-art performance on these two competitive benchmarks. Code is available at https://github.com/zehuichen123/DDOD.
+
+
+

+
+
+## Results and Models
+
+| Model | Backbone | Style | Lr schd | Mem (GB) | box AP | Config | Download |
+| :-------: | :------: | :-----: | :-----: | :------: | :----: | :---------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| DDOD-ATSS | R-50 | pytorch | 1x | 3.4 | 41.7 | [config](./ddod_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/ddod/ddod_r50_fpn_1x_coco/ddod_r50_fpn_1x_coco_20220523_223737-29b2fc67.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/ddod/ddod_r50_fpn_1x_coco/ddod_r50_fpn_1x_coco_20220523_223737.log.json) |
+
+## Citation
+
+```latex
+@inproceedings{chen2021disentangle,
+title={Disentangle Your Dense Object Detector},
+author={Chen, Zehui and Yang, Chenhongyi and Li, Qiaofei and Zhao, Feng and Zha, Zheng-Jun and Wu, Feng},
+booktitle={Proceedings of the 29th ACM International Conference on Multimedia},
+pages={4939--4948},
+year={2021}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddod/ddod_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddod/ddod_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..fed1116b1f92e613517a57aa196839e4de3037dc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddod/ddod_r50_fpn_1x_coco.py
@@ -0,0 +1,72 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ type='DDOD',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output',
+ num_outs=5),
+ bbox_head=dict(
+ type='DDODHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ octave_base_scale=8,
+ scales_per_octave=1,
+ strides=[8, 16, 32, 64, 128]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=2.0),
+ loss_iou=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0)),
+ train_cfg=dict(
+ # assigner is mean cls_assigner
+ assigner=dict(type='ATSSAssigner', topk=9, alpha=0.8),
+ reg_assigner=dict(type='ATSSAssigner', topk=9, alpha=0.5),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddod/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddod/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..c22395002bd614cd0e75d753320c3f9e7ce54bd1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddod/metafile.yml
@@ -0,0 +1,33 @@
+Collections:
+ - Name: DDOD
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - DDOD
+ - FPN
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/pdf/2107.02963.pdf
+ Title: 'Disentangle Your Dense Object Detector'
+ README: configs/ddod/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.25.0/mmdet/models/detectors/ddod.py#L6
+ Version: v2.25.0
+
+Models:
+ - Name: ddod_r50_fpn_1x_coco
+ In Collection: DDOD
+ Config: configs/ddod/ddod_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.4
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/ddod/ddod_r50_fpn_1x_coco/ddod_r50_fpn_1x_coco_20220523_223737-29b2fc67.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddq/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddq/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..3f6f459cbbb48c50d5fbd6abec3c6dbda4d422b4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddq/README.md
@@ -0,0 +1,39 @@
+# DDQ
+
+> [Dense Distinct Query for End-to-End Object Detection](https://arxiv.org/abs/2303.12776)
+
+
+
+## Abstract
+
+
+
+One-to-one label assignment in object detection has successfully obviated the need for non-maximum suppression (NMS) as postprocessing and makes the pipeline end-to-end. However, it triggers a new dilemma as the widely used sparse queries cannot guarantee a high recall, while dense queries inevitably bring more similar queries and encounter optimization difficulties. As both sparse and dense queries are problematic, then what are the expected queries in end-to-end object detection? This paper shows that the solution should be Dense Distinct Queries (DDQ). Concretely, we first lay dense queries like traditional detectors and then select distinct ones for one-to-one assignments. DDQ blends the advantages of traditional and recent end-to-end detectors and significantly improves the performance of various detectors including FCN, R-CNN, and DETRs. Most impressively, DDQ-DETR achieves 52.1 AP on MS-COCO dataset within 12 epochs using a ResNet-50 backbone, outperforming all existing detectors in the same setting. DDQ also shares the benefit of end-to-end detectors in crowded scenes and achieves 93.8 AP on CrowdHuman. We hope DDQ can inspire researchers to consider the complementarity between traditional methods and end-to-end detectors.
+
+
+
+## Results and Models
+
+| Model | Backbone | Lr schd | Augmentation | box AP(val) | Config | Download |
+| :---------------: | :------: | :-----: | :----------: | :---------: | :------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| DDQ DETR-4scale | R-50 | 12e | DETR | 51.4 | [config](./ddq-detr-4scale_r50_8xb2-12e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/ddq/ddq-detr-4scale_r50_8xb2-12e_coco/ddq-detr-4scale_r50_8xb2-12e_coco_20230809_170711-42528127.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/ddq/ddq-detr-4scale_r50_8xb2-12e_coco/ddq-detr-4scale_r50_8xb2-12e_coco_20230809_170711.log.json) |
+| DDQ DETR-5scale\* | R-50 | 12e | DETR | 52.1 | [config](./ddq-detr-5scale_r50_8xb2-12e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/ddq/ddq_detr_5scale_coco_1x.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/ddq/ddq_detr_5scale_coco_1x_20230319_103307.log) |
+| DDQ DETR-4scale\* | Swin-L | 30e | DETR | 58.7 | [config](./ddq-detr-4scale_swinl_8xb2-30e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/ddq/ddq_detr_swinl_30e.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/ddq/ddq_detr_swinl_30e_20230316_221721_20230318_143554.log) |
+
+**Note**
+
+- Models labeled * are not trained by us, but from [DDQ official website](https://github.com/jshilong/DDQ).
+- We find that the performance is unstable and may fluctuate by about 0.2 mAP.
+
+## Citation
+
+```latex
+@InProceedings{Zhang_2023_CVPR,
+ author = {Zhang, Shilong and Wang, Xinjiang and Wang, Jiaqi and Pang, Jiangmiao and Lyu, Chengqi and Zhang, Wenwei and Luo, Ping and Chen, Kai},
+ title = {Dense Distinct Query for End-to-End Object Detection},
+ booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
+ month = {June},
+ year = {2023},
+ pages = {7329-7338}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddq/ddq-detr-4scale_r50_8xb2-12e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddq/ddq-detr-4scale_r50_8xb2-12e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5e64afc087e1ed68b8b5d1474127c832f893cb9b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddq/ddq-detr-4scale_r50_8xb2-12e_coco.py
@@ -0,0 +1,170 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ type='DDQDETR',
+ num_queries=900, # num_matching_queries
+ # ratio of num_dense queries to num_queries
+ dense_topk_ratio=1.5,
+ with_box_refine=True,
+ as_two_stage=True,
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=1),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='ChannelMapper',
+ in_channels=[512, 1024, 2048],
+ kernel_size=1,
+ out_channels=256,
+ act_cfg=None,
+ norm_cfg=dict(type='GN', num_groups=32),
+ num_outs=4),
+ # encoder class name: DeformableDetrTransformerEncoder
+ encoder=dict(
+ num_layers=6,
+ layer_cfg=dict(
+ self_attn_cfg=dict(embed_dims=256, num_levels=4,
+ dropout=0.0), # 0.1 for DeformDETR
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048, # 1024 for DeformDETR
+ ffn_drop=0.0))), # 0.1 for DeformDETR
+ # decoder class name: DDQTransformerDecoder
+ decoder=dict(
+ # `num_layers` >= 2, because attention masks of the last
+ # `num_layers` - 1 layers are used for distinct query selection
+ num_layers=6,
+ return_intermediate=True,
+ layer_cfg=dict(
+ self_attn_cfg=dict(embed_dims=256, num_heads=8,
+ dropout=0.0), # 0.1 for DeformDETR
+ cross_attn_cfg=dict(embed_dims=256, num_levels=4,
+ dropout=0.0), # 0.1 for DeformDETR
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048, # 1024 for DeformDETR
+ ffn_drop=0.0)), # 0.1 for DeformDETR
+ post_norm_cfg=None),
+ positional_encoding=dict(
+ num_feats=128,
+ normalize=True,
+ offset=0.0, # -0.5 for DeformDETR
+ temperature=20), # 10000 for DeformDETR
+ bbox_head=dict(
+ type='DDQDETRHead',
+ num_classes=80,
+ sync_cls_avg_factor=True,
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=5.0),
+ loss_iou=dict(type='GIoULoss', loss_weight=2.0)),
+ dn_cfg=dict(
+ label_noise_scale=0.5,
+ box_noise_scale=1.0,
+ group_cfg=dict(dynamic=True, num_groups=None, num_dn_queries=100)),
+ dqs_cfg=dict(type='nms', iou_threshold=0.8),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='HungarianAssigner',
+ match_costs=[
+ dict(type='FocalLossCost', weight=2.0),
+ dict(type='BBoxL1Cost', weight=5.0, box_format='xywh'),
+ dict(type='IoUCost', iou_mode='giou', weight=2.0)
+ ])),
+ test_cfg=dict(max_per_img=300))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ filter_cfg=dict(filter_empty_gt=False), pipeline=train_pipeline))
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0002, weight_decay=0.05),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(custom_keys={'backbone': dict(lr_mult=0.1)}))
+
+# learning policy
+max_epochs = 12
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=0.0001,
+ by_epoch=False,
+ begin=0,
+ end=2000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[11],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddq/ddq-detr-4scale_swinl_8xb2-30e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddq/ddq-detr-4scale_swinl_8xb2-30e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d863649411e3157373961b3da339990df1e6f267
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddq/ddq-detr-4scale_swinl_8xb2-30e_coco.py
@@ -0,0 +1,177 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py', '../_base_/default_runtime.py'
+]
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_large_patch4_window12_384_22k.pth' # noqa: E501
+model = dict(
+ type='DDQDETR',
+ num_queries=900, # num_matching_queries
+ # ratio of num_dense queries to num_queries
+ dense_topk_ratio=1.5,
+ with_box_refine=True,
+ as_two_stage=True,
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=1),
+ backbone=dict(
+ type='SwinTransformer',
+ pretrain_img_size=384,
+ embed_dims=192,
+ depths=[2, 2, 18, 2],
+ num_heads=[6, 12, 24, 48],
+ window_size=12,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.2,
+ patch_norm=True,
+ out_indices=(1, 2, 3),
+ with_cp=False,
+ convert_weights=True,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ neck=dict(
+ type='ChannelMapper',
+ in_channels=[384, 768, 1536],
+ kernel_size=1,
+ out_channels=256,
+ act_cfg=None,
+ norm_cfg=dict(type='GN', num_groups=32),
+ num_outs=4),
+ # encoder class name: DeformableDetrTransformerEncoder
+ encoder=dict(
+ num_layers=6,
+ layer_cfg=dict(
+ self_attn_cfg=dict(embed_dims=256, num_levels=4,
+ dropout=0.0), # 0.1 for DeformDETR
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048, # 1024 for DeformDETR
+ ffn_drop=0.0))), # 0.1 for DeformDETR
+ # decoder class name: DDQTransformerDecoder
+ decoder=dict(
+ num_layers=6,
+ return_intermediate=True,
+ layer_cfg=dict(
+ self_attn_cfg=dict(embed_dims=256, num_heads=8,
+ dropout=0.0), # 0.1 for DeformDETR
+ cross_attn_cfg=dict(embed_dims=256, num_levels=4,
+ dropout=0.0), # 0.1 for DeformDETR
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048, # 1024 for DeformDETR
+ ffn_drop=0.0)), # 0.1 for DeformDETR
+ post_norm_cfg=None),
+ positional_encoding=dict(
+ num_feats=128,
+ normalize=True,
+ offset=0.0, # -0.5 for DeformDETR
+ temperature=20), # 10000 for DeformDETR
+ bbox_head=dict(
+ type='DDQDETRHead',
+ num_classes=80,
+ sync_cls_avg_factor=True,
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=5.0),
+ loss_iou=dict(type='GIoULoss', loss_weight=2.0)),
+ dn_cfg=dict(
+ label_noise_scale=0.5,
+ box_noise_scale=1.0,
+ group_cfg=dict(dynamic=True, num_groups=None, num_dn_queries=100)),
+ dqs_cfg=dict(type='nms', iou_threshold=0.8),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='HungarianAssigner',
+ match_costs=[
+ dict(type='FocalLossCost', weight=2.0),
+ dict(type='BBoxL1Cost', weight=5.0, box_format='xywh'),
+ dict(type='IoUCost', iou_mode='giou', weight=2.0)
+ ])),
+ test_cfg=dict(max_per_img=300))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ filter_cfg=dict(filter_empty_gt=False), pipeline=train_pipeline))
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0002, weight_decay=0.05),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(custom_keys={'backbone': dict(lr_mult=0.05)}))
+
+# learning policy
+max_epochs = 30
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=0.0001,
+ by_epoch=False,
+ begin=0,
+ end=2000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[20, 26],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddq/ddq-detr-5scale_r50_8xb2-12e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddq/ddq-detr-5scale_r50_8xb2-12e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3c38f553bdd46bc4e0611bbd0fd4bab0c1929825
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddq/ddq-detr-5scale_r50_8xb2-12e_coco.py
@@ -0,0 +1,171 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ type='DDQDETR',
+ num_queries=900, # num_matching_queries
+ # ratio of num_dense queries to num_queries
+ dense_topk_ratio=1.5,
+ with_box_refine=True,
+ as_two_stage=True,
+ num_feature_levels=5,
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=1),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='ChannelMapper',
+ in_channels=[256, 512, 1024, 2048],
+ kernel_size=1,
+ out_channels=256,
+ act_cfg=None,
+ norm_cfg=dict(type='GN', num_groups=32),
+ num_outs=5),
+ # encoder class name: DeformableDetrTransformerEncoder
+ encoder=dict(
+ num_layers=6,
+ layer_cfg=dict(
+ self_attn_cfg=dict(embed_dims=256, num_levels=5,
+ dropout=0.0), # 0.1 for DeformDETR
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048, # 1024 for DeformDETR
+ ffn_drop=0.0))), # 0.1 for DeformDETR
+ # decoder class name: DDQTransformerDecoder
+ decoder=dict(
+ num_layers=6,
+ return_intermediate=True,
+ layer_cfg=dict(
+ self_attn_cfg=dict(embed_dims=256, num_heads=8,
+ dropout=0.0), # 0.1 for DeformDETR
+ cross_attn_cfg=dict(embed_dims=256, num_levels=5,
+ dropout=0.0), # 0.1 for DeformDETR
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048, # 1024 for DeformDETR
+ ffn_drop=0.0)), # 0.1 for DeformDETR
+ post_norm_cfg=None),
+ positional_encoding=dict(
+ num_feats=128,
+ normalize=True,
+ offset=0.0, # -0.5 for DeformDETR
+ temperature=20), # 10000 for DeformDETR
+ bbox_head=dict(
+ type='DDQDETRHead',
+ num_classes=80,
+ sync_cls_avg_factor=True,
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=5.0),
+ loss_iou=dict(type='GIoULoss', loss_weight=2.0)),
+ dn_cfg=dict(
+ label_noise_scale=0.5,
+ box_noise_scale=1.0,
+ group_cfg=dict(dynamic=True, num_groups=None, num_dn_queries=100)),
+ dqs_cfg=dict(type='nms', iou_threshold=0.8),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='HungarianAssigner',
+ match_costs=[
+ dict(type='FocalLossCost', weight=2.0),
+ dict(type='BBoxL1Cost', weight=5.0, box_format='xywh'),
+ dict(type='IoUCost', iou_mode='giou', weight=2.0)
+ ])),
+ test_cfg=dict(max_per_img=300))
+
+# train_pipeline, NOTE the img_scale and the Pad's size_divisor is different
+# from the default setting in mmdet.
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ filter_cfg=dict(filter_empty_gt=False), pipeline=train_pipeline))
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0002, weight_decay=0.05),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(custom_keys={'backbone': dict(lr_mult=0.1)}))
+
+# learning policy
+max_epochs = 12
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=0.0001,
+ by_epoch=False,
+ begin=0,
+ end=2000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[11],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddq/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddq/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..bd33abe1a5122885913a1e8cbee60cb48014239f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ddq/metafile.yml
@@ -0,0 +1,56 @@
+Collections:
+ - Name: DDQ
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - AdamW
+ - Multi Scale Train
+ - Gradient Clip
+ Training Resources: 8x A100 GPUs
+ Architecture:
+ - ResNet
+ - Transformer
+ Paper:
+ URL: https://arxiv.org/abs/2303.12776
+ Title: 'Dense Distinct Query for End-to-End Object Detection'
+ README: configs/ddq/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/dev-3.x/mmdet/models/detectors/ddq_detr.py#L21
+ Version: dev-3.x
+
+Models:
+ - Name: ddq-detr-4scale_r50_8xb2-12e_coco
+ In Collection: DDQ
+ Config: configs/ddq/ddq-detr-4scale_r50_8xb2-12e_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 51.4
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/ddq/ddq-detr-4scale_r50_8xb2-12e_coco/ddq-detr-4scale_r50_8xb2-12e_coco_20230809_170711-42528127.pth
+
+ - Name: ddq-detr-5scale_r50_8xb2-12e_coco
+ In Collection: DDQ
+ Config: configs/dino/ddq-detr-5scale_r50_8xb2-12e_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 52.1
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/ddq/ddq_detr_5scale_coco_1x.pth
+
+ - Name: ddq-detr-4scale_swinl_8xb2-30e_coco
+ In Collection: DDQ
+ Config: configs/dino/ddq-detr-4scale_swinl_8xb2-30e_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 58.7
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/ddq/ddq_detr_swinl_30e.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/deepfashion/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deepfashion/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..844e29d6a72906bc36fd682df270480af5a595c0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deepfashion/README.md
@@ -0,0 +1,70 @@
+# DeepFashion
+
+> [DeepFashion: Powering Robust Clothes Recognition and Retrieval With Rich Annotations](https://openaccess.thecvf.com/content_cvpr_2016/html/Liu_DeepFashion_Powering_Robust_CVPR_2016_paper.html)
+
+
+
+## Abstract
+
+Recent advances in clothes recognition have been driven by the construction of clothes datasets. Existing datasets are limited in the amount of annotations and are difficult to cope with the various challenges in real-world applications. In this work, we introduce DeepFashion, a large-scale clothes dataset with comprehensive annotations. It contains over 800,000 images, which are richly annotated with massive attributes, clothing landmarks, and correspondence of images taken under different scenarios including store, street snapshot, and consumer. Such rich annotations enable the development of powerful algorithms in clothes recognition and facilitating future researches. To demonstrate the advantages of DeepFashion, we propose a new deep model, namely FashionNet, which learns clothing features by jointly predicting clothing attributes and landmarks. The estimated landmarks are then employed to pool or gate the learned features. It is optimized in an iterative manner. Extensive experiments demonstrate the effectiveness of FashionNet and the usefulness of DeepFashion.
+
+
+

+
+
+## Introduction
+
+[MMFashion](https://github.com/open-mmlab/mmfashion) develops "fashion parsing and segmentation" module
+based on the dataset
+[DeepFashion-Inshop](https://drive.google.com/drive/folders/0B7EVK8r0v71pVDZFQXRsMDZCX1E?usp=sharing).
+Its annotation follows COCO style.
+To use it, you need to first download the data. Note that we only use "img_highres" in this task.
+The file tree should be like this:
+
+```sh
+mmdetection
+├── mmdet
+├── tools
+├── configs
+├── data
+│ ├── DeepFashion
+│ │ ├── In-shop
+| │ │ ├── Anno
+| │ │ │ ├── segmentation
+| │ │ │ | ├── DeepFashion_segmentation_train.json
+| │ │ │ | ├── DeepFashion_segmentation_query.json
+| │ │ │ | ├── DeepFashion_segmentation_gallery.json
+| │ │ │ ├── list_bbox_inshop.txt
+| │ │ │ ├── list_description_inshop.json
+| │ │ │ ├── list_item_inshop.txt
+| │ │ │ └── list_landmarks_inshop.txt
+| │ │ ├── Eval
+| │ │ │ └── list_eval_partition.txt
+| │ │ ├── Img
+| │ │ │ ├── img
+| │ │ │ │ ├──XXX.jpg
+| │ │ │ ├── img_highres
+| │ │ │ └── ├──XXX.jpg
+
+```
+
+After that you can train the Mask RCNN r50 on DeepFashion-In-shop dataset by launching training with the `mask_rcnn_r50_fpn_1x.py` config
+or creating your own config file.
+
+## Results and Models
+
+| Backbone | Model type | Dataset | bbox detection Average Precision | segmentation Average Precision | Config | Download (Google) |
+| :------: | :--------: | :-----------------: | :------------------------------: | :----------------------------: | :----------------------------------------------: | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| ResNet50 | Mask RCNN | DeepFashion-In-shop | 0.599 | 0.584 | [config](./mask-rcnn_r50_fpn_15e_deepfashion.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/deepfashion/mask_rcnn_r50_fpn_15e_deepfashion/mask_rcnn_r50_fpn_15e_deepfashion_20200329_192752.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/deepfashion/mask_rcnn_r50_fpn_15e_deepfashion/20200329_192752.log.json) |
+
+## Citation
+
+```latex
+@inproceedings{liuLQWTcvpr16DeepFashion,
+ author = {Liu, Ziwei and Luo, Ping and Qiu, Shi and Wang, Xiaogang and Tang, Xiaoou},
+ title = {DeepFashion: Powering Robust Clothes Recognition and Retrieval with Rich Annotations},
+ booktitle = {Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR)},
+ month = {June},
+ year = {2016}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/deepfashion/mask-rcnn_r50_fpn_15e_deepfashion.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deepfashion/mask-rcnn_r50_fpn_15e_deepfashion.py
new file mode 100644
index 0000000000000000000000000000000000000000..403b18a4ca8ed61aedcb99218ecc79302826ff8c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deepfashion/mask-rcnn_r50_fpn_15e_deepfashion.py
@@ -0,0 +1,23 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/deepfashion.py', '../_base_/schedules/schedule_1x.py',
+ '../_base_/default_runtime.py'
+]
+model = dict(
+ roi_head=dict(
+ bbox_head=dict(num_classes=15), mask_head=dict(num_classes=15)))
+# runtime settings
+max_epochs = 15
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/deepsort/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deepsort/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..e50ec17eb55ffef4fb59dae43175b2688eedfaa9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deepsort/README.md
@@ -0,0 +1,109 @@
+# Simple online and realtime tracking with a deep association metric
+
+## Abstract
+
+
+
+Simple Online and Realtime Tracking (SORT) is a pragmatic approach to multiple object tracking with a focus on simple, effective algorithms. In this paper, we integrate appearance information to improve the performance of SORT. Due to this extension we are able to track objects through longer periods of occlusions, effectively reducing the number of identity switches. In spirit of the original framework we place much of the computational complexity into an offline pre-training stage where we learn a deep association metric on a largescale person re-identification dataset. During online application, we establish measurement-to-track associations using nearest neighbor queries in visual appearance space. Experimental evaluation shows that our extensions reduce the number of identity switches by 45%, achieving overall competitive performance at high frame rates.
+
+
+
+
+

+
+
+## Results and models on MOT17
+
+Currently we do not support training ReID models for DeepSORT.
+We directly use the ReID model from [Tracktor](https://github.com/phil-bergmann/tracking_wo_bnw). These missed features will be supported in the future.
+
+| Method | Detector | ReID | Train Set | Test Set | Public | Inf time (fps) | HOTA | MOTA | IDF1 | FP | FN | IDSw. | Config | Download |
+| :------: | :----------------: | :--: | :--------: | :------: | :----: | :------------: | :--: | :--: | :--: | :---: | :---: | :---: | :--------------------------------------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| DeepSORT | R50-FasterRCNN-FPN | R50 | half-train | half-val | N | 13.8 | 57.0 | 63.7 | 69.5 | 15063 | 40323 | 3276 | [config](deepsort_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py) | [detector](https://download.openmmlab.com/mmtracking/mot/faster_rcnn/faster-rcnn_r50_fpn_4e_mot17-half-64ee2ed4.pth) [reid](https://download.openmmlab.com/mmtracking/mot/reid/tracktor_reid_r50_iter25245-a452f51f.pth) |
+
+## Get started
+
+### 1. Development Environment Setup
+
+Tracking Development Environment Setup can refer to this [document](../../docs/en/get_started.md).
+
+### 2. Dataset Prepare
+
+Tracking Dataset Prepare can refer to this [document](../../docs/en/user_guides/tracking_dataset_prepare.md).
+
+### 3. Training
+
+We implement DeepSORT with independent detector and ReID models.
+Note that, due to the influence of parameters such as learning rate in default configuration file,
+we recommend using 8 GPUs for training in order to reproduce accuracy.
+
+You can train the detector as follows.
+
+```shell script
+# Training Faster R-CNN on mot17-half-train dataset with following command.
+# The number after config file represents the number of GPUs used. Here we use 8 GPUs.
+bash tools/dist_train.sh configs/sort/faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py 8
+```
+
+If you want to know about more detailed usage of `train.py/dist_train.sh/slurm_train.sh`,
+please refer to this [document](../../docs/en/user_guides/tracking_train_test.md).
+
+### 4. Testing and evaluation
+
+### 4.1 Example on MOTxx-halfval dataset
+
+**4.1.1 use separate trained detector and reid model to evaluating and testing**
+
+```shell
+# Example 1: Test on motXX-half-val set.
+# The number after config file represents the number of GPUs used. Here we use 8 GPUs.
+bash tools/dist_test_tracking.sh configs/deepsort/deepsort_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py 8 --detector ${DETECTOR_CHECKPOINT_PATH} --reid ${REID_CHECKPOINT_PATH}
+```
+
+**4.1.2 use video_baesd to evaluating and testing**
+
+we also provide two_ways(img_based or video_based) to evaluating and testing.
+if you want to use video_based to evaluating and testing, you can modify config as follows
+
+```
+val_dataloader = dict(
+ sampler=dict(type='DefaultSampler', shuffle=False, round_up=False))
+```
+
+### 4.2 Example on MOTxx-test dataset
+
+If you want to get the results of the [MOT Challenge](https://motchallenge.net/) test set,
+please use the following command to generate result files that can be used for submission.
+It will be stored in `./mot_17_test_res`, you can modify the saved path in `test_evaluator` of the config.
+
+```shell script
+# Example 2: Test on motxx-test set
+# The number after config file represents the number of GPUs used
+bash tools/dist_test_tracking.sh configs/deepsort/deepsort_faster-rcnn_r50_fpn_8xb2-4e_mot17train_test-mot17test 8 --detector ${DETECTOR_CHECKPOINT_PATH} --reid ${REID_CHECKPOINT_PATH}
+```
+
+If you want to know about more detailed usage of `test_tracking.py/dist_test_tracking.sh/slurm_test_tracking.sh`,
+please refer to this [document](../../docs/en/user_guides/tracking_train_test.md).
+
+### 5.Inference
+
+Use a single GPU to predict a video and save it as a video.
+
+```shell
+python demo/mot_demo.py demo/demo_mot.mp4 configs/deepsort/deepsort_faster-rcnn_r50_fpn_8xb2-4e_mot17train_test-mot17test --detector ${DETECTOR_CHECKPOINT_PATH} --reid ${REID_CHECKPOINT_PATH} --out mot.mp4
+```
+
+## Citation
+
+
+
+```latex
+@inproceedings{wojke2017simple,
+ title={Simple online and realtime tracking with a deep association metric},
+ author={Wojke, Nicolai and Bewley, Alex and Paulus, Dietrich},
+ booktitle={2017 IEEE international conference on image processing (ICIP)},
+ pages={3645--3649},
+ year={2017},
+ organization={IEEE}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/deepsort/deepsort_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deepsort/deepsort_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py
new file mode 100644
index 0000000000000000000000000000000000000000..70d3393829b422740bfba5d1746c7651e9c2d69c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deepsort/deepsort_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py
@@ -0,0 +1,85 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/mot_challenge.py', '../_base_/default_runtime.py'
+]
+
+default_hooks = dict(
+ logger=dict(type='LoggerHook', interval=1),
+ visualization=dict(type='TrackVisualizationHook', draw=False))
+
+vis_backends = [dict(type='LocalVisBackend')]
+visualizer = dict(
+ type='TrackLocalVisualizer', vis_backends=vis_backends, name='visualizer')
+# custom hooks
+custom_hooks = [
+ # Synchronize model buffers such as running_mean and running_var in BN
+ # at the end of each epoch
+ dict(type='SyncBuffersHook')
+]
+
+detector = _base_.model
+detector.pop('data_preprocessor')
+detector.rpn_head.bbox_coder.update(dict(clip_border=False))
+detector.roi_head.bbox_head.update(dict(num_classes=1))
+detector.roi_head.bbox_head.bbox_coder.update(dict(clip_border=False))
+detector['init_cfg'] = dict(
+ type='Pretrained',
+ checkpoint= # noqa: E251
+ 'https://download.openmmlab.com/mmtracking/mot/faster_rcnn/'
+ 'faster-rcnn_r50_fpn_4e_mot17-half-64ee2ed4.pth')
+del _base_.model
+
+model = dict(
+ type='DeepSORT',
+ data_preprocessor=dict(
+ type='TrackDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ detector=detector,
+ reid=dict(
+ type='BaseReID',
+ data_preprocessor=dict(type='mmpretrain.ClsDataPreprocessor'),
+ backbone=dict(
+ type='mmpretrain.ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(3, ),
+ style='pytorch'),
+ neck=dict(type='GlobalAveragePooling', kernel_size=(8, 4), stride=1),
+ head=dict(
+ type='LinearReIDHead',
+ num_fcs=1,
+ in_channels=2048,
+ fc_channels=1024,
+ out_channels=128,
+ num_classes=380,
+ loss_cls=dict(type='mmpretrain.CrossEntropyLoss', loss_weight=1.0),
+ loss_triplet=dict(type='TripletLoss', margin=0.3, loss_weight=1.0),
+ norm_cfg=dict(type='BN1d'),
+ act_cfg=dict(type='ReLU')),
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint= # noqa: E251
+ 'https://download.openmmlab.com/mmtracking/mot/reid/tracktor_reid_r50_iter25245-a452f51f.pth' # noqa: E501
+ )),
+ tracker=dict(
+ type='SORTTracker',
+ motion=dict(type='KalmanFilter', center_only=False),
+ obj_score_thr=0.5,
+ reid=dict(
+ num_samples=10,
+ img_scale=(256, 128),
+ img_norm_cfg=None,
+ match_score_thr=2.0),
+ match_iou_thr=0.5,
+ momentums=None,
+ num_tentatives=2,
+ num_frames_retain=100))
+
+train_dataloader = None
+
+train_cfg = None
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/deepsort/deepsort_faster-rcnn_r50_fpn_8xb2-4e_mot17train_test-mot17test.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deepsort/deepsort_faster-rcnn_r50_fpn_8xb2-4e_mot17train_test-mot17test.py
new file mode 100644
index 0000000000000000000000000000000000000000..687ce7adfcc1742bab75cca939a99df37b43689c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deepsort/deepsort_faster-rcnn_r50_fpn_8xb2-4e_mot17train_test-mot17test.py
@@ -0,0 +1,15 @@
+_base_ = [
+ './deepsort_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain'
+ '_test-mot17halfval.py'
+]
+
+# dataloader
+val_dataloader = dict(
+ dataset=dict(ann_file='annotations/train_cocoformat.json'))
+test_dataloader = dict(
+ dataset=dict(
+ ann_file='annotations/test_cocoformat.json',
+ data_prefix=dict(img_path='test')))
+
+# evaluator
+test_evaluator = dict(format_only=True, outfile_prefix='./mot_17_test_res')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/deepsort/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deepsort/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..2feb358e93d1590f0305e2ed08ae40e18bbd6cb9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deepsort/metafile.yml
@@ -0,0 +1,37 @@
+Collections:
+ - Name: DeepSORT
+ Metadata:
+ Training Techniques:
+ - SGD with Momentum
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNet
+ - FPN
+ Paper:
+ URL: https://arxiv.org/abs/1703.07402
+ Title: Simple Online and Realtime Tracking with a Deep Association Metric
+ README: configs/deepsort/README.md
+
+Models:
+ - Name: deepsort_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval
+ In Collection: DeepSORT
+ Config: configs/deepsort/deepsort_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py
+ Metadata:
+ Training Data: MOT17-half-train
+ inference time (ms/im):
+ - value: 72.5
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (640, 1088)
+ Results:
+ - Task: Multiple Object Tracking
+ Dataset: MOT17-half-val
+ Metrics:
+ MOTA: 63.7
+ IDF1: 69.5
+ HOTA: 57.0
+ Weights:
+ - https://download.openmmlab.com/mmtracking/mot/faster_rcnn/faster-rcnn_r50_fpn_4e_mot17-half-64ee2ed4.pth
+ - https://download.openmmlab.com/mmtracking/mot/reid/tracktor_reid_r50_iter25245-a452f51f.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/deformable_detr/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deformable_detr/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..ca897cdb4cfc17b1d194d2aeaba7feea388839f0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deformable_detr/README.md
@@ -0,0 +1,41 @@
+# Deformable DETR
+
+> [Deformable DETR: Deformable Transformers for End-to-End Object Detection](https://arxiv.org/abs/2010.04159)
+
+
+
+## Abstract
+
+DETR has been recently proposed to eliminate the need for many hand-designed components in object detection while demonstrating good performance. However, it suffers from slow convergence and limited feature spatial resolution, due to the limitation of Transformer attention modules in processing image feature maps. To mitigate these issues, we proposed Deformable DETR, whose attention modules only attend to a small set of key sampling points around a reference. Deformable DETR can achieve better performance than DETR (especially on small objects) with 10 times less training epochs. Extensive experiments on the COCO benchmark demonstrate the effectiveness of our approach.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Model | Lr schd | box AP | Config | Download |
+| :------: | :---------------------------------: | :-----: | :----: | :---------------------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | Deformable DETR | 50e | 44.3 | [config](./deformable-detr_r50_16xb2-50e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/deformable_detr/deformable-detr_r50_16xb2-50e_coco/deformable-detr_r50_16xb2-50e_coco_20221029_210934-6bc7d21b.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/deformable_detr/deformable-detr_r50_16xb2-50e_coco/deformable-detr_r50_16xb2-50e_coco_20221029_210934.log.json) |
+| R-50 | + iterative bounding box refinement | 50e | 46.2 | [config](./deformable-detr-refine_r50_16xb2-50e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/deformable_detr/deformable-detr-refine_r50_16xb2-50e_coco/deformable-detr-refine_r50_16xb2-50e_coco_20221022_225303-844e0f93.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/deformable_detr/deformable-detr-refine_r50_16xb2-50e_coco/deformable-detr-refine_r50_16xb2-50e_coco_20221022_225303.log.json) |
+| R-50 | ++ two-stage Deformable DETR | 50e | 47.0 | [config](./deformable-detr-refine-twostage_r50_16xb2-50e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/deformable_detr/deformable-detr-refine-twostage_r50_16xb2-50e_coco/deformable-detr-refine-twostage_r50_16xb2-50e_coco_20221021_184714-acc8a5ff.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/deformable_detr/deformable-detr-refine-twostage_r50_16xb2-50e_coco/deformable-detr-refine-twostage_r50_16xb2-50e_coco_20221021_184714.log.json) |
+
+### NOTE
+
+1. All models are trained with batch size 32.
+2. The performance is unstable. `Deformable DETR` and `iterative bounding box refinement` may fluctuate about 0.3 mAP. `two-stage Deformable DETR` may fluctuate about 0.2 mAP.
+
+## Citation
+
+We provide the config files for Deformable DETR: [Deformable DETR: Deformable Transformers for End-to-End Object Detection](https://arxiv.org/abs/2010.04159).
+
+```latex
+@inproceedings{
+zhu2021deformable,
+title={Deformable DETR: Deformable Transformers for End-to-End Object Detection},
+author={Xizhou Zhu and Weijie Su and Lewei Lu and Bin Li and Xiaogang Wang and Jifeng Dai},
+booktitle={International Conference on Learning Representations},
+year={2021},
+url={https://openreview.net/forum?id=gZ9hCDWe6ke}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/deformable_detr/deformable-detr-refine-twostage_r50_16xb2-50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deformable_detr/deformable-detr-refine-twostage_r50_16xb2-50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..eeb67fc98486cfd929a8177b9af6be3cdab9aa4b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deformable_detr/deformable-detr-refine-twostage_r50_16xb2-50e_coco.py
@@ -0,0 +1,2 @@
+_base_ = 'deformable-detr-refine_r50_16xb2-50e_coco.py'
+model = dict(as_two_stage=True)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/deformable_detr/deformable-detr-refine_r50_16xb2-50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deformable_detr/deformable-detr-refine_r50_16xb2-50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b968674f4a9fc450803cdba018b0c4e9e6ca422a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deformable_detr/deformable-detr-refine_r50_16xb2-50e_coco.py
@@ -0,0 +1,2 @@
+_base_ = 'deformable-detr_r50_16xb2-50e_coco.py'
+model = dict(with_box_refine=True)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/deformable_detr/deformable-detr_r50_16xb2-50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deformable_detr/deformable-detr_r50_16xb2-50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e0dee411c8e27ab440ccc874e40f4207b24a21e7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deformable_detr/deformable-detr_r50_16xb2-50e_coco.py
@@ -0,0 +1,156 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ type='DeformableDETR',
+ num_queries=300,
+ num_feature_levels=4,
+ with_box_refine=False,
+ as_two_stage=False,
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=1),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='ChannelMapper',
+ in_channels=[512, 1024, 2048],
+ kernel_size=1,
+ out_channels=256,
+ act_cfg=None,
+ norm_cfg=dict(type='GN', num_groups=32),
+ num_outs=4),
+ encoder=dict( # DeformableDetrTransformerEncoder
+ num_layers=6,
+ layer_cfg=dict( # DeformableDetrTransformerEncoderLayer
+ self_attn_cfg=dict( # MultiScaleDeformableAttention
+ embed_dims=256,
+ batch_first=True),
+ ffn_cfg=dict(
+ embed_dims=256, feedforward_channels=1024, ffn_drop=0.1))),
+ decoder=dict( # DeformableDetrTransformerDecoder
+ num_layers=6,
+ return_intermediate=True,
+ layer_cfg=dict( # DeformableDetrTransformerDecoderLayer
+ self_attn_cfg=dict( # MultiheadAttention
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.1,
+ batch_first=True),
+ cross_attn_cfg=dict( # MultiScaleDeformableAttention
+ embed_dims=256,
+ batch_first=True),
+ ffn_cfg=dict(
+ embed_dims=256, feedforward_channels=1024, ffn_drop=0.1)),
+ post_norm_cfg=None),
+ positional_encoding=dict(num_feats=128, normalize=True, offset=-0.5),
+ bbox_head=dict(
+ type='DeformableDETRHead',
+ num_classes=80,
+ sync_cls_avg_factor=True,
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=2.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=5.0),
+ loss_iou=dict(type='GIoULoss', loss_weight=2.0)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='HungarianAssigner',
+ match_costs=[
+ dict(type='FocalLossCost', weight=2.0),
+ dict(type='BBoxL1Cost', weight=5.0, box_format='xywh'),
+ dict(type='IoUCost', iou_mode='giou', weight=2.0)
+ ])),
+ test_cfg=dict(max_per_img=100))
+
+# train_pipeline, NOTE the img_scale and the Pad's size_divisor is different
+# from the default setting in mmdet.
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(
+ dataset=dict(
+ filter_cfg=dict(filter_empty_gt=False), pipeline=train_pipeline))
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0002, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'backbone': dict(lr_mult=0.1),
+ 'sampling_offsets': dict(lr_mult=0.1),
+ 'reference_points': dict(lr_mult=0.1)
+ }))
+
+# learning policy
+max_epochs = 50
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[40],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (16 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=32)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/deformable_detr/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deformable_detr/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..a30c97914baf6f1ec56cea8fd67b5ad1efb574fe
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/deformable_detr/metafile.yml
@@ -0,0 +1,56 @@
+Collections:
+ - Name: Deformable DETR
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - AdamW
+ - Multi Scale Train
+ - Gradient Clip
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNet
+ - Transformer
+ Paper:
+ URL: https://openreview.net/forum?id=gZ9hCDWe6ke
+ Title: 'Deformable DETR: Deformable Transformers for End-to-End Object Detection'
+ README: configs/deformable_detr/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.12.0/mmdet/models/detectors/deformable_detr.py#L6
+ Version: v2.12.0
+
+Models:
+ - Name: deformable-detr_r50_16xb2-50e_coco
+ In Collection: Deformable DETR
+ Config: configs/deformable_detr/deformable-detr_r50_16xb2-50e_coco.py
+ Metadata:
+ Epochs: 50
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.3
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/deformable_detr/deformable-detr_r50_16xb2-50e_coco/deformable-detr_r50_16xb2-50e_coco_20221029_210934-6bc7d21b.pth
+
+ - Name: deformable-detr-refine_r50_16xb2-50e_coco
+ In Collection: Deformable DETR
+ Config: configs/deformable_detr/deformable-detr-refine_r50_16xb2-50e_coco.py
+ Metadata:
+ Epochs: 50
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.2
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/deformable_detr/deformable-detr-refine_r50_16xb2-50e_coco/deformable-detr-refine_r50_16xb2-50e_coco_20221022_225303-844e0f93.pth
+
+ - Name: deformable-detr-refine-twostage_r50_16xb2-50e_coco
+ In Collection: Deformable DETR
+ Config: configs/deformable_detr/deformable-detr-refine-twostage_r50_16xb2-50e_coco.py
+ Metadata:
+ Epochs: 50
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 47.0
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/deformable_detr/deformable-detr-refine-twostage_r50_16xb2-50e_coco/deformable-detr-refine-twostage_r50_16xb2-50e_coco_20221021_184714-acc8a5ff.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..2918d6e4f1072428bbefcfcd05e139fc590766aa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/README.md
@@ -0,0 +1,69 @@
+# DetectoRS
+
+> [DetectoRS: Detecting Objects with Recursive Feature Pyramid and Switchable Atrous Convolution](https://arxiv.org/abs/2006.02334)
+
+
+
+## Abstract
+
+Many modern object detectors demonstrate outstanding performances by using the mechanism of looking and thinking twice. In this paper, we explore this mechanism in the backbone design for object detection. At the macro level, we propose Recursive Feature Pyramid, which incorporates extra feedback connections from Feature Pyramid Networks into the bottom-up backbone layers. At the micro level, we propose Switchable Atrous Convolution, which convolves the features with different atrous rates and gathers the results using switch functions. Combining them results in DetectoRS, which significantly improves the performances of object detection. On COCO test-dev, DetectoRS achieves state-of-the-art 55.7% box AP for object detection, 48.5% mask AP for instance segmentation, and 50.0% PQ for panoptic segmentation.
+
+
+

+
+
+## Introduction
+
+DetectoRS requires COCO and [COCO-stuff](http://calvin.inf.ed.ac.uk/wp-content/uploads/data/cocostuffdataset/stuffthingmaps_trainval2017.zip) dataset for training. You need to download and extract it in the COCO dataset path.
+The directory should be like this.
+
+```none
+mmdetection
+├── mmdet
+├── tools
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ ├── train2017
+│ │ ├── val2017
+│ │ ├── test2017
+| | ├── stuffthingmaps
+```
+
+## Results and Models
+
+DetectoRS includes two major components:
+
+- Recursive Feature Pyramid (RFP).
+- Switchable Atrous Convolution (SAC).
+
+They can be used independently.
+Combining them together results in DetectoRS.
+The results on COCO 2017 val are shown in the below table.
+
+| Method | Detector | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :-------: | :-----------------: | :-----: | :------: | :------------: | :----: | :-----: | :-----------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| RFP | Cascade + ResNet-50 | 1x | 7.5 | - | 44.8 | | [config](./cascade-rcnn_r50-rfp_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/detectors/cascade_rcnn_r50_rfp_1x_coco/cascade_rcnn_r50_rfp_1x_coco-8cf51bfd.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/detectors/cascade_rcnn_r50_rfp_1x_coco/cascade_rcnn_r50_rfp_1x_coco_20200624_104126.log.json) |
+| SAC | Cascade + ResNet-50 | 1x | 5.6 | - | 45.0 | | [config](./cascade-rcnn_r50-sac_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/detectors/cascade_rcnn_r50_sac_1x_coco/cascade_rcnn_r50_sac_1x_coco-24bfda62.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/detectors/cascade_rcnn_r50_sac_1x_coco/cascade_rcnn_r50_sac_1x_coco_20200624_104402.log.json) |
+| DetectoRS | Cascade + ResNet-50 | 1x | 9.9 | - | 47.4 | | [config](./detectors_cascade-rcnn_r50_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/detectors/detectors_cascade_rcnn_r50_1x_coco/detectors_cascade_rcnn_r50_1x_coco-32a10ba0.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/detectors/detectors_cascade_rcnn_r50_1x_coco/detectors_cascade_rcnn_r50_1x_coco_20200706_001203.log.json) |
+| RFP | HTC + ResNet-50 | 1x | 11.2 | - | 46.6 | 40.9 | [config](./htc_r50-rfp_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/detectors/htc_r50_rfp_1x_coco/htc_r50_rfp_1x_coco-8ff87c51.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/detectors/htc_r50_rfp_1x_coco/htc_r50_rfp_1x_coco_20200624_103053.log.json) |
+| SAC | HTC + ResNet-50 | 1x | 9.3 | - | 46.4 | 40.9 | [config](./htc_r50-sac_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/detectors/htc_r50_sac_1x_coco/htc_r50_sac_1x_coco-bfa60c54.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/detectors/htc_r50_sac_1x_coco/htc_r50_sac_1x_coco_20200624_103111.log.json) |
+| DetectoRS | HTC + ResNet-50 | 1x | 13.6 | - | 49.1 | 42.6 | [config](./detectors_htc-r50_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/detectors/detectors_htc_r50_1x_coco/detectors_htc_r50_1x_coco-329b1453.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/detectors/detectors_htc_r50_1x_coco/detectors_htc_r50_1x_coco_20200624_103659.log.json) |
+| DetectoRS | HTC + ResNet-101 | 20e | 19.6 | | 50.5 | 43.9 | [config](./detectors_htc-r101_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/detectors/detectors_htc_r101_20e_coco/detectors_htc_r101_20e_coco_20210419_203638-348d533b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/detectors/detectors_htc_r101_20e_coco/detectors_htc_r101_20e_coco_20210419_203638.log.json) |
+
+*Note*: This is a re-implementation based on MMDetection-V2.
+The original implementation is based on MMDetection-V1.
+
+## Citation
+
+We provide the config files for [DetectoRS: Detecting Objects with Recursive Feature Pyramid and Switchable Atrous Convolution](https://arxiv.org/pdf/2006.02334.pdf).
+
+```latex
+@article{qiao2020detectors,
+ title={DetectoRS: Detecting Objects with Recursive Feature Pyramid and Switchable Atrous Convolution},
+ author={Qiao, Siyuan and Chen, Liang-Chieh and Yuille, Alan},
+ journal={arXiv preprint arXiv:2006.02334},
+ year={2020}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/cascade-rcnn_r50-rfp_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/cascade-rcnn_r50-rfp_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c30c84d74cf68bc4369db16b6b2602626acb6fdf
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/cascade-rcnn_r50-rfp_1x_coco.py
@@ -0,0 +1,28 @@
+_base_ = [
+ '../_base_/models/cascade-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ backbone=dict(
+ type='DetectoRS_ResNet',
+ conv_cfg=dict(type='ConvAWS'),
+ output_img=True),
+ neck=dict(
+ type='RFP',
+ rfp_steps=2,
+ aspp_out_channels=64,
+ aspp_dilations=(1, 3, 6, 1),
+ rfp_backbone=dict(
+ rfp_inplanes=256,
+ type='DetectoRS_ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ conv_cfg=dict(type='ConvAWS'),
+ pretrained='torchvision://resnet50',
+ style='pytorch')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/cascade-rcnn_r50-sac_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/cascade-rcnn_r50-sac_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..24d6cd3a95ecf262caac667cfcc32d6885fa5880
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/cascade-rcnn_r50-sac_1x_coco.py
@@ -0,0 +1,12 @@
+_base_ = [
+ '../_base_/models/cascade-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ backbone=dict(
+ type='DetectoRS_ResNet',
+ conv_cfg=dict(type='ConvAWS'),
+ sac=dict(type='SAC', use_deform=True),
+ stage_with_sac=(False, True, True, True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/detectors_cascade-rcnn_r50_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/detectors_cascade-rcnn_r50_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..19d13d9c8c38b666b7481a58a641918b5d20e0ad
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/detectors_cascade-rcnn_r50_1x_coco.py
@@ -0,0 +1,32 @@
+_base_ = [
+ '../_base_/models/cascade-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ backbone=dict(
+ type='DetectoRS_ResNet',
+ conv_cfg=dict(type='ConvAWS'),
+ sac=dict(type='SAC', use_deform=True),
+ stage_with_sac=(False, True, True, True),
+ output_img=True),
+ neck=dict(
+ type='RFP',
+ rfp_steps=2,
+ aspp_out_channels=64,
+ aspp_dilations=(1, 3, 6, 1),
+ rfp_backbone=dict(
+ rfp_inplanes=256,
+ type='DetectoRS_ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ conv_cfg=dict(type='ConvAWS'),
+ sac=dict(type='SAC', use_deform=True),
+ stage_with_sac=(False, True, True, True),
+ pretrained='torchvision://resnet50',
+ style='pytorch')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/detectors_htc-r101_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/detectors_htc-r101_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..93d7d2b1adeb3fbdb7bac0107edf4433669e8015
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/detectors_htc-r101_20e_coco.py
@@ -0,0 +1,28 @@
+_base_ = '../htc/htc_r101_fpn_20e_coco.py'
+
+model = dict(
+ backbone=dict(
+ type='DetectoRS_ResNet',
+ conv_cfg=dict(type='ConvAWS'),
+ sac=dict(type='SAC', use_deform=True),
+ stage_with_sac=(False, True, True, True),
+ output_img=True),
+ neck=dict(
+ type='RFP',
+ rfp_steps=2,
+ aspp_out_channels=64,
+ aspp_dilations=(1, 3, 6, 1),
+ rfp_backbone=dict(
+ rfp_inplanes=256,
+ type='DetectoRS_ResNet',
+ depth=101,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ conv_cfg=dict(type='ConvAWS'),
+ sac=dict(type='SAC', use_deform=True),
+ stage_with_sac=(False, True, True, True),
+ pretrained='torchvision://resnet101',
+ style='pytorch')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/detectors_htc-r50_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/detectors_htc-r50_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..0d2fc4f77fcca715c1dfb613306d214b636aa0c0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/detectors_htc-r50_1x_coco.py
@@ -0,0 +1,28 @@
+_base_ = '../htc/htc_r50_fpn_1x_coco.py'
+
+model = dict(
+ backbone=dict(
+ type='DetectoRS_ResNet',
+ conv_cfg=dict(type='ConvAWS'),
+ sac=dict(type='SAC', use_deform=True),
+ stage_with_sac=(False, True, True, True),
+ output_img=True),
+ neck=dict(
+ type='RFP',
+ rfp_steps=2,
+ aspp_out_channels=64,
+ aspp_dilations=(1, 3, 6, 1),
+ rfp_backbone=dict(
+ rfp_inplanes=256,
+ type='DetectoRS_ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ conv_cfg=dict(type='ConvAWS'),
+ sac=dict(type='SAC', use_deform=True),
+ stage_with_sac=(False, True, True, True),
+ pretrained='torchvision://resnet50',
+ style='pytorch')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/htc_r50-rfp_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/htc_r50-rfp_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..496104e12550a1985f9c9e3748a343f69d7df6d8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/htc_r50-rfp_1x_coco.py
@@ -0,0 +1,24 @@
+_base_ = '../htc/htc_r50_fpn_1x_coco.py'
+
+model = dict(
+ backbone=dict(
+ type='DetectoRS_ResNet',
+ conv_cfg=dict(type='ConvAWS'),
+ output_img=True),
+ neck=dict(
+ type='RFP',
+ rfp_steps=2,
+ aspp_out_channels=64,
+ aspp_dilations=(1, 3, 6, 1),
+ rfp_backbone=dict(
+ rfp_inplanes=256,
+ type='DetectoRS_ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ conv_cfg=dict(type='ConvAWS'),
+ pretrained='torchvision://resnet50',
+ style='pytorch')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/htc_r50-sac_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/htc_r50-sac_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..72d4db963ffd95851b945911b3db9941426583ab
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/htc_r50-sac_1x_coco.py
@@ -0,0 +1,8 @@
+_base_ = '../htc/htc_r50_fpn_1x_coco.py'
+
+model = dict(
+ backbone=dict(
+ type='DetectoRS_ResNet',
+ conv_cfg=dict(type='ConvAWS'),
+ sac=dict(type='SAC', use_deform=True),
+ stage_with_sac=(False, True, True, True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..196a1cef1751bc9d5812915c4d06de220f62baa1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detectors/metafile.yml
@@ -0,0 +1,114 @@
+Collections:
+ - Name: DetectoRS
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ASPP
+ - FPN
+ - RFP
+ - RPN
+ - ResNet
+ - RoIAlign
+ - SAC
+ Paper:
+ URL: https://arxiv.org/abs/2006.02334
+ Title: 'DetectoRS: Detecting Objects with Recursive Feature Pyramid and Switchable Atrous Convolution'
+ README: configs/detectors/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.2.0/mmdet/models/backbones/detectors_resnet.py#L205
+ Version: v2.2.0
+
+Models:
+ - Name: cascade-rcnn_r50-rfp_1x_coco
+ In Collection: DetectoRS
+ Config: configs/detectors/cascade-rcnn_r50-rfp_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.5
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/detectors/cascade_rcnn_r50_rfp_1x_coco/cascade_rcnn_r50_rfp_1x_coco-8cf51bfd.pth
+
+ - Name: cascade-rcnn_r50-sac_1x_coco
+ In Collection: DetectoRS
+ Config: configs/detectors/cascade-rcnn_r50-sac_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.6
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/detectors/cascade_rcnn_r50_sac_1x_coco/cascade_rcnn_r50_sac_1x_coco-24bfda62.pth
+
+ - Name: detectors_cascade-rcnn_r50_1x_coco
+ In Collection: DetectoRS
+ Config: configs/detectors/detectors_cascade-rcnn_r50_1x_coco.py
+ Metadata:
+ Training Memory (GB): 9.9
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 47.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/detectors/detectors_cascade_rcnn_r50_1x_coco/detectors_cascade_rcnn_r50_1x_coco-32a10ba0.pth
+
+ - Name: htc_r50-rfp_1x_coco
+ In Collection: DetectoRS
+ Config: configs/detectors/htc_r50-rfp_1x_coco.py
+ Metadata:
+ Training Memory (GB): 11.2
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.6
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 40.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/detectors/htc_r50_rfp_1x_coco/htc_r50_rfp_1x_coco-8ff87c51.pth
+
+ - Name: htc_r50-sac_1x_coco
+ In Collection: DetectoRS
+ Config: configs/detectors/htc_r50-sac_1x_coco.py
+ Metadata:
+ Training Memory (GB): 9.3
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 40.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/detectors/htc_r50_sac_1x_coco/htc_r50_sac_1x_coco-bfa60c54.pth
+
+ - Name: detectors_htc-r50_1x_coco
+ In Collection: DetectoRS
+ Config: configs/detectors/detectors_htc-r50_1x_coco.py
+ Metadata:
+ Training Memory (GB): 13.6
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 49.1
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 42.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/detectors/detectors_htc_r50_1x_coco/detectors_htc_r50_1x_coco-329b1453.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/detr/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detr/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..8e843f369be40cac73bbc098d6bb04097de0a722
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detr/README.md
@@ -0,0 +1,37 @@
+# DETR
+
+> [End-to-End Object Detection with Transformers](https://arxiv.org/abs/2005.12872)
+
+
+
+## Abstract
+
+We present a new method that views object detection as a direct set prediction problem. Our approach streamlines the detection pipeline, effectively removing the need for many hand-designed components like a non-maximum suppression procedure or anchor generation that explicitly encode our prior knowledge about the task. The main ingredients of the new framework, called DEtection TRansformer or DETR, are a set-based global loss that forces unique predictions via bipartite matching, and a transformer encoder-decoder architecture. Given a fixed small set of learned object queries, DETR reasons about the relations of the objects and the global image context to directly output the final set of predictions in parallel. The new model is conceptually simple and does not require a specialized library, unlike many other modern detectors. DETR demonstrates accuracy and run-time performance on par with the well-established and highly-optimized Faster RCNN baseline on the challenging COCO object detection dataset. Moreover, DETR can be easily generalized to produce panoptic segmentation in a unified manner. We show that it significantly outperforms competitive baselines.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Model | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :------: | :---: | :-----: | :------: | :------------: | :----: | :------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | DETR | 150e | 7.9 | | 39.9 | [config](./detr_r50_8xb2-150e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/detr/detr_r50_8xb2-150e_coco/detr_r50_8xb2-150e_coco_20221023_153551-436d03e8.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/detr/detr_r50_8xb2-150e_coco/detr_r50_8xb2-150e_coco_20221023_153551.log.json) |
+
+## Citation
+
+We provide the config files for DETR: [End-to-End Object Detection with Transformers](https://arxiv.org/abs/2005.12872).
+
+```latex
+@inproceedings{detr,
+ author = {Nicolas Carion and
+ Francisco Massa and
+ Gabriel Synnaeve and
+ Nicolas Usunier and
+ Alexander Kirillov and
+ Sergey Zagoruyko},
+ title = {End-to-End Object Detection with Transformers},
+ booktitle = {ECCV},
+ year = {2020}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/detr/detr_r101_8xb2-500e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detr/detr_r101_8xb2-500e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6661aacdc54e889aa38b2e759c40fd9797ae44ad
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detr/detr_r101_8xb2-500e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './detr_r50_8xb2-500e_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/detr/detr_r18_8xb2-500e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detr/detr_r18_8xb2-500e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..305b9d6fee8d75273b588f32b2e21582473cb137
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detr/detr_r18_8xb2-500e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './detr_r50_8xb2-500e_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=18,
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet18')),
+ neck=dict(in_channels=[512]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/detr/detr_r50_8xb2-150e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detr/detr_r50_8xb2-150e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..aaa15410532e552cae387ef4eaa57227af1d855d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detr/detr_r50_8xb2-150e_coco.py
@@ -0,0 +1,155 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ type='DETR',
+ num_queries=100,
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=1),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(3, ),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='ChannelMapper',
+ in_channels=[2048],
+ kernel_size=1,
+ out_channels=256,
+ act_cfg=None,
+ norm_cfg=None,
+ num_outs=1),
+ encoder=dict( # DetrTransformerEncoder
+ num_layers=6,
+ layer_cfg=dict( # DetrTransformerEncoderLayer
+ self_attn_cfg=dict( # MultiheadAttention
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.1,
+ batch_first=True),
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048,
+ num_fcs=2,
+ ffn_drop=0.1,
+ act_cfg=dict(type='ReLU', inplace=True)))),
+ decoder=dict( # DetrTransformerDecoder
+ num_layers=6,
+ layer_cfg=dict( # DetrTransformerDecoderLayer
+ self_attn_cfg=dict( # MultiheadAttention
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.1,
+ batch_first=True),
+ cross_attn_cfg=dict( # MultiheadAttention
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.1,
+ batch_first=True),
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048,
+ num_fcs=2,
+ ffn_drop=0.1,
+ act_cfg=dict(type='ReLU', inplace=True))),
+ return_intermediate=True),
+ positional_encoding=dict(num_feats=128, normalize=True),
+ bbox_head=dict(
+ type='DETRHead',
+ num_classes=80,
+ embed_dims=256,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ bg_cls_weight=0.1,
+ use_sigmoid=False,
+ loss_weight=1.0,
+ class_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=5.0),
+ loss_iou=dict(type='GIoULoss', loss_weight=2.0)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='HungarianAssigner',
+ match_costs=[
+ dict(type='ClassificationCost', weight=1.),
+ dict(type='BBoxL1Cost', weight=5.0, box_format='xywh'),
+ dict(type='IoUCost', iou_mode='giou', weight=2.0)
+ ])),
+ test_cfg=dict(max_per_img=100))
+
+# train_pipeline, NOTE the img_scale and the Pad's size_divisor is different
+# from the default setting in mmdet.
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[[
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(400, 1333), (500, 1333), (600, 1333)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333),
+ (576, 1333), (608, 1333), (640, 1333),
+ (672, 1333), (704, 1333), (736, 1333),
+ (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]]),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0001, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(
+ custom_keys={'backbone': dict(lr_mult=0.1, decay_mult=1.0)}))
+
+# learning policy
+max_epochs = 150
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[100],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/detr/detr_r50_8xb2-500e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detr/detr_r50_8xb2-500e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f07d5dce05b08c74aea2059989b45d5d275c53e0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detr/detr_r50_8xb2-500e_coco.py
@@ -0,0 +1,24 @@
+_base_ = './detr_r50_8xb2-150e_coco.py'
+
+# learning policy
+max_epochs = 500
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=10)
+
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[334],
+ gamma=0.1)
+]
+
+# only keep latest 2 checkpoints
+default_hooks = dict(checkpoint=dict(max_keep_ckpts=2))
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/detr/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detr/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..a9132dff0228e31c146ae46ed32445491f4225c1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/detr/metafile.yml
@@ -0,0 +1,33 @@
+Collections:
+ - Name: DETR
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - AdamW
+ - Multi Scale Train
+ - Gradient Clip
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNet
+ - Transformer
+ Paper:
+ URL: https://arxiv.org/abs/2005.12872
+ Title: 'End-to-End Object Detection with Transformers'
+ README: configs/detr/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.7.0/mmdet/models/detectors/detr.py#L7
+ Version: v2.7.0
+
+Models:
+ - Name: detr_r50_8xb2-150e_coco
+ In Collection: DETR
+ Config: configs/detr/detr_r50_8xb2-150e_coco.py
+ Metadata:
+ Training Memory (GB): 7.9
+ Epochs: 150
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.9
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/detr/detr_r50_8xb2-150e_coco/detr_r50_8xb2-150e_coco_20221023_153551-436d03e8.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..d8a01bde25582023ab65c0304faa8ef14340a27a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/README.md
@@ -0,0 +1,40 @@
+# DINO
+
+> [DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection](https://arxiv.org/abs/2203.03605)
+
+
+
+## Abstract
+
+We present DINO (DETR with Improved deNoising anchOr boxes), a state-of-the-art end-to-end object detector. DINO improves over previous DETR-like models in performance and efficiency by using a contrastive way for denoising training, a mixed query selection method for anchor initialization, and a look forward twice scheme for box prediction. DINO achieves 49.4AP in 12 epochs and 51.3AP in 24 epochs on COCO with a ResNet-50 backbone and multi-scale features, yielding a significant improvement of +6.0AP and +2.7AP, respectively, compared to DN-DETR, the previous best DETR-like model. DINO scales well in both model size and data size. Without bells and whistles, after pre-training on the Objects365 dataset with a SwinL backbone, DINO obtains the best results on both COCO val2017 (63.2AP) and test-dev (63.3AP). Compared to other models on the leaderboard, DINO significantly reduces its model size and pre-training data size while achieving better results.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Model | Lr schd | Better-Hyper | box AP | Config | Download |
+| :------: | :---------: | :-----: | :----------: | :----: | :---------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | DINO-4scale | 12e | False | 49.0 | [config](./dino-4scale_r50_8xb2-12e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/dino/dino-4scale_r50_8xb2-12e_coco/dino-4scale_r50_8xb2-12e_coco_20221202_182705-55b2bba2.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/dino/dino-4scale_r50_8xb2-12e_coco/dino-4scale_r50_8xb2-12e_coco_20221202_182705.log.json) |
+| R-50 | DINO-4scale | 12e | True | 50.1 | [config](./dino-4scale_r50_improved_8xb2-12e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/dino/dino-4scale_r50_improved_8xb2-12e_coco/dino-4scale_r50_improved_8xb2-12e_coco_20230818_162607-6f47a913.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/dino/dino-4scale_r50_improved_8xb2-12e_coco/dino-4scale_r50_improved_8xb2-12e_coco_20230818_162607.log.json) |
+| Swin-L | DINO-5scale | 12e | False | 57.2 | [config](./dino-5scale_swin-l_8xb2-12e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/dino/dino-5scale_swin-l_8xb2-12e_coco/dino-5scale_swin-l_8xb2-12e_coco_20230228_072924-a654145f.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/dino/dino-5scale_swin-l_8xb2-12e_coco/dino-5scale_swin-l_8xb2-12e_coco_20230228_072924.log) |
+| Swin-L | DINO-5scale | 36e | False | 58.4 | [config](./dino-5scale_swin-l_8xb2-36e_coco.py) | [model](https://github.com/RistoranteRist/mmlab-weights/releases/download/dino-swinl/dino-5scale_swin-l_8xb2-36e_coco-5486e051.pth) \| [log](https://github.com/RistoranteRist/mmlab-weights/releases/download/dino-swinl/20230307_032359.log) |
+
+### NOTE
+
+The performance is unstable. `DINO-4scale` with `R-50` may fluctuate about 0.4 mAP.
+
+## Citation
+
+We provide the config files for DINO: [DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection](https://arxiv.org/abs/2203.03605).
+
+```latex
+@misc{zhang2022dino,
+ title={DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection},
+ author={Hao Zhang and Feng Li and Shilong Liu and Lei Zhang and Hang Su and Jun Zhu and Lionel M. Ni and Heung-Yeung Shum},
+ year={2022},
+ eprint={2203.03605},
+ archivePrefix={arXiv},
+ primaryClass={cs.CV}}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/dino-4scale_r50_8xb2-12e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/dino-4scale_r50_8xb2-12e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5831f898b4a706accb2b828b6194b2974e78d0fc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/dino-4scale_r50_8xb2-12e_coco.py
@@ -0,0 +1,163 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ type='DINO',
+ num_queries=900, # num_matching_queries
+ with_box_refine=True,
+ as_two_stage=True,
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=1),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='ChannelMapper',
+ in_channels=[512, 1024, 2048],
+ kernel_size=1,
+ out_channels=256,
+ act_cfg=None,
+ norm_cfg=dict(type='GN', num_groups=32),
+ num_outs=4),
+ encoder=dict(
+ num_layers=6,
+ layer_cfg=dict(
+ self_attn_cfg=dict(embed_dims=256, num_levels=4,
+ dropout=0.0), # 0.1 for DeformDETR
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048, # 1024 for DeformDETR
+ ffn_drop=0.0))), # 0.1 for DeformDETR
+ decoder=dict(
+ num_layers=6,
+ return_intermediate=True,
+ layer_cfg=dict(
+ self_attn_cfg=dict(embed_dims=256, num_heads=8,
+ dropout=0.0), # 0.1 for DeformDETR
+ cross_attn_cfg=dict(embed_dims=256, num_levels=4,
+ dropout=0.0), # 0.1 for DeformDETR
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048, # 1024 for DeformDETR
+ ffn_drop=0.0)), # 0.1 for DeformDETR
+ post_norm_cfg=None),
+ positional_encoding=dict(
+ num_feats=128,
+ normalize=True,
+ offset=0.0, # -0.5 for DeformDETR
+ temperature=20), # 10000 for DeformDETR
+ bbox_head=dict(
+ type='DINOHead',
+ num_classes=80,
+ sync_cls_avg_factor=True,
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0), # 2.0 in DeformDETR
+ loss_bbox=dict(type='L1Loss', loss_weight=5.0),
+ loss_iou=dict(type='GIoULoss', loss_weight=2.0)),
+ dn_cfg=dict( # TODO: Move to model.train_cfg ?
+ label_noise_scale=0.5,
+ box_noise_scale=1.0, # 0.4 for DN-DETR
+ group_cfg=dict(dynamic=True, num_groups=None,
+ num_dn_queries=100)), # TODO: half num_dn_queries
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='HungarianAssigner',
+ match_costs=[
+ dict(type='FocalLossCost', weight=2.0),
+ dict(type='BBoxL1Cost', weight=5.0, box_format='xywh'),
+ dict(type='IoUCost', iou_mode='giou', weight=2.0)
+ ])),
+ test_cfg=dict(max_per_img=300)) # 100 for DeformDETR
+
+# train_pipeline, NOTE the img_scale and the Pad's size_divisor is different
+# from the default setting in mmdet.
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(
+ dataset=dict(
+ filter_cfg=dict(filter_empty_gt=False), pipeline=train_pipeline))
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(
+ type='AdamW',
+ lr=0.0001, # 0.0002 for DeformDETR
+ weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(custom_keys={'backbone': dict(lr_mult=0.1)})
+) # custom_keys contains sampling_offsets and reference_points in DeformDETR # noqa
+
+# learning policy
+max_epochs = 12
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[11],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/dino-4scale_r50_8xb2-24e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/dino-4scale_r50_8xb2-24e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8534ac6a7ccc7f3f8c081275b3567a0a0792b7a5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/dino-4scale_r50_8xb2-24e_coco.py
@@ -0,0 +1,13 @@
+_base_ = './dino-4scale_r50_8xb2-12e_coco.py'
+max_epochs = 24
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[20],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/dino-4scale_r50_8xb2-36e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/dino-4scale_r50_8xb2-36e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1c2cf4602d358dfed5b737f8a74843c89a54702d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/dino-4scale_r50_8xb2-36e_coco.py
@@ -0,0 +1,13 @@
+_base_ = './dino-4scale_r50_8xb2-12e_coco.py'
+max_epochs = 36
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[30],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/dino-4scale_r50_improved_8xb2-12e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/dino-4scale_r50_improved_8xb2-12e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6a4a82bacc1f1e990d4720db81cae0af5c012557
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/dino-4scale_r50_improved_8xb2-12e_coco.py
@@ -0,0 +1,18 @@
+_base_ = ['dino-4scale_r50_8xb2-12e_coco.py']
+
+# from deformable detr hyper
+model = dict(
+ backbone=dict(frozen_stages=-1),
+ bbox_head=dict(loss_cls=dict(loss_weight=2.0)),
+ positional_encoding=dict(offset=-0.5, temperature=10000),
+ dn_cfg=dict(group_cfg=dict(num_dn_queries=300)))
+
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(lr=0.0002),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'backbone': dict(lr_mult=0.1),
+ 'sampling_offsets': dict(lr_mult=0.1),
+ 'reference_points': dict(lr_mult=0.1)
+ }))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/dino-5scale_swin-l_8xb2-12e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/dino-5scale_swin-l_8xb2-12e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3d39f22f50926a11137d143976fe4033ec3a8640
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/dino-5scale_swin-l_8xb2-12e_coco.py
@@ -0,0 +1,30 @@
+_base_ = './dino-4scale_r50_8xb2-12e_coco.py'
+
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_large_patch4_window12_384_22k.pth' # noqa
+num_levels = 5
+model = dict(
+ num_feature_levels=num_levels,
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ pretrain_img_size=384,
+ embed_dims=192,
+ depths=[2, 2, 18, 2],
+ num_heads=[6, 12, 24, 48],
+ window_size=12,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.2,
+ patch_norm=True,
+ out_indices=(0, 1, 2, 3),
+ # Please only add indices that would be used
+ # in FPN, otherwise some parameter will not be used
+ with_cp=True,
+ convert_weights=True,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ neck=dict(in_channels=[192, 384, 768, 1536], num_outs=num_levels),
+ encoder=dict(layer_cfg=dict(self_attn_cfg=dict(num_levels=num_levels))),
+ decoder=dict(layer_cfg=dict(cross_attn_cfg=dict(num_levels=num_levels))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/dino-5scale_swin-l_8xb2-36e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/dino-5scale_swin-l_8xb2-36e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d55a38e61d411892c6de819cf46247ba4d41d427
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/dino-5scale_swin-l_8xb2-36e_coco.py
@@ -0,0 +1,13 @@
+_base_ = './dino-5scale_swin-l_8xb2-12e_coco.py'
+max_epochs = 36
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[27, 33],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..f276a04ef557b70443083ac70b6a16671e7fa6e1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dino/metafile.yml
@@ -0,0 +1,85 @@
+Collections:
+ - Name: DINO
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - AdamW
+ - Multi Scale Train
+ - Gradient Clip
+ Training Resources: 8x A100 GPUs
+ Architecture:
+ - ResNet
+ - Transformer
+ Paper:
+ URL: https://arxiv.org/abs/2203.03605
+ Title: 'DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection'
+ README: configs/dino/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/f4112c9e5611468ffbd57cfba548fd1289264b52/mmdet/models/detectors/dino.py#L17
+ Version: v3.0.0rc6
+
+Models:
+ - Name: dino-4scale_r50_8xb2-12e_coco
+ In Collection: DINO
+ Config: configs/dino/dino-4scale_r50_8xb2-12e_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 49.0
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/dino/dino-4scale_r50_8xb2-12e_coco/dino-4scale_r50_8xb2-12e_coco_20221202_182705-55b2bba2.pth
+
+ - Name: dino-4scale_r50_8xb2-24e_coco
+ In Collection: DINO
+ Config: configs/dino/dino-4scale_r50_8xb2-24e_coco.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+
+ - Name: dino-4scale_r50_8xb2-36e_coco
+ In Collection: DINO
+ Config: configs/dino/dino-4scale_r50_8xb2-36e_coco.py
+ Metadata:
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+
+ - Name: dino-5scale_swin-l_8xb2-12e_coco
+ In Collection: DINO
+ Config: configs/dino/dino-5scale_swin-l_8xb2-12e_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 57.2
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/dino/dino-5scale_swin-l_8xb2-12e_coco/dino-5scale_swin-l_8xb2-12e_coco_20230228_072924-a654145f.pth
+
+ - Name: dino-5scale_swin-l_8xb2-36e_coco
+ In Collection: DINO
+ Config: configs/dino/dino-5scale_swin-l_8xb2-36e_coco.py
+ Metadata:
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 58.4
+ Weights: https://github.com/RistoranteRist/mmlab-weights/releases/download/dino-swinl/dino-5scale_swin-l_8xb2-36e_coco-5486e051.pth
+ - Name: dino-4scale_r50_improved_8xb2-12e_coco
+ In Collection: DINO
+ Config: configs/dino/dino-4scale_r50_improved_8xb2-12e_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 50.1
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/dino/dino-4scale_r50_improved_8xb2-12e_coco/dino-4scale_r50_improved_8xb2-12e_coco_20230818_162607-6f47a913.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/double_heads/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/double_heads/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..1b97dbc188df1557814f40e792940ab45a845781
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/double_heads/README.md
@@ -0,0 +1,32 @@
+# Double Heads
+
+> [Rethinking Classification and Localization for Object Detection](https://arxiv.org/abs/1904.06493)
+
+
+
+## Abstract
+
+Two head structures (i.e. fully connected head and convolution head) have been widely used in R-CNN based detectors for classification and localization tasks. However, there is a lack of understanding of how does these two head structures work for these two tasks. To address this issue, we perform a thorough analysis and find an interesting fact that the two head structures have opposite preferences towards the two tasks. Specifically, the fully connected head (fc-head) is more suitable for the classification task, while the convolution head (conv-head) is more suitable for the localization task. Furthermore, we examine the output feature maps of both heads and find that fc-head has more spatial sensitivity than conv-head. Thus, fc-head has more capability to distinguish a complete object from part of an object, but is not robust to regress the whole object. Based upon these findings, we propose a Double-Head method, which has a fully connected head focusing on classification and a convolution head for bounding box regression. Without bells and whistles, our method gains +3.5 and +2.8 AP on MS COCO dataset from Feature Pyramid Network (FPN) baselines with ResNet-50 and ResNet-101 backbones, respectively.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :------: | :-----: | :-----: | :------: | :------------: | :----: | :-------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | pytorch | 1x | 6.8 | 9.5 | 40.0 | [config](./dh-faster-rcnn_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/double_heads/dh_faster_rcnn_r50_fpn_1x_coco/dh_faster_rcnn_r50_fpn_1x_coco_20200130-586b67df.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/double_heads/dh_faster_rcnn_r50_fpn_1x_coco/dh_faster_rcnn_r50_fpn_1x_coco_20200130_220238.log.json) |
+
+## Citation
+
+```latex
+@article{wu2019rethinking,
+ title={Rethinking Classification and Localization for Object Detection},
+ author={Yue Wu and Yinpeng Chen and Lu Yuan and Zicheng Liu and Lijuan Wang and Hongzhi Li and Yun Fu},
+ year={2019},
+ eprint={1904.06493},
+ archivePrefix={arXiv},
+ primaryClass={cs.CV}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/double_heads/dh-faster-rcnn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/double_heads/dh-faster-rcnn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6b9b6e69a12d978a55fbba049fc2b1c5229c1fc5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/double_heads/dh-faster-rcnn_r50_fpn_1x_coco.py
@@ -0,0 +1,23 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ roi_head=dict(
+ type='DoubleHeadRoIHead',
+ reg_roi_scale_factor=1.3,
+ bbox_head=dict(
+ _delete_=True,
+ type='DoubleConvFCBBoxHead',
+ num_convs=4,
+ num_fcs=2,
+ in_channels=256,
+ conv_out_channels=1024,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=False,
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=2.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=2.0))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/double_heads/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/double_heads/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..bb14e7968e259bb6dae1bbd6dad5e1c4e862f228
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/double_heads/metafile.yml
@@ -0,0 +1,41 @@
+Collections:
+ - Name: Rethinking Classification and Localization for Object Detection
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - FPN
+ - RPN
+ - ResNet
+ - RoIAlign
+ Paper:
+ URL: https://arxiv.org/pdf/1904.06493
+ Title: 'Rethinking Classification and Localization for Object Detection'
+ README: configs/double_heads/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/roi_heads/double_roi_head.py#L6
+ Version: v2.0.0
+
+Models:
+ - Name: dh-faster-rcnn_r50_fpn_1x_coco
+ In Collection: Rethinking Classification and Localization for Object Detection
+ Config: configs/double_heads/dh-faster-rcnn_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.8
+ inference time (ms/im):
+ - value: 105.26
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/double_heads/dh_faster_rcnn_r50_fpn_1x_coco/dh_faster_rcnn_r50_fpn_1x_coco_20200130-586b67df.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..f38c3b65ac67ee623eb909acbd1dc8ad3eafa0af
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/README.md
@@ -0,0 +1,63 @@
+# DSDL: Standard Description Language for DataSet
+
+
+
+## 1. Abstract
+
+Data is the cornerstone of artificial intelligence. The efficiency of data acquisition, exchange, and application directly impacts the advances in technologies and applications. Over the long history of AI, a vast quantity of data sets have been developed and distributed. However, these datasets are defined in very different forms, which incurs significant overhead when it comes to exchange, integration, and utilization -- it is often the case that one needs to develop a new customized tool or script in order to incorporate a new dataset into a workflow.
+
+To overcome such difficulties, we develop **Data Set Description Language (DSDL)**. More details please visit our [official documents](https://opendatalab.github.io/dsdl-docs/getting_started/overview/), dsdl datasets can be downloaded from our platform [OpenDataLab](https://opendatalab.com/).
+
+## 2. Steps
+
+- install dsdl:
+
+ install by pip:
+
+ ```
+ pip install dsdl
+ ```
+
+ install by source code:
+
+ ```
+ git clone https://github.com/opendatalab/dsdl-sdk.git -b schema-dsdl
+ cd dsdl-sdk
+ python setup.py install
+ ```
+
+- install mmdet and pytorch:
+ please refer this [installation documents](https://mmdetection.readthedocs.io/en/latest/get_started.html).
+
+- train:
+
+ - using single gpu:
+
+ ```
+ python tools/train.py {config_file}
+ ```
+
+ - using slurm:
+
+ ```
+ ./tools/slurm_train.sh {partition} {job_name} {config_file} {work_dir} {gpu_nums}
+ ```
+
+## 3. Test Results
+
+- detection task:
+
+ | Datasets | Model | box AP | Config |
+ | :--------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: | :----: | :-------------------------: |
+ | VOC07+12 | [model](https://download.openmmlab.com/mmdetection/v2.0/pascal_voc/faster_rcnn_r50_fpn_1x_voc0712/faster_rcnn_r50_fpn_1x_voc0712_20220320_192712-54bef0f3.pth) | 80.3\* | [config](./voc0712.py) |
+ | COCO | [model](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_1x_coco/faster_rcnn_r50_fpn_1x_coco_20200130-047c8118.pth) | 37.4 | [config](./coco.py) |
+ | Objects365 | [model](https://download.openmmlab.com/mmdetection/v2.0/objects365/faster_rcnn_r50_fpn_16x4_1x_obj365v2/faster_rcnn_r50_fpn_16x4_1x_obj365v2_20221220_175040-5910b015.pth) | 19.8 | [config](./objects365v2.py) |
+ | OpenImages | [model](https://download.openmmlab.com/mmdetection/v2.0/openimages/faster_rcnn_r50_fpn_32x2_cas_1x_openimages/faster_rcnn_r50_fpn_32x2_cas_1x_openimages_20220306_202424-98c630e5.pth) | 59.9\* | [config](./openimagesv6.py) |
+
+ \*: box AP in voc metric and openimages metric, actually means AP_50.
+
+- instance segmentation task:
+
+ | Datasets | Model | box AP | mask AP | Config |
+ | :------: | :------------------------------------------------------------------------------------------------------------------------------------------: | :----: | :-----: | :--------------------------: |
+ | COCO | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_fpn_1x_coco/mask_rcnn_r50_fpn_1x_coco_20200205-d4b0c5d6.pth) | 38.1 | 34.7 | [config](./coco_instance.py) |
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3c9e895e53c1588028cf6def2fe79d49fd98d6e1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/coco.py
@@ -0,0 +1,33 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py',
+ '../_base_/datasets/dsdl.py'
+]
+
+# dsdl dataset settings
+
+# please visit our platform [OpenDataLab](https://opendatalab.com/)
+# to downloaded dsdl dataset.
+data_root = 'data/COCO2017'
+img_prefix = 'original'
+train_ann = 'dsdl/set-train/train.yaml'
+val_ann = 'dsdl/set-val/val.yaml'
+specific_key_path = dict(ignore_flag='./annotations/*/iscrowd')
+
+train_dataloader = dict(
+ dataset=dict(
+ specific_key_path=specific_key_path,
+ data_root=data_root,
+ ann_file=train_ann,
+ data_prefix=dict(img_path=img_prefix),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32, bbox_min_size=32),
+ ))
+
+val_dataloader = dict(
+ dataset=dict(
+ specific_key_path=specific_key_path,
+ data_root=data_root,
+ ann_file=val_ann,
+ data_prefix=dict(img_path=img_prefix),
+ ))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/coco_instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/coco_instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..e34f93c97f55f5eeef55f9de73f1a8389f8980c6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/coco_instance.py
@@ -0,0 +1,62 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py',
+ '../_base_/datasets/dsdl.py'
+]
+
+# dsdl dataset settings.
+
+# please visit our platform [OpenDataLab](https://opendatalab.com/)
+# to downloaded dsdl dataset.
+data_root = 'data/COCO2017'
+img_prefix = 'original'
+train_ann = 'dsdl/set-train/train.yaml'
+val_ann = 'dsdl/set-val/val.yaml'
+specific_key_path = dict(ignore_flag='./annotations/*/iscrowd')
+
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'instances'))
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ with_polygon=True,
+ specific_key_path=specific_key_path,
+ data_root=data_root,
+ ann_file=train_ann,
+ data_prefix=dict(img_path=img_prefix),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32, bbox_min_size=32),
+ pipeline=train_pipeline,
+ ))
+
+val_dataloader = dict(
+ dataset=dict(
+ with_polygon=True,
+ specific_key_path=specific_key_path,
+ data_root=data_root,
+ ann_file=val_ann,
+ data_prefix=dict(img_path=img_prefix),
+ pipeline=test_pipeline,
+ ))
+
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric', metric=['bbox', 'segm'], format_only=False)
+
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/objects365v2.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/objects365v2.py
new file mode 100644
index 0000000000000000000000000000000000000000..d25a2323027c22eaf9777f6e62e4992880b29d2c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/objects365v2.py
@@ -0,0 +1,54 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py',
+ '../_base_/datasets/dsdl.py'
+]
+
+model = dict(roi_head=dict(bbox_head=dict(num_classes=365)))
+
+# dsdl dataset settings
+
+# please visit our platform [OpenDataLab](https://opendatalab.com/)
+# to downloaded dsdl dataset.
+data_root = 'data/Objects365'
+img_prefix = 'original'
+train_ann = 'dsdl/set-train/train.yaml'
+val_ann = 'dsdl/set-val/val.yaml'
+specific_key_path = dict(ignore_flag='./annotations/*/iscrowd')
+
+train_dataloader = dict(
+ dataset=dict(
+ specific_key_path=specific_key_path,
+ data_root=data_root,
+ ann_file=train_ann,
+ data_prefix=dict(img_path=img_prefix),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32, bbox_min_size=32),
+ ))
+
+val_dataloader = dict(
+ dataset=dict(
+ specific_key_path=specific_key_path,
+ data_root=data_root,
+ ann_file=val_ann,
+ data_prefix=dict(img_path=img_prefix),
+ test_mode=True,
+ ))
+test_dataloader = val_dataloader
+
+default_hooks = dict(logger=dict(type='LoggerHook', interval=1000), )
+train_cfg = dict(type='EpochBasedTrainLoop', max_epochs=3, val_interval=1)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[1, 2],
+ gamma=0.1)
+]
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/openimagesv6.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/openimagesv6.py
new file mode 100644
index 0000000000000000000000000000000000000000..a65f942a0d4f8cfdaa3cfb712276d6de34d62a84
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/openimagesv6.py
@@ -0,0 +1,94 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/schedules/schedule_1x.py',
+ '../_base_/default_runtime.py',
+]
+
+model = dict(roi_head=dict(bbox_head=dict(num_classes=601)))
+
+# dsdl dataset settings
+
+# please visit our platform [OpenDataLab](https://opendatalab.com/)
+# to downloaded dsdl dataset.
+dataset_type = 'DSDLDetDataset'
+data_root = 'data/OpenImages'
+train_ann = 'dsdl/set-train/train.yaml'
+val_ann = 'dsdl/set-val/val.yaml'
+specific_key_path = dict(
+ image_level_labels='./image_labels/*/label',
+ Label='./objects/*/label',
+ is_group_of='./objects/*/isgroupof',
+)
+
+backend_args = dict(
+ backend='petrel',
+ path_mapping=dict({'data/': 's3://open_dataset_original/'}))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='Resize', scale=(1024, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1024, 800), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'instances', 'image_level_labels'))
+]
+
+train_dataloader = dict(
+ sampler=dict(type='ClassAwareSampler', num_sample_class=1),
+ dataset=dict(
+ type=dataset_type,
+ with_imagelevel_label=True,
+ with_hierarchy=True,
+ specific_key_path=specific_key_path,
+ data_root=data_root,
+ ann_file=train_ann,
+ filter_cfg=dict(filter_empty_gt=True, min_size=32, bbox_min_size=32),
+ pipeline=train_pipeline))
+
+val_dataloader = dict(
+ dataset=dict(
+ type=dataset_type,
+ with_imagelevel_label=True,
+ with_hierarchy=True,
+ specific_key_path=specific_key_path,
+ data_root=data_root,
+ ann_file=val_ann,
+ test_mode=True,
+ pipeline=test_pipeline))
+
+test_dataloader = val_dataloader
+
+default_hooks = dict(logger=dict(type='LoggerHook', interval=1000), )
+train_cfg = dict(type='EpochBasedTrainLoop', max_epochs=3, val_interval=1)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[1, 2],
+ gamma=0.1)
+]
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
+
+val_evaluator = dict(
+ type='OpenImagesMetric',
+ iou_thrs=0.5,
+ ioa_thrs=0.5,
+ use_group_of=True,
+ get_supercategory=True)
+
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/voc07.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/voc07.py
new file mode 100644
index 0000000000000000000000000000000000000000..b7b864714e4987ca9d31eda5fee746e741b7aa10
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/voc07.py
@@ -0,0 +1,94 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py', '../_base_/default_runtime.py'
+]
+
+# model setting
+model = dict(roi_head=dict(bbox_head=dict(num_classes=20)))
+
+# dsdl dataset settings
+
+# please visit our platform [OpenDataLab](https://opendatalab.com/)
+# to downloaded dsdl dataset.
+dataset_type = 'DSDLDetDataset'
+data_root = 'data/VOC07-det'
+img_prefix = 'original'
+train_ann = 'dsdl/set-train/train.yaml'
+val_ann = 'dsdl/set-test/test.yaml'
+
+specific_key_path = dict(ignore_flag='./objects/*/difficult')
+
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='Resize', scale=(1000, 600), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1000, 600), keep_ratio=True),
+ # avoid bboxes being resized
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'instances'))
+]
+train_dataloader = dict(
+ dataset=dict(
+ type=dataset_type,
+ specific_key_path=specific_key_path,
+ data_root=data_root,
+ ann_file=train_ann,
+ data_prefix=dict(img_path=img_prefix),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32, bbox_min_size=32),
+ pipeline=train_pipeline))
+
+val_dataloader = dict(
+ dataset=dict(
+ type=dataset_type,
+ specific_key_path=specific_key_path,
+ data_root=data_root,
+ ann_file=val_ann,
+ data_prefix=dict(img_path=img_prefix),
+ test_mode=True,
+ pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# Pascal VOC2007 uses `11points` as default evaluate mode, while PASCAL
+# VOC2012 defaults to use 'area'.
+val_evaluator = dict(type='VOCMetric', metric='mAP', eval_mode='11points')
+# val_evaluator = dict(type='CocoMetric', metric='bbox')
+test_evaluator = val_evaluator
+
+# training schedule, voc dataset is repeated 3 times, in
+# `_base_/datasets/voc0712.py`, so the actual epoch = 4 * 3 = 12
+max_epochs = 12
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=3)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[9],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/voc0712.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/voc0712.py
new file mode 100644
index 0000000000000000000000000000000000000000..9ec1bb8f98e56d0402c9a80934c3b77bd7919fa4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dsdl/voc0712.py
@@ -0,0 +1,132 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/schedules/schedule_1x.py',
+ '../_base_/default_runtime.py',
+ # '../_base_/datasets/dsdl.py'
+]
+
+# model setting
+model = dict(roi_head=dict(bbox_head=dict(num_classes=20)))
+
+# dsdl dataset settings
+
+# please visit our platform [OpenDataLab](https://opendatalab.com/)
+# to downloaded dsdl dataset.
+dataset_type = 'DSDLDetDataset'
+data_root_07 = 'data/VOC07-det'
+data_root_12 = 'data/VOC12-det'
+img_prefix = 'original'
+
+train_ann = 'dsdl/set-train/train.yaml'
+val_ann = 'dsdl/set-val/val.yaml'
+test_ann = 'dsdl/set-test/test.yaml'
+
+backend_args = None
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='Resize', scale=(1000, 600), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(1000, 600), keep_ratio=True),
+ # If you don't have a gt annotation, delete the pipeline
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'instances'))
+]
+
+specific_key_path = dict(ignore_flag='./objects/*/difficult', )
+
+train_dataloader = dict(
+ dataset=dict(
+ type='RepeatDataset',
+ times=3,
+ dataset=dict(
+ type='ConcatDataset',
+ datasets=[
+ dict(
+ type=dataset_type,
+ specific_key_path=specific_key_path,
+ data_root=data_root_07,
+ ann_file=train_ann,
+ data_prefix=dict(img_path=img_prefix),
+ filter_cfg=dict(
+ filter_empty_gt=True, min_size=32, bbox_min_size=32),
+ pipeline=train_pipeline),
+ dict(
+ type=dataset_type,
+ specific_key_path=specific_key_path,
+ data_root=data_root_07,
+ ann_file=val_ann,
+ data_prefix=dict(img_path=img_prefix),
+ filter_cfg=dict(
+ filter_empty_gt=True, min_size=32, bbox_min_size=32),
+ pipeline=train_pipeline),
+ dict(
+ type=dataset_type,
+ specific_key_path=specific_key_path,
+ data_root=data_root_12,
+ ann_file=train_ann,
+ data_prefix=dict(img_path=img_prefix),
+ filter_cfg=dict(
+ filter_empty_gt=True, min_size=32, bbox_min_size=32),
+ pipeline=train_pipeline),
+ dict(
+ type=dataset_type,
+ specific_key_path=specific_key_path,
+ data_root=data_root_12,
+ ann_file=val_ann,
+ data_prefix=dict(img_path=img_prefix),
+ filter_cfg=dict(
+ filter_empty_gt=True, min_size=32, bbox_min_size=32),
+ pipeline=train_pipeline),
+ ])))
+
+val_dataloader = dict(
+ dataset=dict(
+ type=dataset_type,
+ specific_key_path=specific_key_path,
+ data_root=data_root_07,
+ ann_file=test_ann,
+ test_mode=True,
+ pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(type='CocoMetric', metric='bbox')
+# val_evaluator = dict(type='VOCMetric', metric='mAP', eval_mode='11points')
+test_evaluator = val_evaluator
+
+# training schedule, voc dataset is repeated 3 times, in
+# `_base_/datasets/voc0712.py`, so the actual epoch = 4 * 3 = 12
+max_epochs = 4
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[3],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dyhead/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dyhead/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..decd48051f0b10ef3f9e6de8ad7476e59fb89511
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dyhead/README.md
@@ -0,0 +1,52 @@
+# DyHead
+
+> [Dynamic Head: Unifying Object Detection Heads with Attentions](https://arxiv.org/abs/2106.08322)
+
+
+
+## Abstract
+
+The complex nature of combining localization and classification in object detection has resulted in the flourished development of methods. Previous works tried to improve the performance in various object detection heads but failed to present a unified view. In this paper, we present a novel dynamic head framework to unify object detection heads with attentions. By coherently combining multiple self-attention mechanisms between feature levels for scale-awareness, among spatial locations for spatial-awareness, and within output channels for task-awareness, the proposed approach significantly improves the representation ability of object detection heads without any computational overhead. Further experiments demonstrate that the effectiveness and efficiency of the proposed dynamic head on the COCO benchmark. With a standard ResNeXt-101-DCN backbone, we largely improve the performance over popular object detectors and achieve a new state-of-the-art at 54.0 AP. Furthermore, with latest transformer backbone and extra data, we can push current best COCO result to a new record at 60.6 AP.
+
+
+

+
+
+## Results and Models
+
+| Method | Backbone | Style | Setting | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :----: | :------: | :-----: | :----------: | :-----: | :------: | :------------: | :----: | :----------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| ATSS | R-50 | caffe | reproduction | 1x | 5.4 | 13.2 | 42.5 | [config](./atss_r50-caffe_fpn_dyhead_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/dyhead/atss_r50_fpn_dyhead_for_reproduction_1x_coco/atss_r50_fpn_dyhead_for_reproduction_4x4_1x_coco_20220107_213939-162888e6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/dyhead/atss_r50_fpn_dyhead_for_reproduction_1x_coco/atss_r50_fpn_dyhead_for_reproduction_4x4_1x_coco_20220107_213939.log.json) |
+| ATSS | R-50 | pytorch | simple | 1x | 4.9 | 13.7 | 43.3 | [config](./atss_r50_fpn_dyhead_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/dyhead/atss_r50_fpn_dyhead_4x4_1x_coco/atss_r50_fpn_dyhead_4x4_1x_coco_20211219_023314-eaa620c6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/dyhead/atss_r50_fpn_dyhead_4x4_1x_coco/atss_r50_fpn_dyhead_4x4_1x_coco_20211219_023314.log.json) |
+
+- We trained the above models with 4 GPUs and 4 `samples_per_gpu`.
+- The `reproduction` setting aims to reproduce the official implementation based on Detectron2.
+- The `simple` setting serves as a minimum example to use DyHead in MMDetection. Specifically,
+ - it adds `DyHead` to `neck` after `FPN`
+ - it sets `stacked_convs=0` to `bbox_head`
+- The `simple` setting achieves higher AP than the original implementation.
+ We have not conduct ablation study between the two settings.
+ `dict(type='Pad', size_divisor=128)` may further improve AP by prefer spatial alignment across pyramid levels, although large padding reduces efficiency.
+
+We also trained the model with Swin-L backbone. Results are as below.
+
+| Method | Backbone | Style | Setting | Lr schd | mstrain | box AP | Config | Download |
+| :----: | :------: | :---: | :----------: | :-----: | :------: | :----: | :-----------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| ATSS | Swin-L | caffe | reproduction | 2x | 480~1200 | 56.2 | [config](./atss_swin-l-p4-w12_fpn_dyhead_ms-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/dyhead/atss_swin-l-p4-w12_fpn_dyhead_mstrain_2x_coco/atss_swin-l-p4-w12_fpn_dyhead_mstrain_2x_coco_20220509_100315-bc5b6516.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/dyhead/atss_swin-l-p4-w12_fpn_dyhead_mstrain_2x_coco/atss_swin-l-p4-w12_fpn_dyhead_mstrain_2x_coco_20220509_100315.log.json) |
+
+## Relation to Other Methods
+
+- DyHead can be regarded as an improved [SEPC](https://arxiv.org/abs/2005.03101) with [DyReLU modules](https://arxiv.org/abs/2003.10027) and simplified [SE blocks](https://arxiv.org/abs/1709.01507).
+- Xiyang Dai et al., the author team of DyHead, adopt it for [Dynamic DETR](https://openaccess.thecvf.com/content/ICCV2021/html/Dai_Dynamic_DETR_End-to-End_Object_Detection_With_Dynamic_Attention_ICCV_2021_paper.html).
+ The description of Dynamic Encoder in Sec. 3.2 will help you understand DyHead.
+
+## Citation
+
+```latex
+@inproceedings{DyHead_CVPR2021,
+ author = {Dai, Xiyang and Chen, Yinpeng and Xiao, Bin and Chen, Dongdong and Liu, Mengchen and Yuan, Lu and Zhang, Lei},
+ title = {Dynamic Head: Unifying Object Detection Heads With Attentions},
+ booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
+ year = {2021}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dyhead/atss_r50-caffe_fpn_dyhead_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dyhead/atss_r50-caffe_fpn_dyhead_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8716f1226cb0b37435d0318d62599a74e6126f19
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dyhead/atss_r50-caffe_fpn_dyhead_1x_coco.py
@@ -0,0 +1,103 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ type='ATSS',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=128),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')),
+ neck=[
+ dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output',
+ num_outs=5),
+ dict(
+ type='DyHead',
+ in_channels=256,
+ out_channels=256,
+ num_blocks=6,
+ # disable zero_init_offset to follow official implementation
+ zero_init_offset=False)
+ ],
+ bbox_head=dict(
+ type='ATSSHead',
+ num_classes=80,
+ in_channels=256,
+ pred_kernel_size=1, # follow DyHead official implementation
+ stacked_convs=0,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ octave_base_scale=8,
+ scales_per_octave=1,
+ strides=[8, 16, 32, 64, 128],
+ center_offset=0.5), # follow DyHead official implementation
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=2.0),
+ loss_centerness=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(type='ATSSAssigner', topk=9),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+
+# optimizer
+optim_wrapper = dict(optimizer=dict(lr=0.01))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True, backend='pillow'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True, backend='pillow'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dyhead/atss_r50_fpn_dyhead_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dyhead/atss_r50_fpn_dyhead_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..89e89b98ca437bb13fe5d01acc05cfdcd04e8fa0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dyhead/atss_r50_fpn_dyhead_1x_coco.py
@@ -0,0 +1,72 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ type='ATSS',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=[
+ dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output',
+ num_outs=5),
+ dict(type='DyHead', in_channels=256, out_channels=256, num_blocks=6)
+ ],
+ bbox_head=dict(
+ type='ATSSHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=0,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ octave_base_scale=8,
+ scales_per_octave=1,
+ strides=[8, 16, 32, 64, 128]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=2.0),
+ loss_centerness=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(type='ATSSAssigner', topk=9),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+
+# optimizer
+optim_wrapper = dict(optimizer=dict(lr=0.01))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dyhead/atss_swin-l-p4-w12_fpn_dyhead_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dyhead/atss_swin-l-p4-w12_fpn_dyhead_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f537b9dc9b17aa50f0044b874585fe1e0ba15216
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dyhead/atss_swin-l-p4-w12_fpn_dyhead_ms-2x_coco.py
@@ -0,0 +1,140 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_large_patch4_window12_384_22k.pth' # noqa
+model = dict(
+ type='ATSS',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=128),
+ backbone=dict(
+ type='SwinTransformer',
+ pretrain_img_size=384,
+ embed_dims=192,
+ depths=[2, 2, 18, 2],
+ num_heads=[6, 12, 24, 48],
+ window_size=12,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.2,
+ patch_norm=True,
+ out_indices=(1, 2, 3),
+ # Please only add indices that would be used
+ # in FPN, otherwise some parameter will not be used
+ with_cp=False,
+ convert_weights=True,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ neck=[
+ dict(
+ type='FPN',
+ in_channels=[384, 768, 1536],
+ out_channels=256,
+ start_level=0,
+ add_extra_convs='on_output',
+ num_outs=5),
+ dict(
+ type='DyHead',
+ in_channels=256,
+ out_channels=256,
+ num_blocks=6,
+ # disable zero_init_offset to follow official implementation
+ zero_init_offset=False)
+ ],
+ bbox_head=dict(
+ type='ATSSHead',
+ num_classes=80,
+ in_channels=256,
+ pred_kernel_size=1, # follow DyHead official implementation
+ stacked_convs=0,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ octave_base_scale=8,
+ scales_per_octave=1,
+ strides=[8, 16, 32, 64, 128],
+ center_offset=0.5), # follow DyHead official implementation
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=2.0),
+ loss_centerness=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(type='ATSSAssigner', topk=9),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize',
+ scale=[(2000, 480), (2000, 1200)],
+ keep_ratio=True,
+ backend='pillow'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(2000, 1200), keep_ratio=True, backend='pillow'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ dataset=dict(
+ _delete_=True,
+ type='RepeatDataset',
+ times=2,
+ dataset=dict(
+ type={{_base_.dataset_type}},
+ data_root={{_base_.data_root}},
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args={{_base_.backend_args}})))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# optimizer
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(
+ type='AdamW', lr=0.00005, betas=(0.9, 0.999), weight_decay=0.05),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'relative_position_bias_table': dict(decay_mult=0.),
+ 'norm': dict(decay_mult=0.)
+ }),
+ clip_grad=None)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dyhead/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dyhead/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..28b5a5821c81cea3213494c712910f904ae117f2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dyhead/metafile.yml
@@ -0,0 +1,76 @@
+Collections:
+ - Name: DyHead
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 4x T4 GPUs
+ Architecture:
+ - ATSS
+ - DyHead
+ - FPN
+ - ResNet
+ - Deformable Convolution
+ - Pyramid Convolution
+ Paper:
+ URL: https://arxiv.org/abs/2106.08322
+ Title: 'Dynamic Head: Unifying Object Detection Heads with Attentions'
+ README: configs/dyhead/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.22.0/mmdet/models/necks/dyhead.py#L130
+ Version: v2.22.0
+
+Models:
+ - Name: atss_r50-caffe_fpn_dyhead_1x_coco
+ In Collection: DyHead
+ Config: configs/dyhead/atss_r50-caffe_fpn_dyhead_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.4
+ inference time (ms/im):
+ - value: 75.7
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/dyhead/atss_r50_fpn_dyhead_for_reproduction_1x_coco/atss_r50_fpn_dyhead_for_reproduction_4x4_1x_coco_20220107_213939-162888e6.pth
+
+ - Name: atss_r50_fpn_dyhead_1x_coco
+ In Collection: DyHead
+ Config: configs/dyhead/atss_r50_fpn_dyhead_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.9
+ inference time (ms/im):
+ - value: 73.1
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/dyhead/atss_r50_fpn_dyhead_4x4_1x_coco/atss_r50_fpn_dyhead_4x4_1x_coco_20211219_023314-eaa620c6.pth
+
+ - Name: atss_swin-l-p4-w12_fpn_dyhead_ms-2x_coco
+ In Collection: DyHead
+ Config: configs/dyhead/atss_swin-l-p4-w12_fpn_dyhead_ms-2x_coco.py
+ Metadata:
+ Training Memory (GB): 58.4
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 56.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/dyhead/atss_swin-l-p4-w12_fpn_dyhead_mstrain_2x_coco/atss_swin-l-p4-w12_fpn_dyhead_mstrain_2x_coco_20220509_100315-bc5b6516.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dynamic_rcnn/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dynamic_rcnn/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..b5e803a2f27f07a1e49abfc9195965e33f36b73a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dynamic_rcnn/README.md
@@ -0,0 +1,30 @@
+# Dynamic R-CNN
+
+> [Dynamic R-CNN: Towards High Quality Object Detection via Dynamic Training](https://arxiv.org/abs/2004.06002)
+
+
+
+## Abstract
+
+Although two-stage object detectors have continuously advanced the state-of-the-art performance in recent years, the training process itself is far from crystal. In this work, we first point out the inconsistency problem between the fixed network settings and the dynamic training procedure, which greatly affects the performance. For example, the fixed label assignment strategy and regression loss function cannot fit the distribution change of proposals and thus are harmful to training high quality detectors. Consequently, we propose Dynamic R-CNN to adjust the label assignment criteria (IoU threshold) and the shape of regression loss function (parameters of SmoothL1 Loss) automatically based on the statistics of proposals during training. This dynamic design makes better use of the training samples and pushes the detector to fit more high quality samples. Specifically, our method improves upon ResNet-50-FPN baseline with 1.9% AP and 5.5% AP90 on the MS COCO dataset with no extra overhead.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :------: | :-----: | :-----: | :------: | :------------: | :----: | :-----------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | pytorch | 1x | 3.8 | | 38.9 | [config](./dynamic-rcnn_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/dynamic_rcnn/dynamic_rcnn_r50_fpn_1x/dynamic_rcnn_r50_fpn_1x-62a3f276.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/dynamic_rcnn/dynamic_rcnn_r50_fpn_1x/dynamic_rcnn_r50_fpn_1x_20200618_095048.log.json) |
+
+## Citation
+
+```latex
+@article{DynamicRCNN,
+ author = {Hongkai Zhang and Hong Chang and Bingpeng Ma and Naiyan Wang and Xilin Chen},
+ title = {Dynamic {R-CNN}: Towards High Quality Object Detection via Dynamic Training},
+ journal = {arXiv preprint arXiv:2004.06002},
+ year = {2020}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dynamic_rcnn/dynamic-rcnn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dynamic_rcnn/dynamic-rcnn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f64dfa0b9102d5f7b32793b9d21e19c67afdfc2a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dynamic_rcnn/dynamic-rcnn_r50_fpn_1x_coco.py
@@ -0,0 +1,28 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ roi_head=dict(
+ type='DynamicRoIHead',
+ bbox_head=dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=False,
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0))),
+ train_cfg=dict(
+ rpn_proposal=dict(nms=dict(iou_threshold=0.85)),
+ rcnn=dict(
+ dynamic_rcnn=dict(
+ iou_topk=75,
+ beta_topk=10,
+ update_iter_interval=100,
+ initial_iou=0.4,
+ initial_beta=1.0))),
+ test_cfg=dict(rpn=dict(nms=dict(iou_threshold=0.85))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/dynamic_rcnn/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dynamic_rcnn/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..64ab3b0ce490a25e227b3bcd60442669608fda22
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/dynamic_rcnn/metafile.yml
@@ -0,0 +1,35 @@
+Collections:
+ - Name: Dynamic R-CNN
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Dynamic R-CNN
+ - FPN
+ - RPN
+ - ResNet
+ - RoIAlign
+ Paper:
+ URL: https://arxiv.org/pdf/2004.06002
+ Title: 'Dynamic R-CNN: Towards High Quality Object Detection via Dynamic Training'
+ README: configs/dynamic_rcnn/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.2.0/mmdet/models/roi_heads/dynamic_roi_head.py#L11
+ Version: v2.2.0
+
+Models:
+ - Name: dynamic-rcnn_r50_fpn_1x_coco
+ In Collection: Dynamic R-CNN
+ Config: configs/dynamic_rcnn/dynamic-rcnn_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.8
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/dynamic_rcnn/dynamic_rcnn_r50_fpn_1x/dynamic_rcnn_r50_fpn_1x-62a3f276.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/efficientnet/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/efficientnet/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..941944db4f3fdc887da5ddc9647b3d619138478b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/efficientnet/README.md
@@ -0,0 +1,30 @@
+# EfficientNet
+
+> [EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks](https://arxiv.org/abs/1905.11946v5)
+
+
+
+## Introduction
+
+Convolutional Neural Networks (ConvNets) are commonly developed at a fixed resource budget, and then scaled up for better accuracy if more resources are available. In this paper, we systematically study model scaling and identify that carefully balancing network depth, width, and resolution can lead to better performance. Based on this observation, we propose a new scaling method that uniformly scales all dimensions of depth/width/resolution using a simple yet highly effective compound coefficient. We demonstrate the effectiveness of this method on scaling up MobileNets and ResNet.
+
+To go even further, we use neural architecture search to design a new baseline network and scale it up to obtain a family of models, called EfficientNets, which achieve much better accuracy and efficiency than previous ConvNets. In particular, our EfficientNet-B7 achieves state-of-the-art 84.3% top-1 accuracy on ImageNet, while being 8.4x smaller and 6.1x faster on inference than the best existing ConvNet. Our EfficientNets also transfer well and achieve state-of-the-art accuracy on CIFAR-100 (91.7%), Flowers (98.8%), and 3 other transfer learning datasets, with an order of magnitude fewer parameters.
+
+## Results and Models
+
+### RetinaNet
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :-------------: | :-----: | :-----: | :------: | :------------: | :----: | :-----------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Efficientnet-b3 | pytorch | 1x | - | - | 40.5 | [config](./retinanet_effb3_fpn_8xb4-crop896-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/efficientnet/retinanet_effb3_fpn_crop896_8x4_1x_coco/retinanet_effb3_fpn_crop896_8x4_1x_coco_20220322_234806-615a0dda.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/efficientnet/retinanet_effb3_fpn_crop896_8x4_1x_coco/retinanet_effb3_fpn_crop896_8x4_1x_coco_20220322_234806.log.json) |
+
+## Citation
+
+```latex
+@article{tan2019efficientnet,
+ title={Efficientnet: Rethinking model scaling for convolutional neural networks},
+ author={Tan, Mingxing and Le, Quoc V},
+ journal={arXiv preprint arXiv:1905.11946},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/efficientnet/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/efficientnet/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..6e220c8ad7cd0e25386d950c21616d4b92f8481e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/efficientnet/metafile.yml
@@ -0,0 +1,19 @@
+Models:
+ - Name: retinanet_effb3_fpn_8xb4-crop896-1x_coco
+ In Collection: RetinaNet
+ Config: configs/efficientnet/retinanet_effb3_fpn_8xb4-crop896-1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/efficientnet/retinanet_effb3_fpn_crop896_8x4_1x_coco/retinanet_effb3_fpn_crop896_8x4_1x_coco_20220322_234806-615a0dda.pth
+ Paper:
+ URL: https://arxiv.org/abs/1905.11946v5
+ Title: 'EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks'
+ README: configs/efficientnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.23.0/mmdet/models/backbones/efficientnet.py#L159
+ Version: v2.23.0
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/efficientnet/retinanet_effb3_fpn_8xb4-crop896-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/efficientnet/retinanet_effb3_fpn_8xb4-crop896-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..2d0d9cefd0b565b2cce42117eb872ac9373ea4b9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/efficientnet/retinanet_effb3_fpn_8xb4-crop896-1x_coco.py
@@ -0,0 +1,94 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/schedules/schedule_1x.py',
+ '../_base_/datasets/coco_detection.py', '../_base_/default_runtime.py'
+]
+
+image_size = (896, 896)
+batch_augments = [dict(type='BatchFixedSizePad', size=image_size)]
+norm_cfg = dict(type='BN', requires_grad=True)
+checkpoint = 'https://download.openmmlab.com/mmclassification/v0/efficientnet/efficientnet-b3_3rdparty_8xb32-aa_in1k_20220119-5b4887a0.pth' # noqa
+model = dict(
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32,
+ batch_augments=batch_augments),
+ backbone=dict(
+ _delete_=True,
+ type='EfficientNet',
+ arch='b3',
+ drop_path_rate=0.2,
+ out_indices=(3, 4, 5),
+ frozen_stages=0,
+ norm_cfg=dict(
+ type='SyncBN', requires_grad=True, eps=1e-3, momentum=0.01),
+ norm_eval=False,
+ init_cfg=dict(
+ type='Pretrained', prefix='backbone', checkpoint=checkpoint)),
+ neck=dict(
+ in_channels=[48, 136, 384],
+ start_level=0,
+ out_channels=256,
+ relu_before_extra_convs=True,
+ no_norm_on_lateral=True,
+ norm_cfg=norm_cfg),
+ bbox_head=dict(type='RetinaSepBNHead', num_ins=5, norm_cfg=norm_cfg),
+ # training and testing settings
+ train_cfg=dict(assigner=dict(neg_iou_thr=0.5)))
+
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize',
+ scale=image_size,
+ ratio_range=(0.8, 1.2),
+ keep_ratio=True),
+ dict(type='RandomCrop', crop_size=image_size),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=image_size, keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=4, num_workers=4, dataset=dict(pipeline=train_pipeline))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(lr=0.04),
+ paramwise_cfg=dict(norm_decay_mult=0, bypass_duplicate=True))
+
+# learning policy
+max_epochs = 12
+param_scheduler = [
+ dict(type='LinearLR', start_factor=0.1, by_epoch=False, begin=0, end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs)
+
+# cudnn_benchmark=True can accelerate fix-size training
+env_cfg = dict(cudnn_benchmark=True)
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (4 samples per GPU)
+auto_scale_lr = dict(base_batch_size=32)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/empirical_attention/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/empirical_attention/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..c0b4a68b6c35fefadc886c844d66d871eb90bef6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/empirical_attention/README.md
@@ -0,0 +1,33 @@
+# Empirical Attention
+
+> [An Empirical Study of Spatial Attention Mechanisms in Deep Networks](https://arxiv.org/abs/1904.05873)
+
+
+
+## Abstract
+
+Attention mechanisms have become a popular component in deep neural networks, yet there has been little examination of how different influencing factors and methods for computing attention from these factors affect performance. Toward a better general understanding of attention mechanisms, we present an empirical study that ablates various spatial attention elements within a generalized attention formulation, encompassing the dominant Transformer attention as well as the prevalent deformable convolution and dynamic convolution modules. Conducted on a variety of applications, the study yields significant findings about spatial attention in deep networks, some of which run counter to conventional understanding. For example, we find that the query and key content comparison in Transformer attention is negligible for self-attention, but vital for encoder-decoder attention. A proper combination of deformable convolution with key content only saliency achieves the best accuracy-efficiency tradeoff in self-attention. Our results suggest that there exists much room for improvement in the design of attention mechanisms.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Attention Component | DCN | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :------: | :-----------------: | :-: | :-----: | :------: | :------------: | :----: | :-----------------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | 1111 | N | 1x | 8.0 | 13.8 | 40.0 | [config](./faster-rcnn_r50-attn1111_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/empirical_attention/faster_rcnn_r50_fpn_attention_1111_1x_coco/faster_rcnn_r50_fpn_attention_1111_1x_coco_20200130-403cccba.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/empirical_attention/faster_rcnn_r50_fpn_attention_1111_1x_coco/faster_rcnn_r50_fpn_attention_1111_1x_coco_20200130_210344.log.json) |
+| R-50 | 0010 | N | 1x | 4.2 | 18.4 | 39.1 | [config](./faster-rcnn_r50-attn0010_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/empirical_attention/faster_rcnn_r50_fpn_attention_0010_1x_coco/faster_rcnn_r50_fpn_attention_0010_1x_coco_20200130-7cb0c14d.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/empirical_attention/faster_rcnn_r50_fpn_attention_0010_1x_coco/faster_rcnn_r50_fpn_attention_0010_1x_coco_20200130_210125.log.json) |
+| R-50 | 1111 | Y | 1x | 8.0 | 12.7 | 42.1 | [config](./faster-rcnn_r50-attn1111-dcn_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/empirical_attention/faster_rcnn_r50_fpn_attention_1111_dcn_1x_coco/faster_rcnn_r50_fpn_attention_1111_dcn_1x_coco_20200130-8b2523a6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/empirical_attention/faster_rcnn_r50_fpn_attention_1111_dcn_1x_coco/faster_rcnn_r50_fpn_attention_1111_dcn_1x_coco_20200130_204442.log.json) |
+| R-50 | 0010 | Y | 1x | 4.2 | 17.1 | 42.0 | [config](./faster-rcnn_r50-attn0010-dcn_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/empirical_attention/faster_rcnn_r50_fpn_attention_0010_dcn_1x_coco/faster_rcnn_r50_fpn_attention_0010_dcn_1x_coco_20200130-1a2e831d.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/empirical_attention/faster_rcnn_r50_fpn_attention_0010_dcn_1x_coco/faster_rcnn_r50_fpn_attention_0010_dcn_1x_coco_20200130_210410.log.json) |
+
+## Citation
+
+```latex
+@article{zhu2019empirical,
+ title={An Empirical Study of Spatial Attention Mechanisms in Deep Networks},
+ author={Zhu, Xizhou and Cheng, Dazhi and Zhang, Zheng and Lin, Stephen and Dai, Jifeng},
+ journal={arXiv preprint arXiv:1904.05873},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/empirical_attention/faster-rcnn_r50-attn0010-dcn_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/empirical_attention/faster-rcnn_r50-attn0010-dcn_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e1ae17a7ee4d3516e6aca90697fa165f592cf51e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/empirical_attention/faster-rcnn_r50-attn0010-dcn_fpn_1x_coco.py
@@ -0,0 +1,16 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ plugins=[
+ dict(
+ cfg=dict(
+ type='GeneralizedAttention',
+ spatial_range=-1,
+ num_heads=8,
+ attention_type='0010',
+ kv_stride=2),
+ stages=(False, False, True, True),
+ position='after_conv2')
+ ],
+ dcn=dict(type='DCN', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/empirical_attention/faster-rcnn_r50-attn0010_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/empirical_attention/faster-rcnn_r50-attn0010_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..7336d292eafe8c92407f831e712946a23e231db0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/empirical_attention/faster-rcnn_r50-attn0010_fpn_1x_coco.py
@@ -0,0 +1,13 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(plugins=[
+ dict(
+ cfg=dict(
+ type='GeneralizedAttention',
+ spatial_range=-1,
+ num_heads=8,
+ attention_type='0010',
+ kv_stride=2),
+ stages=(False, False, True, True),
+ position='after_conv2')
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/empirical_attention/faster-rcnn_r50-attn1111-dcn_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/empirical_attention/faster-rcnn_r50-attn1111-dcn_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..980e23d4509a19fe438d5c8494e2905d940705b1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/empirical_attention/faster-rcnn_r50-attn1111-dcn_fpn_1x_coco.py
@@ -0,0 +1,16 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ plugins=[
+ dict(
+ cfg=dict(
+ type='GeneralizedAttention',
+ spatial_range=-1,
+ num_heads=8,
+ attention_type='1111',
+ kv_stride=2),
+ stages=(False, False, True, True),
+ position='after_conv2')
+ ],
+ dcn=dict(type='DCN', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/empirical_attention/faster-rcnn_r50-attn1111_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/empirical_attention/faster-rcnn_r50-attn1111_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..426bc09fd64c16b43b33a5c797265aa9ec2c0c15
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/empirical_attention/faster-rcnn_r50-attn1111_fpn_1x_coco.py
@@ -0,0 +1,13 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(plugins=[
+ dict(
+ cfg=dict(
+ type='GeneralizedAttention',
+ spatial_range=-1,
+ num_heads=8,
+ attention_type='1111',
+ kv_stride=2),
+ stages=(False, False, True, True),
+ position='after_conv2')
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/empirical_attention/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/empirical_attention/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..b488da7d29fbd632da614895272cec2025b5eccc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/empirical_attention/metafile.yml
@@ -0,0 +1,103 @@
+Collections:
+ - Name: Empirical Attention
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Deformable Convolution
+ - FPN
+ - RPN
+ - ResNet
+ - RoIAlign
+ - Spatial Attention
+ Paper:
+ URL: https://arxiv.org/pdf/1904.05873
+ Title: 'An Empirical Study of Spatial Attention Mechanisms in Deep Networks'
+ README: configs/empirical_attention/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/ops/generalized_attention.py#L10
+ Version: v2.0.0
+
+Models:
+ - Name: faster-rcnn_r50_fpn_attention_1111_1x_coco
+ In Collection: Empirical Attention
+ Config: configs/empirical_attention/faster-rcnn_r50-attn1111_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 8.0
+ inference time (ms/im):
+ - value: 72.46
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/empirical_attention/faster_rcnn_r50_fpn_attention_1111_1x_coco/faster_rcnn_r50_fpn_attention_1111_1x_coco_20200130-403cccba.pth
+
+ - Name: faster-rcnn_r50_fpn_attention_0010_1x_coco
+ In Collection: Empirical Attention
+ Config: configs/empirical_attention/faster-rcnn_r50-attn0010_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.2
+ inference time (ms/im):
+ - value: 54.35
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/empirical_attention/faster_rcnn_r50_fpn_attention_0010_1x_coco/faster_rcnn_r50_fpn_attention_0010_1x_coco_20200130-7cb0c14d.pth
+
+ - Name: faster-rcnn_r50_fpn_attention_1111_dcn_1x_coco
+ In Collection: Empirical Attention
+ Config: configs/empirical_attention/faster-rcnn_r50-attn1111-dcn_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 8.0
+ inference time (ms/im):
+ - value: 78.74
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/empirical_attention/faster_rcnn_r50_fpn_attention_1111_dcn_1x_coco/faster_rcnn_r50_fpn_attention_1111_dcn_1x_coco_20200130-8b2523a6.pth
+
+ - Name: faster-rcnn_r50_fpn_attention_0010_dcn_1x_coco
+ In Collection: Empirical Attention
+ Config: configs/empirical_attention/faster-rcnn_r50-attn0010-dcn_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.2
+ inference time (ms/im):
+ - value: 58.48
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/empirical_attention/faster_rcnn_r50_fpn_attention_0010_dcn_1x_coco/faster_rcnn_r50_fpn_attention_0010_dcn_1x_coco_20200130-1a2e831d.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..0bdc9359c7c6e6100fa9f08397aa46e5c9999bac
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/README.md
@@ -0,0 +1,121 @@
+# Fast R-CNN
+
+> [Fast R-CNN](https://arxiv.org/abs/1504.08083)
+
+
+
+## Abstract
+
+This paper proposes a Fast Region-based Convolutional Network method (Fast R-CNN) for object detection. Fast R-CNN builds on previous work to efficiently classify object proposals using deep convolutional networks. Compared to previous work, Fast R-CNN employs several innovations to improve training and testing speed while also increasing detection accuracy. Fast R-CNN trains the very deep VGG16 network 9x faster than R-CNN, is 213x faster at test-time, and achieves a higher mAP on PASCAL VOC 2012. Compared to SPPnet, Fast R-CNN trains VGG16 3x faster, tests 10x faster, and is more accurate.
+
+
+

+
+
+## Introduction
+
+Before training the Fast R-CNN, users should first train an [RPN](../rpn/README.md), and use the RPN to extract the region proposals.
+The region proposals can be obtained by setting `DumpProposals` pseudo metric. The dumped results is a `dict(file_name: pred_instance)`.
+The `pred_instance` is an `InstanceData` containing the sorted boxes and scores predicted by RPN. We provide example of dumping proposals in [RPN config](../rpn/rpn_r50_fpn_1x_coco.py).
+
+- First, it should be obtained the region proposals in both training and validation (or testing) set.
+ change the type of `test_evaluator` to `DumpProposals` in the RPN config to get the region proposals as below:
+
+ The config of get training image region proposals can be set as below:
+
+ ```python
+ # For training set
+ val_dataloader = dict(
+ dataset=dict(
+ ann_file='data/coco/annotations/instances_train2017.json',
+ data_prefix=dict(img='val2017/')))
+ val_dataloader = dict(
+ _delete_=True,
+ type='DumpProposals',
+ output_dir='data/coco/proposals/',
+ proposals_file='rpn_r50_fpn_1x_train2017.pkl')
+ test_dataloader = val_dataloader
+ test_evaluator = val_dataloader
+ ```
+
+ The config of get validation image region proposals can be set as below:
+
+ ```python
+ # For validation set
+ val_dataloader = dict(
+ _delete_=True,
+ type='DumpProposals',
+ output_dir='data/coco/proposals/',
+ proposals_file='rpn_r50_fpn_1x_val2017.pkl')
+ test_evaluator = val_dataloader
+ ```
+
+ Extract the region proposals command can be set as below:
+
+ ```bash
+ ./tools/dist_test.sh \
+ configs/rpn_r50_fpn_1x_coco.py \
+ checkpoints/rpn_r50_fpn_1x_coco_20200218-5525fa2e.pth \
+ 8
+ ```
+
+ Users can refer to [test tutorial](https://mmdetection.readthedocs.io/en/latest/user_guides/test.html) for more details.
+
+- Then, modify the path of `proposal_file` in the dataset and using `ProposalBroadcaster` to process both ground truth bounding boxes and region proposals in pipelines.
+ An example of Fast R-CNN important setting can be seen as below:
+
+ ```python
+ train_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ backend_args={{_base_.backend_args}}),
+ dict(type='LoadProposals', num_max_proposals=2000),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='ProposalBroadcaster',
+ transforms=[
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ ]),
+ dict(type='PackDetInputs')
+ ]
+ test_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ backend_args={{_base_.backend_args}}),
+ dict(type='LoadProposals', num_max_proposals=None),
+ dict(
+ type='ProposalBroadcaster',
+ transforms=[
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ ]),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+ ]
+ train_dataloader = dict(
+ dataset=dict(
+ proposal_file='proposals/rpn_r50_fpn_1x_train2017.pkl',
+ pipeline=train_pipeline))
+ val_dataloader = dict(
+ dataset=dict(
+ proposal_file='proposals/rpn_r50_fpn_1x_val2017.pkl',
+ pipeline=test_pipeline))
+ test_dataloader = val_dataloader
+ ```
+
+- Finally, users can start training the Fast R-CNN.
+
+## Results and Models
+
+## Citation
+
+```latex
+@inproceedings{girshick2015fast,
+ title={Fast r-cnn},
+ author={Girshick, Ross},
+ booktitle={Proceedings of the IEEE international conference on computer vision},
+ year={2015}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/fast-rcnn_r101-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/fast-rcnn_r101-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..02c70296fca04d59b2b87801fa7834c0dc3d30f0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/fast-rcnn_r101-caffe_fpn_1x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './fast-rcnn_r50-caffe_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet101_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/fast-rcnn_r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/fast-rcnn_r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5af6b223c5bf66928a1d79ffba904d86006a3741
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/fast-rcnn_r101_fpn_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './fast-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/fast-rcnn_r101_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/fast-rcnn_r101_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..73425cf1ac3be429c69f6cf6b482fee91a8e2782
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/fast-rcnn_r101_fpn_2x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './fast-rcnn_r50_fpn_2x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/fast-rcnn_r50-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/fast-rcnn_r50-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3110f9fdf590ea665c9d7b7e28a56613cd79b786
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/fast-rcnn_r50-caffe_fpn_1x_coco.py
@@ -0,0 +1,16 @@
+_base_ = './fast-rcnn_r50_fpn_1x_coco.py'
+
+model = dict(
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ norm_cfg=dict(type='BN', requires_grad=False),
+ style='caffe',
+ norm_eval=True,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/fast-rcnn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/fast-rcnn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..daefe2d2d287b865b925263a81c12a6e30c58c4d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/fast-rcnn_r50_fpn_1x_coco.py
@@ -0,0 +1,39 @@
+_base_ = [
+ '../_base_/models/fast-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadProposals', num_max_proposals=2000),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='ProposalBroadcaster',
+ transforms=[
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ ]),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadProposals', num_max_proposals=None),
+ dict(
+ type='ProposalBroadcaster',
+ transforms=[
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ ]),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ dataset=dict(
+ proposal_file='proposals/rpn_r50_fpn_1x_train2017.pkl',
+ pipeline=train_pipeline))
+val_dataloader = dict(
+ dataset=dict(
+ proposal_file='proposals/rpn_r50_fpn_1x_val2017.pkl',
+ pipeline=test_pipeline))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/fast-rcnn_r50_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/fast-rcnn_r50_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d609a7c02d657e15316a4c5747983a4d9a10fc7c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fast_rcnn/fast-rcnn_r50_fpn_2x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './fast-rcnn_r50_fpn_1x_coco.py'
+
+train_cfg = dict(max_epochs=24)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=24,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..8bcdcf6d5120b65cc68c24b46e8d4d35447491fd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/README.md
@@ -0,0 +1,88 @@
+# Faster R-CNN
+
+> [Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks](https://arxiv.org/abs/1506.01497)
+
+
+
+## Abstract
+
+State-of-the-art object detection networks depend on region proposal algorithms to hypothesize object locations. Advances like SPPnet and Fast R-CNN have reduced the running time of these detection networks, exposing region proposal computation as a bottleneck. In this work, we introduce a Region Proposal Network (RPN) that shares full-image convolutional features with the detection network, thus enabling nearly cost-free region proposals. An RPN is a fully convolutional network that simultaneously predicts object bounds and objectness scores at each position. The RPN is trained end-to-end to generate high-quality region proposals, which are used by Fast R-CNN for detection. We further merge RPN and Fast R-CNN into a single network by sharing their convolutional features---using the recently popular terminology of neural networks with 'attention' mechanisms, the RPN component tells the unified network where to look. For the very deep VGG-16 model, our detection system has a frame rate of 5fps (including all steps) on a GPU, while achieving state-of-the-art object detection accuracy on PASCAL VOC 2007, 2012, and MS COCO datasets with only 300 proposals per image. In ILSVRC and COCO 2015 competitions, Faster R-CNN and RPN are the foundations of the 1st-place winning entries in several tracks.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :-------------: | :-----: | :-----: | :------: | :------------: | :----: | :-----------------------------------------------: | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-C4 | caffe | 1x | - | - | 35.6 | [config](./faster-rcnn_r50-caffe_c4-1x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_r50-caffe-c4_1x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_c4_1x_coco/faster_rcnn_r50_caffe_c4_1x_coco_20220316_150152.log.json) |
+| R-50-DC5 | caffe | 1x | - | - | 37.2 | [config](./faster-rcnn_r50-caffe-dc5_1x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_r50-caffe-dc5_1x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_dc5_1x_coco/faster_rcnn_r50_caffe_dc5_1x_coco_20201030_151909.log.json) |
+| R-50-FPN | caffe | 1x | 3.8 | | 37.8 | [config](./faster-rcnn_r50-caffe_fpn_1x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_r50-caffe_fpn_1x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_fpn_1x_coco/faster_rcnn_r50_caffe_fpn_1x_coco_20200504_180032.log.json) |
+| R-50-FPN | pytorch | 1x | 4.0 | 21.4 | 37.4 | [config](./faster-rcnn_r50_fpn_1x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_r50_fpn_1x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_1x_coco/faster_rcnn_r50_fpn_1x_coco_20200130_204655.log.json) |
+| R-50-FPN (FP16) | pytorch | 1x | 3.4 | 28.8 | 37.5 | [config](./faster-rcnn_r50_fpn_amp-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fp16/faster_rcnn_r50_fpn_fp16_1x_coco/faster_rcnn_r50_fpn_fp16_1x_coco_20200204-d4dc1471.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fp16/faster_rcnn_r50_fpn_fp16_1x_coco/faster_rcnn_r50_fpn_fp16_1x_coco_20200204_143530.log.json) |
+| R-50-FPN | pytorch | 2x | - | - | 38.4 | [config](./faster-rcnn_r50_fpn_2x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_r50_fpn_2x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_2x_coco/faster_rcnn_r50_fpn_2x_coco_20200504_210434.log.json) |
+| R-101-FPN | caffe | 1x | 5.7 | | 39.8 | [config](./faster-rcnn_r101-caffe_fpn_1x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_r101-caffe_fpn_1x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r101_caffe_fpn_1x_coco/faster_rcnn_r101_caffe_fpn_1x_coco_20200504_180057.log.json) |
+| R-101-FPN | pytorch | 1x | 6.0 | 15.6 | 39.4 | [config](./faster-rcnn_r101_fpn_1x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_r101_fpn_1x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r101_fpn_1x_coco/faster_rcnn_r101_fpn_1x_coco_20200130_204655.log.json) |
+| R-101-FPN | pytorch | 2x | - | - | 39.8 | [config](./faster-rcnn_r101_fpn_2x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_r101_fpn_2x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r101_fpn_2x_coco/faster_rcnn_r101_fpn_2x_coco_20200504_210455.log.json) |
+| X-101-32x4d-FPN | pytorch | 1x | 7.2 | 13.8 | 41.2 | [config](./faster-rcnn_x101-32x4d_fpn_1x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_x101-32x4d_fpn_1x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_x101_32x4d_fpn_1x_coco/faster_rcnn_x101_32x4d_fpn_1x_coco_20200203_000520.log.json) |
+| X-101-32x4d-FPN | pytorch | 2x | - | - | 41.2 | [config](./faster-rcnn_x101-32x4d_fpn_2x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_x101-32x4d_fpn_2x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_x101_32x4d_fpn_2x_coco/faster_rcnn_x101_32x4d_fpn_2x_coco_20200506_041400.log.json) |
+| X-101-64x4d-FPN | pytorch | 1x | 10.3 | 9.4 | 42.1 | [config](./faster-rcnn_x101-64x4d_fpn_1x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_x101-64x4d_fpn_1x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_x101_64x4d_fpn_1x_coco/faster_rcnn_x101_64x4d_fpn_1x_coco_20200204_134340.log.json) |
+| X-101-64x4d-FPN | pytorch | 2x | - | - | 41.6 | [config](./faster-rcnn_x101-64x4d_fpn_2x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_x101-64x4d_fpn_2x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_x101_64x4d_fpn_2x_coco/faster_rcnn_x101_64x4d_fpn_2x_coco_20200512_161033.log.json) |
+
+## Different regression loss
+
+We trained with R-50-FPN pytorch style backbone for 1x schedule.
+
+| Backbone | Loss type | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :------: | :------------: | :------: | :------------: | :----: | :----------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | L1Loss | 4.0 | 21.4 | 37.4 | [config](./faster-rcnn_r50_fpn_1x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_r50_fpn_1x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_1x_coco/faster_rcnn_r50_fpn_1x_coco_20200130_204655.log.json) |
+| R-50-FPN | IoULoss | | | 37.9 | [config](./faster-rcnn_r50_fpn_iou_1x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_r50_fpn_iou_1x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_iou_1x_coco/faster_rcnn_r50_fpn_iou_1x_coco_20200506_095954.log.json) |
+| R-50-FPN | GIoULoss | | | 37.6 | [config](./faster-rcnn_r50_fpn_giou_1x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_r50_fpn_giou_1x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_1x_coco/faster_rcnn_r50_fpn_giou_1x_coco_20200505_161120.log.json) |
+| R-50-FPN | BoundedIoULoss | | | 37.4 | [config](./faster-rcnn_r50_fpn_bounded-iou_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_1x_coco/faster_rcnn_r50_fpn_bounded_iou_1x_coco-98ad993b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_1x_coco/faster_rcnn_r50_fpn_bounded_iou_1x_coco_20200505_160738.log.json) |
+
+## Pre-trained Models
+
+We also train some models with longer schedules and multi-scale training. The users could finetune them for downstream tasks.
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :-----------------------------------------------------------: | :-----: | :-----: | :------: | :------------: | :----: | :--------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| [R-50-C4](./faster-rcnn_r50-caffe-c4_ms-1x_coco.py) | caffe | 1x | - | | 35.9 | [config](./faster-rcnn_r50-caffe-c4_ms-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_c4_mstrain_1x_coco/faster_rcnn_r50_caffe_c4_mstrain_1x_coco_20220316_150527-db276fed.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_c4_mstrain_1x_coco/faster_rcnn_r50_caffe_c4_mstrain_1x_coco_20220316_150527.log.json) |
+| [R-50-DC5](./faster-rcnn_r50-caffe-dc5_ms-1x_coco.py) | caffe | 1x | - | | 37.4 | [config](./faster-rcnn_r50-caffe-dc5_ms-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_dc5_mstrain_1x_coco/faster_rcnn_r50_caffe_dc5_mstrain_1x_coco_20201028_233851-b33d21b9.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_dc5_mstrain_1x_coco/faster_rcnn_r50_caffe_dc5_mstrain_1x_coco_20201028_233851.log.json) |
+| [R-50-DC5](./faster-rcnn_r50-caffe-dc5_ms-3x_coco.py) | caffe | 3x | - | | 38.7 | [config](./faster-rcnn_r50-caffe-dc5_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_dc5_mstrain_3x_coco/faster_rcnn_r50_caffe_dc5_mstrain_3x_coco_20201028_002107-34a53b2c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_dc5_mstrain_3x_coco/faster_rcnn_r50_caffe_dc5_mstrain_3x_coco_20201028_002107.log.json) |
+| [R-50-FPN](./faster-rcnn_r50-caffe_fpn_ms-2x_coco.py) | caffe | 2x | 3.7 | | 39.7 | [config](./faster-rcnn_r50-caffe_fpn_ms-2x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_r50-caffe_fpn_ms-2x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_fpn_mstrain_2x_coco/faster_rcnn_r50_caffe_fpn_mstrain_2x_coco_20200504_231813.log.json) |
+| [R-50-FPN](./faster-rcnn_r50-caffe_fpn_ms-3x_coco.py) | caffe | 3x | 3.7 | | 39.9 | [config](./faster-rcnn_r50-caffe_fpn_ms-3x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_r50-caffe_fpn_ms-3x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_fpn_mstrain_3x_coco/faster_rcnn_r50_caffe_fpn_mstrain_3x_coco_20210526_095054.log.json) |
+| [R-50-FPN](./faster-rcnn_r50_fpn_ms-3x_coco.py) | pytorch | 3x | 3.9 | | 40.3 | [config](./faster-rcnn_r50_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_mstrain_3x_coco/faster_rcnn_r50_fpn_mstrain_3x_coco_20210524_110822-e10bd31c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_mstrain_3x_coco/faster_rcnn_r50_fpn_mstrain_3x_coco_20210524_110822.log.json) |
+| [R-101-FPN](./faster-rcnn_r101-caffe_fpn_ms-3x_coco.py) | caffe | 3x | 5.6 | | 42.0 | [config](./faster-rcnn_r101-caffe_fpn_ms-3x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_r101-caffe_fpn_ms-3x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r101_caffe_fpn_mstrain_3x_coco/faster_rcnn_r101_caffe_fpn_mstrain_3x_coco_20210526_095742.log.json) |
+| [R-101-FPN](./faster-rcnn_r101_fpn_ms-3x_coco.py) | pytorch | 3x | 5.8 | | 41.8 | [config](./faster-rcnn_r101_fpn_ms-3x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_r101_fpn_ms-3x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r101_fpn_mstrain_3x_coco/faster_rcnn_r101_fpn_mstrain_3x_coco_20210524_110822.log.json) |
+| [X-101-32x4d-FPN](./faster-rcnn_x101-32x4d_fpn_ms-3x_coco.py) | pytorch | 3x | 7.0 | | 42.5 | [config](./faster-rcnn_x101-32x4d_fpn_ms-3x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_x101-32x4d_fpn_ms-3x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_x101_32x4d_fpn_mstrain_3x_coco/faster_rcnn_x101_32x4d_fpn_mstrain_3x_coco_20210524_124151.log.json) |
+| [X-101-32x8d-FPN](./faster-rcnn_x101-32x8d_fpn_ms-3x_coco.py) | pytorch | 3x | 10.1 | | 42.4 | [config](./faster-rcnn_x101-32x8d_fpn_ms-3x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_x101-32x8d_fpn_ms-3x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_x101_32x8d_fpn_mstrain_3x_coco/faster_rcnn_x101_32x8d_fpn_mstrain_3x_coco_20210604_182954.log.json) |
+| [X-101-64x4d-FPN](./faster-rcnn_x101-64x4d_fpn_ms-3x_coco.py) | pytorch | 3x | 10.0 | | 43.1 | [config](./faster-rcnn_x101-64x4d_fpn_ms-3x_coco.py) | [model](https://download.openxlab.org.cn/models/mmdetection/FasterR-CNN/weight/faster-rcnn_x101-64x4d_fpn_ms-3x_coco) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_x101_64x4d_fpn_mstrain_3x_coco/faster_rcnn_x101_64x4d_fpn_mstrain_3x_coco_20210524_124528.log.json) |
+
+We further finetune some pre-trained models on the COCO subsets, which only contain only a few of the 80 categories.
+
+| Backbone | Style | Class name | Pre-traind model | Mem (GB) | box AP | Config | Download |
+| ------------------------------------------------------------------------ | ----- | ------------------ | -------------------------------------------------------------- | -------- | ------ | ---------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
+| [R-50-FPN](./faster-rcnn_r50-caffe_fpn_ms-1x_coco-person.py) | caffe | person | [R-50-FPN-Caffe-3x](./faster-rcnn_r50-caffe_fpn_ms-3x_coco.py) | 3.7 | 55.8 | [config](./faster-rcnn_r50-caffe_fpn_ms-1x_coco-person.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_1x_coco-person/faster_rcnn_r50_fpn_1x_coco-person_20201216_175929-d022e227.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_1x_coco-person/faster_rcnn_r50_fpn_1x_coco-person_20201216_175929.log.json) |
+| [R-50-FPN](./faster-rcnn_r50-caffe_fpn_ms-1x_coco-person-bicycle-car.py) | caffe | person-bicycle-car | [R-50-FPN-Caffe-3x](./faster-rcnn_r50-caffe_fpn_ms-3x_coco.py) | 3.7 | 44.1 | [config](./faster-rcnn_r50-caffe_fpn_ms-1x_coco-person-bicycle-car.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_1x_coco-person-bicycle-car/faster_rcnn_r50_fpn_1x_coco-person-bicycle-car_20201216_173117-6eda6d92.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_1x_coco-person-bicycle-car/faster_rcnn_r50_fpn_1x_coco-person-bicycle-car_20201216_173117.log.json) |
+
+## Torchvision New Receipe (TNR)
+
+Torchvision released its high-precision ResNet models. The training details can be found on the [Pytorch website](https://pytorch.org/blog/how-to-train-state-of-the-art-models-using-torchvision-latest-primitives/). Here, we have done grid searches on learning rate and weight decay and found the optimal hyper-parameter on the detection task.
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :--------------------------------------------------: | :-----: | :-----: | :------: | :------------: | :----: | :------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| [R-50-TNR](./faster-rcnn_r50-tnr-pre_fpn_1x_coco.py) | pytorch | 1x | - | | 40.2 | [config](./faster-rcnn_r50-tnr-pre_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_tnr-pretrain_1x_coco/faster_rcnn_r50_fpn_tnr-pretrain_1x_coco_20220320_085147-efedfda4.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_tnr-pretrain_1x_coco/faster_rcnn_r50_fpn_tnr-pretrain_1x_coco_20220320_085147.log.json) |
+
+## Citation
+
+```latex
+@article{Ren_2017,
+ title={Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks},
+ journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
+ publisher={Institute of Electrical and Electronics Engineers (IEEE)},
+ author={Ren, Shaoqing and He, Kaiming and Girshick, Ross and Sun, Jian},
+ year={2017},
+ month={Jun},
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r101-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r101-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a18f1ada31ed2a2d1023d16470a271ad49c3be2e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r101-caffe_fpn_1x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './faster-rcnn_r50-caffe_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet101_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r101-caffe_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r101-caffe_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1cdb4d4973e364c4f37b80644388a4859f55772e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r101-caffe_fpn_ms-3x_coco.py
@@ -0,0 +1,11 @@
+_base_ = 'faster-rcnn_r50_fpn_ms-3x_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ norm_cfg=dict(requires_grad=False),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet101_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d113ae6295fdc3f3058ef498eb9b675154a05c12
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r101_fpn_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r101_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r101_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b471fb3cbd8a79165e0cd19afc3ba98bbcfeb74e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r101_fpn_2x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './faster-rcnn_r50_fpn_2x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r101_fpn_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r101_fpn_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a71d4afd3246d083bdf0f5a84be2fbf2340f621f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r101_fpn_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './faster-rcnn_r50_fpn_8xb8-amp-lsj-200e_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r101_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r101_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8ef6d1f8ea6b45e9a4bfe438910da827d079479b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r101_fpn_ms-3x_coco.py
@@ -0,0 +1,7 @@
+_base_ = 'faster-rcnn_r50_fpn_ms-3x_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r18_fpn_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r18_fpn_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..65515c9ace8bf4445a77db2485fc8d3f95c263b9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r18_fpn_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './faster-rcnn_r50_fpn_8xb8-amp-lsj-200e_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=18,
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet18')),
+ neck=dict(in_channels=[64, 128, 256, 512]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe-c4_ms-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe-c4_ms-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..7e231e865270acf0383e03a64f151efdbf88c29e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe-c4_ms-1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './faster-rcnn_r50-caffe_c4-1x_coco.py'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+_base_.train_dataloader.dataset.pipeline = train_pipeline
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe-dc5_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe-dc5_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8952a5c9c6c2fe019711968fa2aa7ed2065b13f6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe-dc5_1x_coco.py
@@ -0,0 +1,5 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50-caffe-dc5.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe-dc5_ms-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe-dc5_ms-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..63a68859a85fe5556e927c04aae5cafbef1fc0b6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe-dc5_ms-1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = 'faster-rcnn_r50-caffe-dc5_1x_coco.py'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+_base_.train_dataloader.dataset.pipeline = train_pipeline
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe-dc5_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe-dc5_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..27063468a70436a62a7cc54b8c8efc2de96ec33f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe-dc5_ms-3x_coco.py
@@ -0,0 +1,18 @@
+_base_ = './faster-rcnn_r50-caffe-dc5_ms-1x_coco.py'
+
+# MMEngine support the following two ways, users can choose
+# according to convenience
+# param_scheduler = [
+# dict(
+# type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500), # noqa
+# dict(
+# type='MultiStepLR',
+# begin=0,
+# end=12,
+# by_epoch=True,
+# milestones=[28, 34],
+# gamma=0.1)
+# ]
+_base_.param_scheduler[1].milestones = [28, 34]
+
+train_cfg = dict(max_epochs=36)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_c4-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_c4-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..0888fc01790af82a4c7131280ca5f0247b28d9fd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_c4-1x_coco.py
@@ -0,0 +1,5 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50-caffe-c4.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..9129a9583c52bf8ccab38a65f35c9f14bb128d07
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_1x_coco.py
@@ -0,0 +1,15 @@
+_base_ = './faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ norm_cfg=dict(requires_grad=False),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_90k_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_90k_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..27f49355f3be8f6a53038894405c5f1b3d9b46fa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_90k_coco.py
@@ -0,0 +1,22 @@
+_base_ = 'faster-rcnn_r50-caffe_fpn_1x_coco.py'
+max_iter = 90000
+
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_iter,
+ by_epoch=False,
+ milestones=[60000, 80000],
+ gamma=0.1)
+]
+
+train_cfg = dict(
+ _delete_=True,
+ type='IterBasedTrainLoop',
+ max_iters=max_iter,
+ val_interval=10000)
+default_hooks = dict(checkpoint=dict(by_epoch=False, interval=10000))
+log_processor = dict(by_epoch=False)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-1x_coco-person-bicycle-car.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-1x_coco-person-bicycle-car.py
new file mode 100644
index 0000000000000000000000000000000000000000..f36bb055f87aeadc43aa1233d1d3a7bdc33fbd80
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-1x_coco-person-bicycle-car.py
@@ -0,0 +1,16 @@
+_base_ = './faster-rcnn_r50-caffe_fpn_ms-1x_coco.py'
+model = dict(roi_head=dict(bbox_head=dict(num_classes=3)))
+metainfo = {
+ 'classes': ('person', 'bicycle', 'car'),
+ 'palette': [
+ (220, 20, 60),
+ (119, 11, 32),
+ (0, 0, 142),
+ ]
+}
+
+train_dataloader = dict(dataset=dict(metainfo=metainfo))
+val_dataloader = dict(dataset=dict(metainfo=metainfo))
+test_dataloader = dict(dataset=dict(metainfo=metainfo))
+
+load_from = 'https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_fpn_mstrain_3x_coco/faster_rcnn_r50_caffe_fpn_mstrain_3x_coco_bbox_mAP-0.398_20200504_163323-30042637.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-1x_coco-person.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-1x_coco-person.py
new file mode 100644
index 0000000000000000000000000000000000000000..9528b63f4deabb3610a26af59c856cee62c489c2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-1x_coco-person.py
@@ -0,0 +1,14 @@
+_base_ = './faster-rcnn_r50-caffe_fpn_ms-1x_coco.py'
+model = dict(roi_head=dict(bbox_head=dict(num_classes=1)))
+metainfo = {
+ 'classes': ('person', ),
+ 'palette': [
+ (220, 20, 60),
+ ]
+}
+
+train_dataloader = dict(dataset=dict(metainfo=metainfo))
+val_dataloader = dict(dataset=dict(metainfo=metainfo))
+test_dataloader = dict(dataset=dict(metainfo=metainfo))
+
+load_from = 'https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_fpn_mstrain_3x_coco/faster_rcnn_r50_caffe_fpn_mstrain_3x_coco_bbox_mAP-0.398_20200504_163323-30042637.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..59f1633c807f3eb904657cfaf97113c355df3fca
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-1x_coco.py
@@ -0,0 +1,31 @@
+_base_ = './faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ norm_cfg=dict(requires_grad=False),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+# MMEngine support the following two ways, users can choose
+# according to convenience
+# train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+_base_.train_dataloader.dataset.pipeline = train_pipeline
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..44d320ea01ba53d591ab7db29742e7fffc7c81ce
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-2x_coco.py
@@ -0,0 +1,18 @@
+_base_ = './faster-rcnn_r50-caffe_fpn_ms-1x_coco.py'
+
+# MMEngine support the following two ways, users can choose
+# according to convenience
+# param_scheduler = [
+# dict(
+# type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500), # noqa
+# dict(
+# type='MultiStepLR',
+# begin=0,
+# end=12,
+# by_epoch=True,
+# milestones=[16, 23],
+# gamma=0.1)
+# ]
+_base_.param_scheduler[1].milestones = [16, 23]
+
+train_cfg = dict(max_epochs=24)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..365f6439241c6374554af1fd58a114ef03448877
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-3x_coco.py
@@ -0,0 +1,15 @@
+_base_ = 'faster-rcnn_r50_fpn_ms-3x_coco.py'
+model = dict(
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ norm_cfg=dict(requires_grad=False),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-90k_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-90k_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6b9b3eb0e79b1ffb71d15c21274692d3b85e16ac
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-90k_coco.py
@@ -0,0 +1,23 @@
+_base_ = 'faster-rcnn_r50-caffe_fpn_ms-1x_coco.py'
+
+max_iter = 90000
+
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_iter,
+ by_epoch=False,
+ milestones=[60000, 80000],
+ gamma=0.1)
+]
+
+train_cfg = dict(
+ _delete_=True,
+ type='IterBasedTrainLoop',
+ max_iters=max_iter,
+ val_interval=10000)
+default_hooks = dict(checkpoint=dict(by_epoch=False, interval=10000))
+log_processor = dict(by_epoch=False)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-tnr-pre_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-tnr-pre_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..7b3e5dedbe81b927492dd41b13f017bcc2bd4c92
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50-tnr-pre_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+checkpoint = 'https://download.pytorch.org/models/resnet50-11ad3fa6.pth'
+model = dict(
+ backbone=dict(init_cfg=dict(type='Pretrained', checkpoint=checkpoint)))
+
+# `lr` and `weight_decay` have been searched to be optimal.
+optim_wrapper = dict(
+ optimizer=dict(_delete_=True, type='AdamW', lr=0.0001, weight_decay=0.1),
+ paramwise_cfg=dict(norm_decay_mult=0., bypass_duplicate=True))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8a45417fdd4566241114e20275990a5729486932
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py
@@ -0,0 +1,5 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..2981c6fbe16eb7a8b6ca1202ebb6325e2324c040
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_2x_coco.py
@@ -0,0 +1,5 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_2x.py', '../_base_/default_runtime.py'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3d366f3ba0e5ff098db3e409171a88860f1cf3af
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,20 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../common/lsj-200e_coco-detection.py'
+]
+image_size = (1024, 1024)
+batch_augments = [dict(type='BatchFixedSizePad', size=image_size)]
+
+model = dict(data_preprocessor=dict(batch_augments=batch_augments))
+
+train_dataloader = dict(batch_size=8, num_workers=4)
+# Enable automatic-mixed-precision training with AmpOptimWrapper.
+optim_wrapper = dict(
+ type='AmpOptimWrapper',
+ optimizer=dict(
+ type='SGD', lr=0.02 * 4, momentum=0.9, weight_decay=0.00004))
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_amp-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_amp-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f765deaef1db8a798c44d848c6f759755ccd4c45
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_amp-1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './faster-rcnn_r50_fpn_1x_coco.py'
+
+# MMEngine support the following two ways, users can choose
+# according to convenience
+# optim_wrapper = dict(type='AmpOptimWrapper')
+_base_.optim_wrapper.type = 'AmpOptimWrapper'
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_bounded-iou_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_bounded-iou_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..7758ca80b372e7895be267cad8c4603778d160b3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_bounded-iou_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ roi_head=dict(
+ bbox_head=dict(
+ reg_decoded_bbox=True,
+ loss_bbox=dict(type='BoundedIoULoss', loss_weight=10.0))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_ciou_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_ciou_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e8d8a3042750e8f5f9478b5e8c3111d8b7a10528
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_ciou_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ roi_head=dict(
+ bbox_head=dict(
+ reg_decoded_bbox=True,
+ loss_bbox=dict(type='CIoULoss', loss_weight=12.0))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_fcos-rpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_fcos-rpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b5a34d9f74a60388fa60afd8255d470c45f209f7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_fcos-rpn_1x_coco.py
@@ -0,0 +1,48 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ # copied from configs/fcos/fcos_r50-caffe_fpn_gn-head_1x_coco.py
+ neck=dict(
+ start_level=1,
+ add_extra_convs='on_output', # use P5
+ relu_before_extra_convs=True),
+ rpn_head=dict(
+ _delete_=True, # ignore the unused old settings
+ type='FCOSHead',
+ # num_classes = 1 for rpn,
+ # if num_classes > 1, it will be set to 1 in
+ # TwoStageDetector automatically
+ num_classes=1,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ strides=[8, 16, 32, 64, 128],
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='IoULoss', loss_weight=1.0),
+ loss_centerness=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0)),
+ roi_head=dict( # update featmap_strides
+ bbox_roi_extractor=dict(featmap_strides=[8, 16, 32, 64, 128])))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0,
+ end=1000), # Slowly increase lr, otherwise loss becomes NAN
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_giou_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_giou_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..82b71d77bfc448eceadcd03a6c8cbc4c8f871109
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_giou_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ roi_head=dict(
+ bbox_head=dict(
+ reg_decoded_bbox=True,
+ loss_bbox=dict(type='GIoULoss', loss_weight=10.0))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_iou_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_iou_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e21c43640cb7004e8e4ef189ff8843ad39de3c6f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_iou_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ roi_head=dict(
+ bbox_head=dict(
+ reg_decoded_bbox=True,
+ loss_bbox=dict(type='IoULoss', loss_weight=10.0))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..75dcfeb7a2310938c05cc103fadec6c6e119b90b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_ms-3x_coco.py
@@ -0,0 +1 @@
+_base_ = ['../common/ms_3x_coco.py', '../_base_/models/faster-rcnn_r50_fpn.py']
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_ohem_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_ohem_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..4f804b9be283015d4ec349f0df664e9ca7326c96
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_ohem_1x_coco.py
@@ -0,0 +1,2 @@
+_base_ = './faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(train_cfg=dict(rcnn=dict(sampler=dict(type='OHEMSampler'))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_soft-nms_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_soft-nms_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3775d8e447cb80c0fc28199be2abc4c23383eadd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_r50_fpn_soft-nms_1x_coco.py
@@ -0,0 +1,12 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ test_cfg=dict(
+ rcnn=dict(
+ score_thr=0.05,
+ nms=dict(type='soft_nms', iou_threshold=0.5),
+ max_per_img=100)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-32x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-32x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..395c98cd65cd5f883c9fe206a7b9c99e59acb32e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-32x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-32x4d_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-32x4d_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6232d0edba51f433a930c46d03c49fc27954303f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-32x4d_fpn_2x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './faster-rcnn_r50_fpn_2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-32x4d_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-32x4d_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..88cb40fd62a87a8af13e166df16a348c26e6d29e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-32x4d_fpn_ms-3x_coco.py
@@ -0,0 +1,14 @@
+_base_ = ['../common/ms_3x_coco.py', '../_base_/models/faster-rcnn_r50_fpn.py']
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-32x8d_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-32x8d_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..28d6290be7a75b7cceef8957e872e221fd3e78f5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-32x8d_fpn_ms-3x_coco.py
@@ -0,0 +1,23 @@
+_base_ = ['../common/ms_3x_coco.py', '../_base_/models/faster-rcnn_r50_fpn.py']
+model = dict(
+ # ResNeXt-101-32x8d model trained with Caffe2 at FB,
+ # so the mean and std need to be changed.
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[57.375, 57.120, 58.395],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=8,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnext101_32x8d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-64x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-64x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f39d6322fc3a4729ea7bbfefc207a6975efb4bf4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-64x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-64x4d_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-64x4d_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..97a3c1338fe294f66109fa92de0d8a48686b8a09
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-64x4d_fpn_2x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './faster-rcnn_r50_fpn_2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-64x4d_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-64x4d_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..eeaa218c9dc76123791d9e19b0ebae687cc296c9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/faster-rcnn_x101-64x4d_fpn_ms-3x_coco.py
@@ -0,0 +1,14 @@
+_base_ = ['../common/ms_3x_coco.py', '../_base_/models/faster-rcnn_r50_fpn.py']
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..6a201e177bad065235dd1346c1d36017c4359214
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/faster_rcnn/metafile.yml
@@ -0,0 +1,451 @@
+Collections:
+ - Name: Faster R-CNN
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - FPN
+ - RPN
+ - ResNet
+ - RoIPool
+ Paper:
+ URL: https://arxiv.org/abs/1506.01497
+ Title: "Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks"
+ README: configs/faster_rcnn/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/detectors/faster_rcnn.py#L6
+ Version: v2.0.0
+
+Models:
+ - Name: faster-rcnn_r50-caffe-c4_1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r50-caffe_c4-1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 35.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_c4_1x_coco/faster_rcnn_r50_caffe_c4_1x_coco_20220316_150152-3f885b85.pth
+
+ - Name: faster-rcnn_r50-caffe-c4_mstrain_1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r50-caffe-c4_ms-1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 35.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_c4_mstrain_1x_coco/faster_rcnn_r50_caffe_c4_mstrain_1x_coco_20220316_150527-db276fed.pth
+
+ - Name: faster-rcnn_r50-caffe-dc5_1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r50-caffe-dc5_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_dc5_1x_coco/faster_rcnn_r50_caffe_dc5_1x_coco_20201030_151909-531f0f43.pth
+
+ - Name: faster-rcnn_r50-caffe_fpn_1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.8
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_fpn_1x_coco/faster_rcnn_r50_caffe_fpn_1x_coco_bbox_mAP-0.378_20200504_180032-c5925ee5.pth
+
+ - Name: faster-rcnn_r50_fpn_1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.0
+ inference time (ms/im):
+ - value: 46.73
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_1x_coco/faster_rcnn_r50_fpn_1x_coco_20200130-047c8118.pth
+
+ - Name: faster-rcnn_r50_fpn_fp16_1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r50_fpn_amp-1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.4
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ - Mixed Precision Training
+ inference time (ms/im):
+ - value: 34.72
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP16
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fp16/faster_rcnn_r50_fpn_fp16_1x_coco/faster_rcnn_r50_fpn_fp16_1x_coco_20200204-d4dc1471.pth
+
+ - Name: faster-rcnn_r50_fpn_2x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r50_fpn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 4.0
+ inference time (ms/im):
+ - value: 46.73
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_2x_coco/faster_rcnn_r50_fpn_2x_coco_bbox_mAP-0.384_20200504_210434-a5d8aa15.pth
+
+ - Name: faster-rcnn_r101-caffe_fpn_1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r101-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.7
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r101_caffe_fpn_1x_coco/faster_rcnn_r101_caffe_fpn_1x_coco_bbox_mAP-0.398_20200504_180057-b269e9dd.pth
+
+ - Name: faster-rcnn_r101_fpn_1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r101_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.0
+ inference time (ms/im):
+ - value: 64.1
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r101_fpn_1x_coco/faster_rcnn_r101_fpn_1x_coco_20200130-f513f705.pth
+
+ - Name: faster-rcnn_r101_fpn_2x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r101_fpn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 6.0
+ inference time (ms/im):
+ - value: 64.1
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r101_fpn_2x_coco/faster_rcnn_r101_fpn_2x_coco_bbox_mAP-0.398_20200504_210455-1d2dac9c.pth
+
+ - Name: faster-rcnn_x101-32x4d_fpn_1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_x101-32x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.2
+ inference time (ms/im):
+ - value: 72.46
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_x101_32x4d_fpn_1x_coco/faster_rcnn_x101_32x4d_fpn_1x_coco_20200203-cff10310.pth
+
+ - Name: faster-rcnn_x101-32x4d_fpn_2x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_x101-32x4d_fpn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 7.2
+ inference time (ms/im):
+ - value: 72.46
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_x101_32x4d_fpn_2x_coco/faster_rcnn_x101_32x4d_fpn_2x_coco_bbox_mAP-0.412_20200506_041400-64a12c0b.pth
+
+ - Name: faster-rcnn_x101-64x4d_fpn_1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_x101-64x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 10.3
+ inference time (ms/im):
+ - value: 106.38
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_x101_64x4d_fpn_1x_coco/faster_rcnn_x101_64x4d_fpn_1x_coco_20200204-833ee192.pth
+
+ - Name: faster-rcnn_x101-64x4d_fpn_2x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_x101-64x4d_fpn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 10.3
+ inference time (ms/im):
+ - value: 106.38
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_x101_64x4d_fpn_2x_coco/faster_rcnn_x101_64x4d_fpn_2x_coco_20200512_161033-5961fa95.pth
+
+ - Name: faster-rcnn_r50_fpn_iou_1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r50_fpn_iou_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_iou_1x_coco/faster_rcnn_r50_fpn_iou_1x_coco_20200506_095954-938e81f0.pth
+
+ - Name: faster-rcnn_r50_fpn_giou_1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r50_fpn_giou_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_1x_coco/faster_rcnn_r50_fpn_giou_1x_coco-0eada910.pth
+
+ - Name: faster-rcnn_r50_fpn_bounded_iou_1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r50_fpn_bounded-iou_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_1x_coco/faster_rcnn_r50_fpn_bounded_iou_1x_coco-98ad993b.pth
+
+ - Name: faster-rcnn_r50-caffe-dc5_mstrain_1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r50-caffe-dc5_ms-1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_dc5_mstrain_1x_coco/faster_rcnn_r50_caffe_dc5_mstrain_1x_coco_20201028_233851-b33d21b9.pth
+
+ - Name: faster-rcnn_r50-caffe-dc5_mstrain_3x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r50-caffe-dc5_ms-3x_coco.py
+ Metadata:
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_dc5_mstrain_3x_coco/faster_rcnn_r50_caffe_dc5_mstrain_3x_coco_20201028_002107-34a53b2c.pth
+
+ - Name: faster-rcnn_r50-caffe_fpn_ms-2x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-2x_coco.py
+ Metadata:
+ Training Memory (GB): 4.3
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_fpn_mstrain_2x_coco/faster_rcnn_r50_caffe_fpn_mstrain_2x_coco_bbox_mAP-0.397_20200504_231813-10b2de58.pth
+
+ - Name: faster-rcnn_r50-caffe_fpn_ms-3x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r50-caffe_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 3.7
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_caffe_fpn_mstrain_3x_coco/faster_rcnn_r50_caffe_fpn_mstrain_3x_coco_20210526_095054-1f77628b.pth
+
+ - Name: faster-rcnn_r50_fpn_mstrain_3x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r50_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 3.9
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_mstrain_3x_coco/faster_rcnn_r50_fpn_mstrain_3x_coco_20210524_110822-e10bd31c.pth
+
+ - Name: faster-rcnn_r101-caffe_fpn_ms-3x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r101-caffe_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 5.6
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r101_caffe_fpn_mstrain_3x_coco/faster_rcnn_r101_caffe_fpn_mstrain_3x_coco_20210526_095742-a7ae426d.pth
+
+ - Name: faster-rcnn_r101_fpn_ms-3x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r101_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 5.8
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r101_fpn_mstrain_3x_coco/faster_rcnn_r101_fpn_mstrain_3x_coco_20210524_110822-4d4d2ca8.pth
+
+ - Name: faster-rcnn_x101-32x4d_fpn_ms-3x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_x101-32x4d_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 7.0
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_x101_32x4d_fpn_mstrain_3x_coco/faster_rcnn_x101_32x4d_fpn_mstrain_3x_coco_20210524_124151-16b9b260.pth
+
+ - Name: faster-rcnn_x101-32x8d_fpn_ms-3x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_x101-32x8d_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 10.1
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_x101_32x8d_fpn_mstrain_3x_coco/faster_rcnn_x101_32x8d_fpn_mstrain_3x_coco_20210604_182954-002e082a.pth
+
+ - Name: faster-rcnn_x101-64x4d_fpn_ms-3x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_x101-64x4d_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 10.0
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_x101_64x4d_fpn_mstrain_3x_coco/faster_rcnn_x101_64x4d_fpn_mstrain_3x_coco_20210524_124528-26c63de6.pth
+
+ - Name: faster-rcnn_r50_fpn_tnr-pretrain_1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/faster_rcnn/faster-rcnn_r50-tnr-pre_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.0
+ inference time (ms/im):
+ - value: 46.73
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_tnr-pretrain_1x_coco/faster_rcnn_r50_fpn_tnr-pretrain_1x_coco_20220320_085147-efedfda4.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..8d72237a059793385b43b04b7e77f3392fe30d5e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/README.md
@@ -0,0 +1,45 @@
+# FCOS
+
+> [FCOS: Fully Convolutional One-Stage Object Detection](https://arxiv.org/abs/1904.01355)
+
+
+
+## Abstract
+
+We propose a fully convolutional one-stage object detector (FCOS) to solve object detection in a per-pixel prediction fashion, analogue to semantic segmentation. Almost all state-of-the-art object detectors such as RetinaNet, SSD, YOLOv3, and Faster R-CNN rely on pre-defined anchor boxes. In contrast, our proposed detector FCOS is anchor box free, as well as proposal free. By eliminating the predefined set of anchor boxes, FCOS completely avoids the complicated computation related to anchor boxes such as calculating overlapping during training. More importantly, we also avoid all hyper-parameters related to anchor boxes, which are often very sensitive to the final detection performance. With the only post-processing non-maximum suppression (NMS), FCOS with ResNeXt-64x4d-101 achieves 44.7% in AP with single-model and single-scale testing, surpassing previous one-stage detectors with the advantage of being much simpler. For the first time, we demonstrate a much simpler and flexible detection framework achieving improved detection accuracy. We hope that the proposed FCOS framework can serve as a simple and strong alternative for many other instance-level tasks.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Style | GN | MS train | Tricks | DCN | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :------: | :---: | :-: | :------: | :----: | :-: | :-----: | :------: | :------------: | :----: | :------------------------------------------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | caffe | Y | N | N | N | 1x | 3.6 | 22.7 | 36.6 | [config](./fcos_r50-caffe_fpn_gn-head_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_r50_caffe_fpn_gn-head_1x_coco/fcos_r50_caffe_fpn_gn-head_1x_coco-821213aa.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_r50_caffe_fpn_gn-head_1x_coco/20201227_180009.log.json) |
+| R-50 | caffe | Y | N | Y | N | 1x | 3.7 | - | 38.7 | [config](./fcos_r50-caffe_fpn_gn-head-center-normbbox-centeronreg-giou_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_center-normbbox-centeronreg-giou_r50_caffe_fpn_gn-head_1x_coco/fcos_center-normbbox-centeronreg-giou_r50_caffe_fpn_gn-head_1x_coco-0a0d75a8.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_center-normbbox-centeronreg-giou_r50_caffe_fpn_gn-head_1x_coco/20210105_135818.log.json) |
+| R-50 | caffe | Y | N | Y | Y | 1x | 3.8 | - | 42.3 | [config](./fcos_r50-dcn-caffe_fpn_gn-head-center-normbbox-centeronreg-giou_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_center-normbbox-centeronreg-giou_r50_caffe_fpn_gn-head_dcn_1x_coco/fcos_center-normbbox-centeronreg-giou_r50_caffe_fpn_gn-head_dcn_1x_coco-ae4d8b3d.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_center-normbbox-centeronreg-giou_r50_caffe_fpn_gn-head_dcn_1x_coco/20210105_224556.log.json) |
+| R-101 | caffe | Y | N | N | N | 1x | 5.5 | 17.3 | 39.1 | [config](./fcos_r101-caffe_fpn_gn-head-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_r101_caffe_fpn_gn-head_1x_coco/fcos_r101_caffe_fpn_gn-head_1x_coco-0e37b982.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_r101_caffe_fpn_gn-head_1x_coco/20210103_155046.log.json) |
+
+| Backbone | Style | GN | MS train | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :------: | :-----: | :-: | :------: | :-----: | :------: | :------------: | :----: | :-----------------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | caffe | Y | Y | 2x | 2.6 | 22.9 | 38.5 | [config](./fcos_r50-caffe_fpn_gn-head_ms-640-800-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_r50_caffe_fpn_gn-head_mstrain_640-800_2x_coco/fcos_r50_caffe_fpn_gn-head_mstrain_640-800_2x_coco-d92ceeea.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_r50_caffe_fpn_gn-head_mstrain_640-800_2x_coco/20201227_161900.log.json) |
+| R-101 | caffe | Y | Y | 2x | 5.5 | 17.3 | 40.8 | [config](./fcos_r101-caffe_fpn_gn-head_ms-640-800-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_r101_caffe_fpn_gn-head_mstrain_640-800_2x_coco/fcos_r101_caffe_fpn_gn-head_mstrain_640-800_2x_coco-511424d6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_r101_caffe_fpn_gn-head_mstrain_640-800_2x_coco/20210103_155046.log.json) |
+| X-101 | pytorch | Y | Y | 2x | 10.0 | 9.7 | 42.6 | [config](./fcos_x101-64x4d_fpn_gn-head_ms-640-800-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_x101_64x4d_fpn_gn-head_mstrain_640-800_2x_coco/fcos_x101_64x4d_fpn_gn-head_mstrain_640-800_2x_coco-ede514a8.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_x101_64x4d_fpn_gn-head_mstrain_640-800_2x_coco/20210114_133041.log.json) |
+
+**Notes:**
+
+- The X-101 backbone is X-101-64x4d.
+- Tricks means setting `norm_on_bbox`, `centerness_on_reg`, `center_sampling` as `True`.
+- DCN means using `DCNv2` in both backbone and head.
+
+## Citation
+
+```latex
+@article{tian2019fcos,
+ title={FCOS: Fully Convolutional One-Stage Object Detection},
+ author={Tian, Zhi and Shen, Chunhua and Chen, Hao and He, Tong},
+ journal={arXiv preprint arXiv:1904.01355},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r101-caffe_fpn_gn-head-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r101-caffe_fpn_gn-head-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5380e87483e494b4c0bc6d8846c6892811d581d3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r101-caffe_fpn_gn-head-1x_coco.py
@@ -0,0 +1,9 @@
+_base_ = './fcos_r50-caffe_fpn_gn-head_1x_coco.py'
+
+# model settings
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron/resnet101_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r101-caffe_fpn_gn-head_ms-640-800-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r101-caffe_fpn_gn-head_ms-640-800-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..286a07a2db2c6fc423f6cf039b2609ac81ede73d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r101-caffe_fpn_gn-head_ms-640-800-2x_coco.py
@@ -0,0 +1,38 @@
+_base_ = './fcos_r50-caffe_fpn_gn-head_1x_coco.py'
+
+# model settings
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron/resnet101_caffe')))
+
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+# training schedule for 2x
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(type='ConstantLR', factor=1.0 / 3, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r101_fpn_gn-head-center-normbbox-centeronreg-giou_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r101_fpn_gn-head-center-normbbox-centeronreg-giou_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..77250e6917812d3494c8dabd52a3ed12f5f34483
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r101_fpn_gn-head-center-normbbox-centeronreg-giou_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './fcos_r50_fpn_gn-head-center-normbbox-centeronreg-giou_8xb8-amp-lsj-200e_coco.py' # noqa
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r18_fpn_gn-head-center-normbbox-centeronreg-giou_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r18_fpn_gn-head-center-normbbox-centeronreg-giou_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6f001024bb702c5ed0cb1103c5e10ae3cd7f599b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r18_fpn_gn-head-center-normbbox-centeronreg-giou_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './fcos_r50_fpn_gn-head-center-normbbox-centeronreg-giou_8xb8-amp-lsj-200e_coco.py' # noqa
+
+model = dict(
+ backbone=dict(
+ depth=18,
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet18')),
+ neck=dict(in_channels=[64, 128, 256, 512]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50-caffe_fpn_gn-head-center-normbbox-centeronreg-giou_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50-caffe_fpn_gn-head-center-normbbox-centeronreg-giou_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..2a77641dd87142d5c6d508f2f4a4ba5b70db52c1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50-caffe_fpn_gn-head-center-normbbox-centeronreg-giou_1x_coco.py
@@ -0,0 +1,43 @@
+_base_ = 'fcos_r50-caffe_fpn_gn-head_1x_coco.py'
+
+# model setting
+model = dict(
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')),
+ bbox_head=dict(
+ norm_on_bbox=True,
+ centerness_on_reg=True,
+ dcn_on_last_conv=False,
+ center_sampling=True,
+ conv_bias=True,
+ loss_bbox=dict(type='GIoULoss', loss_weight=1.0)),
+ # training and testing settings
+ test_cfg=dict(nms=dict(type='nms', iou_threshold=0.6)))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 3.0,
+ by_epoch=False,
+ begin=0,
+ end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(clip_grad=None)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50-caffe_fpn_gn-head-center_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50-caffe_fpn_gn-head-center_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..9e4eb1d5981761fab8fe0bb876ff7ef243ac31f9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50-caffe_fpn_gn-head-center_1x_coco.py
@@ -0,0 +1,4 @@
+_base_ = './fcos_r50-caffe_fpn_gn-head_1x_coco.py'
+
+# model settings
+model = dict(bbox_head=dict(center_sampling=True, center_sample_radius=1.5))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50-caffe_fpn_gn-head_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50-caffe_fpn_gn-head_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..928a9b4c92d217822179c0ae00ae50f6f74289b1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50-caffe_fpn_gn-head_1x_coco.py
@@ -0,0 +1,75 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+# model settings
+model = dict(
+ type='FCOS',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[102.9801, 115.9465, 122.7717],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron/resnet50_caffe')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output', # use P5
+ num_outs=5,
+ relu_before_extra_convs=True),
+ bbox_head=dict(
+ type='FCOSHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ strides=[8, 16, 32, 64, 128],
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='IoULoss', loss_weight=1.0),
+ loss_centerness=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0)),
+ # testing settings
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.5),
+ max_per_img=100))
+
+# learning rate
+param_scheduler = [
+ dict(type='ConstantLR', factor=1.0 / 3, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(lr=0.01),
+ paramwise_cfg=dict(bias_lr_mult=2., bias_decay_mult=0.),
+ clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50-caffe_fpn_gn-head_4xb4-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50-caffe_fpn_gn-head_4xb4-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..32358cd3c69800874aa77ba5746ffc0d6f3a219d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50-caffe_fpn_gn-head_4xb4-1x_coco.py
@@ -0,0 +1,5 @@
+# TODO: Remove this config after benchmarking all related configs
+_base_ = 'fcos_r50-caffe_fpn_gn-head_1x_coco.py'
+
+# dataset settings
+train_dataloader = dict(batch_size=4, num_workers=4)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50-caffe_fpn_gn-head_ms-640-800-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50-caffe_fpn_gn-head_ms-640-800-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..4d50b4ec6c4a10b07cbf73475e7af545b058605c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50-caffe_fpn_gn-head_ms-640-800-2x_coco.py
@@ -0,0 +1,30 @@
+_base_ = './fcos_r50-caffe_fpn_gn-head_1x_coco.py'
+
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+# training schedule for 2x
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(type='ConstantLR', factor=1.0 / 3, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50-dcn-caffe_fpn_gn-head-center-normbbox-centeronreg-giou_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50-dcn-caffe_fpn_gn-head-center-normbbox-centeronreg-giou_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a6a6c44f9b4213601b447bc02720e24dc86a53d9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50-dcn-caffe_fpn_gn-head-center-normbbox-centeronreg-giou_1x_coco.py
@@ -0,0 +1,45 @@
+_base_ = 'fcos_r50-caffe_fpn_gn-head_1x_coco.py'
+
+# model settings
+model = dict(
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ dcn=dict(type='DCNv2', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True),
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')),
+ bbox_head=dict(
+ norm_on_bbox=True,
+ centerness_on_reg=True,
+ dcn_on_last_conv=True,
+ center_sampling=True,
+ conv_bias=True,
+ loss_bbox=dict(type='GIoULoss', loss_weight=1.0)),
+ # training and testing settings
+ test_cfg=dict(nms=dict(type='nms', iou_threshold=0.6)))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 3.0,
+ by_epoch=False,
+ begin=0,
+ end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(clip_grad=None)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50_fpn_gn-head-center-normbbox-centeronreg-giou_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50_fpn_gn-head-center-normbbox-centeronreg-giou_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b51556b8eb7f844866d7acff5c7b86c08cb2a054
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_r50_fpn_gn-head-center-normbbox-centeronreg-giou_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,75 @@
+_base_ = '../common/lsj-200e_coco-detection.py'
+
+image_size = (1024, 1024)
+batch_augments = [dict(type='BatchFixedSizePad', size=image_size)]
+
+# model settings
+model = dict(
+ type='FCOS',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32,
+ batch_augments=batch_augments),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output', # use P5
+ num_outs=5,
+ relu_before_extra_convs=True),
+ bbox_head=dict(
+ type='FCOSHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ strides=[8, 16, 32, 64, 128],
+ norm_on_bbox=True,
+ centerness_on_reg=True,
+ dcn_on_last_conv=False,
+ center_sampling=True,
+ conv_bias=True,
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=1.0),
+ loss_centerness=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0)),
+ # testing settings
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+
+train_dataloader = dict(batch_size=8, num_workers=4)
+# Enable automatic-mixed-precision training with AmpOptimWrapper.
+optim_wrapper = dict(
+ type='AmpOptimWrapper',
+ optimizer=dict(
+ type='SGD', lr=0.01 * 4, momentum=0.9, weight_decay=0.00004),
+ paramwise_cfg=dict(bias_lr_mult=2., bias_decay_mult=0.),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_x101-64x4d_fpn_gn-head_ms-640-800-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_x101-64x4d_fpn_gn-head_ms-640-800-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..503c0e1ce79bdbc9f2a32cc65f977b0f1e968927
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/fcos_x101-64x4d_fpn_gn-head_ms-640-800-2x_coco.py
@@ -0,0 +1,52 @@
+_base_ = './fcos_r50-caffe_fpn_gn-head_1x_coco.py'
+
+# model settings
+model = dict(
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
+
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+# training schedule for 2x
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(type='ConstantLR', factor=1.0 / 3, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..fb6527cf2d418762ae1a4a9298ade3da54ece5df
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fcos/metafile.yml
@@ -0,0 +1,146 @@
+Collections:
+ - Name: FCOS
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - FPN
+ - Group Normalization
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/1904.01355
+ Title: 'FCOS: Fully Convolutional One-Stage Object Detection'
+ README: configs/fcos/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/detectors/fcos.py#L6
+ Version: v2.0.0
+
+Models:
+ - Name: fcos_r50-caffe_fpn_gn-head_1x_coco
+ In Collection: FCOS
+ Config: configs/fcos/fcos_r50-caffe_fpn_gn-head_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.6
+ inference time (ms/im):
+ - value: 44.05
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 36.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_r50_caffe_fpn_gn-head_1x_coco/fcos_r50_caffe_fpn_gn-head_1x_coco-821213aa.pth
+
+ - Name: fcos_r50-caffe_fpn_gn-head-center-normbbox-centeronreg-giou_1x_coco
+ In Collection: FCOS
+ Config: configs/fcos/fcos_r50-caffe_fpn_gn-head-center-normbbox-centeronreg-giou_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.7
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_center-normbbox-centeronreg-giou_r50_caffe_fpn_gn-head_1x_coco/fcos_center-normbbox-centeronreg-giou_r50_caffe_fpn_gn-head_1x_coco-0a0d75a8.pth
+
+ - Name: fcos_r50-dcn-caffe_fpn_gn-head-center-normbbox-centeronreg-giou_1x_coco
+ In Collection: FCOS
+ Config: configs/fcos/fcos_r50-dcn-caffe_fpn_gn-head-center-normbbox-centeronreg-giou_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.8
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_center-normbbox-centeronreg-giou_r50_caffe_fpn_gn-head_dcn_1x_coco/fcos_center-normbbox-centeronreg-giou_r50_caffe_fpn_gn-head_dcn_1x_coco-ae4d8b3d.pth
+
+ - Name: fcos_r101-caffe_fpn_gn-head-1x_coco
+ In Collection: FCOS
+ Config: configs/fcos/fcos_r101-caffe_fpn_gn-head-1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.5
+ inference time (ms/im):
+ - value: 57.8
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_r101_caffe_fpn_gn-head_1x_coco/fcos_r101_caffe_fpn_gn-head_1x_coco-0e37b982.pth
+
+ - Name: fcos_r50-caffe_fpn_gn-head_ms-640-800-2x_coco
+ In Collection: FCOS
+ Config: configs/fcos/fcos_r50-caffe_fpn_gn-head_ms-640-800-2x_coco.py
+ Metadata:
+ Training Memory (GB): 2.6
+ inference time (ms/im):
+ - value: 43.67
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_r50_caffe_fpn_gn-head_mstrain_640-800_2x_coco/fcos_r50_caffe_fpn_gn-head_mstrain_640-800_2x_coco-d92ceeea.pth
+
+ - Name: fcos_r101-caffe_fpn_gn-head_ms-640-800-2x_coco
+ In Collection: FCOS
+ Config: configs/fcos/fcos_r101-caffe_fpn_gn-head_ms-640-800-2x_coco.py
+ Metadata:
+ Training Memory (GB): 5.5
+ inference time (ms/im):
+ - value: 57.8
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_r101_caffe_fpn_gn-head_mstrain_640-800_2x_coco/fcos_r101_caffe_fpn_gn-head_mstrain_640-800_2x_coco-511424d6.pth
+
+ - Name: fcos_x101-64x4d_fpn_gn-head_ms-640-800-2x_coco
+ In Collection: FCOS
+ Config: configs/fcos/fcos_x101-64x4d_fpn_gn-head_ms-640-800-2x_coco.py
+ Metadata:
+ Training Memory (GB): 10.0
+ inference time (ms/im):
+ - value: 103.09
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fcos/fcos_x101_64x4d_fpn_gn-head_mstrain_640-800_2x_coco/fcos_x101_64x4d_fpn_gn-head_mstrain_640-800_2x_coco-ede514a8.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..96f1358b11840e5e03d1a640969a8d18d6197588
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/README.md
@@ -0,0 +1,53 @@
+# FoveaBox
+
+> [FoveaBox: Beyond Anchor-based Object Detector](https://arxiv.org/abs/1904.03797)
+
+
+
+## Abstract
+
+We present FoveaBox, an accurate, flexible, and completely anchor-free framework for object detection. While almost all state-of-the-art object detectors utilize predefined anchors to enumerate possible locations, scales and aspect ratios for the search of the objects, their performance and generalization ability are also limited to the design of anchors. Instead, FoveaBox directly learns the object existing possibility and the bounding box coordinates without anchor reference. This is achieved by: (a) predicting category-sensitive semantic maps for the object existing possibility, and (b) producing category-agnostic bounding box for each position that potentially contains an object. The scales of target boxes are naturally associated with feature pyramid representations. In FoveaBox, an instance is assigned to adjacent feature levels to make the model more accurate.We demonstrate its effectiveness on standard benchmarks and report extensive experimental analysis. Without bells and whistles, FoveaBox achieves state-of-the-art single model performance on the standard COCO and Pascal VOC object detection benchmark. More importantly, FoveaBox avoids all computation and hyper-parameters related to anchor boxes, which are often sensitive to the final detection performance. We believe the simple and effective approach will serve as a solid baseline and help ease future research for object detection.
+
+
+

+
+
+## Introduction
+
+FoveaBox is an accurate, flexible and completely anchor-free object detection system for object detection framework, as presented in our paper [https://arxiv.org/abs/1904.03797](https://arxiv.org/abs/1904.03797):
+Different from previous anchor-based methods, FoveaBox directly learns the object existing possibility and the bounding box coordinates without anchor reference. This is achieved by: (a) predicting category-sensitive semantic maps for the object existing possibility, and (b) producing category-agnostic bounding box for each position that potentially contains an object.
+
+## Results and Models
+
+### Results on R50/101-FPN
+
+| Backbone | Style | align | ms-train | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :------: | :-----: | :---: | :------: | :-----: | :------: | :------------: | :----: | :-----------------------------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | pytorch | N | N | 1x | 5.6 | 24.1 | 36.5 | [config](./fovea_r50_fpn_4xb4-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_r50_fpn_4x4_1x_coco/fovea_r50_fpn_4x4_1x_coco_20200219-ee4d5303.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_r50_fpn_4x4_1x_coco/fovea_r50_fpn_4x4_1x_coco_20200219_223025.log.json) |
+| R-50 | pytorch | N | N | 2x | 5.6 | - | 37.2 | [config](./fovea_r50_fpn_4xb4-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_r50_fpn_4x4_2x_coco/fovea_r50_fpn_4x4_2x_coco_20200203-2df792b1.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_r50_fpn_4x4_2x_coco/fovea_r50_fpn_4x4_2x_coco_20200203_112043.log.json) |
+| R-50 | pytorch | Y | N | 2x | 8.1 | 19.4 | 37.9 | [config](./fovea_r50_fpn_gn-head-align_4xb4-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_align_r50_fpn_gn-head_4x4_2x_coco/fovea_align_r50_fpn_gn-head_4x4_2x_coco_20200203-8987880d.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_align_r50_fpn_gn-head_4x4_2x_coco/fovea_align_r50_fpn_gn-head_4x4_2x_coco_20200203_134252.log.json) |
+| R-50 | pytorch | Y | Y | 2x | 8.1 | 18.3 | 40.4 | [config](./fovea_r50_fpn_gn-head-align_ms-640-800-4xb4-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_align_r50_fpn_gn-head_mstrain_640-800_4x4_2x_coco/fovea_align_r50_fpn_gn-head_mstrain_640-800_4x4_2x_coco_20200205-85ce26cb.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_align_r50_fpn_gn-head_mstrain_640-800_4x4_2x_coco/fovea_align_r50_fpn_gn-head_mstrain_640-800_4x4_2x_coco_20200205_112557.log.json) |
+| R-101 | pytorch | N | N | 1x | 9.2 | 17.4 | 38.6 | [config](./fovea_r101_fpn_4xb4-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_r101_fpn_4x4_1x_coco/fovea_r101_fpn_4x4_1x_coco_20200219-05e38f1c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_r101_fpn_4x4_1x_coco/fovea_r101_fpn_4x4_1x_coco_20200219_011740.log.json) |
+| R-101 | pytorch | N | N | 2x | 11.7 | - | 40.0 | [config](./fovea_r101_fpn_4xb4-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_r101_fpn_4x4_2x_coco/fovea_r101_fpn_4x4_2x_coco_20200208-02320ea4.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_r101_fpn_4x4_2x_coco/fovea_r101_fpn_4x4_2x_coco_20200208_202059.log.json) |
+| R-101 | pytorch | Y | N | 2x | 11.7 | 14.7 | 40.0 | [config](./fovea_r101_fpn_gn-head-align_4xb4-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_align_r101_fpn_gn-head_4x4_2x_coco/fovea_align_r101_fpn_gn-head_4x4_2x_coco_20200208-c39a027a.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_align_r101_fpn_gn-head_4x4_2x_coco/fovea_align_r101_fpn_gn-head_4x4_2x_coco_20200208_203337.log.json) |
+| R-101 | pytorch | Y | Y | 2x | 11.7 | 14.7 | 42.0 | [config](./fovea_r101_fpn_gn-head-align_ms-640-800-4xb4-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_align_r101_fpn_gn-head_mstrain_640-800_4x4_2x_coco/fovea_align_r101_fpn_gn-head_mstrain_640-800_4x4_2x_coco_20200208-649c5eb6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_align_r101_fpn_gn-head_mstrain_640-800_4x4_2x_coco/fovea_align_r101_fpn_gn-head_mstrain_640-800_4x4_2x_coco_20200208_202124.log.json) |
+
+\[1\] *1x and 2x mean the model is trained for 12 and 24 epochs, respectively.* \
+\[2\] *Align means utilizing deformable convolution to align the cls branch.* \
+\[3\] *All results are obtained with a single model and without any test time data augmentation.*\
+\[4\] *We use 4 GPUs for training.*
+
+Any pull requests or issues are welcome.
+
+## Citation
+
+Please consider citing our paper in your publications if the project helps your research. BibTeX reference is as follows.
+
+```latex
+@article{kong2019foveabox,
+ title={FoveaBox: Beyond Anchor-based Object Detector},
+ author={Kong, Tao and Sun, Fuchun and Liu, Huaping and Jiang, Yuning and Shi, Jianbo},
+ journal={arXiv preprint arXiv:1904.03797},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r101_fpn_4xb4-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r101_fpn_4xb4-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..7e8ccf910e6317bf576463fa26bfcb330b6ff385
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r101_fpn_4xb4-1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './fovea_r50_fpn_4xb4-1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r101_fpn_4xb4-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r101_fpn_4xb4-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..0dc98515e62b2dba225e822850229f0a2f802d63
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r101_fpn_4xb4-2x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './fovea_r50_fpn_4xb4-2x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r101_fpn_gn-head-align_4xb4-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r101_fpn_gn-head-align_4xb4-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..222671d49d1e3fbc31285e4f13487d86642ebbe3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r101_fpn_gn-head-align_4xb4-2x_coco.py
@@ -0,0 +1,23 @@
+_base_ = './fovea_r50_fpn_4xb4-1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')),
+ bbox_head=dict(
+ with_deform=True,
+ norm_cfg=dict(type='GN', num_groups=32, requires_grad=True)))
+# learning policy
+max_epochs = 24
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r101_fpn_gn-head-align_ms-640-800-4xb4-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r101_fpn_gn-head-align_ms-640-800-4xb4-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e1852d581fcbdd9a1459291fc7f65e51041aa4e6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r101_fpn_gn-head-align_ms-640-800-4xb4-2x_coco.py
@@ -0,0 +1,34 @@
+_base_ = './fovea_r50_fpn_4xb4-1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')),
+ bbox_head=dict(
+ with_deform=True,
+ norm_cfg=dict(type='GN', num_groups=32, requires_grad=True)))
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+# learning policy
+max_epochs = 24
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r50_fpn_4xb4-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r50_fpn_4xb4-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..13cf3ae92b0d2bfd1d84f032f7b202430f095a6a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r50_fpn_4xb4-1x_coco.py
@@ -0,0 +1,59 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+# model settings
+model = dict(
+ type='FOVEA',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ num_outs=5,
+ add_extra_convs='on_input'),
+ bbox_head=dict(
+ type='FoveaHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ strides=[8, 16, 32, 64, 128],
+ base_edge_list=[16, 32, 64, 128, 256],
+ scale_ranges=((1, 64), (32, 128), (64, 256), (128, 512), (256, 2048)),
+ sigma=0.4,
+ with_deform=False,
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=1.50,
+ alpha=0.4,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=0.11, loss_weight=1.0)),
+ # training and testing settings
+ train_cfg=dict(),
+ test_cfg=dict(
+ nms_pre=1000,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.5),
+ max_per_img=100))
+train_dataloader = dict(batch_size=4, num_workers=4)
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r50_fpn_4xb4-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r50_fpn_4xb4-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f9d06ef9f9ba89f202ef13176af39df7e89cb5e6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r50_fpn_4xb4-2x_coco.py
@@ -0,0 +1,15 @@
+_base_ = './fovea_r50_fpn_4xb4-1x_coco.py'
+# learning policy
+max_epochs = 24
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r50_fpn_gn-head-align_4xb4-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r50_fpn_gn-head-align_4xb4-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..877bb4fa4e1c03190a05da4e95558d8534e5e6e8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r50_fpn_gn-head-align_4xb4-2x_coco.py
@@ -0,0 +1,20 @@
+_base_ = './fovea_r50_fpn_4xb4-1x_coco.py'
+model = dict(
+ bbox_head=dict(
+ with_deform=True,
+ norm_cfg=dict(type='GN', num_groups=32, requires_grad=True)))
+# learning policy
+max_epochs = 24
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs)
+optim_wrapper = dict(clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r50_fpn_gn-head-align_ms-640-800-4xb4-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r50_fpn_gn-head-align_ms-640-800-4xb4-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5690bcae08cd0e639afe3c832a46f78036324c08
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/fovea_r50_fpn_gn-head-align_ms-640-800-4xb4-2x_coco.py
@@ -0,0 +1,30 @@
+_base_ = './fovea_r50_fpn_4xb4-1x_coco.py'
+model = dict(
+ bbox_head=dict(
+ with_deform=True,
+ norm_cfg=dict(type='GN', num_groups=32, requires_grad=True)))
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+# learning policy
+max_epochs = 24
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..9ab2f5420323a9eb8c2ace386485c34277d53213
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/foveabox/metafile.yml
@@ -0,0 +1,172 @@
+Collections:
+ - Name: FoveaBox
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 4x V100 GPUs
+ Architecture:
+ - FPN
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/1904.03797
+ Title: 'FoveaBox: Beyond Anchor-based Object Detector'
+ README: configs/foveabox/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/detectors/fovea.py#L6
+ Version: v2.0.0
+
+Models:
+ - Name: fovea_r50_fpn_4xb4-1x_coco
+ In Collection: FoveaBox
+ Config: configs/foveabox/fovea_r50_fpn_4xb4-1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.6
+ inference time (ms/im):
+ - value: 41.49
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 36.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_r50_fpn_4x4_1x_coco/fovea_r50_fpn_4x4_1x_coco_20200219-ee4d5303.pth
+
+ - Name: fovea_r50_fpn_4xb4-2x_coco
+ In Collection: FoveaBox
+ Config: configs/foveabox/fovea_r50_fpn_4xb4-2x_coco.py
+ Metadata:
+ Training Memory (GB): 5.6
+ inference time (ms/im):
+ - value: 41.49
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_r50_fpn_4x4_2x_coco/fovea_r50_fpn_4x4_2x_coco_20200203-2df792b1.pth
+
+ - Name: fovea_r50_fpn_gn-head-align_4xb4-2x_coco
+ In Collection: FoveaBox
+ Config: configs/foveabox/fovea_r50_fpn_gn-head-align_4xb4-2x_coco.py
+ Metadata:
+ Training Memory (GB): 8.1
+ inference time (ms/im):
+ - value: 51.55
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_align_r50_fpn_gn-head_4x4_2x_coco/fovea_align_r50_fpn_gn-head_4x4_2x_coco_20200203-8987880d.pth
+
+ - Name: fovea_r50_fpn_gn-head-align_ms-640-800-4xb4-2x_coco
+ In Collection: FoveaBox
+ Config: configs/foveabox/fovea_r50_fpn_gn-head-align_ms-640-800-4xb4-2x_coco.py
+ Metadata:
+ Training Memory (GB): 8.1
+ inference time (ms/im):
+ - value: 54.64
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_align_r50_fpn_gn-head_mstrain_640-800_4x4_2x_coco/fovea_align_r50_fpn_gn-head_mstrain_640-800_4x4_2x_coco_20200205-85ce26cb.pth
+
+ - Name: fovea_r101_fpn_4xb4-1x_coco
+ In Collection: FoveaBox
+ Config: configs/foveabox/fovea_r101_fpn_4xb4-1x_coco.py
+ Metadata:
+ Training Memory (GB): 9.2
+ inference time (ms/im):
+ - value: 57.47
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_r101_fpn_4x4_1x_coco/fovea_r101_fpn_4x4_1x_coco_20200219-05e38f1c.pth
+
+ - Name: fovea_r101_fpn_4xb4-2x_coco
+ In Collection: FoveaBox
+ Config: configs/foveabox/fovea_r101_fpn_4xb4-2x_coco.py
+ Metadata:
+ Training Memory (GB): 11.7
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_r101_fpn_4x4_2x_coco/fovea_r101_fpn_4x4_2x_coco_20200208-02320ea4.pth
+
+ - Name: fovea_r101_fpn_gn-head-align_4xb4-2x_coco
+ In Collection: FoveaBox
+ Config: configs/foveabox/fovea_r101_fpn_gn-head-align_4xb4-2x_coco.py
+ Metadata:
+ Training Memory (GB): 11.7
+ inference time (ms/im):
+ - value: 68.03
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_align_r101_fpn_gn-head_4x4_2x_coco/fovea_align_r101_fpn_gn-head_4x4_2x_coco_20200208-c39a027a.pth
+
+ - Name: fovea_r101_fpn_gn-head-align_ms-640-800-4xb4-2x_coco
+ In Collection: FoveaBox
+ Config: configs/foveabox/fovea_r101_fpn_gn-head-align_ms-640-800-4xb4-2x_coco.py
+ Metadata:
+ Training Memory (GB): 11.7
+ inference time (ms/im):
+ - value: 68.03
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/foveabox/fovea_align_r101_fpn_gn-head_mstrain_640-800_4x4_2x_coco/fovea_align_r101_fpn_gn-head_mstrain_640-800_4x4_2x_coco_20200208-649c5eb6.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..1e2fd400288d3ebd57741f1b1d18e430a8c62f41
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/README.md
@@ -0,0 +1,43 @@
+# FPG
+
+> [Feature Pyramid Grids](https://arxiv.org/abs/2004.03580)
+
+
+
+## Abstract
+
+Feature pyramid networks have been widely adopted in the object detection literature to improve feature representations for better handling of variations in scale. In this paper, we present Feature Pyramid Grids (FPG), a deep multi-pathway feature pyramid, that represents the feature scale-space as a regular grid of parallel bottom-up pathways which are fused by multi-directional lateral connections. FPG can improve single-pathway feature pyramid networks by significantly increasing its performance at similar computation cost, highlighting importance of deep pyramid representations. In addition to its general and uniform structure, over complicated structures that have been found with neural architecture search, it also compares favorably against such approaches without relying on search. We hope that FPG with its uniform and effective nature can serve as a strong component for future work in object recognition.
+
+
+

+
+
+## Results and Models
+
+We benchmark the new training schedule (crop training, large batch, unfrozen BN, 50 epochs) introduced in NAS-FPN.
+All backbones are Resnet-50 in pytorch style.
+
+| Method | Neck | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :----------: | :--------: | :-----: | :------: | :------------: | :----: | :-----: | :--------------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Faster R-CNN | FPG | 50e | 20.0 | - | 42.3 | - | [config](./faster-rcnn_r50_fpg_crop640-50e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fpg/faster_rcnn_r50_fpg_crop640_50e_coco/faster_rcnn_r50_fpg_crop640_50e_coco_20220311_011856-74109f42.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fpg/faster_rcnn_r50_fpg_crop640_50e_coco/faster_rcnn_r50_fpg_crop640_50e_coco_20220311_011856.log.json) |
+| Faster R-CNN | FPG-chn128 | 50e | 11.9 | - | 41.2 | - | [config](./faster-rcnn_r50_fpg-chn128_crop640-50e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fpg/faster_rcnn_r50_fpg-chn128_crop640_50e_coco/faster_rcnn_r50_fpg-chn128_crop640_50e_coco_20220311_011857-9376aa9d.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fpg/faster_rcnn_r50_fpg-chn128_crop640_50e_coco/faster_rcnn_r50_fpg-chn128_crop640_50e_coco_20220311_011857.log.json) |
+| Faster R-CNN | FPN | 50e | 20.0 | - | 38.9 | - | [config](./faster-rcnn_r50_fpn_crop640-50e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fpg/faster_rcnn_r50_fpn_crop640_50e_coco/faster_rcnn_r50_fpn_crop640_50e_coco_20220311_011857-be7c9f42.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fpg/faster_rcnn_r50_fpn_crop640_50e_coco/faster_rcnn_r50_fpn_crop640_50e_coco_20220311_011857.log.json) |
+| Mask R-CNN | FPG | 50e | 23.2 | - | 43.0 | 38.1 | [config](./mask-rcnn_r50_fpg_crop640-50e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fpg/mask_rcnn_r50_fpg_crop640_50e_coco/mask_rcnn_r50_fpg_crop640_50e_coco_20220311_011857-233b8334.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fpg/mask_rcnn_r50_fpg_crop640_50e_coco/mask_rcnn_r50_fpg_crop640_50e_coco_20220311_011857.log.json) |
+| Mask R-CNN | FPG-chn128 | 50e | 15.3 | - | 41.7 | 37.1 | [config](./mask-rcnn_r50_fpg-chn128_crop640-50e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fpg/mask_rcnn_r50_fpg-chn128_crop640_50e_coco/mask_rcnn_r50_fpg-chn128_crop640_50e_coco_20220311_011859-043c9b4e.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fpg/mask_rcnn_r50_fpg-chn128_crop640_50e_coco/mask_rcnn_r50_fpg-chn128_crop640_50e_coco_20220311_011859.log.json) |
+| Mask R-CNN | FPN | 50e | 23.2 | - | 49.6 | 35.6 | [config](./mask-rcnn_r50_fpn_crop640-50e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fpg/mask_rcnn_r50_fpn_crop640_50e_coco/mask_rcnn_r50_fpn_crop640_50e_coco_20220311_011855-a756664a.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fpg/mask_rcnn_r50_fpn_crop640_50e_coco/mask_rcnn_r50_fpn_crop640_50e_coco_20220311_011855.log.json) |
+| RetinaNet | FPG | 50e | 20.8 | - | 40.5 | - | [config](./retinanet_r50_fpg_crop640_50e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fpg/retinanet_r50_fpg_crop640_50e_coco/retinanet_r50_fpg_crop640_50e_coco_20220311_110809-b0bcf5f4.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fpg/retinanet_r50_fpg_crop640_50e_coco/retinanet_r50_fpg_crop640_50e_coco_20220311_110809.log.json) |
+| RetinaNet | FPG-chn128 | 50e | 19.9 | - | 39.9 | - | [config](./retinanet_r50_fpg-chn128_crop640_50e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fpg/retinanet_r50_fpg-chn128_crop640_50e_coco/retinanet_r50_fpg-chn128_crop640_50e_coco_20220313_104829-ee99a686.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fpg/retinanet_r50_fpg-chn128_crop640_50e_coco/retinanet_r50_fpg-chn128_crop640_50e_coco_20220313_104829.log.json) |
+
+**Note**: Chn128 means to decrease the number of channels of features and convs from 256 (default) to 128 in
+Neck and BBox Head, which can greatly decrease memory consumption without sacrificing much precision.
+
+## Citation
+
+```latex
+@article{chen2020feature,
+ title={Feature pyramid grids},
+ author={Chen, Kai and Cao, Yuhang and Loy, Chen Change and Lin, Dahua and Feichtenhofer, Christoph},
+ journal={arXiv preprint arXiv:2004.03580},
+ year={2020}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/faster-rcnn_r50_fpg-chn128_crop640-50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/faster-rcnn_r50_fpg-chn128_crop640-50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..cb9160f5cc7e118069d7172573018515aa406331
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/faster-rcnn_r50_fpg-chn128_crop640-50e_coco.py
@@ -0,0 +1,9 @@
+_base_ = 'faster-rcnn_r50_fpg_crop640-50e_coco.py'
+
+norm_cfg = dict(type='BN', requires_grad=True)
+model = dict(
+ neck=dict(out_channels=128, inter_channels=128),
+ rpn_head=dict(in_channels=128),
+ roi_head=dict(
+ bbox_roi_extractor=dict(out_channels=128),
+ bbox_head=dict(in_channels=128)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/faster-rcnn_r50_fpg_crop640-50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/faster-rcnn_r50_fpg_crop640-50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d0d366f1f30e5bcc6d52010c46d60183b56386ea
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/faster-rcnn_r50_fpg_crop640-50e_coco.py
@@ -0,0 +1,48 @@
+_base_ = 'faster-rcnn_r50_fpn_crop640-50e_coco.py'
+
+norm_cfg = dict(type='BN', requires_grad=True)
+model = dict(
+ neck=dict(
+ type='FPG',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ inter_channels=256,
+ num_outs=5,
+ stack_times=9,
+ paths=['bu'] * 9,
+ same_down_trans=None,
+ same_up_trans=dict(
+ type='conv',
+ kernel_size=3,
+ stride=2,
+ padding=1,
+ norm_cfg=norm_cfg,
+ inplace=False,
+ order=('act', 'conv', 'norm')),
+ across_lateral_trans=dict(
+ type='conv',
+ kernel_size=1,
+ norm_cfg=norm_cfg,
+ inplace=False,
+ order=('act', 'conv', 'norm')),
+ across_down_trans=dict(
+ type='interpolation_conv',
+ mode='nearest',
+ kernel_size=3,
+ norm_cfg=norm_cfg,
+ order=('act', 'conv', 'norm'),
+ inplace=False),
+ across_up_trans=None,
+ across_skip_trans=dict(
+ type='conv',
+ kernel_size=1,
+ norm_cfg=norm_cfg,
+ inplace=False,
+ order=('act', 'conv', 'norm')),
+ output_trans=dict(
+ type='last_conv',
+ kernel_size=3,
+ order=('act', 'conv', 'norm'),
+ inplace=False),
+ norm_cfg=norm_cfg,
+ skip_inds=[(0, 1, 2, 3), (0, 1, 2), (0, 1), (0, ), ()]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/faster-rcnn_r50_fpn_crop640-50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/faster-rcnn_r50_fpn_crop640-50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..46211de03f34e6a9709a9cfa8561b88a90f69581
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/faster-rcnn_r50_fpn_crop640-50e_coco.py
@@ -0,0 +1,73 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+norm_cfg = dict(type='BN', requires_grad=True)
+image_size = (640, 640)
+batch_augments = [dict(type='BatchFixedSizePad', size=image_size)]
+
+model = dict(
+ data_preprocessor=dict(pad_size_divisor=64, batch_augments=batch_augments),
+ backbone=dict(norm_cfg=norm_cfg, norm_eval=False),
+ neck=dict(norm_cfg=norm_cfg),
+ roi_head=dict(bbox_head=dict(norm_cfg=norm_cfg)))
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize',
+ scale=image_size,
+ ratio_range=(0.8, 1.2),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size,
+ allow_negative_crop=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=image_size, keep_ratio=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=8, num_workers=4, dataset=dict(pipeline=train_pipeline))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# learning policy
+max_epochs = 50
+train_cfg = dict(max_epochs=max_epochs, val_interval=2)
+param_scheduler = [
+ dict(type='LinearLR', start_factor=0.1, by_epoch=False, begin=0, end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[30, 40],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.08, momentum=0.9, weight_decay=0.0001),
+ paramwise_cfg=dict(norm_decay_mult=0, bypass_duplicate=True),
+ clip_grad=None)
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/mask-rcnn_r50_fpg-chn128_crop640-50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/mask-rcnn_r50_fpg-chn128_crop640-50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..804393966c6711a1e5261ace00e9b8b84283fde5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/mask-rcnn_r50_fpg-chn128_crop640-50e_coco.py
@@ -0,0 +1,10 @@
+_base_ = 'mask-rcnn_r50_fpg_crop640-50e_coco.py'
+
+model = dict(
+ neck=dict(out_channels=128, inter_channels=128),
+ rpn_head=dict(in_channels=128),
+ roi_head=dict(
+ bbox_roi_extractor=dict(out_channels=128),
+ bbox_head=dict(in_channels=128),
+ mask_roi_extractor=dict(out_channels=128),
+ mask_head=dict(in_channels=128)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/mask-rcnn_r50_fpg_crop640-50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/mask-rcnn_r50_fpg_crop640-50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..135bb60bb340c40a47a9bd64e5a8afc57ede60db
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/mask-rcnn_r50_fpg_crop640-50e_coco.py
@@ -0,0 +1,48 @@
+_base_ = 'mask-rcnn_r50_fpn_crop640-50e_coco.py'
+
+norm_cfg = dict(type='BN', requires_grad=True)
+model = dict(
+ neck=dict(
+ type='FPG',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ inter_channels=256,
+ num_outs=5,
+ stack_times=9,
+ paths=['bu'] * 9,
+ same_down_trans=None,
+ same_up_trans=dict(
+ type='conv',
+ kernel_size=3,
+ stride=2,
+ padding=1,
+ norm_cfg=norm_cfg,
+ inplace=False,
+ order=('act', 'conv', 'norm')),
+ across_lateral_trans=dict(
+ type='conv',
+ kernel_size=1,
+ norm_cfg=norm_cfg,
+ inplace=False,
+ order=('act', 'conv', 'norm')),
+ across_down_trans=dict(
+ type='interpolation_conv',
+ mode='nearest',
+ kernel_size=3,
+ norm_cfg=norm_cfg,
+ order=('act', 'conv', 'norm'),
+ inplace=False),
+ across_up_trans=None,
+ across_skip_trans=dict(
+ type='conv',
+ kernel_size=1,
+ norm_cfg=norm_cfg,
+ inplace=False,
+ order=('act', 'conv', 'norm')),
+ output_trans=dict(
+ type='last_conv',
+ kernel_size=3,
+ order=('act', 'conv', 'norm'),
+ inplace=False),
+ norm_cfg=norm_cfg,
+ skip_inds=[(0, 1, 2, 3), (0, 1, 2), (0, 1), (0, ), ()]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/mask-rcnn_r50_fpn_crop640-50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/mask-rcnn_r50_fpn_crop640-50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..08ca5b6ffd8b9d166857d3c27bb6f5bde91416cc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/mask-rcnn_r50_fpn_crop640-50e_coco.py
@@ -0,0 +1,79 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+norm_cfg = dict(type='BN', requires_grad=True)
+image_size = (640, 640)
+batch_augments = [dict(type='BatchFixedSizePad', size=image_size)]
+
+model = dict(
+ data_preprocessor=dict(pad_size_divisor=64, batch_augments=batch_augments),
+ backbone=dict(norm_cfg=norm_cfg, norm_eval=False),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ norm_cfg=norm_cfg,
+ num_outs=5),
+ roi_head=dict(
+ bbox_head=dict(norm_cfg=norm_cfg), mask_head=dict(norm_cfg=norm_cfg)))
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomResize',
+ scale=image_size,
+ ratio_range=(0.8, 1.2),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size,
+ allow_negative_crop=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=image_size, keep_ratio=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=8, num_workers=4, dataset=dict(pipeline=train_pipeline))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# learning policy
+max_epochs = 50
+train_cfg = dict(max_epochs=max_epochs, val_interval=2)
+param_scheduler = [
+ dict(type='LinearLR', start_factor=0.1, by_epoch=False, begin=0, end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[30, 40],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.08, momentum=0.9, weight_decay=0.0001),
+ paramwise_cfg=dict(norm_decay_mult=0, bypass_duplicate=True),
+ clip_grad=None)
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..7d7634aec6161a283577059de96d5f995cf1e4bb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/metafile.yml
@@ -0,0 +1,104 @@
+Collections:
+ - Name: Feature Pyramid Grids
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Feature Pyramid Grids
+ Paper:
+ URL: https://arxiv.org/abs/2004.03580
+ Title: 'Feature Pyramid Grids'
+ README: configs/fpg/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.10.0/mmdet/models/necks/fpg.py#L101
+ Version: v2.10.0
+
+Models:
+ - Name: faster-rcnn_r50_fpg_crop640-50e_coco
+ In Collection: Feature Pyramid Grids
+ Config: configs/fpg/faster-rcnn_r50_fpg_crop640-50e_coco.py
+ Metadata:
+ Training Memory (GB): 20.0
+ Epochs: 50
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fpg/faster_rcnn_r50_fpg_crop640_50e_coco/faster_rcnn_r50_fpg_crop640_50e_coco_20220311_011856-74109f42.pth
+
+ - Name: faster-rcnn_r50_fpg-chn128_crop640-50e_coco
+ In Collection: Feature Pyramid Grids
+ Config: configs/fpg/faster-rcnn_r50_fpg-chn128_crop640-50e_coco.py
+ Metadata:
+ Training Memory (GB): 11.9
+ Epochs: 50
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fpg/faster_rcnn_r50_fpg-chn128_crop640_50e_coco/faster_rcnn_r50_fpg-chn128_crop640_50e_coco_20220311_011857-9376aa9d.pth
+
+ - Name: mask-rcnn_r50_fpg_crop640-50e_coco
+ In Collection: Feature Pyramid Grids
+ Config: configs/fpg/mask-rcnn_r50_fpg_crop640-50e_coco.py
+ Metadata:
+ Training Memory (GB): 23.2
+ Epochs: 50
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fpg/mask_rcnn_r50_fpg_crop640_50e_coco/mask_rcnn_r50_fpg_crop640_50e_coco_20220311_011857-233b8334.pth
+
+ - Name: mask-rcnn_r50_fpg-chn128_crop640-50e_coco
+ In Collection: Feature Pyramid Grids
+ Config: configs/fpg/mask-rcnn_r50_fpg-chn128_crop640-50e_coco.py
+ Metadata:
+ Training Memory (GB): 15.3
+ Epochs: 50
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.7
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fpg/mask_rcnn_r50_fpg-chn128_crop640_50e_coco/mask_rcnn_r50_fpg-chn128_crop640_50e_coco_20220311_011859-043c9b4e.pth
+
+ - Name: retinanet_r50_fpg_crop640_50e_coco
+ In Collection: Feature Pyramid Grids
+ Config: configs/fpg/retinanet_r50_fpg_crop640_50e_coco.py
+ Metadata:
+ Training Memory (GB): 20.8
+ Epochs: 50
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fpg/retinanet_r50_fpg_crop640_50e_coco/retinanet_r50_fpg_crop640_50e_coco_20220311_110809-b0bcf5f4.pth
+
+ - Name: retinanet_r50_fpg-chn128_crop640_50e_coco
+ In Collection: Feature Pyramid Grids
+ Config: configs/fpg/retinanet_r50_fpg-chn128_crop640_50e_coco.py
+ Metadata:
+ Training Memory (GB): 19.9
+ Epochs: 50
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fpg/retinanet_r50_fpg-chn128_crop640_50e_coco/retinanet_r50_fpg-chn128_crop640_50e_coco_20220313_104829-ee99a686.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/retinanet_r50_fpg-chn128_crop640_50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/retinanet_r50_fpg-chn128_crop640_50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..9a6cf7e56a4f23a42d3905560a9b8035d6d935ff
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/retinanet_r50_fpg-chn128_crop640_50e_coco.py
@@ -0,0 +1,5 @@
+_base_ = 'retinanet_r50_fpg_crop640_50e_coco.py'
+
+model = dict(
+ neck=dict(out_channels=128, inter_channels=128),
+ bbox_head=dict(in_channels=128))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/retinanet_r50_fpg_crop640_50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/retinanet_r50_fpg_crop640_50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e2aac283992ea9e4595e7594233b21208bd672f5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fpg/retinanet_r50_fpg_crop640_50e_coco.py
@@ -0,0 +1,53 @@
+_base_ = '../nas_fpn/retinanet_r50_nasfpn_crop640-50e_coco.py'
+
+norm_cfg = dict(type='BN', requires_grad=True)
+model = dict(
+ neck=dict(
+ _delete_=True,
+ type='FPG',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ inter_channels=256,
+ num_outs=5,
+ add_extra_convs=True,
+ start_level=1,
+ stack_times=9,
+ paths=['bu'] * 9,
+ same_down_trans=None,
+ same_up_trans=dict(
+ type='conv',
+ kernel_size=3,
+ stride=2,
+ padding=1,
+ norm_cfg=norm_cfg,
+ inplace=False,
+ order=('act', 'conv', 'norm')),
+ across_lateral_trans=dict(
+ type='conv',
+ kernel_size=1,
+ norm_cfg=norm_cfg,
+ inplace=False,
+ order=('act', 'conv', 'norm')),
+ across_down_trans=dict(
+ type='interpolation_conv',
+ mode='nearest',
+ kernel_size=3,
+ norm_cfg=norm_cfg,
+ order=('act', 'conv', 'norm'),
+ inplace=False),
+ across_up_trans=None,
+ across_skip_trans=dict(
+ type='conv',
+ kernel_size=1,
+ norm_cfg=norm_cfg,
+ inplace=False,
+ order=('act', 'conv', 'norm')),
+ output_trans=dict(
+ type='last_conv',
+ kernel_size=3,
+ order=('act', 'conv', 'norm'),
+ inplace=False),
+ norm_cfg=norm_cfg,
+ skip_inds=[(0, 1, 2, 3), (0, 1, 2), (0, 1), (0, ), ()]))
+
+train_cfg = dict(val_interval=2)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/free_anchor/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/free_anchor/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..03dc828319fcfd5368361af8b64de1018a54f638
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/free_anchor/README.md
@@ -0,0 +1,37 @@
+# FreeAnchor
+
+> [FreeAnchor: Learning to Match Anchors for Visual Object Detection](https://arxiv.org/abs/1909.02466)
+
+
+
+## Abstract
+
+Modern CNN-based object detectors assign anchors for ground-truth objects under the restriction of object-anchor Intersection-over-Unit (IoU). In this study, we propose a learning-to-match approach to break IoU restriction, allowing objects to match anchors in a flexible manner. Our approach, referred to as FreeAnchor, updates hand-crafted anchor assignment to "free" anchor matching by formulating detector training as a maximum likelihood estimation (MLE) procedure. FreeAnchor targets at learning features which best explain a class of objects in terms of both classification and localization. FreeAnchor is implemented by optimizing detection customized likelihood and can be fused with CNN-based detectors in a plug-and-play manner. Experiments on COCO demonstrate that FreeAnchor consistently outperforms their counterparts with significant margins.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :---------: | :-----: | :-----: | :------: | :------------: | :----: | :----------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | pytorch | 1x | 4.9 | 18.4 | 38.7 | [config](./freeanchor_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/free_anchor/retinanet_free_anchor_r50_fpn_1x_coco/retinanet_free_anchor_r50_fpn_1x_coco_20200130-0f67375f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/free_anchor/retinanet_free_anchor_r50_fpn_1x_coco/retinanet_free_anchor_r50_fpn_1x_coco_20200130_095625.log.json) |
+| R-101 | pytorch | 1x | 6.8 | 14.9 | 40.3 | [config](./freeanchor_r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/free_anchor/retinanet_free_anchor_r101_fpn_1x_coco/retinanet_free_anchor_r101_fpn_1x_coco_20200130-358324e6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/free_anchor/retinanet_free_anchor_r101_fpn_1x_coco/retinanet_free_anchor_r101_fpn_1x_coco_20200130_100723.log.json) |
+| X-101-32x4d | pytorch | 1x | 8.1 | 11.1 | 41.9 | [config](./freeanchor_x101-32x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/free_anchor/retinanet_free_anchor_x101_32x4d_fpn_1x_coco/retinanet_free_anchor_x101_32x4d_fpn_1x_coco_20200130-d4846968.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/free_anchor/retinanet_free_anchor_x101_32x4d_fpn_1x_coco/retinanet_free_anchor_x101_32x4d_fpn_1x_coco_20200130_095627.log.json) |
+
+**Notes:**
+
+- We use 8 GPUs with 2 images/GPU.
+- For more settings and models, please refer to the [official repo](https://github.com/zhangxiaosong18/FreeAnchor).
+
+## Citation
+
+```latex
+@inproceedings{zhang2019freeanchor,
+ title = {{FreeAnchor}: Learning to Match Anchors for Visual Object Detection},
+ author = {Zhang, Xiaosong and Wan, Fang and Liu, Chang and Ji, Rongrong and Ye, Qixiang},
+ booktitle = {Neural Information Processing Systems},
+ year = {2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/free_anchor/freeanchor_r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/free_anchor/freeanchor_r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..dc323d94f7aa20b38e2204a38ed8e234dd4eadd1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/free_anchor/freeanchor_r101_fpn_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './freeanchor_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/free_anchor/freeanchor_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/free_anchor/freeanchor_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..13f64d14a1ead0431549b8569d031f72669a2e84
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/free_anchor/freeanchor_r50_fpn_1x_coco.py
@@ -0,0 +1,22 @@
+_base_ = '../retinanet/retinanet_r50_fpn_1x_coco.py'
+model = dict(
+ bbox_head=dict(
+ _delete_=True,
+ type='FreeAnchorRetinaHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ octave_base_scale=4,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[8, 16, 32, 64, 128]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ loss_bbox=dict(type='SmoothL1Loss', beta=0.11, loss_weight=0.75)))
+
+optim_wrapper = dict(clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/free_anchor/freeanchor_x101-32x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/free_anchor/freeanchor_x101-32x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8e448bc1123115d37ef9f21a33c8a6b38cd821c3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/free_anchor/freeanchor_x101-32x4d_fpn_1x_coco.py
@@ -0,0 +1,13 @@
+_base_ = './freeanchor_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/free_anchor/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/free_anchor/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..cff19db6c957c2cdc09c1f76ff230c3a611bfc01
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/free_anchor/metafile.yml
@@ -0,0 +1,79 @@
+Collections:
+ - Name: FreeAnchor
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - FreeAnchor
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/1909.02466
+ Title: 'FreeAnchor: Learning to Match Anchors for Visual Object Detection'
+ README: configs/free_anchor/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/dense_heads/free_anchor_retina_head.py#L10
+ Version: v2.0.0
+
+Models:
+ - Name: freeanchor_r50_fpn_1x_coco
+ In Collection: FreeAnchor
+ Config: configs/free_anchor/freeanchor_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.9
+ inference time (ms/im):
+ - value: 54.35
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/free_anchor/retinanet_free_anchor_r50_fpn_1x_coco/retinanet_free_anchor_r50_fpn_1x_coco_20200130-0f67375f.pth
+
+ - Name: freeanchor_r101_fpn_1x_coco
+ In Collection: FreeAnchor
+ Config: configs/free_anchor/freeanchor_r101_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.8
+ inference time (ms/im):
+ - value: 67.11
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/free_anchor/retinanet_free_anchor_r101_fpn_1x_coco/retinanet_free_anchor_r101_fpn_1x_coco_20200130-358324e6.pth
+
+ - Name: freeanchor_x101-32x4d_fpn_1x_coco
+ In Collection: FreeAnchor
+ Config: configs/free_anchor/freeanchor_x101-32x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 8.1
+ inference time (ms/im):
+ - value: 90.09
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/free_anchor/retinanet_free_anchor_x101_32x4d_fpn_1x_coco/retinanet_free_anchor_x101_32x4d_fpn_1x_coco_20200130-d4846968.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fsaf/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fsaf/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..46f60577728d3e9d8785f19d8cda34991bae06d3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fsaf/README.md
@@ -0,0 +1,57 @@
+# FSAF
+
+> [Feature Selective Anchor-Free Module for Single-Shot Object Detection](https://arxiv.org/abs/1903.00621)
+
+
+
+## Abstract
+
+We motivate and present feature selective anchor-free (FSAF) module, a simple and effective building block for single-shot object detectors. It can be plugged into single-shot detectors with feature pyramid structure. The FSAF module addresses two limitations brought up by the conventional anchor-based detection: 1) heuristic-guided feature selection; 2) overlap-based anchor sampling. The general concept of the FSAF module is online feature selection applied to the training of multi-level anchor-free branches. Specifically, an anchor-free branch is attached to each level of the feature pyramid, allowing box encoding and decoding in the anchor-free manner at an arbitrary level. During training, we dynamically assign each instance to the most suitable feature level. At the time of inference, the FSAF module can work jointly with anchor-based branches by outputting predictions in parallel. We instantiate this concept with simple implementations of anchor-free branches and online feature selection strategy. Experimental results on the COCO detection track show that our FSAF module performs better than anchor-based counterparts while being faster. When working jointly with anchor-based branches, the FSAF module robustly improves the baseline RetinaNet by a large margin under various settings, while introducing nearly free inference overhead. And the resulting best model can achieve a state-of-the-art 44.6% mAP, outperforming all existing single-shot detectors on COCO.
+
+
+

+
+
+## Introduction
+
+FSAF is an anchor-free method published in CVPR2019 ([https://arxiv.org/pdf/1903.00621.pdf](https://arxiv.org/pdf/1903.00621.pdf)).
+Actually it is equivalent to the anchor-based method with only one anchor at each feature map position in each FPN level.
+And this is how we implemented it.
+Only the anchor-free branch is released for its better compatibility with the current framework and less computational budget.
+
+In the original paper, feature maps within the central 0.2-0.5 area of a gt box are tagged as ignored. However,
+it is empirically found that a hard threshold (0.2-0.2) gives a further gain on the performance. (see the table below)
+
+## Results and Models
+
+### Results on R50/R101/X101-FPN
+
+| Backbone | ignore range | ms-train | Lr schd | Train Mem (GB) | Train time (s/iter) | Inf time (fps) | box AP | Config | Download |
+| :------: | :----------: | :------: | :-----: | :------------: | :-----------------: | :------------: | :---------: | :----------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | 0.2-0.5 | N | 1x | 3.15 | 0.43 | 12.3 | 36.0 (35.9) | | [model](https://download.openmmlab.com/mmdetection/v2.0/fsaf/fsaf_pscale0.2_nscale0.5_r50_fpn_1x_coco/fsaf_pscale0.2_nscale0.5_r50_fpn_1x_coco_20200715-b555b0e0.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fsaf/fsaf_pscale0.2_nscale0.5_r50_fpn_1x_coco/fsaf_pscale0.2_nscale0.5_r50_fpn_1x_coco_20200715_094657.log.json) |
+| R-50 | 0.2-0.2 | N | 1x | 3.15 | 0.43 | 13.0 | 37.4 | [config](./fsaf_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fsaf/fsaf_r50_fpn_1x_coco/fsaf_r50_fpn_1x_coco-94ccc51f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fsaf/fsaf_r50_fpn_1x_coco/fsaf_r50_fpn_1x_coco_20200428_072327.log.json) |
+| R-101 | 0.2-0.2 | N | 1x | 5.08 | 0.58 | 10.8 | 39.3 (37.9) | [config](./fsaf_r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fsaf/fsaf_r101_fpn_1x_coco/fsaf_r101_fpn_1x_coco-9e71098f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fsaf/fsaf_r101_fpn_1x_coco/fsaf_r101_fpn_1x_coco_20200428_160348.log.json) |
+| X-101 | 0.2-0.2 | N | 1x | 9.38 | 1.23 | 5.6 | 42.4 (41.0) | [config](./fsaf_x101-64x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fsaf/fsaf_x101_64x4d_fpn_1x_coco/fsaf_x101_64x4d_fpn_1x_coco-e3f6e6fd.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fsaf/fsaf_x101_64x4d_fpn_1x_coco/fsaf_x101_64x4d_fpn_1x_coco_20200428_160424.log.json) |
+
+**Notes:**
+
+- *1x means the model is trained for 12 epochs.*
+- *AP values in the brackets represent those reported in the original paper.*
+- *All results are obtained with a single model and single-scale test.*
+- *X-101 backbone represents ResNext-101-64x4d.*
+- *All pretrained backbones use pytorch style.*
+- *All models are trained on 8 Titan-XP gpus and tested on a single gpu.*
+
+## Citation
+
+BibTeX reference is as follows.
+
+```latex
+@inproceedings{zhu2019feature,
+ title={Feature Selective Anchor-Free Module for Single-Shot Object Detection},
+ author={Zhu, Chenchen and He, Yihui and Savvides, Marios},
+ booktitle={Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition},
+ pages={840--849},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fsaf/fsaf_r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fsaf/fsaf_r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..12b49fed5b6cd617aa9c05d76ed737d755992a34
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fsaf/fsaf_r101_fpn_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './fsaf_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fsaf/fsaf_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fsaf/fsaf_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e7165cd63c74ab27ff47f8255836f4c10158cf0e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fsaf/fsaf_r50_fpn_1x_coco.py
@@ -0,0 +1,47 @@
+_base_ = '../retinanet/retinanet_r50_fpn_1x_coco.py'
+# model settings
+model = dict(
+ type='FSAF',
+ bbox_head=dict(
+ type='FSAFHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ reg_decoded_bbox=True,
+ # Only anchor-free branch is implemented. The anchor generator only
+ # generates 1 anchor at each feature point, as a substitute of the
+ # grid of features.
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ octave_base_scale=1,
+ scales_per_octave=1,
+ ratios=[1.0],
+ strides=[8, 16, 32, 64, 128]),
+ bbox_coder=dict(_delete_=True, type='TBLRBBoxCoder', normalizer=4.0),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0,
+ reduction='none'),
+ loss_bbox=dict(
+ _delete_=True,
+ type='IoULoss',
+ eps=1e-6,
+ loss_weight=1.0,
+ reduction='none')),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ _delete_=True,
+ type='CenterRegionAssigner',
+ pos_scale=0.2,
+ neg_scale=0.2,
+ min_pos_iof=0.01),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False))
+
+optim_wrapper = dict(clip_grad=dict(max_norm=10, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fsaf/fsaf_x101-64x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fsaf/fsaf_x101-64x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..89c0c6344aba6e6eae5657eff60745645dd1e8dc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fsaf/fsaf_x101-64x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './fsaf_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/fsaf/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fsaf/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..daaad0d3a864b52df618a95a63c6caeaa1fd76ec
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/fsaf/metafile.yml
@@ -0,0 +1,80 @@
+Collections:
+ - Name: FSAF
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x Titan-XP GPUs
+ Architecture:
+ - FPN
+ - FSAF
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/1903.00621
+ Title: 'Feature Selective Anchor-Free Module for Single-Shot Object Detection'
+ README: configs/fsaf/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/detectors/fsaf.py#L6
+ Version: v2.1.0
+
+Models:
+ - Name: fsaf_r50_fpn_1x_coco
+ In Collection: FSAF
+ Config: configs/fsaf/fsaf_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.15
+ inference time (ms/im):
+ - value: 76.92
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fsaf/fsaf_r50_fpn_1x_coco/fsaf_r50_fpn_1x_coco-94ccc51f.pth
+
+ - Name: fsaf_r101_fpn_1x_coco
+ In Collection: FSAF
+ Config: configs/fsaf/fsaf_r101_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.08
+ inference time (ms/im):
+ - value: 92.59
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fsaf/fsaf_r101_fpn_1x_coco/fsaf_r101_fpn_1x_coco-9e71098f.pth
+
+ - Name: fsaf_x101-64x4d_fpn_1x_coco
+ In Collection: FSAF
+ Config: configs/fsaf/fsaf_x101-64x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 9.38
+ inference time (ms/im):
+ - value: 178.57
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fsaf/fsaf_x101_64x4d_fpn_1x_coco/fsaf_x101_64x4d_fpn_1x_coco-e3f6e6fd.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..1ba6f6f3e4e23d4f68bca2545bba733352d0c498
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/README.md
@@ -0,0 +1,69 @@
+# GCNet
+
+> [GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond](https://arxiv.org/abs/1904.11492)
+
+
+
+## Abstract
+
+The Non-Local Network (NLNet) presents a pioneering approach for capturing long-range dependencies, via aggregating query-specific global context to each query position. However, through a rigorous empirical analysis, we have found that the global contexts modeled by non-local network are almost the same for different query positions within an image. In this paper, we take advantage of this finding to create a simplified network based on a query-independent formulation, which maintains the accuracy of NLNet but with significantly less computation. We further observe that this simplified design shares similar structure with Squeeze-Excitation Network (SENet). Hence we unify them into a three-step general framework for global context modeling. Within the general framework, we design a better instantiation, called the global context (GC) block, which is lightweight and can effectively model the global context. The lightweight property allows us to apply it for multiple layers in a backbone network to construct a global context network (GCNet), which generally outperforms both simplified NLNet and SENet on major benchmarks for various recognition tasks.
+
+
+

+
+
+## Introduction
+
+By [Yue Cao](http://yue-cao.me), [Jiarui Xu](http://jerryxu.net), [Stephen Lin](https://scholar.google.com/citations?user=c3PYmxUAAAAJ&hl=en), Fangyun Wei, [Han Hu](https://sites.google.com/site/hanhushomepage/).
+
+We provide config files to reproduce the results in the paper for
+["GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond"](https://arxiv.org/abs/1904.11492) on COCO object detection.
+
+**GCNet** is initially described in [arxiv](https://arxiv.org/abs/1904.11492). Via absorbing advantages of Non-Local Networks (NLNet) and Squeeze-Excitation Networks (SENet), GCNet provides a simple, fast and effective approach for global context modeling, which generally outperforms both NLNet and SENet on major benchmarks for various recognition tasks.
+
+## Results and Models
+
+The results on COCO 2017val are shown in the below table.
+
+| Backbone | Model | Context | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :-------: | :---: | :------------: | :-----: | :------: | :------------: | :----: | :-----: | :-----------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | Mask | GC(c3-c5, r16) | 1x | 5.0 | | 39.7 | 35.9 | [config](./mask-rcnn_r50-gcb-r16-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r50_fpn_r16_gcb_c3-c5_1x_coco/mask_rcnn_r50_fpn_r16_gcb_c3-c5_1x_coco_20200515_211915-187da160.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r50_fpn_r16_gcb_c3-c5_1x_coco/mask_rcnn_r50_fpn_r16_gcb_c3-c5_1x_coco_20200515_211915.log.json) |
+| R-50-FPN | Mask | GC(c3-c5, r4) | 1x | 5.1 | 15.0 | 39.9 | 36.0 | [config](./mask-rcnn_r50-gcb-r4-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r50_fpn_r4_gcb_c3-c5_1x_coco/mask_rcnn_r50_fpn_r4_gcb_c3-c5_1x_coco_20200204-17235656.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r50_fpn_r4_gcb_c3-c5_1x_coco/mask_rcnn_r50_fpn_r4_gcb_c3-c5_1x_coco_20200204_024626.log.json) |
+| R-101-FPN | Mask | GC(c3-c5, r16) | 1x | 7.6 | 11.4 | 41.3 | 37.2 | [config](./mask-rcnn_r101-gcb-r16-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r101_fpn_r16_gcb_c3-c5_1x_coco/mask_rcnn_r101_fpn_r16_gcb_c3-c5_1x_coco_20200205-e58ae947.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r101_fpn_r16_gcb_c3-c5_1x_coco/mask_rcnn_r101_fpn_r16_gcb_c3-c5_1x_coco_20200205_192835.log.json) |
+| R-101-FPN | Mask | GC(c3-c5, r4) | 1x | 7.8 | 11.6 | 42.2 | 37.8 | [config](./mask-rcnn_r101-gcb-r4-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r101_fpn_r4_gcb_c3-c5_1x_coco/mask_rcnn_r101_fpn_r4_gcb_c3-c5_1x_coco_20200206-af22dc9d.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r101_fpn_r4_gcb_c3-c5_1x_coco/mask_rcnn_r101_fpn_r4_gcb_c3-c5_1x_coco_20200206_112128.log.json) |
+
+| Backbone | Model | Context | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :-------: | :--------------: | :------------: | :-----: | :------: | :------------: | :----: | :-----: | :--------------------------------------------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | Mask | - | 1x | 4.4 | 16.6 | 38.4 | 34.6 | [config](./mask-rcnn_r50-syncbn_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r50_fpn_syncbn-backbone_1x_coco/mask_rcnn_r50_fpn_syncbn-backbone_1x_coco_20200202-bb3eb55c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r50_fpn_syncbn-backbone_1x_coco/mask_rcnn_r50_fpn_syncbn-backbone_1x_coco_20200202_214122.log.json) |
+| R-50-FPN | Mask | GC(c3-c5, r16) | 1x | 5.0 | 15.5 | 40.4 | 36.2 | [config](./mask-rcnn_r50-syncbn-gcb-r16-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r50_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco/mask_rcnn_r50_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco_20200202-587b99aa.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r50_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco/mask_rcnn_r50_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco_20200202_174907.log.json) |
+| R-50-FPN | Mask | GC(c3-c5, r4) | 1x | 5.1 | 15.1 | 40.7 | 36.5 | [config](./mask-rcnn_r50-syncbn-gcb-r4-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r50_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco/mask_rcnn_r50_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco_20200202-50b90e5c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r50_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco/mask_rcnn_r50_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco_20200202_085547.log.json) |
+| R-101-FPN | Mask | - | 1x | 6.4 | 13.3 | 40.5 | 36.3 | [config](./mask-rcnn_r101-syncbn_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r101_fpn_syncbn-backbone_1x_coco/mask_rcnn_r101_fpn_syncbn-backbone_1x_coco_20200210-81658c8a.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r101_fpn_syncbn-backbone_1x_coco/mask_rcnn_r101_fpn_syncbn-backbone_1x_coco_20200210_220422.log.json) |
+| R-101-FPN | Mask | GC(c3-c5, r16) | 1x | 7.6 | 12.0 | 42.2 | 37.8 | [config](./mask-rcnn_r101-syncbn-gcb-r16-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r101_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco/mask_rcnn_r101_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco_20200207-945e77ca.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r101_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco/mask_rcnn_r101_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco_20200207_015330.log.json) |
+| R-101-FPN | Mask | GC(c3-c5, r4) | 1x | 7.8 | 11.8 | 42.2 | 37.8 | [config](./mask-rcnn_r101-syncbn-gcb-r4-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r101_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco/mask_rcnn_r101_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco_20200206-8407a3f0.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r101_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco/mask_rcnn_r101_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco_20200206_142508.log.json) |
+| X-101-FPN | Mask | - | 1x | 7.6 | 11.3 | 42.4 | 37.7 | [config](./mask-rcnn_x101-32x4d-syncbn_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_x101_32x4d_fpn_syncbn-backbone_1x_coco/mask_rcnn_x101_32x4d_fpn_syncbn-backbone_1x_coco_20200211-7584841c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_x101_32x4d_fpn_syncbn-backbone_1x_coco/mask_rcnn_x101_32x4d_fpn_syncbn-backbone_1x_coco_20200211_054326.log.json) |
+| X-101-FPN | Mask | GC(c3-c5, r16) | 1x | 8.8 | 9.8 | 43.5 | 38.6 | [config](./mask-rcnn_x101-32x4d-syncbn-gcb-r16-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco/mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco_20200211-cbed3d2c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco/mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco_20200211_164715.log.json) |
+| X-101-FPN | Mask | GC(c3-c5, r4) | 1x | 9.0 | 9.7 | 43.9 | 39.0 | [config](./mask-rcnn_x101-32x4d-syncbn-gcb-r4-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco/mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco_20200212-68164964.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco/mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco_20200212_070942.log.json) |
+| X-101-FPN | Cascade Mask | - | 1x | 9.2 | 8.4 | 44.7 | 38.6 | [config](./cascade-mask-rcnn_x101-32x4d-syncbn_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_1x_coco_20200310-d5ad2a5e.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_1x_coco_20200310_115217.log.json) |
+| X-101-FPN | Cascade Mask | GC(c3-c5, r16) | 1x | 10.3 | 7.7 | 46.2 | 39.7 | [config](./cascade-mask-rcnn_x101-32x4d-syncbn-r16-gcb-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco_20200211-10bf2463.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco_20200211_184154.log.json) |
+| X-101-FPN | Cascade Mask | GC(c3-c5, r4) | 1x | 10.6 | | 46.4 | 40.1 | [config](./cascade-mask-rcnn_x101-32x4d-syncbn-r4-gcb-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco_20200703_180653-ed035291.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco_20200703_180653.log.json) |
+| X-101-FPN | DCN Cascade Mask | - | 1x | | | 47.5 | 40.9 | [config](./cascade-mask-rcnn_x101-32x4d-syncbn-dconv-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_dconv_c3-c5_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_dconv_c3-c5_1x_coco_20210615_211019-abbc39ea.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_dconv_c3-c5_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_dconv_c3-c5_1x_coco_20210615_211019.log.json) |
+| X-101-FPN | DCN Cascade Mask | GC(c3-c5, r16) | 1x | | | 48.0 | 41.3 | [config](./cascade-mask-rcnn_x101-32x4d-syncbn-dconv-c3-c5-r16-gcb-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_dconv_c3-c5_r16_gcb_c3-c5_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_dconv_c3-c5_r16_gcb_c3-c5_1x_coco_20210615_215648-44aa598a.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_dconv_c3-c5_r16_gcb_c3-c5_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_dconv_c3-c5_r16_gcb_c3-c5_1x_coco_20210615_215648.log.json) |
+| X-101-FPN | DCN Cascade Mask | GC(c3-c5, r4) | 1x | | | 47.9 | 41.1 | [config](./cascade-mask-rcnn_x101-32x4d-syncbn-dconv-c3-c5-r4-gcb-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_dconv_c3-c5_r4_gcb_c3-c5_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_dconv_c3-c5_r4_gcb_c3-c5_1x_coco_20210615_161851-720338ec.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_dconv_c3-c5_r4_gcb_c3-c5_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_dconv_c3-c5_r4_gcb_c3-c5_1x_coco_20210615_161851.log.json) |
+
+**Notes:**
+
+- The `SyncBN` is added in the backbone for all models in **Table 2**.
+- `GC` denotes Global Context (GC) block is inserted after 1x1 conv of backbone.
+- `DCN` denotes replace 3x3 conv with 3x3 Deformable Convolution in `c3-c5` stages of backbone.
+- `r4` and `r16` denote ratio 4 and ratio 16 in GC block respectively.
+
+## Citation
+
+```latex
+@article{cao2019GCNet,
+ title={GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond},
+ author={Cao, Yue and Xu, Jiarui and Lin, Stephen and Wei, Fangyun and Hu, Han},
+ journal={arXiv preprint arXiv:1904.11492},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-dconv-c3-c5-r16-gcb-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-dconv-c3-c5-r16-gcb-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6cf605b666e460aee48adc629b0604af4c64e306
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-dconv-c3-c5-r16-gcb-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,11 @@
+_base_ = '../dcn/cascade-mask-rcnn_x101-32x4d-dconv-c3-c5_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ norm_eval=False,
+ plugins=[
+ dict(
+ cfg=dict(type='ContextBlock', ratio=1. / 16),
+ stages=(False, True, True, True),
+ position='after_conv3')
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-dconv-c3-c5-r4-gcb-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-dconv-c3-c5-r4-gcb-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..95fc687b664b25b754d4ba890ae9c9e982db65fb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-dconv-c3-c5-r4-gcb-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,11 @@
+_base_ = '../dcn/cascade-mask-rcnn_x101-32x4d-dconv-c3-c5_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ norm_eval=False,
+ plugins=[
+ dict(
+ cfg=dict(type='ContextBlock', ratio=1. / 4),
+ stages=(False, True, True, True),
+ position='after_conv3')
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-dconv-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-dconv-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..9b77dc9315f52f9437eb1e39f6d518f1afaa41bb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-dconv-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,4 @@
+_base_ = '../dcn/cascade-mask-rcnn_x101-32x4d-dconv-c3-c5_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ norm_cfg=dict(type='SyncBN', requires_grad=True), norm_eval=False))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-r16-gcb-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-r16-gcb-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8f97972aa2b7d151d5824de40da9cedae9c57535
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-r16-gcb-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,11 @@
+_base_ = '../cascade_rcnn/cascade-mask-rcnn_x101-32x4d_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ norm_eval=False,
+ plugins=[
+ dict(
+ cfg=dict(type='ContextBlock', ratio=1. / 16),
+ stages=(False, True, True, True),
+ position='after_conv3')
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-r4-gcb-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-r4-gcb-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8404cfdaf34e470d2bff57a707ca8183fe442131
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-r4-gcb-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,11 @@
+_base_ = '../cascade_rcnn/cascade-mask-rcnn_x101-32x4d_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ norm_eval=False,
+ plugins=[
+ dict(
+ cfg=dict(type='ContextBlock', ratio=1. / 4),
+ stages=(False, True, True, True),
+ position='after_conv3')
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..87667dee779ee8068075be17638a6d10a9985c7e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn_fpn_1x_coco.py
@@ -0,0 +1,4 @@
+_base_ = '../cascade_rcnn/cascade-mask-rcnn_x101-32x4d_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ norm_cfg=dict(type='SyncBN', requires_grad=True), norm_eval=False))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r101-gcb-r16-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r101-gcb-r16-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..447e2c6d858738db0f0d2e46e57e1fccd2233af3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r101-gcb-r16-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,8 @@
+_base_ = '../mask_rcnn/mask-rcnn_r101_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(plugins=[
+ dict(
+ cfg=dict(type='ContextBlock', ratio=1. / 16),
+ stages=(False, True, True, True),
+ position='after_conv3')
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r101-gcb-r4-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r101-gcb-r4-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..9c723a64b6f686b9dd0f8e7648c7b1b303205168
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r101-gcb-r4-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,8 @@
+_base_ = '../mask_rcnn/mask-rcnn_r101_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(plugins=[
+ dict(
+ cfg=dict(type='ContextBlock', ratio=1. / 4),
+ stages=(False, True, True, True),
+ position='after_conv3')
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r101-syncbn-gcb-r16-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r101-syncbn-gcb-r16-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6f9d03d3f8d94116b4814825ad8377b534a912b1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r101-syncbn-gcb-r16-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,11 @@
+_base_ = '../mask_rcnn/mask-rcnn_r101_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ norm_eval=False,
+ plugins=[
+ dict(
+ cfg=dict(type='ContextBlock', ratio=1. / 16),
+ stages=(False, True, True, True),
+ position='after_conv3')
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r101-syncbn-gcb-r4-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r101-syncbn-gcb-r4-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d07cb0d488c0df76a137bad54123a7583c7da87b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r101-syncbn-gcb-r4-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,11 @@
+_base_ = '../mask_rcnn/mask-rcnn_r101_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ norm_eval=False,
+ plugins=[
+ dict(
+ cfg=dict(type='ContextBlock', ratio=1. / 4),
+ stages=(False, True, True, True),
+ position='after_conv3')
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r101-syncbn_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r101-syncbn_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..957bdf55470017d9ac9fa482b416c2206266af86
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r101-syncbn_fpn_1x_coco.py
@@ -0,0 +1,4 @@
+_base_ = '../mask_rcnn/mask-rcnn_r101_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ norm_cfg=dict(type='SyncBN', requires_grad=True), norm_eval=False))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r50-gcb-r16-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r50-gcb-r16-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c9ec5ac3baf7c46ea95d4c3fcf4f5da4ad7a3dce
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r50-gcb-r16-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,8 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(plugins=[
+ dict(
+ cfg=dict(type='ContextBlock', ratio=1. / 16),
+ stages=(False, True, True, True),
+ position='after_conv3')
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r50-gcb-r4-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r50-gcb-r4-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..42474d5196a8a130999db735989b423664486304
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r50-gcb-r4-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,8 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(plugins=[
+ dict(
+ cfg=dict(type='ContextBlock', ratio=1. / 4),
+ stages=(False, True, True, True),
+ position='after_conv3')
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r50-syncbn-gcb-r16-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r50-syncbn-gcb-r16-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ac1928082405baebfe5ec483f37b9775da21d5ad
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r50-syncbn-gcb-r16-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,11 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ norm_eval=False,
+ plugins=[
+ dict(
+ cfg=dict(type='ContextBlock', ratio=1. / 16),
+ stages=(False, True, True, True),
+ position='after_conv3')
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r50-syncbn-gcb-r4-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r50-syncbn-gcb-r4-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ae29f0cebe4f9fe16f2fea3de53874914186da9b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r50-syncbn-gcb-r4-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,11 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ norm_eval=False,
+ plugins=[
+ dict(
+ cfg=dict(type='ContextBlock', ratio=1. / 4),
+ stages=(False, True, True, True),
+ position='after_conv3')
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r50-syncbn_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r50-syncbn_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f8ef27bad9743cba8f7134f1a77a091af1bca093
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_r50-syncbn_fpn_1x_coco.py
@@ -0,0 +1,4 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ norm_cfg=dict(type='SyncBN', requires_grad=True), norm_eval=False))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_x101-32x4d-syncbn-gcb-r16-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_x101-32x4d-syncbn-gcb-r16-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1a2e2c9f26b25c5aefba912997cd01db60854a5e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_x101-32x4d-syncbn-gcb-r16-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,11 @@
+_base_ = '../mask_rcnn/mask-rcnn_x101-32x4d_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ norm_eval=False,
+ plugins=[
+ dict(
+ cfg=dict(type='ContextBlock', ratio=1. / 16),
+ stages=(False, True, True, True),
+ position='after_conv3')
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_x101-32x4d-syncbn-gcb-r4-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_x101-32x4d-syncbn-gcb-r4-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..65d3f9aadf5f79a4fb9fc9082dfabfdb3de08871
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_x101-32x4d-syncbn-gcb-r4-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,11 @@
+_base_ = '../mask_rcnn/mask-rcnn_x101-32x4d_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ norm_eval=False,
+ plugins=[
+ dict(
+ cfg=dict(type='ContextBlock', ratio=1. / 4),
+ stages=(False, True, True, True),
+ position='after_conv3')
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_x101-32x4d-syncbn_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_x101-32x4d-syncbn_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b5343a6d4596eb82245ef078d36a5a6ce5137aeb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/mask-rcnn_x101-32x4d-syncbn_fpn_1x_coco.py
@@ -0,0 +1,4 @@
+_base_ = '../mask_rcnn/mask-rcnn_x101-32x4d_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ norm_cfg=dict(type='SyncBN', requires_grad=True), norm_eval=False))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..075a94c8fbf4c5f629d9343cc841f94f18472195
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gcnet/metafile.yml
@@ -0,0 +1,440 @@
+Collections:
+ - Name: GCNet
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Global Context Block
+ - FPN
+ - RPN
+ - ResNet
+ - ResNeXt
+ Paper:
+ URL: https://arxiv.org/abs/1904.11492
+ Title: 'GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond'
+ README: configs/gcnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/ops/context_block.py#L13
+ Version: v2.0.0
+
+Models:
+ - Name: mask-rcnn_r50_fpn_r16_gcb_c3-c5_1x_coco
+ In Collection: GCNet
+ Config: configs/gcnet/mask-rcnn_r50-gcb-r16-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.0
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.7
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 35.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r50_fpn_r16_gcb_c3-c5_1x_coco/mask_rcnn_r50_fpn_r16_gcb_c3-c5_1x_coco_20200515_211915-187da160.pth
+
+ - Name: mask-rcnn_r50_fpn_r4_gcb_c3-c5_1x_coco
+ In Collection: GCNet
+ Config: configs/gcnet/mask-rcnn_r50-gcb-r4-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.1
+ inference time (ms/im):
+ - value: 66.67
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.9
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r50_fpn_r4_gcb_c3-c5_1x_coco/mask_rcnn_r50_fpn_r4_gcb_c3-c5_1x_coco_20200204-17235656.pth
+
+ - Name: mask-rcnn_r101-gcb-r16-c3-c5_fpn_1x_coco
+ In Collection: GCNet
+ Config: configs/gcnet/mask-rcnn_r101-gcb-r16-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.6
+ inference time (ms/im):
+ - value: 87.72
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.3
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r101_fpn_r16_gcb_c3-c5_1x_coco/mask_rcnn_r101_fpn_r16_gcb_c3-c5_1x_coco_20200205-e58ae947.pth
+
+ - Name: mask-rcnn_r101-gcb-r4-c3-c5_fpn_1x_coco
+ In Collection: GCNet
+ Config: configs/gcnet/mask-rcnn_r101-gcb-r4-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.8
+ inference time (ms/im):
+ - value: 86.21
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r101_fpn_r4_gcb_c3-c5_1x_coco/mask_rcnn_r101_fpn_r4_gcb_c3-c5_1x_coco_20200206-af22dc9d.pth
+
+ - Name: mask-rcnn_r50_fpn_syncbn-backbone_1x_coco
+ In Collection: GCNet
+ Config: configs/gcnet/mask-rcnn_r50-syncbn_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.4
+ inference time (ms/im):
+ - value: 60.24
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 34.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r50_fpn_syncbn-backbone_1x_coco/mask_rcnn_r50_fpn_syncbn-backbone_1x_coco_20200202-bb3eb55c.pth
+
+ - Name: mask-rcnn_r50_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco
+ In Collection: GCNet
+ Config: configs/gcnet/mask-rcnn_r50-syncbn-gcb-r16-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.0
+ inference time (ms/im):
+ - value: 64.52
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r50_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco/mask_rcnn_r50_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco_20200202-587b99aa.pth
+
+ - Name: mask-rcnn_r50_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco
+ In Collection: GCNet
+ Config: configs/gcnet/mask-rcnn_r50-syncbn-gcb-r4-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.1
+ inference time (ms/im):
+ - value: 66.23
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.7
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r50_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco/mask_rcnn_r50_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco_20200202-50b90e5c.pth
+
+ - Name: mask-rcnn_r101-syncbn_fpn_1x_coco
+ In Collection: GCNet
+ Config: configs/gcnet/mask-rcnn_r101-syncbn_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.4
+ inference time (ms/im):
+ - value: 75.19
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r101_fpn_syncbn-backbone_1x_coco/mask_rcnn_r101_fpn_syncbn-backbone_1x_coco_20200210-81658c8a.pth
+
+ - Name: mask-rcnn_r101-syncbn-gcb-r16-c3-c5_fpn_1x_coco
+ In Collection: GCNet
+ Config: configs/gcnet/mask-rcnn_r101-syncbn-gcb-r16-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.6
+ inference time (ms/im):
+ - value: 83.33
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r101_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco/mask_rcnn_r101_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco_20200207-945e77ca.pth
+
+ - Name: mask-rcnn_r101-syncbn-gcb-r4-c3-c5_fpn_1x_coco
+ In Collection: GCNet
+ Config: configs/gcnet/mask-rcnn_r101-syncbn-gcb-r4-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.8
+ inference time (ms/im):
+ - value: 84.75
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r101_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco/mask_rcnn_r101_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco_20200206-8407a3f0.pth
+
+ - Name: mask-rcnn_x101-32x4d-syncbn_fpn_1x_coco
+ In Collection: GCNet
+ Config: configs/gcnet/mask-rcnn_x101-32x4d-syncbn_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.6
+ inference time (ms/im):
+ - value: 88.5
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_x101_32x4d_fpn_syncbn-backbone_1x_coco/mask_rcnn_x101_32x4d_fpn_syncbn-backbone_1x_coco_20200211-7584841c.pth
+
+ - Name: mask-rcnn_x101-32x4d-syncbn-gcb-r16-c3-c5_fpn_1x_coco
+ In Collection: GCNet
+ Config: configs/gcnet/mask-rcnn_x101-32x4d-syncbn-gcb-r16-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 8.8
+ inference time (ms/im):
+ - value: 102.04
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco/mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco_20200211-cbed3d2c.pth
+
+ - Name: mask-rcnn_x101-32x4d-syncbn-gcb-r4-c3-c5_fpn_1x_coco
+ In Collection: GCNet
+ Config: configs/gcnet/mask-rcnn_x101-32x4d-syncbn-gcb-r4-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 9.0
+ inference time (ms/im):
+ - value: 103.09
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.9
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco/mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco_20200212-68164964.pth
+
+ - Name: cascade-mask-rcnn_x101-32x4d-syncbn_fpn_1x_coco
+ In Collection: GCNet
+ Config: configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 9.2
+ inference time (ms/im):
+ - value: 119.05
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.7
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gcnet/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_1x_coco_20200310-d5ad2a5e.pth
+
+ - Name: cascade-mask-rcnn_x101-32x4d-syncbn-r16-gcb-c3-c5_fpn_1x_coco
+ In Collection: GCNet
+ Config: configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-r16-gcb-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 10.3
+ inference time (ms/im):
+ - value: 129.87
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gcnet/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r16_gcb_c3-c5_1x_coco_20200211-10bf2463.pth
+
+ - Name: cascade-mask-rcnn_x101-32x4d-syncbn-r4-gcb-c3-c5_fpn_1x_coco
+ In Collection: GCNet
+ Config: configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-r4-gcb-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 10.6
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 40.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gcnet/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco_20200703_180653-ed035291.pth
+
+ - Name: cascade-mask-rcnn_x101-32x4d-syncbn-dconv-c3-c5_fpn_1x_coco
+ In Collection: GCNet
+ Config: configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-dconv-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 47.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 40.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gcnet/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_dconv_c3-c5_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_dconv_c3-c5_1x_coco_20210615_211019-abbc39ea.pth
+
+ - Name: cascade-mask-rcnn_x101-32x4d-syncbn-dconv-c3-c5-r16-gcb-c3-c5_fpn_1x_coco
+ In Collection: GCNet
+ Config: configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-dconv-c3-c5-r16-gcb-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 48.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 41.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gcnet/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_dconv_c3-c5_r16_gcb_c3-c5_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_dconv_c3-c5_r16_gcb_c3-c5_1x_coco_20210615_215648-44aa598a.pth
+
+ - Name: cascade-mask-rcnn_x101-32x4d-syncbn-dconv-c3-c5-r4-gcb-c3-c5_fpn_1x_coco
+ In Collection: GCNet
+ Config: configs/gcnet/cascade-mask-rcnn_x101-32x4d-syncbn-dconv-c3-c5-r4-gcb-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 47.9
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 41.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gcnet/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_dconv_c3-c5_r4_gcb_c3-c5_1x_coco/cascade_mask_rcnn_x101_32x4d_fpn_syncbn-backbone_dconv_c3-c5_r4_gcb_c3-c5_1x_coco_20210615_161851-720338ec.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..123f303ab422032aa2bbd2900a7c690d1a496eef
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/README.md
@@ -0,0 +1,42 @@
+# GFL
+
+> [Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object Detection](https://arxiv.org/abs/2006.04388)
+
+
+
+## Abstract
+
+One-stage detector basically formulates object detection as dense classification and localization. The classification is usually optimized by Focal Loss and the box location is commonly learned under Dirac delta distribution. A recent trend for one-stage detectors is to introduce an individual prediction branch to estimate the quality of localization, where the predicted quality facilitates the classification to improve detection performance. This paper delves into the representations of the above three fundamental elements: quality estimation, classification and localization. Two problems are discovered in existing practices, including (1) the inconsistent usage of the quality estimation and classification between training and inference and (2) the inflexible Dirac delta distribution for localization when there is ambiguity and uncertainty in complex scenes. To address the problems, we design new representations for these elements. Specifically, we merge the quality estimation into the class prediction vector to form a joint representation of localization quality and classification, and use a vector to represent arbitrary distribution of box locations. The improved representations eliminate the inconsistency risk and accurately depict the flexible distribution in real data, but contain continuous labels, which is beyond the scope of Focal Loss. We then propose Generalized Focal Loss (GFL) that generalizes Focal Loss from its discrete form to the continuous version for successful optimization. On COCO test-dev, GFL achieves 45.0% AP using ResNet-101 backbone, surpassing state-of-the-art SAPD (43.5%) and ATSS (43.6%) with higher or comparable inference speed, under the same backbone and training settings. Notably, our best model can achieve a single-model single-scale AP of 48.2%, at 10 FPS on a single 2080Ti GPU.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Style | Lr schd | Multi-scale Training | Inf time (fps) | box AP | Config | Download |
+| :---------------: | :-----: | :-----: | :------------------: | :------------: | :----: | :------------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | pytorch | 1x | No | 19.5 | 40.2 | [config](./gfl_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_r50_fpn_1x_coco/gfl_r50_fpn_1x_coco_20200629_121244-25944287.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_r50_fpn_1x_coco/gfl_r50_fpn_1x_coco_20200629_121244.log.json) |
+| R-50 | pytorch | 2x | Yes | 19.5 | 42.9 | [config](./gfl_r50_fpn_ms-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_r50_fpn_mstrain_2x_coco/gfl_r50_fpn_mstrain_2x_coco_20200629_213802-37bb1edc.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_r50_fpn_mstrain_2x_coco/gfl_r50_fpn_mstrain_2x_coco_20200629_213802.log.json) |
+| R-101 | pytorch | 2x | Yes | 14.7 | 44.7 | [config](./gfl_r101_fpn_ms-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_r101_fpn_mstrain_2x_coco/gfl_r101_fpn_mstrain_2x_coco_20200629_200126-dd12f847.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_r101_fpn_mstrain_2x_coco/gfl_r101_fpn_mstrain_2x_coco_20200629_200126.log.json) |
+| R-101-dcnv2 | pytorch | 2x | Yes | 12.9 | 47.1 | [config](./gfl_r101-dconv-c3-c5_fpn_ms-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_r101_fpn_dconv_c3-c5_mstrain_2x_coco/gfl_r101_fpn_dconv_c3-c5_mstrain_2x_coco_20200630_102002-134b07df.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_r101_fpn_dconv_c3-c5_mstrain_2x_coco/gfl_r101_fpn_dconv_c3-c5_mstrain_2x_coco_20200630_102002.log.json) |
+| X-101-32x4d | pytorch | 2x | Yes | 12.1 | 45.9 | [config](./gfl_x101-32x4d_fpn_ms-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_x101_32x4d_fpn_mstrain_2x_coco/gfl_x101_32x4d_fpn_mstrain_2x_coco_20200630_102002-50c1ffdb.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_x101_32x4d_fpn_mstrain_2x_coco/gfl_x101_32x4d_fpn_mstrain_2x_coco_20200630_102002.log.json) |
+| X-101-32x4d-dcnv2 | pytorch | 2x | Yes | 10.7 | 48.1 | [config](./gfl_x101-32x4d-dconv-c4-c5_fpn_ms-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_x101_32x4d_fpn_dconv_c4-c5_mstrain_2x_coco/gfl_x101_32x4d_fpn_dconv_c4-c5_mstrain_2x_coco_20200630_102002-14a2bf25.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_x101_32x4d_fpn_dconv_c4-c5_mstrain_2x_coco/gfl_x101_32x4d_fpn_dconv_c4-c5_mstrain_2x_coco_20200630_102002.log.json) |
+
+\[1\] *1x and 2x mean the model is trained for 90K and 180K iterations, respectively.* \
+\[2\] *All results are obtained with a single model and without any test time data augmentation such as multi-scale, flipping and etc..* \
+\[3\] *`dcnv2` denotes deformable convolutional networks v2.* \
+\[4\] *FPS is tested with a single GeForce RTX 2080Ti GPU, using a batch size of 1.*
+
+## Citation
+
+We provide config files to reproduce the object detection results in the paper [Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object Detection](https://arxiv.org/abs/2006.04388)
+
+```latex
+@article{li2020generalized,
+ title={Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object Detection},
+ author={Li, Xiang and Wang, Wenhai and Wu, Lijun and Chen, Shuo and Hu, Xiaolin and Li, Jun and Tang, Jinhui and Yang, Jian},
+ journal={arXiv preprint arXiv:2006.04388},
+ year={2020}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/gfl_r101-dconv-c3-c5_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/gfl_r101-dconv-c3-c5_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..7f748935b62884fd501af7e6731ad3ef6ce0effb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/gfl_r101-dconv-c3-c5_fpn_ms-2x_coco.py
@@ -0,0 +1,15 @@
+_base_ = './gfl_r50_fpn_ms-2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNet',
+ depth=101,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ dcn=dict(type='DCN', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/gfl_r101_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/gfl_r101_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..10135f161b9e933612d961af12a8e30198cca484
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/gfl_r101_fpn_ms-2x_coco.py
@@ -0,0 +1,13 @@
+_base_ = './gfl_r50_fpn_ms-2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNet',
+ depth=101,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/gfl_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/gfl_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..902382552d58f124bbe2b8c2904ce74ec7b7a4d8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/gfl_r50_fpn_1x_coco.py
@@ -0,0 +1,66 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ type='GFL',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output',
+ num_outs=5),
+ bbox_head=dict(
+ type='GFLHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ octave_base_scale=8,
+ scales_per_octave=1,
+ strides=[8, 16, 32, 64, 128]),
+ loss_cls=dict(
+ type='QualityFocalLoss',
+ use_sigmoid=True,
+ beta=2.0,
+ loss_weight=1.0),
+ loss_dfl=dict(type='DistributionFocalLoss', loss_weight=0.25),
+ reg_max=16,
+ loss_bbox=dict(type='GIoULoss', loss_weight=2.0)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(type='ATSSAssigner', topk=9),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/gfl_r50_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/gfl_r50_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..22770eb101920f9daae750a1b72f5410be395743
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/gfl_r50_fpn_ms-2x_coco.py
@@ -0,0 +1,28 @@
+_base_ = './gfl_r50_fpn_1x_coco.py'
+max_epochs = 24
+
+# learning policy
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs)
+
+# multi-scale training
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize', scale=[(1333, 480), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/gfl_x101-32x4d-dconv-c4-c5_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/gfl_x101-32x4d-dconv-c4-c5_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6aa98eea2d0d25b4df1570aed97cce8475e9104d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/gfl_x101-32x4d-dconv-c4-c5_fpn_ms-2x_coco.py
@@ -0,0 +1,18 @@
+_base_ = './gfl_r50_fpn_ms-2x_coco.py'
+model = dict(
+ type='GFL',
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ dcn=dict(type='DCN', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, False, True, True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/gfl_x101-32x4d_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/gfl_x101-32x4d_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ec629b1f0d5d3317dcb20f1244bc713818518d8a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/gfl_x101-32x4d_fpn_ms-2x_coco.py
@@ -0,0 +1,16 @@
+_base_ = './gfl_r50_fpn_ms-2x_coco.py'
+model = dict(
+ type='GFL',
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..183fc14bdee0492c7ea3fc18ccb7371682dc0066
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gfl/metafile.yml
@@ -0,0 +1,134 @@
+Collections:
+ - Name: Generalized Focal Loss
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Generalized Focal Loss
+ - FPN
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/2006.04388
+ Title: 'Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object Detection'
+ README: configs/gfl/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.2.0/mmdet/models/detectors/gfl.py#L6
+ Version: v2.2.0
+
+Models:
+ - Name: gfl_r50_fpn_1x_coco
+ In Collection: Generalized Focal Loss
+ Config: configs/gfl/gfl_r50_fpn_1x_coco.py
+ Metadata:
+ inference time (ms/im):
+ - value: 51.28
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_r50_fpn_1x_coco/gfl_r50_fpn_1x_coco_20200629_121244-25944287.pth
+
+ - Name: gfl_r50_fpn_ms-2x_coco
+ In Collection: Generalized Focal Loss
+ Config: configs/gfl/gfl_r50_fpn_ms-2x_coco.py
+ Metadata:
+ inference time (ms/im):
+ - value: 51.28
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_r50_fpn_mstrain_2x_coco/gfl_r50_fpn_mstrain_2x_coco_20200629_213802-37bb1edc.pth
+
+ - Name: gfl_r101_fpn_ms-2x_coco
+ In Collection: Generalized Focal Loss
+ Config: configs/gfl/gfl_r101_fpn_ms-2x_coco.py
+ Metadata:
+ inference time (ms/im):
+ - value: 68.03
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_r101_fpn_mstrain_2x_coco/gfl_r101_fpn_mstrain_2x_coco_20200629_200126-dd12f847.pth
+
+ - Name: gfl_r101-dconv-c3-c5_fpn_ms-2x_coco
+ In Collection: Generalized Focal Loss
+ Config: configs/gfl/gfl_r101-dconv-c3-c5_fpn_ms-2x_coco.py
+ Metadata:
+ inference time (ms/im):
+ - value: 77.52
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 47.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_r101_fpn_dconv_c3-c5_mstrain_2x_coco/gfl_r101_fpn_dconv_c3-c5_mstrain_2x_coco_20200630_102002-134b07df.pth
+
+ - Name: gfl_x101-32x4d_fpn_ms-2x_coco
+ In Collection: Generalized Focal Loss
+ Config: configs/gfl/gfl_x101-32x4d_fpn_ms-2x_coco.py
+ Metadata:
+ inference time (ms/im):
+ - value: 82.64
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_x101_32x4d_fpn_mstrain_2x_coco/gfl_x101_32x4d_fpn_mstrain_2x_coco_20200630_102002-50c1ffdb.pth
+
+ - Name: gfl_x101-32x4d-dconv-c4-c5_fpn_ms-2x_coco
+ In Collection: Generalized Focal Loss
+ Config: configs/gfl/gfl_x101-32x4d-dconv-c4-c5_fpn_ms-2x_coco.py
+ Metadata:
+ inference time (ms/im):
+ - value: 93.46
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 48.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_x101_32x4d_fpn_dconv_c4-c5_mstrain_2x_coco/gfl_x101_32x4d_fpn_dconv_c4-c5_mstrain_2x_coco_20200630_102002-14a2bf25.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ghm/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ghm/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..c245cea59d45f2a1a2691ce8019bf12db4af7188
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ghm/README.md
@@ -0,0 +1,33 @@
+# GHM
+
+> [Gradient Harmonized Single-stage Detector](https://arxiv.org/abs/1811.05181)
+
+
+
+## Abstract
+
+Despite the great success of two-stage detectors, single-stage detector is still a more elegant and efficient way, yet suffers from the two well-known disharmonies during training, i.e. the huge difference in quantity between positive and negative examples as well as between easy and hard examples. In this work, we first point out that the essential effect of the two disharmonies can be summarized in term of the gradient. Further, we propose a novel gradient harmonizing mechanism (GHM) to be a hedging for the disharmonies. The philosophy behind GHM can be easily embedded into both classification loss function like cross-entropy (CE) and regression loss function like smooth-L1 (SL1) loss. To this end, two novel loss functions called GHM-C and GHM-R are designed to balancing the gradient flow for anchor classification and bounding box refinement, respectively. Ablation study on MS COCO demonstrates that without laborious hyper-parameter tuning, both GHM-C and GHM-R can bring substantial improvement for single-stage detector. Without any whistles and bells, our model achieves 41.6 mAP on COCO test-dev set which surpasses the state-of-the-art method, Focal Loss (FL) + SL1, by 0.8.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :-------------: | :-----: | :-----: | :------: | :------------: | :----: | :-------------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | pytorch | 1x | 4.0 | 3.3 | 37.0 | [config](./retinanet_r50_fpn_ghm-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/ghm/retinanet_ghm_r50_fpn_1x_coco/retinanet_ghm_r50_fpn_1x_coco_20200130-a437fda3.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/ghm/retinanet_ghm_r50_fpn_1x_coco/retinanet_ghm_r50_fpn_1x_coco_20200130_004213.log.json) |
+| R-101-FPN | pytorch | 1x | 6.0 | 4.4 | 39.1 | [config](./retinanet_r101_fpn_ghm-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/ghm/retinanet_ghm_r101_fpn_1x_coco/retinanet_ghm_r101_fpn_1x_coco_20200130-c148ee8f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/ghm/retinanet_ghm_r101_fpn_1x_coco/retinanet_ghm_r101_fpn_1x_coco_20200130_145259.log.json) |
+| X-101-32x4d-FPN | pytorch | 1x | 7.2 | 5.1 | 40.7 | [config](./retinanet_x101-32x4d_fpn_ghm-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/ghm/retinanet_ghm_x101_32x4d_fpn_1x_coco/retinanet_ghm_x101_32x4d_fpn_1x_coco_20200131-e4333bd0.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/ghm/retinanet_ghm_x101_32x4d_fpn_1x_coco/retinanet_ghm_x101_32x4d_fpn_1x_coco_20200131_113653.log.json) |
+| X-101-64x4d-FPN | pytorch | 1x | 10.3 | 5.2 | 41.4 | [config](./retinanet_x101-64x4d_fpn_ghm-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/ghm/retinanet_ghm_x101_64x4d_fpn_1x_coco/retinanet_ghm_x101_64x4d_fpn_1x_coco_20200131-dd381cef.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/ghm/retinanet_ghm_x101_64x4d_fpn_1x_coco/retinanet_ghm_x101_64x4d_fpn_1x_coco_20200131_113723.log.json) |
+
+## Citation
+
+```latex
+@inproceedings{li2019gradient,
+ title={Gradient Harmonized Single-stage Detector},
+ author={Li, Buyu and Liu, Yu and Wang, Xiaogang},
+ booktitle={AAAI Conference on Artificial Intelligence},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ghm/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ghm/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..63cb48ffe7323686c38fcb279dde9ee6387e9be7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ghm/metafile.yml
@@ -0,0 +1,101 @@
+Collections:
+ - Name: GHM
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - GHM-C
+ - GHM-R
+ - FPN
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/1811.05181
+ Title: 'Gradient Harmonized Single-stage Detector'
+ README: configs/ghm/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/losses/ghm_loss.py#L21
+ Version: v2.0.0
+
+Models:
+ - Name: retinanet_r50_fpn_ghm-1x_coco
+ In Collection: GHM
+ Config: configs/ghm/retinanet_r50_fpn_ghm-1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.0
+ inference time (ms/im):
+ - value: 303.03
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/ghm/retinanet_ghm_r50_fpn_1x_coco/retinanet_ghm_r50_fpn_1x_coco_20200130-a437fda3.pth
+
+ - Name: retinanet_r101_fpn_ghm-1x_coco
+ In Collection: GHM
+ Config: configs/ghm/retinanet_r101_fpn_ghm-1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.0
+ inference time (ms/im):
+ - value: 227.27
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/ghm/retinanet_ghm_r101_fpn_1x_coco/retinanet_ghm_r101_fpn_1x_coco_20200130-c148ee8f.pth
+
+ - Name: retinanet_x101-32x4d_fpn_ghm-1x_coco
+ In Collection: GHM
+ Config: configs/ghm/retinanet_x101-32x4d_fpn_ghm-1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.2
+ inference time (ms/im):
+ - value: 196.08
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/ghm/retinanet_ghm_x101_32x4d_fpn_1x_coco/retinanet_ghm_x101_32x4d_fpn_1x_coco_20200131-e4333bd0.pth
+
+ - Name: retinanet_x101-64x4d_fpn_ghm-1x_coco
+ In Collection: GHM
+ Config: configs/ghm/retinanet_x101-64x4d_fpn_ghm-1x_coco.py
+ Metadata:
+ Training Memory (GB): 10.3
+ inference time (ms/im):
+ - value: 192.31
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/ghm/retinanet_ghm_x101_64x4d_fpn_1x_coco/retinanet_ghm_x101_64x4d_fpn_1x_coco_20200131-dd381cef.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ghm/retinanet_r101_fpn_ghm-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ghm/retinanet_r101_fpn_ghm-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..090221e68f68a95cfcf092b15f2636cd28fc9d87
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ghm/retinanet_r101_fpn_ghm-1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './retinanet_r50_fpn_ghm-1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ghm/retinanet_r50_fpn_ghm-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ghm/retinanet_r50_fpn_ghm-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..42b9aa6d05dc64f3045685a7c23d632a6041249c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ghm/retinanet_r50_fpn_ghm-1x_coco.py
@@ -0,0 +1,18 @@
+_base_ = '../retinanet/retinanet_r50_fpn_1x_coco.py'
+model = dict(
+ bbox_head=dict(
+ loss_cls=dict(
+ _delete_=True,
+ type='GHMC',
+ bins=30,
+ momentum=0.75,
+ use_sigmoid=True,
+ loss_weight=1.0),
+ loss_bbox=dict(
+ _delete_=True,
+ type='GHMR',
+ mu=0.02,
+ bins=10,
+ momentum=0.7,
+ loss_weight=10.0)))
+optim_wrapper = dict(clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ghm/retinanet_x101-32x4d_fpn_ghm-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ghm/retinanet_x101-32x4d_fpn_ghm-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1240545a624a70c7122829e85b426cafcc3f42d2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ghm/retinanet_x101-32x4d_fpn_ghm-1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './retinanet_r50_fpn_ghm-1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ghm/retinanet_x101-64x4d_fpn_ghm-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ghm/retinanet_x101-64x4d_fpn_ghm-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..689d2edcdf1bdffa52ee3aa3a8a4dac7988f6fa5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ghm/retinanet_x101-64x4d_fpn_ghm-1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './retinanet_r50_fpn_ghm-1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..e74e98d1b578824778edc4ae47741b147c420cca
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/README.md
@@ -0,0 +1,173 @@
+# GLIP: Grounded Language-Image Pre-training
+
+> [GLIP: Grounded Language-Image Pre-training](https://arxiv.org/abs/2112.03857)
+
+
+
+## Abstract
+
+This paper presents a grounded language-image pre-training (GLIP) model for learning object-level, language-aware, and semantic-rich visual representations. GLIP unifies object detection and phrase grounding for pre-training. The unification brings two benefits: 1) it allows GLIP to learn from both detection and grounding data to improve both tasks and bootstrap a good grounding model; 2) GLIP can leverage massive image-text pairs by generating grounding boxes in a self-training fashion, making the learned representation semantic-rich. In our experiments, we pre-train GLIP on 27M grounding data, including 3M human-annotated and 24M web-crawled image-text pairs. The learned representations demonstrate strong zero-shot and few-shot transferability to various object-level recognition tasks. 1) When directly evaluated on COCO and LVIS (without seeing any images in COCO during pre-training), GLIP achieves 49.8 AP and 26.9 AP, respectively, surpassing many supervised baselines. 2) After fine-tuned on COCO, GLIP achieves 60.8 AP on val and 61.5 AP on test-dev, surpassing prior SoTA. 3) When transferred to 13 downstream object detection tasks, a 1-shot GLIP rivals with a fully-supervised Dynamic Head.
+
+
+

+
+
+## Installation
+
+```shell
+cd $MMDETROOT
+
+# source installation
+pip install -r requirements/multimodal.txt
+
+# or mim installation
+mim install mmdet[multimodal]
+```
+
+```shell
+cd $MMDETROOT
+
+wget https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_a_mmdet-b3654169.pth
+
+python demo/image_demo.py demo/demo.jpg \
+configs/glip/glip_atss_swin-t_a_fpn_dyhead_pretrain_obj365.py \
+--weights glip_tiny_a_mmdet-b3654169.pth \
+--texts 'bench. car'
+```
+
+
+

+
+
+## NOTE
+
+GLIP utilizes BERT as the language model, which requires access to https://huggingface.co/. If you encounter connection errors due to network access, you can download the required files on a computer with internet access and save them locally. Finally, modify the `lang_model_name` field in the config to the local path. Please refer to the following code:
+
+```python
+from transformers import BertConfig, BertModel
+from transformers import AutoTokenizer
+
+config = BertConfig.from_pretrained("bert-base-uncased")
+model = BertModel.from_pretrained("bert-base-uncased", add_pooling_layer=False, config=config)
+tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
+
+config.save_pretrained("your path/bert-base-uncased")
+model.save_pretrained("your path/bert-base-uncased")
+tokenizer.save_pretrained("your path/bert-base-uncased")
+```
+
+## COCO Results and Models
+
+| Model | Zero-shot or Finetune | COCO mAP | Official COCO mAP | Pre-Train Data | Config | Download |
+| :--------: | :-------------------: | :------: | ----------------: | :------------------------: | :---------------------------------------------------------------------: | :-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| GLIP-T (A) | Zero-shot | 43.0 | 42.9 | O365 | [config](glip_atss_swin-t_a_fpn_dyhead_pretrain_obj365.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_a_mmdet-b3654169.pth) |
+| GLIP-T (A) | Finetune | 53.3 | 52.9 | O365 | [config](glip_atss_swin-t_a_fpn_dyhead_16xb2_ms-2x_funtune_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_atss_swin-t_a_fpn_dyhead_16xb2_ms-2x_funtune_coco/glip_atss_swin-t_a_fpn_dyhead_16xb2_ms-2x_funtune_coco_20230914_180419-e6addd96.pth)\| [log](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_atss_swin-t_a_fpn_dyhead_16xb2_ms-2x_funtune_coco/glip_atss_swin-t_a_fpn_dyhead_16xb2_ms-2x_funtune_coco_20230914_180419.log.json) |
+| GLIP-T (B) | Zero-shot | 44.9 | 44.9 | O365 | [config](glip_atss_swin-t_b_fpn_dyhead_pretrain_obj365.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_b_mmdet-6dfbd102.pth) |
+| GLIP-T (B) | Finetune | 54.1 | 53.8 | O365 | [config](glip_atss_swin-t_b_fpn_dyhead_16xb2_ms-2x_funtune_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_atss_swin-t_b_fpn_dyhead_16xb2_ms-2x_funtune_coco/glip_atss_swin-t_b_fpn_dyhead_16xb2_ms-2x_funtune_coco_20230916_163538-650323ba.pth)\| [log](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_atss_swin-t_b_fpn_dyhead_16xb2_ms-2x_funtune_coco/glip_atss_swin-t_b_fpn_dyhead_16xb2_ms-2x_funtune_coco_20230916_163538.log.json) |
+| GLIP-T (C) | Zero-shot | 46.7 | 46.7 | O365,GoldG | [config](glip_atss_swin-t_c_fpn_dyhead_pretrain_obj365-goldg.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_c_mmdet-2fc427dd.pth) |
+| GLIP-T (C) | Finetune | 55.2 | 55.1 | O365,GoldG | [config](glip_atss_swin-t_c_fpn_dyhead_16xb2_ms-2x_funtune_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_atss_swin-t_c_fpn_dyhead_16xb2_ms-2x_funtune_coco/glip_atss_swin-t_c_fpn_dyhead_16xb2_ms-2x_funtune_coco_20230914_182935-4ba3fc3b.pth)\| [log](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_atss_swin-t_c_fpn_dyhead_16xb2_ms-2x_funtune_coco/glip_atss_swin-t_c_fpn_dyhead_16xb2_ms-2x_funtune_coco_20230914_182935.log.json) |
+| GLIP-T | Zero-shot | 46.6 | 46.6 | O365,GoldG,CC3M,SBU | [config](glip_atss_swin-t_fpn_dyhead_pretrain_obj365-goldg-cc3m-sub.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_mmdet-c24ce662.pth) |
+| GLIP-T | Finetune | 55.4 | 55.2 | O365,GoldG,CC3M,SBU | [config](glip_atss_swin-t_fpn_dyhead_16xb2_ms-2x_funtune_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_atss_swin-t_fpn_dyhead_16xb2_ms-2x_funtune_coco/glip_atss_swin-t_fpn_dyhead_16xb2_ms-2x_funtune_coco_20230914_224410-ba97be24.pth)\| [log](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_atss_swin-t_fpn_dyhead_16xb2_ms-2x_funtune_coco/glip_atss_swin-t_fpn_dyhead_16xb2_ms-2x_funtune_coco_20230914_224410.log.json) |
+| GLIP-L | Zero-shot | 51.3 | 51.4 | FourODs,GoldG,CC3M+12M,SBU | [config](glip_atss_swin-l_fpn_dyhead_pretrain_mixeddata.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_l_mmdet-abfe026b.pth) |
+| GLIP-L | Finetune | 59.4 | | FourODs,GoldG,CC3M+12M,SBU | [config](glip_atss_swin-l_fpn_dyhead_16xb2_ms-2x_funtune_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_atss_swin-l_fpn_dyhead_16xb2_ms-2x_funtune_coco/glip_atss_swin-l_fpn_dyhead_16xb2_ms-2x_funtune_coco_20230910_100800-e9be4274.pth)\| [log](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_atss_swin-l_fpn_dyhead_16xb2_ms-2x_funtune_coco/glip_atss_swin-l_fpn_dyhead_16xb2_ms-2x_funtune_coco_20230910_100800.log.json) |
+
+Note:
+
+1. The weights corresponding to the zero-shot model are adopted from the official weights and converted using the [script](../../tools/model_converters/glip_to_mmdet.py). We have not retrained the model for the time being.
+2. Finetune refers to fine-tuning on the COCO 2017 dataset. The L model is trained using 16 A100 GPUs, while the remaining models are trained using 16 NVIDIA GeForce 3090 GPUs.
+3. Taking the GLIP-T(A) model as an example, I trained it twice using the official code, and the fine-tuning mAP were 52.5 and 52.6. Therefore, the mAP we achieved in our reproduction is higher than the official results. The main reason is that we modified the `weight_decay` parameter.
+4. Our experiments revealed that training for 24 epochs leads to overfitting. Therefore, we chose the best-performing model. If users want to train on a custom dataset, it is advisable to shorten the number of epochs and save the best-performing model.
+5. Due to the official absence of fine-tuning hyperparameters for the GLIP-L model, we have not yet reproduced the official accuracy. I have found that overfitting can also occur, so it may be necessary to consider custom modifications to data augmentation and model enhancement. Given the high cost of training, we have not conducted any research on this matter at the moment.
+
+## LVIS Results
+
+| Model | Official | MiniVal APr | MiniVal APc | MiniVal APf | MiniVal AP | Val1.0 APr | Val1.0 APc | Val1.0 APf | Val1.0 AP | Pre-Train Data | Config | Download |
+| :--------: | :------: | :---------: | :---------: | :---------: | :--------: | :--------: | :--------: | :--------: | :-------: | :------------------------: | :---------------------------------------------------------------------: | :------------------------------------------------------------------------------------------: |
+| GLIP-T (A) | ✔ | | | | | | | | | O365 | [config](lvis/glip_atss_swin-t_a_fpn_dyhead_pretrain_zeroshot_lvis.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_a_mmdet-b3654169.pth) |
+| GLIP-T (A) | | 12.1 | 15.5 | 25.8 | 20.2 | 6.2 | 10.9 | 22.8 | 14.7 | O365 | [config](lvis/glip_atss_swin-t_a_fpn_dyhead_pretrain_zeroshot_lvis.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_a_mmdet-b3654169.pth) |
+| GLIP-T (B) | ✔ | | | | | | | | | O365 | [config](lvis/glip_atss_swin-t_bc_fpn_dyhead_pretrain_zeroshot_lvis.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_b_mmdet-6dfbd102.pth) |
+| GLIP-T (B) | | 8.6 | 13.9 | 26.0 | 19.3 | 4.6 | 9.8 | 22.6 | 13.9 | O365 | [config](lvis/glip_atss_swin-t_bc_fpn_dyhead_pretrain_zeroshot_lvis.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_b_mmdet-6dfbd102.pth) |
+| GLIP-T (C) | ✔ | 14.3 | 19.4 | 31.1 | 24.6 | | | | | O365,GoldG | [config](lvis/glip_atss_swin-t_bc_fpn_dyhead_pretrain_zeroshot_lvis.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_c_mmdet-2fc427dd.pth) |
+| GLIP-T (C) | | 14.4 | 19.8 | 31.9 | 25.2 | 8.3 | 13.2 | 28.1 | 18.2 | O365,GoldG | [config](lvis/glip_atss_swin-t_bc_fpn_dyhead_pretrain_zeroshot_lvis.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_c_mmdet-2fc427dd.pth) |
+| GLIP-T | ✔ | | | | | | | | | O365,GoldG,CC3M,SBU | [config](lvis/glip_atss_swin-t_bc_fpn_dyhead_pretrain_zeroshot_lvis.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_mmdet-c24ce662.pth) |
+| GLIP-T | | 18.1 | 21.2 | 33.1 | 26.7 | 10.8 | 14.7 | 29.0 | 19.6 | O365,GoldG,CC3M,SBU | [config](lvis/glip_atss_swin-t_bc_fpn_dyhead_pretrain_zeroshot_lvis.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_mmdet-c24ce662.pth) |
+| GLIP-L | ✔ | 29.2 | 34.9 | 42.1 | 37.9 | | | | | FourODs,GoldG,CC3M+12M,SBU | [config](lvis/glip_atss_swin-l_fpn_dyhead_pretrain_zeroshot_lvis.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_l_mmdet-abfe026b.pth) |
+| GLIP-L | | 27.9 | 33.7 | 39.7 | 36.1 | 20.2 | 25.8 | 35.3 | 28.5 | FourODs,GoldG,CC3M+12M,SBU | [config](lvis/glip_atss_swin-l_fpn_dyhead_pretrain_zeroshot_lvis.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/glip/glip_l_mmdet-abfe026b.pth) |
+
+Note:
+
+1. The above are zero-shot evaluation results.
+2. The evaluation metric we used is LVIS FixAP. For specific details, please refer to [Evaluating Large-Vocabulary Object Detectors: The Devil is in the Details](https://arxiv.org/pdf/2102.01066.pdf).
+3. We found that the performance on small models is better than the official results, but it is lower on large models. This is mainly due to the incomplete alignment of the GLIP post-processing.
+
+## ODinW (Object Detection in the Wild) Results
+
+Learning visual representations from natural language supervision has recently shown great promise in a number of pioneering works. In general, these language-augmented visual models demonstrate strong transferability to a variety of datasets and tasks. However, it remains challenging to evaluate the transferablity of these models due to the lack of easy-to-use evaluation toolkits and public benchmarks. To tackle this, we build ELEVATER 1 , the first benchmark and toolkit for evaluating (pre-trained) language-augmented visual models. ELEVATER is composed of three components. (i) Datasets. As downstream evaluation suites, it consists of 20 image classification datasets and 35 object detection datasets, each of which is augmented with external knowledge. (ii) Toolkit. An automatic hyper-parameter tuning toolkit is developed to facilitate model evaluation on downstream tasks. (iii) Metrics. A variety of evaluation metrics are used to measure sample-efficiency (zero-shot and few-shot) and parameter-efficiency (linear probing and full model fine-tuning). ELEVATER is platform for Computer Vision in the Wild (CVinW), and is publicly released at https://computer-vision-in-the-wild.github.io/ELEVATER/
+
+### Results and models of ODinW13
+
+| Method | GLIP-T(A) | Official | GLIP-T(B) | Official | GLIP-T(C) | Official | GroundingDINO-T | GroundingDINO-B |
+| --------------------- | --------- | --------- | --------- | --------- | --------- | --------- | --------------- | --------------- |
+| AerialMaritimeDrone | 0.123 | 0.122 | 0.110 | 0.110 | 0.130 | 0.130 | 0.173 | 0.281 |
+| Aquarium | 0.175 | 0.174 | 0.173 | 0.169 | 0.191 | 0.190 | 0.195 | 0.445 |
+| CottontailRabbits | 0.686 | 0.686 | 0.688 | 0.688 | 0.744 | 0.744 | 0.799 | 0.808 |
+| EgoHands | 0.013 | 0.013 | 0.003 | 0.004 | 0.314 | 0.315 | 0.608 | 0.764 |
+| NorthAmericaMushrooms | 0.502 | 0.502 | 0.367 | 0.367 | 0.297 | 0.296 | 0.507 | 0.675 |
+| Packages | 0.589 | 0.589 | 0.083 | 0.083 | 0.699 | 0.699 | 0.687 | 0.670 |
+| PascalVOC | 0.512 | 0.512 | 0.541 | 0.540 | 0.565 | 0.565 | 0.563 | 0.711 |
+| pistols | 0.339 | 0.339 | 0.502 | 0.501 | 0.503 | 0.504 | 0.726 | 0.771 |
+| pothole | 0.007 | 0.007 | 0.030 | 0.030 | 0.058 | 0.058 | 0.215 | 0.478 |
+| Raccoon | 0.075 | 0.074 | 0.285 | 0.288 | 0.241 | 0.244 | 0.549 | 0.541 |
+| ShellfishOpenImages | 0.253 | 0.253 | 0.337 | 0.338 | 0.300 | 0.302 | 0.393 | 0.650 |
+| thermalDogsAndPeople | 0.372 | 0.372 | 0.475 | 0.475 | 0.510 | 0.510 | 0.657 | 0.633 |
+| VehiclesOpenImages | 0.574 | 0.566 | 0.562 | 0.547 | 0.549 | 0.534 | 0.613 | 0.647 |
+| Average | **0.325** | **0.324** | **0.320** | **0.318** | **0.392** | **0.392** | **0.514** | **0.621** |
+
+### Results and models of ODinW35
+
+| Method | GLIP-T(A) | Official | GLIP-T(B) | Official | GLIP-T(C) | Official | GroundingDINO-T | GroundingDINO-B |
+| --------------------------- | --------- | --------- | --------- | --------- | --------- | --------- | --------------- | --------------- |
+| AerialMaritimeDrone_large | 0.123 | 0.122 | 0.110 | 0.110 | 0.130 | 0.130 | 0.173 | 0.281 |
+| AerialMaritimeDrone_tiled | 0.174 | 0.174 | 0.172 | 0.172 | 0.172 | 0.172 | 0.206 | 0.364 |
+| AmericanSignLanguageLetters | 0.001 | 0.001 | 0.003 | 0.003 | 0.009 | 0.009 | 0.002 | 0.096 |
+| Aquarium | 0.175 | 0.175 | 0.173 | 0.171 | 0.192 | 0.182 | 0.195 | 0.445 |
+| BCCD | 0.016 | 0.016 | 0.001 | 0.001 | 0.000 | 0.000 | 0.161 | 0.584 |
+| boggleBoards | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.134 |
+| brackishUnderwater | 0.016 | 0..013 | 0.021 | 0.027 | 0.020 | 0.022 | 0.021 | 0.454 |
+| ChessPieces | 0.001 | 0.001 | 0.000 | 0.000 | 0.001 | 0.001 | 0.000 | 0.000 |
+| CottontailRabbits | 0.710 | 0.709 | 0.683 | 0.683 | 0.752 | 0.752 | 0.806 | 0.797 |
+| dice | 0.005 | 0.005 | 0.004 | 0.004 | 0.004 | 0.004 | 0.004 | 0.082 |
+| DroneControl | 0.016 | 0.017 | 0.006 | 0.008 | 0.005 | 0.007 | 0.042 | 0.638 |
+| EgoHands_generic | 0.009 | 0.010 | 0.005 | 0.006 | 0.510 | 0.508 | 0.608 | 0.764 |
+| EgoHands_specific | 0.001 | 0.001 | 0.004 | 0.006 | 0.003 | 0.004 | 0.002 | 0.687 |
+| HardHatWorkers | 0.029 | 0.029 | 0.023 | 0.023 | 0.033 | 0.033 | 0.046 | 0.439 |
+| MaskWearing | 0.007 | 0.007 | 0.003 | 0.002 | 0.005 | 0.005 | 0.004 | 0.406 |
+| MountainDewCommercial | 0.218 | 0.227 | 0.199 | 0.197 | 0.478 | 0.463 | 0.430 | 0.580 |
+| NorthAmericaMushrooms | 0.502 | 0.502 | 0.450 | 0.450 | 0.497 | 0.497 | 0.471 | 0.501 |
+| openPoetryVision | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.051 |
+| OxfordPets_by_breed | 0.001 | 0.002 | 0.002 | 0.004 | 0.001 | 0.002 | 0.003 | 0.799 |
+| OxfordPets_by_species | 0.016 | 0.011 | 0.012 | 0.009 | 0.013 | 0.009 | 0.011 | 0.872 |
+| PKLot | 0.002 | 0.002 | 0.000 | 0.000 | 0.000 | 0.000 | 0.001 | 0.774 |
+| Packages | 0.569 | 0.569 | 0.279 | 0.279 | 0.712 | 0.712 | 0.695 | 0.728 |
+| PascalVOC | 0.512 | 0.512 | 0.541 | 0.540 | 0.565 | 0.565 | 0.563 | 0.711 |
+| pistols | 0.339 | 0.339 | 0.502 | 0.501 | 0.503 | 0.504 | 0.726 | 0.771 |
+| plantdoc | 0.002 | 0.002 | 0.007 | 0.007 | 0.009 | 0.009 | 0.005 | 0.376 |
+| pothole | 0.007 | 0.010 | 0.024 | 0.025 | 0.085 | 0.101 | 0.215 | 0.478 |
+| Raccoons | 0.075 | 0.074 | 0.285 | 0.288 | 0.241 | 0.244 | 0.549 | 0.541 |
+| selfdrivingCar | 0.071 | 0.072 | 0.074 | 0.074 | 0.081 | 0.080 | 0.089 | 0.318 |
+| ShellfishOpenImages | 0.253 | 0.253 | 0.337 | 0.338 | 0.300 | 0.302 | 0.393 | 0.650 |
+| ThermalCheetah | 0.028 | 0.028 | 0.000 | 0.000 | 0.028 | 0.028 | 0.087 | 0.290 |
+| thermalDogsAndPeople | 0.372 | 0.372 | 0.475 | 0.475 | 0.510 | 0.510 | 0.657 | 0.633 |
+| UnoCards | 0.000 | 0.000 | 0.000 | 0.001 | 0.002 | 0.003 | 0.006 | 0.754 |
+| VehiclesOpenImages | 0.574 | 0.566 | 0.562 | 0.547 | 0.549 | 0.534 | 0.613 | 0.647 |
+| WildfireSmoke | 0.000 | 0.000 | 0.000 | 0.000 | 0.017 | 0.017 | 0.134 | 0.410 |
+| websiteScreenshots | 0.003 | 0.004 | 0.003 | 0.005 | 0.005 | 0.006 | 0.012 | 0.175 |
+| Average | **0.134** | **0.134** | **0.138** | **0.138** | **0.179** | **0.178** | **0.227** | **0.492** |
+
+### Results on Flickr30k
+
+| Model | Official | Pre-Train Data | Val R@1 | Val R@5 | Val R@10 | Test R@1 | Test R@5 | Test R@10 |
+| ------------- | -------- | ------------------- | ------- | ------- | -------- | -------- | -------- | --------- |
+| **GLIP-T(C)** | ✔ | O365, GoldG | 84.8 | 94.9 | 96.3 | 85.5 | 95.4 | 96.6 |
+| **GLIP-T(C)** | | O365, GoldG | 84.9 | 94.9 | 96.3 | 85.6 | 95.4 | 96.7 |
+| **GLIP-T** | | O365,GoldG,CC3M,SBU | 85.3 | 95.5 | 96.9 | 86.0 | 95.9 | 97.2 |
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/flickr30k/glip_atss_swin-t_c_fpn_dyhead_pretrain_obj365-goldg_zeroshot_flickr30k.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/flickr30k/glip_atss_swin-t_c_fpn_dyhead_pretrain_obj365-goldg_zeroshot_flickr30k.py
new file mode 100644
index 0000000000000000000000000000000000000000..14d6e8aaa6372a5272467dd46d33e80979298efc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/flickr30k/glip_atss_swin-t_c_fpn_dyhead_pretrain_obj365-goldg_zeroshot_flickr30k.py
@@ -0,0 +1,61 @@
+_base_ = '../glip_atss_swin-t_a_fpn_dyhead_pretrain_obj365.py'
+
+lang_model_name = 'bert-base-uncased'
+
+model = dict(bbox_head=dict(early_fuse=True))
+
+dataset_type = 'Flickr30kDataset'
+data_root = 'data/flickr30k_entities/'
+
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile', backend_args=None,
+ imdecode_backend='pillow'),
+ dict(
+ type='FixScaleResize',
+ scale=(800, 1333),
+ keep_ratio=True,
+ backend='pillow'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'text', 'custom_entities',
+ 'tokens_positive', 'phrase_ids', 'phrases'))
+]
+
+dataset_Flickr30k_val = dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='final_flickr_separateGT_val.json',
+ data_prefix=dict(img='flickr30k_images/'),
+ pipeline=test_pipeline,
+)
+
+dataset_Flickr30k_test = dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='final_flickr_separateGT_test.json',
+ data_prefix=dict(img='flickr30k_images/'),
+ pipeline=test_pipeline,
+)
+
+val_evaluator_Flickr30k = dict(type='Flickr30kMetric', )
+
+test_evaluator_Flickr30k = dict(type='Flickr30kMetric', )
+
+# ----------Config---------- #
+dataset_prefixes = ['Flickr30kVal', 'Flickr30kTest']
+datasets = [dataset_Flickr30k_val, dataset_Flickr30k_test]
+metrics = [val_evaluator_Flickr30k, test_evaluator_Flickr30k]
+
+val_dataloader = dict(
+ dataset=dict(_delete_=True, type='ConcatDataset', datasets=datasets))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='MultiDatasetsEvaluator',
+ metrics=metrics,
+ dataset_prefixes=dataset_prefixes)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-l_fpn_dyhead_16xb2_ms-2x_funtune_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-l_fpn_dyhead_16xb2_ms-2x_funtune_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..92a85a11d57b6d3d64bfed5f9a691bca739d7ce3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-l_fpn_dyhead_16xb2_ms-2x_funtune_coco.py
@@ -0,0 +1,14 @@
+_base_ = './glip_atss_swin-t_b_fpn_dyhead_16xb2_ms-2x_funtune_coco.py'
+
+model = dict(
+ backbone=dict(
+ embed_dims=192,
+ depths=[2, 2, 18, 2],
+ num_heads=[6, 12, 24, 48],
+ window_size=12,
+ drop_path_rate=0.4,
+ ),
+ neck=dict(in_channels=[384, 768, 1536]),
+ bbox_head=dict(early_fuse=True, num_dyhead_blocks=8, use_checkpoint=True))
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/glip/glip_l_mmdet-abfe026b.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-l_fpn_dyhead_pretrain_mixeddata.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-l_fpn_dyhead_pretrain_mixeddata.py
new file mode 100644
index 0000000000000000000000000000000000000000..546ecfe1d513b4161322f5ffa0e51d01b2775780
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-l_fpn_dyhead_pretrain_mixeddata.py
@@ -0,0 +1,12 @@
+_base_ = './glip_atss_swin-t_a_fpn_dyhead_pretrain_obj365.py'
+
+model = dict(
+ backbone=dict(
+ embed_dims=192,
+ depths=[2, 2, 18, 2],
+ num_heads=[6, 12, 24, 48],
+ window_size=12,
+ drop_path_rate=0.4,
+ ),
+ neck=dict(in_channels=[384, 768, 1536]),
+ bbox_head=dict(early_fuse=True, num_dyhead_blocks=8))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_a_fpn_dyhead_16xb2_ms-2x_funtune_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_a_fpn_dyhead_16xb2_ms-2x_funtune_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..4b280657b315c77dd118ab84880d97dc882102a1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_a_fpn_dyhead_16xb2_ms-2x_funtune_coco.py
@@ -0,0 +1,155 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_a_mmdet-b3654169.pth' # noqa
+lang_model_name = 'bert-base-uncased'
+
+model = dict(
+ type='GLIP',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.53, 116.28, 123.675],
+ std=[57.375, 57.12, 58.395],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='SwinTransformer',
+ embed_dims=96,
+ depths=[2, 2, 6, 2],
+ num_heads=[3, 6, 12, 24],
+ window_size=7,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.2,
+ patch_norm=True,
+ out_indices=(1, 2, 3),
+ with_cp=False,
+ convert_weights=False),
+ neck=dict(
+ type='FPN_DropBlock',
+ in_channels=[192, 384, 768],
+ out_channels=256,
+ start_level=0,
+ relu_before_extra_convs=True,
+ add_extra_convs='on_output',
+ num_outs=5),
+ bbox_head=dict(
+ type='ATSSVLFusionHead',
+ lang_model_name=lang_model_name,
+ num_classes=80,
+ in_channels=256,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ octave_base_scale=8,
+ scales_per_octave=1,
+ strides=[8, 16, 32, 64, 128],
+ center_offset=0.5),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoderForGLIP',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=2.0),
+ loss_centerness=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0)),
+ language_model=dict(type='BertModel', name=lang_model_name),
+ train_cfg=dict(
+ assigner=dict(
+ type='ATSSAssigner',
+ topk=9,
+ iou_calculator=dict(type='BboxOverlaps2D_GLIP')),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+
+# dataset settings
+train_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ imdecode_backend='pillow',
+ backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='GTBoxSubOne_GLIP'),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 480), (1333, 560), (1333, 640), (1333, 720),
+ (1333, 800)],
+ keep_ratio=True,
+ resize_type='FixScaleResize',
+ backend='pillow'),
+ dict(type='RandomFlip_GLIP', prob=0.5),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1, 1)),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities'))
+]
+
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ backend_args=_base_.backend_args,
+ imdecode_backend='pillow'),
+ dict(
+ type='FixScaleResize',
+ scale=(800, 1333),
+ keep_ratio=True,
+ backend='pillow'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'text', 'custom_entities'))
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ _delete_=True,
+ type='RepeatDataset',
+ times=2,
+ dataset=dict(
+ type=_base_.dataset_type,
+ data_root=_base_.data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ return_classes=True,
+ backend_args=_base_.backend_args)))
+
+val_dataloader = dict(
+ dataset=dict(pipeline=test_pipeline, return_classes=True))
+test_dataloader = val_dataloader
+
+# We did not adopt the official 24e optimizer strategy
+# because the results indicate that the current strategy is superior.
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(
+ type='AdamW', lr=0.00002, betas=(0.9, 0.999), weight_decay=0.05),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'relative_position_bias_table': dict(decay_mult=0.),
+ 'norm': dict(decay_mult=0.)
+ }),
+ clip_grad=None)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_a_fpn_dyhead_pretrain_obj365.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_a_fpn_dyhead_pretrain_obj365.py
new file mode 100644
index 0000000000000000000000000000000000000000..34a818caefcbfcdd9e51ec304fb94906c20ceb9a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_a_fpn_dyhead_pretrain_obj365.py
@@ -0,0 +1,90 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+lang_model_name = 'bert-base-uncased'
+
+model = dict(
+ type='GLIP',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.53, 116.28, 123.675],
+ std=[57.375, 57.12, 58.395],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='SwinTransformer',
+ embed_dims=96,
+ depths=[2, 2, 6, 2],
+ num_heads=[3, 6, 12, 24],
+ window_size=7,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.2,
+ patch_norm=True,
+ out_indices=(1, 2, 3),
+ with_cp=False,
+ convert_weights=False),
+ neck=dict(
+ type='FPN',
+ in_channels=[192, 384, 768],
+ out_channels=256,
+ start_level=0,
+ relu_before_extra_convs=True,
+ add_extra_convs='on_output',
+ num_outs=5),
+ bbox_head=dict(
+ type='ATSSVLFusionHead',
+ lang_model_name=lang_model_name,
+ num_classes=80,
+ in_channels=256,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ octave_base_scale=8,
+ scales_per_octave=1,
+ strides=[8, 16, 32, 64, 128],
+ center_offset=0.5),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoderForGLIP',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ ),
+ language_model=dict(type='BertModel', name=lang_model_name),
+ train_cfg=dict(
+ assigner=dict(type='ATSSAssigner', topk=9),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ backend_args=_base_.backend_args,
+ imdecode_backend='pillow'),
+ dict(
+ type='FixScaleResize',
+ scale=(800, 1333),
+ keep_ratio=True,
+ backend='pillow'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'text', 'custom_entities'))
+]
+
+val_dataloader = dict(
+ dataset=dict(pipeline=test_pipeline, return_classes=True))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_b_fpn_dyhead_16xb2_ms-2x_funtune_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_b_fpn_dyhead_16xb2_ms-2x_funtune_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3487de3f3a24077f475e8451722d1b4d252a0084
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_b_fpn_dyhead_16xb2_ms-2x_funtune_coco.py
@@ -0,0 +1,9 @@
+_base_ = './glip_atss_swin-t_a_fpn_dyhead_16xb2_ms-2x_funtune_coco.py'
+
+model = dict(bbox_head=dict(early_fuse=True, use_checkpoint=True))
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_b_mmdet-6dfbd102.pth' # noqa
+
+optim_wrapper = dict(
+ optimizer=dict(lr=0.00001),
+ clip_grad=dict(_delete_=True, max_norm=1, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_b_fpn_dyhead_pretrain_obj365.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_b_fpn_dyhead_pretrain_obj365.py
new file mode 100644
index 0000000000000000000000000000000000000000..6334e5e3b4043a81d154fc03a94594d93d74aed5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_b_fpn_dyhead_pretrain_obj365.py
@@ -0,0 +1,3 @@
+_base_ = './glip_atss_swin-t_a_fpn_dyhead_pretrain_obj365.py'
+
+model = dict(bbox_head=dict(early_fuse=True))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_c_fpn_dyhead_16xb2_ms-2x_funtune_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_c_fpn_dyhead_16xb2_ms-2x_funtune_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5c315e490e7a7e05a6334d4d38ce9be9b70851b3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_c_fpn_dyhead_16xb2_ms-2x_funtune_coco.py
@@ -0,0 +1,3 @@
+_base_ = './glip_atss_swin-t_b_fpn_dyhead_16xb2_ms-2x_funtune_coco.py'
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_c_mmdet-2fc427dd.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_c_fpn_dyhead_pretrain_obj365-goldg.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_c_fpn_dyhead_pretrain_obj365-goldg.py
new file mode 100644
index 0000000000000000000000000000000000000000..24898f4df532cc2e2728265800d2f6a030e8efe0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_c_fpn_dyhead_pretrain_obj365-goldg.py
@@ -0,0 +1 @@
+_base_ = './glip_atss_swin-t_b_fpn_dyhead_pretrain_obj365.py'
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_fpn_dyhead_16xb2_ms-2x_funtune_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_fpn_dyhead_16xb2_ms-2x_funtune_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3391272e608e8098773a6435550e578f462ed886
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_fpn_dyhead_16xb2_ms-2x_funtune_coco.py
@@ -0,0 +1,3 @@
+_base_ = './glip_atss_swin-t_b_fpn_dyhead_16xb2_ms-2x_funtune_coco.py'
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_mmdet-c24ce662.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_fpn_dyhead_pretrain_obj365-goldg-cc3m-sub.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_fpn_dyhead_pretrain_obj365-goldg-cc3m-sub.py
new file mode 100644
index 0000000000000000000000000000000000000000..24898f4df532cc2e2728265800d2f6a030e8efe0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/glip_atss_swin-t_fpn_dyhead_pretrain_obj365-goldg-cc3m-sub.py
@@ -0,0 +1 @@
+_base_ = './glip_atss_swin-t_b_fpn_dyhead_pretrain_obj365.py'
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/lvis/glip_atss_swin-l_fpn_dyhead_pretrain_zeroshot_lvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/lvis/glip_atss_swin-l_fpn_dyhead_pretrain_zeroshot_lvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..1f79e447d3f24e364739740be504bb234adc1e98
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/lvis/glip_atss_swin-l_fpn_dyhead_pretrain_zeroshot_lvis.py
@@ -0,0 +1,12 @@
+_base_ = './glip_atss_swin-t_a_fpn_dyhead_pretrain_zeroshot_lvis.py'
+
+model = dict(
+ backbone=dict(
+ embed_dims=192,
+ depths=[2, 2, 18, 2],
+ num_heads=[6, 12, 24, 48],
+ window_size=12,
+ drop_path_rate=0.4,
+ ),
+ neck=dict(in_channels=[384, 768, 1536]),
+ bbox_head=dict(early_fuse=True, num_dyhead_blocks=8))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/lvis/glip_atss_swin-l_fpn_dyhead_pretrain_zeroshot_mini-lvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/lvis/glip_atss_swin-l_fpn_dyhead_pretrain_zeroshot_mini-lvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..13f1a69082b670632dfe3eb8dc50826549dcf59f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/lvis/glip_atss_swin-l_fpn_dyhead_pretrain_zeroshot_mini-lvis.py
@@ -0,0 +1,12 @@
+_base_ = './glip_atss_swin-t_a_fpn_dyhead_pretrain_zeroshot_mini-lvis.py'
+
+model = dict(
+ backbone=dict(
+ embed_dims=192,
+ depths=[2, 2, 18, 2],
+ num_heads=[6, 12, 24, 48],
+ window_size=12,
+ drop_path_rate=0.4,
+ ),
+ neck=dict(in_channels=[384, 768, 1536]),
+ bbox_head=dict(early_fuse=True, num_dyhead_blocks=8))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/lvis/glip_atss_swin-t_a_fpn_dyhead_pretrain_zeroshot_lvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/lvis/glip_atss_swin-t_a_fpn_dyhead_pretrain_zeroshot_lvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..4d526d59008b39996a147a2852a44d2e936113d2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/lvis/glip_atss_swin-t_a_fpn_dyhead_pretrain_zeroshot_lvis.py
@@ -0,0 +1,24 @@
+_base_ = '../glip_atss_swin-t_a_fpn_dyhead_pretrain_obj365.py'
+
+model = dict(test_cfg=dict(
+ max_per_img=300,
+ chunked_size=40,
+))
+
+dataset_type = 'LVISV1Dataset'
+data_root = 'data/coco/'
+
+val_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ type=dataset_type,
+ ann_file='annotations/lvis_od_val.json',
+ data_prefix=dict(img='')))
+test_dataloader = val_dataloader
+
+# numpy < 1.24.0
+val_evaluator = dict(
+ _delete_=True,
+ type='LVISFixedAPMetric',
+ ann_file=data_root + 'annotations/lvis_od_val.json')
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/lvis/glip_atss_swin-t_a_fpn_dyhead_pretrain_zeroshot_mini-lvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/lvis/glip_atss_swin-t_a_fpn_dyhead_pretrain_zeroshot_mini-lvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..70a57a3f581ca1c374dbae71059c7049a20d3a47
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/lvis/glip_atss_swin-t_a_fpn_dyhead_pretrain_zeroshot_mini-lvis.py
@@ -0,0 +1,25 @@
+_base_ = '../glip_atss_swin-t_a_fpn_dyhead_pretrain_obj365.py'
+
+model = dict(test_cfg=dict(
+ max_per_img=300,
+ chunked_size=40,
+))
+
+dataset_type = 'LVISV1Dataset'
+data_root = 'data/coco/'
+
+val_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ type=dataset_type,
+ ann_file='annotations/lvis_v1_minival_inserted_image_name.json',
+ data_prefix=dict(img='')))
+test_dataloader = val_dataloader
+
+# numpy < 1.24.0
+val_evaluator = dict(
+ _delete_=True,
+ type='LVISFixedAPMetric',
+ ann_file=data_root +
+ 'annotations/lvis_v1_minival_inserted_image_name.json')
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/lvis/glip_atss_swin-t_bc_fpn_dyhead_pretrain_zeroshot_lvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/lvis/glip_atss_swin-t_bc_fpn_dyhead_pretrain_zeroshot_lvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..6dc712b3bcb4f8dd1018b175d3a4e7f59be3a990
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/lvis/glip_atss_swin-t_bc_fpn_dyhead_pretrain_zeroshot_lvis.py
@@ -0,0 +1,3 @@
+_base_ = './glip_atss_swin-t_a_fpn_dyhead_pretrain_zeroshot_lvis.py'
+
+model = dict(bbox_head=dict(early_fuse=True))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/lvis/glip_atss_swin-t_bc_fpn_dyhead_pretrain_zeroshot_mini-lvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/lvis/glip_atss_swin-t_bc_fpn_dyhead_pretrain_zeroshot_mini-lvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..3babb91101a6dc283ada78911672c7c7433f67ac
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/lvis/glip_atss_swin-t_bc_fpn_dyhead_pretrain_zeroshot_mini-lvis.py
@@ -0,0 +1,3 @@
+_base_ = './glip_atss_swin-t_a_fpn_dyhead_pretrain_zeroshot_mini-lvis.py'
+
+model = dict(bbox_head=dict(early_fuse=True))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..fbbf718b9fff3061a4e02a7d39a6c95252beb603
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/metafile.yml
@@ -0,0 +1,111 @@
+Collections:
+ - Name: GLIP
+ Metadata:
+ Training Data: Objects365, GoldG, CC3M, SBU and COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: A100 GPUs
+ Architecture:
+ - Swin Transformer
+ - DYHead
+ - BERT
+ Paper:
+ URL: https://arxiv.org/abs/2112.03857
+ Title: 'GLIP: Grounded Language-Image Pre-training'
+ README: configs/glip/README.md
+ Code:
+ URL:
+ Version: v3.0.0
+
+Models:
+ - Name: glip_atss_swin-t_a_fpn_dyhead_pretrain_obj365
+ In Collection: GLIP
+ Config: configs/glip/glip_atss_swin-t_a_fpn_dyhead_pretrain_obj365.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.0
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_a_mmdet-b3654169.pth
+ - Name: glip_atss_swin-t_b_fpn_dyhead_pretrain_obj365
+ In Collection: GLIP
+ Config: configs/glip/glip_atss_swin-t_b_fpn_dyhead_pretrain_obj365.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.9
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_b_mmdet-6dfbd102.pth
+ - Name: glip_atss_swin-t_c_fpn_dyhead_pretrain_obj365-goldg
+ In Collection: GLIP
+ Config: configs/glip/glip_atss_swin-t_c_fpn_dyhead_pretrain_obj365-goldg.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.7
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_c_mmdet-2fc427dd.pth
+ - Name: glip_atss_swin-t_fpn_dyhead_pretrain_obj365-goldg-cc3m-sub
+ In Collection: GLIP
+ Config: configs/glip/glip_atss_swin-t_fpn_dyhead_pretrain_obj365-goldg-cc3m-sub.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.4
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/glip/glip_tiny_mmdet-c24ce662.pth
+ - Name: glip_atss_swin-l_fpn_dyhead_pretrain_mixeddata
+ In Collection: GLIP
+ Config: configs/glip/glip_atss_swin-l_fpn_dyhead_pretrain_mixeddata.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 51.3
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/glip/glip_l_mmdet-abfe026b.pth
+ - Name: glip_atss_swin-t_a_fpn_dyhead_16xb2_ms-2x_funtune_coco
+ In Collection: GLIP
+ Config: configs/glip/glip_atss_swin-t_a_fpn_dyhead_16xb2_ms-2x_funtune_coco.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 53.3
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/glip/glip_atss_swin-t_a_fpn_dyhead_16xb2_ms-2x_funtune_coco/glip_atss_swin-t_a_fpn_dyhead_16xb2_ms-2x_funtune_coco_20230914_180419-e6addd96.pth
+ - Name: glip_atss_swin-t_b_fpn_dyhead_16xb2_ms-2x_funtune_coco
+ In Collection: GLIP
+ Config: configs/glip/glip_atss_swin-t_b_fpn_dyhead_16xb2_ms-2x_funtune_coco.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 54.1
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/glip/glip_atss_swin-t_b_fpn_dyhead_16xb2_ms-2x_funtune_coco/glip_atss_swin-t_b_fpn_dyhead_16xb2_ms-2x_funtune_coco_20230916_163538-650323ba.pth
+ - Name: glip_atss_swin-t_c_fpn_dyhead_16xb2_ms-2x_funtune_coco
+ In Collection: GLIP
+ Config: configs/glip/glip_atss_swin-t_c_fpn_dyhead_16xb2_ms-2x_funtune_coco.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 55.2
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/glip/glip_atss_swin-t_c_fpn_dyhead_16xb2_ms-2x_funtune_coco/glip_atss_swin-t_c_fpn_dyhead_16xb2_ms-2x_funtune_coco_20230914_182935-4ba3fc3b.pth
+ - Name: glip_atss_swin-t_fpn_dyhead_16xb2_ms-2x_funtune_coco
+ In Collection: GLIP
+ Config: configs/glip/glip_atss_swin-t_fpn_dyhead_16xb2_ms-2x_funtune_coco.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 55.4
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/glip/glip_atss_swin-t_fpn_dyhead_16xb2_ms-2x_funtune_coco/glip_atss_swin-t_fpn_dyhead_16xb2_ms-2x_funtune_coco_20230914_224410-ba97be24.pth
+ - Name: glip_atss_swin-l_fpn_dyhead_16xb2_ms-2x_funtune_coco
+ In Collection: GLIP
+ Config: configs/glip/glip_atss_swin-l_fpn_dyhead_16xb2_ms-2x_funtune_coco.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 59.4
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/glip/glip_atss_swin-l_fpn_dyhead_16xb2_ms-2x_funtune_coco/glip_atss_swin-l_fpn_dyhead_16xb2_ms-2x_funtune_coco_20230910_100800-e9be4274.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/odinw/glip_atss_swin-t_a_fpn_dyhead_pretrain_odinw13.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/odinw/glip_atss_swin-t_a_fpn_dyhead_pretrain_odinw13.py
new file mode 100644
index 0000000000000000000000000000000000000000..d38effba8c1333a2403c6bc0f20b7fde21c4c47d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/odinw/glip_atss_swin-t_a_fpn_dyhead_pretrain_odinw13.py
@@ -0,0 +1,338 @@
+_base_ = '../glip_atss_swin-t_a_fpn_dyhead_pretrain_obj365.py'
+
+dataset_type = 'CocoDataset'
+data_root = 'data/odinw/'
+
+base_test_pipeline = _base_.test_pipeline
+base_test_pipeline[-1]['meta_keys'] = ('img_id', 'img_path', 'ori_shape',
+ 'img_shape', 'scale_factor', 'text',
+ 'custom_entities', 'caption_prompt')
+
+# ---------------------1 AerialMaritimeDrone---------------------#
+class_name = ('boat', 'car', 'dock', 'jetski', 'lift')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'AerialMaritimeDrone/large/'
+dataset_AerialMaritimeDrone = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ test_mode=True,
+ pipeline=base_test_pipeline,
+ return_classes=True)
+val_evaluator_AerialMaritimeDrone = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------2 Aquarium---------------------#
+class_name = ('fish', 'jellyfish', 'penguin', 'puffin', 'shark', 'starfish',
+ 'stingray')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Aquarium/Aquarium Combined.v2-raw-1024.coco/'
+
+caption_prompt = None
+# caption_prompt = {
+# 'penguin': {
+# 'suffix': ', which is black and white'
+# },
+# 'puffin': {
+# 'suffix': ' with orange beaks'
+# },
+# 'stingray': {
+# 'suffix': ' which is flat and round'
+# },
+# }
+dataset_Aquarium = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Aquarium = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------3 CottontailRabbits---------------------#
+class_name = ('Cottontail-Rabbit', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'CottontailRabbits/'
+
+caption_prompt = None
+# caption_prompt = {'Cottontail-Rabbit': {'name': 'rabbit'}}
+
+dataset_CottontailRabbits = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_CottontailRabbits = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------4 EgoHands---------------------#
+class_name = ('hand', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'EgoHands/generic/'
+
+caption_prompt = None
+# caption_prompt = {'hand': {'suffix': ' of a person'}}
+
+dataset_EgoHands = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_EgoHands = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------5 NorthAmericaMushrooms---------------------#
+class_name = ('CoW', 'chanterelle')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'NorthAmericaMushrooms/North American Mushrooms.v1-416x416.coco/' # noqa
+
+caption_prompt = None
+# caption_prompt = {
+# 'CoW': {
+# 'name': 'flat mushroom'
+# },
+# 'chanterelle': {
+# 'name': 'yellow mushroom'
+# }
+# }
+
+dataset_NorthAmericaMushrooms = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_NorthAmericaMushrooms = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------6 Packages---------------------#
+class_name = ('package', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Packages/Raw/'
+
+caption_prompt = None
+# caption_prompt = {
+# 'package': {
+# 'prefix': 'there is a ',
+# 'suffix': ' on the porch'
+# }
+# }
+
+dataset_Packages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Packages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------7 PascalVOC---------------------#
+class_name = ('aeroplane', 'bicycle', 'bird', 'boat', 'bottle', 'bus', 'car',
+ 'cat', 'chair', 'cow', 'diningtable', 'dog', 'horse',
+ 'motorbike', 'person', 'pottedplant', 'sheep', 'sofa', 'train',
+ 'tvmonitor')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'PascalVOC/'
+dataset_PascalVOC = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_PascalVOC = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------8 pistols---------------------#
+class_name = ('pistol', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'pistols/export/'
+dataset_pistols = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_pistols = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------9 pothole---------------------#
+class_name = ('pothole', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'pothole/'
+
+caption_prompt = None
+# caption_prompt = {
+# 'pothole': {
+# 'prefix': 'there are some ',
+# 'name': 'holes',
+# 'suffix': ' on the road'
+# }
+# }
+
+dataset_pothole = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_pothole = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------10 Raccoon---------------------#
+class_name = ('raccoon', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Raccoon/Raccoon.v2-raw.coco/'
+dataset_Raccoon = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Raccoon = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------11 ShellfishOpenImages---------------------#
+class_name = ('Crab', 'Lobster', 'Shrimp')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'ShellfishOpenImages/raw/'
+dataset_ShellfishOpenImages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_ShellfishOpenImages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------12 thermalDogsAndPeople---------------------#
+class_name = ('dog', 'person')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'thermalDogsAndPeople/'
+dataset_thermalDogsAndPeople = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_thermalDogsAndPeople = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------13 VehiclesOpenImages---------------------#
+class_name = ('Ambulance', 'Bus', 'Car', 'Motorcycle', 'Truck')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'VehiclesOpenImages/416x416/'
+dataset_VehiclesOpenImages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_VehiclesOpenImages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# --------------------- Config---------------------#
+dataset_prefixes = [
+ 'AerialMaritimeDrone', 'Aquarium', 'CottontailRabbits', 'EgoHands',
+ 'NorthAmericaMushrooms', 'Packages', 'PascalVOC', 'pistols', 'pothole',
+ 'Raccoon', 'ShellfishOpenImages', 'thermalDogsAndPeople',
+ 'VehiclesOpenImages'
+]
+datasets = [
+ dataset_AerialMaritimeDrone, dataset_Aquarium, dataset_CottontailRabbits,
+ dataset_EgoHands, dataset_NorthAmericaMushrooms, dataset_Packages,
+ dataset_PascalVOC, dataset_pistols, dataset_pothole, dataset_Raccoon,
+ dataset_ShellfishOpenImages, dataset_thermalDogsAndPeople,
+ dataset_VehiclesOpenImages
+]
+metrics = [
+ val_evaluator_AerialMaritimeDrone, val_evaluator_Aquarium,
+ val_evaluator_CottontailRabbits, val_evaluator_EgoHands,
+ val_evaluator_NorthAmericaMushrooms, val_evaluator_Packages,
+ val_evaluator_PascalVOC, val_evaluator_pistols, val_evaluator_pothole,
+ val_evaluator_Raccoon, val_evaluator_ShellfishOpenImages,
+ val_evaluator_thermalDogsAndPeople, val_evaluator_VehiclesOpenImages
+]
+
+# -------------------------------------------------#
+val_dataloader = dict(
+ dataset=dict(_delete_=True, type='ConcatDataset', datasets=datasets))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='MultiDatasetsEvaluator',
+ metrics=metrics,
+ dataset_prefixes=dataset_prefixes)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/odinw/glip_atss_swin-t_a_fpn_dyhead_pretrain_odinw35.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/odinw/glip_atss_swin-t_a_fpn_dyhead_pretrain_odinw35.py
new file mode 100644
index 0000000000000000000000000000000000000000..2eaf09ed771978397b9d67048b371724418e50aa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/odinw/glip_atss_swin-t_a_fpn_dyhead_pretrain_odinw35.py
@@ -0,0 +1,794 @@
+_base_ = '../glip_atss_swin-t_a_fpn_dyhead_pretrain_obj365.py'
+
+dataset_type = 'CocoDataset'
+data_root = 'data/odinw/'
+
+base_test_pipeline = _base_.test_pipeline
+base_test_pipeline[-1]['meta_keys'] = ('img_id', 'img_path', 'ori_shape',
+ 'img_shape', 'scale_factor', 'text',
+ 'custom_entities', 'caption_prompt')
+
+# ---------------------1 AerialMaritimeDrone_large---------------------#
+class_name = ('boat', 'car', 'dock', 'jetski', 'lift')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'AerialMaritimeDrone/large/'
+dataset_AerialMaritimeDrone_large = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_AerialMaritimeDrone_large = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------2 AerialMaritimeDrone_tiled---------------------#
+class_name = ('boat', 'car', 'dock', 'jetski', 'lift')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'AerialMaritimeDrone/tiled/'
+dataset_AerialMaritimeDrone_tiled = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_AerialMaritimeDrone_tiled = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------3 AmericanSignLanguageLetters---------------------#
+class_name = ('A', 'B', 'C', 'D', 'E', 'F', 'G', 'H', 'I', 'J', 'K', 'L', 'M',
+ 'N', 'O', 'P', 'Q', 'R', 'S', 'T', 'U', 'V', 'W', 'X', 'Y', 'Z')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'AmericanSignLanguageLetters/American Sign Language Letters.v1-v1.coco/' # noqa
+dataset_AmericanSignLanguageLetters = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_AmericanSignLanguageLetters = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------4 Aquarium---------------------#
+class_name = ('fish', 'jellyfish', 'penguin', 'puffin', 'shark', 'starfish',
+ 'stingray')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Aquarium/Aquarium Combined.v2-raw-1024.coco/'
+dataset_Aquarium = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Aquarium = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------5 BCCD---------------------#
+class_name = ('Platelets', 'RBC', 'WBC')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'BCCD/BCCD.v3-raw.coco/'
+dataset_BCCD = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_BCCD = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------6 boggleBoards---------------------#
+class_name = ('Q', 'a', 'an', 'b', 'c', 'd', 'e', 'er', 'f', 'g', 'h', 'he',
+ 'i', 'in', 'j', 'k', 'l', 'm', 'n', 'o', 'o ', 'p', 'q', 'qu',
+ 'r', 's', 't', 't\\', 'th', 'u', 'v', 'w', 'wild', 'x', 'y', 'z')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'boggleBoards/416x416AutoOrient/export/'
+dataset_boggleBoards = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_boggleBoards = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------7 brackishUnderwater---------------------#
+class_name = ('crab', 'fish', 'jellyfish', 'shrimp', 'small_fish', 'starfish')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'brackishUnderwater/960x540/'
+dataset_brackishUnderwater = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_brackishUnderwater = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------8 ChessPieces---------------------#
+class_name = (' ', 'black bishop', 'black king', 'black knight', 'black pawn',
+ 'black queen', 'black rook', 'white bishop', 'white king',
+ 'white knight', 'white pawn', 'white queen', 'white rook')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'ChessPieces/Chess Pieces.v23-raw.coco/'
+dataset_ChessPieces = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/new_annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_ChessPieces = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/new_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------9 CottontailRabbits---------------------#
+class_name = ('rabbit', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'CottontailRabbits/'
+dataset_CottontailRabbits = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/new_annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_CottontailRabbits = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/new_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------10 dice---------------------#
+class_name = ('1', '2', '3', '4', '5', '6')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'dice/mediumColor/export/'
+dataset_dice = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_dice = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------11 DroneControl---------------------#
+class_name = ('follow', 'follow_hand', 'land', 'land_hand', 'null', 'object',
+ 'takeoff', 'takeoff-hand')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'DroneControl/Drone Control.v3-raw.coco/'
+dataset_DroneControl = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_DroneControl = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------12 EgoHands_generic---------------------#
+class_name = ('hand', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'EgoHands/generic/'
+caption_prompt = {'hand': {'suffix': ' of a person'}}
+dataset_EgoHands_generic = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_EgoHands_generic = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------13 EgoHands_specific---------------------#
+class_name = ('myleft', 'myright', 'yourleft', 'yourright')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'EgoHands/specific/'
+dataset_EgoHands_specific = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_EgoHands_specific = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------14 HardHatWorkers---------------------#
+class_name = ('head', 'helmet', 'person')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'HardHatWorkers/raw/'
+dataset_HardHatWorkers = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_HardHatWorkers = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------15 MaskWearing---------------------#
+class_name = ('mask', 'no-mask')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'MaskWearing/raw/'
+dataset_MaskWearing = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_MaskWearing = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------16 MountainDewCommercial---------------------#
+class_name = ('bottle', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'MountainDewCommercial/'
+dataset_MountainDewCommercial = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_MountainDewCommercial = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------17 NorthAmericaMushrooms---------------------#
+class_name = ('flat mushroom', 'yellow mushroom')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'NorthAmericaMushrooms/North American Mushrooms.v1-416x416.coco/' # noqa
+dataset_NorthAmericaMushrooms = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/new_annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_NorthAmericaMushrooms = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/new_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------18 openPoetryVision---------------------#
+class_name = ('American Typewriter', 'Andale Mono', 'Apple Chancery', 'Arial',
+ 'Avenir', 'Baskerville', 'Big Caslon', 'Bradley Hand',
+ 'Brush Script MT', 'Chalkboard', 'Comic Sans MS', 'Copperplate',
+ 'Courier', 'Didot', 'Futura', 'Geneva', 'Georgia', 'Gill Sans',
+ 'Helvetica', 'Herculanum', 'Impact', 'Kefa', 'Lucida Grande',
+ 'Luminari', 'Marker Felt', 'Menlo', 'Monaco', 'Noteworthy',
+ 'Optima', 'PT Sans', 'PT Serif', 'Palatino', 'Papyrus',
+ 'Phosphate', 'Rockwell', 'SF Pro', 'SignPainter', 'Skia',
+ 'Snell Roundhand', 'Tahoma', 'Times New Roman', 'Trebuchet MS',
+ 'Verdana')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'openPoetryVision/512x512/'
+dataset_openPoetryVision = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_openPoetryVision = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------19 OxfordPets_by_breed---------------------#
+class_name = ('cat-Abyssinian', 'cat-Bengal', 'cat-Birman', 'cat-Bombay',
+ 'cat-British_Shorthair', 'cat-Egyptian_Mau', 'cat-Maine_Coon',
+ 'cat-Persian', 'cat-Ragdoll', 'cat-Russian_Blue', 'cat-Siamese',
+ 'cat-Sphynx', 'dog-american_bulldog',
+ 'dog-american_pit_bull_terrier', 'dog-basset_hound',
+ 'dog-beagle', 'dog-boxer', 'dog-chihuahua',
+ 'dog-english_cocker_spaniel', 'dog-english_setter',
+ 'dog-german_shorthaired', 'dog-great_pyrenees', 'dog-havanese',
+ 'dog-japanese_chin', 'dog-keeshond', 'dog-leonberger',
+ 'dog-miniature_pinscher', 'dog-newfoundland', 'dog-pomeranian',
+ 'dog-pug', 'dog-saint_bernard', 'dog-samoyed',
+ 'dog-scottish_terrier', 'dog-shiba_inu',
+ 'dog-staffordshire_bull_terrier', 'dog-wheaten_terrier',
+ 'dog-yorkshire_terrier')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'OxfordPets/by-breed/' # noqa
+dataset_OxfordPets_by_breed = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_OxfordPets_by_breed = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------20 OxfordPets_by_species---------------------#
+class_name = ('cat', 'dog')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'OxfordPets/by-species/' # noqa
+dataset_OxfordPets_by_species = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_OxfordPets_by_species = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------21 PKLot---------------------#
+class_name = ('space-empty', 'space-occupied')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'PKLot/640/' # noqa
+dataset_PKLot = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_PKLot = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------22 Packages---------------------#
+class_name = ('package', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Packages/Raw/'
+caption_prompt = {
+ 'package': {
+ 'prefix': 'there is a ',
+ 'suffix': ' on the porch'
+ }
+}
+dataset_Packages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Packages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------23 PascalVOC---------------------#
+class_name = ('aeroplane', 'bicycle', 'bird', 'boat', 'bottle', 'bus', 'car',
+ 'cat', 'chair', 'cow', 'diningtable', 'dog', 'horse',
+ 'motorbike', 'person', 'pottedplant', 'sheep', 'sofa', 'train',
+ 'tvmonitor')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'PascalVOC/'
+dataset_PascalVOC = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_PascalVOC = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------24 pistols---------------------#
+class_name = ('pistol', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'pistols/export/'
+dataset_pistols = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_pistols = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------25 plantdoc---------------------#
+class_name = ('Apple Scab Leaf', 'Apple leaf', 'Apple rust leaf',
+ 'Bell_pepper leaf', 'Bell_pepper leaf spot', 'Blueberry leaf',
+ 'Cherry leaf', 'Corn Gray leaf spot', 'Corn leaf blight',
+ 'Corn rust leaf', 'Peach leaf', 'Potato leaf',
+ 'Potato leaf early blight', 'Potato leaf late blight',
+ 'Raspberry leaf', 'Soyabean leaf', 'Soybean leaf',
+ 'Squash Powdery mildew leaf', 'Strawberry leaf',
+ 'Tomato Early blight leaf', 'Tomato Septoria leaf spot',
+ 'Tomato leaf', 'Tomato leaf bacterial spot',
+ 'Tomato leaf late blight', 'Tomato leaf mosaic virus',
+ 'Tomato leaf yellow virus', 'Tomato mold leaf',
+ 'Tomato two spotted spider mites leaf', 'grape leaf',
+ 'grape leaf black rot')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'plantdoc/416x416/'
+dataset_plantdoc = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_plantdoc = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------26 pothole---------------------#
+class_name = ('pothole', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'pothole/'
+caption_prompt = {
+ 'pothole': {
+ 'name': 'holes',
+ 'prefix': 'there are some ',
+ 'suffix': ' on the road'
+ }
+}
+dataset_pothole = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ caption_prompt=caption_prompt,
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_pothole = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------27 Raccoon---------------------#
+class_name = ('raccoon', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Raccoon/Raccoon.v2-raw.coco/'
+dataset_Raccoon = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Raccoon = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------28 selfdrivingCar---------------------#
+class_name = ('biker', 'car', 'pedestrian', 'trafficLight',
+ 'trafficLight-Green', 'trafficLight-GreenLeft',
+ 'trafficLight-Red', 'trafficLight-RedLeft',
+ 'trafficLight-Yellow', 'trafficLight-YellowLeft', 'truck')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'selfdrivingCar/fixedLarge/export/'
+dataset_selfdrivingCar = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_selfdrivingCar = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------29 ShellfishOpenImages---------------------#
+class_name = ('Crab', 'Lobster', 'Shrimp')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'ShellfishOpenImages/raw/'
+dataset_ShellfishOpenImages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_ShellfishOpenImages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------30 ThermalCheetah---------------------#
+class_name = ('cheetah', 'human')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'ThermalCheetah/'
+dataset_ThermalCheetah = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_ThermalCheetah = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------31 thermalDogsAndPeople---------------------#
+class_name = ('dog', 'person')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'thermalDogsAndPeople/'
+dataset_thermalDogsAndPeople = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_thermalDogsAndPeople = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------32 UnoCards---------------------#
+class_name = ('0', '1', '2', '3', '4', '5', '6', '7', '8', '9', '10', '11',
+ '12', '13', '14')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'UnoCards/raw/'
+dataset_UnoCards = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_UnoCards = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------33 VehiclesOpenImages---------------------#
+class_name = ('Ambulance', 'Bus', 'Car', 'Motorcycle', 'Truck')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'VehiclesOpenImages/416x416/'
+dataset_VehiclesOpenImages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_VehiclesOpenImages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------34 WildfireSmoke---------------------#
+class_name = ('smoke', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'WildfireSmoke/'
+dataset_WildfireSmoke = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_WildfireSmoke = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------35 websiteScreenshots---------------------#
+class_name = ('button', 'field', 'heading', 'iframe', 'image', 'label', 'link',
+ 'text')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'websiteScreenshots/'
+dataset_websiteScreenshots = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_websiteScreenshots = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# --------------------- Config---------------------#
+
+dataset_prefixes = [
+ 'AerialMaritimeDrone_large',
+ 'AerialMaritimeDrone_tiled',
+ 'AmericanSignLanguageLetters',
+ 'Aquarium',
+ 'BCCD',
+ 'boggleBoards',
+ 'brackishUnderwater',
+ 'ChessPieces',
+ 'CottontailRabbits',
+ 'dice',
+ 'DroneControl',
+ 'EgoHands_generic',
+ 'EgoHands_specific',
+ 'HardHatWorkers',
+ 'MaskWearing',
+ 'MountainDewCommercial',
+ 'NorthAmericaMushrooms',
+ 'openPoetryVision',
+ 'OxfordPets_by_breed',
+ 'OxfordPets_by_species',
+ 'PKLot',
+ 'Packages',
+ 'PascalVOC',
+ 'pistols',
+ 'plantdoc',
+ 'pothole',
+ 'Raccoons',
+ 'selfdrivingCar',
+ 'ShellfishOpenImages',
+ 'ThermalCheetah',
+ 'thermalDogsAndPeople',
+ 'UnoCards',
+ 'VehiclesOpenImages',
+ 'WildfireSmoke',
+ 'websiteScreenshots',
+]
+
+datasets = [
+ dataset_AerialMaritimeDrone_large, dataset_AerialMaritimeDrone_tiled,
+ dataset_AmericanSignLanguageLetters, dataset_Aquarium, dataset_BCCD,
+ dataset_boggleBoards, dataset_brackishUnderwater, dataset_ChessPieces,
+ dataset_CottontailRabbits, dataset_dice, dataset_DroneControl,
+ dataset_EgoHands_generic, dataset_EgoHands_specific,
+ dataset_HardHatWorkers, dataset_MaskWearing, dataset_MountainDewCommercial,
+ dataset_NorthAmericaMushrooms, dataset_openPoetryVision,
+ dataset_OxfordPets_by_breed, dataset_OxfordPets_by_species, dataset_PKLot,
+ dataset_Packages, dataset_PascalVOC, dataset_pistols, dataset_plantdoc,
+ dataset_pothole, dataset_Raccoon, dataset_selfdrivingCar,
+ dataset_ShellfishOpenImages, dataset_ThermalCheetah,
+ dataset_thermalDogsAndPeople, dataset_UnoCards, dataset_VehiclesOpenImages,
+ dataset_WildfireSmoke, dataset_websiteScreenshots
+]
+
+metrics = [
+ val_evaluator_AerialMaritimeDrone_large,
+ val_evaluator_AerialMaritimeDrone_tiled,
+ val_evaluator_AmericanSignLanguageLetters, val_evaluator_Aquarium,
+ val_evaluator_BCCD, val_evaluator_boggleBoards,
+ val_evaluator_brackishUnderwater, val_evaluator_ChessPieces,
+ val_evaluator_CottontailRabbits, val_evaluator_dice,
+ val_evaluator_DroneControl, val_evaluator_EgoHands_generic,
+ val_evaluator_EgoHands_specific, val_evaluator_HardHatWorkers,
+ val_evaluator_MaskWearing, val_evaluator_MountainDewCommercial,
+ val_evaluator_NorthAmericaMushrooms, val_evaluator_openPoetryVision,
+ val_evaluator_OxfordPets_by_breed, val_evaluator_OxfordPets_by_species,
+ val_evaluator_PKLot, val_evaluator_Packages, val_evaluator_PascalVOC,
+ val_evaluator_pistols, val_evaluator_plantdoc, val_evaluator_pothole,
+ val_evaluator_Raccoon, val_evaluator_selfdrivingCar,
+ val_evaluator_ShellfishOpenImages, val_evaluator_ThermalCheetah,
+ val_evaluator_thermalDogsAndPeople, val_evaluator_UnoCards,
+ val_evaluator_VehiclesOpenImages, val_evaluator_WildfireSmoke,
+ val_evaluator_websiteScreenshots
+]
+
+# -------------------------------------------------#
+val_dataloader = dict(
+ dataset=dict(_delete_=True, type='ConcatDataset', datasets=datasets))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='MultiDatasetsEvaluator',
+ metrics=metrics,
+ dataset_prefixes=dataset_prefixes)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/odinw/glip_atss_swin-t_bc_fpn_dyhead_pretrain_odinw13.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/odinw/glip_atss_swin-t_bc_fpn_dyhead_pretrain_odinw13.py
new file mode 100644
index 0000000000000000000000000000000000000000..c3479b62b781fa38282b26ab69763d1766301dc7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/odinw/glip_atss_swin-t_bc_fpn_dyhead_pretrain_odinw13.py
@@ -0,0 +1,3 @@
+_base_ = './glip_atss_swin-t_a_fpn_dyhead_pretrain_odinw13.py'
+
+model = dict(bbox_head=dict(early_fuse=True))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/odinw/glip_atss_swin-t_bc_fpn_dyhead_pretrain_odinw35.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/odinw/glip_atss_swin-t_bc_fpn_dyhead_pretrain_odinw35.py
new file mode 100644
index 0000000000000000000000000000000000000000..182afc66c93441da85d7e0116970e45a58c492d0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/odinw/glip_atss_swin-t_bc_fpn_dyhead_pretrain_odinw35.py
@@ -0,0 +1,3 @@
+_base_ = './glip_atss_swin-t_a_fpn_dyhead_pretrain_odinw35.py'
+
+model = dict(bbox_head=dict(early_fuse=True))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/odinw/override_category.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/odinw/override_category.py
new file mode 100644
index 0000000000000000000000000000000000000000..9ff05fc6e5e4d0989cf7fcf7af4dc902ee99f3a3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/glip/odinw/override_category.py
@@ -0,0 +1,109 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import argparse
+
+import mmengine
+
+
+def parse_args():
+ parser = argparse.ArgumentParser(description='Override Category')
+ parser.add_argument('data_root')
+ return parser.parse_args()
+
+
+def main():
+ args = parse_args()
+
+ ChessPieces = [{
+ 'id': 1,
+ 'name': ' ',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 2,
+ 'name': 'black bishop',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 3,
+ 'name': 'black king',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 4,
+ 'name': 'black knight',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 5,
+ 'name': 'black pawn',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 6,
+ 'name': 'black queen',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 7,
+ 'name': 'black rook',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 8,
+ 'name': 'white bishop',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 9,
+ 'name': 'white king',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 10,
+ 'name': 'white knight',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 11,
+ 'name': 'white pawn',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 12,
+ 'name': 'white queen',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 13,
+ 'name': 'white rook',
+ 'supercategory': 'pieces'
+ }]
+
+ _data_root = args.data_root + 'ChessPieces/Chess Pieces.v23-raw.coco/'
+ json_data = mmengine.load(_data_root +
+ 'valid/annotations_without_background.json')
+ json_data['categories'] = ChessPieces
+ mmengine.dump(json_data,
+ _data_root + 'valid/new_annotations_without_background.json')
+
+ CottontailRabbits = [{
+ 'id': 1,
+ 'name': 'rabbit',
+ 'supercategory': 'Cottontail-Rabbit'
+ }]
+
+ _data_root = args.data_root + 'CottontailRabbits/'
+ json_data = mmengine.load(_data_root +
+ 'valid/annotations_without_background.json')
+ json_data['categories'] = CottontailRabbits
+ mmengine.dump(json_data,
+ _data_root + 'valid/new_annotations_without_background.json')
+
+ NorthAmericaMushrooms = [{
+ 'id': 1,
+ 'name': 'flat mushroom',
+ 'supercategory': 'mushroom'
+ }, {
+ 'id': 2,
+ 'name': 'yellow mushroom',
+ 'supercategory': 'mushroom'
+ }]
+
+ _data_root = args.data_root + 'NorthAmericaMushrooms/North American Mushrooms.v1-416x416.coco/' # noqa
+ json_data = mmengine.load(_data_root +
+ 'valid/annotations_without_background.json')
+ json_data['categories'] = NorthAmericaMushrooms
+ mmengine.dump(json_data,
+ _data_root + 'valid/new_annotations_without_background.json')
+
+
+if __name__ == '__main__':
+ main()
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..ef8cfc812c40712db9006f7c25d0d3a1f1a8a12c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/README.md
@@ -0,0 +1,54 @@
+# GN + WS
+
+> [Weight Standardization](https://arxiv.org/abs/1903.10520)
+
+
+
+## Abstract
+
+Batch Normalization (BN) has become an out-of-box technique to improve deep network training. However, its effectiveness is limited for micro-batch training, i.e., each GPU typically has only 1-2 images for training, which is inevitable for many computer vision tasks, e.g., object detection and semantic segmentation, constrained by memory consumption. To address this issue, we propose Weight Standardization (WS) and Batch-Channel Normalization (BCN) to bring two success factors of BN into micro-batch training: 1) the smoothing effects on the loss landscape and 2) the ability to avoid harmful elimination singularities along the training trajectory. WS standardizes the weights in convolutional layers to smooth the loss landscape by reducing the Lipschitz constants of the loss and the gradients; BCN combines batch and channel normalizations and leverages estimated statistics of the activations in convolutional layers to keep networks away from elimination singularities. We validate WS and BCN on comprehensive computer vision tasks, including image classification, object detection, instance segmentation, video recognition and semantic segmentation. All experimental results consistently show that WS and BCN improve micro-batch training significantly. Moreover, using WS and BCN with micro-batch training is even able to match or outperform the performances of BN with large-batch training.
+
+
+

+
+
+## Results and Models
+
+Faster R-CNN
+
+| Backbone | Style | Normalization | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :-------------: | :-----: | :-----------: | :-----: | :------: | :------------: | :----: | :-----: | :---------------------------------------------------------: | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | pytorch | GN+WS | 1x | 5.9 | 11.7 | 39.7 | - | [config](./faster-rcnn_r50_fpn_gn-ws-all_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/faster_rcnn_r50_fpn_gn_ws-all_1x_coco/faster_rcnn_r50_fpn_gn_ws-all_1x_coco_20200130-613d9fe2.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/faster_rcnn_r50_fpn_gn_ws-all_1x_coco/faster_rcnn_r50_fpn_gn_ws-all_1x_coco_20200130_210936.log.json) |
+| R-101-FPN | pytorch | GN+WS | 1x | 8.9 | 9.0 | 41.7 | - | [config](./faster-rcnn_r101_fpn_gn-ws-all_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/faster_rcnn_r101_fpn_gn_ws-all_1x_coco/faster_rcnn_r101_fpn_gn_ws-all_1x_coco_20200205-a93b0d75.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/faster_rcnn_r101_fpn_gn_ws-all_1x_coco/faster_rcnn_r101_fpn_gn_ws-all_1x_coco_20200205_232146.log.json) |
+| X-50-32x4d-FPN | pytorch | GN+WS | 1x | 7.0 | 10.3 | 40.7 | - | [config](./faster-rcnn_x50-32x4d_fpn_gn-ws-all_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/faster_rcnn_x50_32x4d_fpn_gn_ws-all_1x_coco/faster_rcnn_x50_32x4d_fpn_gn_ws-all_1x_coco_20200203-839c5d9d.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/faster_rcnn_x50_32x4d_fpn_gn_ws-all_1x_coco/faster_rcnn_x50_32x4d_fpn_gn_ws-all_1x_coco_20200203_220113.log.json) |
+| X-101-32x4d-FPN | pytorch | GN+WS | 1x | 10.8 | 7.6 | 42.1 | - | [config](./faster-rcnn_x101-32x4d_fpn_gn-ws-all_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/faster_rcnn_x101_32x4d_fpn_gn_ws-all_1x_coco/faster_rcnn_x101_32x4d_fpn_gn_ws-all_1x_coco_20200212-27da1bc2.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/faster_rcnn_x101_32x4d_fpn_gn_ws-all_1x_coco/faster_rcnn_x101_32x4d_fpn_gn_ws-all_1x_coco_20200212_195302.log.json) |
+
+Mask R-CNN
+
+| Backbone | Style | Normalization | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :-------------: | :-----: | :-----------: | :-------: | :------: | :------------: | :----: | :-----: | :--------------------------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | pytorch | GN+WS | 2x | 7.3 | 10.5 | 40.6 | 36.6 | [config](./mask-rcnn_r50_fpn_gn-ws-all_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_r50_fpn_gn_ws-all_2x_coco/mask_rcnn_r50_fpn_gn_ws-all_2x_coco_20200226-16acb762.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_r50_fpn_gn_ws-all_2x_coco/mask_rcnn_r50_fpn_gn_ws-all_2x_coco_20200226_062128.log.json) |
+| R-101-FPN | pytorch | GN+WS | 2x | 10.3 | 8.6 | 42.0 | 37.7 | [config](./mask-rcnn_r101_fpn_gn-ws-all_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_r101_fpn_gn_ws-all_2x_coco/mask_rcnn_r101_fpn_gn_ws-all_2x_coco_20200212-ea357cd9.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_r101_fpn_gn_ws-all_2x_coco/mask_rcnn_r101_fpn_gn_ws-all_2x_coco_20200212_213627.log.json) |
+| X-50-32x4d-FPN | pytorch | GN+WS | 2x | 8.4 | 9.3 | 41.1 | 37.0 | [config](./mask-rcnn_x50-32x4d_fpn_gn-ws-all_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_x50_32x4d_fpn_gn_ws-all_2x_coco/mask_rcnn_x50_32x4d_fpn_gn_ws-all_2x_coco_20200216-649fdb6f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_x50_32x4d_fpn_gn_ws-all_2x_coco/mask_rcnn_x50_32x4d_fpn_gn_ws-all_2x_coco_20200216_201500.log.json) |
+| X-101-32x4d-FPN | pytorch | GN+WS | 2x | 12.2 | 7.1 | 42.1 | 37.9 | [config](./mask-rcnn_x101-32x4d_fpn_gn-ws-all_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_x101_32x4d_fpn_gn_ws-all_2x_coco/mask_rcnn_x101_32x4d_fpn_gn_ws-all_2x_coco_20200319-33fb95b5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_x101_32x4d_fpn_gn_ws-all_2x_coco/mask_rcnn_x101_32x4d_fpn_gn_ws-all_2x_coco_20200319_104101.log.json) |
+| R-50-FPN | pytorch | GN+WS | 20-23-24e | 7.3 | - | 41.1 | 37.1 | [config](./mask-rcnn_r50_fpn_gn-ws-all_20-23-24e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_r50_fpn_gn_ws-all_20_23_24e_coco/mask_rcnn_r50_fpn_gn_ws-all_20_23_24e_coco_20200213-487d1283.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_r50_fpn_gn_ws-all_20_23_24e_coco/mask_rcnn_r50_fpn_gn_ws-all_20_23_24e_coco_20200213_035123.log.json) |
+| R-101-FPN | pytorch | GN+WS | 20-23-24e | 10.3 | - | 43.1 | 38.6 | [config](./mask-rcnn_r101_fpn_gn-ws-all_20-23-24e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_r101_fpn_gn_ws-all_20_23_24e_coco/mask_rcnn_r101_fpn_gn_ws-all_20_23_24e_coco_20200213-57b5a50f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_r101_fpn_gn_ws-all_20_23_24e_coco/mask_rcnn_r101_fpn_gn_ws-all_20_23_24e_coco_20200213_130142.log.json) |
+| X-50-32x4d-FPN | pytorch | GN+WS | 20-23-24e | 8.4 | - | 42.1 | 38.0 | [config](./mask-rcnn_x50-32x4d_fpn_gn-ws-all_20-23-24e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_x50_32x4d_fpn_gn_ws-all_20_23_24e_coco/mask_rcnn_x50_32x4d_fpn_gn_ws-all_20_23_24e_coco_20200226-969bcb2c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_x50_32x4d_fpn_gn_ws-all_20_23_24e_coco/mask_rcnn_x50_32x4d_fpn_gn_ws-all_20_23_24e_coco_20200226_093732.log.json) |
+| X-101-32x4d-FPN | pytorch | GN+WS | 20-23-24e | 12.2 | - | 42.7 | 38.5 | [config](./mask-rcnn_x101-32x4d_fpn_gn-ws-all_20-23-24e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_x101_32x4d_fpn_gn_ws-all_20_23_24e_coco/mask_rcnn_x101_32x4d_fpn_gn_ws-all_20_23_24e_coco_20200316-e6cd35ef.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_x101_32x4d_fpn_gn_ws-all_20_23_24e_coco/mask_rcnn_x101_32x4d_fpn_gn_ws-all_20_23_24e_coco_20200316_013741.log.json) |
+
+Note:
+
+- GN+WS requires about 5% more memory than GN, and it is only 5% slower than GN.
+- In the paper, a 20-23-24e lr schedule is used instead of 2x.
+- The X-50-GN and X-101-GN pretrained models are also shared by the authors.
+
+## Citation
+
+```latex
+@article{weightstandardization,
+ author = {Siyuan Qiao and Huiyu Wang and Chenxi Liu and Wei Shen and Alan Yuille},
+ title = {Weight Standardization},
+ journal = {arXiv preprint arXiv:1903.10520},
+ year = {2019},
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/faster-rcnn_r101_fpn_gn-ws-all_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/faster-rcnn_r101_fpn_gn-ws-all_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a4cb8281ac6d4b43a615ba1a05903770d8ee2f69
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/faster-rcnn_r101_fpn_gn-ws-all_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './faster-rcnn_r50_fpn_gn-ws-all_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://jhu/resnet101_gn_ws')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/faster-rcnn_r50_fpn_gn-ws-all_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/faster-rcnn_r50_fpn_gn-ws-all_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1a044c99a2e84de71822edb62543570891141b25
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/faster-rcnn_r50_fpn_gn-ws-all_1x_coco.py
@@ -0,0 +1,16 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+conv_cfg = dict(type='ConvWS')
+norm_cfg = dict(type='GN', num_groups=32, requires_grad=True)
+model = dict(
+ backbone=dict(
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://jhu/resnet50_gn_ws')),
+ neck=dict(conv_cfg=conv_cfg, norm_cfg=norm_cfg),
+ roi_head=dict(
+ bbox_head=dict(
+ type='Shared4Conv1FCBBoxHead',
+ conv_out_channels=256,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/faster-rcnn_x101-32x4d_fpn_gn-ws-all_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/faster-rcnn_x101-32x4d_fpn_gn-ws-all_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b2a317d2ac830d95788084eaa8d374838b34a365
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/faster-rcnn_x101-32x4d_fpn_gn-ws-all_1x_coco.py
@@ -0,0 +1,18 @@
+_base_ = './faster-rcnn_r50_fpn_gn-ws-all_1x_coco.py'
+conv_cfg = dict(type='ConvWS')
+norm_cfg = dict(type='GN', num_groups=32, requires_grad=True)
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ style='pytorch',
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://jhu/resnext101_32x4d_gn_ws')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/faster-rcnn_x50-32x4d_fpn_gn-ws-all_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/faster-rcnn_x50-32x4d_fpn_gn-ws-all_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..dd75a2c004b8cc04411d47d8b9db6ba0ec4ffcb0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/faster-rcnn_x50-32x4d_fpn_gn-ws-all_1x_coco.py
@@ -0,0 +1,18 @@
+_base_ = './faster-rcnn_r50_fpn_gn-ws-all_1x_coco.py'
+conv_cfg = dict(type='ConvWS')
+norm_cfg = dict(type='GN', num_groups=32, requires_grad=True)
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=50,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ style='pytorch',
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://jhu/resnext50_32x4d_gn_ws')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_r101_fpn_gn-ws-all_20-23-24e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_r101_fpn_gn-ws-all_20-23-24e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1815e3f85b9fd5d7204b08cd60a13980a382fd51
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_r101_fpn_gn-ws-all_20-23-24e_coco.py
@@ -0,0 +1,17 @@
+_base_ = './mask-rcnn_r101_fpn_gn-ws-all_2x_coco.py'
+# learning policy
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[20, 23],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_r101_fpn_gn-ws-all_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_r101_fpn_gn-ws-all_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5de37dee5e86e202c211464eaa08dd295dba44b2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_r101_fpn_gn-ws-all_2x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './mask-rcnn_r50_fpn_gn-ws-all_2x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://jhu/resnet101_gn_ws')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_r50_fpn_gn-ws-all_20-23-24e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_r50_fpn_gn-ws-all_20-23-24e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..287c652045d6230411043f2abab34be4f6106687
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_r50_fpn_gn-ws-all_20-23-24e_coco.py
@@ -0,0 +1,17 @@
+_base_ = './mask-rcnn_r50_fpn_gn-ws-all_2x_coco.py'
+# learning policy
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[20, 23],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_r50_fpn_gn-ws-all_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_r50_fpn_gn-ws-all_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ed8b1b73fe8695fc6bbb4054405192fca995cf81
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_r50_fpn_gn-ws-all_2x_coco.py
@@ -0,0 +1,33 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+conv_cfg = dict(type='ConvWS')
+norm_cfg = dict(type='GN', num_groups=32, requires_grad=True)
+model = dict(
+ backbone=dict(
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://jhu/resnet50_gn_ws')),
+ neck=dict(conv_cfg=conv_cfg, norm_cfg=norm_cfg),
+ roi_head=dict(
+ bbox_head=dict(
+ type='Shared4Conv1FCBBoxHead',
+ conv_out_channels=256,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg),
+ mask_head=dict(conv_cfg=conv_cfg, norm_cfg=norm_cfg)))
+# learning policy
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_x101-32x4d_fpn_gn-ws-all_20-23-24e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_x101-32x4d_fpn_gn-ws-all_20-23-24e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8ce9193579b914f8dc0804cb73c3d8e41b153655
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_x101-32x4d_fpn_gn-ws-all_20-23-24e_coco.py
@@ -0,0 +1,17 @@
+_base_ = './mask-rcnn_x101-32x4d_fpn_gn-ws-all_2x_coco.py'
+# learning policy
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[20, 23],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_x101-32x4d_fpn_gn-ws-all_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_x101-32x4d_fpn_gn-ws-all_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..bcfc371e774470ede7d171b4268db919385775ab
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_x101-32x4d_fpn_gn-ws-all_2x_coco.py
@@ -0,0 +1,19 @@
+_base_ = './mask-rcnn_r50_fpn_gn-ws-all_2x_coco.py'
+# model settings
+conv_cfg = dict(type='ConvWS')
+norm_cfg = dict(type='GN', num_groups=32, requires_grad=True)
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ style='pytorch',
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://jhu/resnext101_32x4d_gn_ws')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_x50-32x4d_fpn_gn-ws-all_20-23-24e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_x50-32x4d_fpn_gn-ws-all_20-23-24e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..af9ea5ab476b8ea3247062261726bef6b6bc1b0c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_x50-32x4d_fpn_gn-ws-all_20-23-24e_coco.py
@@ -0,0 +1,17 @@
+_base_ = './mask-rcnn_x50-32x4d_fpn_gn-ws-all_2x_coco.py'
+# learning policy
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[20, 23],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_x50-32x4d_fpn_gn-ws-all_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_x50-32x4d_fpn_gn-ws-all_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ab2b14042e9510ab14698e7a64c68d6ff60835e1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/mask-rcnn_x50-32x4d_fpn_gn-ws-all_2x_coco.py
@@ -0,0 +1,19 @@
+_base_ = './mask-rcnn_r50_fpn_gn-ws-all_2x_coco.py'
+# model settings
+conv_cfg = dict(type='ConvWS')
+norm_cfg = dict(type='GN', num_groups=32, requires_grad=True)
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=50,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ style='pytorch',
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://jhu/resnext50_32x4d_gn_ws')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..89b91072924a31e53db1e95df30b47636a67b74b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn+ws/metafile.yml
@@ -0,0 +1,263 @@
+Collections:
+ - Name: Weight Standardization
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Group Normalization
+ - Weight Standardization
+ Paper:
+ URL: https://arxiv.org/abs/1903.10520
+ Title: 'Weight Standardization'
+ README: configs/gn+ws/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/configs/gn%2Bws/mask-rcnn_r50_fpn_gn-ws-all_2x_coco.py
+ Version: v2.0.0
+
+Models:
+ - Name: faster-rcnn_r50_fpn_gn_ws-all_1x_coco
+ In Collection: Weight Standardization
+ Config: configs/gn%2Bws/faster-rcnn_r50_fpn_gn-ws-all_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.9
+ inference time (ms/im):
+ - value: 85.47
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/faster_rcnn_r50_fpn_gn_ws-all_1x_coco/faster_rcnn_r50_fpn_gn_ws-all_1x_coco_20200130-613d9fe2.pth
+
+ - Name: faster-rcnn_r101_fpn_gn-ws-all_1x_coco
+ In Collection: Weight Standardization
+ Config: configs/gn%2Bws/faster-rcnn_r101_fpn_gn-ws-all_1x_coco.py
+ Metadata:
+ Training Memory (GB): 8.9
+ inference time (ms/im):
+ - value: 111.11
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/faster_rcnn_r101_fpn_gn_ws-all_1x_coco/faster_rcnn_r101_fpn_gn_ws-all_1x_coco_20200205-a93b0d75.pth
+
+ - Name: faster-rcnn_x50-32x4d_fpn_gn-ws-all_1x_coco
+ In Collection: Weight Standardization
+ Config: configs/gn%2Bws/faster-rcnn_x50-32x4d_fpn_gn-ws-all_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.0
+ inference time (ms/im):
+ - value: 97.09
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/faster_rcnn_x50_32x4d_fpn_gn_ws-all_1x_coco/faster_rcnn_x50_32x4d_fpn_gn_ws-all_1x_coco_20200203-839c5d9d.pth
+
+ - Name: faster-rcnn_x101-32x4d_fpn_gn-ws-all_1x_coco
+ In Collection: Weight Standardization
+ Config: configs/gn%2Bws/faster-rcnn_x101-32x4d_fpn_gn-ws-all_1x_coco.py
+ Metadata:
+ Training Memory (GB): 10.8
+ inference time (ms/im):
+ - value: 131.58
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/faster_rcnn_x101_32x4d_fpn_gn_ws-all_1x_coco/faster_rcnn_x101_32x4d_fpn_gn_ws-all_1x_coco_20200212-27da1bc2.pth
+
+ - Name: mask-rcnn_r50_fpn_gn_ws-all_2x_coco
+ In Collection: Weight Standardization
+ Config: configs/gn%2Bws/mask-rcnn_r50_fpn_gn-ws-all_2x_coco.py
+ Metadata:
+ Training Memory (GB): 7.3
+ inference time (ms/im):
+ - value: 95.24
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.6
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_r50_fpn_gn_ws-all_2x_coco/mask_rcnn_r50_fpn_gn_ws-all_2x_coco_20200226-16acb762.pth
+
+ - Name: mask-rcnn_r101_fpn_gn-ws-all_2x_coco
+ In Collection: Weight Standardization
+ Config: configs/gn%2Bws/mask-rcnn_r101_fpn_gn-ws-all_2x_coco.py
+ Metadata:
+ Training Memory (GB): 10.3
+ inference time (ms/im):
+ - value: 116.28
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_r101_fpn_gn_ws-all_2x_coco/mask_rcnn_r101_fpn_gn_ws-all_2x_coco_20200212-ea357cd9.pth
+
+ - Name: mask-rcnn_x50-32x4d_fpn_gn-ws-all_2x_coco
+ In Collection: Weight Standardization
+ Config: configs/gn%2Bws/mask-rcnn_x50-32x4d_fpn_gn-ws-all_2x_coco.py
+ Metadata:
+ Training Memory (GB): 8.4
+ inference time (ms/im):
+ - value: 107.53
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.1
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_x50_32x4d_fpn_gn_ws-all_2x_coco/mask_rcnn_x50_32x4d_fpn_gn_ws-all_2x_coco_20200216-649fdb6f.pth
+
+ - Name: mask-rcnn_x101-32x4d_fpn_gn-ws-all_2x_coco
+ In Collection: Weight Standardization
+ Config: configs/gn%2Bws/mask-rcnn_x101-32x4d_fpn_gn-ws-all_2x_coco.py
+ Metadata:
+ Training Memory (GB): 12.2
+ inference time (ms/im):
+ - value: 140.85
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.1
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_x101_32x4d_fpn_gn_ws-all_2x_coco/mask_rcnn_x101_32x4d_fpn_gn_ws-all_2x_coco_20200319-33fb95b5.pth
+
+ - Name: mask-rcnn_r50_fpn_gn_ws-all_20_23_24e_coco
+ In Collection: Weight Standardization
+ Config: configs/gn%2Bws/mask-rcnn_r50_fpn_gn-ws-all_20-23-24e_coco.py
+ Metadata:
+ Training Memory (GB): 7.3
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.1
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_r50_fpn_gn_ws-all_20_23_24e_coco/mask_rcnn_r50_fpn_gn_ws-all_20_23_24e_coco_20200213-487d1283.pth
+
+ - Name: mask-rcnn_r101_fpn_gn-ws-all_20-23-24e_coco
+ In Collection: Weight Standardization
+ Config: configs/gn%2Bws/mask-rcnn_r101_fpn_gn-ws-all_20-23-24e_coco.py
+ Metadata:
+ Training Memory (GB): 10.3
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.1
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_r101_fpn_gn_ws-all_20_23_24e_coco/mask_rcnn_r101_fpn_gn_ws-all_20_23_24e_coco_20200213-57b5a50f.pth
+
+ - Name: mask-rcnn_x50-32x4d_fpn_gn-ws-all_20-23-24e_coco
+ In Collection: Weight Standardization
+ Config: configs/gn%2Bws/mask-rcnn_x50-32x4d_fpn_gn-ws-all_20-23-24e_coco.py
+ Metadata:
+ Training Memory (GB): 8.4
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.1
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_x50_32x4d_fpn_gn_ws-all_20_23_24e_coco/mask_rcnn_x50_32x4d_fpn_gn_ws-all_20_23_24e_coco_20200226-969bcb2c.pth
+
+ - Name: mask-rcnn_x101-32x4d_fpn_gn-ws-all_20-23-24e_coco
+ In Collection: Weight Standardization
+ Config: configs/gn%2Bws/mask-rcnn_x101-32x4d_fpn_gn-ws-all_20-23-24e_coco.py
+ Metadata:
+ Training Memory (GB): 12.2
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.7
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gn%2Bws/mask_rcnn_x101_32x4d_fpn_gn_ws-all_20_23_24e_coco/mask_rcnn_x101_32x4d_fpn_gn_ws-all_20_23_24e_coco_20200316-e6cd35ef.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..1bc8192f24a56b11449944fc3d949302dfa781b6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/README.md
@@ -0,0 +1,41 @@
+# GN
+
+> [Group Normalization](https://arxiv.org/abs/1803.08494)
+
+
+
+## Abstract
+
+Batch Normalization (BN) is a milestone technique in the development of deep learning, enabling various networks to train. However, normalizing along the batch dimension introduces problems --- BN's error increases rapidly when the batch size becomes smaller, caused by inaccurate batch statistics estimation. This limits BN's usage for training larger models and transferring features to computer vision tasks including detection, segmentation, and video, which require small batches constrained by memory consumption. In this paper, we present Group Normalization (GN) as a simple alternative to BN. GN divides the channels into groups and computes within each group the mean and variance for normalization. GN's computation is independent of batch sizes, and its accuracy is stable in a wide range of batch sizes. On ResNet-50 trained in ImageNet, GN has 10.6% lower error than its BN counterpart when using a batch size of 2; when using typical batch sizes, GN is comparably good with BN and outperforms other normalization variants. Moreover, GN can be naturally transferred from pre-training to fine-tuning. GN can outperform its BN-based counterparts for object detection and segmentation in COCO, and for video classification in Kinetics, showing that GN can effectively replace the powerful BN in a variety of tasks. GN can be easily implemented by a few lines of code in modern libraries.
+
+
+

+
+
+## Results and Models
+
+| Backbone | model | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :-----------: | :--------: | :-----: | :------: | :------------: | :----: | :-----: | :-----------------------------------------------------: | :-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN (d) | Mask R-CNN | 2x | 7.1 | 11.0 | 40.2 | 36.4 | [config](./mask-rcnn_r50_fpn_gn-all_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gn/mask_rcnn_r50_fpn_gn-all_2x_coco/mask_rcnn_r50_fpn_gn-all_2x_coco_20200206-8eee02a6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gn/mask_rcnn_r50_fpn_gn-all_2x_coco/mask_rcnn_r50_fpn_gn-all_2x_coco_20200206_050355.log.json) |
+| R-50-FPN (d) | Mask R-CNN | 3x | 7.1 | - | 40.5 | 36.7 | [config](./mask-rcnn_r50_fpn_gn-all_3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gn/mask_rcnn_r50_fpn_gn-all_3x_coco/mask_rcnn_r50_fpn_gn-all_3x_coco_20200214-8b23b1e5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gn/mask_rcnn_r50_fpn_gn-all_3x_coco/mask_rcnn_r50_fpn_gn-all_3x_coco_20200214_063512.log.json) |
+| R-101-FPN (d) | Mask R-CNN | 2x | 9.9 | 9.0 | 41.9 | 37.6 | [config](./mask-rcnn_r101_fpn_gn-all_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gn/mask_rcnn_r101_fpn_gn-all_2x_coco/mask_rcnn_r101_fpn_gn-all_2x_coco_20200205-d96b1b50.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gn/mask_rcnn_r101_fpn_gn-all_2x_coco/mask_rcnn_r101_fpn_gn-all_2x_coco_20200205_234402.log.json) |
+| R-101-FPN (d) | Mask R-CNN | 3x | 9.9 | | 42.1 | 38.0 | [config](./mask-rcnn_r101_fpn_gn-all_3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gn/mask_rcnn_r101_fpn_gn-all_3x_coco/mask_rcnn_r101_fpn_gn-all_3x_coco_20200513_181609-0df864f4.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gn/mask_rcnn_r101_fpn_gn-all_3x_coco/mask_rcnn_r101_fpn_gn-all_3x_coco_20200513_181609.log.json) |
+| R-50-FPN (c) | Mask R-CNN | 2x | 7.1 | 10.9 | 40.0 | 36.1 | [config](./mask-rcnn_r50-contrib_fpn_gn-all_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gn/mask_rcnn_r50_fpn_gn-all_contrib_2x_coco/mask_rcnn_r50_fpn_gn-all_contrib_2x_coco_20200207-20d3e849.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gn/mask_rcnn_r50_fpn_gn-all_contrib_2x_coco/mask_rcnn_r50_fpn_gn-all_contrib_2x_coco_20200207_225832.log.json) |
+| R-50-FPN (c) | Mask R-CNN | 3x | 7.1 | - | 40.1 | 36.2 | [config](./mask-rcnn_r50-contrib_fpn_gn-all_3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gn/mask_rcnn_r50_fpn_gn-all_contrib_3x_coco/mask_rcnn_r50_fpn_gn-all_contrib_3x_coco_20200225-542aefbc.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gn/mask_rcnn_r50_fpn_gn-all_contrib_3x_coco/mask_rcnn_r50_fpn_gn-all_contrib_3x_coco_20200225_235135.log.json) |
+
+**Notes:**
+
+- (d) means pretrained model converted from Detectron, and (c) means the contributed model pretrained by [@thangvubk](https://github.com/thangvubk).
+- The `3x` schedule is epoch \[28, 34, 36\].
+- **Memory, Train/Inf time is outdated.**
+
+## Citation
+
+```latex
+@inproceedings{wu2018group,
+ title={Group Normalization},
+ author={Wu, Yuxin and He, Kaiming},
+ booktitle={Proceedings of the European Conference on Computer Vision (ECCV)},
+ year={2018}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/mask-rcnn_r101_fpn_gn-all_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/mask-rcnn_r101_fpn_gn-all_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..54f57d8d0855d07c696907d8c7c0758e4c13a573
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/mask-rcnn_r101_fpn_gn-all_2x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './mask-rcnn_r50_fpn_gn-all_2x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron/resnet101_gn')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/mask-rcnn_r101_fpn_gn-all_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/mask-rcnn_r101_fpn_gn-all_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a94e063ecd2a5e2fd83eb78aa4d7ddd8f51e2b9e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/mask-rcnn_r101_fpn_gn-all_3x_coco.py
@@ -0,0 +1,18 @@
+_base_ = './mask-rcnn_r101_fpn_gn-all_2x_coco.py'
+
+# learning policy
+max_epochs = 36
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[28, 34],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/mask-rcnn_r50-contrib_fpn_gn-all_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/mask-rcnn_r50-contrib_fpn_gn-all_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5515ec14a47a0dfa58acf6c46bc40d77ce39ac3d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/mask-rcnn_r50-contrib_fpn_gn-all_2x_coco.py
@@ -0,0 +1,31 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+norm_cfg = dict(type='GN', num_groups=32, requires_grad=True)
+model = dict(
+ backbone=dict(
+ norm_cfg=norm_cfg,
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://contrib/resnet50_gn')),
+ neck=dict(norm_cfg=norm_cfg),
+ roi_head=dict(
+ bbox_head=dict(
+ type='Shared4Conv1FCBBoxHead',
+ conv_out_channels=256,
+ norm_cfg=norm_cfg),
+ mask_head=dict(norm_cfg=norm_cfg)))
+
+# learning policy
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/mask-rcnn_r50-contrib_fpn_gn-all_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/mask-rcnn_r50-contrib_fpn_gn-all_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e6f7a97e8e0482836b225e832be2e3de4ae99947
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/mask-rcnn_r50-contrib_fpn_gn-all_3x_coco.py
@@ -0,0 +1,18 @@
+_base_ = './mask-rcnn_r50-contrib_fpn_gn-all_2x_coco.py'
+
+# learning policy
+max_epochs = 36
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[28, 34],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/mask-rcnn_r50_fpn_gn-all_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/mask-rcnn_r50_fpn_gn-all_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1313b22e4795239d5148fb8d665cdadb5fac8e4f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/mask-rcnn_r50_fpn_gn-all_2x_coco.py
@@ -0,0 +1,36 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+norm_cfg = dict(type='GN', num_groups=32, requires_grad=True)
+model = dict(
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False),
+ backbone=dict(
+ norm_cfg=norm_cfg,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron/resnet50_gn')),
+ neck=dict(norm_cfg=norm_cfg),
+ roi_head=dict(
+ bbox_head=dict(
+ type='Shared4Conv1FCBBoxHead',
+ conv_out_channels=256,
+ norm_cfg=norm_cfg),
+ mask_head=dict(norm_cfg=norm_cfg)))
+
+# learning policy
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/mask-rcnn_r50_fpn_gn-all_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/mask-rcnn_r50_fpn_gn-all_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e425de951bb0419d1d1596e45637be1d914a8034
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/mask-rcnn_r50_fpn_gn-all_3x_coco.py
@@ -0,0 +1,18 @@
+_base_ = './mask-rcnn_r50_fpn_gn-all_2x_coco.py'
+
+# learning policy
+max_epochs = 36
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[28, 34],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..9781dc9393f17b89a8e4228ef905a06dfdbc7eb5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/gn/metafile.yml
@@ -0,0 +1,162 @@
+Collections:
+ - Name: Group Normalization
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Group Normalization
+ Paper:
+ URL: https://arxiv.org/abs/1803.08494
+ Title: 'Group Normalization'
+ README: configs/gn/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/configs/gn/mask-rcnn_r50_fpn_gn-all_2x_coco.py
+ Version: v2.0.0
+
+Models:
+ - Name: mask-rcnn_r50_fpn_gn-all_2x_coco
+ In Collection: Group Normalization
+ Config: configs/gn/mask-rcnn_r50_fpn_gn-all_2x_coco.py
+ Metadata:
+ Training Memory (GB): 7.1
+ inference time (ms/im):
+ - value: 90.91
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gn/mask_rcnn_r50_fpn_gn-all_2x_coco/mask_rcnn_r50_fpn_gn-all_2x_coco_20200206-8eee02a6.pth
+
+ - Name: mask-rcnn_r50_fpn_gn-all_3x_coco
+ In Collection: Group Normalization
+ Config: configs/gn/mask-rcnn_r50_fpn_gn-all_3x_coco.py
+ Metadata:
+ Training Memory (GB): 7.1
+ inference time (ms/im):
+ - value: 90.91
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gn/mask_rcnn_r50_fpn_gn-all_3x_coco/mask_rcnn_r50_fpn_gn-all_3x_coco_20200214-8b23b1e5.pth
+
+ - Name: mask-rcnn_r101_fpn_gn-all_2x_coco
+ In Collection: Group Normalization
+ Config: configs/gn/mask-rcnn_r101_fpn_gn-all_2x_coco.py
+ Metadata:
+ Training Memory (GB): 9.9
+ inference time (ms/im):
+ - value: 111.11
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.9
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gn/mask_rcnn_r101_fpn_gn-all_2x_coco/mask_rcnn_r101_fpn_gn-all_2x_coco_20200205-d96b1b50.pth
+
+ - Name: mask-rcnn_r101_fpn_gn-all_3x_coco
+ In Collection: Group Normalization
+ Config: configs/gn/mask-rcnn_r101_fpn_gn-all_3x_coco.py
+ Metadata:
+ Training Memory (GB): 9.9
+ inference time (ms/im):
+ - value: 111.11
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.1
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gn/mask_rcnn_r101_fpn_gn-all_3x_coco/mask_rcnn_r101_fpn_gn-all_3x_coco_20200513_181609-0df864f4.pth
+
+ - Name: mask-rcnn_r50_fpn_gn-all_contrib_2x_coco
+ In Collection: Group Normalization
+ Config: configs/gn/mask-rcnn_r50-contrib_fpn_gn-all_2x_coco.py
+ Metadata:
+ Training Memory (GB): 7.1
+ inference time (ms/im):
+ - value: 91.74
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gn/mask_rcnn_r50_fpn_gn-all_contrib_2x_coco/mask_rcnn_r50_fpn_gn-all_contrib_2x_coco_20200207-20d3e849.pth
+
+ - Name: mask-rcnn_r50_fpn_gn-all_contrib_3x_coco
+ In Collection: Group Normalization
+ Config: configs/gn/mask-rcnn_r50-contrib_fpn_gn-all_3x_coco.py
+ Metadata:
+ Training Memory (GB): 7.1
+ inference time (ms/im):
+ - value: 91.74
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.1
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/gn/mask_rcnn_r50_fpn_gn-all_contrib_3x_coco/mask_rcnn_r50_fpn_gn-all_contrib_3x_coco_20200225-542aefbc.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..3de810afc66c29df6ab9bd1728d0cb8b57316acf
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/README.md
@@ -0,0 +1,47 @@
+# Grid R-CNN
+
+> [Grid R-CNN](https://arxiv.org/abs/1811.12030)
+
+
+
+## Abstract
+
+This paper proposes a novel object detection framework named Grid R-CNN, which adopts a grid guided localization mechanism for accurate object detection. Different from the traditional regression based methods, the Grid R-CNN captures the spatial information explicitly and enjoys the position sensitive property of fully convolutional architecture. Instead of using only two independent points, we design a multi-point supervision formulation to encode more clues in order to reduce the impact of inaccurate prediction of specific points. To take the full advantage of the correlation of points in a grid, we propose a two-stage information fusion strategy to fuse feature maps of neighbor grid points. The grid guided localization approach is easy to be extended to different state-of-the-art detection frameworks. Grid R-CNN leads to high quality object localization, and experiments demonstrate that it achieves a 4.1% AP gain at IoU=0.8 and a 10.0% AP gain at IoU=0.9 on COCO benchmark compared to Faster R-CNN with Res50 backbone and FPN architecture.
+
+Grid R-CNN is a well-performed objection detection framework. It transforms the traditional box offset regression problem into a grid point estimation problem. With the guidance of the grid points, it can obtain high-quality localization results. However, the speed of Grid R-CNN is not so satisfactory. In this technical report we present Grid R-CNN Plus, a better and faster version of Grid R-CNN. We have made several updates that significantly speed up the framework and simultaneously improve the accuracy. On COCO dataset, the Res50-FPN based Grid R-CNN Plus detector achieves an mAP of 40.4%, outperforming the baseline on the same model by 3.0 points with similar inference time.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :---------: | :-----: | :------: | :------------: | :----: | :-----------------------------------------------------: | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | 2x | 5.1 | 15.0 | 40.4 | [config](./grid-rcnn_r50_fpn_gn-head_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/grid_rcnn/grid_rcnn_r50_fpn_gn-head_2x_coco/grid_rcnn_r50_fpn_gn-head_2x_coco_20200130-6cca8223.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/grid_rcnn/grid_rcnn_r50_fpn_gn-head_2x_coco/grid_rcnn_r50_fpn_gn-head_2x_coco_20200130_221140.log.json) |
+| R-101 | 2x | 7.0 | 12.6 | 41.5 | [config](./grid-rcnn_r101_fpn_gn-head_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/grid_rcnn/grid_rcnn_r101_fpn_gn-head_2x_coco/grid_rcnn_r101_fpn_gn-head_2x_coco_20200309-d6eca030.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/grid_rcnn/grid_rcnn_r101_fpn_gn-head_2x_coco/grid_rcnn_r101_fpn_gn-head_2x_coco_20200309_164224.log.json) |
+| X-101-32x4d | 2x | 8.3 | 10.8 | 42.9 | [config](./grid-rcnn_x101-32x4d_fpn_gn-head_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/grid_rcnn/grid_rcnn_x101_32x4d_fpn_gn-head_2x_coco/grid_rcnn_x101_32x4d_fpn_gn-head_2x_coco_20200130-d8f0e3ff.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/grid_rcnn/grid_rcnn_x101_32x4d_fpn_gn-head_2x_coco/grid_rcnn_x101_32x4d_fpn_gn-head_2x_coco_20200130_215413.log.json) |
+| X-101-64x4d | 2x | 11.3 | 7.7 | 43.0 | [config](./grid-rcnn_x101-64x4d_fpn_gn-head_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/grid_rcnn/grid_rcnn_x101_64x4d_fpn_gn-head_2x_coco/grid_rcnn_x101_64x4d_fpn_gn-head_2x_coco_20200204-ec76a754.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/grid_rcnn/grid_rcnn_x101_64x4d_fpn_gn-head_2x_coco/grid_rcnn_x101_64x4d_fpn_gn-head_2x_coco_20200204_080641.log.json) |
+
+**Notes:**
+
+- All models are trained with 8 GPUs instead of 32 GPUs in the original paper.
+- The warming up lasts for 1 epoch and `2x` here indicates 25 epochs.
+
+## Citation
+
+```latex
+@inproceedings{lu2019grid,
+ title={Grid r-cnn},
+ author={Lu, Xin and Li, Buyu and Yue, Yuxin and Li, Quanquan and Yan, Junjie},
+ booktitle={Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition},
+ year={2019}
+}
+
+@article{lu2019grid,
+ title={Grid R-CNN Plus: Faster and Better},
+ author={Lu, Xin and Li, Buyu and Yue, Yuxin and Li, Quanquan and Yan, Junjie},
+ journal={arXiv preprint arXiv:1906.05688},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/grid-rcnn_r101_fpn_gn-head_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/grid-rcnn_r101_fpn_gn-head_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..46d41ed4ed5d1d6345e98434221cc5b07c60767d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/grid-rcnn_r101_fpn_gn-head_2x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './grid-rcnn_r50_fpn_gn-head_2x_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/grid-rcnn_r50_fpn_gn-head_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/grid-rcnn_r50_fpn_gn-head_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..358280630fa96e40ac7834cbda6b1ad3dc689c55
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/grid-rcnn_r50_fpn_gn-head_1x_coco.py
@@ -0,0 +1,19 @@
+_base_ = './grid-rcnn_r50_fpn_gn-head_2x_coco.py'
+
+# training schedule
+max_epochs = 12
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.0001, by_epoch=False, begin=0,
+ end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/grid-rcnn_r50_fpn_gn-head_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/grid-rcnn_r50_fpn_gn-head_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..228fca2323ceec2052a3835089d987a2643c53c1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/grid-rcnn_r50_fpn_gn-head_2x_coco.py
@@ -0,0 +1,160 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py', '../_base_/default_runtime.py'
+]
+# model settings
+model = dict(
+ type='GridRCNN',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5),
+ rpn_head=dict(
+ type='RPNHead',
+ in_channels=256,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ scales=[8],
+ ratios=[0.5, 1.0, 2.0],
+ strides=[4, 8, 16, 32, 64]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0 / 9.0, loss_weight=1.0)),
+ roi_head=dict(
+ type='GridRoIHead',
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=7, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ bbox_head=dict(
+ type='Shared2FCBBoxHead',
+ with_reg=False,
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=False),
+ grid_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=14, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ grid_head=dict(
+ type='GridHead',
+ grid_points=9,
+ num_convs=8,
+ in_channels=256,
+ point_feat_channels=64,
+ norm_cfg=dict(type='GN', num_groups=36),
+ loss_grid=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=15))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=0,
+ pos_weight=-1,
+ debug=False),
+ rpn_proposal=dict(
+ nms_pre=2000,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ pos_radius=1,
+ pos_weight=-1,
+ max_num_grid=192,
+ debug=False)),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=1000,
+ max_per_img=1000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ score_thr=0.03,
+ nms=dict(type='nms', iou_threshold=0.3),
+ max_per_img=100)))
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.0001))
+
+# training schedule
+max_epochs = 25
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 80,
+ by_epoch=False,
+ begin=0,
+ end=3665),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[17, 23],
+ gamma=0.1)
+]
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/grid-rcnn_x101-32x4d_fpn_gn-head_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/grid-rcnn_x101-32x4d_fpn_gn-head_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..dddf157beb6667887d0cd920cb2803e340d43183
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/grid-rcnn_x101-32x4d_fpn_gn-head_2x_coco.py
@@ -0,0 +1,13 @@
+_base_ = './grid-rcnn_r50_fpn_gn-head_2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/grid-rcnn_x101-64x4d_fpn_gn-head_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/grid-rcnn_x101-64x4d_fpn_gn-head_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e4ff50f546ae660cf398c2cb1c6f67ca20848c0f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/grid-rcnn_x101-64x4d_fpn_gn-head_2x_coco.py
@@ -0,0 +1,13 @@
+_base_ = './grid-rcnn_x101-32x4d_fpn_gn-head_2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..cee91e3b88e7bafa27e705713f2bc45d0dc872d0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grid_rcnn/metafile.yml
@@ -0,0 +1,101 @@
+Collections:
+ - Name: Grid R-CNN
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RPN
+ - Dilated Convolution
+ - ResNet
+ - RoIAlign
+ Paper:
+ URL: https://arxiv.org/abs/1906.05688
+ Title: 'Grid R-CNN'
+ README: configs/grid_rcnn/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/detectors/grid_rcnn.py#L6
+ Version: v2.0.0
+
+Models:
+ - Name: grid-rcnn_r50_fpn_gn-head_2x_coco
+ In Collection: Grid R-CNN
+ Config: configs/grid_rcnn/grid-rcnn_r50_fpn_gn-head_2x_coco.py
+ Metadata:
+ Training Memory (GB): 5.1
+ inference time (ms/im):
+ - value: 66.67
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/grid_rcnn/grid_rcnn_r50_fpn_gn-head_2x_coco/grid_rcnn_r50_fpn_gn-head_2x_coco_20200130-6cca8223.pth
+
+ - Name: grid-rcnn_r101_fpn_gn-head_2x_coco
+ In Collection: Grid R-CNN
+ Config: configs/grid_rcnn/grid-rcnn_r101_fpn_gn-head_2x_coco.py
+ Metadata:
+ Training Memory (GB): 7.0
+ inference time (ms/im):
+ - value: 79.37
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/grid_rcnn/grid_rcnn_r101_fpn_gn-head_2x_coco/grid_rcnn_r101_fpn_gn-head_2x_coco_20200309-d6eca030.pth
+
+ - Name: grid-rcnn_x101-32x4d_fpn_gn-head_2x_coco
+ In Collection: Grid R-CNN
+ Config: configs/grid_rcnn/grid-rcnn_x101-32x4d_fpn_gn-head_2x_coco.py
+ Metadata:
+ Training Memory (GB): 8.3
+ inference time (ms/im):
+ - value: 92.59
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/grid_rcnn/grid_rcnn_x101_32x4d_fpn_gn-head_2x_coco/grid_rcnn_x101_32x4d_fpn_gn-head_2x_coco_20200130-d8f0e3ff.pth
+
+ - Name: grid-rcnn_x101-64x4d_fpn_gn-head_2x_coco
+ In Collection: Grid R-CNN
+ Config: configs/grid_rcnn/grid-rcnn_x101-64x4d_fpn_gn-head_2x_coco.py
+ Metadata:
+ Training Memory (GB): 11.3
+ inference time (ms/im):
+ - value: 129.87
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/grid_rcnn/grid_rcnn_x101_64x4d_fpn_gn-head_2x_coco/grid_rcnn_x101_64x4d_fpn_gn-head_2x_coco_20200204-ec76a754.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..9792df93c1e9093d298467ee3037991c09fd1dae
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/README.md
@@ -0,0 +1,72 @@
+# GRoIE
+
+> [A novel Region of Interest Extraction Layer for Instance Segmentation](https://arxiv.org/abs/2004.13665)
+
+
+
+## Abstract
+
+Given the wide diffusion of deep neural network architectures for computer vision tasks, several new applications are nowadays more and more feasible. Among them, a particular attention has been recently given to instance segmentation, by exploiting the results achievable by two-stage networks (such as Mask R-CNN or Faster R-CNN), derived from R-CNN. In these complex architectures, a crucial role is played by the Region of Interest (RoI) extraction layer, devoted to extracting a coherent subset of features from a single Feature Pyramid Network (FPN) layer attached on top of a backbone.
+This paper is motivated by the need to overcome the limitations of existing RoI extractors which select only one (the best) layer from FPN. Our intuition is that all the layers of FPN retain useful information. Therefore, the proposed layer (called Generic RoI Extractor - GRoIE) introduces non-local building blocks and attention mechanisms to boost the performance.
+A comprehensive ablation study at component level is conducted to find the best set of algorithms and parameters for the GRoIE layer. Moreover, GRoIE can be integrated seamlessly with every two-stage architecture for both object detection and instance segmentation tasks. Therefore, the improvements brought about by the use of GRoIE in different state-of-the-art architectures are also evaluated. The proposed layer leads up to gain a 1.1% AP improvement on bounding box detection and 1.7% AP improvement on instance segmentation.
+
+
+

+
+
+## Introduction
+
+By Leonardo Rossi, Akbar Karimi and Andrea Prati from
+[IMPLab](http://implab.ce.unipr.it/).
+
+We provide configs to reproduce the results in the paper for
+"*A novel Region of Interest Extraction Layer for Instance Segmentation*"
+on COCO object detection.
+
+This paper is motivated by the need to overcome to the limitations of existing
+RoI extractors which select only one (the best) layer from FPN.
+
+Our intuition is that all the layers of FPN retain useful information.
+
+Therefore, the proposed layer (called Generic RoI Extractor - **GRoIE**)
+introduces non-local building blocks and attention mechanisms to boost the
+performance.
+
+## Results and Models
+
+The results on COCO 2017 minival (5k images) are shown in the below table.
+
+### Application of GRoIE to different architectures
+
+| Backbone | Method | Lr schd | box AP | mask AP | Config | Download |
+| :-------: | :-------------: | :-----: | :----: | :-----: | :------------------------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | Faster Original | 1x | 37.4 | | [config](../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_1x_coco/faster_rcnn_r50_fpn_1x_coco_20200130-047c8118.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_1x_coco/faster_rcnn_r50_fpn_1x_coco_20200130_204655.log.json) |
+| R-50-FPN | + GRoIE | 1x | 38.3 | | [config](./faste-rcnn_r50_fpn_groie_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/groie/faster_rcnn_r50_fpn_groie_1x_coco/faster_rcnn_r50_fpn_groie_1x_coco_20200604_211715-66ee9516.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/groie/faster_rcnn_r50_fpn_groie_1x_coco/faster_rcnn_r50_fpn_groie_1x_coco_20200604_211715.log.json) |
+| R-50-FPN | Grid R-CNN | 1x | 39.1 | | [config](./grid-rcnn_r50_fpn_gn-head-groie_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/groie/grid_rcnn_r50_fpn_gn-head_groie_1x_coco/grid_rcnn_r50_fpn_gn-head_groie_1x_coco_20200605_202059-4b75d86f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/groie/grid_rcnn_r50_fpn_gn-head_groie_1x_coco/grid_rcnn_r50_fpn_gn-head_groie_1x_coco_20200605_202059.log.json) |
+| R-50-FPN | + GRoIE | 1x | | | [config](./grid-rcnn_r50_fpn_gn-head-groie_1x_coco.py) | |
+| R-50-FPN | Mask R-CNN | 1x | 38.2 | 34.7 | [config](../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_fpn_1x_coco/mask_rcnn_r50_fpn_1x_coco_20200205-d4b0c5d6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_fpn_1x_coco/mask_rcnn_r50_fpn_1x_coco_20200205_050542.log.json) |
+| R-50-FPN | + GRoIE | 1x | 39.0 | 36.0 | [config](./mask-rcnn_r50_fpn_groie_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/groie/mask_rcnn_r50_fpn_groie_1x_coco/mask_rcnn_r50_fpn_groie_1x_coco_20200604_211715-50d90c74.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/groie/mask_rcnn_r50_fpn_groie_1x_coco/mask_rcnn_r50_fpn_groie_1x_coco_20200604_211715.log.json) |
+| R-50-FPN | GC-Net | 1x | 40.7 | 36.5 | [config](../gcnet/mask-rcnn_r50-syncbn-gcb-r4-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r50_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco/mask_rcnn_r50_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco_20200202-50b90e5c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r50_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco/mask_rcnn_r50_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco_20200202_085547.log.json) |
+| R-50-FPN | + GRoIE | 1x | 41.0 | 37.8 | [config](./mask-rcnn_r50_fpn_syncbn-r4-gcb-c3-c5-groie_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/groie/mask_rcnn_r50_fpn_syncbn-backbone_r4_gcb_c3-c5_groie_1x_coco/mask_rcnn_r50_fpn_syncbn-backbone_r4_gcb_c3-c5_groie_1x_coco_20200604_211715-42eb79e1.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/groie/mask_rcnn_r50_fpn_syncbn-backbone_r4_gcb_c3-c5_groie_1x_coco/mask_rcnn_r50_fpn_syncbn-backbone_r4_gcb_c3-c5_groie_1x_coco_20200604_211715-42eb79e1.pth) |
+| R-101-FPN | GC-Net | 1x | 42.2 | 37.8 | [config](../gcnet/mask-rcnn_r101-syncbn-gcb-r4-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r101_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco/mask_rcnn_r101_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco_20200206-8407a3f0.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/gcnet/mask_rcnn_r101_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco/mask_rcnn_r101_fpn_syncbn-backbone_r4_gcb_c3-c5_1x_coco_20200206_142508.log.json) |
+| R-101-FPN | + GRoIE | 1x | 42.6 | 38.7 | [config](./mask-rcnn_r101_fpn_syncbn-r4-gcb_c3-c5-groie_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/groie/mask_rcnn_r101_fpn_syncbn-backbone_r4_gcb_c3-c5_groie_1x_coco/mask_rcnn_r101_fpn_syncbn-backbone_r4_gcb_c3-c5_groie_1x_coco_20200607_224507-8daae01c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/groie/mask_rcnn_r101_fpn_syncbn-backbone_r4_gcb_c3-c5_groie_1x_coco/mask_rcnn_r101_fpn_syncbn-backbone_r4_gcb_c3-c5_groie_1x_coco_20200607_224507.log.json) |
+
+## Citation
+
+If you use this work or benchmark in your research, please cite this project.
+
+```latex
+@inproceedings{rossi2021novel,
+ title={A novel region of interest extraction layer for instance segmentation},
+ author={Rossi, Leonardo and Karimi, Akbar and Prati, Andrea},
+ booktitle={2020 25th International Conference on Pattern Recognition (ICPR)},
+ pages={2203--2209},
+ year={2021},
+ organization={IEEE}
+}
+```
+
+## Contact
+
+The implementation of GRoIE is currently maintained by
+[Leonardo Rossi](https://github.com/hachreak/).
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/faste-rcnn_r50_fpn_groie_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/faste-rcnn_r50_fpn_groie_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..0fbe8a32c3a81e9b312a02f79f3495171387d9f0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/faste-rcnn_r50_fpn_groie_1x_coco.py
@@ -0,0 +1,25 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+# model settings
+model = dict(
+ roi_head=dict(
+ bbox_roi_extractor=dict(
+ type='GenericRoIExtractor',
+ aggregation='sum',
+ roi_layer=dict(type='RoIAlign', output_size=7, sampling_ratio=2),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32],
+ pre_cfg=dict(
+ type='ConvModule',
+ in_channels=256,
+ out_channels=256,
+ kernel_size=5,
+ padding=2,
+ inplace=False,
+ ),
+ post_cfg=dict(
+ type='GeneralizedAttention',
+ in_channels=256,
+ spatial_range=-1,
+ num_heads=6,
+ attention_type='0100',
+ kv_stride=2))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/grid-rcnn_r50_fpn_gn-head-groie_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/grid-rcnn_r50_fpn_gn-head-groie_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..dadccb79c2288f16eb4a1fa33269e4a8f5a55c9b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/grid-rcnn_r50_fpn_gn-head-groie_1x_coco.py
@@ -0,0 +1,45 @@
+_base_ = '../grid_rcnn/grid-rcnn_r50_fpn_gn-head_1x_coco.py'
+# model settings
+model = dict(
+ roi_head=dict(
+ bbox_roi_extractor=dict(
+ type='GenericRoIExtractor',
+ aggregation='sum',
+ roi_layer=dict(type='RoIAlign', output_size=7, sampling_ratio=2),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32],
+ pre_cfg=dict(
+ type='ConvModule',
+ in_channels=256,
+ out_channels=256,
+ kernel_size=5,
+ padding=2,
+ inplace=False,
+ ),
+ post_cfg=dict(
+ type='GeneralizedAttention',
+ in_channels=256,
+ spatial_range=-1,
+ num_heads=6,
+ attention_type='0100',
+ kv_stride=2)),
+ grid_roi_extractor=dict(
+ type='GenericRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=14, sampling_ratio=2),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32],
+ pre_cfg=dict(
+ type='ConvModule',
+ in_channels=256,
+ out_channels=256,
+ kernel_size=5,
+ padding=2,
+ inplace=False,
+ ),
+ post_cfg=dict(
+ type='GeneralizedAttention',
+ in_channels=256,
+ spatial_range=-1,
+ num_heads=6,
+ attention_type='0100',
+ kv_stride=2))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/mask-rcnn_r101_fpn_syncbn-r4-gcb_c3-c5-groie_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/mask-rcnn_r101_fpn_syncbn-r4-gcb_c3-c5-groie_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5699b4284a76fe633afd81acb0b047a81df6afd2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/mask-rcnn_r101_fpn_syncbn-r4-gcb_c3-c5-groie_1x_coco.py
@@ -0,0 +1,45 @@
+_base_ = '../gcnet/mask-rcnn_r101-syncbn-gcb-r4-c3-c5_fpn_1x_coco.py'
+# model settings
+model = dict(
+ roi_head=dict(
+ bbox_roi_extractor=dict(
+ type='GenericRoIExtractor',
+ aggregation='sum',
+ roi_layer=dict(type='RoIAlign', output_size=7, sampling_ratio=2),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32],
+ pre_cfg=dict(
+ type='ConvModule',
+ in_channels=256,
+ out_channels=256,
+ kernel_size=5,
+ padding=2,
+ inplace=False,
+ ),
+ post_cfg=dict(
+ type='GeneralizedAttention',
+ in_channels=256,
+ spatial_range=-1,
+ num_heads=6,
+ attention_type='0100',
+ kv_stride=2)),
+ mask_roi_extractor=dict(
+ type='GenericRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=14, sampling_ratio=2),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32],
+ pre_cfg=dict(
+ type='ConvModule',
+ in_channels=256,
+ out_channels=256,
+ kernel_size=5,
+ padding=2,
+ inplace=False,
+ ),
+ post_cfg=dict(
+ type='GeneralizedAttention',
+ in_channels=256,
+ spatial_range=-1,
+ num_heads=6,
+ attention_type='0100',
+ kv_stride=2))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/mask-rcnn_r50_fpn_groie_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/mask-rcnn_r50_fpn_groie_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..4c9521e2f5730b74efc51f2051f861bfe5f8192d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/mask-rcnn_r50_fpn_groie_1x_coco.py
@@ -0,0 +1,45 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+# model settings
+model = dict(
+ roi_head=dict(
+ bbox_roi_extractor=dict(
+ type='GenericRoIExtractor',
+ aggregation='sum',
+ roi_layer=dict(type='RoIAlign', output_size=7, sampling_ratio=2),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32],
+ pre_cfg=dict(
+ type='ConvModule',
+ in_channels=256,
+ out_channels=256,
+ kernel_size=5,
+ padding=2,
+ inplace=False,
+ ),
+ post_cfg=dict(
+ type='GeneralizedAttention',
+ in_channels=256,
+ spatial_range=-1,
+ num_heads=6,
+ attention_type='0100',
+ kv_stride=2)),
+ mask_roi_extractor=dict(
+ type='GenericRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=14, sampling_ratio=2),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32],
+ pre_cfg=dict(
+ type='ConvModule',
+ in_channels=256,
+ out_channels=256,
+ kernel_size=5,
+ padding=2,
+ inplace=False,
+ ),
+ post_cfg=dict(
+ type='GeneralizedAttention',
+ in_channels=256,
+ spatial_range=-1,
+ num_heads=6,
+ attention_type='0100',
+ kv_stride=2))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/mask-rcnn_r50_fpn_syncbn-r4-gcb-c3-c5-groie_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/mask-rcnn_r50_fpn_syncbn-r4-gcb-c3-c5-groie_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..22e97b6959a0bd13ae4432c806c61ca3d899f9ea
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/mask-rcnn_r50_fpn_syncbn-r4-gcb-c3-c5-groie_1x_coco.py
@@ -0,0 +1,45 @@
+_base_ = '../gcnet/mask-rcnn_r50-syncbn-gcb-r4-c3-c5_fpn_1x_coco.py'
+# model settings
+model = dict(
+ roi_head=dict(
+ bbox_roi_extractor=dict(
+ type='GenericRoIExtractor',
+ aggregation='sum',
+ roi_layer=dict(type='RoIAlign', output_size=7, sampling_ratio=2),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32],
+ pre_cfg=dict(
+ type='ConvModule',
+ in_channels=256,
+ out_channels=256,
+ kernel_size=5,
+ padding=2,
+ inplace=False,
+ ),
+ post_cfg=dict(
+ type='GeneralizedAttention',
+ in_channels=256,
+ spatial_range=-1,
+ num_heads=6,
+ attention_type='0100',
+ kv_stride=2)),
+ mask_roi_extractor=dict(
+ type='GenericRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=14, sampling_ratio=2),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32],
+ pre_cfg=dict(
+ type='ConvModule',
+ in_channels=256,
+ out_channels=256,
+ kernel_size=5,
+ padding=2,
+ inplace=False,
+ ),
+ post_cfg=dict(
+ type='GeneralizedAttention',
+ in_channels=256,
+ spatial_range=-1,
+ num_heads=6,
+ attention_type='0100',
+ kv_stride=2))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..ce957004719cb542a51c48e7e07a3d94d6bdee18
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/groie/metafile.yml
@@ -0,0 +1,94 @@
+Collections:
+ - Name: GRoIE
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Generic RoI Extractor
+ - FPN
+ - RPN
+ - ResNet
+ - RoIAlign
+ Paper:
+ URL: https://arxiv.org/abs/2004.13665
+ Title: 'A novel Region of Interest Extraction Layer for Instance Segmentation'
+ README: configs/groie/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/roi_heads/roi_extractors/groie.py#L15
+ Version: v2.1.0
+
+Models:
+ - Name: faster-rcnn_r50_fpn_groie_1x_coco
+ In Collection: GRoIE
+ Config: configs/groie/faste-rcnn_r50_fpn_groie_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/groie/faster_rcnn_r50_fpn_groie_1x_coco/faster_rcnn_r50_fpn_groie_1x_coco_20200604_211715-66ee9516.pth
+
+ - Name: grid-rcnn_r50_fpn_gn-head-groie_1x_coco
+ In Collection: GRoIE
+ Config: configs/groie/grid-rcnn_r50_fpn_gn-head-groie_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/groie/grid_rcnn_r50_fpn_gn-head_groie_1x_coco/grid_rcnn_r50_fpn_gn-head_groie_1x_coco_20200605_202059-4b75d86f.pth
+
+ - Name: mask-rcnn_r50_fpn_groie_1x_coco
+ In Collection: GRoIE
+ Config: configs/groie/mask-rcnn_r50_fpn_groie_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/groie/mask_rcnn_r50_fpn_groie_1x_coco/mask_rcnn_r50_fpn_groie_1x_coco_20200604_211715-50d90c74.pth
+
+ - Name: mask-rcnn_r50_fpn_syncbn-backbone_r4_gcb_c3-c5_groie_1x_coco
+ In Collection: GRoIE
+ Config: configs/groie/mask-rcnn_r50_fpn_syncbn-r4-gcb-c3-c5-groie_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/groie/mask_rcnn_r50_fpn_syncbn-backbone_r4_gcb_c3-c5_groie_1x_coco/mask_rcnn_r50_fpn_syncbn-backbone_r4_gcb_c3-c5_groie_1x_coco_20200604_211715-42eb79e1.pth
+
+ - Name: mask-rcnn_r101_fpn_syncbn-r4-gcb_c3-c5-groie_1x_coco
+ In Collection: GRoIE
+ Config: configs/groie/mask-rcnn_r101_fpn_syncbn-r4-gcb_c3-c5-groie_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.6
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/groie/mask_rcnn_r101_fpn_syncbn-backbone_r4_gcb_c3-c5_groie_1x_coco/mask_rcnn_r101_fpn_syncbn-backbone_r4_gcb_c3-c5_groie_1x_coco_20200607_224507-8daae01c.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..2a527828a467df069bbdbe624b55c1afcaa3521f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/README.md
@@ -0,0 +1,317 @@
+# Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
+
+[Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection](https://arxiv.org/abs/2303.05499)
+
+
+
+## Abstract
+
+In this paper, we present an open-set object detector, called Grounding DINO, by marrying Transformer-based detector DINO with grounded pre-training, which can detect arbitrary objects with human inputs such as category names or referring expressions. The key solution of open-set object detection is introducing language to a closed-set detector for open-set concept generalization. To effectively fuse language and vision modalities, we conceptually divide a closed-set detector into three phases and propose a tight fusion solution, which includes a feature enhancer, a language-guided query selection, and a cross-modality decoder for cross-modality fusion. While previous works mainly evaluate open-set object detection on novel categories, we propose to also perform evaluations on referring expression comprehension for objects specified with attributes. Grounding DINO performs remarkably well on all three settings, including benchmarks on COCO, LVIS, ODinW, and RefCOCO/+/g. Grounding DINO achieves a 52.5 AP on the COCO detection zero-shot transfer benchmark, i.e., without any training data from COCO. It sets a new record on the ODinW zero-shot benchmark with a mean 26.1 AP.
+
+
+

+
+
+## Installation
+
+```shell
+cd $MMDETROOT
+
+# source installation
+pip install -r requirements/multimodal.txt
+
+# or mim installation
+mim install mmdet[multimodal]
+```
+
+## NOTE
+
+Grounding DINO utilizes BERT as the language model, which requires access to https://huggingface.co/. If you encounter connection errors due to network access, you can download the required files on a computer with internet access and save them locally. Finally, modify the `lang_model_name` field in the config to the local path. Please refer to the following code:
+
+```python
+from transformers import BertConfig, BertModel
+from transformers import AutoTokenizer
+
+config = BertConfig.from_pretrained("bert-base-uncased")
+model = BertModel.from_pretrained("bert-base-uncased", add_pooling_layer=False, config=config)
+tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
+
+config.save_pretrained("your path/bert-base-uncased")
+model.save_pretrained("your path/bert-base-uncased")
+tokenizer.save_pretrained("your path/bert-base-uncased")
+```
+
+## Inference
+
+```
+cd $MMDETROOT
+
+wget https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/groundingdino_swint_ogc_mmdet-822d7e9d.pth
+
+python demo/image_demo.py \
+ demo/demo.jpg \
+ configs/grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_cap4m.py \
+ --weights groundingdino_swint_ogc_mmdet-822d7e9d.pth \
+ --texts 'bench . car .'
+```
+
+
+

+
+
+## COCO Results and Models
+
+| Model | Backbone | Style | COCO mAP | Official COCO mAP | Pre-Train Data | Config | Download |
+| :----------------: | :------: | :-------: | :--------: | :---------------: | :----------------------------------------------: | :------------------------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Grounding DINO-T | Swin-T | Zero-shot | 48.5 | 48.4 | O365,GoldG,Cap4M | [config](grounding_dino_swin-t_pretrain_obj365_goldg_cap4m.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/groundingdino_swint_ogc_mmdet-822d7e9d.pth) |
+| Grounding DINO-T | Swin-T | Finetune | 58.1(+0.9) | 57.2 | O365,GoldG,Cap4M | [config](grounding_dino_swin-t_finetune_16xb2_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/grounding_dino_swin-t_finetune_16xb2_1x_coco/grounding_dino_swin-t_finetune_16xb2_1x_coco_20230921_152544-5f234b20.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/grounding_dino_swin-t_finetune_16xb2_1x_coco/grounding_dino_swin-t_finetune_16xb2_1x_coco_20230921_152544.log.json) |
+| Grounding DINO-B | Swin-B | Zero-shot | 56.9 | 56.7 | COCO,O365,GoldG,Cap4M,OpenImage,ODinW-35,RefCOCO | [config](grounding_dino_swin-b_pretrain_mixeddata.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/groundingdino_swinb_cogcoor_mmdet-55949c9c.pth) |
+| Grounding DINO-B | Swin-B | Finetune | 59.7 | | COCO,O365,GoldG,Cap4M,OpenImage,ODinW-35,RefCOCO | [config](grounding_dino_swin-b_finetune_16xb2_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/grounding_dino_swin-b_finetune_16xb2_1x_coco/grounding_dino_swin-b_finetune_16xb2_1x_coco_20230921_153201-f219e0c0.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/grounding_dino_swin-b_finetune_16xb2_1x_coco/grounding_dino_swin-b_finetune_16xb2_1x_coco_20230921_153201.log.json) |
+| Grounding DINO-R50 | R50 | Scratch | 48.9(+0.8) | 48.1 | | [config](grounding_dino_r50_scratch_8xb2_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/grounding_dino_r50_scratch_8xb2_1x_coco/grounding_dino_r50_scratch_1x_coco-fe0002f2.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/grounding_dino_r50_scratch_8xb2_1x_coco/20230922_114218.json) |
+
+Note:
+
+1. The weights corresponding to the zero-shot model are adopted from the official weights and converted using the [script](../../tools/model_converters/groundingdino_to_mmdet.py). We have not retrained the model for the time being.
+2. Finetune refers to fine-tuning on the COCO 2017 dataset. The R50 model is trained using 8 NVIDIA GeForce 3090 GPUs, while the remaining models are trained using 16 NVIDIA GeForce 3090 GPUs. The GPU memory usage is approximately 8.5GB.
+3. Our performance is higher than the official model due to two reasons: we modified the initialization strategy and introduced a log scaler.
+
+## LVIS Results
+
+| Model | MiniVal APr | MiniVal APc | MiniVal APf | MiniVal AP | Val1.0 APr | Val1.0 APc | Val1.0 APf | Val1.0 AP | Pre-Train Data | Config | Download |
+| :--------------: | :---------: | :---------: | :---------: | :--------: | :--------: | :--------: | :--------: | :-------: | :----------------------------------------------: | :-----------------------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------: |
+| Grounding DINO-T | 18.8 | 24.2 | 34.7 | 28.8 | 10.1 | 15.3 | 29.9 | 20.1 | O365,GoldG,Cap4M | [config](lvis/grounding_dino_swin-t_pretrain_zeroshot_mini-lvis.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/groundingdino_swint_ogc_mmdet-822d7e9d.pth) |
+| Grounding DINO-B | 27.9 | 33.4 | 37.2 | 34.7 | 19.0 | 24.1 | 32.9 | 26.7 | COCO,O365,GoldG,Cap4M,OpenImage,ODinW-35,RefCOCO | [config](lvis/grounding_dino_swin-b_pretrain_zeroshot_mini-lvis.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/groundingdino_swinb_cogcoor_mmdet-55949c9c.pth) |
+
+Note:
+
+1. The above are zero-shot evaluation results.
+2. The evaluation metric we used is LVIS FixAP. For specific details, please refer to [Evaluating Large-Vocabulary Object Detectors: The Devil is in the Details](https://arxiv.org/pdf/2102.01066.pdf).
+
+## ODinW (Object Detection in the Wild) Results
+
+Learning visual representations from natural language supervision has recently shown great promise in a number of pioneering works. In general, these language-augmented visual models demonstrate strong transferability to a variety of datasets and tasks. However, it remains challenging to evaluate the transferablity of these models due to the lack of easy-to-use evaluation toolkits and public benchmarks. To tackle this, we build ELEVATER 1 , the first benchmark and toolkit for evaluating (pre-trained) language-augmented visual models. ELEVATER is composed of three components. (i) Datasets. As downstream evaluation suites, it consists of 20 image classification datasets and 35 object detection datasets, each of which is augmented with external knowledge. (ii) Toolkit. An automatic hyper-parameter tuning toolkit is developed to facilitate model evaluation on downstream tasks. (iii) Metrics. A variety of evaluation metrics are used to measure sample-efficiency (zero-shot and few-shot) and parameter-efficiency (linear probing and full model fine-tuning). ELEVATER is platform for Computer Vision in the Wild (CVinW), and is publicly released at https://computer-vision-in-the-wild.github.io/ELEVATER/
+
+### Results and models of ODinW13
+
+| Method | GLIP-T(A) | Official | GLIP-T(B) | Official | GLIP-T(C) | Official | GroundingDINO-T | GroundingDINO-B |
+| --------------------- | --------- | --------- | --------- | --------- | --------- | --------- | --------------- | --------------- |
+| AerialMaritimeDrone | 0.123 | 0.122 | 0.110 | 0.110 | 0.130 | 0.130 | 0.173 | 0.281 |
+| Aquarium | 0.175 | 0.174 | 0.173 | 0.169 | 0.191 | 0.190 | 0.195 | 0.445 |
+| CottontailRabbits | 0.686 | 0.686 | 0.688 | 0.688 | 0.744 | 0.744 | 0.799 | 0.808 |
+| EgoHands | 0.013 | 0.013 | 0.003 | 0.004 | 0.314 | 0.315 | 0.608 | 0.764 |
+| NorthAmericaMushrooms | 0.502 | 0.502 | 0.367 | 0.367 | 0.297 | 0.296 | 0.507 | 0.675 |
+| Packages | 0.589 | 0.589 | 0.083 | 0.083 | 0.699 | 0.699 | 0.687 | 0.670 |
+| PascalVOC | 0.512 | 0.512 | 0.541 | 0.540 | 0.565 | 0.565 | 0.563 | 0.711 |
+| pistols | 0.339 | 0.339 | 0.502 | 0.501 | 0.503 | 0.504 | 0.726 | 0.771 |
+| pothole | 0.007 | 0.007 | 0.030 | 0.030 | 0.058 | 0.058 | 0.215 | 0.478 |
+| Raccoon | 0.075 | 0.074 | 0.285 | 0.288 | 0.241 | 0.244 | 0.549 | 0.541 |
+| ShellfishOpenImages | 0.253 | 0.253 | 0.337 | 0.338 | 0.300 | 0.302 | 0.393 | 0.650 |
+| thermalDogsAndPeople | 0.372 | 0.372 | 0.475 | 0.475 | 0.510 | 0.510 | 0.657 | 0.633 |
+| VehiclesOpenImages | 0.574 | 0.566 | 0.562 | 0.547 | 0.549 | 0.534 | 0.613 | 0.647 |
+| Average | **0.325** | **0.324** | **0.320** | **0.318** | **0.392** | **0.392** | **0.514** | **0.621** |
+
+### Results and models of ODinW35
+
+| Method | GLIP-T(A) | Official | GLIP-T(B) | Official | GLIP-T(C) | Official | GroundingDINO-T | GroundingDINO-B |
+| --------------------------- | --------- | --------- | --------- | --------- | --------- | --------- | --------------- | --------------- |
+| AerialMaritimeDrone_large | 0.123 | 0.122 | 0.110 | 0.110 | 0.130 | 0.130 | 0.173 | 0.281 |
+| AerialMaritimeDrone_tiled | 0.174 | 0.174 | 0.172 | 0.172 | 0.172 | 0.172 | 0.206 | 0.364 |
+| AmericanSignLanguageLetters | 0.001 | 0.001 | 0.003 | 0.003 | 0.009 | 0.009 | 0.002 | 0.096 |
+| Aquarium | 0.175 | 0.175 | 0.173 | 0.171 | 0.192 | 0.182 | 0.195 | 0.445 |
+| BCCD | 0.016 | 0.016 | 0.001 | 0.001 | 0.000 | 0.000 | 0.161 | 0.584 |
+| boggleBoards | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.134 |
+| brackishUnderwater | 0.016 | 0..013 | 0.021 | 0.027 | 0.020 | 0.022 | 0.021 | 0.454 |
+| ChessPieces | 0.001 | 0.001 | 0.000 | 0.000 | 0.001 | 0.001 | 0.000 | 0.000 |
+| CottontailRabbits | 0.710 | 0.709 | 0.683 | 0.683 | 0.752 | 0.752 | 0.806 | 0.797 |
+| dice | 0.005 | 0.005 | 0.004 | 0.004 | 0.004 | 0.004 | 0.004 | 0.082 |
+| DroneControl | 0.016 | 0.017 | 0.006 | 0.008 | 0.005 | 0.007 | 0.042 | 0.638 |
+| EgoHands_generic | 0.009 | 0.010 | 0.005 | 0.006 | 0.510 | 0.508 | 0.608 | 0.764 |
+| EgoHands_specific | 0.001 | 0.001 | 0.004 | 0.006 | 0.003 | 0.004 | 0.002 | 0.687 |
+| HardHatWorkers | 0.029 | 0.029 | 0.023 | 0.023 | 0.033 | 0.033 | 0.046 | 0.439 |
+| MaskWearing | 0.007 | 0.007 | 0.003 | 0.002 | 0.005 | 0.005 | 0.004 | 0.406 |
+| MountainDewCommercial | 0.218 | 0.227 | 0.199 | 0.197 | 0.478 | 0.463 | 0.430 | 0.580 |
+| NorthAmericaMushrooms | 0.502 | 0.502 | 0.450 | 0.450 | 0.497 | 0.497 | 0.471 | 0.501 |
+| openPoetryVision | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 | 0.051 |
+| OxfordPets_by_breed | 0.001 | 0.002 | 0.002 | 0.004 | 0.001 | 0.002 | 0.003 | 0.799 |
+| OxfordPets_by_species | 0.016 | 0.011 | 0.012 | 0.009 | 0.013 | 0.009 | 0.011 | 0.872 |
+| PKLot | 0.002 | 0.002 | 0.000 | 0.000 | 0.000 | 0.000 | 0.001 | 0.774 |
+| Packages | 0.569 | 0.569 | 0.279 | 0.279 | 0.712 | 0.712 | 0.695 | 0.728 |
+| PascalVOC | 0.512 | 0.512 | 0.541 | 0.540 | 0.565 | 0.565 | 0.563 | 0.711 |
+| pistols | 0.339 | 0.339 | 0.502 | 0.501 | 0.503 | 0.504 | 0.726 | 0.771 |
+| plantdoc | 0.002 | 0.002 | 0.007 | 0.007 | 0.009 | 0.009 | 0.005 | 0.376 |
+| pothole | 0.007 | 0.010 | 0.024 | 0.025 | 0.085 | 0.101 | 0.215 | 0.478 |
+| Raccoons | 0.075 | 0.074 | 0.285 | 0.288 | 0.241 | 0.244 | 0.549 | 0.541 |
+| selfdrivingCar | 0.071 | 0.072 | 0.074 | 0.074 | 0.081 | 0.080 | 0.089 | 0.318 |
+| ShellfishOpenImages | 0.253 | 0.253 | 0.337 | 0.338 | 0.300 | 0.302 | 0.393 | 0.650 |
+| ThermalCheetah | 0.028 | 0.028 | 0.000 | 0.000 | 0.028 | 0.028 | 0.087 | 0.290 |
+| thermalDogsAndPeople | 0.372 | 0.372 | 0.475 | 0.475 | 0.510 | 0.510 | 0.657 | 0.633 |
+| UnoCards | 0.000 | 0.000 | 0.000 | 0.001 | 0.002 | 0.003 | 0.006 | 0.754 |
+| VehiclesOpenImages | 0.574 | 0.566 | 0.562 | 0.547 | 0.549 | 0.534 | 0.613 | 0.647 |
+| WildfireSmoke | 0.000 | 0.000 | 0.000 | 0.000 | 0.017 | 0.017 | 0.134 | 0.410 |
+| websiteScreenshots | 0.003 | 0.004 | 0.003 | 0.005 | 0.005 | 0.006 | 0.012 | 0.175 |
+| Average | **0.134** | **0.134** | **0.138** | **0.138** | **0.179** | **0.178** | **0.227** | **0.492** |
+
+## Flickr30k Results
+
+| Model | Pre-Train Data | Val R@1 | Val R@5 | Val R@10 | Tesst R@1 | Test R@5 | Test R@10 | Config | Download |
+| :--------------: | :--------------: | ------- | ------- | -------- | --------- | -------- | --------- | :-------------------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Grounding DINO-T | O365,GoldG,Cap4M | 87.8 | 96.6 | 98.0 | 88.1 | 96.9 | 98.2 | [config](grounding_dino_swin-t_finetune_16xb2_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/grounding_dino_swin-t_finetune_16xb2_1x_coco/grounding_dino_swin-t_finetune_16xb2_1x_coco_20230921_152544-5f234b20.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/grounding_dino_swin-t_finetune_16xb2_1x_coco/grounding_dino_swin-t_finetune_16xb2_1x_coco_20230921_152544.log.json) |
+
+Note:
+
+1. `@1,5,10` refers to precision at the top 1, 5, and 10 positions in a predicted ranked list.
+2. The pretraining data used by Grounding DINO-T is `O365,GoldG,Cap4M`, and the corresponding evaluation configuration is (grounding_dino_swin-t_pretrain_zeroshot_refcoco)\[refcoco/grounding_dino_swin-t_pretrain_zeroshot_refcoco.py\].
+
+Test Command
+
+```shell
+cd mmdetection
+bash tools/dist_test.sh configs/grounding_dino/flickr30k/grounding_dino_swin-t-pretrain_zeroshot_flickr30k.py checkpoints/groundingdino_swint_ogc_mmdet-822d7e9d.pth 8
+```
+
+## Referring Expression Comprehension Results
+
+| Method | Grounding DINO-T
(O365,GoldG,Cap4M) | Grounding DINO-B
(COCO,O365,GoldG,Cap4M,OpenImage,ODinW-35,RefCOCO) |
+| --------------------------------------- | ----------------------------------------- | ------------------------------------------------------------------------- |
+| RefCOCO val @1,5,10 | 50.77/89.45/94.86 | 84.61/97.88/99.10 |
+| RefCOCO testA @1,5,10 | 57.45/91.29/95.62 | 88.65/98.89/99.63 |
+| RefCOCO testB @1,5,10 | 44.97/86.54/92.88 | 80.51/96.64/98.51 |
+| RefCOCO+ val @1,5,10 | 51.64/86.35/92.57 | 73.67/96.60/98.65 |
+| RefCOCO+ testA @1,5,10 | 57.25/86.74/92.65 | 82.19/97.92/99.09 |
+| RefCOCO+ testB @1,5,10 | 46.35/84.05/90.67 | 64.10/94.25/97.46 |
+| RefCOCOg val @1,5,10 | 60.42/92.10/96.18 | 78.33/97.28/98.57 |
+| RefCOCOg test @1,5,10 | 59.74/92.08/96.28 | 78.11/97.06/98.65 |
+| gRefCOCO val Pr@(F1=1, IoU≥0.5),N-acc | 41.32/91.82 | 46.18/81.44 |
+| gRefCOCO testA Pr@(F1=1, IoU≥0.5),N-acc | 27.23/90.24 | 38.60/76.06 |
+| gRefCOCO testB Pr@(F1=1, IoU≥0.5),N-acc | 29.70/93.49 | 35.87/80.58 |
+
+Note:
+
+1. `@1,5,10` refers to precision at the top 1, 5, and 10 positions in a predicted ranked list.
+2. `Pr@(F1=1, IoU≥0.5),N-acc` from the paper [GREC: Generalized Referring Expression Comprehension](https://arxiv.org/pdf/2308.16182.pdf)
+3. The pretraining data used by Grounding DINO-T is `O365,GoldG,Cap4M`, and the corresponding evaluation configuration is (grounding_dino_swin-t_pretrain_zeroshot_refcoco)\[refcoco/grounding_dino_swin-t_pretrain_zeroshot_refcoco.py\].
+4. The pretraining data used by Grounding DINO-B is `COCO,O365,GoldG,Cap4M,OpenImage,ODinW-35,RefCOCO`, and the corresponding evaluation configuration is (grounding_dino_swin-t_pretrain_zeroshot_refcoco)\[refcoco/grounding_dino_swin-b_pretrain_zeroshot_refcoco.py\].
+
+Test Command
+
+```shell
+cd mmdetection
+./tools/dist_test.sh configs/grounding_dino/refcoco/grounding_dino_swin-t_pretrain_zeroshot_refexp.py https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/groundingdino_swint_ogc_mmdet-822d7e9d.pth 8
+./tools/dist_test.sh configs/grounding_dino/refcoco/grounding_dino_swin-b_pretrain_zeroshot_refexp.py https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/groundingdino_swinb_cogcoor_mmdet-55949c9c.pth 8
+```
+
+## Description Detection Dataset
+
+```shell
+pip install ddd-dataset
+```
+
+| Method | mode | Grounding DINO-T
(O365,GoldG,Cap4M) | Grounding DINO-B
(COCO,O365,GoldG,Cap4M,OpenImage,ODinW-35,RefCOCO) |
+| -------------------------------- | -------- | ----------------------------------------- | ------------------------------------------------------------------------- |
+| FULL/short/middle/long/very long | concat | 17.2/18.0/18.7/14.8/16.3 | 20.2/20.4/21.1/18.8/19.8 |
+| FULL/short/middle/long/very long | parallel | 22.3/28.2/24.8/19.1/13.9 | 25.0/26.4/27.2/23.5/19.7 |
+| PRES/short/middle/long/very long | concat | 17.8/18.3/19.2/15.2/17.3 | 20.7/21.7/21.4/19.1/20.3 |
+| PRES/short/middle/long/very long | parallel | 21.0/27.0/22.8/17.5/12.5 | 23.7/25.8/25.1/21.9/19.3 |
+| ABS/short/middle/long/very long | concat | 15.4/17.1/16.4/13.6/14.9 | 18.6/16.1/19.7/18.1/19.1 |
+| ABS/short/middle/long/very long | parallel | 26.0/32.0/33.0/23.6/15.5 | 28.8/28.1/35.8/28.2/20.2 |
+
+Note:
+
+1. Considering that the evaluation time for Inter-scenario is very long and the performance is low, it is temporarily not supported. The mentioned metrics are for Intra-scenario.
+2. `concat` is the default inference mode for Grounding DINO, where it concatenates multiple sub-sentences with "." to form a single sentence for inference. On the other hand, "parallel" performs inference on each sub-sentence in a for-loop.
+
+## Custom Dataset
+
+To facilitate fine-tuning on custom datasets, we use a simple cat dataset as an example, as shown in the following steps.
+
+### 1. Dataset Preparation
+
+```shell
+cd mmdetection
+wget https://download.openmmlab.com/mmyolo/data/cat_dataset.zip
+unzip cat_dataset.zip -d data/cat/
+```
+
+cat dataset is a single-category dataset with 144 images, which has been converted to coco format.
+
+
+

+
+
+### 2. Config Preparation
+
+Due to the simplicity and small number of cat datasets, we use 8 cards to train 20 epochs, scale the learning rate accordingly, and do not train the language model, only the visual model.
+
+The Details of the configuration can be found in [grounding_dino_swin-t_finetune_8xb2_20e_cat](grounding_dino_swin-t_finetune_8xb2_20e_cat.py)
+
+### 3. Visualization and Evaluation
+
+Due to the Grounding DINO is an open detection model, so it can be detected and evaluated even if it is not trained on the cat dataset.
+
+The single image visualization is as follows:
+
+```shell
+cd mmdetection
+python demo/image_demo.py data/cat/images/IMG_20211205_120756.jpg configs/grounding_dino/grounding_dino_swin-t_finetune_8xb2_20e_cat.py --weights https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/groundingdino_swint_ogc_mmdet-822d7e9d.pth --texts cat.
+```
+
+
+

+
+
+The test dataset evaluation on single card is as follows:
+
+```shell
+python tools/test.py configs/grounding_dino/grounding_dino_swin-t_finetune_8xb2_20e_cat.py https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/groundingdino_swint_ogc_mmdet-822d7e9d.pth
+```
+
+```text
+ Average Precision (AP) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.867
+ Average Precision (AP) @[ IoU=0.50 | area= all | maxDets=1000 ] = 1.000
+ Average Precision (AP) @[ IoU=0.75 | area= all | maxDets=1000 ] = 0.931
+ Average Precision (AP) @[ IoU=0.50:0.95 | area= small | maxDets=1000 ] = -1.000
+ Average Precision (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=1000 ] = -1.000
+ Average Precision (AP) @[ IoU=0.50:0.95 | area= large | maxDets=1000 ] = 0.867
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.903
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=300 ] = 0.907
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=1000 ] = 0.907
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= small | maxDets=1000 ] = -1.000
+ Average Recall (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=1000 ] = -1.000
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= large | maxDets=1000 ] = 0.907
+```
+
+### 4. Model Training and Visualization
+
+```shell
+./tools/dist_train.sh configs/grounding_dino/grounding_dino_swin-t_finetune_8xb2_20e_cat.py 8 --work-dir cat_work_dir
+```
+
+The model will be saved based on the best performance on the test set. The performance of the best model (at epoch 16) is as follows:
+
+```text
+ Average Precision (AP) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.905
+ Average Precision (AP) @[ IoU=0.50 | area= all | maxDets=1000 ] = 1.000
+ Average Precision (AP) @[ IoU=0.75 | area= all | maxDets=1000 ] = 0.923
+ Average Precision (AP) @[ IoU=0.50:0.95 | area= small | maxDets=1000 ] = -1.000
+ Average Precision (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=1000 ] = -1.000
+ Average Precision (AP) @[ IoU=0.50:0.95 | area= large | maxDets=1000 ] = 0.905
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.927
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=300 ] = 0.937
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=1000 ] = 0.937
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= small | maxDets=1000 ] = -1.000
+ Average Recall (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=1000 ] = -1.000
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= large | maxDets=1000 ] = 0.937
+```
+
+We can find that after fine-tuning training, the training of the cat dataset is increased from 86.7 to 90.5.
+
+If we do single image inference visualization again, the result is as follows:
+
+```shell
+cd mmdetection
+python demo/image_demo.py data/cat/images/IMG_20211205_120756.jpg configs/grounding_dino/grounding_dino_swin-t_finetune_8xb2_20e_cat.py --weights cat_work_dir/best_coco_bbox_mAP_epoch_16.pth --texts cat.
+```
+
+
+

+
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/dod/grounding_dino_swin-b_pretrain_zeroshot_concat_dod.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/dod/grounding_dino_swin-b_pretrain_zeroshot_concat_dod.py
new file mode 100644
index 0000000000000000000000000000000000000000..ac655b74aa664ef912b6b1f509e4eb9341ccd62a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/dod/grounding_dino_swin-b_pretrain_zeroshot_concat_dod.py
@@ -0,0 +1,14 @@
+_base_ = 'grounding_dino_swin-t_pretrain_zeroshot_concat_dod.py'
+
+model = dict(
+ type='GroundingDINO',
+ backbone=dict(
+ pretrain_img_size=384,
+ embed_dims=128,
+ depths=[2, 2, 18, 2],
+ num_heads=[4, 8, 16, 32],
+ window_size=12,
+ drop_path_rate=0.3,
+ patch_norm=True),
+ neck=dict(in_channels=[256, 512, 1024]),
+)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/dod/grounding_dino_swin-b_pretrain_zeroshot_parallel_dod.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/dod/grounding_dino_swin-b_pretrain_zeroshot_parallel_dod.py
new file mode 100644
index 0000000000000000000000000000000000000000..9a1c8f2ac740c6c64a01a1a6a8f7dd57622bedf6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/dod/grounding_dino_swin-b_pretrain_zeroshot_parallel_dod.py
@@ -0,0 +1,3 @@
+_base_ = 'grounding_dino_swin-b_pretrain_zeroshot_concat_dod.py'
+
+model = dict(test_cfg=dict(chunked_size=1))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/dod/grounding_dino_swin-t_pretrain_zeroshot_concat_dod.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/dod/grounding_dino_swin-t_pretrain_zeroshot_concat_dod.py
new file mode 100644
index 0000000000000000000000000000000000000000..bb418011bf489c259f3696589aa56c5b8296256c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/dod/grounding_dino_swin-t_pretrain_zeroshot_concat_dod.py
@@ -0,0 +1,78 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365_goldg_cap4m.py'
+
+data_root = 'data/d3/'
+
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile', backend_args=None,
+ imdecode_backend='pillow'),
+ dict(
+ type='FixScaleResize',
+ scale=(800, 1333),
+ keep_ratio=True,
+ backend='pillow'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'text', 'custom_entities', 'sent_ids'))
+]
+
+# -------------------------------------------------#
+val_dataset_full = dict(
+ type='DODDataset',
+ data_root=data_root,
+ ann_file='d3_json/d3_full_annotations.json',
+ data_prefix=dict(img='d3_images/', anno='d3_pkl'),
+ pipeline=test_pipeline,
+ test_mode=True,
+ backend_args=None,
+ return_classes=True)
+
+val_evaluator_full = dict(
+ type='DODCocoMetric',
+ ann_file=data_root + 'd3_json/d3_full_annotations.json')
+
+# -------------------------------------------------#
+val_dataset_pres = dict(
+ type='DODDataset',
+ data_root=data_root,
+ ann_file='d3_json/d3_pres_annotations.json',
+ data_prefix=dict(img='d3_images/', anno='d3_pkl'),
+ pipeline=test_pipeline,
+ test_mode=True,
+ backend_args=None,
+ return_classes=True)
+val_evaluator_pres = dict(
+ type='DODCocoMetric',
+ ann_file=data_root + 'd3_json/d3_pres_annotations.json')
+
+# -------------------------------------------------#
+val_dataset_abs = dict(
+ type='DODDataset',
+ data_root=data_root,
+ ann_file='d3_json/d3_abs_annotations.json',
+ data_prefix=dict(img='d3_images/', anno='d3_pkl'),
+ pipeline=test_pipeline,
+ test_mode=True,
+ backend_args=None,
+ return_classes=True)
+val_evaluator_abs = dict(
+ type='DODCocoMetric',
+ ann_file=data_root + 'd3_json/d3_abs_annotations.json')
+
+# -------------------------------------------------#
+datasets = [val_dataset_full, val_dataset_pres, val_dataset_abs]
+dataset_prefixes = ['FULL', 'PRES', 'ABS']
+metrics = [val_evaluator_full, val_evaluator_pres, val_evaluator_abs]
+
+val_dataloader = dict(
+ dataset=dict(_delete_=True, type='ConcatDataset', datasets=datasets))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='MultiDatasetsEvaluator',
+ metrics=metrics,
+ dataset_prefixes=dataset_prefixes)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/dod/grounding_dino_swin-t_pretrain_zeroshot_parallel_dod.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/dod/grounding_dino_swin-t_pretrain_zeroshot_parallel_dod.py
new file mode 100644
index 0000000000000000000000000000000000000000..3d680091162e5ac96c15c76b58a18764e85d3233
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/dod/grounding_dino_swin-t_pretrain_zeroshot_parallel_dod.py
@@ -0,0 +1,3 @@
+_base_ = 'grounding_dino_swin-t_pretrain_zeroshot_concat_dod.py'
+
+model = dict(test_cfg=dict(chunked_size=1))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/flickr30k/grounding_dino_swin-t-pretrain_zeroshot_flickr30k.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/flickr30k/grounding_dino_swin-t-pretrain_zeroshot_flickr30k.py
new file mode 100644
index 0000000000000000000000000000000000000000..c1996567588842f82c0af83e3a9ab84c81e7c25d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/flickr30k/grounding_dino_swin-t-pretrain_zeroshot_flickr30k.py
@@ -0,0 +1,57 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365_goldg_cap4m.py'
+
+dataset_type = 'Flickr30kDataset'
+data_root = 'data/flickr30k_entities/'
+
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile', backend_args=None,
+ imdecode_backend='pillow'),
+ dict(
+ type='FixScaleResize',
+ scale=(800, 1333),
+ keep_ratio=True,
+ backend='pillow'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'text', 'custom_entities',
+ 'tokens_positive', 'phrase_ids', 'phrases'))
+]
+
+dataset_Flickr30k_val = dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='final_flickr_separateGT_val.json',
+ data_prefix=dict(img='flickr30k_images/'),
+ pipeline=test_pipeline,
+)
+
+dataset_Flickr30k_test = dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='final_flickr_separateGT_test.json',
+ data_prefix=dict(img='flickr30k_images/'),
+ pipeline=test_pipeline,
+)
+
+val_evaluator_Flickr30k = dict(type='Flickr30kMetric')
+
+test_evaluator_Flickr30k = dict(type='Flickr30kMetric')
+
+# ----------Config---------- #
+dataset_prefixes = ['Flickr30kVal', 'Flickr30kTest']
+datasets = [dataset_Flickr30k_val, dataset_Flickr30k_test]
+metrics = [val_evaluator_Flickr30k, test_evaluator_Flickr30k]
+
+val_dataloader = dict(
+ dataset=dict(_delete_=True, type='ConcatDataset', datasets=datasets))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='MultiDatasetsEvaluator',
+ metrics=metrics,
+ dataset_prefixes=dataset_prefixes)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/grounding_dino_r50_scratch_8xb2_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/grounding_dino_r50_scratch_8xb2_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..623a29b87adfd6734e980e814766e873b2b89d05
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/grounding_dino_r50_scratch_8xb2_1x_coco.py
@@ -0,0 +1,208 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+lang_model_name = 'bert-base-uncased'
+
+model = dict(
+ type='GroundingDINO',
+ num_queries=900,
+ with_box_refine=True,
+ as_two_stage=True,
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_mask=False,
+ ),
+ language_model=dict(
+ type='BertModel',
+ name=lang_model_name,
+ pad_to_max=False,
+ use_sub_sentence_represent=True,
+ special_tokens_list=['[CLS]', '[SEP]', '.', '?'],
+ add_pooling_layer=False,
+ ),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='ChannelMapper',
+ in_channels=[512, 1024, 2048],
+ kernel_size=1,
+ out_channels=256,
+ act_cfg=None,
+ bias=True,
+ norm_cfg=dict(type='GN', num_groups=32),
+ num_outs=4),
+ encoder=dict(
+ num_layers=6,
+ num_cp=6,
+ # visual layer config
+ layer_cfg=dict(
+ self_attn_cfg=dict(embed_dims=256, num_levels=4, dropout=0.0),
+ ffn_cfg=dict(
+ embed_dims=256, feedforward_channels=2048, ffn_drop=0.0)),
+ # text layer config
+ text_layer_cfg=dict(
+ self_attn_cfg=dict(num_heads=4, embed_dims=256, dropout=0.0),
+ ffn_cfg=dict(
+ embed_dims=256, feedforward_channels=1024, ffn_drop=0.0)),
+ # fusion layer config
+ fusion_layer_cfg=dict(
+ v_dim=256,
+ l_dim=256,
+ embed_dim=1024,
+ num_heads=4,
+ init_values=1e-4),
+ ),
+ decoder=dict(
+ num_layers=6,
+ return_intermediate=True,
+ layer_cfg=dict(
+ # query self attention layer
+ self_attn_cfg=dict(embed_dims=256, num_heads=8, dropout=0.0),
+ # cross attention layer query to text
+ cross_attn_text_cfg=dict(embed_dims=256, num_heads=8, dropout=0.0),
+ # cross attention layer query to image
+ cross_attn_cfg=dict(embed_dims=256, num_heads=8, dropout=0.0),
+ ffn_cfg=dict(
+ embed_dims=256, feedforward_channels=2048, ffn_drop=0.0)),
+ post_norm_cfg=None),
+ positional_encoding=dict(
+ num_feats=128, normalize=True, offset=0.0, temperature=20),
+ bbox_head=dict(
+ type='GroundingDINOHead',
+ num_classes=80,
+ sync_cls_avg_factor=True,
+ contrastive_cfg=dict(max_text_len=256, log_scale='auto', bias=True),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0), # 2.0 in DeformDETR
+ loss_bbox=dict(type='L1Loss', loss_weight=5.0),
+ loss_iou=dict(type='GIoULoss', loss_weight=2.0)),
+ dn_cfg=dict( # TODO: Move to model.train_cfg ?
+ label_noise_scale=0.5,
+ box_noise_scale=1.0, # 0.4 for DN-DETR
+ group_cfg=dict(dynamic=True, num_groups=None,
+ num_dn_queries=100)), # TODO: half num_dn_queries
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='HungarianAssigner',
+ match_costs=[
+ dict(type='BinaryFocalLossCost', weight=2.0),
+ dict(type='BBoxL1Cost', weight=5.0, box_format='xywh'),
+ dict(type='IoUCost', iou_mode='giou', weight=2.0)
+ ])),
+ test_cfg=dict(max_per_img=300))
+
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities'))
+]
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='FixScaleResize', scale=(800, 1333), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'text', 'custom_entities'))
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=train_pipeline,
+ return_classes=True))
+val_dataloader = dict(
+ dataset=dict(pipeline=test_pipeline, return_classes=True))
+test_dataloader = val_dataloader
+
+# We did not adopt the official 24e optimizer strategy
+# because the results indicate that the current strategy is superior.
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(
+ type='AdamW',
+ lr=0.0001, # 0.0002 for DeformDETR
+ weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'backbone': dict(lr_mult=0.1)
+ }))
+# learning policy
+max_epochs = 12
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[11],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/grounding_dino_swin-b_finetune_16xb2_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/grounding_dino_swin-b_finetune_16xb2_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3554ee245ffe4312fc7f2cdd83755b1a0731aab9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/grounding_dino_swin-b_finetune_16xb2_1x_coco.py
@@ -0,0 +1,17 @@
+_base_ = [
+ './grounding_dino_swin-t_finetune_16xb2_1x_coco.py',
+]
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/groundingdino_swinb_cogcoor_mmdet-55949c9c.pth' # noqa
+model = dict(
+ type='GroundingDINO',
+ backbone=dict(
+ pretrain_img_size=384,
+ embed_dims=128,
+ depths=[2, 2, 18, 2],
+ num_heads=[4, 8, 16, 32],
+ window_size=12,
+ drop_path_rate=0.3,
+ patch_norm=True),
+ neck=dict(in_channels=[256, 512, 1024]),
+)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/grounding_dino_swin-b_pretrain_mixeddata.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/grounding_dino_swin-b_pretrain_mixeddata.py
new file mode 100644
index 0000000000000000000000000000000000000000..92f327fef8311f0f72d7f75149bfc163863e913c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/grounding_dino_swin-b_pretrain_mixeddata.py
@@ -0,0 +1,16 @@
+_base_ = [
+ './grounding_dino_swin-t_pretrain_obj365_goldg_cap4m.py',
+]
+
+model = dict(
+ type='GroundingDINO',
+ backbone=dict(
+ pretrain_img_size=384,
+ embed_dims=128,
+ depths=[2, 2, 18, 2],
+ num_heads=[4, 8, 16, 32],
+ window_size=12,
+ drop_path_rate=0.3,
+ patch_norm=True),
+ neck=dict(in_channels=[256, 512, 1024]),
+)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/grounding_dino_swin-t_finetune_16xb2_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/grounding_dino_swin-t_finetune_16xb2_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..0c6403ee66d9e5782723117191176efbadec2a90
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/grounding_dino_swin-t_finetune_16xb2_1x_coco.py
@@ -0,0 +1,204 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/groundingdino_swint_ogc_mmdet-822d7e9d.pth' # noqa
+lang_model_name = 'bert-base-uncased'
+
+model = dict(
+ type='GroundingDINO',
+ num_queries=900,
+ with_box_refine=True,
+ as_two_stage=True,
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_mask=False,
+ ),
+ language_model=dict(
+ type='BertModel',
+ name=lang_model_name,
+ pad_to_max=False,
+ use_sub_sentence_represent=True,
+ special_tokens_list=['[CLS]', '[SEP]', '.', '?'],
+ add_pooling_layer=False,
+ ),
+ backbone=dict(
+ type='SwinTransformer',
+ embed_dims=96,
+ depths=[2, 2, 6, 2],
+ num_heads=[3, 6, 12, 24],
+ window_size=7,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.2,
+ patch_norm=True,
+ out_indices=(1, 2, 3),
+ with_cp=True,
+ convert_weights=False),
+ neck=dict(
+ type='ChannelMapper',
+ in_channels=[192, 384, 768],
+ kernel_size=1,
+ out_channels=256,
+ act_cfg=None,
+ bias=True,
+ norm_cfg=dict(type='GN', num_groups=32),
+ num_outs=4),
+ encoder=dict(
+ num_layers=6,
+ num_cp=6,
+ # visual layer config
+ layer_cfg=dict(
+ self_attn_cfg=dict(embed_dims=256, num_levels=4, dropout=0.0),
+ ffn_cfg=dict(
+ embed_dims=256, feedforward_channels=2048, ffn_drop=0.0)),
+ # text layer config
+ text_layer_cfg=dict(
+ self_attn_cfg=dict(num_heads=4, embed_dims=256, dropout=0.0),
+ ffn_cfg=dict(
+ embed_dims=256, feedforward_channels=1024, ffn_drop=0.0)),
+ # fusion layer config
+ fusion_layer_cfg=dict(
+ v_dim=256,
+ l_dim=256,
+ embed_dim=1024,
+ num_heads=4,
+ init_values=1e-4),
+ ),
+ decoder=dict(
+ num_layers=6,
+ return_intermediate=True,
+ layer_cfg=dict(
+ # query self attention layer
+ self_attn_cfg=dict(embed_dims=256, num_heads=8, dropout=0.0),
+ # cross attention layer query to text
+ cross_attn_text_cfg=dict(embed_dims=256, num_heads=8, dropout=0.0),
+ # cross attention layer query to image
+ cross_attn_cfg=dict(embed_dims=256, num_heads=8, dropout=0.0),
+ ffn_cfg=dict(
+ embed_dims=256, feedforward_channels=2048, ffn_drop=0.0)),
+ post_norm_cfg=None),
+ positional_encoding=dict(
+ num_feats=128, normalize=True, offset=0.0, temperature=20),
+ bbox_head=dict(
+ type='GroundingDINOHead',
+ num_classes=80,
+ sync_cls_avg_factor=True,
+ contrastive_cfg=dict(max_text_len=256, log_scale=0.0, bias=False),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0), # 2.0 in DeformDETR
+ loss_bbox=dict(type='L1Loss', loss_weight=5.0),
+ loss_iou=dict(type='GIoULoss', loss_weight=2.0)),
+ dn_cfg=dict( # TODO: Move to model.train_cfg ?
+ label_noise_scale=0.5,
+ box_noise_scale=1.0, # 0.4 for DN-DETR
+ group_cfg=dict(dynamic=True, num_groups=None,
+ num_dn_queries=100)), # TODO: half num_dn_queries
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='HungarianAssigner',
+ match_costs=[
+ dict(type='BinaryFocalLossCost', weight=2.0),
+ dict(type='BBoxL1Cost', weight=5.0, box_format='xywh'),
+ dict(type='IoUCost', iou_mode='giou', weight=2.0)
+ ])),
+ test_cfg=dict(max_per_img=300))
+
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities'))
+]
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='FixScaleResize', scale=(800, 1333), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'text', 'custom_entities'))
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=train_pipeline,
+ return_classes=True))
+val_dataloader = dict(
+ dataset=dict(pipeline=test_pipeline, return_classes=True))
+test_dataloader = val_dataloader
+
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0001, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'backbone': dict(lr_mult=0.1)
+ }))
+# learning policy
+max_epochs = 12
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[11],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (16 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=32)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/grounding_dino_swin-t_finetune_8xb2_20e_cat.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/grounding_dino_swin-t_finetune_8xb2_20e_cat.py
new file mode 100644
index 0000000000000000000000000000000000000000..c2265e86730f68ed69af246a5e0e87fa2cb5e570
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/grounding_dino_swin-t_finetune_8xb2_20e_cat.py
@@ -0,0 +1,56 @@
+_base_ = 'grounding_dino_swin-t_finetune_16xb2_1x_coco.py'
+
+data_root = 'data/cat/'
+class_name = ('cat', )
+num_classes = len(class_name)
+metainfo = dict(classes=class_name, palette=[(220, 20, 60)])
+
+model = dict(bbox_head=dict(num_classes=num_classes))
+
+train_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ metainfo=metainfo,
+ ann_file='annotations/trainval.json',
+ data_prefix=dict(img='images/')))
+
+val_dataloader = dict(
+ dataset=dict(
+ metainfo=metainfo,
+ data_root=data_root,
+ ann_file='annotations/test.json',
+ data_prefix=dict(img='images/')))
+
+test_dataloader = val_dataloader
+
+val_evaluator = dict(ann_file=data_root + 'annotations/test.json')
+test_evaluator = val_evaluator
+
+max_epoch = 20
+
+default_hooks = dict(
+ checkpoint=dict(interval=1, max_keep_ckpts=1, save_best='auto'),
+ logger=dict(type='LoggerHook', interval=5))
+train_cfg = dict(max_epochs=max_epoch, val_interval=1)
+
+param_scheduler = [
+ dict(type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=30),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epoch,
+ by_epoch=True,
+ milestones=[15],
+ gamma=0.1)
+]
+
+optim_wrapper = dict(
+ optimizer=dict(lr=0.00005),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'backbone': dict(lr_mult=0.1),
+ 'language_model': dict(lr_mult=0),
+ }))
+
+auto_scale_lr = dict(base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_cap4m.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_cap4m.py
new file mode 100644
index 0000000000000000000000000000000000000000..7448764ef7ed4fb91bbca981e8006b412e74c414
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_cap4m.py
@@ -0,0 +1,128 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+lang_model_name = 'bert-base-uncased'
+
+model = dict(
+ type='GroundingDINO',
+ num_queries=900,
+ with_box_refine=True,
+ as_two_stage=True,
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_mask=False,
+ ),
+ language_model=dict(
+ type='BertModel',
+ name=lang_model_name,
+ pad_to_max=False,
+ use_sub_sentence_represent=True,
+ special_tokens_list=['[CLS]', '[SEP]', '.', '?'],
+ add_pooling_layer=True,
+ ),
+ backbone=dict(
+ type='SwinTransformer',
+ embed_dims=96,
+ depths=[2, 2, 6, 2],
+ num_heads=[3, 6, 12, 24],
+ window_size=7,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.2,
+ patch_norm=True,
+ out_indices=(1, 2, 3),
+ with_cp=False,
+ convert_weights=False),
+ neck=dict(
+ type='ChannelMapper',
+ in_channels=[192, 384, 768],
+ kernel_size=1,
+ out_channels=256,
+ act_cfg=None,
+ bias=True,
+ norm_cfg=dict(type='GN', num_groups=32),
+ num_outs=4),
+ encoder=dict(
+ num_layers=6,
+ # visual layer config
+ layer_cfg=dict(
+ self_attn_cfg=dict(embed_dims=256, num_levels=4, dropout=0.0),
+ ffn_cfg=dict(
+ embed_dims=256, feedforward_channels=2048, ffn_drop=0.0)),
+ # text layer config
+ text_layer_cfg=dict(
+ self_attn_cfg=dict(num_heads=4, embed_dims=256, dropout=0.0),
+ ffn_cfg=dict(
+ embed_dims=256, feedforward_channels=1024, ffn_drop=0.0)),
+ # fusion layer config
+ fusion_layer_cfg=dict(
+ v_dim=256,
+ l_dim=256,
+ embed_dim=1024,
+ num_heads=4,
+ init_values=1e-4),
+ ),
+ decoder=dict(
+ num_layers=6,
+ return_intermediate=True,
+ layer_cfg=dict(
+ # query self attention layer
+ self_attn_cfg=dict(embed_dims=256, num_heads=8, dropout=0.0),
+ # cross attention layer query to text
+ cross_attn_text_cfg=dict(embed_dims=256, num_heads=8, dropout=0.0),
+ # cross attention layer query to image
+ cross_attn_cfg=dict(embed_dims=256, num_heads=8, dropout=0.0),
+ ffn_cfg=dict(
+ embed_dims=256, feedforward_channels=2048, ffn_drop=0.0)),
+ post_norm_cfg=None),
+ positional_encoding=dict(
+ num_feats=128, normalize=True, offset=0.0, temperature=20),
+ bbox_head=dict(
+ type='GroundingDINOHead',
+ num_classes=80,
+ sync_cls_avg_factor=True,
+ contrastive_cfg=dict(max_text_len=256),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0), # 2.0 in DeformDETR
+ loss_bbox=dict(type='L1Loss', loss_weight=5.0)),
+ dn_cfg=dict( # TODO: Move to model.train_cfg ?
+ label_noise_scale=0.5,
+ box_noise_scale=1.0, # 0.4 for DN-DETR
+ group_cfg=dict(dynamic=True, num_groups=None,
+ num_dn_queries=100)), # TODO: half num_dn_queries
+ # training and testing settings
+ train_cfg=None,
+ test_cfg=dict(max_per_img=300))
+
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile', backend_args=None,
+ imdecode_backend='pillow'),
+ dict(
+ type='FixScaleResize',
+ scale=(800, 1333),
+ keep_ratio=True,
+ backend='pillow'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'text', 'custom_entities',
+ 'tokens_positive'))
+]
+
+val_dataloader = dict(
+ dataset=dict(pipeline=test_pipeline, return_classes=True))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/lvis/grounding_dino_swin-b_pretrain_zeroshot_lvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/lvis/grounding_dino_swin-b_pretrain_zeroshot_lvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..6084159044e8c0e8642a1226c6a9efd85c7d27d2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/lvis/grounding_dino_swin-b_pretrain_zeroshot_lvis.py
@@ -0,0 +1,14 @@
+_base_ = './grounding_dino_swin-t_pretrain_zeroshot_lvis.py'
+
+model = dict(
+ type='GroundingDINO',
+ backbone=dict(
+ pretrain_img_size=384,
+ embed_dims=128,
+ depths=[2, 2, 18, 2],
+ num_heads=[4, 8, 16, 32],
+ window_size=12,
+ drop_path_rate=0.3,
+ patch_norm=True),
+ neck=dict(in_channels=[256, 512, 1024]),
+)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/lvis/grounding_dino_swin-b_pretrain_zeroshot_mini-lvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/lvis/grounding_dino_swin-b_pretrain_zeroshot_mini-lvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..68467a7237ca893aa79eb5b0acc9d159f7082968
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/lvis/grounding_dino_swin-b_pretrain_zeroshot_mini-lvis.py
@@ -0,0 +1,14 @@
+_base_ = './grounding_dino_swin-t_pretrain_zeroshot_mini-lvis.py'
+
+model = dict(
+ type='GroundingDINO',
+ backbone=dict(
+ pretrain_img_size=384,
+ embed_dims=128,
+ depths=[2, 2, 18, 2],
+ num_heads=[4, 8, 16, 32],
+ window_size=12,
+ drop_path_rate=0.3,
+ patch_norm=True),
+ neck=dict(in_channels=[256, 512, 1024]),
+)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/lvis/grounding_dino_swin-t_pretrain_zeroshot_lvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/lvis/grounding_dino_swin-t_pretrain_zeroshot_lvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..3d05f0ce1c0cb095c0c9f9a65bd7666cba57afe7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/lvis/grounding_dino_swin-t_pretrain_zeroshot_lvis.py
@@ -0,0 +1,24 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365_goldg_cap4m.py'
+
+model = dict(test_cfg=dict(
+ max_per_img=300,
+ chunked_size=40,
+))
+
+dataset_type = 'LVISV1Dataset'
+data_root = 'data/coco/'
+
+val_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ type=dataset_type,
+ ann_file='annotations/lvis_od_val.json',
+ data_prefix=dict(img='')))
+test_dataloader = val_dataloader
+
+# numpy < 1.24.0
+val_evaluator = dict(
+ _delete_=True,
+ type='LVISFixedAPMetric',
+ ann_file=data_root + 'annotations/lvis_od_val.json')
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/lvis/grounding_dino_swin-t_pretrain_zeroshot_mini-lvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/lvis/grounding_dino_swin-t_pretrain_zeroshot_mini-lvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..0aac6cf33a92827c9c350175977bb1a595d2c0c8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/lvis/grounding_dino_swin-t_pretrain_zeroshot_mini-lvis.py
@@ -0,0 +1,25 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365_goldg_cap4m.py'
+
+model = dict(test_cfg=dict(
+ max_per_img=300,
+ chunked_size=40,
+))
+
+dataset_type = 'LVISV1Dataset'
+data_root = 'data/coco/'
+
+val_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ type=dataset_type,
+ ann_file='annotations/lvis_v1_minival_inserted_image_name.json',
+ data_prefix=dict(img='')))
+test_dataloader = val_dataloader
+
+# numpy < 1.24.0
+val_evaluator = dict(
+ _delete_=True,
+ type='LVISFixedAPMetric',
+ ann_file=data_root +
+ 'annotations/lvis_v1_minival_inserted_image_name.json')
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..dcb5ebf82846d3cfbc2fa345cc89468ba269fd84
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/metafile.yml
@@ -0,0 +1,67 @@
+Collections:
+ - Name: Grounding DINO
+ Metadata:
+ Training Data: Objects365, GoldG, CC3M and COCO
+ Training Techniques:
+ - AdamW
+ - Multi Scale Train
+ - Gradient Clip
+ Training Resources: 3090 GPUs
+ Architecture:
+ - Swin Transformer
+ - BERT
+ Paper:
+ URL: https://arxiv.org/abs/2303.05499
+ Title: 'Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
+'
+ README: configs/grounding_dino/README.md
+ Code:
+ URL:
+ Version: v3.0.0
+
+Models:
+ - Name: grounding_dino_swin-t_pretrain_obj365_goldg_cap4m
+ In Collection: Grounding DINO
+ Config: configs/grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_cap4m.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 48.5
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/groundingdino_swint_ogc_mmdet-822d7e9d.pth
+ - Name: grounding_dino_swin-b_pretrain_mixeddata
+ In Collection: Grounding DINO
+ Config: configs/grounding_dino/grounding_dino_swin-b_pretrain_mixeddata.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 56.9
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/groundingdino_swinb_cogcoor_mmdet-55949c9c.pth
+ - Name: grounding_dino_swin-t_finetune_16xb2_1x_coco
+ In Collection: Grounding DINO
+ Config: configs/grounding_dino/grounding_dino_swin-t_finetune_16xb2_1x_coco.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 58.1
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/grounding_dino_swin-t_finetune_16xb2_1x_coco/grounding_dino_swin-t_finetune_16xb2_1x_coco_20230921_152544-5f234b20.pth
+ - Name: grounding_dino_swin-b_finetune_16xb2_1x_coco
+ In Collection: Grounding DINO
+ Config: configs/grounding_dino/grounding_dino_swin-b_finetune_16xb2_1x_coco.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 59.7
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/grounding_dino_swin-b_finetune_16xb2_1x_coco/grounding_dino_swin-b_finetune_16xb2_1x_coco_20230921_153201-f219e0c0.pth
+ - Name: grounding_dino_r50_scratch_8xb2_1x_coco
+ In Collection: Grounding DINO
+ Config: configs/grounding_dino/grounding_dino_r50_scratch_8xb2_1x_coco.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 48.9
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/grounding_dino_r50_scratch_8xb2_1x_coco/grounding_dino_r50_scratch_1x_coco-fe0002f2.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/odinw/grounding_dino_swin-b_pretrain_odinw13.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/odinw/grounding_dino_swin-b_pretrain_odinw13.py
new file mode 100644
index 0000000000000000000000000000000000000000..65a6bc2a078a9ea5123c745aa72ba22466ea6e58
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/odinw/grounding_dino_swin-b_pretrain_odinw13.py
@@ -0,0 +1,338 @@
+_base_ = '../grounding_dino_swin-b_pretrain_mixeddata.py'
+
+dataset_type = 'CocoDataset'
+data_root = 'data/odinw/'
+
+base_test_pipeline = _base_.test_pipeline
+base_test_pipeline[-1]['meta_keys'] = ('img_id', 'img_path', 'ori_shape',
+ 'img_shape', 'scale_factor', 'text',
+ 'custom_entities', 'caption_prompt')
+
+# ---------------------1 AerialMaritimeDrone---------------------#
+class_name = ('boat', 'car', 'dock', 'jetski', 'lift')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'AerialMaritimeDrone/large/'
+dataset_AerialMaritimeDrone = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ test_mode=True,
+ pipeline=base_test_pipeline,
+ return_classes=True)
+val_evaluator_AerialMaritimeDrone = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------2 Aquarium---------------------#
+class_name = ('fish', 'jellyfish', 'penguin', 'puffin', 'shark', 'starfish',
+ 'stingray')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Aquarium/Aquarium Combined.v2-raw-1024.coco/'
+
+caption_prompt = None
+# caption_prompt = {
+# 'penguin': {
+# 'suffix': ', which is black and white'
+# },
+# 'puffin': {
+# 'suffix': ' with orange beaks'
+# },
+# 'stingray': {
+# 'suffix': ' which is flat and round'
+# },
+# }
+dataset_Aquarium = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Aquarium = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------3 CottontailRabbits---------------------#
+class_name = ('Cottontail-Rabbit', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'CottontailRabbits/'
+
+caption_prompt = None
+# caption_prompt = {'Cottontail-Rabbit': {'name': 'rabbit'}}
+
+dataset_CottontailRabbits = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_CottontailRabbits = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------4 EgoHands---------------------#
+class_name = ('hand', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'EgoHands/generic/'
+
+caption_prompt = None
+# caption_prompt = {'hand': {'suffix': ' of a person'}}
+
+dataset_EgoHands = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_EgoHands = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------5 NorthAmericaMushrooms---------------------#
+class_name = ('CoW', 'chanterelle')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'NorthAmericaMushrooms/North American Mushrooms.v1-416x416.coco/' # noqa
+
+caption_prompt = None
+# caption_prompt = {
+# 'CoW': {
+# 'name': 'flat mushroom'
+# },
+# 'chanterelle': {
+# 'name': 'yellow mushroom'
+# }
+# }
+
+dataset_NorthAmericaMushrooms = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_NorthAmericaMushrooms = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------6 Packages---------------------#
+class_name = ('package', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Packages/Raw/'
+
+caption_prompt = None
+# caption_prompt = {
+# 'package': {
+# 'prefix': 'there is a ',
+# 'suffix': ' on the porch'
+# }
+# }
+
+dataset_Packages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Packages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------7 PascalVOC---------------------#
+class_name = ('aeroplane', 'bicycle', 'bird', 'boat', 'bottle', 'bus', 'car',
+ 'cat', 'chair', 'cow', 'diningtable', 'dog', 'horse',
+ 'motorbike', 'person', 'pottedplant', 'sheep', 'sofa', 'train',
+ 'tvmonitor')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'PascalVOC/'
+dataset_PascalVOC = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_PascalVOC = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------8 pistols---------------------#
+class_name = ('pistol', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'pistols/export/'
+dataset_pistols = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_pistols = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------9 pothole---------------------#
+class_name = ('pothole', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'pothole/'
+
+caption_prompt = None
+# caption_prompt = {
+# 'pothole': {
+# 'prefix': 'there are some ',
+# 'name': 'holes',
+# 'suffix': ' on the road'
+# }
+# }
+
+dataset_pothole = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_pothole = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------10 Raccoon---------------------#
+class_name = ('raccoon', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Raccoon/Raccoon.v2-raw.coco/'
+dataset_Raccoon = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Raccoon = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------11 ShellfishOpenImages---------------------#
+class_name = ('Crab', 'Lobster', 'Shrimp')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'ShellfishOpenImages/raw/'
+dataset_ShellfishOpenImages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_ShellfishOpenImages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------12 thermalDogsAndPeople---------------------#
+class_name = ('dog', 'person')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'thermalDogsAndPeople/'
+dataset_thermalDogsAndPeople = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_thermalDogsAndPeople = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------13 VehiclesOpenImages---------------------#
+class_name = ('Ambulance', 'Bus', 'Car', 'Motorcycle', 'Truck')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'VehiclesOpenImages/416x416/'
+dataset_VehiclesOpenImages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_VehiclesOpenImages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# --------------------- Config---------------------#
+dataset_prefixes = [
+ 'AerialMaritimeDrone', 'Aquarium', 'CottontailRabbits', 'EgoHands',
+ 'NorthAmericaMushrooms', 'Packages', 'PascalVOC', 'pistols', 'pothole',
+ 'Raccoon', 'ShellfishOpenImages', 'thermalDogsAndPeople',
+ 'VehiclesOpenImages'
+]
+datasets = [
+ dataset_AerialMaritimeDrone, dataset_Aquarium, dataset_CottontailRabbits,
+ dataset_EgoHands, dataset_NorthAmericaMushrooms, dataset_Packages,
+ dataset_PascalVOC, dataset_pistols, dataset_pothole, dataset_Raccoon,
+ dataset_ShellfishOpenImages, dataset_thermalDogsAndPeople,
+ dataset_VehiclesOpenImages
+]
+metrics = [
+ val_evaluator_AerialMaritimeDrone, val_evaluator_Aquarium,
+ val_evaluator_CottontailRabbits, val_evaluator_EgoHands,
+ val_evaluator_NorthAmericaMushrooms, val_evaluator_Packages,
+ val_evaluator_PascalVOC, val_evaluator_pistols, val_evaluator_pothole,
+ val_evaluator_Raccoon, val_evaluator_ShellfishOpenImages,
+ val_evaluator_thermalDogsAndPeople, val_evaluator_VehiclesOpenImages
+]
+
+# -------------------------------------------------#
+val_dataloader = dict(
+ dataset=dict(_delete_=True, type='ConcatDataset', datasets=datasets))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='MultiDatasetsEvaluator',
+ metrics=metrics,
+ dataset_prefixes=dataset_prefixes)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/odinw/grounding_dino_swin-b_pretrain_odinw35.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/odinw/grounding_dino_swin-b_pretrain_odinw35.py
new file mode 100644
index 0000000000000000000000000000000000000000..e73cd8e61ba20f4baff6f7c85477a8fae3735e44
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/odinw/grounding_dino_swin-b_pretrain_odinw35.py
@@ -0,0 +1,796 @@
+_base_ = '../grounding_dino_swin-b_pretrain_mixeddata.py'
+
+dataset_type = 'CocoDataset'
+data_root = 'data/odinw/'
+
+base_test_pipeline = _base_.test_pipeline
+base_test_pipeline[-1]['meta_keys'] = ('img_id', 'img_path', 'ori_shape',
+ 'img_shape', 'scale_factor', 'text',
+ 'custom_entities', 'caption_prompt')
+
+# ---------------------1 AerialMaritimeDrone_large---------------------#
+class_name = ('boat', 'car', 'dock', 'jetski', 'lift')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'AerialMaritimeDrone/large/'
+dataset_AerialMaritimeDrone_large = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_AerialMaritimeDrone_large = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------2 AerialMaritimeDrone_tiled---------------------#
+class_name = ('boat', 'car', 'dock', 'jetski', 'lift')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'AerialMaritimeDrone/tiled/'
+dataset_AerialMaritimeDrone_tiled = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_AerialMaritimeDrone_tiled = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------3 AmericanSignLanguageLetters---------------------#
+class_name = ('A', 'B', 'C', 'D', 'E', 'F', 'G', 'H', 'I', 'J', 'K', 'L', 'M',
+ 'N', 'O', 'P', 'Q', 'R', 'S', 'T', 'U', 'V', 'W', 'X', 'Y', 'Z')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'AmericanSignLanguageLetters/American Sign Language Letters.v1-v1.coco/' # noqa
+dataset_AmericanSignLanguageLetters = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_AmericanSignLanguageLetters = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------4 Aquarium---------------------#
+class_name = ('fish', 'jellyfish', 'penguin', 'puffin', 'shark', 'starfish',
+ 'stingray')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Aquarium/Aquarium Combined.v2-raw-1024.coco/'
+dataset_Aquarium = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Aquarium = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------5 BCCD---------------------#
+class_name = ('Platelets', 'RBC', 'WBC')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'BCCD/BCCD.v3-raw.coco/'
+dataset_BCCD = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_BCCD = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------6 boggleBoards---------------------#
+class_name = ('Q', 'a', 'an', 'b', 'c', 'd', 'e', 'er', 'f', 'g', 'h', 'he',
+ 'i', 'in', 'j', 'k', 'l', 'm', 'n', 'o', 'o ', 'p', 'q', 'qu',
+ 'r', 's', 't', 't\\', 'th', 'u', 'v', 'w', 'wild', 'x', 'y', 'z')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'boggleBoards/416x416AutoOrient/export/'
+dataset_boggleBoards = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_boggleBoards = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------7 brackishUnderwater---------------------#
+class_name = ('crab', 'fish', 'jellyfish', 'shrimp', 'small_fish', 'starfish')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'brackishUnderwater/960x540/'
+dataset_brackishUnderwater = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_brackishUnderwater = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------8 ChessPieces---------------------#
+class_name = (' ', 'black bishop', 'black king', 'black knight', 'black pawn',
+ 'black queen', 'black rook', 'white bishop', 'white king',
+ 'white knight', 'white pawn', 'white queen', 'white rook')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'ChessPieces/Chess Pieces.v23-raw.coco/'
+dataset_ChessPieces = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/new_annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_ChessPieces = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/new_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------9 CottontailRabbits---------------------#
+class_name = ('rabbit', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'CottontailRabbits/'
+dataset_CottontailRabbits = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/new_annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_CottontailRabbits = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/new_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------10 dice---------------------#
+class_name = ('1', '2', '3', '4', '5', '6')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'dice/mediumColor/export/'
+dataset_dice = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_dice = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------11 DroneControl---------------------#
+class_name = ('follow', 'follow_hand', 'land', 'land_hand', 'null', 'object',
+ 'takeoff', 'takeoff-hand')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'DroneControl/Drone Control.v3-raw.coco/'
+dataset_DroneControl = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_DroneControl = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------12 EgoHands_generic---------------------#
+class_name = ('hand', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'EgoHands/generic/'
+caption_prompt = {'hand': {'suffix': ' of a person'}}
+dataset_EgoHands_generic = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ # NOTE w. prompt 0.548; wo. prompt 0.764
+ # caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_EgoHands_generic = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------13 EgoHands_specific---------------------#
+class_name = ('myleft', 'myright', 'yourleft', 'yourright')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'EgoHands/specific/'
+dataset_EgoHands_specific = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_EgoHands_specific = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------14 HardHatWorkers---------------------#
+class_name = ('head', 'helmet', 'person')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'HardHatWorkers/raw/'
+dataset_HardHatWorkers = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_HardHatWorkers = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------15 MaskWearing---------------------#
+class_name = ('mask', 'no-mask')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'MaskWearing/raw/'
+dataset_MaskWearing = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_MaskWearing = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------16 MountainDewCommercial---------------------#
+class_name = ('bottle', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'MountainDewCommercial/'
+dataset_MountainDewCommercial = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_MountainDewCommercial = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------17 NorthAmericaMushrooms---------------------#
+class_name = ('flat mushroom', 'yellow mushroom')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'NorthAmericaMushrooms/North American Mushrooms.v1-416x416.coco/' # noqa
+dataset_NorthAmericaMushrooms = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/new_annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_NorthAmericaMushrooms = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/new_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------18 openPoetryVision---------------------#
+class_name = ('American Typewriter', 'Andale Mono', 'Apple Chancery', 'Arial',
+ 'Avenir', 'Baskerville', 'Big Caslon', 'Bradley Hand',
+ 'Brush Script MT', 'Chalkboard', 'Comic Sans MS', 'Copperplate',
+ 'Courier', 'Didot', 'Futura', 'Geneva', 'Georgia', 'Gill Sans',
+ 'Helvetica', 'Herculanum', 'Impact', 'Kefa', 'Lucida Grande',
+ 'Luminari', 'Marker Felt', 'Menlo', 'Monaco', 'Noteworthy',
+ 'Optima', 'PT Sans', 'PT Serif', 'Palatino', 'Papyrus',
+ 'Phosphate', 'Rockwell', 'SF Pro', 'SignPainter', 'Skia',
+ 'Snell Roundhand', 'Tahoma', 'Times New Roman', 'Trebuchet MS',
+ 'Verdana')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'openPoetryVision/512x512/'
+dataset_openPoetryVision = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_openPoetryVision = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------19 OxfordPets_by_breed---------------------#
+class_name = ('cat-Abyssinian', 'cat-Bengal', 'cat-Birman', 'cat-Bombay',
+ 'cat-British_Shorthair', 'cat-Egyptian_Mau', 'cat-Maine_Coon',
+ 'cat-Persian', 'cat-Ragdoll', 'cat-Russian_Blue', 'cat-Siamese',
+ 'cat-Sphynx', 'dog-american_bulldog',
+ 'dog-american_pit_bull_terrier', 'dog-basset_hound',
+ 'dog-beagle', 'dog-boxer', 'dog-chihuahua',
+ 'dog-english_cocker_spaniel', 'dog-english_setter',
+ 'dog-german_shorthaired', 'dog-great_pyrenees', 'dog-havanese',
+ 'dog-japanese_chin', 'dog-keeshond', 'dog-leonberger',
+ 'dog-miniature_pinscher', 'dog-newfoundland', 'dog-pomeranian',
+ 'dog-pug', 'dog-saint_bernard', 'dog-samoyed',
+ 'dog-scottish_terrier', 'dog-shiba_inu',
+ 'dog-staffordshire_bull_terrier', 'dog-wheaten_terrier',
+ 'dog-yorkshire_terrier')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'OxfordPets/by-breed/' # noqa
+dataset_OxfordPets_by_breed = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_OxfordPets_by_breed = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------20 OxfordPets_by_species---------------------#
+class_name = ('cat', 'dog')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'OxfordPets/by-species/' # noqa
+dataset_OxfordPets_by_species = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_OxfordPets_by_species = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------21 PKLot---------------------#
+class_name = ('space-empty', 'space-occupied')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'PKLot/640/' # noqa
+dataset_PKLot = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_PKLot = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------22 Packages---------------------#
+class_name = ('package', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Packages/Raw/'
+caption_prompt = {
+ 'package': {
+ 'prefix': 'there is a ',
+ 'suffix': ' on the porch'
+ }
+}
+dataset_Packages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt, # NOTE w. prompt 0.728; wo. prompt 0.670
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Packages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------23 PascalVOC---------------------#
+class_name = ('aeroplane', 'bicycle', 'bird', 'boat', 'bottle', 'bus', 'car',
+ 'cat', 'chair', 'cow', 'diningtable', 'dog', 'horse',
+ 'motorbike', 'person', 'pottedplant', 'sheep', 'sofa', 'train',
+ 'tvmonitor')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'PascalVOC/'
+dataset_PascalVOC = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_PascalVOC = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------24 pistols---------------------#
+class_name = ('pistol', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'pistols/export/'
+dataset_pistols = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_pistols = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------25 plantdoc---------------------#
+class_name = ('Apple Scab Leaf', 'Apple leaf', 'Apple rust leaf',
+ 'Bell_pepper leaf', 'Bell_pepper leaf spot', 'Blueberry leaf',
+ 'Cherry leaf', 'Corn Gray leaf spot', 'Corn leaf blight',
+ 'Corn rust leaf', 'Peach leaf', 'Potato leaf',
+ 'Potato leaf early blight', 'Potato leaf late blight',
+ 'Raspberry leaf', 'Soyabean leaf', 'Soybean leaf',
+ 'Squash Powdery mildew leaf', 'Strawberry leaf',
+ 'Tomato Early blight leaf', 'Tomato Septoria leaf spot',
+ 'Tomato leaf', 'Tomato leaf bacterial spot',
+ 'Tomato leaf late blight', 'Tomato leaf mosaic virus',
+ 'Tomato leaf yellow virus', 'Tomato mold leaf',
+ 'Tomato two spotted spider mites leaf', 'grape leaf',
+ 'grape leaf black rot')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'plantdoc/416x416/'
+dataset_plantdoc = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_plantdoc = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------26 pothole---------------------#
+class_name = ('pothole', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'pothole/'
+caption_prompt = {
+ 'pothole': {
+ 'name': 'holes',
+ 'prefix': 'there are some ',
+ 'suffix': ' on the road'
+ }
+}
+dataset_pothole = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ # NOTE w. prompt 0.221; wo. prompt 0.478
+ # caption_prompt=caption_prompt,
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_pothole = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------27 Raccoon---------------------#
+class_name = ('raccoon', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Raccoon/Raccoon.v2-raw.coco/'
+dataset_Raccoon = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Raccoon = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------28 selfdrivingCar---------------------#
+class_name = ('biker', 'car', 'pedestrian', 'trafficLight',
+ 'trafficLight-Green', 'trafficLight-GreenLeft',
+ 'trafficLight-Red', 'trafficLight-RedLeft',
+ 'trafficLight-Yellow', 'trafficLight-YellowLeft', 'truck')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'selfdrivingCar/fixedLarge/export/'
+dataset_selfdrivingCar = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_selfdrivingCar = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------29 ShellfishOpenImages---------------------#
+class_name = ('Crab', 'Lobster', 'Shrimp')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'ShellfishOpenImages/raw/'
+dataset_ShellfishOpenImages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_ShellfishOpenImages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------30 ThermalCheetah---------------------#
+class_name = ('cheetah', 'human')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'ThermalCheetah/'
+dataset_ThermalCheetah = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_ThermalCheetah = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------31 thermalDogsAndPeople---------------------#
+class_name = ('dog', 'person')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'thermalDogsAndPeople/'
+dataset_thermalDogsAndPeople = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_thermalDogsAndPeople = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------32 UnoCards---------------------#
+class_name = ('0', '1', '2', '3', '4', '5', '6', '7', '8', '9', '10', '11',
+ '12', '13', '14')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'UnoCards/raw/'
+dataset_UnoCards = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_UnoCards = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------33 VehiclesOpenImages---------------------#
+class_name = ('Ambulance', 'Bus', 'Car', 'Motorcycle', 'Truck')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'VehiclesOpenImages/416x416/'
+dataset_VehiclesOpenImages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_VehiclesOpenImages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------34 WildfireSmoke---------------------#
+class_name = ('smoke', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'WildfireSmoke/'
+dataset_WildfireSmoke = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_WildfireSmoke = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------35 websiteScreenshots---------------------#
+class_name = ('button', 'field', 'heading', 'iframe', 'image', 'label', 'link',
+ 'text')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'websiteScreenshots/'
+dataset_websiteScreenshots = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_websiteScreenshots = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# --------------------- Config---------------------#
+
+dataset_prefixes = [
+ 'AerialMaritimeDrone_large',
+ 'AerialMaritimeDrone_tiled',
+ 'AmericanSignLanguageLetters',
+ 'Aquarium',
+ 'BCCD',
+ 'boggleBoards',
+ 'brackishUnderwater',
+ 'ChessPieces',
+ 'CottontailRabbits',
+ 'dice',
+ 'DroneControl',
+ 'EgoHands_generic',
+ 'EgoHands_specific',
+ 'HardHatWorkers',
+ 'MaskWearing',
+ 'MountainDewCommercial',
+ 'NorthAmericaMushrooms',
+ 'openPoetryVision',
+ 'OxfordPets_by_breed',
+ 'OxfordPets_by_species',
+ 'PKLot',
+ 'Packages',
+ 'PascalVOC',
+ 'pistols',
+ 'plantdoc',
+ 'pothole',
+ 'Raccoons',
+ 'selfdrivingCar',
+ 'ShellfishOpenImages',
+ 'ThermalCheetah',
+ 'thermalDogsAndPeople',
+ 'UnoCards',
+ 'VehiclesOpenImages',
+ 'WildfireSmoke',
+ 'websiteScreenshots',
+]
+
+datasets = [
+ dataset_AerialMaritimeDrone_large, dataset_AerialMaritimeDrone_tiled,
+ dataset_AmericanSignLanguageLetters, dataset_Aquarium, dataset_BCCD,
+ dataset_boggleBoards, dataset_brackishUnderwater, dataset_ChessPieces,
+ dataset_CottontailRabbits, dataset_dice, dataset_DroneControl,
+ dataset_EgoHands_generic, dataset_EgoHands_specific,
+ dataset_HardHatWorkers, dataset_MaskWearing, dataset_MountainDewCommercial,
+ dataset_NorthAmericaMushrooms, dataset_openPoetryVision,
+ dataset_OxfordPets_by_breed, dataset_OxfordPets_by_species, dataset_PKLot,
+ dataset_Packages, dataset_PascalVOC, dataset_pistols, dataset_plantdoc,
+ dataset_pothole, dataset_Raccoon, dataset_selfdrivingCar,
+ dataset_ShellfishOpenImages, dataset_ThermalCheetah,
+ dataset_thermalDogsAndPeople, dataset_UnoCards, dataset_VehiclesOpenImages,
+ dataset_WildfireSmoke, dataset_websiteScreenshots
+]
+
+metrics = [
+ val_evaluator_AerialMaritimeDrone_large,
+ val_evaluator_AerialMaritimeDrone_tiled,
+ val_evaluator_AmericanSignLanguageLetters, val_evaluator_Aquarium,
+ val_evaluator_BCCD, val_evaluator_boggleBoards,
+ val_evaluator_brackishUnderwater, val_evaluator_ChessPieces,
+ val_evaluator_CottontailRabbits, val_evaluator_dice,
+ val_evaluator_DroneControl, val_evaluator_EgoHands_generic,
+ val_evaluator_EgoHands_specific, val_evaluator_HardHatWorkers,
+ val_evaluator_MaskWearing, val_evaluator_MountainDewCommercial,
+ val_evaluator_NorthAmericaMushrooms, val_evaluator_openPoetryVision,
+ val_evaluator_OxfordPets_by_breed, val_evaluator_OxfordPets_by_species,
+ val_evaluator_PKLot, val_evaluator_Packages, val_evaluator_PascalVOC,
+ val_evaluator_pistols, val_evaluator_plantdoc, val_evaluator_pothole,
+ val_evaluator_Raccoon, val_evaluator_selfdrivingCar,
+ val_evaluator_ShellfishOpenImages, val_evaluator_ThermalCheetah,
+ val_evaluator_thermalDogsAndPeople, val_evaluator_UnoCards,
+ val_evaluator_VehiclesOpenImages, val_evaluator_WildfireSmoke,
+ val_evaluator_websiteScreenshots
+]
+
+# -------------------------------------------------#
+val_dataloader = dict(
+ dataset=dict(_delete_=True, type='ConcatDataset', datasets=datasets))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='MultiDatasetsEvaluator',
+ metrics=metrics,
+ dataset_prefixes=dataset_prefixes)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/odinw/grounding_dino_swin-t_pretrain_odinw13.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/odinw/grounding_dino_swin-t_pretrain_odinw13.py
new file mode 100644
index 0000000000000000000000000000000000000000..216b8059726b8fbe9dff3b2a43718bc563502aab
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/odinw/grounding_dino_swin-t_pretrain_odinw13.py
@@ -0,0 +1,338 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365_goldg_cap4m.py' # noqa
+
+dataset_type = 'CocoDataset'
+data_root = 'data/odinw/'
+
+base_test_pipeline = _base_.test_pipeline
+base_test_pipeline[-1]['meta_keys'] = ('img_id', 'img_path', 'ori_shape',
+ 'img_shape', 'scale_factor', 'text',
+ 'custom_entities', 'caption_prompt')
+
+# ---------------------1 AerialMaritimeDrone---------------------#
+class_name = ('boat', 'car', 'dock', 'jetski', 'lift')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'AerialMaritimeDrone/large/'
+dataset_AerialMaritimeDrone = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ test_mode=True,
+ pipeline=base_test_pipeline,
+ return_classes=True)
+val_evaluator_AerialMaritimeDrone = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------2 Aquarium---------------------#
+class_name = ('fish', 'jellyfish', 'penguin', 'puffin', 'shark', 'starfish',
+ 'stingray')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Aquarium/Aquarium Combined.v2-raw-1024.coco/'
+
+caption_prompt = None
+# caption_prompt = {
+# 'penguin': {
+# 'suffix': ', which is black and white'
+# },
+# 'puffin': {
+# 'suffix': ' with orange beaks'
+# },
+# 'stingray': {
+# 'suffix': ' which is flat and round'
+# },
+# }
+dataset_Aquarium = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Aquarium = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------3 CottontailRabbits---------------------#
+class_name = ('Cottontail-Rabbit', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'CottontailRabbits/'
+
+caption_prompt = None
+# caption_prompt = {'Cottontail-Rabbit': {'name': 'rabbit'}}
+
+dataset_CottontailRabbits = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_CottontailRabbits = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------4 EgoHands---------------------#
+class_name = ('hand', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'EgoHands/generic/'
+
+caption_prompt = None
+# caption_prompt = {'hand': {'suffix': ' of a person'}}
+
+dataset_EgoHands = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_EgoHands = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------5 NorthAmericaMushrooms---------------------#
+class_name = ('CoW', 'chanterelle')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'NorthAmericaMushrooms/North American Mushrooms.v1-416x416.coco/' # noqa
+
+caption_prompt = None
+# caption_prompt = {
+# 'CoW': {
+# 'name': 'flat mushroom'
+# },
+# 'chanterelle': {
+# 'name': 'yellow mushroom'
+# }
+# }
+
+dataset_NorthAmericaMushrooms = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_NorthAmericaMushrooms = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------6 Packages---------------------#
+class_name = ('package', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Packages/Raw/'
+
+caption_prompt = None
+# caption_prompt = {
+# 'package': {
+# 'prefix': 'there is a ',
+# 'suffix': ' on the porch'
+# }
+# }
+
+dataset_Packages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Packages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------7 PascalVOC---------------------#
+class_name = ('aeroplane', 'bicycle', 'bird', 'boat', 'bottle', 'bus', 'car',
+ 'cat', 'chair', 'cow', 'diningtable', 'dog', 'horse',
+ 'motorbike', 'person', 'pottedplant', 'sheep', 'sofa', 'train',
+ 'tvmonitor')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'PascalVOC/'
+dataset_PascalVOC = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_PascalVOC = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------8 pistols---------------------#
+class_name = ('pistol', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'pistols/export/'
+dataset_pistols = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_pistols = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------9 pothole---------------------#
+class_name = ('pothole', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'pothole/'
+
+caption_prompt = None
+# caption_prompt = {
+# 'pothole': {
+# 'prefix': 'there are some ',
+# 'name': 'holes',
+# 'suffix': ' on the road'
+# }
+# }
+
+dataset_pothole = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_pothole = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------10 Raccoon---------------------#
+class_name = ('raccoon', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Raccoon/Raccoon.v2-raw.coco/'
+dataset_Raccoon = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Raccoon = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------11 ShellfishOpenImages---------------------#
+class_name = ('Crab', 'Lobster', 'Shrimp')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'ShellfishOpenImages/raw/'
+dataset_ShellfishOpenImages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_ShellfishOpenImages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------12 thermalDogsAndPeople---------------------#
+class_name = ('dog', 'person')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'thermalDogsAndPeople/'
+dataset_thermalDogsAndPeople = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_thermalDogsAndPeople = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------13 VehiclesOpenImages---------------------#
+class_name = ('Ambulance', 'Bus', 'Car', 'Motorcycle', 'Truck')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'VehiclesOpenImages/416x416/'
+dataset_VehiclesOpenImages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_VehiclesOpenImages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# --------------------- Config---------------------#
+dataset_prefixes = [
+ 'AerialMaritimeDrone', 'Aquarium', 'CottontailRabbits', 'EgoHands',
+ 'NorthAmericaMushrooms', 'Packages', 'PascalVOC', 'pistols', 'pothole',
+ 'Raccoon', 'ShellfishOpenImages', 'thermalDogsAndPeople',
+ 'VehiclesOpenImages'
+]
+datasets = [
+ dataset_AerialMaritimeDrone, dataset_Aquarium, dataset_CottontailRabbits,
+ dataset_EgoHands, dataset_NorthAmericaMushrooms, dataset_Packages,
+ dataset_PascalVOC, dataset_pistols, dataset_pothole, dataset_Raccoon,
+ dataset_ShellfishOpenImages, dataset_thermalDogsAndPeople,
+ dataset_VehiclesOpenImages
+]
+metrics = [
+ val_evaluator_AerialMaritimeDrone, val_evaluator_Aquarium,
+ val_evaluator_CottontailRabbits, val_evaluator_EgoHands,
+ val_evaluator_NorthAmericaMushrooms, val_evaluator_Packages,
+ val_evaluator_PascalVOC, val_evaluator_pistols, val_evaluator_pothole,
+ val_evaluator_Raccoon, val_evaluator_ShellfishOpenImages,
+ val_evaluator_thermalDogsAndPeople, val_evaluator_VehiclesOpenImages
+]
+
+# -------------------------------------------------#
+val_dataloader = dict(
+ dataset=dict(_delete_=True, type='ConcatDataset', datasets=datasets))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='MultiDatasetsEvaluator',
+ metrics=metrics,
+ dataset_prefixes=dataset_prefixes)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/odinw/grounding_dino_swin-t_pretrain_odinw35.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/odinw/grounding_dino_swin-t_pretrain_odinw35.py
new file mode 100644
index 0000000000000000000000000000000000000000..3df0394a204061684cbb9bb66adb08d92a784efb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/odinw/grounding_dino_swin-t_pretrain_odinw35.py
@@ -0,0 +1,796 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365_goldg_cap4m.py' # noqa
+
+dataset_type = 'CocoDataset'
+data_root = 'data/odinw/'
+
+base_test_pipeline = _base_.test_pipeline
+base_test_pipeline[-1]['meta_keys'] = ('img_id', 'img_path', 'ori_shape',
+ 'img_shape', 'scale_factor', 'text',
+ 'custom_entities', 'caption_prompt')
+
+# ---------------------1 AerialMaritimeDrone_large---------------------#
+class_name = ('boat', 'car', 'dock', 'jetski', 'lift')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'AerialMaritimeDrone/large/'
+dataset_AerialMaritimeDrone_large = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_AerialMaritimeDrone_large = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------2 AerialMaritimeDrone_tiled---------------------#
+class_name = ('boat', 'car', 'dock', 'jetski', 'lift')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'AerialMaritimeDrone/tiled/'
+dataset_AerialMaritimeDrone_tiled = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_AerialMaritimeDrone_tiled = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------3 AmericanSignLanguageLetters---------------------#
+class_name = ('A', 'B', 'C', 'D', 'E', 'F', 'G', 'H', 'I', 'J', 'K', 'L', 'M',
+ 'N', 'O', 'P', 'Q', 'R', 'S', 'T', 'U', 'V', 'W', 'X', 'Y', 'Z')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'AmericanSignLanguageLetters/American Sign Language Letters.v1-v1.coco/' # noqa
+dataset_AmericanSignLanguageLetters = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_AmericanSignLanguageLetters = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------4 Aquarium---------------------#
+class_name = ('fish', 'jellyfish', 'penguin', 'puffin', 'shark', 'starfish',
+ 'stingray')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Aquarium/Aquarium Combined.v2-raw-1024.coco/'
+dataset_Aquarium = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Aquarium = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------5 BCCD---------------------#
+class_name = ('Platelets', 'RBC', 'WBC')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'BCCD/BCCD.v3-raw.coco/'
+dataset_BCCD = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_BCCD = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------6 boggleBoards---------------------#
+class_name = ('Q', 'a', 'an', 'b', 'c', 'd', 'e', 'er', 'f', 'g', 'h', 'he',
+ 'i', 'in', 'j', 'k', 'l', 'm', 'n', 'o', 'o ', 'p', 'q', 'qu',
+ 'r', 's', 't', 't\\', 'th', 'u', 'v', 'w', 'wild', 'x', 'y', 'z')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'boggleBoards/416x416AutoOrient/export/'
+dataset_boggleBoards = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_boggleBoards = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------7 brackishUnderwater---------------------#
+class_name = ('crab', 'fish', 'jellyfish', 'shrimp', 'small_fish', 'starfish')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'brackishUnderwater/960x540/'
+dataset_brackishUnderwater = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_brackishUnderwater = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------8 ChessPieces---------------------#
+class_name = (' ', 'black bishop', 'black king', 'black knight', 'black pawn',
+ 'black queen', 'black rook', 'white bishop', 'white king',
+ 'white knight', 'white pawn', 'white queen', 'white rook')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'ChessPieces/Chess Pieces.v23-raw.coco/'
+dataset_ChessPieces = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/new_annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_ChessPieces = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/new_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------9 CottontailRabbits---------------------#
+class_name = ('rabbit', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'CottontailRabbits/'
+dataset_CottontailRabbits = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/new_annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_CottontailRabbits = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/new_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------10 dice---------------------#
+class_name = ('1', '2', '3', '4', '5', '6')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'dice/mediumColor/export/'
+dataset_dice = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_dice = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------11 DroneControl---------------------#
+class_name = ('follow', 'follow_hand', 'land', 'land_hand', 'null', 'object',
+ 'takeoff', 'takeoff-hand')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'DroneControl/Drone Control.v3-raw.coco/'
+dataset_DroneControl = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_DroneControl = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------12 EgoHands_generic---------------------#
+class_name = ('hand', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'EgoHands/generic/'
+caption_prompt = {'hand': {'suffix': ' of a person'}}
+dataset_EgoHands_generic = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ # NOTE w. prompt 0.526, wo. prompt 0.608
+ # caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_EgoHands_generic = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------13 EgoHands_specific---------------------#
+class_name = ('myleft', 'myright', 'yourleft', 'yourright')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'EgoHands/specific/'
+dataset_EgoHands_specific = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_EgoHands_specific = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------14 HardHatWorkers---------------------#
+class_name = ('head', 'helmet', 'person')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'HardHatWorkers/raw/'
+dataset_HardHatWorkers = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_HardHatWorkers = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------15 MaskWearing---------------------#
+class_name = ('mask', 'no-mask')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'MaskWearing/raw/'
+dataset_MaskWearing = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_MaskWearing = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------16 MountainDewCommercial---------------------#
+class_name = ('bottle', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'MountainDewCommercial/'
+dataset_MountainDewCommercial = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_MountainDewCommercial = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------17 NorthAmericaMushrooms---------------------#
+class_name = ('flat mushroom', 'yellow mushroom')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'NorthAmericaMushrooms/North American Mushrooms.v1-416x416.coco/' # noqa
+dataset_NorthAmericaMushrooms = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/new_annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_NorthAmericaMushrooms = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/new_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------18 openPoetryVision---------------------#
+class_name = ('American Typewriter', 'Andale Mono', 'Apple Chancery', 'Arial',
+ 'Avenir', 'Baskerville', 'Big Caslon', 'Bradley Hand',
+ 'Brush Script MT', 'Chalkboard', 'Comic Sans MS', 'Copperplate',
+ 'Courier', 'Didot', 'Futura', 'Geneva', 'Georgia', 'Gill Sans',
+ 'Helvetica', 'Herculanum', 'Impact', 'Kefa', 'Lucida Grande',
+ 'Luminari', 'Marker Felt', 'Menlo', 'Monaco', 'Noteworthy',
+ 'Optima', 'PT Sans', 'PT Serif', 'Palatino', 'Papyrus',
+ 'Phosphate', 'Rockwell', 'SF Pro', 'SignPainter', 'Skia',
+ 'Snell Roundhand', 'Tahoma', 'Times New Roman', 'Trebuchet MS',
+ 'Verdana')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'openPoetryVision/512x512/'
+dataset_openPoetryVision = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_openPoetryVision = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------19 OxfordPets_by_breed---------------------#
+class_name = ('cat-Abyssinian', 'cat-Bengal', 'cat-Birman', 'cat-Bombay',
+ 'cat-British_Shorthair', 'cat-Egyptian_Mau', 'cat-Maine_Coon',
+ 'cat-Persian', 'cat-Ragdoll', 'cat-Russian_Blue', 'cat-Siamese',
+ 'cat-Sphynx', 'dog-american_bulldog',
+ 'dog-american_pit_bull_terrier', 'dog-basset_hound',
+ 'dog-beagle', 'dog-boxer', 'dog-chihuahua',
+ 'dog-english_cocker_spaniel', 'dog-english_setter',
+ 'dog-german_shorthaired', 'dog-great_pyrenees', 'dog-havanese',
+ 'dog-japanese_chin', 'dog-keeshond', 'dog-leonberger',
+ 'dog-miniature_pinscher', 'dog-newfoundland', 'dog-pomeranian',
+ 'dog-pug', 'dog-saint_bernard', 'dog-samoyed',
+ 'dog-scottish_terrier', 'dog-shiba_inu',
+ 'dog-staffordshire_bull_terrier', 'dog-wheaten_terrier',
+ 'dog-yorkshire_terrier')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'OxfordPets/by-breed/' # noqa
+dataset_OxfordPets_by_breed = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_OxfordPets_by_breed = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------20 OxfordPets_by_species---------------------#
+class_name = ('cat', 'dog')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'OxfordPets/by-species/' # noqa
+dataset_OxfordPets_by_species = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_OxfordPets_by_species = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------21 PKLot---------------------#
+class_name = ('space-empty', 'space-occupied')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'PKLot/640/' # noqa
+dataset_PKLot = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_PKLot = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------22 Packages---------------------#
+class_name = ('package', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Packages/Raw/'
+caption_prompt = {
+ 'package': {
+ 'prefix': 'there is a ',
+ 'suffix': ' on the porch'
+ }
+}
+dataset_Packages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt, # NOTE w. prompt 0.695; wo. prompt 0.687
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Packages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------23 PascalVOC---------------------#
+class_name = ('aeroplane', 'bicycle', 'bird', 'boat', 'bottle', 'bus', 'car',
+ 'cat', 'chair', 'cow', 'diningtable', 'dog', 'horse',
+ 'motorbike', 'person', 'pottedplant', 'sheep', 'sofa', 'train',
+ 'tvmonitor')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'PascalVOC/'
+dataset_PascalVOC = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_PascalVOC = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------24 pistols---------------------#
+class_name = ('pistol', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'pistols/export/'
+dataset_pistols = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_pistols = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------25 plantdoc---------------------#
+class_name = ('Apple Scab Leaf', 'Apple leaf', 'Apple rust leaf',
+ 'Bell_pepper leaf', 'Bell_pepper leaf spot', 'Blueberry leaf',
+ 'Cherry leaf', 'Corn Gray leaf spot', 'Corn leaf blight',
+ 'Corn rust leaf', 'Peach leaf', 'Potato leaf',
+ 'Potato leaf early blight', 'Potato leaf late blight',
+ 'Raspberry leaf', 'Soyabean leaf', 'Soybean leaf',
+ 'Squash Powdery mildew leaf', 'Strawberry leaf',
+ 'Tomato Early blight leaf', 'Tomato Septoria leaf spot',
+ 'Tomato leaf', 'Tomato leaf bacterial spot',
+ 'Tomato leaf late blight', 'Tomato leaf mosaic virus',
+ 'Tomato leaf yellow virus', 'Tomato mold leaf',
+ 'Tomato two spotted spider mites leaf', 'grape leaf',
+ 'grape leaf black rot')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'plantdoc/416x416/'
+dataset_plantdoc = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_plantdoc = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------26 pothole---------------------#
+class_name = ('pothole', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'pothole/'
+caption_prompt = {
+ 'pothole': {
+ 'name': 'holes',
+ 'prefix': 'there are some ',
+ 'suffix': ' on the road'
+ }
+}
+dataset_pothole = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ # NOTE w. prompt 0.137; wo. prompt 0.215
+ # caption_prompt=caption_prompt,
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_pothole = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------27 Raccoon---------------------#
+class_name = ('raccoon', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Raccoon/Raccoon.v2-raw.coco/'
+dataset_Raccoon = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Raccoon = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------28 selfdrivingCar---------------------#
+class_name = ('biker', 'car', 'pedestrian', 'trafficLight',
+ 'trafficLight-Green', 'trafficLight-GreenLeft',
+ 'trafficLight-Red', 'trafficLight-RedLeft',
+ 'trafficLight-Yellow', 'trafficLight-YellowLeft', 'truck')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'selfdrivingCar/fixedLarge/export/'
+dataset_selfdrivingCar = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_selfdrivingCar = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------29 ShellfishOpenImages---------------------#
+class_name = ('Crab', 'Lobster', 'Shrimp')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'ShellfishOpenImages/raw/'
+dataset_ShellfishOpenImages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_ShellfishOpenImages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------30 ThermalCheetah---------------------#
+class_name = ('cheetah', 'human')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'ThermalCheetah/'
+dataset_ThermalCheetah = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_ThermalCheetah = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------31 thermalDogsAndPeople---------------------#
+class_name = ('dog', 'person')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'thermalDogsAndPeople/'
+dataset_thermalDogsAndPeople = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_thermalDogsAndPeople = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------32 UnoCards---------------------#
+class_name = ('0', '1', '2', '3', '4', '5', '6', '7', '8', '9', '10', '11',
+ '12', '13', '14')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'UnoCards/raw/'
+dataset_UnoCards = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_UnoCards = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------33 VehiclesOpenImages---------------------#
+class_name = ('Ambulance', 'Bus', 'Car', 'Motorcycle', 'Truck')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'VehiclesOpenImages/416x416/'
+dataset_VehiclesOpenImages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_VehiclesOpenImages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------34 WildfireSmoke---------------------#
+class_name = ('smoke', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'WildfireSmoke/'
+dataset_WildfireSmoke = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_WildfireSmoke = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------35 websiteScreenshots---------------------#
+class_name = ('button', 'field', 'heading', 'iframe', 'image', 'label', 'link',
+ 'text')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'websiteScreenshots/'
+dataset_websiteScreenshots = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_websiteScreenshots = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# --------------------- Config---------------------#
+
+dataset_prefixes = [
+ 'AerialMaritimeDrone_large',
+ 'AerialMaritimeDrone_tiled',
+ 'AmericanSignLanguageLetters',
+ 'Aquarium',
+ 'BCCD',
+ 'boggleBoards',
+ 'brackishUnderwater',
+ 'ChessPieces',
+ 'CottontailRabbits',
+ 'dice',
+ 'DroneControl',
+ 'EgoHands_generic',
+ 'EgoHands_specific',
+ 'HardHatWorkers',
+ 'MaskWearing',
+ 'MountainDewCommercial',
+ 'NorthAmericaMushrooms',
+ 'openPoetryVision',
+ 'OxfordPets_by_breed',
+ 'OxfordPets_by_species',
+ 'PKLot',
+ 'Packages',
+ 'PascalVOC',
+ 'pistols',
+ 'plantdoc',
+ 'pothole',
+ 'Raccoons',
+ 'selfdrivingCar',
+ 'ShellfishOpenImages',
+ 'ThermalCheetah',
+ 'thermalDogsAndPeople',
+ 'UnoCards',
+ 'VehiclesOpenImages',
+ 'WildfireSmoke',
+ 'websiteScreenshots',
+]
+
+datasets = [
+ dataset_AerialMaritimeDrone_large, dataset_AerialMaritimeDrone_tiled,
+ dataset_AmericanSignLanguageLetters, dataset_Aquarium, dataset_BCCD,
+ dataset_boggleBoards, dataset_brackishUnderwater, dataset_ChessPieces,
+ dataset_CottontailRabbits, dataset_dice, dataset_DroneControl,
+ dataset_EgoHands_generic, dataset_EgoHands_specific,
+ dataset_HardHatWorkers, dataset_MaskWearing, dataset_MountainDewCommercial,
+ dataset_NorthAmericaMushrooms, dataset_openPoetryVision,
+ dataset_OxfordPets_by_breed, dataset_OxfordPets_by_species, dataset_PKLot,
+ dataset_Packages, dataset_PascalVOC, dataset_pistols, dataset_plantdoc,
+ dataset_pothole, dataset_Raccoon, dataset_selfdrivingCar,
+ dataset_ShellfishOpenImages, dataset_ThermalCheetah,
+ dataset_thermalDogsAndPeople, dataset_UnoCards, dataset_VehiclesOpenImages,
+ dataset_WildfireSmoke, dataset_websiteScreenshots
+]
+
+metrics = [
+ val_evaluator_AerialMaritimeDrone_large,
+ val_evaluator_AerialMaritimeDrone_tiled,
+ val_evaluator_AmericanSignLanguageLetters, val_evaluator_Aquarium,
+ val_evaluator_BCCD, val_evaluator_boggleBoards,
+ val_evaluator_brackishUnderwater, val_evaluator_ChessPieces,
+ val_evaluator_CottontailRabbits, val_evaluator_dice,
+ val_evaluator_DroneControl, val_evaluator_EgoHands_generic,
+ val_evaluator_EgoHands_specific, val_evaluator_HardHatWorkers,
+ val_evaluator_MaskWearing, val_evaluator_MountainDewCommercial,
+ val_evaluator_NorthAmericaMushrooms, val_evaluator_openPoetryVision,
+ val_evaluator_OxfordPets_by_breed, val_evaluator_OxfordPets_by_species,
+ val_evaluator_PKLot, val_evaluator_Packages, val_evaluator_PascalVOC,
+ val_evaluator_pistols, val_evaluator_plantdoc, val_evaluator_pothole,
+ val_evaluator_Raccoon, val_evaluator_selfdrivingCar,
+ val_evaluator_ShellfishOpenImages, val_evaluator_ThermalCheetah,
+ val_evaluator_thermalDogsAndPeople, val_evaluator_UnoCards,
+ val_evaluator_VehiclesOpenImages, val_evaluator_WildfireSmoke,
+ val_evaluator_websiteScreenshots
+]
+
+# -------------------------------------------------#
+val_dataloader = dict(
+ dataset=dict(_delete_=True, type='ConcatDataset', datasets=datasets))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='MultiDatasetsEvaluator',
+ metrics=metrics,
+ dataset_prefixes=dataset_prefixes)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/odinw/override_category.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/odinw/override_category.py
new file mode 100644
index 0000000000000000000000000000000000000000..9ff05fc6e5e4d0989cf7fcf7af4dc902ee99f3a3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/odinw/override_category.py
@@ -0,0 +1,109 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import argparse
+
+import mmengine
+
+
+def parse_args():
+ parser = argparse.ArgumentParser(description='Override Category')
+ parser.add_argument('data_root')
+ return parser.parse_args()
+
+
+def main():
+ args = parse_args()
+
+ ChessPieces = [{
+ 'id': 1,
+ 'name': ' ',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 2,
+ 'name': 'black bishop',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 3,
+ 'name': 'black king',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 4,
+ 'name': 'black knight',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 5,
+ 'name': 'black pawn',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 6,
+ 'name': 'black queen',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 7,
+ 'name': 'black rook',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 8,
+ 'name': 'white bishop',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 9,
+ 'name': 'white king',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 10,
+ 'name': 'white knight',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 11,
+ 'name': 'white pawn',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 12,
+ 'name': 'white queen',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 13,
+ 'name': 'white rook',
+ 'supercategory': 'pieces'
+ }]
+
+ _data_root = args.data_root + 'ChessPieces/Chess Pieces.v23-raw.coco/'
+ json_data = mmengine.load(_data_root +
+ 'valid/annotations_without_background.json')
+ json_data['categories'] = ChessPieces
+ mmengine.dump(json_data,
+ _data_root + 'valid/new_annotations_without_background.json')
+
+ CottontailRabbits = [{
+ 'id': 1,
+ 'name': 'rabbit',
+ 'supercategory': 'Cottontail-Rabbit'
+ }]
+
+ _data_root = args.data_root + 'CottontailRabbits/'
+ json_data = mmengine.load(_data_root +
+ 'valid/annotations_without_background.json')
+ json_data['categories'] = CottontailRabbits
+ mmengine.dump(json_data,
+ _data_root + 'valid/new_annotations_without_background.json')
+
+ NorthAmericaMushrooms = [{
+ 'id': 1,
+ 'name': 'flat mushroom',
+ 'supercategory': 'mushroom'
+ }, {
+ 'id': 2,
+ 'name': 'yellow mushroom',
+ 'supercategory': 'mushroom'
+ }]
+
+ _data_root = args.data_root + 'NorthAmericaMushrooms/North American Mushrooms.v1-416x416.coco/' # noqa
+ json_data = mmengine.load(_data_root +
+ 'valid/annotations_without_background.json')
+ json_data['categories'] = NorthAmericaMushrooms
+ mmengine.dump(json_data,
+ _data_root + 'valid/new_annotations_without_background.json')
+
+
+if __name__ == '__main__':
+ main()
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/refcoco/grounding_dino_swin-b_pretrain_zeroshot_refexp.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/refcoco/grounding_dino_swin-b_pretrain_zeroshot_refexp.py
new file mode 100644
index 0000000000000000000000000000000000000000..dea0bad08c0ebf6455211fadb268b07868ab4ded
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/refcoco/grounding_dino_swin-b_pretrain_zeroshot_refexp.py
@@ -0,0 +1,14 @@
+_base_ = './grounding_dino_swin-t_pretrain_zeroshot_refexp.py'
+
+model = dict(
+ type='GroundingDINO',
+ backbone=dict(
+ pretrain_img_size=384,
+ embed_dims=128,
+ depths=[2, 2, 18, 2],
+ num_heads=[4, 8, 16, 32],
+ window_size=12,
+ drop_path_rate=0.3,
+ patch_norm=True),
+ neck=dict(in_channels=[256, 512, 1024]),
+)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/refcoco/grounding_dino_swin-t_pretrain_zeroshot_refexp.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/refcoco/grounding_dino_swin-t_pretrain_zeroshot_refexp.py
new file mode 100644
index 0000000000000000000000000000000000000000..4b5c46574a30bbb2253fc69f79edbcf0cb016505
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/grounding_dino/refcoco/grounding_dino_swin-t_pretrain_zeroshot_refexp.py
@@ -0,0 +1,228 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365_goldg_cap4m.py'
+
+# 30 is an empirical value, just set it to the maximum value
+# without affecting the evaluation result
+model = dict(test_cfg=dict(max_per_img=30))
+
+data_root = 'data/coco/'
+
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile', backend_args=None,
+ imdecode_backend='pillow'),
+ dict(
+ type='FixScaleResize',
+ scale=(800, 1333),
+ keep_ratio=True,
+ backend='pillow'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'text', 'custom_entities',
+ 'tokens_positive'))
+]
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/final_refexp_val.json'
+val_dataset_all_val = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=test_pipeline,
+ backend_args=None)
+val_evaluator_all_val = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_refcoco_testA.json'
+val_dataset_refcoco_testA = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=test_pipeline,
+ backend_args=None)
+
+val_evaluator_refcoco_testA = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_refcoco_testB.json'
+val_dataset_refcoco_testB = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=test_pipeline,
+ backend_args=None)
+
+val_evaluator_refcoco_testB = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_refcoco+_testA.json'
+val_dataset_refcoco_plus_testA = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=test_pipeline,
+ backend_args=None)
+
+val_evaluator_refcoco_plus_testA = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_refcoco+_testB.json'
+val_dataset_refcoco_plus_testB = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=test_pipeline,
+ backend_args=None)
+
+val_evaluator_refcoco_plus_testB = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_refcocog_test.json'
+val_dataset_refcocog_test = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=test_pipeline,
+ backend_args=None)
+
+val_evaluator_refcocog_test = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_grefcoco_val.json'
+val_dataset_grefcoco_val = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=test_pipeline,
+ backend_args=None)
+
+val_evaluator_grefcoco_val = dict(
+ type='gRefCOCOMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ thresh_score=0.7,
+ thresh_f1=1.0)
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_grefcoco_testA.json'
+val_dataset_grefcoco_testA = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=test_pipeline,
+ backend_args=None)
+
+val_evaluator_grefcoco_testA = dict(
+ type='gRefCOCOMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ thresh_score=0.7,
+ thresh_f1=1.0)
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_grefcoco_testB.json'
+val_dataset_grefcoco_testB = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=test_pipeline,
+ backend_args=None)
+
+val_evaluator_grefcoco_testB = dict(
+ type='gRefCOCOMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ thresh_score=0.7,
+ thresh_f1=1.0)
+
+# -------------------------------------------------#
+datasets = [
+ val_dataset_all_val, val_dataset_refcoco_testA, val_dataset_refcoco_testB,
+ val_dataset_refcoco_plus_testA, val_dataset_refcoco_plus_testB,
+ val_dataset_refcocog_test, val_dataset_grefcoco_val,
+ val_dataset_grefcoco_testA, val_dataset_grefcoco_testB
+]
+dataset_prefixes = [
+ 'val', 'refcoco_testA', 'refcoco_testB', 'refcoco+_testA',
+ 'refcoco+_testB', 'refcocog_test', 'grefcoco_val', 'grefcoco_testA',
+ 'grefcoco_testB'
+]
+metrics = [
+ val_evaluator_all_val, val_evaluator_refcoco_testA,
+ val_evaluator_refcoco_testB, val_evaluator_refcoco_plus_testA,
+ val_evaluator_refcoco_plus_testB, val_evaluator_refcocog_test,
+ val_evaluator_grefcoco_val, val_evaluator_grefcoco_testA,
+ val_evaluator_grefcoco_testB
+]
+
+val_dataloader = dict(
+ dataset=dict(_delete_=True, type='ConcatDataset', datasets=datasets))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='MultiDatasetsEvaluator',
+ metrics=metrics,
+ dataset_prefixes=dataset_prefixes)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..1a5e505d2888f4c521c29d9c8bc6079fac077590
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/README.md
@@ -0,0 +1,59 @@
+# Guided Anchoring
+
+> [Region Proposal by Guided Anchoring](https://arxiv.org/abs/1901.03278)
+
+
+
+## Abstract
+
+Region anchors are the cornerstone of modern object detection techniques. State-of-the-art detectors mostly rely on a dense anchoring scheme, where anchors are sampled uniformly over the spatial domain with a predefined set of scales and aspect ratios. In this paper, we revisit this foundational stage. Our study shows that it can be done much more effectively and efficiently. Specifically, we present an alternative scheme, named Guided Anchoring, which leverages semantic features to guide the anchoring. The proposed method jointly predicts the locations where the center of objects of interest are likely to exist as well as the scales and aspect ratios at different locations. On top of predicted anchor shapes, we mitigate the feature inconsistency with a feature adaption module. We also study the use of high-quality proposals to improve detection performance. The anchoring scheme can be seamlessly integrated into proposal methods and detectors. With Guided Anchoring, we achieve 9.1% higher recall on MS COCO with 90% fewer anchors than the RPN baseline. We also adopt Guided Anchoring in Fast R-CNN, Faster R-CNN and RetinaNet, respectively improving the detection mAP by 2.2%, 2.7% and 1.2%.
+
+
+

+
+
+## Results and Models
+
+The results on COCO 2017 val is shown in the below table. (results on test-dev are usually slightly higher than val).
+
+| Method | Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | AR 1000 | Config | Download |
+| :----: | :-------------: | :-----: | :-----: | :------: | :------------: | :-----: | :------------------------------------------: | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| GA-RPN | R-50-FPN | caffe | 1x | 5.3 | 15.8 | 68.4 | [config](./ga-rpn_r50-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_rpn_r50_caffe_fpn_1x_coco/ga_rpn_r50_caffe_fpn_1x_coco_20200531-899008a6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_rpn_r50_caffe_fpn_1x_coco/ga_rpn_r50_caffe_fpn_1x_coco_20200531_011819.log.json) |
+| GA-RPN | R-101-FPN | caffe | 1x | 7.3 | 13.0 | 69.5 | [config](./ga-rpn_r101-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_rpn_r101_caffe_fpn_1x_coco/ga_rpn_r101_caffe_fpn_1x_coco_20200531-ca9ba8fb.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_rpn_r101_caffe_fpn_1x_coco/ga_rpn_r101_caffe_fpn_1x_coco_20200531_011812.log.json) |
+| GA-RPN | X-101-32x4d-FPN | pytorch | 1x | 8.5 | 10.0 | 70.6 | [config](./ga-rpn_x101-32x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_rpn_x101_32x4d_fpn_1x_coco/ga_rpn_x101_32x4d_fpn_1x_coco_20200220-c28d1b18.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_rpn_x101_32x4d_fpn_1x_coco/ga_rpn_x101_32x4d_fpn_1x_coco_20200220_221326.log.json) |
+| GA-RPN | X-101-64x4d-FPN | pytorch | 1x | 7.1 | 7.5 | 71.2 | [config](./ga-rpn_x101-64x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_rpn_x101_64x4d_fpn_1x_coco/ga_rpn_x101_64x4d_fpn_1x_coco_20200225-3c6e1aa2.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_rpn_x101_64x4d_fpn_1x_coco/ga_rpn_x101_64x4d_fpn_1x_coco_20200225_152704.log.json) |
+
+| Method | Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :------------: | :-------------: | :-----: | :-----: | :------: | :------------: | :----: | :--------------------------------------------------: | :-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| GA-Faster RCNN | R-50-FPN | caffe | 1x | 5.5 | | 39.6 | [config](./ga-faster-rcnn_r50-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_faster_r50_caffe_fpn_1x_coco/ga_faster_r50_caffe_fpn_1x_coco_20200702_000718-a11ccfe6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_faster_r50_caffe_fpn_1x_coco/ga_faster_r50_caffe_fpn_1x_coco_20200702_000718.log.json) |
+| GA-Faster RCNN | R-101-FPN | caffe | 1x | 7.5 | | 41.5 | [config](./ga-faster-rcnn_r101-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_faster_r101_caffe_fpn_1x_coco/ga_faster_r101_caffe_fpn_1x_coco_bbox_mAP-0.415_20200505_115528-fb82e499.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_faster_r101_caffe_fpn_1x_coco/ga_faster_r101_caffe_fpn_1x_coco_20200505_115528.log.json) |
+| GA-Faster RCNN | X-101-32x4d-FPN | pytorch | 1x | 8.7 | 9.7 | 43.0 | [config](./ga-faster-rcnn_x101-32x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_faster_x101_32x4d_fpn_1x_coco/ga_faster_x101_32x4d_fpn_1x_coco_20200215-1ded9da3.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_faster_x101_32x4d_fpn_1x_coco/ga_faster_x101_32x4d_fpn_1x_coco_20200215_184547.log.json) |
+| GA-Faster RCNN | X-101-64x4d-FPN | pytorch | 1x | 11.8 | 7.3 | 43.9 | [config](./ga-faster-rcnn_x101-64x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_faster_x101_64x4d_fpn_1x_coco/ga_faster_x101_64x4d_fpn_1x_coco_20200215-0fa7bde7.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_faster_x101_64x4d_fpn_1x_coco/ga_faster_x101_64x4d_fpn_1x_coco_20200215_104455.log.json) |
+| GA-RetinaNet | R-50-FPN | caffe | 1x | 3.5 | 16.8 | 36.9 | [config](./ga-retinanet_r50-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_retinanet_r50_caffe_fpn_1x_coco/ga_retinanet_r50_caffe_fpn_1x_coco_20201020-39581c6f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_retinanet_r50_caffe_fpn_1x_coco/ga_retinanet_r50_caffe_fpn_1x_coco_20201020_225450.log.json) |
+| GA-RetinaNet | R-101-FPN | caffe | 1x | 5.5 | 12.9 | 39.0 | [config](./ga-retinanet_r101-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_retinanet_r101_caffe_fpn_1x_coco/ga_retinanet_r101_caffe_fpn_1x_coco_20200531-6266453c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_retinanet_r101_caffe_fpn_1x_coco/ga_retinanet_r101_caffe_fpn_1x_coco_20200531_012847.log.json) |
+| GA-RetinaNet | X-101-32x4d-FPN | pytorch | 1x | 6.9 | 10.6 | 40.5 | [config](./ga-retinanet_x101-32x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_retinanet_x101_32x4d_fpn_1x_coco/ga_retinanet_x101_32x4d_fpn_1x_coco_20200219-40c56caa.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_retinanet_x101_32x4d_fpn_1x_coco/ga_retinanet_x101_32x4d_fpn_1x_coco_20200219_223025.log.json) |
+| GA-RetinaNet | X-101-64x4d-FPN | pytorch | 1x | 9.9 | 7.7 | 41.3 | [config](./ga-retinanet_x101-64x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_retinanet_x101_64x4d_fpn_1x_coco/ga_retinanet_x101_64x4d_fpn_1x_coco_20200226-ef9f7f1f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_retinanet_x101_64x4d_fpn_1x_coco/ga_retinanet_x101_64x4d_fpn_1x_coco_20200226_221123.log.json) |
+
+- In the Guided Anchoring paper, `score_thr` is set to 0.001 in Fast/Faster RCNN and 0.05 in RetinaNet for both baselines and Guided Anchoring.
+
+- Performance on COCO test-dev benchmark are shown as follows.
+
+| Method | Backbone | Style | Lr schd | Aug Train | Score thr | AP | AP_50 | AP_75 | AP_small | AP_medium | AP_large | Download |
+| :------------: | :-------: | :---: | :-----: | :-------: | :-------: | :-: | :---: | :---: | :------: | :-------: | :------: | :------: |
+| GA-Faster RCNN | R-101-FPN | caffe | 1x | F | 0.05 | | | | | | | |
+| GA-Faster RCNN | R-101-FPN | caffe | 1x | F | 0.001 | | | | | | | |
+| GA-RetinaNet | R-101-FPN | caffe | 1x | F | 0.05 | | | | | | | |
+| GA-RetinaNet | R-101-FPN | caffe | 2x | T | 0.05 | | | | | | | |
+
+## Citation
+
+We provide config files to reproduce the results in the CVPR 2019 paper for [Region Proposal by Guided Anchoring](https://arxiv.org/abs/1901.03278).
+
+```latex
+@inproceedings{wang2019region,
+ title={Region Proposal by Guided Anchoring},
+ author={Jiaqi Wang and Kai Chen and Shuo Yang and Chen Change Loy and Dahua Lin},
+ booktitle={IEEE Conference on Computer Vision and Pattern Recognition},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-fast-rcnn_r50-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-fast-rcnn_r50-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..2d0579c53cb23d71d0bec57387f413cc39449e93
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-fast-rcnn_r50-caffe_fpn_1x_coco.py
@@ -0,0 +1,66 @@
+_base_ = '../fast_rcnn/fast-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')),
+ roi_head=dict(
+ bbox_head=dict(bbox_coder=dict(target_stds=[0.05, 0.05, 0.1, 0.1]))),
+ # model training and testing settings
+ train_cfg=dict(
+ rcnn=dict(
+ assigner=dict(pos_iou_thr=0.6, neg_iou_thr=0.6, min_pos_iou=0.6),
+ sampler=dict(num=256))),
+ test_cfg=dict(rcnn=dict(score_thr=1e-3)))
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+img_norm_cfg = dict(
+ mean=[103.530, 116.280, 123.675], std=[1.0, 1.0, 1.0], to_rgb=False)
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadProposals', num_max_proposals=300),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='Resize', img_scale=(1333, 800), keep_ratio=True),
+ dict(type='RandomFlip', flip_ratio=0.5),
+ dict(type='Normalize', **img_norm_cfg),
+ dict(type='Pad', size_divisor=32),
+ dict(type='DefaultFormatBundle'),
+ dict(type='Collect', keys=['img', 'proposals', 'gt_bboxes', 'gt_labels']),
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadProposals', num_max_proposals=None),
+ dict(
+ type='MultiScaleFlipAug',
+ img_scale=(1333, 800),
+ flip=False,
+ transforms=[
+ dict(type='Resize', keep_ratio=True),
+ dict(type='RandomFlip'),
+ dict(type='Normalize', **img_norm_cfg),
+ dict(type='Pad', size_divisor=32),
+ dict(type='ImageToTensor', keys=['img']),
+ dict(type='Collect', keys=['img', 'proposals']),
+ ])
+]
+# TODO: support loading proposals
+data = dict(
+ train=dict(
+ proposal_file=data_root + 'proposals/ga_rpn_r50_fpn_1x_train2017.pkl',
+ pipeline=train_pipeline),
+ val=dict(
+ proposal_file=data_root + 'proposals/ga_rpn_r50_fpn_1x_val2017.pkl',
+ pipeline=test_pipeline),
+ test=dict(
+ proposal_file=data_root + 'proposals/ga_rpn_r50_fpn_1x_val2017.pkl',
+ pipeline=test_pipeline))
+optimizer_config = dict(
+ _delete_=True, grad_clip=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-faster-rcnn_r101-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-faster-rcnn_r101-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f585dc355ac7dc10e75875f6b9f739fe669912bb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-faster-rcnn_r101-caffe_fpn_1x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './ga-faster-rcnn_r50-caffe_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet101_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-faster-rcnn_r50-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-faster-rcnn_r50-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6cd44de557bfb20b4298099bd0972e3327b410cb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-faster-rcnn_r50-caffe_fpn_1x_coco.py
@@ -0,0 +1,64 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50-caffe_fpn_1x_coco.py'
+model = dict(
+ rpn_head=dict(
+ _delete_=True,
+ type='GARPNHead',
+ in_channels=256,
+ feat_channels=256,
+ approx_anchor_generator=dict(
+ type='AnchorGenerator',
+ octave_base_scale=8,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[4, 8, 16, 32, 64]),
+ square_anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ scales=[8],
+ strides=[4, 8, 16, 32, 64]),
+ anchor_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.07, 0.07, 0.14, 0.14]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.07, 0.07, 0.11, 0.11]),
+ loc_filter_thr=0.01,
+ loss_loc=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_shape=dict(type='BoundedIoULoss', beta=0.2, loss_weight=1.0),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0)),
+ roi_head=dict(
+ bbox_head=dict(bbox_coder=dict(target_stds=[0.05, 0.05, 0.1, 0.1]))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ ga_assigner=dict(
+ type='ApproxMaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ ignore_iof_thr=-1),
+ ga_sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=-1,
+ center_ratio=0.2,
+ ignore_ratio=0.5),
+ rpn_proposal=dict(nms_post=1000, max_per_img=300),
+ rcnn=dict(
+ assigner=dict(pos_iou_thr=0.6, neg_iou_thr=0.6, min_pos_iou=0.6),
+ sampler=dict(type='RandomSampler', num=256))),
+ test_cfg=dict(
+ rpn=dict(nms_post=1000, max_per_img=300), rcnn=dict(score_thr=1e-3)))
+optim_wrapper = dict(clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-faster-rcnn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-faster-rcnn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3007fbec42016fa8c6b90ba5b0b4e772d0e865f7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-faster-rcnn_r50_fpn_1x_coco.py
@@ -0,0 +1,64 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ rpn_head=dict(
+ _delete_=True,
+ type='GARPNHead',
+ in_channels=256,
+ feat_channels=256,
+ approx_anchor_generator=dict(
+ type='AnchorGenerator',
+ octave_base_scale=8,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[4, 8, 16, 32, 64]),
+ square_anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ scales=[8],
+ strides=[4, 8, 16, 32, 64]),
+ anchor_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.07, 0.07, 0.14, 0.14]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.07, 0.07, 0.11, 0.11]),
+ loc_filter_thr=0.01,
+ loss_loc=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_shape=dict(type='BoundedIoULoss', beta=0.2, loss_weight=1.0),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0)),
+ roi_head=dict(
+ bbox_head=dict(bbox_coder=dict(target_stds=[0.05, 0.05, 0.1, 0.1]))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ ga_assigner=dict(
+ type='ApproxMaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ ignore_iof_thr=-1),
+ ga_sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=-1,
+ center_ratio=0.2,
+ ignore_ratio=0.5),
+ rpn_proposal=dict(nms_post=1000, max_per_img=300),
+ rcnn=dict(
+ assigner=dict(pos_iou_thr=0.6, neg_iou_thr=0.6, min_pos_iou=0.6),
+ sampler=dict(type='RandomSampler', num=256))),
+ test_cfg=dict(
+ rpn=dict(nms_post=1000, max_per_img=300), rcnn=dict(score_thr=1e-3)))
+optim_wrapper = dict(clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-faster-rcnn_x101-32x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-faster-rcnn_x101-32x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8a22a1ec01e66854c68968f65802dc117aa59953
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-faster-rcnn_x101-32x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './ga-faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-faster-rcnn_x101-64x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-faster-rcnn_x101-64x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3d6aaeaa7187deaa2c0da73a89bf14980a3405db
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-faster-rcnn_x101-64x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './ga-faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-retinanet_r101-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-retinanet_r101-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..9adbae55eea2311800ccbc8e01e3f41521c7040b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-retinanet_r101-caffe_fpn_1x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './ga-retinanet_r50-caffe_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet101_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-retinanet_r101-caffe_fpn_ms-2x.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-retinanet_r101-caffe_fpn_ms-2x.py
new file mode 100644
index 0000000000000000000000000000000000000000..012e89b8338c69c4ffdf4182827a185233945288
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-retinanet_r101-caffe_fpn_ms-2x.py
@@ -0,0 +1,34 @@
+_base_ = './ga-retinanet_r101-caffe_fpn_1x_coco.py'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize', scale=[(1333, 480), (1333, 960)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+# learning policy
+max_epochs = 24
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 3.0,
+ by_epoch=False,
+ begin=0,
+ end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-retinanet_r50-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-retinanet_r50-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b62aba62c64870977c7c8fe4021a361c8871b633
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-retinanet_r50-caffe_fpn_1x_coco.py
@@ -0,0 +1,61 @@
+_base_ = '../retinanet/retinanet_r50-caffe_fpn_1x_coco.py'
+model = dict(
+ bbox_head=dict(
+ _delete_=True,
+ type='GARetinaHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ approx_anchor_generator=dict(
+ type='AnchorGenerator',
+ octave_base_scale=4,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[8, 16, 32, 64, 128]),
+ square_anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ scales=[4],
+ strides=[8, 16, 32, 64, 128]),
+ anchor_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loc_filter_thr=0.01,
+ loss_loc=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_shape=dict(type='BoundedIoULoss', beta=0.2, loss_weight=1.0),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=0.04, loss_weight=1.0)),
+ # training and testing settings
+ train_cfg=dict(
+ ga_assigner=dict(
+ type='ApproxMaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.4,
+ min_pos_iou=0.4,
+ ignore_iof_thr=-1),
+ ga_sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ assigner=dict(neg_iou_thr=0.5, min_pos_iou=0.0),
+ center_ratio=0.2,
+ ignore_ratio=0.5))
+optim_wrapper = dict(clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-retinanet_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-retinanet_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..da39c7005b26d65cca0ae122bf078db2d8ad2786
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-retinanet_r50_fpn_1x_coco.py
@@ -0,0 +1,61 @@
+_base_ = '../retinanet/retinanet_r50_fpn_1x_coco.py'
+model = dict(
+ bbox_head=dict(
+ _delete_=True,
+ type='GARetinaHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ approx_anchor_generator=dict(
+ type='AnchorGenerator',
+ octave_base_scale=4,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[8, 16, 32, 64, 128]),
+ square_anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ scales=[4],
+ strides=[8, 16, 32, 64, 128]),
+ anchor_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loc_filter_thr=0.01,
+ loss_loc=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_shape=dict(type='BoundedIoULoss', beta=0.2, loss_weight=1.0),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=0.04, loss_weight=1.0)),
+ # training and testing settings
+ train_cfg=dict(
+ ga_assigner=dict(
+ type='ApproxMaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.4,
+ min_pos_iou=0.4,
+ ignore_iof_thr=-1),
+ ga_sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ assigner=dict(neg_iou_thr=0.5, min_pos_iou=0.0),
+ center_ratio=0.2,
+ ignore_ratio=0.5))
+optim_wrapper = dict(clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-retinanet_x101-32x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-retinanet_x101-32x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..478a8e5e4a2192e23329564ac688ac40c93110dd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-retinanet_x101-32x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './ga-retinanet_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-retinanet_x101-64x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-retinanet_x101-64x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..cb7721d3a604277977b102d431076d6d58a7d457
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-retinanet_x101-64x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './ga-retinanet_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-rpn_r101-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-rpn_r101-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b375c874ac8cabf5ad29aacc51e1065d14d83ee1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-rpn_r101-caffe_fpn_1x_coco.py
@@ -0,0 +1,8 @@
+_base_ = './ga-rpn_r50-caffe_fpn_1x_coco.py'
+# model settings
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet101_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-rpn_r50-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-rpn_r50-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..aa58426effe8bedbe9ffb907153b98d51bef5ef2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-rpn_r50-caffe_fpn_1x_coco.py
@@ -0,0 +1,57 @@
+_base_ = '../rpn/rpn_r50-caffe_fpn_1x_coco.py'
+model = dict(
+ rpn_head=dict(
+ _delete_=True,
+ type='GARPNHead',
+ in_channels=256,
+ feat_channels=256,
+ approx_anchor_generator=dict(
+ type='AnchorGenerator',
+ octave_base_scale=8,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[4, 8, 16, 32, 64]),
+ square_anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ scales=[8],
+ strides=[4, 8, 16, 32, 64]),
+ anchor_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.07, 0.07, 0.14, 0.14]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.07, 0.07, 0.11, 0.11]),
+ loc_filter_thr=0.01,
+ loss_loc=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_shape=dict(type='BoundedIoULoss', beta=0.2, loss_weight=1.0),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0)),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ ga_assigner=dict(
+ type='ApproxMaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ ignore_iof_thr=-1),
+ ga_sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=-1,
+ center_ratio=0.2,
+ ignore_ratio=0.5)),
+ test_cfg=dict(rpn=dict(nms_post=1000)))
+optim_wrapper = dict(clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-rpn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-rpn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..2973f272b740c8deec74f6c24798a2d80d917946
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-rpn_r50_fpn_1x_coco.py
@@ -0,0 +1,57 @@
+_base_ = '../rpn/rpn_r50_fpn_1x_coco.py'
+model = dict(
+ rpn_head=dict(
+ _delete_=True,
+ type='GARPNHead',
+ in_channels=256,
+ feat_channels=256,
+ approx_anchor_generator=dict(
+ type='AnchorGenerator',
+ octave_base_scale=8,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[4, 8, 16, 32, 64]),
+ square_anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ scales=[8],
+ strides=[4, 8, 16, 32, 64]),
+ anchor_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.07, 0.07, 0.14, 0.14]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.07, 0.07, 0.11, 0.11]),
+ loc_filter_thr=0.01,
+ loss_loc=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_shape=dict(type='BoundedIoULoss', beta=0.2, loss_weight=1.0),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0)),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ ga_assigner=dict(
+ type='ApproxMaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ ignore_iof_thr=-1),
+ ga_sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=-1,
+ center_ratio=0.2,
+ ignore_ratio=0.5)),
+ test_cfg=dict(rpn=dict(nms_post=1000)))
+optim_wrapper = dict(clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-rpn_x101-32x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-rpn_x101-32x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..276d45d8c21fa1eba130e834671bdddd794fa1f5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-rpn_x101-32x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './ga-rpn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-rpn_x101-64x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-rpn_x101-64x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f29fe9aa20054f3152e290df5ca75363dff6a4ce
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/ga-rpn_x101-64x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './ga-rpn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..516b3e93fc2b10fb563de1b377144da103ef4523
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/guided_anchoring/metafile.yml
@@ -0,0 +1,246 @@
+Collections:
+ - Name: Guided Anchoring
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - FPN
+ - Guided Anchoring
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/1901.03278
+ Title: 'Region Proposal by Guided Anchoring'
+ README: configs/guided_anchoring/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/dense_heads/ga_retina_head.py#L10
+ Version: v2.0.0
+
+Models:
+ - Name: ga-rpn_r50-caffe_fpn_1x_coco
+ In Collection: Guided Anchoring
+ Config: configs/guided_anchoring/ga-rpn_r50-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.3
+ inference time (ms/im):
+ - value: 63.29
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Region Proposal
+ Dataset: COCO
+ Metrics:
+ AR@1000: 68.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_rpn_r50_caffe_fpn_1x_coco/ga_rpn_r50_caffe_fpn_1x_coco_20200531-899008a6.pth
+
+ - Name: ga-rpn_r101-caffe_fpn_1x_coco
+ In Collection: Guided Anchoring
+ Config: configs/guided_anchoring/ga-rpn_r101-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.3
+ inference time (ms/im):
+ - value: 76.92
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Region Proposal
+ Dataset: COCO
+ Metrics:
+ AR@1000: 69.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_rpn_r101_caffe_fpn_1x_coco/ga_rpn_r101_caffe_fpn_1x_coco_20200531-ca9ba8fb.pth
+
+ - Name: ga-rpn_x101-32x4d_fpn_1x_coco
+ In Collection: Guided Anchoring
+ Config: configs/guided_anchoring/ga-rpn_x101-32x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 8.5
+ inference time (ms/im):
+ - value: 100
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Region Proposal
+ Dataset: COCO
+ Metrics:
+ AR@1000: 70.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_rpn_x101_32x4d_fpn_1x_coco/ga_rpn_x101_32x4d_fpn_1x_coco_20200220-c28d1b18.pth
+
+ - Name: ga-rpn_x101-64x4d_fpn_1x_coco
+ In Collection: Guided Anchoring
+ Config: configs/guided_anchoring/ga-rpn_x101-64x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.1
+ inference time (ms/im):
+ - value: 133.33
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Region Proposal
+ Dataset: COCO
+ Metrics:
+ AR@1000: 70.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_rpn_x101_64x4d_fpn_1x_coco/ga_rpn_x101_64x4d_fpn_1x_coco_20200225-3c6e1aa2.pth
+
+ - Name: ga-faster-rcnn_r50-caffe_fpn_1x_coco
+ In Collection: Guided Anchoring
+ Config: configs/guided_anchoring/ga-faster-rcnn_r50-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.5
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_faster_r50_caffe_fpn_1x_coco/ga_faster_r50_caffe_fpn_1x_coco_20200702_000718-a11ccfe6.pth
+
+ - Name: ga-faster-rcnn_r101-caffe_fpn_1x_coco
+ In Collection: Guided Anchoring
+ Config: configs/guided_anchoring/ga-faster-rcnn_r101-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.5
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_faster_r101_caffe_fpn_1x_coco/ga_faster_r101_caffe_fpn_1x_coco_bbox_mAP-0.415_20200505_115528-fb82e499.pth
+
+ - Name: ga-faster-rcnn_x101-32x4d_fpn_1x_coco
+ In Collection: Guided Anchoring
+ Config: configs/guided_anchoring/ga-faster-rcnn_x101-32x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 8.7
+ inference time (ms/im):
+ - value: 103.09
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_faster_x101_32x4d_fpn_1x_coco/ga_faster_x101_32x4d_fpn_1x_coco_20200215-1ded9da3.pth
+
+ - Name: ga-faster-rcnn_x101-64x4d_fpn_1x_coco
+ In Collection: Guided Anchoring
+ Config: configs/guided_anchoring/ga-faster-rcnn_x101-64x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 11.8
+ inference time (ms/im):
+ - value: 136.99
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_faster_x101_64x4d_fpn_1x_coco/ga_faster_x101_64x4d_fpn_1x_coco_20200215-0fa7bde7.pth
+
+ - Name: ga-retinanet_r50-caffe_fpn_1x_coco
+ In Collection: Guided Anchoring
+ Config: configs/guided_anchoring/ga-retinanet_r50-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.5
+ inference time (ms/im):
+ - value: 59.52
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 36.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_retinanet_r50_caffe_fpn_1x_coco/ga_retinanet_r50_caffe_fpn_1x_coco_20201020-39581c6f.pth
+
+ - Name: ga-retinanet_r101-caffe_fpn_1x_coco
+ In Collection: Guided Anchoring
+ Config: configs/guided_anchoring/ga-retinanet_r101-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.5
+ inference time (ms/im):
+ - value: 77.52
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_retinanet_r101_caffe_fpn_1x_coco/ga_retinanet_r101_caffe_fpn_1x_coco_20200531-6266453c.pth
+
+ - Name: ga-retinanet_x101-32x4d_fpn_1x_coco
+ In Collection: Guided Anchoring
+ Config: configs/guided_anchoring/ga-retinanet_x101-32x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.9
+ inference time (ms/im):
+ - value: 94.34
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_retinanet_x101_32x4d_fpn_1x_coco/ga_retinanet_x101_32x4d_fpn_1x_coco_20200219-40c56caa.pth
+
+ - Name: ga-retinanet_x101-64x4d_fpn_1x_coco
+ In Collection: Guided Anchoring
+ Config: configs/guided_anchoring/ga-retinanet_x101-64x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 9.9
+ inference time (ms/im):
+ - value: 129.87
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/guided_anchoring/ga_retinanet_x101_64x4d_fpn_1x_coco/ga_retinanet_x101_64x4d_fpn_1x_coco_20200226-ef9f7f1f.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..fc1ed0cc94e778ad56504b9fa8050ad8237c4c11
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/README.md
@@ -0,0 +1,101 @@
+# HRNet
+
+> [Deep High-Resolution Representation Learning for Human Pose Estimation](https://arxiv.org/abs/1902.09212)
+
+
+
+## Abstract
+
+This is an official pytorch implementation of Deep High-Resolution Representation Learning for Human Pose Estimation. In this work, we are interested in the human pose estimation problem with a focus on learning reliable high-resolution representations. Most existing methods recover high-resolution representations from low-resolution representations produced by a high-to-low resolution network. Instead, our proposed network maintains high-resolution representations through the whole process. We start from a high-resolution subnetwork as the first stage, gradually add high-to-low resolution subnetworks one by one to form more stages, and connect the mutli-resolution subnetworks in parallel. We conduct repeated multi-scale fusions such that each of the high-to-low resolution representations receives information from other parallel representations over and over, leading to rich high-resolution representations. As a result, the predicted keypoint heatmap is potentially more accurate and spatially more precise. We empirically demonstrate the effectiveness of our network through the superior pose estimation results over two benchmark datasets: the COCO keypoint detection dataset and the MPII Human Pose dataset.
+
+High-resolution representation learning plays an essential role in many vision problems, e.g., pose estimation and semantic segmentation. The high-resolution network (HRNet), recently developed for human pose estimation, maintains high-resolution representations through the whole process by connecting high-to-low resolution convolutions in parallel and produces strong high-resolution representations by repeatedly conducting fusions across parallel convolutions.
+In this paper, we conduct a further study on high-resolution representations by introducing a simple yet effective modification and apply it to a wide range of vision tasks. We augment the high-resolution representation by aggregating the (upsampled) representations from all the parallel convolutions rather than only the representation from the high-resolution convolution as done in HRNet. This simple modification leads to stronger representations, evidenced by superior results. We show top results in semantic segmentation on Cityscapes, LIP, and PASCAL Context, and facial landmark detection on AFLW, COFW, 300W, and WFLW. In addition, we build a multi-level representation from the high-resolution representation and apply it to the Faster R-CNN object detection framework and the extended frameworks. The proposed approach achieves superior results to existing single-model networks on COCO object detection.
+
+
+

+
+
+## Results and Models
+
+### Faster R-CNN
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :----------: | :-----: | :-----: | :------: | :------------: | :----: | :---------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| HRNetV2p-W18 | pytorch | 1x | 6.6 | 13.4 | 36.9 | [config](./faster-rcnn_hrnetv2p-w18-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/faster_rcnn_hrnetv2p_w18_1x_coco/faster_rcnn_hrnetv2p_w18_1x_coco_20200130-56651a6d.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/faster_rcnn_hrnetv2p_w18_1x_coco/faster_rcnn_hrnetv2p_w18_1x_coco_20200130_211246.log.json) |
+| HRNetV2p-W18 | pytorch | 2x | 6.6 | - | 38.9 | [config](./faster-rcnn_hrnetv2p-w18-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/faster_rcnn_hrnetv2p_w18_2x_coco/faster_rcnn_hrnetv2p_w18_2x_coco_20200702_085731-a4ec0611.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/faster_rcnn_hrnetv2p_w18_2x_coco/faster_rcnn_hrnetv2p_w18_2x_coco_20200702_085731.log.json) |
+| HRNetV2p-W32 | pytorch | 1x | 9.0 | 12.4 | 40.2 | [config](./faster-rcnn_hrnetv2p-w32-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/faster_rcnn_hrnetv2p_w32_1x_coco/faster_rcnn_hrnetv2p_w32_1x_coco_20200130-6e286425.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/faster_rcnn_hrnetv2p_w32_1x_coco/faster_rcnn_hrnetv2p_w32_1x_coco_20200130_204442.log.json) |
+| HRNetV2p-W32 | pytorch | 2x | 9.0 | - | 41.4 | [config](./faster-rcnn_hrnetv2p-w32_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/faster_rcnn_hrnetv2p_w32_2x_coco/faster_rcnn_hrnetv2p_w32_2x_coco_20200529_015927-976a9c15.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/faster_rcnn_hrnetv2p_w32_2x_coco/faster_rcnn_hrnetv2p_w32_2x_coco_20200529_015927.log.json) |
+| HRNetV2p-W40 | pytorch | 1x | 10.4 | 10.5 | 41.2 | [config](./faster-rcnn_hrnetv2p-w40-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/faster_rcnn_hrnetv2p_w40_1x_coco/faster_rcnn_hrnetv2p_w40_1x_coco_20200210-95c1f5ce.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/faster_rcnn_hrnetv2p_w40_1x_coco/faster_rcnn_hrnetv2p_w40_1x_coco_20200210_125315.log.json) |
+| HRNetV2p-W40 | pytorch | 2x | 10.4 | - | 42.1 | [config](./faster-rcnn_hrnetv2p-w40_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/faster_rcnn_hrnetv2p_w40_2x_coco/faster_rcnn_hrnetv2p_w40_2x_coco_20200512_161033-0f236ef4.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/faster_rcnn_hrnetv2p_w40_2x_coco/faster_rcnn_hrnetv2p_w40_2x_coco_20200512_161033.log.json) |
+
+### Mask R-CNN
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :----------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :-------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| HRNetV2p-W18 | pytorch | 1x | 7.0 | 11.7 | 37.7 | 34.2 | [config](./mask-rcnn_hrnetv2p-w18-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/mask_rcnn_hrnetv2p_w18_1x_coco/mask_rcnn_hrnetv2p_w18_1x_coco_20200205-1c3d78ed.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/mask_rcnn_hrnetv2p_w18_1x_coco/mask_rcnn_hrnetv2p_w18_1x_coco_20200205_232523.log.json) |
+| HRNetV2p-W18 | pytorch | 2x | 7.0 | - | 39.8 | 36.0 | [config](./mask-rcnn_hrnetv2p-w18-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/mask_rcnn_hrnetv2p_w18_2x_coco/mask_rcnn_hrnetv2p_w18_2x_coco_20200212-b3c825b1.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/mask_rcnn_hrnetv2p_w18_2x_coco/mask_rcnn_hrnetv2p_w18_2x_coco_20200212_134222.log.json) |
+| HRNetV2p-W32 | pytorch | 1x | 9.4 | 11.3 | 41.2 | 37.1 | [config](./mask-rcnn_hrnetv2p-w32-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/mask_rcnn_hrnetv2p_w32_1x_coco/mask_rcnn_hrnetv2p_w32_1x_coco_20200207-b29f616e.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/mask_rcnn_hrnetv2p_w32_1x_coco/mask_rcnn_hrnetv2p_w32_1x_coco_20200207_055017.log.json) |
+| HRNetV2p-W32 | pytorch | 2x | 9.4 | - | 42.5 | 37.8 | [config](./mask-rcnn_hrnetv2p-w32-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/mask_rcnn_hrnetv2p_w32_2x_coco/mask_rcnn_hrnetv2p_w32_2x_coco_20200213-45b75b4d.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/mask_rcnn_hrnetv2p_w32_2x_coco/mask_rcnn_hrnetv2p_w32_2x_coco_20200213_150518.log.json) |
+| HRNetV2p-W40 | pytorch | 1x | 10.9 | | 42.1 | 37.5 | [config](./mask-rcnn_hrnetv2p-w40_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/mask_rcnn_hrnetv2p_w40_1x_coco/mask_rcnn_hrnetv2p_w40_1x_coco_20200511_015646-66738b35.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/mask_rcnn_hrnetv2p_w40_1x_coco/mask_rcnn_hrnetv2p_w40_1x_coco_20200511_015646.log.json) |
+| HRNetV2p-W40 | pytorch | 2x | 10.9 | | 42.8 | 38.2 | [config](./mask-rcnn_hrnetv2p-w40-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/mask_rcnn_hrnetv2p_w40_2x_coco/mask_rcnn_hrnetv2p_w40_2x_coco_20200512_163732-aed5e4ab.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/mask_rcnn_hrnetv2p_w40_2x_coco/mask_rcnn_hrnetv2p_w40_2x_coco_20200512_163732.log.json) |
+
+### Cascade R-CNN
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :----------: | :-----: | :-----: | :------: | :------------: | :----: | :-----------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| HRNetV2p-W18 | pytorch | 20e | 7.0 | 11.0 | 41.2 | [config](./cascade-rcnn_hrnetv2p-w18-20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/cascade_rcnn_hrnetv2p_w18_20e_coco/cascade_rcnn_hrnetv2p_w18_20e_coco_20200210-434be9d7.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/cascade_rcnn_hrnetv2p_w18_20e_coco/cascade_rcnn_hrnetv2p_w18_20e_coco_20200210_105632.log.json) |
+| HRNetV2p-W32 | pytorch | 20e | 9.4 | 11.0 | 43.3 | [config](./cascade-rcnn_hrnetv2p-w32-20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/cascade_rcnn_hrnetv2p_w32_20e_coco/cascade_rcnn_hrnetv2p_w32_20e_coco_20200208-928455a4.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/cascade_rcnn_hrnetv2p_w32_20e_coco/cascade_rcnn_hrnetv2p_w32_20e_coco_20200208_160511.log.json) |
+| HRNetV2p-W40 | pytorch | 20e | 10.8 | | 43.8 | [config](./cascade-rcnn_hrnetv2p-w40-20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/cascade_rcnn_hrnetv2p_w40_20e_coco/cascade_rcnn_hrnetv2p_w40_20e_coco_20200512_161112-75e47b04.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/cascade_rcnn_hrnetv2p_w40_20e_coco/cascade_rcnn_hrnetv2p_w40_20e_coco_20200512_161112.log.json) |
+
+### Cascade Mask R-CNN
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :----------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :----------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| HRNetV2p-W18 | pytorch | 20e | 8.5 | 8.5 | 41.6 | 36.4 | [config](./cascade-mask-rcnn_hrnetv2p-w18_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/cascade_mask_rcnn_hrnetv2p_w18_20e_coco/cascade_mask_rcnn_hrnetv2p_w18_20e_coco_20200210-b543cd2b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/cascade_mask_rcnn_hrnetv2p_w18_20e_coco/cascade_mask_rcnn_hrnetv2p_w18_20e_coco_20200210_093149.log.json) |
+| HRNetV2p-W32 | pytorch | 20e | | 8.3 | 44.3 | 38.6 | [config](./cascade-mask-rcnn_hrnetv2p-w32_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/cascade_mask_rcnn_hrnetv2p_w32_20e_coco/cascade_mask_rcnn_hrnetv2p_w32_20e_coco_20200512_154043-39d9cf7b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/cascade_mask_rcnn_hrnetv2p_w32_20e_coco/cascade_mask_rcnn_hrnetv2p_w32_20e_coco_20200512_154043.log.json) |
+| HRNetV2p-W40 | pytorch | 20e | 12.5 | | 45.1 | 39.3 | [config](./cascade-mask-rcnn_hrnetv2p-w40-20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/cascade_mask_rcnn_hrnetv2p_w40_20e_coco/cascade_mask_rcnn_hrnetv2p_w40_20e_coco_20200527_204922-969c4610.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/cascade_mask_rcnn_hrnetv2p_w40_20e_coco/cascade_mask_rcnn_hrnetv2p_w40_20e_coco_20200527_204922.log.json) |
+
+### Hybrid Task Cascade (HTC)
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :----------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :--------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| HRNetV2p-W18 | pytorch | 20e | 10.8 | 4.7 | 42.8 | 37.9 | [config](./htc_hrnetv2p-w18_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/htc_hrnetv2p_w18_20e_coco/htc_hrnetv2p_w18_20e_coco_20200210-b266988c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/htc_hrnetv2p_w18_20e_coco/htc_hrnetv2p_w18_20e_coco_20200210_182735.log.json) |
+| HRNetV2p-W32 | pytorch | 20e | 13.1 | 4.9 | 45.4 | 39.9 | [config](./htc_hrnetv2p-w32_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/htc_hrnetv2p_w32_20e_coco/htc_hrnetv2p_w32_20e_coco_20200207-7639fa12.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/htc_hrnetv2p_w32_20e_coco/htc_hrnetv2p_w32_20e_coco_20200207_193153.log.json) |
+| HRNetV2p-W40 | pytorch | 20e | 14.6 | | 46.4 | 40.8 | [config](./htc_hrnetv2p-w40_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/htc_hrnetv2p_w40_20e_coco/htc_hrnetv2p_w40_20e_coco_20200529_183411-417c4d5b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/htc_hrnetv2p_w40_20e_coco/htc_hrnetv2p_w40_20e_coco_20200529_183411.log.json) |
+
+### FCOS
+
+| Backbone | Style | GN | MS train | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :----------: | :-----: | :-: | :------: | :-----: | :------: | :------------: | :----: | :--------------------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| HRNetV2p-W18 | pytorch | Y | N | 1x | 13.0 | 12.9 | 35.3 | [config](./fcos_hrnetv2p-w18-gn-head_4xb4-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w18_gn-head_4x4_1x_coco/fcos_hrnetv2p_w18_gn-head_4x4_1x_coco_20201212_100710-4ad151de.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w18_gn-head_4x4_1x_coco/fcos_hrnetv2p_w18_gn-head_4x4_1x_coco_20201212_100710.log.json) |
+| HRNetV2p-W18 | pytorch | Y | N | 2x | 13.0 | - | 38.2 | [config](./fcos_hrnetv2p-w18-gn-head_4xb4-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w18_gn-head_4x4_2x_coco/fcos_hrnetv2p_w18_gn-head_4x4_2x_coco_20201212_101110-5c575fa5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w18_gn-head_4x4_2x_coco/fcos_hrnetv2p_w18_gn-head_4x4_2x_coco_20201212_101110.log.json) |
+| HRNetV2p-W32 | pytorch | Y | N | 1x | 17.5 | 12.9 | 39.5 | [config](./fcos_hrnetv2p-w32-gn-head_4xb4-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w32_gn-head_4x4_1x_coco/fcos_hrnetv2p_w32_gn-head_4x4_1x_coco_20201211_134730-cb8055c0.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w32_gn-head_4x4_1x_coco/fcos_hrnetv2p_w32_gn-head_4x4_1x_coco_20201211_134730.log.json) |
+| HRNetV2p-W32 | pytorch | Y | N | 2x | 17.5 | - | 40.8 | [config](./fcos_hrnetv2p-w32-gn-head_4xb4-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w32_gn-head_4x4_2x_coco/fcos_hrnetv2p_w32_gn-head_4x4_2x_coco_20201212_112133-77b6b9bb.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w32_gn-head_4x4_2x_coco/fcos_hrnetv2p_w32_gn-head_4x4_2x_coco_20201212_112133.log.json) |
+| HRNetV2p-W18 | pytorch | Y | Y | 2x | 13.0 | 12.9 | 38.3 | [config](./fcos_hrnetv2p-w18-gn-head_ms-640-800-4xb4-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w18_gn-head_mstrain_640-800_4x4_2x_coco/fcos_hrnetv2p_w18_gn-head_mstrain_640-800_4x4_2x_coco_20201212_111651-441e9d9f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w18_gn-head_mstrain_640-800_4x4_2x_coco/fcos_hrnetv2p_w18_gn-head_mstrain_640-800_4x4_2x_coco_20201212_111651.log.json) |
+| HRNetV2p-W32 | pytorch | Y | Y | 2x | 17.5 | 12.4 | 41.9 | [config](./fcos_hrnetv2p-w32-gn-head_ms-640-800-4xb4-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w32_gn-head_mstrain_640-800_4x4_2x_coco/fcos_hrnetv2p_w32_gn-head_mstrain_640-800_4x4_2x_coco_20201212_090846-b6f2b49f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w32_gn-head_mstrain_640-800_4x4_2x_coco/fcos_hrnetv2p_w32_gn-head_mstrain_640-800_4x4_2x_coco_20201212_090846.log.json) |
+| HRNetV2p-W48 | pytorch | Y | Y | 2x | 20.3 | 10.8 | 42.7 | [config](./fcos_hrnetv2p-w40-gn-head_ms-640-800-4xb4-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w40_gn-head_mstrain_640-800_4x4_2x_coco/fcos_hrnetv2p_w40_gn-head_mstrain_640-800_4x4_2x_coco_20201212_124752-f22d2ce5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w40_gn-head_mstrain_640-800_4x4_2x_coco/fcos_hrnetv2p_w40_gn-head_mstrain_640-800_4x4_2x_coco_20201212_124752.log.json) |
+
+**Note:**
+
+- The `28e` schedule in HTC indicates decreasing the lr at 24 and 27 epochs, with a total of 28 epochs.
+- HRNetV2 ImageNet pretrained models are in [HRNets for Image Classification](https://github.com/HRNet/HRNet-Image-Classification).
+
+## Citation
+
+```latex
+@inproceedings{SunXLW19,
+ title={Deep High-Resolution Representation Learning for Human Pose Estimation},
+ author={Ke Sun and Bin Xiao and Dong Liu and Jingdong Wang},
+ booktitle={CVPR},
+ year={2019}
+}
+
+@article{SunZJCXLMWLW19,
+ title={High-Resolution Representations for Labeling Pixels and Regions},
+ author={Ke Sun and Yang Zhao and Borui Jiang and Tianheng Cheng and Bin Xiao
+ and Dong Liu and Yadong Mu and Xinggang Wang and Wenyu Liu and Jingdong Wang},
+ journal = {CoRR},
+ volume = {abs/1904.04514},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/cascade-mask-rcnn_hrnetv2p-w18_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/cascade-mask-rcnn_hrnetv2p-w18_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5ca0ebfe43b00886b22ffc426c5ac89a50f4fda6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/cascade-mask-rcnn_hrnetv2p-w18_20e_coco.py
@@ -0,0 +1,11 @@
+_base_ = './cascade-mask-rcnn_hrnetv2p-w32_20e_coco.py'
+# model settings
+model = dict(
+ backbone=dict(
+ extra=dict(
+ stage2=dict(num_channels=(18, 36)),
+ stage3=dict(num_channels=(18, 36, 72)),
+ stage4=dict(num_channels=(18, 36, 72, 144))),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://msra/hrnetv2_w18')),
+ neck=dict(type='HRFPN', in_channels=[18, 36, 72, 144], out_channels=256))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/cascade-mask-rcnn_hrnetv2p-w32_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/cascade-mask-rcnn_hrnetv2p-w32_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1ffedc3916748c3c6b333023110e56895de7e4bd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/cascade-mask-rcnn_hrnetv2p-w32_20e_coco.py
@@ -0,0 +1,51 @@
+_base_ = '../cascade_rcnn/cascade-mask-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='HRNet',
+ extra=dict(
+ stage1=dict(
+ num_modules=1,
+ num_branches=1,
+ block='BOTTLENECK',
+ num_blocks=(4, ),
+ num_channels=(64, )),
+ stage2=dict(
+ num_modules=1,
+ num_branches=2,
+ block='BASIC',
+ num_blocks=(4, 4),
+ num_channels=(32, 64)),
+ stage3=dict(
+ num_modules=4,
+ num_branches=3,
+ block='BASIC',
+ num_blocks=(4, 4, 4),
+ num_channels=(32, 64, 128)),
+ stage4=dict(
+ num_modules=3,
+ num_branches=4,
+ block='BASIC',
+ num_blocks=(4, 4, 4, 4),
+ num_channels=(32, 64, 128, 256))),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://msra/hrnetv2_w32')),
+ neck=dict(
+ _delete_=True,
+ type='HRFPN',
+ in_channels=[32, 64, 128, 256],
+ out_channels=256))
+# learning policy
+max_epochs = 20
+train_cfg = dict(max_epochs=max_epochs)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 19],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/cascade-mask-rcnn_hrnetv2p-w40-20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/cascade-mask-rcnn_hrnetv2p-w40-20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..4a51a02412871905d947bcbb648b1a24e5033f56
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/cascade-mask-rcnn_hrnetv2p-w40-20e_coco.py
@@ -0,0 +1,12 @@
+_base_ = './cascade-mask-rcnn_hrnetv2p-w32_20e_coco.py'
+# model settings
+model = dict(
+ backbone=dict(
+ type='HRNet',
+ extra=dict(
+ stage2=dict(num_channels=(40, 80)),
+ stage3=dict(num_channels=(40, 80, 160)),
+ stage4=dict(num_channels=(40, 80, 160, 320))),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://msra/hrnetv2_w40')),
+ neck=dict(type='HRFPN', in_channels=[40, 80, 160, 320], out_channels=256))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/cascade-rcnn_hrnetv2p-w18-20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/cascade-rcnn_hrnetv2p-w18-20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8834c1d4ac7973a0e5ceb9f794786c0d706f343a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/cascade-rcnn_hrnetv2p-w18-20e_coco.py
@@ -0,0 +1,11 @@
+_base_ = './cascade-rcnn_hrnetv2p-w32-20e_coco.py'
+# model settings
+model = dict(
+ backbone=dict(
+ extra=dict(
+ stage2=dict(num_channels=(18, 36)),
+ stage3=dict(num_channels=(18, 36, 72)),
+ stage4=dict(num_channels=(18, 36, 72, 144))),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://msra/hrnetv2_w18')),
+ neck=dict(type='HRFPN', in_channels=[18, 36, 72, 144], out_channels=256))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/cascade-rcnn_hrnetv2p-w32-20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/cascade-rcnn_hrnetv2p-w32-20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..afeb75dbe13c5a8425924e280b250208aaec872f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/cascade-rcnn_hrnetv2p-w32-20e_coco.py
@@ -0,0 +1,51 @@
+_base_ = '../cascade_rcnn/cascade-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='HRNet',
+ extra=dict(
+ stage1=dict(
+ num_modules=1,
+ num_branches=1,
+ block='BOTTLENECK',
+ num_blocks=(4, ),
+ num_channels=(64, )),
+ stage2=dict(
+ num_modules=1,
+ num_branches=2,
+ block='BASIC',
+ num_blocks=(4, 4),
+ num_channels=(32, 64)),
+ stage3=dict(
+ num_modules=4,
+ num_branches=3,
+ block='BASIC',
+ num_blocks=(4, 4, 4),
+ num_channels=(32, 64, 128)),
+ stage4=dict(
+ num_modules=3,
+ num_branches=4,
+ block='BASIC',
+ num_blocks=(4, 4, 4, 4),
+ num_channels=(32, 64, 128, 256))),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://msra/hrnetv2_w32')),
+ neck=dict(
+ _delete_=True,
+ type='HRFPN',
+ in_channels=[32, 64, 128, 256],
+ out_channels=256))
+# learning policy
+max_epochs = 20
+train_cfg = dict(max_epochs=max_epochs)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 19],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/cascade-rcnn_hrnetv2p-w40-20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/cascade-rcnn_hrnetv2p-w40-20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..66f8882a0030ae82f7a74f67963bbd1da3422a48
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/cascade-rcnn_hrnetv2p-w40-20e_coco.py
@@ -0,0 +1,12 @@
+_base_ = './cascade-rcnn_hrnetv2p-w32-20e_coco.py'
+# model settings
+model = dict(
+ backbone=dict(
+ type='HRNet',
+ extra=dict(
+ stage2=dict(num_channels=(40, 80)),
+ stage3=dict(num_channels=(40, 80, 160)),
+ stage4=dict(num_channels=(40, 80, 160, 320))),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://msra/hrnetv2_w40')),
+ neck=dict(type='HRFPN', in_channels=[40, 80, 160, 320], out_channels=256))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/faster-rcnn_hrnetv2p-w18-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/faster-rcnn_hrnetv2p-w18-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ee9a698699a6674c90011b4037843560459462db
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/faster-rcnn_hrnetv2p-w18-1x_coco.py
@@ -0,0 +1,11 @@
+_base_ = './faster-rcnn_hrnetv2p-w32-1x_coco.py'
+# model settings
+model = dict(
+ backbone=dict(
+ extra=dict(
+ stage2=dict(num_channels=(18, 36)),
+ stage3=dict(num_channels=(18, 36, 72)),
+ stage4=dict(num_channels=(18, 36, 72, 144))),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://msra/hrnetv2_w18')),
+ neck=dict(type='HRFPN', in_channels=[18, 36, 72, 144], out_channels=256))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/faster-rcnn_hrnetv2p-w18-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/faster-rcnn_hrnetv2p-w18-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..0b72c68f8cbbc83d16313c6d3ab3faf0ac86926f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/faster-rcnn_hrnetv2p-w18-2x_coco.py
@@ -0,0 +1,16 @@
+_base_ = './faster-rcnn_hrnetv2p-w18-1x_coco.py'
+
+# learning policy
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/faster-rcnn_hrnetv2p-w32-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/faster-rcnn_hrnetv2p-w32-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a27ad06c5c169c84c6368f767b79b0a817d99fa1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/faster-rcnn_hrnetv2p-w32-1x_coco.py
@@ -0,0 +1,37 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='HRNet',
+ extra=dict(
+ stage1=dict(
+ num_modules=1,
+ num_branches=1,
+ block='BOTTLENECK',
+ num_blocks=(4, ),
+ num_channels=(64, )),
+ stage2=dict(
+ num_modules=1,
+ num_branches=2,
+ block='BASIC',
+ num_blocks=(4, 4),
+ num_channels=(32, 64)),
+ stage3=dict(
+ num_modules=4,
+ num_branches=3,
+ block='BASIC',
+ num_blocks=(4, 4, 4),
+ num_channels=(32, 64, 128)),
+ stage4=dict(
+ num_modules=3,
+ num_branches=4,
+ block='BASIC',
+ num_blocks=(4, 4, 4, 4),
+ num_channels=(32, 64, 128, 256))),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://msra/hrnetv2_w32')),
+ neck=dict(
+ _delete_=True,
+ type='HRFPN',
+ in_channels=[32, 64, 128, 256],
+ out_channels=256))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/faster-rcnn_hrnetv2p-w32_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/faster-rcnn_hrnetv2p-w32_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c9568ce65c142f86ec6181236464454106d7de99
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/faster-rcnn_hrnetv2p-w32_2x_coco.py
@@ -0,0 +1,16 @@
+_base_ = './faster-rcnn_hrnetv2p-w32-1x_coco.py'
+
+# learning policy
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/faster-rcnn_hrnetv2p-w40-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/faster-rcnn_hrnetv2p-w40-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b36200230b76269a9644cc7852cec6ce62eac5c3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/faster-rcnn_hrnetv2p-w40-1x_coco.py
@@ -0,0 +1,11 @@
+_base_ = './faster-rcnn_hrnetv2p-w32-1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='HRNet',
+ extra=dict(
+ stage2=dict(num_channels=(40, 80)),
+ stage3=dict(num_channels=(40, 80, 160)),
+ stage4=dict(num_channels=(40, 80, 160, 320))),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://msra/hrnetv2_w40')),
+ neck=dict(type='HRFPN', in_channels=[40, 80, 160, 320], out_channels=256))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/faster-rcnn_hrnetv2p-w40_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/faster-rcnn_hrnetv2p-w40_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d1b45355db1de7c649136438b91fec5199e08141
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/faster-rcnn_hrnetv2p-w40_2x_coco.py
@@ -0,0 +1,16 @@
+_base_ = './faster-rcnn_hrnetv2p-w40-1x_coco.py'
+
+# learning policy
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w18-gn-head_4xb4-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w18-gn-head_4xb4-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c20ca7767364e14e552b5b8af68a8124f6a1253e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w18-gn-head_4xb4-1x_coco.py
@@ -0,0 +1,10 @@
+_base_ = './fcos_hrnetv2p-w32-gn-head_4xb4-1x_coco.py'
+model = dict(
+ backbone=dict(
+ extra=dict(
+ stage2=dict(num_channels=(18, 36)),
+ stage3=dict(num_channels=(18, 36, 72)),
+ stage4=dict(num_channels=(18, 36, 72, 144))),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://msra/hrnetv2_w18')),
+ neck=dict(type='HRFPN', in_channels=[18, 36, 72, 144], out_channels=256))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w18-gn-head_4xb4-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w18-gn-head_4xb4-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f5b67f6a12e294455829dddb89d05e281f2d7dc0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w18-gn-head_4xb4-2x_coco.py
@@ -0,0 +1,16 @@
+_base_ = './fcos_hrnetv2p-w18-gn-head_4xb4-1x_coco.py'
+
+# learning policy
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w18-gn-head_ms-640-800-4xb4-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w18-gn-head_ms-640-800-4xb4-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c5332d65d129255117f459f45369d5e13ed6653c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w18-gn-head_ms-640-800-4xb4-2x_coco.py
@@ -0,0 +1,10 @@
+_base_ = './fcos_hrnetv2p-w32-gn-head_ms-640-800-4xb4-2x_coco.py'
+model = dict(
+ backbone=dict(
+ extra=dict(
+ stage2=dict(num_channels=(18, 36)),
+ stage3=dict(num_channels=(18, 36, 72)),
+ stage4=dict(num_channels=(18, 36, 72, 144))),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://msra/hrnetv2_w18')),
+ neck=dict(type='HRFPN', in_channels=[18, 36, 72, 144], out_channels=256))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w32-gn-head_4xb4-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w32-gn-head_4xb4-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..159d96d712ae047efd7988bc53ae65006291478f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w32-gn-head_4xb4-1x_coco.py
@@ -0,0 +1,43 @@
+_base_ = '../fcos/fcos_r50-caffe_fpn_gn-head_4xb4-1x_coco.py'
+model = dict(
+ data_preprocessor=dict(
+ mean=[103.53, 116.28, 123.675],
+ std=[57.375, 57.12, 58.395],
+ bgr_to_rgb=False),
+ backbone=dict(
+ _delete_=True,
+ type='HRNet',
+ extra=dict(
+ stage1=dict(
+ num_modules=1,
+ num_branches=1,
+ block='BOTTLENECK',
+ num_blocks=(4, ),
+ num_channels=(64, )),
+ stage2=dict(
+ num_modules=1,
+ num_branches=2,
+ block='BASIC',
+ num_blocks=(4, 4),
+ num_channels=(32, 64)),
+ stage3=dict(
+ num_modules=4,
+ num_branches=3,
+ block='BASIC',
+ num_blocks=(4, 4, 4),
+ num_channels=(32, 64, 128)),
+ stage4=dict(
+ num_modules=3,
+ num_branches=4,
+ block='BASIC',
+ num_blocks=(4, 4, 4, 4),
+ num_channels=(32, 64, 128, 256))),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://msra/hrnetv2_w32')),
+ neck=dict(
+ _delete_=True,
+ type='HRFPN',
+ in_channels=[32, 64, 128, 256],
+ out_channels=256,
+ stride=2,
+ num_outs=5))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w32-gn-head_4xb4-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w32-gn-head_4xb4-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..73fd80e979d88840a57c68ca2fad6cb2e82a26bd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w32-gn-head_4xb4-2x_coco.py
@@ -0,0 +1,16 @@
+_base_ = './fcos_hrnetv2p-w32-gn-head_4xb4-1x_coco.py'
+
+# learning policy
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w32-gn-head_ms-640-800-4xb4-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w32-gn-head_ms-640-800-4xb4-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..4c977bf31ed2fb0ef062108cea97c1cd235b89d3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w32-gn-head_ms-640-800-4xb4-2x_coco.py
@@ -0,0 +1,35 @@
+_base_ = './fcos_hrnetv2p-w32-gn-head_4xb4-1x_coco.py'
+
+model = dict(
+ data_preprocessor=dict(
+ mean=[103.53, 116.28, 123.675],
+ std=[57.375, 57.12, 58.395],
+ bgr_to_rgb=False))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+# learning policy
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w40-gn-head_ms-640-800-4xb4-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w40-gn-head_ms-640-800-4xb4-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..bb0ff6d6ce80e702f6e88b556a770345a23afca4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/fcos_hrnetv2p-w40-gn-head_ms-640-800-4xb4-2x_coco.py
@@ -0,0 +1,11 @@
+_base_ = './fcos_hrnetv2p-w32-gn-head_ms-640-800-4xb4-2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='HRNet',
+ extra=dict(
+ stage2=dict(num_channels=(40, 80)),
+ stage3=dict(num_channels=(40, 80, 160)),
+ stage4=dict(num_channels=(40, 80, 160, 320))),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://msra/hrnetv2_w40')),
+ neck=dict(type='HRFPN', in_channels=[40, 80, 160, 320], out_channels=256))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/htc_hrnetv2p-w18_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/htc_hrnetv2p-w18_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..55255d52a3541c99660dcddfba96da27c99f841d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/htc_hrnetv2p-w18_20e_coco.py
@@ -0,0 +1,10 @@
+_base_ = './htc_hrnetv2p-w32_20e_coco.py'
+model = dict(
+ backbone=dict(
+ extra=dict(
+ stage2=dict(num_channels=(18, 36)),
+ stage3=dict(num_channels=(18, 36, 72)),
+ stage4=dict(num_channels=(18, 36, 72, 144))),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://msra/hrnetv2_w18')),
+ neck=dict(type='HRFPN', in_channels=[18, 36, 72, 144], out_channels=256))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/htc_hrnetv2p-w32_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/htc_hrnetv2p-w32_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..545cb83eaca50f9d5de1fa6b3f3e569faab7d5f2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/htc_hrnetv2p-w32_20e_coco.py
@@ -0,0 +1,37 @@
+_base_ = '../htc/htc_r50_fpn_20e_coco.py'
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='HRNet',
+ extra=dict(
+ stage1=dict(
+ num_modules=1,
+ num_branches=1,
+ block='BOTTLENECK',
+ num_blocks=(4, ),
+ num_channels=(64, )),
+ stage2=dict(
+ num_modules=1,
+ num_branches=2,
+ block='BASIC',
+ num_blocks=(4, 4),
+ num_channels=(32, 64)),
+ stage3=dict(
+ num_modules=4,
+ num_branches=3,
+ block='BASIC',
+ num_blocks=(4, 4, 4),
+ num_channels=(32, 64, 128)),
+ stage4=dict(
+ num_modules=3,
+ num_branches=4,
+ block='BASIC',
+ num_blocks=(4, 4, 4, 4),
+ num_channels=(32, 64, 128, 256))),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://msra/hrnetv2_w32')),
+ neck=dict(
+ _delete_=True,
+ type='HRFPN',
+ in_channels=[32, 64, 128, 256],
+ out_channels=256))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/htc_hrnetv2p-w40_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/htc_hrnetv2p-w40_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b09256a08ee16893bcc0dd6518714daece294e0d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/htc_hrnetv2p-w40_20e_coco.py
@@ -0,0 +1,11 @@
+_base_ = './htc_hrnetv2p-w32_20e_coco.py'
+model = dict(
+ backbone=dict(
+ type='HRNet',
+ extra=dict(
+ stage2=dict(num_channels=(40, 80)),
+ stage3=dict(num_channels=(40, 80, 160)),
+ stage4=dict(num_channels=(40, 80, 160, 320))),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://msra/hrnetv2_w40')),
+ neck=dict(type='HRFPN', in_channels=[40, 80, 160, 320], out_channels=256))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/htc_hrnetv2p-w40_28e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/htc_hrnetv2p-w40_28e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1c13b58a1a0690d19239fef40915489ddaff408e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/htc_hrnetv2p-w40_28e_coco.py
@@ -0,0 +1,16 @@
+_base_ = './htc_hrnetv2p-w40_20e_coco.py'
+
+# learning policy
+max_epochs = 28
+train_cfg = dict(max_epochs=max_epochs)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[24, 27],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/htc_x101-64x4d_fpn_16xb1-28e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/htc_x101-64x4d_fpn_16xb1-28e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1f1304e5f963351667c28cb264ca5434bc81f744
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/htc_x101-64x4d_fpn_16xb1-28e_coco.py
@@ -0,0 +1,16 @@
+_base_ = '../htc/htc_x101-64x4d_fpn_16xb1-20e_coco.py'
+
+# learning policy
+max_epochs = 28
+train_cfg = dict(max_epochs=max_epochs)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[24, 27],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/mask-rcnn_hrnetv2p-w18-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/mask-rcnn_hrnetv2p-w18-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5d5a463a66bed51d73a42eafffea654a18c111ce
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/mask-rcnn_hrnetv2p-w18-1x_coco.py
@@ -0,0 +1,10 @@
+_base_ = './mask-rcnn_hrnetv2p-w32-1x_coco.py'
+model = dict(
+ backbone=dict(
+ extra=dict(
+ stage2=dict(num_channels=(18, 36)),
+ stage3=dict(num_channels=(18, 36, 72)),
+ stage4=dict(num_channels=(18, 36, 72, 144))),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://msra/hrnetv2_w18')),
+ neck=dict(type='HRFPN', in_channels=[18, 36, 72, 144], out_channels=256))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/mask-rcnn_hrnetv2p-w18-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/mask-rcnn_hrnetv2p-w18-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8abc55924a3eb8e06f9e1e5eeed503890542f6f6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/mask-rcnn_hrnetv2p-w18-2x_coco.py
@@ -0,0 +1,16 @@
+_base_ = './mask-rcnn_hrnetv2p-w18-1x_coco.py'
+
+# learning policy
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/mask-rcnn_hrnetv2p-w32-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/mask-rcnn_hrnetv2p-w32-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..208b037807dfa9cab1d33ac58ac785ff72e400c1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/mask-rcnn_hrnetv2p-w32-1x_coco.py
@@ -0,0 +1,37 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='HRNet',
+ extra=dict(
+ stage1=dict(
+ num_modules=1,
+ num_branches=1,
+ block='BOTTLENECK',
+ num_blocks=(4, ),
+ num_channels=(64, )),
+ stage2=dict(
+ num_modules=1,
+ num_branches=2,
+ block='BASIC',
+ num_blocks=(4, 4),
+ num_channels=(32, 64)),
+ stage3=dict(
+ num_modules=4,
+ num_branches=3,
+ block='BASIC',
+ num_blocks=(4, 4, 4),
+ num_channels=(32, 64, 128)),
+ stage4=dict(
+ num_modules=3,
+ num_branches=4,
+ block='BASIC',
+ num_blocks=(4, 4, 4, 4),
+ num_channels=(32, 64, 128, 256))),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://msra/hrnetv2_w32')),
+ neck=dict(
+ _delete_=True,
+ type='HRFPN',
+ in_channels=[32, 64, 128, 256],
+ out_channels=256))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/mask-rcnn_hrnetv2p-w32-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/mask-rcnn_hrnetv2p-w32-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d3741c820a6a0ca622ce6bbf80cb3e922107efb6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/mask-rcnn_hrnetv2p-w32-2x_coco.py
@@ -0,0 +1,16 @@
+_base_ = './mask-rcnn_hrnetv2p-w32-1x_coco.py'
+
+# learning policy
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/mask-rcnn_hrnetv2p-w40-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/mask-rcnn_hrnetv2p-w40-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..360420c56d42814ed6f4d84775f1a19dfa96574a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/mask-rcnn_hrnetv2p-w40-2x_coco.py
@@ -0,0 +1,16 @@
+_base_ = './mask-rcnn_hrnetv2p-w40_1x_coco.py'
+
+# learning policy
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/mask-rcnn_hrnetv2p-w40_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/mask-rcnn_hrnetv2p-w40_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..36e2305a520fd8305f9fd1358f5cbcb01027e40d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/mask-rcnn_hrnetv2p-w40_1x_coco.py
@@ -0,0 +1,11 @@
+_base_ = './mask-rcnn_hrnetv2p-w18-1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='HRNet',
+ extra=dict(
+ stage2=dict(num_channels=(40, 80)),
+ stage3=dict(num_channels=(40, 80, 160)),
+ stage4=dict(num_channels=(40, 80, 160, 320))),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://msra/hrnetv2_w40')),
+ neck=dict(type='HRFPN', in_channels=[40, 80, 160, 320], out_channels=256))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..54c624793291dc9a713c9a6fa6df50499136768c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/hrnet/metafile.yml
@@ -0,0 +1,971 @@
+Models:
+ - Name: faster-rcnn_hrnetv2p-w18-1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/hrnet/faster-rcnn_hrnetv2p-w18-1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.6
+ inference time (ms/im):
+ - value: 74.63
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 36.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/faster_rcnn_hrnetv2p_w18_1x_coco/faster_rcnn_hrnetv2p_w18_1x_coco_20200130-56651a6d.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: faster-rcnn_hrnetv2p-w18-2x_coco
+ In Collection: Faster R-CNN
+ Config: configs/hrnet/faster-rcnn_hrnetv2p-w18-2x_coco.py
+ Metadata:
+ Training Memory (GB): 6.6
+ inference time (ms/im):
+ - value: 74.63
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/faster_rcnn_hrnetv2p_w18_2x_coco/faster_rcnn_hrnetv2p_w18_2x_coco_20200702_085731-a4ec0611.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: faster-rcnn_hrnetv2p-w32-1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/hrnet/faster-rcnn_hrnetv2p-w32-1x_coco.py
+ Metadata:
+ Training Memory (GB): 9.0
+ inference time (ms/im):
+ - value: 80.65
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/faster_rcnn_hrnetv2p_w32_1x_coco/faster_rcnn_hrnetv2p_w32_1x_coco_20200130-6e286425.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: faster-rcnn_hrnetv2p-w32_2x_coco
+ In Collection: Faster R-CNN
+ Config: configs/hrnet/faster-rcnn_hrnetv2p-w32_2x_coco.py
+ Metadata:
+ Training Memory (GB): 9.0
+ inference time (ms/im):
+ - value: 80.65
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/faster_rcnn_hrnetv2p_w32_2x_coco/faster_rcnn_hrnetv2p_w32_2x_coco_20200529_015927-976a9c15.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: faster-rcnn_hrnetv2p-w40-1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/hrnet/faster-rcnn_hrnetv2p-w40-1x_coco.py
+ Metadata:
+ Training Memory (GB): 10.4
+ inference time (ms/im):
+ - value: 95.24
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/faster_rcnn_hrnetv2p_w40_1x_coco/faster_rcnn_hrnetv2p_w40_1x_coco_20200210-95c1f5ce.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: faster-rcnn_hrnetv2p-w40_2x_coco
+ In Collection: Faster R-CNN
+ Config: configs/hrnet/faster-rcnn_hrnetv2p-w40_2x_coco.py
+ Metadata:
+ Training Memory (GB): 10.4
+ inference time (ms/im):
+ - value: 95.24
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/faster_rcnn_hrnetv2p_w40_2x_coco/faster_rcnn_hrnetv2p_w40_2x_coco_20200512_161033-0f236ef4.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: mask-rcnn_hrnetv2p-w18-1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/hrnet/mask-rcnn_hrnetv2p-w18-1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.0
+ inference time (ms/im):
+ - value: 85.47
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.7
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 34.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/mask_rcnn_hrnetv2p_w18_1x_coco/mask_rcnn_hrnetv2p_w18_1x_coco_20200205-1c3d78ed.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: mask-rcnn_hrnetv2p-w18-2x_coco
+ In Collection: Mask R-CNN
+ Config: configs/hrnet/mask-rcnn_hrnetv2p-w18-2x_coco.py
+ Metadata:
+ Training Memory (GB): 7.0
+ inference time (ms/im):
+ - value: 85.47
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/mask_rcnn_hrnetv2p_w18_2x_coco/mask_rcnn_hrnetv2p_w18_2x_coco_20200212-b3c825b1.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: mask-rcnn_hrnetv2p-w32-1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/hrnet/mask-rcnn_hrnetv2p-w32-1x_coco.py
+ Metadata:
+ Training Memory (GB): 9.4
+ inference time (ms/im):
+ - value: 88.5
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/mask_rcnn_hrnetv2p_w32_1x_coco/mask_rcnn_hrnetv2p_w32_1x_coco_20200207-b29f616e.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: mask-rcnn_hrnetv2p-w32-2x_coco
+ In Collection: Mask R-CNN
+ Config: configs/hrnet/mask-rcnn_hrnetv2p-w32-2x_coco.py
+ Metadata:
+ Training Memory (GB): 9.4
+ inference time (ms/im):
+ - value: 88.5
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/mask_rcnn_hrnetv2p_w32_2x_coco/mask_rcnn_hrnetv2p_w32_2x_coco_20200213-45b75b4d.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: mask-rcnn_hrnetv2p-w40_1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/hrnet/mask-rcnn_hrnetv2p-w40_1x_coco.py
+ Metadata:
+ Training Memory (GB): 10.9
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.1
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/mask_rcnn_hrnetv2p_w40_1x_coco/mask_rcnn_hrnetv2p_w40_1x_coco_20200511_015646-66738b35.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: mask-rcnn_hrnetv2p-w40-2x_coco
+ In Collection: Mask R-CNN
+ Config: configs/hrnet/mask-rcnn_hrnetv2p-w40-2x_coco.py
+ Metadata:
+ Training Memory (GB): 10.9
+ Epochs: 24
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/mask_rcnn_hrnetv2p_w40_2x_coco/mask_rcnn_hrnetv2p_w40_2x_coco_20200512_163732-aed5e4ab.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: cascade-rcnn_hrnetv2p-w18-20e_coco
+ In Collection: Cascade R-CNN
+ Config: configs/hrnet/cascade-rcnn_hrnetv2p-w18-20e_coco.py
+ Metadata:
+ Training Memory (GB): 7.0
+ inference time (ms/im):
+ - value: 90.91
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 20
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/cascade_rcnn_hrnetv2p_w18_20e_coco/cascade_rcnn_hrnetv2p_w18_20e_coco_20200210-434be9d7.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: cascade-rcnn_hrnetv2p-w32-20e_coco
+ In Collection: Cascade R-CNN
+ Config: configs/hrnet/cascade-rcnn_hrnetv2p-w32-20e_coco.py
+ Metadata:
+ Training Memory (GB): 9.4
+ inference time (ms/im):
+ - value: 90.91
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 20
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/cascade_rcnn_hrnetv2p_w32_20e_coco/cascade_rcnn_hrnetv2p_w32_20e_coco_20200208-928455a4.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: cascade-rcnn_hrnetv2p-w40-20e_coco
+ In Collection: Cascade R-CNN
+ Config: configs/hrnet/cascade-rcnn_hrnetv2p-w40-20e_coco.py
+ Metadata:
+ Training Memory (GB): 10.8
+ Epochs: 20
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/cascade_rcnn_hrnetv2p_w40_20e_coco/cascade_rcnn_hrnetv2p_w40_20e_coco_20200512_161112-75e47b04.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: cascade-mask-rcnn_hrnetv2p-w18_20e_coco
+ In Collection: Cascade R-CNN
+ Config: configs/hrnet/cascade-mask-rcnn_hrnetv2p-w18_20e_coco.py
+ Metadata:
+ Training Memory (GB): 8.5
+ inference time (ms/im):
+ - value: 117.65
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 20
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.6
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/cascade_mask_rcnn_hrnetv2p_w18_20e_coco/cascade_mask_rcnn_hrnetv2p_w18_20e_coco_20200210-b543cd2b.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: cascade-mask-rcnn_hrnetv2p-w32_20e_coco
+ In Collection: Cascade R-CNN
+ Config: configs/hrnet/cascade-mask-rcnn_hrnetv2p-w32_20e_coco.py
+ Metadata:
+ inference time (ms/im):
+ - value: 120.48
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 20
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.3
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/cascade_mask_rcnn_hrnetv2p_w32_20e_coco/cascade_mask_rcnn_hrnetv2p_w32_20e_coco_20200512_154043-39d9cf7b.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: cascade-mask-rcnn_hrnetv2p-w40-20e_coco
+ In Collection: Cascade R-CNN
+ Config: configs/hrnet/cascade-mask-rcnn_hrnetv2p-w40-20e_coco.py
+ Metadata:
+ Training Memory (GB): 12.5
+ Epochs: 20
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.1
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/cascade_mask_rcnn_hrnetv2p_w40_20e_coco/cascade_mask_rcnn_hrnetv2p_w40_20e_coco_20200527_204922-969c4610.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: htc_hrnetv2p-w18_20e_coco
+ In Collection: HTC
+ Config: configs/hrnet/htc_hrnetv2p-w18_20e_coco.py
+ Metadata:
+ Training Memory (GB): 10.8
+ inference time (ms/im):
+ - value: 212.77
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 20
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/htc_hrnetv2p_w18_20e_coco/htc_hrnetv2p_w18_20e_coco_20200210-b266988c.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: htc_hrnetv2p-w32_20e_coco
+ In Collection: HTC
+ Config: configs/hrnet/htc_hrnetv2p-w32_20e_coco.py
+ Metadata:
+ Training Memory (GB): 13.1
+ inference time (ms/im):
+ - value: 204.08
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 20
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/htc_hrnetv2p_w32_20e_coco/htc_hrnetv2p_w32_20e_coco_20200207-7639fa12.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: htc_hrnetv2p-w40_20e_coco
+ In Collection: HTC
+ Config: configs/hrnet/htc_hrnetv2p-w40_20e_coco.py
+ Metadata:
+ Training Memory (GB): 14.6
+ Epochs: 20
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 40.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/htc_hrnetv2p_w40_20e_coco/htc_hrnetv2p_w40_20e_coco_20200529_183411-417c4d5b.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: fcos_hrnetv2p-w18-gn-head_4xb4-1x_coco
+ In Collection: FCOS
+ Config: configs/hrnet/fcos_hrnetv2p-w18-gn-head_4xb4-1x_coco.py
+ Metadata:
+ Training Resources: 4x V100 GPUs
+ Batch Size: 16
+ Training Memory (GB): 13.0
+ inference time (ms/im):
+ - value: 77.52
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 35.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w18_gn-head_4x4_1x_coco/fcos_hrnetv2p_w18_gn-head_4x4_1x_coco_20201212_100710-4ad151de.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: fcos_hrnetv2p-w18-gn-head_4xb4-2x_coco
+ In Collection: FCOS
+ Config: configs/hrnet/fcos_hrnetv2p-w18-gn-head_4xb4-2x_coco.py
+ Metadata:
+ Training Resources: 4x V100 GPUs
+ Batch Size: 16
+ Training Memory (GB): 13.0
+ inference time (ms/im):
+ - value: 77.52
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w18_gn-head_4x4_2x_coco/fcos_hrnetv2p_w18_gn-head_4x4_2x_coco_20201212_101110-5c575fa5.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: fcos_hrnetv2p-w32-gn-head_4xb4-1x_coco
+ In Collection: FCOS
+ Config: configs/hrnet/fcos_hrnetv2p-w32-gn-head_4xb4-1x_coco.py
+ Metadata:
+ Training Resources: 4x V100 GPUs
+ Batch Size: 16
+ Training Memory (GB): 17.5
+ inference time (ms/im):
+ - value: 77.52
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w32_gn-head_4x4_1x_coco/fcos_hrnetv2p_w32_gn-head_4x4_1x_coco_20201211_134730-cb8055c0.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: fcos_hrnetv2p-w32-gn-head_4xb4-2x_coco
+ In Collection: FCOS
+ Config: configs/hrnet/fcos_hrnetv2p-w32-gn-head_4xb4-2x_coco.py
+ Metadata:
+ Training Resources: 4x V100 GPUs
+ Batch Size: 16
+ Training Memory (GB): 17.5
+ inference time (ms/im):
+ - value: 77.52
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w32_gn-head_4x4_2x_coco/fcos_hrnetv2p_w32_gn-head_4x4_2x_coco_20201212_112133-77b6b9bb.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: fcos_hrnetv2p-w18-gn-head_ms-640-800-4xb4-2x_coco
+ In Collection: FCOS
+ Config: configs/hrnet/fcos_hrnetv2p-w18-gn-head_ms-640-800-4xb4-2x_coco.py
+ Metadata:
+ Training Resources: 4x V100 GPUs
+ Batch Size: 16
+ Training Memory (GB): 13.0
+ inference time (ms/im):
+ - value: 77.52
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w18_gn-head_mstrain_640-800_4x4_2x_coco/fcos_hrnetv2p_w18_gn-head_mstrain_640-800_4x4_2x_coco_20201212_111651-441e9d9f.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: fcos_hrnetv2p-w32-gn-head_ms-640-800-4xb4-2x_coco
+ In Collection: FCOS
+ Config: configs/hrnet/fcos_hrnetv2p-w32-gn-head_ms-640-800-4xb4-2x_coco.py
+ Metadata:
+ Training Resources: 4x V100 GPUs
+ Batch Size: 16
+ Training Memory (GB): 17.5
+ inference time (ms/im):
+ - value: 80.65
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w32_gn-head_mstrain_640-800_4x4_2x_coco/fcos_hrnetv2p_w32_gn-head_mstrain_640-800_4x4_2x_coco_20201212_090846-b6f2b49f.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
+
+ - Name: fcos_hrnetv2p-w40-gn-head_ms-640-800-4xb4-2x_coco
+ In Collection: FCOS
+ Config: configs/hrnet/fcos_hrnetv2p-w40-gn-head_ms-640-800-4xb4-2x_coco.py
+ Metadata:
+ Training Resources: 4x V100 GPUs
+ Batch Size: 16
+ Training Memory (GB): 20.3
+ inference time (ms/im):
+ - value: 92.59
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Architecture:
+ - HRNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/hrnet/fcos_hrnetv2p_w40_gn-head_mstrain_640-800_4x4_2x_coco/fcos_hrnetv2p_w40_gn-head_mstrain_640-800_4x4_2x_coco_20201212_124752-f22d2ce5.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.04514
+ Title: 'Deep High-Resolution Representation Learning for Visual Recognition'
+ README: configs/hrnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/backbones/hrnet.py#L195
+ Version: v2.0.0
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..a6b77ce4754f5f88e6effcd47dcdbbe4cd739757
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/README.md
@@ -0,0 +1,67 @@
+# HTC
+
+> [Hybrid Task Cascade for Instance Segmentation](https://arxiv.org/abs/1901.07518)
+
+
+
+## Abstract
+
+Cascade is a classic yet powerful architecture that has boosted performance on various tasks. However, how to introduce cascade to instance segmentation remains an open question. A simple combination of Cascade R-CNN and Mask R-CNN only brings limited gain. In exploring a more effective approach, we find that the key to a successful instance segmentation cascade is to fully leverage the reciprocal relationship between detection and segmentation. In this work, we propose a new framework, Hybrid Task Cascade (HTC), which differs in two important aspects: (1) instead of performing cascaded refinement on these two tasks separately, it interweaves them for a joint multi-stage processing; (2) it adopts a fully convolutional branch to provide spatial context, which can help distinguishing hard foreground from cluttered background. Overall, this framework can learn more discriminative features progressively while integrating complementary features together in each stage. Without bells and whistles, a single HTC obtains 38.4 and 1.5 improvement over a strong Cascade Mask R-CNN baseline on MSCOCO dataset. Moreover, our overall system achieves 48.6 mask AP on the test-challenge split, ranking 1st in the COCO 2018 Challenge Object Detection Task.
+
+
+

+
+
+## Introduction
+
+HTC requires COCO and [COCO-stuff](http://calvin.inf.ed.ac.uk/wp-content/uploads/data/cocostuffdataset/stuffthingmaps_trainval2017.zip) dataset for training. You need to download and extract it in the COCO dataset path.
+The directory should be like this.
+
+```none
+mmdetection
+├── mmdet
+├── tools
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ ├── train2017
+│ │ ├── val2017
+│ │ ├── test2017
+| | ├── stuffthingmaps
+```
+
+## Results and Models
+
+The results on COCO 2017val are shown in the below table. (results on test-dev are usually slightly higher than val)
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :-------------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :----------------------------------------------: | :-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | pytorch | 1x | 8.2 | 5.8 | 42.3 | 37.4 | [config](./htc_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/htc/htc_r50_fpn_1x_coco/htc_r50_fpn_1x_coco_20200317-7332cf16.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/htc/htc_r50_fpn_1x_coco/htc_r50_fpn_1x_coco_20200317_070435.log.json) |
+| R-50-FPN | pytorch | 20e | 8.2 | - | 43.3 | 38.3 | [config](./htc_r50_fpn_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/htc/htc_r50_fpn_20e_coco/htc_r50_fpn_20e_coco_20200319-fe28c577.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/htc/htc_r50_fpn_20e_coco/htc_r50_fpn_20e_coco_20200319_070313.log.json) |
+| R-101-FPN | pytorch | 20e | 10.2 | 5.5 | 44.8 | 39.6 | [config](./htc_r101_fpn_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/htc/htc_r101_fpn_20e_coco/htc_r101_fpn_20e_coco_20200317-9b41b48f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/htc/htc_r101_fpn_20e_coco/htc_r101_fpn_20e_coco_20200317_153107.log.json) |
+| X-101-32x4d-FPN | pytorch | 20e | 11.4 | 5.0 | 46.1 | 40.5 | [config](./htc_x101-32x4d_fpn_16xb1-20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/htc/htc_x101_32x4d_fpn_16x1_20e_coco/htc_x101_32x4d_fpn_16x1_20e_coco_20200318-de97ae01.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/htc/htc_x101_32x4d_fpn_16x1_20e_coco/htc_x101_32x4d_fpn_16x1_20e_coco_20200318_034519.log.json) |
+| X-101-64x4d-FPN | pytorch | 20e | 14.5 | 4.4 | 47.0 | 41.4 | [config](./htc_x101-64x4d_fpn_16xb1-20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/htc/htc_x101_64x4d_fpn_16x1_20e_coco/htc_x101_64x4d_fpn_16x1_20e_coco_20200318-b181fd7a.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/htc/htc_x101_64x4d_fpn_16x1_20e_coco/htc_x101_64x4d_fpn_16x1_20e_coco_20200318_081711.log.json) |
+
+- In the HTC paper and COCO 2018 Challenge, `score_thr` is set to 0.001 for both baselines and HTC.
+- We use 8 GPUs with 2 images/GPU for R-50 and R-101 models, and 16 GPUs with 1 image/GPU for X-101 models.
+ If you would like to train X-101 HTC with 8 GPUs, you need to change the lr from 0.02 to 0.01.
+
+We also provide a powerful HTC with DCN and multi-scale training model. No testing augmentation is used.
+
+| Backbone | Style | DCN | training scales | Lr schd | box AP | mask AP | Config | Download |
+| :-------------: | :-----: | :---: | :-------------: | :-----: | :----: | :-----: | :----------------------------------------------------------------------: | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| X-101-64x4d-FPN | pytorch | c3-c5 | 400~1400 | 20e | 50.4 | 43.8 | [config](./htc_x101-64x4d-dconv-c3-c5_fpn_ms-400-1400-16xb1-20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/htc/htc_x101_64x4d_fpn_dconv_c3-c5_mstrain_400_1400_16x1_20e_coco/htc_x101_64x4d_fpn_dconv_c3-c5_mstrain_400_1400_16x1_20e_coco_20200312-946fd751.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/htc/htc_x101_64x4d_fpn_dconv_c3-c5_mstrain_400_1400_16x1_20e_coco/htc_x101_64x4d_fpn_dconv_c3-c5_mstrain_400_1400_16x1_20e_coco_20200312_203410.log.json) |
+
+## Citation
+
+We provide config files to reproduce the results in the CVPR 2019 paper for [Hybrid Task Cascade](https://arxiv.org/abs/1901.07518).
+
+```latex
+@inproceedings{chen2019hybrid,
+ title={Hybrid task cascade for instance segmentation},
+ author={Chen, Kai and Pang, Jiangmiao and Wang, Jiaqi and Xiong, Yu and Li, Xiaoxiao and Sun, Shuyang and Feng, Wansen and Liu, Ziwei and Shi, Jianping and Ouyang, Wanli and Chen Change Loy and Dahua Lin},
+ booktitle={IEEE Conference on Computer Vision and Pattern Recognition},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc-without-semantic_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc-without-semantic_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..791f4eb25b53e122cd4876a71e84a4a9d2f67e26
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc-without-semantic_r50_fpn_1x_coco.py
@@ -0,0 +1,223 @@
+_base_ = [
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+# model settings
+model = dict(
+ type='HybridTaskCascade',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5),
+ rpn_head=dict(
+ type='RPNHead',
+ in_channels=256,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ scales=[8],
+ ratios=[0.5, 1.0, 2.0],
+ strides=[4, 8, 16, 32, 64]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0 / 9.0, loss_weight=1.0)),
+ roi_head=dict(
+ type='HybridTaskCascadeRoIHead',
+ interleaved=True,
+ mask_info_flow=True,
+ num_stages=3,
+ stage_loss_weights=[1, 0.5, 0.25],
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=7, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ bbox_head=[
+ dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0,
+ loss_weight=1.0)),
+ dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.05, 0.05, 0.1, 0.1]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0,
+ loss_weight=1.0)),
+ dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.033, 0.033, 0.067, 0.067]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0))
+ ],
+ mask_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=14, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ mask_head=[
+ dict(
+ type='HTCMaskHead',
+ with_conv_res=False,
+ num_convs=4,
+ in_channels=256,
+ conv_out_channels=256,
+ num_classes=80,
+ loss_mask=dict(
+ type='CrossEntropyLoss', use_mask=True, loss_weight=1.0)),
+ dict(
+ type='HTCMaskHead',
+ num_convs=4,
+ in_channels=256,
+ conv_out_channels=256,
+ num_classes=80,
+ loss_mask=dict(
+ type='CrossEntropyLoss', use_mask=True, loss_weight=1.0)),
+ dict(
+ type='HTCMaskHead',
+ num_convs=4,
+ in_channels=256,
+ conv_out_channels=256,
+ num_classes=80,
+ loss_mask=dict(
+ type='CrossEntropyLoss', use_mask=True, loss_weight=1.0))
+ ]),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=0,
+ pos_weight=-1,
+ debug=False),
+ rpn_proposal=dict(
+ nms_pre=2000,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=[
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ mask_size=28,
+ pos_weight=-1,
+ debug=False),
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.6,
+ neg_iou_thr=0.6,
+ min_pos_iou=0.6,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ mask_size=28,
+ pos_weight=-1,
+ debug=False),
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.7,
+ min_pos_iou=0.7,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ mask_size=28,
+ pos_weight=-1,
+ debug=False)
+ ]),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=1000,
+ max_per_img=1000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ score_thr=0.001,
+ nms=dict(type='nms', iou_threshold=0.5),
+ max_per_img=100,
+ mask_thr_binary=0.5)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc_r101_fpn_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc_r101_fpn_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..28091aad31029109c29941404f2c3cc47f9c1092
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc_r101_fpn_20e_coco.py
@@ -0,0 +1,6 @@
+_base_ = './htc_r50_fpn_20e_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3573f1f698095585f4a1de692d0e45a21429822e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc_r50_fpn_1x_coco.py
@@ -0,0 +1,33 @@
+_base_ = './htc-without-semantic_r50_fpn_1x_coco.py'
+model = dict(
+ data_preprocessor=dict(pad_seg=True),
+ roi_head=dict(
+ semantic_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=14, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[8]),
+ semantic_head=dict(
+ type='FusedSemanticHead',
+ num_ins=5,
+ fusion_level=1,
+ seg_scale_factor=1 / 8,
+ num_convs=4,
+ in_channels=256,
+ conv_out_channels=256,
+ num_classes=183,
+ loss_seg=dict(
+ type='CrossEntropyLoss', ignore_index=255, loss_weight=0.2))))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(
+ type='LoadAnnotations', with_bbox=True, with_mask=True, with_seg=True),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(
+ dataset=dict(
+ data_prefix=dict(img='train2017/', seg='stuffthingmaps/train2017/'),
+ pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc_r50_fpn_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc_r50_fpn_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..9f510fa6eec210381707f4d1b01264e72e0d0f76
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc_r50_fpn_20e_coco.py
@@ -0,0 +1,16 @@
+_base_ = './htc_r50_fpn_1x_coco.py'
+
+# learning policy
+max_epochs = 20
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 19],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc_x101-32x4d_fpn_16xb1-20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc_x101-32x4d_fpn_16xb1-20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..396d3a0e2b72acc1d9601706ec4629720a46a738
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc_x101-32x4d_fpn_16xb1-20e_coco.py
@@ -0,0 +1,32 @@
+_base_ = './htc_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
+
+train_dataloader = dict(batch_size=1, num_workers=1)
+
+# learning policy
+max_epochs = 20
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 19],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc_x101-64x4d-dconv-c3-c5_fpn_ms-400-1400-16xb1-20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc_x101-64x4d-dconv-c3-c5_fpn_ms-400-1400-16xb1-20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..26d68e7e2cda2a711e4d16899ae85b100afc60a0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc_x101-64x4d-dconv-c3-c5_fpn_ms-400-1400-16xb1-20e_coco.py
@@ -0,0 +1,20 @@
+_base_ = './htc_x101-64x4d_fpn_16xb1-20e_coco.py'
+
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCN', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)))
+
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(
+ type='LoadAnnotations', with_bbox=True, with_mask=True, with_seg=True),
+ dict(
+ type='RandomResize',
+ scale=[(1600, 400), (1600, 1400)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc_x101-64x4d_fpn_16xb1-20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc_x101-64x4d_fpn_16xb1-20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a600ddb0ebd2287cdaa0d00a6008db636d79be76
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/htc_x101-64x4d_fpn_16xb1-20e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './htc_x101-32x4d_fpn_16xb1-20e_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ groups=64,
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..2f0f74d2d06a0f6053fa7f0b9bb73024f8dcaac5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/htc/metafile.yml
@@ -0,0 +1,165 @@
+Collections:
+ - Name: HTC
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - FPN
+ - HTC
+ - RPN
+ - ResNet
+ - ResNeXt
+ - RoIAlign
+ Paper:
+ URL: https://arxiv.org/abs/1901.07518
+ Title: 'Hybrid Task Cascade for Instance Segmentation'
+ README: configs/htc/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/detectors/htc.py#L6
+ Version: v2.0.0
+
+Models:
+ - Name: htc_r50_fpn_1x_coco
+ In Collection: HTC
+ Config: configs/htc/htc_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 8.2
+ inference time (ms/im):
+ - value: 172.41
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.3
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/htc/htc_r50_fpn_1x_coco/htc_r50_fpn_1x_coco_20200317-7332cf16.pth
+
+ - Name: htc_r50_fpn_20e_coco
+ In Collection: HTC
+ Config: configs/htc/htc_r50_fpn_20e_coco.py
+ Metadata:
+ Training Memory (GB): 8.2
+ inference time (ms/im):
+ - value: 172.41
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 20
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.3
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/htc/htc_r50_fpn_20e_coco/htc_r50_fpn_20e_coco_20200319-fe28c577.pth
+
+ - Name: htc_r101_fpn_20e_coco
+ In Collection: HTC
+ Config: configs/htc/htc_r101_fpn_20e_coco.py
+ Metadata:
+ Training Memory (GB): 10.2
+ inference time (ms/im):
+ - value: 181.82
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 20
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/htc/htc_r101_fpn_20e_coco/htc_r101_fpn_20e_coco_20200317-9b41b48f.pth
+
+ - Name: htc_x101-32x4d_fpn_16xb1-20e_coco
+ In Collection: HTC
+ Config: configs/htc/htc_x101-32x4d_fpn_16xb1-20e_coco.py
+ Metadata:
+ Training Resources: 16x V100 GPUs
+ Batch Size: 16
+ Training Memory (GB): 11.4
+ inference time (ms/im):
+ - value: 200
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 20
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.1
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 40.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/htc/htc_x101_32x4d_fpn_16x1_20e_coco/htc_x101_32x4d_fpn_16x1_20e_coco_20200318-de97ae01.pth
+
+ - Name: htc_x101-64x4d_fpn_16xb1-20e_coco
+ In Collection: HTC
+ Config: configs/htc/htc_x101-64x4d_fpn_16xb1-20e_coco.py
+ Metadata:
+ Training Resources: 16x V100 GPUs
+ Batch Size: 16
+ Training Memory (GB): 14.5
+ inference time (ms/im):
+ - value: 227.27
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 20
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 47.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 41.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/htc/htc_x101_64x4d_fpn_16x1_20e_coco/htc_x101_64x4d_fpn_16x1_20e_coco_20200318-b181fd7a.pth
+
+ - Name: htc_x101-64x4d-dconv-c3-c5_fpn_ms-400-1400-16xb1-20e_coco
+ In Collection: HTC
+ Config: configs/htc/htc_x101-64x4d-dconv-c3-c5_fpn_ms-400-1400-16xb1-20e_coco.py
+ Metadata:
+ Training Resources: 16x V100 GPUs
+ Batch Size: 16
+ Epochs: 20
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 50.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 43.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/htc/htc_x101_64x4d_fpn_dconv_c3-c5_mstrain_400_1400_16x1_20e_coco/htc_x101_64x4d_fpn_dconv_c3-c5_mstrain_400_1400_16x1_20e_coco_20200312-946fd751.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..34132341833308e2d5d3dcb65bd5d8ba0b4e23bd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/README.md
@@ -0,0 +1,58 @@
+# Instaboost
+
+> [Instaboost: Boosting instance segmentation via probability map guided copy-pasting](https://arxiv.org/abs/1908.07801)
+
+
+
+## Abstract
+
+Instance segmentation requires a large number of training samples to achieve satisfactory performance and benefits from proper data augmentation. To enlarge the training set and increase the diversity, previous methods have investigated using data annotation from other domain (e.g. bbox, point) in a weakly supervised mechanism. In this paper, we present a simple, efficient and effective method to augment the training set using the existing instance mask annotations. Exploiting the pixel redundancy of the background, we are able to improve the performance of Mask R-CNN for 1.7 mAP on COCO dataset and 3.3 mAP on Pascal VOC dataset by simply introducing random jittering to objects. Furthermore, we propose a location probability map based approach to explore the feasible locations that objects can be placed based on local appearance similarity. With the guidance of such map, we boost the performance of R101-Mask R-CNN on instance segmentation from 35.7 mAP to 37.9 mAP without modifying the backbone or network structure. Our method is simple to implement and does not increase the computational complexity. It can be integrated into the training pipeline of any instance segmentation model without affecting the training and inference efficiency.
+
+
+

+
+
+## Introduction
+
+Configs in this directory is the implementation for ICCV2019 paper "InstaBoost: Boosting Instance Segmentation Via Probability Map Guided Copy-Pasting" and provided by the authors of the paper. InstaBoost is a data augmentation method for object detection and instance segmentation. The paper has been released on [`arXiv`](https://arxiv.org/abs/1908.07801).
+
+## Usage
+
+### Requirements
+
+You need to install `instaboostfast` before using it.
+
+```shell
+pip install instaboostfast
+```
+
+The code and more details can be found [here](https://github.com/GothicAi/Instaboost).
+
+### Integration with MMDetection
+
+InstaBoost have been already integrated in the data pipeline, thus all you need is to add or change **InstaBoost** configurations after **LoadImageFromFile**. We have provided examples like [this](mask_rcnn_r50_fpn_instaboost_4x#L121). You can refer to [`InstaBoostConfig`](https://github.com/GothicAi/InstaBoost-pypi#instaboostconfig) for more details.
+
+## Results and Models
+
+- All models were trained on `coco_2017_train` and tested on `coco_2017_val` for convenience of evaluation and comparison. In the paper, the results are obtained from `test-dev`.
+- To balance accuracy and training time when using InstaBoost, models released in this page are all trained for 48 Epochs. Other training and testing configs strictly follow the original framework.
+- For results and models in MMDetection V1.x, please refer to [Instaboost](https://github.com/GothicAi/Instaboost).
+
+| Network | Backbone | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :-----------: | :-------------: | :-----: | :------: | :------------: | :----: | :-----: | :---------------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Mask R-CNN | R-50-FPN | 4x | 4.4 | 17.5 | 40.6 | 36.6 | [config](./mask-rcnn_r50_fpn_instaboost-4x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/instaboost/mask_rcnn_r50_fpn_instaboost_4x_coco/mask_rcnn_r50_fpn_instaboost_4x_coco_20200307-d025f83a.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/instaboost/mask_rcnn_r50_fpn_instaboost_4x_coco/mask_rcnn_r50_fpn_instaboost_4x_coco_20200307_223635.log.json) |
+| Mask R-CNN | R-101-FPN | 4x | 6.4 | | 42.5 | 38.0 | [config](./mask-rcnn_r101_fpn_instaboost-4x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/instaboost/mask_rcnn_r101_fpn_instaboost_4x_coco/mask_rcnn_r101_fpn_instaboost_4x_coco_20200703_235738-f23f3a5f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/instaboost/mask_rcnn_r101_fpn_instaboost_4x_coco/mask_rcnn_r101_fpn_instaboost_4x_coco_20200703_235738.log.json) |
+| Mask R-CNN | X-101-64x4d-FPN | 4x | 10.7 | | 44.7 | 39.7 | [config](./mask-rcnn_x101-64x4d_fpn_instaboost-4x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/instaboost/mask_rcnn_x101_64x4d_fpn_instaboost_4x_coco/mask_rcnn_x101_64x4d_fpn_instaboost_4x_coco_20200515_080947-8ed58c1b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/instaboost/mask_rcnn_x101_64x4d_fpn_instaboost_4x_coco/mask_rcnn_x101_64x4d_fpn_instaboost_4x_coco_20200515_080947.log.json) |
+| Cascade R-CNN | R-101-FPN | 4x | 6.0 | 12.0 | 43.7 | 38.0 | [config](./cascade-mask-rcnn_r50_fpn_instaboost-4x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/instaboost/cascade_mask_rcnn_r50_fpn_instaboost_4x_coco/cascade_mask_rcnn_r50_fpn_instaboost_4x_coco_20200307-c19d98d9.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/instaboost/cascade_mask_rcnn_r50_fpn_instaboost_4x_coco/cascade_mask_rcnn_r50_fpn_instaboost_4x_coco_20200307_223646.log.json) |
+
+## Citation
+
+```latex
+@inproceedings{fang2019instaboost,
+ title={Instaboost: Boosting instance segmentation via probability map guided copy-pasting},
+ author={Fang, Hao-Shu and Sun, Jianhua and Wang, Runzhong and Gou, Minghao and Li, Yong-Lu and Lu, Cewu},
+ booktitle={Proceedings of the IEEE International Conference on Computer Vision},
+ pages={682--691},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/cascade-mask-rcnn_r101_fpn_instaboost-4x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/cascade-mask-rcnn_r101_fpn_instaboost-4x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..53e33b890cad86fcc64e6ea6eefe39138241c8e7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/cascade-mask-rcnn_r101_fpn_instaboost-4x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './cascade-mask-rcnn_r50_fpn_instaboost-4x_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/cascade-mask-rcnn_r50_fpn_instaboost-4x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/cascade-mask-rcnn_r50_fpn_instaboost-4x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f7736cf5756676944c543b7e8412997ac81c2745
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/cascade-mask-rcnn_r50_fpn_instaboost-4x_coco.py
@@ -0,0 +1,40 @@
+_base_ = '../cascade_rcnn/cascade-mask-rcnn_r50_fpn_1x_coco.py'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(
+ type='InstaBoost',
+ action_candidate=('normal', 'horizontal', 'skip'),
+ action_prob=(1, 0, 0),
+ scale=(0.8, 1.2),
+ dx=15,
+ dy=15,
+ theta=(-1, 1),
+ color_prob=0.5,
+ hflag=False,
+ aug_ratio=0.5),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+max_epochs = 48
+
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[32, 44],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs)
+
+# only keep latest 3 checkpoints
+default_hooks = dict(checkpoint=dict(max_keep_ckpts=3))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/cascade-mask-rcnn_x101-64x4d_fpn_instaboost-4x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/cascade-mask-rcnn_x101-64x4d_fpn_instaboost-4x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c7938d9e00e3a9c030b788ca83b1a6ddee208aed
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/cascade-mask-rcnn_x101-64x4d_fpn_instaboost-4x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './cascade-mask-rcnn_r50_fpn_instaboost-4x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/mask-rcnn_r101_fpn_instaboost-4x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/mask-rcnn_r101_fpn_instaboost-4x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..55bfa9fefa4db9d6d69fb3c4a285d04592168398
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/mask-rcnn_r101_fpn_instaboost-4x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './mask-rcnn_r50_fpn_instaboost-4x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/mask-rcnn_r50_fpn_instaboost-4x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/mask-rcnn_r50_fpn_instaboost-4x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..0a8c9be81f03f98f97975aca47922575555e3844
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/mask-rcnn_r50_fpn_instaboost-4x_coco.py
@@ -0,0 +1,40 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(
+ type='InstaBoost',
+ action_candidate=('normal', 'horizontal', 'skip'),
+ action_prob=(1, 0, 0),
+ scale=(0.8, 1.2),
+ dx=15,
+ dy=15,
+ theta=(-1, 1),
+ color_prob=0.5,
+ hflag=False,
+ aug_ratio=0.5),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+max_epochs = 48
+
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[32, 44],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs)
+
+# only keep latest 3 checkpoints
+default_hooks = dict(checkpoint=dict(max_keep_ckpts=3))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/mask-rcnn_x101-64x4d_fpn_instaboost-4x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/mask-rcnn_x101-64x4d_fpn_instaboost-4x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..9ba2ada6011dd77ea2dcac2133bef8d92e522381
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/mask-rcnn_x101-64x4d_fpn_instaboost-4x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './mask-rcnn_r50_fpn_instaboost-4x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..228f31b7301e6a5f9d2206e10be07bc7ea3b70be
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/instaboost/metafile.yml
@@ -0,0 +1,99 @@
+Collections:
+ - Name: InstaBoost
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - InstaBoost
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Paper:
+ URL: https://arxiv.org/abs/1908.07801
+ Title: 'Instaboost: Boosting instance segmentation via probability map guided copy-pasting'
+ README: configs/instaboost/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/datasets/pipelines/instaboost.py#L7
+ Version: v2.0.0
+
+Models:
+ - Name: mask-rcnn_r50_fpn_instaboost_4x_coco
+ In Collection: InstaBoost
+ Config: configs/instaboost/mask-rcnn_r50_fpn_instaboost-4x_coco.py
+ Metadata:
+ Training Memory (GB): 4.4
+ inference time (ms/im):
+ - value: 57.14
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 48
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.6
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/instaboost/mask_rcnn_r50_fpn_instaboost_4x_coco/mask_rcnn_r50_fpn_instaboost_4x_coco_20200307-d025f83a.pth
+
+ - Name: mask-rcnn_r101_fpn_instaboost-4x_coco
+ In Collection: InstaBoost
+ Config: configs/instaboost/mask-rcnn_r101_fpn_instaboost-4x_coco.py
+ Metadata:
+ Training Memory (GB): 6.4
+ Epochs: 48
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/instaboost/mask_rcnn_r101_fpn_instaboost_4x_coco/mask_rcnn_r101_fpn_instaboost_4x_coco_20200703_235738-f23f3a5f.pth
+
+ - Name: mask-rcnn_x101-64x4d_fpn_instaboost-4x_coco
+ In Collection: InstaBoost
+ Config: configs/instaboost/mask-rcnn_x101-64x4d_fpn_instaboost-4x_coco.py
+ Metadata:
+ Training Memory (GB): 10.7
+ Epochs: 48
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.7
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/instaboost/mask_rcnn_x101_64x4d_fpn_instaboost_4x_coco/mask_rcnn_x101_64x4d_fpn_instaboost_4x_coco_20200515_080947-8ed58c1b.pth
+
+ - Name: cascade-mask-rcnn_r50_fpn_instaboost_4x_coco
+ In Collection: InstaBoost
+ Config: configs/instaboost/cascade-mask-rcnn_r50_fpn_instaboost-4x_coco.py
+ Metadata:
+ Training Memory (GB): 6.0
+ inference time (ms/im):
+ - value: 83.33
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 48
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.7
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/instaboost/cascade_mask_rcnn_r50_fpn_instaboost_4x_coco/cascade_mask_rcnn_r50_fpn_instaboost_4x_coco_20200307-c19d98d9.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/lad/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lad/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..3c3b6b4bb4d9a86d87c7843dabb23b4e5d0abc66
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lad/README.md
@@ -0,0 +1,45 @@
+# LAD
+
+> [Improving Object Detection by Label Assignment Distillation](https://arxiv.org/abs/2108.10520)
+
+
+
+## Abstract
+
+Label assignment in object detection aims to assign targets, foreground or background, to sampled regions in an image. Unlike labeling for image classification, this problem is not well defined due to the object's bounding box. In this paper, we investigate the problem from a perspective of distillation, hence we call Label Assignment Distillation (LAD). Our initial motivation is very simple, we use a teacher network to generate labels for the student. This can be achieved in two ways: either using the teacher's prediction as the direct targets (soft label), or through the hard labels dynamically assigned by the teacher (LAD). Our experiments reveal that: (i) LAD is more effective than soft-label, but they are complementary. (ii) Using LAD, a smaller teacher can also improve a larger student significantly, while soft-label can't. We then introduce Co-learning LAD, in which two networks simultaneously learn from scratch and the role of teacher and student are dynamically interchanged. Using PAA-ResNet50 as a teacher, our LAD techniques can improve detectors PAA-ResNet101 and PAA-ResNeXt101 to 46AP and 47.5AP on the COCO test-dev set. With a stronger teacher PAA-SwinB, we improve the students PAA-ResNet50 to 43.7AP by only 1x schedule training and standard setting, and PAA-ResNet101 to 47.9AP, significantly surpassing the current methods.
+
+
+

+
+
+## Results and Models
+
+We provide config files to reproduce the object detection results in the
+WACV 2022 paper for Improving Object Detection by Label Assignment
+Distillation.
+
+### PAA with LAD
+
+| Teacher | Student | Training schedule | AP (val) | Config | Download |
+| :-----: | :-----: | :---------------: | :------: | :----------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| -- | R-50 | 1x | 40.4 | [config](../paa/paa_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r50_fpn_1x_coco/paa_r50_fpn_1x_coco_20200821-936edec3.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r50_fpn_1x_coco/paa_r50_fpn_1x_coco_20200821-936edec3.log.json) |
+| -- | R-101 | 1x | 42.6 | [config](../paa/paa_r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r101_fpn_1x_coco/paa_r101_fpn_1x_coco_20200821-0a1825a4.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r101_fpn_1x_coco/paa_r101_fpn_1x_coco_20200821-0a1825a4.log.json) |
+| R-101 | R-50 | 1x | 41.4 | [config](./lad_r50-paa-r101_fpn_2xb8_coco_1x.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/lad/lad_r50_paa_r101_fpn_coco_1x/lad_r50_paa_r101_fpn_coco_1x_20220708_124246-74c76ff0.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/lad/lad_r50_paa_r101_fpn_coco_1x/lad_r50_paa_r101_fpn_coco_1x_20220708_124246.log.json) |
+| R-50 | R-101 | 1x | 43.2 | [config](./lad_r101-paa-r50_fpn_2xb8_coco_1x.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/lad/lad_r101_paa_r50_fpn_coco_1x/lad_r101_paa_r50_fpn_coco_1x_20220708_124357-9407ac54.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/lad/lad_r101_paa_r50_fpn_coco_1x/lad_r101_paa_r50_fpn_coco_1x_20220708_124357.log.json) |
+
+## Note
+
+- Meaning of Config name: lad_r50(student model)\_paa(based on paa)\_r101(teacher model)\_fpn(neck)\_coco(dataset)\_1x(12 epoch).py
+- Results may fluctuate by about 0.2 mAP.
+- 2 GPUs are used, 8 samples per GPU.
+
+## Citation
+
+```latex
+@inproceedings{nguyen2021improving,
+ title={Improving Object Detection by Label Assignment Distillation},
+ author={Chuong H. Nguyen and Thuy C. Nguyen and Tuan N. Tang and Nam L. H. Phan},
+ booktitle = {WACV},
+ year={2022}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/lad/lad_r101-paa-r50_fpn_2xb8_coco_1x.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lad/lad_r101-paa-r50_fpn_2xb8_coco_1x.py
new file mode 100644
index 0000000000000000000000000000000000000000..d61d08638a073f3dad71d7499221e3ef62ff90f3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lad/lad_r101-paa-r50_fpn_2xb8_coco_1x.py
@@ -0,0 +1,127 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+teacher_ckpt = 'https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r50_fpn_1x_coco/paa_r50_fpn_1x_coco_20200821-936edec3.pth' # noqa
+
+model = dict(
+ type='LAD',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ # student
+ backbone=dict(
+ type='ResNet',
+ depth=101,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output',
+ num_outs=5),
+ bbox_head=dict(
+ type='LADHead',
+ reg_decoded_bbox=True,
+ score_voting=True,
+ topk=9,
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ octave_base_scale=8,
+ scales_per_octave=1,
+ strides=[8, 16, 32, 64, 128]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=1.3),
+ loss_centerness=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=0.5)),
+ # teacher
+ teacher_ckpt=teacher_ckpt,
+ teacher_backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch'),
+ teacher_neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output',
+ num_outs=5),
+ teacher_bbox_head=dict(
+ type='LADHead',
+ reg_decoded_bbox=True,
+ score_voting=True,
+ topk=9,
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ octave_base_scale=8,
+ scales_per_octave=1,
+ strides=[8, 16, 32, 64, 128]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=1.3),
+ loss_centerness=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=0.5)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.1,
+ neg_iou_thr=0.1,
+ min_pos_iou=0,
+ ignore_iof_thr=-1),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ score_voting=True,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+train_dataloader = dict(batch_size=8, num_workers=4)
+optim_wrapper = dict(type='AmpOptimWrapper', optimizer=dict(lr=0.01))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/lad/lad_r50-paa-r101_fpn_2xb8_coco_1x.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lad/lad_r50-paa-r101_fpn_2xb8_coco_1x.py
new file mode 100644
index 0000000000000000000000000000000000000000..f7eaf2bfba1c41b42836e94ffe2714978dffd20a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lad/lad_r50-paa-r101_fpn_2xb8_coco_1x.py
@@ -0,0 +1,126 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+teacher_ckpt = 'http://download.openmmlab.com/mmdetection/v2.0/paa/paa_r101_fpn_1x_coco/paa_r101_fpn_1x_coco_20200821-0a1825a4.pth' # noqa
+
+model = dict(
+ type='LAD',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ # student
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output',
+ num_outs=5),
+ bbox_head=dict(
+ type='LADHead',
+ reg_decoded_bbox=True,
+ score_voting=True,
+ topk=9,
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ octave_base_scale=8,
+ scales_per_octave=1,
+ strides=[8, 16, 32, 64, 128]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=1.3),
+ loss_centerness=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=0.5)),
+ # teacher
+ teacher_ckpt=teacher_ckpt,
+ teacher_backbone=dict(
+ type='ResNet',
+ depth=101,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch'),
+ teacher_neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output',
+ num_outs=5),
+ teacher_bbox_head=dict(
+ type='LADHead',
+ reg_decoded_bbox=True,
+ score_voting=True,
+ topk=9,
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ octave_base_scale=8,
+ scales_per_octave=1,
+ strides=[8, 16, 32, 64, 128]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=1.3),
+ loss_centerness=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=0.5)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.1,
+ neg_iou_thr=0.1,
+ min_pos_iou=0,
+ ignore_iof_thr=-1),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ score_voting=True,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+train_dataloader = dict(batch_size=8, num_workers=4)
+optim_wrapper = dict(type='AmpOptimWrapper', optimizer=dict(lr=0.01))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/lad/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lad/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..230132e63c06c77e16902450c282cf9a25150751
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lad/metafile.yml
@@ -0,0 +1,45 @@
+Collections:
+ - Name: Label Assignment Distillation
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - Label Assignment Distillation
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 2x V100 GPUs
+ Architecture:
+ - FPN
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/2108.10520
+ Title: 'Improving Object Detection by Label Assignment Distillation'
+ README: configs/lad/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.19.0/mmdet/models/detectors/lad.py#L10
+ Version: v2.19.0
+
+Models:
+ - Name: lad_r101-paa-r50_fpn_2xb8_coco_1x
+ In Collection: Label Assignment Distillation
+ Config: configs/lad/lad_r101-paa-r50_fpn_2xb8_coco_1x.py
+ Metadata:
+ Training Memory (GB): 12.4
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/lad/lad_r101_paa_r50_fpn_coco_1x/lad_r101_paa_r50_fpn_coco_1x_20220708_124357-9407ac54.pth
+ - Name: lad_r50-paa-r101_fpn_2xb8_coco_1x
+ In Collection: Label Assignment Distillation
+ Config: configs/lad/lad_r50-paa-r101_fpn_2xb8_coco_1x.py
+ Metadata:
+ Training Memory (GB): 8.9
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/lad/lad_r50_paa_r101_fpn_coco_1x/lad_r50_paa_r101_fpn_coco_1x_20220708_124246-74c76ff0.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ld/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ld/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..65e16c79d9ce4072f46c1473f0f208a533a3a300
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ld/README.md
@@ -0,0 +1,43 @@
+# LD
+
+> [Localization Distillation for Dense Object Detection](https://arxiv.org/abs/2102.12252)
+
+
+
+## Abstract
+
+Knowledge distillation (KD) has witnessed its powerful capability in learning compact models in object detection. Previous KD methods for object detection mostly focus on imitating deep features within the imitation regions instead of mimicking classification logits due to its inefficiency in distilling localization information. In this paper, by reformulating the knowledge distillation process on localization, we present a novel localization distillation (LD) method which can efficiently transfer the localization knowledge from the teacher to the student. Moreover, we also heuristically introduce the concept of valuable localization region that can aid to selectively distill the semantic and localization knowledge for a certain region. Combining these two new components, for the first time, we show that logit mimicking can outperform feature imitation and localization knowledge distillation is more important and efficient than semantic knowledge for distilling object detectors. Our distillation scheme is simple as well as effective and can be easily applied to different dense object detectors. Experiments show that our LD can boost the AP score of GFocal-ResNet-50 with a single-scale 1× training schedule from 40.1 to 42.1 on the COCO benchmark without any sacrifice on the inference speed.
+
+
+

+
+
+## Results and Models
+
+### GFocalV1 with LD
+
+| Teacher | Student | Training schedule | Mini-batch size | AP (val) | Config | Download |
+| :-------: | :-----: | :---------------: | :-------------: | :------: | :-----------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| -- | R-18 | 1x | 6 | 35.8 | | |
+| R-101 | R-18 | 1x | 6 | 36.5 | [config](./ld_r18-gflv1-r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/ld/ld_r18_gflv1_r101_fpn_coco_1x/ld_r18_gflv1_r101_fpn_coco_1x_20220702_062206-330e6332.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/ld/ld_r18_gflv1_r101_fpn_coco_1x/ld_r18_gflv1_r101_fpn_coco_1x_20220702_062206.log.json) |
+| -- | R-34 | 1x | 6 | 38.9 | | |
+| R-101 | R-34 | 1x | 6 | 39.9 | [config](./ld_r34-gflv1-r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/ld/ld_r34_gflv1_r101_fpn_coco_1x/ld_r34_gflv1_r101_fpn_coco_1x_20220630_134007-9bc69413.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/ld/ld_r34_gflv1_r101_fpn_coco_1x/ld_r34_gflv1_r101_fpn_coco_1x_20220630_134007.log.json) |
+| -- | R-50 | 1x | 6 | 40.1 | | |
+| R-101 | R-50 | 1x | 6 | 41.0 | [config](./ld_r50-gflv1-r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/ld/ld_r50_gflv1_r101_fpn_coco_1x/ld_r50_gflv1_r101_fpn_coco_1x_20220629_145355-8dc5bad8.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/ld/ld_r50_gflv1_r101_fpn_coco_1x/ld_r50_gflv1_r101_fpn_coco_1x_20220629_145355.log.json) |
+| -- | R-101 | 2x | 6 | 44.6 | | |
+| R-101-DCN | R-101 | 2x | 6 | 45.5 | [config](./ld_r101-gflv1-r101-dcn_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/ld/ld_r101_gflv1_r101dcn_fpn_coco_2x/ld_r101_gflv1_r101dcn_fpn_coco_2x_20220629_185920-9e658426.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/ld/ld_r101_gflv1_r101dcn_fpn_coco_2x/ld_r101_gflv1_r101dcn_fpn_coco_2x_20220629_185920.log.json) |
+
+## Note
+
+- Meaning of Config name: ld_r18(student model)\_gflv1(based on gflv1)\_r101(teacher model)\_fpn(neck)\_coco(dataset)\_1x(12 epoch).py
+
+## Citation
+
+```latex
+@Inproceedings{zheng2022LD,
+ title={Localization Distillation for Dense Object Detection},
+ author= {Zheng, Zhaohui and Ye, Rongguang and Wang, Ping and Ren, Dongwei and Zuo, Wangmeng and Hou, Qibin and Cheng, Mingming},
+ booktitle={CVPR},
+ year={2022}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ld/ld_r101-gflv1-r101-dcn_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ld/ld_r101-gflv1-r101-dcn_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a7e928bdc2325825d836bd939f163d71e972c238
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ld/ld_r101-gflv1-r101-dcn_fpn_2x_coco.py
@@ -0,0 +1,49 @@
+_base_ = ['./ld_r18-gflv1-r101_fpn_1x_coco.py']
+teacher_ckpt = 'https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_r101_fpn_dconv_c3-c5_mstrain_2x_coco/gfl_r101_fpn_dconv_c3-c5_mstrain_2x_coco_20200630_102002-134b07df.pth' # noqa
+model = dict(
+ teacher_config='configs/gfl/gfl_r101-dconv-c3-c5_fpn_ms-2x_coco.py',
+ teacher_ckpt=teacher_ckpt,
+ backbone=dict(
+ type='ResNet',
+ depth=101,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output',
+ num_outs=5))
+
+max_epochs = 24
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs)
+
+# multi-scale training
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize', scale=[(1333, 480), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ld/ld_r18-gflv1-r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ld/ld_r18-gflv1-r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f18bb1d3620f3caecdc870ea8a3346424729225c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ld/ld_r18-gflv1-r101_fpn_1x_coco.py
@@ -0,0 +1,70 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+teacher_ckpt = 'https://download.openmmlab.com/mmdetection/v2.0/gfl/gfl_r101_fpn_mstrain_2x_coco/gfl_r101_fpn_mstrain_2x_coco_20200629_200126-dd12f847.pth' # noqa
+model = dict(
+ type='KnowledgeDistillationSingleStageDetector',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ teacher_config='configs/gfl/gfl_r101_fpn_ms-2x_coco.py',
+ teacher_ckpt=teacher_ckpt,
+ backbone=dict(
+ type='ResNet',
+ depth=18,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet18')),
+ neck=dict(
+ type='FPN',
+ in_channels=[64, 128, 256, 512],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output',
+ num_outs=5),
+ bbox_head=dict(
+ type='LDHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ octave_base_scale=8,
+ scales_per_octave=1,
+ strides=[8, 16, 32, 64, 128]),
+ loss_cls=dict(
+ type='QualityFocalLoss',
+ use_sigmoid=True,
+ beta=2.0,
+ loss_weight=1.0),
+ loss_dfl=dict(type='DistributionFocalLoss', loss_weight=0.25),
+ loss_ld=dict(
+ type='KnowledgeDistillationKLDivLoss', loss_weight=0.25, T=10),
+ reg_max=16,
+ loss_bbox=dict(type='GIoULoss', loss_weight=2.0)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(type='ATSSAssigner', topk=9),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ld/ld_r34-gflv1-r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ld/ld_r34-gflv1-r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..2198adc82cfc98fca139e120ea0487989ac8bae7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ld/ld_r34-gflv1-r101_fpn_1x_coco.py
@@ -0,0 +1,19 @@
+_base_ = ['./ld_r18-gflv1-r101_fpn_1x_coco.py']
+model = dict(
+ backbone=dict(
+ type='ResNet',
+ depth=34,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet34')),
+ neck=dict(
+ type='FPN',
+ in_channels=[64, 128, 256, 512],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output',
+ num_outs=5))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ld/ld_r50-gflv1-r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ld/ld_r50-gflv1-r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..89ab5796969b88080f96f3afcc24183b0c11c730
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ld/ld_r50-gflv1-r101_fpn_1x_coco.py
@@ -0,0 +1,19 @@
+_base_ = ['./ld_r18-gflv1-r101_fpn_1x_coco.py']
+model = dict(
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output',
+ num_outs=5))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ld/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ld/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..a807d1b816e78734839cc1482c9c3d4afe59d6ac
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ld/metafile.yml
@@ -0,0 +1,69 @@
+Collections:
+ - Name: Localization Distillation
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - Localization Distillation
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - FPN
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/2102.12252
+ Title: 'Localization Distillation for Dense Object Detection'
+ README: configs/ld/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.11.0/mmdet/models/dense_heads/ld_head.py#L11
+ Version: v2.11.0
+
+Models:
+ - Name: ld_r18-gflv1-r101_fpn_1x_coco
+ In Collection: Localization Distillation
+ Config: configs/ld/ld_r18-gflv1-r101_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 1.8
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 36.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/ld/ld_r18_gflv1_r101_fpn_coco_1x/ld_r18_gflv1_r101_fpn_coco_1x_20220702_062206-330e6332.pth
+ - Name: ld_r34-gflv1-r101_fpn_1x_coco
+ In Collection: Localization Distillation
+ Config: configs/ld/ld_r34-gflv1-r101_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 2.2
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/ld/ld_r34_gflv1_r101_fpn_coco_1x/ld_r34_gflv1_r101_fpn_coco_1x_20220630_134007-9bc69413.pth
+ - Name: ld_r50-gflv1-r101_fpn_1x_coco
+ In Collection: Localization Distillation
+ Config: configs/ld/ld_r50-gflv1-r101_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.6
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/ld/ld_r50_gflv1_r101_fpn_coco_1x/ld_r50_gflv1_r101_fpn_coco_1x_20220629_145355-8dc5bad8.pth
+ - Name: ld_r101-gflv1-r101-dcn_fpn_2x_coco
+ In Collection: Localization Distillation
+ Config: configs/ld/ld_r101-gflv1-r101-dcn_fpn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 5.5
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/ld/ld_r101_gflv1_r101dcn_fpn_coco_2x/ld_r101_gflv1_r101dcn_fpn_coco_2x_20220629_185920-9e658426.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..443a0a71b46f4d2eda45571d6b7e108af6528d02
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/README.md
@@ -0,0 +1,54 @@
+# Legacy Configs in MMDetection V1.x
+
+
+
+Configs in this directory implement the legacy configs used by MMDetection V1.x and its model zoos.
+
+To help users convert their models from V1.x to MMDetection V2.0, we provide v1.x configs to inference the converted v1.x models.
+Due to the BC-breaking changes in MMDetection V2.0 from MMDetection V1.x, running inference with the same model weights in these two version will produce different results. The difference will cause within 1% AP absolute difference as can be found in the following table.
+
+## Usage
+
+To upgrade the model version, the users need to do the following steps.
+
+### 1. Convert model weights
+
+There are three main difference in the model weights between V1.x and V2.0 codebases.
+
+1. Since the class order in all the detector's classification branch is reordered, all the legacy model weights need to go through the conversion process.
+2. The regression and segmentation head no longer contain the background channel. Weights in these background channels should be removed to fix in the current codebase.
+3. For two-stage detectors, their wegihts need to be upgraded since MMDetection V2.0 refactors all the two-stage detectors with `RoIHead`.
+
+The users can do the same modification as mentioned above for the self-implemented
+detectors. We provide a scripts `tools/model_converters/upgrade_model_version.py` to convert the model weights in the V1.x model zoo.
+
+```bash
+python tools/model_converters/upgrade_model_version.py ${OLD_MODEL_PATH} ${NEW_MODEL_PATH} --num-classes ${NUM_CLASSES}
+
+```
+
+- OLD_MODEL_PATH: the path to load the model weights in 1.x version.
+- NEW_MODEL_PATH: the path to save the converted model weights in 2.0 version.
+- NUM_CLASSES: number of classes of the original model weights. Usually it is 81 for COCO dataset, 21 for VOC dataset.
+ The number of classes in V2.0 models should be equal to that in V1.x models - 1.
+
+### 2. Use configs with legacy settings
+
+After converting the model weights, checkout to the v1.2 release to find the corresponding config file that uses the legacy settings.
+The V1.x models usually need these three legacy modules: `LegacyAnchorGenerator`, `LegacyDeltaXYWHBBoxCoder`, and `RoIAlign(align=False)`.
+For models using ResNet Caffe backbones, they also need to change the pretrain name and the corresponding `img_norm_cfg`.
+An example is in [`retinanet_r50-caffe_fpn_1x_coco_v1.py`](retinanet_r50-caffe_fpn_1x_coco_v1.py)
+Then use the config to test the model weights. For most models, the obtained results should be close to that in V1.x.
+We provide configs of some common structures in this directory.
+
+## Performance
+
+The performance change after converting the models in this directory are listed as the following.
+
+| Method | Style | Lr schd | V1.x box AP | V1.x mask AP | V2.0 box AP | V2.0 mask AP | Config | Download |
+| :-------------------------: | :-----: | :-----: | :---------: | :----------: | :---------: | :----------: | :-------------------------------------------------: | :-------------------------------------------------------------------------------------------------------------------------------: |
+| Mask R-CNN R-50-FPN | pytorch | 1x | 37.3 | 34.2 | 36.8 | 33.9 | [config](./mask-rcnn_r50_fpn_1x_coco_v1.py) | [model](https://s3.ap-northeast-2.amazonaws.com/open-mmlab/mmdetection/models/mask_rcnn_r50_fpn_1x_20181010-069fa190.pth) |
+| RetinaNet R-50-FPN | caffe | 1x | 35.8 | - | 35.4 | - | [config](./retinanet_r50-caffe_fpn_1x_coco_v1.py) | |
+| RetinaNet R-50-FPN | pytorch | 1x | 35.6 | - | 35.2 | - | [config](./retinanet_r50_fpn_1x_coco_v1.py) | [model](https://s3.ap-northeast-2.amazonaws.com/open-mmlab/mmdetection/models/retinanet_r50_fpn_1x_20181125-7b0c2548.pth) |
+| Cascade Mask R-CNN R-50-FPN | pytorch | 1x | 41.2 | 35.7 | 40.8 | 35.6 | [config](./cascade-mask-rcnn_r50_fpn_1x_coco_v1.py) | [model](https://s3.ap-northeast-2.amazonaws.com/open-mmlab/mmdetection/models/cascade_mask_rcnn_r50_fpn_1x_20181123-88b170c9.pth) |
+| SSD300-VGG16 | caffe | 120e | 25.7 | - | 25.4 | - | [config](./ssd300_coco_v1.py) | [model](https://s3.ap-northeast-2.amazonaws.com/open-mmlab/mmdetection/models/ssd300_coco_vgg16_caffe_120e_20181221-84d7110b.pth) |
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/cascade-mask-rcnn_r50_fpn_1x_coco_v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/cascade-mask-rcnn_r50_fpn_1x_coco_v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..f948a7a9c10f618438e8ff54bdf3333335577e90
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/cascade-mask-rcnn_r50_fpn_1x_coco_v1.py
@@ -0,0 +1,78 @@
+_base_ = [
+ '../_base_/models/cascade-mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ type='CascadeRCNN',
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5),
+ rpn_head=dict(
+ anchor_generator=dict(type='LegacyAnchorGenerator', center_offset=0.5),
+ bbox_coder=dict(
+ type='LegacyDeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0])),
+ roi_head=dict(
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(
+ type='RoIAlign',
+ output_size=7,
+ sampling_ratio=2,
+ aligned=False)),
+ bbox_head=[
+ dict(
+ type='Shared2FCBBoxHead',
+ reg_class_agnostic=True,
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='LegacyDeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2])),
+ dict(
+ type='Shared2FCBBoxHead',
+ reg_class_agnostic=True,
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='LegacyDeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.05, 0.05, 0.1, 0.1])),
+ dict(
+ type='Shared2FCBBoxHead',
+ reg_class_agnostic=True,
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='LegacyDeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.033, 0.033, 0.067, 0.067])),
+ ],
+ mask_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(
+ type='RoIAlign',
+ output_size=14,
+ sampling_ratio=2,
+ aligned=False))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/faster-rcnn_r50_fpn_1x_coco_v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/faster-rcnn_r50_fpn_1x_coco_v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..66bf9713793c4a0a951273d037253f930fbb31a6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/faster-rcnn_r50_fpn_1x_coco_v1.py
@@ -0,0 +1,38 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ type='FasterRCNN',
+ backbone=dict(
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ rpn_head=dict(
+ type='RPNHead',
+ anchor_generator=dict(
+ type='LegacyAnchorGenerator',
+ center_offset=0.5,
+ scales=[8],
+ ratios=[0.5, 1.0, 2.0],
+ strides=[4, 8, 16, 32, 64]),
+ bbox_coder=dict(type='LegacyDeltaXYWHBBoxCoder'),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0 / 9.0, loss_weight=1.0)),
+ roi_head=dict(
+ type='StandardRoIHead',
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(
+ type='RoIAlign',
+ output_size=7,
+ sampling_ratio=2,
+ aligned=False),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ bbox_head=dict(
+ bbox_coder=dict(type='LegacyDeltaXYWHBBoxCoder'),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn_proposal=dict(max_per_img=2000),
+ rcnn=dict(assigner=dict(match_low_quality=True))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/mask-rcnn_r50_fpn_1x_coco_v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/mask-rcnn_r50_fpn_1x_coco_v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..690802598493e64821aaf98111161e36b169e475
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/mask-rcnn_r50_fpn_1x_coco_v1.py
@@ -0,0 +1,34 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ rpn_head=dict(
+ anchor_generator=dict(type='LegacyAnchorGenerator', center_offset=0.5),
+ bbox_coder=dict(type='LegacyDeltaXYWHBBoxCoder'),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0 / 9.0, loss_weight=1.0)),
+ roi_head=dict(
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(
+ type='RoIAlign',
+ output_size=7,
+ sampling_ratio=2,
+ aligned=False)),
+ mask_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(
+ type='RoIAlign',
+ output_size=14,
+ sampling_ratio=2,
+ aligned=False)),
+ bbox_head=dict(
+ bbox_coder=dict(type='LegacyDeltaXYWHBBoxCoder'),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0))),
+
+ # model training and testing settings
+ train_cfg=dict(
+ rpn_proposal=dict(max_per_img=2000),
+ rcnn=dict(assigner=dict(match_low_quality=True))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/retinanet_r50-caffe_fpn_1x_coco_v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/retinanet_r50-caffe_fpn_1x_coco_v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..49abc31a002f56147cacf1b7707140a14b784a99
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/retinanet_r50-caffe_fpn_1x_coco_v1.py
@@ -0,0 +1,16 @@
+_base_ = './retinanet_r50_fpn_1x_coco_v1.py'
+model = dict(
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ # use caffe img_norm
+ mean=[102.9801, 115.9465, 122.7717],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ norm_cfg=dict(requires_grad=False),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron/resnet50_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/retinanet_r50_fpn_1x_coco_v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/retinanet_r50_fpn_1x_coco_v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..6198b9717957374ce734ca74de5f54dda44123b9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/retinanet_r50_fpn_1x_coco_v1.py
@@ -0,0 +1,17 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ bbox_head=dict(
+ type='RetinaHead',
+ anchor_generator=dict(
+ type='LegacyAnchorGenerator',
+ center_offset=0.5,
+ octave_base_scale=4,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[8, 16, 32, 64, 128]),
+ bbox_coder=dict(type='LegacyDeltaXYWHBBoxCoder'),
+ loss_bbox=dict(type='SmoothL1Loss', beta=0.11, loss_weight=1.0)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/ssd300_coco_v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/ssd300_coco_v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..e5ffc633a9b4773d7116bed7cbf8bcab7fb3110d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/legacy_1.x/ssd300_coco_v1.py
@@ -0,0 +1,20 @@
+_base_ = [
+ '../_base_/models/ssd300.py', '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_2x.py', '../_base_/default_runtime.py'
+]
+# model settings
+input_size = 300
+model = dict(
+ bbox_head=dict(
+ type='SSDHead',
+ anchor_generator=dict(
+ type='LegacySSDAnchorGenerator',
+ scale_major=False,
+ input_size=input_size,
+ basesize_ratio_range=(0.15, 0.9),
+ strides=[8, 16, 32, 64, 100, 300],
+ ratios=[[2], [2, 3], [2, 3], [2, 3], [2], [2]]),
+ bbox_coder=dict(
+ type='LegacyDeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2])))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..ee8015ba12286a9bf940bf2b690441505e39ec0e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/README.md
@@ -0,0 +1,53 @@
+# Libra R-CNN
+
+> [Libra R-CNN: Towards Balanced Learning for Object Detection](https://arxiv.org/abs/1904.02701)
+
+
+
+## Abstract
+
+Compared with model architectures, the training process, which is also crucial to the success of detectors, has received relatively less attention in object detection. In this work, we carefully revisit the standard training practice of detectors, and find that the detection performance is often limited by the imbalance during the training process, which generally consists in three levels - sample level, feature level, and objective level. To mitigate the adverse effects caused thereby, we propose Libra R-CNN, a simple but effective framework towards balanced learning for object detection. It integrates three novel components: IoU-balanced sampling, balanced feature pyramid, and balanced L1 loss, respectively for reducing the imbalance at sample, feature, and objective level. Benefitted from the overall balanced design, Libra R-CNN significantly improves the detection performance. Without bells and whistles, it achieves 2.5 points and 2.0 points higher Average Precision (AP) than FPN Faster R-CNN and RetinaNet respectively on MSCOCO.
+
+Instance recognition is rapidly advanced along with the developments of various deep convolutional neural networks. Compared to the architectures of networks, the training process, which is also crucial to the success of detectors, has received relatively less attention. In this work, we carefully revisit the standard training practice of detectors, and find that the detection performance is often limited by the imbalance during the training process, which generally consists in three levels - sample level, feature level, and objective level. To mitigate the adverse effects caused thereby, we propose Libra R-CNN, a simple yet effective framework towards balanced learning for instance recognition. It integrates IoU-balanced sampling, balanced feature pyramid, and objective re-weighting, respectively for reducing the imbalance at sample, feature, and objective level. Extensive experiments conducted on MS COCO, LVIS and Pascal VOC datasets prove the effectiveness of the overall balanced design.
+
+
+

+
+
+## Results and Models
+
+The results on COCO 2017val are shown in the below table. (results on test-dev are usually slightly higher than val)
+
+| Architecture | Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :----------: | :-------------: | :-----: | :-----: | :------: | :------------: | :----: | :-----------------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Faster R-CNN | R-50-FPN | pytorch | 1x | 4.6 | 19.0 | 38.3 | [config](./libra-faster-rcnn_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/libra_rcnn/libra_faster_rcnn_r50_fpn_1x_coco/libra_faster_rcnn_r50_fpn_1x_coco_20200130-3afee3a9.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/libra_rcnn/libra_faster_rcnn_r50_fpn_1x_coco/libra_faster_rcnn_r50_fpn_1x_coco_20200130_204655.log.json) |
+| Fast R-CNN | R-50-FPN | pytorch | 1x | | | | | |
+| Faster R-CNN | R-101-FPN | pytorch | 1x | 6.5 | 14.4 | 40.1 | [config](./libra-faster-rcnn_r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/libra_rcnn/libra_faster_rcnn_r101_fpn_1x_coco/libra_faster_rcnn_r101_fpn_1x_coco_20200203-8dba6a5a.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/libra_rcnn/libra_faster_rcnn_r101_fpn_1x_coco/libra_faster_rcnn_r101_fpn_1x_coco_20200203_001405.log.json) |
+| Faster R-CNN | X-101-64x4d-FPN | pytorch | 1x | 10.8 | 8.5 | 42.7 | [config](./libra-faster-rcnn_x101-64x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/libra_rcnn/libra_faster_rcnn_x101_64x4d_fpn_1x_coco/libra_faster_rcnn_x101_64x4d_fpn_1x_coco_20200315-3a7d0488.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/libra_rcnn/libra_faster_rcnn_x101_64x4d_fpn_1x_coco/libra_faster_rcnn_x101_64x4d_fpn_1x_coco_20200315_231625.log.json) |
+| RetinaNet | R-50-FPN | pytorch | 1x | 4.2 | 17.7 | 37.6 | [config](./libra-retinanet_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/libra_rcnn/libra_retinanet_r50_fpn_1x_coco/libra_retinanet_r50_fpn_1x_coco_20200205-804d94ce.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/libra_rcnn/libra_retinanet_r50_fpn_1x_coco/libra_retinanet_r50_fpn_1x_coco_20200205_112757.log.json) |
+
+## Citation
+
+We provide config files to reproduce the results in the CVPR 2019 paper [Libra R-CNN](https://arxiv.org/pdf/1904.02701.pdf).
+
+The extended version of [Libra R-CNN](https://arxiv.org/pdf/2108.10175.pdf) is accpeted by IJCV.
+
+```latex
+@inproceedings{pang2019libra,
+ title={Libra R-CNN: Towards Balanced Learning for Object Detection},
+ author={Pang, Jiangmiao and Chen, Kai and Shi, Jianping and Feng, Huajun and Ouyang, Wanli and Dahua Lin},
+ booktitle={IEEE Conference on Computer Vision and Pattern Recognition},
+ year={2019}
+}
+
+@article{pang2021towards,
+ title={Towards Balanced Learning for Instance Recognition},
+ author={Pang, Jiangmiao and Chen, Kai and Li, Qi and Xu, Zhihai and Feng, Huajun and Shi, Jianping and Ouyang, Wanli and Lin, Dahua},
+ journal={International Journal of Computer Vision},
+ volume={129},
+ number={5},
+ pages={1376--1393},
+ year={2021},
+ publisher={Springer}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/libra-fast-rcnn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/libra-fast-rcnn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..2efe440ce361d5bc5855c76001a5ff6b661a568a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/libra-fast-rcnn_r50_fpn_1x_coco.py
@@ -0,0 +1,52 @@
+_base_ = '../fast_rcnn/fast-rcnn_r50_fpn_1x_coco.py'
+# model settings
+model = dict(
+ neck=[
+ dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5),
+ dict(
+ type='BFP',
+ in_channels=256,
+ num_levels=5,
+ refine_level=2,
+ refine_type='non_local')
+ ],
+ roi_head=dict(
+ bbox_head=dict(
+ loss_bbox=dict(
+ _delete_=True,
+ type='BalancedL1Loss',
+ alpha=0.5,
+ gamma=1.5,
+ beta=1.0,
+ loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rcnn=dict(
+ sampler=dict(
+ _delete_=True,
+ type='CombinedSampler',
+ num=512,
+ pos_fraction=0.25,
+ add_gt_as_proposals=True,
+ pos_sampler=dict(type='InstanceBalancedPosSampler'),
+ neg_sampler=dict(
+ type='IoUBalancedNegSampler',
+ floor_thr=-1,
+ floor_fraction=0,
+ num_bins=3)))))
+
+# MMEngine support the following two ways, users can choose
+# according to convenience
+# _base_.train_dataloader.dataset.proposal_file = 'libra_proposals/rpn_r50_fpn_1x_train2017.pkl' # noqa
+train_dataloader = dict(
+ dataset=dict(proposal_file='libra_proposals/rpn_r50_fpn_1x_train2017.pkl'))
+
+# _base_.val_dataloader.dataset.proposal_file = 'libra_proposals/rpn_r50_fpn_1x_val2017.pkl' # noqa
+# test_dataloader = _base_.val_dataloader
+val_dataloader = dict(
+ dataset=dict(proposal_file='libra_proposals/rpn_r50_fpn_1x_val2017.pkl'))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/libra-faster-rcnn_r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/libra-faster-rcnn_r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..985df64cb437e233f76235ee9be4b788ec8f701c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/libra-faster-rcnn_r101_fpn_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './libra-faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/libra-faster-rcnn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/libra-faster-rcnn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f9ee507d26338b49eca004ee195fd2b1954c32d9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/libra-faster-rcnn_r50_fpn_1x_coco.py
@@ -0,0 +1,41 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+# model settings
+model = dict(
+ neck=[
+ dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5),
+ dict(
+ type='BFP',
+ in_channels=256,
+ num_levels=5,
+ refine_level=2,
+ refine_type='non_local')
+ ],
+ roi_head=dict(
+ bbox_head=dict(
+ loss_bbox=dict(
+ _delete_=True,
+ type='BalancedL1Loss',
+ alpha=0.5,
+ gamma=1.5,
+ beta=1.0,
+ loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(sampler=dict(neg_pos_ub=5), allowed_border=-1),
+ rcnn=dict(
+ sampler=dict(
+ _delete_=True,
+ type='CombinedSampler',
+ num=512,
+ pos_fraction=0.25,
+ add_gt_as_proposals=True,
+ pos_sampler=dict(type='InstanceBalancedPosSampler'),
+ neg_sampler=dict(
+ type='IoUBalancedNegSampler',
+ floor_thr=-1,
+ floor_fraction=0,
+ num_bins=3)))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/libra-faster-rcnn_x101-64x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/libra-faster-rcnn_x101-64x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..158e238ed14d9c56b7d02d17f0061b08d4116282
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/libra-faster-rcnn_x101-64x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './libra-faster-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/libra-retinanet_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/libra-retinanet_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..be2742098fb8f1e46bbb16c9d3e2e20c2e3083aa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/libra-retinanet_r50_fpn_1x_coco.py
@@ -0,0 +1,26 @@
+_base_ = '../retinanet/retinanet_r50_fpn_1x_coco.py'
+# model settings
+model = dict(
+ neck=[
+ dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_input',
+ num_outs=5),
+ dict(
+ type='BFP',
+ in_channels=256,
+ num_levels=5,
+ refine_level=1,
+ refine_type='non_local')
+ ],
+ bbox_head=dict(
+ loss_bbox=dict(
+ _delete_=True,
+ type='BalancedL1Loss',
+ alpha=0.5,
+ gamma=1.5,
+ beta=0.11,
+ loss_weight=1.0)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..f01bd02bb7a55dd899bc64a56346357f2951f6d5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/libra_rcnn/metafile.yml
@@ -0,0 +1,99 @@
+Collections:
+ - Name: Libra R-CNN
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - IoU-Balanced Sampling
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Balanced Feature Pyramid
+ Paper:
+ URL: https://arxiv.org/abs/1904.02701
+ Title: 'Libra R-CNN: Towards Balanced Learning for Object Detection'
+ README: configs/libra_rcnn/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/necks/bfp.py#L10
+ Version: v2.0.0
+
+Models:
+ - Name: libra-faster-rcnn_r50_fpn_1x_coco
+ In Collection: Libra R-CNN
+ Config: configs/libra_rcnn/libra-faster-rcnn_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.6
+ inference time (ms/im):
+ - value: 52.63
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/libra_rcnn/libra_faster_rcnn_r50_fpn_1x_coco/libra_faster_rcnn_r50_fpn_1x_coco_20200130-3afee3a9.pth
+
+ - Name: libra-faster-rcnn_r101_fpn_1x_coco
+ In Collection: Libra R-CNN
+ Config: configs/libra_rcnn/libra-faster-rcnn_r101_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.5
+ inference time (ms/im):
+ - value: 69.44
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/libra_rcnn/libra_faster_rcnn_r101_fpn_1x_coco/libra_faster_rcnn_r101_fpn_1x_coco_20200203-8dba6a5a.pth
+
+ - Name: libra-faster-rcnn_x101-64x4d_fpn_1x_coco
+ In Collection: Libra R-CNN
+ Config: configs/libra_rcnn/libra-faster-rcnn_x101-64x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 10.8
+ inference time (ms/im):
+ - value: 117.65
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/libra_rcnn/libra_faster_rcnn_x101_64x4d_fpn_1x_coco/libra_faster_rcnn_x101_64x4d_fpn_1x_coco_20200315-3a7d0488.pth
+
+ - Name: libra-retinanet_r50_fpn_1x_coco
+ In Collection: Libra R-CNN
+ Config: configs/libra_rcnn/libra-retinanet_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.2
+ inference time (ms/im):
+ - value: 56.5
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/libra_rcnn/libra_retinanet_r50_fpn_1x_coco/libra_retinanet_r50_fpn_1x_coco_20200205-804d94ce.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..57aeda438b3cb55e7c3c0d22cddc27a41e6fa3ae
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/README.md
@@ -0,0 +1,56 @@
+# LVIS
+
+> [LVIS: A Dataset for Large Vocabulary Instance Segmentation](https://arxiv.org/abs/1908.03195)
+
+
+
+## Abstract
+
+Progress on object detection is enabled by datasets that focus the research community's attention on open challenges. This process led us from simple images to complex scenes and from bounding boxes to segmentation masks. In this work, we introduce LVIS (pronounced \`el-vis'): a new dataset for Large Vocabulary Instance Segmentation. We plan to collect ~2 million high-quality instance segmentation masks for over 1000 entry-level object categories in 164k images. Due to the Zipfian distribution of categories in natural images, LVIS naturally has a long tail of categories with few training samples. Given that state-of-the-art deep learning methods for object detection perform poorly in the low-sample regime, we believe that our dataset poses an important and exciting new scientific challenge.
+
+
+

+
+
+## Common Setting
+
+- Please follow [install guide](../../docs/get_started.md#install-mmdetection) to install open-mmlab forked cocoapi first.
+
+- Run following scripts to install our forked lvis-api.
+
+ ```shell
+ pip install git+https://github.com/lvis-dataset/lvis-api.git
+ ```
+
+- All experiments use oversample strategy [here](../../docs/tutorials/customize_dataset.md#class-balanced-dataset) with oversample threshold `1e-3`.
+
+- The size of LVIS v0.5 is half of COCO, so schedule `2x` in LVIS is roughly the same iterations as `1x` in COCO.
+
+## Results and models of LVIS v0.5
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :-------------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :----------------------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | pytorch | 2x | - | - | 26.1 | 25.9 | [config](./mask-rcnn_r50_fpn_sample1e-3_ms-2x_lvis-v0.5.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_r50_fpn_sample1e-3_mstrain_2x_lvis/mask_rcnn_r50_fpn_sample1e-3_mstrain_2x_lvis-dbd06831.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_r50_fpn_sample1e-3_mstrain_2x_lvis/mask_rcnn_r50_fpn_sample1e-3_mstrain_2x_lvis_20200531_160435.log.json) |
+| R-101-FPN | pytorch | 2x | - | - | 27.1 | 27.0 | [config](./mask-rcnn_r101_fpn_sample1e-3_ms-2x_lvis-v0.5.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_r101_fpn_sample1e-3_mstrain_2x_lvis/mask_rcnn_r101_fpn_sample1e-3_mstrain_2x_lvis-54582ee2.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_r101_fpn_sample1e-3_mstrain_2x_lvis/mask_rcnn_r101_fpn_sample1e-3_mstrain_2x_lvis_20200601_134748.log.json) |
+| X-101-32x4d-FPN | pytorch | 2x | - | - | 26.7 | 26.9 | [config](./mask-rcnn_x101-32x4d_fpn_sample1e-3_ms-2x_lvis-v0.5.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_x101_32x4d_fpn_sample1e-3_mstrain_2x_lvis/mask_rcnn_x101_32x4d_fpn_sample1e-3_mstrain_2x_lvis-3cf55ea2.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_x101_32x4d_fpn_sample1e-3_mstrain_2x_lvis/mask_rcnn_x101_32x4d_fpn_sample1e-3_mstrain_2x_lvis_20200531_221749.log.json) |
+| X-101-64x4d-FPN | pytorch | 2x | - | - | 26.4 | 26.0 | [config](./mask-rcnn_x101-64x4d_fpn_sample1e-3_ms-2x_lvis-v0.5.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_x101_64x4d_fpn_sample1e-3_mstrain_2x_lvis/mask_rcnn_x101_64x4d_fpn_sample1e-3_mstrain_2x_lvis-1c99a5ad.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_x101_64x4d_fpn_sample1e-3_mstrain_2x_lvis/mask_rcnn_x101_64x4d_fpn_sample1e-3_mstrain_2x_lvis_20200601_194651.log.json) |
+
+## Results and models of LVIS v1
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :-------------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :--------------------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | pytorch | 1x | 9.1 | - | 22.5 | 21.7 | [config](./mask-rcnn_r50_fpn_sample1e-3_ms-1x_lvis-v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_r50_fpn_sample1e-3_mstrain_1x_lvis_v1/mask_rcnn_r50_fpn_sample1e-3_mstrain_1x_lvis_v1-aa78ac3d.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_r50_fpn_sample1e-3_mstrain_1x_lvis_v1/mask_rcnn_r50_fpn_sample1e-3_mstrain_1x_lvis_v1-20200829_061305.log.json) |
+| R-101-FPN | pytorch | 1x | 10.8 | - | 24.6 | 23.6 | [config](./mask-rcnn_r101_fpn_sample1e-3_ms-1x_lvis-v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_r101_fpn_sample1e-3_mstrain_1x_lvis_v1/mask_rcnn_r101_fpn_sample1e-3_mstrain_1x_lvis_v1-ec55ce32.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_r101_fpn_sample1e-3_mstrain_1x_lvis_v1/mask_rcnn_r101_fpn_sample1e-3_mstrain_1x_lvis_v1-20200829_070959.log.json) |
+| X-101-32x4d-FPN | pytorch | 1x | 11.8 | - | 26.7 | 25.5 | [config](./mask-rcnn_x101-32x4d_fpn_sample1e-3_ms-1x_lvis-v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_x101_32x4d_fpn_sample1e-3_mstrain_1x_lvis_v1/mask_rcnn_x101_32x4d_fpn_sample1e-3_mstrain_1x_lvis_v1-ebbc5c81.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_x101_32x4d_fpn_sample1e-3_mstrain_1x_lvis_v1/mask_rcnn_x101_32x4d_fpn_sample1e-3_mstrain_1x_lvis_v1-20200829_071317.log.json) |
+| X-101-64x4d-FPN | pytorch | 1x | 14.6 | - | 27.2 | 25.8 | [config](./mask-rcnn_x101-64x4d_fpn_sample1e-3_ms-1x_lvis-v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_x101_64x4d_fpn_sample1e-3_mstrain_1x_lvis_v1/mask_rcnn_x101_64x4d_fpn_sample1e-3_mstrain_1x_lvis_v1-43d9edfe.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_x101_64x4d_fpn_sample1e-3_mstrain_1x_lvis_v1/mask_rcnn_x101_64x4d_fpn_sample1e-3_mstrain_1x_lvis_v1-20200830_060206.log.json) |
+
+## Citation
+
+```latex
+@inproceedings{gupta2019lvis,
+ title={{LVIS}: A Dataset for Large Vocabulary Instance Segmentation},
+ author={Gupta, Agrim and Dollar, Piotr and Girshick, Ross},
+ booktitle={Proceedings of the {IEEE} Conference on Computer Vision and Pattern Recognition},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_r101_fpn_sample1e-3_ms-1x_lvis-v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_r101_fpn_sample1e-3_ms-1x_lvis-v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..3994d75a81aaa5368bd42c591fa770b05b665e25
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_r101_fpn_sample1e-3_ms-1x_lvis-v1.py
@@ -0,0 +1,6 @@
+_base_ = './mask-rcnn_r50_fpn_sample1e-3_ms-1x_lvis-v1.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_r101_fpn_sample1e-3_ms-2x_lvis-v0.5.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_r101_fpn_sample1e-3_ms-2x_lvis-v0.5.py
new file mode 100644
index 0000000000000000000000000000000000000000..ed8b3639a0046e14d5c11a98f9d7dc38eb4badec
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_r101_fpn_sample1e-3_ms-2x_lvis-v0.5.py
@@ -0,0 +1,6 @@
+_base_ = './mask-rcnn_r50_fpn_sample1e-3_ms-2x_lvis-v0.5.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_r50_fpn_sample1e-3_ms-1x_lvis-v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_r50_fpn_sample1e-3_ms-1x_lvis-v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..cdd3683e3005dd09ada78827825da516bfd4c66e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_r50_fpn_sample1e-3_ms-1x_lvis-v1.py
@@ -0,0 +1,13 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/lvis_v1_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ roi_head=dict(
+ bbox_head=dict(num_classes=1203), mask_head=dict(num_classes=1203)),
+ test_cfg=dict(
+ rcnn=dict(
+ score_thr=0.0001,
+ # LVIS allows up to 300
+ max_per_img=300)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_r50_fpn_sample1e-3_ms-2x_lvis-v0.5.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_r50_fpn_sample1e-3_ms-2x_lvis-v0.5.py
new file mode 100644
index 0000000000000000000000000000000000000000..b36b6c17fef7da3646654e494fa715302b1b050e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_r50_fpn_sample1e-3_ms-2x_lvis-v0.5.py
@@ -0,0 +1,13 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/lvis_v0.5_instance.py',
+ '../_base_/schedules/schedule_2x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ roi_head=dict(
+ bbox_head=dict(num_classes=1230), mask_head=dict(num_classes=1230)),
+ test_cfg=dict(
+ rcnn=dict(
+ score_thr=0.0001,
+ # LVIS allows up to 300
+ max_per_img=300)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_x101-32x4d_fpn_sample1e-3_ms-1x_lvis-v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_x101-32x4d_fpn_sample1e-3_ms-1x_lvis-v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..9da3ab6db04ec6ee772202270a47179171a9d13c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_x101-32x4d_fpn_sample1e-3_ms-1x_lvis-v1.py
@@ -0,0 +1,14 @@
+_base_ = './mask-rcnn_r50_fpn_sample1e-3_ms-1x_lvis-v1.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_x101-32x4d_fpn_sample1e-3_ms-2x_lvis-v0.5.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_x101-32x4d_fpn_sample1e-3_ms-2x_lvis-v0.5.py
new file mode 100644
index 0000000000000000000000000000000000000000..9a097c94c7e2d7c7b583027ce6000aba8205d490
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_x101-32x4d_fpn_sample1e-3_ms-2x_lvis-v0.5.py
@@ -0,0 +1,14 @@
+_base_ = './mask-rcnn_r50_fpn_sample1e-3_ms-2x_lvis-v0.5.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_x101-64x4d_fpn_sample1e-3_ms-1x_lvis-v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_x101-64x4d_fpn_sample1e-3_ms-1x_lvis-v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..b0819b3ec60d710205a643305edd2a27db977d9b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_x101-64x4d_fpn_sample1e-3_ms-1x_lvis-v1.py
@@ -0,0 +1,14 @@
+_base_ = './mask-rcnn_r50_fpn_sample1e-3_ms-1x_lvis-v1.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_x101-64x4d_fpn_sample1e-3_ms-2x_lvis-v0.5.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_x101-64x4d_fpn_sample1e-3_ms-2x_lvis-v0.5.py
new file mode 100644
index 0000000000000000000000000000000000000000..9d2720089181f066bcaa04b73903836b64b97bb9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/mask-rcnn_x101-64x4d_fpn_sample1e-3_ms-2x_lvis-v0.5.py
@@ -0,0 +1,14 @@
+_base_ = './mask-rcnn_r50_fpn_sample1e-3_ms-2x_lvis-v0.5.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..f8def96c7e5404bba0b40f4f00ce9efabfe0a891
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/lvis/metafile.yml
@@ -0,0 +1,128 @@
+Models:
+ - Name: mask-rcnn_r50_fpn_sample1e-3_ms-2x_lvis-v0.5
+ In Collection: Mask R-CNN
+ Config: configs/lvis/mask-rcnn_r50_fpn_sample1e-3_ms-2x_lvis-v0.5.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v0.5
+ Metrics:
+ box AP: 26.1
+ - Task: Instance Segmentation
+ Dataset: LVIS v0.5
+ Metrics:
+ mask AP: 25.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_r50_fpn_sample1e-3_mstrain_2x_lvis/mask_rcnn_r50_fpn_sample1e-3_mstrain_2x_lvis-dbd06831.pth
+
+ - Name: mask-rcnn_r101_fpn_sample1e-3_ms-2x_lvis-v0.5
+ In Collection: Mask R-CNN
+ Config: configs/lvis/mask-rcnn_r101_fpn_sample1e-3_ms-2x_lvis-v0.5.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v0.5
+ Metrics:
+ box AP: 27.1
+ - Task: Instance Segmentation
+ Dataset: LVIS v0.5
+ Metrics:
+ mask AP: 27.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_r101_fpn_sample1e-3_mstrain_2x_lvis/mask_rcnn_r101_fpn_sample1e-3_mstrain_2x_lvis-54582ee2.pth
+
+ - Name: mask-rcnn_x101-32x4d_fpn_sample1e-3_ms-2x_lvis-v0.5
+ In Collection: Mask R-CNN
+ Config: configs/lvis/mask-rcnn_x101-32x4d_fpn_sample1e-3_ms-2x_lvis-v0.5.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v0.5
+ Metrics:
+ box AP: 26.7
+ - Task: Instance Segmentation
+ Dataset: LVIS v0.5
+ Metrics:
+ mask AP: 26.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_x101_32x4d_fpn_sample1e-3_mstrain_2x_lvis/mask_rcnn_x101_32x4d_fpn_sample1e-3_mstrain_2x_lvis-3cf55ea2.pth
+
+ - Name: mask-rcnn_x101-64x4d_fpn_sample1e-3_ms-2x_lvis-v0.5
+ In Collection: Mask R-CNN
+ Config: configs/lvis/mask-rcnn_x101-64x4d_fpn_sample1e-3_ms-2x_lvis-v0.5.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v0.5
+ Metrics:
+ box AP: 26.4
+ - Task: Instance Segmentation
+ Dataset: LVIS v0.5
+ Metrics:
+ mask AP: 26.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_x101_64x4d_fpn_sample1e-3_mstrain_2x_lvis/mask_rcnn_x101_64x4d_fpn_sample1e-3_mstrain_2x_lvis-1c99a5ad.pth
+
+ - Name: mask-rcnn_r50_fpn_sample1e-3_ms-1x_lvis-v1
+ In Collection: Mask R-CNN
+ Config: configs/lvis/mask-rcnn_r50_fpn_sample1e-3_ms-1x_lvis-v1.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v1
+ Metrics:
+ box AP: 22.5
+ - Task: Instance Segmentation
+ Dataset: LVIS v1
+ Metrics:
+ mask AP: 21.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_r50_fpn_sample1e-3_mstrain_1x_lvis_v1/mask_rcnn_r50_fpn_sample1e-3_mstrain_1x_lvis_v1-aa78ac3d.pth
+
+ - Name: mask-rcnn_r101_fpn_sample1e-3_ms-1x_lvis-v1
+ In Collection: Mask R-CNN
+ Config: configs/lvis/mask-rcnn_r101_fpn_sample1e-3_ms-1x_lvis-v1.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v1
+ Metrics:
+ box AP: 24.6
+ - Task: Instance Segmentation
+ Dataset: LVIS v1
+ Metrics:
+ mask AP: 23.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_r101_fpn_sample1e-3_mstrain_1x_lvis_v1/mask_rcnn_r101_fpn_sample1e-3_mstrain_1x_lvis_v1-ec55ce32.pth
+
+ - Name: mask-rcnn_x101-32x4d_fpn_sample1e-3_ms-1x_lvis-v1
+ In Collection: Mask R-CNN
+ Config: configs/lvis/mask-rcnn_x101-32x4d_fpn_sample1e-3_ms-1x_lvis-v1.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v1
+ Metrics:
+ box AP: 26.7
+ - Task: Instance Segmentation
+ Dataset: LVIS v1
+ Metrics:
+ mask AP: 25.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_x101_32x4d_fpn_sample1e-3_mstrain_1x_lvis_v1/mask_rcnn_x101_32x4d_fpn_sample1e-3_mstrain_1x_lvis_v1-ebbc5c81.pth
+
+ - Name: mask-rcnn_x101-64x4d_fpn_sample1e-3_ms-1x_lvis-v1
+ In Collection: Mask R-CNN
+ Config: configs/lvis/mask-rcnn_x101-64x4d_fpn_sample1e-3_ms-1x_lvis-v1.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v1
+ Metrics:
+ box AP: 27.2
+ - Task: Instance Segmentation
+ Dataset: LVIS v1
+ Metrics:
+ mask AP: 25.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/lvis/mask_rcnn_x101_64x4d_fpn_sample1e-3_mstrain_1x_lvis_v1/mask_rcnn_x101_64x4d_fpn_sample1e-3_mstrain_1x_lvis_v1-43d9edfe.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..94b0821e7a2f3a467f48f8f7581e6c10d1571404
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/README.md
@@ -0,0 +1,76 @@
+# Mask2Former
+
+> [Masked-attention Mask Transformer for Universal Image Segmentation](http://arxiv.org/abs/2112.01527)
+
+
+
+## Abstract
+
+Image segmentation is about grouping pixels with different semantics, e.g., category or instance membership, where each choice of semantics defines a task. While only the semantics of each task differ, current research focuses on designing specialized architectures for each task. We present Masked-attention Mask Transformer (Mask2Former), a new architecture capable of addressing any image segmentation task (panoptic, instance or semantic). Its key components include masked attention, which extracts localized features by constraining cross-attention within predicted mask regions. In addition to reducing the research effort by at least three times, it outperforms the best specialized architectures by a significant margin on four popular datasets. Most notably, Mask2Former sets a new state-of-the-art for panoptic segmentation (57.8 PQ on COCO), instance segmentation (50.1 AP on COCO) and semantic segmentation (57.7 mIoU on ADE20K).
+
+
+

+
+
+## Introduction
+
+Mask2Former requires COCO and [COCO-panoptic](http://images.cocodataset.org/annotations/panoptic_annotations_trainval2017.zip) dataset for training and evaluation. You need to download and extract it in the COCO dataset path.
+The directory should be like this.
+
+```none
+mmdetection
+├── mmdet
+├── tools
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+| | | ├── instances_train2017.json
+| | | ├── instances_val2017.json
+│ │ │ ├── panoptic_train2017.json
+│ │ │ ├── panoptic_train2017
+│ │ │ ├── panoptic_val2017.json
+│ │ │ ├── panoptic_val2017
+│ │ ├── train2017
+│ │ ├── val2017
+│ │ ├── test2017
+```
+
+## Results and Models
+
+### Panoptic segmentation
+
+| Backbone | style | Pretrain | Lr schd | Mem (GB) | Inf time (fps) | PQ | box mAP | mask mAP | Config | Download |
+| -------- | ------- | ------------ | ------- | -------- | -------------- | ---- | ------- | -------- | ------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
+| R-50 | pytorch | ImageNet-1K | 50e | 13.9 | - | 52.0 | 44.5 | 41.8 | [config](./mask2former_r50_8xb2-lsj-50e_coco-panoptic.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_r50_8xb2-lsj-50e_coco-panoptic/mask2former_r50_8xb2-lsj-50e_coco-panoptic_20230118_125535-54df384a.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_r50_8xb2-lsj-50e_coco-panoptic/mask2former_r50_8xb2-lsj-50e_coco-panoptic_20230118_125535.log.json) |
+| R-101 | pytorch | ImageNet-1K | 50e | 16.1 | - | 52.4 | 45.3 | 42.4 | [config](./mask2former_r101_8xb2-lsj-50e_coco-panoptic.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_r101_8xb2-lsj-50e_coco-panoptic/mask2former_r101_8xb2-lsj-50e_coco-panoptic_20220329_225104-c74d4d71.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask2former/mask2former_r101_lsj_8x2_50e_coco-panoptic/mask2former_r101_lsj_8x2_50e_coco-panoptic_20220329_225104.log.json) |
+| Swin-T | - | ImageNet-1K | 50e | 15.9 | - | 53.4 | 46.3 | 43.4 | [config](./mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco-panoptic.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco-panoptic/mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco-panoptic_20220326_224553-3ec9e0ae.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask2former/mask2former_swin-t-p4-w7-224_lsj_8x2_50e_coco-panoptic/mask2former_swin-t-p4-w7-224_lsj_8x2_50e_coco-panoptic_20220326_224553.log.json) |
+| Swin-S | - | ImageNet-1K | 50e | 19.1 | - | 54.5 | 47.8 | 44.5 | [config](./mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco-panoptic.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco-panoptic/mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco-panoptic_20220329_225200-4a16ded7.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask2former/mask2former_swin-s-p4-w7-224_lsj_8x2_50e_coco-panoptic/mask2former_swin-s-p4-w7-224_lsj_8x2_50e_coco-panoptic_20220329_225200.log.json) |
+| Swin-B | - | ImageNet-1K | 50e | 26.0 | - | 55.1 | 48.2 | 44.9 | [config](./mask2former_swin-b-p4-w12-384_8xb2-lsj-50e_coco-panoptic.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_swin-b-p4-w12-384_8xb2-lsj-50e_coco-panoptic/mask2former_swin-b-p4-w12-384_8xb2-lsj-50e_coco-panoptic_20220331_002244-8a651d82.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask2former/mask2former_swin-b-p4-w12-384_lsj_8x2_50e_coco-panoptic/mask2former_swin-b-p4-w12-384_lsj_8x2_50e_coco-panoptic_20220331_002244.log.json) |
+| Swin-B | - | ImageNet-21K | 50e | 25.8 | - | 56.3 | 50.0 | 46.3 | [config](./mask2former_swin-b-p4-w12-384-in21k_8xb2-lsj-50e_coco-panoptic.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_swin-b-p4-w12-384-in21k_8xb2-lsj-50e_coco-panoptic/mask2former_swin-b-p4-w12-384-in21k_8xb2-lsj-50e_coco-panoptic_20220329_230021-05ec7315.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask2former/mask2former_swin-b-p4-w12-384-in21k_lsj_8x2_50e_coco-panoptic/mask2former_swin-b-p4-w12-384-in21k_lsj_8x2_50e_coco-panoptic_20220329_230021.log.json) |
+| Swin-L | - | ImageNet-21K | 100e | 21.1 | - | 57.6 | 52.2 | 48.5 | [config](./mask2former_swin-l-p4-w12-384-in21k_16xb1-lsj-100e_coco-panoptic.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_swin-l-p4-w12-384-in21k_16xb1-lsj-100e_coco-panoptic/mask2former_swin-l-p4-w12-384-in21k_16xb1-lsj-100e_coco-panoptic_20220407_104949-82f8d28d.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask2former/mask2former_swin-l-p4-w12-384-in21k_lsj_16x1_100e_coco-panoptic/mask2former_swin-l-p4-w12-384-in21k_lsj_16x1_100e_coco-panoptic_20220407_104949.log.json) |
+
+### Instance segmentation
+
+| Backbone | style | Pretrain | Lr schd | Mem (GB) | Inf time (fps) | box mAP | mask mAP | Config | Download |
+| -------- | ------- | ----------- | ------- | -------- | -------------- | ------- | -------- | ------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
+| R-50 | pytorch | ImageNet-1K | 50e | 13.7 | - | 45.7 | 42.9 | [config](./mask2former_r50_8xb2-lsj-50e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_r50_8xb2-lsj-50e_coco/mask2former_r50_8xb2-lsj-50e_coco_20220506_191028-41b088b6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask2former/mask2former_r50_lsj_8x2_50e_coco/mask2former_r50_lsj_8x2_50e_coco_20220506_191028.log.json) |
+| R-101 | pytorch | ImageNet-1K | 50e | 15.5 | - | 46.7 | 44.0 | [config](./mask2former_r101_8xb2-lsj-50e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_r101_8xb2-lsj-50e_coco/mask2former_r101_8xb2-lsj-50e_coco_20220426_100250-ecf181e2.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask2former/mask2former_r101_lsj_8x2_50e_coco/mask2former_r101_lsj_8x2_50e_coco_20220426_100250.log.json) |
+| Swin-T | - | ImageNet-1K | 50e | 15.3 | - | 47.7 | 44.7 | [config](./mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco/mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco_20220508_091649-01b0f990.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask2former/mask2former_swin-t-p4-w7-224_lsj_8x2_50e_coco/mask2former_swin-t-p4-w7-224_lsj_8x2_50e_coco_20220508_091649.log.json) |
+| Swin-S | - | ImageNet-1K | 50e | 18.8 | - | 49.3 | 46.1 | [config](./mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco/mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco_20220504_001756-c9d0c4f2.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask2former/mask2former_swin-s-p4-w7-224_lsj_8x2_50e_coco/mask2former_swin-s-p4-w7-224_lsj_8x2_50e_coco_20220504_001756.log.json) |
+
+### Note
+
+1. The performance is unstable. The `Mask2Former-R50-coco-panoptic` may fluctuate about 0.2 PQ. The models other than `Mask2Former-R50-coco-panoptic` were trained with mmdet 2.x and have been converted for mmdet 3.x.
+2. We have trained the instance segmentation models many times (see more details in [PR 7571](https://github.com/open-mmlab/mmdetection/pull/7571)). The results of the trained models are relatively stable (+- 0.2), and have a certain gap (about 0.2 AP) in comparison with the results in the [paper](http://arxiv.org/abs/2112.01527). However, the performance of the model trained with the official code is unstable and may also be slightly lower than the reported results as mentioned in the [issue](https://github.com/facebookresearch/Mask2Former/issues/46).
+
+## Citation
+
+```latex
+@article{cheng2021mask2former,
+ title={Masked-attention Mask Transformer for Universal Image Segmentation},
+ author={Bowen Cheng and Ishan Misra and Alexander G. Schwing and Alexander Kirillov and Rohit Girdhar},
+ journal={arXiv},
+ year={2021}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_r101_8xb2-lsj-50e_coco-panoptic.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_r101_8xb2-lsj-50e_coco-panoptic.py
new file mode 100644
index 0000000000000000000000000000000000000000..66685a2fca9c0e165ba0024e242d5eabf5d565c9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_r101_8xb2-lsj-50e_coco-panoptic.py
@@ -0,0 +1,7 @@
+_base_ = './mask2former_r50_8xb2-lsj-50e_coco-panoptic.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_r101_8xb2-lsj-50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_r101_8xb2-lsj-50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f4c29906d9fc6ce47ce928fb73dcb1bb6c6f7ba9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_r101_8xb2-lsj-50e_coco.py
@@ -0,0 +1,7 @@
+_base_ = ['./mask2former_r50_8xb2-lsj-50e_coco.py']
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_r50_8xb2-lsj-50e_coco-panoptic.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_r50_8xb2-lsj-50e_coco-panoptic.py
new file mode 100644
index 0000000000000000000000000000000000000000..c53e981bf0d5081c3735676be922f64298a8fc80
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_r50_8xb2-lsj-50e_coco-panoptic.py
@@ -0,0 +1,251 @@
+_base_ = [
+ '../_base_/datasets/coco_panoptic.py', '../_base_/default_runtime.py'
+]
+image_size = (1024, 1024)
+batch_augments = [
+ dict(
+ type='BatchFixedSizePad',
+ size=image_size,
+ img_pad_value=0,
+ pad_mask=True,
+ mask_pad_value=0,
+ pad_seg=True,
+ seg_pad_value=255)
+]
+data_preprocessor = dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32,
+ pad_mask=True,
+ mask_pad_value=0,
+ pad_seg=True,
+ seg_pad_value=255,
+ batch_augments=batch_augments)
+
+num_things_classes = 80
+num_stuff_classes = 53
+num_classes = num_things_classes + num_stuff_classes
+model = dict(
+ type='Mask2Former',
+ data_preprocessor=data_preprocessor,
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=-1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ panoptic_head=dict(
+ type='Mask2FormerHead',
+ in_channels=[256, 512, 1024, 2048], # pass to pixel_decoder inside
+ strides=[4, 8, 16, 32],
+ feat_channels=256,
+ out_channels=256,
+ num_things_classes=num_things_classes,
+ num_stuff_classes=num_stuff_classes,
+ num_queries=100,
+ num_transformer_feat_level=3,
+ pixel_decoder=dict(
+ type='MSDeformAttnPixelDecoder',
+ num_outs=3,
+ norm_cfg=dict(type='GN', num_groups=32),
+ act_cfg=dict(type='ReLU'),
+ encoder=dict( # DeformableDetrTransformerEncoder
+ num_layers=6,
+ layer_cfg=dict( # DeformableDetrTransformerEncoderLayer
+ self_attn_cfg=dict( # MultiScaleDeformableAttention
+ embed_dims=256,
+ num_heads=8,
+ num_levels=3,
+ num_points=4,
+ dropout=0.0,
+ batch_first=True),
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=1024,
+ num_fcs=2,
+ ffn_drop=0.0,
+ act_cfg=dict(type='ReLU', inplace=True)))),
+ positional_encoding=dict(num_feats=128, normalize=True)),
+ enforce_decoder_input_project=False,
+ positional_encoding=dict(num_feats=128, normalize=True),
+ transformer_decoder=dict( # Mask2FormerTransformerDecoder
+ return_intermediate=True,
+ num_layers=9,
+ layer_cfg=dict( # Mask2FormerTransformerDecoderLayer
+ self_attn_cfg=dict( # MultiheadAttention
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.0,
+ batch_first=True),
+ cross_attn_cfg=dict( # MultiheadAttention
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.0,
+ batch_first=True),
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048,
+ num_fcs=2,
+ ffn_drop=0.0,
+ act_cfg=dict(type='ReLU', inplace=True))),
+ init_cfg=None),
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=2.0,
+ reduction='mean',
+ class_weight=[1.0] * num_classes + [0.1]),
+ loss_mask=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ reduction='mean',
+ loss_weight=5.0),
+ loss_dice=dict(
+ type='DiceLoss',
+ use_sigmoid=True,
+ activate=True,
+ reduction='mean',
+ naive_dice=True,
+ eps=1.0,
+ loss_weight=5.0)),
+ panoptic_fusion_head=dict(
+ type='MaskFormerFusionHead',
+ num_things_classes=num_things_classes,
+ num_stuff_classes=num_stuff_classes,
+ loss_panoptic=None,
+ init_cfg=None),
+ train_cfg=dict(
+ num_points=12544,
+ oversample_ratio=3.0,
+ importance_sample_ratio=0.75,
+ assigner=dict(
+ type='HungarianAssigner',
+ match_costs=[
+ dict(type='ClassificationCost', weight=2.0),
+ dict(
+ type='CrossEntropyLossCost', weight=5.0, use_sigmoid=True),
+ dict(type='DiceCost', weight=5.0, pred_act=True, eps=1.0)
+ ]),
+ sampler=dict(type='MaskPseudoSampler')),
+ test_cfg=dict(
+ panoptic_on=True,
+ # For now, the dataset does not support
+ # evaluating semantic segmentation metric.
+ semantic_on=False,
+ instance_on=True,
+ # max_per_image is for instance segmentation.
+ max_per_image=100,
+ iou_thr=0.8,
+ # In Mask2Former's panoptic postprocessing,
+ # it will filter mask area where score is less than 0.5 .
+ filter_low_score=True),
+ init_cfg=None)
+
+# dataset settings
+data_root = 'data/coco/'
+train_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ to_float32=True,
+ backend_args={{_base_.backend_args}}),
+ dict(
+ type='LoadPanopticAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ with_seg=True,
+ backend_args={{_base_.backend_args}}),
+ dict(type='RandomFlip', prob=0.5),
+ # large scale jittering
+ dict(
+ type='RandomResize',
+ scale=image_size,
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_size=image_size,
+ crop_type='absolute',
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+val_evaluator = [
+ dict(
+ type='CocoPanopticMetric',
+ ann_file=data_root + 'annotations/panoptic_val2017.json',
+ seg_prefix=data_root + 'annotations/panoptic_val2017/',
+ backend_args={{_base_.backend_args}}),
+ dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric=['bbox', 'segm'],
+ backend_args={{_base_.backend_args}})
+]
+test_evaluator = val_evaluator
+
+# optimizer
+embed_multi = dict(lr_mult=1.0, decay_mult=0.0)
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(
+ type='AdamW',
+ lr=0.0001,
+ weight_decay=0.05,
+ eps=1e-8,
+ betas=(0.9, 0.999)),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'backbone': dict(lr_mult=0.1, decay_mult=1.0),
+ 'query_embed': embed_multi,
+ 'query_feat': embed_multi,
+ 'level_embed': embed_multi,
+ },
+ norm_decay_mult=0.0),
+ clip_grad=dict(max_norm=0.01, norm_type=2))
+
+# learning policy
+max_iters = 368750
+param_scheduler = dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_iters,
+ by_epoch=False,
+ milestones=[327778, 355092],
+ gamma=0.1)
+
+# Before 365001th iteration, we do evaluation every 5000 iterations.
+# After 365000th iteration, we do evaluation every 368750 iterations,
+# which means that we do evaluation at the end of training.
+interval = 5000
+dynamic_intervals = [(max_iters // interval * interval + 1, max_iters)]
+train_cfg = dict(
+ type='IterBasedTrainLoop',
+ max_iters=max_iters,
+ val_interval=interval,
+ dynamic_intervals=dynamic_intervals)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+default_hooks = dict(
+ checkpoint=dict(
+ type='CheckpointHook',
+ by_epoch=False,
+ save_last=True,
+ max_keep_ckpts=3,
+ interval=interval))
+log_processor = dict(type='LogProcessor', window_size=50, by_epoch=False)
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_r50_8xb2-lsj-50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_r50_8xb2-lsj-50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..24a17f58c54a2e8694a8bf960d10ebc918acdddc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_r50_8xb2-lsj-50e_coco.py
@@ -0,0 +1,100 @@
+_base_ = ['./mask2former_r50_8xb2-lsj-50e_coco-panoptic.py']
+
+num_things_classes = 80
+num_stuff_classes = 0
+num_classes = num_things_classes + num_stuff_classes
+image_size = (1024, 1024)
+batch_augments = [
+ dict(
+ type='BatchFixedSizePad',
+ size=image_size,
+ img_pad_value=0,
+ pad_mask=True,
+ mask_pad_value=0,
+ pad_seg=False)
+]
+data_preprocessor = dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32,
+ pad_mask=True,
+ mask_pad_value=0,
+ pad_seg=False,
+ batch_augments=batch_augments)
+model = dict(
+ data_preprocessor=data_preprocessor,
+ panoptic_head=dict(
+ num_things_classes=num_things_classes,
+ num_stuff_classes=num_stuff_classes,
+ loss_cls=dict(class_weight=[1.0] * num_classes + [0.1])),
+ panoptic_fusion_head=dict(
+ num_things_classes=num_things_classes,
+ num_stuff_classes=num_stuff_classes),
+ test_cfg=dict(panoptic_on=False))
+
+# dataset settings
+train_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ to_float32=True,
+ backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(type='RandomFlip', prob=0.5),
+ # large scale jittering
+ dict(
+ type='RandomResize',
+ scale=image_size,
+ ratio_range=(0.1, 2.0),
+ resize_type='Resize',
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_size=image_size,
+ crop_type='absolute',
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-5, 1e-5), by_mask=True),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ to_float32=True,
+ backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ # If you don't have a gt annotation, delete the pipeline
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+
+train_dataloader = dict(
+ dataset=dict(
+ type=dataset_type,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ pipeline=train_pipeline))
+val_dataloader = dict(
+ dataset=dict(
+ type=dataset_type,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric=['bbox', 'segm'],
+ format_only=False,
+ backend_args={{_base_.backend_args}})
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-b-p4-w12-384-in21k_8xb2-lsj-50e_coco-panoptic.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-b-p4-w12-384-in21k_8xb2-lsj-50e_coco-panoptic.py
new file mode 100644
index 0000000000000000000000000000000000000000..b275f23175e8d8294b8bb76e9708dd014ef7030b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-b-p4-w12-384-in21k_8xb2-lsj-50e_coco-panoptic.py
@@ -0,0 +1,5 @@
+_base_ = ['./mask2former_swin-b-p4-w12-384_8xb2-lsj-50e_coco-panoptic.py']
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_base_patch4_window12_384_22k.pth' # noqa
+
+model = dict(
+ backbone=dict(init_cfg=dict(type='Pretrained', checkpoint=pretrained)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-b-p4-w12-384_8xb2-lsj-50e_coco-panoptic.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-b-p4-w12-384_8xb2-lsj-50e_coco-panoptic.py
new file mode 100644
index 0000000000000000000000000000000000000000..bd59400b4aed1aac97795e474633d5581705b899
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-b-p4-w12-384_8xb2-lsj-50e_coco-panoptic.py
@@ -0,0 +1,42 @@
+_base_ = ['./mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco-panoptic.py']
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_base_patch4_window12_384.pth' # noqa
+
+depths = [2, 2, 18, 2]
+model = dict(
+ backbone=dict(
+ pretrain_img_size=384,
+ embed_dims=128,
+ depths=depths,
+ num_heads=[4, 8, 16, 32],
+ window_size=12,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ panoptic_head=dict(in_channels=[128, 256, 512, 1024]))
+
+# set all layers in backbone to lr_mult=0.1
+# set all norm layers, position_embeding,
+# query_embeding, level_embeding to decay_multi=0.0
+backbone_norm_multi = dict(lr_mult=0.1, decay_mult=0.0)
+backbone_embed_multi = dict(lr_mult=0.1, decay_mult=0.0)
+embed_multi = dict(lr_mult=1.0, decay_mult=0.0)
+custom_keys = {
+ 'backbone': dict(lr_mult=0.1, decay_mult=1.0),
+ 'backbone.patch_embed.norm': backbone_norm_multi,
+ 'backbone.norm': backbone_norm_multi,
+ 'absolute_pos_embed': backbone_embed_multi,
+ 'relative_position_bias_table': backbone_embed_multi,
+ 'query_embed': embed_multi,
+ 'query_feat': embed_multi,
+ 'level_embed': embed_multi
+}
+custom_keys.update({
+ f'backbone.stages.{stage_id}.blocks.{block_id}.norm': backbone_norm_multi
+ for stage_id, num_blocks in enumerate(depths)
+ for block_id in range(num_blocks)
+})
+custom_keys.update({
+ f'backbone.stages.{stage_id}.downsample.norm': backbone_norm_multi
+ for stage_id in range(len(depths) - 1)
+})
+# optimizer
+optim_wrapper = dict(
+ paramwise_cfg=dict(custom_keys=custom_keys, norm_decay_mult=0.0))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-l-p4-w12-384-in21k_16xb1-lsj-100e_coco-panoptic.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-l-p4-w12-384-in21k_16xb1-lsj-100e_coco-panoptic.py
new file mode 100644
index 0000000000000000000000000000000000000000..e203ffc96c40098e4cf0788fc47b4438ebffbb41
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-l-p4-w12-384-in21k_16xb1-lsj-100e_coco-panoptic.py
@@ -0,0 +1,25 @@
+_base_ = ['./mask2former_swin-b-p4-w12-384_8xb2-lsj-50e_coco-panoptic.py']
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_large_patch4_window12_384_22k.pth' # noqa
+
+model = dict(
+ backbone=dict(
+ embed_dims=192,
+ num_heads=[6, 12, 24, 48],
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ panoptic_head=dict(num_queries=200, in_channels=[192, 384, 768, 1536]))
+
+train_dataloader = dict(batch_size=1, num_workers=1)
+
+# learning policy
+max_iters = 737500
+param_scheduler = dict(end=max_iters, milestones=[655556, 710184])
+
+# Before 735001th iteration, we do evaluation every 5000 iterations.
+# After 735000th iteration, we do evaluation every 737500 iterations,
+# which means that we do evaluation at the end of training.'
+interval = 5000
+dynamic_intervals = [(max_iters // interval * interval + 1, max_iters)]
+train_cfg = dict(
+ max_iters=max_iters,
+ val_interval=interval,
+ dynamic_intervals=dynamic_intervals)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco-panoptic.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco-panoptic.py
new file mode 100644
index 0000000000000000000000000000000000000000..f9d081db58a74dd02b3b715c3777f077d42de7ca
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco-panoptic.py
@@ -0,0 +1,37 @@
+_base_ = ['./mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco-panoptic.py']
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_small_patch4_window7_224.pth' # noqa
+
+depths = [2, 2, 18, 2]
+model = dict(
+ backbone=dict(
+ depths=depths, init_cfg=dict(type='Pretrained',
+ checkpoint=pretrained)))
+
+# set all layers in backbone to lr_mult=0.1
+# set all norm layers, position_embeding,
+# query_embeding, level_embeding to decay_multi=0.0
+backbone_norm_multi = dict(lr_mult=0.1, decay_mult=0.0)
+backbone_embed_multi = dict(lr_mult=0.1, decay_mult=0.0)
+embed_multi = dict(lr_mult=1.0, decay_mult=0.0)
+custom_keys = {
+ 'backbone': dict(lr_mult=0.1, decay_mult=1.0),
+ 'backbone.patch_embed.norm': backbone_norm_multi,
+ 'backbone.norm': backbone_norm_multi,
+ 'absolute_pos_embed': backbone_embed_multi,
+ 'relative_position_bias_table': backbone_embed_multi,
+ 'query_embed': embed_multi,
+ 'query_feat': embed_multi,
+ 'level_embed': embed_multi
+}
+custom_keys.update({
+ f'backbone.stages.{stage_id}.blocks.{block_id}.norm': backbone_norm_multi
+ for stage_id, num_blocks in enumerate(depths)
+ for block_id in range(num_blocks)
+})
+custom_keys.update({
+ f'backbone.stages.{stage_id}.downsample.norm': backbone_norm_multi
+ for stage_id in range(len(depths) - 1)
+})
+# optimizer
+optim_wrapper = dict(
+ paramwise_cfg=dict(custom_keys=custom_keys, norm_decay_mult=0.0))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..69d5e8c6f96434973e3e9f3498155e385af815be
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco.py
@@ -0,0 +1,37 @@
+_base_ = ['./mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco.py']
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_small_patch4_window7_224.pth' # noqa
+
+depths = [2, 2, 18, 2]
+model = dict(
+ backbone=dict(
+ depths=depths, init_cfg=dict(type='Pretrained',
+ checkpoint=pretrained)))
+
+# set all layers in backbone to lr_mult=0.1
+# set all norm layers, position_embeding,
+# query_embeding, level_embeding to decay_multi=0.0
+backbone_norm_multi = dict(lr_mult=0.1, decay_mult=0.0)
+backbone_embed_multi = dict(lr_mult=0.1, decay_mult=0.0)
+embed_multi = dict(lr_mult=1.0, decay_mult=0.0)
+custom_keys = {
+ 'backbone': dict(lr_mult=0.1, decay_mult=1.0),
+ 'backbone.patch_embed.norm': backbone_norm_multi,
+ 'backbone.norm': backbone_norm_multi,
+ 'absolute_pos_embed': backbone_embed_multi,
+ 'relative_position_bias_table': backbone_embed_multi,
+ 'query_embed': embed_multi,
+ 'query_feat': embed_multi,
+ 'level_embed': embed_multi
+}
+custom_keys.update({
+ f'backbone.stages.{stage_id}.blocks.{block_id}.norm': backbone_norm_multi
+ for stage_id, num_blocks in enumerate(depths)
+ for block_id in range(num_blocks)
+})
+custom_keys.update({
+ f'backbone.stages.{stage_id}.downsample.norm': backbone_norm_multi
+ for stage_id in range(len(depths) - 1)
+})
+# optimizer
+optim_wrapper = dict(
+ paramwise_cfg=dict(custom_keys=custom_keys, norm_decay_mult=0.0))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco-panoptic.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco-panoptic.py
new file mode 100644
index 0000000000000000000000000000000000000000..1c00d7a697f07ad618a0b4735432a0a74d4992a9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco-panoptic.py
@@ -0,0 +1,58 @@
+_base_ = ['./mask2former_r50_8xb2-lsj-50e_coco-panoptic.py']
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_tiny_patch4_window7_224.pth' # noqa
+
+depths = [2, 2, 6, 2]
+model = dict(
+ type='Mask2Former',
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ embed_dims=96,
+ depths=depths,
+ num_heads=[3, 6, 12, 24],
+ window_size=7,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(0, 1, 2, 3),
+ with_cp=False,
+ convert_weights=True,
+ frozen_stages=-1,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ panoptic_head=dict(
+ type='Mask2FormerHead', in_channels=[96, 192, 384, 768]),
+ init_cfg=None)
+
+# set all layers in backbone to lr_mult=0.1
+# set all norm layers, position_embeding,
+# query_embeding, level_embeding to decay_multi=0.0
+backbone_norm_multi = dict(lr_mult=0.1, decay_mult=0.0)
+backbone_embed_multi = dict(lr_mult=0.1, decay_mult=0.0)
+embed_multi = dict(lr_mult=1.0, decay_mult=0.0)
+custom_keys = {
+ 'backbone': dict(lr_mult=0.1, decay_mult=1.0),
+ 'backbone.patch_embed.norm': backbone_norm_multi,
+ 'backbone.norm': backbone_norm_multi,
+ 'absolute_pos_embed': backbone_embed_multi,
+ 'relative_position_bias_table': backbone_embed_multi,
+ 'query_embed': embed_multi,
+ 'query_feat': embed_multi,
+ 'level_embed': embed_multi
+}
+custom_keys.update({
+ f'backbone.stages.{stage_id}.blocks.{block_id}.norm': backbone_norm_multi
+ for stage_id, num_blocks in enumerate(depths)
+ for block_id in range(num_blocks)
+})
+custom_keys.update({
+ f'backbone.stages.{stage_id}.downsample.norm': backbone_norm_multi
+ for stage_id in range(len(depths) - 1)
+})
+
+# optimizer
+optim_wrapper = dict(
+ paramwise_cfg=dict(custom_keys=custom_keys, norm_decay_mult=0.0))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5bb9c21858ebe065691a8a963bf5dec85542fb57
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco.py
@@ -0,0 +1,56 @@
+_base_ = ['./mask2former_r50_8xb2-lsj-50e_coco.py']
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_tiny_patch4_window7_224.pth' # noqa
+depths = [2, 2, 6, 2]
+model = dict(
+ type='Mask2Former',
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ embed_dims=96,
+ depths=depths,
+ num_heads=[3, 6, 12, 24],
+ window_size=7,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(0, 1, 2, 3),
+ with_cp=False,
+ convert_weights=True,
+ frozen_stages=-1,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ panoptic_head=dict(
+ type='Mask2FormerHead', in_channels=[96, 192, 384, 768]),
+ init_cfg=None)
+
+# set all layers in backbone to lr_mult=0.1
+# set all norm layers, position_embeding,
+# query_embeding, level_embeding to decay_multi=0.0
+backbone_norm_multi = dict(lr_mult=0.1, decay_mult=0.0)
+backbone_embed_multi = dict(lr_mult=0.1, decay_mult=0.0)
+embed_multi = dict(lr_mult=1.0, decay_mult=0.0)
+custom_keys = {
+ 'backbone': dict(lr_mult=0.1, decay_mult=1.0),
+ 'backbone.patch_embed.norm': backbone_norm_multi,
+ 'backbone.norm': backbone_norm_multi,
+ 'absolute_pos_embed': backbone_embed_multi,
+ 'relative_position_bias_table': backbone_embed_multi,
+ 'query_embed': embed_multi,
+ 'query_feat': embed_multi,
+ 'level_embed': embed_multi
+}
+custom_keys.update({
+ f'backbone.stages.{stage_id}.blocks.{block_id}.norm': backbone_norm_multi
+ for stage_id, num_blocks in enumerate(depths)
+ for block_id in range(num_blocks)
+})
+custom_keys.update({
+ f'backbone.stages.{stage_id}.downsample.norm': backbone_norm_multi
+ for stage_id in range(len(depths) - 1)
+})
+# optimizer
+optim_wrapper = dict(
+ paramwise_cfg=dict(custom_keys=custom_keys, norm_decay_mult=0.0))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..3321239213f7345084b63b77cf02b0525a534585
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former/metafile.yml
@@ -0,0 +1,223 @@
+Collections:
+ - Name: Mask2Former
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - AdamW
+ - Weight Decay
+ Training Resources: 8x A100 GPUs
+ Architecture:
+ - Mask2Former
+ Paper:
+ URL: https://arxiv.org/pdf/2112.01527
+ Title: 'Masked-attention Mask Transformer for Universal Image Segmentation'
+ README: configs/mask2former/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.23.0/mmdet/models/detectors/mask2former.py#L7
+ Version: v2.23.0
+
+Models:
+- Name: mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco-panoptic
+ In Collection: Mask2Former
+ Config: configs/mask2former/mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco-panoptic.py
+ Metadata:
+ Training Memory (GB): 19.1
+ Iterations: 368750
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 47.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 44.5
+ - Task: Panoptic Segmentation
+ Dataset: COCO
+ Metrics:
+ PQ: 54.5
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco-panoptic/mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco-panoptic_20220329_225200-4a16ded7.pth
+- Name: mask2former_r101_8xb2-lsj-50e_coco
+ In Collection: Mask2Former
+ Config: configs/mask2former/mask2former_r101_8xb2-lsj-50e_coco.py
+ Metadata:
+ Training Memory (GB): 15.5
+ Iterations: 368750
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.7
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 44.0
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_r101_8xb2-lsj-50e_coco/mask2former_r101_8xb2-lsj-50e_coco_20220426_100250-ecf181e2.pth
+- Name: mask2former_r101_8xb2-lsj-50e_coco-panoptic
+ In Collection: Mask2Former
+ Config: configs/mask2former/mask2former_r101_8xb2-lsj-50e_coco-panoptic.py
+ Metadata:
+ Training Memory (GB): 16.1
+ Iterations: 368750
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.3
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 42.4
+ - Task: Panoptic Segmentation
+ Dataset: COCO
+ Metrics:
+ PQ: 52.4
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_r101_8xb2-lsj-50e_coco-panoptic/mask2former_r101_8xb2-lsj-50e_coco-panoptic_20220329_225104-c74d4d71.pth
+- Name: mask2former_r50_8xb2-lsj-50e_coco-panoptic
+ In Collection: Mask2Former
+ Config: configs/mask2former/mask2former_r50_8xb2-lsj-50e_coco-panoptic.py
+ Metadata:
+ Training Memory (GB): 13.9
+ Iterations: 368750
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 41.8
+ - Task: Panoptic Segmentation
+ Dataset: COCO
+ Metrics:
+ PQ: 52.0
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_r50_8xb2-lsj-50e_coco-panoptic/mask2former_r50_8xb2-lsj-50e_coco-panoptic_20230118_125535-54df384a.pth
+- Name: mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco-panoptic
+ In Collection: Mask2Former
+ Config: configs/mask2former/mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco-panoptic.py
+ Metadata:
+ Training Memory (GB): 15.9
+ Iterations: 368750
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.3
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 43.4
+ - Task: Panoptic Segmentation
+ Dataset: COCO
+ Metrics:
+ PQ: 53.4
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco-panoptic/mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco-panoptic_20220326_224553-3ec9e0ae.pth
+- Name: mask2former_r50_8xb2-lsj-50e_coco
+ In Collection: Mask2Former
+ Config: configs/mask2former/mask2former_r50_8xb2-lsj-50e_coco.py
+ Metadata:
+ Training Memory (GB): 13.7
+ Iterations: 368750
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.7
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 42.9
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_r50_8xb2-lsj-50e_coco/mask2former_r50_8xb2-lsj-50e_coco_20220506_191028-41b088b6.pth
+- Name: mask2former_swin-l-p4-w12-384-in21k_16xb1-lsj-100e_coco-panoptic
+ In Collection: Mask2Former
+ Config: configs/mask2former/mask2former_swin-l-p4-w12-384-in21k_16xb1-lsj-100e_coco-panoptic.py
+ Metadata:
+ Training Memory (GB): 21.1
+ Iterations: 737500
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 52.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 48.5
+ - Task: Panoptic Segmentation
+ Dataset: COCO
+ Metrics:
+ PQ: 57.6
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_swin-l-p4-w12-384-in21k_16xb1-lsj-100e_coco-panoptic/mask2former_swin-l-p4-w12-384-in21k_16xb1-lsj-100e_coco-panoptic_20220407_104949-82f8d28d.pth
+- Name: mask2former_swin-b-p4-w12-384-in21k_8xb2-lsj-50e_coco-panoptic
+ In Collection: Mask2Former
+ Config: configs/mask2former/mask2former_swin-b-p4-w12-384-in21k_8xb2-lsj-50e_coco-panoptic.py
+ Metadata:
+ Training Memory (GB): 25.8
+ Iterations: 368750
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 50.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 46.3
+ - Task: Panoptic Segmentation
+ Dataset: COCO
+ Metrics:
+ PQ: 56.3
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_swin-b-p4-w12-384-in21k_8xb2-lsj-50e_coco-panoptic/mask2former_swin-b-p4-w12-384-in21k_8xb2-lsj-50e_coco-panoptic_20220329_230021-05ec7315.pth
+- Name: mask2former_swin-b-p4-w12-384_8xb2-lsj-50e_coco-panoptic
+ In Collection: Mask2Former
+ Config: configs/mask2former/mask2former_swin-b-p4-w12-384_8xb2-lsj-50e_coco-panoptic.py
+ Metadata:
+ Training Memory (GB): 26.0
+ Iterations: 368750
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 48.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 44.9
+ - Task: Panoptic Segmentation
+ Dataset: COCO
+ Metrics:
+ PQ: 55.1
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_swin-b-p4-w12-384_8xb2-lsj-50e_coco-panoptic/mask2former_swin-b-p4-w12-384_8xb2-lsj-50e_coco-panoptic_20220331_002244-8a651d82.pth
+- Name: mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco
+ In Collection: Mask2Former
+ Config: configs/mask2former/mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco.py
+ Metadata:
+ Training Memory (GB): 15.3
+ Iterations: 368750
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 47.7
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 44.7
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco/mask2former_swin-t-p4-w7-224_8xb2-lsj-50e_coco_20220508_091649-01b0f990.pth
+- Name: mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco
+ In Collection: Mask2Former
+ Config: configs/mask2former/mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco.py
+ Metadata:
+ Training Memory (GB): 18.8
+ Iterations: 368750
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 49.3
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 46.1
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mask2former/mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco/mask2former_swin-s-p4-w7-224_8xb2-lsj-50e_coco_20220504_001756-c9d0c4f2.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..699657290896f1d2ccb36ffe60ec6471f68043fd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/README.md
@@ -0,0 +1,81 @@
+# Mask2Former for Video Instance Segmentation
+
+## Abstract
+
+
+
+We find Mask2Former also achieves state-of-the-art performance on video instance segmentation without modifying the architecture, the loss or even the training pipeline. In this report, we show universal image segmentation architectures trivially generalize to video segmentation by directly predicting 3D segmentation volumes. Specifically, Mask2Former sets a new state-of-the-art of 60.4 AP on YouTubeVIS-2019 and 52.6 AP on YouTubeVIS-2021. We believe Mask2Former is also capable of handling video semantic and panoptic segmentation, given its versatility in image segmentation. We hope this will make state-of-theart video segmentation research more accessible and bring more attention to designing universal image and video segmentation architectures.
+
+
+
+
+

+
+
+## Citation
+
+
+
+```latex
+@inproceedings{cheng2021mask2former,
+ title={Masked-attention Mask Transformer for Universal Image Segmentation},
+ author={Bowen Cheng and Ishan Misra and Alexander G. Schwing and Alexander Kirillov and Rohit Girdhar},
+ journal={CVPR},
+ year={2022}
+}
+```
+
+## Results and models of Mask2Former on YouTube-VIS 2021 validation dataset
+
+Note: Codalab has closed the evaluation portal of `YouTube-VIS 2019`, so we do not provide the results of `YouTube-VIS 2019` at present. If you want to evaluate the results of `YouTube-VIS 2021`, at present, you can submit the result to the evaluation portal of `YouTube-VIS 2022`. The value of `AP_S` is the result of `YouTube-VIS 2021`.
+
+| Method | Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | AP | Config | Download |
+| :----------------------: | :------: | :-----: | :-----: | :------: | :------------: | :--: | :---------------------------------------------------------------------: | :-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Mask2Former | R-50 | pytorch | 8e | 6.0 | - | 41.3 | [config](mask2former_r50_8xb2-8e_youtubevis2021.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mask2former_vis/mask2former_r50_8xb2-8e_youtubevis2021/mask2former_r50_8xb2-8e_youtubevis2021_20230426_131833-5d215283.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/mask2former_vis/mask2former_r50_8xb2-8e_youtubevis2021/mask2former_r50_8xb2-8e_youtubevis2021_20230426_131833.json) |
+| Mask2Former | R-101 | pytorch | 8e | 7.5 | - | 42.3 | [config](mask2former_r101_8xb2-8e_youtubevis2021.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mask2former_vis/mask2former_r101_8xb2-8e_youtubevis2021/mask2former_r101_8xb2-8e_youtubevis2021_20220823_092747-8077d115.pth) \| [log](https://download.openmmlab.com/mmtracking/vis/mask2former/mask2former_r101_8xb2-8e_youtubevis2021_20220823_092747.json) |
+| Mask2Former(200 queries) | Swin-L | pytorch | 8e | 18.5 | - | 52.3 | [config](mask2former_swin-l-p4-w12-384-in21k_8xb2-8e_youtubevis2021.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mask2former_vis/mask2former_swin-l-p4-w12-384-in21k_8xb2-8e_youtubevis2021/mask2former_swin-l-p4-w12-384-in21k_8xb2-8e_youtubevis2021_20220907_124752-48252603.pth) \| [log](https://download.openmmlab.com/mmtracking/vis/mask2former/mask2former_swin-l-p4-w12-384-in21k_8xb2-8e_youtubevis2021_20220907_124752.json) |
+
+## Get started
+
+### 1. Development Environment Setup
+
+Tracking Development Environment Setup can refer to this [document](../../docs/en/get_started.md).
+
+### 2. Dataset Prepare
+
+Tracking Dataset Prepare can refer to this [document](../../docs/en/user_guides/tracking_dataset_prepare.md).
+
+### 3. Training
+
+Due to the influence of parameters such as learning rate in default configuration file, we recommend using 8 GPUs for training in order to reproduce accuracy. You can use the following command to start the training.
+
+```shell
+# Training Mask2Former on YouTube-VIS-2021 dataset with following command.
+# The number after config file represents the number of GPUs used. Here we use 8 GPUs.
+bash tools/dist_train.sh configs/mask2former_vis/mask2former_r50_8xb2-8e_youtubevis2021.py 8
+```
+
+If you want to know about more detailed usage of `train.py/dist_train.sh/slurm_train.sh`,
+please refer to this [document](../../docs/en/user_guides/tracking_train_test.md).
+
+### 4. Testing and evaluation
+
+If you want to get the results of the [YouTube-VOS](https://youtube-vos.org/dataset/vis/) val/test set, please use the following command to generate result files that can be used for submission. It will be stored in `./youtube_vis_results.submission_file.zip`, you can modify the saved path in `test_evaluator` of the config.
+
+```shell
+# The number after config file represents the number of GPUs used.
+bash tools/dist_test_tracking.sh configs/mask2former_vis/mask2former_r50_8xb2-8e_youtubevis2021.py --checkpoint ${CHECKPOINT_PATH}
+```
+
+If you want to know about more detailed usage of `test_tracking.py/dist_test_tracking.sh/slurm_test_tracking.sh`,
+please refer to this [document](../../docs/en/user_guides/tracking_train_test.md).
+
+### 5.Inference
+
+Use a single GPU to predict a video and save it as a video.
+
+```shell
+python demo/mot_demo.py demo/demo_mot.mp4 configs/mask2former_vis/mask2former_r50_8xb2-8e_youtubevis2021.py --checkpoint {CHECKPOINT_PATH} --out vis.mp4
+```
+
+If you want to know about more detailed usage of `mot_demo.py`, please refer to this [document](../../docs/en/user_guides/tracking_inference.md).
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/mask2former_r101_8xb2-8e_youtubevis2019.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/mask2former_r101_8xb2-8e_youtubevis2019.py
new file mode 100644
index 0000000000000000000000000000000000000000..3ba4aea8eac72f347940fb12ac964e9bf67c2e0e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/mask2former_r101_8xb2-8e_youtubevis2019.py
@@ -0,0 +1,12 @@
+_base_ = './mask2former_r50_8xb2-8e_youtubevis2019.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')),
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='https://download.openmmlab.com/mmdetection/v3.0/'
+ 'mask2former/mask2former_r101_8xb2-lsj-50e_coco/'
+ 'mask2former_r101_8xb2-lsj-50e_coco_20220426_100250-ecf181e2.pth'))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/mask2former_r101_8xb2-8e_youtubevis2021.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/mask2former_r101_8xb2-8e_youtubevis2021.py
new file mode 100644
index 0000000000000000000000000000000000000000..95f9ceeb38833aeef342e12178703db6901fe5f6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/mask2former_r101_8xb2-8e_youtubevis2021.py
@@ -0,0 +1,12 @@
+_base_ = './mask2former_r50_8xb2-8e_youtubevis2021.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')),
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='https://download.openmmlab.com/mmdetection/v3.0/'
+ 'mask2former/mask2former_r101_8xb2-lsj-50e_coco/'
+ 'mask2former_r101_8xb2-lsj-50e_coco_20220426_100250-ecf181e2.pth'))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/mask2former_r50_8xb2-8e_youtubevis2019.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/mask2former_r50_8xb2-8e_youtubevis2019.py
new file mode 100644
index 0000000000000000000000000000000000000000..8dc03bf97a2ed2b90e097bbd9637a42bf4d64c35
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/mask2former_r50_8xb2-8e_youtubevis2019.py
@@ -0,0 +1,174 @@
+_base_ = ['../_base_/datasets/youtube_vis.py', '../_base_/default_runtime.py']
+
+num_classes = 40
+num_frames = 2
+model = dict(
+ type='Mask2FormerVideo',
+ data_preprocessor=dict(
+ type='TrackDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_mask=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=-1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ track_head=dict(
+ type='Mask2FormerTrackHead',
+ in_channels=[256, 512, 1024, 2048], # pass to pixel_decoder inside
+ strides=[4, 8, 16, 32],
+ feat_channels=256,
+ out_channels=256,
+ num_classes=num_classes,
+ num_queries=100,
+ num_frames=num_frames,
+ num_transformer_feat_level=3,
+ pixel_decoder=dict(
+ type='MSDeformAttnPixelDecoder',
+ num_outs=3,
+ norm_cfg=dict(type='GN', num_groups=32),
+ act_cfg=dict(type='ReLU'),
+ encoder=dict( # DeformableDetrTransformerEncoder
+ num_layers=6,
+ layer_cfg=dict( # DeformableDetrTransformerEncoderLayer
+ self_attn_cfg=dict( # MultiScaleDeformableAttention
+ embed_dims=256,
+ num_heads=8,
+ num_levels=3,
+ num_points=4,
+ im2col_step=128,
+ dropout=0.0,
+ batch_first=True),
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=1024,
+ num_fcs=2,
+ ffn_drop=0.0,
+ act_cfg=dict(type='ReLU', inplace=True)))),
+ positional_encoding=dict(num_feats=128, normalize=True)),
+ enforce_decoder_input_project=False,
+ positional_encoding=dict(
+ type='SinePositionalEncoding3D', num_feats=128, normalize=True),
+ transformer_decoder=dict( # Mask2FormerTransformerDecoder
+ return_intermediate=True,
+ num_layers=9,
+ layer_cfg=dict( # Mask2FormerTransformerDecoderLayer
+ self_attn_cfg=dict( # MultiheadAttention
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.0,
+ batch_first=True),
+ cross_attn_cfg=dict( # MultiheadAttention
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.0,
+ batch_first=True),
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048,
+ num_fcs=2,
+ ffn_drop=0.0,
+ act_cfg=dict(type='ReLU', inplace=True))),
+ init_cfg=None),
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=2.0,
+ reduction='mean',
+ class_weight=[1.0] * num_classes + [0.1]),
+ loss_mask=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ reduction='mean',
+ loss_weight=5.0),
+ loss_dice=dict(
+ type='DiceLoss',
+ use_sigmoid=True,
+ activate=True,
+ reduction='mean',
+ naive_dice=True,
+ eps=1.0,
+ loss_weight=5.0),
+ train_cfg=dict(
+ num_points=12544,
+ oversample_ratio=3.0,
+ importance_sample_ratio=0.75,
+ assigner=dict(
+ type='HungarianAssigner',
+ match_costs=[
+ dict(type='ClassificationCost', weight=2.0),
+ dict(
+ type='CrossEntropyLossCost',
+ weight=5.0,
+ use_sigmoid=True),
+ dict(type='DiceCost', weight=5.0, pred_act=True, eps=1.0)
+ ]),
+ sampler=dict(type='MaskPseudoSampler'))),
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='https://download.openmmlab.com/mmdetection/v3.0/'
+ 'mask2former/mask2former_r50_8xb2-lsj-50e_coco/'
+ 'mask2former_r50_8xb2-lsj-50e_coco_20220506_191028-41b088b6.pth'))
+
+# optimizer
+embed_multi = dict(lr_mult=1.0, decay_mult=0.0)
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(
+ type='AdamW',
+ lr=0.0001,
+ weight_decay=0.05,
+ eps=1e-8,
+ betas=(0.9, 0.999)),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'backbone': dict(lr_mult=0.1, decay_mult=1.0),
+ 'query_embed': embed_multi,
+ 'query_feat': embed_multi,
+ 'level_embed': embed_multi,
+ },
+ norm_decay_mult=0.0),
+ clip_grad=dict(max_norm=0.01, norm_type=2))
+
+# learning policy
+max_iters = 6000
+param_scheduler = dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_iters,
+ by_epoch=False,
+ milestones=[
+ 4000,
+ ],
+ gamma=0.1)
+# runtime settings
+train_cfg = dict(
+ type='IterBasedTrainLoop', max_iters=max_iters, val_interval=6001)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+vis_backends = [dict(type='LocalVisBackend')]
+visualizer = dict(
+ type='TrackLocalVisualizer', vis_backends=vis_backends, name='visualizer')
+
+default_hooks = dict(
+ checkpoint=dict(
+ type='CheckpointHook', by_epoch=False, save_last=True, interval=2000),
+ visualization=dict(type='TrackVisualizationHook', draw=False))
+log_processor = dict(type='LogProcessor', window_size=50, by_epoch=False)
+
+# evaluator
+val_evaluator = dict(
+ type='YouTubeVISMetric',
+ metric='youtube_vis_ap',
+ outfile_prefix='./youtube_vis_results',
+ format_only=True)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/mask2former_r50_8xb2-8e_youtubevis2021.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/mask2former_r50_8xb2-8e_youtubevis2021.py
new file mode 100644
index 0000000000000000000000000000000000000000..158fe52d20fccf162cb66202fbc9069ba0f4cb68
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/mask2former_r50_8xb2-8e_youtubevis2021.py
@@ -0,0 +1,37 @@
+_base_ = './mask2former_r50_8xb2-8e_youtubevis2019.py'
+
+dataset_type = 'YouTubeVISDataset'
+data_root = 'data/youtube_vis_2021/'
+dataset_version = data_root[-5:-1] # 2019 or 2021
+
+train_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ dataset_version=dataset_version,
+ ann_file='annotations/youtube_vis_2021_train.json'))
+
+val_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ dataset_version=dataset_version,
+ ann_file='annotations/youtube_vis_2021_valid.json'))
+test_dataloader = val_dataloader
+
+# learning policy
+max_iters = 8000
+param_scheduler = dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_iters,
+ by_epoch=False,
+ milestones=[
+ 5500,
+ ],
+ gamma=0.1)
+# runtime settings
+train_cfg = dict(
+ type='IterBasedTrainLoop', max_iters=max_iters, val_interval=8001)
+
+default_hooks = dict(
+ checkpoint=dict(
+ type='CheckpointHook', by_epoch=False, save_last=True, interval=500))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/mask2former_swin-l-p4-w12-384-in21k_8xb2-8e_youtubevis2021.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/mask2former_swin-l-p4-w12-384-in21k_8xb2-8e_youtubevis2021.py
new file mode 100644
index 0000000000000000000000000000000000000000..94dcccf408dfb989ea264536a617a48ecc13171c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/mask2former_swin-l-p4-w12-384-in21k_8xb2-8e_youtubevis2021.py
@@ -0,0 +1,64 @@
+_base_ = ['./mask2former_r50_8xb2-8e_youtubevis2021.py']
+depths = [2, 2, 18, 2]
+model = dict(
+ type='Mask2FormerVideo',
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ pretrain_img_size=384,
+ embed_dims=192,
+ depths=depths,
+ num_heads=[6, 12, 24, 48],
+ window_size=12,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(0, 1, 2, 3),
+ with_cp=False,
+ convert_weights=True,
+ frozen_stages=-1,
+ init_cfg=None),
+ track_head=dict(
+ type='Mask2FormerTrackHead',
+ in_channels=[192, 384, 768, 1536],
+ num_queries=200),
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint= # noqa: E251
+ 'https://download.openmmlab.com/mmdetection/v3.0/mask2former/'
+ 'mask2former_swin-l-p4-w12-384-in21k_16xb1-lsj-100e_coco-panoptic/'
+ 'mask2former_swin-l-p4-w12-384-in21k_16xb1-lsj-100e_coco-panoptic_'
+ '20220407_104949-82f8d28d.pth'))
+
+# set all layers in backbone to lr_mult=0.1
+# set all norm layers, position_embeding,
+# query_embeding, level_embeding to decay_multi=0.0
+backbone_norm_multi = dict(lr_mult=0.1, decay_mult=0.0)
+backbone_embed_multi = dict(lr_mult=0.1, decay_mult=0.0)
+embed_multi = dict(lr_mult=1.0, decay_mult=0.0)
+custom_keys = {
+ 'backbone': dict(lr_mult=0.1, decay_mult=1.0),
+ 'backbone.patch_embed.norm': backbone_norm_multi,
+ 'backbone.norm': backbone_norm_multi,
+ 'absolute_pos_embed': backbone_embed_multi,
+ 'relative_position_bias_table': backbone_embed_multi,
+ 'query_embed': embed_multi,
+ 'query_feat': embed_multi,
+ 'level_embed': embed_multi
+}
+custom_keys.update({
+ f'backbone.stages.{stage_id}.blocks.{block_id}.norm': backbone_norm_multi
+ for stage_id, num_blocks in enumerate(depths)
+ for block_id in range(num_blocks)
+})
+custom_keys.update({
+ f'backbone.stages.{stage_id}.downsample.norm': backbone_norm_multi
+ for stage_id in range(len(depths) - 1)
+})
+# optimizer
+optim_wrapper = dict(
+ paramwise_cfg=dict(custom_keys=custom_keys, norm_decay_mult=0.0))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..f5f4bd7c5775820f283a7544bf5978fe0aa1abc5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask2former_vis/metafile.yml
@@ -0,0 +1,53 @@
+Collections:
+ - Name: Mask2Former
+ Metadata:
+ Training Techniques:
+ - AdamW
+ - Weight Decay
+ Training Resources: 8x A100 GPUs
+ Architecture:
+ - Mask2Former
+ Paper:
+ URL: https://arxiv.org/pdf/2112.10764.pdf
+ Title: Mask2Former for Video Instance Segmentation
+ README: configs/mask2former/README.md
+
+Models:
+ - Name: mask2former_r50_8xb2-8e_youtubevis2021
+ In Collection: Mask2Former
+ Config: configs/mask2former_vis/mask2former_r50_8xb2-8e_youtubevis2021.py
+ Metadata:
+ Training Data: YouTube-VIS 2021
+ Training Memory (GB): 6.0
+ Results:
+ - Task: Video Instance Segmentation
+ Dataset: YouTube-VIS 2021
+ Metrics:
+ AP: 41.3
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mask2former_vis/mask2former_r50_8xb2-8e_youtubevis2021/mask2former_r50_8xb2-8e_youtubevis2021_20230426_131833-5d215283.pth
+
+ - Name: mask2former_r101_8xb2-8e_youtubevis2021
+ In Collection: Mask2Former
+ Config: configs/mask2former_vis/mask2former_r101_8xb2-8e_youtubevis2021.py
+ Metadata:
+ Training Data: YouTube-VIS 2021
+ Training Memory (GB): 7.5
+ Results:
+ - Task: Video Instance Segmentation
+ Dataset: YouTube-VIS 2021
+ Metrics:
+ AP: 42.3
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mask2former_vis/mask2former_r101_8xb2-8e_youtubevis2021/mask2former_r101_8xb2-8e_youtubevis2021_20220823_092747-8077d115.pth
+
+ - Name: mask2former_swin-l-p4-w12-384-in21k_8xb2-8e_youtubevis2021.py
+ In Collection: Mask2Former
+ Config: configs/mask2former_vis/mask2former_swin-l-p4-w12-384-in21k_8xb2-8e_youtubevis2021.py
+ Metadata:
+ Training Data: YouTube-VIS 2021
+ Training Memory (GB): 18.5
+ Results:
+ - Task: Video Instance Segmentation
+ Dataset: YouTube-VIS 2021
+ Metrics:
+ AP: 52.3
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mask2former_vis/mask2former_swin-l-p4-w12-384-in21k_8xb2-8e_youtubevis2021/mask2former_swin-l-p4-w12-384-in21k_8xb2-8e_youtubevis2021_20220907_124752-48252603.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..afc5c3c92c683947ca01ad05456b0d7ff77be5e9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/README.md
@@ -0,0 +1,59 @@
+# Mask R-CNN
+
+> [Mask R-CNN](https://arxiv.org/abs/1703.06870)
+
+
+
+## Abstract
+
+We present a conceptually simple, flexible, and general framework for object instance segmentation. Our approach efficiently detects objects in an image while simultaneously generating a high-quality segmentation mask for each instance. The method, called Mask R-CNN, extends Faster R-CNN by adding a branch for predicting an object mask in parallel with the existing branch for bounding box recognition. Mask R-CNN is simple to train and adds only a small overhead to Faster R-CNN, running at 5 fps. Moreover, Mask R-CNN is easy to generalize to other tasks, e.g., allowing us to estimate human poses in the same framework. We show top results in all three tracks of the COCO suite of challenges, including instance segmentation, bounding-box object detection, and person keypoint detection. Without bells and whistles, Mask R-CNN outperforms all existing, single-model entries on every task, including the COCO 2016 challenge winners. We hope our simple and effective approach will serve as a solid baseline and help ease future research in instance-level recognition.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :-------------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :---------------------------------------------: | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | caffe | 1x | 4.3 | | 38.0 | 34.4 | [config](./mask-rcnn_r50-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_caffe_fpn_1x_coco/mask_rcnn_r50_caffe_fpn_1x_coco_bbox_mAP-0.38__segm_mAP-0.344_20200504_231812-0ebd1859.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_caffe_fpn_1x_coco/mask_rcnn_r50_caffe_fpn_1x_coco_20200504_231812.log.json) |
+| R-50-FPN | pytorch | 1x | 4.4 | 16.1 | 38.2 | 34.7 | [config](./mask-rcnn_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_fpn_1x_coco/mask_rcnn_r50_fpn_1x_coco_20200205-d4b0c5d6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_fpn_1x_coco/mask_rcnn_r50_fpn_1x_coco_20200205_050542.log.json) |
+| R-50-FPN (FP16) | pytorch | 1x | 3.6 | 24.1 | 38.1 | 34.7 | [config](./mask-rcnn_r50_fpn_amp-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fp16/mask_rcnn_r50_fpn_fp16_1x_coco/mask_rcnn_r50_fpn_fp16_1x_coco_20200205-59faf7e4.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fp16/mask_rcnn_r50_fpn_fp16_1x_coco/mask_rcnn_r50_fpn_fp16_1x_coco_20200205_130539.log.json) |
+| R-50-FPN | pytorch | 2x | - | - | 39.2 | 35.4 | [config](./mask-rcnn_r50_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_fpn_2x_coco/mask_rcnn_r50_fpn_2x_coco_bbox_mAP-0.392__segm_mAP-0.354_20200505_003907-3e542a40.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_fpn_2x_coco/mask_rcnn_r50_fpn_2x_coco_20200505_003907.log.json) |
+| R-101-FPN | caffe | 1x | | | 40.4 | 36.4 | [config](./mask-rcnn_r101-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r101_caffe_fpn_1x_coco/mask_rcnn_r101_caffe_fpn_1x_coco_20200601_095758-805e06c1.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r101_caffe_fpn_1x_coco/mask_rcnn_r101_caffe_fpn_1x_coco_20200601_095758.log.json) |
+| R-101-FPN | pytorch | 1x | 6.4 | 13.5 | 40.0 | 36.1 | [config](./mask-rcnn_r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r101_fpn_1x_coco/mask_rcnn_r101_fpn_1x_coco_20200204-1efe0ed5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r101_fpn_1x_coco/mask_rcnn_r101_fpn_1x_coco_20200204_144809.log.json) |
+| R-101-FPN | pytorch | 2x | - | - | 40.8 | 36.6 | [config](./mask-rcnn_r101_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r101_fpn_2x_coco/mask_rcnn_r101_fpn_2x_coco_bbox_mAP-0.408__segm_mAP-0.366_20200505_071027-14b391c7.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r101_fpn_2x_coco/mask_rcnn_r101_fpn_2x_coco_20200505_071027.log.json) |
+| X-101-32x4d-FPN | pytorch | 1x | 7.6 | 11.3 | 41.9 | 37.5 | [config](./mask-rcnn_x101-32x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x4d_fpn_1x_coco/mask_rcnn_x101_32x4d_fpn_1x_coco_20200205-478d0b67.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x4d_fpn_1x_coco/mask_rcnn_x101_32x4d_fpn_1x_coco_20200205_034906.log.json) |
+| X-101-32x4d-FPN | pytorch | 2x | - | - | 42.2 | 37.8 | [config](./mask-rcnn_x101-32x4d_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x4d_fpn_2x_coco/mask_rcnn_x101_32x4d_fpn_2x_coco_bbox_mAP-0.422__segm_mAP-0.378_20200506_004702-faef898c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x4d_fpn_2x_coco/mask_rcnn_x101_32x4d_fpn_2x_coco_20200506_004702.log.json) |
+| X-101-64x4d-FPN | pytorch | 1x | 10.7 | 8.0 | 42.8 | 38.4 | [config](./mask-rcnn_x101-64x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_64x4d_fpn_1x_coco/mask_rcnn_x101_64x4d_fpn_1x_coco_20200201-9352eb0d.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_64x4d_fpn_1x_coco/mask_rcnn_x101_64x4d_fpn_1x_coco_20200201_124310.log.json) |
+| X-101-64x4d-FPN | pytorch | 2x | - | - | 42.7 | 38.1 | [config](./mask-rcnn_x101-64x4d_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_64x4d_fpn_2x_coco/mask_rcnn_x101_64x4d_fpn_2x_coco_20200509_224208-39d6f70c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_64x4d_fpn_2x_coco/mask_rcnn_x101_64x4d_fpn_2x_coco_20200509_224208.log.json) |
+| X-101-32x8d-FPN | pytorch | 1x | 10.6 | - | 42.8 | 38.3 | [config](./mask-rcnn_x101-32x8d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x8d_fpn_1x_coco/mask_rcnn_x101_32x8d_fpn_1x_coco_20220630_173841-0aaf329e.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x8d_fpn_1x_coco/mask_rcnn_x101_32x8d_fpn_1x_coco_20220630_173841.log.json) |
+
+## Pre-trained Models
+
+We also train some models with longer schedules and multi-scale training. The users could finetune them for downstream tasks.
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :--------------------------------------------------------------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :-----------------------------------------------------: | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| [R-50-FPN](./mask-rcnn_r50-caffe_fpn_ms-poly-2x_coco.py) | caffe | 2x | 4.3 | | 40.3 | 36.5 | [config](./mask-rcnn_r50-caffe_fpn_ms-poly-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_caffe_fpn_mstrain-poly_2x_coco/mask_rcnn_r50_caffe_fpn_mstrain-poly_2x_coco_bbox_mAP-0.403__segm_mAP-0.365_20200504_231822-a75c98ce.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_caffe_fpn_mstrain-poly_2x_coco/mask_rcnn_r50_caffe_fpn_mstrain-poly_2x_coco_20200504_231822.log.json) |
+| [R-50-FPN](./mask-rcnn_r50-caffe_fpn_ms-poly-3x_coco.py) | caffe | 3x | 4.3 | | 40.8 | 37.0 | [config](./mask-rcnn_r50-caffe_fpn_ms-poly-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_caffe_fpn_mstrain-poly_3x_coco/mask_rcnn_r50_caffe_fpn_mstrain-poly_3x_coco_bbox_mAP-0.408__segm_mAP-0.37_20200504_163245-42aa3d00.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_caffe_fpn_mstrain-poly_3x_coco/mask_rcnn_r50_caffe_fpn_mstrain-poly_3x_coco_20200504_163245.log.json) |
+| [R-50-FPN](./mask-rcnn_r50_fpn_ms-poly-3x_coco.py) | pytorch | 3x | 4.1 | | 40.9 | 37.1 | [config](./mask-rcnn_r50_fpn_ms-poly-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_fpn_mstrain-poly_3x_coco/mask_rcnn_r50_fpn_mstrain-poly_3x_coco_20210524_201154-21b550bb.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_fpn_mstrain-poly_3x_coco/mask_rcnn_r50_fpn_mstrain-poly_3x_coco_20210524_201154.log.json) |
+| [R-101-FPN](./mask-rcnn_r101-caffe_fpn_ms-poly-3x_coco.py) | caffe | 3x | 5.9 | | 42.9 | 38.5 | [config](./mask-rcnn_r101-caffe_fpn_ms-poly-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r101_caffe_fpn_mstrain-poly_3x_coco/mask_rcnn_r101_caffe_fpn_mstrain-poly_3x_coco_20210526_132339-3c33ce02.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn_r101_caffe_fpn_mstrain-poly_3x_coco/mask_rcnn_r101_caffe_fpn_mstrain-poly_3x_coco_20210526_132339.log.json) |
+| [R-101-FPN](./mask-rcnn_r101_fpn_ms-poly-3x_coco.py) | pytorch | 3x | 6.1 | | 42.7 | 38.5 | [config](./mask-rcnn_r101_fpn_ms-poly-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r101_fpn_mstrain-poly_3x_coco/mask_rcnn_r101_fpn_mstrain-poly_3x_coco_20210524_200244-5675c317.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r101_fpn_mstrain-poly_3x_coco/mask_rcnn_r101_fpn_mstrain-poly_3x_coco_20210524_200244.log.json) |
+| [x101-32x4d-FPN](./mask-rcnn_x101-32x4d_fpn_ms-poly-3x_coco.py) | pytorch | 3x | 7.3 | | 43.6 | 39.0 | [config](./mask-rcnn_x101-32x4d_fpn_ms-poly-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x4d_fpn_mstrain-poly_3x_coco/mask_rcnn_x101_32x4d_fpn_mstrain-poly_3x_coco_20210524_201410-abcd7859.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x4d_fpn_mstrain-poly_3x_coco/mask_rcnn_x101_32x4d_fpn_mstrain-poly_3x_coco_20210524_201410.log.json) |
+| [X-101-32x8d-FPN](./mask-rcnn_x101-32x8d_fpn_ms-poly-3x_coco.py) | pytorch | 1x | 10.4 | | 43.4 | 39.0 | [config](./mask-rcnn_x101-32x8d_fpn_ms-poly-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x8d_fpn_mstrain-poly_1x_coco/mask_rcnn_x101_32x8d_fpn_mstrain-poly_1x_coco_20220630_170346-b4637974.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x8d_fpn_mstrain-poly_1x_coco/mask_rcnn_x101_32x8d_fpn_mstrain-poly_1x_coco_20220630_170346.log.json) |
+| [X-101-32x8d-FPN](./mask-rcnn_x101-32x8d_fpn_ms-poly-3x_coco.py) | pytorch | 3x | 10.3 | | 44.3 | 39.5 | [config](./mask-rcnn_x101-32x8d_fpn_ms-poly-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x8d_fpn_mstrain-poly_3x_coco/mask_rcnn_x101_32x8d_fpn_mstrain-poly_3x_coco_20210607_161042-8bd2c639.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x8d_fpn_mstrain-poly_3x_coco/mask_rcnn_x101_32x8d_fpn_mstrain-poly_3x_coco_20210607_161042.log.json) |
+| [X-101-64x4d-FPN](./mask-rcnn_x101-64x4d_fpn_ms-poly_3x_coco.py) | pytorch | 3x | 10.4 | | 44.5 | 39.7 | [config](./mask-rcnn_x101-64x4d_fpn_ms-poly_3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_64x4d_fpn_mstrain-poly_3x_coco/mask_rcnn_x101_64x4d_fpn_mstrain-poly_3x_coco_20210526_120447-c376f129.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_64x4d_fpn_mstrain-poly_3x_coco/mask_rcnn_x101_64x4d_fpn_mstrain-poly_3x_coco_20210526_120447.log.json) |
+
+## Citation
+
+```latex
+@article{He_2017,
+ title={Mask R-CNN},
+ journal={2017 IEEE International Conference on Computer Vision (ICCV)},
+ publisher={IEEE},
+ author={He, Kaiming and Gkioxari, Georgia and Dollar, Piotr and Girshick, Ross},
+ year={2017},
+ month={Oct}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r101-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r101-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..09808e4bcada43b1e935d5393894c7ba3401fc3d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r101-caffe_fpn_1x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './mask-rcnn_r50-caffe_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet101_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r101-caffe_fpn_ms-poly-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r101-caffe_fpn_ms-poly-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e723aea81ff82dfa842d7468e166f42ee9291669
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r101-caffe_fpn_ms-poly-3x_coco.py
@@ -0,0 +1,19 @@
+_base_ = [
+ '../common/ms-poly_3x_coco-instance.py',
+ '../_base_/models/mask-rcnn_r50_fpn.py'
+]
+
+model = dict(
+ # use caffe img_norm
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False),
+ backbone=dict(
+ depth=101,
+ norm_cfg=dict(requires_grad=False),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet101_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..af91ff0b8349b0e9e658b69cf4c5dd138b7b8a5a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r101_fpn_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './mask-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r101_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r101_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a5599e7c4942b523d6500e2c7c8ad4638cab45c6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r101_fpn_2x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './mask-rcnn_r50_fpn_2x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r101_fpn_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r101_fpn_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..452351050238a4d4411b2bf6fc916e2d69804766
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r101_fpn_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './mask-rcnn_r50_fpn_8xb8-amp-lsj-200e_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r101_fpn_ms-poly-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r101_fpn_ms-poly-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..384f6dcd3ca33cd91755b48dd525d747a358ee02
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r101_fpn_ms-poly-3x_coco.py
@@ -0,0 +1,10 @@
+_base_ = [
+ '../common/ms-poly_3x_coco-instance.py',
+ '../_base_/models/mask-rcnn_r50_fpn.py'
+]
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r18_fpn_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r18_fpn_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5b9219c9c1da8ca68cf7ada0881419b371a26a87
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r18_fpn_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './mask-rcnn_r50_fpn_8xb8-amp-lsj-200e_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=18,
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet18')),
+ neck=dict(in_channels=[64, 128, 256, 512]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe-c4_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe-c4_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..9919f11c3fc7b68528bf6f690e39185d703aff43
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe-c4_1x_coco.py
@@ -0,0 +1,5 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50-caffe-c4.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..4124f138d874def6810cea6c884a02eaacdf5f71
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_1x_coco.py
@@ -0,0 +1,13 @@
+_base_ = './mask-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ # use caffe img_norm
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False),
+ backbone=dict(
+ norm_cfg=dict(requires_grad=False),
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_ms-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_ms-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..7702ae14a9cc54686df6a3eadec5bc8cfeb8e0a8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_ms-1x_coco.py
@@ -0,0 +1,28 @@
+_base_ = './mask-rcnn_r50_fpn_1x_coco.py'
+
+model = dict(
+ # use caffe img_norm
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False),
+ backbone=dict(
+ norm_cfg=dict(requires_grad=False),
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs'),
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_ms-poly-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_ms-poly-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..94d94dd3613e0599f51f113ccf12e568a5b29f8f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_ms-poly-1x_coco.py
@@ -0,0 +1,31 @@
+_base_ = './mask-rcnn_r50_fpn_1x_coco.py'
+
+model = dict(
+ # use caffe img_norm
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False),
+ backbone=dict(
+ norm_cfg=dict(requires_grad=False),
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')))
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ poly2mask=False),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_ms-poly-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_ms-poly-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..dbf87bb8346dd351c8f16700df7b9640bcfa984a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_ms-poly-2x_coco.py
@@ -0,0 +1,15 @@
+_base_ = './mask-rcnn_r50-caffe_fpn_ms-poly-1x_coco.py'
+
+train_cfg = dict(max_epochs=24)
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=24,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_ms-poly-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_ms-poly-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..45260e2e39b53c0107e257ef2d05a14f5d5c0323
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_ms-poly-3x_coco.py
@@ -0,0 +1,15 @@
+_base_ = './mask-rcnn_r50-caffe_fpn_ms-poly-1x_coco.py'
+
+train_cfg = dict(max_epochs=36)
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=24,
+ by_epoch=True,
+ milestones=[28, 34],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_poly-1x_coco_v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_poly-1x_coco_v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..3baf00140ecfa57ea54b68b85ac826e14490daa4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_poly-1x_coco_v1.py
@@ -0,0 +1,31 @@
+_base_ = './mask-rcnn_r50_fpn_1x_coco.py'
+
+model = dict(
+ # use caffe img_norm
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False),
+ backbone=dict(
+ norm_cfg=dict(requires_grad=False),
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')),
+ rpn_head=dict(
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0 / 9.0, loss_weight=1.0)),
+ roi_head=dict(
+ bbox_roi_extractor=dict(
+ roi_layer=dict(
+ type='RoIAlign',
+ output_size=7,
+ sampling_ratio=2,
+ aligned=False)),
+ bbox_head=dict(
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0)),
+ mask_roi_extractor=dict(
+ roi_layer=dict(
+ type='RoIAlign',
+ output_size=14,
+ sampling_ratio=2,
+ aligned=False))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_1x-wandb_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_1x-wandb_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..28b125ccb94869aff2bb283e6533fd693c79a76e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_1x-wandb_coco.py
@@ -0,0 +1,16 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+vis_backends = [dict(type='LocalVisBackend'), dict(type='WandbVisBackend')]
+visualizer = dict(vis_backends=vis_backends)
+
+# MMEngine support the following two ways, users can choose
+# according to convenience
+# default_hooks = dict(checkpoint=dict(interval=4))
+_base_.default_hooks.checkpoint.interval = 4
+
+# train_cfg = dict(val_interval=2)
+_base_.train_cfg.val_interval = 2
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..0fc6b91aa895e044b3fc62a3cdedbc12a052e91b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py
@@ -0,0 +1,5 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..87cb8b4bb7d2fbfcfe667e7bd6cfc08e01e28c1a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_2x_coco.py
@@ -0,0 +1,5 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_2x.py', '../_base_/default_runtime.py'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..7371b3646fdda7bdc1fcfcd44cf8a20df27c40b5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,22 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../common/lsj-100e_coco-instance.py'
+]
+image_size = (1024, 1024)
+batch_augments = [
+ dict(type='BatchFixedSizePad', size=image_size, pad_mask=True)
+]
+
+model = dict(data_preprocessor=dict(batch_augments=batch_augments))
+
+train_dataloader = dict(batch_size=8, num_workers=4)
+# Enable automatic-mixed-precision training with AmpOptimWrapper.
+optim_wrapper = dict(
+ type='AmpOptimWrapper',
+ optimizer=dict(
+ type='SGD', lr=0.02 * 4, momentum=0.9, weight_decay=0.00004))
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_amp-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_amp-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a139c48b2091a3a40943ce7ec8301b06cea01d4f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_amp-1x_coco.py
@@ -0,0 +1,4 @@
+_base_ = './mask-rcnn_r50_fpn_1x_coco.py'
+
+# Enable automatic-mixed-precision training with AmpOptimWrapper.
+optim_wrapper = dict(type='AmpOptimWrapper')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_ms-poly-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_ms-poly-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..417adc3cebb3acbcc987b3f0453a78204dde1ea9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_ms-poly-3x_coco.py
@@ -0,0 +1,4 @@
+_base_ = [
+ '../common/ms-poly_3x_coco-instance.py',
+ '../_base_/models/mask-rcnn_r50_fpn.py'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_poly-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_poly-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..826180ce0a831a1ee6206bd52ffa516df766136c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_r50_fpn_poly-1x_coco.py
@@ -0,0 +1,18 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ poly2mask=False),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs'),
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-32x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-32x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..921ade81e30afb60a3a6f03d2f2aecef85767da8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-32x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './mask-rcnn_r101_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-32x4d_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-32x4d_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..db8157f80fac23f6216afbeefed6cb80398f7e0d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-32x4d_fpn_2x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './mask-rcnn_r101_fpn_2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-32x4d_fpn_ms-poly-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-32x4d_fpn_ms-poly-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..83e5451f38cb01d3d30712f22633fed6234d06c9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-32x4d_fpn_ms-poly-3x_coco.py
@@ -0,0 +1,18 @@
+_base_ = [
+ '../common/ms-poly_3x_coco-instance.py',
+ '../_base_/models/mask-rcnn_r50_fpn.py'
+]
+
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-32x8d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-32x8d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3e9b1b6fe8fcb152d9ad22bc403da6e62e936f77
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-32x8d_fpn_1x_coco.py
@@ -0,0 +1,22 @@
+_base_ = './mask-rcnn_r101_fpn_1x_coco.py'
+
+model = dict(
+ # ResNeXt-101-32x8d model trained with Caffe2 at FB,
+ # so the mean and std need to be changed.
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[57.375, 57.120, 58.395],
+ bgr_to_rgb=False),
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=8,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnext101_32x8d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-32x8d_fpn_ms-poly-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-32x8d_fpn_ms-poly-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6ee204d90001edd3e8e08e4a59ba25dd1ec4195c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-32x8d_fpn_ms-poly-1x_coco.py
@@ -0,0 +1,40 @@
+_base_ = './mask-rcnn_r101_fpn_1x_coco.py'
+
+model = dict(
+ # ResNeXt-101-32x8d model trained with Caffe2 at FB,
+ # so the mean and std need to be changed.
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[57.375, 57.120, 58.395],
+ bgr_to_rgb=False),
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=8,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnext101_32x8d')))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ poly2mask=False),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs'),
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-32x8d_fpn_ms-poly-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-32x8d_fpn_ms-poly-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..999a30c39fc083f26fe0cd9e2ec13bb4f6063268
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-32x8d_fpn_ms-poly-3x_coco.py
@@ -0,0 +1,25 @@
+_base_ = [
+ '../common/ms-poly_3x_coco-instance.py',
+ '../_base_/models/mask-rcnn_r50_fpn.py'
+]
+
+model = dict(
+ # ResNeXt-101-32x8d model trained with Caffe2 at FB,
+ # so the mean and std need to be changed.
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[57.375, 57.120, 58.395],
+ bgr_to_rgb=False),
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=8,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnext101_32x8d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-64x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-64x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..2cbb658c1b053d6674694c1a09101e965d5724ba
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-64x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './mask-rcnn_x101-32x4d_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-64x4d_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-64x4d_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f21a55b00db77a3cf2386a738a3b8fb39bf2fa44
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-64x4d_fpn_2x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './mask-rcnn_x101-32x4d_fpn_2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-64x4d_fpn_ms-poly_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-64x4d_fpn_ms-poly_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..09b49d47740b70c4a192d94a95b994d0a303f2d1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/mask-rcnn_x101-64x4d_fpn_ms-poly_3x_coco.py
@@ -0,0 +1,18 @@
+_base_ = [
+ '../common/ms-poly_3x_coco-instance.py',
+ '../_base_/models/mask-rcnn_r50_fpn.py'
+]
+
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..ddf85c872bc8681a849c59c917a4b5ca0151d21a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mask_rcnn/metafile.yml
@@ -0,0 +1,443 @@
+Collections:
+ - Name: Mask R-CNN
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Softmax
+ - RPN
+ - Convolution
+ - Dense Connections
+ - FPN
+ - ResNet
+ - RoIAlign
+ Paper:
+ URL: https://arxiv.org/abs/1703.06870v3
+ Title: "Mask R-CNN"
+ README: configs/mask_rcnn/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/detectors/mask_rcnn.py#L6
+ Version: v2.0.0
+
+Models:
+ - Name: mask-rcnn_r50-caffe_fpn_1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.3
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 34.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_caffe_fpn_1x_coco/mask_rcnn_r50_caffe_fpn_1x_coco_bbox_mAP-0.38__segm_mAP-0.344_20200504_231812-0ebd1859.pth
+
+ - Name: mask-rcnn_r50_fpn_1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.4
+ inference time (ms/im):
+ - value: 62.11
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 34.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_fpn_1x_coco/mask_rcnn_r50_fpn_1x_coco_20200205-d4b0c5d6.pth
+
+ - Name: mask-rcnn_r50_fpn_fp16_1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_r50_fpn_amp-1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.6
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ - Mixed Precision Training
+ inference time (ms/im):
+ - value: 41.49
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP16
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.1
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 34.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fp16/mask_rcnn_r50_fpn_fp16_1x_coco/mask_rcnn_r50_fpn_fp16_1x_coco_20200205-59faf7e4.pth
+
+ - Name: mask-rcnn_r50_fpn_2x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_r50_fpn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 4.4
+ inference time (ms/im):
+ - value: 62.11
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 35.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_fpn_2x_coco/mask_rcnn_r50_fpn_2x_coco_bbox_mAP-0.392__segm_mAP-0.354_20200505_003907-3e542a40.pth
+
+ - Name: mask-rcnn_r101-caffe_fpn_1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_r101-caffe_fpn_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r101_caffe_fpn_1x_coco/mask_rcnn_r101_caffe_fpn_1x_coco_20200601_095758-805e06c1.pth
+
+ - Name: mask-rcnn_r101_fpn_1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_r101_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.4
+ inference time (ms/im):
+ - value: 74.07
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r101_fpn_1x_coco/mask_rcnn_r101_fpn_1x_coco_20200204-1efe0ed5.pth
+
+ - Name: mask-rcnn_r101_fpn_2x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_r101_fpn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 6.4
+ inference time (ms/im):
+ - value: 74.07
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r101_fpn_2x_coco/mask_rcnn_r101_fpn_2x_coco_bbox_mAP-0.408__segm_mAP-0.366_20200505_071027-14b391c7.pth
+
+ - Name: mask-rcnn_x101-32x4d_fpn_1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_x101-32x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.6
+ inference time (ms/im):
+ - value: 88.5
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.9
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x4d_fpn_1x_coco/mask_rcnn_x101_32x4d_fpn_1x_coco_20200205-478d0b67.pth
+
+ - Name: mask-rcnn_x101-32x4d_fpn_2x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_x101-32x4d_fpn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 7.6
+ inference time (ms/im):
+ - value: 88.5
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x4d_fpn_2x_coco/mask_rcnn_x101_32x4d_fpn_2x_coco_bbox_mAP-0.422__segm_mAP-0.378_20200506_004702-faef898c.pth
+
+ - Name: mask-rcnn_x101-64x4d_fpn_1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_x101-64x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 10.7
+ inference time (ms/im):
+ - value: 125
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_64x4d_fpn_1x_coco/mask_rcnn_x101_64x4d_fpn_1x_coco_20200201-9352eb0d.pth
+
+ - Name: mask-rcnn_x101-64x4d_fpn_2x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_x101-64x4d_fpn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 10.7
+ inference time (ms/im):
+ - value: 125
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.7
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_64x4d_fpn_2x_coco/mask_rcnn_x101_64x4d_fpn_2x_coco_20200509_224208-39d6f70c.pth
+
+ - Name: mask-rcnn_x101-32x8d_fpn_1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_x101-32x8d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 10.6
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x8d_fpn_1x_coco/mask_rcnn_x101_32x8d_fpn_1x_coco_20220630_173841-0aaf329e.pth
+
+ - Name: mask-rcnn_r50-caffe_fpn_ms-poly-2x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_ms-poly-2x_coco.py
+ Metadata:
+ Training Memory (GB): 4.3
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.3
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_caffe_fpn_mstrain-poly_2x_coco/mask_rcnn_r50_caffe_fpn_mstrain-poly_2x_coco_bbox_mAP-0.403__segm_mAP-0.365_20200504_231822-a75c98ce.pth
+
+ - Name: mask-rcnn_r50-caffe_fpn_ms-poly-3x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_r50-caffe_fpn_ms-poly-3x_coco.py
+ Metadata:
+ Training Memory (GB): 4.3
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_caffe_fpn_mstrain-poly_3x_coco/mask_rcnn_r50_caffe_fpn_mstrain-poly_3x_coco_bbox_mAP-0.408__segm_mAP-0.37_20200504_163245-42aa3d00.pth
+
+ - Name: mask-rcnn_r50_fpn_mstrain-poly_3x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_r50_fpn_ms-poly-3x_coco.py
+ Metadata:
+ Training Memory (GB): 4.1
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.9
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_fpn_mstrain-poly_3x_coco/mask_rcnn_r50_fpn_mstrain-poly_3x_coco_20210524_201154-21b550bb.pth
+
+ - Name: mask-rcnn_r101_fpn_ms-poly-3x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_r101_fpn_ms-poly-3x_coco.py
+ Metadata:
+ Training Memory (GB): 6.1
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.7
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r101_fpn_mstrain-poly_3x_coco/mask_rcnn_r101_fpn_mstrain-poly_3x_coco_20210524_200244-5675c317.pth
+
+ - Name: mask-rcnn_r101-caffe_fpn_ms-poly-3x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_r101-caffe_fpn_ms-poly-3x_coco.py
+ Metadata:
+ Training Memory (GB): 5.9
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.9
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r101_caffe_fpn_mstrain-poly_3x_coco/mask_rcnn_r101_caffe_fpn_mstrain-poly_3x_coco_20210526_132339-3c33ce02.pth
+
+ - Name: mask-rcnn_x101-32x4d_fpn_ms-poly-3x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_x101-32x4d_fpn_ms-poly-3x_coco.py
+ Metadata:
+ Training Memory (GB): 7.3
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.6
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x4d_fpn_mstrain-poly_3x_coco/mask_rcnn_x101_32x4d_fpn_mstrain-poly_3x_coco_20210524_201410-abcd7859.pth
+
+ - Name: mask-rcnn_x101-32x8d_fpn_ms-poly-1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_x101-32x8d_fpn_ms-poly-1x_coco.py
+ Metadata:
+ Training Memory (GB): 10.4
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x8d_fpn_mstrain-poly_1x_coco/mask_rcnn_x101_32x8d_fpn_mstrain-poly_1x_coco_20220630_170346-b4637974.pth
+
+ - Name: mask-rcnn_x101-32x8d_fpn_ms-poly-3x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_x101-32x8d_fpn_ms-poly-3x_coco.py
+ Metadata:
+ Training Memory (GB): 10.3
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x8d_fpn_mstrain-poly_3x_coco/mask_rcnn_x101_32x8d_fpn_mstrain-poly_3x_coco_20210607_161042-8bd2c639.pth
+
+ - Name: mask-rcnn_x101-64x4d_fpn_ms-poly_3x_coco
+ In Collection: Mask R-CNN
+ Config: configs/mask_rcnn/mask-rcnn_x101-64x4d_fpn_ms-poly_3x_coco.py
+ Metadata:
+ Epochs: 36
+ Training Memory (GB): 10.4
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_64x4d_fpn_mstrain-poly_3x_coco/mask_rcnn_x101_64x4d_fpn_mstrain-poly_3x_coco_20210526_120447-c376f129.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/maskformer/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/maskformer/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..ca5ce320e1eb42f9cc12b4192fecb038fff71113
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/maskformer/README.md
@@ -0,0 +1,58 @@
+# MaskFormer
+
+> [Per-Pixel Classification is Not All You Need for Semantic Segmentation](https://arxiv.org/abs/2107.06278)
+
+
+
+## Abstract
+
+Modern approaches typically formulate semantic segmentation as a per-pixel classification task, while instance-level segmentation is handled with an alternative mask classification. Our key insight: mask classification is sufficiently general to solve both semantic- and instance-level segmentation tasks in a unified manner using the exact same model, loss, and training procedure. Following this observation, we propose MaskFormer, a simple mask classification model which predicts a set of binary masks, each associated with a single global class label prediction. Overall, the proposed mask classification-based method simplifies the landscape of effective approaches to semantic and panoptic segmentation tasks and shows excellent empirical results. In particular, we observe that MaskFormer outperforms per-pixel classification baselines when the number of classes is large. Our mask classification-based method outperforms both current state-of-the-art semantic (55.6 mIoU on ADE20K) and panoptic segmentation (52.7 PQ on COCO) models.
+
+
+

+
+
+## Introduction
+
+MaskFormer requires COCO and [COCO-panoptic](http://images.cocodataset.org/annotations/panoptic_annotations_trainval2017.zip) dataset for training and evaluation. You need to download and extract it in the COCO dataset path.
+The directory should be like this.
+
+```none
+mmdetection
+├── mmdet
+├── tools
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── panoptic_train2017.json
+│ │ │ ├── panoptic_train2017
+│ │ │ ├── panoptic_val2017.json
+│ │ │ ├── panoptic_val2017
+│ │ ├── train2017
+│ │ ├── val2017
+│ │ ├── test2017
+```
+
+## Results and Models
+
+| Backbone | style | Lr schd | Mem (GB) | Inf time (fps) | PQ | SQ | RQ | PQ_th | SQ_th | RQ_th | PQ_st | SQ_st | RQ_st | Config | Download |
+| :------: | :-----: | :-----: | :------: | :------------: | :----: | :----: | :----: | :----: | :----: | :----: | :----: | :----: | :----: | :--------------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | pytorch | 75e | 16.2 | - | 46.757 | 80.297 | 57.176 | 50.829 | 81.125 | 61.798 | 40.610 | 79.048 | 50.199 | [config](./maskformer_r50_ms-16xb1-75e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/maskformer/maskformer_r50_ms-16xb1-75e_coco/maskformer_r50_ms-16xb1-75e_coco_20230116_095226-baacd858.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/maskformer/maskformer_r50_ms-16xb1-75e_coco/maskformer_r50_ms-16xb1-75e_coco_20230116_095226.log.json) |
+| Swin-L | pytorch | 300e | 27.2 | - | 53.249 | 81.704 | 64.231 | 58.798 | 82.923 | 70.282 | 44.874 | 79.863 | 55.097 | [config](./maskformer_swin-l-p4-w12_64xb1-ms-300e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/maskformer/maskformer_swin-l-p4-w12_64xb1-ms-300e_coco/maskformer_swin-l-p4-w12_64xb1-ms-300e_coco_20220326_221612-c63ab967.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/maskformer/maskformer_swin-l-p4-w12_mstrain_64x1_300e_coco/maskformer_swin-l-p4-w12_mstrain_64x1_300e_coco_20220326_221612.log.json) |
+
+### Note
+
+1. The `R-50` version was mentioned in Table XI, in paper [Masked-attention Mask Transformer for Universal Image Segmentation](https://arxiv.org/abs/2112.01527).
+2. The models were trained with mmdet 2.x and have been converted for mmdet 3.x.
+
+## Citation
+
+```latex
+@inproceedings{cheng2021maskformer,
+ title={Per-Pixel Classification is Not All You Need for Semantic Segmentation},
+ author={Bowen Cheng and Alexander G. Schwing and Alexander Kirillov},
+ journal={NeurIPS},
+ year={2021}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/maskformer/maskformer_r50_ms-16xb1-75e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/maskformer/maskformer_r50_ms-16xb1-75e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..784ee7767bf1318e967444461028b49a38dc3dbc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/maskformer/maskformer_r50_ms-16xb1-75e_coco.py
@@ -0,0 +1,216 @@
+_base_ = [
+ '../_base_/datasets/coco_panoptic.py', '../_base_/default_runtime.py'
+]
+
+data_preprocessor = dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=1,
+ pad_mask=True,
+ mask_pad_value=0,
+ pad_seg=True,
+ seg_pad_value=255)
+
+num_things_classes = 80
+num_stuff_classes = 53
+num_classes = num_things_classes + num_stuff_classes
+model = dict(
+ type='MaskFormer',
+ data_preprocessor=data_preprocessor,
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=-1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ panoptic_head=dict(
+ type='MaskFormerHead',
+ in_channels=[256, 512, 1024, 2048], # pass to pixel_decoder inside
+ feat_channels=256,
+ out_channels=256,
+ num_things_classes=num_things_classes,
+ num_stuff_classes=num_stuff_classes,
+ num_queries=100,
+ pixel_decoder=dict(
+ type='TransformerEncoderPixelDecoder',
+ norm_cfg=dict(type='GN', num_groups=32),
+ act_cfg=dict(type='ReLU'),
+ encoder=dict( # DetrTransformerEncoder
+ num_layers=6,
+ layer_cfg=dict( # DetrTransformerEncoderLayer
+ self_attn_cfg=dict( # MultiheadAttention
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.1,
+ batch_first=True),
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048,
+ num_fcs=2,
+ ffn_drop=0.1,
+ act_cfg=dict(type='ReLU', inplace=True)))),
+ positional_encoding=dict(num_feats=128, normalize=True)),
+ enforce_decoder_input_project=False,
+ positional_encoding=dict(num_feats=128, normalize=True),
+ transformer_decoder=dict( # DetrTransformerDecoder
+ num_layers=6,
+ layer_cfg=dict( # DetrTransformerDecoderLayer
+ self_attn_cfg=dict( # MultiheadAttention
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.1,
+ batch_first=True),
+ cross_attn_cfg=dict( # MultiheadAttention
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.1,
+ batch_first=True),
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048,
+ num_fcs=2,
+ ffn_drop=0.1,
+ act_cfg=dict(type='ReLU', inplace=True))),
+ return_intermediate=True),
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0,
+ reduction='mean',
+ class_weight=[1.0] * num_classes + [0.1]),
+ loss_mask=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ reduction='mean',
+ loss_weight=20.0),
+ loss_dice=dict(
+ type='DiceLoss',
+ use_sigmoid=True,
+ activate=True,
+ reduction='mean',
+ naive_dice=True,
+ eps=1.0,
+ loss_weight=1.0)),
+ panoptic_fusion_head=dict(
+ type='MaskFormerFusionHead',
+ num_things_classes=num_things_classes,
+ num_stuff_classes=num_stuff_classes,
+ loss_panoptic=None,
+ init_cfg=None),
+ train_cfg=dict(
+ assigner=dict(
+ type='HungarianAssigner',
+ match_costs=[
+ dict(type='ClassificationCost', weight=1.0),
+ dict(type='FocalLossCost', weight=20.0, binary_input=True),
+ dict(type='DiceCost', weight=1.0, pred_act=True, eps=1.0)
+ ]),
+ sampler=dict(type='MaskPseudoSampler')),
+ test_cfg=dict(
+ panoptic_on=True,
+ # For now, the dataset does not support
+ # evaluating semantic segmentation metric.
+ semantic_on=False,
+ instance_on=False,
+ # max_per_image is for instance segmentation.
+ max_per_image=100,
+ object_mask_thr=0.8,
+ iou_thr=0.8,
+ # In MaskFormer's panoptic postprocessing,
+ # it will not filter masks whose score is smaller than 0.5 .
+ filter_low_score=False),
+ init_cfg=None)
+
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(
+ type='LoadPanopticAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ with_seg=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[[
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(400, 1333), (500, 1333), (600, 1333)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333),
+ (576, 1333), (608, 1333), (640, 1333),
+ (672, 1333), (704, 1333), (736, 1333),
+ (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]]),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(
+ batch_size=1, num_workers=1, dataset=dict(pipeline=train_pipeline))
+
+val_dataloader = dict(batch_size=1, num_workers=1)
+
+test_dataloader = val_dataloader
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(
+ type='AdamW',
+ lr=0.0001,
+ weight_decay=0.0001,
+ eps=1e-8,
+ betas=(0.9, 0.999)),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'backbone': dict(lr_mult=0.1, decay_mult=1.0),
+ 'query_embed': dict(lr_mult=1.0, decay_mult=0.0)
+ },
+ norm_decay_mult=0.0),
+ clip_grad=dict(max_norm=0.01, norm_type=2))
+
+max_epochs = 75
+
+# learning rate
+param_scheduler = dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[50],
+ gamma=0.1)
+
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (16 GPUs) x (1 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/maskformer/maskformer_swin-l-p4-w12_64xb1-ms-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/maskformer/maskformer_swin-l-p4-w12_64xb1-ms-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..9e4897f26d47c049f8791169867c2df307b87f61
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/maskformer/maskformer_swin-l-p4-w12_64xb1-ms-300e_coco.py
@@ -0,0 +1,73 @@
+_base_ = './maskformer_r50_ms-16xb1-75e_coco.py'
+
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_large_patch4_window12_384_22k.pth' # noqa
+depths = [2, 2, 18, 2]
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ pretrain_img_size=384,
+ embed_dims=192,
+ patch_size=4,
+ window_size=12,
+ mlp_ratio=4,
+ depths=depths,
+ num_heads=[6, 12, 24, 48],
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(0, 1, 2, 3),
+ with_cp=False,
+ convert_weights=True,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ panoptic_head=dict(
+ in_channels=[192, 384, 768, 1536], # pass to pixel_decoder inside
+ pixel_decoder=dict(
+ _delete_=True,
+ type='PixelDecoder',
+ norm_cfg=dict(type='GN', num_groups=32),
+ act_cfg=dict(type='ReLU')),
+ enforce_decoder_input_project=True))
+
+# optimizer
+
+# weight_decay = 0.01
+# norm_weight_decay = 0.0
+# embed_weight_decay = 0.0
+embed_multi = dict(lr_mult=1.0, decay_mult=0.0)
+norm_multi = dict(lr_mult=1.0, decay_mult=0.0)
+custom_keys = {
+ 'norm': norm_multi,
+ 'absolute_pos_embed': embed_multi,
+ 'relative_position_bias_table': embed_multi,
+ 'query_embed': embed_multi
+}
+
+optim_wrapper = dict(
+ optimizer=dict(lr=6e-5, weight_decay=0.01),
+ paramwise_cfg=dict(custom_keys=custom_keys, norm_decay_mult=0.0))
+
+max_epochs = 300
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=1e-6, by_epoch=False, begin=0, end=1500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[250],
+ gamma=0.1)
+]
+
+train_cfg = dict(max_epochs=max_epochs)
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (64 GPUs) x (1 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/maskformer/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/maskformer/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..fa58269d51c3e936f6acfaa664766afb84e7e0b6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/maskformer/metafile.yml
@@ -0,0 +1,43 @@
+Collections:
+ - Name: MaskFormer
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - AdamW
+ - Weight Decay
+ Training Resources: 16x V100 GPUs
+ Architecture:
+ - MaskFormer
+ Paper:
+ URL: https://arxiv.org/pdf/2107.06278
+ Title: 'Per-Pixel Classification is Not All You Need for Semantic Segmentation'
+ README: configs/maskformer/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.22.0/mmdet/models/detectors/maskformer.py#L7
+ Version: v2.22.0
+
+Models:
+ - Name: maskformer_r50_ms-16xb1-75e_coco
+ In Collection: MaskFormer
+ Config: configs/maskformer/maskformer_r50_ms-16xb1-75e_coco.py
+ Metadata:
+ Training Memory (GB): 16.2
+ Epochs: 75
+ Results:
+ - Task: Panoptic Segmentation
+ Dataset: COCO
+ Metrics:
+ PQ: 46.9
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/maskformer/maskformer_r50_ms-16xb1-75e_coco/maskformer_r50_ms-16xb1-75e_coco_20230116_095226-baacd858.pth
+ - Name: maskformer_swin-l-p4-w12_64xb1-ms-300e_coco
+ In Collection: MaskFormer
+ Config: configs/maskformer/maskformer_swin-l-p4-w12_64xb1-ms-300e_coco.py
+ Metadata:
+ Training Memory (GB): 27.2
+ Epochs: 300
+ Results:
+ - Task: Panoptic Segmentation
+ Dataset: COCO
+ Metrics:
+ PQ: 53.2
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/maskformer/maskformer_swin-l-p4-w12_64xb1-ms-300e_coco/maskformer_swin-l-p4-w12_64xb1-ms-300e_coco_20220326_221612-c63ab967.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..5cef692a382635e88732b2dc38985cfbc3c773e7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/README.md
@@ -0,0 +1,93 @@
+# Video Instance Segmentation
+
+## Abstract
+
+
+
+In this paper we present a new computer vision task, named video instance segmentation. The goal of this new task is simultaneous detection, segmentation and tracking of instances in videos. In words, it is the first time that the image instance segmentation problem is extended to the video domain. To facilitate research on this new task, we propose a large-scale benchmark called YouTube-VIS, which consists of 2883 high-resolution YouTube videos, a 40-category label set and 131k high-quality instance masks. In addition, we propose a novel algorithm called MaskTrack R-CNN for this task. Our new method introduces a new tracking branch to Mask R-CNN to jointly perform the detection, segmentation and tracking tasks simultaneously. Finally, we evaluate the proposed method and several strong baselines on our new dataset. Experimental results clearly demonstrate the advantages of the proposed algorithm and reveal insight for future improvement. We believe the video instance segmentation task will motivate the community along the line of research for video understanding.
+
+
+
+
+

+
+
+## Citation
+
+
+
+```latex
+@inproceedings{yang2019video,
+ title={Video instance segmentation},
+ author={Yang, Linjie and Fan, Yuchen and Xu, Ning},
+ booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
+ pages={5188--5197},
+ year={2019}
+}
+```
+
+## Results and models of MaskTrack R-CNN on YouTube-VIS 2019 validation dataset
+
+As mentioned in [Issues #6](https://github.com/youtubevos/MaskTrackRCNN/issues/6#issuecomment-502503505) in MaskTrack R-CNN, the result is kind of unstable for different trials, which ranges from 28 AP to 31 AP when using R-50-FPN as backbone.
+The checkpoint provided below is the best one from two experiments.
+
+| Method | Base detector | Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | AP | Config | Download |
+| :-------------: | :-----------: | :-------: | :-----: | :-----: | :------: | :------------: | :--: | :--------------------------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| MaskTrack R-CNN | Mask R-CNN | R-50-FPN | pytorch | 12e | 1.61 | - | 30.2 | [config](masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2019.py) | [model](https://download.openmmlab.com/mmtracking/vis/masktrack_rcnn/masktrack_rcnn_r50_fpn_12e_youtubevis2019/masktrack_rcnn_r50_fpn_12e_youtubevis2019_20211022_194830-6ca6b91e.pth) \| [log](https://download.openmmlab.com/mmtracking/vis/masktrack_rcnn/masktrack_rcnn_r50_fpn_12e_youtubevis2019/masktrack_rcnn_r50_fpn_12e_youtubevis2019_20211022_194830.log.json) |
+| MaskTrack R-CNN | Mask R-CNN | R-101-FPN | pytorch | 12e | 2.27 | - | 32.2 | [config](masktrack-rcnn_mask-rcnn_r101_fpn_8xb1-12e_youtubevis2019.py) | [model](https://download.openmmlab.com/mmtracking/vis/masktrack_rcnn/masktrack_rcnn_r101_fpn_12e_youtubevis2019/masktrack_rcnn_r101_fpn_12e_youtubevis2019_20211023_150038-454dc48b.pth) \| [log](https://download.openmmlab.com/mmtracking/vis/masktrack_rcnn/masktrack_rcnn_r101_fpn_12e_youtubevis2019/masktrack_rcnn_r101_fpn_12e_youtubevis2019_20211023_150038.log.json) |
+| MaskTrack R-CNN | Mask R-CNN | X-101-FPN | pytorch | 12e | 3.69 | - | 34.7 | [config](masktrack-rcnn_mask-rcnn_x101_fpn_8xb1-12e_youtubevis2019.py) | [model](https://download.openmmlab.com/mmtracking/vis/masktrack_rcnn/masktrack_rcnn_x101_fpn_12e_youtubevis2019/masktrack_rcnn_x101_fpn_12e_youtubevis2019_20211023_153205-fff7a102.pth) \| [log](https://download.openmmlab.com/mmtracking/vis/masktrack_rcnn/masktrack_rcnn_x101_fpn_12e_youtubevis2019/masktrack_rcnn_x101_fpn_12e_youtubevis2019_20211023_153205.log.json) |
+
+## Results and models of MaskTrack R-CNN on YouTube-VIS 2021 validation dataset
+
+The checkpoint provided below is the best one from two experiments.
+
+| Method | Base detector | Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | AP | Config | Download |
+| :-------------: | :-----------: | :-------: | :-----: | :-----: | :------: | :------------: | :--: | :--------------------------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| MaskTrack R-CNN | Mask R-CNN | R-50-FPN | pytorch | 12e | 1.61 | - | 28.7 | [config](masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2021.py) | [model](https://download.openmmlab.com/mmtracking/vis/masktrack_rcnn/masktrack_rcnn_r50_fpn_12e_youtubevis2021/masktrack_rcnn_r50_fpn_12e_youtubevis2021_20211026_044948-10da90d9.pth) \| [log](https://download.openmmlab.com/mmtracking/vis/masktrack_rcnn/masktrack_rcnn_r50_fpn_12e_youtubevis2021/masktrack_rcnn_r50_fpn_12e_youtubevis2021_20211026_044948.log.json) |
+| MaskTrack R-CNN | Mask R-CNN | R-101-FPN | pytorch | 12e | 2.27 | - | 31.3 | [config](masktrack-rcnn_mask-rcnn_r101_fpn_8xb1-12e_youtubevis2021.py) | [model](https://download.openmmlab.com/mmtracking/vis/masktrack_rcnn/masktrack_rcnn_r101_fpn_12e_youtubevis2021/masktrack_rcnn_r101_fpn_12e_youtubevis2021_20211026_045509-3c49e4f3.pth) \| [log](https://download.openmmlab.com/mmtracking/vis/masktrack_rcnn/masktrack_rcnn_r101_fpn_12e_youtubevis2021/masktrack_rcnn_r101_fpn_12e_youtubevis2021_20211026_045509.log.json) |
+| MaskTrack R-CNN | Mask R-CNN | X-101-FPN | pytorch | 12e | 3.69 | - | 33.5 | [config](masktrack-rcnn_mask-rcnn_x101_fpn_8xb1-12e_youtubevis2021.py) | [model](https://download.openmmlab.com/mmtracking/vis/masktrack_rcnn/masktrack_rcnn_x101_fpn_12e_youtubevis2021/masktrack_rcnn_x101_fpn_12e_youtubevis2021_20211026_095943-90831df4.pth) \| [log](https://download.openmmlab.com/mmtracking/vis/masktrack_rcnn/masktrack_rcnn_x101_fpn_12e_youtubevis2021/masktrack_rcnn_x101_fpn_12e_youtubevis2021_20211026_095943.log.json) |
+
+## Get started
+
+### 1. Development Environment Setup
+
+Tracking Development Environment Setup can refer to this [document](../../docs/en/get_started.md).
+
+### 2. Dataset Prepare
+
+Tracking Dataset Prepare can refer to this [document](../../docs/en/user_guides/tracking_dataset_prepare.md).
+
+### 3. Training
+
+Due to the influence of parameters such as learning rate in default configuration file, we recommend using 8 GPUs for training in order to reproduce accuracy. You can use the following command to start the training.
+
+```shell
+# Training MaskTrack R-CNN on YouTube-VIS-2021 dataset with following command.
+# The number after config file represents the number of GPUs used. Here we use 8 GPUs.
+bash tools/dist_train.sh configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2021.py 8
+```
+
+If you want to know about more detailed usage of `train.py/dist_train.sh/slurm_train.sh`,
+please refer to this [document](../../docs/en/user_guides/tracking_train_test.md).
+
+### 4. Testing and evaluation
+
+If you want to get the results of the [YouTube-VOS](https://youtube-vos.org/dataset/vis/) val/test set, please use the following command to generate result files that can be used for submission. It will be stored in `./youtube_vis_results.submission_file.zip`, you can modify the saved path in `test_evaluator` of the config.
+
+```shell
+# The number after config file represents the number of GPUs used.
+bash tools/dist_test_tracking.sh configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2021.py 8 --checkpoint ${CHECKPOINT_PATH}
+```
+
+If you want to know about more detailed usage of `train.py/dist_train.sh/slurm_train.sh`,
+please refer to this [document](../../docs/en/user_guides/tracking_train_test.md).
+
+### 5.Inference
+
+Use a single GPU to predict a video and save it as a video.
+
+```shell
+python demo/mot_demo.py demo/demo_mot.mp4 configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2021.py --checkpoint {CHECKPOINT_PATH} --out vis.mp4
+```
+
+If you want to know about more detailed usage of `mot_demo.py`, please refer to this [document](../../docs/en/user_guides/tracking_inference.md).
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_r101_fpn_8xb1-12e_youtubevis2019.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_r101_fpn_8xb1-12e_youtubevis2019.py
new file mode 100644
index 0000000000000000000000000000000000000000..4be492d5419b8598120faa29eed44eada0fb5ba2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_r101_fpn_8xb1-12e_youtubevis2019.py
@@ -0,0 +1,12 @@
+_base_ = ['./masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2019.py']
+model = dict(
+ detector=dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained', checkpoint='torchvision://resnet101')),
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint= # noqa: E251
+ 'https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r101_fpn_1x_coco/mask_rcnn_r101_fpn_1x_coco_20200204-1efe0ed5.pth' # noqa: E501
+ )))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_r101_fpn_8xb1-12e_youtubevis2021.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_r101_fpn_8xb1-12e_youtubevis2021.py
new file mode 100644
index 0000000000000000000000000000000000000000..81bae4af8d8945a024cd498a001e52059741f8a9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_r101_fpn_8xb1-12e_youtubevis2021.py
@@ -0,0 +1,28 @@
+_base_ = ['./masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2019.py']
+model = dict(
+ detector=dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained', checkpoint='torchvision://resnet101')),
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint= # noqa: E251
+ 'https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r101_fpn_1x_coco/mask_rcnn_r101_fpn_1x_coco_20200204-1efe0ed5.pth' # noqa: E501
+ )))
+
+data_root = 'data/youtube_vis_2021/'
+dataset_version = data_root[-5:-1]
+
+# dataloader
+train_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ dataset_version=dataset_version,
+ ann_file='annotations/youtube_vis_2021_train.json'))
+val_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ dataset_version=dataset_version,
+ ann_file='annotations/youtube_vis_2021_valid.json'))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2019.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2019.py
new file mode 100644
index 0000000000000000000000000000000000000000..db1be7b0ddf00a07ce6e06e4e179059e68c103a3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2019.py
@@ -0,0 +1,130 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/youtube_vis.py', '../_base_/default_runtime.py'
+]
+
+detector = _base_.model
+detector.pop('data_preprocessor')
+detector.roi_head.bbox_head.update(dict(num_classes=40))
+detector.roi_head.mask_head.update(dict(num_classes=40))
+detector.train_cfg.rpn.sampler.update(dict(num=64))
+detector.train_cfg.rpn_proposal.update(dict(nms_pre=200, max_per_img=200))
+detector.train_cfg.rcnn.sampler.update(dict(num=128))
+detector.test_cfg.rpn.update(dict(nms_pre=200, max_per_img=200))
+detector.test_cfg.rcnn.update(dict(score_thr=0.01))
+detector['init_cfg'] = dict(
+ type='Pretrained',
+ checkpoint= # noqa: E251
+ 'https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_fpn_1x_coco/mask_rcnn_r50_fpn_1x_coco_20200205-d4b0c5d6.pth' # noqa: E501
+)
+del _base_.model
+
+model = dict(
+ type='MaskTrackRCNN',
+ data_preprocessor=dict(
+ type='TrackDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_mask=True,
+ pad_size_divisor=32),
+ detector=detector,
+ track_head=dict(
+ type='RoITrackHead',
+ roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=7, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ embed_head=dict(
+ type='RoIEmbedHead',
+ num_fcs=2,
+ roi_feat_size=7,
+ in_channels=256,
+ fc_out_channels=1024),
+ train_cfg=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ match_low_quality=True,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=128,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ pos_weight=-1,
+ debug=False)),
+ tracker=dict(
+ type='MaskTrackRCNNTracker',
+ match_weights=dict(det_score=1.0, iou=2.0, det_label=10.0),
+ num_frames_retain=20))
+
+dataset_type = 'YouTubeVISDataset'
+data_root = 'data/youtube_vis_2019/'
+dataset_version = data_root[-5:-1] # 2019 or 2021
+
+# train_dataloader
+train_dataloader = dict(
+ _delete_=True,
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='TrackImgSampler'), # image-based sampling
+ batch_sampler=dict(type='TrackAspectRatioBatchSampler'),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ dataset_version=dataset_version,
+ ann_file='annotations/youtube_vis_2019_train.json',
+ data_prefix=dict(img_path='train/JPEGImages'),
+ pipeline=_base_.train_pipeline))
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.00125, momentum=0.9, weight_decay=0.0001),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+# learning policy
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 3.0,
+ by_epoch=False,
+ begin=0,
+ end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+
+# visualizer
+default_hooks = dict(
+ visualization=dict(type='TrackVisualizationHook', draw=False))
+
+vis_backends = [dict(type='LocalVisBackend')]
+visualizer = dict(
+ type='TrackLocalVisualizer', vis_backends=vis_backends, name='visualizer')
+
+# runtime settings
+train_cfg = dict(type='EpochBasedTrainLoop', max_epochs=12, val_begin=13)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# evaluator
+val_evaluator = dict(
+ type='YouTubeVISMetric',
+ metric='youtube_vis_ap',
+ outfile_prefix='./youtube_vis_results',
+ format_only=True)
+test_evaluator = val_evaluator
+
+del detector
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2021.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2021.py
new file mode 100644
index 0000000000000000000000000000000000000000..47263d5091c3b5b76056373558ce9a0a97bb071b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2021.py
@@ -0,0 +1,17 @@
+_base_ = ['./masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2019.py']
+
+data_root = 'data/youtube_vis_2021/'
+dataset_version = data_root[-5:-1]
+
+# dataloader
+train_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ dataset_version=dataset_version,
+ ann_file='annotations/youtube_vis_2021_train.json'))
+val_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ dataset_version=dataset_version,
+ ann_file='annotations/youtube_vis_2021_valid.json'))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_x101_fpn_8xb1-12e_youtubevis2019.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_x101_fpn_8xb1-12e_youtubevis2019.py
new file mode 100644
index 0000000000000000000000000000000000000000..e7e3f11e13a3a20ba8e4311963db558a9e4fd247
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_x101_fpn_8xb1-12e_youtubevis2019.py
@@ -0,0 +1,16 @@
+_base_ = ['./masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2019.py']
+model = dict(
+ detector=dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://resnext101_64x4d')),
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint= # noqa: E251
+ 'https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_64x4d_fpn_1x_coco/mask_rcnn_x101_64x4d_fpn_1x_coco_20200201-9352eb0d.pth' # noqa: E501
+ )))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_x101_fpn_8xb1-12e_youtubevis2021.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_x101_fpn_8xb1-12e_youtubevis2021.py
new file mode 100644
index 0000000000000000000000000000000000000000..ea4c8b92483292cc7de1b2f321d4d514427f3cb5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_x101_fpn_8xb1-12e_youtubevis2021.py
@@ -0,0 +1,32 @@
+_base_ = ['./masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2019.py']
+model = dict(
+ detector=dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://resnext101_64x4d')),
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint= # noqa: E251
+ 'https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_64x4d_fpn_1x_coco/mask_rcnn_x101_64x4d_fpn_1x_coco_20200201-9352eb0d.pth' # noqa: E501
+ )))
+
+data_root = 'data/youtube_vis_2021/'
+dataset_version = data_root[-5:-1]
+
+# dataloader
+train_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ dataset_version=dataset_version,
+ ann_file='annotations/youtube_vis_2021_train.json'))
+val_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ dataset_version=dataset_version,
+ ann_file='annotations/youtube_vis_2021_valid.json'))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..7a1d71d582dc31f3c05f721c6ea8a225d0e0ce33
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/masktrack_rcnn/metafile.yml
@@ -0,0 +1,91 @@
+Collections:
+ - Name: MaskTrack R-CNN
+ Metadata:
+ Training Techniques:
+ - SGD with Momentum
+ Training Resources: 8x TiTanXP GPUs
+ Architecture:
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/pdf/1905.04804.pdf
+ Title: Video Instance Segmentation
+ README: configs/masktrack_rcnn/README.md
+
+Models:
+ - Name: masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2019
+ In Collection: MaskTrack R-CNN
+ Config: configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2019.py
+ Metadata:
+ Training Data: YouTube-VIS 2019
+ Training Memory (GB): 1.16
+ Results:
+ - Task: Video Instance Segmentation
+ Dataset: YouTube-VIS 2019
+ Metrics:
+ AP: 30.2
+ Weights: https://download.openmmlab.com/mmtracking/vis/masktrack_rcnn/masktrack_rcnn_r50_fpn_12e_youtubevis2019/masktrack_rcnn_r50_fpn_12e_youtubevis2019_20211022_194830-6ca6b91e.pth
+
+ - Name: masktrack-rcnn_mask-rcnn_r101_fpn_8xb1-12e_youtubevis2019
+ In Collection: MaskTrack R-CNN
+ Config: configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_r101_fpn_8xb1-12e_youtubevis2019.py
+ Metadata:
+ Training Data: YouTube-VIS 2019
+ Training Memory (GB): 2.27
+ Results:
+ - Task: Video Instance Segmentation
+ Dataset: YouTube-VIS 2019
+ Metrics:
+ AP: 32.2
+ Weights: https://download.openmmlab.com/mmtracking/vis/masktrack_rcnn/masktrack_rcnn_r101_fpn_12e_youtubevis2019/masktrack_rcnn_r101_fpn_12e_youtubevis2019_20211023_150038-454dc48b.pth
+
+ - Name: masktrack-rcnn_mask-rcnn_x101_fpn_8xb1-12e_youtubevis2019
+ In Collection: MaskTrack R-CNN
+ Config: configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_x101_fpn_8xb1-12e_youtubevis2019.py
+ Metadata:
+ Training Data: YouTube-VIS 2019
+ Training Memory (GB): 3.69
+ Results:
+ - Task: Video Instance Segmentation
+ Dataset: YouTube-VIS 2019
+ Metrics:
+ AP: 34.7
+ Weights: https://download.openmmlab.com/mmtracking/vis/masktrack_rcnn/masktrack_rcnn_x101_fpn_12e_youtubevis2019/masktrack_rcnn_x101_fpn_12e_youtubevis2019_20211023_153205-fff7a102.pth
+
+ - Name: masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2021
+ In Collection: MaskTrack R-CNN
+ Config: configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_r50_fpn_8xb1-12e_youtubevis2021.py
+ Metadata:
+ Training Data: YouTube-VIS 2021
+ Training Memory (GB): 1.16
+ Results:
+ - Task: Video Instance Segmentation
+ Dataset: YouTube-VIS 2021
+ Metrics:
+ AP: 28.7
+ Weights: https://download.openmmlab.com/mmtracking/vis/masktrack_rcnn/masktrack_rcnn_r50_fpn_12e_youtubevis2021/masktrack_rcnn_r50_fpn_12e_youtubevis2021_20211026_044948-10da90d9.pth
+
+ - Name: masktrack-rcnn_mask-rcnn_r101_fpn_8xb1-12e_youtubevis2021
+ In Collection: MaskTrack R-CNN
+ Config: configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_r101_fpn_8xb1-12e_youtubevis2021.py
+ Metadata:
+ Training Data: YouTube-VIS 2021
+ Training Memory (GB): 2.27
+ Results:
+ - Task: Video Instance Segmentation
+ Dataset: YouTube-VIS 2021
+ Metrics:
+ AP: 31.3
+ Weights: https://download.openmmlab.com/mmtracking/vis/masktrack_rcnn/masktrack_rcnn_r101_fpn_12e_youtubevis2021/masktrack_rcnn_r101_fpn_12e_youtubevis2021_20211026_045509-3c49e4f3.pth
+
+ - Name: masktrack-rcnn_mask-rcnn_x101_fpn_8xb1-12e_youtubevis2021
+ In Collection: MaskTrack R-CNN
+ Config: configs/masktrack_rcnn/masktrack-rcnn_mask-rcnn_x101_fpn_8xb1-12e_youtubevis2021.py
+ Metadata:
+ Training Data: YouTube-VIS 2021
+ Training Memory (GB): 3.69
+ Results:
+ - Task: Video Instance Segmentation
+ Dataset: YouTube-VIS 2021
+ Metrics:
+ AP: 33.5
+ Weights: https://download.openmmlab.com/mmtracking/vis/masktrack_rcnn/masktrack_rcnn_x101_fpn_12e_youtubevis2021/masktrack_rcnn_x101_fpn_12e_youtubevis2021_20211026_095943-90831df4.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/misc/d2_faster-rcnn_r50-caffe_fpn_ms-90k_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/misc/d2_faster-rcnn_r50-caffe_fpn_ms-90k_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d93e1562606b3d6bd657454c99220d329c526f30
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/misc/d2_faster-rcnn_r50-caffe_fpn_ms-90k_coco.py
@@ -0,0 +1,75 @@
+_base_ = '../common/ms-90k_coco.py'
+
+# model settings
+model = dict(
+ type='Detectron2Wrapper',
+ bgr_to_rgb=False,
+ detector=dict(
+ # The settings in `d2_detector` will merged into default settings
+ # in detectron2. More details please refer to
+ # https://github.com/facebookresearch/detectron2/blob/main/detectron2/config/defaults.py # noqa
+ meta_architecture='GeneralizedRCNN',
+ # If you want to finetune the detector, you can use the
+ # checkpoint released by detectron2, for example:
+ # weights='detectron2://COCO-Detection/faster_rcnn_R_50_FPN_1x/137257794/model_final_b275ba.pkl' # noqa
+ weights='detectron2://ImageNetPretrained/MSRA/R-50.pkl',
+ mask_on=False,
+ pixel_mean=[103.530, 116.280, 123.675],
+ pixel_std=[1.0, 1.0, 1.0],
+ backbone=dict(name='build_resnet_fpn_backbone', freeze_at=2),
+ resnets=dict(
+ depth=50,
+ out_features=['res2', 'res3', 'res4', 'res5'],
+ num_groups=1,
+ norm='FrozenBN'),
+ fpn=dict(
+ in_features=['res2', 'res3', 'res4', 'res5'], out_channels=256),
+ anchor_generator=dict(
+ name='DefaultAnchorGenerator',
+ sizes=[[32], [64], [128], [256], [512]],
+ aspect_ratios=[[0.5, 1.0, 2.0]],
+ angles=[[-90, 0, 90]]),
+ proposal_generator=dict(name='RPN'),
+ rpn=dict(
+ head_name='StandardRPNHead',
+ in_features=['p2', 'p3', 'p4', 'p5', 'p6'],
+ iou_thresholds=[0.3, 0.7],
+ iou_labels=[0, -1, 1],
+ batch_size_per_image=256,
+ positive_fraction=0.5,
+ bbox_reg_loss_type='smooth_l1',
+ bbox_reg_loss_weight=1.0,
+ bbox_reg_weights=(1.0, 1.0, 1.0, 1.0),
+ smooth_l1_beta=0.0,
+ loss_weight=1.0,
+ boundary_thresh=-1,
+ pre_nms_topk_train=2000,
+ post_nms_topk_train=1000,
+ pre_nms_topk_test=1000,
+ post_nms_topk_test=1000,
+ nms_thresh=0.7,
+ conv_dims=[-1]),
+ roi_heads=dict(
+ name='StandardROIHeads',
+ num_classes=80,
+ in_features=['p2', 'p3', 'p4', 'p5'],
+ iou_thresholds=[0.5],
+ iou_labels=[0, 1],
+ batch_size_per_image=512,
+ positive_fraction=0.25,
+ score_thresh_test=0.05,
+ nms_thresh_test=0.5,
+ proposal_append_gt=True),
+ roi_box_head=dict(
+ name='FastRCNNConvFCHead',
+ num_fc=2,
+ fc_dim=1024,
+ conv_dim=256,
+ pooler_type='ROIAlignV2',
+ pooler_resolution=7,
+ pooler_sampling_ratio=0,
+ bbox_reg_loss_type='smooth_l1',
+ bbox_reg_loss_weight=1.0,
+ bbox_reg_weights=(10.0, 10.0, 5.0, 5.0),
+ smooth_l1_beta=0.0,
+ cls_agnostic_bbox_reg=False)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/misc/d2_mask-rcnn_r50-caffe_fpn_ms-90k_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/misc/d2_mask-rcnn_r50-caffe_fpn_ms-90k_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c0919c4593f028445dc033e85314320f88409a54
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/misc/d2_mask-rcnn_r50-caffe_fpn_ms-90k_coco.py
@@ -0,0 +1,83 @@
+_base_ = '../common/ms-poly-90k_coco-instance.py'
+
+# model settings
+model = dict(
+ type='Detectron2Wrapper',
+ bgr_to_rgb=False,
+ detector=dict(
+ # The settings in `d2_detector` will merged into default settings
+ # in detectron2. More details please refer to
+ # https://github.com/facebookresearch/detectron2/blob/main/detectron2/config/defaults.py # noqa
+ meta_architecture='GeneralizedRCNN',
+ # If you want to finetune the detector, you can use the
+ # checkpoint released by detectron2, for example:
+ # weights='detectron2://COCO-InstanceSegmentation/mask_rcnn_R_50_FPN_1x/137260431/model_final_a54504.pkl' # noqa
+ weights='detectron2://ImageNetPretrained/MSRA/R-50.pkl',
+ mask_on=True,
+ pixel_mean=[103.530, 116.280, 123.675],
+ pixel_std=[1.0, 1.0, 1.0],
+ backbone=dict(name='build_resnet_fpn_backbone', freeze_at=2),
+ resnets=dict(
+ depth=50,
+ out_features=['res2', 'res3', 'res4', 'res5'],
+ num_groups=1,
+ norm='FrozenBN'),
+ fpn=dict(
+ in_features=['res2', 'res3', 'res4', 'res5'], out_channels=256),
+ anchor_generator=dict(
+ name='DefaultAnchorGenerator',
+ sizes=[[32], [64], [128], [256], [512]],
+ aspect_ratios=[[0.5, 1.0, 2.0]],
+ angles=[[-90, 0, 90]]),
+ proposal_generator=dict(name='RPN'),
+ rpn=dict(
+ head_name='StandardRPNHead',
+ in_features=['p2', 'p3', 'p4', 'p5', 'p6'],
+ iou_thresholds=[0.3, 0.7],
+ iou_labels=[0, -1, 1],
+ batch_size_per_image=256,
+ positive_fraction=0.5,
+ bbox_reg_loss_type='smooth_l1',
+ bbox_reg_loss_weight=1.0,
+ bbox_reg_weights=(1.0, 1.0, 1.0, 1.0),
+ smooth_l1_beta=0.0,
+ loss_weight=1.0,
+ boundary_thresh=-1,
+ pre_nms_topk_train=2000,
+ post_nms_topk_train=1000,
+ pre_nms_topk_test=1000,
+ post_nms_topk_test=1000,
+ nms_thresh=0.7,
+ conv_dims=[-1]),
+ roi_heads=dict(
+ name='StandardROIHeads',
+ num_classes=80,
+ in_features=['p2', 'p3', 'p4', 'p5'],
+ iou_thresholds=[0.5],
+ iou_labels=[0, 1],
+ batch_size_per_image=512,
+ positive_fraction=0.25,
+ score_thresh_test=0.05,
+ nms_thresh_test=0.5,
+ proposal_append_gt=True),
+ roi_box_head=dict(
+ name='FastRCNNConvFCHead',
+ num_fc=2,
+ fc_dim=1024,
+ conv_dim=256,
+ pooler_type='ROIAlignV2',
+ pooler_resolution=7,
+ pooler_sampling_ratio=0,
+ bbox_reg_loss_type='smooth_l1',
+ bbox_reg_loss_weight=1.0,
+ bbox_reg_weights=(10.0, 10.0, 5.0, 5.0),
+ smooth_l1_beta=0.0,
+ cls_agnostic_bbox_reg=False),
+ roi_mask_head=dict(
+ name='MaskRCNNConvUpsampleHead',
+ conv_dim=256,
+ num_conv=4,
+ pooler_type='ROIAlignV2',
+ pooler_resolution=14,
+ pooler_sampling_ratio=0,
+ cls_agnostic_mask=False)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/misc/d2_retinanet_r50-caffe_fpn_ms-90k_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/misc/d2_retinanet_r50-caffe_fpn_ms-90k_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d3f7587648bde1d15b5c3c1e1ace6c35bb7c20b0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/misc/d2_retinanet_r50-caffe_fpn_ms-90k_coco.py
@@ -0,0 +1,48 @@
+_base_ = '../common/ms-90k_coco.py'
+
+# model settings
+model = dict(
+ type='Detectron2Wrapper',
+ bgr_to_rgb=False,
+ detector=dict(
+ # The settings in `d2_detector` will merged into default settings
+ # in detectron2. More details please refer to
+ # https://github.com/facebookresearch/detectron2/blob/main/detectron2/config/defaults.py # noqa
+ meta_architecture='RetinaNet',
+ # If you want to finetune the detector, you can use the
+ # checkpoint released by detectron2, for example:
+ # weights='detectron2://COCO-Detection/retinanet_R_50_FPN_1x/190397773/model_final_bfca0b.pkl' # noqa
+ weights='detectron2://ImageNetPretrained/MSRA/R-50.pkl',
+ mask_on=False,
+ pixel_mean=[103.530, 116.280, 123.675],
+ pixel_std=[1.0, 1.0, 1.0],
+ backbone=dict(name='build_retinanet_resnet_fpn_backbone', freeze_at=2),
+ resnets=dict(
+ depth=50,
+ out_features=['res3', 'res4', 'res5'],
+ num_groups=1,
+ norm='FrozenBN'),
+ fpn=dict(in_features=['res3', 'res4', 'res5'], out_channels=256),
+ anchor_generator=dict(
+ name='DefaultAnchorGenerator',
+ sizes=[[x, x * 2**(1.0 / 3), x * 2**(2.0 / 3)]
+ for x in [32, 64, 128, 256, 512]],
+ aspect_ratios=[[0.5, 1.0, 2.0]],
+ angles=[[-90, 0, 90]]),
+ retinanet=dict(
+ num_classes=80,
+ in_features=['p3', 'p4', 'p5', 'p6', 'p7'],
+ num_convs=4,
+ iou_thresholds=[0.4, 0.5],
+ iou_labels=[0, -1, 1],
+ bbox_reg_weights=(1.0, 1.0, 1.0, 1.0),
+ bbox_reg_loss_type='smooth_l1',
+ smooth_l1_loss_beta=0.0,
+ focal_loss_gamma=2.0,
+ focal_loss_alpha=0.25,
+ prior_prob=0.01,
+ score_thresh_test=0.05,
+ topk_candidates_test=1000,
+ nms_thresh_test=0.5)))
+
+optim_wrapper = dict(optimizer=dict(lr=0.01))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..c88cb1c902667e4bb480eb143d7b1268c35433dd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/README.md
@@ -0,0 +1,387 @@
+# MM Grounding DINO
+
+> [An Open and Comprehensive Pipeline for Unified Object Grounding and Detection](https://arxiv.org/abs/2401.02361)
+
+
+
+## Abstract
+
+Grounding-DINO is a state-of-the-art open-set detection model that tackles multiple vision tasks including Open-Vocabulary Detection (OVD), Phrase Grounding (PG), and Referring Expression Comprehension (REC). Its effectiveness has led to its widespread adoption as a mainstream architecture for various downstream applications. However, despite its significance, the original Grounding-DINO model lacks comprehensive public technical details due to the unavailability of its training code. To bridge this gap, we present MM-Grounding-DINO, an open-source, comprehensive, and user-friendly baseline, which is built with the MMDetection toolbox. It adopts abundant vision datasets for pre-training and various detection and grounding datasets for fine-tuning. We give a comprehensive analysis of each reported result and detailed settings for reproduction. The extensive experiments on the benchmarks mentioned demonstrate that our MM-Grounding-DINO-Tiny outperforms the Grounding-DINO-Tiny baseline. We release all our models to the research community.
+
+
+

+
+
+
+

+
+
+## Dataset Preparation
+
+Please refer to [dataset_prepare.md](dataset_prepare.md) or [中文版数据准备](dataset_prepare_zh-CN.md)
+
+## ✨ What's New
+
+💎 **We have released the pre-trained weights for Swin-B and Swin-L, welcome to try and give feedback.**
+
+## Usage
+
+Please refer to [usage.md](usage.md) or [中文版用法说明](usage_zh-CN.md)
+
+## Zero-Shot COCO Results and Models
+
+| Model | Backbone | Style | COCO mAP | Pre-Train Data | Config | Download |
+| :----------: | :------: | :-------: | :--------: | :----------------------: | :------------------------------------------------------------------------------: | :-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| GDINO-T | Swin-T | Zero-shot | 46.7 | O365 | | |
+| GDINO-T | Swin-T | Zero-shot | 48.1 | O365,GoldG | | |
+| GDINO-T | Swin-T | Zero-shot | 48.4 | O365,GoldG,Cap4M | [config](../grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_cap4m.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/grounding_dino/groundingdino_swint_ogc_mmdet-822d7e9d.pth) |
+| MM-GDINO-T | Swin-T | Zero-shot | 48.5(+1.8) | O365 | [config](grounding_dino_swin-t_pretrain_obj365.py) | |
+| MM-GDINO-T | Swin-T | Zero-shot | 50.4(+2.3) | O365,GoldG | [config](grounding_dino_swin-t_pretrain_obj365_goldg.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg/grounding_dino_swin-t_pretrain_obj365_goldg_20231122_132602-4ea751ce.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg/grounding_dino_swin-t_pretrain_obj365_goldg_20231122_132602.log.json) |
+| MM-GDINO-T | Swin-T | Zero-shot | 50.5(+2.1) | O365,GoldG,GRIT | [config](grounding_dino_swin-t_pretrain_obj365_goldg_grit9m.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_20231128_200818-169cc352.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_20231128_200818.log.json) |
+| MM-GDINO-T | Swin-T | Zero-shot | 50.6(+2.2) | O365,GoldG,V3Det | [config](grounding_dino_swin-t_pretrain_obj365_goldg_v3det.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_v3det_20231218_095741-e316e297.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_v3det_20231218_095741.log.json) |
+| MM-GDINO-T | Swin-T | Zero-shot | 50.4(+2.0) | O365,GoldG,GRIT,V3Det | [config](grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047.log.json) |
+| MM-GDINO-B | Swin-B | Zero-shot | 52.5 | O365,GoldG,V3Det | [config](grounding_dino_swin-b_pretrain_obj365_goldg_v3det.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-b_pretrain_obj365_goldg_v3det/grounding_dino_swin-b_pretrain_obj365_goldg_v3de-f83eef00.pth) \| [log](<>) |
+| MM-GDINO-B\* | Swin-B | - | 59.5 | O365,ALL | [config](grounding_dino_swin-b_pretrain_all.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-b_pretrain_all/grounding_dino_swin-b_pretrain_all-f9818a7c.pth) \| [log](<>) |
+| MM-GDINO-L | Swin-L | Zero-shot | 53.0 | O365V2,OpenImageV6,GoldG | [config](grounding_dino_swin-l_pretrain_obj365_goldg.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-l_pretrain_obj365_goldg/grounding_dino_swin-l_pretrain_obj365_goldg-34dcdc53.pth) \| [log](<>) |
+| MM-GDINO-L\* | Swin-L | - | 60.3 | O365V2,OpenImageV6,ALL | [config](grounding_dino_swin-l_pretrain_all.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-l_pretrain_all/grounding_dino_swin-l_pretrain_all-56d69e78.pth) \| [log](<>) |
+
+- This * indicates that the model has not been fully trained yet. We will release the final weights in the future.
+- ALL: GoldG,V3det,COCO2017,LVISV1,COCO2014,GRIT,RefCOCO,RefCOCO+,RefCOCOg,gRefCOCO.
+
+## Zero-Shot LVIS Results
+
+| Model | MiniVal APr | MiniVal APc | MiniVal APf | MiniVal AP | Val1.0 APr | Val1.0 APc | Val1.0 APf | Val1.0 AP | Pre-Train Data |
+| :--------: | :---------: | :---------: | :---------: | :---------: | :--------: | :--------: | :--------: | :---------: | :-------------------: |
+| GDINO-T | 18.8 | 24.2 | 34.7 | 28.8 | 10.1 | 15.3 | 29.9 | 20.1 | O365,GoldG,Cap4M |
+| MM-GDINO-T | 28.1 | 30.2 | 42.0 | 35.7(+6.9) | 17.1 | 22.4 | 36.5 | 27.0(+6.9) | O365,GoldG |
+| MM-GDINO-T | 26.6 | 32.4 | 41.8 | 36.5(+7.7) | 17.3 | 22.6 | 36.4 | 27.1(+7.0) | O365,GoldG,GRIT |
+| MM-GDINO-T | 33.0 | 36.0 | 45.9 | 40.5(+11.7) | 21.5 | 25.5 | 40.2 | 30.6(+10.5) | O365,GoldG,V3Det |
+| MM-GDINO-T | 34.2 | 37.4 | 46.2 | 41.4(+12.6) | 23.6 | 27.6 | 40.5 | 31.9(+11.8) | O365,GoldG,GRIT,V3Det |
+
+- The MM-GDINO-T config file is [mini-lvis](lvis/grounding_dino_swin-t_pretrain_zeroshot_mini-lvis.py) and [lvis 1.0](lvis/grounding_dino_swin-t_pretrain_zeroshot_lvis.py)
+
+## Zero-Shot ODinW (Object Detection in the Wild) Results
+
+### Results and models of ODinW13
+
+| Method | GDINO-T
(O365,GoldG,Cap4M) | MM-GDINO-T
(O365,GoldG) | MM-GDINO-T
(O365,GoldG,GRIT) | MM-GDINO-T
(O365,GoldG,V3Det) | MM-GDINO-T
(O365,GoldG,GRIT,V3Det) |
+| --------------------- | -------------------------------- | ----------------------------- | ---------------------------------- | ----------------------------------- | ---------------------------------------- |
+| AerialMaritimeDrone | 0.173 | 0.133 | 0.155 | 0.177 | 0.151 |
+| Aquarium | 0.195 | 0.252 | 0.261 | 0.266 | 0.283 |
+| CottontailRabbits | 0.799 | 0.771 | 0.810 | 0.778 | 0.786 |
+| EgoHands | 0.608 | 0.499 | 0.537 | 0.506 | 0.519 |
+| NorthAmericaMushrooms | 0.507 | 0.331 | 0.462 | 0.669 | 0.767 |
+| Packages | 0.687 | 0.707 | 0.687 | 0.710 | 0.706 |
+| PascalVOC | 0.563 | 0.565 | 0.580 | 0.556 | 0.566 |
+| pistols | 0.726 | 0.585 | 0.709 | 0.671 | 0.729 |
+| pothole | 0.215 | 0.136 | 0.285 | 0.199 | 0.243 |
+| Raccoon | 0.549 | 0.469 | 0.511 | 0.553 | 0.535 |
+| ShellfishOpenImages | 0.393 | 0.321 | 0.437 | 0.519 | 0.488 |
+| thermalDogsAndPeople | 0.657 | 0.556 | 0.603 | 0.493 | 0.542 |
+| VehiclesOpenImages | 0.613 | 0.566 | 0.603 | 0.614 | 0.615 |
+| Average | **0.514** | **0.453** | **0.511** | **0.516** | **0.533** |
+
+- The MM-GDINO-T config file is [odinw13](odinw/grounding_dino_swin-t_pretrain_odinw13.py)
+
+### Results and models of ODinW35
+
+| Method | GDINO-T
(O365,GoldG,Cap4M) | MM-GDINO-T
(O365,GoldG) | MM-GDINO-T
(O365,GoldG,GRIT) | MM-GDINO-T
(O365,GoldG,V3Det) | MM-GDINO-T
(O365,GoldG,GRIT,V3Det) |
+| --------------------------- | -------------------------------- | ----------------------------- | ---------------------------------- | ----------------------------------- | ---------------------------------------- |
+| AerialMaritimeDrone_large | 0.173 | 0.133 | 0.155 | 0.177 | 0.151 |
+| AerialMaritimeDrone_tiled | 0.206 | 0.170 | 0.225 | 0.184 | 0.206 |
+| AmericanSignLanguageLetters | 0.002 | 0.016 | 0.020 | 0.011 | 0.007 |
+| Aquarium | 0.195 | 0.252 | 0.261 | 0.266 | 0.283 |
+| BCCD | 0.161 | 0.069 | 0.118 | 0.083 | 0.077 |
+| boggleBoards | 0.000 | 0.002 | 0.001 | 0.001 | 0.002 |
+| brackishUnderwater | 0.021 | 0.033 | 0.021 | 0.025 | 0.025 |
+| ChessPieces | 0.000 | 0.000 | 0.000 | 0.000 | 0.000 |
+| CottontailRabbits | 0.806 | 0.771 | 0.810 | 0.778 | 0.786 |
+| dice | 0.004 | 0.002 | 0.005 | 0.001 | 0.001 |
+| DroneControl | 0.042 | 0.047 | 0.097 | 0.088 | 0.074 |
+| EgoHands_generic | 0.608 | 0.527 | 0.537 | 0.506 | 0.519 |
+| EgoHands_specific | 0.002 | 0.001 | 0.005 | 0.007 | 0.003 |
+| HardHatWorkers | 0.046 | 0.048 | 0.070 | 0.070 | 0.108 |
+| MaskWearing | 0.004 | 0.009 | 0.004 | 0.011 | 0.009 |
+| MountainDewCommercial | 0.430 | 0.453 | 0.465 | 0.194 | 0.430 |
+| NorthAmericaMushrooms | 0.471 | 0.331 | 0.462 | 0.669 | 0.767 |
+| openPoetryVision | 0.000 | 0.001 | 0.000 | 0.000 | 0.000 |
+| OxfordPets_by_breed | 0.003 | 0.002 | 0.004 | 0.006 | 0.004 |
+| OxfordPets_by_species | 0.011 | 0.019 | 0.016 | 0.020 | 0.015 |
+| PKLot | 0.001 | 0.004 | 0.002 | 0.008 | 0.007 |
+| Packages | 0.695 | 0.707 | 0.687 | 0.710 | 0.706 |
+| PascalVOC | 0.563 | 0.565 | 0.580 | 0.566 | 0.566 |
+| pistols | 0.726 | 0.585 | 0.709 | 0.671 | 0.729 |
+| plantdoc | 0.005 | 0.005 | 0.007 | 0.008 | 0.011 |
+| pothole | 0.215 | 0.136 | 0.219 | 0.077 | 0.168 |
+| Raccoons | 0.549 | 0.469 | 0.511 | 0.553 | 0.535 |
+| selfdrivingCar | 0.089 | 0.091 | 0.076 | 0.094 | 0.083 |
+| ShellfishOpenImages | 0.393 | 0.321 | 0.437 | 0.519 | 0.488 |
+| ThermalCheetah | 0.087 | 0.063 | 0.081 | 0.030 | 0.045 |
+| thermalDogsAndPeople | 0.657 | 0.556 | 0.603 | 0.493 | 0.543 |
+| UnoCards | 0.006 | 0.012 | 0.010 | 0.009 | 0.005 |
+| VehiclesOpenImages | 0.613 | 0.566 | 0.603 | 0.614 | 0.615 |
+| WildfireSmoke | 0.134 | 0.106 | 0.154 | 0.042 | 0.127 |
+| websiteScreenshots | 0.012 | 0.02 | 0.016 | 0.016 | 0.016 |
+| Average | **0.227** | **0.202** | **0.228** | **0.214** | **0.284** |
+
+- The MM-GDINO-T config file is [odinw35](odinw/grounding_dino_swin-t_pretrain_odinw35.py)
+
+## Zero-Shot Referring Expression Comprehension Results
+
+| Method | GDINO-T
(O365,GoldG,Cap4M) | MM-GDINO-T
(O365,GoldG) | MM-GDINO-T
(O365,GoldG,GRIT) | MM-GDINO-T
(O365,GoldG,V3Det) | MM-GDINO-T
(O365,GoldG,GRIT,V3Det) |
+| ---------------------- | -------------------------------- | ----------------------------- | ---------------------------------- | ----------------------------------- | ---------------------------------------- |
+| RefCOCO val @1,5,10 | 50.8/89.5/94.9 | 53.1/89.9/94.7 | 53.4/90.3/95.5 | 52.1/89.8/95.0 | 53.1/89.7/95.1 |
+| RefCOCO testA @1,5,10 | 57.4/91.3/95.6 | 59.7/91.5/95.9 | 58.8/91.70/96.2 | 58.4/86.8/95.6 | 59.1/91.0/95.5 |
+| RefCOCO testB @1,5,10 | 45.0/86.5/92.9 | 46.4/86.9/92.2 | 46.8/87.7/93.3 | 45.4/86.2/92.6 | 46.8/87.8/93.6 |
+| RefCOCO+ val @1,5,10 | 51.6/86.4/92.6 | 53.1/87.0/92.8 | 53.5/88.0/93.7 | 52.5/86.8/93.2 | 52.7/87.7/93.5 |
+| RefCOCO+ testA @1,5,10 | 57.3/86.7/92.7 | 58.9/87.3/92.9 | 59.0/88.1/93.7 | 58.1/86.7/93.5 | 58.7/87.2/93.1 |
+| RefCOCO+ testB @1,5,10 | 46.4/84.1/90.7 | 47.9/84.3/91.0 | 47.9/85.5/92.7 | 46.9/83.7/91.5 | 48.4/85.8/92.1 |
+| RefCOCOg val @1,5,10 | 60.4/92.1/96.2 | 61.2/92.6/96.1 | 62.7/93.3/97.0 | 61.7/92.9/96.6 | 62.9/93.3/97.2 |
+| RefCOCOg test @1,5,10 | 59.7/92.1/96.3 | 61.1/93.3/96.7 | 62.6/94.9/97.1 | 61.0/93.1/96.8 | 62.9/93.9/97.4 |
+
+| Method | thresh_score | GDINO-T
(O365,GoldG,Cap4M) | MM-GDINO-T
(O365,GoldG) | MM-GDINO-T
(O365,GoldG,GRIT) | MM-GDINO-T
(O365,GoldG,V3Det) | MM-GDINO-T
(O365,GoldG,GRIT,V3Det) |
+| --------------------------------------- | ------------ | -------------------------------- | ----------------------------- | ---------------------------------- | ----------------------------------- | ---------------------------------------- |
+| gRefCOCO val Pr@(F1=1, IoU≥0.5),N-acc | 0.5 | 39.3/70.4 | | | | 39.4/67.5 |
+| gRefCOCO val Pr@(F1=1, IoU≥0.5),N-acc | 0.6 | 40.5/83.8 | | | | 40.6/83.1 |
+| gRefCOCO val Pr@(F1=1, IoU≥0.5),N-acc | 0.7 | 41.3/91.8 | 39.8/84.7 | 40.7/89.7 | 40.3/88.8 | 41.0/91.3 |
+| gRefCOCO val Pr@(F1=1, IoU≥0.5),N-acc | 0.8 | 41.5/96.8 | | | | 41.1/96.4 |
+| gRefCOCO testA Pr@(F1=1, IoU≥0.5),N-acc | 0.5 | 31.9/70.4 | | | | 33.1/69.5 |
+| gRefCOCO testA Pr@(F1=1, IoU≥0.5),N-acc | 0.6 | 29.3/82.9 | | | | 29.2/84.3 |
+| gRefCOCO testA Pr@(F1=1, IoU≥0.5),N-acc | 0.7 | 27.2/90.2 | 26.3/89.0 | 26.0/91.9 | 25.4/91.8 | 26.1/93.0 |
+| gRefCOCO testA Pr@(F1=1, IoU≥0.5),N-acc | 0.8 | 25.1/96.3 | | | | 23.8/97.2 |
+| gRefCOCO testB Pr@(F1=1, IoU≥0.5),N-acc | 0.5 | 30.9/72.5 | | | | 33.0/69.6 |
+| gRefCOCO testB Pr@(F1=1, IoU≥0.5),N-acc | 0.6 | 30.0/86.1 | | | | 31.6/96.7 |
+| gRefCOCO testB Pr@(F1=1, IoU≥0.5),N-acc | 0.7 | 29.7/93.5 | 31.3/84.8 | 30.6/90.2 | 30.7/89.9 | 30.4/92.3 |
+| gRefCOCO testB Pr@(F1=1, IoU≥0.5),N-acc | 0.8 | 29.1/97.4 | | | | 29.5/84.2 |
+
+- The MM-GDINO-T config file is [here](refcoco/grounding_dino_swin-t_pretrain_zeroshot_refexp.py)
+
+## Zero-Shot Description Detection Dataset(DOD)
+
+```shell
+pip install ddd-dataset
+```
+
+| Method | mode | GDINO-T
(O365,GoldG,Cap4M) | MM-GDINO-T
(O365,GoldG) | MM-GDINO-T
(O365,GoldG,GRIT) | MM-GDINO-T
(O365,GoldG,V3Det) | MM-GDINO-T
(O365,GoldG,GRIT,V3Det) |
+| -------------------------------- | -------- | -------------------------------- | ----------------------------- | ---------------------------------- | ----------------------------------- | ---------------------------------------- |
+| FULL/short/middle/long/very long | concat | 17.2/18.0/18.7/14.8/16.3 | 15.6/17.3/16.7/14.3/13.1 | 17.0/17.7/18.0/15.7/15.7 | 16.2/17.4/16.8/14.9/15.4 | 17.5/23.4/18.3/14.7/13.8 |
+| FULL/short/middle/long/very long | parallel | 22.3/28.2/24.8/19.1/13.9 | 21.7/24.7/24.0/20.2/13.7 | 22.5/25.6/25.1/20.5/14.9 | 22.3/25.6/24.5/20.6/14.7 | 22.9/28.1/25.4/20.4/14.4 |
+| PRES/short/middle/long/very long | concat | 17.8/18.3/19.2/15.2/17.3 | 16.4/18.4/17.3/14.5/14.2 | 17.9/19.0/18.3/16.5/17.5 | 16.6/18.8/17.1/15.1/15.0 | 18.0/23.7/18.6/15.4/13.3 |
+| PRES/short/middle/long/very long | parallel | 21.0/27.0/22.8/17.5/12.5 | 21.3/25.5/22.8/19.2/12.9 | 21.5/25.2/23.0/19.0/15.0 | 21.6/25.7/23.0/19.5/14.8 | 21.9/27.4/23.2/19.1/14.2 |
+| ABS/short/middle/long/very long | concat | 15.4/17.1/16.4/13.6/14.9 | 13.4/13.4/14.5/13.5/11.9 | 14.5/13.1/16.7/13.6/13.3 | 14.8/12.5/15.6/14.3/15.8 | 15.9/22.2/17.1/12.5/14.4 |
+| ABS/short/middle/long/very long | parallel | 26.0/32.0/33.0/23.6/15.5 | 22.8/22.2/28.7/22.9/14.7 | 25.6/26.8/33.9/24.5/14.7 | 24.1/24.9/30.7/23.8/14.7 | 26.0/30.3/34.1/23.9/14.6 |
+
+Note:
+
+1. Considering that the evaluation time for Inter-scenario is very long and the performance is low, it is temporarily not supported. The mentioned metrics are for Intra-scenario.
+2. `concat` is the default inference mode for Grounding DINO, where it concatenates multiple sub-sentences with "." to form a single sentence for inference. On the other hand, "parallel" performs inference on each sub-sentence in a for-loop.
+3. The MM-GDINO-T config file is [concat_dod](dod/grounding_dino_swin-t_pretrain_zeroshot_concat_dod.py) and [parallel_dod](dod/grounding_dino_swin-t_pretrain_zeroshot_parallel_dod.py)
+
+## Pretrain Flickr30k Results
+
+| Model | Pre-Train Data | Val R@1 | Val R@5 | Val R@10 | Test R@1 | Test R@5 | Test R@10 |
+| :--------: | :-------------------: | ------- | ------- | -------- | -------- | -------- | --------- |
+| GLIP-T | O365,GoldG | 84.9 | 94.9 | 96.3 | 85.6 | 95.4 | 96.7 |
+| GLIP-T | O365,GoldG,CC3M,SBU | 85.3 | 95.5 | 96.9 | 86.0 | 95.9 | 97.2 |
+| GDINO-T | O365,GoldG,Cap4M | 87.8 | 96.6 | 98.0 | 88.1 | 96.9 | 98.2 |
+| MM-GDINO-T | O365,GoldG | 85.5 | 95.6 | 97.2 | 86.2 | 95.7 | 97.4 |
+| MM-GDINO-T | O365,GoldG,GRIT | 86.7 | 95.8 | 97.6 | 87.0 | 96.2 | 97.7 |
+| MM-GDINO-T | O365,GoldG,V3Det | 85.9 | 95.7 | 97.4 | 86.3 | 95.7 | 97.4 |
+| MM-GDINO-T | O365,GoldG,GRIT,V3Det | 86.7 | 96.0 | 97.6 | 87.2 | 96.2 | 97.7 |
+
+Note:
+
+1. `@1,5,10` refers to precision at the top 1, 5, and 10 positions in a predicted ranked list.
+2. The MM-GDINO-T config file is [here](flickr30k/grounding_dino_swin-t-pretrain_flickr30k.py)
+
+## Validating the generalization of a pre-trained model through fine-tuning
+
+### RTTS
+
+| Architecture | Backbone | Lr schd | box AP |
+| :-----------------: | :------: | ------- | -------- |
+| Faster R-CNN | R-50 | 1x | 48.1 |
+| Cascade R-CNN | R-50 | 1x | 50.8 |
+| ATSS | R-50 | 1x | 48.2 |
+| TOOD | R-50 | 1X | 50.8 |
+| MM-GDINO(zero-shot) | Swin-T | | 49.8 |
+| MM-GDINO | Swin-T | 1x | **69.1** |
+
+- The reference metrics come from https://github.com/BIGWangYuDong/lqit/tree/main/configs/detection/rtts_dataset
+- The MM-GDINO-T config file is [here](rtts/grounding_dino_swin-t_finetune_8xb4_1x_rtts.py)
+
+### RUOD
+
+| Architecture | Backbone | Lr schd | box AP |
+| :-----------------: | :------: | ------- | -------- |
+| Faster R-CNN | R-50 | 1x | 52.4 |
+| Cascade R-CNN | R-50 | 1x | 55.3 |
+| ATSS | R-50 | 1x | 55.7 |
+| TOOD | R-50 | 1X | 57.4 |
+| MM-GDINO(zero-shot) | Swin-T | | 29.8 |
+| MM-GDINO | Swin-T | 1x | **65.5** |
+
+- The reference metrics come from https://github.com/BIGWangYuDong/lqit/tree/main/configs/detection/ruod_dataset
+- The MM-GDINO-T config file is [here](ruod/grounding_dino_swin-t_finetune_8xb4_1x_ruod.py)
+
+### Brain Tumor
+
+| Architecture | Backbone | Lr schd | box AP |
+| :-----------: | :------: | ------- | ------ |
+| Faster R-CNN | R-50 | 50e | 43.5 |
+| Cascade R-CNN | R-50 | 50e | 46.2 |
+| DINO | R-50 | 50e | 46.4 |
+| Cascade-DINO | R-50 | 50e | 48.6 |
+| MM-GDINO | Swin-T | 50e | 47.5 |
+
+- The reference metrics come from https://arxiv.org/abs/2307.11035
+- The MM-GDINO-T config file is [here](brain_tumor/grounding_dino_swin-t_finetune_8xb4_50e_brain_tumor.py)
+
+### Cityscapes
+
+| Architecture | Backbone | Lr schd | box AP |
+| :-----------------: | :------: | ------- | -------- |
+| Faster R-CNN | R-50 | 50e | 30.1 |
+| Cascade R-CNN | R-50 | 50e | 31.8 |
+| DINO | R-50 | 50e | 34.5 |
+| Cascade-DINO | R-50 | 50e | 34.8 |
+| MM-GDINO(zero-shot) | Swin-T | | 34.2 |
+| MM-GDINO | Swin-T | 50e | **51.5** |
+
+- The reference metrics come from https://arxiv.org/abs/2307.11035
+- The MM-GDINO-T config file is [here](cityscapes/grounding_dino_swin-t_finetune_8xb4_50e_cityscapes.py)
+
+### People in Painting
+
+| Architecture | Backbone | Lr schd | box AP |
+| :-----------------: | :------: | ------- | -------- |
+| Faster R-CNN | R-50 | 50e | 17.0 |
+| Cascade R-CNN | R-50 | 50e | 18.0 |
+| DINO | R-50 | 50e | 12.0 |
+| Cascade-DINO | R-50 | 50e | 13.4 |
+| MM-GDINO(zero-shot) | Swin-T | | 23.1 |
+| MM-GDINO | Swin-T | 50e | **38.9** |
+
+- The reference metrics come from https://arxiv.org/abs/2307.11035
+- The MM-GDINO-T config file is [here](people_in_painting/grounding_dino_swin-t_finetune_8xb4_50e_people_in_painting.py)
+
+### COCO
+
+**(1) Closed-set performance**
+
+| Architecture | Backbone | Lr schd | box AP |
+| :-----------------: | :------: | ------- | ------ |
+| Faster R-CNN | R-50 | 1x | 37.4 |
+| Cascade R-CNN | R-50 | 1x | 40.3 |
+| ATSS | R-50 | 1x | 39.4 |
+| TOOD | R-50 | 1X | 42.4 |
+| DINO | R-50 | 1X | 50.1 |
+| GLIP(zero-shot) | Swin-T | | 46.6 |
+| GDINO(zero-shot) | Swin-T | | 48.5 |
+| MM-GDINO(zero-shot) | Swin-T | | 50.4 |
+| GLIP | Swin-T | 1x | 55.4 |
+| GDINO | Swin-T | 1x | 58.1 |
+| MM-GDINO | Swin-T | 1x | 58.2 |
+
+- The MM-GDINO-T config file is [here](coco/grounding_dino_swin-t_finetune_16xb4_1x_coco.py)
+
+**(2) Open-set continuing pretraining performance**
+
+| Architecture | Backbone | Lr schd | box AP |
+| :-----------------: | :------: | :-----: | :----: |
+| GLIP(zero-shot) | Swin-T | | 46.7 |
+| GDINO(zero-shot) | Swin-T | | 48.5 |
+| MM-GDINO(zero-shot) | Swin-T | | 50.4 |
+| MM-GDINO | Swin-T | 1x | 54.7 |
+
+- The MM-GDINO-T config file is [here](coco/grounding_dino_swin-t_finetune_16xb4_1x_sft_coco.py)
+- Due to the small size of the COCO dataset, continuing pretraining solely on COCO can easily lead to overfitting. The results shown above are from the third epoch. I do not recommend you train using this approach.
+
+**(3) Open vocabulary performance**
+
+| Architecture | Backbone | Lr schd | box AP | Base box AP | Novel box AP | box AP@50 | Base box AP@50 | Novel box AP@50 |
+| :-----------------: | :------: | :-----: | :----: | :---------: | :----------: | :-------: | :------------: | :-------------: |
+| MM-GDINO(zero-shot) | Swin-T | | 51.1 | 48.4 | 58.9 | 66.7 | 64.0 | 74.2 |
+| MM-GDINO | Swin-T | 1x | 57.2 | 56.1 | 60.4 | 73.6 | 73.0 | 75.3 |
+
+- The MM-GDINO-T config file is [here](coco/grounding_dino_swin-t_finetune_16xb4_1x_coco_48_17.py)
+
+### LVIS 1.0
+
+**(1) Open-set continuing pretraining performance**
+
+| Architecture | Backbone | Lr schd | MiniVal APr | MiniVal APc | MiniVal APf | MiniVal AP | Val1.0 APr | Val1.0 APc | Val1.0 APf | Val1.0 AP |
+| :-----------------: | :------: | :-----: | :---------: | :---------: | :---------: | :--------: | :--------: | :--------: | :--------: | :-------: |
+| GLIP(zero-shot) | Swin-T | | 18.1 | 21.2 | 33.1 | 26.7 | 10.8 | 14.7 | 29.0 | 19.6 |
+| GDINO(zero-shot) | Swin-T | | 18.8 | 24.2 | 34.7 | 28.8 | 10.1 | 15.3 | 29.9 | 20.1 |
+| MM-GDINO(zero-shot) | Swin-T | | 34.2 | 37.4 | 46.2 | 41.4 | 23.6 | 27.6 | 40.5 | 31.9 |
+| MM-GDINO | Swin-T | 1x | 50.7 | 58.8 | 60.1 | 58.7 | 45.2 | 50.2 | 56.1 | 51.7 |
+
+- The MM-GDINO-T config file is [here](lvis/grounding_dino_swin-t_finetune_16xb4_1x_lvis.py)
+
+**(2) Open vocabulary performance**
+
+| Architecture | Backbone | Lr schd | MiniVal APr | MiniVal APc | MiniVal APf | MiniVal AP |
+| :-----------------: | :------: | :-----: | :---------: | :---------: | :---------: | :--------: |
+| MM-GDINO(zero-shot) | Swin-T | | 34.2 | 37.4 | 46.2 | 41.4 |
+| MM-GDINO | Swin-T | 1x | 43.2 | 57.4 | 59.3 | 57.1 |
+
+- The MM-GDINO-T config file is [here](lvis/grounding_dino_swin-t_finetune_16xb4_1x_lvis_866_337.py)
+
+### RefEXP
+
+#### RefCOCO
+
+| Architecture | Backbone | Lr schd | val @1 | val @5 | val @10 | testA @1 | testA @5 | testA @10 | testB @1 | testB @5 | testB @10 |
+| :-----------------: | :------: | :-----: | :----: | :----: | :-----: | :------: | :------: | :-------: | :------: | :------: | :-------: |
+| GDINO(zero-shot) | Swin-T | | 50.8 | 89.5 | 94.9 | 57.5 | 91.3 | 95.6 | 45.0 | 86.5 | 92.9 |
+| MM-GDINO(zero-shot) | Swin-T | | 53.1 | 89.7 | 95.1 | 59.1 | 91.0 | 95.5 | 46.8 | 87.8 | 93.6 |
+| GDINO | Swin-T | UNK | 89.2 | | | 91.9 | | | 86.0 | | |
+| MM-GDINO | Swin-T | 5e | 89.5 | 98.6 | 99.4 | 91.4 | 99.2 | 99.8 | 86.6 | 97.9 | 99.1 |
+
+- The MM-GDINO-T config file is [here](refcoco/grounding_dino_swin-t_finetune_8xb4_5e_refcoco.py)
+
+#### RefCOCO+
+
+| Architecture | Backbone | Lr schd | val @1 | val @5 | val @10 | testA @1 | testA @5 | testA @10 | testB @1 | testB @5 | testB @10 |
+| :-----------------: | :------: | :-----: | :----: | :----: | :-----: | :------: | :------: | :-------: | :------: | :------: | :-------: |
+| GDINO(zero-shot) | Swin-T | | 51.6 | 86.4 | 92.6 | 57.3 | 86.7 | 92.7 | 46.4 | 84.1 | 90.7 |
+| MM-GDINO(zero-shot) | Swin-T | | 52.7 | 87.7 | 93.5 | 58.7 | 87.2 | 93.1 | 48.4 | 85.8 | 92.1 |
+| GDINO | Swin-T | UNK | 81.1 | | | 87.4 | | | 74.7 | | |
+| MM-GDINO | Swin-T | 5e | 82.1 | 97.8 | 99.2 | 87.5 | 99.2 | 99.7 | 74.0 | 96.3 | 96.4 |
+
+- The MM-GDINO-T config file is [here](refcoco/grounding_dino_swin-t_finetune_8xb4_5e_refcoco_plus.py)
+
+#### RefCOCOg
+
+| Architecture | Backbone | Lr schd | val @1 | val @5 | val @10 | test @1 | test @5 | test @10 |
+| :-----------------: | :------: | :-----: | :----: | :----: | :-----: | :-----: | :-----: | :------: |
+| GDINO(zero-shot) | Swin-T | | 60.4 | 92.1 | 96.2 | 59.7 | 92.1 | 96.3 |
+| MM-GDINO(zero-shot) | Swin-T | | 62.9 | 93.3 | 97.2 | 62.9 | 93.9 | 97.4 |
+| GDINO | Swin-T | UNK | 84.2 | | | 84.9 | | |
+| MM-GDINO | Swin-T | 5e | 85.5 | 98.4 | 99.4 | 85.8 | 98.6 | 99.4 |
+
+- The MM-GDINO-T config file is [here](refcoco/grounding_dino_swin-t_finetune_8xb4_5e_refcocog.py)
+
+#### gRefCOCO
+
+| Architecture | Backbone | Lr schd | val Pr@(F1=1, IoU≥0.5) | val N-acc | testA Pr@(F1=1, IoU≥0.5) | testA N-acc | testB Pr@(F1=1, IoU≥0.5) | testB N-acc |
+| :-----------------: | :------: | :-----: | :--------------------: | :-------: | :----------------------: | :---------: | :----------------------: | :---------: |
+| GDINO(zero-shot) | Swin-T | | 41.3 | 91.8 | 27.2 | 90.2 | 29.7 | 93.5 |
+| MM-GDINO(zero-shot) | Swin-T | | 41.0 | 91.3 | 26.1 | 93.0 | 30.4 | 92.3 |
+| MM-GDINO | Swin-T | 5e | 45.1 | 64.7 | 42.5 | 65.5 | 40.3 | 63.2 |
+
+- The MM-GDINO-T config file is [here](refcoco/grounding_dino_swin-t_finetune_8xb4_5e_grefcoco.py)
+
+## Citation
+
+If you find this project useful in your research, please consider citing:
+
+```latex
+@article{zhao2024open,
+ title={An Open and Comprehensive Pipeline for Unified Object Grounding and Detection},
+ author={Zhao, Xiangyu and Chen, Yicheng and Xu, Shilin and Li, Xiangtai and Wang, Xinjiang and Li, Yining and Huang, Haian},
+ journal={arXiv preprint arXiv:2401.02361},
+ year={2024}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/brain_tumor/grounding_dino_swin-t_finetune_8xb4_50e_brain_tumor.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/brain_tumor/grounding_dino_swin-t_finetune_8xb4_50e_brain_tumor.py
new file mode 100644
index 0000000000000000000000000000000000000000..1172da5b64102413eec11f223f467ad4c03a7cdf
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/brain_tumor/grounding_dino_swin-t_finetune_8xb4_50e_brain_tumor.py
@@ -0,0 +1,112 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py'
+
+# https://universe.roboflow.com/roboflow-100/brain-tumor-m2pbp/dataset/2
+data_root = 'data/brain_tumor_v2/'
+class_name = ('label0', 'label1', 'label2')
+label_name = '_annotations.coco.json'
+
+palette = [(220, 20, 60), (255, 0, 0), (0, 0, 142)]
+
+metainfo = dict(classes=class_name, palette=palette)
+
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities'))
+]
+
+train_dataloader = dict(
+ sampler=dict(_delete_=True, type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ _delete_=True,
+ type='RepeatDataset',
+ times=10,
+ dataset=dict(
+ type='CocoDataset',
+ data_root=data_root,
+ metainfo=metainfo,
+ filter_cfg=dict(filter_empty_gt=False, min_size=32),
+ pipeline=train_pipeline,
+ return_classes=True,
+ data_prefix=dict(img='train/'),
+ ann_file='train/' + label_name)))
+
+val_dataloader = dict(
+ dataset=dict(
+ metainfo=metainfo,
+ data_root=data_root,
+ return_classes=True,
+ ann_file='valid/' + label_name,
+ data_prefix=dict(img='valid/')))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'valid/' + label_name,
+ metric='bbox',
+ format_only=False)
+test_evaluator = val_evaluator
+
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0001, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'backbone': dict(lr_mult=0.1)
+ }))
+
+# learning policy
+max_epochs = 5
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[4],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs, val_interval=1)
+
+default_hooks = dict(checkpoint=dict(max_keep_ckpts=1, save_best='auto'))
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/cityscapes/grounding_dino_swin-t_finetune_8xb4_50e_cityscapes.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/cityscapes/grounding_dino_swin-t_finetune_8xb4_50e_cityscapes.py
new file mode 100644
index 0000000000000000000000000000000000000000..c4283413c4ba0c060144d7fb85f7d064a60577c7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/cityscapes/grounding_dino_swin-t_finetune_8xb4_50e_cityscapes.py
@@ -0,0 +1,110 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py'
+
+data_root = 'data/cityscapes/'
+class_name = ('person', 'rider', 'car', 'truck', 'bus', 'train', 'motorcycle',
+ 'bicycle')
+palette = [(220, 20, 60), (255, 0, 0), (0, 0, 142), (0, 0, 70), (0, 60, 100),
+ (0, 80, 100), (0, 0, 230), (119, 11, 32)]
+
+metainfo = dict(classes=class_name, palette=palette)
+
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities'))
+]
+
+train_dataloader = dict(
+ sampler=dict(_delete_=True, type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ _delete_=True,
+ type='RepeatDataset',
+ times=10,
+ dataset=dict(
+ type='CocoDataset',
+ data_root=data_root,
+ metainfo=metainfo,
+ filter_cfg=dict(filter_empty_gt=False, min_size=32),
+ pipeline=train_pipeline,
+ return_classes=True,
+ data_prefix=dict(img='leftImg8bit/train/'),
+ ann_file='annotations/instancesonly_filtered_gtFine_train.json')))
+
+val_dataloader = dict(
+ dataset=dict(
+ metainfo=metainfo,
+ data_root=data_root,
+ return_classes=True,
+ ann_file='annotations/instancesonly_filtered_gtFine_val.json',
+ data_prefix=dict(img='leftImg8bit/val/')))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/instancesonly_filtered_gtFine_val.json',
+ metric='bbox',
+ format_only=False)
+test_evaluator = val_evaluator
+
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0001, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'backbone': dict(lr_mult=0.1)
+ }))
+
+# learning policy
+max_epochs = 5
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[4],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs, val_interval=1)
+default_hooks = dict(checkpoint=dict(max_keep_ckpts=1, save_best='auto'))
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/coco/grounding_dino_swin-t_finetune_16xb4_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/coco/grounding_dino_swin-t_finetune_16xb4_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..792297accd302d390f865bee294b1294863d6ac1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/coco/grounding_dino_swin-t_finetune_16xb4_1x_coco.py
@@ -0,0 +1,85 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py'
+
+data_root = 'data/coco/'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities'))
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ _delete_=True,
+ type='CocoDataset',
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ return_classes=True,
+ filter_cfg=dict(filter_empty_gt=False, min_size=32),
+ pipeline=train_pipeline))
+
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0002, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'backbone': dict(lr_mult=0.1),
+ 'language_model': dict(lr_mult=0.1),
+ }))
+
+# learning policy
+max_epochs = 12
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs, val_interval=1)
+
+default_hooks = dict(checkpoint=dict(max_keep_ckpts=1, save_best='auto'))
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/coco/grounding_dino_swin-t_finetune_16xb4_1x_coco_48_17.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/coco/grounding_dino_swin-t_finetune_16xb4_1x_coco_48_17.py
new file mode 100644
index 0000000000000000000000000000000000000000..e68afbb43286af24612321129042e7d0e0f34b29
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/coco/grounding_dino_swin-t_finetune_16xb4_1x_coco_48_17.py
@@ -0,0 +1,157 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py'
+
+data_root = 'data/coco/'
+base_classes = ('person', 'bicycle', 'car', 'motorcycle', 'train', 'truck',
+ 'boat', 'bench', 'bird', 'horse', 'sheep', 'bear', 'zebra',
+ 'giraffe', 'backpack', 'handbag', 'suitcase', 'frisbee',
+ 'skis', 'kite', 'surfboard', 'bottle', 'fork', 'spoon', 'bowl',
+ 'banana', 'apple', 'sandwich', 'orange', 'broccoli', 'carrot',
+ 'pizza', 'donut', 'chair', 'bed', 'toilet', 'tv', 'laptop',
+ 'mouse', 'remote', 'microwave', 'oven', 'toaster',
+ 'refrigerator', 'book', 'clock', 'vase', 'toothbrush') # 48
+novel_classes = ('airplane', 'bus', 'cat', 'dog', 'cow', 'elephant',
+ 'umbrella', 'tie', 'snowboard', 'skateboard', 'cup', 'knife',
+ 'cake', 'couch', 'keyboard', 'sink', 'scissors') # 17
+all_classes = (
+ 'person', 'bicycle', 'car', 'motorcycle', 'airplane', 'bus', 'train',
+ 'truck', 'boat', 'bench', 'bird', 'cat', 'dog', 'horse', 'sheep', 'cow',
+ 'elephant', 'bear', 'zebra', 'giraffe', 'backpack', 'umbrella', 'handbag',
+ 'tie', 'suitcase', 'frisbee', 'skis', 'snowboard', 'kite', 'skateboard',
+ 'surfboard', 'bottle', 'cup', 'fork', 'knife', 'spoon', 'bowl', 'banana',
+ 'apple', 'sandwich', 'orange', 'broccoli', 'carrot', 'pizza', 'donut',
+ 'cake', 'chair', 'couch', 'bed', 'toilet', 'tv', 'laptop', 'mouse',
+ 'remote', 'keyboard', 'microwave', 'oven', 'toaster', 'sink',
+ 'refrigerator', 'book', 'clock', 'vase', 'scissors', 'toothbrush') # 65
+
+train_metainfo = dict(classes=base_classes)
+test_metainfo = dict(
+ classes=all_classes,
+ base_classes=base_classes,
+ novel_classes=novel_classes)
+
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities'))
+]
+
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile', backend_args=None,
+ imdecode_backend='pillow'),
+ dict(
+ type='FixScaleResize',
+ scale=(800, 1333),
+ keep_ratio=True,
+ backend='pillow'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'text', 'custom_entities',
+ 'tokens_positive'))
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ _delete_=True,
+ type='CocoDataset',
+ metainfo=train_metainfo,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017_seen_2.json',
+ data_prefix=dict(img='train2017/'),
+ return_classes=True,
+ filter_cfg=dict(filter_empty_gt=False, min_size=32),
+ pipeline=train_pipeline))
+
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type='CocoDataset',
+ metainfo=test_metainfo,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017_all_2.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ return_classes=True,
+ ))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='OVCocoMetric',
+ ann_file=data_root + 'annotations/instances_val2017_all_2.json',
+ metric='bbox',
+ format_only=False)
+test_evaluator = val_evaluator
+
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.00005, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'backbone': dict(lr_mult=0.1),
+ # 'language_model': dict(lr_mult=0),
+ }))
+
+# learning policy
+max_epochs = 12
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs, val_interval=1)
+
+default_hooks = dict(
+ checkpoint=dict(
+ max_keep_ckpts=1, save_best='coco/novel_ap50', rule='greater'))
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/coco/grounding_dino_swin-t_finetune_16xb4_1x_sft_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/coco/grounding_dino_swin-t_finetune_16xb4_1x_sft_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5505df58b8b103a93570519c20aaf0fcc144e91c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/coco/grounding_dino_swin-t_finetune_16xb4_1x_sft_coco.py
@@ -0,0 +1,93 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py'
+
+data_root = 'data/coco/'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=_base_.lang_model_name,
+ num_sample_negative=20, # ======= important =====
+ label_map_file='data/coco/annotations/coco2017_label_map.json',
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ _delete_=True,
+ type='ODVGDataset',
+ need_text=False,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017_od.json',
+ label_map_file='annotations/coco2017_label_map.json',
+ data_prefix=dict(img='train2017/'),
+ return_classes=True,
+ filter_cfg=dict(filter_empty_gt=False, min_size=32),
+ pipeline=train_pipeline))
+
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.00005, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'backbone': dict(lr_mult=0.1),
+ 'language_model': dict(lr_mult=0.0),
+ }))
+
+# learning policy
+max_epochs = 12
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs, val_interval=1)
+
+default_hooks = dict(checkpoint=dict(max_keep_ckpts=1, save_best='auto'))
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/dataset_prepare.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/dataset_prepare.md
new file mode 100644
index 0000000000000000000000000000000000000000..af60a8bf4bf7ebc0dde342a7a9ec0bd05dc1fadd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/dataset_prepare.md
@@ -0,0 +1,1193 @@
+# Data Prepare and Process
+
+## MM-GDINO-T Pre-train Dataset
+
+For the MM-GDINO-T model, we provide a total of 5 different data combination pre-training configurations. The data is trained in a progressive accumulation manner, so users can prepare it according to their actual needs.
+
+### 1 Objects365v1
+
+The corresponding training config is [grounding_dino_swin-t_pretrain_obj365](./grounding_dino_swin-t_pretrain_obj365.py)
+
+Objects365v1 can be downloaded from [opendatalab](https://opendatalab.com/OpenDataLab/Objects365_v1). It offers two methods of download: CLI and SDK.
+
+After downloading and unzipping, place the dataset or create a symbolic link to the `data/objects365v1` directory. The directory structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── objects365v1
+│ │ ├── objects365_train.json
+│ │ ├── objects365_val.json
+│ │ ├── train
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+│ │ ├── test
+```
+
+Then, use [coco2odvg.py](../../tools/dataset_converters/coco2odvg.py) to convert it into the ODVG format required for training.
+
+```shell
+python tools/dataset_converters/coco2odvg.py data/objects365v1/objects365_train.json -d o365v1
+```
+
+After the program runs successfully, it will create two new files, `o365v1_train_od.json` and `o365v1_label_map.json`, in the `data/objects365v1` directory. The complete structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── objects365v1
+│ │ ├── objects365_train.json
+│ │ ├── objects365_val.json
+│ │ ├── o365v1_train_od.json
+│ │ ├── o365v1_label_map.json
+│ │ ├── train
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+│ │ ├── test
+```
+
+### 2 COCO 2017
+
+The above configuration will evaluate the performance on the COCO 2017 dataset during the training process. Therefore, it is necessary to prepare the COCO 2017 dataset. You can download it from the [COCO](https://cocodataset.org/) official website or from [opendatalab](https://opendatalab.com/OpenDataLab/COCO_2017).
+
+After downloading and unzipping, place the dataset or create a symbolic link to the `data/coco` directory. The directory structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+### 3 GoldG
+
+After downloading the dataset, you can start training with the [grounding_dino_swin-t_pretrain_obj365_goldg](./grounding_dino_swin-t_pretrain_obj365_goldg.py) configuration.
+
+The GoldG dataset includes the `GQA` and `Flickr30k` datasets, which are part of the MixedGrounding dataset mentioned in the GLIP paper, excluding the COCO dataset. The download links are [mdetr_annotations](https://huggingface.co/GLIPModel/GLIP/tree/main/mdetr_annotations), and the specific files currently needed are `mdetr_annotations/final_mixed_train_no_coco.json` and `mdetr_annotations/final_flickr_separateGT_train.json`.
+
+Then download the [GQA images](https://nlp.stanford.edu/data/gqa/images.zip). After downloading and unzipping, place the dataset or create a symbolic link to them in the `data/gqa` directory, with the following directory structure:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── gqa
+| | ├── final_mixed_train_no_coco.json
+│ │ ├── images
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+Then download the [Flickr30k images](http://shannon.cs.illinois.edu/DenotationGraph/). You need to apply for access to this dataset and then download it using the provided link. After downloading and unzipping, place the dataset or create a symbolic link to them in the `data/flickr30k_entities` directory, with the following directory structure:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── flickr30k_entities
+│ │ ├── final_flickr_separateGT_train.json
+│ │ ├── flickr30k_images
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+For the GQA dataset, you need to use [goldg2odvg.py](../../tools/dataset_converters/goldg2odvg.py) to convert it into the ODVG format required for training:
+
+```shell
+python tools/dataset_converters/goldg2odvg.py data/gqa/final_mixed_train_no_coco.json
+```
+
+After the program has run, a new file `final_mixed_train_no_coco_vg.json` will be created in the `data/gqa` directory, with the complete structure as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── gqa
+| | ├── final_mixed_train_no_coco.json
+| | ├── final_mixed_train_no_coco_vg.json
+│ │ ├── images
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+For the Flickr30k dataset, you need to use [goldg2odvg.py](../../tools/dataset_converters/goldg2odvg.py) to convert it into the ODVG format required for training:
+
+```shell
+python tools/dataset_converters/goldg2odvg.py data/flickr30k_entities/final_flickr_separateGT_train.json
+```
+
+After the program has run, a new file `final_flickr_separateGT_train_vg.json` will be created in the `data/flickr30k_entities` directory, with the complete structure as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── flickr30k_entities
+│ │ ├── final_flickr_separateGT_train.json
+│ │ ├── final_flickr_separateGT_train_vg.json
+│ │ ├── flickr30k_images
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+### 4 GRIT-20M
+
+The corresponding training configuration is [grounding_dino_swin-t_pretrain_obj365_goldg_grit9m](./grounding_dino_swin-t_pretrain_obj365_goldg_grit9m.py).
+
+The GRIT dataset can be downloaded using the img2dataset package from [GRIT](https://huggingface.co/datasets/zzliang/GRIT#download-image). By default, the dataset size is 1.1T, and downloading and processing it may require at least 2T of disk space, depending on your available storage capacity. After downloading, the dataset is in its original format, which includes:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── grit_raw
+│ │ ├── 00000_stats.json
+│ │ ├── 00000.parquet
+│ │ ├── 00000.tar
+│ │ ├── 00001_stats.json
+│ │ ├── 00001.parquet
+│ │ ├── 00001.tar
+│ │ ├── ...
+```
+
+After downloading, further format processing is required:
+
+```shell
+python tools/dataset_converters/grit_processing.py data/grit_raw data/grit_processed
+```
+
+The processed format is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── grit_processed
+│ │ ├── annotations
+│ │ │ ├── 00000.json
+│ │ │ ├── 00001.json
+│ │ │ ├── ...
+│ │ ├── images
+│ │ │ ├── 00000
+│ │ │ │ ├── 000000000.jpg
+│ │ │ │ ├── 000000003.jpg
+│ │ │ │ ├── 000000004.jpg
+│ │ │ │ ├── ...
+│ │ │ ├── 00001
+│ │ │ ├── ...
+```
+
+As for the GRIT dataset, you need to use [grit2odvg.py](../../tools/dataset_converters/grit2odvg.py) to convert it to the format of ODVG:
+
+```shell
+python tools/dataset_converters/grit2odvg.py data/grit_processed/
+```
+
+After the program has run, a new file `grit20m_vg.json` will be created in the `data/grit_processed` directory, which has about 9M data, with the complete structure as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── grit_processed
+| | ├── grit20m_vg.json
+│ │ ├── annotations
+│ │ │ ├── 00000.json
+│ │ │ ├── 00001.json
+│ │ │ ├── ...
+│ │ ├── images
+│ │ │ ├── 00000
+│ │ │ │ ├── 000000000.jpg
+│ │ │ │ ├── 000000003.jpg
+│ │ │ │ ├── 000000004.jpg
+│ │ │ │ ├── ...
+│ │ │ ├── 00001
+│ │ │ ├── ...
+```
+
+### 5 V3Det
+
+The corresponding training configurations are:
+
+- [grounding_dino_swin-t_pretrain_obj365_goldg_v3det](./grounding_dino_swin-t_pretrain_obj365_goldg_v3det.py)
+- [grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det](./grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det.py)
+
+The V3Det dataset can be downloaded from [opendatalab](https://opendatalab.com/V3Det/V3Det). After downloading and unzipping, place the dataset or create a symbolic link to it in the `data/v3det` directory, with the following directory structure:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── v3det
+│ │ ├── annotations
+│ │ | ├── v3det_2023_v1_train.json
+│ │ ├── images
+│ │ │ ├── a00000066
+│ │ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+Then use [coco2odvg.py](../../tools/dataset_converters/coco2odvg.py) to convert it into the ODVG format required for training:
+
+```shell
+python tools/dataset_converters/coco2odvg.py data/v3det/annotations/v3det_2023_v1_train.json -d v3det
+```
+
+After the program has run, two new files `v3det_2023_v1_train_od.json` and `v3det_2023_v1_label_map.json` will be created in the `data/v3det/annotations` directory, with the complete structure as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── v3det
+│ │ ├── annotations
+│ │ | ├── v3det_2023_v1_train.json
+│ │ | ├── v3det_2023_v1_train_od.json
+│ │ | ├── v3det_2023_v1_label_map.json
+│ │ ├── images
+│ │ │ ├── a00000066
+│ │ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+### 6 Data Splitting and Visualization
+
+Considering that users need to prepare many datasets, which is inconvenient for confirming images and annotations before training, we provide a data splitting and visualization tool. This tool can split the dataset into a tiny version and then use a visualization script to check the correctness of the images and labels.
+
+1. Splitting the Dataset
+
+The script is located [here](../../tools/misc/split_odvg.py). Taking `Object365 v1` as an example, the command to split the dataset is as follows:
+
+```shell
+python tools/misc/split_odvg.py data/object365_v1/ o365v1_train_od.json train your_output_dir --label-map-file o365v1_label_map.json -n 200
+```
+
+After running the above script, it will create a folder structure in the `your_output_dir` directory identical to `data/object365_v1/`, but it will only save 200 training images and their corresponding json files for convenient user review.
+
+2. Visualizing the Original Dataset
+
+The script is located [here](../../tools/analysis_tools/browse_grounding_raw.py). Taking `Object365 v1` as an example, the command to visualize the dataset is as follows:
+
+```shell
+python tools/analysis_tools/browse_grounding_raw.py data/object365_v1/ o365v1_train_od.json train --label-map-file o365v1_label_map.json -o your_output_dir --not-show
+```
+
+After running the above script, it will generate images in the `your_output_dir` directory that include both the pictures and their labels, making it convenient for users to review.
+
+3. Visualizing the Output Dataset
+
+The script is located [here](../../tools/analysis_tools/browse_grounding_dataset.py). Users can use this script to view the results of the dataset output, including the results of data augmentation. Taking `Object365 v1` as an example, the command to visualize the dataset is as follows:
+
+```shell
+python tools/analysis_tools/browse_grounding_dataset.py configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py -o your_output_dir --not-show
+```
+
+After running the above script, it will generate images in the `your_output_dir` directory that include both the pictures and their labels, making it convenient for users to review.
+
+## MM-GDINO-L Pre-training Data Preparation and Processing
+
+### 1 Object365 v2
+
+Objects365_v2 can be downloaded from [opendatalab](https://opendatalab.com/OpenDataLab/Objects365). It offers two download methods: CLI and SDK.
+
+After downloading and unzipping, place the dataset or create a symbolic link to it in the `data/objects365v2` directory, with the following directory structure:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── objects365v2
+│ │ ├── annotations
+│ │ │ ├── zhiyuan_objv2_train.json
+│ │ ├── train
+│ │ │ ├── patch0
+│ │ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+Since some category names in Objects365v2 are incorrect, it is necessary to correct them first.
+
+```shell
+python tools/dataset_converters/fix_o365_names.py
+```
+
+A new annotation file `zhiyuan_objv2_train_fixname.json` will be generated in the `data/objects365v2/annotations` directory.
+
+Then use [coco2odvg.py](../../tools/dataset_converters/coco2odvg.py) to convert it into the ODVG format required for training:
+
+```shell
+python tools/dataset_converters/coco2odvg.py data/objects365v2/annotations/zhiyuan_objv2_train_fixname.json -d o365v2
+```
+
+After the program has run, two new files `zhiyuan_objv2_train_fixname_od.json` and `o365v2_label_map.json` will be created in the `data/objects365v2` directory, with the complete structure as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── objects365v2
+│ │ ├── annotations
+│ │ │ ├── zhiyuan_objv2_train.json
+│ │ │ ├── zhiyuan_objv2_train_fixname.json
+│ │ │ ├── zhiyuan_objv2_train_fixname_od.json
+│ │ │ ├── o365v2_label_map.json
+│ │ ├── train
+│ │ │ ├── patch0
+│ │ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+### 2 OpenImages v6
+
+OpenImages v6 can be downloaded from the [official website](https://storage.googleapis.com/openimages/web/download_v6.html). Due to the large size of the dataset, it may take some time to download. After completion, the file structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── OpenImages
+│ │ ├── annotations
+| │ │ ├── oidv6-train-annotations-bbox.csv
+| │ │ ├── class-descriptions-boxable.csv
+│ │ ├── OpenImages
+│ │ │ ├── train
+│ │ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+Then use [openimages2odvg.py](../../tools/dataset_converters/openimages2odvg.py) to convert it into the ODVG format required for training:
+
+```shell
+python tools/dataset_converters/openimages2odvg.py data/OpenImages/annotations
+```
+
+After the program has run, two new files `oidv6-train-annotation_od.json` and `openimages_label_map.json` will be created in the `data/OpenImages/annotations` directory, with the complete structure as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── OpenImages
+│ │ ├── annotations
+| │ │ ├── oidv6-train-annotations-bbox.csv
+| │ │ ├── class-descriptions-boxable.csv
+| │ │ ├── oidv6-train-annotations_od.json
+| │ │ ├── openimages_label_map.json
+│ │ ├── OpenImages
+│ │ │ ├── train
+│ │ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+### 3 V3Det
+
+Referring to the data preparation section of the previously mentioned MM-GDINO-T pre-training data preparation and processing, the complete dataset structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── v3det
+│ │ ├── annotations
+│ │ | ├── v3det_2023_v1_train.json
+│ │ | ├── v3det_2023_v1_train_od.json
+│ │ | ├── v3det_2023_v1_label_map.json
+│ │ ├── images
+│ │ │ ├── a00000066
+│ │ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+### 4 LVIS 1.0
+
+Please refer to the `2 LVIS 1.0` section of the later `Fine-tuning Dataset Preparation`. The complete dataset structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── lvis_v1_train.json
+│ │ │ ├── lvis_v1_val.json
+│ │ │ ├── lvis_v1_train_od.json
+│ │ │ ├── lvis_v1_label_map.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── lvis_v1_minival_inserted_image_name.json
+│ │ │ ├── lvis_od_val.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+### 5 COCO2017 OD
+
+You can refer to the earlier section `MM-GDINO-T Pre-training Data Preparation and Processing` for data preparation. For convenience in subsequent processing, please create a symbolic link or move the downloaded [mdetr_annotations](https://huggingface.co/GLIPModel/GLIP/tree/main/mdetr_annotations) folder to the `data/coco` path. The complete dataset structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ ├── mdetr_annotations
+│ │ │ ├── final_refexp_val.json
+│ │ │ ├── finetune_refcoco_testA.json
+│ │ │ ├── ...
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+Due to some overlap between COCO2017 train and RefCOCO/RefCOCO+/RefCOCOg/gRefCOCO val, if not removed in advance, there will be data leakage when evaluating RefExp.
+
+```shell
+python tools/dataset_converters/remove_cocotrain2017_from_refcoco.py data/coco/mdetr_annotations data/coco/annotations/instances_train2017.json
+```
+
+A new file `instances_train2017_norefval.json` will be created in the `data/coco/annotations` directory. Finally, use [coco2odvg.py](../../tools/dataset_converters/coco2odvg.py) to convert it into the ODVG format required for training:
+
+```shell
+python tools/dataset_converters/coco2odvg.py data/coco/annotations/instances_train2017_norefval.json -d coco
+```
+
+Two new files `instances_train2017_norefval_od.json` and `coco_label_map.json` will be created in the `data/coco/annotations` directory, with the complete structure as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── instances_train2017_norefval_od.json
+│ │ │ ├── coco_label_map.json
+│ │ ├── mdetr_annotations
+│ │ │ ├── final_refexp_val.json
+│ │ │ ├── finetune_refcoco_testA.json
+│ │ │ ├── ...
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+Note: There are 15,000 images that overlap between the COCO2017 train and LVIS 1.0 val datasets. Therefore, if the COCO2017 train dataset is used in training, the evaluation results of LVIS 1.0 val will have a data leakage issue. However, LVIS 1.0 minival does not have this problem.
+
+### 6 GoldG
+
+Please refer to the section on `MM-GDINO-T Pre-training Data Preparation and Processing`.
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── flickr30k_entities
+│ │ ├── final_flickr_separateGT_train.json
+│ │ ├── final_flickr_separateGT_train_vg.json
+│ │ ├── flickr30k_images
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ ├── gqa
+| | ├── final_mixed_train_no_coco.json
+| | ├── final_mixed_train_no_coco_vg.json
+│ │ ├── images
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+### 7 COCO2014 VG
+
+MDetr provides a Phrase Grounding version of the COCO2014 train annotations. The original annotation file is named `final_mixed_train.json`, and similar to the previous structure, the file structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ ├── mdetr_annotations
+│ │ │ ├── final_mixed_train.json
+│ │ │ ├── ...
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── train2014
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+We can extract the COCO portion of the data from `final_mixed_train.json`.
+
+```shell
+python tools/dataset_converters/extract_coco_from_mixed.py data/coco/mdetr_annotations/final_mixed_train.json
+```
+
+A new file named `final_mixed_train_only_coco.json` will be created in the `data/coco/mdetr_annotations` directory. Finally, use [goldg2odvg.py](../../tools/dataset_converters/goldg2odvg.py) to convert it into the ODVG format required for training:
+
+```shell
+python tools/dataset_converters/goldg2odvg.py data/coco/mdetr_annotations/final_mixed_train_only_coco.json
+```
+
+A new file named `final_mixed_train_only_coco_vg.json` will be created in the `data/coco/mdetr_annotations` directory, with the complete structure as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ ├── mdetr_annotations
+│ │ │ ├── final_mixed_train.json
+│ │ │ ├── final_mixed_train_only_coco.json
+│ │ │ ├── final_mixed_train_only_coco_vg.json
+│ │ │ ├── ...
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── train2014
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+Note: COCO2014 train and COCO2017 val do not have duplicate images, so there is no need to worry about data leakage issues in COCO evaluation.
+
+### 8 Referring Expression Comprehension
+
+There are a total of 4 datasets included. For data preparation, please refer to the `Fine-tuning Dataset Preparation` section.
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── instances_train2014.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+│ │ ├── train2014
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── mdetr_annotations
+│ │ │ ├── final_refexp_val.json
+│ │ │ ├── finetune_refcoco_testA.json
+│ │ │ ├── finetune_refcoco_testB.json
+│ │ │ ├── finetune_refcoco+_testA.json
+│ │ │ ├── finetune_refcoco+_testB.json
+│ │ │ ├── finetune_refcocog_test.json
+│ │ │ ├── finetune_refcoco_train_vg.json
+│ │ │ ├── finetune_refcoco+_train_vg.json
+│ │ │ ├── finetune_refcocog_train_vg.json
+│ │ │ ├── finetune_grefcoco_train_vg.json
+```
+
+### 9 GRIT-20M
+
+Please refer to the `MM-GDINO-T Pre-training Data Preparation and Processing` section.
+
+## Preparation of Evaluation Dataset
+
+### 1 COCO 2017
+
+The data preparation process is consistent with the previous descriptions, and the final structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+### 2 LVIS 1.0
+
+The LVIS 1.0 val dataset includes both mini and full versions. The significance of the mini version is:
+
+1. The full LVIS val evaluation dataset is quite large, and conducting an evaluation with it can take a significant amount of time.
+2. In the full LVIS val dataset, there are 15,000 images from the COCO2017 train dataset. If a user has used the COCO2017 data for training, there can be a data leakage issue when evaluating on the full LVIS val dataset
+
+The LVIS 1.0 dataset contains images that are exactly the same as the COCO2017 dataset, with the addition of new annotations. You can download the minival annotation file from [here](https://huggingface.co/GLIPModel/GLIP/blob/main/lvis_v1_minival_inserted_image_name.json), and the val 1.0 annotation file from [here](https://huggingface.co/GLIPModel/GLIP/blob/main/lvis_od_val.json). The final structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── lvis_v1_minival_inserted_image_name.json
+│ │ │ ├── lvis_od_val.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+### 3 ODinW
+
+ODinW, which stands for Object Detection in the Wild, is a dataset used to evaluate the generalization capability of grounding pre-trained models in different real-world scenarios. It consists of two subsets, ODinW13 and ODinW35, representing datasets composed of 13 and 35 different datasets, respectively. You can download it from [here](https://huggingface.co/GLIPModel/GLIP/tree/main/odinw_35), and then unzip each file. The final structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── odinw
+│ │ ├── AerialMaritimeDrone
+│ │ | |── large
+│ │ | | ├── test
+│ │ | | ├── train
+│ │ | | ├── valid
+│ │ | |── tiled
+│ │ ├── AmericanSignLanguageLetters
+│ │ ├── Aquarium
+│ │ ├── BCCD
+│ │ ├── ...
+```
+
+When evaluating ODinW35, custom prompts are required. Therefore, it's necessary to preprocess the annotated JSON files in advance. You can use the [override_category.py](./odinw/override_category.py) script for this purpose. After processing, it will generate new annotation files without overwriting the original ones.
+
+```shell
+python configs/mm_grounding_dino/odinw/override_category.py data/odinw/
+```
+
+### 4 DOD
+
+DOD stands for Described Object Detection, and it is introduced in the paper titled [Described Object Detection: Liberating Object Detection with Flexible Expressions](https://arxiv.org/abs/2307.12813). You can download the dataset from [here](https://github.com/shikras/d-cube?tab=readme-ov-file). The final structure of the dataset is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── d3
+│ │ ├── d3_images
+│ │ ├── d3_json
+│ │ ├── d3_pkl
+```
+
+### 5 Flickr30k Entities
+
+In the previous GoldG data preparation section, we downloaded the necessary files for training with Flickr30k. For evaluation, you will need 2 JSON files, which you can download from [here](https://huggingface.co/GLIPModel/GLIP/blob/main/mdetr_annotations/final_flickr_separateGT_val.json) and [here](https://huggingface.co/GLIPModel/GLIP/blob/main/mdetr_annotations/final_flickr_separateGT_test.json). The final structure of the dataset is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── flickr30k_entities
+│ │ ├── final_flickr_separateGT_train.json
+│ │ ├── final_flickr_separateGT_val.json
+│ │ ├── final_flickr_separateGT_test.json
+│ │ ├── final_flickr_separateGT_train_vg.json
+│ │ ├── flickr30k_images
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+### 6 Referring Expression Comprehension
+
+Referential Expression Comprehension includes 4 datasets: RefCOCO, RefCOCO+, RefCOCOg, and gRefCOCO. The images used in these 4 datasets are from COCO2014 train, similar to COCO2017. You can download the images from the official COCO website or opendatalab. The annotations can be directly downloaded from [here](https://huggingface.co/GLIPModel/GLIP/tree/main/mdetr_annotations). The mdetr_annotations folder contains a large number of annotations, so you can choose to download only the JSON files you need. The final structure of the dataset is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── instances_train2014.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+│ │ ├── train2014
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── mdetr_annotations
+│ │ │ ├── final_refexp_val.json
+│ │ │ ├── finetune_refcoco_testA.json
+│ │ │ ├── finetune_refcoco_testB.json
+│ │ │ ├── finetune_refcoco+_testA.json
+│ │ │ ├── finetune_refcoco+_testB.json
+│ │ │ ├── finetune_refcocog_test.json
+│ │ │ ├── finetune_refcocog_test.json
+```
+
+Please note that gRefCOCO is introduced in [GREC: Generalized Referring Expression Comprehension](https://arxiv.org/abs/2308.16182) and is not available in the `mdetr_annotations` folder. You will need to handle it separately. Here are the specific steps:
+
+1. Download [gRefCOCO](https://github.com/henghuiding/gRefCOCO?tab=readme-ov-file) and unzip it into the `data/coco/` folder.
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── instances_train2014.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+│ │ ├── train2014
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── mdetr_annotations
+│ │ ├── grefs
+│ │ │ ├── grefs(unc).json
+│ │ │ ├── instances.json
+```
+
+2. Convert to COCO format
+
+You can use the official [conversion script](https://github.com/henghuiding/gRefCOCO/blob/b4b1e55b4d3a41df26d6b7d843ea011d581127d4/mdetr/scripts/fine-tuning/grefexp_coco_format.py) provided by gRefCOCO. Please note that you need to uncomment line 161 and comment out line 160 in the script to obtain the full JSON file.
+
+```shell
+# you need to clone the official repo
+git clone https://github.com/henghuiding/gRefCOCO.git
+cd gRefCOCO/mdetr
+python scripts/fine-tuning/grefexp_coco_format.py --data_path ../../data/coco/grefs --out_path ../../data/coco/mdetr_annotations/ --coco_path ../../data/coco
+```
+
+Four JSON files will be generated in the `data/coco/mdetr_annotations/` folder. The complete dataset structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── instances_train2014.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+│ │ ├── train2014
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── mdetr_annotations
+│ │ │ ├── final_refexp_val.json
+│ │ │ ├── finetune_refcoco_testA.json
+│ │ │ ├── finetune_refcoco_testB.json
+│ │ │ ├── finetune_grefcoco_train.json
+│ │ │ ├── finetune_grefcoco_val.json
+│ │ │ ├── finetune_grefcoco_testA.json
+│ │ │ ├── finetune_grefcoco_testB.json
+```
+
+## Fine-Tuning Dataset Preparation
+
+### 1 COCO 2017
+
+COCO is the most commonly used dataset in the field of object detection, and we aim to explore its fine-tuning modes more comprehensively. From current developments, there are a total of three fine-tuning modes:
+
+1. Closed-set fine-tuning, where the description on the text side cannot be modified after fine-tuning, transforms into a closed-set algorithm. This approach maximizes performance on COCO but loses generality.
+2. Open-set continued pretraining fine-tuning involves using pretraining methods consistent with the COCO dataset. There are two approaches to this: the first is to reduce the learning rate and fix certain modules, fine-tuning only on the COCO dataset; the second is to mix COCO data with some of the pre-trained data. The goal of both approaches is to improve performance on the COCO dataset as much as possible without compromising generalization.
+3. Open-vocabulary fine-tuning involves adopting a common practice in the OVD (Open-Vocabulary Detection) domain. It divides COCO categories into base classes and novel classes. During training, fine-tuning is performed only on the base classes, while evaluation is conducted on both base and novel classes. This approach allows for the assessment of COCO OVD capabilities, with the goal of improving COCO dataset performance without compromising generalization as much as possible.
+
+\*\*(1) Closed-set Fine-tuning \*\*
+
+This section does not require data preparation; you can directly use the data you have prepared previously.
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+**(2) Open-set Continued Pretraining Fine-tuning**
+To use this approach, you need to convert the COCO training data into ODVG format. You can use the following command for conversion:
+
+```shell
+python tools/dataset_converters/coco2odvg.py data/coco/annotations/instances_train2017.json -d coco
+```
+
+This will generate new files, `instances_train2017_od.json` and `coco2017_label_map.json`, in the `data/coco/annotations/` directory. The complete dataset structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_train2017_od.json
+│ │ │ ├── coco2017_label_map.json
+│ │ │ ├── instances_val2017.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+Once you have obtained the data, you can choose whether to perform individual pretraining or mixed pretraining.
+
+**(3) Open-vocabulary Fine-tuning**
+For this approach, you need to convert the COCO training data into OVD (Open-Vocabulary Detection) format. You can use the following command for conversion:
+
+```shell
+python tools/dataset_converters/coco2ovd.py data/coco/
+```
+
+This will generate new files, `instances_val2017_all_2.json` and `instances_val2017_seen_2.json`, in the `data/coco/annotations/` directory. The complete dataset structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_train2017_od.json
+│ │ │ ├── instances_val2017_all_2.json
+│ │ │ ├── instances_val2017_seen_2.json
+│ │ │ ├── coco2017_label_map.json
+│ │ │ ├── instances_val2017.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+You can then proceed to train and test directly using the [configuration](coco/grounding_dino_swin-t_finetune_16xb4_1x_coco_48_17.py).
+
+### 2 LVIS 1.0
+
+LVIS is a dataset that includes 1,203 classes, making it a valuable dataset for fine-tuning. Due to its large number of classes, it's not feasible to perform closed-set fine-tuning. Therefore, we can only use open-set continued pretraining fine-tuning and open-vocabulary fine-tuning on LVIS.
+
+You need to prepare the LVIS training JSON files first, which you can download from [here](https://www.lvisdataset.org/dataset). We only need `lvis_v1_train.json` and `lvis_v1_val.json`. After downloading them, place them in the `data/coco/annotations/` directory, and then run the following command:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── lvis_v1_train.json
+│ │ │ ├── lvis_v1_val.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── lvis_v1_minival_inserted_image_name.json
+│ │ │ ├── lvis_od_val.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+(1) Open-set continued pretraining fine-tuning
+
+Convert to ODVG format using the following command:
+
+```shell
+python tools/dataset_converters/lvis2odvg.py data/coco/annotations/lvis_v1_train.json
+```
+
+It will generate new files, `lvis_v1_train_od.json` and `lvis_v1_label_map.json`, in the `data/coco/annotations/` directory, and the complete dataset structure will look like this:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── lvis_v1_train.json
+│ │ │ ├── lvis_v1_val.json
+│ │ │ ├── lvis_v1_train_od.json
+│ │ │ ├── lvis_v1_label_map.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── lvis_v1_minival_inserted_image_name.json
+│ │ │ ├── lvis_od_val.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+You can directly use the provided [configuration](lvis/grounding_dino_swin-t_finetune_16xb4_1x_lvis.py) for training and testing, or you can modify the configuration to mix it with some of the pretraining datasets as needed.
+
+**(2) Open Vocabulary Fine-tuning**
+
+Convert to OVD format using the following command:
+
+```shell
+python tools/dataset_converters/lvis2ovd.py data/coco/
+```
+
+New `lvis_v1_train_od_norare.json` and `lvis_v1_label_map_norare.json` will be generated under `data/coco/annotations/`, and the complete dataset structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── lvis_v1_train.json
+│ │ │ ├── lvis_v1_val.json
+│ │ │ ├── lvis_v1_train_od.json
+│ │ │ ├── lvis_v1_label_map.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── lvis_v1_minival_inserted_image_name.json
+│ │ │ ├── lvis_od_val.json
+│ │ │ ├── lvis_v1_train_od_norare.json
+│ │ │ ├── lvis_v1_label_map_norare.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+然Then you can directly use the [configuration](lvis/grounding_dino_swin-t_finetune_16xb4_1x_lvis_866_337.py) for training and testing.
+
+### 3 RTTS
+
+RTTS is a foggy weather dataset, which contains 4,322 foggy images, including five classes: bicycle, bus, car, motorbike, and person. It can be downloaded from [here](https://drive.google.com/file/d/15Ei1cHGVqR1mXFep43BO7nkHq1IEGh1e/view), and then extracted to the `data/RTTS/` folder. The complete dataset structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── RTTS
+│ │ ├── annotations_json
+│ │ ├── annotations_xml
+│ │ ├── ImageSets
+│ │ ├── JPEGImages
+```
+
+### 4 RUOD
+
+RUOD is an underwater object detection dataset. You can download it from [here](https://drive.google.com/file/d/1hxtbdgfVveUm_DJk5QXkNLokSCTa_E5o/view), and then extract it to the `data/RUOD/` folder. The complete dataset structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── RUOD
+│ │ ├── Environment_pic
+│ │ ├── Environmet_ANN
+│ │ ├── RUOD_ANN
+│ │ ├── RUOD_pic
+```
+
+### 5 Brain Tumor
+
+Brain Tumor is a 2D detection dataset in the medical field. You can download it from [here](https://universe.roboflow.com/roboflow-100/brain-tumor-m2pbp/dataset/2), please make sure to choose the `COCO JSON` format. Then extract it to the `data/brain_tumor_v2/` folder. The complete dataset structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── brain_tumor_v2
+│ │ ├── test
+│ │ ├── train
+│ │ ├── valid
+```
+
+### 6 Cityscapes
+
+Cityscapes is an urban street scene dataset. You can download it from [here](https://www.cityscapes-dataset.com/) or from opendatalab, and then extract it to the `data/cityscapes/` folder. The complete dataset structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── cityscapes
+│ │ ├── annotations
+│ │ ├── leftImg8bit
+│ │ │ ├── train
+│ │ │ ├── val
+│ │ ├── gtFine
+│ │ │ ├── train
+│ │ │ ├── val
+```
+
+After downloading, you can use the [cityscapes.py](../../tools/dataset_converters/cityscapes.py) script to generate the required JSON format.
+
+```shell
+python tools/dataset_converters/cityscapes.py data/cityscapes/
+```
+
+Three new JSON files will be generated in the annotations directory. The complete dataset structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── cityscapes
+│ │ ├── annotations
+│ │ │ ├── instancesonly_filtered_gtFine_train.json
+│ │ │ ├── instancesonly_filtered_gtFine_val.json
+│ │ │ ├── instancesonly_filtered_gtFine_test.json
+│ │ ├── leftImg8bit
+│ │ │ ├── train
+│ │ │ ├── val
+│ │ ├── gtFine
+│ │ │ ├── train
+│ │ │ ├── val
+```
+
+### 7 People in Painting
+
+People in Painting is an oil painting dataset that you can download from [here](https://universe.roboflow.com/roboflow-100/people-in-paintings/dataset/2). Please make sure to choose the `COCO JSON` format. After downloading, unzip the dataset to the `data/people_in_painting_v2/` folder. The complete dataset structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── people_in_painting_v2
+│ │ ├── test
+│ │ ├── train
+│ │ ├── valid
+```
+
+### 8 Referring Expression Comprehension
+
+Fine-tuning for Referential Expression Comprehension is similar to what was described earlier and includes four datasets. The dataset preparation for evaluation has already been organized. The complete dataset structure is as follows:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── instances_train2014.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+│ │ ├── train2014
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── mdetr_annotations
+│ │ │ ├── final_refexp_val.json
+│ │ │ ├── finetune_refcoco_testA.json
+│ │ │ ├── finetune_refcoco_testB.json
+│ │ │ ├── finetune_refcoco+_testA.json
+│ │ │ ├── finetune_refcoco+_testB.json
+│ │ │ ├── finetune_refcocog_test.json
+│ │ │ ├── finetune_refcocog_test.json
+```
+
+Then we need to convert it to the required ODVG format. Please use the [refcoco2odvg.py](../../tools/dataset_converters/refcoco2odvg.py) script to perform the conversion.
+
+```shell
+python tools/dataset_converters/refcoco2odvg.py data/coco/mdetr_annotations
+```
+
+The converted dataset structure will include 4 new JSON files in the `data/coco/mdetr_annotations` directory. Here is the structure of the converted dataset:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── instances_train2014.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+│ │ ├── train2014
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── mdetr_annotations
+│ │ │ ├── final_refexp_val.json
+│ │ │ ├── finetune_refcoco_testA.json
+│ │ │ ├── finetune_refcoco_testB.json
+│ │ │ ├── finetune_refcoco+_testA.json
+│ │ │ ├── finetune_refcoco+_testB.json
+│ │ │ ├── finetune_refcocog_test.json
+│ │ │ ├── finetune_refcoco_train_vg.json
+│ │ │ ├── finetune_refcoco+_train_vg.json
+│ │ │ ├── finetune_refcocog_train_vg.json
+│ │ │ ├── finetune_grefcoco_train_vg.json
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/dataset_prepare_zh-CN.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/dataset_prepare_zh-CN.md
new file mode 100644
index 0000000000000000000000000000000000000000..10520b02fe54cda845335b55ac5bc6fa8bfdac65
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/dataset_prepare_zh-CN.md
@@ -0,0 +1,1194 @@
+# 数据准备和处理
+
+## MM-GDINO-T 预训练数据准备和处理
+
+MM-GDINO-T 模型中我们一共提供了 5 种不同数据组合的预训练配置,数据采用逐步累加的方式进行训练,因此用户可以根据自己的实际需求准备数据。
+
+### 1 Objects365 v1
+
+对应的训练配置为 [grounding_dino_swin-t_pretrain_obj365](./grounding_dino_swin-t_pretrain_obj365.py)
+
+Objects365_v1 可以从 [opendatalab](https://opendatalab.com/OpenDataLab/Objects365_v1) 下载,其提供了 CLI 和 SDK 两者下载方式。
+
+下载并解压后,将其放置或者软链接到 `data/objects365v1` 目录下,目录结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── objects365v1
+│ │ ├── objects365_train.json
+│ │ ├── objects365_val.json
+│ │ ├── train
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+│ │ ├── test
+```
+
+然后使用 [coco2odvg.py](../../tools/dataset_converters/coco2odvg.py) 转换为训练所需的 ODVG 格式:
+
+```shell
+python tools/dataset_converters/coco2odvg.py data/objects365v1/objects365_train.json -d o365v1
+```
+
+程序运行完成后会在 `data/objects365v1` 目录下创建 `o365v1_train_od.json` 和 `o365v1_label_map.json` 两个新文件,完整结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── objects365v1
+│ │ ├── objects365_train.json
+│ │ ├── objects365_val.json
+│ │ ├── o365v1_train_od.json
+│ │ ├── o365v1_label_map.json
+│ │ ├── train
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+│ │ ├── test
+```
+
+### 2 COCO 2017
+
+上述配置在训练过程中会评估 COCO 2017 数据集的性能,因此需要准备 COCO 2017 数据集。你可以从 [COCO](https://cocodataset.org/) 官网下载或者从 [opendatalab](https://opendatalab.com/OpenDataLab/COCO_2017) 下载
+
+下载并解压后,将其放置或者软链接到 `data/coco` 目录下,目录结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+### 3 GoldG
+
+下载该数据集后就可以训练 [grounding_dino_swin-t_pretrain_obj365_goldg](./grounding_dino_swin-t_pretrain_obj365_goldg.py) 配置了。
+
+GoldG 数据集包括 `GQA` 和 `Flickr30k` 两个数据集,来自 GLIP 论文中提到的 MixedGrounding 数据集,其排除了 COCO 数据集。下载链接为 [mdetr_annotations](https://huggingface.co/GLIPModel/GLIP/tree/main/mdetr_annotations),我们目前需要的是 `mdetr_annotations/final_mixed_train_no_coco.json` 和 `mdetr_annotations/final_flickr_separateGT_train.json` 文件。
+
+然后下载 [GQA images](https://nlp.stanford.edu/data/gqa/images.zip) 图片。下载并解压后,将其放置或者软链接到 `data/gqa` 目录下,目录结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── gqa
+| | ├── final_mixed_train_no_coco.json
+│ │ ├── images
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+然后下载 [Flickr30k images](http://shannon.cs.illinois.edu/DenotationGraph/) 图片。这个数据下载需要先申请,再获得下载链接后才可以下载。下载并解压后,将其放置或者软链接到 `data/flickr30k_entities` 目录下,目录结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── flickr30k_entities
+│ │ ├── final_flickr_separateGT_train.json
+│ │ ├── flickr30k_images
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+对于 GQA 数据集,你需要使用 [goldg2odvg.py](../../tools/dataset_converters/goldg2odvg.py) 转换为训练所需的 ODVG 格式:
+
+```shell
+python tools/dataset_converters/goldg2odvg.py data/gqa/final_mixed_train_no_coco.json
+```
+
+程序运行完成后会在 `data/gqa` 目录下创建 `final_mixed_train_no_coco_vg.json` 新文件,完整结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── gqa
+| | ├── final_mixed_train_no_coco.json
+| | ├── final_mixed_train_no_coco_vg.json
+│ │ ├── images
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+对于 Flickr30k 数据集,你需要使用 [goldg2odvg.py](../../tools/dataset_converters/goldg2odvg.py) 转换为训练所需的 ODVG 格式:
+
+```shell
+python tools/dataset_converters/goldg2odvg.py data/flickr30k_entities/final_flickr_separateGT_train.json
+```
+
+程序运行完成后会在 `data/flickr30k_entities` 目录下创建 `final_flickr_separateGT_train_vg.json` 新文件,完整结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── flickr30k_entities
+│ │ ├── final_flickr_separateGT_train.json
+│ │ ├── final_flickr_separateGT_train_vg.json
+│ │ ├── flickr30k_images
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+### 4 GRIT-20M
+
+对应的训练配置为 [grounding_dino_swin-t_pretrain_obj365_goldg_grit9m](./grounding_dino_swin-t_pretrain_obj365_goldg_grit9m.py)
+
+GRIT数据集可以从 [GRIT](https://huggingface.co/datasets/zzliang/GRIT#download-image) 中使用 img2dataset 包下载,默认指令下载后数据集大小为 1.1T,下载和处理预估需要至少 2T 硬盘空间,可根据硬盘容量酌情下载。下载后原始格式为:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── grit_raw
+│ │ ├── 00000_stats.json
+│ │ ├── 00000.parquet
+│ │ ├── 00000.tar
+│ │ ├── 00001_stats.json
+│ │ ├── 00001.parquet
+│ │ ├── 00001.tar
+│ │ ├── ...
+```
+
+下载后需要对格式进行进一步处理:
+
+```shell
+python tools/dataset_converters/grit_processing.py data/grit_raw data/grit_processed
+```
+
+处理后的格式为:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── grit_processed
+│ │ ├── annotations
+│ │ │ ├── 00000.json
+│ │ │ ├── 00001.json
+│ │ │ ├── ...
+│ │ ├── images
+│ │ │ ├── 00000
+│ │ │ │ ├── 000000000.jpg
+│ │ │ │ ├── 000000003.jpg
+│ │ │ │ ├── 000000004.jpg
+│ │ │ │ ├── ...
+│ │ │ ├── 00001
+│ │ │ ├── ...
+```
+
+对于 GRIT 数据集,你需要使用 [grit2odvg.py](../../tools/dataset_converters/grit2odvg.py) 转化成需要的 ODVG 格式:
+
+```shell
+python tools/dataset_converters/grit2odvg.py data/grit_processed/
+```
+
+程序运行完成后会在 `data/grit_processed` 目录下创建 `grit20m_vg.json` 新文件,大概包含 9M 条数据,完整结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── grit_processed
+| | ├── grit20m_vg.json
+│ │ ├── annotations
+│ │ │ ├── 00000.json
+│ │ │ ├── 00001.json
+│ │ │ ├── ...
+│ │ ├── images
+│ │ │ ├── 00000
+│ │ │ │ ├── 000000000.jpg
+│ │ │ │ ├── 000000003.jpg
+│ │ │ │ ├── 000000004.jpg
+│ │ │ │ ├── ...
+│ │ │ ├── 00001
+│ │ │ ├── ...
+```
+
+### 5 V3Det
+
+对应的训练配置为
+
+- [grounding_dino_swin-t_pretrain_obj365_goldg_v3det](./grounding_dino_swin-t_pretrain_obj365_goldg_v3det.py)
+- [grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det](./grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det.py)
+
+V3Det 数据集下载可以从 [opendatalab](https://opendatalab.com/V3Det/V3Det) 下载,下载并解压后,将其放置或者软链接到 `data/v3det` 目录下,目录结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── v3det
+│ │ ├── annotations
+│ │ | ├── v3det_2023_v1_train.json
+│ │ ├── images
+│ │ │ ├── a00000066
+│ │ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+然后使用 [coco2odvg.py](../../tools/dataset_converters/coco2odvg.py) 转换为训练所需的 ODVG 格式:
+
+```shell
+python tools/dataset_converters/coco2odvg.py data/v3det/annotations/v3det_2023_v1_train.json -d v3det
+```
+
+程序运行完成后会在 `data/v3det/annotations` 目录下创建目录下创建 `v3det_2023_v1_train_od.json` 和 `v3det_2023_v1_label_map.json` 两个新文件,完整结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── v3det
+│ │ ├── annotations
+│ │ | ├── v3det_2023_v1_train.json
+│ │ | ├── v3det_2023_v1_train_od.json
+│ │ | ├── v3det_2023_v1_label_map.json
+│ │ ├── images
+│ │ │ ├── a00000066
+│ │ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+### 6 数据切分和可视化
+
+考虑到用户需要准备的数据集过多,不方便对图片和标注进行训练前确认,因此我们提供了一个数据切分和可视化的工具,可以将数据集切分为 tiny 版本,然后使用可视化脚本查看图片和标签正确性。
+
+1. 切分数据集
+
+脚本位于 [这里](../../tools/misc/split_odvg.py), 以 `Object365 v1` 为例,切分数据集的命令如下:
+
+```shell
+python tools/misc/split_odvg.py data/object365_v1/ o365v1_train_od.json train your_output_dir --label-map-file o365v1_label_map.json -n 200
+```
+
+上述脚本运行后会在 `your_output_dir` 目录下创建和 `data/object365_v1/` 一样的文件夹结构,但是只会保存 200 张训练图片和对应的 json,方便用户查看。
+
+2. 可视化原始数据集
+
+脚本位于 [这里](../../tools/analysis_tools/browse_grounding_raw.py), 以 `Object365 v1` 为例,可视化数据集的命令如下:
+
+```shell
+python tools/analysis_tools/browse_grounding_raw.py data/object365_v1/ o365v1_train_od.json train --label-map-file o365v1_label_map.json -o your_output_dir --not-show
+```
+
+上述脚本运行后会在 `your_output_dir` 目录下生成同时包括图片和标签的图片,方便用户查看。
+
+3. 可视化 dataset 输出的数据集
+
+脚本位于 [这里](../../tools/analysis_tools/browse_grounding_dataset.py), 用户可以通过该脚本查看 dataset 输出的结果即包括了数据增强的结果。 以 `Object365 v1` 为例,可视化数据集的命令如下:
+
+```shell
+python tools/analysis_tools/browse_grounding_dataset.py configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py -o your_output_dir --not-show
+```
+
+上述脚本运行后会在 `your_output_dir` 目录下生成同时包括图片和标签的图片,方便用户查看。
+
+## MM-GDINO-L 预训练数据准备和处理
+
+### 1 Object365 v2
+
+Objects365_v2 可以从 [opendatalab](https://opendatalab.com/OpenDataLab/Objects365) 下载,其提供了 CLI 和 SDK 两者下载方式。
+
+下载并解压后,将其放置或者软链接到 `data/objects365v2` 目录下,目录结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── objects365v2
+│ │ ├── annotations
+│ │ │ ├── zhiyuan_objv2_train.json
+│ │ ├── train
+│ │ │ ├── patch0
+│ │ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+由于 objects365v2 类别中有部分类名是错误的,因此需要先进行修正。
+
+```shell
+python tools/dataset_converters/fix_o365_names.py
+```
+
+会在 `data/objects365v2/annotations` 下生成新的标注文件 `zhiyuan_objv2_train_fixname.json`。
+
+然后使用 [coco2odvg.py](../../tools/dataset_converters/coco2odvg.py) 转换为训练所需的 ODVG 格式:
+
+```shell
+python tools/dataset_converters/coco2odvg.py data/objects365v2/annotations/zhiyuan_objv2_train_fixname.json -d o365v2
+```
+
+程序运行完成后会在 `data/objects365v2` 目录下创建 `zhiyuan_objv2_train_fixname_od.json` 和 `o365v2_label_map.json` 两个新文件,完整结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── objects365v2
+│ │ ├── annotations
+│ │ │ ├── zhiyuan_objv2_train.json
+│ │ │ ├── zhiyuan_objv2_train_fixname.json
+│ │ │ ├── zhiyuan_objv2_train_fixname_od.json
+│ │ │ ├── o365v2_label_map.json
+│ │ ├── train
+│ │ │ ├── patch0
+│ │ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+### 2 OpenImages v6
+
+OpenImages v6 可以从 [官网](https://storage.googleapis.com/openimages/web/download_v6.html) 下载,由于数据集比较大,需要花费一定的时间,下载完成后文件结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── OpenImages
+│ │ ├── annotations
+| │ │ ├── oidv6-train-annotations-bbox.csv
+| │ │ ├── class-descriptions-boxable.csv
+│ │ ├── OpenImages
+│ │ │ ├── train
+│ │ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+然后使用 [openimages2odvg.py](../../tools/dataset_converters/openimages2odvg.py) 转换为训练所需的 ODVG 格式:
+
+```shell
+python tools/dataset_converters/openimages2odvg.py data/OpenImages/annotations
+```
+
+程序运行完成后会在 `data/OpenImages/annotations` 目录下创建 `oidv6-train-annotation_od.json` 和 `openimages_label_map.json` 两个新文件,完整结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── OpenImages
+│ │ ├── annotations
+| │ │ ├── oidv6-train-annotations-bbox.csv
+| │ │ ├── class-descriptions-boxable.csv
+| │ │ ├── oidv6-train-annotations_od.json
+| │ │ ├── openimages_label_map.json
+│ │ ├── OpenImages
+│ │ │ ├── train
+│ │ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+### 3 V3Det
+
+参见前面的 MM-GDINO-T 预训练数据准备和处理 数据准备部分,完整数据集结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── v3det
+│ │ ├── annotations
+│ │ | ├── v3det_2023_v1_train.json
+│ │ | ├── v3det_2023_v1_train_od.json
+│ │ | ├── v3det_2023_v1_label_map.json
+│ │ ├── images
+│ │ │ ├── a00000066
+│ │ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+### 4 LVIS 1.0
+
+参见后面的 `微调数据集准备` 的 `2 LVIS 1.0` 部分。完整数据集结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── lvis_v1_train.json
+│ │ │ ├── lvis_v1_val.json
+│ │ │ ├── lvis_v1_train_od.json
+│ │ │ ├── lvis_v1_label_map.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── lvis_v1_minival_inserted_image_name.json
+│ │ │ ├── lvis_od_val.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+### 5 COCO2017 OD
+
+数据准备可以参考前面的 `MM-GDINO-T 预训练数据准备和处理` 部分。为了方便后续处理,请将下载的 [mdetr_annotations](https://huggingface.co/GLIPModel/GLIP/tree/main/mdetr_annotations) 文件夹软链接或者移动到 `data/coco` 路径下
+完整数据集结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ ├── mdetr_annotations
+│ │ │ ├── final_refexp_val.json
+│ │ │ ├── finetune_refcoco_testA.json
+│ │ │ ├── ...
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+由于 COCO2017 train 和 RefCOCO/RefCOCO+/RefCOCOg/gRefCOCO val 中存在部分重叠,如果不提前移除,在评测 RefExp 时候会存在数据泄露。
+
+```shell
+python tools/dataset_converters/remove_cocotrain2017_from_refcoco.py data/coco/mdetr_annotations data/coco/annotations/instances_train2017.json
+```
+
+会在 `data/coco/annotations` 目录下创建 `instances_train2017_norefval.json` 新文件。最后使用 [coco2odvg.py](../../tools/dataset_converters/coco2odvg.py) 转换为训练所需的 ODVG 格式:
+
+```shell
+python tools/dataset_converters/coco2odvg.py data/coco/annotations/instances_train2017_norefval.json -d coco
+```
+
+会在 `data/coco/annotations` 目录下创建 `instances_train2017_norefval_od.json` 和 `coco_label_map.json` 两个新文件,完整结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── instances_train2017_norefval_od.json
+│ │ │ ├── coco_label_map.json
+│ │ ├── mdetr_annotations
+│ │ │ ├── final_refexp_val.json
+│ │ │ ├── finetune_refcoco_testA.json
+│ │ │ ├── ...
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+注意: COCO2017 train 和 LVIS 1.0 val 数据集有 15000 张图片重复,因此一旦在训练中使用了 COCO2017 train,那么 LVIS 1.0 val 的评测结果就存在数据泄露问题,LVIS 1.0 minival 没有这个问题。
+
+### 6 GoldG
+
+参见 MM-GDINO-T 预训练数据准备和处理 部分
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── flickr30k_entities
+│ │ ├── final_flickr_separateGT_train.json
+│ │ ├── final_flickr_separateGT_train_vg.json
+│ │ ├── flickr30k_images
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ ├── gqa
+| | ├── final_mixed_train_no_coco.json
+| | ├── final_mixed_train_no_coco_vg.json
+│ │ ├── images
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+### 7 COCO2014 VG
+
+MDetr 中提供了 COCO2014 train 的 Phrase Grounding 版本标注, 最原始标注文件为 `final_mixed_train.json`,和之前类似,文件结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ ├── mdetr_annotations
+│ │ │ ├── final_mixed_train.json
+│ │ │ ├── ...
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── train2014
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+我们可以从 `final_mixed_train.json` 中提取出 COCO 部分数据
+
+```shell
+python tools/dataset_converters/extract_coco_from_mixed.py data/coco/mdetr_annotations/final_mixed_train.json
+```
+
+会在 `data/coco/mdetr_annotations` 目录下创建 `final_mixed_train_only_coco.json` 新文件,最后使用 [goldg2odvg.py](../../tools/dataset_converters/goldg2odvg.py) 转换为训练所需的 ODVG 格式:
+
+```shell
+python tools/dataset_converters/goldg2odvg.py data/coco/mdetr_annotations/final_mixed_train_only_coco.json
+```
+
+会在 `data/coco/mdetr_annotations` 目录下创建 `final_mixed_train_only_coco_vg.json` 新文件,完整结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ ├── mdetr_annotations
+│ │ │ ├── final_mixed_train.json
+│ │ │ ├── final_mixed_train_only_coco.json
+│ │ │ ├── final_mixed_train_only_coco_vg.json
+│ │ │ ├── ...
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── train2014
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+注意: COCO2014 train 和 COCO2017 val 没有重复图片,因此不用担心 COCO 评测的数据泄露问题。
+
+### 8 Referring Expression Comprehension
+
+其一共包括 4 个数据集。数据准备部分请参见 微调数据集准备 部分。
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── instances_train2014.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+│ │ ├── train2014
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── mdetr_annotations
+│ │ │ ├── final_refexp_val.json
+│ │ │ ├── finetune_refcoco_testA.json
+│ │ │ ├── finetune_refcoco_testB.json
+│ │ │ ├── finetune_refcoco+_testA.json
+│ │ │ ├── finetune_refcoco+_testB.json
+│ │ │ ├── finetune_refcocog_test.json
+│ │ │ ├── finetune_refcoco_train_vg.json
+│ │ │ ├── finetune_refcoco+_train_vg.json
+│ │ │ ├── finetune_refcocog_train_vg.json
+│ │ │ ├── finetune_grefcoco_train_vg.json
+```
+
+### 9 GRIT-20M
+
+参见 MM-GDINO-T 预训练数据准备和处理 部分
+
+## 评测数据集准备
+
+### 1 COCO 2017
+
+数据准备流程和前面描述一致,最终结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+### 2 LVIS 1.0
+
+LVIS 1.0 val 数据集包括 mini 和全量两个版本,mini 版本存在的意义是:
+
+1. LVIS val 全量评测数据集比较大,评测一次需要比较久的时间
+2. LVIS val 全量数据集中包括了 15000 张 COCO2017 train, 如果用户使用了 COCO2017 数据进行训练,那么将存在数据泄露问题
+
+LVIS 1.0 图片和 COCO2017 数据集图片完全一样,只是提供了新的标注而已,minival 标注文件可以从 [这里](https://huggingface.co/GLIPModel/GLIP/blob/main/lvis_v1_minival_inserted_image_name.json)下载, val 1.0 标注文件可以从 [这里](https://huggingface.co/GLIPModel/GLIP/blob/main/lvis_od_val.json) 下载。 最终结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── lvis_v1_minival_inserted_image_name.json
+│ │ │ ├── lvis_od_val.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+### 3 ODinW
+
+ODinw 全称为 Object Detection in the Wild,是用于验证 grounding 预训练模型在不同实际场景中的泛化能力的数据集,其包括两个子集,分别是 ODinW13 和 ODinW35,代表是由 13 和 35 个数据集组成的。你可以从 [这里](https://huggingface.co/GLIPModel/GLIP/tree/main/odinw_35)下载,然后对每个文件进行解压,最终结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── odinw
+│ │ ├── AerialMaritimeDrone
+│ │ | |── large
+│ │ | | ├── test
+│ │ | | ├── train
+│ │ | | ├── valid
+│ │ | |── tiled
+│ │ ├── AmericanSignLanguageLetters
+│ │ ├── Aquarium
+│ │ ├── BCCD
+│ │ ├── ...
+```
+
+在评测 ODinW3535 时候由于需要自定义 prompt,因此需要提前对标注的 json 文件进行处理,你可以使用 [override_category.py](./odinw/override_category.py) 脚本进行处理,处理后会生成新的标注文件,不会覆盖原先的标注文件。
+
+```shell
+python configs/mm_grounding_dino/odinw/override_category.py data/odinw/
+```
+
+### 4 DOD
+
+DOD 来自 [Described Object Detection: Liberating Object Detection with Flexible Expressions](https://arxiv.org/abs/2307.12813)。其数据集可以从 [这里](https://github.com/shikras/d-cube?tab=readme-ov-file#download)下载,最终的数据集结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── d3
+│ │ ├── d3_images
+│ │ ├── d3_json
+│ │ ├── d3_pkl
+```
+
+### 5 Flickr30k Entities
+
+在前面 GoldG 数据准备章节中我们已经下载了 Flickr30k 训练所需文件,评估所需的文件是 2 个 json 文件,你可以从 [这里](https://huggingface.co/GLIPModel/GLIP/blob/main/mdetr_annotations/final_flickr_separateGT_val.json) 和 [这里](https://huggingface.co/GLIPModel/GLIP/blob/main/mdetr_annotations/final_flickr_separateGT_test.json)下载,最终的数据集结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── flickr30k_entities
+│ │ ├── final_flickr_separateGT_train.json
+│ │ ├── final_flickr_separateGT_val.json
+│ │ ├── final_flickr_separateGT_test.json
+│ │ ├── final_flickr_separateGT_train_vg.json
+│ │ ├── flickr30k_images
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+```
+
+### 6 Referring Expression Comprehension
+
+指代性表达式理解包括 4 个数据集: RefCOCO, RefCOCO+, RefCOCOg, gRefCOCO。这 4 个数据集所采用的图片都来自于 COCO2014 train,和 COCO2017 类似,你可以从 COCO 官方或者 opendatalab 中下载,而标注可以直接从 [这里](https://huggingface.co/GLIPModel/GLIP/tree/main/mdetr_annotations) 下载,mdetr_annotations 文件夹里面包括了其他大量的标注,你如果觉得数量过多,可以只下载所需要的几个 json 文件即可。最终的数据集结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── instances_train2014.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+│ │ ├── train2014
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── mdetr_annotations
+│ │ │ ├── final_refexp_val.json
+│ │ │ ├── finetune_refcoco_testA.json
+│ │ │ ├── finetune_refcoco_testB.json
+│ │ │ ├── finetune_refcoco+_testA.json
+│ │ │ ├── finetune_refcoco+_testB.json
+│ │ │ ├── finetune_refcocog_test.json
+│ │ │ ├── finetune_refcocog_test.json
+```
+
+注意 gRefCOCO 是在 [GREC: Generalized Referring Expression Comprehension](https://arxiv.org/abs/2308.16182) 被提出,并不在 `mdetr_annotations` 文件夹中,需要自行处理。具体步骤为:
+
+1. 下载 [gRefCOCO](https://github.com/henghuiding/gRefCOCO?tab=readme-ov-file#grefcoco-dataset-download),并解压到 data/coco/ 文件夹中
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── instances_train2014.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+│ │ ├── train2014
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── mdetr_annotations
+│ │ ├── grefs
+│ │ │ ├── grefs(unc).json
+│ │ │ ├── instances.json
+```
+
+2. 转换为 coco 格式
+
+你可以使用 gRefCOCO 官方提供的[转换脚本](https://github.com/henghuiding/gRefCOCO/blob/b4b1e55b4d3a41df26d6b7d843ea011d581127d4/mdetr/scripts/fine-tuning/grefexp_coco_format.py)。注意需要将被注释的 161 行打开,并注释 160 行才可以得到全量的 json 文件。
+
+```shell
+# 需要克隆官方 repo
+git clone https://github.com/henghuiding/gRefCOCO.git
+cd gRefCOCO/mdetr
+python scripts/fine-tuning/grefexp_coco_format.py --data_path ../../data/coco/grefs --out_path ../../data/coco/mdetr_annotations/ --coco_path ../../data/coco
+```
+
+会在 `data/coco/mdetr_annotations/` 文件夹中生成 4 个 json 文件,完整的数据集结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── instances_train2014.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+│ │ ├── train2014
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── mdetr_annotations
+│ │ │ ├── final_refexp_val.json
+│ │ │ ├── finetune_refcoco_testA.json
+│ │ │ ├── finetune_refcoco_testB.json
+│ │ │ ├── finetune_grefcoco_train.json
+│ │ │ ├── finetune_grefcoco_val.json
+│ │ │ ├── finetune_grefcoco_testA.json
+│ │ │ ├── finetune_grefcoco_testB.json
+```
+
+## 微调数据集准备
+
+### 1 COCO 2017
+
+COCO 是检测领域最常用的数据集,我们希望能够更充分探索其微调模式。从目前发展来看,一共有 3 种微调方式:
+
+1. 闭集微调,即微调后文本端将无法修改描述,转变为闭集算法,在 COCO 上性能能够最大化,但是失去了通用性。
+2. 开集继续预训练微调,即对 COCO 数据集采用和预训练一致的预训练手段。此时有两种做法,第一种是降低学习率并固定某些模块,仅仅在 COCO 数据上预训练,第二种是将 COCO 数据和部分预训练数据混合一起训练,两种方式的目的都是在尽可能不降低泛化性时提高 COCO 数据集性能
+3. 开放词汇微调,即采用 OVD 领域常用做法,将 COCO 类别分成 base 类和 novel 类,训练时候仅仅在 base 类上进行,评测在 base 和 novel 类上进行。这种方式可以验证 COCO OVD 能力,目的也是在尽可能不降低泛化性时提高 COCO 数据集性能
+
+**(1) 闭集微调**
+
+这个部分无需准备数据,直接用之前的数据即可。
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+**(2) 开集继续预训练微调**
+这种方式需要将 COCO 训练数据转换为 ODVG 格式,你可以使用如下命令转换:
+
+```shell
+python tools/dataset_converters/coco2odvg.py data/coco/annotations/instances_train2017.json -d coco
+```
+
+会在 `data/coco/annotations/` 下生成新的 `instances_train2017_od.json` 和 `coco2017_label_map.json`,完整的数据集结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_train2017_od.json
+│ │ │ ├── coco2017_label_map.json
+│ │ │ ├── instances_val2017.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+在得到数据后,你可以自行选择单独预习还是混合预训练方式。
+
+**(3) 开放词汇微调**
+这种方式需要将 COCO 训练数据转换为 OVD 格式,你可以使用如下命令转换:
+
+```shell
+python tools/dataset_converters/coco2ovd.py data/coco/
+```
+
+会在 `data/coco/annotations/` 下生成新的 `instances_val2017_all_2.json` 和 `instances_val2017_seen_2.json`,完整的数据集结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_train2017_od.json
+│ │ │ ├── instances_val2017_all_2.json
+│ │ │ ├── instances_val2017_seen_2.json
+│ │ │ ├── coco2017_label_map.json
+│ │ │ ├── instances_val2017.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+然后可以直接使用 [配置](coco/grounding_dino_swin-t_finetune_16xb4_1x_coco_48_17.py) 进行训练和测试。
+
+### 2 LVIS 1.0
+
+LVIS 是一个包括 1203 类的数据集,同时也是一个长尾联邦数据集,对其进行微调很有意义。 由于其类别过多,我们无法对其进行闭集微调,因此只能采用开集继续预训练微调和开放词汇微调。
+
+你需要先准备好 LVIS 训练 JSON 文件,你可以从 [这里](https://www.lvisdataset.org/dataset) 下载,我们只需要 `lvis_v1_train.json` 和 `lvis_v1_val.json`,然后将其放到 `data/coco/annotations/` 下,然后运行如下命令:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── lvis_v1_train.json
+│ │ │ ├── lvis_v1_val.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── lvis_v1_minival_inserted_image_name.json
+│ │ │ ├── lvis_od_val.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+(1) 开集继续预训练微调
+
+使用如下命令转换为 ODVG 格式:
+
+```shell
+python tools/dataset_converters/lvis2odvg.py data/coco/annotations/lvis_v1_train.json
+```
+
+会在 `data/coco/annotations/` 下生成新的 `lvis_v1_train_od.json` 和 `lvis_v1_label_map.json`,完整的数据集结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── lvis_v1_train.json
+│ │ │ ├── lvis_v1_val.json
+│ │ │ ├── lvis_v1_train_od.json
+│ │ │ ├── lvis_v1_label_map.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── lvis_v1_minival_inserted_image_name.json
+│ │ │ ├── lvis_od_val.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+然后可以直接使用 [配置](lvis/grounding_dino_swin-t_finetune_16xb4_1x_lvis.py) 进行训练测试,或者你修改配置将其和部分预训练数据集混合使用。
+
+**(2) 开放词汇微调**
+
+使用如下命令转换为 OVD 格式:
+
+```shell
+python tools/dataset_converters/lvis2ovd.py data/coco/
+```
+
+会在 `data/coco/annotations/` 下生成新的 `lvis_v1_train_od_norare.json` 和 `lvis_v1_label_map_norare.json`,完整的数据集结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── lvis_v1_train.json
+│ │ │ ├── lvis_v1_val.json
+│ │ │ ├── lvis_v1_train_od.json
+│ │ │ ├── lvis_v1_label_map.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── lvis_v1_minival_inserted_image_name.json
+│ │ │ ├── lvis_od_val.json
+│ │ │ ├── lvis_v1_train_od_norare.json
+│ │ │ ├── lvis_v1_label_map_norare.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+```
+
+然后可以直接使用 [配置](lvis/grounding_dino_swin-t_finetune_16xb4_1x_lvis_866_337.py) 进行训练测试
+
+### 3 RTTS
+
+RTTS 是一个浓雾天气数据集,该数据集包含 4,322 张雾天图像,包含五个类:自行车 (bicycle)、公共汽车 (bus)、汽车 (car)、摩托车 (motorbike) 和人 (person)。可以从 [这里](https://drive.google.com/file/d/15Ei1cHGVqR1mXFep43BO7nkHq1IEGh1e/view)下载, 然后解压到 `data/RTTS/` 文件夹中。完整的数据集结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── RTTS
+│ │ ├── annotations_json
+│ │ ├── annotations_xml
+│ │ ├── ImageSets
+│ │ ├── JPEGImages
+```
+
+### 4 RUOD
+
+RUOD 是一个水下目标检测数据集,你可以从 [这里](https://drive.google.com/file/d/1hxtbdgfVveUm_DJk5QXkNLokSCTa_E5o/view)下载, 然后解压到 `data/RUOD/` 文件夹中。完整的数据集结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── RUOD
+│ │ ├── Environment_pic
+│ │ ├── Environmet_ANN
+│ │ ├── RUOD_ANN
+│ │ ├── RUOD_pic
+```
+
+### 5 Brain Tumor
+
+Brain Tumor 是一个医学领域的 2d 检测数据集,你可以从 [这里](https://universe.roboflow.com/roboflow-100/brain-tumor-m2pbp/dataset/2)下载, 请注意选择 `COCO JSON` 格式。然后解压到 `data/brain_tumor_v2/` 文件夹中。完整的数据集结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── brain_tumor_v2
+│ │ ├── test
+│ │ ├── train
+│ │ ├── valid
+```
+
+### 6 Cityscapes
+
+Cityscapes 是一个城市街景数据集,你可以从 [这里](https://www.cityscapes-dataset.com/) 或者 opendatalab 中下载, 然后解压到 `data/cityscapes/` 文件夹中。完整的数据集结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── cityscapes
+│ │ ├── annotations
+│ │ ├── leftImg8bit
+│ │ │ ├── train
+│ │ │ ├── val
+│ │ ├── gtFine
+│ │ │ ├── train
+│ │ │ ├── val
+```
+
+在下载后,然后使用 [cityscapes.py](../../tools/dataset_converters/cityscapes.py) 脚本生成我们所需要的 json 格式
+
+```shell
+python tools/dataset_converters/cityscapes.py data/cityscapes/
+```
+
+会在 annotations 中生成 3 个新的 json 文件。完整的数据集结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── cityscapes
+│ │ ├── annotations
+│ │ │ ├── instancesonly_filtered_gtFine_train.json
+│ │ │ ├── instancesonly_filtered_gtFine_val.json
+│ │ │ ├── instancesonly_filtered_gtFine_test.json
+│ │ ├── leftImg8bit
+│ │ │ ├── train
+│ │ │ ├── val
+│ │ ├── gtFine
+│ │ │ ├── train
+│ │ │ ├── val
+```
+
+### 7 People in Painting
+
+People in Painting 是一个油画数据集,你可以从 [这里](https://universe.roboflow.com/roboflow-100/people-in-paintings/dataset/2), 请注意选择 `COCO JSON` 格式。然后解压到 `data/people_in_painting_v2/` 文件夹中。完整的数据集结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── people_in_painting_v2
+│ │ ├── test
+│ │ ├── train
+│ │ ├── valid
+```
+
+### 8 Referring Expression Comprehension
+
+指代性表达式理解的微调和前面一样,也是包括 4 个数据集,在评测数据准备阶段已经全部整理好了,完整的数据集结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── instances_train2014.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+│ │ ├── train2014
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── mdetr_annotations
+│ │ │ ├── final_refexp_val.json
+│ │ │ ├── finetune_refcoco_testA.json
+│ │ │ ├── finetune_refcoco_testB.json
+│ │ │ ├── finetune_refcoco+_testA.json
+│ │ │ ├── finetune_refcoco+_testB.json
+│ │ │ ├── finetune_refcocog_test.json
+│ │ │ ├── finetune_refcocog_test.json
+```
+
+然后我们需要将其转换为所需的 ODVG 格式,请使用 [refcoco2odvg.py](../../tools/dataset_converters/refcoco2odvg.py) 脚本转换,
+
+```shell
+python tools/dataset_converters/refcoco2odvg.py data/coco/mdetr_annotations
+```
+
+会在 `data/coco/mdetr_annotations` 中生成新的 4 个 json 文件。 转换后的数据集结构如下:
+
+```text
+mmdetection
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── instances_train2017.json
+│ │ │ ├── instances_val2017.json
+│ │ │ ├── instances_train2014.json
+│ │ ├── train2017
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── val2017
+│ │ │ ├── xxxx.jpg
+│ │ │ ├── ...
+│ │ ├── train2014
+│ │ │ ├── xxx.jpg
+│ │ │ ├── ...
+│ │ ├── mdetr_annotations
+│ │ │ ├── final_refexp_val.json
+│ │ │ ├── finetune_refcoco_testA.json
+│ │ │ ├── finetune_refcoco_testB.json
+│ │ │ ├── finetune_refcoco+_testA.json
+│ │ │ ├── finetune_refcoco+_testB.json
+│ │ │ ├── finetune_refcocog_test.json
+│ │ │ ├── finetune_refcoco_train_vg.json
+│ │ │ ├── finetune_refcoco+_train_vg.json
+│ │ │ ├── finetune_refcocog_train_vg.json
+│ │ │ ├── finetune_grefcoco_train_vg.json
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/dod/grounding_dino_swin-t_pretrain_zeroshot_concat_dod.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/dod/grounding_dino_swin-t_pretrain_zeroshot_concat_dod.py
new file mode 100644
index 0000000000000000000000000000000000000000..e59a0a52518aa125d556aab12f8076a95f39ec22
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/dod/grounding_dino_swin-t_pretrain_zeroshot_concat_dod.py
@@ -0,0 +1,78 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py'
+
+data_root = 'data/d3/'
+
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile', backend_args=None,
+ imdecode_backend='pillow'),
+ dict(
+ type='FixScaleResize',
+ scale=(800, 1333),
+ keep_ratio=True,
+ backend='pillow'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'text', 'custom_entities', 'sent_ids'))
+]
+
+# -------------------------------------------------#
+val_dataset_full = dict(
+ type='DODDataset',
+ data_root=data_root,
+ ann_file='d3_json/d3_full_annotations.json',
+ data_prefix=dict(img='d3_images/', anno='d3_pkl'),
+ pipeline=test_pipeline,
+ test_mode=True,
+ backend_args=None,
+ return_classes=True)
+
+val_evaluator_full = dict(
+ type='DODCocoMetric',
+ ann_file=data_root + 'd3_json/d3_full_annotations.json')
+
+# -------------------------------------------------#
+val_dataset_pres = dict(
+ type='DODDataset',
+ data_root=data_root,
+ ann_file='d3_json/d3_pres_annotations.json',
+ data_prefix=dict(img='d3_images/', anno='d3_pkl'),
+ pipeline=test_pipeline,
+ test_mode=True,
+ backend_args=None,
+ return_classes=True)
+val_evaluator_pres = dict(
+ type='DODCocoMetric',
+ ann_file=data_root + 'd3_json/d3_pres_annotations.json')
+
+# -------------------------------------------------#
+val_dataset_abs = dict(
+ type='DODDataset',
+ data_root=data_root,
+ ann_file='d3_json/d3_abs_annotations.json',
+ data_prefix=dict(img='d3_images/', anno='d3_pkl'),
+ pipeline=test_pipeline,
+ test_mode=True,
+ backend_args=None,
+ return_classes=True)
+val_evaluator_abs = dict(
+ type='DODCocoMetric',
+ ann_file=data_root + 'd3_json/d3_abs_annotations.json')
+
+# -------------------------------------------------#
+datasets = [val_dataset_full, val_dataset_pres, val_dataset_abs]
+dataset_prefixes = ['FULL', 'PRES', 'ABS']
+metrics = [val_evaluator_full, val_evaluator_pres, val_evaluator_abs]
+
+val_dataloader = dict(
+ dataset=dict(_delete_=True, type='ConcatDataset', datasets=datasets))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='MultiDatasetsEvaluator',
+ metrics=metrics,
+ dataset_prefixes=dataset_prefixes)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/dod/grounding_dino_swin-t_pretrain_zeroshot_parallel_dod.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/dod/grounding_dino_swin-t_pretrain_zeroshot_parallel_dod.py
new file mode 100644
index 0000000000000000000000000000000000000000..3d680091162e5ac96c15c76b58a18764e85d3233
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/dod/grounding_dino_swin-t_pretrain_zeroshot_parallel_dod.py
@@ -0,0 +1,3 @@
+_base_ = 'grounding_dino_swin-t_pretrain_zeroshot_concat_dod.py'
+
+model = dict(test_cfg=dict(chunked_size=1))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/flickr30k/grounding_dino_swin-t-pretrain_flickr30k.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/flickr30k/grounding_dino_swin-t-pretrain_flickr30k.py
new file mode 100644
index 0000000000000000000000000000000000000000..e9eb783da97a6d665002cc9192f740010282870e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/flickr30k/grounding_dino_swin-t-pretrain_flickr30k.py
@@ -0,0 +1,57 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py'
+
+dataset_type = 'Flickr30kDataset'
+data_root = 'data/flickr30k_entities/'
+
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile', backend_args=None,
+ imdecode_backend='pillow'),
+ dict(
+ type='FixScaleResize',
+ scale=(800, 1333),
+ keep_ratio=True,
+ backend='pillow'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'text', 'custom_entities',
+ 'tokens_positive', 'phrase_ids', 'phrases'))
+]
+
+dataset_Flickr30k_val = dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='final_flickr_separateGT_val.json',
+ data_prefix=dict(img='flickr30k_images/'),
+ pipeline=test_pipeline,
+)
+
+dataset_Flickr30k_test = dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='final_flickr_separateGT_test.json',
+ data_prefix=dict(img='flickr30k_images/'),
+ pipeline=test_pipeline,
+)
+
+val_evaluator_Flickr30k = dict(type='Flickr30kMetric')
+
+test_evaluator_Flickr30k = dict(type='Flickr30kMetric')
+
+# ----------Config---------- #
+dataset_prefixes = ['Flickr30kVal', 'Flickr30kTest']
+datasets = [dataset_Flickr30k_val, dataset_Flickr30k_test]
+metrics = [val_evaluator_Flickr30k, test_evaluator_Flickr30k]
+
+val_dataloader = dict(
+ dataset=dict(_delete_=True, type='ConcatDataset', datasets=datasets))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='MultiDatasetsEvaluator',
+ metrics=metrics,
+ dataset_prefixes=dataset_prefixes)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-b_pretrain_all.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-b_pretrain_all.py
new file mode 100644
index 0000000000000000000000000000000000000000..eff58bba6b192fe43e62cb1e3ae40a546e1a3ddf
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-b_pretrain_all.py
@@ -0,0 +1,335 @@
+_base_ = 'grounding_dino_swin-t_pretrain_obj365.py'
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-b_pretrain_obj365_goldg_v3det/grounding_dino_swin-b_pretrain_obj365_goldg_v3de-f83eef00.pth' # noqa
+
+model = dict(
+ use_autocast=True,
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ pretrain_img_size=384,
+ embed_dims=128,
+ depths=[2, 2, 18, 2],
+ num_heads=[4, 8, 16, 32],
+ window_size=12,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(1, 2, 3),
+ with_cp=True,
+ convert_weights=True,
+ frozen_stages=-1,
+ init_cfg=None),
+ neck=dict(in_channels=[256, 512, 1024]),
+)
+
+o365v1_od_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/objects365v1/',
+ ann_file='o365v1_train_odvg.json',
+ label_map_file='o365v1_label_map.json',
+ data_prefix=dict(img='train/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None,
+)
+
+flickr30k_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/flickr30k_entities/',
+ ann_file='final_flickr_separateGT_train_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='flickr30k_images/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+gqa_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/gqa/',
+ ann_file='final_mixed_train_no_coco_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='images/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+v3d_train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=_base_.lang_model_name,
+ num_sample_negative=85,
+ # change this
+ label_map_file='data/V3Det/annotations/v3det_2023_v1_label_map.json',
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+v3det_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/V3Det/',
+ ann_file='annotations/v3det_2023_v1_train_od.json',
+ label_map_file='annotations/v3det_2023_v1_label_map.json',
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=False),
+ need_text=False, # change this
+ pipeline=v3d_train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+grit_dataset = dict(
+ type='ODVGDataset',
+ data_root='grit_processed/',
+ ann_file='grit20m_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+# --------------------------- lvis od dataset---------------------------
+lvis_train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=_base_.lang_model_name,
+ num_sample_negative=85,
+ # change this
+ label_map_file='data/coco/annotations/lvis_v1_label_map.json',
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+lvis_dataset = dict(
+ type='ClassBalancedDataset',
+ oversample_thr=1e-3,
+ dataset=dict(
+ type='ODVGDataset',
+ data_root='data/coco/',
+ ann_file='annotations/lvis_v1_train_od.json',
+ label_map_file='annotations/lvis_v1_label_map.json',
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=False),
+ need_text=False, # change this
+ pipeline=lvis_train_pipeline,
+ return_classes=True,
+ backend_args=None))
+
+# --------------------------- coco2017 od dataset---------------------------
+coco2017_train_dataset = dict(
+ type='RepeatDataset',
+ times=2,
+ dataset=dict(
+ type='ODVGDataset',
+ data_root='data/coco/',
+ ann_file='annotations/instance_train2017_norefval_od.json',
+ label_map_file='annotations/coco2017_label_map.json',
+ data_prefix=dict(img='train2017'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None))
+
+# --------------------------- coco2014 vg dataset---------------------------
+coco2014_vg_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/coco/',
+ ann_file='mdetr_annotations/final_mixed_train_only_coco_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='train2014/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+# --------------------------- refcoco vg dataset---------------------------
+refcoco_dataset = dict(
+ type='RepeatDataset',
+ times=2,
+ dataset=dict(
+ type='ODVGDataset',
+ data_root='data/coco/',
+ ann_file='mdetr_annotations/finetune_refcoco_train_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='train2014'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None))
+
+# --------------------------- refcoco+ vg dataset---------------------------
+refcoco_plus_dataset = dict(
+ type='RepeatDataset',
+ times=2,
+ dataset=dict(
+ type='ODVGDataset',
+ data_root='data/coco/',
+ ann_file='mdetr_annotations/finetune_refcoco+_train_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='train2014'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None))
+
+# --------------------------- refcocog vg dataset---------------------------
+refcocog_dataset = dict(
+ type='RepeatDataset',
+ times=3,
+ dataset=dict(
+ type='ODVGDataset',
+ data_root='data/coco/',
+ ann_file='mdetr_annotations/finetune_refcocog_train_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='train2014'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None))
+
+# --------------------------- grefcoco vg dataset---------------------------
+grefcoco_dataset = dict(
+ type='RepeatDataset',
+ times=2,
+ dataset=dict(
+ type='ODVGDataset',
+ data_root='data/coco/',
+ ann_file='mdetr_annotations/finetune_grefcoco_train_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='train2014'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None))
+
+# --------------------------- dataloader---------------------------
+train_dataloader = dict(
+ batch_size=4,
+ num_workers=4,
+ sampler=dict(
+ _delete_=True,
+ type='CustomSampleSizeSampler',
+ ratio_mode=True,
+ dataset_size=[-1, -1, 0.07, -1, -1, -1, -1, -1, -1, -1, -1, -1]),
+ dataset=dict(datasets=[
+ o365v1_od_dataset, # 1.74M
+ v3det_dataset, #
+ grit_dataset,
+ lvis_dataset,
+ coco2017_train_dataset, # 0.12M
+ flickr30k_dataset, # 0.15M
+ gqa_dataset, # 0.62M
+ coco2014_vg_dataset, # 0.49M
+ refcoco_dataset, # 0.12M
+ refcoco_plus_dataset, # 0.12M
+ refcocog_dataset, # 0.08M
+ grefcoco_dataset, # 0.19M
+ ]))
+
+optim_wrapper = dict(optimizer=dict(lr=0.0001))
+
+# learning policy
+max_iter = 304680
+train_cfg = dict(
+ _delete_=True,
+ type='IterBasedTrainLoop',
+ max_iters=max_iter,
+ val_interval=10000)
+
+param_scheduler = [
+ dict(type='LinearLR', start_factor=0.1, by_epoch=False, begin=0, end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_iter,
+ by_epoch=False,
+ milestones=[228510],
+ gamma=0.1)
+]
+
+default_hooks = dict(
+ checkpoint=dict(by_epoch=False, interval=10000, max_keep_ckpts=20))
+log_processor = dict(by_epoch=False)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-b_pretrain_obj365_goldg_v3det.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-b_pretrain_obj365_goldg_v3det.py
new file mode 100644
index 0000000000000000000000000000000000000000..743d02cffbe9c38977edad2bce8a53bd6a8594af
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-b_pretrain_obj365_goldg_v3det.py
@@ -0,0 +1,143 @@
+_base_ = 'grounding_dino_swin-t_pretrain_obj365.py'
+
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_base_patch4_window12_384_22k.pth' # noqa
+model = dict(
+ use_autocast=True,
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ pretrain_img_size=384,
+ embed_dims=128,
+ depths=[2, 2, 18, 2],
+ num_heads=[4, 8, 16, 32],
+ window_size=12,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(1, 2, 3),
+ with_cp=True,
+ convert_weights=True,
+ frozen_stages=-1,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ neck=dict(in_channels=[256, 512, 1024]),
+)
+
+o365v1_od_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/objects365v1/',
+ ann_file='o365v1_train_odvg.json',
+ label_map_file='o365v1_label_map.json',
+ data_prefix=dict(img='train/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None,
+)
+
+flickr30k_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/flickr30k_entities/',
+ ann_file='final_flickr_separateGT_train_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='flickr30k_images/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+gqa_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/gqa/',
+ ann_file='final_mixed_train_no_coco_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='images/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+v3d_train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=_base_.lang_model_name,
+ num_sample_negative=85,
+ # change this
+ label_map_file='data/V3Det/annotations/v3det_2023_v1_label_map.json',
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+v3det_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/V3Det/',
+ ann_file='annotations/v3det_2023_v1_train_od.json',
+ label_map_file='annotations/v3det_2023_v1_label_map.json',
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=False),
+ need_text=False, # change this
+ pipeline=v3d_train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+train_dataloader = dict(
+ dataset=dict(datasets=[
+ o365v1_od_dataset, flickr30k_dataset, gqa_dataset, v3det_dataset
+ ]))
+
+# learning policy
+max_epochs = 18
+param_scheduler = [
+ dict(type='LinearLR', start_factor=0.1, by_epoch=False, begin=0, end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[13, 16],
+ gamma=0.1)
+]
+
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-l_pretrain_all.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-l_pretrain_all.py
new file mode 100644
index 0000000000000000000000000000000000000000..a17f2344e14d8af81bd267d8bd47662f7e6e059d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-l_pretrain_all.py
@@ -0,0 +1,540 @@
+_base_ = 'grounding_dino_swin-t_pretrain_obj365.py'
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-l_pretrain_obj365_goldg/grounding_dino_swin-l_pretrain_obj365_goldg-34dcdc53.pth' # noqa
+
+num_levels = 5
+model = dict(
+ use_autocast=True,
+ num_feature_levels=num_levels,
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ pretrain_img_size=384,
+ embed_dims=192,
+ depths=[2, 2, 18, 2],
+ num_heads=[6, 12, 24, 48],
+ window_size=12,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.2,
+ patch_norm=True,
+ out_indices=(0, 1, 2, 3),
+ # Please only add indices that would be used
+ # in FPN, otherwise some parameter will not be used
+ with_cp=True,
+ convert_weights=True,
+ frozen_stages=-1,
+ init_cfg=None),
+ neck=dict(in_channels=[192, 384, 768, 1536], num_outs=num_levels),
+ encoder=dict(layer_cfg=dict(self_attn_cfg=dict(num_levels=num_levels))),
+ decoder=dict(layer_cfg=dict(cross_attn_cfg=dict(num_levels=num_levels))))
+
+# --------------------------- object365v2 od dataset---------------------------
+# objv2_backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/objects365v2/': 'yudong:s3://wangyudong/obj365_v2/',
+# 'data/objects365v2/': 'yudong:s3://wangyudong/obj365_v2/'
+# }))
+objv2_backend_args = None
+
+objv2_train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=objv2_backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=_base_.lang_model_name,
+ num_sample_negative=85,
+ # change this
+ label_map_file='data/objects365v2/annotations/o365v2_label_map.json',
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+
+o365v2_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/objects365v2/',
+ ann_file='annotations/zhiyuan_objv2_train_od.json',
+ label_map_file='annotations/o365v2_label_map.json',
+ data_prefix=dict(img='train/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=objv2_train_pipeline,
+ return_classes=True,
+ need_text=False,
+ backend_args=None,
+)
+
+# --------------------------- openimagev6 od dataset---------------------------
+# oi_backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+oi_backend_args = None
+
+oi_train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=oi_backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=_base_.lang_model_name,
+ num_sample_negative=85,
+ # change this
+ label_map_file='data/OpenImages/annotations/openimages_label_map.json',
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+
+oiv6_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/OpenImages/',
+ ann_file='annotations/oidv6-train-annotations_od.json',
+ label_map_file='annotations/openimages_label_map.json',
+ data_prefix=dict(img='OpenImages/train/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ need_text=False,
+ pipeline=oi_train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+# --------------------------- v3det od dataset---------------------------
+v3d_train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=_base_.lang_model_name,
+ num_sample_negative=85,
+ # change this
+ label_map_file='data/V3Det/annotations/v3det_2023_v1_label_map.json',
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+v3det_dataset = dict(
+ type='RepeatDataset',
+ times=2,
+ dataset=dict(
+ type='ODVGDataset',
+ data_root='data/V3Det/',
+ ann_file='annotations/v3det_2023_v1_train_od.json',
+ label_map_file='annotations/v3det_2023_v1_label_map.json',
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=False),
+ need_text=False,
+ pipeline=v3d_train_pipeline,
+ return_classes=True,
+ backend_args=None))
+
+# --------------------------- lvis od dataset---------------------------
+lvis_train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=_base_.lang_model_name,
+ num_sample_negative=85,
+ # change this
+ label_map_file='data/coco/annotations/lvis_v1_label_map.json',
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+lvis_dataset = dict(
+ type='ClassBalancedDataset',
+ oversample_thr=1e-3,
+ dataset=dict(
+ type='ODVGDataset',
+ data_root='data/coco/',
+ ann_file='annotations/lvis_v1_train_od.json',
+ label_map_file='annotations/lvis_v1_label_map.json',
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=False),
+ need_text=False, # change this
+ pipeline=lvis_train_pipeline,
+ return_classes=True,
+ backend_args=None))
+
+# --------------------------- coco2017 od dataset---------------------------
+coco2017_train_dataset = dict(
+ type='RepeatDataset',
+ times=2,
+ dataset=dict(
+ type='ODVGDataset',
+ data_root='data/coco/',
+ ann_file='annotations/instance_train2017_norefval_od.json',
+ label_map_file='annotations/coco2017_label_map.json',
+ data_prefix=dict(img='train2017'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None))
+
+# --------------------------- flickr30k vg dataset---------------------------
+flickr30k_dataset = dict(
+ type='RepeatDataset',
+ times=2,
+ dataset=dict(
+ type='ODVGDataset',
+ data_root='data/flickr30k_entities/',
+ ann_file='final_flickr_separateGT_train_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='flickr30k_images/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None))
+
+# --------------------------- gqa vg dataset---------------------------
+gqa_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/gqa/',
+ ann_file='final_mixed_train_no_coco_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='images/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+# --------------------------- coco2014 vg dataset---------------------------
+coco2014_vg_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/coco/',
+ ann_file='mdetr_annotations/final_mixed_train_only_coco_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='train2014/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+# --------------------------- refcoco vg dataset---------------------------
+refcoco_dataset = dict(
+ type='RepeatDataset',
+ times=2,
+ dataset=dict(
+ type='ODVGDataset',
+ data_root='data/coco/',
+ ann_file='mdetr_annotations/finetune_refcoco_train_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='train2014'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None))
+
+# --------------------------- refcoco+ vg dataset---------------------------
+refcoco_plus_dataset = dict(
+ type='RepeatDataset',
+ times=2,
+ dataset=dict(
+ type='ODVGDataset',
+ data_root='data/coco/',
+ ann_file='mdetr_annotations/finetune_refcoco+_train_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='train2014'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None))
+
+# --------------------------- refcocog vg dataset---------------------------
+refcocog_dataset = dict(
+ type='RepeatDataset',
+ times=3,
+ dataset=dict(
+ type='ODVGDataset',
+ data_root='data/coco/',
+ ann_file='mdetr_annotations/finetune_refcocog_train_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='train2014'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None))
+
+# --------------------------- grefcoco vg dataset---------------------------
+grefcoco_dataset = dict(
+ type='RepeatDataset',
+ times=2,
+ dataset=dict(
+ type='ODVGDataset',
+ data_root='data/coco/',
+ ann_file='mdetr_annotations/finetune_grefcoco_train_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='train2014'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None))
+
+# --------------------------- grit vg dataset---------------------------
+# grit_backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/grit/': 'yichen:s3://chenyicheng/grit/',
+# 'data/grit/': 'yichen:s3://chenyicheng/grit/'
+# }))
+grit_backend_args = None
+
+grit_train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=grit_backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=_base_.lang_model_name,
+ num_sample_negative=85,
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+
+grit_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/grit/',
+ ann_file='grit20m_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=grit_train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+# --------------------------- dataloader---------------------------
+train_dataloader = dict(
+ batch_size=4,
+ num_workers=4,
+ sampler=dict(
+ _delete_=True,
+ type='CustomSampleSizeSampler',
+ ratio_mode=True,
+ # OD ~ 1.74+1.67*0.5+0.18*2+0.12*2+0.1=3.2
+ # vg ~ 0.15*2+0.62*1+0.49*1+0.12*2+0.12*2+0.08*3+0.19*2+9*0.09=3.3
+ dataset_size=[-1, 0.5, -1, -1, -1, -1, -1, -1, -1, -1, -1, -1, 0.09]),
+ dataset=dict(datasets=[
+ o365v2_dataset, # 1.74M
+ oiv6_dataset, # 1.67M
+ v3det_dataset, # 0.18M
+ coco2017_train_dataset, # 0.12M
+ lvis_dataset, # 0.1M
+ flickr30k_dataset, # 0.15M
+ gqa_dataset, # 0.62M
+ coco2014_vg_dataset, # 0.49M
+ refcoco_dataset, # 0.12M
+ refcoco_plus_dataset, # 0.12M
+ refcocog_dataset, # 0.08M
+ grefcoco_dataset, # 0.19M
+ grit_dataset # 9M
+ ]))
+
+# 4NODES * 8GPU
+optim_wrapper = dict(optimizer=dict(lr=0.0001))
+
+max_iter = 250000
+train_cfg = dict(
+ _delete_=True,
+ type='IterBasedTrainLoop',
+ max_iters=max_iter,
+ val_interval=13000)
+
+param_scheduler = [
+ dict(type='LinearLR', start_factor=0.1, by_epoch=False, begin=0, end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_iter,
+ by_epoch=False,
+ milestones=[210000],
+ gamma=0.1)
+]
+
+default_hooks = dict(
+ checkpoint=dict(by_epoch=False, interval=13000, max_keep_ckpts=30))
+log_processor = dict(by_epoch=False)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-l_pretrain_obj365_goldg.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-l_pretrain_obj365_goldg.py
new file mode 100644
index 0000000000000000000000000000000000000000..85d43f96b3bdf79081dfb091c1cc8b6c03de7252
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-l_pretrain_obj365_goldg.py
@@ -0,0 +1,227 @@
+_base_ = 'grounding_dino_swin-t_pretrain_obj365.py'
+
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_large_patch4_window12_384_22k.pth' # noqa
+num_levels = 5
+model = dict(
+ use_autocast=True,
+ num_feature_levels=num_levels,
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ pretrain_img_size=384,
+ embed_dims=192,
+ depths=[2, 2, 18, 2],
+ num_heads=[6, 12, 24, 48],
+ window_size=12,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.2,
+ patch_norm=True,
+ out_indices=(0, 1, 2, 3),
+ # Please only add indices that would be used
+ # in FPN, otherwise some parameter will not be used
+ with_cp=True,
+ convert_weights=True,
+ frozen_stages=-1,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ neck=dict(in_channels=[192, 384, 768, 1536], num_outs=num_levels),
+ encoder=dict(layer_cfg=dict(self_attn_cfg=dict(num_levels=num_levels))),
+ decoder=dict(layer_cfg=dict(cross_attn_cfg=dict(num_levels=num_levels))))
+
+# --------------------------- object365v2 od dataset---------------------------
+# objv2_backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/objects365v2/': 'yudong:s3://wangyudong/obj365_v2/',
+# 'data/objects365v2/': 'yudong:s3://wangyudong/obj365_v2/'
+# }))
+objv2_backend_args = None
+
+objv2_train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=objv2_backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=_base_.lang_model_name,
+ num_sample_negative=85,
+ # change this
+ label_map_file='data/objects365v2/annotations/o365v2_label_map.json',
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+
+o365v2_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/objects365v2/',
+ ann_file='annotations/zhiyuan_objv2_train_od.json',
+ label_map_file='annotations/o365v2_label_map.json',
+ data_prefix=dict(img='train/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=objv2_train_pipeline,
+ return_classes=True,
+ need_text=False,
+ backend_args=None,
+)
+
+# --------------------------- openimagev6 od dataset---------------------------
+# oi_backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+oi_backend_args = None
+
+oi_train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=oi_backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=_base_.lang_model_name,
+ num_sample_negative=85,
+ # change this
+ label_map_file='data/OpenImages/annotations/openimages_label_map.json',
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+
+oiv6_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/OpenImages/',
+ ann_file='annotations/oidv6-train-annotations_od.json',
+ label_map_file='annotations/openimages_label_map.json',
+ data_prefix=dict(img='OpenImages/train/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ need_text=False,
+ pipeline=oi_train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+flickr30k_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/flickr30k_entities/',
+ ann_file='final_flickr_separateGT_train_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='flickr30k_images/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+gqa_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/gqa/',
+ ann_file='final_mixed_train_no_coco_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='images/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+train_dataloader = dict(
+ dataset=dict(datasets=[
+ o365v2_dataset, oiv6_dataset, flickr30k_dataset, gqa_dataset
+ ]))
+
+# 4Nodex8GPU
+optim_wrapper = dict(optimizer=dict(lr=0.0002))
+
+max_iter = 200000
+train_cfg = dict(
+ _delete_=True,
+ type='IterBasedTrainLoop',
+ max_iters=max_iter,
+ val_interval=13000)
+
+param_scheduler = [
+ dict(type='LinearLR', start_factor=0.1, by_epoch=False, begin=0, end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_iter,
+ by_epoch=False,
+ milestones=[156100],
+ gamma=0.5)
+]
+
+default_hooks = dict(
+ checkpoint=dict(by_epoch=False, interval=13000, max_keep_ckpts=30))
+log_processor = dict(by_epoch=False)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_finetune_8xb4_20e_cat.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_finetune_8xb4_20e_cat.py
new file mode 100644
index 0000000000000000000000000000000000000000..bf3b35894eb5fcee6db9f02c2ab8a837cd6da20b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_finetune_8xb4_20e_cat.py
@@ -0,0 +1,102 @@
+_base_ = 'grounding_dino_swin-t_pretrain_obj365.py'
+
+data_root = 'data/cat/'
+class_name = ('cat', )
+num_classes = len(class_name)
+metainfo = dict(classes=class_name, palette=[(220, 20, 60)])
+
+model = dict(bbox_head=dict(num_classes=num_classes))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities'))
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ _delete_=True,
+ type='CocoDataset',
+ data_root=data_root,
+ metainfo=metainfo,
+ return_classes=True,
+ pipeline=train_pipeline,
+ filter_cfg=dict(filter_empty_gt=False, min_size=32),
+ ann_file='annotations/trainval.json',
+ data_prefix=dict(img='images/')))
+
+val_dataloader = dict(
+ dataset=dict(
+ metainfo=metainfo,
+ data_root=data_root,
+ ann_file='annotations/test.json',
+ data_prefix=dict(img='images/')))
+
+test_dataloader = val_dataloader
+
+val_evaluator = dict(ann_file=data_root + 'annotations/test.json')
+test_evaluator = val_evaluator
+
+max_epoch = 20
+
+default_hooks = dict(
+ checkpoint=dict(interval=1, max_keep_ckpts=1, save_best='auto'),
+ logger=dict(type='LoggerHook', interval=5))
+train_cfg = dict(max_epochs=max_epoch, val_interval=1)
+
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epoch,
+ by_epoch=True,
+ milestones=[15],
+ gamma=0.1)
+]
+
+optim_wrapper = dict(
+ optimizer=dict(lr=0.0001),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'backbone': dict(lr_mult=0.0),
+ 'language_model': dict(lr_mult=0.0)
+ }))
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py
new file mode 100644
index 0000000000000000000000000000000000000000..66060f45ea735ab5bbd8e1852c035ea20adcbd80
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py
@@ -0,0 +1,247 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_tiny_patch4_window7_224.pth' # noqa
+lang_model_name = 'bert-base-uncased'
+
+model = dict(
+ type='GroundingDINO',
+ num_queries=900,
+ with_box_refine=True,
+ as_two_stage=True,
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_mask=False,
+ ),
+ language_model=dict(
+ type='BertModel',
+ name=lang_model_name,
+ max_tokens=256,
+ pad_to_max=False,
+ use_sub_sentence_represent=True,
+ special_tokens_list=['[CLS]', '[SEP]', '.', '?'],
+ add_pooling_layer=False,
+ ),
+ backbone=dict(
+ type='SwinTransformer',
+ embed_dims=96,
+ depths=[2, 2, 6, 2],
+ num_heads=[3, 6, 12, 24],
+ window_size=7,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.2,
+ patch_norm=True,
+ out_indices=(1, 2, 3),
+ with_cp=True,
+ convert_weights=True,
+ frozen_stages=-1,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ neck=dict(
+ type='ChannelMapper',
+ in_channels=[192, 384, 768],
+ kernel_size=1,
+ out_channels=256,
+ act_cfg=None,
+ bias=True,
+ norm_cfg=dict(type='GN', num_groups=32),
+ num_outs=4),
+ encoder=dict(
+ num_layers=6,
+ num_cp=6,
+ # visual layer config
+ layer_cfg=dict(
+ self_attn_cfg=dict(embed_dims=256, num_levels=4, dropout=0.0),
+ ffn_cfg=dict(
+ embed_dims=256, feedforward_channels=2048, ffn_drop=0.0)),
+ # text layer config
+ text_layer_cfg=dict(
+ self_attn_cfg=dict(num_heads=4, embed_dims=256, dropout=0.0),
+ ffn_cfg=dict(
+ embed_dims=256, feedforward_channels=1024, ffn_drop=0.0)),
+ # fusion layer config
+ fusion_layer_cfg=dict(
+ v_dim=256,
+ l_dim=256,
+ embed_dim=1024,
+ num_heads=4,
+ init_values=1e-4),
+ ),
+ decoder=dict(
+ num_layers=6,
+ return_intermediate=True,
+ layer_cfg=dict(
+ # query self attention layer
+ self_attn_cfg=dict(embed_dims=256, num_heads=8, dropout=0.0),
+ # cross attention layer query to text
+ cross_attn_text_cfg=dict(embed_dims=256, num_heads=8, dropout=0.0),
+ # cross attention layer query to image
+ cross_attn_cfg=dict(embed_dims=256, num_heads=8, dropout=0.0),
+ ffn_cfg=dict(
+ embed_dims=256, feedforward_channels=2048, ffn_drop=0.0)),
+ post_norm_cfg=None),
+ positional_encoding=dict(
+ num_feats=128, normalize=True, offset=0.0, temperature=20),
+ bbox_head=dict(
+ type='GroundingDINOHead',
+ num_classes=256,
+ sync_cls_avg_factor=True,
+ contrastive_cfg=dict(max_text_len=256, log_scale='auto', bias=True),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0), # 2.0 in DeformDETR
+ loss_bbox=dict(type='L1Loss', loss_weight=5.0)),
+ dn_cfg=dict( # TODO: Move to model.train_cfg ?
+ label_noise_scale=0.5,
+ box_noise_scale=1.0, # 0.4 for DN-DETR
+ group_cfg=dict(dynamic=True, num_groups=None,
+ num_dn_queries=100)), # TODO: half num_dn_queries
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='HungarianAssigner',
+ match_costs=[
+ dict(type='BinaryFocalLossCost', weight=2.0),
+ dict(type='BBoxL1Cost', weight=5.0, box_format='xywh'),
+ dict(type='IoUCost', iou_mode='giou', weight=2.0)
+ ])),
+ test_cfg=dict(max_per_img=300))
+
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=lang_model_name,
+ num_sample_negative=85,
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile', backend_args=None,
+ imdecode_backend='pillow'),
+ dict(
+ type='FixScaleResize',
+ scale=(800, 1333),
+ keep_ratio=True,
+ backend='pillow'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'text', 'custom_entities',
+ 'tokens_positive'))
+]
+
+dataset_type = 'ODVGDataset'
+data_root = 'data/objects365v1/'
+
+coco_od_dataset = dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='o365v1_train_odvg.json',
+ label_map_file='o365v1_label_map.json',
+ data_prefix=dict(img='train/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+train_dataloader = dict(
+ _delete_=True,
+ batch_size=4,
+ num_workers=4,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(type='ConcatDataset', datasets=[coco_od_dataset]))
+
+val_dataloader = dict(
+ dataset=dict(pipeline=test_pipeline, return_classes=True))
+test_dataloader = val_dataloader
+
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0004,
+ weight_decay=0.0001), # bs=16 0.0001
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'backbone': dict(lr_mult=0.1),
+ 'language_model': dict(lr_mult=0.1),
+ }))
+
+# learning policy
+max_epochs = 30
+param_scheduler = [
+ dict(type='LinearLR', start_factor=0.1, by_epoch=False, begin=0, end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[19, 26],
+ gamma=0.1)
+]
+
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (16 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
+
+default_hooks = dict(visualization=dict(type='GroundingVisualizationHook'))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg.py
new file mode 100644
index 0000000000000000000000000000000000000000..b7f388bdd4e8b61d1e7b6fd19445b3628164c4a0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg.py
@@ -0,0 +1,38 @@
+_base_ = 'grounding_dino_swin-t_pretrain_obj365.py'
+
+o365v1_od_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/objects365v1/',
+ ann_file='o365v1_train_odvg.json',
+ label_map_file='o365v1_label_map.json',
+ data_prefix=dict(img='train/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None,
+)
+
+flickr30k_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/flickr30k_entities/',
+ ann_file='final_flickr_separateGT_train_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='flickr30k_images/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+gqa_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/gqa/',
+ ann_file='final_mixed_train_no_coco_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='images/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+train_dataloader = dict(
+ dataset=dict(datasets=[o365v1_od_dataset, flickr30k_dataset, gqa_dataset]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m.py
new file mode 100644
index 0000000000000000000000000000000000000000..8e9f5ca4aaba7afb631f76b8a575101868fed2a4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m.py
@@ -0,0 +1,55 @@
+_base_ = 'grounding_dino_swin-t_pretrain_obj365.py'
+
+o365v1_od_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/objects365v1/',
+ ann_file='o365v1_train_odvg.json',
+ label_map_file='o365v1_label_map.json',
+ data_prefix=dict(img='train/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None,
+)
+
+flickr30k_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/flickr30k_entities/',
+ ann_file='final_flickr_separateGT_train_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='flickr30k_images/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+gqa_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/gqa/',
+ ann_file='final_mixed_train_no_coco_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='images/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+grit_dataset = dict(
+ type='ODVGDataset',
+ data_root='grit_processed/',
+ ann_file='grit20m_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+train_dataloader = dict(
+ sampler=dict(
+ _delete_=True,
+ type='CustomSampleSizeSampler',
+ dataset_size=[-1, -1, -1, 500000]),
+ dataset=dict(datasets=[
+ o365v1_od_dataset, flickr30k_dataset, gqa_dataset, grit_dataset
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det.py
new file mode 100644
index 0000000000000000000000000000000000000000..56e500c86932a8e61dba88fde2bfc00c0ced5585
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det.py
@@ -0,0 +1,117 @@
+_base_ = 'grounding_dino_swin-t_pretrain_obj365.py'
+
+o365v1_od_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/objects365v1/',
+ ann_file='o365v1_train_odvg.json',
+ label_map_file='o365v1_label_map.json',
+ data_prefix=dict(img='train/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None,
+)
+
+flickr30k_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/flickr30k_entities/',
+ ann_file='final_flickr_separateGT_train_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='flickr30k_images/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+gqa_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/gqa/',
+ ann_file='final_mixed_train_no_coco_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='images/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+v3d_train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=_base_.lang_model_name,
+ num_sample_negative=85,
+ # change this
+ label_map_file='data/V3Det/annotations/v3det_2023_v1_label_map.json',
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+v3det_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/V3Det/',
+ ann_file='annotations/v3det_2023_v1_train_od.json',
+ label_map_file='annotations/v3det_2023_v1_label_map.json',
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=False),
+ need_text=False, # change this
+ pipeline=v3d_train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+grit_dataset = dict(
+ type='ODVGDataset',
+ data_root='grit_processed/',
+ ann_file='grit20m_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+train_dataloader = dict(
+ sampler=dict(
+ _delete_=True,
+ type='CustomSampleSizeSampler',
+ dataset_size=[-1, -1, -1, -1, 500000]),
+ dataset=dict(datasets=[
+ o365v1_od_dataset, flickr30k_dataset, gqa_dataset, v3det_dataset,
+ grit_dataset
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_v3det.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_v3det.py
new file mode 100644
index 0000000000000000000000000000000000000000..c89014fbbe43a1e7787fa46d7d850d42a64ff8a9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_v3det.py
@@ -0,0 +1,101 @@
+_base_ = 'grounding_dino_swin-t_pretrain_obj365.py'
+
+o365v1_od_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/objects365v1/',
+ ann_file='o365v1_train_odvg.json',
+ label_map_file='o365v1_label_map.json',
+ data_prefix=dict(img='train/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None,
+)
+
+flickr30k_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/flickr30k_entities/',
+ ann_file='final_flickr_separateGT_train_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='flickr30k_images/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+gqa_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/gqa/',
+ ann_file='final_mixed_train_no_coco_vg.json',
+ label_map_file=None,
+ data_prefix=dict(img='images/'),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=_base_.train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+v3d_train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=_base_.lang_model_name,
+ num_sample_negative=85,
+ # change this
+ label_map_file='data/V3Det/annotations/v3det_2023_v1_label_map.json',
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+v3det_dataset = dict(
+ type='ODVGDataset',
+ data_root='data/V3Det/',
+ ann_file='annotations/v3det_2023_v1_train_od.json',
+ label_map_file='annotations/v3det_2023_v1_label_map.json',
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=False),
+ need_text=False, # change this
+ pipeline=v3d_train_pipeline,
+ return_classes=True,
+ backend_args=None)
+
+train_dataloader = dict(
+ dataset=dict(datasets=[
+ o365v1_od_dataset, flickr30k_dataset, gqa_dataset, v3det_dataset
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_pseudo-labeling_cat.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_pseudo-labeling_cat.py
new file mode 100644
index 0000000000000000000000000000000000000000..6dc8dcd8df4b98a3fdb3aa26d73ce353b9251f50
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_pseudo-labeling_cat.py
@@ -0,0 +1,43 @@
+_base_ = 'grounding_dino_swin-t_pretrain_obj365.py'
+
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile', backend_args=None,
+ imdecode_backend='pillow'),
+ dict(
+ type='FixScaleResize',
+ scale=(800, 1333),
+ keep_ratio=True,
+ backend='pillow'),
+ dict(type='LoadTextAnnotations'),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'text', 'custom_entities',
+ 'tokens_positive'))
+]
+
+data_root = 'data/cat/'
+
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=False,
+ dataset=dict(
+ type='ODVGDataset',
+ data_root=data_root,
+ label_map_file='cat_label_map.json',
+ ann_file='cat_train_od.json',
+ data_prefix=dict(img='images/'),
+ pipeline=test_pipeline,
+ return_classes=True))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ outfile_path=data_root + 'cat_train_od_v1.json',
+ img_prefix=data_root + 'images/',
+ score_thr=0.7,
+ nms_thr=0.5,
+ type='DumpODVGResults')
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_pseudo-labeling_flickr30k.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_pseudo-labeling_flickr30k.py
new file mode 100644
index 0000000000000000000000000000000000000000..78bf1c344bf7c795ace08283b745527dfc9b15f7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_pseudo-labeling_flickr30k.py
@@ -0,0 +1,42 @@
+_base_ = 'grounding_dino_swin-t_pretrain_obj365.py'
+
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile', backend_args=None,
+ imdecode_backend='pillow'),
+ dict(
+ type='FixScaleResize',
+ scale=(800, 1333),
+ keep_ratio=True,
+ backend='pillow'),
+ dict(type='LoadTextAnnotations'),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'text', 'custom_entities',
+ 'tokens_positive'))
+]
+
+data_root = 'data/flickr30k_entities/'
+
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=False,
+ dataset=dict(
+ type='ODVGDataset',
+ data_root=data_root,
+ ann_file='flickr_simple_train_vg.json',
+ data_prefix=dict(img='flickr30k_images/'),
+ pipeline=test_pipeline,
+ return_classes=True))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ outfile_path=data_root + 'flickr_simple_train_vg_v1.json',
+ img_prefix=data_root + 'flickr30k_images/',
+ score_thr=0.4,
+ nms_thr=0.5,
+ type='DumpODVGResults')
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/lvis/grounding_dino_swin-t_finetune_16xb4_1x_lvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/lvis/grounding_dino_swin-t_finetune_16xb4_1x_lvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..3ba12c9067511b00b616781ca0cf2e477e5e689e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/lvis/grounding_dino_swin-t_finetune_16xb4_1x_lvis.py
@@ -0,0 +1,120 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py'
+
+data_root = 'data/coco/'
+
+model = dict(test_cfg=dict(
+ max_per_img=300,
+ chunked_size=40,
+))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=_base_.lang_model_name,
+ num_sample_negative=85,
+ # change this
+ label_map_file='data/coco/annotations/lvis_v1_label_map.json',
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ _delete_=True,
+ type='ClassBalancedDataset',
+ oversample_thr=1e-3,
+ dataset=dict(
+ type='ODVGDataset',
+ data_root=data_root,
+ need_text=False,
+ label_map_file='annotations/lvis_v1_label_map.json',
+ ann_file='annotations/lvis_v1_train_od.json',
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=False, min_size=32),
+ return_classes=True,
+ pipeline=train_pipeline)))
+
+val_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ type='LVISV1Dataset',
+ ann_file='annotations/lvis_v1_minival_inserted_image_name.json',
+ data_prefix=dict(img='')))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='LVISFixedAPMetric',
+ ann_file=data_root +
+ 'annotations/lvis_v1_minival_inserted_image_name.json')
+test_evaluator = val_evaluator
+
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0002, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'backbone': dict(lr_mult=0.1),
+ # 'language_model': dict(lr_mult=0),
+ }))
+
+# learning policy
+max_epochs = 12
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[11],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs, val_interval=3)
+
+default_hooks = dict(
+ checkpoint=dict(
+ max_keep_ckpts=1, save_best='lvis_fixed_ap/AP', rule='greater'))
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/lvis/grounding_dino_swin-t_finetune_16xb4_1x_lvis_866_337.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/lvis/grounding_dino_swin-t_finetune_16xb4_1x_lvis_866_337.py
new file mode 100644
index 0000000000000000000000000000000000000000..28d0141d3e2c0feba26ae4ed924000960c311bf5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/lvis/grounding_dino_swin-t_finetune_16xb4_1x_lvis_866_337.py
@@ -0,0 +1,120 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py'
+
+data_root = 'data/coco/'
+
+model = dict(test_cfg=dict(
+ max_per_img=300,
+ chunked_size=40,
+))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=_base_.lang_model_name,
+ num_sample_negative=85,
+ # change this
+ label_map_file='data/coco/annotations/lvis_v1_label_map_norare.json',
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ _delete_=True,
+ type='ClassBalancedDataset',
+ oversample_thr=1e-3,
+ dataset=dict(
+ type='ODVGDataset',
+ data_root=data_root,
+ need_text=False,
+ label_map_file='annotations/lvis_v1_label_map_norare.json',
+ ann_file='annotations/lvis_v1_train_od_norare.json',
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=False, min_size=32),
+ return_classes=True,
+ pipeline=train_pipeline)))
+
+val_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ type='LVISV1Dataset',
+ ann_file='annotations/lvis_v1_minival_inserted_image_name.json',
+ data_prefix=dict(img='')))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='LVISFixedAPMetric',
+ ann_file=data_root +
+ 'annotations/lvis_v1_minival_inserted_image_name.json')
+test_evaluator = val_evaluator
+
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.00005, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'backbone': dict(lr_mult=0.1),
+ # 'language_model': dict(lr_mult=0),
+ }))
+
+# learning policy
+max_epochs = 12
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs, val_interval=3)
+
+default_hooks = dict(
+ checkpoint=dict(
+ max_keep_ckpts=3, save_best='lvis_fixed_ap/AP', rule='greater'))
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/lvis/grounding_dino_swin-t_pretrain_zeroshot_lvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/lvis/grounding_dino_swin-t_pretrain_zeroshot_lvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..fb4ed438e0b59ca4c991836310cf7103cc02f0f2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/lvis/grounding_dino_swin-t_pretrain_zeroshot_lvis.py
@@ -0,0 +1,24 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py'
+
+model = dict(test_cfg=dict(
+ max_per_img=300,
+ chunked_size=40,
+))
+
+dataset_type = 'LVISV1Dataset'
+data_root = 'data/coco/'
+
+val_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ type=dataset_type,
+ ann_file='annotations/lvis_od_val.json',
+ data_prefix=dict(img='')))
+test_dataloader = val_dataloader
+
+# numpy < 1.24.0
+val_evaluator = dict(
+ _delete_=True,
+ type='LVISFixedAPMetric',
+ ann_file=data_root + 'annotations/lvis_od_val.json')
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/lvis/grounding_dino_swin-t_pretrain_zeroshot_mini-lvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/lvis/grounding_dino_swin-t_pretrain_zeroshot_mini-lvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..406a39a4264a0d6ea5d7950a205b0bac72e8f846
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/lvis/grounding_dino_swin-t_pretrain_zeroshot_mini-lvis.py
@@ -0,0 +1,25 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py'
+
+model = dict(test_cfg=dict(
+ max_per_img=300,
+ chunked_size=40,
+))
+
+dataset_type = 'LVISV1Dataset'
+data_root = 'data/coco/'
+
+val_dataloader = dict(
+ dataset=dict(
+ data_root=data_root,
+ type=dataset_type,
+ ann_file='annotations/lvis_v1_minival_inserted_image_name.json',
+ data_prefix=dict(img='')))
+test_dataloader = val_dataloader
+
+# numpy < 1.24.0
+val_evaluator = dict(
+ _delete_=True,
+ type='LVISFixedAPMetric',
+ ann_file=data_root +
+ 'annotations/lvis_v1_minival_inserted_image_name.json')
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..c104ac051363ab1ed033061e7b01274404d300d1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/metafile.yml
@@ -0,0 +1,90 @@
+Collections:
+ - Name: MM Grounding DINO
+ Metadata:
+ Training Data: Objects365, GoldG, GRIT and V3Det
+ Training Techniques:
+ - AdamW
+ - Multi Scale Train
+ - Gradient Clip
+ Training Resources: 3090 GPUs
+ Architecture:
+ - Swin Transformer
+ - BERT
+ README: configs/mm_grounding_dino/README.md
+ Code:
+ URL:
+ Version: v3.0.0
+
+Models:
+ - Name: grounding_dino_swin-t_pretrain_obj365_goldg
+ In Collection: MM Grounding DINO
+ Config: configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 50.4
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg/grounding_dino_swin-t_pretrain_obj365_goldg_20231122_132602-4ea751ce.pth
+ - Name: grounding_dino_swin-t_pretrain_obj365_goldg_grit9m
+ In Collection: MM Grounding DINO
+ Config: configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 50.5
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_20231128_200818-169cc352.pth
+ - Name: grounding_dino_swin-t_pretrain_obj365_goldg_v3det
+ In Collection: MM Grounding DINO
+ Config: configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_v3det.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 50.6
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_v3det_20231218_095741-e316e297.pth
+ - Name: grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det
+ In Collection: MM Grounding DINO
+ Config: configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 50.4
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth
+ - Name: grounding_dino_swin-b_pretrain_obj365_goldg_v3det
+ In Collection: MM Grounding DINO
+ Config: configs/mm_grounding_dino/grounding_dino_swin-b_pretrain_obj365_goldg_v3det.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 52.5
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-b_pretrain_obj365_goldg_v3det/grounding_dino_swin-b_pretrain_obj365_goldg_v3de-f83eef00.pth
+ - Name: grounding_dino_swin-b_pretrain_all
+ In Collection: MM Grounding DINO
+ Config: configs/mm_grounding_dino/grounding_dino_swin-b_pretrain_all.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 59.5
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-b_pretrain_all/grounding_dino_swin-b_pretrain_all-f9818a7c.pth
+ - Name: grounding_dino_swin-l_pretrain_obj365_goldg
+ In Collection: MM Grounding DINO
+ Config: configs/mm_grounding_dino/grounding_dino_swin-l_pretrain_obj365_goldg.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 53.0
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-l_pretrain_obj365_goldg/grounding_dino_swin-l_pretrain_obj365_goldg-34dcdc53.pth
+ - Name: grounding_dino_swin-l_pretrain_all
+ In Collection: MM Grounding DINO
+ Config: configs/mm_grounding_dino/grounding_dino_swin-l_pretrain_all.py
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 60.3
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-l_pretrain_all/grounding_dino_swin-l_pretrain_all-56d69e78.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/odinw/grounding_dino_swin-t_pretrain_odinw13.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/odinw/grounding_dino_swin-t_pretrain_odinw13.py
new file mode 100644
index 0000000000000000000000000000000000000000..d87ca7ca1ea48a3cff83e15f3e2ad66927598d7f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/odinw/grounding_dino_swin-t_pretrain_odinw13.py
@@ -0,0 +1,338 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py' # noqa
+
+dataset_type = 'CocoDataset'
+data_root = 'data/odinw/'
+
+base_test_pipeline = _base_.test_pipeline
+base_test_pipeline[-1]['meta_keys'] = ('img_id', 'img_path', 'ori_shape',
+ 'img_shape', 'scale_factor', 'text',
+ 'custom_entities', 'caption_prompt')
+
+# ---------------------1 AerialMaritimeDrone---------------------#
+class_name = ('boat', 'car', 'dock', 'jetski', 'lift')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'AerialMaritimeDrone/large/'
+dataset_AerialMaritimeDrone = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ test_mode=True,
+ pipeline=base_test_pipeline,
+ return_classes=True)
+val_evaluator_AerialMaritimeDrone = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------2 Aquarium---------------------#
+class_name = ('fish', 'jellyfish', 'penguin', 'puffin', 'shark', 'starfish',
+ 'stingray')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Aquarium/Aquarium Combined.v2-raw-1024.coco/'
+
+caption_prompt = None
+# caption_prompt = {
+# 'penguin': {
+# 'suffix': ', which is black and white'
+# },
+# 'puffin': {
+# 'suffix': ' with orange beaks'
+# },
+# 'stingray': {
+# 'suffix': ' which is flat and round'
+# },
+# }
+dataset_Aquarium = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Aquarium = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------3 CottontailRabbits---------------------#
+class_name = ('Cottontail-Rabbit', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'CottontailRabbits/'
+
+# caption_prompt = None
+caption_prompt = {'Cottontail-Rabbit': {'name': 'rabbit'}}
+
+dataset_CottontailRabbits = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_CottontailRabbits = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------4 EgoHands---------------------#
+class_name = ('hand', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'EgoHands/generic/'
+
+# caption_prompt = None
+caption_prompt = {'hand': {'suffix': ' of a person'}}
+
+dataset_EgoHands = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_EgoHands = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------5 NorthAmericaMushrooms---------------------#
+class_name = ('CoW', 'chanterelle')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'NorthAmericaMushrooms/North American Mushrooms.v1-416x416.coco/' # noqa
+
+# caption_prompt = None
+caption_prompt = {
+ 'CoW': {
+ 'name': 'flat mushroom'
+ },
+ 'chanterelle': {
+ 'name': 'yellow mushroom'
+ }
+}
+
+dataset_NorthAmericaMushrooms = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_NorthAmericaMushrooms = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------6 Packages---------------------#
+class_name = ('package', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Packages/Raw/'
+
+# caption_prompt = None
+caption_prompt = {
+ 'package': {
+ 'prefix': 'there is a ',
+ 'suffix': ' on the porch'
+ }
+}
+
+dataset_Packages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Packages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------7 PascalVOC---------------------#
+class_name = ('aeroplane', 'bicycle', 'bird', 'boat', 'bottle', 'bus', 'car',
+ 'cat', 'chair', 'cow', 'diningtable', 'dog', 'horse',
+ 'motorbike', 'person', 'pottedplant', 'sheep', 'sofa', 'train',
+ 'tvmonitor')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'PascalVOC/'
+dataset_PascalVOC = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_PascalVOC = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------8 pistols---------------------#
+class_name = ('pistol', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'pistols/export/'
+dataset_pistols = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_pistols = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------9 pothole---------------------#
+class_name = ('pothole', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'pothole/'
+
+# caption_prompt = None
+caption_prompt = {
+ 'pothole': {
+ 'prefix': 'there are some ',
+ 'name': 'holes',
+ 'suffix': ' on the road'
+ }
+}
+
+dataset_pothole = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_pothole = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------10 Raccoon---------------------#
+class_name = ('raccoon', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Raccoon/Raccoon.v2-raw.coco/'
+dataset_Raccoon = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Raccoon = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------11 ShellfishOpenImages---------------------#
+class_name = ('Crab', 'Lobster', 'Shrimp')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'ShellfishOpenImages/raw/'
+dataset_ShellfishOpenImages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_ShellfishOpenImages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------12 thermalDogsAndPeople---------------------#
+class_name = ('dog', 'person')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'thermalDogsAndPeople/'
+dataset_thermalDogsAndPeople = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_thermalDogsAndPeople = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------13 VehiclesOpenImages---------------------#
+class_name = ('Ambulance', 'Bus', 'Car', 'Motorcycle', 'Truck')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'VehiclesOpenImages/416x416/'
+dataset_VehiclesOpenImages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_VehiclesOpenImages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# --------------------- Config---------------------#
+dataset_prefixes = [
+ 'AerialMaritimeDrone', 'Aquarium', 'CottontailRabbits', 'EgoHands',
+ 'NorthAmericaMushrooms', 'Packages', 'PascalVOC', 'pistols', 'pothole',
+ 'Raccoon', 'ShellfishOpenImages', 'thermalDogsAndPeople',
+ 'VehiclesOpenImages'
+]
+datasets = [
+ dataset_AerialMaritimeDrone, dataset_Aquarium, dataset_CottontailRabbits,
+ dataset_EgoHands, dataset_NorthAmericaMushrooms, dataset_Packages,
+ dataset_PascalVOC, dataset_pistols, dataset_pothole, dataset_Raccoon,
+ dataset_ShellfishOpenImages, dataset_thermalDogsAndPeople,
+ dataset_VehiclesOpenImages
+]
+metrics = [
+ val_evaluator_AerialMaritimeDrone, val_evaluator_Aquarium,
+ val_evaluator_CottontailRabbits, val_evaluator_EgoHands,
+ val_evaluator_NorthAmericaMushrooms, val_evaluator_Packages,
+ val_evaluator_PascalVOC, val_evaluator_pistols, val_evaluator_pothole,
+ val_evaluator_Raccoon, val_evaluator_ShellfishOpenImages,
+ val_evaluator_thermalDogsAndPeople, val_evaluator_VehiclesOpenImages
+]
+
+# -------------------------------------------------#
+val_dataloader = dict(
+ dataset=dict(_delete_=True, type='ConcatDataset', datasets=datasets))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='MultiDatasetsEvaluator',
+ metrics=metrics,
+ dataset_prefixes=dataset_prefixes)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/odinw/grounding_dino_swin-t_pretrain_odinw35.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/odinw/grounding_dino_swin-t_pretrain_odinw35.py
new file mode 100644
index 0000000000000000000000000000000000000000..a6b8566aed486ef48653b6e54200cb8817910f2f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/odinw/grounding_dino_swin-t_pretrain_odinw35.py
@@ -0,0 +1,794 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py' # noqa
+
+dataset_type = 'CocoDataset'
+data_root = 'data/odinw/'
+
+base_test_pipeline = _base_.test_pipeline
+base_test_pipeline[-1]['meta_keys'] = ('img_id', 'img_path', 'ori_shape',
+ 'img_shape', 'scale_factor', 'text',
+ 'custom_entities', 'caption_prompt')
+
+# ---------------------1 AerialMaritimeDrone_large---------------------#
+class_name = ('boat', 'car', 'dock', 'jetski', 'lift')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'AerialMaritimeDrone/large/'
+dataset_AerialMaritimeDrone_large = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_AerialMaritimeDrone_large = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------2 AerialMaritimeDrone_tiled---------------------#
+class_name = ('boat', 'car', 'dock', 'jetski', 'lift')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'AerialMaritimeDrone/tiled/'
+dataset_AerialMaritimeDrone_tiled = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_AerialMaritimeDrone_tiled = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------3 AmericanSignLanguageLetters---------------------#
+class_name = ('A', 'B', 'C', 'D', 'E', 'F', 'G', 'H', 'I', 'J', 'K', 'L', 'M',
+ 'N', 'O', 'P', 'Q', 'R', 'S', 'T', 'U', 'V', 'W', 'X', 'Y', 'Z')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'AmericanSignLanguageLetters/American Sign Language Letters.v1-v1.coco/' # noqa
+dataset_AmericanSignLanguageLetters = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_AmericanSignLanguageLetters = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------4 Aquarium---------------------#
+class_name = ('fish', 'jellyfish', 'penguin', 'puffin', 'shark', 'starfish',
+ 'stingray')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Aquarium/Aquarium Combined.v2-raw-1024.coco/'
+dataset_Aquarium = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Aquarium = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------5 BCCD---------------------#
+class_name = ('Platelets', 'RBC', 'WBC')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'BCCD/BCCD.v3-raw.coco/'
+dataset_BCCD = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_BCCD = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------6 boggleBoards---------------------#
+class_name = ('Q', 'a', 'an', 'b', 'c', 'd', 'e', 'er', 'f', 'g', 'h', 'he',
+ 'i', 'in', 'j', 'k', 'l', 'm', 'n', 'o', 'o ', 'p', 'q', 'qu',
+ 'r', 's', 't', 't\\', 'th', 'u', 'v', 'w', 'wild', 'x', 'y', 'z')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'boggleBoards/416x416AutoOrient/export/'
+dataset_boggleBoards = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_boggleBoards = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------7 brackishUnderwater---------------------#
+class_name = ('crab', 'fish', 'jellyfish', 'shrimp', 'small_fish', 'starfish')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'brackishUnderwater/960x540/'
+dataset_brackishUnderwater = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_brackishUnderwater = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------8 ChessPieces---------------------#
+class_name = (' ', 'black bishop', 'black king', 'black knight', 'black pawn',
+ 'black queen', 'black rook', 'white bishop', 'white king',
+ 'white knight', 'white pawn', 'white queen', 'white rook')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'ChessPieces/Chess Pieces.v23-raw.coco/'
+dataset_ChessPieces = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/new_annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_ChessPieces = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/new_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------9 CottontailRabbits---------------------#
+class_name = ('rabbit', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'CottontailRabbits/'
+dataset_CottontailRabbits = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/new_annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_CottontailRabbits = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/new_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------10 dice---------------------#
+class_name = ('1', '2', '3', '4', '5', '6')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'dice/mediumColor/export/'
+dataset_dice = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_dice = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------11 DroneControl---------------------#
+class_name = ('follow', 'follow_hand', 'land', 'land_hand', 'null', 'object',
+ 'takeoff', 'takeoff-hand')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'DroneControl/Drone Control.v3-raw.coco/'
+dataset_DroneControl = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_DroneControl = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------12 EgoHands_generic---------------------#
+class_name = ('hand', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'EgoHands/generic/'
+caption_prompt = {'hand': {'suffix': ' of a person'}}
+dataset_EgoHands_generic = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_EgoHands_generic = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------13 EgoHands_specific---------------------#
+class_name = ('myleft', 'myright', 'yourleft', 'yourright')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'EgoHands/specific/'
+dataset_EgoHands_specific = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_EgoHands_specific = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------14 HardHatWorkers---------------------#
+class_name = ('head', 'helmet', 'person')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'HardHatWorkers/raw/'
+dataset_HardHatWorkers = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_HardHatWorkers = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------15 MaskWearing---------------------#
+class_name = ('mask', 'no-mask')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'MaskWearing/raw/'
+dataset_MaskWearing = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_MaskWearing = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------16 MountainDewCommercial---------------------#
+class_name = ('bottle', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'MountainDewCommercial/'
+dataset_MountainDewCommercial = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_MountainDewCommercial = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------17 NorthAmericaMushrooms---------------------#
+class_name = ('flat mushroom', 'yellow mushroom')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'NorthAmericaMushrooms/North American Mushrooms.v1-416x416.coco/' # noqa
+dataset_NorthAmericaMushrooms = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/new_annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_NorthAmericaMushrooms = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/new_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------18 openPoetryVision---------------------#
+class_name = ('American Typewriter', 'Andale Mono', 'Apple Chancery', 'Arial',
+ 'Avenir', 'Baskerville', 'Big Caslon', 'Bradley Hand',
+ 'Brush Script MT', 'Chalkboard', 'Comic Sans MS', 'Copperplate',
+ 'Courier', 'Didot', 'Futura', 'Geneva', 'Georgia', 'Gill Sans',
+ 'Helvetica', 'Herculanum', 'Impact', 'Kefa', 'Lucida Grande',
+ 'Luminari', 'Marker Felt', 'Menlo', 'Monaco', 'Noteworthy',
+ 'Optima', 'PT Sans', 'PT Serif', 'Palatino', 'Papyrus',
+ 'Phosphate', 'Rockwell', 'SF Pro', 'SignPainter', 'Skia',
+ 'Snell Roundhand', 'Tahoma', 'Times New Roman', 'Trebuchet MS',
+ 'Verdana')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'openPoetryVision/512x512/'
+dataset_openPoetryVision = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_openPoetryVision = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------19 OxfordPets_by_breed---------------------#
+class_name = ('cat-Abyssinian', 'cat-Bengal', 'cat-Birman', 'cat-Bombay',
+ 'cat-British_Shorthair', 'cat-Egyptian_Mau', 'cat-Maine_Coon',
+ 'cat-Persian', 'cat-Ragdoll', 'cat-Russian_Blue', 'cat-Siamese',
+ 'cat-Sphynx', 'dog-american_bulldog',
+ 'dog-american_pit_bull_terrier', 'dog-basset_hound',
+ 'dog-beagle', 'dog-boxer', 'dog-chihuahua',
+ 'dog-english_cocker_spaniel', 'dog-english_setter',
+ 'dog-german_shorthaired', 'dog-great_pyrenees', 'dog-havanese',
+ 'dog-japanese_chin', 'dog-keeshond', 'dog-leonberger',
+ 'dog-miniature_pinscher', 'dog-newfoundland', 'dog-pomeranian',
+ 'dog-pug', 'dog-saint_bernard', 'dog-samoyed',
+ 'dog-scottish_terrier', 'dog-shiba_inu',
+ 'dog-staffordshire_bull_terrier', 'dog-wheaten_terrier',
+ 'dog-yorkshire_terrier')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'OxfordPets/by-breed/' # noqa
+dataset_OxfordPets_by_breed = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_OxfordPets_by_breed = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------20 OxfordPets_by_species---------------------#
+class_name = ('cat', 'dog')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'OxfordPets/by-species/' # noqa
+dataset_OxfordPets_by_species = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_OxfordPets_by_species = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------21 PKLot---------------------#
+class_name = ('space-empty', 'space-occupied')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'PKLot/640/' # noqa
+dataset_PKLot = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_PKLot = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------22 Packages---------------------#
+class_name = ('package', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Packages/Raw/'
+caption_prompt = {
+ 'package': {
+ 'prefix': 'there is a ',
+ 'suffix': ' on the porch'
+ }
+}
+dataset_Packages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=base_test_pipeline,
+ caption_prompt=caption_prompt,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Packages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------23 PascalVOC---------------------#
+class_name = ('aeroplane', 'bicycle', 'bird', 'boat', 'bottle', 'bus', 'car',
+ 'cat', 'chair', 'cow', 'diningtable', 'dog', 'horse',
+ 'motorbike', 'person', 'pottedplant', 'sheep', 'sofa', 'train',
+ 'tvmonitor')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'PascalVOC/'
+dataset_PascalVOC = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_PascalVOC = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------24 pistols---------------------#
+class_name = ('pistol', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'pistols/export/'
+dataset_pistols = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_pistols = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------25 plantdoc---------------------#
+class_name = ('Apple Scab Leaf', 'Apple leaf', 'Apple rust leaf',
+ 'Bell_pepper leaf', 'Bell_pepper leaf spot', 'Blueberry leaf',
+ 'Cherry leaf', 'Corn Gray leaf spot', 'Corn leaf blight',
+ 'Corn rust leaf', 'Peach leaf', 'Potato leaf',
+ 'Potato leaf early blight', 'Potato leaf late blight',
+ 'Raspberry leaf', 'Soyabean leaf', 'Soybean leaf',
+ 'Squash Powdery mildew leaf', 'Strawberry leaf',
+ 'Tomato Early blight leaf', 'Tomato Septoria leaf spot',
+ 'Tomato leaf', 'Tomato leaf bacterial spot',
+ 'Tomato leaf late blight', 'Tomato leaf mosaic virus',
+ 'Tomato leaf yellow virus', 'Tomato mold leaf',
+ 'Tomato two spotted spider mites leaf', 'grape leaf',
+ 'grape leaf black rot')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'plantdoc/416x416/'
+dataset_plantdoc = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_plantdoc = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------26 pothole---------------------#
+class_name = ('pothole', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'pothole/'
+caption_prompt = {
+ 'pothole': {
+ 'name': 'holes',
+ 'prefix': 'there are some ',
+ 'suffix': ' on the road'
+ }
+}
+dataset_pothole = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ caption_prompt=caption_prompt,
+ pipeline=base_test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_pothole = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------27 Raccoon---------------------#
+class_name = ('raccoon', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'Raccoon/Raccoon.v2-raw.coco/'
+dataset_Raccoon = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_Raccoon = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------28 selfdrivingCar---------------------#
+class_name = ('biker', 'car', 'pedestrian', 'trafficLight',
+ 'trafficLight-Green', 'trafficLight-GreenLeft',
+ 'trafficLight-Red', 'trafficLight-RedLeft',
+ 'trafficLight-Yellow', 'trafficLight-YellowLeft', 'truck')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'selfdrivingCar/fixedLarge/export/'
+dataset_selfdrivingCar = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='val_annotations_without_background.json',
+ data_prefix=dict(img=''),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_selfdrivingCar = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'val_annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------29 ShellfishOpenImages---------------------#
+class_name = ('Crab', 'Lobster', 'Shrimp')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'ShellfishOpenImages/raw/'
+dataset_ShellfishOpenImages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_ShellfishOpenImages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------30 ThermalCheetah---------------------#
+class_name = ('cheetah', 'human')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'ThermalCheetah/'
+dataset_ThermalCheetah = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_ThermalCheetah = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------31 thermalDogsAndPeople---------------------#
+class_name = ('dog', 'person')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'thermalDogsAndPeople/'
+dataset_thermalDogsAndPeople = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_thermalDogsAndPeople = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------32 UnoCards---------------------#
+class_name = ('0', '1', '2', '3', '4', '5', '6', '7', '8', '9', '10', '11',
+ '12', '13', '14')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'UnoCards/raw/'
+dataset_UnoCards = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_UnoCards = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------33 VehiclesOpenImages---------------------#
+class_name = ('Ambulance', 'Bus', 'Car', 'Motorcycle', 'Truck')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'VehiclesOpenImages/416x416/'
+dataset_VehiclesOpenImages = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_VehiclesOpenImages = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------34 WildfireSmoke---------------------#
+class_name = ('smoke', )
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'WildfireSmoke/'
+dataset_WildfireSmoke = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_WildfireSmoke = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# ---------------------35 websiteScreenshots---------------------#
+class_name = ('button', 'field', 'heading', 'iframe', 'image', 'label', 'link',
+ 'text')
+metainfo = dict(classes=class_name)
+_data_root = data_root + 'websiteScreenshots/'
+dataset_websiteScreenshots = dict(
+ type=dataset_type,
+ metainfo=metainfo,
+ data_root=_data_root,
+ ann_file='valid/annotations_without_background.json',
+ data_prefix=dict(img='valid/'),
+ pipeline=_base_.test_pipeline,
+ test_mode=True,
+ return_classes=True)
+val_evaluator_websiteScreenshots = dict(
+ type='CocoMetric',
+ ann_file=_data_root + 'valid/annotations_without_background.json',
+ metric='bbox')
+
+# --------------------- Config---------------------#
+
+dataset_prefixes = [
+ 'AerialMaritimeDrone_large',
+ 'AerialMaritimeDrone_tiled',
+ 'AmericanSignLanguageLetters',
+ 'Aquarium',
+ 'BCCD',
+ 'boggleBoards',
+ 'brackishUnderwater',
+ 'ChessPieces',
+ 'CottontailRabbits',
+ 'dice',
+ 'DroneControl',
+ 'EgoHands_generic',
+ 'EgoHands_specific',
+ 'HardHatWorkers',
+ 'MaskWearing',
+ 'MountainDewCommercial',
+ 'NorthAmericaMushrooms',
+ 'openPoetryVision',
+ 'OxfordPets_by_breed',
+ 'OxfordPets_by_species',
+ 'PKLot',
+ 'Packages',
+ 'PascalVOC',
+ 'pistols',
+ 'plantdoc',
+ 'pothole',
+ 'Raccoons',
+ 'selfdrivingCar',
+ 'ShellfishOpenImages',
+ 'ThermalCheetah',
+ 'thermalDogsAndPeople',
+ 'UnoCards',
+ 'VehiclesOpenImages',
+ 'WildfireSmoke',
+ 'websiteScreenshots',
+]
+
+datasets = [
+ dataset_AerialMaritimeDrone_large, dataset_AerialMaritimeDrone_tiled,
+ dataset_AmericanSignLanguageLetters, dataset_Aquarium, dataset_BCCD,
+ dataset_boggleBoards, dataset_brackishUnderwater, dataset_ChessPieces,
+ dataset_CottontailRabbits, dataset_dice, dataset_DroneControl,
+ dataset_EgoHands_generic, dataset_EgoHands_specific,
+ dataset_HardHatWorkers, dataset_MaskWearing, dataset_MountainDewCommercial,
+ dataset_NorthAmericaMushrooms, dataset_openPoetryVision,
+ dataset_OxfordPets_by_breed, dataset_OxfordPets_by_species, dataset_PKLot,
+ dataset_Packages, dataset_PascalVOC, dataset_pistols, dataset_plantdoc,
+ dataset_pothole, dataset_Raccoon, dataset_selfdrivingCar,
+ dataset_ShellfishOpenImages, dataset_ThermalCheetah,
+ dataset_thermalDogsAndPeople, dataset_UnoCards, dataset_VehiclesOpenImages,
+ dataset_WildfireSmoke, dataset_websiteScreenshots
+]
+
+metrics = [
+ val_evaluator_AerialMaritimeDrone_large,
+ val_evaluator_AerialMaritimeDrone_tiled,
+ val_evaluator_AmericanSignLanguageLetters, val_evaluator_Aquarium,
+ val_evaluator_BCCD, val_evaluator_boggleBoards,
+ val_evaluator_brackishUnderwater, val_evaluator_ChessPieces,
+ val_evaluator_CottontailRabbits, val_evaluator_dice,
+ val_evaluator_DroneControl, val_evaluator_EgoHands_generic,
+ val_evaluator_EgoHands_specific, val_evaluator_HardHatWorkers,
+ val_evaluator_MaskWearing, val_evaluator_MountainDewCommercial,
+ val_evaluator_NorthAmericaMushrooms, val_evaluator_openPoetryVision,
+ val_evaluator_OxfordPets_by_breed, val_evaluator_OxfordPets_by_species,
+ val_evaluator_PKLot, val_evaluator_Packages, val_evaluator_PascalVOC,
+ val_evaluator_pistols, val_evaluator_plantdoc, val_evaluator_pothole,
+ val_evaluator_Raccoon, val_evaluator_selfdrivingCar,
+ val_evaluator_ShellfishOpenImages, val_evaluator_ThermalCheetah,
+ val_evaluator_thermalDogsAndPeople, val_evaluator_UnoCards,
+ val_evaluator_VehiclesOpenImages, val_evaluator_WildfireSmoke,
+ val_evaluator_websiteScreenshots
+]
+
+# -------------------------------------------------#
+val_dataloader = dict(
+ dataset=dict(_delete_=True, type='ConcatDataset', datasets=datasets))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='MultiDatasetsEvaluator',
+ metrics=metrics,
+ dataset_prefixes=dataset_prefixes)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/odinw/override_category.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/odinw/override_category.py
new file mode 100644
index 0000000000000000000000000000000000000000..9ff05fc6e5e4d0989cf7fcf7af4dc902ee99f3a3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/odinw/override_category.py
@@ -0,0 +1,109 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import argparse
+
+import mmengine
+
+
+def parse_args():
+ parser = argparse.ArgumentParser(description='Override Category')
+ parser.add_argument('data_root')
+ return parser.parse_args()
+
+
+def main():
+ args = parse_args()
+
+ ChessPieces = [{
+ 'id': 1,
+ 'name': ' ',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 2,
+ 'name': 'black bishop',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 3,
+ 'name': 'black king',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 4,
+ 'name': 'black knight',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 5,
+ 'name': 'black pawn',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 6,
+ 'name': 'black queen',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 7,
+ 'name': 'black rook',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 8,
+ 'name': 'white bishop',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 9,
+ 'name': 'white king',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 10,
+ 'name': 'white knight',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 11,
+ 'name': 'white pawn',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 12,
+ 'name': 'white queen',
+ 'supercategory': 'pieces'
+ }, {
+ 'id': 13,
+ 'name': 'white rook',
+ 'supercategory': 'pieces'
+ }]
+
+ _data_root = args.data_root + 'ChessPieces/Chess Pieces.v23-raw.coco/'
+ json_data = mmengine.load(_data_root +
+ 'valid/annotations_without_background.json')
+ json_data['categories'] = ChessPieces
+ mmengine.dump(json_data,
+ _data_root + 'valid/new_annotations_without_background.json')
+
+ CottontailRabbits = [{
+ 'id': 1,
+ 'name': 'rabbit',
+ 'supercategory': 'Cottontail-Rabbit'
+ }]
+
+ _data_root = args.data_root + 'CottontailRabbits/'
+ json_data = mmengine.load(_data_root +
+ 'valid/annotations_without_background.json')
+ json_data['categories'] = CottontailRabbits
+ mmengine.dump(json_data,
+ _data_root + 'valid/new_annotations_without_background.json')
+
+ NorthAmericaMushrooms = [{
+ 'id': 1,
+ 'name': 'flat mushroom',
+ 'supercategory': 'mushroom'
+ }, {
+ 'id': 2,
+ 'name': 'yellow mushroom',
+ 'supercategory': 'mushroom'
+ }]
+
+ _data_root = args.data_root + 'NorthAmericaMushrooms/North American Mushrooms.v1-416x416.coco/' # noqa
+ json_data = mmengine.load(_data_root +
+ 'valid/annotations_without_background.json')
+ json_data['categories'] = NorthAmericaMushrooms
+ mmengine.dump(json_data,
+ _data_root + 'valid/new_annotations_without_background.json')
+
+
+if __name__ == '__main__':
+ main()
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/people_in_painting/grounding_dino_swin-t_finetune_8xb4_50e_people_in_painting.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/people_in_painting/grounding_dino_swin-t_finetune_8xb4_50e_people_in_painting.py
new file mode 100644
index 0000000000000000000000000000000000000000..449d8682f896c3857e6a50b16a13b43acc77ebc2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/people_in_painting/grounding_dino_swin-t_finetune_8xb4_50e_people_in_painting.py
@@ -0,0 +1,109 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py'
+
+# https://universe.roboflow.com/roboflow-100/people-in-paintings/dataset/2
+data_root = 'data/people_in_painting_v2/'
+class_name = ('Human', )
+palette = [(220, 20, 60)]
+
+metainfo = dict(classes=class_name, palette=palette)
+
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities'))
+]
+
+train_dataloader = dict(
+ sampler=dict(_delete_=True, type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ _delete_=True,
+ type='RepeatDataset',
+ times=10,
+ dataset=dict(
+ type='CocoDataset',
+ data_root=data_root,
+ metainfo=metainfo,
+ filter_cfg=dict(filter_empty_gt=False, min_size=32),
+ pipeline=train_pipeline,
+ return_classes=True,
+ data_prefix=dict(img='train/'),
+ ann_file='train/_annotations.coco.json')))
+
+val_dataloader = dict(
+ dataset=dict(
+ metainfo=metainfo,
+ data_root=data_root,
+ return_classes=True,
+ ann_file='valid/_annotations.coco.json',
+ data_prefix=dict(img='valid/')))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'valid/_annotations.coco.json',
+ metric='bbox',
+ format_only=False)
+test_evaluator = val_evaluator
+
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0001, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'backbone': dict(lr_mult=0.1)
+ }))
+
+# learning policy
+max_epochs = 5
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[4],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs, val_interval=1)
+default_hooks = dict(checkpoint=dict(max_keep_ckpts=1, save_best='auto'))
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/refcoco/grounding_dino_swin-t_finetune_8xb4_5e_grefcoco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/refcoco/grounding_dino_swin-t_finetune_8xb4_5e_grefcoco.py
new file mode 100644
index 0000000000000000000000000000000000000000..983ffe5c6f3f6e59cf1616a0b22c17f065e08437
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/refcoco/grounding_dino_swin-t_finetune_8xb4_5e_grefcoco.py
@@ -0,0 +1,170 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py'
+
+data_root = 'data/coco/'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ # change this
+ dict(type='RandomFlip', prob=0.0),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=_base_.lang_model_name,
+ num_sample_negative=85,
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ _delete_=True,
+ type='ODVGDataset',
+ data_root=data_root,
+ ann_file='mdetr_annotations/finetune_grefcoco_train_vg.json',
+ data_prefix=dict(img='train2014/'),
+ filter_cfg=dict(filter_empty_gt=False, min_size=32),
+ return_classes=True,
+ pipeline=train_pipeline))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_grefcoco_val.json'
+val_dataset_all_val = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=_base_.test_pipeline,
+ backend_args=None)
+val_evaluator_all_val = dict(
+ type='gRefCOCOMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ thresh_score=0.7,
+ thresh_f1=1.0)
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_grefcoco_testA.json'
+val_dataset_refcoco_testA = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=_base_.test_pipeline,
+ backend_args=None)
+
+val_evaluator_refcoco_testA = dict(
+ type='gRefCOCOMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ thresh_score=0.7,
+ thresh_f1=1.0)
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_grefcoco_testB.json'
+val_dataset_refcoco_testB = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=_base_.test_pipeline,
+ backend_args=None)
+
+val_evaluator_refcoco_testB = dict(
+ type='gRefCOCOMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ thresh_score=0.7,
+ thresh_f1=1.0)
+
+# -------------------------------------------------#
+datasets = [
+ val_dataset_all_val, val_dataset_refcoco_testA, val_dataset_refcoco_testB
+]
+dataset_prefixes = ['grefcoco_val', 'grefcoco_testA', 'grefcoco_testB']
+metrics = [
+ val_evaluator_all_val, val_evaluator_refcoco_testA,
+ val_evaluator_refcoco_testB
+]
+
+val_dataloader = dict(
+ dataset=dict(_delete_=True, type='ConcatDataset', datasets=datasets))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='MultiDatasetsEvaluator',
+ metrics=metrics,
+ dataset_prefixes=dataset_prefixes)
+test_evaluator = val_evaluator
+
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0002, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'backbone': dict(lr_mult=0.1),
+ # 'language_model': dict(lr_mult=0),
+ }))
+
+# learning policy
+max_epochs = 5
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[3],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs, val_interval=1)
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/refcoco/grounding_dino_swin-t_finetune_8xb4_5e_refcoco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/refcoco/grounding_dino_swin-t_finetune_8xb4_5e_refcoco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d91af473a239f2f48a09a272d926e00c52da987b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/refcoco/grounding_dino_swin-t_finetune_8xb4_5e_refcoco.py
@@ -0,0 +1,167 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py'
+
+data_root = 'data/coco/'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ # change this
+ dict(type='RandomFlip', prob=0.0),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=_base_.lang_model_name,
+ num_sample_negative=85,
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ _delete_=True,
+ type='ODVGDataset',
+ data_root=data_root,
+ ann_file='mdetr_annotations/finetune_refcoco_train_vg.json',
+ data_prefix=dict(img='train2014/'),
+ filter_cfg=dict(filter_empty_gt=False, min_size=32),
+ return_classes=True,
+ pipeline=train_pipeline))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_refcoco_val.json'
+val_dataset_all_val = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=_base_.test_pipeline,
+ backend_args=None)
+val_evaluator_all_val = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_refcoco_testA.json'
+val_dataset_refcoco_testA = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=_base_.test_pipeline,
+ backend_args=None)
+
+val_evaluator_refcoco_testA = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_refcoco_testB.json'
+val_dataset_refcoco_testB = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=_base_.test_pipeline,
+ backend_args=None)
+
+val_evaluator_refcoco_testB = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+datasets = [
+ val_dataset_all_val, val_dataset_refcoco_testA, val_dataset_refcoco_testB
+]
+dataset_prefixes = ['refcoco_val', 'refcoco_testA', 'refcoco_testB']
+metrics = [
+ val_evaluator_all_val, val_evaluator_refcoco_testA,
+ val_evaluator_refcoco_testB
+]
+
+val_dataloader = dict(
+ dataset=dict(_delete_=True, type='ConcatDataset', datasets=datasets))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='MultiDatasetsEvaluator',
+ metrics=metrics,
+ dataset_prefixes=dataset_prefixes)
+test_evaluator = val_evaluator
+
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0002, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'backbone': dict(lr_mult=0.1),
+ # 'language_model': dict(lr_mult=0),
+ }))
+
+# learning policy
+max_epochs = 5
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[3],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs, val_interval=1)
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/refcoco/grounding_dino_swin-t_finetune_8xb4_5e_refcoco_plus.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/refcoco/grounding_dino_swin-t_finetune_8xb4_5e_refcoco_plus.py
new file mode 100644
index 0000000000000000000000000000000000000000..871adc8efb48532fb5e0fbfa07e6019c37911712
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/refcoco/grounding_dino_swin-t_finetune_8xb4_5e_refcoco_plus.py
@@ -0,0 +1,167 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py'
+
+data_root = 'data/coco/'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ # change this
+ dict(type='RandomFlip', prob=0.0),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=_base_.lang_model_name,
+ num_sample_negative=85,
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ _delete_=True,
+ type='ODVGDataset',
+ data_root=data_root,
+ ann_file='mdetr_annotations/finetune_refcoco+_train_vg.json',
+ data_prefix=dict(img='train2014/'),
+ filter_cfg=dict(filter_empty_gt=False, min_size=32),
+ return_classes=True,
+ pipeline=train_pipeline))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_refcoco+_val.json'
+val_dataset_all_val = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=_base_.test_pipeline,
+ backend_args=None)
+val_evaluator_all_val = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_refcoco+_testA.json'
+val_dataset_refcoco_testA = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=_base_.test_pipeline,
+ backend_args=None)
+
+val_evaluator_refcoco_testA = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_refcoco+_testB.json'
+val_dataset_refcoco_testB = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=_base_.test_pipeline,
+ backend_args=None)
+
+val_evaluator_refcoco_testB = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+datasets = [
+ val_dataset_all_val, val_dataset_refcoco_testA, val_dataset_refcoco_testB
+]
+dataset_prefixes = ['refcoco+_val', 'refcoco+_testA', 'refcoco+_testB']
+metrics = [
+ val_evaluator_all_val, val_evaluator_refcoco_testA,
+ val_evaluator_refcoco_testB
+]
+
+val_dataloader = dict(
+ dataset=dict(_delete_=True, type='ConcatDataset', datasets=datasets))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='MultiDatasetsEvaluator',
+ metrics=metrics,
+ dataset_prefixes=dataset_prefixes)
+test_evaluator = val_evaluator
+
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0002, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'backbone': dict(lr_mult=0.1),
+ # 'language_model': dict(lr_mult=0),
+ }))
+
+# learning policy
+max_epochs = 5
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[3],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs, val_interval=1)
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/refcoco/grounding_dino_swin-t_finetune_8xb4_5e_refcocog.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/refcoco/grounding_dino_swin-t_finetune_8xb4_5e_refcocog.py
new file mode 100644
index 0000000000000000000000000000000000000000..a351d6f9d123fc8f2000990a5e6d02adbb3eb2fa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/refcoco/grounding_dino_swin-t_finetune_8xb4_5e_refcocog.py
@@ -0,0 +1,145 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py'
+
+data_root = 'data/coco/'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ # change this
+ dict(type='RandomFlip', prob=0.0),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(
+ type='RandomSamplingNegPos',
+ tokenizer_name=_base_.lang_model_name,
+ num_sample_negative=85,
+ max_tokens=256),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities', 'tokens_positive', 'dataset_mode'))
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ _delete_=True,
+ type='ODVGDataset',
+ data_root=data_root,
+ ann_file='mdetr_annotations/finetune_refcocog_train_vg.json',
+ data_prefix=dict(img='train2014/'),
+ filter_cfg=dict(filter_empty_gt=False, min_size=32),
+ return_classes=True,
+ pipeline=train_pipeline))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_refcocog_val.json'
+val_dataset_all_val = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=_base_.test_pipeline,
+ backend_args=None)
+val_evaluator_all_val = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_refcocog_test.json'
+val_dataset_refcoco_test = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=_base_.test_pipeline,
+ backend_args=None)
+
+val_evaluator_refcoco_test = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+datasets = [val_dataset_all_val, val_dataset_refcoco_test]
+dataset_prefixes = ['refcocog_val', 'refcocog_test']
+metrics = [val_evaluator_all_val, val_evaluator_refcoco_test]
+
+val_dataloader = dict(
+ dataset=dict(_delete_=True, type='ConcatDataset', datasets=datasets))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='MultiDatasetsEvaluator',
+ metrics=metrics,
+ dataset_prefixes=dataset_prefixes)
+test_evaluator = val_evaluator
+
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0002, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'backbone': dict(lr_mult=0.1),
+ # 'language_model': dict(lr_mult=0),
+ }))
+
+# learning policy
+max_epochs = 5
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[3],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs, val_interval=1)
+
+default_hooks = dict(checkpoint=dict(max_keep_ckpts=1, save_best='auto'))
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/refcoco/grounding_dino_swin-t_pretrain_zeroshot_refexp.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/refcoco/grounding_dino_swin-t_pretrain_zeroshot_refexp.py
new file mode 100644
index 0000000000000000000000000000000000000000..437d71c6b357eda85d13b5efd4c81d4d32f91120
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/refcoco/grounding_dino_swin-t_pretrain_zeroshot_refexp.py
@@ -0,0 +1,228 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py'
+
+# 30 is an empirical value, just set it to the maximum value
+# without affecting the evaluation result
+model = dict(test_cfg=dict(max_per_img=30))
+
+data_root = 'data/coco/'
+
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile', backend_args=None,
+ imdecode_backend='pillow'),
+ dict(
+ type='FixScaleResize',
+ scale=(800, 1333),
+ keep_ratio=True,
+ backend='pillow'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'text', 'custom_entities',
+ 'tokens_positive'))
+]
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/final_refexp_val.json'
+val_dataset_all_val = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=test_pipeline,
+ backend_args=None)
+val_evaluator_all_val = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_refcoco_testA.json'
+val_dataset_refcoco_testA = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=test_pipeline,
+ backend_args=None)
+
+val_evaluator_refcoco_testA = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_refcoco_testB.json'
+val_dataset_refcoco_testB = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=test_pipeline,
+ backend_args=None)
+
+val_evaluator_refcoco_testB = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_refcoco+_testA.json'
+val_dataset_refcoco_plus_testA = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=test_pipeline,
+ backend_args=None)
+
+val_evaluator_refcoco_plus_testA = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_refcoco+_testB.json'
+val_dataset_refcoco_plus_testB = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=test_pipeline,
+ backend_args=None)
+
+val_evaluator_refcoco_plus_testB = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_refcocog_test.json'
+val_dataset_refcocog_test = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=test_pipeline,
+ backend_args=None)
+
+val_evaluator_refcocog_test = dict(
+ type='RefExpMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ topk=(1, 5, 10))
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_grefcoco_val.json'
+val_dataset_grefcoco_val = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=test_pipeline,
+ backend_args=None)
+
+val_evaluator_grefcoco_val = dict(
+ type='gRefCOCOMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ thresh_score=0.7,
+ thresh_f1=1.0)
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_grefcoco_testA.json'
+val_dataset_grefcoco_testA = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=test_pipeline,
+ backend_args=None)
+
+val_evaluator_grefcoco_testA = dict(
+ type='gRefCOCOMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ thresh_score=0.7,
+ thresh_f1=1.0)
+
+# -------------------------------------------------#
+ann_file = 'mdetr_annotations/finetune_grefcoco_testB.json'
+val_dataset_grefcoco_testB = dict(
+ type='MDETRStyleRefCocoDataset',
+ data_root=data_root,
+ ann_file=ann_file,
+ data_prefix=dict(img='train2014/'),
+ test_mode=True,
+ return_classes=True,
+ pipeline=test_pipeline,
+ backend_args=None)
+
+val_evaluator_grefcoco_testB = dict(
+ type='gRefCOCOMetric',
+ ann_file=data_root + ann_file,
+ metric='bbox',
+ iou_thrs=0.5,
+ thresh_score=0.7,
+ thresh_f1=1.0)
+
+# -------------------------------------------------#
+datasets = [
+ val_dataset_all_val, val_dataset_refcoco_testA, val_dataset_refcoco_testB,
+ val_dataset_refcoco_plus_testA, val_dataset_refcoco_plus_testB,
+ val_dataset_refcocog_test, val_dataset_grefcoco_val,
+ val_dataset_grefcoco_testA, val_dataset_grefcoco_testB
+]
+dataset_prefixes = [
+ 'val', 'refcoco_testA', 'refcoco_testB', 'refcoco+_testA',
+ 'refcoco+_testB', 'refcocog_test', 'grefcoco_val', 'grefcoco_testA',
+ 'grefcoco_testB'
+]
+metrics = [
+ val_evaluator_all_val, val_evaluator_refcoco_testA,
+ val_evaluator_refcoco_testB, val_evaluator_refcoco_plus_testA,
+ val_evaluator_refcoco_plus_testB, val_evaluator_refcocog_test,
+ val_evaluator_grefcoco_val, val_evaluator_grefcoco_testA,
+ val_evaluator_grefcoco_testB
+]
+
+val_dataloader = dict(
+ dataset=dict(_delete_=True, type='ConcatDataset', datasets=datasets))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ _delete_=True,
+ type='MultiDatasetsEvaluator',
+ metrics=metrics,
+ dataset_prefixes=dataset_prefixes)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/rtts/grounding_dino_swin-t_finetune_8xb4_1x_rtts.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/rtts/grounding_dino_swin-t_finetune_8xb4_1x_rtts.py
new file mode 100644
index 0000000000000000000000000000000000000000..95c2be058e2c407fc92de93f4b79ec8b36e25c18
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/rtts/grounding_dino_swin-t_finetune_8xb4_1x_rtts.py
@@ -0,0 +1,106 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py'
+
+data_root = 'data/RTTS/'
+class_name = ('bicycle', 'bus', 'car', 'motorbike', 'person')
+palette = [(255, 97, 0), (0, 201, 87), (176, 23, 31), (138, 43, 226),
+ (30, 144, 255)]
+
+metainfo = dict(classes=class_name, palette=palette)
+
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities'))
+]
+
+train_dataloader = dict(
+ sampler=dict(_delete_=True, type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ _delete_=True,
+ type='CocoDataset',
+ data_root=data_root,
+ metainfo=metainfo,
+ filter_cfg=dict(filter_empty_gt=False, min_size=32),
+ pipeline=train_pipeline,
+ return_classes=True,
+ ann_file='annotations_json/rtts_train.json',
+ data_prefix=dict(img='')))
+
+val_dataloader = dict(
+ dataset=dict(
+ metainfo=metainfo,
+ data_root=data_root,
+ return_classes=True,
+ ann_file='annotations_json/rtts_val.json',
+ data_prefix=dict(img='')))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations_json/rtts_val.json',
+ metric='bbox',
+ format_only=False)
+test_evaluator = val_evaluator
+
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0001, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'backbone': dict(lr_mult=0.1)
+ }))
+
+# learning policy
+max_epochs = 12
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[11],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs, val_interval=1)
+default_hooks = dict(checkpoint=dict(max_keep_ckpts=1, save_best='auto'))
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/ruod/grounding_dino_swin-t_finetune_8xb4_1x_ruod.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/ruod/grounding_dino_swin-t_finetune_8xb4_1x_ruod.py
new file mode 100644
index 0000000000000000000000000000000000000000..f57682b29d970fb6d46c2f459f773b03e803695d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/ruod/grounding_dino_swin-t_finetune_8xb4_1x_ruod.py
@@ -0,0 +1,108 @@
+_base_ = '../grounding_dino_swin-t_pretrain_obj365.py'
+
+data_root = 'data/RUOD/'
+class_name = ('holothurian', 'echinus', 'scallop', 'starfish', 'fish',
+ 'corals', 'diver', 'cuttlefish', 'turtle', 'jellyfish')
+palette = [(235, 211, 70), (106, 90, 205), (160, 32, 240), (176, 23, 31),
+ (142, 0, 0), (230, 0, 0), (106, 0, 228), (60, 100, 0), (80, 100, 0),
+ (70, 0, 0)]
+
+metainfo = dict(classes=class_name, palette=palette)
+
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction', 'text',
+ 'custom_entities'))
+]
+
+train_dataloader = dict(
+ sampler=dict(_delete_=True, type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ _delete_=True,
+ type='CocoDataset',
+ data_root=data_root,
+ metainfo=metainfo,
+ filter_cfg=dict(filter_empty_gt=False, min_size=32),
+ pipeline=train_pipeline,
+ return_classes=True,
+ ann_file='RUOD_ANN/instances_train.json',
+ data_prefix=dict(img='RUOD_pic/train/')))
+
+val_dataloader = dict(
+ dataset=dict(
+ metainfo=metainfo,
+ data_root=data_root,
+ return_classes=True,
+ ann_file='RUOD_ANN/instances_test.json',
+ data_prefix=dict(img='RUOD_pic/test/')))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'RUOD_ANN/instances_test.json',
+ metric='bbox',
+ format_only=False)
+test_evaluator = val_evaluator
+
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0001, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'backbone': dict(lr_mult=0.1)
+ }))
+
+# learning policy
+max_epochs = 12
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[11],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs, val_interval=1)
+default_hooks = dict(checkpoint=dict(max_keep_ckpts=1, save_best='auto'))
+
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth' # noqa
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/usage.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/usage.md
new file mode 100644
index 0000000000000000000000000000000000000000..123c6638cbea2cad01d935994f08eab252f35cbf
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/usage.md
@@ -0,0 +1,491 @@
+# Usage
+
+## Install
+
+After installing MMDet according to the instructions in the [get_started](../../docs/zh_cn/get_started.md) section, you need to install additional dependency packages:
+
+```shell
+cd $MMDETROOT
+
+pip install -r requirements/multimodal.txt
+pip install emoji ddd-dataset
+pip install git+https://github.com/lvis-dataset/lvis-api.git"
+```
+
+Please note that since the LVIS third-party library does not currently support numpy 1.24, ensure that your numpy version meets the requirements. It is recommended to install numpy version 1.23.
+
+## Instructions
+
+### Download BERT Weight
+
+MM Grounding DINO uses BERT as its language model and requires access to https://huggingface.co/. If you encounter connection errors due to network access issues, you can download the necessary files on a computer with network access and save them locally. Finally, modify the `lang_model_name` field in the configuration file to the local path. For specific instructions, please refer to the following code:
+
+```python
+from transformers import BertConfig, BertModel
+from transformers import AutoTokenizer
+
+config = BertConfig.from_pretrained("bert-base-uncased")
+model = BertModel.from_pretrained("bert-base-uncased", add_pooling_layer=False, config=config)
+tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
+
+config.save_pretrained("your path/bert-base-uncased")
+model.save_pretrained("your path/bert-base-uncased")
+tokenizer.save_pretrained("your path/bert-base-uncased")
+```
+
+### Download NLTK Weight
+
+When MM Grounding DINO performs Phrase Grounding inference, it may extract noun phrases. Although it downloads specific models at runtime, considering that some users' running environments cannot connect to the internet, it is possible to download them in advance to the `~/nltk_data` path.
+
+```python
+import nltk
+nltk.download('punkt', download_dir='~/nltk_data')
+nltk.download('averaged_perceptron_tagger', download_dir='~/nltk_data')
+```
+
+### Download MM Grounding DINO-T Weight
+
+For convenience in demonstration, you can download the MM Grounding DINO-T model weights in advance to the current path.
+
+```shell
+wget load_from = 'https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth' # noqa
+```
+
+## Inference
+
+Before inference, for a better experience of the inference effects on different images, it is recommended that you first download [these images](https://github.com/microsoft/X-Decoder/tree/main/inference_demo/images) to the current path.
+
+MM Grounding DINO supports four types of inference methods: Closed-Set Object Detection, Open Vocabulary Object Detection, Phrase Grounding, and Referential Expression Comprehension. The details are explained below.
+
+**(1) Closed-Set Object Detection**
+
+Since MM Grounding DINO is a pretrained model, it can theoretically be applied to any closed-set detection dataset. Currently, we support commonly used datasets such as coco/voc/cityscapes/objects365v1/lvis, etc. Below, we will use coco as an example.
+
+```shell
+python demo/image_demo.py images/animals.png \
+ configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py \
+ --weights grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth \
+ --texts '$: coco'
+```
+
+The predictions for `outputs/vis/animals.png` will be generated in the current directory, as shown in the following image.
+
+
+

+
+
+Since ostrich is not one of the 80 classes in COCO, it will not be detected.
+
+It's important to note that Objects365v1 and LVIS have a large number of categories. If you try to input all category names directly into the network, it may exceed 256 tokens, leading to poor model predictions. In such cases, you can use the `--chunked-size` parameter to perform chunked predictions. However, please be aware that chunked predictions may take longer to complete due to the large number of categories.
+
+```shell
+python demo/image_demo.py images/animals.png \
+ configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py \
+ --weights grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth \
+ --texts '$: lvis' --chunked-size 70 \
+ --palette random
+```
+
+
+

+
+
+Different `--chunked-size` values can lead to different prediction results. You can experiment with different chunked sizes to find the one that works best for your specific task and dataset.
+
+**(2) Open Vocabulary Object Detection**
+
+Open vocabulary object detection refers to the ability to input arbitrary class names during inference.
+
+```shell
+python demo/image_demo.py images/animals.png \
+ configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py \
+ --weights grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth \
+ --texts 'zebra. giraffe' -c
+```
+
+
+

+
+
+**(3) Phrase Grounding**
+
+Phrase Grounding refers to the process where a user inputs a natural language description, and the model automatically detects the corresponding bounding boxes for the mentioned noun phrases. It can be used in two ways:
+
+1. Automatically extracting noun phrases using the NLTK library and then performing detection.
+
+```shell
+python demo/image_demo.py images/apples.jpg \
+ configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py \
+ --weights grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth \
+ --texts 'There are many apples here.'
+```
+
+
+

+
+
+The program will automatically split `many apples` as a noun phrase and then detect the corresponding objects. Different input descriptions can have a significant impact on the prediction results.
+
+2. Users can manually specify which parts of the sentence are noun phrases to avoid errors in NLTK extraction.
+
+```shell
+python demo/image_demo.py images/fruit.jpg \
+ configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py \
+ --weights grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth \
+ --texts 'The picture contains watermelon, flower, and a white bottle.' \
+ --tokens-positive "[[[21,31]], [[45,59]]]" --pred-score-thr 0.12
+```
+
+The noun phrase corresponding to positions 21-31 is `watermelon`, and the noun phrase corresponding to positions 45-59 is `a white bottle`.
+
+
+

+
+
+**(4) Referential Expression Comprehension**
+
+Referential expression understanding refers to the model automatically comprehending the referential expressions involved in a user's language description without the need for noun phrase extraction.
+
+```shell
+python demo/image_demo.py images/apples.jpg \
+ configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py \
+ --weights grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth \
+ --texts 'red apple.' \
+ --tokens-positive -1
+```
+
+
+

+
+
+## Evaluation
+
+Our provided evaluation scripts are unified, and you only need to prepare the data in advance and then run the relevant configuration.
+
+(1) Zero-Shot COCO2017 val
+
+```shell
+# single GPU
+python tools/test.py configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py \
+ grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth
+
+# 8 GPUs
+./tools/dist_test.sh configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py \
+ grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth 8
+```
+
+(2) Zero-Shot ODinW13
+
+```shell
+# single GPU
+python tools/test.py configs/mm_grounding_dino/odinw/grounding_dino_swin-t_pretrain_odinw13.py \
+ grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth
+
+# 8 GPUs
+./tools/dist_test.sh configs/mm_grounding_dino/odinw/grounding_dino_swin-t_pretrain_odinw13.py \
+ grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth 8
+```
+
+## Visualization of Evaluation Results
+
+For the convenience of visualizing and analyzing model prediction results, we provide support for visualizing evaluation dataset prediction results. Taking referential expression understanding as an example, the usage is as follows:
+
+```shell
+python tools/test.py configs/mm_grounding_dino/refcoco/grounding_dino_swin-t_pretrain_zeroshot_refexp \
+ grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth --work-dir refcoco_result --show-dir save_path
+```
+
+During the inference process, it will save the visualization results to the `refcoco_result/{current_timestamp}/save_path` directory. For other evaluation dataset visualizations, you only need to replace the configuration file.
+
+Here are some visualization results for various datasets. The left image represents the Ground Truth (GT). The right image represents the Predicted Result.
+
+1. COCO2017 val Results:
+
+
+

+
+
+2. Flickr30k Entities Results:
+
+
+

+
+
+3. DOD Results:
+
+
+

+
+
+4. RefCOCO val Results:
+
+
+

+
+
+5. RefCOCO testA Results:
+
+
+

+
+
+6. gRefCOCO val Results:
+
+
+

+
+
+## Training
+
+If you want to reproduce our results, you can train the model by using the following command after preparing the dataset:
+
+```shell
+# Training on a single machine with 8 GPUs for obj365v1 dataset
+./tools/dist_train.sh configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py 8
+# Training on a single machine with 8 GPUs for datasets like obj365v1, goldg, grit, v3det, and other datasets is similar.
+./tools/dist_train.sh configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det.py 8
+```
+
+For multi-machine training, please refer to [train.md](../../docs/zh_cn/user_guides/train.md). The MM-Grounding-DINO T model is designed to work with 32 GPUs (specifically, 3090Ti GPUs). If your total batch size is not 32x4=128, you will need to manually adjust the learning rate accordingly.
+
+### Pretraining Custom Format Explanation
+
+In order to standardize the pretraining formats for different datasets, we refer to the format design proposed by [Open-GroundingDino](https://github.com/longzw1997/Open-GroundingDino). Specifically, it is divided into two formats.
+
+**(1) Object Detection Format (OD)**
+
+```text
+{"filename": "obj365_train_000000734304.jpg",
+ "height": 512,
+ "width": 769,
+ "detection": {
+ "instances": [
+ {"bbox": [109.4768676992, 346.0190429696, 135.1918335098, 365.3641967616], "label": 2, "category": "chair"},
+ {"bbox": [58.612365705900004, 323.2281494016, 242.6005859067, 451.4166870016], "label": 8, "category": "car"}
+ ]
+ }
+}
+```
+
+The numerical values corresponding to labels in the label dictionary should match the respective label_map. Each item in the instances list corresponds to a bounding box (in the format x1y1x2y2).
+
+**(2) Phrase Grounding Format (VG)**
+
+```text
+{"filename": "2405116.jpg",
+ "height": 375,
+ "width": 500,
+ "grounding":
+ {"caption": "Two surfers walking down the shore. sand on the beach.",
+ "regions": [
+ {"bbox": [206, 156, 282, 248], "phrase": "Two surfers", "tokens_positive": [[0, 3], [4, 11]]},
+ {"bbox": [303, 338, 443, 343], "phrase": "sand", "tokens_positive": [[36, 40]]},
+ {"bbox": [[327, 223, 421, 282], [300, 200, 400, 210]], "phrase": "beach", "tokens_positive": [[48, 53]]}
+ ]
+ }
+```
+
+The `tokens_positive` field indicates the character positions of the current phrase within the caption.
+
+## Example of Fine-tuning Custom Dataset
+
+In order to facilitate downstream fine-tuning on custom datasets, we have provided a fine-tuning example using the simple "cat" dataset as an illustration.
+
+### 1 Data Preparation
+
+```shell
+cd mmdetection
+wget https://download.openmmlab.com/mmyolo/data/cat_dataset.zip
+unzip cat_dataset.zip -d data/cat/
+```
+
+The "cat" dataset is a single-category dataset consisting of 144 images, already converted to the COCO format.
+
+
+

+
+
+### 2 Configuration Preparation
+
+Due to the simplicity and small size of the "cat" dataset, we trained it for 20 epochs using 8 GPUs, with corresponding learning rate scaling. We did not train the language model, only the visual model.
+
+Detailed configuration information can be found in [grounding_dino_swin-t_finetune_8xb4_20e_cat](grounding_dino_swin-t_finetune_8xb4_20e_cat.py).
+
+### 3 Visualization and Evaluation of Zero-Shot Results
+
+Due to MM Grounding DINO being an open-set detection model, you can perform detection and evaluation even if it was not trained on the cat dataset.
+
+Visualization of a single image:
+
+```shell
+cd mmdetection
+python demo/image_demo.py data/cat/images/IMG_20211205_120756.jpg configs/mm_grounding_dino/grounding_dino_swin-t_finetune_8xb4_20e_cat.py --weights grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth --texts cat.
+```
+
+Evaluation results of Zero-shot on test dataset:
+
+```shell
+python tools/test.py configs/mm_grounding_dino/grounding_dino_swin-t_finetune_8xb4_20e_cat.py grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth
+```
+
+```text
+ Average Precision (AP) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.881
+ Average Precision (AP) @[ IoU=0.50 | area= all | maxDets=1000 ] = 1.000
+ Average Precision (AP) @[ IoU=0.75 | area= all | maxDets=1000 ] = 0.929
+ Average Precision (AP) @[ IoU=0.50:0.95 | area= small | maxDets=1000 ] = -1.000
+ Average Precision (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=1000 ] = -1.000
+ Average Precision (AP) @[ IoU=0.50:0.95 | area= large | maxDets=1000 ] = 0.881
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.913
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=300 ] = 0.913
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=1000 ] = 0.913
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= small | maxDets=1000 ] = -1.000
+ Average Recall (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=1000 ] = -1.000
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= large | maxDets=1000 ] = 0.913
+```
+
+### 4 Fine-tuning
+
+```shell
+./tools/dist_train.sh configs/mm_grounding_dino/grounding_dino_swin-t_finetune_8xb4_20e_cat.py 8 --work-dir cat_work_dir
+```
+
+The model will save the best-performing checkpoint. It achieved its best performance at the 16th epoch, with the following results:
+
+```text
+ Average Precision (AP) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.901
+ Average Precision (AP) @[ IoU=0.50 | area= all | maxDets=1000 ] = 1.000
+ Average Precision (AP) @[ IoU=0.75 | area= all | maxDets=1000 ] = 0.930
+ Average Precision (AP) @[ IoU=0.50:0.95 | area= small | maxDets=1000 ] = -1.000
+ Average Precision (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=1000 ] = -1.000
+ Average Precision (AP) @[ IoU=0.50:0.95 | area= large | maxDets=1000 ] = 0.901
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.967
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=300 ] = 0.967
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=1000 ] = 0.967
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= small | maxDets=1000 ] = -1.000
+ Average Recall (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=1000 ] = -1.000
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= large | maxDets=1000 ] = 0.967
+```
+
+We can observe that after fine-tuning, the training performance on the cat dataset improved from 88.1 to 90.1. However, due to the small dataset size, the evaluation metrics show some fluctuations.
+
+## Iterative Generation and Optimization Pipeline of Model Self-training Pseduo Label
+
+To facilitate users in creating their own datasets from scratch or those who want to leverage the model's inference capabilities for iterative pseudo-label generation and optimization, continuously modifying pseudo-labels to improve model performance, we have provided relevant pipelines.
+
+Since we have defined two data formats, we will provide separate explanations for demonstration purposes.
+
+### 1 Object Detection Format
+
+Here, we continue to use the aforementioned cat dataset as an example. Let's assume that we currently have a series of images and predefined categories but no annotations.
+
+1. Generate initial `odvg` format file
+
+```python
+import os
+import cv2
+import json
+import jsonlines
+
+data_root = 'data/cat'
+images_path = os.path.join(data_root, 'images')
+out_path = os.path.join(data_root, 'cat_train_od.json')
+metas = []
+for files in os.listdir(images_path):
+ img = cv2.imread(os.path.join(images_path, files))
+ height, width, _ = img.shape
+ metas.append({"filename": files, "height": height, "width": width})
+
+with jsonlines.open(out_path, mode='w') as writer:
+ writer.write_all(metas)
+
+# 生成 label_map.json,由于只有一个类别,所以只需要写一个 cat 即可
+label_map_path = os.path.join(data_root, 'cat_label_map.json')
+with open(label_map_path, 'w') as f:
+ json.dump({'0': 'cat'}, f)
+```
+
+Two files, `cat_train_od.json` and `cat_label_map.json`, will be generated in the `data/cat` directory.
+
+2. Inference with pre-trained model and save the results
+
+We provide a readily usable [configuration](grounding_dino_swin-t_pretrain_pseudo-labeling_cat.py). If you are using a different dataset, you can refer to this configuration for modifications.
+
+```shell
+python tools/test.py configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_pseudo-labeling_cat.py \
+ grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth
+```
+
+A new file `cat_train_od_v1.json` will be generated in the `data/cat` directory. You can manually open it to confirm or use the provided [script](../../tools/analysis_tools/browse_grounding_raw.py) to visualize the results.
+
+```shell
+python tools/analysis_tools/browse_grounding_raw.py data/cat/ cat_train_od_v1.json images --label-map-file cat_label_map.json -o your_output_dir --not-show
+```
+
+The visualization results will be generated in the `your_output_dir` directory.
+
+3. Continue training to boost performance
+
+After obtaining pseudo-labels, you can mix them with some pre-training data for further pre-training to improve the model's performance on the current dataset. Then, you can repeat step 2 to obtain more accurate pseudo-labels, and continue this iterative process.
+
+### 2 Phrase Grounding Format
+
+1. Generate initial `odvg` format file
+
+The bootstrapping process of Phrase Grounding requires providing captions corresponding to each image and pre-segmented phrase information initially. Taking flickr30k entities images as an example, the generated typical file should look like this:
+
+```text
+[
+{"filename": "3028766968.jpg",
+ "height": 375,
+ "width": 500,
+ "grounding":
+ {"caption": "Man with a black shirt on sit behind a desk sorting threw a giant stack of people work with a smirk on his face .",
+ "regions": [
+ {"bbox": [0, 0, 1, 1], "phrase": "a giant stack of people", "tokens_positive": [[58, 81]]},
+ {"bbox": [0, 0, 1, 1], "phrase": "a black shirt", "tokens_positive": [[9, 22]]},
+ {"bbox": [0, 0, 1, 1], "phrase": "a desk", "tokens_positive": [[37, 43]]},
+ {"bbox": [0, 0, 1, 1], "phrase": "his face", "tokens_positive": [[103, 111]]},
+ {"bbox": [0, 0, 1, 1], "phrase": "Man", "tokens_positive": [[0, 3]]}]}}
+{"filename": "6944134083.jpg",
+ "height": 319,
+ "width": 500,
+ "grounding":
+ {"caption": "Two men are competing in a horse race .",
+ "regions": [
+ {"bbox": [0, 0, 1, 1], "phrase": "Two men", "tokens_positive": [[0, 7]]}]}}
+]
+```
+
+Bbox needs to be set to `[0, 0, 1, 1]` for initialization to make sure the programme could run, but this value would not be utilized.
+
+```text
+{"filename": "3028766968.jpg", "height": 375, "width": 500, "grounding": {"caption": "Man with a black shirt on sit behind a desk sorting threw a giant stack of people work with a smirk on his face .", "regions": [{"bbox": [0, 0, 1, 1], "phrase": "a giant stack of people", "tokens_positive": [[58, 81]]}, {"bbox": [0, 0, 1, 1], "phrase": "a black shirt", "tokens_positive": [[9, 22]]}, {"bbox": [0, 0, 1, 1], "phrase": "a desk", "tokens_positive": [[37, 43]]}, {"bbox": [0, 0, 1, 1], "phrase": "his face", "tokens_positive": [[103, 111]]}, {"bbox": [0, 0, 1, 1], "phrase": "Man", "tokens_positive": [[0, 3]]}]}}
+{"filename": "6944134083.jpg", "height": 319, "width": 500, "grounding": {"caption": "Two men are competing in a horse race .", "regions": [{"bbox": [0, 0, 1, 1], "phrase": "Two men", "tokens_positive": [[0, 7]]}]}}
+```
+
+You can directly copy the text above, and assume that the text content is pasted into a file named `flickr_simple_train_vg.json`, which is placed in the pre-prepared `data/flickr30k_entities` dataset directory, as detailed in the data preparation document.
+
+2. Inference with pre-trained model and save the results
+
+We provide a directly usable [configuration](https://chat.openai.com/c/grounding_dino_swin-t_pretrain_pseudo-labeling_flickr30k.py). If you are using a different dataset, you can refer to this configuration for modifications.
+
+```shell
+python tools/test.py configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_pseudo-labeling_flickr30k.py \
+ grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth
+```
+
+The translation of your text from Chinese to English is: "A new file `flickr_simple_train_vg_v1.json` will be generated in the `data/flickr30k_entities` directory. You can manually open it to confirm or use the [script](../../tools/analysis_tools/browse_grounding_raw.py) to visualize the effects
+
+```shell
+python tools/analysis_tools/browse_grounding_raw.py data/flickr30k_entities/ flickr_simple_train_vg_v1.json flickr30k_images -o your_output_dir --not-show
+```
+
+The visualization results will be generated in the `your_output_dir` directory, as shown in the following image:
+
+
+

+
+
+3. Continue training to boost performance
+
+After obtaining the pseudo-labels, you can mix some pre-training data to continue pre-training jointly, which enhances the model's performance on the current dataset. Then, rerun step 2 to obtain more accurate pseudo-labels, and repeat this cycle iteratively.
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/usage_zh-CN.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/usage_zh-CN.md
new file mode 100644
index 0000000000000000000000000000000000000000..5f625ea6ca8dc09225aebbe00c424fc0128cf736
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/mm_grounding_dino/usage_zh-CN.md
@@ -0,0 +1,491 @@
+# 用法说明
+
+## 安装
+
+在按照 [get_started](../../docs/zh_cn/get_started.md) 一节的说明安装好 MMDet 之后,需要安装额外的依赖包:
+
+```shell
+cd $MMDETROOT
+
+pip install -r requirements/multimodal.txt
+pip install emoji ddd-dataset
+pip install git+https://github.com/lvis-dataset/lvis-api.git"
+```
+
+请注意由于 LVIS 第三方库暂时不支持 numpy 1.24,因此请确保您的 numpy 版本符合要求。建议安装 numpy 1.23 版本。
+
+## 说明
+
+### BERT 权重下载
+
+MM Grounding DINO 采用了 BERT 作为语言模型,需要访问 https://huggingface.co/, 如果您因为网络访问问题遇到连接错误,可以在有网络访问权限的电脑上下载所需文件并保存在本地。最后,修改配置文件中的 `lang_model_name` 字段为本地路径即可。具体请参考以下代码:
+
+```python
+from transformers import BertConfig, BertModel
+from transformers import AutoTokenizer
+
+config = BertConfig.from_pretrained("bert-base-uncased")
+model = BertModel.from_pretrained("bert-base-uncased", add_pooling_layer=False, config=config)
+tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
+
+config.save_pretrained("your path/bert-base-uncased")
+model.save_pretrained("your path/bert-base-uncased")
+tokenizer.save_pretrained("your path/bert-base-uncased")
+```
+
+### NLTK 权重下载
+
+MM Grounding DINO 在进行 Phrase Grounding 推理时候可能会进行名词短语提取,虽然会在运行时候下载特定的模型,但是考虑到有些用户运行环境无法联网,因此可以提前下载到 `~/nltk_data` 路径下
+
+```python
+import nltk
+nltk.download('punkt', download_dir='~/nltk_data')
+nltk.download('averaged_perceptron_tagger', download_dir='~/nltk_data')
+```
+
+### MM Grounding DINO-T 模型权重下载
+
+为了方便演示,您可以提前下载 MM Grounding DINO-T 模型权重到当前路径下
+
+```shell
+wget load_from = 'https://download.openmmlab.com/mmdetection/v3.0/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth' # noqa
+```
+
+## 推理
+
+在推理前,为了更好的体验不同图片的推理效果,建议您先下载 [这些图片](https://github.com/microsoft/X-Decoder/tree/main/inference_demo/images) 到当前路径下
+
+MM Grounding DINO 支持了闭集目标检测,开放词汇目标检测,Phrase Grounding 和指代性表达式理解 4 种推理方式,下面详细说明。
+
+**(1) 闭集目标检测**
+
+由于 MM Grounding DINO 是预训练模型,理论上可以应用于任何闭集检测数据集,目前我们支持了常用的 coco/voc/cityscapes/objects365v1/lvis 等,下面以 coco 为例
+
+```shell
+python demo/image_demo.py images/animals.png \
+ configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py \
+ --weights grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth \
+ --texts '$: coco'
+```
+
+会在当前路径下生成 `outputs/vis/animals.png` 的预测结果,如下图所示
+
+
+

+
+
+由于鸵鸟并不在 COCO 80 类中, 因此不会检测出来。
+
+需要注意,由于 objects365v1 和 lvis 类别很多,如果直接将类别名全部输入到网络中,会超过 256 个 token 导致模型预测效果极差,此时我们需要通过 `--chunked-size` 参数进行截断预测, 同时预测时间会比较长。
+
+```shell
+python demo/image_demo.py images/animals.png \
+ configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py \
+ --weights grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth \
+ --texts '$: lvis' --chunked-size 70 \
+ --palette random
+```
+
+
+

+
+
+不同的 `--chunked-size` 会导致不同的预测效果,您可以自行尝试。
+
+**(2) 开放词汇目标检测**
+
+开放词汇目标检测是指在推理时候,可以输入任意的类别名
+
+```shell
+python demo/image_demo.py images/animals.png \
+ configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py \
+ --weights grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth \
+ --texts 'zebra. giraffe' -c
+```
+
+
+

+
+
+**(3) Phrase Grounding**
+
+Phrase Grounding 是指的用户输入一句语言描述,模型自动对其涉及到的名词短语想对应的 bbox 进行检测,有两种用法
+
+1. 通过 NLTK 库自动提取名词短语,然后进行检测
+
+```shell
+python demo/image_demo.py images/apples.jpg \
+ configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py \
+ --weights grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth \
+ --texts 'There are many apples here.'
+```
+
+
+

+
+
+程序内部会自动切分出 `many apples` 作为名词短语,然后检测出对应物体。不同的输入描述对预测结果影响很大。
+
+2. 用户自己指定句子中哪些为名词短语,避免 NLTK 提取错误的情况
+
+```shell
+python demo/image_demo.py images/fruit.jpg \
+ configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py \
+ --weights grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth \
+ --texts 'The picture contains watermelon, flower, and a white bottle.' \
+ --tokens-positive "[[[21,31]], [[45,59]]]" --pred-score-thr 0.12
+```
+
+21,31 对应的名词短语为 `watermelon`,45,59 对应的名词短语为 `a white bottle`。
+
+
+

+
+
+**(4) 指代性表达式理解**
+
+指代性表达式理解是指的用户输入一句语言描述,模型自动对其涉及到的指代性表达式进行理解, 不需要进行名词短语提取。
+
+```shell
+python demo/image_demo.py images/apples.jpg \
+ configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py \
+ --weights grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth \
+ --texts 'red apple.' \
+ --tokens-positive -1
+```
+
+
+

+
+
+## 评测
+
+我们所提供的评测脚本都是统一的,你只需要提前准备好数据,然后运行相关配置就可以了
+
+(1) Zero-Shot COCO2017 val
+
+```shell
+# 单卡
+python tools/test.py configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py \
+ grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth
+
+# 8 卡
+./tools/dist_test.sh configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py \
+ grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth 8
+```
+
+(2) Zero-Shot ODinW13
+
+```shell
+# 单卡
+python tools/test.py configs/mm_grounding_dino/odinw/grounding_dino_swin-t_pretrain_odinw13.py \
+ grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth
+
+# 8 卡
+./tools/dist_test.sh configs/mm_grounding_dino/odinw/grounding_dino_swin-t_pretrain_odinw13.py \
+ grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth 8
+```
+
+## 评测数据集结果可视化
+
+为了方便大家对模型预测结果进行可视化和分析,我们支持了评测数据集预测结果可视化,以指代性表达式理解为例用法如下:
+
+```shell
+python tools/test.py configs/mm_grounding_dino/refcoco/grounding_dino_swin-t_pretrain_zeroshot_refexp \
+ grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth --work-dir refcoco_result --show-dir save_path
+```
+
+模型在推理过程中会将可视化结果保存到 `refcoco_result/{当前时间戳}/save_path` 路径下。其余评测数据集可视化只需要替换配置文件即可。
+
+下面展示一些数据集的可视化结果: 左图为 GT,右图为预测结果
+
+1. COCO2017 val 结果:
+
+
+

+
+
+2. Flickr30k Entities 结果:
+
+
+

+
+
+3. DOD 结果:
+
+
+

+
+
+4. RefCOCO val 结果:
+
+
+

+
+
+5. RefCOCO testA 结果:
+
+
+

+
+
+6. gRefCOCO val 结果:
+
+
+

+
+
+## 模型训练
+
+如果想复现我们的结果,你可以在准备好数据集后,直接通过如下命令进行训练
+
+```shell
+# 单机 8 卡训练仅包括 obj365v1 数据集
+./tools/dist_train.sh configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365.py 8
+# 单机 8 卡训练包括 obj365v1/goldg/grit/v3det 数据集,其余数据集类似
+./tools/dist_train.sh configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det.py 8
+```
+
+多机训练的用法请参考 [train.md](../../docs/zh_cn/user_guides/train.md)。MM-Grounding-DINO T 模型默认采用的是 32 张 3090Ti,如果你的总 bs 数不是 32x4=128,那么你需要手动的线性调整学习率。
+
+### 预训练自定义格式说明
+
+为了统一不同数据集的预训练格式,我们参考 [Open-GroundingDino](https://github.com/longzw1997/Open-GroundingDino) 所设计的格式。具体来说分成 2 种格式
+
+**(1) 目标检测数据格式 OD**
+
+```text
+{"filename": "obj365_train_000000734304.jpg",
+ "height": 512,
+ "width": 769,
+ "detection": {
+ "instances": [
+ {"bbox": [109.4768676992, 346.0190429696, 135.1918335098, 365.3641967616], "label": 2, "category": "chair"},
+ {"bbox": [58.612365705900004, 323.2281494016, 242.6005859067, 451.4166870016], "label": 8, "category": "car"}
+ ]
+ }
+}
+```
+
+label字典中所对应的数值需要和相应的 label_map 一致。 instances 列表中的每一项都对应一个 bbox (x1y1x2y2 格式)。
+
+**(2) phrase grounding 数据格式 VG**
+
+```text
+{"filename": "2405116.jpg",
+ "height": 375,
+ "width": 500,
+ "grounding":
+ {"caption": "Two surfers walking down the shore. sand on the beach.",
+ "regions": [
+ {"bbox": [206, 156, 282, 248], "phrase": "Two surfers", "tokens_positive": [[0, 3], [4, 11]]},
+ {"bbox": [303, 338, 443, 343], "phrase": "sand", "tokens_positive": [[36, 40]]},
+ {"bbox": [[327, 223, 421, 282], [300, 200, 400, 210]], "phrase": "beach", "tokens_positive": [[48, 53]]}
+ ]
+ }
+```
+
+tokens_positive 表示当前 phrase 在 caption 中的字符位置。
+
+## 自定义数据集微调训练案例
+
+为了方便用户针对自定义数据集进行下游微调,我们特意提供了以简单的 cat 数据集为例的微调训练案例。
+
+### 1 数据准备
+
+```shell
+cd mmdetection
+wget https://download.openmmlab.com/mmyolo/data/cat_dataset.zip
+unzip cat_dataset.zip -d data/cat/
+```
+
+cat 数据集是一个单类别数据集,包含 144 张图片,已经转换为 coco 格式。
+
+
+

+
+
+### 2 配置准备
+
+由于 cat 数据集的简单性和数量较少,我们使用 8 卡训练 20 个 epoch,相应的缩放学习率,不训练语言模型,只训练视觉模型。
+
+详细的配置信息可以在 [grounding_dino_swin-t_finetune_8xb4_20e_cat](grounding_dino_swin-t_finetune_8xb4_20e_cat.py) 中找到。
+
+### 3 可视化和 Zero-Shot 评估
+
+由于 MM Grounding DINO 是一个开放的检测模型,所以即使没有在 cat 数据集上训练,也可以进行检测和评估。
+
+单张图片的可视化结果如下:
+
+```shell
+cd mmdetection
+python demo/image_demo.py data/cat/images/IMG_20211205_120756.jpg configs/mm_grounding_dino/grounding_dino_swin-t_finetune_8xb4_20e_cat.py --weights grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth --texts cat.
+```
+
+测试集上的 Zero-Shot 评估结果如下:
+
+```shell
+python tools/test.py configs/mm_grounding_dino/grounding_dino_swin-t_finetune_8xb4_20e_cat.py grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth
+```
+
+```text
+ Average Precision (AP) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.881
+ Average Precision (AP) @[ IoU=0.50 | area= all | maxDets=1000 ] = 1.000
+ Average Precision (AP) @[ IoU=0.75 | area= all | maxDets=1000 ] = 0.929
+ Average Precision (AP) @[ IoU=0.50:0.95 | area= small | maxDets=1000 ] = -1.000
+ Average Precision (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=1000 ] = -1.000
+ Average Precision (AP) @[ IoU=0.50:0.95 | area= large | maxDets=1000 ] = 0.881
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.913
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=300 ] = 0.913
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=1000 ] = 0.913
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= small | maxDets=1000 ] = -1.000
+ Average Recall (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=1000 ] = -1.000
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= large | maxDets=1000 ] = 0.913
+```
+
+### 4 模型训练
+
+```shell
+./tools/dist_train.sh configs/mm_grounding_dino/grounding_dino_swin-t_finetune_8xb4_20e_cat.py 8 --work-dir cat_work_dir
+```
+
+模型将会保存性能最佳的模型。在第 16 epoch 时候达到最佳,性能如下所示:
+
+```text
+ Average Precision (AP) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.901
+ Average Precision (AP) @[ IoU=0.50 | area= all | maxDets=1000 ] = 1.000
+ Average Precision (AP) @[ IoU=0.75 | area= all | maxDets=1000 ] = 0.930
+ Average Precision (AP) @[ IoU=0.50:0.95 | area= small | maxDets=1000 ] = -1.000
+ Average Precision (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=1000 ] = -1.000
+ Average Precision (AP) @[ IoU=0.50:0.95 | area= large | maxDets=1000 ] = 0.901
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.967
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=300 ] = 0.967
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets=1000 ] = 0.967
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= small | maxDets=1000 ] = -1.000
+ Average Recall (AR) @[ IoU=0.50:0.95 | area=medium | maxDets=1000 ] = -1.000
+ Average Recall (AR) @[ IoU=0.50:0.95 | area= large | maxDets=1000 ] = 0.967
+```
+
+我们可以发现,经过微调训练后,cat 数据集的训练性能从 88.1 提升到了 90.1。同时由于数据集比较小,评估指标波动比较大。
+
+## 模型自训练伪标签迭代生成和优化 pipeline
+
+为了方便用户从头构建自己的数据集或者希望利用模型推理能力进行自举式伪标签迭代生成和优化,不断修改伪标签来提升模型性能,我们特意提供了相关的 pipeline。
+
+由于我们定义了两种数据格式,为了演示我们也将分别进行说明。
+
+### 1 目标检测格式
+
+此处我们依然采用上述的 cat 数据集为例,假设我们目前只有一系列图片和预定义的类别,并不存在标注。
+
+1. 生成初始 odvg 格式文件
+
+```python
+import os
+import cv2
+import json
+import jsonlines
+
+data_root = 'data/cat'
+images_path = os.path.join(data_root, 'images')
+out_path = os.path.join(data_root, 'cat_train_od.json')
+metas = []
+for files in os.listdir(images_path):
+ img = cv2.imread(os.path.join(images_path, files))
+ height, width, _ = img.shape
+ metas.append({"filename": files, "height": height, "width": width})
+
+with jsonlines.open(out_path, mode='w') as writer:
+ writer.write_all(metas)
+
+# 生成 label_map.json,由于只有一个类别,所以只需要写一个 cat 即可
+label_map_path = os.path.join(data_root, 'cat_label_map.json')
+with open(label_map_path, 'w') as f:
+ json.dump({'0': 'cat'}, f)
+```
+
+会在 `data/cat` 目录下生成 `cat_train_od.json` 和 `cat_label_map.json` 两个文件。
+
+2. 使用预训练模型进行推理,并保存结果
+
+我们提供了直接可用的 [配置](grounding_dino_swin-t_pretrain_pseudo-labeling_cat.py), 如果你是其他数据集可以参考这个配置进行修改。
+
+```shell
+python tools/test.py configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_pseudo-labeling_cat.py \
+ grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth
+```
+
+会在 `data/cat` 目录下新生成 `cat_train_od_v1.json` 文件,你可以手动打开确认或者使用 [脚本](../../tools/analysis_tools/browse_grounding_raw.py) 可视化效果
+
+```shell
+python tools/analysis_tools/browse_grounding_raw.py data/cat/ cat_train_od_v1.json images --label-map-file cat_label_map.json -o your_output_dir --not-show
+```
+
+会在 your_output_dir 目录下生成可视化结果
+
+3. 继续训练提高性能
+
+在得到伪标签后,你可以混合一些预训练数据联合进行继续预训练,提升模型在当前数据集上的性能,然后重新运行 2 步骤,得到更准确的伪标签,如此循环迭代即可。
+
+### 2 Phrase Grounding 格式
+
+1. 生成初始 odvg 格式文件
+
+Phrase Grounding 的自举流程要求初始时候提供每张图片对应的 caption 和提前切割好的 phrase 信息。以 flickr30k entities 图片为例,生成的典型的文件应该如下所示:
+
+```text
+[
+{"filename": "3028766968.jpg",
+ "height": 375,
+ "width": 500,
+ "grounding":
+ {"caption": "Man with a black shirt on sit behind a desk sorting threw a giant stack of people work with a smirk on his face .",
+ "regions": [
+ {"bbox": [0, 0, 1, 1], "phrase": "a giant stack of people", "tokens_positive": [[58, 81]]},
+ {"bbox": [0, 0, 1, 1], "phrase": "a black shirt", "tokens_positive": [[9, 22]]},
+ {"bbox": [0, 0, 1, 1], "phrase": "a desk", "tokens_positive": [[37, 43]]},
+ {"bbox": [0, 0, 1, 1], "phrase": "his face", "tokens_positive": [[103, 111]]},
+ {"bbox": [0, 0, 1, 1], "phrase": "Man", "tokens_positive": [[0, 3]]}]}}
+{"filename": "6944134083.jpg",
+ "height": 319,
+ "width": 500,
+ "grounding":
+ {"caption": "Two men are competing in a horse race .",
+ "regions": [
+ {"bbox": [0, 0, 1, 1], "phrase": "Two men", "tokens_positive": [[0, 7]]}]}}
+]
+```
+
+初始时候 bbox 必须要设置为 `[0, 0, 1, 1]`,因为这能确保程序正常运行,但是 bbox 的值并不会被使用。
+
+```text
+{"filename": "3028766968.jpg", "height": 375, "width": 500, "grounding": {"caption": "Man with a black shirt on sit behind a desk sorting threw a giant stack of people work with a smirk on his face .", "regions": [{"bbox": [0, 0, 1, 1], "phrase": "a giant stack of people", "tokens_positive": [[58, 81]]}, {"bbox": [0, 0, 1, 1], "phrase": "a black shirt", "tokens_positive": [[9, 22]]}, {"bbox": [0, 0, 1, 1], "phrase": "a desk", "tokens_positive": [[37, 43]]}, {"bbox": [0, 0, 1, 1], "phrase": "his face", "tokens_positive": [[103, 111]]}, {"bbox": [0, 0, 1, 1], "phrase": "Man", "tokens_positive": [[0, 3]]}]}}
+{"filename": "6944134083.jpg", "height": 319, "width": 500, "grounding": {"caption": "Two men are competing in a horse race .", "regions": [{"bbox": [0, 0, 1, 1], "phrase": "Two men", "tokens_positive": [[0, 7]]}]}}
+```
+
+你可直接复制上面的文本,并假设将文本内容粘贴到命名为 `flickr_simple_train_vg.json` 文件中,并放置于提前准备好的 `data/flickr30k_entities` 数据集目录下,具体见数据准备文档。
+
+2. 使用预训练模型进行推理,并保存结果
+
+我们提供了直接可用的 [配置](grounding_dino_swin-t_pretrain_pseudo-labeling_flickr30k.py), 如果你是其他数据集可以参考这个配置进行修改。
+
+```shell
+python tools/test.py configs/mm_grounding_dino/grounding_dino_swin-t_pretrain_pseudo-labeling_flickr30k.py \
+ grounding_dino_swin-t_pretrain_obj365_goldg_grit9m_v3det_20231204_095047-b448804b.pth
+```
+
+会在 `data/flickr30k_entities` 目录下新生成 `flickr_simple_train_vg_v1.json` 文件,你可以手动打开确认或者使用 [脚本](../../tools/analysis_tools/browse_grounding_raw.py) 可视化效果
+
+```shell
+python tools/analysis_tools/browse_grounding_raw.py data/flickr30k_entities/ flickr_simple_train_vg_v1.json flickr30k_images -o your_output_dir --not-show
+```
+
+会在 `your_output_dir` 目录下生成可视化结果,如下图所示:
+
+
+

+
+
+3. 继续训练提高性能
+
+在得到伪标签后,你可以混合一些预训练数据联合进行继续预训练,提升模型在当前数据集上的性能,然后重新运行 2 步骤,得到更准确的伪标签,如此循环迭代即可。
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..abbec9b6851ee135f61a82b82a7a58423b204b97
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/README.md
@@ -0,0 +1,36 @@
+# MS R-CNN
+
+> [Mask Scoring R-CNN](https://arxiv.org/abs/1903.00241)
+
+
+
+## Abstract
+
+Letting a deep network be aware of the quality of its own predictions is an interesting yet important problem. In the task of instance segmentation, the confidence of instance classification is used as mask quality score in most instance segmentation frameworks. However, the mask quality, quantified as the IoU between the instance mask and its ground truth, is usually not well correlated with classification score. In this paper, we study this problem and propose Mask Scoring R-CNN which contains a network block to learn the quality of the predicted instance masks. The proposed network block takes the instance feature and the corresponding predicted mask together to regress the mask IoU. The mask scoring strategy calibrates the misalignment between mask quality and mask score, and improves instance segmentation performance by prioritizing more accurate mask predictions during COCO AP evaluation. By extensive evaluations on the COCO dataset, Mask Scoring R-CNN brings consistent and noticeable gain with different models, and outperforms the state-of-the-art Mask R-CNN. We hope our simple and effective approach will provide a new direction for improving instance segmentation.
+
+
+

+
+
+## Results and Models
+
+| Backbone | style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :----------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :-------------------------------------------: | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | caffe | 1x | 4.5 | | 38.2 | 36.0 | [config](./ms-rcnn_r50-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_r50_caffe_fpn_1x_coco/ms_rcnn_r50_caffe_fpn_1x_coco_20200702_180848-61c9355e.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_r50_caffe_fpn_1x_coco/ms_rcnn_r50_caffe_fpn_1x_coco_20200702_180848.log.json) |
+| R-50-FPN | caffe | 2x | - | - | 38.8 | 36.3 | [config](./ms-rcnn_r50-caffe_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_r50_caffe_fpn_2x_coco/ms_rcnn_r50_caffe_fpn_2x_coco_bbox_mAP-0.388__segm_mAP-0.363_20200506_004738-ee87b137.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_r50_caffe_fpn_2x_coco/ms_rcnn_r50_caffe_fpn_2x_coco_20200506_004738.log.json) |
+| R-101-FPN | caffe | 1x | 6.5 | | 40.4 | 37.6 | [config](./ms-rcnn_r101-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_r101_caffe_fpn_1x_coco/ms_rcnn_r101_caffe_fpn_1x_coco_bbox_mAP-0.404__segm_mAP-0.376_20200506_004755-b9b12a37.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_r101_caffe_fpn_1x_coco/ms_rcnn_r101_caffe_fpn_1x_coco_20200506_004755.log.json) |
+| R-101-FPN | caffe | 2x | - | - | 41.1 | 38.1 | [config](./ms-rcnn_r101-caffe_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_r101_caffe_fpn_2x_coco/ms_rcnn_r101_caffe_fpn_2x_coco_bbox_mAP-0.411__segm_mAP-0.381_20200506_011134-5f3cc74f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_r101_caffe_fpn_2x_coco/ms_rcnn_r101_caffe_fpn_2x_coco_20200506_011134.log.json) |
+| R-X101-32x4d | pytorch | 2x | 7.9 | 11.0 | 41.8 | 38.7 | [config](./ms-rcnn_x101-32x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_x101_32x4d_fpn_1x_coco/ms_rcnn_x101_32x4d_fpn_1x_coco_20200206-81fd1740.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_x101_32x4d_fpn_1x_coco/ms_rcnn_x101_32x4d_fpn_1x_coco_20200206_100113.log.json) |
+| R-X101-64x4d | pytorch | 1x | 11.0 | 8.0 | 43.0 | 39.5 | [config](./ms-rcnn_x101-64x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_x101_64x4d_fpn_1x_coco/ms_rcnn_x101_64x4d_fpn_1x_coco_20200206-86ba88d2.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_x101_64x4d_fpn_1x_coco/ms_rcnn_x101_64x4d_fpn_1x_coco_20200206_091744.log.json) |
+| R-X101-64x4d | pytorch | 2x | 11.0 | 8.0 | 42.6 | 39.5 | [config](./ms-rcnn_x101-64x4d_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_x101_64x4d_fpn_2x_coco/ms_rcnn_x101_64x4d_fpn_2x_coco_20200308-02a445e2.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_x101_64x4d_fpn_2x_coco/ms_rcnn_x101_64x4d_fpn_2x_coco_20200308_012247.log.json) |
+
+## Citation
+
+```latex
+@inproceedings{huang2019msrcnn,
+ title={Mask Scoring R-CNN},
+ author={Zhaojin Huang and Lichao Huang and Yongchao Gong and Chang Huang and Xinggang Wang},
+ booktitle={IEEE Conference on Computer Vision and Pattern Recognition},
+ year={2019},
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..290f05436949c68d226d8bc2f107e480acbd6b4c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/metafile.yml
@@ -0,0 +1,159 @@
+Collections:
+ - Name: Mask Scoring R-CNN
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RPN
+ - FPN
+ - ResNet
+ - RoIAlign
+ Paper:
+ URL: https://arxiv.org/abs/1903.00241
+ Title: 'Mask Scoring R-CNN'
+ README: configs/ms_rcnn/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/detectors/mask_scoring_rcnn.py#L6
+ Version: v2.0.0
+
+Models:
+ - Name: ms-rcnn_r50-caffe_fpn_1x_coco
+ In Collection: Mask Scoring R-CNN
+ Config: configs/ms_rcnn/ms-rcnn_r50-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.5
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_r50_caffe_fpn_1x_coco/ms_rcnn_r50_caffe_fpn_1x_coco_20200702_180848-61c9355e.pth
+
+ - Name: ms-rcnn_r50-caffe_fpn_2x_coco
+ In Collection: Mask Scoring R-CNN
+ Config: configs/ms_rcnn/ms-rcnn_r50-caffe_fpn_2x_coco.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_r50_caffe_fpn_2x_coco/ms_rcnn_r50_caffe_fpn_2x_coco_bbox_mAP-0.388__segm_mAP-0.363_20200506_004738-ee87b137.pth
+
+ - Name: ms-rcnn_r101-caffe_fpn_1x_coco
+ In Collection: Mask Scoring R-CNN
+ Config: configs/ms_rcnn/ms-rcnn_r101-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.5
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_r101_caffe_fpn_1x_coco/ms_rcnn_r101_caffe_fpn_1x_coco_bbox_mAP-0.404__segm_mAP-0.376_20200506_004755-b9b12a37.pth
+
+ - Name: ms-rcnn_r101-caffe_fpn_2x_coco
+ In Collection: Mask Scoring R-CNN
+ Config: configs/ms_rcnn/ms-rcnn_r101-caffe_fpn_2x_coco.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.1
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_r101_caffe_fpn_2x_coco/ms_rcnn_r101_caffe_fpn_2x_coco_bbox_mAP-0.411__segm_mAP-0.381_20200506_011134-5f3cc74f.pth
+
+ - Name: ms-rcnn_x101-32x4d_fpn_1x_coco
+ In Collection: Mask Scoring R-CNN
+ Config: configs/ms_rcnn/ms-rcnn_x101-32x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.9
+ inference time (ms/im):
+ - value: 90.91
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_x101_32x4d_fpn_1x_coco/ms_rcnn_x101_32x4d_fpn_1x_coco_20200206-81fd1740.pth
+
+ - Name: ms-rcnn_x101-64x4d_fpn_1x_coco
+ In Collection: Mask Scoring R-CNN
+ Config: configs/ms_rcnn/ms-rcnn_x101-64x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 11.0
+ inference time (ms/im):
+ - value: 125
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_x101_64x4d_fpn_1x_coco/ms_rcnn_x101_64x4d_fpn_1x_coco_20200206-86ba88d2.pth
+
+ - Name: ms-rcnn_x101-64x4d_fpn_2x_coco
+ In Collection: Mask Scoring R-CNN
+ Config: configs/ms_rcnn/ms-rcnn_x101-64x4d_fpn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 11.0
+ inference time (ms/im):
+ - value: 125
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.6
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/ms_rcnn/ms_rcnn_x101_64x4d_fpn_2x_coco/ms_rcnn_x101_64x4d_fpn_2x_coco_20200308-02a445e2.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_r101-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_r101-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..2ff4f2d66ae6de88ba9d5d8fb5cf31abaa4cb3c5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_r101-caffe_fpn_1x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './ms-rcnn_r50-caffe_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet101_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_r101-caffe_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_r101-caffe_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..54b29e4f7aea547e2b26782b71ada8053930d325
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_r101-caffe_fpn_2x_coco.py
@@ -0,0 +1,17 @@
+_base_ = './ms-rcnn_r101-caffe_fpn_1x_coco.py'
+# learning policy
+max_epochs = 24
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_r50-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_r50-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e7fbc51f1ba431ca7c22ff3d2c74cfc9e1263ffb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_r50-caffe_fpn_1x_coco.py
@@ -0,0 +1,16 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50-caffe_fpn_1x_coco.py'
+model = dict(
+ type='MaskScoringRCNN',
+ roi_head=dict(
+ type='MaskScoringRoIHead',
+ mask_iou_head=dict(
+ type='MaskIoUHead',
+ num_convs=4,
+ num_fcs=2,
+ roi_feat_size=14,
+ in_channels=256,
+ conv_out_channels=256,
+ fc_out_channels=1024,
+ num_classes=80)),
+ # model training and testing settings
+ train_cfg=dict(rcnn=dict(mask_thr_binary=0.5)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_r50-caffe_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_r50-caffe_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..033488229220e5b044c30c43f5e72f8468f68224
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_r50-caffe_fpn_2x_coco.py
@@ -0,0 +1,17 @@
+_base_ = './ms-rcnn_r50-caffe_fpn_1x_coco.py'
+# learning policy
+max_epochs = 24
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..0ae47d1c38daa4430de4b4264bbb2aef0eb7f7ea
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_r50_fpn_1x_coco.py
@@ -0,0 +1,16 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ type='MaskScoringRCNN',
+ roi_head=dict(
+ type='MaskScoringRoIHead',
+ mask_iou_head=dict(
+ type='MaskIoUHead',
+ num_convs=4,
+ num_fcs=2,
+ roi_feat_size=14,
+ in_channels=256,
+ conv_out_channels=256,
+ fc_out_channels=1024,
+ num_classes=80)),
+ # model training and testing settings
+ train_cfg=dict(rcnn=dict(mask_thr_binary=0.5)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_x101-32x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_x101-32x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1a5d0d0f3188e8e661cc9ab7a731fc631dd950ac
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_x101-32x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './ms-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_x101-64x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_x101-64x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..16290076c07d7a97108b89e4a41b5ff51cbbcdc1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_x101-64x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './ms-rcnn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_x101-64x4d_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_x101-64x4d_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..7aec1874394692a63dc8caeef2609cf01b7bfd7c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ms_rcnn/ms-rcnn_x101-64x4d_fpn_2x_coco.py
@@ -0,0 +1,17 @@
+_base_ = './ms-rcnn_x101-64x4d_fpn_1x_coco.py'
+# learning policy
+max_epochs = 24
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fcos/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fcos/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..a0ec77c8f118f8aeb47ef4cb0efb0022790fa270
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fcos/README.md
@@ -0,0 +1,35 @@
+# NAS-FCOS
+
+> [NAS-FCOS: Fast Neural Architecture Search for Object Detection](https://arxiv.org/abs/1906.04423)
+
+
+
+## Abstract
+
+The success of deep neural networks relies on significant architecture engineering. Recently neural architecture search (NAS) has emerged as a promise to greatly reduce manual effort in network design by automatically searching for optimal architectures, although typically such algorithms need an excessive amount of computational resources, e.g., a few thousand GPU-days. To date, on challenging vision tasks such as object detection, NAS, especially fast versions of NAS, is less studied. Here we propose to search for the decoder structure of object detectors with search efficiency being taken into consideration. To be more specific, we aim to efficiently search for the feature pyramid network (FPN) as well as the prediction head of a simple anchor-free object detector, namely FCOS, using a tailored reinforcement learning paradigm. With carefully designed search space, search algorithms and strategies for evaluating network quality, we are able to efficiently search a top-performing detection architecture within 4 days using 8 V100 GPUs. The discovered architecture surpasses state-of-the-art object detection models (such as Faster R-CNN, RetinaNet and FCOS) by 1.5 to 3.5 points in AP on the COCO dataset, with comparable computation complexity and memory footprint, demonstrating the efficacy of the proposed NAS for object detection.
+
+
+

+
+
+## Results and Models
+
+| Head | Backbone | Style | GN-head | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :----------: | :------: | :---: | :-----: | :-----: | :------: | :------------: | :----: | :-----------------------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| NAS-FCOSHead | R-50 | caffe | Y | 1x | | | 39.4 | [config](./nas-fcos_r50-caffe_fpn_nashead-gn-head_4xb4-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/nas_fcos/nas_fcos_nashead_r50_caffe_fpn_gn-head_4x4_1x_coco/nas_fcos_nashead_r50_caffe_fpn_gn-head_4x4_1x_coco_20200520-1bdba3ce.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/nas_fcos/nas_fcos_nashead_r50_caffe_fpn_gn-head_4x4_1x_coco/nas_fcos_nashead_r50_caffe_fpn_gn-head_4x4_1x_coco_20200520.log.json) |
+| FCOSHead | R-50 | caffe | Y | 1x | | | 38.5 | [config](./nas-fcos_r50-caffe_fpn_fcoshead-gn-head_4xb4-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/nas_fcos/nas_fcos_fcoshead_r50_caffe_fpn_gn-head_4x4_1x_coco/nas_fcos_fcoshead_r50_caffe_fpn_gn-head_4x4_1x_coco_20200521-7fdcbce0.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/nas_fcos/nas_fcos_fcoshead_r50_caffe_fpn_gn-head_4x4_1x_coco/nas_fcos_fcoshead_r50_caffe_fpn_gn-head_4x4_1x_coco_20200521.log.json) |
+
+**Notes:**
+
+- To be consistent with the author's implementation, we use 4 GPUs with 4 images/GPU.
+
+## Citation
+
+```latex
+@article{wang2019fcos,
+ title={Nas-fcos: Fast neural architecture search for object detection},
+ author={Wang, Ning and Gao, Yang and Chen, Hao and Wang, Peng and Tian, Zhi and Shen, Chunhua},
+ journal={arXiv preprint arXiv:1906.04423},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fcos/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fcos/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..02292a41516b6b2d5ab87e629f2bd2672e61e0fb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fcos/metafile.yml
@@ -0,0 +1,44 @@
+Collections:
+ - Name: NAS-FCOS
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 4x V100 GPUs
+ Architecture:
+ - FPN
+ - NAS-FCOS
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/1906.04423
+ Title: 'NAS-FCOS: Fast Neural Architecture Search for Object Detection'
+ README: configs/nas_fcos/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/detectors/nasfcos.py#L6
+ Version: v2.1.0
+
+Models:
+ - Name: nas-fcos_r50-caffe_fpn_nashead-gn-head_4xb4-1x_coco
+ In Collection: NAS-FCOS
+ Config: configs/nas_fcos/nas-fcos_r50-caffe_fpn_nashead-gn-head_4xb4-1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/nas_fcos/nas_fcos_nashead_r50_caffe_fpn_gn-head_4x4_1x_coco/nas_fcos_nashead_r50_caffe_fpn_gn-head_4x4_1x_coco_20200520-1bdba3ce.pth
+
+ - Name: nas-fcos_r50-caffe_fpn_fcoshead-gn-head_4xb4-1x_coco
+ In Collection: NAS-FCOS
+ Config: configs/nas_fcos/nas-fcos_r50-caffe_fpn_fcoshead-gn-head_4xb4-1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/nas_fcos/nas_fcos_fcoshead_r50_caffe_fpn_gn-head_4x4_1x_coco/nas_fcos_fcoshead_r50_caffe_fpn_gn-head_4x4_1x_coco_20200521-7fdcbce0.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fcos/nas-fcos_r50-caffe_fpn_fcoshead-gn-head_4xb4-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fcos/nas-fcos_r50-caffe_fpn_fcoshead-gn-head_4xb4-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ba207c9fbdddc5cd30e4d4d86add2c98664e7ffb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fcos/nas-fcos_r50-caffe_fpn_fcoshead-gn-head_4xb4-1x_coco.py
@@ -0,0 +1,75 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+# model settings
+model = dict(
+ type='NASFCOS',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False, eps=0),
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')),
+ neck=dict(
+ type='NASFCOS_FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs=True,
+ num_outs=5,
+ norm_cfg=dict(type='BN'),
+ conv_cfg=dict(type='DCNv2', deform_groups=2)),
+ bbox_head=dict(
+ type='FCOSHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ strides=[8, 16, 32, 64, 128],
+ norm_cfg=dict(type='GN', num_groups=32),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='IoULoss', loss_weight=1.0),
+ loss_centerness=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0)),
+ train_cfg=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.4,
+ min_pos_iou=0,
+ ignore_iof_thr=-1),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+
+# dataset settings
+train_dataloader = dict(batch_size=4, num_workers=2)
+
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(lr=0.01),
+ paramwise_cfg=dict(bias_lr_mult=2., bias_decay_mult=0.))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fcos/nas-fcos_r50-caffe_fpn_nashead-gn-head_4xb4-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fcos/nas-fcos_r50-caffe_fpn_nashead-gn-head_4xb4-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..329f34c45ca0ea3f95e8da8505717df86b7c79c0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fcos/nas-fcos_r50-caffe_fpn_nashead-gn-head_4xb4-1x_coco.py
@@ -0,0 +1,74 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+# model settings
+model = dict(
+ type='NASFCOS',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False, eps=0),
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')),
+ neck=dict(
+ type='NASFCOS_FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs=True,
+ num_outs=5,
+ norm_cfg=dict(type='BN'),
+ conv_cfg=dict(type='DCNv2', deform_groups=2)),
+ bbox_head=dict(
+ type='NASFCOSHead',
+ num_classes=80,
+ in_channels=256,
+ feat_channels=256,
+ strides=[8, 16, 32, 64, 128],
+ norm_cfg=dict(type='GN', num_groups=32),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='IoULoss', loss_weight=1.0),
+ loss_centerness=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0)),
+ train_cfg=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.4,
+ min_pos_iou=0,
+ ignore_iof_thr=-1),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+
+# dataset settings
+train_dataloader = dict(batch_size=4, num_workers=2)
+
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(lr=0.01),
+ paramwise_cfg=dict(bias_lr_mult=2., bias_decay_mult=0.))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fpn/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fpn/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..260ec470fda46ae8d41dd768c5924da59803eb94
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fpn/README.md
@@ -0,0 +1,36 @@
+# NAS-FPN
+
+> [NAS-FPN: Learning Scalable Feature Pyramid Architecture for Object Detection](https://arxiv.org/abs/1904.07392)
+
+
+
+## Abstract
+
+Current state-of-the-art convolutional architectures for object detection are manually designed. Here we aim to learn a better architecture of feature pyramid network for object detection. We adopt Neural Architecture Search and discover a new feature pyramid architecture in a novel scalable search space covering all cross-scale connections. The discovered architecture, named NAS-FPN, consists of a combination of top-down and bottom-up connections to fuse features across scales. NAS-FPN, combined with various backbone models in the RetinaNet framework, achieves better accuracy and latency tradeoff compared to state-of-the-art object detection models. NAS-FPN improves mobile detection accuracy by 2 AP compared to state-of-the-art SSDLite with MobileNetV2 model in \[32\] and achieves 48.3 AP which surpasses Mask R-CNN \[10\] detection accuracy with less computation time.
+
+
+

+
+
+## Results and Models
+
+We benchmark the new training schedule (crop training, large batch, unfrozen BN, 50 epochs) introduced in NAS-FPN. RetinaNet is used in the paper.
+
+| Backbone | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :---------: | :-----: | :------: | :------------: | :----: | :--------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | 50e | 12.9 | 22.9 | 37.9 | [config](./retinanet_r50_fpn_crop640-50e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/nas_fpn/retinanet_r50_fpn_crop640_50e_coco/retinanet_r50_fpn_crop640_50e_coco-9b953d76.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/nas_fpn/retinanet_r50_fpn_crop640_50e_coco/retinanet_r50_fpn_crop640_50e_coco_20200529_095329.log.json) |
+| R-50-NASFPN | 50e | 13.2 | 23.0 | 40.5 | [config](./retinanet_r50_nasfpn_crop640-50e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/nas_fpn/retinanet_r50_nasfpn_crop640_50e_coco/retinanet_r50_nasfpn_crop640_50e_coco-0ad1f644.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/nas_fpn/retinanet_r50_nasfpn_crop640_50e_coco/retinanet_r50_nasfpn_crop640_50e_coco_20200528_230008.log.json) |
+
+**Note**: We find that it is unstable to train NAS-FPN and there is a small chance that results can be 3% mAP lower.
+
+## Citation
+
+```latex
+@inproceedings{ghiasi2019fpn,
+ title={Nas-fpn: Learning scalable feature pyramid architecture for object detection},
+ author={Ghiasi, Golnaz and Lin, Tsung-Yi and Le, Quoc V},
+ booktitle={Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition},
+ pages={7036--7045},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fpn/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fpn/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..aef0df6d7f38c71d691526004c0f1d19d66744b0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fpn/metafile.yml
@@ -0,0 +1,59 @@
+Collections:
+ - Name: NAS-FPN
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - NAS-FPN
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/1904.07392
+ Title: 'NAS-FPN: Learning Scalable Feature Pyramid Architecture for Object Detection'
+ README: configs/nas_fpn/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/necks/nas_fpn.py#L67
+ Version: v2.0.0
+
+Models:
+ - Name: retinanet_r50_fpn_crop640-50e_coco
+ In Collection: NAS-FPN
+ Config: configs/nas_fpn/retinanet_r50_fpn_crop640-50e_coco.py
+ Metadata:
+ Training Memory (GB): 12.9
+ inference time (ms/im):
+ - value: 43.67
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 50
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/nas_fpn/retinanet_r50_fpn_crop640_50e_coco/retinanet_r50_fpn_crop640_50e_coco-9b953d76.pth
+
+ - Name: retinanet_r50_nasfpn_crop640-50e_coco
+ In Collection: NAS-FPN
+ Config: configs/nas_fpn/retinanet_r50_nasfpn_crop640-50e_coco.py
+ Metadata:
+ Training Memory (GB): 13.2
+ inference time (ms/im):
+ - value: 43.48
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 50
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/nas_fpn/retinanet_r50_nasfpn_crop640_50e_coco/retinanet_r50_nasfpn_crop640_50e_coco-0ad1f644.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fpn/retinanet_r50_fpn_crop640-50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fpn/retinanet_r50_fpn_crop640-50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..11c34f6758a4862571e3f840424341c3964115be
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fpn/retinanet_r50_fpn_crop640-50e_coco.py
@@ -0,0 +1,78 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+norm_cfg = dict(type='BN', requires_grad=True)
+model = dict(
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=64,
+ batch_augments=[dict(type='BatchFixedSizePad', size=(640, 640))]),
+ backbone=dict(norm_eval=False),
+ neck=dict(
+ relu_before_extra_convs=True,
+ no_norm_on_lateral=True,
+ norm_cfg=norm_cfg),
+ bbox_head=dict(type='RetinaSepBNHead', num_ins=5, norm_cfg=norm_cfg),
+ # training and testing settings
+ train_cfg=dict(assigner=dict(neg_iou_thr=0.5)))
+
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize',
+ scale=(640, 640),
+ ratio_range=(0.8, 1.2),
+ keep_ratio=True),
+ dict(type='RandomCrop', crop_size=(640, 640)),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(640, 640), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=8, num_workers=4, dataset=dict(pipeline=train_pipeline))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# training schedule for 50e
+max_epochs = 50
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(type='LinearLR', start_factor=0.1, by_epoch=False, begin=0, end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[30, 40],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.08, momentum=0.9, weight_decay=0.0001),
+ paramwise_cfg=dict(norm_decay_mult=0, bypass_duplicate=True))
+
+env_cfg = dict(cudnn_benchmark=True)
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fpn/retinanet_r50_nasfpn_crop640-50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fpn/retinanet_r50_nasfpn_crop640-50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a851b745defb72aa05df289a3002c1534655d118
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/nas_fpn/retinanet_r50_nasfpn_crop640-50e_coco.py
@@ -0,0 +1,16 @@
+_base_ = './retinanet_r50_fpn_crop640-50e_coco.py'
+
+# model settings
+model = dict(
+ # `pad_size_divisor=128` ensures the feature maps sizes
+ # in `NAS_FPN` won't mismatch.
+ data_preprocessor=dict(pad_size_divisor=128),
+ neck=dict(
+ _delete_=True,
+ type='NASFPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5,
+ stack_times=7,
+ start_level=1,
+ norm_cfg=dict(type='BN', requires_grad=True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..fca0dbfc94505b437598a02ba9e2c6cf10778834
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/README.md
@@ -0,0 +1,102 @@
+# Objects365 Dataset
+
+> [Objects365 Dataset](https://openaccess.thecvf.com/content_ICCV_2019/papers/Shao_Objects365_A_Large-Scale_High-Quality_Dataset_for_Object_Detection_ICCV_2019_paper.pdf)
+
+
+
+## Abstract
+
+
+
+#### Objects365 Dataset V1
+
+[Objects365 Dataset V1](http://www.objects365.org/overview.html) is a brand new dataset,
+designed to spur object detection research with a focus on diverse objects in the Wild.
+It has 365 object categories over 600K training images. More than 10 million, high-quality bounding boxes are manually labeled through a three-step, carefully designed annotation pipeline. It is the largest object detection dataset (with full annotation) so far and establishes a more challenging benchmark for the community. Objects365 can serve as a better feature learning dataset for localization-sensitive tasks like object detection
+and semantic segmentation.
+
+
+
+
+

+
+
+#### Objects365 Dataset V2
+
+[Objects365 Dataset V2](http://www.objects365.org/overview.html) is based on the V1 release of the Objects365 dataset.
+Objects 365 annotated 365 object classes on more than 1800k images, with more than 29 million bounding boxes in the training set, surpassing PASCAL VOC, ImageNet, and COCO datasets.
+Objects 365 includes 11 categories of people, clothing, living room, bathroom, kitchen, office/medical, electrical appliances, transportation, food, animals, sports/musical instruments, and each category has dozens of subcategories.
+
+## Citation
+
+```
+@inproceedings{shao2019objects365,
+ title={Objects365: A large-scale, high-quality dataset for object detection},
+ author={Shao, Shuai and Li, Zeming and Zhang, Tianyuan and Peng, Chao and Yu, Gang and Zhang, Xiangyu and Li, Jing and Sun, Jian},
+ booktitle={Proceedings of the IEEE/CVF international conference on computer vision},
+ pages={8430--8439},
+ year={2019}
+}
+```
+
+## Prepare Dataset
+
+1. You need to download and extract Objects365 dataset. Users can download Objects365 V2 by using `tools/misc/download_dataset.py`.
+
+ **Usage**
+
+ ```shell
+ python tools/misc/download_dataset.py --dataset-name objects365v2 \
+ --save-dir ${SAVING PATH} \
+ --unzip \
+ --delete # Optional, delete the download zip file
+ ```
+
+ **Note:** There is no download link for Objects365 V1 right now. If you would like to download Objects365-V1, please visit [official website](http://www.objects365.org/) to concat the author.
+
+2. The directory should be like this:
+
+ ```none
+ mmdetection
+ ├── mmdet
+ ├── tools
+ ├── configs
+ ├── data
+ │ ├── Objects365
+ │ │ ├── Obj365_v1
+ │ │ │ ├── annotations
+ │ │ │ │ ├── objects365_train.json
+ │ │ │ │ ├── objects365_val.json
+ │ │ │ ├── train # training images
+ │ │ │ ├── val # validation images
+ │ │ ├── Obj365_v2
+ │ │ │ ├── annotations
+ │ │ │ │ ├── zhiyuan_objv2_train.json
+ │ │ │ │ ├── zhiyuan_objv2_val.json
+ │ │ │ ├── train # training images
+ │ │ │ │ ├── patch0
+ │ │ │ │ ├── patch1
+ │ │ │ │ ├── ...
+ │ │ │ ├── val # validation images
+ │ │ │ │ ├── patch0
+ │ │ │ │ ├── patch1
+ │ │ │ │ ├── ...
+ ```
+
+## Results and Models
+
+### Objects365 V1
+
+| Architecture | Backbone | Style | Lr schd | Mem (GB) | box AP | Config | Download |
+| :----------: | :------: | :-----: | :-----: | :------: | :----: | :-------------------------------------------------------------------------------------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Faster R-CNN | R-50 | pytorch | 1x | - | 19.6 | [config](https://github.com/open-mmlab/mmdetection/tree/main/configs/objects365/faster-rcnn_r50_fpn_16xb4-1x_objects365v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/objects365/faster_rcnn_r50_fpn_16x4_1x_obj365v1/faster_rcnn_r50_fpn_16x4_1x_obj365v1_20221219_181226-9ff10f95.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/objects365/faster_rcnn_r50_fpn_16x4_1x_obj365v1/faster_rcnn_r50_fpn_16x4_1x_obj365v1_20221219_181226.log.json) |
+| Faster R-CNN | R-50 | pytorch | 1350K | - | 22.3 | [config](https://github.com/open-mmlab/mmdetection/tree/main/configs/objects365/faster-rcnn_r50-syncbn_fpn_1350k_objects365v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/objects365/faster_rcnn_r50_fpn_syncbn_1350k_obj365v1/faster_rcnn_r50_fpn_syncbn_1350k_obj365v1_20220510_142457-337d8965.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/objects365/faster_rcnn_r50_fpn_syncbn_1350k_obj365v1/faster_rcnn_r50_fpn_syncbn_1350k_obj365v1_20220510_142457.log.json) |
+| Retinanet | R-50 | pytorch | 1x | - | 14.8 | [config](https://github.com/open-mmlab/mmdetection/tree/main/configs/objects365/retinanet_r50_fpn_1x_objects365v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/objects365/retinanet_r50_fpn_1x_obj365v1/retinanet_r50_fpn_1x_obj365v1_20221219_181859-ba3e3dd5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/objects365/retinanet_r50_fpn_1x_obj365v1/retinanet_r50_fpn_1x_obj365v1_20221219_181859.log.json) |
+| Retinanet | R-50 | pytorch | 1350K | - | 18.0 | [config](https://github.com/open-mmlab/mmdetection/tree/main/configs/objects365/retinanet_r50-syncbn_fpn_1350k_objects365v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/objects365/retinanet_r50_fpn_syncbn_1350k_obj365v1/retinanet_r50_fpn_syncbn_1350k_obj365v1_20220513_111237-7517c576.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/objects365/retinanet_r50_fpn_syncbn_1350k_obj365v1/retinanet_r50_fpn_syncbn_1350k_obj365v1_20220513_111237.log.json) |
+
+### Objects365 V2
+
+| Architecture | Backbone | Style | Lr schd | Mem (GB) | box AP | Config | Download |
+| :----------: | :------: | :-----: | :-----: | :------: | :----: | :---------------------------------------------------------------------------------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Faster R-CNN | R-50 | pytorch | 1x | - | 19.8 | [config](https://github.com/open-mmlab/mmdetection/tree/main/configs/objects365/faster-rcnn_r50_fpn_16xb4-1x_objects365v2.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/objects365/faster_rcnn_r50_fpn_16x4_1x_obj365v2/faster_rcnn_r50_fpn_16x4_1x_obj365v2_20221220_175040-5910b015.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/objects365/faster_rcnn_r50_fpn_16x4_1x_obj365v2/faster_rcnn_r50_fpn_16x4_1x_obj365v2_20221220_175040.log.json) |
+| Retinanet | R-50 | pytorch | 1x | - | 16.7 | [config](https://github.com/open-mmlab/mmdetection/tree/main/configs/objects365/retinanet_r50_fpn_1x_objects365v2.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/objects365/retinanet_r50_fpn_1x_obj365v2/retinanet_r50_fpn_1x_obj365v2_20221223_122105-d9b191f1.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/objects365/retinanet_r50_fpn_1x_obj365v2/retinanet_r50_fpn_1x_obj365v2_20221223_122105.log.json) |
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/faster-rcnn_r50-syncbn_fpn_1350k_objects365v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/faster-rcnn_r50-syncbn_fpn_1350k_objects365v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..ff7d0a360b95b1a72f779a8f7ad22a7e03235720
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/faster-rcnn_r50-syncbn_fpn_1350k_objects365v1.py
@@ -0,0 +1,49 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/objects365v2_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ backbone=dict(norm_cfg=dict(type='SyncBN', requires_grad=True)),
+ roi_head=dict(bbox_head=dict(num_classes=365)))
+
+# training schedule for 1350K
+train_cfg = dict(
+ _delete_=True,
+ type='IterBasedTrainLoop',
+ max_iters=1350000, # 36 epochs
+ val_interval=150000)
+
+# Using 8 GPUS while training
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.0001),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+# learning rate policy
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 1000,
+ by_epoch=False,
+ begin=0,
+ end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=1350000,
+ by_epoch=False,
+ milestones=[900000, 1200000],
+ gamma=0.1)
+]
+
+train_dataloader = dict(sampler=dict(type='InfiniteSampler'))
+default_hooks = dict(checkpoint=dict(by_epoch=False, interval=150000))
+
+log_processor = dict(by_epoch=False)
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/faster-rcnn_r50_fpn_16xb4-1x_objects365v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/faster-rcnn_r50_fpn_16xb4-1x_objects365v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..bc0d96fa22920a34f9ab9437a0f15cc93f46d0fa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/faster-rcnn_r50_fpn_16xb4-1x_objects365v1.py
@@ -0,0 +1,39 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/objects365v1_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(roi_head=dict(bbox_head=dict(num_classes=365)))
+
+train_dataloader = dict(
+ batch_size=4, # using 16 GPUS while training. total batch size is 16 x 4)
+)
+
+# Using 32 GPUS while training
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.08, momentum=0.9, weight_decay=0.0001),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 1000,
+ by_epoch=False,
+ begin=0,
+ end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (32 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/faster-rcnn_r50_fpn_16xb4-1x_objects365v2.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/faster-rcnn_r50_fpn_16xb4-1x_objects365v2.py
new file mode 100644
index 0000000000000000000000000000000000000000..1090678f652444c82a627fbf8bdda39fe0077f1e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/faster-rcnn_r50_fpn_16xb4-1x_objects365v2.py
@@ -0,0 +1,39 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/objects365v2_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(roi_head=dict(bbox_head=dict(num_classes=365)))
+
+train_dataloader = dict(
+ batch_size=4, # using 16 GPUS while training. total batch size is 16 x 4)
+)
+
+# Using 32 GPUS while training
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.08, momentum=0.9, weight_decay=0.0001),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 1000,
+ by_epoch=False,
+ begin=0,
+ end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (32 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..d43e8bde9d2aad9516f5383cd4152faf8f097660
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/metafile.yml
@@ -0,0 +1,101 @@
+- Name: retinanet_r50_fpn_1x_objects365v1
+ In Collection: RetinaNet
+ Config: configs/objects365/retinanet_r50_fpn_1x_objects365v1.py
+ Metadata:
+ Training Memory (GB): 7.4
+ Epochs: 12
+ Training Data: Objects365 v1
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Results:
+ - Task: Object Detection
+ Dataset: Objects365 v1
+ Metrics:
+ box AP: 14.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/objects365/retinanet_r50_fpn_1x_obj365v1/retinanet_r50_fpn_1x_obj365v1_20221219_181859-ba3e3dd5.pth
+
+- Name: retinanet_r50-syncbn_fpn_1350k_objects365v1
+ In Collection: RetinaNet
+ Config: configs/objects365/retinanet_r50-syncbn_fpn_1350k_objects365v1.py
+ Metadata:
+ Training Memory (GB): 7.6
+ Iterations: 1350000
+ Training Data: Objects365 v1
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Results:
+ - Task: Object Detection
+ Dataset: Objects365 v1
+ Metrics:
+ box AP: 18.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/objects365/retinanet_r50_fpn_syncbn_1350k_obj365v1/retinanet_r50_fpn_syncbn_1350k_obj365v1_20220513_111237-7517c576.pth
+
+- Name: retinanet_r50_fpn_1x_objects365v2
+ In Collection: RetinaNet
+ Config: configs/objects365/retinanet_r50_fpn_1x_objects365v2.py
+ Metadata:
+ Training Memory (GB): 7.2
+ Epochs: 12
+ Training Data: Objects365 v2
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Results:
+ - Task: Object Detection
+ Dataset: Objects365 v2
+ Metrics:
+ box AP: 16.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/objects365/retinanet_r50_fpn_1x_obj365v2/retinanet_r50_fpn_1x_obj365v2_20221223_122105-d9b191f1.pth
+
+- Name: faster-rcnn_r50_fpn_16xb4-1x_objects365v1
+ In Collection: Faster R-CNN
+ Config: configs/objects365/faster-rcnn_r50_fpn_16xb4-1x_objects365v1.py
+ Metadata:
+ Training Memory (GB): 11.4
+ Epochs: 12
+ Training Data: Objects365 v1
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Results:
+ - Task: Object Detection
+ Dataset: Objects365 v1
+ Metrics:
+ box AP: 19.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/objects365/faster_rcnn_r50_fpn_16x4_1x_obj365v1/faster_rcnn_r50_fpn_16x4_1x_obj365v1_20221219_181226-9ff10f95.pth
+
+- Name: faster-rcnn_r50-syncbn_fpn_1350k_objects365v1
+ In Collection: Faster R-CNN
+ Config: configs/objects365/faster-rcnn_r50-syncbn_fpn_1350k_objects365v1.py
+ Metadata:
+ Training Memory (GB): 8.6
+ Iterations: 1350000
+ Training Data: Objects365 v1
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Results:
+ - Task: Object Detection
+ Dataset: Objects365 v1
+ Metrics:
+ box AP: 22.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/objects365/faster_rcnn_r50_fpn_syncbn_1350k_obj365v1/faster_rcnn_r50_fpn_syncbn_1350k_obj365v1_20220510_142457-337d8965.pth
+
+- Name: faster-rcnn_r50_fpn_16xb4-1x_objects365v2
+ In Collection: Faster R-CNN
+ Config: configs/objects365/faster-rcnn_r50_fpn_16xb4-1x_objects365v2.py
+ Metadata:
+ Training Memory (GB): 10.8
+ Epochs: 12
+ Training Data: Objects365 v1
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Results:
+ - Task: Object Detection
+ Dataset: Objects365 v2
+ Metrics:
+ box AP: 19.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/objects365/faster_rcnn_r50_fpn_16x4_1x_obj365v2/faster_rcnn_r50_fpn_16x4_1x_obj365v2_20221220_175040-5910b015.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/retinanet_r50-syncbn_fpn_1350k_objects365v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/retinanet_r50-syncbn_fpn_1350k_objects365v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..c41dfce8bc67e7f4d18434a2c10a33c66da403c1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/retinanet_r50-syncbn_fpn_1350k_objects365v1.py
@@ -0,0 +1,49 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/objects365v2_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ backbone=dict(norm_cfg=dict(type='SyncBN', requires_grad=True)),
+ bbox_head=dict(num_classes=365))
+
+# training schedule for 1350K
+train_cfg = dict(
+ _delete_=True,
+ type='IterBasedTrainLoop',
+ max_iters=1350000, # 36 epochs
+ val_interval=150000)
+
+# Using 8 GPUS while training
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+# learning rate policy
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 1000,
+ by_epoch=False,
+ begin=0,
+ end=10000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=1350000,
+ by_epoch=False,
+ milestones=[900000, 1200000],
+ gamma=0.1)
+]
+
+train_dataloader = dict(sampler=dict(type='InfiniteSampler'))
+default_hooks = dict(checkpoint=dict(by_epoch=False, interval=150000))
+
+log_processor = dict(by_epoch=False)
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/retinanet_r50_fpn_1x_objects365v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/retinanet_r50_fpn_1x_objects365v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..72144192aaa36d757053a982ed7ad2a886916b75
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/retinanet_r50_fpn_1x_objects365v1.py
@@ -0,0 +1,35 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/objects365v1_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(bbox_head=dict(num_classes=365))
+
+# Using 8 GPUS while training
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 1000,
+ by_epoch=False,
+ begin=0,
+ end=10000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/retinanet_r50_fpn_1x_objects365v2.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/retinanet_r50_fpn_1x_objects365v2.py
new file mode 100644
index 0000000000000000000000000000000000000000..219544126ab0ab6e93d50f1962ffaf40f25b14f0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/objects365/retinanet_r50_fpn_1x_objects365v2.py
@@ -0,0 +1,35 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/objects365v2_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(bbox_head=dict(num_classes=365))
+
+# Using 8 GPUS while training
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 1000,
+ by_epoch=False,
+ begin=0,
+ end=10000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ocsort/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ocsort/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..e9b86c6c6c1ca167b875c5d3241af28dc4919358
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ocsort/README.md
@@ -0,0 +1,56 @@
+# Observation-Centric SORT: Rethinking SORT for Robust Multi-Object Tracking
+
+## Abstract
+
+
+
+Multi-Object Tracking (MOT) has rapidly progressed with the development of object detection and re-identification. However, motion modeling, which facilitates object association by forecasting short-term trajec- tories with past observations, has been relatively under-explored in recent years. Current motion models in MOT typically assume that the object motion is linear in a small time window and needs continuous observations, so these methods are sensitive to occlusions and non-linear motion and require high frame-rate videos. In this work, we show that a simple motion model can obtain state-of-the-art tracking performance without other cues like appearance. We emphasize the role of “observation” when recovering tracks from being lost and reducing the error accumulated by linear motion models during the lost period. We thus name the proposed method as Observation-Centric SORT, OC-SORT for short. It remains simple, online, and real-time but improves robustness over occlusion and non-linear motion. It achieves 63.2 and 62.1 HOTA on MOT17 and MOT20, respectively, surpassing all published methods. It also sets new states of the art on KITTI Pedestrian Tracking and DanceTrack where the object motion is highly non-linear
+
+
+
+
+

+
+
+## Citation
+
+
+
+```latex
+@article{cao2022observation,
+ title={Observation-Centric SORT: Rethinking SORT for Robust Multi-Object Tracking},
+ author={Cao, Jinkun and Weng, Xinshuo and Khirodkar, Rawal and Pang, Jiangmiao and Kitani, Kris},
+ journal={arXiv preprint arXiv:2203.14360},
+ year={2022}
+}
+```
+
+## Results and models on MOT17
+
+The performance on `MOT17-half-val` is comparable with the performance from [the OC-SORT official implementation](https://github.com/noahcao/OC_SORT). We use the same YOLO-X detector weights as in [ByteTrack](https://github.com/open-mmlab/mmtracking/tree/master/configs/mot/bytetrack).
+
+| Method | Detector | Train Set | Test Set | Public | Inf time (fps) | HOTA | MOTA | IDF1 | FP | FN | IDSw. | Config | Download |
+| :-----: | :------: | :---------------------: | :------: | :----: | :------------: | :--: | :--: | :--: | :---: | :---: | :---: | :-------------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| OC-SORT | YOLOX-X | CrowdHuman + half-train | half-val | N | - | 67.5 | 77.5 | 78.2 | 15987 | 19590 | 855 | [config](ocsort_yolox_x_crowdhuman_mot17-private-half.py) | [model](https://download.openmmlab.com/mmtracking/mot/ocsort/mot_dataset/ocsort_yolox_x_crowdhuman_mot17-private-half_20220813_101618-fe150582.pth) \| [log](https://download.openmmlab.com/mmtracking/mot/ocsort/mot_dataset/ocsort_yolox_x_crowdhuman_mot17-private-half_20220813_101618.log.json) |
+
+## Get started
+
+### 1. Development Environment Setup
+
+Tracking Development Environment Setup can refer to this [document](../../docs/en/get_started.md).
+
+### 2. Dataset Prepare
+
+Tracking Dataset Prepare can refer to this [document](../../docs/en/user_guides/tracking_dataset_prepare.md).
+
+### 3. Training
+
+OCSORT training is same as Bytetrack, please refer to [document](../../configs/bytetrack/README.md).
+
+### 4. Testing and evaluation
+
+OCSORT evaluation and test are same as Bytetrack, please refer to [document](../../configs/bytetrack/README.md).
+
+### 5.Inference
+
+OCSORT inference is same as Bytetrack, please refer to [document](../../configs/bytetrack/README.md).
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ocsort/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ocsort/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..0a31ef108ea7c594d3566970763ff704234d4e0c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ocsort/metafile.yml
@@ -0,0 +1,27 @@
+Collections:
+ - Name: OCSORT
+ Metadata:
+ Training Techniques:
+ - SGD with Momentum
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - YOLOX
+ Paper:
+ URL: https://arxiv.org/abs/2203.14360
+ Title: Observation-Centric SORT Rethinking SORT for Robust Multi-Object Tracking
+ README: configs/ocsort/README.md
+
+Models:
+ - Name: ocsort_yolox_x_crowdhuman_mot17-private-half
+ In Collection: OCSORT
+ Config: configs/ocsort/ocsort_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17halfval.py
+ Metadata:
+ Training Data: CrowdHuman + MOT17-half-train
+ Results:
+ - Task: Multiple Object Tracking
+ Dataset: MOT17-half-val
+ Metrics:
+ HOTA: 67.5
+ MOTA: 77.5
+ IDF1: 78.2
+ Weights: https://download.openmmlab.com/mmtracking/mot/ocsort/mot_dataset/ocsort_yolox_x_crowdhuman_mot17-private-half_20220813_101618-fe150582.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ocsort/ocsort_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17halfval.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ocsort/ocsort_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17halfval.py
new file mode 100644
index 0000000000000000000000000000000000000000..ea04923d6aec237c51b7e23d0348c487cb9d697b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ocsort/ocsort_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17halfval.py
@@ -0,0 +1,18 @@
+_base_ = [
+ '../bytetrack/bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17halfval.py', # noqa: E501
+]
+
+model = dict(
+ type='OCSORT',
+ tracker=dict(
+ _delete_=True,
+ type='OCSORTTracker',
+ motion=dict(type='KalmanFilter'),
+ obj_score_thr=0.3,
+ init_track_thr=0.7,
+ weight_iou_with_det_scores=True,
+ match_iou_thr=0.3,
+ num_tentatives=3,
+ vel_consist_weight=0.2,
+ vel_delta_t=3,
+ num_frames_retain=30))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ocsort/ocsort_yolox_x_8xb4-amp-80e_crowdhuman-mot20train_test-mot20test.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ocsort/ocsort_yolox_x_8xb4-amp-80e_crowdhuman-mot20train_test-mot20test.py
new file mode 100644
index 0000000000000000000000000000000000000000..ea04923d6aec237c51b7e23d0348c487cb9d697b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ocsort/ocsort_yolox_x_8xb4-amp-80e_crowdhuman-mot20train_test-mot20test.py
@@ -0,0 +1,18 @@
+_base_ = [
+ '../bytetrack/bytetrack_yolox_x_8xb4-amp-80e_crowdhuman-mot17halftrain_test-mot17halfval.py', # noqa: E501
+]
+
+model = dict(
+ type='OCSORT',
+ tracker=dict(
+ _delete_=True,
+ type='OCSORTTracker',
+ motion=dict(type='KalmanFilter'),
+ obj_score_thr=0.3,
+ init_track_thr=0.7,
+ weight_iou_with_det_scores=True,
+ match_iou_thr=0.3,
+ num_tentatives=3,
+ vel_consist_weight=0.2,
+ vel_delta_t=3,
+ num_frames_retain=30))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..ccfc721da568833222038000ac1a5ea12e9bb732
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/README.md
@@ -0,0 +1,149 @@
+# Open Images Dataset
+
+> [Open Images Dataset](https://arxiv.org/abs/1811.00982)
+
+
+
+## Abstract
+
+
+
+#### Open Images v6
+
+[Open Images](https://storage.googleapis.com/openimages/web/index.html) is a dataset of ~9M images annotated with image-level labels,
+object bounding boxes, object segmentation masks, visual relationships,
+and localized narratives:
+
+- It contains a total of 16M bounding boxes for 600 object classes on
+ 1.9M images, making it the largest existing dataset with object location
+ annotations. The boxes have been largely manually drawn by professional
+ annotators to ensure accuracy and consistency. The images are very diverse
+ and often contain complex scenes with several objects (8.3 per image on
+ average).
+
+- Open Images also offers visual relationship annotations, indicating pairs
+ of objects in particular relations (e.g. "woman playing guitar", "beer on
+ table"), object properties (e.g. "table is wooden"), and human actions (e.g.
+ "woman is jumping"). In total it has 3.3M annotations from 1,466 distinct
+ relationship triplets.
+
+- In V5 we added segmentation masks for 2.8M object instances in 350 classes.
+ Segmentation masks mark the outline of objects, which characterizes their
+ spatial extent to a much higher level of detail.
+
+- In V6 we added 675k localized narratives: multimodal descriptions of images
+ consisting of synchronized voice, text, and mouse traces over the objects being
+ described. (Note we originally launched localized narratives only on train in V6,
+ but since July 2020 we also have validation and test covered.)
+
+- Finally, the dataset is annotated with 59.9M image-level labels spanning 19,957
+ classes.
+
+We believe that having a single dataset with unified annotations for image
+classification, object detection, visual relationship detection, instance
+segmentation, and multimodal image descriptions will enable to study these
+tasks jointly and stimulate progress towards genuine scene understanding.
+
+
+
+
+

+
+
+#### Open Images Challenge 2019
+
+[Open Images Challenges 2019](https://storage.googleapis.com/openimages/web/challenge2019.html) is based on the V5 release of the Open
+Images dataset. The images of the dataset are very varied and
+often contain complex scenes with several objects (explore the dataset).
+
+## Citation
+
+```
+@article{OpenImages,
+ author = {Alina Kuznetsova and Hassan Rom and Neil Alldrin and Jasper Uijlings and Ivan Krasin and Jordi Pont-Tuset and Shahab Kamali and Stefan Popov and Matteo Malloci and Alexander Kolesnikov and Tom Duerig and Vittorio Ferrari},
+ title = {The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale},
+ year = {2020},
+ journal = {IJCV}
+}
+```
+
+## Prepare Dataset
+
+1. You need to download and extract Open Images dataset.
+
+2. The Open Images dataset does not have image metas (width and height of the image),
+ which will be used during training and testing (evaluation). We suggest to get test image metas before
+ training/testing by using `tools/misc/get_image_metas.py`.
+
+ **Usage**
+
+ ```shell
+ python tools/misc/get_image_metas.py ${CONFIG} \
+ --dataset ${DATASET TYPE} \ # train or val or test
+ --out ${OUTPUT FILE NAME}
+ ```
+
+3. The directory should be like this:
+
+ ```none
+ mmdetection
+ ├── mmdet
+ ├── tools
+ ├── configs
+ ├── data
+ │ ├── OpenImages
+ │ │ ├── annotations
+ │ │ │ ├── bbox_labels_600_hierarchy.json
+ │ │ │ ├── class-descriptions-boxable.csv
+ │ │ │ ├── oidv6-train-annotations-bbox.scv
+ │ │ │ ├── validation-annotations-bbox.csv
+ │ │ │ ├── validation-annotations-human-imagelabels-boxable.csv
+ │ │ │ ├── validation-image-metas.pkl # get from script
+ │ │ ├── challenge2019
+ │ │ │ ├── challenge-2019-train-detection-bbox.txt
+ │ │ │ ├── challenge-2019-validation-detection-bbox.txt
+ │ │ │ ├── class_label_tree.np
+ │ │ │ ├── class_sample_train.pkl
+ │ │ │ ├── challenge-2019-validation-detection-human-imagelabels.csv # download from official website
+ │ │ │ ├── challenge-2019-validation-metas.pkl # get from script
+ │ │ ├── OpenImages
+ │ │ │ ├── train # training images
+ │ │ │ ├── test # testing images
+ │ │ │ ├── validation # validation images
+ ```
+
+**Note**:
+
+1. The training and validation images of Open Images Challenge dataset are based on
+ Open Images v6, but the test images are different.
+2. The Open Images Challenges annotations are obtained from [TSD](https://github.com/Sense-X/TSD).
+ You can also download the annotations from [official website](https://storage.googleapis.com/openimages/web/challenge2019_downloads.html),
+ and set data.train.type=OpenImagesDataset, data.val.type=OpenImagesDataset, and data.test.type=OpenImagesDataset in the config
+3. If users do not want to use `validation-annotations-human-imagelabels-boxable.csv` and `challenge-2019-validation-detection-human-imagelabels.csv`
+ users can set `test_dataloader.dataset.image_level_ann_file=None` and `test_dataloader.dataset.image_level_ann_file=None` in the config.
+ Please note that loading image-levels label is the default of Open Images evaluation metric.
+ More details please refer to the [official website](https://storage.googleapis.com/openimages/web/evaluation.html)
+
+## Results and Models
+
+| Architecture | Backbone | Style | Lr schd | Sampler | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :---------------------------: | :------: | :-----: | :-----: | :-----------------: | :------: | :------------: | :----: | :------------------------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Faster R-CNN | R-50 | pytorch | 1x | Group Sampler | 7.7 | - | 51.6 | [config](./faster-rcnn_r50_fpn_32xb2-1x_openimages.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/openimages/faster_rcnn_r50_fpn_32x2_1x_openimages/faster_rcnn_r50_fpn_32x2_1x_openimages_20211130_231159-e87ab7ce.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/openimages/faster_rcnn_r50_fpn_32x2_1x_openimages/faster_rcnn_r50_fpn_32x2_1x_openimages_20211130_231159.log.json) |
+| Faster R-CNN | R-50 | pytorch | 1x | Class Aware Sampler | 7.7 | - | 60.0 | [config](./faster-rcnn_r50_fpn_32xb2-cas-1x_openimages.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/openimages/faster_rcnn_r50_fpn_32x2_cas_1x_openimages/faster_rcnn_r50_fpn_32x2_cas_1x_openimages_20220306_202424-98c630e5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/openimages/faster_rcnn_r50_fpn_32x2_cas_1x_openimages/faster_rcnn_r50_fpn_32x2_cas_1x_openimages_20220306_202424.log.json) |
+| Faster R-CNN (Challenge 2019) | R-50 | pytorch | 1x | Group Sampler | 7.7 | - | 54.9 | [config](./faster-rcnn_r50_fpn_32xb2-1x_openimages-challenge.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/openimages/faster_rcnn_r50_fpn_32x2_1x_openimages_challenge/faster_rcnn_r50_fpn_32x2_1x_openimages_challenge_20220114_045100-0e79e5df.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/openimages/faster_rcnn_r50_fpn_32x2_1x_openimages_challenge/faster_rcnn_r50_fpn_32x2_1x_openimages_challenge_20220114_045100.log.json) |
+| Faster R-CNN (Challenge 2019) | R-50 | pytorch | 1x | Class Aware Sampler | 7.1 | - | 65.0 | [config](./faster-rcnn_r50_fpn_32xb2-cas-1x_openimages-challenge.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/openimages/faster_rcnn_r50_fpn_32x2_cas_1x_openimages_challenge/faster_rcnn_r50_fpn_32x2_cas_1x_openimages_challenge_20220221_192021-34c402d9.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/openimages/faster_rcnn_r50_fpn_32x2_cas_1x_openimages_challenge/faster_rcnn_r50_fpn_32x2_cas_1x_openimages_challenge_20220221_192021.log.json) |
+| Retinanet | R-50 | pytorch | 1x | Group Sampler | 6.6 | - | 61.5 | [config](./retinanet_r50_fpn_32xb2-1x_openimages.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/openimages/retinanet_r50_fpn_32x2_1x_openimages/retinanet_r50_fpn_32x2_1x_openimages_20211223_071954-d2ae5462.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/openimages/retinanet_r50_fpn_32x2_1x_openimages/retinanet_r50_fpn_32x2_1x_openimages_20211223_071954.log.json) |
+| SSD | VGG16 | pytorch | 36e | Group Sampler | 10.8 | - | 35.4 | [config](./ssd300_32xb8-36e_openimages.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/openimages/ssd300_32x8_36e_openimages/ssd300_32x8_36e_openimages_20211224_000232-dce93846.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/openimages/ssd300_32x8_36e_openimages/ssd300_32x8_36e_openimages_20211224_000232.log.json) |
+
+**Notes:**
+
+- 'cas' is short for 'Class Aware Sampler'
+
+### Results of consider image level labels
+
+| Architecture | Sampler | Consider Image Level Labels | box AP |
+| :-------------------------------: | :-----------------: | :-------------------------: | :----: |
+| Faster R-CNN r50 (Challenge 2019) | Group Sampler | w/o | 62.19 |
+| Faster R-CNN r50 (Challenge 2019) | Group Sampler | w/ | 54.87 |
+| Faster R-CNN r50 (Challenge 2019) | Class Aware Sampler | w/o | 71.77 |
+| Faster R-CNN r50 (Challenge 2019) | Class Aware Sampler | w/ | 64.98 |
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/faster-rcnn_r50_fpn_32xb2-1x_openimages-challenge.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/faster-rcnn_r50_fpn_32xb2-1x_openimages-challenge.py
new file mode 100644
index 0000000000000000000000000000000000000000..e79a92cccb2e432e5dd60bc080dab76781eb32bc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/faster-rcnn_r50_fpn_32xb2-1x_openimages-challenge.py
@@ -0,0 +1,39 @@
+_base_ = ['faster-rcnn_r50_fpn_32xb2-1x_openimages.py']
+
+model = dict(
+ roi_head=dict(bbox_head=dict(num_classes=500)),
+ test_cfg=dict(rcnn=dict(score_thr=0.01)))
+
+# dataset settings
+dataset_type = 'OpenImagesChallengeDataset'
+train_dataloader = dict(
+ dataset=dict(
+ type=dataset_type,
+ ann_file='challenge2019/challenge-2019-train-detection-bbox.txt',
+ label_file='challenge2019/cls-label-description.csv',
+ hierarchy_file='challenge2019/class_label_tree.np',
+ meta_file='challenge2019/challenge-2019-train-metas.pkl'))
+val_dataloader = dict(
+ dataset=dict(
+ type=dataset_type,
+ ann_file='challenge2019/challenge-2019-validation-detection-bbox.txt',
+ data_prefix=dict(img='OpenImages/'),
+ label_file='challenge2019/cls-label-description.csv',
+ hierarchy_file='challenge2019/class_label_tree.np',
+ meta_file='challenge2019/challenge-2019-validation-metas.pkl',
+ image_level_ann_file='challenge2019/challenge-2019-validation-'
+ 'detection-human-imagelabels.csv'))
+test_dataloader = dict(
+ dataset=dict(
+ type=dataset_type,
+ ann_file='challenge2019/challenge-2019-validation-detection-bbox.txt',
+ label_file='challenge2019/cls-label-description.csv',
+ hierarchy_file='challenge2019/class_label_tree.np',
+ meta_file='challenge2019/challenge-2019-validation-metas.pkl',
+ image_level_ann_file='challenge2019/challenge-2019-validation-'
+ 'detection-human-imagelabels.csv'))
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (32 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/faster-rcnn_r50_fpn_32xb2-1x_openimages.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/faster-rcnn_r50_fpn_32xb2-1x_openimages.py
new file mode 100644
index 0000000000000000000000000000000000000000..f3f0aa0a0ff0ef16cd6e55543a72b5fe405ec5a8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/faster-rcnn_r50_fpn_32xb2-1x_openimages.py
@@ -0,0 +1,35 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/openimages_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(roi_head=dict(bbox_head=dict(num_classes=601)))
+
+# Using 32 GPUS while training
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.08, momentum=0.9, weight_decay=0.0001),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 64,
+ by_epoch=False,
+ begin=0,
+ end=26000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (32 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/faster-rcnn_r50_fpn_32xb2-cas-1x_openimages-challenge.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/faster-rcnn_r50_fpn_32xb2-cas-1x_openimages-challenge.py
new file mode 100644
index 0000000000000000000000000000000000000000..9e428725bcc39d2c009a2382c191fa53fe5ce284
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/faster-rcnn_r50_fpn_32xb2-cas-1x_openimages-challenge.py
@@ -0,0 +1,5 @@
+_base_ = ['faster-rcnn_r50_fpn_32xb2-1x_openimages-challenge.py']
+
+# Use ClassAwareSampler
+train_dataloader = dict(
+ sampler=dict(_delete_=True, type='ClassAwareSampler', num_sample_class=1))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/faster-rcnn_r50_fpn_32xb2-cas-1x_openimages.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/faster-rcnn_r50_fpn_32xb2-cas-1x_openimages.py
new file mode 100644
index 0000000000000000000000000000000000000000..803190abfee63ea87e70dfe1b0fddca02f3556b8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/faster-rcnn_r50_fpn_32xb2-cas-1x_openimages.py
@@ -0,0 +1,5 @@
+_base_ = ['faster-rcnn_r50_fpn_32xb2-1x_openimages.py']
+
+# Use ClassAwareSampler
+train_dataloader = dict(
+ sampler=dict(_delete_=True, type='ClassAwareSampler', num_sample_class=1))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..76c1209471921610f791a074ed7a6863cd0709c0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/metafile.yml
@@ -0,0 +1,102 @@
+Models:
+ - Name: faster-rcnn_r50_fpn_32x2_1x_openimages
+ In Collection: Faster R-CNN
+ Config: configs/openimages/faster-rcnn_r50_fpn_32xb2-1x_openimages.py
+ Metadata:
+ Training Memory (GB): 7.7
+ Epochs: 12
+ Training Data: Open Images v6
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Results:
+ - Task: Object Detection
+ Dataset: Open Images v6
+ Metrics:
+ box AP: 51.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/openimages/faster_rcnn_r50_fpn_32x2_1x_openimages/faster_rcnn_r50_fpn_32x2_1x_openimages_20211130_231159-e87ab7ce.pth
+
+ - Name: retinanet_r50_fpn_32xb2-1x_openimages
+ In Collection: RetinaNet
+ Config: configs/openimages/retinanet_r50_fpn_32xb2-1x_openimages.py
+ Metadata:
+ Training Memory (GB): 6.6
+ Epochs: 12
+ Training Data: Open Images v6
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Results:
+ - Task: Object Detection
+ Dataset: Open Images v6
+ Metrics:
+ box AP: 61.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/openimages/retinanet_r50_fpn_32x2_1x_openimages/retinanet_r50_fpn_32x2_1x_openimages_20211223_071954-d2ae5462.pth
+
+ - Name: ssd300_32xb8-36e_openimages
+ In Collection: SSD
+ Config: configs/openimages/ssd300_32xb8-36e_openimages.py
+ Metadata:
+ Training Memory (GB): 10.8
+ Epochs: 36
+ Training Data: Open Images v6
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Results:
+ - Task: Object Detection
+ Dataset: Open Images v6
+ Metrics:
+ box AP: 35.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/openimages/ssd300_32x8_36e_openimages/ssd300_32x8_36e_openimages_20211224_000232-dce93846.pth
+
+ - Name: faster-rcnn_r50_fpn_32x2_1x_openimages_challenge
+ In Collection: Faster R-CNN
+ Config: configs/openimages/faster-rcnn_r50_fpn_32xb2-1x_openimages-challenge.py
+ Metadata:
+ Training Memory (GB): 7.7
+ Epochs: 12
+ Training Data: Open Images Challenge 2019
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Results:
+ - Task: Object Detection
+ Dataset: Open Images Challenge 2019
+ Metrics:
+ box AP: 54.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/openimages/faster_rcnn_r50_fpn_32x2_1x_openimages_challenge/faster_rcnn_r50_fpn_32x2_1x_openimages_challenge_20220114_045100-0e79e5df.pth
+
+ - Name: faster-rcnn_r50_fpn_32x2_cas_1x_openimages
+ In Collection: Faster R-CNN
+ Config: configs/openimages/faster-rcnn_r50_fpn_32xb2-cas-1x_openimages.py
+ Metadata:
+ Training Memory (GB): 7.7
+ Epochs: 12
+ Training Data: Open Images Challenge 2019
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Results:
+ - Task: Object Detection
+ Dataset: Open Images Challenge 2019
+ Metrics:
+ box AP: 60.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/openimages/faster_rcnn_r50_fpn_32x2_cas_1x_openimages/faster_rcnn_r50_fpn_32x2_cas_1x_openimages_20220306_202424-98c630e5.pth
+
+ - Name: faster-rcnn_r50_fpn_32x2_cas_1x_openimages_challenge
+ In Collection: Faster R-CNN
+ Config: configs/openimages/faster-rcnn_r50_fpn_32xb2-cas-1x_openimages-challenge.py
+ Metadata:
+ Training Memory (GB): 7.1
+ Epochs: 12
+ Training Data: Open Images Challenge 2019
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Results:
+ - Task: Object Detection
+ Dataset: Open Images Challenge 2019
+ Metrics:
+ box AP: 65.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/openimages/faster_rcnn_r50_fpn_32x2_cas_1x_openimages_challenge/faster_rcnn_r50_fpn_32x2_cas_1x_openimages_challenge_20220221_192021-34c402d9.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/retinanet_r50_fpn_32xb2-1x_openimages.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/retinanet_r50_fpn_32xb2-1x_openimages.py
new file mode 100644
index 0000000000000000000000000000000000000000..97a0eb075c730ceeaa494190e0b8369706c7d7c3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/retinanet_r50_fpn_32xb2-1x_openimages.py
@@ -0,0 +1,35 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/openimages_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(bbox_head=dict(num_classes=601))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 64,
+ by_epoch=False,
+ begin=0,
+ end=26000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.08, momentum=0.9, weight_decay=0.0001),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (32 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/ssd300_32xb8-36e_openimages.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/ssd300_32xb8-36e_openimages.py
new file mode 100644
index 0000000000000000000000000000000000000000..9cb51cae00a8707c0a901b99620851132e9eaccf
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/openimages/ssd300_32xb8-36e_openimages.py
@@ -0,0 +1,88 @@
+_base_ = [
+ '../_base_/models/ssd300.py', '../_base_/datasets/openimages_detection.py',
+ '../_base_/default_runtime.py', '../_base_/schedules/schedule_1x.py'
+]
+model = dict(
+ bbox_head=dict(
+ num_classes=601,
+ anchor_generator=dict(basesize_ratio_range=(0.2, 0.9))))
+# dataset settings
+dataset_type = 'OpenImagesDataset'
+data_root = 'data/OpenImages/'
+input_size = 300
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PhotoMetricDistortion',
+ brightness_delta=32,
+ contrast_range=(0.5, 1.5),
+ saturation_range=(0.5, 1.5),
+ hue_delta=18),
+ dict(
+ type='Expand',
+ mean={{_base_.model.data_preprocessor.mean}},
+ to_rgb={{_base_.model.data_preprocessor.bgr_to_rgb}},
+ ratio_range=(1, 4)),
+ dict(
+ type='MinIoURandomCrop',
+ min_ious=(0.1, 0.3, 0.5, 0.7, 0.9),
+ min_crop_size=0.3),
+ dict(type='Resize', scale=(input_size, input_size), keep_ratio=False),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(input_size, input_size), keep_ratio=False),
+ # avoid bboxes being resized
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'instances'))
+]
+
+train_dataloader = dict(
+ batch_size=8, # using 32 GPUS while training. total batch size is 32 x 8
+ batch_sampler=None,
+ dataset=dict(
+ _delete_=True,
+ type='RepeatDataset',
+ times=3, # repeat 3 times, total epochs are 12 x 3
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/oidv6-train-annotations-bbox.csv',
+ data_prefix=dict(img='OpenImages/train/'),
+ label_file='annotations/class-descriptions-boxable.csv',
+ hierarchy_file='annotations/bbox_labels_600_hierarchy.json',
+ meta_file='annotations/train-image-metas.pkl',
+ pipeline=train_pipeline)))
+val_dataloader = dict(batch_size=8, dataset=dict(pipeline=test_pipeline))
+test_dataloader = dict(batch_size=8, dataset=dict(pipeline=test_pipeline))
+
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.04, momentum=0.9, weight_decay=5e-4))
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=0.001,
+ by_epoch=False,
+ begin=0,
+ end=20000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (32 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=256)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..625aacf24516087cefe1082271f25baf536bc03d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/README.md
@@ -0,0 +1,47 @@
+# PAA
+
+> [Probabilistic Anchor Assignment with IoU Prediction for Object Detection](https://arxiv.org/abs/2007.08103)
+
+
+
+## Abstract
+
+In object detection, determining which anchors to assign as positive or negative samples, known as anchor assignment, has been revealed as a core procedure that can significantly affect a model's performance. In this paper we propose a novel anchor assignment strategy that adaptively separates anchors into positive and negative samples for a ground truth bounding box according to the model's learning status such that it is able to reason about the separation in a probabilistic manner. To do so we first calculate the scores of anchors conditioned on the model and fit a probability distribution to these scores. The model is then trained with anchors separated into positive and negative samples according to their probabilities. Moreover, we investigate the gap between the training and testing objectives and propose to predict the Intersection-over-Unions of detected boxes as a measure of localization quality to reduce the discrepancy. The combined score of classification and localization qualities serving as a box selection metric in non-maximum suppression well aligns with the proposed anchor assignment strategy and leads significant performance improvements. The proposed methods only add a single convolutional layer to RetinaNet baseline and does not require multiple anchors per location, so are efficient. Experimental results verify the effectiveness of the proposed methods. Especially, our models set new records for single-stage detectors on MS COCO test-dev dataset with various backbones.
+
+
+

+
+
+## Results and Models
+
+We provide config files to reproduce the object detection results in the
+ECCV 2020 paper for Probabilistic Anchor Assignment with IoU
+Prediction for Object Detection.
+
+| Backbone | Lr schd | Mem (GB) | Score voting | box AP | Config | Download |
+| :-------: | :-----: | :------: | :----------: | :----: | :------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | 12e | 3.7 | True | 40.4 | [config](./paa_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r50_fpn_1x_coco/paa_r50_fpn_1x_coco_20200821-936edec3.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r50_fpn_1x_coco/paa_r50_fpn_1x_coco_20200821-936edec3.log.json) |
+| R-50-FPN | 12e | 3.7 | False | 40.2 | - | |
+| R-50-FPN | 18e | 3.7 | True | 41.4 | [config](./paa_r50_fpn_1.5x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r50_fpn_1.5x_coco/paa_r50_fpn_1.5x_coco_20200823-805d6078.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r50_fpn_1.5x_coco/paa_r50_fpn_1.5x_coco_20200823-805d6078.log.json) |
+| R-50-FPN | 18e | 3.7 | False | 41.2 | - | |
+| R-50-FPN | 24e | 3.7 | True | 41.6 | [config](./paa_r50_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r50_fpn_2x_coco/paa_r50_fpn_2x_coco_20200821-c98bfc4e.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r50_fpn_2x_coco/paa_r50_fpn_2x_coco_20200821-c98bfc4e.log.json) |
+| R-50-FPN | 36e | 3.7 | True | 43.3 | [config](./paa_r50_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r50_fpn_mstrain_3x_coco/paa_r50_fpn_mstrain_3x_coco_20210121_145722-06a6880b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r50_fpn_mstrain_3x_coco/paa_r50_fpn_mstrain_3x_coco_20210121_145722.log.json) |
+| R-101-FPN | 12e | 6.2 | True | 42.6 | [config](./paa_r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r101_fpn_1x_coco/paa_r101_fpn_1x_coco_20200821-0a1825a4.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r101_fpn_1x_coco/paa_r101_fpn_1x_coco_20200821-0a1825a4.log.json) |
+| R-101-FPN | 12e | 6.2 | False | 42.4 | - | |
+| R-101-FPN | 24e | 6.2 | True | 43.5 | [config](./paa_r101_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r101_fpn_2x_coco/paa_r101_fpn_2x_coco_20200821-6829f96b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r101_fpn_2x_coco/paa_r101_fpn_2x_coco_20200821-6829f96b.log.json) |
+| R-101-FPN | 36e | 6.2 | True | 45.1 | [config](./paa_r101_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r101_fpn_mstrain_3x_coco/paa_r101_fpn_mstrain_3x_coco_20210122_084202-83250d22.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r101_fpn_mstrain_3x_coco/paa_r101_fpn_mstrain_3x_coco_20210122_084202.log.json) |
+
+**Note**:
+
+1. We find that the performance is unstable with 1x setting and may fluctuate by about 0.2 mAP. We report the best results.
+
+## Citation
+
+```latex
+@inproceedings{paa-eccv2020,
+ title={Probabilistic Anchor Assignment with IoU Prediction for Object Detection},
+ author={Kim, Kang and Lee, Hee Seok},
+ booktitle = {ECCV},
+ year={2020}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..078b974971d3a3faf537cc52937278488923667e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/metafile.yml
@@ -0,0 +1,111 @@
+Collections:
+ - Name: PAA
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - FPN
+ - Probabilistic Anchor Assignment
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/2007.08103
+ Title: 'Probabilistic Anchor Assignment with IoU Prediction for Object Detection'
+ README: configs/paa/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.4.0/mmdet/models/detectors/paa.py#L6
+ Version: v2.4.0
+
+Models:
+ - Name: paa_r50_fpn_1x_coco
+ In Collection: PAA
+ Config: configs/paa/paa_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.7
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r50_fpn_1x_coco/paa_r50_fpn_1x_coco_20200821-936edec3.pth
+
+ - Name: paa_r50_fpn_1.5x_coco
+ In Collection: PAA
+ Config: configs/paa/paa_r50_fpn_1.5x_coco.py
+ Metadata:
+ Training Memory (GB): 3.7
+ Epochs: 18
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r50_fpn_1.5x_coco/paa_r50_fpn_1.5x_coco_20200823-805d6078.pth
+
+ - Name: paa_r50_fpn_2x_coco
+ In Collection: PAA
+ Config: configs/paa/paa_r50_fpn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 3.7
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r50_fpn_2x_coco/paa_r50_fpn_2x_coco_20200821-c98bfc4e.pth
+
+ - Name: paa_r50_fpn_mstrain_3x_coco
+ In Collection: PAA
+ Config: configs/paa/paa_r50_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 3.7
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r50_fpn_mstrain_3x_coco/paa_r50_fpn_mstrain_3x_coco_20210121_145722-06a6880b.pth
+
+ - Name: paa_r101_fpn_1x_coco
+ In Collection: PAA
+ Config: configs/paa/paa_r101_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.2
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r101_fpn_1x_coco/paa_r101_fpn_1x_coco_20200821-0a1825a4.pth
+
+ - Name: paa_r101_fpn_2x_coco
+ In Collection: PAA
+ Config: configs/paa/paa_r101_fpn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 6.2
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r101_fpn_2x_coco/paa_r101_fpn_2x_coco_20200821-6829f96b.pth
+
+ - Name: paa_r101_fpn_mstrain_3x_coco
+ In Collection: PAA
+ Config: configs/paa/paa_r101_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 6.2
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/paa/paa_r101_fpn_mstrain_3x_coco/paa_r101_fpn_mstrain_3x_coco_20210122_084202-83250d22.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..94f1c278dc16c1befbca510ca0ac5ba407969f6d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r101_fpn_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './paa_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r101_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r101_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c6136f3bb404df6a6fc18536e6770116738af6c7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r101_fpn_2x_coco.py
@@ -0,0 +1,18 @@
+_base_ = './paa_r101_fpn_1x_coco.py'
+max_epochs = 24
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
+
+# training schedule for 2x
+train_cfg = dict(max_epochs=max_epochs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r101_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r101_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8529dcdb90adb2b02162f4d2268088f5f376fcb0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r101_fpn_ms-3x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './paa_r50_fpn_ms-3x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r50_fpn_1.5x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r50_fpn_1.5x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ae993b5c4370c8fc3e450f84fb7058528b853727
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r50_fpn_1.5x_coco.py
@@ -0,0 +1,18 @@
+_base_ = './paa_r50_fpn_1x_coco.py'
+max_epochs = 18
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[12, 16],
+ gamma=0.1)
+]
+
+# training schedule for 1.5x
+train_cfg = dict(max_epochs=max_epochs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f806a3ea65ffb9ee8b898122fb678b94ef212637
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r50_fpn_1x_coco.py
@@ -0,0 +1,80 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+# model settings
+model = dict(
+ type='PAA',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output',
+ num_outs=5),
+ bbox_head=dict(
+ type='PAAHead',
+ reg_decoded_bbox=True,
+ score_voting=True,
+ topk=9,
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ octave_base_scale=8,
+ scales_per_octave=1,
+ strides=[8, 16, 32, 64, 128]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=1.3),
+ loss_centerness=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=0.5)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.1,
+ neg_iou_thr=0.1,
+ min_pos_iou=0,
+ ignore_iof_thr=-1),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r50_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r50_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6908e4eb97fcfa92a20d486ceab9a7ddfaf480b7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r50_fpn_2x_coco.py
@@ -0,0 +1,18 @@
+_base_ = './paa_r50_fpn_1x_coco.py'
+max_epochs = 24
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
+
+# training schedule for 2x
+train_cfg = dict(max_epochs=max_epochs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r50_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r50_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..fed8b90a0fde7a1d344160a6658be04d1f9c654e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/paa/paa_r50_fpn_ms-3x_coco.py
@@ -0,0 +1,29 @@
+_base_ = './paa_r50_fpn_1x_coco.py'
+max_epochs = 36
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[28, 34],
+ gamma=0.1)
+]
+
+# training schedule for 3x
+train_cfg = dict(max_epochs=max_epochs)
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize', scale=[(1333, 640), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pafpn/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pafpn/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..36cd6e9fd5d6e31ac94e59c06cc1055be8480d21
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pafpn/README.md
@@ -0,0 +1,34 @@
+# PAFPN
+
+> [Path Aggregation Network for Instance Segmentation](https://arxiv.org/abs/1803.01534)
+
+
+
+## Abstract
+
+The way that information propagates in neural networks is of great importance. In this paper, we propose Path Aggregation Network (PANet) aiming at boosting information flow in proposal-based instance segmentation framework. Specifically, we enhance the entire feature hierarchy with accurate localization signals in lower layers by bottom-up path augmentation, which shortens the information path between lower layers and topmost feature. We present adaptive feature pooling, which links feature grid and all feature levels to make useful information in each feature level propagate directly to following proposal subnetworks. A complementary branch capturing different views for each proposal is created to further improve mask prediction. These improvements are simple to implement, with subtle extra computational overhead. Our PANet reaches the 1st place in the COCO 2017 Challenge Instance Segmentation task and the 2nd place in Object Detection task without large-batch training. It is also state-of-the-art on MVD and Cityscapes.
+
+
+

+
+
+## Results and Models
+
+| Backbone | style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :------------------------------------------: | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | pytorch | 1x | 4.0 | 17.2 | 37.5 | | [config](./faster-rcnn_r50_pafpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pafpn/faster_rcnn_r50_pafpn_1x_coco/faster_rcnn_r50_pafpn_1x_coco_bbox_mAP-0.375_20200503_105836-b7b4b9bd.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pafpn/faster_rcnn_r50_pafpn_1x_coco/faster_rcnn_r50_pafpn_1x_coco_20200503_105836.log.json) |
+
+## Citation
+
+```latex
+@inproceedings{liu2018path,
+ author = {Shu Liu and
+ Lu Qi and
+ Haifang Qin and
+ Jianping Shi and
+ Jiaya Jia},
+ title = {Path Aggregation Network for Instance Segmentation},
+ booktitle = {Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR)},
+ year = {2018}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pafpn/faster-rcnn_r50_pafpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pafpn/faster-rcnn_r50_pafpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1452baeca7e680b11f9b2ec654abe689d3e53042
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pafpn/faster-rcnn_r50_pafpn_1x_coco.py
@@ -0,0 +1,8 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+
+model = dict(
+ neck=dict(
+ type='PAFPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pafpn/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pafpn/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..7772d276ab6f0da685ed8ea5e58efd8fc5164529
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pafpn/metafile.yml
@@ -0,0 +1,38 @@
+Collections:
+ - Name: PAFPN
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - PAFPN
+ Paper:
+ URL: https://arxiv.org/abs/1803.01534
+ Title: 'Path Aggregation Network for Instance Segmentation'
+ README: configs/pafpn/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/necks/pafpn.py#L11
+ Version: v2.0.0
+
+Models:
+ - Name: faster-rcnn_r50_pafpn_1x_coco
+ In Collection: PAFPN
+ Config: configs/pafpn/faster-rcnn_r50_pafpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.0
+ inference time (ms/im):
+ - value: 58.14
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/pafpn/faster_rcnn_r50_pafpn_1x_coco/faster_rcnn_r50_pafpn_1x_coco_bbox_mAP-0.375_20200503_105836-b7b4b9bd.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/panoptic_fpn/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/panoptic_fpn/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..0321fb7ce1db42868c7753ce56fb330fef7e4764
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/panoptic_fpn/README.md
@@ -0,0 +1,62 @@
+# Panoptic FPN
+
+> [Panoptic feature pyramid networks](https://arxiv.org/abs/1901.02446)
+
+
+
+## Abstract
+
+The recently introduced panoptic segmentation task has renewed our community's interest in unifying the tasks of instance segmentation (for thing classes) and semantic segmentation (for stuff classes). However, current state-of-the-art methods for this joint task use separate and dissimilar networks for instance and semantic segmentation, without performing any shared computation. In this work, we aim to unify these methods at the architectural level, designing a single network for both tasks. Our approach is to endow Mask R-CNN, a popular instance segmentation method, with a semantic segmentation branch using a shared Feature Pyramid Network (FPN) backbone. Surprisingly, this simple baseline not only remains effective for instance segmentation, but also yields a lightweight, top-performing method for semantic segmentation. In this work, we perform a detailed study of this minimally extended version of Mask R-CNN with FPN, which we refer to as Panoptic FPN, and show it is a robust and accurate baseline for both tasks. Given its effectiveness and conceptual simplicity, we hope our method can serve as a strong baseline and aid future research in panoptic segmentation.
+
+
+

+
+
+## Dataset
+
+PanopticFPN requires COCO and [COCO-panoptic](http://images.cocodataset.org/annotations/panoptic_annotations_trainval2017.zip) dataset for training and evaluation. You need to download and extract it in the COCO dataset path.
+The directory should be like this.
+
+```none
+mmdetection
+├── mmdet
+├── tools
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ │ ├── panoptic_train2017.json
+│ │ │ ├── panoptic_train2017
+│ │ │ ├── panoptic_val2017.json
+│ │ │ ├── panoptic_val2017
+│ │ ├── train2017
+│ │ ├── val2017
+│ │ ├── test2017
+```
+
+## Results and Models
+
+| Backbone | style | Lr schd | Mem (GB) | Inf time (fps) | PQ | SQ | RQ | PQ_th | SQ_th | RQ_th | PQ_st | SQ_st | RQ_st | Config | Download |
+| :-------: | :-----: | :-----: | :------: | :------------: | :--: | :--: | :--: | :---: | :---: | :---: | :---: | :---: | :---: | :---------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | pytorch | 1x | 4.7 | | 40.2 | 77.8 | 49.3 | 47.8 | 80.9 | 57.5 | 28.9 | 73.1 | 37.0 | [config](./panoptic-fpn_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/panoptic_fpn/panoptic_fpn_r50_fpn_1x_coco/panoptic_fpn_r50_fpn_1x_coco_20210821_101153-9668fd13.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/panoptic_fpn/panoptic_fpn_r50_fpn_1x_coco/panoptic_fpn_r50_fpn_1x_coco_20210821_101153.log.json) |
+| R-50-FPN | pytorch | 3x | - | - | 42.5 | 78.1 | 51.7 | 50.3 | 81.5 | 60.3 | 30.7 | 73.0 | 38.8 | [config](./panoptic-fpn_r50_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/panoptic_fpn/panoptic_fpn_r50_fpn_mstrain_3x_coco/panoptic_fpn_r50_fpn_mstrain_3x_coco_20210824_171155-5650f98b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/panoptic_fpn/panoptic_fpn_r50_fpn_mstrain_3x_coco/panoptic_fpn_r50_fpn_mstrain_3x_coco_20210824_171155.log.json) |
+| R-101-FPN | pytorch | 1x | 6.7 | | 42.2 | 78.3 | 51.4 | 50.1 | 81.4 | 59.9 | 30.3 | 73.6 | 38.5 | [config](./panoptic-fpn_r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/panoptic_fpn/panoptic_fpn_r101_fpn_1x_coco/panoptic_fpn_r101_fpn_1x_coco_20210820_193950-ab9157a2.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/panoptic_fpn/panoptic_fpn_r101_fpn_1x_coco/panoptic_fpn_r101_fpn_1x_coco_20210820_193950.log.json) |
+| R-101-FPN | pytorch | 3x | - | - | 44.1 | 78.9 | 53.6 | 52.1 | 81.7 | 62.3 | 32.0 | 74.6 | 40.3 | [config](./panoptic-fpn_r101_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/panoptic_fpn/panoptic_fpn_r101_fpn_mstrain_3x_coco/panoptic_fpn_r101_fpn_mstrain_3x_coco_20210823_114712-9c99acc4.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/panoptic_fpn/panoptic_fpn_r101_fpn_mstrain_3x_coco/panoptic_fpn_r101_fpn_mstrain_3x_coco_20210823_114712.log.json) |
+
+## Citation
+
+The base method for panoptic segmentation task.
+
+```latex
+@inproceedings{kirillov2018panopticfpn,
+ author = {
+ Alexander Kirillov,
+ Ross Girshick,
+ Kaiming He,
+ Piotr Dollar,
+ },
+ title = {Panoptic Feature Pyramid Networks},
+ booktitle = {Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR)},
+ year = {2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/panoptic_fpn/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/panoptic_fpn/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..c99275ec3f37f47db756b96a4603c466d5fbd946
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/panoptic_fpn/metafile.yml
@@ -0,0 +1,70 @@
+Collections:
+ - Name: PanopticFPN
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - PanopticFPN
+ Paper:
+ URL: https://arxiv.org/pdf/1901.02446
+ Title: 'Panoptic feature pyramid networks'
+ README: configs/panoptic_fpn/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.16.0/mmdet/models/detectors/panoptic_fpn.py#L7
+ Version: v2.16.0
+
+Models:
+ - Name: panoptic_fpn_r50_fpn_1x_coco
+ In Collection: PanopticFPN
+ Config: configs/panoptic_fpn/panoptic-fpn_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.6
+ Epochs: 12
+ Results:
+ - Task: Panoptic Segmentation
+ Dataset: COCO
+ Metrics:
+ PQ: 40.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/panoptic_fpn/panoptic_fpn_r50_fpn_1x_coco/panoptic_fpn_r50_fpn_1x_coco_20210821_101153-9668fd13.pth
+
+ - Name: panoptic_fpn_r50_fpn_mstrain_3x_coco
+ In Collection: PanopticFPN
+ Config: configs/panoptic_fpn/panoptic-fpn_r50_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 4.6
+ Epochs: 36
+ Results:
+ - Task: Panoptic Segmentation
+ Dataset: COCO
+ Metrics:
+ PQ: 42.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/panoptic_fpn/panoptic_fpn_r50_fpn_mstrain_3x_coco/panoptic_fpn_r50_fpn_mstrain_3x_coco_20210824_171155-5650f98b.pth
+
+ - Name: panoptic_fpn_r101_fpn_1x_coco
+ In Collection: PanopticFPN
+ Config: configs/panoptic_fpn/panoptic-fpn_r101_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.5
+ Epochs: 12
+ Results:
+ - Task: Panoptic Segmentation
+ Dataset: COCO
+ Metrics:
+ PQ: 42.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/panoptic_fpn/panoptic_fpn_r101_fpn_1x_coco/panoptic_fpn_r101_fpn_1x_coco_20210820_193950-ab9157a2.pth
+
+ - Name: panoptic_fpn_r101_fpn_mstrain_3x_coco
+ In Collection: PanopticFPN
+ Config: configs/panoptic_fpn/panoptic-fpn_r101_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 6.5
+ Epochs: 36
+ Results:
+ - Task: Panoptic Segmentation
+ Dataset: COCO
+ Metrics:
+ PQ: 44.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/panoptic_fpn/panoptic_fpn_r101_fpn_mstrain_3x_coco/panoptic_fpn_r101_fpn_mstrain_3x_coco_20210823_114712-9c99acc4.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/panoptic_fpn/panoptic-fpn_r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/panoptic_fpn/panoptic-fpn_r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b960254ef5ecfac1de790a66a5378535114e9ba3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/panoptic_fpn/panoptic-fpn_r101_fpn_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './panoptic-fpn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/panoptic_fpn/panoptic-fpn_r101_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/panoptic_fpn/panoptic-fpn_r101_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..268782ee2cca31796e43423300319176556cfef7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/panoptic_fpn/panoptic-fpn_r101_fpn_ms-3x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './panoptic-fpn_r50_fpn_ms-3x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/panoptic_fpn/panoptic-fpn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/panoptic_fpn/panoptic-fpn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c2c89ef520124a43c910b35a4808153e4c455d3a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/panoptic_fpn/panoptic-fpn_r50_fpn_1x_coco.py
@@ -0,0 +1,45 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_panoptic.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ type='PanopticFPN',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32,
+ pad_mask=True,
+ mask_pad_value=0,
+ pad_seg=True,
+ seg_pad_value=255),
+ semantic_head=dict(
+ type='PanopticFPNHead',
+ num_things_classes=80,
+ num_stuff_classes=53,
+ in_channels=256,
+ inner_channels=128,
+ start_level=0,
+ end_level=4,
+ norm_cfg=dict(type='GN', num_groups=32, requires_grad=True),
+ conv_cfg=None,
+ loss_seg=dict(
+ type='CrossEntropyLoss', ignore_index=255, loss_weight=0.5)),
+ panoptic_fusion_head=dict(
+ type='HeuristicFusionHead',
+ num_things_classes=80,
+ num_stuff_classes=53),
+ test_cfg=dict(
+ rcnn=dict(
+ score_thr=0.6,
+ nms=dict(type='nms', iou_threshold=0.5, class_agnostic=True),
+ max_per_img=100,
+ mask_thr_binary=0.5),
+ # used in HeuristicFusionHead
+ panoptic=dict(mask_overlap=0.5, stuff_area_limit=4096)))
+
+# Forced to remove NumClassCheckHook
+custom_hooks = []
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/panoptic_fpn/panoptic-fpn_r50_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/panoptic_fpn/panoptic-fpn_r50_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b18a8f8dd7eb6c49e277346ffe71c6e36c9d3b68
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/panoptic_fpn/panoptic-fpn_r50_fpn_ms-3x_coco.py
@@ -0,0 +1,35 @@
+_base_ = './panoptic-fpn_r50_fpn_1x_coco.py'
+
+# In mstrain 3x config, img_scale=[(1333, 640), (1333, 800)],
+# multiscale_mode='range'
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(
+ type='LoadPanopticAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ with_seg=True),
+ dict(
+ type='RandomResize', scale=[(1333, 640), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+# TODO: Use RepeatDataset to speed up training
+# training schedule for 3x
+train_cfg = dict(max_epochs=36, val_interval=3)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=36,
+ by_epoch=True,
+ milestones=[24, 33],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..2ead3add79ec914d9720562ce4c4c121fac15a7e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/README.md
@@ -0,0 +1,40 @@
+# Pascal VOC
+
+> [The Pascal Visual Object Classes (VOC) Challenge](https://link.springer.com/article/10.1007/s11263-009-0275-4)
+
+
+
+## Abstract
+
+The Pascal Visual Object Classes (VOC) challenge is a benchmark in visual object category recognition and detection, providing the vision and machine learning communities with a standard dataset of images and annotation, and standard evaluation procedures. Organised annually from 2005 to present, the challenge and its associated dataset has become accepted as the benchmark for object detection.
+
+This paper describes the dataset and evaluation procedure. We review the state-of-the-art in evaluated methods for both classification and detection, analyse whether the methods are statistically different, what they are learning from the images (e.g. the object or its context), and what the methods find easy or confuse. The paper concludes with lessons learnt in the three year history of the challenge, and proposes directions for future improvement and extension.
+
+
+

+
+
+## Results and Models
+
+| Architecture | Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :-------------: | :------: | :-----: | :-----: | :------: | :------------: | :----: | :----------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Faster R-CNN C4 | R-50 | caffe | 18k | | - | 80.9 | [config](./faster-rcnn_r50-caffe-c4_ms-18k_voc0712.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pascal_voc/faster_rcnn_r50_caffe_c4_mstrain_18k_voc0712//home/dong/code_sensetime/2022Q1/mmdetection/work_dirs/prepare_voc/gather/pascal_voc/faster_rcnn_r50_caffe_c4_mstrain_18k_voc0712/faster_rcnn_r50_caffe_c4_mstrain_18k_voc0712_20220314_234327-847a14d2.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pascal_voc/faster_rcnn_r50_caffe_c4_mstrain_18k_voc0712/faster_rcnn_r50_caffe_c4_mstrain_18k_voc0712_20220314_234327.log.json) |
+| Faster R-CNN | R-50 | pytorch | 1x | 2.6 | - | 80.4 | [config](./faster-rcnn_r50_fpn_1x_voc0712.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pascal_voc/faster_rcnn_r50_fpn_1x_voc0712/faster_rcnn_r50_fpn_1x_voc0712_20220320_192712-54bef0f3.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pascal_voc/faster_rcnn_r50_fpn_1x_voc0712/faster_rcnn_r50_fpn_1x_voc0712_20220320_192712.log.json) |
+| Retinanet | R-50 | pytorch | 1x | 2.1 | - | 77.3 | [config](./retinanet_r50_fpn_1x_voc0712.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pascal_voc/retinanet_r50_fpn_1x_voc0712/retinanet_r50_fpn_1x_voc0712_20200617-47cbdd0e.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pascal_voc/retinanet_r50_fpn_1x_voc0712/retinanet_r50_fpn_1x_voc0712_20200616_014642.log.json) |
+| SSD300 | VGG16 | - | 120e | - | - | 76.5 | [config](./ssd300_voc0712.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pascal_voc/ssd300_voc0712/ssd300_voc0712_20220320_194658-17edda1b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pascal_voc/ssd300_voc0712/ssd300_voc0712_20220320_194658.log.json) |
+| SSD512 | VGG16 | - | 120e | - | - | 79.5 | [config](./ssd512_voc0712.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pascal_voc/ssd512_voc0712/ssd512_voc0712_20220320_194717-03cefefe.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pascal_voc/ssd512_voc0712/ssd512_voc0712_20220320_194717.log.json) |
+
+## Citation
+
+```latex
+@Article{Everingham10,
+ author = "Everingham, M. and Van~Gool, L. and Williams, C. K. I. and Winn, J. and Zisserman, A.",
+ title = "The Pascal Visual Object Classes (VOC) Challenge",
+ journal = "International Journal of Computer Vision",
+ volume = "88",
+ year = "2010",
+ number = "2",
+ month = jun,
+ pages = "303--338",
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/faster-rcnn_r50-caffe-c4_ms-18k_voc0712.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/faster-rcnn_r50-caffe-c4_ms-18k_voc0712.py
new file mode 100644
index 0000000000000000000000000000000000000000..dddc0bbdf33948478e11bb701f844a8473ddf165
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/faster-rcnn_r50-caffe-c4_ms-18k_voc0712.py
@@ -0,0 +1,86 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50-caffe-c4.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/datasets/voc0712.py',
+ '../_base_/default_runtime.py'
+]
+model = dict(roi_head=dict(bbox_head=dict(num_classes=20)))
+
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 480), (1333, 512), (1333, 544), (1333, 576),
+ (1333, 608), (1333, 640), (1333, 672), (1333, 704),
+ (1333, 736), (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ # avoid bboxes being resized
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ sampler=dict(type='InfiniteSampler', shuffle=True),
+ dataset=dict(
+ _delete_=True,
+ type='ConcatDataset',
+ datasets=[
+ dict(
+ type='VOCDataset',
+ data_root={{_base_.data_root}},
+ ann_file='VOC2007/ImageSets/Main/trainval.txt',
+ data_prefix=dict(sub_data_root='VOC2007/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args={{_base_.backend_args}}),
+ dict(
+ type='VOCDataset',
+ data_root={{_base_.data_root}},
+ ann_file='VOC2012/ImageSets/Main/trainval.txt',
+ data_prefix=dict(sub_data_root='VOC2012/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args={{_base_.backend_args}})
+ ]))
+
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# training schedule for 18k
+max_iter = 18000
+train_cfg = dict(
+ _delete_=True,
+ type='IterBasedTrainLoop',
+ max_iters=max_iter,
+ val_interval=3000)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=100),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_iter,
+ by_epoch=False,
+ milestones=[12000, 16000],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.0001))
+
+default_hooks = dict(checkpoint=dict(by_epoch=False, interval=3000))
+log_processor = dict(by_epoch=False)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/faster-rcnn_r50_fpn_1x_voc0712-cocofmt.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/faster-rcnn_r50_fpn_1x_voc0712-cocofmt.py
new file mode 100644
index 0000000000000000000000000000000000000000..0b0aa41d67fc4edfde6d534e2e54a135f5de6e44
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/faster-rcnn_r50_fpn_1x_voc0712-cocofmt.py
@@ -0,0 +1,100 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py', '../_base_/datasets/voc0712.py',
+ '../_base_/default_runtime.py'
+]
+model = dict(roi_head=dict(bbox_head=dict(num_classes=20)))
+
+METAINFO = {
+ 'classes':
+ ('aeroplane', 'bicycle', 'bird', 'boat', 'bottle', 'bus', 'car', 'cat',
+ 'chair', 'cow', 'diningtable', 'dog', 'horse', 'motorbike', 'person',
+ 'pottedplant', 'sheep', 'sofa', 'train', 'tvmonitor'),
+ # palette is a list of color tuples, which is used for visualization.
+ 'palette': [(106, 0, 228), (119, 11, 32), (165, 42, 42), (0, 0, 192),
+ (197, 226, 255), (0, 60, 100), (0, 0, 142), (255, 77, 255),
+ (153, 69, 1), (120, 166, 157), (0, 182, 199), (0, 226, 252),
+ (182, 182, 255), (0, 0, 230), (220, 20, 60), (163, 255, 0),
+ (0, 82, 0), (3, 95, 161), (0, 80, 100), (183, 130, 88)]
+}
+
+# dataset settings
+dataset_type = 'CocoDataset'
+data_root = 'data/VOCdevkit/'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='Resize', scale=(1000, 600), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(1000, 600), keep_ratio=True),
+ # avoid bboxes being resized
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ dataset=dict(
+ type='RepeatDataset',
+ times=3,
+ dataset=dict(
+ _delete_=True,
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/voc0712_trainval.json',
+ data_prefix=dict(img=''),
+ metainfo=METAINFO,
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args={{_base_.backend_args}})))
+val_dataloader = dict(
+ dataset=dict(
+ type=dataset_type,
+ ann_file='annotations/voc07_test.json',
+ data_prefix=dict(img=''),
+ metainfo=METAINFO,
+ pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/voc07_test.json',
+ metric='bbox',
+ format_only=False,
+ backend_args={{_base_.backend_args}})
+test_evaluator = val_evaluator
+
+# training schedule, the dataset is repeated 3 times, so the
+# actual epoch = 4 * 3 = 12
+max_epochs = 4
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[3],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/faster-rcnn_r50_fpn_1x_voc0712.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/faster-rcnn_r50_fpn_1x_voc0712.py
new file mode 100644
index 0000000000000000000000000000000000000000..07391667b35c9db9e352a03624411bb568f5396a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/faster-rcnn_r50_fpn_1x_voc0712.py
@@ -0,0 +1,35 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py', '../_base_/datasets/voc0712.py',
+ '../_base_/default_runtime.py'
+]
+model = dict(roi_head=dict(bbox_head=dict(num_classes=20)))
+
+# training schedule, voc dataset is repeated 3 times, in
+# `_base_/datasets/voc0712.py`, so the actual epoch = 4 * 3 = 12
+max_epochs = 4
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[3],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/retinanet_r50_fpn_1x_voc0712.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/retinanet_r50_fpn_1x_voc0712.py
new file mode 100644
index 0000000000000000000000000000000000000000..c86a6f199c9317804692189975f3abaff24f6aff
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/retinanet_r50_fpn_1x_voc0712.py
@@ -0,0 +1,34 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py', '../_base_/datasets/voc0712.py',
+ '../_base_/default_runtime.py'
+]
+model = dict(bbox_head=dict(num_classes=20))
+
+# training schedule, voc dataset is repeated 3 times, in
+# `_base_/datasets/voc0712.py`, so the actual epoch = 4 * 3 = 12
+max_epochs = 4
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[3],
+ gamma=0.1)
+]
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/ssd300_voc0712.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/ssd300_voc0712.py
new file mode 100644
index 0000000000000000000000000000000000000000..ff7a1368b76aa53700bd81a912b54e84ab58e53a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/ssd300_voc0712.py
@@ -0,0 +1,102 @@
+_base_ = [
+ '../_base_/models/ssd300.py', '../_base_/datasets/voc0712.py',
+ '../_base_/schedules/schedule_2x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ bbox_head=dict(
+ num_classes=20, anchor_generator=dict(basesize_ratio_range=(0.2,
+ 0.9))))
+# dataset settings
+dataset_type = 'VOCDataset'
+data_root = 'data/VOCdevkit/'
+input_size = 300
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='Expand',
+ mean={{_base_.model.data_preprocessor.mean}},
+ to_rgb={{_base_.model.data_preprocessor.bgr_to_rgb}},
+ ratio_range=(1, 4)),
+ dict(
+ type='MinIoURandomCrop',
+ min_ious=(0.1, 0.3, 0.5, 0.7, 0.9),
+ min_crop_size=0.3),
+ dict(type='Resize', scale=(input_size, input_size), keep_ratio=False),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='PhotoMetricDistortion',
+ brightness_delta=32,
+ contrast_range=(0.5, 1.5),
+ saturation_range=(0.5, 1.5),
+ hue_delta=18),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='Resize', scale=(input_size, input_size), keep_ratio=False),
+ # avoid bboxes being resized
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=8,
+ num_workers=3,
+ dataset=dict( # RepeatDataset
+ # the dataset is repeated 10 times, and the training schedule is 2x,
+ # so the actual epoch = 12 * 10 = 120.
+ times=10,
+ dataset=dict( # ConcatDataset
+ # VOCDataset will add different `dataset_type` in dataset.metainfo,
+ # which will get error if using ConcatDataset. Adding
+ # `ignore_keys` can avoid this error.
+ ignore_keys=['dataset_type'],
+ datasets=[
+ dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='VOC2007/ImageSets/Main/trainval.txt',
+ data_prefix=dict(sub_data_root='VOC2007/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline),
+ dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='VOC2012/ImageSets/Main/trainval.txt',
+ data_prefix=dict(sub_data_root='VOC2012/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline)
+ ])))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+custom_hooks = [
+ dict(type='NumClassCheckHook'),
+ dict(type='CheckInvalidLossHook', interval=50, priority='VERY_LOW')
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=1e-3, momentum=0.9, weight_decay=5e-4))
+
+# learning policy
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=24,
+ by_epoch=True,
+ milestones=[16, 20],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/ssd512_voc0712.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/ssd512_voc0712.py
new file mode 100644
index 0000000000000000000000000000000000000000..6c4dc8a3eec86ccced7d44120b254463d18c00f5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pascal_voc/ssd512_voc0712.py
@@ -0,0 +1,82 @@
+_base_ = 'ssd300_voc0712.py'
+
+input_size = 512
+model = dict(
+ neck=dict(
+ out_channels=(512, 1024, 512, 256, 256, 256, 256),
+ level_strides=(2, 2, 2, 2, 1),
+ level_paddings=(1, 1, 1, 1, 1),
+ last_kernel_size=4),
+ bbox_head=dict(
+ in_channels=(512, 1024, 512, 256, 256, 256, 256),
+ anchor_generator=dict(
+ input_size=input_size,
+ strides=[8, 16, 32, 64, 128, 256, 512],
+ basesize_ratio_range=(0.15, 0.9),
+ ratios=([2], [2, 3], [2, 3], [2, 3], [2, 3], [2], [2]))))
+
+# dataset settings
+dataset_type = 'VOCDataset'
+data_root = 'data/VOCdevkit/'
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='Expand',
+ mean={{_base_.model.data_preprocessor.mean}},
+ to_rgb={{_base_.model.data_preprocessor.bgr_to_rgb}},
+ ratio_range=(1, 4)),
+ dict(
+ type='MinIoURandomCrop',
+ min_ious=(0.1, 0.3, 0.5, 0.7, 0.9),
+ min_crop_size=0.3),
+ dict(type='Resize', scale=(input_size, input_size), keep_ratio=False),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='PhotoMetricDistortion',
+ brightness_delta=32,
+ contrast_range=(0.5, 1.5),
+ saturation_range=(0.5, 1.5),
+ hue_delta=18),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='Resize', scale=(input_size, input_size), keep_ratio=False),
+ # avoid bboxes being resized
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=8,
+ num_workers=3,
+ dataset=dict( # RepeatDataset
+ # the dataset is repeated 10 times, and the training schedule is 2x,
+ # so the actual epoch = 12 * 10 = 120.
+ times=10,
+ dataset=dict( # ConcatDataset
+ # VOCDataset will add different `dataset_type` in dataset.metainfo,
+ # which will get error if using ConcatDataset. Adding
+ # `ignore_keys` can avoid this error.
+ ignore_keys=['dataset_type'],
+ datasets=[
+ dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='VOC2007/ImageSets/Main/trainval.txt',
+ data_prefix=dict(sub_data_root='VOC2007/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline),
+ dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='VOC2012/ImageSets/Main/trainval.txt',
+ data_prefix=dict(sub_data_root='VOC2012/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline)
+ ])))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..39f79ecd1b9b007b6bbf1417e6fd809d47141470
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/README.md
@@ -0,0 +1,50 @@
+# PISA
+
+> [Prime Sample Attention in Object Detection](https://arxiv.org/abs/1904.04821)
+
+
+
+## Abstract
+
+It is a common paradigm in object detection frameworks to treat all samples equally and target at maximizing the performance on average. In this work, we revisit this paradigm through a careful study on how different samples contribute to the overall performance measured in terms of mAP. Our study suggests that the samples in each mini-batch are neither independent nor equally important, and therefore a better classifier on average does not necessarily mean higher mAP. Motivated by this study, we propose the notion of Prime Samples, those that play a key role in driving the detection performance. We further develop a simple yet effective sampling and learning strategy called PrIme Sample Attention (PISA) that directs the focus of the training process towards such samples. Our experiments demonstrate that it is often more effective to focus on prime samples than hard samples when training a detector. Particularly, On the MSCOCO dataset, PISA outperforms the random sampling baseline and hard mining schemes, e.g., OHEM and Focal Loss, consistently by around 2% on both single-stage and two-stage detectors, even with a strong backbone ResNeXt-101.
+
+
+

+
+
+## Results and Models
+
+| PISA | Network | Backbone | Lr schd | box AP | mask AP | Config | Download |
+| :--: | :----------: | :------------: | :-----: | :----: | :-----: | :----------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| × | Faster R-CNN | R-50-FPN | 1x | 36.4 | | - | |
+| √ | Faster R-CNN | R-50-FPN | 1x | 38.4 | | [config](./faster-rcnn_r50_fpn_pisa_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_faster_rcnn_r50_fpn_1x_coco/pisa_faster_rcnn_r50_fpn_1x_coco-dea93523.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_faster_rcnn_r50_fpn_1x_coco/pisa_faster_rcnn_r50_fpn_1x_coco_20200506_185619.log.json) |
+| × | Faster R-CNN | X101-32x4d-FPN | 1x | 40.1 | | - | |
+| √ | Faster R-CNN | X101-32x4d-FPN | 1x | 41.9 | | [config](./faster-rcnn_x101-32x4d_fpn_pisa_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_faster_rcnn_x101_32x4d_fpn_1x_coco/pisa_faster_rcnn_x101_32x4d_fpn_1x_coco-e4accec4.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_faster_rcnn_x101_32x4d_fpn_1x_coco/pisa_faster_rcnn_x101_32x4d_fpn_1x_coco_20200505_181503.log.json) |
+| × | Mask R-CNN | R-50-FPN | 1x | 37.3 | 34.2 | - | |
+| √ | Mask R-CNN | R-50-FPN | 1x | 39.1 | 35.2 | [config](./mask-rcnn_r50_fpn_pisa_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_mask_rcnn_r50_fpn_1x_coco/pisa_mask_rcnn_r50_fpn_1x_coco-dfcedba6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_mask_rcnn_r50_fpn_1x_coco/pisa_mask_rcnn_r50_fpn_1x_coco_20200508_150500.log.json) |
+| × | Mask R-CNN | X101-32x4d-FPN | 1x | 41.1 | 37.1 | - | |
+| √ | Mask R-CNN | X101-32x4d-FPN | 1x | | | | |
+| × | RetinaNet | R-50-FPN | 1x | 35.6 | | - | |
+| √ | RetinaNet | R-50-FPN | 1x | 36.9 | | [config](./retinanet-r50_fpn_pisa_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_retinanet_r50_fpn_1x_coco/pisa_retinanet_r50_fpn_1x_coco-76409952.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_retinanet_r50_fpn_1x_coco/pisa_retinanet_r50_fpn_1x_coco_20200504_014311.log.json) |
+| × | RetinaNet | X101-32x4d-FPN | 1x | 39.0 | | - | |
+| √ | RetinaNet | X101-32x4d-FPN | 1x | 40.7 | | [config](./retinanet_x101-32x4d_fpn_pisa_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_retinanet_x101_32x4d_fpn_1x_coco/pisa_retinanet_x101_32x4d_fpn_1x_coco-a0c13c73.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_retinanet_x101_32x4d_fpn_1x_coco/pisa_retinanet_x101_32x4d_fpn_1x_coco_20200505_001404.log.json) |
+| × | SSD300 | VGG16 | 1x | 25.6 | | - | |
+| √ | SSD300 | VGG16 | 1x | 27.6 | | [config](./ssd300_pisa_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_ssd300_coco/pisa_ssd300_coco-710e3ac9.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_ssd300_coco/pisa_ssd300_coco_20200504_144325.log.json) |
+| × | SSD512 | VGG16 | 1x | 29.3 | | - | |
+| √ | SSD512 | VGG16 | 1x | 31.8 | | [config](./ssd512_pisa_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_ssd512_coco/pisa_ssd512_coco-247addee.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_ssd512_coco/pisa_ssd512_coco_20200508_131030.log.json) |
+
+**Notes:**
+
+- In the original paper, all models are trained and tested on mmdet v1.x, thus results may not be exactly the same with this release on v2.0.
+- It is noted PISA only modifies the training pipeline so the inference time remains the same with the baseline.
+
+## Citation
+
+```latex
+@inproceedings{cao2019prime,
+ title={Prime sample attention in object detection},
+ author={Cao, Yuhang and Chen, Kai and Loy, Chen Change and Lin, Dahua},
+ booktitle={IEEE Conference on Computer Vision and Pattern Recognition},
+ year={2020}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/faster-rcnn_r50_fpn_pisa_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/faster-rcnn_r50_fpn_pisa_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..237a3b13aa5e61f04579670af01df8f481d80dd1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/faster-rcnn_r50_fpn_pisa_1x_coco.py
@@ -0,0 +1,30 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+
+model = dict(
+ roi_head=dict(
+ type='PISARoIHead',
+ bbox_head=dict(
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0))),
+ train_cfg=dict(
+ rpn_proposal=dict(
+ nms_pre=2000,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ sampler=dict(
+ type='ScoreHLRSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True,
+ k=0.5,
+ bias=0.),
+ isr=dict(k=2, bias=0),
+ carl=dict(k=1, bias=0.2))),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=2000,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/faster-rcnn_x101-32x4d_fpn_pisa_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/faster-rcnn_x101-32x4d_fpn_pisa_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..4b2c8d9a20ac7adf1965bb3d98e868c785cb23c3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/faster-rcnn_x101-32x4d_fpn_pisa_1x_coco.py
@@ -0,0 +1,30 @@
+_base_ = '../faster_rcnn/faster-rcnn_x101-32x4d_fpn_1x_coco.py'
+
+model = dict(
+ roi_head=dict(
+ type='PISARoIHead',
+ bbox_head=dict(
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0))),
+ train_cfg=dict(
+ rpn_proposal=dict(
+ nms_pre=2000,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ sampler=dict(
+ type='ScoreHLRSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True,
+ k=0.5,
+ bias=0.),
+ isr=dict(k=2, bias=0),
+ carl=dict(k=1, bias=0.2))),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=2000,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/mask-rcnn_r50_fpn_pisa_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/mask-rcnn_r50_fpn_pisa_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d6a6823591b1d7780c7f9d49029579afede239aa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/mask-rcnn_r50_fpn_pisa_1x_coco.py
@@ -0,0 +1,30 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+
+model = dict(
+ roi_head=dict(
+ type='PISARoIHead',
+ bbox_head=dict(
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0))),
+ train_cfg=dict(
+ rpn_proposal=dict(
+ nms_pre=2000,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ sampler=dict(
+ type='ScoreHLRSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True,
+ k=0.5,
+ bias=0.),
+ isr=dict(k=2, bias=0),
+ carl=dict(k=1, bias=0.2))),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=2000,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/mask-rcnn_x101-32x4d_fpn_pisa_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/mask-rcnn_x101-32x4d_fpn_pisa_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f2ac19fe75ba8c5b2440772eced16397e2273735
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/mask-rcnn_x101-32x4d_fpn_pisa_1x_coco.py
@@ -0,0 +1,30 @@
+_base_ = '../mask_rcnn/mask-rcnn_x101-32x4d_fpn_1x_coco.py'
+
+model = dict(
+ roi_head=dict(
+ type='PISARoIHead',
+ bbox_head=dict(
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0))),
+ train_cfg=dict(
+ rpn_proposal=dict(
+ nms_pre=2000,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ sampler=dict(
+ type='ScoreHLRSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True,
+ k=0.5,
+ bias=0.),
+ isr=dict(k=2, bias=0),
+ carl=dict(k=1, bias=0.2))),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=2000,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..3be5c3baf6d386d246b8fdc39035245d7dbbaad5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/metafile.yml
@@ -0,0 +1,110 @@
+Collections:
+ - Name: PISA
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - FPN
+ - PISA
+ - RPN
+ - ResNet
+ - RoIPool
+ Paper:
+ URL: https://arxiv.org/abs/1904.04821
+ Title: 'Prime Sample Attention in Object Detection'
+ README: configs/pisa/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/roi_heads/pisa_roi_head.py#L8
+ Version: v2.1.0
+
+Models:
+ - Name: pisa_faster_rcnn_r50_fpn_1x_coco
+ In Collection: PISA
+ Config: configs/pisa/faster-rcnn_r50_fpn_pisa_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_faster_rcnn_r50_fpn_1x_coco/pisa_faster_rcnn_r50_fpn_1x_coco-dea93523.pth
+
+ - Name: pisa_faster_rcnn_x101_32x4d_fpn_1x_coco
+ In Collection: PISA
+ Config: configs/pisa/faster-rcnn_x101-32x4d_fpn_pisa_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_faster_rcnn_x101_32x4d_fpn_1x_coco/pisa_faster_rcnn_x101_32x4d_fpn_1x_coco-e4accec4.pth
+
+ - Name: pisa_mask_rcnn_r50_fpn_1x_coco
+ In Collection: PISA
+ Config: configs/pisa/mask-rcnn_r50_fpn_pisa_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.1
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 35.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_mask_rcnn_r50_fpn_1x_coco/pisa_mask_rcnn_r50_fpn_1x_coco-dfcedba6.pth
+
+ - Name: pisa_retinanet_r50_fpn_1x_coco
+ In Collection: PISA
+ Config: configs/pisa/retinanet-r50_fpn_pisa_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 36.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_retinanet_r50_fpn_1x_coco/pisa_retinanet_r50_fpn_1x_coco-76409952.pth
+
+ - Name: pisa_retinanet_x101_32x4d_fpn_1x_coco
+ In Collection: PISA
+ Config: configs/pisa/retinanet_x101-32x4d_fpn_pisa_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_retinanet_x101_32x4d_fpn_1x_coco/pisa_retinanet_x101_32x4d_fpn_1x_coco-a0c13c73.pth
+
+ - Name: pisa_ssd300_coco
+ In Collection: PISA
+ Config: configs/pisa/ssd300_pisa_coco.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 27.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_ssd300_coco/pisa_ssd300_coco-710e3ac9.pth
+
+ - Name: pisa_ssd512_coco
+ In Collection: PISA
+ Config: configs/pisa/ssd512_pisa_coco.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 31.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/pisa/pisa_ssd512_coco/pisa_ssd512_coco-247addee.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/retinanet-r50_fpn_pisa_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/retinanet-r50_fpn_pisa_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..70f89e227ec64b5c7224375aac0cf7ae3a10a29e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/retinanet-r50_fpn_pisa_1x_coco.py
@@ -0,0 +1,7 @@
+_base_ = '../retinanet/retinanet_r50_fpn_1x_coco.py'
+
+model = dict(
+ bbox_head=dict(
+ type='PISARetinaHead',
+ loss_bbox=dict(type='SmoothL1Loss', beta=0.11, loss_weight=1.0)),
+ train_cfg=dict(isr=dict(k=2., bias=0.), carl=dict(k=1., bias=0.2)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/retinanet_x101-32x4d_fpn_pisa_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/retinanet_x101-32x4d_fpn_pisa_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..9caad45d34a9cde84a3c29ad45e3080bb831bb76
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/retinanet_x101-32x4d_fpn_pisa_1x_coco.py
@@ -0,0 +1,7 @@
+_base_ = '../retinanet/retinanet_x101-32x4d_fpn_1x_coco.py'
+
+model = dict(
+ bbox_head=dict(
+ type='PISARetinaHead',
+ loss_bbox=dict(type='SmoothL1Loss', beta=0.11, loss_weight=1.0)),
+ train_cfg=dict(isr=dict(k=2., bias=0.), carl=dict(k=1., bias=0.2)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/ssd300_pisa_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/ssd300_pisa_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b10236baeb1925483c2fdb025d86c45d51ba0276
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/ssd300_pisa_coco.py
@@ -0,0 +1,7 @@
+_base_ = '../ssd/ssd300_coco.py'
+
+model = dict(
+ bbox_head=dict(type='PISASSDHead'),
+ train_cfg=dict(isr=dict(k=2., bias=0.), carl=dict(k=1., bias=0.2)))
+
+optim_wrapper = dict(clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/ssd512_pisa_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/ssd512_pisa_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..939c7f453d4d881324c3b0443b0696eb96b3df4f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pisa/ssd512_pisa_coco.py
@@ -0,0 +1,7 @@
+_base_ = '../ssd/ssd512_coco.py'
+
+model = dict(
+ bbox_head=dict(type='PISASSDHead'),
+ train_cfg=dict(isr=dict(k=2., bias=0.), carl=dict(k=1., bias=0.2)))
+
+optim_wrapper = dict(clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/point_rend/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/point_rend/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..efa1dcac214adafb6f7a9b9c6aba97e9ecd7b51c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/point_rend/README.md
@@ -0,0 +1,33 @@
+# PointRend
+
+> [PointRend: Image Segmentation as Rendering](https://arxiv.org/abs/1912.08193)
+
+
+
+## Abstract
+
+We present a new method for efficient high-quality image segmentation of objects and scenes. By analogizing classical computer graphics methods for efficient rendering with over- and undersampling challenges faced in pixel labeling tasks, we develop a unique perspective of image segmentation as a rendering problem. From this vantage, we present the PointRend (Point-based Rendering) neural network module: a module that performs point-based segmentation predictions at adaptively selected locations based on an iterative subdivision algorithm. PointRend can be flexibly applied to both instance and semantic segmentation tasks by building on top of existing state-of-the-art models. While many concrete implementations of the general idea are possible, we show that a simple design already achieves excellent results. Qualitatively, PointRend outputs crisp object boundaries in regions that are over-smoothed by previous methods. Quantitatively, PointRend yields significant gains on COCO and Cityscapes, for both instance and semantic segmentation. PointRend's efficiency enables output resolutions that are otherwise impractical in terms of memory or computation compared to existing approaches.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :------: | :---: | :-----: | :------: | :------------: | :----: | :-----: | :------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | caffe | 1x | 4.6 | | 38.4 | 36.3 | [config](./point-rend_r50-caffe_fpn_ms-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/point_rend/point_rend_r50_caffe_fpn_mstrain_1x_coco/point_rend_r50_caffe_fpn_mstrain_1x_coco-1bcb5fb4.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/point_rend/point_rend_r50_caffe_fpn_mstrain_1x_coco/point_rend_r50_caffe_fpn_mstrain_1x_coco_20200612_161407.log.json) |
+| R-50-FPN | caffe | 3x | 4.6 | | 41.0 | 38.0 | [config](./point-rend_r50-caffe_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/point_rend/point_rend_r50_caffe_fpn_mstrain_3x_coco/point_rend_r50_caffe_fpn_mstrain_3x_coco-e0ebb6b7.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/point_rend/point_rend_r50_caffe_fpn_mstrain_3x_coco/point_rend_r50_caffe_fpn_mstrain_3x_coco_20200614_002632.log.json) |
+
+Note: All models are trained with multi-scale, the input image shorter side is randomly scaled to one of (640, 672, 704, 736, 768, 800).
+
+## Citation
+
+```latex
+@InProceedings{kirillov2019pointrend,
+ title={{PointRend}: Image Segmentation as Rendering},
+ author={Alexander Kirillov and Yuxin Wu and Kaiming He and Ross Girshick},
+ journal={ArXiv:1912.08193},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/point_rend/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/point_rend/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..f54f8a860b7951c1e99471b1f10e69c4685d998b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/point_rend/metafile.yml
@@ -0,0 +1,54 @@
+Collections:
+ - Name: PointRend
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - PointRend
+ - FPN
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/1912.08193
+ Title: 'PointRend: Image Segmentation as Rendering'
+ README: configs/point_rend/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.2.0/mmdet/models/detectors/point_rend.py#L6
+ Version: v2.2.0
+
+Models:
+ - Name: point_rend_r50_caffe_fpn_mstrain_1x_coco
+ In Collection: PointRend
+ Config: configs/point_rend/point-rend_r50-caffe_fpn_ms-1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.6
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/point_rend/point_rend_r50_caffe_fpn_mstrain_1x_coco/point_rend_r50_caffe_fpn_mstrain_1x_coco-1bcb5fb4.pth
+
+ - Name: point_rend_r50_caffe_fpn_mstrain_3x_coco
+ In Collection: PointRend
+ Config: configs/point_rend/point-rend_r50-caffe_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 4.6
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/point_rend/point_rend_r50_caffe_fpn_mstrain_3x_coco/point_rend_r50_caffe_fpn_mstrain_3x_coco-e0ebb6b7.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/point_rend/point-rend_r50-caffe_fpn_ms-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/point_rend/point-rend_r50-caffe_fpn_ms-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8b17f5a340bad54a8fe9b366ccc7d5574f687b17
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/point_rend/point-rend_r50-caffe_fpn_ms-1x_coco.py
@@ -0,0 +1,44 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50-caffe_fpn_ms-1x_coco.py'
+# model settings
+model = dict(
+ type='PointRend',
+ roi_head=dict(
+ type='PointRendRoIHead',
+ mask_roi_extractor=dict(
+ type='GenericRoIExtractor',
+ aggregation='concat',
+ roi_layer=dict(
+ _delete_=True, type='SimpleRoIAlign', output_size=14),
+ out_channels=256,
+ featmap_strides=[4]),
+ mask_head=dict(
+ _delete_=True,
+ type='CoarseMaskHead',
+ num_fcs=2,
+ in_channels=256,
+ conv_out_channels=256,
+ fc_out_channels=1024,
+ num_classes=80,
+ loss_mask=dict(
+ type='CrossEntropyLoss', use_mask=True, loss_weight=1.0)),
+ point_head=dict(
+ type='MaskPointHead',
+ num_fcs=3,
+ in_channels=256,
+ fc_channels=256,
+ num_classes=80,
+ coarse_pred_each_layer=True,
+ loss_point=dict(
+ type='CrossEntropyLoss', use_mask=True, loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rcnn=dict(
+ mask_size=7,
+ num_points=14 * 14,
+ oversample_ratio=3,
+ importance_sample_ratio=0.75)),
+ test_cfg=dict(
+ rcnn=dict(
+ subdivision_steps=5,
+ subdivision_num_points=28 * 28,
+ scale_factor=2)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/point_rend/point-rend_r50-caffe_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/point_rend/point-rend_r50-caffe_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b11faaa98ebc5b61f086a2297debda6769dc6270
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/point_rend/point-rend_r50-caffe_fpn_ms-3x_coco.py
@@ -0,0 +1,18 @@
+_base_ = './point-rend_r50-caffe_fpn_ms-1x_coco.py'
+
+max_epochs = 36
+
+# learning policy
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[28, 34],
+ gamma=0.1)
+]
+
+train_cfg = dict(max_epochs=max_epochs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..fccad4f6b8b7e6ac89e937fac6d7858ecbfa881b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/README.md
@@ -0,0 +1,57 @@
+# PVT
+
+> [Pyramid vision transformer: A versatile backbone for dense prediction without convolutions](https://arxiv.org/abs/2102.12122)
+
+
+
+## Abstract
+
+Although using convolutional neural networks (CNNs) as backbones achieves great successes in computer vision, this work investigates a simple backbone network useful for many dense prediction tasks without convolutions. Unlike the recently-proposed Transformer model (e.g., ViT) that is specially designed for image classification, we propose Pyramid Vision Transformer~(PVT), which overcomes the difficulties of porting Transformer to various dense prediction tasks. PVT has several merits compared to prior arts. (1) Different from ViT that typically has low-resolution outputs and high computational and memory cost, PVT can be not only trained on dense partitions of the image to achieve high output resolution, which is important for dense predictions but also using a progressive shrinking pyramid to reduce computations of large feature maps. (2) PVT inherits the advantages from both CNN and Transformer, making it a unified backbone in various vision tasks without convolutions by simply replacing CNN backbones. (3) We validate PVT by conducting extensive experiments, showing that it boosts the performance of many downstream tasks, e.g., object detection, semantic, and instance segmentation. For example, with a comparable number of parameters, RetinaNet+PVT achieves 40.4 AP on the COCO dataset, surpassing RetinNet+ResNet50 (36.3 AP) by 4.1 absolute AP. We hope PVT could serve as an alternative and useful backbone for pixel-level predictions and facilitate future researches.
+
+Transformer recently has shown encouraging progresses in computer vision. In this work, we present new baselines by improving the original Pyramid Vision Transformer (abbreviated as PVTv1) by adding three designs, including (1) overlapping patch embedding, (2) convolutional feed-forward networks, and (3) linear complexity attention layers.
+With these modifications, our PVTv2 significantly improves PVTv1 on three tasks e.g., classification, detection, and segmentation. Moreover, PVTv2 achieves comparable or better performances than recent works such as Swin Transformer. We hope this work will facilitate state-of-the-art Transformer researches in computer vision.
+
+
+

+
+
+## Results and Models
+
+### RetinaNet (PVTv1)
+
+| Backbone | Lr schd | Mem (GB) | box AP | Config | Download |
+| :--------: | :-----: | :------: | :----: | :----------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| PVT-Tiny | 12e | 8.5 | 36.6 | [config](./retinanet_pvt-t_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvt-t_fpn_1x_coco/retinanet_pvt-t_fpn_1x_coco_20210831_103110-17b566bd.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvt-t_fpn_1x_coco/retinanet_pvt-t_fpn_1x_coco_20210831_103110.log.json) |
+| PVT-Small | 12e | 14.5 | 40.4 | [config](./retinanet_pvt-s_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvt-s_fpn_1x_coco/retinanet_pvt-s_fpn_1x_coco_20210906_142921-b6c94a5b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvt-s_fpn_1x_coco/retinanet_pvt-s_fpn_1x_coco_20210906_142921.log.json) |
+| PVT-Medium | 12e | 20.9 | 41.7 | [config](./retinanet_pvt-m_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvt-m_fpn_1x_coco/retinanet_pvt-m_fpn_1x_coco_20210831_103243-55effa1b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvt-m_fpn_1x_coco/retinanet_pvt-m_fpn_1x_coco_20210831_103243.log.json) |
+
+### RetinaNet (PVTv2)
+
+| Backbone | Lr schd | Mem (GB) | box AP | Config | Download |
+| :------: | :-----: | :------: | :----: | :-------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| PVTv2-B0 | 12e | 7.4 | 37.1 | [config](./retinanet_pvtv2-b0_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvtv2-b0_fpn_1x_coco/retinanet_pvtv2-b0_fpn_1x_coco_20210831_103157-13e9aabe.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvtv2-b0_fpn_1x_coco/retinanet_pvtv2-b0_fpn_1x_coco_20210831_103157.log.json) |
+| PVTv2-B1 | 12e | 9.5 | 41.2 | [config](./retinanet_pvtv2-b1_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvtv2-b1_fpn_1x_coco/retinanet_pvtv2-b1_fpn_1x_coco_20210831_103318-7e169a7d.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvtv2-b1_fpn_1x_coco/retinanet_pvtv2-b1_fpn_1x_coco_20210831_103318.log.json) |
+| PVTv2-B2 | 12e | 16.2 | 44.6 | [config](./retinanet_pvtv2-b2_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvtv2-b2_fpn_1x_coco/retinanet_pvtv2-b2_fpn_1x_coco_20210901_174843-529f0b9a.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvtv2-b2_fpn_1x_coco/retinanet_pvtv2-b2_fpn_1x_coco_20210901_174843.log.json) |
+| PVTv2-B3 | 12e | 23.0 | 46.0 | [config](./retinanet_pvtv2-b3_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvtv2-b3_fpn_1x_coco/retinanet_pvtv2-b3_fpn_1x_coco_20210903_151512-8357deff.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvtv2-b3_fpn_1x_coco/retinanet_pvtv2-b3_fpn_1x_coco_20210903_151512.log.json) |
+| PVTv2-B4 | 12e | 17.0 | 46.3 | [config](./retinanet_pvtv2-b4_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvtv2-b4_fpn_1x_coco/retinanet_pvtv2-b4_fpn_1x_coco_20210901_170151-83795c86.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvtv2-b4_fpn_1x_coco/retinanet_pvtv2-b4_fpn_1x_coco_20210901_170151.log.json) |
+| PVTv2-B5 | 12e | 18.7 | 46.1 | [config](./retinanet_pvtv2-b5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvtv2-b5_fpn_1x_coco/retinanet_pvtv2-b5_fpn_1x_coco_20210902_201800-3420eb57.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvtv2-b5_fpn_1x_coco/retinanet_pvtv2-b5_fpn_1x_coco_20210902_201800.log.json) |
+
+## Citation
+
+```latex
+@article{wang2021pyramid,
+ title={Pyramid vision transformer: A versatile backbone for dense prediction without convolutions},
+ author={Wang, Wenhai and Xie, Enze and Li, Xiang and Fan, Deng-Ping and Song, Kaitao and Liang, Ding and Lu, Tong and Luo, Ping and Shao, Ling},
+ journal={arXiv preprint arXiv:2102.12122},
+ year={2021}
+}
+```
+
+```latex
+@article{wang2021pvtv2,
+ title={PVTv2: Improved Baselines with Pyramid Vision Transformer},
+ author={Wang, Wenhai and Xie, Enze and Li, Xiang and Fan, Deng-Ping and Song, Kaitao and Liang, Ding and Lu, Tong and Luo, Ping and Shao, Ling},
+ journal={arXiv preprint arXiv:2106.13797},
+ year={2021}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..58843784955f3f4be7aeebf7caa9b50b7891f4c5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/metafile.yml
@@ -0,0 +1,243 @@
+Models:
+ - Name: retinanet_pvt-t_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/pvt/retinanet_pvt-t_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 8.5
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x NVIDIA V100 GPUs
+ Architecture:
+ - PyramidVisionTransformer
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 36.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvt-t_fpn_1x_coco/retinanet_pvt-t_fpn_1x_coco_20210831_103110-17b566bd.pth
+ Paper:
+ URL: https://arxiv.org/abs/2102.12122
+ Title: "Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions"
+ README: configs/pvt/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.17.0/mmdet/models/backbones/pvt.py#L315
+ Version: 2.17.0
+
+ - Name: retinanet_pvt-s_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/pvt/retinanet_pvt-s_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 14.5
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x NVIDIA V100 GPUs
+ Architecture:
+ - PyramidVisionTransformer
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvt-s_fpn_1x_coco/retinanet_pvt-s_fpn_1x_coco_20210906_142921-b6c94a5b.pth
+ Paper:
+ URL: https://arxiv.org/abs/2102.12122
+ Title: "Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions"
+ README: configs/pvt/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.17.0/mmdet/models/backbones/pvt.py#L315
+ Version: 2.17.0
+
+ - Name: retinanet_pvt-m_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/pvt/retinanet_pvt-m_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 20.9
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x NVIDIA V100 GPUs
+ Architecture:
+ - PyramidVisionTransformer
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvt-m_fpn_1x_coco/retinanet_pvt-m_fpn_1x_coco_20210831_103243-55effa1b.pth
+ Paper:
+ URL: https://arxiv.org/abs/2102.12122
+ Title: "Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions"
+ README: configs/pvt/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.17.0/mmdet/models/backbones/pvt.py#L315
+ Version: 2.17.0
+
+ - Name: retinanet_pvtv2-b0_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/pvt/retinanet_pvtv2-b0_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.4
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x NVIDIA V100 GPUs
+ Architecture:
+ - PyramidVisionTransformerV2
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvtv2-b0_fpn_1x_coco/retinanet_pvtv2-b0_fpn_1x_coco_20210831_103157-13e9aabe.pth
+ Paper:
+ URL: https://arxiv.org/abs/2106.13797
+ Title: "PVTv2: Improved Baselines with Pyramid Vision Transformer"
+ README: configs/pvt/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.17.0/mmdet/models/backbones/pvt.py#L543
+ Version: 2.17.0
+
+ - Name: retinanet_pvtv2-b1_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/pvt/retinanet_pvtv2-b1_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 9.5
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x NVIDIA V100 GPUs
+ Architecture:
+ - PyramidVisionTransformerV2
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvtv2-b1_fpn_1x_coco/retinanet_pvtv2-b1_fpn_1x_coco_20210831_103318-7e169a7d.pth
+ Paper:
+ URL: https://arxiv.org/abs/2106.13797
+ Title: "PVTv2: Improved Baselines with Pyramid Vision Transformer"
+ README: configs/pvt/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.17.0/mmdet/models/backbones/pvt.py#L543
+ Version: 2.17.0
+
+ - Name: retinanet_pvtv2-b2_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/pvt/retinanet_pvtv2-b2_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 16.2
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x NVIDIA V100 GPUs
+ Architecture:
+ - PyramidVisionTransformerV2
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvtv2-b2_fpn_1x_coco/retinanet_pvtv2-b2_fpn_1x_coco_20210901_174843-529f0b9a.pth
+ Paper:
+ URL: https://arxiv.org/abs/2106.13797
+ Title: "PVTv2: Improved Baselines with Pyramid Vision Transformer"
+ README: configs/pvt/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.17.0/mmdet/models/backbones/pvt.py#L543
+ Version: 2.17.0
+
+ - Name: retinanet_pvtv2-b3_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/pvt/retinanet_pvtv2-b3_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 23.0
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x NVIDIA V100 GPUs
+ Architecture:
+ - PyramidVisionTransformerV2
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvtv2-b3_fpn_1x_coco/retinanet_pvtv2-b3_fpn_1x_coco_20210903_151512-8357deff.pth
+ Paper:
+ URL: https://arxiv.org/abs/2106.13797
+ Title: "PVTv2: Improved Baselines with Pyramid Vision Transformer"
+ README: configs/pvt/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.17.0/mmdet/models/backbones/pvt.py#L543
+ Version: 2.17.0
+
+ - Name: retinanet_pvtv2-b4_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/pvt/retinanet_pvtv2-b4_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 17.0
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x NVIDIA V100 GPUs
+ Architecture:
+ - PyramidVisionTransformerV2
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvtv2-b4_fpn_1x_coco/retinanet_pvtv2-b4_fpn_1x_coco_20210901_170151-83795c86.pth
+ Paper:
+ URL: https://arxiv.org/abs/2106.13797
+ Title: "PVTv2: Improved Baselines with Pyramid Vision Transformer"
+ README: configs/pvt/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.17.0/mmdet/models/backbones/pvt.py#L543
+ Version: 2.17.0
+
+ - Name: retinanet_pvtv2-b5_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/pvt/retinanet_pvtv2-b5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 18.7
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x NVIDIA V100 GPUs
+ Architecture:
+ - PyramidVisionTransformerV2
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/pvt/retinanet_pvtv2-b5_fpn_1x_coco/retinanet_pvtv2-b5_fpn_1x_coco_20210902_201800-3420eb57.pth
+ Paper:
+ URL: https://arxiv.org/abs/2106.13797
+ Title: "PVTv2: Improved Baselines with Pyramid Vision Transformer"
+ README: configs/pvt/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.17.0/mmdet/models/backbones/pvt.py#L543
+ Version: 2.17.0
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvt-l_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvt-l_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1a6f604bdb367106bc75680808ce6fabc2740ed1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvt-l_fpn_1x_coco.py
@@ -0,0 +1,8 @@
+_base_ = 'retinanet_pvt-t_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ num_layers=[3, 8, 27, 3],
+ init_cfg=dict(checkpoint='https://github.com/whai362/PVT/'
+ 'releases/download/v2/pvt_large.pth')))
+# Enable automatic-mixed-precision training with AmpOptimWrapper.
+optim_wrapper = dict(type='AmpOptimWrapper')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvt-m_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvt-m_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b888f788b6c7310491751774238451bb7107dccc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvt-m_fpn_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = 'retinanet_pvt-t_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ num_layers=[3, 4, 18, 3],
+ init_cfg=dict(checkpoint='https://github.com/whai362/PVT/'
+ 'releases/download/v2/pvt_medium.pth')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvt-s_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvt-s_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..46603488bb3ceb4fc1052139da53340a3d595256
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvt-s_fpn_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = 'retinanet_pvt-t_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ num_layers=[3, 4, 6, 3],
+ init_cfg=dict(checkpoint='https://github.com/whai362/PVT/'
+ 'releases/download/v2/pvt_small.pth')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvt-t_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvt-t_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5f67c444f262613d615b8b7331991ca7e2f57935
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvt-t_fpn_1x_coco.py
@@ -0,0 +1,18 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ type='RetinaNet',
+ backbone=dict(
+ _delete_=True,
+ type='PyramidVisionTransformer',
+ num_layers=[2, 2, 2, 2],
+ init_cfg=dict(checkpoint='https://github.com/whai362/PVT/'
+ 'releases/download/v2/pvt_tiny.pth')),
+ neck=dict(in_channels=[64, 128, 320, 512]))
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(
+ _delete_=True, type='AdamW', lr=0.0001, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvtv2-b0_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvtv2-b0_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..cbebf90fb89d81bd2f4c0874dc2c82cf7c7393d0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvtv2-b0_fpn_1x_coco.py
@@ -0,0 +1,19 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ type='RetinaNet',
+ backbone=dict(
+ _delete_=True,
+ type='PyramidVisionTransformerV2',
+ embed_dims=32,
+ num_layers=[2, 2, 2, 2],
+ init_cfg=dict(checkpoint='https://github.com/whai362/PVT/'
+ 'releases/download/v2/pvt_v2_b0.pth')),
+ neck=dict(in_channels=[32, 64, 160, 256]))
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(
+ _delete_=True, type='AdamW', lr=0.0001, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvtv2-b1_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvtv2-b1_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5374c50925f5c7ed8a761eda40dc4bf374df3aeb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvtv2-b1_fpn_1x_coco.py
@@ -0,0 +1,7 @@
+_base_ = 'retinanet_pvtv2-b0_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ embed_dims=64,
+ init_cfg=dict(checkpoint='https://github.com/whai362/PVT/'
+ 'releases/download/v2/pvt_v2_b1.pth')),
+ neck=dict(in_channels=[64, 128, 320, 512]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvtv2-b2_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvtv2-b2_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..cf9a18debbe5f8b9918e0d086ad6d54d203ef310
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvtv2-b2_fpn_1x_coco.py
@@ -0,0 +1,8 @@
+_base_ = 'retinanet_pvtv2-b0_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ embed_dims=64,
+ num_layers=[3, 4, 6, 3],
+ init_cfg=dict(checkpoint='https://github.com/whai362/PVT/'
+ 'releases/download/v2/pvt_v2_b2.pth')),
+ neck=dict(in_channels=[64, 128, 320, 512]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvtv2-b3_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvtv2-b3_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..7a47f820324af7fecf773640d7d1829b0c115471
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvtv2-b3_fpn_1x_coco.py
@@ -0,0 +1,8 @@
+_base_ = 'retinanet_pvtv2-b0_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ embed_dims=64,
+ num_layers=[3, 4, 18, 3],
+ init_cfg=dict(checkpoint='https://github.com/whai362/PVT/'
+ 'releases/download/v2/pvt_v2_b3.pth')),
+ neck=dict(in_channels=[64, 128, 320, 512]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvtv2-b4_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvtv2-b4_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5faf4c507ba89ffe614b2b9d34d452e4c106b0fe
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvtv2-b4_fpn_1x_coco.py
@@ -0,0 +1,20 @@
+_base_ = 'retinanet_pvtv2-b0_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ embed_dims=64,
+ num_layers=[3, 8, 27, 3],
+ init_cfg=dict(checkpoint='https://github.com/whai362/PVT/'
+ 'releases/download/v2/pvt_v2_b4.pth')),
+ neck=dict(in_channels=[64, 128, 320, 512]))
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(
+ _delete_=True, type='AdamW', lr=0.0001 / 1.4, weight_decay=0.0001))
+
+# dataset settings
+train_dataloader = dict(batch_size=1, num_workers=1)
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (1 samples per GPU)
+auto_scale_lr = dict(base_batch_size=8)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvtv2-b5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvtv2-b5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..afff8719ece41dbfbbe23e2259b9973bb29871f6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/pvt/retinanet_pvtv2-b5_fpn_1x_coco.py
@@ -0,0 +1,21 @@
+_base_ = 'retinanet_pvtv2-b0_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ embed_dims=64,
+ num_layers=[3, 6, 40, 3],
+ mlp_ratios=(4, 4, 4, 4),
+ init_cfg=dict(checkpoint='https://github.com/whai362/PVT/'
+ 'releases/download/v2/pvt_v2_b5.pth')),
+ neck=dict(in_channels=[64, 128, 320, 512]))
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(
+ _delete_=True, type='AdamW', lr=0.0001 / 1.4, weight_decay=0.0001))
+
+# dataset settings
+train_dataloader = dict(batch_size=1, num_workers=1)
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (1 samples per GPU)
+auto_scale_lr = dict(base_batch_size=8)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/qdtrack/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/qdtrack/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..5a6efe7d3fd62d61b328d8ce248e1fd9132f5792
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/qdtrack/README.md
@@ -0,0 +1,89 @@
+# Quasi-Dense Similarity Learning for Multiple Object Tracking
+
+## Abstract
+
+
+
+Similarity learning has been recognized as a crucial step for object tracking. However, existing multiple object tracking methods only use sparse ground truth matching as the training objective, while ignoring the majority of the informative regions on the images. In this paper, we present Quasi-Dense Similarity Learning, which densely samples hundreds of region proposals on a pair of images for contrastive learning. We can directly combine this similarity learning with existing detection methods to build Quasi-Dense Tracking (QDTrack) without turning to displacementregression or motion priors. We also find that the resulting distinctive feature space admits a simple nearest neighbor search at the inference time. Despite its simplicity, QD-Track outperforms all existing methods on MOT, BDD100K, Waymo, and TAO tracking benchmarks. It achieves 68.7 MOTA at 20.3 FPS on MOT17 without using external training data. Compared to methods with similar detectors, it boosts almost 10 points of MOTA and significantly decreases the number of ID switches on BDD100K and Waymo datasets.
+
+
+
+
+

+

+
+
+## Results and models on MOT17
+
+| Method | Detector | Train Set | Test Set | Public | Inf time (fps) | HOTA | MOTA | IDF1 | FP | FN | IDSw. | Config | Download |
+| :-----: | :----------: | :--------: | :------: | :----: | :------------: | :--: | :--: | :--: | :--: | :---: | :---: | :-------------------------------------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| QDTrack | Faster R-CNN | half-train | half-val | N | - | 57.1 | 68.1 | 68.6 | 7707 | 42732 | 1083 | [config](qdtrack_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py) | [model](https://download.openmmlab.com/mmtracking/mot/qdtrack/mot_dataset/qdtrack_faster-rcnn_r50_fpn_4e_mot17_20220315_145635-76f295ef.pth) \| [log](https://download.openmmlab.com/mmtracking/mot/qdtrack/mot_dataset/qdtrack_faster-rcnn_r50_fpn_4e_mot17_20220315_145635.log.json) |
+
+## Get started
+
+### 1. Development Environment Setup
+
+Tracking Development Environment Setup can refer to this [document](../../docs/en/get_started.md).
+
+### 2. Dataset Prepare
+
+Tracking Dataset Prepare can refer to this [document](../../docs/en/user_guides/tracking_dataset_prepare.md).
+
+### 3. Training
+
+Due to the influence of parameters such as learning rate in default configuration file, we recommend using 8 GPUs for training in order to reproduce accuracy. You can use the following command to start the training.
+
+```shell
+# Training QDTrack on mot17-half-train dataset with following command.
+# The number after config file represents the number of GPUs used. Here we use 8 GPUs.
+bash tools/dist_train.sh configs/qdtrack/qdtrack_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py 8
+```
+
+If you want to know about more detailed usage of `train.py/dist_train.sh/slurm_train.sh`,
+please refer to this [document](../../docs/en/user_guides/tracking_train_test.md).
+
+### 4. Testing and evaluation
+
+**4.1 Example on MOTxx-halfval dataset**
+
+```shell
+# Example 1: Test on motXX-half-val set
+# The number after config file represents the number of GPUs used. Here we use 8 GPUs.
+bash tools/dist_test_tracking.sh configs/qdtrack/qdtrack_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py 8 --checkpoint ${CHECKPOINT_PATH}
+```
+
+**4.2 use video_baesd to evaluating and testing**
+we also provide two_ways(img_based or video_based) to evaluating and testing.
+if you want to use video_based to evaluating and testing, you can modify config as follows
+
+```
+val_dataloader = dict(
+ sampler=dict(type='DefaultSampler', shuffle=False, round_up=False))
+```
+
+If you want to know about more detailed usage of `test_tracking.py/dist_test_tracking.sh/slurm_test_tracking.sh`,
+please refer to this [document](../../docs/en/user_guides/tracking_train_test.md).
+
+### 5.Inference
+
+Use a single GPU to predict a video and save it as a video.
+
+```shell
+python demo/mot_demo.py demo/demo_mot.mp4 configs/qdtrack/qdtrack_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py --checkpoint ${CHECKPOINT_PATH} --out mot.mp4
+```
+
+If you want to know about more detailed usage of `mot_demo.py`, please refer to this [document](../../docs/en/user_guides/tracking_inference.md).
+
+## Citation
+
+
+
+```latex
+@inproceedings{pang2021quasi,
+ title={Quasi-dense similarity learning for multiple object tracking},
+ author={Pang, Jiangmiao and Qiu, Linlu and Li, Xia and Chen, Haofeng and Li, Qi and Darrell, Trevor and Yu, Fisher},
+ booktitle={Proceedings of the IEEE/CVF conference on computer vision and pattern recognition},
+ pages={164--173},
+ year={2021}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/qdtrack/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/qdtrack/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..e5c5504d1bd00e43bdba7f28efcbf9dd23555342
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/qdtrack/metafile.yml
@@ -0,0 +1,30 @@
+Collections:
+ - Name: QDTrack
+ Metadata:
+ Training Data: MOT17, crowdhuman
+ Training Techniques:
+ - SGD
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/pdf/2006.06664.pdf
+ Title: Quasi-Dense Similarity Learning for Multiple Object Tracking
+ README: configs/qdtrack/README.md
+
+Models:
+ - Name: qdtrack_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval
+ In Collection: QDTrack
+ Config: configs/qdtrack/qdtrack_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py
+ Metadata:
+ Training Data: MOT17
+ Training Memory (GB): 5.83
+ Epochs: 4
+ Results:
+ - Task: Multi-object Tracking
+ Dataset: MOT17
+ Metrics:
+ HOTA: 57.1
+ MOTA: 68.1
+ IDF1: 68.6
+ Weights: https://download.openmmlab.com/mmtracking/mot/qdtrack/mot_dataset/qdtrack_faster-rcnn_r50_fpn_4e_mot17_20220315_145635-76f295ef.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/qdtrack/qdtrack_faster-rcnn_r50_fpn_4e_base.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/qdtrack/qdtrack_faster-rcnn_r50_fpn_4e_base.py
new file mode 100644
index 0000000000000000000000000000000000000000..e3c17c3eb97eedef88949c841364b858a3a1d6e9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/qdtrack/qdtrack_faster-rcnn_r50_fpn_4e_base.py
@@ -0,0 +1,118 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py', '../_base_/default_runtime.py'
+]
+
+detector = _base_.model
+detector.pop('data_preprocessor')
+
+detector['backbone'].update(
+ dict(
+ norm_cfg=dict(type='BN', requires_grad=False),
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')))
+detector.rpn_head.loss_bbox.update(
+ dict(type='SmoothL1Loss', beta=1.0 / 9.0, loss_weight=1.0))
+detector.rpn_head.bbox_coder.update(dict(clip_border=False))
+detector.roi_head.bbox_head.update(dict(num_classes=1))
+detector.roi_head.bbox_head.bbox_coder.update(dict(clip_border=False))
+detector['init_cfg'] = dict(
+ type='Pretrained',
+ checkpoint= # noqa: E251
+ 'https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/'
+ 'faster_rcnn_r50_fpn_1x_coco-person/'
+ 'faster_rcnn_r50_fpn_1x_coco-person_20201216_175929-d022e227.pth'
+ # noqa: E501
+)
+del _base_.model
+
+model = dict(
+ type='QDTrack',
+ data_preprocessor=dict(
+ type='TrackDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ detector=detector,
+ track_head=dict(
+ type='QuasiDenseTrackHead',
+ roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=7, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ embed_head=dict(
+ type='QuasiDenseEmbedHead',
+ num_convs=4,
+ num_fcs=1,
+ embed_channels=256,
+ norm_cfg=dict(type='GN', num_groups=32),
+ loss_track=dict(type='MultiPosCrossEntropyLoss', loss_weight=0.25),
+ loss_track_aux=dict(
+ type='MarginL2Loss',
+ neg_pos_ub=3,
+ pos_margin=0,
+ neg_margin=0.1,
+ hard_mining=True,
+ loss_weight=1.0)),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0),
+ train_cfg=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='CombinedSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=3,
+ add_gt_as_proposals=True,
+ pos_sampler=dict(type='InstanceBalancedPosSampler'),
+ neg_sampler=dict(type='RandomSampler')))),
+ tracker=dict(
+ type='QuasiDenseTracker',
+ init_score_thr=0.9,
+ obj_score_thr=0.5,
+ match_score_thr=0.5,
+ memo_tracklet_frames=30,
+ memo_backdrop_frames=1,
+ memo_momentum=0.8,
+ nms_conf_thr=0.5,
+ nms_backdrop_iou_thr=0.3,
+ nms_class_iou_thr=0.7,
+ with_cats=True,
+ match_metric='bisoftmax'))
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.0001),
+ clip_grad=dict(max_norm=35, norm_type=2))
+# learning policy
+param_scheduler = [
+ dict(type='MultiStepLR', begin=0, end=4, by_epoch=True, milestones=[3])
+]
+
+# runtime settings
+train_cfg = dict(type='EpochBasedTrainLoop', max_epochs=4, val_interval=4)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+default_hooks = dict(
+ logger=dict(type='LoggerHook', interval=50),
+ visualization=dict(type='TrackVisualizationHook', draw=False))
+
+vis_backends = [dict(type='LocalVisBackend')]
+visualizer = dict(
+ type='TrackLocalVisualizer', vis_backends=vis_backends, name='visualizer')
+
+# custom hooks
+custom_hooks = [
+ # Synchronize model buffers such as running_mean and running_var in BN
+ # at the end of each epoch
+ dict(type='SyncBuffersHook')
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/qdtrack/qdtrack_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/qdtrack/qdtrack_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py
new file mode 100644
index 0000000000000000000000000000000000000000..d87604dad6bf39028a8111708307482186118b19
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/qdtrack/qdtrack_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py
@@ -0,0 +1,14 @@
+_base_ = [
+ './qdtrack_faster-rcnn_r50_fpn_4e_base.py',
+ '../_base_/datasets/mot_challenge.py',
+]
+
+# evaluator
+val_evaluator = [
+ dict(type='CocoVideoMetric', metric=['bbox'], classwise=True),
+ dict(type='MOTChallengeMetric', metric=['HOTA', 'CLEAR', 'Identity'])
+]
+
+test_evaluator = val_evaluator
+# The fluctuation of HOTA is about +-1.
+randomness = dict(seed=6)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..ee62ccbf8a3b77a6ddbb62c8ba3740bc509d8ae8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/README.md
@@ -0,0 +1,36 @@
+# QueryInst
+
+> [Instances as Queries](https://openaccess.thecvf.com/content/ICCV2021/html/Fang_Instances_As_Queries_ICCV_2021_paper.html)
+
+
+
+## Abstract
+
+We present QueryInst, a new perspective for instance segmentation. QueryInst is a multi-stage end-to-end system that treats instances of interest as learnable queries, enabling query based object detectors, e.g., Sparse R-CNN, to have strong instance segmentation performance. The attributes of instances such as categories, bounding boxes, instance masks, and instance association embeddings are represented by queries in a unified manner. In QueryInst, a query is shared by both detection and segmentation via dynamic convolutions and driven by parallelly-supervised multi-stage learning. We conduct extensive experiments on three challenging benchmarks, i.e., COCO, CityScapes, and YouTube-VIS to evaluate the effectiveness of QueryInst in object detection, instance segmentation, and video instance segmentation tasks. For the first time, we demonstrate that a simple end-to-end query based framework can achieve the state-of-the-art performance in various instance-level recognition tasks.
+
+
+

+
+
+## Results and Models
+
+| Model | Backbone | Style | Lr schd | Number of Proposals | Multi-Scale | RandomCrop | box AP | mask AP | Config | Download |
+| :-------: | :-------: | :-----: | :-----: | :-----------------: | :---------: | :--------: | :----: | :-----: | :---------------------------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| QueryInst | R-50-FPN | pytorch | 1x | 100 | False | False | 42.0 | 37.5 | [config](./queryinst_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/queryinst/queryinst_r50_fpn_1x_coco/queryinst_r50_fpn_1x_coco_20210907_084916-5a8f1998.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/queryinst/queryinst_r50_fpn_1x_coco/queryinst_r50_fpn_1x_coco_20210907_084916.log.json) |
+| QueryInst | R-50-FPN | pytorch | 3x | 100 | True | False | 44.8 | 39.8 | [config](./queryinst_r50_fpn_ms-480-800-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/queryinst/queryinst_r50_fpn_mstrain_480-800_3x_coco/queryinst_r50_fpn_mstrain_480-800_3x_coco_20210901_103643-7837af86.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/queryinst/queryinst_r50_fpn_mstrain_480-800_3x_coco/queryinst_r50_fpn_mstrain_480-800_3x_coco_20210901_103643.log.json) |
+| QueryInst | R-50-FPN | pytorch | 3x | 300 | True | True | 47.5 | 41.7 | [config](./queryinst_r50_fpn_300-proposals_crop-ms-480-800-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/queryinst/queryinst_r50_fpn_300_proposals_crop_mstrain_480-800_3x_coco/queryinst_r50_fpn_300_proposals_crop_mstrain_480-800_3x_coco_20210904_101802-85cffbd8.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/queryinst/queryinst_r50_fpn_300_proposals_crop_mstrain_480-800_3x_coco/queryinst_r50_fpn_300_proposals_crop_mstrain_480-800_3x_coco_20210904_101802.log.json) |
+| QueryInst | R-101-FPN | pytorch | 3x | 100 | True | False | 46.4 | 41.0 | [config](./queryinst_r101_fpn_ms-480-800-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/queryinst/queryinst_r101_fpn_mstrain_480-800_3x_coco/queryinst_r101_fpn_mstrain_480-800_3x_coco_20210904_104048-91f9995b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/queryinst/queryinst_r101_fpn_mstrain_480-800_3x_coco/queryinst_r101_fpn_mstrain_480-800_3x_coco_20210904_104048.log.json) |
+| QueryInst | R-101-FPN | pytorch | 3x | 300 | True | True | 49.0 | 42.9 | [config](./queryinst_r101_fpn_300-proposals_crop-ms-480-800-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/queryinst/queryinst_r101_fpn_300_proposals_crop_mstrain_480-800_3x_coco/queryinst_r101_fpn_300_proposals_crop_mstrain_480-800_3x_coco_20210904_153621-76cce59f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/queryinst/queryinst_r101_fpn_300_proposals_crop_mstrain_480-800_3x_coco/queryinst_r101_fpn_300_proposals_crop_mstrain_480-800_3x_coco_20210904_153621.log.json) |
+
+## Citation
+
+```latex
+@InProceedings{Fang_2021_ICCV,
+ author = {Fang, Yuxin and Yang, Shusheng and Wang, Xinggang and Li, Yu and Fang, Chen and Shan, Ying and Feng, Bin and Liu, Wenyu},
+ title = {Instances As Queries},
+ booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
+ month = {October},
+ year = {2021},
+ pages = {6910-6919}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..3ea3b00a945c8856b8c63f68a0ec6a48c70a933f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/metafile.yml
@@ -0,0 +1,100 @@
+Collections:
+ - Name: QueryInst
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - AdamW
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - FPN
+ - ResNet
+ - QueryInst
+ Paper:
+ URL: https://openaccess.thecvf.com/content/ICCV2021/papers/Fang_Instances_As_Queries_ICCV_2021_paper.pdf
+ Title: 'Instances as Queries'
+ README: configs/queryinst/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/main/mmdet/models/detectors/queryinst.py
+ Version: v2.18.0
+
+Models:
+ - Name: queryinst_r50_fpn_1x_coco
+ In Collection: QueryInst
+ Config: configs/queryinst/queryinst_r50_fpn_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/queryinst/queryinst_r50_fpn_1x_coco/queryinst_r50_fpn_1x_coco_20210907_084916-5a8f1998.pth
+
+ - Name: queryinst_r50_fpn_ms-480-800-3x_coco
+ In Collection: QueryInst
+ Config: configs/queryinst/queryinst_r50_fpn_ms-480-800-3x_coco.py
+ Metadata:
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/queryinst/queryinst_r50_fpn_mstrain_480-800_3x_coco/queryinst_r50_fpn_mstrain_480-800_3x_coco_20210901_103643-7837af86.pth
+
+ - Name: queryinst_r50_fpn_300-proposals_crop-ms-480-800-3x_coco
+ In Collection: QueryInst
+ Config: configs/queryinst/queryinst_r50_fpn_300-proposals_crop-ms-480-800-3x_coco.py
+ Metadata:
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 47.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 41.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/queryinst/queryinst_r50_fpn_300_proposals_crop_mstrain_480-800_3x_coco/queryinst_r50_fpn_300_proposals_crop_mstrain_480-800_3x_coco_20210904_101802-85cffbd8.pth
+
+ - Name: queryinst_r101_fpn_ms-480-800-3x_coco
+ In Collection: QueryInst
+ Config: configs/queryinst/queryinst_r101_fpn_ms-480-800-3x_coco.py
+ Metadata:
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 41.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/queryinst/queryinst_r101_fpn_mstrain_480-800_3x_coco/queryinst_r101_fpn_mstrain_480-800_3x_coco_20210904_104048-91f9995b.pth
+
+ - Name: queryinst_r101_fpn_300-proposals_crop-ms-480-800-3x_coco
+ In Collection: QueryInst
+ Config: configs/queryinst/queryinst_r101_fpn_300-proposals_crop-ms-480-800-3x_coco.py
+ Metadata:
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 49.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 42.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/queryinst/queryinst_r101_fpn_300_proposals_crop_mstrain_480-800_3x_coco/queryinst_r101_fpn_300_proposals_crop_mstrain_480-800_3x_coco_20210904_153621-76cce59f.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/queryinst_r101_fpn_300-proposals_crop-ms-480-800-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/queryinst_r101_fpn_300-proposals_crop-ms-480-800-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1692c134698a98da33612487a9fb703117fdb8b6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/queryinst_r101_fpn_300-proposals_crop-ms-480-800-3x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './queryinst_r50_fpn_300-proposals_crop-ms-480-800-3x_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/queryinst_r101_fpn_ms-480-800-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/queryinst_r101_fpn_ms-480-800-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..dd5b7f452e583eb362e0bb05f272a771d68b6e48
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/queryinst_r101_fpn_ms-480-800-3x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './queryinst_r50_fpn_ms-480-800-3x_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/queryinst_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/queryinst_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..63d61d78872b452bdd8d2607fc03181b169ea845
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/queryinst_r50_fpn_1x_coco.py
@@ -0,0 +1,155 @@
+_base_ = [
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+num_stages = 6
+num_proposals = 100
+model = dict(
+ type='QueryInst',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_mask=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=0,
+ add_extra_convs='on_input',
+ num_outs=4),
+ rpn_head=dict(
+ type='EmbeddingRPNHead',
+ num_proposals=num_proposals,
+ proposal_feature_channel=256),
+ roi_head=dict(
+ type='SparseRoIHead',
+ num_stages=num_stages,
+ stage_loss_weights=[1] * num_stages,
+ proposal_feature_channel=256,
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=7, sampling_ratio=2),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ mask_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=14, sampling_ratio=2),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ bbox_head=[
+ dict(
+ type='DIIHead',
+ num_classes=80,
+ num_ffn_fcs=2,
+ num_heads=8,
+ num_cls_fcs=1,
+ num_reg_fcs=3,
+ feedforward_channels=2048,
+ in_channels=256,
+ dropout=0.0,
+ ffn_act_cfg=dict(type='ReLU', inplace=True),
+ dynamic_conv_cfg=dict(
+ type='DynamicConv',
+ in_channels=256,
+ feat_channels=64,
+ out_channels=256,
+ input_feat_shape=7,
+ act_cfg=dict(type='ReLU', inplace=True),
+ norm_cfg=dict(type='LN')),
+ loss_bbox=dict(type='L1Loss', loss_weight=5.0),
+ loss_iou=dict(type='GIoULoss', loss_weight=2.0),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=2.0),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ clip_border=False,
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.5, 0.5, 1., 1.])) for _ in range(num_stages)
+ ],
+ mask_head=[
+ dict(
+ type='DynamicMaskHead',
+ dynamic_conv_cfg=dict(
+ type='DynamicConv',
+ in_channels=256,
+ feat_channels=64,
+ out_channels=256,
+ input_feat_shape=14,
+ with_proj=False,
+ act_cfg=dict(type='ReLU', inplace=True),
+ norm_cfg=dict(type='LN')),
+ num_convs=4,
+ num_classes=80,
+ roi_feat_size=14,
+ in_channels=256,
+ conv_kernel_size=3,
+ conv_out_channels=256,
+ class_agnostic=False,
+ norm_cfg=dict(type='BN'),
+ upsample_cfg=dict(type='deconv', scale_factor=2),
+ loss_mask=dict(
+ type='DiceLoss',
+ loss_weight=8.0,
+ use_sigmoid=True,
+ activate=False,
+ eps=1e-5)) for _ in range(num_stages)
+ ]),
+ # training and testing settings
+ train_cfg=dict(
+ rpn=None,
+ rcnn=[
+ dict(
+ assigner=dict(
+ type='HungarianAssigner',
+ match_costs=[
+ dict(type='FocalLossCost', weight=2.0),
+ dict(type='BBoxL1Cost', weight=5.0, box_format='xyxy'),
+ dict(type='IoUCost', iou_mode='giou', weight=2.0)
+ ]),
+ sampler=dict(type='PseudoSampler'),
+ pos_weight=1,
+ mask_size=28,
+ ) for _ in range(num_stages)
+ ]),
+ test_cfg=dict(
+ rpn=None, rcnn=dict(max_per_img=num_proposals, mask_thr_binary=0.5)))
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(
+ _delete_=True, type='AdamW', lr=0.0001, weight_decay=0.0001),
+ paramwise_cfg=dict(
+ custom_keys={'backbone': dict(lr_mult=0.1, decay_mult=1.0)}),
+ clip_grad=dict(max_norm=0.1, norm_type=2))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0,
+ end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/queryinst_r50_fpn_300-proposals_crop-ms-480-800-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/queryinst_r50_fpn_300-proposals_crop-ms-480-800-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..33ab061267bc9753f490acc57ed8d4193f1250b4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/queryinst_r50_fpn_300-proposals_crop-ms-480-800-3x_coco.py
@@ -0,0 +1,45 @@
+_base_ = './queryinst_r50_fpn_ms-480-800-3x_coco.py'
+num_proposals = 300
+model = dict(
+ rpn_head=dict(num_proposals=num_proposals),
+ test_cfg=dict(
+ _delete_=True,
+ rpn=None,
+ rcnn=dict(max_per_img=num_proposals, mask_thr_binary=0.5)))
+
+# augmentation strategy originates from DETR.
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[[
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(400, 1333), (500, 1333), (600, 1333)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333),
+ (576, 1333), (608, 1333), (640, 1333),
+ (672, 1333), (704, 1333), (736, 1333),
+ (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]]),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/queryinst_r50_fpn_ms-480-800-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/queryinst_r50_fpn_ms-480-800-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6b99374ef4364dc76a60c2dd74377f92c15780ed
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/queryinst/queryinst_r50_fpn_ms-480-800-3x_coco.py
@@ -0,0 +1,32 @@
+_base_ = './queryinst_r50_fpn_1x_coco.py'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+# learning policy
+max_epochs = 36
+train_cfg = dict(type='EpochBasedTrainLoop', max_epochs=max_epochs)
+
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[27, 33],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..0bfcec1891ccb468bcccf975b9bd26bca53e0a7f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/README.md
@@ -0,0 +1,121 @@
+# RegNet
+
+> [Designing Network Design Spaces](https://arxiv.org/abs/2003.13678)
+
+
+
+## Abstract
+
+In this work, we present a new network design paradigm. Our goal is to help advance the understanding of network design and discover design principles that generalize across settings. Instead of focusing on designing individual network instances, we design network design spaces that parametrize populations of networks. The overall process is analogous to classic manual design of networks, but elevated to the design space level. Using our methodology we explore the structure aspect of network design and arrive at a low-dimensional design space consisting of simple, regular networks that we call RegNet. The core insight of the RegNet parametrization is surprisingly simple: widths and depths of good networks can be explained by a quantized linear function. We analyze the RegNet design space and arrive at interesting findings that do not match the current practice of network design. The RegNet design space provides simple and fast networks that work well across a wide range of flop regimes. Under comparable training settings and flops, the RegNet models outperform the popular EfficientNet models while being up to 5x faster on GPUs.
+
+
+

+
+
+## Introduction
+
+We implement RegNetX and RegNetY models in detection systems and provide their first results on Mask R-CNN, Faster R-CNN and RetinaNet.
+
+The pre-trained models are converted from [model zoo of pycls](https://github.com/facebookresearch/pycls/blob/master/MODEL_ZOO.md).
+
+## Usage
+
+To use a regnet model, there are two steps to do:
+
+1. Convert the model to ResNet-style supported by MMDetection
+2. Modify backbone and neck in config accordingly
+
+### Convert model
+
+We already prepare models of FLOPs from 400M to 12G in our model zoo.
+
+For more general usage, we also provide script `regnet2mmdet.py` in the tools directory to convert the key of models pretrained by [pycls](https://github.com/facebookresearch/pycls/) to
+ResNet-style checkpoints used in MMDetection.
+
+```bash
+python -u tools/model_converters/regnet2mmdet.py ${PRETRAIN_PATH} ${STORE_PATH}
+```
+
+This script convert model from `PRETRAIN_PATH` and store the converted model in `STORE_PATH`.
+
+### Modify config
+
+The users can modify the config's `depth` of backbone and corresponding keys in `arch` according to the configs in the [pycls model zoo](https://github.com/facebookresearch/pycls/blob/master/MODEL_ZOO.md).
+The parameter `in_channels` in FPN can be found in the Figure 15 & 16 of the paper (`wi` in the legend).
+This directory already provides some configs with their performance, using RegNetX from 800MF to 12GF level.
+For other pre-trained models or self-implemented regnet models, the users are responsible to check these parameters by themselves.
+
+**Note**: Although Fig. 15 & 16 also provide `w0`, `wa`, `wm`, `group_w`, and `bot_mul` for `arch`, they are quantized thus inaccurate, using them sometimes produces different backbone that does not match the key in the pre-trained model.
+
+## Results and Models
+
+### Mask R-CNN
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :----------------------------------------------------------------------------------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :-------------------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| [R-50-FPN](../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py) | pytorch | 1x | 4.4 | 12.0 | 38.2 | 34.7 | [config](../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_fpn_1x_coco/mask_rcnn_r50_fpn_1x_coco_20200205-d4b0c5d6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r50_fpn_1x_coco/mask_rcnn_r50_fpn_1x_coco_20200205_050542.log.json) |
+| [RegNetX-3.2GF-FPN](./mask-rcnn_regnetx-3.2GF_fpn_1x_coco.py) | pytorch | 1x | 5.0 | | 40.3 | 36.6 | [config](./mask-rcnn_regnetx-3.2GF_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-3.2GF_fpn_1x_coco/mask_rcnn_regnetx-3.2GF_fpn_1x_coco_20200520_163141-2a9d1814.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-3.2GF_fpn_1x_coco/mask_rcnn_regnetx-3.2GF_fpn_1x_coco_20200520_163141.log.json) |
+| [RegNetX-4.0GF-FPN](./mask-rcnn_regnetx-4GF_fpn_1x_coco.py) | pytorch | 1x | 5.5 | | 41.5 | 37.4 | [config](./mask-rcnn_regnetx-4GF_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-4GF_fpn_1x_coco/mask_rcnn_regnetx-4GF_fpn_1x_coco_20200517_180217-32e9c92d.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-4GF_fpn_1x_coco/mask_rcnn_regnetx-4GF_fpn_1x_coco_20200517_180217.log.json) |
+| [R-101-FPN](../mask_rcnn/mask-rcnn_r101_fpn_1x_coco.py) | pytorch | 1x | 6.4 | 10.3 | 40.0 | 36.1 | [config](../mask_rcnn/mask-rcnn_r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r101_fpn_1x_coco/mask_rcnn_r101_fpn_1x_coco_20200204-1efe0ed5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_r101_fpn_1x_coco/mask_rcnn_r101_fpn_1x_coco_20200204_144809.log.json) |
+| [RegNetX-6.4GF-FPN](./mask-rcnn_regnetx-6.4GF_fpn_1x_coco.py) | pytorch | 1x | 6.1 | | 41.0 | 37.1 | [config](./mask-rcnn_regnetx-6.4GF_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-6.4GF_fpn_1x_coco/mask_rcnn_regnetx-6.4GF_fpn_1x_coco_20200517_180439-3a7aae83.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-6.4GF_fpn_1x_coco/mask_rcnn_regnetx-6.4GF_fpn_1x_coco_20200517_180439.log.json) |
+| [X-101-32x4d-FPN](../mask_rcnn/mask-rcnn_x101-32x4d_fpn_1x_coco.py) | pytorch | 1x | 7.6 | 9.4 | 41.9 | 37.5 | [config](../mask_rcnn/mask-rcnn_x101-32x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x4d_fpn_1x_coco/mask_rcnn_x101_32x4d_fpn_1x_coco_20200205-478d0b67.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/mask_rcnn/mask_rcnn_x101_32x4d_fpn_1x_coco/mask_rcnn_x101_32x4d_fpn_1x_coco_20200205_034906.log.json) |
+| [RegNetX-8.0GF-FPN](./mask-rcnn_regnetx-8GF_fpn_1x_coco.py) | pytorch | 1x | 6.4 | | 41.7 | 37.5 | [config](./mask-rcnn_regnetx-8GF_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-8GF_fpn_1x_coco/mask_rcnn_regnetx-8GF_fpn_1x_coco_20200517_180515-09daa87e.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-8GF_fpn_1x_coco/mask_rcnn_regnetx-8GF_fpn_1x_coco_20200517_180515.log.json) |
+| [RegNetX-12GF-FPN](./mask-rcnn_regnetx-12GF_fpn_1x_coco.py) | pytorch | 1x | 7.4 | | 42.2 | 38 | [config](./mask-rcnn_regnetx-12GF_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-12GF_fpn_1x_coco/mask_rcnn_regnetx-12GF_fpn_1x_coco_20200517_180552-b538bd8b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-12GF_fpn_1x_coco/mask_rcnn_regnetx-12GF_fpn_1x_coco_20200517_180552.log.json) |
+| [RegNetX-3.2GF-FPN-DCN-C3-C5](./mask-rcnn_regnetx-3.2GF-mdconv-c3-c5_fpn_1x_coco.py) | pytorch | 1x | 5.0 | | 40.3 | 36.6 | [config](./mask-rcnn_regnetx-3.2GF-mdconv-c3-c5_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-3.2GF_fpn_mdconv_c3-c5_1x_coco/mask_rcnn_regnetx-3.2GF_fpn_mdconv_c3-c5_1x_coco_20200520_172726-75f40794.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-3.2GF_fpn_mdconv_c3-c5_1x_coco/mask_rcnn_regnetx-3.2GF_fpn_mdconv_c3-c5_1x_coco_20200520_172726.log.json) |
+
+### Faster R-CNN
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :-------------------------------------------------------------: | :-----: | :-----: | :------: | :------------: | :----: | :-----------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| [R-50-FPN](../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py) | pytorch | 1x | 4.0 | 18.2 | 37.4 | [config](../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_1x_coco/faster_rcnn_r50_fpn_1x_coco_20200130-047c8118.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_1x_coco/faster_rcnn_r50_fpn_1x_coco_20200130_204655.log.json) |
+| [RegNetX-3.2GF-FPN](./faster-rcnn_regnetx-3.2GF_fpn_1x_coco.py) | pytorch | 1x | 4.5 | | 39.9 | [config](./faster-rcnn_regnetx-3.2GF_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-3.2GF_fpn_1x_coco/faster_rcnn_regnetx-3.2GF_fpn_1x_coco_20200517_175927-126fd9bf.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-3.2GF_fpn_1x_coco/faster_rcnn_regnetx-3.2GF_fpn_1x_coco_20200517_175927.log.json) |
+| [RegNetX-3.2GF-FPN](./faster-rcnn_regnetx-3.2GF_fpn_2x_coco.py) | pytorch | 2x | 4.5 | | 41.1 | [config](./faster-rcnn_regnetx-3.2GF_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-3.2GF_fpn_2x_coco/faster_rcnn_regnetx-3.2GF_fpn_2x_coco_20200520_223955-e2081918.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-3.2GF_fpn_2x_coco/faster_rcnn_regnetx-3.2GF_fpn_2x_coco_20200520_223955.log.json) |
+
+### RetinaNet
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :-----------------------------------------------------------: | :-----: | :-----: | :------: | :------------: | :----: | :-------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| [R-50-FPN](../retinanet/retinanet_r50_fpn_1x_coco.py) | pytorch | 1x | 3.8 | 16.6 | 36.5 | [config](../retinanet/retinanet_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r50_fpn_1x_coco/retinanet_r50_fpn_1x_coco_20200130-c2398f9e.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r50_fpn_1x_coco/retinanet_r50_fpn_1x_coco_20200130_002941.log.json) |
+| [RegNetX-800MF-FPN](./retinanet_regnetx-800MF_fpn_1x_coco.py) | pytorch | 1x | 2.5 | | 35.6 | [config](./retinanet_regnetx-800MF_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/retinanet_regnetx-800MF_fpn_1x_coco/retinanet_regnetx-800MF_fpn_1x_coco_20200517_191403-f6f91d10.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/retinanet_regnetx-800MF_fpn_1x_coco/retinanet_regnetx-800MF_fpn_1x_coco_20200517_191403.log.json) |
+| [RegNetX-1.6GF-FPN](./retinanet_regnetx-1.6GF_fpn_1x_coco.py) | pytorch | 1x | 3.3 | | 37.3 | [config](./retinanet_regnetx-1.6GF_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/retinanet_regnetx-1.6GF_fpn_1x_coco/retinanet_regnetx-1.6GF_fpn_1x_coco_20200517_191403-37009a9d.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/retinanet_regnetx-1.6GF_fpn_1x_coco/retinanet_regnetx-1.6GF_fpn_1x_coco_20200517_191403.log.json) |
+| [RegNetX-3.2GF-FPN](./retinanet_regnetx-3.2GF_fpn_1x_coco.py) | pytorch | 1x | 4.2 | | 39.1 | [config](./retinanet_regnetx-3.2GF_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/retinanet_regnetx-3.2GF_fpn_1x_coco/retinanet_regnetx-3.2GF_fpn_1x_coco_20200520_163141-cb1509e8.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/retinanet_regnetx-3.2GF_fpn_1x_coco/retinanet_regnetx-3.2GF_fpn_1x_coco_20200520_163141.log.json) |
+
+### Pre-trained models
+
+We also train some models with longer schedules and multi-scale training. The users could finetune them for downstream tasks.
+
+| Method | Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :---------------: | :----------------------------------------------------------------------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :-----------------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Faster RCNN | [RegNetX-400MF-FPN](./faster-rcnn_regnetx-400MF_fpn_ms-3x_coco.py) | pytorch | 3x | 2.3 | | 37.1 | - | [config](./faster-rcnn_regnetx-400MF_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-400MF_fpn_mstrain_3x_coco/faster_rcnn_regnetx-400MF_fpn_mstrain_3x_coco_20210526_095112-e1967c37.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-400MF_fpn_mstrain_3x_coco/faster_rcnn_regnetx-400MF_fpn_mstrain_3x_coco_20210526_095112.log.json) |
+| Faster RCNN | [RegNetX-800MF-FPN](./faster-rcnn_regnetx-800MF_fpn_ms-3x_coco.py) | pytorch | 3x | 2.8 | | 38.8 | - | [config](./faster-rcnn_regnetx-800MF_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-800MF_fpn_mstrain_3x_coco/faster_rcnn_regnetx-800MF_fpn_mstrain_3x_coco_20210526_095118-a2c70b20.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-800MF_fpn_mstrain_3x_coco/faster_rcnn_regnetx-800MF_fpn_mstrain_3x_coco_20210526_095118.log.json) |
+| Faster RCNN | [RegNetX-1.6GF-FPN](./faster-rcnn_regnetx-1.6GF_fpn_ms-3x_coco.py) | pytorch | 3x | 3.4 | | 40.5 | - | [config](./faster-rcnn_regnetx-1.6GF_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-1.6GF_fpn_mstrain_3x_coco/faster_rcnn_regnetx-1_20210526_095325-94aa46cc.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-1.6GF_fpn_mstrain_3x_coco/faster_rcnn_regnetx-1_20210526_095325.log.json) |
+| Faster RCNN | [RegNetX-3.2GF-FPN](./faster-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py) | pytorch | 3x | 4.4 | | 42.3 | - | [config](./faster-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-3.2GF_fpn_mstrain_3x_coco/faster_rcnn_regnetx-3_20210526_095152-e16a5227.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-3.2GF_fpn_mstrain_3x_coco/faster_rcnn_regnetx-3_20210526_095152.log.json) |
+| Faster RCNN | [RegNetX-4GF-FPN](./faster-rcnn_regnetx-4GF_fpn_ms-3x_coco.py) | pytorch | 3x | 4.9 | | 42.8 | - | [config](./faster-rcnn_regnetx-4GF_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-4GF_fpn_mstrain_3x_coco/faster_rcnn_regnetx-4GF_fpn_mstrain_3x_coco_20210526_095201-65eaf841.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-4GF_fpn_mstrain_3x_coco/faster_rcnn_regnetx-4GF_fpn_mstrain_3x_coco_20210526_095201.log.json) |
+| Mask RCNN | [RegNetX-400MF-FPN](./mask-rcnn_regnetx-400MF_fpn_ms-poly-3x_coco.py) | pytorch | 3x | 2.5 | | 37.6 | 34.4 | [config](./mask-rcnn_regnetx-400MF_fpn_ms-poly-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-400MF_fpn_mstrain-poly_3x_coco/mask_rcnn_regnetx-400MF_fpn_mstrain-poly_3x_coco_20210601_235443-8aac57a4.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-400MF_fpn_mstrain-poly_3x_coco/mask_rcnn_regnetx-400MF_fpn_mstrain-poly_3x_coco_20210601_235443.log.json) |
+| Mask RCNN | [RegNetX-800MF-FPN](./mask-rcnn_regnetx-800MF_fpn_ms-poly-3x_coco.py) | pytorch | 3x | 2.9 | | 39.5 | 36.1 | [config](./mask-rcnn_regnetx-800MF_fpn_ms-poly-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-800MF_fpn_mstrain-poly_3x_coco/mask_rcnn_regnetx-800MF_fpn_mstrain-poly_3x_coco_20210602_210641-715d51f5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-800MF_fpn_mstrain-poly_3x_coco/mask_rcnn_regnetx-800MF_fpn_mstrain-poly_3x_coco_20210602_210641.log.json) |
+| Mask RCNN | [RegNetX-1.6GF-FPN](./mask-rcnn_regnetx-1.6GF_fpn_ms-poly-3x_coco.py) | pytorch | 3x | 3.6 | | 40.9 | 37.5 | [config](./mask-rcnn_regnetx-1.6GF_fpn_ms-poly-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-1.6GF_fpn_mstrain-poly_3x_coco/mask_rcnn_regnetx-1_20210602_210641-6764cff5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-1.6GF_fpn_mstrain-poly_3x_coco/mask_rcnn_regnetx-1_20210602_210641.log.json) |
+| Mask RCNN | [RegNetX-3.2GF-FPN](./mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py) | pytorch | 3x | 5.0 | | 43.1 | 38.7 | [config](./mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-3.2GF_fpn_mstrain_3x_coco/mask_rcnn_regnetx-3.2GF_fpn_mstrain_3x_coco_20200521_202221-99879813.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-3.2GF_fpn_mstrain_3x_coco/mask_rcnn_regnetx-3.2GF_fpn_mstrain_3x_coco_20200521_202221.log.json) |
+| Mask RCNN | [RegNetX-4GF-FPN](./mask-rcnn_regnetx-4GF_fpn_ms-poly-3x_coco.py) | pytorch | 3x | 5.1 | | 43.4 | 39.2 | [config](./mask-rcnn_regnetx-4GF_fpn_ms-poly-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-4GF_fpn_mstrain-poly_3x_coco/mask_rcnn_regnetx-4GF_fpn_mstrain-poly_3x_coco_20210602_032621-00f0331c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-4GF_fpn_mstrain-poly_3x_coco/mask_rcnn_regnetx-4GF_fpn_mstrain-poly_3x_coco_20210602_032621.log.json) |
+| Cascade Mask RCNN | [RegNetX-400MF-FPN](./cascade-mask-rcnn_regnetx-400MF_fpn_ms-3x_coco.py) | pytorch | 3x | 4.3 | | 41.6 | 36.4 | [config](./cascade-mask-rcnn_regnetx-400MF_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/cascade_mask_rcnn_regnetx-400MF_fpn_mstrain_3x_coco/cascade_mask_rcnn_regnetx-400MF_fpn_mstrain_3x_coco_20210715_211619-5142f449.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/cascade_mask_rcnn_regnetx-400MF_fpn_mstrain_3x_coco/cascade_mask_rcnn_regnetx-400MF_fpn_mstrain_3x_coco_20210715_211619.log.json) |
+| Cascade Mask RCNN | [RegNetX-800MF-FPN](./cascade-mask-rcnn_regnetx-800MF_fpn_ms-3x_coco.py) | pytorch | 3x | 4.8 | | 42.8 | 37.6 | [config](./cascade-mask-rcnn_regnetx-800MF_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/cascade_mask_rcnn_regnetx-800MF_fpn_mstrain_3x_coco/cascade_mask_rcnn_regnetx-800MF_fpn_mstrain_3x_coco_20210715_211616-dcbd13f4.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/cascade_mask_rcnn_regnetx-800MF_fpn_mstrain_3x_coco/cascade_mask_rcnn_regnetx-800MF_fpn_mstrain_3x_coco_20210715_211616.log.json) |
+| Cascade Mask RCNN | [RegNetX-1.6GF-FPN](./cascade-mask-rcnn_regnetx-1.6GF_fpn_ms-3x_coco.py) | pytorch | 3x | 5.4 | | 44.5 | 39.0 | [config](./cascade-mask-rcnn_regnetx-1.6GF_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/cascade_mask_rcnn_regnetx-1.6GF_fpn_mstrain_3x_coco/cascade_mask_rcnn_regnetx-1_20210715_211616-75f29a61.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/cascade_mask_rcnn_regnetx-1.6GF_fpn_mstrain_3x_coco/cascade_mask_rcnn_regnetx-1_20210715_211616.log.json) |
+| Cascade Mask RCNN | [RegNetX-3.2GF-FPN](./cascade-mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py) | pytorch | 3x | 6.4 | | 45.8 | 40.0 | [config](./cascade-mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/cascade_mask_rcnn_regnetx-3.2GF_fpn_mstrain_3x_coco/cascade_mask_rcnn_regnetx-3_20210715_211616-b9c2c58b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/cascade_mask_rcnn_regnetx-3.2GF_fpn_mstrain_3x_coco/cascade_mask_rcnn_regnetx-3_20210715_211616.log.json) |
+| Cascade Mask RCNN | [RegNetX-4GF-FPN](./cascade-mask-rcnn_regnetx-4GF_fpn_ms-3x_coco.py) | pytorch | 3x | 6.9 | | 45.8 | 40.0 | [config](./cascade-mask-rcnn_regnetx-4GF_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/regnet/cascade_mask_rcnn_regnetx-4GF_fpn_mstrain_3x_coco/cascade_mask_rcnn_regnetx-4GF_fpn_mstrain_3x_coco_20210715_212034-cbb1be4c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/regnet/cascade_mask_rcnn_regnetx-4GF_fpn_mstrain_3x_coco/cascade_mask_rcnn_regnetx-4GF_fpn_mstrain_3x_coco_20210715_212034.log.json) |
+
+### Notice
+
+1. The models are trained using a different weight decay, i.e., `weight_decay=5e-5` according to the setting in ImageNet training. This brings improvement of at least 0.7 AP absolute but does not improve the model using ResNet-50.
+2. RetinaNets using RegNets are trained with learning rate 0.02 with gradient clip. We find that using learning rate 0.02 could improve the results by at least 0.7 AP absolute and gradient clip is necessary to stabilize the training. However, this does not improve the performance of ResNet-50-FPN RetinaNet.
+
+## Citation
+
+```latex
+@article{radosavovic2020designing,
+ title={Designing Network Design Spaces},
+ author={Ilija Radosavovic and Raj Prateek Kosaraju and Ross Girshick and Kaiming He and Piotr Dollár},
+ year={2020},
+ eprint={2003.13678},
+ archivePrefix={arXiv},
+ primaryClass={cs.CV}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/cascade-mask-rcnn_regnetx-1.6GF_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/cascade-mask-rcnn_regnetx-1.6GF_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..74e6adaba5c262d45aaec876d1225b0061bb290b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/cascade-mask-rcnn_regnetx-1.6GF_fpn_ms-3x_coco.py
@@ -0,0 +1,17 @@
+_base_ = 'cascade-mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py'
+model = dict(
+ backbone=dict(
+ type='RegNet',
+ arch='regnetx_1.6gf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_1.6gf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[72, 168, 408, 912],
+ out_channels=256,
+ num_outs=5))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/cascade-mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/cascade-mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ea219021260b6aa3a844eb6b4780e9669e50ed3b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/cascade-mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py
@@ -0,0 +1,28 @@
+_base_ = [
+ '../common/ms_3x_coco-instance.py',
+ '../_base_/models/cascade-mask-rcnn_r50_fpn.py'
+]
+model = dict(
+ data_preprocessor=dict(
+ # The mean and std are used in PyCls when training RegNets
+ mean=[103.53, 116.28, 123.675],
+ std=[57.375, 57.12, 58.395],
+ bgr_to_rgb=False),
+ backbone=dict(
+ _delete_=True,
+ type='RegNet',
+ arch='regnetx_3.2gf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_3.2gf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[96, 192, 432, 1008],
+ out_channels=256,
+ num_outs=5))
+
+optim_wrapper = dict(optimizer=dict(weight_decay=0.00005))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/cascade-mask-rcnn_regnetx-400MF_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/cascade-mask-rcnn_regnetx-400MF_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3fe47f837437163710ecd28f1bb217c643464965
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/cascade-mask-rcnn_regnetx-400MF_fpn_ms-3x_coco.py
@@ -0,0 +1,17 @@
+_base_ = 'cascade-mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py'
+model = dict(
+ backbone=dict(
+ type='RegNet',
+ arch='regnetx_400mf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_400mf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[32, 64, 160, 384],
+ out_channels=256,
+ num_outs=5))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/cascade-mask-rcnn_regnetx-4GF_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/cascade-mask-rcnn_regnetx-4GF_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e22886a80f92ba4269477a307b2689c45468381c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/cascade-mask-rcnn_regnetx-4GF_fpn_ms-3x_coco.py
@@ -0,0 +1,17 @@
+_base_ = 'cascade-mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py'
+model = dict(
+ backbone=dict(
+ type='RegNet',
+ arch='regnetx_4.0gf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_4.0gf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[80, 240, 560, 1360],
+ out_channels=256,
+ num_outs=5))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/cascade-mask-rcnn_regnetx-800MF_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/cascade-mask-rcnn_regnetx-800MF_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..655bdc60c772875e0a1ed871bd6bf02aab8e39cc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/cascade-mask-rcnn_regnetx-800MF_fpn_ms-3x_coco.py
@@ -0,0 +1,17 @@
+_base_ = 'cascade-mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py'
+model = dict(
+ backbone=dict(
+ type='RegNet',
+ arch='regnetx_800mf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_800mf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[64, 128, 288, 672],
+ out_channels=256,
+ num_outs=5))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-1.6GF_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-1.6GF_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e9e8302bdd1537b825f36777e3211d27dec8fb0c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-1.6GF_fpn_ms-3x_coco.py
@@ -0,0 +1,17 @@
+_base_ = 'faster-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py'
+model = dict(
+ backbone=dict(
+ type='RegNet',
+ arch='regnetx_1.6gf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_1.6gf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[72, 168, 408, 912],
+ out_channels=256,
+ num_outs=5))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-3.2GF_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-3.2GF_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..db49092e2fb7e1cf3dbcad2bb99aa08396ea35e7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-3.2GF_fpn_1x_coco.py
@@ -0,0 +1,30 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ data_preprocessor=dict(
+ # The mean and std are used in PyCls when training RegNets
+ mean=[103.53, 116.28, 123.675],
+ std=[57.375, 57.12, 58.395],
+ bgr_to_rgb=False),
+ backbone=dict(
+ _delete_=True,
+ type='RegNet',
+ arch='regnetx_3.2gf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_3.2gf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[96, 192, 432, 1008],
+ out_channels=256,
+ num_outs=5))
+
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.00005))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-3.2GF_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-3.2GF_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..be533603085a89b65556b47f5e333fdde734bbd1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-3.2GF_fpn_2x_coco.py
@@ -0,0 +1,16 @@
+_base_ = './faster-rcnn_regnetx-3.2GF_fpn_1x_coco.py'
+
+# learning policy
+max_epochs = 24
+train_cfg = dict(max_epochs=max_epochs)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d3d5d5d689162d805c0cfb4d84f9a128faf90c25
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py
@@ -0,0 +1,25 @@
+_base_ = ['../common/ms_3x_coco.py', '../_base_/models/faster-rcnn_r50_fpn.py']
+model = dict(
+ data_preprocessor=dict(
+ # The mean and std are used in PyCls when training RegNets
+ mean=[103.53, 116.28, 123.675],
+ std=[57.375, 57.12, 58.395],
+ bgr_to_rgb=False),
+ backbone=dict(
+ _delete_=True,
+ type='RegNet',
+ arch='regnetx_3.2gf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_3.2gf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[96, 192, 432, 1008],
+ out_channels=256,
+ num_outs=5))
+
+optim_wrapper = dict(optimizer=dict(weight_decay=0.00005))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-400MF_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-400MF_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..2edeff9c1f5a794ed14dc8723917986ac26e3d36
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-400MF_fpn_ms-3x_coco.py
@@ -0,0 +1,17 @@
+_base_ = 'faster-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py'
+model = dict(
+ backbone=dict(
+ type='RegNet',
+ arch='regnetx_400mf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_400mf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[32, 64, 160, 384],
+ out_channels=256,
+ num_outs=5))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-4GF_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-4GF_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..afcbb5d5d1a8aee47267d1f82fff8d40fa0d8e9b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-4GF_fpn_ms-3x_coco.py
@@ -0,0 +1,17 @@
+_base_ = 'faster-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py'
+model = dict(
+ backbone=dict(
+ type='RegNet',
+ arch='regnetx_4.0gf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_4.0gf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[80, 240, 560, 1360],
+ out_channels=256,
+ num_outs=5))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-800MF_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-800MF_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f659ec9689068afd94aa3bc545d4fed91ffb5eb4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/faster-rcnn_regnetx-800MF_fpn_ms-3x_coco.py
@@ -0,0 +1,17 @@
+_base_ = 'faster-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py'
+model = dict(
+ backbone=dict(
+ type='RegNet',
+ arch='regnetx_800mf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_800mf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[64, 128, 288, 672],
+ out_channels=256,
+ num_outs=5))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-1.6GF_fpn_ms-poly-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-1.6GF_fpn_ms-poly-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..60874c66dbc37df824a9c44bb8c28a441f7f84e4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-1.6GF_fpn_ms-poly-3x_coco.py
@@ -0,0 +1,26 @@
+_base_ = [
+ '../common/ms-poly_3x_coco-instance.py',
+ '../_base_/models/mask-rcnn_r50_fpn.py'
+]
+
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='RegNet',
+ arch='regnetx_1.6gf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_1.6gf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[72, 168, 408, 912],
+ out_channels=256,
+ num_outs=5))
+
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.00005),
+ clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-12GF_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-12GF_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e82cecea010fb32143f809add198a052285a6897
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-12GF_fpn_1x_coco.py
@@ -0,0 +1,17 @@
+_base_ = './mask-rcnn_regnetx-3.2GF_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='RegNet',
+ arch='regnetx_12gf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_12gf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[224, 448, 896, 2240],
+ out_channels=256,
+ num_outs=5))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-3.2GF-mdconv-c3-c5_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-3.2GF-mdconv-c3-c5_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c7c1d1ac3a7bd87bd210b4cd2194dd7e430f8d96
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-3.2GF-mdconv-c3-c5_fpn_1x_coco.py
@@ -0,0 +1,7 @@
+_base_ = 'mask-rcnn_regnetx-3.2GF_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCNv2', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_3.2gf')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-3.2GF_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-3.2GF_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c52bf13ff6df5cda353c21ac32a950602620dbde
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-3.2GF_fpn_1x_coco.py
@@ -0,0 +1,30 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ data_preprocessor=dict(
+ # The mean and std are used in PyCls when training RegNets
+ mean=[103.53, 116.28, 123.675],
+ std=[57.375, 57.12, 58.395],
+ bgr_to_rgb=False),
+ backbone=dict(
+ _delete_=True,
+ type='RegNet',
+ arch='regnetx_3.2gf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_3.2gf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[96, 192, 432, 1008],
+ out_channels=256,
+ num_outs=5))
+
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.00005))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..36482c939dc3e600171b98bc159440e5fb740ffa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py
@@ -0,0 +1,60 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ data_preprocessor=dict(
+ # The mean and std are used in PyCls when training RegNets
+ mean=[103.53, 116.28, 123.675],
+ std=[57.375, 57.12, 58.395],
+ bgr_to_rgb=False),
+ backbone=dict(
+ _delete_=True,
+ type='RegNet',
+ arch='regnetx_3.2gf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_3.2gf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[96, 192, 432, 1008],
+ out_channels=256,
+ num_outs=5))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.00005),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+# learning policy
+max_epochs = 36
+train_cfg = dict(max_epochs=max_epochs)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[28, 34],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-400MF_fpn_ms-poly-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-400MF_fpn_ms-poly-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b96e1921f0dae8ad6656a7785d9d4655f9f349b3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-400MF_fpn_ms-poly-3x_coco.py
@@ -0,0 +1,26 @@
+_base_ = [
+ '../common/ms-poly_3x_coco-instance.py',
+ '../_base_/models/mask-rcnn_r50_fpn.py'
+]
+
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='RegNet',
+ arch='regnetx_400mf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_400mf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[32, 64, 160, 384],
+ out_channels=256,
+ num_outs=5))
+
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.00005),
+ clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-4GF_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-4GF_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ce9f8ef4ffbcce66ec0184b3ff06a92425231597
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-4GF_fpn_1x_coco.py
@@ -0,0 +1,17 @@
+_base_ = './mask-rcnn_regnetx-3.2GF_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='RegNet',
+ arch='regnetx_4.0gf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_4.0gf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[80, 240, 560, 1360],
+ out_channels=256,
+ num_outs=5))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-4GF_fpn_ms-poly-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-4GF_fpn_ms-poly-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f160ccf66700d98a6403ed736928e529368e800c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-4GF_fpn_ms-poly-3x_coco.py
@@ -0,0 +1,26 @@
+_base_ = [
+ '../common/ms-poly_3x_coco-instance.py',
+ '../_base_/models/mask-rcnn_r50_fpn.py'
+]
+
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='RegNet',
+ arch='regnetx_4.0gf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_4.0gf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[80, 240, 560, 1360],
+ out_channels=256,
+ num_outs=5))
+
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.00005),
+ clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-6.4GF_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-6.4GF_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e17a3d7695fa7ba9e135d7a436118aae29be4747
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-6.4GF_fpn_1x_coco.py
@@ -0,0 +1,17 @@
+_base_ = './mask-rcnn_regnetx-3.2GF_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='RegNet',
+ arch='regnetx_6.4gf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_6.4gf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[168, 392, 784, 1624],
+ out_channels=256,
+ num_outs=5))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-800MF_fpn_ms-poly-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-800MF_fpn_ms-poly-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..93851fdbb99e5d8e3a58062c7ad83d2acad14ac6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-800MF_fpn_ms-poly-3x_coco.py
@@ -0,0 +1,26 @@
+_base_ = [
+ '../common/ms-poly_3x_coco-instance.py',
+ '../_base_/models/mask-rcnn_r50_fpn.py'
+]
+
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='RegNet',
+ arch='regnetx_800mf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_800mf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[64, 128, 288, 672],
+ out_channels=256,
+ num_outs=5))
+
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.00005),
+ clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-8GF_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-8GF_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..62a4c931512e6b46093b03fd4e80741a93151c6a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/mask-rcnn_regnetx-8GF_fpn_1x_coco.py
@@ -0,0 +1,17 @@
+_base_ = './mask-rcnn_regnetx-3.2GF_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='RegNet',
+ arch='regnetx_8.0gf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_8.0gf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[80, 240, 720, 1920],
+ out_channels=256,
+ num_outs=5))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..19fbba80f0396e1dad7a330ef769d98ad1a0c4d2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/metafile.yml
@@ -0,0 +1,797 @@
+Models:
+ - Name: mask-rcnn_regnetx-3.2GF_fpn_1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/regnet/mask-rcnn_regnetx-3.2GF_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.0
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.3
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-3.2GF_fpn_1x_coco/mask_rcnn_regnetx-3.2GF_fpn_1x_coco_20200520_163141-2a9d1814.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: mask-rcnn_regnetx-4GF_fpn_1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/regnet/mask-rcnn_regnetx-4GF_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.5
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-4GF_fpn_1x_coco/mask_rcnn_regnetx-4GF_fpn_1x_coco_20200517_180217-32e9c92d.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: mask-rcnn_regnetx-6.4GF_fpn_1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/regnet/mask-rcnn_regnetx-6.4GF_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.1
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-6.4GF_fpn_1x_coco/mask_rcnn_regnetx-6.4GF_fpn_1x_coco_20200517_180439-3a7aae83.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: mask-rcnn_regnetx-8GF_fpn_1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/regnet/mask-rcnn_regnetx-8GF_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.4
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.7
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-8GF_fpn_1x_coco/mask_rcnn_regnetx-8GF_fpn_1x_coco_20200517_180515-09daa87e.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: mask-rcnn_regnetx-12GF_fpn_1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/regnet/mask-rcnn_regnetx-12GF_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.4
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-12GF_fpn_1x_coco/mask_rcnn_regnetx-12GF_fpn_1x_coco_20200517_180552-b538bd8b.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: mask-rcnn_regnetx-3.2GF-mdconv-c3-c5_fpn_1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/regnet/mask-rcnn_regnetx-3.2GF-mdconv-c3-c5_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.0
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.3
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-3.2GF_fpn_mdconv_c3-c5_1x_coco/mask_rcnn_regnetx-3.2GF_fpn_mdconv_c3-c5_1x_coco_20200520_172726-75f40794.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: faster-rcnn_regnetx-3.2GF_fpn_1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/regnet/faster-rcnn_regnetx-3.2GF_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.5
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-3.2GF_fpn_1x_coco/faster_rcnn_regnetx-3.2GF_fpn_1x_coco_20200517_175927-126fd9bf.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: faster-rcnn_regnetx-3.2GF_fpn_2x_coco
+ In Collection: Faster R-CNN
+ Config: configs/regnet/faster-rcnn_regnetx-3.2GF_fpn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 4.5
+ Epochs: 24
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-3.2GF_fpn_2x_coco/faster_rcnn_regnetx-3.2GF_fpn_2x_coco_20200520_223955-e2081918.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: retinanet_regnetx-800MF_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/regnet/retinanet_regnetx-800MF_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 2.5
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 35.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/retinanet_regnetx-800MF_fpn_1x_coco/retinanet_regnetx-800MF_fpn_1x_coco_20200517_191403-f6f91d10.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: retinanet_regnetx-1.6GF_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/regnet/retinanet_regnetx-1.6GF_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.3
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/retinanet_regnetx-1.6GF_fpn_1x_coco/retinanet_regnetx-1.6GF_fpn_1x_coco_20200517_191403-37009a9d.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: retinanet_regnetx-3.2GF_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/regnet/retinanet_regnetx-3.2GF_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.2
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/retinanet_regnetx-3.2GF_fpn_1x_coco/retinanet_regnetx-3.2GF_fpn_1x_coco_20200520_163141-cb1509e8.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: faster-rcnn_regnetx-400MF_fpn_ms-3x_coco
+ In Collection: Faster R-CNN
+ Config: configs/regnet/faster-rcnn_regnetx-400MF_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 2.3
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-400MF_fpn_mstrain_3x_coco/faster_rcnn_regnetx-400MF_fpn_mstrain_3x_coco_20210526_095112-e1967c37.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: faster-rcnn_regnetx-800MF_fpn_ms-3x_coco
+ In Collection: Faster R-CNN
+ Config: configs/regnet/faster-rcnn_regnetx-800MF_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 2.8
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-800MF_fpn_mstrain_3x_coco/faster_rcnn_regnetx-800MF_fpn_mstrain_3x_coco_20210526_095118-a2c70b20.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: faster-rcnn_regnetx-1.6GF_fpn_ms-3x_coco
+ In Collection: Faster R-CNN
+ Config: configs/regnet/faster-rcnn_regnetx-1.6GF_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 3.4
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-1.6GF_fpn_mstrain_3x_coco/faster_rcnn_regnetx-1_20210526_095325-94aa46cc.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: faster-rcnn_regnetx-3.2GF_fpn_ms-3x_coco
+ In Collection: Faster R-CNN
+ Config: configs/regnet/faster-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 4.4
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-3.2GF_fpn_mstrain_3x_coco/faster_rcnn_regnetx-3_20210526_095152-e16a5227.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: faster-rcnn_regnetx-4GF_fpn_ms-3x_coco
+ In Collection: Faster R-CNN
+ Config: configs/regnet/faster-rcnn_regnetx-4GF_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 4.9
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/faster_rcnn_regnetx-4GF_fpn_mstrain_3x_coco/faster_rcnn_regnetx-4GF_fpn_mstrain_3x_coco_20210526_095201-65eaf841.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco
+ In Collection: Mask R-CNN
+ Config: configs/regnet/mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 5.0
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.1
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-3.2GF_fpn_mstrain_3x_coco/mask_rcnn_regnetx-3.2GF_fpn_mstrain_3x_coco_20200521_202221-99879813.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: mask-rcnn_regnetx-400MF_fpn_ms-poly-3x_coco
+ In Collection: Mask R-CNN
+ Config: configs/regnet/mask-rcnn_regnetx-400MF_fpn_ms-poly-3x_coco.py
+ Metadata:
+ Training Memory (GB): 2.5
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.6
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 34.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-400MF_fpn_mstrain-poly_3x_coco/mask_rcnn_regnetx-400MF_fpn_mstrain-poly_3x_coco_20210601_235443-8aac57a4.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: mask-rcnn_regnetx-800MF_fpn_ms-poly-3x_coco
+ In Collection: Mask R-CNN
+ Config: configs/regnet/mask-rcnn_regnetx-800MF_fpn_ms-poly-3x_coco.py
+ Metadata:
+ Training Memory (GB): 2.9
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-800MF_fpn_mstrain-poly_3x_coco/mask_rcnn_regnetx-800MF_fpn_mstrain-poly_3x_coco_20210602_210641-715d51f5.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: mask-rcnn_regnetx-1.6GF_fpn_ms-poly-3x_coco
+ In Collection: Mask R-CNN
+ Config: configs/regnet/mask-rcnn_regnetx-1.6GF_fpn_ms-poly-3x_coco.py
+ Metadata:
+ Training Memory (GB): 3.6
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.9
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-1.6GF_fpn_mstrain-poly_3x_coco/mask_rcnn_regnetx-1_20210602_210641-6764cff5.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco
+ In Collection: Mask R-CNN
+ Config: configs/regnet/mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 5.0
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.1
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-3.2GF_fpn_mstrain_3x_coco/mask_rcnn_regnetx-3.2GF_fpn_mstrain_3x_coco_20200521_202221-99879813.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: mask-rcnn_regnetx-4GF_fpn_ms-poly-3x_coco
+ In Collection: Mask R-CNN
+ Config: configs/regnet/mask-rcnn_regnetx-4GF_fpn_ms-poly-3x_coco.py
+ Metadata:
+ Training Memory (GB): 5.1
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/mask_rcnn_regnetx-4GF_fpn_mstrain-poly_3x_coco/mask_rcnn_regnetx-4GF_fpn_mstrain-poly_3x_coco_20210602_032621-00f0331c.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: cascade-mask-rcnn_regnetx-400MF_fpn_ms-3x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/regnet/cascade-mask-rcnn_regnetx-400MF_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 4.3
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.6
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/cascade_mask_rcnn_regnetx-400MF_fpn_mstrain_3x_coco/cascade_mask_rcnn_regnetx-400MF_fpn_mstrain_3x_coco_20210715_211619-5142f449.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: cascade-mask-rcnn_regnetx-800MF_fpn_ms-3x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/regnet/cascade-mask-rcnn_regnetx-800MF_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 4.8
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/cascade_mask_rcnn_regnetx-800MF_fpn_mstrain_3x_coco/cascade_mask_rcnn_regnetx-800MF_fpn_mstrain_3x_coco_20210715_211616-dcbd13f4.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: cascade-mask-rcnn_regnetx-1.6GF_fpn_ms-3x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/regnet/cascade-mask-rcnn_regnetx-1.6GF_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 5.4
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/cascade_mask_rcnn_regnetx-1.6GF_fpn_mstrain_3x_coco/cascade_mask_rcnn_regnetx-1_20210715_211616-75f29a61.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: cascade-mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/regnet/cascade-mask-rcnn_regnetx-3.2GF_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 6.4
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 40.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/cascade_mask_rcnn_regnetx-3.2GF_fpn_mstrain_3x_coco/cascade_mask_rcnn_regnetx-3_20210715_211616-b9c2c58b.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
+
+ - Name: cascade-mask-rcnn_regnetx-4GF_fpn_ms-3x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/regnet/cascade-mask-rcnn_regnetx-4GF_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 6.9
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - RegNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 40.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/regnet/cascade_mask_rcnn_regnetx-4GF_fpn_mstrain_3x_coco/cascade_mask_rcnn_regnetx-4GF_fpn_mstrain_3x_coco_20210715_212034-cbb1be4c.pth
+ Paper:
+ URL: https://arxiv.org/abs/2003.13678
+ Title: 'Designing Network Design Spaces'
+ README: configs/regnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/regnet.py#L11
+ Version: v2.1.0
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/retinanet_regnetx-1.6GF_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/retinanet_regnetx-1.6GF_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..7395c1bfbfa16670294c721f9f3135da9b9e69ae
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/retinanet_regnetx-1.6GF_fpn_1x_coco.py
@@ -0,0 +1,17 @@
+_base_ = './retinanet_regnetx-3.2GF_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='RegNet',
+ arch='regnetx_1.6gf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_1.6gf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[72, 168, 408, 912],
+ out_channels=256,
+ num_outs=5))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/retinanet_regnetx-3.2GF_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/retinanet_regnetx-3.2GF_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8b8a32cec195901e2f1326bf62f4fa4508e744d2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/retinanet_regnetx-3.2GF_fpn_1x_coco.py
@@ -0,0 +1,31 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ data_preprocessor=dict(
+ # The mean and std are used in PyCls when training RegNets
+ mean=[103.53, 116.28, 123.675],
+ std=[57.375, 57.12, 58.395],
+ bgr_to_rgb=False),
+ backbone=dict(
+ _delete_=True,
+ type='RegNet',
+ arch='regnetx_3.2gf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_3.2gf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[96, 192, 432, 1008],
+ out_channels=256,
+ num_outs=5))
+
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.00005),
+ clip_grad=dict(max_norm=35, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/retinanet_regnetx-800MF_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/retinanet_regnetx-800MF_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f6f8989320d6ffbcd55148471f62a962c52f9131
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/regnet/retinanet_regnetx-800MF_fpn_1x_coco.py
@@ -0,0 +1,17 @@
+_base_ = './retinanet_regnetx-3.2GF_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='RegNet',
+ arch='regnetx_800mf',
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://regnetx_800mf')),
+ neck=dict(
+ type='FPN',
+ in_channels=[64, 128, 288, 672],
+ out_channels=256,
+ num_outs=5))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/reid/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reid/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..a5bfe5ec49947e939a3261fa9938d77cc04df44f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reid/README.md
@@ -0,0 +1,135 @@
+# Training a ReID Model
+
+You may want to train a ReID model for multiple object tracking or other applications. We support ReID model training in MMDetection, which is built upon [MMPretrain](https://github.com/open-mmlab/mmpretrain).
+
+### 1. Development Environment Setup
+
+Tracking Development Environment Setup can refer to this [document](../../docs/en/get_started.md).
+
+### 2. Dataset Preparation
+
+This section will show how to train a ReID model on standard datasets i.e. MOT17.
+
+We need to download datasets following docs. We use [ReIDDataset](mmdet/datasets/reid_dataset.py) to maintain standard datasets. In this case, you need to convert the official dataset to this style. We provide scripts and the usages as follow:
+
+```python
+python tools/dataset_converters/mot2reid.py -i ./data/MOT17/ -o ./data/MOT17/reid --val-split 0.2 --vis-threshold 0.3
+```
+
+Arguments:
+
+- `--val-split`: Proportion of the validation dataset to the whole ReID dataset.
+- `--vis-threshold`: Threshold of visibility for each person.
+
+The directory of the converted datasets is as follows:
+
+```
+MOT17
+├── train
+├── test
+├── reid
+│ ├── imgs
+│ │ ├── MOT17-02-FRCNN_000002
+│ │ │ ├── 000000.jpg
+│ │ │ ├── 000001.jpg
+│ │ │ ├── ...
+│ │ ├── MOT17-02-FRCNN_000003
+│ │ │ ├── 000000.jpg
+│ │ │ ├── 000001.jpg
+│ │ │ ├── ...
+│ ├── meta
+│ │ ├── train_80.txt
+│ │ ├── val_20.txt
+```
+
+Note: `80` in `train_80.txt` means the proportion of the training dataset to the whole ReID dataset is eighty percent. While the proportion of the validation dataset is twenty percent.
+
+For training, we provide a annotation list `train_80.txt`. Each line of the list constraints a filename and its corresponding ground-truth labels. The format is as follows:
+
+```
+MOT17-05-FRCNN_000110/000018.jpg 0
+MOT17-13-FRCNN_000146/000014.jpg 1
+MOT17-05-FRCNN_000088/000004.jpg 2
+MOT17-02-FRCNN_000009/000081.jpg 3
+```
+
+For validation, The annotation list `val_20.txt` remains the same as format above.
+
+Note: Images in `MOT17/reid/imgs` are cropped from raw images in `MOT17/train` by the corresponding `gt.txt`. The value of ground-truth labels should fall in range `[0, num_classes - 1]`.
+
+### 3. Training
+
+#### Training on a single GPU
+
+```shell
+python tools/train.py configs/reid/reid_r50_8xb32-6e_mot17train80_test-mot17val20.py
+```
+
+#### Training on multiple GPUs
+
+We provide `tools/dist_train.sh` to launch training on multiple GPUs.
+The basic usage is as follows.
+
+```shell
+bash tools/dist_train.sh configs/reid/reid_r50_8xb32-6e_mot17train80_test-mot17val20.py 8
+```
+
+### 4. Customize Dataset
+
+This section will show how to train a ReID model on customize datasets.
+
+### 4.1 Dataset Preparation
+
+You need to convert your customize datasets to existing dataset format.
+
+#### An example of customized dataset
+
+Assume we are going to implement a `Filelist` dataset, which takes filelists for both training and testing. The directory of the dataset is as follows:
+
+```
+Filelist
+├── imgs
+│ ├── person1
+│ │ ├── 000000.jpg
+│ │ ├── 000001.jpg
+│ │ ├── ...
+│ ├── person2
+│ │ ├── 000000.jpg
+│ │ ├── 000001.jpg
+│ │ ├── ...
+├── meta
+│ ├── train.txt
+│ ├── val.txt
+```
+
+The format of annotation list is as follows:
+
+```
+person1/000000.jpg 0
+person1/000001.jpg 0
+person2/000000.jpg 1
+person2/000001.jpg 1
+```
+
+You can directly use [ReIDDataset](mmdet/datasets/reid_dataset.py). In this case, you only need to modify the config as follows:
+
+```python
+# modify the path of annotation files and the image path prefix
+data = dict(
+ train=dict(
+ data_prefix='data/Filelist/imgs',
+ ann_file='data/Filelist/meta/train.txt'),
+ val=dict(
+ data_prefix='data/Filelist/imgs',
+ ann_file='data/Filelist/meta/val.txt'),
+ test=dict(
+ data_prefix='data/Filelist/imgs',
+ ann_file='data/Filelist/meta/val.txt'),
+)
+# modify the number of classes, assume your training set has 100 classes
+model = dict(reid=dict(head=dict(num_classes=100)))
+```
+
+### 4.2 Training
+
+The training stage is the same as `Standard Dataset`.
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/reid/reid_r50_8xb32-6e_mot15train80_test-mot15val20.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reid/reid_r50_8xb32-6e_mot15train80_test-mot15val20.py
new file mode 100644
index 0000000000000000000000000000000000000000..4e30b22964d0504771678dbd0a551bc16a0714ea
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reid/reid_r50_8xb32-6e_mot15train80_test-mot15val20.py
@@ -0,0 +1,7 @@
+_base_ = ['./reid_r50_8xb32-6e_mot17train80_test-mot17val20.py']
+model = dict(head=dict(num_classes=368))
+# data
+data_root = 'data/MOT15/'
+train_dataloader = dict(dataset=dict(data_root=data_root))
+val_dataloader = dict(dataset=dict(data_root=data_root))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/reid/reid_r50_8xb32-6e_mot16train80_test-mot16val20.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reid/reid_r50_8xb32-6e_mot16train80_test-mot16val20.py
new file mode 100644
index 0000000000000000000000000000000000000000..468b9bfb2453f97c83282cc2f383c7592694269c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reid/reid_r50_8xb32-6e_mot16train80_test-mot16val20.py
@@ -0,0 +1,7 @@
+_base_ = ['./reid_r50_8xb32-6e_mot17train80_test-mot17val20.py']
+model = dict(head=dict(num_classes=371))
+# data
+data_root = 'data/MOT16/'
+train_dataloader = dict(dataset=dict(data_root=data_root))
+val_dataloader = dict(dataset=dict(data_root=data_root))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/reid/reid_r50_8xb32-6e_mot17train80_test-mot17val20.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reid/reid_r50_8xb32-6e_mot17train80_test-mot17val20.py
new file mode 100644
index 0000000000000000000000000000000000000000..83669de7c170c5de0e2054808ef7a76878bc1f24
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reid/reid_r50_8xb32-6e_mot17train80_test-mot17val20.py
@@ -0,0 +1,61 @@
+_base_ = [
+ '../_base_/datasets/mot_challenge_reid.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ type='BaseReID',
+ data_preprocessor=dict(
+ type='ReIDDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ to_rgb=True),
+ backbone=dict(
+ type='mmpretrain.ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(3, ),
+ style='pytorch'),
+ neck=dict(type='GlobalAveragePooling', kernel_size=(8, 4), stride=1),
+ head=dict(
+ type='LinearReIDHead',
+ num_fcs=1,
+ in_channels=2048,
+ fc_channels=1024,
+ out_channels=128,
+ num_classes=380,
+ loss_cls=dict(type='mmpretrain.CrossEntropyLoss', loss_weight=1.0),
+ loss_triplet=dict(type='TripletLoss', margin=0.3, loss_weight=1.0),
+ norm_cfg=dict(type='BN1d'),
+ act_cfg=dict(type='ReLU')),
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint= # noqa: E251
+ 'https://download.openmmlab.com/mmclassification/v0/resnet/resnet50_batch256_imagenet_20200708-cfb998bf.pth' # noqa: E501
+ ))
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ clip_grad=None,
+ optimizer=dict(type='SGD', lr=0.1, momentum=0.9, weight_decay=0.0001))
+
+# learning policy
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 1000,
+ by_epoch=False,
+ begin=0,
+ end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=6,
+ by_epoch=True,
+ milestones=[5],
+ gamma=0.1)
+]
+
+# train, val, test setting
+train_cfg = dict(type='EpochBasedTrainLoop', max_epochs=6, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/reid/reid_r50_8xb32-6e_mot20train80_test-mot20val20.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reid/reid_r50_8xb32-6e_mot20train80_test-mot20val20.py
new file mode 100644
index 0000000000000000000000000000000000000000..8a807996186c35f91e23f6e0ec95a2191479c15b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reid/reid_r50_8xb32-6e_mot20train80_test-mot20val20.py
@@ -0,0 +1,10 @@
+_base_ = ['./reid_r50_8xb32-6e_mot17train80_test-mot17val20.py']
+model = dict(head=dict(num_classes=1701))
+# data
+data_root = 'data/MOT20/'
+train_dataloader = dict(dataset=dict(data_root=data_root))
+val_dataloader = dict(dataset=dict(data_root=data_root))
+test_dataloader = val_dataloader
+
+# train, val, test setting
+train_cfg = dict(type='EpochBasedTrainLoop', max_epochs=6, val_interval=7)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..03cb86bef4e24298075d67b5acb4a2e30bafef7e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/README.md
@@ -0,0 +1,59 @@
+# RepPoints
+
+> [RepPoints: Point Set Representation for Object Detection](https://arxiv.org/abs/1904.11490)
+
+
+
+## Abstract
+
+Modern object detectors rely heavily on rectangular bounding boxes, such as anchors, proposals and the final predictions, to represent objects at various recognition stages. The bounding box is convenient to use but provides only a coarse localization of objects and leads to a correspondingly coarse extraction of object features. In this paper, we present RepPoints(representative points), a new finer representation of objects as a set of sample points useful for both localization and recognition. Given ground truth localization and recognition targets for training, RepPoints learn to automatically arrange themselves in a manner that bounds the spatial extent of an object and indicates semantically significant local areas. They furthermore do not require the use of anchors to sample a space of bounding boxes. We show that an anchor-free object detector based on RepPoints can be as effective as the state-of-the-art anchor-based detection methods, with 46.5 AP and 67.4 AP50 on the COCO test-dev detection benchmark, using ResNet-101 model.
+
+
+

+
+
+## Introdution
+
+By [Ze Yang](https://yangze.tech/), [Shaohui Liu](http://b1ueber2y.me/), and [Han Hu](https://ancientmooner.github.io/).
+
+We provide code support and configuration files to reproduce the results in the paper for
+["RepPoints: Point Set Representation for Object Detection"](https://arxiv.org/abs/1904.11490) on COCO object detection.
+
+**RepPoints**, initially described in [arXiv](https://arxiv.org/abs/1904.11490), is a new representation method for visual objects, on which visual understanding tasks are typically centered. Visual object representation, aiming at both geometric description and appearance feature extraction, is conventionally achieved by `bounding box + RoIPool (RoIAlign)`. The bounding box representation is convenient to use; however, it provides only a rectangular localization of objects that lacks geometric precision and may consequently degrade feature quality. Our new representation, RepPoints, models objects by a `point set` instead of a `bounding box`, which learns to adaptively position themselves over an object in a manner that circumscribes the object’s `spatial extent` and enables `semantically aligned feature extraction`. This richer and more flexible representation maintains the convenience of bounding boxes while facilitating various visual understanding applications. This repo demonstrated the effectiveness of RepPoints for COCO object detection.
+
+Another feature of this repo is the demonstration of an `anchor-free detector`, which can be as effective as state-of-the-art anchor-based detection methods. The anchor-free detector can utilize either `bounding box` or `RepPoints` as the basic object representation.
+
+## Results and Models
+
+The results on COCO 2017val are shown in the table below.
+
+| Method | Backbone | GN | Anchor | convert func | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :-------: | :-----------: | :-: | :----: | :----------: | :-----: | :------: | :------------: | :----: | :---------------------------------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| BBox | R-50-FPN | Y | single | - | 1x | 3.9 | 15.9 | 36.4 | [config](./reppoints-bbox_r50_fpn-gn_head-gn-grid_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/reppoints/bbox_r50_grid_fpn_gn-neck%2Bhead_1x_coco/bbox_r50_grid_fpn_gn-neck%2Bhead_1x_coco_20200329_145916-0eedf8d1.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/reppoints/bbox_r50_grid_fpn_gn-neck%2Bhead_1x_coco/bbox_r50_grid_fpn_gn-neck%2Bhead_1x_coco_20200329_145916.log.json) |
+| BBox | R-50-FPN | Y | none | - | 1x | 3.9 | 15.4 | 37.4 | [config](./reppoints-bbox_r50-center_fpn-gn_head-gn-grid_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/reppoints/bbox_r50_grid_fpn_gn-neck%2Bhead_1x_coco/bbox_r50_grid_fpn_gn-neck%2Bhead_1x_coco_20200329_145916-0eedf8d1.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/reppoints/bbox_r50_grid_fpn_gn-neck%2Bhead_1x_coco/bbox_r50_grid_fpn_gn-neck%2Bhead_1x_coco_20200329_145916.log.json) |
+| RepPoints | R-50-FPN | N | none | moment | 1x | 3.3 | 18.5 | 37.0 | [config](./reppoints-moment_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/reppoints/reppoints_moment_r50_fpn_1x_coco/reppoints_moment_r50_fpn_1x_coco_20200330-b73db8d1.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/reppoints/reppoints_moment_r50_fpn_1x_coco/reppoints_moment_r50_fpn_1x_coco_20200330_233609.log.json) |
+| RepPoints | R-50-FPN | Y | none | moment | 1x | 3.9 | 17.5 | 38.1 | [config](./reppoints-moment_r50_fpn-gn_head-gn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/reppoints/reppoints_moment_r50_fpn_gn-neck%2Bhead_1x_coco/reppoints_moment_r50_fpn_gn-neck%2Bhead_1x_coco_20200329_145952-3e51b550.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/reppoints/reppoints_moment_r50_fpn_gn-neck%2Bhead_1x_coco/reppoints_moment_r50_fpn_gn-neck%2Bhead_1x_coco_20200329_145952.log.json) |
+| RepPoints | R-50-FPN | Y | none | moment | 2x | 3.9 | - | 38.6 | [config](./reppoints-moment_r50_fpn-gn_head-gn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/reppoints/reppoints_moment_r50_fpn_gn-neck%2Bhead_2x_coco/reppoints_moment_r50_fpn_gn-neck%2Bhead_2x_coco_20200329-91babaa2.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/reppoints/reppoints_moment_r50_fpn_gn-neck%2Bhead_2x_coco/reppoints_moment_r50_fpn_gn-neck%2Bhead_2x_coco_20200329_150020.log.json) |
+| RepPoints | R-101-FPN | Y | none | moment | 2x | 5.8 | 13.7 | 40.5 | [config](./reppoints-moment_r101_fpn-gn_head-gn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/reppoints/reppoints_moment_r101_fpn_gn-neck%2Bhead_2x_coco/reppoints_moment_r101_fpn_gn-neck%2Bhead_2x_coco_20200329-4fbc7310.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/reppoints/reppoints_moment_r101_fpn_gn-neck%2Bhead_2x_coco/reppoints_moment_r101_fpn_gn-neck%2Bhead_2x_coco_20200329_132205.log.json) |
+| RepPoints | R-101-FPN-DCN | Y | none | moment | 2x | 5.9 | 12.1 | 42.9 | [config](./reppoints-moment_r101-dconv-c3-c5_fpn-gn_head-gn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/reppoints/reppoints_moment_r101_fpn_dconv_c3-c5_gn-neck%2Bhead_2x_coco/reppoints_moment_r101_fpn_dconv_c3-c5_gn-neck%2Bhead_2x_coco_20200329-3309fbf2.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/reppoints/reppoints_moment_r101_fpn_dconv_c3-c5_gn-neck%2Bhead_2x_coco/reppoints_moment_r101_fpn_dconv_c3-c5_gn-neck%2Bhead_2x_coco_20200329_132134.log.json) |
+| RepPoints | X-101-FPN-DCN | Y | none | moment | 2x | 7.1 | 9.3 | 44.2 | [config](./reppoints-moment_x101-dconv-c3-c5_fpn-gn_head-gn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/reppoints/reppoints_moment_x101_fpn_dconv_c3-c5_gn-neck%2Bhead_2x_coco/reppoints_moment_x101_fpn_dconv_c3-c5_gn-neck%2Bhead_2x_coco_20200329-f87da1ea.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/reppoints/reppoints_moment_x101_fpn_dconv_c3-c5_gn-neck%2Bhead_2x_coco/reppoints_moment_x101_fpn_dconv_c3-c5_gn-neck%2Bhead_2x_coco_20200329_132201.log.json) |
+
+**Notes:**
+
+- `R-xx`, `X-xx` denote the ResNet and ResNeXt architectures, respectively.
+- `DCN` denotes replacing 3x3 conv with the 3x3 deformable convolution in `c3-c5` stages of backbone.
+- `none` in the `anchor` column means 2-d `center point` (x,y) is used to represent the initial object hypothesis. `single` denotes one 4-d anchor box (x,y,w,h) with IoU based label assign criterion is adopted.
+- `moment`, `partial MinMax`, `MinMax` in the `convert func` column are three functions to convert a point set to a pseudo box.
+- Note the results here are slightly different from those reported in the paper, due to framework change. While the original paper uses an [MXNet](https://mxnet.apache.org/) implementation, we re-implement the method in [PyTorch](https://pytorch.org/) based on mmdetection.
+
+## Citation
+
+```latex
+@inproceedings{yang2019reppoints,
+ title={RepPoints: Point Set Representation for Object Detection},
+ author={Yang, Ze and Liu, Shaohui and Hu, Han and Wang, Liwei and Lin, Stephen},
+ booktitle={The IEEE International Conference on Computer Vision (ICCV)},
+ month={Oct},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..732d541fb548f6eed00d6ba0fb4ffe3854b4f9c5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/metafile.yml
@@ -0,0 +1,181 @@
+Collections:
+ - Name: RepPoints
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Group Normalization
+ - FPN
+ - RepPoints
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/1904.11490
+ Title: 'RepPoints: Point Set Representation for Object Detection'
+ README: configs/reppoints/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/detectors/reppoints_detector.py#L9
+ Version: v2.0.0
+
+Models:
+ - Name: reppoints-bbox_r50_fpn-gn_head-gn-grid_1x_coco
+ In Collection: RepPoints
+ Config: configs/reppoints/reppoints-bbox_r50_fpn-gn_head-gn-grid_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.9
+ inference time (ms/im):
+ - value: 62.89
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 36.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/reppoints/bbox_r50_grid_fpn_gn-neck%2Bhead_1x_coco/bbox_r50_grid_fpn_gn-neck%2Bhead_1x_coco_20200329_145916-0eedf8d1.pth
+
+ - Name: reppoints-bbox_r50-center_fpn-gn_head-gn-grid_1x_coco
+ In Collection: RepPoints
+ Config: configs/reppoints/reppoints-bbox_r50-center_fpn-gn_head-gn-grid_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.9
+ inference time (ms/im):
+ - value: 64.94
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/reppoints/bbox_r50_grid_fpn_gn-neck%2Bhead_1x_coco/bbox_r50_grid_fpn_gn-neck%2Bhead_1x_coco_20200329_145916-0eedf8d1.pth
+
+ - Name: reppoints-moment_r50_fpn_1x_coco
+ In Collection: RepPoints
+ Config: configs/reppoints/reppoints-moment_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.3
+ inference time (ms/im):
+ - value: 54.05
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/reppoints/reppoints_moment_r50_fpn_1x_coco/reppoints_moment_r50_fpn_1x_coco_20200330-b73db8d1.pth
+
+ - Name: reppoints-moment_r50_fpn-gn_head-gn_1x_coco
+ In Collection: RepPoints
+ Config: configs/reppoints/reppoints-moment_r50_fpn-gn_head-gn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.9
+ inference time (ms/im):
+ - value: 57.14
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/reppoints/reppoints_moment_r50_fpn_gn-neck%2Bhead_1x_coco/reppoints_moment_r50_fpn_gn-neck%2Bhead_1x_coco_20200329_145952-3e51b550.pth
+
+ - Name: reppoints-moment_r50_fpn-gn_head-gn_2x_coco
+ In Collection: RepPoints
+ Config: configs/reppoints/reppoints-moment_r50_fpn-gn_head-gn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 3.9
+ inference time (ms/im):
+ - value: 57.14
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/reppoints/reppoints_moment_r50_fpn_gn-neck%2Bhead_2x_coco/reppoints_moment_r50_fpn_gn-neck%2Bhead_2x_coco_20200329-91babaa2.pth
+
+ - Name: reppoints-moment_r101_fpn-gn_head-gn_2x_coco
+ In Collection: RepPoints
+ Config: configs/reppoints/reppoints-moment_r101_fpn-gn_head-gn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 5.8
+ inference time (ms/im):
+ - value: 72.99
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/reppoints/reppoints_moment_r101_fpn_gn-neck%2Bhead_2x_coco/reppoints_moment_r101_fpn_gn-neck%2Bhead_2x_coco_20200329-4fbc7310.pth
+
+ - Name: reppoints-moment_r101-dconv-c3-c5_fpn-gn_head-gn_2x_coco
+ In Collection: RepPoints
+ Config: configs/reppoints/reppoints-moment_r101-dconv-c3-c5_fpn-gn_head-gn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 5.9
+ inference time (ms/im):
+ - value: 82.64
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/reppoints/reppoints_moment_r101_fpn_dconv_c3-c5_gn-neck%2Bhead_2x_coco/reppoints_moment_r101_fpn_dconv_c3-c5_gn-neck%2Bhead_2x_coco_20200329-3309fbf2.pth
+
+ - Name: reppoints-moment_x101-dconv-c3-c5_fpn-gn_head-gn_2x_coco
+ In Collection: RepPoints
+ Config: configs/reppoints/reppoints-moment_x101-dconv-c3-c5_fpn-gn_head-gn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 7.1
+ inference time (ms/im):
+ - value: 107.53
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/reppoints/reppoints_moment_x101_fpn_dconv_c3-c5_gn-neck%2Bhead_2x_coco/reppoints_moment_x101_fpn_dconv_c3-c5_gn-neck%2Bhead_2x_coco_20200329-f87da1ea.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-bbox_r50-center_fpn-gn_head-gn-grid_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-bbox_r50-center_fpn-gn_head-gn-grid_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f116e53f6ded9468098733c1bab938831fee041d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-bbox_r50-center_fpn-gn_head-gn-grid_1x_coco.py
@@ -0,0 +1,2 @@
+_base_ = './reppoints-moment_r50_fpn-gn_head-gn_1x_coco.py'
+model = dict(bbox_head=dict(transform_method='minmax', use_grid_points=True))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-bbox_r50_fpn-gn_head-gn-grid_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-bbox_r50_fpn-gn_head-gn-grid_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..76be39b8de8f52d48c6cdd4626f23221e35164ab
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-bbox_r50_fpn-gn_head-gn-grid_1x_coco.py
@@ -0,0 +1,13 @@
+_base_ = './reppoints-moment_r50_fpn-gn_head-gn_1x_coco.py'
+model = dict(
+ bbox_head=dict(transform_method='minmax', use_grid_points=True),
+ # training and testing settings
+ train_cfg=dict(
+ init=dict(
+ assigner=dict(
+ _delete_=True,
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.4,
+ min_pos_iou=0,
+ ignore_iof_thr=-1))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-minmax_r50_fpn-gn_head-gn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-minmax_r50_fpn-gn_head-gn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..0e7dffe77a062268737205fd86ab23f22cd85479
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-minmax_r50_fpn-gn_head-gn_1x_coco.py
@@ -0,0 +1,2 @@
+_base_ = './reppoints-moment_r50_fpn-gn_head-gn_1x_coco.py'
+model = dict(bbox_head=dict(transform_method='minmax'))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-moment_r101-dconv-c3-c5_fpn-gn_head-gn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-moment_r101-dconv-c3-c5_fpn-gn_head-gn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5c2bfab40020d7508ba90029ad29b24da8a7ad78
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-moment_r101-dconv-c3-c5_fpn-gn_head-gn_2x_coco.py
@@ -0,0 +1,8 @@
+_base_ = './reppoints-moment_r50_fpn-gn_head-gn_2x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ dcn=dict(type='DCN', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True),
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-moment_r101_fpn-gn_head-gn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-moment_r101_fpn-gn_head-gn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..02c447ada075ca6b076a5e7ff2ed74fb3b80c30d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-moment_r101_fpn-gn_head-gn_2x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './reppoints-moment_r50_fpn-gn_head-gn_2x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-moment_r50_fpn-gn_head-gn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-moment_r50_fpn-gn_head-gn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..cedf2226b5ecd2e5dd207041523ab4a2627a1734
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-moment_r50_fpn-gn_head-gn_1x_coco.py
@@ -0,0 +1,3 @@
+_base_ = './reppoints-moment_r50_fpn_1x_coco.py'
+norm_cfg = dict(type='GN', num_groups=32, requires_grad=True)
+model = dict(neck=dict(norm_cfg=norm_cfg), bbox_head=dict(norm_cfg=norm_cfg))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-moment_r50_fpn-gn_head-gn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-moment_r50_fpn-gn_head-gn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..4490d4496af6d680fbed2eedcaf73e138afff0cc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-moment_r50_fpn-gn_head-gn_2x_coco.py
@@ -0,0 +1,17 @@
+_base_ = './reppoints-moment_r50_fpn-gn_head-gn_1x_coco.py'
+
+max_epochs = 24
+
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-moment_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-moment_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..df7e72a80c66f42fe8554cfb344fee87ee5fe24a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-moment_r50_fpn_1x_coco.py
@@ -0,0 +1,74 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ type='RepPointsDetector',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_input',
+ num_outs=5),
+ bbox_head=dict(
+ type='RepPointsHead',
+ num_classes=80,
+ in_channels=256,
+ feat_channels=256,
+ point_feat_channels=256,
+ stacked_convs=3,
+ num_points=9,
+ gradient_mul=0.1,
+ point_strides=[8, 16, 32, 64, 128],
+ point_base_scale=4,
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox_init=dict(type='SmoothL1Loss', beta=0.11, loss_weight=0.5),
+ loss_bbox_refine=dict(type='SmoothL1Loss', beta=0.11, loss_weight=1.0),
+ transform_method='moment'),
+ # training and testing settings
+ train_cfg=dict(
+ init=dict(
+ assigner=dict(type='PointAssigner', scale=4, pos_num=1),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ refine=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.4,
+ min_pos_iou=0,
+ ignore_iof_thr=-1),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False)),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.5),
+ max_per_img=100))
+
+optim_wrapper = dict(optimizer=dict(lr=0.01))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-moment_x101-dconv-c3-c5_fpn-gn_head-gn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-moment_x101-dconv-c3-c5_fpn-gn_head-gn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a9909efe511da9423859de6ce096b1b1524a9b6f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-moment_x101-dconv-c3-c5_fpn-gn_head-gn_2x_coco.py
@@ -0,0 +1,16 @@
+_base_ = './reppoints-moment_r50_fpn-gn_head-gn_2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ dcn=dict(type='DCN', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-partial-minmax_r50_fpn-gn_head-gn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-partial-minmax_r50_fpn-gn_head-gn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..30f7844b8344110896c5d885bd0ca340322045e4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints-partial-minmax_r50_fpn-gn_head-gn_1x_coco.py
@@ -0,0 +1,2 @@
+_base_ = './reppoints-moment_r50_fpn-gn_head-gn_1x_coco.py'
+model = dict(bbox_head=dict(transform_method='partial_minmax'))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints.png b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints.png
new file mode 100644
index 0000000000000000000000000000000000000000..16d491b9ec62835d91b474b7d69c46bd25da25e5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/reppoints/reppoints.png
@@ -0,0 +1,3 @@
+version https://git-lfs.github.com/spec/v1
+oid sha256:c8c4c485b83297b7972632a0fc8dbc2b27a3620afecbc7b42aaf2183e3f98f6b
+size 1198109
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..cd6732b60aff3d80eeb23f14a97657f57344a480
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/README.md
@@ -0,0 +1,77 @@
+# Res2Net
+
+> [Res2Net: A New Multi-scale Backbone Architecture](https://arxiv.org/abs/1904.01169)
+
+
+
+## Abstract
+
+Representing features at multiple scales is of great importance for numerous vision tasks. Recent advances in backbone convolutional neural networks (CNNs) continually demonstrate stronger multi-scale representation ability, leading to consistent performance gains on a wide range of applications. However, most existing methods represent the multi-scale features in a layer-wise manner. In this paper, we propose a novel building block for CNNs, namely Res2Net, by constructing hierarchical residual-like connections within one single residual block. The Res2Net represents multi-scale features at a granular level and increases the range of receptive fields for each network layer. The proposed Res2Net block can be plugged into the state-of-the-art backbone CNN models, e.g., ResNet, ResNeXt, and DLA. We evaluate the Res2Net block on all these models and demonstrate consistent performance gains over baseline models on widely-used datasets, e.g., CIFAR-100 and ImageNet. Further ablation studies and experimental results on representative computer vision tasks, i.e., object detection, class activation mapping, and salient object detection, further verify the superiority of the Res2Net over the state-of-the-art baseline methods.
+
+
+

+
+
+## Introduction
+
+We propose a novel building block for CNNs, namely Res2Net, by constructing hierarchical residual-like connections within one single residual block. The Res2Net represents multi-scale features at a granular level and increases the range of receptive fields for each network layer.
+
+| Backbone | Params. | GFLOPs | top-1 err. | top-5 err. |
+| :---------------: | :-----: | :----: | :--------: | :--------: |
+| ResNet-101 | 44.6 M | 7.8 | 22.63 | 6.44 |
+| ResNeXt-101-64x4d | 83.5M | 15.5 | 20.40 | - |
+| HRNetV2p-W48 | 77.5M | 16.1 | 20.70 | 5.50 |
+| Res2Net-101 | 45.2M | 8.3 | 18.77 | 4.64 |
+
+Compared with other backbone networks, Res2Net requires fewer parameters and FLOPs.
+
+**Note:**
+
+- GFLOPs for classification are calculated with image size (224x224).
+
+## Results and Models
+
+### Faster R-CNN
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :--------: | :-----: | :-----: | :------: | :------------: | :----: | :------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R2-101-FPN | pytorch | 2x | 7.4 | - | 43.0 | [config](./faster-rcnn_res2net-101_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/res2net/faster_rcnn_r2_101_fpn_2x_coco/faster_rcnn_r2_101_fpn_2x_coco-175f1da6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/res2net/faster_rcnn_r2_101_fpn_2x_coco/faster_rcnn_r2_101_fpn_2x_coco_20200514_231734.log.json) |
+
+### Mask R-CNN
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :--------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :----------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R2-101-FPN | pytorch | 2x | 7.9 | - | 43.6 | 38.7 | [config](./mask-rcnn_res2net-101_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/res2net/mask_rcnn_r2_101_fpn_2x_coco/mask_rcnn_r2_101_fpn_2x_coco-17f061e8.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/res2net/mask_rcnn_r2_101_fpn_2x_coco/mask_rcnn_r2_101_fpn_2x_coco_20200515_002413.log.json) |
+
+### Cascade R-CNN
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :--------: | :-----: | :-----: | :------: | :------------: | :----: | :--------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R2-101-FPN | pytorch | 20e | 7.8 | - | 45.7 | [config](./cascade-rcnn_res2net-101_fpn_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/res2net/cascade_rcnn_r2_101_fpn_20e_coco/cascade_rcnn_r2_101_fpn_20e_coco-f4b7b7db.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/res2net/cascade_rcnn_r2_101_fpn_20e_coco/cascade_rcnn_r2_101_fpn_20e_coco_20200515_091644.log.json) |
+
+### Cascade Mask R-CNN
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :--------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :-------------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R2-101-FPN | pytorch | 20e | 9.5 | - | 46.4 | 40.0 | [config](./cascade-mask-rcnn_res2net-101_fpn_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/res2net/cascade_mask_rcnn_r2_101_fpn_20e_coco/cascade_mask_rcnn_r2_101_fpn_20e_coco-8a7b41e1.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/res2net/cascade_mask_rcnn_r2_101_fpn_20e_coco/cascade_mask_rcnn_r2_101_fpn_20e_coco_20200515_091645.log.json) |
+
+### Hybrid Task Cascade (HTC)
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :--------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :-----------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R2-101-FPN | pytorch | 20e | - | - | 47.5 | 41.6 | [config](./htc_res2net-101_fpn_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/res2net/htc_r2_101_fpn_20e_coco/htc_r2_101_fpn_20e_coco-3a8d2112.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/res2net/htc_r2_101_fpn_20e_coco/htc_r2_101_fpn_20e_coco_20200515_150029.log.json) |
+
+- Res2Net ImageNet pretrained models are in [Res2Net-PretrainedModels](https://github.com/Res2Net/Res2Net-PretrainedModels).
+- More applications of Res2Net are in [Res2Net-Github](https://github.com/Res2Net/).
+
+## Citation
+
+```latex
+@article{gao2019res2net,
+ title={Res2Net: A New Multi-scale Backbone Architecture},
+ author={Gao, Shang-Hua and Cheng, Ming-Ming and Zhao, Kai and Zhang, Xin-Yu and Yang, Ming-Hsuan and Torr, Philip},
+ journal={IEEE TPAMI},
+ year={2020},
+ doi={10.1109/TPAMI.2019.2938758},
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/cascade-mask-rcnn_res2net-101_fpn_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/cascade-mask-rcnn_res2net-101_fpn_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..21b6d2ea1c0167b8dd643211b520ac89ddd63e10
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/cascade-mask-rcnn_res2net-101_fpn_20e_coco.py
@@ -0,0 +1,10 @@
+_base_ = '../cascade_rcnn/cascade-mask-rcnn_r50_fpn_20e_coco.py'
+model = dict(
+ backbone=dict(
+ type='Res2Net',
+ depth=101,
+ scales=4,
+ base_width=26,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://res2net101_v1d_26w_4s')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/cascade-rcnn_res2net-101_fpn_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/cascade-rcnn_res2net-101_fpn_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..670a77454e060f8f639dbdc40064b71cd82520e9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/cascade-rcnn_res2net-101_fpn_20e_coco.py
@@ -0,0 +1,10 @@
+_base_ = '../cascade_rcnn/cascade-rcnn_r50_fpn_20e_coco.py'
+model = dict(
+ backbone=dict(
+ type='Res2Net',
+ depth=101,
+ scales=4,
+ base_width=26,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://res2net101_v1d_26w_4s')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/faster-rcnn_res2net-101_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/faster-rcnn_res2net-101_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..033cf574962f51a75c3fce1e74a22efb9c6320f2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/faster-rcnn_res2net-101_fpn_2x_coco.py
@@ -0,0 +1,10 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='Res2Net',
+ depth=101,
+ scales=4,
+ base_width=26,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://res2net101_v1d_26w_4s')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/htc_res2net-101_fpn_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/htc_res2net-101_fpn_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d5542fda4c8181a417f14817180296e84944b832
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/htc_res2net-101_fpn_20e_coco.py
@@ -0,0 +1,10 @@
+_base_ = '../htc/htc_r50_fpn_20e_coco.py'
+model = dict(
+ backbone=dict(
+ type='Res2Net',
+ depth=101,
+ scales=4,
+ base_width=26,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://res2net101_v1d_26w_4s')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/mask-rcnn_res2net-101_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/mask-rcnn_res2net-101_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3a2d57304d07d9b3dbc58ee9a5d8f2355c6b4427
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/mask-rcnn_res2net-101_fpn_2x_coco.py
@@ -0,0 +1,10 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='Res2Net',
+ depth=101,
+ scales=4,
+ base_width=26,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://res2net101_v1d_26w_4s')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..1d9f9ea023d895cd8a93b0f48b3bc4dee5a93e6b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/res2net/metafile.yml
@@ -0,0 +1,146 @@
+Models:
+ - Name: faster-rcnn_res2net-101_fpn_2x_coco
+ In Collection: Faster R-CNN
+ Config: configs/res2net/faster-rcnn_res2net-101_fpn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 7.4
+ Epochs: 24
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Res2Net
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/res2net/faster_rcnn_r2_101_fpn_2x_coco/faster_rcnn_r2_101_fpn_2x_coco-175f1da6.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.01169
+ Title: 'Res2Net for object detection and instance segmentation'
+ README: configs/res2net/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/res2net.py#L239
+ Version: v2.1.0
+
+ - Name: mask-rcnn_res2net-101_fpn_2x_coco
+ In Collection: Mask R-CNN
+ Config: configs/res2net/mask-rcnn_res2net-101_fpn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 7.9
+ Epochs: 24
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Res2Net
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.6
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/res2net/mask_rcnn_r2_101_fpn_2x_coco/mask_rcnn_r2_101_fpn_2x_coco-17f061e8.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.01169
+ Title: 'Res2Net for object detection and instance segmentation'
+ README: configs/res2net/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/res2net.py#L239
+ Version: v2.1.0
+
+ - Name: cascade-rcnn_res2net-101_fpn_20e_coco
+ In Collection: Cascade R-CNN
+ Config: configs/res2net/cascade-rcnn_res2net-101_fpn_20e_coco.py
+ Metadata:
+ Training Memory (GB): 7.8
+ Epochs: 20
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Res2Net
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/res2net/cascade_rcnn_r2_101_fpn_20e_coco/cascade_rcnn_r2_101_fpn_20e_coco-f4b7b7db.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.01169
+ Title: 'Res2Net for object detection and instance segmentation'
+ README: configs/res2net/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/res2net.py#L239
+ Version: v2.1.0
+
+ - Name: cascade-mask-rcnn_res2net-101_fpn_20e_coco
+ In Collection: Cascade R-CNN
+ Config: configs/res2net/cascade-mask-rcnn_res2net-101_fpn_20e_coco.py
+ Metadata:
+ Training Memory (GB): 9.5
+ Epochs: 20
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Res2Net
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 40.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/res2net/cascade_mask_rcnn_r2_101_fpn_20e_coco/cascade_mask_rcnn_r2_101_fpn_20e_coco-8a7b41e1.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.01169
+ Title: 'Res2Net for object detection and instance segmentation'
+ README: configs/res2net/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/res2net.py#L239
+ Version: v2.1.0
+
+ - Name: htc_res2net-101_fpn_20e_coco
+ In Collection: HTC
+ Config: configs/res2net/htc_res2net-101_fpn_20e_coco.py
+ Metadata:
+ Epochs: 20
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Res2Net
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 47.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 41.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/res2net/htc_r2_101_fpn_20e_coco/htc_r2_101_fpn_20e_coco-3a8d2112.pth
+ Paper:
+ URL: https://arxiv.org/abs/1904.01169
+ Title: 'Res2Net for object detection and instance segmentation'
+ README: configs/res2net/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.1.0/mmdet/models/backbones/res2net.py#L239
+ Version: v2.1.0
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..a72f842357999af4bf48e0b26edd2581d01d7a80
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/README.md
@@ -0,0 +1,54 @@
+# ResNeSt
+
+> [ResNeSt: Split-Attention Networks](https://arxiv.org/abs/2004.08955)
+
+
+
+## Abstract
+
+It is well known that featuremap attention and multi-path representation are important for visual recognition. In this paper, we present a modularized architecture, which applies the channel-wise attention on different network branches to leverage their success in capturing cross-feature interactions and learning diverse representations. Our design results in a simple and unified computation block, which can be parameterized using only a few variables. Our model, named ResNeSt, outperforms EfficientNet in accuracy and latency trade-off on image classification. In addition, ResNeSt has achieved superior transfer learning results on several public benchmarks serving as the backbone, and has been adopted by the winning entries of COCO-LVIS challenge.
+
+
+

+
+
+## Results and Models
+
+### Faster R-CNN
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :-------: | :-----: | :-----: | :------: | :------------: | :----: | :-----------------------------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| S-50-FPN | pytorch | 1x | 4.8 | - | 42.0 | [config](./faster-rcnn_s50_fpn_syncbn-backbone+head_ms-range-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/resnest/faster_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco/faster_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco_20200926_125502-20289c16.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/resnest/faster_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco/faster_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco-20200926_125502.log.json) |
+| S-101-FPN | pytorch | 1x | 7.1 | - | 44.5 | [config](./faster-rcnn_s101_fpn_syncbn-backbone+head_ms-range-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/resnest/faster_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco/faster_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco_20201006_021058-421517f1.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/resnest/faster_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco/faster_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco-20201006_021058.log.json) |
+
+### Mask R-CNN
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :-------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :---------------------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| S-50-FPN | pytorch | 1x | 5.5 | - | 42.6 | 38.1 | [config](./mask-rcnn_s50_fpn_syncbn-backbone+head_ms-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/resnest/mask_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco/mask_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco_20200926_125503-8a2c3d47.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/resnest/mask_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco/mask_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco-20200926_125503.log.json) |
+| S-101-FPN | pytorch | 1x | 7.8 | - | 45.2 | 40.2 | [config](./mask-rcnn_s101_fpn_syncbn-backbone+head_ms-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/resnest/mask_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco/mask_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco_20201005_215831-af60cdf9.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/resnest/mask_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco/mask_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco-20201005_215831.log.json) |
+
+### Cascade R-CNN
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :-------: | :-----: | :-----: | :------: | :------------: | :----: | :------------------------------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| S-50-FPN | pytorch | 1x | - | - | 44.5 | [config](./cascade-rcnn_s50_fpn_syncbn-backbone+head_ms-range-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/resnest/cascade_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco/cascade_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco_20201122_213640-763cc7b5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/resnest/cascade_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco/cascade_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco-20201005_113242.log.json) |
+| S-101-FPN | pytorch | 1x | 8.4 | - | 46.8 | [config](./cascade-rcnn_s101_fpn_syncbn-backbone+head_ms-range-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/resnest/cascade_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco/cascade_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco_20201005_113242-b9459f8f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/resnest/cascade_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco/cascade_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco-20201122_213640.log.json) |
+
+### Cascade Mask R-CNN
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :-------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :-----------------------------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| S-50-FPN | pytorch | 1x | - | - | 45.4 | 39.5 | [config](./cascade-mask-rcnn_s50_fpn_syncbn-backbone+head_ms-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/resnest/cascade_mask_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco/cascade_mask_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco_20201122_104428-99eca4c7.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/resnest/cascade_mask_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco/cascade_mask_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco-20201122_104428.log.json) |
+| S-101-FPN | pytorch | 1x | 10.5 | - | 47.7 | 41.4 | [config](./cascade-mask-rcnn_s101_fpn_syncbn-backbone+head_ms-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/resnest/cascade_mask_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco/cascade_mask_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco_20201005_113243-42607475.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/resnest/cascade_mask_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco/cascade_mask_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco-20201005_113243.log.json) |
+
+## Citation
+
+```latex
+@article{zhang2020resnest,
+title={ResNeSt: Split-Attention Networks},
+author={Zhang, Hang and Wu, Chongruo and Zhang, Zhongyue and Zhu, Yi and Zhang, Zhi and Lin, Haibin and Sun, Yue and He, Tong and Muller, Jonas and Manmatha, R. and Li, Mu and Smola, Alexander},
+journal={arXiv preprint arXiv:2004.08955},
+year={2020}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/cascade-mask-rcnn_s101_fpn_syncbn-backbone+head_ms-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/cascade-mask-rcnn_s101_fpn_syncbn-backbone+head_ms-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f4f19925788acc357e9720513d4f388598927a70
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/cascade-mask-rcnn_s101_fpn_syncbn-backbone+head_ms-1x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './cascade-mask-rcnn_s50_fpn_syncbn-backbone+head_ms-1x_coco.py'
+model = dict(
+ backbone=dict(
+ stem_channels=128,
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='open-mmlab://resnest101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/cascade-mask-rcnn_s50_fpn_syncbn-backbone+head_ms-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/cascade-mask-rcnn_s50_fpn_syncbn-backbone+head_ms-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c6ef41c05cd97d19320c02fb065b0cde1dda54d7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/cascade-mask-rcnn_s50_fpn_syncbn-backbone+head_ms-1x_coco.py
@@ -0,0 +1,101 @@
+_base_ = '../cascade_rcnn/cascade-mask-rcnn_r50_fpn_1x_coco.py'
+norm_cfg = dict(type='SyncBN', requires_grad=True)
+
+model = dict(
+ # use ResNeSt img_norm
+ data_preprocessor=dict(
+ mean=[123.68, 116.779, 103.939],
+ std=[58.393, 57.12, 57.375],
+ bgr_to_rgb=True),
+ backbone=dict(
+ type='ResNeSt',
+ stem_channels=64,
+ depth=50,
+ radix=2,
+ reduction_factor=4,
+ avg_down_stride=True,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=norm_cfg,
+ norm_eval=False,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='open-mmlab://resnest50')),
+ roi_head=dict(
+ bbox_head=[
+ dict(
+ type='Shared4Conv1FCBBoxHead',
+ in_channels=256,
+ conv_out_channels=256,
+ fc_out_channels=1024,
+ norm_cfg=norm_cfg,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0,
+ loss_weight=1.0)),
+ dict(
+ type='Shared4Conv1FCBBoxHead',
+ in_channels=256,
+ conv_out_channels=256,
+ fc_out_channels=1024,
+ norm_cfg=norm_cfg,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.05, 0.05, 0.1, 0.1]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0,
+ loss_weight=1.0)),
+ dict(
+ type='Shared4Conv1FCBBoxHead',
+ in_channels=256,
+ conv_out_channels=256,
+ fc_out_channels=1024,
+ norm_cfg=norm_cfg,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.033, 0.033, 0.067, 0.067]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0))
+ ],
+ mask_head=dict(norm_cfg=norm_cfg)))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ poly2mask=False),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/cascade-rcnn_s101_fpn_syncbn-backbone+head_ms-range-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/cascade-rcnn_s101_fpn_syncbn-backbone+head_ms-range-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..9dbf3fae5ffb9382b053852c35e263f109668020
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/cascade-rcnn_s101_fpn_syncbn-backbone+head_ms-range-1x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './cascade-rcnn_s50_fpn_syncbn-backbone+head_ms-range-1x_coco.py'
+model = dict(
+ backbone=dict(
+ stem_channels=128,
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='open-mmlab://resnest101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/cascade-rcnn_s50_fpn_syncbn-backbone+head_ms-range-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/cascade-rcnn_s50_fpn_syncbn-backbone+head_ms-range-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..7ce7b56320a6511376237710c25061edd44b17dd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/cascade-rcnn_s50_fpn_syncbn-backbone+head_ms-range-1x_coco.py
@@ -0,0 +1,93 @@
+_base_ = '../cascade_rcnn/cascade-rcnn_r50_fpn_1x_coco.py'
+norm_cfg = dict(type='SyncBN', requires_grad=True)
+model = dict(
+ # use ResNeSt img_norm
+ data_preprocessor=dict(
+ mean=[123.68, 116.779, 103.939],
+ std=[58.393, 57.12, 57.375],
+ bgr_to_rgb=True),
+ backbone=dict(
+ type='ResNeSt',
+ stem_channels=64,
+ depth=50,
+ radix=2,
+ reduction_factor=4,
+ avg_down_stride=True,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=norm_cfg,
+ norm_eval=False,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='open-mmlab://resnest50')),
+ roi_head=dict(
+ bbox_head=[
+ dict(
+ type='Shared4Conv1FCBBoxHead',
+ in_channels=256,
+ conv_out_channels=256,
+ fc_out_channels=1024,
+ norm_cfg=norm_cfg,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0,
+ loss_weight=1.0)),
+ dict(
+ type='Shared4Conv1FCBBoxHead',
+ in_channels=256,
+ conv_out_channels=256,
+ fc_out_channels=1024,
+ norm_cfg=norm_cfg,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.05, 0.05, 0.1, 0.1]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0,
+ loss_weight=1.0)),
+ dict(
+ type='Shared4Conv1FCBBoxHead',
+ in_channels=256,
+ conv_out_channels=256,
+ fc_out_channels=1024,
+ norm_cfg=norm_cfg,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.033, 0.033, 0.067, 0.067]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0))
+ ], ))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize', scale=[(1333, 640), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/faster-rcnn_s101_fpn_syncbn-backbone+head_ms-range-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/faster-rcnn_s101_fpn_syncbn-backbone+head_ms-range-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f1e16321adff643d593268f868c09f5a318e7e93
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/faster-rcnn_s101_fpn_syncbn-backbone+head_ms-range-1x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './faster-rcnn_s50_fpn_syncbn-backbone+head_ms-range-1x_coco.py'
+model = dict(
+ backbone=dict(
+ stem_channels=128,
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='open-mmlab://resnest101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/faster-rcnn_s50_fpn_syncbn-backbone+head_ms-range-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/faster-rcnn_s50_fpn_syncbn-backbone+head_ms-range-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8f0ec6e07af1fcd250171cb769252eeb03f92da8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/faster-rcnn_s50_fpn_syncbn-backbone+head_ms-range-1x_coco.py
@@ -0,0 +1,39 @@
+_base_ = '../faster_rcnn/faster-rcnn_r50_fpn_1x_coco.py'
+norm_cfg = dict(type='SyncBN', requires_grad=True)
+model = dict(
+ # use ResNeSt img_norm
+ data_preprocessor=dict(
+ mean=[123.68, 116.779, 103.939],
+ std=[58.393, 57.12, 57.375],
+ bgr_to_rgb=True),
+ backbone=dict(
+ type='ResNeSt',
+ stem_channels=64,
+ depth=50,
+ radix=2,
+ reduction_factor=4,
+ avg_down_stride=True,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=norm_cfg,
+ norm_eval=False,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='open-mmlab://resnest50')),
+ roi_head=dict(
+ bbox_head=dict(
+ type='Shared4Conv1FCBBoxHead',
+ conv_out_channels=256,
+ norm_cfg=norm_cfg)))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize', scale=[(1333, 640), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/mask-rcnn_s101_fpn_syncbn-backbone+head_ms-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/mask-rcnn_s101_fpn_syncbn-backbone+head_ms-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3edf49f052f1f3c875cca2c061276cc1aca77604
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/mask-rcnn_s101_fpn_syncbn-backbone+head_ms-1x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './mask-rcnn_s50_fpn_syncbn-backbone+head_ms-1x_coco.py'
+model = dict(
+ backbone=dict(
+ stem_channels=128,
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='open-mmlab://resnest101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/mask-rcnn_s50_fpn_syncbn-backbone+head_ms-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/mask-rcnn_s50_fpn_syncbn-backbone+head_ms-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c6f27000862d74e23a665f3bf8caae0ec4a3d6f5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/mask-rcnn_s50_fpn_syncbn-backbone+head_ms-1x_coco.py
@@ -0,0 +1,46 @@
+_base_ = '../mask_rcnn/mask-rcnn_r50_fpn_1x_coco.py'
+norm_cfg = dict(type='SyncBN', requires_grad=True)
+model = dict(
+ # use ResNeSt img_norm
+ data_preprocessor=dict(
+ mean=[123.68, 116.779, 103.939],
+ std=[58.393, 57.12, 57.375],
+ bgr_to_rgb=True),
+ backbone=dict(
+ type='ResNeSt',
+ stem_channels=64,
+ depth=50,
+ radix=2,
+ reduction_factor=4,
+ avg_down_stride=True,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=norm_cfg,
+ norm_eval=False,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='open-mmlab://resnest50')),
+ roi_head=dict(
+ bbox_head=dict(
+ type='Shared4Conv1FCBBoxHead',
+ conv_out_channels=256,
+ norm_cfg=norm_cfg),
+ mask_head=dict(norm_cfg=norm_cfg)))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ poly2mask=False),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..265c94094975858ff0cc0ceac3870c9b4f9b9a84
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnest/metafile.yml
@@ -0,0 +1,230 @@
+Models:
+ - Name: faster-rcnn_s50_fpn_syncbn-backbone+head_ms-range-1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/resnest/faster-rcnn_s50_fpn_syncbn-backbone+head_ms-range-1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.8
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNeSt
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/resnest/faster_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco/faster_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco_20200926_125502-20289c16.pth
+ Paper:
+ URL: https://arxiv.org/abs/2004.08955
+ Title: 'ResNeSt: Split-Attention Networks'
+ README: configs/resnest/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.7.0/mmdet/models/backbones/resnest.py#L273
+ Version: v2.7.0
+
+ - Name: faster-rcnn_s101_fpn_syncbn-backbone+head_ms-range-1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/resnest/faster-rcnn_s101_fpn_syncbn-backbone+head_ms-range-1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.1
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNeSt
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/resnest/faster_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco/faster_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco_20201006_021058-421517f1.pth
+ Paper:
+ URL: https://arxiv.org/abs/2004.08955
+ Title: 'ResNeSt: Split-Attention Networks'
+ README: configs/resnest/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.7.0/mmdet/models/backbones/resnest.py#L273
+ Version: v2.7.0
+
+ - Name: mask-rcnn_s50_fpn_syncbn-backbone+head_ms-1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/resnest/mask-rcnn_s50_fpn_syncbn-backbone+head_ms-1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.5
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNeSt
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.6
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/resnest/mask_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco/mask_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco_20200926_125503-8a2c3d47.pth
+ Paper:
+ URL: https://arxiv.org/abs/2004.08955
+ Title: 'ResNeSt: Split-Attention Networks'
+ README: configs/resnest/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.7.0/mmdet/models/backbones/resnest.py#L273
+ Version: v2.7.0
+
+ - Name: mask-rcnn_s101_fpn_syncbn-backbone+head_ms-1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/resnest/mask-rcnn_s101_fpn_syncbn-backbone+head_ms-1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.8
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNeSt
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 40.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/resnest/mask_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco/mask_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco_20201005_215831-af60cdf9.pth
+ Paper:
+ URL: https://arxiv.org/abs/2004.08955
+ Title: 'ResNeSt: Split-Attention Networks'
+ README: configs/resnest/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.7.0/mmdet/models/backbones/resnest.py#L273
+ Version: v2.7.0
+
+ - Name: cascade-rcnn_s50_fpn_syncbn-backbone+head_ms-range-1x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/resnest/cascade-rcnn_s50_fpn_syncbn-backbone+head_ms-range-1x_coco.py
+ Metadata:
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNeSt
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/resnest/cascade_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco/cascade_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco_20201122_213640-763cc7b5.pth
+ Paper:
+ URL: https://arxiv.org/abs/2004.08955
+ Title: 'ResNeSt: Split-Attention Networks'
+ README: configs/resnest/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.7.0/mmdet/models/backbones/resnest.py#L273
+ Version: v2.7.0
+
+ - Name: cascade-rcnn_s101_fpn_syncbn-backbone+head_ms-range-1x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/resnest/cascade-rcnn_s101_fpn_syncbn-backbone+head_ms-range-1x_coco.py
+ Metadata:
+ Training Memory (GB): 8.4
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNeSt
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/resnest/cascade_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco/cascade_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain-range_1x_coco_20201005_113242-b9459f8f.pth
+ Paper:
+ URL: https://arxiv.org/abs/2004.08955
+ Title: 'ResNeSt: Split-Attention Networks'
+ README: configs/resnest/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.7.0/mmdet/models/backbones/resnest.py#L273
+ Version: v2.7.0
+
+ - Name: cascade-mask-rcnn_s50_fpn_syncbn-backbone+head_ms-1x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/resnest/cascade-mask-rcnn_s50_fpn_syncbn-backbone+head_ms-1x_coco.py
+ Metadata:
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNeSt
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/resnest/cascade_mask_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco/cascade_mask_rcnn_s50_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco_20201122_104428-99eca4c7.pth
+ Paper:
+ URL: https://arxiv.org/abs/2004.08955
+ Title: 'ResNeSt: Split-Attention Networks'
+ README: configs/resnest/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.7.0/mmdet/models/backbones/resnest.py#L273
+ Version: v2.7.0
+
+ - Name: cascade-mask-rcnn_s101_fpn_syncbn-backbone+head_ms-1x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/resnest/cascade-mask-rcnn_s101_fpn_syncbn-backbone+head_ms-1x_coco.py
+ Metadata:
+ Training Memory (GB): 10.5
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNeSt
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 47.7
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 41.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/resnest/cascade_mask_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco/cascade_mask_rcnn_s101_fpn_syncbn-backbone%2Bhead_mstrain_1x_coco_20201005_113243-42607475.pth
+ Paper:
+ URL: https://arxiv.org/abs/2004.08955
+ Title: 'ResNeSt: Split-Attention Networks'
+ README: configs/resnest/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.7.0/mmdet/models/backbones/resnest.py#L273
+ Version: v2.7.0
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnet_strikes_back/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnet_strikes_back/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..f015729a8d4ae4d78a909185a9b93b619e0f0f04
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnet_strikes_back/README.md
@@ -0,0 +1,40 @@
+# ResNet strikes back
+
+> [ResNet strikes back: An improved training procedure in timm](https://arxiv.org/abs/2110.00476)
+
+
+
+## Abstract
+
+The influential Residual Networks designed by He et al. remain the gold-standard architecture in numerous scientific publications. They typically serve as the default architecture in studies, or as baselines when new architectures are proposed. Yet there has been significant progress on best practices for training neural networks since the inception of the ResNet architecture in 2015. Novel optimization & dataaugmentation have increased the effectiveness of the training recipes.
+
+In this paper, we re-evaluate the performance of the vanilla ResNet-50 when trained with a procedure that integrates such advances. We share competitive training settings and pre-trained models in the timm open-source library, with the hope that they will serve as better baselines for future work. For instance, with our more demanding training setting, a vanilla ResNet-50 reaches 80.4% top-1 accuracy at resolution 224×224 on ImageNet-val without extra data or distillation. We also report the performance achieved with popular models with our training procedure.
+
+
+

+
+
+## Results and Models
+
+| Method | Backbone | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :----------------: | :------: | :-----: | :------: | :------------: | :---------: | :---------: | :------------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Faster R-CNN | R-50 rsb | 1x | 3.9 | - | 40.8 (+3.4) | - | [Config](./faster-rcnn_r50-rsb-pre_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/resnet_strikes_back/faster_rcnn_r50_fpn_rsb-pretrain_1x_coco/faster_rcnn_r50_fpn_rsb-pretrain_1x_coco_20220113_162229-32ae82a9.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/resnet_strikes_back/faster_rcnn_r50_fpn_rsb-pretrain_1x_coco/faster_rcnn_r50_fpn_rsb-pretrain_1x_coco_20220113_162229.log.json) |
+| Mask R-CNN | R-50 rsb | 1x | 4.5 | - | 41.2 (+3.0) | 38.2 (+3.0) | [Config](./mask-rcnn_r50-rsb-pre_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/resnet_strikes_back/mask_rcnn_r50_fpn_rsb-pretrain_1x_coco/mask_rcnn_r50_fpn_rsb-pretrain_1x_coco_20220113_174054-06ce8ba0.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/resnet_strikes_back/mask_rcnn_r50_fpn_rsb-pretrain_1x_coco/mask_rcnn_r50_fpn_rsb-pretrain_1x_coco_20220113_174054.log.json) |
+| Cascade Mask R-CNN | R-50 rsb | 1x | 6.2 | - | 44.8 (+3.6) | 39.9 (+3.6) | [Config](./cascade-mask-rcnn_r50-rsb-pre_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/resnet_strikes_back/cascade_mask_rcnn_r50_fpn_rsb-pretrain_1x_coco/cascade_mask_rcnn_r50_fpn_rsb-pretrain_1x_coco_20220113_193636-8b9ad50f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/resnet_strikes_back/cascade_mask_rcnn_r50_fpn_rsb-pretrain_1x_coco/cascade_mask_rcnn_r50_fpn_rsb-pretrain_1x_coco_20220113_193636.log.json) |
+| RetinaNet | R-50 rsb | 1x | 3.8 | - | 39.0 (+2.5) | - | [Config](./retinanet_r50-rsb-pre_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/resnet_strikes_back/retinanet_r50_fpn_rsb-pretrain_1x_coco/retinanet_r50_fpn_rsb-pretrain_1x_coco_20220113_175432-bd24aae9.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/resnet_strikes_back/retinanet_r50_fpn_rsb-pretrain_1x_coco/retinanet_r50_fpn_rsb-pretrain_1x_coco_20220113_175432.log.json) |
+
+**Notes:**
+
+- 'rsb' is short for 'resnet strikes back'
+- We have done some grid searches on learning rate and weight decay and get these optimal hyper-parameters.
+
+## Citation
+
+```latex
+@article{wightman2021resnet,
+title={Resnet strikes back: An improved training procedure in timm},
+author={Ross Wightman, Hugo Touvron, Hervé Jégou},
+journal={arXiv preprint arXiv:2110.00476},
+year={2021}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnet_strikes_back/cascade-mask-rcnn_r50-rsb-pre_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnet_strikes_back/cascade-mask-rcnn_r50-rsb-pre_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..de7b95b0863d1ea89382fd9fa5852eccf0f34150
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnet_strikes_back/cascade-mask-rcnn_r50-rsb-pre_fpn_1x_coco.py
@@ -0,0 +1,15 @@
+_base_ = [
+ '../_base_/models/cascade-mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+checkpoint = 'https://download.openmmlab.com/mmclassification/v0/resnet/resnet50_8xb256-rsb-a1-600e_in1k_20211228-20e21305.pth' # noqa
+model = dict(
+ backbone=dict(
+ init_cfg=dict(
+ type='Pretrained', prefix='backbone.', checkpoint=checkpoint)))
+
+optim_wrapper = dict(
+ optimizer=dict(_delete_=True, type='AdamW', lr=0.0002, weight_decay=0.05),
+ paramwise_cfg=dict(norm_decay_mult=0., bypass_duplicate=True))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnet_strikes_back/faster-rcnn_r50-rsb-pre_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnet_strikes_back/faster-rcnn_r50-rsb-pre_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8c60f66a7ba8e5b6a7ee6af06e771b3c6ad71f6c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnet_strikes_back/faster-rcnn_r50-rsb-pre_fpn_1x_coco.py
@@ -0,0 +1,15 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+checkpoint = 'https://download.openmmlab.com/mmclassification/v0/resnet/resnet50_8xb256-rsb-a1-600e_in1k_20211228-20e21305.pth' # noqa
+model = dict(
+ backbone=dict(
+ init_cfg=dict(
+ type='Pretrained', prefix='backbone.', checkpoint=checkpoint)))
+
+optim_wrapper = dict(
+ optimizer=dict(_delete_=True, type='AdamW', lr=0.0002, weight_decay=0.05),
+ paramwise_cfg=dict(norm_decay_mult=0., bypass_duplicate=True))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnet_strikes_back/mask-rcnn_r50-rsb-pre_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnet_strikes_back/mask-rcnn_r50-rsb-pre_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..85e25d392359b1a7811fb0c933ede5edacbfb9c3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnet_strikes_back/mask-rcnn_r50-rsb-pre_fpn_1x_coco.py
@@ -0,0 +1,15 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+checkpoint = 'https://download.openmmlab.com/mmclassification/v0/resnet/resnet50_8xb256-rsb-a1-600e_in1k_20211228-20e21305.pth' # noqa
+model = dict(
+ backbone=dict(
+ init_cfg=dict(
+ type='Pretrained', prefix='backbone.', checkpoint=checkpoint)))
+
+optim_wrapper = dict(
+ optimizer=dict(_delete_=True, type='AdamW', lr=0.0002, weight_decay=0.05),
+ paramwise_cfg=dict(norm_decay_mult=0., bypass_duplicate=True))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnet_strikes_back/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnet_strikes_back/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..74b152107d7a6d96f671c52d5273c79751122bfa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnet_strikes_back/metafile.yml
@@ -0,0 +1,116 @@
+Models:
+ - Name: faster-rcnn_r50_fpn_rsb-pretrain_1x_coco
+ In Collection: Faster R-CNN
+ Config: configs/resnet_strikes_back/faster-rcnn_r50-rsb-pre_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.9
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/resnet_strikes_back/faster_rcnn_r50_fpn_rsb-pretrain_1x_coco/faster_rcnn_r50_fpn_rsb-pretrain_1x_coco_20220113_162229-32ae82a9.pth
+ Paper:
+ URL: https://arxiv.org/abs/2110.00476
+ Title: 'ResNet strikes back: An improved training procedure in timm'
+ README: configs/resnet_strikes_back/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.22.0/configs/resnet_strikes_back/README.md
+ Version: v2.22.0
+
+ - Name: cascade-mask-rcnn_r50_fpn_rsb-pretrain_1x_coco
+ In Collection: Cascade R-CNN
+ Config: configs/resnet_strikes_back/cascade-mask-rcnn_r50-rsb-pre_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 6.2
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/resnet_strikes_back/cascade_mask_rcnn_r50_fpn_rsb-pretrain_1x_coco/cascade_mask_rcnn_r50_fpn_rsb-pretrain_1x_coco_20220113_193636-8b9ad50f.pth
+ Paper:
+ URL: https://arxiv.org/abs/2110.00476
+ Title: 'ResNet strikes back: An improved training procedure in timm'
+ README: configs/resnet_strikes_back/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.22.0/configs/resnet_strikes_back/README.md
+ Version: v2.22.0
+
+ - Name: retinanet_r50-rsb-pre_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/resnet_strikes_back/retinanet_r50-rsb-pre_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.8
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/resnet_strikes_back/retinanet_r50_fpn_rsb-pretrain_1x_coco/retinanet_r50_fpn_rsb-pretrain_1x_coco_20220113_175432-bd24aae9.pth
+ Paper:
+ URL: https://arxiv.org/abs/2110.00476
+ Title: 'ResNet strikes back: An improved training procedure in timm'
+ README: configs/resnet_strikes_back/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.22.0/configs/resnet_strikes_back/README.md
+ Version: v2.22.0
+
+ - Name: mask-rcnn_r50_fpn_rsb-pretrain_1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/resnet_strikes_back/mask-rcnn_r50-rsb-pre_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.5
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNet
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/resnet_strikes_back/mask_rcnn_r50_fpn_rsb-pretrain_1x_coco/mask_rcnn_r50_fpn_rsb-pretrain_1x_coco_20220113_174054-06ce8ba0.pth
+ Paper:
+ URL: https://arxiv.org/abs/2110.00476
+ Title: 'ResNet strikes back: An improved training procedure in timm'
+ README: configs/resnet_strikes_back/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.22.0/configs/resnet_strikes_back/README.md
+ Version: v2.22.0
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnet_strikes_back/retinanet_r50-rsb-pre_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnet_strikes_back/retinanet_r50-rsb-pre_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..7ce7bfd87d6b41a36acc4ff207695e38ef89700c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/resnet_strikes_back/retinanet_r50-rsb-pre_fpn_1x_coco.py
@@ -0,0 +1,15 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+checkpoint = 'https://download.openmmlab.com/mmclassification/v0/resnet/resnet50_8xb256-rsb-a1-600e_in1k_20211228-20e21305.pth' # noqa
+model = dict(
+ backbone=dict(
+ init_cfg=dict(
+ type='Pretrained', prefix='backbone.', checkpoint=checkpoint)))
+
+optim_wrapper = dict(
+ optimizer=dict(_delete_=True, type='AdamW', lr=0.0001, weight_decay=0.05),
+ paramwise_cfg=dict(norm_decay_mult=0., bypass_duplicate=True))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..b38335a3ce3585918cd45f70a18a2c703d201e9b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/README.md
@@ -0,0 +1,53 @@
+# RetinaNet
+
+> [Focal Loss for Dense Object Detection](https://arxiv.org/abs/1708.02002)
+
+
+
+## Abstract
+
+The highest accuracy object detectors to date are based on a two-stage approach popularized by R-CNN, where a classifier is applied to a sparse set of candidate object locations. In contrast, one-stage detectors that are applied over a regular, dense sampling of possible object locations have the potential to be faster and simpler, but have trailed the accuracy of two-stage detectors thus far. In this paper, we investigate why this is the case. We discover that the extreme foreground-background class imbalance encountered during training of dense detectors is the central cause. We propose to address this class imbalance by reshaping the standard cross entropy loss such that it down-weights the loss assigned to well-classified examples. Our novel Focal Loss focuses training on a sparse set of hard examples and prevents the vast number of easy negatives from overwhelming the detector during training. To evaluate the effectiveness of our loss, we design and train a simple dense detector we call RetinaNet. Our results show that when trained with the focal loss, RetinaNet is able to match the speed of previous one-stage detectors while surpassing the accuracy of all existing state-of-the-art two-stage detectors.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :-------------: | :-----: | :----------: | :------: | :------------: | :----: | :---------------------------------------------: | :-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-18-FPN | pytorch | 1x | 1.7 | | 31.7 | [config](./retinanet_r18_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r18_fpn_1x_coco/retinanet_r18_fpn_1x_coco_20220407_171055-614fd399.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r18_fpn_1x_coco/retinanet_r18_fpn_1x_coco_20220407_171055.log.json) |
+| R-18-FPN | pytorch | 1x(1 x 8 BS) | 5.0 | | 31.7 | [config](./retinanet_r18_fpn_1xb8-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r18_fpn_1x8_1x_coco/retinanet_r18_fpn_1x8_1x_coco_20220407_171255-4ea310d7.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r18_fpn_1x8_1x_coco/retinanet_r18_fpn_1x8_1x_coco_20220407_171255.log.json) |
+| R-50-FPN | caffe | 1x | 3.5 | 18.6 | 36.3 | [config](./retinanet_r50-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r50_caffe_fpn_1x_coco/retinanet_r50_caffe_fpn_1x_coco_20200531-f11027c5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r50_caffe_fpn_1x_coco/retinanet_r50_caffe_fpn_1x_coco_20200531_012518.log.json) |
+| R-50-FPN | pytorch | 1x | 3.8 | 19.0 | 36.5 | [config](./retinanet_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r50_fpn_1x_coco/retinanet_r50_fpn_1x_coco_20200130-c2398f9e.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r50_fpn_1x_coco/retinanet_r50_fpn_1x_coco_20200130_002941.log.json) |
+| R-50-FPN (FP16) | pytorch | 1x | 2.8 | 31.6 | 36.4 | [config](./retinanet_r50_fpn_amp-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/fp16/retinanet_r50_fpn_fp16_1x_coco/retinanet_r50_fpn_fp16_1x_coco_20200702-0dbfb212.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/fp16/retinanet_r50_fpn_fp16_1x_coco/retinanet_r50_fpn_fp16_1x_coco_20200702_020127.log.json) |
+| R-50-FPN | pytorch | 2x | - | - | 37.4 | [config](./retinanet_r50_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r50_fpn_2x_coco/retinanet_r50_fpn_2x_coco_20200131-fdb43119.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r50_fpn_2x_coco/retinanet_r50_fpn_2x_coco_20200131_114738.log.json) |
+| R-101-FPN | caffe | 1x | 5.5 | 14.7 | 38.5 | [config](./retinanet_r101-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r101_caffe_fpn_1x_coco/retinanet_r101_caffe_fpn_1x_coco_20200531-b428fa0f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r101_caffe_fpn_1x_coco/retinanet_r101_caffe_fpn_1x_coco_20200531_012536.log.json) |
+| R-101-FPN | pytorch | 1x | 5.7 | 15.0 | 38.5 | [config](./retinanet_r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r101_fpn_1x_coco/retinanet_r101_fpn_1x_coco_20200130-7a93545f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r101_fpn_1x_coco/retinanet_r101_fpn_1x_coco_20200130_003055.log.json) |
+| R-101-FPN | pytorch | 2x | - | - | 38.9 | [config](./retinanet_r101_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r101_fpn_2x_coco/retinanet_r101_fpn_2x_coco_20200131-5560aee8.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r101_fpn_2x_coco/retinanet_r101_fpn_2x_coco_20200131_114859.log.json) |
+| X-101-32x4d-FPN | pytorch | 1x | 7.0 | 12.1 | 39.9 | [config](./retinanet_x101-32x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_x101_32x4d_fpn_1x_coco/retinanet_x101_32x4d_fpn_1x_coco_20200130-5c8b7ec4.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_x101_32x4d_fpn_1x_coco/retinanet_x101_32x4d_fpn_1x_coco_20200130_003004.log.json) |
+| X-101-32x4d-FPN | pytorch | 2x | - | - | 40.1 | [config](./retinanet_x101-32x4d_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_x101_32x4d_fpn_2x_coco/retinanet_x101_32x4d_fpn_2x_coco_20200131-237fc5e1.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_x101_32x4d_fpn_2x_coco/retinanet_x101_32x4d_fpn_2x_coco_20200131_114812.log.json) |
+| X-101-64x4d-FPN | pytorch | 1x | 10.0 | 8.7 | 41.0 | [config](./retinanet_x101-64x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_x101_64x4d_fpn_1x_coco/retinanet_x101_64x4d_fpn_1x_coco_20200130-366f5af1.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_x101_64x4d_fpn_1x_coco/retinanet_x101_64x4d_fpn_1x_coco_20200130_003008.log.json) |
+| X-101-64x4d-FPN | pytorch | 2x | - | - | 40.8 | [config](./retinanet_x101-64x4d_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_x101_64x4d_fpn_2x_coco/retinanet_x101_64x4d_fpn_2x_coco_20200131-bca068ab.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_x101_64x4d_fpn_2x_coco/retinanet_x101_64x4d_fpn_2x_coco_20200131_114833.log.json) |
+
+## Pre-trained Models
+
+We also train some models with longer schedules and multi-scale training. The users could finetune them for downstream tasks.
+
+| Backbone | Style | Lr schd | Mem (GB) | box AP | Config | Download |
+| :-------------: | :-----: | :-----: | :------: | :----: | :--------------------------------------------------------: | :-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | pytorch | 3x | 3.5 | 39.5 | [config](./retinanet_r50_fpn_ms-640-800-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r50_fpn_mstrain_3x_coco/retinanet_r50_fpn_mstrain_3x_coco_20210718_220633-88476508.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r50_fpn_mstrain_3x_coco/retinanet_r50_fpn_mstrain_3x_coco_20210718_220633-88476508.log.json) |
+| R-101-FPN | caffe | 3x | 5.4 | 40.7 | [config](./retinanet_r101-caffe_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r101_caffe_fpn_mstrain_3x_coco/retinanet_r101_caffe_fpn_mstrain_3x_coco_20210721_063439-88a8a944.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r101_caffe_fpn_mstrain_3x_coco/retinanet_r101_caffe_fpn_mstrain_3x_coco_20210721_063439-88a8a944.log.json) |
+| R-101-FPN | pytorch | 3x | 5.4 | 41 | [config](./retinanet_r101_fpn_ms-640-800-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r101_fpn_mstrain_3x_coco/retinanet_r101_fpn_mstrain_3x_coco_20210720_214650-7ee888e0.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r101_fpn_mstrain_3x_coco/retinanet_r101_fpn_mstrain_3x_coco_20210720_214650-7ee888e0.log.json) |
+| X-101-64x4d-FPN | pytorch | 3x | 9.8 | 41.6 | [config](./retinanet_x101-64x4d_fpn_ms-640-800-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_x101_64x4d_fpn_mstrain_3x_coco/retinanet_x101_64x4d_fpn_mstrain_3x_coco_20210719_051838-022c2187.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_x101_64x4d_fpn_mstrain_3x_coco/retinanet_x101_64x4d_fpn_mstrain_3x_coco_20210719_051838-022c2187.log.json) |
+
+## Citation
+
+```latex
+@inproceedings{lin2017focal,
+ title={Focal loss for dense object detection},
+ author={Lin, Tsung-Yi and Goyal, Priya and Girshick, Ross and He, Kaiming and Doll{\'a}r, Piotr},
+ booktitle={Proceedings of the IEEE international conference on computer vision},
+ year={2017}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..0551541c59100d3cc8fb361cc8895c2dbd4cf8f3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/metafile.yml
@@ -0,0 +1,312 @@
+Collections:
+ - Name: RetinaNet
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Focal Loss
+ - FPN
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/1708.02002
+ Title: "Focal Loss for Dense Object Detection"
+ README: configs/retinanet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/detectors/retinanet.py#L6
+ Version: v2.0.0
+
+Models:
+ - Name: retinanet_r18_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/retinanet/retinanet_r18_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 1.7
+ Training Resources: 8x V100 GPUs
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 31.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r18_fpn_1x_coco/retinanet_r18_fpn_1x_coco_20220407_171055-614fd399.pth
+
+ - Name: retinanet_r18_fpn_1xb8-1x_coco
+ In Collection: RetinaNet
+ Config: configs/retinanet/retinanet_r18_fpn_1xb8-1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.0
+ Training Resources: 1x V100 GPUs
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 31.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r18_fpn_1x8_1x_coco/retinanet_r18_fpn_1x8_1x_coco_20220407_171255-4ea310d7.pth
+
+ - Name: retinanet_r50-caffe_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/retinanet/retinanet_r50-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.5
+ inference time (ms/im):
+ - value: 53.76
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 36.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r50_caffe_fpn_1x_coco/retinanet_r50_caffe_fpn_1x_coco_20200531-f11027c5.pth
+
+ - Name: retinanet_r50_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/retinanet/retinanet_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.8
+ inference time (ms/im):
+ - value: 52.63
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 36.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r50_fpn_1x_coco/retinanet_r50_fpn_1x_coco_20200130-c2398f9e.pth
+
+ - Name: retinanet_r50_fpn_amp-1x_coco
+ In Collection: RetinaNet
+ Config: configs/retinanet/retinanet_r50_fpn_amp-1x_coco.py
+ Metadata:
+ Training Memory (GB): 2.8
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ - Mixed Precision Training
+ inference time (ms/im):
+ - value: 31.65
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP16
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 36.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/fp16/retinanet_r50_fpn_fp16_1x_coco/retinanet_r50_fpn_fp16_1x_coco_20200702-0dbfb212.pth
+
+ - Name: retinanet_r50_fpn_2x_coco
+ In Collection: RetinaNet
+ Config: configs/retinanet/retinanet_r50_fpn_2x_coco.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r50_fpn_2x_coco/retinanet_r50_fpn_2x_coco_20200131-fdb43119.pth
+
+ - Name: retinanet_r50_fpn_ms-640-800-3x_coco
+ In Collection: RetinaNet
+ Config: configs/retinanet/retinanet_r50_fpn_ms-640-800-3x_coco.py
+ Metadata:
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r50_fpn_mstrain_3x_coco/retinanet_r50_fpn_mstrain_3x_coco_20210718_220633-88476508.pth
+
+ - Name: retinanet_r101-caffe_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/retinanet/retinanet_r101-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.5
+ inference time (ms/im):
+ - value: 68.03
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r101_caffe_fpn_1x_coco/retinanet_r101_caffe_fpn_1x_coco_20200531-b428fa0f.pth
+
+ - Name: retinanet_r101-caffe_fpn_ms-3x_coco
+ In Collection: RetinaNet
+ Config: configs/retinanet/retinanet_r101-caffe_fpn_ms-3x_coco.py
+ Metadata:
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r101_caffe_fpn_mstrain_3x_coco/retinanet_r101_caffe_fpn_mstrain_3x_coco_20210721_063439-88a8a944.pth
+
+ - Name: retinanet_r101_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/retinanet/retinanet_r101_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.7
+ inference time (ms/im):
+ - value: 66.67
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r101_fpn_1x_coco/retinanet_r101_fpn_1x_coco_20200130-7a93545f.pth
+
+ - Name: retinanet_r101_fpn_2x_coco
+ In Collection: RetinaNet
+ Config: configs/retinanet/retinanet_r101_fpn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 5.7
+ inference time (ms/im):
+ - value: 66.67
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r101_fpn_2x_coco/retinanet_r101_fpn_2x_coco_20200131-5560aee8.pth
+
+ - Name: retinanet_r101_fpn_ms-640-800-3x_coco
+ In Collection: RetinaNet
+ Config: configs/retinanet/retinanet_r101_fpn_ms-640-800-3x_coco.py
+ Metadata:
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_r101_fpn_mstrain_3x_coco/retinanet_r101_fpn_mstrain_3x_coco_20210720_214650-7ee888e0.pth
+
+ - Name: retinanet_x101-32x4d_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/retinanet/retinanet_x101-32x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.0
+ inference time (ms/im):
+ - value: 82.64
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_x101_32x4d_fpn_1x_coco/retinanet_x101_32x4d_fpn_1x_coco_20200130-5c8b7ec4.pth
+
+ - Name: retinanet_x101-32x4d_fpn_2x_coco
+ In Collection: RetinaNet
+ Config: configs/retinanet/retinanet_x101-32x4d_fpn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 7.0
+ inference time (ms/im):
+ - value: 82.64
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_x101_32x4d_fpn_2x_coco/retinanet_x101_32x4d_fpn_2x_coco_20200131-237fc5e1.pth
+
+ - Name: retinanet_x101-64x4d_fpn_1x_coco
+ In Collection: RetinaNet
+ Config: configs/retinanet/retinanet_x101-64x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 10.0
+ inference time (ms/im):
+ - value: 114.94
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_x101_64x4d_fpn_1x_coco/retinanet_x101_64x4d_fpn_1x_coco_20200130-366f5af1.pth
+
+ - Name: retinanet_x101-64x4d_fpn_2x_coco
+ In Collection: RetinaNet
+ Config: configs/retinanet/retinanet_x101-64x4d_fpn_2x_coco.py
+ Metadata:
+ Training Memory (GB): 10.0
+ inference time (ms/im):
+ - value: 114.94
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_x101_64x4d_fpn_2x_coco/retinanet_x101_64x4d_fpn_2x_coco_20200131-bca068ab.pth
+
+ - Name: retinanet_x101-64x4d_fpn_ms-640-800-3x_coco
+ In Collection: RetinaNet
+ Config: configs/retinanet/retinanet_x101-64x4d_fpn_ms-640-800-3x_coco.py
+ Metadata:
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/retinanet/retinanet_x101_64x4d_fpn_mstrain_3x_coco/retinanet_x101_64x4d_fpn_mstrain_3x_coco_20210719_051838-022c2187.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r101-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r101-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1f3a4487103eea868eafe8539517b38455025bbe
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r101-caffe_fpn_1x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './retinanet_r50-caffe_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet101_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r101-caffe_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r101-caffe_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..cfe773459c2529079274b241f5f99ae66d8906ad
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r101-caffe_fpn_ms-3x_coco.py
@@ -0,0 +1,8 @@
+_base_ = './retinanet_r50-caffe_fpn_ms-3x_coco.py'
+# learning policy
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet101_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a7f06002413dcdf2716975655a582a3eefaf007a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r101_fpn_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './retinanet_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r101_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r101_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..721112a221953bb86dc3259e3991d7f0f740b26c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r101_fpn_2x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './retinanet_r50_fpn_2x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r101_fpn_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r101_fpn_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..be018eaac672a4c1c3a61eac9940c4d28ea4fb40
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r101_fpn_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './retinanet_r50_fpn_8xb8-amp-lsj-200e_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r101_fpn_ms-640-800-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r101_fpn_ms-640-800-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..566397227f7861a268c4cc4e111279b95b620ab8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r101_fpn_ms-640-800-3x_coco.py
@@ -0,0 +1,9 @@
+_base_ = ['../_base_/models/retinanet_r50_fpn.py', '../common/ms_3x_coco.py']
+# optimizer
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r18_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r18_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..960211806756d38cf74eed998addcca3f8467a4d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r18_fpn_1x_coco.py
@@ -0,0 +1,20 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+# model
+model = dict(
+ backbone=dict(
+ depth=18,
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet18')),
+ neck=dict(in_channels=[64, 128, 256, 512]))
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
+
+# TODO: support auto scaling lr
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (2 samples per GPU)
+# auto_scale_lr = dict(base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r18_fpn_1xb8-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r18_fpn_1xb8-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d2e88d68e3366671e402b1766d3b456593262a9b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r18_fpn_1xb8-1x_coco.py
@@ -0,0 +1,24 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+# data
+train_dataloader = dict(batch_size=8)
+
+# model
+model = dict(
+ backbone=dict(
+ depth=18,
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet18')),
+ neck=dict(in_channels=[64, 128, 256, 512]))
+
+# Note: If the learning rate is set to 0.0025, the mAP will be 32.4.
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.005, momentum=0.9, weight_decay=0.0001))
+# TODO: support auto scaling lr
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (1 GPUs) x (8 samples per GPU)
+# auto_scale_lr = dict(base_batch_size=8)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r18_fpn_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r18_fpn_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d6833f3f4711ec28a25ae8a51687fc4ac13ffb89
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r18_fpn_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './retinanet_r50_fpn_8xb8-amp-lsj-200e_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=18,
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet18')),
+ neck=dict(in_channels=[64, 128, 256, 512]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6ba1cdddc4707b40f549189f768457312635669d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50-caffe_fpn_1x_coco.py
@@ -0,0 +1,16 @@
+_base_ = './retinanet_r50_fpn_1x_coco.py'
+model = dict(
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ # use caffe img_norm
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ norm_cfg=dict(requires_grad=False),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50-caffe_fpn_ms-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50-caffe_fpn_ms-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..93687d8c27b73ae2a172b45a733345e5fc036f03
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50-caffe_fpn_ms-1x_coco.py
@@ -0,0 +1,15 @@
+_base_ = './retinanet_r50-caffe_fpn_1x_coco.py'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50-caffe_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50-caffe_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6d1604fb9efd5deb11ffc04f6f9685739f82aea9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50-caffe_fpn_ms-2x_coco.py
@@ -0,0 +1,16 @@
+_base_ = './retinanet_r50-caffe_fpn_ms-1x_coco.py'
+# training schedule for 2x
+train_cfg = dict(max_epochs=24)
+
+# learning rate policy
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=24,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50-caffe_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50-caffe_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5a6d42a13c27d5fc0b8072e2c96ef5d15a0f248c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50-caffe_fpn_ms-3x_coco.py
@@ -0,0 +1,17 @@
+_base_ = './retinanet_r50-caffe_fpn_ms-1x_coco.py'
+
+# training schedule for 2x
+train_cfg = dict(max_epochs=36)
+
+# learning rate policy
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=36,
+ by_epoch=True,
+ milestones=[28, 34],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..00d2567b245dba2b2be815a92146ea1364e1e799
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50_fpn_1x_coco.py
@@ -0,0 +1,10 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py',
+ './retinanet_tta.py'
+]
+
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..47511b78ed2edb43121de2fc27986f6bb81abcfa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50_fpn_2x_coco.py
@@ -0,0 +1,25 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+# training schedule for 2x
+train_cfg = dict(max_epochs=24)
+
+# learning rate policy
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=24,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50_fpn_8xb8-amp-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50_fpn_8xb8-amp-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..2f10db2f3c84d4b1970f13f54c563408487d04af
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50_fpn_8xb8-amp-lsj-200e_coco.py
@@ -0,0 +1,21 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../common/lsj-200e_coco-detection.py'
+]
+
+image_size = (1024, 1024)
+batch_augments = [dict(type='BatchFixedSizePad', size=image_size)]
+
+model = dict(data_preprocessor=dict(batch_augments=batch_augments))
+
+train_dataloader = dict(batch_size=8, num_workers=4)
+# Enable automatic-mixed-precision training with AmpOptimWrapper.
+optim_wrapper = dict(
+ type='AmpOptimWrapper',
+ optimizer=dict(
+ type='SGD', lr=0.01 * 4, momentum=0.9, weight_decay=0.00004))
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50_fpn_90k_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50_fpn_90k_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1e1b2fd950a0293220cc93ce3f3b377b4163f3aa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50_fpn_90k_coco.py
@@ -0,0 +1,24 @@
+_base_ = 'retinanet_r50_fpn_1x_coco.py'
+
+# training schedule for 90k
+train_cfg = dict(
+ _delete_=True,
+ type='IterBasedTrainLoop',
+ max_iters=90000,
+ val_interval=10000)
+# learning rate policy
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=90000,
+ by_epoch=False,
+ milestones=[60000, 80000],
+ gamma=0.1)
+]
+train_dataloader = dict(sampler=dict(type='InfiniteSampler'))
+default_hooks = dict(checkpoint=dict(by_epoch=False, interval=10000))
+
+log_processor = dict(by_epoch=False)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50_fpn_amp-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50_fpn_amp-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..acf5266337b8e73957a1cdf2b06076c1733b4d56
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50_fpn_amp-1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './retinanet_r50_fpn_1x_coco.py'
+
+# MMEngine support the following two ways, users can choose
+# according to convenience
+# optim_wrapper = dict(type='AmpOptimWrapper')
+_base_.optim_wrapper.type = 'AmpOptimWrapper'
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50_fpn_ms-640-800-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50_fpn_ms-640-800-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d91cf8ce0df15968706631d7eac76e834cba93dc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_r50_fpn_ms-640-800-3x_coco.py
@@ -0,0 +1,4 @@
+_base_ = ['../_base_/models/retinanet_r50_fpn.py', '../common/ms_3x_coco.py']
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_tta.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_tta.py
new file mode 100644
index 0000000000000000000000000000000000000000..d0f37e0ab25e2aff1ad55e76a7ee02777293d507
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_tta.py
@@ -0,0 +1,23 @@
+tta_model = dict(
+ type='DetTTAModel',
+ tta_cfg=dict(nms=dict(type='nms', iou_threshold=0.5), max_per_img=100))
+
+img_scales = [(1333, 800), (666, 400), (2000, 1200)]
+tta_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=None),
+ dict(
+ type='TestTimeAug',
+ transforms=[[
+ dict(type='Resize', scale=s, keep_ratio=True) for s in img_scales
+ ], [
+ dict(type='RandomFlip', prob=1.),
+ dict(type='RandomFlip', prob=0.)
+ ], [dict(type='LoadAnnotations', with_bbox=True)],
+ [
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape',
+ 'img_shape', 'scale_factor', 'flip',
+ 'flip_direction'))
+ ]])
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_x101-32x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_x101-32x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..765a4c2cc0f69bf13891bf371c94c17b6cd5f30c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_x101-32x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './retinanet_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_x101-32x4d_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_x101-32x4d_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..14de96faf70180d7828a670630a8f48a3cd1081d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_x101-32x4d_fpn_2x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './retinanet_r50_fpn_2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_x101-64x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_x101-64x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..948cd18e4d995d18d947b345ba7229b5cad60eb1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_x101-64x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './retinanet_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_x101-64x4d_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_x101-64x4d_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ad04b6eea793add40c81d1d7096481597357d5bd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_x101-64x4d_fpn_2x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './retinanet_r50_fpn_2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_x101-64x4d_fpn_ms-640-800-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_x101-64x4d_fpn_ms-640-800-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..853134160cd2128cac7954cca7e008444522fd2c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/retinanet/retinanet_x101-64x4d_fpn_ms-640-800-3x_coco.py
@@ -0,0 +1,11 @@
+_base_ = ['../_base_/models/retinanet_r50_fpn.py', '../common/ms_3x_coco.py']
+# optimizer
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
+optim_wrapper = dict(optimizer=dict(type='SGD', lr=0.01))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..bd328b4746d4125f68554eeeca3d2d765c638a5a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/README.md
@@ -0,0 +1,39 @@
+# RPN
+
+> [Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks](https://arxiv.org/abs/1506.01497)
+
+
+
+## Abstract
+
+State-of-the-art object detection networks depend on region proposal algorithms to hypothesize object locations. Advances like SPPnet and Fast R-CNN have reduced the running time of these detection networks, exposing region proposal computation as a bottleneck. In this work, we introduce a Region Proposal Network (RPN) that shares full-image convolutional features with the detection network, thus enabling nearly cost-free region proposals. An RPN is a fully convolutional network that simultaneously predicts object bounds and objectness scores at each position. The RPN is trained end-to-end to generate high-quality region proposals, which are used by Fast R-CNN for detection. We further merge RPN and Fast R-CNN into a single network by sharing their convolutional features---using the recently popular terminology of neural networks with 'attention' mechanisms, the RPN component tells the unified network where to look. For the very deep VGG-16 model, our detection system has a frame rate of 5fps (including all steps) on a GPU, while achieving state-of-the-art object detection accuracy on PASCAL VOC 2007, 2012, and MS COCO datasets with only 300 proposals per image. In ILSVRC and COCO 2015 competitions, Faster R-CNN and RPN are the foundations of the 1st-place winning entries in several tracks.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | AR1000 | Config | Download |
+| :-------------: | :-----: | :-----: | :------: | :------------: | :----: | :---------------------------------------: | :-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | caffe | 1x | 3.5 | 22.6 | 58.7 | [config](./rpn_r50-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_r50_caffe_fpn_1x_coco/rpn_r50_caffe_fpn_1x_coco_20200531-5b903a37.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_r50_caffe_fpn_1x_coco/rpn_r50_caffe_fpn_1x_coco_20200531_012334.log.json) |
+| R-50-FPN | pytorch | 1x | 3.8 | 22.3 | 58.2 | [config](./rpn_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_r50_fpn_1x_coco/rpn_r50_fpn_1x_coco_20200218-5525fa2e.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_r50_fpn_1x_coco/rpn_r50_fpn_1x_coco_20200218_151240.log.json) |
+| R-50-FPN | pytorch | 2x | - | - | 58.6 | [config](./rpn_r50_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_r50_fpn_2x_coco/rpn_r50_fpn_2x_coco_20200131-0728c9b3.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_r50_fpn_2x_coco/rpn_r50_fpn_2x_coco_20200131_190631.log.json) |
+| R-101-FPN | caffe | 1x | 5.4 | 17.3 | 60.0 | [config](./rpn_r101-caffe_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_r101_caffe_fpn_1x_coco/rpn_r101_caffe_fpn_1x_coco_20200531-0629a2e2.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_r101_caffe_fpn_1x_coco/rpn_r101_caffe_fpn_1x_coco_20200531_012345.log.json) |
+| R-101-FPN | pytorch | 1x | 5.8 | 16.5 | 59.7 | [config](./rpn_r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_r101_fpn_1x_coco/rpn_r101_fpn_1x_coco_20200131-2ace2249.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_r101_fpn_1x_coco/rpn_r101_fpn_1x_coco_20200131_191000.log.json) |
+| R-101-FPN | pytorch | 2x | - | - | 60.2 | [config](./rpn_r101_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_r101_fpn_2x_coco/rpn_r101_fpn_2x_coco_20200131-24e3db1a.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_r101_fpn_2x_coco/rpn_r101_fpn_2x_coco_20200131_191106.log.json) |
+| X-101-32x4d-FPN | pytorch | 1x | 7.0 | 13.0 | 60.6 | [config](./rpn_x101-32x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_x101_32x4d_fpn_1x_coco/rpn_x101_32x4d_fpn_1x_coco_20200219-b02646c6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_x101_32x4d_fpn_1x_coco/rpn_x101_32x4d_fpn_1x_coco_20200219_012037.log.json) |
+| X-101-32x4d-FPN | pytorch | 2x | - | - | 61.1 | [config](./rpn_x101-32x4d_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_x101_32x4d_fpn_2x_coco/rpn_x101_32x4d_fpn_2x_coco_20200208-d22bd0bb.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_x101_32x4d_fpn_2x_coco/rpn_x101_32x4d_fpn_2x_coco_20200208_200752.log.json) |
+| X-101-64x4d-FPN | pytorch | 1x | 10.1 | 9.1 | 61.0 | [config](./rpn_x101-64x4d_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_x101_64x4d_fpn_1x_coco/rpn_x101_64x4d_fpn_1x_coco_20200208-cde6f7dd.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_x101_64x4d_fpn_1x_coco/rpn_x101_64x4d_fpn_1x_coco_20200208_200752.log.json) |
+| X-101-64x4d-FPN | pytorch | 2x | - | - | 61.5 | [config](./rpn_x101-64x4d_fpn_2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_x101_64x4d_fpn_2x_coco/rpn_x101_64x4d_fpn_2x_coco_20200208-c65f524f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_x101_64x4d_fpn_2x_coco/rpn_x101_64x4d_fpn_2x_coco_20200208_200752.log.json) |
+
+## Citation
+
+```latex
+@inproceedings{ren2015faster,
+ title={Faster r-cnn: Towards real-time object detection with region proposal networks},
+ author={Ren, Shaoqing and He, Kaiming and Girshick, Ross and Sun, Jian},
+ booktitle={Advances in neural information processing systems},
+ year={2015}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..9796ead6d2ed28f0e10e16165103e31c289dae26
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/metafile.yml
@@ -0,0 +1,127 @@
+Collections:
+ - Name: RPN
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - FPN
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/1506.01497
+ Title: "Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks"
+ README: configs/rpn/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/mmdet/models/detectors/rpn.py#L6
+ Version: v2.0.0
+
+Models:
+ - Name: rpn_r50-caffe_fpn_1x_coco
+ In Collection: RPN
+ Config: configs/rpn/rpn_r50-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.5
+ Training Resources: 8x V100 GPUs
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ AR@1000: 58.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_r50_caffe_fpn_1x_coco/rpn_r50_caffe_fpn_1x_coco_20200531-5b903a37.pth
+
+ - Name: rpn_r50_fpn_1x_coco
+ In Collection: RPN
+ Config: configs/rpn/rpn_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 3.8
+ Training Resources: 8x V100 GPUs
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ AR@1000: 58.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_r50_fpn_1x_coco/rpn_r50_fpn_1x_coco_20200218-5525fa2e.pth
+
+ - Name: rpn_r50_fpn_2x_coco
+ In Collection: RPN
+ Config: rpn_r50_fpn_2x_coco.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ AR@1000: 58.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_r50_fpn_2x_coco/rpn_r50_fpn_2x_coco_20200131-0728c9b3.pth
+
+ - Name: rpn_r101-caffe_fpn_1x_coco
+ In Collection: RPN
+ Config: configs/rpn/rpn_r101-caffe_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.4
+ Training Resources: 8x V100 GPUs
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ AR@1000: 60.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_r101_caffe_fpn_1x_coco/rpn_r101_caffe_fpn_1x_coco_20200531-0629a2e2.pth
+
+ - Name: rpn_x101-32x4d_fpn_1x_coco
+ In Collection: RPN
+ Config: configs/rpn/rpn_x101-32x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.0
+ Training Resources: 8x V100 GPUs
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ AR@1000: 60.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_x101_32x4d_fpn_1x_coco/rpn_x101_32x4d_fpn_1x_coco_20200219-b02646c6.pth
+
+ - Name: rpn_x101-32x4d_fpn_2x_coco
+ In Collection: RPN
+ Config: configs/rpn/rpn_x101-32x4d_fpn_2x_coco.py
+ Metadata:
+ Training Resources: 8x V100 GPUs
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ AR@1000: 61.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_x101_32x4d_fpn_2x_coco/rpn_x101_32x4d_fpn_2x_coco_20200208-d22bd0bb.pth
+
+ - Name: rpn_x101-64x4d_fpn_1x_coco
+ In Collection: RPN
+ Config: configs/rpn/rpn_x101-64x4d_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 10.1
+ Training Resources: 8x V100 GPUs
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ AR@1000: 61.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_x101_64x4d_fpn_1x_coco/rpn_x101_64x4d_fpn_1x_coco_20200208-cde6f7dd.pth
+
+ - Name: rpn_x101-64x4d_fpn_2x_coco
+ In Collection: RPN
+ Config: configs/rpn/rpn_x101-64x4d_fpn_2x_coco.py
+ Metadata:
+ Training Resources: 8x V100 GPUs
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ AR@1000: 61.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/rpn/rpn_x101_64x4d_fpn_2x_coco/rpn_x101_64x4d_fpn_2x_coco_20200208-c65f524f.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r101-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r101-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..22977af8cb761f9415c55f8fa6d458937a00ba06
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r101-caffe_fpn_1x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './rpn_r50-caffe_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet101_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..962728ff08abb4652c617a085649575b6cfdcbf8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r101_fpn_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './rpn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r101_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r101_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ac7671c1c2421c0caa7b42d012cc3a2edc068934
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r101_fpn_2x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './rpn_r50_fpn_2x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r50-caffe-c4_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r50-caffe-c4_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..76b878c874d6545e537ee8a9618e83bb095de281
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r50-caffe-c4_1x_coco.py
@@ -0,0 +1,8 @@
+_base_ = [
+ '../_base_/models/rpn_r50-caffe-c4.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+val_evaluator = dict(metric='proposal_fast')
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r50-caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r50-caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..530f365210572f9bf55ca2775bfdbeba98567076
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r50-caffe_fpn_1x_coco.py
@@ -0,0 +1,16 @@
+_base_ = './rpn_r50_fpn_1x_coco.py'
+# use caffe img_norm
+model = dict(
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ norm_cfg=dict(requires_grad=False),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..7fe88d395b8a32e7513ede3c0c724e29b3554da6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r50_fpn_1x_coco.py
@@ -0,0 +1,36 @@
+_base_ = [
+ '../_base_/models/rpn_r50_fpn.py', '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+val_evaluator = dict(metric='proposal_fast')
+test_evaluator = val_evaluator
+
+# inference on val dataset and dump the proposals with evaluate metric
+# data_root = 'data/coco/'
+# test_evaluator = [
+# dict(
+# type='DumpProposals',
+# output_dir=data_root + 'proposals/',
+# proposals_file='rpn_r50_fpn_1x_val2017.pkl'),
+# dict(
+# type='CocoMetric',
+# ann_file=data_root + 'annotations/instances_val2017.json',
+# metric='proposal_fast',
+# backend_args={{_base_.backend_args}},
+# format_only=False)
+# ]
+
+# inference on training dataset and dump the proposals without evaluate metric
+# data_root = 'data/coco/'
+# test_dataloader = dict(
+# dataset=dict(
+# ann_file='annotations/instances_train2017.json',
+# data_prefix=dict(img='train2017/')))
+#
+# test_evaluator = [
+# dict(
+# type='DumpProposals',
+# output_dir=data_root + 'proposals/',
+# proposals_file='rpn_r50_fpn_1x_train2017.pkl'),
+# ]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r50_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r50_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..0ebccbcfaf394fcbb4fbdaea51abdd583f628cac
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_r50_fpn_2x_coco.py
@@ -0,0 +1,17 @@
+_base_ = './rpn_r50_fpn_1x_coco.py'
+
+# learning policy
+max_epochs = 24
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_x101-32x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_x101-32x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d0c73948ac56afa34b9d6c8d22d6158271306b8c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_x101-32x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './rpn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_x101-32x4d_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_x101-32x4d_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c6880b762abc8f5d3bf12f278054d76958756fb2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_x101-32x4d_fpn_2x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './rpn_r50_fpn_2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_x101-64x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_x101-64x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..96e691a912c424f09add038c75631a2e1fefeffc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_x101-64x4d_fpn_1x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './rpn_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_x101-64x4d_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_x101-64x4d_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..4182a39667c47d774a1df9d34a1bc2fe60b45538
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rpn/rpn_x101-64x4d_fpn_2x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './rpn_r50_fpn_2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..1677184af761a5b6ac5d643ddf7e2d802f96723e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/README.md
@@ -0,0 +1,457 @@
+# RTMDet: An Empirical Study of Designing Real-Time Object Detectors
+
+> [RTMDet: An Empirical Study of Designing Real-Time Object Detectors](https://arxiv.org/abs/2212.07784)
+
+[](https://paperswithcode.com/sota/real-time-instance-segmentation-on-mscoco?p=rtmdet-an-empirical-study-of-designing-real)
+[](https://paperswithcode.com/sota/object-detection-in-aerial-images-on-dota-1?p=rtmdet-an-empirical-study-of-designing-real)
+[](https://paperswithcode.com/sota/object-detection-in-aerial-images-on-hrsc2016?p=rtmdet-an-empirical-study-of-designing-real)
+
+
+
+## Abstract
+
+In this paper, we aim to design an efficient real-time object detector that exceeds the YOLO series and is easily extensible for many object recognition tasks such as instance segmentation and rotated object detection. To obtain a more efficient model architecture, we explore an architecture that has compatible capacities in the backbone and neck, constructed by a basic building block that consists of large-kernel depth-wise convolutions. We further introduce soft labels when calculating matching costs in the dynamic label assignment to improve accuracy. Together with better training techniques, the resulting object detector, named RTMDet, achieves 52.8% AP on COCO with 300+ FPS on an NVIDIA 3090 GPU, outperforming the current mainstream industrial detectors. RTMDet achieves the best parameter-accuracy trade-off with tiny/small/medium/large/extra-large model sizes for various application scenarios, and obtains new state-of-the-art performance on real-time instance segmentation and rotated object detection. We hope the experimental results can provide new insights into designing versatile real-time object detectors for many object recognition tasks.
+
+
+

+
+
+## Results and Models
+
+### Object Detection
+
+| Model | size | box AP | Params(M) | FLOPS(G) | TRT-FP16-Latency(ms)
RTX3090 | TRT-FP16-Latency(ms)
T4 | Config | Download |
+| :-----------------: | :--: | :----: | :-------: | :------: | :-----------------------------: | :------------------------: | :------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| RTMDet-tiny | 640 | 41.1 | 4.8 | 8.1 | 0.98 | 2.34 | [config](./rtmdet_tiny_8xb32-300e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet_tiny_8xb32-300e_coco/rtmdet_tiny_8xb32-300e_coco_20220902_112414-78e30dcc.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet_tiny_8xb32-300e_coco/rtmdet_tiny_8xb32-300e_coco_20220902_112414.log.json) |
+| RTMDet-s | 640 | 44.6 | 8.89 | 14.8 | 1.22 | 2.96 | [config](./rtmdet_s_8xb32-300e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet_s_8xb32-300e_coco/rtmdet_s_8xb32-300e_coco_20220905_161602-387a891e.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet_s_8xb32-300e_coco/rtmdet_s_8xb32-300e_coco_20220905_161602.log.json) |
+| RTMDet-m | 640 | 49.4 | 24.71 | 39.27 | 1.62 | 6.41 | [config](./rtmdet_m_8xb32-300e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet_m_8xb32-300e_coco/rtmdet_m_8xb32-300e_coco_20220719_112220-229f527c.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet_m_8xb32-300e_coco/rtmdet_m_8xb32-300e_coco_20220719_112220.log.json) |
+| RTMDet-l | 640 | 51.5 | 52.3 | 80.23 | 2.44 | 10.32 | [config](./rtmdet_l_8xb32-300e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet_l_8xb32-300e_coco/rtmdet_l_8xb32-300e_coco_20220719_112030-5a0be7c4.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet_l_8xb32-300e_coco/rtmdet_l_8xb32-300e_coco_20220719_112030.log.json) |
+| RTMDet-x | 640 | 52.8 | 94.86 | 141.67 | 3.10 | 18.80 | [config](./rtmdet_x_8xb32-300e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet_x_8xb32-300e_coco/rtmdet_x_8xb32-300e_coco_20220715_230555-cc79b9ae.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet_x_8xb32-300e_coco/rtmdet_x_8xb32-300e_coco_20220715_230555.log.json) |
+| RTMDet-x-P6 | 1280 | 54.9 | | | | | [config](./rtmdet_x_p6_4xb8-300e_coco.py) | [model](https://github.com/orange0-jp/orange-weights/releases/download/v0.1.0rtmdet-p6/rtmdet_x_p6_4xb8-300e_coco-bf32be58.pth) |
+| RTMDet-l-ConvNeXt-B | 640 | 53.1 | | | | | [config](./rtmdet_l_convnext_b_4xb32-100e_coco.py) | [model](https://github.com/orange0-jp/orange-weights/releases/download/v0.1.0rtmdet-swin-convnext/rtmdet_l_convnext_b_4xb32-100e_coco-d4731b3d.pth) |
+| RTMDet-l-Swin-B | 640 | 52.4 | | | | | [config](./rtmdet_l_swin_b_4xb32-100e_coco.py) | [model](https://github.com/orange0-jp/orange-weights/releases/download/v0.1.0rtmdet-swin-convnext/rtmdet_l_swin_b_4xb32-100e_coco-0828ce5d.pth) |
+| RTMDet-l-Swin-B-P6 | 1280 | 56.4 | | | | | [config](./rtmdet_l_swin_b_p6_4xb16-100e_coco.py) | [model](https://github.com/orange0-jp/orange-weights/releases/download/v0.1.0rtmdet-swin-convnext/rtmdet_l_swin_b_p6_4xb16-100e_coco-a1486b6f.pth) |
+
+**Note**:
+
+1. We implement a fast training version of RTMDet in [MMYOLO](https://github.com/open-mmlab/mmyolo). Its training speed is **2.6 times faster** and memory requirement is lower! Try it [here](https://github.com/open-mmlab/mmyolo/tree/main/configs/rtmdet)!
+2. The inference speed of RTMDet is measured with TensorRT 8.4.3, cuDNN 8.2.0, FP16, batch size=1, and without NMS.
+3. For a fair comparison, the config of bbox postprocessing is changed to be consistent with YOLOv5/6/7 after [PR#9494](https://github.com/open-mmlab/mmdetection/pull/9494), bringing about 0.1~0.3% AP improvement.
+
+### Instance Segmentation
+
+RTMDet-Ins is the state-of-the-art real-time instance segmentation on coco dataset:
+
+[](https://paperswithcode.com/sota/real-time-instance-segmentation-on-mscoco?p=rtmdet-an-empirical-study-of-designing-real)
+
+| Model | size | box AP | mask AP | Params(M) | FLOPS(G) | TRT-FP16-Latency(ms) | Config | Download |
+| :-------------: | :--: | :----: | :-----: | :-------: | :------: | :------------------: | :--------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| RTMDet-Ins-tiny | 640 | 40.5 | 35.4 | 5.6 | 11.8 | 1.70 | [config](./rtmdet-ins_tiny_8xb32-300e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet-ins_tiny_8xb32-300e_coco/rtmdet-ins_tiny_8xb32-300e_coco_20221130_151727-ec670f7e.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet-ins_tiny_8xb32-300e_coco/rtmdet-ins_tiny_8xb32-300e_coco_20221130_151727.log.json) |
+| RTMDet-Ins-s | 640 | 44.0 | 38.7 | 10.18 | 21.5 | 1.93 | [config](./rtmdet-ins_s_8xb32-300e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet-ins_s_8xb32-300e_coco/rtmdet-ins_s_8xb32-300e_coco_20221121_212604-fdc5d7ec.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet-ins_s_8xb32-300e_coco/rtmdet-ins_s_8xb32-300e_coco_20221121_212604.log.json) |
+| RTMDet-Ins-m | 640 | 48.8 | 42.1 | 27.58 | 54.13 | 2.69 | [config](./rtmdet-ins_m_8xb32-300e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet-ins_m_8xb32-300e_coco/rtmdet-ins_m_8xb32-300e_coco_20221123_001039-6eba602e.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet-ins_m_8xb32-300e_coco/rtmdet-ins_m_8xb32-300e_coco_20221123_001039.log.json) |
+| RTMDet-Ins-l | 640 | 51.2 | 43.7 | 57.37 | 106.56 | 3.68 | [config](./rtmdet-ins_l_8xb32-300e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet-ins_l_8xb32-300e_coco/rtmdet-ins_l_8xb32-300e_coco_20221124_103237-78d1d652.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet-ins_l_8xb32-300e_coco/rtmdet-ins_l_8xb32-300e_coco_20221124_103237.log.json) |
+| RTMDet-Ins-x | 640 | 52.4 | 44.6 | 102.7 | 182.7 | 5.31 | [config](./rtmdet-ins_x_8xb16-300e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet-ins_x_8xb16-300e_coco/rtmdet-ins_x_8xb16-300e_coco_20221124_111313-33d4595b.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet-ins_x_8xb16-300e_coco/rtmdet-ins_x_8xb16-300e_coco_20221124_111313.log.json) |
+
+**Note**:
+
+1. The inference speed of RTMDet-Ins is measured on an NVIDIA 3090 GPU with TensorRT 8.4.3, cuDNN 8.2.0, FP16, batch size=1. Top 100 masks are kept and the post process latency is included.
+
+### Rotated Object Detection
+
+RTMDet-R achieves state-of-the-art on various remote sensing datasets.
+
+[](https://paperswithcode.com/sota/object-detection-in-aerial-images-on-dota-1?p=rtmdet-an-empirical-study-of-designing-real)
+
+[](https://paperswithcode.com/sota/one-stage-anchor-free-oriented-object-1?p=rtmdet-an-empirical-study-of-designing-real)
+
+[](https://paperswithcode.com/sota/object-detection-in-aerial-images-on-hrsc2016?p=rtmdet-an-empirical-study-of-designing-real)
+
+[](https://paperswithcode.com/sota/one-stage-anchor-free-oriented-object-3?p=rtmdet-an-empirical-study-of-designing-real)
+
+Models and configs of RTMDet-R are available in [MMRotate](https://github.com/open-mmlab/mmrotate/tree/1.x/configs/rotated_rtmdet).
+
+| Backbone | pretrain | Aug | mmAP | mAP50 | mAP75 | Params(M) | FLOPS(G) | TRT-FP16-Latency(ms) | Config | Download |
+| :---------: | :------: | :---: | :---: | :---: | :---: | :-------: | :------: | :------------------: | :---------------------------------------------------------------------------------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| RTMDet-tiny | IN | RR | 47.37 | 75.36 | 50.64 | 4.88 | 20.45 | 4.40 | [config](https://github.com/open-mmlab/mmrotate/edit/1.x/configs/rotated_rtmdet/rotated_rtmdet_tiny-3x-dota.py) | [model](https://download.openmmlab.com/mmrotate/v1.0/rotated_rtmdet/rotated_rtmdet_tiny-3x-dota/rotated_rtmdet_tiny-3x-dota-9d821076.pth) \| [log](https://download.openmmlab.com/mmrotate/v1.0/rotated_rtmdet/rotated_rtmdet_tiny-3x-dota/rotated_rtmdet_tiny-3x-dota_20221201_120814.json) |
+| RTMDet-tiny | IN | MS+RR | 53.59 | 79.82 | 58.87 | 4.88 | 20.45 | 4.40 | [config](https://github.com/open-mmlab/mmrotate/edit/1.x/configs/rotated_rtmdet/rotated_rtmdet_tiny-3x-dota_ms.py) | [model](https://download.openmmlab.com/mmrotate/v1.0/rotated_rtmdet/rotated_rtmdet_tiny-3x-dota_ms/rotated_rtmdet_tiny-3x-dota_ms-f12286ff.pth) \| [log](https://download.openmmlab.com/mmrotate/v1.0/rotated_rtmdet/rotated_rtmdet_tiny-3x-dota_ms/rotated_rtmdet_tiny-3x-dota_ms_20221113_201235.log) |
+| RTMDet-s | IN | RR | 48.16 | 76.93 | 50.59 | 8.86 | 37.62 | 4.86 | [config](https://github.com/open-mmlab/mmrotate/edit/1.x/configs/rotated_rtmdet/rotated_rtmdet_s-3x-dota.py) | [model](https://download.openmmlab.com/mmrotate/v1.0/rotated_rtmdet/rotated_rtmdet_s-3x-dota/rotated_rtmdet_s-3x-dota-11f6ccf5.pth) \| [log](https://download.openmmlab.com/mmrotate/v1.0/rotated_rtmdet/rotated_rtmdet_s-3x-dota/rotated_rtmdet_s-3x-dota_20221124_081442.json) |
+| RTMDet-s | IN | MS+RR | 54.43 | 79.98 | 60.07 | 8.86 | 37.62 | 4.86 | [config](https://github.com/open-mmlab/mmrotate/edit/1.x/configs/rotated_rtmdet/rotated_rtmdet_s-3x-dota_ms.py) | [model](https://download.openmmlab.com/mmrotate/v1.0/rotated_rtmdet/rotated_rtmdet_s-3x-dota_ms/rotated_rtmdet_s-3x-dota_ms-20ead048.pth) \| [log](https://download.openmmlab.com/mmrotate/v1.0/rotated_rtmdet/rotated_rtmdet_s-3x-dota_ms/rotated_rtmdet_s-3x-dota_ms_20221113_201055.json) |
+| RTMDet-m | IN | RR | 50.56 | 78.24 | 54.47 | 24.67 | 99.76 | 7.82 | [config](https://github.com/open-mmlab/mmrotate/edit/1.x/configs/rotated_rtmdet/rotated_rtmdet_m-3x-dota.py) | [model](https://download.openmmlab.com/mmrotate/v1.0/rotated_rtmdet/rotated_rtmdet_m-3x-dota/rotated_rtmdet_m-3x-dota-beeadda6.pth) \| [log](https://download.openmmlab.com/mmrotate/v1.0/rotated_rtmdet/rotated_rtmdet_m-3x-dota/rotated_rtmdet_m-3x-dota_20221122_011234.json) |
+| RTMDet-m | IN | MS+RR | 55.00 | 80.26 | 61.26 | 24.67 | 99.76 | 7.82 | [config](https://github.com/open-mmlab/mmrotate/edit/1.x/configs/rotated_rtmdet/rotated_rtmdet_m-3x-dota_ms.py) | [model](https://download.openmmlab.com/mmrotate/v1.0/rotated_rtmdet/rotated_rtmdet_m-3x-dota_ms/rotated_rtmdet_m-3x-dota_ms-c71eb375.pth) \| [log](https://download.openmmlab.com/mmrotate/v1.0/rotated_rtmdet/rotated_rtmdet_m-3x-dota_ms/rotated_rtmdet_m-3x-dota_ms_20221122_011234.json) |
+| RTMDet-l | IN | RR | 51.01 | 78.85 | 55.21 | 52.27 | 204.21 | 10.82 | [config](https://github.com/open-mmlab/mmrotate/edit/1.x/configs/rotated_rtmdet/rotated_rtmdet_l-3x-dota.py) | [model](https://download.openmmlab.com/mmrotate/v1.0/rotated_rtmdet/rotated_rtmdet_l-3x-dota/rotated_rtmdet_l-3x-dota-23992372.pth) \| [log](https://download.openmmlab.com/mmrotate/v1.0/rotated_rtmdet/rotated_rtmdet_l-3x-dota/rotated_rtmdet_l-3x-dota_20221122_011241.json) |
+| RTMDet-l | IN | MS+RR | 55.52 | 80.54 | 61.47 | 52.27 | 204.21 | 10.82 | [config](https://github.com/open-mmlab/mmrotate/edit/1.x/configs/rotated_rtmdet/rotated_rtmdet_l-3x-dota_ms.py) | [model](https://download.openmmlab.com/mmrotate/v1.0/rotated_rtmdet/rotated_rtmdet_l-3x-dota_ms/rotated_rtmdet_l-3x-dota_ms-2738da34.pth) \| [log](https://download.openmmlab.com/mmrotate/v1.0/rotated_rtmdet/rotated_rtmdet_l-3x-dota_ms/rotated_rtmdet_l-3x-dota_ms_20221122_011241.json) |
+| RTMDet-l | COCO | MS+RR | 56.74 | 81.33 | 63.45 | 52.27 | 204.21 | 10.82 | [config](https://github.com/open-mmlab/mmrotate/edit/1.x/configs/rotated_rtmdet/rotated_rtmdet_l-coco_pretrain-3x-dota_ms.py) | [model](https://download.openmmlab.com/mmrotate/v1.0/rotated_rtmdet/rotated_rtmdet_l-coco_pretrain-3x-dota_ms/rotated_rtmdet_l-coco_pretrain-3x-dota_ms-06d248a2.pth) \| [log](https://download.openmmlab.com/mmrotate/v1.0/rotated_rtmdet/rotated_rtmdet_l-coco_pretrain-3x-dota_ms/rotated_rtmdet_l-coco_pretrain-3x-dota_ms_20221113_202010.json) |
+
+### Classification
+
+We also provide the imagenet classification configs of the RTMDet backbone. Find more details in the [classification folder](./classification).
+
+| Model | resolution | Params(M) | Flops(G) | Top-1 (%) | Top-5 (%) | Download |
+| :----------: | :--------: | :-------: | :------: | :-------: | :-------: | :---------------------------------------------------------------------------------------------------------------------------------: |
+| CSPNeXt-tiny | 224x224 | 2.73 | 0.34 | 69.44 | 89.45 | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/cspnext_rsb_pretrain/cspnext-tiny_imagenet_600e-3a2dd350.pth) |
+| CSPNeXt-s | 224x224 | 4.89 | 0.66 | 74.41 | 92.23 | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/cspnext_rsb_pretrain/cspnext-s_imagenet_600e-ea671761.pth) |
+| CSPNeXt-m | 224x224 | 13.05 | 1.93 | 79.27 | 94.79 | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/cspnext_rsb_pretrain/cspnext-m_8xb256-rsb-a1-600e_in1k-ecb3bbd9.pth) |
+| CSPNeXt-l | 224x224 | 27.16 | 4.19 | 81.30 | 95.62 | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/cspnext_rsb_pretrain/cspnext-l_8xb256-rsb-a1-600e_in1k-6a760974.pth) |
+| CSPNeXt-x | 224x224 | 48.85 | 7.76 | 82.10 | 95.69 | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/cspnext_rsb_pretrain/cspnext-x_8xb256-rsb-a1-600e_in1k-b3f78edd.pth) |
+
+## Citation
+
+```latex
+@misc{lyu2022rtmdet,
+ title={RTMDet: An Empirical Study of Designing Real-Time Object Detectors},
+ author={Chengqi Lyu and Wenwei Zhang and Haian Huang and Yue Zhou and Yudong Wang and Yanyi Liu and Shilong Zhang and Kai Chen},
+ year={2022},
+ eprint={2212.07784},
+ archivePrefix={arXiv},
+ primaryClass={cs.CV}
+}
+```
+
+## Visualization
+
+
+

+
+
+## Deployment Tutorial
+
+Here is a basic example of deploy RTMDet with [MMDeploy-1.x](https://github.com/open-mmlab/mmdeploy/tree/1.x).
+
+### Step1. Install MMDeploy
+
+Before starting the deployment, please make sure you install MMDetection and MMDeploy-1.x correctly.
+
+- Install MMDetection, please refer to the [MMDetection installation guide](https://mmdetection.readthedocs.io/en/latest/get_started.html).
+- Install MMDeploy-1.x, please refer to the [MMDeploy-1.x installation guide](https://mmdeploy.readthedocs.io/en/1.x/get_started.html#installation).
+
+If you want to deploy RTMDet with ONNXRuntime, TensorRT, or other inference engine,
+please make sure you have installed the corresponding dependencies and MMDeploy precompiled packages.
+
+### Step2. Convert Model
+
+After the installation, you can enjoy the model deployment journey starting from converting PyTorch model to backend model by running MMDeploy's `tools/deploy.py`.
+
+The detailed model conversion tutorial please refer to the [MMDeploy document](https://mmdeploy.readthedocs.io/en/1.x/02-how-to-run/convert_model.html).
+Here we only give the example of converting RTMDet.
+
+MMDeploy supports converting dynamic and static models. Dynamic models support different input shape, but the inference speed is slower than static models.
+To achieve the best performance, we suggest converting RTMDet with static setting.
+
+- If you only want to use ONNX, please use [`configs/mmdet/detection/detection_onnxruntime_static.py`](https://github.com/open-mmlab/mmdeploy/blob/1.x/configs/mmdet/detection/detection_onnxruntime_static.py) as the deployment config.
+- If you want to use TensorRT, please use [`configs/mmdet/detection/detection_tensorrt_static-640x640.py`](https://github.com/open-mmlab/mmdeploy/blob/1.x/configs/mmdet/detection/detection_tensorrt_static-640x640.py).
+
+If you want to customize the settings in the deployment config for your requirements, please refer to [MMDeploy config tutorial](https://mmdeploy.readthedocs.io/en/1.x/02-how-to-run/write_config.html).
+
+After preparing the deployment config, you can run the `tools/deploy.py` script to convert your model.
+Here we take converting RTMDet-s to TensorRT as an example:
+
+```shell
+# go to the mmdeploy folder
+cd ${PATH_TO_MMDEPLOY}
+
+# download RTMDet-s checkpoint
+wget -P checkpoint https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet_s_8xb32-300e_coco/rtmdet_s_8xb32-300e_coco_20220905_161602-387a891e.pth
+
+# run the command to start model conversion
+python tools/deploy.py \
+ configs/mmdet/detection/detection_tensorrt_static-640x640.py \
+ ${PATH_TO_MMDET}/configs/rtmdet/rtmdet_s_8xb32-300e_coco.py \
+ checkpoint/rtmdet_s_8xb32-300e_coco_20220905_161602-387a891e.pth \
+ demo/resources/det.jpg \
+ --work-dir ./work_dirs/rtmdet \
+ --device cuda:0 \
+ --show
+```
+
+If the script runs successfully, you will see the following files:
+
+```
+|----work_dirs
+ |----rtmdet
+ |----end2end.onnx # ONNX model
+ |----end2end.engine # TensorRT engine file
+```
+
+After this, you can check the inference results with MMDeploy Model Converter API:
+
+```python
+from mmdeploy.apis import inference_model
+
+result = inference_model(
+ model_cfg='${PATH_TO_MMDET}/configs/rtmdet/rtmdet_s_8xb32-300e_coco.py',
+ deploy_cfg='${PATH_TO_MMDEPLOY}/configs/mmdet/detection/detection_tensorrt_static-640x640.py',
+ backend_files=['work_dirs/rtmdet/end2end.engine'],
+ img='demo/resources/det.jpg',
+ device='cuda:0')
+```
+
+#### Advanced Setting
+
+To convert the model with TRT-FP16, you can enable the fp16 mode in your deploy config:
+
+```python
+# in MMDeploy config
+backend_config = dict(
+ type='tensorrt',
+ common_config=dict(
+ fp16_mode=True # enable fp16
+ ))
+```
+
+To reduce the end to end inference speed with the inference engine, we suggest you to adjust the post-processing setting of the model.
+We set a very low score threshold during training and testing to achieve better COCO mAP.
+However, in actual usage scenarios, a relatively high score threshold (e.g. 0.3) is usually used.
+
+You can adjust the score threshold and the number of detection boxes in your model config according to the actual usage to reduce the time-consuming of post-processing.
+
+```python
+# in MMDetection config
+model = dict(
+ test_cfg=dict(
+ nms_pre=1000, # keep top-k score bboxes before nms
+ min_bbox_size=0,
+ score_thr=0.3, # score threshold to filter bboxes
+ nms=dict(type='nms', iou_threshold=0.65),
+ max_per_img=100) # only keep top-100 as the final results.
+)
+```
+
+### Step3. Inference with SDK
+
+We provide both Python and C++ inference API with MMDeploy SDK.
+
+To use SDK, you need to dump the required info during converting the model. Just add `--dump-info` to the model conversion command:
+
+```shell
+python tools/deploy.py \
+ configs/mmdet/detection/detection_tensorrt_static-640x640.py \
+ ${PATH_TO_MMDET}/configs/rtmdet/rtmdet_s_8xb32-300e_coco.py \
+ checkpoint/rtmdet_s_8xb32-300e_coco_20220905_161602-387a891e.pth \
+ demo/resources/det.jpg \
+ --work-dir ./work_dirs/rtmdet-sdk \
+ --device cuda:0 \
+ --show \
+ --dump-info # dump sdk info
+```
+
+After running the command, it will dump 3 json files additionally for the SDK:
+
+```
+|----work_dirs
+ |----rtmdet-sdk
+ |----end2end.onnx # ONNX model
+ |----end2end.engine # TensorRT engine file
+ # json files for the SDK
+ |----pipeline.json
+ |----deploy.json
+ |----detail.json
+```
+
+#### Python API
+
+Here is a basic example of SDK Python API:
+
+```python
+from mmdeploy_python import Detector
+import cv2
+
+img = cv2.imread('demo/resources/det.jpg')
+
+# create a detector
+detector = Detector(model_path='work_dirs/rtmdet-sdk', device_name='cuda', device_id=0)
+# run the inference
+bboxes, labels, _ = detector(img)
+# Filter the result according to threshold
+indices = [i for i in range(len(bboxes))]
+for index, bbox, label_id in zip(indices, bboxes, labels):
+ [left, top, right, bottom], score = bbox[0:4].astype(int), bbox[4]
+ if score < 0.3:
+ continue
+ # draw bbox
+ cv2.rectangle(img, (left, top), (right, bottom), (0, 255, 0))
+
+cv2.imwrite('output_detection.png', img)
+```
+
+#### C++ API
+
+Here is a basic example of SDK C++ API:
+
+```C++
+#include
+#include
+#include "mmdeploy/detector.hpp"
+
+int main() {
+ const char* device_name = "cuda";
+ int device_id = 0;
+ std::string model_path = "work_dirs/rtmdet-sdk";
+ std::string image_path = "demo/resources/det.jpg";
+
+ // 1. load model
+ mmdeploy::Model model(model_path);
+ // 2. create predictor
+ mmdeploy::Detector detector(model, mmdeploy::Device{device_name, device_id});
+ // 3. read image
+ cv::Mat img = cv::imread(image_path);
+ // 4. inference
+ auto dets = detector.Apply(img);
+ // 5. deal with the result. Here we choose to visualize it
+ for (int i = 0; i < dets.size(); ++i) {
+ const auto& box = dets[i].bbox;
+ fprintf(stdout, "box %d, left=%.2f, top=%.2f, right=%.2f, bottom=%.2f, label=%d, score=%.4f\n",
+ i, box.left, box.top, box.right, box.bottom, dets[i].label_id, dets[i].score);
+ if (bboxes[i].score < 0.3) {
+ continue;
+ }
+ cv::rectangle(img, cv::Point{(int)box.left, (int)box.top},
+ cv::Point{(int)box.right, (int)box.bottom}, cv::Scalar{0, 255, 0});
+ }
+ cv::imwrite("output_detection.png", img);
+ return 0;
+}
+```
+
+To build C++ example, please add MMDeploy package in your CMake project as following:
+
+```cmake
+find_package(MMDeploy REQUIRED)
+target_link_libraries(${name} PRIVATE mmdeploy ${OpenCV_LIBS})
+```
+
+#### Other languages
+
+- [C# API Examples](https://github.com/open-mmlab/mmdeploy/tree/1.x/demo/csharp)
+- [JAVA API Examples](https://github.com/open-mmlab/mmdeploy/tree/1.x/demo/java)
+
+### Deploy RTMDet Instance Segmentation Model
+
+We support RTMDet-Ins ONNXRuntime and TensorRT deployment after [MMDeploy v1.0.0rc2](https://github.com/open-mmlab/mmdeploy/tree/v1.0.0rc2). And its deployment process is almost consistent with the detection model.
+
+#### Step1. Install MMDeploy >= v1.0.0rc2
+
+Please refer to the [MMDeploy-1.x installation guide](https://mmdeploy.readthedocs.io/en/1.x/get_started.html#installation) to install the latest version.
+Please remember to replace the pre-built package with the latest version.
+The v1.0.0rc2 package can be downloaded from [v1.0.0rc2 release page](https://github.com/open-mmlab/mmdeploy/releases/tag/v1.0.0rc2).
+
+Step2. Convert Model
+
+This step has no difference with the previous tutorial. The only thing you need to change is switching to the RTMDet-Ins deploy config:
+
+- If you want to use ONNXRuntime, please use [`configs/mmdet/instance-seg/instance-seg_rtmdet-ins_onnxruntime_static-640x640.py`](https://github.com/open-mmlab/mmdeploy/blob/dev-1.x/configs/mmdet/instance-seg/instance-seg_rtmdet-ins_onnxruntime_static-640x640.py) as the deployment config.
+- If you want to use TensorRT, please use [`configs/mmdet/instance-seg/instance-seg_rtmdet-ins_tensorrt_static-640x640.py`](https://github.com/open-mmlab/mmdeploy/blob/dev-1.x/configs/mmdet/instance-seg/instance-seg_rtmdet-ins_tensorrt_static-640x640.py).
+
+Here we take converting RTMDet-Ins-s to TensorRT as an example:
+
+```shell
+# go to the mmdeploy folder
+cd ${PATH_TO_MMDEPLOY}
+
+# download RTMDet-s checkpoint
+wget -P checkpoint https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet-ins_s_8xb32-300e_coco/rtmdet-ins_s_8xb32-300e_coco_20221121_212604-fdc5d7ec.pth
+
+# run the command to start model conversion
+python tools/deploy.py \
+ configs/mmdet/instance-seg/instance-seg_rtmdet-ins_tensorrt_static-640x640.py \
+ ${PATH_TO_MMDET}/configs/rtmdet/rtmdet-ins_s_8xb32-300e_coco.py \
+ checkpoint/rtmdet-ins_s_8xb32-300e_coco_20221121_212604-fdc5d7ec.pth \
+ demo/resources/det.jpg \
+ --work-dir ./work_dirs/rtmdet-ins \
+ --device cuda:0 \
+ --show
+```
+
+If the script runs successfully, you will see the following files:
+
+```
+|----work_dirs
+ |----rtmdet-ins
+ |----end2end.onnx # ONNX model
+ |----end2end.engine # TensorRT engine file
+```
+
+After this, you can check the inference results with MMDeploy Model Converter API:
+
+```python
+from mmdeploy.apis import inference_model
+
+result = inference_model(
+ model_cfg='${PATH_TO_MMDET}/configs/rtmdet/rtmdet-ins_s_8xb32-300e_coco.py',
+ deploy_cfg='${PATH_TO_MMDEPLOY}/configs/mmdet/instance-seg/instance-seg_rtmdet-ins_tensorrt_static-640x640.py',
+ backend_files=['work_dirs/rtmdet-ins/end2end.engine'],
+ img='demo/resources/det.jpg',
+ device='cuda:0')
+```
+
+### Model Config
+
+In MMDetection's config, we use `model` to set up detection algorithm components. In addition to neural network components such as `backbone`, `neck`, etc, it also requires `data_preprocessor`, `train_cfg`, and `test_cfg`. `data_preprocessor` is responsible for processing a batch of data output by dataloader. `train_cfg`, and `test_cfg` in the model config are for training and testing hyperparameters of the components.Taking RTMDet as an example, we will introduce each field in the config according to different function modules:
+
+```python
+model = dict(
+ type='RTMDet', # The name of detector
+ data_preprocessor=dict( # The config of data preprocessor, usually includes image normalization and padding
+ type='DetDataPreprocessor', # The type of the data preprocessor. Refer to https://mmdetection.readthedocs.io/en/latest/api.html#mmdet.models.data_preprocessors.DetDataPreprocessor
+ mean=[103.53, 116.28, 123.675], # Mean values used to pre-training the pre-trained backbone models, ordered in R, G, B
+ std=[57.375, 57.12, 58.395], # Standard variance used to pre-training the pre-trained backbone models, ordered in R, G, B
+ bgr_to_rgb=False, # whether to convert image from BGR to RGB
+ batch_augments=None), # Batch-level augmentations
+ backbone=dict( # The config of backbone
+ type='CSPNeXt', # The type of backbone network. Refer to https://mmdetection.readthedocs.io/en/latest/api.html#mmdet.models.backbones.CSPNeXt
+ arch='P5', # Architecture of CSPNeXt, from {P5, P6}. Defaults to P5
+ expand_ratio=0.5, # Ratio to adjust the number of channels of the hidden layer. Defaults to 0.5
+ deepen_factor=1, # Depth multiplier, multiply number of blocks in CSP layer by this amount. Defaults to 1.0
+ widen_factor=1, # Width multiplier, multiply number of channels in each layer by this amount. Defaults to 1.0
+ channel_attention=True, # Whether to add channel attention in each stage. Defaults to True
+ norm_cfg=dict(type='SyncBN'), # Dictionary to construct and config norm layer. Defaults to dict(type=’BN’, requires_grad=True)
+ act_cfg=dict(type='SiLU', inplace=True)), # Config dict for activation layer. Defaults to dict(type=’SiLU’)
+ neck=dict(
+ type='CSPNeXtPAFPN', # The type of neck is CSPNeXtPAFPN. Refer to https://mmdetection.readthedocs.io/en/latest/api.html#mmdet.models.necks.CSPNeXtPAFPN
+ in_channels=[256, 512, 1024], # Number of input channels per scale
+ out_channels=256, # Number of output channels (used at each scale)
+ num_csp_blocks=3, # Number of bottlenecks in CSPLayer. Defaults to 3
+ expand_ratio=0.5, # Ratio to adjust the number of channels of the hidden layer. Default: 0.5
+ norm_cfg=dict(type='SyncBN'), # Config dict for normalization layer. Default: dict(type=’BN’)
+ act_cfg=dict(type='SiLU', inplace=True)), # Config dict for activation layer. Default: dict(type=’Swish’)
+ bbox_head=dict(
+ type='RTMDetSepBNHead', # The type of bbox_head is RTMDetSepBNHead. RTMDetHead with separated BN layers and shared conv layers. Refer to https://mmdetection.readthedocs.io/en/latest/api.html#mmdet.models.dense_heads.RTMDetSepBNHead
+ num_classes=80, # Number of categories excluding the background category
+ in_channels=256, # Number of channels in the input feature map
+ stacked_convs=2, # Whether to share conv layers between stages. Defaults to True
+ feat_channels=256, # Feature channels of convolutional layers in the head
+ anchor_generator=dict( # The config of anchor generator
+ type='MlvlPointGenerator', # The methods use MlvlPointGenerator. Refer to https://github.com/open-mmlab/mmdetection/blob/main/mmdet/models/task_modules/prior_generators/point_generator.py#L92
+ offset=0, # The offset of points, the value is normalized with corresponding stride. Defaults to 0.5
+ strides=[8, 16, 32]), # Strides of anchors in multiple feature levels in order (w, h)
+ bbox_coder=dict(type='DistancePointBBoxCoder'), # Distance Point BBox coder.This coder encodes gt bboxes (x1, y1, x2, y2) into (top, bottom, left,right) and decode it back to the original. Refer to https://github.com/open-mmlab/mmdetection/blob/main/mmdet/models/task_modules/coders/distance_point_bbox_coder.py#L9
+ loss_cls=dict( # Config of loss function for the classification branch
+ type='QualityFocalLoss', # Type of loss for classification branch. Refer to https://mmdetection.readthedocs.io/en/latest/api.html#mmdet.models.losses.QualityFocalLoss
+ use_sigmoid=True, # Whether sigmoid operation is conducted in QFL. Defaults to True
+ beta=2.0, # The beta parameter for calculating the modulating factor. Defaults to 2.0
+ loss_weight=1.0), # Loss weight of current loss
+ loss_bbox=dict( # Config of loss function for the regression branch
+ type='GIoULoss', # Type of loss. Refer to https://mmdetection.readthedocs.io/en/latest/api.html#mmdet.models.losses.GIoULoss
+ loss_weight=2.0), # Loss weight of the regression branch
+ with_objectness=False, # Whether to add an objectness branch. Defaults to True
+ exp_on_reg=True, # Whether to use .exp() in regression
+ share_conv=True, # Whether to share conv layers between stages. Defaults to True
+ pred_kernel_size=1, # Kernel size of prediction layer. Defaults to 1
+ norm_cfg=dict(type='SyncBN'), # Config dict for normalization layer. Defaults to dict(type='BN', momentum=0.03, eps=0.001)
+ act_cfg=dict(type='SiLU', inplace=True)), # Config dict for activation layer. Defaults to dict(type='SiLU')
+ train_cfg=dict( # Config of training hyperparameters for ATSS
+ assigner=dict( # Config of assigner
+ type='DynamicSoftLabelAssigner', # Type of assigner. DynamicSoftLabelAssigner computes matching between predictions and ground truth with dynamic soft label assignment. Refer to https://github.com/open-mmlab/mmdetection/blob/main/mmdet/models/task_modules/assigners/dynamic_soft_label_assigner.py#L40
+ topk=13), # Select top-k predictions to calculate dynamic k best matches for each gt. Defaults to 13
+ allowed_border=-1, # The border allowed after padding for valid anchors
+ pos_weight=-1, # The weight of positive samples during training
+ debug=False), # Whether to set the debug mode
+ test_cfg=dict( # Config for testing hyperparameters for ATSS
+ nms_pre=30000, # The number of boxes before NMS
+ min_bbox_size=0, # The allowed minimal box size
+ score_thr=0.001, # Threshold to filter out boxes
+ nms=dict( # Config of NMS in the second stage
+ type='nms', # Type of NMS
+ iou_threshold=0.65), # NMS threshold
+ max_per_img=300), # Max number of detections of each image
+)
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/classification/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/classification/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..acc127db2ca82b2cbc5fe93495306c2776acaf33
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/classification/README.md
@@ -0,0 +1,56 @@
+# CSPNeXt ImageNet Pre-training
+
+In this folder, we provide the imagenet pre-training config of RTMDet's backbone CSPNeXt.
+
+## Requirements
+
+To train with these configs, please install [MMPreTrain](https://github.com/open-mmlab/mmpretrain) first.
+
+Install by MIM:
+
+```shell
+mim install mmpretrain
+```
+
+or install by pip:
+
+```shell
+pip install mmpretrain
+```
+
+## Prepare Dataset
+
+To pre-train on ImageNet, you need to prepare the dataset first. Please refer to the [guide](https://mmpretrain.readthedocs.io/en/latest/user_guides/dataset_prepare.html#imagenet).
+
+## How to Train
+
+You can use the classification config in the same way as the detection config.
+
+For single-GPU training, run:
+
+```shell
+python tools/train.py \
+ ${CONFIG_FILE} \
+ [optional arguments]
+```
+
+For multi-GPU training, run:
+
+```shell
+bash ./tools/dist_train.sh \
+ ${CONFIG_FILE} \
+ ${GPU_NUM} \
+ [optional arguments]
+```
+
+More details can be found in [user guides](https://mmdetection.readthedocs.io/en/latest/user_guides/train.html).
+
+## Results and Models
+
+| Model | resolution | Params(M) | Flops(G) | Top-1 (%) | Top-5 (%) | Download |
+| :----------: | :--------: | :-------: | :------: | :-------: | :-------: | :---------------------------------------------------------------------------------------------------------------------------------: |
+| CSPNeXt-tiny | 224x224 | 2.73 | 0.34 | 69.44 | 89.45 | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/cspnext_rsb_pretrain/cspnext-tiny_imagenet_600e-3a2dd350.pth) |
+| CSPNeXt-s | 224x224 | 4.89 | 0.66 | 74.41 | 92.23 | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/cspnext_rsb_pretrain/cspnext-s_imagenet_600e-ea671761.pth) |
+| CSPNeXt-m | 224x224 | 13.05 | 1.93 | 79.27 | 94.79 | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/cspnext_rsb_pretrain/cspnext-m_8xb256-rsb-a1-600e_in1k-ecb3bbd9.pth) |
+| CSPNeXt-l | 224x224 | 27.16 | 4.19 | 81.30 | 95.62 | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/cspnext_rsb_pretrain/cspnext-l_8xb256-rsb-a1-600e_in1k-6a760974.pth) |
+| CSPNeXt-x | 224x224 | 48.85 | 7.76 | 82.10 | 95.69 | [model](https://download.openmmlab.com/mmdetection/v3.0/rtmdet/cspnext_rsb_pretrain/cspnext-x_8xb256-rsb-a1-600e_in1k-b3f78edd.pth) |
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/classification/cspnext-l_8xb256-rsb-a1-600e_in1k.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/classification/cspnext-l_8xb256-rsb-a1-600e_in1k.py
new file mode 100644
index 0000000000000000000000000000000000000000..d2e70539f05da69cca53f273d11e3296c87c4eda
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/classification/cspnext-l_8xb256-rsb-a1-600e_in1k.py
@@ -0,0 +1,5 @@
+_base_ = './cspnext-s_8xb256-rsb-a1-600e_in1k.py'
+
+model = dict(
+ backbone=dict(deepen_factor=1, widen_factor=1),
+ head=dict(in_channels=1024))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/classification/cspnext-m_8xb256-rsb-a1-600e_in1k.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/classification/cspnext-m_8xb256-rsb-a1-600e_in1k.py
new file mode 100644
index 0000000000000000000000000000000000000000..e1b1352dd91a803eeafe80f587203f96a247c27f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/classification/cspnext-m_8xb256-rsb-a1-600e_in1k.py
@@ -0,0 +1,5 @@
+_base_ = './cspnext-s_8xb256-rsb-a1-600e_in1k.py'
+
+model = dict(
+ backbone=dict(deepen_factor=0.67, widen_factor=0.75),
+ head=dict(in_channels=768))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/classification/cspnext-s_8xb256-rsb-a1-600e_in1k.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/classification/cspnext-s_8xb256-rsb-a1-600e_in1k.py
new file mode 100644
index 0000000000000000000000000000000000000000..dcfd2ea47d54408ef6d2fe225b57c5c9e540918a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/classification/cspnext-s_8xb256-rsb-a1-600e_in1k.py
@@ -0,0 +1,64 @@
+_base_ = [
+ 'mmpretrain::_base_/datasets/imagenet_bs256_rsb_a12.py',
+ 'mmpretrain::_base_/schedules/imagenet_bs2048_rsb.py',
+ 'mmpretrain::_base_/default_runtime.py'
+]
+
+model = dict(
+ type='ImageClassifier',
+ backbone=dict(
+ type='mmdet.CSPNeXt',
+ arch='P5',
+ out_indices=(4, ),
+ expand_ratio=0.5,
+ deepen_factor=0.33,
+ widen_factor=0.5,
+ channel_attention=True,
+ norm_cfg=dict(type='BN'),
+ act_cfg=dict(type='mmdet.SiLU')),
+ neck=dict(type='GlobalAveragePooling'),
+ head=dict(
+ type='LinearClsHead',
+ num_classes=1000,
+ in_channels=512,
+ loss=dict(
+ type='LabelSmoothLoss',
+ label_smooth_val=0.1,
+ mode='original',
+ loss_weight=1.0),
+ topk=(1, 5)),
+ train_cfg=dict(augments=[
+ dict(type='Mixup', alpha=0.2),
+ dict(type='CutMix', alpha=1.0)
+ ]))
+
+# dataset settings
+train_dataloader = dict(sampler=dict(type='RepeatAugSampler', shuffle=True))
+
+# schedule settings
+optim_wrapper = dict(
+ optimizer=dict(weight_decay=0.01),
+ paramwise_cfg=dict(bias_decay_mult=0., norm_decay_mult=0.),
+)
+
+param_scheduler = [
+ # warm up learning rate scheduler
+ dict(
+ type='LinearLR',
+ start_factor=0.0001,
+ by_epoch=True,
+ begin=0,
+ end=5,
+ # update by iter
+ convert_to_iter_based=True),
+ # main learning rate scheduler
+ dict(
+ type='CosineAnnealingLR',
+ T_max=595,
+ eta_min=1.0e-6,
+ by_epoch=True,
+ begin=5,
+ end=600)
+]
+
+train_cfg = dict(by_epoch=True, max_epochs=600)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/classification/cspnext-tiny_8xb256-rsb-a1-600e_in1k.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/classification/cspnext-tiny_8xb256-rsb-a1-600e_in1k.py
new file mode 100644
index 0000000000000000000000000000000000000000..af3170bdc51778c4601d4426aa88cc27c608f100
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/classification/cspnext-tiny_8xb256-rsb-a1-600e_in1k.py
@@ -0,0 +1,5 @@
+_base_ = './cspnext-s_8xb256-rsb-a1-600e_in1k.py'
+
+model = dict(
+ backbone=dict(deepen_factor=0.167, widen_factor=0.375),
+ head=dict(in_channels=384))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/classification/cspnext-x_8xb256-rsb-a1-600e_in1k.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/classification/cspnext-x_8xb256-rsb-a1-600e_in1k.py
new file mode 100644
index 0000000000000000000000000000000000000000..edec48d78dbefdb7783c5dd50e97873e29ea6497
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/classification/cspnext-x_8xb256-rsb-a1-600e_in1k.py
@@ -0,0 +1,5 @@
+_base_ = './cspnext-s_8xb256-rsb-a1-600e_in1k.py'
+
+model = dict(
+ backbone=dict(deepen_factor=1.33, widen_factor=1.25),
+ head=dict(in_channels=1280))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..a62abcb2faabb2e7d6c4a6c7d3b492392eba9775
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/metafile.yml
@@ -0,0 +1,242 @@
+Collections:
+ - Name: RTMDet
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - AdamW
+ - Flat Cosine Annealing
+ Training Resources: 8x A100 GPUs
+ Architecture:
+ - CSPNeXt
+ - CSPNeXtPAFPN
+ README: configs/rtmdet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v3.0.0rc1/mmdet/models/detectors/rtmdet.py#L6
+ Version: v3.0.0rc1
+
+Models:
+ - Name: rtmdet_tiny_8xb32-300e_coco
+ Alias:
+ - rtmdet-t
+ In Collection: RTMDet
+ Config: configs/rtmdet/rtmdet_tiny_8xb32-300e_coco.py
+ Metadata:
+ Training Memory (GB): 11.7
+ Epochs: 300
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.9
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet_tiny_8xb32-300e_coco/rtmdet_tiny_8xb32-300e_coco_20220902_112414-78e30dcc.pth
+
+ - Name: rtmdet_s_8xb32-300e_coco
+ Alias:
+ - rtmdet-s
+ In Collection: RTMDet
+ Config: configs/rtmdet/rtmdet_s_8xb32-300e_coco.py
+ Metadata:
+ Training Memory (GB): 15.9
+ Epochs: 300
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.5
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet_s_8xb32-300e_coco/rtmdet_s_8xb32-300e_coco_20220905_161602-387a891e.pth
+
+ - Name: rtmdet_m_8xb32-300e_coco
+ Alias:
+ - rtmdet-m
+ In Collection: RTMDet
+ Config: configs/rtmdet/rtmdet_m_8xb32-300e_coco.py
+ Metadata:
+ Training Memory (GB): 27.8
+ Epochs: 300
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 49.1
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet_m_8xb32-300e_coco/rtmdet_m_8xb32-300e_coco_20220719_112220-229f527c.pth
+
+ - Name: rtmdet_l_8xb32-300e_coco
+ Alias:
+ - rtmdet-l
+ In Collection: RTMDet
+ Config: configs/rtmdet/rtmdet_l_8xb32-300e_coco.py
+ Metadata:
+ Training Memory (GB): 43.2
+ Epochs: 300
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 51.3
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet_l_8xb32-300e_coco/rtmdet_l_8xb32-300e_coco_20220719_112030-5a0be7c4.pth
+
+ - Name: rtmdet_x_8xb32-300e_coco
+ Alias:
+ - rtmdet-x
+ In Collection: RTMDet
+ Config: configs/rtmdet/rtmdet_x_8xb32-300e_coco.py
+ Metadata:
+ Training Memory (GB): 61.1
+ Epochs: 300
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 52.6
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet_x_8xb32-300e_coco/rtmdet_x_8xb32-300e_coco_20220715_230555-cc79b9ae.pth
+
+ - Name: rtmdet_x_p6_4xb8-300e_coco
+ Alias:
+ - rtmdet-x_p6
+ In Collection: RTMDet
+ Config: configs/rtmdet/rtmdet_x_p6_4xb8-300e_coco.py
+ Metadata:
+ Epochs: 300
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 54.9
+ Weights: https://github.com/orange0-jp/orange-weights/releases/download/v0.1.0rtmdet-p6/rtmdet_x_p6_4xb8-300e_coco-bf32be58.pth
+
+ - Name: rtmdet_l_convnext_b_4xb32-100e_coco
+ Alias:
+ - rtmdet-l_convnext_b
+ In Collection: RTMDet
+ Config: configs/rtmdet/rtmdet_l_convnext_b_4xb32-100e_coco.py
+ Metadata:
+ Epochs: 100
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 53.1
+ Weights: https://github.com/orange0-jp/orange-weights/releases/download/v0.1.0rtmdet-swin-convnext/rtmdet_l_convnext_b_4xb32-100e_coco-d4731b3d.pth
+
+ - Name: rtmdet_l_swin_b_4xb32-100e_coco
+ Alias:
+ - rtmdet-l_swin_b
+ In Collection: RTMDet
+ Config: configs/rtmdet/rtmdet_l_swin_b_4xb32-100e_coco.py
+ Metadata:
+ Epochs: 100
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 52.4
+ Weights: https://github.com/orange0-jp/orange-weights/releases/download/v0.1.0rtmdet-swin-convnext/rtmdet_l_swin_b_4xb32-100e_coco-0828ce5d.pth
+
+ - Name: rtmdet_l_swin_b_p6_4xb16-100e_coco
+ Alias:
+ - rtmdet-l_swin_b_p6
+ In Collection: RTMDet
+ Config: configs/rtmdet/rtmdet_l_swin_b_p6_4xb16-100e_coco.py
+ Metadata:
+ Epochs: 100
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 56.4
+ Weights: https://github.com/orange0-jp/orange-weights/releases/download/v0.1.0rtmdet-swin-convnext/rtmdet_l_swin_b_p6_4xb16-100e_coco-a1486b6f.pth
+
+ - Name: rtmdet-ins_tiny_8xb32-300e_coco
+ Alias:
+ - rtmdet-ins-t
+ In Collection: RTMDet
+ Config: configs/rtmdet/rtmdet-ins_tiny_8xb32-300e_coco.py
+ Metadata:
+ Training Memory (GB): 18.4
+ Epochs: 300
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 35.4
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet-ins_tiny_8xb32-300e_coco/rtmdet-ins_tiny_8xb32-300e_coco_20221130_151727-ec670f7e.pth
+
+ - Name: rtmdet-ins_s_8xb32-300e_coco
+ Alias:
+ - rtmdet-ins-s
+ In Collection: RTMDet
+ Config: configs/rtmdet/rtmdet-ins_s_8xb32-300e_coco.py
+ Metadata:
+ Training Memory (GB): 27.6
+ Epochs: 300
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 38.7
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet-ins_s_8xb32-300e_coco/rtmdet-ins_s_8xb32-300e_coco_20221121_212604-fdc5d7ec.pth
+
+ - Name: rtmdet-ins_m_8xb32-300e_coco
+ Alias:
+ - rtmdet-ins-m
+ In Collection: RTMDet
+ Config: configs/rtmdet/rtmdet-ins_m_8xb32-300e_coco.py
+ Metadata:
+ Training Memory (GB): 42.5
+ Epochs: 300
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 48.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 42.1
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet-ins_m_8xb32-300e_coco/rtmdet-ins_m_8xb32-300e_coco_20221123_001039-6eba602e.pth
+
+ - Name: rtmdet-ins_l_8xb32-300e_coco
+ Alias:
+ - rtmdet-ins-l
+ In Collection: RTMDet
+ Config: configs/rtmdet/rtmdet-ins_l_8xb32-300e_coco.py
+ Metadata:
+ Training Memory (GB): 59.8
+ Epochs: 300
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 51.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 43.7
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet-ins_l_8xb32-300e_coco/rtmdet-ins_l_8xb32-300e_coco_20221124_103237-78d1d652.pth
+
+ - Name: rtmdet-ins_x_8xb16-300e_coco
+ Alias:
+ - rtmdet-ins-x
+ In Collection: RTMDet
+ Config: configs/rtmdet/rtmdet-ins_x_8xb16-300e_coco.py
+ Metadata:
+ Training Memory (GB): 33.7
+ Epochs: 300
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 52.4
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 44.6
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/rtmdet/rtmdet-ins_x_8xb16-300e_coco/rtmdet-ins_x_8xb16-300e_coco_20221124_111313-33d4595b.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet-ins_l_8xb32-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet-ins_l_8xb32-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6b4b9240a64d39d8a16352ef87de53af9e81ac96
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet-ins_l_8xb32-300e_coco.py
@@ -0,0 +1,104 @@
+_base_ = './rtmdet_l_8xb32-300e_coco.py'
+model = dict(
+ bbox_head=dict(
+ _delete_=True,
+ type='RTMDetInsSepBNHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=2,
+ share_conv=True,
+ pred_kernel_size=1,
+ feat_channels=256,
+ act_cfg=dict(type='SiLU', inplace=True),
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ anchor_generator=dict(
+ type='MlvlPointGenerator', offset=0, strides=[8, 16, 32]),
+ bbox_coder=dict(type='DistancePointBBoxCoder'),
+ loss_cls=dict(
+ type='QualityFocalLoss',
+ use_sigmoid=True,
+ beta=2.0,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=2.0),
+ loss_mask=dict(
+ type='DiceLoss', loss_weight=2.0, eps=5e-6, reduction='mean')),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100,
+ mask_thr_binary=0.5),
+)
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ poly2mask=False),
+ dict(type='CachedMosaic', img_scale=(640, 640), pad_val=114.0),
+ dict(
+ type='RandomResize',
+ scale=(1280, 1280),
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_size=(640, 640),
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(
+ type='CachedMixUp',
+ img_scale=(640, 640),
+ ratio_range=(1.0, 1.0),
+ max_cached_images=20,
+ pad_val=(114, 114, 114)),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1, 1)),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(pin_memory=True, dataset=dict(pipeline=train_pipeline))
+
+train_pipeline_stage2 = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ poly2mask=False),
+ dict(
+ type='RandomResize',
+ scale=(640, 640),
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_size=(640, 640),
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1, 1)),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(type='PackDetInputs')
+]
+custom_hooks = [
+ dict(
+ type='EMAHook',
+ ema_type='ExpMomentumEMA',
+ momentum=0.0002,
+ update_buffers=True,
+ priority=49),
+ dict(
+ type='PipelineSwitchHook',
+ switch_epoch=280,
+ switch_pipeline=train_pipeline_stage2)
+]
+
+val_evaluator = dict(metric=['bbox', 'segm'])
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet-ins_m_8xb32-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet-ins_m_8xb32-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..66da9148775b425c6b0052beb04f9c8ca17257d9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet-ins_m_8xb32-300e_coco.py
@@ -0,0 +1,6 @@
+_base_ = './rtmdet-ins_l_8xb32-300e_coco.py'
+
+model = dict(
+ backbone=dict(deepen_factor=0.67, widen_factor=0.75),
+ neck=dict(in_channels=[192, 384, 768], out_channels=192, num_csp_blocks=2),
+ bbox_head=dict(in_channels=192, feat_channels=192))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet-ins_s_8xb32-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet-ins_s_8xb32-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..28bc21cc93bb36d2d2fc8601b06bb0f0c58d6d49
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet-ins_s_8xb32-300e_coco.py
@@ -0,0 +1,80 @@
+_base_ = './rtmdet-ins_l_8xb32-300e_coco.py'
+checkpoint = 'https://download.openmmlab.com/mmdetection/v3.0/rtmdet/cspnext_rsb_pretrain/cspnext-s_imagenet_600e.pth' # noqa
+model = dict(
+ backbone=dict(
+ deepen_factor=0.33,
+ widen_factor=0.5,
+ init_cfg=dict(
+ type='Pretrained', prefix='backbone.', checkpoint=checkpoint)),
+ neck=dict(in_channels=[128, 256, 512], out_channels=128, num_csp_blocks=1),
+ bbox_head=dict(in_channels=128, feat_channels=128))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ poly2mask=False),
+ dict(type='CachedMosaic', img_scale=(640, 640), pad_val=114.0),
+ dict(
+ type='RandomResize',
+ scale=(1280, 1280),
+ ratio_range=(0.5, 2.0),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_size=(640, 640),
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(
+ type='CachedMixUp',
+ img_scale=(640, 640),
+ ratio_range=(1.0, 1.0),
+ max_cached_images=20,
+ pad_val=(114, 114, 114)),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1, 1)),
+ dict(type='PackDetInputs')
+]
+
+train_pipeline_stage2 = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ poly2mask=False),
+ dict(
+ type='RandomResize',
+ scale=(640, 640),
+ ratio_range=(0.5, 2.0),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_size=(640, 640),
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1, 1)),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+custom_hooks = [
+ dict(
+ type='EMAHook',
+ ema_type='ExpMomentumEMA',
+ momentum=0.0002,
+ update_buffers=True,
+ priority=49),
+ dict(
+ type='PipelineSwitchHook',
+ switch_epoch=280,
+ switch_pipeline=train_pipeline_stage2)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet-ins_tiny_8xb32-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet-ins_tiny_8xb32-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..954f911614e75eb9910effbf1bbc1d7b01120276
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet-ins_tiny_8xb32-300e_coco.py
@@ -0,0 +1,48 @@
+_base_ = './rtmdet-ins_s_8xb32-300e_coco.py'
+
+checkpoint = 'https://download.openmmlab.com/mmdetection/v3.0/rtmdet/cspnext_rsb_pretrain/cspnext-tiny_imagenet_600e.pth' # noqa
+
+model = dict(
+ backbone=dict(
+ deepen_factor=0.167,
+ widen_factor=0.375,
+ init_cfg=dict(
+ type='Pretrained', prefix='backbone.', checkpoint=checkpoint)),
+ neck=dict(in_channels=[96, 192, 384], out_channels=96, num_csp_blocks=1),
+ bbox_head=dict(in_channels=96, feat_channels=96))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ poly2mask=False),
+ dict(
+ type='CachedMosaic',
+ img_scale=(640, 640),
+ pad_val=114.0,
+ max_cached_images=20,
+ random_pop=False),
+ dict(
+ type='RandomResize',
+ scale=(1280, 1280),
+ ratio_range=(0.5, 2.0),
+ keep_ratio=True),
+ dict(type='RandomCrop', crop_size=(640, 640)),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(
+ type='CachedMixUp',
+ img_scale=(640, 640),
+ ratio_range=(1.0, 1.0),
+ max_cached_images=10,
+ random_pop=False,
+ pad_val=(114, 114, 114),
+ prob=0.5),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1, 1)),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet-ins_x_8xb16-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet-ins_x_8xb16-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..daaa640edac6b2114caf13b650d99d7c7632629a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet-ins_x_8xb16-300e_coco.py
@@ -0,0 +1,31 @@
+_base_ = './rtmdet-ins_l_8xb32-300e_coco.py'
+
+model = dict(
+ backbone=dict(deepen_factor=1.33, widen_factor=1.25),
+ neck=dict(
+ in_channels=[320, 640, 1280], out_channels=320, num_csp_blocks=4),
+ bbox_head=dict(in_channels=320, feat_channels=320))
+
+base_lr = 0.002
+
+# optimizer
+optim_wrapper = dict(optimizer=dict(lr=base_lr))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0e-5,
+ by_epoch=False,
+ begin=0,
+ end=1000),
+ dict(
+ # use cosine lr from 150 to 300 epoch
+ type='CosineAnnealingLR',
+ eta_min=base_lr * 0.05,
+ begin=_base_.max_epochs // 2,
+ end=_base_.max_epochs,
+ T_max=_base_.max_epochs // 2,
+ by_epoch=True,
+ convert_to_iter_based=True),
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_l_8xb32-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_l_8xb32-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1cce4d89c84a81d7aa22197cd6dd70fe08637a35
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_l_8xb32-300e_coco.py
@@ -0,0 +1,179 @@
+_base_ = [
+ '../_base_/default_runtime.py', '../_base_/schedules/schedule_1x.py',
+ '../_base_/datasets/coco_detection.py', './rtmdet_tta.py'
+]
+model = dict(
+ type='RTMDet',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.53, 116.28, 123.675],
+ std=[57.375, 57.12, 58.395],
+ bgr_to_rgb=False,
+ batch_augments=None),
+ backbone=dict(
+ type='CSPNeXt',
+ arch='P5',
+ expand_ratio=0.5,
+ deepen_factor=1,
+ widen_factor=1,
+ channel_attention=True,
+ norm_cfg=dict(type='SyncBN'),
+ act_cfg=dict(type='SiLU', inplace=True)),
+ neck=dict(
+ type='CSPNeXtPAFPN',
+ in_channels=[256, 512, 1024],
+ out_channels=256,
+ num_csp_blocks=3,
+ expand_ratio=0.5,
+ norm_cfg=dict(type='SyncBN'),
+ act_cfg=dict(type='SiLU', inplace=True)),
+ bbox_head=dict(
+ type='RTMDetSepBNHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=2,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='MlvlPointGenerator', offset=0, strides=[8, 16, 32]),
+ bbox_coder=dict(type='DistancePointBBoxCoder'),
+ loss_cls=dict(
+ type='QualityFocalLoss',
+ use_sigmoid=True,
+ beta=2.0,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=2.0),
+ with_objectness=False,
+ exp_on_reg=True,
+ share_conv=True,
+ pred_kernel_size=1,
+ norm_cfg=dict(type='SyncBN'),
+ act_cfg=dict(type='SiLU', inplace=True)),
+ train_cfg=dict(
+ assigner=dict(type='DynamicSoftLabelAssigner', topk=13),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=30000,
+ min_bbox_size=0,
+ score_thr=0.001,
+ nms=dict(type='nms', iou_threshold=0.65),
+ max_per_img=300),
+)
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='CachedMosaic', img_scale=(640, 640), pad_val=114.0),
+ dict(
+ type='RandomResize',
+ scale=(1280, 1280),
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(type='RandomCrop', crop_size=(640, 640)),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(
+ type='CachedMixUp',
+ img_scale=(640, 640),
+ ratio_range=(1.0, 1.0),
+ max_cached_images=20,
+ pad_val=(114, 114, 114)),
+ dict(type='PackDetInputs')
+]
+
+train_pipeline_stage2 = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize',
+ scale=(640, 640),
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(type='RandomCrop', crop_size=(640, 640)),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(640, 640), keep_ratio=True),
+ dict(type='Pad', size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=32,
+ num_workers=10,
+ batch_sampler=None,
+ pin_memory=True,
+ dataset=dict(pipeline=train_pipeline))
+val_dataloader = dict(
+ batch_size=5, num_workers=10, dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+max_epochs = 300
+stage2_num_epochs = 20
+base_lr = 0.004
+interval = 10
+
+train_cfg = dict(
+ max_epochs=max_epochs,
+ val_interval=interval,
+ dynamic_intervals=[(max_epochs - stage2_num_epochs, 1)])
+
+val_evaluator = dict(proposal_nums=(100, 1, 10))
+test_evaluator = val_evaluator
+
+# optimizer
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=base_lr, weight_decay=0.05),
+ paramwise_cfg=dict(
+ norm_decay_mult=0, bias_decay_mult=0, bypass_duplicate=True))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0e-5,
+ by_epoch=False,
+ begin=0,
+ end=1000),
+ dict(
+ # use cosine lr from 150 to 300 epoch
+ type='CosineAnnealingLR',
+ eta_min=base_lr * 0.05,
+ begin=max_epochs // 2,
+ end=max_epochs,
+ T_max=max_epochs // 2,
+ by_epoch=True,
+ convert_to_iter_based=True),
+]
+
+# hooks
+default_hooks = dict(
+ checkpoint=dict(
+ interval=interval,
+ max_keep_ckpts=3 # only keep latest 3 checkpoints
+ ))
+custom_hooks = [
+ dict(
+ type='EMAHook',
+ ema_type='ExpMomentumEMA',
+ momentum=0.0002,
+ update_buffers=True,
+ priority=49),
+ dict(
+ type='PipelineSwitchHook',
+ switch_epoch=max_epochs - stage2_num_epochs,
+ switch_pipeline=train_pipeline_stage2)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_l_convnext_b_4xb32-100e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_l_convnext_b_4xb32-100e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..85af292bcaba2e1853ed4f3a3f5818c0c0d5813e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_l_convnext_b_4xb32-100e_coco.py
@@ -0,0 +1,81 @@
+_base_ = './rtmdet_l_8xb32-300e_coco.py'
+
+custom_imports = dict(
+ imports=['mmpretrain.models'], allow_failed_imports=False)
+
+norm_cfg = dict(type='GN', num_groups=32)
+checkpoint_file = 'https://download.openmmlab.com/mmclassification/v0/convnext/convnext-base_in21k-pre-3rdparty_in1k-384px_20221219-4570f792.pth' # noqa
+model = dict(
+ type='RTMDet',
+ data_preprocessor=dict(
+ _delete_=True,
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ batch_augments=None),
+ backbone=dict(
+ _delete_=True,
+ type='mmpretrain.ConvNeXt',
+ arch='base',
+ out_indices=[1, 2, 3],
+ drop_path_rate=0.7,
+ layer_scale_init_value=1.0,
+ gap_before_final_norm=False,
+ with_cp=True,
+ init_cfg=dict(
+ type='Pretrained', checkpoint=checkpoint_file,
+ prefix='backbone.')),
+ neck=dict(in_channels=[256, 512, 1024], norm_cfg=norm_cfg),
+ bbox_head=dict(norm_cfg=norm_cfg))
+
+max_epochs = 100
+stage2_num_epochs = 10
+interval = 10
+base_lr = 0.001
+
+train_cfg = dict(
+ max_epochs=max_epochs,
+ val_interval=interval,
+ dynamic_intervals=[(max_epochs - stage2_num_epochs, 1)])
+
+optim_wrapper = dict(
+ constructor='LearningRateDecayOptimizerConstructor',
+ paramwise_cfg={
+ 'decay_rate': 0.8,
+ 'decay_type': 'layer_wise',
+ 'num_layers': 12
+ },
+ optimizer=dict(lr=base_lr))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0e-5,
+ by_epoch=False,
+ begin=0,
+ end=1000),
+ dict(
+ # use cosine lr from 50 to 100 epoch
+ type='CosineAnnealingLR',
+ eta_min=base_lr * 0.05,
+ begin=max_epochs // 2,
+ end=max_epochs,
+ T_max=max_epochs // 2,
+ by_epoch=True,
+ convert_to_iter_based=True),
+]
+
+custom_hooks = [
+ dict(
+ type='EMAHook',
+ ema_type='ExpMomentumEMA',
+ momentum=0.0002,
+ update_buffers=True,
+ priority=49),
+ dict(
+ type='PipelineSwitchHook',
+ switch_epoch=max_epochs - stage2_num_epochs,
+ switch_pipeline={{_base_.train_pipeline_stage2}})
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_l_swin_b_4xb32-100e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_l_swin_b_4xb32-100e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..84b0e0fa7d18848a4c1e305985e33e69e3196790
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_l_swin_b_4xb32-100e_coco.py
@@ -0,0 +1,78 @@
+_base_ = './rtmdet_l_8xb32-300e_coco.py'
+
+norm_cfg = dict(type='GN', num_groups=32)
+checkpoint = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_base_patch4_window12_384_22k.pth' # noqa
+model = dict(
+ type='RTMDet',
+ data_preprocessor=dict(
+ _delete_=True,
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ batch_augments=None),
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ pretrain_img_size=384,
+ embed_dims=128,
+ depths=[2, 2, 18, 2],
+ num_heads=[4, 8, 16, 32],
+ window_size=12,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(1, 2, 3),
+ with_cp=True,
+ convert_weights=True,
+ init_cfg=dict(type='Pretrained', checkpoint=checkpoint)),
+ neck=dict(in_channels=[256, 512, 1024], norm_cfg=norm_cfg),
+ bbox_head=dict(norm_cfg=norm_cfg))
+
+max_epochs = 100
+stage2_num_epochs = 10
+interval = 10
+base_lr = 0.001
+
+train_cfg = dict(
+ max_epochs=max_epochs,
+ val_interval=interval,
+ dynamic_intervals=[(max_epochs - stage2_num_epochs, 1)])
+
+optim_wrapper = dict(optimizer=dict(lr=base_lr))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0e-5,
+ by_epoch=False,
+ begin=0,
+ end=1000),
+ dict(
+ # use cosine lr from 50 to 100 epoch
+ type='CosineAnnealingLR',
+ eta_min=base_lr * 0.05,
+ begin=max_epochs // 2,
+ end=max_epochs,
+ T_max=max_epochs // 2,
+ by_epoch=True,
+ convert_to_iter_based=True),
+]
+
+custom_hooks = [
+ dict(
+ type='EMAHook',
+ ema_type='ExpMomentumEMA',
+ momentum=0.0002,
+ update_buffers=True,
+ priority=49),
+ dict(
+ type='PipelineSwitchHook',
+ switch_epoch=max_epochs - stage2_num_epochs,
+ switch_pipeline={{_base_.train_pipeline_stage2}})
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_l_swin_b_p6_4xb16-100e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_l_swin_b_p6_4xb16-100e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..37d4215c3f014ef20c7817875cbc1689186e0766
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_l_swin_b_p6_4xb16-100e_coco.py
@@ -0,0 +1,114 @@
+_base_ = './rtmdet_l_swin_b_4xb32-100e_coco.py'
+
+model = dict(
+ backbone=dict(
+ depths=[2, 2, 18, 2, 1],
+ num_heads=[4, 8, 16, 32, 64],
+ strides=(4, 2, 2, 2, 2),
+ out_indices=(1, 2, 3, 4)),
+ neck=dict(in_channels=[256, 512, 1024, 2048]),
+ bbox_head=dict(
+ anchor_generator=dict(
+ type='MlvlPointGenerator', offset=0, strides=[8, 16, 32, 64])))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='CachedMosaic', img_scale=(1280, 1280), pad_val=114.0),
+ dict(
+ type='RandomResize',
+ scale=(2560, 2560),
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(type='RandomCrop', crop_size=(1280, 1280)),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size=(1280, 1280), pad_val=dict(img=(114, 114, 114))),
+ dict(
+ type='CachedMixUp',
+ img_scale=(1280, 1280),
+ ratio_range=(1.0, 1.0),
+ max_cached_images=20,
+ pad_val=(114, 114, 114)),
+ dict(type='PackDetInputs')
+]
+
+train_pipeline_stage2 = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize',
+ scale=(1280, 1280),
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(type='RandomCrop', crop_size=(1280, 1280)),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size=(1280, 1280), pad_val=dict(img=(114, 114, 114))),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(1280, 1280), keep_ratio=True),
+ dict(type='Pad', size=(1280, 1280), pad_val=dict(img=(114, 114, 114))),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=16, num_workers=20, dataset=dict(pipeline=train_pipeline))
+val_dataloader = dict(num_workers=20, dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+max_epochs = 100
+stage2_num_epochs = 10
+
+custom_hooks = [
+ dict(
+ type='EMAHook',
+ ema_type='ExpMomentumEMA',
+ momentum=0.0002,
+ update_buffers=True,
+ priority=49),
+ dict(
+ type='PipelineSwitchHook',
+ switch_epoch=max_epochs - stage2_num_epochs,
+ switch_pipeline=train_pipeline_stage2)
+]
+
+img_scales = [(1280, 1280), (640, 640), (1920, 1920)]
+tta_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=None),
+ dict(
+ type='TestTimeAug',
+ transforms=[
+ [
+ dict(type='Resize', scale=s, keep_ratio=True)
+ for s in img_scales
+ ],
+ [
+ # ``RandomFlip`` must be placed before ``Pad``, otherwise
+ # bounding box coordinates after flipping cannot be
+ # recovered correctly.
+ dict(type='RandomFlip', prob=1.),
+ dict(type='RandomFlip', prob=0.)
+ ],
+ [
+ dict(
+ type='Pad',
+ size=(1920, 1920),
+ pad_val=dict(img=(114, 114, 114))),
+ ],
+ [dict(type='LoadAnnotations', with_bbox=True)],
+ [
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction'))
+ ]
+ ])
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_m_8xb32-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_m_8xb32-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c83f5a60bd7d9f85f46574ee4cd19027391b5e1e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_m_8xb32-300e_coco.py
@@ -0,0 +1,6 @@
+_base_ = './rtmdet_l_8xb32-300e_coco.py'
+
+model = dict(
+ backbone=dict(deepen_factor=0.67, widen_factor=0.75),
+ neck=dict(in_channels=[192, 384, 768], out_channels=192, num_csp_blocks=2),
+ bbox_head=dict(in_channels=192, feat_channels=192))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_s_8xb32-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_s_8xb32-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..cbf76247b74e94735eea0dd70ce6ac9e57f4dadf
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_s_8xb32-300e_coco.py
@@ -0,0 +1,62 @@
+_base_ = './rtmdet_l_8xb32-300e_coco.py'
+checkpoint = 'https://download.openmmlab.com/mmdetection/v3.0/rtmdet/cspnext_rsb_pretrain/cspnext-s_imagenet_600e.pth' # noqa
+model = dict(
+ backbone=dict(
+ deepen_factor=0.33,
+ widen_factor=0.5,
+ init_cfg=dict(
+ type='Pretrained', prefix='backbone.', checkpoint=checkpoint)),
+ neck=dict(in_channels=[128, 256, 512], out_channels=128, num_csp_blocks=1),
+ bbox_head=dict(in_channels=128, feat_channels=128, exp_on_reg=False))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='CachedMosaic', img_scale=(640, 640), pad_val=114.0),
+ dict(
+ type='RandomResize',
+ scale=(1280, 1280),
+ ratio_range=(0.5, 2.0),
+ keep_ratio=True),
+ dict(type='RandomCrop', crop_size=(640, 640)),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(
+ type='CachedMixUp',
+ img_scale=(640, 640),
+ ratio_range=(1.0, 1.0),
+ max_cached_images=20,
+ pad_val=(114, 114, 114)),
+ dict(type='PackDetInputs')
+]
+
+train_pipeline_stage2 = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize',
+ scale=(640, 640),
+ ratio_range=(0.5, 2.0),
+ keep_ratio=True),
+ dict(type='RandomCrop', crop_size=(640, 640)),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+custom_hooks = [
+ dict(
+ type='EMAHook',
+ ema_type='ExpMomentumEMA',
+ momentum=0.0002,
+ update_buffers=True,
+ priority=49),
+ dict(
+ type='PipelineSwitchHook',
+ switch_epoch=280,
+ switch_pipeline=train_pipeline_stage2)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_tiny_8xb32-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_tiny_8xb32-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a686f4a7f0c4c3bed956c2a3fa504ea8863c669d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_tiny_8xb32-300e_coco.py
@@ -0,0 +1,43 @@
+_base_ = './rtmdet_s_8xb32-300e_coco.py'
+
+checkpoint = 'https://download.openmmlab.com/mmdetection/v3.0/rtmdet/cspnext_rsb_pretrain/cspnext-tiny_imagenet_600e.pth' # noqa
+
+model = dict(
+ backbone=dict(
+ deepen_factor=0.167,
+ widen_factor=0.375,
+ init_cfg=dict(
+ type='Pretrained', prefix='backbone.', checkpoint=checkpoint)),
+ neck=dict(in_channels=[96, 192, 384], out_channels=96, num_csp_blocks=1),
+ bbox_head=dict(in_channels=96, feat_channels=96, exp_on_reg=False))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='CachedMosaic',
+ img_scale=(640, 640),
+ pad_val=114.0,
+ max_cached_images=20,
+ random_pop=False),
+ dict(
+ type='RandomResize',
+ scale=(1280, 1280),
+ ratio_range=(0.5, 2.0),
+ keep_ratio=True),
+ dict(type='RandomCrop', crop_size=(640, 640)),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(
+ type='CachedMixUp',
+ img_scale=(640, 640),
+ ratio_range=(1.0, 1.0),
+ max_cached_images=10,
+ random_pop=False,
+ pad_val=(114, 114, 114),
+ prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_tta.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_tta.py
new file mode 100644
index 0000000000000000000000000000000000000000..6dde36de3ff06576944a351de9daf53746103f21
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_tta.py
@@ -0,0 +1,36 @@
+tta_model = dict(
+ type='DetTTAModel',
+ tta_cfg=dict(nms=dict(type='nms', iou_threshold=0.6), max_per_img=100))
+
+img_scales = [(640, 640), (320, 320), (960, 960)]
+tta_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=None),
+ dict(
+ type='TestTimeAug',
+ transforms=[
+ [
+ dict(type='Resize', scale=s, keep_ratio=True)
+ for s in img_scales
+ ],
+ [
+ # ``RandomFlip`` must be placed before ``Pad``, otherwise
+ # bounding box coordinates after flipping cannot be
+ # recovered correctly.
+ dict(type='RandomFlip', prob=1.),
+ dict(type='RandomFlip', prob=0.)
+ ],
+ [
+ dict(
+ type='Pad',
+ size=(960, 960),
+ pad_val=dict(img=(114, 114, 114))),
+ ],
+ [dict(type='LoadAnnotations', with_bbox=True)],
+ [
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction'))
+ ]
+ ])
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_x_8xb32-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_x_8xb32-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..16a33632c00b19b270b237f5dcd8f603350ac0c9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_x_8xb32-300e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './rtmdet_l_8xb32-300e_coco.py'
+
+model = dict(
+ backbone=dict(deepen_factor=1.33, widen_factor=1.25),
+ neck=dict(
+ in_channels=[320, 640, 1280], out_channels=320, num_csp_blocks=4),
+ bbox_head=dict(in_channels=320, feat_channels=320))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_x_p6_4xb8-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_x_p6_4xb8-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d1bb7fa6a78812e5a415acfb60eccedae9b884e2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/rtmdet/rtmdet_x_p6_4xb8-300e_coco.py
@@ -0,0 +1,132 @@
+_base_ = './rtmdet_x_8xb32-300e_coco.py'
+
+model = dict(
+ backbone=dict(arch='P6', out_indices=(2, 3, 4, 5)),
+ neck=dict(in_channels=[320, 640, 960, 1280]),
+ bbox_head=dict(
+ anchor_generator=dict(
+ type='MlvlPointGenerator', offset=0, strides=[8, 16, 32, 64])))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='CachedMosaic', img_scale=(1280, 1280), pad_val=114.0),
+ dict(
+ type='RandomResize',
+ scale=(2560, 2560),
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(type='RandomCrop', crop_size=(1280, 1280)),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size=(1280, 1280), pad_val=dict(img=(114, 114, 114))),
+ dict(
+ type='CachedMixUp',
+ img_scale=(1280, 1280),
+ ratio_range=(1.0, 1.0),
+ max_cached_images=20,
+ pad_val=(114, 114, 114)),
+ dict(type='PackDetInputs')
+]
+
+train_pipeline_stage2 = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize',
+ scale=(1280, 1280),
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(type='RandomCrop', crop_size=(1280, 1280)),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size=(1280, 1280), pad_val=dict(img=(114, 114, 114))),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(1280, 1280), keep_ratio=True),
+ dict(type='Pad', size=(1280, 1280), pad_val=dict(img=(114, 114, 114))),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=8, num_workers=20, dataset=dict(pipeline=train_pipeline))
+val_dataloader = dict(
+ batch_size=5, num_workers=20, dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+max_epochs = 300
+stage2_num_epochs = 20
+
+base_lr = 0.004 * 32 / 256
+optim_wrapper = dict(optimizer=dict(lr=base_lr))
+
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0e-5,
+ by_epoch=False,
+ begin=0,
+ end=1000),
+ dict(
+ # use cosine lr from 150 to 300 epoch
+ type='CosineAnnealingLR',
+ eta_min=base_lr * 0.05,
+ begin=max_epochs // 2,
+ end=max_epochs,
+ T_max=max_epochs // 2,
+ by_epoch=True,
+ convert_to_iter_based=True),
+]
+
+custom_hooks = [
+ dict(
+ type='EMAHook',
+ ema_type='ExpMomentumEMA',
+ momentum=0.0002,
+ update_buffers=True,
+ priority=49),
+ dict(
+ type='PipelineSwitchHook',
+ switch_epoch=max_epochs - stage2_num_epochs,
+ switch_pipeline=train_pipeline_stage2)
+]
+
+img_scales = [(1280, 1280), (640, 640), (1920, 1920)]
+tta_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=None),
+ dict(
+ type='TestTimeAug',
+ transforms=[
+ [
+ dict(type='Resize', scale=s, keep_ratio=True)
+ for s in img_scales
+ ],
+ [
+ # ``RandomFlip`` must be placed before ``Pad``, otherwise
+ # bounding box coordinates after flipping cannot be
+ # recovered correctly.
+ dict(type='RandomFlip', prob=1.),
+ dict(type='RandomFlip', prob=0.)
+ ],
+ [
+ dict(
+ type='Pad',
+ size=(1920, 1920),
+ pad_val=dict(img=(114, 114, 114))),
+ ],
+ [dict(type='LoadAnnotations', with_bbox=True)],
+ [
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction'))
+ ]
+ ])
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..c730729cfc72a7e3efe885f814ce18c16d2f4a6d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/README.md
@@ -0,0 +1,47 @@
+# SABL
+
+> [Side-Aware Boundary Localization for More Precise Object Detection](https://arxiv.org/abs/1912.04260)
+
+
+
+## Abstract
+
+Current object detection frameworks mainly rely on bounding box regression to localize objects. Despite the remarkable progress in recent years, the precision of bounding box regression remains unsatisfactory, hence limiting performance in object detection. We observe that precise localization requires careful placement of each side of the bounding box. However, the mainstream approach, which focuses on predicting centers and sizes, is not the most effective way to accomplish this task, especially when there exists displacements with large variance between the anchors and the targets. In this paper, we propose an alternative approach, named as Side-Aware Boundary Localization (SABL), where each side of the bounding box is respectively localized with a dedicated network branch. To tackle the difficulty of precise localization in the presence of displacements with large variance, we further propose a two-step localization scheme, which first predicts a range of movement through bucket prediction and then pinpoints the precise position within the predicted bucket. We test the proposed method on both two-stage and single-stage detection frameworks. Replacing the standard bounding box regression branch with the proposed design leads to significant improvements on Faster R-CNN, RetinaNet, and Cascade R-CNN, by 3.0%, 1.7%, and 0.9%, respectively.
+
+
+

+
+
+## Results and Models
+
+The results on COCO 2017 val is shown in the below table. (results on test-dev are usually slightly higher than val).
+Single-scale testing (1333x800) is adopted in all results.
+
+| Method | Backbone | Lr schd | ms-train | box AP | Config | Download |
+| :----------------: | :-------: | :-----: | :------: | :----: | :-----------------------------------------------: | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| SABL Faster R-CNN | R-50-FPN | 1x | N | 39.9 | [config](./sabl-faster-rcnn_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_faster_rcnn_r50_fpn_1x_coco/sabl_faster_rcnn_r50_fpn_1x_coco-e867595b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_faster_rcnn_r50_fpn_1x_coco/20200830_130324.log.json) |
+| SABL Faster R-CNN | R-101-FPN | 1x | N | 41.7 | [config](./sabl-faster-rcnn_r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_faster_rcnn_r101_fpn_1x_coco/sabl_faster_rcnn_r101_fpn_1x_coco-f804c6c1.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_faster_rcnn_r101_fpn_1x_coco/20200830_183949.log.json) |
+| SABL Cascade R-CNN | R-50-FPN | 1x | N | 41.6 | [config](./sabl-cascade-rcnn_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_cascade_rcnn_r50_fpn_1x_coco/sabl_cascade_rcnn_r50_fpn_1x_coco-e1748e5e.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_cascade_rcnn_r50_fpn_1x_coco/20200831_033726.log.json) |
+| SABL Cascade R-CNN | R-101-FPN | 1x | N | 43.0 | [config](./sabl-cascade-rcnn_r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_cascade_rcnn_r101_fpn_1x_coco/sabl_cascade_rcnn_r101_fpn_1x_coco-2b83e87c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_cascade_rcnn_r101_fpn_1x_coco/20200831_141745.log.json) |
+
+| Method | Backbone | GN | Lr schd | ms-train | box AP | Config | Download |
+| :------------: | :-------: | :-: | :-----: | :---------: | :----: | :----------------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| SABL RetinaNet | R-50-FPN | N | 1x | N | 37.7 | [config](./sabl-retinanet_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_retinanet_r50_fpn_1x_coco/sabl_retinanet_r50_fpn_1x_coco-6c54fd4f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_retinanet_r50_fpn_1x_coco/20200830_053451.log.json) |
+| SABL RetinaNet | R-50-FPN | Y | 1x | N | 38.8 | [config](./sabl-retinanet_r50-gn_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_retinanet_r50_fpn_gn_1x_coco/sabl_retinanet_r50_fpn_gn_1x_coco-e16dfcf1.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_retinanet_r50_fpn_gn_1x_coco/20200831_141955.log.json) |
+| SABL RetinaNet | R-101-FPN | N | 1x | N | 39.7 | [config](./sabl-retinanet_r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_retinanet_r101_fpn_1x_coco/sabl_retinanet_r101_fpn_1x_coco-42026904.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_retinanet_r101_fpn_1x_coco/20200831_034256.log.json) |
+| SABL RetinaNet | R-101-FPN | Y | 1x | N | 40.5 | [config](./sabl-retinanet_r101-gn_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_retinanet_r101_fpn_gn_1x_coco/sabl_retinanet_r101_fpn_gn_1x_coco-40a893e8.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_retinanet_r101_fpn_gn_1x_coco/20200830_201422.log.json) |
+| SABL RetinaNet | R-101-FPN | Y | 2x | Y (640~800) | 42.9 | [config](./sabl-retinanet_r101-gn_fpn_ms-640-800-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_retinanet_r101_fpn_gn_2x_ms_640_800_coco/sabl_retinanet_r101_fpn_gn_2x_ms_640_800_coco-1e63382c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_retinanet_r101_fpn_gn_2x_ms_640_800_coco/20200830_144807.log.json) |
+| SABL RetinaNet | R-101-FPN | Y | 2x | Y (480~960) | 43.6 | [config](./sabl-retinanet_r101-gn_fpn_ms-480-960-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_retinanet_r101_fpn_gn_2x_ms_480_960_coco/sabl_retinanet_r101_fpn_gn_2x_ms_480_960_coco-5342f857.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_retinanet_r101_fpn_gn_2x_ms_480_960_coco/20200830_164537.log.json) |
+
+## Citation
+
+We provide config files to reproduce the object detection results in the ECCV 2020 Spotlight paper for [Side-Aware Boundary Localization for More Precise Object Detection](https://arxiv.org/abs/1912.04260).
+
+```latex
+@inproceedings{Wang_2020_ECCV,
+ title = {Side-Aware Boundary Localization for More Precise Object Detection},
+ author = {Jiaqi Wang and Wenwei Zhang and Yuhang Cao and Kai Chen and Jiangmiao Pang and Tao Gong and Jianping Shi and Chen Change Loy and Dahua Lin},
+ booktitle = {ECCV},
+ year = {2020}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..632b869cc4bec559d442410b1d3a4f18d74556ed
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/metafile.yml
@@ -0,0 +1,140 @@
+Collections:
+ - Name: SABL
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - FPN
+ - ResNet
+ - SABL
+ Paper:
+ URL: https://arxiv.org/abs/1912.04260
+ Title: 'Side-Aware Boundary Localization for More Precise Object Detection'
+ README: configs/sabl/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.4.0/mmdet/models/roi_heads/bbox_heads/sabl_head.py#L14
+ Version: v2.4.0
+
+Models:
+ - Name: sabl-faster-rcnn_r50_fpn_1x_coco
+ In Collection: SABL
+ Config: configs/sabl/sabl-faster-rcnn_r50_fpn_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_faster_rcnn_r50_fpn_1x_coco/sabl_faster_rcnn_r50_fpn_1x_coco-e867595b.pth
+
+ - Name: sabl-faster-rcnn_r101_fpn_1x_coco
+ In Collection: SABL
+ Config: configs/sabl/sabl-faster-rcnn_r101_fpn_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_faster_rcnn_r101_fpn_1x_coco/sabl_faster_rcnn_r101_fpn_1x_coco-f804c6c1.pth
+
+ - Name: sabl-cascade-rcnn_r50_fpn_1x_coco
+ In Collection: SABL
+ Config: configs/sabl/sabl-cascade-rcnn_r50_fpn_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_cascade_rcnn_r50_fpn_1x_coco/sabl_cascade_rcnn_r50_fpn_1x_coco-e1748e5e.pth
+
+ - Name: sabl-cascade-rcnn_r101_fpn_1x_coco
+ In Collection: SABL
+ Config: configs/sabl/sabl-cascade-rcnn_r101_fpn_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_cascade_rcnn_r101_fpn_1x_coco/sabl_cascade_rcnn_r101_fpn_1x_coco-2b83e87c.pth
+
+ - Name: sabl-retinanet_r50_fpn_1x_coco
+ In Collection: SABL
+ Config: configs/sabl/sabl-retinanet_r50_fpn_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_retinanet_r50_fpn_1x_coco/sabl_retinanet_r50_fpn_1x_coco-6c54fd4f.pth
+
+ - Name: sabl-retinanet_r50-gn_fpn_1x_coco
+ In Collection: SABL
+ Config: configs/sabl/sabl-retinanet_r50-gn_fpn_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 38.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_retinanet_r50_fpn_gn_1x_coco/sabl_retinanet_r50_fpn_gn_1x_coco-e16dfcf1.pth
+
+ - Name: sabl-retinanet_r101_fpn_1x_coco
+ In Collection: SABL
+ Config: configs/sabl/sabl-retinanet_r101_fpn_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 39.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_retinanet_r101_fpn_1x_coco/sabl_retinanet_r101_fpn_1x_coco-42026904.pth
+
+ - Name: sabl-retinanet_r101-gn_fpn_1x_coco
+ In Collection: SABL
+ Config: configs/sabl/sabl-retinanet_r101-gn_fpn_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_retinanet_r101_fpn_gn_1x_coco/sabl_retinanet_r101_fpn_gn_1x_coco-40a893e8.pth
+
+ - Name: sabl-retinanet_r101-gn_fpn_ms-640-800-2x_coco
+ In Collection: SABL
+ Config: configs/sabl/sabl-retinanet_r101-gn_fpn_ms-640-800-2x_coco.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_retinanet_r101_fpn_gn_2x_ms_640_800_coco/sabl_retinanet_r101_fpn_gn_2x_ms_640_800_coco-1e63382c.pth
+
+ - Name: sabl-retinanet_r101-gn_fpn_ms-480-960-2x_coco
+ In Collection: SABL
+ Config: configs/sabl/sabl-retinanet_r101-gn_fpn_ms-480-960-2x_coco.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/sabl/sabl_retinanet_r101_fpn_gn_2x_ms_480_960_coco/sabl_retinanet_r101_fpn_gn_2x_ms_480_960_coco-5342f857.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-cascade-rcnn_r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-cascade-rcnn_r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..404e7fcb2ac52773c9bc74f411e66584114f378e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-cascade-rcnn_r101_fpn_1x_coco.py
@@ -0,0 +1,90 @@
+_base_ = [
+ '../_base_/models/cascade-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+# model settings
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')),
+ roi_head=dict(bbox_head=[
+ dict(
+ type='SABLHead',
+ num_classes=80,
+ cls_in_channels=256,
+ reg_in_channels=256,
+ roi_feat_size=7,
+ reg_feat_up_ratio=2,
+ reg_pre_kernel=3,
+ reg_post_kernel=3,
+ reg_pre_num=2,
+ reg_post_num=1,
+ cls_out_channels=1024,
+ reg_offset_out_channels=256,
+ reg_cls_out_channels=256,
+ num_cls_fcs=1,
+ num_reg_fcs=0,
+ reg_class_agnostic=True,
+ norm_cfg=None,
+ bbox_coder=dict(
+ type='BucketingBBoxCoder', num_buckets=14, scale_factor=1.7),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.0),
+ loss_bbox_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox_reg=dict(type='SmoothL1Loss', beta=0.1,
+ loss_weight=1.0)),
+ dict(
+ type='SABLHead',
+ num_classes=80,
+ cls_in_channels=256,
+ reg_in_channels=256,
+ roi_feat_size=7,
+ reg_feat_up_ratio=2,
+ reg_pre_kernel=3,
+ reg_post_kernel=3,
+ reg_pre_num=2,
+ reg_post_num=1,
+ cls_out_channels=1024,
+ reg_offset_out_channels=256,
+ reg_cls_out_channels=256,
+ num_cls_fcs=1,
+ num_reg_fcs=0,
+ reg_class_agnostic=True,
+ norm_cfg=None,
+ bbox_coder=dict(
+ type='BucketingBBoxCoder', num_buckets=14, scale_factor=1.5),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.0),
+ loss_bbox_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox_reg=dict(type='SmoothL1Loss', beta=0.1,
+ loss_weight=1.0)),
+ dict(
+ type='SABLHead',
+ num_classes=80,
+ cls_in_channels=256,
+ reg_in_channels=256,
+ roi_feat_size=7,
+ reg_feat_up_ratio=2,
+ reg_pre_kernel=3,
+ reg_post_kernel=3,
+ reg_pre_num=2,
+ reg_post_num=1,
+ cls_out_channels=1024,
+ reg_offset_out_channels=256,
+ reg_cls_out_channels=256,
+ num_cls_fcs=1,
+ num_reg_fcs=0,
+ reg_class_agnostic=True,
+ norm_cfg=None,
+ bbox_coder=dict(
+ type='BucketingBBoxCoder', num_buckets=14, scale_factor=1.3),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.0),
+ loss_bbox_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox_reg=dict(type='SmoothL1Loss', beta=0.1, loss_weight=1.0))
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-cascade-rcnn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-cascade-rcnn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..69c59ca20d6c16e458292a55b8e4258a3d9a06bb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-cascade-rcnn_r50_fpn_1x_coco.py
@@ -0,0 +1,86 @@
+_base_ = [
+ '../_base_/models/cascade-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+# model settings
+model = dict(
+ roi_head=dict(bbox_head=[
+ dict(
+ type='SABLHead',
+ num_classes=80,
+ cls_in_channels=256,
+ reg_in_channels=256,
+ roi_feat_size=7,
+ reg_feat_up_ratio=2,
+ reg_pre_kernel=3,
+ reg_post_kernel=3,
+ reg_pre_num=2,
+ reg_post_num=1,
+ cls_out_channels=1024,
+ reg_offset_out_channels=256,
+ reg_cls_out_channels=256,
+ num_cls_fcs=1,
+ num_reg_fcs=0,
+ reg_class_agnostic=True,
+ norm_cfg=None,
+ bbox_coder=dict(
+ type='BucketingBBoxCoder', num_buckets=14, scale_factor=1.7),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.0),
+ loss_bbox_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox_reg=dict(type='SmoothL1Loss', beta=0.1,
+ loss_weight=1.0)),
+ dict(
+ type='SABLHead',
+ num_classes=80,
+ cls_in_channels=256,
+ reg_in_channels=256,
+ roi_feat_size=7,
+ reg_feat_up_ratio=2,
+ reg_pre_kernel=3,
+ reg_post_kernel=3,
+ reg_pre_num=2,
+ reg_post_num=1,
+ cls_out_channels=1024,
+ reg_offset_out_channels=256,
+ reg_cls_out_channels=256,
+ num_cls_fcs=1,
+ num_reg_fcs=0,
+ reg_class_agnostic=True,
+ norm_cfg=None,
+ bbox_coder=dict(
+ type='BucketingBBoxCoder', num_buckets=14, scale_factor=1.5),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.0),
+ loss_bbox_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox_reg=dict(type='SmoothL1Loss', beta=0.1,
+ loss_weight=1.0)),
+ dict(
+ type='SABLHead',
+ num_classes=80,
+ cls_in_channels=256,
+ reg_in_channels=256,
+ roi_feat_size=7,
+ reg_feat_up_ratio=2,
+ reg_pre_kernel=3,
+ reg_post_kernel=3,
+ reg_pre_num=2,
+ reg_post_num=1,
+ cls_out_channels=1024,
+ reg_offset_out_channels=256,
+ reg_cls_out_channels=256,
+ num_cls_fcs=1,
+ num_reg_fcs=0,
+ reg_class_agnostic=True,
+ norm_cfg=None,
+ bbox_coder=dict(
+ type='BucketingBBoxCoder', num_buckets=14, scale_factor=1.3),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.0),
+ loss_bbox_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox_reg=dict(type='SmoothL1Loss', beta=0.1, loss_weight=1.0))
+ ]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-faster-rcnn_r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-faster-rcnn_r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d1bf8b9c8cf1ac62d351456e7b19f75259ec0625
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-faster-rcnn_r101_fpn_1x_coco.py
@@ -0,0 +1,38 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')),
+ roi_head=dict(
+ bbox_head=dict(
+ _delete_=True,
+ type='SABLHead',
+ num_classes=80,
+ cls_in_channels=256,
+ reg_in_channels=256,
+ roi_feat_size=7,
+ reg_feat_up_ratio=2,
+ reg_pre_kernel=3,
+ reg_post_kernel=3,
+ reg_pre_num=2,
+ reg_post_num=1,
+ cls_out_channels=1024,
+ reg_offset_out_channels=256,
+ reg_cls_out_channels=256,
+ num_cls_fcs=1,
+ num_reg_fcs=0,
+ reg_class_agnostic=True,
+ norm_cfg=None,
+ bbox_coder=dict(
+ type='BucketingBBoxCoder', num_buckets=14, scale_factor=1.7),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.0),
+ loss_bbox_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox_reg=dict(type='SmoothL1Loss', beta=0.1,
+ loss_weight=1.0))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-faster-rcnn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-faster-rcnn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a727bd6d3da09c86908c3c584509c5313cf732b5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-faster-rcnn_r50_fpn_1x_coco.py
@@ -0,0 +1,34 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ roi_head=dict(
+ bbox_head=dict(
+ _delete_=True,
+ type='SABLHead',
+ num_classes=80,
+ cls_in_channels=256,
+ reg_in_channels=256,
+ roi_feat_size=7,
+ reg_feat_up_ratio=2,
+ reg_pre_kernel=3,
+ reg_post_kernel=3,
+ reg_pre_num=2,
+ reg_post_num=1,
+ cls_out_channels=1024,
+ reg_offset_out_channels=256,
+ reg_cls_out_channels=256,
+ num_cls_fcs=1,
+ num_reg_fcs=0,
+ reg_class_agnostic=True,
+ norm_cfg=None,
+ bbox_coder=dict(
+ type='BucketingBBoxCoder', num_buckets=14, scale_factor=1.7),
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.0),
+ loss_bbox_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox_reg=dict(type='SmoothL1Loss', beta=0.1,
+ loss_weight=1.0))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-retinanet_r101-gn_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-retinanet_r101-gn_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f181ad6813e4c6e3729ff80b3b8d915d84b53bf2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-retinanet_r101-gn_fpn_1x_coco.py
@@ -0,0 +1,57 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+# model settings
+norm_cfg = dict(type='GN', num_groups=32, requires_grad=True)
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')),
+ bbox_head=dict(
+ _delete_=True,
+ type='SABLRetinaHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ approx_anchor_generator=dict(
+ type='AnchorGenerator',
+ octave_base_scale=4,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[8, 16, 32, 64, 128]),
+ square_anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ scales=[4],
+ strides=[8, 16, 32, 64, 128]),
+ norm_cfg=norm_cfg,
+ bbox_coder=dict(
+ type='BucketingBBoxCoder', num_buckets=14, scale_factor=3.0),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.5),
+ loss_bbox_reg=dict(
+ type='SmoothL1Loss', beta=1.0 / 9.0, loss_weight=1.5)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='ApproxMaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.4,
+ min_pos_iou=0.0,
+ ignore_iof_thr=-1),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False))
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-retinanet_r101-gn_fpn_ms-480-960-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-retinanet_r101-gn_fpn_ms-480-960-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..dc7209aebad3efcb88945460cf20b36e6ec4b419
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-retinanet_r101-gn_fpn_ms-480-960-2x_coco.py
@@ -0,0 +1,68 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_2x.py', '../_base_/default_runtime.py'
+]
+# model settings
+norm_cfg = dict(type='GN', num_groups=32, requires_grad=True)
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')),
+ bbox_head=dict(
+ _delete_=True,
+ type='SABLRetinaHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ approx_anchor_generator=dict(
+ type='AnchorGenerator',
+ octave_base_scale=4,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[8, 16, 32, 64, 128]),
+ square_anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ scales=[4],
+ strides=[8, 16, 32, 64, 128]),
+ norm_cfg=norm_cfg,
+ bbox_coder=dict(
+ type='BucketingBBoxCoder', num_buckets=14, scale_factor=3.0),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.5),
+ loss_bbox_reg=dict(
+ type='SmoothL1Loss', beta=1.0 / 9.0, loss_weight=1.5)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='ApproxMaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.4,
+ min_pos_iou=0.0,
+ ignore_iof_thr=-1),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False))
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize', scale=[(1333, 480), (1333, 960)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-retinanet_r101-gn_fpn_ms-640-800-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-retinanet_r101-gn_fpn_ms-640-800-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ac5f6d9811dc8e45cfc036b3a3d4a04e7fa5ee60
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-retinanet_r101-gn_fpn_ms-640-800-2x_coco.py
@@ -0,0 +1,68 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_2x.py', '../_base_/default_runtime.py'
+]
+# model settings
+norm_cfg = dict(type='GN', num_groups=32, requires_grad=True)
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')),
+ bbox_head=dict(
+ _delete_=True,
+ type='SABLRetinaHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ approx_anchor_generator=dict(
+ type='AnchorGenerator',
+ octave_base_scale=4,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[8, 16, 32, 64, 128]),
+ square_anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ scales=[4],
+ strides=[8, 16, 32, 64, 128]),
+ norm_cfg=norm_cfg,
+ bbox_coder=dict(
+ type='BucketingBBoxCoder', num_buckets=14, scale_factor=3.0),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.5),
+ loss_bbox_reg=dict(
+ type='SmoothL1Loss', beta=1.0 / 9.0, loss_weight=1.5)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='ApproxMaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.4,
+ min_pos_iou=0.0,
+ ignore_iof_thr=-1),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False))
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize', scale=[(1333, 480), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-retinanet_r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-retinanet_r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..409695b5dbccfe20bb6e85ee16231211c2ebcdba
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-retinanet_r101_fpn_1x_coco.py
@@ -0,0 +1,55 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+# model settings
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')),
+ bbox_head=dict(
+ _delete_=True,
+ type='SABLRetinaHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ approx_anchor_generator=dict(
+ type='AnchorGenerator',
+ octave_base_scale=4,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[8, 16, 32, 64, 128]),
+ square_anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ scales=[4],
+ strides=[8, 16, 32, 64, 128]),
+ bbox_coder=dict(
+ type='BucketingBBoxCoder', num_buckets=14, scale_factor=3.0),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.5),
+ loss_bbox_reg=dict(
+ type='SmoothL1Loss', beta=1.0 / 9.0, loss_weight=1.5)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='ApproxMaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.4,
+ min_pos_iou=0.0,
+ ignore_iof_thr=-1),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False))
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-retinanet_r50-gn_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-retinanet_r50-gn_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..4facdb6aaab05fd04b95e8c3ba2f0460090b1d6c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-retinanet_r50-gn_fpn_1x_coco.py
@@ -0,0 +1,53 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+# model settings
+norm_cfg = dict(type='GN', num_groups=32, requires_grad=True)
+model = dict(
+ bbox_head=dict(
+ _delete_=True,
+ type='SABLRetinaHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ approx_anchor_generator=dict(
+ type='AnchorGenerator',
+ octave_base_scale=4,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[8, 16, 32, 64, 128]),
+ square_anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ scales=[4],
+ strides=[8, 16, 32, 64, 128]),
+ norm_cfg=norm_cfg,
+ bbox_coder=dict(
+ type='BucketingBBoxCoder', num_buckets=14, scale_factor=3.0),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.5),
+ loss_bbox_reg=dict(
+ type='SmoothL1Loss', beta=1.0 / 9.0, loss_weight=1.5)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='ApproxMaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.4,
+ min_pos_iou=0.0,
+ ignore_iof_thr=-1),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False))
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-retinanet_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-retinanet_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..9073d6f002fcb49aecc280f318b8769b477d2d82
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sabl/sabl-retinanet_r50_fpn_1x_coco.py
@@ -0,0 +1,51 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+# model settings
+model = dict(
+ bbox_head=dict(
+ _delete_=True,
+ type='SABLRetinaHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ approx_anchor_generator=dict(
+ type='AnchorGenerator',
+ octave_base_scale=4,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[8, 16, 32, 64, 128]),
+ square_anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ scales=[4],
+ strides=[8, 16, 32, 64, 128]),
+ bbox_coder=dict(
+ type='BucketingBBoxCoder', num_buckets=14, scale_factor=3.0),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.5),
+ loss_bbox_reg=dict(
+ type='SmoothL1Loss', beta=1.0 / 9.0, loss_weight=1.5)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='ApproxMaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.4,
+ min_pos_iou=0.0,
+ ignore_iof_thr=-1),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False))
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..08dbfa87f5625ba6500c731910c178a5e2684e0f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/README.md
@@ -0,0 +1,63 @@
+# SCNet
+
+> [SCNet: Training Inference Sample Consistency for Instance Segmentation](https://arxiv.org/abs/2012.10150)
+
+
+
+## Abstract
+
+
+
+Cascaded architectures have brought significant performance improvement in object detection and instance segmentation. However, there are lingering issues regarding the disparity in the Intersection-over-Union (IoU) distribution of the samples between training and inference. This disparity can potentially exacerbate detection accuracy. This paper proposes an architecture referred to as Sample Consistency Network (SCNet) to ensure that the IoU distribution of the samples at training time is close to that at inference time. Furthermore, SCNet incorporates feature relay and utilizes global contextual information to further reinforce the reciprocal relationships among classifying, detecting, and segmenting sub-tasks. Extensive experiments on the standard COCO dataset reveal the effectiveness of the proposed method over multiple evaluation metrics, including box AP, mask AP, and inference speed. In particular, while running 38% faster, the proposed SCNet improves the AP of the box and mask predictions by respectively 1.3 and 2.3 points compared to the strong Cascade Mask R-CNN baseline.
+
+
+

+
+
+## Dataset
+
+SCNet requires COCO and [COCO-stuff](http://calvin.inf.ed.ac.uk/wp-content/uploads/data/cocostuffdataset/stuffthingmaps_trainval2017.zip) dataset for training. You need to download and extract it in the COCO dataset path.
+The directory should be like this.
+
+```none
+mmdetection
+├── mmdet
+├── tools
+├── configs
+├── data
+│ ├── coco
+│ │ ├── annotations
+│ │ ├── train2017
+│ │ ├── val2017
+│ │ ├── test2017
+| | ├── stuffthingmaps
+```
+
+## Results and Models
+
+The results on COCO 2017val are shown in the below table. (results on test-dev are usually slightly higher than val)
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf speed (fps) | box AP | mask AP | TTA box AP | TTA mask AP | Config | Download |
+| :-------------: | :-----: | :-----: | :------: | :-------------: | :----: | :-----: | :--------: | :---------: | :------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-FPN | pytorch | 1x | 7.0 | 6.2 | 43.5 | 39.2 | 44.8 | 40.9 | [config](./scnet_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/scnet/scnet_r50_fpn_1x_coco/scnet_r50_fpn_1x_coco-c3f09857.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/scnet/scnet_r50_fpn_1x_coco/scnet_r50_fpn_1x_coco_20210117_192725.log.json) |
+| R-50-FPN | pytorch | 20e | 7.0 | 6.2 | 44.5 | 40.0 | 45.8 | 41.5 | [config](./scnet_r50_fpn_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/scnet/scnet_r50_fpn_20e_coco/scnet_r50_fpn_20e_coco-a569f645.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/scnet/scnet_r50_fpn_20e_coco/scnet_r50_fpn_20e_coco_20210116_060148.log.json) |
+| R-101-FPN | pytorch | 20e | 8.9 | 5.8 | 45.8 | 40.9 | 47.3 | 42.7 | [config](./scnet_r101_fpn_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/scnet/scnet_r101_fpn_20e_coco/scnet_r101_fpn_20e_coco-294e312c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/scnet/scnet_r101_fpn_20e_coco/scnet_r101_fpn_20e_coco_20210118_175824.log.json) |
+| X-101-64x4d-FPN | pytorch | 20e | 13.2 | 4.9 | 47.5 | 42.3 | 48.9 | 44.0 | [config](./scnet_x101-64x4d_fpn_20e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/scnet/scnet_x101_64x4d_fpn_20e_coco/scnet_x101_64x4d_fpn_20e_coco-fb09dec9.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/scnet/scnet_x101_64x4d_fpn_20e_coco/scnet_x101_64x4d_fpn_20e_coco_20210120_045959.log.json) |
+
+### Notes
+
+- Training hyper-parameters are identical to those of [HTC](https://github.com/open-mmlab/mmdetection/tree/main/configs/htc).
+- TTA means Test Time Augmentation, which applies horizontal flip and multi-scale testing. Refer to [config](./scnet_r50_fpn_1x_coco.py).
+
+## Citation
+
+We provide the code for reproducing experiment results of [SCNet](https://arxiv.org/abs/2012.10150).
+
+```latex
+@inproceedings{vu2019cascade,
+ title={SCNet: Training Inference Sample Consistency for Instance Segmentation},
+ author={Vu, Thang and Haeyong, Kang and Yoo, Chang D},
+ booktitle={AAAI},
+ year={2021}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..936d38960a8f423198702194f64a9eb46c770979
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/metafile.yml
@@ -0,0 +1,116 @@
+Collections:
+ - Name: SCNet
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - FPN
+ - ResNet
+ - SCNet
+ Paper:
+ URL: https://arxiv.org/abs/2012.10150
+ Title: 'SCNet: Training Inference Sample Consistency for Instance Segmentation'
+ README: configs/scnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.9.0/mmdet/models/detectors/scnet.py#L6
+ Version: v2.9.0
+
+Models:
+ - Name: scnet_r50_fpn_1x_coco
+ In Collection: SCNet
+ Config: configs/scnet/scnet_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.0
+ inference time (ms/im):
+ - value: 161.29
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/scnet/scnet_r50_fpn_1x_coco/scnet_r50_fpn_1x_coco-c3f09857.pth
+
+ - Name: scnet_r50_fpn_20e_coco
+ In Collection: SCNet
+ Config: configs/scnet/scnet_r50_fpn_20e_coco.py
+ Metadata:
+ Training Memory (GB): 7.0
+ inference time (ms/im):
+ - value: 161.29
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 20
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 40.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/scnet/scnet_r50_fpn_20e_coco/scnet_r50_fpn_20e_coco-a569f645.pth
+
+ - Name: scnet_r101_fpn_20e_coco
+ In Collection: SCNet
+ Config: configs/scnet/scnet_r101_fpn_20e_coco.py
+ Metadata:
+ Training Memory (GB): 8.9
+ inference time (ms/im):
+ - value: 172.41
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 20
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 40.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/scnet/scnet_r101_fpn_20e_coco/scnet_r101_fpn_20e_coco-294e312c.pth
+
+ - Name: scnet_x101-64x4d_fpn_20e_coco
+ In Collection: SCNet
+ Config: configs/scnet/scnet_x101-64x4d_fpn_20e_coco.py
+ Metadata:
+ Training Memory (GB): 13.2
+ inference time (ms/im):
+ - value: 204.08
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (800, 1333)
+ Epochs: 20
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 47.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 42.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/scnet/scnet_x101_64x4d_fpn_20e_coco/scnet_x101_64x4d_fpn_20e_coco-fb09dec9.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/scnet_r101_fpn_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/scnet_r101_fpn_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ebba52978b23c07a68e3563033c860a95dd515b6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/scnet_r101_fpn_20e_coco.py
@@ -0,0 +1,6 @@
+_base_ = './scnet_r50_fpn_20e_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/scnet_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/scnet_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a0210fdb456c26b2c05d99a2435da14fc30f088d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/scnet_r50_fpn_1x_coco.py
@@ -0,0 +1,138 @@
+_base_ = '../htc/htc_r50_fpn_1x_coco.py'
+# model settings
+model = dict(
+ type='SCNet',
+ roi_head=dict(
+ _delete_=True,
+ type='SCNetRoIHead',
+ num_stages=3,
+ stage_loss_weights=[1, 0.5, 0.25],
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=7, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ bbox_head=[
+ dict(
+ type='SCNetBBoxHead',
+ num_shared_fcs=2,
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0,
+ loss_weight=1.0)),
+ dict(
+ type='SCNetBBoxHead',
+ num_shared_fcs=2,
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.05, 0.05, 0.1, 0.1]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0,
+ loss_weight=1.0)),
+ dict(
+ type='SCNetBBoxHead',
+ num_shared_fcs=2,
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.033, 0.033, 0.067, 0.067]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0))
+ ],
+ mask_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=14, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ mask_head=dict(
+ type='SCNetMaskHead',
+ num_convs=12,
+ in_channels=256,
+ conv_out_channels=256,
+ num_classes=80,
+ conv_to_res=True,
+ loss_mask=dict(
+ type='CrossEntropyLoss', use_mask=True, loss_weight=1.0)),
+ semantic_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=14, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[8]),
+ semantic_head=dict(
+ type='SCNetSemanticHead',
+ num_ins=5,
+ fusion_level=1,
+ seg_scale_factor=1 / 8,
+ num_convs=4,
+ in_channels=256,
+ conv_out_channels=256,
+ num_classes=183,
+ loss_seg=dict(
+ type='CrossEntropyLoss', ignore_index=255, loss_weight=0.2),
+ conv_to_res=True),
+ glbctx_head=dict(
+ type='GlobalContextHead',
+ num_convs=4,
+ in_channels=256,
+ conv_out_channels=256,
+ num_classes=80,
+ loss_weight=3.0,
+ conv_to_res=True),
+ feat_relay_head=dict(
+ type='FeatureRelayHead',
+ in_channels=1024,
+ out_conv_channels=256,
+ roi_feat_size=7,
+ scale_factor=2)))
+
+# TODO
+# uncomment below code to enable test time augmentations
+# img_norm_cfg = dict(
+# mean=[123.675, 116.28, 103.53], std=[58.395, 57.12, 57.375], to_rgb=True)
+# test_pipeline = [
+# dict(type='LoadImageFromFile'),
+# dict(
+# type='MultiScaleFlipAug',
+# img_scale=[(600, 900), (800, 1200), (1000, 1500), (1200, 1800),
+# (1400, 2100)],
+# flip=True,
+# transforms=[
+# dict(type='Resize', keep_ratio=True),
+# dict(type='RandomFlip', flip_ratio=0.5),
+# dict(type='Normalize', **img_norm_cfg),
+# dict(type='Pad', size_divisor=32),
+# dict(type='ImageToTensor', keys=['img']),
+# dict(type='Collect', keys=['img']),
+# ])
+# ]
+# data = dict(
+# val=dict(pipeline=test_pipeline),
+# test=dict(pipeline=test_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/scnet_r50_fpn_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/scnet_r50_fpn_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..533e1b5f3253387788fbf1a9d6d7a38c7c5c5f30
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/scnet_r50_fpn_20e_coco.py
@@ -0,0 +1,15 @@
+_base_ = './scnet_r50_fpn_1x_coco.py'
+# learning policy
+max_epochs = 20
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 19],
+ gamma=0.1)
+]
+train_cfg = dict(max_epochs=max_epochs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/scnet_x101-64x4d_fpn_20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/scnet_x101-64x4d_fpn_20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1e54b030fa68f76f22edf66e3594d66a13c2c672
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/scnet_x101-64x4d_fpn_20e_coco.py
@@ -0,0 +1,15 @@
+_base_ = './scnet_r50_fpn_20e_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/scnet_x101-64x4d_fpn_8xb1-20e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/scnet_x101-64x4d_fpn_8xb1-20e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3cdce7d54248e77e98639d68490cc30dfd625c87
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scnet/scnet_x101-64x4d_fpn_8xb1-20e_coco.py
@@ -0,0 +1,8 @@
+_base_ = './scnet_x101-64x4d_fpn_20e_coco.py'
+train_dataloader = dict(batch_size=1, num_workers=1)
+
+optim_wrapper = dict(optimizer=dict(lr=0.01))
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (1 samples per GPU)
+auto_scale_lr = dict(base_batch_size=8)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/scratch/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scratch/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..7bdd8ff9f20a0b222a37eebfb44311150c130b15
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scratch/README.md
@@ -0,0 +1,35 @@
+# Scratch
+
+> [Rethinking ImageNet Pre-training](https://arxiv.org/abs/1811.08883)
+
+
+
+## Abstract
+
+We report competitive results on object detection and instance segmentation on the COCO dataset using standard models trained from random initialization. The results are no worse than their ImageNet pre-training counterparts even when using the hyper-parameters of the baseline system (Mask R-CNN) that were optimized for fine-tuning pre-trained models, with the sole exception of increasing the number of training iterations so the randomly initialized models may converge. Training from random initialization is surprisingly robust; our results hold even when: (i) using only 10% of the training data, (ii) for deeper and wider models, and (iii) for multiple tasks and metrics. Experiments show that ImageNet pre-training speeds up convergence early in training, but does not necessarily provide regularization or improve final target task accuracy. To push the envelope we demonstrate 50.9 AP on COCO object detection without using any external data---a result on par with the top COCO 2017 competition results that used ImageNet pre-training. These observations challenge the conventional wisdom of ImageNet pre-training for dependent tasks and we expect these discoveries will encourage people to rethink the current de facto paradigm of \`pre-training and fine-tuning' in computer vision.
+
+
+

+
+
+## Results and Models
+
+| Model | Backbone | Style | Lr schd | box AP | mask AP | Config | Download |
+| :----------: | :------: | :-----: | :-----: | :----: | :-----: | :-------------------------------------------------------: | :-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Faster R-CNN | R-50-FPN | pytorch | 6x | 40.7 | | [config](./faster-rcnn_r50-scratch_fpn_gn-all_6x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/scratch/faster_rcnn_r50_fpn_gn-all_scratch_6x_coco/scratch_faster_rcnn_r50_fpn_gn_6x_bbox_mAP-0.407_20200201_193013-90813d01.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/scratch/faster_rcnn_r50_fpn_gn-all_scratch_6x_coco/scratch_faster_rcnn_r50_fpn_gn_6x_20200201_193013.log.json) |
+| Mask R-CNN | R-50-FPN | pytorch | 6x | 41.2 | 37.4 | [config](./mask-rcnn_r50-scratch_fpn_gn-all_6x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/scratch/mask_rcnn_r50_fpn_gn-all_scratch_6x_coco/scratch_mask_rcnn_r50_fpn_gn_6x_bbox_mAP-0.412__segm_mAP-0.374_20200201_193051-1e190a40.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/scratch/mask_rcnn_r50_fpn_gn-all_scratch_6x_coco/scratch_mask_rcnn_r50_fpn_gn_6x_20200201_193051.log.json) |
+
+Note:
+
+- The above models are trained with 16 GPUs.
+
+## Citation
+
+```latex
+@article{he2018rethinking,
+ title={Rethinking imagenet pre-training},
+ author={He, Kaiming and Girshick, Ross and Doll{\'a}r, Piotr},
+ journal={arXiv preprint arXiv:1811.08883},
+ year={2018}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/scratch/faster-rcnn_r50-scratch_fpn_gn-all_6x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scratch/faster-rcnn_r50-scratch_fpn_gn-all_6x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6e632b9a150871a44b698dfdb0fdc3f07308ef81
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scratch/faster-rcnn_r50-scratch_fpn_gn-all_6x_coco.py
@@ -0,0 +1,39 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+norm_cfg = dict(type='GN', num_groups=32, requires_grad=True)
+model = dict(
+ backbone=dict(
+ frozen_stages=-1,
+ zero_init_residual=False,
+ norm_cfg=norm_cfg,
+ init_cfg=None),
+ neck=dict(norm_cfg=norm_cfg),
+ roi_head=dict(
+ bbox_head=dict(
+ type='Shared4Conv1FCBBoxHead',
+ conv_out_channels=256,
+ norm_cfg=norm_cfg)))
+
+optim_wrapper = dict(paramwise_cfg=dict(norm_decay_mult=0.))
+
+max_epochs = 73
+
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[65, 71],
+ gamma=0.1)
+]
+
+train_cfg = dict(max_epochs=max_epochs)
+
+# only keep latest 3 checkpoints
+default_hooks = dict(checkpoint=dict(max_keep_ckpts=3))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/scratch/mask-rcnn_r50-scratch_fpn_gn-all_6x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scratch/mask-rcnn_r50-scratch_fpn_gn-all_6x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..9796f504b677a841919bb058ded414de25e74a50
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scratch/mask-rcnn_r50-scratch_fpn_gn-all_6x_coco.py
@@ -0,0 +1,40 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+norm_cfg = dict(type='GN', num_groups=32, requires_grad=True)
+model = dict(
+ backbone=dict(
+ frozen_stages=-1,
+ zero_init_residual=False,
+ norm_cfg=norm_cfg,
+ init_cfg=None),
+ neck=dict(norm_cfg=norm_cfg),
+ roi_head=dict(
+ bbox_head=dict(
+ type='Shared4Conv1FCBBoxHead',
+ conv_out_channels=256,
+ norm_cfg=norm_cfg),
+ mask_head=dict(norm_cfg=norm_cfg)))
+
+optim_wrapper = dict(paramwise_cfg=dict(norm_decay_mult=0.))
+
+max_epochs = 73
+
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[65, 71],
+ gamma=0.1)
+]
+
+train_cfg = dict(max_epochs=max_epochs)
+
+# only keep latest 3 checkpoints
+default_hooks = dict(checkpoint=dict(max_keep_ckpts=3))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/scratch/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scratch/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..977b8e5bfc2b6319793ae8abdeb71e5e04d7cb1b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/scratch/metafile.yml
@@ -0,0 +1,48 @@
+Collections:
+ - Name: Rethinking ImageNet Pre-training
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - FPN
+ - RPN
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/1811.08883
+ Title: 'Rethinking ImageNet Pre-training'
+ README: configs/scratch/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.0.0/configs/scratch/faster-rcnn_r50-scratch_fpn_gn-all_6x_coco.py
+ Version: v2.0.0
+
+Models:
+ - Name: faster-rcnn_r50_fpn_gn-all_scratch_6x_coco
+ In Collection: Rethinking ImageNet Pre-training
+ Config: configs/scratch/faster-rcnn_r50-scratch_fpn_gn-all_6x_coco.py
+ Metadata:
+ Epochs: 72
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/scratch/faster_rcnn_r50_fpn_gn-all_scratch_6x_coco/scratch_faster_rcnn_r50_fpn_gn_6x_bbox_mAP-0.407_20200201_193013-90813d01.pth
+
+ - Name: mask-rcnn_r50_fpn_gn-all_scratch_6x_coco
+ In Collection: Rethinking ImageNet Pre-training
+ Config: configs/scratch/mask-rcnn_r50-scratch_fpn_gn-all_6x_coco.py
+ Metadata:
+ Epochs: 72
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/scratch/mask_rcnn_r50_fpn_gn-all_scratch_6x_coco/scratch_mask_rcnn_r50_fpn_gn_6x_bbox_mAP-0.412__segm_mAP-0.374_20200201_193051-1e190a40.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..7077d75351caf0ca21760939eb0e2cea2fee5f85
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/README.md
@@ -0,0 +1,47 @@
+# Seesaw Loss
+
+> [Seesaw Loss for Long-Tailed Instance Segmentation](https://arxiv.org/abs/2008.10032)
+
+
+
+## Abstract
+
+Instance segmentation has witnessed a remarkable progress on class-balanced benchmarks. However, they fail to perform as accurately in real-world scenarios, where the category distribution of objects naturally comes with a long tail. Instances of head classes dominate a long-tailed dataset and they serve as negative samples of tail categories. The overwhelming gradients of negative samples on tail classes lead to a biased learning process for classifiers. Consequently, objects of tail categories are more likely to be misclassified as backgrounds or head categories. To tackle this problem, we propose Seesaw Loss to dynamically re-balance gradients of positive and negative samples for each category, with two complementary factors, i.e., mitigation factor and compensation factor. The mitigation factor reduces punishments to tail categories w.r.t. the ratio of cumulative training instances between different categories. Meanwhile, the compensation factor increases the penalty of misclassified instances to avoid false positives of tail categories. We conduct extensive experiments on Seesaw Loss with mainstream frameworks and different data sampling strategies. With a simple end-to-end training pipeline, Seesaw Loss obtains significant gains over Cross-Entropy Loss, and achieves state-of-the-art performance on LVIS dataset without bells and whistles.
+
+
+

+
+
+- Please setup [LVIS dataset](../lvis/README.md) for MMDetection.
+
+- RFS indicates to use oversample strategy [here](../../docs/tutorials/customipredataset.md#class-balanced-dataset) with oversample threshold `1e-3`.
+
+## Results and models of Seasaw Loss on LVIS v1 dataset
+
+| Method | Backbone | Style | Lr schd | Data Sampler | Norm Mask | box AP | mask AP | Config | Download |
+| :----------------: | :-------: | :-----: | :-----: | :----------: | :-------: | :----: | :-----: | :----------------------------------------------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Mask R-CNN | R-50-FPN | pytorch | 2x | random | N | 25.6 | 25.0 | [config](./mask-rcnn_r50_fpn_seesaw-loss_random-ms-2x_lvis-v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r50_fpn_random_seesaw_loss_mstrain_2x_lvis_v1-a698dd3d.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r50_fpn_random_seesaw_loss_mstrain_2x_lvis_v1.log.json) |
+| Mask R-CNN | R-50-FPN | pytorch | 2x | random | Y | 25.6 | 25.4 | [config](./mask-rcnn_r50_fpn_seesaw-loss-normed-mask_random-ms-2x_lvis-v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r50_fpn_random_seesaw_loss_normed_mask_mstrain_2x_lvis_v1-a1c11314.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r50_fpn_random_seesaw_loss_normed_mask_mstrain_2x_lvis_v1.log.json) |
+| Mask R-CNN | R-101-FPN | pytorch | 2x | random | N | 27.4 | 26.7 | [config](./mask-rcnn_r101_fpn_seesaw-loss_random-ms-2x_lvis-v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r101_fpn_random_seesaw_loss_mstrain_2x_lvis_v1-8e6e6dd5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r101_fpn_random_seesaw_loss_mstrain_2x_lvis_v1.log.json) |
+| Mask R-CNN | R-101-FPN | pytorch | 2x | random | Y | 27.2 | 27.3 | [config](./mask-rcnn_r101_fpn_seesaw-loss-normed-mask_random-ms-2x_lvis-v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r101_fpn_random_seesaw_loss_normed_mask_mstrain_2x_lvis_v1-a0b59c42.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r101_fpn_random_seesaw_loss_normed_mask_mstrain_2x_lvis_v1.log.json) |
+| Mask R-CNN | R-50-FPN | pytorch | 2x | RFS | N | 27.6 | 26.4 | [config](./mask-rcnn_r50_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r50_fpn_sample1e-3_seesaw_loss_mstrain_2x_lvis_v1-392a804b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r50_fpn_sample1e-3_seesaw_loss_mstrain_2x_lvis_v1.log.json) |
+| Mask R-CNN | R-50-FPN | pytorch | 2x | RFS | Y | 27.6 | 26.8 | [config](./mask-rcnn_r50_fpn_seesaw-loss-normed-mask_sample1e-3-ms-2x_lvis-v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r50_fpn_sample1e-3_seesaw_loss_normed_mask_mstrain_2x_lvis_v1-cd0f6a12.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r50_fpn_sample1e-3_seesaw_loss_normed_mask_mstrain_2x_lvis_v1.log.json) |
+| Mask R-CNN | R-101-FPN | pytorch | 2x | RFS | N | 28.9 | 27.6 | [config](./mask-rcnn_r101_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r101_fpn_sample1e-3_seesaw_loss_mstrain_2x_lvis_v1-e68eb464.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r101_fpn_sample1e-3_seesaw_loss_mstrain_2x_lvis_v1.log.json) |
+| Mask R-CNN | R-101-FPN | pytorch | 2x | RFS | Y | 28.9 | 28.2 | [config](./mask-rcnn_r101_fpn_seesaw-loss-normed-mask_sample1e-3-ms-2x_lvis-v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r101_fpn_sample1e-3_seesaw_loss_normed_mask_mstrain_2x_lvis_v1-1d817139.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r101_fpn_sample1e-3_seesaw_loss_normed_mask_mstrain_2x_lvis_v1.log.json) |
+| Cascade Mask R-CNN | R-101-FPN | pytorch | 2x | random | N | 33.1 | 29.2 | [config](./cascade-mask-rcnn_r101_fpn_seesaw-loss_random-ms-2x_lvis-v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/cascade_mask_rcnn_r101_fpn_random_seesaw_loss_mstrain_2x_lvis_v1-71e2215e.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/cascade_mask_rcnn_r101_fpn_random_seesaw_loss_mstrain_2x_lvis_v1.log.json) |
+| Cascade Mask R-CNN | R-101-FPN | pytorch | 2x | random | Y | 33.0 | 30.0 | [config](./cascade-mask-rcnn_r101_fpn_seesaw-loss-normed-mask_random-ms-2x_lvis-v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/cascade_mask_rcnn_r101_fpn_random_seesaw_loss_normed_mask_mstrain_2x_lvis_v1-8b5a6745.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/cascade_mask_rcnn_r101_fpn_random_seesaw_loss_normed_mask_mstrain_2x_lvis_v1.log.json) |
+| Cascade Mask R-CNN | R-101-FPN | pytorch | 2x | RFS | N | 30.0 | 29.3 | [config](./cascade-mask-rcnn_r101_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/cascade_mask_rcnn_r101_fpn_sample1e-3_seesaw_loss_mstrain_2x_lvis_v1-5d8ca2a4.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/cascade_mask_rcnn_r101_fpn_sample1e-3_seesaw_loss_mstrain_2x_lvis_v1.log.json) |
+| Cascade Mask R-CNN | R-101-FPN | pytorch | 2x | RFS | Y | 32.8 | 30.1 | [config](./cascade-mask-rcnn_r101_fpn_seesaw-loss-normed-mask_sample1e-3-ms-2x_lvis-v1.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/cascade_mask_rcnn_r101_fpn_sample1e-3_seesaw_loss_normed_mask_mstrain_2x_lvis_v1-c8551505.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/cascade_mask_rcnn_r101_fpn_sample1e-3_seesaw_loss_normed_mask_mstrain_2x_lvis_v1.log.json) |
+
+## Citation
+
+We provide config files to reproduce the instance segmentation performance in the CVPR 2021 paper for [Seesaw Loss for Long-Tailed Instance Segmentation](https://arxiv.org/abs/2008.10032).
+
+```latex
+@inproceedings{wang2021seesaw,
+ title={Seesaw Loss for Long-Tailed Instance Segmentation},
+ author={Jiaqi Wang and Wenwei Zhang and Yuhang Zang and Yuhang Cao and Jiangmiao Pang and Tao Gong and Kai Chen and Ziwei Liu and Chen Change Loy and Dahua Lin},
+ booktitle={Proceedings of the {IEEE} Conference on Computer Vision and Pattern Recognition},
+ year={2021}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/cascade-mask-rcnn_r101_fpn_seesaw-loss-normed-mask_random-ms-2x_lvis-v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/cascade-mask-rcnn_r101_fpn_seesaw-loss-normed-mask_random-ms-2x_lvis-v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..2de87dcca59ccac7fc96c10c2a069fcf0464aeff
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/cascade-mask-rcnn_r101_fpn_seesaw-loss-normed-mask_random-ms-2x_lvis-v1.py
@@ -0,0 +1,5 @@
+_base_ = './cascade-mask-rcnn_r101_fpn_seesaw-loss_random-ms-2x_lvis-v1.py' # noqa: E501
+model = dict(
+ roi_head=dict(
+ mask_head=dict(
+ predictor_cfg=dict(type='NormedConv2d', tempearture=20))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/cascade-mask-rcnn_r101_fpn_seesaw-loss-normed-mask_sample1e-3-ms-2x_lvis-v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/cascade-mask-rcnn_r101_fpn_seesaw-loss-normed-mask_sample1e-3-ms-2x_lvis-v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..4d67ad7d4817a32b365bc2567937f69b68a9c97c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/cascade-mask-rcnn_r101_fpn_seesaw-loss-normed-mask_sample1e-3-ms-2x_lvis-v1.py
@@ -0,0 +1,5 @@
+_base_ = './cascade-mask-rcnn_r101_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1.py' # noqa: E501
+model = dict(
+ roi_head=dict(
+ mask_head=dict(
+ predictor_cfg=dict(type='NormedConv2d', tempearture=20))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/cascade-mask-rcnn_r101_fpn_seesaw-loss_random-ms-2x_lvis-v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/cascade-mask-rcnn_r101_fpn_seesaw-loss_random-ms-2x_lvis-v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..2a1a87d4203a12a78a26fd873bd6017fafb49cdf
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/cascade-mask-rcnn_r101_fpn_seesaw-loss_random-ms-2x_lvis-v1.py
@@ -0,0 +1,116 @@
+_base_ = [
+ '../_base_/models/cascade-mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_2x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')),
+ roi_head=dict(
+ bbox_head=[
+ dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=1203,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=True,
+ cls_predictor_cfg=dict(type='NormedLinear', tempearture=20),
+ loss_cls=dict(
+ type='SeesawLoss',
+ p=0.8,
+ q=2.0,
+ num_classes=1203,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0,
+ loss_weight=1.0)),
+ dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=1203,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.05, 0.05, 0.1, 0.1]),
+ reg_class_agnostic=True,
+ cls_predictor_cfg=dict(type='NormedLinear', tempearture=20),
+ loss_cls=dict(
+ type='SeesawLoss',
+ p=0.8,
+ q=2.0,
+ num_classes=1203,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0,
+ loss_weight=1.0)),
+ dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=1203,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.033, 0.033, 0.067, 0.067]),
+ reg_class_agnostic=True,
+ cls_predictor_cfg=dict(type='NormedLinear', tempearture=20),
+ loss_cls=dict(
+ type='SeesawLoss',
+ p=0.8,
+ q=2.0,
+ num_classes=1203,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0))
+ ],
+ mask_head=dict(num_classes=1203)),
+ test_cfg=dict(
+ rcnn=dict(
+ score_thr=0.0001,
+ # LVIS allows up to 300
+ max_per_img=300)))
+
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+dataset_type = 'LVISV1Dataset'
+data_root = 'data/lvis_v1/'
+train_dataloader = dict(
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/lvis_v1_train.json',
+ data_prefix=dict(img=''),
+ pipeline=train_pipeline))
+val_dataloader = dict(
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/lvis_v1_val.json',
+ data_prefix=dict(img='')))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='LVISMetric',
+ ann_file=data_root + 'annotations/lvis_v1_val.json',
+ metric=['bbox', 'segm'])
+test_evaluator = val_evaluator
+
+train_cfg = dict(val_interval=24)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/cascade-mask-rcnn_r101_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/cascade-mask-rcnn_r101_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..0e7b4df91368d23092a68f16ba4a35660ea23130
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/cascade-mask-rcnn_r101_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1.py
@@ -0,0 +1,95 @@
+_base_ = [
+ '../_base_/models/cascade-mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/lvis_v1_instance.py',
+ '../_base_/schedules/schedule_2x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')),
+ roi_head=dict(
+ bbox_head=[
+ dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=1203,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=True,
+ cls_predictor_cfg=dict(type='NormedLinear', tempearture=20),
+ loss_cls=dict(
+ type='SeesawLoss',
+ p=0.8,
+ q=2.0,
+ num_classes=1203,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0,
+ loss_weight=1.0)),
+ dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=1203,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.05, 0.05, 0.1, 0.1]),
+ reg_class_agnostic=True,
+ cls_predictor_cfg=dict(type='NormedLinear', tempearture=20),
+ loss_cls=dict(
+ type='SeesawLoss',
+ p=0.8,
+ q=2.0,
+ num_classes=1203,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0,
+ loss_weight=1.0)),
+ dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=1203,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.033, 0.033, 0.067, 0.067]),
+ reg_class_agnostic=True,
+ cls_predictor_cfg=dict(type='NormedLinear', tempearture=20),
+ loss_cls=dict(
+ type='SeesawLoss',
+ p=0.8,
+ q=2.0,
+ num_classes=1203,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0))
+ ],
+ mask_head=dict(num_classes=1203)),
+ test_cfg=dict(
+ rcnn=dict(
+ score_thr=0.0001,
+ # LVIS allows up to 300
+ max_per_img=300)))
+
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(dataset=dict(pipeline=train_pipeline)))
+
+train_cfg = dict(val_interval=24)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r101_fpn_seesaw-loss-normed-mask_random-ms-2x_lvis-v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r101_fpn_seesaw-loss-normed-mask_random-ms-2x_lvis-v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..b518c2135acb39a3d1119a8892c72816910ca496
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r101_fpn_seesaw-loss-normed-mask_random-ms-2x_lvis-v1.py
@@ -0,0 +1,6 @@
+_base_ = './mask-rcnn_r50_fpn_seesaw-loss-normed-mask_random-ms-2x_lvis-v1.py' # noqa: E501
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r101_fpn_seesaw-loss-normed-mask_sample1e-3-ms-2x_lvis-v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r101_fpn_seesaw-loss-normed-mask_sample1e-3-ms-2x_lvis-v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..008bbcae6eb8d189bdd0688b42d663eeba2a661e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r101_fpn_seesaw-loss-normed-mask_sample1e-3-ms-2x_lvis-v1.py
@@ -0,0 +1,6 @@
+_base_ = './mask-rcnn_r50_fpn_seesaw-loss-normed-mask_sample1e-3-ms-2x_lvis-v1.py' # noqa: E501
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r101_fpn_seesaw-loss_random-ms-2x_lvis-v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r101_fpn_seesaw-loss_random-ms-2x_lvis-v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..8a0b6755bf6f218c337d9ee16677e3e64886c019
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r101_fpn_seesaw-loss_random-ms-2x_lvis-v1.py
@@ -0,0 +1,6 @@
+_base_ = './mask-rcnn_r50_fpn_seesaw-loss_random-ms-2x_lvis-v1.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r101_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r101_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..6143231918e028523b6bb1792887ef7ce16dde02
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r101_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1.py
@@ -0,0 +1,6 @@
+_base_ = './mask-rcnn_r50_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r50_fpn_seesaw-loss-normed-mask_random-ms-2x_lvis-v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r50_fpn_seesaw-loss-normed-mask_random-ms-2x_lvis-v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..06d2438cf7c351a2fb352f787bc434cc6afc3ebb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r50_fpn_seesaw-loss-normed-mask_random-ms-2x_lvis-v1.py
@@ -0,0 +1,5 @@
+_base_ = './mask-rcnn_r50_fpn_seesaw-loss_random-ms-2x_lvis-v1.py'
+model = dict(
+ roi_head=dict(
+ mask_head=dict(
+ predictor_cfg=dict(type='NormedConv2d', tempearture=20))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r50_fpn_seesaw-loss-normed-mask_sample1e-3-ms-2x_lvis-v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r50_fpn_seesaw-loss-normed-mask_sample1e-3-ms-2x_lvis-v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..5fc68d3df32015e0fc8d5dd2bc92df416a8fc5fd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r50_fpn_seesaw-loss-normed-mask_sample1e-3-ms-2x_lvis-v1.py
@@ -0,0 +1,5 @@
+_base_ = './mask-rcnn_r50_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1.py'
+model = dict(
+ roi_head=dict(
+ mask_head=dict(
+ predictor_cfg=dict(type='NormedConv2d', tempearture=20))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r50_fpn_seesaw-loss_random-ms-2x_lvis-v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r50_fpn_seesaw-loss_random-ms-2x_lvis-v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..25c646c9c75c4468e71442049876a77382528e02
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r50_fpn_seesaw-loss_random-ms-2x_lvis-v1.py
@@ -0,0 +1,59 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_2x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ roi_head=dict(
+ bbox_head=dict(
+ num_classes=1203,
+ cls_predictor_cfg=dict(type='NormedLinear', tempearture=20),
+ loss_cls=dict(
+ type='SeesawLoss',
+ p=0.8,
+ q=2.0,
+ num_classes=1203,
+ loss_weight=1.0)),
+ mask_head=dict(num_classes=1203)),
+ test_cfg=dict(
+ rcnn=dict(
+ score_thr=0.0001,
+ # LVIS allows up to 300
+ max_per_img=300)))
+
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+dataset_type = 'LVISV1Dataset'
+data_root = 'data/lvis_v1/'
+train_dataloader = dict(
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/lvis_v1_train.json',
+ data_prefix=dict(img=''),
+ pipeline=train_pipeline))
+val_dataloader = dict(
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/lvis_v1_val.json',
+ data_prefix=dict(img='')))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='LVISMetric',
+ ann_file=data_root + 'annotations/lvis_v1_val.json',
+ metric=['bbox', 'segm'])
+test_evaluator = val_evaluator
+
+train_cfg = dict(val_interval=24)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r50_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r50_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..d60320e0b78035d24adb86f3aa184433951481fe
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/mask-rcnn_r50_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1.py
@@ -0,0 +1,38 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/lvis_v1_instance.py',
+ '../_base_/schedules/schedule_2x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ roi_head=dict(
+ bbox_head=dict(
+ num_classes=1203,
+ cls_predictor_cfg=dict(type='NormedLinear', tempearture=20),
+ loss_cls=dict(
+ type='SeesawLoss',
+ p=0.8,
+ q=2.0,
+ num_classes=1203,
+ loss_weight=1.0)),
+ mask_head=dict(num_classes=1203)),
+ test_cfg=dict(
+ rcnn=dict(
+ score_thr=0.0001,
+ # LVIS allows up to 300
+ max_per_img=300)))
+
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(dataset=dict(pipeline=train_pipeline)))
+
+train_cfg = dict(val_interval=24)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..374b9cde64ab1ff3c5f23971467846804738b0aa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/seesaw_loss/metafile.yml
@@ -0,0 +1,203 @@
+Collections:
+ - Name: Seesaw Loss
+ Metadata:
+ Training Data: LVIS
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Softmax
+ - RPN
+ - Convolution
+ - Dense Connections
+ - FPN
+ - ResNet
+ - RoIAlign
+ - Seesaw Loss
+ Paper:
+ URL: https://arxiv.org/abs/2008.10032
+ Title: 'Seesaw Loss for Long-Tailed Instance Segmentation'
+ README: configs/seesaw_loss/README.md
+
+Models:
+ - Name: mask-rcnn_r50_fpn_random_seesaw_loss_mstrain_2x_lvis_v1
+ In Collection: Seesaw Loss
+ Config: configs/seesaw_loss/mask-rcnn_r50_fpn_seesaw-loss_random-ms-2x_lvis-v1.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v1
+ Metrics:
+ box AP: 25.6
+ - Task: Instance Segmentation
+ Dataset: LVIS v1
+ Metrics:
+ mask AP: 25.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r50_fpn_random_seesaw_loss_mstrain_2x_lvis_v1-a698dd3d.pth
+ - Name: mask-rcnn_r50_fpn_random_seesaw_loss_normed_mask_mstrain_2x_lvis_v1
+ In Collection: Seesaw Loss
+ Config: configs/seesaw_loss/mask-rcnn_r50_fpn_seesaw-loss-normed-mask_random-ms-2x_lvis-v1.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v1
+ Metrics:
+ box AP: 25.6
+ - Task: Instance Segmentation
+ Dataset: LVIS v1
+ Metrics:
+ mask AP: 25.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r50_fpn_random_seesaw_loss_normed_mask_mstrain_2x_lvis_v1-a1c11314.pth
+ - Name: mask-rcnn_r101_fpn_seesaw-loss_random-ms-2x_lvis-v1
+ In Collection: Seesaw Loss
+ Config: configs/seesaw_loss/mask-rcnn_r101_fpn_seesaw-loss_random-ms-2x_lvis-v1.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v1
+ Metrics:
+ box AP: 27.4
+ - Task: Instance Segmentation
+ Dataset: LVIS v1
+ Metrics:
+ mask AP: 26.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r101_fpn_random_seesaw_loss_mstrain_2x_lvis_v1-8e6e6dd5.pth
+ - Name: mask-rcnn_r101_fpn_seesaw-loss-normed-mask_random-ms-2x_lvis-v1
+ In Collection: Seesaw Loss
+ Config: configs/seesaw_loss/mask-rcnn_r101_fpn_seesaw-loss-normed-mask_random-ms-2x_lvis-v1.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v1
+ Metrics:
+ box AP: 27.2
+ - Task: Instance Segmentation
+ Dataset: LVIS v1
+ Metrics:
+ mask AP: 27.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r101_fpn_random_seesaw_loss_normed_mask_mstrain_2x_lvis_v1-a0b59c42.pth
+ - Name: mask-rcnn_r50_fpn_sample1e-3_seesaw_loss_mstrain_2x_lvis_v1
+ In Collection: Seesaw Loss
+ Config: configs/seesaw_loss/mask-rcnn_r50_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v1
+ Metrics:
+ box AP: 27.6
+ - Task: Instance Segmentation
+ Dataset: LVIS v1
+ Metrics:
+ mask AP: 26.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r50_fpn_sample1e-3_seesaw_loss_mstrain_2x_lvis_v1-392a804b.pth
+ - Name: mask-rcnn_r50_fpn_sample1e-3_seesaw_loss_normed_mask_mstrain_2x_lvis_v1
+ In Collection: Seesaw Loss
+ Config: configs/seesaw_loss/mask-rcnn_r50_fpn_seesaw-loss-normed-mask_sample1e-3-ms-2x_lvis-v1.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v1
+ Metrics:
+ box AP: 27.6
+ - Task: Instance Segmentation
+ Dataset: LVIS v1
+ Metrics:
+ mask AP: 26.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r50_fpn_sample1e-3_seesaw_loss_normed_mask_mstrain_2x_lvis_v1-cd0f6a12.pth
+ - Name: mask-rcnn_r101_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1
+ In Collection: Seesaw Loss
+ Config: configs/seesaw_loss/mask-rcnn_r101_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v1
+ Metrics:
+ box AP: 28.9
+ - Task: Instance Segmentation
+ Dataset: LVIS v1
+ Metrics:
+ mask AP: 27.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r101_fpn_sample1e-3_seesaw_loss_mstrain_2x_lvis_v1-e68eb464.pth
+ - Name: mask-rcnn_r101_fpn_seesaw-loss-normed-mask_sample1e-3-ms-2x_lvis-v1
+ In Collection: Seesaw Loss
+ Config: configs/seesaw_loss/mask-rcnn_r101_fpn_seesaw-loss-normed-mask_sample1e-3-ms-2x_lvis-v1.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v1
+ Metrics:
+ box AP: 28.9
+ - Task: Instance Segmentation
+ Dataset: LVIS v1
+ Metrics:
+ mask AP: 28.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/mask_rcnn_r101_fpn_sample1e-3_seesaw_loss_normed_mask_mstrain_2x_lvis_v1-1d817139.pth
+ - Name: cascade-mask-rcnn_r101_fpn_seesaw-loss_random-ms-2x_lvis-v1
+ In Collection: Seesaw Loss
+ Config: configs/seesaw_loss/cascade-mask-rcnn_r101_fpn_seesaw-loss_random-ms-2x_lvis-v1.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v1
+ Metrics:
+ box AP: 33.1
+ - Task: Instance Segmentation
+ Dataset: LVIS v1
+ Metrics:
+ mask AP: 29.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/cascade_mask_rcnn_r101_fpn_random_seesaw_loss_mstrain_2x_lvis_v1-71e2215e.pth
+ - Name: cascade-mask-rcnn_r101_fpn_seesaw-loss-normed-mask_random-ms-2x_lvis-v1
+ In Collection: Seesaw Loss
+ Config: configs/seesaw_loss/cascade-mask-rcnn_r101_fpn_seesaw-loss-normed-mask_random-ms-2x_lvis-v1.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v1
+ Metrics:
+ box AP: 33.0
+ - Task: Instance Segmentation
+ Dataset: LVIS v1
+ Metrics:
+ mask AP: 30.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/cascade_mask_rcnn_r101_fpn_random_seesaw_loss_normed_mask_mstrain_2x_lvis_v1-8b5a6745.pth
+ - Name: cascade-mask-rcnn_r101_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1
+ In Collection: Seesaw Loss
+ Config: configs/seesaw_loss/cascade-mask-rcnn_r101_fpn_seesaw-loss_sample1e-3-ms-2x_lvis-v1.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v1
+ Metrics:
+ box AP: 30.0
+ - Task: Instance Segmentation
+ Dataset: LVIS v1
+ Metrics:
+ mask AP: 29.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/cascade_mask_rcnn_r101_fpn_sample1e-3_seesaw_loss_mstrain_2x_lvis_v1-5d8ca2a4.pth
+ - Name: cascade-mask-rcnn_r101_fpn_seesaw-loss-normed-mask_sample1e-3-ms-2x_lvis-v1
+ In Collection: Seesaw Loss
+ Config: configs/seesaw_loss/cascade-mask-rcnn_r101_fpn_seesaw-loss-normed-mask_sample1e-3-ms-2x_lvis-v1.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: LVIS v1
+ Metrics:
+ box AP: 32.8
+ - Task: Instance Segmentation
+ Dataset: LVIS v1
+ Metrics:
+ mask AP: 30.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/seesaw_loss/cascade_mask_rcnn_r101_fpn_sample1e-3_seesaw_loss_normed_mask_mstrain_2x_lvis_v1-c8551505.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/selfsup_pretrain/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/selfsup_pretrain/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..57537dddaca80756b7a6fc582808907edc8d850a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/selfsup_pretrain/README.md
@@ -0,0 +1,109 @@
+# Backbones Trained by Self-Supervise Algorithms
+
+
+
+## Abstract
+
+Unsupervised image representations have significantly reduced the gap with supervised pretraining, notably with the recent achievements of contrastive learning methods. These contrastive methods typically work online and rely on a large number of explicit pairwise feature comparisons, which is computationally challenging. In this paper, we propose an online algorithm, SwAV, that takes advantage of contrastive methods without requiring to compute pairwise comparisons. Specifically, our method simultaneously clusters the data while enforcing consistency between cluster assignments produced for different augmentations (or views) of the same image, instead of comparing features directly as in contrastive learning. Simply put, we use a swapped prediction mechanism where we predict the cluster assignment of a view from the representation of another view. Our method can be trained with large and small batches and can scale to unlimited amounts of data. Compared to previous contrastive methods, our method is more memory efficient since it does not require a large memory bank or a special momentum network. In addition, we also propose a new data augmentation strategy, multi-crop, that uses a mix of views with different resolutions in place of two full-resolution views, without increasing the memory or compute requirements much. We validate our findings by achieving 75.3% top-1 accuracy on ImageNet with ResNet-50, as well as surpassing supervised pretraining on all the considered transfer tasks.
+
+
+

+
+
+We present Momentum Contrast (MoCo) for unsupervised visual representation learning. From a perspective on contrastive learning as dictionary look-up, we build a dynamic dictionary with a queue and a moving-averaged encoder. This enables building a large and consistent dictionary on-the-fly that facilitates contrastive unsupervised learning. MoCo provides competitive results under the common linear protocol on ImageNet classification. More importantly, the representations learned by MoCo transfer well to downstream tasks. MoCo can outperform its supervised pre-training counterpart in 7 detection/segmentation tasks on PASCAL VOC, COCO, and other datasets, sometimes surpassing it by large margins. This suggests that the gap between unsupervised and supervised representation learning has been largely closed in many vision tasks.
+
+
+

+
+
+## Usage
+
+To use a self-supervisely pretrained backbone, there are two steps to do:
+
+1. Download and convert the model to PyTorch-style supported by MMDetection
+2. Modify the config and change the training setting accordingly
+
+### Convert model
+
+For more general usage, we also provide script `selfsup2mmdet.py` in the tools directory to convert the key of models pretrained by different self-supervised methods to PyTorch-style checkpoints used in MMDetection.
+
+```bash
+python -u tools/model_converters/selfsup2mmdet.py ${PRETRAIN_PATH} ${STORE_PATH} --selfsup ${method}
+```
+
+This script convert model from `PRETRAIN_PATH` and store the converted model in `STORE_PATH`.
+
+For example, to use a ResNet-50 backbone released by MoCo, you can download it from [here](https://dl.fbaipublicfiles.com/moco/moco_checkpoints/moco_v2_800ep/moco_v2_800ep_pretrain.pth.tar) and use the following command
+
+```bash
+python -u tools/model_converters/selfsup2mmdet.py ./moco_v2_800ep_pretrain.pth.tar mocov2_r50_800ep_pretrain.pth --selfsup moco
+```
+
+To use the ResNet-50 backbone released by SwAV, you can download it from [here](https://dl.fbaipublicfiles.com/deepcluster/swav_800ep_pretrain.pth.tar)
+
+### Modify config
+
+The backbone requires SyncBN and the `frozen_stages` need to be changed. A config that use the moco backbone is as below
+
+```python
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ pretrained='./mocov2_r50_800ep_pretrain.pth',
+ backbone=dict(
+ frozen_stages=0,
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ norm_eval=False))
+
+```
+
+## Results and Models
+
+| Method | Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :-------: | :------------------------------------------------------------: | :-----: | :------------: | :------: | :------------: | :----: | :-----: | :----------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Mask RCNN | [R50 by MoCo v2](./mask-rcnn_r50-mocov2-pre_fpn_1x_coco.py) | pytorch | 1x | | | 38.0 | 34.3 | [config](./mask-rcnn_r50-mocov2-pre_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/selfsup_pretrain/mask_rcnn_r50_fpn_mocov2-pretrain_1x_coco/mask_rcnn_r50_fpn_mocov2-pretrain_1x_coco_20210604_114614-a8b63483.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/selfsup_pretrain/mask_rcnn_r50_fpn_mocov2-pretrain_1x_coco/mask_rcnn_r50_fpn_mocov2-pretrain_1x_coco_20210604_114614.log.json) |
+| Mask RCNN | [R50 by MoCo v2](./mask-rcnn_r50-mocov2-pre_fpn_ms-2x_coco.py) | pytorch | multi-scale 2x | | | 40.8 | 36.8 | [config](./mask-rcnn_r50-mocov2-pre_fpn_ms-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/selfsup_pretrain/mask_rcnn_r50_fpn_mocov2-pretrain_ms-2x_coco/mask_rcnn_r50_fpn_mocov2-pretrain_ms-2x_coco_20210605_163717-d95df20a.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/selfsup_pretrain/mask_rcnn_r50_fpn_mocov2-pretrain_ms-2x_coco/mask_rcnn_r50_fpn_mocov2-pretrain_ms-2x_coco_20210605_163717.log.json) |
+| Mask RCNN | [R50 by SwAV](./mask-rcnn_r50-swav-pre_fpn_1x_coco.py) | pytorch | 1x | | | 39.1 | 35.7 | [config](./mask-rcnn_r50-swav-pre_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/selfsup_pretrain/mask_rcnn_r50_fpn_swav-pretrain_1x_coco/mask_rcnn_r50_fpn_swav-pretrain_1x_coco_20210604_114640-7b9baf28.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/selfsup_pretrain/mask_rcnn_r50_fpn_swav-pretrain_1x_coco/mask_rcnn_r50_fpn_swav-pretrain_1x_coco_20210604_114640.log.json) |
+| Mask RCNN | [R50 by SwAV](./mask-rcnn_r50-swav-pre_fpn_ms-2x_coco.py) | pytorch | multi-scale 2x | | | 41.3 | 37.3 | [config](./mask-rcnn_r50-swav-pre_fpn_ms-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/selfsup_pretrain/mask_rcnn_r50_fpn_swav-pretrain_ms-2x_coco/mask_rcnn_r50_fpn_swav-pretrain_ms-2x_coco_20210605_163717-08e26fca.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/selfsup_pretrain/mask_rcnn_r50_fpn_swav-pretrain_ms-2x_coco/mask_rcnn_r50_fpn_swav-pretrain_ms-2x_coco_20210605_163717.log.json) |
+
+### Notice
+
+1. We only provide single-scale 1x and multi-scale 2x configs as examples to show how to use backbones trained by self-supervised algorithms. We will try to reproduce the results in their corresponding paper using the released backbone in the future. Please stay tuned.
+
+## Citation
+
+We support to apply the backbone models pre-trained by different self-supervised methods in detection systems and provide their results on Mask R-CNN.
+
+The pre-trained models are converted from [MoCo](https://github.com/facebookresearch/moco) and downloaded from [SwAV](https://github.com/facebookresearch/swav).
+
+For SwAV, please cite
+
+```latex
+@article{caron2020unsupervised,
+ title={Unsupervised Learning of Visual Features by Contrasting Cluster Assignments},
+ author={Caron, Mathilde and Misra, Ishan and Mairal, Julien and Goyal, Priya and Bojanowski, Piotr and Joulin, Armand},
+ booktitle={Proceedings of Advances in Neural Information Processing Systems (NeurIPS)},
+ year={2020}
+}
+```
+
+For MoCo, please cite
+
+```latex
+@Article{he2019moco,
+ author = {Kaiming He and Haoqi Fan and Yuxin Wu and Saining Xie and Ross Girshick},
+ title = {Momentum Contrast for Unsupervised Visual Representation Learning},
+ journal = {arXiv preprint arXiv:1911.05722},
+ year = {2019},
+}
+@Article{chen2020mocov2,
+ author = {Xinlei Chen and Haoqi Fan and Ross Girshick and Kaiming He},
+ title = {Improved Baselines with Momentum Contrastive Learning},
+ journal = {arXiv preprint arXiv:2003.04297},
+ year = {2020},
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/selfsup_pretrain/mask-rcnn_r50-mocov2-pre_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/selfsup_pretrain/mask-rcnn_r50-mocov2-pre_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..91d45add8aba54de4b25fba11ecf5e18bca0084f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/selfsup_pretrain/mask-rcnn_r50-mocov2-pre_fpn_1x_coco.py
@@ -0,0 +1,13 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ backbone=dict(
+ frozen_stages=0,
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ norm_eval=False,
+ init_cfg=dict(
+ type='Pretrained', checkpoint='./mocov2_r50_800ep_pretrain.pth')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/selfsup_pretrain/mask-rcnn_r50-mocov2-pre_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/selfsup_pretrain/mask-rcnn_r50-mocov2-pre_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ddaebf5558a22680d556aa8b3fe79541d634d910
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/selfsup_pretrain/mask-rcnn_r50-mocov2-pre_fpn_ms-2x_coco.py
@@ -0,0 +1,25 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_2x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ backbone=dict(
+ frozen_stages=0,
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ norm_eval=False,
+ init_cfg=dict(
+ type='Pretrained', checkpoint='./mocov2_r50_800ep_pretrain.pth')))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomResize', scale=[(1333, 640), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/selfsup_pretrain/mask-rcnn_r50-swav-pre_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/selfsup_pretrain/mask-rcnn_r50-swav-pre_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..785c80ec9d14c8e4b54b2e3359f9b4c680eaca17
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/selfsup_pretrain/mask-rcnn_r50-swav-pre_fpn_1x_coco.py
@@ -0,0 +1,13 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ backbone=dict(
+ frozen_stages=0,
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ norm_eval=False,
+ init_cfg=dict(
+ type='Pretrained', checkpoint='./swav_800ep_pretrain.pth.tar')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/selfsup_pretrain/mask-rcnn_r50-swav-pre_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/selfsup_pretrain/mask-rcnn_r50-swav-pre_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c393e0b36047f731c91c3f0963ef90347a0910e9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/selfsup_pretrain/mask-rcnn_r50-swav-pre_fpn_ms-2x_coco.py
@@ -0,0 +1,25 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_2x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ backbone=dict(
+ frozen_stages=0,
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ norm_eval=False,
+ init_cfg=dict(
+ type='Pretrained', checkpoint='./swav_800ep_pretrain.pth.tar')))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomResize', scale=[(1333, 640), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/simple_copy_paste/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/simple_copy_paste/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..23b09ce5dbb3e2cba41cad7b6b45fccd95996fb1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/simple_copy_paste/README.md
@@ -0,0 +1,38 @@
+# SimpleCopyPaste
+
+> [Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation](https://arxiv.org/abs/2012.07177)
+
+
+
+## Abstract
+
+Building instance segmentation models that are data-efficient and can handle rare object categories is an important challenge in computer vision. Leveraging data augmentations is a promising direction towards addressing this challenge. Here, we perform a systematic study of the Copy-Paste augmentation (\[13, 12\]) for instance segmentation where we randomly paste objects onto an image. Prior studies on Copy-Paste relied on modeling the surrounding visual context for pasting the objects. However, we find that the simple mechanism of pasting objects randomly is good enough and can provide solid gains on top of strong baselines. Furthermore, we show Copy-Paste is additive with semi-supervised methods that leverage extra data through pseudo labeling (e.g. self-training). On COCO instance segmentation, we achieve 49.1 mask AP and 57.3 box AP, an improvement of +0.6 mask AP and +1.5 box AP over the previous state-of-the-art. We further demonstrate that Copy-Paste can lead to significant improvements on the LVIS benchmark. Our baseline model outperforms the LVIS 2020 Challenge winning entry by +3.6 mask AP on rare categories.
+
+
+

+
+
+## Results and Models
+
+### Mask R-CNN with Standard Scale Jittering (SSJ) and Simple Copy-Paste(SCP)
+
+Standard Scale Jittering(SSJ) resizes and crops an image with a resize range of 0.8 to 1.25 of the original image size, and Simple Copy-Paste(SCP) selects a random subset of objects from one of the images and pastes them onto the other image.
+
+| Backbone | Training schedule | Augmentation | batch size | box AP | mask AP | Config | Download |
+| :------: | :---------------: | :----------: | :--------: | :----: | :-----: | :------------------------------------------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | 90k | SSJ | 64 | 43.3 | 39.0 | [config](./mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-90k_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/simple_copy_paste/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_32x2_90k_coco/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_32x2_90k_coco_20220316_181409-f79c84c5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/simple_copy_paste/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_32x2_90k_coco/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_32x2_90k_coco_20220316_181409.log.json) |
+| R-50 | 90k | SSJ+SCP | 64 | 43.8 | 39.2 | [config](./mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-scp-90k_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/simple_copy_paste/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_scp_32x2_90k_coco/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_scp_32x2_90k_coco_20220316_181307-6bc5726f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/simple_copy_paste/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_scp_32x2_90k_coco/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_scp_32x2_90k_coco_20220316_181307.log.json) |
+| R-50 | 270k | SSJ | 64 | 43.5 | 39.1 | [config](./mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-270k_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/simple_copy_paste/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_32x2_270k_coco/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_32x2_270k_coco_20220324_182940-33a100c5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/simple_copy_paste/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_32x2_270k_coco/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_32x2_270k_coco_20220324_182940.log.json) |
+| R-50 | 270k | SSJ+SCP | 64 | 45.1 | 40.3 | [config](./mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-scp-270k_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/simple_copy_paste/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_scp_32x2_270k_coco/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_scp_32x2_270k_coco_20220324_201229-80ee90b7.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/simple_copy_paste/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_scp_32x2_270k_coco/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_scp_32x2_270k_coco_20220324_201229.log.json) |
+
+## Citation
+
+```latex
+@inproceedings{ghiasi2021simple,
+ title={Simple copy-paste is a strong data augmentation method for instance segmentation},
+ author={Ghiasi, Golnaz and Cui, Yin and Srinivas, Aravind and Qian, Rui and Lin, Tsung-Yi and Cubuk, Ekin D and Le, Quoc V and Zoph, Barret},
+ booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
+ pages={2918--2928},
+ year={2021}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/simple_copy_paste/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-270k_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/simple_copy_paste/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-270k_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..0c6e081e860e1240f8d35efa8176563a8b5be845
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/simple_copy_paste/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-270k_coco.py
@@ -0,0 +1,31 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ # 270k iterations with batch_size 64 is roughly equivalent to 144 epochs
+ '../common/ssj_270k_coco-instance.py',
+]
+
+image_size = (1024, 1024)
+batch_augments = [
+ dict(type='BatchFixedSizePad', size=image_size, pad_mask=True)
+]
+norm_cfg = dict(type='SyncBN', requires_grad=True)
+# Use MMSyncBN that handles empty tensor in head. It can be changed to
+# SyncBN after https://github.com/pytorch/pytorch/issues/36530 is fixed
+head_norm_cfg = dict(type='MMSyncBN', requires_grad=True)
+model = dict(
+ # the model is trained from scratch, so init_cfg is None
+ data_preprocessor=dict(
+ # pad_size_divisor=32 is unnecessary in training but necessary
+ # in testing.
+ pad_size_divisor=32,
+ batch_augments=batch_augments),
+ backbone=dict(
+ frozen_stages=-1, norm_eval=False, norm_cfg=norm_cfg, init_cfg=None),
+ neck=dict(norm_cfg=norm_cfg),
+ rpn_head=dict(num_convs=2), # leads to 0.1+ mAP
+ roi_head=dict(
+ bbox_head=dict(
+ type='Shared4Conv1FCBBoxHead',
+ conv_out_channels=256,
+ norm_cfg=head_norm_cfg),
+ mask_head=dict(norm_cfg=head_norm_cfg)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/simple_copy_paste/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-90k_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/simple_copy_paste/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-90k_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..abe8962ac69184241e30628242e5313c52f503f4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/simple_copy_paste/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-90k_coco.py
@@ -0,0 +1,18 @@
+_base_ = 'mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-270k_coco.py' # noqa
+
+# training schedule for 90k
+max_iters = 90000
+
+# learning rate policy
+# lr steps at [0.9, 0.95, 0.975] of the maximum iterations
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.067, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=90000,
+ by_epoch=False,
+ milestones=[81000, 85500, 87750],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/simple_copy_paste/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-scp-270k_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/simple_copy_paste/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-scp-270k_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f0ea57d19728d7c563e56d139888059dd9c81317
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/simple_copy_paste/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-scp-270k_coco.py
@@ -0,0 +1,31 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ # 270k iterations with batch_size 64 is roughly equivalent to 144 epochs
+ '../common/ssj_scp_270k_coco-instance.py'
+]
+
+image_size = (1024, 1024)
+batch_augments = [
+ dict(type='BatchFixedSizePad', size=image_size, pad_mask=True)
+]
+norm_cfg = dict(type='SyncBN', requires_grad=True)
+# Use MMSyncBN that handles empty tensor in head. It can be changed to
+# SyncBN after https://github.com/pytorch/pytorch/issues/36530 is fixed
+head_norm_cfg = dict(type='MMSyncBN', requires_grad=True)
+model = dict(
+ # the model is trained from scratch, so init_cfg is None
+ data_preprocessor=dict(
+ # pad_size_divisor=32 is unnecessary in training but necessary
+ # in testing.
+ pad_size_divisor=32,
+ batch_augments=batch_augments),
+ backbone=dict(
+ frozen_stages=-1, norm_eval=False, norm_cfg=norm_cfg, init_cfg=None),
+ neck=dict(norm_cfg=norm_cfg),
+ rpn_head=dict(num_convs=2), # leads to 0.1+ mAP
+ roi_head=dict(
+ bbox_head=dict(
+ type='Shared4Conv1FCBBoxHead',
+ conv_out_channels=256,
+ norm_cfg=head_norm_cfg),
+ mask_head=dict(norm_cfg=head_norm_cfg)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/simple_copy_paste/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-scp-90k_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/simple_copy_paste/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-scp-90k_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e158b5c05aae3345ba9d4d1a55d1bbb82a789726
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/simple_copy_paste/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-scp-90k_coco.py
@@ -0,0 +1,18 @@
+_base_ = 'mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-scp-270k_coco.py' # noqa
+
+# training schedule for 90k
+max_iters = 90000
+
+# learning rate policy
+# lr steps at [0.9, 0.95, 0.975] of the maximum iterations
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.067, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=90000,
+ by_epoch=False,
+ milestones=[81000, 85500, 87750],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/simple_copy_paste/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/simple_copy_paste/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..8a40b658feeefd870300e62934ea21315218bfba
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/simple_copy_paste/metafile.yml
@@ -0,0 +1,92 @@
+Collections:
+ - Name: SimpleCopyPaste
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 32x A100 GPUs
+ Architecture:
+ - Softmax
+ - RPN
+ - Convolution
+ - Dense Connections
+ - FPN
+ - ResNet
+ - RoIAlign
+ Paper:
+ URL: https://arxiv.org/abs/2012.07177
+ Title: "Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation"
+ README: configs/simple_copy_paste/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.25.0/mmdet/datasets/pipelines/transforms.py#L2762
+ Version: v2.25.0
+
+Models:
+ - Name: mask-rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_32x2_270k_coco
+ In Collection: SimpleCopyPaste
+ Config: configs/simple_copy_paste/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-270k_coco.py
+ Metadata:
+ Training Memory (GB): 7.2
+ Iterations: 270000
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.5
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/simple_copy_paste/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_32x2_270k_coco/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_32x2_270k_coco_20220324_182940-33a100c5.pth
+
+ - Name: mask-rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_32x2_90k_coco
+ In Collection: SimpleCopyPaste
+ Config: configs/simple_copy_paste/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-90k_coco.py
+ Metadata:
+ Training Memory (GB): 7.2
+ Iterations: 90000
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.3
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/simple_copy_paste/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_32x2_90k_coco/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_32x2_90k_coco_20220316_181409-f79c84c5.pth
+
+ - Name: mask-rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_scp_32x2_270k_coco
+ In Collection: SimpleCopyPaste
+ Config: configs/simple_copy_paste/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-scp-270k_coco.py
+ Metadata:
+ Training Memory (GB): 7.2
+ Iterations: 270000
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.1
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 40.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/simple_copy_paste/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_scp_32x2_270k_coco/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_scp_32x2_270k_coco_20220324_201229-80ee90b7.pth
+
+ - Name: mask-rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_scp_32x2_90k_coco
+ In Collection: SimpleCopyPaste
+ Config: configs/simple_copy_paste/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_32xb2-ssj-scp-90k_coco.py
+ Metadata:
+ Training Memory (GB): 7.2
+ Iterations: 90000
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.8
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/simple_copy_paste/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_scp_32x2_90k_coco/mask_rcnn_r50_fpn_syncbn-all_rpn-2conv_ssj_scp_32x2_90k_coco_20220316_181307-6bc5726f.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/soft_teacher/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/soft_teacher/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..1fd3d84dc36b8f7e4a0342e951f81979f1a9dce9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/soft_teacher/README.md
@@ -0,0 +1,33 @@
+# SoftTeacher
+
+> [End-to-End Semi-Supervised Object Detection with Soft Teacher](https://arxiv.org/abs/2106.09018)
+
+
+
+## Abstract
+
+This paper presents an end-to-end semi-supervised object detection approach, in contrast to previous more complex multi-stage methods. The end-to-end training gradually improves pseudo label qualities during the curriculum, and the more and more accurate pseudo labels in turn benefit object detection training. We also propose two simple yet effective techniques within this framework: a soft teacher mechanism where the classification loss of each unlabeled bounding box is weighed by the classification score produced by the teacher network; a box jittering approach to select reliable pseudo boxes for the learning of box regression. On the COCO benchmark, the proposed approach outperforms previous methods by a large margin under various labeling ratios, i.e. 1%, 5% and 10%. Moreover, our approach proves to perform also well when the amount of labeled data is relatively large. For example, it can improve a 40.9 mAP baseline detector trained using the full COCO training set by +3.6 mAP, reaching 44.5 mAP, by leveraging the 123K unlabeled images of COCO. On the state-of-the-art Swin Transformer based object detector (58.9 mAP on test-dev), it can still significantly improve the detection accuracy by +1.5 mAP, reaching 60.4 mAP, and improve the instance segmentation accuracy by +1.2 mAP, reaching 52.4 mAP. Further incorporating with the Object365 pre-trained model, the detection accuracy reaches 61.3 mAP and the instance segmentation accuracy reaches 53.0 mAP, pushing the new state-of-the-art.
+
+
+

+
+
+## Results and Models
+
+| Model | Detector | Labeled Dataset | Iteration | box AP | Config | Download |
+| :---------: | :----------: | :-------------: | :-------: | :----: | :-----------------------------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| SoftTeacher | Faster R-CNN | COCO-1% | 180k | 19.9 | [config](./soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.01-coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.01-coco/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0_20230330_233412-3c8f6d4a.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.01-coco/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0_20230330_233412.log.json) |
+| SoftTeacher | Faster R-CNN | COCO-2% | 180k | 24.9 | [config](./soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.02-coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.02-coco/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0_20230331_020244-c0d2c3aa.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.02-coco/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0_20230331_020244.log.json) |
+| SoftTeacher | Faster R-CNN | COCO-5% | 180k | 30.4 | [config](./soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.05-coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.05-coco/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0_20230331_070656-308798ad.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.05-coco/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0_20230331_070656.log.json) |
+| SoftTeacher | Faster R-CNN | COCO-10% | 180k | 33.8 | [config](./soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.1-coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.1-coco/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0_20230330_232113-b46f78d0.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.1-coco/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0_20230330_232113.log.json) |
+
+## Citation
+
+```latex
+@article{xu2021end,
+ title={End-to-End Semi-Supervised Object Detection with Soft Teacher},
+ author={Xu, Mengde and Zhang, Zheng and Hu, Han and Wang, Jianfeng and Wang, Lijuan and Wei, Fangyun and Bai, Xiang and Liu, Zicheng},
+ journal={Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
+ year={2021}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/soft_teacher/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/soft_teacher/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..9622acec93ad3138daff09930ecfa2807dc7748a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/soft_teacher/metafile.yml
@@ -0,0 +1,67 @@
+Collections:
+ - Name: SoftTeacher
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x A100 GPUs
+ Architecture:
+ - FPN
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/2106.09018
+ Title: "End-to-End Semi-Supervised Object Detection with Soft Teacher"
+ README: configs/soft_teacher/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v3.0.0rc1/mmdet/models/detectors/soft_teacher.py#L20
+ Version: v3.0.0rc1
+
+Models:
+ - Name: soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.01-coco.py
+ In Collection: SoftTeacher
+ Config: configs/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.01-coco.py
+ Metadata:
+ Iterations: 180000
+ Results:
+ - Task: Semi-Supervised Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 19.9
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.01-coco/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0_20230330_233412-3c8f6d4a.pth
+
+ - Name: soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.02-coco.py
+ In Collection: SoftTeacher
+ Config: configs/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.02-coco.py
+ Metadata:
+ Iterations: 180000
+ Results:
+ - Task: Semi-Supervised Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 24.9
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.02-coco/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0_20230331_020244-c0d2c3aa.pth
+
+ - Name: soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.05-coco.py
+ In Collection: SoftTeacher
+ Config: configs/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.05-coco.py
+ Metadata:
+ Iterations: 180000
+ Results:
+ - Task: Semi-Supervised Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 30.4
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.05-coco/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0_20230331_070656-308798ad.pth
+
+ - Name: soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.1-coco.py
+ In Collection: SoftTeacher
+ Config: configs/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.1-coco.py
+ Metadata:
+ Iterations: 180000
+ Results:
+ - Task: Semi-Supervised Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 33.8
+ Weights: https://download.openmmlab.com/mmdetection/v3.0/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.1-coco/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0_20230330_232113-b46f78d0.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.01-coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.01-coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..2bd09645598204482e9f88f6baf00d32eba9cab6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.01-coco.py
@@ -0,0 +1,9 @@
+_base_ = ['soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.1-coco.py']
+
+# 1% coco train2017 is set as labeled dataset
+labeled_dataset = _base_.labeled_dataset
+unlabeled_dataset = _base_.unlabeled_dataset
+labeled_dataset.ann_file = 'semi_anns/instances_train2017.1@1.json'
+unlabeled_dataset.ann_file = 'semi_anns/instances_train2017.1@1-unlabeled.json'
+train_dataloader = dict(
+ dataset=dict(datasets=[labeled_dataset, unlabeled_dataset]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.02-coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.02-coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8ca38c931926cef33321f931b0c6d5c66824ff55
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.02-coco.py
@@ -0,0 +1,9 @@
+_base_ = ['soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.1-coco.py']
+
+# 2% coco train2017 is set as labeled dataset
+labeled_dataset = _base_.labeled_dataset
+unlabeled_dataset = _base_.unlabeled_dataset
+labeled_dataset.ann_file = 'semi_anns/instances_train2017.1@2.json'
+unlabeled_dataset.ann_file = 'semi_anns/instances_train2017.1@2-unlabeled.json'
+train_dataloader = dict(
+ dataset=dict(datasets=[labeled_dataset, unlabeled_dataset]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.05-coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.05-coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..750b7ed6df6c91bab8f68f58f339b2f3696fa693
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.05-coco.py
@@ -0,0 +1,9 @@
+_base_ = ['soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.1-coco.py']
+
+# 5% coco train2017 is set as labeled dataset
+labeled_dataset = _base_.labeled_dataset
+unlabeled_dataset = _base_.unlabeled_dataset
+labeled_dataset.ann_file = 'semi_anns/instances_train2017.1@5.json'
+unlabeled_dataset.ann_file = 'semi_anns/instances_train2017.1@5-unlabeled.json'
+train_dataloader = dict(
+ dataset=dict(datasets=[labeled_dataset, unlabeled_dataset]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.1-coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.1-coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3713aef442f4add55efafde08b2c98da1773bab0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/soft_teacher/soft-teacher_faster-rcnn_r50-caffe_fpn_180k_semi-0.1-coco.py
@@ -0,0 +1,84 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py', '../_base_/default_runtime.py',
+ '../_base_/datasets/semi_coco_detection.py'
+]
+
+detector = _base_.model
+detector.data_preprocessor = dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32)
+detector.backbone = dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe'))
+
+model = dict(
+ _delete_=True,
+ type='SoftTeacher',
+ detector=detector,
+ data_preprocessor=dict(
+ type='MultiBranchDataPreprocessor',
+ data_preprocessor=detector.data_preprocessor),
+ semi_train_cfg=dict(
+ freeze_teacher=True,
+ sup_weight=1.0,
+ unsup_weight=4.0,
+ pseudo_label_initial_score_thr=0.5,
+ rpn_pseudo_thr=0.9,
+ cls_pseudo_thr=0.9,
+ reg_pseudo_thr=0.02,
+ jitter_times=10,
+ jitter_scale=0.06,
+ min_pseudo_bbox_wh=(1e-2, 1e-2)),
+ semi_test_cfg=dict(predict_on='teacher'))
+
+# 10% coco train2017 is set as labeled dataset
+labeled_dataset = _base_.labeled_dataset
+unlabeled_dataset = _base_.unlabeled_dataset
+labeled_dataset.ann_file = 'semi_anns/instances_train2017.1@10.json'
+unlabeled_dataset.ann_file = 'semi_anns/' \
+ 'instances_train2017.1@10-unlabeled.json'
+unlabeled_dataset.data_prefix = dict(img='train2017/')
+train_dataloader = dict(
+ dataset=dict(datasets=[labeled_dataset, unlabeled_dataset]))
+
+# training schedule for 180k
+train_cfg = dict(
+ type='IterBasedTrainLoop', max_iters=180000, val_interval=5000)
+val_cfg = dict(type='TeacherStudentValLoop')
+test_cfg = dict(type='TestLoop')
+
+# learning rate policy
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=180000,
+ by_epoch=False,
+ milestones=[120000, 160000],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
+
+default_hooks = dict(
+ checkpoint=dict(by_epoch=False, interval=10000, max_keep_ckpts=2))
+log_processor = dict(by_epoch=False)
+
+custom_hooks = [dict(type='MeanTeacherHook')]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..4a36676b1b5e0fafd3bfb1cbe4a6cef5fd549c57
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/README.md
@@ -0,0 +1,54 @@
+# SOLO
+
+> [SOLO: Segmenting Objects by Locations](https://arxiv.org/abs/1912.04488)
+
+
+
+## Abstract
+
+We present a new, embarrassingly simple approach to instance segmentation in images. Compared to many other dense prediction tasks, e.g., semantic segmentation, it is the arbitrary number of instances that have made instance segmentation much more challenging. In order to predict a mask for each instance, mainstream approaches either follow the 'detect-thensegment' strategy as used by Mask R-CNN, or predict category masks first then use clustering techniques to group pixels into individual instances. We view the task of instance segmentation from a completely new perspective by introducing the notion of "instance categories", which assigns categories to each pixel within an instance according to the instance's location and size, thus nicely converting instance mask segmentation into a classification-solvable problem. Now instance segmentation is decomposed into two classification tasks. We demonstrate a much simpler and flexible instance segmentation framework with strong performance, achieving on par accuracy with Mask R-CNN and outperforming recent singleshot instance segmenters in accuracy. We hope that this very simple and strong framework can serve as a baseline for many instance-level recognition tasks besides instance segmentation.
+
+
+

+
+
+## Results and Models
+
+### SOLO
+
+| Backbone | Style | MS train | Lr schd | Mem (GB) | Inf time (fps) | mask AP | Download |
+| :------: | :-----: | :------: | :-----: | :------: | :------------: | :-----: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | pytorch | N | 1x | 8.0 | 14.0 | 33.1 | [model](https://download.openmmlab.com/mmdetection/v2.0/solo/solo_r50_fpn_1x_coco/solo_r50_fpn_1x_coco_20210821_035055-2290a6b8.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/solo/solo_r50_fpn_1x_coco/solo_r50_fpn_1x_coco_20210821_035055.log.json) |
+| R-50 | pytorch | Y | 3x | 7.4 | 14.0 | 35.9 | [model](https://download.openmmlab.com/mmdetection/v2.0/solo/solo_r50_fpn_3x_coco/solo_r50_fpn_3x_coco_20210901_012353-11d224d7.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/solo/solo_r50_fpn_3x_coco/solo_r50_fpn_3x_coco_20210901_012353.log.json) |
+
+### Decoupled SOLO
+
+| Backbone | Style | MS train | Lr schd | Mem (GB) | Inf time (fps) | mask AP | Download |
+| :------: | :-----: | :------: | :-----: | :------: | :------------: | :-----: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | pytorch | N | 1x | 7.8 | 12.5 | 33.9 | [model](https://download.openmmlab.com/mmdetection/v2.0/solo/decoupled_solo_r50_fpn_1x_coco/decoupled_solo_r50_fpn_1x_coco_20210820_233348-6337c589.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/solo/decoupled_solo_r50_fpn_1x_coco/decoupled_solo_r50_fpn_1x_coco_20210820_233348.log.json) |
+| R-50 | pytorch | Y | 3x | 7.9 | 12.5 | 36.7 | [model](https://download.openmmlab.com/mmdetection/v2.0/solo/decoupled_solo_r50_fpn_3x_coco/decoupled_solo_r50_fpn_3x_coco_20210821_042504-7b3301ec.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/solo/decoupled_solo_r50_fpn_3x_coco/decoupled_solo_r50_fpn_3x_coco_20210821_042504.log.json) |
+
+- Decoupled SOLO has a decoupled head which is different from SOLO head.
+ Decoupled SOLO serves as an efficient and equivalent variant in accuracy
+ of SOLO. Please refer to the corresponding config files for details.
+
+### Decoupled Light SOLO
+
+| Backbone | Style | MS train | Lr schd | Mem (GB) | Inf time (fps) | mask AP | Download |
+| :------: | :-----: | :------: | :-----: | :------: | :------------: | :-----: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | pytorch | Y | 3x | 2.2 | 31.2 | 32.9 | [model](https://download.openmmlab.com/mmdetection/v2.0/solo/decoupled_solo_light_r50_fpn_3x_coco/decoupled_solo_light_r50_fpn_3x_coco_20210906_142703-e70e226f.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/solo/decoupled_solo_light_r50_fpn_3x_coco/decoupled_solo_light_r50_fpn_3x_coco_20210906_142703.log.json) |
+
+- Decoupled Light SOLO using decoupled structure similar to Decoupled
+ SOLO head, with light-weight head and smaller input size, Please refer
+ to the corresponding config files for details.
+
+## Citation
+
+```latex
+@inproceedings{wang2020solo,
+ title = {{SOLO}: Segmenting Objects by Locations},
+ author = {Wang, Xinlong and Kong, Tao and Shen, Chunhua and Jiang, Yuning and Li, Lei},
+ booktitle = {Proc. Eur. Conf. Computer Vision (ECCV)},
+ year = {2020}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/decoupled-solo-light_r50_fpn_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/decoupled-solo-light_r50_fpn_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..fc35df3c3cbbd70532e066de27b06418549eb906
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/decoupled-solo-light_r50_fpn_3x_coco.py
@@ -0,0 +1,50 @@
+_base_ = './decoupled-solo_r50_fpn_3x_coco.py'
+
+# model settings
+model = dict(
+ mask_head=dict(
+ type='DecoupledSOLOLightHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ strides=[8, 8, 16, 32, 32],
+ scale_ranges=((1, 64), (32, 128), (64, 256), (128, 512), (256, 2048)),
+ pos_scale=0.2,
+ num_grids=[40, 36, 24, 16, 12],
+ cls_down_index=0,
+ loss_mask=dict(
+ type='DiceLoss', use_sigmoid=True, activate=False,
+ loss_weight=3.0),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ norm_cfg=dict(type='GN', num_groups=32, requires_grad=True)))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(852, 512), (852, 480), (852, 448), (852, 416), (852, 384),
+ (852, 352)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(852, 512), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/decoupled-solo_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/decoupled-solo_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6d7f4b90c19d9fdcc3c895deb4101cf7acd7bd8e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/decoupled-solo_r50_fpn_1x_coco.py
@@ -0,0 +1,24 @@
+_base_ = './solo_r50_fpn_1x_coco.py'
+# model settings
+model = dict(
+ mask_head=dict(
+ type='DecoupledSOLOHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=7,
+ feat_channels=256,
+ strides=[8, 8, 16, 32, 32],
+ scale_ranges=((1, 96), (48, 192), (96, 384), (192, 768), (384, 2048)),
+ pos_scale=0.2,
+ num_grids=[40, 36, 24, 16, 12],
+ cls_down_index=0,
+ loss_mask=dict(
+ type='DiceLoss', use_sigmoid=True, activate=False,
+ loss_weight=3.0),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ norm_cfg=dict(type='GN', num_groups=32, requires_grad=True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/decoupled-solo_r50_fpn_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/decoupled-solo_r50_fpn_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..4a8c19decb72a3d904a277faac06670999f6b322
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/decoupled-solo_r50_fpn_3x_coco.py
@@ -0,0 +1,25 @@
+_base_ = './solo_r50_fpn_3x_coco.py'
+
+# model settings
+model = dict(
+ mask_head=dict(
+ type='DecoupledSOLOHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=7,
+ feat_channels=256,
+ strides=[8, 8, 16, 32, 32],
+ scale_ranges=((1, 96), (48, 192), (96, 384), (192, 768), (384, 2048)),
+ pos_scale=0.2,
+ num_grids=[40, 36, 24, 16, 12],
+ cls_down_index=0,
+ loss_mask=dict(
+ type='DiceLoss', use_sigmoid=True, activate=False,
+ loss_weight=3.0),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ norm_cfg=dict(type='GN', num_groups=32, requires_grad=True)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..aa38b8c07b3db7eb018bb769b6eca6e010a1d764
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/metafile.yml
@@ -0,0 +1,115 @@
+Collections:
+ - Name: SOLO
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - FPN
+ - Convolution
+ - ResNet
+ Paper: https://arxiv.org/abs/1912.04488
+ README: configs/solo/README.md
+
+Models:
+ - Name: decoupled-solo_r50_fpn_1x_coco
+ In Collection: SOLO
+ Config: configs/solo/decoupled-solo_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.8
+ Epochs: 12
+ inference time (ms/im):
+ - value: 116.4
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (1333, 800)
+ Results:
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 33.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/solo/decoupled_solo_r50_fpn_1x_coco/decoupled_solo_r50_fpn_1x_coco_20210820_233348-6337c589.pth
+
+ - Name: decoupled-solo_r50_fpn_3x_coco
+ In Collection: SOLO
+ Config: configs/solo/decoupled-solo_r50_fpn_3x_coco.py
+ Metadata:
+ Training Memory (GB): 7.9
+ Epochs: 36
+ inference time (ms/im):
+ - value: 117.2
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (1333, 800)
+ Results:
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 36.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/solo/decoupled_solo_r50_fpn_3x_coco/decoupled_solo_r50_fpn_3x_coco_20210821_042504-7b3301ec.pth
+
+ - Name: decoupled-solo-light_r50_fpn_3x_coco
+ In Collection: SOLO
+ Config: configs/solo/decoupled-solo-light_r50_fpn_3x_coco.py
+ Metadata:
+ Training Memory (GB): 2.2
+ Epochs: 36
+ inference time (ms/im):
+ - value: 35.0
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (852, 512)
+ Results:
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 32.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/solo/decoupled_solo_light_r50_fpn_3x_coco/decoupled_solo_light_r50_fpn_3x_coco_20210906_142703-e70e226f.pth
+
+ - Name: solo_r50_fpn_3x_coco
+ In Collection: SOLO
+ Config: configs/solo/solo_r50_fpn_3x_coco.py
+ Metadata:
+ Training Memory (GB): 7.4
+ Epochs: 36
+ inference time (ms/im):
+ - value: 94.2
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (1333, 800)
+ Results:
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 35.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/solo/solo_r50_fpn_3x_coco/solo_r50_fpn_3x_coco_20210901_012353-11d224d7.pth
+
+ - Name: solo_r50_fpn_1x_coco
+ In Collection: SOLO
+ Config: configs/solo/solo_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 8.0
+ Epochs: 12
+ inference time (ms/im):
+ - value: 95.1
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (1333, 800)
+ Results:
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 33.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/solo/solo_r50_fpn_1x_coco/solo_r50_fpn_1x_coco_20210821_035055-2290a6b8.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/solo_r101_fpn_8xb8-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/solo_r101_fpn_8xb8-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..0f49c5c1ce67973d15b3fad3ad8c966af8203af7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/solo_r101_fpn_8xb8-lsj-200e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './solo_r50_fpn_8xb8-lsj-200e_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/solo_r18_fpn_8xb8-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/solo_r18_fpn_8xb8-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..977ae54dc28e56802289ac552ce20815b7d1d761
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/solo_r18_fpn_8xb8-lsj-200e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './solo_r50_fpn_8xb8-lsj-200e_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=18,
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet18')),
+ neck=dict(in_channels=[64, 128, 256, 512]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/solo_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/solo_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..595e9ffe148be84dcc3d5c89e5315e8ef3a24477
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/solo_r50_fpn_1x_coco.py
@@ -0,0 +1,62 @@
+_base_ = [
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+# model settings
+model = dict(
+ type='SOLO',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_mask=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50'),
+ style='pytorch'),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=0,
+ num_outs=5),
+ mask_head=dict(
+ type='SOLOHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=7,
+ feat_channels=256,
+ strides=[8, 8, 16, 32, 32],
+ scale_ranges=((1, 96), (48, 192), (96, 384), (192, 768), (384, 2048)),
+ pos_scale=0.2,
+ num_grids=[40, 36, 24, 16, 12],
+ cls_down_index=0,
+ loss_mask=dict(type='DiceLoss', use_sigmoid=True, loss_weight=3.0),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ norm_cfg=dict(type='GN', num_groups=32, requires_grad=True)),
+ # model training and testing settings
+ test_cfg=dict(
+ nms_pre=500,
+ score_thr=0.1,
+ mask_thr=0.5,
+ filter_thr=0.05,
+ kernel='gaussian', # gaussian/linear
+ sigma=2.0,
+ max_per_img=100))
+
+# optimizer
+optim_wrapper = dict(optimizer=dict(lr=0.01))
+
+val_evaluator = dict(metric='segm')
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/solo_r50_fpn_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/solo_r50_fpn_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..0d5abbd2f4d4e1fdc2e3cb92c8e0157188b0aa9a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/solo_r50_fpn_3x_coco.py
@@ -0,0 +1,35 @@
+_base_ = './solo_r50_fpn_1x_coco.py'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 800), (1333, 768), (1333, 736), (1333, 704),
+ (1333, 672), (1333, 640)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+# training schedule for 3x
+max_epochs = 36
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 3,
+ by_epoch=False,
+ begin=0,
+ end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=36,
+ by_epoch=True,
+ milestones=[27, 33],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/solo_r50_fpn_8xb8-lsj-200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/solo_r50_fpn_8xb8-lsj-200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d46bf391c907707d222756e9450b661b6edd6985
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solo/solo_r50_fpn_8xb8-lsj-200e_coco.py
@@ -0,0 +1,71 @@
+_base_ = '../common/lsj-200e_coco-instance.py'
+
+image_size = (1024, 1024)
+batch_augments = [dict(type='BatchFixedSizePad', size=image_size)]
+
+# model settings
+model = dict(
+ type='SOLO',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32,
+ batch_augments=batch_augments),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50'),
+ style='pytorch'),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=0,
+ num_outs=5),
+ mask_head=dict(
+ type='SOLOHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=7,
+ feat_channels=256,
+ strides=[8, 8, 16, 32, 32],
+ scale_ranges=((1, 96), (48, 192), (96, 384), (192, 768), (384, 2048)),
+ pos_scale=0.2,
+ num_grids=[40, 36, 24, 16, 12],
+ cls_down_index=0,
+ loss_mask=dict(type='DiceLoss', use_sigmoid=True, loss_weight=3.0),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ norm_cfg=dict(type='GN', num_groups=32, requires_grad=True)),
+ # model training and testing settings
+ test_cfg=dict(
+ nms_pre=500,
+ score_thr=0.1,
+ mask_thr=0.5,
+ filter_thr=0.05,
+ kernel='gaussian', # gaussian/linear
+ sigma=2.0,
+ max_per_img=100))
+
+train_dataloader = dict(batch_size=8, num_workers=4)
+
+# Enable automatic-mixed-precision training with AmpOptimWrapper.
+optim_wrapper = dict(
+ type='AmpOptimWrapper',
+ optimizer=dict(
+ type='SGD', lr=0.01 * 4, momentum=0.9, weight_decay=0.00004),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..b216913126e7ee86fc474c2cb1cc8b6023e251d1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/README.md
@@ -0,0 +1,59 @@
+# SOLOv2
+
+> [SOLOv2: Dynamic and Fast Instance Segmentation](https://arxiv.org/abs/2003.10152)
+
+
+
+## Abstract
+
+In this work, we aim at building a simple, direct, and fast instance segmentation
+framework with strong performance. We follow the principle of the SOLO method of
+Wang et al. "SOLO: segmenting objects by locations". Importantly, we take one
+step further by dynamically learning the mask head of the object segmenter such
+that the mask head is conditioned on the location. Specifically, the mask branch
+is decoupled into a mask kernel branch and mask feature branch, which are
+responsible for learning the convolution kernel and the convolved features
+respectively. Moreover, we propose Matrix NMS (non maximum suppression) to
+significantly reduce the inference time overhead due to NMS of masks. Our
+Matrix NMS performs NMS with parallel matrix operations in one shot, and
+yields better results. We demonstrate a simple direct instance segmentation
+system, outperforming a few state-of-the-art methods in both speed and accuracy.
+A light-weight version of SOLOv2 executes at 31.3 FPS and yields 37.1% AP.
+Moreover, our state-of-the-art results in object detection (from our mask byproduct)
+and panoptic segmentation show the potential to serve as a new strong baseline
+for many instance-level recognition tasks besides instance segmentation.
+
+
+

+
+
+## Results and Models
+
+### SOLOv2
+
+| Backbone | Style | MS train | Lr schd | Mem (GB) | mask AP | Config | Download |
+| :--------: | :-----: | :------: | :-----: | :------: | :-----: | :-------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | pytorch | N | 1x | 5.1 | 34.8 | [config](./solov2_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_r50_fpn_1x_coco/solov2_r50_fpn_1x_coco_20220512_125858-a357fa23.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_r50_fpn_1x_coco/solov2_r50_fpn_1x_coco_20220512_125858.log.json) |
+| R-50 | pytorch | Y | 3x | 5.1 | 37.5 | [config](./solov2_r50_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_r50_fpn_3x_coco/solov2_r50_fpn_3x_coco_20220512_125856-fed092d4.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_r50_fpn_3x_coco/solov2_r50_fpn_3x_coco_20220512_125856.log.json) |
+| R-101 | pytorch | Y | 3x | 6.9 | 39.1 | [config](./solov2_r101_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_r101_fpn_3x_coco/solov2_r101_fpn_3x_coco_20220511_095119-c559a076.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_r101_fpn_3x_coco/solov2_r101_fpn_3x_coco_20220511_095119.log.json) |
+| R-101(DCN) | pytorch | Y | 3x | 7.1 | 41.2 | [config](./solov2_r101-dcn_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_r101_dcn_fpn_3x_coco/solov2_r101_dcn_fpn_3x_coco_20220513_214734-16c966cb.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_r101_dcn_fpn_3x_coco/solov2_r101_dcn_fpn_3x_coco_20220513_214734.log.json) |
+| X-101(DCN) | pytorch | Y | 3x | 11.3 | 42.4 | [config](./solov2_x101-dcn_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_x101_dcn_fpn_3x_coco/solov2_x101_dcn_fpn_3x_coco_20220513_214337-aef41095.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_x101_dcn_fpn_3x_coco/solov2_x101_dcn_fpn_3x_coco_20220513_214337.log.json) |
+
+### Light SOLOv2
+
+| Backbone | Style | MS train | Lr schd | Mem (GB) | mask AP | Config | Download |
+| :------: | :-----: | :------: | :-----: | :------: | :-----: | :--------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-18 | pytorch | Y | 3x | 9.1 | 29.7 | [config](./solov2-light_r18_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_light_r18_fpn_3x_coco/solov2_light_r18_fpn_3x_coco_20220511_083717-75fa355b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_light_r18_fpn_3x_coco/solov2_light_r18_fpn_3x_coco_20220511_083717.log.json) |
+| R-34 | pytorch | Y | 3x | 9.3 | 31.9 | [config](./solov2-light_r34_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_light_r34_fpn_3x_coco/solov2_light_r34_fpn_3x_coco_20220511_091839-e51659d3.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_light_r34_fpn_3x_coco/solov2_light_r34_fpn_3x_coco_20220511_091839.log.json) |
+| R-50 | pytorch | Y | 3x | 9.9 | 33.7 | [config](./solov2-light_r50_fpn_ms-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_light_r50_fpn_3x_coco/solov2_light_r50_fpn_3x_coco_20220512_165256-c93a6074.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_light_r50_fpn_3x_coco/solov2_light_r50_fpn_3x_coco_20220512_165256.log.json) |
+
+## Citation
+
+```latex
+@article{wang2020solov2,
+ title={SOLOv2: Dynamic and Fast Instance Segmentation},
+ author={Wang, Xinlong and Zhang, Rufeng and Kong, Tao and Li, Lei and Shen, Chunhua},
+ journal={Proc. Advances in Neural Information Processing Systems (NeurIPS)},
+ year={2020}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..d0156b2b40cf62537cdc62af4fa57d644a7978ad
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/metafile.yml
@@ -0,0 +1,93 @@
+Collections:
+ - Name: SOLOv2
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x A100 GPUs
+ Architecture:
+ - FPN
+ - Convolution
+ - ResNet
+ Paper: https://arxiv.org/abs/2003.10152
+ README: configs/solov2/README.md
+
+Models:
+ - Name: solov2_r50_fpn_1x_coco
+ In Collection: SOLOv2
+ Config: configs/solov2/solov2_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 5.1
+ Epochs: 12
+ Results:
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 34.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_r50_fpn_1x_coco/solov2_r50_fpn_1x_coco_20220512_125858-a357fa23.pth
+
+ - Name: solov2_r50_fpn_ms-3x_coco
+ In Collection: SOLOv2
+ Config: configs/solov2/solov2_r50_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 5.1
+ Epochs: 36
+ Results:
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 37.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_r50_fpn_3x_coco/solov2_r50_fpn_3x_coco_20220512_125856-fed092d4.pth
+
+ - Name: solov2_r101-dcn_fpn_ms-3x_coco
+ In Collection: SOLOv2
+ Config: configs/solov2/solov2_r101-dcn_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 7.1
+ Epochs: 36
+ Results:
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 41.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_r101_dcn_fpn_3x_coco/solov2_r101_dcn_fpn_3x_coco_20220513_214734-16c966cb.pth
+
+ - Name: solov2_x101-dcn_fpn_ms-3x_coco
+ In Collection: SOLOv2
+ Config: configs/solov2/solov2_x101-dcn_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 11.3
+ Epochs: 36
+ Results:
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 42.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_x101_dcn_fpn_3x_coco/solov2_x101_dcn_fpn_3x_coco_20220513_214337-aef41095.pth
+
+ - Name: solov2-light_r18_fpn_ms-3x_coco
+ In Collection: SOLOv2
+ Config: configs/solov2/solov2-light_r18_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 9.1
+ Epochs: 36
+ Results:
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 29.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_light_r18_fpn_3x_coco/solov2_light_r18_fpn_3x_coco_20220511_083717-75fa355b.pth
+
+ - Name: solov2-light_r50_fpn_ms-3x_coco
+ In Collection: SOLOv2
+ Config: configs/solov2/solov2-light_r50_fpn_ms-3x_coco.py
+ Metadata:
+ Training Memory (GB): 9.9
+ Epochs: 36
+ Results:
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 33.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/solov2/solov2_light_r50_fpn_3x_coco/solov2_light_r50_fpn_3x_coco_20220512_165256-c93a6074.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2-light_r18_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2-light_r18_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f8fc53e0aed9dd4479f9cd8dcc98ca61db2e50bf
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2-light_r18_fpn_ms-3x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './solov2-light_r50_fpn_ms-3x_coco.py'
+
+# model settings
+model = dict(
+ backbone=dict(
+ depth=18, init_cfg=dict(checkpoint='torchvision://resnet18')),
+ neck=dict(in_channels=[64, 128, 256, 512]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2-light_r34_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2-light_r34_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..149b336655349c70233e78d03f72d7ee3f1a75f3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2-light_r34_fpn_ms-3x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './solov2-light_r50_fpn_ms-3x_coco.py'
+
+# model settings
+model = dict(
+ backbone=dict(
+ depth=34, init_cfg=dict(checkpoint='torchvision://resnet34')),
+ neck=dict(in_channels=[64, 128, 256, 512]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2-light_r50-dcn_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2-light_r50-dcn_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..05391944b683985ab975dc8f66be0c8a12f7d255
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2-light_r50-dcn_fpn_ms-3x_coco.py
@@ -0,0 +1,14 @@
+_base_ = './solov2-light_r50_fpn_ms-3x_coco.py'
+
+# model settings
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCNv2', deformable_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)),
+ mask_head=dict(
+ feat_channels=256,
+ stacked_convs=3,
+ scale_ranges=((1, 64), (32, 128), (64, 256), (128, 512), (256, 2048)),
+ mask_feature_head=dict(out_channels=128),
+ dcn_cfg=dict(type='DCNv2'),
+ dcn_apply_to_all_conv=False)) # light solov2 head
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2-light_r50_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2-light_r50_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..cf0a7f779c0f587d11c86a31aca19b2663f79a57
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2-light_r50_fpn_ms-3x_coco.py
@@ -0,0 +1,56 @@
+_base_ = './solov2_r50_fpn_1x_coco.py'
+
+# model settings
+model = dict(
+ mask_head=dict(
+ stacked_convs=2,
+ feat_channels=256,
+ scale_ranges=((1, 56), (28, 112), (56, 224), (112, 448), (224, 896)),
+ mask_feature_head=dict(out_channels=128)))
+
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(768, 512), (768, 480), (768, 448), (768, 416), (768, 384),
+ (768, 352)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(448, 768), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# training schedule for 3x
+max_epochs = 36
+train_cfg = dict(by_epoch=True, max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 3,
+ by_epoch=False,
+ begin=0,
+ end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=36,
+ by_epoch=True,
+ milestones=[27, 33],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2_r101-dcn_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2_r101-dcn_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..370a4eb7db811b285cc55282e4b66360ca338a31
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2_r101-dcn_fpn_ms-3x_coco.py
@@ -0,0 +1,13 @@
+_base_ = './solov2_r50_fpn_ms-3x_coco.py'
+
+# model settings
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(checkpoint='torchvision://resnet101'),
+ dcn=dict(type='DCNv2', deformable_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)),
+ mask_head=dict(
+ mask_feature_head=dict(conv_cfg=dict(type='DCNv2')),
+ dcn_cfg=dict(type='DCNv2'),
+ dcn_apply_to_all_conv=True))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2_r101_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2_r101_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..96aaac0a7c2689a125ac0a68edaff2a76dfc773d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2_r101_fpn_ms-3x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './solov2_r50_fpn_ms-3x_coco.py'
+
+# model settings
+model = dict(
+ backbone=dict(
+ depth=101, init_cfg=dict(checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..138ca010b5f3f96a4f296ffbe66cb1be3add7ec2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2_r50_fpn_1x_coco.py
@@ -0,0 +1,70 @@
+_base_ = [
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+# model settings
+model = dict(
+ type='SOLOv2',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_mask=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50'),
+ style='pytorch'),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=0,
+ num_outs=5),
+ mask_head=dict(
+ type='SOLOV2Head',
+ num_classes=80,
+ in_channels=256,
+ feat_channels=512,
+ stacked_convs=4,
+ strides=[8, 8, 16, 32, 32],
+ scale_ranges=((1, 96), (48, 192), (96, 384), (192, 768), (384, 2048)),
+ pos_scale=0.2,
+ num_grids=[40, 36, 24, 16, 12],
+ cls_down_index=0,
+ mask_feature_head=dict(
+ feat_channels=128,
+ start_level=0,
+ end_level=3,
+ out_channels=256,
+ mask_stride=4,
+ norm_cfg=dict(type='GN', num_groups=32, requires_grad=True)),
+ loss_mask=dict(type='DiceLoss', use_sigmoid=True, loss_weight=3.0),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0)),
+ # model training and testing settings
+ test_cfg=dict(
+ nms_pre=500,
+ score_thr=0.1,
+ mask_thr=0.5,
+ filter_thr=0.05,
+ kernel='gaussian', # gaussian/linear
+ sigma=2.0,
+ max_per_img=100))
+
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(lr=0.01), clip_grad=dict(max_norm=35, norm_type=2))
+
+val_evaluator = dict(metric='segm')
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2_r50_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2_r50_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d6f09827efbe4e135a784b0808604dbc855ed47e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2_r50_fpn_ms-3x_coco.py
@@ -0,0 +1,35 @@
+_base_ = './solov2_r50_fpn_1x_coco.py'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 800), (1333, 768), (1333, 736), (1333, 704),
+ (1333, 672), (1333, 640)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+# training schedule for 3x
+max_epochs = 36
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 3,
+ by_epoch=False,
+ begin=0,
+ end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=36,
+ by_epoch=True,
+ milestones=[27, 33],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2_x101-dcn_fpn_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2_x101-dcn_fpn_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..612c45eb437efc481948edb660ef1a3eebbcfebe
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/solov2/solov2_x101-dcn_fpn_ms-3x_coco.py
@@ -0,0 +1,17 @@
+_base_ = './solov2_r50_fpn_ms-3x_coco.py'
+
+# model settings
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ dcn=dict(type='DCNv2', deformable_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')),
+ mask_head=dict(
+ mask_feature_head=dict(conv_cfg=dict(type='DCNv2')),
+ dcn_cfg=dict(type='DCNv2'),
+ dcn_apply_to_all_conv=True))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..8f035fded78e53fbe5ee50df8dce7ad97319cc6c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/README.md
@@ -0,0 +1,108 @@
+# Simple online and realtime tracking
+
+## Abstract
+
+
+
+This paper explores a pragmatic approach to multiple object tracking where the main focus is to associate objects efficiently for online and realtime applications. To this end, detection quality is identified as a key factor influencing tracking performance, where changing the detector can improve tracking by up to 18.9%. Despite only using a rudimentary combination of familiar techniques such as the Kalman Filter and Hungarian algorithm for the tracking components, this approach achieves an accuracy comparable to state-of-the-art online trackers. Furthermore, due to the simplicity of our tracking method, the tracker updates at a rate of 260 Hz which is over 20x faster than other state-of-the-art trackers.
+
+
+
+
+

+
+
+## Citation
+
+
+
+```latex
+@inproceedings{bewley2016simple,
+ title={Simple online and realtime tracking},
+ author={Bewley, Alex and Ge, Zongyuan and Ott, Lionel and Ramos, Fabio and Upcroft, Ben},
+ booktitle={2016 IEEE International Conference on Image Processing (ICIP)},
+ pages={3464--3468},
+ year={2016},
+ organization={IEEE}
+}
+```
+
+## Results and models on MOT17
+
+| Method | Detector | ReID | Train Set | Test Set | Public | Inf time (fps) | HOTA | MOTA | IDF1 | FP | FN | IDSw. | Config | Download |
+| :----: | :----------------: | :--: | :--------: | :------: | :----: | :------------: | :--: | :--: | :--: | :---: | :---: | :---: | :----------------------------------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------: |
+| SORT | R50-FasterRCNN-FPN | - | half-train | half-val | N | 18.6 | 52.0 | 62.0 | 57.8 | 15150 | 40410 | 5847 | [config](sort_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py) | [detector](https://download.openmmlab.com/mmtracking/mot/faster_rcnn/faster-rcnn_r50_fpn_4e_mot17-half-64ee2ed4.pth) |
+
+## Get started
+
+### 1. Development Environment Setup
+
+Tracking Development Environment Setup can refer to this [document](../../docs/en/get_started.md).
+
+### 2. Dataset Prepare
+
+Tracking Dataset Prepare can refer to this [document](../../docs/en/user_guides/tracking_dataset_prepare.md).
+
+### 3. Training
+
+We implement SORT with independent detector models.
+Note that, due to the influence of parameters such as learning rate in default configuration file,
+we recommend using 8 GPUs for training in order to reproduce accuracy.
+
+You can train the detector as follows.
+
+```shell script
+# Training Faster R-CNN on mot17-half-train dataset with following command.
+# The number after config file represents the number of GPUs used. Here we use 8 GPUs.
+bash tools/dist_train.sh configs/sort/faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py 8
+```
+
+If you want to know about more detailed usage of `train.py/dist_train.sh/slurm_train.sh`,
+please refer to this [document](../../docs/en/user_guides/tracking_train_test.md).
+
+### 4. Testing and evaluation
+
+### 4.1 Example on MOTxx-halfval dataset
+
+**4.1.1 use separate trained detector model to evaluating and testing**\*
+
+```shell script
+# Example 1: Test on motXX-half-val set.
+# The number after config file represents the number of GPUs used. Here we use 8 GPUs.
+bash tools/dist_test_tracking.sh configs/sort/sort_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py 8 --detector ${DETECTOR_CHECKPOINT_PATH}
+```
+
+**4.1.2 use video_baesd to evaluating and testing**
+
+we also provide two_ways(img_based or video_based) to evaluating and testing.
+if you want to use video_based to evaluating and testing, you can modify config as follows
+
+```
+val_dataloader = dict(
+ sampler=dict(type='DefaultSampler', shuffle=False, round_up=False))
+```
+
+### 4.2 Example on MOTxx-test dataset
+
+If you want to get the results of the [MOT Challenge](https://motchallenge.net/) test set,
+please use the following command to generate result files that can be used for submission.
+It will be stored in `./mot_17_test_res`, you can modify the saved path in `test_evaluator` of the config.
+
+```shell script
+# Example 2: Test on motxx-test set
+# The number after config file represents the number of GPUs used
+bash tools/dist_test_tracking.sh configs/sort/sort_faster-rcnn_r50_fpn_8xb2-4e_mot17train_test-mot17test.py 8 --detector ${DETECTOR_CHECKPOINT_PATH}
+```
+
+If you want to know about more detailed usage of `test_tracking.py/dist_test_tracking.sh/slurm_test_tracking.sh`,
+please refer to this [document](../../docs/en/user_guides/tracking_train_test.md).
+
+### 5.Inference
+
+Use a single GPU to predict a video and save it as a video.
+
+```shell
+python demo/mot_demo.py demo/demo_mot.mp4 configs/sort/sort_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py --detector ${DETECTOR_CHECKPOINT_PATH} --out mot.mp4
+```
+
+If you want to know about more detailed usage of `mot_demo.py`, please refer to this [document](../../docs/en/user_guides/tracking_inference.md).
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py
new file mode 100644
index 0000000000000000000000000000000000000000..f1d5b72ce3fff73504a0c032867d246bc4e30123
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py
@@ -0,0 +1,41 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/mot_challenge_det.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ rpn_head=dict(
+ bbox_coder=dict(clip_border=False),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0 / 9.0, loss_weight=1.0)),
+ roi_head=dict(
+ bbox_head=dict(
+ num_classes=1,
+ bbox_coder=dict(clip_border=False),
+ loss_bbox=dict(type='SmoothL1Loss', loss_weight=1.0))),
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint= # noqa: E251
+ 'http://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/faster_rcnn_r50_fpn_2x_coco/faster_rcnn_r50_fpn_2x_coco_bbox_mAP-0.384_20200504_210434-a5d8aa15.pth' # noqa: E501
+ ))
+
+# training schedule for 4e
+train_cfg = dict(type='EpochBasedTrainLoop', max_epochs=4, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# learning rate
+param_scheduler = [
+ dict(type='LinearLR', start_factor=0.01, by_epoch=False, begin=0, end=100),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=4,
+ by_epoch=True,
+ milestones=[3],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.02, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/faster-rcnn_r50_fpn_8xb2-4e_mot17train_test-mot17train.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/faster-rcnn_r50_fpn_8xb2-4e_mot17train_test-mot17train.py
new file mode 100644
index 0000000000000000000000000000000000000000..83647061c7f59dc8a6e8d033cdb8dc81de648df4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/faster-rcnn_r50_fpn_8xb2-4e_mot17train_test-mot17train.py
@@ -0,0 +1,11 @@
+_base_ = ['./faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval']
+# data
+data_root = 'data/MOT17/'
+train_dataloader = dict(
+ dataset=dict(ann_file='annotations/train_cocoformat.json'))
+val_dataloader = dict(
+ dataset=dict(ann_file='annotations/train_cocoformat.json'))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(ann_file=data_root + 'annotations/train_cocoformat.json')
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/faster-rcnn_r50_fpn_8xb2-8e_mot20halftrain_test-mot20halfval.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/faster-rcnn_r50_fpn_8xb2-8e_mot20halftrain_test-mot20halfval.py
new file mode 100644
index 0000000000000000000000000000000000000000..a6d14ad8be2a939bce168f4f09f08dde50f140c8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/faster-rcnn_r50_fpn_8xb2-8e_mot20halftrain_test-mot20halfval.py
@@ -0,0 +1,29 @@
+_base_ = ['./faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval']
+model = dict(
+ rpn_head=dict(bbox_coder=dict(clip_border=True)),
+ roi_head=dict(
+ bbox_head=dict(bbox_coder=dict(clip_border=True), num_classes=1)))
+# data
+data_root = 'data/MOT20/'
+train_dataloader = dict(dataset=dict(data_root=data_root))
+val_dataloader = dict(dataset=dict(data_root=data_root))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(ann_file=data_root +
+ 'annotations/half-val_cocoformat.json')
+test_evaluator = val_evaluator
+
+# training schedule for 8e
+train_cfg = dict(type='EpochBasedTrainLoop', max_epochs=8, val_interval=1)
+
+# learning rate
+param_scheduler = [
+ dict(type='LinearLR', start_factor=0.01, by_epoch=False, begin=0, end=100),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=8,
+ by_epoch=True,
+ milestones=[6],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/faster-rcnn_r50_fpn_8xb2-8e_mot20train_test-mot20train.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/faster-rcnn_r50_fpn_8xb2-8e_mot20train_test-mot20train.py
new file mode 100644
index 0000000000000000000000000000000000000000..85c859732cb3e4742d3003d555f72f4cc7ac2e05
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/faster-rcnn_r50_fpn_8xb2-8e_mot20train_test-mot20train.py
@@ -0,0 +1,32 @@
+_base_ = ['./faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval']
+model = dict(
+ rpn_head=dict(bbox_coder=dict(clip_border=True)),
+ roi_head=dict(
+ bbox_head=dict(bbox_coder=dict(clip_border=True), num_classes=1)))
+# data
+data_root = 'data/MOT20/'
+train_dataloader = dict(
+ dataset=dict(
+ data_root=data_root, ann_file='annotations/train_cocoformat.json'))
+val_dataloader = dict(
+ dataset=dict(
+ data_root=data_root, ann_file='annotations/train_cocoformat.json'))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(ann_file=data_root + 'annotations/train_cocoformat.json')
+test_evaluator = val_evaluator
+
+# training schedule for 8e
+train_cfg = dict(type='EpochBasedTrainLoop', max_epochs=8, val_interval=1)
+
+# learning rate
+param_scheduler = [
+ dict(type='LinearLR', start_factor=0.01, by_epoch=False, begin=0, end=100),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=8,
+ by_epoch=True,
+ milestones=[6],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..c582ce353df6344aaa2fe25e0f410bb458e50803
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/metafile.yml
@@ -0,0 +1,35 @@
+Collections:
+ - Name: SORT
+ Metadata:
+ Training Techniques:
+ - SGD with Momentum
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNet
+ - FPN
+ Paper:
+ URL: https://arxiv.org/abs/1602.00763
+ Title: Simple Online and Realtime Tracking
+ README: configs/sort/README.md
+
+Models:
+ - Name: sort_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval
+ In Collection: SORT
+ Config: configs/mot/sort/sort_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py
+ Metadata:
+ Training Data: MOT17-half-train
+ inference time (ms/im):
+ - value: 53.8
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (640, 1088)
+ Results:
+ - Task: Multiple Object Tracking
+ Dataset: MOT17-half-val
+ Metrics:
+ MOTA: 62.0
+ IDF1: 57.8
+ HOTA: 52.0
+ Weights: https://download.openmmlab.com/mmtracking/mot/faster_rcnn/faster-rcnn_r50_fpn_4e_mot17-half-64ee2ed4.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/sort_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/sort_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py
new file mode 100644
index 0000000000000000000000000000000000000000..78acb774ec22b7555e633b541c21fe20beb75ce9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/sort_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py
@@ -0,0 +1,54 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py',
+ '../_base_/datasets/mot_challenge.py', '../_base_/default_runtime.py'
+]
+
+default_hooks = dict(
+ logger=dict(type='LoggerHook', interval=1),
+ visualization=dict(type='TrackVisualizationHook', draw=False))
+
+vis_backends = [dict(type='LocalVisBackend')]
+visualizer = dict(
+ type='TrackLocalVisualizer', vis_backends=vis_backends, name='visualizer')
+
+# custom hooks
+custom_hooks = [
+ # Synchronize model buffers such as running_mean and running_var in BN
+ # at the end of each epoch
+ dict(type='SyncBuffersHook')
+]
+
+detector = _base_.model
+detector.pop('data_preprocessor')
+detector.rpn_head.bbox_coder.update(dict(clip_border=False))
+detector.roi_head.bbox_head.update(dict(num_classes=1))
+detector.roi_head.bbox_head.bbox_coder.update(dict(clip_border=False))
+detector['init_cfg'] = dict(
+ type='Pretrained',
+ checkpoint= # noqa: E251
+ 'https://download.openmmlab.com/mmtracking/mot/'
+ 'faster_rcnn/faster-rcnn_r50_fpn_4e_mot17-half-64ee2ed4.pth') # noqa: E501
+del _base_.model
+
+model = dict(
+ type='DeepSORT',
+ data_preprocessor=dict(
+ type='TrackDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ rgb_to_bgr=False,
+ pad_size_divisor=32),
+ detector=detector,
+ tracker=dict(
+ type='SORTTracker',
+ motion=dict(type='KalmanFilter', center_only=False),
+ obj_score_thr=0.5,
+ match_iou_thr=0.5,
+ reid=None))
+
+train_dataloader = None
+
+train_cfg = None
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/sort_faster-rcnn_r50_fpn_8xb2-4e_mot17train_test-mot17test.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/sort_faster-rcnn_r50_fpn_8xb2-4e_mot17train_test-mot17test.py
new file mode 100644
index 0000000000000000000000000000000000000000..921652c4430ccf63cd5850884b2a064e8dc73251
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sort/sort_faster-rcnn_r50_fpn_8xb2-4e_mot17train_test-mot17test.py
@@ -0,0 +1,15 @@
+_base_ = [
+ './sort_faster-rcnn_r50_fpn_8xb2-4e_mot17halftrain'
+ '_test-mot17halfval.py'
+]
+
+# dataloader
+val_dataloader = dict(
+ dataset=dict(ann_file='annotations/train_cocoformat.json'))
+test_dataloader = dict(
+ dataset=dict(
+ ann_file='annotations/test_cocoformat.json',
+ data_prefix=dict(img_path='test')))
+
+# evaluator
+test_evaluator = dict(format_only=True, outfile_prefix='./mot_17_test_res')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..2e8e365b3df2476bb2d8f9acfe76f24fcf7756ea
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/README.md
@@ -0,0 +1,38 @@
+# Sparse R-CNN
+
+> [Sparse R-CNN: End-to-End Object Detection with Learnable Proposals](https://arxiv.org/abs/2011.12450)
+
+
+
+## Abstract
+
+We present Sparse R-CNN, a purely sparse method for object detection in images. Existing works on object detection heavily rely on dense object candidates, such as k anchor boxes pre-defined on all grids of image feature map of size H×W. In our method, however, a fixed sparse set of learned object proposals, total length of N, are provided to object recognition head to perform classification and location. By eliminating HWk (up to hundreds of thousands) hand-designed object candidates to N (e.g. 100) learnable proposals, Sparse R-CNN completely avoids all efforts related to object candidates design and many-to-one label assignment. More importantly, final predictions are directly output without non-maximum suppression post-procedure. Sparse R-CNN demonstrates accuracy, run-time and training convergence performance on par with the well-established detector baselines on the challenging COCO dataset, e.g., achieving 45.0 AP in standard 3× training schedule and running at 22 fps using ResNet-50 FPN model. We hope our work could inspire re-thinking the convention of dense prior in object detectors.
+
+
+

+
+
+## Results and Models
+
+| Model | Backbone | Style | Lr schd | Number of Proposals | Multi-Scale | RandomCrop | box AP | Config | Download |
+| :----------: | :-------: | :-----: | :-----: | :-----------------: | :---------: | :--------: | :----: | :-----------------------------------------------------------------------: | :-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Sparse R-CNN | R-50-FPN | pytorch | 1x | 100 | False | False | 37.9 | [config](./sparse-rcnn_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/sparse_rcnn/sparse_rcnn_r50_fpn_1x_coco/sparse_rcnn_r50_fpn_1x_coco_20201222_214453-dc79b137.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/sparse_rcnn/sparse_rcnn_r50_fpn_1x_coco/sparse_rcnn_r50_fpn_1x_coco_20201222_214453-dc79b137.log.json) |
+| Sparse R-CNN | R-50-FPN | pytorch | 3x | 100 | True | False | 42.8 | [config](./sparse-rcnn_r50_fpn_ms-480-800-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/sparse_rcnn/sparse_rcnn_r50_fpn_mstrain_480-800_3x_coco/sparse_rcnn_r50_fpn_mstrain_480-800_3x_coco_20201218_154234-7bc5c054.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/sparse_rcnn/sparse_rcnn_r50_fpn_mstrain_480-800_3x_coco/sparse_rcnn_r50_fpn_mstrain_480-800_3x_coco_20201218_154234-7bc5c054.log.json) |
+| Sparse R-CNN | R-50-FPN | pytorch | 3x | 300 | True | True | 45.0 | [config](./sparse-rcnn_r50_fpn_300-proposals_crop-ms-480-800-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/sparse_rcnn/sparse_rcnn_r50_fpn_300_proposals_crop_mstrain_480-800_3x_coco/sparse_rcnn_r50_fpn_300_proposals_crop_mstrain_480-800_3x_coco_20201223_024605-9fe92701.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/sparse_rcnn/sparse_rcnn_r50_fpn_300_proposals_crop_mstrain_480-800_3x_coco/sparse_rcnn_r50_fpn_300_proposals_crop_mstrain_480-800_3x_coco_20201223_024605-9fe92701.log.json) |
+| Sparse R-CNN | R-101-FPN | pytorch | 3x | 100 | True | False | 44.2 | [config](./sparse-rcnn_r101_fpn_ms-480-800-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/sparse_rcnn/sparse_rcnn_r101_fpn_mstrain_480-800_3x_coco/sparse_rcnn_r101_fpn_mstrain_480-800_3x_coco_20201223_121552-6c46c9d6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/sparse_rcnn/sparse_rcnn_r101_fpn_mstrain_480-800_3x_coco/sparse_rcnn_r101_fpn_mstrain_480-800_3x_coco_20201223_121552-6c46c9d6.log.json) |
+| Sparse R-CNN | R-101-FPN | pytorch | 3x | 300 | True | True | 46.2 | [config](./sparse-rcnn_r101_fpn_300-proposals_crop-ms-480-800-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/sparse_rcnn/sparse_rcnn_r101_fpn_300_proposals_crop_mstrain_480-800_3x_coco/sparse_rcnn_r101_fpn_300_proposals_crop_mstrain_480-800_3x_coco_20201223_023452-c23c3564.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/sparse_rcnn/sparse_rcnn_r101_fpn_300_proposals_crop_mstrain_480-800_3x_coco/sparse_rcnn_r101_fpn_300_proposals_crop_mstrain_480-800_3x_coco_20201223_023452-c23c3564.log.json) |
+
+### Notes
+
+We observe about 0.3 AP noise especially when using ResNet-101 as the backbone.
+
+## Citation
+
+```latex
+@article{peize2020sparse,
+ title = {{SparseR-CNN}: End-to-End Object Detection with Learnable Proposals},
+ author = {Peize Sun and Rufeng Zhang and Yi Jiang and Tao Kong and Chenfeng Xu and Wei Zhan and Masayoshi Tomizuka and Lei Li and Zehuan Yuan and Changhu Wang and Ping Luo},
+ journal = {arXiv preprint arXiv:2011.12450},
+ year = {2020}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..8fe2531893b99662bd9e5dbbc1d6f9a6ced00325
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/metafile.yml
@@ -0,0 +1,80 @@
+Collections:
+ - Name: Sparse R-CNN
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - FPN
+ - ResNet
+ - Sparse R-CNN
+ Paper:
+ URL: https://arxiv.org/abs/2011.12450
+ Title: 'Sparse R-CNN: End-to-End Object Detection with Learnable Proposals'
+ README: configs/sparse_rcnn/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.9.0/mmdet/models/detectors/sparse_rcnn.py#L6
+ Version: v2.9.0
+
+Models:
+ - Name: sparse-rcnn_r50_fpn_1x_coco
+ In Collection: Sparse R-CNN
+ Config: configs/sparse_rcnn/sparse-rcnn_r50_fpn_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/sparse_rcnn/sparse_rcnn_r50_fpn_1x_coco/sparse_rcnn_r50_fpn_1x_coco_20201222_214453-dc79b137.pth
+
+ - Name: sparse-rcnn_r50_fpn_ms-480-800-3x_coco
+ In Collection: Sparse R-CNN
+ Config: configs/sparse_rcnn/sparse-rcnn_r50_fpn_ms-480-800-3x_coco.py
+ Metadata:
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/sparse_rcnn/sparse_rcnn_r50_fpn_mstrain_480-800_3x_coco/sparse_rcnn_r50_fpn_mstrain_480-800_3x_coco_20201218_154234-7bc5c054.pth
+
+ - Name: sparse-rcnn_r50_fpn_300-proposals_crop-ms-480-800-3x_coco
+ In Collection: Sparse R-CNN
+ Config: configs/sparse_rcnn/sparse-rcnn_r50_fpn_300-proposals_crop-ms-480-800-3x_coco.py
+ Metadata:
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 45.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/sparse_rcnn/sparse_rcnn_r50_fpn_300_proposals_crop_mstrain_480-800_3x_coco/sparse_rcnn_r50_fpn_300_proposals_crop_mstrain_480-800_3x_coco_20201223_024605-9fe92701.pth
+
+ - Name: sparse-rcnn_r101_fpn_ms-480-800-3x_coco
+ In Collection: Sparse R-CNN
+ Config: configs/sparse_rcnn/sparse-rcnn_r101_fpn_ms-480-800-3x_coco.py
+ Metadata:
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/sparse_rcnn/sparse_rcnn_r101_fpn_mstrain_480-800_3x_coco/sparse_rcnn_r101_fpn_mstrain_480-800_3x_coco_20201223_121552-6c46c9d6.pth
+
+ - Name: sparse-rcnn_r101_fpn_300-proposals_crop-ms-480-800-3x_coco
+ In Collection: Sparse R-CNN
+ Config: configs/sparse_rcnn/sparse-rcnn_r101_fpn_300-proposals_crop-ms-480-800-3x_coco.py
+ Metadata:
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/sparse_rcnn/sparse_rcnn_r101_fpn_300_proposals_crop_mstrain_480-800_3x_coco/sparse_rcnn_r101_fpn_300_proposals_crop_mstrain_480-800_3x_coco_20201223_023452-c23c3564.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/sparse-rcnn_r101_fpn_300-proposals_crop-ms-480-800-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/sparse-rcnn_r101_fpn_300-proposals_crop-ms-480-800-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..09c11c6565ea2444fe8ffc930ca49fbffff3e8fa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/sparse-rcnn_r101_fpn_300-proposals_crop-ms-480-800-3x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './sparse-rcnn_r50_fpn_300-proposals_crop-ms-480-800-3x_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/sparse-rcnn_r101_fpn_ms-480-800-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/sparse-rcnn_r101_fpn_ms-480-800-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a51f11ce5b6d55b2037461a93aa2bd18c8f2639d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/sparse-rcnn_r101_fpn_ms-480-800-3x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './sparse-rcnn_r50_fpn_ms-480-800-3x_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/sparse-rcnn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/sparse-rcnn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..88354427b4138f4f5587f2a4a047bad654693780
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/sparse-rcnn_r50_fpn_1x_coco.py
@@ -0,0 +1,101 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+num_stages = 6
+num_proposals = 100
+model = dict(
+ type='SparseRCNN',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=0,
+ add_extra_convs='on_input',
+ num_outs=4),
+ rpn_head=dict(
+ type='EmbeddingRPNHead',
+ num_proposals=num_proposals,
+ proposal_feature_channel=256),
+ roi_head=dict(
+ type='SparseRoIHead',
+ num_stages=num_stages,
+ stage_loss_weights=[1] * num_stages,
+ proposal_feature_channel=256,
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=7, sampling_ratio=2),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ bbox_head=[
+ dict(
+ type='DIIHead',
+ num_classes=80,
+ num_ffn_fcs=2,
+ num_heads=8,
+ num_cls_fcs=1,
+ num_reg_fcs=3,
+ feedforward_channels=2048,
+ in_channels=256,
+ dropout=0.0,
+ ffn_act_cfg=dict(type='ReLU', inplace=True),
+ dynamic_conv_cfg=dict(
+ type='DynamicConv',
+ in_channels=256,
+ feat_channels=64,
+ out_channels=256,
+ input_feat_shape=7,
+ act_cfg=dict(type='ReLU', inplace=True),
+ norm_cfg=dict(type='LN')),
+ loss_bbox=dict(type='L1Loss', loss_weight=5.0),
+ loss_iou=dict(type='GIoULoss', loss_weight=2.0),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=2.0),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ clip_border=False,
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.5, 0.5, 1., 1.])) for _ in range(num_stages)
+ ]),
+ # training and testing settings
+ train_cfg=dict(
+ rpn=None,
+ rcnn=[
+ dict(
+ assigner=dict(
+ type='HungarianAssigner',
+ match_costs=[
+ dict(type='FocalLossCost', weight=2.0),
+ dict(type='BBoxL1Cost', weight=5.0, box_format='xyxy'),
+ dict(type='IoUCost', iou_mode='giou', weight=2.0)
+ ]),
+ sampler=dict(type='PseudoSampler'),
+ pos_weight=1) for _ in range(num_stages)
+ ]),
+ test_cfg=dict(rpn=None, rcnn=dict(max_per_img=num_proposals)))
+
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(
+ _delete_=True, type='AdamW', lr=0.000025, weight_decay=0.0001),
+ clip_grad=dict(max_norm=1, norm_type=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/sparse-rcnn_r50_fpn_300-proposals_crop-ms-480-800-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/sparse-rcnn_r50_fpn_300-proposals_crop-ms-480-800-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..93edc0314b510c635f703f82e39c446ed056c6ea
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/sparse-rcnn_r50_fpn_300-proposals_crop-ms-480-800-3x_coco.py
@@ -0,0 +1,43 @@
+_base_ = './sparse-rcnn_r50_fpn_ms-480-800-3x_coco.py'
+num_proposals = 300
+model = dict(
+ rpn_head=dict(num_proposals=num_proposals),
+ test_cfg=dict(
+ _delete_=True, rpn=None, rcnn=dict(max_per_img=num_proposals)))
+
+# augmentation strategy originates from DETR.
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[[
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(400, 1333), (500, 1333), (600, 1333)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333),
+ (576, 1333), (608, 1333), (640, 1333),
+ (672, 1333), (704, 1333), (736, 1333),
+ (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]]),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/sparse-rcnn_r50_fpn_ms-480-800-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/sparse-rcnn_r50_fpn_ms-480-800-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..156028d7cdd22c32c00a765c6cf86b8f9e2df48b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/sparse_rcnn/sparse-rcnn_r50_fpn_ms-480-800-3x_coco.py
@@ -0,0 +1,32 @@
+_base_ = './sparse-rcnn_r50_fpn_1x_coco.py'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+# learning policy
+max_epochs = 36
+train_cfg = dict(type='EpochBasedTrainLoop', max_epochs=max_epochs)
+
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[27, 33],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ssd/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ssd/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..8b3ca9128fd483841eaac6943e9fac68a116eb25
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ssd/README.md
@@ -0,0 +1,62 @@
+# SSD
+
+> [SSD: Single Shot MultiBox Detector](https://arxiv.org/abs/1512.02325)
+
+
+
+## Abstract
+
+We present a method for detecting objects in images using a single deep neural network. Our approach, named SSD, discretizes the output space of bounding boxes into a set of default boxes over different aspect ratios and scales per feature map location. At prediction time, the network generates scores for the presence of each object category in each default box and produces adjustments to the box to better match the object shape. Additionally, the network combines predictions from multiple feature maps with different resolutions to naturally handle objects of various sizes. Our SSD model is simple relative to methods that require object proposals because it completely eliminates proposal generation and subsequent pixel or feature resampling stage and encapsulates all computation in a single network. This makes SSD easy to train and straightforward to integrate into systems that require a detection component. Experimental results on the PASCAL VOC, MS COCO, and ILSVRC datasets confirm that SSD has comparable accuracy to methods that utilize an additional object proposal step and is much faster, while providing a unified framework for both training and inference. Compared to other single stage methods, SSD has much better accuracy, even with a smaller input image size. For 300×300 input, SSD achieves 72.1% mAP on VOC2007 test at 58 FPS on a Nvidia Titan X and for 500×500 input, SSD achieves 75.1% mAP, outperforming a comparable state of the art Faster R-CNN model.
+
+
+

+
+
+## Results and models of SSD
+
+| Backbone | Size | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :------: | :--: | :---: | :-----: | :------: | :------------: | :----: | :------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| VGG16 | 300 | caffe | 120e | 9.9 | 43.7 | 25.5 | [config](./ssd300_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/ssd/ssd300_coco/ssd300_coco_20210803_015428-d231a06e.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/ssd/ssd300_coco/ssd300_coco_20210803_015428.log.json) |
+| VGG16 | 512 | caffe | 120e | 19.4 | 30.7 | 29.5 | [config](./ssd512_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/ssd/ssd512_coco/ssd512_coco_20210803_022849-0a47a1ca.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/ssd/ssd512_coco/ssd512_coco_20210803_022849.log.json) |
+
+## Results and models of SSD-Lite
+
+| Backbone | Size | Training from scratch | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :---------: | :--: | :-------------------: | :-----: | :------: | :------------: | :----: | :--------------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| MobileNetV2 | 320 | yes | 600e | 4.0 | 69.9 | 21.3 | [config](./ssdlite_mobilenetv2-scratch_8xb24-600e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/ssd/ssdlite_mobilenetv2_scratch_600e_coco/ssdlite_mobilenetv2_scratch_600e_coco_20210629_110627-974d9307.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/ssd/ssdlite_mobilenetv2_scratch_600e_coco/ssdlite_mobilenetv2_scratch_600e_coco_20210629_110627.log.json) |
+
+## Notice
+
+### Compatibility
+
+In v2.14.0, [PR5291](https://github.com/open-mmlab/mmdetection/pull/5291) refactored SSD neck and head for more
+flexible usage. If users want to use the SSD checkpoint trained in the older versions, we provide a scripts
+`tools/model_converters/upgrade_ssd_version.py` to convert the model weights.
+
+```bash
+python tools/model_converters/upgrade_ssd_version.py ${OLD_MODEL_PATH} ${NEW_MODEL_PATH}
+
+```
+
+- OLD_MODEL_PATH: the path to load the old version SSD model.
+- NEW_MODEL_PATH: the path to save the converted model weights.
+
+### SSD-Lite training settings
+
+There are some differences between our implementation of MobileNetV2 SSD-Lite and the one in [TensorFlow 1.x detection model zoo](https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/tf1_detection_zoo.md) .
+
+1. Use 320x320 as input size instead of 300x300.
+2. The anchor sizes are different.
+3. The C4 feature map is taken from the last layer of stage 4 instead of the middle of the block.
+4. The model in TensorFlow1.x is trained on coco 2014 and validated on coco minival2014, but we trained and validated the model on coco 2017. The mAP on val2017 is usually a little lower than minival2014 (refer to the results in TensorFlow Object Detection API, e.g., MobileNetV2 SSD gets 22 mAP on minival2014 but 20.2 mAP on val2017).
+
+## Citation
+
+```latex
+@article{Liu_2016,
+ title={SSD: Single Shot MultiBox Detector},
+ journal={ECCV},
+ author={Liu, Wei and Anguelov, Dragomir and Erhan, Dumitru and Szegedy, Christian and Reed, Scott and Fu, Cheng-Yang and Berg, Alexander C.},
+ year={2016},
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ssd/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ssd/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..190a207ccc9b62a002d026f917d66778e5cee8b7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ssd/metafile.yml
@@ -0,0 +1,78 @@
+Collections:
+ - Name: SSD
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - VGG
+ Paper:
+ URL: https://arxiv.org/abs/1512.02325
+ Title: 'SSD: Single Shot MultiBox Detector'
+ README: configs/ssd/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.14.0/mmdet/models/dense_heads/ssd_head.py#L16
+ Version: v2.14.0
+
+Models:
+ - Name: ssd300_coco
+ In Collection: SSD
+ Config: configs/ssd/ssd300_coco.py
+ Metadata:
+ Training Memory (GB): 9.9
+ inference time (ms/im):
+ - value: 22.88
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (300, 300)
+ Epochs: 120
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 25.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/ssd/ssd300_coco/ssd300_coco_20210803_015428-d231a06e.pth
+
+ - Name: ssd512_coco
+ In Collection: SSD
+ Config: configs/ssd/ssd512_coco.py
+ Metadata:
+ Training Memory (GB): 19.4
+ inference time (ms/im):
+ - value: 32.57
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (512, 512)
+ Epochs: 120
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 29.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/ssd/ssd512_coco/ssd512_coco_20210803_022849-0a47a1ca.pth
+
+ - Name: ssdlite_mobilenetv2-scratch_8xb24-600e_coco
+ In Collection: SSD
+ Config: configs/ssd/ssdlite_mobilenetv2-scratch_8xb24-600e_coco.py
+ Metadata:
+ Training Memory (GB): 4.0
+ inference time (ms/im):
+ - value: 14.3
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (320, 320)
+ Epochs: 600
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 21.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/ssd/ssdlite_mobilenetv2_scratch_600e_coco/ssdlite_mobilenetv2_scratch_600e_coco_20210629_110627-974d9307.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ssd/ssd300_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ssd/ssd300_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..796d25c905350a8ed263b9cd1d2f8027b8c9a3ca
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ssd/ssd300_coco.py
@@ -0,0 +1,71 @@
+_base_ = [
+ '../_base_/models/ssd300.py', '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_2x.py', '../_base_/default_runtime.py'
+]
+
+# dataset settings
+input_size = 300
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='Expand',
+ mean={{_base_.model.data_preprocessor.mean}},
+ to_rgb={{_base_.model.data_preprocessor.bgr_to_rgb}},
+ ratio_range=(1, 4)),
+ dict(
+ type='MinIoURandomCrop',
+ min_ious=(0.1, 0.3, 0.5, 0.7, 0.9),
+ min_crop_size=0.3),
+ dict(type='Resize', scale=(input_size, input_size), keep_ratio=False),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='PhotoMetricDistortion',
+ brightness_delta=32,
+ contrast_range=(0.5, 1.5),
+ saturation_range=(0.5, 1.5),
+ hue_delta=18),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(input_size, input_size), keep_ratio=False),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=8,
+ num_workers=2,
+ batch_sampler=None,
+ dataset=dict(
+ _delete_=True,
+ type='RepeatDataset',
+ times=5,
+ dataset=dict(
+ type={{_base_.dataset_type}},
+ data_root={{_base_.data_root}},
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args={{_base_.backend_args}})))
+val_dataloader = dict(batch_size=8, dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=2e-3, momentum=0.9, weight_decay=5e-4))
+
+custom_hooks = [
+ dict(type='NumClassCheckHook'),
+ dict(type='CheckInvalidLossHook', interval=50, priority='VERY_LOW')
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ssd/ssd512_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ssd/ssd512_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..7acd6144202e8fee232e3ed49a557d3cf7c53e15
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ssd/ssd512_coco.py
@@ -0,0 +1,60 @@
+_base_ = 'ssd300_coco.py'
+
+# model settings
+input_size = 512
+model = dict(
+ neck=dict(
+ out_channels=(512, 1024, 512, 256, 256, 256, 256),
+ level_strides=(2, 2, 2, 2, 1),
+ level_paddings=(1, 1, 1, 1, 1),
+ last_kernel_size=4),
+ bbox_head=dict(
+ in_channels=(512, 1024, 512, 256, 256, 256, 256),
+ anchor_generator=dict(
+ type='SSDAnchorGenerator',
+ scale_major=False,
+ input_size=input_size,
+ basesize_ratio_range=(0.1, 0.9),
+ strides=[8, 16, 32, 64, 128, 256, 512],
+ ratios=[[2], [2, 3], [2, 3], [2, 3], [2, 3], [2], [2]])))
+
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='Expand',
+ mean={{_base_.model.data_preprocessor.mean}},
+ to_rgb={{_base_.model.data_preprocessor.bgr_to_rgb}},
+ ratio_range=(1, 4)),
+ dict(
+ type='MinIoURandomCrop',
+ min_ious=(0.1, 0.3, 0.5, 0.7, 0.9),
+ min_crop_size=0.3),
+ dict(type='Resize', scale=(input_size, input_size), keep_ratio=False),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='PhotoMetricDistortion',
+ brightness_delta=32,
+ contrast_range=(0.5, 1.5),
+ saturation_range=(0.5, 1.5),
+ hue_delta=18),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(input_size, input_size), keep_ratio=False),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(dataset=dict(dataset=dict(pipeline=train_pipeline)))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/ssd/ssdlite_mobilenetv2-scratch_8xb24-600e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ssd/ssdlite_mobilenetv2-scratch_8xb24-600e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..4e508f20ecf33e58ddfe6ff8ee94f516d3e03f79
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/ssd/ssdlite_mobilenetv2-scratch_8xb24-600e_coco.py
@@ -0,0 +1,158 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+# model settings
+data_preprocessor = dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=1)
+model = dict(
+ type='SingleStageDetector',
+ data_preprocessor=data_preprocessor,
+ backbone=dict(
+ type='MobileNetV2',
+ out_indices=(4, 7),
+ norm_cfg=dict(type='BN', eps=0.001, momentum=0.03),
+ init_cfg=dict(type='TruncNormal', layer='Conv2d', std=0.03)),
+ neck=dict(
+ type='SSDNeck',
+ in_channels=(96, 1280),
+ out_channels=(96, 1280, 512, 256, 256, 128),
+ level_strides=(2, 2, 2, 2),
+ level_paddings=(1, 1, 1, 1),
+ l2_norm_scale=None,
+ use_depthwise=True,
+ norm_cfg=dict(type='BN', eps=0.001, momentum=0.03),
+ act_cfg=dict(type='ReLU6'),
+ init_cfg=dict(type='TruncNormal', layer='Conv2d', std=0.03)),
+ bbox_head=dict(
+ type='SSDHead',
+ in_channels=(96, 1280, 512, 256, 256, 128),
+ num_classes=80,
+ use_depthwise=True,
+ norm_cfg=dict(type='BN', eps=0.001, momentum=0.03),
+ act_cfg=dict(type='ReLU6'),
+ init_cfg=dict(type='Normal', layer='Conv2d', std=0.001),
+
+ # set anchor size manually instead of using the predefined
+ # SSD300 setting.
+ anchor_generator=dict(
+ type='SSDAnchorGenerator',
+ scale_major=False,
+ strides=[16, 32, 64, 107, 160, 320],
+ ratios=[[2, 3], [2, 3], [2, 3], [2, 3], [2, 3], [2, 3]],
+ min_sizes=[48, 100, 150, 202, 253, 304],
+ max_sizes=[100, 150, 202, 253, 304, 320]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2])),
+ # model training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.,
+ ignore_iof_thr=-1,
+ gt_max_assign_all=False),
+ sampler=dict(type='PseudoSampler'),
+ smoothl1_beta=1.,
+ allowed_border=-1,
+ pos_weight=-1,
+ neg_pos_ratio=3,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ nms=dict(type='nms', iou_threshold=0.45),
+ min_bbox_size=0,
+ score_thr=0.02,
+ max_per_img=200))
+env_cfg = dict(cudnn_benchmark=True)
+
+# dataset settings
+input_size = 320
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='Expand',
+ mean=data_preprocessor['mean'],
+ to_rgb=data_preprocessor['bgr_to_rgb'],
+ ratio_range=(1, 4)),
+ dict(
+ type='MinIoURandomCrop',
+ min_ious=(0.1, 0.3, 0.5, 0.7, 0.9),
+ min_crop_size=0.3),
+ dict(type='Resize', scale=(input_size, input_size), keep_ratio=False),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='PhotoMetricDistortion',
+ brightness_delta=32,
+ contrast_range=(0.5, 1.5),
+ saturation_range=(0.5, 1.5),
+ hue_delta=18),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='Resize', scale=(input_size, input_size), keep_ratio=False),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=24,
+ num_workers=4,
+ batch_sampler=None,
+ dataset=dict(
+ _delete_=True,
+ type='RepeatDataset',
+ times=5,
+ dataset=dict(
+ type={{_base_.dataset_type}},
+ data_root={{_base_.data_root}},
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline)))
+val_dataloader = dict(batch_size=8, dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# training schedule
+max_epochs = 120
+train_cfg = dict(max_epochs=max_epochs, val_interval=5)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='CosineAnnealingLR',
+ begin=0,
+ T_max=max_epochs,
+ end=max_epochs,
+ by_epoch=True,
+ eta_min=0)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.015, momentum=0.9, weight_decay=4.0e-5))
+
+custom_hooks = [
+ dict(type='NumClassCheckHook'),
+ dict(type='CheckInvalidLossHook', interval=50, priority='VERY_LOW')
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (24 samples per GPU)
+auto_scale_lr = dict(base_batch_size=192)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..e5db3e08e0774060913382b5b25cfe515bd7ead5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/README.md
@@ -0,0 +1,20 @@
+# Strong Baselines
+
+
+
+We train Mask R-CNN with large-scale jitter and longer schedule as strong baselines.
+The modifications follow those in [Detectron2](https://github.com/facebookresearch/detectron2/tree/master/configs/new_baselines).
+
+## Results and Models
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :------: | :-----: | :-----: | :------: | :------------: | :----: | :-----: | :--------------------------------------------------------------------------------: | :----------------------: |
+| R-50-FPN | pytorch | 50e | | | | | [config](./mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-50e_coco.py) | [model](<>) \| [log](<>) |
+| R-50-FPN | pytorch | 100e | | | | | [config](./mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-100e_coco.py) | [model](<>) \| [log](<>) |
+| R-50-FPN | caffe | 100e | | | 44.7 | 40.4 | [config](./mask-rcnn_r50-caffe_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-100e_coco.py) | [model](<>) \| [log](<>) |
+| R-50-FPN | caffe | 400e | | | | | [config](./mask-rcnn_r50-caffe_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-400e_coco.py) | [model](<>) \| [log](<>) |
+
+## Notice
+
+When using large-scale jittering, there are sometimes empty proposals in the box and mask heads during training.
+This requires MMSyncBN that allows empty tensors. Therefore, please use mmcv-full>=1.3.14 to train models supported in this directory.
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/mask-rcnn_r50-caffe_fpn_rpn-2conv_4conv1fc_syncbn-all_amp-lsj-100e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/mask-rcnn_r50-caffe_fpn_rpn-2conv_4conv1fc_syncbn-all_amp-lsj-100e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b004d740a8f1e303bc4ad32593baad021ccae710
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/mask-rcnn_r50-caffe_fpn_rpn-2conv_4conv1fc_syncbn-all_amp-lsj-100e_coco.py
@@ -0,0 +1,4 @@
+_base_ = 'mask-rcnn_r50-caffe_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-100e_coco.py' # noqa
+
+# Enable automatic-mixed-precision training with AmpOptimWrapper.
+optim_wrapper = dict(type='AmpOptimWrapper')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/mask-rcnn_r50-caffe_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-100e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/mask-rcnn_r50-caffe_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-100e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..70e92a82e0cd1f083fbb87035f61877da4c11022
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/mask-rcnn_r50-caffe_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-100e_coco.py
@@ -0,0 +1,68 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../common/lsj-100e_coco-instance.py'
+]
+image_size = (1024, 1024)
+batch_augments = [
+ dict(type='BatchFixedSizePad', size=image_size, pad_mask=True)
+]
+norm_cfg = dict(type='SyncBN', requires_grad=True)
+# Use MMSyncBN that handles empty tensor in head. It can be changed to
+# SyncBN after https://github.com/pytorch/pytorch/issues/36530 is fixed
+head_norm_cfg = dict(type='MMSyncBN', requires_grad=True)
+model = dict(
+ # use caffe norm
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+
+ # pad_size_divisor=32 is unnecessary in training but necessary
+ # in testing.
+ pad_size_divisor=32,
+ batch_augments=batch_augments),
+ backbone=dict(
+ frozen_stages=-1,
+ norm_eval=False,
+ norm_cfg=norm_cfg,
+ init_cfg=None,
+ style='caffe'),
+ neck=dict(norm_cfg=norm_cfg),
+ rpn_head=dict(num_convs=2),
+ roi_head=dict(
+ bbox_head=dict(
+ type='Shared4Conv1FCBBoxHead',
+ conv_out_channels=256,
+ norm_cfg=head_norm_cfg),
+ mask_head=dict(norm_cfg=head_norm_cfg)))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomResize',
+ scale=image_size,
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size,
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+# Use RepeatDataset to speed up training
+train_dataloader = dict(dataset=dict(dataset=dict(pipeline=train_pipeline)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/mask-rcnn_r50-caffe_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-400e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/mask-rcnn_r50-caffe_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-400e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..cb64c9b6865634412c8b9d951b588cf0fb8cd32b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/mask-rcnn_r50-caffe_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-400e_coco.py
@@ -0,0 +1,20 @@
+_base_ = './mask-rcnn_r50-caffe_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-100e_coco.py' # noqa
+
+# Use RepeatDataset to speed up training
+# change repeat time from 4 (for 100 epochs) to 16 (for 400 epochs)
+train_dataloader = dict(dataset=dict(times=4 * 4))
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=0.067,
+ by_epoch=False,
+ begin=0,
+ end=500 * 4),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[22, 24],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_amp-lsj-100e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_amp-lsj-100e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..7fab2c72114cbe8a4d6cd3bdddb4e7c3b8dc2d0c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_amp-lsj-100e_coco.py
@@ -0,0 +1,4 @@
+_base_ = 'mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-100e_coco.py'
+
+# Enable automatic-mixed-precision training with AmpOptimWrapper.
+optim_wrapper = dict(type='AmpOptimWrapper')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-100e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-100e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8e06587fb03d42958142cac9ce7b15e7a19a9f6d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-100e_coco.py
@@ -0,0 +1,30 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../common/lsj-100e_coco-instance.py'
+]
+
+image_size = (1024, 1024)
+batch_augments = [
+ dict(type='BatchFixedSizePad', size=image_size, pad_mask=True)
+]
+norm_cfg = dict(type='SyncBN', requires_grad=True)
+# Use MMSyncBN that handles empty tensor in head. It can be changed to
+# SyncBN after https://github.com/pytorch/pytorch/issues/36530 is fixed
+head_norm_cfg = dict(type='MMSyncBN', requires_grad=True)
+model = dict(
+ # the model is trained from scratch, so init_cfg is None
+ data_preprocessor=dict(
+ # pad_size_divisor=32 is unnecessary in training but necessary
+ # in testing.
+ pad_size_divisor=32,
+ batch_augments=batch_augments),
+ backbone=dict(
+ frozen_stages=-1, norm_eval=False, norm_cfg=norm_cfg, init_cfg=None),
+ neck=dict(norm_cfg=norm_cfg),
+ rpn_head=dict(num_convs=2), # leads to 0.1+ mAP
+ roi_head=dict(
+ bbox_head=dict(
+ type='Shared4Conv1FCBBoxHead',
+ conv_out_channels=256,
+ norm_cfg=head_norm_cfg),
+ mask_head=dict(norm_cfg=head_norm_cfg)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6621d28c0a80bd669fa857ce4eb7058a6f82296c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-50e_coco.py
@@ -0,0 +1,5 @@
+_base_ = 'mask-rcnn_r50_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-100e_coco.py'
+
+# Use RepeatDataset to speed up training
+# change repeat time from 4 (for 100 epochs) to 2 (for 50 epochs)
+train_dataloader = dict(dataset=dict(times=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..f72c07e64b6e72dc0c71ae114877ce5c8513be7b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strong_baselines/metafile.yml
@@ -0,0 +1,24 @@
+Models:
+ - Name: mask-rcnn_r50-caffe_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-100e_coco
+ In Collection: Mask R-CNN
+ Config: configs/strong_baselines/mask-rcnn_r50-caffe_fpn_rpn-2conv_4conv1fc_syncbn-all_lsj-100e_coco.py
+ Metadata:
+ Epochs: 100
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ - LSJ
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNet
+ - FPN
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.7
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ box AP: 40.4
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/strongsort/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strongsort/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..8e08413cbc04d6b552b911b1d9fb6ad2e4205a35
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strongsort/README.md
@@ -0,0 +1,108 @@
+# StrongSORT: Make DeepSORT Great Again
+
+## Abstract
+
+
+
+Existing Multi-Object Tracking (MOT) methods can be roughly classified as tracking-by-detection and joint-detection-association paradigms. Although the latter has elicited more attention and demonstrates comparable performance relative to the former, we claim that the tracking-by-detection paradigm is still the optimal solution in terms of tracking accuracy. In this paper, we revisit the classic tracker DeepSORT and upgrade it from various aspects, i.e., detection, embedding and association. The resulting tracker, called StrongSORT, sets new HOTA and IDF1 records on MOT17 and MOT20. We also present two lightweight and plug-and-play algorithms to further refine the tracking results. Firstly, an appearance-free link model (AFLink) is proposed to associate short tracklets into complete trajectories. To the best of our knowledge, this is the first global link model without appearance information. Secondly, we propose Gaussian-smoothed interpolation (GSI) to compensate for missing detections. Instead of ignoring motion information like linear interpolation, GSI is based on the Gaussian process regression algorithm and can achieve more accurate localizations. Moreover, AFLink and GSI can be plugged into various trackers with a negligible extra computational cost (591.9 and 140.9 Hz, respectively, on MOT17). By integrating StrongSORT with the two algorithms, the final tracker StrongSORT++ ranks first on MOT17 and MOT20 in terms of HOTA and IDF1 metrics and surpasses the second-place one by 1.3 - 2.2. Code will be released soon.
+
+
+
+
+

+
+
+## Citation
+
+
+
+```latex
+@article{du2022strongsort,
+ title={Strongsort: Make deepsort great again},
+ author={Du, Yunhao and Song, Yang and Yang, Bo and Zhao, Yanyun},
+ journal={arXiv preprint arXiv:2202.13514},
+ year={2022}
+}
+```
+
+## Results and models on MOT17
+
+| Method | Detector | ReID | Train Set | Test Set | Public | Inf time (fps) | HOTA | MOTA | IDF1 | FP | FN | IDSw. | Config | Download |
+| :----------: | :------: | :--: | :---------------------------: | :------------: | :----: | :------------: | :--: | :--: | :--: | :---: | :---: | :---: | :----------------------------------------------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| StrongSORT++ | YOLOX-X | R50 | CrowdHuman + MOT17-half-train | MOT17-half-val | N | - | 70.9 | 78.4 | 83.3 | 15237 | 19035 | 582 | [config](strongsort_yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval.py) | [detector](https://download.openmmlab.com/mmtracking/mot/strongsort/mot_dataset/yolox_x_crowdhuman_mot17-private-half_20220812_192036-b6c9ce9a.pth) [reid](https://download.openmmlab.com/mmtracking/mot/reid/reid_r50_6e_mot17-4bf6b63d.pth) [AFLink](https://download.openmmlab.com/mmtracking/mot/strongsort/mot_dataset/aflink_motchallenge_20220812_190310-a7578ad3.pth) |
+
+## Results and models on MOT20
+
+| Method | Detector | ReID | Train Set | Test Set | Public | Inf time (fps) | HOTA | MOTA | IDF1 | FP | FN | IDSw. | Config | Download |
+| :----------: | :------: | :--: | :----------------------: | :--------: | :----: | :------------: | :--: | :--: | :--: | :---: | :---: | :---: | :---------------------------------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| StrongSORT++ | YOLOX-X | R50 | CrowdHuman + MOT20-train | MOT20-test | N | - | 62.9 | 75.5 | 77.3 | 29043 | 96155 | 1640 | [config](strongsort_yolox_x_8xb4-80e_crowdhuman-mot20train_test-mot20test.py) | [detector](https://download.openmmlab.com/mmtracking/mot/strongsort/mot_dataset/yolox_x_crowdhuman_mot20-private_20220812_192123-77c014de.pth) [reid](https://download.openmmlab.com/mmtracking/mot/reid/reid_r50_6e_mot20_20210803_212426-c83b1c01.pth) [AFLink](https://download.openmmlab.com/mmtracking/mot/strongsort/mot_dataset/aflink_motchallenge_20220812_190310-a7578ad3.pth) |
+
+## Get started
+
+### 1. Development Environment Setup
+
+Tracking Development Environment Setup can refer to this [document](../../docs/en/get_started.md).
+
+### 2. Dataset Prepare
+
+Tracking Dataset Prepare can refer to this [document](../../docs/en/user_guides/tracking_dataset_prepare.md).
+
+### 3. Training
+
+We implement StrongSORT with independent detector and ReID models.
+Note that, due to the influence of parameters such as learning rate in default configuration file,
+we recommend using 8 GPUs for training in order to reproduce accuracy.
+
+You can train the detector as follows.
+
+```shell script
+# Training YOLOX-X on crowdhuman and mot17-half-train dataset with following command.
+# The number after config file represents the number of GPUs used. Here we use 8 GPUs.
+bash tools/dist_train.sh configs/det/yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval.py 8
+```
+
+And you can train the ReID model as follows.
+
+```shell script
+# Training ReID model on mot17-train80 dataset with following command.
+# The number after config file represents the number of GPUs used. Here we use 8 GPUs.
+bash tools/dist_train.sh configs/reid/reid_r50_8xb32-6e_mot17train80_test-mot17val20.py 8
+```
+
+If you want to know about more detailed usage of `train.py/dist_train.sh/slurm_train.sh`,
+please refer to this [document](../../docs/en/user_guides/tracking_train_test.md).
+
+### 4. Testing and evaluation
+
+**2.1 Example on MOTxx-halfval dataset**
+
+```shell script
+# Example 1: Test on motXX-half-val set.
+# The number after config file represents the number of GPUs used. Here we use 8 GPUs.
+bash tools/dist_test_tracking.sh configs/strongsort/strongsort_yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval.py 8 --detector ${CHECKPOINT_PATH} --reid ${CHECKPOINT_PATH}
+```
+
+**2.2 Example on MOTxx-test dataset**
+
+If you want to get the results of the [MOT Challenge](https://motchallenge.net/) test set,
+please use the following command to generate result files that can be used for submission.
+It will be stored in `./mot_20_test_res`, you can modify the saved path in `test_evaluator` of the config.
+
+```shell script
+# Example 2: Test on motxx-test set
+# The number after config file represents the number of GPUs used
+bash tools/dist_test_tracking.sh configs/strongsort/strongsort_yolox_x_8xb4-80e_crowdhuman-mot20train_test-mot20test.py 8 --detector ${CHECKPOINT_PATH} --reid ${CHECKPOINT_PATH}
+```
+
+If you want to know about more detailed usage of `test_tracking.py/dist_test_tracking.sh/slurm_test_tracking.sh`,
+please refer to this [document](../../docs/en/user_guides/tracking_train_test.md).
+
+### 3.Inference
+
+Use a single GPU to predict a video and save it as a video.
+
+```shell
+python demo/mot_demo.py demo/demo_mot.mp4 configs/strongsort/strongsort_yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval.py --detector ${CHECKPOINT_FILE} --reid ${CHECKPOINT_PATH} --out mot.mp4
+```
+
+If you want to know about more detailed usage of `mot_demo.py`, please refer to this [document](../../docs/en/user_guides/tracking_inference.md).
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/strongsort/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strongsort/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..08a564b77b866ebe55e2b634faa919817a1de09a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strongsort/metafile.yml
@@ -0,0 +1,48 @@
+Collections:
+ - Name: StrongSORT++
+ Metadata:
+ Training Techniques:
+ - SGD with Momentum
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNet
+ - YOLOX
+ Paper:
+ URL: https://arxiv.org/abs/2202.13514
+ Title: "StrongSORT: Make DeepSORT Great Again"
+ README: configs/strongsort/README.md
+
+Models:
+ - Name: strongsort_yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval
+ In Collection: StrongSORT++
+ Config: configs/strongsort/strongsort_yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval.py
+ Metadata:
+ Training Data: CrowdHuman + MOT17-half-train
+ Results:
+ - Task: Multiple Object Tracking
+ Dataset: MOT17-half-val
+ Metrics:
+ MOTA: 78.3
+ IDF1: 83.2
+ HOTA: 70.9
+ Weights:
+ - https://download.openmmlab.com/mmtracking/mot/strongsort/mot_dataset/yolox_x_crowdhuman_mot17-private-half_20220812_192036-b6c9ce9a.pth
+ - https://download.openmmlab.com/mmtracking/mot/reid/reid_r50_6e_mot17-4bf6b63d.pth
+ - https://download.openmmlab.com/mmtracking/mot/strongsort/mot_dataset/aflink_motchallenge_20220812_190310-a7578ad3.pth
+
+ - Name: strongsort_yolox_x_8xb4-80e_crowdhuman-mot20train_test-mot20test
+ In Collection: StrongSORT++
+ Config: configs/strongsort/strongsort_yolox_x_8xb4-80e_crowdhuman-mot20train_test-mot20test.py
+ Metadata:
+ Training Data: CrowdHuman + MOT20-train
+ Results:
+ - Task: Multiple Object Tracking
+ Dataset: MOT20-test
+ Metrics:
+ MOTA: 75.5
+ IDF1: 77.3
+ HOTA: 62.9
+ Weights:
+ - https://download.openmmlab.com/mmtracking/mot/strongsort/mot_dataset/yolox_x_crowdhuman_mot20-private_20220812_192123-77c014de.pth
+ - https://download.openmmlab.com/mmtracking/mot/reid/reid_r50_6e_mot20_20210803_212426-c83b1c01.pth
+ - https://download.openmmlab.com/mmtracking/mot/strongsort/mot_dataset/aflink_motchallenge_20220812_190310-a7578ad3.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/strongsort/strongsort_yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strongsort/strongsort_yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval.py
new file mode 100644
index 0000000000000000000000000000000000000000..532e2aee718fb481bc81759a2853ac0fddf80e0e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strongsort/strongsort_yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval.py
@@ -0,0 +1,130 @@
+_base_ = [
+ './yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval.py', # noqa: E501
+]
+
+dataset_type = 'MOTChallengeDataset'
+detector = _base_.model
+detector.pop('data_preprocessor')
+del _base_.model
+
+model = dict(
+ type='StrongSORT',
+ data_preprocessor=dict(
+ type='TrackDataPreprocessor',
+ pad_size_divisor=32,
+ batch_augments=[
+ dict(
+ type='BatchSyncRandomResize',
+ random_size_range=(576, 1024),
+ size_divisor=32,
+ interval=10)
+ ]),
+ detector=detector,
+ reid=dict(
+ type='BaseReID',
+ data_preprocessor=dict(type='mmpretrain.ClsDataPreprocessor'),
+ backbone=dict(
+ type='mmpretrain.ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(3, ),
+ style='pytorch'),
+ neck=dict(type='GlobalAveragePooling', kernel_size=(8, 4), stride=1),
+ head=dict(
+ type='LinearReIDHead',
+ num_fcs=1,
+ in_channels=2048,
+ fc_channels=1024,
+ out_channels=128,
+ num_classes=380,
+ loss_cls=dict(type='mmpretrain.CrossEntropyLoss', loss_weight=1.0),
+ loss_triplet=dict(type='TripletLoss', margin=0.3, loss_weight=1.0),
+ norm_cfg=dict(type='BN1d'),
+ act_cfg=dict(type='ReLU'))),
+ cmc=dict(
+ type='CameraMotionCompensation',
+ warp_mode='cv2.MOTION_EUCLIDEAN',
+ num_iters=100,
+ stop_eps=0.00001),
+ tracker=dict(
+ type='StrongSORTTracker',
+ motion=dict(type='KalmanFilter', center_only=False, use_nsa=True),
+ obj_score_thr=0.6,
+ reid=dict(
+ num_samples=None,
+ img_scale=(256, 128),
+ img_norm_cfg=dict(
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ to_rgb=True),
+ match_score_thr=0.3,
+ motion_weight=0.02,
+ ),
+ match_iou_thr=0.7,
+ momentums=dict(embeds=0.1, ),
+ num_tentatives=2,
+ num_frames_retain=100),
+ postprocess_model=dict(
+ type='AppearanceFreeLink',
+ checkpoint= # noqa: E251
+ 'https://download.openmmlab.com/mmtracking/mot/strongsort/mot_dataset/aflink_motchallenge_20220812_190310-a7578ad3.pth', # noqa: E501
+ temporal_threshold=(0, 30),
+ spatial_threshold=50,
+ confidence_threshold=0.95,
+ ))
+
+train_pipeline = None
+test_pipeline = [
+ dict(
+ type='TransformBroadcaster',
+ transforms=[
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='Resize', scale=_base_.img_scale, keep_ratio=True),
+ dict(
+ type='Pad',
+ size_divisor=32,
+ pad_val=dict(img=(114.0, 114.0, 114.0))),
+ dict(type='LoadTrackAnnotations'),
+ ]),
+ dict(type='PackTrackInputs')
+]
+
+train_dataloader = None
+val_dataloader = dict(
+ # Now StrongSORT only support video_based sampling
+ sampler=dict(type='DefaultSampler', shuffle=False, round_up=False),
+ dataset=dict(
+ _delete_=True,
+ type=dataset_type,
+ data_root=_base_.data_root,
+ ann_file='annotations/half-val_cocoformat.json',
+ data_prefix=dict(img_path='train'),
+ # when you evaluate track performance, you need to remove metainfo
+ test_mode=True,
+ pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+train_cfg = None
+optim_wrapper = None
+
+# evaluator
+val_evaluator = dict(
+ _delete_=True,
+ type='MOTChallengeMetric',
+ metric=['HOTA', 'CLEAR', 'Identity'],
+ # use_postprocess to support AppearanceFreeLink in val_evaluator
+ use_postprocess=True,
+ postprocess_tracklet_cfg=[
+ dict(
+ type='InterpolateTracklets',
+ min_num_frames=5,
+ max_num_frames=20,
+ use_gsi=True,
+ smooth_tau=10)
+ ])
+test_evaluator = val_evaluator
+
+default_hooks = dict(logger=dict(type='LoggerHook', interval=1))
+
+del _base_.param_scheduler
+del _base_.custom_hooks
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/strongsort/strongsort_yolox_x_8xb4-80e_crowdhuman-mot20train_test-mot20test.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strongsort/strongsort_yolox_x_8xb4-80e_crowdhuman-mot20train_test-mot20test.py
new file mode 100644
index 0000000000000000000000000000000000000000..eab97063932528df7e17c7d65bf9f0d13f5dfa73
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strongsort/strongsort_yolox_x_8xb4-80e_crowdhuman-mot20train_test-mot20test.py
@@ -0,0 +1,44 @@
+_base_ = [
+ './strongsort_yolox_x_8xb4-80e_crowdhuman-mot17halftrain'
+ '_test-mot17halfval.py'
+]
+
+img_scale = (1600, 896) # width, height
+
+model = dict(
+ data_preprocessor=dict(
+ type='TrackDataPreprocessor',
+ pad_size_divisor=32,
+ batch_augments=[
+ dict(type='BatchSyncRandomResize', random_size_range=(640, 1152))
+ ]))
+
+test_pipeline = [
+ dict(
+ type='TransformBroadcaster',
+ transforms=[
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='Resize', scale=img_scale, keep_ratio=True),
+ dict(
+ type='Pad',
+ size_divisor=32,
+ pad_val=dict(img=(114.0, 114.0, 114.0))),
+ dict(type='LoadTrackAnnotations'),
+ ]),
+ dict(type='PackTrackInputs')
+]
+
+val_dataloader = dict(
+ dataset=dict(
+ data_root='data/MOT17',
+ ann_file='annotations/train_cocoformat.json',
+ data_prefix=dict(img_path='train'),
+ pipeline=test_pipeline))
+test_dataloader = dict(
+ dataset=dict(
+ data_root='data/MOT20',
+ ann_file='annotations/test_cocoformat.json',
+ data_prefix=dict(img_path='test'),
+ pipeline=test_pipeline))
+
+test_evaluator = dict(format_only=True, outfile_prefix='./mot_20_test_res')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/strongsort/yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strongsort/yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval.py
new file mode 100644
index 0000000000000000000000000000000000000000..59a52e4394b5825d40a99e08793147fe836b4c19
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strongsort/yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval.py
@@ -0,0 +1,188 @@
+_base_ = ['../yolox/yolox_x_8xb8-300e_coco.py']
+
+data_root = 'data/MOT17/'
+
+img_scale = (1440, 800) # width, height
+batch_size = 4
+
+# model settings
+model = dict(
+ bbox_head=dict(num_classes=1),
+ test_cfg=dict(nms=dict(iou_threshold=0.7)),
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint= # noqa: E251
+ 'https://download.openmmlab.com/mmdetection/v2.0/yolox/yolox_x_8x8_300e_coco/yolox_x_8x8_300e_coco_20211126_140254-1ef88d67.pth' # noqa: E501
+ ))
+
+train_pipeline = [
+ dict(
+ type='Mosaic',
+ img_scale=img_scale,
+ pad_val=114.0,
+ bbox_clip_border=False),
+ dict(
+ type='RandomAffine',
+ scaling_ratio_range=(0.1, 2),
+ border=(-img_scale[0] // 2, -img_scale[1] // 2),
+ bbox_clip_border=False),
+ dict(
+ type='MixUp',
+ img_scale=img_scale,
+ ratio_range=(0.8, 1.6),
+ pad_val=114.0,
+ bbox_clip_border=False),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='Resize',
+ scale=img_scale,
+ keep_ratio=True,
+ clip_object_border=False),
+ dict(type='Pad', size_divisor=32, pad_val=dict(img=(114.0, 114.0, 114.0))),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1, 1), keep_empty=False),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='Resize', scale=img_scale, keep_ratio=True),
+ dict(type='Pad', size_divisor=32, pad_val=dict(img=(114.0, 114.0, 114.0))),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ _delete_=True,
+ batch_size=batch_size,
+ num_workers=4,
+ persistent_workers=True,
+ pin_memory=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ dataset=dict(
+ type='MultiImageMixDataset',
+ dataset=dict(
+ type='ConcatDataset',
+ datasets=[
+ dict(
+ type='CocoDataset',
+ data_root=data_root,
+ ann_file='annotations/half-train_cocoformat.json',
+ data_prefix=dict(img='train'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ metainfo=dict(classes=('pedestrian', )),
+ pipeline=[
+ dict(
+ type='LoadImageFromFile',
+ backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ ]),
+ dict(
+ type='CocoDataset',
+ data_root='data/crowdhuman',
+ ann_file='annotations/crowdhuman_train.json',
+ data_prefix=dict(img='train'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ metainfo=dict(classes=('pedestrian', )),
+ pipeline=[
+ dict(
+ type='LoadImageFromFile',
+ backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ ]),
+ dict(
+ type='CocoDataset',
+ data_root='data/crowdhuman',
+ ann_file='annotations/crowdhuman_val.json',
+ data_prefix=dict(img='val'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ metainfo=dict(classes=('pedestrian', )),
+ pipeline=[
+ dict(
+ type='LoadImageFromFile',
+ backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ ]),
+ ]),
+ pipeline=train_pipeline))
+
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ dataset=dict(
+ data_root=data_root,
+ ann_file='annotations/half-val_cocoformat.json',
+ data_prefix=dict(img='train'),
+ metainfo=dict(classes=('pedestrian', )),
+ pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# training settings
+max_epochs = 80
+num_last_epochs = 10
+interval = 5
+
+train_cfg = dict(max_epochs=max_epochs, val_begin=75, val_interval=1)
+
+# optimizer
+# default 8 gpu
+base_lr = 0.001 / 8 * batch_size
+optim_wrapper = dict(optimizer=dict(lr=base_lr))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='QuadraticWarmupLR',
+ by_epoch=True,
+ begin=0,
+ end=1,
+ convert_to_iter_based=True),
+ dict(
+ type='CosineAnnealingLR',
+ eta_min=base_lr * 0.05,
+ begin=1,
+ T_max=max_epochs - num_last_epochs,
+ end=max_epochs - num_last_epochs,
+ by_epoch=True,
+ convert_to_iter_based=True),
+ dict(
+ type='ConstantLR',
+ by_epoch=True,
+ factor=1,
+ begin=max_epochs - num_last_epochs,
+ end=max_epochs,
+ )
+]
+
+default_hooks = dict(
+ checkpoint=dict(
+ interval=1,
+ max_keep_ckpts=5 # only keep latest 5 checkpoints
+ ))
+
+custom_hooks = [
+ dict(
+ type='YOLOXModeSwitchHook',
+ num_last_epochs=num_last_epochs,
+ priority=48),
+ dict(type='SyncNormHook', priority=48),
+ dict(
+ type='EMAHook',
+ ema_type='ExpMomentumEMA',
+ momentum=0.0001,
+ update_buffers=True,
+ priority=49)
+]
+
+# evaluator
+val_evaluator = dict(
+ ann_file=data_root + 'annotations/half-val_cocoformat.json',
+ format_only=False)
+test_evaluator = val_evaluator
+
+del _base_.tta_model
+del _base_.tta_pipeline
+del _base_.train_dataset
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/strongsort/yolox_x_8xb4-80e_crowdhuman-mot20train_test-mot20test.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strongsort/yolox_x_8xb4-80e_crowdhuman-mot20train_test-mot20test.py
new file mode 100644
index 0000000000000000000000000000000000000000..d4eb3cb2c9804f0219ba91d0b5d460da342ab668
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/strongsort/yolox_x_8xb4-80e_crowdhuman-mot20train_test-mot20test.py
@@ -0,0 +1,108 @@
+_base_ = ['./yolox_x_8xb4-80e_crowdhuman-mot17halftrain_test-mot17halfval.py']
+
+data_root = 'data/MOT20/'
+
+img_scale = (1600, 896) # width, height
+
+# model settings
+model = dict(
+ data_preprocessor=dict(batch_augments=[
+ dict(type='BatchSyncRandomResize', random_size_range=(640, 1152))
+ ]))
+
+train_pipeline = [
+ dict(
+ type='Mosaic',
+ img_scale=img_scale,
+ pad_val=114.0,
+ bbox_clip_border=True),
+ dict(
+ type='RandomAffine',
+ scaling_ratio_range=(0.1, 2),
+ border=(-img_scale[0] // 2, -img_scale[1] // 2),
+ bbox_clip_border=True),
+ dict(
+ type='MixUp',
+ img_scale=img_scale,
+ ratio_range=(0.8, 1.6),
+ pad_val=114.0,
+ bbox_clip_border=True),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='Resize',
+ scale=img_scale,
+ keep_ratio=True,
+ clip_object_border=True),
+ dict(type='Pad', size_divisor=32, pad_val=dict(img=(114.0, 114.0, 114.0))),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1, 1), keep_empty=False),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='Resize', scale=img_scale, keep_ratio=True),
+ dict(type='Pad', size_divisor=32, pad_val=dict(img=(114.0, 114.0, 114.0))),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ type='MultiImageMixDataset',
+ dataset=dict(
+ type='ConcatDataset',
+ datasets=[
+ dict(
+ type='CocoDataset',
+ data_root=data_root,
+ ann_file='annotations/train_cocoformat.json',
+ data_prefix=dict(img='train'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ metainfo=dict(classes=('pedestrian', )),
+ pipeline=[
+ dict(
+ type='LoadImageFromFile',
+ backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ ]),
+ dict(
+ type='CocoDataset',
+ data_root='data/crowdhuman',
+ ann_file='annotations/crowdhuman_train.json',
+ data_prefix=dict(img='train'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ metainfo=dict(classes=('pedestrian', )),
+ pipeline=[
+ dict(
+ type='LoadImageFromFile',
+ backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ ]),
+ dict(
+ type='CocoDataset',
+ data_root='data/crowdhuman',
+ ann_file='annotations/crowdhuman_val.json',
+ data_prefix=dict(img='val'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ metainfo=dict(classes=('pedestrian', )),
+ pipeline=[
+ dict(
+ type='LoadImageFromFile',
+ backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ ]),
+ ]),
+ pipeline=train_pipeline))
+
+val_dataloader = dict(
+ dataset=dict(
+ data_root='data/MOT17', ann_file='annotations/train_cocoformat.json'))
+test_dataloader = val_dataloader
+
+# evaluator
+val_evaluator = dict(ann_file='data/MOT17/annotations/train_cocoformat.json')
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..99bcf6ed7102ac7cd9801a7350c7e4070b60cbf4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/README.md
@@ -0,0 +1,41 @@
+# Swin
+
+> [Swin Transformer: Hierarchical Vision Transformer using Shifted Windows](https://arxiv.org/abs/2103.14030)
+
+
+
+## Abstract
+
+This paper presents a new vision Transformer, called Swin Transformer, that capably serves as a general-purpose backbone for computer vision. Challenges in adapting Transformer from language to vision arise from differences between the two domains, such as large variations in the scale of visual entities and the high resolution of pixels in images compared to words in text. To address these differences, we propose a hierarchical Transformer whose representation is computed with Shifted windows. The shifted windowing scheme brings greater efficiency by limiting self-attention computation to non-overlapping local windows while also allowing for cross-window connection. This hierarchical architecture has the flexibility to model at various scales and has linear computational complexity with respect to image size. These qualities of Swin Transformer make it compatible with a broad range of vision tasks, including image classification (87.3 top-1 accuracy on ImageNet-1K) and dense prediction tasks such as object detection (58.7 box AP and 51.1 mask AP on COCO test-dev) and semantic segmentation (53.5 mIoU on ADE20K val). Its performance surpasses the previous state-of-the-art by a large margin of +2.7 box AP and +2.6 mask AP on COCO, and +3.2 mIoU on ADE20K, demonstrating the potential of Transformer-based models as vision backbones. The hierarchical design and the shifted window approach also prove beneficial for all-MLP architectures.
+
+
+

+
+
+## Results and Models
+
+### Mask R-CNN
+
+| Backbone | Pretrain | Lr schd | Multi-scale crop | FP16 | Mem (GB) | Inf time (fps) | box AP | mask AP | Config | Download |
+| :------: | :---------: | :-----: | :--------------: | :--: | :------: | :------------: | :----: | :-----: | :-----------------------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Swin-T | ImageNet-1K | 1x | no | no | 7.6 | | 42.7 | 39.3 | [config](./mask-rcnn_swin-t-p4-w7_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/swin/mask_rcnn_swin-t-p4-w7_fpn_1x_coco/mask_rcnn_swin-t-p4-w7_fpn_1x_coco_20210902_120937-9d6b7cfa.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/swin/mask_rcnn_swin-t-p4-w7_fpn_1x_coco/mask_rcnn_swin-t-p4-w7_fpn_1x_coco_20210902_120937.log.json) |
+| Swin-T | ImageNet-1K | 3x | yes | no | 10.2 | | 46.0 | 41.6 | [config](./mask-rcnn_swin-t-p4-w7_fpn_ms-crop-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/swin/mask_rcnn_swin-t-p4-w7_fpn_ms-crop-3x_coco/mask_rcnn_swin-t-p4-w7_fpn_ms-crop-3x_coco_20210906_131725-bacf6f7b.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/swin/mask_rcnn_swin-t-p4-w7_fpn_ms-crop-3x_coco/mask_rcnn_swin-t-p4-w7_fpn_ms-crop-3x_coco_20210906_131725.log.json) |
+| Swin-T | ImageNet-1K | 3x | yes | yes | 7.8 | | 46.0 | 41.7 | [config](./mask-rcnn_swin-t-p4-w7_fpn_amp-ms-crop-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/swin/mask_rcnn_swin-t-p4-w7_fpn_fp16_ms-crop-3x_coco/mask_rcnn_swin-t-p4-w7_fpn_fp16_ms-crop-3x_coco_20210908_165006-90a4008c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/swin/mask_rcnn_swin-t-p4-w7_fpn_fp16_ms-crop-3x_coco/mask_rcnn_swin-t-p4-w7_fpn_fp16_ms-crop-3x_coco_20210908_165006.log.json) |
+| Swin-S | ImageNet-1K | 3x | yes | yes | 11.9 | | 48.2 | 43.2 | [config](./mask-rcnn_swin-s-p4-w7_fpn_amp-ms-crop-3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/swin/mask_rcnn_swin-s-p4-w7_fpn_fp16_ms-crop-3x_coco/mask_rcnn_swin-s-p4-w7_fpn_fp16_ms-crop-3x_coco_20210903_104808-b92c91f1.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/swin/mask_rcnn_swin-s-p4-w7_fpn_fp16_ms-crop-3x_coco/mask_rcnn_swin-s-p4-w7_fpn_fp16_ms-crop-3x_coco_20210903_104808.log.json) |
+
+### Notice
+
+Please follow the example
+of `retinanet_swin-t-p4-w7_fpn_1x_coco.py` when you want to combine Swin Transformer with
+the one-stage detector. Because there is a layer norm at the outs of Swin Transformer, you must set `start_level` as 0 in FPN, so we have to set the `out_indices` of backbone as `[1,2,3]`.
+
+## Citation
+
+```latex
+@article{liu2021Swin,
+ title={Swin Transformer: Hierarchical Vision Transformer using Shifted Windows},
+ author={Liu, Ze and Lin, Yutong and Cao, Yue and Hu, Han and Wei, Yixuan and Zhang, Zheng and Lin, Stephen and Guo, Baining},
+ journal={arXiv preprint arXiv:2103.14030},
+ year={2021}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/mask-rcnn_swin-s-p4-w7_fpn_amp-ms-crop-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/mask-rcnn_swin-s-p4-w7_fpn_amp-ms-crop-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..4a3e8ad900553c38d11ddc7747cbc0f244f6b4c7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/mask-rcnn_swin-s-p4-w7_fpn_amp-ms-crop-3x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './mask-rcnn_swin-t-p4-w7_fpn_amp-ms-crop-3x_coco.py'
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_small_patch4_window7_224.pth' # noqa
+model = dict(
+ backbone=dict(
+ depths=[2, 2, 18, 2],
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/mask-rcnn_swin-t-p4-w7_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/mask-rcnn_swin-t-p4-w7_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5471caa139c0b7670f995501347ddf80383e9268
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/mask-rcnn_swin-t-p4-w7_fpn_1x_coco.py
@@ -0,0 +1,60 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_tiny_patch4_window7_224.pth' # noqa
+model = dict(
+ type='MaskRCNN',
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ embed_dims=96,
+ depths=[2, 2, 6, 2],
+ num_heads=[3, 6, 12, 24],
+ window_size=7,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.2,
+ patch_norm=True,
+ out_indices=(0, 1, 2, 3),
+ with_cp=False,
+ convert_weights=True,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ neck=dict(in_channels=[96, 192, 384, 768]))
+
+max_epochs = 12
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0,
+ end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ paramwise_cfg=dict(
+ custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'relative_position_bias_table': dict(decay_mult=0.),
+ 'norm': dict(decay_mult=0.)
+ }),
+ optimizer=dict(
+ _delete_=True,
+ type='AdamW',
+ lr=0.0001,
+ betas=(0.9, 0.999),
+ weight_decay=0.05))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/mask-rcnn_swin-t-p4-w7_fpn_amp-ms-crop-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/mask-rcnn_swin-t-p4-w7_fpn_amp-ms-crop-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..622087ba7164fda53a70eb927b9258572b7c8ef0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/mask-rcnn_swin-t-p4-w7_fpn_amp-ms-crop-3x_coco.py
@@ -0,0 +1,3 @@
+_base_ = './mask-rcnn_swin-t-p4-w7_fpn_ms-crop-3x_coco.py'
+# Enable automatic-mixed-precision training with AmpOptimWrapper.
+optim_wrapper = dict(type='AmpOptimWrapper')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/mask-rcnn_swin-t-p4-w7_fpn_ms-crop-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/mask-rcnn_swin-t-p4-w7_fpn_ms-crop-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..7024b73249ca8c77da89ab9e4653757f36a1d1d2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/mask-rcnn_swin-t-p4-w7_fpn_ms-crop-3x_coco.py
@@ -0,0 +1,99 @@
+_base_ = [
+ '../_base_/models/mask-rcnn_r50_fpn.py',
+ '../_base_/datasets/coco_instance.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_tiny_patch4_window7_224.pth' # noqa
+
+model = dict(
+ type='MaskRCNN',
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ embed_dims=96,
+ depths=[2, 2, 6, 2],
+ num_heads=[3, 6, 12, 24],
+ window_size=7,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.2,
+ patch_norm=True,
+ out_indices=(0, 1, 2, 3),
+ with_cp=False,
+ convert_weights=True,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ neck=dict(in_channels=[96, 192, 384, 768]))
+
+# augmentation strategy originates from DETR / Sparse RCNN
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[[
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(400, 1333), (500, 1333), (600, 1333)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333),
+ (576, 1333), (608, 1333), (640, 1333),
+ (672, 1333), (704, 1333), (736, 1333),
+ (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]]),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+
+max_epochs = 36
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0,
+ end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[27, 33],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ paramwise_cfg=dict(
+ custom_keys={
+ 'absolute_pos_embed': dict(decay_mult=0.),
+ 'relative_position_bias_table': dict(decay_mult=0.),
+ 'norm': dict(decay_mult=0.)
+ }),
+ optimizer=dict(
+ _delete_=True,
+ type='AdamW',
+ lr=0.0001,
+ betas=(0.9, 0.999),
+ weight_decay=0.05))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..763f9300d44bcc3f9348951f3640ada171c3ce05
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/metafile.yml
@@ -0,0 +1,120 @@
+Models:
+ - Name: mask-rcnn_swin-s-p4-w7_fpn_amp-ms-crop-3x_coco
+ In Collection: Mask R-CNN
+ Config: configs/swin/mask-rcnn_swin-s-p4-w7_fpn_amp-ms-crop-3x_coco.py
+ Metadata:
+ Training Memory (GB): 11.9
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - AdamW
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Swin Transformer
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 48.2
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 43.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/swin/mask_rcnn_swin-s-p4-w7_fpn_fp16_ms-crop-3x_coco/mask_rcnn_swin-s-p4-w7_fpn_fp16_ms-crop-3x_coco_20210903_104808-b92c91f1.pth
+ Paper:
+ URL: https://arxiv.org/abs/2107.08430
+ Title: 'Swin Transformer: Hierarchical Vision Transformer using Shifted Windows'
+ README: configs/swin/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.16.0/mmdet/models/backbones/swin.py#L465
+ Version: v2.16.0
+
+ - Name: mask-rcnn_swin-t-p4-w7_fpn_ms-crop-3x_coco
+ In Collection: Mask R-CNN
+ Config: configs/swin/mask-rcnn_swin-t-p4-w7_fpn_ms-crop-3x_coco.py
+ Metadata:
+ Training Memory (GB): 10.2
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - AdamW
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Swin Transformer
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 41.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/swin/mask_rcnn_swin-t-p4-w7_fpn_ms-crop-3x_coco/mask_rcnn_swin-t-p4-w7_fpn_ms-crop-3x_coco_20210906_131725-bacf6f7b.pth
+ Paper:
+ URL: https://arxiv.org/abs/2107.08430
+ Title: 'Swin Transformer: Hierarchical Vision Transformer using Shifted Windows'
+ README: configs/swin/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.16.0/mmdet/models/backbones/swin.py#L465
+ Version: v2.16.0
+
+ - Name: mask-rcnn_swin-t-p4-w7_fpn_1x_coco
+ In Collection: Mask R-CNN
+ Config: configs/swin/mask-rcnn_swin-t-p4-w7_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 7.6
+ Epochs: 12
+ Training Data: COCO
+ Training Techniques:
+ - AdamW
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Swin Transformer
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.7
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 39.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/swin/mask_rcnn_swin-t-p4-w7_fpn_1x_coco/mask_rcnn_swin-t-p4-w7_fpn_1x_coco_20210902_120937-9d6b7cfa.pth
+ Paper:
+ URL: https://arxiv.org/abs/2107.08430
+ Title: 'Swin Transformer: Hierarchical Vision Transformer using Shifted Windows'
+ README: configs/swin/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.16.0/mmdet/models/backbones/swin.py#L465
+ Version: v2.16.0
+
+ - Name: mask-rcnn_swin-t-p4-w7_fpn_amp-ms-crop-3x_coco
+ In Collection: Mask R-CNN
+ Config: configs/swin/mask-rcnn_swin-t-p4-w7_fpn_amp-ms-crop-3x_coco.py
+ Metadata:
+ Training Memory (GB): 7.8
+ Epochs: 36
+ Training Data: COCO
+ Training Techniques:
+ - AdamW
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Swin Transformer
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.0
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 41.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/swin/mask_rcnn_swin-t-p4-w7_fpn_fp16_ms-crop-3x_coco/mask_rcnn_swin-t-p4-w7_fpn_fp16_ms-crop-3x_coco_20210908_165006-90a4008c.pth
+ Paper:
+ URL: https://arxiv.org/abs/2107.08430
+ Title: 'Swin Transformer: Hierarchical Vision Transformer using Shifted Windows'
+ README: configs/swin/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.16.0/mmdet/models/backbones/swin.py#L465
+ Version: v2.16.0
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/retinanet_swin-t-p4-w7_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/retinanet_swin-t-p4-w7_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..2f40a87e8cf8593edd92f024d0bb0ed43a87b4fb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/swin/retinanet_swin-t-p4-w7_fpn_1x_coco.py
@@ -0,0 +1,31 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_tiny_patch4_window7_224.pth' # noqa
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ embed_dims=96,
+ depths=[2, 2, 6, 2],
+ num_heads=[3, 6, 12, 24],
+ window_size=7,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.2,
+ patch_norm=True,
+ out_indices=(1, 2, 3),
+ # Please only add indices that would be used
+ # in FPN, otherwise some parameter will not be used
+ with_cp=False,
+ convert_weights=True,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ neck=dict(in_channels=[192, 384, 768], start_level=0, num_outs=5))
+
+# optimizer
+optim_wrapper = dict(optimizer=dict(lr=0.01))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/timm_example/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/timm_example/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..848f8d3c269cc0de2fad5fa60a62ed44bfd9b29e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/timm_example/README.md
@@ -0,0 +1,62 @@
+# Timm Example
+
+> [PyTorch Image Models](https://github.com/rwightman/pytorch-image-models)
+
+
+
+## Abstract
+
+Py**T**orch **Im**age **M**odels (`timm`) is a collection of image models, layers, utilities, optimizers, schedulers, data-loaders / augmentations, and reference training / validation scripts that aim to pull together a wide variety of SOTA models with ability to reproduce ImageNet training results.
+
+
+
+## Results and Models
+
+### RetinaNet
+
+| Backbone | Style | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :-------------: | :-----: | :-----: | :------: | :------------: | :----: | :-------------------------------------------------------: | :------: |
+| R-50 | pytorch | 1x | | | | [config](./retinanet_timm-tv-resnet50_fpn_1x_coco.py) | |
+| EfficientNet-B1 | - | 1x | | | | [config](./retinanet_timm-efficientnet-b1_fpn_1x_coco.py) | |
+
+## Usage
+
+### Install additional requirements
+
+MMDetection supports timm backbones via `TIMMBackbone`, a wrapper class in MMPretrain.
+Thus, you need to install `mmpretrain` in addition to timm.
+If you have already installed requirements for mmdet, run
+
+```shell
+pip install 'dataclasses; python_version<"3.7"'
+pip install timm
+pip install mmpretrain
+```
+
+See [this document](https://mmpretrain.readthedocs.io/en/latest/get_started.html#installation) for the details of MMPretrain installation.
+
+### Edit config
+
+- See example configs for basic usage.
+- See the documents of [timm feature extraction](https://rwightman.github.io/pytorch-image-models/feature_extraction/#multi-scale-feature-maps-feature-pyramid) and [TIMMBackbone](https://mmpretrain.readthedocs.io/en/latest/api/generated/mmpretrain.models.backbones.TIMMBackbone.html#mmpretrain.models.backbones.TIMMBackbone) for details.
+- Which feature map is output depends on the backbone.
+ Please check `backbone out_channels` and `backbone out_strides` in your log, and modify `model.neck.in_channels` and `model.backbone.out_indices` if necessary.
+- If you use Vision Transformer models that do not support `features_only=True`, add `custom_hooks = []` to your config to disable `NumClassCheckHook`.
+
+## Citation
+
+```latex
+@misc{rw2019timm,
+ author = {Ross Wightman},
+ title = {PyTorch Image Models},
+ year = {2019},
+ publisher = {GitHub},
+ journal = {GitHub repository},
+ doi = {10.5281/zenodo.4414861},
+ howpublished = {\url{https://github.com/rwightman/pytorch-image-models}}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/timm_example/retinanet_timm-efficientnet-b1_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/timm_example/retinanet_timm-efficientnet-b1_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b87dddf50f7179dc143b9ab9aecb07d09d4dea4b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/timm_example/retinanet_timm-efficientnet-b1_fpn_1x_coco.py
@@ -0,0 +1,23 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+# please install mmpretrain
+# import mmpretrain.models to trigger register_module in mmpretrain
+custom_imports = dict(
+ imports=['mmpretrain.models'], allow_failed_imports=False)
+
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='mmpretrain.TIMMBackbone',
+ model_name='efficientnet_b1',
+ features_only=True,
+ pretrained=True,
+ out_indices=(1, 2, 3, 4)),
+ neck=dict(in_channels=[24, 40, 112, 320]))
+
+# optimizer
+optim_wrapper = dict(optimizer=dict(lr=0.01))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/timm_example/retinanet_timm-tv-resnet50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/timm_example/retinanet_timm-tv-resnet50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..74e43506959574abbf08feb44848f4bfa8d65719
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/timm_example/retinanet_timm-tv-resnet50_fpn_1x_coco.py
@@ -0,0 +1,22 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+# please install mmpretrain
+# import mmpretrain.models to trigger register_module in mmpretrain
+custom_imports = dict(
+ imports=['mmpretrain.models'], allow_failed_imports=False)
+
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='mmpretrain.TIMMBackbone',
+ model_name='tv_resnet50', # ResNet-50 with torchvision weights
+ features_only=True,
+ pretrained=True,
+ out_indices=(1, 2, 3, 4)))
+
+# optimizer
+optim_wrapper = dict(optimizer=dict(lr=0.01))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..9371d9d783ffdca321fa9befc3c93279d45673a7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/README.md
@@ -0,0 +1,40 @@
+# TOOD
+
+> [TOOD: Task-aligned One-stage Object Detection](https://arxiv.org/abs/2108.07755)
+
+
+
+## Abstract
+
+One-stage object detection is commonly implemented by optimizing two sub-tasks: object classification and localization, using heads with two parallel branches, which might lead to a certain level of spatial misalignment in predictions between the two tasks. In this work, we propose a Task-aligned One-stage Object Detection (TOOD) that explicitly aligns the two tasks in a learning-based manner. First, we design a novel Task-aligned Head (T-Head) which offers a better balance between learning task-interactive and task-specific features, as well as a greater flexibility to learn the alignment via a task-aligned predictor. Second, we propose Task Alignment Learning (TAL) to explicitly pull closer (or even unify) the optimal anchors for the two tasks during training via a designed sample assignment scheme and a task-aligned loss. Extensive experiments are conducted on MS-COCO, where TOOD achieves a 51.1 AP at single-model single-scale testing. This surpasses the recent one-stage detectors by a large margin, such as ATSS (47.7 AP), GFL (48.2 AP), and PAA (49.0 AP), with fewer parameters and FLOPs. Qualitative results also demonstrate the effectiveness of TOOD for better aligning the tasks of object classification and localization.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Style | Anchor Type | Lr schd | Multi-scale Training | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :---------------: | :-----: | :----------: | :-----: | :------------------: | :------: | :------------: | :----: | :-------------------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | pytorch | Anchor-free | 1x | N | 4.1 | | 42.4 | [config](./tood_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/tood/tood_r50_fpn_1x_coco/tood_r50_fpn_1x_coco_20211210_103425-20e20746.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/tood/tood_r50_fpn_1x_coco/tood_r50_fpn_1x_coco_20211210_103425.log) |
+| R-50 | pytorch | Anchor-based | 1x | N | 4.1 | | 42.4 | [config](./tood_r50_fpn_anchor-based_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/tood/tood_r50_fpn_anchor_based_1x_coco/tood_r50_fpn_anchor_based_1x_coco_20211214_100105-b776c134.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/tood/tood_r50_fpn_anchor_based_1x_coco/tood_r50_fpn_anchor_based_1x_coco_20211214_100105.log) |
+| R-50 | pytorch | Anchor-free | 2x | Y | 4.1 | | 44.5 | [config](./tood_r50_fpn_ms-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/tood/tood_r50_fpn_mstrain_2x_coco/tood_r50_fpn_mstrain_2x_coco_20211210_144231-3b23174c.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/tood/tood_r50_fpn_mstrain_2x_coco/tood_r50_fpn_mstrain_2x_coco_20211210_144231.log) |
+| R-101 | pytorch | Anchor-free | 2x | Y | 6.0 | | 46.1 | [config](./tood_r101_fpn_ms-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/tood/tood_r101_fpn_mstrain_2x_coco/tood_r101_fpn_mstrain_2x_coco_20211210_144232-a18f53c8.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/tood/tood_r101_fpn_mstrain_2x_coco/tood_r101_fpn_mstrain_2x_coco_20211210_144232.log) |
+| R-101-dcnv2 | pytorch | Anchor-free | 2x | Y | 6.2 | | 49.3 | [config](./tood_r101-dconv-c3-c5_fpn_ms-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/tood/tood_r101_fpn_dconv_c3-c5_mstrain_2x_coco/tood_r101_fpn_dconv_c3-c5_mstrain_2x_coco_20211210_213728-4a824142.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/tood/tood_r101_fpn_dconv_c3-c5_mstrain_2x_coco/tood_r101_fpn_dconv_c3-c5_mstrain_2x_coco_20211210_213728.log) |
+| X-101-64x4d | pytorch | Anchor-free | 2x | Y | 10.2 | | 47.6 | [config](./tood_x101-64x4d_fpn_ms-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/tood/tood_x101_64x4d_fpn_mstrain_2x_coco/tood_x101_64x4d_fpn_mstrain_2x_coco_20211211_003519-a4f36113.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/tood/tood_x101_64x4d_fpn_mstrain_2x_coco/tood_x101_64x4d_fpn_mstrain_2x_coco_20211211_003519.log) |
+| X-101-64x4d-dcnv2 | pytorch | Anchor-free | 2x | Y | | | | [config](./tood_x101-64x4d-dconv-c4-c5_fpn_ms-2x_coco.py) | [model](<>) \| [log](<>) |
+
+\[1\] *1x and 2x mean the model is trained for 90K and 180K iterations, respectively.* \
+\[2\] *All results are obtained with a single model and without any test time data augmentation such as multi-scale, flipping and etc..* \
+\[3\] *`dcnv2` denotes deformable convolutional networks v2.* \\
+
+## Citation
+
+```latex
+@inproceedings{feng2021tood,
+ title={TOOD: Task-aligned One-stage Object Detection},
+ author={Feng, Chengjian and Zhong, Yujie and Gao, Yu and Scott, Matthew R and Huang, Weilin},
+ booktitle={ICCV},
+ year={2021}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..d2bc08073a10ef153b9c97f4d2742e5f85015aa5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/metafile.yml
@@ -0,0 +1,95 @@
+Collections:
+ - Name: TOOD
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - TOOD
+ Paper:
+ URL: https://arxiv.org/abs/2108.07755
+ Title: 'TOOD: Task-aligned One-stage Object Detection'
+ README: configs/tood/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.20.0/mmdet/models/detectors/tood.py#L7
+ Version: v2.20.0
+
+Models:
+ - Name: tood_r101_fpn_ms-2x_coco
+ In Collection: TOOD
+ Config: configs/tood/tood_r101_fpn_ms-2x_coco.py
+ Metadata:
+ Training Memory (GB): 6.0
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.1
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/tood/tood_r101_fpn_mstrain_2x_coco/tood_r101_fpn_mstrain_2x_coco_20211210_144232-a18f53c8.pth
+
+ - Name: tood_x101-64x4d_fpn_ms-2x_coco
+ In Collection: TOOD
+ Config: configs/tood/tood_x101-64x4d_fpn_ms-2x_coco.py
+ Metadata:
+ Training Memory (GB): 10.2
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 47.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/tood/tood_x101_64x4d_fpn_mstrain_2x_coco/tood_x101_64x4d_fpn_mstrain_2x_coco_20211211_003519-a4f36113.pth
+
+ - Name: tood_r101-dconv-c3-c5_fpn_ms-2x_coco
+ In Collection: TOOD
+ Config: configs/tood/tood_r101-dconv-c3-c5_fpn_ms-2x_coco.py
+ Metadata:
+ Training Memory (GB): 6.2
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 49.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/tood/tood_r101_fpn_dconv_c3-c5_mstrain_2x_coco/tood_r101_fpn_dconv_c3-c5_mstrain_2x_coco_20211210_213728-4a824142.pth
+
+ - Name: tood_r50_fpn_anchor-based_1x_coco
+ In Collection: TOOD
+ Config: configs/tood/tood_r50_fpn_anchor-based_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.1
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/tood/tood_r50_fpn_anchor_based_1x_coco/tood_r50_fpn_anchor_based_1x_coco_20211214_100105-b776c134.pth
+
+ - Name: tood_r50_fpn_1x_coco
+ In Collection: TOOD
+ Config: configs/tood/tood_r50_fpn_1x_coco.py
+ Metadata:
+ Training Memory (GB): 4.1
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 42.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/tood/tood_r50_fpn_1x_coco/tood_r50_fpn_1x_coco_20211210_103425-20e20746.pth
+
+ - Name: tood_r50_fpn_ms-2x_coco
+ In Collection: TOOD
+ Config: configs/tood/tood_r50_fpn_ms-2x_coco.py
+ Metadata:
+ Training Memory (GB): 4.1
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/tood/tood_r50_fpn_mstrain_2x_coco/tood_r50_fpn_mstrain_2x_coco_20211210_144231-3b23174c.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_r101-dconv-c3-c5_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_r101-dconv-c3-c5_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..45030a6832db39a329d0901dde4a5320f34a9b6e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_r101-dconv-c3-c5_fpn_ms-2x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './tood_r101_fpn_ms-2x_coco.py'
+
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCNv2', deformable_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)),
+ bbox_head=dict(num_dcn=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_r101_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_r101_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..fc6ae5d942e05ac90162ca9ac67adb311d581e5b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_r101_fpn_ms-2x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './tood_r50_fpn_ms-2x_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e4839d9d77e64d61b504ed8789bda225cc878da1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_r50_fpn_1x_coco.py
@@ -0,0 +1,80 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+# model settings
+model = dict(
+ type='TOOD',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output',
+ num_outs=5),
+ bbox_head=dict(
+ type='TOODHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=6,
+ feat_channels=256,
+ anchor_type='anchor_free',
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ octave_base_scale=8,
+ scales_per_octave=1,
+ strides=[8, 16, 32, 64, 128]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ initial_loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ activated=True, # use probability instead of logit as input
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_cls=dict(
+ type='QualityFocalLoss',
+ use_sigmoid=True,
+ activated=True, # use probability instead of logit as input
+ beta=2.0,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=2.0)),
+ train_cfg=dict(
+ initial_epoch=4,
+ initial_assigner=dict(type='ATSSAssigner', topk=9),
+ assigner=dict(type='TaskAlignedAssigner', topk=13),
+ alpha=1,
+ beta=6,
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_r50_fpn_anchor-based_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_r50_fpn_anchor-based_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c7fbf6aff197b821de07f8d4a73f9c72e5f76288
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_r50_fpn_anchor-based_1x_coco.py
@@ -0,0 +1,2 @@
+_base_ = './tood_r50_fpn_1x_coco.py'
+model = dict(bbox_head=dict(anchor_type='anchor_based'))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_r50_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_r50_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ffb296dccee30438977bac61b970f5844d647cfa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_r50_fpn_ms-2x_coco.py
@@ -0,0 +1,30 @@
+_base_ = './tood_r50_fpn_1x_coco.py'
+max_epochs = 24
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
+
+# training schedule for 2x
+train_cfg = dict(max_epochs=max_epochs)
+
+# multi-scale training
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize', scale=[(1333, 480), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_x101-64x4d-dconv-c4-c5_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_x101-64x4d-dconv-c4-c5_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..43405196184715923bb22499958c74fe9bf4a2da
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_x101-64x4d-dconv-c4-c5_fpn_ms-2x_coco.py
@@ -0,0 +1,7 @@
+_base_ = './tood_x101-64x4d_fpn_ms-2x_coco.py'
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCNv2', deformable_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, False, True, True),
+ ),
+ bbox_head=dict(num_dcn=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_x101-64x4d_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_x101-64x4d_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1651542c7562553f206ba763fb9a43838e042450
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tood/tood_x101-64x4d_fpn_ms-2x_coco.py
@@ -0,0 +1,16 @@
+_base_ = './tood_r50_fpn_ms-2x_coco.py'
+
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/tridentnet/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tridentnet/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..b972b3a3c9b2de5409af9f76622e8947fd6eace1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tridentnet/README.md
@@ -0,0 +1,38 @@
+# TridentNet
+
+> [Scale-Aware Trident Networks for Object Detection](https://arxiv.org/abs/1901.01892)
+
+
+
+## Abstract
+
+Scale variation is one of the key challenges in object detection. In this work, we first present a controlled experiment to investigate the effect of receptive fields for scale variation in object detection. Based on the findings from the exploration experiments, we propose a novel Trident Network (TridentNet) aiming to generate scale-specific feature maps with a uniform representational power. We construct a parallel multi-branch architecture in which each branch shares the same transformation parameters but with different receptive fields. Then, we adopt a scale-aware training scheme to specialize each branch by sampling object instances of proper scales for training. As a bonus, a fast approximation version of TridentNet could achieve significant improvements without any additional parameters and computational cost compared with the vanilla detector. On the COCO dataset, our TridentNet with ResNet-101 backbone achieves state-of-the-art single-model results of 48.4 mAP.
+
+
+

+
+
+## Results and Models
+
+We reports the test results using only one branch for inference.
+
+| Backbone | Style | mstrain | Lr schd | Mem (GB) | Inf time (fps) | box AP | Download |
+| :------: | :---: | :-----: | :-----: | :------: | :------------: | :----: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | caffe | N | 1x | | | 37.7 | [model](https://download.openmmlab.com/mmdetection/v2.0/tridentnet/tridentnet_r50_caffe_1x_coco/tridentnet_r50_caffe_1x_coco_20201230_141838-2ec0b530.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/tridentnet/tridentnet_r50_caffe_1x_coco/tridentnet_r50_caffe_1x_coco_20201230_141838.log.json) |
+| R-50 | caffe | Y | 1x | | | 37.6 | [model](https://download.openmmlab.com/mmdetection/v2.0/tridentnet/tridentnet_r50_caffe_mstrain_1x_coco/tridentnet_r50_caffe_mstrain_1x_coco_20201230_141839-6ce55ccb.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/tridentnet/tridentnet_r50_caffe_mstrain_1x_coco/tridentnet_r50_caffe_mstrain_1x_coco_20201230_141839.log.json) |
+| R-50 | caffe | Y | 3x | | | 40.3 | [model](https://download.openmmlab.com/mmdetection/v2.0/tridentnet/tridentnet_r50_caffe_mstrain_3x_coco/tridentnet_r50_caffe_mstrain_3x_coco_20201130_100539-46d227ba.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/tridentnet/tridentnet_r50_caffe_mstrain_3x_coco/tridentnet_r50_caffe_mstrain_3x_coco_20201130_100539.log.json) |
+
+**Note**
+
+Similar to [Detectron2](https://github.com/facebookresearch/detectron2/tree/master/projects/TridentNet), we haven't implemented the Scale-aware Training Scheme in section 4.2 of the paper.
+
+## Citation
+
+```latex
+@InProceedings{li2019scale,
+ title={Scale-Aware Trident Networks for Object Detection},
+ author={Li, Yanghao and Chen, Yuntao and Wang, Naiyan and Zhang, Zhaoxiang},
+ journal={The International Conference on Computer Vision (ICCV)},
+ year={2019}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/tridentnet/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tridentnet/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..c0081c5be02986efbfdad9f199aa8ccd4b599d0f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tridentnet/metafile.yml
@@ -0,0 +1,55 @@
+Collections:
+ - Name: TridentNet
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - ResNet
+ - TridentNet Block
+ Paper:
+ URL: https://arxiv.org/abs/1901.01892
+ Title: 'Scale-Aware Trident Networks for Object Detection'
+ README: configs/tridentnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.8.0/mmdet/models/detectors/trident_faster_rcnn.py#L6
+ Version: v2.8.0
+
+Models:
+ - Name: tridentnet_r50-caffe_1x_coco
+ In Collection: TridentNet
+ Config: configs/tridentnet/tridentnet_r50-caffe_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/tridentnet/tridentnet_r50_caffe_1x_coco/tridentnet_r50_caffe_1x_coco_20201230_141838-2ec0b530.pth
+
+ - Name: tridentnet_r50-caffe_ms-1x_coco
+ In Collection: TridentNet
+ Config: configs/tridentnet/tridentnet_r50-caffe_ms-1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/tridentnet/tridentnet_r50_caffe_mstrain_1x_coco/tridentnet_r50_caffe_mstrain_1x_coco_20201230_141839-6ce55ccb.pth
+
+ - Name: tridentnet_r50-caffe_ms-3x_coco
+ In Collection: TridentNet
+ Config: configs/tridentnet/tridentnet_r50-caffe_ms-3x_coco.py
+ Metadata:
+ Epochs: 36
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.3
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/tridentnet/tridentnet_r50_caffe_mstrain_3x_coco/tridentnet_r50_caffe_mstrain_3x_coco_20201130_100539-46d227ba.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/tridentnet/tridentnet_r50-caffe_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tridentnet/tridentnet_r50-caffe_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..26a4c12316ee80c7dfae1624af3f4146dba0a414
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tridentnet/tridentnet_r50-caffe_1x_coco.py
@@ -0,0 +1,22 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50-caffe-c4.py',
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+
+model = dict(
+ type='TridentFasterRCNN',
+ backbone=dict(
+ type='TridentResNet',
+ trident_dilations=(1, 2, 3),
+ num_branch=3,
+ test_branch_idx=1,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')),
+ roi_head=dict(type='TridentRoIHead', num_branch=3, test_branch_idx=1),
+ train_cfg=dict(
+ rpn_proposal=dict(max_per_img=500),
+ rcnn=dict(
+ sampler=dict(num=128, pos_fraction=0.5,
+ add_gt_as_proposals=False))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/tridentnet/tridentnet_r50-caffe_ms-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tridentnet/tridentnet_r50-caffe_ms-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..806d20b90c96be9357eccd9f9ca8c880b0716cae
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tridentnet/tridentnet_r50-caffe_ms-1x_coco.py
@@ -0,0 +1,15 @@
+_base_ = 'tridentnet_r50-caffe_1x_coco.py'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/tridentnet/tridentnet_r50-caffe_ms-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tridentnet/tridentnet_r50-caffe_ms-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..4de249c60c234a9d301658594f7b072b0b48017b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/tridentnet/tridentnet_r50-caffe_ms-3x_coco.py
@@ -0,0 +1,18 @@
+_base_ = 'tridentnet_r50-caffe_ms-1x_coco.py'
+
+# learning rate
+max_epochs = 36
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[28, 34],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..36879316f4fe0066707fecb95af4329852fe55fc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/README.md
@@ -0,0 +1,86 @@
+
+
+
+#

V3Det: Vast Vocabulary Visual Detection Dataset
+
+
+
Jiaqi Wang*,
+
Pan Zhang*,
+ Tao Chu*,
+ Yuhang Cao*,
+ Yujie Zhou,
+
Tong Wu,
+ Bin Wang,
+ Conghui He,
+
Dahua Lin
+ (* equal contribution)
+
Accepted to ICCV 2023 (Oral)
+
+
+
+
+
+
+
+
+

+
+
+
+
+## Abstract
+
+Recent advances in detecting arbitrary objects in the real world are trained and evaluated on object detection datasets with a relatively restricted vocabulary. To facilitate the development of more general visual object detection, we propose V3Det, a vast vocabulary visual detection dataset with precisely annotated bounding boxes on massive images. V3Det has several appealing properties: 1) Vast Vocabulary: It contains bounding boxes of objects from 13,204 categories on real-world images, which is 10 times larger than the existing large vocabulary object detection dataset, e.g., LVIS. 2) Hierarchical Category Organization: The vast vocabulary of V3Det is organized by a hierarchical category tree which annotates the inclusion relationship among categories, encouraging the exploration of category relationships in vast and open vocabulary object detection. 3) Rich Annotations: V3Det comprises precisely annotated objects in 243k images and professional descriptions of each category written by human experts and a powerful chatbot. By offering a vast exploration space, V3Det enables extensive benchmarks on both vast and open vocabulary object detection, leading to new observations, practices, and insights for future research. It has the potential to serve as a cornerstone dataset for developing more general visual perception systems. V3Det is available at https://v3det.openxlab.org.cn/.
+
+## Prepare Dataset
+
+Please download and prepare V3Det Dataset at [V3Det Homepage](https://v3det.openxlab.org.cn/) and [V3Det Github](https://github.com/V3Det/V3Det).
+
+The data includes a training set, a validation set, comprising 13,204 categories. The training set consists of 183,354 images, while the validation set has 29,821 images. The data organization is:
+
+```
+data/
+ V3Det/
+ images/
+ /
+ |────.png
+ ...
+ ...
+ annotations/
+ |────v3det_2023_v1_category_tree.json # Category tree
+ |────category_name_13204_v3det_2023_v1.txt # Category name
+ |────v3det_2023_v1_train.json # Train set
+ |────v3det_2023_v1_val.json # Validation set
+```
+
+## Results and Models
+
+| Backbone | Model | Lr schd | box AP | Config | Download |
+| :------: | :-------------: | :-----: | :----: | :----------------------------------------------------------------------------: | :-------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | Faster R-CNN | 2x | 25.4 | [config](./faster_rcnn_r50_fpn_8x4_sample1e-3_mstrain_v3det_2x.py) | [model](https://download.openxlab.org.cn/models/V3Det/V3Det/weight//faster_rcnn_r50_fpn_8x4_sample1e-3_mstrain_v3det_2x) |
+| R-50 | Cascade R-CNN | 2x | 31.6 | [config](./cascade_rcnn_r50_fpn_8x4_sample1e-3_mstrain_v3det_2x.py) | [model](https://download.openxlab.org.cn/models/V3Det/V3Det/weight//cascade_rcnn_r50_fpn_8x4_sample1e-3_mstrain_v3det_2x) |
+| R-50 | FCOS | 2x | 9.4 | [config](./fcos_r50_fpn_8x4_sample1e-3_mstrain_v3det_2x.py) | [model](https://download.openxlab.org.cn/models/V3Det/V3Det/weight//fcos_r50_fpn_8x4_sample1e-3_mstrain_v3det_2x) |
+| R-50 | Deformable-DETR | 50e | 34.4 | [config](./deformable-detr-refine-twostage_r50_8xb4_sample1e-3_v3det_50e.py) | [model](https://download.openxlab.org.cn/models/V3Det/V3Det/weight/Deformable_DETR_V3Det_R50) |
+| R-50 | DINO | 36e | 33.5 | [config](./dino-4scale_r50_8xb2_sample1e-3_v3det_36e.py) | [model](https://download.openxlab.org.cn/models/V3Det/V3Det/weight/DINO_V3Det_R50) |
+| Swin-B | Faster R-CNN | 2x | 37.6 | [config](./faster_rcnn_swinb_fpn_8x4_sample1e-3_mstrain_v3det_2x.py) | [model](https://download.openxlab.org.cn/models/V3Det/V3Det/weight//faster_rcnn_swinb_fpn_8x4_sample1e-3_mstrain_v3det_2x) |
+| Swin-B | Cascade R-CNN | 2x | 42.5 | [config](./cascade_rcnn_swinb_fpn_8x4_sample1e-3_mstrain_v3det_2x.py) | [model](https://download.openxlab.org.cn/models/V3Det/V3Det/weight//cascade_rcnn_swinb_fpn_8x4_sample1e-3_mstrain_v3det_2x) |
+| Swin-B | FCOS | 2x | 21.0 | [config](./fcos_swinb_fpn_8x4_sample1e-3_mstrain_v3det_2x.py) | [model](https://download.openxlab.org.cn/models/V3Det/V3Det/weight//fcos_swinb_fpn_8x4_sample1e-3_mstrain_v3det_2x) |
+| Swin-B | Deformable-DETR | 50e | 42.5 | [config](./deformable-detr-refine-twostage_swin_16xb2_sample1e-3_v3det_50e.py) | [model](https://download.openxlab.org.cn/models/V3Det/V3Det/weight/Deformable_DETR_V3Det_SwinB) |
+| Swin-B | DINO | 36e | 42.0 | [config](./dino-4scale_swin_16xb1_sample1e-3_v3det_36e.py) | [model](https://download.openxlab.org.cn/models/V3Det/V3Det/weight/DINO_V3Det_SwinB) |
+
+## Citation
+
+```latex
+@inproceedings{wang2023v3det,
+ title = {V3Det: Vast Vocabulary Visual Detection Dataset},
+ author = {Wang, Jiaqi and Zhang, Pan and Chu, Tao and Cao, Yuhang and Zhou, Yujie and Wu, Tong and Wang, Bin and He, Conghui and Lin, Dahua},
+ booktitle = {The IEEE International Conference on Computer Vision (ICCV)},
+ month = {October},
+ year = {2023}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/cascade_rcnn_r50_fpn_8x4_sample1e-3_mstrain_v3det_2x.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/cascade_rcnn_r50_fpn_8x4_sample1e-3_mstrain_v3det_2x.py
new file mode 100644
index 0000000000000000000000000000000000000000..567c31bd0e986e071b50ff2aac9cb896d4daf6fd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/cascade_rcnn_r50_fpn_8x4_sample1e-3_mstrain_v3det_2x.py
@@ -0,0 +1,171 @@
+_base_ = [
+ '../_base_/models/cascade-rcnn_r50_fpn.py', '../_base_/datasets/v3det.py',
+ '../_base_/schedules/schedule_2x.py', '../_base_/default_runtime.py'
+]
+# model settings
+model = dict(
+ rpn_head=dict(
+ loss_bbox=dict(_delete_=True, type='L1Loss', loss_weight=1.0)),
+ roi_head=dict(bbox_head=[
+ dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=13204,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=True,
+ cls_predictor_cfg=dict(
+ type='NormedLinear', tempearture=50, bias=True),
+ loss_cls=dict(
+ type='CrossEntropyCustomLoss',
+ num_classes=13204,
+ use_sigmoid=True,
+ loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0)),
+ dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=13204,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.05, 0.05, 0.1, 0.1]),
+ reg_class_agnostic=True,
+ cls_predictor_cfg=dict(
+ type='NormedLinear', tempearture=50, bias=True),
+ loss_cls=dict(
+ type='CrossEntropyCustomLoss',
+ num_classes=13204,
+ use_sigmoid=True,
+ loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0)),
+ dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=13204,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.033, 0.033, 0.067, 0.067]),
+ reg_class_agnostic=True,
+ cls_predictor_cfg=dict(
+ type='NormedLinear', tempearture=50, bias=True),
+ loss_cls=dict(
+ type='CrossEntropyCustomLoss',
+ num_classes=13204,
+ use_sigmoid=True,
+ loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0))
+ ]),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn_proposal=dict(nms_pre=4000, max_per_img=2000),
+ rcnn=[
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ match_low_quality=False,
+ ignore_iof_thr=-1,
+ perm_repeat_gt_cfg=dict(iou_thr=0.7, perm_range=0.01)),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ pos_weight=-1,
+ debug=False),
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.6,
+ neg_iou_thr=0.6,
+ min_pos_iou=0.6,
+ match_low_quality=False,
+ ignore_iof_thr=-1,
+ perm_repeat_gt_cfg=dict(iou_thr=0.7, perm_range=0.01)),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ pos_weight=-1,
+ debug=False),
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.7,
+ min_pos_iou=0.7,
+ match_low_quality=False,
+ ignore_iof_thr=-1,
+ perm_repeat_gt_cfg=dict(iou_thr=0.7, perm_range=0.01)),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ pos_weight=-1,
+ debug=False)
+ ]),
+ test_cfg=dict(
+ rcnn=dict(
+ score_thr=0.0001,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=300)))
+# dataset settings
+train_dataloader = dict(batch_size=4, num_workers=8)
+
+# training schedule for 1x
+max_iter = 68760 * 2
+train_cfg = dict(
+ _delete_=True,
+ type='IterBasedTrainLoop',
+ max_iters=max_iter,
+ val_interval=max_iter)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 2048,
+ by_epoch=False,
+ begin=0,
+ end=5000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_iter,
+ by_epoch=False,
+ milestones=[45840 * 2, 63030 * 2],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(_delete_=True, type='AdamW', lr=1e-4 * 1, weight_decay=0.1),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=32)
+
+default_hooks = dict(
+ checkpoint=dict(type='CheckpointHook', by_epoch=False, interval=5730 * 2))
+log_processor = dict(type='LogProcessor', window_size=50, by_epoch=False)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/cascade_rcnn_swinb_fpn_8x4_sample1e-3_mstrain_v3det_2x.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/cascade_rcnn_swinb_fpn_8x4_sample1e-3_mstrain_v3det_2x.py
new file mode 100644
index 0000000000000000000000000000000000000000..f6493323ba8d92d2628fb4784f5a12dd564460be
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/cascade_rcnn_swinb_fpn_8x4_sample1e-3_mstrain_v3det_2x.py
@@ -0,0 +1,27 @@
+_base_ = [
+ './cascade_rcnn_r50_fpn_8x4_sample1e-3_mstrain_v3det_2x.py',
+]
+
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_base_patch4_window7_224.pth' # noqa
+
+# model settings
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ embed_dims=128,
+ depths=[2, 2, 18, 2],
+ num_heads=[4, 8, 16, 32],
+ window_size=7,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(0, 1, 2, 3),
+ with_cp=False,
+ convert_weights=True,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ neck=dict(in_channels=[128, 256, 512, 1024]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/deformable-detr-refine-twostage_r50_8xb4_sample1e-3_v3det_50e.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/deformable-detr-refine-twostage_r50_8xb4_sample1e-3_v3det_50e.py
new file mode 100644
index 0000000000000000000000000000000000000000..97544a27edfd75eef4ba25fd12a122f03b392c1f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/deformable-detr-refine-twostage_r50_8xb4_sample1e-3_v3det_50e.py
@@ -0,0 +1,108 @@
+_base_ = '../deformable_detr/deformable-detr-refine-twostage_r50_16xb2-50e_coco.py' # noqa
+
+model = dict(
+ bbox_head=dict(num_classes=13204),
+ test_cfg=dict(max_per_img=300),
+)
+
+data_root = 'data/V3Det/'
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(
+ _delete_=True,
+ batch_size=4,
+ num_workers=4,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type='ClassBalancedDataset',
+ oversample_thr=1e-3,
+ dataset=dict(
+ type='V3DetDataset',
+ data_root=data_root,
+ ann_file='annotations/v3det_2023_v1_train.json',
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=train_pipeline,
+ backend_args=None)))
+val_dataloader = dict(
+ dataset=dict(
+ type='V3DetDataset',
+ data_root=data_root,
+ ann_file='annotations/v3det_2023_v1_val.json',
+ data_prefix=dict(img='')))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ ann_file=data_root + 'annotations/v3det_2023_v1_val.json',
+ use_mp_eval=True,
+ proposal_nums=[300])
+test_evaluator = val_evaluator
+
+# training schedule for 50e
+# when using RFS, bs32, each epoch ~ 5730 iter
+max_iter = 286500
+train_cfg = dict(
+ _delete_=True,
+ type='IterBasedTrainLoop',
+ max_iters=max_iter,
+ val_interval=max_iter / 5)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_iter,
+ by_epoch=False,
+ milestones=[229200], # 40e
+ gamma=0.1)
+]
+
+default_hooks = dict(
+ timer=dict(type='IterTimerHook'),
+ param_scheduler=dict(type='ParamSchedulerHook'),
+ checkpoint=dict(
+ type='CheckpointHook', by_epoch=False, interval=5730,
+ max_keep_ckpts=3))
+
+log_processor = dict(type='LogProcessor', window_size=50, by_epoch=False)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/deformable-detr-refine-twostage_swin_16xb2_sample1e-3_v3det_50e.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/deformable-detr-refine-twostage_swin_16xb2_sample1e-3_v3det_50e.py
new file mode 100644
index 0000000000000000000000000000000000000000..e640cd604a97813a70588d5ffe23701543ab0087
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/deformable-detr-refine-twostage_swin_16xb2_sample1e-3_v3det_50e.py
@@ -0,0 +1,27 @@
+_base_ = 'deformable-detr-refine-twostage_r50_8xb4_sample1e-3_v3det_50e.py'
+
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_base_patch4_window7_224.pth' # noqa
+
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ embed_dims=128,
+ depths=[2, 2, 18, 2],
+ num_heads=[4, 8, 16, 32],
+ window_size=7,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(1, 2, 3),
+ with_cp=False,
+ convert_weights=True,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ neck=dict(in_channels=[256, 512, 1024]),
+)
+
+train_dataloader = dict(batch_size=2, num_workers=2)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/dino-4scale_r50_8xb2_sample1e-3_v3det_36e.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/dino-4scale_r50_8xb2_sample1e-3_v3det_36e.py
new file mode 100644
index 0000000000000000000000000000000000000000..d9e6e6be0715512b111171c4b60cca7433f8ca34
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/dino-4scale_r50_8xb2_sample1e-3_v3det_36e.py
@@ -0,0 +1,109 @@
+_base_ = '../dino/dino-4scale_r50_8xb2-36e_coco.py'
+
+model = dict(
+ bbox_head=dict(num_classes=13204),
+ test_cfg=dict(max_per_img=300),
+)
+
+data_root = 'data/V3Det/'
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(
+ _delete_=True,
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type='ClassBalancedDataset',
+ oversample_thr=1e-3,
+ dataset=dict(
+ type='V3DetDataset',
+ data_root=data_root,
+ ann_file='annotations/v3det_2023_v1_train.json',
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=False),
+ pipeline=train_pipeline,
+ backend_args=None)))
+val_dataloader = dict(
+ dataset=dict(
+ type='V3DetDataset',
+ data_root=data_root,
+ ann_file='annotations/v3det_2023_v1_val.json',
+ data_prefix=dict(img='')))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ ann_file=data_root + 'annotations/v3det_2023_v1_val.json',
+ use_mp_eval=True,
+ proposal_nums=[300])
+test_evaluator = val_evaluator
+
+# training schedule for 36e
+# when using RFS, bs16, each epoch ~ 11460 iter
+max_iter = 412560
+train_cfg = dict(
+ _delete_=True,
+ type='IterBasedTrainLoop',
+ max_iters=max_iter,
+ val_interval=max_iter / 5)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_iter,
+ by_epoch=False,
+ milestones=[343800], # 30e
+ gamma=0.1)
+]
+
+default_hooks = dict(
+ timer=dict(type='IterTimerHook'),
+ param_scheduler=dict(type='ParamSchedulerHook'),
+ checkpoint=dict(
+ type='CheckpointHook',
+ by_epoch=False,
+ interval=11460,
+ max_keep_ckpts=3))
+
+log_processor = dict(type='LogProcessor', window_size=50, by_epoch=False)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/dino-4scale_swin_16xb1_sample1e-3_v3det_36e.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/dino-4scale_swin_16xb1_sample1e-3_v3det_36e.py
new file mode 100644
index 0000000000000000000000000000000000000000..100c4ba4b8cb2c0ac3e44f5e9ddcfc37bbfe6b55
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/dino-4scale_swin_16xb1_sample1e-3_v3det_36e.py
@@ -0,0 +1,27 @@
+_base_ = 'dino-4scale_r50_8xb2_sample1e-3_v3det_36e.py'
+
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_base_patch4_window7_224.pth' # noqa
+
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ embed_dims=128,
+ depths=[2, 2, 18, 2],
+ num_heads=[4, 8, 16, 32],
+ window_size=7,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(1, 2, 3),
+ with_cp=False,
+ convert_weights=True,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ neck=dict(in_channels=[256, 512, 1024]),
+)
+
+train_dataloader = dict(batch_size=1)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/faster_rcnn_r50_fpn_8x4_sample1e-3_mstrain_v3det_2x.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/faster_rcnn_r50_fpn_8x4_sample1e-3_mstrain_v3det_2x.py
new file mode 100644
index 0000000000000000000000000000000000000000..3d306fb094806d75ec614b52a43bf6614d13eed4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/faster_rcnn_r50_fpn_8x4_sample1e-3_mstrain_v3det_2x.py
@@ -0,0 +1,72 @@
+_base_ = [
+ '../_base_/models/faster-rcnn_r50_fpn.py', '../_base_/datasets/v3det.py',
+ '../_base_/schedules/schedule_2x.py', '../_base_/default_runtime.py'
+]
+# model settings
+model = dict(
+ roi_head=dict(
+ bbox_head=dict(
+ num_classes=13204,
+ reg_class_agnostic=True,
+ cls_predictor_cfg=dict(
+ type='NormedLinear', tempearture=50, bias=True),
+ loss_cls=dict(
+ type='CrossEntropyCustomLoss',
+ num_classes=13204,
+ use_sigmoid=True,
+ loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn_proposal=dict(nms_pre=4000, max_per_img=2000),
+ rcnn=dict(
+ assigner=dict(
+ perm_repeat_gt_cfg=dict(iou_thr=0.7, perm_range=0.01)))),
+ test_cfg=dict(
+ rcnn=dict(
+ score_thr=0.0001,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=300)))
+# dataset settings
+train_dataloader = dict(batch_size=4, num_workers=8)
+
+# training schedule for 2x
+max_iter = 68760 * 2
+train_cfg = dict(
+ _delete_=True,
+ type='IterBasedTrainLoop',
+ max_iters=max_iter,
+ val_interval=max_iter)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 2048,
+ by_epoch=False,
+ begin=0,
+ end=5000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_iter,
+ by_epoch=False,
+ milestones=[45840 * 2, 63030 * 2],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(_delete_=True, type='AdamW', lr=1e-4 * 1, weight_decay=0.1),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=32)
+
+default_hooks = dict(
+ checkpoint=dict(type='CheckpointHook', by_epoch=False, interval=5730 * 2))
+log_processor = dict(type='LogProcessor', window_size=50, by_epoch=False)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/faster_rcnn_swinb_fpn_8x4_sample1e-3_mstrain_v3det_2x.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/faster_rcnn_swinb_fpn_8x4_sample1e-3_mstrain_v3det_2x.py
new file mode 100644
index 0000000000000000000000000000000000000000..b0b1110811230b4bda27da9fd2e58067c7326c52
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/faster_rcnn_swinb_fpn_8x4_sample1e-3_mstrain_v3det_2x.py
@@ -0,0 +1,27 @@
+_base_ = [
+ './faster_rcnn_r50_fpn_8x4_sample1e-3_mstrain_v3det_2x.py',
+]
+
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_base_patch4_window7_224.pth' # noqa
+
+# model settings
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ embed_dims=128,
+ depths=[2, 2, 18, 2],
+ num_heads=[4, 8, 16, 32],
+ window_size=7,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(0, 1, 2, 3),
+ with_cp=False,
+ convert_weights=True,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ neck=dict(in_channels=[128, 256, 512, 1024]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/fcos_r50_fpn_8x4_sample1e-3_mstrain_v3det_2x.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/fcos_r50_fpn_8x4_sample1e-3_mstrain_v3det_2x.py
new file mode 100644
index 0000000000000000000000000000000000000000..b78e38c93cb0fdedff3948f1ce7b5b7787efcaea
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/fcos_r50_fpn_8x4_sample1e-3_mstrain_v3det_2x.py
@@ -0,0 +1,116 @@
+_base_ = [
+ '../_base_/datasets/v3det.py', '../_base_/schedules/schedule_2x.py',
+ '../_base_/default_runtime.py'
+]
+# model settings
+model = dict(
+ type='FCOS',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output', # use P5
+ num_outs=5,
+ relu_before_extra_convs=True),
+ bbox_head=dict(
+ type='FCOSHead',
+ num_classes=13204,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ strides=[8, 16, 32, 64, 128],
+ cls_predictor_cfg=dict(type='NormedLinear', tempearture=50, bias=True),
+ loss_cls=dict(
+ type='FocalCustomLoss',
+ use_sigmoid=True,
+ num_classes=13204,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='IoULoss', loss_weight=1.0),
+ loss_centerness=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0)),
+ # model training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.4,
+ min_pos_iou=0,
+ ignore_iof_thr=-1,
+ perm_repeat_gt_cfg=dict(iou_thr=0.7, perm_range=0.01)),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.0001,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=300))
+# dataset settings
+
+backend_args = None
+
+train_dataloader = dict(batch_size=2, num_workers=8)
+
+# training schedule for 2x
+max_iter = 68760 * 2 * 2
+train_cfg = dict(
+ _delete_=True,
+ type='IterBasedTrainLoop',
+ max_iters=max_iter,
+ val_interval=max_iter)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=1.0 / 2048,
+ by_epoch=False,
+ begin=0,
+ end=5000 * 2),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_iter,
+ by_epoch=False,
+ milestones=[45840 * 2 * 2, 63030 * 2 * 2],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(
+ _delete_=True, type='AdamW', lr=1e-4 * 0.25, weight_decay=0.1),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=32)
+
+default_hooks = dict(
+ checkpoint=dict(type='CheckpointHook', by_epoch=False, interval=5730 * 2))
+log_processor = dict(type='LogProcessor', window_size=50, by_epoch=False)
+
+find_unused_parameters = True
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/fcos_swinb_fpn_8x4_sample1e-3_mstrain_v3det_2x.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/fcos_swinb_fpn_8x4_sample1e-3_mstrain_v3det_2x.py
new file mode 100644
index 0000000000000000000000000000000000000000..6ca952a28fc08ae9b14ad30308eff823b1bba55e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/fcos_swinb_fpn_8x4_sample1e-3_mstrain_v3det_2x.py
@@ -0,0 +1,27 @@
+_base_ = [
+ './fcos_r50_fpn_8x4_sample1e-3_mstrain_v3det_2x.py',
+]
+
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_base_patch4_window7_224.pth' # noqa
+
+# model settings
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ embed_dims=128,
+ depths=[2, 2, 18, 2],
+ num_heads=[4, 8, 16, 32],
+ window_size=7,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.0,
+ attn_drop_rate=0.0,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(0, 1, 2, 3),
+ with_cp=False,
+ convert_weights=True,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ neck=dict(in_channels=[128, 256, 512, 1024], force_grad_on_level=True))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/v3det_icon.jpg b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/v3det_icon.jpg
new file mode 100644
index 0000000000000000000000000000000000000000..b25be6fd1f2dd1ddcba0739004e836fb80081656
Binary files /dev/null and b/scripts/evaluation/FasterRCNN_score-mmdet/configs/v3det/v3det_icon.jpg differ
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..73b5c07be9e9eb3419fd363a5becf5f3c2b91641
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/README.md
@@ -0,0 +1,48 @@
+# VarifocalNet
+
+> [VarifocalNet: An IoU-aware Dense Object Detector](https://arxiv.org/abs/2008.13367)
+
+
+
+## Abstract
+
+Accurately ranking the vast number of candidate detections is crucial for dense object detectors to achieve high performance. Prior work uses the classification score or a combination of classification and predicted localization scores to rank candidates. However, neither option results in a reliable ranking, thus degrading detection performance. In this paper, we propose to learn an Iou-aware Classification Score (IACS) as a joint representation of object presence confidence and localization accuracy. We show that dense object detectors can achieve a more accurate ranking of candidate detections based on the IACS. We design a new loss function, named Varifocal Loss, to train a dense object detector to predict the IACS, and propose a new star-shaped bounding box feature representation for IACS prediction and bounding box refinement. Combining these two new components and a bounding box refinement branch, we build an IoU-aware dense object detector based on the FCOS+ATSS architecture, that we call VarifocalNet or VFNet for short. Extensive experiments on MS COCO show that our VFNet consistently surpasses the strong baseline by ∼2.0 AP with different backbones. Our best model VFNet-X-1200 with Res2Net-101-DCN achieves a single-model single-scale AP of 55.1 on COCO test-dev, which is state-of-the-art among various object detectors.
+
+
+

+
+
+## Introduction
+
+**VarifocalNet (VFNet)** learns to predict the IoU-aware classification score which mixes the object presence confidence and localization accuracy together as the detection score for a bounding box. The learning is supervised by the proposed Varifocal Loss (VFL), based on a new star-shaped bounding box feature representation (the features at nine yellow sampling points). Given the new representation, the object localization accuracy is further improved by refining the initially regressed bounding box. The full paper is available at: [https://arxiv.org/abs/2008.13367](https://arxiv.org/abs/2008.13367).
+
+## Results and Models
+
+| Backbone | Style | DCN | MS train | Lr schd | Inf time (fps) | box AP (val) | box AP (test-dev) | Config | Download |
+| :---------: | :-----: | :-: | :------: | :-----: | :------------: | :----------: | :---------------: | :---------------------------------------------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | pytorch | N | N | 1x | - | 41.6 | 41.6 | [config](./vfnet_r50_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_r50_fpn_1x_coco/vfnet_r50_fpn_1x_coco_20201027-38db6f58.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_r50_fpn_1x_coco/vfnet_r50_fpn_1x_coco.json) |
+| R-50 | pytorch | N | Y | 2x | - | 44.5 | 44.8 | [config](./vfnet_r50_fpn_ms-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_r50_fpn_mstrain_2x_coco/vfnet_r50_fpn_mstrain_2x_coco_20201027-7cc75bd2.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_r50_fpn_mstrain_2x_coco/vfnet_r50_fpn_mstrain_2x_coco.json) |
+| R-50 | pytorch | Y | Y | 2x | - | 47.8 | 48.0 | [config](./vfnet_r50-mdconv-c3-c5_fpn_ms-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_r50_fpn_mdconv_c3-c5_mstrain_2x_coco/vfnet_r50_fpn_mdconv_c3-c5_mstrain_2x_coco_20201027pth-6879c318.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_r50_fpn_mdconv_c3-c5_mstrain_2x_coco/vfnet_r50_fpn_mdconv_c3-c5_mstrain_2x_coco.json) |
+| R-101 | pytorch | N | N | 1x | - | 43.0 | 43.6 | [config](./vfnet_r101_fpn_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_r101_fpn_1x_coco/vfnet_r101_fpn_1x_coco_20201027pth-c831ece7.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_r101_fpn_1x_coco/vfnet_r101_fpn_1x_coco.json) |
+| R-101 | pytorch | N | Y | 2x | - | 46.2 | 46.7 | [config](./vfnet_r101_fpn_ms-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_r101_fpn_mstrain_2x_coco/vfnet_r101_fpn_mstrain_2x_coco_20201027pth-4a5d53f1.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_r101_fpn_mstrain_2x_coco/vfnet_r101_fpn_mstrain_2x_coco.json) |
+| R-101 | pytorch | Y | Y | 2x | - | 49.0 | 49.2 | [config](./vfnet_r101-mdconv-c3-c5_fpn_ms-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_r101_fpn_mdconv_c3-c5_mstrain_2x_coco/vfnet_r101_fpn_mdconv_c3-c5_mstrain_2x_coco_20201027pth-7729adb5.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_r101_fpn_mdconv_c3-c5_mstrain_2x_coco/vfnet_r101_fpn_mdconv_c3-c5_mstrain_2x_coco.json) |
+| X-101-32x4d | pytorch | Y | Y | 2x | - | 49.7 | 50.0 | [config](./vfnet_x101-32x4d-mdconv-c3-c5_fpn_ms-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_x101_32x4d_fpn_mdconv_c3-c5_mstrain_2x_coco/vfnet_x101_32x4d_fpn_mdconv_c3-c5_mstrain_2x_coco_20201027pth-d300a6fc.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_x101_32x4d_fpn_mdconv_c3-c5_mstrain_2x_coco/vfnet_x101_32x4d_fpn_mdconv_c3-c5_mstrain_2x_coco.json) |
+| X-101-64x4d | pytorch | Y | Y | 2x | - | 50.4 | 50.8 | [config](./vfnet_x101-64x4d-mdconv-c3-c5_fpn_ms-2x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_x101_64x4d_fpn_mdconv_c3-c5_mstrain_2x_coco/vfnet_x101_64x4d_fpn_mdconv_c3-c5_mstrain_2x_coco_20201027pth-b5f6da5e.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_x101_64x4d_fpn_mdconv_c3-c5_mstrain_2x_coco/vfnet_x101_64x4d_fpn_mdconv_c3-c5_mstrain_2x_coco.json) |
+
+**Notes:**
+
+- The MS-train scale range is 1333x\[480:960\] (`range` mode) and the inference scale keeps 1333x800.
+- DCN means using `DCNv2` in both backbone and head.
+- Inference time will be updated soon.
+- More results and pre-trained models can be found in [VarifocalNet-Github](https://github.com/hyz-xmaster/VarifocalNet)
+
+## Citation
+
+```latex
+@article{zhang2020varifocalnet,
+ title={VarifocalNet: An IoU-aware Dense Object Detector},
+ author={Zhang, Haoyang and Wang, Ying and Dayoub, Feras and S{\"u}nderhauf, Niko},
+ journal={arXiv preprint arXiv:2008.13367},
+ year={2020}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..1b791d01d50ad8a28bff225fa1d3f5af8d348207
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/metafile.yml
@@ -0,0 +1,116 @@
+Collections:
+ - Name: VFNet
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - FPN
+ - ResNet
+ - Varifocal Loss
+ Paper:
+ URL: https://arxiv.org/abs/2008.13367
+ Title: 'VarifocalNet: An IoU-aware Dense Object Detector'
+ README: configs/vfnet/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.6.0/mmdet/models/detectors/vfnet.py#L6
+ Version: v2.6.0
+
+Models:
+ - Name: vfnet_r50_fpn_1x_coco
+ In Collection: VFNet
+ Config: configs/vfnet/vfnet_r50_fpn_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 41.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_r50_fpn_1x_coco/vfnet_r50_fpn_1x_coco_20201027-38db6f58.pth
+
+ - Name: vfnet_r50_fpn_ms-2x_coco
+ In Collection: VFNet
+ Config: configs/vfnet/vfnet_r50_fpn_ms-2x_coco.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 44.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_r50_fpn_mstrain_2x_coco/vfnet_r50_fpn_mstrain_2x_coco_20201027-7cc75bd2.pth
+
+ - Name: vfnet_r50-mdconv-c3-c5_fpn_ms-2x_coco
+ In Collection: VFNet
+ Config: configs/vfnet/vfnet_r50-mdconv-c3-c5_fpn_ms-2x_coco.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 48.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_r50_fpn_mdconv_c3-c5_mstrain_2x_coco/vfnet_r50_fpn_mdconv_c3-c5_mstrain_2x_coco_20201027pth-6879c318.pth
+
+ - Name: vfnet_r101_fpn_1x_coco
+ In Collection: VFNet
+ Config: configs/vfnet/vfnet_r101_fpn_1x_coco.py
+ Metadata:
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 43.6
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_r101_fpn_1x_coco/vfnet_r101_fpn_1x_coco_20201027pth-c831ece7.pth
+
+ - Name: vfnet_r101_fpn_ms-2x_coco
+ In Collection: VFNet
+ Config: configs/vfnet/vfnet_r101_fpn_ms-2x_coco.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 46.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_r101_fpn_mstrain_2x_coco/vfnet_r101_fpn_mstrain_2x_coco_20201027pth-4a5d53f1.pth
+
+ - Name: vfnet_r101-mdconv-c3-c5_fpn_ms-2x_coco
+ In Collection: VFNet
+ Config: configs/vfnet/vfnet_r101-mdconv-c3-c5_fpn_ms-2x_coco.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 49.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_r101_fpn_mdconv_c3-c5_mstrain_2x_coco/vfnet_r101_fpn_mdconv_c3-c5_mstrain_2x_coco_20201027pth-7729adb5.pth
+
+ - Name: vfnet_x101-32x4d-mdconv-c3-c5_fpn_ms-2x_coco
+ In Collection: VFNet
+ Config: configs/vfnet/vfnet_x101-32x4d-mdconv-c3-c5_fpn_ms-2x_coco.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 50.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_x101_32x4d_fpn_mdconv_c3-c5_mstrain_2x_coco/vfnet_x101_32x4d_fpn_mdconv_c3-c5_mstrain_2x_coco_20201027pth-d300a6fc.pth
+
+ - Name: vfnet_x101-64x4d-mdconv-c3-c5_fpn_ms-2x_coco
+ In Collection: VFNet
+ Config: configs/vfnet/vfnet_x101-64x4d-mdconv-c3-c5_fpn_ms-2x_coco.py
+ Metadata:
+ Epochs: 24
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 50.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/vfnet/vfnet_x101_64x4d_fpn_mdconv_c3-c5_mstrain_2x_coco/vfnet_x101_64x4d_fpn_mdconv_c3-c5_mstrain_2x_coco_20201027pth-b5f6da5e.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r101-mdconv-c3-c5_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r101-mdconv-c3-c5_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..2dd67a3bcce3bbb66531997133880d65af0c856a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r101-mdconv-c3-c5_fpn_ms-2x_coco.py
@@ -0,0 +1,15 @@
+_base_ = './vfnet_r50-mdconv-c3-c5_fpn_ms-2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNet',
+ depth=101,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ dcn=dict(type='DCNv2', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True),
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b296a07959e43517d792f36f356404a232fb0dc3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r101_fpn_1x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './vfnet_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r101_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r101_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..37a7bacb5e409a75ae2cd71fc022837f09537aa7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r101_fpn_2x_coco.py
@@ -0,0 +1,20 @@
+_base_ = './vfnet_r50_fpn_1x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
+# learning policy
+max_epochs = 24
+param_scheduler = [
+ dict(type='LinearLR', start_factor=0.1, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
+
+train_cfg = dict(max_epochs=max_epochs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r101_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r101_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..62f064b7473f4e6fec3ac50962240ac1f828753f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r101_fpn_ms-2x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './vfnet_r50_fpn_ms-2x_coco.py'
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r50-mdconv-c3-c5_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r50-mdconv-c3-c5_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..08adf927599b7759dea0e2d14c37ce716482b301
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r50-mdconv-c3-c5_fpn_ms-2x_coco.py
@@ -0,0 +1,6 @@
+_base_ = './vfnet_r50_fpn_ms-2x_coco.py'
+model = dict(
+ backbone=dict(
+ dcn=dict(type='DCNv2', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True)),
+ bbox_head=dict(dcn_on_last_conv=True))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..99bc3b5f4c78c7a7cda11e20f209ea40af7dfd80
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r50_fpn_1x_coco.py
@@ -0,0 +1,104 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+# model settings
+model = dict(
+ type='VFNet',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_output', # use P5
+ num_outs=5,
+ relu_before_extra_convs=True),
+ bbox_head=dict(
+ type='VFNetHead',
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=3,
+ feat_channels=256,
+ strides=[8, 16, 32, 64, 128],
+ center_sampling=False,
+ dcn_on_last_conv=False,
+ use_atss=True,
+ use_vfl=True,
+ loss_cls=dict(
+ type='VarifocalLoss',
+ use_sigmoid=True,
+ alpha=0.75,
+ gamma=2.0,
+ iou_weighted=True,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=1.5),
+ loss_bbox_refine=dict(type='GIoULoss', loss_weight=2.0)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(type='ATSSAssigner', topk=9),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+
+# data setting
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(lr=0.01),
+ paramwise_cfg=dict(bias_lr_mult=2., bias_decay_mult=0.),
+ clip_grad=None)
+# learning rate
+max_epochs = 12
+param_scheduler = [
+ dict(type='LinearLR', start_factor=0.1, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+
+train_cfg = dict(max_epochs=max_epochs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r50_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r50_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..0f8eed298e81967582420ac45a241b2726c47f6a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_r50_fpn_ms-2x_coco.py
@@ -0,0 +1,36 @@
+_base_ = './vfnet_r50_fpn_1x_coco.py'
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='RandomResize', scale=[(1333, 480), (1333, 960)],
+ keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+# learning policy
+max_epochs = 24
+param_scheduler = [
+ dict(type='LinearLR', start_factor=0.1, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
+
+train_cfg = dict(max_epochs=max_epochs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_res2net-101_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_res2net-101_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..94288e8e80e5be2c6e8effd38e30e239cd1e3c5f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_res2net-101_fpn_ms-2x_coco.py
@@ -0,0 +1,16 @@
+_base_ = './vfnet_r50_fpn_ms-2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='Res2Net',
+ depth=101,
+ scales=4,
+ base_width=26,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://res2net101_v1d_26w_4s')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_res2net101-mdconv-c3-c5_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_res2net101-mdconv-c3-c5_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..269330d3d8c218e51c3e65b550e4afc3296f2ec4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_res2net101-mdconv-c3-c5_fpn_ms-2x_coco.py
@@ -0,0 +1,18 @@
+_base_ = './vfnet_r50-mdconv-c3-c5_fpn_ms-2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='Res2Net',
+ depth=101,
+ scales=4,
+ base_width=26,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ dcn=dict(type='DCNv2', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True),
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://res2net101_v1d_26w_4s')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_x101-32x4d-mdconv-c3-c5_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_x101-32x4d-mdconv-c3-c5_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..465da0cbdf4c4ae34d648349f4f9fa2d3fb13fe6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_x101-32x4d-mdconv-c3-c5_fpn_ms-2x_coco.py
@@ -0,0 +1,17 @@
+_base_ = './vfnet_r50-mdconv-c3-c5_fpn_ms-2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ dcn=dict(type='DCNv2', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_x101-32x4d_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_x101-32x4d_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..486bcfe5ebd85f8c4ac3b211694e7dd9d13aa302
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_x101-32x4d_fpn_ms-2x_coco.py
@@ -0,0 +1,15 @@
+_base_ = './vfnet_r50_fpn_ms-2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_x101-64x4d-mdconv-c3-c5_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_x101-64x4d-mdconv-c3-c5_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..14a070e73ff54d6833aced096e2d94da4171ca42
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_x101-64x4d-mdconv-c3-c5_fpn_ms-2x_coco.py
@@ -0,0 +1,17 @@
+_base_ = './vfnet_r50-mdconv-c3-c5_fpn_ms-2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ dcn=dict(type='DCNv2', deform_groups=1, fallback_on_stride=False),
+ stage_with_dcn=(False, True, True, True),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_x101-64x4d_fpn_ms-2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_x101-64x4d_fpn_ms-2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..92e3f71df6818a5653ec9c0475c277d89a1adb47
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/vfnet/vfnet_x101-64x4d_fpn_ms-2x_coco.py
@@ -0,0 +1,15 @@
+_base_ = './vfnet_r50_fpn_ms-2x_coco.py'
+model = dict(
+ backbone=dict(
+ type='ResNeXt',
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/wider_face/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/wider_face/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..1904506c64a893f2bfd3881c7e95bd7100fcc6f4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/wider_face/README.md
@@ -0,0 +1,57 @@
+# WIDER FACE
+
+> [WIDER FACE: A Face Detection Benchmark](https://arxiv.org/abs/1511.06523)
+
+
+
+## Abstract
+
+Face detection is one of the most studied topics in the computer vision community. Much of the progresses have been made by the availability of face detection benchmark datasets. We show that there is a gap between current face detection performance and the real world requirements. To facilitate future face detection research, we introduce the WIDER FACE dataset, which is 10 times larger than existing datasets. The dataset contains rich annotations, including occlusions, poses, event categories, and face bounding boxes. Faces in the proposed dataset are extremely challenging due to large variations in scale, pose and occlusion, as shown in Fig. 1. Furthermore, we show that WIDER FACE dataset is an effective training source for face detection. We benchmark several representative detection systems, providing an overview of state-of-the-art performance and propose a solution to deal with large scale variation. Finally, we discuss common failure cases that worth to be further investigated.
+
+
+

+
+
+## Introduction
+
+To use the WIDER Face dataset you need to download it
+and extract to the `data/WIDERFace` folder. Annotation in the VOC format
+can be found in this [repo](https://github.com/sovrasov/wider-face-pascal-voc-annotations.git).
+You should move the annotation files from `WIDER_train_annotations` and `WIDER_val_annotations` folders
+to the `Annotation` folders inside the corresponding directories `WIDER_train` and `WIDER_val`.
+Also annotation lists `val.txt` and `train.txt` should be copied to `data/WIDERFace` from `WIDER_train_annotations` and `WIDER_val_annotations`.
+The directory should be like this:
+
+```
+mmdetection
+├── mmdet
+├── tools
+├── configs
+├── data
+│ ├── WIDERFace
+│ │ ├── WIDER_train
+│ | │ ├──0--Parade
+│ | │ ├── ...
+│ | │ ├── Annotations
+│ │ ├── WIDER_val
+│ | │ ├──0--Parade
+│ | │ ├── ...
+│ | │ ├── Annotations
+│ │ ├── val.txt
+│ │ ├── train.txt
+
+```
+
+After that you can train the SSD300 on WIDER by launching training with the `ssd300_wider_face.py` config or
+create your own config based on the presented one.
+
+## Citation
+
+```latex
+@inproceedings{yang2016wider,
+ Author = {Yang, Shuo and Luo, Ping and Loy, Chen Change and Tang, Xiaoou},
+ Booktitle = {IEEE Conference on Computer Vision and Pattern Recognition (CVPR)},
+ Title = {WIDER FACE: A Face Detection Benchmark},
+ Year = {2016}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/wider_face/retinanet_r50_fpn_1x_widerface.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/wider_face/retinanet_r50_fpn_1x_widerface.py
new file mode 100644
index 0000000000000000000000000000000000000000..78067255f8f69f9d193e8d3ae2fe8a685e4defe1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/wider_face/retinanet_r50_fpn_1x_widerface.py
@@ -0,0 +1,10 @@
+_base_ = [
+ '../_base_/models/retinanet_r50_fpn.py',
+ '../_base_/datasets/wider_face.py', '../_base_/schedules/schedule_1x.py',
+ '../_base_/default_runtime.py'
+]
+# model settings
+model = dict(bbox_head=dict(num_classes=1))
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.01, momentum=0.9, weight_decay=0.0001))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/wider_face/ssd300_8xb32-24e_widerface.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/wider_face/ssd300_8xb32-24e_widerface.py
new file mode 100644
index 0000000000000000000000000000000000000000..02c3c927f78ff022b03bf180789ce91d6061ec9e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/wider_face/ssd300_8xb32-24e_widerface.py
@@ -0,0 +1,64 @@
+_base_ = [
+ '../_base_/models/ssd300.py', '../_base_/datasets/wider_face.py',
+ '../_base_/default_runtime.py', '../_base_/schedules/schedule_2x.py'
+]
+model = dict(bbox_head=dict(num_classes=1))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PhotoMetricDistortion',
+ brightness_delta=32,
+ contrast_range=(0.5, 1.5),
+ saturation_range=(0.5, 1.5),
+ hue_delta=18),
+ dict(
+ type='Expand',
+ mean={{_base_.model.data_preprocessor.mean}},
+ to_rgb={{_base_.model.data_preprocessor.bgr_to_rgb}},
+ ratio_range=(1, 4)),
+ dict(
+ type='MinIoURandomCrop',
+ min_ious=(0.1, 0.3, 0.5, 0.7, 0.9),
+ min_crop_size=0.3),
+ dict(type='Resize', scale=(300, 300), keep_ratio=False),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='Resize', scale=(300, 300), keep_ratio=False),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+dataset_type = 'WIDERFaceDataset'
+data_root = 'data/WIDERFace/'
+train_dataloader = dict(
+ batch_size=32, num_workers=8, dataset=dict(pipeline=train_pipeline))
+
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0,
+ end=1000),
+ dict(type='MultiStepLR', by_epoch=True, milestones=[16, 20], gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(lr=0.012, momentum=0.9, weight_decay=5e-4),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (32 samples per GPU)
+auto_scale_lr = dict(base_batch_size=256)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolact/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolact/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..e884ad65e7181503efd129e7444391e7ea8e2e51
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolact/README.md
@@ -0,0 +1,75 @@
+# YOLACT
+
+> [YOLACT: Real-time Instance Segmentation](https://arxiv.org/abs/1904.02689)
+
+
+
+## Abstract
+
+We present a simple, fully-convolutional model for real-time instance segmentation that achieves 29.8 mAP on MS COCO at 33.5 fps evaluated on a single Titan Xp, which is significantly faster than any previous competitive approach. Moreover, we obtain this result after training on only one GPU. We accomplish this by breaking instance segmentation into two parallel subtasks: (1) generating a set of prototype masks and (2) predicting per-instance mask coefficients. Then we produce instance masks by linearly combining the prototypes with the mask coefficients. We find that because this process doesn't depend on repooling, this approach produces very high-quality masks and exhibits temporal stability for free. Furthermore, we analyze the emergent behavior of our prototypes and show they learn to localize instances on their own in a translation variant manner, despite being fully-convolutional. Finally, we also propose Fast NMS, a drop-in 12 ms faster replacement for standard NMS that only has a marginal performance penalty.
+
+
+

+
+
+## Introduction
+
+A simple, fully convolutional model for real-time instance segmentation. This is the code for our paper:
+
+- [YOLACT: Real-time Instance Segmentation](https://arxiv.org/abs/1904.02689)
+
+
+
+For a real-time demo, check out our ICCV video:
+[](https://www.youtube.com/watch?v=0pMfmo8qfpQ)
+
+## Evaluation
+
+Here are our YOLACT models along with their FPS on a Titan Xp and mAP on COCO's `val`:
+
+| Image Size | GPU x BS | Backbone | \*FPS | mAP | Weights | Configs | Download |
+| :--------: | :------: | :-----------: | :---: | :--: | :-----: | :--------------------------------------: | :-----------------------------------------------------------------------------------------------------------------------------: |
+| 550 | 1x8 | Resnet50-FPN | 42.5 | 29.0 | | [config](./yolact_r50_1xb8-55e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/yolact/yolact_r50_1x8_coco/yolact_r50_1x8_coco_20200908-f38d58df.pth) |
+| 550 | 8x8 | Resnet50-FPN | 42.5 | 28.4 | | [config](./yolact_r50_8xb8-55e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/yolact/yolact_r50_8x8_coco/yolact_r50_8x8_coco_20200908-ca34f5db.pth) |
+| 550 | 1x8 | Resnet101-FPN | 33.5 | 30.4 | | [config](./yolact_r101_1xb8-55e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/yolact/yolact_r101_1x8_coco/yolact_r101_1x8_coco_20200908-4cbe9101.pth) |
+
+\*Note: The FPS is evaluated by the [original implementation](https://github.com/dbolya/yolact). When calculating FPS, only the model inference time is taken into account. Data loading and post-processing operations such as converting masks to RLE code, generating COCO JSON results, image rendering are not included.
+
+## Training
+
+All the aforementioned models are trained with a single GPU. It typically takes ~12GB VRAM when using resnet-101 as the backbone. If you want to try multiple GPUs training, you may have to modify the configuration files accordingly, such as adjusting the training schedule and freezing batch norm.
+
+```Shell
+# Trains using the resnet-101 backbone with a batch size of 8 on a single GPU.
+./tools/dist_train.sh configs/yolact/yolact_r101.py 1
+```
+
+## Testing
+
+Please refer to [mmdetection/docs/getting_started.md](https://mmdetection.readthedocs.io/en/latest/1_exist_data_model.html#test-existing-models).
+
+## Citation
+
+If you use YOLACT or this code base in your work, please cite
+
+```latex
+@inproceedings{yolact-iccv2019,
+ author = {Daniel Bolya and Chong Zhou and Fanyi Xiao and Yong Jae Lee},
+ title = {YOLACT: {Real-time} Instance Segmentation},
+ booktitle = {ICCV},
+ year = {2019},
+}
+```
+
+
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolact/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolact/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..9ca76b3d3910f497e97275d0f25b1b1c3062d12b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolact/metafile.yml
@@ -0,0 +1,81 @@
+Collections:
+ - Name: YOLACT
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - FPN
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/1904.02689
+ Title: 'YOLACT: Real-time Instance Segmentation'
+ README: configs/yolact/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.5.0/mmdet/models/detectors/yolact.py#L9
+ Version: v2.5.0
+
+Models:
+ - Name: yolact_r50_1x8_coco
+ In Collection: YOLACT
+ Config: configs/yolact/yolact_r50_1xb8-55e_coco.py
+ Metadata:
+ Training Resources: 1x V100 GPU
+ Batch Size: 8
+ Epochs: 55
+ inference time (ms/im):
+ - value: 23.53
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (550, 550)
+ Results:
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 29.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/yolact/yolact_r50_1x8_coco/yolact_r50_1x8_coco_20200908-f38d58df.pth
+
+ - Name: yolact_r50_8x8_coco
+ In Collection: YOLACT
+ Config: configs/yolact/yolact_r50_8xb8-55e_coco.py
+ Metadata:
+ Batch Size: 64
+ Epochs: 55
+ inference time (ms/im):
+ - value: 23.53
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (550, 550)
+ Results:
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 28.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/yolact/yolact_r50_8x8_coco/yolact_r50_8x8_coco_20200908-ca34f5db.pth
+
+ - Name: yolact_r101_1x8_coco
+ In Collection: YOLACT
+ Config: configs/yolact/yolact_r101_1xb8-55e_coco.py
+ Metadata:
+ Training Resources: 1x V100 GPU
+ Batch Size: 8
+ Epochs: 55
+ inference time (ms/im):
+ - value: 29.85
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (550, 550)
+ Results:
+ - Task: Instance Segmentation
+ Dataset: COCO
+ Metrics:
+ mask AP: 30.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/yolact/yolact_r101_1x8_coco/yolact_r101_1x8_coco_20200908-4cbe9101.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolact/yolact_r101_1xb8-55e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolact/yolact_r101_1xb8-55e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e6ffe29627ff5bd24b8e53be8d7defaa9eb91df7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolact/yolact_r101_1xb8-55e_coco.py
@@ -0,0 +1,7 @@
+_base_ = './yolact_r50_1xb8-55e_coco.py'
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(type='Pretrained',
+ checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolact/yolact_r50_1xb8-55e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolact/yolact_r50_1xb8-55e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b7dabf1548a733cbf18b8007ae2fa9033a340af6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolact/yolact_r50_1xb8-55e_coco.py
@@ -0,0 +1,170 @@
+_base_ = [
+ '../_base_/datasets/coco_instance.py', '../_base_/default_runtime.py'
+]
+img_norm_cfg = dict(
+ mean=[123.68, 116.78, 103.94], std=[58.40, 57.12, 57.38], to_rgb=True)
+# model settings
+input_size = 550
+model = dict(
+ type='YOLACT',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=img_norm_cfg['mean'],
+ std=img_norm_cfg['std'],
+ bgr_to_rgb=img_norm_cfg['to_rgb'],
+ pad_mask=True),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=-1, # do not freeze stem
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=False, # update the statistics of bn
+ zero_init_residual=False,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_input',
+ num_outs=5,
+ upsample_cfg=dict(mode='bilinear')),
+ bbox_head=dict(
+ type='YOLACTHead',
+ num_classes=80,
+ in_channels=256,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ octave_base_scale=3,
+ scales_per_octave=1,
+ base_sizes=[8, 16, 32, 64, 128],
+ ratios=[0.5, 1.0, 2.0],
+ strides=[550.0 / x for x in [69, 35, 18, 9, 5]],
+ centers=[(550 * 0.5 / x, 550 * 0.5 / x)
+ for x in [69, 35, 18, 9, 5]]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ reduction='none',
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.5),
+ num_head_convs=1,
+ num_protos=32,
+ use_ohem=True),
+ mask_head=dict(
+ type='YOLACTProtonet',
+ in_channels=256,
+ num_protos=32,
+ num_classes=80,
+ max_masks_to_train=100,
+ loss_mask_weight=6.125,
+ with_seg_branch=True,
+ loss_segm=dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.4,
+ min_pos_iou=0.,
+ ignore_iof_thr=-1,
+ gt_max_assign_all=False),
+ sampler=dict(type='PseudoSampler'), # YOLACT should use PseudoSampler
+ # smoothl1_beta=1.,
+ allowed_border=-1,
+ pos_weight=-1,
+ neg_pos_ratio=3,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ mask_thr=0.5,
+ iou_thr=0.5,
+ top_k=200,
+ max_per_img=100,
+ mask_thr_binary=0.5))
+# dataset settings
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(4.0, 4.0)),
+ dict(
+ type='Expand',
+ mean=img_norm_cfg['mean'],
+ to_rgb=img_norm_cfg['to_rgb'],
+ ratio_range=(1, 4)),
+ dict(
+ type='MinIoURandomCrop',
+ min_ious=(0.1, 0.3, 0.5, 0.7, 0.9),
+ min_crop_size=0.3),
+ dict(type='Resize', scale=(input_size, input_size), keep_ratio=False),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='PhotoMetricDistortion',
+ brightness_delta=32,
+ contrast_range=(0.5, 1.5),
+ saturation_range=(0.5, 1.5),
+ hue_delta=18),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(input_size, input_size), keep_ratio=False),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=8,
+ num_workers=4,
+ batch_sampler=None,
+ dataset=dict(pipeline=train_pipeline))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+max_epochs = 55
+# training schedule for 55e
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# learning rate
+param_scheduler = [
+ dict(type='LinearLR', start_factor=0.1, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[20, 42, 49, 52],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=1e-3, momentum=0.9, weight_decay=5e-4))
+
+custom_hooks = [
+ dict(type='CheckInvalidLossHook', interval=50, priority='VERY_LOW')
+]
+
+env_cfg = dict(cudnn_benchmark=True)
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (1 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=8)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolact/yolact_r50_8xb8-55e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolact/yolact_r50_8xb8-55e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e39c285da10ef4821343ebf3c0d0d4c094a97198
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolact/yolact_r50_8xb8-55e_coco.py
@@ -0,0 +1,23 @@
+_base_ = 'yolact_r50_1xb8-55e_coco.py'
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(lr=8e-3),
+ clip_grad=dict(max_norm=35, norm_type=2))
+# learning rate
+max_epochs = 55
+param_scheduler = [
+ dict(type='LinearLR', start_factor=0.1, by_epoch=False, begin=0, end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[20, 42, 49, 52],
+ gamma=0.1)
+]
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..9cb47bcc81a1221dcb4a31b278e7bd62eebf1307
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/README.md
@@ -0,0 +1,55 @@
+# YOLOv3
+
+> [YOLOv3: An Incremental Improvement](https://arxiv.org/abs/1804.02767)
+
+
+
+## Abstract
+
+We present some updates to YOLO! We made a bunch of little design changes to make it better. We also trained this new network that's pretty swell. It's a little bigger than last time but more accurate. It's still fast though, don't worry. At 320x320 YOLOv3 runs in 22 ms at 28.2 mAP, as accurate as SSD but three times faster. When we look at the old .5 IOU mAP detection metric YOLOv3 is quite good. It achieves 57.9 mAP@50 in 51 ms on a Titan X, compared to 57.5 mAP@50 in 198 ms by RetinaNet, similar performance but 3.8x faster.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Scale | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :--------: | :---: | :-----: | :------: | :------------: | :----: | :---------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| DarkNet-53 | 320 | 273e | 2.7 | 63.9 | 27.9 | [config](./yolov3_d53_8xb8-320-273e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/yolo/yolov3_d53_320_273e_coco/yolov3_d53_320_273e_coco-421362b6.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/yolo/yolov3_d53_320_273e_coco/yolov3_d53_320_273e_coco-20200819_172101.log.json) |
+| DarkNet-53 | 416 | 273e | 3.8 | 61.2 | 30.9 | [config](./yolov3_d53_8xb8-ms-416-273e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/yolo/yolov3_d53_mstrain-416_273e_coco/yolov3_d53_mstrain-416_273e_coco-2b60fcd9.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/yolo/yolov3_d53_mstrain-416_273e_coco/yolov3_d53_mstrain-416_273e_coco-20200819_173424.log.json) |
+| DarkNet-53 | 608 | 273e | 7.4 | 48.1 | 33.7 | [config](./yolov3_d53_8xb8-ms-608-273e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/yolo/yolov3_d53_mstrain-608_273e_coco/yolov3_d53_mstrain-608_273e_coco_20210518_115020-a2c3acb8.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/yolo/yolov3_d53_mstrain-608_273e_coco/yolov3_d53_mstrain-608_273e_coco_20210518_115020.log.json) |
+
+## Mixed Precision Training
+
+We also train YOLOv3 with mixed precision training.
+
+| Backbone | Scale | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :--------: | :---: | :-----: | :------: | :------------: | :----: | :-------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| DarkNet-53 | 608 | 273e | 4.7 | 48.1 | 33.8 | [config](./yolov3_d53_8xb8-amp-ms-608-273e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/yolo/yolov3_d53_fp16_mstrain-608_273e_coco/yolov3_d53_fp16_mstrain-608_273e_coco_20210517_213542-4bc34944.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/yolo/yolov3_d53_fp16_mstrain-608_273e_coco/yolov3_d53_fp16_mstrain-608_273e_coco_20210517_213542.log.json) |
+
+## Lightweight models
+
+| Backbone | Scale | Lr schd | Mem (GB) | Inf time (fps) | box AP | Config | Download |
+| :---------: | :---: | :-----: | :------: | :------------: | :----: | :------------------------------------------------------: | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| MobileNetV2 | 416 | 300e | 5.3 | | 23.9 | [config](./yolov3_mobilenetv2_8xb24-ms-416-300e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/yolo/yolov3_mobilenetv2_mstrain-416_300e_coco/yolov3_mobilenetv2_mstrain-416_300e_coco_20210718_010823-f68a07b3.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/yolo/yolov3_mobilenetv2_mstrain-416_300e_coco/yolov3_mobilenetv2_mstrain-416_300e_coco_20210718_010823.log.json) |
+| MobileNetV2 | 320 | 300e | 3.2 | | 22.2 | [config](./yolov3_mobilenetv2_8xb24-320-300e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/yolo/yolov3_mobilenetv2_320_300e_coco/yolov3_mobilenetv2_320_300e_coco_20210719_215349-d18dff72.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/yolo/yolov3_mobilenetv2_320_300e_coco/yolov3_mobilenetv2_320_300e_coco_20210719_215349.log.json) |
+
+Notice: We reduce the number of channels to 96 in both head and neck. It can reduce the flops and parameters, which makes these models more suitable for edge devices.
+
+## Credit
+
+This implementation originates from the project of Haoyu Wu(@wuhy08) at Western Digital.
+
+## Citation
+
+```latex
+@misc{redmon2018yolov3,
+ title={YOLOv3: An Incremental Improvement},
+ author={Joseph Redmon and Ali Farhadi},
+ year={2018},
+ eprint={1804.02767},
+ archivePrefix={arXiv},
+ primaryClass={cs.CV}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..627e70c4d368728d3632f4fda6b68475c3a0fa66
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/metafile.yml
@@ -0,0 +1,124 @@
+Collections:
+ - Name: YOLOv3
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - DarkNet
+ Paper:
+ URL: https://arxiv.org/abs/1804.02767
+ Title: 'YOLOv3: An Incremental Improvement'
+ README: configs/yolo/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.4.0/mmdet/models/detectors/yolo.py#L8
+ Version: v2.4.0
+
+Models:
+ - Name: yolov3_d53_320_273e_coco
+ In Collection: YOLOv3
+ Config: configs/yolo/yolov3_d53_8xb8-320-273e_coco.py
+ Metadata:
+ Training Memory (GB): 2.7
+ inference time (ms/im):
+ - value: 15.65
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (320, 320)
+ Epochs: 273
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 27.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/yolo/yolov3_d53_320_273e_coco/yolov3_d53_320_273e_coco-421362b6.pth
+
+ - Name: yolov3_d53_mstrain-416_273e_coco
+ In Collection: YOLOv3
+ Config: configs/yolo/yolov3_d53_8xb8-ms-416-273e_coco.py
+ Metadata:
+ Training Memory (GB): 3.8
+ inference time (ms/im):
+ - value: 16.34
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (416, 416)
+ Epochs: 273
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 30.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/yolo/yolov3_d53_mstrain-416_273e_coco/yolov3_d53_mstrain-416_273e_coco-2b60fcd9.pth
+
+ - Name: yolov3_d53_mstrain-608_273e_coco
+ In Collection: YOLOv3
+ Config: configs/yolo/yolov3_d53_8xb8-ms-608-273e_coco.py
+ Metadata:
+ Training Memory (GB): 7.4
+ inference time (ms/im):
+ - value: 20.79
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP32
+ resolution: (608, 608)
+ Epochs: 273
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 33.7
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/yolo/yolov3_d53_mstrain-608_273e_coco/yolov3_d53_mstrain-608_273e_coco_20210518_115020-a2c3acb8.pth
+
+ - Name: yolov3_d53_fp16_mstrain-608_273e_coco
+ In Collection: YOLOv3
+ Config: configs/yolo/yolov3_d53_8xb8-amp-ms-608-273e_coco.py
+ Metadata:
+ Training Memory (GB): 4.7
+ inference time (ms/im):
+ - value: 20.79
+ hardware: V100
+ backend: PyTorch
+ batch size: 1
+ mode: FP16
+ resolution: (608, 608)
+ Epochs: 273
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 33.8
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/yolo/yolov3_d53_fp16_mstrain-608_273e_coco/yolov3_d53_fp16_mstrain-608_273e_coco_20210517_213542-4bc34944.pth
+
+ - Name: yolov3_mobilenetv2_8xb24-320-300e_coco
+ In Collection: YOLOv3
+ Config: configs/yolo/yolov3_mobilenetv2_8xb24-320-300e_coco.py
+ Metadata:
+ Training Memory (GB): 3.2
+ Epochs: 300
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 22.2
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/yolo/yolov3_mobilenetv2_320_300e_coco/yolov3_mobilenetv2_320_300e_coco_20210719_215349-d18dff72.pth
+
+ - Name: yolov3_mobilenetv2_8xb24-ms-416-300e_coco
+ In Collection: YOLOv3
+ Config: configs/yolo/yolov3_mobilenetv2_8xb24-ms-416-300e_coco.py
+ Metadata:
+ Training Memory (GB): 5.3
+ Epochs: 300
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 23.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/yolo/yolov3_mobilenetv2_mstrain-416_300e_coco/yolov3_mobilenetv2_mstrain-416_300e_coco_20210718_010823-f68a07b3.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/yolov3_d53_8xb8-320-273e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/yolov3_d53_8xb8-320-273e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a3d08dd7706e5ba5bec5fc9e8da6fab120ed813d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/yolov3_d53_8xb8-320-273e_coco.py
@@ -0,0 +1,29 @@
+_base_ = './yolov3_d53_8xb8-ms-608-273e_coco.py'
+
+input_size = (320, 320)
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ # `mean` and `to_rgb` should be the same with the `preprocess_cfg`
+ dict(type='Expand', mean=[0, 0, 0], to_rgb=True, ratio_range=(1, 2)),
+ dict(
+ type='MinIoURandomCrop',
+ min_ious=(0.4, 0.5, 0.6, 0.7, 0.8, 0.9),
+ min_crop_size=0.3),
+ dict(type='Resize', scale=input_size, keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PhotoMetricDistortion'),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=input_size, keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/yolov3_d53_8xb8-amp-ms-608-273e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/yolov3_d53_8xb8-amp-ms-608-273e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..173d8ee22227b3c3f4aa0488cb4e6f131d7dbee4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/yolov3_d53_8xb8-amp-ms-608-273e_coco.py
@@ -0,0 +1,3 @@
+_base_ = './yolov3_d53_8xb8-ms-608-273e_coco.py'
+# fp16 settings
+optim_wrapper = dict(type='AmpOptimWrapper', loss_scale='dynamic')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/yolov3_d53_8xb8-ms-416-273e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/yolov3_d53_8xb8-ms-416-273e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ca0127e83edaeb8d5851ed089f6bd6d7385a1f86
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/yolov3_d53_8xb8-ms-416-273e_coco.py
@@ -0,0 +1,28 @@
+_base_ = './yolov3_d53_8xb8-ms-608-273e_coco.py'
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ # `mean` and `to_rgb` should be the same with the `preprocess_cfg`
+ dict(type='Expand', mean=[0, 0, 0], to_rgb=True, ratio_range=(1, 2)),
+ dict(
+ type='MinIoURandomCrop',
+ min_ious=(0.4, 0.5, 0.6, 0.7, 0.8, 0.9),
+ min_crop_size=0.3),
+ dict(type='RandomResize', scale=[(320, 320), (416, 416)], keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PhotoMetricDistortion'),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(416, 416), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/yolov3_d53_8xb8-ms-608-273e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/yolov3_d53_8xb8-ms-608-273e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d4a36dfdaaf9b9e013882a6c28d42cca5942be20
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/yolov3_d53_8xb8-ms-608-273e_coco.py
@@ -0,0 +1,167 @@
+_base_ = ['../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py']
+# model settings
+data_preprocessor = dict(
+ type='DetDataPreprocessor',
+ mean=[0, 0, 0],
+ std=[255., 255., 255.],
+ bgr_to_rgb=True,
+ pad_size_divisor=32)
+model = dict(
+ type='YOLOV3',
+ data_preprocessor=data_preprocessor,
+ backbone=dict(
+ type='Darknet',
+ depth=53,
+ out_indices=(3, 4, 5),
+ init_cfg=dict(type='Pretrained', checkpoint='open-mmlab://darknet53')),
+ neck=dict(
+ type='YOLOV3Neck',
+ num_scales=3,
+ in_channels=[1024, 512, 256],
+ out_channels=[512, 256, 128]),
+ bbox_head=dict(
+ type='YOLOV3Head',
+ num_classes=80,
+ in_channels=[512, 256, 128],
+ out_channels=[1024, 512, 256],
+ anchor_generator=dict(
+ type='YOLOAnchorGenerator',
+ base_sizes=[[(116, 90), (156, 198), (373, 326)],
+ [(30, 61), (62, 45), (59, 119)],
+ [(10, 13), (16, 30), (33, 23)]],
+ strides=[32, 16, 8]),
+ bbox_coder=dict(type='YOLOBBoxCoder'),
+ featmap_strides=[32, 16, 8],
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ loss_weight=1.0,
+ reduction='sum'),
+ loss_conf=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ loss_weight=1.0,
+ reduction='sum'),
+ loss_xy=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ loss_weight=2.0,
+ reduction='sum'),
+ loss_wh=dict(type='MSELoss', loss_weight=2.0, reduction='sum')),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='GridAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0)),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ conf_thr=0.005,
+ nms=dict(type='nms', iou_threshold=0.45),
+ max_per_img=100))
+# dataset settings
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='Expand',
+ mean=data_preprocessor['mean'],
+ to_rgb=data_preprocessor['bgr_to_rgb'],
+ ratio_range=(1, 2)),
+ dict(
+ type='MinIoURandomCrop',
+ min_ious=(0.4, 0.5, 0.6, 0.7, 0.8, 0.9),
+ min_crop_size=0.3),
+ dict(type='RandomResize', scale=[(320, 320), (608, 608)], keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PhotoMetricDistortion'),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(608, 608), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=8,
+ num_workers=4,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric='bbox',
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+train_cfg = dict(max_epochs=273, val_interval=7)
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.001, momentum=0.9, weight_decay=0.0005),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+# learning policy
+param_scheduler = [
+ dict(type='LinearLR', start_factor=0.1, by_epoch=False, begin=0, end=2000),
+ dict(type='MultiStepLR', by_epoch=True, milestones=[218, 246], gamma=0.1)
+]
+
+default_hooks = dict(checkpoint=dict(type='CheckpointHook', interval=7))
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/yolov3_mobilenetv2_8xb24-320-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/yolov3_mobilenetv2_8xb24-320-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..07b393734329fd3ed5f4bd11fbc15b4abf7846bb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/yolov3_mobilenetv2_8xb24-320-300e_coco.py
@@ -0,0 +1,42 @@
+_base_ = ['./yolov3_mobilenetv2_8xb24-ms-416-300e_coco.py']
+
+# yapf:disable
+model = dict(
+ bbox_head=dict(
+ anchor_generator=dict(
+ base_sizes=[[(220, 125), (128, 222), (264, 266)],
+ [(35, 87), (102, 96), (60, 170)],
+ [(10, 15), (24, 36), (72, 42)]])))
+# yapf:enable
+
+input_size = (320, 320)
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ # `mean` and `to_rgb` should be the same with the `preprocess_cfg`
+ dict(
+ type='Expand',
+ mean=[123.675, 116.28, 103.53],
+ to_rgb=True,
+ ratio_range=(1, 2)),
+ dict(
+ type='MinIoURandomCrop',
+ min_ious=(0.4, 0.5, 0.6, 0.7, 0.8, 0.9),
+ min_crop_size=0.3),
+ dict(type='Resize', scale=input_size, keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PhotoMetricDistortion'),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=input_size, keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(dataset=dict(dataset=dict(pipeline=train_pipeline)))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/yolov3_mobilenetv2_8xb24-ms-416-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/yolov3_mobilenetv2_8xb24-ms-416-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..9a161b66fe92666e904a9580ab5a1ff16d630ab7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolo/yolov3_mobilenetv2_8xb24-ms-416-300e_coco.py
@@ -0,0 +1,176 @@
+_base_ = ['../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py']
+# model settings
+data_preprocessor = dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32)
+model = dict(
+ type='YOLOV3',
+ data_preprocessor=data_preprocessor,
+ backbone=dict(
+ type='MobileNetV2',
+ out_indices=(2, 4, 6),
+ act_cfg=dict(type='LeakyReLU', negative_slope=0.1),
+ init_cfg=dict(
+ type='Pretrained', checkpoint='open-mmlab://mmdet/mobilenet_v2')),
+ neck=dict(
+ type='YOLOV3Neck',
+ num_scales=3,
+ in_channels=[320, 96, 32],
+ out_channels=[96, 96, 96]),
+ bbox_head=dict(
+ type='YOLOV3Head',
+ num_classes=80,
+ in_channels=[96, 96, 96],
+ out_channels=[96, 96, 96],
+ anchor_generator=dict(
+ type='YOLOAnchorGenerator',
+ base_sizes=[[(116, 90), (156, 198), (373, 326)],
+ [(30, 61), (62, 45), (59, 119)],
+ [(10, 13), (16, 30), (33, 23)]],
+ strides=[32, 16, 8]),
+ bbox_coder=dict(type='YOLOBBoxCoder'),
+ featmap_strides=[32, 16, 8],
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ loss_weight=1.0,
+ reduction='sum'),
+ loss_conf=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ loss_weight=1.0,
+ reduction='sum'),
+ loss_xy=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ loss_weight=2.0,
+ reduction='sum'),
+ loss_wh=dict(type='MSELoss', loss_weight=2.0, reduction='sum')),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='GridAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0)),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ conf_thr=0.005,
+ nms=dict(type='nms', iou_threshold=0.45),
+ max_per_img=100))
+# dataset settings
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='Expand',
+ mean=data_preprocessor['mean'],
+ to_rgb=data_preprocessor['bgr_to_rgb'],
+ ratio_range=(1, 2)),
+ dict(
+ type='MinIoURandomCrop',
+ min_ious=(0.4, 0.5, 0.6, 0.7, 0.8, 0.9),
+ min_crop_size=0.3),
+ dict(type='RandomResize', scale=[(320, 320), (416, 416)], keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PhotoMetricDistortion'),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=(416, 416), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=24,
+ num_workers=4,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type='RepeatDataset', # use RepeatDataset to speed up training
+ times=10,
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args)))
+val_dataloader = dict(
+ batch_size=24,
+ num_workers=4,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric='bbox',
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+train_cfg = dict(max_epochs=30)
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='SGD', lr=0.003, momentum=0.9, weight_decay=0.0005),
+ clip_grad=dict(max_norm=35, norm_type=2))
+
+# learning policy
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=0.0001,
+ by_epoch=False,
+ begin=0,
+ end=4000),
+ dict(type='MultiStepLR', by_epoch=True, milestones=[24, 28], gamma=0.1)
+]
+
+find_unused_parameters = True
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (24 samples per GPU)
+auto_scale_lr = dict(base_batch_size=192)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolof/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolof/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..b9167f6e6e34a64022b82b212e4bc81808dc3395
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolof/README.md
@@ -0,0 +1,35 @@
+# YOLOF
+
+> [You Only Look One-level Feature](https://arxiv.org/abs/2103.09460)
+
+
+
+## Abstract
+
+This paper revisits feature pyramids networks (FPN) for one-stage detectors and points out that the success of FPN is due to its divide-and-conquer solution to the optimization problem in object detection rather than multi-scale feature fusion. From the perspective of optimization, we introduce an alternative way to address the problem instead of adopting the complex feature pyramids - {\\em utilizing only one-level feature for detection}. Based on the simple and efficient solution, we present You Only Look One-level Feature (YOLOF). In our method, two key components, Dilated Encoder and Uniform Matching, are proposed and bring considerable improvements. Extensive experiments on the COCO benchmark prove the effectiveness of the proposed model. Our YOLOF achieves comparable results with its feature pyramids counterpart RetinaNet while being 2.5× faster. Without transformer layers, YOLOF can match the performance of DETR in a single-level feature manner with 7× less training epochs. With an image size of 608×608, YOLOF achieves 44.3 mAP running at 60 fps on 2080Ti, which is 13% faster than YOLOv4.
+
+
+

+
+
+## Results and Models
+
+| Backbone | Style | Epoch | Lr schd | Mem (GB) | box AP | Config | Download |
+| :------: | :---: | :---: | :-----: | :------: | :----: | :--------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50-C5 | caffe | Y | 1x | 8.3 | 37.5 | [config](./yolof_r50-c5_8xb8-1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/yolof/yolof_r50_c5_8x8_1x_coco/yolof_r50_c5_8x8_1x_coco_20210425_024427-8e864411.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/yolof/yolof_r50_c5_8x8_1x_coco/yolof_r50_c5_8x8_1x_coco_20210425_024427.log.json) |
+
+**Note**:
+
+1. We find that the performance is unstable and may fluctuate by about 0.3 mAP. mAP 37.4 ~ 37.7 is acceptable in YOLOF_R_50_C5_1x. Such fluctuation can also be found in the [original implementation](https://github.com/chensnathan/YOLOF).
+2. In addition to instability issues, sometimes there are large loss fluctuations and NAN, so there may still be problems with this project, which will be improved subsequently.
+
+## Citation
+
+```latex
+@inproceedings{chen2021you,
+ title={You Only Look One-level Feature},
+ author={Chen, Qiang and Wang, Yingming and Yang, Tong and Zhang, Xiangyu and Cheng, Jian and Sun, Jian},
+ booktitle={IEEE Conference on Computer Vision and Pattern Recognition},
+ year={2021}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolof/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolof/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..b3b7b7f8d5d3d7faec0cd04984ede59a99d06f38
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolof/metafile.yml
@@ -0,0 +1,32 @@
+Collections:
+ - Name: YOLOF
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Momentum
+ - Weight Decay
+ Training Resources: 8x V100 GPUs
+ Architecture:
+ - Dilated Encoder
+ - ResNet
+ Paper:
+ URL: https://arxiv.org/abs/2103.09460
+ Title: 'You Only Look One-level Feature'
+ README: configs/yolof/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.12.0/mmdet/models/detectors/yolof.py#L6
+ Version: v2.12.0
+
+Models:
+ - Name: yolof_r50_c5_8x8_1x_coco
+ In Collection: YOLOF
+ Config: configs/yolof/yolof_r50-c5_8xb8-1x_coco.py
+ Metadata:
+ Training Memory (GB): 8.3
+ Epochs: 12
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 37.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/yolof/yolof_r50_c5_8x8_1x_coco/yolof_r50_c5_8x8_1x_coco_20210425_024427-8e864411.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolof/yolof_r50-c5_8xb8-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolof/yolof_r50-c5_8xb8-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5ea228e3e3270e07a4e5b171ab544c704fb172f3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolof/yolof_r50-c5_8xb8-1x_coco.py
@@ -0,0 +1,116 @@
+_base_ = [
+ '../_base_/datasets/coco_detection.py',
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py'
+]
+model = dict(
+ type='YOLOF',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(3, ),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='open-mmlab://detectron/resnet50_caffe')),
+ neck=dict(
+ type='DilatedEncoder',
+ in_channels=2048,
+ out_channels=512,
+ block_mid_channels=128,
+ num_residual_blocks=4,
+ block_dilations=[2, 4, 6, 8]),
+ bbox_head=dict(
+ type='YOLOFHead',
+ num_classes=80,
+ in_channels=512,
+ reg_decoded_bbox=True,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ scales=[1, 2, 4, 8, 16],
+ strides=[32]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1., 1., 1., 1.],
+ add_ctr_clamp=True,
+ ctr_clamp=32),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=1.0)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='UniformAssigner', pos_ignore_thr=0.15, neg_ignore_thr=0.7),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100))
+# optimizer
+optim_wrapper = dict(
+ optimizer=dict(type='SGD', lr=0.12, momentum=0.9, weight_decay=0.0001),
+ paramwise_cfg=dict(
+ norm_decay_mult=0., custom_keys={'backbone': dict(lr_mult=1. / 3)}))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=0.00066667,
+ by_epoch=False,
+ begin=0,
+ end=1500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='RandomShift', prob=0.5, max_shift_px=32),
+ dict(type='PackDetInputs')
+]
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=8, num_workers=8, dataset=dict(pipeline=train_pipeline))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolof/yolof_r50-c5_8xb8-iter-1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolof/yolof_r50-c5_8xb8-iter-1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..466a820099e3ac1760371e8352a89f93fbeef5ee
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolof/yolof_r50-c5_8xb8-iter-1x_coco.py
@@ -0,0 +1,32 @@
+_base_ = './yolof_r50-c5_8xb8-1x_coco.py'
+
+# We implemented the iter-based config according to the source code.
+# COCO dataset has 117266 images after filtering. We use 8 gpu and
+# 8 batch size training, so 22500 is equivalent to
+# 22500/(117266/(8x8))=12.3 epoch, 15000 is equivalent to 8.2 epoch,
+# 20000 is equivalent to 10.9 epoch. Due to lr(0.12) is large,
+# the iter-based and epoch-based setting have about 0.2 difference on
+# the mAP evaluation value.
+
+train_cfg = dict(
+ _delete_=True,
+ type='IterBasedTrainLoop',
+ max_iters=22500,
+ val_interval=4500)
+
+# learning rate policy
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=22500,
+ by_epoch=False,
+ milestones=[15000, 20000],
+ gamma=0.1)
+]
+train_dataloader = dict(sampler=dict(type='InfiniteSampler'))
+default_hooks = dict(checkpoint=dict(by_epoch=False, interval=2500))
+
+log_processor = dict(by_epoch=False)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..0cde192676db90e8dbd92de80b55d540493e17e5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/README.md
@@ -0,0 +1,39 @@
+# YOLOX
+
+> [YOLOX: Exceeding YOLO Series in 2021](https://arxiv.org/abs/2107.08430)
+
+
+
+## Abstract
+
+In this report, we present some experienced improvements to YOLO series, forming a new high-performance detector -- YOLOX. We switch the YOLO detector to an anchor-free manner and conduct other advanced detection techniques, i.e., a decoupled head and the leading label assignment strategy SimOTA to achieve state-of-the-art results across a large scale range of models: For YOLO-Nano with only 0.91M parameters and 1.08G FLOPs, we get 25.3% AP on COCO, surpassing NanoDet by 1.8% AP; for YOLOv3, one of the most widely used detectors in industry, we boost it to 47.3% AP on COCO, outperforming the current best practice by 3.0% AP; for YOLOX-L with roughly the same amount of parameters as YOLOv4-CSP, YOLOv5-L, we achieve 50.0% AP on COCO at a speed of 68.9 FPS on Tesla V100, exceeding YOLOv5-L by 1.8% AP. Further, we won the 1st Place on Streaming Perception Challenge (Workshop on Autonomous Driving at CVPR 2021) using a single YOLOX-L model. We hope this report can provide useful experience for developers and researchers in practical scenes, and we also provide deploy versions with ONNX, TensorRT, NCNN, and Openvino supported.
+
+
+

+
+
+## Results and Models
+
+| Backbone | size | Mem (GB) | box AP | Config | Download |
+| :--------: | :--: | :------: | :----: | :--------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| YOLOX-tiny | 416 | 3.5 | 32.0 | [config](./yolox_tiny_8xb8-300e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/yolox/yolox_tiny_8x8_300e_coco/yolox_tiny_8x8_300e_coco_20211124_171234-b4047906.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/yolox/yolox_tiny_8x8_300e_coco/yolox_tiny_8x8_300e_coco_20211124_171234.log.json) |
+| YOLOX-s | 640 | 7.6 | 40.5 | [config](./yolox_s_8xb8-300e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/yolox/yolox_s_8x8_300e_coco/yolox_s_8x8_300e_coco_20211121_095711-4592a793.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/yolox/yolox_s_8x8_300e_coco/yolox_s_8x8_300e_coco_20211121_095711.log.json) |
+| YOLOX-l | 640 | 19.9 | 49.4 | [config](./yolox_l_8xb8-300e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/yolox/yolox_l_8x8_300e_coco/yolox_l_8x8_300e_coco_20211126_140236-d3bd2b23.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/yolox/yolox_l_8x8_300e_coco/yolox_l_8x8_300e_coco_20211126_140236.log.json) |
+| YOLOX-x | 640 | 28.1 | 50.9 | [config](./yolox_x_8xb8-300e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v2.0/yolox/yolox_x_8x8_300e_coco/yolox_x_8x8_300e_coco_20211126_140254-1ef88d67.pth) \| [log](https://download.openmmlab.com/mmdetection/v2.0/yolox/yolox_x_8x8_300e_coco/yolox_x_8x8_300e_coco_20211126_140254.log.json) |
+
+**Note**:
+
+1. The test score threshold is 0.001, and the box AP indicates the best AP.
+2. Due to the need for pre-training weights, we cannot reproduce the performance of the `yolox-nano` model. Please refer to https://github.com/Megvii-BaseDetection/YOLOX/issues/674 for more information.
+3. We also trained the model by the official release of YOLOX based on [Megvii-BaseDetection/YOLOX#735](https://github.com/Megvii-BaseDetection/YOLOX/issues/735) with commit ID [38c633](https://github.com/Megvii-BaseDetection/YOLOX/tree/38c633bf176462ee42b110c70e4ffe17b5753208). We found that the best AP of `YOLOX-tiny`, `YOLOX-s`, `YOLOX-l`, and `YOLOX-x` is 31.8, 40.3, 49.2, and 50.9, respectively. The performance is consistent with that of our re-implementation (see Table above) but still has a gap (0.3~0.8 AP) in comparison with the reported performance in their [README](https://github.com/Megvii-BaseDetection/YOLOX/blob/38c633bf176462ee42b110c70e4ffe17b5753208/README.md#benchmark).
+
+## Citation
+
+```latex
+@article{yolox2021,
+ title={{YOLOX}: Exceeding YOLO Series in 2021},
+ author={Ge, Zheng and Liu, Songtao and Wang, Feng and Li, Zeming and Sun, Jian},
+ journal={arXiv preprint arXiv:2107.08430},
+ year={2021}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/metafile.yml b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/metafile.yml
new file mode 100644
index 0000000000000000000000000000000000000000..2f64450e94cae436a05f46da67d3a1264235ffbd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/metafile.yml
@@ -0,0 +1,70 @@
+Collections:
+ - Name: YOLOX
+ Metadata:
+ Training Data: COCO
+ Training Techniques:
+ - SGD with Nesterov
+ - Weight Decay
+ - Cosine Annealing Lr Updater
+ Training Resources: 8x TITANXp GPUs
+ Architecture:
+ - CSPDarkNet
+ - PAFPN
+ Paper:
+ URL: https://arxiv.org/abs/2107.08430
+ Title: 'YOLOX: Exceeding YOLO Series in 2021'
+ README: configs/yolox/README.md
+ Code:
+ URL: https://github.com/open-mmlab/mmdetection/blob/v2.15.1/mmdet/models/detectors/yolox.py#L6
+ Version: v2.15.1
+
+
+Models:
+ - Name: yolox_s_8x8_300e_coco
+ In Collection: YOLOX
+ Config: configs/yolox/yolox_s_8xb8-300e_coco.py
+ Metadata:
+ Training Memory (GB): 7.6
+ Epochs: 300
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 40.5
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/yolox/yolox_s_8x8_300e_coco/yolox_s_8x8_300e_coco_20211121_095711-4592a793.pth
+ - Name: yolox_l_8x8_300e_coco
+ In Collection: YOLOX
+ Config: configs/yolox/yolox_l_8xb8-300e_coco.py
+ Metadata:
+ Training Memory (GB): 19.9
+ Epochs: 300
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 49.4
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/yolox/yolox_l_8x8_300e_coco/yolox_l_8x8_300e_coco_20211126_140236-d3bd2b23.pth
+ - Name: yolox_x_8x8_300e_coco
+ In Collection: YOLOX
+ Config: configs/yolox/yolox_x_8xb8-300e_coco.py
+ Metadata:
+ Training Memory (GB): 28.1
+ Epochs: 300
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 50.9
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/yolox/yolox_x_8x8_300e_coco/yolox_x_8x8_300e_coco_20211126_140254-1ef88d67.pth
+ - Name: yolox_tiny_8x8_300e_coco
+ In Collection: YOLOX
+ Config: configs/yolox/yolox_tiny_8xb8-300e_coco.py
+ Metadata:
+ Training Memory (GB): 3.5
+ Epochs: 300
+ Results:
+ - Task: Object Detection
+ Dataset: COCO
+ Metrics:
+ box AP: 32.0
+ Weights: https://download.openmmlab.com/mmdetection/v2.0/yolox/yolox_tiny_8x8_300e_coco/yolox_tiny_8x8_300e_coco_20211124_171234-b4047906.pth
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_l_8xb8-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_l_8xb8-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..2a4b287bad595db65df69b7d6f80163bd4a49e44
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_l_8xb8-300e_coco.py
@@ -0,0 +1,8 @@
+_base_ = './yolox_s_8xb8-300e_coco.py'
+
+# model settings
+model = dict(
+ backbone=dict(deepen_factor=1.0, widen_factor=1.0),
+ neck=dict(
+ in_channels=[256, 512, 1024], out_channels=256, num_csp_blocks=3),
+ bbox_head=dict(in_channels=256, feat_channels=256))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_m_8xb8-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_m_8xb8-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d82f9e98f1fcd4a1c6089807adc3cca2b48d6b5e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_m_8xb8-300e_coco.py
@@ -0,0 +1,8 @@
+_base_ = './yolox_s_8xb8-300e_coco.py'
+
+# model settings
+model = dict(
+ backbone=dict(deepen_factor=0.67, widen_factor=0.75),
+ neck=dict(in_channels=[192, 384, 768], out_channels=192, num_csp_blocks=2),
+ bbox_head=dict(in_channels=192, feat_channels=192),
+)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_nano_8xb8-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_nano_8xb8-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3f7a1c5ab066439c78ffa005a2a60c9057223849
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_nano_8xb8-300e_coco.py
@@ -0,0 +1,11 @@
+_base_ = './yolox_tiny_8xb8-300e_coco.py'
+
+# model settings
+model = dict(
+ backbone=dict(deepen_factor=0.33, widen_factor=0.25, use_depthwise=True),
+ neck=dict(
+ in_channels=[64, 128, 256],
+ out_channels=64,
+ num_csp_blocks=1,
+ use_depthwise=True),
+ bbox_head=dict(in_channels=64, feat_channels=64, use_depthwise=True))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_s_8xb8-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_s_8xb8-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3e324eb5b99202fd42c8d67847a1be1c165b4057
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_s_8xb8-300e_coco.py
@@ -0,0 +1,250 @@
+_base_ = [
+ '../_base_/schedules/schedule_1x.py', '../_base_/default_runtime.py',
+ './yolox_tta.py'
+]
+
+img_scale = (640, 640) # width, height
+
+# model settings
+model = dict(
+ type='YOLOX',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ pad_size_divisor=32,
+ batch_augments=[
+ dict(
+ type='BatchSyncRandomResize',
+ random_size_range=(480, 800),
+ size_divisor=32,
+ interval=10)
+ ]),
+ backbone=dict(
+ type='CSPDarknet',
+ deepen_factor=0.33,
+ widen_factor=0.5,
+ out_indices=(2, 3, 4),
+ use_depthwise=False,
+ spp_kernal_sizes=(5, 9, 13),
+ norm_cfg=dict(type='BN', momentum=0.03, eps=0.001),
+ act_cfg=dict(type='Swish'),
+ ),
+ neck=dict(
+ type='YOLOXPAFPN',
+ in_channels=[128, 256, 512],
+ out_channels=128,
+ num_csp_blocks=1,
+ use_depthwise=False,
+ upsample_cfg=dict(scale_factor=2, mode='nearest'),
+ norm_cfg=dict(type='BN', momentum=0.03, eps=0.001),
+ act_cfg=dict(type='Swish')),
+ bbox_head=dict(
+ type='YOLOXHead',
+ num_classes=80,
+ in_channels=128,
+ feat_channels=128,
+ stacked_convs=2,
+ strides=(8, 16, 32),
+ use_depthwise=False,
+ norm_cfg=dict(type='BN', momentum=0.03, eps=0.001),
+ act_cfg=dict(type='Swish'),
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ reduction='sum',
+ loss_weight=1.0),
+ loss_bbox=dict(
+ type='IoULoss',
+ mode='square',
+ eps=1e-16,
+ reduction='sum',
+ loss_weight=5.0),
+ loss_obj=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ reduction='sum',
+ loss_weight=1.0),
+ loss_l1=dict(type='L1Loss', reduction='sum', loss_weight=1.0)),
+ train_cfg=dict(assigner=dict(type='SimOTAAssigner', center_radius=2.5)),
+ # In order to align the source code, the threshold of the val phase is
+ # 0.01, and the threshold of the test phase is 0.001.
+ test_cfg=dict(score_thr=0.01, nms=dict(type='nms', iou_threshold=0.65)))
+
+# dataset settings
+data_root = 'data/coco/'
+dataset_type = 'CocoDataset'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type='Mosaic', img_scale=img_scale, pad_val=114.0),
+ dict(
+ type='RandomAffine',
+ scaling_ratio_range=(0.1, 2),
+ # img_scale is (width, height)
+ border=(-img_scale[0] // 2, -img_scale[1] // 2)),
+ dict(
+ type='MixUp',
+ img_scale=img_scale,
+ ratio_range=(0.8, 1.6),
+ pad_val=114.0),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ # According to the official implementation, multi-scale
+ # training is not considered here but in the
+ # 'mmdet/models/detectors/yolox.py'.
+ # Resize and Pad are for the last 15 epochs when Mosaic,
+ # RandomAffine, and MixUp are closed by YOLOXModeSwitchHook.
+ dict(type='Resize', scale=img_scale, keep_ratio=True),
+ dict(
+ type='Pad',
+ pad_to_square=True,
+ # If the image is three-channel, the pad value needs
+ # to be set separately for each channel.
+ pad_val=dict(img=(114.0, 114.0, 114.0))),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1, 1), keep_empty=False),
+ dict(type='PackDetInputs')
+]
+
+train_dataset = dict(
+ # use MultiImageMixDataset wrapper to support mosaic and mixup
+ type='MultiImageMixDataset',
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ pipeline=[
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True)
+ ],
+ filter_cfg=dict(filter_empty_gt=False, min_size=32),
+ backend_args=backend_args),
+ pipeline=train_pipeline)
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='Resize', scale=img_scale, keep_ratio=True),
+ dict(
+ type='Pad',
+ pad_to_square=True,
+ pad_val=dict(img=(114.0, 114.0, 114.0))),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=8,
+ num_workers=4,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ dataset=train_dataset)
+val_dataloader = dict(
+ batch_size=8,
+ num_workers=4,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type='CocoMetric',
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric='bbox',
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+# training settings
+max_epochs = 300
+num_last_epochs = 15
+interval = 10
+
+train_cfg = dict(max_epochs=max_epochs, val_interval=interval)
+
+# optimizer
+# default 8 gpu
+base_lr = 0.01
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(
+ type='SGD', lr=base_lr, momentum=0.9, weight_decay=5e-4,
+ nesterov=True),
+ paramwise_cfg=dict(norm_decay_mult=0., bias_decay_mult=0.))
+
+# learning rate
+param_scheduler = [
+ dict(
+ # use quadratic formula to warm up 5 epochs
+ # and lr is updated by iteration
+ # TODO: fix default scope in get function
+ type='mmdet.QuadraticWarmupLR',
+ by_epoch=True,
+ begin=0,
+ end=5,
+ convert_to_iter_based=True),
+ dict(
+ # use cosine lr from 5 to 285 epoch
+ type='CosineAnnealingLR',
+ eta_min=base_lr * 0.05,
+ begin=5,
+ T_max=max_epochs - num_last_epochs,
+ end=max_epochs - num_last_epochs,
+ by_epoch=True,
+ convert_to_iter_based=True),
+ dict(
+ # use fixed lr during last 15 epochs
+ type='ConstantLR',
+ by_epoch=True,
+ factor=1,
+ begin=max_epochs - num_last_epochs,
+ end=max_epochs,
+ )
+]
+
+default_hooks = dict(
+ checkpoint=dict(
+ interval=interval,
+ max_keep_ckpts=3 # only keep latest 3 checkpoints
+ ))
+
+custom_hooks = [
+ dict(
+ type='YOLOXModeSwitchHook',
+ num_last_epochs=num_last_epochs,
+ priority=48),
+ dict(type='SyncNormHook', priority=48),
+ dict(
+ type='EMAHook',
+ ema_type='ExpMomentumEMA',
+ momentum=0.0001,
+ update_buffers=True,
+ priority=49)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_tiny_8xb8-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_tiny_8xb8-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..86f7e9a6191066ab9b672d548b93a29e64746f29
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_tiny_8xb8-300e_coco.py
@@ -0,0 +1,54 @@
+_base_ = './yolox_s_8xb8-300e_coco.py'
+
+# model settings
+model = dict(
+ data_preprocessor=dict(batch_augments=[
+ dict(
+ type='BatchSyncRandomResize',
+ random_size_range=(320, 640),
+ size_divisor=32,
+ interval=10)
+ ]),
+ backbone=dict(deepen_factor=0.33, widen_factor=0.375),
+ neck=dict(in_channels=[96, 192, 384], out_channels=96),
+ bbox_head=dict(in_channels=96, feat_channels=96))
+
+img_scale = (640, 640) # width, height
+
+train_pipeline = [
+ dict(type='Mosaic', img_scale=img_scale, pad_val=114.0),
+ dict(
+ type='RandomAffine',
+ scaling_ratio_range=(0.5, 1.5),
+ # img_scale is (width, height)
+ border=(-img_scale[0] // 2, -img_scale[1] // 2)),
+ dict(type='YOLOXHSVRandomAug'),
+ dict(type='RandomFlip', prob=0.5),
+ # Resize and Pad are for the last 15 epochs when Mosaic and
+ # RandomAffine are closed by YOLOXModeSwitchHook.
+ dict(type='Resize', scale=img_scale, keep_ratio=True),
+ dict(
+ type='Pad',
+ pad_to_square=True,
+ pad_val=dict(img=(114.0, 114.0, 114.0))),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1, 1), keep_empty=False),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args={{_base_.backend_args}}),
+ dict(type='Resize', scale=(416, 416), keep_ratio=True),
+ dict(
+ type='Pad',
+ pad_to_square=True,
+ pad_val=dict(img=(114.0, 114.0, 114.0))),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_tta.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_tta.py
new file mode 100644
index 0000000000000000000000000000000000000000..e65244be6e1bb70393d111ef4d25334d3b2ce8a6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_tta.py
@@ -0,0 +1,36 @@
+tta_model = dict(
+ type='DetTTAModel',
+ tta_cfg=dict(nms=dict(type='nms', iou_threshold=0.65), max_per_img=100))
+
+img_scales = [(640, 640), (320, 320), (960, 960)]
+tta_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=None),
+ dict(
+ type='TestTimeAug',
+ transforms=[
+ [
+ dict(type='Resize', scale=s, keep_ratio=True)
+ for s in img_scales
+ ],
+ [
+ # ``RandomFlip`` must be placed before ``Pad``, otherwise
+ # bounding box coordinates after flipping cannot be
+ # recovered correctly.
+ dict(type='RandomFlip', prob=1.),
+ dict(type='RandomFlip', prob=0.)
+ ],
+ [
+ dict(
+ type='Pad',
+ pad_to_square=True,
+ pad_val=dict(img=(114.0, 114.0, 114.0))),
+ ],
+ [dict(type='LoadAnnotations', with_bbox=True)],
+ [
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction'))
+ ]
+ ])
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_x_8xb8-300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_x_8xb8-300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..34828e0363a2f282af59da74e805e59772dfeb69
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/configs/yolox/yolox_x_8xb8-300e_coco.py
@@ -0,0 +1,8 @@
+_base_ = './yolox_s_8xb8-300e_coco.py'
+
+# model settings
+model = dict(
+ backbone=dict(deepen_factor=1.33, widen_factor=1.25),
+ neck=dict(
+ in_channels=[320, 640, 1280], out_channels=320, num_csp_blocks=4),
+ bbox_head=dict(in_channels=320, feat_channels=320))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/gt_jsons/dior/xml_result_novel.json b/scripts/evaluation/FasterRCNN_score-mmdet/gt_jsons/dior/xml_result_novel.json
new file mode 100644
index 0000000000000000000000000000000000000000..447ec384d0c709dd3a2855cc9f5da260cdf43a2b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/gt_jsons/dior/xml_result_novel.json
@@ -0,0 +1,92209 @@
+{
+ "images": [
+ {
+ "file_name": "16375.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 0
+ },
+ {
+ "file_name": "20774.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1
+ },
+ {
+ "file_name": "18648.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2
+ },
+ {
+ "file_name": "20099.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 3
+ },
+ {
+ "file_name": "14830.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 4
+ },
+ {
+ "file_name": "13730.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 5
+ },
+ {
+ "file_name": "23410.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 6
+ },
+ {
+ "file_name": "13868.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 7
+ },
+ {
+ "file_name": "20988.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 8
+ },
+ {
+ "file_name": "21308.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 9
+ },
+ {
+ "file_name": "13279.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 10
+ },
+ {
+ "file_name": "16718.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 11
+ },
+ {
+ "file_name": "14158.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 12
+ },
+ {
+ "file_name": "20521.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 13
+ },
+ {
+ "file_name": "20038.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 14
+ },
+ {
+ "file_name": "14644.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 15
+ },
+ {
+ "file_name": "22188.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 16
+ },
+ {
+ "file_name": "12287.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 17
+ },
+ {
+ "file_name": "19222.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 18
+ },
+ {
+ "file_name": "17328.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 19
+ },
+ {
+ "file_name": "13212.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 20
+ },
+ {
+ "file_name": "12030.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 21
+ },
+ {
+ "file_name": "17820.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 22
+ },
+ {
+ "file_name": "13237.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 23
+ },
+ {
+ "file_name": "22447.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 24
+ },
+ {
+ "file_name": "11974.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 25
+ },
+ {
+ "file_name": "14655.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 26
+ },
+ {
+ "file_name": "19297.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 27
+ },
+ {
+ "file_name": "15867.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 28
+ },
+ {
+ "file_name": "16676.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 29
+ },
+ {
+ "file_name": "22723.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 30
+ },
+ {
+ "file_name": "22224.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 31
+ },
+ {
+ "file_name": "21262.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 32
+ },
+ {
+ "file_name": "16080.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 33
+ },
+ {
+ "file_name": "21727.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 34
+ },
+ {
+ "file_name": "16158.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 35
+ },
+ {
+ "file_name": "12232.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 36
+ },
+ {
+ "file_name": "14712.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 37
+ },
+ {
+ "file_name": "22900.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 38
+ },
+ {
+ "file_name": "16713.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 39
+ },
+ {
+ "file_name": "23420.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 40
+ },
+ {
+ "file_name": "17104.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 41
+ },
+ {
+ "file_name": "22751.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 42
+ },
+ {
+ "file_name": "19315.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 43
+ },
+ {
+ "file_name": "15538.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 44
+ },
+ {
+ "file_name": "18546.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 45
+ },
+ {
+ "file_name": "15488.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 46
+ },
+ {
+ "file_name": "17510.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 47
+ },
+ {
+ "file_name": "17815.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 48
+ },
+ {
+ "file_name": "16312.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 49
+ },
+ {
+ "file_name": "19725.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 50
+ },
+ {
+ "file_name": "14307.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 51
+ },
+ {
+ "file_name": "12804.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 52
+ },
+ {
+ "file_name": "15094.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 53
+ },
+ {
+ "file_name": "17530.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 54
+ },
+ {
+ "file_name": "18204.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 55
+ },
+ {
+ "file_name": "20059.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 56
+ },
+ {
+ "file_name": "18233.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 57
+ },
+ {
+ "file_name": "14788.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 58
+ },
+ {
+ "file_name": "22980.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 59
+ },
+ {
+ "file_name": "14469.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 60
+ },
+ {
+ "file_name": "16863.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 61
+ },
+ {
+ "file_name": "13715.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 62
+ },
+ {
+ "file_name": "15727.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 63
+ },
+ {
+ "file_name": "12584.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 64
+ },
+ {
+ "file_name": "13570.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 65
+ },
+ {
+ "file_name": "12698.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 66
+ },
+ {
+ "file_name": "20525.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 67
+ },
+ {
+ "file_name": "22020.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 68
+ },
+ {
+ "file_name": "15960.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 69
+ },
+ {
+ "file_name": "18381.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 70
+ },
+ {
+ "file_name": "19859.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 71
+ },
+ {
+ "file_name": "12901.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 72
+ },
+ {
+ "file_name": "12004.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 73
+ },
+ {
+ "file_name": "22684.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 74
+ },
+ {
+ "file_name": "12767.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 75
+ },
+ {
+ "file_name": "18009.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 76
+ },
+ {
+ "file_name": "14267.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 77
+ },
+ {
+ "file_name": "17624.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 78
+ },
+ {
+ "file_name": "14138.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 79
+ },
+ {
+ "file_name": "16254.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 80
+ },
+ {
+ "file_name": "21392.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 81
+ },
+ {
+ "file_name": "19209.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 82
+ },
+ {
+ "file_name": "23088.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 83
+ },
+ {
+ "file_name": "21086.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 84
+ },
+ {
+ "file_name": "22370.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 85
+ },
+ {
+ "file_name": "15273.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 86
+ },
+ {
+ "file_name": "14364.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 87
+ },
+ {
+ "file_name": "18585.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 88
+ },
+ {
+ "file_name": "20898.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 89
+ },
+ {
+ "file_name": "18583.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 90
+ },
+ {
+ "file_name": "16238.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 91
+ },
+ {
+ "file_name": "15805.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 92
+ },
+ {
+ "file_name": "20358.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 93
+ },
+ {
+ "file_name": "13836.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 94
+ },
+ {
+ "file_name": "15790.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 95
+ },
+ {
+ "file_name": "12390.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 96
+ },
+ {
+ "file_name": "19330.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 97
+ },
+ {
+ "file_name": "22625.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 98
+ },
+ {
+ "file_name": "18437.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 99
+ },
+ {
+ "file_name": "22442.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 100
+ },
+ {
+ "file_name": "12816.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 101
+ },
+ {
+ "file_name": "15528.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 102
+ },
+ {
+ "file_name": "22199.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 103
+ },
+ {
+ "file_name": "16469.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 104
+ },
+ {
+ "file_name": "19041.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 105
+ },
+ {
+ "file_name": "22124.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 106
+ },
+ {
+ "file_name": "21562.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 107
+ },
+ {
+ "file_name": "16071.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 108
+ },
+ {
+ "file_name": "20284.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 109
+ },
+ {
+ "file_name": "13313.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 110
+ },
+ {
+ "file_name": "11825.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 111
+ },
+ {
+ "file_name": "17989.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 112
+ },
+ {
+ "file_name": "23214.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 113
+ },
+ {
+ "file_name": "12613.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 114
+ },
+ {
+ "file_name": "20772.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 115
+ },
+ {
+ "file_name": "22285.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 116
+ },
+ {
+ "file_name": "19691.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 117
+ },
+ {
+ "file_name": "20434.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 118
+ },
+ {
+ "file_name": "21229.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 119
+ },
+ {
+ "file_name": "18793.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 120
+ },
+ {
+ "file_name": "12843.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 121
+ },
+ {
+ "file_name": "15941.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 122
+ },
+ {
+ "file_name": "23169.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 123
+ },
+ {
+ "file_name": "21634.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 124
+ },
+ {
+ "file_name": "20156.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 125
+ },
+ {
+ "file_name": "20799.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 126
+ },
+ {
+ "file_name": "11823.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 127
+ },
+ {
+ "file_name": "13871.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 128
+ },
+ {
+ "file_name": "15499.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 129
+ },
+ {
+ "file_name": "17737.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 130
+ },
+ {
+ "file_name": "21974.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 131
+ },
+ {
+ "file_name": "15686.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 132
+ },
+ {
+ "file_name": "16399.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 133
+ },
+ {
+ "file_name": "19162.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 134
+ },
+ {
+ "file_name": "12491.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 135
+ },
+ {
+ "file_name": "20665.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 136
+ },
+ {
+ "file_name": "23256.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 137
+ },
+ {
+ "file_name": "19363.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 138
+ },
+ {
+ "file_name": "14063.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 139
+ },
+ {
+ "file_name": "14670.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 140
+ },
+ {
+ "file_name": "20221.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 141
+ },
+ {
+ "file_name": "23442.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 142
+ },
+ {
+ "file_name": "22105.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 143
+ },
+ {
+ "file_name": "13377.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 144
+ },
+ {
+ "file_name": "14639.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 145
+ },
+ {
+ "file_name": "17918.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 146
+ },
+ {
+ "file_name": "14749.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 147
+ },
+ {
+ "file_name": "14595.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 148
+ },
+ {
+ "file_name": "22185.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 149
+ },
+ {
+ "file_name": "21291.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 150
+ },
+ {
+ "file_name": "18747.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 151
+ },
+ {
+ "file_name": "12539.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 152
+ },
+ {
+ "file_name": "16630.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 153
+ },
+ {
+ "file_name": "22417.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 154
+ },
+ {
+ "file_name": "17856.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 155
+ },
+ {
+ "file_name": "12148.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 156
+ },
+ {
+ "file_name": "20796.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 157
+ },
+ {
+ "file_name": "18549.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 158
+ },
+ {
+ "file_name": "18453.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 159
+ },
+ {
+ "file_name": "20412.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 160
+ },
+ {
+ "file_name": "23454.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 161
+ },
+ {
+ "file_name": "21182.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 162
+ },
+ {
+ "file_name": "13255.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 163
+ },
+ {
+ "file_name": "17178.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 164
+ },
+ {
+ "file_name": "18817.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 165
+ },
+ {
+ "file_name": "17915.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 166
+ },
+ {
+ "file_name": "17769.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 167
+ },
+ {
+ "file_name": "22685.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 168
+ },
+ {
+ "file_name": "16023.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 169
+ },
+ {
+ "file_name": "20182.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 170
+ },
+ {
+ "file_name": "21561.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 171
+ },
+ {
+ "file_name": "21930.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 172
+ },
+ {
+ "file_name": "18467.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 173
+ },
+ {
+ "file_name": "15014.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 174
+ },
+ {
+ "file_name": "19908.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 175
+ },
+ {
+ "file_name": "18777.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 176
+ },
+ {
+ "file_name": "18417.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 177
+ },
+ {
+ "file_name": "22677.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 178
+ },
+ {
+ "file_name": "13032.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 179
+ },
+ {
+ "file_name": "21105.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 180
+ },
+ {
+ "file_name": "15783.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 181
+ },
+ {
+ "file_name": "14589.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 182
+ },
+ {
+ "file_name": "17672.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 183
+ },
+ {
+ "file_name": "21862.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 184
+ },
+ {
+ "file_name": "13793.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 185
+ },
+ {
+ "file_name": "22969.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 186
+ },
+ {
+ "file_name": "12759.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 187
+ },
+ {
+ "file_name": "14265.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 188
+ },
+ {
+ "file_name": "21195.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 189
+ },
+ {
+ "file_name": "19612.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 190
+ },
+ {
+ "file_name": "13563.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 191
+ },
+ {
+ "file_name": "14521.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 192
+ },
+ {
+ "file_name": "22047.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 193
+ },
+ {
+ "file_name": "22136.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 194
+ },
+ {
+ "file_name": "22210.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 195
+ },
+ {
+ "file_name": "19712.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 196
+ },
+ {
+ "file_name": "19211.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 197
+ },
+ {
+ "file_name": "11860.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 198
+ },
+ {
+ "file_name": "20680.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 199
+ },
+ {
+ "file_name": "23433.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 200
+ },
+ {
+ "file_name": "21667.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 201
+ },
+ {
+ "file_name": "18601.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 202
+ },
+ {
+ "file_name": "16895.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 203
+ },
+ {
+ "file_name": "14197.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 204
+ },
+ {
+ "file_name": "21869.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 205
+ },
+ {
+ "file_name": "16796.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 206
+ },
+ {
+ "file_name": "17077.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 207
+ },
+ {
+ "file_name": "15355.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 208
+ },
+ {
+ "file_name": "13950.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 209
+ },
+ {
+ "file_name": "20971.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 210
+ },
+ {
+ "file_name": "16030.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 211
+ },
+ {
+ "file_name": "20896.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 212
+ },
+ {
+ "file_name": "18325.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 213
+ },
+ {
+ "file_name": "12986.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 214
+ },
+ {
+ "file_name": "15991.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 215
+ },
+ {
+ "file_name": "17211.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 216
+ },
+ {
+ "file_name": "22741.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 217
+ },
+ {
+ "file_name": "20283.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 218
+ },
+ {
+ "file_name": "21244.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 219
+ },
+ {
+ "file_name": "23456.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 220
+ },
+ {
+ "file_name": "13241.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 221
+ },
+ {
+ "file_name": "22176.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 222
+ },
+ {
+ "file_name": "13386.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 223
+ },
+ {
+ "file_name": "22916.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 224
+ },
+ {
+ "file_name": "20094.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 225
+ },
+ {
+ "file_name": "13897.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 226
+ },
+ {
+ "file_name": "21875.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 227
+ },
+ {
+ "file_name": "21326.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 228
+ },
+ {
+ "file_name": "14600.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 229
+ },
+ {
+ "file_name": "23431.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 230
+ },
+ {
+ "file_name": "16907.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 231
+ },
+ {
+ "file_name": "15791.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 232
+ },
+ {
+ "file_name": "17692.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 233
+ },
+ {
+ "file_name": "21497.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 234
+ },
+ {
+ "file_name": "15044.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 235
+ },
+ {
+ "file_name": "20614.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 236
+ },
+ {
+ "file_name": "18582.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 237
+ },
+ {
+ "file_name": "22202.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 238
+ },
+ {
+ "file_name": "12967.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 239
+ },
+ {
+ "file_name": "15263.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 240
+ },
+ {
+ "file_name": "18901.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 241
+ },
+ {
+ "file_name": "15239.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 242
+ },
+ {
+ "file_name": "13333.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 243
+ },
+ {
+ "file_name": "16100.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 244
+ },
+ {
+ "file_name": "17278.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 245
+ },
+ {
+ "file_name": "12649.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 246
+ },
+ {
+ "file_name": "13496.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 247
+ },
+ {
+ "file_name": "22855.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 248
+ },
+ {
+ "file_name": "19262.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 249
+ },
+ {
+ "file_name": "23226.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 250
+ },
+ {
+ "file_name": "12705.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 251
+ },
+ {
+ "file_name": "19631.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 252
+ },
+ {
+ "file_name": "20135.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 253
+ },
+ {
+ "file_name": "13441.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 254
+ },
+ {
+ "file_name": "11734.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 255
+ },
+ {
+ "file_name": "22892.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 256
+ },
+ {
+ "file_name": "22956.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 257
+ },
+ {
+ "file_name": "19627.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 258
+ },
+ {
+ "file_name": "21171.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 259
+ },
+ {
+ "file_name": "21341.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 260
+ },
+ {
+ "file_name": "18943.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 261
+ },
+ {
+ "file_name": "20366.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 262
+ },
+ {
+ "file_name": "15151.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 263
+ },
+ {
+ "file_name": "17294.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 264
+ },
+ {
+ "file_name": "21069.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 265
+ },
+ {
+ "file_name": "20617.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 266
+ },
+ {
+ "file_name": "21824.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 267
+ },
+ {
+ "file_name": "20632.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 268
+ },
+ {
+ "file_name": "21273.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 269
+ },
+ {
+ "file_name": "12255.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 270
+ },
+ {
+ "file_name": "14581.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 271
+ },
+ {
+ "file_name": "14994.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 272
+ },
+ {
+ "file_name": "20802.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 273
+ },
+ {
+ "file_name": "18493.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 274
+ },
+ {
+ "file_name": "15910.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 275
+ },
+ {
+ "file_name": "22318.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 276
+ },
+ {
+ "file_name": "17564.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 277
+ },
+ {
+ "file_name": "12971.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 278
+ },
+ {
+ "file_name": "22129.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 279
+ },
+ {
+ "file_name": "17449.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 280
+ },
+ {
+ "file_name": "15311.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 281
+ },
+ {
+ "file_name": "14204.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 282
+ },
+ {
+ "file_name": "14533.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 283
+ },
+ {
+ "file_name": "16574.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 284
+ },
+ {
+ "file_name": "18111.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 285
+ },
+ {
+ "file_name": "17671.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 286
+ },
+ {
+ "file_name": "19081.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 287
+ },
+ {
+ "file_name": "16627.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 288
+ },
+ {
+ "file_name": "23057.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 289
+ },
+ {
+ "file_name": "20750.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 290
+ },
+ {
+ "file_name": "19798.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 291
+ },
+ {
+ "file_name": "12914.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 292
+ },
+ {
+ "file_name": "18166.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 293
+ },
+ {
+ "file_name": "21402.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 294
+ },
+ {
+ "file_name": "13385.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 295
+ },
+ {
+ "file_name": "21153.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 296
+ },
+ {
+ "file_name": "17862.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 297
+ },
+ {
+ "file_name": "18174.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 298
+ },
+ {
+ "file_name": "20108.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 299
+ },
+ {
+ "file_name": "19583.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 300
+ },
+ {
+ "file_name": "22701.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 301
+ },
+ {
+ "file_name": "21737.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 302
+ },
+ {
+ "file_name": "13285.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 303
+ },
+ {
+ "file_name": "13448.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 304
+ },
+ {
+ "file_name": "12308.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 305
+ },
+ {
+ "file_name": "15731.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 306
+ },
+ {
+ "file_name": "12683.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 307
+ },
+ {
+ "file_name": "15775.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 308
+ },
+ {
+ "file_name": "17063.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 309
+ },
+ {
+ "file_name": "17685.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 310
+ },
+ {
+ "file_name": "15677.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 311
+ },
+ {
+ "file_name": "21124.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 312
+ },
+ {
+ "file_name": "12678.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 313
+ },
+ {
+ "file_name": "18267.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 314
+ },
+ {
+ "file_name": "22996.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 315
+ },
+ {
+ "file_name": "16314.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 316
+ },
+ {
+ "file_name": "12755.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 317
+ },
+ {
+ "file_name": "12181.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 318
+ },
+ {
+ "file_name": "12398.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 319
+ },
+ {
+ "file_name": "15390.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 320
+ },
+ {
+ "file_name": "22303.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 321
+ },
+ {
+ "file_name": "14986.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 322
+ },
+ {
+ "file_name": "20607.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 323
+ },
+ {
+ "file_name": "12939.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 324
+ },
+ {
+ "file_name": "15170.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 325
+ },
+ {
+ "file_name": "20174.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 326
+ },
+ {
+ "file_name": "13045.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 327
+ },
+ {
+ "file_name": "18302.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 328
+ },
+ {
+ "file_name": "21209.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 329
+ },
+ {
+ "file_name": "20371.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 330
+ },
+ {
+ "file_name": "21178.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 331
+ },
+ {
+ "file_name": "16648.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 332
+ },
+ {
+ "file_name": "15269.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 333
+ },
+ {
+ "file_name": "21972.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 334
+ },
+ {
+ "file_name": "18336.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 335
+ },
+ {
+ "file_name": "16656.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 336
+ },
+ {
+ "file_name": "17318.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 337
+ },
+ {
+ "file_name": "11770.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 338
+ },
+ {
+ "file_name": "15603.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 339
+ },
+ {
+ "file_name": "20436.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 340
+ },
+ {
+ "file_name": "18470.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 341
+ },
+ {
+ "file_name": "16948.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 342
+ },
+ {
+ "file_name": "19428.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 343
+ },
+ {
+ "file_name": "18715.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 344
+ },
+ {
+ "file_name": "13962.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 345
+ },
+ {
+ "file_name": "16842.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 346
+ },
+ {
+ "file_name": "12279.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 347
+ },
+ {
+ "file_name": "22585.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 348
+ },
+ {
+ "file_name": "12575.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 349
+ },
+ {
+ "file_name": "12736.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 350
+ },
+ {
+ "file_name": "16603.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 351
+ },
+ {
+ "file_name": "18634.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 352
+ },
+ {
+ "file_name": "17038.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 353
+ },
+ {
+ "file_name": "14961.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 354
+ },
+ {
+ "file_name": "17489.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 355
+ },
+ {
+ "file_name": "20165.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 356
+ },
+ {
+ "file_name": "11775.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 357
+ },
+ {
+ "file_name": "21295.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 358
+ },
+ {
+ "file_name": "15994.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 359
+ },
+ {
+ "file_name": "17578.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 360
+ },
+ {
+ "file_name": "12888.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 361
+ },
+ {
+ "file_name": "22817.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 362
+ },
+ {
+ "file_name": "22002.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 363
+ },
+ {
+ "file_name": "19652.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 364
+ },
+ {
+ "file_name": "16819.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 365
+ },
+ {
+ "file_name": "21101.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 366
+ },
+ {
+ "file_name": "11966.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 367
+ },
+ {
+ "file_name": "21108.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 368
+ },
+ {
+ "file_name": "14682.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 369
+ },
+ {
+ "file_name": "20372.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 370
+ },
+ {
+ "file_name": "16845.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 371
+ },
+ {
+ "file_name": "17751.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 372
+ },
+ {
+ "file_name": "13782.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 373
+ },
+ {
+ "file_name": "12078.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 374
+ },
+ {
+ "file_name": "16005.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 375
+ },
+ {
+ "file_name": "12294.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 376
+ },
+ {
+ "file_name": "13996.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 377
+ },
+ {
+ "file_name": "12766.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 378
+ },
+ {
+ "file_name": "20362.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 379
+ },
+ {
+ "file_name": "17740.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 380
+ },
+ {
+ "file_name": "22248.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 381
+ },
+ {
+ "file_name": "12100.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 382
+ },
+ {
+ "file_name": "14522.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 383
+ },
+ {
+ "file_name": "22830.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 384
+ },
+ {
+ "file_name": "12906.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 385
+ },
+ {
+ "file_name": "14124.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 386
+ },
+ {
+ "file_name": "14889.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 387
+ },
+ {
+ "file_name": "17025.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 388
+ },
+ {
+ "file_name": "22825.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 389
+ },
+ {
+ "file_name": "14175.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 390
+ },
+ {
+ "file_name": "11999.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 391
+ },
+ {
+ "file_name": "17625.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 392
+ },
+ {
+ "file_name": "11729.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 393
+ },
+ {
+ "file_name": "20316.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 394
+ },
+ {
+ "file_name": "21642.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 395
+ },
+ {
+ "file_name": "11874.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 396
+ },
+ {
+ "file_name": "13378.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 397
+ },
+ {
+ "file_name": "12068.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 398
+ },
+ {
+ "file_name": "13307.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 399
+ },
+ {
+ "file_name": "13046.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 400
+ },
+ {
+ "file_name": "14534.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 401
+ },
+ {
+ "file_name": "20671.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 402
+ },
+ {
+ "file_name": "18235.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 403
+ },
+ {
+ "file_name": "14423.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 404
+ },
+ {
+ "file_name": "12844.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 405
+ },
+ {
+ "file_name": "22495.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 406
+ },
+ {
+ "file_name": "14130.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 407
+ },
+ {
+ "file_name": "18028.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 408
+ },
+ {
+ "file_name": "15124.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 409
+ },
+ {
+ "file_name": "20914.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 410
+ },
+ {
+ "file_name": "13624.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 411
+ },
+ {
+ "file_name": "12900.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 412
+ },
+ {
+ "file_name": "15626.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 413
+ },
+ {
+ "file_name": "19684.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 414
+ },
+ {
+ "file_name": "11782.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 415
+ },
+ {
+ "file_name": "17325.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 416
+ },
+ {
+ "file_name": "15491.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 417
+ },
+ {
+ "file_name": "17661.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 418
+ },
+ {
+ "file_name": "22450.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 419
+ },
+ {
+ "file_name": "11976.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 420
+ },
+ {
+ "file_name": "12642.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 421
+ },
+ {
+ "file_name": "14060.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 422
+ },
+ {
+ "file_name": "14388.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 423
+ },
+ {
+ "file_name": "15030.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 424
+ },
+ {
+ "file_name": "19842.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 425
+ },
+ {
+ "file_name": "15690.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 426
+ },
+ {
+ "file_name": "16750.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 427
+ },
+ {
+ "file_name": "16329.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 428
+ },
+ {
+ "file_name": "12336.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 429
+ },
+ {
+ "file_name": "18838.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 430
+ },
+ {
+ "file_name": "22554.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 431
+ },
+ {
+ "file_name": "23149.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 432
+ },
+ {
+ "file_name": "13796.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 433
+ },
+ {
+ "file_name": "12325.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 434
+ },
+ {
+ "file_name": "21729.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 435
+ },
+ {
+ "file_name": "19819.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 436
+ },
+ {
+ "file_name": "19737.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 437
+ },
+ {
+ "file_name": "12717.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 438
+ },
+ {
+ "file_name": "17582.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 439
+ },
+ {
+ "file_name": "18903.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 440
+ },
+ {
+ "file_name": "18503.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 441
+ },
+ {
+ "file_name": "18933.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 442
+ },
+ {
+ "file_name": "14480.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 443
+ },
+ {
+ "file_name": "16872.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 444
+ },
+ {
+ "file_name": "18320.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 445
+ },
+ {
+ "file_name": "15152.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 446
+ },
+ {
+ "file_name": "15091.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 447
+ },
+ {
+ "file_name": "15900.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 448
+ },
+ {
+ "file_name": "17103.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 449
+ },
+ {
+ "file_name": "20997.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 450
+ },
+ {
+ "file_name": "15978.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 451
+ },
+ {
+ "file_name": "19582.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 452
+ },
+ {
+ "file_name": "14531.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 453
+ },
+ {
+ "file_name": "17326.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 454
+ },
+ {
+ "file_name": "15127.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 455
+ },
+ {
+ "file_name": "22881.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 456
+ },
+ {
+ "file_name": "16331.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 457
+ },
+ {
+ "file_name": "16567.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 458
+ },
+ {
+ "file_name": "13323.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 459
+ },
+ {
+ "file_name": "17164.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 460
+ },
+ {
+ "file_name": "20587.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 461
+ },
+ {
+ "file_name": "16580.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 462
+ },
+ {
+ "file_name": "21510.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 463
+ },
+ {
+ "file_name": "16754.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 464
+ },
+ {
+ "file_name": "18024.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 465
+ },
+ {
+ "file_name": "17529.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 466
+ },
+ {
+ "file_name": "14683.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 467
+ },
+ {
+ "file_name": "22934.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 468
+ },
+ {
+ "file_name": "16788.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 469
+ },
+ {
+ "file_name": "20652.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 470
+ },
+ {
+ "file_name": "20515.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 471
+ },
+ {
+ "file_name": "19617.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 472
+ },
+ {
+ "file_name": "19248.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 473
+ },
+ {
+ "file_name": "13802.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 474
+ },
+ {
+ "file_name": "23014.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 475
+ },
+ {
+ "file_name": "13362.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 476
+ },
+ {
+ "file_name": "18021.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 477
+ },
+ {
+ "file_name": "15361.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 478
+ },
+ {
+ "file_name": "23136.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 479
+ },
+ {
+ "file_name": "12903.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 480
+ },
+ {
+ "file_name": "15384.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 481
+ },
+ {
+ "file_name": "15485.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 482
+ },
+ {
+ "file_name": "15286.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 483
+ },
+ {
+ "file_name": "21154.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 484
+ },
+ {
+ "file_name": "14405.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 485
+ },
+ {
+ "file_name": "19520.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 486
+ },
+ {
+ "file_name": "15578.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 487
+ },
+ {
+ "file_name": "21726.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 488
+ },
+ {
+ "file_name": "14128.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 489
+ },
+ {
+ "file_name": "12718.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 490
+ },
+ {
+ "file_name": "12623.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 491
+ },
+ {
+ "file_name": "15426.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 492
+ },
+ {
+ "file_name": "22613.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 493
+ },
+ {
+ "file_name": "17757.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 494
+ },
+ {
+ "file_name": "16421.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 495
+ },
+ {
+ "file_name": "17069.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 496
+ },
+ {
+ "file_name": "13072.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 497
+ },
+ {
+ "file_name": "14105.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 498
+ },
+ {
+ "file_name": "13865.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 499
+ },
+ {
+ "file_name": "19331.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 500
+ },
+ {
+ "file_name": "16123.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 501
+ },
+ {
+ "file_name": "12242.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 502
+ },
+ {
+ "file_name": "15041.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 503
+ },
+ {
+ "file_name": "21944.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 504
+ },
+ {
+ "file_name": "19395.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 505
+ },
+ {
+ "file_name": "20714.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 506
+ },
+ {
+ "file_name": "12732.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 507
+ },
+ {
+ "file_name": "13819.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 508
+ },
+ {
+ "file_name": "13536.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 509
+ },
+ {
+ "file_name": "20524.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 510
+ },
+ {
+ "file_name": "12577.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 511
+ },
+ {
+ "file_name": "13491.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 512
+ },
+ {
+ "file_name": "19086.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 513
+ },
+ {
+ "file_name": "16857.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 514
+ },
+ {
+ "file_name": "14694.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 515
+ },
+ {
+ "file_name": "16265.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 516
+ },
+ {
+ "file_name": "17035.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 517
+ },
+ {
+ "file_name": "17778.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 518
+ },
+ {
+ "file_name": "19422.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 519
+ },
+ {
+ "file_name": "16463.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 520
+ },
+ {
+ "file_name": "19957.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 521
+ },
+ {
+ "file_name": "18608.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 522
+ },
+ {
+ "file_name": "19836.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 523
+ },
+ {
+ "file_name": "19425.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 524
+ },
+ {
+ "file_name": "15070.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 525
+ },
+ {
+ "file_name": "12618.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 526
+ },
+ {
+ "file_name": "15276.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 527
+ },
+ {
+ "file_name": "14945.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 528
+ },
+ {
+ "file_name": "21170.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 529
+ },
+ {
+ "file_name": "17720.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 530
+ },
+ {
+ "file_name": "17463.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 531
+ },
+ {
+ "file_name": "19059.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 532
+ },
+ {
+ "file_name": "15743.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 533
+ },
+ {
+ "file_name": "16465.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 534
+ },
+ {
+ "file_name": "18177.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 535
+ },
+ {
+ "file_name": "17482.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 536
+ },
+ {
+ "file_name": "13652.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 537
+ },
+ {
+ "file_name": "19473.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 538
+ },
+ {
+ "file_name": "14955.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 539
+ },
+ {
+ "file_name": "12837.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 540
+ },
+ {
+ "file_name": "20540.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 541
+ },
+ {
+ "file_name": "20365.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 542
+ },
+ {
+ "file_name": "11810.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 543
+ },
+ {
+ "file_name": "11952.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 544
+ },
+ {
+ "file_name": "16325.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 545
+ },
+ {
+ "file_name": "15319.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 546
+ },
+ {
+ "file_name": "21722.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 547
+ },
+ {
+ "file_name": "12421.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 548
+ },
+ {
+ "file_name": "17951.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 549
+ },
+ {
+ "file_name": "18808.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 550
+ },
+ {
+ "file_name": "14096.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 551
+ },
+ {
+ "file_name": "15638.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 552
+ },
+ {
+ "file_name": "19995.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 553
+ },
+ {
+ "file_name": "19258.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 554
+ },
+ {
+ "file_name": "22365.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 555
+ },
+ {
+ "file_name": "19574.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 556
+ },
+ {
+ "file_name": "17894.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 557
+ },
+ {
+ "file_name": "20471.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 558
+ },
+ {
+ "file_name": "16705.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 559
+ },
+ {
+ "file_name": "17235.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 560
+ },
+ {
+ "file_name": "15702.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 561
+ },
+ {
+ "file_name": "13969.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 562
+ },
+ {
+ "file_name": "13734.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 563
+ },
+ {
+ "file_name": "18155.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 564
+ },
+ {
+ "file_name": "23434.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 565
+ },
+ {
+ "file_name": "20849.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 566
+ },
+ {
+ "file_name": "15597.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 567
+ },
+ {
+ "file_name": "20222.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 568
+ },
+ {
+ "file_name": "19522.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 569
+ },
+ {
+ "file_name": "21302.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 570
+ },
+ {
+ "file_name": "12161.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 571
+ },
+ {
+ "file_name": "16361.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 572
+ },
+ {
+ "file_name": "13458.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 573
+ },
+ {
+ "file_name": "16640.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 574
+ },
+ {
+ "file_name": "16571.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 575
+ },
+ {
+ "file_name": "18720.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 576
+ },
+ {
+ "file_name": "17083.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 577
+ },
+ {
+ "file_name": "12142.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 578
+ },
+ {
+ "file_name": "22798.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 579
+ },
+ {
+ "file_name": "17363.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 580
+ },
+ {
+ "file_name": "17024.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 581
+ },
+ {
+ "file_name": "12008.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 582
+ },
+ {
+ "file_name": "18187.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 583
+ },
+ {
+ "file_name": "13303.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 584
+ },
+ {
+ "file_name": "15510.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 585
+ },
+ {
+ "file_name": "22378.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 586
+ },
+ {
+ "file_name": "14088.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 587
+ },
+ {
+ "file_name": "13169.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 588
+ },
+ {
+ "file_name": "12890.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 589
+ },
+ {
+ "file_name": "21463.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 590
+ },
+ {
+ "file_name": "17161.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 591
+ },
+ {
+ "file_name": "14724.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 592
+ },
+ {
+ "file_name": "22617.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 593
+ },
+ {
+ "file_name": "14013.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 594
+ },
+ {
+ "file_name": "12467.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 595
+ },
+ {
+ "file_name": "22108.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 596
+ },
+ {
+ "file_name": "16582.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 597
+ },
+ {
+ "file_name": "11790.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 598
+ },
+ {
+ "file_name": "22480.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 599
+ },
+ {
+ "file_name": "16428.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 600
+ },
+ {
+ "file_name": "12791.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 601
+ },
+ {
+ "file_name": "13075.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 602
+ },
+ {
+ "file_name": "20506.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 603
+ },
+ {
+ "file_name": "15798.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 604
+ },
+ {
+ "file_name": "14964.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 605
+ },
+ {
+ "file_name": "14061.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 606
+ },
+ {
+ "file_name": "20470.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 607
+ },
+ {
+ "file_name": "19512.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 608
+ },
+ {
+ "file_name": "16472.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 609
+ },
+ {
+ "file_name": "17952.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 610
+ },
+ {
+ "file_name": "22530.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 611
+ },
+ {
+ "file_name": "13471.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 612
+ },
+ {
+ "file_name": "16225.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 613
+ },
+ {
+ "file_name": "14442.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 614
+ },
+ {
+ "file_name": "17066.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 615
+ },
+ {
+ "file_name": "12359.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 616
+ },
+ {
+ "file_name": "19031.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 617
+ },
+ {
+ "file_name": "19195.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 618
+ },
+ {
+ "file_name": "12849.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 619
+ },
+ {
+ "file_name": "13187.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 620
+ },
+ {
+ "file_name": "18227.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 621
+ },
+ {
+ "file_name": "23451.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 622
+ },
+ {
+ "file_name": "23369.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 623
+ },
+ {
+ "file_name": "14657.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 624
+ },
+ {
+ "file_name": "19467.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 625
+ },
+ {
+ "file_name": "17968.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 626
+ },
+ {
+ "file_name": "18821.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 627
+ },
+ {
+ "file_name": "14366.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 628
+ },
+ {
+ "file_name": "15561.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 629
+ },
+ {
+ "file_name": "20883.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 630
+ },
+ {
+ "file_name": "16711.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 631
+ },
+ {
+ "file_name": "21469.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 632
+ },
+ {
+ "file_name": "20456.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 633
+ },
+ {
+ "file_name": "13404.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 634
+ },
+ {
+ "file_name": "19408.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 635
+ },
+ {
+ "file_name": "18968.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 636
+ },
+ {
+ "file_name": "23128.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 637
+ },
+ {
+ "file_name": "22092.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 638
+ },
+ {
+ "file_name": "22005.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 639
+ },
+ {
+ "file_name": "19471.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 640
+ },
+ {
+ "file_name": "12448.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 641
+ },
+ {
+ "file_name": "19283.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 642
+ },
+ {
+ "file_name": "19067.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 643
+ },
+ {
+ "file_name": "16987.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 644
+ },
+ {
+ "file_name": "20249.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 645
+ },
+ {
+ "file_name": "13392.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 646
+ },
+ {
+ "file_name": "21809.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 647
+ },
+ {
+ "file_name": "16856.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 648
+ },
+ {
+ "file_name": "22326.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 649
+ },
+ {
+ "file_name": "15118.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 650
+ },
+ {
+ "file_name": "15017.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 651
+ },
+ {
+ "file_name": "12042.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 652
+ },
+ {
+ "file_name": "14679.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 653
+ },
+ {
+ "file_name": "21424.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 654
+ },
+ {
+ "file_name": "18398.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 655
+ },
+ {
+ "file_name": "14574.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 656
+ },
+ {
+ "file_name": "16778.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 657
+ },
+ {
+ "file_name": "18245.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 658
+ },
+ {
+ "file_name": "13182.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 659
+ },
+ {
+ "file_name": "19340.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 660
+ },
+ {
+ "file_name": "13417.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 661
+ },
+ {
+ "file_name": "20943.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 662
+ },
+ {
+ "file_name": "16212.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 663
+ },
+ {
+ "file_name": "18560.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 664
+ },
+ {
+ "file_name": "18708.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 665
+ },
+ {
+ "file_name": "22797.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 666
+ },
+ {
+ "file_name": "19489.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 667
+ },
+ {
+ "file_name": "13759.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 668
+ },
+ {
+ "file_name": "20827.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 669
+ },
+ {
+ "file_name": "16179.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 670
+ },
+ {
+ "file_name": "13945.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 671
+ },
+ {
+ "file_name": "15625.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 672
+ },
+ {
+ "file_name": "17096.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 673
+ },
+ {
+ "file_name": "19001.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 674
+ },
+ {
+ "file_name": "12315.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 675
+ },
+ {
+ "file_name": "21970.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 676
+ },
+ {
+ "file_name": "23370.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 677
+ },
+ {
+ "file_name": "13367.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 678
+ },
+ {
+ "file_name": "17950.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 679
+ },
+ {
+ "file_name": "16932.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 680
+ },
+ {
+ "file_name": "22580.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 681
+ },
+ {
+ "file_name": "19372.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 682
+ },
+ {
+ "file_name": "15693.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 683
+ },
+ {
+ "file_name": "14369.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 684
+ },
+ {
+ "file_name": "19325.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 685
+ },
+ {
+ "file_name": "16777.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 686
+ },
+ {
+ "file_name": "15121.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 687
+ },
+ {
+ "file_name": "21009.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 688
+ },
+ {
+ "file_name": "17585.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 689
+ },
+ {
+ "file_name": "16994.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 690
+ },
+ {
+ "file_name": "20017.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 691
+ },
+ {
+ "file_name": "16073.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 692
+ },
+ {
+ "file_name": "12157.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 693
+ },
+ {
+ "file_name": "18786.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 694
+ },
+ {
+ "file_name": "18750.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 695
+ },
+ {
+ "file_name": "20280.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 696
+ },
+ {
+ "file_name": "14796.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 697
+ },
+ {
+ "file_name": "22688.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 698
+ },
+ {
+ "file_name": "13003.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 699
+ },
+ {
+ "file_name": "13965.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 700
+ },
+ {
+ "file_name": "19023.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 701
+ },
+ {
+ "file_name": "21304.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 702
+ },
+ {
+ "file_name": "13893.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 703
+ },
+ {
+ "file_name": "16698.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 704
+ },
+ {
+ "file_name": "22937.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 705
+ },
+ {
+ "file_name": "19693.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 706
+ },
+ {
+ "file_name": "15530.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 707
+ },
+ {
+ "file_name": "17557.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 708
+ },
+ {
+ "file_name": "13778.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 709
+ },
+ {
+ "file_name": "12667.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 710
+ },
+ {
+ "file_name": "20937.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 711
+ },
+ {
+ "file_name": "19354.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 712
+ },
+ {
+ "file_name": "17837.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 713
+ },
+ {
+ "file_name": "20150.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 714
+ },
+ {
+ "file_name": "17773.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 715
+ },
+ {
+ "file_name": "16283.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 716
+ },
+ {
+ "file_name": "19463.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 717
+ },
+ {
+ "file_name": "20347.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 718
+ },
+ {
+ "file_name": "18385.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 719
+ },
+ {
+ "file_name": "19399.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 720
+ },
+ {
+ "file_name": "15997.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 721
+ },
+ {
+ "file_name": "20708.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 722
+ },
+ {
+ "file_name": "22227.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 723
+ },
+ {
+ "file_name": "15982.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 724
+ },
+ {
+ "file_name": "19893.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 725
+ },
+ {
+ "file_name": "17675.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 726
+ },
+ {
+ "file_name": "20509.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 727
+ },
+ {
+ "file_name": "18609.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 728
+ },
+ {
+ "file_name": "21447.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 729
+ },
+ {
+ "file_name": "18677.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 730
+ },
+ {
+ "file_name": "22041.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 731
+ },
+ {
+ "file_name": "20669.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 732
+ },
+ {
+ "file_name": "12028.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 733
+ },
+ {
+ "file_name": "22302.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 734
+ },
+ {
+ "file_name": "13214.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 735
+ },
+ {
+ "file_name": "13216.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 736
+ },
+ {
+ "file_name": "22026.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 737
+ },
+ {
+ "file_name": "11840.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 738
+ },
+ {
+ "file_name": "17702.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 739
+ },
+ {
+ "file_name": "23050.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 740
+ },
+ {
+ "file_name": "14372.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 741
+ },
+ {
+ "file_name": "15344.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 742
+ },
+ {
+ "file_name": "16836.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 743
+ },
+ {
+ "file_name": "12828.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 744
+ },
+ {
+ "file_name": "13432.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 745
+ },
+ {
+ "file_name": "17360.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 746
+ },
+ {
+ "file_name": "17397.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 747
+ },
+ {
+ "file_name": "13899.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 748
+ },
+ {
+ "file_name": "21501.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 749
+ },
+ {
+ "file_name": "19529.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 750
+ },
+ {
+ "file_name": "11989.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 751
+ },
+ {
+ "file_name": "15214.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 752
+ },
+ {
+ "file_name": "18346.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 753
+ },
+ {
+ "file_name": "22657.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 754
+ },
+ {
+ "file_name": "17884.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 755
+ },
+ {
+ "file_name": "20603.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 756
+ },
+ {
+ "file_name": "18697.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 757
+ },
+ {
+ "file_name": "12569.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 758
+ },
+ {
+ "file_name": "13813.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 759
+ },
+ {
+ "file_name": "18631.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 760
+ },
+ {
+ "file_name": "19869.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 761
+ },
+ {
+ "file_name": "21542.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 762
+ },
+ {
+ "file_name": "16002.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 763
+ },
+ {
+ "file_name": "14705.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 764
+ },
+ {
+ "file_name": "16107.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 765
+ },
+ {
+ "file_name": "13981.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 766
+ },
+ {
+ "file_name": "14780.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 767
+ },
+ {
+ "file_name": "12483.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 768
+ },
+ {
+ "file_name": "22437.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 769
+ },
+ {
+ "file_name": "16548.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 770
+ },
+ {
+ "file_name": "16494.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 771
+ },
+ {
+ "file_name": "11949.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 772
+ },
+ {
+ "file_name": "17481.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 773
+ },
+ {
+ "file_name": "12793.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 774
+ },
+ {
+ "file_name": "14129.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 775
+ },
+ {
+ "file_name": "15248.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 776
+ },
+ {
+ "file_name": "14832.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 777
+ },
+ {
+ "file_name": "14968.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 778
+ },
+ {
+ "file_name": "13226.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 779
+ },
+ {
+ "file_name": "21057.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 780
+ },
+ {
+ "file_name": "15217.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 781
+ },
+ {
+ "file_name": "12942.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 782
+ },
+ {
+ "file_name": "20570.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 783
+ },
+ {
+ "file_name": "17966.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 784
+ },
+ {
+ "file_name": "12434.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 785
+ },
+ {
+ "file_name": "13835.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 786
+ },
+ {
+ "file_name": "15135.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 787
+ },
+ {
+ "file_name": "21686.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 788
+ },
+ {
+ "file_name": "20865.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 789
+ },
+ {
+ "file_name": "19969.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 790
+ },
+ {
+ "file_name": "15207.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 791
+ },
+ {
+ "file_name": "19381.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 792
+ },
+ {
+ "file_name": "16919.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 793
+ },
+ {
+ "file_name": "12703.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 794
+ },
+ {
+ "file_name": "15105.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 795
+ },
+ {
+ "file_name": "14386.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 796
+ },
+ {
+ "file_name": "16847.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 797
+ },
+ {
+ "file_name": "22151.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 798
+ },
+ {
+ "file_name": "13701.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 799
+ },
+ {
+ "file_name": "18391.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 800
+ },
+ {
+ "file_name": "22911.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 801
+ },
+ {
+ "file_name": "18230.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 802
+ },
+ {
+ "file_name": "12305.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 803
+ },
+ {
+ "file_name": "19563.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 804
+ },
+ {
+ "file_name": "17330.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 805
+ },
+ {
+ "file_name": "18471.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 806
+ },
+ {
+ "file_name": "19936.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 807
+ },
+ {
+ "file_name": "12361.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 808
+ },
+ {
+ "file_name": "22505.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 809
+ },
+ {
+ "file_name": "14858.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 810
+ },
+ {
+ "file_name": "16229.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 811
+ },
+ {
+ "file_name": "16826.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 812
+ },
+ {
+ "file_name": "19910.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 813
+ },
+ {
+ "file_name": "16288.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 814
+ },
+ {
+ "file_name": "16626.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 815
+ },
+ {
+ "file_name": "17133.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 816
+ },
+ {
+ "file_name": "19736.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 817
+ },
+ {
+ "file_name": "21707.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 818
+ },
+ {
+ "file_name": "14178.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 819
+ },
+ {
+ "file_name": "18242.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 820
+ },
+ {
+ "file_name": "17015.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 821
+ },
+ {
+ "file_name": "17173.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 822
+ },
+ {
+ "file_name": "14526.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 823
+ },
+ {
+ "file_name": "12321.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 824
+ },
+ {
+ "file_name": "13272.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 825
+ },
+ {
+ "file_name": "12211.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 826
+ },
+ {
+ "file_name": "18216.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 827
+ },
+ {
+ "file_name": "17380.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 828
+ },
+ {
+ "file_name": "17591.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 829
+ },
+ {
+ "file_name": "17072.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 830
+ },
+ {
+ "file_name": "22065.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 831
+ },
+ {
+ "file_name": "22736.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 832
+ },
+ {
+ "file_name": "15763.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 833
+ },
+ {
+ "file_name": "13869.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 834
+ },
+ {
+ "file_name": "21041.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 835
+ },
+ {
+ "file_name": "18893.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 836
+ },
+ {
+ "file_name": "22542.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 837
+ },
+ {
+ "file_name": "15650.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 838
+ },
+ {
+ "file_name": "21166.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 839
+ },
+ {
+ "file_name": "14612.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 840
+ },
+ {
+ "file_name": "18431.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 841
+ },
+ {
+ "file_name": "20192.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 842
+ },
+ {
+ "file_name": "18809.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 843
+ },
+ {
+ "file_name": "17139.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 844
+ },
+ {
+ "file_name": "12093.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 845
+ },
+ {
+ "file_name": "15955.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 846
+ },
+ {
+ "file_name": "16830.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 847
+ },
+ {
+ "file_name": "21103.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 848
+ },
+ {
+ "file_name": "20064.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 849
+ },
+ {
+ "file_name": "17469.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 850
+ },
+ {
+ "file_name": "15879.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 851
+ },
+ {
+ "file_name": "22513.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 852
+ },
+ {
+ "file_name": "16678.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 853
+ },
+ {
+ "file_name": "18287.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 854
+ },
+ {
+ "file_name": "13644.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 855
+ },
+ {
+ "file_name": "18405.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 856
+ },
+ {
+ "file_name": "22414.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 857
+ },
+ {
+ "file_name": "18477.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 858
+ },
+ {
+ "file_name": "16589.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 859
+ },
+ {
+ "file_name": "19245.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 860
+ },
+ {
+ "file_name": "20894.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 861
+ },
+ {
+ "file_name": "15131.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 862
+ },
+ {
+ "file_name": "22367.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 863
+ },
+ {
+ "file_name": "15157.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 864
+ },
+ {
+ "file_name": "18806.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 865
+ },
+ {
+ "file_name": "22995.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 866
+ },
+ {
+ "file_name": "18285.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 867
+ },
+ {
+ "file_name": "16931.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 868
+ },
+ {
+ "file_name": "23426.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 869
+ },
+ {
+ "file_name": "15098.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 870
+ },
+ {
+ "file_name": "20778.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 871
+ },
+ {
+ "file_name": "15745.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 872
+ },
+ {
+ "file_name": "13848.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 873
+ },
+ {
+ "file_name": "15335.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 874
+ },
+ {
+ "file_name": "22832.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 875
+ },
+ {
+ "file_name": "20196.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 876
+ },
+ {
+ "file_name": "13992.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 877
+ },
+ {
+ "file_name": "14461.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 878
+ },
+ {
+ "file_name": "12533.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 879
+ },
+ {
+ "file_name": "11984.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 880
+ },
+ {
+ "file_name": "21152.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 881
+ },
+ {
+ "file_name": "21053.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 882
+ },
+ {
+ "file_name": "19673.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 883
+ },
+ {
+ "file_name": "12882.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 884
+ },
+ {
+ "file_name": "15196.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 885
+ },
+ {
+ "file_name": "19795.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 886
+ },
+ {
+ "file_name": "15609.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 887
+ },
+ {
+ "file_name": "18035.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 888
+ },
+ {
+ "file_name": "13924.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 889
+ },
+ {
+ "file_name": "22527.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 890
+ },
+ {
+ "file_name": "20930.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 891
+ },
+ {
+ "file_name": "22851.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 892
+ },
+ {
+ "file_name": "22607.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 893
+ },
+ {
+ "file_name": "20087.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 894
+ },
+ {
+ "file_name": "14358.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 895
+ },
+ {
+ "file_name": "14629.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 896
+ },
+ {
+ "file_name": "19299.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 897
+ },
+ {
+ "file_name": "16289.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 898
+ },
+ {
+ "file_name": "18207.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 899
+ },
+ {
+ "file_name": "18160.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 900
+ },
+ {
+ "file_name": "17219.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 901
+ },
+ {
+ "file_name": "12668.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 902
+ },
+ {
+ "file_name": "19793.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 903
+ },
+ {
+ "file_name": "15559.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 904
+ },
+ {
+ "file_name": "20594.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 905
+ },
+ {
+ "file_name": "18670.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 906
+ },
+ {
+ "file_name": "21470.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 907
+ },
+ {
+ "file_name": "22621.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 908
+ },
+ {
+ "file_name": "23435.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 909
+ },
+ {
+ "file_name": "23116.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 910
+ },
+ {
+ "file_name": "13100.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 911
+ },
+ {
+ "file_name": "19333.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 912
+ },
+ {
+ "file_name": "14081.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 913
+ },
+ {
+ "file_name": "22369.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 914
+ },
+ {
+ "file_name": "16160.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 915
+ },
+ {
+ "file_name": "14932.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 916
+ },
+ {
+ "file_name": "19962.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 917
+ },
+ {
+ "file_name": "19233.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 918
+ },
+ {
+ "file_name": "16859.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 919
+ },
+ {
+ "file_name": "17478.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 920
+ },
+ {
+ "file_name": "15317.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 921
+ },
+ {
+ "file_name": "14377.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 922
+ },
+ {
+ "file_name": "17000.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 923
+ },
+ {
+ "file_name": "19830.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 924
+ },
+ {
+ "file_name": "15673.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 925
+ },
+ {
+ "file_name": "18659.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 926
+ },
+ {
+ "file_name": "16960.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 927
+ },
+ {
+ "file_name": "18687.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 928
+ },
+ {
+ "file_name": "16751.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 929
+ },
+ {
+ "file_name": "18876.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 930
+ },
+ {
+ "file_name": "20270.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 931
+ },
+ {
+ "file_name": "13625.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 932
+ },
+ {
+ "file_name": "12716.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 933
+ },
+ {
+ "file_name": "19952.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 934
+ },
+ {
+ "file_name": "13512.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 935
+ },
+ {
+ "file_name": "16935.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 936
+ },
+ {
+ "file_name": "19308.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 937
+ },
+ {
+ "file_name": "12909.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 938
+ },
+ {
+ "file_name": "12802.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 939
+ },
+ {
+ "file_name": "15622.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 940
+ },
+ {
+ "file_name": "13657.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 941
+ },
+ {
+ "file_name": "16016.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 942
+ },
+ {
+ "file_name": "14258.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 943
+ },
+ {
+ "file_name": "19311.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 944
+ },
+ {
+ "file_name": "15283.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 945
+ },
+ {
+ "file_name": "18162.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 946
+ },
+ {
+ "file_name": "13955.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 947
+ },
+ {
+ "file_name": "21871.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 948
+ },
+ {
+ "file_name": "21038.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 949
+ },
+ {
+ "file_name": "20923.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 950
+ },
+ {
+ "file_name": "14941.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 951
+ },
+ {
+ "file_name": "12605.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 952
+ },
+ {
+ "file_name": "19070.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 953
+ },
+ {
+ "file_name": "17925.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 954
+ },
+ {
+ "file_name": "22678.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 955
+ },
+ {
+ "file_name": "12650.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 956
+ },
+ {
+ "file_name": "22854.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 957
+ },
+ {
+ "file_name": "14122.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 958
+ },
+ {
+ "file_name": "20733.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 959
+ },
+ {
+ "file_name": "17098.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 960
+ },
+ {
+ "file_name": "15166.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 961
+ },
+ {
+ "file_name": "11727.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 962
+ },
+ {
+ "file_name": "19200.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 963
+ },
+ {
+ "file_name": "16328.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 964
+ },
+ {
+ "file_name": "14462.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 965
+ },
+ {
+ "file_name": "19731.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 966
+ },
+ {
+ "file_name": "20636.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 967
+ },
+ {
+ "file_name": "22680.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 968
+ },
+ {
+ "file_name": "20724.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 969
+ },
+ {
+ "file_name": "19014.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 970
+ },
+ {
+ "file_name": "18450.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 971
+ },
+ {
+ "file_name": "19082.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 972
+ },
+ {
+ "file_name": "13309.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 973
+ },
+ {
+ "file_name": "16530.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 974
+ },
+ {
+ "file_name": "15952.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 975
+ },
+ {
+ "file_name": "23079.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 976
+ },
+ {
+ "file_name": "18373.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 977
+ },
+ {
+ "file_name": "15085.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 978
+ },
+ {
+ "file_name": "22788.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 979
+ },
+ {
+ "file_name": "19419.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 980
+ },
+ {
+ "file_name": "12338.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 981
+ },
+ {
+ "file_name": "12477.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 982
+ },
+ {
+ "file_name": "16629.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 983
+ },
+ {
+ "file_name": "14668.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 984
+ },
+ {
+ "file_name": "21275.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 985
+ },
+ {
+ "file_name": "13934.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 986
+ },
+ {
+ "file_name": "22578.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 987
+ },
+ {
+ "file_name": "14435.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 988
+ },
+ {
+ "file_name": "20566.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 989
+ },
+ {
+ "file_name": "21597.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 990
+ },
+ {
+ "file_name": "12907.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 991
+ },
+ {
+ "file_name": "16197.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 992
+ },
+ {
+ "file_name": "23333.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 993
+ },
+ {
+ "file_name": "19178.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 994
+ },
+ {
+ "file_name": "17152.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 995
+ },
+ {
+ "file_name": "18016.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 996
+ },
+ {
+ "file_name": "21566.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 997
+ },
+ {
+ "file_name": "20149.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 998
+ },
+ {
+ "file_name": "18510.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 999
+ },
+ {
+ "file_name": "20929.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1000
+ },
+ {
+ "file_name": "13789.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1001
+ },
+ {
+ "file_name": "17018.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1002
+ },
+ {
+ "file_name": "15737.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1003
+ },
+ {
+ "file_name": "20813.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1004
+ },
+ {
+ "file_name": "11887.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1005
+ },
+ {
+ "file_name": "14473.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1006
+ },
+ {
+ "file_name": "16609.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1007
+ },
+ {
+ "file_name": "19724.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1008
+ },
+ {
+ "file_name": "18487.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1009
+ },
+ {
+ "file_name": "16974.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1010
+ },
+ {
+ "file_name": "16019.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1011
+ },
+ {
+ "file_name": "19807.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1012
+ },
+ {
+ "file_name": "12326.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1013
+ },
+ {
+ "file_name": "21886.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1014
+ },
+ {
+ "file_name": "16015.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1015
+ },
+ {
+ "file_name": "15154.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1016
+ },
+ {
+ "file_name": "21695.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1017
+ },
+ {
+ "file_name": "14142.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1018
+ },
+ {
+ "file_name": "14095.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1019
+ },
+ {
+ "file_name": "20019.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1020
+ },
+ {
+ "file_name": "15425.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1021
+ },
+ {
+ "file_name": "21602.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1022
+ },
+ {
+ "file_name": "13380.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1023
+ },
+ {
+ "file_name": "19223.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1024
+ },
+ {
+ "file_name": "12355.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1025
+ },
+ {
+ "file_name": "12457.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1026
+ },
+ {
+ "file_name": "21515.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1027
+ },
+ {
+ "file_name": "16062.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1028
+ },
+ {
+ "file_name": "14300.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1029
+ },
+ {
+ "file_name": "16069.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1030
+ },
+ {
+ "file_name": "19057.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1031
+ },
+ {
+ "file_name": "13138.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1032
+ },
+ {
+ "file_name": "18913.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1033
+ },
+ {
+ "file_name": "12456.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1034
+ },
+ {
+ "file_name": "22456.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1035
+ },
+ {
+ "file_name": "19500.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1036
+ },
+ {
+ "file_name": "17682.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1037
+ },
+ {
+ "file_name": "19933.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1038
+ },
+ {
+ "file_name": "18554.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1039
+ },
+ {
+ "file_name": "16146.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1040
+ },
+ {
+ "file_name": "18725.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1041
+ },
+ {
+ "file_name": "11929.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1042
+ },
+ {
+ "file_name": "12188.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1043
+ },
+ {
+ "file_name": "13818.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1044
+ },
+ {
+ "file_name": "16552.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1045
+ },
+ {
+ "file_name": "14979.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1046
+ },
+ {
+ "file_name": "14837.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1047
+ },
+ {
+ "file_name": "20349.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1048
+ },
+ {
+ "file_name": "18854.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1049
+ },
+ {
+ "file_name": "16074.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1050
+ },
+ {
+ "file_name": "20058.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1051
+ },
+ {
+ "file_name": "21825.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1052
+ },
+ {
+ "file_name": "21650.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1053
+ },
+ {
+ "file_name": "12578.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1054
+ },
+ {
+ "file_name": "17130.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1055
+ },
+ {
+ "file_name": "22943.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1056
+ },
+ {
+ "file_name": "12963.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1057
+ },
+ {
+ "file_name": "18987.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1058
+ },
+ {
+ "file_name": "16164.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1059
+ },
+ {
+ "file_name": "20954.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1060
+ },
+ {
+ "file_name": "15937.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1061
+ },
+ {
+ "file_name": "21306.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1062
+ },
+ {
+ "file_name": "12210.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1063
+ },
+ {
+ "file_name": "16191.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1064
+ },
+ {
+ "file_name": "14299.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1065
+ },
+ {
+ "file_name": "23084.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1066
+ },
+ {
+ "file_name": "21526.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1067
+ },
+ {
+ "file_name": "20210.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1068
+ },
+ {
+ "file_name": "19690.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1069
+ },
+ {
+ "file_name": "16383.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1070
+ },
+ {
+ "file_name": "18541.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1071
+ },
+ {
+ "file_name": "17406.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1072
+ },
+ {
+ "file_name": "13291.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1073
+ },
+ {
+ "file_name": "19064.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1074
+ },
+ {
+ "file_name": "17859.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1075
+ },
+ {
+ "file_name": "15966.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1076
+ },
+ {
+ "file_name": "13165.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1077
+ },
+ {
+ "file_name": "22700.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1078
+ },
+ {
+ "file_name": "17924.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1079
+ },
+ {
+ "file_name": "19254.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1080
+ },
+ {
+ "file_name": "22913.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1081
+ },
+ {
+ "file_name": "19161.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1082
+ },
+ {
+ "file_name": "13978.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1083
+ },
+ {
+ "file_name": "21094.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1084
+ },
+ {
+ "file_name": "16391.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1085
+ },
+ {
+ "file_name": "11791.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1086
+ },
+ {
+ "file_name": "17517.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1087
+ },
+ {
+ "file_name": "20000.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1088
+ },
+ {
+ "file_name": "20622.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1089
+ },
+ {
+ "file_name": "21919.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1090
+ },
+ {
+ "file_name": "17053.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1091
+ },
+ {
+ "file_name": "13215.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1092
+ },
+ {
+ "file_name": "15106.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1093
+ },
+ {
+ "file_name": "14285.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1094
+ },
+ {
+ "file_name": "13576.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1095
+ },
+ {
+ "file_name": "14546.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1096
+ },
+ {
+ "file_name": "21529.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1097
+ },
+ {
+ "file_name": "20965.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1098
+ },
+ {
+ "file_name": "17978.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1099
+ },
+ {
+ "file_name": "22743.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1100
+ },
+ {
+ "file_name": "16117.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1101
+ },
+ {
+ "file_name": "12050.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1102
+ },
+ {
+ "file_name": "23003.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1103
+ },
+ {
+ "file_name": "14323.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1104
+ },
+ {
+ "file_name": "12580.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1105
+ },
+ {
+ "file_name": "19376.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1106
+ },
+ {
+ "file_name": "11972.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1107
+ },
+ {
+ "file_name": "18363.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1108
+ },
+ {
+ "file_name": "16477.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1109
+ },
+ {
+ "file_name": "21253.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1110
+ },
+ {
+ "file_name": "20554.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1111
+ },
+ {
+ "file_name": "22950.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1112
+ },
+ {
+ "file_name": "17534.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1113
+ },
+ {
+ "file_name": "14775.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1114
+ },
+ {
+ "file_name": "12673.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1115
+ },
+ {
+ "file_name": "20158.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1116
+ },
+ {
+ "file_name": "17289.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1117
+ },
+ {
+ "file_name": "21917.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1118
+ },
+ {
+ "file_name": "22309.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1119
+ },
+ {
+ "file_name": "13287.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1120
+ },
+ {
+ "file_name": "19996.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1121
+ },
+ {
+ "file_name": "16363.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1122
+ },
+ {
+ "file_name": "17515.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1123
+ },
+ {
+ "file_name": "21367.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1124
+ },
+ {
+ "file_name": "21514.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1125
+ },
+ {
+ "file_name": "17940.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1126
+ },
+ {
+ "file_name": "21474.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1127
+ },
+ {
+ "file_name": "13217.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1128
+ },
+ {
+ "file_name": "13719.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1129
+ },
+ {
+ "file_name": "16414.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1130
+ },
+ {
+ "file_name": "13544.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1131
+ },
+ {
+ "file_name": "11758.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1132
+ },
+ {
+ "file_name": "16445.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1133
+ },
+ {
+ "file_name": "12443.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1134
+ },
+ {
+ "file_name": "15566.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1135
+ },
+ {
+ "file_name": "18744.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1136
+ },
+ {
+ "file_name": "14698.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1137
+ },
+ {
+ "file_name": "16061.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1138
+ },
+ {
+ "file_name": "14037.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1139
+ },
+ {
+ "file_name": "14056.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1140
+ },
+ {
+ "file_name": "18254.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1141
+ },
+ {
+ "file_name": "12292.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1142
+ },
+ {
+ "file_name": "17388.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1143
+ },
+ {
+ "file_name": "21932.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1144
+ },
+ {
+ "file_name": "17723.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1145
+ },
+ {
+ "file_name": "15414.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1146
+ },
+ {
+ "file_name": "12918.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1147
+ },
+ {
+ "file_name": "11759.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1148
+ },
+ {
+ "file_name": "22786.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1149
+ },
+ {
+ "file_name": "20161.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1150
+ },
+ {
+ "file_name": "15696.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1151
+ },
+ {
+ "file_name": "13935.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1152
+ },
+ {
+ "file_name": "21174.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1153
+ },
+ {
+ "file_name": "16547.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1154
+ },
+ {
+ "file_name": "20961.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1155
+ },
+ {
+ "file_name": "18495.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1156
+ },
+ {
+ "file_name": "21386.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1157
+ },
+ {
+ "file_name": "16601.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1158
+ },
+ {
+ "file_name": "22525.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1159
+ },
+ {
+ "file_name": "14328.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1160
+ },
+ {
+ "file_name": "23126.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1161
+ },
+ {
+ "file_name": "18783.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1162
+ },
+ {
+ "file_name": "20396.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1163
+ },
+ {
+ "file_name": "17831.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1164
+ },
+ {
+ "file_name": "14762.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1165
+ },
+ {
+ "file_name": "13788.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1166
+ },
+ {
+ "file_name": "14642.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1167
+ },
+ {
+ "file_name": "12780.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1168
+ },
+ {
+ "file_name": "20015.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1169
+ },
+ {
+ "file_name": "23266.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1170
+ },
+ {
+ "file_name": "21836.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1171
+ },
+ {
+ "file_name": "22024.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1172
+ },
+ {
+ "file_name": "13508.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1173
+ },
+ {
+ "file_name": "15318.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1174
+ },
+ {
+ "file_name": "18740.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1175
+ },
+ {
+ "file_name": "15221.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1176
+ },
+ {
+ "file_name": "18623.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1177
+ },
+ {
+ "file_name": "13286.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1178
+ },
+ {
+ "file_name": "12653.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1179
+ },
+ {
+ "file_name": "21879.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1180
+ },
+ {
+ "file_name": "18804.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1181
+ },
+ {
+ "file_name": "18237.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1182
+ },
+ {
+ "file_name": "17089.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1183
+ },
+ {
+ "file_name": "13368.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1184
+ },
+ {
+ "file_name": "12811.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1185
+ },
+ {
+ "file_name": "13864.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1186
+ },
+ {
+ "file_name": "19029.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1187
+ },
+ {
+ "file_name": "21518.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1188
+ },
+ {
+ "file_name": "22668.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1189
+ },
+ {
+ "file_name": "14697.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1190
+ },
+ {
+ "file_name": "21079.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1191
+ },
+ {
+ "file_name": "13342.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1192
+ },
+ {
+ "file_name": "14167.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1193
+ },
+ {
+ "file_name": "17851.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1194
+ },
+ {
+ "file_name": "17084.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1195
+ },
+ {
+ "file_name": "21318.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1196
+ },
+ {
+ "file_name": "20847.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1197
+ },
+ {
+ "file_name": "13455.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1198
+ },
+ {
+ "file_name": "17155.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1199
+ },
+ {
+ "file_name": "17826.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1200
+ },
+ {
+ "file_name": "15142.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1201
+ },
+ {
+ "file_name": "15475.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1202
+ },
+ {
+ "file_name": "13334.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1203
+ },
+ {
+ "file_name": "20978.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1204
+ },
+ {
+ "file_name": "18833.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1205
+ },
+ {
+ "file_name": "22664.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1206
+ },
+ {
+ "file_name": "22665.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1207
+ },
+ {
+ "file_name": "16988.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1208
+ },
+ {
+ "file_name": "22775.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1209
+ },
+ {
+ "file_name": "15405.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1210
+ },
+ {
+ "file_name": "13600.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1211
+ },
+ {
+ "file_name": "18969.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1212
+ },
+ {
+ "file_name": "15498.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1213
+ },
+ {
+ "file_name": "13251.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1214
+ },
+ {
+ "file_name": "13257.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1215
+ },
+ {
+ "file_name": "19506.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1216
+ },
+ {
+ "file_name": "20667.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1217
+ },
+ {
+ "file_name": "14897.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1218
+ },
+ {
+ "file_name": "15555.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1219
+ },
+ {
+ "file_name": "21701.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1220
+ },
+ {
+ "file_name": "12198.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1221
+ },
+ {
+ "file_name": "13825.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1222
+ },
+ {
+ "file_name": "21016.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1223
+ },
+ {
+ "file_name": "11836.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1224
+ },
+ {
+ "file_name": "14015.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1225
+ },
+ {
+ "file_name": "15234.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1226
+ },
+ {
+ "file_name": "12576.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1227
+ },
+ {
+ "file_name": "20538.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1228
+ },
+ {
+ "file_name": "18000.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1229
+ },
+ {
+ "file_name": "21044.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1230
+ },
+ {
+ "file_name": "20755.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1231
+ },
+ {
+ "file_name": "22086.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1232
+ },
+ {
+ "file_name": "22649.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1233
+ },
+ {
+ "file_name": "20944.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1234
+ },
+ {
+ "file_name": "19851.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1235
+ },
+ {
+ "file_name": "15259.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1236
+ },
+ {
+ "file_name": "17698.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1237
+ },
+ {
+ "file_name": "20510.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1238
+ },
+ {
+ "file_name": "22594.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1239
+ },
+ {
+ "file_name": "19685.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1240
+ },
+ {
+ "file_name": "12935.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1241
+ },
+ {
+ "file_name": "11735.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1242
+ },
+ {
+ "file_name": "13405.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1243
+ },
+ {
+ "file_name": "20286.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1244
+ },
+ {
+ "file_name": "19105.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1245
+ },
+ {
+ "file_name": "20818.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1246
+ },
+ {
+ "file_name": "14373.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1247
+ },
+ {
+ "file_name": "21241.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1248
+ },
+ {
+ "file_name": "14890.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1249
+ },
+ {
+ "file_name": "14789.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1250
+ },
+ {
+ "file_name": "13140.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1251
+ },
+ {
+ "file_name": "20863.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1252
+ },
+ {
+ "file_name": "21578.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1253
+ },
+ {
+ "file_name": "16720.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1254
+ },
+ {
+ "file_name": "21104.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1255
+ },
+ {
+ "file_name": "19432.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1256
+ },
+ {
+ "file_name": "12102.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1257
+ },
+ {
+ "file_name": "20504.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1258
+ },
+ {
+ "file_name": "16096.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1259
+ },
+ {
+ "file_name": "14848.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1260
+ },
+ {
+ "file_name": "21237.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1261
+ },
+ {
+ "file_name": "15313.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1262
+ },
+ {
+ "file_name": "14582.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1263
+ },
+ {
+ "file_name": "21199.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1264
+ },
+ {
+ "file_name": "14016.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1265
+ },
+ {
+ "file_name": "13960.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1266
+ },
+ {
+ "file_name": "16741.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1267
+ },
+ {
+ "file_name": "14563.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1268
+ },
+ {
+ "file_name": "16256.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1269
+ },
+ {
+ "file_name": "20095.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1270
+ },
+ {
+ "file_name": "13001.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1271
+ },
+ {
+ "file_name": "12604.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1272
+ },
+ {
+ "file_name": "17574.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1273
+ },
+ {
+ "file_name": "15898.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1274
+ },
+ {
+ "file_name": "14628.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1275
+ },
+ {
+ "file_name": "18249.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1276
+ },
+ {
+ "file_name": "20067.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1277
+ },
+ {
+ "file_name": "17621.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1278
+ },
+ {
+ "file_name": "22423.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1279
+ },
+ {
+ "file_name": "17982.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1280
+ },
+ {
+ "file_name": "15682.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1281
+ },
+ {
+ "file_name": "18972.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1282
+ },
+ {
+ "file_name": "13071.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1283
+ },
+ {
+ "file_name": "22637.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1284
+ },
+ {
+ "file_name": "13021.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1285
+ },
+ {
+ "file_name": "13530.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1286
+ },
+ {
+ "file_name": "17542.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1287
+ },
+ {
+ "file_name": "22284.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1288
+ },
+ {
+ "file_name": "19895.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1289
+ },
+ {
+ "file_name": "18742.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1290
+ },
+ {
+ "file_name": "17028.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1291
+ },
+ {
+ "file_name": "21741.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1292
+ },
+ {
+ "file_name": "22084.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1293
+ },
+ {
+ "file_name": "18449.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1294
+ },
+ {
+ "file_name": "20294.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1295
+ },
+ {
+ "file_name": "12631.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1296
+ },
+ {
+ "file_name": "14587.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1297
+ },
+ {
+ "file_name": "18418.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1298
+ },
+ {
+ "file_name": "17803.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1299
+ },
+ {
+ "file_name": "16278.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1300
+ },
+ {
+ "file_name": "23425.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1301
+ },
+ {
+ "file_name": "17767.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1302
+ },
+ {
+ "file_name": "23177.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1303
+ },
+ {
+ "file_name": "19888.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1304
+ },
+ {
+ "file_name": "16712.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1305
+ },
+ {
+ "file_name": "22295.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1306
+ },
+ {
+ "file_name": "19130.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1307
+ },
+ {
+ "file_name": "11960.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1308
+ },
+ {
+ "file_name": "20591.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1309
+ },
+ {
+ "file_name": "15860.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1310
+ },
+ {
+ "file_name": "15396.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1311
+ },
+ {
+ "file_name": "16241.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1312
+ },
+ {
+ "file_name": "18527.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1313
+ },
+ {
+ "file_name": "14988.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1314
+ },
+ {
+ "file_name": "15428.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1315
+ },
+ {
+ "file_name": "22852.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1316
+ },
+ {
+ "file_name": "20445.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1317
+ },
+ {
+ "file_name": "15252.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1318
+ },
+ {
+ "file_name": "11812.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1319
+ },
+ {
+ "file_name": "19598.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1320
+ },
+ {
+ "file_name": "17228.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1321
+ },
+ {
+ "file_name": "15602.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1322
+ },
+ {
+ "file_name": "18071.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1323
+ },
+ {
+ "file_name": "20518.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1324
+ },
+ {
+ "file_name": "18986.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1325
+ },
+ {
+ "file_name": "14869.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1326
+ },
+ {
+ "file_name": "18410.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1327
+ },
+ {
+ "file_name": "17548.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1328
+ },
+ {
+ "file_name": "22149.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1329
+ },
+ {
+ "file_name": "21690.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1330
+ },
+ {
+ "file_name": "13180.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1331
+ },
+ {
+ "file_name": "23390.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1332
+ },
+ {
+ "file_name": "22965.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1333
+ },
+ {
+ "file_name": "21845.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1334
+ },
+ {
+ "file_name": "23440.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1335
+ },
+ {
+ "file_name": "14904.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1336
+ },
+ {
+ "file_name": "23198.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1337
+ },
+ {
+ "file_name": "16538.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1338
+ },
+ {
+ "file_name": "20056.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1339
+ },
+ {
+ "file_name": "19179.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1340
+ },
+ {
+ "file_name": "17381.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1341
+ },
+ {
+ "file_name": "12289.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1342
+ },
+ {
+ "file_name": "22126.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1343
+ },
+ {
+ "file_name": "13812.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1344
+ },
+ {
+ "file_name": "20143.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1345
+ },
+ {
+ "file_name": "14744.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1346
+ },
+ {
+ "file_name": "17122.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1347
+ },
+ {
+ "file_name": "17222.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1348
+ },
+ {
+ "file_name": "21884.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1349
+ },
+ {
+ "file_name": "13297.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1350
+ },
+ {
+ "file_name": "21278.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1351
+ },
+ {
+ "file_name": "11926.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1352
+ },
+ {
+ "file_name": "19109.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1353
+ },
+ {
+ "file_name": "15502.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1354
+ },
+ {
+ "file_name": "21082.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1355
+ },
+ {
+ "file_name": "20072.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1356
+ },
+ {
+ "file_name": "20664.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1357
+ },
+ {
+ "file_name": "15039.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1358
+ },
+ {
+ "file_name": "14810.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1359
+ },
+ {
+ "file_name": "16025.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1360
+ },
+ {
+ "file_name": "12985.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1361
+ },
+ {
+ "file_name": "19435.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1362
+ },
+ {
+ "file_name": "21122.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1363
+ },
+ {
+ "file_name": "15916.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1364
+ },
+ {
+ "file_name": "19768.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1365
+ },
+ {
+ "file_name": "13269.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1366
+ },
+ {
+ "file_name": "16092.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1367
+ },
+ {
+ "file_name": "17229.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1368
+ },
+ {
+ "file_name": "20088.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1369
+ },
+ {
+ "file_name": "21337.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1370
+ },
+ {
+ "file_name": "15063.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1371
+ },
+ {
+ "file_name": "22194.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1372
+ },
+ {
+ "file_name": "22819.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1373
+ },
+ {
+ "file_name": "18283.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1374
+ },
+ {
+ "file_name": "17493.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1375
+ },
+ {
+ "file_name": "16753.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1376
+ },
+ {
+ "file_name": "13547.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1377
+ },
+ {
+ "file_name": "12262.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1378
+ },
+ {
+ "file_name": "14021.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1379
+ },
+ {
+ "file_name": "15909.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1380
+ },
+ {
+ "file_name": "17097.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1381
+ },
+ {
+ "file_name": "14460.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1382
+ },
+ {
+ "file_name": "19868.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1383
+ },
+ {
+ "file_name": "13528.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1384
+ },
+ {
+ "file_name": "18660.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1385
+ },
+ {
+ "file_name": "13631.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1386
+ },
+ {
+ "file_name": "14472.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1387
+ },
+ {
+ "file_name": "13821.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1388
+ },
+ {
+ "file_name": "22334.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1389
+ },
+ {
+ "file_name": "21450.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1390
+ },
+ {
+ "file_name": "16612.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1391
+ },
+ {
+ "file_name": "19069.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1392
+ },
+ {
+ "file_name": "19277.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1393
+ },
+ {
+ "file_name": "13774.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1394
+ },
+ {
+ "file_name": "14843.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1395
+ },
+ {
+ "file_name": "14089.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1396
+ },
+ {
+ "file_name": "22589.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1397
+ },
+ {
+ "file_name": "22938.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1398
+ },
+ {
+ "file_name": "16735.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1399
+ },
+ {
+ "file_name": "19098.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1400
+ },
+ {
+ "file_name": "17459.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1401
+ },
+ {
+ "file_name": "17362.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1402
+ },
+ {
+ "file_name": "15733.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1403
+ },
+ {
+ "file_name": "15028.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1404
+ },
+ {
+ "file_name": "22355.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1405
+ },
+ {
+ "file_name": "12790.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1406
+ },
+ {
+ "file_name": "21664.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1407
+ },
+ {
+ "file_name": "14266.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1408
+ },
+ {
+ "file_name": "18675.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1409
+ },
+ {
+ "file_name": "16813.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1410
+ },
+ {
+ "file_name": "14079.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1411
+ },
+ {
+ "file_name": "13336.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1412
+ },
+ {
+ "file_name": "21426.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1413
+ },
+ {
+ "file_name": "19892.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1414
+ },
+ {
+ "file_name": "19942.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1415
+ },
+ {
+ "file_name": "11920.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1416
+ },
+ {
+ "file_name": "17724.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1417
+ },
+ {
+ "file_name": "21914.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1418
+ },
+ {
+ "file_name": "13919.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1419
+ },
+ {
+ "file_name": "21952.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1420
+ },
+ {
+ "file_name": "15809.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1421
+ },
+ {
+ "file_name": "20219.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1422
+ },
+ {
+ "file_name": "19844.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1423
+ },
+ {
+ "file_name": "20689.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1424
+ },
+ {
+ "file_name": "22063.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1425
+ },
+ {
+ "file_name": "21176.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1426
+ },
+ {
+ "file_name": "13506.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1427
+ },
+ {
+ "file_name": "23395.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1428
+ },
+ {
+ "file_name": "17719.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1429
+ },
+ {
+ "file_name": "19298.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1430
+ },
+ {
+ "file_name": "21368.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1431
+ },
+ {
+ "file_name": "12865.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1432
+ },
+ {
+ "file_name": "18801.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1433
+ },
+ {
+ "file_name": "23328.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1434
+ },
+ {
+ "file_name": "17850.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1435
+ },
+ {
+ "file_name": "20070.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1436
+ },
+ {
+ "file_name": "21752.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1437
+ },
+ {
+ "file_name": "18704.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1438
+ },
+ {
+ "file_name": "22857.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1439
+ },
+ {
+ "file_name": "15944.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1440
+ },
+ {
+ "file_name": "17304.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1441
+ },
+ {
+ "file_name": "22168.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1442
+ },
+ {
+ "file_name": "15192.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1443
+ },
+ {
+ "file_name": "20476.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1444
+ },
+ {
+ "file_name": "20981.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1445
+ },
+ {
+ "file_name": "23123.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1446
+ },
+ {
+ "file_name": "15614.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1447
+ },
+ {
+ "file_name": "18377.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1448
+ },
+ {
+ "file_name": "22531.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1449
+ },
+ {
+ "file_name": "16926.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1450
+ },
+ {
+ "file_name": "13242.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1451
+ },
+ {
+ "file_name": "15931.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1452
+ },
+ {
+ "file_name": "18779.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1453
+ },
+ {
+ "file_name": "18321.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1454
+ },
+ {
+ "file_name": "19239.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1455
+ },
+ {
+ "file_name": "22951.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1456
+ },
+ {
+ "file_name": "16867.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1457
+ },
+ {
+ "file_name": "11971.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1458
+ },
+ {
+ "file_name": "19862.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1459
+ },
+ {
+ "file_name": "18959.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1460
+ },
+ {
+ "file_name": "18988.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1461
+ },
+ {
+ "file_name": "17188.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1462
+ },
+ {
+ "file_name": "12788.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1463
+ },
+ {
+ "file_name": "22144.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1464
+ },
+ {
+ "file_name": "16768.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1465
+ },
+ {
+ "file_name": "16366.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1466
+ },
+ {
+ "file_name": "19198.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1467
+ },
+ {
+ "file_name": "18589.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1468
+ },
+ {
+ "file_name": "23350.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1469
+ },
+ {
+ "file_name": "18834.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1470
+ },
+ {
+ "file_name": "21631.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1471
+ },
+ {
+ "file_name": "17403.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1472
+ },
+ {
+ "file_name": "21960.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1473
+ },
+ {
+ "file_name": "22543.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1474
+ },
+ {
+ "file_name": "12570.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1475
+ },
+ {
+ "file_name": "13389.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1476
+ },
+ {
+ "file_name": "15260.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1477
+ },
+ {
+ "file_name": "21818.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1478
+ },
+ {
+ "file_name": "23154.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1479
+ },
+ {
+ "file_name": "17626.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1480
+ },
+ {
+ "file_name": "15722.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1481
+ },
+ {
+ "file_name": "20800.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1482
+ },
+ {
+ "file_name": "12932.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1483
+ },
+ {
+ "file_name": "11957.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1484
+ },
+ {
+ "file_name": "17307.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1485
+ },
+ {
+ "file_name": "13521.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1486
+ },
+ {
+ "file_name": "14334.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1487
+ },
+ {
+ "file_name": "22707.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1488
+ },
+ {
+ "file_name": "19546.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1489
+ },
+ {
+ "file_name": "19013.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1490
+ },
+ {
+ "file_name": "18060.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1491
+ },
+ {
+ "file_name": "13296.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1492
+ },
+ {
+ "file_name": "16509.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1493
+ },
+ {
+ "file_name": "22760.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1494
+ },
+ {
+ "file_name": "15648.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1495
+ },
+ {
+ "file_name": "12286.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1496
+ },
+ {
+ "file_name": "20116.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1497
+ },
+ {
+ "file_name": "15964.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1498
+ },
+ {
+ "file_name": "13236.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1499
+ },
+ {
+ "file_name": "23124.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1500
+ },
+ {
+ "file_name": "19286.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1501
+ },
+ {
+ "file_name": "21833.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1502
+ },
+ {
+ "file_name": "17473.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1503
+ },
+ {
+ "file_name": "14866.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1504
+ },
+ {
+ "file_name": "22432.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1505
+ },
+ {
+ "file_name": "15178.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1506
+ },
+ {
+ "file_name": "17342.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1507
+ },
+ {
+ "file_name": "16583.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1508
+ },
+ {
+ "file_name": "19532.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1509
+ },
+ {
+ "file_name": "15119.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1510
+ },
+ {
+ "file_name": "15272.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1511
+ },
+ {
+ "file_name": "20759.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1512
+ },
+ {
+ "file_name": "18116.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1513
+ },
+ {
+ "file_name": "18384.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1514
+ },
+ {
+ "file_name": "14411.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1515
+ },
+ {
+ "file_name": "21549.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1516
+ },
+ {
+ "file_name": "13917.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1517
+ },
+ {
+ "file_name": "22874.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1518
+ },
+ {
+ "file_name": "13348.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1519
+ },
+ {
+ "file_name": "21637.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1520
+ },
+ {
+ "file_name": "15784.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1521
+ },
+ {
+ "file_name": "17267.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1522
+ },
+ {
+ "file_name": "19455.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1523
+ },
+ {
+ "file_name": "20251.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1524
+ },
+ {
+ "file_name": "20101.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1525
+ },
+ {
+ "file_name": "22223.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1526
+ },
+ {
+ "file_name": "17953.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1527
+ },
+ {
+ "file_name": "22348.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1528
+ },
+ {
+ "file_name": "19438.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1529
+ },
+ {
+ "file_name": "12912.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1530
+ },
+ {
+ "file_name": "19459.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1531
+ },
+ {
+ "file_name": "15954.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1532
+ },
+ {
+ "file_name": "23157.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1533
+ },
+ {
+ "file_name": "16782.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1534
+ },
+ {
+ "file_name": "17819.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1535
+ },
+ {
+ "file_name": "17374.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1536
+ },
+ {
+ "file_name": "18115.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1537
+ },
+ {
+ "file_name": "23023.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1538
+ },
+ {
+ "file_name": "20974.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1539
+ },
+ {
+ "file_name": "18134.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1540
+ },
+ {
+ "file_name": "13564.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1541
+ },
+ {
+ "file_name": "16154.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1542
+ },
+ {
+ "file_name": "21574.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1543
+ },
+ {
+ "file_name": "17605.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1544
+ },
+ {
+ "file_name": "20080.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1545
+ },
+ {
+ "file_name": "13824.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1546
+ },
+ {
+ "file_name": "12415.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1547
+ },
+ {
+ "file_name": "19054.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1548
+ },
+ {
+ "file_name": "14374.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1549
+ },
+ {
+ "file_name": "22399.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1550
+ },
+ {
+ "file_name": "17445.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1551
+ },
+ {
+ "file_name": "18523.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1552
+ },
+ {
+ "file_name": "21113.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1553
+ },
+ {
+ "file_name": "19092.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1554
+ },
+ {
+ "file_name": "20785.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1555
+ },
+ {
+ "file_name": "14399.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1556
+ },
+ {
+ "file_name": "13692.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1557
+ },
+ {
+ "file_name": "19353.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1558
+ },
+ {
+ "file_name": "22473.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1559
+ },
+ {
+ "file_name": "22052.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1560
+ },
+ {
+ "file_name": "17145.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1561
+ },
+ {
+ "file_name": "14751.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1562
+ },
+ {
+ "file_name": "16449.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1563
+ },
+ {
+ "file_name": "15703.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1564
+ },
+ {
+ "file_name": "19716.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1565
+ },
+ {
+ "file_name": "16579.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1566
+ },
+ {
+ "file_name": "16787.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1567
+ },
+ {
+ "file_name": "22820.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1568
+ },
+ {
+ "file_name": "22791.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1569
+ },
+ {
+ "file_name": "17900.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1570
+ },
+ {
+ "file_name": "13345.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1571
+ },
+ {
+ "file_name": "22203.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1572
+ },
+ {
+ "file_name": "15359.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1573
+ },
+ {
+ "file_name": "22964.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1574
+ },
+ {
+ "file_name": "15642.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1575
+ },
+ {
+ "file_name": "16531.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1576
+ },
+ {
+ "file_name": "15213.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1577
+ },
+ {
+ "file_name": "19509.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1578
+ },
+ {
+ "file_name": "22298.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1579
+ },
+ {
+ "file_name": "22919.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1580
+ },
+ {
+ "file_name": "17641.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1581
+ },
+ {
+ "file_name": "23160.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1582
+ },
+ {
+ "file_name": "18472.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1583
+ },
+ {
+ "file_name": "20333.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1584
+ },
+ {
+ "file_name": "14253.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1585
+ },
+ {
+ "file_name": "17303.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1586
+ },
+ {
+ "file_name": "19632.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1587
+ },
+ {
+ "file_name": "16885.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1588
+ },
+ {
+ "file_name": "22658.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1589
+ },
+ {
+ "file_name": "23391.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1590
+ },
+ {
+ "file_name": "15327.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1591
+ },
+ {
+ "file_name": "18593.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1592
+ },
+ {
+ "file_name": "12983.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1593
+ },
+ {
+ "file_name": "11895.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1594
+ },
+ {
+ "file_name": "15976.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1595
+ },
+ {
+ "file_name": "12997.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1596
+ },
+ {
+ "file_name": "17202.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1597
+ },
+ {
+ "file_name": "19696.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1598
+ },
+ {
+ "file_name": "15516.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1599
+ },
+ {
+ "file_name": "20653.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1600
+ },
+ {
+ "file_name": "13125.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1601
+ },
+ {
+ "file_name": "12519.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1602
+ },
+ {
+ "file_name": "12960.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1603
+ },
+ {
+ "file_name": "18312.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1604
+ },
+ {
+ "file_name": "16098.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1605
+ },
+ {
+ "file_name": "21772.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1606
+ },
+ {
+ "file_name": "22933.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1607
+ },
+ {
+ "file_name": "12984.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1608
+ },
+ {
+ "file_name": "12194.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1609
+ },
+ {
+ "file_name": "11733.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1610
+ },
+ {
+ "file_name": "14882.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1611
+ },
+ {
+ "file_name": "17109.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1612
+ },
+ {
+ "file_name": "18300.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1613
+ },
+ {
+ "file_name": "13085.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1614
+ },
+ {
+ "file_name": "14794.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1615
+ },
+ {
+ "file_name": "13033.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1616
+ },
+ {
+ "file_name": "18094.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1617
+ },
+ {
+ "file_name": "13814.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1618
+ },
+ {
+ "file_name": "19313.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1619
+ },
+ {
+ "file_name": "18847.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1620
+ },
+ {
+ "file_name": "20414.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1621
+ },
+ {
+ "file_name": "20379.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1622
+ },
+ {
+ "file_name": "14791.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1623
+ },
+ {
+ "file_name": "13268.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1624
+ },
+ {
+ "file_name": "16690.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1625
+ },
+ {
+ "file_name": "12438.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1626
+ },
+ {
+ "file_name": "16587.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1627
+ },
+ {
+ "file_name": "19122.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1628
+ },
+ {
+ "file_name": "12784.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1629
+ },
+ {
+ "file_name": "13888.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1630
+ },
+ {
+ "file_name": "17749.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1631
+ },
+ {
+ "file_name": "17653.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1632
+ },
+ {
+ "file_name": "12762.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1633
+ },
+ {
+ "file_name": "13928.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1634
+ },
+ {
+ "file_name": "22520.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1635
+ },
+ {
+ "file_name": "13833.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1636
+ },
+ {
+ "file_name": "17986.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1637
+ },
+ {
+ "file_name": "16192.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1638
+ },
+ {
+ "file_name": "14826.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1639
+ },
+ {
+ "file_name": "13397.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1640
+ },
+ {
+ "file_name": "22197.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1641
+ },
+ {
+ "file_name": "21070.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1642
+ },
+ {
+ "file_name": "14584.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1643
+ },
+ {
+ "file_name": "12989.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1644
+ },
+ {
+ "file_name": "17507.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1645
+ },
+ {
+ "file_name": "18335.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1646
+ },
+ {
+ "file_name": "15198.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1647
+ },
+ {
+ "file_name": "11796.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1648
+ },
+ {
+ "file_name": "12250.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1649
+ },
+ {
+ "file_name": "12134.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1650
+ },
+ {
+ "file_name": "18701.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1651
+ },
+ {
+ "file_name": "22643.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1652
+ },
+ {
+ "file_name": "17151.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1653
+ },
+ {
+ "file_name": "12462.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1654
+ },
+ {
+ "file_name": "17485.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1655
+ },
+ {
+ "file_name": "18871.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1656
+ },
+ {
+ "file_name": "18576.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1657
+ },
+ {
+ "file_name": "22976.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1658
+ },
+ {
+ "file_name": "21765.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1659
+ },
+ {
+ "file_name": "23346.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1660
+ },
+ {
+ "file_name": "13979.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1661
+ },
+ {
+ "file_name": "20145.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1662
+ },
+ {
+ "file_name": "15676.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1663
+ },
+ {
+ "file_name": "12872.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1664
+ },
+ {
+ "file_name": "22918.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1665
+ },
+ {
+ "file_name": "19839.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1666
+ },
+ {
+ "file_name": "20511.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1667
+ },
+ {
+ "file_name": "20319.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1668
+ },
+ {
+ "file_name": "23335.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1669
+ },
+ {
+ "file_name": "13543.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1670
+ },
+ {
+ "file_name": "11854.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1671
+ },
+ {
+ "file_name": "20972.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1672
+ },
+ {
+ "file_name": "19591.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1673
+ },
+ {
+ "file_name": "19872.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1674
+ },
+ {
+ "file_name": "12563.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1675
+ },
+ {
+ "file_name": "17721.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1676
+ },
+ {
+ "file_name": "14723.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1677
+ },
+ {
+ "file_name": "21056.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1678
+ },
+ {
+ "file_name": "17784.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1679
+ },
+ {
+ "file_name": "12701.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1680
+ },
+ {
+ "file_name": "19733.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1681
+ },
+ {
+ "file_name": "15256.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1682
+ },
+ {
+ "file_name": "13186.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1683
+ },
+ {
+ "file_name": "14451.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1684
+ },
+ {
+ "file_name": "13111.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1685
+ },
+ {
+ "file_name": "12480.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1686
+ },
+ {
+ "file_name": "22829.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1687
+ },
+ {
+ "file_name": "19630.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1688
+ },
+ {
+ "file_name": "19785.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1689
+ },
+ {
+ "file_name": "19540.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1690
+ },
+ {
+ "file_name": "23275.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1691
+ },
+ {
+ "file_name": "22342.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1692
+ },
+ {
+ "file_name": "15155.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1693
+ },
+ {
+ "file_name": "20956.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1694
+ },
+ {
+ "file_name": "20851.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1695
+ },
+ {
+ "file_name": "13374.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1696
+ },
+ {
+ "file_name": "12710.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1697
+ },
+ {
+ "file_name": "23308.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1698
+ },
+ {
+ "file_name": "22988.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1699
+ },
+ {
+ "file_name": "23242.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1700
+ },
+ {
+ "file_name": "16142.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1701
+ },
+ {
+ "file_name": "17990.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1702
+ },
+ {
+ "file_name": "17387.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1703
+ },
+ {
+ "file_name": "13964.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1704
+ },
+ {
+ "file_name": "20259.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1705
+ },
+ {
+ "file_name": "13653.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1706
+ },
+ {
+ "file_name": "16041.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1707
+ },
+ {
+ "file_name": "17251.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1708
+ },
+ {
+ "file_name": "22201.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1709
+ },
+ {
+ "file_name": "21854.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1710
+ },
+ {
+ "file_name": "15449.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1711
+ },
+ {
+ "file_name": "16643.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1712
+ },
+ {
+ "file_name": "20846.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1713
+ },
+ {
+ "file_name": "14301.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1714
+ },
+ {
+ "file_name": "21345.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1715
+ },
+ {
+ "file_name": "12054.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1716
+ },
+ {
+ "file_name": "22924.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1717
+ },
+ {
+ "file_name": "14336.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1718
+ },
+ {
+ "file_name": "17051.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1719
+ },
+ {
+ "file_name": "16976.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1720
+ },
+ {
+ "file_name": "22256.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1721
+ },
+ {
+ "file_name": "15249.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1722
+ },
+ {
+ "file_name": "20756.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1723
+ },
+ {
+ "file_name": "13535.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1724
+ },
+ {
+ "file_name": "16942.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1725
+ },
+ {
+ "file_name": "18282.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1726
+ },
+ {
+ "file_name": "13827.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1727
+ },
+ {
+ "file_name": "16338.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1728
+ },
+ {
+ "file_name": "22847.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1729
+ },
+ {
+ "file_name": "19410.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1730
+ },
+ {
+ "file_name": "14770.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1731
+ },
+ {
+ "file_name": "14244.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1732
+ },
+ {
+ "file_name": "16779.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1733
+ },
+ {
+ "file_name": "15250.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1734
+ },
+ {
+ "file_name": "22492.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1735
+ },
+ {
+ "file_name": "18194.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1736
+ },
+ {
+ "file_name": "20089.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1737
+ },
+ {
+ "file_name": "21807.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1738
+ },
+ {
+ "file_name": "12459.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1739
+ },
+ {
+ "file_name": "16588.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1740
+ },
+ {
+ "file_name": "13760.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1741
+ },
+ {
+ "file_name": "12602.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1742
+ },
+ {
+ "file_name": "17439.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1743
+ },
+ {
+ "file_name": "12046.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1744
+ },
+ {
+ "file_name": "17636.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1745
+ },
+ {
+ "file_name": "21397.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1746
+ },
+ {
+ "file_name": "20555.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1747
+ },
+ {
+ "file_name": "23209.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1748
+ },
+ {
+ "file_name": "21357.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1749
+ },
+ {
+ "file_name": "12444.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1750
+ },
+ {
+ "file_name": "20206.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1751
+ },
+ {
+ "file_name": "13379.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1752
+ },
+ {
+ "file_name": "15458.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1753
+ },
+ {
+ "file_name": "13671.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1754
+ },
+ {
+ "file_name": "14282.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1755
+ },
+ {
+ "file_name": "12521.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1756
+ },
+ {
+ "file_name": "17759.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1757
+ },
+ {
+ "file_name": "17923.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1758
+ },
+ {
+ "file_name": "20326.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1759
+ },
+ {
+ "file_name": "13094.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1760
+ },
+ {
+ "file_name": "15808.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1761
+ },
+ {
+ "file_name": "19606.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1762
+ },
+ {
+ "file_name": "16983.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1763
+ },
+ {
+ "file_name": "22278.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1764
+ },
+ {
+ "file_name": "17583.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1765
+ },
+ {
+ "file_name": "18407.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1766
+ },
+ {
+ "file_name": "18305.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1767
+ },
+ {
+ "file_name": "21189.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1768
+ },
+ {
+ "file_name": "20762.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1769
+ },
+ {
+ "file_name": "17708.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1770
+ },
+ {
+ "file_name": "13048.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1771
+ },
+ {
+ "file_name": "22713.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1772
+ },
+ {
+ "file_name": "14324.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1773
+ },
+ {
+ "file_name": "19980.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1774
+ },
+ {
+ "file_name": "19519.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1775
+ },
+ {
+ "file_name": "19464.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1776
+ },
+ {
+ "file_name": "21956.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1777
+ },
+ {
+ "file_name": "21248.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1778
+ },
+ {
+ "file_name": "16306.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1779
+ },
+ {
+ "file_name": "17431.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1780
+ },
+ {
+ "file_name": "16485.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1781
+ },
+ {
+ "file_name": "22300.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1782
+ },
+ {
+ "file_name": "21250.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1783
+ },
+ {
+ "file_name": "18138.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1784
+ },
+ {
+ "file_name": "12090.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1785
+ },
+ {
+ "file_name": "16239.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1786
+ },
+ {
+ "file_name": "19997.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1787
+ },
+ {
+ "file_name": "21524.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1788
+ },
+ {
+ "file_name": "12064.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1789
+ },
+ {
+ "file_name": "21630.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1790
+ },
+ {
+ "file_name": "17435.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1791
+ },
+ {
+ "file_name": "14094.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1792
+ },
+ {
+ "file_name": "17960.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1793
+ },
+ {
+ "file_name": "13994.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1794
+ },
+ {
+ "file_name": "19303.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1795
+ },
+ {
+ "file_name": "22299.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1796
+ },
+ {
+ "file_name": "14807.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1797
+ },
+ {
+ "file_name": "15794.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1798
+ },
+ {
+ "file_name": "13009.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1799
+ },
+ {
+ "file_name": "21216.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1800
+ },
+ {
+ "file_name": "16806.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1801
+ },
+ {
+ "file_name": "15802.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1802
+ },
+ {
+ "file_name": "17806.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1803
+ },
+ {
+ "file_name": "12681.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1804
+ },
+ {
+ "file_name": "18780.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1805
+ },
+ {
+ "file_name": "18872.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1806
+ },
+ {
+ "file_name": "12324.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1807
+ },
+ {
+ "file_name": "12632.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1808
+ },
+ {
+ "file_name": "16658.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1809
+ },
+ {
+ "file_name": "15420.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1810
+ },
+ {
+ "file_name": "12682.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1811
+ },
+ {
+ "file_name": "13873.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1812
+ },
+ {
+ "file_name": "22419.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1813
+ },
+ {
+ "file_name": "13290.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1814
+ },
+ {
+ "file_name": "21580.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1815
+ },
+ {
+ "file_name": "14064.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1816
+ },
+ {
+ "file_name": "20049.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1817
+ },
+ {
+ "file_name": "14452.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1818
+ },
+ {
+ "file_name": "20585.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1819
+ },
+ {
+ "file_name": "18925.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1820
+ },
+ {
+ "file_name": "15227.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1821
+ },
+ {
+ "file_name": "12953.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1822
+ },
+ {
+ "file_name": "17314.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1823
+ },
+ {
+ "file_name": "14347.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1824
+ },
+ {
+ "file_name": "17136.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1825
+ },
+ {
+ "file_name": "14547.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1826
+ },
+ {
+ "file_name": "15980.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1827
+ },
+ {
+ "file_name": "17258.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1828
+ },
+ {
+ "file_name": "11774.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1829
+ },
+ {
+ "file_name": "18557.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1830
+ },
+ {
+ "file_name": "13677.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1831
+ },
+ {
+ "file_name": "21100.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1832
+ },
+ {
+ "file_name": "22490.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1833
+ },
+ {
+ "file_name": "19204.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1834
+ },
+ {
+ "file_name": "16947.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1835
+ },
+ {
+ "file_name": "13163.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1836
+ },
+ {
+ "file_name": "22305.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1837
+ },
+ {
+ "file_name": "16869.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1838
+ },
+ {
+ "file_name": "20275.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1839
+ },
+ {
+ "file_name": "19063.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1840
+ },
+ {
+ "file_name": "15520.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1841
+ },
+ {
+ "file_name": "14139.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1842
+ },
+ {
+ "file_name": "18416.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1843
+ },
+ {
+ "file_name": "18058.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1844
+ },
+ {
+ "file_name": "18317.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1845
+ },
+ {
+ "file_name": "15734.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1846
+ },
+ {
+ "file_name": "21462.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1847
+ },
+ {
+ "file_name": "18568.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1848
+ },
+ {
+ "file_name": "18196.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1849
+ },
+ {
+ "file_name": "21534.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1850
+ },
+ {
+ "file_name": "21623.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1851
+ },
+ {
+ "file_name": "14042.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1852
+ },
+ {
+ "file_name": "22183.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1853
+ },
+ {
+ "file_name": "18545.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1854
+ },
+ {
+ "file_name": "15695.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1855
+ },
+ {
+ "file_name": "18099.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1856
+ },
+ {
+ "file_name": "20246.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1857
+ },
+ {
+ "file_name": "18413.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1858
+ },
+ {
+ "file_name": "21779.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1859
+ },
+ {
+ "file_name": "13343.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1860
+ },
+ {
+ "file_name": "15764.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1861
+ },
+ {
+ "file_name": "18909.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1862
+ },
+ {
+ "file_name": "12502.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1863
+ },
+ {
+ "file_name": "15474.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1864
+ },
+ {
+ "file_name": "16120.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1865
+ },
+ {
+ "file_name": "16242.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1866
+ },
+ {
+ "file_name": "20128.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1867
+ },
+ {
+ "file_name": "17183.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1868
+ },
+ {
+ "file_name": "21075.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1869
+ },
+ {
+ "file_name": "22060.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1870
+ },
+ {
+ "file_name": "12592.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1871
+ },
+ {
+ "file_name": "20175.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1872
+ },
+ {
+ "file_name": "20760.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1873
+ },
+ {
+ "file_name": "18526.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1874
+ },
+ {
+ "file_name": "16554.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1875
+ },
+ {
+ "file_name": "17674.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1876
+ },
+ {
+ "file_name": "18882.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1877
+ },
+ {
+ "file_name": "23276.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1878
+ },
+ {
+ "file_name": "17742.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1879
+ },
+ {
+ "file_name": "20093.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1880
+ },
+ {
+ "file_name": "19418.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1881
+ },
+ {
+ "file_name": "16385.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1882
+ },
+ {
+ "file_name": "12838.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1883
+ },
+ {
+ "file_name": "21855.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1884
+ },
+ {
+ "file_name": "23180.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1885
+ },
+ {
+ "file_name": "19637.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1886
+ },
+ {
+ "file_name": "20635.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1887
+ },
+ {
+ "file_name": "20900.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1888
+ },
+ {
+ "file_name": "14228.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1889
+ },
+ {
+ "file_name": "23082.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1890
+ },
+ {
+ "file_name": "21943.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1891
+ },
+ {
+ "file_name": "17628.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1892
+ },
+ {
+ "file_name": "14676.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1893
+ },
+ {
+ "file_name": "12271.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1894
+ },
+ {
+ "file_name": "20890.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1895
+ },
+ {
+ "file_name": "19507.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1896
+ },
+ {
+ "file_name": "20443.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1897
+ },
+ {
+ "file_name": "20204.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1898
+ },
+ {
+ "file_name": "23376.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1899
+ },
+ {
+ "file_name": "20630.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1900
+ },
+ {
+ "file_name": "22382.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1901
+ },
+ {
+ "file_name": "14685.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1902
+ },
+ {
+ "file_name": "12051.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1903
+ },
+ {
+ "file_name": "21222.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1904
+ },
+ {
+ "file_name": "17454.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1905
+ },
+ {
+ "file_name": "14696.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1906
+ },
+ {
+ "file_name": "18108.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1907
+ },
+ {
+ "file_name": "17442.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1908
+ },
+ {
+ "file_name": "15975.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1909
+ },
+ {
+ "file_name": "14767.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1910
+ },
+ {
+ "file_name": "22828.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1911
+ },
+ {
+ "file_name": "22496.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1912
+ },
+ {
+ "file_name": "20018.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1913
+ },
+ {
+ "file_name": "23125.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1914
+ },
+ {
+ "file_name": "13763.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1915
+ },
+ {
+ "file_name": "13427.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1916
+ },
+ {
+ "file_name": "15440.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1917
+ },
+ {
+ "file_name": "17094.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1918
+ },
+ {
+ "file_name": "13687.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1919
+ },
+ {
+ "file_name": "17558.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1920
+ },
+ {
+ "file_name": "23093.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1921
+ },
+ {
+ "file_name": "19495.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1922
+ },
+ {
+ "file_name": "23352.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1923
+ },
+ {
+ "file_name": "23173.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1924
+ },
+ {
+ "file_name": "20351.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1925
+ },
+ {
+ "file_name": "17074.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1926
+ },
+ {
+ "file_name": "13709.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1927
+ },
+ {
+ "file_name": "14652.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1928
+ },
+ {
+ "file_name": "23315.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1929
+ },
+ {
+ "file_name": "18939.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1930
+ },
+ {
+ "file_name": "18445.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1931
+ },
+ {
+ "file_name": "23246.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1932
+ },
+ {
+ "file_name": "17334.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1933
+ },
+ {
+ "file_name": "13648.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1934
+ },
+ {
+ "file_name": "13745.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1935
+ },
+ {
+ "file_name": "13645.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1936
+ },
+ {
+ "file_name": "11867.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1937
+ },
+ {
+ "file_name": "21691.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1938
+ },
+ {
+ "file_name": "13573.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1939
+ },
+ {
+ "file_name": "18573.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1940
+ },
+ {
+ "file_name": "20279.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1941
+ },
+ {
+ "file_name": "17908.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1942
+ },
+ {
+ "file_name": "12276.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1943
+ },
+ {
+ "file_name": "15549.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1944
+ },
+ {
+ "file_name": "21954.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1945
+ },
+ {
+ "file_name": "13470.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1946
+ },
+ {
+ "file_name": "12335.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1947
+ },
+ {
+ "file_name": "13646.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1948
+ },
+ {
+ "file_name": "19644.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1949
+ },
+ {
+ "file_name": "17733.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1950
+ },
+ {
+ "file_name": "21594.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1951
+ },
+ {
+ "file_name": "19950.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1952
+ },
+ {
+ "file_name": "22172.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1953
+ },
+ {
+ "file_name": "19514.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1954
+ },
+ {
+ "file_name": "16423.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1955
+ },
+ {
+ "file_name": "22748.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1956
+ },
+ {
+ "file_name": "17142.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1957
+ },
+ {
+ "file_name": "16860.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1958
+ },
+ {
+ "file_name": "14758.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1959
+ },
+ {
+ "file_name": "13190.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1960
+ },
+ {
+ "file_name": "20324.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1961
+ },
+ {
+ "file_name": "17689.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1962
+ },
+ {
+ "file_name": "17447.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1963
+ },
+ {
+ "file_name": "22880.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1964
+ },
+ {
+ "file_name": "19436.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1965
+ },
+ {
+ "file_name": "17233.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1966
+ },
+ {
+ "file_name": "21681.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1967
+ },
+ {
+ "file_name": "14211.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1968
+ },
+ {
+ "file_name": "13621.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1969
+ },
+ {
+ "file_name": "18869.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1970
+ },
+ {
+ "file_name": "15836.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1971
+ },
+ {
+ "file_name": "21682.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1972
+ },
+ {
+ "file_name": "19568.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1973
+ },
+ {
+ "file_name": "22975.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1974
+ },
+ {
+ "file_name": "14991.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1975
+ },
+ {
+ "file_name": "18694.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1976
+ },
+ {
+ "file_name": "14425.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1977
+ },
+ {
+ "file_name": "18424.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1978
+ },
+ {
+ "file_name": "16285.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1979
+ },
+ {
+ "file_name": "22559.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1980
+ },
+ {
+ "file_name": "22745.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1981
+ },
+ {
+ "file_name": "12551.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1982
+ },
+ {
+ "file_name": "19328.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1983
+ },
+ {
+ "file_name": "12153.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1984
+ },
+ {
+ "file_name": "12634.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1985
+ },
+ {
+ "file_name": "19028.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1986
+ },
+ {
+ "file_name": "15948.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1987
+ },
+ {
+ "file_name": "22430.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1988
+ },
+ {
+ "file_name": "19051.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1989
+ },
+ {
+ "file_name": "18479.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1990
+ },
+ {
+ "file_name": "13980.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1991
+ },
+ {
+ "file_name": "20747.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1992
+ },
+ {
+ "file_name": "17617.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1993
+ },
+ {
+ "file_name": "17997.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1994
+ },
+ {
+ "file_name": "13911.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1995
+ },
+ {
+ "file_name": "17276.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1996
+ },
+ {
+ "file_name": "16424.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1997
+ },
+ {
+ "file_name": "18743.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1998
+ },
+ {
+ "file_name": "23062.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 1999
+ },
+ {
+ "file_name": "15996.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2000
+ },
+ {
+ "file_name": "22463.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2001
+ },
+ {
+ "file_name": "22576.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2002
+ },
+ {
+ "file_name": "20218.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2003
+ },
+ {
+ "file_name": "22733.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2004
+ },
+ {
+ "file_name": "21863.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2005
+ },
+ {
+ "file_name": "19510.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2006
+ },
+ {
+ "file_name": "18189.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2007
+ },
+ {
+ "file_name": "20995.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2008
+ },
+ {
+ "file_name": "19427.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2009
+ },
+ {
+ "file_name": "14305.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2010
+ },
+ {
+ "file_name": "14170.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2011
+ },
+ {
+ "file_name": "13594.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2012
+ },
+ {
+ "file_name": "23252.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2013
+ },
+ {
+ "file_name": "22265.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2014
+ },
+ {
+ "file_name": "19896.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2015
+ },
+ {
+ "file_name": "18139.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2016
+ },
+ {
+ "file_name": "20831.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2017
+ },
+ {
+ "file_name": "22395.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2018
+ },
+ {
+ "file_name": "22932.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2019
+ },
+ {
+ "file_name": "15635.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2020
+ },
+ {
+ "file_name": "16670.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2021
+ },
+ {
+ "file_name": "18007.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2022
+ },
+ {
+ "file_name": "16851.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2023
+ },
+ {
+ "file_name": "19671.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2024
+ },
+ {
+ "file_name": "12887.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2025
+ },
+ {
+ "file_name": "19491.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2026
+ },
+ {
+ "file_name": "22632.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2027
+ },
+ {
+ "file_name": "22049.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2028
+ },
+ {
+ "file_name": "13428.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2029
+ },
+ {
+ "file_name": "11947.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2030
+ },
+ {
+ "file_name": "22166.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2031
+ },
+ {
+ "file_name": "23437.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2032
+ },
+ {
+ "file_name": "19638.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2033
+ },
+ {
+ "file_name": "18455.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2034
+ },
+ {
+ "file_name": "22198.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2035
+ },
+ {
+ "file_name": "22401.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2036
+ },
+ {
+ "file_name": "14561.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2037
+ },
+ {
+ "file_name": "13647.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2038
+ },
+ {
+ "file_name": "20074.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2039
+ },
+ {
+ "file_name": "12557.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2040
+ },
+ {
+ "file_name": "12825.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2041
+ },
+ {
+ "file_name": "19575.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2042
+ },
+ {
+ "file_name": "18087.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2043
+ },
+ {
+ "file_name": "18529.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2044
+ },
+ {
+ "file_name": "19120.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2045
+ },
+ {
+ "file_name": "17020.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2046
+ },
+ {
+ "file_name": "15182.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2047
+ },
+ {
+ "file_name": "16844.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2048
+ },
+ {
+ "file_name": "21405.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2049
+ },
+ {
+ "file_name": "15451.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2050
+ },
+ {
+ "file_name": "15203.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2051
+ },
+ {
+ "file_name": "14733.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2052
+ },
+ {
+ "file_name": "12637.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2053
+ },
+ {
+ "file_name": "14551.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2054
+ },
+ {
+ "file_name": "20244.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2055
+ },
+ {
+ "file_name": "15923.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2056
+ },
+ {
+ "file_name": "22471.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2057
+ },
+ {
+ "file_name": "21029.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2058
+ },
+ {
+ "file_name": "16748.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2059
+ },
+ {
+ "file_name": "14636.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2060
+ },
+ {
+ "file_name": "12393.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2061
+ },
+ {
+ "file_name": "18671.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2062
+ },
+ {
+ "file_name": "11859.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2063
+ },
+ {
+ "file_name": "18040.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2064
+ },
+ {
+ "file_name": "20638.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2065
+ },
+ {
+ "file_name": "20599.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2066
+ },
+ {
+ "file_name": "12265.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2067
+ },
+ {
+ "file_name": "12806.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2068
+ },
+ {
+ "file_name": "18239.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2069
+ },
+ {
+ "file_name": "14365.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2070
+ },
+ {
+ "file_name": "17647.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2071
+ },
+ {
+ "file_name": "22905.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2072
+ },
+ {
+ "file_name": "16969.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2073
+ },
+ {
+ "file_name": "17004.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2074
+ },
+ {
+ "file_name": "16966.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2075
+ },
+ {
+ "file_name": "22444.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2076
+ },
+ {
+ "file_name": "21046.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2077
+ },
+ {
+ "file_name": "12743.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2078
+ },
+ {
+ "file_name": "14330.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2079
+ },
+ {
+ "file_name": "16703.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2080
+ },
+ {
+ "file_name": "19229.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2081
+ },
+ {
+ "file_name": "12722.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2082
+ },
+ {
+ "file_name": "23317.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2083
+ },
+ {
+ "file_name": "23048.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2084
+ },
+ {
+ "file_name": "20695.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2085
+ },
+ {
+ "file_name": "17563.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2086
+ },
+ {
+ "file_name": "22536.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2087
+ },
+ {
+ "file_name": "13499.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2088
+ },
+ {
+ "file_name": "15751.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2089
+ },
+ {
+ "file_name": "20906.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2090
+ },
+ {
+ "file_name": "19688.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2091
+ },
+ {
+ "file_name": "13851.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2092
+ },
+ {
+ "file_name": "21135.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2093
+ },
+ {
+ "file_name": "18112.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2094
+ },
+ {
+ "file_name": "19466.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2095
+ },
+ {
+ "file_name": "16689.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2096
+ },
+ {
+ "file_name": "14673.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2097
+ },
+ {
+ "file_name": "19723.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2098
+ },
+ {
+ "file_name": "21853.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2099
+ },
+ {
+ "file_name": "21903.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2100
+ },
+ {
+ "file_name": "16544.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2101
+ },
+ {
+ "file_name": "22154.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2102
+ },
+ {
+ "file_name": "22413.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2103
+ },
+ {
+ "file_name": "14392.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2104
+ },
+ {
+ "file_name": "19282.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2105
+ },
+ {
+ "file_name": "21512.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2106
+ },
+ {
+ "file_name": "17588.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2107
+ },
+ {
+ "file_name": "18928.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2108
+ },
+ {
+ "file_name": "23301.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2109
+ },
+ {
+ "file_name": "18706.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2110
+ },
+ {
+ "file_name": "16091.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2111
+ },
+ {
+ "file_name": "21922.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2112
+ },
+ {
+ "file_name": "22120.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2113
+ },
+ {
+ "file_name": "21710.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2114
+ },
+ {
+ "file_name": "22090.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2115
+ },
+ {
+ "file_name": "16490.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2116
+ },
+ {
+ "file_name": "20681.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2117
+ },
+ {
+ "file_name": "14091.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2118
+ },
+ {
+ "file_name": "21173.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2119
+ },
+ {
+ "file_name": "17604.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2120
+ },
+ {
+ "file_name": "15706.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2121
+ },
+ {
+ "file_name": "14489.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2122
+ },
+ {
+ "file_name": "22418.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2123
+ },
+ {
+ "file_name": "12164.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2124
+ },
+ {
+ "file_name": "13015.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2125
+ },
+ {
+ "file_name": "22017.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2126
+ },
+ {
+ "file_name": "21794.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2127
+ },
+ {
+ "file_name": "16214.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2128
+ },
+ {
+ "file_name": "12313.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2129
+ },
+ {
+ "file_name": "16433.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2130
+ },
+ {
+ "file_name": "23299.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2131
+ },
+ {
+ "file_name": "20687.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2132
+ },
+ {
+ "file_name": "17673.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2133
+ },
+ {
+ "file_name": "18463.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2134
+ },
+ {
+ "file_name": "13134.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2135
+ },
+ {
+ "file_name": "19157.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2136
+ },
+ {
+ "file_name": "19718.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2137
+ },
+ {
+ "file_name": "22568.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2138
+ },
+ {
+ "file_name": "14464.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2139
+ },
+ {
+ "file_name": "19947.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2140
+ },
+ {
+ "file_name": "13010.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2141
+ },
+ {
+ "file_name": "21774.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2142
+ },
+ {
+ "file_name": "21012.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2143
+ },
+ {
+ "file_name": "20137.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2144
+ },
+ {
+ "file_name": "20458.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2145
+ },
+ {
+ "file_name": "14813.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2146
+ },
+ {
+ "file_name": "20718.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2147
+ },
+ {
+ "file_name": "21546.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2148
+ },
+ {
+ "file_name": "18186.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2149
+ },
+ {
+ "file_name": "14792.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2150
+ },
+ {
+ "file_name": "16723.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2151
+ },
+ {
+ "file_name": "16771.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2152
+ },
+ {
+ "file_name": "22796.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2153
+ },
+ {
+ "file_name": "19726.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2154
+ },
+ {
+ "file_name": "17889.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2155
+ },
+ {
+ "file_name": "18564.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2156
+ },
+ {
+ "file_name": "22793.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2157
+ },
+ {
+ "file_name": "16970.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2158
+ },
+ {
+ "file_name": "16332.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2159
+ },
+ {
+ "file_name": "17857.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2160
+ },
+ {
+ "file_name": "17060.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2161
+ },
+ {
+ "file_name": "12179.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2162
+ },
+ {
+ "file_name": "16116.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2163
+ },
+ {
+ "file_name": "20758.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2164
+ },
+ {
+ "file_name": "12758.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2165
+ },
+ {
+ "file_name": "12136.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2166
+ },
+ {
+ "file_name": "16037.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2167
+ },
+ {
+ "file_name": "14542.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2168
+ },
+ {
+ "file_name": "18125.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2169
+ },
+ {
+ "file_name": "23380.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2170
+ },
+ {
+ "file_name": "18517.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2171
+ },
+ {
+ "file_name": "20124.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2172
+ },
+ {
+ "file_name": "17855.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2173
+ },
+ {
+ "file_name": "16563.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2174
+ },
+ {
+ "file_name": "19121.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2175
+ },
+ {
+ "file_name": "23130.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2176
+ },
+ {
+ "file_name": "16585.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2177
+ },
+ {
+ "file_name": "16995.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2178
+ },
+ {
+ "file_name": "19883.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2179
+ },
+ {
+ "file_name": "18148.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2180
+ },
+ {
+ "file_name": "20852.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2181
+ },
+ {
+ "file_name": "16556.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2182
+ },
+ {
+ "file_name": "21841.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2183
+ },
+ {
+ "file_name": "21439.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2184
+ },
+ {
+ "file_name": "20793.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2185
+ },
+ {
+ "file_name": "21747.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2186
+ },
+ {
+ "file_name": "11950.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2187
+ },
+ {
+ "file_name": "23358.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2188
+ },
+ {
+ "file_name": "19152.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2189
+ },
+ {
+ "file_name": "20234.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2190
+ },
+ {
+ "file_name": "12033.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2191
+ },
+ {
+ "file_name": "15961.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2192
+ },
+ {
+ "file_name": "16605.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2193
+ },
+ {
+ "file_name": "12553.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2194
+ },
+ {
+ "file_name": "19794.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2195
+ },
+ {
+ "file_name": "15508.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2196
+ },
+ {
+ "file_name": "22022.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2197
+ },
+ {
+ "file_name": "13629.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2198
+ },
+ {
+ "file_name": "17869.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2199
+ },
+ {
+ "file_name": "15669.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2200
+ },
+ {
+ "file_name": "19706.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2201
+ },
+ {
+ "file_name": "12512.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2202
+ },
+ {
+ "file_name": "18963.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2203
+ },
+ {
+ "file_name": "23354.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2204
+ },
+ {
+ "file_name": "14140.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2205
+ },
+ {
+ "file_name": "13025.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2206
+ },
+ {
+ "file_name": "15529.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2207
+ },
+ {
+ "file_name": "17799.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2208
+ },
+ {
+ "file_name": "20361.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2209
+ },
+ {
+ "file_name": "12568.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2210
+ },
+ {
+ "file_name": "15612.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2211
+ },
+ {
+ "file_name": "12353.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2212
+ },
+ {
+ "file_name": "18772.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2213
+ },
+ {
+ "file_name": "18073.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2214
+ },
+ {
+ "file_name": "16104.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2215
+ },
+ {
+ "file_name": "23220.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2216
+ },
+ {
+ "file_name": "14174.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2217
+ },
+ {
+ "file_name": "22731.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2218
+ },
+ {
+ "file_name": "16591.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2219
+ },
+ {
+ "file_name": "17064.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2220
+ },
+ {
+ "file_name": "15145.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2221
+ },
+ {
+ "file_name": "15199.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2222
+ },
+ {
+ "file_name": "23418.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2223
+ },
+ {
+ "file_name": "13852.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2224
+ },
+ {
+ "file_name": "23018.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2225
+ },
+ {
+ "file_name": "17378.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2226
+ },
+ {
+ "file_name": "15646.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2227
+ },
+ {
+ "file_name": "21084.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2228
+ },
+ {
+ "file_name": "23438.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2229
+ },
+ {
+ "file_name": "15845.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2230
+ },
+ {
+ "file_name": "17691.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2231
+ },
+ {
+ "file_name": "22824.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2232
+ },
+ {
+ "file_name": "14857.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2233
+ },
+ {
+ "file_name": "18439.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2234
+ },
+ {
+ "file_name": "16103.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2235
+ },
+ {
+ "file_name": "21723.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2236
+ },
+ {
+ "file_name": "15558.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2237
+ },
+ {
+ "file_name": "17632.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2238
+ },
+ {
+ "file_name": "15387.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2239
+ },
+ {
+ "file_name": "20948.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2240
+ },
+ {
+ "file_name": "12259.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2241
+ },
+ {
+ "file_name": "20833.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2242
+ },
+ {
+ "file_name": "14634.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2243
+ },
+ {
+ "file_name": "16320.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2244
+ },
+ {
+ "file_name": "23185.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2245
+ },
+ {
+ "file_name": "15104.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2246
+ },
+ {
+ "file_name": "16168.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2247
+ },
+ {
+ "file_name": "16600.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2248
+ },
+ {
+ "file_name": "15599.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2249
+ },
+ {
+ "file_name": "16243.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2250
+ },
+ {
+ "file_name": "20276.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2251
+ },
+ {
+ "file_name": "16506.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2252
+ },
+ {
+ "file_name": "11827.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2253
+ },
+ {
+ "file_name": "12532.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2254
+ },
+ {
+ "file_name": "18818.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2255
+ },
+ {
+ "file_name": "12373.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2256
+ },
+ {
+ "file_name": "23363.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2257
+ },
+ {
+ "file_name": "13022.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2258
+ },
+ {
+ "file_name": "17279.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2259
+ },
+ {
+ "file_name": "12025.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2260
+ },
+ {
+ "file_name": "21495.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2261
+ },
+ {
+ "file_name": "19511.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2262
+ },
+ {
+ "file_name": "17619.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2263
+ },
+ {
+ "file_name": "14020.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2264
+ },
+ {
+ "file_name": "18093.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2265
+ },
+ {
+ "file_name": "16911.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2266
+ },
+ {
+ "file_name": "14795.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2267
+ },
+ {
+ "file_name": "21205.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2268
+ },
+ {
+ "file_name": "14255.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2269
+ },
+ {
+ "file_name": "13328.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2270
+ },
+ {
+ "file_name": "20982.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2271
+ },
+ {
+ "file_name": "19876.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2272
+ },
+ {
+ "file_name": "11802.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2273
+ },
+ {
+ "file_name": "11776.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2274
+ },
+ {
+ "file_name": "23417.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2275
+ },
+ {
+ "file_name": "21407.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2276
+ },
+ {
+ "file_name": "12056.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2277
+ },
+ {
+ "file_name": "16055.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2278
+ },
+ {
+ "file_name": "14947.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2279
+ },
+ {
+ "file_name": "12586.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2280
+ },
+ {
+ "file_name": "22257.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2281
+ },
+ {
+ "file_name": "23188.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2282
+ },
+ {
+ "file_name": "18348.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2283
+ },
+ {
+ "file_name": "11740.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2284
+ },
+ {
+ "file_name": "17410.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2285
+ },
+ {
+ "file_name": "17931.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2286
+ },
+ {
+ "file_name": "14876.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2287
+ },
+ {
+ "file_name": "18992.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2288
+ },
+ {
+ "file_name": "12868.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2289
+ },
+ {
+ "file_name": "23424.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2290
+ },
+ {
+ "file_name": "21217.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2291
+ },
+ {
+ "file_name": "21197.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2292
+ },
+ {
+ "file_name": "16650.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2293
+ },
+ {
+ "file_name": "22987.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2294
+ },
+ {
+ "file_name": "21964.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2295
+ },
+ {
+ "file_name": "22899.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2296
+ },
+ {
+ "file_name": "16048.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2297
+ },
+ {
+ "file_name": "12451.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2298
+ },
+ {
+ "file_name": "12162.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2299
+ },
+ {
+ "file_name": "13462.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2300
+ },
+ {
+ "file_name": "20640.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2301
+ },
+ {
+ "file_name": "18279.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2302
+ },
+ {
+ "file_name": "12765.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2303
+ },
+ {
+ "file_name": "16390.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2304
+ },
+ {
+ "file_name": "22966.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2305
+ },
+ {
+ "file_name": "17756.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2306
+ },
+ {
+ "file_name": "16373.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2307
+ },
+ {
+ "file_name": "11745.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2308
+ },
+ {
+ "file_name": "18971.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2309
+ },
+ {
+ "file_name": "13548.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2310
+ },
+ {
+ "file_name": "15545.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2311
+ },
+ {
+ "file_name": "20022.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2312
+ },
+ {
+ "file_name": "12488.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2313
+ },
+ {
+ "file_name": "12952.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2314
+ },
+ {
+ "file_name": "14026.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2315
+ },
+ {
+ "file_name": "20494.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2316
+ },
+ {
+ "file_name": "16004.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2317
+ },
+ {
+ "file_name": "15379.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2318
+ },
+ {
+ "file_name": "14502.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2319
+ },
+ {
+ "file_name": "17496.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2320
+ },
+ {
+ "file_name": "14671.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2321
+ },
+ {
+ "file_name": "18420.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2322
+ },
+ {
+ "file_name": "16957.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2323
+ },
+ {
+ "file_name": "21049.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2324
+ },
+ {
+ "file_name": "18724.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2325
+ },
+ {
+ "file_name": "20055.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2326
+ },
+ {
+ "file_name": "16437.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2327
+ },
+ {
+ "file_name": "13467.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2328
+ },
+ {
+ "file_name": "17941.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2329
+ },
+ {
+ "file_name": "20069.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2330
+ },
+ {
+ "file_name": "16481.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2331
+ },
+ {
+ "file_name": "16546.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2332
+ },
+ {
+ "file_name": "15679.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2333
+ },
+ {
+ "file_name": "12408.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2334
+ },
+ {
+ "file_name": "19085.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2335
+ },
+ {
+ "file_name": "17323.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2336
+ },
+ {
+ "file_name": "13937.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2337
+ },
+ {
+ "file_name": "12751.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2338
+ },
+ {
+ "file_name": "12037.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2339
+ },
+ {
+ "file_name": "15230.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2340
+ },
+ {
+ "file_name": "22161.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2341
+ },
+ {
+ "file_name": "18760.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2342
+ },
+ {
+ "file_name": "17486.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2343
+ },
+ {
+ "file_name": "17808.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2344
+ },
+ {
+ "file_name": "15034.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2345
+ },
+ {
+ "file_name": "14448.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2346
+ },
+ {
+ "file_name": "19421.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2347
+ },
+ {
+ "file_name": "18362.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2348
+ },
+ {
+ "file_name": "14564.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2349
+ },
+ {
+ "file_name": "21010.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2350
+ },
+ {
+ "file_name": "14415.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2351
+ },
+ {
+ "file_name": "16351.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2352
+ },
+ {
+ "file_name": "17007.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2353
+ },
+ {
+ "file_name": "19534.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2354
+ },
+ {
+ "file_name": "21985.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2355
+ },
+ {
+ "file_name": "17926.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2356
+ },
+ {
+ "file_name": "22941.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2357
+ },
+ {
+ "file_name": "14613.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2358
+ },
+ {
+ "file_name": "12509.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2359
+ },
+ {
+ "file_name": "16696.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2360
+ },
+ {
+ "file_name": "17105.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2361
+ },
+ {
+ "file_name": "17717.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2362
+ },
+ {
+ "file_name": "20057.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2363
+ },
+ {
+ "file_name": "16173.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2364
+ },
+ {
+ "file_name": "16147.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2365
+ },
+ {
+ "file_name": "13055.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2366
+ },
+ {
+ "file_name": "14769.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2367
+ },
+ {
+ "file_name": "13012.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2368
+ },
+ {
+ "file_name": "20328.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2369
+ },
+ {
+ "file_name": "15187.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2370
+ },
+ {
+ "file_name": "13756.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2371
+ },
+ {
+ "file_name": "12515.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2372
+ },
+ {
+ "file_name": "12542.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2373
+ },
+ {
+ "file_name": "21121.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2374
+ },
+ {
+ "file_name": "18278.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2375
+ },
+ {
+ "file_name": "13596.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2376
+ },
+ {
+ "file_name": "14635.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2377
+ },
+ {
+ "file_name": "14104.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2378
+ },
+ {
+ "file_name": "19056.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2379
+ },
+ {
+ "file_name": "15003.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2380
+ },
+ {
+ "file_name": "21559.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2381
+ },
+ {
+ "file_name": "17361.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2382
+ },
+ {
+ "file_name": "12655.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2383
+ },
+ {
+ "file_name": "13260.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2384
+ },
+ {
+ "file_name": "19073.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2385
+ },
+ {
+ "file_name": "22316.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2386
+ },
+ {
+ "file_name": "15720.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2387
+ },
+ {
+ "file_name": "23445.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2388
+ },
+ {
+ "file_name": "14914.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2389
+ },
+ {
+ "file_name": "17508.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2390
+ },
+ {
+ "file_name": "15164.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2391
+ },
+ {
+ "file_name": "17174.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2392
+ },
+ {
+ "file_name": "22171.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2393
+ },
+ {
+ "file_name": "18427.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2394
+ },
+ {
+ "file_name": "20836.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2395
+ },
+ {
+ "file_name": "17207.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2396
+ },
+ {
+ "file_name": "15539.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2397
+ },
+ {
+ "file_name": "12734.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2398
+ },
+ {
+ "file_name": "14421.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2399
+ },
+ {
+ "file_name": "19336.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2400
+ },
+ {
+ "file_name": "19558.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2401
+ },
+ {
+ "file_name": "13790.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2402
+ },
+ {
+ "file_name": "16541.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2403
+ },
+ {
+ "file_name": "20514.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2404
+ },
+ {
+ "file_name": "13498.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2405
+ },
+ {
+ "file_name": "19404.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2406
+ },
+ {
+ "file_name": "18998.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2407
+ },
+ {
+ "file_name": "18121.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2408
+ },
+ {
+ "file_name": "21527.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2409
+ },
+ {
+ "file_name": "12511.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2410
+ },
+ {
+ "file_name": "17121.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2411
+ },
+ {
+ "file_name": "12505.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2412
+ },
+ {
+ "file_name": "12428.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2413
+ },
+ {
+ "file_name": "17590.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2414
+ },
+ {
+ "file_name": "13355.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2415
+ },
+ {
+ "file_name": "21446.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2416
+ },
+ {
+ "file_name": "17594.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2417
+ },
+ {
+ "file_name": "17158.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2418
+ },
+ {
+ "file_name": "22738.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2419
+ },
+ {
+ "file_name": "18610.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2420
+ },
+ {
+ "file_name": "22759.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2421
+ },
+ {
+ "file_name": "12412.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2422
+ },
+ {
+ "file_name": "16290.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2423
+ },
+ {
+ "file_name": "19792.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2424
+ },
+ {
+ "file_name": "20704.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2425
+ },
+ {
+ "file_name": "18114.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2426
+ },
+ {
+ "file_name": "12366.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2427
+ },
+ {
+ "file_name": "23110.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2428
+ },
+ {
+ "file_name": "16081.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2429
+ },
+ {
+ "file_name": "14660.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2430
+ },
+ {
+ "file_name": "17714.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2431
+ },
+ {
+ "file_name": "15049.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2432
+ },
+ {
+ "file_name": "19307.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2433
+ },
+ {
+ "file_name": "19088.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2434
+ },
+ {
+ "file_name": "17112.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2435
+ },
+ {
+ "file_name": "12529.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2436
+ },
+ {
+ "file_name": "12611.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2437
+ },
+ {
+ "file_name": "19119.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2438
+ },
+ {
+ "file_name": "15109.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2439
+ },
+ {
+ "file_name": "20567.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2440
+ },
+ {
+ "file_name": "16248.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2441
+ },
+ {
+ "file_name": "20522.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2442
+ },
+ {
+ "file_name": "13702.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2443
+ },
+ {
+ "file_name": "12291.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2444
+ },
+ {
+ "file_name": "15193.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2445
+ },
+ {
+ "file_name": "18256.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2446
+ },
+ {
+ "file_name": "19431.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2447
+ },
+ {
+ "file_name": "21619.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2448
+ },
+ {
+ "file_name": "16470.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2449
+ },
+ {
+ "file_name": "21665.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2450
+ },
+ {
+ "file_name": "21117.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2451
+ },
+ {
+ "file_name": "16106.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2452
+ },
+ {
+ "file_name": "15072.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2453
+ },
+ {
+ "file_name": "18241.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2454
+ },
+ {
+ "file_name": "20181.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2455
+ },
+ {
+ "file_name": "17514.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2456
+ },
+ {
+ "file_name": "23027.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2457
+ },
+ {
+ "file_name": "16878.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2458
+ },
+ {
+ "file_name": "20565.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2459
+ },
+ {
+ "file_name": "17760.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2460
+ },
+ {
+ "file_name": "22036.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2461
+ },
+ {
+ "file_name": "13686.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2462
+ },
+ {
+ "file_name": "11877.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2463
+ },
+ {
+ "file_name": "22991.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2464
+ },
+ {
+ "file_name": "22164.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2465
+ },
+ {
+ "file_name": "17734.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2466
+ },
+ {
+ "file_name": "13968.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2467
+ },
+ {
+ "file_name": "23182.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2468
+ },
+ {
+ "file_name": "14825.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2469
+ },
+ {
+ "file_name": "16412.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2470
+ },
+ {
+ "file_name": "18782.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2471
+ },
+ {
+ "file_name": "13500.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2472
+ },
+ {
+ "file_name": "18553.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2473
+ },
+ {
+ "file_name": "16939.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2474
+ },
+ {
+ "file_name": "17913.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2475
+ },
+ {
+ "file_name": "17793.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2476
+ },
+ {
+ "file_name": "15785.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2477
+ },
+ {
+ "file_name": "15369.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2478
+ },
+ {
+ "file_name": "17646.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2479
+ },
+ {
+ "file_name": "15949.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2480
+ },
+ {
+ "file_name": "15882.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2481
+ },
+ {
+ "file_name": "19351.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2482
+ },
+ {
+ "file_name": "22454.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2483
+ },
+ {
+ "file_name": "23371.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2484
+ },
+ {
+ "file_name": "14691.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2485
+ },
+ {
+ "file_name": "16140.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2486
+ },
+ {
+ "file_name": "14223.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2487
+ },
+ {
+ "file_name": "16201.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2488
+ },
+ {
+ "file_name": "22385.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2489
+ },
+ {
+ "file_name": "15395.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2490
+ },
+ {
+ "file_name": "14859.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2491
+ },
+ {
+ "file_name": "13918.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2492
+ },
+ {
+ "file_name": "19914.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2493
+ },
+ {
+ "file_name": "17565.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2494
+ },
+ {
+ "file_name": "21531.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2495
+ },
+ {
+ "file_name": "22119.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2496
+ },
+ {
+ "file_name": "19649.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2497
+ },
+ {
+ "file_name": "11975.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2498
+ },
+ {
+ "file_name": "19592.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2499
+ },
+ {
+ "file_name": "17232.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2500
+ },
+ {
+ "file_name": "15586.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2501
+ },
+ {
+ "file_name": "15640.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2502
+ },
+ {
+ "file_name": "22294.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2503
+ },
+ {
+ "file_name": "22038.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2504
+ },
+ {
+ "file_name": "22747.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2505
+ },
+ {
+ "file_name": "14893.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2506
+ },
+ {
+ "file_name": "16608.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2507
+ },
+ {
+ "file_name": "16800.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2508
+ },
+ {
+ "file_name": "15807.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2509
+ },
+ {
+ "file_name": "18756.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2510
+ },
+ {
+ "file_name": "15401.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2511
+ },
+ {
+ "file_name": "12005.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2512
+ },
+ {
+ "file_name": "14151.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2513
+ },
+ {
+ "file_name": "19071.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2514
+ },
+ {
+ "file_name": "21554.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2515
+ },
+ {
+ "file_name": "16480.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2516
+ },
+ {
+ "file_name": "15296.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2517
+ },
+ {
+ "file_name": "20848.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2518
+ },
+ {
+ "file_name": "15467.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2519
+ },
+ {
+ "file_name": "15291.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2520
+ },
+ {
+ "file_name": "19148.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2521
+ },
+ {
+ "file_name": "15241.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2522
+ },
+ {
+ "file_name": "22518.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2523
+ },
+ {
+ "file_name": "17944.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2524
+ },
+ {
+ "file_name": "12440.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2525
+ },
+ {
+ "file_name": "13414.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2526
+ },
+ {
+ "file_name": "18294.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2527
+ },
+ {
+ "file_name": "22784.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2528
+ },
+ {
+ "file_name": "16922.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2529
+ },
+ {
+ "file_name": "15878.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2530
+ },
+ {
+ "file_name": "19112.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2531
+ },
+ {
+ "file_name": "18550.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2532
+ },
+ {
+ "file_name": "13450.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2533
+ },
+ {
+ "file_name": "19698.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2534
+ },
+ {
+ "file_name": "16247.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2535
+ },
+ {
+ "file_name": "16755.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2536
+ },
+ {
+ "file_name": "17597.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2537
+ },
+ {
+ "file_name": "18794.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2538
+ },
+ {
+ "file_name": "15645.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2539
+ },
+ {
+ "file_name": "13383.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2540
+ },
+ {
+ "file_name": "15436.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2541
+ },
+ {
+ "file_name": "20121.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2542
+ },
+ {
+ "file_name": "17124.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2543
+ },
+ {
+ "file_name": "13522.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2544
+ },
+ {
+ "file_name": "18225.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2545
+ },
+ {
+ "file_name": "17839.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2546
+ },
+ {
+ "file_name": "18571.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2547
+ },
+ {
+ "file_name": "21848.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2548
+ },
+ {
+ "file_name": "14990.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2549
+ },
+ {
+ "file_name": "19487.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2550
+ },
+ {
+ "file_name": "17194.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2551
+ },
+ {
+ "file_name": "12988.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2552
+ },
+ {
+ "file_name": "19402.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2553
+ },
+ {
+ "file_name": "17223.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2554
+ },
+ {
+ "file_name": "14004.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2555
+ },
+ {
+ "file_name": "23398.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2556
+ },
+ {
+ "file_name": "14333.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2557
+ },
+ {
+ "file_name": "14607.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2558
+ },
+ {
+ "file_name": "22259.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2559
+ },
+ {
+ "file_name": "15890.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2560
+ },
+ {
+ "file_name": "15288.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2561
+ },
+ {
+ "file_name": "21749.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2562
+ },
+ {
+ "file_name": "17552.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2563
+ },
+ {
+ "file_name": "12007.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2564
+ },
+ {
+ "file_name": "19990.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2565
+ },
+ {
+ "file_name": "21890.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2566
+ },
+ {
+ "file_name": "19231.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2567
+ },
+ {
+ "file_name": "23367.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2568
+ },
+ {
+ "file_name": "23340.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2569
+ },
+ {
+ "file_name": "19138.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2570
+ },
+ {
+ "file_name": "12709.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2571
+ },
+ {
+ "file_name": "15189.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2572
+ },
+ {
+ "file_name": "22537.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2573
+ },
+ {
+ "file_name": "19849.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2574
+ },
+ {
+ "file_name": "21984.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2575
+ },
+ {
+ "file_name": "15511.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2576
+ },
+ {
+ "file_name": "20901.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2577
+ },
+ {
+ "file_name": "11747.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2578
+ },
+ {
+ "file_name": "18018.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2579
+ },
+ {
+ "file_name": "17049.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2580
+ },
+ {
+ "file_name": "14921.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2581
+ },
+ {
+ "file_name": "19588.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2582
+ },
+ {
+ "file_name": "14361.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2583
+ },
+ {
+ "file_name": "15864.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2584
+ },
+ {
+ "file_name": "19272.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2585
+ },
+ {
+ "file_name": "21798.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2586
+ },
+ {
+ "file_name": "21796.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2587
+ },
+ {
+ "file_name": "16920.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2588
+ },
+ {
+ "file_name": "17942.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2589
+ },
+ {
+ "file_name": "19420.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2590
+ },
+ {
+ "file_name": "21036.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2591
+ },
+ {
+ "file_name": "16536.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2592
+ },
+ {
+ "file_name": "13057.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2593
+ },
+ {
+ "file_name": "19735.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2594
+ },
+ {
+ "file_name": "20360.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2595
+ },
+ {
+ "file_name": "14312.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2596
+ },
+ {
+ "file_name": "14704.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2597
+ },
+ {
+ "file_name": "14797.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2598
+ },
+ {
+ "file_name": "19439.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2599
+ },
+ {
+ "file_name": "17668.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2600
+ },
+ {
+ "file_name": "17964.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2601
+ },
+ {
+ "file_name": "22782.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2602
+ },
+ {
+ "file_name": "22614.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2603
+ },
+ {
+ "file_name": "15056.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2604
+ },
+ {
+ "file_name": "18723.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2605
+ },
+ {
+ "file_name": "14240.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2606
+ },
+ {
+ "file_name": "19368.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2607
+ },
+ {
+ "file_name": "17955.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2608
+ },
+ {
+ "file_name": "12261.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2609
+ },
+ {
+ "file_name": "18900.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2610
+ },
+ {
+ "file_name": "19847.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2611
+ },
+ {
+ "file_name": "15889.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2612
+ },
+ {
+ "file_name": "22244.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2613
+ },
+ {
+ "file_name": "16789.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2614
+ },
+ {
+ "file_name": "22422.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2615
+ },
+ {
+ "file_name": "17291.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2616
+ },
+ {
+ "file_name": "15758.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2617
+ },
+ {
+ "file_name": "21655.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2618
+ },
+ {
+ "file_name": "13632.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2619
+ },
+ {
+ "file_name": "23036.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2620
+ },
+ {
+ "file_name": "14666.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2621
+ },
+ {
+ "file_name": "17881.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2622
+ },
+ {
+ "file_name": "12859.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2623
+ },
+ {
+ "file_name": "17172.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2624
+ },
+ {
+ "file_name": "18907.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2625
+ },
+ {
+ "file_name": "11983.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2626
+ },
+ {
+ "file_name": "18543.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2627
+ },
+ {
+ "file_name": "19559.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2628
+ },
+ {
+ "file_name": "23368.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2629
+ },
+ {
+ "file_name": "20228.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2630
+ },
+ {
+ "file_name": "20743.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2631
+ },
+ {
+ "file_name": "16359.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2632
+ },
+ {
+ "file_name": "21591.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2633
+ },
+ {
+ "file_name": "18964.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2634
+ },
+ {
+ "file_name": "23105.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2635
+ },
+ {
+ "file_name": "12160.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2636
+ },
+ {
+ "file_name": "13850.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2637
+ },
+ {
+ "file_name": "17614.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2638
+ },
+ {
+ "file_name": "14938.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2639
+ },
+ {
+ "file_name": "14341.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2640
+ },
+ {
+ "file_name": "18940.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2641
+ },
+ {
+ "file_name": "11938.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2642
+ },
+ {
+ "file_name": "15285.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2643
+ },
+ {
+ "file_name": "13116.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2644
+ },
+ {
+ "file_name": "20026.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2645
+ },
+ {
+ "file_name": "18855.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2646
+ },
+ {
+ "file_name": "13196.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2647
+ },
+ {
+ "file_name": "19776.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2648
+ },
+ {
+ "file_name": "16917.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2649
+ },
+ {
+ "file_name": "20013.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2650
+ },
+ {
+ "file_name": "22558.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2651
+ },
+ {
+ "file_name": "20674.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2652
+ },
+ {
+ "file_name": "12150.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2653
+ },
+ {
+ "file_name": "12213.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2654
+ },
+ {
+ "file_name": "20584.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2655
+ },
+ {
+ "file_name": "17962.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2656
+ },
+ {
+ "file_name": "12088.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2657
+ },
+ {
+ "file_name": "23111.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2658
+ },
+ {
+ "file_name": "20187.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2659
+ },
+ {
+ "file_name": "13634.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2660
+ },
+ {
+ "file_name": "14310.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2661
+ },
+ {
+ "file_name": "22642.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2662
+ },
+ {
+ "file_name": "15675.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2663
+ },
+ {
+ "file_name": "12478.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2664
+ },
+ {
+ "file_name": "18934.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2665
+ },
+ {
+ "file_name": "11912.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2666
+ },
+ {
+ "file_name": "21198.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2667
+ },
+ {
+ "file_name": "16021.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2668
+ },
+ {
+ "file_name": "21130.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2669
+ },
+ {
+ "file_name": "14189.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2670
+ },
+ {
+ "file_name": "13089.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2671
+ },
+ {
+ "file_name": "14186.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2672
+ },
+ {
+ "file_name": "14261.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2673
+ },
+ {
+ "file_name": "13527.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2674
+ },
+ {
+ "file_name": "15433.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2675
+ },
+ {
+ "file_name": "12425.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2676
+ },
+ {
+ "file_name": "20170.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2677
+ },
+ {
+ "file_name": "18154.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2678
+ },
+ {
+ "file_name": "17413.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2679
+ },
+ {
+ "file_name": "14485.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2680
+ },
+ {
+ "file_name": "13039.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2681
+ },
+ {
+ "file_name": "13620.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2682
+ },
+ {
+ "file_name": "14284.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2683
+ },
+ {
+ "file_name": "19758.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2684
+ },
+ {
+ "file_name": "18395.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2685
+ },
+ {
+ "file_name": "18930.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2686
+ },
+ {
+ "file_name": "16846.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2687
+ },
+ {
+ "file_name": "13227.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2688
+ },
+ {
+ "file_name": "17697.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2689
+ },
+ {
+ "file_name": "14038.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2690
+ },
+ {
+ "file_name": "18931.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2691
+ },
+ {
+ "file_name": "22364.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2692
+ },
+ {
+ "file_name": "16272.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2693
+ },
+ {
+ "file_name": "19386.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2694
+ },
+ {
+ "file_name": "16807.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2695
+ },
+ {
+ "file_name": "16486.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2696
+ },
+ {
+ "file_name": "17805.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2697
+ },
+ {
+ "file_name": "15831.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2698
+ },
+ {
+ "file_name": "13406.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2699
+ },
+ {
+ "file_name": "22073.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2700
+ },
+ {
+ "file_name": "15332.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2701
+ },
+ {
+ "file_name": "18604.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2702
+ },
+ {
+ "file_name": "21622.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2703
+ },
+ {
+ "file_name": "21734.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2704
+ },
+ {
+ "file_name": "21264.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2705
+ },
+ {
+ "file_name": "15911.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2706
+ },
+ {
+ "file_name": "13289.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2707
+ },
+ {
+ "file_name": "20260.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2708
+ },
+ {
+ "file_name": "17115.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2709
+ },
+ {
+ "file_name": "14279.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2710
+ },
+ {
+ "file_name": "17849.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2711
+ },
+ {
+ "file_name": "16124.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2712
+ },
+ {
+ "file_name": "17561.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2713
+ },
+ {
+ "file_name": "23049.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2714
+ },
+ {
+ "file_name": "19301.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2715
+ },
+ {
+ "file_name": "19572.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2716
+ },
+ {
+ "file_name": "14233.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2717
+ },
+ {
+ "file_name": "12214.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2718
+ },
+ {
+ "file_name": "17424.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2719
+ },
+ {
+ "file_name": "19800.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2720
+ },
+ {
+ "file_name": "13608.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2721
+ },
+ {
+ "file_name": "14548.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2722
+ },
+ {
+ "file_name": "19475.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2723
+ },
+ {
+ "file_name": "12323.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2724
+ },
+ {
+ "file_name": "19744.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2725
+ },
+ {
+ "file_name": "16295.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2726
+ },
+ {
+ "file_name": "17545.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2727
+ },
+ {
+ "file_name": "16901.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2728
+ },
+ {
+ "file_name": "12595.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2729
+ },
+ {
+ "file_name": "21923.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2730
+ },
+ {
+ "file_name": "19759.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2731
+ },
+ {
+ "file_name": "19863.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2732
+ },
+ {
+ "file_name": "12723.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2733
+ },
+ {
+ "file_name": "23032.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2734
+ },
+ {
+ "file_name": "19164.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2735
+ },
+ {
+ "file_name": "19006.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2736
+ },
+ {
+ "file_name": "19750.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2737
+ },
+ {
+ "file_name": "20432.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2738
+ },
+ {
+ "file_name": "19806.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2739
+ },
+ {
+ "file_name": "22110.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2740
+ },
+ {
+ "file_name": "19392.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2741
+ },
+ {
+ "file_name": "13985.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2742
+ },
+ {
+ "file_name": "13837.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2743
+ },
+ {
+ "file_name": "19810.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2744
+ },
+ {
+ "file_name": "20834.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2745
+ },
+ {
+ "file_name": "13847.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2746
+ },
+ {
+ "file_name": "19452.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2747
+ },
+ {
+ "file_name": "18532.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2748
+ },
+ {
+ "file_name": "11933.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2749
+ },
+ {
+ "file_name": "18044.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2750
+ },
+ {
+ "file_name": "18356.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2751
+ },
+ {
+ "file_name": "18350.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2752
+ },
+ {
+ "file_name": "17648.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2753
+ },
+ {
+ "file_name": "16184.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2754
+ },
+ {
+ "file_name": "19169.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2755
+ },
+ {
+ "file_name": "21613.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2756
+ },
+ {
+ "file_name": "17127.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2757
+ },
+ {
+ "file_name": "13501.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2758
+ },
+ {
+ "file_name": "21169.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2759
+ },
+ {
+ "file_name": "19972.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2760
+ },
+ {
+ "file_name": "15536.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2761
+ },
+ {
+ "file_name": "20102.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2762
+ },
+ {
+ "file_name": "20226.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2763
+ },
+ {
+ "file_name": "12079.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2764
+ },
+ {
+ "file_name": "15570.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2765
+ },
+ {
+ "file_name": "14378.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2766
+ },
+ {
+ "file_name": "16523.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2767
+ },
+ {
+ "file_name": "18731.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2768
+ },
+ {
+ "file_name": "12273.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2769
+ },
+ {
+ "file_name": "18578.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2770
+ },
+ {
+ "file_name": "19025.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2771
+ },
+ {
+ "file_name": "13780.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2772
+ },
+ {
+ "file_name": "15153.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2773
+ },
+ {
+ "file_name": "16237.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2774
+ },
+ {
+ "file_name": "16077.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2775
+ },
+ {
+ "file_name": "20417.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2776
+ },
+ {
+ "file_name": "19979.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2777
+ },
+ {
+ "file_name": "16513.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2778
+ },
+ {
+ "file_name": "21279.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2779
+ },
+ {
+ "file_name": "15628.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2780
+ },
+ {
+ "file_name": "17141.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2781
+ },
+ {
+ "file_name": "14784.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2782
+ },
+ {
+ "file_name": "12684.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2783
+ },
+ {
+ "file_name": "13949.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2784
+ },
+ {
+ "file_name": "22856.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2785
+ },
+ {
+ "file_name": "22044.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2786
+ },
+ {
+ "file_name": "19316.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2787
+ },
+ {
+ "file_name": "14664.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2788
+ },
+ {
+ "file_name": "20625.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2789
+ },
+ {
+ "file_name": "14044.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2790
+ },
+ {
+ "file_name": "11820.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2791
+ },
+ {
+ "file_name": "22132.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2792
+ },
+ {
+ "file_name": "20505.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2793
+ },
+ {
+ "file_name": "15144.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2794
+ },
+ {
+ "file_name": "21231.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2795
+ },
+ {
+ "file_name": "20634.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2796
+ },
+ {
+ "file_name": "19513.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2797
+ },
+ {
+ "file_name": "22304.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2798
+ },
+ {
+ "file_name": "11757.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2799
+ },
+ {
+ "file_name": "21868.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2800
+ },
+ {
+ "file_name": "16816.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2801
+ },
+ {
+ "file_name": "12489.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2802
+ },
+ {
+ "file_name": "13604.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2803
+ },
+ {
+ "file_name": "11916.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2804
+ },
+ {
+ "file_name": "19757.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2805
+ },
+ {
+ "file_name": "17536.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2806
+ },
+ {
+ "file_name": "12764.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2807
+ },
+ {
+ "file_name": "13326.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2808
+ },
+ {
+ "file_name": "19911.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2809
+ },
+ {
+ "file_name": "19129.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2810
+ },
+ {
+ "file_name": "22311.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2811
+ },
+ {
+ "file_name": "23423.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2812
+ },
+ {
+ "file_name": "16909.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2813
+ },
+ {
+ "file_name": "12796.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2814
+ },
+ {
+ "file_name": "11749.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2815
+ },
+ {
+ "file_name": "16264.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2816
+ },
+ {
+ "file_name": "12414.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2817
+ },
+ {
+ "file_name": "20512.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2818
+ },
+ {
+ "file_name": "21558.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2819
+ },
+ {
+ "file_name": "22794.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2820
+ },
+ {
+ "file_name": "18878.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2821
+ },
+ {
+ "file_name": "21296.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2822
+ },
+ {
+ "file_name": "16938.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2823
+ },
+ {
+ "file_name": "18387.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2824
+ },
+ {
+ "file_name": "15103.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2825
+ },
+ {
+ "file_name": "19978.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2826
+ },
+ {
+ "file_name": "14192.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2827
+ },
+ {
+ "file_name": "17875.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2828
+ },
+ {
+ "file_name": "20860.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2829
+ },
+ {
+ "file_name": "13408.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2830
+ },
+ {
+ "file_name": "21107.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2831
+ },
+ {
+ "file_name": "14036.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2832
+ },
+ {
+ "file_name": "22590.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2833
+ },
+ {
+ "file_name": "15793.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2834
+ },
+ {
+ "file_name": "20399.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2835
+ },
+ {
+ "file_name": "20431.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2836
+ },
+ {
+ "file_name": "11885.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2837
+ },
+ {
+ "file_name": "18457.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2838
+ },
+ {
+ "file_name": "11866.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2839
+ },
+ {
+ "file_name": "22270.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2840
+ },
+ {
+ "file_name": "21018.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2841
+ },
+ {
+ "file_name": "21547.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2842
+ },
+ {
+ "file_name": "15284.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2843
+ },
+ {
+ "file_name": "16686.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2844
+ },
+ {
+ "file_name": "18081.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2845
+ },
+ {
+ "file_name": "15305.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2846
+ },
+ {
+ "file_name": "16139.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2847
+ },
+ {
+ "file_name": "23068.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2848
+ },
+ {
+ "file_name": "15175.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2849
+ },
+ {
+ "file_name": "21612.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2850
+ },
+ {
+ "file_name": "11934.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2851
+ },
+ {
+ "file_name": "17220.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2852
+ },
+ {
+ "file_name": "11994.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2853
+ },
+ {
+ "file_name": "20317.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2854
+ },
+ {
+ "file_name": "13642.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2855
+ },
+ {
+ "file_name": "14550.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2856
+ },
+ {
+ "file_name": "21401.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2857
+ },
+ {
+ "file_name": "17385.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2858
+ },
+ {
+ "file_name": "20844.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2859
+ },
+ {
+ "file_name": "21054.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2860
+ },
+ {
+ "file_name": "16014.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2861
+ },
+ {
+ "file_name": "13038.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2862
+ },
+ {
+ "file_name": "13638.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2863
+ },
+ {
+ "file_name": "18468.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2864
+ },
+ {
+ "file_name": "20235.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2865
+ },
+ {
+ "file_name": "22252.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2866
+ },
+ {
+ "file_name": "20884.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2867
+ },
+ {
+ "file_name": "14427.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2868
+ },
+ {
+ "file_name": "18496.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2869
+ },
+ {
+ "file_name": "17665.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2870
+ },
+ {
+ "file_name": "15386.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2871
+ },
+ {
+ "file_name": "13773.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2872
+ },
+ {
+ "file_name": "20466.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2873
+ },
+ {
+ "file_name": "13034.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2874
+ },
+ {
+ "file_name": "20065.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2875
+ },
+ {
+ "file_name": "22753.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2876
+ },
+ {
+ "file_name": "23065.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2877
+ },
+ {
+ "file_name": "15503.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2878
+ },
+ {
+ "file_name": "12590.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2879
+ },
+ {
+ "file_name": "16094.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2880
+ },
+ {
+ "file_name": "19992.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2881
+ },
+ {
+ "file_name": "18638.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2882
+ },
+ {
+ "file_name": "12621.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2883
+ },
+ {
+ "file_name": "22272.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2884
+ },
+ {
+ "file_name": "20062.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2885
+ },
+ {
+ "file_name": "21792.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2886
+ },
+ {
+ "file_name": "18332.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2887
+ },
+ {
+ "file_name": "15522.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2888
+ },
+ {
+ "file_name": "19556.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2889
+ },
+ {
+ "file_name": "17987.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2890
+ },
+ {
+ "file_name": "13580.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2891
+ },
+ {
+ "file_name": "17768.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2892
+ },
+ {
+ "file_name": "22583.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2893
+ },
+ {
+ "file_name": "11865.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2894
+ },
+ {
+ "file_name": "12058.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2895
+ },
+ {
+ "file_name": "14476.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2896
+ },
+ {
+ "file_name": "13423.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2897
+ },
+ {
+ "file_name": "20465.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2898
+ },
+ {
+ "file_name": "14631.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2899
+ },
+ {
+ "file_name": "16156.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2900
+ },
+ {
+ "file_name": "12581.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2901
+ },
+ {
+ "file_name": "11969.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2902
+ },
+ {
+ "file_name": "13533.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2903
+ },
+ {
+ "file_name": "19709.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2904
+ },
+ {
+ "file_name": "13513.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2905
+ },
+ {
+ "file_name": "20474.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2906
+ },
+ {
+ "file_name": "21849.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2907
+ },
+ {
+ "file_name": "13880.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2908
+ },
+ {
+ "file_name": "17531.jpg",
+ "height": 512,
+ "width": 512,
+ "id": 2909
+ }
+ ],
+ "annotations": [
+ {
+ "image_id": 0,
+ "bbox": [
+ 157,
+ 167,
+ 160,
+ 174
+ ],
+ "category_id": 10,
+ "id": 1,
+ "area": 27840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1,
+ "bbox": [
+ 67,
+ 293,
+ 272,
+ 124
+ ],
+ "category_id": 19,
+ "id": 2,
+ "area": 33728,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2,
+ "bbox": [
+ 47,
+ 157,
+ 395,
+ 156
+ ],
+ "category_id": 16,
+ "id": 3,
+ "area": 61620,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 3,
+ "bbox": [
+ 279,
+ 223,
+ 49,
+ 22
+ ],
+ "category_id": 16,
+ "id": 4,
+ "area": 1078,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 4,
+ "bbox": [
+ 198,
+ 79,
+ 181,
+ 317
+ ],
+ "category_id": 10,
+ "id": 5,
+ "area": 57377,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 5,
+ "bbox": [
+ 246,
+ 265,
+ 173,
+ 97
+ ],
+ "category_id": 4,
+ "id": 6,
+ "area": 16781,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 6,
+ "bbox": [
+ 59,
+ 141,
+ 397,
+ 99
+ ],
+ "category_id": 19,
+ "id": 7,
+ "area": 39303,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 7,
+ "bbox": [
+ 305,
+ 46,
+ 124,
+ 198
+ ],
+ "category_id": 10,
+ "id": 8,
+ "area": 24552,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 8,
+ "bbox": [
+ 86,
+ 212,
+ 122,
+ 136
+ ],
+ "category_id": 15,
+ "id": 9,
+ "area": 16592,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 8,
+ "bbox": [
+ 244,
+ 177,
+ 143,
+ 162
+ ],
+ "category_id": 15,
+ "id": 10,
+ "area": 23166,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 9,
+ "bbox": [
+ 30,
+ 165,
+ 117,
+ 140
+ ],
+ "category_id": 19,
+ "id": 11,
+ "area": 16380,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 10,
+ "bbox": [
+ 142,
+ 122,
+ 144,
+ 171
+ ],
+ "category_id": 15,
+ "id": 12,
+ "area": 24624,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 10,
+ "bbox": [
+ 312,
+ 258,
+ 144,
+ 173
+ ],
+ "category_id": 15,
+ "id": 13,
+ "area": 24912,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 11,
+ "bbox": [
+ 22,
+ 39,
+ 22,
+ 32
+ ],
+ "category_id": 4,
+ "id": 14,
+ "area": 704,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 11,
+ "bbox": [
+ 36,
+ 177,
+ 27,
+ 31
+ ],
+ "category_id": 4,
+ "id": 15,
+ "area": 837,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 11,
+ "bbox": [
+ 199,
+ 348,
+ 28,
+ 37
+ ],
+ "category_id": 4,
+ "id": 16,
+ "area": 1036,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 11,
+ "bbox": [
+ 242,
+ 76,
+ 27,
+ 36
+ ],
+ "category_id": 4,
+ "id": 17,
+ "area": 972,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 11,
+ "bbox": [
+ 457,
+ 197,
+ 19,
+ 35
+ ],
+ "category_id": 4,
+ "id": 18,
+ "area": 665,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 11,
+ "bbox": [
+ 389,
+ 449,
+ 25,
+ 25
+ ],
+ "category_id": 4,
+ "id": 19,
+ "area": 625,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 12,
+ "bbox": [
+ 3,
+ 312,
+ 62,
+ 41
+ ],
+ "category_id": 16,
+ "id": 20,
+ "area": 2542,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 13,
+ "bbox": [
+ 203,
+ 182,
+ 93,
+ 112
+ ],
+ "category_id": 4,
+ "id": 21,
+ "area": 10416,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 14,
+ "bbox": [
+ 88,
+ 122,
+ 133,
+ 146
+ ],
+ "category_id": 15,
+ "id": 22,
+ "area": 19418,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 14,
+ "bbox": [
+ 97,
+ 298,
+ 129,
+ 144
+ ],
+ "category_id": 15,
+ "id": 23,
+ "area": 18576,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 14,
+ "bbox": [
+ 257,
+ 301,
+ 134,
+ 140
+ ],
+ "category_id": 15,
+ "id": 24,
+ "area": 18760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 15,
+ "bbox": [
+ 146,
+ 16,
+ 31,
+ 16
+ ],
+ "category_id": 4,
+ "id": 25,
+ "area": 496,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 15,
+ "bbox": [
+ 296,
+ 58,
+ 37,
+ 19
+ ],
+ "category_id": 4,
+ "id": 26,
+ "area": 703,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 15,
+ "bbox": [
+ 316,
+ 364,
+ 39,
+ 22
+ ],
+ "category_id": 4,
+ "id": 27,
+ "area": 858,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 15,
+ "bbox": [
+ 8,
+ 483,
+ 38,
+ 19
+ ],
+ "category_id": 4,
+ "id": 28,
+ "area": 722,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 16,
+ "bbox": [
+ 36,
+ 178,
+ 414,
+ 118
+ ],
+ "category_id": 10,
+ "id": 29,
+ "area": 48852,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 17,
+ "bbox": [
+ 85,
+ 80,
+ 222,
+ 274
+ ],
+ "category_id": 10,
+ "id": 30,
+ "area": 60828,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 18,
+ "bbox": [
+ 115,
+ 99,
+ 77,
+ 42
+ ],
+ "category_id": 4,
+ "id": 31,
+ "area": 3234,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 18,
+ "bbox": [
+ 202,
+ 257,
+ 80,
+ 43
+ ],
+ "category_id": 4,
+ "id": 32,
+ "area": 3440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 18,
+ "bbox": [
+ 406,
+ 341,
+ 75,
+ 38
+ ],
+ "category_id": 4,
+ "id": 33,
+ "area": 2850,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 19,
+ "bbox": [
+ 40,
+ 453,
+ 27,
+ 23
+ ],
+ "category_id": 16,
+ "id": 34,
+ "area": 621,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 19,
+ "bbox": [
+ 398,
+ 93,
+ 34,
+ 18
+ ],
+ "category_id": 16,
+ "id": 35,
+ "area": 612,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 20,
+ "bbox": [
+ 49,
+ 193,
+ 112,
+ 147
+ ],
+ "category_id": 15,
+ "id": 36,
+ "area": 16464,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 20,
+ "bbox": [
+ 249,
+ 211,
+ 108,
+ 144
+ ],
+ "category_id": 15,
+ "id": 37,
+ "area": 15552,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 21,
+ "bbox": [
+ 220,
+ 91,
+ 207,
+ 352
+ ],
+ "category_id": 10,
+ "id": 38,
+ "area": 72864,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 22,
+ "bbox": [
+ 142,
+ 49,
+ 16,
+ 15
+ ],
+ "category_id": 4,
+ "id": 39,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 22,
+ "bbox": [
+ 425,
+ 215,
+ 21,
+ 11
+ ],
+ "category_id": 4,
+ "id": 40,
+ "area": 231,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 22,
+ "bbox": [
+ 286,
+ 136,
+ 18,
+ 14
+ ],
+ "category_id": 4,
+ "id": 41,
+ "area": 252,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 22,
+ "bbox": [
+ 56,
+ 186,
+ 25,
+ 15
+ ],
+ "category_id": 4,
+ "id": 42,
+ "area": 375,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 22,
+ "bbox": [
+ 120,
+ 332,
+ 24,
+ 21
+ ],
+ "category_id": 4,
+ "id": 43,
+ "area": 504,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 22,
+ "bbox": [
+ 221,
+ 438,
+ 24,
+ 17
+ ],
+ "category_id": 4,
+ "id": 44,
+ "area": 408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 23,
+ "bbox": [
+ 33,
+ 119,
+ 18,
+ 24
+ ],
+ "category_id": 4,
+ "id": 45,
+ "area": 432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 23,
+ "bbox": [
+ 113,
+ 361,
+ 22,
+ 24
+ ],
+ "category_id": 4,
+ "id": 46,
+ "area": 528,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 23,
+ "bbox": [
+ 49,
+ 484,
+ 18,
+ 26
+ ],
+ "category_id": 4,
+ "id": 47,
+ "area": 468,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 23,
+ "bbox": [
+ 252,
+ 55,
+ 21,
+ 22
+ ],
+ "category_id": 4,
+ "id": 48,
+ "area": 462,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 23,
+ "bbox": [
+ 460,
+ 183,
+ 20,
+ 27
+ ],
+ "category_id": 4,
+ "id": 49,
+ "area": 540,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 23,
+ "bbox": [
+ 437,
+ 302,
+ 20,
+ 28
+ ],
+ "category_id": 4,
+ "id": 50,
+ "area": 560,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 23,
+ "bbox": [
+ 440,
+ 426,
+ 17,
+ 28
+ ],
+ "category_id": 4,
+ "id": 51,
+ "area": 476,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 24,
+ "bbox": [
+ 118,
+ 83,
+ 26,
+ 52
+ ],
+ "category_id": 4,
+ "id": 52,
+ "area": 1352,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 24,
+ "bbox": [
+ 156,
+ 437,
+ 23,
+ 51
+ ],
+ "category_id": 4,
+ "id": 53,
+ "area": 1173,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 24,
+ "bbox": [
+ 357,
+ 172,
+ 21,
+ 52
+ ],
+ "category_id": 4,
+ "id": 54,
+ "area": 1092,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 25,
+ "bbox": [
+ 465,
+ 168,
+ 35,
+ 6
+ ],
+ "category_id": 15,
+ "id": 55,
+ "area": 210,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 25,
+ "bbox": [
+ 413,
+ 172,
+ 40,
+ 9
+ ],
+ "category_id": 15,
+ "id": 56,
+ "area": 360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 26,
+ "bbox": [
+ 188,
+ 146,
+ 192,
+ 261
+ ],
+ "category_id": 16,
+ "id": 57,
+ "area": 50112,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 27,
+ "bbox": [
+ 151,
+ 115,
+ 268,
+ 245
+ ],
+ "category_id": 16,
+ "id": 58,
+ "area": 65660,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 28,
+ "bbox": [
+ 210,
+ 85,
+ 256,
+ 267
+ ],
+ "category_id": 16,
+ "id": 59,
+ "area": 68352,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 29,
+ "bbox": [
+ 263,
+ 171,
+ 245,
+ 305
+ ],
+ "category_id": 10,
+ "id": 60,
+ "area": 74725,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 30,
+ "bbox": [
+ 53,
+ 122,
+ 259,
+ 233
+ ],
+ "category_id": 16,
+ "id": 61,
+ "area": 60347,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 31,
+ "bbox": [
+ 371,
+ 88,
+ 79,
+ 329
+ ],
+ "category_id": 10,
+ "id": 62,
+ "area": 25991,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 32,
+ "bbox": [
+ 111,
+ 243,
+ 337,
+ 84
+ ],
+ "category_id": 19,
+ "id": 63,
+ "area": 28308,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 33,
+ "bbox": [
+ 168,
+ 223,
+ 217,
+ 219
+ ],
+ "category_id": 16,
+ "id": 64,
+ "area": 47523,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 34,
+ "bbox": [
+ 149,
+ 35,
+ 232,
+ 430
+ ],
+ "category_id": 16,
+ "id": 65,
+ "area": 99760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 35,
+ "bbox": [
+ 286,
+ 184,
+ 143,
+ 34
+ ],
+ "category_id": 4,
+ "id": 66,
+ "area": 4862,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 36,
+ "bbox": [
+ 214,
+ 261,
+ 166,
+ 46
+ ],
+ "category_id": 19,
+ "id": 67,
+ "area": 7636,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 37,
+ "bbox": [
+ 148,
+ 316,
+ 85,
+ 77
+ ],
+ "category_id": 15,
+ "id": 68,
+ "area": 6545,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 37,
+ "bbox": [
+ 113,
+ 234,
+ 81,
+ 72
+ ],
+ "category_id": 15,
+ "id": 69,
+ "area": 5832,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 37,
+ "bbox": [
+ 157,
+ 159,
+ 79,
+ 77
+ ],
+ "category_id": 15,
+ "id": 70,
+ "area": 6083,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 37,
+ "bbox": [
+ 327,
+ 191,
+ 73,
+ 70
+ ],
+ "category_id": 15,
+ "id": 71,
+ "area": 5110,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 38,
+ "bbox": [
+ 197,
+ 122,
+ 50,
+ 308
+ ],
+ "category_id": 19,
+ "id": 72,
+ "area": 15400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 39,
+ "bbox": [
+ 218,
+ 116,
+ 83,
+ 38
+ ],
+ "category_id": 16,
+ "id": 73,
+ "area": 3154,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 40,
+ "bbox": [
+ 27,
+ 228,
+ 459,
+ 92
+ ],
+ "category_id": 10,
+ "id": 74,
+ "area": 42228,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 41,
+ "bbox": [
+ 64,
+ 108,
+ 26,
+ 35
+ ],
+ "category_id": 4,
+ "id": 75,
+ "area": 910,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 41,
+ "bbox": [
+ 88,
+ 346,
+ 26,
+ 37
+ ],
+ "category_id": 4,
+ "id": 76,
+ "area": 962,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 41,
+ "bbox": [
+ 446,
+ 265,
+ 29,
+ 46
+ ],
+ "category_id": 4,
+ "id": 77,
+ "area": 1334,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 42,
+ "bbox": [
+ 83,
+ 7,
+ 247,
+ 422
+ ],
+ "category_id": 10,
+ "id": 78,
+ "area": 104234,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 43,
+ "bbox": [
+ 67,
+ 319,
+ 373,
+ 83
+ ],
+ "category_id": 10,
+ "id": 79,
+ "area": 30959,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 44,
+ "bbox": [
+ 78,
+ 204,
+ 296,
+ 116
+ ],
+ "category_id": 10,
+ "id": 80,
+ "area": 34336,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 45,
+ "bbox": [
+ 227,
+ 184,
+ 37,
+ 101
+ ],
+ "category_id": 16,
+ "id": 81,
+ "area": 3737,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 46,
+ "bbox": [
+ 239,
+ 61,
+ 150,
+ 395
+ ],
+ "category_id": 10,
+ "id": 82,
+ "area": 59250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 47,
+ "bbox": [
+ 176,
+ 95,
+ 153,
+ 321
+ ],
+ "category_id": 16,
+ "id": 83,
+ "area": 49113,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 48,
+ "bbox": [
+ 211,
+ 190,
+ 211,
+ 252
+ ],
+ "category_id": 15,
+ "id": 84,
+ "area": 53172,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 49,
+ "bbox": [
+ 184,
+ 245,
+ 208,
+ 41
+ ],
+ "category_id": 19,
+ "id": 85,
+ "area": 8528,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 50,
+ "bbox": [
+ 1,
+ 228,
+ 501,
+ 82
+ ],
+ "category_id": 10,
+ "id": 86,
+ "area": 41082,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 51,
+ "bbox": [
+ 180,
+ 197,
+ 77,
+ 39
+ ],
+ "category_id": 4,
+ "id": 87,
+ "area": 3003,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 51,
+ "bbox": [
+ 298,
+ 335,
+ 71,
+ 40
+ ],
+ "category_id": 4,
+ "id": 88,
+ "area": 2840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 52,
+ "bbox": [
+ 122,
+ 205,
+ 220,
+ 86
+ ],
+ "category_id": 19,
+ "id": 89,
+ "area": 18920,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 53,
+ "bbox": [
+ 117,
+ 192,
+ 172,
+ 191
+ ],
+ "category_id": 16,
+ "id": 90,
+ "area": 32852,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 54,
+ "bbox": [
+ 104,
+ 189,
+ 399,
+ 64
+ ],
+ "category_id": 19,
+ "id": 91,
+ "area": 25536,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 55,
+ "bbox": [
+ 66,
+ 263,
+ 254,
+ 142
+ ],
+ "category_id": 16,
+ "id": 92,
+ "area": 36068,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 56,
+ "bbox": [
+ 186,
+ 70,
+ 48,
+ 322
+ ],
+ "category_id": 19,
+ "id": 93,
+ "area": 15456,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 57,
+ "bbox": [
+ 94,
+ 191,
+ 202,
+ 166
+ ],
+ "category_id": 10,
+ "id": 94,
+ "area": 33532,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 58,
+ "bbox": [
+ 159,
+ 218,
+ 209,
+ 64
+ ],
+ "category_id": 19,
+ "id": 95,
+ "area": 13376,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 59,
+ "bbox": [
+ 90,
+ 86,
+ 214,
+ 329
+ ],
+ "category_id": 16,
+ "id": 96,
+ "area": 70406,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 60,
+ "bbox": [
+ 215,
+ 49,
+ 225,
+ 235
+ ],
+ "category_id": 19,
+ "id": 97,
+ "area": 52875,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 61,
+ "bbox": [
+ 133,
+ 111,
+ 91,
+ 88
+ ],
+ "category_id": 15,
+ "id": 98,
+ "area": 8008,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 61,
+ "bbox": [
+ 81,
+ 231,
+ 92,
+ 90
+ ],
+ "category_id": 15,
+ "id": 99,
+ "area": 8280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 62,
+ "bbox": [
+ 96,
+ 134,
+ 103,
+ 118
+ ],
+ "category_id": 15,
+ "id": 100,
+ "area": 12154,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 62,
+ "bbox": [
+ 261,
+ 139,
+ 103,
+ 118
+ ],
+ "category_id": 15,
+ "id": 101,
+ "area": 12154,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 62,
+ "bbox": [
+ 85,
+ 285,
+ 101,
+ 123
+ ],
+ "category_id": 15,
+ "id": 102,
+ "area": 12423,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 63,
+ "bbox": [
+ 156,
+ 236,
+ 229,
+ 176
+ ],
+ "category_id": 16,
+ "id": 103,
+ "area": 40304,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 64,
+ "bbox": [
+ 202,
+ 84,
+ 129,
+ 134
+ ],
+ "category_id": 15,
+ "id": 104,
+ "area": 17286,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 64,
+ "bbox": [
+ 179,
+ 268,
+ 127,
+ 135
+ ],
+ "category_id": 15,
+ "id": 105,
+ "area": 17145,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 65,
+ "bbox": [
+ 147,
+ 79,
+ 58,
+ 59
+ ],
+ "category_id": 16,
+ "id": 106,
+ "area": 3422,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 66,
+ "bbox": [
+ 144,
+ 88,
+ 81,
+ 32
+ ],
+ "category_id": 16,
+ "id": 107,
+ "area": 2592,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 67,
+ "bbox": [
+ 135,
+ 85,
+ 140,
+ 314
+ ],
+ "category_id": 10,
+ "id": 108,
+ "area": 43960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 68,
+ "bbox": [
+ 214,
+ 191,
+ 50,
+ 152
+ ],
+ "category_id": 19,
+ "id": 109,
+ "area": 7600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 69,
+ "bbox": [
+ 151,
+ 102,
+ 325,
+ 241
+ ],
+ "category_id": 10,
+ "id": 110,
+ "area": 78325,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 70,
+ "bbox": [
+ 177,
+ 189,
+ 244,
+ 205
+ ],
+ "category_id": 10,
+ "id": 111,
+ "area": 50020,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 71,
+ "bbox": [
+ 215,
+ 120,
+ 78,
+ 177
+ ],
+ "category_id": 19,
+ "id": 112,
+ "area": 13806,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 72,
+ "bbox": [
+ 92,
+ 126,
+ 144,
+ 159
+ ],
+ "category_id": 19,
+ "id": 113,
+ "area": 22896,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 73,
+ "bbox": [
+ 76,
+ 169,
+ 325,
+ 217
+ ],
+ "category_id": 16,
+ "id": 114,
+ "area": 70525,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 74,
+ "bbox": [
+ 197,
+ 42,
+ 206,
+ 345
+ ],
+ "category_id": 10,
+ "id": 115,
+ "area": 71070,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 75,
+ "bbox": [
+ 178,
+ 18,
+ 102,
+ 113
+ ],
+ "category_id": 15,
+ "id": 116,
+ "area": 11526,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 75,
+ "bbox": [
+ 212,
+ 170,
+ 102,
+ 117
+ ],
+ "category_id": 15,
+ "id": 117,
+ "area": 11934,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 75,
+ "bbox": [
+ 239,
+ 288,
+ 97,
+ 109
+ ],
+ "category_id": 15,
+ "id": 118,
+ "area": 10573,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 75,
+ "bbox": [
+ 259,
+ 398,
+ 101,
+ 111
+ ],
+ "category_id": 15,
+ "id": 119,
+ "area": 11211,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 75,
+ "bbox": [
+ 384,
+ 201,
+ 110,
+ 136
+ ],
+ "category_id": 15,
+ "id": 120,
+ "area": 14960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 75,
+ "bbox": [
+ 441,
+ 441,
+ 67,
+ 71
+ ],
+ "category_id": 15,
+ "id": 121,
+ "area": 4757,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 76,
+ "bbox": [
+ 234,
+ 111,
+ 64,
+ 59
+ ],
+ "category_id": 4,
+ "id": 122,
+ "area": 3776,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 76,
+ "bbox": [
+ 275,
+ 420,
+ 47,
+ 52
+ ],
+ "category_id": 4,
+ "id": 123,
+ "area": 2444,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 77,
+ "bbox": [
+ 83,
+ 312,
+ 30,
+ 31
+ ],
+ "category_id": 4,
+ "id": 124,
+ "area": 930,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 77,
+ "bbox": [
+ 374,
+ 266,
+ 34,
+ 30
+ ],
+ "category_id": 4,
+ "id": 125,
+ "area": 1020,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 77,
+ "bbox": [
+ 461,
+ 110,
+ 30,
+ 27
+ ],
+ "category_id": 4,
+ "id": 126,
+ "area": 810,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 77,
+ "bbox": [
+ 326,
+ 435,
+ 29,
+ 25
+ ],
+ "category_id": 4,
+ "id": 127,
+ "area": 725,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 78,
+ "bbox": [
+ 40,
+ 113,
+ 141,
+ 141
+ ],
+ "category_id": 15,
+ "id": 128,
+ "area": 19881,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 78,
+ "bbox": [
+ 131,
+ 282,
+ 141,
+ 142
+ ],
+ "category_id": 15,
+ "id": 129,
+ "area": 20022,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 79,
+ "bbox": [
+ 210,
+ 191,
+ 159,
+ 154
+ ],
+ "category_id": 15,
+ "id": 130,
+ "area": 24486,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 80,
+ "bbox": [
+ 65,
+ 107,
+ 147,
+ 174
+ ],
+ "category_id": 15,
+ "id": 131,
+ "area": 25578,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 80,
+ "bbox": [
+ 248,
+ 72,
+ 150,
+ 180
+ ],
+ "category_id": 15,
+ "id": 132,
+ "area": 27000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 80,
+ "bbox": [
+ 99,
+ 304,
+ 153,
+ 179
+ ],
+ "category_id": 15,
+ "id": 133,
+ "area": 27387,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 80,
+ "bbox": [
+ 288,
+ 274,
+ 149,
+ 174
+ ],
+ "category_id": 15,
+ "id": 134,
+ "area": 25926,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 81,
+ "bbox": [
+ 235,
+ 171,
+ 184,
+ 179
+ ],
+ "category_id": 10,
+ "id": 135,
+ "area": 32936,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 82,
+ "bbox": [
+ 22,
+ 94,
+ 17,
+ 10
+ ],
+ "category_id": 4,
+ "id": 136,
+ "area": 170,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 82,
+ "bbox": [
+ 184,
+ 250,
+ 18,
+ 13
+ ],
+ "category_id": 4,
+ "id": 137,
+ "area": 234,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 82,
+ "bbox": [
+ 88,
+ 430,
+ 13,
+ 14
+ ],
+ "category_id": 4,
+ "id": 138,
+ "area": 182,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 82,
+ "bbox": [
+ 269,
+ 410,
+ 18,
+ 12
+ ],
+ "category_id": 4,
+ "id": 139,
+ "area": 216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 82,
+ "bbox": [
+ 481,
+ 383,
+ 18,
+ 10
+ ],
+ "category_id": 4,
+ "id": 140,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 83,
+ "bbox": [
+ 248,
+ 197,
+ 82,
+ 88
+ ],
+ "category_id": 4,
+ "id": 141,
+ "area": 7216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 84,
+ "bbox": [
+ 127,
+ 212,
+ 244,
+ 106
+ ],
+ "category_id": 19,
+ "id": 142,
+ "area": 25864,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 85,
+ "bbox": [
+ 48,
+ 225,
+ 36,
+ 50
+ ],
+ "category_id": 4,
+ "id": 143,
+ "area": 1800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 85,
+ "bbox": [
+ 277,
+ 293,
+ 41,
+ 40
+ ],
+ "category_id": 4,
+ "id": 144,
+ "area": 1640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 85,
+ "bbox": [
+ 404,
+ 154,
+ 60,
+ 32
+ ],
+ "category_id": 4,
+ "id": 145,
+ "area": 1920,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 86,
+ "bbox": [
+ 144,
+ 200,
+ 224,
+ 121
+ ],
+ "category_id": 19,
+ "id": 146,
+ "area": 27104,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 87,
+ "bbox": [
+ 12,
+ 252,
+ 119,
+ 210
+ ],
+ "category_id": 19,
+ "id": 147,
+ "area": 24990,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 88,
+ "bbox": [
+ 24,
+ 83,
+ 424,
+ 244
+ ],
+ "category_id": 10,
+ "id": 148,
+ "area": 103456,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 89,
+ "bbox": [
+ 66,
+ 120,
+ 414,
+ 173
+ ],
+ "category_id": 10,
+ "id": 149,
+ "area": 71622,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 90,
+ "bbox": [
+ 208,
+ 83,
+ 54,
+ 374
+ ],
+ "category_id": 19,
+ "id": 150,
+ "area": 20196,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 91,
+ "bbox": [
+ 60,
+ 78,
+ 123,
+ 119
+ ],
+ "category_id": 10,
+ "id": 151,
+ "area": 14637,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 92,
+ "bbox": [
+ 181,
+ 122,
+ 66,
+ 281
+ ],
+ "category_id": 10,
+ "id": 152,
+ "area": 18546,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 93,
+ "bbox": [
+ 73,
+ 32,
+ 343,
+ 180
+ ],
+ "category_id": 10,
+ "id": 153,
+ "area": 61740,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 94,
+ "bbox": [
+ 263,
+ 205,
+ 78,
+ 112
+ ],
+ "category_id": 15,
+ "id": 154,
+ "area": 8736,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 95,
+ "bbox": [
+ 38,
+ 74,
+ 408,
+ 308
+ ],
+ "category_id": 10,
+ "id": 155,
+ "area": 125664,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 96,
+ "bbox": [
+ 195,
+ 276,
+ 130,
+ 159
+ ],
+ "category_id": 4,
+ "id": 156,
+ "area": 20670,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 97,
+ "bbox": [
+ 166,
+ 179,
+ 194,
+ 132
+ ],
+ "category_id": 16,
+ "id": 157,
+ "area": 25608,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 98,
+ "bbox": [
+ 43,
+ 203,
+ 426,
+ 221
+ ],
+ "category_id": 16,
+ "id": 158,
+ "area": 94146,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 99,
+ "bbox": [
+ 256,
+ 355,
+ 90,
+ 70
+ ],
+ "category_id": 15,
+ "id": 159,
+ "area": 6300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 100,
+ "bbox": [
+ 33,
+ 63,
+ 44,
+ 230
+ ],
+ "category_id": 10,
+ "id": 160,
+ "area": 10120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 101,
+ "bbox": [
+ 310,
+ 51,
+ 102,
+ 421
+ ],
+ "category_id": 10,
+ "id": 161,
+ "area": 42942,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 102,
+ "bbox": [
+ 282,
+ 137,
+ 39,
+ 256
+ ],
+ "category_id": 10,
+ "id": 162,
+ "area": 9984,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 103,
+ "bbox": [
+ 154,
+ 229,
+ 115,
+ 35
+ ],
+ "category_id": 19,
+ "id": 163,
+ "area": 4025,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 104,
+ "bbox": [
+ 87,
+ 237,
+ 366,
+ 17
+ ],
+ "category_id": 10,
+ "id": 164,
+ "area": 6222,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 105,
+ "bbox": [
+ 217,
+ 191,
+ 56,
+ 123
+ ],
+ "category_id": 16,
+ "id": 165,
+ "area": 6888,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 106,
+ "bbox": [
+ 444,
+ 108,
+ 16,
+ 15
+ ],
+ "category_id": 4,
+ "id": 166,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 106,
+ "bbox": [
+ 263,
+ 60,
+ 16,
+ 11
+ ],
+ "category_id": 4,
+ "id": 167,
+ "area": 176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 106,
+ "bbox": [
+ 344,
+ 226,
+ 15,
+ 13
+ ],
+ "category_id": 4,
+ "id": 168,
+ "area": 195,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 106,
+ "bbox": [
+ 17,
+ 442,
+ 17,
+ 11
+ ],
+ "category_id": 4,
+ "id": 169,
+ "area": 187,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 106,
+ "bbox": [
+ 218,
+ 397,
+ 18,
+ 10
+ ],
+ "category_id": 4,
+ "id": 170,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 106,
+ "bbox": [
+ 488,
+ 346,
+ 15,
+ 13
+ ],
+ "category_id": 4,
+ "id": 171,
+ "area": 195,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 106,
+ "bbox": [
+ 451,
+ 465,
+ 14,
+ 11
+ ],
+ "category_id": 4,
+ "id": 172,
+ "area": 154,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 107,
+ "bbox": [
+ 264,
+ 328,
+ 176,
+ 88
+ ],
+ "category_id": 19,
+ "id": 173,
+ "area": 15488,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 108,
+ "bbox": [
+ 370,
+ 19,
+ 121,
+ 138
+ ],
+ "category_id": 15,
+ "id": 174,
+ "area": 16698,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 108,
+ "bbox": [
+ 376,
+ 236,
+ 119,
+ 137
+ ],
+ "category_id": 15,
+ "id": 175,
+ "area": 16303,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 108,
+ "bbox": [
+ 192,
+ 75,
+ 161,
+ 165
+ ],
+ "category_id": 15,
+ "id": 176,
+ "area": 26565,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 108,
+ "bbox": [
+ 197,
+ 268,
+ 155,
+ 167
+ ],
+ "category_id": 15,
+ "id": 177,
+ "area": 25885,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 109,
+ "bbox": [
+ 230,
+ 179,
+ 26,
+ 181
+ ],
+ "category_id": 19,
+ "id": 178,
+ "area": 4706,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 110,
+ "bbox": [
+ 211,
+ 98,
+ 128,
+ 289
+ ],
+ "category_id": 10,
+ "id": 179,
+ "area": 36992,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 111,
+ "bbox": [
+ 35,
+ 147,
+ 182,
+ 207
+ ],
+ "category_id": 15,
+ "id": 180,
+ "area": 37674,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 111,
+ "bbox": [
+ 296,
+ 198,
+ 181,
+ 204
+ ],
+ "category_id": 15,
+ "id": 181,
+ "area": 36924,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 112,
+ "bbox": [
+ 207,
+ 355,
+ 94,
+ 136
+ ],
+ "category_id": 15,
+ "id": 182,
+ "area": 12784,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 113,
+ "bbox": [
+ 233,
+ 68,
+ 51,
+ 31
+ ],
+ "category_id": 4,
+ "id": 183,
+ "area": 1581,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 113,
+ "bbox": [
+ 56,
+ 355,
+ 26,
+ 40
+ ],
+ "category_id": 4,
+ "id": 184,
+ "area": 1040,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 113,
+ "bbox": [
+ 435,
+ 382,
+ 25,
+ 25
+ ],
+ "category_id": 4,
+ "id": 185,
+ "area": 625,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 114,
+ "bbox": [
+ 161,
+ 245,
+ 119,
+ 110
+ ],
+ "category_id": 15,
+ "id": 186,
+ "area": 13090,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 114,
+ "bbox": [
+ 321,
+ 220,
+ 127,
+ 119
+ ],
+ "category_id": 15,
+ "id": 187,
+ "area": 15113,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 115,
+ "bbox": [
+ 79,
+ 289,
+ 19,
+ 51
+ ],
+ "category_id": 4,
+ "id": 188,
+ "area": 969,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 115,
+ "bbox": [
+ 413,
+ 127,
+ 19,
+ 50
+ ],
+ "category_id": 4,
+ "id": 189,
+ "area": 950,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 116,
+ "bbox": [
+ 234,
+ 300,
+ 66,
+ 14
+ ],
+ "category_id": 15,
+ "id": 190,
+ "area": 924,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 116,
+ "bbox": [
+ 144,
+ 287,
+ 76,
+ 15
+ ],
+ "category_id": 15,
+ "id": 191,
+ "area": 1140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 117,
+ "bbox": [
+ 204,
+ 55,
+ 212,
+ 305
+ ],
+ "category_id": 10,
+ "id": 192,
+ "area": 64660,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 118,
+ "bbox": [
+ 134,
+ 106,
+ 15,
+ 9
+ ],
+ "category_id": 4,
+ "id": 193,
+ "area": 135,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 118,
+ "bbox": [
+ 10,
+ 217,
+ 14,
+ 8
+ ],
+ "category_id": 4,
+ "id": 194,
+ "area": 112,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 118,
+ "bbox": [
+ 120,
+ 371,
+ 15,
+ 9
+ ],
+ "category_id": 4,
+ "id": 195,
+ "area": 135,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 118,
+ "bbox": [
+ 246,
+ 309,
+ 17,
+ 9
+ ],
+ "category_id": 4,
+ "id": 196,
+ "area": 153,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 118,
+ "bbox": [
+ 401,
+ 238,
+ 15,
+ 10
+ ],
+ "category_id": 4,
+ "id": 197,
+ "area": 150,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 118,
+ "bbox": [
+ 414,
+ 312,
+ 18,
+ 9
+ ],
+ "category_id": 4,
+ "id": 198,
+ "area": 162,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 118,
+ "bbox": [
+ 487,
+ 389,
+ 15,
+ 8
+ ],
+ "category_id": 4,
+ "id": 199,
+ "area": 120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 119,
+ "bbox": [
+ 130,
+ 154,
+ 249,
+ 190
+ ],
+ "category_id": 10,
+ "id": 200,
+ "area": 47310,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 120,
+ "bbox": [
+ 60,
+ 134,
+ 364,
+ 205
+ ],
+ "category_id": 10,
+ "id": 201,
+ "area": 74620,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 121,
+ "bbox": [
+ 237,
+ 151,
+ 138,
+ 177
+ ],
+ "category_id": 19,
+ "id": 202,
+ "area": 24426,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 122,
+ "bbox": [
+ 126,
+ 124,
+ 296,
+ 290
+ ],
+ "category_id": 10,
+ "id": 203,
+ "area": 85840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 123,
+ "bbox": [
+ 127,
+ 175,
+ 260,
+ 201
+ ],
+ "category_id": 16,
+ "id": 204,
+ "area": 52260,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 124,
+ "bbox": [
+ 183,
+ 120,
+ 103,
+ 228
+ ],
+ "category_id": 16,
+ "id": 205,
+ "area": 23484,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 125,
+ "bbox": [
+ 120,
+ 204,
+ 276,
+ 66
+ ],
+ "category_id": 19,
+ "id": 206,
+ "area": 18216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 126,
+ "bbox": [
+ 208,
+ 42,
+ 28,
+ 14
+ ],
+ "category_id": 4,
+ "id": 207,
+ "area": 392,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 126,
+ "bbox": [
+ 212,
+ 165,
+ 29,
+ 15
+ ],
+ "category_id": 4,
+ "id": 208,
+ "area": 435,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 126,
+ "bbox": [
+ 222,
+ 304,
+ 23,
+ 11
+ ],
+ "category_id": 4,
+ "id": 209,
+ "area": 253,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 126,
+ "bbox": [
+ 270,
+ 423,
+ 30,
+ 15
+ ],
+ "category_id": 4,
+ "id": 210,
+ "area": 450,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 127,
+ "bbox": [
+ 115,
+ 150,
+ 364,
+ 261
+ ],
+ "category_id": 10,
+ "id": 211,
+ "area": 95004,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 128,
+ "bbox": [
+ 184,
+ 88,
+ 118,
+ 304
+ ],
+ "category_id": 19,
+ "id": 212,
+ "area": 35872,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 129,
+ "bbox": [
+ 66,
+ 205,
+ 403,
+ 60
+ ],
+ "category_id": 10,
+ "id": 213,
+ "area": 24180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 130,
+ "bbox": [
+ 72,
+ 112,
+ 164,
+ 151
+ ],
+ "category_id": 15,
+ "id": 214,
+ "area": 24764,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 130,
+ "bbox": [
+ 293,
+ 105,
+ 164,
+ 151
+ ],
+ "category_id": 15,
+ "id": 215,
+ "area": 24764,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 130,
+ "bbox": [
+ 61,
+ 325,
+ 124,
+ 110
+ ],
+ "category_id": 15,
+ "id": 216,
+ "area": 13640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 130,
+ "bbox": [
+ 332,
+ 376,
+ 126,
+ 98
+ ],
+ "category_id": 15,
+ "id": 217,
+ "area": 12348,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 131,
+ "bbox": [
+ 295,
+ 95,
+ 97,
+ 90
+ ],
+ "category_id": 15,
+ "id": 218,
+ "area": 8730,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 131,
+ "bbox": [
+ 246,
+ 214,
+ 98,
+ 93
+ ],
+ "category_id": 15,
+ "id": 219,
+ "area": 9114,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 131,
+ "bbox": [
+ 188,
+ 325,
+ 77,
+ 85
+ ],
+ "category_id": 15,
+ "id": 220,
+ "area": 6545,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 132,
+ "bbox": [
+ 183,
+ 33,
+ 238,
+ 455
+ ],
+ "category_id": 10,
+ "id": 221,
+ "area": 108290,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 133,
+ "bbox": [
+ 140,
+ 167,
+ 268,
+ 217
+ ],
+ "category_id": 10,
+ "id": 222,
+ "area": 58156,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 134,
+ "bbox": [
+ 17,
+ 250,
+ 465,
+ 26
+ ],
+ "category_id": 10,
+ "id": 223,
+ "area": 12090,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 135,
+ "bbox": [
+ 218,
+ 214,
+ 48,
+ 118
+ ],
+ "category_id": 4,
+ "id": 224,
+ "area": 5664,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 136,
+ "bbox": [
+ 218,
+ 84,
+ 57,
+ 44
+ ],
+ "category_id": 4,
+ "id": 225,
+ "area": 2508,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 136,
+ "bbox": [
+ 208,
+ 439,
+ 57,
+ 40
+ ],
+ "category_id": 4,
+ "id": 226,
+ "area": 2280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 137,
+ "bbox": [
+ 236,
+ 61,
+ 17,
+ 31
+ ],
+ "category_id": 4,
+ "id": 227,
+ "area": 527,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 137,
+ "bbox": [
+ 51,
+ 150,
+ 18,
+ 34
+ ],
+ "category_id": 4,
+ "id": 228,
+ "area": 612,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 137,
+ "bbox": [
+ 61,
+ 446,
+ 15,
+ 26
+ ],
+ "category_id": 4,
+ "id": 229,
+ "area": 390,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 137,
+ "bbox": [
+ 194,
+ 380,
+ 17,
+ 33
+ ],
+ "category_id": 4,
+ "id": 230,
+ "area": 561,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 137,
+ "bbox": [
+ 330,
+ 312,
+ 20,
+ 32
+ ],
+ "category_id": 4,
+ "id": 231,
+ "area": 640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 138,
+ "bbox": [
+ 75,
+ 350,
+ 42,
+ 58
+ ],
+ "category_id": 16,
+ "id": 232,
+ "area": 2436,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 139,
+ "bbox": [
+ 118,
+ 147,
+ 48,
+ 50
+ ],
+ "category_id": 4,
+ "id": 233,
+ "area": 2400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 139,
+ "bbox": [
+ 424,
+ 332,
+ 66,
+ 57
+ ],
+ "category_id": 4,
+ "id": 234,
+ "area": 3762,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 140,
+ "bbox": [
+ 237,
+ 104,
+ 143,
+ 161
+ ],
+ "category_id": 10,
+ "id": 235,
+ "area": 23023,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 141,
+ "bbox": [
+ 210,
+ 112,
+ 126,
+ 122
+ ],
+ "category_id": 15,
+ "id": 236,
+ "area": 15372,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 141,
+ "bbox": [
+ 259,
+ 276,
+ 122,
+ 121
+ ],
+ "category_id": 15,
+ "id": 237,
+ "area": 14762,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 142,
+ "bbox": [
+ 144,
+ 158,
+ 224,
+ 179
+ ],
+ "category_id": 10,
+ "id": 238,
+ "area": 40096,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 143,
+ "bbox": [
+ 160,
+ 158,
+ 179,
+ 202
+ ],
+ "category_id": 10,
+ "id": 239,
+ "area": 36158,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 144,
+ "bbox": [
+ 84,
+ 279,
+ 75,
+ 39
+ ],
+ "category_id": 4,
+ "id": 240,
+ "area": 2925,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 144,
+ "bbox": [
+ 386,
+ 154,
+ 80,
+ 46
+ ],
+ "category_id": 4,
+ "id": 241,
+ "area": 3680,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 145,
+ "bbox": [
+ 74,
+ 140,
+ 308,
+ 231
+ ],
+ "category_id": 10,
+ "id": 242,
+ "area": 71148,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 146,
+ "bbox": [
+ 139,
+ 63,
+ 50,
+ 31
+ ],
+ "category_id": 4,
+ "id": 243,
+ "area": 1550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 146,
+ "bbox": [
+ 81,
+ 232,
+ 46,
+ 31
+ ],
+ "category_id": 4,
+ "id": 244,
+ "area": 1426,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 146,
+ "bbox": [
+ 61,
+ 376,
+ 45,
+ 31
+ ],
+ "category_id": 4,
+ "id": 245,
+ "area": 1395,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 146,
+ "bbox": [
+ 286,
+ 420,
+ 49,
+ 36
+ ],
+ "category_id": 4,
+ "id": 246,
+ "area": 1764,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 146,
+ "bbox": [
+ 389,
+ 76,
+ 53,
+ 18
+ ],
+ "category_id": 4,
+ "id": 247,
+ "area": 954,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 147,
+ "bbox": [
+ 45,
+ 124,
+ 171,
+ 175
+ ],
+ "category_id": 15,
+ "id": 248,
+ "area": 29925,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 147,
+ "bbox": [
+ 288,
+ 167,
+ 172,
+ 187
+ ],
+ "category_id": 15,
+ "id": 249,
+ "area": 32164,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 148,
+ "bbox": [
+ 265,
+ 105,
+ 193,
+ 293
+ ],
+ "category_id": 10,
+ "id": 250,
+ "area": 56549,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 149,
+ "bbox": [
+ 119,
+ 227,
+ 355,
+ 75
+ ],
+ "category_id": 10,
+ "id": 251,
+ "area": 26625,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 150,
+ "bbox": [
+ 50,
+ 229,
+ 422,
+ 47
+ ],
+ "category_id": 10,
+ "id": 252,
+ "area": 19834,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 151,
+ "bbox": [
+ 316,
+ 55,
+ 62,
+ 49
+ ],
+ "category_id": 4,
+ "id": 253,
+ "area": 3038,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 151,
+ "bbox": [
+ 172,
+ 376,
+ 79,
+ 53
+ ],
+ "category_id": 4,
+ "id": 254,
+ "area": 4187,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 152,
+ "bbox": [
+ 51,
+ 67,
+ 127,
+ 59
+ ],
+ "category_id": 10,
+ "id": 255,
+ "area": 7493,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 153,
+ "bbox": [
+ 108,
+ 32,
+ 130,
+ 171
+ ],
+ "category_id": 15,
+ "id": 256,
+ "area": 22230,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 153,
+ "bbox": [
+ 259,
+ 98,
+ 147,
+ 198
+ ],
+ "category_id": 15,
+ "id": 257,
+ "area": 29106,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 153,
+ "bbox": [
+ 13,
+ 160,
+ 134,
+ 178
+ ],
+ "category_id": 15,
+ "id": 258,
+ "area": 23852,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 153,
+ "bbox": [
+ 255,
+ 290,
+ 126,
+ 168
+ ],
+ "category_id": 15,
+ "id": 259,
+ "area": 21168,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 154,
+ "bbox": [
+ 39,
+ 94,
+ 43,
+ 40
+ ],
+ "category_id": 4,
+ "id": 260,
+ "area": 1720,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 154,
+ "bbox": [
+ 140,
+ 399,
+ 44,
+ 36
+ ],
+ "category_id": 4,
+ "id": 261,
+ "area": 1584,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 154,
+ "bbox": [
+ 400,
+ 428,
+ 49,
+ 42
+ ],
+ "category_id": 4,
+ "id": 262,
+ "area": 2058,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 155,
+ "bbox": [
+ 186,
+ 190,
+ 94,
+ 130
+ ],
+ "category_id": 15,
+ "id": 263,
+ "area": 12220,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 156,
+ "bbox": [
+ 88,
+ 129,
+ 379,
+ 215
+ ],
+ "category_id": 19,
+ "id": 264,
+ "area": 81485,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 157,
+ "bbox": [
+ 228,
+ 94,
+ 28,
+ 51
+ ],
+ "category_id": 4,
+ "id": 265,
+ "area": 1428,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 157,
+ "bbox": [
+ 357,
+ 360,
+ 21,
+ 46
+ ],
+ "category_id": 4,
+ "id": 266,
+ "area": 966,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 158,
+ "bbox": [
+ 156,
+ 226,
+ 263,
+ 110
+ ],
+ "category_id": 16,
+ "id": 267,
+ "area": 28930,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 159,
+ "bbox": [
+ 189,
+ 177,
+ 164,
+ 171
+ ],
+ "category_id": 16,
+ "id": 268,
+ "area": 28044,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 160,
+ "bbox": [
+ 115,
+ 58,
+ 299,
+ 395
+ ],
+ "category_id": 10,
+ "id": 269,
+ "area": 118105,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 161,
+ "bbox": [
+ 144,
+ 225,
+ 251,
+ 184
+ ],
+ "category_id": 16,
+ "id": 270,
+ "area": 46184,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 162,
+ "bbox": [
+ 177,
+ 117,
+ 208,
+ 281
+ ],
+ "category_id": 16,
+ "id": 271,
+ "area": 58448,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 163,
+ "bbox": [
+ 183,
+ 267,
+ 68,
+ 83
+ ],
+ "category_id": 4,
+ "id": 272,
+ "area": 5644,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 164,
+ "bbox": [
+ 211,
+ 233,
+ 96,
+ 92
+ ],
+ "category_id": 4,
+ "id": 273,
+ "area": 8832,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 165,
+ "bbox": [
+ 42,
+ 254,
+ 66,
+ 47
+ ],
+ "category_id": 4,
+ "id": 274,
+ "area": 3102,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 165,
+ "bbox": [
+ 414,
+ 126,
+ 65,
+ 25
+ ],
+ "category_id": 4,
+ "id": 275,
+ "area": 1625,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 166,
+ "bbox": [
+ 193,
+ 101,
+ 55,
+ 274
+ ],
+ "category_id": 19,
+ "id": 276,
+ "area": 15070,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 167,
+ "bbox": [
+ 48,
+ 80,
+ 19,
+ 9
+ ],
+ "category_id": 4,
+ "id": 277,
+ "area": 171,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 167,
+ "bbox": [
+ 266,
+ 55,
+ 19,
+ 13
+ ],
+ "category_id": 4,
+ "id": 278,
+ "area": 247,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 167,
+ "bbox": [
+ 474,
+ 43,
+ 20,
+ 13
+ ],
+ "category_id": 4,
+ "id": 279,
+ "area": 260,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 167,
+ "bbox": [
+ 131,
+ 208,
+ 14,
+ 10
+ ],
+ "category_id": 4,
+ "id": 280,
+ "area": 140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 167,
+ "bbox": [
+ 337,
+ 321,
+ 23,
+ 10
+ ],
+ "category_id": 4,
+ "id": 281,
+ "area": 230,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 167,
+ "bbox": [
+ 10,
+ 364,
+ 18,
+ 12
+ ],
+ "category_id": 4,
+ "id": 282,
+ "area": 216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 168,
+ "bbox": [
+ 428,
+ 60,
+ 40,
+ 176
+ ],
+ "category_id": 10,
+ "id": 283,
+ "area": 7040,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 169,
+ "bbox": [
+ 244,
+ 124,
+ 142,
+ 60
+ ],
+ "category_id": 15,
+ "id": 284,
+ "area": 8520,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 169,
+ "bbox": [
+ 197,
+ 325,
+ 170,
+ 74
+ ],
+ "category_id": 15,
+ "id": 285,
+ "area": 12580,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 170,
+ "bbox": [
+ 57,
+ 250,
+ 383,
+ 58
+ ],
+ "category_id": 16,
+ "id": 286,
+ "area": 22214,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 171,
+ "bbox": [
+ 195,
+ 133,
+ 95,
+ 270
+ ],
+ "category_id": 19,
+ "id": 287,
+ "area": 25650,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 172,
+ "bbox": [
+ 144,
+ 55,
+ 36,
+ 15
+ ],
+ "category_id": 16,
+ "id": 288,
+ "area": 540,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 172,
+ "bbox": [
+ 396,
+ 141,
+ 36,
+ 16
+ ],
+ "category_id": 16,
+ "id": 289,
+ "area": 576,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 173,
+ "bbox": [
+ 78,
+ 238,
+ 343,
+ 90
+ ],
+ "category_id": 19,
+ "id": 290,
+ "area": 30870,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 174,
+ "bbox": [
+ 119,
+ 185,
+ 151,
+ 280
+ ],
+ "category_id": 16,
+ "id": 291,
+ "area": 42280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 175,
+ "bbox": [
+ 231,
+ 206,
+ 51,
+ 152
+ ],
+ "category_id": 19,
+ "id": 292,
+ "area": 7752,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 176,
+ "bbox": [
+ 328,
+ 85,
+ 168,
+ 30
+ ],
+ "category_id": 10,
+ "id": 293,
+ "area": 5040,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 177,
+ "bbox": [
+ 54,
+ 192,
+ 82,
+ 82
+ ],
+ "category_id": 15,
+ "id": 294,
+ "area": 6724,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 178,
+ "bbox": [
+ 231,
+ 126,
+ 124,
+ 122
+ ],
+ "category_id": 15,
+ "id": 295,
+ "area": 15128,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 178,
+ "bbox": [
+ 224,
+ 307,
+ 124,
+ 117
+ ],
+ "category_id": 15,
+ "id": 296,
+ "area": 14508,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 179,
+ "bbox": [
+ 317,
+ 72,
+ 20,
+ 29
+ ],
+ "category_id": 4,
+ "id": 297,
+ "area": 580,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 179,
+ "bbox": [
+ 39,
+ 336,
+ 16,
+ 13
+ ],
+ "category_id": 4,
+ "id": 298,
+ "area": 208,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 179,
+ "bbox": [
+ 236,
+ 251,
+ 14,
+ 17
+ ],
+ "category_id": 4,
+ "id": 299,
+ "area": 238,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 179,
+ "bbox": [
+ 428,
+ 295,
+ 27,
+ 17
+ ],
+ "category_id": 4,
+ "id": 300,
+ "area": 459,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 179,
+ "bbox": [
+ 136,
+ 446,
+ 25,
+ 15
+ ],
+ "category_id": 4,
+ "id": 301,
+ "area": 375,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 180,
+ "bbox": [
+ 49,
+ 284,
+ 17,
+ 9
+ ],
+ "category_id": 4,
+ "id": 302,
+ "area": 153,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 180,
+ "bbox": [
+ 423,
+ 411,
+ 17,
+ 10
+ ],
+ "category_id": 4,
+ "id": 303,
+ "area": 170,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 180,
+ "bbox": [
+ 224,
+ 453,
+ 17,
+ 11
+ ],
+ "category_id": 4,
+ "id": 304,
+ "area": 187,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 181,
+ "bbox": [
+ 241,
+ 40,
+ 19,
+ 9
+ ],
+ "category_id": 4,
+ "id": 305,
+ "area": 171,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 181,
+ "bbox": [
+ 10,
+ 32,
+ 17,
+ 10
+ ],
+ "category_id": 4,
+ "id": 306,
+ "area": 170,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 181,
+ "bbox": [
+ 202,
+ 211,
+ 16,
+ 12
+ ],
+ "category_id": 4,
+ "id": 307,
+ "area": 192,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 181,
+ "bbox": [
+ 433,
+ 168,
+ 15,
+ 12
+ ],
+ "category_id": 4,
+ "id": 308,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 181,
+ "bbox": [
+ 23,
+ 368,
+ 16,
+ 10
+ ],
+ "category_id": 4,
+ "id": 309,
+ "area": 160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 181,
+ "bbox": [
+ 416,
+ 334,
+ 19,
+ 12
+ ],
+ "category_id": 4,
+ "id": 310,
+ "area": 228,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 182,
+ "bbox": [
+ 343,
+ 270,
+ 105,
+ 110
+ ],
+ "category_id": 19,
+ "id": 311,
+ "area": 11550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 183,
+ "bbox": [
+ 111,
+ 149,
+ 255,
+ 258
+ ],
+ "category_id": 10,
+ "id": 312,
+ "area": 65790,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 184,
+ "bbox": [
+ 88,
+ 70,
+ 12,
+ 20
+ ],
+ "category_id": 4,
+ "id": 313,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 184,
+ "bbox": [
+ 268,
+ 228,
+ 14,
+ 23
+ ],
+ "category_id": 4,
+ "id": 314,
+ "area": 322,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 184,
+ "bbox": [
+ 439,
+ 378,
+ 13,
+ 23
+ ],
+ "category_id": 4,
+ "id": 315,
+ "area": 299,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 185,
+ "bbox": [
+ 31,
+ 344,
+ 27,
+ 52
+ ],
+ "category_id": 4,
+ "id": 316,
+ "area": 1404,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 185,
+ "bbox": [
+ 435,
+ 346,
+ 26,
+ 48
+ ],
+ "category_id": 4,
+ "id": 317,
+ "area": 1248,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 185,
+ "bbox": [
+ 407,
+ 63,
+ 26,
+ 35
+ ],
+ "category_id": 4,
+ "id": 318,
+ "area": 910,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 186,
+ "bbox": [
+ 262,
+ 173,
+ 159,
+ 259
+ ],
+ "category_id": 16,
+ "id": 319,
+ "area": 41181,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 187,
+ "bbox": [
+ 89,
+ 190,
+ 350,
+ 95
+ ],
+ "category_id": 16,
+ "id": 320,
+ "area": 33250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 188,
+ "bbox": [
+ 162,
+ 395,
+ 78,
+ 31
+ ],
+ "category_id": 15,
+ "id": 321,
+ "area": 2418,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 188,
+ "bbox": [
+ 261,
+ 300,
+ 85,
+ 36
+ ],
+ "category_id": 15,
+ "id": 322,
+ "area": 3060,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 188,
+ "bbox": [
+ 171,
+ 205,
+ 85,
+ 73
+ ],
+ "category_id": 15,
+ "id": 323,
+ "area": 6205,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 188,
+ "bbox": [
+ 277,
+ 203,
+ 86,
+ 74
+ ],
+ "category_id": 15,
+ "id": 324,
+ "area": 6364,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 189,
+ "bbox": [
+ 81,
+ 238,
+ 352,
+ 152
+ ],
+ "category_id": 10,
+ "id": 325,
+ "area": 53504,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 190,
+ "bbox": [
+ 316,
+ 152,
+ 127,
+ 243
+ ],
+ "category_id": 10,
+ "id": 326,
+ "area": 30861,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 191,
+ "bbox": [
+ 222,
+ 144,
+ 213,
+ 262
+ ],
+ "category_id": 10,
+ "id": 327,
+ "area": 55806,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 192,
+ "bbox": [
+ 108,
+ 227,
+ 82,
+ 90
+ ],
+ "category_id": 15,
+ "id": 328,
+ "area": 7380,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 192,
+ "bbox": [
+ 243,
+ 339,
+ 77,
+ 89
+ ],
+ "category_id": 15,
+ "id": 329,
+ "area": 6853,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 193,
+ "bbox": [
+ 246,
+ 244,
+ 157,
+ 75
+ ],
+ "category_id": 4,
+ "id": 330,
+ "area": 11775,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 194,
+ "bbox": [
+ 232,
+ 232,
+ 76,
+ 114
+ ],
+ "category_id": 4,
+ "id": 331,
+ "area": 8664,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 195,
+ "bbox": [
+ 69,
+ 245,
+ 28,
+ 16
+ ],
+ "category_id": 4,
+ "id": 332,
+ "area": 448,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 195,
+ "bbox": [
+ 494,
+ 276,
+ 15,
+ 11
+ ],
+ "category_id": 4,
+ "id": 333,
+ "area": 165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 195,
+ "bbox": [
+ 344,
+ 204,
+ 16,
+ 10
+ ],
+ "category_id": 4,
+ "id": 334,
+ "area": 160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 195,
+ "bbox": [
+ 3,
+ 320,
+ 25,
+ 17
+ ],
+ "category_id": 4,
+ "id": 335,
+ "area": 425,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 195,
+ "bbox": [
+ 178,
+ 404,
+ 25,
+ 20
+ ],
+ "category_id": 4,
+ "id": 336,
+ "area": 500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 195,
+ "bbox": [
+ 257,
+ 368,
+ 23,
+ 17
+ ],
+ "category_id": 4,
+ "id": 337,
+ "area": 391,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 195,
+ "bbox": [
+ 380,
+ 377,
+ 18,
+ 9
+ ],
+ "category_id": 4,
+ "id": 338,
+ "area": 162,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 195,
+ "bbox": [
+ 480,
+ 433,
+ 29,
+ 15
+ ],
+ "category_id": 4,
+ "id": 339,
+ "area": 435,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 196,
+ "bbox": [
+ 39,
+ 277,
+ 14,
+ 9
+ ],
+ "category_id": 16,
+ "id": 340,
+ "area": 126,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 196,
+ "bbox": [
+ 346,
+ 87,
+ 15,
+ 9
+ ],
+ "category_id": 16,
+ "id": 341,
+ "area": 135,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 197,
+ "bbox": [
+ 372,
+ 90,
+ 7,
+ 17
+ ],
+ "category_id": 4,
+ "id": 342,
+ "area": 119,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 197,
+ "bbox": [
+ 298,
+ 108,
+ 8,
+ 19
+ ],
+ "category_id": 4,
+ "id": 343,
+ "area": 152,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 197,
+ "bbox": [
+ 42,
+ 282,
+ 6,
+ 10
+ ],
+ "category_id": 4,
+ "id": 344,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 197,
+ "bbox": [
+ 72,
+ 258,
+ 6,
+ 12
+ ],
+ "category_id": 4,
+ "id": 345,
+ "area": 72,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 197,
+ "bbox": [
+ 213,
+ 356,
+ 9,
+ 12
+ ],
+ "category_id": 4,
+ "id": 346,
+ "area": 108,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 197,
+ "bbox": [
+ 152,
+ 301,
+ 6,
+ 9
+ ],
+ "category_id": 4,
+ "id": 347,
+ "area": 54,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 197,
+ "bbox": [
+ 251,
+ 421,
+ 6,
+ 11
+ ],
+ "category_id": 4,
+ "id": 348,
+ "area": 66,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 197,
+ "bbox": [
+ 328,
+ 422,
+ 5,
+ 9
+ ],
+ "category_id": 4,
+ "id": 349,
+ "area": 45,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 197,
+ "bbox": [
+ 484,
+ 56,
+ 6,
+ 7
+ ],
+ "category_id": 4,
+ "id": 350,
+ "area": 42,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 198,
+ "bbox": [
+ 20,
+ 184,
+ 20,
+ 47
+ ],
+ "category_id": 4,
+ "id": 351,
+ "area": 940,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 198,
+ "bbox": [
+ 304,
+ 439,
+ 21,
+ 42
+ ],
+ "category_id": 4,
+ "id": 352,
+ "area": 882,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 198,
+ "bbox": [
+ 477,
+ 100,
+ 19,
+ 45
+ ],
+ "category_id": 4,
+ "id": 353,
+ "area": 855,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 199,
+ "bbox": [
+ 235,
+ 219,
+ 92,
+ 69
+ ],
+ "category_id": 4,
+ "id": 354,
+ "area": 6348,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 200,
+ "bbox": [
+ 254,
+ 217,
+ 99,
+ 79
+ ],
+ "category_id": 15,
+ "id": 355,
+ "area": 7821,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 200,
+ "bbox": [
+ 256,
+ 309,
+ 99,
+ 75
+ ],
+ "category_id": 15,
+ "id": 356,
+ "area": 7425,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 201,
+ "bbox": [
+ 199,
+ 61,
+ 237,
+ 357
+ ],
+ "category_id": 10,
+ "id": 357,
+ "area": 84609,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 202,
+ "bbox": [
+ 69,
+ 191,
+ 24,
+ 14
+ ],
+ "category_id": 4,
+ "id": 358,
+ "area": 336,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 202,
+ "bbox": [
+ 165,
+ 91,
+ 30,
+ 17
+ ],
+ "category_id": 4,
+ "id": 359,
+ "area": 510,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 202,
+ "bbox": [
+ 346,
+ 107,
+ 20,
+ 10
+ ],
+ "category_id": 4,
+ "id": 360,
+ "area": 200,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 202,
+ "bbox": [
+ 338,
+ 280,
+ 26,
+ 14
+ ],
+ "category_id": 4,
+ "id": 361,
+ "area": 364,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 202,
+ "bbox": [
+ 383,
+ 394,
+ 34,
+ 17
+ ],
+ "category_id": 4,
+ "id": 362,
+ "area": 578,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 202,
+ "bbox": [
+ 310,
+ 464,
+ 26,
+ 12
+ ],
+ "category_id": 4,
+ "id": 363,
+ "area": 312,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 203,
+ "bbox": [
+ 169,
+ 99,
+ 198,
+ 285
+ ],
+ "category_id": 16,
+ "id": 364,
+ "area": 56430,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 204,
+ "bbox": [
+ 195,
+ 135,
+ 206,
+ 232
+ ],
+ "category_id": 16,
+ "id": 365,
+ "area": 47792,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 205,
+ "bbox": [
+ 133,
+ 39,
+ 14,
+ 9
+ ],
+ "category_id": 4,
+ "id": 366,
+ "area": 126,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 205,
+ "bbox": [
+ 481,
+ 110,
+ 17,
+ 15
+ ],
+ "category_id": 4,
+ "id": 367,
+ "area": 255,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 205,
+ "bbox": [
+ 311,
+ 201,
+ 17,
+ 11
+ ],
+ "category_id": 4,
+ "id": 368,
+ "area": 187,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 205,
+ "bbox": [
+ 166,
+ 335,
+ 21,
+ 14
+ ],
+ "category_id": 4,
+ "id": 369,
+ "area": 294,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 205,
+ "bbox": [
+ 413,
+ 331,
+ 16,
+ 15
+ ],
+ "category_id": 4,
+ "id": 370,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 205,
+ "bbox": [
+ 361,
+ 448,
+ 18,
+ 12
+ ],
+ "category_id": 4,
+ "id": 371,
+ "area": 216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 205,
+ "bbox": [
+ 49,
+ 472,
+ 18,
+ 11
+ ],
+ "category_id": 4,
+ "id": 372,
+ "area": 198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 206,
+ "bbox": [
+ 52,
+ 190,
+ 371,
+ 222
+ ],
+ "category_id": 16,
+ "id": 373,
+ "area": 82362,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 207,
+ "bbox": [
+ 18,
+ 30,
+ 204,
+ 163
+ ],
+ "category_id": 10,
+ "id": 374,
+ "area": 33252,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 208,
+ "bbox": [
+ 293,
+ 150,
+ 81,
+ 236
+ ],
+ "category_id": 16,
+ "id": 375,
+ "area": 19116,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 209,
+ "bbox": [
+ 156,
+ 85,
+ 19,
+ 19
+ ],
+ "category_id": 4,
+ "id": 376,
+ "area": 361,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 209,
+ "bbox": [
+ 399,
+ 33,
+ 17,
+ 18
+ ],
+ "category_id": 4,
+ "id": 377,
+ "area": 306,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 209,
+ "bbox": [
+ 285,
+ 152,
+ 21,
+ 23
+ ],
+ "category_id": 4,
+ "id": 378,
+ "area": 483,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 209,
+ "bbox": [
+ 46,
+ 232,
+ 18,
+ 16
+ ],
+ "category_id": 4,
+ "id": 379,
+ "area": 288,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 209,
+ "bbox": [
+ 27,
+ 468,
+ 17,
+ 12
+ ],
+ "category_id": 4,
+ "id": 380,
+ "area": 204,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 209,
+ "bbox": [
+ 167,
+ 453,
+ 20,
+ 16
+ ],
+ "category_id": 4,
+ "id": 381,
+ "area": 320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 210,
+ "bbox": [
+ 196,
+ 138,
+ 96,
+ 156
+ ],
+ "category_id": 19,
+ "id": 382,
+ "area": 14976,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 211,
+ "bbox": [
+ 184,
+ 58,
+ 38,
+ 46
+ ],
+ "category_id": 4,
+ "id": 383,
+ "area": 1748,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 211,
+ "bbox": [
+ 257,
+ 421,
+ 39,
+ 50
+ ],
+ "category_id": 4,
+ "id": 384,
+ "area": 1950,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 212,
+ "bbox": [
+ 224,
+ 42,
+ 76,
+ 379
+ ],
+ "category_id": 10,
+ "id": 385,
+ "area": 28804,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 213,
+ "bbox": [
+ 171,
+ 60,
+ 107,
+ 382
+ ],
+ "category_id": 10,
+ "id": 386,
+ "area": 40874,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 214,
+ "bbox": [
+ 206,
+ 57,
+ 121,
+ 324
+ ],
+ "category_id": 16,
+ "id": 387,
+ "area": 39204,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 215,
+ "bbox": [
+ 10,
+ 364,
+ 129,
+ 82
+ ],
+ "category_id": 19,
+ "id": 388,
+ "area": 10578,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 216,
+ "bbox": [
+ 42,
+ 324,
+ 124,
+ 52
+ ],
+ "category_id": 19,
+ "id": 389,
+ "area": 6448,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 217,
+ "bbox": [
+ 216,
+ 261,
+ 85,
+ 94
+ ],
+ "category_id": 15,
+ "id": 390,
+ "area": 7990,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 218,
+ "bbox": [
+ 78,
+ 211,
+ 316,
+ 219
+ ],
+ "category_id": 16,
+ "id": 391,
+ "area": 69204,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 219,
+ "bbox": [
+ 208,
+ 184,
+ 95,
+ 191
+ ],
+ "category_id": 19,
+ "id": 392,
+ "area": 18145,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 220,
+ "bbox": [
+ 81,
+ 163,
+ 396,
+ 240
+ ],
+ "category_id": 10,
+ "id": 393,
+ "area": 95040,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 221,
+ "bbox": [
+ 188,
+ 177,
+ 77,
+ 81
+ ],
+ "category_id": 4,
+ "id": 394,
+ "area": 6237,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 222,
+ "bbox": [
+ 236,
+ 420,
+ 67,
+ 48
+ ],
+ "category_id": 4,
+ "id": 395,
+ "area": 3216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 222,
+ "bbox": [
+ 298,
+ 116,
+ 53,
+ 46
+ ],
+ "category_id": 4,
+ "id": 396,
+ "area": 2438,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 223,
+ "bbox": [
+ 20,
+ 1,
+ 10,
+ 16
+ ],
+ "category_id": 4,
+ "id": 397,
+ "area": 160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 223,
+ "bbox": [
+ 238,
+ 198,
+ 26,
+ 24
+ ],
+ "category_id": 4,
+ "id": 398,
+ "area": 624,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 223,
+ "bbox": [
+ 151,
+ 300,
+ 12,
+ 16
+ ],
+ "category_id": 4,
+ "id": 399,
+ "area": 192,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 223,
+ "bbox": [
+ 64,
+ 405,
+ 12,
+ 17
+ ],
+ "category_id": 4,
+ "id": 400,
+ "area": 204,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 223,
+ "bbox": [
+ 366,
+ 497,
+ 10,
+ 15
+ ],
+ "category_id": 4,
+ "id": 401,
+ "area": 150,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 224,
+ "bbox": [
+ 79,
+ 253,
+ 15,
+ 15
+ ],
+ "category_id": 4,
+ "id": 402,
+ "area": 225,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 224,
+ "bbox": [
+ 30,
+ 452,
+ 32,
+ 20
+ ],
+ "category_id": 4,
+ "id": 403,
+ "area": 640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 224,
+ "bbox": [
+ 222,
+ 423,
+ 24,
+ 16
+ ],
+ "category_id": 4,
+ "id": 404,
+ "area": 384,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 224,
+ "bbox": [
+ 383,
+ 129,
+ 19,
+ 18
+ ],
+ "category_id": 4,
+ "id": 405,
+ "area": 342,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 225,
+ "bbox": [
+ 355,
+ 261,
+ 126,
+ 233
+ ],
+ "category_id": 10,
+ "id": 406,
+ "area": 29358,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 226,
+ "bbox": [
+ 161,
+ 105,
+ 96,
+ 203
+ ],
+ "category_id": 19,
+ "id": 407,
+ "area": 19488,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 227,
+ "bbox": [
+ 237,
+ 83,
+ 11,
+ 25
+ ],
+ "category_id": 4,
+ "id": 408,
+ "area": 275,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 227,
+ "bbox": [
+ 48,
+ 225,
+ 12,
+ 21
+ ],
+ "category_id": 4,
+ "id": 409,
+ "area": 252,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 227,
+ "bbox": [
+ 220,
+ 213,
+ 26,
+ 22
+ ],
+ "category_id": 4,
+ "id": 410,
+ "area": 572,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 227,
+ "bbox": [
+ 440,
+ 280,
+ 15,
+ 27
+ ],
+ "category_id": 4,
+ "id": 411,
+ "area": 405,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 227,
+ "bbox": [
+ 134,
+ 447,
+ 14,
+ 21
+ ],
+ "category_id": 4,
+ "id": 412,
+ "area": 294,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 227,
+ "bbox": [
+ 369,
+ 405,
+ 13,
+ 27
+ ],
+ "category_id": 4,
+ "id": 413,
+ "area": 351,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 228,
+ "bbox": [
+ 178,
+ 395,
+ 52,
+ 31
+ ],
+ "category_id": 4,
+ "id": 414,
+ "area": 1612,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 228,
+ "bbox": [
+ 261,
+ 85,
+ 72,
+ 27
+ ],
+ "category_id": 4,
+ "id": 415,
+ "area": 1944,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 229,
+ "bbox": [
+ 133,
+ 156,
+ 29,
+ 83
+ ],
+ "category_id": 16,
+ "id": 416,
+ "area": 2407,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 230,
+ "bbox": [
+ 124,
+ 208,
+ 177,
+ 165
+ ],
+ "category_id": 16,
+ "id": 417,
+ "area": 29205,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 231,
+ "bbox": [
+ 128,
+ 135,
+ 79,
+ 81
+ ],
+ "category_id": 15,
+ "id": 418,
+ "area": 6399,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 231,
+ "bbox": [
+ 303,
+ 326,
+ 45,
+ 24
+ ],
+ "category_id": 15,
+ "id": 419,
+ "area": 1080,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 232,
+ "bbox": [
+ 231,
+ 201,
+ 47,
+ 200
+ ],
+ "category_id": 19,
+ "id": 420,
+ "area": 9400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 233,
+ "bbox": [
+ 22,
+ 13,
+ 26,
+ 17
+ ],
+ "category_id": 4,
+ "id": 421,
+ "area": 442,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 233,
+ "bbox": [
+ 193,
+ 55,
+ 31,
+ 16
+ ],
+ "category_id": 4,
+ "id": 422,
+ "area": 496,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 233,
+ "bbox": [
+ 106,
+ 261,
+ 32,
+ 18
+ ],
+ "category_id": 4,
+ "id": 423,
+ "area": 576,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 233,
+ "bbox": [
+ 208,
+ 451,
+ 32,
+ 16
+ ],
+ "category_id": 4,
+ "id": 424,
+ "area": 512,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 234,
+ "bbox": [
+ 86,
+ 206,
+ 369,
+ 96
+ ],
+ "category_id": 10,
+ "id": 425,
+ "area": 35424,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 235,
+ "bbox": [
+ 105,
+ 143,
+ 105,
+ 104
+ ],
+ "category_id": 15,
+ "id": 426,
+ "area": 10920,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 235,
+ "bbox": [
+ 105,
+ 289,
+ 105,
+ 105
+ ],
+ "category_id": 15,
+ "id": 427,
+ "area": 11025,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 236,
+ "bbox": [
+ 122,
+ 65,
+ 228,
+ 137
+ ],
+ "category_id": 19,
+ "id": 428,
+ "area": 31236,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 237,
+ "bbox": [
+ 94,
+ 187,
+ 32,
+ 40
+ ],
+ "category_id": 4,
+ "id": 429,
+ "area": 1280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 237,
+ "bbox": [
+ 286,
+ 172,
+ 30,
+ 36
+ ],
+ "category_id": 4,
+ "id": 430,
+ "area": 1080,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 237,
+ "bbox": [
+ 446,
+ 271,
+ 28,
+ 38
+ ],
+ "category_id": 4,
+ "id": 431,
+ "area": 1064,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 238,
+ "bbox": [
+ 190,
+ 24,
+ 89,
+ 48
+ ],
+ "category_id": 19,
+ "id": 432,
+ "area": 4272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 239,
+ "bbox": [
+ 319,
+ 48,
+ 162,
+ 126
+ ],
+ "category_id": 10,
+ "id": 433,
+ "area": 20412,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 240,
+ "bbox": [
+ 200,
+ 298,
+ 97,
+ 94
+ ],
+ "category_id": 4,
+ "id": 434,
+ "area": 9118,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 241,
+ "bbox": [
+ 163,
+ 154,
+ 245,
+ 278
+ ],
+ "category_id": 16,
+ "id": 435,
+ "area": 68110,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 242,
+ "bbox": [
+ 45,
+ 239,
+ 86,
+ 87
+ ],
+ "category_id": 15,
+ "id": 436,
+ "area": 7482,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 242,
+ "bbox": [
+ 108,
+ 339,
+ 87,
+ 87
+ ],
+ "category_id": 15,
+ "id": 437,
+ "area": 7569,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 243,
+ "bbox": [
+ 241,
+ 148,
+ 127,
+ 154
+ ],
+ "category_id": 15,
+ "id": 438,
+ "area": 19558,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 243,
+ "bbox": [
+ 120,
+ 141,
+ 61,
+ 95
+ ],
+ "category_id": 15,
+ "id": 439,
+ "area": 5795,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 244,
+ "bbox": [
+ 11,
+ 85,
+ 28,
+ 16
+ ],
+ "category_id": 4,
+ "id": 440,
+ "area": 448,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 244,
+ "bbox": [
+ 92,
+ 236,
+ 24,
+ 9
+ ],
+ "category_id": 4,
+ "id": 441,
+ "area": 216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 244,
+ "bbox": [
+ 181,
+ 254,
+ 30,
+ 16
+ ],
+ "category_id": 4,
+ "id": 442,
+ "area": 480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 244,
+ "bbox": [
+ 297,
+ 285,
+ 19,
+ 9
+ ],
+ "category_id": 4,
+ "id": 443,
+ "area": 171,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 244,
+ "bbox": [
+ 460,
+ 285,
+ 30,
+ 16
+ ],
+ "category_id": 4,
+ "id": 444,
+ "area": 480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 245,
+ "bbox": [
+ 33,
+ 79,
+ 130,
+ 68
+ ],
+ "category_id": 10,
+ "id": 445,
+ "area": 8840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 246,
+ "bbox": [
+ 216,
+ 98,
+ 24,
+ 44
+ ],
+ "category_id": 4,
+ "id": 446,
+ "area": 1056,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 246,
+ "bbox": [
+ 289,
+ 375,
+ 29,
+ 46
+ ],
+ "category_id": 4,
+ "id": 447,
+ "area": 1334,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 247,
+ "bbox": [
+ 96,
+ 129,
+ 55,
+ 32
+ ],
+ "category_id": 4,
+ "id": 448,
+ "area": 1760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 247,
+ "bbox": [
+ 408,
+ 248,
+ 64,
+ 43
+ ],
+ "category_id": 4,
+ "id": 449,
+ "area": 2752,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 248,
+ "bbox": [
+ 78,
+ 145,
+ 38,
+ 32
+ ],
+ "category_id": 4,
+ "id": 450,
+ "area": 1216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 248,
+ "bbox": [
+ 404,
+ 295,
+ 48,
+ 38
+ ],
+ "category_id": 4,
+ "id": 451,
+ "area": 1824,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 249,
+ "bbox": [
+ 97,
+ 144,
+ 57,
+ 30
+ ],
+ "category_id": 4,
+ "id": 452,
+ "area": 1710,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 249,
+ "bbox": [
+ 403,
+ 360,
+ 56,
+ 30
+ ],
+ "category_id": 4,
+ "id": 453,
+ "area": 1680,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 250,
+ "bbox": [
+ 193,
+ 163,
+ 239,
+ 302
+ ],
+ "category_id": 16,
+ "id": 454,
+ "area": 72178,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 251,
+ "bbox": [
+ 72,
+ 176,
+ 340,
+ 164
+ ],
+ "category_id": 16,
+ "id": 455,
+ "area": 55760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 252,
+ "bbox": [
+ 85,
+ 113,
+ 32,
+ 59
+ ],
+ "category_id": 4,
+ "id": 456,
+ "area": 1888,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 252,
+ "bbox": [
+ 416,
+ 275,
+ 21,
+ 50
+ ],
+ "category_id": 4,
+ "id": 457,
+ "area": 1050,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 253,
+ "bbox": [
+ 100,
+ 79,
+ 167,
+ 188
+ ],
+ "category_id": 19,
+ "id": 458,
+ "area": 31396,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 254,
+ "bbox": [
+ 167,
+ 221,
+ 108,
+ 86
+ ],
+ "category_id": 15,
+ "id": 459,
+ "area": 9288,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 255,
+ "bbox": [
+ 55,
+ 266,
+ 39,
+ 30
+ ],
+ "category_id": 4,
+ "id": 460,
+ "area": 1170,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 255,
+ "bbox": [
+ 250,
+ 34,
+ 28,
+ 36
+ ],
+ "category_id": 4,
+ "id": 461,
+ "area": 1008,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 255,
+ "bbox": [
+ 65,
+ 440,
+ 18,
+ 29
+ ],
+ "category_id": 4,
+ "id": 462,
+ "area": 522,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 255,
+ "bbox": [
+ 242,
+ 460,
+ 36,
+ 31
+ ],
+ "category_id": 4,
+ "id": 463,
+ "area": 1116,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 255,
+ "bbox": [
+ 417,
+ 336,
+ 31,
+ 23
+ ],
+ "category_id": 4,
+ "id": 464,
+ "area": 713,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 255,
+ "bbox": [
+ 458,
+ 155,
+ 26,
+ 29
+ ],
+ "category_id": 4,
+ "id": 465,
+ "area": 754,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 256,
+ "bbox": [
+ 200,
+ 110,
+ 48,
+ 253
+ ],
+ "category_id": 19,
+ "id": 466,
+ "area": 12144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 257,
+ "bbox": [
+ 151,
+ 151,
+ 196,
+ 174
+ ],
+ "category_id": 19,
+ "id": 467,
+ "area": 34104,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 258,
+ "bbox": [
+ 253,
+ 66,
+ 216,
+ 130
+ ],
+ "category_id": 10,
+ "id": 468,
+ "area": 28080,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 259,
+ "bbox": [
+ 184,
+ 169,
+ 201,
+ 161
+ ],
+ "category_id": 16,
+ "id": 469,
+ "area": 32361,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 260,
+ "bbox": [
+ 38,
+ 140,
+ 429,
+ 84
+ ],
+ "category_id": 19,
+ "id": 470,
+ "area": 36036,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 261,
+ "bbox": [
+ 180,
+ 132,
+ 127,
+ 179
+ ],
+ "category_id": 19,
+ "id": 471,
+ "area": 22733,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 262,
+ "bbox": [
+ 50,
+ 176,
+ 388,
+ 211
+ ],
+ "category_id": 10,
+ "id": 472,
+ "area": 81868,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 263,
+ "bbox": [
+ 162,
+ 91,
+ 74,
+ 45
+ ],
+ "category_id": 4,
+ "id": 473,
+ "area": 3330,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 263,
+ "bbox": [
+ 316,
+ 404,
+ 73,
+ 48
+ ],
+ "category_id": 4,
+ "id": 474,
+ "area": 3504,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 264,
+ "bbox": [
+ 203,
+ 112,
+ 76,
+ 283
+ ],
+ "category_id": 19,
+ "id": 475,
+ "area": 21508,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 265,
+ "bbox": [
+ 46,
+ 28,
+ 133,
+ 153
+ ],
+ "category_id": 10,
+ "id": 476,
+ "area": 20349,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 266,
+ "bbox": [
+ 51,
+ 196,
+ 381,
+ 125
+ ],
+ "category_id": 10,
+ "id": 477,
+ "area": 47625,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 267,
+ "bbox": [
+ 228,
+ 201,
+ 97,
+ 103
+ ],
+ "category_id": 4,
+ "id": 478,
+ "area": 9991,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 268,
+ "bbox": [
+ 83,
+ 119,
+ 386,
+ 282
+ ],
+ "category_id": 10,
+ "id": 479,
+ "area": 108852,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 269,
+ "bbox": [
+ 176,
+ 167,
+ 77,
+ 223
+ ],
+ "category_id": 19,
+ "id": 480,
+ "area": 17171,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 270,
+ "bbox": [
+ 169,
+ 197,
+ 165,
+ 73
+ ],
+ "category_id": 4,
+ "id": 481,
+ "area": 12045,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 271,
+ "bbox": [
+ 165,
+ 247,
+ 238,
+ 128
+ ],
+ "category_id": 16,
+ "id": 482,
+ "area": 30464,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 272,
+ "bbox": [
+ 147,
+ 83,
+ 97,
+ 73
+ ],
+ "category_id": 19,
+ "id": 483,
+ "area": 7081,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 273,
+ "bbox": [
+ 96,
+ 304,
+ 231,
+ 188
+ ],
+ "category_id": 15,
+ "id": 484,
+ "area": 43428,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 273,
+ "bbox": [
+ 211,
+ 72,
+ 229,
+ 186
+ ],
+ "category_id": 15,
+ "id": 485,
+ "area": 42594,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 274,
+ "bbox": [
+ 257,
+ 174,
+ 45,
+ 126
+ ],
+ "category_id": 19,
+ "id": 486,
+ "area": 5670,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 275,
+ "bbox": [
+ 206,
+ 106,
+ 27,
+ 272
+ ],
+ "category_id": 19,
+ "id": 487,
+ "area": 7344,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 276,
+ "bbox": [
+ 56,
+ 201,
+ 326,
+ 132
+ ],
+ "category_id": 16,
+ "id": 488,
+ "area": 43032,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 277,
+ "bbox": [
+ 182,
+ 116,
+ 255,
+ 254
+ ],
+ "category_id": 16,
+ "id": 489,
+ "area": 64770,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 278,
+ "bbox": [
+ 358,
+ 346,
+ 76,
+ 33
+ ],
+ "category_id": 4,
+ "id": 490,
+ "area": 2508,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 278,
+ "bbox": [
+ 142,
+ 117,
+ 60,
+ 26
+ ],
+ "category_id": 4,
+ "id": 491,
+ "area": 1560,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 279,
+ "bbox": [
+ 260,
+ 129,
+ 70,
+ 282
+ ],
+ "category_id": 10,
+ "id": 492,
+ "area": 19740,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 280,
+ "bbox": [
+ 15,
+ 252,
+ 484,
+ 180
+ ],
+ "category_id": 16,
+ "id": 493,
+ "area": 87120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 281,
+ "bbox": [
+ 60,
+ 319,
+ 38,
+ 20
+ ],
+ "category_id": 4,
+ "id": 494,
+ "area": 760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 281,
+ "bbox": [
+ 69,
+ 68,
+ 32,
+ 27
+ ],
+ "category_id": 4,
+ "id": 495,
+ "area": 864,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 281,
+ "bbox": [
+ 339,
+ 51,
+ 49,
+ 16
+ ],
+ "category_id": 4,
+ "id": 496,
+ "area": 784,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 281,
+ "bbox": [
+ 250,
+ 432,
+ 46,
+ 22
+ ],
+ "category_id": 4,
+ "id": 497,
+ "area": 1012,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 281,
+ "bbox": [
+ 282,
+ 308,
+ 43,
+ 15
+ ],
+ "category_id": 4,
+ "id": 498,
+ "area": 645,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 281,
+ "bbox": [
+ 404,
+ 375,
+ 49,
+ 22
+ ],
+ "category_id": 4,
+ "id": 499,
+ "area": 1078,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 282,
+ "bbox": [
+ 140,
+ 240,
+ 241,
+ 54
+ ],
+ "category_id": 19,
+ "id": 500,
+ "area": 13014,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 283,
+ "bbox": [
+ 203,
+ 174,
+ 145,
+ 213
+ ],
+ "category_id": 10,
+ "id": 501,
+ "area": 30885,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 284,
+ "bbox": [
+ 53,
+ 217,
+ 169,
+ 121
+ ],
+ "category_id": 15,
+ "id": 502,
+ "area": 20449,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 284,
+ "bbox": [
+ 242,
+ 223,
+ 168,
+ 126
+ ],
+ "category_id": 15,
+ "id": 503,
+ "area": 21168,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 285,
+ "bbox": [
+ 304,
+ 48,
+ 188,
+ 142
+ ],
+ "category_id": 10,
+ "id": 504,
+ "area": 26696,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 286,
+ "bbox": [
+ 143,
+ 206,
+ 145,
+ 168
+ ],
+ "category_id": 16,
+ "id": 505,
+ "area": 24360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 286,
+ "bbox": [
+ 35,
+ 294,
+ 52,
+ 58
+ ],
+ "category_id": 16,
+ "id": 506,
+ "area": 3016,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 287,
+ "bbox": [
+ 74,
+ 222,
+ 392,
+ 147
+ ],
+ "category_id": 16,
+ "id": 507,
+ "area": 57624,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 288,
+ "bbox": [
+ 128,
+ 81,
+ 346,
+ 358
+ ],
+ "category_id": 10,
+ "id": 508,
+ "area": 123868,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 289,
+ "bbox": [
+ 109,
+ 266,
+ 320,
+ 108
+ ],
+ "category_id": 19,
+ "id": 509,
+ "area": 34560,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 290,
+ "bbox": [
+ 98,
+ 103,
+ 40,
+ 147
+ ],
+ "category_id": 10,
+ "id": 510,
+ "area": 5880,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 291,
+ "bbox": [
+ 32,
+ 44,
+ 135,
+ 117
+ ],
+ "category_id": 10,
+ "id": 511,
+ "area": 15795,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 292,
+ "bbox": [
+ 226,
+ 145,
+ 65,
+ 248
+ ],
+ "category_id": 19,
+ "id": 512,
+ "area": 16120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 293,
+ "bbox": [
+ 91,
+ 199,
+ 38,
+ 29
+ ],
+ "category_id": 4,
+ "id": 513,
+ "area": 1102,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 293,
+ "bbox": [
+ 396,
+ 286,
+ 43,
+ 33
+ ],
+ "category_id": 4,
+ "id": 514,
+ "area": 1419,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 294,
+ "bbox": [
+ 215,
+ 202,
+ 49,
+ 82
+ ],
+ "category_id": 19,
+ "id": 515,
+ "area": 4018,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 295,
+ "bbox": [
+ 149,
+ 230,
+ 223,
+ 30
+ ],
+ "category_id": 19,
+ "id": 516,
+ "area": 6690,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 296,
+ "bbox": [
+ 200,
+ 245,
+ 109,
+ 55
+ ],
+ "category_id": 4,
+ "id": 517,
+ "area": 5995,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 297,
+ "bbox": [
+ 257,
+ 201,
+ 71,
+ 131
+ ],
+ "category_id": 4,
+ "id": 518,
+ "area": 9301,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 298,
+ "bbox": [
+ 97,
+ 55,
+ 366,
+ 378
+ ],
+ "category_id": 10,
+ "id": 519,
+ "area": 138348,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 299,
+ "bbox": [
+ 147,
+ 57,
+ 178,
+ 425
+ ],
+ "category_id": 16,
+ "id": 520,
+ "area": 75650,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 300,
+ "bbox": [
+ 320,
+ 336,
+ 123,
+ 82
+ ],
+ "category_id": 4,
+ "id": 521,
+ "area": 10086,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 301,
+ "bbox": [
+ 144,
+ 86,
+ 260,
+ 377
+ ],
+ "category_id": 10,
+ "id": 522,
+ "area": 98020,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 302,
+ "bbox": [
+ 56,
+ 277,
+ 66,
+ 35
+ ],
+ "category_id": 4,
+ "id": 523,
+ "area": 2310,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 302,
+ "bbox": [
+ 368,
+ 214,
+ 67,
+ 33
+ ],
+ "category_id": 4,
+ "id": 524,
+ "area": 2211,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 303,
+ "bbox": [
+ 26,
+ 139,
+ 424,
+ 299
+ ],
+ "category_id": 16,
+ "id": 525,
+ "area": 126776,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 304,
+ "bbox": [
+ 194,
+ 165,
+ 132,
+ 149
+ ],
+ "category_id": 15,
+ "id": 526,
+ "area": 19668,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 305,
+ "bbox": [
+ 48,
+ 349,
+ 61,
+ 55
+ ],
+ "category_id": 4,
+ "id": 527,
+ "area": 3355,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 305,
+ "bbox": [
+ 406,
+ 152,
+ 59,
+ 55
+ ],
+ "category_id": 4,
+ "id": 528,
+ "area": 3245,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 306,
+ "bbox": [
+ 106,
+ 180,
+ 352,
+ 235
+ ],
+ "category_id": 10,
+ "id": 529,
+ "area": 82720,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 307,
+ "bbox": [
+ 247,
+ 197,
+ 138,
+ 105
+ ],
+ "category_id": 4,
+ "id": 530,
+ "area": 14490,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 308,
+ "bbox": [
+ 137,
+ 214,
+ 113,
+ 155
+ ],
+ "category_id": 15,
+ "id": 531,
+ "area": 17515,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 309,
+ "bbox": [
+ 222,
+ 204,
+ 39,
+ 126
+ ],
+ "category_id": 4,
+ "id": 532,
+ "area": 4914,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 310,
+ "bbox": [
+ 61,
+ 297,
+ 8,
+ 14
+ ],
+ "category_id": 4,
+ "id": 533,
+ "area": 112,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 310,
+ "bbox": [
+ 60,
+ 256,
+ 5,
+ 13
+ ],
+ "category_id": 4,
+ "id": 534,
+ "area": 65,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 310,
+ "bbox": [
+ 202,
+ 261,
+ 7,
+ 12
+ ],
+ "category_id": 4,
+ "id": 535,
+ "area": 84,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 310,
+ "bbox": [
+ 325,
+ 325,
+ 7,
+ 11
+ ],
+ "category_id": 4,
+ "id": 536,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 310,
+ "bbox": [
+ 496,
+ 241,
+ 8,
+ 11
+ ],
+ "category_id": 4,
+ "id": 537,
+ "area": 88,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 310,
+ "bbox": [
+ 376,
+ 45,
+ 10,
+ 13
+ ],
+ "category_id": 4,
+ "id": 538,
+ "area": 130,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 311,
+ "bbox": [
+ 52,
+ 63,
+ 28,
+ 23
+ ],
+ "category_id": 4,
+ "id": 539,
+ "area": 644,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 311,
+ "bbox": [
+ 172,
+ 133,
+ 32,
+ 22
+ ],
+ "category_id": 4,
+ "id": 540,
+ "area": 704,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 311,
+ "bbox": [
+ 424,
+ 340,
+ 19,
+ 18
+ ],
+ "category_id": 4,
+ "id": 541,
+ "area": 342,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 312,
+ "bbox": [
+ 179,
+ 140,
+ 148,
+ 118
+ ],
+ "category_id": 15,
+ "id": 542,
+ "area": 17464,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 312,
+ "bbox": [
+ 215,
+ 301,
+ 148,
+ 118
+ ],
+ "category_id": 15,
+ "id": 543,
+ "area": 17464,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 313,
+ "bbox": [
+ 30,
+ 343,
+ 40,
+ 55
+ ],
+ "category_id": 4,
+ "id": 544,
+ "area": 2200,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 313,
+ "bbox": [
+ 324,
+ 85,
+ 54,
+ 52
+ ],
+ "category_id": 4,
+ "id": 545,
+ "area": 2808,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 313,
+ "bbox": [
+ 402,
+ 396,
+ 40,
+ 57
+ ],
+ "category_id": 4,
+ "id": 546,
+ "area": 2280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 314,
+ "bbox": [
+ 301,
+ 40,
+ 11,
+ 18
+ ],
+ "category_id": 4,
+ "id": 547,
+ "area": 198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 314,
+ "bbox": [
+ 65,
+ 261,
+ 11,
+ 16
+ ],
+ "category_id": 4,
+ "id": 548,
+ "area": 176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 314,
+ "bbox": [
+ 149,
+ 481,
+ 13,
+ 17
+ ],
+ "category_id": 4,
+ "id": 549,
+ "area": 221,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 315,
+ "bbox": [
+ 456,
+ 233,
+ 25,
+ 18
+ ],
+ "category_id": 4,
+ "id": 550,
+ "area": 450,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 315,
+ "bbox": [
+ 314,
+ 81,
+ 15,
+ 17
+ ],
+ "category_id": 4,
+ "id": 551,
+ "area": 255,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 315,
+ "bbox": [
+ 193,
+ 160,
+ 17,
+ 16
+ ],
+ "category_id": 4,
+ "id": 552,
+ "area": 272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 315,
+ "bbox": [
+ 172,
+ 272,
+ 25,
+ 17
+ ],
+ "category_id": 4,
+ "id": 553,
+ "area": 425,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 315,
+ "bbox": [
+ 114,
+ 371,
+ 20,
+ 18
+ ],
+ "category_id": 4,
+ "id": 554,
+ "area": 360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 315,
+ "bbox": [
+ 16,
+ 456,
+ 17,
+ 18
+ ],
+ "category_id": 4,
+ "id": 555,
+ "area": 306,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 316,
+ "bbox": [
+ 73,
+ 165,
+ 140,
+ 130
+ ],
+ "category_id": 15,
+ "id": 556,
+ "area": 18200,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 316,
+ "bbox": [
+ 110,
+ 318,
+ 144,
+ 137
+ ],
+ "category_id": 15,
+ "id": 557,
+ "area": 19728,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 317,
+ "bbox": [
+ 186,
+ 133,
+ 59,
+ 251
+ ],
+ "category_id": 19,
+ "id": 558,
+ "area": 14809,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 318,
+ "bbox": [
+ 175,
+ 84,
+ 146,
+ 360
+ ],
+ "category_id": 16,
+ "id": 559,
+ "area": 52560,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 319,
+ "bbox": [
+ 40,
+ 60,
+ 50,
+ 43
+ ],
+ "category_id": 4,
+ "id": 560,
+ "area": 2150,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 319,
+ "bbox": [
+ 120,
+ 446,
+ 46,
+ 44
+ ],
+ "category_id": 4,
+ "id": 561,
+ "area": 2024,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 319,
+ "bbox": [
+ 424,
+ 227,
+ 39,
+ 46
+ ],
+ "category_id": 4,
+ "id": 562,
+ "area": 1794,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 320,
+ "bbox": [
+ 23,
+ 140,
+ 35,
+ 12
+ ],
+ "category_id": 4,
+ "id": 563,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 320,
+ "bbox": [
+ 410,
+ 91,
+ 34,
+ 9
+ ],
+ "category_id": 4,
+ "id": 564,
+ "area": 306,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 320,
+ "bbox": [
+ 168,
+ 223,
+ 13,
+ 11
+ ],
+ "category_id": 4,
+ "id": 565,
+ "area": 143,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 320,
+ "bbox": [
+ 39,
+ 332,
+ 24,
+ 18
+ ],
+ "category_id": 4,
+ "id": 566,
+ "area": 432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 320,
+ "bbox": [
+ 240,
+ 432,
+ 32,
+ 19
+ ],
+ "category_id": 4,
+ "id": 567,
+ "area": 608,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 320,
+ "bbox": [
+ 86,
+ 476,
+ 18,
+ 14
+ ],
+ "category_id": 4,
+ "id": 568,
+ "area": 252,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 321,
+ "bbox": [
+ 360,
+ 234,
+ 39,
+ 45
+ ],
+ "category_id": 4,
+ "id": 569,
+ "area": 1755,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 322,
+ "bbox": [
+ 77,
+ 261,
+ 35,
+ 37
+ ],
+ "category_id": 4,
+ "id": 570,
+ "area": 1295,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 322,
+ "bbox": [
+ 449,
+ 207,
+ 39,
+ 38
+ ],
+ "category_id": 4,
+ "id": 571,
+ "area": 1482,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 323,
+ "bbox": [
+ 151,
+ 389,
+ 48,
+ 47
+ ],
+ "category_id": 4,
+ "id": 572,
+ "area": 2256,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 323,
+ "bbox": [
+ 338,
+ 88,
+ 65,
+ 41
+ ],
+ "category_id": 4,
+ "id": 573,
+ "area": 2665,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 324,
+ "bbox": [
+ 228,
+ 79,
+ 202,
+ 319
+ ],
+ "category_id": 10,
+ "id": 574,
+ "area": 64438,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 325,
+ "bbox": [
+ 147,
+ 144,
+ 110,
+ 128
+ ],
+ "category_id": 4,
+ "id": 575,
+ "area": 14080,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 326,
+ "bbox": [
+ 93,
+ 220,
+ 214,
+ 133
+ ],
+ "category_id": 16,
+ "id": 576,
+ "area": 28462,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 327,
+ "bbox": [
+ 138,
+ 166,
+ 215,
+ 49
+ ],
+ "category_id": 19,
+ "id": 577,
+ "area": 10535,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 328,
+ "bbox": [
+ 117,
+ 104,
+ 26,
+ 14
+ ],
+ "category_id": 4,
+ "id": 578,
+ "area": 364,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 328,
+ "bbox": [
+ 344,
+ 66,
+ 21,
+ 10
+ ],
+ "category_id": 4,
+ "id": 579,
+ "area": 210,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 328,
+ "bbox": [
+ 269,
+ 167,
+ 24,
+ 12
+ ],
+ "category_id": 4,
+ "id": 580,
+ "area": 288,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 328,
+ "bbox": [
+ 245,
+ 388,
+ 30,
+ 16
+ ],
+ "category_id": 4,
+ "id": 581,
+ "area": 480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 328,
+ "bbox": [
+ 368,
+ 364,
+ 24,
+ 12
+ ],
+ "category_id": 4,
+ "id": 582,
+ "area": 288,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 329,
+ "bbox": [
+ 273,
+ 20,
+ 15,
+ 17
+ ],
+ "category_id": 4,
+ "id": 583,
+ "area": 255,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 329,
+ "bbox": [
+ 210,
+ 172,
+ 15,
+ 14
+ ],
+ "category_id": 4,
+ "id": 584,
+ "area": 210,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 329,
+ "bbox": [
+ 103,
+ 256,
+ 17,
+ 24
+ ],
+ "category_id": 4,
+ "id": 585,
+ "area": 408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 329,
+ "bbox": [
+ 295,
+ 340,
+ 16,
+ 21
+ ],
+ "category_id": 4,
+ "id": 586,
+ "area": 336,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 329,
+ "bbox": [
+ 444,
+ 187,
+ 18,
+ 21
+ ],
+ "category_id": 4,
+ "id": 587,
+ "area": 378,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 329,
+ "bbox": [
+ 305,
+ 474,
+ 18,
+ 22
+ ],
+ "category_id": 4,
+ "id": 588,
+ "area": 396,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 330,
+ "bbox": [
+ 176,
+ 234,
+ 192,
+ 126
+ ],
+ "category_id": 16,
+ "id": 589,
+ "area": 24192,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 331,
+ "bbox": [
+ 206,
+ 42,
+ 211,
+ 117
+ ],
+ "category_id": 19,
+ "id": 590,
+ "area": 24687,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 332,
+ "bbox": [
+ 176,
+ 141,
+ 261,
+ 205
+ ],
+ "category_id": 16,
+ "id": 591,
+ "area": 53505,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 333,
+ "bbox": [
+ 174,
+ 76,
+ 256,
+ 271
+ ],
+ "category_id": 10,
+ "id": 592,
+ "area": 69376,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 334,
+ "bbox": [
+ 231,
+ 300,
+ 69,
+ 48
+ ],
+ "category_id": 15,
+ "id": 593,
+ "area": 3312,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 334,
+ "bbox": [
+ 273,
+ 136,
+ 71,
+ 52
+ ],
+ "category_id": 15,
+ "id": 594,
+ "area": 3692,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 335,
+ "bbox": [
+ 136,
+ 203,
+ 160,
+ 70
+ ],
+ "category_id": 19,
+ "id": 595,
+ "area": 11200,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 336,
+ "bbox": [
+ 152,
+ 233,
+ 120,
+ 120
+ ],
+ "category_id": 15,
+ "id": 596,
+ "area": 14400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 336,
+ "bbox": [
+ 307,
+ 240,
+ 116,
+ 117
+ ],
+ "category_id": 15,
+ "id": 597,
+ "area": 13572,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 337,
+ "bbox": [
+ 179,
+ 237,
+ 142,
+ 31
+ ],
+ "category_id": 19,
+ "id": 598,
+ "area": 4402,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 338,
+ "bbox": [
+ 146,
+ 136,
+ 266,
+ 184
+ ],
+ "category_id": 10,
+ "id": 599,
+ "area": 48944,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 339,
+ "bbox": [
+ 177,
+ 151,
+ 89,
+ 76
+ ],
+ "category_id": 15,
+ "id": 600,
+ "area": 6764,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 339,
+ "bbox": [
+ 256,
+ 209,
+ 88,
+ 79
+ ],
+ "category_id": 15,
+ "id": 601,
+ "area": 6952,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 340,
+ "bbox": [
+ 158,
+ 76,
+ 18,
+ 9
+ ],
+ "category_id": 4,
+ "id": 602,
+ "area": 162,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 340,
+ "bbox": [
+ 128,
+ 143,
+ 18,
+ 10
+ ],
+ "category_id": 4,
+ "id": 603,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 340,
+ "bbox": [
+ 88,
+ 198,
+ 18,
+ 11
+ ],
+ "category_id": 4,
+ "id": 604,
+ "area": 198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 340,
+ "bbox": [
+ 79,
+ 261,
+ 18,
+ 11
+ ],
+ "category_id": 4,
+ "id": 605,
+ "area": 198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 340,
+ "bbox": [
+ 472,
+ 39,
+ 14,
+ 10
+ ],
+ "category_id": 4,
+ "id": 606,
+ "area": 140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 340,
+ "bbox": [
+ 393,
+ 83,
+ 15,
+ 11
+ ],
+ "category_id": 4,
+ "id": 607,
+ "area": 165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 340,
+ "bbox": [
+ 386,
+ 208,
+ 18,
+ 13
+ ],
+ "category_id": 4,
+ "id": 608,
+ "area": 234,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 340,
+ "bbox": [
+ 412,
+ 285,
+ 20,
+ 12
+ ],
+ "category_id": 4,
+ "id": 609,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 340,
+ "bbox": [
+ 161,
+ 331,
+ 16,
+ 13
+ ],
+ "category_id": 4,
+ "id": 610,
+ "area": 208,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 340,
+ "bbox": [
+ 293,
+ 366,
+ 18,
+ 14
+ ],
+ "category_id": 4,
+ "id": 611,
+ "area": 252,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 341,
+ "bbox": [
+ 376,
+ 104,
+ 26,
+ 14
+ ],
+ "category_id": 15,
+ "id": 612,
+ "area": 364,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 341,
+ "bbox": [
+ 389,
+ 40,
+ 20,
+ 6
+ ],
+ "category_id": 15,
+ "id": 613,
+ "area": 120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 341,
+ "bbox": [
+ 360,
+ 205,
+ 24,
+ 24
+ ],
+ "category_id": 15,
+ "id": 614,
+ "area": 576,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 341,
+ "bbox": [
+ 314,
+ 481,
+ 17,
+ 16
+ ],
+ "category_id": 15,
+ "id": 615,
+ "area": 272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 341,
+ "bbox": [
+ 446,
+ 444,
+ 4,
+ 8
+ ],
+ "category_id": 15,
+ "id": 616,
+ "area": 32,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 341,
+ "bbox": [
+ 452,
+ 460,
+ 4,
+ 8
+ ],
+ "category_id": 15,
+ "id": 617,
+ "area": 32,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 341,
+ "bbox": [
+ 449,
+ 474,
+ 5,
+ 9
+ ],
+ "category_id": 15,
+ "id": 618,
+ "area": 45,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 341,
+ "bbox": [
+ 434,
+ 487,
+ 4,
+ 8
+ ],
+ "category_id": 15,
+ "id": 619,
+ "area": 32,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 341,
+ "bbox": [
+ 433,
+ 499,
+ 2,
+ 7
+ ],
+ "category_id": 15,
+ "id": 620,
+ "area": 14,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 341,
+ "bbox": [
+ 464,
+ 356,
+ 2,
+ 3
+ ],
+ "category_id": 15,
+ "id": 621,
+ "area": 6,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 341,
+ "bbox": [
+ 463,
+ 362,
+ 2,
+ 3
+ ],
+ "category_id": 15,
+ "id": 622,
+ "area": 6,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 341,
+ "bbox": [
+ 462,
+ 367,
+ 3,
+ 3
+ ],
+ "category_id": 15,
+ "id": 623,
+ "area": 9,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 341,
+ "bbox": [
+ 462,
+ 372,
+ 2,
+ 4
+ ],
+ "category_id": 15,
+ "id": 624,
+ "area": 8,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 341,
+ "bbox": [
+ 460,
+ 378,
+ 4,
+ 4
+ ],
+ "category_id": 15,
+ "id": 625,
+ "area": 16,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 341,
+ "bbox": [
+ 456,
+ 404,
+ 4,
+ 4
+ ],
+ "category_id": 15,
+ "id": 626,
+ "area": 16,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 342,
+ "bbox": [
+ 221,
+ 191,
+ 207,
+ 110
+ ],
+ "category_id": 19,
+ "id": 627,
+ "area": 22770,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 343,
+ "bbox": [
+ 166,
+ 87,
+ 164,
+ 367
+ ],
+ "category_id": 10,
+ "id": 628,
+ "area": 60188,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 344,
+ "bbox": [
+ 68,
+ 236,
+ 15,
+ 17
+ ],
+ "category_id": 4,
+ "id": 629,
+ "area": 255,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 344,
+ "bbox": [
+ 369,
+ 229,
+ 20,
+ 12
+ ],
+ "category_id": 4,
+ "id": 630,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 344,
+ "bbox": [
+ 430,
+ 132,
+ 12,
+ 18
+ ],
+ "category_id": 4,
+ "id": 631,
+ "area": 216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 344,
+ "bbox": [
+ 466,
+ 51,
+ 14,
+ 17
+ ],
+ "category_id": 4,
+ "id": 632,
+ "area": 238,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 344,
+ "bbox": [
+ 133,
+ 51,
+ 11,
+ 22
+ ],
+ "category_id": 4,
+ "id": 633,
+ "area": 242,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 344,
+ "bbox": [
+ 383,
+ 332,
+ 24,
+ 20
+ ],
+ "category_id": 4,
+ "id": 634,
+ "area": 480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 345,
+ "bbox": [
+ 115,
+ 264,
+ 253,
+ 71
+ ],
+ "category_id": 19,
+ "id": 635,
+ "area": 17963,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 346,
+ "bbox": [
+ 72,
+ 70,
+ 33,
+ 32
+ ],
+ "category_id": 4,
+ "id": 636,
+ "area": 1056,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 346,
+ "bbox": [
+ 26,
+ 180,
+ 29,
+ 31
+ ],
+ "category_id": 4,
+ "id": 637,
+ "area": 899,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 346,
+ "bbox": [
+ 176,
+ 364,
+ 29,
+ 33
+ ],
+ "category_id": 4,
+ "id": 638,
+ "area": 957,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 346,
+ "bbox": [
+ 182,
+ 461,
+ 24,
+ 31
+ ],
+ "category_id": 4,
+ "id": 639,
+ "area": 744,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 346,
+ "bbox": [
+ 445,
+ 278,
+ 24,
+ 34
+ ],
+ "category_id": 4,
+ "id": 640,
+ "area": 816,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 346,
+ "bbox": [
+ 479,
+ 352,
+ 24,
+ 31
+ ],
+ "category_id": 4,
+ "id": 641,
+ "area": 744,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 347,
+ "bbox": [
+ 297,
+ 53,
+ 58,
+ 420
+ ],
+ "category_id": 10,
+ "id": 642,
+ "area": 24360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 348,
+ "bbox": [
+ 58,
+ 161,
+ 292,
+ 319
+ ],
+ "category_id": 16,
+ "id": 643,
+ "area": 93148,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 349,
+ "bbox": [
+ 145,
+ 216,
+ 231,
+ 80
+ ],
+ "category_id": 19,
+ "id": 644,
+ "area": 18480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 350,
+ "bbox": [
+ 87,
+ 195,
+ 42,
+ 58
+ ],
+ "category_id": 4,
+ "id": 645,
+ "area": 2436,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 350,
+ "bbox": [
+ 351,
+ 160,
+ 48,
+ 61
+ ],
+ "category_id": 4,
+ "id": 646,
+ "area": 2928,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 351,
+ "bbox": [
+ 160,
+ 163,
+ 263,
+ 210
+ ],
+ "category_id": 15,
+ "id": 647,
+ "area": 55230,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 352,
+ "bbox": [
+ 291,
+ 94,
+ 83,
+ 84
+ ],
+ "category_id": 15,
+ "id": 648,
+ "area": 6972,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 352,
+ "bbox": [
+ 217,
+ 239,
+ 86,
+ 84
+ ],
+ "category_id": 15,
+ "id": 649,
+ "area": 7224,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 352,
+ "bbox": [
+ 149,
+ 372,
+ 86,
+ 85
+ ],
+ "category_id": 15,
+ "id": 650,
+ "area": 7310,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 352,
+ "bbox": [
+ 365,
+ 268,
+ 33,
+ 30
+ ],
+ "category_id": 15,
+ "id": 651,
+ "area": 990,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 352,
+ "bbox": [
+ 396,
+ 214,
+ 32,
+ 27
+ ],
+ "category_id": 15,
+ "id": 652,
+ "area": 864,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 352,
+ "bbox": [
+ 422,
+ 159,
+ 33,
+ 27
+ ],
+ "category_id": 15,
+ "id": 653,
+ "area": 891,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 353,
+ "bbox": [
+ 113,
+ 314,
+ 67,
+ 67
+ ],
+ "category_id": 4,
+ "id": 654,
+ "area": 4489,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 353,
+ "bbox": [
+ 49,
+ 72,
+ 47,
+ 66
+ ],
+ "category_id": 4,
+ "id": 655,
+ "area": 3102,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 353,
+ "bbox": [
+ 403,
+ 55,
+ 37,
+ 65
+ ],
+ "category_id": 4,
+ "id": 656,
+ "area": 2405,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 354,
+ "bbox": [
+ 82,
+ 353,
+ 37,
+ 29
+ ],
+ "category_id": 4,
+ "id": 657,
+ "area": 1073,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 354,
+ "bbox": [
+ 261,
+ 143,
+ 28,
+ 29
+ ],
+ "category_id": 4,
+ "id": 658,
+ "area": 812,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 354,
+ "bbox": [
+ 356,
+ 446,
+ 30,
+ 22
+ ],
+ "category_id": 4,
+ "id": 659,
+ "area": 660,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 355,
+ "bbox": [
+ 153,
+ 128,
+ 240,
+ 288
+ ],
+ "category_id": 16,
+ "id": 660,
+ "area": 69120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 356,
+ "bbox": [
+ 91,
+ 36,
+ 236,
+ 40
+ ],
+ "category_id": 10,
+ "id": 661,
+ "area": 9440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 357,
+ "bbox": [
+ 212,
+ 90,
+ 40,
+ 50
+ ],
+ "category_id": 4,
+ "id": 662,
+ "area": 2000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 357,
+ "bbox": [
+ 189,
+ 395,
+ 41,
+ 51
+ ],
+ "category_id": 4,
+ "id": 663,
+ "area": 2091,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 358,
+ "bbox": [
+ 327,
+ 374,
+ 113,
+ 88
+ ],
+ "category_id": 10,
+ "id": 664,
+ "area": 9944,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 359,
+ "bbox": [
+ 72,
+ 424,
+ 58,
+ 35
+ ],
+ "category_id": 4,
+ "id": 665,
+ "area": 2030,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 359,
+ "bbox": [
+ 310,
+ 397,
+ 52,
+ 59
+ ],
+ "category_id": 4,
+ "id": 666,
+ "area": 3068,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 359,
+ "bbox": [
+ 326,
+ 65,
+ 50,
+ 60
+ ],
+ "category_id": 4,
+ "id": 667,
+ "area": 3000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 360,
+ "bbox": [
+ 56,
+ 233,
+ 57,
+ 69
+ ],
+ "category_id": 4,
+ "id": 668,
+ "area": 3933,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 360,
+ "bbox": [
+ 410,
+ 147,
+ 65,
+ 62
+ ],
+ "category_id": 4,
+ "id": 669,
+ "area": 4030,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 361,
+ "bbox": [
+ 140,
+ 189,
+ 350,
+ 175
+ ],
+ "category_id": 10,
+ "id": 670,
+ "area": 61250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 362,
+ "bbox": [
+ 27,
+ 188,
+ 462,
+ 164
+ ],
+ "category_id": 10,
+ "id": 671,
+ "area": 75768,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 363,
+ "bbox": [
+ 181,
+ 168,
+ 248,
+ 255
+ ],
+ "category_id": 10,
+ "id": 672,
+ "area": 63240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 364,
+ "bbox": [
+ 49,
+ 449,
+ 235,
+ 51
+ ],
+ "category_id": 10,
+ "id": 673,
+ "area": 11985,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 365,
+ "bbox": [
+ 87,
+ 230,
+ 305,
+ 128
+ ],
+ "category_id": 19,
+ "id": 674,
+ "area": 39040,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 366,
+ "bbox": [
+ 174,
+ 52,
+ 24,
+ 28
+ ],
+ "category_id": 16,
+ "id": 675,
+ "area": 672,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 366,
+ "bbox": [
+ 76,
+ 330,
+ 64,
+ 30
+ ],
+ "category_id": 16,
+ "id": 676,
+ "area": 1920,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 367,
+ "bbox": [
+ 236,
+ 112,
+ 67,
+ 338
+ ],
+ "category_id": 16,
+ "id": 677,
+ "area": 22646,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 368,
+ "bbox": [
+ 83,
+ 231,
+ 37,
+ 47
+ ],
+ "category_id": 4,
+ "id": 678,
+ "area": 1739,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 368,
+ "bbox": [
+ 442,
+ 222,
+ 32,
+ 26
+ ],
+ "category_id": 4,
+ "id": 679,
+ "area": 832,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 368,
+ "bbox": [
+ 402,
+ 392,
+ 30,
+ 19
+ ],
+ "category_id": 4,
+ "id": 680,
+ "area": 570,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 369,
+ "bbox": [
+ 48,
+ 189,
+ 417,
+ 173
+ ],
+ "category_id": 10,
+ "id": 681,
+ "area": 72141,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 370,
+ "bbox": [
+ 37,
+ 119,
+ 18,
+ 101
+ ],
+ "category_id": 10,
+ "id": 682,
+ "area": 1818,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 371,
+ "bbox": [
+ 208,
+ 23,
+ 14,
+ 9
+ ],
+ "category_id": 4,
+ "id": 683,
+ "area": 126,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 371,
+ "bbox": [
+ 410,
+ 48,
+ 15,
+ 10
+ ],
+ "category_id": 4,
+ "id": 684,
+ "area": 150,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 371,
+ "bbox": [
+ 316,
+ 94,
+ 15,
+ 9
+ ],
+ "category_id": 4,
+ "id": 685,
+ "area": 135,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 371,
+ "bbox": [
+ 137,
+ 132,
+ 15,
+ 9
+ ],
+ "category_id": 4,
+ "id": 686,
+ "area": 135,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 371,
+ "bbox": [
+ 270,
+ 234,
+ 18,
+ 11
+ ],
+ "category_id": 4,
+ "id": 687,
+ "area": 198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 371,
+ "bbox": [
+ 114,
+ 288,
+ 15,
+ 11
+ ],
+ "category_id": 4,
+ "id": 688,
+ "area": 165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 371,
+ "bbox": [
+ 16,
+ 431,
+ 15,
+ 9
+ ],
+ "category_id": 4,
+ "id": 689,
+ "area": 135,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 371,
+ "bbox": [
+ 186,
+ 472,
+ 19,
+ 9
+ ],
+ "category_id": 4,
+ "id": 690,
+ "area": 171,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 372,
+ "bbox": [
+ 56,
+ 128,
+ 432,
+ 222
+ ],
+ "category_id": 10,
+ "id": 691,
+ "area": 95904,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 373,
+ "bbox": [
+ 424,
+ 35,
+ 21,
+ 21
+ ],
+ "category_id": 4,
+ "id": 692,
+ "area": 441,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 373,
+ "bbox": [
+ 64,
+ 299,
+ 17,
+ 15
+ ],
+ "category_id": 4,
+ "id": 693,
+ "area": 255,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 373,
+ "bbox": [
+ 389,
+ 230,
+ 11,
+ 25
+ ],
+ "category_id": 4,
+ "id": 694,
+ "area": 275,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 373,
+ "bbox": [
+ 391,
+ 364,
+ 17,
+ 26
+ ],
+ "category_id": 4,
+ "id": 695,
+ "area": 442,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 373,
+ "bbox": [
+ 466,
+ 449,
+ 15,
+ 31
+ ],
+ "category_id": 4,
+ "id": 696,
+ "area": 465,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 373,
+ "bbox": [
+ 236,
+ 456,
+ 19,
+ 25
+ ],
+ "category_id": 4,
+ "id": 697,
+ "area": 475,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 374,
+ "bbox": [
+ 226,
+ 129,
+ 72,
+ 189
+ ],
+ "category_id": 16,
+ "id": 698,
+ "area": 13608,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 374,
+ "bbox": [
+ 248,
+ 327,
+ 84,
+ 139
+ ],
+ "category_id": 16,
+ "id": 699,
+ "area": 11676,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 375,
+ "bbox": [
+ 110,
+ 83,
+ 128,
+ 118
+ ],
+ "category_id": 10,
+ "id": 700,
+ "area": 15104,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 376,
+ "bbox": [
+ 306,
+ 116,
+ 6,
+ 14
+ ],
+ "category_id": 4,
+ "id": 701,
+ "area": 84,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 376,
+ "bbox": [
+ 392,
+ 133,
+ 8,
+ 11
+ ],
+ "category_id": 4,
+ "id": 702,
+ "area": 88,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 376,
+ "bbox": [
+ 417,
+ 186,
+ 8,
+ 11
+ ],
+ "category_id": 4,
+ "id": 703,
+ "area": 88,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 376,
+ "bbox": [
+ 108,
+ 221,
+ 6,
+ 7
+ ],
+ "category_id": 4,
+ "id": 704,
+ "area": 42,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 376,
+ "bbox": [
+ 151,
+ 432,
+ 7,
+ 12
+ ],
+ "category_id": 4,
+ "id": 705,
+ "area": 84,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 376,
+ "bbox": [
+ 70,
+ 392,
+ 10,
+ 12
+ ],
+ "category_id": 4,
+ "id": 706,
+ "area": 120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 376,
+ "bbox": [
+ 37,
+ 456,
+ 7,
+ 14
+ ],
+ "category_id": 4,
+ "id": 707,
+ "area": 98,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 376,
+ "bbox": [
+ 204,
+ 106,
+ 8,
+ 6
+ ],
+ "category_id": 4,
+ "id": 708,
+ "area": 48,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 377,
+ "bbox": [
+ 52,
+ 89,
+ 22,
+ 26
+ ],
+ "category_id": 4,
+ "id": 709,
+ "area": 572,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 377,
+ "bbox": [
+ 139,
+ 305,
+ 17,
+ 25
+ ],
+ "category_id": 4,
+ "id": 710,
+ "area": 425,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 377,
+ "bbox": [
+ 219,
+ 103,
+ 19,
+ 23
+ ],
+ "category_id": 4,
+ "id": 711,
+ "area": 437,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 377,
+ "bbox": [
+ 369,
+ 130,
+ 23,
+ 31
+ ],
+ "category_id": 4,
+ "id": 712,
+ "area": 713,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 377,
+ "bbox": [
+ 414,
+ 268,
+ 19,
+ 28
+ ],
+ "category_id": 4,
+ "id": 713,
+ "area": 532,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 377,
+ "bbox": [
+ 274,
+ 398,
+ 23,
+ 28
+ ],
+ "category_id": 4,
+ "id": 714,
+ "area": 644,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 378,
+ "bbox": [
+ 297,
+ 90,
+ 57,
+ 333
+ ],
+ "category_id": 10,
+ "id": 715,
+ "area": 18981,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 379,
+ "bbox": [
+ 39,
+ 125,
+ 455,
+ 153
+ ],
+ "category_id": 10,
+ "id": 716,
+ "area": 69615,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 380,
+ "bbox": [
+ 90,
+ 104,
+ 98,
+ 61
+ ],
+ "category_id": 15,
+ "id": 717,
+ "area": 5978,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 380,
+ "bbox": [
+ 209,
+ 184,
+ 98,
+ 62
+ ],
+ "category_id": 15,
+ "id": 718,
+ "area": 6076,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 380,
+ "bbox": [
+ 326,
+ 254,
+ 102,
+ 60
+ ],
+ "category_id": 15,
+ "id": 719,
+ "area": 6120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 381,
+ "bbox": [
+ 247,
+ 216,
+ 7,
+ 9
+ ],
+ "category_id": 4,
+ "id": 720,
+ "area": 63,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 381,
+ "bbox": [
+ 227,
+ 259,
+ 5,
+ 8
+ ],
+ "category_id": 4,
+ "id": 721,
+ "area": 40,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 381,
+ "bbox": [
+ 184,
+ 382,
+ 8,
+ 7
+ ],
+ "category_id": 4,
+ "id": 722,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 381,
+ "bbox": [
+ 172,
+ 427,
+ 5,
+ 8
+ ],
+ "category_id": 4,
+ "id": 723,
+ "area": 40,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 381,
+ "bbox": [
+ 171,
+ 477,
+ 3,
+ 7
+ ],
+ "category_id": 4,
+ "id": 724,
+ "area": 21,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 382,
+ "bbox": [
+ 108,
+ 172,
+ 32,
+ 31
+ ],
+ "category_id": 4,
+ "id": 725,
+ "area": 992,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 382,
+ "bbox": [
+ 232,
+ 386,
+ 32,
+ 38
+ ],
+ "category_id": 4,
+ "id": 726,
+ "area": 1216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 383,
+ "bbox": [
+ 71,
+ 164,
+ 389,
+ 95
+ ],
+ "category_id": 10,
+ "id": 727,
+ "area": 36955,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 384,
+ "bbox": [
+ 308,
+ 94,
+ 100,
+ 315
+ ],
+ "category_id": 10,
+ "id": 728,
+ "area": 31500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 385,
+ "bbox": [
+ 102,
+ 253,
+ 109,
+ 117
+ ],
+ "category_id": 15,
+ "id": 729,
+ "area": 12753,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 385,
+ "bbox": [
+ 303,
+ 319,
+ 97,
+ 105
+ ],
+ "category_id": 15,
+ "id": 730,
+ "area": 10185,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 386,
+ "bbox": [
+ 65,
+ 133,
+ 376,
+ 262
+ ],
+ "category_id": 16,
+ "id": 731,
+ "area": 98512,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 387,
+ "bbox": [
+ 113,
+ 152,
+ 258,
+ 324
+ ],
+ "category_id": 16,
+ "id": 732,
+ "area": 83592,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 388,
+ "bbox": [
+ 115,
+ 30,
+ 21,
+ 24
+ ],
+ "category_id": 4,
+ "id": 733,
+ "area": 504,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 388,
+ "bbox": [
+ 97,
+ 175,
+ 16,
+ 26
+ ],
+ "category_id": 4,
+ "id": 734,
+ "area": 416,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 388,
+ "bbox": [
+ 283,
+ 47,
+ 20,
+ 27
+ ],
+ "category_id": 4,
+ "id": 735,
+ "area": 540,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 388,
+ "bbox": [
+ 412,
+ 255,
+ 17,
+ 30
+ ],
+ "category_id": 4,
+ "id": 736,
+ "area": 510,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 388,
+ "bbox": [
+ 411,
+ 456,
+ 13,
+ 20
+ ],
+ "category_id": 4,
+ "id": 737,
+ "area": 260,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 389,
+ "bbox": [
+ 180,
+ 101,
+ 116,
+ 284
+ ],
+ "category_id": 19,
+ "id": 738,
+ "area": 32944,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 390,
+ "bbox": [
+ 327,
+ 56,
+ 12,
+ 35
+ ],
+ "category_id": 4,
+ "id": 739,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 390,
+ "bbox": [
+ 336,
+ 240,
+ 18,
+ 31
+ ],
+ "category_id": 4,
+ "id": 740,
+ "area": 558,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 390,
+ "bbox": [
+ 449,
+ 353,
+ 13,
+ 36
+ ],
+ "category_id": 4,
+ "id": 741,
+ "area": 468,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 390,
+ "bbox": [
+ 133,
+ 373,
+ 23,
+ 35
+ ],
+ "category_id": 4,
+ "id": 742,
+ "area": 805,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 391,
+ "bbox": [
+ 0,
+ 403,
+ 19,
+ 19
+ ],
+ "category_id": 15,
+ "id": 743,
+ "area": 361,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 392,
+ "bbox": [
+ 325,
+ 245,
+ 18,
+ 12
+ ],
+ "category_id": 15,
+ "id": 744,
+ "area": 216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 393,
+ "bbox": [
+ 212,
+ 108,
+ 47,
+ 244
+ ],
+ "category_id": 19,
+ "id": 745,
+ "area": 11468,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 394,
+ "bbox": [
+ 111,
+ 5,
+ 15,
+ 9
+ ],
+ "category_id": 4,
+ "id": 746,
+ "area": 135,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 394,
+ "bbox": [
+ 137,
+ 82,
+ 16,
+ 10
+ ],
+ "category_id": 4,
+ "id": 747,
+ "area": 160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 394,
+ "bbox": [
+ 60,
+ 125,
+ 15,
+ 12
+ ],
+ "category_id": 4,
+ "id": 748,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 394,
+ "bbox": [
+ 55,
+ 253,
+ 15,
+ 11
+ ],
+ "category_id": 4,
+ "id": 749,
+ "area": 165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 394,
+ "bbox": [
+ 79,
+ 328,
+ 18,
+ 15
+ ],
+ "category_id": 4,
+ "id": 750,
+ "area": 270,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 394,
+ "bbox": [
+ 385,
+ 39,
+ 16,
+ 12
+ ],
+ "category_id": 4,
+ "id": 751,
+ "area": 192,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 394,
+ "bbox": [
+ 488,
+ 5,
+ 17,
+ 12
+ ],
+ "category_id": 4,
+ "id": 752,
+ "area": 204,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 394,
+ "bbox": [
+ 368,
+ 129,
+ 19,
+ 13
+ ],
+ "category_id": 4,
+ "id": 753,
+ "area": 247,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 394,
+ "bbox": [
+ 205,
+ 192,
+ 17,
+ 15
+ ],
+ "category_id": 4,
+ "id": 754,
+ "area": 255,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 394,
+ "bbox": [
+ 259,
+ 298,
+ 20,
+ 14
+ ],
+ "category_id": 4,
+ "id": 755,
+ "area": 280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 394,
+ "bbox": [
+ 426,
+ 225,
+ 18,
+ 13
+ ],
+ "category_id": 4,
+ "id": 756,
+ "area": 234,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 394,
+ "bbox": [
+ 335,
+ 339,
+ 23,
+ 13
+ ],
+ "category_id": 4,
+ "id": 757,
+ "area": 299,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 394,
+ "bbox": [
+ 350,
+ 384,
+ 18,
+ 11
+ ],
+ "category_id": 4,
+ "id": 758,
+ "area": 198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 394,
+ "bbox": [
+ 480,
+ 387,
+ 18,
+ 12
+ ],
+ "category_id": 4,
+ "id": 759,
+ "area": 216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 395,
+ "bbox": [
+ 116,
+ 135,
+ 19,
+ 30
+ ],
+ "category_id": 4,
+ "id": 760,
+ "area": 570,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 395,
+ "bbox": [
+ 293,
+ 424,
+ 14,
+ 25
+ ],
+ "category_id": 4,
+ "id": 761,
+ "area": 350,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 395,
+ "bbox": [
+ 165,
+ 265,
+ 18,
+ 28
+ ],
+ "category_id": 4,
+ "id": 762,
+ "area": 504,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 395,
+ "bbox": [
+ 379,
+ 224,
+ 17,
+ 26
+ ],
+ "category_id": 4,
+ "id": 763,
+ "area": 442,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 396,
+ "bbox": [
+ 131,
+ 81,
+ 75,
+ 354
+ ],
+ "category_id": 19,
+ "id": 764,
+ "area": 26550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 397,
+ "bbox": [
+ 88,
+ 388,
+ 48,
+ 45
+ ],
+ "category_id": 4,
+ "id": 765,
+ "area": 2160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 397,
+ "bbox": [
+ 391,
+ 177,
+ 51,
+ 58
+ ],
+ "category_id": 4,
+ "id": 766,
+ "area": 2958,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 398,
+ "bbox": [
+ 165,
+ 168,
+ 151,
+ 182
+ ],
+ "category_id": 19,
+ "id": 767,
+ "area": 27482,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 399,
+ "bbox": [
+ 102,
+ 43,
+ 17,
+ 22
+ ],
+ "category_id": 4,
+ "id": 768,
+ "area": 374,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 399,
+ "bbox": [
+ 183,
+ 159,
+ 17,
+ 24
+ ],
+ "category_id": 4,
+ "id": 769,
+ "area": 408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 399,
+ "bbox": [
+ 299,
+ 189,
+ 29,
+ 23
+ ],
+ "category_id": 4,
+ "id": 770,
+ "area": 667,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 399,
+ "bbox": [
+ 353,
+ 349,
+ 22,
+ 16
+ ],
+ "category_id": 4,
+ "id": 771,
+ "area": 352,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 399,
+ "bbox": [
+ 364,
+ 472,
+ 17,
+ 16
+ ],
+ "category_id": 4,
+ "id": 772,
+ "area": 272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 399,
+ "bbox": [
+ 102,
+ 340,
+ 28,
+ 17
+ ],
+ "category_id": 4,
+ "id": 773,
+ "area": 476,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 400,
+ "bbox": [
+ 268,
+ 126,
+ 48,
+ 308
+ ],
+ "category_id": 16,
+ "id": 774,
+ "area": 14784,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 401,
+ "bbox": [
+ 453,
+ 97,
+ 25,
+ 25
+ ],
+ "category_id": 4,
+ "id": 775,
+ "area": 625,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 401,
+ "bbox": [
+ 272,
+ 145,
+ 22,
+ 20
+ ],
+ "category_id": 4,
+ "id": 776,
+ "area": 440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 401,
+ "bbox": [
+ 60,
+ 293,
+ 28,
+ 15
+ ],
+ "category_id": 4,
+ "id": 777,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 401,
+ "bbox": [
+ 31,
+ 401,
+ 27,
+ 23
+ ],
+ "category_id": 4,
+ "id": 778,
+ "area": 621,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 401,
+ "bbox": [
+ 251,
+ 273,
+ 31,
+ 20
+ ],
+ "category_id": 4,
+ "id": 779,
+ "area": 620,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 402,
+ "bbox": [
+ 133,
+ 141,
+ 113,
+ 124
+ ],
+ "category_id": 15,
+ "id": 780,
+ "area": 14012,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 403,
+ "bbox": [
+ 234,
+ 147,
+ 6,
+ 10
+ ],
+ "category_id": 4,
+ "id": 781,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 403,
+ "bbox": [
+ 216,
+ 203,
+ 6,
+ 10
+ ],
+ "category_id": 4,
+ "id": 782,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 403,
+ "bbox": [
+ 375,
+ 340,
+ 7,
+ 10
+ ],
+ "category_id": 4,
+ "id": 783,
+ "area": 70,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 403,
+ "bbox": [
+ 300,
+ 412,
+ 5,
+ 12
+ ],
+ "category_id": 4,
+ "id": 784,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 404,
+ "bbox": [
+ 159,
+ 240,
+ 159,
+ 118
+ ],
+ "category_id": 4,
+ "id": 785,
+ "area": 18762,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 405,
+ "bbox": [
+ 33,
+ 239,
+ 167,
+ 44
+ ],
+ "category_id": 19,
+ "id": 786,
+ "area": 7348,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 405,
+ "bbox": [
+ 190,
+ 160,
+ 128,
+ 79
+ ],
+ "category_id": 19,
+ "id": 787,
+ "area": 10112,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 406,
+ "bbox": [
+ 322,
+ 140,
+ 76,
+ 107
+ ],
+ "category_id": 15,
+ "id": 788,
+ "area": 8132,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 406,
+ "bbox": [
+ 316,
+ 261,
+ 83,
+ 114
+ ],
+ "category_id": 15,
+ "id": 789,
+ "area": 9462,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 407,
+ "bbox": [
+ 377,
+ 313,
+ 114,
+ 83
+ ],
+ "category_id": 15,
+ "id": 790,
+ "area": 9462,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 407,
+ "bbox": [
+ 262,
+ 300,
+ 97,
+ 77
+ ],
+ "category_id": 15,
+ "id": 791,
+ "area": 7469,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 407,
+ "bbox": [
+ 148,
+ 275,
+ 102,
+ 82
+ ],
+ "category_id": 15,
+ "id": 792,
+ "area": 8364,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 408,
+ "bbox": [
+ 66,
+ 305,
+ 86,
+ 52
+ ],
+ "category_id": 10,
+ "id": 793,
+ "area": 4472,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 409,
+ "bbox": [
+ 104,
+ 22,
+ 17,
+ 8
+ ],
+ "category_id": 4,
+ "id": 794,
+ "area": 136,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 409,
+ "bbox": [
+ 12,
+ 103,
+ 28,
+ 16
+ ],
+ "category_id": 4,
+ "id": 795,
+ "area": 448,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 409,
+ "bbox": [
+ 288,
+ 112,
+ 16,
+ 10
+ ],
+ "category_id": 4,
+ "id": 796,
+ "area": 160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 409,
+ "bbox": [
+ 165,
+ 189,
+ 12,
+ 7
+ ],
+ "category_id": 4,
+ "id": 797,
+ "area": 84,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 409,
+ "bbox": [
+ 45,
+ 306,
+ 19,
+ 12
+ ],
+ "category_id": 4,
+ "id": 798,
+ "area": 228,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 409,
+ "bbox": [
+ 371,
+ 265,
+ 17,
+ 10
+ ],
+ "category_id": 4,
+ "id": 799,
+ "area": 170,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 409,
+ "bbox": [
+ 82,
+ 481,
+ 31,
+ 12
+ ],
+ "category_id": 4,
+ "id": 800,
+ "area": 372,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 409,
+ "bbox": [
+ 279,
+ 492,
+ 24,
+ 10
+ ],
+ "category_id": 4,
+ "id": 801,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 409,
+ "bbox": [
+ 378,
+ 451,
+ 21,
+ 9
+ ],
+ "category_id": 4,
+ "id": 802,
+ "area": 189,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 410,
+ "bbox": [
+ 293,
+ 164,
+ 8,
+ 6
+ ],
+ "category_id": 4,
+ "id": 803,
+ "area": 48,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 410,
+ "bbox": [
+ 350,
+ 200,
+ 3,
+ 8
+ ],
+ "category_id": 4,
+ "id": 804,
+ "area": 24,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 410,
+ "bbox": [
+ 456,
+ 147,
+ 7,
+ 8
+ ],
+ "category_id": 4,
+ "id": 805,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 410,
+ "bbox": [
+ 475,
+ 200,
+ 6,
+ 7
+ ],
+ "category_id": 4,
+ "id": 806,
+ "area": 42,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 410,
+ "bbox": [
+ 482,
+ 278,
+ 5,
+ 10
+ ],
+ "category_id": 4,
+ "id": 807,
+ "area": 50,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 410,
+ "bbox": [
+ 221,
+ 328,
+ 8,
+ 7
+ ],
+ "category_id": 4,
+ "id": 808,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 410,
+ "bbox": [
+ 169,
+ 364,
+ 7,
+ 9
+ ],
+ "category_id": 4,
+ "id": 809,
+ "area": 63,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 411,
+ "bbox": [
+ 172,
+ 235,
+ 145,
+ 153
+ ],
+ "category_id": 4,
+ "id": 810,
+ "area": 22185,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 412,
+ "bbox": [
+ 28,
+ 56,
+ 238,
+ 121
+ ],
+ "category_id": 19,
+ "id": 811,
+ "area": 28798,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 413,
+ "bbox": [
+ 88,
+ 124,
+ 26,
+ 16
+ ],
+ "category_id": 4,
+ "id": 812,
+ "area": 416,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 413,
+ "bbox": [
+ 243,
+ 197,
+ 25,
+ 26
+ ],
+ "category_id": 4,
+ "id": 813,
+ "area": 650,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 413,
+ "bbox": [
+ 56,
+ 385,
+ 24,
+ 23
+ ],
+ "category_id": 4,
+ "id": 814,
+ "area": 552,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 413,
+ "bbox": [
+ 168,
+ 346,
+ 25,
+ 33
+ ],
+ "category_id": 4,
+ "id": 815,
+ "area": 825,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 413,
+ "bbox": [
+ 408,
+ 332,
+ 30,
+ 26
+ ],
+ "category_id": 4,
+ "id": 816,
+ "area": 780,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 414,
+ "bbox": [
+ 176,
+ 170,
+ 183,
+ 207
+ ],
+ "category_id": 16,
+ "id": 817,
+ "area": 37881,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 415,
+ "bbox": [
+ 49,
+ 245,
+ 15,
+ 7
+ ],
+ "category_id": 16,
+ "id": 818,
+ "area": 105,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 416,
+ "bbox": [
+ 63,
+ 194,
+ 283,
+ 210
+ ],
+ "category_id": 16,
+ "id": 819,
+ "area": 59430,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 417,
+ "bbox": [
+ 178,
+ 429,
+ 40,
+ 25
+ ],
+ "category_id": 4,
+ "id": 820,
+ "area": 1000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 417,
+ "bbox": [
+ 270,
+ 148,
+ 37,
+ 24
+ ],
+ "category_id": 4,
+ "id": 821,
+ "area": 888,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 418,
+ "bbox": [
+ 233,
+ 173,
+ 103,
+ 245
+ ],
+ "category_id": 16,
+ "id": 822,
+ "area": 25235,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 419,
+ "bbox": [
+ 192,
+ 323,
+ 108,
+ 112
+ ],
+ "category_id": 15,
+ "id": 823,
+ "area": 12096,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 420,
+ "bbox": [
+ 425,
+ 204,
+ 51,
+ 100
+ ],
+ "category_id": 10,
+ "id": 824,
+ "area": 5100,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 421,
+ "bbox": [
+ 281,
+ 327,
+ 10,
+ 13
+ ],
+ "category_id": 4,
+ "id": 825,
+ "area": 130,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 421,
+ "bbox": [
+ 398,
+ 296,
+ 9,
+ 15
+ ],
+ "category_id": 4,
+ "id": 826,
+ "area": 135,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 421,
+ "bbox": [
+ 186,
+ 395,
+ 10,
+ 15
+ ],
+ "category_id": 4,
+ "id": 827,
+ "area": 150,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 421,
+ "bbox": [
+ 62,
+ 460,
+ 11,
+ 12
+ ],
+ "category_id": 4,
+ "id": 828,
+ "area": 132,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 422,
+ "bbox": [
+ 74,
+ 204,
+ 384,
+ 90
+ ],
+ "category_id": 10,
+ "id": 829,
+ "area": 34560,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 423,
+ "bbox": [
+ 16,
+ 125,
+ 362,
+ 168
+ ],
+ "category_id": 19,
+ "id": 830,
+ "area": 60816,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 424,
+ "bbox": [
+ 68,
+ 295,
+ 35,
+ 42
+ ],
+ "category_id": 4,
+ "id": 831,
+ "area": 1470,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 424,
+ "bbox": [
+ 242,
+ 103,
+ 43,
+ 49
+ ],
+ "category_id": 4,
+ "id": 832,
+ "area": 2107,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 425,
+ "bbox": [
+ 280,
+ 263,
+ 98,
+ 126
+ ],
+ "category_id": 4,
+ "id": 833,
+ "area": 12348,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 426,
+ "bbox": [
+ 197,
+ 99,
+ 167,
+ 332
+ ],
+ "category_id": 10,
+ "id": 834,
+ "area": 55444,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 427,
+ "bbox": [
+ 99,
+ 64,
+ 19,
+ 62
+ ],
+ "category_id": 4,
+ "id": 835,
+ "area": 1178,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 427,
+ "bbox": [
+ 297,
+ 342,
+ 17,
+ 61
+ ],
+ "category_id": 4,
+ "id": 836,
+ "area": 1037,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 428,
+ "bbox": [
+ 190,
+ 124,
+ 128,
+ 148
+ ],
+ "category_id": 15,
+ "id": 837,
+ "area": 18944,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 429,
+ "bbox": [
+ 59,
+ 128,
+ 339,
+ 251
+ ],
+ "category_id": 10,
+ "id": 838,
+ "area": 85089,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 430,
+ "bbox": [
+ 162,
+ 142,
+ 166,
+ 221
+ ],
+ "category_id": 19,
+ "id": 839,
+ "area": 36686,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 431,
+ "bbox": [
+ 147,
+ 133,
+ 67,
+ 85
+ ],
+ "category_id": 15,
+ "id": 840,
+ "area": 5695,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 431,
+ "bbox": [
+ 229,
+ 179,
+ 66,
+ 77
+ ],
+ "category_id": 15,
+ "id": 841,
+ "area": 5082,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 431,
+ "bbox": [
+ 295,
+ 242,
+ 69,
+ 74
+ ],
+ "category_id": 15,
+ "id": 842,
+ "area": 5106,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 431,
+ "bbox": [
+ 374,
+ 288,
+ 70,
+ 80
+ ],
+ "category_id": 15,
+ "id": 843,
+ "area": 5600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 431,
+ "bbox": [
+ 290,
+ 126,
+ 33,
+ 73
+ ],
+ "category_id": 15,
+ "id": 844,
+ "area": 2409,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 431,
+ "bbox": [
+ 408,
+ 210,
+ 26,
+ 55
+ ],
+ "category_id": 15,
+ "id": 845,
+ "area": 1430,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 432,
+ "bbox": [
+ 189,
+ 226,
+ 166,
+ 197
+ ],
+ "category_id": 16,
+ "id": 846,
+ "area": 32702,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 433,
+ "bbox": [
+ 133,
+ 209,
+ 174,
+ 91
+ ],
+ "category_id": 19,
+ "id": 847,
+ "area": 15834,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 434,
+ "bbox": [
+ 39,
+ 293,
+ 71,
+ 87
+ ],
+ "category_id": 10,
+ "id": 848,
+ "area": 6177,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 435,
+ "bbox": [
+ 14,
+ 146,
+ 18,
+ 19
+ ],
+ "category_id": 4,
+ "id": 849,
+ "area": 342,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 435,
+ "bbox": [
+ 340,
+ 134,
+ 21,
+ 15
+ ],
+ "category_id": 4,
+ "id": 850,
+ "area": 315,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 435,
+ "bbox": [
+ 455,
+ 1,
+ 14,
+ 20
+ ],
+ "category_id": 4,
+ "id": 851,
+ "area": 280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 435,
+ "bbox": [
+ 275,
+ 241,
+ 18,
+ 15
+ ],
+ "category_id": 4,
+ "id": 852,
+ "area": 270,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 435,
+ "bbox": [
+ 471,
+ 275,
+ 20,
+ 22
+ ],
+ "category_id": 4,
+ "id": 853,
+ "area": 440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 435,
+ "bbox": [
+ 135,
+ 407,
+ 21,
+ 18
+ ],
+ "category_id": 4,
+ "id": 854,
+ "area": 378,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 435,
+ "bbox": [
+ 383,
+ 408,
+ 19,
+ 18
+ ],
+ "category_id": 4,
+ "id": 855,
+ "area": 342,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 436,
+ "bbox": [
+ 48,
+ 184,
+ 27,
+ 45
+ ],
+ "category_id": 4,
+ "id": 856,
+ "area": 1215,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 436,
+ "bbox": [
+ 414,
+ 380,
+ 26,
+ 57
+ ],
+ "category_id": 4,
+ "id": 857,
+ "area": 1482,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 437,
+ "bbox": [
+ 147,
+ 129,
+ 73,
+ 252
+ ],
+ "category_id": 19,
+ "id": 858,
+ "area": 18396,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 438,
+ "bbox": [
+ 30,
+ 206,
+ 450,
+ 153
+ ],
+ "category_id": 10,
+ "id": 859,
+ "area": 68850,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 439,
+ "bbox": [
+ 55,
+ 339,
+ 46,
+ 61
+ ],
+ "category_id": 4,
+ "id": 860,
+ "area": 2806,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 439,
+ "bbox": [
+ 405,
+ 200,
+ 64,
+ 47
+ ],
+ "category_id": 4,
+ "id": 861,
+ "area": 3008,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 440,
+ "bbox": [
+ 218,
+ 250,
+ 116,
+ 119
+ ],
+ "category_id": 15,
+ "id": 862,
+ "area": 13804,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 441,
+ "bbox": [
+ 226,
+ 242,
+ 114,
+ 98
+ ],
+ "category_id": 15,
+ "id": 863,
+ "area": 11172,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 441,
+ "bbox": [
+ 144,
+ 404,
+ 89,
+ 63
+ ],
+ "category_id": 15,
+ "id": 864,
+ "area": 5607,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 442,
+ "bbox": [
+ 59,
+ 136,
+ 58,
+ 40
+ ],
+ "category_id": 4,
+ "id": 865,
+ "area": 2320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 442,
+ "bbox": [
+ 302,
+ 128,
+ 44,
+ 51
+ ],
+ "category_id": 4,
+ "id": 866,
+ "area": 2244,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 442,
+ "bbox": [
+ 320,
+ 360,
+ 73,
+ 31
+ ],
+ "category_id": 4,
+ "id": 867,
+ "area": 2263,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 443,
+ "bbox": [
+ 60,
+ 134,
+ 46,
+ 67
+ ],
+ "category_id": 16,
+ "id": 868,
+ "area": 3082,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 444,
+ "bbox": [
+ 116,
+ 251,
+ 139,
+ 56
+ ],
+ "category_id": 19,
+ "id": 869,
+ "area": 7784,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 445,
+ "bbox": [
+ 259,
+ 87,
+ 23,
+ 18
+ ],
+ "category_id": 4,
+ "id": 870,
+ "area": 414,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 445,
+ "bbox": [
+ 119,
+ 268,
+ 21,
+ 14
+ ],
+ "category_id": 4,
+ "id": 871,
+ "area": 294,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 445,
+ "bbox": [
+ 247,
+ 391,
+ 16,
+ 13
+ ],
+ "category_id": 4,
+ "id": 872,
+ "area": 208,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 445,
+ "bbox": [
+ 432,
+ 193,
+ 24,
+ 19
+ ],
+ "category_id": 4,
+ "id": 873,
+ "area": 456,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 446,
+ "bbox": [
+ 192,
+ 135,
+ 116,
+ 149
+ ],
+ "category_id": 15,
+ "id": 874,
+ "area": 17284,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 446,
+ "bbox": [
+ 257,
+ 308,
+ 119,
+ 139
+ ],
+ "category_id": 15,
+ "id": 875,
+ "area": 16541,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 447,
+ "bbox": [
+ 77,
+ 71,
+ 43,
+ 96
+ ],
+ "category_id": 10,
+ "id": 876,
+ "area": 4128,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 447,
+ "bbox": [
+ 192,
+ 97,
+ 48,
+ 119
+ ],
+ "category_id": 10,
+ "id": 877,
+ "area": 5712,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 448,
+ "bbox": [
+ 115,
+ 0,
+ 20,
+ 33
+ ],
+ "category_id": 15,
+ "id": 878,
+ "area": 660,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 448,
+ "bbox": [
+ 141,
+ 224,
+ 111,
+ 148
+ ],
+ "category_id": 15,
+ "id": 879,
+ "area": 16428,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 448,
+ "bbox": [
+ 337,
+ 96,
+ 111,
+ 141
+ ],
+ "category_id": 15,
+ "id": 880,
+ "area": 15651,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 448,
+ "bbox": [
+ 288,
+ 262,
+ 112,
+ 151
+ ],
+ "category_id": 15,
+ "id": 881,
+ "area": 16912,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 448,
+ "bbox": [
+ 258,
+ 0,
+ 17,
+ 37
+ ],
+ "category_id": 15,
+ "id": 882,
+ "area": 629,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 449,
+ "bbox": [
+ 102,
+ 135,
+ 254,
+ 278
+ ],
+ "category_id": 10,
+ "id": 883,
+ "area": 70612,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 450,
+ "bbox": [
+ 246,
+ 269,
+ 88,
+ 112
+ ],
+ "category_id": 4,
+ "id": 884,
+ "area": 9856,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 451,
+ "bbox": [
+ 216,
+ 104,
+ 52,
+ 287
+ ],
+ "category_id": 19,
+ "id": 885,
+ "area": 14924,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 452,
+ "bbox": [
+ 245,
+ 264,
+ 85,
+ 73
+ ],
+ "category_id": 4,
+ "id": 886,
+ "area": 6205,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 453,
+ "bbox": [
+ 81,
+ 201,
+ 196,
+ 234
+ ],
+ "category_id": 16,
+ "id": 887,
+ "area": 45864,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 454,
+ "bbox": [
+ 140,
+ 345,
+ 58,
+ 73
+ ],
+ "category_id": 15,
+ "id": 888,
+ "area": 4234,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 454,
+ "bbox": [
+ 194,
+ 293,
+ 58,
+ 70
+ ],
+ "category_id": 15,
+ "id": 889,
+ "area": 4060,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 454,
+ "bbox": [
+ 255,
+ 231,
+ 57,
+ 72
+ ],
+ "category_id": 15,
+ "id": 890,
+ "area": 4104,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 454,
+ "bbox": [
+ 310,
+ 173,
+ 58,
+ 73
+ ],
+ "category_id": 15,
+ "id": 891,
+ "area": 4234,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 454,
+ "bbox": [
+ 86,
+ 139,
+ 26,
+ 73
+ ],
+ "category_id": 15,
+ "id": 892,
+ "area": 1898,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 454,
+ "bbox": [
+ 158,
+ 67,
+ 26,
+ 71
+ ],
+ "category_id": 15,
+ "id": 893,
+ "area": 1846,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 455,
+ "bbox": [
+ 189,
+ 186,
+ 221,
+ 162
+ ],
+ "category_id": 16,
+ "id": 894,
+ "area": 35802,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 456,
+ "bbox": [
+ 18,
+ 135,
+ 42,
+ 249
+ ],
+ "category_id": 19,
+ "id": 895,
+ "area": 10458,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 457,
+ "bbox": [
+ 37,
+ 308,
+ 42,
+ 51
+ ],
+ "category_id": 4,
+ "id": 896,
+ "area": 2142,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 457,
+ "bbox": [
+ 184,
+ 140,
+ 34,
+ 60
+ ],
+ "category_id": 4,
+ "id": 897,
+ "area": 2040,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 457,
+ "bbox": [
+ 432,
+ 318,
+ 53,
+ 47
+ ],
+ "category_id": 4,
+ "id": 898,
+ "area": 2491,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 458,
+ "bbox": [
+ 8,
+ 287,
+ 18,
+ 31
+ ],
+ "category_id": 4,
+ "id": 899,
+ "area": 558,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 458,
+ "bbox": [
+ 110,
+ 38,
+ 15,
+ 18
+ ],
+ "category_id": 4,
+ "id": 900,
+ "area": 270,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 458,
+ "bbox": [
+ 256,
+ 307,
+ 19,
+ 25
+ ],
+ "category_id": 4,
+ "id": 901,
+ "area": 475,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 458,
+ "bbox": [
+ 487,
+ 397,
+ 20,
+ 26
+ ],
+ "category_id": 4,
+ "id": 902,
+ "area": 520,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 458,
+ "bbox": [
+ 417,
+ 126,
+ 22,
+ 25
+ ],
+ "category_id": 4,
+ "id": 903,
+ "area": 550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 459,
+ "bbox": [
+ 59,
+ 268,
+ 86,
+ 74
+ ],
+ "category_id": 4,
+ "id": 904,
+ "area": 6364,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 459,
+ "bbox": [
+ 389,
+ 124,
+ 70,
+ 48
+ ],
+ "category_id": 4,
+ "id": 905,
+ "area": 3360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 460,
+ "bbox": [
+ 244,
+ 372,
+ 82,
+ 46
+ ],
+ "category_id": 16,
+ "id": 906,
+ "area": 3772,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 461,
+ "bbox": [
+ 325,
+ 131,
+ 46,
+ 160
+ ],
+ "category_id": 19,
+ "id": 907,
+ "area": 7360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 462,
+ "bbox": [
+ 419,
+ 61,
+ 65,
+ 209
+ ],
+ "category_id": 10,
+ "id": 908,
+ "area": 13585,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 463,
+ "bbox": [
+ 135,
+ 224,
+ 219,
+ 109
+ ],
+ "category_id": 16,
+ "id": 909,
+ "area": 23871,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 464,
+ "bbox": [
+ 293,
+ 73,
+ 178,
+ 379
+ ],
+ "category_id": 10,
+ "id": 910,
+ "area": 67462,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 464,
+ "bbox": [
+ 140,
+ 344,
+ 171,
+ 68
+ ],
+ "category_id": 10,
+ "id": 911,
+ "area": 11628,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 465,
+ "bbox": [
+ 73,
+ 202,
+ 51,
+ 42
+ ],
+ "category_id": 4,
+ "id": 912,
+ "area": 2142,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 465,
+ "bbox": [
+ 432,
+ 172,
+ 51,
+ 43
+ ],
+ "category_id": 4,
+ "id": 913,
+ "area": 2193,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 466,
+ "bbox": [
+ 83,
+ 37,
+ 199,
+ 199
+ ],
+ "category_id": 15,
+ "id": 914,
+ "area": 39601,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 466,
+ "bbox": [
+ 40,
+ 250,
+ 195,
+ 201
+ ],
+ "category_id": 15,
+ "id": 915,
+ "area": 39195,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 466,
+ "bbox": [
+ 307,
+ 46,
+ 193,
+ 199
+ ],
+ "category_id": 15,
+ "id": 916,
+ "area": 38407,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 466,
+ "bbox": [
+ 373,
+ 236,
+ 94,
+ 96
+ ],
+ "category_id": 15,
+ "id": 917,
+ "area": 9024,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 467,
+ "bbox": [
+ 96,
+ 428,
+ 35,
+ 14
+ ],
+ "category_id": 15,
+ "id": 918,
+ "area": 490,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 467,
+ "bbox": [
+ 184,
+ 460,
+ 40,
+ 20
+ ],
+ "category_id": 15,
+ "id": 919,
+ "area": 800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 467,
+ "bbox": [
+ 133,
+ 188,
+ 69,
+ 60
+ ],
+ "category_id": 15,
+ "id": 920,
+ "area": 4140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 467,
+ "bbox": [
+ 308,
+ 240,
+ 81,
+ 67
+ ],
+ "category_id": 15,
+ "id": 921,
+ "area": 5427,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 468,
+ "bbox": [
+ 186,
+ 172,
+ 203,
+ 226
+ ],
+ "category_id": 10,
+ "id": 922,
+ "area": 45878,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 469,
+ "bbox": [
+ 58,
+ 188,
+ 72,
+ 39
+ ],
+ "category_id": 4,
+ "id": 923,
+ "area": 2808,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 469,
+ "bbox": [
+ 336,
+ 140,
+ 61,
+ 60
+ ],
+ "category_id": 4,
+ "id": 924,
+ "area": 3660,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 469,
+ "bbox": [
+ 393,
+ 407,
+ 90,
+ 33
+ ],
+ "category_id": 4,
+ "id": 925,
+ "area": 2970,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 470,
+ "bbox": [
+ 306,
+ 166,
+ 70,
+ 70
+ ],
+ "category_id": 15,
+ "id": 926,
+ "area": 4900,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 471,
+ "bbox": [
+ 155,
+ 254,
+ 82,
+ 25
+ ],
+ "category_id": 19,
+ "id": 927,
+ "area": 2050,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 472,
+ "bbox": [
+ 215,
+ 117,
+ 76,
+ 373
+ ],
+ "category_id": 19,
+ "id": 928,
+ "area": 28348,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 473,
+ "bbox": [
+ 116,
+ 369,
+ 82,
+ 45
+ ],
+ "category_id": 4,
+ "id": 929,
+ "area": 3690,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 473,
+ "bbox": [
+ 314,
+ 69,
+ 82,
+ 44
+ ],
+ "category_id": 4,
+ "id": 930,
+ "area": 3608,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 474,
+ "bbox": [
+ 163,
+ 77,
+ 71,
+ 20
+ ],
+ "category_id": 4,
+ "id": 931,
+ "area": 1420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 474,
+ "bbox": [
+ 353,
+ 380,
+ 57,
+ 25
+ ],
+ "category_id": 4,
+ "id": 932,
+ "area": 1425,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 475,
+ "bbox": [
+ 136,
+ 53,
+ 20,
+ 33
+ ],
+ "category_id": 4,
+ "id": 933,
+ "area": 660,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 475,
+ "bbox": [
+ 79,
+ 240,
+ 12,
+ 33
+ ],
+ "category_id": 4,
+ "id": 934,
+ "area": 396,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 475,
+ "bbox": [
+ 233,
+ 215,
+ 20,
+ 33
+ ],
+ "category_id": 4,
+ "id": 935,
+ "area": 660,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 475,
+ "bbox": [
+ 426,
+ 296,
+ 13,
+ 32
+ ],
+ "category_id": 4,
+ "id": 936,
+ "area": 416,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 475,
+ "bbox": [
+ 90,
+ 391,
+ 18,
+ 30
+ ],
+ "category_id": 4,
+ "id": 937,
+ "area": 540,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 476,
+ "bbox": [
+ 117,
+ 195,
+ 346,
+ 210
+ ],
+ "category_id": 10,
+ "id": 938,
+ "area": 72660,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 477,
+ "bbox": [
+ 179,
+ 146,
+ 115,
+ 83
+ ],
+ "category_id": 16,
+ "id": 939,
+ "area": 9545,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 478,
+ "bbox": [
+ 112,
+ 112,
+ 124,
+ 266
+ ],
+ "category_id": 16,
+ "id": 940,
+ "area": 32984,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 479,
+ "bbox": [
+ 23,
+ 412,
+ 20,
+ 24
+ ],
+ "category_id": 4,
+ "id": 941,
+ "area": 480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 479,
+ "bbox": [
+ 182,
+ 424,
+ 18,
+ 21
+ ],
+ "category_id": 4,
+ "id": 942,
+ "area": 378,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 479,
+ "bbox": [
+ 392,
+ 435,
+ 18,
+ 18
+ ],
+ "category_id": 4,
+ "id": 943,
+ "area": 324,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 479,
+ "bbox": [
+ 451,
+ 467,
+ 17,
+ 21
+ ],
+ "category_id": 4,
+ "id": 944,
+ "area": 357,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 479,
+ "bbox": [
+ 438,
+ 295,
+ 22,
+ 22
+ ],
+ "category_id": 4,
+ "id": 945,
+ "area": 484,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 479,
+ "bbox": [
+ 304,
+ 207,
+ 20,
+ 23
+ ],
+ "category_id": 4,
+ "id": 946,
+ "area": 460,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 480,
+ "bbox": [
+ 14,
+ 189,
+ 19,
+ 15
+ ],
+ "category_id": 4,
+ "id": 947,
+ "area": 285,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 480,
+ "bbox": [
+ 151,
+ 102,
+ 19,
+ 17
+ ],
+ "category_id": 4,
+ "id": 948,
+ "area": 323,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 480,
+ "bbox": [
+ 163,
+ 321,
+ 17,
+ 19
+ ],
+ "category_id": 4,
+ "id": 949,
+ "area": 323,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 480,
+ "bbox": [
+ 359,
+ 8,
+ 17,
+ 16
+ ],
+ "category_id": 4,
+ "id": 950,
+ "area": 272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 480,
+ "bbox": [
+ 462,
+ 161,
+ 14,
+ 11
+ ],
+ "category_id": 4,
+ "id": 951,
+ "area": 154,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 480,
+ "bbox": [
+ 308,
+ 406,
+ 24,
+ 18
+ ],
+ "category_id": 4,
+ "id": 952,
+ "area": 432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 480,
+ "bbox": [
+ 19,
+ 492,
+ 15,
+ 11
+ ],
+ "category_id": 4,
+ "id": 953,
+ "area": 165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 481,
+ "bbox": [
+ 182,
+ 56,
+ 158,
+ 421
+ ],
+ "category_id": 10,
+ "id": 954,
+ "area": 66518,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 482,
+ "bbox": [
+ 181,
+ 232,
+ 162,
+ 75
+ ],
+ "category_id": 10,
+ "id": 955,
+ "area": 12150,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 483,
+ "bbox": [
+ 223,
+ 227,
+ 56,
+ 144
+ ],
+ "category_id": 4,
+ "id": 956,
+ "area": 8064,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 484,
+ "bbox": [
+ 245,
+ 62,
+ 45,
+ 263
+ ],
+ "category_id": 19,
+ "id": 957,
+ "area": 11835,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 485,
+ "bbox": [
+ 145,
+ 171,
+ 194,
+ 187
+ ],
+ "category_id": 19,
+ "id": 958,
+ "area": 36278,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 486,
+ "bbox": [
+ 105,
+ 344,
+ 67,
+ 79
+ ],
+ "category_id": 16,
+ "id": 959,
+ "area": 5293,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 487,
+ "bbox": [
+ 176,
+ 151,
+ 178,
+ 151
+ ],
+ "category_id": 19,
+ "id": 960,
+ "area": 26878,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 488,
+ "bbox": [
+ 62,
+ 229,
+ 390,
+ 89
+ ],
+ "category_id": 19,
+ "id": 961,
+ "area": 34710,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 489,
+ "bbox": [
+ 103,
+ 10,
+ 28,
+ 30
+ ],
+ "category_id": 4,
+ "id": 962,
+ "area": 840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 489,
+ "bbox": [
+ 15,
+ 155,
+ 28,
+ 28
+ ],
+ "category_id": 4,
+ "id": 963,
+ "area": 784,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 489,
+ "bbox": [
+ 471,
+ 115,
+ 14,
+ 17
+ ],
+ "category_id": 4,
+ "id": 964,
+ "area": 238,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 489,
+ "bbox": [
+ 392,
+ 355,
+ 29,
+ 23
+ ],
+ "category_id": 4,
+ "id": 965,
+ "area": 667,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 489,
+ "bbox": [
+ 345,
+ 460,
+ 16,
+ 17
+ ],
+ "category_id": 4,
+ "id": 966,
+ "area": 272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 490,
+ "bbox": [
+ 353,
+ 389,
+ 132,
+ 57
+ ],
+ "category_id": 10,
+ "id": 967,
+ "area": 7524,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 491,
+ "bbox": [
+ 49,
+ 113,
+ 37,
+ 139
+ ],
+ "category_id": 19,
+ "id": 968,
+ "area": 5143,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 492,
+ "bbox": [
+ 210,
+ 282,
+ 130,
+ 141
+ ],
+ "category_id": 15,
+ "id": 969,
+ "area": 18330,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 492,
+ "bbox": [
+ 375,
+ 192,
+ 101,
+ 108
+ ],
+ "category_id": 15,
+ "id": 970,
+ "area": 10908,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 493,
+ "bbox": [
+ 156,
+ 243,
+ 137,
+ 149
+ ],
+ "category_id": 15,
+ "id": 971,
+ "area": 20413,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 494,
+ "bbox": [
+ 194,
+ 106,
+ 181,
+ 130
+ ],
+ "category_id": 15,
+ "id": 972,
+ "area": 23530,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 494,
+ "bbox": [
+ 267,
+ 284,
+ 188,
+ 132
+ ],
+ "category_id": 15,
+ "id": 973,
+ "area": 24816,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 495,
+ "bbox": [
+ 83,
+ 96,
+ 31,
+ 62
+ ],
+ "category_id": 16,
+ "id": 974,
+ "area": 1922,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 496,
+ "bbox": [
+ 218,
+ 403,
+ 50,
+ 63
+ ],
+ "category_id": 15,
+ "id": 975,
+ "area": 3150,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 496,
+ "bbox": [
+ 357,
+ 468,
+ 29,
+ 38
+ ],
+ "category_id": 15,
+ "id": 976,
+ "area": 1102,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 497,
+ "bbox": [
+ 170,
+ 236,
+ 185,
+ 60
+ ],
+ "category_id": 19,
+ "id": 977,
+ "area": 11100,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 498,
+ "bbox": [
+ 139,
+ 87,
+ 65,
+ 20
+ ],
+ "category_id": 4,
+ "id": 978,
+ "area": 1300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 498,
+ "bbox": [
+ 318,
+ 369,
+ 68,
+ 24
+ ],
+ "category_id": 4,
+ "id": 979,
+ "area": 1632,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 499,
+ "bbox": [
+ 216,
+ 109,
+ 32,
+ 36
+ ],
+ "category_id": 4,
+ "id": 980,
+ "area": 1152,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 499,
+ "bbox": [
+ 66,
+ 326,
+ 43,
+ 42
+ ],
+ "category_id": 4,
+ "id": 981,
+ "area": 1806,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 499,
+ "bbox": [
+ 421,
+ 414,
+ 31,
+ 47
+ ],
+ "category_id": 4,
+ "id": 982,
+ "area": 1457,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 500,
+ "bbox": [
+ 64,
+ 147,
+ 48,
+ 37
+ ],
+ "category_id": 4,
+ "id": 983,
+ "area": 1776,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 500,
+ "bbox": [
+ 207,
+ 367,
+ 60,
+ 43
+ ],
+ "category_id": 4,
+ "id": 984,
+ "area": 2580,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 500,
+ "bbox": [
+ 401,
+ 81,
+ 64,
+ 41
+ ],
+ "category_id": 4,
+ "id": 985,
+ "area": 2624,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 501,
+ "bbox": [
+ 83,
+ 296,
+ 165,
+ 141
+ ],
+ "category_id": 10,
+ "id": 986,
+ "area": 23265,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 502,
+ "bbox": [
+ 50,
+ 184,
+ 437,
+ 91
+ ],
+ "category_id": 10,
+ "id": 987,
+ "area": 39767,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 503,
+ "bbox": [
+ 109,
+ 85,
+ 311,
+ 262
+ ],
+ "category_id": 10,
+ "id": 988,
+ "area": 81482,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 504,
+ "bbox": [
+ 128,
+ 60,
+ 103,
+ 104
+ ],
+ "category_id": 10,
+ "id": 989,
+ "area": 10712,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 505,
+ "bbox": [
+ 26,
+ 53,
+ 225,
+ 149
+ ],
+ "category_id": 10,
+ "id": 990,
+ "area": 33525,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 506,
+ "bbox": [
+ 53,
+ 407,
+ 57,
+ 61
+ ],
+ "category_id": 4,
+ "id": 991,
+ "area": 3477,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 506,
+ "bbox": [
+ 382,
+ 356,
+ 61,
+ 58
+ ],
+ "category_id": 4,
+ "id": 992,
+ "area": 3538,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 506,
+ "bbox": [
+ 360,
+ 83,
+ 62,
+ 61
+ ],
+ "category_id": 4,
+ "id": 993,
+ "area": 3782,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 507,
+ "bbox": [
+ 256,
+ 212,
+ 65,
+ 81
+ ],
+ "category_id": 15,
+ "id": 994,
+ "area": 5265,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 508,
+ "bbox": [
+ 26,
+ 194,
+ 459,
+ 59
+ ],
+ "category_id": 10,
+ "id": 995,
+ "area": 27081,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 509,
+ "bbox": [
+ 81,
+ 167,
+ 146,
+ 170
+ ],
+ "category_id": 15,
+ "id": 996,
+ "area": 24820,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 509,
+ "bbox": [
+ 284,
+ 206,
+ 146,
+ 170
+ ],
+ "category_id": 15,
+ "id": 997,
+ "area": 24820,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 510,
+ "bbox": [
+ 22,
+ 179,
+ 253,
+ 185
+ ],
+ "category_id": 16,
+ "id": 998,
+ "area": 46805,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 511,
+ "bbox": [
+ 362,
+ 232,
+ 118,
+ 45
+ ],
+ "category_id": 19,
+ "id": 999,
+ "area": 5310,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 512,
+ "bbox": [
+ 226,
+ 100,
+ 44,
+ 250
+ ],
+ "category_id": 19,
+ "id": 1000,
+ "area": 11000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 513,
+ "bbox": [
+ 161,
+ 246,
+ 197,
+ 44
+ ],
+ "category_id": 19,
+ "id": 1001,
+ "area": 8668,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 514,
+ "bbox": [
+ 151,
+ 76,
+ 46,
+ 28
+ ],
+ "category_id": 4,
+ "id": 1002,
+ "area": 1288,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 514,
+ "bbox": [
+ 405,
+ 69,
+ 55,
+ 17
+ ],
+ "category_id": 4,
+ "id": 1003,
+ "area": 935,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 514,
+ "bbox": [
+ 98,
+ 247,
+ 49,
+ 33
+ ],
+ "category_id": 4,
+ "id": 1004,
+ "area": 1617,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 514,
+ "bbox": [
+ 91,
+ 398,
+ 42,
+ 27
+ ],
+ "category_id": 4,
+ "id": 1005,
+ "area": 1134,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 514,
+ "bbox": [
+ 319,
+ 425,
+ 49,
+ 39
+ ],
+ "category_id": 4,
+ "id": 1006,
+ "area": 1911,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 515,
+ "bbox": [
+ 64,
+ 170,
+ 40,
+ 34
+ ],
+ "category_id": 4,
+ "id": 1007,
+ "area": 1360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 515,
+ "bbox": [
+ 410,
+ 225,
+ 46,
+ 39
+ ],
+ "category_id": 4,
+ "id": 1008,
+ "area": 1794,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 516,
+ "bbox": [
+ 140,
+ 161,
+ 37,
+ 109
+ ],
+ "category_id": 16,
+ "id": 1009,
+ "area": 4033,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 517,
+ "bbox": [
+ 77,
+ 257,
+ 161,
+ 218
+ ],
+ "category_id": 10,
+ "id": 1010,
+ "area": 35098,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 518,
+ "bbox": [
+ 103,
+ 171,
+ 296,
+ 68
+ ],
+ "category_id": 19,
+ "id": 1011,
+ "area": 20128,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 519,
+ "bbox": [
+ 114,
+ 391,
+ 35,
+ 33
+ ],
+ "category_id": 4,
+ "id": 1012,
+ "area": 1155,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 519,
+ "bbox": [
+ 278,
+ 167,
+ 33,
+ 36
+ ],
+ "category_id": 4,
+ "id": 1013,
+ "area": 1188,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 519,
+ "bbox": [
+ 411,
+ 269,
+ 34,
+ 35
+ ],
+ "category_id": 4,
+ "id": 1014,
+ "area": 1190,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 520,
+ "bbox": [
+ 167,
+ 212,
+ 177,
+ 143
+ ],
+ "category_id": 16,
+ "id": 1015,
+ "area": 25311,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 521,
+ "bbox": [
+ 248,
+ 55,
+ 130,
+ 134
+ ],
+ "category_id": 15,
+ "id": 1016,
+ "area": 17420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 521,
+ "bbox": [
+ 204,
+ 180,
+ 275,
+ 274
+ ],
+ "category_id": 15,
+ "id": 1017,
+ "area": 75350,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 522,
+ "bbox": [
+ 244,
+ 96,
+ 132,
+ 188
+ ],
+ "category_id": 16,
+ "id": 1018,
+ "area": 24816,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 523,
+ "bbox": [
+ 190,
+ 221,
+ 159,
+ 114
+ ],
+ "category_id": 16,
+ "id": 1019,
+ "area": 18126,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 524,
+ "bbox": [
+ 160,
+ 107,
+ 197,
+ 242
+ ],
+ "category_id": 16,
+ "id": 1020,
+ "area": 47674,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 525,
+ "bbox": [
+ 156,
+ 280,
+ 23,
+ 95
+ ],
+ "category_id": 16,
+ "id": 1021,
+ "area": 2185,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 526,
+ "bbox": [
+ 33,
+ 284,
+ 30,
+ 15
+ ],
+ "category_id": 4,
+ "id": 1022,
+ "area": 450,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 526,
+ "bbox": [
+ 371,
+ 299,
+ 30,
+ 15
+ ],
+ "category_id": 4,
+ "id": 1023,
+ "area": 450,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 526,
+ "bbox": [
+ 448,
+ 352,
+ 18,
+ 28
+ ],
+ "category_id": 4,
+ "id": 1024,
+ "area": 504,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 526,
+ "bbox": [
+ 245,
+ 407,
+ 32,
+ 19
+ ],
+ "category_id": 4,
+ "id": 1025,
+ "area": 608,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 526,
+ "bbox": [
+ 172,
+ 476,
+ 35,
+ 15
+ ],
+ "category_id": 4,
+ "id": 1026,
+ "area": 525,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 527,
+ "bbox": [
+ 295,
+ 103,
+ 61,
+ 42
+ ],
+ "category_id": 19,
+ "id": 1027,
+ "area": 2562,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 528,
+ "bbox": [
+ 129,
+ 240,
+ 249,
+ 74
+ ],
+ "category_id": 19,
+ "id": 1028,
+ "area": 18426,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 529,
+ "bbox": [
+ 210,
+ 182,
+ 180,
+ 159
+ ],
+ "category_id": 10,
+ "id": 1029,
+ "area": 28620,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 530,
+ "bbox": [
+ 109,
+ 255,
+ 374,
+ 232
+ ],
+ "category_id": 10,
+ "id": 1030,
+ "area": 86768,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 531,
+ "bbox": [
+ 131,
+ 190,
+ 199,
+ 195
+ ],
+ "category_id": 10,
+ "id": 1031,
+ "area": 38805,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 532,
+ "bbox": [
+ 129,
+ 149,
+ 206,
+ 310
+ ],
+ "category_id": 16,
+ "id": 1032,
+ "area": 63860,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 533,
+ "bbox": [
+ 37,
+ 254,
+ 14,
+ 16
+ ],
+ "category_id": 4,
+ "id": 1033,
+ "area": 224,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 533,
+ "bbox": [
+ 216,
+ 372,
+ 14,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1034,
+ "area": 322,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 533,
+ "bbox": [
+ 271,
+ 181,
+ 13,
+ 18
+ ],
+ "category_id": 4,
+ "id": 1035,
+ "area": 234,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 533,
+ "bbox": [
+ 467,
+ 172,
+ 12,
+ 14
+ ],
+ "category_id": 4,
+ "id": 1036,
+ "area": 168,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 534,
+ "bbox": [
+ 53,
+ 357,
+ 36,
+ 32
+ ],
+ "category_id": 15,
+ "id": 1037,
+ "area": 1152,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 534,
+ "bbox": [
+ 99,
+ 118,
+ 107,
+ 82
+ ],
+ "category_id": 15,
+ "id": 1038,
+ "area": 8774,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 534,
+ "bbox": [
+ 200,
+ 232,
+ 109,
+ 98
+ ],
+ "category_id": 15,
+ "id": 1039,
+ "area": 10682,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 535,
+ "bbox": [
+ 275,
+ 204,
+ 82,
+ 214
+ ],
+ "category_id": 16,
+ "id": 1040,
+ "area": 17548,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 536,
+ "bbox": [
+ 37,
+ 216,
+ 40,
+ 31
+ ],
+ "category_id": 4,
+ "id": 1041,
+ "area": 1240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 536,
+ "bbox": [
+ 435,
+ 235,
+ 43,
+ 37
+ ],
+ "category_id": 4,
+ "id": 1042,
+ "area": 1591,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 537,
+ "bbox": [
+ 435,
+ 435,
+ 13,
+ 25
+ ],
+ "category_id": 16,
+ "id": 1043,
+ "area": 325,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 538,
+ "bbox": [
+ 126,
+ 232,
+ 122,
+ 73
+ ],
+ "category_id": 10,
+ "id": 1044,
+ "area": 8906,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 539,
+ "bbox": [
+ 64,
+ 218,
+ 40,
+ 62
+ ],
+ "category_id": 4,
+ "id": 1045,
+ "area": 2480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 539,
+ "bbox": [
+ 305,
+ 73,
+ 44,
+ 53
+ ],
+ "category_id": 4,
+ "id": 1046,
+ "area": 2332,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 539,
+ "bbox": [
+ 395,
+ 414,
+ 41,
+ 48
+ ],
+ "category_id": 4,
+ "id": 1047,
+ "area": 1968,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 540,
+ "bbox": [
+ 151,
+ 184,
+ 336,
+ 264
+ ],
+ "category_id": 15,
+ "id": 1048,
+ "area": 88704,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 541,
+ "bbox": [
+ 370,
+ 227,
+ 23,
+ 39
+ ],
+ "category_id": 16,
+ "id": 1049,
+ "area": 897,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 541,
+ "bbox": [
+ 450,
+ 44,
+ 24,
+ 20
+ ],
+ "category_id": 16,
+ "id": 1050,
+ "area": 480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 541,
+ "bbox": [
+ 383,
+ 471,
+ 34,
+ 16
+ ],
+ "category_id": 16,
+ "id": 1051,
+ "area": 544,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 542,
+ "bbox": [
+ 234,
+ 257,
+ 71,
+ 55
+ ],
+ "category_id": 4,
+ "id": 1052,
+ "area": 3905,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 543,
+ "bbox": [
+ 31,
+ 152,
+ 174,
+ 168
+ ],
+ "category_id": 15,
+ "id": 1053,
+ "area": 29232,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 543,
+ "bbox": [
+ 281,
+ 192,
+ 169,
+ 160
+ ],
+ "category_id": 15,
+ "id": 1054,
+ "area": 27040,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 544,
+ "bbox": [
+ 55,
+ 37,
+ 134,
+ 308
+ ],
+ "category_id": 10,
+ "id": 1055,
+ "area": 41272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 545,
+ "bbox": [
+ 307,
+ 76,
+ 34,
+ 21
+ ],
+ "category_id": 4,
+ "id": 1056,
+ "area": 714,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 545,
+ "bbox": [
+ 144,
+ 223,
+ 42,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1057,
+ "area": 1008,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 545,
+ "bbox": [
+ 272,
+ 417,
+ 35,
+ 27
+ ],
+ "category_id": 4,
+ "id": 1058,
+ "area": 945,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 546,
+ "bbox": [
+ 383,
+ 343,
+ 57,
+ 111
+ ],
+ "category_id": 10,
+ "id": 1059,
+ "area": 6327,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 546,
+ "bbox": [
+ 310,
+ 318,
+ 51,
+ 90
+ ],
+ "category_id": 10,
+ "id": 1060,
+ "area": 4590,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 547,
+ "bbox": [
+ 215,
+ 99,
+ 248,
+ 366
+ ],
+ "category_id": 10,
+ "id": 1061,
+ "area": 90768,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 548,
+ "bbox": [
+ 232,
+ 238,
+ 54,
+ 89
+ ],
+ "category_id": 4,
+ "id": 1062,
+ "area": 4806,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 549,
+ "bbox": [
+ 141,
+ 106,
+ 89,
+ 52
+ ],
+ "category_id": 19,
+ "id": 1063,
+ "area": 4628,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 550,
+ "bbox": [
+ 19,
+ 241,
+ 320,
+ 164
+ ],
+ "category_id": 16,
+ "id": 1064,
+ "area": 52480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 551,
+ "bbox": [
+ 182,
+ 221,
+ 224,
+ 172
+ ],
+ "category_id": 16,
+ "id": 1065,
+ "area": 38528,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 552,
+ "bbox": [
+ 150,
+ 384,
+ 246,
+ 99
+ ],
+ "category_id": 10,
+ "id": 1066,
+ "area": 24354,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 553,
+ "bbox": [
+ 353,
+ 423,
+ 59,
+ 57
+ ],
+ "category_id": 15,
+ "id": 1067,
+ "area": 3363,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 553,
+ "bbox": [
+ 480,
+ 459,
+ 32,
+ 31
+ ],
+ "category_id": 15,
+ "id": 1068,
+ "area": 992,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 553,
+ "bbox": [
+ 177,
+ 105,
+ 84,
+ 87
+ ],
+ "category_id": 15,
+ "id": 1069,
+ "area": 7308,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 553,
+ "bbox": [
+ 288,
+ 80,
+ 88,
+ 87
+ ],
+ "category_id": 15,
+ "id": 1070,
+ "area": 7656,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 553,
+ "bbox": [
+ 182,
+ 354,
+ 90,
+ 88
+ ],
+ "category_id": 15,
+ "id": 1071,
+ "area": 7920,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 553,
+ "bbox": [
+ 208,
+ 236,
+ 88,
+ 91
+ ],
+ "category_id": 15,
+ "id": 1072,
+ "area": 8008,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 554,
+ "bbox": [
+ 82,
+ 250,
+ 194,
+ 190
+ ],
+ "category_id": 15,
+ "id": 1073,
+ "area": 36860,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 554,
+ "bbox": [
+ 235,
+ 69,
+ 120,
+ 120
+ ],
+ "category_id": 15,
+ "id": 1074,
+ "area": 14400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 554,
+ "bbox": [
+ 314,
+ 176,
+ 118,
+ 119
+ ],
+ "category_id": 15,
+ "id": 1075,
+ "area": 14042,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 555,
+ "bbox": [
+ 154,
+ 180,
+ 244,
+ 189
+ ],
+ "category_id": 16,
+ "id": 1076,
+ "area": 46116,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 556,
+ "bbox": [
+ 166,
+ 118,
+ 249,
+ 192
+ ],
+ "category_id": 15,
+ "id": 1077,
+ "area": 47808,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 557,
+ "bbox": [
+ 321,
+ 83,
+ 33,
+ 13
+ ],
+ "category_id": 16,
+ "id": 1078,
+ "area": 429,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 557,
+ "bbox": [
+ 385,
+ 25,
+ 25,
+ 21
+ ],
+ "category_id": 16,
+ "id": 1079,
+ "area": 525,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 558,
+ "bbox": [
+ 161,
+ 90,
+ 279,
+ 246
+ ],
+ "category_id": 10,
+ "id": 1080,
+ "area": 68634,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 559,
+ "bbox": [
+ 200,
+ 421,
+ 36,
+ 39
+ ],
+ "category_id": 4,
+ "id": 1081,
+ "area": 1404,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 559,
+ "bbox": [
+ 348,
+ 98,
+ 38,
+ 50
+ ],
+ "category_id": 4,
+ "id": 1082,
+ "area": 1900,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 560,
+ "bbox": [
+ 29,
+ 88,
+ 463,
+ 292
+ ],
+ "category_id": 10,
+ "id": 1083,
+ "area": 135196,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 561,
+ "bbox": [
+ 183,
+ 96,
+ 292,
+ 337
+ ],
+ "category_id": 10,
+ "id": 1084,
+ "area": 98404,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 562,
+ "bbox": [
+ 169,
+ 182,
+ 70,
+ 45
+ ],
+ "category_id": 4,
+ "id": 1085,
+ "area": 3150,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 562,
+ "bbox": [
+ 407,
+ 372,
+ 62,
+ 54
+ ],
+ "category_id": 4,
+ "id": 1086,
+ "area": 3348,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 563,
+ "bbox": [
+ 141,
+ 103,
+ 79,
+ 33
+ ],
+ "category_id": 19,
+ "id": 1087,
+ "area": 2607,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 564,
+ "bbox": [
+ 318,
+ 119,
+ 66,
+ 276
+ ],
+ "category_id": 19,
+ "id": 1088,
+ "area": 18216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 565,
+ "bbox": [
+ 478,
+ 109,
+ 10,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1089,
+ "area": 230,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 565,
+ "bbox": [
+ 37,
+ 283,
+ 14,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1090,
+ "area": 336,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 565,
+ "bbox": [
+ 410,
+ 224,
+ 15,
+ 28
+ ],
+ "category_id": 4,
+ "id": 1091,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 565,
+ "bbox": [
+ 179,
+ 363,
+ 11,
+ 26
+ ],
+ "category_id": 4,
+ "id": 1092,
+ "area": 286,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 565,
+ "bbox": [
+ 343,
+ 332,
+ 14,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1093,
+ "area": 322,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 565,
+ "bbox": [
+ 122,
+ 465,
+ 14,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1094,
+ "area": 350,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 566,
+ "bbox": [
+ 73,
+ 248,
+ 15,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1095,
+ "area": 345,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 566,
+ "bbox": [
+ 172,
+ 13,
+ 16,
+ 17
+ ],
+ "category_id": 4,
+ "id": 1096,
+ "area": 272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 566,
+ "bbox": [
+ 378,
+ 122,
+ 17,
+ 22
+ ],
+ "category_id": 4,
+ "id": 1097,
+ "area": 374,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 566,
+ "bbox": [
+ 292,
+ 352,
+ 14,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1098,
+ "area": 336,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 566,
+ "bbox": [
+ 471,
+ 344,
+ 14,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1099,
+ "area": 350,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 567,
+ "bbox": [
+ 114,
+ 357,
+ 59,
+ 29
+ ],
+ "category_id": 4,
+ "id": 1100,
+ "area": 1711,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 567,
+ "bbox": [
+ 368,
+ 94,
+ 55,
+ 32
+ ],
+ "category_id": 4,
+ "id": 1101,
+ "area": 1760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 568,
+ "bbox": [
+ 166,
+ 56,
+ 26,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1102,
+ "area": 650,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 568,
+ "bbox": [
+ 402,
+ 170,
+ 24,
+ 30
+ ],
+ "category_id": 4,
+ "id": 1103,
+ "area": 720,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 568,
+ "bbox": [
+ 207,
+ 407,
+ 23,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1104,
+ "area": 529,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 569,
+ "bbox": [
+ 59,
+ 224,
+ 419,
+ 78
+ ],
+ "category_id": 19,
+ "id": 1105,
+ "area": 32682,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 570,
+ "bbox": [
+ 279,
+ 175,
+ 88,
+ 107
+ ],
+ "category_id": 15,
+ "id": 1106,
+ "area": 9416,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 570,
+ "bbox": [
+ 261,
+ 317,
+ 90,
+ 109
+ ],
+ "category_id": 15,
+ "id": 1107,
+ "area": 9810,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 571,
+ "bbox": [
+ 165,
+ 128,
+ 192,
+ 296
+ ],
+ "category_id": 19,
+ "id": 1108,
+ "area": 56832,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 572,
+ "bbox": [
+ 74,
+ 203,
+ 17,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1109,
+ "area": 187,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 572,
+ "bbox": [
+ 53,
+ 87,
+ 9,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1110,
+ "area": 72,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 572,
+ "bbox": [
+ 186,
+ 33,
+ 13,
+ 7
+ ],
+ "category_id": 4,
+ "id": 1111,
+ "area": 91,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 572,
+ "bbox": [
+ 82,
+ 19,
+ 9,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1112,
+ "area": 72,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 572,
+ "bbox": [
+ 284,
+ 342,
+ 5,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1113,
+ "area": 55,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 572,
+ "bbox": [
+ 204,
+ 330,
+ 7,
+ 5
+ ],
+ "category_id": 4,
+ "id": 1114,
+ "area": 35,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 572,
+ "bbox": [
+ 346,
+ 343,
+ 7,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1115,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 572,
+ "bbox": [
+ 196,
+ 204,
+ 12,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1116,
+ "area": 96,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 572,
+ "bbox": [
+ 460,
+ 88,
+ 12,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1117,
+ "area": 108,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 572,
+ "bbox": [
+ 319,
+ 12,
+ 13,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1118,
+ "area": 143,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 572,
+ "bbox": [
+ 277,
+ 69,
+ 12,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1119,
+ "area": 96,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 572,
+ "bbox": [
+ 382,
+ 133,
+ 10,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1120,
+ "area": 90,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 573,
+ "bbox": [
+ 86,
+ 35,
+ 330,
+ 423
+ ],
+ "category_id": 16,
+ "id": 1121,
+ "area": 139590,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 574,
+ "bbox": [
+ 156,
+ 132,
+ 249,
+ 98
+ ],
+ "category_id": 19,
+ "id": 1122,
+ "area": 24402,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 575,
+ "bbox": [
+ 65,
+ 328,
+ 42,
+ 68
+ ],
+ "category_id": 4,
+ "id": 1123,
+ "area": 2856,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 575,
+ "bbox": [
+ 438,
+ 168,
+ 36,
+ 60
+ ],
+ "category_id": 4,
+ "id": 1124,
+ "area": 2160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 576,
+ "bbox": [
+ 275,
+ 202,
+ 48,
+ 98
+ ],
+ "category_id": 4,
+ "id": 1125,
+ "area": 4704,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 577,
+ "bbox": [
+ 316,
+ 136,
+ 26,
+ 19
+ ],
+ "category_id": 4,
+ "id": 1126,
+ "area": 494,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 577,
+ "bbox": [
+ 169,
+ 263,
+ 18,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1127,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 577,
+ "bbox": [
+ 51,
+ 402,
+ 16,
+ 14
+ ],
+ "category_id": 4,
+ "id": 1128,
+ "area": 224,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 577,
+ "bbox": [
+ 432,
+ 273,
+ 17,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1129,
+ "area": 153,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 578,
+ "bbox": [
+ 65,
+ 58,
+ 194,
+ 43
+ ],
+ "category_id": 10,
+ "id": 1130,
+ "area": 8342,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 579,
+ "bbox": [
+ 38,
+ 174,
+ 425,
+ 115
+ ],
+ "category_id": 19,
+ "id": 1131,
+ "area": 48875,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 580,
+ "bbox": [
+ 281,
+ 102,
+ 108,
+ 297
+ ],
+ "category_id": 16,
+ "id": 1132,
+ "area": 32076,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 581,
+ "bbox": [
+ 176,
+ 60,
+ 225,
+ 395
+ ],
+ "category_id": 10,
+ "id": 1133,
+ "area": 88875,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 582,
+ "bbox": [
+ 128,
+ 196,
+ 23,
+ 12
+ ],
+ "category_id": 4,
+ "id": 1134,
+ "area": 276,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 582,
+ "bbox": [
+ 98,
+ 265,
+ 33,
+ 12
+ ],
+ "category_id": 4,
+ "id": 1135,
+ "area": 396,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 582,
+ "bbox": [
+ 3,
+ 296,
+ 21,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1136,
+ "area": 231,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 583,
+ "bbox": [
+ 178,
+ 135,
+ 197,
+ 270
+ ],
+ "category_id": 16,
+ "id": 1137,
+ "area": 53190,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 584,
+ "bbox": [
+ 66,
+ 115,
+ 18,
+ 15
+ ],
+ "category_id": 4,
+ "id": 1138,
+ "area": 270,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 584,
+ "bbox": [
+ 214,
+ 99,
+ 25,
+ 20
+ ],
+ "category_id": 4,
+ "id": 1139,
+ "area": 500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 584,
+ "bbox": [
+ 270,
+ 286,
+ 32,
+ 19
+ ],
+ "category_id": 4,
+ "id": 1140,
+ "area": 608,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 584,
+ "bbox": [
+ 402,
+ 218,
+ 19,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1141,
+ "area": 437,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 584,
+ "bbox": [
+ 406,
+ 403,
+ 26,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1142,
+ "area": 650,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 585,
+ "bbox": [
+ 228,
+ 134,
+ 38,
+ 234
+ ],
+ "category_id": 19,
+ "id": 1143,
+ "area": 8892,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 586,
+ "bbox": [
+ 73,
+ 122,
+ 152,
+ 88
+ ],
+ "category_id": 15,
+ "id": 1144,
+ "area": 13376,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 586,
+ "bbox": [
+ 289,
+ 243,
+ 164,
+ 96
+ ],
+ "category_id": 15,
+ "id": 1145,
+ "area": 15744,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 587,
+ "bbox": [
+ 247,
+ 11,
+ 22,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1146,
+ "area": 506,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 587,
+ "bbox": [
+ 328,
+ 156,
+ 20,
+ 20
+ ],
+ "category_id": 4,
+ "id": 1147,
+ "area": 400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 587,
+ "bbox": [
+ 103,
+ 273,
+ 30,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1148,
+ "area": 690,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 587,
+ "bbox": [
+ 204,
+ 389,
+ 24,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1149,
+ "area": 576,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 587,
+ "bbox": [
+ 315,
+ 488,
+ 24,
+ 19
+ ],
+ "category_id": 4,
+ "id": 1150,
+ "area": 456,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 588,
+ "bbox": [
+ 188,
+ 333,
+ 197,
+ 53
+ ],
+ "category_id": 16,
+ "id": 1151,
+ "area": 10441,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 589,
+ "bbox": [
+ 291,
+ 78,
+ 44,
+ 44
+ ],
+ "category_id": 4,
+ "id": 1152,
+ "area": 1936,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 589,
+ "bbox": [
+ 191,
+ 358,
+ 48,
+ 50
+ ],
+ "category_id": 4,
+ "id": 1153,
+ "area": 2400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 590,
+ "bbox": [
+ 383,
+ 197,
+ 76,
+ 94
+ ],
+ "category_id": 19,
+ "id": 1154,
+ "area": 7144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 591,
+ "bbox": [
+ 61,
+ 73,
+ 50,
+ 59
+ ],
+ "category_id": 4,
+ "id": 1155,
+ "area": 2950,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 591,
+ "bbox": [
+ 295,
+ 197,
+ 88,
+ 88
+ ],
+ "category_id": 4,
+ "id": 1156,
+ "area": 7744,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 591,
+ "bbox": [
+ 410,
+ 227,
+ 45,
+ 58
+ ],
+ "category_id": 4,
+ "id": 1157,
+ "area": 2610,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 592,
+ "bbox": [
+ 159,
+ 149,
+ 34,
+ 17
+ ],
+ "category_id": 4,
+ "id": 1158,
+ "area": 578,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 592,
+ "bbox": [
+ 25,
+ 285,
+ 26,
+ 20
+ ],
+ "category_id": 4,
+ "id": 1159,
+ "area": 520,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 592,
+ "bbox": [
+ 411,
+ 195,
+ 29,
+ 20
+ ],
+ "category_id": 4,
+ "id": 1160,
+ "area": 580,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 593,
+ "bbox": [
+ 69,
+ 122,
+ 119,
+ 260
+ ],
+ "category_id": 19,
+ "id": 1161,
+ "area": 30940,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 594,
+ "bbox": [
+ 292,
+ 143,
+ 157,
+ 226
+ ],
+ "category_id": 19,
+ "id": 1162,
+ "area": 35482,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 595,
+ "bbox": [
+ 102,
+ 48,
+ 338,
+ 384
+ ],
+ "category_id": 10,
+ "id": 1163,
+ "area": 129792,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 596,
+ "bbox": [
+ 37,
+ 114,
+ 23,
+ 14
+ ],
+ "category_id": 4,
+ "id": 1164,
+ "area": 322,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 596,
+ "bbox": [
+ 392,
+ 5,
+ 20,
+ 16
+ ],
+ "category_id": 4,
+ "id": 1165,
+ "area": 320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 596,
+ "bbox": [
+ 242,
+ 190,
+ 15,
+ 16
+ ],
+ "category_id": 4,
+ "id": 1166,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 596,
+ "bbox": [
+ 160,
+ 318,
+ 17,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1167,
+ "area": 153,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 596,
+ "bbox": [
+ 314,
+ 408,
+ 18,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1168,
+ "area": 198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 596,
+ "bbox": [
+ 181,
+ 502,
+ 18,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1169,
+ "area": 162,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 597,
+ "bbox": [
+ 124,
+ 136,
+ 315,
+ 323
+ ],
+ "category_id": 10,
+ "id": 1170,
+ "area": 101745,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 598,
+ "bbox": [
+ 112,
+ 317,
+ 227,
+ 86
+ ],
+ "category_id": 16,
+ "id": 1171,
+ "area": 19522,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 599,
+ "bbox": [
+ 237,
+ 69,
+ 32,
+ 34
+ ],
+ "category_id": 4,
+ "id": 1172,
+ "area": 1088,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 599,
+ "bbox": [
+ 186,
+ 392,
+ 30,
+ 35
+ ],
+ "category_id": 4,
+ "id": 1173,
+ "area": 1050,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 600,
+ "bbox": [
+ 136,
+ 180,
+ 126,
+ 200
+ ],
+ "category_id": 16,
+ "id": 1174,
+ "area": 25200,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 601,
+ "bbox": [
+ 108,
+ 22,
+ 26,
+ 26
+ ],
+ "category_id": 4,
+ "id": 1175,
+ "area": 676,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 601,
+ "bbox": [
+ 108,
+ 144,
+ 23,
+ 22
+ ],
+ "category_id": 4,
+ "id": 1176,
+ "area": 506,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 601,
+ "bbox": [
+ 80,
+ 276,
+ 24,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1177,
+ "area": 600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 601,
+ "bbox": [
+ 81,
+ 400,
+ 20,
+ 21
+ ],
+ "category_id": 4,
+ "id": 1178,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 601,
+ "bbox": [
+ 417,
+ 85,
+ 25,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1179,
+ "area": 575,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 601,
+ "bbox": [
+ 412,
+ 208,
+ 21,
+ 19
+ ],
+ "category_id": 4,
+ "id": 1180,
+ "area": 399,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 601,
+ "bbox": [
+ 390,
+ 343,
+ 26,
+ 27
+ ],
+ "category_id": 4,
+ "id": 1181,
+ "area": 702,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 601,
+ "bbox": [
+ 390,
+ 460,
+ 24,
+ 28
+ ],
+ "category_id": 4,
+ "id": 1182,
+ "area": 672,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 602,
+ "bbox": [
+ 192,
+ 69,
+ 192,
+ 358
+ ],
+ "category_id": 16,
+ "id": 1183,
+ "area": 68736,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 603,
+ "bbox": [
+ 202,
+ 24,
+ 33,
+ 13
+ ],
+ "category_id": 4,
+ "id": 1184,
+ "area": 429,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 603,
+ "bbox": [
+ 5,
+ 193,
+ 31,
+ 15
+ ],
+ "category_id": 4,
+ "id": 1185,
+ "area": 465,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 603,
+ "bbox": [
+ 413,
+ 58,
+ 28,
+ 14
+ ],
+ "category_id": 4,
+ "id": 1186,
+ "area": 392,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 603,
+ "bbox": [
+ 309,
+ 235,
+ 25,
+ 22
+ ],
+ "category_id": 4,
+ "id": 1187,
+ "area": 550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 603,
+ "bbox": [
+ 471,
+ 254,
+ 29,
+ 14
+ ],
+ "category_id": 4,
+ "id": 1188,
+ "area": 406,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 603,
+ "bbox": [
+ 266,
+ 358,
+ 30,
+ 20
+ ],
+ "category_id": 4,
+ "id": 1189,
+ "area": 600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 603,
+ "bbox": [
+ 21,
+ 414,
+ 30,
+ 16
+ ],
+ "category_id": 4,
+ "id": 1190,
+ "area": 480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 603,
+ "bbox": [
+ 318,
+ 488,
+ 28,
+ 18
+ ],
+ "category_id": 4,
+ "id": 1191,
+ "area": 504,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 604,
+ "bbox": [
+ 43,
+ 203,
+ 50,
+ 33
+ ],
+ "category_id": 4,
+ "id": 1192,
+ "area": 1650,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 604,
+ "bbox": [
+ 385,
+ 164,
+ 49,
+ 40
+ ],
+ "category_id": 4,
+ "id": 1193,
+ "area": 1960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 604,
+ "bbox": [
+ 293,
+ 417,
+ 48,
+ 32
+ ],
+ "category_id": 4,
+ "id": 1194,
+ "area": 1536,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 605,
+ "bbox": [
+ 94,
+ 55,
+ 14,
+ 21
+ ],
+ "category_id": 4,
+ "id": 1195,
+ "area": 294,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 605,
+ "bbox": [
+ 405,
+ 36,
+ 16,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1196,
+ "area": 384,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 605,
+ "bbox": [
+ 314,
+ 94,
+ 18,
+ 29
+ ],
+ "category_id": 4,
+ "id": 1197,
+ "area": 522,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 605,
+ "bbox": [
+ 204,
+ 168,
+ 17,
+ 27
+ ],
+ "category_id": 4,
+ "id": 1198,
+ "area": 459,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 605,
+ "bbox": [
+ 84,
+ 243,
+ 17,
+ 30
+ ],
+ "category_id": 4,
+ "id": 1199,
+ "area": 510,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 605,
+ "bbox": [
+ 471,
+ 264,
+ 16,
+ 29
+ ],
+ "category_id": 4,
+ "id": 1200,
+ "area": 464,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 605,
+ "bbox": [
+ 356,
+ 339,
+ 16,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1201,
+ "area": 368,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 605,
+ "bbox": [
+ 241,
+ 389,
+ 17,
+ 35
+ ],
+ "category_id": 4,
+ "id": 1202,
+ "area": 595,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 606,
+ "bbox": [
+ 176,
+ 158,
+ 213,
+ 198
+ ],
+ "category_id": 19,
+ "id": 1203,
+ "area": 42174,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 607,
+ "bbox": [
+ 79,
+ 172,
+ 344,
+ 124
+ ],
+ "category_id": 19,
+ "id": 1204,
+ "area": 42656,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 608,
+ "bbox": [
+ 96,
+ 229,
+ 299,
+ 169
+ ],
+ "category_id": 10,
+ "id": 1205,
+ "area": 50531,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 609,
+ "bbox": [
+ 257,
+ 166,
+ 32,
+ 170
+ ],
+ "category_id": 19,
+ "id": 1206,
+ "area": 5440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 610,
+ "bbox": [
+ 154,
+ 236,
+ 202,
+ 64
+ ],
+ "category_id": 19,
+ "id": 1207,
+ "area": 12928,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 611,
+ "bbox": [
+ 288,
+ 311,
+ 24,
+ 10
+ ],
+ "category_id": 16,
+ "id": 1208,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 612,
+ "bbox": [
+ 111,
+ 172,
+ 24,
+ 19
+ ],
+ "category_id": 16,
+ "id": 1209,
+ "area": 456,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 613,
+ "bbox": [
+ 182,
+ 199,
+ 92,
+ 97
+ ],
+ "category_id": 15,
+ "id": 1210,
+ "area": 8924,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 614,
+ "bbox": [
+ 249,
+ 237,
+ 63,
+ 131
+ ],
+ "category_id": 16,
+ "id": 1211,
+ "area": 8253,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 615,
+ "bbox": [
+ 114,
+ 176,
+ 179,
+ 145
+ ],
+ "category_id": 16,
+ "id": 1212,
+ "area": 25955,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 616,
+ "bbox": [
+ 152,
+ 207,
+ 65,
+ 67
+ ],
+ "category_id": 15,
+ "id": 1213,
+ "area": 4355,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 616,
+ "bbox": [
+ 218,
+ 252,
+ 66,
+ 68
+ ],
+ "category_id": 15,
+ "id": 1214,
+ "area": 4488,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 616,
+ "bbox": [
+ 289,
+ 298,
+ 49,
+ 53
+ ],
+ "category_id": 15,
+ "id": 1215,
+ "area": 2597,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 616,
+ "bbox": [
+ 340,
+ 334,
+ 53,
+ 57
+ ],
+ "category_id": 15,
+ "id": 1216,
+ "area": 3021,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 617,
+ "bbox": [
+ 73,
+ 285,
+ 33,
+ 60
+ ],
+ "category_id": 16,
+ "id": 1217,
+ "area": 1980,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 618,
+ "bbox": [
+ 62,
+ 292,
+ 146,
+ 83
+ ],
+ "category_id": 10,
+ "id": 1218,
+ "area": 12118,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 619,
+ "bbox": [
+ 35,
+ 148,
+ 423,
+ 91
+ ],
+ "category_id": 10,
+ "id": 1219,
+ "area": 38493,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 620,
+ "bbox": [
+ 96,
+ 46,
+ 51,
+ 116
+ ],
+ "category_id": 10,
+ "id": 1220,
+ "area": 5916,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 621,
+ "bbox": [
+ 145,
+ 21,
+ 102,
+ 187
+ ],
+ "category_id": 19,
+ "id": 1221,
+ "area": 19074,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 622,
+ "bbox": [
+ 58,
+ 91,
+ 48,
+ 36
+ ],
+ "category_id": 4,
+ "id": 1222,
+ "area": 1728,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 622,
+ "bbox": [
+ 264,
+ 419,
+ 47,
+ 39
+ ],
+ "category_id": 4,
+ "id": 1223,
+ "area": 1833,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 622,
+ "bbox": [
+ 434,
+ 129,
+ 58,
+ 43
+ ],
+ "category_id": 4,
+ "id": 1224,
+ "area": 2494,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 623,
+ "bbox": [
+ 193,
+ 240,
+ 234,
+ 96
+ ],
+ "category_id": 16,
+ "id": 1225,
+ "area": 22464,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 624,
+ "bbox": [
+ 186,
+ 139,
+ 51,
+ 196
+ ],
+ "category_id": 19,
+ "id": 1226,
+ "area": 9996,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 625,
+ "bbox": [
+ 108,
+ 257,
+ 299,
+ 51
+ ],
+ "category_id": 19,
+ "id": 1227,
+ "area": 15249,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 626,
+ "bbox": [
+ 94,
+ 139,
+ 60,
+ 49
+ ],
+ "category_id": 4,
+ "id": 1228,
+ "area": 2940,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 626,
+ "bbox": [
+ 405,
+ 361,
+ 58,
+ 49
+ ],
+ "category_id": 4,
+ "id": 1229,
+ "area": 2842,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 627,
+ "bbox": [
+ 145,
+ 92,
+ 11,
+ 7
+ ],
+ "category_id": 4,
+ "id": 1230,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 627,
+ "bbox": [
+ 193,
+ 122,
+ 6,
+ 7
+ ],
+ "category_id": 4,
+ "id": 1231,
+ "area": 42,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 627,
+ "bbox": [
+ 198,
+ 207,
+ 7,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1232,
+ "area": 63,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 627,
+ "bbox": [
+ 261,
+ 222,
+ 8,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1233,
+ "area": 64,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 627,
+ "bbox": [
+ 299,
+ 289,
+ 7,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1234,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 627,
+ "bbox": [
+ 105,
+ 375,
+ 5,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1235,
+ "area": 40,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 627,
+ "bbox": [
+ 378,
+ 396,
+ 4,
+ 7
+ ],
+ "category_id": 4,
+ "id": 1236,
+ "area": 28,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 627,
+ "bbox": [
+ 405,
+ 344,
+ 7,
+ 13
+ ],
+ "category_id": 4,
+ "id": 1237,
+ "area": 91,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 628,
+ "bbox": [
+ 71,
+ 147,
+ 339,
+ 201
+ ],
+ "category_id": 16,
+ "id": 1238,
+ "area": 68139,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 629,
+ "bbox": [
+ 98,
+ 301,
+ 56,
+ 148
+ ],
+ "category_id": 10,
+ "id": 1239,
+ "area": 8288,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 630,
+ "bbox": [
+ 147,
+ 243,
+ 70,
+ 61
+ ],
+ "category_id": 15,
+ "id": 1240,
+ "area": 4270,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 631,
+ "bbox": [
+ 252,
+ 168,
+ 70,
+ 191
+ ],
+ "category_id": 16,
+ "id": 1241,
+ "area": 13370,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 632,
+ "bbox": [
+ 383,
+ 90,
+ 13,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1242,
+ "area": 299,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 632,
+ "bbox": [
+ 178,
+ 177,
+ 12,
+ 22
+ ],
+ "category_id": 4,
+ "id": 1243,
+ "area": 264,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 632,
+ "bbox": [
+ 88,
+ 286,
+ 13,
+ 21
+ ],
+ "category_id": 4,
+ "id": 1244,
+ "area": 273,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 632,
+ "bbox": [
+ 250,
+ 319,
+ 12,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1245,
+ "area": 276,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 632,
+ "bbox": [
+ 449,
+ 369,
+ 11,
+ 29
+ ],
+ "category_id": 4,
+ "id": 1246,
+ "area": 319,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 633,
+ "bbox": [
+ 240,
+ 165,
+ 85,
+ 60
+ ],
+ "category_id": 4,
+ "id": 1247,
+ "area": 5100,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 633,
+ "bbox": [
+ 328,
+ 389,
+ 84,
+ 60
+ ],
+ "category_id": 4,
+ "id": 1248,
+ "area": 5040,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 634,
+ "bbox": [
+ 74,
+ 124,
+ 366,
+ 117
+ ],
+ "category_id": 19,
+ "id": 1249,
+ "area": 42822,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 635,
+ "bbox": [
+ 133,
+ 163,
+ 298,
+ 131
+ ],
+ "category_id": 16,
+ "id": 1250,
+ "area": 39038,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 636,
+ "bbox": [
+ 97,
+ 332,
+ 35,
+ 32
+ ],
+ "category_id": 4,
+ "id": 1251,
+ "area": 1120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 636,
+ "bbox": [
+ 400,
+ 346,
+ 32,
+ 30
+ ],
+ "category_id": 4,
+ "id": 1252,
+ "area": 960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 637,
+ "bbox": [
+ 198,
+ 173,
+ 109,
+ 92
+ ],
+ "category_id": 15,
+ "id": 1253,
+ "area": 10028,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 638,
+ "bbox": [
+ 85,
+ 163,
+ 396,
+ 151
+ ],
+ "category_id": 10,
+ "id": 1254,
+ "area": 59796,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 639,
+ "bbox": [
+ 344,
+ 342,
+ 35,
+ 72
+ ],
+ "category_id": 16,
+ "id": 1255,
+ "area": 2520,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 640,
+ "bbox": [
+ 272,
+ 14,
+ 84,
+ 63
+ ],
+ "category_id": 4,
+ "id": 1256,
+ "area": 5292,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 640,
+ "bbox": [
+ 196,
+ 378,
+ 95,
+ 68
+ ],
+ "category_id": 4,
+ "id": 1257,
+ "area": 6460,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 641,
+ "bbox": [
+ 19,
+ 226,
+ 481,
+ 81
+ ],
+ "category_id": 10,
+ "id": 1258,
+ "area": 38961,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 642,
+ "bbox": [
+ 133,
+ 149,
+ 151,
+ 140
+ ],
+ "category_id": 15,
+ "id": 1259,
+ "area": 21140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 642,
+ "bbox": [
+ 129,
+ 353,
+ 153,
+ 135
+ ],
+ "category_id": 15,
+ "id": 1260,
+ "area": 20655,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 643,
+ "bbox": [
+ 206,
+ 135,
+ 148,
+ 283
+ ],
+ "category_id": 16,
+ "id": 1261,
+ "area": 41884,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 644,
+ "bbox": [
+ 113,
+ 80,
+ 45,
+ 64
+ ],
+ "category_id": 4,
+ "id": 1262,
+ "area": 2880,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 644,
+ "bbox": [
+ 399,
+ 350,
+ 46,
+ 60
+ ],
+ "category_id": 4,
+ "id": 1263,
+ "area": 2760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 645,
+ "bbox": [
+ 144,
+ 140,
+ 261,
+ 270
+ ],
+ "category_id": 10,
+ "id": 1264,
+ "area": 70470,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 646,
+ "bbox": [
+ 286,
+ 55,
+ 66,
+ 146
+ ],
+ "category_id": 19,
+ "id": 1265,
+ "area": 9636,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 647,
+ "bbox": [
+ 261,
+ 232,
+ 73,
+ 82
+ ],
+ "category_id": 4,
+ "id": 1266,
+ "area": 5986,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 648,
+ "bbox": [
+ 108,
+ 85,
+ 53,
+ 60
+ ],
+ "category_id": 4,
+ "id": 1267,
+ "area": 3180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 648,
+ "bbox": [
+ 333,
+ 364,
+ 55,
+ 44
+ ],
+ "category_id": 4,
+ "id": 1268,
+ "area": 2420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 649,
+ "bbox": [
+ 58,
+ 163,
+ 275,
+ 129
+ ],
+ "category_id": 10,
+ "id": 1269,
+ "area": 35475,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 650,
+ "bbox": [
+ 199,
+ 289,
+ 102,
+ 121
+ ],
+ "category_id": 15,
+ "id": 1270,
+ "area": 12342,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 651,
+ "bbox": [
+ 36,
+ 161,
+ 35,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1271,
+ "area": 805,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 651,
+ "bbox": [
+ 43,
+ 346,
+ 25,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1272,
+ "area": 625,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 651,
+ "bbox": [
+ 457,
+ 64,
+ 29,
+ 21
+ ],
+ "category_id": 4,
+ "id": 1273,
+ "area": 609,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 651,
+ "bbox": [
+ 460,
+ 180,
+ 24,
+ 26
+ ],
+ "category_id": 4,
+ "id": 1274,
+ "area": 624,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 651,
+ "bbox": [
+ 456,
+ 296,
+ 25,
+ 22
+ ],
+ "category_id": 4,
+ "id": 1275,
+ "area": 550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 651,
+ "bbox": [
+ 462,
+ 395,
+ 22,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1276,
+ "area": 528,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 652,
+ "bbox": [
+ 346,
+ 51,
+ 20,
+ 7
+ ],
+ "category_id": 4,
+ "id": 1277,
+ "area": 140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 652,
+ "bbox": [
+ 109,
+ 208,
+ 16,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1278,
+ "area": 144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 652,
+ "bbox": [
+ 391,
+ 229,
+ 17,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1279,
+ "area": 153,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 652,
+ "bbox": [
+ 184,
+ 315,
+ 20,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1280,
+ "area": 200,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 652,
+ "bbox": [
+ 402,
+ 425,
+ 17,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1281,
+ "area": 153,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 652,
+ "bbox": [
+ 227,
+ 423,
+ 20,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1282,
+ "area": 200,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 652,
+ "bbox": [
+ 33,
+ 300,
+ 18,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1283,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 653,
+ "bbox": [
+ 53,
+ 61,
+ 448,
+ 399
+ ],
+ "category_id": 16,
+ "id": 1284,
+ "area": 178752,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 654,
+ "bbox": [
+ 193,
+ 67,
+ 27,
+ 19
+ ],
+ "category_id": 4,
+ "id": 1285,
+ "area": 513,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 654,
+ "bbox": [
+ 425,
+ 80,
+ 19,
+ 16
+ ],
+ "category_id": 4,
+ "id": 1286,
+ "area": 304,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 654,
+ "bbox": [
+ 56,
+ 184,
+ 20,
+ 15
+ ],
+ "category_id": 4,
+ "id": 1287,
+ "area": 300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 654,
+ "bbox": [
+ 250,
+ 230,
+ 16,
+ 14
+ ],
+ "category_id": 4,
+ "id": 1288,
+ "area": 224,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 654,
+ "bbox": [
+ 469,
+ 196,
+ 23,
+ 15
+ ],
+ "category_id": 4,
+ "id": 1289,
+ "area": 345,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 654,
+ "bbox": [
+ 51,
+ 295,
+ 22,
+ 14
+ ],
+ "category_id": 4,
+ "id": 1290,
+ "area": 308,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 654,
+ "bbox": [
+ 51,
+ 410,
+ 20,
+ 13
+ ],
+ "category_id": 4,
+ "id": 1291,
+ "area": 260,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 654,
+ "bbox": [
+ 232,
+ 376,
+ 19,
+ 17
+ ],
+ "category_id": 4,
+ "id": 1292,
+ "area": 323,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 654,
+ "bbox": [
+ 454,
+ 306,
+ 17,
+ 15
+ ],
+ "category_id": 4,
+ "id": 1293,
+ "area": 255,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 654,
+ "bbox": [
+ 459,
+ 417,
+ 24,
+ 18
+ ],
+ "category_id": 4,
+ "id": 1294,
+ "area": 432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 655,
+ "bbox": [
+ 14,
+ 268,
+ 66,
+ 18
+ ],
+ "category_id": 16,
+ "id": 1295,
+ "area": 1188,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 656,
+ "bbox": [
+ 128,
+ 160,
+ 250,
+ 212
+ ],
+ "category_id": 16,
+ "id": 1296,
+ "area": 53000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 657,
+ "bbox": [
+ 60,
+ 287,
+ 46,
+ 34
+ ],
+ "category_id": 15,
+ "id": 1297,
+ "area": 1564,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 657,
+ "bbox": [
+ 62,
+ 337,
+ 48,
+ 37
+ ],
+ "category_id": 15,
+ "id": 1298,
+ "area": 1776,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 657,
+ "bbox": [
+ 208,
+ 318,
+ 26,
+ 18
+ ],
+ "category_id": 15,
+ "id": 1299,
+ "area": 468,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 657,
+ "bbox": [
+ 234,
+ 319,
+ 25,
+ 20
+ ],
+ "category_id": 15,
+ "id": 1300,
+ "area": 500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 657,
+ "bbox": [
+ 330,
+ 361,
+ 27,
+ 23
+ ],
+ "category_id": 15,
+ "id": 1301,
+ "area": 621,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 657,
+ "bbox": [
+ 351,
+ 375,
+ 27,
+ 21
+ ],
+ "category_id": 15,
+ "id": 1302,
+ "area": 567,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 658,
+ "bbox": [
+ 127,
+ 176,
+ 180,
+ 60
+ ],
+ "category_id": 19,
+ "id": 1303,
+ "area": 10800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 659,
+ "bbox": [
+ 143,
+ 137,
+ 253,
+ 340
+ ],
+ "category_id": 16,
+ "id": 1304,
+ "area": 86020,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 660,
+ "bbox": [
+ 102,
+ 128,
+ 156,
+ 194
+ ],
+ "category_id": 15,
+ "id": 1305,
+ "area": 30264,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 660,
+ "bbox": [
+ 293,
+ 175,
+ 155,
+ 188
+ ],
+ "category_id": 15,
+ "id": 1306,
+ "area": 29140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 661,
+ "bbox": [
+ 80,
+ 193,
+ 369,
+ 144
+ ],
+ "category_id": 19,
+ "id": 1307,
+ "area": 53136,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 662,
+ "bbox": [
+ 192,
+ 81,
+ 189,
+ 338
+ ],
+ "category_id": 10,
+ "id": 1308,
+ "area": 63882,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 663,
+ "bbox": [
+ 51,
+ 243,
+ 48,
+ 36
+ ],
+ "category_id": 4,
+ "id": 1309,
+ "area": 1728,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 663,
+ "bbox": [
+ 427,
+ 277,
+ 37,
+ 47
+ ],
+ "category_id": 4,
+ "id": 1310,
+ "area": 1739,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 664,
+ "bbox": [
+ 311,
+ 91,
+ 50,
+ 325
+ ],
+ "category_id": 10,
+ "id": 1311,
+ "area": 16250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 665,
+ "bbox": [
+ 151,
+ 201,
+ 261,
+ 98
+ ],
+ "category_id": 16,
+ "id": 1312,
+ "area": 25578,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 666,
+ "bbox": [
+ 141,
+ 97,
+ 143,
+ 267
+ ],
+ "category_id": 16,
+ "id": 1313,
+ "area": 38181,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 667,
+ "bbox": [
+ 254,
+ 188,
+ 192,
+ 265
+ ],
+ "category_id": 10,
+ "id": 1314,
+ "area": 50880,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 667,
+ "bbox": [
+ 377,
+ 33,
+ 72,
+ 127
+ ],
+ "category_id": 10,
+ "id": 1315,
+ "area": 9144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 668,
+ "bbox": [
+ 442,
+ 164,
+ 32,
+ 17
+ ],
+ "category_id": 16,
+ "id": 1316,
+ "area": 544,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 669,
+ "bbox": [
+ 392,
+ 46,
+ 16,
+ 12
+ ],
+ "category_id": 4,
+ "id": 1317,
+ "area": 192,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 669,
+ "bbox": [
+ 270,
+ 147,
+ 17,
+ 12
+ ],
+ "category_id": 4,
+ "id": 1318,
+ "area": 204,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 669,
+ "bbox": [
+ 435,
+ 212,
+ 14,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1319,
+ "area": 140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 669,
+ "bbox": [
+ 203,
+ 238,
+ 17,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1320,
+ "area": 170,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 669,
+ "bbox": [
+ 83,
+ 277,
+ 18,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1321,
+ "area": 198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 669,
+ "bbox": [
+ 22,
+ 420,
+ 15,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1322,
+ "area": 135,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 669,
+ "bbox": [
+ 352,
+ 322,
+ 16,
+ 12
+ ],
+ "category_id": 4,
+ "id": 1323,
+ "area": 192,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 669,
+ "bbox": [
+ 373,
+ 444,
+ 16,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1324,
+ "area": 176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 669,
+ "bbox": [
+ 187,
+ 480,
+ 15,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1325,
+ "area": 150,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 670,
+ "bbox": [
+ 225,
+ 133,
+ 159,
+ 142
+ ],
+ "category_id": 15,
+ "id": 1326,
+ "area": 22578,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 670,
+ "bbox": [
+ 58,
+ 200,
+ 160,
+ 141
+ ],
+ "category_id": 15,
+ "id": 1327,
+ "area": 22560,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 671,
+ "bbox": [
+ 239,
+ 176,
+ 71,
+ 149
+ ],
+ "category_id": 4,
+ "id": 1328,
+ "area": 10579,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 672,
+ "bbox": [
+ 127,
+ 317,
+ 56,
+ 67
+ ],
+ "category_id": 4,
+ "id": 1329,
+ "area": 3752,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 672,
+ "bbox": [
+ 372,
+ 131,
+ 46,
+ 49
+ ],
+ "category_id": 4,
+ "id": 1330,
+ "area": 2254,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 673,
+ "bbox": [
+ 150,
+ 115,
+ 295,
+ 269
+ ],
+ "category_id": 10,
+ "id": 1331,
+ "area": 79355,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 674,
+ "bbox": [
+ 78,
+ 240,
+ 97,
+ 136
+ ],
+ "category_id": 19,
+ "id": 1332,
+ "area": 13192,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 675,
+ "bbox": [
+ 127,
+ 405,
+ 16,
+ 32
+ ],
+ "category_id": 4,
+ "id": 1333,
+ "area": 512,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 675,
+ "bbox": [
+ 423,
+ 344,
+ 12,
+ 37
+ ],
+ "category_id": 4,
+ "id": 1334,
+ "area": 444,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 675,
+ "bbox": [
+ 375,
+ 48,
+ 13,
+ 33
+ ],
+ "category_id": 4,
+ "id": 1335,
+ "area": 429,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 676,
+ "bbox": [
+ 108,
+ 211,
+ 255,
+ 60
+ ],
+ "category_id": 19,
+ "id": 1336,
+ "area": 15300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 677,
+ "bbox": [
+ 209,
+ 181,
+ 139,
+ 199
+ ],
+ "category_id": 10,
+ "id": 1337,
+ "area": 27661,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 678,
+ "bbox": [
+ 248,
+ 87,
+ 52,
+ 284
+ ],
+ "category_id": 19,
+ "id": 1338,
+ "area": 14768,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 679,
+ "bbox": [
+ 52,
+ 261,
+ 9,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1339,
+ "area": 81,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 679,
+ "bbox": [
+ 208,
+ 240,
+ 10,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1340,
+ "area": 80,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 679,
+ "bbox": [
+ 193,
+ 304,
+ 11,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1341,
+ "area": 88,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 679,
+ "bbox": [
+ 122,
+ 293,
+ 11,
+ 7
+ ],
+ "category_id": 4,
+ "id": 1342,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 679,
+ "bbox": [
+ 472,
+ 419,
+ 12,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1343,
+ "area": 96,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 679,
+ "bbox": [
+ 352,
+ 230,
+ 8,
+ 6
+ ],
+ "category_id": 4,
+ "id": 1344,
+ "area": 48,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 679,
+ "bbox": [
+ 399,
+ 179,
+ 9,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1345,
+ "area": 81,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 679,
+ "bbox": [
+ 442,
+ 119,
+ 6,
+ 7
+ ],
+ "category_id": 4,
+ "id": 1346,
+ "area": 42,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 680,
+ "bbox": [
+ 197,
+ 122,
+ 62,
+ 226
+ ],
+ "category_id": 19,
+ "id": 1347,
+ "area": 14012,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 681,
+ "bbox": [
+ 118,
+ 298,
+ 249,
+ 39
+ ],
+ "category_id": 19,
+ "id": 1348,
+ "area": 9711,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 682,
+ "bbox": [
+ 20,
+ 148,
+ 20,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1349,
+ "area": 480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 682,
+ "bbox": [
+ 336,
+ 72,
+ 17,
+ 31
+ ],
+ "category_id": 4,
+ "id": 1350,
+ "area": 527,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 682,
+ "bbox": [
+ 378,
+ 243,
+ 18,
+ 30
+ ],
+ "category_id": 4,
+ "id": 1351,
+ "area": 540,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 682,
+ "bbox": [
+ 319,
+ 428,
+ 14,
+ 31
+ ],
+ "category_id": 4,
+ "id": 1352,
+ "area": 434,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 682,
+ "bbox": [
+ 477,
+ 404,
+ 20,
+ 34
+ ],
+ "category_id": 4,
+ "id": 1353,
+ "area": 680,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 683,
+ "bbox": [
+ 51,
+ 152,
+ 193,
+ 235
+ ],
+ "category_id": 16,
+ "id": 1354,
+ "area": 45355,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 684,
+ "bbox": [
+ 40,
+ 21,
+ 21,
+ 32
+ ],
+ "category_id": 4,
+ "id": 1355,
+ "area": 672,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 684,
+ "bbox": [
+ 40,
+ 142,
+ 20,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1356,
+ "area": 500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 684,
+ "bbox": [
+ 39,
+ 266,
+ 17,
+ 27
+ ],
+ "category_id": 4,
+ "id": 1357,
+ "area": 459,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 684,
+ "bbox": [
+ 39,
+ 389,
+ 17,
+ 21
+ ],
+ "category_id": 4,
+ "id": 1358,
+ "area": 357,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 684,
+ "bbox": [
+ 479,
+ 85,
+ 17,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1359,
+ "area": 425,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 684,
+ "bbox": [
+ 475,
+ 208,
+ 19,
+ 27
+ ],
+ "category_id": 4,
+ "id": 1360,
+ "area": 513,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 684,
+ "bbox": [
+ 476,
+ 328,
+ 16,
+ 28
+ ],
+ "category_id": 4,
+ "id": 1361,
+ "area": 448,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 684,
+ "bbox": [
+ 474,
+ 451,
+ 16,
+ 29
+ ],
+ "category_id": 4,
+ "id": 1362,
+ "area": 464,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 685,
+ "bbox": [
+ 272,
+ 56,
+ 46,
+ 43
+ ],
+ "category_id": 4,
+ "id": 1363,
+ "area": 1978,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 685,
+ "bbox": [
+ 252,
+ 371,
+ 43,
+ 45
+ ],
+ "category_id": 4,
+ "id": 1364,
+ "area": 1935,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 686,
+ "bbox": [
+ 58,
+ 236,
+ 433,
+ 92
+ ],
+ "category_id": 10,
+ "id": 1365,
+ "area": 39836,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 687,
+ "bbox": [
+ 212,
+ 165,
+ 91,
+ 267
+ ],
+ "category_id": 10,
+ "id": 1366,
+ "area": 24297,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 688,
+ "bbox": [
+ 119,
+ 187,
+ 358,
+ 255
+ ],
+ "category_id": 10,
+ "id": 1367,
+ "area": 91290,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 689,
+ "bbox": [
+ 186,
+ 156,
+ 153,
+ 246
+ ],
+ "category_id": 16,
+ "id": 1368,
+ "area": 37638,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 690,
+ "bbox": [
+ 300,
+ 254,
+ 92,
+ 82
+ ],
+ "category_id": 4,
+ "id": 1369,
+ "area": 7544,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 691,
+ "bbox": [
+ 82,
+ 126,
+ 324,
+ 322
+ ],
+ "category_id": 10,
+ "id": 1370,
+ "area": 104328,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 692,
+ "bbox": [
+ 73,
+ 251,
+ 19,
+ 12
+ ],
+ "category_id": 4,
+ "id": 1371,
+ "area": 228,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 692,
+ "bbox": [
+ 25,
+ 368,
+ 31,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1372,
+ "area": 744,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 692,
+ "bbox": [
+ 232,
+ 330,
+ 22,
+ 28
+ ],
+ "category_id": 4,
+ "id": 1373,
+ "area": 616,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 692,
+ "bbox": [
+ 238,
+ 151,
+ 19,
+ 17
+ ],
+ "category_id": 4,
+ "id": 1374,
+ "area": 323,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 692,
+ "bbox": [
+ 481,
+ 259,
+ 20,
+ 14
+ ],
+ "category_id": 4,
+ "id": 1375,
+ "area": 280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 693,
+ "bbox": [
+ 269,
+ 310,
+ 175,
+ 174
+ ],
+ "category_id": 10,
+ "id": 1376,
+ "area": 30450,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 694,
+ "bbox": [
+ 35,
+ 159,
+ 413,
+ 182
+ ],
+ "category_id": 10,
+ "id": 1377,
+ "area": 75166,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 695,
+ "bbox": [
+ 119,
+ 237,
+ 241,
+ 79
+ ],
+ "category_id": 19,
+ "id": 1378,
+ "area": 19039,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 696,
+ "bbox": [
+ 48,
+ 83,
+ 391,
+ 287
+ ],
+ "category_id": 16,
+ "id": 1379,
+ "area": 112217,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 697,
+ "bbox": [
+ 168,
+ 115,
+ 237,
+ 318
+ ],
+ "category_id": 10,
+ "id": 1380,
+ "area": 75366,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 698,
+ "bbox": [
+ 90,
+ 265,
+ 44,
+ 44
+ ],
+ "category_id": 4,
+ "id": 1381,
+ "area": 1936,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 698,
+ "bbox": [
+ 436,
+ 344,
+ 42,
+ 35
+ ],
+ "category_id": 4,
+ "id": 1382,
+ "area": 1470,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 699,
+ "bbox": [
+ 197,
+ 107,
+ 174,
+ 320
+ ],
+ "category_id": 16,
+ "id": 1383,
+ "area": 55680,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 700,
+ "bbox": [
+ 379,
+ 57,
+ 122,
+ 76
+ ],
+ "category_id": 19,
+ "id": 1384,
+ "area": 9272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 701,
+ "bbox": [
+ 336,
+ 97,
+ 14,
+ 15
+ ],
+ "category_id": 4,
+ "id": 1385,
+ "area": 210,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 701,
+ "bbox": [
+ 247,
+ 168,
+ 16,
+ 19
+ ],
+ "category_id": 4,
+ "id": 1386,
+ "area": 304,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 701,
+ "bbox": [
+ 183,
+ 259,
+ 15,
+ 13
+ ],
+ "category_id": 4,
+ "id": 1387,
+ "area": 195,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 701,
+ "bbox": [
+ 421,
+ 238,
+ 13,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1388,
+ "area": 117,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 701,
+ "bbox": [
+ 94,
+ 440,
+ 15,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1389,
+ "area": 165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 701,
+ "bbox": [
+ 272,
+ 440,
+ 16,
+ 7
+ ],
+ "category_id": 4,
+ "id": 1390,
+ "area": 112,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 702,
+ "bbox": [
+ 115,
+ 115,
+ 288,
+ 236
+ ],
+ "category_id": 15,
+ "id": 1391,
+ "area": 67968,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 702,
+ "bbox": [
+ 18,
+ 226,
+ 208,
+ 119
+ ],
+ "category_id": 15,
+ "id": 1392,
+ "area": 24752,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 703,
+ "bbox": [
+ 97,
+ 177,
+ 302,
+ 108
+ ],
+ "category_id": 19,
+ "id": 1393,
+ "area": 32616,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 704,
+ "bbox": [
+ 173,
+ 328,
+ 86,
+ 86
+ ],
+ "category_id": 10,
+ "id": 1394,
+ "area": 7396,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 705,
+ "bbox": [
+ 122,
+ 99,
+ 249,
+ 221
+ ],
+ "category_id": 16,
+ "id": 1395,
+ "area": 55029,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 706,
+ "bbox": [
+ 184,
+ 218,
+ 116,
+ 93
+ ],
+ "category_id": 15,
+ "id": 1396,
+ "area": 10788,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 706,
+ "bbox": [
+ 203,
+ 378,
+ 61,
+ 16
+ ],
+ "category_id": 15,
+ "id": 1397,
+ "area": 976,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 707,
+ "bbox": [
+ 56,
+ 161,
+ 415,
+ 101
+ ],
+ "category_id": 10,
+ "id": 1398,
+ "area": 41915,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 708,
+ "bbox": [
+ 94,
+ 39,
+ 207,
+ 350
+ ],
+ "category_id": 16,
+ "id": 1399,
+ "area": 72450,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 709,
+ "bbox": [
+ 105,
+ 148,
+ 22,
+ 33
+ ],
+ "category_id": 4,
+ "id": 1400,
+ "area": 726,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 709,
+ "bbox": [
+ 132,
+ 221,
+ 23,
+ 27
+ ],
+ "category_id": 4,
+ "id": 1401,
+ "area": 621,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 709,
+ "bbox": [
+ 26,
+ 411,
+ 24,
+ 28
+ ],
+ "category_id": 4,
+ "id": 1402,
+ "area": 672,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 709,
+ "bbox": [
+ 256,
+ 270,
+ 24,
+ 29
+ ],
+ "category_id": 4,
+ "id": 1403,
+ "area": 696,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 709,
+ "bbox": [
+ 372,
+ 234,
+ 27,
+ 29
+ ],
+ "category_id": 4,
+ "id": 1404,
+ "area": 783,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 709,
+ "bbox": [
+ 474,
+ 241,
+ 24,
+ 27
+ ],
+ "category_id": 4,
+ "id": 1405,
+ "area": 648,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 710,
+ "bbox": [
+ 183,
+ 131,
+ 74,
+ 108
+ ],
+ "category_id": 19,
+ "id": 1406,
+ "area": 7992,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 711,
+ "bbox": [
+ 373,
+ 27,
+ 3,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1407,
+ "area": 27,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 711,
+ "bbox": [
+ 375,
+ 125,
+ 1,
+ 6
+ ],
+ "category_id": 4,
+ "id": 1408,
+ "area": 6,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 711,
+ "bbox": [
+ 384,
+ 75,
+ 6,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1409,
+ "area": 54,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 711,
+ "bbox": [
+ 356,
+ 173,
+ 4,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1410,
+ "area": 44,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 711,
+ "bbox": [
+ 352,
+ 213,
+ 5,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1411,
+ "area": 50,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 711,
+ "bbox": [
+ 302,
+ 229,
+ 9,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1412,
+ "area": 99,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 711,
+ "bbox": [
+ 224,
+ 257,
+ 5,
+ 13
+ ],
+ "category_id": 4,
+ "id": 1413,
+ "area": 65,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 711,
+ "bbox": [
+ 172,
+ 220,
+ 6,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1414,
+ "area": 48,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 711,
+ "bbox": [
+ 216,
+ 318,
+ 6,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1415,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 711,
+ "bbox": [
+ 92,
+ 368,
+ 9,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1416,
+ "area": 72,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 711,
+ "bbox": [
+ 141,
+ 372,
+ 6,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1417,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 711,
+ "bbox": [
+ 178,
+ 339,
+ 4,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1418,
+ "area": 32,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 711,
+ "bbox": [
+ 67,
+ 398,
+ 5,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1419,
+ "area": 45,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 712,
+ "bbox": [
+ 88,
+ 113,
+ 378,
+ 324
+ ],
+ "category_id": 15,
+ "id": 1420,
+ "area": 122472,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 712,
+ "bbox": [
+ 359,
+ 35,
+ 153,
+ 78
+ ],
+ "category_id": 15,
+ "id": 1421,
+ "area": 11934,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 713,
+ "bbox": [
+ 304,
+ 132,
+ 26,
+ 29
+ ],
+ "category_id": 4,
+ "id": 1422,
+ "area": 754,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 713,
+ "bbox": [
+ 213,
+ 380,
+ 27,
+ 30
+ ],
+ "category_id": 4,
+ "id": 1423,
+ "area": 810,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 714,
+ "bbox": [
+ 291,
+ 202,
+ 105,
+ 105
+ ],
+ "category_id": 4,
+ "id": 1424,
+ "area": 11025,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 715,
+ "bbox": [
+ 162,
+ 196,
+ 191,
+ 48
+ ],
+ "category_id": 19,
+ "id": 1425,
+ "area": 9168,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 716,
+ "bbox": [
+ 165,
+ 185,
+ 314,
+ 169
+ ],
+ "category_id": 10,
+ "id": 1426,
+ "area": 53066,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 717,
+ "bbox": [
+ 140,
+ 202,
+ 121,
+ 183
+ ],
+ "category_id": 15,
+ "id": 1427,
+ "area": 22143,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 717,
+ "bbox": [
+ 284,
+ 211,
+ 122,
+ 183
+ ],
+ "category_id": 15,
+ "id": 1428,
+ "area": 22326,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 718,
+ "bbox": [
+ 209,
+ 155,
+ 180,
+ 207
+ ],
+ "category_id": 10,
+ "id": 1429,
+ "area": 37260,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 719,
+ "bbox": [
+ 161,
+ 65,
+ 114,
+ 336
+ ],
+ "category_id": 19,
+ "id": 1430,
+ "area": 38304,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 720,
+ "bbox": [
+ 95,
+ 92,
+ 320,
+ 304
+ ],
+ "category_id": 16,
+ "id": 1431,
+ "area": 97280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 721,
+ "bbox": [
+ 201,
+ 101,
+ 108,
+ 119
+ ],
+ "category_id": 15,
+ "id": 1432,
+ "area": 12852,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 721,
+ "bbox": [
+ 113,
+ 210,
+ 108,
+ 122
+ ],
+ "category_id": 15,
+ "id": 1433,
+ "area": 13176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 721,
+ "bbox": [
+ 252,
+ 270,
+ 81,
+ 101
+ ],
+ "category_id": 15,
+ "id": 1434,
+ "area": 8181,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 722,
+ "bbox": [
+ 119,
+ 66,
+ 89,
+ 52
+ ],
+ "category_id": 4,
+ "id": 1435,
+ "area": 4628,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 722,
+ "bbox": [
+ 372,
+ 364,
+ 71,
+ 52
+ ],
+ "category_id": 4,
+ "id": 1436,
+ "area": 3692,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 723,
+ "bbox": [
+ 144,
+ 101,
+ 206,
+ 362
+ ],
+ "category_id": 16,
+ "id": 1437,
+ "area": 74572,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 724,
+ "bbox": [
+ 226,
+ 120,
+ 62,
+ 144
+ ],
+ "category_id": 19,
+ "id": 1438,
+ "area": 8928,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 725,
+ "bbox": [
+ 71,
+ 172,
+ 379,
+ 198
+ ],
+ "category_id": 16,
+ "id": 1439,
+ "area": 75042,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 726,
+ "bbox": [
+ 284,
+ 172,
+ 43,
+ 251
+ ],
+ "category_id": 15,
+ "id": 1440,
+ "area": 10793,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 727,
+ "bbox": [
+ 189,
+ 103,
+ 162,
+ 327
+ ],
+ "category_id": 10,
+ "id": 1441,
+ "area": 52974,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 728,
+ "bbox": [
+ 44,
+ 108,
+ 52,
+ 27
+ ],
+ "category_id": 4,
+ "id": 1442,
+ "area": 1404,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 728,
+ "bbox": [
+ 209,
+ 382,
+ 57,
+ 29
+ ],
+ "category_id": 4,
+ "id": 1443,
+ "area": 1653,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 728,
+ "bbox": [
+ 430,
+ 191,
+ 60,
+ 32
+ ],
+ "category_id": 4,
+ "id": 1444,
+ "area": 1920,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 729,
+ "bbox": [
+ 350,
+ 265,
+ 138,
+ 170
+ ],
+ "category_id": 10,
+ "id": 1445,
+ "area": 23460,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 730,
+ "bbox": [
+ 227,
+ 65,
+ 94,
+ 57
+ ],
+ "category_id": 4,
+ "id": 1446,
+ "area": 5358,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 730,
+ "bbox": [
+ 253,
+ 373,
+ 78,
+ 57
+ ],
+ "category_id": 4,
+ "id": 1447,
+ "area": 4446,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 731,
+ "bbox": [
+ 136,
+ 136,
+ 130,
+ 192
+ ],
+ "category_id": 16,
+ "id": 1448,
+ "area": 24960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 732,
+ "bbox": [
+ 44,
+ 128,
+ 41,
+ 58
+ ],
+ "category_id": 4,
+ "id": 1449,
+ "area": 2378,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 732,
+ "bbox": [
+ 381,
+ 340,
+ 58,
+ 52
+ ],
+ "category_id": 4,
+ "id": 1450,
+ "area": 3016,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 733,
+ "bbox": [
+ 239,
+ 186,
+ 51,
+ 43
+ ],
+ "category_id": 4,
+ "id": 1451,
+ "area": 2193,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 733,
+ "bbox": [
+ 328,
+ 417,
+ 59,
+ 47
+ ],
+ "category_id": 4,
+ "id": 1452,
+ "area": 2773,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 734,
+ "bbox": [
+ 64,
+ 160,
+ 92,
+ 116
+ ],
+ "category_id": 15,
+ "id": 1453,
+ "area": 10672,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 734,
+ "bbox": [
+ 47,
+ 290,
+ 90,
+ 115
+ ],
+ "category_id": 15,
+ "id": 1454,
+ "area": 10350,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 734,
+ "bbox": [
+ 380,
+ 291,
+ 21,
+ 105
+ ],
+ "category_id": 15,
+ "id": 1455,
+ "area": 2205,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 735,
+ "bbox": [
+ 465,
+ 480,
+ 18,
+ 12
+ ],
+ "category_id": 15,
+ "id": 1456,
+ "area": 216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 735,
+ "bbox": [
+ 490,
+ 471,
+ 20,
+ 13
+ ],
+ "category_id": 15,
+ "id": 1457,
+ "area": 260,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 736,
+ "bbox": [
+ 291,
+ 105,
+ 123,
+ 143
+ ],
+ "category_id": 19,
+ "id": 1458,
+ "area": 17589,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 737,
+ "bbox": [
+ 47,
+ 141,
+ 34,
+ 18
+ ],
+ "category_id": 4,
+ "id": 1459,
+ "area": 612,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 737,
+ "bbox": [
+ 265,
+ 85,
+ 42,
+ 20
+ ],
+ "category_id": 4,
+ "id": 1460,
+ "area": 840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 737,
+ "bbox": [
+ 445,
+ 215,
+ 46,
+ 21
+ ],
+ "category_id": 4,
+ "id": 1461,
+ "area": 966,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 737,
+ "bbox": [
+ 366,
+ 353,
+ 35,
+ 21
+ ],
+ "category_id": 4,
+ "id": 1462,
+ "area": 735,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 738,
+ "bbox": [
+ 247,
+ 94,
+ 67,
+ 318
+ ],
+ "category_id": 10,
+ "id": 1463,
+ "area": 21306,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 739,
+ "bbox": [
+ 194,
+ 180,
+ 80,
+ 38
+ ],
+ "category_id": 15,
+ "id": 1464,
+ "area": 3040,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 739,
+ "bbox": [
+ 206,
+ 371,
+ 79,
+ 35
+ ],
+ "category_id": 15,
+ "id": 1465,
+ "area": 2765,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 740,
+ "bbox": [
+ 119,
+ 62,
+ 49,
+ 21
+ ],
+ "category_id": 4,
+ "id": 1466,
+ "area": 1029,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 740,
+ "bbox": [
+ 201,
+ 223,
+ 44,
+ 30
+ ],
+ "category_id": 4,
+ "id": 1467,
+ "area": 1320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 740,
+ "bbox": [
+ 390,
+ 386,
+ 38,
+ 28
+ ],
+ "category_id": 4,
+ "id": 1468,
+ "area": 1064,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 741,
+ "bbox": [
+ 249,
+ 138,
+ 117,
+ 144
+ ],
+ "category_id": 15,
+ "id": 1469,
+ "area": 16848,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 741,
+ "bbox": [
+ 457,
+ 19,
+ 55,
+ 118
+ ],
+ "category_id": 15,
+ "id": 1470,
+ "area": 6490,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 742,
+ "bbox": [
+ 83,
+ 319,
+ 43,
+ 49
+ ],
+ "category_id": 4,
+ "id": 1471,
+ "area": 2107,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 742,
+ "bbox": [
+ 417,
+ 192,
+ 37,
+ 55
+ ],
+ "category_id": 4,
+ "id": 1472,
+ "area": 2035,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 743,
+ "bbox": [
+ 248,
+ 300,
+ 50,
+ 37
+ ],
+ "category_id": 15,
+ "id": 1473,
+ "area": 1850,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 743,
+ "bbox": [
+ 135,
+ 363,
+ 85,
+ 33
+ ],
+ "category_id": 15,
+ "id": 1474,
+ "area": 2805,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 744,
+ "bbox": [
+ 57,
+ 35,
+ 124,
+ 53
+ ],
+ "category_id": 10,
+ "id": 1475,
+ "area": 6572,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 745,
+ "bbox": [
+ 52,
+ 223,
+ 447,
+ 114
+ ],
+ "category_id": 10,
+ "id": 1476,
+ "area": 50958,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 746,
+ "bbox": [
+ 268,
+ 231,
+ 53,
+ 158
+ ],
+ "category_id": 16,
+ "id": 1477,
+ "area": 8374,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 747,
+ "bbox": [
+ 338,
+ 87,
+ 65,
+ 33
+ ],
+ "category_id": 4,
+ "id": 1478,
+ "area": 2145,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 747,
+ "bbox": [
+ 216,
+ 280,
+ 63,
+ 41
+ ],
+ "category_id": 4,
+ "id": 1479,
+ "area": 2583,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 747,
+ "bbox": [
+ 158,
+ 456,
+ 49,
+ 34
+ ],
+ "category_id": 4,
+ "id": 1480,
+ "area": 1666,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 748,
+ "bbox": [
+ 63,
+ 8,
+ 25,
+ 26
+ ],
+ "category_id": 4,
+ "id": 1481,
+ "area": 650,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 748,
+ "bbox": [
+ 62,
+ 128,
+ 26,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1482,
+ "area": 624,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 748,
+ "bbox": [
+ 70,
+ 248,
+ 18,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1483,
+ "area": 432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 748,
+ "bbox": [
+ 61,
+ 366,
+ 22,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1484,
+ "area": 528,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 748,
+ "bbox": [
+ 51,
+ 476,
+ 26,
+ 26
+ ],
+ "category_id": 4,
+ "id": 1485,
+ "area": 676,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 749,
+ "bbox": [
+ 123,
+ 163,
+ 276,
+ 125
+ ],
+ "category_id": 10,
+ "id": 1486,
+ "area": 34500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 750,
+ "bbox": [
+ 82,
+ 159,
+ 355,
+ 139
+ ],
+ "category_id": 10,
+ "id": 1487,
+ "area": 49345,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 751,
+ "bbox": [
+ 388,
+ 313,
+ 77,
+ 170
+ ],
+ "category_id": 10,
+ "id": 1488,
+ "area": 13090,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 752,
+ "bbox": [
+ 53,
+ 229,
+ 380,
+ 76
+ ],
+ "category_id": 10,
+ "id": 1489,
+ "area": 28880,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 753,
+ "bbox": [
+ 79,
+ 120,
+ 25,
+ 54
+ ],
+ "category_id": 4,
+ "id": 1490,
+ "area": 1350,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 753,
+ "bbox": [
+ 303,
+ 229,
+ 24,
+ 55
+ ],
+ "category_id": 4,
+ "id": 1491,
+ "area": 1320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 753,
+ "bbox": [
+ 312,
+ 461,
+ 25,
+ 49
+ ],
+ "category_id": 4,
+ "id": 1492,
+ "area": 1225,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 754,
+ "bbox": [
+ 129,
+ 103,
+ 323,
+ 350
+ ],
+ "category_id": 10,
+ "id": 1493,
+ "area": 113050,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 755,
+ "bbox": [
+ 240,
+ 107,
+ 195,
+ 328
+ ],
+ "category_id": 10,
+ "id": 1494,
+ "area": 63960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 756,
+ "bbox": [
+ 120,
+ 254,
+ 27,
+ 33
+ ],
+ "category_id": 4,
+ "id": 1495,
+ "area": 891,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 756,
+ "bbox": [
+ 456,
+ 208,
+ 27,
+ 28
+ ],
+ "category_id": 4,
+ "id": 1496,
+ "area": 756,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 757,
+ "bbox": [
+ 81,
+ 423,
+ 11,
+ 73
+ ],
+ "category_id": 10,
+ "id": 1497,
+ "area": 803,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 758,
+ "bbox": [
+ 257,
+ 178,
+ 155,
+ 154
+ ],
+ "category_id": 16,
+ "id": 1498,
+ "area": 23870,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 759,
+ "bbox": [
+ 197,
+ 125,
+ 222,
+ 266
+ ],
+ "category_id": 10,
+ "id": 1499,
+ "area": 59052,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 760,
+ "bbox": [
+ 73,
+ 185,
+ 371,
+ 220
+ ],
+ "category_id": 16,
+ "id": 1500,
+ "area": 81620,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 761,
+ "bbox": [
+ 171,
+ 94,
+ 86,
+ 348
+ ],
+ "category_id": 19,
+ "id": 1501,
+ "area": 29928,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 762,
+ "bbox": [
+ 26,
+ 335,
+ 71,
+ 63
+ ],
+ "category_id": 4,
+ "id": 1502,
+ "area": 4473,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 762,
+ "bbox": [
+ 392,
+ 224,
+ 80,
+ 58
+ ],
+ "category_id": 4,
+ "id": 1503,
+ "area": 4640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 763,
+ "bbox": [
+ 404,
+ 344,
+ 68,
+ 133
+ ],
+ "category_id": 10,
+ "id": 1504,
+ "area": 9044,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 764,
+ "bbox": [
+ 472,
+ 150,
+ 16,
+ 15
+ ],
+ "category_id": 4,
+ "id": 1505,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 764,
+ "bbox": [
+ 285,
+ 109,
+ 18,
+ 13
+ ],
+ "category_id": 4,
+ "id": 1506,
+ "area": 234,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 764,
+ "bbox": [
+ 95,
+ 147,
+ 18,
+ 15
+ ],
+ "category_id": 4,
+ "id": 1507,
+ "area": 270,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 764,
+ "bbox": [
+ 265,
+ 325,
+ 22,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1508,
+ "area": 242,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 764,
+ "bbox": [
+ 385,
+ 459,
+ 18,
+ 12
+ ],
+ "category_id": 4,
+ "id": 1509,
+ "area": 216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 764,
+ "bbox": [
+ 195,
+ 476,
+ 21,
+ 14
+ ],
+ "category_id": 4,
+ "id": 1510,
+ "area": 294,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 765,
+ "bbox": [
+ 325,
+ 121,
+ 123,
+ 313
+ ],
+ "category_id": 10,
+ "id": 1511,
+ "area": 38499,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 766,
+ "bbox": [
+ 166,
+ 202,
+ 84,
+ 128
+ ],
+ "category_id": 15,
+ "id": 1512,
+ "area": 10752,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 766,
+ "bbox": [
+ 37,
+ 280,
+ 41,
+ 109
+ ],
+ "category_id": 15,
+ "id": 1513,
+ "area": 4469,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 767,
+ "bbox": [
+ 374,
+ 393,
+ 42,
+ 101
+ ],
+ "category_id": 19,
+ "id": 1514,
+ "area": 4242,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 768,
+ "bbox": [
+ 105,
+ 108,
+ 7,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1515,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 768,
+ "bbox": [
+ 72,
+ 167,
+ 8,
+ 12
+ ],
+ "category_id": 4,
+ "id": 1516,
+ "area": 96,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 768,
+ "bbox": [
+ 179,
+ 145,
+ 7,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1517,
+ "area": 70,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 768,
+ "bbox": [
+ 275,
+ 232,
+ 7,
+ 12
+ ],
+ "category_id": 4,
+ "id": 1518,
+ "area": 84,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 768,
+ "bbox": [
+ 431,
+ 179,
+ 4,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1519,
+ "area": 32,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 768,
+ "bbox": [
+ 463,
+ 83,
+ 4,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1520,
+ "area": 32,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 768,
+ "bbox": [
+ 421,
+ 43,
+ 6,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1521,
+ "area": 48,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 768,
+ "bbox": [
+ 318,
+ 91,
+ 5,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1522,
+ "area": 45,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 768,
+ "bbox": [
+ 280,
+ 63,
+ 8,
+ 7
+ ],
+ "category_id": 4,
+ "id": 1523,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 768,
+ "bbox": [
+ 245,
+ 37,
+ 6,
+ 5
+ ],
+ "category_id": 4,
+ "id": 1524,
+ "area": 30,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 768,
+ "bbox": [
+ 343,
+ 309,
+ 5,
+ 7
+ ],
+ "category_id": 4,
+ "id": 1525,
+ "area": 35,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 768,
+ "bbox": [
+ 424,
+ 471,
+ 9,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1526,
+ "area": 99,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 768,
+ "bbox": [
+ 331,
+ 413,
+ 8,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1527,
+ "area": 80,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 769,
+ "bbox": [
+ 178,
+ 35,
+ 76,
+ 390
+ ],
+ "category_id": 10,
+ "id": 1528,
+ "area": 29640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 770,
+ "bbox": [
+ 4,
+ 115,
+ 27,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1529,
+ "area": 675,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 770,
+ "bbox": [
+ 25,
+ 221,
+ 27,
+ 26
+ ],
+ "category_id": 4,
+ "id": 1530,
+ "area": 702,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 770,
+ "bbox": [
+ 49,
+ 317,
+ 25,
+ 29
+ ],
+ "category_id": 4,
+ "id": 1531,
+ "area": 725,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 770,
+ "bbox": [
+ 186,
+ 464,
+ 22,
+ 30
+ ],
+ "category_id": 4,
+ "id": 1532,
+ "area": 660,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 771,
+ "bbox": [
+ 255,
+ 134,
+ 131,
+ 214
+ ],
+ "category_id": 16,
+ "id": 1533,
+ "area": 28034,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 772,
+ "bbox": [
+ 87,
+ 186,
+ 352,
+ 47
+ ],
+ "category_id": 10,
+ "id": 1534,
+ "area": 16544,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 773,
+ "bbox": [
+ 84,
+ 223,
+ 330,
+ 189
+ ],
+ "category_id": 16,
+ "id": 1535,
+ "area": 62370,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 774,
+ "bbox": [
+ 176,
+ 167,
+ 147,
+ 118
+ ],
+ "category_id": 10,
+ "id": 1536,
+ "area": 17346,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 775,
+ "bbox": [
+ 45,
+ 160,
+ 417,
+ 159
+ ],
+ "category_id": 16,
+ "id": 1537,
+ "area": 66303,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 776,
+ "bbox": [
+ 120,
+ 122,
+ 112,
+ 130
+ ],
+ "category_id": 15,
+ "id": 1538,
+ "area": 14560,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 776,
+ "bbox": [
+ 138,
+ 277,
+ 112,
+ 133
+ ],
+ "category_id": 15,
+ "id": 1539,
+ "area": 14896,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 777,
+ "bbox": [
+ 126,
+ 338,
+ 41,
+ 50
+ ],
+ "category_id": 16,
+ "id": 1540,
+ "area": 2050,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 778,
+ "bbox": [
+ 125,
+ 270,
+ 35,
+ 44
+ ],
+ "category_id": 16,
+ "id": 1541,
+ "area": 1540,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 779,
+ "bbox": [
+ 142,
+ 134,
+ 254,
+ 267
+ ],
+ "category_id": 10,
+ "id": 1542,
+ "area": 67818,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 780,
+ "bbox": [
+ 20,
+ 181,
+ 411,
+ 263
+ ],
+ "category_id": 16,
+ "id": 1543,
+ "area": 108093,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 781,
+ "bbox": [
+ 127,
+ 143,
+ 141,
+ 202
+ ],
+ "category_id": 16,
+ "id": 1544,
+ "area": 28482,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 782,
+ "bbox": [
+ 60,
+ 72,
+ 368,
+ 331
+ ],
+ "category_id": 10,
+ "id": 1545,
+ "area": 121808,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 783,
+ "bbox": [
+ 167,
+ 105,
+ 82,
+ 83
+ ],
+ "category_id": 15,
+ "id": 1546,
+ "area": 6806,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 783,
+ "bbox": [
+ 66,
+ 175,
+ 108,
+ 109
+ ],
+ "category_id": 15,
+ "id": 1547,
+ "area": 11772,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 783,
+ "bbox": [
+ 396,
+ 103,
+ 76,
+ 66
+ ],
+ "category_id": 15,
+ "id": 1548,
+ "area": 5016,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 784,
+ "bbox": [
+ 58,
+ 206,
+ 33,
+ 32
+ ],
+ "category_id": 4,
+ "id": 1549,
+ "area": 1056,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 784,
+ "bbox": [
+ 412,
+ 136,
+ 36,
+ 29
+ ],
+ "category_id": 4,
+ "id": 1550,
+ "area": 1044,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 784,
+ "bbox": [
+ 350,
+ 399,
+ 35,
+ 26
+ ],
+ "category_id": 4,
+ "id": 1551,
+ "area": 910,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 785,
+ "bbox": [
+ 23,
+ 98,
+ 454,
+ 54
+ ],
+ "category_id": 10,
+ "id": 1552,
+ "area": 24516,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 786,
+ "bbox": [
+ 115,
+ 122,
+ 18,
+ 18
+ ],
+ "category_id": 4,
+ "id": 1553,
+ "area": 324,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 786,
+ "bbox": [
+ 323,
+ 108,
+ 20,
+ 18
+ ],
+ "category_id": 4,
+ "id": 1554,
+ "area": 360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 786,
+ "bbox": [
+ 424,
+ 101,
+ 23,
+ 19
+ ],
+ "category_id": 4,
+ "id": 1555,
+ "area": 437,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 786,
+ "bbox": [
+ 198,
+ 206,
+ 15,
+ 14
+ ],
+ "category_id": 4,
+ "id": 1556,
+ "area": 210,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 786,
+ "bbox": [
+ 151,
+ 302,
+ 16,
+ 17
+ ],
+ "category_id": 4,
+ "id": 1557,
+ "area": 272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 786,
+ "bbox": [
+ 105,
+ 435,
+ 15,
+ 14
+ ],
+ "category_id": 4,
+ "id": 1558,
+ "area": 210,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 786,
+ "bbox": [
+ 440,
+ 266,
+ 18,
+ 16
+ ],
+ "category_id": 4,
+ "id": 1559,
+ "area": 288,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 787,
+ "bbox": [
+ 169,
+ 56,
+ 60,
+ 38
+ ],
+ "category_id": 4,
+ "id": 1560,
+ "area": 2280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 787,
+ "bbox": [
+ 110,
+ 237,
+ 53,
+ 35
+ ],
+ "category_id": 4,
+ "id": 1561,
+ "area": 1855,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 787,
+ "bbox": [
+ 381,
+ 353,
+ 61,
+ 43
+ ],
+ "category_id": 4,
+ "id": 1562,
+ "area": 2623,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 788,
+ "bbox": [
+ 179,
+ 192,
+ 77,
+ 80
+ ],
+ "category_id": 15,
+ "id": 1563,
+ "area": 6160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 788,
+ "bbox": [
+ 427,
+ 274,
+ 28,
+ 33
+ ],
+ "category_id": 15,
+ "id": 1564,
+ "area": 924,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 789,
+ "bbox": [
+ 145,
+ 154,
+ 23,
+ 57
+ ],
+ "category_id": 4,
+ "id": 1565,
+ "area": 1311,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 789,
+ "bbox": [
+ 382,
+ 316,
+ 26,
+ 55
+ ],
+ "category_id": 4,
+ "id": 1566,
+ "area": 1430,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 790,
+ "bbox": [
+ 137,
+ 109,
+ 139,
+ 326
+ ],
+ "category_id": 16,
+ "id": 1567,
+ "area": 45314,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 791,
+ "bbox": [
+ 97,
+ 268,
+ 75,
+ 47
+ ],
+ "category_id": 4,
+ "id": 1568,
+ "area": 3525,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 791,
+ "bbox": [
+ 398,
+ 144,
+ 67,
+ 44
+ ],
+ "category_id": 4,
+ "id": 1569,
+ "area": 2948,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 792,
+ "bbox": [
+ 133,
+ 115,
+ 264,
+ 246
+ ],
+ "category_id": 19,
+ "id": 1570,
+ "area": 64944,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 793,
+ "bbox": [
+ 211,
+ 140,
+ 32,
+ 43
+ ],
+ "category_id": 4,
+ "id": 1571,
+ "area": 1376,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 793,
+ "bbox": [
+ 132,
+ 363,
+ 28,
+ 36
+ ],
+ "category_id": 4,
+ "id": 1572,
+ "area": 1008,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 793,
+ "bbox": [
+ 474,
+ 330,
+ 27,
+ 35
+ ],
+ "category_id": 4,
+ "id": 1573,
+ "area": 945,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 794,
+ "bbox": [
+ 76,
+ 189,
+ 307,
+ 91
+ ],
+ "category_id": 19,
+ "id": 1574,
+ "area": 27937,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 795,
+ "bbox": [
+ 91,
+ 212,
+ 296,
+ 45
+ ],
+ "category_id": 19,
+ "id": 1575,
+ "area": 13320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 796,
+ "bbox": [
+ 249,
+ 146,
+ 119,
+ 264
+ ],
+ "category_id": 16,
+ "id": 1576,
+ "area": 31416,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 797,
+ "bbox": [
+ 96,
+ 71,
+ 18,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1577,
+ "area": 162,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 797,
+ "bbox": [
+ 71,
+ 236,
+ 15,
+ 6
+ ],
+ "category_id": 4,
+ "id": 1578,
+ "area": 90,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 797,
+ "bbox": [
+ 30,
+ 408,
+ 19,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1579,
+ "area": 171,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 797,
+ "bbox": [
+ 259,
+ 411,
+ 16,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1580,
+ "area": 128,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 797,
+ "bbox": [
+ 364,
+ 273,
+ 19,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1581,
+ "area": 209,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 797,
+ "bbox": [
+ 432,
+ 155,
+ 18,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1582,
+ "area": 198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 798,
+ "bbox": [
+ 187,
+ 479,
+ 11,
+ 11
+ ],
+ "category_id": 15,
+ "id": 1583,
+ "area": 121,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 798,
+ "bbox": [
+ 210,
+ 440,
+ 12,
+ 13
+ ],
+ "category_id": 15,
+ "id": 1584,
+ "area": 156,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 798,
+ "bbox": [
+ 252,
+ 424,
+ 20,
+ 22
+ ],
+ "category_id": 15,
+ "id": 1585,
+ "area": 440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 798,
+ "bbox": [
+ 280,
+ 434,
+ 20,
+ 22
+ ],
+ "category_id": 15,
+ "id": 1586,
+ "area": 440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 799,
+ "bbox": [
+ 103,
+ 330,
+ 49,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1587,
+ "area": 1176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 799,
+ "bbox": [
+ 178,
+ 63,
+ 50,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1588,
+ "area": 1250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 799,
+ "bbox": [
+ 418,
+ 418,
+ 53,
+ 31
+ ],
+ "category_id": 4,
+ "id": 1589,
+ "area": 1643,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 800,
+ "bbox": [
+ 144,
+ 151,
+ 236,
+ 122
+ ],
+ "category_id": 16,
+ "id": 1590,
+ "area": 28792,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 801,
+ "bbox": [
+ 266,
+ 330,
+ 158,
+ 89
+ ],
+ "category_id": 10,
+ "id": 1591,
+ "area": 14062,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 802,
+ "bbox": [
+ 44,
+ 220,
+ 416,
+ 118
+ ],
+ "category_id": 16,
+ "id": 1592,
+ "area": 49088,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 803,
+ "bbox": [
+ 133,
+ 93,
+ 347,
+ 324
+ ],
+ "category_id": 10,
+ "id": 1593,
+ "area": 112428,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 804,
+ "bbox": [
+ 223,
+ 88,
+ 67,
+ 82
+ ],
+ "category_id": 15,
+ "id": 1594,
+ "area": 5494,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 804,
+ "bbox": [
+ 43,
+ 90,
+ 47,
+ 77
+ ],
+ "category_id": 15,
+ "id": 1595,
+ "area": 3619,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 804,
+ "bbox": [
+ 30,
+ 195,
+ 44,
+ 71
+ ],
+ "category_id": 15,
+ "id": 1596,
+ "area": 3124,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 804,
+ "bbox": [
+ 205,
+ 198,
+ 67,
+ 80
+ ],
+ "category_id": 15,
+ "id": 1597,
+ "area": 5360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 804,
+ "bbox": [
+ 190,
+ 304,
+ 71,
+ 83
+ ],
+ "category_id": 15,
+ "id": 1598,
+ "area": 5893,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 804,
+ "bbox": [
+ 178,
+ 391,
+ 69,
+ 81
+ ],
+ "category_id": 15,
+ "id": 1599,
+ "area": 5589,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 805,
+ "bbox": [
+ 211,
+ 252,
+ 82,
+ 62
+ ],
+ "category_id": 4,
+ "id": 1600,
+ "area": 5084,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 806,
+ "bbox": [
+ 139,
+ 192,
+ 222,
+ 145
+ ],
+ "category_id": 19,
+ "id": 1601,
+ "area": 32190,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 807,
+ "bbox": [
+ 373,
+ 38,
+ 64,
+ 411
+ ],
+ "category_id": 19,
+ "id": 1602,
+ "area": 26304,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 808,
+ "bbox": [
+ 239,
+ 90,
+ 68,
+ 290
+ ],
+ "category_id": 19,
+ "id": 1603,
+ "area": 19720,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 809,
+ "bbox": [
+ 296,
+ 82,
+ 176,
+ 106
+ ],
+ "category_id": 10,
+ "id": 1604,
+ "area": 18656,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 810,
+ "bbox": [
+ 197,
+ 243,
+ 114,
+ 93
+ ],
+ "category_id": 4,
+ "id": 1605,
+ "area": 10602,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 811,
+ "bbox": [
+ 144,
+ 136,
+ 54,
+ 55
+ ],
+ "category_id": 4,
+ "id": 1606,
+ "area": 2970,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 811,
+ "bbox": [
+ 385,
+ 353,
+ 57,
+ 62
+ ],
+ "category_id": 4,
+ "id": 1607,
+ "area": 3534,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 812,
+ "bbox": [
+ 135,
+ 169,
+ 69,
+ 200
+ ],
+ "category_id": 19,
+ "id": 1608,
+ "area": 13800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 813,
+ "bbox": [
+ 110,
+ 253,
+ 8,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1609,
+ "area": 64,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 813,
+ "bbox": [
+ 176,
+ 239,
+ 7,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1610,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 813,
+ "bbox": [
+ 161,
+ 295,
+ 6,
+ 6
+ ],
+ "category_id": 4,
+ "id": 1611,
+ "area": 36,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 813,
+ "bbox": [
+ 186,
+ 335,
+ 8,
+ 4
+ ],
+ "category_id": 4,
+ "id": 1612,
+ "area": 32,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 813,
+ "bbox": [
+ 117,
+ 307,
+ 3,
+ 5
+ ],
+ "category_id": 4,
+ "id": 1613,
+ "area": 15,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 813,
+ "bbox": [
+ 223,
+ 205,
+ 8,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1614,
+ "area": 64,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 813,
+ "bbox": [
+ 283,
+ 254,
+ 10,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1615,
+ "area": 80,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 813,
+ "bbox": [
+ 311,
+ 222,
+ 8,
+ 7
+ ],
+ "category_id": 4,
+ "id": 1616,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 813,
+ "bbox": [
+ 278,
+ 184,
+ 6,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1617,
+ "area": 54,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 813,
+ "bbox": [
+ 284,
+ 403,
+ 11,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1618,
+ "area": 88,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 813,
+ "bbox": [
+ 275,
+ 475,
+ 10,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1619,
+ "area": 100,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 813,
+ "bbox": [
+ 257,
+ 435,
+ 9,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1620,
+ "area": 72,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 813,
+ "bbox": [
+ 289,
+ 79,
+ 7,
+ 7
+ ],
+ "category_id": 4,
+ "id": 1621,
+ "area": 49,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 814,
+ "bbox": [
+ 78,
+ 352,
+ 56,
+ 45
+ ],
+ "category_id": 4,
+ "id": 1622,
+ "area": 2520,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 814,
+ "bbox": [
+ 399,
+ 69,
+ 64,
+ 39
+ ],
+ "category_id": 4,
+ "id": 1623,
+ "area": 2496,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 815,
+ "bbox": [
+ 298,
+ 220,
+ 51,
+ 36
+ ],
+ "category_id": 15,
+ "id": 1624,
+ "area": 1836,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 815,
+ "bbox": [
+ 204,
+ 173,
+ 44,
+ 38
+ ],
+ "category_id": 15,
+ "id": 1625,
+ "area": 1672,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 815,
+ "bbox": [
+ 236,
+ 224,
+ 43,
+ 37
+ ],
+ "category_id": 15,
+ "id": 1626,
+ "area": 1591,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 815,
+ "bbox": [
+ 258,
+ 261,
+ 40,
+ 36
+ ],
+ "category_id": 15,
+ "id": 1627,
+ "area": 1440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 815,
+ "bbox": [
+ 278,
+ 297,
+ 41,
+ 35
+ ],
+ "category_id": 15,
+ "id": 1628,
+ "area": 1435,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 815,
+ "bbox": [
+ 302,
+ 336,
+ 65,
+ 58
+ ],
+ "category_id": 15,
+ "id": 1629,
+ "area": 3770,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 815,
+ "bbox": [
+ 333,
+ 391,
+ 64,
+ 56
+ ],
+ "category_id": 15,
+ "id": 1630,
+ "area": 3584,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 815,
+ "bbox": [
+ 415,
+ 444,
+ 59,
+ 44
+ ],
+ "category_id": 15,
+ "id": 1631,
+ "area": 2596,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 815,
+ "bbox": [
+ 343,
+ 294,
+ 51,
+ 38
+ ],
+ "category_id": 15,
+ "id": 1632,
+ "area": 1938,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 816,
+ "bbox": [
+ 361,
+ 34,
+ 33,
+ 12
+ ],
+ "category_id": 4,
+ "id": 1633,
+ "area": 396,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 816,
+ "bbox": [
+ 266,
+ 117,
+ 34,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1634,
+ "area": 306,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 816,
+ "bbox": [
+ 85,
+ 195,
+ 33,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1635,
+ "area": 330,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 816,
+ "bbox": [
+ 33,
+ 350,
+ 22,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1636,
+ "area": 506,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 816,
+ "bbox": [
+ 444,
+ 417,
+ 35,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1637,
+ "area": 315,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 817,
+ "bbox": [
+ 80,
+ 115,
+ 26,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1638,
+ "area": 650,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 817,
+ "bbox": [
+ 22,
+ 209,
+ 22,
+ 31
+ ],
+ "category_id": 4,
+ "id": 1639,
+ "area": 682,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 817,
+ "bbox": [
+ 366,
+ 248,
+ 30,
+ 32
+ ],
+ "category_id": 4,
+ "id": 1640,
+ "area": 960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 817,
+ "bbox": [
+ 464,
+ 297,
+ 28,
+ 33
+ ],
+ "category_id": 4,
+ "id": 1641,
+ "area": 924,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 818,
+ "bbox": [
+ 147,
+ 58,
+ 68,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1642,
+ "area": 1700,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 818,
+ "bbox": [
+ 248,
+ 396,
+ 63,
+ 39
+ ],
+ "category_id": 4,
+ "id": 1643,
+ "area": 2457,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 819,
+ "bbox": [
+ 229,
+ 164,
+ 167,
+ 220
+ ],
+ "category_id": 10,
+ "id": 1644,
+ "area": 36740,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 820,
+ "bbox": [
+ 171,
+ 271,
+ 122,
+ 111
+ ],
+ "category_id": 19,
+ "id": 1645,
+ "area": 13542,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 821,
+ "bbox": [
+ 131,
+ 373,
+ 174,
+ 51
+ ],
+ "category_id": 19,
+ "id": 1646,
+ "area": 8874,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 822,
+ "bbox": [
+ 106,
+ 131,
+ 16,
+ 57
+ ],
+ "category_id": 4,
+ "id": 1647,
+ "area": 912,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 822,
+ "bbox": [
+ 138,
+ 407,
+ 19,
+ 53
+ ],
+ "category_id": 4,
+ "id": 1648,
+ "area": 1007,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 822,
+ "bbox": [
+ 382,
+ 85,
+ 28,
+ 55
+ ],
+ "category_id": 4,
+ "id": 1649,
+ "area": 1540,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 823,
+ "bbox": [
+ 115,
+ 240,
+ 253,
+ 61
+ ],
+ "category_id": 19,
+ "id": 1650,
+ "area": 15433,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 824,
+ "bbox": [
+ 85,
+ 59,
+ 59,
+ 31
+ ],
+ "category_id": 4,
+ "id": 1651,
+ "area": 1829,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 824,
+ "bbox": [
+ 154,
+ 451,
+ 61,
+ 37
+ ],
+ "category_id": 4,
+ "id": 1652,
+ "area": 2257,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 824,
+ "bbox": [
+ 343,
+ 289,
+ 56,
+ 36
+ ],
+ "category_id": 4,
+ "id": 1653,
+ "area": 2016,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 825,
+ "bbox": [
+ 227,
+ 184,
+ 42,
+ 131
+ ],
+ "category_id": 19,
+ "id": 1654,
+ "area": 5502,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 826,
+ "bbox": [
+ 300,
+ 419,
+ 189,
+ 36
+ ],
+ "category_id": 10,
+ "id": 1655,
+ "area": 6804,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 827,
+ "bbox": [
+ 188,
+ 137,
+ 123,
+ 147
+ ],
+ "category_id": 15,
+ "id": 1656,
+ "area": 18081,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 827,
+ "bbox": [
+ 131,
+ 286,
+ 119,
+ 146
+ ],
+ "category_id": 15,
+ "id": 1657,
+ "area": 17374,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 828,
+ "bbox": [
+ 232,
+ 112,
+ 25,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1658,
+ "area": 600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 828,
+ "bbox": [
+ 158,
+ 467,
+ 19,
+ 21
+ ],
+ "category_id": 4,
+ "id": 1659,
+ "area": 399,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 828,
+ "bbox": [
+ 360,
+ 457,
+ 25,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1660,
+ "area": 600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 829,
+ "bbox": [
+ 159,
+ 80,
+ 236,
+ 183
+ ],
+ "category_id": 16,
+ "id": 1661,
+ "area": 43188,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 830,
+ "bbox": [
+ 267,
+ 325,
+ 204,
+ 115
+ ],
+ "category_id": 10,
+ "id": 1662,
+ "area": 23460,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 831,
+ "bbox": [
+ 135,
+ 201,
+ 105,
+ 94
+ ],
+ "category_id": 15,
+ "id": 1663,
+ "area": 9870,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 831,
+ "bbox": [
+ 55,
+ 316,
+ 110,
+ 103
+ ],
+ "category_id": 15,
+ "id": 1664,
+ "area": 11330,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 832,
+ "bbox": [
+ 26,
+ 200,
+ 19,
+ 68
+ ],
+ "category_id": 4,
+ "id": 1665,
+ "area": 1292,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 832,
+ "bbox": [
+ 481,
+ 192,
+ 18,
+ 80
+ ],
+ "category_id": 4,
+ "id": 1666,
+ "area": 1440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 833,
+ "bbox": [
+ 239,
+ 109,
+ 131,
+ 339
+ ],
+ "category_id": 10,
+ "id": 1667,
+ "area": 44409,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 834,
+ "bbox": [
+ 250,
+ 120,
+ 120,
+ 161
+ ],
+ "category_id": 15,
+ "id": 1668,
+ "area": 19320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 835,
+ "bbox": [
+ 32,
+ 190,
+ 457,
+ 49
+ ],
+ "category_id": 10,
+ "id": 1669,
+ "area": 22393,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 836,
+ "bbox": [
+ 82,
+ 373,
+ 18,
+ 60
+ ],
+ "category_id": 10,
+ "id": 1670,
+ "area": 1080,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 837,
+ "bbox": [
+ 241,
+ 259,
+ 73,
+ 60
+ ],
+ "category_id": 4,
+ "id": 1671,
+ "area": 4380,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 838,
+ "bbox": [
+ 104,
+ 111,
+ 12,
+ 18
+ ],
+ "category_id": 4,
+ "id": 1672,
+ "area": 216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 838,
+ "bbox": [
+ 364,
+ 84,
+ 21,
+ 13
+ ],
+ "category_id": 4,
+ "id": 1673,
+ "area": 273,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 838,
+ "bbox": [
+ 37,
+ 387,
+ 12,
+ 21
+ ],
+ "category_id": 4,
+ "id": 1674,
+ "area": 252,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 838,
+ "bbox": [
+ 467,
+ 297,
+ 14,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1675,
+ "area": 350,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 839,
+ "bbox": [
+ 149,
+ 97,
+ 274,
+ 338
+ ],
+ "category_id": 10,
+ "id": 1676,
+ "area": 92612,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 840,
+ "bbox": [
+ 168,
+ 90,
+ 201,
+ 347
+ ],
+ "category_id": 10,
+ "id": 1677,
+ "area": 69747,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 841,
+ "bbox": [
+ 52,
+ 284,
+ 404,
+ 39
+ ],
+ "category_id": 10,
+ "id": 1678,
+ "area": 15756,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 842,
+ "bbox": [
+ 53,
+ 134,
+ 363,
+ 259
+ ],
+ "category_id": 10,
+ "id": 1679,
+ "area": 94017,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 843,
+ "bbox": [
+ 143,
+ 177,
+ 299,
+ 214
+ ],
+ "category_id": 10,
+ "id": 1680,
+ "area": 63986,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 844,
+ "bbox": [
+ 140,
+ 42,
+ 18,
+ 41
+ ],
+ "category_id": 4,
+ "id": 1681,
+ "area": 738,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 844,
+ "bbox": [
+ 109,
+ 240,
+ 24,
+ 35
+ ],
+ "category_id": 4,
+ "id": 1682,
+ "area": 840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 844,
+ "bbox": [
+ 58,
+ 413,
+ 20,
+ 36
+ ],
+ "category_id": 4,
+ "id": 1683,
+ "area": 720,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 844,
+ "bbox": [
+ 245,
+ 437,
+ 23,
+ 30
+ ],
+ "category_id": 4,
+ "id": 1684,
+ "area": 690,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 844,
+ "bbox": [
+ 364,
+ 247,
+ 21,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1685,
+ "area": 525,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 844,
+ "bbox": [
+ 357,
+ 72,
+ 37,
+ 31
+ ],
+ "category_id": 4,
+ "id": 1686,
+ "area": 1147,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 845,
+ "bbox": [
+ 30,
+ 254,
+ 7,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1687,
+ "area": 70,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 845,
+ "bbox": [
+ 100,
+ 257,
+ 7,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1688,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 845,
+ "bbox": [
+ 149,
+ 261,
+ 9,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1689,
+ "area": 90,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 845,
+ "bbox": [
+ 112,
+ 326,
+ 7,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1690,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 845,
+ "bbox": [
+ 80,
+ 371,
+ 8,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1691,
+ "area": 88,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 845,
+ "bbox": [
+ 218,
+ 380,
+ 4,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1692,
+ "area": 36,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 845,
+ "bbox": [
+ 296,
+ 177,
+ 7,
+ 13
+ ],
+ "category_id": 4,
+ "id": 1693,
+ "area": 91,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 845,
+ "bbox": [
+ 332,
+ 154,
+ 7,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1694,
+ "area": 63,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 845,
+ "bbox": [
+ 424,
+ 168,
+ 8,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1695,
+ "area": 72,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 845,
+ "bbox": [
+ 464,
+ 122,
+ 10,
+ 14
+ ],
+ "category_id": 4,
+ "id": 1696,
+ "area": 140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 845,
+ "bbox": [
+ 464,
+ 31,
+ 7,
+ 13
+ ],
+ "category_id": 4,
+ "id": 1697,
+ "area": 91,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 845,
+ "bbox": [
+ 475,
+ 65,
+ 7,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1698,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 846,
+ "bbox": [
+ 250,
+ 237,
+ 79,
+ 109
+ ],
+ "category_id": 15,
+ "id": 1699,
+ "area": 8611,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 847,
+ "bbox": [
+ 345,
+ 192,
+ 96,
+ 118
+ ],
+ "category_id": 15,
+ "id": 1700,
+ "area": 11328,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 847,
+ "bbox": [
+ 403,
+ 306,
+ 98,
+ 117
+ ],
+ "category_id": 15,
+ "id": 1701,
+ "area": 11466,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 848,
+ "bbox": [
+ 110,
+ 254,
+ 56,
+ 206
+ ],
+ "category_id": 10,
+ "id": 1702,
+ "area": 11536,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 849,
+ "bbox": [
+ 54,
+ 302,
+ 405,
+ 158
+ ],
+ "category_id": 10,
+ "id": 1703,
+ "area": 63990,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 850,
+ "bbox": [
+ 85,
+ 110,
+ 327,
+ 300
+ ],
+ "category_id": 10,
+ "id": 1704,
+ "area": 98100,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 851,
+ "bbox": [
+ 230,
+ 201,
+ 69,
+ 190
+ ],
+ "category_id": 16,
+ "id": 1705,
+ "area": 13110,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 851,
+ "bbox": [
+ 250,
+ 66,
+ 37,
+ 138
+ ],
+ "category_id": 16,
+ "id": 1706,
+ "area": 5106,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 852,
+ "bbox": [
+ 62,
+ 85,
+ 261,
+ 213
+ ],
+ "category_id": 19,
+ "id": 1707,
+ "area": 55593,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 853,
+ "bbox": [
+ 202,
+ 209,
+ 44,
+ 91
+ ],
+ "category_id": 19,
+ "id": 1708,
+ "area": 4004,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 854,
+ "bbox": [
+ 211,
+ 216,
+ 110,
+ 70
+ ],
+ "category_id": 4,
+ "id": 1709,
+ "area": 7700,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 855,
+ "bbox": [
+ 410,
+ 76,
+ 58,
+ 263
+ ],
+ "category_id": 19,
+ "id": 1710,
+ "area": 15254,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 856,
+ "bbox": [
+ 87,
+ 325,
+ 33,
+ 35
+ ],
+ "category_id": 4,
+ "id": 1711,
+ "area": 1155,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 856,
+ "bbox": [
+ 454,
+ 220,
+ 34,
+ 35
+ ],
+ "category_id": 4,
+ "id": 1712,
+ "area": 1190,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 857,
+ "bbox": [
+ 199,
+ 141,
+ 138,
+ 269
+ ],
+ "category_id": 16,
+ "id": 1713,
+ "area": 37122,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 858,
+ "bbox": [
+ 218,
+ 85,
+ 46,
+ 296
+ ],
+ "category_id": 19,
+ "id": 1714,
+ "area": 13616,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 859,
+ "bbox": [
+ 273,
+ 49,
+ 189,
+ 100
+ ],
+ "category_id": 10,
+ "id": 1715,
+ "area": 18900,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 860,
+ "bbox": [
+ 30,
+ 391,
+ 220,
+ 86
+ ],
+ "category_id": 10,
+ "id": 1716,
+ "area": 18920,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 861,
+ "bbox": [
+ 88,
+ 389,
+ 24,
+ 46
+ ],
+ "category_id": 4,
+ "id": 1717,
+ "area": 1104,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 861,
+ "bbox": [
+ 210,
+ 236,
+ 22,
+ 43
+ ],
+ "category_id": 4,
+ "id": 1718,
+ "area": 946,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 861,
+ "bbox": [
+ 432,
+ 122,
+ 22,
+ 48
+ ],
+ "category_id": 4,
+ "id": 1719,
+ "area": 1056,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 862,
+ "bbox": [
+ 33,
+ 48,
+ 119,
+ 83
+ ],
+ "category_id": 10,
+ "id": 1720,
+ "area": 9877,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 863,
+ "bbox": [
+ 190,
+ 237,
+ 201,
+ 122
+ ],
+ "category_id": 16,
+ "id": 1721,
+ "area": 24522,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 864,
+ "bbox": [
+ 270,
+ 291,
+ 23,
+ 21
+ ],
+ "category_id": 16,
+ "id": 1722,
+ "area": 483,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 864,
+ "bbox": [
+ 252,
+ 385,
+ 23,
+ 37
+ ],
+ "category_id": 16,
+ "id": 1723,
+ "area": 851,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 865,
+ "bbox": [
+ 193,
+ 428,
+ 16,
+ 17
+ ],
+ "category_id": 15,
+ "id": 1724,
+ "area": 272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 865,
+ "bbox": [
+ 208,
+ 441,
+ 19,
+ 21
+ ],
+ "category_id": 15,
+ "id": 1725,
+ "area": 399,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 865,
+ "bbox": [
+ 476,
+ 322,
+ 15,
+ 24
+ ],
+ "category_id": 15,
+ "id": 1726,
+ "area": 360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 866,
+ "bbox": [
+ 154,
+ 216,
+ 221,
+ 238
+ ],
+ "category_id": 16,
+ "id": 1727,
+ "area": 52598,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 867,
+ "bbox": [
+ 96,
+ 96,
+ 185,
+ 303
+ ],
+ "category_id": 10,
+ "id": 1728,
+ "area": 56055,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 868,
+ "bbox": [
+ 206,
+ 82,
+ 100,
+ 336
+ ],
+ "category_id": 16,
+ "id": 1729,
+ "area": 33600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 869,
+ "bbox": [
+ 37,
+ 234,
+ 198,
+ 158
+ ],
+ "category_id": 15,
+ "id": 1730,
+ "area": 31284,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 869,
+ "bbox": [
+ 261,
+ 230,
+ 190,
+ 155
+ ],
+ "category_id": 15,
+ "id": 1731,
+ "area": 29450,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 870,
+ "bbox": [
+ 147,
+ 117,
+ 274,
+ 300
+ ],
+ "category_id": 10,
+ "id": 1732,
+ "area": 82200,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 871,
+ "bbox": [
+ 77,
+ 64,
+ 347,
+ 307
+ ],
+ "category_id": 16,
+ "id": 1733,
+ "area": 106529,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 872,
+ "bbox": [
+ 373,
+ 59,
+ 6,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1734,
+ "area": 54,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 872,
+ "bbox": [
+ 332,
+ 84,
+ 4,
+ 7
+ ],
+ "category_id": 4,
+ "id": 1735,
+ "area": 28,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 872,
+ "bbox": [
+ 289,
+ 131,
+ 7,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1736,
+ "area": 70,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 872,
+ "bbox": [
+ 272,
+ 192,
+ 7,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1737,
+ "area": 70,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 872,
+ "bbox": [
+ 203,
+ 240,
+ 5,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1738,
+ "area": 40,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 872,
+ "bbox": [
+ 222,
+ 364,
+ 7,
+ 12
+ ],
+ "category_id": 4,
+ "id": 1739,
+ "area": 84,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 872,
+ "bbox": [
+ 144,
+ 364,
+ 5,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1740,
+ "area": 55,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 872,
+ "bbox": [
+ 222,
+ 418,
+ 5,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1741,
+ "area": 40,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 872,
+ "bbox": [
+ 202,
+ 458,
+ 6,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1742,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 873,
+ "bbox": [
+ 89,
+ 108,
+ 121,
+ 113
+ ],
+ "category_id": 15,
+ "id": 1743,
+ "area": 13673,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 873,
+ "bbox": [
+ 206,
+ 188,
+ 119,
+ 112
+ ],
+ "category_id": 15,
+ "id": 1744,
+ "area": 13328,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 874,
+ "bbox": [
+ 274,
+ 152,
+ 84,
+ 160
+ ],
+ "category_id": 15,
+ "id": 1745,
+ "area": 13440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 875,
+ "bbox": [
+ 246,
+ 295,
+ 94,
+ 154
+ ],
+ "category_id": 16,
+ "id": 1746,
+ "area": 14476,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 876,
+ "bbox": [
+ 268,
+ 112,
+ 80,
+ 323
+ ],
+ "category_id": 10,
+ "id": 1747,
+ "area": 25840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 877,
+ "bbox": [
+ 185,
+ 97,
+ 121,
+ 100
+ ],
+ "category_id": 15,
+ "id": 1748,
+ "area": 12100,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 877,
+ "bbox": [
+ 75,
+ 185,
+ 130,
+ 108
+ ],
+ "category_id": 15,
+ "id": 1749,
+ "area": 14040,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 877,
+ "bbox": [
+ 152,
+ 301,
+ 148,
+ 117
+ ],
+ "category_id": 15,
+ "id": 1750,
+ "area": 17316,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 877,
+ "bbox": [
+ 267,
+ 206,
+ 149,
+ 117
+ ],
+ "category_id": 15,
+ "id": 1751,
+ "area": 17433,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 878,
+ "bbox": [
+ 197,
+ 90,
+ 178,
+ 274
+ ],
+ "category_id": 19,
+ "id": 1752,
+ "area": 48772,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 879,
+ "bbox": [
+ 95,
+ 200,
+ 325,
+ 209
+ ],
+ "category_id": 10,
+ "id": 1753,
+ "area": 67925,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 880,
+ "bbox": [
+ 135,
+ 209,
+ 211,
+ 172
+ ],
+ "category_id": 16,
+ "id": 1754,
+ "area": 36292,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 881,
+ "bbox": [
+ 391,
+ 37,
+ 24,
+ 28
+ ],
+ "category_id": 4,
+ "id": 1755,
+ "area": 672,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 881,
+ "bbox": [
+ 432,
+ 127,
+ 21,
+ 27
+ ],
+ "category_id": 4,
+ "id": 1756,
+ "area": 567,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 881,
+ "bbox": [
+ 149,
+ 62,
+ 26,
+ 22
+ ],
+ "category_id": 4,
+ "id": 1757,
+ "area": 572,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 881,
+ "bbox": [
+ 160,
+ 142,
+ 19,
+ 21
+ ],
+ "category_id": 4,
+ "id": 1758,
+ "area": 399,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 881,
+ "bbox": [
+ 176,
+ 221,
+ 24,
+ 27
+ ],
+ "category_id": 4,
+ "id": 1759,
+ "area": 648,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 881,
+ "bbox": [
+ 188,
+ 307,
+ 24,
+ 22
+ ],
+ "category_id": 4,
+ "id": 1760,
+ "area": 528,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 881,
+ "bbox": [
+ 346,
+ 279,
+ 25,
+ 36
+ ],
+ "category_id": 4,
+ "id": 1761,
+ "area": 900,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 881,
+ "bbox": [
+ 154,
+ 421,
+ 24,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1762,
+ "area": 576,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 881,
+ "bbox": [
+ 64,
+ 476,
+ 22,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1763,
+ "area": 506,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 882,
+ "bbox": [
+ 113,
+ 172,
+ 66,
+ 50
+ ],
+ "category_id": 4,
+ "id": 1764,
+ "area": 3300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 882,
+ "bbox": [
+ 397,
+ 305,
+ 45,
+ 54
+ ],
+ "category_id": 4,
+ "id": 1765,
+ "area": 2430,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 883,
+ "bbox": [
+ 40,
+ 76,
+ 361,
+ 252
+ ],
+ "category_id": 10,
+ "id": 1766,
+ "area": 90972,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 884,
+ "bbox": [
+ 141,
+ 28,
+ 271,
+ 345
+ ],
+ "category_id": 10,
+ "id": 1767,
+ "area": 93495,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 885,
+ "bbox": [
+ 300,
+ 145,
+ 44,
+ 80
+ ],
+ "category_id": 4,
+ "id": 1768,
+ "area": 3520,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 885,
+ "bbox": [
+ 75,
+ 376,
+ 57,
+ 66
+ ],
+ "category_id": 4,
+ "id": 1769,
+ "area": 3762,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 886,
+ "bbox": [
+ 70,
+ 178,
+ 173,
+ 211
+ ],
+ "category_id": 15,
+ "id": 1770,
+ "area": 36503,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 886,
+ "bbox": [
+ 253,
+ 182,
+ 166,
+ 202
+ ],
+ "category_id": 15,
+ "id": 1771,
+ "area": 33532,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 887,
+ "bbox": [
+ 234,
+ 136,
+ 39,
+ 259
+ ],
+ "category_id": 19,
+ "id": 1772,
+ "area": 10101,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 888,
+ "bbox": [
+ 73,
+ 101,
+ 170,
+ 165
+ ],
+ "category_id": 15,
+ "id": 1773,
+ "area": 28050,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 888,
+ "bbox": [
+ 63,
+ 312,
+ 173,
+ 164
+ ],
+ "category_id": 15,
+ "id": 1774,
+ "area": 28372,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 889,
+ "bbox": [
+ 344,
+ 100,
+ 38,
+ 47
+ ],
+ "category_id": 4,
+ "id": 1775,
+ "area": 1786,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 889,
+ "bbox": [
+ 142,
+ 381,
+ 47,
+ 37
+ ],
+ "category_id": 4,
+ "id": 1776,
+ "area": 1739,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 890,
+ "bbox": [
+ 86,
+ 42,
+ 17,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1777,
+ "area": 425,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 890,
+ "bbox": [
+ 100,
+ 156,
+ 14,
+ 30
+ ],
+ "category_id": 4,
+ "id": 1778,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 890,
+ "bbox": [
+ 96,
+ 279,
+ 18,
+ 27
+ ],
+ "category_id": 4,
+ "id": 1779,
+ "area": 486,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 890,
+ "bbox": [
+ 97,
+ 401,
+ 15,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1780,
+ "area": 375,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 890,
+ "bbox": [
+ 442,
+ 97,
+ 18,
+ 31
+ ],
+ "category_id": 4,
+ "id": 1781,
+ "area": 558,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 890,
+ "bbox": [
+ 440,
+ 218,
+ 17,
+ 30
+ ],
+ "category_id": 4,
+ "id": 1782,
+ "area": 510,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 890,
+ "bbox": [
+ 435,
+ 344,
+ 20,
+ 29
+ ],
+ "category_id": 4,
+ "id": 1783,
+ "area": 580,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 890,
+ "bbox": [
+ 438,
+ 465,
+ 18,
+ 31
+ ],
+ "category_id": 4,
+ "id": 1784,
+ "area": 558,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 891,
+ "bbox": [
+ 223,
+ 58,
+ 68,
+ 89
+ ],
+ "category_id": 15,
+ "id": 1785,
+ "area": 6052,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 891,
+ "bbox": [
+ 218,
+ 154,
+ 69,
+ 93
+ ],
+ "category_id": 15,
+ "id": 1786,
+ "area": 6417,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 891,
+ "bbox": [
+ 214,
+ 261,
+ 63,
+ 87
+ ],
+ "category_id": 15,
+ "id": 1787,
+ "area": 5481,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 891,
+ "bbox": [
+ 207,
+ 348,
+ 62,
+ 85
+ ],
+ "category_id": 15,
+ "id": 1788,
+ "area": 5270,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 891,
+ "bbox": [
+ 472,
+ 127,
+ 18,
+ 97
+ ],
+ "category_id": 15,
+ "id": 1789,
+ "area": 1746,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 891,
+ "bbox": [
+ 464,
+ 243,
+ 16,
+ 98
+ ],
+ "category_id": 15,
+ "id": 1790,
+ "area": 1568,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 892,
+ "bbox": [
+ 300,
+ 69,
+ 133,
+ 275
+ ],
+ "category_id": 10,
+ "id": 1791,
+ "area": 36575,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 893,
+ "bbox": [
+ 69,
+ 240,
+ 131,
+ 209
+ ],
+ "category_id": 19,
+ "id": 1792,
+ "area": 27379,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 894,
+ "bbox": [
+ 185,
+ 286,
+ 208,
+ 137
+ ],
+ "category_id": 16,
+ "id": 1793,
+ "area": 28496,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 895,
+ "bbox": [
+ 288,
+ 107,
+ 55,
+ 300
+ ],
+ "category_id": 10,
+ "id": 1794,
+ "area": 16500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 896,
+ "bbox": [
+ 128,
+ 54,
+ 156,
+ 163
+ ],
+ "category_id": 15,
+ "id": 1795,
+ "area": 25428,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 896,
+ "bbox": [
+ 113,
+ 288,
+ 159,
+ 160
+ ],
+ "category_id": 15,
+ "id": 1796,
+ "area": 25440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 897,
+ "bbox": [
+ 69,
+ 140,
+ 408,
+ 262
+ ],
+ "category_id": 10,
+ "id": 1797,
+ "area": 106896,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 898,
+ "bbox": [
+ 449,
+ 90,
+ 20,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1798,
+ "area": 500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 898,
+ "bbox": [
+ 306,
+ 74,
+ 10,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1799,
+ "area": 250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 898,
+ "bbox": [
+ 207,
+ 154,
+ 19,
+ 26
+ ],
+ "category_id": 4,
+ "id": 1800,
+ "area": 494,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 898,
+ "bbox": [
+ 165,
+ 318,
+ 15,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1801,
+ "area": 360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 898,
+ "bbox": [
+ 62,
+ 428,
+ 19,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1802,
+ "area": 475,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 899,
+ "bbox": [
+ 143,
+ 121,
+ 171,
+ 279
+ ],
+ "category_id": 16,
+ "id": 1803,
+ "area": 47709,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 900,
+ "bbox": [
+ 210,
+ 27,
+ 15,
+ 13
+ ],
+ "category_id": 4,
+ "id": 1804,
+ "area": 195,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 900,
+ "bbox": [
+ 300,
+ 63,
+ 16,
+ 14
+ ],
+ "category_id": 4,
+ "id": 1805,
+ "area": 224,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 900,
+ "bbox": [
+ 410,
+ 144,
+ 18,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1806,
+ "area": 198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 900,
+ "bbox": [
+ 257,
+ 129,
+ 16,
+ 12
+ ],
+ "category_id": 4,
+ "id": 1807,
+ "area": 192,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 900,
+ "bbox": [
+ 438,
+ 236,
+ 33,
+ 18
+ ],
+ "category_id": 4,
+ "id": 1808,
+ "area": 594,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 900,
+ "bbox": [
+ 288,
+ 284,
+ 24,
+ 12
+ ],
+ "category_id": 4,
+ "id": 1809,
+ "area": 288,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 900,
+ "bbox": [
+ 209,
+ 428,
+ 35,
+ 12
+ ],
+ "category_id": 4,
+ "id": 1810,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 900,
+ "bbox": [
+ 469,
+ 466,
+ 21,
+ 18
+ ],
+ "category_id": 4,
+ "id": 1811,
+ "area": 378,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 901,
+ "bbox": [
+ 47,
+ 218,
+ 33,
+ 62
+ ],
+ "category_id": 4,
+ "id": 1812,
+ "area": 2046,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 901,
+ "bbox": [
+ 436,
+ 204,
+ 36,
+ 65
+ ],
+ "category_id": 4,
+ "id": 1813,
+ "area": 2340,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 902,
+ "bbox": [
+ 184,
+ 94,
+ 56,
+ 350
+ ],
+ "category_id": 19,
+ "id": 1814,
+ "area": 19600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 903,
+ "bbox": [
+ 75,
+ 333,
+ 66,
+ 31
+ ],
+ "category_id": 4,
+ "id": 1815,
+ "area": 2046,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 903,
+ "bbox": [
+ 420,
+ 245,
+ 76,
+ 30
+ ],
+ "category_id": 4,
+ "id": 1816,
+ "area": 2280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 904,
+ "bbox": [
+ 45,
+ 289,
+ 409,
+ 30
+ ],
+ "category_id": 10,
+ "id": 1817,
+ "area": 12270,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 905,
+ "bbox": [
+ 88,
+ 333,
+ 139,
+ 153
+ ],
+ "category_id": 10,
+ "id": 1818,
+ "area": 21267,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 906,
+ "bbox": [
+ 104,
+ 136,
+ 18,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1819,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 906,
+ "bbox": [
+ 327,
+ 114,
+ 14,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1820,
+ "area": 154,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 906,
+ "bbox": [
+ 187,
+ 263,
+ 15,
+ 12
+ ],
+ "category_id": 4,
+ "id": 1821,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 906,
+ "bbox": [
+ 395,
+ 378,
+ 18,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1822,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 906,
+ "bbox": [
+ 69,
+ 423,
+ 16,
+ 12
+ ],
+ "category_id": 4,
+ "id": 1823,
+ "area": 192,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 907,
+ "bbox": [
+ 115,
+ 180,
+ 135,
+ 185
+ ],
+ "category_id": 16,
+ "id": 1824,
+ "area": 24975,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 908,
+ "bbox": [
+ 55,
+ 224,
+ 424,
+ 37
+ ],
+ "category_id": 10,
+ "id": 1825,
+ "area": 15688,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 909,
+ "bbox": [
+ 37,
+ 31,
+ 27,
+ 436
+ ],
+ "category_id": 10,
+ "id": 1826,
+ "area": 11772,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 910,
+ "bbox": [
+ 243,
+ 159,
+ 121,
+ 276
+ ],
+ "category_id": 16,
+ "id": 1827,
+ "area": 33396,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 911,
+ "bbox": [
+ 73,
+ 229,
+ 9,
+ 34
+ ],
+ "category_id": 15,
+ "id": 1828,
+ "area": 306,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 911,
+ "bbox": [
+ 116,
+ 211,
+ 40,
+ 47
+ ],
+ "category_id": 15,
+ "id": 1829,
+ "area": 1880,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 911,
+ "bbox": [
+ 154,
+ 165,
+ 44,
+ 49
+ ],
+ "category_id": 15,
+ "id": 1830,
+ "area": 2156,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 912,
+ "bbox": [
+ 24,
+ 22,
+ 22,
+ 15
+ ],
+ "category_id": 4,
+ "id": 1831,
+ "area": 330,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 912,
+ "bbox": [
+ 296,
+ 32,
+ 20,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1832,
+ "area": 220,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 912,
+ "bbox": [
+ 82,
+ 195,
+ 22,
+ 16
+ ],
+ "category_id": 4,
+ "id": 1833,
+ "area": 352,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 912,
+ "bbox": [
+ 62,
+ 423,
+ 21,
+ 16
+ ],
+ "category_id": 4,
+ "id": 1834,
+ "area": 336,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 912,
+ "bbox": [
+ 487,
+ 413,
+ 19,
+ 12
+ ],
+ "category_id": 4,
+ "id": 1835,
+ "area": 228,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 912,
+ "bbox": [
+ 332,
+ 478,
+ 20,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1836,
+ "area": 220,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 913,
+ "bbox": [
+ 190,
+ 119,
+ 79,
+ 247
+ ],
+ "category_id": 19,
+ "id": 1837,
+ "area": 19513,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 914,
+ "bbox": [
+ 158,
+ 180,
+ 244,
+ 136
+ ],
+ "category_id": 16,
+ "id": 1838,
+ "area": 33184,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 915,
+ "bbox": [
+ 262,
+ 141,
+ 109,
+ 179
+ ],
+ "category_id": 15,
+ "id": 1839,
+ "area": 19511,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 915,
+ "bbox": [
+ 293,
+ 303,
+ 111,
+ 184
+ ],
+ "category_id": 15,
+ "id": 1840,
+ "area": 20424,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 916,
+ "bbox": [
+ 298,
+ 33,
+ 193,
+ 260
+ ],
+ "category_id": 10,
+ "id": 1841,
+ "area": 50180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 917,
+ "bbox": [
+ 263,
+ 193,
+ 37,
+ 208
+ ],
+ "category_id": 19,
+ "id": 1842,
+ "area": 7696,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 918,
+ "bbox": [
+ 128,
+ 263,
+ 303,
+ 35
+ ],
+ "category_id": 10,
+ "id": 1843,
+ "area": 10605,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 919,
+ "bbox": [
+ 467,
+ 113,
+ 16,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1844,
+ "area": 144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 919,
+ "bbox": [
+ 202,
+ 315,
+ 16,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1845,
+ "area": 144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 919,
+ "bbox": [
+ 481,
+ 420,
+ 16,
+ 8
+ ],
+ "category_id": 4,
+ "id": 1846,
+ "area": 128,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 919,
+ "bbox": [
+ 53,
+ 467,
+ 17,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1847,
+ "area": 187,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 920,
+ "bbox": [
+ 394,
+ 59,
+ 30,
+ 18
+ ],
+ "category_id": 4,
+ "id": 1848,
+ "area": 540,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 920,
+ "bbox": [
+ 369,
+ 167,
+ 29,
+ 22
+ ],
+ "category_id": 4,
+ "id": 1849,
+ "area": 638,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 920,
+ "bbox": [
+ 307,
+ 269,
+ 25,
+ 21
+ ],
+ "category_id": 4,
+ "id": 1850,
+ "area": 525,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 920,
+ "bbox": [
+ 484,
+ 428,
+ 23,
+ 18
+ ],
+ "category_id": 4,
+ "id": 1851,
+ "area": 414,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 920,
+ "bbox": [
+ 339,
+ 460,
+ 30,
+ 14
+ ],
+ "category_id": 4,
+ "id": 1852,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 920,
+ "bbox": [
+ 248,
+ 350,
+ 24,
+ 21
+ ],
+ "category_id": 4,
+ "id": 1853,
+ "area": 504,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 920,
+ "bbox": [
+ 97,
+ 353,
+ 25,
+ 16
+ ],
+ "category_id": 4,
+ "id": 1854,
+ "area": 400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 920,
+ "bbox": [
+ 12,
+ 392,
+ 30,
+ 14
+ ],
+ "category_id": 4,
+ "id": 1855,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 921,
+ "bbox": [
+ 106,
+ 160,
+ 12,
+ 18
+ ],
+ "category_id": 4,
+ "id": 1856,
+ "area": 216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 921,
+ "bbox": [
+ 56,
+ 310,
+ 22,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1857,
+ "area": 528,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 921,
+ "bbox": [
+ 18,
+ 471,
+ 15,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1858,
+ "area": 345,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 921,
+ "bbox": [
+ 138,
+ 444,
+ 18,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1859,
+ "area": 450,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 921,
+ "bbox": [
+ 289,
+ 279,
+ 20,
+ 21
+ ],
+ "category_id": 4,
+ "id": 1860,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 921,
+ "bbox": [
+ 284,
+ 440,
+ 15,
+ 17
+ ],
+ "category_id": 4,
+ "id": 1861,
+ "area": 255,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 921,
+ "bbox": [
+ 484,
+ 69,
+ 22,
+ 26
+ ],
+ "category_id": 4,
+ "id": 1862,
+ "area": 572,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 921,
+ "bbox": [
+ 395,
+ 211,
+ 22,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1863,
+ "area": 550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 922,
+ "bbox": [
+ 152,
+ 198,
+ 177,
+ 127
+ ],
+ "category_id": 16,
+ "id": 1864,
+ "area": 22479,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 923,
+ "bbox": [
+ 498,
+ 99,
+ 14,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1865,
+ "area": 322,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 923,
+ "bbox": [
+ 304,
+ 215,
+ 20,
+ 30
+ ],
+ "category_id": 4,
+ "id": 1866,
+ "area": 600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 923,
+ "bbox": [
+ 176,
+ 161,
+ 19,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1867,
+ "area": 437,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 923,
+ "bbox": [
+ 43,
+ 339,
+ 22,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1868,
+ "area": 550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 923,
+ "bbox": [
+ 421,
+ 492,
+ 27,
+ 20
+ ],
+ "category_id": 4,
+ "id": 1869,
+ "area": 540,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 924,
+ "bbox": [
+ 72,
+ 196,
+ 376,
+ 236
+ ],
+ "category_id": 10,
+ "id": 1870,
+ "area": 88736,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 925,
+ "bbox": [
+ 128,
+ 100,
+ 281,
+ 292
+ ],
+ "category_id": 10,
+ "id": 1871,
+ "area": 82052,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 926,
+ "bbox": [
+ 246,
+ 246,
+ 152,
+ 207
+ ],
+ "category_id": 16,
+ "id": 1872,
+ "area": 31464,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 927,
+ "bbox": [
+ 209,
+ 273,
+ 167,
+ 43
+ ],
+ "category_id": 16,
+ "id": 1873,
+ "area": 7181,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 928,
+ "bbox": [
+ 425,
+ 152,
+ 77,
+ 162
+ ],
+ "category_id": 19,
+ "id": 1874,
+ "area": 12474,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 929,
+ "bbox": [
+ 73,
+ 204,
+ 411,
+ 135
+ ],
+ "category_id": 10,
+ "id": 1875,
+ "area": 55485,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 930,
+ "bbox": [
+ 103,
+ 402,
+ 132,
+ 74
+ ],
+ "category_id": 10,
+ "id": 1876,
+ "area": 9768,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 931,
+ "bbox": [
+ 34,
+ 108,
+ 411,
+ 348
+ ],
+ "category_id": 15,
+ "id": 1877,
+ "area": 143028,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 932,
+ "bbox": [
+ 132,
+ 106,
+ 324,
+ 247
+ ],
+ "category_id": 10,
+ "id": 1878,
+ "area": 80028,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 933,
+ "bbox": [
+ 75,
+ 253,
+ 70,
+ 54
+ ],
+ "category_id": 16,
+ "id": 1879,
+ "area": 3780,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 934,
+ "bbox": [
+ 133,
+ 128,
+ 259,
+ 135
+ ],
+ "category_id": 16,
+ "id": 1880,
+ "area": 34965,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 935,
+ "bbox": [
+ 71,
+ 340,
+ 30,
+ 83
+ ],
+ "category_id": 4,
+ "id": 1881,
+ "area": 2490,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 935,
+ "bbox": [
+ 397,
+ 200,
+ 28,
+ 72
+ ],
+ "category_id": 4,
+ "id": 1882,
+ "area": 2016,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 935,
+ "bbox": [
+ 336,
+ 422,
+ 29,
+ 66
+ ],
+ "category_id": 4,
+ "id": 1883,
+ "area": 1914,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 936,
+ "bbox": [
+ 95,
+ 38,
+ 89,
+ 182
+ ],
+ "category_id": 19,
+ "id": 1884,
+ "area": 16198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 936,
+ "bbox": [
+ 229,
+ 361,
+ 65,
+ 94
+ ],
+ "category_id": 19,
+ "id": 1885,
+ "area": 6110,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 937,
+ "bbox": [
+ 396,
+ 378,
+ 44,
+ 39
+ ],
+ "category_id": 16,
+ "id": 1886,
+ "area": 1716,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 937,
+ "bbox": [
+ 177,
+ 120,
+ 19,
+ 20
+ ],
+ "category_id": 16,
+ "id": 1887,
+ "area": 380,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 938,
+ "bbox": [
+ 157,
+ 125,
+ 92,
+ 115
+ ],
+ "category_id": 15,
+ "id": 1888,
+ "area": 10580,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 938,
+ "bbox": [
+ 288,
+ 126,
+ 88,
+ 113
+ ],
+ "category_id": 15,
+ "id": 1889,
+ "area": 9944,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 938,
+ "bbox": [
+ 156,
+ 256,
+ 87,
+ 102
+ ],
+ "category_id": 15,
+ "id": 1890,
+ "area": 8874,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 938,
+ "bbox": [
+ 284,
+ 256,
+ 90,
+ 101
+ ],
+ "category_id": 15,
+ "id": 1891,
+ "area": 9090,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 939,
+ "bbox": [
+ 215,
+ 195,
+ 130,
+ 277
+ ],
+ "category_id": 16,
+ "id": 1892,
+ "area": 36010,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 940,
+ "bbox": [
+ 61,
+ 225,
+ 351,
+ 171
+ ],
+ "category_id": 16,
+ "id": 1893,
+ "area": 60021,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 941,
+ "bbox": [
+ 88,
+ 69,
+ 31,
+ 18
+ ],
+ "category_id": 4,
+ "id": 1894,
+ "area": 558,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 941,
+ "bbox": [
+ 364,
+ 55,
+ 22,
+ 14
+ ],
+ "category_id": 4,
+ "id": 1895,
+ "area": 308,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 941,
+ "bbox": [
+ 108,
+ 248,
+ 36,
+ 20
+ ],
+ "category_id": 4,
+ "id": 1896,
+ "area": 720,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 941,
+ "bbox": [
+ 276,
+ 282,
+ 24,
+ 15
+ ],
+ "category_id": 4,
+ "id": 1897,
+ "area": 360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 941,
+ "bbox": [
+ 158,
+ 433,
+ 28,
+ 18
+ ],
+ "category_id": 4,
+ "id": 1898,
+ "area": 504,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 941,
+ "bbox": [
+ 445,
+ 392,
+ 26,
+ 13
+ ],
+ "category_id": 4,
+ "id": 1899,
+ "area": 338,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 942,
+ "bbox": [
+ 317,
+ 186,
+ 95,
+ 262
+ ],
+ "category_id": 16,
+ "id": 1900,
+ "area": 24890,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 943,
+ "bbox": [
+ 138,
+ 150,
+ 228,
+ 75
+ ],
+ "category_id": 19,
+ "id": 1901,
+ "area": 17100,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 944,
+ "bbox": [
+ 119,
+ 392,
+ 66,
+ 38
+ ],
+ "category_id": 15,
+ "id": 1902,
+ "area": 2508,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 944,
+ "bbox": [
+ 274,
+ 192,
+ 113,
+ 89
+ ],
+ "category_id": 15,
+ "id": 1903,
+ "area": 10057,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 945,
+ "bbox": [
+ 31,
+ 270,
+ 415,
+ 49
+ ],
+ "category_id": 10,
+ "id": 1904,
+ "area": 20335,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 946,
+ "bbox": [
+ 73,
+ 119,
+ 381,
+ 309
+ ],
+ "category_id": 19,
+ "id": 1905,
+ "area": 117729,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 947,
+ "bbox": [
+ 96,
+ 199,
+ 167,
+ 218
+ ],
+ "category_id": 15,
+ "id": 1906,
+ "area": 36406,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 947,
+ "bbox": [
+ 339,
+ 206,
+ 161,
+ 211
+ ],
+ "category_id": 15,
+ "id": 1907,
+ "area": 33971,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 948,
+ "bbox": [
+ 63,
+ 170,
+ 93,
+ 142
+ ],
+ "category_id": 15,
+ "id": 1908,
+ "area": 13206,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 948,
+ "bbox": [
+ 124,
+ 289,
+ 94,
+ 140
+ ],
+ "category_id": 15,
+ "id": 1909,
+ "area": 13160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 949,
+ "bbox": [
+ 88,
+ 157,
+ 260,
+ 246
+ ],
+ "category_id": 16,
+ "id": 1910,
+ "area": 63960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 950,
+ "bbox": [
+ 200,
+ 116,
+ 111,
+ 186
+ ],
+ "category_id": 19,
+ "id": 1911,
+ "area": 20646,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 951,
+ "bbox": [
+ 59,
+ 96,
+ 49,
+ 34
+ ],
+ "category_id": 4,
+ "id": 1912,
+ "area": 1666,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 951,
+ "bbox": [
+ 257,
+ 259,
+ 47,
+ 30
+ ],
+ "category_id": 4,
+ "id": 1913,
+ "area": 1410,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 951,
+ "bbox": [
+ 336,
+ 407,
+ 43,
+ 26
+ ],
+ "category_id": 4,
+ "id": 1914,
+ "area": 1118,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 952,
+ "bbox": [
+ 112,
+ 203,
+ 348,
+ 193
+ ],
+ "category_id": 10,
+ "id": 1915,
+ "area": 67164,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 953,
+ "bbox": [
+ 79,
+ 396,
+ 23,
+ 53
+ ],
+ "category_id": 4,
+ "id": 1916,
+ "area": 1219,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 953,
+ "bbox": [
+ 428,
+ 155,
+ 20,
+ 68
+ ],
+ "category_id": 4,
+ "id": 1917,
+ "area": 1360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 954,
+ "bbox": [
+ 303,
+ 58,
+ 155,
+ 243
+ ],
+ "category_id": 10,
+ "id": 1918,
+ "area": 37665,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 955,
+ "bbox": [
+ 205,
+ 150,
+ 71,
+ 175
+ ],
+ "category_id": 19,
+ "id": 1919,
+ "area": 12425,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 956,
+ "bbox": [
+ 158,
+ 226,
+ 180,
+ 73
+ ],
+ "category_id": 19,
+ "id": 1920,
+ "area": 13140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 957,
+ "bbox": [
+ 336,
+ 243,
+ 46,
+ 82
+ ],
+ "category_id": 4,
+ "id": 1921,
+ "area": 3772,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 958,
+ "bbox": [
+ 167,
+ 110,
+ 108,
+ 101
+ ],
+ "category_id": 15,
+ "id": 1922,
+ "area": 10908,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 958,
+ "bbox": [
+ 95,
+ 335,
+ 110,
+ 93
+ ],
+ "category_id": 15,
+ "id": 1923,
+ "area": 10230,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 959,
+ "bbox": [
+ 314,
+ 80,
+ 118,
+ 165
+ ],
+ "category_id": 19,
+ "id": 1924,
+ "area": 19470,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 960,
+ "bbox": [
+ 182,
+ 88,
+ 274,
+ 370
+ ],
+ "category_id": 10,
+ "id": 1925,
+ "area": 101380,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 961,
+ "bbox": [
+ 128,
+ 42,
+ 284,
+ 409
+ ],
+ "category_id": 10,
+ "id": 1926,
+ "area": 116156,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 962,
+ "bbox": [
+ 172,
+ 118,
+ 73,
+ 66
+ ],
+ "category_id": 15,
+ "id": 1927,
+ "area": 4818,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 962,
+ "bbox": [
+ 231,
+ 210,
+ 75,
+ 65
+ ],
+ "category_id": 15,
+ "id": 1928,
+ "area": 4875,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 962,
+ "bbox": [
+ 266,
+ 403,
+ 76,
+ 66
+ ],
+ "category_id": 15,
+ "id": 1929,
+ "area": 5016,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 963,
+ "bbox": [
+ 126,
+ 99,
+ 35,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1930,
+ "area": 805,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 963,
+ "bbox": [
+ 285,
+ 240,
+ 61,
+ 28
+ ],
+ "category_id": 4,
+ "id": 1931,
+ "area": 1708,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 963,
+ "bbox": [
+ 416,
+ 362,
+ 44,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1932,
+ "area": 1056,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 964,
+ "bbox": [
+ 176,
+ 203,
+ 158,
+ 160
+ ],
+ "category_id": 16,
+ "id": 1933,
+ "area": 25280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 965,
+ "bbox": [
+ 279,
+ 147,
+ 59,
+ 207
+ ],
+ "category_id": 19,
+ "id": 1934,
+ "area": 12213,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 966,
+ "bbox": [
+ 136,
+ 165,
+ 39,
+ 31
+ ],
+ "category_id": 4,
+ "id": 1935,
+ "area": 1209,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 966,
+ "bbox": [
+ 380,
+ 83,
+ 41,
+ 32
+ ],
+ "category_id": 4,
+ "id": 1936,
+ "area": 1312,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 966,
+ "bbox": [
+ 274,
+ 419,
+ 45,
+ 31
+ ],
+ "category_id": 4,
+ "id": 1937,
+ "area": 1395,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 967,
+ "bbox": [
+ 108,
+ 197,
+ 281,
+ 211
+ ],
+ "category_id": 16,
+ "id": 1938,
+ "area": 59291,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 968,
+ "bbox": [
+ 176,
+ 138,
+ 162,
+ 189
+ ],
+ "category_id": 19,
+ "id": 1939,
+ "area": 30618,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 969,
+ "bbox": [
+ 45,
+ 154,
+ 387,
+ 194
+ ],
+ "category_id": 16,
+ "id": 1940,
+ "area": 75078,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 970,
+ "bbox": [
+ 134,
+ 115,
+ 42,
+ 48
+ ],
+ "category_id": 4,
+ "id": 1941,
+ "area": 2016,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 970,
+ "bbox": [
+ 399,
+ 391,
+ 41,
+ 34
+ ],
+ "category_id": 4,
+ "id": 1942,
+ "area": 1394,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 971,
+ "bbox": [
+ 160,
+ 106,
+ 220,
+ 306
+ ],
+ "category_id": 16,
+ "id": 1943,
+ "area": 67320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 972,
+ "bbox": [
+ 52,
+ 349,
+ 49,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1944,
+ "area": 1176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 972,
+ "bbox": [
+ 211,
+ 120,
+ 45,
+ 29
+ ],
+ "category_id": 4,
+ "id": 1945,
+ "area": 1305,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 972,
+ "bbox": [
+ 445,
+ 307,
+ 43,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1946,
+ "area": 989,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 973,
+ "bbox": [
+ 83,
+ 147,
+ 74,
+ 185
+ ],
+ "category_id": 19,
+ "id": 1947,
+ "area": 13690,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 973,
+ "bbox": [
+ 402,
+ 0,
+ 108,
+ 204
+ ],
+ "category_id": 10,
+ "id": 1948,
+ "area": 22032,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 974,
+ "bbox": [
+ 31,
+ 351,
+ 24,
+ 13
+ ],
+ "category_id": 4,
+ "id": 1949,
+ "area": 312,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 974,
+ "bbox": [
+ 143,
+ 328,
+ 21,
+ 9
+ ],
+ "category_id": 4,
+ "id": 1950,
+ "area": 189,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 974,
+ "bbox": [
+ 231,
+ 332,
+ 22,
+ 11
+ ],
+ "category_id": 4,
+ "id": 1951,
+ "area": 242,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 974,
+ "bbox": [
+ 357,
+ 339,
+ 18,
+ 14
+ ],
+ "category_id": 4,
+ "id": 1952,
+ "area": 252,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 974,
+ "bbox": [
+ 432,
+ 337,
+ 20,
+ 12
+ ],
+ "category_id": 4,
+ "id": 1953,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 974,
+ "bbox": [
+ 375,
+ 258,
+ 19,
+ 10
+ ],
+ "category_id": 4,
+ "id": 1954,
+ "area": 190,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 975,
+ "bbox": [
+ 284,
+ 84,
+ 33,
+ 67
+ ],
+ "category_id": 4,
+ "id": 1955,
+ "area": 2211,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 975,
+ "bbox": [
+ 165,
+ 447,
+ 39,
+ 56
+ ],
+ "category_id": 4,
+ "id": 1956,
+ "area": 2184,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 976,
+ "bbox": [
+ 92,
+ 158,
+ 36,
+ 50
+ ],
+ "category_id": 4,
+ "id": 1957,
+ "area": 1800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 976,
+ "bbox": [
+ 410,
+ 278,
+ 40,
+ 46
+ ],
+ "category_id": 4,
+ "id": 1958,
+ "area": 1840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 977,
+ "bbox": [
+ 102,
+ 239,
+ 293,
+ 130
+ ],
+ "category_id": 16,
+ "id": 1959,
+ "area": 38090,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 978,
+ "bbox": [
+ 158,
+ 152,
+ 133,
+ 139
+ ],
+ "category_id": 15,
+ "id": 1960,
+ "area": 18487,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 978,
+ "bbox": [
+ 327,
+ 190,
+ 109,
+ 110
+ ],
+ "category_id": 15,
+ "id": 1961,
+ "area": 11990,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 979,
+ "bbox": [
+ 46,
+ 23,
+ 349,
+ 439
+ ],
+ "category_id": 10,
+ "id": 1962,
+ "area": 153211,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 980,
+ "bbox": [
+ 98,
+ 224,
+ 288,
+ 111
+ ],
+ "category_id": 19,
+ "id": 1963,
+ "area": 31968,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 981,
+ "bbox": [
+ 63,
+ 120,
+ 19,
+ 19
+ ],
+ "category_id": 4,
+ "id": 1964,
+ "area": 361,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 981,
+ "bbox": [
+ 127,
+ 461,
+ 23,
+ 19
+ ],
+ "category_id": 4,
+ "id": 1965,
+ "area": 437,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 981,
+ "bbox": [
+ 178,
+ 291,
+ 22,
+ 21
+ ],
+ "category_id": 4,
+ "id": 1966,
+ "area": 462,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 981,
+ "bbox": [
+ 343,
+ 176,
+ 24,
+ 17
+ ],
+ "category_id": 4,
+ "id": 1967,
+ "area": 408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 981,
+ "bbox": [
+ 464,
+ 61,
+ 24,
+ 20
+ ],
+ "category_id": 4,
+ "id": 1968,
+ "area": 480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 982,
+ "bbox": [
+ 85,
+ 136,
+ 371,
+ 223
+ ],
+ "category_id": 10,
+ "id": 1969,
+ "area": 82733,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 983,
+ "bbox": [
+ 52,
+ 136,
+ 439,
+ 196
+ ],
+ "category_id": 10,
+ "id": 1970,
+ "area": 86044,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 984,
+ "bbox": [
+ 343,
+ 87,
+ 62,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1971,
+ "area": 1550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 984,
+ "bbox": [
+ 192,
+ 406,
+ 62,
+ 26
+ ],
+ "category_id": 4,
+ "id": 1972,
+ "area": 1612,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 985,
+ "bbox": [
+ 120,
+ 254,
+ 21,
+ 13
+ ],
+ "category_id": 15,
+ "id": 1973,
+ "area": 273,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 985,
+ "bbox": [
+ 135,
+ 343,
+ 19,
+ 15
+ ],
+ "category_id": 15,
+ "id": 1974,
+ "area": 285,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 985,
+ "bbox": [
+ 88,
+ 304,
+ 20,
+ 14
+ ],
+ "category_id": 15,
+ "id": 1975,
+ "area": 280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 986,
+ "bbox": [
+ 230,
+ 232,
+ 186,
+ 54
+ ],
+ "category_id": 19,
+ "id": 1976,
+ "area": 10044,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 987,
+ "bbox": [
+ 225,
+ 247,
+ 155,
+ 106
+ ],
+ "category_id": 10,
+ "id": 1977,
+ "area": 16430,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 988,
+ "bbox": [
+ 186,
+ 218,
+ 160,
+ 46
+ ],
+ "category_id": 19,
+ "id": 1978,
+ "area": 7360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 989,
+ "bbox": [
+ 16,
+ 199,
+ 72,
+ 159
+ ],
+ "category_id": 15,
+ "id": 1979,
+ "area": 11448,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 989,
+ "bbox": [
+ 83,
+ 223,
+ 91,
+ 115
+ ],
+ "category_id": 15,
+ "id": 1980,
+ "area": 10465,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 989,
+ "bbox": [
+ 190,
+ 224,
+ 145,
+ 171
+ ],
+ "category_id": 15,
+ "id": 1981,
+ "area": 24795,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 990,
+ "bbox": [
+ 197,
+ 108,
+ 225,
+ 354
+ ],
+ "category_id": 10,
+ "id": 1982,
+ "area": 79650,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 991,
+ "bbox": [
+ 22,
+ 88,
+ 213,
+ 271
+ ],
+ "category_id": 15,
+ "id": 1983,
+ "area": 57723,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 991,
+ "bbox": [
+ 282,
+ 184,
+ 218,
+ 271
+ ],
+ "category_id": 15,
+ "id": 1984,
+ "area": 59078,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 992,
+ "bbox": [
+ 107,
+ 88,
+ 296,
+ 409
+ ],
+ "category_id": 19,
+ "id": 1985,
+ "area": 121064,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 993,
+ "bbox": [
+ 98,
+ 268,
+ 210,
+ 32
+ ],
+ "category_id": 19,
+ "id": 1986,
+ "area": 6720,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 994,
+ "bbox": [
+ 119,
+ 124,
+ 39,
+ 31
+ ],
+ "category_id": 4,
+ "id": 1987,
+ "area": 1209,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 994,
+ "bbox": [
+ 280,
+ 422,
+ 30,
+ 25
+ ],
+ "category_id": 4,
+ "id": 1988,
+ "area": 750,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 994,
+ "bbox": [
+ 449,
+ 92,
+ 26,
+ 35
+ ],
+ "category_id": 4,
+ "id": 1989,
+ "area": 910,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 995,
+ "bbox": [
+ 97,
+ 83,
+ 135,
+ 283
+ ],
+ "category_id": 10,
+ "id": 1990,
+ "area": 38205,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 996,
+ "bbox": [
+ 304,
+ 95,
+ 77,
+ 52
+ ],
+ "category_id": 4,
+ "id": 1991,
+ "area": 4004,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 996,
+ "bbox": [
+ 210,
+ 366,
+ 78,
+ 56
+ ],
+ "category_id": 4,
+ "id": 1992,
+ "area": 4368,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 997,
+ "bbox": [
+ 33,
+ 184,
+ 427,
+ 91
+ ],
+ "category_id": 10,
+ "id": 1993,
+ "area": 38857,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 998,
+ "bbox": [
+ 56,
+ 72,
+ 17,
+ 29
+ ],
+ "category_id": 4,
+ "id": 1994,
+ "area": 493,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 998,
+ "bbox": [
+ 199,
+ 159,
+ 17,
+ 24
+ ],
+ "category_id": 4,
+ "id": 1995,
+ "area": 408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 998,
+ "bbox": [
+ 248,
+ 240,
+ 20,
+ 23
+ ],
+ "category_id": 4,
+ "id": 1996,
+ "area": 460,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 998,
+ "bbox": [
+ 244,
+ 392,
+ 23,
+ 22
+ ],
+ "category_id": 4,
+ "id": 1997,
+ "area": 506,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 998,
+ "bbox": [
+ 423,
+ 435,
+ 29,
+ 17
+ ],
+ "category_id": 4,
+ "id": 1998,
+ "area": 493,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 999,
+ "bbox": [
+ 268,
+ 242,
+ 84,
+ 84
+ ],
+ "category_id": 15,
+ "id": 1999,
+ "area": 7056,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1000,
+ "bbox": [
+ 299,
+ 68,
+ 21,
+ 22
+ ],
+ "category_id": 4,
+ "id": 2000,
+ "area": 462,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1000,
+ "bbox": [
+ 269,
+ 182,
+ 23,
+ 22
+ ],
+ "category_id": 4,
+ "id": 2001,
+ "area": 506,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1000,
+ "bbox": [
+ 127,
+ 161,
+ 24,
+ 23
+ ],
+ "category_id": 4,
+ "id": 2002,
+ "area": 552,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1000,
+ "bbox": [
+ 289,
+ 326,
+ 31,
+ 26
+ ],
+ "category_id": 4,
+ "id": 2003,
+ "area": 806,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1000,
+ "bbox": [
+ 272,
+ 433,
+ 23,
+ 21
+ ],
+ "category_id": 4,
+ "id": 2004,
+ "area": 483,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1000,
+ "bbox": [
+ 165,
+ 440,
+ 27,
+ 28
+ ],
+ "category_id": 4,
+ "id": 2005,
+ "area": 756,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1001,
+ "bbox": [
+ 116,
+ 296,
+ 50,
+ 48
+ ],
+ "category_id": 4,
+ "id": 2006,
+ "area": 2400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1001,
+ "bbox": [
+ 87,
+ 154,
+ 52,
+ 46
+ ],
+ "category_id": 4,
+ "id": 2007,
+ "area": 2392,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1001,
+ "bbox": [
+ 432,
+ 325,
+ 56,
+ 47
+ ],
+ "category_id": 4,
+ "id": 2008,
+ "area": 2632,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1002,
+ "bbox": [
+ 162,
+ 98,
+ 31,
+ 51
+ ],
+ "category_id": 4,
+ "id": 2009,
+ "area": 1581,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1002,
+ "bbox": [
+ 41,
+ 373,
+ 28,
+ 57
+ ],
+ "category_id": 4,
+ "id": 2010,
+ "area": 1596,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1002,
+ "bbox": [
+ 421,
+ 350,
+ 50,
+ 51
+ ],
+ "category_id": 4,
+ "id": 2011,
+ "area": 2550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1003,
+ "bbox": [
+ 362,
+ 346,
+ 45,
+ 44
+ ],
+ "category_id": 19,
+ "id": 2012,
+ "area": 1980,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1004,
+ "bbox": [
+ 218,
+ 275,
+ 85,
+ 62
+ ],
+ "category_id": 4,
+ "id": 2013,
+ "area": 5270,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1005,
+ "bbox": [
+ 160,
+ 247,
+ 129,
+ 121
+ ],
+ "category_id": 10,
+ "id": 2014,
+ "area": 15609,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1006,
+ "bbox": [
+ 267,
+ 74,
+ 72,
+ 331
+ ],
+ "category_id": 10,
+ "id": 2015,
+ "area": 23832,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1007,
+ "bbox": [
+ 417,
+ 47,
+ 25,
+ 27
+ ],
+ "category_id": 4,
+ "id": 2016,
+ "area": 675,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1007,
+ "bbox": [
+ 190,
+ 137,
+ 21,
+ 29
+ ],
+ "category_id": 4,
+ "id": 2017,
+ "area": 609,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1007,
+ "bbox": [
+ 364,
+ 135,
+ 14,
+ 23
+ ],
+ "category_id": 4,
+ "id": 2018,
+ "area": 322,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1007,
+ "bbox": [
+ 382,
+ 302,
+ 29,
+ 32
+ ],
+ "category_id": 4,
+ "id": 2019,
+ "area": 928,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1007,
+ "bbox": [
+ 216,
+ 435,
+ 26,
+ 29
+ ],
+ "category_id": 4,
+ "id": 2020,
+ "area": 754,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1007,
+ "bbox": [
+ 382,
+ 449,
+ 22,
+ 32
+ ],
+ "category_id": 4,
+ "id": 2021,
+ "area": 704,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1008,
+ "bbox": [
+ 59,
+ 364,
+ 92,
+ 36
+ ],
+ "category_id": 4,
+ "id": 2022,
+ "area": 3312,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1008,
+ "bbox": [
+ 319,
+ 26,
+ 95,
+ 38
+ ],
+ "category_id": 4,
+ "id": 2023,
+ "area": 3610,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1009,
+ "bbox": [
+ 200,
+ 206,
+ 146,
+ 119
+ ],
+ "category_id": 16,
+ "id": 2024,
+ "area": 17374,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1010,
+ "bbox": [
+ 42,
+ 287,
+ 37,
+ 37
+ ],
+ "category_id": 4,
+ "id": 2025,
+ "area": 1369,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1010,
+ "bbox": [
+ 240,
+ 109,
+ 37,
+ 38
+ ],
+ "category_id": 4,
+ "id": 2026,
+ "area": 1406,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1010,
+ "bbox": [
+ 427,
+ 410,
+ 33,
+ 30
+ ],
+ "category_id": 4,
+ "id": 2027,
+ "area": 990,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1011,
+ "bbox": [
+ 97,
+ 236,
+ 161,
+ 211
+ ],
+ "category_id": 16,
+ "id": 2028,
+ "area": 33971,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1012,
+ "bbox": [
+ 180,
+ 180,
+ 127,
+ 106
+ ],
+ "category_id": 19,
+ "id": 2029,
+ "area": 13462,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1013,
+ "bbox": [
+ 195,
+ 245,
+ 213,
+ 154
+ ],
+ "category_id": 16,
+ "id": 2030,
+ "area": 32802,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1014,
+ "bbox": [
+ 139,
+ 92,
+ 145,
+ 375
+ ],
+ "category_id": 16,
+ "id": 2031,
+ "area": 54375,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1015,
+ "bbox": [
+ 143,
+ 160,
+ 107,
+ 112
+ ],
+ "category_id": 15,
+ "id": 2032,
+ "area": 11984,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1015,
+ "bbox": [
+ 156,
+ 323,
+ 106,
+ 109
+ ],
+ "category_id": 15,
+ "id": 2033,
+ "area": 11554,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1016,
+ "bbox": [
+ 268,
+ 116,
+ 228,
+ 239
+ ],
+ "category_id": 10,
+ "id": 2034,
+ "area": 54492,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1016,
+ "bbox": [
+ 24,
+ 7,
+ 203,
+ 191
+ ],
+ "category_id": 10,
+ "id": 2035,
+ "area": 38773,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1017,
+ "bbox": [
+ 134,
+ 430,
+ 273,
+ 57
+ ],
+ "category_id": 10,
+ "id": 2036,
+ "area": 15561,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1018,
+ "bbox": [
+ 67,
+ 0,
+ 119,
+ 237
+ ],
+ "category_id": 15,
+ "id": 2037,
+ "area": 28203,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1018,
+ "bbox": [
+ 131,
+ 111,
+ 317,
+ 371
+ ],
+ "category_id": 15,
+ "id": 2038,
+ "area": 117607,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1019,
+ "bbox": [
+ 3,
+ 203,
+ 168,
+ 168
+ ],
+ "category_id": 15,
+ "id": 2039,
+ "area": 28224,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1019,
+ "bbox": [
+ 282,
+ 145,
+ 180,
+ 175
+ ],
+ "category_id": 15,
+ "id": 2040,
+ "area": 31500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1019,
+ "bbox": [
+ 288,
+ 339,
+ 99,
+ 105
+ ],
+ "category_id": 15,
+ "id": 2041,
+ "area": 10395,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1020,
+ "bbox": [
+ 438,
+ 93,
+ 25,
+ 33
+ ],
+ "category_id": 4,
+ "id": 2042,
+ "area": 825,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1020,
+ "bbox": [
+ 140,
+ 62,
+ 25,
+ 36
+ ],
+ "category_id": 4,
+ "id": 2043,
+ "area": 900,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1020,
+ "bbox": [
+ 61,
+ 157,
+ 23,
+ 27
+ ],
+ "category_id": 4,
+ "id": 2044,
+ "area": 621,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1020,
+ "bbox": [
+ 71,
+ 223,
+ 25,
+ 31
+ ],
+ "category_id": 4,
+ "id": 2045,
+ "area": 775,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1020,
+ "bbox": [
+ 90,
+ 307,
+ 23,
+ 28
+ ],
+ "category_id": 4,
+ "id": 2046,
+ "area": 644,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1020,
+ "bbox": [
+ 139,
+ 399,
+ 19,
+ 31
+ ],
+ "category_id": 4,
+ "id": 2047,
+ "area": 589,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1020,
+ "bbox": [
+ 135,
+ 477,
+ 23,
+ 20
+ ],
+ "category_id": 4,
+ "id": 2048,
+ "area": 460,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1021,
+ "bbox": [
+ 39,
+ 229,
+ 51,
+ 44
+ ],
+ "category_id": 4,
+ "id": 2049,
+ "area": 2244,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1021,
+ "bbox": [
+ 449,
+ 236,
+ 46,
+ 47
+ ],
+ "category_id": 4,
+ "id": 2050,
+ "area": 2162,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1022,
+ "bbox": [
+ 113,
+ 135,
+ 43,
+ 80
+ ],
+ "category_id": 16,
+ "id": 2051,
+ "area": 3440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1023,
+ "bbox": [
+ 96,
+ 163,
+ 275,
+ 141
+ ],
+ "category_id": 19,
+ "id": 2052,
+ "area": 38775,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1024,
+ "bbox": [
+ 147,
+ 231,
+ 201,
+ 177
+ ],
+ "category_id": 16,
+ "id": 2053,
+ "area": 35577,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1025,
+ "bbox": [
+ 100,
+ 175,
+ 263,
+ 93
+ ],
+ "category_id": 19,
+ "id": 2054,
+ "area": 24459,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1026,
+ "bbox": [
+ 455,
+ 320,
+ 33,
+ 8
+ ],
+ "category_id": 16,
+ "id": 2055,
+ "area": 264,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1027,
+ "bbox": [
+ 92,
+ 87,
+ 351,
+ 303
+ ],
+ "category_id": 10,
+ "id": 2056,
+ "area": 106353,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1028,
+ "bbox": [
+ 204,
+ 222,
+ 245,
+ 49
+ ],
+ "category_id": 10,
+ "id": 2057,
+ "area": 12005,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1029,
+ "bbox": [
+ 37,
+ 106,
+ 110,
+ 154
+ ],
+ "category_id": 10,
+ "id": 2058,
+ "area": 16940,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1030,
+ "bbox": [
+ 146,
+ 109,
+ 37,
+ 63
+ ],
+ "category_id": 16,
+ "id": 2059,
+ "area": 2331,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1031,
+ "bbox": [
+ 56,
+ 157,
+ 245,
+ 168
+ ],
+ "category_id": 16,
+ "id": 2060,
+ "area": 41160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1031,
+ "bbox": [
+ 1,
+ 278,
+ 29,
+ 41
+ ],
+ "category_id": 16,
+ "id": 2061,
+ "area": 1189,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1032,
+ "bbox": [
+ 14,
+ 221,
+ 469,
+ 77
+ ],
+ "category_id": 10,
+ "id": 2062,
+ "area": 36113,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1033,
+ "bbox": [
+ 110,
+ 67,
+ 333,
+ 336
+ ],
+ "category_id": 10,
+ "id": 2063,
+ "area": 111888,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1034,
+ "bbox": [
+ 143,
+ 112,
+ 79,
+ 25
+ ],
+ "category_id": 4,
+ "id": 2064,
+ "area": 1975,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1034,
+ "bbox": [
+ 314,
+ 439,
+ 69,
+ 26
+ ],
+ "category_id": 4,
+ "id": 2065,
+ "area": 1794,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1035,
+ "bbox": [
+ 154,
+ 231,
+ 123,
+ 90
+ ],
+ "category_id": 15,
+ "id": 2066,
+ "area": 11070,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1035,
+ "bbox": [
+ 268,
+ 203,
+ 123,
+ 91
+ ],
+ "category_id": 15,
+ "id": 2067,
+ "area": 11193,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1036,
+ "bbox": [
+ 104,
+ 161,
+ 94,
+ 35
+ ],
+ "category_id": 4,
+ "id": 2068,
+ "area": 3290,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1036,
+ "bbox": [
+ 420,
+ 398,
+ 72,
+ 58
+ ],
+ "category_id": 4,
+ "id": 2069,
+ "area": 4176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1037,
+ "bbox": [
+ 34,
+ 268,
+ 70,
+ 221
+ ],
+ "category_id": 10,
+ "id": 2070,
+ "area": 15470,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1038,
+ "bbox": [
+ 137,
+ 211,
+ 135,
+ 114
+ ],
+ "category_id": 15,
+ "id": 2071,
+ "area": 15390,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1038,
+ "bbox": [
+ 284,
+ 208,
+ 132,
+ 117
+ ],
+ "category_id": 15,
+ "id": 2072,
+ "area": 15444,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1039,
+ "bbox": [
+ 27,
+ 28,
+ 349,
+ 21
+ ],
+ "category_id": 10,
+ "id": 2073,
+ "area": 7329,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1040,
+ "bbox": [
+ 197,
+ 7,
+ 139,
+ 316
+ ],
+ "category_id": 10,
+ "id": 2074,
+ "area": 43924,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1041,
+ "bbox": [
+ 108,
+ 206,
+ 138,
+ 114
+ ],
+ "category_id": 15,
+ "id": 2075,
+ "area": 15732,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1041,
+ "bbox": [
+ 258,
+ 216,
+ 141,
+ 112
+ ],
+ "category_id": 15,
+ "id": 2076,
+ "area": 15792,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1042,
+ "bbox": [
+ 126,
+ 172,
+ 55,
+ 21
+ ],
+ "category_id": 4,
+ "id": 2077,
+ "area": 1155,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1042,
+ "bbox": [
+ 352,
+ 353,
+ 55,
+ 41
+ ],
+ "category_id": 4,
+ "id": 2078,
+ "area": 2255,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1043,
+ "bbox": [
+ 310,
+ 144,
+ 44,
+ 39
+ ],
+ "category_id": 16,
+ "id": 2079,
+ "area": 1716,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1043,
+ "bbox": [
+ 147,
+ 428,
+ 44,
+ 31
+ ],
+ "category_id": 16,
+ "id": 2080,
+ "area": 1364,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1044,
+ "bbox": [
+ 221,
+ 216,
+ 24,
+ 84
+ ],
+ "category_id": 4,
+ "id": 2081,
+ "area": 2016,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1045,
+ "bbox": [
+ 224,
+ 95,
+ 129,
+ 305
+ ],
+ "category_id": 16,
+ "id": 2082,
+ "area": 39345,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1046,
+ "bbox": [
+ 197,
+ 136,
+ 85,
+ 262
+ ],
+ "category_id": 19,
+ "id": 2083,
+ "area": 22270,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1047,
+ "bbox": [
+ 268,
+ 261,
+ 172,
+ 46
+ ],
+ "category_id": 19,
+ "id": 2084,
+ "area": 7912,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1048,
+ "bbox": [
+ 172,
+ 115,
+ 79,
+ 34
+ ],
+ "category_id": 4,
+ "id": 2085,
+ "area": 2686,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1048,
+ "bbox": [
+ 59,
+ 324,
+ 74,
+ 50
+ ],
+ "category_id": 4,
+ "id": 2086,
+ "area": 3700,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1048,
+ "bbox": [
+ 371,
+ 292,
+ 78,
+ 47
+ ],
+ "category_id": 4,
+ "id": 2087,
+ "area": 3666,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1049,
+ "bbox": [
+ 147,
+ 81,
+ 187,
+ 319
+ ],
+ "category_id": 16,
+ "id": 2088,
+ "area": 59653,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1050,
+ "bbox": [
+ 314,
+ 78,
+ 25,
+ 42
+ ],
+ "category_id": 4,
+ "id": 2089,
+ "area": 1050,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1050,
+ "bbox": [
+ 115,
+ 239,
+ 25,
+ 47
+ ],
+ "category_id": 4,
+ "id": 2090,
+ "area": 1175,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1050,
+ "bbox": [
+ 398,
+ 405,
+ 28,
+ 37
+ ],
+ "category_id": 4,
+ "id": 2091,
+ "area": 1036,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1051,
+ "bbox": [
+ 33,
+ 61,
+ 38,
+ 25
+ ],
+ "category_id": 4,
+ "id": 2092,
+ "area": 950,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1051,
+ "bbox": [
+ 184,
+ 2,
+ 28,
+ 22
+ ],
+ "category_id": 4,
+ "id": 2093,
+ "area": 616,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1051,
+ "bbox": [
+ 172,
+ 147,
+ 27,
+ 18
+ ],
+ "category_id": 4,
+ "id": 2094,
+ "area": 486,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1051,
+ "bbox": [
+ 36,
+ 300,
+ 22,
+ 25
+ ],
+ "category_id": 4,
+ "id": 2095,
+ "area": 550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1051,
+ "bbox": [
+ 144,
+ 362,
+ 33,
+ 18
+ ],
+ "category_id": 4,
+ "id": 2096,
+ "area": 594,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1051,
+ "bbox": [
+ 419,
+ 193,
+ 21,
+ 23
+ ],
+ "category_id": 4,
+ "id": 2097,
+ "area": 483,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1051,
+ "bbox": [
+ 335,
+ 339,
+ 26,
+ 16
+ ],
+ "category_id": 4,
+ "id": 2098,
+ "area": 416,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1051,
+ "bbox": [
+ 325,
+ 490,
+ 19,
+ 16
+ ],
+ "category_id": 4,
+ "id": 2099,
+ "area": 304,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1052,
+ "bbox": [
+ 281,
+ 154,
+ 115,
+ 310
+ ],
+ "category_id": 16,
+ "id": 2100,
+ "area": 35650,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1053,
+ "bbox": [
+ 192,
+ 63,
+ 51,
+ 35
+ ],
+ "category_id": 19,
+ "id": 2101,
+ "area": 1785,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1054,
+ "bbox": [
+ 166,
+ 115,
+ 196,
+ 302
+ ],
+ "category_id": 15,
+ "id": 2102,
+ "area": 59192,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1055,
+ "bbox": [
+ 238,
+ 362,
+ 68,
+ 57
+ ],
+ "category_id": 4,
+ "id": 2103,
+ "area": 3876,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1056,
+ "bbox": [
+ 309,
+ 93,
+ 96,
+ 118
+ ],
+ "category_id": 15,
+ "id": 2104,
+ "area": 11328,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1056,
+ "bbox": [
+ 252,
+ 204,
+ 96,
+ 121
+ ],
+ "category_id": 15,
+ "id": 2105,
+ "area": 11616,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1057,
+ "bbox": [
+ 65,
+ 247,
+ 29,
+ 49
+ ],
+ "category_id": 4,
+ "id": 2106,
+ "area": 1421,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1057,
+ "bbox": [
+ 415,
+ 199,
+ 36,
+ 49
+ ],
+ "category_id": 4,
+ "id": 2107,
+ "area": 1764,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1058,
+ "bbox": [
+ 147,
+ 256,
+ 141,
+ 128
+ ],
+ "category_id": 15,
+ "id": 2108,
+ "area": 18048,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1059,
+ "bbox": [
+ 330,
+ 349,
+ 123,
+ 55
+ ],
+ "category_id": 10,
+ "id": 2109,
+ "area": 6765,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1060,
+ "bbox": [
+ 149,
+ 167,
+ 36,
+ 72
+ ],
+ "category_id": 16,
+ "id": 2110,
+ "area": 2592,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1061,
+ "bbox": [
+ 446,
+ 286,
+ 29,
+ 10
+ ],
+ "category_id": 15,
+ "id": 2111,
+ "area": 290,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1062,
+ "bbox": [
+ 124,
+ 149,
+ 101,
+ 135
+ ],
+ "category_id": 15,
+ "id": 2112,
+ "area": 13635,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1062,
+ "bbox": [
+ 225,
+ 226,
+ 82,
+ 113
+ ],
+ "category_id": 15,
+ "id": 2113,
+ "area": 9266,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1062,
+ "bbox": [
+ 307,
+ 143,
+ 69,
+ 109
+ ],
+ "category_id": 15,
+ "id": 2114,
+ "area": 7521,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1063,
+ "bbox": [
+ 291,
+ 176,
+ 94,
+ 122
+ ],
+ "category_id": 4,
+ "id": 2115,
+ "area": 11468,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1064,
+ "bbox": [
+ 81,
+ 188,
+ 315,
+ 235
+ ],
+ "category_id": 16,
+ "id": 2116,
+ "area": 74025,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1065,
+ "bbox": [
+ 112,
+ 247,
+ 241,
+ 72
+ ],
+ "category_id": 19,
+ "id": 2117,
+ "area": 17352,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1066,
+ "bbox": [
+ 21,
+ 56,
+ 16,
+ 27
+ ],
+ "category_id": 4,
+ "id": 2118,
+ "area": 432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1066,
+ "bbox": [
+ 87,
+ 284,
+ 17,
+ 30
+ ],
+ "category_id": 4,
+ "id": 2119,
+ "area": 510,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1066,
+ "bbox": [
+ 186,
+ 218,
+ 14,
+ 23
+ ],
+ "category_id": 4,
+ "id": 2120,
+ "area": 322,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1066,
+ "bbox": [
+ 290,
+ 149,
+ 15,
+ 31
+ ],
+ "category_id": 4,
+ "id": 2121,
+ "area": 465,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1066,
+ "bbox": [
+ 392,
+ 80,
+ 18,
+ 26
+ ],
+ "category_id": 4,
+ "id": 2122,
+ "area": 468,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1066,
+ "bbox": [
+ 369,
+ 314,
+ 17,
+ 25
+ ],
+ "category_id": 4,
+ "id": 2123,
+ "area": 425,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1066,
+ "bbox": [
+ 264,
+ 378,
+ 16,
+ 26
+ ],
+ "category_id": 4,
+ "id": 2124,
+ "area": 416,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1067,
+ "bbox": [
+ 17,
+ 33,
+ 15,
+ 9
+ ],
+ "category_id": 4,
+ "id": 2125,
+ "area": 135,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1067,
+ "bbox": [
+ 296,
+ 94,
+ 15,
+ 7
+ ],
+ "category_id": 4,
+ "id": 2126,
+ "area": 105,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1067,
+ "bbox": [
+ 132,
+ 118,
+ 15,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2127,
+ "area": 150,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1067,
+ "bbox": [
+ 204,
+ 247,
+ 16,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2128,
+ "area": 160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1067,
+ "bbox": [
+ 358,
+ 369,
+ 15,
+ 11
+ ],
+ "category_id": 4,
+ "id": 2129,
+ "area": 165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1067,
+ "bbox": [
+ 481,
+ 477,
+ 15,
+ 13
+ ],
+ "category_id": 4,
+ "id": 2130,
+ "area": 195,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1067,
+ "bbox": [
+ 169,
+ 459,
+ 15,
+ 8
+ ],
+ "category_id": 4,
+ "id": 2131,
+ "area": 120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1068,
+ "bbox": [
+ 408,
+ 23,
+ 77,
+ 146
+ ],
+ "category_id": 10,
+ "id": 2132,
+ "area": 11242,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1069,
+ "bbox": [
+ 17,
+ 58,
+ 28,
+ 16
+ ],
+ "category_id": 4,
+ "id": 2133,
+ "area": 448,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1069,
+ "bbox": [
+ 98,
+ 208,
+ 21,
+ 12
+ ],
+ "category_id": 4,
+ "id": 2134,
+ "area": 252,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1069,
+ "bbox": [
+ 187,
+ 227,
+ 28,
+ 15
+ ],
+ "category_id": 4,
+ "id": 2135,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1069,
+ "bbox": [
+ 303,
+ 259,
+ 18,
+ 7
+ ],
+ "category_id": 4,
+ "id": 2136,
+ "area": 126,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1069,
+ "bbox": [
+ 468,
+ 256,
+ 28,
+ 17
+ ],
+ "category_id": 4,
+ "id": 2137,
+ "area": 476,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1070,
+ "bbox": [
+ 127,
+ 288,
+ 254,
+ 78
+ ],
+ "category_id": 16,
+ "id": 2138,
+ "area": 19812,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1071,
+ "bbox": [
+ 183,
+ 88,
+ 54,
+ 31
+ ],
+ "category_id": 4,
+ "id": 2139,
+ "area": 1674,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1071,
+ "bbox": [
+ 333,
+ 413,
+ 75,
+ 26
+ ],
+ "category_id": 4,
+ "id": 2140,
+ "area": 1950,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1072,
+ "bbox": [
+ 113,
+ 250,
+ 66,
+ 230
+ ],
+ "category_id": 10,
+ "id": 2141,
+ "area": 15180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1073,
+ "bbox": [
+ 110,
+ 233,
+ 273,
+ 78
+ ],
+ "category_id": 19,
+ "id": 2142,
+ "area": 21294,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1074,
+ "bbox": [
+ 209,
+ 154,
+ 182,
+ 228
+ ],
+ "category_id": 16,
+ "id": 2143,
+ "area": 41496,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1075,
+ "bbox": [
+ 158,
+ 96,
+ 281,
+ 321
+ ],
+ "category_id": 10,
+ "id": 2144,
+ "area": 90201,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1076,
+ "bbox": [
+ 156,
+ 309,
+ 174,
+ 41
+ ],
+ "category_id": 10,
+ "id": 2145,
+ "area": 7134,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1077,
+ "bbox": [
+ 51,
+ 154,
+ 20,
+ 21
+ ],
+ "category_id": 4,
+ "id": 2146,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1077,
+ "bbox": [
+ 206,
+ 168,
+ 22,
+ 22
+ ],
+ "category_id": 4,
+ "id": 2147,
+ "area": 484,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1077,
+ "bbox": [
+ 353,
+ 83,
+ 21,
+ 22
+ ],
+ "category_id": 4,
+ "id": 2148,
+ "area": 462,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1077,
+ "bbox": [
+ 93,
+ 368,
+ 22,
+ 22
+ ],
+ "category_id": 4,
+ "id": 2149,
+ "area": 484,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1077,
+ "bbox": [
+ 311,
+ 355,
+ 21,
+ 21
+ ],
+ "category_id": 4,
+ "id": 2150,
+ "area": 441,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1077,
+ "bbox": [
+ 461,
+ 280,
+ 20,
+ 23
+ ],
+ "category_id": 4,
+ "id": 2151,
+ "area": 460,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1078,
+ "bbox": [
+ 62,
+ 252,
+ 361,
+ 44
+ ],
+ "category_id": 10,
+ "id": 2152,
+ "area": 15884,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1079,
+ "bbox": [
+ 146,
+ 81,
+ 267,
+ 341
+ ],
+ "category_id": 10,
+ "id": 2153,
+ "area": 91047,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1080,
+ "bbox": [
+ 61,
+ 113,
+ 63,
+ 151
+ ],
+ "category_id": 10,
+ "id": 2154,
+ "area": 9513,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1081,
+ "bbox": [
+ 142,
+ 87,
+ 80,
+ 46
+ ],
+ "category_id": 4,
+ "id": 2155,
+ "area": 3680,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1081,
+ "bbox": [
+ 312,
+ 410,
+ 90,
+ 43
+ ],
+ "category_id": 4,
+ "id": 2156,
+ "area": 3870,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1082,
+ "bbox": [
+ 147,
+ 373,
+ 71,
+ 103
+ ],
+ "category_id": 16,
+ "id": 2157,
+ "area": 7313,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1082,
+ "bbox": [
+ 285,
+ 268,
+ 29,
+ 98
+ ],
+ "category_id": 16,
+ "id": 2158,
+ "area": 2842,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1082,
+ "bbox": [
+ 342,
+ 168,
+ 75,
+ 74
+ ],
+ "category_id": 16,
+ "id": 2159,
+ "area": 5550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1083,
+ "bbox": [
+ 72,
+ 128,
+ 112,
+ 135
+ ],
+ "category_id": 15,
+ "id": 2160,
+ "area": 15120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1083,
+ "bbox": [
+ 64,
+ 269,
+ 113,
+ 137
+ ],
+ "category_id": 15,
+ "id": 2161,
+ "area": 15481,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1084,
+ "bbox": [
+ 282,
+ 116,
+ 120,
+ 279
+ ],
+ "category_id": 10,
+ "id": 2162,
+ "area": 33480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1085,
+ "bbox": [
+ 95,
+ 325,
+ 10,
+ 11
+ ],
+ "category_id": 4,
+ "id": 2163,
+ "area": 110,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1085,
+ "bbox": [
+ 10,
+ 301,
+ 5,
+ 9
+ ],
+ "category_id": 4,
+ "id": 2164,
+ "area": 45,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1085,
+ "bbox": [
+ 176,
+ 410,
+ 10,
+ 7
+ ],
+ "category_id": 4,
+ "id": 2165,
+ "area": 70,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1085,
+ "bbox": [
+ 61,
+ 424,
+ 8,
+ 8
+ ],
+ "category_id": 4,
+ "id": 2166,
+ "area": 64,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1085,
+ "bbox": [
+ 151,
+ 472,
+ 10,
+ 8
+ ],
+ "category_id": 4,
+ "id": 2167,
+ "area": 80,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1085,
+ "bbox": [
+ 308,
+ 316,
+ 10,
+ 12
+ ],
+ "category_id": 4,
+ "id": 2168,
+ "area": 120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1085,
+ "bbox": [
+ 352,
+ 339,
+ 12,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2169,
+ "area": 120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1085,
+ "bbox": [
+ 371,
+ 221,
+ 8,
+ 6
+ ],
+ "category_id": 4,
+ "id": 2170,
+ "area": 48,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1085,
+ "bbox": [
+ 325,
+ 115,
+ 8,
+ 7
+ ],
+ "category_id": 4,
+ "id": 2171,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1085,
+ "bbox": [
+ 408,
+ 81,
+ 13,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2172,
+ "area": 130,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1086,
+ "bbox": [
+ 186,
+ 59,
+ 200,
+ 312
+ ],
+ "category_id": 10,
+ "id": 2173,
+ "area": 62400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1087,
+ "bbox": [
+ 21,
+ 387,
+ 51,
+ 43
+ ],
+ "category_id": 4,
+ "id": 2174,
+ "area": 2193,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1087,
+ "bbox": [
+ 270,
+ 171,
+ 36,
+ 44
+ ],
+ "category_id": 4,
+ "id": 2175,
+ "area": 1584,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1087,
+ "bbox": [
+ 456,
+ 286,
+ 32,
+ 54
+ ],
+ "category_id": 4,
+ "id": 2176,
+ "area": 1728,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1088,
+ "bbox": [
+ 216,
+ 151,
+ 118,
+ 217
+ ],
+ "category_id": 16,
+ "id": 2177,
+ "area": 25606,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1089,
+ "bbox": [
+ 192,
+ 284,
+ 114,
+ 101
+ ],
+ "category_id": 15,
+ "id": 2178,
+ "area": 11514,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1089,
+ "bbox": [
+ 279,
+ 282,
+ 92,
+ 73
+ ],
+ "category_id": 15,
+ "id": 2179,
+ "area": 6716,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1090,
+ "bbox": [
+ 68,
+ 113,
+ 52,
+ 53
+ ],
+ "category_id": 4,
+ "id": 2180,
+ "area": 2756,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1090,
+ "bbox": [
+ 408,
+ 75,
+ 48,
+ 49
+ ],
+ "category_id": 4,
+ "id": 2181,
+ "area": 2352,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1090,
+ "bbox": [
+ 372,
+ 440,
+ 67,
+ 49
+ ],
+ "category_id": 4,
+ "id": 2182,
+ "area": 3283,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1091,
+ "bbox": [
+ 25,
+ 162,
+ 469,
+ 291
+ ],
+ "category_id": 10,
+ "id": 2183,
+ "area": 136479,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1092,
+ "bbox": [
+ 366,
+ 33,
+ 55,
+ 160
+ ],
+ "category_id": 10,
+ "id": 2184,
+ "area": 8800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1093,
+ "bbox": [
+ 139,
+ 96,
+ 57,
+ 332
+ ],
+ "category_id": 19,
+ "id": 2185,
+ "area": 18924,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1094,
+ "bbox": [
+ 263,
+ 178,
+ 111,
+ 245
+ ],
+ "category_id": 10,
+ "id": 2186,
+ "area": 27195,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1095,
+ "bbox": [
+ 270,
+ 339,
+ 23,
+ 23
+ ],
+ "category_id": 15,
+ "id": 2187,
+ "area": 529,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1096,
+ "bbox": [
+ 158,
+ 76,
+ 322,
+ 417
+ ],
+ "category_id": 10,
+ "id": 2188,
+ "area": 134274,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1097,
+ "bbox": [
+ 87,
+ 104,
+ 6,
+ 9
+ ],
+ "category_id": 4,
+ "id": 2189,
+ "area": 54,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1097,
+ "bbox": [
+ 58,
+ 161,
+ 8,
+ 9
+ ],
+ "category_id": 4,
+ "id": 2190,
+ "area": 72,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1097,
+ "bbox": [
+ 418,
+ 17,
+ 9,
+ 13
+ ],
+ "category_id": 4,
+ "id": 2191,
+ "area": 117,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1098,
+ "bbox": [
+ 185,
+ 227,
+ 198,
+ 205
+ ],
+ "category_id": 16,
+ "id": 2192,
+ "area": 40590,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1099,
+ "bbox": [
+ 208,
+ 69,
+ 32,
+ 21
+ ],
+ "category_id": 4,
+ "id": 2193,
+ "area": 672,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1099,
+ "bbox": [
+ 140,
+ 233,
+ 28,
+ 19
+ ],
+ "category_id": 4,
+ "id": 2194,
+ "area": 532,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1099,
+ "bbox": [
+ 414,
+ 197,
+ 33,
+ 22
+ ],
+ "category_id": 4,
+ "id": 2195,
+ "area": 726,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1099,
+ "bbox": [
+ 90,
+ 393,
+ 30,
+ 19
+ ],
+ "category_id": 4,
+ "id": 2196,
+ "area": 570,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1099,
+ "bbox": [
+ 277,
+ 427,
+ 34,
+ 19
+ ],
+ "category_id": 4,
+ "id": 2197,
+ "area": 646,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1100,
+ "bbox": [
+ 140,
+ 224,
+ 219,
+ 79
+ ],
+ "category_id": 19,
+ "id": 2198,
+ "area": 17301,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1101,
+ "bbox": [
+ 158,
+ 227,
+ 181,
+ 39
+ ],
+ "category_id": 19,
+ "id": 2199,
+ "area": 7059,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1102,
+ "bbox": [
+ 135,
+ 40,
+ 14,
+ 13
+ ],
+ "category_id": 4,
+ "id": 2200,
+ "area": 182,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1102,
+ "bbox": [
+ 76,
+ 198,
+ 16,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2201,
+ "area": 160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1102,
+ "bbox": [
+ 111,
+ 273,
+ 17,
+ 13
+ ],
+ "category_id": 4,
+ "id": 2202,
+ "area": 221,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1102,
+ "bbox": [
+ 42,
+ 389,
+ 14,
+ 9
+ ],
+ "category_id": 4,
+ "id": 2203,
+ "area": 126,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1102,
+ "bbox": [
+ 429,
+ 70,
+ 15,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2204,
+ "area": 150,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1102,
+ "bbox": [
+ 350,
+ 145,
+ 18,
+ 9
+ ],
+ "category_id": 4,
+ "id": 2205,
+ "area": 162,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1102,
+ "bbox": [
+ 263,
+ 227,
+ 15,
+ 9
+ ],
+ "category_id": 4,
+ "id": 2206,
+ "area": 135,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1102,
+ "bbox": [
+ 236,
+ 400,
+ 17,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2207,
+ "area": 170,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1102,
+ "bbox": [
+ 446,
+ 275,
+ 16,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2208,
+ "area": 160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1102,
+ "bbox": [
+ 364,
+ 334,
+ 16,
+ 9
+ ],
+ "category_id": 4,
+ "id": 2209,
+ "area": 144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1103,
+ "bbox": [
+ 212,
+ 149,
+ 49,
+ 167
+ ],
+ "category_id": 19,
+ "id": 2210,
+ "area": 8183,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1104,
+ "bbox": [
+ 81,
+ 149,
+ 260,
+ 243
+ ],
+ "category_id": 10,
+ "id": 2211,
+ "area": 63180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1105,
+ "bbox": [
+ 442,
+ 155,
+ 48,
+ 60
+ ],
+ "category_id": 4,
+ "id": 2212,
+ "area": 2880,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1105,
+ "bbox": [
+ 77,
+ 250,
+ 86,
+ 85
+ ],
+ "category_id": 4,
+ "id": 2213,
+ "area": 7310,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1105,
+ "bbox": [
+ 106,
+ 339,
+ 63,
+ 42
+ ],
+ "category_id": 4,
+ "id": 2214,
+ "area": 2646,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1106,
+ "bbox": [
+ 225,
+ 261,
+ 68,
+ 92
+ ],
+ "category_id": 4,
+ "id": 2215,
+ "area": 6256,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1107,
+ "bbox": [
+ 101,
+ 284,
+ 135,
+ 144
+ ],
+ "category_id": 15,
+ "id": 2216,
+ "area": 19440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1107,
+ "bbox": [
+ 247,
+ 147,
+ 140,
+ 138
+ ],
+ "category_id": 15,
+ "id": 2217,
+ "area": 19320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1108,
+ "bbox": [
+ 145,
+ 218,
+ 108,
+ 158
+ ],
+ "category_id": 19,
+ "id": 2218,
+ "area": 17064,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1109,
+ "bbox": [
+ 176,
+ 57,
+ 14,
+ 8
+ ],
+ "category_id": 4,
+ "id": 2219,
+ "area": 112,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1109,
+ "bbox": [
+ 65,
+ 49,
+ 13,
+ 9
+ ],
+ "category_id": 4,
+ "id": 2220,
+ "area": 117,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1109,
+ "bbox": [
+ 7,
+ 420,
+ 14,
+ 9
+ ],
+ "category_id": 4,
+ "id": 2221,
+ "area": 126,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1109,
+ "bbox": [
+ 120,
+ 440,
+ 14,
+ 7
+ ],
+ "category_id": 4,
+ "id": 2222,
+ "area": 98,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1110,
+ "bbox": [
+ 33,
+ 75,
+ 230,
+ 99
+ ],
+ "category_id": 10,
+ "id": 2223,
+ "area": 22770,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1111,
+ "bbox": [
+ 42,
+ 245,
+ 142,
+ 185
+ ],
+ "category_id": 15,
+ "id": 2224,
+ "area": 26270,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1111,
+ "bbox": [
+ 223,
+ 128,
+ 139,
+ 178
+ ],
+ "category_id": 15,
+ "id": 2225,
+ "area": 24742,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1112,
+ "bbox": [
+ 162,
+ 225,
+ 149,
+ 40
+ ],
+ "category_id": 19,
+ "id": 2226,
+ "area": 5960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1113,
+ "bbox": [
+ 158,
+ 205,
+ 185,
+ 53
+ ],
+ "category_id": 19,
+ "id": 2227,
+ "area": 9805,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1114,
+ "bbox": [
+ 180,
+ 186,
+ 177,
+ 174
+ ],
+ "category_id": 16,
+ "id": 2228,
+ "area": 30798,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1115,
+ "bbox": [
+ 151,
+ 387,
+ 45,
+ 65
+ ],
+ "category_id": 4,
+ "id": 2229,
+ "area": 2925,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1115,
+ "bbox": [
+ 335,
+ 38,
+ 44,
+ 60
+ ],
+ "category_id": 4,
+ "id": 2230,
+ "area": 2640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1116,
+ "bbox": [
+ 136,
+ 106,
+ 311,
+ 333
+ ],
+ "category_id": 10,
+ "id": 2231,
+ "area": 103563,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1117,
+ "bbox": [
+ 159,
+ 186,
+ 236,
+ 185
+ ],
+ "category_id": 10,
+ "id": 2232,
+ "area": 43660,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1118,
+ "bbox": [
+ 133,
+ 143,
+ 94,
+ 129
+ ],
+ "category_id": 15,
+ "id": 2233,
+ "area": 12126,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1118,
+ "bbox": [
+ 104,
+ 303,
+ 89,
+ 126
+ ],
+ "category_id": 15,
+ "id": 2234,
+ "area": 11214,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1119,
+ "bbox": [
+ 124,
+ 192,
+ 278,
+ 107
+ ],
+ "category_id": 19,
+ "id": 2235,
+ "area": 29746,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1120,
+ "bbox": [
+ 107,
+ 418,
+ 58,
+ 28
+ ],
+ "category_id": 16,
+ "id": 2236,
+ "area": 1624,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1121,
+ "bbox": [
+ 174,
+ 112,
+ 65,
+ 67
+ ],
+ "category_id": 4,
+ "id": 2237,
+ "area": 4355,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1121,
+ "bbox": [
+ 367,
+ 418,
+ 77,
+ 25
+ ],
+ "category_id": 4,
+ "id": 2238,
+ "area": 1925,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1122,
+ "bbox": [
+ 30,
+ 179,
+ 464,
+ 224
+ ],
+ "category_id": 10,
+ "id": 2239,
+ "area": 103936,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1123,
+ "bbox": [
+ 282,
+ 97,
+ 105,
+ 319
+ ],
+ "category_id": 10,
+ "id": 2240,
+ "area": 33495,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1124,
+ "bbox": [
+ 131,
+ 369,
+ 45,
+ 48
+ ],
+ "category_id": 4,
+ "id": 2241,
+ "area": 2160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1124,
+ "bbox": [
+ 339,
+ 141,
+ 48,
+ 57
+ ],
+ "category_id": 4,
+ "id": 2242,
+ "area": 2736,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1125,
+ "bbox": [
+ 406,
+ 51,
+ 23,
+ 25
+ ],
+ "category_id": 4,
+ "id": 2243,
+ "area": 575,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1125,
+ "bbox": [
+ 223,
+ 145,
+ 23,
+ 23
+ ],
+ "category_id": 4,
+ "id": 2244,
+ "area": 529,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1125,
+ "bbox": [
+ 285,
+ 250,
+ 27,
+ 22
+ ],
+ "category_id": 4,
+ "id": 2245,
+ "area": 594,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1125,
+ "bbox": [
+ 91,
+ 353,
+ 28,
+ 32
+ ],
+ "category_id": 4,
+ "id": 2246,
+ "area": 896,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1125,
+ "bbox": [
+ 44,
+ 442,
+ 15,
+ 29
+ ],
+ "category_id": 4,
+ "id": 2247,
+ "area": 435,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1126,
+ "bbox": [
+ 42,
+ 79,
+ 411,
+ 315
+ ],
+ "category_id": 10,
+ "id": 2248,
+ "area": 129465,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1127,
+ "bbox": [
+ 198,
+ 119,
+ 98,
+ 263
+ ],
+ "category_id": 19,
+ "id": 2249,
+ "area": 25774,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1128,
+ "bbox": [
+ 459,
+ 178,
+ 32,
+ 28
+ ],
+ "category_id": 15,
+ "id": 2250,
+ "area": 896,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1129,
+ "bbox": [
+ 62,
+ 74,
+ 36,
+ 54
+ ],
+ "category_id": 4,
+ "id": 2251,
+ "area": 1944,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1129,
+ "bbox": [
+ 106,
+ 401,
+ 38,
+ 50
+ ],
+ "category_id": 4,
+ "id": 2252,
+ "area": 1900,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1129,
+ "bbox": [
+ 410,
+ 446,
+ 50,
+ 57
+ ],
+ "category_id": 4,
+ "id": 2253,
+ "area": 2850,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1130,
+ "bbox": [
+ 30,
+ 209,
+ 37,
+ 51
+ ],
+ "category_id": 4,
+ "id": 2254,
+ "area": 1887,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1130,
+ "bbox": [
+ 256,
+ 310,
+ 37,
+ 44
+ ],
+ "category_id": 4,
+ "id": 2255,
+ "area": 1628,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1130,
+ "bbox": [
+ 405,
+ 183,
+ 59,
+ 41
+ ],
+ "category_id": 4,
+ "id": 2256,
+ "area": 2419,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1131,
+ "bbox": [
+ 97,
+ 131,
+ 276,
+ 250
+ ],
+ "category_id": 16,
+ "id": 2257,
+ "area": 69000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1132,
+ "bbox": [
+ 73,
+ 76,
+ 36,
+ 34
+ ],
+ "category_id": 4,
+ "id": 2258,
+ "area": 1224,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1132,
+ "bbox": [
+ 381,
+ 152,
+ 29,
+ 44
+ ],
+ "category_id": 4,
+ "id": 2259,
+ "area": 1276,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1132,
+ "bbox": [
+ 279,
+ 428,
+ 31,
+ 46
+ ],
+ "category_id": 4,
+ "id": 2260,
+ "area": 1426,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1133,
+ "bbox": [
+ 177,
+ 183,
+ 231,
+ 214
+ ],
+ "category_id": 10,
+ "id": 2261,
+ "area": 49434,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1134,
+ "bbox": [
+ 218,
+ 265,
+ 146,
+ 67
+ ],
+ "category_id": 10,
+ "id": 2262,
+ "area": 9782,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1135,
+ "bbox": [
+ 236,
+ 37,
+ 24,
+ 9
+ ],
+ "category_id": 4,
+ "id": 2263,
+ "area": 216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1135,
+ "bbox": [
+ 96,
+ 108,
+ 19,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2264,
+ "area": 190,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1135,
+ "bbox": [
+ 248,
+ 256,
+ 16,
+ 8
+ ],
+ "category_id": 4,
+ "id": 2265,
+ "area": 128,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1135,
+ "bbox": [
+ 97,
+ 478,
+ 21,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2266,
+ "area": 210,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1135,
+ "bbox": [
+ 263,
+ 445,
+ 22,
+ 11
+ ],
+ "category_id": 4,
+ "id": 2267,
+ "area": 242,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1135,
+ "bbox": [
+ 439,
+ 382,
+ 21,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2268,
+ "area": 210,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1136,
+ "bbox": [
+ 64,
+ 145,
+ 157,
+ 113
+ ],
+ "category_id": 15,
+ "id": 2269,
+ "area": 17741,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1136,
+ "bbox": [
+ 293,
+ 332,
+ 133,
+ 93
+ ],
+ "category_id": 15,
+ "id": 2270,
+ "area": 12369,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1136,
+ "bbox": [
+ 0,
+ 464,
+ 50,
+ 19
+ ],
+ "category_id": 15,
+ "id": 2271,
+ "area": 950,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1137,
+ "bbox": [
+ 53,
+ 126,
+ 370,
+ 288
+ ],
+ "category_id": 15,
+ "id": 2272,
+ "area": 106560,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1138,
+ "bbox": [
+ 156,
+ 102,
+ 37,
+ 52
+ ],
+ "category_id": 4,
+ "id": 2273,
+ "area": 1924,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1138,
+ "bbox": [
+ 414,
+ 253,
+ 35,
+ 45
+ ],
+ "category_id": 4,
+ "id": 2274,
+ "area": 1575,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1138,
+ "bbox": [
+ 222,
+ 464,
+ 40,
+ 42
+ ],
+ "category_id": 4,
+ "id": 2275,
+ "area": 1680,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1139,
+ "bbox": [
+ 194,
+ 172,
+ 61,
+ 172
+ ],
+ "category_id": 19,
+ "id": 2276,
+ "area": 10492,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1140,
+ "bbox": [
+ 169,
+ 203,
+ 92,
+ 111
+ ],
+ "category_id": 16,
+ "id": 2277,
+ "area": 10212,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1141,
+ "bbox": [
+ 26,
+ 298,
+ 143,
+ 123
+ ],
+ "category_id": 15,
+ "id": 2278,
+ "area": 17589,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1141,
+ "bbox": [
+ 203,
+ 168,
+ 149,
+ 105
+ ],
+ "category_id": 15,
+ "id": 2279,
+ "area": 15645,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1142,
+ "bbox": [
+ 291,
+ 240,
+ 101,
+ 108
+ ],
+ "category_id": 15,
+ "id": 2280,
+ "area": 10908,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1142,
+ "bbox": [
+ 467,
+ 389,
+ 45,
+ 60
+ ],
+ "category_id": 15,
+ "id": 2281,
+ "area": 2700,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1143,
+ "bbox": [
+ 80,
+ 130,
+ 26,
+ 30
+ ],
+ "category_id": 16,
+ "id": 2282,
+ "area": 780,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1144,
+ "bbox": [
+ 31,
+ 212,
+ 70,
+ 54
+ ],
+ "category_id": 4,
+ "id": 2283,
+ "area": 3780,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1144,
+ "bbox": [
+ 438,
+ 287,
+ 62,
+ 52
+ ],
+ "category_id": 4,
+ "id": 2284,
+ "area": 3224,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1145,
+ "bbox": [
+ 103,
+ 104,
+ 250,
+ 299
+ ],
+ "category_id": 16,
+ "id": 2285,
+ "area": 74750,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1146,
+ "bbox": [
+ 169,
+ 405,
+ 55,
+ 35
+ ],
+ "category_id": 4,
+ "id": 2286,
+ "area": 1925,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1146,
+ "bbox": [
+ 359,
+ 106,
+ 52,
+ 28
+ ],
+ "category_id": 4,
+ "id": 2287,
+ "area": 1456,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1147,
+ "bbox": [
+ 144,
+ 76,
+ 198,
+ 169
+ ],
+ "category_id": 15,
+ "id": 2288,
+ "area": 33462,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1148,
+ "bbox": [
+ 90,
+ 48,
+ 32,
+ 25
+ ],
+ "category_id": 4,
+ "id": 2289,
+ "area": 800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1148,
+ "bbox": [
+ 114,
+ 401,
+ 54,
+ 39
+ ],
+ "category_id": 4,
+ "id": 2290,
+ "area": 2106,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1148,
+ "bbox": [
+ 417,
+ 261,
+ 47,
+ 30
+ ],
+ "category_id": 4,
+ "id": 2291,
+ "area": 1410,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1149,
+ "bbox": [
+ 53,
+ 117,
+ 10,
+ 25
+ ],
+ "category_id": 15,
+ "id": 2292,
+ "area": 250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1149,
+ "bbox": [
+ 18,
+ 110,
+ 13,
+ 39
+ ],
+ "category_id": 15,
+ "id": 2293,
+ "area": 507,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1149,
+ "bbox": [
+ 82,
+ 86,
+ 13,
+ 17
+ ],
+ "category_id": 15,
+ "id": 2294,
+ "area": 221,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1149,
+ "bbox": [
+ 86,
+ 73,
+ 10,
+ 14
+ ],
+ "category_id": 15,
+ "id": 2295,
+ "area": 140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1150,
+ "bbox": [
+ 14,
+ 119,
+ 492,
+ 17
+ ],
+ "category_id": 10,
+ "id": 2296,
+ "area": 8364,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1151,
+ "bbox": [
+ 143,
+ 179,
+ 80,
+ 99
+ ],
+ "category_id": 15,
+ "id": 2297,
+ "area": 7920,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1151,
+ "bbox": [
+ 257,
+ 192,
+ 108,
+ 129
+ ],
+ "category_id": 15,
+ "id": 2298,
+ "area": 13932,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1151,
+ "bbox": [
+ 346,
+ 307,
+ 105,
+ 128
+ ],
+ "category_id": 15,
+ "id": 2299,
+ "area": 13440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1151,
+ "bbox": [
+ 304,
+ 143,
+ 55,
+ 118
+ ],
+ "category_id": 15,
+ "id": 2300,
+ "area": 6490,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1152,
+ "bbox": [
+ 278,
+ 23,
+ 19,
+ 11
+ ],
+ "category_id": 4,
+ "id": 2301,
+ "area": 209,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1152,
+ "bbox": [
+ 157,
+ 110,
+ 18,
+ 12
+ ],
+ "category_id": 4,
+ "id": 2302,
+ "area": 216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1152,
+ "bbox": [
+ 223,
+ 273,
+ 17,
+ 11
+ ],
+ "category_id": 4,
+ "id": 2303,
+ "area": 187,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1152,
+ "bbox": [
+ 429,
+ 162,
+ 18,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2304,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1152,
+ "bbox": [
+ 417,
+ 370,
+ 22,
+ 12
+ ],
+ "category_id": 4,
+ "id": 2305,
+ "area": 264,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1152,
+ "bbox": [
+ 268,
+ 445,
+ 23,
+ 12
+ ],
+ "category_id": 4,
+ "id": 2306,
+ "area": 276,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1153,
+ "bbox": [
+ 88,
+ 81,
+ 333,
+ 166
+ ],
+ "category_id": 10,
+ "id": 2307,
+ "area": 55278,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1154,
+ "bbox": [
+ 268,
+ 88,
+ 43,
+ 49
+ ],
+ "category_id": 4,
+ "id": 2308,
+ "area": 2107,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1154,
+ "bbox": [
+ 289,
+ 417,
+ 48,
+ 47
+ ],
+ "category_id": 4,
+ "id": 2309,
+ "area": 2256,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1155,
+ "bbox": [
+ 112,
+ 88,
+ 245,
+ 224
+ ],
+ "category_id": 16,
+ "id": 2310,
+ "area": 54880,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1156,
+ "bbox": [
+ 95,
+ 303,
+ 170,
+ 100
+ ],
+ "category_id": 16,
+ "id": 2311,
+ "area": 17000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1157,
+ "bbox": [
+ 106,
+ 56,
+ 19,
+ 12
+ ],
+ "category_id": 4,
+ "id": 2312,
+ "area": 228,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1157,
+ "bbox": [
+ 76,
+ 183,
+ 27,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2313,
+ "area": 270,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1157,
+ "bbox": [
+ 30,
+ 324,
+ 30,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2314,
+ "area": 300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1157,
+ "bbox": [
+ 167,
+ 372,
+ 23,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2315,
+ "area": 230,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1157,
+ "bbox": [
+ 339,
+ 216,
+ 23,
+ 13
+ ],
+ "category_id": 4,
+ "id": 2316,
+ "area": 299,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1157,
+ "bbox": [
+ 422,
+ 305,
+ 27,
+ 13
+ ],
+ "category_id": 4,
+ "id": 2317,
+ "area": 351,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1157,
+ "bbox": [
+ 407,
+ 432,
+ 23,
+ 12
+ ],
+ "category_id": 4,
+ "id": 2318,
+ "area": 276,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1157,
+ "bbox": [
+ 469,
+ 113,
+ 26,
+ 13
+ ],
+ "category_id": 4,
+ "id": 2319,
+ "area": 338,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1158,
+ "bbox": [
+ 212,
+ 165,
+ 70,
+ 217
+ ],
+ "category_id": 19,
+ "id": 2320,
+ "area": 15190,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1159,
+ "bbox": [
+ 376,
+ 343,
+ 52,
+ 46
+ ],
+ "category_id": 16,
+ "id": 2321,
+ "area": 2392,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1160,
+ "bbox": [
+ 209,
+ 222,
+ 93,
+ 57
+ ],
+ "category_id": 4,
+ "id": 2322,
+ "area": 5301,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1161,
+ "bbox": [
+ 16,
+ 83,
+ 27,
+ 77
+ ],
+ "category_id": 4,
+ "id": 2323,
+ "area": 2079,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1161,
+ "bbox": [
+ 34,
+ 421,
+ 44,
+ 61
+ ],
+ "category_id": 4,
+ "id": 2324,
+ "area": 2684,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1161,
+ "bbox": [
+ 471,
+ 325,
+ 28,
+ 81
+ ],
+ "category_id": 4,
+ "id": 2325,
+ "area": 2268,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1162,
+ "bbox": [
+ 324,
+ 273,
+ 56,
+ 90
+ ],
+ "category_id": 4,
+ "id": 2326,
+ "area": 5040,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1163,
+ "bbox": [
+ 49,
+ 100,
+ 374,
+ 323
+ ],
+ "category_id": 10,
+ "id": 2327,
+ "area": 120802,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1164,
+ "bbox": [
+ 148,
+ 156,
+ 137,
+ 141
+ ],
+ "category_id": 15,
+ "id": 2328,
+ "area": 19317,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1165,
+ "bbox": [
+ 37,
+ 271,
+ 21,
+ 21
+ ],
+ "category_id": 4,
+ "id": 2329,
+ "area": 441,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1165,
+ "bbox": [
+ 171,
+ 208,
+ 22,
+ 23
+ ],
+ "category_id": 4,
+ "id": 2330,
+ "area": 506,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1165,
+ "bbox": [
+ 110,
+ 70,
+ 21,
+ 21
+ ],
+ "category_id": 4,
+ "id": 2331,
+ "area": 441,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1165,
+ "bbox": [
+ 263,
+ 104,
+ 22,
+ 27
+ ],
+ "category_id": 4,
+ "id": 2332,
+ "area": 594,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1165,
+ "bbox": [
+ 397,
+ 179,
+ 25,
+ 21
+ ],
+ "category_id": 4,
+ "id": 2333,
+ "area": 525,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1165,
+ "bbox": [
+ 446,
+ 357,
+ 18,
+ 24
+ ],
+ "category_id": 4,
+ "id": 2334,
+ "area": 432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1166,
+ "bbox": [
+ 169,
+ 188,
+ 196,
+ 192
+ ],
+ "category_id": 16,
+ "id": 2335,
+ "area": 37632,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1167,
+ "bbox": [
+ 306,
+ 297,
+ 140,
+ 97
+ ],
+ "category_id": 4,
+ "id": 2336,
+ "area": 13580,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1168,
+ "bbox": [
+ 127,
+ 207,
+ 140,
+ 183
+ ],
+ "category_id": 15,
+ "id": 2337,
+ "area": 25620,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1168,
+ "bbox": [
+ 329,
+ 154,
+ 138,
+ 189
+ ],
+ "category_id": 15,
+ "id": 2338,
+ "area": 26082,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1169,
+ "bbox": [
+ 139,
+ 227,
+ 193,
+ 102
+ ],
+ "category_id": 19,
+ "id": 2339,
+ "area": 19686,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1170,
+ "bbox": [
+ 222,
+ 212,
+ 68,
+ 88
+ ],
+ "category_id": 15,
+ "id": 2340,
+ "area": 5984,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1170,
+ "bbox": [
+ 385,
+ 305,
+ 33,
+ 75
+ ],
+ "category_id": 15,
+ "id": 2341,
+ "area": 2475,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1171,
+ "bbox": [
+ 269,
+ 154,
+ 205,
+ 150
+ ],
+ "category_id": 19,
+ "id": 2342,
+ "area": 30750,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1172,
+ "bbox": [
+ 114,
+ 98,
+ 315,
+ 266
+ ],
+ "category_id": 16,
+ "id": 2343,
+ "area": 83790,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1173,
+ "bbox": [
+ 207,
+ 159,
+ 161,
+ 178
+ ],
+ "category_id": 10,
+ "id": 2344,
+ "area": 28658,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1174,
+ "bbox": [
+ 55,
+ 145,
+ 329,
+ 87
+ ],
+ "category_id": 16,
+ "id": 2345,
+ "area": 28623,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1175,
+ "bbox": [
+ 414,
+ 357,
+ 75,
+ 120
+ ],
+ "category_id": 10,
+ "id": 2346,
+ "area": 9000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1176,
+ "bbox": [
+ 212,
+ 193,
+ 13,
+ 12
+ ],
+ "category_id": 4,
+ "id": 2347,
+ "area": 156,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1176,
+ "bbox": [
+ 238,
+ 259,
+ 9,
+ 11
+ ],
+ "category_id": 4,
+ "id": 2348,
+ "area": 99,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1176,
+ "bbox": [
+ 391,
+ 331,
+ 9,
+ 9
+ ],
+ "category_id": 4,
+ "id": 2349,
+ "area": 81,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1176,
+ "bbox": [
+ 369,
+ 168,
+ 12,
+ 8
+ ],
+ "category_id": 4,
+ "id": 2350,
+ "area": 96,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1176,
+ "bbox": [
+ 418,
+ 261,
+ 5,
+ 7
+ ],
+ "category_id": 4,
+ "id": 2351,
+ "area": 35,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1176,
+ "bbox": [
+ 319,
+ 293,
+ 6,
+ 8
+ ],
+ "category_id": 4,
+ "id": 2352,
+ "area": 48,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1177,
+ "bbox": [
+ 33,
+ 184,
+ 196,
+ 209
+ ],
+ "category_id": 15,
+ "id": 2353,
+ "area": 40964,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1177,
+ "bbox": [
+ 241,
+ 133,
+ 198,
+ 206
+ ],
+ "category_id": 15,
+ "id": 2354,
+ "area": 40788,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1177,
+ "bbox": [
+ 77,
+ 455,
+ 61,
+ 57
+ ],
+ "category_id": 15,
+ "id": 2355,
+ "area": 3477,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1178,
+ "bbox": [
+ 101,
+ 112,
+ 86,
+ 48
+ ],
+ "category_id": 16,
+ "id": 2356,
+ "area": 4128,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1179,
+ "bbox": [
+ 153,
+ 263,
+ 56,
+ 116
+ ],
+ "category_id": 19,
+ "id": 2357,
+ "area": 6496,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1180,
+ "bbox": [
+ 222,
+ 197,
+ 131,
+ 160
+ ],
+ "category_id": 16,
+ "id": 2358,
+ "area": 20960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1181,
+ "bbox": [
+ 168,
+ 92,
+ 30,
+ 28
+ ],
+ "category_id": 4,
+ "id": 2359,
+ "area": 840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1181,
+ "bbox": [
+ 402,
+ 87,
+ 26,
+ 33
+ ],
+ "category_id": 4,
+ "id": 2360,
+ "area": 858,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1181,
+ "bbox": [
+ 422,
+ 206,
+ 24,
+ 34
+ ],
+ "category_id": 4,
+ "id": 2361,
+ "area": 816,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1181,
+ "bbox": [
+ 261,
+ 403,
+ 25,
+ 29
+ ],
+ "category_id": 4,
+ "id": 2362,
+ "area": 725,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1181,
+ "bbox": [
+ 119,
+ 453,
+ 25,
+ 28
+ ],
+ "category_id": 4,
+ "id": 2363,
+ "area": 700,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1182,
+ "bbox": [
+ 33,
+ 81,
+ 226,
+ 347
+ ],
+ "category_id": 10,
+ "id": 2364,
+ "area": 78422,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1183,
+ "bbox": [
+ 204,
+ 121,
+ 222,
+ 268
+ ],
+ "category_id": 10,
+ "id": 2365,
+ "area": 59496,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1184,
+ "bbox": [
+ 119,
+ 200,
+ 240,
+ 128
+ ],
+ "category_id": 16,
+ "id": 2366,
+ "area": 30720,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1185,
+ "bbox": [
+ 215,
+ 123,
+ 88,
+ 249
+ ],
+ "category_id": 19,
+ "id": 2367,
+ "area": 21912,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1186,
+ "bbox": [
+ 194,
+ 218,
+ 134,
+ 45
+ ],
+ "category_id": 19,
+ "id": 2368,
+ "area": 6030,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1187,
+ "bbox": [
+ 83,
+ 195,
+ 47,
+ 41
+ ],
+ "category_id": 4,
+ "id": 2369,
+ "area": 1927,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1187,
+ "bbox": [
+ 375,
+ 273,
+ 47,
+ 37
+ ],
+ "category_id": 4,
+ "id": 2370,
+ "area": 1739,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1188,
+ "bbox": [
+ 58,
+ 177,
+ 67,
+ 73
+ ],
+ "category_id": 19,
+ "id": 2371,
+ "area": 4891,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1189,
+ "bbox": [
+ 398,
+ 61,
+ 23,
+ 15
+ ],
+ "category_id": 4,
+ "id": 2372,
+ "area": 345,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1189,
+ "bbox": [
+ 248,
+ 119,
+ 21,
+ 17
+ ],
+ "category_id": 4,
+ "id": 2373,
+ "area": 357,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1189,
+ "bbox": [
+ 91,
+ 167,
+ 21,
+ 14
+ ],
+ "category_id": 4,
+ "id": 2374,
+ "area": 294,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1189,
+ "bbox": [
+ 245,
+ 300,
+ 27,
+ 16
+ ],
+ "category_id": 4,
+ "id": 2375,
+ "area": 432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1189,
+ "bbox": [
+ 416,
+ 253,
+ 19,
+ 13
+ ],
+ "category_id": 4,
+ "id": 2376,
+ "area": 247,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1189,
+ "bbox": [
+ 82,
+ 356,
+ 28,
+ 15
+ ],
+ "category_id": 4,
+ "id": 2377,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1189,
+ "bbox": [
+ 273,
+ 483,
+ 26,
+ 18
+ ],
+ "category_id": 4,
+ "id": 2378,
+ "area": 468,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1190,
+ "bbox": [
+ 236,
+ 307,
+ 103,
+ 101
+ ],
+ "category_id": 15,
+ "id": 2379,
+ "area": 10403,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1190,
+ "bbox": [
+ 375,
+ 333,
+ 101,
+ 108
+ ],
+ "category_id": 15,
+ "id": 2380,
+ "area": 10908,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1191,
+ "bbox": [
+ 193,
+ 78,
+ 32,
+ 44
+ ],
+ "category_id": 4,
+ "id": 2381,
+ "area": 1408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1191,
+ "bbox": [
+ 272,
+ 408,
+ 30,
+ 46
+ ],
+ "category_id": 4,
+ "id": 2382,
+ "area": 1380,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1192,
+ "bbox": [
+ 66,
+ 16,
+ 30,
+ 12
+ ],
+ "category_id": 4,
+ "id": 2383,
+ "area": 360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1192,
+ "bbox": [
+ 6,
+ 99,
+ 31,
+ 14
+ ],
+ "category_id": 4,
+ "id": 2384,
+ "area": 434,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1192,
+ "bbox": [
+ 73,
+ 366,
+ 33,
+ 17
+ ],
+ "category_id": 4,
+ "id": 2385,
+ "area": 561,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1192,
+ "bbox": [
+ 288,
+ 453,
+ 33,
+ 13
+ ],
+ "category_id": 4,
+ "id": 2386,
+ "area": 429,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1192,
+ "bbox": [
+ 474,
+ 483,
+ 35,
+ 14
+ ],
+ "category_id": 4,
+ "id": 2387,
+ "area": 490,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1193,
+ "bbox": [
+ 165,
+ 183,
+ 293,
+ 178
+ ],
+ "category_id": 10,
+ "id": 2388,
+ "area": 52154,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1194,
+ "bbox": [
+ 228,
+ 71,
+ 6,
+ 9
+ ],
+ "category_id": 4,
+ "id": 2389,
+ "area": 54,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1194,
+ "bbox": [
+ 239,
+ 128,
+ 8,
+ 8
+ ],
+ "category_id": 4,
+ "id": 2390,
+ "area": 64,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1194,
+ "bbox": [
+ 278,
+ 164,
+ 7,
+ 5
+ ],
+ "category_id": 4,
+ "id": 2391,
+ "area": 35,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1194,
+ "bbox": [
+ 210,
+ 181,
+ 6,
+ 5
+ ],
+ "category_id": 4,
+ "id": 2392,
+ "area": 30,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1194,
+ "bbox": [
+ 289,
+ 244,
+ 9,
+ 6
+ ],
+ "category_id": 4,
+ "id": 2393,
+ "area": 54,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1194,
+ "bbox": [
+ 202,
+ 250,
+ 12,
+ 7
+ ],
+ "category_id": 4,
+ "id": 2394,
+ "area": 84,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1194,
+ "bbox": [
+ 202,
+ 312,
+ 9,
+ 7
+ ],
+ "category_id": 4,
+ "id": 2395,
+ "area": 63,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1194,
+ "bbox": [
+ 203,
+ 354,
+ 8,
+ 6
+ ],
+ "category_id": 4,
+ "id": 2396,
+ "area": 48,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1194,
+ "bbox": [
+ 169,
+ 429,
+ 9,
+ 6
+ ],
+ "category_id": 4,
+ "id": 2397,
+ "area": 54,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1194,
+ "bbox": [
+ 109,
+ 435,
+ 7,
+ 5
+ ],
+ "category_id": 4,
+ "id": 2398,
+ "area": 35,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1195,
+ "bbox": [
+ 37,
+ 179,
+ 70,
+ 165
+ ],
+ "category_id": 10,
+ "id": 2399,
+ "area": 11550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1196,
+ "bbox": [
+ 250,
+ 396,
+ 43,
+ 36
+ ],
+ "category_id": 4,
+ "id": 2400,
+ "area": 1548,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1196,
+ "bbox": [
+ 243,
+ 84,
+ 42,
+ 32
+ ],
+ "category_id": 4,
+ "id": 2401,
+ "area": 1344,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1197,
+ "bbox": [
+ 81,
+ 205,
+ 329,
+ 117
+ ],
+ "category_id": 19,
+ "id": 2402,
+ "area": 38493,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1198,
+ "bbox": [
+ 131,
+ 128,
+ 269,
+ 124
+ ],
+ "category_id": 16,
+ "id": 2403,
+ "area": 33356,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1199,
+ "bbox": [
+ 344,
+ 68,
+ 37,
+ 32
+ ],
+ "category_id": 4,
+ "id": 2404,
+ "area": 1184,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1199,
+ "bbox": [
+ 40,
+ 353,
+ 42,
+ 32
+ ],
+ "category_id": 4,
+ "id": 2405,
+ "area": 1344,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1199,
+ "bbox": [
+ 212,
+ 353,
+ 41,
+ 33
+ ],
+ "category_id": 4,
+ "id": 2406,
+ "area": 1353,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1200,
+ "bbox": [
+ 131,
+ 51,
+ 162,
+ 165
+ ],
+ "category_id": 10,
+ "id": 2407,
+ "area": 26730,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1201,
+ "bbox": [
+ 236,
+ 116,
+ 73,
+ 232
+ ],
+ "category_id": 16,
+ "id": 2408,
+ "area": 16936,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1202,
+ "bbox": [
+ 112,
+ 110,
+ 324,
+ 345
+ ],
+ "category_id": 10,
+ "id": 2409,
+ "area": 111780,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1203,
+ "bbox": [
+ 239,
+ 73,
+ 45,
+ 170
+ ],
+ "category_id": 10,
+ "id": 2410,
+ "area": 7650,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1204,
+ "bbox": [
+ 389,
+ 83,
+ 28,
+ 29
+ ],
+ "category_id": 4,
+ "id": 2411,
+ "area": 812,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1204,
+ "bbox": [
+ 268,
+ 363,
+ 29,
+ 36
+ ],
+ "category_id": 4,
+ "id": 2412,
+ "area": 1044,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1205,
+ "bbox": [
+ 205,
+ 257,
+ 145,
+ 25
+ ],
+ "category_id": 19,
+ "id": 2413,
+ "area": 3625,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1206,
+ "bbox": [
+ 55,
+ 140,
+ 237,
+ 212
+ ],
+ "category_id": 10,
+ "id": 2414,
+ "area": 50244,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1207,
+ "bbox": [
+ 55,
+ 261,
+ 79,
+ 39
+ ],
+ "category_id": 4,
+ "id": 2415,
+ "area": 3081,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1207,
+ "bbox": [
+ 392,
+ 154,
+ 80,
+ 57
+ ],
+ "category_id": 4,
+ "id": 2416,
+ "area": 4560,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1208,
+ "bbox": [
+ 338,
+ 282,
+ 97,
+ 107
+ ],
+ "category_id": 15,
+ "id": 2417,
+ "area": 10379,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1209,
+ "bbox": [
+ 155,
+ 115,
+ 100,
+ 221
+ ],
+ "category_id": 19,
+ "id": 2418,
+ "area": 22100,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1210,
+ "bbox": [
+ 203,
+ 126,
+ 69,
+ 247
+ ],
+ "category_id": 19,
+ "id": 2419,
+ "area": 17043,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1211,
+ "bbox": [
+ 176,
+ 77,
+ 95,
+ 318
+ ],
+ "category_id": 19,
+ "id": 2420,
+ "area": 30210,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1212,
+ "bbox": [
+ 264,
+ 160,
+ 75,
+ 21
+ ],
+ "category_id": 16,
+ "id": 2421,
+ "area": 1575,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1213,
+ "bbox": [
+ 203,
+ 88,
+ 205,
+ 244
+ ],
+ "category_id": 16,
+ "id": 2422,
+ "area": 50020,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1214,
+ "bbox": [
+ 60,
+ 155,
+ 11,
+ 17
+ ],
+ "category_id": 4,
+ "id": 2423,
+ "area": 187,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1214,
+ "bbox": [
+ 466,
+ 323,
+ 10,
+ 16
+ ],
+ "category_id": 4,
+ "id": 2424,
+ "area": 160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1214,
+ "bbox": [
+ 233,
+ 385,
+ 10,
+ 16
+ ],
+ "category_id": 4,
+ "id": 2425,
+ "area": 160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1214,
+ "bbox": [
+ 55,
+ 446,
+ 10,
+ 16
+ ],
+ "category_id": 4,
+ "id": 2426,
+ "area": 160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1214,
+ "bbox": [
+ 493,
+ 497,
+ 11,
+ 14
+ ],
+ "category_id": 4,
+ "id": 2427,
+ "area": 154,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1215,
+ "bbox": [
+ 303,
+ 236,
+ 120,
+ 105
+ ],
+ "category_id": 15,
+ "id": 2428,
+ "area": 12600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1216,
+ "bbox": [
+ 220,
+ 168,
+ 100,
+ 88
+ ],
+ "category_id": 15,
+ "id": 2429,
+ "area": 8800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1216,
+ "bbox": [
+ 342,
+ 209,
+ 98,
+ 87
+ ],
+ "category_id": 15,
+ "id": 2430,
+ "area": 8526,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1216,
+ "bbox": [
+ 167,
+ 312,
+ 160,
+ 148
+ ],
+ "category_id": 15,
+ "id": 2431,
+ "area": 23680,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1217,
+ "bbox": [
+ 39,
+ 303,
+ 105,
+ 114
+ ],
+ "category_id": 16,
+ "id": 2432,
+ "area": 11970,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1217,
+ "bbox": [
+ 280,
+ 308,
+ 68,
+ 74
+ ],
+ "category_id": 16,
+ "id": 2433,
+ "area": 5032,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1218,
+ "bbox": [
+ 200,
+ 72,
+ 285,
+ 285
+ ],
+ "category_id": 16,
+ "id": 2434,
+ "area": 81225,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1219,
+ "bbox": [
+ 205,
+ 122,
+ 53,
+ 57
+ ],
+ "category_id": 4,
+ "id": 2435,
+ "area": 3021,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1219,
+ "bbox": [
+ 204,
+ 391,
+ 55,
+ 58
+ ],
+ "category_id": 4,
+ "id": 2436,
+ "area": 3190,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1220,
+ "bbox": [
+ 195,
+ 165,
+ 141,
+ 147
+ ],
+ "category_id": 15,
+ "id": 2437,
+ "area": 20727,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1221,
+ "bbox": [
+ 158,
+ 229,
+ 195,
+ 75
+ ],
+ "category_id": 19,
+ "id": 2438,
+ "area": 14625,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1222,
+ "bbox": [
+ 217,
+ 224,
+ 113,
+ 88
+ ],
+ "category_id": 16,
+ "id": 2439,
+ "area": 9944,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1223,
+ "bbox": [
+ 216,
+ 340,
+ 84,
+ 60
+ ],
+ "category_id": 15,
+ "id": 2440,
+ "area": 5040,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1223,
+ "bbox": [
+ 390,
+ 246,
+ 79,
+ 55
+ ],
+ "category_id": 15,
+ "id": 2441,
+ "area": 4345,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1223,
+ "bbox": [
+ 320,
+ 205,
+ 74,
+ 52
+ ],
+ "category_id": 15,
+ "id": 2442,
+ "area": 3848,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1223,
+ "bbox": [
+ 245,
+ 182,
+ 75,
+ 55
+ ],
+ "category_id": 15,
+ "id": 2443,
+ "area": 4125,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1223,
+ "bbox": [
+ 171,
+ 162,
+ 72,
+ 54
+ ],
+ "category_id": 15,
+ "id": 2444,
+ "area": 3888,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1223,
+ "bbox": [
+ 96,
+ 142,
+ 73,
+ 53
+ ],
+ "category_id": 15,
+ "id": 2445,
+ "area": 3869,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1223,
+ "bbox": [
+ 128,
+ 83,
+ 65,
+ 27
+ ],
+ "category_id": 15,
+ "id": 2446,
+ "area": 1755,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1223,
+ "bbox": [
+ 217,
+ 110,
+ 72,
+ 26
+ ],
+ "category_id": 15,
+ "id": 2447,
+ "area": 1872,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1224,
+ "bbox": [
+ 184,
+ 115,
+ 69,
+ 274
+ ],
+ "category_id": 19,
+ "id": 2448,
+ "area": 18906,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1225,
+ "bbox": [
+ 372,
+ 160,
+ 90,
+ 112
+ ],
+ "category_id": 15,
+ "id": 2449,
+ "area": 10080,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1225,
+ "bbox": [
+ 290,
+ 254,
+ 94,
+ 110
+ ],
+ "category_id": 15,
+ "id": 2450,
+ "area": 10340,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1225,
+ "bbox": [
+ 142,
+ 239,
+ 84,
+ 106
+ ],
+ "category_id": 15,
+ "id": 2451,
+ "area": 8904,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1225,
+ "bbox": [
+ 107,
+ 170,
+ 58,
+ 82
+ ],
+ "category_id": 15,
+ "id": 2452,
+ "area": 4756,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1225,
+ "bbox": [
+ 87,
+ 89,
+ 59,
+ 81
+ ],
+ "category_id": 15,
+ "id": 2453,
+ "area": 4779,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1226,
+ "bbox": [
+ 247,
+ 256,
+ 85,
+ 84
+ ],
+ "category_id": 4,
+ "id": 2454,
+ "area": 7140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1227,
+ "bbox": [
+ 124,
+ 144,
+ 292,
+ 159
+ ],
+ "category_id": 16,
+ "id": 2455,
+ "area": 46428,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1228,
+ "bbox": [
+ 184,
+ 86,
+ 47,
+ 24
+ ],
+ "category_id": 4,
+ "id": 2456,
+ "area": 1128,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1228,
+ "bbox": [
+ 240,
+ 397,
+ 69,
+ 17
+ ],
+ "category_id": 4,
+ "id": 2457,
+ "area": 1173,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1229,
+ "bbox": [
+ 71,
+ 262,
+ 27,
+ 23
+ ],
+ "category_id": 4,
+ "id": 2458,
+ "area": 621,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1229,
+ "bbox": [
+ 298,
+ 250,
+ 20,
+ 20
+ ],
+ "category_id": 4,
+ "id": 2459,
+ "area": 400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1229,
+ "bbox": [
+ 197,
+ 420,
+ 21,
+ 16
+ ],
+ "category_id": 4,
+ "id": 2460,
+ "area": 336,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1229,
+ "bbox": [
+ 446,
+ 196,
+ 32,
+ 16
+ ],
+ "category_id": 4,
+ "id": 2461,
+ "area": 512,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1230,
+ "bbox": [
+ 106,
+ 81,
+ 378,
+ 202
+ ],
+ "category_id": 10,
+ "id": 2462,
+ "area": 76356,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1231,
+ "bbox": [
+ 70,
+ 247,
+ 248,
+ 78
+ ],
+ "category_id": 19,
+ "id": 2463,
+ "area": 19344,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1232,
+ "bbox": [
+ 56,
+ 232,
+ 221,
+ 64
+ ],
+ "category_id": 19,
+ "id": 2464,
+ "area": 14144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1233,
+ "bbox": [
+ 138,
+ 380,
+ 68,
+ 33
+ ],
+ "category_id": 16,
+ "id": 2465,
+ "area": 2244,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1234,
+ "bbox": [
+ 147,
+ 176,
+ 108,
+ 91
+ ],
+ "category_id": 15,
+ "id": 2466,
+ "area": 9828,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1234,
+ "bbox": [
+ 252,
+ 157,
+ 109,
+ 93
+ ],
+ "category_id": 15,
+ "id": 2467,
+ "area": 10137,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1235,
+ "bbox": [
+ 102,
+ 67,
+ 281,
+ 407
+ ],
+ "category_id": 15,
+ "id": 2468,
+ "area": 114367,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1236,
+ "bbox": [
+ 220,
+ 167,
+ 151,
+ 163
+ ],
+ "category_id": 15,
+ "id": 2469,
+ "area": 24613,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1237,
+ "bbox": [
+ 133,
+ 65,
+ 105,
+ 337
+ ],
+ "category_id": 16,
+ "id": 2470,
+ "area": 35385,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1238,
+ "bbox": [
+ 50,
+ 236,
+ 104,
+ 89
+ ],
+ "category_id": 15,
+ "id": 2471,
+ "area": 9256,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1238,
+ "bbox": [
+ 165,
+ 266,
+ 135,
+ 102
+ ],
+ "category_id": 15,
+ "id": 2472,
+ "area": 13770,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1239,
+ "bbox": [
+ 168,
+ 193,
+ 134,
+ 68
+ ],
+ "category_id": 19,
+ "id": 2473,
+ "area": 9112,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1240,
+ "bbox": [
+ 261,
+ 128,
+ 146,
+ 145
+ ],
+ "category_id": 15,
+ "id": 2474,
+ "area": 21170,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1241,
+ "bbox": [
+ 24,
+ 18,
+ 267,
+ 124
+ ],
+ "category_id": 10,
+ "id": 2475,
+ "area": 33108,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1242,
+ "bbox": [
+ 205,
+ 190,
+ 202,
+ 182
+ ],
+ "category_id": 16,
+ "id": 2476,
+ "area": 36764,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1243,
+ "bbox": [
+ 127,
+ 44,
+ 212,
+ 412
+ ],
+ "category_id": 16,
+ "id": 2477,
+ "area": 87344,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1244,
+ "bbox": [
+ 239,
+ 133,
+ 142,
+ 98
+ ],
+ "category_id": 15,
+ "id": 2478,
+ "area": 13916,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1244,
+ "bbox": [
+ 193,
+ 271,
+ 143,
+ 100
+ ],
+ "category_id": 15,
+ "id": 2479,
+ "area": 14300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1244,
+ "bbox": [
+ 99,
+ 461,
+ 141,
+ 37
+ ],
+ "category_id": 15,
+ "id": 2480,
+ "area": 5217,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1245,
+ "bbox": [
+ 79,
+ 71,
+ 110,
+ 127
+ ],
+ "category_id": 10,
+ "id": 2481,
+ "area": 13970,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1246,
+ "bbox": [
+ 56,
+ 12,
+ 41,
+ 12
+ ],
+ "category_id": 4,
+ "id": 2482,
+ "area": 492,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1246,
+ "bbox": [
+ 46,
+ 134,
+ 34,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2483,
+ "area": 340,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1246,
+ "bbox": [
+ 85,
+ 248,
+ 32,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2484,
+ "area": 320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1246,
+ "bbox": [
+ 134,
+ 370,
+ 32,
+ 8
+ ],
+ "category_id": 4,
+ "id": 2485,
+ "area": 256,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1246,
+ "bbox": [
+ 181,
+ 483,
+ 39,
+ 8
+ ],
+ "category_id": 4,
+ "id": 2486,
+ "area": 312,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1247,
+ "bbox": [
+ 94,
+ 56,
+ 57,
+ 209
+ ],
+ "category_id": 10,
+ "id": 2487,
+ "area": 11913,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1248,
+ "bbox": [
+ 45,
+ 147,
+ 19,
+ 20
+ ],
+ "category_id": 4,
+ "id": 2488,
+ "area": 380,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1248,
+ "bbox": [
+ 419,
+ 64,
+ 19,
+ 26
+ ],
+ "category_id": 4,
+ "id": 2489,
+ "area": 494,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1248,
+ "bbox": [
+ 80,
+ 330,
+ 24,
+ 22
+ ],
+ "category_id": 4,
+ "id": 2490,
+ "area": 528,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1249,
+ "bbox": [
+ 69,
+ 126,
+ 403,
+ 275
+ ],
+ "category_id": 10,
+ "id": 2491,
+ "area": 110825,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1250,
+ "bbox": [
+ 284,
+ 217,
+ 100,
+ 104
+ ],
+ "category_id": 15,
+ "id": 2492,
+ "area": 10400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1251,
+ "bbox": [
+ 85,
+ 101,
+ 393,
+ 269
+ ],
+ "category_id": 10,
+ "id": 2493,
+ "area": 105717,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1252,
+ "bbox": [
+ 326,
+ 140,
+ 24,
+ 61
+ ],
+ "category_id": 4,
+ "id": 2494,
+ "area": 1464,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1252,
+ "bbox": [
+ 160,
+ 344,
+ 27,
+ 62
+ ],
+ "category_id": 4,
+ "id": 2495,
+ "area": 1674,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1253,
+ "bbox": [
+ 309,
+ 176,
+ 93,
+ 89
+ ],
+ "category_id": 4,
+ "id": 2496,
+ "area": 8277,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1254,
+ "bbox": [
+ 211,
+ 138,
+ 105,
+ 138
+ ],
+ "category_id": 15,
+ "id": 2497,
+ "area": 14490,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1254,
+ "bbox": [
+ 186,
+ 273,
+ 107,
+ 135
+ ],
+ "category_id": 15,
+ "id": 2498,
+ "area": 14445,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1255,
+ "bbox": [
+ 140,
+ 404,
+ 301,
+ 35
+ ],
+ "category_id": 10,
+ "id": 2499,
+ "area": 10535,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1256,
+ "bbox": [
+ 46,
+ 260,
+ 340,
+ 74
+ ],
+ "category_id": 19,
+ "id": 2500,
+ "area": 25160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1257,
+ "bbox": [
+ 286,
+ 221,
+ 149,
+ 86
+ ],
+ "category_id": 19,
+ "id": 2501,
+ "area": 12814,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1258,
+ "bbox": [
+ 26,
+ 234,
+ 422,
+ 108
+ ],
+ "category_id": 16,
+ "id": 2502,
+ "area": 45576,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1259,
+ "bbox": [
+ 97,
+ 321,
+ 21,
+ 47
+ ],
+ "category_id": 15,
+ "id": 2503,
+ "area": 987,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1260,
+ "bbox": [
+ 94,
+ 193,
+ 279,
+ 127
+ ],
+ "category_id": 19,
+ "id": 2504,
+ "area": 35433,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1261,
+ "bbox": [
+ 312,
+ 296,
+ 59,
+ 75
+ ],
+ "category_id": 4,
+ "id": 2505,
+ "area": 4425,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1262,
+ "bbox": [
+ 285,
+ 259,
+ 31,
+ 16
+ ],
+ "category_id": 4,
+ "id": 2506,
+ "area": 496,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1262,
+ "bbox": [
+ 481,
+ 181,
+ 30,
+ 17
+ ],
+ "category_id": 4,
+ "id": 2507,
+ "area": 510,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1262,
+ "bbox": [
+ 126,
+ 375,
+ 30,
+ 17
+ ],
+ "category_id": 4,
+ "id": 2508,
+ "area": 510,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1262,
+ "bbox": [
+ 0,
+ 356,
+ 21,
+ 20
+ ],
+ "category_id": 4,
+ "id": 2509,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1262,
+ "bbox": [
+ 479,
+ 422,
+ 11,
+ 25
+ ],
+ "category_id": 4,
+ "id": 2510,
+ "area": 275,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1263,
+ "bbox": [
+ 374,
+ 12,
+ 60,
+ 174
+ ],
+ "category_id": 10,
+ "id": 2511,
+ "area": 10440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1264,
+ "bbox": [
+ 223,
+ 206,
+ 86,
+ 113
+ ],
+ "category_id": 15,
+ "id": 2512,
+ "area": 9718,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1264,
+ "bbox": [
+ 316,
+ 0,
+ 13,
+ 35
+ ],
+ "category_id": 15,
+ "id": 2513,
+ "area": 455,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1265,
+ "bbox": [
+ 119,
+ 183,
+ 220,
+ 151
+ ],
+ "category_id": 16,
+ "id": 2514,
+ "area": 33220,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1266,
+ "bbox": [
+ 224,
+ 70,
+ 92,
+ 420
+ ],
+ "category_id": 10,
+ "id": 2515,
+ "area": 38640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1267,
+ "bbox": [
+ 33,
+ 194,
+ 55,
+ 60
+ ],
+ "category_id": 4,
+ "id": 2516,
+ "area": 3300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1267,
+ "bbox": [
+ 358,
+ 148,
+ 45,
+ 48
+ ],
+ "category_id": 4,
+ "id": 2517,
+ "area": 2160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1267,
+ "bbox": [
+ 432,
+ 410,
+ 43,
+ 43
+ ],
+ "category_id": 4,
+ "id": 2518,
+ "area": 1849,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1268,
+ "bbox": [
+ 133,
+ 211,
+ 71,
+ 183
+ ],
+ "category_id": 19,
+ "id": 2519,
+ "area": 12993,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1268,
+ "bbox": [
+ 226,
+ 159,
+ 80,
+ 277
+ ],
+ "category_id": 19,
+ "id": 2520,
+ "area": 22160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1269,
+ "bbox": [
+ 31,
+ 219,
+ 420,
+ 45
+ ],
+ "category_id": 19,
+ "id": 2521,
+ "area": 18900,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1270,
+ "bbox": [
+ 277,
+ 305,
+ 102,
+ 159
+ ],
+ "category_id": 4,
+ "id": 2522,
+ "area": 16218,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1271,
+ "bbox": [
+ 225,
+ 204,
+ 128,
+ 138
+ ],
+ "category_id": 15,
+ "id": 2523,
+ "area": 17664,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1272,
+ "bbox": [
+ 87,
+ 28,
+ 127,
+ 136
+ ],
+ "category_id": 19,
+ "id": 2524,
+ "area": 17272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1273,
+ "bbox": [
+ 433,
+ 53,
+ 42,
+ 48
+ ],
+ "category_id": 4,
+ "id": 2525,
+ "area": 2016,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1273,
+ "bbox": [
+ 264,
+ 196,
+ 54,
+ 44
+ ],
+ "category_id": 4,
+ "id": 2526,
+ "area": 2376,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1273,
+ "bbox": [
+ 211,
+ 405,
+ 45,
+ 39
+ ],
+ "category_id": 4,
+ "id": 2527,
+ "area": 1755,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1274,
+ "bbox": [
+ 179,
+ 436,
+ 52,
+ 52
+ ],
+ "category_id": 4,
+ "id": 2528,
+ "area": 2704,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1274,
+ "bbox": [
+ 410,
+ 152,
+ 46,
+ 35
+ ],
+ "category_id": 4,
+ "id": 2529,
+ "area": 1610,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1275,
+ "bbox": [
+ 28,
+ 244,
+ 20,
+ 13
+ ],
+ "category_id": 4,
+ "id": 2530,
+ "area": 260,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1275,
+ "bbox": [
+ 174,
+ 202,
+ 21,
+ 15
+ ],
+ "category_id": 4,
+ "id": 2531,
+ "area": 315,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1275,
+ "bbox": [
+ 236,
+ 80,
+ 36,
+ 18
+ ],
+ "category_id": 4,
+ "id": 2532,
+ "area": 648,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1275,
+ "bbox": [
+ 464,
+ 464,
+ 21,
+ 11
+ ],
+ "category_id": 4,
+ "id": 2533,
+ "area": 231,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1276,
+ "bbox": [
+ 69,
+ 79,
+ 346,
+ 160
+ ],
+ "category_id": 10,
+ "id": 2534,
+ "area": 55360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1277,
+ "bbox": [
+ 213,
+ 113,
+ 151,
+ 319
+ ],
+ "category_id": 16,
+ "id": 2535,
+ "area": 48169,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1278,
+ "bbox": [
+ 190,
+ 99,
+ 186,
+ 281
+ ],
+ "category_id": 16,
+ "id": 2536,
+ "area": 52266,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1279,
+ "bbox": [
+ 65,
+ 144,
+ 370,
+ 216
+ ],
+ "category_id": 10,
+ "id": 2537,
+ "area": 79920,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1280,
+ "bbox": [
+ 48,
+ 167,
+ 14,
+ 52
+ ],
+ "category_id": 4,
+ "id": 2538,
+ "area": 728,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1280,
+ "bbox": [
+ 200,
+ 400,
+ 15,
+ 49
+ ],
+ "category_id": 4,
+ "id": 2539,
+ "area": 735,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1280,
+ "bbox": [
+ 457,
+ 76,
+ 13,
+ 50
+ ],
+ "category_id": 4,
+ "id": 2540,
+ "area": 650,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1281,
+ "bbox": [
+ 96,
+ 251,
+ 320,
+ 52
+ ],
+ "category_id": 10,
+ "id": 2541,
+ "area": 16640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1282,
+ "bbox": [
+ 239,
+ 155,
+ 41,
+ 146
+ ],
+ "category_id": 19,
+ "id": 2542,
+ "area": 5986,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1283,
+ "bbox": [
+ 45,
+ 128,
+ 408,
+ 225
+ ],
+ "category_id": 16,
+ "id": 2543,
+ "area": 91800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1284,
+ "bbox": [
+ 141,
+ 154,
+ 103,
+ 126
+ ],
+ "category_id": 15,
+ "id": 2544,
+ "area": 12978,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1284,
+ "bbox": [
+ 284,
+ 160,
+ 100,
+ 128
+ ],
+ "category_id": 15,
+ "id": 2545,
+ "area": 12800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1285,
+ "bbox": [
+ 184,
+ 92,
+ 59,
+ 42
+ ],
+ "category_id": 4,
+ "id": 2546,
+ "area": 2478,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1285,
+ "bbox": [
+ 336,
+ 392,
+ 57,
+ 47
+ ],
+ "category_id": 4,
+ "id": 2547,
+ "area": 2679,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1285,
+ "bbox": [
+ 451,
+ 115,
+ 41,
+ 34
+ ],
+ "category_id": 4,
+ "id": 2548,
+ "area": 1394,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1286,
+ "bbox": [
+ 111,
+ 226,
+ 288,
+ 42
+ ],
+ "category_id": 19,
+ "id": 2549,
+ "area": 12096,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1287,
+ "bbox": [
+ 268,
+ 67,
+ 92,
+ 117
+ ],
+ "category_id": 15,
+ "id": 2550,
+ "area": 10764,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1287,
+ "bbox": [
+ 220,
+ 195,
+ 90,
+ 117
+ ],
+ "category_id": 15,
+ "id": 2551,
+ "area": 10530,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1287,
+ "bbox": [
+ 172,
+ 309,
+ 95,
+ 123
+ ],
+ "category_id": 15,
+ "id": 2552,
+ "area": 11685,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1288,
+ "bbox": [
+ 226,
+ 136,
+ 40,
+ 295
+ ],
+ "category_id": 19,
+ "id": 2553,
+ "area": 11800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1289,
+ "bbox": [
+ 140,
+ 106,
+ 337,
+ 279
+ ],
+ "category_id": 10,
+ "id": 2554,
+ "area": 94023,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1290,
+ "bbox": [
+ 185,
+ 103,
+ 131,
+ 262
+ ],
+ "category_id": 16,
+ "id": 2555,
+ "area": 34322,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1291,
+ "bbox": [
+ 211,
+ 161,
+ 29,
+ 151
+ ],
+ "category_id": 19,
+ "id": 2556,
+ "area": 4379,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1292,
+ "bbox": [
+ 205,
+ 243,
+ 130,
+ 50
+ ],
+ "category_id": 4,
+ "id": 2557,
+ "area": 6500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1293,
+ "bbox": [
+ 346,
+ 83,
+ 82,
+ 92
+ ],
+ "category_id": 19,
+ "id": 2558,
+ "area": 7544,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1294,
+ "bbox": [
+ 154,
+ 199,
+ 230,
+ 92
+ ],
+ "category_id": 16,
+ "id": 2559,
+ "area": 21160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1295,
+ "bbox": [
+ 248,
+ 218,
+ 40,
+ 100
+ ],
+ "category_id": 4,
+ "id": 2560,
+ "area": 4000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1296,
+ "bbox": [
+ 307,
+ 195,
+ 97,
+ 99
+ ],
+ "category_id": 15,
+ "id": 2561,
+ "area": 9603,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1297,
+ "bbox": [
+ 264,
+ 174,
+ 65,
+ 35
+ ],
+ "category_id": 4,
+ "id": 2562,
+ "area": 2275,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1297,
+ "bbox": [
+ 96,
+ 429,
+ 71,
+ 33
+ ],
+ "category_id": 4,
+ "id": 2563,
+ "area": 2343,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1298,
+ "bbox": [
+ 129,
+ 74,
+ 30,
+ 75
+ ],
+ "category_id": 4,
+ "id": 2564,
+ "area": 2250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1298,
+ "bbox": [
+ 392,
+ 405,
+ 30,
+ 63
+ ],
+ "category_id": 4,
+ "id": 2565,
+ "area": 1890,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1299,
+ "bbox": [
+ 169,
+ 204,
+ 138,
+ 107
+ ],
+ "category_id": 19,
+ "id": 2566,
+ "area": 14766,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1300,
+ "bbox": [
+ 343,
+ 42,
+ 146,
+ 224
+ ],
+ "category_id": 10,
+ "id": 2567,
+ "area": 32704,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1301,
+ "bbox": [
+ 85,
+ 210,
+ 125,
+ 147
+ ],
+ "category_id": 15,
+ "id": 2568,
+ "area": 18375,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1301,
+ "bbox": [
+ 285,
+ 88,
+ 127,
+ 157
+ ],
+ "category_id": 15,
+ "id": 2569,
+ "area": 19939,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1302,
+ "bbox": [
+ 46,
+ 391,
+ 190,
+ 69
+ ],
+ "category_id": 10,
+ "id": 2570,
+ "area": 13110,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1303,
+ "bbox": [
+ 217,
+ 55,
+ 21,
+ 20
+ ],
+ "category_id": 4,
+ "id": 2571,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1303,
+ "bbox": [
+ 215,
+ 172,
+ 19,
+ 25
+ ],
+ "category_id": 4,
+ "id": 2572,
+ "area": 475,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1303,
+ "bbox": [
+ 56,
+ 399,
+ 21,
+ 22
+ ],
+ "category_id": 4,
+ "id": 2573,
+ "area": 462,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1303,
+ "bbox": [
+ 417,
+ 321,
+ 27,
+ 33
+ ],
+ "category_id": 4,
+ "id": 2574,
+ "area": 891,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1303,
+ "bbox": [
+ 436,
+ 483,
+ 23,
+ 27
+ ],
+ "category_id": 4,
+ "id": 2575,
+ "area": 621,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1304,
+ "bbox": [
+ 17,
+ 435,
+ 469,
+ 20
+ ],
+ "category_id": 10,
+ "id": 2576,
+ "area": 9380,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1305,
+ "bbox": [
+ 163,
+ 84,
+ 215,
+ 259
+ ],
+ "category_id": 16,
+ "id": 2577,
+ "area": 55685,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1306,
+ "bbox": [
+ 197,
+ 154,
+ 51,
+ 213
+ ],
+ "category_id": 19,
+ "id": 2578,
+ "area": 10863,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1307,
+ "bbox": [
+ 86,
+ 282,
+ 52,
+ 57
+ ],
+ "category_id": 4,
+ "id": 2579,
+ "area": 2964,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1307,
+ "bbox": [
+ 342,
+ 144,
+ 50,
+ 71
+ ],
+ "category_id": 4,
+ "id": 2580,
+ "area": 3550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1308,
+ "bbox": [
+ 27,
+ 293,
+ 183,
+ 94
+ ],
+ "category_id": 19,
+ "id": 2581,
+ "area": 17202,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1309,
+ "bbox": [
+ 417,
+ 5,
+ 27,
+ 17
+ ],
+ "category_id": 4,
+ "id": 2582,
+ "area": 459,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1309,
+ "bbox": [
+ 272,
+ 40,
+ 31,
+ 16
+ ],
+ "category_id": 4,
+ "id": 2583,
+ "area": 496,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1309,
+ "bbox": [
+ 106,
+ 55,
+ 30,
+ 16
+ ],
+ "category_id": 4,
+ "id": 2584,
+ "area": 480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1309,
+ "bbox": [
+ 10,
+ 140,
+ 27,
+ 13
+ ],
+ "category_id": 4,
+ "id": 2585,
+ "area": 351,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1309,
+ "bbox": [
+ 322,
+ 206,
+ 30,
+ 18
+ ],
+ "category_id": 4,
+ "id": 2586,
+ "area": 540,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1309,
+ "bbox": [
+ 268,
+ 331,
+ 30,
+ 15
+ ],
+ "category_id": 4,
+ "id": 2587,
+ "area": 450,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1309,
+ "bbox": [
+ 220,
+ 437,
+ 13,
+ 23
+ ],
+ "category_id": 4,
+ "id": 2588,
+ "area": 299,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1309,
+ "bbox": [
+ 22,
+ 468,
+ 31,
+ 17
+ ],
+ "category_id": 4,
+ "id": 2589,
+ "area": 527,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1310,
+ "bbox": [
+ 276,
+ 122,
+ 51,
+ 320
+ ],
+ "category_id": 10,
+ "id": 2590,
+ "area": 16320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1311,
+ "bbox": [
+ 96,
+ 334,
+ 40,
+ 35
+ ],
+ "category_id": 4,
+ "id": 2591,
+ "area": 1400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1311,
+ "bbox": [
+ 398,
+ 177,
+ 50,
+ 35
+ ],
+ "category_id": 4,
+ "id": 2592,
+ "area": 1750,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1312,
+ "bbox": [
+ 112,
+ 295,
+ 20,
+ 69
+ ],
+ "category_id": 16,
+ "id": 2593,
+ "area": 1380,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1313,
+ "bbox": [
+ 30,
+ 255,
+ 46,
+ 49
+ ],
+ "category_id": 4,
+ "id": 2594,
+ "area": 2254,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1313,
+ "bbox": [
+ 428,
+ 236,
+ 47,
+ 41
+ ],
+ "category_id": 4,
+ "id": 2595,
+ "area": 1927,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1314,
+ "bbox": [
+ 33,
+ 30,
+ 137,
+ 188
+ ],
+ "category_id": 10,
+ "id": 2596,
+ "area": 25756,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1315,
+ "bbox": [
+ 265,
+ 112,
+ 124,
+ 224
+ ],
+ "category_id": 16,
+ "id": 2597,
+ "area": 27776,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1316,
+ "bbox": [
+ 285,
+ 237,
+ 98,
+ 82
+ ],
+ "category_id": 4,
+ "id": 2598,
+ "area": 8036,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1317,
+ "bbox": [
+ 72,
+ 347,
+ 17,
+ 8
+ ],
+ "category_id": 4,
+ "id": 2599,
+ "area": 136,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1317,
+ "bbox": [
+ 66,
+ 458,
+ 17,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2600,
+ "area": 170,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1317,
+ "bbox": [
+ 186,
+ 422,
+ 19,
+ 6
+ ],
+ "category_id": 4,
+ "id": 2601,
+ "area": 114,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1317,
+ "bbox": [
+ 460,
+ 181,
+ 28,
+ 17
+ ],
+ "category_id": 4,
+ "id": 2602,
+ "area": 476,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1317,
+ "bbox": [
+ 316,
+ 217,
+ 18,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2603,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1318,
+ "bbox": [
+ 255,
+ 185,
+ 119,
+ 209
+ ],
+ "category_id": 16,
+ "id": 2604,
+ "area": 24871,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1319,
+ "bbox": [
+ 74,
+ 127,
+ 54,
+ 28
+ ],
+ "category_id": 4,
+ "id": 2605,
+ "area": 1512,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1319,
+ "bbox": [
+ 215,
+ 353,
+ 53,
+ 29
+ ],
+ "category_id": 4,
+ "id": 2606,
+ "area": 1537,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1319,
+ "bbox": [
+ 416,
+ 120,
+ 55,
+ 31
+ ],
+ "category_id": 4,
+ "id": 2607,
+ "area": 1705,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1320,
+ "bbox": [
+ 202,
+ 240,
+ 100,
+ 79
+ ],
+ "category_id": 15,
+ "id": 2608,
+ "area": 7900,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1320,
+ "bbox": [
+ 311,
+ 230,
+ 99,
+ 81
+ ],
+ "category_id": 15,
+ "id": 2609,
+ "area": 8019,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1321,
+ "bbox": [
+ 204,
+ 128,
+ 242,
+ 328
+ ],
+ "category_id": 10,
+ "id": 2610,
+ "area": 79376,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1322,
+ "bbox": [
+ 167,
+ 166,
+ 103,
+ 102
+ ],
+ "category_id": 15,
+ "id": 2611,
+ "area": 10506,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1322,
+ "bbox": [
+ 294,
+ 154,
+ 101,
+ 104
+ ],
+ "category_id": 15,
+ "id": 2612,
+ "area": 10504,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1323,
+ "bbox": [
+ 60,
+ 55,
+ 14,
+ 20
+ ],
+ "category_id": 4,
+ "id": 2613,
+ "area": 280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1323,
+ "bbox": [
+ 422,
+ 109,
+ 11,
+ 21
+ ],
+ "category_id": 4,
+ "id": 2614,
+ "area": 231,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1323,
+ "bbox": [
+ 482,
+ 263,
+ 11,
+ 22
+ ],
+ "category_id": 4,
+ "id": 2615,
+ "area": 242,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1323,
+ "bbox": [
+ 330,
+ 379,
+ 11,
+ 24
+ ],
+ "category_id": 4,
+ "id": 2616,
+ "area": 264,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1323,
+ "bbox": [
+ 134,
+ 211,
+ 11,
+ 19
+ ],
+ "category_id": 4,
+ "id": 2617,
+ "area": 209,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1323,
+ "bbox": [
+ 75,
+ 312,
+ 13,
+ 23
+ ],
+ "category_id": 4,
+ "id": 2618,
+ "area": 299,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1323,
+ "bbox": [
+ 19,
+ 401,
+ 14,
+ 25
+ ],
+ "category_id": 4,
+ "id": 2619,
+ "area": 350,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1324,
+ "bbox": [
+ 199,
+ 133,
+ 70,
+ 257
+ ],
+ "category_id": 19,
+ "id": 2620,
+ "area": 17990,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1325,
+ "bbox": [
+ 7,
+ 67,
+ 176,
+ 334
+ ],
+ "category_id": 16,
+ "id": 2621,
+ "area": 58784,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1326,
+ "bbox": [
+ 48,
+ 254,
+ 419,
+ 70
+ ],
+ "category_id": 19,
+ "id": 2622,
+ "area": 29330,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1327,
+ "bbox": [
+ 143,
+ 184,
+ 72,
+ 120
+ ],
+ "category_id": 19,
+ "id": 2623,
+ "area": 8640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1328,
+ "bbox": [
+ 134,
+ 252,
+ 303,
+ 170
+ ],
+ "category_id": 16,
+ "id": 2624,
+ "area": 51510,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1329,
+ "bbox": [
+ 172,
+ 72,
+ 112,
+ 101
+ ],
+ "category_id": 15,
+ "id": 2625,
+ "area": 11312,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1329,
+ "bbox": [
+ 236,
+ 203,
+ 110,
+ 101
+ ],
+ "category_id": 15,
+ "id": 2626,
+ "area": 11110,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1329,
+ "bbox": [
+ 293,
+ 330,
+ 117,
+ 103
+ ],
+ "category_id": 15,
+ "id": 2627,
+ "area": 12051,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1330,
+ "bbox": [
+ 41,
+ 213,
+ 309,
+ 150
+ ],
+ "category_id": 10,
+ "id": 2628,
+ "area": 46350,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1331,
+ "bbox": [
+ 235,
+ 122,
+ 17,
+ 21
+ ],
+ "category_id": 4,
+ "id": 2629,
+ "area": 357,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1331,
+ "bbox": [
+ 353,
+ 120,
+ 18,
+ 24
+ ],
+ "category_id": 4,
+ "id": 2630,
+ "area": 432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1331,
+ "bbox": [
+ 46,
+ 381,
+ 18,
+ 23
+ ],
+ "category_id": 4,
+ "id": 2631,
+ "area": 414,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1331,
+ "bbox": [
+ 197,
+ 356,
+ 14,
+ 22
+ ],
+ "category_id": 4,
+ "id": 2632,
+ "area": 308,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1331,
+ "bbox": [
+ 385,
+ 259,
+ 15,
+ 27
+ ],
+ "category_id": 4,
+ "id": 2633,
+ "area": 405,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1331,
+ "bbox": [
+ 440,
+ 439,
+ 14,
+ 22
+ ],
+ "category_id": 4,
+ "id": 2634,
+ "area": 308,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1332,
+ "bbox": [
+ 287,
+ 261,
+ 125,
+ 82
+ ],
+ "category_id": 4,
+ "id": 2635,
+ "area": 10250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1333,
+ "bbox": [
+ 166,
+ 120,
+ 283,
+ 321
+ ],
+ "category_id": 10,
+ "id": 2636,
+ "area": 90843,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1334,
+ "bbox": [
+ 128,
+ 141,
+ 143,
+ 134
+ ],
+ "category_id": 15,
+ "id": 2637,
+ "area": 19162,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1334,
+ "bbox": [
+ 306,
+ 147,
+ 143,
+ 135
+ ],
+ "category_id": 15,
+ "id": 2638,
+ "area": 19305,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1334,
+ "bbox": [
+ 122,
+ 318,
+ 143,
+ 137
+ ],
+ "category_id": 15,
+ "id": 2639,
+ "area": 19591,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1334,
+ "bbox": [
+ 298,
+ 321,
+ 146,
+ 144
+ ],
+ "category_id": 15,
+ "id": 2640,
+ "area": 21024,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1335,
+ "bbox": [
+ 205,
+ 205,
+ 198,
+ 177
+ ],
+ "category_id": 15,
+ "id": 2641,
+ "area": 35046,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1336,
+ "bbox": [
+ 279,
+ 126,
+ 67,
+ 231
+ ],
+ "category_id": 19,
+ "id": 2642,
+ "area": 15477,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1337,
+ "bbox": [
+ 209,
+ 90,
+ 222,
+ 375
+ ],
+ "category_id": 10,
+ "id": 2643,
+ "area": 83250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1338,
+ "bbox": [
+ 287,
+ 24,
+ 221,
+ 216
+ ],
+ "category_id": 15,
+ "id": 2644,
+ "area": 47736,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1338,
+ "bbox": [
+ 35,
+ 166,
+ 104,
+ 104
+ ],
+ "category_id": 15,
+ "id": 2645,
+ "area": 10816,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1338,
+ "bbox": [
+ 3,
+ 263,
+ 217,
+ 217
+ ],
+ "category_id": 15,
+ "id": 2646,
+ "area": 47089,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1338,
+ "bbox": [
+ 245,
+ 263,
+ 221,
+ 213
+ ],
+ "category_id": 15,
+ "id": 2647,
+ "area": 47073,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1339,
+ "bbox": [
+ 279,
+ 68,
+ 103,
+ 377
+ ],
+ "category_id": 10,
+ "id": 2648,
+ "area": 38831,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1340,
+ "bbox": [
+ 83,
+ 140,
+ 31,
+ 18
+ ],
+ "category_id": 4,
+ "id": 2649,
+ "area": 558,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1340,
+ "bbox": [
+ 366,
+ 101,
+ 17,
+ 9
+ ],
+ "category_id": 4,
+ "id": 2650,
+ "area": 153,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1340,
+ "bbox": [
+ 16,
+ 217,
+ 28,
+ 14
+ ],
+ "category_id": 4,
+ "id": 2651,
+ "area": 392,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1340,
+ "bbox": [
+ 278,
+ 263,
+ 22,
+ 18
+ ],
+ "category_id": 4,
+ "id": 2652,
+ "area": 396,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1340,
+ "bbox": [
+ 195,
+ 300,
+ 29,
+ 15
+ ],
+ "category_id": 4,
+ "id": 2653,
+ "area": 435,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1340,
+ "bbox": [
+ 407,
+ 273,
+ 16,
+ 9
+ ],
+ "category_id": 4,
+ "id": 2654,
+ "area": 144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1341,
+ "bbox": [
+ 432,
+ 151,
+ 25,
+ 35
+ ],
+ "category_id": 4,
+ "id": 2655,
+ "area": 875,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1341,
+ "bbox": [
+ 186,
+ 251,
+ 29,
+ 29
+ ],
+ "category_id": 4,
+ "id": 2656,
+ "area": 841,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1341,
+ "bbox": [
+ 412,
+ 383,
+ 33,
+ 38
+ ],
+ "category_id": 4,
+ "id": 2657,
+ "area": 1254,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1342,
+ "bbox": [
+ 65,
+ 231,
+ 348,
+ 228
+ ],
+ "category_id": 19,
+ "id": 2658,
+ "area": 79344,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1343,
+ "bbox": [
+ 44,
+ 315,
+ 67,
+ 28
+ ],
+ "category_id": 4,
+ "id": 2659,
+ "area": 1876,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1343,
+ "bbox": [
+ 408,
+ 122,
+ 61,
+ 29
+ ],
+ "category_id": 4,
+ "id": 2660,
+ "area": 1769,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1344,
+ "bbox": [
+ 172,
+ 129,
+ 42,
+ 44
+ ],
+ "category_id": 4,
+ "id": 2661,
+ "area": 1848,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1344,
+ "bbox": [
+ 183,
+ 407,
+ 42,
+ 45
+ ],
+ "category_id": 4,
+ "id": 2662,
+ "area": 1890,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1345,
+ "bbox": [
+ 88,
+ 37,
+ 360,
+ 373
+ ],
+ "category_id": 10,
+ "id": 2663,
+ "area": 134280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1346,
+ "bbox": [
+ 168,
+ 179,
+ 216,
+ 22
+ ],
+ "category_id": 10,
+ "id": 2664,
+ "area": 4752,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1347,
+ "bbox": [
+ 83,
+ 68,
+ 22,
+ 25
+ ],
+ "category_id": 4,
+ "id": 2665,
+ "area": 550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1347,
+ "bbox": [
+ 216,
+ 145,
+ 20,
+ 21
+ ],
+ "category_id": 4,
+ "id": 2666,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1347,
+ "bbox": [
+ 427,
+ 127,
+ 18,
+ 23
+ ],
+ "category_id": 4,
+ "id": 2667,
+ "area": 414,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1347,
+ "bbox": [
+ 170,
+ 264,
+ 20,
+ 16
+ ],
+ "category_id": 4,
+ "id": 2668,
+ "area": 320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1347,
+ "bbox": [
+ 51,
+ 361,
+ 21,
+ 15
+ ],
+ "category_id": 4,
+ "id": 2669,
+ "area": 315,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1348,
+ "bbox": [
+ 215,
+ 43,
+ 10,
+ 16
+ ],
+ "category_id": 4,
+ "id": 2670,
+ "area": 160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1348,
+ "bbox": [
+ 66,
+ 88,
+ 12,
+ 24
+ ],
+ "category_id": 4,
+ "id": 2671,
+ "area": 288,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1348,
+ "bbox": [
+ 122,
+ 207,
+ 14,
+ 28
+ ],
+ "category_id": 4,
+ "id": 2672,
+ "area": 392,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1348,
+ "bbox": [
+ 62,
+ 371,
+ 13,
+ 25
+ ],
+ "category_id": 4,
+ "id": 2673,
+ "area": 325,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1348,
+ "bbox": [
+ 226,
+ 456,
+ 14,
+ 24
+ ],
+ "category_id": 4,
+ "id": 2674,
+ "area": 336,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1348,
+ "bbox": [
+ 318,
+ 255,
+ 16,
+ 20
+ ],
+ "category_id": 4,
+ "id": 2675,
+ "area": 320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1348,
+ "bbox": [
+ 454,
+ 147,
+ 10,
+ 25
+ ],
+ "category_id": 4,
+ "id": 2676,
+ "area": 250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1349,
+ "bbox": [
+ 62,
+ 289,
+ 44,
+ 34
+ ],
+ "category_id": 4,
+ "id": 2677,
+ "area": 1496,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1349,
+ "bbox": [
+ 434,
+ 190,
+ 44,
+ 32
+ ],
+ "category_id": 4,
+ "id": 2678,
+ "area": 1408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1350,
+ "bbox": [
+ 306,
+ 111,
+ 26,
+ 66
+ ],
+ "category_id": 4,
+ "id": 2679,
+ "area": 1716,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1350,
+ "bbox": [
+ 78,
+ 408,
+ 30,
+ 61
+ ],
+ "category_id": 4,
+ "id": 2680,
+ "area": 1830,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1351,
+ "bbox": [
+ 220,
+ 350,
+ 78,
+ 55
+ ],
+ "category_id": 16,
+ "id": 2681,
+ "area": 4290,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1352,
+ "bbox": [
+ 291,
+ 94,
+ 208,
+ 313
+ ],
+ "category_id": 10,
+ "id": 2682,
+ "area": 65104,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1353,
+ "bbox": [
+ 209,
+ 422,
+ 55,
+ 49
+ ],
+ "category_id": 4,
+ "id": 2683,
+ "area": 2695,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1353,
+ "bbox": [
+ 363,
+ 65,
+ 56,
+ 50
+ ],
+ "category_id": 4,
+ "id": 2684,
+ "area": 2800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1354,
+ "bbox": [
+ 115,
+ 127,
+ 261,
+ 159
+ ],
+ "category_id": 16,
+ "id": 2685,
+ "area": 41499,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1355,
+ "bbox": [
+ 378,
+ 167,
+ 86,
+ 101
+ ],
+ "category_id": 15,
+ "id": 2686,
+ "area": 8686,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1355,
+ "bbox": [
+ 158,
+ 290,
+ 85,
+ 93
+ ],
+ "category_id": 15,
+ "id": 2687,
+ "area": 7905,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1356,
+ "bbox": [
+ 163,
+ 371,
+ 67,
+ 57
+ ],
+ "category_id": 4,
+ "id": 2688,
+ "area": 3819,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1356,
+ "bbox": [
+ 343,
+ 76,
+ 75,
+ 48
+ ],
+ "category_id": 4,
+ "id": 2689,
+ "area": 3600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1357,
+ "bbox": [
+ 421,
+ 272,
+ 44,
+ 49
+ ],
+ "category_id": 15,
+ "id": 2690,
+ "area": 2156,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1357,
+ "bbox": [
+ 257,
+ 442,
+ 34,
+ 38
+ ],
+ "category_id": 15,
+ "id": 2691,
+ "area": 1292,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1357,
+ "bbox": [
+ 289,
+ 422,
+ 19,
+ 18
+ ],
+ "category_id": 15,
+ "id": 2692,
+ "area": 342,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1357,
+ "bbox": [
+ 241,
+ 456,
+ 15,
+ 18
+ ],
+ "category_id": 15,
+ "id": 2693,
+ "area": 270,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1358,
+ "bbox": [
+ 165,
+ 237,
+ 171,
+ 58
+ ],
+ "category_id": 19,
+ "id": 2694,
+ "area": 9918,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1359,
+ "bbox": [
+ 103,
+ 214,
+ 75,
+ 54
+ ],
+ "category_id": 15,
+ "id": 2695,
+ "area": 4050,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1360,
+ "bbox": [
+ 0,
+ 154,
+ 55,
+ 25
+ ],
+ "category_id": 15,
+ "id": 2696,
+ "area": 1375,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1360,
+ "bbox": [
+ 252,
+ 175,
+ 184,
+ 158
+ ],
+ "category_id": 15,
+ "id": 2697,
+ "area": 29072,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1360,
+ "bbox": [
+ 240,
+ 339,
+ 102,
+ 103
+ ],
+ "category_id": 15,
+ "id": 2698,
+ "area": 10506,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1361,
+ "bbox": [
+ 207,
+ 101,
+ 54,
+ 285
+ ],
+ "category_id": 19,
+ "id": 2699,
+ "area": 15390,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1362,
+ "bbox": [
+ 72,
+ 227,
+ 25,
+ 29
+ ],
+ "category_id": 4,
+ "id": 2700,
+ "area": 725,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1362,
+ "bbox": [
+ 254,
+ 38,
+ 21,
+ 26
+ ],
+ "category_id": 4,
+ "id": 2701,
+ "area": 546,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1362,
+ "bbox": [
+ 246,
+ 395,
+ 22,
+ 25
+ ],
+ "category_id": 4,
+ "id": 2702,
+ "area": 550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1362,
+ "bbox": [
+ 438,
+ 318,
+ 29,
+ 29
+ ],
+ "category_id": 4,
+ "id": 2703,
+ "area": 841,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1363,
+ "bbox": [
+ 174,
+ 179,
+ 154,
+ 232
+ ],
+ "category_id": 10,
+ "id": 2704,
+ "area": 35728,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1364,
+ "bbox": [
+ 133,
+ 116,
+ 288,
+ 205
+ ],
+ "category_id": 19,
+ "id": 2705,
+ "area": 59040,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1365,
+ "bbox": [
+ 111,
+ 181,
+ 322,
+ 227
+ ],
+ "category_id": 10,
+ "id": 2706,
+ "area": 73094,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1366,
+ "bbox": [
+ 172,
+ 84,
+ 75,
+ 34
+ ],
+ "category_id": 4,
+ "id": 2707,
+ "area": 2550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1366,
+ "bbox": [
+ 264,
+ 432,
+ 76,
+ 39
+ ],
+ "category_id": 4,
+ "id": 2708,
+ "area": 2964,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1367,
+ "bbox": [
+ 122,
+ 86,
+ 78,
+ 267
+ ],
+ "category_id": 19,
+ "id": 2709,
+ "area": 20826,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1368,
+ "bbox": [
+ 19,
+ 38,
+ 478,
+ 286
+ ],
+ "category_id": 10,
+ "id": 2710,
+ "area": 136708,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1369,
+ "bbox": [
+ 213,
+ 295,
+ 35,
+ 32
+ ],
+ "category_id": 16,
+ "id": 2711,
+ "area": 1120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1370,
+ "bbox": [
+ 44,
+ 95,
+ 403,
+ 312
+ ],
+ "category_id": 15,
+ "id": 2712,
+ "area": 125736,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1371,
+ "bbox": [
+ 86,
+ 229,
+ 357,
+ 114
+ ],
+ "category_id": 19,
+ "id": 2713,
+ "area": 40698,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1372,
+ "bbox": [
+ 61,
+ 8,
+ 323,
+ 169
+ ],
+ "category_id": 10,
+ "id": 2714,
+ "area": 54587,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1373,
+ "bbox": [
+ 64,
+ 120,
+ 296,
+ 285
+ ],
+ "category_id": 16,
+ "id": 2715,
+ "area": 84360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1374,
+ "bbox": [
+ 183,
+ 143,
+ 184,
+ 250
+ ],
+ "category_id": 10,
+ "id": 2716,
+ "area": 46000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1375,
+ "bbox": [
+ 150,
+ 357,
+ 23,
+ 25
+ ],
+ "category_id": 4,
+ "id": 2717,
+ "area": 575,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1375,
+ "bbox": [
+ 169,
+ 52,
+ 24,
+ 22
+ ],
+ "category_id": 4,
+ "id": 2718,
+ "area": 528,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1375,
+ "bbox": [
+ 415,
+ 382,
+ 24,
+ 23
+ ],
+ "category_id": 4,
+ "id": 2719,
+ "area": 552,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1376,
+ "bbox": [
+ 295,
+ 234,
+ 83,
+ 158
+ ],
+ "category_id": 16,
+ "id": 2720,
+ "area": 13114,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1377,
+ "bbox": [
+ 188,
+ 76,
+ 143,
+ 169
+ ],
+ "category_id": 15,
+ "id": 2721,
+ "area": 24167,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1377,
+ "bbox": [
+ 191,
+ 284,
+ 141,
+ 144
+ ],
+ "category_id": 15,
+ "id": 2722,
+ "area": 20304,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1378,
+ "bbox": [
+ 50,
+ 229,
+ 94,
+ 98
+ ],
+ "category_id": 15,
+ "id": 2723,
+ "area": 9212,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1378,
+ "bbox": [
+ 5,
+ 368,
+ 110,
+ 108
+ ],
+ "category_id": 15,
+ "id": 2724,
+ "area": 11880,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1378,
+ "bbox": [
+ 163,
+ 186,
+ 100,
+ 102
+ ],
+ "category_id": 15,
+ "id": 2725,
+ "area": 10200,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1378,
+ "bbox": [
+ 303,
+ 156,
+ 80,
+ 84
+ ],
+ "category_id": 15,
+ "id": 2726,
+ "area": 6720,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1378,
+ "bbox": [
+ 412,
+ 160,
+ 80,
+ 83
+ ],
+ "category_id": 15,
+ "id": 2727,
+ "area": 6640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1378,
+ "bbox": [
+ 372,
+ 474,
+ 19,
+ 38
+ ],
+ "category_id": 15,
+ "id": 2728,
+ "area": 722,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1379,
+ "bbox": [
+ 84,
+ 104,
+ 185,
+ 158
+ ],
+ "category_id": 15,
+ "id": 2729,
+ "area": 29230,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1379,
+ "bbox": [
+ 267,
+ 233,
+ 205,
+ 197
+ ],
+ "category_id": 15,
+ "id": 2730,
+ "area": 40385,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1380,
+ "bbox": [
+ 182,
+ 143,
+ 274,
+ 208
+ ],
+ "category_id": 10,
+ "id": 2731,
+ "area": 56992,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1381,
+ "bbox": [
+ 89,
+ 192,
+ 250,
+ 166
+ ],
+ "category_id": 16,
+ "id": 2732,
+ "area": 41500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1382,
+ "bbox": [
+ 208,
+ 122,
+ 263,
+ 339
+ ],
+ "category_id": 10,
+ "id": 2733,
+ "area": 89157,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1383,
+ "bbox": [
+ 30,
+ 128,
+ 51,
+ 170
+ ],
+ "category_id": 19,
+ "id": 2734,
+ "area": 8670,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1384,
+ "bbox": [
+ 112,
+ 223,
+ 88,
+ 123
+ ],
+ "category_id": 16,
+ "id": 2735,
+ "area": 10824,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1384,
+ "bbox": [
+ 231,
+ 361,
+ 30,
+ 32
+ ],
+ "category_id": 16,
+ "id": 2736,
+ "area": 960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1385,
+ "bbox": [
+ 239,
+ 188,
+ 69,
+ 63
+ ],
+ "category_id": 4,
+ "id": 2737,
+ "area": 4347,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1386,
+ "bbox": [
+ 71,
+ 201,
+ 315,
+ 209
+ ],
+ "category_id": 16,
+ "id": 2738,
+ "area": 65835,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1387,
+ "bbox": [
+ 114,
+ 56,
+ 216,
+ 323
+ ],
+ "category_id": 19,
+ "id": 2739,
+ "area": 69768,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1388,
+ "bbox": [
+ 299,
+ 373,
+ 22,
+ 24
+ ],
+ "category_id": 15,
+ "id": 2740,
+ "area": 528,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1388,
+ "bbox": [
+ 288,
+ 367,
+ 12,
+ 16
+ ],
+ "category_id": 15,
+ "id": 2741,
+ "area": 192,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1389,
+ "bbox": [
+ 87,
+ 76,
+ 253,
+ 247
+ ],
+ "category_id": 19,
+ "id": 2742,
+ "area": 62491,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1390,
+ "bbox": [
+ 18,
+ 207,
+ 343,
+ 178
+ ],
+ "category_id": 10,
+ "id": 2743,
+ "area": 61054,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1391,
+ "bbox": [
+ 104,
+ 209,
+ 281,
+ 80
+ ],
+ "category_id": 19,
+ "id": 2744,
+ "area": 22480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1392,
+ "bbox": [
+ 119,
+ 186,
+ 242,
+ 78
+ ],
+ "category_id": 19,
+ "id": 2745,
+ "area": 18876,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1393,
+ "bbox": [
+ 201,
+ 106,
+ 110,
+ 304
+ ],
+ "category_id": 10,
+ "id": 2746,
+ "area": 33440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1394,
+ "bbox": [
+ 35,
+ 161,
+ 66,
+ 39
+ ],
+ "category_id": 4,
+ "id": 2747,
+ "area": 2574,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1394,
+ "bbox": [
+ 417,
+ 280,
+ 71,
+ 40
+ ],
+ "category_id": 4,
+ "id": 2748,
+ "area": 2840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1395,
+ "bbox": [
+ 55,
+ 161,
+ 92,
+ 86
+ ],
+ "category_id": 4,
+ "id": 2749,
+ "area": 7912,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1395,
+ "bbox": [
+ 376,
+ 293,
+ 114,
+ 51
+ ],
+ "category_id": 4,
+ "id": 2750,
+ "area": 5814,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1396,
+ "bbox": [
+ 213,
+ 139,
+ 98,
+ 207
+ ],
+ "category_id": 19,
+ "id": 2751,
+ "area": 20286,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1397,
+ "bbox": [
+ 170,
+ 250,
+ 212,
+ 145
+ ],
+ "category_id": 19,
+ "id": 2752,
+ "area": 30740,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1398,
+ "bbox": [
+ 231,
+ 169,
+ 62,
+ 145
+ ],
+ "category_id": 19,
+ "id": 2753,
+ "area": 8990,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1399,
+ "bbox": [
+ 154,
+ 197,
+ 255,
+ 210
+ ],
+ "category_id": 16,
+ "id": 2754,
+ "area": 53550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1400,
+ "bbox": [
+ 215,
+ 152,
+ 187,
+ 292
+ ],
+ "category_id": 16,
+ "id": 2755,
+ "area": 54604,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1401,
+ "bbox": [
+ 312,
+ 103,
+ 192,
+ 51
+ ],
+ "category_id": 19,
+ "id": 2756,
+ "area": 9792,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1402,
+ "bbox": [
+ 216,
+ 279,
+ 34,
+ 12
+ ],
+ "category_id": 4,
+ "id": 2757,
+ "area": 408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1402,
+ "bbox": [
+ 297,
+ 330,
+ 14,
+ 28
+ ],
+ "category_id": 4,
+ "id": 2758,
+ "area": 392,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1402,
+ "bbox": [
+ 91,
+ 385,
+ 32,
+ 20
+ ],
+ "category_id": 4,
+ "id": 2759,
+ "area": 640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1402,
+ "bbox": [
+ 17,
+ 456,
+ 36,
+ 11
+ ],
+ "category_id": 4,
+ "id": 2760,
+ "area": 396,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1403,
+ "bbox": [
+ 152,
+ 257,
+ 213,
+ 59
+ ],
+ "category_id": 19,
+ "id": 2761,
+ "area": 12567,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1404,
+ "bbox": [
+ 61,
+ 204,
+ 402,
+ 41
+ ],
+ "category_id": 10,
+ "id": 2762,
+ "area": 16482,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1405,
+ "bbox": [
+ 37,
+ 172,
+ 370,
+ 173
+ ],
+ "category_id": 10,
+ "id": 2763,
+ "area": 64010,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1406,
+ "bbox": [
+ 222,
+ 104,
+ 126,
+ 163
+ ],
+ "category_id": 15,
+ "id": 2764,
+ "area": 20538,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1406,
+ "bbox": [
+ 211,
+ 250,
+ 133,
+ 160
+ ],
+ "category_id": 15,
+ "id": 2765,
+ "area": 21280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1407,
+ "bbox": [
+ 88,
+ 103,
+ 71,
+ 42
+ ],
+ "category_id": 16,
+ "id": 2766,
+ "area": 2982,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1408,
+ "bbox": [
+ 236,
+ 149,
+ 201,
+ 306
+ ],
+ "category_id": 10,
+ "id": 2767,
+ "area": 61506,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1409,
+ "bbox": [
+ 11,
+ 122,
+ 175,
+ 171
+ ],
+ "category_id": 19,
+ "id": 2768,
+ "area": 29925,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1410,
+ "bbox": [
+ 339,
+ 161,
+ 96,
+ 132
+ ],
+ "category_id": 19,
+ "id": 2769,
+ "area": 12672,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1411,
+ "bbox": [
+ 103,
+ 227,
+ 121,
+ 96
+ ],
+ "category_id": 15,
+ "id": 2770,
+ "area": 11616,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1411,
+ "bbox": [
+ 305,
+ 299,
+ 102,
+ 81
+ ],
+ "category_id": 15,
+ "id": 2771,
+ "area": 8262,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1412,
+ "bbox": [
+ 174,
+ 218,
+ 104,
+ 47
+ ],
+ "category_id": 16,
+ "id": 2772,
+ "area": 4888,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1413,
+ "bbox": [
+ 26,
+ 175,
+ 475,
+ 18
+ ],
+ "category_id": 10,
+ "id": 2773,
+ "area": 8550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1414,
+ "bbox": [
+ 92,
+ 198,
+ 59,
+ 27
+ ],
+ "category_id": 4,
+ "id": 2774,
+ "area": 1593,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1414,
+ "bbox": [
+ 348,
+ 168,
+ 59,
+ 30
+ ],
+ "category_id": 4,
+ "id": 2775,
+ "area": 1770,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1415,
+ "bbox": [
+ 154,
+ 80,
+ 192,
+ 299
+ ],
+ "category_id": 19,
+ "id": 2776,
+ "area": 57408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1416,
+ "bbox": [
+ 232,
+ 172,
+ 185,
+ 257
+ ],
+ "category_id": 10,
+ "id": 2777,
+ "area": 47545,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1417,
+ "bbox": [
+ 211,
+ 144,
+ 84,
+ 259
+ ],
+ "category_id": 19,
+ "id": 2778,
+ "area": 21756,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1418,
+ "bbox": [
+ 61,
+ 263,
+ 111,
+ 97
+ ],
+ "category_id": 15,
+ "id": 2779,
+ "area": 10767,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1418,
+ "bbox": [
+ 188,
+ 151,
+ 113,
+ 92
+ ],
+ "category_id": 15,
+ "id": 2780,
+ "area": 10396,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1418,
+ "bbox": [
+ 184,
+ 278,
+ 114,
+ 98
+ ],
+ "category_id": 15,
+ "id": 2781,
+ "area": 11172,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1419,
+ "bbox": [
+ 42,
+ 86,
+ 153,
+ 118
+ ],
+ "category_id": 19,
+ "id": 2782,
+ "area": 18054,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1420,
+ "bbox": [
+ 51,
+ 150,
+ 26,
+ 36
+ ],
+ "category_id": 4,
+ "id": 2783,
+ "area": 936,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1420,
+ "bbox": [
+ 84,
+ 335,
+ 28,
+ 34
+ ],
+ "category_id": 4,
+ "id": 2784,
+ "area": 952,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1420,
+ "bbox": [
+ 366,
+ 368,
+ 25,
+ 37
+ ],
+ "category_id": 4,
+ "id": 2785,
+ "area": 925,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1420,
+ "bbox": [
+ 443,
+ 49,
+ 23,
+ 32
+ ],
+ "category_id": 4,
+ "id": 2786,
+ "area": 736,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1421,
+ "bbox": [
+ 213,
+ 215,
+ 168,
+ 192
+ ],
+ "category_id": 15,
+ "id": 2787,
+ "area": 32256,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1421,
+ "bbox": [
+ 19,
+ 272,
+ 185,
+ 195
+ ],
+ "category_id": 15,
+ "id": 2788,
+ "area": 36075,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1422,
+ "bbox": [
+ 177,
+ 174,
+ 163,
+ 228
+ ],
+ "category_id": 10,
+ "id": 2789,
+ "area": 37164,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1423,
+ "bbox": [
+ 65,
+ 289,
+ 246,
+ 154
+ ],
+ "category_id": 19,
+ "id": 2790,
+ "area": 37884,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1424,
+ "bbox": [
+ 126,
+ 134,
+ 22,
+ 30
+ ],
+ "category_id": 4,
+ "id": 2791,
+ "area": 660,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1424,
+ "bbox": [
+ 122,
+ 298,
+ 22,
+ 24
+ ],
+ "category_id": 4,
+ "id": 2792,
+ "area": 528,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1424,
+ "bbox": [
+ 80,
+ 450,
+ 22,
+ 27
+ ],
+ "category_id": 4,
+ "id": 2793,
+ "area": 594,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1424,
+ "bbox": [
+ 338,
+ 190,
+ 22,
+ 28
+ ],
+ "category_id": 4,
+ "id": 2794,
+ "area": 616,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1424,
+ "bbox": [
+ 363,
+ 48,
+ 19,
+ 28
+ ],
+ "category_id": 4,
+ "id": 2795,
+ "area": 532,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1425,
+ "bbox": [
+ 316,
+ 364,
+ 128,
+ 78
+ ],
+ "category_id": 10,
+ "id": 2796,
+ "area": 9984,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1426,
+ "bbox": [
+ 271,
+ 66,
+ 22,
+ 50
+ ],
+ "category_id": 4,
+ "id": 2797,
+ "area": 1100,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1426,
+ "bbox": [
+ 234,
+ 418,
+ 14,
+ 43
+ ],
+ "category_id": 4,
+ "id": 2798,
+ "area": 602,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1427,
+ "bbox": [
+ 288,
+ 175,
+ 126,
+ 66
+ ],
+ "category_id": 16,
+ "id": 2799,
+ "area": 8316,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1428,
+ "bbox": [
+ 80,
+ 218,
+ 291,
+ 223
+ ],
+ "category_id": 16,
+ "id": 2800,
+ "area": 64893,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1429,
+ "bbox": [
+ 237,
+ 77,
+ 81,
+ 356
+ ],
+ "category_id": 19,
+ "id": 2801,
+ "area": 28836,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1430,
+ "bbox": [
+ 78,
+ 261,
+ 155,
+ 111
+ ],
+ "category_id": 15,
+ "id": 2802,
+ "area": 17205,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1430,
+ "bbox": [
+ 274,
+ 152,
+ 152,
+ 111
+ ],
+ "category_id": 15,
+ "id": 2803,
+ "area": 16872,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1431,
+ "bbox": [
+ 208,
+ 270,
+ 199,
+ 26
+ ],
+ "category_id": 10,
+ "id": 2804,
+ "area": 5174,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1432,
+ "bbox": [
+ 288,
+ 39,
+ 18,
+ 12
+ ],
+ "category_id": 4,
+ "id": 2805,
+ "area": 216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1432,
+ "bbox": [
+ 368,
+ 135,
+ 15,
+ 11
+ ],
+ "category_id": 4,
+ "id": 2806,
+ "area": 165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1432,
+ "bbox": [
+ 13,
+ 321,
+ 19,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2807,
+ "area": 190,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1433,
+ "bbox": [
+ 40,
+ 72,
+ 35,
+ 25
+ ],
+ "category_id": 4,
+ "id": 2808,
+ "area": 875,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1433,
+ "bbox": [
+ 226,
+ 81,
+ 31,
+ 21
+ ],
+ "category_id": 4,
+ "id": 2809,
+ "area": 651,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1433,
+ "bbox": [
+ 362,
+ 229,
+ 25,
+ 17
+ ],
+ "category_id": 4,
+ "id": 2810,
+ "area": 425,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1433,
+ "bbox": [
+ 440,
+ 360,
+ 25,
+ 19
+ ],
+ "category_id": 4,
+ "id": 2811,
+ "area": 475,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1433,
+ "bbox": [
+ 312,
+ 448,
+ 27,
+ 16
+ ],
+ "category_id": 4,
+ "id": 2812,
+ "area": 432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1434,
+ "bbox": [
+ 105,
+ 124,
+ 279,
+ 297
+ ],
+ "category_id": 15,
+ "id": 2813,
+ "area": 82863,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1435,
+ "bbox": [
+ 43,
+ 165,
+ 440,
+ 110
+ ],
+ "category_id": 19,
+ "id": 2814,
+ "area": 48400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1436,
+ "bbox": [
+ 185,
+ 129,
+ 52,
+ 258
+ ],
+ "category_id": 10,
+ "id": 2815,
+ "area": 13416,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1437,
+ "bbox": [
+ 182,
+ 30,
+ 88,
+ 356
+ ],
+ "category_id": 19,
+ "id": 2816,
+ "area": 31328,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1438,
+ "bbox": [
+ 238,
+ 304,
+ 151,
+ 90
+ ],
+ "category_id": 19,
+ "id": 2817,
+ "area": 13590,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1439,
+ "bbox": [
+ 107,
+ 251,
+ 281,
+ 44
+ ],
+ "category_id": 19,
+ "id": 2818,
+ "area": 12364,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1440,
+ "bbox": [
+ 337,
+ 349,
+ 58,
+ 154
+ ],
+ "category_id": 15,
+ "id": 2819,
+ "area": 8932,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1440,
+ "bbox": [
+ 136,
+ 224,
+ 118,
+ 150
+ ],
+ "category_id": 15,
+ "id": 2820,
+ "area": 17700,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1440,
+ "bbox": [
+ 13,
+ 304,
+ 117,
+ 151
+ ],
+ "category_id": 15,
+ "id": 2821,
+ "area": 17667,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1440,
+ "bbox": [
+ 263,
+ 143,
+ 115,
+ 150
+ ],
+ "category_id": 15,
+ "id": 2822,
+ "area": 17250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1440,
+ "bbox": [
+ 389,
+ 61,
+ 116,
+ 153
+ ],
+ "category_id": 15,
+ "id": 2823,
+ "area": 17748,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1441,
+ "bbox": [
+ 60,
+ 167,
+ 122,
+ 101
+ ],
+ "category_id": 15,
+ "id": 2824,
+ "area": 12322,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1441,
+ "bbox": [
+ 56,
+ 310,
+ 118,
+ 102
+ ],
+ "category_id": 15,
+ "id": 2825,
+ "area": 12036,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1441,
+ "bbox": [
+ 208,
+ 316,
+ 123,
+ 100
+ ],
+ "category_id": 15,
+ "id": 2826,
+ "area": 12300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1441,
+ "bbox": [
+ 352,
+ 322,
+ 124,
+ 101
+ ],
+ "category_id": 15,
+ "id": 2827,
+ "area": 12524,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1442,
+ "bbox": [
+ 194,
+ 49,
+ 134,
+ 333
+ ],
+ "category_id": 10,
+ "id": 2828,
+ "area": 44622,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1443,
+ "bbox": [
+ 268,
+ 33,
+ 25,
+ 15
+ ],
+ "category_id": 4,
+ "id": 2829,
+ "area": 375,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1443,
+ "bbox": [
+ 273,
+ 168,
+ 27,
+ 16
+ ],
+ "category_id": 4,
+ "id": 2830,
+ "area": 432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1443,
+ "bbox": [
+ 322,
+ 291,
+ 30,
+ 18
+ ],
+ "category_id": 4,
+ "id": 2831,
+ "area": 540,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1443,
+ "bbox": [
+ 320,
+ 419,
+ 42,
+ 14
+ ],
+ "category_id": 4,
+ "id": 2832,
+ "area": 588,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1443,
+ "bbox": [
+ 24,
+ 440,
+ 24,
+ 19
+ ],
+ "category_id": 4,
+ "id": 2833,
+ "area": 456,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1444,
+ "bbox": [
+ 51,
+ 37,
+ 52,
+ 267
+ ],
+ "category_id": 10,
+ "id": 2834,
+ "area": 13884,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1445,
+ "bbox": [
+ 165,
+ 110,
+ 268,
+ 206
+ ],
+ "category_id": 16,
+ "id": 2835,
+ "area": 55208,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1446,
+ "bbox": [
+ 63,
+ 157,
+ 63,
+ 58
+ ],
+ "category_id": 4,
+ "id": 2836,
+ "area": 3654,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1446,
+ "bbox": [
+ 58,
+ 279,
+ 26,
+ 45
+ ],
+ "category_id": 4,
+ "id": 2837,
+ "area": 1170,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1446,
+ "bbox": [
+ 353,
+ 125,
+ 39,
+ 40
+ ],
+ "category_id": 4,
+ "id": 2838,
+ "area": 1560,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1447,
+ "bbox": [
+ 161,
+ 256,
+ 234,
+ 182
+ ],
+ "category_id": 16,
+ "id": 2839,
+ "area": 42588,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1448,
+ "bbox": [
+ 225,
+ 235,
+ 93,
+ 115
+ ],
+ "category_id": 15,
+ "id": 2840,
+ "area": 10695,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1449,
+ "bbox": [
+ 40,
+ 39,
+ 111,
+ 439
+ ],
+ "category_id": 19,
+ "id": 2841,
+ "area": 48729,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1450,
+ "bbox": [
+ 142,
+ 3,
+ 32,
+ 14
+ ],
+ "category_id": 4,
+ "id": 2842,
+ "area": 448,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1450,
+ "bbox": [
+ 352,
+ 36,
+ 30,
+ 15
+ ],
+ "category_id": 4,
+ "id": 2843,
+ "area": 450,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1450,
+ "bbox": [
+ 247,
+ 215,
+ 26,
+ 18
+ ],
+ "category_id": 4,
+ "id": 2844,
+ "area": 468,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1450,
+ "bbox": [
+ 411,
+ 232,
+ 30,
+ 13
+ ],
+ "category_id": 4,
+ "id": 2845,
+ "area": 390,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1450,
+ "bbox": [
+ 207,
+ 337,
+ 26,
+ 19
+ ],
+ "category_id": 4,
+ "id": 2846,
+ "area": 494,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1450,
+ "bbox": [
+ 445,
+ 464,
+ 33,
+ 12
+ ],
+ "category_id": 4,
+ "id": 2847,
+ "area": 396,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1450,
+ "bbox": [
+ 257,
+ 465,
+ 28,
+ 22
+ ],
+ "category_id": 4,
+ "id": 2848,
+ "area": 616,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1450,
+ "bbox": [
+ 121,
+ 492,
+ 30,
+ 16
+ ],
+ "category_id": 4,
+ "id": 2849,
+ "area": 480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1451,
+ "bbox": [
+ 23,
+ 56,
+ 56,
+ 34
+ ],
+ "category_id": 4,
+ "id": 2850,
+ "area": 1904,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1451,
+ "bbox": [
+ 108,
+ 429,
+ 56,
+ 31
+ ],
+ "category_id": 4,
+ "id": 2851,
+ "area": 1736,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1451,
+ "bbox": [
+ 449,
+ 464,
+ 49,
+ 24
+ ],
+ "category_id": 4,
+ "id": 2852,
+ "area": 1176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1452,
+ "bbox": [
+ 44,
+ 72,
+ 404,
+ 405
+ ],
+ "category_id": 16,
+ "id": 2853,
+ "area": 163620,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1453,
+ "bbox": [
+ 71,
+ 297,
+ 361,
+ 87
+ ],
+ "category_id": 16,
+ "id": 2854,
+ "area": 31407,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1454,
+ "bbox": [
+ 157,
+ 168,
+ 134,
+ 146
+ ],
+ "category_id": 19,
+ "id": 2855,
+ "area": 19564,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1455,
+ "bbox": [
+ 142,
+ 115,
+ 194,
+ 278
+ ],
+ "category_id": 16,
+ "id": 2856,
+ "area": 53932,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1456,
+ "bbox": [
+ 42,
+ 300,
+ 39,
+ 37
+ ],
+ "category_id": 4,
+ "id": 2857,
+ "area": 1443,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1456,
+ "bbox": [
+ 449,
+ 179,
+ 39,
+ 35
+ ],
+ "category_id": 4,
+ "id": 2858,
+ "area": 1365,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1457,
+ "bbox": [
+ 23,
+ 67,
+ 163,
+ 341
+ ],
+ "category_id": 19,
+ "id": 2859,
+ "area": 55583,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1458,
+ "bbox": [
+ 73,
+ 231,
+ 15,
+ 15
+ ],
+ "category_id": 4,
+ "id": 2860,
+ "area": 225,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1458,
+ "bbox": [
+ 109,
+ 38,
+ 16,
+ 11
+ ],
+ "category_id": 4,
+ "id": 2861,
+ "area": 176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1458,
+ "bbox": [
+ 282,
+ 44,
+ 14,
+ 12
+ ],
+ "category_id": 4,
+ "id": 2862,
+ "area": 168,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1458,
+ "bbox": [
+ 343,
+ 215,
+ 17,
+ 14
+ ],
+ "category_id": 4,
+ "id": 2863,
+ "area": 238,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1458,
+ "bbox": [
+ 455,
+ 47,
+ 17,
+ 11
+ ],
+ "category_id": 4,
+ "id": 2864,
+ "area": 187,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1458,
+ "bbox": [
+ 282,
+ 433,
+ 19,
+ 13
+ ],
+ "category_id": 4,
+ "id": 2865,
+ "area": 247,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1459,
+ "bbox": [
+ 206,
+ 172,
+ 125,
+ 106
+ ],
+ "category_id": 16,
+ "id": 2866,
+ "area": 13250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1460,
+ "bbox": [
+ 210,
+ 128,
+ 241,
+ 311
+ ],
+ "category_id": 10,
+ "id": 2867,
+ "area": 74951,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1461,
+ "bbox": [
+ 136,
+ 24,
+ 97,
+ 146
+ ],
+ "category_id": 15,
+ "id": 2868,
+ "area": 14162,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1461,
+ "bbox": [
+ 101,
+ 195,
+ 129,
+ 185
+ ],
+ "category_id": 15,
+ "id": 2869,
+ "area": 23865,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1462,
+ "bbox": [
+ 238,
+ 263,
+ 160,
+ 47
+ ],
+ "category_id": 19,
+ "id": 2870,
+ "area": 7520,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1463,
+ "bbox": [
+ 145,
+ 78,
+ 105,
+ 211
+ ],
+ "category_id": 10,
+ "id": 2871,
+ "area": 22155,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1464,
+ "bbox": [
+ 120,
+ 220,
+ 132,
+ 60
+ ],
+ "category_id": 19,
+ "id": 2872,
+ "area": 7920,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1465,
+ "bbox": [
+ 141,
+ 52,
+ 162,
+ 70
+ ],
+ "category_id": 19,
+ "id": 2873,
+ "area": 11340,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1466,
+ "bbox": [
+ 62,
+ 204,
+ 384,
+ 65
+ ],
+ "category_id": 19,
+ "id": 2874,
+ "area": 24960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1467,
+ "bbox": [
+ 199,
+ 273,
+ 145,
+ 129
+ ],
+ "category_id": 16,
+ "id": 2875,
+ "area": 18705,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1468,
+ "bbox": [
+ 120,
+ 120,
+ 170,
+ 310
+ ],
+ "category_id": 16,
+ "id": 2876,
+ "area": 52700,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1469,
+ "bbox": [
+ 61,
+ 196,
+ 105,
+ 175
+ ],
+ "category_id": 19,
+ "id": 2877,
+ "area": 18375,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1470,
+ "bbox": [
+ 186,
+ 188,
+ 94,
+ 181
+ ],
+ "category_id": 19,
+ "id": 2878,
+ "area": 17014,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1471,
+ "bbox": [
+ 24,
+ 167,
+ 51,
+ 47
+ ],
+ "category_id": 4,
+ "id": 2879,
+ "area": 2397,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1471,
+ "bbox": [
+ 471,
+ 376,
+ 41,
+ 50
+ ],
+ "category_id": 4,
+ "id": 2880,
+ "area": 2050,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1472,
+ "bbox": [
+ 9,
+ 301,
+ 24,
+ 18
+ ],
+ "category_id": 4,
+ "id": 2881,
+ "area": 432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1472,
+ "bbox": [
+ 368,
+ 364,
+ 19,
+ 9
+ ],
+ "category_id": 4,
+ "id": 2882,
+ "area": 171,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1472,
+ "bbox": [
+ 459,
+ 334,
+ 31,
+ 8
+ ],
+ "category_id": 4,
+ "id": 2883,
+ "area": 248,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1472,
+ "bbox": [
+ 487,
+ 263,
+ 19,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2884,
+ "area": 190,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1473,
+ "bbox": [
+ 115,
+ 113,
+ 189,
+ 318
+ ],
+ "category_id": 16,
+ "id": 2885,
+ "area": 60102,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1474,
+ "bbox": [
+ 241,
+ 128,
+ 35,
+ 197
+ ],
+ "category_id": 19,
+ "id": 2886,
+ "area": 6895,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1475,
+ "bbox": [
+ 61,
+ 142,
+ 11,
+ 31
+ ],
+ "category_id": 4,
+ "id": 2887,
+ "area": 341,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1475,
+ "bbox": [
+ 396,
+ 65,
+ 16,
+ 36
+ ],
+ "category_id": 4,
+ "id": 2888,
+ "area": 576,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1475,
+ "bbox": [
+ 424,
+ 380,
+ 16,
+ 34
+ ],
+ "category_id": 4,
+ "id": 2889,
+ "area": 544,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1476,
+ "bbox": [
+ 143,
+ 241,
+ 267,
+ 36
+ ],
+ "category_id": 19,
+ "id": 2890,
+ "area": 9612,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1477,
+ "bbox": [
+ 109,
+ 217,
+ 118,
+ 82
+ ],
+ "category_id": 15,
+ "id": 2891,
+ "area": 9676,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1477,
+ "bbox": [
+ 344,
+ 236,
+ 116,
+ 77
+ ],
+ "category_id": 15,
+ "id": 2892,
+ "area": 8932,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1478,
+ "bbox": [
+ 115,
+ 20,
+ 23,
+ 17
+ ],
+ "category_id": 4,
+ "id": 2893,
+ "area": 391,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1478,
+ "bbox": [
+ 10,
+ 103,
+ 32,
+ 17
+ ],
+ "category_id": 4,
+ "id": 2894,
+ "area": 544,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1478,
+ "bbox": [
+ 144,
+ 241,
+ 22,
+ 16
+ ],
+ "category_id": 4,
+ "id": 2895,
+ "area": 352,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1478,
+ "bbox": [
+ 61,
+ 373,
+ 20,
+ 16
+ ],
+ "category_id": 4,
+ "id": 2896,
+ "area": 320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1478,
+ "bbox": [
+ 321,
+ 110,
+ 28,
+ 15
+ ],
+ "category_id": 4,
+ "id": 2897,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1478,
+ "bbox": [
+ 481,
+ 51,
+ 31,
+ 18
+ ],
+ "category_id": 4,
+ "id": 2898,
+ "area": 558,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1478,
+ "bbox": [
+ 417,
+ 347,
+ 27,
+ 21
+ ],
+ "category_id": 4,
+ "id": 2899,
+ "area": 567,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1479,
+ "bbox": [
+ 167,
+ 40,
+ 19,
+ 23
+ ],
+ "category_id": 4,
+ "id": 2900,
+ "area": 437,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1479,
+ "bbox": [
+ 135,
+ 151,
+ 13,
+ 24
+ ],
+ "category_id": 4,
+ "id": 2901,
+ "area": 312,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1479,
+ "bbox": [
+ 101,
+ 358,
+ 21,
+ 24
+ ],
+ "category_id": 4,
+ "id": 2902,
+ "area": 504,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1479,
+ "bbox": [
+ 83,
+ 481,
+ 12,
+ 25
+ ],
+ "category_id": 4,
+ "id": 2903,
+ "area": 300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1479,
+ "bbox": [
+ 288,
+ 291,
+ 30,
+ 11
+ ],
+ "category_id": 4,
+ "id": 2904,
+ "area": 330,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1479,
+ "bbox": [
+ 433,
+ 196,
+ 13,
+ 27
+ ],
+ "category_id": 4,
+ "id": 2905,
+ "area": 351,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1480,
+ "bbox": [
+ 209,
+ 56,
+ 96,
+ 412
+ ],
+ "category_id": 10,
+ "id": 2906,
+ "area": 39552,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1481,
+ "bbox": [
+ 236,
+ 84,
+ 46,
+ 39
+ ],
+ "category_id": 4,
+ "id": 2907,
+ "area": 1794,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1481,
+ "bbox": [
+ 168,
+ 387,
+ 43,
+ 34
+ ],
+ "category_id": 4,
+ "id": 2908,
+ "area": 1462,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1481,
+ "bbox": [
+ 390,
+ 417,
+ 39,
+ 37
+ ],
+ "category_id": 4,
+ "id": 2909,
+ "area": 1443,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1482,
+ "bbox": [
+ 102,
+ 27,
+ 24,
+ 15
+ ],
+ "category_id": 4,
+ "id": 2910,
+ "area": 360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1482,
+ "bbox": [
+ 222,
+ 132,
+ 23,
+ 19
+ ],
+ "category_id": 4,
+ "id": 2911,
+ "area": 437,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1482,
+ "bbox": [
+ 389,
+ 102,
+ 23,
+ 20
+ ],
+ "category_id": 4,
+ "id": 2912,
+ "area": 460,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1482,
+ "bbox": [
+ 338,
+ 290,
+ 33,
+ 19
+ ],
+ "category_id": 4,
+ "id": 2913,
+ "area": 627,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1482,
+ "bbox": [
+ 312,
+ 414,
+ 20,
+ 21
+ ],
+ "category_id": 4,
+ "id": 2914,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1482,
+ "bbox": [
+ 62,
+ 414,
+ 20,
+ 23
+ ],
+ "category_id": 4,
+ "id": 2915,
+ "area": 460,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1483,
+ "bbox": [
+ 105,
+ 172,
+ 103,
+ 109
+ ],
+ "category_id": 19,
+ "id": 2916,
+ "area": 11227,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1484,
+ "bbox": [
+ 44,
+ 105,
+ 20,
+ 22
+ ],
+ "category_id": 4,
+ "id": 2917,
+ "area": 440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1484,
+ "bbox": [
+ 177,
+ 72,
+ 27,
+ 18
+ ],
+ "category_id": 4,
+ "id": 2918,
+ "area": 486,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1484,
+ "bbox": [
+ 154,
+ 176,
+ 21,
+ 19
+ ],
+ "category_id": 4,
+ "id": 2919,
+ "area": 399,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1484,
+ "bbox": [
+ 440,
+ 90,
+ 12,
+ 27
+ ],
+ "category_id": 4,
+ "id": 2920,
+ "area": 324,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1484,
+ "bbox": [
+ 419,
+ 248,
+ 14,
+ 24
+ ],
+ "category_id": 4,
+ "id": 2921,
+ "area": 336,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1484,
+ "bbox": [
+ 282,
+ 410,
+ 14,
+ 27
+ ],
+ "category_id": 4,
+ "id": 2922,
+ "area": 378,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1485,
+ "bbox": [
+ 99,
+ 197,
+ 27,
+ 219
+ ],
+ "category_id": 10,
+ "id": 2923,
+ "area": 5913,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1486,
+ "bbox": [
+ 204,
+ 259,
+ 91,
+ 93
+ ],
+ "category_id": 15,
+ "id": 2924,
+ "area": 8463,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1487,
+ "bbox": [
+ 125,
+ 143,
+ 142,
+ 127
+ ],
+ "category_id": 15,
+ "id": 2925,
+ "area": 18034,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1487,
+ "bbox": [
+ 268,
+ 229,
+ 144,
+ 126
+ ],
+ "category_id": 15,
+ "id": 2926,
+ "area": 18144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1488,
+ "bbox": [
+ 119,
+ 142,
+ 65,
+ 258
+ ],
+ "category_id": 19,
+ "id": 2927,
+ "area": 16770,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1489,
+ "bbox": [
+ 122,
+ 205,
+ 274,
+ 198
+ ],
+ "category_id": 19,
+ "id": 2928,
+ "area": 54252,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1490,
+ "bbox": [
+ 193,
+ 246,
+ 71,
+ 51
+ ],
+ "category_id": 4,
+ "id": 2929,
+ "area": 3621,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1491,
+ "bbox": [
+ 224,
+ 44,
+ 98,
+ 36
+ ],
+ "category_id": 10,
+ "id": 2930,
+ "area": 3528,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1491,
+ "bbox": [
+ 265,
+ 89,
+ 63,
+ 26
+ ],
+ "category_id": 10,
+ "id": 2931,
+ "area": 1638,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1492,
+ "bbox": [
+ 233,
+ 52,
+ 160,
+ 215
+ ],
+ "category_id": 15,
+ "id": 2932,
+ "area": 34400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1492,
+ "bbox": [
+ 135,
+ 229,
+ 162,
+ 215
+ ],
+ "category_id": 15,
+ "id": 2933,
+ "area": 34830,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1493,
+ "bbox": [
+ 206,
+ 168,
+ 105,
+ 127
+ ],
+ "category_id": 19,
+ "id": 2934,
+ "area": 13335,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1494,
+ "bbox": [
+ 102,
+ 199,
+ 145,
+ 112
+ ],
+ "category_id": 19,
+ "id": 2935,
+ "area": 16240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1495,
+ "bbox": [
+ 154,
+ 157,
+ 268,
+ 125
+ ],
+ "category_id": 19,
+ "id": 2936,
+ "area": 33500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1496,
+ "bbox": [
+ 296,
+ 112,
+ 198,
+ 376
+ ],
+ "category_id": 19,
+ "id": 2937,
+ "area": 74448,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1497,
+ "bbox": [
+ 108,
+ 327,
+ 117,
+ 17
+ ],
+ "category_id": 19,
+ "id": 2938,
+ "area": 1989,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1498,
+ "bbox": [
+ 186,
+ 223,
+ 168,
+ 91
+ ],
+ "category_id": 10,
+ "id": 2939,
+ "area": 15288,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1499,
+ "bbox": [
+ 318,
+ 138,
+ 49,
+ 14
+ ],
+ "category_id": 15,
+ "id": 2940,
+ "area": 686,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1500,
+ "bbox": [
+ 161,
+ 37,
+ 47,
+ 30
+ ],
+ "category_id": 4,
+ "id": 2941,
+ "area": 1410,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1500,
+ "bbox": [
+ 51,
+ 131,
+ 28,
+ 38
+ ],
+ "category_id": 4,
+ "id": 2942,
+ "area": 1064,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1500,
+ "bbox": [
+ 302,
+ 205,
+ 31,
+ 36
+ ],
+ "category_id": 4,
+ "id": 2943,
+ "area": 1116,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1500,
+ "bbox": [
+ 280,
+ 371,
+ 26,
+ 36
+ ],
+ "category_id": 4,
+ "id": 2944,
+ "area": 936,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1501,
+ "bbox": [
+ 74,
+ 377,
+ 113,
+ 80
+ ],
+ "category_id": 10,
+ "id": 2945,
+ "area": 9040,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1502,
+ "bbox": [
+ 248,
+ 234,
+ 122,
+ 132
+ ],
+ "category_id": 16,
+ "id": 2946,
+ "area": 16104,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1503,
+ "bbox": [
+ 29,
+ 382,
+ 11,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2947,
+ "area": 110,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1503,
+ "bbox": [
+ 64,
+ 248,
+ 8,
+ 8
+ ],
+ "category_id": 4,
+ "id": 2948,
+ "area": 64,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1503,
+ "bbox": [
+ 65,
+ 188,
+ 10,
+ 8
+ ],
+ "category_id": 4,
+ "id": 2949,
+ "area": 80,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1503,
+ "bbox": [
+ 197,
+ 127,
+ 8,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2950,
+ "area": 80,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1503,
+ "bbox": [
+ 197,
+ 172,
+ 10,
+ 7
+ ],
+ "category_id": 4,
+ "id": 2951,
+ "area": 70,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1503,
+ "bbox": [
+ 202,
+ 219,
+ 11,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2952,
+ "area": 110,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1503,
+ "bbox": [
+ 129,
+ 293,
+ 8,
+ 5
+ ],
+ "category_id": 4,
+ "id": 2953,
+ "area": 40,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1503,
+ "bbox": [
+ 334,
+ 165,
+ 8,
+ 8
+ ],
+ "category_id": 4,
+ "id": 2954,
+ "area": 64,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1503,
+ "bbox": [
+ 501,
+ 157,
+ 7,
+ 8
+ ],
+ "category_id": 4,
+ "id": 2955,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1504,
+ "bbox": [
+ 151,
+ 60,
+ 228,
+ 356
+ ],
+ "category_id": 10,
+ "id": 2956,
+ "area": 81168,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1505,
+ "bbox": [
+ 42,
+ 177,
+ 16,
+ 13
+ ],
+ "category_id": 4,
+ "id": 2957,
+ "area": 208,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1505,
+ "bbox": [
+ 236,
+ 12,
+ 15,
+ 11
+ ],
+ "category_id": 4,
+ "id": 2958,
+ "area": 165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1505,
+ "bbox": [
+ 272,
+ 140,
+ 19,
+ 12
+ ],
+ "category_id": 4,
+ "id": 2959,
+ "area": 228,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1505,
+ "bbox": [
+ 195,
+ 206,
+ 22,
+ 12
+ ],
+ "category_id": 4,
+ "id": 2960,
+ "area": 264,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1505,
+ "bbox": [
+ 37,
+ 272,
+ 17,
+ 8
+ ],
+ "category_id": 4,
+ "id": 2961,
+ "area": 136,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1505,
+ "bbox": [
+ 334,
+ 272,
+ 20,
+ 10
+ ],
+ "category_id": 4,
+ "id": 2962,
+ "area": 200,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1505,
+ "bbox": [
+ 116,
+ 410,
+ 13,
+ 11
+ ],
+ "category_id": 4,
+ "id": 2963,
+ "area": 143,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1505,
+ "bbox": [
+ 312,
+ 389,
+ 18,
+ 14
+ ],
+ "category_id": 4,
+ "id": 2964,
+ "area": 252,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1505,
+ "bbox": [
+ 410,
+ 462,
+ 16,
+ 15
+ ],
+ "category_id": 4,
+ "id": 2965,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1505,
+ "bbox": [
+ 236,
+ 496,
+ 15,
+ 13
+ ],
+ "category_id": 4,
+ "id": 2966,
+ "area": 195,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1506,
+ "bbox": [
+ 135,
+ 183,
+ 336,
+ 294
+ ],
+ "category_id": 10,
+ "id": 2967,
+ "area": 98784,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1507,
+ "bbox": [
+ 141,
+ 173,
+ 100,
+ 226
+ ],
+ "category_id": 16,
+ "id": 2968,
+ "area": 22600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1508,
+ "bbox": [
+ 106,
+ 126,
+ 185,
+ 112
+ ],
+ "category_id": 19,
+ "id": 2969,
+ "area": 20720,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1509,
+ "bbox": [
+ 142,
+ 137,
+ 259,
+ 288
+ ],
+ "category_id": 10,
+ "id": 2970,
+ "area": 74592,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1510,
+ "bbox": [
+ 67,
+ 359,
+ 155,
+ 88
+ ],
+ "category_id": 10,
+ "id": 2971,
+ "area": 13640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1511,
+ "bbox": [
+ 196,
+ 139,
+ 46,
+ 102
+ ],
+ "category_id": 19,
+ "id": 2972,
+ "area": 4692,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1512,
+ "bbox": [
+ 92,
+ 13,
+ 64,
+ 124
+ ],
+ "category_id": 10,
+ "id": 2973,
+ "area": 7936,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1513,
+ "bbox": [
+ 204,
+ 22,
+ 278,
+ 448
+ ],
+ "category_id": 10,
+ "id": 2974,
+ "area": 124544,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1514,
+ "bbox": [
+ 152,
+ 120,
+ 91,
+ 326
+ ],
+ "category_id": 16,
+ "id": 2975,
+ "area": 29666,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1515,
+ "bbox": [
+ 133,
+ 192,
+ 172,
+ 95
+ ],
+ "category_id": 19,
+ "id": 2976,
+ "area": 16340,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1516,
+ "bbox": [
+ 52,
+ 34,
+ 40,
+ 179
+ ],
+ "category_id": 10,
+ "id": 2977,
+ "area": 7160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1517,
+ "bbox": [
+ 70,
+ 35,
+ 203,
+ 200
+ ],
+ "category_id": 10,
+ "id": 2978,
+ "area": 40600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1518,
+ "bbox": [
+ 140,
+ 15,
+ 44,
+ 448
+ ],
+ "category_id": 10,
+ "id": 2979,
+ "area": 19712,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1519,
+ "bbox": [
+ 147,
+ 74,
+ 10,
+ 27
+ ],
+ "category_id": 16,
+ "id": 2980,
+ "area": 270,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1519,
+ "bbox": [
+ 413,
+ 417,
+ 15,
+ 28
+ ],
+ "category_id": 16,
+ "id": 2981,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1520,
+ "bbox": [
+ 222,
+ 190,
+ 187,
+ 135
+ ],
+ "category_id": 15,
+ "id": 2982,
+ "area": 25245,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1520,
+ "bbox": [
+ 94,
+ 320,
+ 196,
+ 128
+ ],
+ "category_id": 15,
+ "id": 2983,
+ "area": 25088,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1521,
+ "bbox": [
+ 168,
+ 269,
+ 133,
+ 107
+ ],
+ "category_id": 4,
+ "id": 2984,
+ "area": 14231,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1522,
+ "bbox": [
+ 169,
+ 101,
+ 263,
+ 221
+ ],
+ "category_id": 15,
+ "id": 2985,
+ "area": 58123,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1523,
+ "bbox": [
+ 22,
+ 140,
+ 15,
+ 14
+ ],
+ "category_id": 4,
+ "id": 2986,
+ "area": 210,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1523,
+ "bbox": [
+ 284,
+ 177,
+ 16,
+ 12
+ ],
+ "category_id": 4,
+ "id": 2987,
+ "area": 192,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1523,
+ "bbox": [
+ 48,
+ 301,
+ 17,
+ 11
+ ],
+ "category_id": 4,
+ "id": 2988,
+ "area": 187,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1523,
+ "bbox": [
+ 195,
+ 359,
+ 18,
+ 11
+ ],
+ "category_id": 4,
+ "id": 2989,
+ "area": 198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1524,
+ "bbox": [
+ 204,
+ 160,
+ 113,
+ 220
+ ],
+ "category_id": 19,
+ "id": 2990,
+ "area": 24860,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1525,
+ "bbox": [
+ 117,
+ 162,
+ 274,
+ 202
+ ],
+ "category_id": 16,
+ "id": 2991,
+ "area": 55348,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1526,
+ "bbox": [
+ 85,
+ 5,
+ 25,
+ 25
+ ],
+ "category_id": 4,
+ "id": 2992,
+ "area": 625,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1526,
+ "bbox": [
+ 67,
+ 126,
+ 23,
+ 23
+ ],
+ "category_id": 4,
+ "id": 2993,
+ "area": 529,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1526,
+ "bbox": [
+ 82,
+ 245,
+ 22,
+ 23
+ ],
+ "category_id": 4,
+ "id": 2994,
+ "area": 506,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1526,
+ "bbox": [
+ 90,
+ 364,
+ 23,
+ 27
+ ],
+ "category_id": 4,
+ "id": 2995,
+ "area": 621,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1526,
+ "bbox": [
+ 79,
+ 478,
+ 24,
+ 23
+ ],
+ "category_id": 4,
+ "id": 2996,
+ "area": 552,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1526,
+ "bbox": [
+ 405,
+ 70,
+ 21,
+ 24
+ ],
+ "category_id": 4,
+ "id": 2997,
+ "area": 504,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1526,
+ "bbox": [
+ 392,
+ 187,
+ 24,
+ 27
+ ],
+ "category_id": 4,
+ "id": 2998,
+ "area": 648,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1526,
+ "bbox": [
+ 389,
+ 305,
+ 23,
+ 27
+ ],
+ "category_id": 4,
+ "id": 2999,
+ "area": 621,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1526,
+ "bbox": [
+ 384,
+ 424,
+ 30,
+ 26
+ ],
+ "category_id": 4,
+ "id": 3000,
+ "area": 780,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1527,
+ "bbox": [
+ 204,
+ 76,
+ 26,
+ 23
+ ],
+ "category_id": 4,
+ "id": 3001,
+ "area": 598,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1527,
+ "bbox": [
+ 293,
+ 238,
+ 22,
+ 23
+ ],
+ "category_id": 4,
+ "id": 3002,
+ "area": 506,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1527,
+ "bbox": [
+ 300,
+ 463,
+ 24,
+ 26
+ ],
+ "category_id": 4,
+ "id": 3003,
+ "area": 624,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1528,
+ "bbox": [
+ 167,
+ 219,
+ 142,
+ 61
+ ],
+ "category_id": 4,
+ "id": 3004,
+ "area": 8662,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1529,
+ "bbox": [
+ 172,
+ 124,
+ 250,
+ 196
+ ],
+ "category_id": 16,
+ "id": 3005,
+ "area": 49000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1530,
+ "bbox": [
+ 71,
+ 97,
+ 17,
+ 68
+ ],
+ "category_id": 4,
+ "id": 3006,
+ "area": 1156,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1530,
+ "bbox": [
+ 392,
+ 331,
+ 23,
+ 67
+ ],
+ "category_id": 4,
+ "id": 3007,
+ "area": 1541,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1531,
+ "bbox": [
+ 152,
+ 298,
+ 316,
+ 41
+ ],
+ "category_id": 10,
+ "id": 3008,
+ "area": 12956,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1532,
+ "bbox": [
+ 179,
+ 266,
+ 55,
+ 105
+ ],
+ "category_id": 15,
+ "id": 3009,
+ "area": 5775,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1532,
+ "bbox": [
+ 274,
+ 252,
+ 109,
+ 124
+ ],
+ "category_id": 15,
+ "id": 3010,
+ "area": 13516,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1533,
+ "bbox": [
+ 407,
+ 195,
+ 21,
+ 20
+ ],
+ "category_id": 4,
+ "id": 3011,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1533,
+ "bbox": [
+ 197,
+ 174,
+ 18,
+ 17
+ ],
+ "category_id": 4,
+ "id": 3012,
+ "area": 306,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1533,
+ "bbox": [
+ 168,
+ 320,
+ 18,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3013,
+ "area": 378,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1534,
+ "bbox": [
+ 135,
+ 244,
+ 217,
+ 142
+ ],
+ "category_id": 16,
+ "id": 3014,
+ "area": 30814,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1535,
+ "bbox": [
+ 104,
+ 403,
+ 51,
+ 61
+ ],
+ "category_id": 4,
+ "id": 3015,
+ "area": 3111,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1535,
+ "bbox": [
+ 255,
+ 101,
+ 41,
+ 39
+ ],
+ "category_id": 4,
+ "id": 3016,
+ "area": 1599,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1536,
+ "bbox": [
+ 138,
+ 236,
+ 253,
+ 145
+ ],
+ "category_id": 16,
+ "id": 3017,
+ "area": 36685,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1537,
+ "bbox": [
+ 225,
+ 48,
+ 16,
+ 19
+ ],
+ "category_id": 4,
+ "id": 3018,
+ "area": 304,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1537,
+ "bbox": [
+ 48,
+ 198,
+ 17,
+ 16
+ ],
+ "category_id": 4,
+ "id": 3019,
+ "area": 272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1537,
+ "bbox": [
+ 151,
+ 354,
+ 17,
+ 20
+ ],
+ "category_id": 4,
+ "id": 3020,
+ "area": 340,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1537,
+ "bbox": [
+ 156,
+ 467,
+ 18,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3021,
+ "area": 378,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1537,
+ "bbox": [
+ 334,
+ 350,
+ 19,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3022,
+ "area": 399,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1538,
+ "bbox": [
+ 126,
+ 94,
+ 323,
+ 264
+ ],
+ "category_id": 10,
+ "id": 3023,
+ "area": 85272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1539,
+ "bbox": [
+ 119,
+ 110,
+ 314,
+ 133
+ ],
+ "category_id": 16,
+ "id": 3024,
+ "area": 41762,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1540,
+ "bbox": [
+ 148,
+ 144,
+ 46,
+ 211
+ ],
+ "category_id": 19,
+ "id": 3025,
+ "area": 9706,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1540,
+ "bbox": [
+ 232,
+ 251,
+ 156,
+ 58
+ ],
+ "category_id": 19,
+ "id": 3026,
+ "area": 9048,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1541,
+ "bbox": [
+ 245,
+ 76,
+ 28,
+ 356
+ ],
+ "category_id": 19,
+ "id": 3027,
+ "area": 9968,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1542,
+ "bbox": [
+ 54,
+ 179,
+ 390,
+ 159
+ ],
+ "category_id": 10,
+ "id": 3028,
+ "area": 62010,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1543,
+ "bbox": [
+ 171,
+ 210,
+ 197,
+ 79
+ ],
+ "category_id": 16,
+ "id": 3029,
+ "area": 15563,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1544,
+ "bbox": [
+ 128,
+ 211,
+ 79,
+ 95
+ ],
+ "category_id": 15,
+ "id": 3030,
+ "area": 7505,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1544,
+ "bbox": [
+ 229,
+ 224,
+ 78,
+ 97
+ ],
+ "category_id": 15,
+ "id": 3031,
+ "area": 7566,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1544,
+ "bbox": [
+ 113,
+ 316,
+ 80,
+ 97
+ ],
+ "category_id": 15,
+ "id": 3032,
+ "area": 7760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1544,
+ "bbox": [
+ 215,
+ 331,
+ 78,
+ 95
+ ],
+ "category_id": 15,
+ "id": 3033,
+ "area": 7410,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1544,
+ "bbox": [
+ 478,
+ 112,
+ 16,
+ 79
+ ],
+ "category_id": 15,
+ "id": 3034,
+ "area": 1264,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1544,
+ "bbox": [
+ 461,
+ 236,
+ 16,
+ 67
+ ],
+ "category_id": 15,
+ "id": 3035,
+ "area": 1072,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1545,
+ "bbox": [
+ 30,
+ 240,
+ 26,
+ 22
+ ],
+ "category_id": 4,
+ "id": 3036,
+ "area": 572,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1545,
+ "bbox": [
+ 93,
+ 140,
+ 23,
+ 31
+ ],
+ "category_id": 4,
+ "id": 3037,
+ "area": 713,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1545,
+ "bbox": [
+ 257,
+ 94,
+ 21,
+ 28
+ ],
+ "category_id": 4,
+ "id": 3038,
+ "area": 588,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1545,
+ "bbox": [
+ 367,
+ 215,
+ 22,
+ 26
+ ],
+ "category_id": 4,
+ "id": 3039,
+ "area": 572,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1545,
+ "bbox": [
+ 293,
+ 401,
+ 23,
+ 29
+ ],
+ "category_id": 4,
+ "id": 3040,
+ "area": 667,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1545,
+ "bbox": [
+ 472,
+ 390,
+ 21,
+ 26
+ ],
+ "category_id": 4,
+ "id": 3041,
+ "area": 546,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1546,
+ "bbox": [
+ 109,
+ 203,
+ 216,
+ 117
+ ],
+ "category_id": 19,
+ "id": 3042,
+ "area": 25272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1547,
+ "bbox": [
+ 93,
+ 101,
+ 338,
+ 279
+ ],
+ "category_id": 10,
+ "id": 3043,
+ "area": 94302,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1547,
+ "bbox": [
+ 161,
+ 385,
+ 48,
+ 29
+ ],
+ "category_id": 16,
+ "id": 3044,
+ "area": 1392,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1548,
+ "bbox": [
+ 170,
+ 66,
+ 23,
+ 33
+ ],
+ "category_id": 4,
+ "id": 3045,
+ "area": 759,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1548,
+ "bbox": [
+ 443,
+ 71,
+ 18,
+ 26
+ ],
+ "category_id": 4,
+ "id": 3046,
+ "area": 468,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1548,
+ "bbox": [
+ 343,
+ 161,
+ 21,
+ 29
+ ],
+ "category_id": 4,
+ "id": 3047,
+ "area": 609,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1548,
+ "bbox": [
+ 237,
+ 260,
+ 20,
+ 29
+ ],
+ "category_id": 4,
+ "id": 3048,
+ "area": 580,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1548,
+ "bbox": [
+ 100,
+ 397,
+ 16,
+ 24
+ ],
+ "category_id": 4,
+ "id": 3049,
+ "area": 384,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1548,
+ "bbox": [
+ 329,
+ 456,
+ 21,
+ 25
+ ],
+ "category_id": 4,
+ "id": 3050,
+ "area": 525,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1549,
+ "bbox": [
+ 211,
+ 40,
+ 55,
+ 432
+ ],
+ "category_id": 10,
+ "id": 3051,
+ "area": 23760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1550,
+ "bbox": [
+ 231,
+ 49,
+ 163,
+ 411
+ ],
+ "category_id": 10,
+ "id": 3052,
+ "area": 66993,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1551,
+ "bbox": [
+ 52,
+ 168,
+ 24,
+ 20
+ ],
+ "category_id": 15,
+ "id": 3053,
+ "area": 480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1552,
+ "bbox": [
+ 485,
+ 164,
+ 17,
+ 6
+ ],
+ "category_id": 4,
+ "id": 3054,
+ "area": 102,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1552,
+ "bbox": [
+ 254,
+ 192,
+ 28,
+ 16
+ ],
+ "category_id": 4,
+ "id": 3055,
+ "area": 448,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1552,
+ "bbox": [
+ 126,
+ 290,
+ 27,
+ 18
+ ],
+ "category_id": 4,
+ "id": 3056,
+ "area": 486,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1552,
+ "bbox": [
+ 430,
+ 342,
+ 16,
+ 10
+ ],
+ "category_id": 4,
+ "id": 3057,
+ "area": 160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1552,
+ "bbox": [
+ 18,
+ 399,
+ 17,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3058,
+ "area": 153,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1553,
+ "bbox": [
+ 67,
+ 79,
+ 21,
+ 14
+ ],
+ "category_id": 4,
+ "id": 3059,
+ "area": 294,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1553,
+ "bbox": [
+ 261,
+ 151,
+ 19,
+ 14
+ ],
+ "category_id": 4,
+ "id": 3060,
+ "area": 266,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1553,
+ "bbox": [
+ 398,
+ 262,
+ 21,
+ 18
+ ],
+ "category_id": 4,
+ "id": 3061,
+ "area": 378,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1553,
+ "bbox": [
+ 263,
+ 373,
+ 22,
+ 16
+ ],
+ "category_id": 4,
+ "id": 3062,
+ "area": 352,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1553,
+ "bbox": [
+ 474,
+ 449,
+ 21,
+ 18
+ ],
+ "category_id": 4,
+ "id": 3063,
+ "area": 378,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1554,
+ "bbox": [
+ 233,
+ 166,
+ 201,
+ 259
+ ],
+ "category_id": 16,
+ "id": 3064,
+ "area": 52059,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1555,
+ "bbox": [
+ 155,
+ 190,
+ 61,
+ 105
+ ],
+ "category_id": 19,
+ "id": 3065,
+ "area": 6405,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1556,
+ "bbox": [
+ 321,
+ 219,
+ 64,
+ 248
+ ],
+ "category_id": 19,
+ "id": 3066,
+ "area": 15872,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1557,
+ "bbox": [
+ 110,
+ 357,
+ 48,
+ 67
+ ],
+ "category_id": 16,
+ "id": 3067,
+ "area": 3216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1558,
+ "bbox": [
+ 166,
+ 21,
+ 248,
+ 477
+ ],
+ "category_id": 10,
+ "id": 3068,
+ "area": 118296,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1559,
+ "bbox": [
+ 268,
+ 136,
+ 105,
+ 111
+ ],
+ "category_id": 15,
+ "id": 3069,
+ "area": 11655,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1559,
+ "bbox": [
+ 272,
+ 285,
+ 108,
+ 115
+ ],
+ "category_id": 15,
+ "id": 3070,
+ "area": 12420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1560,
+ "bbox": [
+ 186,
+ 174,
+ 169,
+ 222
+ ],
+ "category_id": 10,
+ "id": 3071,
+ "area": 37518,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1561,
+ "bbox": [
+ 20,
+ 55,
+ 26,
+ 16
+ ],
+ "category_id": 4,
+ "id": 3072,
+ "area": 416,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1561,
+ "bbox": [
+ 128,
+ 165,
+ 33,
+ 14
+ ],
+ "category_id": 4,
+ "id": 3073,
+ "area": 462,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1561,
+ "bbox": [
+ 255,
+ 247,
+ 36,
+ 14
+ ],
+ "category_id": 4,
+ "id": 3074,
+ "area": 504,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1561,
+ "bbox": [
+ 465,
+ 99,
+ 36,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3075,
+ "area": 324,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1561,
+ "bbox": [
+ 439,
+ 344,
+ 26,
+ 12
+ ],
+ "category_id": 4,
+ "id": 3076,
+ "area": 312,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1562,
+ "bbox": [
+ 130,
+ 160,
+ 218,
+ 215
+ ],
+ "category_id": 15,
+ "id": 3077,
+ "area": 46870,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1563,
+ "bbox": [
+ 16,
+ 172,
+ 475,
+ 117
+ ],
+ "category_id": 10,
+ "id": 3078,
+ "area": 55575,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1564,
+ "bbox": [
+ 130,
+ 232,
+ 208,
+ 87
+ ],
+ "category_id": 16,
+ "id": 3079,
+ "area": 18096,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1565,
+ "bbox": [
+ 221,
+ 125,
+ 100,
+ 131
+ ],
+ "category_id": 15,
+ "id": 3080,
+ "area": 13100,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1565,
+ "bbox": [
+ 320,
+ 209,
+ 97,
+ 133
+ ],
+ "category_id": 15,
+ "id": 3081,
+ "area": 12901,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1565,
+ "bbox": [
+ 0,
+ 396,
+ 26,
+ 55
+ ],
+ "category_id": 15,
+ "id": 3082,
+ "area": 1430,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1566,
+ "bbox": [
+ 69,
+ 356,
+ 74,
+ 116
+ ],
+ "category_id": 10,
+ "id": 3083,
+ "area": 8584,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1567,
+ "bbox": [
+ 23,
+ 148,
+ 25,
+ 20
+ ],
+ "category_id": 4,
+ "id": 3084,
+ "area": 500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1567,
+ "bbox": [
+ 130,
+ 325,
+ 22,
+ 14
+ ],
+ "category_id": 4,
+ "id": 3085,
+ "area": 308,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1567,
+ "bbox": [
+ 289,
+ 246,
+ 26,
+ 24
+ ],
+ "category_id": 4,
+ "id": 3086,
+ "area": 624,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1567,
+ "bbox": [
+ 338,
+ 351,
+ 23,
+ 18
+ ],
+ "category_id": 4,
+ "id": 3087,
+ "area": 414,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1567,
+ "bbox": [
+ 485,
+ 365,
+ 22,
+ 18
+ ],
+ "category_id": 4,
+ "id": 3088,
+ "area": 396,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1568,
+ "bbox": [
+ 193,
+ 75,
+ 54,
+ 229
+ ],
+ "category_id": 19,
+ "id": 3089,
+ "area": 12366,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1569,
+ "bbox": [
+ 184,
+ 137,
+ 120,
+ 293
+ ],
+ "category_id": 19,
+ "id": 3090,
+ "area": 35160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1570,
+ "bbox": [
+ 112,
+ 54,
+ 160,
+ 172
+ ],
+ "category_id": 15,
+ "id": 3091,
+ "area": 27520,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1570,
+ "bbox": [
+ 124,
+ 273,
+ 161,
+ 173
+ ],
+ "category_id": 15,
+ "id": 3092,
+ "area": 27853,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1571,
+ "bbox": [
+ 195,
+ 217,
+ 141,
+ 169
+ ],
+ "category_id": 15,
+ "id": 3093,
+ "area": 23829,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1572,
+ "bbox": [
+ 232,
+ 181,
+ 86,
+ 315
+ ],
+ "category_id": 19,
+ "id": 3094,
+ "area": 27090,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1572,
+ "bbox": [
+ 140,
+ 81,
+ 96,
+ 410
+ ],
+ "category_id": 19,
+ "id": 3095,
+ "area": 39360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1573,
+ "bbox": [
+ 83,
+ 310,
+ 370,
+ 61
+ ],
+ "category_id": 16,
+ "id": 3096,
+ "area": 22570,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1574,
+ "bbox": [
+ 318,
+ 202,
+ 96,
+ 107
+ ],
+ "category_id": 4,
+ "id": 3097,
+ "area": 10272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1575,
+ "bbox": [
+ 247,
+ 136,
+ 38,
+ 206
+ ],
+ "category_id": 19,
+ "id": 3098,
+ "area": 7828,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1576,
+ "bbox": [
+ 114,
+ 141,
+ 61,
+ 32
+ ],
+ "category_id": 4,
+ "id": 3099,
+ "area": 1952,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1576,
+ "bbox": [
+ 59,
+ 384,
+ 60,
+ 39
+ ],
+ "category_id": 4,
+ "id": 3100,
+ "area": 2340,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1576,
+ "bbox": [
+ 439,
+ 260,
+ 47,
+ 40
+ ],
+ "category_id": 4,
+ "id": 3101,
+ "area": 1880,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1577,
+ "bbox": [
+ 216,
+ 113,
+ 191,
+ 320
+ ],
+ "category_id": 10,
+ "id": 3102,
+ "area": 61120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1578,
+ "bbox": [
+ 222,
+ 189,
+ 58,
+ 164
+ ],
+ "category_id": 19,
+ "id": 3103,
+ "area": 9512,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1579,
+ "bbox": [
+ 69,
+ 75,
+ 48,
+ 29
+ ],
+ "category_id": 4,
+ "id": 3104,
+ "area": 1392,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1579,
+ "bbox": [
+ 65,
+ 405,
+ 64,
+ 28
+ ],
+ "category_id": 4,
+ "id": 3105,
+ "area": 1792,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1579,
+ "bbox": [
+ 426,
+ 151,
+ 55,
+ 31
+ ],
+ "category_id": 4,
+ "id": 3106,
+ "area": 1705,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1580,
+ "bbox": [
+ 16,
+ 69,
+ 240,
+ 201
+ ],
+ "category_id": 15,
+ "id": 3107,
+ "area": 48240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1580,
+ "bbox": [
+ 263,
+ 217,
+ 236,
+ 204
+ ],
+ "category_id": 15,
+ "id": 3108,
+ "area": 48144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1581,
+ "bbox": [
+ 191,
+ 73,
+ 41,
+ 39
+ ],
+ "category_id": 4,
+ "id": 3109,
+ "area": 1599,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1581,
+ "bbox": [
+ 171,
+ 448,
+ 35,
+ 35
+ ],
+ "category_id": 4,
+ "id": 3110,
+ "area": 1225,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1581,
+ "bbox": [
+ 324,
+ 311,
+ 42,
+ 39
+ ],
+ "category_id": 4,
+ "id": 3111,
+ "area": 1638,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1582,
+ "bbox": [
+ 152,
+ 120,
+ 91,
+ 91
+ ],
+ "category_id": 15,
+ "id": 3112,
+ "area": 8281,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1582,
+ "bbox": [
+ 266,
+ 152,
+ 87,
+ 85
+ ],
+ "category_id": 15,
+ "id": 3113,
+ "area": 7395,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1582,
+ "bbox": [
+ 248,
+ 271,
+ 100,
+ 95
+ ],
+ "category_id": 15,
+ "id": 3114,
+ "area": 9500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1583,
+ "bbox": [
+ 26,
+ 17,
+ 212,
+ 59
+ ],
+ "category_id": 10,
+ "id": 3115,
+ "area": 12508,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1584,
+ "bbox": [
+ 88,
+ 447,
+ 34,
+ 31
+ ],
+ "category_id": 15,
+ "id": 3116,
+ "area": 1054,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1585,
+ "bbox": [
+ 432,
+ 211,
+ 52,
+ 71
+ ],
+ "category_id": 19,
+ "id": 3117,
+ "area": 3692,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1586,
+ "bbox": [
+ 69,
+ 243,
+ 311,
+ 153
+ ],
+ "category_id": 16,
+ "id": 3118,
+ "area": 47583,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1587,
+ "bbox": [
+ 225,
+ 101,
+ 40,
+ 281
+ ],
+ "category_id": 19,
+ "id": 3119,
+ "area": 11240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1588,
+ "bbox": [
+ 101,
+ 105,
+ 273,
+ 337
+ ],
+ "category_id": 15,
+ "id": 3120,
+ "area": 92001,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1589,
+ "bbox": [
+ 217,
+ 253,
+ 90,
+ 91
+ ],
+ "category_id": 15,
+ "id": 3121,
+ "area": 8190,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1590,
+ "bbox": [
+ 156,
+ 282,
+ 57,
+ 59
+ ],
+ "category_id": 4,
+ "id": 3122,
+ "area": 3363,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1590,
+ "bbox": [
+ 396,
+ 298,
+ 47,
+ 50
+ ],
+ "category_id": 4,
+ "id": 3123,
+ "area": 2350,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1591,
+ "bbox": [
+ 208,
+ 118,
+ 46,
+ 50
+ ],
+ "category_id": 4,
+ "id": 3124,
+ "area": 2300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1591,
+ "bbox": [
+ 44,
+ 427,
+ 52,
+ 63
+ ],
+ "category_id": 4,
+ "id": 3125,
+ "area": 3276,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1591,
+ "bbox": [
+ 383,
+ 297,
+ 53,
+ 72
+ ],
+ "category_id": 4,
+ "id": 3126,
+ "area": 3816,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1592,
+ "bbox": [
+ 80,
+ 74,
+ 129,
+ 89
+ ],
+ "category_id": 10,
+ "id": 3127,
+ "area": 11481,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1592,
+ "bbox": [
+ 103,
+ 278,
+ 25,
+ 19
+ ],
+ "category_id": 19,
+ "id": 3128,
+ "area": 475,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1593,
+ "bbox": [
+ 161,
+ 267,
+ 76,
+ 93
+ ],
+ "category_id": 15,
+ "id": 3129,
+ "area": 7068,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1593,
+ "bbox": [
+ 256,
+ 173,
+ 79,
+ 91
+ ],
+ "category_id": 15,
+ "id": 3130,
+ "area": 7189,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1594,
+ "bbox": [
+ 86,
+ 58,
+ 225,
+ 102
+ ],
+ "category_id": 19,
+ "id": 3131,
+ "area": 22950,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1595,
+ "bbox": [
+ 74,
+ 296,
+ 189,
+ 142
+ ],
+ "category_id": 10,
+ "id": 3132,
+ "area": 26838,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1596,
+ "bbox": [
+ 380,
+ 74,
+ 78,
+ 99
+ ],
+ "category_id": 15,
+ "id": 3133,
+ "area": 7722,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1596,
+ "bbox": [
+ 346,
+ 192,
+ 47,
+ 80
+ ],
+ "category_id": 15,
+ "id": 3134,
+ "area": 3760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1597,
+ "bbox": [
+ 232,
+ 48,
+ 16,
+ 12
+ ],
+ "category_id": 4,
+ "id": 3135,
+ "area": 192,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1597,
+ "bbox": [
+ 16,
+ 272,
+ 17,
+ 10
+ ],
+ "category_id": 4,
+ "id": 3136,
+ "area": 170,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1597,
+ "bbox": [
+ 210,
+ 251,
+ 17,
+ 8
+ ],
+ "category_id": 4,
+ "id": 3137,
+ "area": 136,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1597,
+ "bbox": [
+ 296,
+ 204,
+ 16,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3138,
+ "area": 176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1597,
+ "bbox": [
+ 395,
+ 248,
+ 19,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3139,
+ "area": 209,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1597,
+ "bbox": [
+ 472,
+ 404,
+ 25,
+ 19
+ ],
+ "category_id": 4,
+ "id": 3140,
+ "area": 475,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1598,
+ "bbox": [
+ 77,
+ 180,
+ 138,
+ 121
+ ],
+ "category_id": 15,
+ "id": 3141,
+ "area": 16698,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1598,
+ "bbox": [
+ 248,
+ 167,
+ 136,
+ 122
+ ],
+ "category_id": 15,
+ "id": 3142,
+ "area": 16592,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1599,
+ "bbox": [
+ 37,
+ 186,
+ 396,
+ 191
+ ],
+ "category_id": 16,
+ "id": 3143,
+ "area": 75636,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1600,
+ "bbox": [
+ 296,
+ 277,
+ 183,
+ 171
+ ],
+ "category_id": 10,
+ "id": 3144,
+ "area": 31293,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1601,
+ "bbox": [
+ 65,
+ 275,
+ 55,
+ 46
+ ],
+ "category_id": 4,
+ "id": 3145,
+ "area": 2530,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1601,
+ "bbox": [
+ 428,
+ 328,
+ 60,
+ 54
+ ],
+ "category_id": 4,
+ "id": 3146,
+ "area": 3240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1602,
+ "bbox": [
+ 474,
+ 186,
+ 30,
+ 16
+ ],
+ "category_id": 4,
+ "id": 3147,
+ "area": 480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1602,
+ "bbox": [
+ 395,
+ 266,
+ 28,
+ 19
+ ],
+ "category_id": 4,
+ "id": 3148,
+ "area": 532,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1603,
+ "bbox": [
+ 134,
+ 105,
+ 29,
+ 31
+ ],
+ "category_id": 4,
+ "id": 3149,
+ "area": 899,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1603,
+ "bbox": [
+ 339,
+ 375,
+ 32,
+ 38
+ ],
+ "category_id": 4,
+ "id": 3150,
+ "area": 1216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1604,
+ "bbox": [
+ 130,
+ 200,
+ 223,
+ 240
+ ],
+ "category_id": 16,
+ "id": 3151,
+ "area": 53520,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1605,
+ "bbox": [
+ 231,
+ 246,
+ 83,
+ 97
+ ],
+ "category_id": 15,
+ "id": 3152,
+ "area": 8051,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1605,
+ "bbox": [
+ 193,
+ 348,
+ 65,
+ 96
+ ],
+ "category_id": 15,
+ "id": 3153,
+ "area": 6240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1606,
+ "bbox": [
+ 186,
+ 144,
+ 193,
+ 202
+ ],
+ "category_id": 10,
+ "id": 3154,
+ "area": 38986,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1607,
+ "bbox": [
+ 199,
+ 168,
+ 90,
+ 171
+ ],
+ "category_id": 19,
+ "id": 3155,
+ "area": 15390,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1608,
+ "bbox": [
+ 40,
+ 229,
+ 9,
+ 24
+ ],
+ "category_id": 4,
+ "id": 3156,
+ "area": 216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1608,
+ "bbox": [
+ 210,
+ 264,
+ 12,
+ 20
+ ],
+ "category_id": 4,
+ "id": 3157,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1608,
+ "bbox": [
+ 192,
+ 81,
+ 13,
+ 23
+ ],
+ "category_id": 4,
+ "id": 3158,
+ "area": 299,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1608,
+ "bbox": [
+ 371,
+ 85,
+ 13,
+ 22
+ ],
+ "category_id": 4,
+ "id": 3159,
+ "area": 286,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1608,
+ "bbox": [
+ 237,
+ 469,
+ 11,
+ 23
+ ],
+ "category_id": 4,
+ "id": 3160,
+ "area": 253,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1608,
+ "bbox": [
+ 436,
+ 385,
+ 12,
+ 22
+ ],
+ "category_id": 4,
+ "id": 3161,
+ "area": 264,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1609,
+ "bbox": [
+ 45,
+ 164,
+ 35,
+ 27
+ ],
+ "category_id": 4,
+ "id": 3162,
+ "area": 945,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1609,
+ "bbox": [
+ 410,
+ 88,
+ 30,
+ 28
+ ],
+ "category_id": 4,
+ "id": 3163,
+ "area": 840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1609,
+ "bbox": [
+ 244,
+ 425,
+ 35,
+ 27
+ ],
+ "category_id": 4,
+ "id": 3164,
+ "area": 945,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1610,
+ "bbox": [
+ 321,
+ 60,
+ 74,
+ 359
+ ],
+ "category_id": 10,
+ "id": 3165,
+ "area": 26566,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1611,
+ "bbox": [
+ 151,
+ 222,
+ 206,
+ 83
+ ],
+ "category_id": 19,
+ "id": 3166,
+ "area": 17098,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1612,
+ "bbox": [
+ 245,
+ 76,
+ 147,
+ 377
+ ],
+ "category_id": 10,
+ "id": 3167,
+ "area": 55419,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1613,
+ "bbox": [
+ 30,
+ 474,
+ 18,
+ 22
+ ],
+ "category_id": 4,
+ "id": 3168,
+ "area": 396,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1613,
+ "bbox": [
+ 126,
+ 372,
+ 17,
+ 20
+ ],
+ "category_id": 4,
+ "id": 3169,
+ "area": 340,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1613,
+ "bbox": [
+ 284,
+ 323,
+ 20,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3170,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1613,
+ "bbox": [
+ 446,
+ 314,
+ 21,
+ 33
+ ],
+ "category_id": 4,
+ "id": 3171,
+ "area": 693,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1613,
+ "bbox": [
+ 352,
+ 163,
+ 23,
+ 18
+ ],
+ "category_id": 4,
+ "id": 3172,
+ "area": 414,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1614,
+ "bbox": [
+ 194,
+ 67,
+ 147,
+ 374
+ ],
+ "category_id": 16,
+ "id": 3173,
+ "area": 54978,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1615,
+ "bbox": [
+ 302,
+ 58,
+ 5,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3174,
+ "area": 55,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1615,
+ "bbox": [
+ 462,
+ 192,
+ 7,
+ 16
+ ],
+ "category_id": 4,
+ "id": 3175,
+ "area": 112,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1615,
+ "bbox": [
+ 385,
+ 267,
+ 7,
+ 13
+ ],
+ "category_id": 4,
+ "id": 3176,
+ "area": 91,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1615,
+ "bbox": [
+ 54,
+ 352,
+ 7,
+ 14
+ ],
+ "category_id": 4,
+ "id": 3177,
+ "area": 98,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1615,
+ "bbox": [
+ 23,
+ 408,
+ 12,
+ 15
+ ],
+ "category_id": 4,
+ "id": 3178,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1616,
+ "bbox": [
+ 86,
+ 399,
+ 34,
+ 33
+ ],
+ "category_id": 4,
+ "id": 3179,
+ "area": 1122,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1616,
+ "bbox": [
+ 205,
+ 167,
+ 34,
+ 26
+ ],
+ "category_id": 4,
+ "id": 3180,
+ "area": 884,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1616,
+ "bbox": [
+ 448,
+ 365,
+ 38,
+ 33
+ ],
+ "category_id": 4,
+ "id": 3181,
+ "area": 1254,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1617,
+ "bbox": [
+ 220,
+ 186,
+ 50,
+ 186
+ ],
+ "category_id": 19,
+ "id": 3182,
+ "area": 9300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1618,
+ "bbox": [
+ 243,
+ 82,
+ 90,
+ 72
+ ],
+ "category_id": 15,
+ "id": 3183,
+ "area": 6480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1618,
+ "bbox": [
+ 224,
+ 197,
+ 92,
+ 73
+ ],
+ "category_id": 15,
+ "id": 3184,
+ "area": 6716,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1618,
+ "bbox": [
+ 216,
+ 307,
+ 88,
+ 69
+ ],
+ "category_id": 15,
+ "id": 3185,
+ "area": 6072,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1619,
+ "bbox": [
+ 54,
+ 271,
+ 399,
+ 70
+ ],
+ "category_id": 19,
+ "id": 3186,
+ "area": 27930,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1620,
+ "bbox": [
+ 60,
+ 197,
+ 359,
+ 231
+ ],
+ "category_id": 16,
+ "id": 3187,
+ "area": 82929,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1621,
+ "bbox": [
+ 345,
+ 135,
+ 133,
+ 199
+ ],
+ "category_id": 19,
+ "id": 3188,
+ "area": 26467,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1622,
+ "bbox": [
+ 171,
+ 110,
+ 181,
+ 286
+ ],
+ "category_id": 16,
+ "id": 3189,
+ "area": 51766,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1623,
+ "bbox": [
+ 21,
+ 279,
+ 151,
+ 172
+ ],
+ "category_id": 10,
+ "id": 3190,
+ "area": 25972,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1624,
+ "bbox": [
+ 299,
+ 92,
+ 38,
+ 9
+ ],
+ "category_id": 16,
+ "id": 3191,
+ "area": 342,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1625,
+ "bbox": [
+ 88,
+ 158,
+ 15,
+ 13
+ ],
+ "category_id": 15,
+ "id": 3192,
+ "area": 195,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1626,
+ "bbox": [
+ 33,
+ 346,
+ 193,
+ 116
+ ],
+ "category_id": 10,
+ "id": 3193,
+ "area": 22388,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1627,
+ "bbox": [
+ 241,
+ 144,
+ 223,
+ 301
+ ],
+ "category_id": 10,
+ "id": 3194,
+ "area": 67123,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1628,
+ "bbox": [
+ 337,
+ 66,
+ 49,
+ 24
+ ],
+ "category_id": 4,
+ "id": 3195,
+ "area": 1176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1628,
+ "bbox": [
+ 278,
+ 404,
+ 43,
+ 22
+ ],
+ "category_id": 4,
+ "id": 3196,
+ "area": 946,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1629,
+ "bbox": [
+ 360,
+ 86,
+ 19,
+ 51
+ ],
+ "category_id": 4,
+ "id": 3197,
+ "area": 969,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1629,
+ "bbox": [
+ 80,
+ 374,
+ 48,
+ 43
+ ],
+ "category_id": 4,
+ "id": 3198,
+ "area": 2064,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1630,
+ "bbox": [
+ 42,
+ 181,
+ 434,
+ 78
+ ],
+ "category_id": 10,
+ "id": 3199,
+ "area": 33852,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1631,
+ "bbox": [
+ 132,
+ 287,
+ 296,
+ 17
+ ],
+ "category_id": 10,
+ "id": 3200,
+ "area": 5032,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1632,
+ "bbox": [
+ 62,
+ 222,
+ 7,
+ 6
+ ],
+ "category_id": 4,
+ "id": 3201,
+ "area": 42,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1632,
+ "bbox": [
+ 83,
+ 167,
+ 4,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3202,
+ "area": 28,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1632,
+ "bbox": [
+ 132,
+ 152,
+ 4,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3203,
+ "area": 36,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1632,
+ "bbox": [
+ 194,
+ 98,
+ 7,
+ 10
+ ],
+ "category_id": 4,
+ "id": 3204,
+ "area": 70,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1632,
+ "bbox": [
+ 127,
+ 92,
+ 6,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3205,
+ "area": 66,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1632,
+ "bbox": [
+ 307,
+ 120,
+ 7,
+ 15
+ ],
+ "category_id": 4,
+ "id": 3206,
+ "area": 105,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1632,
+ "bbox": [
+ 389,
+ 83,
+ 5,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3207,
+ "area": 55,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1632,
+ "bbox": [
+ 193,
+ 231,
+ 7,
+ 14
+ ],
+ "category_id": 4,
+ "id": 3208,
+ "area": 98,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1632,
+ "bbox": [
+ 188,
+ 289,
+ 6,
+ 10
+ ],
+ "category_id": 4,
+ "id": 3209,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1632,
+ "bbox": [
+ 381,
+ 193,
+ 5,
+ 15
+ ],
+ "category_id": 4,
+ "id": 3210,
+ "area": 75,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1632,
+ "bbox": [
+ 467,
+ 232,
+ 5,
+ 8
+ ],
+ "category_id": 4,
+ "id": 3211,
+ "area": 40,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1633,
+ "bbox": [
+ 193,
+ 250,
+ 128,
+ 167
+ ],
+ "category_id": 19,
+ "id": 3212,
+ "area": 21376,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1634,
+ "bbox": [
+ 172,
+ 64,
+ 215,
+ 344
+ ],
+ "category_id": 10,
+ "id": 3213,
+ "area": 73960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1635,
+ "bbox": [
+ 214,
+ 284,
+ 168,
+ 147
+ ],
+ "category_id": 19,
+ "id": 3214,
+ "area": 24696,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1636,
+ "bbox": [
+ 143,
+ 96,
+ 93,
+ 135
+ ],
+ "category_id": 15,
+ "id": 3215,
+ "area": 12555,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1636,
+ "bbox": [
+ 263,
+ 103,
+ 99,
+ 129
+ ],
+ "category_id": 15,
+ "id": 3216,
+ "area": 12771,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1636,
+ "bbox": [
+ 335,
+ 241,
+ 95,
+ 122
+ ],
+ "category_id": 15,
+ "id": 3217,
+ "area": 11590,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1636,
+ "bbox": [
+ 195,
+ 226,
+ 111,
+ 155
+ ],
+ "category_id": 15,
+ "id": 3218,
+ "area": 17205,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1636,
+ "bbox": [
+ 79,
+ 265,
+ 39,
+ 55
+ ],
+ "category_id": 15,
+ "id": 3219,
+ "area": 2145,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1637,
+ "bbox": [
+ 232,
+ 47,
+ 176,
+ 145
+ ],
+ "category_id": 15,
+ "id": 3220,
+ "area": 25520,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1637,
+ "bbox": [
+ 284,
+ 144,
+ 131,
+ 121
+ ],
+ "category_id": 15,
+ "id": 3221,
+ "area": 15851,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1637,
+ "bbox": [
+ 161,
+ 214,
+ 129,
+ 117
+ ],
+ "category_id": 15,
+ "id": 3222,
+ "area": 15093,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1637,
+ "bbox": [
+ 272,
+ 355,
+ 151,
+ 138
+ ],
+ "category_id": 15,
+ "id": 3223,
+ "area": 20838,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1638,
+ "bbox": [
+ 42,
+ 156,
+ 58,
+ 53
+ ],
+ "category_id": 4,
+ "id": 3224,
+ "area": 3074,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1638,
+ "bbox": [
+ 428,
+ 348,
+ 68,
+ 49
+ ],
+ "category_id": 4,
+ "id": 3225,
+ "area": 3332,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1639,
+ "bbox": [
+ 122,
+ 30,
+ 63,
+ 205
+ ],
+ "category_id": 19,
+ "id": 3226,
+ "area": 12915,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1640,
+ "bbox": [
+ 122,
+ 137,
+ 297,
+ 296
+ ],
+ "category_id": 15,
+ "id": 3227,
+ "area": 87912,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1641,
+ "bbox": [
+ 43,
+ 97,
+ 410,
+ 258
+ ],
+ "category_id": 16,
+ "id": 3228,
+ "area": 105780,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1642,
+ "bbox": [
+ 117,
+ 126,
+ 308,
+ 226
+ ],
+ "category_id": 16,
+ "id": 3229,
+ "area": 69608,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1643,
+ "bbox": [
+ 195,
+ 197,
+ 98,
+ 110
+ ],
+ "category_id": 15,
+ "id": 3230,
+ "area": 10780,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1643,
+ "bbox": [
+ 182,
+ 327,
+ 98,
+ 114
+ ],
+ "category_id": 15,
+ "id": 3231,
+ "area": 11172,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1644,
+ "bbox": [
+ 94,
+ 305,
+ 41,
+ 164
+ ],
+ "category_id": 10,
+ "id": 3232,
+ "area": 6724,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1645,
+ "bbox": [
+ 186,
+ 5,
+ 28,
+ 38
+ ],
+ "category_id": 16,
+ "id": 3233,
+ "area": 1064,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1645,
+ "bbox": [
+ 216,
+ 303,
+ 16,
+ 29
+ ],
+ "category_id": 16,
+ "id": 3234,
+ "area": 464,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1645,
+ "bbox": [
+ 392,
+ 420,
+ 24,
+ 40
+ ],
+ "category_id": 16,
+ "id": 3235,
+ "area": 960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1646,
+ "bbox": [
+ 156,
+ 222,
+ 103,
+ 67
+ ],
+ "category_id": 15,
+ "id": 3236,
+ "area": 6901,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1646,
+ "bbox": [
+ 307,
+ 190,
+ 109,
+ 72
+ ],
+ "category_id": 15,
+ "id": 3237,
+ "area": 7848,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1646,
+ "bbox": [
+ 120,
+ 50,
+ 96,
+ 16
+ ],
+ "category_id": 15,
+ "id": 3238,
+ "area": 1536,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1647,
+ "bbox": [
+ 90,
+ 210,
+ 373,
+ 157
+ ],
+ "category_id": 10,
+ "id": 3239,
+ "area": 58561,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1648,
+ "bbox": [
+ 280,
+ 386,
+ 176,
+ 37
+ ],
+ "category_id": 10,
+ "id": 3240,
+ "area": 6512,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1649,
+ "bbox": [
+ 216,
+ 131,
+ 82,
+ 277
+ ],
+ "category_id": 16,
+ "id": 3241,
+ "area": 22714,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1650,
+ "bbox": [
+ 51,
+ 74,
+ 132,
+ 214
+ ],
+ "category_id": 10,
+ "id": 3242,
+ "area": 28248,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1651,
+ "bbox": [
+ 207,
+ 173,
+ 200,
+ 160
+ ],
+ "category_id": 16,
+ "id": 3243,
+ "area": 32000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1652,
+ "bbox": [
+ 329,
+ 22,
+ 51,
+ 434
+ ],
+ "category_id": 10,
+ "id": 3244,
+ "area": 22134,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1653,
+ "bbox": [
+ 183,
+ 66,
+ 215,
+ 246
+ ],
+ "category_id": 16,
+ "id": 3245,
+ "area": 52890,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1654,
+ "bbox": [
+ 78,
+ 94,
+ 393,
+ 364
+ ],
+ "category_id": 10,
+ "id": 3246,
+ "area": 143052,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1655,
+ "bbox": [
+ 59,
+ 73,
+ 424,
+ 306
+ ],
+ "category_id": 10,
+ "id": 3247,
+ "area": 129744,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1656,
+ "bbox": [
+ 199,
+ 115,
+ 51,
+ 249
+ ],
+ "category_id": 19,
+ "id": 3248,
+ "area": 12699,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1657,
+ "bbox": [
+ 227,
+ 67,
+ 194,
+ 392
+ ],
+ "category_id": 10,
+ "id": 3249,
+ "area": 76048,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1658,
+ "bbox": [
+ 88,
+ 12,
+ 28,
+ 19
+ ],
+ "category_id": 15,
+ "id": 3250,
+ "area": 532,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1659,
+ "bbox": [
+ 250,
+ 133,
+ 157,
+ 137
+ ],
+ "category_id": 15,
+ "id": 3251,
+ "area": 21509,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1659,
+ "bbox": [
+ 257,
+ 297,
+ 153,
+ 132
+ ],
+ "category_id": 15,
+ "id": 3252,
+ "area": 20196,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1660,
+ "bbox": [
+ 211,
+ 99,
+ 141,
+ 262
+ ],
+ "category_id": 16,
+ "id": 3253,
+ "area": 36942,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1661,
+ "bbox": [
+ 445,
+ 311,
+ 45,
+ 106
+ ],
+ "category_id": 10,
+ "id": 3254,
+ "area": 4770,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1662,
+ "bbox": [
+ 177,
+ 68,
+ 22,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3255,
+ "area": 462,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1662,
+ "bbox": [
+ 378,
+ 44,
+ 18,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3256,
+ "area": 198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1662,
+ "bbox": [
+ 480,
+ 251,
+ 20,
+ 12
+ ],
+ "category_id": 4,
+ "id": 3257,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1662,
+ "bbox": [
+ 101,
+ 264,
+ 18,
+ 18
+ ],
+ "category_id": 4,
+ "id": 3258,
+ "area": 324,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1662,
+ "bbox": [
+ 305,
+ 284,
+ 20,
+ 19
+ ],
+ "category_id": 4,
+ "id": 3259,
+ "area": 380,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1662,
+ "bbox": [
+ 233,
+ 416,
+ 21,
+ 14
+ ],
+ "category_id": 4,
+ "id": 3260,
+ "area": 294,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1662,
+ "bbox": [
+ 56,
+ 494,
+ 18,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3261,
+ "area": 162,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1663,
+ "bbox": [
+ 395,
+ 224,
+ 44,
+ 39
+ ],
+ "category_id": 15,
+ "id": 3262,
+ "area": 1716,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1664,
+ "bbox": [
+ 222,
+ 81,
+ 22,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3263,
+ "area": 462,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1664,
+ "bbox": [
+ 207,
+ 190,
+ 22,
+ 16
+ ],
+ "category_id": 4,
+ "id": 3264,
+ "area": 352,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1664,
+ "bbox": [
+ 87,
+ 323,
+ 19,
+ 16
+ ],
+ "category_id": 4,
+ "id": 3265,
+ "area": 304,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1664,
+ "bbox": [
+ 24,
+ 393,
+ 24,
+ 15
+ ],
+ "category_id": 4,
+ "id": 3266,
+ "area": 360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1664,
+ "bbox": [
+ 198,
+ 401,
+ 18,
+ 17
+ ],
+ "category_id": 4,
+ "id": 3267,
+ "area": 306,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1664,
+ "bbox": [
+ 357,
+ 267,
+ 19,
+ 13
+ ],
+ "category_id": 4,
+ "id": 3268,
+ "area": 247,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1664,
+ "bbox": [
+ 432,
+ 471,
+ 17,
+ 13
+ ],
+ "category_id": 4,
+ "id": 3269,
+ "area": 221,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1665,
+ "bbox": [
+ 309,
+ 32,
+ 35,
+ 28
+ ],
+ "category_id": 15,
+ "id": 3270,
+ "area": 980,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1665,
+ "bbox": [
+ 324,
+ 85,
+ 35,
+ 28
+ ],
+ "category_id": 15,
+ "id": 3271,
+ "area": 980,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1665,
+ "bbox": [
+ 337,
+ 136,
+ 38,
+ 27
+ ],
+ "category_id": 15,
+ "id": 3272,
+ "area": 1026,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1666,
+ "bbox": [
+ 204,
+ 80,
+ 185,
+ 319
+ ],
+ "category_id": 16,
+ "id": 3273,
+ "area": 59015,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1667,
+ "bbox": [
+ 332,
+ 64,
+ 99,
+ 371
+ ],
+ "category_id": 10,
+ "id": 3274,
+ "area": 36729,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1668,
+ "bbox": [
+ 87,
+ 339,
+ 71,
+ 117
+ ],
+ "category_id": 10,
+ "id": 3275,
+ "area": 8307,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1669,
+ "bbox": [
+ 159,
+ 314,
+ 118,
+ 116
+ ],
+ "category_id": 15,
+ "id": 3276,
+ "area": 13688,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1669,
+ "bbox": [
+ 193,
+ 125,
+ 19,
+ 33
+ ],
+ "category_id": 15,
+ "id": 3277,
+ "area": 627,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1670,
+ "bbox": [
+ 252,
+ 39,
+ 225,
+ 138
+ ],
+ "category_id": 10,
+ "id": 3278,
+ "area": 31050,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1671,
+ "bbox": [
+ 141,
+ 190,
+ 257,
+ 181
+ ],
+ "category_id": 16,
+ "id": 3279,
+ "area": 46517,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1672,
+ "bbox": [
+ 98,
+ 307,
+ 150,
+ 155
+ ],
+ "category_id": 10,
+ "id": 3280,
+ "area": 23250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1673,
+ "bbox": [
+ 165,
+ 251,
+ 277,
+ 68
+ ],
+ "category_id": 19,
+ "id": 3281,
+ "area": 18836,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1674,
+ "bbox": [
+ 92,
+ 21,
+ 82,
+ 249
+ ],
+ "category_id": 10,
+ "id": 3282,
+ "area": 20418,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1675,
+ "bbox": [
+ 160,
+ 192,
+ 236,
+ 184
+ ],
+ "category_id": 16,
+ "id": 3283,
+ "area": 43424,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1676,
+ "bbox": [
+ 157,
+ 79,
+ 141,
+ 179
+ ],
+ "category_id": 15,
+ "id": 3284,
+ "area": 25239,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1676,
+ "bbox": [
+ 152,
+ 284,
+ 144,
+ 179
+ ],
+ "category_id": 15,
+ "id": 3285,
+ "area": 25776,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1677,
+ "bbox": [
+ 195,
+ 128,
+ 204,
+ 302
+ ],
+ "category_id": 16,
+ "id": 3286,
+ "area": 61608,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1678,
+ "bbox": [
+ 117,
+ 107,
+ 289,
+ 250
+ ],
+ "category_id": 16,
+ "id": 3287,
+ "area": 72250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1679,
+ "bbox": [
+ 147,
+ 220,
+ 146,
+ 166
+ ],
+ "category_id": 15,
+ "id": 3288,
+ "area": 24236,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1680,
+ "bbox": [
+ 288,
+ 122,
+ 77,
+ 312
+ ],
+ "category_id": 16,
+ "id": 3289,
+ "area": 24024,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1681,
+ "bbox": [
+ 48,
+ 365,
+ 67,
+ 24
+ ],
+ "category_id": 4,
+ "id": 3290,
+ "area": 1608,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1681,
+ "bbox": [
+ 368,
+ 58,
+ 32,
+ 69
+ ],
+ "category_id": 4,
+ "id": 3291,
+ "area": 2208,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1682,
+ "bbox": [
+ 261,
+ 151,
+ 114,
+ 251
+ ],
+ "category_id": 16,
+ "id": 3292,
+ "area": 28614,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1683,
+ "bbox": [
+ 62,
+ 223,
+ 264,
+ 68
+ ],
+ "category_id": 19,
+ "id": 3293,
+ "area": 17952,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1684,
+ "bbox": [
+ 55,
+ 241,
+ 10,
+ 10
+ ],
+ "category_id": 4,
+ "id": 3294,
+ "area": 100,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1684,
+ "bbox": [
+ 172,
+ 300,
+ 7,
+ 8
+ ],
+ "category_id": 4,
+ "id": 3295,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1684,
+ "bbox": [
+ 122,
+ 379,
+ 7,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3296,
+ "area": 63,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1684,
+ "bbox": [
+ 154,
+ 485,
+ 8,
+ 6
+ ],
+ "category_id": 4,
+ "id": 3297,
+ "area": 48,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1684,
+ "bbox": [
+ 279,
+ 241,
+ 6,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3298,
+ "area": 66,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1684,
+ "bbox": [
+ 215,
+ 140,
+ 8,
+ 10
+ ],
+ "category_id": 4,
+ "id": 3299,
+ "area": 80,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1684,
+ "bbox": [
+ 51,
+ 42,
+ 7,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3300,
+ "area": 49,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1684,
+ "bbox": [
+ 469,
+ 186,
+ 12,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3301,
+ "area": 132,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1684,
+ "bbox": [
+ 421,
+ 71,
+ 8,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3302,
+ "area": 72,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1685,
+ "bbox": [
+ 309,
+ 172,
+ 55,
+ 194
+ ],
+ "category_id": 19,
+ "id": 3303,
+ "area": 10670,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1686,
+ "bbox": [
+ 311,
+ 208,
+ 103,
+ 86
+ ],
+ "category_id": 4,
+ "id": 3304,
+ "area": 8858,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1687,
+ "bbox": [
+ 68,
+ 258,
+ 382,
+ 72
+ ],
+ "category_id": 19,
+ "id": 3305,
+ "area": 27504,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1688,
+ "bbox": [
+ 106,
+ 118,
+ 54,
+ 43
+ ],
+ "category_id": 4,
+ "id": 3306,
+ "area": 2322,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1688,
+ "bbox": [
+ 50,
+ 336,
+ 49,
+ 43
+ ],
+ "category_id": 4,
+ "id": 3307,
+ "area": 2107,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1688,
+ "bbox": [
+ 428,
+ 279,
+ 58,
+ 46
+ ],
+ "category_id": 4,
+ "id": 3308,
+ "area": 2668,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1689,
+ "bbox": [
+ 253,
+ 81,
+ 35,
+ 41
+ ],
+ "category_id": 4,
+ "id": 3309,
+ "area": 1435,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1689,
+ "bbox": [
+ 193,
+ 413,
+ 34,
+ 41
+ ],
+ "category_id": 4,
+ "id": 3310,
+ "area": 1394,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1690,
+ "bbox": [
+ 240,
+ 232,
+ 133,
+ 181
+ ],
+ "category_id": 16,
+ "id": 3311,
+ "area": 24073,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1691,
+ "bbox": [
+ 179,
+ 382,
+ 78,
+ 36
+ ],
+ "category_id": 4,
+ "id": 3312,
+ "area": 2808,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1691,
+ "bbox": [
+ 291,
+ 61,
+ 78,
+ 40
+ ],
+ "category_id": 4,
+ "id": 3313,
+ "area": 3120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1692,
+ "bbox": [
+ 268,
+ 208,
+ 132,
+ 26
+ ],
+ "category_id": 4,
+ "id": 3314,
+ "area": 3432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1693,
+ "bbox": [
+ 229,
+ 227,
+ 39,
+ 209
+ ],
+ "category_id": 16,
+ "id": 3315,
+ "area": 8151,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1694,
+ "bbox": [
+ 257,
+ 227,
+ 171,
+ 241
+ ],
+ "category_id": 10,
+ "id": 3316,
+ "area": 41211,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1695,
+ "bbox": [
+ 264,
+ 456,
+ 24,
+ 29
+ ],
+ "category_id": 16,
+ "id": 3317,
+ "area": 696,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1696,
+ "bbox": [
+ 167,
+ 90,
+ 215,
+ 275
+ ],
+ "category_id": 16,
+ "id": 3318,
+ "area": 59125,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1697,
+ "bbox": [
+ 89,
+ 263,
+ 97,
+ 155
+ ],
+ "category_id": 15,
+ "id": 3319,
+ "area": 15035,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1697,
+ "bbox": [
+ 208,
+ 200,
+ 122,
+ 159
+ ],
+ "category_id": 15,
+ "id": 3320,
+ "area": 19398,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1697,
+ "bbox": [
+ 364,
+ 176,
+ 87,
+ 118
+ ],
+ "category_id": 15,
+ "id": 3321,
+ "area": 10266,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1698,
+ "bbox": [
+ 237,
+ 247,
+ 56,
+ 89
+ ],
+ "category_id": 4,
+ "id": 3322,
+ "area": 4984,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1699,
+ "bbox": [
+ 98,
+ 80,
+ 19,
+ 14
+ ],
+ "category_id": 4,
+ "id": 3323,
+ "area": 266,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1699,
+ "bbox": [
+ 245,
+ 81,
+ 33,
+ 22
+ ],
+ "category_id": 4,
+ "id": 3324,
+ "area": 726,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1699,
+ "bbox": [
+ 80,
+ 427,
+ 24,
+ 19
+ ],
+ "category_id": 4,
+ "id": 3325,
+ "area": 456,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1699,
+ "bbox": [
+ 237,
+ 425,
+ 24,
+ 19
+ ],
+ "category_id": 4,
+ "id": 3326,
+ "area": 456,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1699,
+ "bbox": [
+ 420,
+ 408,
+ 24,
+ 24
+ ],
+ "category_id": 4,
+ "id": 3327,
+ "area": 576,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1700,
+ "bbox": [
+ 58,
+ 227,
+ 287,
+ 118
+ ],
+ "category_id": 16,
+ "id": 3328,
+ "area": 33866,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1701,
+ "bbox": [
+ 186,
+ 147,
+ 30,
+ 27
+ ],
+ "category_id": 4,
+ "id": 3329,
+ "area": 810,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1701,
+ "bbox": [
+ 97,
+ 385,
+ 29,
+ 28
+ ],
+ "category_id": 4,
+ "id": 3330,
+ "area": 812,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1701,
+ "bbox": [
+ 418,
+ 347,
+ 24,
+ 22
+ ],
+ "category_id": 4,
+ "id": 3331,
+ "area": 528,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1702,
+ "bbox": [
+ 144,
+ 188,
+ 115,
+ 150
+ ],
+ "category_id": 15,
+ "id": 3332,
+ "area": 17250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1702,
+ "bbox": [
+ 262,
+ 284,
+ 114,
+ 157
+ ],
+ "category_id": 15,
+ "id": 3333,
+ "area": 17898,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1703,
+ "bbox": [
+ 147,
+ 215,
+ 73,
+ 86
+ ],
+ "category_id": 15,
+ "id": 3334,
+ "area": 6278,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1703,
+ "bbox": [
+ 112,
+ 296,
+ 72,
+ 86
+ ],
+ "category_id": 15,
+ "id": 3335,
+ "area": 6192,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1703,
+ "bbox": [
+ 249,
+ 316,
+ 33,
+ 39
+ ],
+ "category_id": 15,
+ "id": 3336,
+ "area": 1287,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1703,
+ "bbox": [
+ 235,
+ 238,
+ 69,
+ 85
+ ],
+ "category_id": 15,
+ "id": 3337,
+ "area": 5865,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1703,
+ "bbox": [
+ 148,
+ 454,
+ 41,
+ 58
+ ],
+ "category_id": 15,
+ "id": 3338,
+ "area": 2378,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1703,
+ "bbox": [
+ 226,
+ 472,
+ 28,
+ 40
+ ],
+ "category_id": 15,
+ "id": 3339,
+ "area": 1120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1704,
+ "bbox": [
+ 227,
+ 176,
+ 46,
+ 177
+ ],
+ "category_id": 19,
+ "id": 3340,
+ "area": 8142,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1705,
+ "bbox": [
+ 263,
+ 91,
+ 31,
+ 19
+ ],
+ "category_id": 4,
+ "id": 3341,
+ "area": 589,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1705,
+ "bbox": [
+ 0,
+ 240,
+ 79,
+ 44
+ ],
+ "category_id": 4,
+ "id": 3342,
+ "area": 3476,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1705,
+ "bbox": [
+ 453,
+ 56,
+ 40,
+ 20
+ ],
+ "category_id": 4,
+ "id": 3343,
+ "area": 800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1705,
+ "bbox": [
+ 477,
+ 261,
+ 31,
+ 19
+ ],
+ "category_id": 4,
+ "id": 3344,
+ "area": 589,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1705,
+ "bbox": [
+ 456,
+ 398,
+ 33,
+ 18
+ ],
+ "category_id": 4,
+ "id": 3345,
+ "area": 594,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1705,
+ "bbox": [
+ 327,
+ 462,
+ 37,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3346,
+ "area": 777,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1706,
+ "bbox": [
+ 175,
+ 109,
+ 149,
+ 269
+ ],
+ "category_id": 19,
+ "id": 3347,
+ "area": 40081,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1707,
+ "bbox": [
+ 83,
+ 194,
+ 32,
+ 73
+ ],
+ "category_id": 4,
+ "id": 3348,
+ "area": 2336,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1707,
+ "bbox": [
+ 432,
+ 205,
+ 28,
+ 67
+ ],
+ "category_id": 4,
+ "id": 3349,
+ "area": 1876,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1708,
+ "bbox": [
+ 131,
+ 124,
+ 295,
+ 259
+ ],
+ "category_id": 10,
+ "id": 3350,
+ "area": 76405,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1709,
+ "bbox": [
+ 177,
+ 138,
+ 103,
+ 201
+ ],
+ "category_id": 19,
+ "id": 3351,
+ "area": 20703,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1710,
+ "bbox": [
+ 130,
+ 187,
+ 325,
+ 89
+ ],
+ "category_id": 19,
+ "id": 3352,
+ "area": 28925,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1711,
+ "bbox": [
+ 26,
+ 108,
+ 36,
+ 39
+ ],
+ "category_id": 4,
+ "id": 3353,
+ "area": 1404,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1711,
+ "bbox": [
+ 212,
+ 305,
+ 38,
+ 39
+ ],
+ "category_id": 4,
+ "id": 3354,
+ "area": 1482,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1711,
+ "bbox": [
+ 456,
+ 233,
+ 46,
+ 37
+ ],
+ "category_id": 4,
+ "id": 3355,
+ "area": 1702,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1712,
+ "bbox": [
+ 219,
+ 120,
+ 42,
+ 233
+ ],
+ "category_id": 19,
+ "id": 3356,
+ "area": 9786,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1713,
+ "bbox": [
+ 97,
+ 245,
+ 292,
+ 94
+ ],
+ "category_id": 19,
+ "id": 3357,
+ "area": 27448,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1714,
+ "bbox": [
+ 139,
+ 173,
+ 294,
+ 193
+ ],
+ "category_id": 15,
+ "id": 3358,
+ "area": 56742,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1715,
+ "bbox": [
+ 38,
+ 236,
+ 38,
+ 62
+ ],
+ "category_id": 4,
+ "id": 3359,
+ "area": 2356,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1715,
+ "bbox": [
+ 450,
+ 177,
+ 43,
+ 61
+ ],
+ "category_id": 4,
+ "id": 3360,
+ "area": 2623,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1716,
+ "bbox": [
+ 23,
+ 220,
+ 33,
+ 27
+ ],
+ "category_id": 4,
+ "id": 3361,
+ "area": 891,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1716,
+ "bbox": [
+ 296,
+ 379,
+ 32,
+ 29
+ ],
+ "category_id": 4,
+ "id": 3362,
+ "area": 928,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1716,
+ "bbox": [
+ 458,
+ 155,
+ 29,
+ 31
+ ],
+ "category_id": 4,
+ "id": 3363,
+ "area": 899,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1717,
+ "bbox": [
+ 186,
+ 119,
+ 39,
+ 42
+ ],
+ "category_id": 4,
+ "id": 3364,
+ "area": 1638,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1717,
+ "bbox": [
+ 298,
+ 406,
+ 46,
+ 48
+ ],
+ "category_id": 4,
+ "id": 3365,
+ "area": 2208,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1718,
+ "bbox": [
+ 161,
+ 24,
+ 64,
+ 290
+ ],
+ "category_id": 19,
+ "id": 3366,
+ "area": 18560,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1719,
+ "bbox": [
+ 205,
+ 109,
+ 67,
+ 333
+ ],
+ "category_id": 19,
+ "id": 3367,
+ "area": 22311,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1720,
+ "bbox": [
+ 26,
+ 269,
+ 54,
+ 33
+ ],
+ "category_id": 4,
+ "id": 3368,
+ "area": 1782,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1720,
+ "bbox": [
+ 267,
+ 71,
+ 62,
+ 37
+ ],
+ "category_id": 4,
+ "id": 3369,
+ "area": 2294,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1720,
+ "bbox": [
+ 442,
+ 329,
+ 61,
+ 40
+ ],
+ "category_id": 4,
+ "id": 3370,
+ "area": 2440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1721,
+ "bbox": [
+ 119,
+ 174,
+ 220,
+ 144
+ ],
+ "category_id": 19,
+ "id": 3371,
+ "area": 31680,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1722,
+ "bbox": [
+ 155,
+ 87,
+ 180,
+ 274
+ ],
+ "category_id": 19,
+ "id": 3372,
+ "area": 49320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1723,
+ "bbox": [
+ 234,
+ 68,
+ 167,
+ 339
+ ],
+ "category_id": 10,
+ "id": 3373,
+ "area": 56613,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1724,
+ "bbox": [
+ 225,
+ 83,
+ 57,
+ 49
+ ],
+ "category_id": 4,
+ "id": 3374,
+ "area": 2793,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1724,
+ "bbox": [
+ 382,
+ 421,
+ 60,
+ 49
+ ],
+ "category_id": 4,
+ "id": 3375,
+ "area": 2940,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1725,
+ "bbox": [
+ 360,
+ 156,
+ 62,
+ 314
+ ],
+ "category_id": 16,
+ "id": 3376,
+ "area": 19468,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1726,
+ "bbox": [
+ 199,
+ 58,
+ 185,
+ 136
+ ],
+ "category_id": 15,
+ "id": 3377,
+ "area": 25160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1726,
+ "bbox": [
+ 40,
+ 148,
+ 182,
+ 136
+ ],
+ "category_id": 15,
+ "id": 3378,
+ "area": 24752,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1726,
+ "bbox": [
+ 122,
+ 304,
+ 186,
+ 136
+ ],
+ "category_id": 15,
+ "id": 3379,
+ "area": 25296,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1726,
+ "bbox": [
+ 290,
+ 214,
+ 177,
+ 136
+ ],
+ "category_id": 15,
+ "id": 3380,
+ "area": 24072,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1727,
+ "bbox": [
+ 363,
+ 176,
+ 81,
+ 58
+ ],
+ "category_id": 4,
+ "id": 3381,
+ "area": 4698,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1727,
+ "bbox": [
+ 126,
+ 396,
+ 68,
+ 66
+ ],
+ "category_id": 4,
+ "id": 3382,
+ "area": 4488,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1728,
+ "bbox": [
+ 158,
+ 201,
+ 265,
+ 168
+ ],
+ "category_id": 10,
+ "id": 3383,
+ "area": 44520,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1729,
+ "bbox": [
+ 102,
+ 353,
+ 20,
+ 38
+ ],
+ "category_id": 16,
+ "id": 3384,
+ "area": 760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1730,
+ "bbox": [
+ 79,
+ 124,
+ 402,
+ 313
+ ],
+ "category_id": 10,
+ "id": 3385,
+ "area": 125826,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1731,
+ "bbox": [
+ 273,
+ 77,
+ 198,
+ 337
+ ],
+ "category_id": 10,
+ "id": 3386,
+ "area": 66726,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1732,
+ "bbox": [
+ 295,
+ 129,
+ 127,
+ 105
+ ],
+ "category_id": 19,
+ "id": 3387,
+ "area": 13335,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1733,
+ "bbox": [
+ 101,
+ 245,
+ 295,
+ 206
+ ],
+ "category_id": 16,
+ "id": 3388,
+ "area": 60770,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1734,
+ "bbox": [
+ 113,
+ 478,
+ 7,
+ 14
+ ],
+ "category_id": 4,
+ "id": 3389,
+ "area": 98,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1734,
+ "bbox": [
+ 391,
+ 408,
+ 8,
+ 10
+ ],
+ "category_id": 4,
+ "id": 3390,
+ "area": 80,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1734,
+ "bbox": [
+ 353,
+ 355,
+ 6,
+ 8
+ ],
+ "category_id": 4,
+ "id": 3391,
+ "area": 48,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1734,
+ "bbox": [
+ 299,
+ 280,
+ 5,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3392,
+ "area": 35,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1734,
+ "bbox": [
+ 251,
+ 213,
+ 8,
+ 12
+ ],
+ "category_id": 4,
+ "id": 3393,
+ "area": 96,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1734,
+ "bbox": [
+ 199,
+ 286,
+ 8,
+ 10
+ ],
+ "category_id": 4,
+ "id": 3394,
+ "area": 80,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1734,
+ "bbox": [
+ 250,
+ 379,
+ 9,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3395,
+ "area": 81,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1734,
+ "bbox": [
+ 151,
+ 198,
+ 8,
+ 10
+ ],
+ "category_id": 4,
+ "id": 3396,
+ "area": 80,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1734,
+ "bbox": [
+ 107,
+ 118,
+ 6,
+ 10
+ ],
+ "category_id": 4,
+ "id": 3397,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1734,
+ "bbox": [
+ 72,
+ 55,
+ 6,
+ 6
+ ],
+ "category_id": 4,
+ "id": 3398,
+ "area": 36,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1734,
+ "bbox": [
+ 133,
+ 49,
+ 5,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3399,
+ "area": 45,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1735,
+ "bbox": [
+ 85,
+ 250,
+ 329,
+ 37
+ ],
+ "category_id": 19,
+ "id": 3400,
+ "area": 12173,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1736,
+ "bbox": [
+ 144,
+ 126,
+ 129,
+ 98
+ ],
+ "category_id": 19,
+ "id": 3401,
+ "area": 12642,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1737,
+ "bbox": [
+ 169,
+ 167,
+ 175,
+ 241
+ ],
+ "category_id": 16,
+ "id": 3402,
+ "area": 42175,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1738,
+ "bbox": [
+ 129,
+ 26,
+ 180,
+ 162
+ ],
+ "category_id": 15,
+ "id": 3403,
+ "area": 29160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1738,
+ "bbox": [
+ 320,
+ 115,
+ 180,
+ 163
+ ],
+ "category_id": 15,
+ "id": 3404,
+ "area": 29340,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1739,
+ "bbox": [
+ 70,
+ 138,
+ 407,
+ 273
+ ],
+ "category_id": 10,
+ "id": 3405,
+ "area": 111111,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1740,
+ "bbox": [
+ 139,
+ 124,
+ 61,
+ 51
+ ],
+ "category_id": 4,
+ "id": 3406,
+ "area": 3111,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1740,
+ "bbox": [
+ 361,
+ 421,
+ 63,
+ 49
+ ],
+ "category_id": 4,
+ "id": 3407,
+ "area": 3087,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1741,
+ "bbox": [
+ 171,
+ 191,
+ 173,
+ 126
+ ],
+ "category_id": 19,
+ "id": 3408,
+ "area": 21798,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1742,
+ "bbox": [
+ 126,
+ 186,
+ 202,
+ 84
+ ],
+ "category_id": 19,
+ "id": 3409,
+ "area": 16968,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1743,
+ "bbox": [
+ 147,
+ 46,
+ 13,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3410,
+ "area": 91,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1743,
+ "bbox": [
+ 250,
+ 44,
+ 7,
+ 8
+ ],
+ "category_id": 4,
+ "id": 3411,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1743,
+ "bbox": [
+ 271,
+ 119,
+ 9,
+ 10
+ ],
+ "category_id": 4,
+ "id": 3412,
+ "area": 90,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1743,
+ "bbox": [
+ 390,
+ 58,
+ 6,
+ 6
+ ],
+ "category_id": 4,
+ "id": 3413,
+ "area": 36,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1743,
+ "bbox": [
+ 458,
+ 76,
+ 10,
+ 5
+ ],
+ "category_id": 4,
+ "id": 3414,
+ "area": 50,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1743,
+ "bbox": [
+ 504,
+ 167,
+ 6,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3415,
+ "area": 42,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1743,
+ "bbox": [
+ 109,
+ 263,
+ 10,
+ 6
+ ],
+ "category_id": 4,
+ "id": 3416,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1743,
+ "bbox": [
+ 140,
+ 209,
+ 11,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3417,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1743,
+ "bbox": [
+ 33,
+ 267,
+ 9,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3418,
+ "area": 81,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1743,
+ "bbox": [
+ 183,
+ 349,
+ 8,
+ 6
+ ],
+ "category_id": 4,
+ "id": 3419,
+ "area": 48,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1743,
+ "bbox": [
+ 137,
+ 396,
+ 10,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3420,
+ "area": 70,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1743,
+ "bbox": [
+ 245,
+ 431,
+ 9,
+ 6
+ ],
+ "category_id": 4,
+ "id": 3421,
+ "area": 54,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1743,
+ "bbox": [
+ 158,
+ 433,
+ 10,
+ 6
+ ],
+ "category_id": 4,
+ "id": 3422,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1743,
+ "bbox": [
+ 472,
+ 396,
+ 9,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3423,
+ "area": 63,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1743,
+ "bbox": [
+ 419,
+ 359,
+ 9,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3424,
+ "area": 81,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1743,
+ "bbox": [
+ 352,
+ 298,
+ 8,
+ 6
+ ],
+ "category_id": 4,
+ "id": 3425,
+ "area": 48,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1744,
+ "bbox": [
+ 239,
+ 105,
+ 49,
+ 60
+ ],
+ "category_id": 4,
+ "id": 3426,
+ "area": 2940,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1744,
+ "bbox": [
+ 271,
+ 410,
+ 47,
+ 61
+ ],
+ "category_id": 4,
+ "id": 3427,
+ "area": 2867,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1745,
+ "bbox": [
+ 387,
+ 40,
+ 35,
+ 13
+ ],
+ "category_id": 4,
+ "id": 3428,
+ "area": 455,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1745,
+ "bbox": [
+ 83,
+ 126,
+ 33,
+ 13
+ ],
+ "category_id": 4,
+ "id": 3429,
+ "area": 429,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1745,
+ "bbox": [
+ 124,
+ 203,
+ 31,
+ 12
+ ],
+ "category_id": 4,
+ "id": 3430,
+ "area": 372,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1745,
+ "bbox": [
+ 131,
+ 312,
+ 37,
+ 17
+ ],
+ "category_id": 4,
+ "id": 3431,
+ "area": 629,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1745,
+ "bbox": [
+ 102,
+ 439,
+ 28,
+ 17
+ ],
+ "category_id": 4,
+ "id": 3432,
+ "area": 476,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1746,
+ "bbox": [
+ 62,
+ 358,
+ 50,
+ 38
+ ],
+ "category_id": 4,
+ "id": 3433,
+ "area": 1900,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1746,
+ "bbox": [
+ 434,
+ 396,
+ 43,
+ 35
+ ],
+ "category_id": 4,
+ "id": 3434,
+ "area": 1505,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1746,
+ "bbox": [
+ 300,
+ 120,
+ 35,
+ 41
+ ],
+ "category_id": 4,
+ "id": 3435,
+ "area": 1435,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1747,
+ "bbox": [
+ 58,
+ 41,
+ 103,
+ 401
+ ],
+ "category_id": 19,
+ "id": 3436,
+ "area": 41303,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1748,
+ "bbox": [
+ 94,
+ 150,
+ 320,
+ 166
+ ],
+ "category_id": 16,
+ "id": 3437,
+ "area": 53120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1749,
+ "bbox": [
+ 264,
+ 300,
+ 82,
+ 68
+ ],
+ "category_id": 16,
+ "id": 3438,
+ "area": 5576,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1750,
+ "bbox": [
+ 172,
+ 122,
+ 156,
+ 279
+ ],
+ "category_id": 10,
+ "id": 3439,
+ "area": 43524,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1751,
+ "bbox": [
+ 67,
+ 202,
+ 75,
+ 32
+ ],
+ "category_id": 16,
+ "id": 3440,
+ "area": 2400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1752,
+ "bbox": [
+ 80,
+ 87,
+ 141,
+ 174
+ ],
+ "category_id": 15,
+ "id": 3441,
+ "area": 24534,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1752,
+ "bbox": [
+ 271,
+ 60,
+ 154,
+ 186
+ ],
+ "category_id": 15,
+ "id": 3442,
+ "area": 28644,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1752,
+ "bbox": [
+ 97,
+ 275,
+ 152,
+ 185
+ ],
+ "category_id": 15,
+ "id": 3443,
+ "area": 28120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1752,
+ "bbox": [
+ 284,
+ 261,
+ 151,
+ 177
+ ],
+ "category_id": 15,
+ "id": 3444,
+ "area": 26727,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1753,
+ "bbox": [
+ 188,
+ 52,
+ 36,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3445,
+ "area": 756,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1753,
+ "bbox": [
+ 127,
+ 238,
+ 28,
+ 16
+ ],
+ "category_id": 4,
+ "id": 3446,
+ "area": 448,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1753,
+ "bbox": [
+ 88,
+ 414,
+ 32,
+ 18
+ ],
+ "category_id": 4,
+ "id": 3447,
+ "area": 576,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1753,
+ "bbox": [
+ 296,
+ 433,
+ 35,
+ 22
+ ],
+ "category_id": 4,
+ "id": 3448,
+ "area": 770,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1753,
+ "bbox": [
+ 422,
+ 173,
+ 36,
+ 24
+ ],
+ "category_id": 4,
+ "id": 3449,
+ "area": 864,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1754,
+ "bbox": [
+ 101,
+ 73,
+ 89,
+ 368
+ ],
+ "category_id": 10,
+ "id": 3450,
+ "area": 32752,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1755,
+ "bbox": [
+ 108,
+ 114,
+ 252,
+ 317
+ ],
+ "category_id": 10,
+ "id": 3451,
+ "area": 79884,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1756,
+ "bbox": [
+ 225,
+ 368,
+ 68,
+ 27
+ ],
+ "category_id": 4,
+ "id": 3452,
+ "area": 1836,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1756,
+ "bbox": [
+ 350,
+ 116,
+ 64,
+ 33
+ ],
+ "category_id": 4,
+ "id": 3453,
+ "area": 2112,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1757,
+ "bbox": [
+ 66,
+ 118,
+ 82,
+ 109
+ ],
+ "category_id": 15,
+ "id": 3454,
+ "area": 8938,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1757,
+ "bbox": [
+ 94,
+ 252,
+ 84,
+ 112
+ ],
+ "category_id": 15,
+ "id": 3455,
+ "area": 9408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1757,
+ "bbox": [
+ 159,
+ 361,
+ 84,
+ 108
+ ],
+ "category_id": 15,
+ "id": 3456,
+ "area": 9072,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1757,
+ "bbox": [
+ 403,
+ 243,
+ 18,
+ 97
+ ],
+ "category_id": 15,
+ "id": 3457,
+ "area": 1746,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1757,
+ "bbox": [
+ 435,
+ 82,
+ 18,
+ 106
+ ],
+ "category_id": 15,
+ "id": 3458,
+ "area": 1908,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1758,
+ "bbox": [
+ 160,
+ 301,
+ 64,
+ 76
+ ],
+ "category_id": 15,
+ "id": 3459,
+ "area": 4864,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1758,
+ "bbox": [
+ 243,
+ 304,
+ 61,
+ 76
+ ],
+ "category_id": 15,
+ "id": 3460,
+ "area": 4636,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1758,
+ "bbox": [
+ 224,
+ 408,
+ 56,
+ 74
+ ],
+ "category_id": 15,
+ "id": 3461,
+ "area": 4144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1759,
+ "bbox": [
+ 88,
+ 206,
+ 322,
+ 53
+ ],
+ "category_id": 19,
+ "id": 3462,
+ "area": 17066,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1760,
+ "bbox": [
+ 353,
+ 113,
+ 79,
+ 97
+ ],
+ "category_id": 19,
+ "id": 3463,
+ "area": 7663,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1761,
+ "bbox": [
+ 174,
+ 220,
+ 132,
+ 110
+ ],
+ "category_id": 15,
+ "id": 3464,
+ "area": 14520,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1761,
+ "bbox": [
+ 456,
+ 248,
+ 56,
+ 33
+ ],
+ "category_id": 15,
+ "id": 3465,
+ "area": 1848,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1762,
+ "bbox": [
+ 197,
+ 57,
+ 149,
+ 180
+ ],
+ "category_id": 15,
+ "id": 3466,
+ "area": 26820,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1762,
+ "bbox": [
+ 243,
+ 278,
+ 153,
+ 180
+ ],
+ "category_id": 15,
+ "id": 3467,
+ "area": 27540,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1763,
+ "bbox": [
+ 200,
+ 259,
+ 90,
+ 87
+ ],
+ "category_id": 4,
+ "id": 3468,
+ "area": 7830,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1764,
+ "bbox": [
+ 156,
+ 147,
+ 68,
+ 84
+ ],
+ "category_id": 16,
+ "id": 3469,
+ "area": 5712,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1765,
+ "bbox": [
+ 201,
+ 79,
+ 106,
+ 328
+ ],
+ "category_id": 10,
+ "id": 3470,
+ "area": 34768,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1766,
+ "bbox": [
+ 240,
+ 140,
+ 54,
+ 46
+ ],
+ "category_id": 4,
+ "id": 3471,
+ "area": 2484,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1766,
+ "bbox": [
+ 289,
+ 406,
+ 49,
+ 40
+ ],
+ "category_id": 4,
+ "id": 3472,
+ "area": 1960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1767,
+ "bbox": [
+ 264,
+ 92,
+ 40,
+ 290
+ ],
+ "category_id": 19,
+ "id": 3473,
+ "area": 11600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1768,
+ "bbox": [
+ 8,
+ 145,
+ 429,
+ 183
+ ],
+ "category_id": 19,
+ "id": 3474,
+ "area": 78507,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1769,
+ "bbox": [
+ 217,
+ 113,
+ 210,
+ 299
+ ],
+ "category_id": 10,
+ "id": 3475,
+ "area": 62790,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1770,
+ "bbox": [
+ 91,
+ 128,
+ 144,
+ 31
+ ],
+ "category_id": 16,
+ "id": 3476,
+ "area": 4464,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1771,
+ "bbox": [
+ 103,
+ 94,
+ 11,
+ 19
+ ],
+ "category_id": 4,
+ "id": 3477,
+ "area": 209,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1771,
+ "bbox": [
+ 203,
+ 205,
+ 15,
+ 23
+ ],
+ "category_id": 4,
+ "id": 3478,
+ "area": 345,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1771,
+ "bbox": [
+ 151,
+ 312,
+ 10,
+ 23
+ ],
+ "category_id": 4,
+ "id": 3479,
+ "area": 230,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1771,
+ "bbox": [
+ 34,
+ 440,
+ 19,
+ 27
+ ],
+ "category_id": 4,
+ "id": 3480,
+ "area": 513,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1771,
+ "bbox": [
+ 478,
+ 97,
+ 21,
+ 30
+ ],
+ "category_id": 4,
+ "id": 3481,
+ "area": 630,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1772,
+ "bbox": [
+ 169,
+ 116,
+ 240,
+ 264
+ ],
+ "category_id": 10,
+ "id": 3482,
+ "area": 63360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1773,
+ "bbox": [
+ 162,
+ 42,
+ 25,
+ 29
+ ],
+ "category_id": 4,
+ "id": 3483,
+ "area": 725,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1773,
+ "bbox": [
+ 216,
+ 113,
+ 27,
+ 32
+ ],
+ "category_id": 4,
+ "id": 3484,
+ "area": 864,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1773,
+ "bbox": [
+ 450,
+ 186,
+ 24,
+ 25
+ ],
+ "category_id": 4,
+ "id": 3485,
+ "area": 600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1773,
+ "bbox": [
+ 448,
+ 326,
+ 27,
+ 31
+ ],
+ "category_id": 4,
+ "id": 3486,
+ "area": 837,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1773,
+ "bbox": [
+ 171,
+ 407,
+ 21,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3487,
+ "area": 441,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1773,
+ "bbox": [
+ 58,
+ 467,
+ 25,
+ 24
+ ],
+ "category_id": 4,
+ "id": 3488,
+ "area": 600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1774,
+ "bbox": [
+ 185,
+ 99,
+ 209,
+ 368
+ ],
+ "category_id": 10,
+ "id": 3489,
+ "area": 76912,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1775,
+ "bbox": [
+ 171,
+ 129,
+ 162,
+ 207
+ ],
+ "category_id": 10,
+ "id": 3490,
+ "area": 33534,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1776,
+ "bbox": [
+ 35,
+ 230,
+ 10,
+ 8
+ ],
+ "category_id": 4,
+ "id": 3491,
+ "area": 80,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1776,
+ "bbox": [
+ 19,
+ 300,
+ 9,
+ 8
+ ],
+ "category_id": 4,
+ "id": 3492,
+ "area": 72,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1776,
+ "bbox": [
+ 88,
+ 337,
+ 8,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3493,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1776,
+ "bbox": [
+ 130,
+ 337,
+ 10,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3494,
+ "area": 70,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1776,
+ "bbox": [
+ 174,
+ 390,
+ 10,
+ 6
+ ],
+ "category_id": 4,
+ "id": 3495,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1776,
+ "bbox": [
+ 209,
+ 435,
+ 9,
+ 8
+ ],
+ "category_id": 4,
+ "id": 3496,
+ "area": 72,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1776,
+ "bbox": [
+ 61,
+ 179,
+ 9,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3497,
+ "area": 63,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1776,
+ "bbox": [
+ 72,
+ 123,
+ 11,
+ 6
+ ],
+ "category_id": 4,
+ "id": 3498,
+ "area": 66,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1776,
+ "bbox": [
+ 190,
+ 228,
+ 7,
+ 8
+ ],
+ "category_id": 4,
+ "id": 3499,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1776,
+ "bbox": [
+ 243,
+ 224,
+ 11,
+ 8
+ ],
+ "category_id": 4,
+ "id": 3500,
+ "area": 88,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1776,
+ "bbox": [
+ 304,
+ 241,
+ 12,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3501,
+ "area": 84,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1776,
+ "bbox": [
+ 445,
+ 293,
+ 11,
+ 10
+ ],
+ "category_id": 4,
+ "id": 3502,
+ "area": 110,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1776,
+ "bbox": [
+ 197,
+ 23,
+ 8,
+ 5
+ ],
+ "category_id": 4,
+ "id": 3503,
+ "area": 40,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1776,
+ "bbox": [
+ 236,
+ 85,
+ 11,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3504,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1776,
+ "bbox": [
+ 285,
+ 168,
+ 11,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3505,
+ "area": 121,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1776,
+ "bbox": [
+ 298,
+ 62,
+ 9,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3506,
+ "area": 81,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1776,
+ "bbox": [
+ 390,
+ 87,
+ 10,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3507,
+ "area": 70,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1776,
+ "bbox": [
+ 403,
+ 48,
+ 11,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3508,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1776,
+ "bbox": [
+ 457,
+ 111,
+ 10,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3509,
+ "area": 70,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1777,
+ "bbox": [
+ 118,
+ 49,
+ 319,
+ 269
+ ],
+ "category_id": 19,
+ "id": 3510,
+ "area": 85811,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1778,
+ "bbox": [
+ 378,
+ 110,
+ 62,
+ 42
+ ],
+ "category_id": 4,
+ "id": 3511,
+ "area": 2604,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1778,
+ "bbox": [
+ 149,
+ 381,
+ 55,
+ 36
+ ],
+ "category_id": 4,
+ "id": 3512,
+ "area": 1980,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1779,
+ "bbox": [
+ 243,
+ 175,
+ 144,
+ 124
+ ],
+ "category_id": 15,
+ "id": 3513,
+ "area": 17856,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1779,
+ "bbox": [
+ 262,
+ 336,
+ 122,
+ 122
+ ],
+ "category_id": 15,
+ "id": 3514,
+ "area": 14884,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1780,
+ "bbox": [
+ 166,
+ 180,
+ 251,
+ 207
+ ],
+ "category_id": 16,
+ "id": 3515,
+ "area": 51957,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1781,
+ "bbox": [
+ 218,
+ 40,
+ 56,
+ 47
+ ],
+ "category_id": 4,
+ "id": 3516,
+ "area": 2632,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1781,
+ "bbox": [
+ 13,
+ 296,
+ 51,
+ 46
+ ],
+ "category_id": 4,
+ "id": 3517,
+ "area": 2346,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1781,
+ "bbox": [
+ 314,
+ 460,
+ 48,
+ 46
+ ],
+ "category_id": 4,
+ "id": 3518,
+ "area": 2208,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1782,
+ "bbox": [
+ 174,
+ 92,
+ 34,
+ 34
+ ],
+ "category_id": 4,
+ "id": 3519,
+ "area": 1156,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1782,
+ "bbox": [
+ 309,
+ 397,
+ 24,
+ 38
+ ],
+ "category_id": 4,
+ "id": 3520,
+ "area": 912,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1783,
+ "bbox": [
+ 305,
+ 215,
+ 47,
+ 116
+ ],
+ "category_id": 4,
+ "id": 3521,
+ "area": 5452,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1784,
+ "bbox": [
+ 113,
+ 27,
+ 15,
+ 13
+ ],
+ "category_id": 4,
+ "id": 3522,
+ "area": 195,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1784,
+ "bbox": [
+ 336,
+ 13,
+ 16,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3523,
+ "area": 176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1784,
+ "bbox": [
+ 476,
+ 5,
+ 12,
+ 12
+ ],
+ "category_id": 4,
+ "id": 3524,
+ "area": 144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1784,
+ "bbox": [
+ 227,
+ 117,
+ 15,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3525,
+ "area": 165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1784,
+ "bbox": [
+ 383,
+ 122,
+ 15,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3526,
+ "area": 165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1784,
+ "bbox": [
+ 133,
+ 220,
+ 19,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3527,
+ "area": 209,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1784,
+ "bbox": [
+ 282,
+ 369,
+ 16,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3528,
+ "area": 176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1784,
+ "bbox": [
+ 213,
+ 426,
+ 16,
+ 14
+ ],
+ "category_id": 4,
+ "id": 3529,
+ "area": 224,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1784,
+ "bbox": [
+ 121,
+ 491,
+ 14,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3530,
+ "area": 154,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1785,
+ "bbox": [
+ 192,
+ 268,
+ 95,
+ 88
+ ],
+ "category_id": 4,
+ "id": 3531,
+ "area": 8360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1786,
+ "bbox": [
+ 156,
+ 221,
+ 100,
+ 90
+ ],
+ "category_id": 4,
+ "id": 3532,
+ "area": 9000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1787,
+ "bbox": [
+ 265,
+ 116,
+ 104,
+ 155
+ ],
+ "category_id": 15,
+ "id": 3533,
+ "area": 16120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1787,
+ "bbox": [
+ 105,
+ 163,
+ 106,
+ 162
+ ],
+ "category_id": 15,
+ "id": 3534,
+ "area": 17172,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1788,
+ "bbox": [
+ 262,
+ 320,
+ 146,
+ 55
+ ],
+ "category_id": 4,
+ "id": 3535,
+ "area": 8030,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1789,
+ "bbox": [
+ 67,
+ 205,
+ 247,
+ 169
+ ],
+ "category_id": 16,
+ "id": 3536,
+ "area": 41743,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1790,
+ "bbox": [
+ 221,
+ 92,
+ 113,
+ 144
+ ],
+ "category_id": 15,
+ "id": 3537,
+ "area": 16272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1790,
+ "bbox": [
+ 76,
+ 157,
+ 114,
+ 143
+ ],
+ "category_id": 15,
+ "id": 3538,
+ "area": 16302,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1790,
+ "bbox": [
+ 293,
+ 240,
+ 114,
+ 140
+ ],
+ "category_id": 15,
+ "id": 3539,
+ "area": 15960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1791,
+ "bbox": [
+ 177,
+ 221,
+ 144,
+ 121
+ ],
+ "category_id": 15,
+ "id": 3540,
+ "area": 17424,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1792,
+ "bbox": [
+ 220,
+ 105,
+ 128,
+ 156
+ ],
+ "category_id": 15,
+ "id": 3541,
+ "area": 19968,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1792,
+ "bbox": [
+ 204,
+ 278,
+ 132,
+ 157
+ ],
+ "category_id": 15,
+ "id": 3542,
+ "area": 20724,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1793,
+ "bbox": [
+ 244,
+ 198,
+ 143,
+ 124
+ ],
+ "category_id": 15,
+ "id": 3543,
+ "area": 17732,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1794,
+ "bbox": [
+ 257,
+ 100,
+ 100,
+ 310
+ ],
+ "category_id": 19,
+ "id": 3544,
+ "area": 31000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1795,
+ "bbox": [
+ 208,
+ 253,
+ 147,
+ 57
+ ],
+ "category_id": 4,
+ "id": 3545,
+ "area": 8379,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1796,
+ "bbox": [
+ 90,
+ 45,
+ 70,
+ 90
+ ],
+ "category_id": 15,
+ "id": 3546,
+ "area": 6300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1796,
+ "bbox": [
+ 84,
+ 139,
+ 72,
+ 91
+ ],
+ "category_id": 15,
+ "id": 3547,
+ "area": 6552,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1796,
+ "bbox": [
+ 97,
+ 231,
+ 96,
+ 121
+ ],
+ "category_id": 15,
+ "id": 3548,
+ "area": 11616,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1796,
+ "bbox": [
+ 256,
+ 297,
+ 104,
+ 119
+ ],
+ "category_id": 15,
+ "id": 3549,
+ "area": 12376,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1796,
+ "bbox": [
+ 375,
+ 218,
+ 103,
+ 122
+ ],
+ "category_id": 15,
+ "id": 3550,
+ "area": 12566,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1797,
+ "bbox": [
+ 71,
+ 236,
+ 362,
+ 90
+ ],
+ "category_id": 16,
+ "id": 3551,
+ "area": 32580,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1798,
+ "bbox": [
+ 236,
+ 19,
+ 182,
+ 151
+ ],
+ "category_id": 15,
+ "id": 3552,
+ "area": 27482,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1798,
+ "bbox": [
+ 163,
+ 227,
+ 121,
+ 100
+ ],
+ "category_id": 15,
+ "id": 3553,
+ "area": 12100,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1798,
+ "bbox": [
+ 142,
+ 358,
+ 110,
+ 102
+ ],
+ "category_id": 15,
+ "id": 3554,
+ "area": 11220,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1798,
+ "bbox": [
+ 250,
+ 236,
+ 98,
+ 32
+ ],
+ "category_id": 15,
+ "id": 3555,
+ "area": 3136,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1798,
+ "bbox": [
+ 372,
+ 0,
+ 115,
+ 26
+ ],
+ "category_id": 15,
+ "id": 3556,
+ "area": 2990,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1799,
+ "bbox": [
+ 34,
+ 252,
+ 33,
+ 76
+ ],
+ "category_id": 19,
+ "id": 3557,
+ "area": 2508,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1800,
+ "bbox": [
+ 103,
+ 228,
+ 325,
+ 52
+ ],
+ "category_id": 10,
+ "id": 3558,
+ "area": 16900,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1801,
+ "bbox": [
+ 221,
+ 275,
+ 139,
+ 123
+ ],
+ "category_id": 15,
+ "id": 3559,
+ "area": 17097,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1801,
+ "bbox": [
+ 133,
+ 254,
+ 57,
+ 35
+ ],
+ "category_id": 15,
+ "id": 3560,
+ "area": 1995,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1802,
+ "bbox": [
+ 103,
+ 146,
+ 356,
+ 241
+ ],
+ "category_id": 10,
+ "id": 3561,
+ "area": 85796,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1803,
+ "bbox": [
+ 251,
+ 191,
+ 28,
+ 118
+ ],
+ "category_id": 19,
+ "id": 3562,
+ "area": 3304,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1804,
+ "bbox": [
+ 116,
+ 216,
+ 218,
+ 38
+ ],
+ "category_id": 19,
+ "id": 3563,
+ "area": 8284,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1805,
+ "bbox": [
+ 67,
+ 96,
+ 117,
+ 341
+ ],
+ "category_id": 10,
+ "id": 3564,
+ "area": 39897,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1806,
+ "bbox": [
+ 227,
+ 134,
+ 131,
+ 119
+ ],
+ "category_id": 15,
+ "id": 3565,
+ "area": 15589,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1806,
+ "bbox": [
+ 236,
+ 335,
+ 127,
+ 115
+ ],
+ "category_id": 15,
+ "id": 3566,
+ "area": 14605,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1806,
+ "bbox": [
+ 351,
+ 275,
+ 80,
+ 64
+ ],
+ "category_id": 15,
+ "id": 3567,
+ "area": 5120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1807,
+ "bbox": [
+ 181,
+ 104,
+ 265,
+ 313
+ ],
+ "category_id": 10,
+ "id": 3568,
+ "area": 82945,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1808,
+ "bbox": [
+ 140,
+ 58,
+ 346,
+ 37
+ ],
+ "category_id": 10,
+ "id": 3569,
+ "area": 12802,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1809,
+ "bbox": [
+ 154,
+ 209,
+ 206,
+ 59
+ ],
+ "category_id": 10,
+ "id": 3570,
+ "area": 12154,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1810,
+ "bbox": [
+ 80,
+ 128,
+ 313,
+ 222
+ ],
+ "category_id": 10,
+ "id": 3571,
+ "area": 69486,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1811,
+ "bbox": [
+ 120,
+ 106,
+ 238,
+ 238
+ ],
+ "category_id": 16,
+ "id": 3572,
+ "area": 56644,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1812,
+ "bbox": [
+ 275,
+ 110,
+ 40,
+ 289
+ ],
+ "category_id": 10,
+ "id": 3573,
+ "area": 11560,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1813,
+ "bbox": [
+ 94,
+ 105,
+ 400,
+ 231
+ ],
+ "category_id": 10,
+ "id": 3574,
+ "area": 92400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1814,
+ "bbox": [
+ 215,
+ 195,
+ 218,
+ 189
+ ],
+ "category_id": 10,
+ "id": 3575,
+ "area": 41202,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1815,
+ "bbox": [
+ 240,
+ 72,
+ 125,
+ 338
+ ],
+ "category_id": 19,
+ "id": 3576,
+ "area": 42250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1816,
+ "bbox": [
+ 209,
+ 388,
+ 46,
+ 85
+ ],
+ "category_id": 16,
+ "id": 3577,
+ "area": 3910,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1817,
+ "bbox": [
+ 240,
+ 437,
+ 29,
+ 12
+ ],
+ "category_id": 16,
+ "id": 3578,
+ "area": 348,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1818,
+ "bbox": [
+ 208,
+ 160,
+ 113,
+ 173
+ ],
+ "category_id": 10,
+ "id": 3579,
+ "area": 19549,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1819,
+ "bbox": [
+ 136,
+ 186,
+ 312,
+ 136
+ ],
+ "category_id": 16,
+ "id": 3580,
+ "area": 42432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1820,
+ "bbox": [
+ 120,
+ 231,
+ 228,
+ 42
+ ],
+ "category_id": 19,
+ "id": 3581,
+ "area": 9576,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1821,
+ "bbox": [
+ 268,
+ 157,
+ 167,
+ 129
+ ],
+ "category_id": 19,
+ "id": 3582,
+ "area": 21543,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1822,
+ "bbox": [
+ 76,
+ 253,
+ 148,
+ 115
+ ],
+ "category_id": 15,
+ "id": 3583,
+ "area": 17020,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1822,
+ "bbox": [
+ 241,
+ 199,
+ 111,
+ 24
+ ],
+ "category_id": 15,
+ "id": 3584,
+ "area": 2664,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1822,
+ "bbox": [
+ 240,
+ 303,
+ 67,
+ 15
+ ],
+ "category_id": 15,
+ "id": 3585,
+ "area": 1005,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1822,
+ "bbox": [
+ 288,
+ 115,
+ 82,
+ 60
+ ],
+ "category_id": 15,
+ "id": 3586,
+ "area": 4920,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1822,
+ "bbox": [
+ 371,
+ 133,
+ 83,
+ 62
+ ],
+ "category_id": 15,
+ "id": 3587,
+ "area": 5146,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1823,
+ "bbox": [
+ 30,
+ 115,
+ 35,
+ 40
+ ],
+ "category_id": 4,
+ "id": 3588,
+ "area": 1400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1823,
+ "bbox": [
+ 254,
+ 81,
+ 21,
+ 41
+ ],
+ "category_id": 4,
+ "id": 3589,
+ "area": 861,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1823,
+ "bbox": [
+ 443,
+ 363,
+ 24,
+ 44
+ ],
+ "category_id": 4,
+ "id": 3590,
+ "area": 1056,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1824,
+ "bbox": [
+ 249,
+ 252,
+ 102,
+ 112
+ ],
+ "category_id": 15,
+ "id": 3591,
+ "area": 11424,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1825,
+ "bbox": [
+ 273,
+ 91,
+ 145,
+ 324
+ ],
+ "category_id": 19,
+ "id": 3592,
+ "area": 46980,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1826,
+ "bbox": [
+ 89,
+ 294,
+ 67,
+ 153
+ ],
+ "category_id": 10,
+ "id": 3593,
+ "area": 10251,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1827,
+ "bbox": [
+ 65,
+ 306,
+ 31,
+ 17
+ ],
+ "category_id": 4,
+ "id": 3594,
+ "area": 527,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1827,
+ "bbox": [
+ 196,
+ 327,
+ 17,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3595,
+ "area": 119,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1827,
+ "bbox": [
+ 285,
+ 347,
+ 18,
+ 10
+ ],
+ "category_id": 4,
+ "id": 3596,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1828,
+ "bbox": [
+ 202,
+ 83,
+ 83,
+ 278
+ ],
+ "category_id": 19,
+ "id": 3597,
+ "area": 23074,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1829,
+ "bbox": [
+ 157,
+ 44,
+ 142,
+ 444
+ ],
+ "category_id": 19,
+ "id": 3598,
+ "area": 63048,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1830,
+ "bbox": [
+ 112,
+ 173,
+ 65,
+ 83
+ ],
+ "category_id": 4,
+ "id": 3599,
+ "area": 5395,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1830,
+ "bbox": [
+ 347,
+ 267,
+ 73,
+ 87
+ ],
+ "category_id": 4,
+ "id": 3600,
+ "area": 6351,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1831,
+ "bbox": [
+ 138,
+ 383,
+ 13,
+ 27
+ ],
+ "category_id": 16,
+ "id": 3601,
+ "area": 351,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1831,
+ "bbox": [
+ 407,
+ 154,
+ 23,
+ 23
+ ],
+ "category_id": 16,
+ "id": 3602,
+ "area": 529,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1832,
+ "bbox": [
+ 76,
+ 245,
+ 409,
+ 35
+ ],
+ "category_id": 10,
+ "id": 3603,
+ "area": 14315,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1833,
+ "bbox": [
+ 51,
+ 88,
+ 29,
+ 12
+ ],
+ "category_id": 4,
+ "id": 3604,
+ "area": 348,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1833,
+ "bbox": [
+ 208,
+ 149,
+ 28,
+ 13
+ ],
+ "category_id": 4,
+ "id": 3605,
+ "area": 364,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1833,
+ "bbox": [
+ 286,
+ 212,
+ 28,
+ 15
+ ],
+ "category_id": 4,
+ "id": 3606,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1833,
+ "bbox": [
+ 408,
+ 239,
+ 31,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3607,
+ "area": 341,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1833,
+ "bbox": [
+ 440,
+ 76,
+ 29,
+ 14
+ ],
+ "category_id": 4,
+ "id": 3608,
+ "area": 406,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1834,
+ "bbox": [
+ 327,
+ 280,
+ 112,
+ 42
+ ],
+ "category_id": 19,
+ "id": 3609,
+ "area": 4704,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1835,
+ "bbox": [
+ 108,
+ 175,
+ 153,
+ 185
+ ],
+ "category_id": 15,
+ "id": 3610,
+ "area": 28305,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1835,
+ "bbox": [
+ 314,
+ 195,
+ 155,
+ 183
+ ],
+ "category_id": 15,
+ "id": 3611,
+ "area": 28365,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1836,
+ "bbox": [
+ 272,
+ 181,
+ 28,
+ 37
+ ],
+ "category_id": 15,
+ "id": 3612,
+ "area": 1036,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1837,
+ "bbox": [
+ 204,
+ 369,
+ 108,
+ 126
+ ],
+ "category_id": 15,
+ "id": 3613,
+ "area": 13608,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1838,
+ "bbox": [
+ 125,
+ 184,
+ 242,
+ 162
+ ],
+ "category_id": 16,
+ "id": 3614,
+ "area": 39204,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1839,
+ "bbox": [
+ 68,
+ 96,
+ 79,
+ 37
+ ],
+ "category_id": 4,
+ "id": 3615,
+ "area": 2923,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1839,
+ "bbox": [
+ 387,
+ 293,
+ 87,
+ 50
+ ],
+ "category_id": 4,
+ "id": 3616,
+ "area": 4350,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1840,
+ "bbox": [
+ 99,
+ 167,
+ 388,
+ 156
+ ],
+ "category_id": 10,
+ "id": 3617,
+ "area": 60528,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1841,
+ "bbox": [
+ 90,
+ 103,
+ 377,
+ 281
+ ],
+ "category_id": 10,
+ "id": 3618,
+ "area": 105937,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1842,
+ "bbox": [
+ 124,
+ 204,
+ 267,
+ 96
+ ],
+ "category_id": 19,
+ "id": 3619,
+ "area": 25632,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1843,
+ "bbox": [
+ 214,
+ 143,
+ 44,
+ 260
+ ],
+ "category_id": 19,
+ "id": 3620,
+ "area": 11440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1844,
+ "bbox": [
+ 205,
+ 70,
+ 72,
+ 41
+ ],
+ "category_id": 4,
+ "id": 3621,
+ "area": 2952,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1844,
+ "bbox": [
+ 286,
+ 384,
+ 23,
+ 81
+ ],
+ "category_id": 4,
+ "id": 3622,
+ "area": 1863,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1845,
+ "bbox": [
+ 193,
+ 87,
+ 57,
+ 47
+ ],
+ "category_id": 4,
+ "id": 3623,
+ "area": 2679,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1845,
+ "bbox": [
+ 329,
+ 414,
+ 63,
+ 57
+ ],
+ "category_id": 4,
+ "id": 3624,
+ "area": 3591,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1846,
+ "bbox": [
+ 204,
+ 251,
+ 55,
+ 88
+ ],
+ "category_id": 4,
+ "id": 3625,
+ "area": 4840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1847,
+ "bbox": [
+ 40,
+ 131,
+ 27,
+ 18
+ ],
+ "category_id": 4,
+ "id": 3626,
+ "area": 486,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1847,
+ "bbox": [
+ 487,
+ 174,
+ 24,
+ 13
+ ],
+ "category_id": 4,
+ "id": 3627,
+ "area": 312,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1847,
+ "bbox": [
+ 149,
+ 243,
+ 35,
+ 14
+ ],
+ "category_id": 4,
+ "id": 3628,
+ "area": 490,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1847,
+ "bbox": [
+ 273,
+ 325,
+ 37,
+ 12
+ ],
+ "category_id": 4,
+ "id": 3629,
+ "area": 444,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1847,
+ "bbox": [
+ 459,
+ 421,
+ 30,
+ 14
+ ],
+ "category_id": 4,
+ "id": 3630,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1848,
+ "bbox": [
+ 41,
+ 175,
+ 136,
+ 161
+ ],
+ "category_id": 15,
+ "id": 3631,
+ "area": 21896,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1848,
+ "bbox": [
+ 220,
+ 211,
+ 137,
+ 161
+ ],
+ "category_id": 15,
+ "id": 3632,
+ "area": 22057,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1849,
+ "bbox": [
+ 169,
+ 147,
+ 231,
+ 103
+ ],
+ "category_id": 16,
+ "id": 3633,
+ "area": 23793,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1850,
+ "bbox": [
+ 76,
+ 94,
+ 50,
+ 25
+ ],
+ "category_id": 4,
+ "id": 3634,
+ "area": 1250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1850,
+ "bbox": [
+ 335,
+ 181,
+ 46,
+ 60
+ ],
+ "category_id": 4,
+ "id": 3635,
+ "area": 2760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1850,
+ "bbox": [
+ 392,
+ 253,
+ 43,
+ 39
+ ],
+ "category_id": 4,
+ "id": 3636,
+ "area": 1677,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1850,
+ "bbox": [
+ 202,
+ 416,
+ 37,
+ 37
+ ],
+ "category_id": 4,
+ "id": 3637,
+ "area": 1369,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1851,
+ "bbox": [
+ 60,
+ 354,
+ 37,
+ 30
+ ],
+ "category_id": 4,
+ "id": 3638,
+ "area": 1110,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1851,
+ "bbox": [
+ 382,
+ 308,
+ 32,
+ 29
+ ],
+ "category_id": 4,
+ "id": 3639,
+ "area": 928,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1851,
+ "bbox": [
+ 430,
+ 174,
+ 24,
+ 28
+ ],
+ "category_id": 4,
+ "id": 3640,
+ "area": 672,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1852,
+ "bbox": [
+ 400,
+ 28,
+ 6,
+ 16
+ ],
+ "category_id": 4,
+ "id": 3641,
+ "area": 96,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1852,
+ "bbox": [
+ 385,
+ 70,
+ 6,
+ 10
+ ],
+ "category_id": 4,
+ "id": 3642,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1852,
+ "bbox": [
+ 383,
+ 122,
+ 7,
+ 14
+ ],
+ "category_id": 4,
+ "id": 3643,
+ "area": 98,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1852,
+ "bbox": [
+ 376,
+ 184,
+ 8,
+ 13
+ ],
+ "category_id": 4,
+ "id": 3644,
+ "area": 104,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1852,
+ "bbox": [
+ 310,
+ 257,
+ 7,
+ 8
+ ],
+ "category_id": 4,
+ "id": 3645,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1852,
+ "bbox": [
+ 371,
+ 314,
+ 5,
+ 13
+ ],
+ "category_id": 4,
+ "id": 3646,
+ "area": 65,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1852,
+ "bbox": [
+ 348,
+ 408,
+ 7,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3647,
+ "area": 49,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1853,
+ "bbox": [
+ 173,
+ 251,
+ 150,
+ 63
+ ],
+ "category_id": 19,
+ "id": 3648,
+ "area": 9450,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1854,
+ "bbox": [
+ 37,
+ 140,
+ 444,
+ 224
+ ],
+ "category_id": 16,
+ "id": 3649,
+ "area": 99456,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1855,
+ "bbox": [
+ 192,
+ 25,
+ 106,
+ 36
+ ],
+ "category_id": 15,
+ "id": 3650,
+ "area": 3816,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1855,
+ "bbox": [
+ 320,
+ 145,
+ 147,
+ 123
+ ],
+ "category_id": 15,
+ "id": 3651,
+ "area": 18081,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1855,
+ "bbox": [
+ 210,
+ 250,
+ 152,
+ 105
+ ],
+ "category_id": 15,
+ "id": 3652,
+ "area": 15960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1855,
+ "bbox": [
+ 50,
+ 252,
+ 104,
+ 86
+ ],
+ "category_id": 15,
+ "id": 3653,
+ "area": 8944,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1855,
+ "bbox": [
+ 107,
+ 333,
+ 103,
+ 88
+ ],
+ "category_id": 15,
+ "id": 3654,
+ "area": 9064,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1856,
+ "bbox": [
+ 166,
+ 138,
+ 151,
+ 237
+ ],
+ "category_id": 16,
+ "id": 3655,
+ "area": 35787,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1857,
+ "bbox": [
+ 27,
+ 229,
+ 82,
+ 274
+ ],
+ "category_id": 10,
+ "id": 3656,
+ "area": 22468,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1858,
+ "bbox": [
+ 180,
+ 120,
+ 107,
+ 287
+ ],
+ "category_id": 16,
+ "id": 3657,
+ "area": 30709,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1859,
+ "bbox": [
+ 197,
+ 184,
+ 40,
+ 121
+ ],
+ "category_id": 19,
+ "id": 3658,
+ "area": 4840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1860,
+ "bbox": [
+ 158,
+ 110,
+ 181,
+ 292
+ ],
+ "category_id": 10,
+ "id": 3659,
+ "area": 52852,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1861,
+ "bbox": [
+ 289,
+ 62,
+ 80,
+ 375
+ ],
+ "category_id": 10,
+ "id": 3660,
+ "area": 30000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1862,
+ "bbox": [
+ 70,
+ 138,
+ 19,
+ 13
+ ],
+ "category_id": 4,
+ "id": 3661,
+ "area": 247,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1862,
+ "bbox": [
+ 37,
+ 267,
+ 18,
+ 13
+ ],
+ "category_id": 4,
+ "id": 3662,
+ "area": 234,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1862,
+ "bbox": [
+ 158,
+ 389,
+ 13,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3663,
+ "area": 117,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1862,
+ "bbox": [
+ 64,
+ 430,
+ 16,
+ 13
+ ],
+ "category_id": 4,
+ "id": 3664,
+ "area": 208,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1862,
+ "bbox": [
+ 442,
+ 20,
+ 20,
+ 17
+ ],
+ "category_id": 4,
+ "id": 3665,
+ "area": 340,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1862,
+ "bbox": [
+ 437,
+ 129,
+ 19,
+ 17
+ ],
+ "category_id": 4,
+ "id": 3666,
+ "area": 323,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1862,
+ "bbox": [
+ 419,
+ 312,
+ 24,
+ 18
+ ],
+ "category_id": 4,
+ "id": 3667,
+ "area": 432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1862,
+ "bbox": [
+ 331,
+ 375,
+ 13,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3668,
+ "area": 117,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1862,
+ "bbox": [
+ 322,
+ 446,
+ 18,
+ 12
+ ],
+ "category_id": 4,
+ "id": 3669,
+ "area": 216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1862,
+ "bbox": [
+ 484,
+ 412,
+ 17,
+ 13
+ ],
+ "category_id": 4,
+ "id": 3670,
+ "area": 221,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1862,
+ "bbox": [
+ 419,
+ 485,
+ 17,
+ 13
+ ],
+ "category_id": 4,
+ "id": 3671,
+ "area": 221,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1863,
+ "bbox": [
+ 176,
+ 124,
+ 120,
+ 106
+ ],
+ "category_id": 15,
+ "id": 3672,
+ "area": 12720,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1863,
+ "bbox": [
+ 162,
+ 282,
+ 120,
+ 105
+ ],
+ "category_id": 15,
+ "id": 3673,
+ "area": 12600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1864,
+ "bbox": [
+ 46,
+ 144,
+ 413,
+ 22
+ ],
+ "category_id": 10,
+ "id": 3674,
+ "area": 9086,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1865,
+ "bbox": [
+ 110,
+ 218,
+ 251,
+ 96
+ ],
+ "category_id": 19,
+ "id": 3675,
+ "area": 24096,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1866,
+ "bbox": [
+ 58,
+ 142,
+ 163,
+ 130
+ ],
+ "category_id": 15,
+ "id": 3676,
+ "area": 21190,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1866,
+ "bbox": [
+ 202,
+ 49,
+ 134,
+ 119
+ ],
+ "category_id": 15,
+ "id": 3677,
+ "area": 15946,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1866,
+ "bbox": [
+ 288,
+ 194,
+ 175,
+ 141
+ ],
+ "category_id": 15,
+ "id": 3678,
+ "area": 24675,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1866,
+ "bbox": [
+ 138,
+ 291,
+ 172,
+ 142
+ ],
+ "category_id": 15,
+ "id": 3679,
+ "area": 24424,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1867,
+ "bbox": [
+ 126,
+ 211,
+ 257,
+ 71
+ ],
+ "category_id": 19,
+ "id": 3680,
+ "area": 18247,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1868,
+ "bbox": [
+ 175,
+ 132,
+ 32,
+ 17
+ ],
+ "category_id": 4,
+ "id": 3681,
+ "area": 544,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1868,
+ "bbox": [
+ 62,
+ 213,
+ 36,
+ 22
+ ],
+ "category_id": 4,
+ "id": 3682,
+ "area": 792,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1868,
+ "bbox": [
+ 39,
+ 375,
+ 32,
+ 19
+ ],
+ "category_id": 4,
+ "id": 3683,
+ "area": 608,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1868,
+ "bbox": [
+ 411,
+ 171,
+ 33,
+ 19
+ ],
+ "category_id": 4,
+ "id": 3684,
+ "area": 627,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1868,
+ "bbox": [
+ 462,
+ 362,
+ 32,
+ 20
+ ],
+ "category_id": 4,
+ "id": 3685,
+ "area": 640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1869,
+ "bbox": [
+ 83,
+ 165,
+ 365,
+ 114
+ ],
+ "category_id": 16,
+ "id": 3686,
+ "area": 41610,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1870,
+ "bbox": [
+ 201,
+ 396,
+ 38,
+ 94
+ ],
+ "category_id": 15,
+ "id": 3687,
+ "area": 3572,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1871,
+ "bbox": [
+ 140,
+ 248,
+ 49,
+ 167
+ ],
+ "category_id": 16,
+ "id": 3688,
+ "area": 8183,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1872,
+ "bbox": [
+ 88,
+ 135,
+ 275,
+ 320
+ ],
+ "category_id": 16,
+ "id": 3689,
+ "area": 88000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1873,
+ "bbox": [
+ 46,
+ 73,
+ 416,
+ 332
+ ],
+ "category_id": 10,
+ "id": 3690,
+ "area": 138112,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1874,
+ "bbox": [
+ 166,
+ 216,
+ 206,
+ 153
+ ],
+ "category_id": 16,
+ "id": 3691,
+ "area": 31518,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1875,
+ "bbox": [
+ 47,
+ 108,
+ 36,
+ 35
+ ],
+ "category_id": 4,
+ "id": 3692,
+ "area": 1260,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1875,
+ "bbox": [
+ 294,
+ 93,
+ 42,
+ 37
+ ],
+ "category_id": 4,
+ "id": 3693,
+ "area": 1554,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1875,
+ "bbox": [
+ 332,
+ 417,
+ 44,
+ 41
+ ],
+ "category_id": 4,
+ "id": 3694,
+ "area": 1804,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1876,
+ "bbox": [
+ 0,
+ 166,
+ 221,
+ 261
+ ],
+ "category_id": 15,
+ "id": 3695,
+ "area": 57681,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1876,
+ "bbox": [
+ 272,
+ 148,
+ 213,
+ 266
+ ],
+ "category_id": 15,
+ "id": 3696,
+ "area": 56658,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1877,
+ "bbox": [
+ 40,
+ 64,
+ 57,
+ 46
+ ],
+ "category_id": 4,
+ "id": 3697,
+ "area": 2622,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1877,
+ "bbox": [
+ 337,
+ 145,
+ 47,
+ 16
+ ],
+ "category_id": 4,
+ "id": 3698,
+ "area": 752,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1877,
+ "bbox": [
+ 87,
+ 418,
+ 47,
+ 57
+ ],
+ "category_id": 4,
+ "id": 3699,
+ "area": 2679,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1878,
+ "bbox": [
+ 165,
+ 119,
+ 27,
+ 29
+ ],
+ "category_id": 4,
+ "id": 3700,
+ "area": 783,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1878,
+ "bbox": [
+ 474,
+ 100,
+ 25,
+ 26
+ ],
+ "category_id": 4,
+ "id": 3701,
+ "area": 650,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1878,
+ "bbox": [
+ 268,
+ 213,
+ 20,
+ 30
+ ],
+ "category_id": 4,
+ "id": 3702,
+ "area": 600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1878,
+ "bbox": [
+ 21,
+ 398,
+ 22,
+ 23
+ ],
+ "category_id": 4,
+ "id": 3703,
+ "area": 506,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1878,
+ "bbox": [
+ 433,
+ 254,
+ 19,
+ 30
+ ],
+ "category_id": 4,
+ "id": 3704,
+ "area": 570,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1879,
+ "bbox": [
+ 208,
+ 136,
+ 84,
+ 208
+ ],
+ "category_id": 19,
+ "id": 3705,
+ "area": 17472,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1880,
+ "bbox": [
+ 71,
+ 170,
+ 149,
+ 109
+ ],
+ "category_id": 15,
+ "id": 3706,
+ "area": 16241,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1880,
+ "bbox": [
+ 236,
+ 167,
+ 130,
+ 109
+ ],
+ "category_id": 15,
+ "id": 3707,
+ "area": 14170,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1881,
+ "bbox": [
+ 231,
+ 114,
+ 173,
+ 340
+ ],
+ "category_id": 16,
+ "id": 3708,
+ "area": 58820,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1882,
+ "bbox": [
+ 221,
+ 175,
+ 122,
+ 59
+ ],
+ "category_id": 19,
+ "id": 3709,
+ "area": 7198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1883,
+ "bbox": [
+ 246,
+ 269,
+ 68,
+ 95
+ ],
+ "category_id": 15,
+ "id": 3710,
+ "area": 6460,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1883,
+ "bbox": [
+ 42,
+ 268,
+ 31,
+ 39
+ ],
+ "category_id": 15,
+ "id": 3711,
+ "area": 1209,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1883,
+ "bbox": [
+ 86,
+ 272,
+ 31,
+ 40
+ ],
+ "category_id": 15,
+ "id": 3712,
+ "area": 1240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1883,
+ "bbox": [
+ 149,
+ 392,
+ 11,
+ 70
+ ],
+ "category_id": 15,
+ "id": 3713,
+ "area": 770,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1883,
+ "bbox": [
+ 218,
+ 419,
+ 9,
+ 51
+ ],
+ "category_id": 15,
+ "id": 3714,
+ "area": 459,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1884,
+ "bbox": [
+ 44,
+ 51,
+ 26,
+ 32
+ ],
+ "category_id": 4,
+ "id": 3715,
+ "area": 832,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1884,
+ "bbox": [
+ 10,
+ 263,
+ 30,
+ 27
+ ],
+ "category_id": 4,
+ "id": 3716,
+ "area": 810,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1884,
+ "bbox": [
+ 34,
+ 414,
+ 30,
+ 29
+ ],
+ "category_id": 4,
+ "id": 3717,
+ "area": 870,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1884,
+ "bbox": [
+ 469,
+ 147,
+ 30,
+ 28
+ ],
+ "category_id": 4,
+ "id": 3718,
+ "area": 840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1884,
+ "bbox": [
+ 466,
+ 332,
+ 24,
+ 30
+ ],
+ "category_id": 4,
+ "id": 3719,
+ "area": 720,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1885,
+ "bbox": [
+ 71,
+ 215,
+ 335,
+ 95
+ ],
+ "category_id": 19,
+ "id": 3720,
+ "area": 31825,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1886,
+ "bbox": [
+ 95,
+ 248,
+ 43,
+ 37
+ ],
+ "category_id": 4,
+ "id": 3721,
+ "area": 1591,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1886,
+ "bbox": [
+ 326,
+ 160,
+ 37,
+ 37
+ ],
+ "category_id": 4,
+ "id": 3722,
+ "area": 1369,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1886,
+ "bbox": [
+ 426,
+ 344,
+ 45,
+ 35
+ ],
+ "category_id": 4,
+ "id": 3723,
+ "area": 1575,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1887,
+ "bbox": [
+ 165,
+ 177,
+ 271,
+ 197
+ ],
+ "category_id": 10,
+ "id": 3724,
+ "area": 53387,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1888,
+ "bbox": [
+ 236,
+ 99,
+ 135,
+ 166
+ ],
+ "category_id": 15,
+ "id": 3725,
+ "area": 22410,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1888,
+ "bbox": [
+ 92,
+ 242,
+ 137,
+ 165
+ ],
+ "category_id": 15,
+ "id": 3726,
+ "area": 22605,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1889,
+ "bbox": [
+ 359,
+ 119,
+ 127,
+ 141
+ ],
+ "category_id": 10,
+ "id": 3727,
+ "area": 17907,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1890,
+ "bbox": [
+ 49,
+ 26,
+ 127,
+ 121
+ ],
+ "category_id": 10,
+ "id": 3728,
+ "area": 15367,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1891,
+ "bbox": [
+ 163,
+ 124,
+ 194,
+ 312
+ ],
+ "category_id": 16,
+ "id": 3729,
+ "area": 60528,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1892,
+ "bbox": [
+ 239,
+ 118,
+ 66,
+ 214
+ ],
+ "category_id": 10,
+ "id": 3730,
+ "area": 14124,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1893,
+ "bbox": [
+ 113,
+ 321,
+ 47,
+ 14
+ ],
+ "category_id": 15,
+ "id": 3731,
+ "area": 658,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1894,
+ "bbox": [
+ 174,
+ 145,
+ 80,
+ 212
+ ],
+ "category_id": 19,
+ "id": 3732,
+ "area": 16960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1895,
+ "bbox": [
+ 169,
+ 230,
+ 184,
+ 70
+ ],
+ "category_id": 19,
+ "id": 3733,
+ "area": 12880,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1896,
+ "bbox": [
+ 24,
+ 76,
+ 9,
+ 13
+ ],
+ "category_id": 4,
+ "id": 3734,
+ "area": 117,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1896,
+ "bbox": [
+ 331,
+ 15,
+ 9,
+ 15
+ ],
+ "category_id": 4,
+ "id": 3735,
+ "area": 135,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1896,
+ "bbox": [
+ 439,
+ 108,
+ 9,
+ 14
+ ],
+ "category_id": 4,
+ "id": 3736,
+ "area": 126,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1896,
+ "bbox": [
+ 136,
+ 218,
+ 9,
+ 17
+ ],
+ "category_id": 4,
+ "id": 3737,
+ "area": 153,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1896,
+ "bbox": [
+ 297,
+ 368,
+ 7,
+ 17
+ ],
+ "category_id": 4,
+ "id": 3738,
+ "area": 119,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1896,
+ "bbox": [
+ 364,
+ 489,
+ 11,
+ 14
+ ],
+ "category_id": 4,
+ "id": 3739,
+ "area": 154,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1897,
+ "bbox": [
+ 154,
+ 44,
+ 126,
+ 124
+ ],
+ "category_id": 15,
+ "id": 3740,
+ "area": 15624,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1897,
+ "bbox": [
+ 311,
+ 419,
+ 73,
+ 69
+ ],
+ "category_id": 15,
+ "id": 3741,
+ "area": 5037,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1898,
+ "bbox": [
+ 158,
+ 221,
+ 258,
+ 123
+ ],
+ "category_id": 10,
+ "id": 3742,
+ "area": 31734,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1899,
+ "bbox": [
+ 312,
+ 295,
+ 168,
+ 112
+ ],
+ "category_id": 10,
+ "id": 3743,
+ "area": 18816,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1900,
+ "bbox": [
+ 275,
+ 313,
+ 117,
+ 165
+ ],
+ "category_id": 19,
+ "id": 3744,
+ "area": 19305,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1901,
+ "bbox": [
+ 192,
+ 167,
+ 147,
+ 168
+ ],
+ "category_id": 15,
+ "id": 3745,
+ "area": 24696,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1902,
+ "bbox": [
+ 56,
+ 71,
+ 238,
+ 59
+ ],
+ "category_id": 19,
+ "id": 3746,
+ "area": 14042,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1903,
+ "bbox": [
+ 62,
+ 293,
+ 153,
+ 194
+ ],
+ "category_id": 10,
+ "id": 3747,
+ "area": 29682,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1904,
+ "bbox": [
+ 193,
+ 97,
+ 128,
+ 311
+ ],
+ "category_id": 19,
+ "id": 3748,
+ "area": 39808,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1905,
+ "bbox": [
+ 73,
+ 284,
+ 73,
+ 185
+ ],
+ "category_id": 10,
+ "id": 3749,
+ "area": 13505,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1906,
+ "bbox": [
+ 307,
+ 79,
+ 61,
+ 359
+ ],
+ "category_id": 10,
+ "id": 3750,
+ "area": 21899,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1907,
+ "bbox": [
+ 296,
+ 72,
+ 111,
+ 338
+ ],
+ "category_id": 10,
+ "id": 3751,
+ "area": 37518,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1908,
+ "bbox": [
+ 159,
+ 135,
+ 197,
+ 314
+ ],
+ "category_id": 19,
+ "id": 3752,
+ "area": 61858,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1909,
+ "bbox": [
+ 109,
+ 168,
+ 234,
+ 132
+ ],
+ "category_id": 19,
+ "id": 3753,
+ "area": 30888,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1910,
+ "bbox": [
+ 313,
+ 192,
+ 171,
+ 108
+ ],
+ "category_id": 10,
+ "id": 3754,
+ "area": 18468,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1911,
+ "bbox": [
+ 235,
+ 84,
+ 159,
+ 282
+ ],
+ "category_id": 10,
+ "id": 3755,
+ "area": 44838,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1912,
+ "bbox": [
+ 46,
+ 106,
+ 428,
+ 283
+ ],
+ "category_id": 10,
+ "id": 3756,
+ "area": 121124,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1913,
+ "bbox": [
+ 11,
+ 67,
+ 300,
+ 395
+ ],
+ "category_id": 10,
+ "id": 3757,
+ "area": 118500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1914,
+ "bbox": [
+ 12,
+ 241,
+ 18,
+ 23
+ ],
+ "category_id": 4,
+ "id": 3758,
+ "area": 414,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1914,
+ "bbox": [
+ 175,
+ 154,
+ 11,
+ 22
+ ],
+ "category_id": 4,
+ "id": 3759,
+ "area": 242,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1914,
+ "bbox": [
+ 197,
+ 321,
+ 18,
+ 19
+ ],
+ "category_id": 4,
+ "id": 3760,
+ "area": 342,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1914,
+ "bbox": [
+ 243,
+ 432,
+ 20,
+ 18
+ ],
+ "category_id": 4,
+ "id": 3761,
+ "area": 360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1914,
+ "bbox": [
+ 435,
+ 284,
+ 20,
+ 20
+ ],
+ "category_id": 4,
+ "id": 3762,
+ "area": 400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1914,
+ "bbox": [
+ 469,
+ 151,
+ 19,
+ 18
+ ],
+ "category_id": 4,
+ "id": 3763,
+ "area": 342,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1915,
+ "bbox": [
+ 119,
+ 115,
+ 47,
+ 39
+ ],
+ "category_id": 16,
+ "id": 3764,
+ "area": 1833,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1916,
+ "bbox": [
+ 175,
+ 216,
+ 249,
+ 152
+ ],
+ "category_id": 19,
+ "id": 3765,
+ "area": 37848,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1917,
+ "bbox": [
+ 44,
+ 154,
+ 40,
+ 32
+ ],
+ "category_id": 4,
+ "id": 3766,
+ "area": 1280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1917,
+ "bbox": [
+ 406,
+ 401,
+ 36,
+ 31
+ ],
+ "category_id": 4,
+ "id": 3767,
+ "area": 1116,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1918,
+ "bbox": [
+ 190,
+ 152,
+ 81,
+ 155
+ ],
+ "category_id": 19,
+ "id": 3768,
+ "area": 12555,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1919,
+ "bbox": [
+ 236,
+ 270,
+ 89,
+ 61
+ ],
+ "category_id": 4,
+ "id": 3769,
+ "area": 5429,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1920,
+ "bbox": [
+ 74,
+ 341,
+ 278,
+ 114
+ ],
+ "category_id": 19,
+ "id": 3770,
+ "area": 31692,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1921,
+ "bbox": [
+ 132,
+ 382,
+ 77,
+ 50
+ ],
+ "category_id": 4,
+ "id": 3771,
+ "area": 3850,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1921,
+ "bbox": [
+ 375,
+ 115,
+ 72,
+ 52
+ ],
+ "category_id": 4,
+ "id": 3772,
+ "area": 3744,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1922,
+ "bbox": [
+ 128,
+ 121,
+ 246,
+ 250
+ ],
+ "category_id": 16,
+ "id": 3773,
+ "area": 61500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1923,
+ "bbox": [
+ 286,
+ 179,
+ 73,
+ 221
+ ],
+ "category_id": 10,
+ "id": 3774,
+ "area": 16133,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1924,
+ "bbox": [
+ 211,
+ 245,
+ 19,
+ 12
+ ],
+ "category_id": 4,
+ "id": 3775,
+ "area": 228,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1924,
+ "bbox": [
+ 309,
+ 170,
+ 18,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3776,
+ "area": 198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1924,
+ "bbox": [
+ 460,
+ 40,
+ 18,
+ 13
+ ],
+ "category_id": 4,
+ "id": 3777,
+ "area": 234,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1924,
+ "bbox": [
+ 330,
+ 272,
+ 18,
+ 12
+ ],
+ "category_id": 4,
+ "id": 3778,
+ "area": 216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1924,
+ "bbox": [
+ 430,
+ 135,
+ 15,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3779,
+ "area": 165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1924,
+ "bbox": [
+ 172,
+ 293,
+ 17,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3780,
+ "area": 153,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1924,
+ "bbox": [
+ 106,
+ 330,
+ 16,
+ 14
+ ],
+ "category_id": 4,
+ "id": 3781,
+ "area": 224,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1924,
+ "bbox": [
+ 40,
+ 405,
+ 18,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3782,
+ "area": 198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1925,
+ "bbox": [
+ 104,
+ 113,
+ 36,
+ 41
+ ],
+ "category_id": 4,
+ "id": 3783,
+ "area": 1476,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1925,
+ "bbox": [
+ 171,
+ 362,
+ 41,
+ 35
+ ],
+ "category_id": 4,
+ "id": 3784,
+ "area": 1435,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1925,
+ "bbox": [
+ 443,
+ 357,
+ 28,
+ 33
+ ],
+ "category_id": 4,
+ "id": 3785,
+ "area": 924,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1926,
+ "bbox": [
+ 67,
+ 132,
+ 70,
+ 40
+ ],
+ "category_id": 4,
+ "id": 3786,
+ "area": 2800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1926,
+ "bbox": [
+ 353,
+ 171,
+ 55,
+ 57
+ ],
+ "category_id": 4,
+ "id": 3787,
+ "area": 3135,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1926,
+ "bbox": [
+ 216,
+ 396,
+ 81,
+ 60
+ ],
+ "category_id": 4,
+ "id": 3788,
+ "area": 4860,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1927,
+ "bbox": [
+ 119,
+ 304,
+ 33,
+ 61
+ ],
+ "category_id": 4,
+ "id": 3789,
+ "area": 2013,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1927,
+ "bbox": [
+ 385,
+ 184,
+ 27,
+ 57
+ ],
+ "category_id": 4,
+ "id": 3790,
+ "area": 1539,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1928,
+ "bbox": [
+ 12,
+ 223,
+ 98,
+ 133
+ ],
+ "category_id": 10,
+ "id": 3791,
+ "area": 13034,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1929,
+ "bbox": [
+ 250,
+ 73,
+ 52,
+ 37
+ ],
+ "category_id": 4,
+ "id": 3792,
+ "area": 1924,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1929,
+ "bbox": [
+ 194,
+ 403,
+ 70,
+ 53
+ ],
+ "category_id": 4,
+ "id": 3793,
+ "area": 3710,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1930,
+ "bbox": [
+ 72,
+ 30,
+ 20,
+ 51
+ ],
+ "category_id": 16,
+ "id": 3794,
+ "area": 1020,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1931,
+ "bbox": [
+ 37,
+ 165,
+ 389,
+ 167
+ ],
+ "category_id": 10,
+ "id": 3795,
+ "area": 64963,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1932,
+ "bbox": [
+ 122,
+ 211,
+ 230,
+ 221
+ ],
+ "category_id": 16,
+ "id": 3796,
+ "area": 50830,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1933,
+ "bbox": [
+ 157,
+ 81,
+ 186,
+ 262
+ ],
+ "category_id": 19,
+ "id": 3797,
+ "area": 48732,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1934,
+ "bbox": [
+ 161,
+ 98,
+ 35,
+ 35
+ ],
+ "category_id": 4,
+ "id": 3798,
+ "area": 1225,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1934,
+ "bbox": [
+ 93,
+ 432,
+ 34,
+ 35
+ ],
+ "category_id": 4,
+ "id": 3799,
+ "area": 1190,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1934,
+ "bbox": [
+ 211,
+ 300,
+ 30,
+ 29
+ ],
+ "category_id": 4,
+ "id": 3800,
+ "area": 870,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1935,
+ "bbox": [
+ 154,
+ 55,
+ 250,
+ 411
+ ],
+ "category_id": 10,
+ "id": 3801,
+ "area": 102750,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1936,
+ "bbox": [
+ 105,
+ 48,
+ 175,
+ 158
+ ],
+ "category_id": 15,
+ "id": 3802,
+ "area": 27650,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1936,
+ "bbox": [
+ 110,
+ 270,
+ 172,
+ 155
+ ],
+ "category_id": 15,
+ "id": 3803,
+ "area": 26660,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1936,
+ "bbox": [
+ 320,
+ 119,
+ 81,
+ 32
+ ],
+ "category_id": 15,
+ "id": 3804,
+ "area": 2592,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1936,
+ "bbox": [
+ 325,
+ 252,
+ 80,
+ 30
+ ],
+ "category_id": 15,
+ "id": 3805,
+ "area": 2400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1937,
+ "bbox": [
+ 274,
+ 111,
+ 83,
+ 247
+ ],
+ "category_id": 10,
+ "id": 3806,
+ "area": 20501,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1938,
+ "bbox": [
+ 20,
+ 207,
+ 429,
+ 169
+ ],
+ "category_id": 10,
+ "id": 3807,
+ "area": 72501,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1939,
+ "bbox": [
+ 133,
+ 359,
+ 65,
+ 30
+ ],
+ "category_id": 19,
+ "id": 3808,
+ "area": 1950,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1940,
+ "bbox": [
+ 265,
+ 211,
+ 71,
+ 130
+ ],
+ "category_id": 10,
+ "id": 3809,
+ "area": 9230,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1941,
+ "bbox": [
+ 1,
+ 316,
+ 511,
+ 196
+ ],
+ "category_id": 10,
+ "id": 3810,
+ "area": 100156,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1942,
+ "bbox": [
+ 21,
+ 88,
+ 299,
+ 253
+ ],
+ "category_id": 16,
+ "id": 3811,
+ "area": 75647,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1943,
+ "bbox": [
+ 140,
+ 240,
+ 238,
+ 65
+ ],
+ "category_id": 19,
+ "id": 3812,
+ "area": 15470,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1944,
+ "bbox": [
+ 112,
+ 235,
+ 298,
+ 127
+ ],
+ "category_id": 10,
+ "id": 3813,
+ "area": 37846,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1945,
+ "bbox": [
+ 10,
+ 225,
+ 477,
+ 41
+ ],
+ "category_id": 10,
+ "id": 3814,
+ "area": 19557,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1946,
+ "bbox": [
+ 67,
+ 65,
+ 25,
+ 26
+ ],
+ "category_id": 4,
+ "id": 3815,
+ "area": 650,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1946,
+ "bbox": [
+ 70,
+ 184,
+ 20,
+ 20
+ ],
+ "category_id": 4,
+ "id": 3816,
+ "area": 400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1946,
+ "bbox": [
+ 64,
+ 294,
+ 26,
+ 24
+ ],
+ "category_id": 4,
+ "id": 3817,
+ "area": 624,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1946,
+ "bbox": [
+ 63,
+ 422,
+ 24,
+ 23
+ ],
+ "category_id": 4,
+ "id": 3818,
+ "area": 552,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1946,
+ "bbox": [
+ 382,
+ 8,
+ 22,
+ 24
+ ],
+ "category_id": 4,
+ "id": 3819,
+ "area": 528,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1946,
+ "bbox": [
+ 374,
+ 128,
+ 24,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3820,
+ "area": 504,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1946,
+ "bbox": [
+ 369,
+ 244,
+ 23,
+ 25
+ ],
+ "category_id": 4,
+ "id": 3821,
+ "area": 575,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1946,
+ "bbox": [
+ 373,
+ 364,
+ 22,
+ 23
+ ],
+ "category_id": 4,
+ "id": 3822,
+ "area": 506,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1946,
+ "bbox": [
+ 369,
+ 484,
+ 24,
+ 28
+ ],
+ "category_id": 4,
+ "id": 3823,
+ "area": 672,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1947,
+ "bbox": [
+ 65,
+ 245,
+ 36,
+ 65
+ ],
+ "category_id": 4,
+ "id": 3824,
+ "area": 2340,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1947,
+ "bbox": [
+ 439,
+ 216,
+ 25,
+ 46
+ ],
+ "category_id": 4,
+ "id": 3825,
+ "area": 1150,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1947,
+ "bbox": [
+ 352,
+ 421,
+ 26,
+ 43
+ ],
+ "category_id": 4,
+ "id": 3826,
+ "area": 1118,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1948,
+ "bbox": [
+ 216,
+ 206,
+ 15,
+ 15
+ ],
+ "category_id": 4,
+ "id": 3827,
+ "area": 225,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1948,
+ "bbox": [
+ 338,
+ 255,
+ 15,
+ 31
+ ],
+ "category_id": 4,
+ "id": 3828,
+ "area": 465,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1948,
+ "bbox": [
+ 237,
+ 392,
+ 17,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3829,
+ "area": 357,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1948,
+ "bbox": [
+ 40,
+ 371,
+ 11,
+ 14
+ ],
+ "category_id": 4,
+ "id": 3830,
+ "area": 154,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1948,
+ "bbox": [
+ 386,
+ 69,
+ 10,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3831,
+ "area": 90,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1949,
+ "bbox": [
+ 160,
+ 92,
+ 67,
+ 255
+ ],
+ "category_id": 19,
+ "id": 3832,
+ "area": 17085,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1950,
+ "bbox": [
+ 61,
+ 167,
+ 219,
+ 248
+ ],
+ "category_id": 16,
+ "id": 3833,
+ "area": 54312,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1951,
+ "bbox": [
+ 40,
+ 225,
+ 100,
+ 42
+ ],
+ "category_id": 19,
+ "id": 3834,
+ "area": 4200,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1951,
+ "bbox": [
+ 125,
+ 132,
+ 237,
+ 198
+ ],
+ "category_id": 19,
+ "id": 3835,
+ "area": 46926,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1952,
+ "bbox": [
+ 226,
+ 141,
+ 37,
+ 242
+ ],
+ "category_id": 16,
+ "id": 3836,
+ "area": 8954,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1953,
+ "bbox": [
+ 310,
+ 104,
+ 187,
+ 245
+ ],
+ "category_id": 19,
+ "id": 3837,
+ "area": 45815,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1954,
+ "bbox": [
+ 305,
+ 76,
+ 93,
+ 60
+ ],
+ "category_id": 19,
+ "id": 3838,
+ "area": 5580,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1955,
+ "bbox": [
+ 89,
+ 161,
+ 60,
+ 118
+ ],
+ "category_id": 15,
+ "id": 3839,
+ "area": 7080,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1955,
+ "bbox": [
+ 348,
+ 203,
+ 81,
+ 161
+ ],
+ "category_id": 15,
+ "id": 3840,
+ "area": 13041,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1956,
+ "bbox": [
+ 192,
+ 173,
+ 203,
+ 202
+ ],
+ "category_id": 16,
+ "id": 3841,
+ "area": 41006,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1957,
+ "bbox": [
+ 51,
+ 262,
+ 31,
+ 50
+ ],
+ "category_id": 4,
+ "id": 3842,
+ "area": 1550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1957,
+ "bbox": [
+ 446,
+ 111,
+ 29,
+ 39
+ ],
+ "category_id": 4,
+ "id": 3843,
+ "area": 1131,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1957,
+ "bbox": [
+ 353,
+ 376,
+ 34,
+ 47
+ ],
+ "category_id": 4,
+ "id": 3844,
+ "area": 1598,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1958,
+ "bbox": [
+ 166,
+ 220,
+ 15,
+ 25
+ ],
+ "category_id": 16,
+ "id": 3845,
+ "area": 375,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1958,
+ "bbox": [
+ 163,
+ 270,
+ 16,
+ 29
+ ],
+ "category_id": 16,
+ "id": 3846,
+ "area": 464,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1958,
+ "bbox": [
+ 73,
+ 399,
+ 26,
+ 24
+ ],
+ "category_id": 16,
+ "id": 3847,
+ "area": 624,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1959,
+ "bbox": [
+ 179,
+ 75,
+ 148,
+ 254
+ ],
+ "category_id": 16,
+ "id": 3848,
+ "area": 37592,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1960,
+ "bbox": [
+ 209,
+ 186,
+ 123,
+ 85
+ ],
+ "category_id": 4,
+ "id": 3849,
+ "area": 10455,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1961,
+ "bbox": [
+ 135,
+ 430,
+ 25,
+ 19
+ ],
+ "category_id": 16,
+ "id": 3850,
+ "area": 475,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1962,
+ "bbox": [
+ 122,
+ 65,
+ 32,
+ 18
+ ],
+ "category_id": 4,
+ "id": 3851,
+ "area": 576,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1962,
+ "bbox": [
+ 240,
+ 195,
+ 32,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3852,
+ "area": 672,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1962,
+ "bbox": [
+ 36,
+ 352,
+ 34,
+ 17
+ ],
+ "category_id": 4,
+ "id": 3853,
+ "area": 578,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1962,
+ "bbox": [
+ 369,
+ 357,
+ 32,
+ 18
+ ],
+ "category_id": 4,
+ "id": 3854,
+ "area": 576,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1962,
+ "bbox": [
+ 471,
+ 179,
+ 38,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3855,
+ "area": 798,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1963,
+ "bbox": [
+ 48,
+ 177,
+ 408,
+ 98
+ ],
+ "category_id": 10,
+ "id": 3856,
+ "area": 39984,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1964,
+ "bbox": [
+ 222,
+ 72,
+ 119,
+ 399
+ ],
+ "category_id": 10,
+ "id": 3857,
+ "area": 47481,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1965,
+ "bbox": [
+ 114,
+ 157,
+ 259,
+ 275
+ ],
+ "category_id": 16,
+ "id": 3858,
+ "area": 71225,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1966,
+ "bbox": [
+ 116,
+ 76,
+ 286,
+ 327
+ ],
+ "category_id": 10,
+ "id": 3859,
+ "area": 93522,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1967,
+ "bbox": [
+ 118,
+ 376,
+ 33,
+ 42
+ ],
+ "category_id": 4,
+ "id": 3860,
+ "area": 1386,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1967,
+ "bbox": [
+ 318,
+ 101,
+ 42,
+ 43
+ ],
+ "category_id": 4,
+ "id": 3861,
+ "area": 1806,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1967,
+ "bbox": [
+ 429,
+ 364,
+ 43,
+ 49
+ ],
+ "category_id": 4,
+ "id": 3862,
+ "area": 2107,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1968,
+ "bbox": [
+ 76,
+ 277,
+ 37,
+ 34
+ ],
+ "category_id": 4,
+ "id": 3863,
+ "area": 1258,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1968,
+ "bbox": [
+ 398,
+ 229,
+ 41,
+ 36
+ ],
+ "category_id": 4,
+ "id": 3864,
+ "area": 1476,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1969,
+ "bbox": [
+ 19,
+ 42,
+ 182,
+ 132
+ ],
+ "category_id": 19,
+ "id": 3865,
+ "area": 24024,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1970,
+ "bbox": [
+ 190,
+ 156,
+ 98,
+ 42
+ ],
+ "category_id": 4,
+ "id": 3866,
+ "area": 4116,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1970,
+ "bbox": [
+ 322,
+ 440,
+ 91,
+ 50
+ ],
+ "category_id": 4,
+ "id": 3867,
+ "area": 4550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1971,
+ "bbox": [
+ 323,
+ 67,
+ 27,
+ 391
+ ],
+ "category_id": 10,
+ "id": 3868,
+ "area": 10557,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1972,
+ "bbox": [
+ 19,
+ 185,
+ 141,
+ 117
+ ],
+ "category_id": 10,
+ "id": 3869,
+ "area": 16497,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1973,
+ "bbox": [
+ 80,
+ 37,
+ 155,
+ 126
+ ],
+ "category_id": 15,
+ "id": 3870,
+ "area": 19530,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1973,
+ "bbox": [
+ 249,
+ 113,
+ 163,
+ 140
+ ],
+ "category_id": 15,
+ "id": 3871,
+ "area": 22820,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1973,
+ "bbox": [
+ 64,
+ 211,
+ 154,
+ 126
+ ],
+ "category_id": 15,
+ "id": 3872,
+ "area": 19404,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1973,
+ "bbox": [
+ 227,
+ 295,
+ 162,
+ 133
+ ],
+ "category_id": 15,
+ "id": 3873,
+ "area": 21546,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1974,
+ "bbox": [
+ 74,
+ 85,
+ 98,
+ 61
+ ],
+ "category_id": 19,
+ "id": 3874,
+ "area": 5978,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1975,
+ "bbox": [
+ 158,
+ 327,
+ 302,
+ 110
+ ],
+ "category_id": 19,
+ "id": 3875,
+ "area": 33220,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1976,
+ "bbox": [
+ 15,
+ 28,
+ 18,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3876,
+ "area": 198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1976,
+ "bbox": [
+ 82,
+ 241,
+ 16,
+ 10
+ ],
+ "category_id": 4,
+ "id": 3877,
+ "area": 160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1976,
+ "bbox": [
+ 318,
+ 350,
+ 16,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3878,
+ "area": 176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1977,
+ "bbox": [
+ 85,
+ 125,
+ 72,
+ 33
+ ],
+ "category_id": 4,
+ "id": 3879,
+ "area": 2376,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1977,
+ "bbox": [
+ 396,
+ 331,
+ 72,
+ 29
+ ],
+ "category_id": 4,
+ "id": 3880,
+ "area": 2088,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1978,
+ "bbox": [
+ 197,
+ 141,
+ 110,
+ 129
+ ],
+ "category_id": 15,
+ "id": 3881,
+ "area": 14190,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1978,
+ "bbox": [
+ 217,
+ 322,
+ 110,
+ 128
+ ],
+ "category_id": 15,
+ "id": 3882,
+ "area": 14080,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1979,
+ "bbox": [
+ 161,
+ 196,
+ 105,
+ 138
+ ],
+ "category_id": 19,
+ "id": 3883,
+ "area": 14490,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1980,
+ "bbox": [
+ 247,
+ 215,
+ 157,
+ 124
+ ],
+ "category_id": 15,
+ "id": 3884,
+ "area": 19468,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1981,
+ "bbox": [
+ 157,
+ 129,
+ 262,
+ 130
+ ],
+ "category_id": 16,
+ "id": 3885,
+ "area": 34060,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1981,
+ "bbox": [
+ 24,
+ 339,
+ 66,
+ 46
+ ],
+ "category_id": 16,
+ "id": 3886,
+ "area": 3036,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1982,
+ "bbox": [
+ 124,
+ 154,
+ 257,
+ 300
+ ],
+ "category_id": 16,
+ "id": 3887,
+ "area": 77100,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1983,
+ "bbox": [
+ 196,
+ 46,
+ 47,
+ 50
+ ],
+ "category_id": 4,
+ "id": 3888,
+ "area": 2350,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1983,
+ "bbox": [
+ 146,
+ 420,
+ 53,
+ 57
+ ],
+ "category_id": 4,
+ "id": 3889,
+ "area": 3021,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1984,
+ "bbox": [
+ 139,
+ 85,
+ 45,
+ 113
+ ],
+ "category_id": 16,
+ "id": 3890,
+ "area": 5085,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1985,
+ "bbox": [
+ 199,
+ 192,
+ 69,
+ 221
+ ],
+ "category_id": 16,
+ "id": 3891,
+ "area": 15249,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1986,
+ "bbox": [
+ 220,
+ 273,
+ 95,
+ 143
+ ],
+ "category_id": 4,
+ "id": 3892,
+ "area": 13585,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1987,
+ "bbox": [
+ 97,
+ 369,
+ 47,
+ 57
+ ],
+ "category_id": 4,
+ "id": 3893,
+ "area": 2679,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1987,
+ "bbox": [
+ 370,
+ 129,
+ 51,
+ 50
+ ],
+ "category_id": 4,
+ "id": 3894,
+ "area": 2550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1988,
+ "bbox": [
+ 11,
+ 348,
+ 72,
+ 80
+ ],
+ "category_id": 4,
+ "id": 3895,
+ "area": 5760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1988,
+ "bbox": [
+ 413,
+ 233,
+ 83,
+ 72
+ ],
+ "category_id": 4,
+ "id": 3896,
+ "area": 5976,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1989,
+ "bbox": [
+ 12,
+ 343,
+ 96,
+ 22
+ ],
+ "category_id": 15,
+ "id": 3897,
+ "area": 2112,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1989,
+ "bbox": [
+ 120,
+ 383,
+ 103,
+ 20
+ ],
+ "category_id": 15,
+ "id": 3898,
+ "area": 2060,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1989,
+ "bbox": [
+ 288,
+ 243,
+ 113,
+ 93
+ ],
+ "category_id": 15,
+ "id": 3899,
+ "area": 10509,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1989,
+ "bbox": [
+ 193,
+ 116,
+ 82,
+ 70
+ ],
+ "category_id": 15,
+ "id": 3900,
+ "area": 5740,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1989,
+ "bbox": [
+ 287,
+ 135,
+ 80,
+ 67
+ ],
+ "category_id": 15,
+ "id": 3901,
+ "area": 5360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1989,
+ "bbox": [
+ 267,
+ 380,
+ 110,
+ 94
+ ],
+ "category_id": 15,
+ "id": 3902,
+ "area": 10340,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1990,
+ "bbox": [
+ 208,
+ 58,
+ 172,
+ 333
+ ],
+ "category_id": 10,
+ "id": 3903,
+ "area": 57276,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1991,
+ "bbox": [
+ 138,
+ 225,
+ 329,
+ 52
+ ],
+ "category_id": 10,
+ "id": 3904,
+ "area": 17108,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1992,
+ "bbox": [
+ 83,
+ 268,
+ 27,
+ 15
+ ],
+ "category_id": 4,
+ "id": 3905,
+ "area": 405,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1992,
+ "bbox": [
+ 312,
+ 244,
+ 17,
+ 23
+ ],
+ "category_id": 4,
+ "id": 3906,
+ "area": 391,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1992,
+ "bbox": [
+ 297,
+ 69,
+ 21,
+ 19
+ ],
+ "category_id": 4,
+ "id": 3907,
+ "area": 399,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1992,
+ "bbox": [
+ 407,
+ 360,
+ 17,
+ 23
+ ],
+ "category_id": 4,
+ "id": 3908,
+ "area": 391,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1993,
+ "bbox": [
+ 24,
+ 350,
+ 38,
+ 20
+ ],
+ "category_id": 4,
+ "id": 3909,
+ "area": 760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1993,
+ "bbox": [
+ 416,
+ 39,
+ 32,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3910,
+ "area": 672,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1993,
+ "bbox": [
+ 218,
+ 297,
+ 38,
+ 19
+ ],
+ "category_id": 4,
+ "id": 3911,
+ "area": 722,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1993,
+ "bbox": [
+ 444,
+ 167,
+ 35,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3912,
+ "area": 735,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1993,
+ "bbox": [
+ 419,
+ 293,
+ 37,
+ 22
+ ],
+ "category_id": 4,
+ "id": 3913,
+ "area": 814,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1993,
+ "bbox": [
+ 204,
+ 490,
+ 38,
+ 18
+ ],
+ "category_id": 4,
+ "id": 3914,
+ "area": 684,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1994,
+ "bbox": [
+ 250,
+ 64,
+ 86,
+ 385
+ ],
+ "category_id": 10,
+ "id": 3915,
+ "area": 33110,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1995,
+ "bbox": [
+ 165,
+ 215,
+ 168,
+ 143
+ ],
+ "category_id": 16,
+ "id": 3916,
+ "area": 24024,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1996,
+ "bbox": [
+ 272,
+ 88,
+ 150,
+ 342
+ ],
+ "category_id": 10,
+ "id": 3917,
+ "area": 51300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1996,
+ "bbox": [
+ 0,
+ 250,
+ 32,
+ 11
+ ],
+ "category_id": 16,
+ "id": 3918,
+ "area": 352,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1997,
+ "bbox": [
+ 21,
+ 40,
+ 155,
+ 176
+ ],
+ "category_id": 10,
+ "id": 3919,
+ "area": 27280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1998,
+ "bbox": [
+ 112,
+ 153,
+ 243,
+ 238
+ ],
+ "category_id": 16,
+ "id": 3920,
+ "area": 57834,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 1999,
+ "bbox": [
+ 137,
+ 188,
+ 167,
+ 188
+ ],
+ "category_id": 16,
+ "id": 3921,
+ "area": 31396,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2000,
+ "bbox": [
+ 202,
+ 194,
+ 176,
+ 117
+ ],
+ "category_id": 4,
+ "id": 3922,
+ "area": 20592,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2001,
+ "bbox": [
+ 26,
+ 226,
+ 453,
+ 55
+ ],
+ "category_id": 10,
+ "id": 3923,
+ "area": 24915,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2002,
+ "bbox": [
+ 349,
+ 130,
+ 40,
+ 39
+ ],
+ "category_id": 4,
+ "id": 3924,
+ "area": 1560,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2002,
+ "bbox": [
+ 55,
+ 255,
+ 48,
+ 42
+ ],
+ "category_id": 4,
+ "id": 3925,
+ "area": 2016,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2002,
+ "bbox": [
+ 428,
+ 401,
+ 38,
+ 37
+ ],
+ "category_id": 4,
+ "id": 3926,
+ "area": 1406,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2003,
+ "bbox": [
+ 92,
+ 60,
+ 34,
+ 186
+ ],
+ "category_id": 10,
+ "id": 3927,
+ "area": 6324,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2004,
+ "bbox": [
+ 105,
+ 234,
+ 79,
+ 85
+ ],
+ "category_id": 15,
+ "id": 3928,
+ "area": 6715,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2005,
+ "bbox": [
+ 329,
+ 94,
+ 42,
+ 339
+ ],
+ "category_id": 10,
+ "id": 3929,
+ "area": 14238,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2006,
+ "bbox": [
+ 227,
+ 44,
+ 21,
+ 39
+ ],
+ "category_id": 4,
+ "id": 3930,
+ "area": 819,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2006,
+ "bbox": [
+ 427,
+ 272,
+ 21,
+ 43
+ ],
+ "category_id": 4,
+ "id": 3931,
+ "area": 903,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2006,
+ "bbox": [
+ 94,
+ 471,
+ 26,
+ 37
+ ],
+ "category_id": 4,
+ "id": 3932,
+ "area": 962,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2007,
+ "bbox": [
+ 200,
+ 148,
+ 56,
+ 27
+ ],
+ "category_id": 4,
+ "id": 3933,
+ "area": 1512,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2007,
+ "bbox": [
+ 331,
+ 382,
+ 67,
+ 26
+ ],
+ "category_id": 4,
+ "id": 3934,
+ "area": 1742,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2008,
+ "bbox": [
+ 97,
+ 40,
+ 22,
+ 25
+ ],
+ "category_id": 4,
+ "id": 3935,
+ "area": 550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2008,
+ "bbox": [
+ 184,
+ 155,
+ 17,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3936,
+ "area": 357,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2008,
+ "bbox": [
+ 76,
+ 386,
+ 14,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3937,
+ "area": 294,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2008,
+ "bbox": [
+ 241,
+ 328,
+ 19,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3938,
+ "area": 399,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2008,
+ "bbox": [
+ 378,
+ 413,
+ 13,
+ 17
+ ],
+ "category_id": 4,
+ "id": 3939,
+ "area": 221,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2009,
+ "bbox": [
+ 33,
+ 63,
+ 456,
+ 332
+ ],
+ "category_id": 15,
+ "id": 3940,
+ "area": 151392,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2010,
+ "bbox": [
+ 195,
+ 174,
+ 172,
+ 169
+ ],
+ "category_id": 10,
+ "id": 3941,
+ "area": 29068,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2011,
+ "bbox": [
+ 193,
+ 145,
+ 79,
+ 192
+ ],
+ "category_id": 19,
+ "id": 3942,
+ "area": 15168,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2012,
+ "bbox": [
+ 122,
+ 142,
+ 53,
+ 48
+ ],
+ "category_id": 4,
+ "id": 3943,
+ "area": 2544,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2012,
+ "bbox": [
+ 404,
+ 148,
+ 60,
+ 51
+ ],
+ "category_id": 4,
+ "id": 3944,
+ "area": 3060,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2013,
+ "bbox": [
+ 379,
+ 259,
+ 102,
+ 172
+ ],
+ "category_id": 10,
+ "id": 3945,
+ "area": 17544,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2014,
+ "bbox": [
+ 34,
+ 33,
+ 20,
+ 227
+ ],
+ "category_id": 10,
+ "id": 3946,
+ "area": 4540,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2015,
+ "bbox": [
+ 232,
+ 158,
+ 102,
+ 208
+ ],
+ "category_id": 19,
+ "id": 3947,
+ "area": 21216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2016,
+ "bbox": [
+ 332,
+ 191,
+ 98,
+ 110
+ ],
+ "category_id": 15,
+ "id": 3948,
+ "area": 10780,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2016,
+ "bbox": [
+ 166,
+ 389,
+ 63,
+ 74
+ ],
+ "category_id": 15,
+ "id": 3949,
+ "area": 4662,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2016,
+ "bbox": [
+ 423,
+ 352,
+ 63,
+ 70
+ ],
+ "category_id": 15,
+ "id": 3950,
+ "area": 4410,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2017,
+ "bbox": [
+ 61,
+ 145,
+ 410,
+ 148
+ ],
+ "category_id": 10,
+ "id": 3951,
+ "area": 60680,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2017,
+ "bbox": [
+ 105,
+ 337,
+ 67,
+ 37
+ ],
+ "category_id": 19,
+ "id": 3952,
+ "area": 2479,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2018,
+ "bbox": [
+ 92,
+ 211,
+ 389,
+ 64
+ ],
+ "category_id": 10,
+ "id": 3953,
+ "area": 24896,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2019,
+ "bbox": [
+ 143,
+ 228,
+ 121,
+ 100
+ ],
+ "category_id": 19,
+ "id": 3954,
+ "area": 12100,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2020,
+ "bbox": [
+ 430,
+ 267,
+ 66,
+ 176
+ ],
+ "category_id": 10,
+ "id": 3955,
+ "area": 11616,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2021,
+ "bbox": [
+ 368,
+ 179,
+ 10,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3956,
+ "area": 90,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2021,
+ "bbox": [
+ 474,
+ 225,
+ 8,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3957,
+ "area": 72,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2021,
+ "bbox": [
+ 465,
+ 344,
+ 13,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3958,
+ "area": 91,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2022,
+ "bbox": [
+ 24,
+ 355,
+ 9,
+ 16
+ ],
+ "category_id": 4,
+ "id": 3959,
+ "area": 144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2022,
+ "bbox": [
+ 62,
+ 314,
+ 7,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3960,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2022,
+ "bbox": [
+ 97,
+ 179,
+ 8,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3961,
+ "area": 88,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2022,
+ "bbox": [
+ 113,
+ 125,
+ 7,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3962,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2022,
+ "bbox": [
+ 133,
+ 17,
+ 8,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3963,
+ "area": 72,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2022,
+ "bbox": [
+ 293,
+ 96,
+ 7,
+ 13
+ ],
+ "category_id": 4,
+ "id": 3964,
+ "area": 91,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2022,
+ "bbox": [
+ 307,
+ 155,
+ 8,
+ 8
+ ],
+ "category_id": 4,
+ "id": 3965,
+ "area": 64,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2022,
+ "bbox": [
+ 403,
+ 106,
+ 7,
+ 7
+ ],
+ "category_id": 4,
+ "id": 3966,
+ "area": 49,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2023,
+ "bbox": [
+ 8,
+ 134,
+ 486,
+ 74
+ ],
+ "category_id": 19,
+ "id": 3967,
+ "area": 35964,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2024,
+ "bbox": [
+ 133,
+ 197,
+ 117,
+ 158
+ ],
+ "category_id": 15,
+ "id": 3968,
+ "area": 18486,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2024,
+ "bbox": [
+ 301,
+ 145,
+ 115,
+ 157
+ ],
+ "category_id": 15,
+ "id": 3969,
+ "area": 18055,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2025,
+ "bbox": [
+ 33,
+ 83,
+ 111,
+ 43
+ ],
+ "category_id": 4,
+ "id": 3970,
+ "area": 4773,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2025,
+ "bbox": [
+ 371,
+ 212,
+ 88,
+ 41
+ ],
+ "category_id": 4,
+ "id": 3971,
+ "area": 3608,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2026,
+ "bbox": [
+ 439,
+ 31,
+ 37,
+ 22
+ ],
+ "category_id": 4,
+ "id": 3972,
+ "area": 814,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2026,
+ "bbox": [
+ 245,
+ 126,
+ 35,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3973,
+ "area": 735,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2026,
+ "bbox": [
+ 104,
+ 259,
+ 31,
+ 18
+ ],
+ "category_id": 4,
+ "id": 3974,
+ "area": 558,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2026,
+ "bbox": [
+ 256,
+ 268,
+ 38,
+ 20
+ ],
+ "category_id": 4,
+ "id": 3975,
+ "area": 760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2026,
+ "bbox": [
+ 469,
+ 158,
+ 39,
+ 19
+ ],
+ "category_id": 4,
+ "id": 3976,
+ "area": 741,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2026,
+ "bbox": [
+ 9,
+ 337,
+ 37,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3977,
+ "area": 777,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2026,
+ "bbox": [
+ 193,
+ 358,
+ 36,
+ 18
+ ],
+ "category_id": 4,
+ "id": 3978,
+ "area": 648,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2026,
+ "bbox": [
+ 13,
+ 485,
+ 31,
+ 20
+ ],
+ "category_id": 4,
+ "id": 3979,
+ "area": 620,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2027,
+ "bbox": [
+ 63,
+ 85,
+ 16,
+ 12
+ ],
+ "category_id": 4,
+ "id": 3980,
+ "area": 192,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2027,
+ "bbox": [
+ 387,
+ 146,
+ 17,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3981,
+ "area": 153,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2027,
+ "bbox": [
+ 295,
+ 192,
+ 21,
+ 9
+ ],
+ "category_id": 4,
+ "id": 3982,
+ "area": 189,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2027,
+ "bbox": [
+ 145,
+ 238,
+ 23,
+ 12
+ ],
+ "category_id": 4,
+ "id": 3983,
+ "area": 276,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2027,
+ "bbox": [
+ 359,
+ 368,
+ 24,
+ 17
+ ],
+ "category_id": 4,
+ "id": 3984,
+ "area": 408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2027,
+ "bbox": [
+ 472,
+ 375,
+ 27,
+ 18
+ ],
+ "category_id": 4,
+ "id": 3985,
+ "area": 486,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2027,
+ "bbox": [
+ 333,
+ 444,
+ 22,
+ 21
+ ],
+ "category_id": 4,
+ "id": 3986,
+ "area": 462,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2027,
+ "bbox": [
+ 154,
+ 421,
+ 22,
+ 12
+ ],
+ "category_id": 4,
+ "id": 3987,
+ "area": 264,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2027,
+ "bbox": [
+ 56,
+ 465,
+ 19,
+ 11
+ ],
+ "category_id": 4,
+ "id": 3988,
+ "area": 209,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2028,
+ "bbox": [
+ 263,
+ 198,
+ 147,
+ 157
+ ],
+ "category_id": 15,
+ "id": 3989,
+ "area": 23079,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2029,
+ "bbox": [
+ 58,
+ 176,
+ 155,
+ 136
+ ],
+ "category_id": 15,
+ "id": 3990,
+ "area": 21080,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2029,
+ "bbox": [
+ 282,
+ 186,
+ 155,
+ 135
+ ],
+ "category_id": 15,
+ "id": 3991,
+ "area": 20925,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2029,
+ "bbox": [
+ 174,
+ 0,
+ 92,
+ 44
+ ],
+ "category_id": 15,
+ "id": 3992,
+ "area": 4048,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2030,
+ "bbox": [
+ 27,
+ 294,
+ 74,
+ 39
+ ],
+ "category_id": 4,
+ "id": 3993,
+ "area": 2886,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2030,
+ "bbox": [
+ 406,
+ 290,
+ 66,
+ 44
+ ],
+ "category_id": 4,
+ "id": 3994,
+ "area": 2904,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2031,
+ "bbox": [
+ 97,
+ 122,
+ 168,
+ 262
+ ],
+ "category_id": 16,
+ "id": 3995,
+ "area": 44016,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2032,
+ "bbox": [
+ 241,
+ 144,
+ 183,
+ 189
+ ],
+ "category_id": 16,
+ "id": 3996,
+ "area": 34587,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2033,
+ "bbox": [
+ 109,
+ 407,
+ 31,
+ 49
+ ],
+ "category_id": 4,
+ "id": 3997,
+ "area": 1519,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2033,
+ "bbox": [
+ 248,
+ 293,
+ 26,
+ 49
+ ],
+ "category_id": 4,
+ "id": 3998,
+ "area": 1274,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2033,
+ "bbox": [
+ 376,
+ 106,
+ 26,
+ 46
+ ],
+ "category_id": 4,
+ "id": 3999,
+ "area": 1196,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2034,
+ "bbox": [
+ 286,
+ 279,
+ 90,
+ 90
+ ],
+ "category_id": 4,
+ "id": 4000,
+ "area": 8100,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2035,
+ "bbox": [
+ 153,
+ 339,
+ 87,
+ 51
+ ],
+ "category_id": 16,
+ "id": 4001,
+ "area": 4437,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2036,
+ "bbox": [
+ 243,
+ 132,
+ 45,
+ 90
+ ],
+ "category_id": 4,
+ "id": 4002,
+ "area": 4050,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2037,
+ "bbox": [
+ 272,
+ 58,
+ 7,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4003,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2037,
+ "bbox": [
+ 222,
+ 145,
+ 7,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4004,
+ "area": 63,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2037,
+ "bbox": [
+ 223,
+ 200,
+ 6,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4005,
+ "area": 48,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2037,
+ "bbox": [
+ 197,
+ 248,
+ 5,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4006,
+ "area": 45,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2037,
+ "bbox": [
+ 211,
+ 339,
+ 8,
+ 12
+ ],
+ "category_id": 4,
+ "id": 4007,
+ "area": 96,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2037,
+ "bbox": [
+ 236,
+ 382,
+ 4,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4008,
+ "area": 36,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2037,
+ "bbox": [
+ 250,
+ 424,
+ 3,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4009,
+ "area": 21,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2037,
+ "bbox": [
+ 146,
+ 454,
+ 5,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4010,
+ "area": 35,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2038,
+ "bbox": [
+ 62,
+ 159,
+ 35,
+ 43
+ ],
+ "category_id": 4,
+ "id": 4011,
+ "area": 1505,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2038,
+ "bbox": [
+ 407,
+ 329,
+ 45,
+ 53
+ ],
+ "category_id": 4,
+ "id": 4012,
+ "area": 2385,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2039,
+ "bbox": [
+ 34,
+ 184,
+ 462,
+ 80
+ ],
+ "category_id": 19,
+ "id": 4013,
+ "area": 36960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2040,
+ "bbox": [
+ 366,
+ 417,
+ 123,
+ 28
+ ],
+ "category_id": 10,
+ "id": 4014,
+ "area": 3444,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2041,
+ "bbox": [
+ 108,
+ 72,
+ 60,
+ 33
+ ],
+ "category_id": 4,
+ "id": 4015,
+ "area": 1980,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2041,
+ "bbox": [
+ 272,
+ 410,
+ 55,
+ 32
+ ],
+ "category_id": 4,
+ "id": 4016,
+ "area": 1760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2042,
+ "bbox": [
+ 235,
+ 224,
+ 82,
+ 117
+ ],
+ "category_id": 19,
+ "id": 4017,
+ "area": 9594,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2043,
+ "bbox": [
+ 61,
+ 169,
+ 56,
+ 79
+ ],
+ "category_id": 15,
+ "id": 4018,
+ "area": 4424,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2043,
+ "bbox": [
+ 262,
+ 270,
+ 54,
+ 76
+ ],
+ "category_id": 15,
+ "id": 4019,
+ "area": 4104,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2044,
+ "bbox": [
+ 171,
+ 131,
+ 72,
+ 89
+ ],
+ "category_id": 15,
+ "id": 4020,
+ "area": 6408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2044,
+ "bbox": [
+ 166,
+ 319,
+ 104,
+ 123
+ ],
+ "category_id": 15,
+ "id": 4021,
+ "area": 12792,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2045,
+ "bbox": [
+ 54,
+ 261,
+ 55,
+ 53
+ ],
+ "category_id": 4,
+ "id": 4022,
+ "area": 2915,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2045,
+ "bbox": [
+ 417,
+ 200,
+ 57,
+ 51
+ ],
+ "category_id": 4,
+ "id": 4023,
+ "area": 2907,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2046,
+ "bbox": [
+ 240,
+ 182,
+ 44,
+ 152
+ ],
+ "category_id": 19,
+ "id": 4024,
+ "area": 6688,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2047,
+ "bbox": [
+ 220,
+ 141,
+ 194,
+ 223
+ ],
+ "category_id": 16,
+ "id": 4025,
+ "area": 43262,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2048,
+ "bbox": [
+ 28,
+ 255,
+ 28,
+ 18
+ ],
+ "category_id": 4,
+ "id": 4026,
+ "area": 504,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2048,
+ "bbox": [
+ 137,
+ 270,
+ 31,
+ 15
+ ],
+ "category_id": 4,
+ "id": 4027,
+ "area": 465,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2048,
+ "bbox": [
+ 267,
+ 291,
+ 18,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4028,
+ "area": 144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2048,
+ "bbox": [
+ 355,
+ 312,
+ 21,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4029,
+ "area": 210,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2049,
+ "bbox": [
+ 39,
+ 209,
+ 342,
+ 84
+ ],
+ "category_id": 10,
+ "id": 4030,
+ "area": 28728,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2050,
+ "bbox": [
+ 197,
+ 70,
+ 11,
+ 23
+ ],
+ "category_id": 4,
+ "id": 4031,
+ "area": 253,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2050,
+ "bbox": [
+ 49,
+ 272,
+ 17,
+ 24
+ ],
+ "category_id": 4,
+ "id": 4032,
+ "area": 408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2050,
+ "bbox": [
+ 250,
+ 229,
+ 13,
+ 27
+ ],
+ "category_id": 4,
+ "id": 4033,
+ "area": 351,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2050,
+ "bbox": [
+ 264,
+ 428,
+ 11,
+ 24
+ ],
+ "category_id": 4,
+ "id": 4034,
+ "area": 264,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2050,
+ "bbox": [
+ 456,
+ 407,
+ 16,
+ 28
+ ],
+ "category_id": 4,
+ "id": 4035,
+ "area": 448,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2051,
+ "bbox": [
+ 44,
+ 208,
+ 68,
+ 35
+ ],
+ "category_id": 4,
+ "id": 4036,
+ "area": 2380,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2051,
+ "bbox": [
+ 430,
+ 200,
+ 67,
+ 30
+ ],
+ "category_id": 4,
+ "id": 4037,
+ "area": 2010,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2052,
+ "bbox": [
+ 202,
+ 253,
+ 144,
+ 128
+ ],
+ "category_id": 16,
+ "id": 4038,
+ "area": 18432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2053,
+ "bbox": [
+ 410,
+ 382,
+ 40,
+ 92
+ ],
+ "category_id": 10,
+ "id": 4039,
+ "area": 3680,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2054,
+ "bbox": [
+ 188,
+ 121,
+ 226,
+ 324
+ ],
+ "category_id": 10,
+ "id": 4040,
+ "area": 73224,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2055,
+ "bbox": [
+ 239,
+ 268,
+ 151,
+ 117
+ ],
+ "category_id": 4,
+ "id": 4041,
+ "area": 17667,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2056,
+ "bbox": [
+ 64,
+ 339,
+ 11,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4042,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2056,
+ "bbox": [
+ 85,
+ 415,
+ 8,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4043,
+ "area": 64,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2056,
+ "bbox": [
+ 203,
+ 355,
+ 5,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4044,
+ "area": 45,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2056,
+ "bbox": [
+ 271,
+ 371,
+ 9,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4045,
+ "area": 63,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2056,
+ "bbox": [
+ 318,
+ 461,
+ 9,
+ 6
+ ],
+ "category_id": 4,
+ "id": 4046,
+ "area": 54,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2056,
+ "bbox": [
+ 396,
+ 503,
+ 11,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4047,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2056,
+ "bbox": [
+ 488,
+ 433,
+ 9,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4048,
+ "area": 63,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2056,
+ "bbox": [
+ 322,
+ 284,
+ 12,
+ 6
+ ],
+ "category_id": 4,
+ "id": 4049,
+ "area": 72,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2056,
+ "bbox": [
+ 112,
+ 241,
+ 6,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4050,
+ "area": 54,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2056,
+ "bbox": [
+ 326,
+ 158,
+ 9,
+ 6
+ ],
+ "category_id": 4,
+ "id": 4051,
+ "area": 54,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2056,
+ "bbox": [
+ 442,
+ 187,
+ 9,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4052,
+ "area": 63,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2056,
+ "bbox": [
+ 452,
+ 90,
+ 9,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4053,
+ "area": 63,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2057,
+ "bbox": [
+ 197,
+ 172,
+ 104,
+ 119
+ ],
+ "category_id": 19,
+ "id": 4054,
+ "area": 12376,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2058,
+ "bbox": [
+ 53,
+ 194,
+ 34,
+ 29
+ ],
+ "category_id": 4,
+ "id": 4055,
+ "area": 986,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2058,
+ "bbox": [
+ 336,
+ 376,
+ 24,
+ 23
+ ],
+ "category_id": 4,
+ "id": 4056,
+ "area": 552,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2058,
+ "bbox": [
+ 321,
+ 69,
+ 26,
+ 21
+ ],
+ "category_id": 4,
+ "id": 4057,
+ "area": 546,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2059,
+ "bbox": [
+ 257,
+ 186,
+ 31,
+ 38
+ ],
+ "category_id": 15,
+ "id": 4058,
+ "area": 1178,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2060,
+ "bbox": [
+ 156,
+ 172,
+ 100,
+ 103
+ ],
+ "category_id": 15,
+ "id": 4059,
+ "area": 10300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2060,
+ "bbox": [
+ 301,
+ 314,
+ 95,
+ 107
+ ],
+ "category_id": 15,
+ "id": 4060,
+ "area": 10165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2061,
+ "bbox": [
+ 124,
+ 181,
+ 235,
+ 195
+ ],
+ "category_id": 15,
+ "id": 4061,
+ "area": 45825,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2062,
+ "bbox": [
+ 98,
+ 204,
+ 315,
+ 95
+ ],
+ "category_id": 19,
+ "id": 4062,
+ "area": 29925,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2063,
+ "bbox": [
+ 97,
+ 236,
+ 339,
+ 43
+ ],
+ "category_id": 19,
+ "id": 4063,
+ "area": 14577,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2064,
+ "bbox": [
+ 216,
+ 81,
+ 46,
+ 337
+ ],
+ "category_id": 19,
+ "id": 4064,
+ "area": 15502,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2065,
+ "bbox": [
+ 256,
+ 203,
+ 97,
+ 56
+ ],
+ "category_id": 4,
+ "id": 4065,
+ "area": 5432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2066,
+ "bbox": [
+ 227,
+ 131,
+ 194,
+ 281
+ ],
+ "category_id": 15,
+ "id": 4066,
+ "area": 54514,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2067,
+ "bbox": [
+ 122,
+ 258,
+ 340,
+ 71
+ ],
+ "category_id": 19,
+ "id": 4067,
+ "area": 24140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2068,
+ "bbox": [
+ 28,
+ 387,
+ 455,
+ 29
+ ],
+ "category_id": 10,
+ "id": 4068,
+ "area": 13195,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2069,
+ "bbox": [
+ 129,
+ 211,
+ 81,
+ 64
+ ],
+ "category_id": 15,
+ "id": 4069,
+ "area": 5184,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2069,
+ "bbox": [
+ 229,
+ 93,
+ 55,
+ 45
+ ],
+ "category_id": 15,
+ "id": 4070,
+ "area": 2475,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2069,
+ "bbox": [
+ 181,
+ 272,
+ 81,
+ 66
+ ],
+ "category_id": 15,
+ "id": 4071,
+ "area": 5346,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2069,
+ "bbox": [
+ 254,
+ 339,
+ 92,
+ 76
+ ],
+ "category_id": 15,
+ "id": 4072,
+ "area": 6992,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2070,
+ "bbox": [
+ 225,
+ 112,
+ 87,
+ 306
+ ],
+ "category_id": 16,
+ "id": 4073,
+ "area": 26622,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2071,
+ "bbox": [
+ 174,
+ 112,
+ 234,
+ 268
+ ],
+ "category_id": 16,
+ "id": 4074,
+ "area": 62712,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2072,
+ "bbox": [
+ 48,
+ 204,
+ 364,
+ 51
+ ],
+ "category_id": 19,
+ "id": 4075,
+ "area": 18564,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2073,
+ "bbox": [
+ 221,
+ 202,
+ 161,
+ 167
+ ],
+ "category_id": 15,
+ "id": 4076,
+ "area": 26887,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2074,
+ "bbox": [
+ 146,
+ 88,
+ 5,
+ 13
+ ],
+ "category_id": 4,
+ "id": 4077,
+ "area": 65,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2074,
+ "bbox": [
+ 56,
+ 69,
+ 9,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4078,
+ "area": 126,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2074,
+ "bbox": [
+ 214,
+ 123,
+ 7,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4079,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2074,
+ "bbox": [
+ 274,
+ 65,
+ 8,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4080,
+ "area": 88,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2074,
+ "bbox": [
+ 280,
+ 161,
+ 7,
+ 13
+ ],
+ "category_id": 4,
+ "id": 4081,
+ "area": 91,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2074,
+ "bbox": [
+ 264,
+ 222,
+ 8,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4082,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2074,
+ "bbox": [
+ 9,
+ 336,
+ 8,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4083,
+ "area": 64,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2074,
+ "bbox": [
+ 82,
+ 318,
+ 5,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4084,
+ "area": 50,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2074,
+ "bbox": [
+ 317,
+ 247,
+ 7,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4085,
+ "area": 70,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2074,
+ "bbox": [
+ 364,
+ 207,
+ 7,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4086,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2074,
+ "bbox": [
+ 391,
+ 321,
+ 7,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4087,
+ "area": 63,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2074,
+ "bbox": [
+ 444,
+ 325,
+ 5,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4088,
+ "area": 40,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2074,
+ "bbox": [
+ 406,
+ 435,
+ 5,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4089,
+ "area": 35,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2074,
+ "bbox": [
+ 401,
+ 491,
+ 6,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4090,
+ "area": 48,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2074,
+ "bbox": [
+ 307,
+ 484,
+ 5,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4091,
+ "area": 55,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2074,
+ "bbox": [
+ 257,
+ 493,
+ 5,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4092,
+ "area": 40,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2075,
+ "bbox": [
+ 60,
+ 154,
+ 155,
+ 224
+ ],
+ "category_id": 15,
+ "id": 4093,
+ "area": 34720,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2075,
+ "bbox": [
+ 277,
+ 160,
+ 159,
+ 215
+ ],
+ "category_id": 15,
+ "id": 4094,
+ "area": 34185,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2076,
+ "bbox": [
+ 45,
+ 218,
+ 49,
+ 57
+ ],
+ "category_id": 19,
+ "id": 4095,
+ "area": 2793,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2077,
+ "bbox": [
+ 99,
+ 207,
+ 336,
+ 223
+ ],
+ "category_id": 10,
+ "id": 4096,
+ "area": 74928,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2078,
+ "bbox": [
+ 60,
+ 169,
+ 153,
+ 255
+ ],
+ "category_id": 16,
+ "id": 4097,
+ "area": 39015,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2079,
+ "bbox": [
+ 267,
+ 99,
+ 18,
+ 65
+ ],
+ "category_id": 4,
+ "id": 4098,
+ "area": 1170,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2079,
+ "bbox": [
+ 126,
+ 358,
+ 46,
+ 48
+ ],
+ "category_id": 4,
+ "id": 4099,
+ "area": 2208,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2080,
+ "bbox": [
+ 231,
+ 307,
+ 103,
+ 87
+ ],
+ "category_id": 16,
+ "id": 4100,
+ "area": 8961,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2081,
+ "bbox": [
+ 257,
+ 142,
+ 82,
+ 282
+ ],
+ "category_id": 16,
+ "id": 4101,
+ "area": 23124,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2082,
+ "bbox": [
+ 141,
+ 186,
+ 88,
+ 86
+ ],
+ "category_id": 15,
+ "id": 4102,
+ "area": 7568,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2082,
+ "bbox": [
+ 359,
+ 237,
+ 65,
+ 60
+ ],
+ "category_id": 15,
+ "id": 4103,
+ "area": 3900,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2083,
+ "bbox": [
+ 249,
+ 72,
+ 121,
+ 345
+ ],
+ "category_id": 16,
+ "id": 4104,
+ "area": 41745,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2084,
+ "bbox": [
+ 87,
+ 51,
+ 85,
+ 57
+ ],
+ "category_id": 4,
+ "id": 4105,
+ "area": 4845,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2084,
+ "bbox": [
+ 411,
+ 368,
+ 80,
+ 60
+ ],
+ "category_id": 4,
+ "id": 4106,
+ "area": 4800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2085,
+ "bbox": [
+ 139,
+ 135,
+ 79,
+ 72
+ ],
+ "category_id": 15,
+ "id": 4107,
+ "area": 5688,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2085,
+ "bbox": [
+ 130,
+ 215,
+ 85,
+ 78
+ ],
+ "category_id": 15,
+ "id": 4108,
+ "area": 6630,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2085,
+ "bbox": [
+ 256,
+ 217,
+ 78,
+ 64
+ ],
+ "category_id": 15,
+ "id": 4109,
+ "area": 4992,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2085,
+ "bbox": [
+ 238,
+ 140,
+ 91,
+ 86
+ ],
+ "category_id": 15,
+ "id": 4110,
+ "area": 7826,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2086,
+ "bbox": [
+ 55,
+ 420,
+ 160,
+ 34
+ ],
+ "category_id": 10,
+ "id": 4111,
+ "area": 5440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2087,
+ "bbox": [
+ 55,
+ 224,
+ 426,
+ 55
+ ],
+ "category_id": 10,
+ "id": 4112,
+ "area": 23430,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2088,
+ "bbox": [
+ 104,
+ 159,
+ 40,
+ 202
+ ],
+ "category_id": 19,
+ "id": 4113,
+ "area": 8080,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2089,
+ "bbox": [
+ 136,
+ 256,
+ 152,
+ 63
+ ],
+ "category_id": 4,
+ "id": 4114,
+ "area": 9576,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2090,
+ "bbox": [
+ 106,
+ 64,
+ 53,
+ 79
+ ],
+ "category_id": 15,
+ "id": 4115,
+ "area": 4187,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2090,
+ "bbox": [
+ 137,
+ 145,
+ 54,
+ 80
+ ],
+ "category_id": 15,
+ "id": 4116,
+ "area": 4320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2090,
+ "bbox": [
+ 195,
+ 223,
+ 84,
+ 117
+ ],
+ "category_id": 15,
+ "id": 4117,
+ "area": 9828,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2090,
+ "bbox": [
+ 240,
+ 339,
+ 83,
+ 122
+ ],
+ "category_id": 15,
+ "id": 4118,
+ "area": 10126,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2091,
+ "bbox": [
+ 212,
+ 113,
+ 24,
+ 27
+ ],
+ "category_id": 4,
+ "id": 4119,
+ "area": 648,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2091,
+ "bbox": [
+ 109,
+ 261,
+ 26,
+ 32
+ ],
+ "category_id": 4,
+ "id": 4120,
+ "area": 832,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2091,
+ "bbox": [
+ 250,
+ 360,
+ 30,
+ 33
+ ],
+ "category_id": 4,
+ "id": 4121,
+ "area": 990,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2091,
+ "bbox": [
+ 403,
+ 464,
+ 27,
+ 32
+ ],
+ "category_id": 4,
+ "id": 4122,
+ "area": 864,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2092,
+ "bbox": [
+ 301,
+ 146,
+ 66,
+ 81
+ ],
+ "category_id": 15,
+ "id": 4123,
+ "area": 5346,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2092,
+ "bbox": [
+ 124,
+ 284,
+ 111,
+ 131
+ ],
+ "category_id": 15,
+ "id": 4124,
+ "area": 14541,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2093,
+ "bbox": [
+ 166,
+ 166,
+ 171,
+ 178
+ ],
+ "category_id": 19,
+ "id": 4125,
+ "area": 30438,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2094,
+ "bbox": [
+ 325,
+ 221,
+ 133,
+ 108
+ ],
+ "category_id": 10,
+ "id": 4126,
+ "area": 14364,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2095,
+ "bbox": [
+ 437,
+ 471,
+ 16,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4127,
+ "area": 144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2095,
+ "bbox": [
+ 122,
+ 355,
+ 15,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4128,
+ "area": 135,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2095,
+ "bbox": [
+ 370,
+ 371,
+ 17,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4129,
+ "area": 119,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2095,
+ "bbox": [
+ 55,
+ 217,
+ 15,
+ 6
+ ],
+ "category_id": 4,
+ "id": 4130,
+ "area": 90,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2095,
+ "bbox": [
+ 183,
+ 109,
+ 14,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4131,
+ "area": 126,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2096,
+ "bbox": [
+ 296,
+ 12,
+ 14,
+ 13
+ ],
+ "category_id": 4,
+ "id": 4132,
+ "area": 182,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2096,
+ "bbox": [
+ 169,
+ 44,
+ 17,
+ 12
+ ],
+ "category_id": 4,
+ "id": 4133,
+ "area": 204,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2096,
+ "bbox": [
+ 293,
+ 119,
+ 19,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4134,
+ "area": 304,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2096,
+ "bbox": [
+ 417,
+ 144,
+ 16,
+ 12
+ ],
+ "category_id": 4,
+ "id": 4135,
+ "area": 192,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2096,
+ "bbox": [
+ 217,
+ 285,
+ 15,
+ 17
+ ],
+ "category_id": 4,
+ "id": 4136,
+ "area": 255,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2096,
+ "bbox": [
+ 91,
+ 235,
+ 19,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4137,
+ "area": 304,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2096,
+ "bbox": [
+ 66,
+ 299,
+ 19,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4138,
+ "area": 266,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2096,
+ "bbox": [
+ 3,
+ 402,
+ 18,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4139,
+ "area": 252,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2096,
+ "bbox": [
+ 273,
+ 397,
+ 17,
+ 13
+ ],
+ "category_id": 4,
+ "id": 4140,
+ "area": 221,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2096,
+ "bbox": [
+ 397,
+ 423,
+ 17,
+ 12
+ ],
+ "category_id": 4,
+ "id": 4141,
+ "area": 204,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2096,
+ "bbox": [
+ 242,
+ 478,
+ 21,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4142,
+ "area": 294,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2096,
+ "bbox": [
+ 359,
+ 487,
+ 23,
+ 13
+ ],
+ "category_id": 4,
+ "id": 4143,
+ "area": 299,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2097,
+ "bbox": [
+ 21,
+ 243,
+ 27,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4144,
+ "area": 378,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2097,
+ "bbox": [
+ 232,
+ 97,
+ 28,
+ 15
+ ],
+ "category_id": 4,
+ "id": 4145,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2097,
+ "bbox": [
+ 219,
+ 357,
+ 31,
+ 15
+ ],
+ "category_id": 4,
+ "id": 4146,
+ "area": 465,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2097,
+ "bbox": [
+ 455,
+ 55,
+ 19,
+ 24
+ ],
+ "category_id": 4,
+ "id": 4147,
+ "area": 456,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2097,
+ "bbox": [
+ 476,
+ 362,
+ 30,
+ 19
+ ],
+ "category_id": 4,
+ "id": 4148,
+ "area": 570,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2098,
+ "bbox": [
+ 184,
+ 0,
+ 109,
+ 407
+ ],
+ "category_id": 16,
+ "id": 4149,
+ "area": 44363,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2099,
+ "bbox": [
+ 122,
+ 255,
+ 264,
+ 161
+ ],
+ "category_id": 16,
+ "id": 4150,
+ "area": 42504,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2100,
+ "bbox": [
+ 158,
+ 222,
+ 209,
+ 46
+ ],
+ "category_id": 19,
+ "id": 4151,
+ "area": 9614,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2101,
+ "bbox": [
+ 55,
+ 272,
+ 312,
+ 95
+ ],
+ "category_id": 19,
+ "id": 4152,
+ "area": 29640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2102,
+ "bbox": [
+ 16,
+ 80,
+ 124,
+ 44
+ ],
+ "category_id": 19,
+ "id": 4153,
+ "area": 5456,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2103,
+ "bbox": [
+ 65,
+ 172,
+ 62,
+ 39
+ ],
+ "category_id": 4,
+ "id": 4154,
+ "area": 2418,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2103,
+ "bbox": [
+ 414,
+ 247,
+ 63,
+ 39
+ ],
+ "category_id": 4,
+ "id": 4155,
+ "area": 2457,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2104,
+ "bbox": [
+ 156,
+ 140,
+ 302,
+ 265
+ ],
+ "category_id": 10,
+ "id": 4156,
+ "area": 80030,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2105,
+ "bbox": [
+ 246,
+ 41,
+ 162,
+ 133
+ ],
+ "category_id": 15,
+ "id": 4157,
+ "area": 21546,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2105,
+ "bbox": [
+ 145,
+ 178,
+ 162,
+ 133
+ ],
+ "category_id": 15,
+ "id": 4158,
+ "area": 21546,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2105,
+ "bbox": [
+ 288,
+ 279,
+ 157,
+ 131
+ ],
+ "category_id": 15,
+ "id": 4159,
+ "area": 20567,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2105,
+ "bbox": [
+ 48,
+ 318,
+ 162,
+ 130
+ ],
+ "category_id": 15,
+ "id": 4160,
+ "area": 21060,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2106,
+ "bbox": [
+ 71,
+ 194,
+ 397,
+ 73
+ ],
+ "category_id": 19,
+ "id": 4161,
+ "area": 28981,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2107,
+ "bbox": [
+ 216,
+ 264,
+ 72,
+ 79
+ ],
+ "category_id": 4,
+ "id": 4162,
+ "area": 5688,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2108,
+ "bbox": [
+ 220,
+ 112,
+ 169,
+ 305
+ ],
+ "category_id": 10,
+ "id": 4163,
+ "area": 51545,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2109,
+ "bbox": [
+ 86,
+ 316,
+ 35,
+ 100
+ ],
+ "category_id": 15,
+ "id": 4164,
+ "area": 3500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2109,
+ "bbox": [
+ 267,
+ 160,
+ 97,
+ 116
+ ],
+ "category_id": 15,
+ "id": 4165,
+ "area": 11252,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2109,
+ "bbox": [
+ 259,
+ 304,
+ 98,
+ 118
+ ],
+ "category_id": 15,
+ "id": 4166,
+ "area": 11564,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2110,
+ "bbox": [
+ 97,
+ 106,
+ 304,
+ 306
+ ],
+ "category_id": 10,
+ "id": 4167,
+ "area": 93024,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2111,
+ "bbox": [
+ 101,
+ 231,
+ 73,
+ 168
+ ],
+ "category_id": 19,
+ "id": 4168,
+ "area": 12264,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2112,
+ "bbox": [
+ 131,
+ 156,
+ 269,
+ 224
+ ],
+ "category_id": 10,
+ "id": 4169,
+ "area": 60256,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2113,
+ "bbox": [
+ 233,
+ 86,
+ 99,
+ 94
+ ],
+ "category_id": 15,
+ "id": 4170,
+ "area": 9306,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2113,
+ "bbox": [
+ 254,
+ 220,
+ 99,
+ 97
+ ],
+ "category_id": 15,
+ "id": 4171,
+ "area": 9603,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2114,
+ "bbox": [
+ 186,
+ 149,
+ 101,
+ 286
+ ],
+ "category_id": 16,
+ "id": 4172,
+ "area": 28886,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2115,
+ "bbox": [
+ 68,
+ 138,
+ 68,
+ 32
+ ],
+ "category_id": 4,
+ "id": 4173,
+ "area": 2176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2115,
+ "bbox": [
+ 104,
+ 357,
+ 46,
+ 40
+ ],
+ "category_id": 4,
+ "id": 4174,
+ "area": 1840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2116,
+ "bbox": [
+ 37,
+ 146,
+ 75,
+ 74
+ ],
+ "category_id": 15,
+ "id": 4175,
+ "area": 5550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2116,
+ "bbox": [
+ 151,
+ 154,
+ 80,
+ 77
+ ],
+ "category_id": 15,
+ "id": 4176,
+ "area": 6160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2116,
+ "bbox": [
+ 233,
+ 323,
+ 96,
+ 96
+ ],
+ "category_id": 15,
+ "id": 4177,
+ "area": 9216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2116,
+ "bbox": [
+ 261,
+ 216,
+ 98,
+ 98
+ ],
+ "category_id": 15,
+ "id": 4178,
+ "area": 9604,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2116,
+ "bbox": [
+ 368,
+ 252,
+ 95,
+ 88
+ ],
+ "category_id": 15,
+ "id": 4179,
+ "area": 8360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2116,
+ "bbox": [
+ 340,
+ 352,
+ 95,
+ 95
+ ],
+ "category_id": 15,
+ "id": 4180,
+ "area": 9025,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2117,
+ "bbox": [
+ 120,
+ 230,
+ 231,
+ 50
+ ],
+ "category_id": 19,
+ "id": 4181,
+ "area": 11550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2118,
+ "bbox": [
+ 335,
+ 36,
+ 56,
+ 37
+ ],
+ "category_id": 4,
+ "id": 4182,
+ "area": 2072,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2118,
+ "bbox": [
+ 67,
+ 144,
+ 53,
+ 31
+ ],
+ "category_id": 4,
+ "id": 4183,
+ "area": 1643,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2118,
+ "bbox": [
+ 84,
+ 438,
+ 56,
+ 38
+ ],
+ "category_id": 4,
+ "id": 4184,
+ "area": 2128,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2119,
+ "bbox": [
+ 139,
+ 147,
+ 93,
+ 133
+ ],
+ "category_id": 15,
+ "id": 4185,
+ "area": 12369,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2119,
+ "bbox": [
+ 287,
+ 136,
+ 93,
+ 130
+ ],
+ "category_id": 15,
+ "id": 4186,
+ "area": 12090,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2119,
+ "bbox": [
+ 275,
+ 284,
+ 148,
+ 201
+ ],
+ "category_id": 15,
+ "id": 4187,
+ "area": 29748,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2120,
+ "bbox": [
+ 88,
+ 175,
+ 66,
+ 77
+ ],
+ "category_id": 15,
+ "id": 4188,
+ "area": 5082,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2120,
+ "bbox": [
+ 184,
+ 143,
+ 71,
+ 83
+ ],
+ "category_id": 15,
+ "id": 4189,
+ "area": 5893,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2120,
+ "bbox": [
+ 206,
+ 248,
+ 66,
+ 79
+ ],
+ "category_id": 15,
+ "id": 4190,
+ "area": 5214,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2120,
+ "bbox": [
+ 221,
+ 39,
+ 69,
+ 73
+ ],
+ "category_id": 15,
+ "id": 4191,
+ "area": 5037,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2121,
+ "bbox": [
+ 189,
+ 105,
+ 186,
+ 350
+ ],
+ "category_id": 16,
+ "id": 4192,
+ "area": 65100,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2122,
+ "bbox": [
+ 192,
+ 171,
+ 16,
+ 22
+ ],
+ "category_id": 4,
+ "id": 4193,
+ "area": 352,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2122,
+ "bbox": [
+ 72,
+ 446,
+ 19,
+ 21
+ ],
+ "category_id": 4,
+ "id": 4194,
+ "area": 399,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2122,
+ "bbox": [
+ 229,
+ 393,
+ 19,
+ 20
+ ],
+ "category_id": 4,
+ "id": 4195,
+ "area": 380,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2122,
+ "bbox": [
+ 343,
+ 410,
+ 18,
+ 22
+ ],
+ "category_id": 4,
+ "id": 4196,
+ "area": 396,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2122,
+ "bbox": [
+ 450,
+ 239,
+ 21,
+ 20
+ ],
+ "category_id": 4,
+ "id": 4197,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2122,
+ "bbox": [
+ 382,
+ 61,
+ 21,
+ 22
+ ],
+ "category_id": 4,
+ "id": 4198,
+ "area": 462,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2123,
+ "bbox": [
+ 122,
+ 112,
+ 64,
+ 264
+ ],
+ "category_id": 10,
+ "id": 4199,
+ "area": 16896,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2124,
+ "bbox": [
+ 141,
+ 240,
+ 282,
+ 77
+ ],
+ "category_id": 16,
+ "id": 4200,
+ "area": 21714,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2125,
+ "bbox": [
+ 74,
+ 272,
+ 146,
+ 154
+ ],
+ "category_id": 19,
+ "id": 4201,
+ "area": 22484,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2126,
+ "bbox": [
+ 225,
+ 58,
+ 77,
+ 367
+ ],
+ "category_id": 19,
+ "id": 4202,
+ "area": 28259,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2127,
+ "bbox": [
+ 111,
+ 227,
+ 296,
+ 132
+ ],
+ "category_id": 16,
+ "id": 4203,
+ "area": 39072,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2128,
+ "bbox": [
+ 418,
+ 396,
+ 42,
+ 14
+ ],
+ "category_id": 16,
+ "id": 4204,
+ "area": 588,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2129,
+ "bbox": [
+ 95,
+ 60,
+ 361,
+ 269
+ ],
+ "category_id": 10,
+ "id": 4205,
+ "area": 97109,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2130,
+ "bbox": [
+ 72,
+ 349,
+ 29,
+ 62
+ ],
+ "category_id": 4,
+ "id": 4206,
+ "area": 1798,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2130,
+ "bbox": [
+ 431,
+ 161,
+ 29,
+ 68
+ ],
+ "category_id": 4,
+ "id": 4207,
+ "area": 1972,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2131,
+ "bbox": [
+ 357,
+ 120,
+ 53,
+ 235
+ ],
+ "category_id": 16,
+ "id": 4208,
+ "area": 12455,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2131,
+ "bbox": [
+ 185,
+ 414,
+ 115,
+ 39
+ ],
+ "category_id": 16,
+ "id": 4209,
+ "area": 4485,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2132,
+ "bbox": [
+ 240,
+ 105,
+ 163,
+ 275
+ ],
+ "category_id": 16,
+ "id": 4210,
+ "area": 44825,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2133,
+ "bbox": [
+ 122,
+ 24,
+ 337,
+ 43
+ ],
+ "category_id": 10,
+ "id": 4211,
+ "area": 14491,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2134,
+ "bbox": [
+ 282,
+ 76,
+ 36,
+ 43
+ ],
+ "category_id": 4,
+ "id": 4212,
+ "area": 1548,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2134,
+ "bbox": [
+ 213,
+ 397,
+ 42,
+ 48
+ ],
+ "category_id": 4,
+ "id": 4213,
+ "area": 2016,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2135,
+ "bbox": [
+ 208,
+ 128,
+ 217,
+ 303
+ ],
+ "category_id": 10,
+ "id": 4214,
+ "area": 65751,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2136,
+ "bbox": [
+ 113,
+ 303,
+ 127,
+ 102
+ ],
+ "category_id": 10,
+ "id": 4215,
+ "area": 12954,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2137,
+ "bbox": [
+ 81,
+ 108,
+ 374,
+ 313
+ ],
+ "category_id": 10,
+ "id": 4216,
+ "area": 117062,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2138,
+ "bbox": [
+ 194,
+ 207,
+ 130,
+ 65
+ ],
+ "category_id": 19,
+ "id": 4217,
+ "area": 8450,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2139,
+ "bbox": [
+ 114,
+ 261,
+ 91,
+ 75
+ ],
+ "category_id": 15,
+ "id": 4218,
+ "area": 6825,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2139,
+ "bbox": [
+ 243,
+ 252,
+ 89,
+ 82
+ ],
+ "category_id": 15,
+ "id": 4219,
+ "area": 7298,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2140,
+ "bbox": [
+ 37,
+ 83,
+ 21,
+ 20
+ ],
+ "category_id": 4,
+ "id": 4220,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2140,
+ "bbox": [
+ 215,
+ 112,
+ 21,
+ 28
+ ],
+ "category_id": 4,
+ "id": 4221,
+ "area": 588,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2140,
+ "bbox": [
+ 163,
+ 329,
+ 20,
+ 23
+ ],
+ "category_id": 4,
+ "id": 4222,
+ "area": 460,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2140,
+ "bbox": [
+ 422,
+ 163,
+ 17,
+ 20
+ ],
+ "category_id": 4,
+ "id": 4223,
+ "area": 340,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2140,
+ "bbox": [
+ 326,
+ 456,
+ 20,
+ 22
+ ],
+ "category_id": 4,
+ "id": 4224,
+ "area": 440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2141,
+ "bbox": [
+ 52,
+ 270,
+ 124,
+ 190
+ ],
+ "category_id": 10,
+ "id": 4225,
+ "area": 23560,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2142,
+ "bbox": [
+ 178,
+ 398,
+ 96,
+ 50
+ ],
+ "category_id": 15,
+ "id": 4226,
+ "area": 4800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2143,
+ "bbox": [
+ 376,
+ 63,
+ 28,
+ 25
+ ],
+ "category_id": 4,
+ "id": 4227,
+ "area": 700,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2143,
+ "bbox": [
+ 398,
+ 181,
+ 25,
+ 27
+ ],
+ "category_id": 4,
+ "id": 4228,
+ "area": 675,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2143,
+ "bbox": [
+ 397,
+ 300,
+ 26,
+ 28
+ ],
+ "category_id": 4,
+ "id": 4229,
+ "area": 728,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2143,
+ "bbox": [
+ 398,
+ 419,
+ 22,
+ 23
+ ],
+ "category_id": 4,
+ "id": 4230,
+ "area": 506,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2144,
+ "bbox": [
+ 173,
+ 268,
+ 129,
+ 57
+ ],
+ "category_id": 19,
+ "id": 4231,
+ "area": 7353,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2145,
+ "bbox": [
+ 208,
+ 172,
+ 37,
+ 155
+ ],
+ "category_id": 19,
+ "id": 4232,
+ "area": 5735,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2146,
+ "bbox": [
+ 275,
+ 81,
+ 36,
+ 31
+ ],
+ "category_id": 4,
+ "id": 4233,
+ "area": 1116,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2146,
+ "bbox": [
+ 56,
+ 286,
+ 39,
+ 34
+ ],
+ "category_id": 4,
+ "id": 4234,
+ "area": 1326,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2146,
+ "bbox": [
+ 417,
+ 404,
+ 64,
+ 54
+ ],
+ "category_id": 4,
+ "id": 4235,
+ "area": 3456,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2147,
+ "bbox": [
+ 233,
+ 109,
+ 7,
+ 6
+ ],
+ "category_id": 4,
+ "id": 4236,
+ "area": 42,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2147,
+ "bbox": [
+ 197,
+ 197,
+ 11,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4237,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2147,
+ "bbox": [
+ 210,
+ 257,
+ 11,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4238,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2147,
+ "bbox": [
+ 264,
+ 304,
+ 10,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4239,
+ "area": 80,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2147,
+ "bbox": [
+ 449,
+ 159,
+ 8,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4240,
+ "area": 64,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2148,
+ "bbox": [
+ 383,
+ 350,
+ 22,
+ 27
+ ],
+ "category_id": 16,
+ "id": 4241,
+ "area": 594,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2149,
+ "bbox": [
+ 119,
+ 104,
+ 140,
+ 304
+ ],
+ "category_id": 16,
+ "id": 4242,
+ "area": 42560,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2150,
+ "bbox": [
+ 126,
+ 187,
+ 247,
+ 141
+ ],
+ "category_id": 19,
+ "id": 4243,
+ "area": 34827,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2151,
+ "bbox": [
+ 39,
+ 32,
+ 169,
+ 85
+ ],
+ "category_id": 10,
+ "id": 4244,
+ "area": 14365,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2152,
+ "bbox": [
+ 26,
+ 253,
+ 142,
+ 93
+ ],
+ "category_id": 10,
+ "id": 4245,
+ "area": 13206,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2153,
+ "bbox": [
+ 54,
+ 65,
+ 26,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4246,
+ "area": 416,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2153,
+ "bbox": [
+ 328,
+ 74,
+ 15,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4247,
+ "area": 165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2153,
+ "bbox": [
+ 410,
+ 227,
+ 18,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4248,
+ "area": 252,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2153,
+ "bbox": [
+ 85,
+ 268,
+ 21,
+ 12
+ ],
+ "category_id": 4,
+ "id": 4249,
+ "area": 252,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2153,
+ "bbox": [
+ 11,
+ 376,
+ 22,
+ 21
+ ],
+ "category_id": 4,
+ "id": 4250,
+ "area": 462,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2153,
+ "bbox": [
+ 121,
+ 445,
+ 32,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4251,
+ "area": 288,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2153,
+ "bbox": [
+ 318,
+ 456,
+ 19,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4252,
+ "area": 171,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2153,
+ "bbox": [
+ 419,
+ 416,
+ 20,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4253,
+ "area": 140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2153,
+ "bbox": [
+ 202,
+ 149,
+ 16,
+ 12
+ ],
+ "category_id": 4,
+ "id": 4254,
+ "area": 192,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2154,
+ "bbox": [
+ 223,
+ 102,
+ 94,
+ 129
+ ],
+ "category_id": 15,
+ "id": 4255,
+ "area": 12126,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2155,
+ "bbox": [
+ 200,
+ 82,
+ 70,
+ 326
+ ],
+ "category_id": 19,
+ "id": 4256,
+ "area": 22820,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2156,
+ "bbox": [
+ 51,
+ 201,
+ 99,
+ 100
+ ],
+ "category_id": 15,
+ "id": 4257,
+ "area": 9900,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2156,
+ "bbox": [
+ 174,
+ 305,
+ 119,
+ 132
+ ],
+ "category_id": 15,
+ "id": 4258,
+ "area": 15708,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2157,
+ "bbox": [
+ 161,
+ 30,
+ 23,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4259,
+ "area": 161,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2157,
+ "bbox": [
+ 368,
+ 158,
+ 17,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4260,
+ "area": 153,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2157,
+ "bbox": [
+ 478,
+ 278,
+ 14,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4261,
+ "area": 154,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2157,
+ "bbox": [
+ 104,
+ 350,
+ 24,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4262,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2157,
+ "bbox": [
+ 141,
+ 492,
+ 22,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4263,
+ "area": 242,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2158,
+ "bbox": [
+ 180,
+ 146,
+ 113,
+ 127
+ ],
+ "category_id": 15,
+ "id": 4264,
+ "area": 14351,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2159,
+ "bbox": [
+ 58,
+ 67,
+ 334,
+ 39
+ ],
+ "category_id": 10,
+ "id": 4265,
+ "area": 13026,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2160,
+ "bbox": [
+ 73,
+ 215,
+ 370,
+ 47
+ ],
+ "category_id": 19,
+ "id": 4266,
+ "area": 17390,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2161,
+ "bbox": [
+ 9,
+ 117,
+ 474,
+ 227
+ ],
+ "category_id": 10,
+ "id": 4267,
+ "area": 107598,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2162,
+ "bbox": [
+ 203,
+ 254,
+ 74,
+ 85
+ ],
+ "category_id": 4,
+ "id": 4268,
+ "area": 6290,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2163,
+ "bbox": [
+ 108,
+ 83,
+ 54,
+ 46
+ ],
+ "category_id": 10,
+ "id": 4269,
+ "area": 2484,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2163,
+ "bbox": [
+ 330,
+ 353,
+ 75,
+ 29
+ ],
+ "category_id": 10,
+ "id": 4270,
+ "area": 2175,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2164,
+ "bbox": [
+ 71,
+ 263,
+ 403,
+ 124
+ ],
+ "category_id": 16,
+ "id": 4271,
+ "area": 49972,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2165,
+ "bbox": [
+ 131,
+ 132,
+ 24,
+ 29
+ ],
+ "category_id": 4,
+ "id": 4272,
+ "area": 696,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2165,
+ "bbox": [
+ 15,
+ 332,
+ 32,
+ 31
+ ],
+ "category_id": 4,
+ "id": 4273,
+ "area": 992,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2165,
+ "bbox": [
+ 236,
+ 386,
+ 32,
+ 29
+ ],
+ "category_id": 4,
+ "id": 4274,
+ "area": 928,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2165,
+ "bbox": [
+ 307,
+ 75,
+ 32,
+ 26
+ ],
+ "category_id": 4,
+ "id": 4275,
+ "area": 832,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2165,
+ "bbox": [
+ 459,
+ 86,
+ 23,
+ 33
+ ],
+ "category_id": 4,
+ "id": 4276,
+ "area": 759,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2166,
+ "bbox": [
+ 270,
+ 253,
+ 23,
+ 93
+ ],
+ "category_id": 4,
+ "id": 4277,
+ "area": 2139,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2167,
+ "bbox": [
+ 125,
+ 310,
+ 87,
+ 78
+ ],
+ "category_id": 10,
+ "id": 4278,
+ "area": 6786,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2168,
+ "bbox": [
+ 371,
+ 95,
+ 11,
+ 21
+ ],
+ "category_id": 15,
+ "id": 4279,
+ "area": 231,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2169,
+ "bbox": [
+ 200,
+ 57,
+ 120,
+ 76
+ ],
+ "category_id": 15,
+ "id": 4280,
+ "area": 9120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2169,
+ "bbox": [
+ 169,
+ 218,
+ 127,
+ 104
+ ],
+ "category_id": 15,
+ "id": 4281,
+ "area": 13208,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2170,
+ "bbox": [
+ 90,
+ 303,
+ 183,
+ 129
+ ],
+ "category_id": 19,
+ "id": 4282,
+ "area": 23607,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2171,
+ "bbox": [
+ 185,
+ 122,
+ 217,
+ 265
+ ],
+ "category_id": 10,
+ "id": 4283,
+ "area": 57505,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2172,
+ "bbox": [
+ 103,
+ 256,
+ 354,
+ 70
+ ],
+ "category_id": 10,
+ "id": 4284,
+ "area": 24780,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2173,
+ "bbox": [
+ 33,
+ 140,
+ 375,
+ 308
+ ],
+ "category_id": 16,
+ "id": 4285,
+ "area": 115500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2174,
+ "bbox": [
+ 182,
+ 204,
+ 118,
+ 124
+ ],
+ "category_id": 15,
+ "id": 4286,
+ "area": 14632,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2174,
+ "bbox": [
+ 342,
+ 216,
+ 122,
+ 123
+ ],
+ "category_id": 15,
+ "id": 4287,
+ "area": 15006,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2175,
+ "bbox": [
+ 329,
+ 234,
+ 103,
+ 90
+ ],
+ "category_id": 16,
+ "id": 4288,
+ "area": 9270,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2176,
+ "bbox": [
+ 452,
+ 327,
+ 5,
+ 5
+ ],
+ "category_id": 4,
+ "id": 4289,
+ "area": 25,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2176,
+ "bbox": [
+ 329,
+ 457,
+ 6,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4290,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2176,
+ "bbox": [
+ 279,
+ 400,
+ 6,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4291,
+ "area": 42,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2176,
+ "bbox": [
+ 176,
+ 312,
+ 7,
+ 13
+ ],
+ "category_id": 4,
+ "id": 4292,
+ "area": 91,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2176,
+ "bbox": [
+ 110,
+ 283,
+ 6,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4293,
+ "area": 66,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2176,
+ "bbox": [
+ 93,
+ 65,
+ 6,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4294,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2176,
+ "bbox": [
+ 14,
+ 75,
+ 5,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4295,
+ "area": 40,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2177,
+ "bbox": [
+ 263,
+ 77,
+ 67,
+ 369
+ ],
+ "category_id": 10,
+ "id": 4296,
+ "area": 24723,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2178,
+ "bbox": [
+ 73,
+ 271,
+ 10,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4297,
+ "area": 80,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2178,
+ "bbox": [
+ 136,
+ 280,
+ 9,
+ 6
+ ],
+ "category_id": 4,
+ "id": 4298,
+ "area": 54,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2178,
+ "bbox": [
+ 295,
+ 246,
+ 7,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4299,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2178,
+ "bbox": [
+ 372,
+ 201,
+ 8,
+ 5
+ ],
+ "category_id": 4,
+ "id": 4300,
+ "area": 40,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2178,
+ "bbox": [
+ 442,
+ 151,
+ 9,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4301,
+ "area": 72,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2178,
+ "bbox": [
+ 492,
+ 103,
+ 7,
+ 6
+ ],
+ "category_id": 4,
+ "id": 4302,
+ "area": 42,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2178,
+ "bbox": [
+ 447,
+ 467,
+ 6,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4303,
+ "area": 42,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2178,
+ "bbox": [
+ 350,
+ 461,
+ 8,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4304,
+ "area": 64,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2178,
+ "bbox": [
+ 184,
+ 459,
+ 7,
+ 5
+ ],
+ "category_id": 4,
+ "id": 4305,
+ "area": 35,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2179,
+ "bbox": [
+ 224,
+ 293,
+ 89,
+ 55
+ ],
+ "category_id": 16,
+ "id": 4306,
+ "area": 4895,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2180,
+ "bbox": [
+ 288,
+ 172,
+ 107,
+ 94
+ ],
+ "category_id": 15,
+ "id": 4307,
+ "area": 10058,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2181,
+ "bbox": [
+ 260,
+ 154,
+ 36,
+ 18
+ ],
+ "category_id": 4,
+ "id": 4308,
+ "area": 648,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2181,
+ "bbox": [
+ 468,
+ 42,
+ 34,
+ 20
+ ],
+ "category_id": 4,
+ "id": 4309,
+ "area": 680,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2181,
+ "bbox": [
+ 441,
+ 165,
+ 31,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4310,
+ "area": 496,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2181,
+ "bbox": [
+ 358,
+ 304,
+ 34,
+ 20
+ ],
+ "category_id": 4,
+ "id": 4311,
+ "area": 680,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2181,
+ "bbox": [
+ 3,
+ 293,
+ 39,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4312,
+ "area": 624,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2181,
+ "bbox": [
+ 316,
+ 468,
+ 36,
+ 22
+ ],
+ "category_id": 4,
+ "id": 4313,
+ "area": 792,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2182,
+ "bbox": [
+ 165,
+ 220,
+ 88,
+ 100
+ ],
+ "category_id": 15,
+ "id": 4314,
+ "area": 8800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2182,
+ "bbox": [
+ 289,
+ 211,
+ 92,
+ 111
+ ],
+ "category_id": 15,
+ "id": 4315,
+ "area": 10212,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2183,
+ "bbox": [
+ 216,
+ 194,
+ 47,
+ 135
+ ],
+ "category_id": 19,
+ "id": 4316,
+ "area": 6345,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2184,
+ "bbox": [
+ 150,
+ 174,
+ 113,
+ 128
+ ],
+ "category_id": 15,
+ "id": 4317,
+ "area": 14464,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2184,
+ "bbox": [
+ 120,
+ 339,
+ 118,
+ 128
+ ],
+ "category_id": 15,
+ "id": 4318,
+ "area": 15104,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2185,
+ "bbox": [
+ 285,
+ 87,
+ 40,
+ 58
+ ],
+ "category_id": 4,
+ "id": 4319,
+ "area": 2320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2185,
+ "bbox": [
+ 120,
+ 301,
+ 50,
+ 58
+ ],
+ "category_id": 4,
+ "id": 4320,
+ "area": 2900,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2186,
+ "bbox": [
+ 3,
+ 404,
+ 24,
+ 19
+ ],
+ "category_id": 4,
+ "id": 4321,
+ "area": 456,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2186,
+ "bbox": [
+ 196,
+ 364,
+ 25,
+ 19
+ ],
+ "category_id": 4,
+ "id": 4322,
+ "area": 475,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2186,
+ "bbox": [
+ 476,
+ 268,
+ 30,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4323,
+ "area": 480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2186,
+ "bbox": [
+ 322,
+ 380,
+ 27,
+ 18
+ ],
+ "category_id": 4,
+ "id": 4324,
+ "area": 486,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2187,
+ "bbox": [
+ 134,
+ 318,
+ 46,
+ 54
+ ],
+ "category_id": 4,
+ "id": 4325,
+ "area": 2484,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2187,
+ "bbox": [
+ 289,
+ 140,
+ 42,
+ 53
+ ],
+ "category_id": 4,
+ "id": 4326,
+ "area": 2226,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2188,
+ "bbox": [
+ 64,
+ 290,
+ 247,
+ 46
+ ],
+ "category_id": 19,
+ "id": 4327,
+ "area": 11362,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2188,
+ "bbox": [
+ 18,
+ 214,
+ 101,
+ 58
+ ],
+ "category_id": 19,
+ "id": 4328,
+ "area": 5858,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2189,
+ "bbox": [
+ 119,
+ 389,
+ 77,
+ 38
+ ],
+ "category_id": 4,
+ "id": 4329,
+ "area": 2926,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2189,
+ "bbox": [
+ 400,
+ 162,
+ 67,
+ 46
+ ],
+ "category_id": 4,
+ "id": 4330,
+ "area": 3082,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2190,
+ "bbox": [
+ 275,
+ 44,
+ 37,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4331,
+ "area": 592,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2190,
+ "bbox": [
+ 442,
+ 271,
+ 33,
+ 24
+ ],
+ "category_id": 4,
+ "id": 4332,
+ "area": 792,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2190,
+ "bbox": [
+ 293,
+ 416,
+ 35,
+ 23
+ ],
+ "category_id": 4,
+ "id": 4333,
+ "area": 805,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2190,
+ "bbox": [
+ 56,
+ 429,
+ 30,
+ 19
+ ],
+ "category_id": 4,
+ "id": 4334,
+ "area": 570,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2191,
+ "bbox": [
+ 115,
+ 261,
+ 142,
+ 92
+ ],
+ "category_id": 16,
+ "id": 4335,
+ "area": 13064,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2192,
+ "bbox": [
+ 21,
+ 70,
+ 75,
+ 153
+ ],
+ "category_id": 10,
+ "id": 4336,
+ "area": 11475,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2193,
+ "bbox": [
+ 136,
+ 99,
+ 260,
+ 260
+ ],
+ "category_id": 10,
+ "id": 4337,
+ "area": 67600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2194,
+ "bbox": [
+ 57,
+ 26,
+ 26,
+ 18
+ ],
+ "category_id": 4,
+ "id": 4338,
+ "area": 468,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2194,
+ "bbox": [
+ 403,
+ 83,
+ 11,
+ 17
+ ],
+ "category_id": 4,
+ "id": 4339,
+ "area": 187,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2194,
+ "bbox": [
+ 438,
+ 263,
+ 27,
+ 21
+ ],
+ "category_id": 4,
+ "id": 4340,
+ "area": 567,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2194,
+ "bbox": [
+ 284,
+ 291,
+ 27,
+ 21
+ ],
+ "category_id": 4,
+ "id": 4341,
+ "area": 567,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2194,
+ "bbox": [
+ 129,
+ 302,
+ 26,
+ 23
+ ],
+ "category_id": 4,
+ "id": 4342,
+ "area": 598,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2194,
+ "bbox": [
+ 385,
+ 469,
+ 17,
+ 12
+ ],
+ "category_id": 4,
+ "id": 4343,
+ "area": 204,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2195,
+ "bbox": [
+ 102,
+ 188,
+ 86,
+ 53
+ ],
+ "category_id": 4,
+ "id": 4344,
+ "area": 4558,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2195,
+ "bbox": [
+ 395,
+ 369,
+ 94,
+ 57
+ ],
+ "category_id": 4,
+ "id": 4345,
+ "area": 5358,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2196,
+ "bbox": [
+ 76,
+ 189,
+ 118,
+ 54
+ ],
+ "category_id": 4,
+ "id": 4346,
+ "area": 6372,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2196,
+ "bbox": [
+ 376,
+ 299,
+ 106,
+ 42
+ ],
+ "category_id": 4,
+ "id": 4347,
+ "area": 4452,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2197,
+ "bbox": [
+ 157,
+ 67,
+ 143,
+ 149
+ ],
+ "category_id": 15,
+ "id": 4348,
+ "area": 21307,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2197,
+ "bbox": [
+ 317,
+ 71,
+ 144,
+ 150
+ ],
+ "category_id": 15,
+ "id": 4349,
+ "area": 21600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2197,
+ "bbox": [
+ 124,
+ 268,
+ 182,
+ 187
+ ],
+ "category_id": 15,
+ "id": 4350,
+ "area": 34034,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2198,
+ "bbox": [
+ 175,
+ 99,
+ 267,
+ 307
+ ],
+ "category_id": 10,
+ "id": 4351,
+ "area": 81969,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2199,
+ "bbox": [
+ 252,
+ 167,
+ 139,
+ 177
+ ],
+ "category_id": 16,
+ "id": 4352,
+ "area": 24603,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2200,
+ "bbox": [
+ 246,
+ 197,
+ 188,
+ 78
+ ],
+ "category_id": 19,
+ "id": 4353,
+ "area": 14664,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2201,
+ "bbox": [
+ 76,
+ 297,
+ 64,
+ 95
+ ],
+ "category_id": 4,
+ "id": 4354,
+ "area": 6080,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2201,
+ "bbox": [
+ 368,
+ 231,
+ 56,
+ 84
+ ],
+ "category_id": 4,
+ "id": 4355,
+ "area": 4704,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2202,
+ "bbox": [
+ 256,
+ 113,
+ 76,
+ 327
+ ],
+ "category_id": 16,
+ "id": 4356,
+ "area": 24852,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2203,
+ "bbox": [
+ 341,
+ 374,
+ 23,
+ 24
+ ],
+ "category_id": 16,
+ "id": 4357,
+ "area": 552,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2204,
+ "bbox": [
+ 144,
+ 215,
+ 170,
+ 71
+ ],
+ "category_id": 19,
+ "id": 4358,
+ "area": 12070,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2205,
+ "bbox": [
+ 207,
+ 150,
+ 278,
+ 232
+ ],
+ "category_id": 10,
+ "id": 4359,
+ "area": 64496,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2206,
+ "bbox": [
+ 122,
+ 428,
+ 49,
+ 18
+ ],
+ "category_id": 4,
+ "id": 4360,
+ "area": 882,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2206,
+ "bbox": [
+ 314,
+ 295,
+ 49,
+ 26
+ ],
+ "category_id": 4,
+ "id": 4361,
+ "area": 1274,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2206,
+ "bbox": [
+ 304,
+ 62,
+ 47,
+ 23
+ ],
+ "category_id": 4,
+ "id": 4362,
+ "area": 1081,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2207,
+ "bbox": [
+ 129,
+ 27,
+ 24,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4363,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2207,
+ "bbox": [
+ 28,
+ 155,
+ 15,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4364,
+ "area": 165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2207,
+ "bbox": [
+ 355,
+ 224,
+ 20,
+ 12
+ ],
+ "category_id": 4,
+ "id": 4365,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2207,
+ "bbox": [
+ 162,
+ 308,
+ 19,
+ 15
+ ],
+ "category_id": 4,
+ "id": 4366,
+ "area": 285,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2207,
+ "bbox": [
+ 140,
+ 462,
+ 18,
+ 19
+ ],
+ "category_id": 4,
+ "id": 4367,
+ "area": 342,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2207,
+ "bbox": [
+ 364,
+ 408,
+ 20,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4368,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2208,
+ "bbox": [
+ 88,
+ 68,
+ 13,
+ 26
+ ],
+ "category_id": 4,
+ "id": 4369,
+ "area": 338,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2208,
+ "bbox": [
+ 370,
+ 88,
+ 12,
+ 31
+ ],
+ "category_id": 4,
+ "id": 4370,
+ "area": 372,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2208,
+ "bbox": [
+ 241,
+ 251,
+ 10,
+ 28
+ ],
+ "category_id": 4,
+ "id": 4371,
+ "area": 280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2208,
+ "bbox": [
+ 436,
+ 289,
+ 10,
+ 26
+ ],
+ "category_id": 4,
+ "id": 4372,
+ "area": 260,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2208,
+ "bbox": [
+ 128,
+ 439,
+ 13,
+ 29
+ ],
+ "category_id": 4,
+ "id": 4373,
+ "area": 377,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2209,
+ "bbox": [
+ 216,
+ 211,
+ 8,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4374,
+ "area": 88,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2209,
+ "bbox": [
+ 196,
+ 241,
+ 8,
+ 13
+ ],
+ "category_id": 4,
+ "id": 4375,
+ "area": 104,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2209,
+ "bbox": [
+ 154,
+ 318,
+ 7,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4376,
+ "area": 70,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2209,
+ "bbox": [
+ 184,
+ 290,
+ 6,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4377,
+ "area": 48,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2209,
+ "bbox": [
+ 200,
+ 350,
+ 4,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4378,
+ "area": 40,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2209,
+ "bbox": [
+ 282,
+ 265,
+ 5,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4379,
+ "area": 40,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2209,
+ "bbox": [
+ 326,
+ 267,
+ 8,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4380,
+ "area": 64,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2209,
+ "bbox": [
+ 353,
+ 241,
+ 8,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4381,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2210,
+ "bbox": [
+ 79,
+ 246,
+ 418,
+ 45
+ ],
+ "category_id": 10,
+ "id": 4382,
+ "area": 18810,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2211,
+ "bbox": [
+ 211,
+ 233,
+ 73,
+ 117
+ ],
+ "category_id": 15,
+ "id": 4383,
+ "area": 8541,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2212,
+ "bbox": [
+ 71,
+ 208,
+ 65,
+ 39
+ ],
+ "category_id": 4,
+ "id": 4384,
+ "area": 2535,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2212,
+ "bbox": [
+ 420,
+ 343,
+ 56,
+ 43
+ ],
+ "category_id": 4,
+ "id": 4385,
+ "area": 2408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2213,
+ "bbox": [
+ 186,
+ 155,
+ 118,
+ 205
+ ],
+ "category_id": 10,
+ "id": 4386,
+ "area": 24190,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2214,
+ "bbox": [
+ 44,
+ 167,
+ 339,
+ 300
+ ],
+ "category_id": 16,
+ "id": 4387,
+ "area": 101700,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2215,
+ "bbox": [
+ 332,
+ 497,
+ 27,
+ 15
+ ],
+ "category_id": 16,
+ "id": 4388,
+ "area": 405,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2216,
+ "bbox": [
+ 168,
+ 165,
+ 200,
+ 263
+ ],
+ "category_id": 16,
+ "id": 4389,
+ "area": 52600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2217,
+ "bbox": [
+ 211,
+ 124,
+ 123,
+ 124
+ ],
+ "category_id": 15,
+ "id": 4390,
+ "area": 15252,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2217,
+ "bbox": [
+ 218,
+ 266,
+ 122,
+ 134
+ ],
+ "category_id": 15,
+ "id": 4391,
+ "area": 16348,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2218,
+ "bbox": [
+ 225,
+ 98,
+ 61,
+ 50
+ ],
+ "category_id": 4,
+ "id": 4392,
+ "area": 3050,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2218,
+ "bbox": [
+ 186,
+ 430,
+ 53,
+ 44
+ ],
+ "category_id": 4,
+ "id": 4393,
+ "area": 2332,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2219,
+ "bbox": [
+ 78,
+ 177,
+ 119,
+ 113
+ ],
+ "category_id": 15,
+ "id": 4394,
+ "area": 13447,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2219,
+ "bbox": [
+ 296,
+ 128,
+ 155,
+ 145
+ ],
+ "category_id": 15,
+ "id": 4395,
+ "area": 22475,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2219,
+ "bbox": [
+ 78,
+ 336,
+ 117,
+ 112
+ ],
+ "category_id": 15,
+ "id": 4396,
+ "area": 13104,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2219,
+ "bbox": [
+ 295,
+ 325,
+ 152,
+ 144
+ ],
+ "category_id": 15,
+ "id": 4397,
+ "area": 21888,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2220,
+ "bbox": [
+ 216,
+ 84,
+ 35,
+ 319
+ ],
+ "category_id": 19,
+ "id": 4398,
+ "area": 11165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2221,
+ "bbox": [
+ 74,
+ 119,
+ 388,
+ 142
+ ],
+ "category_id": 10,
+ "id": 4399,
+ "area": 55096,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2222,
+ "bbox": [
+ 34,
+ 208,
+ 17,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4400,
+ "area": 272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2222,
+ "bbox": [
+ 198,
+ 193,
+ 18,
+ 17
+ ],
+ "category_id": 4,
+ "id": 4401,
+ "area": 306,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2222,
+ "bbox": [
+ 241,
+ 27,
+ 16,
+ 17
+ ],
+ "category_id": 4,
+ "id": 4402,
+ "area": 272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2222,
+ "bbox": [
+ 389,
+ 103,
+ 15,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4403,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2222,
+ "bbox": [
+ 291,
+ 339,
+ 21,
+ 17
+ ],
+ "category_id": 4,
+ "id": 4404,
+ "area": 357,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2222,
+ "bbox": [
+ 193,
+ 414,
+ 17,
+ 17
+ ],
+ "category_id": 4,
+ "id": 4405,
+ "area": 289,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2222,
+ "bbox": [
+ 397,
+ 465,
+ 20,
+ 20
+ ],
+ "category_id": 4,
+ "id": 4406,
+ "area": 400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2223,
+ "bbox": [
+ 287,
+ 33,
+ 120,
+ 434
+ ],
+ "category_id": 10,
+ "id": 4407,
+ "area": 52080,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2224,
+ "bbox": [
+ 75,
+ 62,
+ 9,
+ 25
+ ],
+ "category_id": 4,
+ "id": 4408,
+ "area": 225,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2224,
+ "bbox": [
+ 267,
+ 155,
+ 19,
+ 22
+ ],
+ "category_id": 4,
+ "id": 4409,
+ "area": 418,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2224,
+ "bbox": [
+ 464,
+ 168,
+ 15,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4410,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2224,
+ "bbox": [
+ 23,
+ 391,
+ 14,
+ 24
+ ],
+ "category_id": 4,
+ "id": 4411,
+ "area": 336,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2224,
+ "bbox": [
+ 250,
+ 395,
+ 14,
+ 24
+ ],
+ "category_id": 4,
+ "id": 4412,
+ "area": 336,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2224,
+ "bbox": [
+ 440,
+ 425,
+ 14,
+ 22
+ ],
+ "category_id": 4,
+ "id": 4413,
+ "area": 308,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2225,
+ "bbox": [
+ 184,
+ 211,
+ 82,
+ 89
+ ],
+ "category_id": 19,
+ "id": 4414,
+ "area": 7298,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2226,
+ "bbox": [
+ 172,
+ 410,
+ 58,
+ 39
+ ],
+ "category_id": 4,
+ "id": 4415,
+ "area": 2262,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2226,
+ "bbox": [
+ 314,
+ 51,
+ 49,
+ 32
+ ],
+ "category_id": 4,
+ "id": 4416,
+ "area": 1568,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2227,
+ "bbox": [
+ 124,
+ 93,
+ 41,
+ 33
+ ],
+ "category_id": 4,
+ "id": 4417,
+ "area": 1353,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2227,
+ "bbox": [
+ 401,
+ 405,
+ 58,
+ 31
+ ],
+ "category_id": 4,
+ "id": 4418,
+ "area": 1798,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2228,
+ "bbox": [
+ 64,
+ 126,
+ 49,
+ 50
+ ],
+ "category_id": 4,
+ "id": 4419,
+ "area": 2450,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2228,
+ "bbox": [
+ 246,
+ 442,
+ 50,
+ 50
+ ],
+ "category_id": 4,
+ "id": 4420,
+ "area": 2500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2228,
+ "bbox": [
+ 431,
+ 169,
+ 52,
+ 60
+ ],
+ "category_id": 4,
+ "id": 4421,
+ "area": 3120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2229,
+ "bbox": [
+ 396,
+ 115,
+ 43,
+ 143
+ ],
+ "category_id": 19,
+ "id": 4422,
+ "area": 6149,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2230,
+ "bbox": [
+ 131,
+ 112,
+ 266,
+ 303
+ ],
+ "category_id": 10,
+ "id": 4423,
+ "area": 80598,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2231,
+ "bbox": [
+ 130,
+ 232,
+ 195,
+ 59
+ ],
+ "category_id": 19,
+ "id": 4424,
+ "area": 11505,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2232,
+ "bbox": [
+ 255,
+ 146,
+ 125,
+ 312
+ ],
+ "category_id": 16,
+ "id": 4425,
+ "area": 39000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2233,
+ "bbox": [
+ 286,
+ 56,
+ 20,
+ 63
+ ],
+ "category_id": 4,
+ "id": 4426,
+ "area": 1260,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2233,
+ "bbox": [
+ 218,
+ 424,
+ 20,
+ 57
+ ],
+ "category_id": 4,
+ "id": 4427,
+ "area": 1140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2234,
+ "bbox": [
+ 114,
+ 204,
+ 147,
+ 122
+ ],
+ "category_id": 15,
+ "id": 4428,
+ "area": 17934,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2234,
+ "bbox": [
+ 274,
+ 206,
+ 143,
+ 119
+ ],
+ "category_id": 15,
+ "id": 4429,
+ "area": 17017,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2235,
+ "bbox": [
+ 280,
+ 97,
+ 60,
+ 37
+ ],
+ "category_id": 4,
+ "id": 4430,
+ "area": 2220,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2235,
+ "bbox": [
+ 246,
+ 350,
+ 79,
+ 31
+ ],
+ "category_id": 4,
+ "id": 4431,
+ "area": 2449,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2236,
+ "bbox": [
+ 172,
+ 218,
+ 123,
+ 33
+ ],
+ "category_id": 19,
+ "id": 4432,
+ "area": 4059,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2237,
+ "bbox": [
+ 120,
+ 162,
+ 25,
+ 67
+ ],
+ "category_id": 16,
+ "id": 4433,
+ "area": 1675,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2238,
+ "bbox": [
+ 108,
+ 124,
+ 238,
+ 320
+ ],
+ "category_id": 16,
+ "id": 4434,
+ "area": 76160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2239,
+ "bbox": [
+ 290,
+ 268,
+ 126,
+ 182
+ ],
+ "category_id": 15,
+ "id": 4435,
+ "area": 22932,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2239,
+ "bbox": [
+ 284,
+ 131,
+ 117,
+ 167
+ ],
+ "category_id": 15,
+ "id": 4436,
+ "area": 19539,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2239,
+ "bbox": [
+ 177,
+ 280,
+ 85,
+ 123
+ ],
+ "category_id": 15,
+ "id": 4437,
+ "area": 10455,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2239,
+ "bbox": [
+ 71,
+ 288,
+ 83,
+ 118
+ ],
+ "category_id": 15,
+ "id": 4438,
+ "area": 9794,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2240,
+ "bbox": [
+ 195,
+ 72,
+ 53,
+ 345
+ ],
+ "category_id": 19,
+ "id": 4439,
+ "area": 18285,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2241,
+ "bbox": [
+ 33,
+ 87,
+ 13,
+ 26
+ ],
+ "category_id": 4,
+ "id": 4440,
+ "area": 338,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2241,
+ "bbox": [
+ 256,
+ 244,
+ 12,
+ 26
+ ],
+ "category_id": 4,
+ "id": 4441,
+ "area": 312,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2241,
+ "bbox": [
+ 459,
+ 181,
+ 12,
+ 33
+ ],
+ "category_id": 4,
+ "id": 4442,
+ "area": 396,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2241,
+ "bbox": [
+ 398,
+ 348,
+ 14,
+ 41
+ ],
+ "category_id": 4,
+ "id": 4443,
+ "area": 574,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2242,
+ "bbox": [
+ 52,
+ 179,
+ 433,
+ 192
+ ],
+ "category_id": 16,
+ "id": 4444,
+ "area": 83136,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2243,
+ "bbox": [
+ 83,
+ 116,
+ 39,
+ 27
+ ],
+ "category_id": 4,
+ "id": 4445,
+ "area": 1053,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2243,
+ "bbox": [
+ 116,
+ 307,
+ 43,
+ 22
+ ],
+ "category_id": 4,
+ "id": 4446,
+ "area": 946,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2243,
+ "bbox": [
+ 370,
+ 130,
+ 38,
+ 25
+ ],
+ "category_id": 4,
+ "id": 4447,
+ "area": 950,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2243,
+ "bbox": [
+ 413,
+ 361,
+ 39,
+ 28
+ ],
+ "category_id": 4,
+ "id": 4448,
+ "area": 1092,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2244,
+ "bbox": [
+ 261,
+ 259,
+ 65,
+ 94
+ ],
+ "category_id": 4,
+ "id": 4449,
+ "area": 6110,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2245,
+ "bbox": [
+ 219,
+ 222,
+ 32,
+ 156
+ ],
+ "category_id": 19,
+ "id": 4450,
+ "area": 4992,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2246,
+ "bbox": [
+ 133,
+ 102,
+ 132,
+ 140
+ ],
+ "category_id": 15,
+ "id": 4451,
+ "area": 18480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2247,
+ "bbox": [
+ 105,
+ 368,
+ 42,
+ 40
+ ],
+ "category_id": 4,
+ "id": 4452,
+ "area": 1680,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2247,
+ "bbox": [
+ 390,
+ 159,
+ 38,
+ 39
+ ],
+ "category_id": 4,
+ "id": 4453,
+ "area": 1482,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2248,
+ "bbox": [
+ 205,
+ 202,
+ 125,
+ 78
+ ],
+ "category_id": 10,
+ "id": 4454,
+ "area": 9750,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2249,
+ "bbox": [
+ 271,
+ 293,
+ 112,
+ 100
+ ],
+ "category_id": 4,
+ "id": 4455,
+ "area": 11200,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2250,
+ "bbox": [
+ 144,
+ 248,
+ 213,
+ 71
+ ],
+ "category_id": 16,
+ "id": 4456,
+ "area": 15123,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2251,
+ "bbox": [
+ 237,
+ 147,
+ 193,
+ 237
+ ],
+ "category_id": 16,
+ "id": 4457,
+ "area": 45741,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2252,
+ "bbox": [
+ 146,
+ 115,
+ 219,
+ 240
+ ],
+ "category_id": 16,
+ "id": 4458,
+ "area": 52560,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2253,
+ "bbox": [
+ 190,
+ 144,
+ 158,
+ 205
+ ],
+ "category_id": 10,
+ "id": 4459,
+ "area": 32390,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2254,
+ "bbox": [
+ 250,
+ 214,
+ 51,
+ 95
+ ],
+ "category_id": 4,
+ "id": 4460,
+ "area": 4845,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2255,
+ "bbox": [
+ 41,
+ 150,
+ 55,
+ 44
+ ],
+ "category_id": 4,
+ "id": 4461,
+ "area": 2420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2255,
+ "bbox": [
+ 428,
+ 211,
+ 66,
+ 54
+ ],
+ "category_id": 4,
+ "id": 4462,
+ "area": 3564,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2256,
+ "bbox": [
+ 365,
+ 133,
+ 16,
+ 27
+ ],
+ "category_id": 4,
+ "id": 4463,
+ "area": 432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2256,
+ "bbox": [
+ 247,
+ 195,
+ 18,
+ 30
+ ],
+ "category_id": 4,
+ "id": 4464,
+ "area": 540,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2256,
+ "bbox": [
+ 136,
+ 262,
+ 18,
+ 25
+ ],
+ "category_id": 4,
+ "id": 4465,
+ "area": 450,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2256,
+ "bbox": [
+ 246,
+ 343,
+ 26,
+ 30
+ ],
+ "category_id": 4,
+ "id": 4466,
+ "area": 780,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2256,
+ "bbox": [
+ 67,
+ 445,
+ 21,
+ 26
+ ],
+ "category_id": 4,
+ "id": 4467,
+ "area": 546,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2257,
+ "bbox": [
+ 192,
+ 92,
+ 22,
+ 55
+ ],
+ "category_id": 16,
+ "id": 4468,
+ "area": 1210,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2258,
+ "bbox": [
+ 387,
+ 0,
+ 80,
+ 60
+ ],
+ "category_id": 15,
+ "id": 4469,
+ "area": 4800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2258,
+ "bbox": [
+ 195,
+ 216,
+ 114,
+ 98
+ ],
+ "category_id": 15,
+ "id": 4470,
+ "area": 11172,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2258,
+ "bbox": [
+ 314,
+ 246,
+ 75,
+ 67
+ ],
+ "category_id": 15,
+ "id": 4471,
+ "area": 5025,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2258,
+ "bbox": [
+ 490,
+ 37,
+ 22,
+ 23
+ ],
+ "category_id": 15,
+ "id": 4472,
+ "area": 506,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2259,
+ "bbox": [
+ 330,
+ 86,
+ 64,
+ 334
+ ],
+ "category_id": 10,
+ "id": 4473,
+ "area": 21376,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2260,
+ "bbox": [
+ 17,
+ 138,
+ 62,
+ 229
+ ],
+ "category_id": 10,
+ "id": 4474,
+ "area": 14198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2261,
+ "bbox": [
+ 93,
+ 259,
+ 202,
+ 153
+ ],
+ "category_id": 16,
+ "id": 4475,
+ "area": 30906,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2262,
+ "bbox": [
+ 72,
+ 198,
+ 69,
+ 42
+ ],
+ "category_id": 4,
+ "id": 4476,
+ "area": 2898,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2262,
+ "bbox": [
+ 385,
+ 212,
+ 68,
+ 31
+ ],
+ "category_id": 4,
+ "id": 4477,
+ "area": 2108,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2263,
+ "bbox": [
+ 21,
+ 342,
+ 16,
+ 15
+ ],
+ "category_id": 4,
+ "id": 4478,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2263,
+ "bbox": [
+ 74,
+ 278,
+ 19,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4479,
+ "area": 266,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2263,
+ "bbox": [
+ 240,
+ 252,
+ 16,
+ 17
+ ],
+ "category_id": 4,
+ "id": 4480,
+ "area": 272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2263,
+ "bbox": [
+ 343,
+ 296,
+ 18,
+ 15
+ ],
+ "category_id": 4,
+ "id": 4481,
+ "area": 270,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2263,
+ "bbox": [
+ 349,
+ 202,
+ 18,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4482,
+ "area": 288,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2263,
+ "bbox": [
+ 462,
+ 321,
+ 15,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4483,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2263,
+ "bbox": [
+ 421,
+ 464,
+ 11,
+ 15
+ ],
+ "category_id": 4,
+ "id": 4484,
+ "area": 165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2264,
+ "bbox": [
+ 71,
+ 85,
+ 58,
+ 50
+ ],
+ "category_id": 4,
+ "id": 4485,
+ "area": 2900,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2264,
+ "bbox": [
+ 131,
+ 412,
+ 57,
+ 48
+ ],
+ "category_id": 4,
+ "id": 4486,
+ "area": 2736,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2264,
+ "bbox": [
+ 406,
+ 312,
+ 50,
+ 41
+ ],
+ "category_id": 4,
+ "id": 4487,
+ "area": 2050,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2265,
+ "bbox": [
+ 126,
+ 67,
+ 116,
+ 103
+ ],
+ "category_id": 15,
+ "id": 4488,
+ "area": 11948,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2265,
+ "bbox": [
+ 56,
+ 174,
+ 69,
+ 32
+ ],
+ "category_id": 15,
+ "id": 4489,
+ "area": 2208,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2266,
+ "bbox": [
+ 262,
+ 78,
+ 125,
+ 322
+ ],
+ "category_id": 10,
+ "id": 4490,
+ "area": 40250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2267,
+ "bbox": [
+ 235,
+ 93,
+ 60,
+ 154
+ ],
+ "category_id": 15,
+ "id": 4491,
+ "area": 9240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2267,
+ "bbox": [
+ 108,
+ 179,
+ 209,
+ 234
+ ],
+ "category_id": 15,
+ "id": 4492,
+ "area": 48906,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2268,
+ "bbox": [
+ 293,
+ 112,
+ 65,
+ 52
+ ],
+ "category_id": 4,
+ "id": 4493,
+ "area": 3380,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2268,
+ "bbox": [
+ 126,
+ 344,
+ 55,
+ 52
+ ],
+ "category_id": 4,
+ "id": 4494,
+ "area": 2860,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2268,
+ "bbox": [
+ 371,
+ 433,
+ 65,
+ 51
+ ],
+ "category_id": 4,
+ "id": 4495,
+ "area": 3315,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2269,
+ "bbox": [
+ 83,
+ 186,
+ 402,
+ 73
+ ],
+ "category_id": 10,
+ "id": 4496,
+ "area": 29346,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2270,
+ "bbox": [
+ 161,
+ 63,
+ 68,
+ 50
+ ],
+ "category_id": 4,
+ "id": 4497,
+ "area": 3400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2270,
+ "bbox": [
+ 120,
+ 397,
+ 64,
+ 61
+ ],
+ "category_id": 4,
+ "id": 4498,
+ "area": 3904,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2271,
+ "bbox": [
+ 148,
+ 179,
+ 111,
+ 218
+ ],
+ "category_id": 16,
+ "id": 4499,
+ "area": 24198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2272,
+ "bbox": [
+ 124,
+ 191,
+ 43,
+ 55
+ ],
+ "category_id": 4,
+ "id": 4500,
+ "area": 2365,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2272,
+ "bbox": [
+ 424,
+ 374,
+ 43,
+ 46
+ ],
+ "category_id": 4,
+ "id": 4501,
+ "area": 1978,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2273,
+ "bbox": [
+ 271,
+ 78,
+ 196,
+ 318
+ ],
+ "category_id": 10,
+ "id": 4502,
+ "area": 62328,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2273,
+ "bbox": [
+ 117,
+ 160,
+ 119,
+ 54
+ ],
+ "category_id": 16,
+ "id": 4503,
+ "area": 6426,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2274,
+ "bbox": [
+ 99,
+ 88,
+ 324,
+ 386
+ ],
+ "category_id": 10,
+ "id": 4504,
+ "area": 125064,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2275,
+ "bbox": [
+ 139,
+ 228,
+ 250,
+ 37
+ ],
+ "category_id": 19,
+ "id": 4505,
+ "area": 9250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2276,
+ "bbox": [
+ 73,
+ 245,
+ 360,
+ 82
+ ],
+ "category_id": 16,
+ "id": 4506,
+ "area": 29520,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2277,
+ "bbox": [
+ 56,
+ 164,
+ 424,
+ 156
+ ],
+ "category_id": 10,
+ "id": 4507,
+ "area": 66144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2278,
+ "bbox": [
+ 209,
+ 267,
+ 69,
+ 87
+ ],
+ "category_id": 4,
+ "id": 4508,
+ "area": 6003,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2279,
+ "bbox": [
+ 134,
+ 124,
+ 123,
+ 212
+ ],
+ "category_id": 19,
+ "id": 4509,
+ "area": 26076,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2280,
+ "bbox": [
+ 42,
+ 185,
+ 429,
+ 94
+ ],
+ "category_id": 10,
+ "id": 4510,
+ "area": 40326,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2281,
+ "bbox": [
+ 65,
+ 216,
+ 167,
+ 114
+ ],
+ "category_id": 15,
+ "id": 4511,
+ "area": 19038,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2282,
+ "bbox": [
+ 214,
+ 115,
+ 134,
+ 313
+ ],
+ "category_id": 19,
+ "id": 4512,
+ "area": 41942,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2283,
+ "bbox": [
+ 266,
+ 112,
+ 105,
+ 131
+ ],
+ "category_id": 15,
+ "id": 4513,
+ "area": 13755,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2283,
+ "bbox": [
+ 211,
+ 224,
+ 103,
+ 131
+ ],
+ "category_id": 15,
+ "id": 4514,
+ "area": 13493,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2284,
+ "bbox": [
+ 262,
+ 240,
+ 71,
+ 94
+ ],
+ "category_id": 15,
+ "id": 4515,
+ "area": 6674,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2285,
+ "bbox": [
+ 172,
+ 89,
+ 259,
+ 352
+ ],
+ "category_id": 10,
+ "id": 4516,
+ "area": 91168,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2286,
+ "bbox": [
+ 235,
+ 130,
+ 126,
+ 214
+ ],
+ "category_id": 19,
+ "id": 4517,
+ "area": 26964,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2287,
+ "bbox": [
+ 344,
+ 297,
+ 127,
+ 171
+ ],
+ "category_id": 10,
+ "id": 4518,
+ "area": 21717,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2288,
+ "bbox": [
+ 33,
+ 99,
+ 457,
+ 130
+ ],
+ "category_id": 10,
+ "id": 4519,
+ "area": 59410,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2289,
+ "bbox": [
+ 163,
+ 165,
+ 92,
+ 154
+ ],
+ "category_id": 16,
+ "id": 4520,
+ "area": 14168,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2290,
+ "bbox": [
+ 78,
+ 330,
+ 50,
+ 57
+ ],
+ "category_id": 4,
+ "id": 4521,
+ "area": 2850,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2290,
+ "bbox": [
+ 411,
+ 169,
+ 46,
+ 64
+ ],
+ "category_id": 4,
+ "id": 4522,
+ "area": 2944,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2291,
+ "bbox": [
+ 220,
+ 208,
+ 64,
+ 60
+ ],
+ "category_id": 4,
+ "id": 4523,
+ "area": 3840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2292,
+ "bbox": [
+ 174,
+ 120,
+ 50,
+ 60
+ ],
+ "category_id": 4,
+ "id": 4524,
+ "area": 3000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2292,
+ "bbox": [
+ 149,
+ 270,
+ 79,
+ 77
+ ],
+ "category_id": 4,
+ "id": 4525,
+ "area": 6083,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2292,
+ "bbox": [
+ 374,
+ 389,
+ 56,
+ 46
+ ],
+ "category_id": 4,
+ "id": 4526,
+ "area": 2576,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2293,
+ "bbox": [
+ 98,
+ 65,
+ 44,
+ 41
+ ],
+ "category_id": 4,
+ "id": 4527,
+ "area": 1804,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2293,
+ "bbox": [
+ 343,
+ 293,
+ 27,
+ 46
+ ],
+ "category_id": 4,
+ "id": 4528,
+ "area": 1242,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2293,
+ "bbox": [
+ 186,
+ 424,
+ 30,
+ 43
+ ],
+ "category_id": 4,
+ "id": 4529,
+ "area": 1290,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2294,
+ "bbox": [
+ 280,
+ 176,
+ 187,
+ 196
+ ],
+ "category_id": 15,
+ "id": 4530,
+ "area": 36652,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2295,
+ "bbox": [
+ 164,
+ 136,
+ 49,
+ 205
+ ],
+ "category_id": 19,
+ "id": 4531,
+ "area": 10045,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2296,
+ "bbox": [
+ 83,
+ 151,
+ 42,
+ 44
+ ],
+ "category_id": 4,
+ "id": 4532,
+ "area": 1848,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2296,
+ "bbox": [
+ 450,
+ 353,
+ 51,
+ 41
+ ],
+ "category_id": 4,
+ "id": 4533,
+ "area": 2091,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2297,
+ "bbox": [
+ 67,
+ 128,
+ 29,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4534,
+ "area": 464,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2297,
+ "bbox": [
+ 254,
+ 150,
+ 27,
+ 18
+ ],
+ "category_id": 4,
+ "id": 4535,
+ "area": 486,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2297,
+ "bbox": [
+ 402,
+ 49,
+ 30,
+ 17
+ ],
+ "category_id": 4,
+ "id": 4536,
+ "area": 510,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2297,
+ "bbox": [
+ 51,
+ 298,
+ 29,
+ 18
+ ],
+ "category_id": 4,
+ "id": 4537,
+ "area": 522,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2297,
+ "bbox": [
+ 164,
+ 376,
+ 29,
+ 19
+ ],
+ "category_id": 4,
+ "id": 4538,
+ "area": 551,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2298,
+ "bbox": [
+ 289,
+ 61,
+ 84,
+ 347
+ ],
+ "category_id": 10,
+ "id": 4539,
+ "area": 29148,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2299,
+ "bbox": [
+ 229,
+ 76,
+ 226,
+ 394
+ ],
+ "category_id": 10,
+ "id": 4540,
+ "area": 89044,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2300,
+ "bbox": [
+ 188,
+ 151,
+ 110,
+ 209
+ ],
+ "category_id": 19,
+ "id": 4541,
+ "area": 22990,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2301,
+ "bbox": [
+ 101,
+ 398,
+ 32,
+ 48
+ ],
+ "category_id": 4,
+ "id": 4542,
+ "area": 1536,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2301,
+ "bbox": [
+ 374,
+ 339,
+ 33,
+ 46
+ ],
+ "category_id": 4,
+ "id": 4543,
+ "area": 1518,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2301,
+ "bbox": [
+ 254,
+ 40,
+ 24,
+ 46
+ ],
+ "category_id": 4,
+ "id": 4544,
+ "area": 1104,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2302,
+ "bbox": [
+ 215,
+ 41,
+ 203,
+ 387
+ ],
+ "category_id": 10,
+ "id": 4545,
+ "area": 78561,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2303,
+ "bbox": [
+ 124,
+ 3,
+ 22,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4546,
+ "area": 220,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2303,
+ "bbox": [
+ 33,
+ 48,
+ 18,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4547,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2303,
+ "bbox": [
+ 359,
+ 96,
+ 21,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4548,
+ "area": 210,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2303,
+ "bbox": [
+ 301,
+ 192,
+ 15,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4549,
+ "area": 105,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2303,
+ "bbox": [
+ 211,
+ 234,
+ 28,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4550,
+ "area": 448,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2303,
+ "bbox": [
+ 97,
+ 225,
+ 23,
+ 18
+ ],
+ "category_id": 4,
+ "id": 4551,
+ "area": 414,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2303,
+ "bbox": [
+ 71,
+ 303,
+ 23,
+ 20
+ ],
+ "category_id": 4,
+ "id": 4552,
+ "area": 460,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2303,
+ "bbox": [
+ 432,
+ 419,
+ 19,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4553,
+ "area": 209,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2303,
+ "bbox": [
+ 335,
+ 407,
+ 18,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4554,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2304,
+ "bbox": [
+ 58,
+ 181,
+ 431,
+ 53
+ ],
+ "category_id": 10,
+ "id": 4555,
+ "area": 22843,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2305,
+ "bbox": [
+ 370,
+ 210,
+ 111,
+ 218
+ ],
+ "category_id": 10,
+ "id": 4556,
+ "area": 24198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2306,
+ "bbox": [
+ 108,
+ 156,
+ 174,
+ 187
+ ],
+ "category_id": 19,
+ "id": 4557,
+ "area": 32538,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2307,
+ "bbox": [
+ 170,
+ 267,
+ 120,
+ 104
+ ],
+ "category_id": 15,
+ "id": 4558,
+ "area": 12480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2307,
+ "bbox": [
+ 308,
+ 263,
+ 122,
+ 111
+ ],
+ "category_id": 15,
+ "id": 4559,
+ "area": 13542,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2308,
+ "bbox": [
+ 224,
+ 336,
+ 34,
+ 25
+ ],
+ "category_id": 15,
+ "id": 4560,
+ "area": 850,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2309,
+ "bbox": [
+ 282,
+ 131,
+ 98,
+ 87
+ ],
+ "category_id": 15,
+ "id": 4561,
+ "area": 8526,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2309,
+ "bbox": [
+ 168,
+ 163,
+ 100,
+ 84
+ ],
+ "category_id": 15,
+ "id": 4562,
+ "area": 8400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2309,
+ "bbox": [
+ 182,
+ 295,
+ 108,
+ 89
+ ],
+ "category_id": 15,
+ "id": 4563,
+ "area": 9612,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2309,
+ "bbox": [
+ 314,
+ 250,
+ 118,
+ 96
+ ],
+ "category_id": 15,
+ "id": 4564,
+ "area": 11328,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2310,
+ "bbox": [
+ 122,
+ 117,
+ 299,
+ 264
+ ],
+ "category_id": 10,
+ "id": 4565,
+ "area": 78936,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2311,
+ "bbox": [
+ 96,
+ 136,
+ 296,
+ 120
+ ],
+ "category_id": 16,
+ "id": 4566,
+ "area": 35520,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2312,
+ "bbox": [
+ 420,
+ 87,
+ 16,
+ 21
+ ],
+ "category_id": 4,
+ "id": 4567,
+ "area": 336,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2312,
+ "bbox": [
+ 293,
+ 124,
+ 17,
+ 20
+ ],
+ "category_id": 4,
+ "id": 4568,
+ "area": 340,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2312,
+ "bbox": [
+ 129,
+ 267,
+ 13,
+ 22
+ ],
+ "category_id": 4,
+ "id": 4569,
+ "area": 286,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2312,
+ "bbox": [
+ 31,
+ 301,
+ 18,
+ 23
+ ],
+ "category_id": 4,
+ "id": 4570,
+ "area": 414,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2312,
+ "bbox": [
+ 347,
+ 447,
+ 13,
+ 23
+ ],
+ "category_id": 4,
+ "id": 4571,
+ "area": 299,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2313,
+ "bbox": [
+ 212,
+ 117,
+ 52,
+ 50
+ ],
+ "category_id": 4,
+ "id": 4572,
+ "area": 2600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2313,
+ "bbox": [
+ 257,
+ 427,
+ 53,
+ 45
+ ],
+ "category_id": 4,
+ "id": 4573,
+ "area": 2385,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2314,
+ "bbox": [
+ 52,
+ 117,
+ 68,
+ 92
+ ],
+ "category_id": 19,
+ "id": 4574,
+ "area": 6256,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2315,
+ "bbox": [
+ 79,
+ 141,
+ 345,
+ 287
+ ],
+ "category_id": 10,
+ "id": 4575,
+ "area": 99015,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2316,
+ "bbox": [
+ 108,
+ 103,
+ 32,
+ 26
+ ],
+ "category_id": 4,
+ "id": 4576,
+ "area": 832,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2316,
+ "bbox": [
+ 151,
+ 421,
+ 33,
+ 43
+ ],
+ "category_id": 4,
+ "id": 4577,
+ "area": 1419,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2316,
+ "bbox": [
+ 386,
+ 225,
+ 33,
+ 27
+ ],
+ "category_id": 4,
+ "id": 4578,
+ "area": 891,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2317,
+ "bbox": [
+ 444,
+ 46,
+ 10,
+ 19
+ ],
+ "category_id": 4,
+ "id": 4579,
+ "area": 190,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2317,
+ "bbox": [
+ 309,
+ 115,
+ 9,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4580,
+ "area": 126,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2317,
+ "bbox": [
+ 166,
+ 134,
+ 8,
+ 15
+ ],
+ "category_id": 4,
+ "id": 4581,
+ "area": 120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2317,
+ "bbox": [
+ 36,
+ 224,
+ 12,
+ 17
+ ],
+ "category_id": 4,
+ "id": 4582,
+ "area": 204,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2318,
+ "bbox": [
+ 282,
+ 95,
+ 77,
+ 363
+ ],
+ "category_id": 10,
+ "id": 4583,
+ "area": 27951,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2319,
+ "bbox": [
+ 6,
+ 428,
+ 7,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4584,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2319,
+ "bbox": [
+ 74,
+ 459,
+ 7,
+ 12
+ ],
+ "category_id": 4,
+ "id": 4585,
+ "area": 84,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2319,
+ "bbox": [
+ 305,
+ 335,
+ 7,
+ 12
+ ],
+ "category_id": 4,
+ "id": 4586,
+ "area": 84,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2319,
+ "bbox": [
+ 222,
+ 268,
+ 10,
+ 12
+ ],
+ "category_id": 4,
+ "id": 4587,
+ "area": 120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2319,
+ "bbox": [
+ 176,
+ 314,
+ 7,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4588,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2319,
+ "bbox": [
+ 144,
+ 386,
+ 6,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4589,
+ "area": 66,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2319,
+ "bbox": [
+ 101,
+ 272,
+ 5,
+ 15
+ ],
+ "category_id": 4,
+ "id": 4590,
+ "area": 75,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2319,
+ "bbox": [
+ 132,
+ 187,
+ 5,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4591,
+ "area": 50,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2319,
+ "bbox": [
+ 40,
+ 119,
+ 8,
+ 12
+ ],
+ "category_id": 4,
+ "id": 4592,
+ "area": 96,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2319,
+ "bbox": [
+ 24,
+ 55,
+ 9,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4593,
+ "area": 90,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2319,
+ "bbox": [
+ 332,
+ 51,
+ 9,
+ 15
+ ],
+ "category_id": 4,
+ "id": 4594,
+ "area": 135,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2319,
+ "bbox": [
+ 342,
+ 172,
+ 8,
+ 18
+ ],
+ "category_id": 4,
+ "id": 4595,
+ "area": 144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2320,
+ "bbox": [
+ 237,
+ 297,
+ 50,
+ 84
+ ],
+ "category_id": 15,
+ "id": 4596,
+ "area": 4200,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2321,
+ "bbox": [
+ 326,
+ 49,
+ 9,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4597,
+ "area": 81,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2321,
+ "bbox": [
+ 353,
+ 243,
+ 4,
+ 5
+ ],
+ "category_id": 4,
+ "id": 4598,
+ "area": 20,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2321,
+ "bbox": [
+ 364,
+ 286,
+ 5,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4599,
+ "area": 35,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2321,
+ "bbox": [
+ 318,
+ 311,
+ 6,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4600,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2321,
+ "bbox": [
+ 267,
+ 339,
+ 8,
+ 13
+ ],
+ "category_id": 4,
+ "id": 4601,
+ "area": 104,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2321,
+ "bbox": [
+ 215,
+ 333,
+ 4,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4602,
+ "area": 40,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2322,
+ "bbox": [
+ 333,
+ 397,
+ 13,
+ 45
+ ],
+ "category_id": 4,
+ "id": 4603,
+ "area": 585,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2322,
+ "bbox": [
+ 134,
+ 81,
+ 11,
+ 44
+ ],
+ "category_id": 4,
+ "id": 4604,
+ "area": 484,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2323,
+ "bbox": [
+ 79,
+ 193,
+ 64,
+ 35
+ ],
+ "category_id": 4,
+ "id": 4605,
+ "area": 2240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2323,
+ "bbox": [
+ 373,
+ 152,
+ 51,
+ 41
+ ],
+ "category_id": 4,
+ "id": 4606,
+ "area": 2091,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2323,
+ "bbox": [
+ 282,
+ 389,
+ 67,
+ 39
+ ],
+ "category_id": 4,
+ "id": 4607,
+ "area": 2613,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2324,
+ "bbox": [
+ 160,
+ 60,
+ 21,
+ 14
+ ],
+ "category_id": 15,
+ "id": 4608,
+ "area": 294,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2324,
+ "bbox": [
+ 208,
+ 46,
+ 19,
+ 14
+ ],
+ "category_id": 15,
+ "id": 4609,
+ "area": 266,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2325,
+ "bbox": [
+ 271,
+ 211,
+ 24,
+ 110
+ ],
+ "category_id": 4,
+ "id": 4610,
+ "area": 2640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2326,
+ "bbox": [
+ 346,
+ 320,
+ 45,
+ 51
+ ],
+ "category_id": 16,
+ "id": 4611,
+ "area": 2295,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2327,
+ "bbox": [
+ 213,
+ 166,
+ 48,
+ 224
+ ],
+ "category_id": 19,
+ "id": 4612,
+ "area": 10752,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2328,
+ "bbox": [
+ 90,
+ 90,
+ 25,
+ 28
+ ],
+ "category_id": 4,
+ "id": 4613,
+ "area": 700,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2328,
+ "bbox": [
+ 234,
+ 336,
+ 27,
+ 49
+ ],
+ "category_id": 4,
+ "id": 4614,
+ "area": 1323,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2328,
+ "bbox": [
+ 423,
+ 127,
+ 33,
+ 45
+ ],
+ "category_id": 4,
+ "id": 4615,
+ "area": 1485,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2329,
+ "bbox": [
+ 287,
+ 90,
+ 51,
+ 79
+ ],
+ "category_id": 4,
+ "id": 4616,
+ "area": 4029,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2330,
+ "bbox": [
+ 124,
+ 74,
+ 245,
+ 372
+ ],
+ "category_id": 10,
+ "id": 4617,
+ "area": 91140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2331,
+ "bbox": [
+ 192,
+ 133,
+ 223,
+ 248
+ ],
+ "category_id": 10,
+ "id": 4618,
+ "area": 55304,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2332,
+ "bbox": [
+ 71,
+ 35,
+ 154,
+ 206
+ ],
+ "category_id": 10,
+ "id": 4619,
+ "area": 31724,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2333,
+ "bbox": [
+ 97,
+ 161,
+ 283,
+ 198
+ ],
+ "category_id": 10,
+ "id": 4620,
+ "area": 56034,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2334,
+ "bbox": [
+ 101,
+ 73,
+ 43,
+ 51
+ ],
+ "category_id": 4,
+ "id": 4621,
+ "area": 2193,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2334,
+ "bbox": [
+ 5,
+ 371,
+ 43,
+ 55
+ ],
+ "category_id": 4,
+ "id": 4622,
+ "area": 2365,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2334,
+ "bbox": [
+ 463,
+ 145,
+ 35,
+ 48
+ ],
+ "category_id": 4,
+ "id": 4623,
+ "area": 1680,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2335,
+ "bbox": [
+ 366,
+ 408,
+ 59,
+ 69
+ ],
+ "category_id": 19,
+ "id": 4624,
+ "area": 4071,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2336,
+ "bbox": [
+ 154,
+ 197,
+ 132,
+ 140
+ ],
+ "category_id": 10,
+ "id": 4625,
+ "area": 18480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2337,
+ "bbox": [
+ 152,
+ 254,
+ 203,
+ 45
+ ],
+ "category_id": 19,
+ "id": 4626,
+ "area": 9135,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2338,
+ "bbox": [
+ 131,
+ 273,
+ 76,
+ 99
+ ],
+ "category_id": 16,
+ "id": 4627,
+ "area": 7524,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2338,
+ "bbox": [
+ 238,
+ 130,
+ 199,
+ 126
+ ],
+ "category_id": 16,
+ "id": 4628,
+ "area": 25074,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2339,
+ "bbox": [
+ 222,
+ 111,
+ 110,
+ 272
+ ],
+ "category_id": 16,
+ "id": 4629,
+ "area": 29920,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2340,
+ "bbox": [
+ 280,
+ 45,
+ 43,
+ 67
+ ],
+ "category_id": 15,
+ "id": 4630,
+ "area": 2881,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2340,
+ "bbox": [
+ 198,
+ 188,
+ 140,
+ 176
+ ],
+ "category_id": 15,
+ "id": 4631,
+ "area": 24640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2341,
+ "bbox": [
+ 101,
+ 204,
+ 130,
+ 100
+ ],
+ "category_id": 15,
+ "id": 4632,
+ "area": 13000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2342,
+ "bbox": [
+ 127,
+ 112,
+ 271,
+ 280
+ ],
+ "category_id": 16,
+ "id": 4633,
+ "area": 75880,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2343,
+ "bbox": [
+ 40,
+ 332,
+ 447,
+ 116
+ ],
+ "category_id": 10,
+ "id": 4634,
+ "area": 51852,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2344,
+ "bbox": [
+ 195,
+ 77,
+ 155,
+ 381
+ ],
+ "category_id": 10,
+ "id": 4635,
+ "area": 59055,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2345,
+ "bbox": [
+ 275,
+ 78,
+ 144,
+ 368
+ ],
+ "category_id": 19,
+ "id": 4636,
+ "area": 52992,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2345,
+ "bbox": [
+ 36,
+ 126,
+ 44,
+ 163
+ ],
+ "category_id": 19,
+ "id": 4637,
+ "area": 7172,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2346,
+ "bbox": [
+ 57,
+ 151,
+ 425,
+ 282
+ ],
+ "category_id": 16,
+ "id": 4638,
+ "area": 119850,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2347,
+ "bbox": [
+ 208,
+ 258,
+ 110,
+ 126
+ ],
+ "category_id": 16,
+ "id": 4639,
+ "area": 13860,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2348,
+ "bbox": [
+ 65,
+ 252,
+ 433,
+ 174
+ ],
+ "category_id": 10,
+ "id": 4640,
+ "area": 75342,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2349,
+ "bbox": [
+ 122,
+ 103,
+ 12,
+ 17
+ ],
+ "category_id": 4,
+ "id": 4641,
+ "area": 204,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2349,
+ "bbox": [
+ 471,
+ 159,
+ 8,
+ 15
+ ],
+ "category_id": 4,
+ "id": 4642,
+ "area": 120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2349,
+ "bbox": [
+ 97,
+ 423,
+ 13,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4643,
+ "area": 208,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2349,
+ "bbox": [
+ 268,
+ 374,
+ 10,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4644,
+ "area": 160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2350,
+ "bbox": [
+ 378,
+ 350,
+ 11,
+ 21
+ ],
+ "category_id": 15,
+ "id": 4645,
+ "area": 231,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2351,
+ "bbox": [
+ 42,
+ 77,
+ 151,
+ 45
+ ],
+ "category_id": 10,
+ "id": 4646,
+ "area": 6795,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2352,
+ "bbox": [
+ 37,
+ 133,
+ 58,
+ 23
+ ],
+ "category_id": 4,
+ "id": 4647,
+ "area": 1334,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2352,
+ "bbox": [
+ 312,
+ 369,
+ 73,
+ 53
+ ],
+ "category_id": 4,
+ "id": 4648,
+ "area": 3869,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2352,
+ "bbox": [
+ 408,
+ 184,
+ 70,
+ 25
+ ],
+ "category_id": 4,
+ "id": 4649,
+ "area": 1750,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2353,
+ "bbox": [
+ 212,
+ 144,
+ 83,
+ 100
+ ],
+ "category_id": 15,
+ "id": 4650,
+ "area": 8300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2353,
+ "bbox": [
+ 135,
+ 231,
+ 76,
+ 89
+ ],
+ "category_id": 15,
+ "id": 4651,
+ "area": 6764,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2354,
+ "bbox": [
+ 127,
+ 272,
+ 279,
+ 108
+ ],
+ "category_id": 16,
+ "id": 4652,
+ "area": 30132,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2355,
+ "bbox": [
+ 363,
+ 242,
+ 43,
+ 58
+ ],
+ "category_id": 15,
+ "id": 4653,
+ "area": 2494,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2355,
+ "bbox": [
+ 189,
+ 176,
+ 68,
+ 78
+ ],
+ "category_id": 15,
+ "id": 4654,
+ "area": 5304,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2356,
+ "bbox": [
+ 177,
+ 442,
+ 231,
+ 39
+ ],
+ "category_id": 10,
+ "id": 4655,
+ "area": 9009,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2357,
+ "bbox": [
+ 233,
+ 327,
+ 150,
+ 133
+ ],
+ "category_id": 15,
+ "id": 4656,
+ "area": 19950,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2357,
+ "bbox": [
+ 92,
+ 210,
+ 148,
+ 133
+ ],
+ "category_id": 15,
+ "id": 4657,
+ "area": 19684,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2357,
+ "bbox": [
+ 192,
+ 81,
+ 148,
+ 136
+ ],
+ "category_id": 15,
+ "id": 4658,
+ "area": 20128,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2357,
+ "bbox": [
+ 334,
+ 200,
+ 148,
+ 132
+ ],
+ "category_id": 15,
+ "id": 4659,
+ "area": 19536,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2358,
+ "bbox": [
+ 108,
+ 231,
+ 308,
+ 49
+ ],
+ "category_id": 19,
+ "id": 4660,
+ "area": 15092,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2359,
+ "bbox": [
+ 159,
+ 113,
+ 231,
+ 299
+ ],
+ "category_id": 10,
+ "id": 4661,
+ "area": 69069,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2360,
+ "bbox": [
+ 205,
+ 169,
+ 75,
+ 291
+ ],
+ "category_id": 16,
+ "id": 4662,
+ "area": 21825,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2361,
+ "bbox": [
+ 64,
+ 213,
+ 89,
+ 103
+ ],
+ "category_id": 15,
+ "id": 4663,
+ "area": 9167,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2361,
+ "bbox": [
+ 176,
+ 208,
+ 82,
+ 98
+ ],
+ "category_id": 15,
+ "id": 4664,
+ "area": 8036,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2361,
+ "bbox": [
+ 287,
+ 204,
+ 83,
+ 101
+ ],
+ "category_id": 15,
+ "id": 4665,
+ "area": 8383,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2362,
+ "bbox": [
+ 300,
+ 453,
+ 172,
+ 24
+ ],
+ "category_id": 10,
+ "id": 4666,
+ "area": 4128,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2362,
+ "bbox": [
+ 160,
+ 152,
+ 26,
+ 13
+ ],
+ "category_id": 16,
+ "id": 4667,
+ "area": 338,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2363,
+ "bbox": [
+ 192,
+ 17,
+ 65,
+ 398
+ ],
+ "category_id": 10,
+ "id": 4668,
+ "area": 25870,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2364,
+ "bbox": [
+ 207,
+ 113,
+ 132,
+ 268
+ ],
+ "category_id": 19,
+ "id": 4669,
+ "area": 35376,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2365,
+ "bbox": [
+ 80,
+ 225,
+ 385,
+ 103
+ ],
+ "category_id": 10,
+ "id": 4670,
+ "area": 39655,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2366,
+ "bbox": [
+ 152,
+ 230,
+ 56,
+ 53
+ ],
+ "category_id": 10,
+ "id": 4671,
+ "area": 2968,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2367,
+ "bbox": [
+ 205,
+ 147,
+ 77,
+ 158
+ ],
+ "category_id": 10,
+ "id": 4672,
+ "area": 12166,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2368,
+ "bbox": [
+ 236,
+ 239,
+ 41,
+ 121
+ ],
+ "category_id": 4,
+ "id": 4673,
+ "area": 4961,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2369,
+ "bbox": [
+ 252,
+ 184,
+ 44,
+ 130
+ ],
+ "category_id": 19,
+ "id": 4674,
+ "area": 5720,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2370,
+ "bbox": [
+ 136,
+ 214,
+ 246,
+ 166
+ ],
+ "category_id": 16,
+ "id": 4675,
+ "area": 40836,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2371,
+ "bbox": [
+ 177,
+ 151,
+ 115,
+ 177
+ ],
+ "category_id": 10,
+ "id": 4676,
+ "area": 20355,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2372,
+ "bbox": [
+ 71,
+ 107,
+ 398,
+ 251
+ ],
+ "category_id": 10,
+ "id": 4677,
+ "area": 99898,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2373,
+ "bbox": [
+ 169,
+ 68,
+ 19,
+ 19
+ ],
+ "category_id": 4,
+ "id": 4678,
+ "area": 361,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2373,
+ "bbox": [
+ 438,
+ 168,
+ 30,
+ 23
+ ],
+ "category_id": 4,
+ "id": 4679,
+ "area": 690,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2373,
+ "bbox": [
+ 291,
+ 300,
+ 19,
+ 19
+ ],
+ "category_id": 4,
+ "id": 4680,
+ "area": 361,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2373,
+ "bbox": [
+ 97,
+ 412,
+ 20,
+ 26
+ ],
+ "category_id": 4,
+ "id": 4681,
+ "area": 520,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2374,
+ "bbox": [
+ 126,
+ 161,
+ 310,
+ 202
+ ],
+ "category_id": 10,
+ "id": 4682,
+ "area": 62620,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2375,
+ "bbox": [
+ 0,
+ 76,
+ 13,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4683,
+ "area": 130,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2375,
+ "bbox": [
+ 144,
+ 272,
+ 19,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4684,
+ "area": 190,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2375,
+ "bbox": [
+ 466,
+ 382,
+ 15,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4685,
+ "area": 120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2376,
+ "bbox": [
+ 203,
+ 204,
+ 86,
+ 75
+ ],
+ "category_id": 4,
+ "id": 4686,
+ "area": 6450,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2377,
+ "bbox": [
+ 136,
+ 182,
+ 333,
+ 200
+ ],
+ "category_id": 10,
+ "id": 4687,
+ "area": 66600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2378,
+ "bbox": [
+ 220,
+ 118,
+ 172,
+ 203
+ ],
+ "category_id": 16,
+ "id": 4688,
+ "area": 34916,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2379,
+ "bbox": [
+ 72,
+ 223,
+ 25,
+ 104
+ ],
+ "category_id": 16,
+ "id": 4689,
+ "area": 2600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2380,
+ "bbox": [
+ 140,
+ 280,
+ 315,
+ 123
+ ],
+ "category_id": 16,
+ "id": 4690,
+ "area": 38745,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2381,
+ "bbox": [
+ 159,
+ 324,
+ 35,
+ 54
+ ],
+ "category_id": 4,
+ "id": 4691,
+ "area": 1890,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2381,
+ "bbox": [
+ 319,
+ 135,
+ 24,
+ 55
+ ],
+ "category_id": 4,
+ "id": 4692,
+ "area": 1320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2382,
+ "bbox": [
+ 76,
+ 43,
+ 18,
+ 19
+ ],
+ "category_id": 4,
+ "id": 4693,
+ "area": 342,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2382,
+ "bbox": [
+ 87,
+ 208,
+ 14,
+ 21
+ ],
+ "category_id": 4,
+ "id": 4694,
+ "area": 294,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2382,
+ "bbox": [
+ 67,
+ 319,
+ 15,
+ 27
+ ],
+ "category_id": 4,
+ "id": 4695,
+ "area": 405,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2382,
+ "bbox": [
+ 272,
+ 213,
+ 15,
+ 19
+ ],
+ "category_id": 4,
+ "id": 4696,
+ "area": 285,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2382,
+ "bbox": [
+ 209,
+ 408,
+ 13,
+ 25
+ ],
+ "category_id": 4,
+ "id": 4697,
+ "area": 325,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2382,
+ "bbox": [
+ 371,
+ 301,
+ 11,
+ 30
+ ],
+ "category_id": 4,
+ "id": 4698,
+ "area": 330,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2382,
+ "bbox": [
+ 471,
+ 172,
+ 20,
+ 30
+ ],
+ "category_id": 4,
+ "id": 4699,
+ "area": 600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2383,
+ "bbox": [
+ 23,
+ 21,
+ 30,
+ 12
+ ],
+ "category_id": 4,
+ "id": 4700,
+ "area": 360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2383,
+ "bbox": [
+ 39,
+ 170,
+ 27,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4701,
+ "area": 297,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2383,
+ "bbox": [
+ 101,
+ 311,
+ 30,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4702,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2383,
+ "bbox": [
+ 439,
+ 135,
+ 12,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4703,
+ "area": 132,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2383,
+ "bbox": [
+ 462,
+ 314,
+ 12,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4704,
+ "area": 168,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2383,
+ "bbox": [
+ 465,
+ 485,
+ 13,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4705,
+ "area": 182,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2384,
+ "bbox": [
+ 225,
+ 151,
+ 168,
+ 127
+ ],
+ "category_id": 15,
+ "id": 4706,
+ "area": 21336,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2384,
+ "bbox": [
+ 114,
+ 261,
+ 175,
+ 122
+ ],
+ "category_id": 15,
+ "id": 4707,
+ "area": 21350,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2385,
+ "bbox": [
+ 106,
+ 162,
+ 336,
+ 184
+ ],
+ "category_id": 16,
+ "id": 4708,
+ "area": 61824,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2386,
+ "bbox": [
+ 194,
+ 72,
+ 160,
+ 160
+ ],
+ "category_id": 15,
+ "id": 4709,
+ "area": 25600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2386,
+ "bbox": [
+ 146,
+ 290,
+ 161,
+ 163
+ ],
+ "category_id": 15,
+ "id": 4710,
+ "area": 26243,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2387,
+ "bbox": [
+ 85,
+ 124,
+ 249,
+ 192
+ ],
+ "category_id": 10,
+ "id": 4711,
+ "area": 47808,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2388,
+ "bbox": [
+ 204,
+ 72,
+ 215,
+ 388
+ ],
+ "category_id": 10,
+ "id": 4712,
+ "area": 83420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2389,
+ "bbox": [
+ 75,
+ 69,
+ 75,
+ 71
+ ],
+ "category_id": 19,
+ "id": 4713,
+ "area": 5325,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2390,
+ "bbox": [
+ 155,
+ 24,
+ 26,
+ 22
+ ],
+ "category_id": 4,
+ "id": 4714,
+ "area": 572,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2390,
+ "bbox": [
+ 298,
+ 11,
+ 22,
+ 22
+ ],
+ "category_id": 4,
+ "id": 4715,
+ "area": 484,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2390,
+ "bbox": [
+ 455,
+ 91,
+ 23,
+ 24
+ ],
+ "category_id": 4,
+ "id": 4716,
+ "area": 552,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2390,
+ "bbox": [
+ 280,
+ 176,
+ 16,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4717,
+ "area": 176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2390,
+ "bbox": [
+ 248,
+ 261,
+ 18,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4718,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2390,
+ "bbox": [
+ 400,
+ 361,
+ 18,
+ 17
+ ],
+ "category_id": 4,
+ "id": 4719,
+ "area": 306,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2390,
+ "bbox": [
+ 324,
+ 490,
+ 27,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4720,
+ "area": 378,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2391,
+ "bbox": [
+ 103,
+ 192,
+ 348,
+ 214
+ ],
+ "category_id": 10,
+ "id": 4721,
+ "area": 74472,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2392,
+ "bbox": [
+ 300,
+ 106,
+ 126,
+ 157
+ ],
+ "category_id": 16,
+ "id": 4722,
+ "area": 19782,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2393,
+ "bbox": [
+ 128,
+ 340,
+ 95,
+ 86
+ ],
+ "category_id": 15,
+ "id": 4723,
+ "area": 8170,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2393,
+ "bbox": [
+ 215,
+ 280,
+ 129,
+ 120
+ ],
+ "category_id": 15,
+ "id": 4724,
+ "area": 15480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2393,
+ "bbox": [
+ 232,
+ 188,
+ 93,
+ 88
+ ],
+ "category_id": 15,
+ "id": 4725,
+ "area": 8184,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2393,
+ "bbox": [
+ 341,
+ 229,
+ 130,
+ 128
+ ],
+ "category_id": 15,
+ "id": 4726,
+ "area": 16640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2394,
+ "bbox": [
+ 102,
+ 268,
+ 61,
+ 82
+ ],
+ "category_id": 15,
+ "id": 4727,
+ "area": 5002,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2394,
+ "bbox": [
+ 375,
+ 245,
+ 47,
+ 75
+ ],
+ "category_id": 15,
+ "id": 4728,
+ "area": 3525,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2395,
+ "bbox": [
+ 144,
+ 311,
+ 262,
+ 104
+ ],
+ "category_id": 16,
+ "id": 4729,
+ "area": 27248,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2396,
+ "bbox": [
+ 125,
+ 88,
+ 228,
+ 340
+ ],
+ "category_id": 16,
+ "id": 4730,
+ "area": 77520,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2397,
+ "bbox": [
+ 37,
+ 112,
+ 325,
+ 330
+ ],
+ "category_id": 16,
+ "id": 4731,
+ "area": 107250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2398,
+ "bbox": [
+ 109,
+ 441,
+ 182,
+ 42
+ ],
+ "category_id": 10,
+ "id": 4732,
+ "area": 7644,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2399,
+ "bbox": [
+ 237,
+ 105,
+ 65,
+ 312
+ ],
+ "category_id": 19,
+ "id": 4733,
+ "area": 20280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2400,
+ "bbox": [
+ 347,
+ 302,
+ 65,
+ 49
+ ],
+ "category_id": 16,
+ "id": 4734,
+ "area": 3185,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2400,
+ "bbox": [
+ 220,
+ 465,
+ 25,
+ 25
+ ],
+ "category_id": 16,
+ "id": 4735,
+ "area": 625,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2400,
+ "bbox": [
+ 188,
+ 186,
+ 44,
+ 27
+ ],
+ "category_id": 16,
+ "id": 4736,
+ "area": 1188,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2401,
+ "bbox": [
+ 195,
+ 215,
+ 62,
+ 85
+ ],
+ "category_id": 4,
+ "id": 4737,
+ "area": 5270,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2402,
+ "bbox": [
+ 149,
+ 245,
+ 91,
+ 101
+ ],
+ "category_id": 4,
+ "id": 4738,
+ "area": 9191,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2403,
+ "bbox": [
+ 459,
+ 307,
+ 23,
+ 23
+ ],
+ "category_id": 16,
+ "id": 4739,
+ "area": 529,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2404,
+ "bbox": [
+ 92,
+ 141,
+ 55,
+ 56
+ ],
+ "category_id": 15,
+ "id": 4740,
+ "area": 3080,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2404,
+ "bbox": [
+ 115,
+ 57,
+ 171,
+ 168
+ ],
+ "category_id": 15,
+ "id": 4741,
+ "area": 28728,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2404,
+ "bbox": [
+ 296,
+ 95,
+ 170,
+ 173
+ ],
+ "category_id": 15,
+ "id": 4742,
+ "area": 29410,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2404,
+ "bbox": [
+ 152,
+ 216,
+ 208,
+ 206
+ ],
+ "category_id": 15,
+ "id": 4743,
+ "area": 42848,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2405,
+ "bbox": [
+ 101,
+ 120,
+ 4,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4744,
+ "area": 44,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2405,
+ "bbox": [
+ 184,
+ 99,
+ 6,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4745,
+ "area": 54,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2405,
+ "bbox": [
+ 101,
+ 170,
+ 7,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4746,
+ "area": 49,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2405,
+ "bbox": [
+ 226,
+ 200,
+ 10,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4747,
+ "area": 90,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2405,
+ "bbox": [
+ 386,
+ 179,
+ 9,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4748,
+ "area": 90,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2405,
+ "bbox": [
+ 492,
+ 124,
+ 10,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4749,
+ "area": 70,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2405,
+ "bbox": [
+ 476,
+ 62,
+ 5,
+ 6
+ ],
+ "category_id": 4,
+ "id": 4750,
+ "area": 30,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2405,
+ "bbox": [
+ 131,
+ 275,
+ 5,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4751,
+ "area": 35,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2405,
+ "bbox": [
+ 92,
+ 321,
+ 7,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4752,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2405,
+ "bbox": [
+ 112,
+ 361,
+ 4,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4753,
+ "area": 28,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2405,
+ "bbox": [
+ 386,
+ 373,
+ 8,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4754,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2405,
+ "bbox": [
+ 314,
+ 349,
+ 10,
+ 6
+ ],
+ "category_id": 4,
+ "id": 4755,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2405,
+ "bbox": [
+ 254,
+ 376,
+ 3,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4756,
+ "area": 24,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2405,
+ "bbox": [
+ 166,
+ 399,
+ 7,
+ 5
+ ],
+ "category_id": 4,
+ "id": 4757,
+ "area": 35,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2405,
+ "bbox": [
+ 110,
+ 431,
+ 7,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4758,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2406,
+ "bbox": [
+ 236,
+ 116,
+ 227,
+ 230
+ ],
+ "category_id": 16,
+ "id": 4759,
+ "area": 52210,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2407,
+ "bbox": [
+ 145,
+ 69,
+ 20,
+ 13
+ ],
+ "category_id": 4,
+ "id": 4760,
+ "area": 260,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2407,
+ "bbox": [
+ 67,
+ 218,
+ 21,
+ 15
+ ],
+ "category_id": 4,
+ "id": 4761,
+ "area": 315,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2407,
+ "bbox": [
+ 213,
+ 368,
+ 20,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4762,
+ "area": 320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2407,
+ "bbox": [
+ 434,
+ 309,
+ 20,
+ 19
+ ],
+ "category_id": 4,
+ "id": 4763,
+ "area": 380,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2407,
+ "bbox": [
+ 389,
+ 471,
+ 18,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4764,
+ "area": 288,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2408,
+ "bbox": [
+ 248,
+ 389,
+ 53,
+ 56
+ ],
+ "category_id": 4,
+ "id": 4765,
+ "area": 2968,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2408,
+ "bbox": [
+ 203,
+ 76,
+ 57,
+ 44
+ ],
+ "category_id": 4,
+ "id": 4766,
+ "area": 2508,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2409,
+ "bbox": [
+ 20,
+ 37,
+ 51,
+ 153
+ ],
+ "category_id": 10,
+ "id": 4767,
+ "area": 7803,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2410,
+ "bbox": [
+ 340,
+ 93,
+ 77,
+ 336
+ ],
+ "category_id": 19,
+ "id": 4768,
+ "area": 25872,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2411,
+ "bbox": [
+ 183,
+ 177,
+ 11,
+ 6
+ ],
+ "category_id": 4,
+ "id": 4769,
+ "area": 66,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2411,
+ "bbox": [
+ 294,
+ 188,
+ 10,
+ 7
+ ],
+ "category_id": 4,
+ "id": 4770,
+ "area": 70,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2411,
+ "bbox": [
+ 318,
+ 229,
+ 10,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4771,
+ "area": 90,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2411,
+ "bbox": [
+ 378,
+ 234,
+ 11,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4772,
+ "area": 99,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2411,
+ "bbox": [
+ 301,
+ 365,
+ 10,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4773,
+ "area": 80,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2412,
+ "bbox": [
+ 178,
+ 35,
+ 113,
+ 122
+ ],
+ "category_id": 15,
+ "id": 4774,
+ "area": 13786,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2412,
+ "bbox": [
+ 291,
+ 76,
+ 108,
+ 121
+ ],
+ "category_id": 15,
+ "id": 4775,
+ "area": 13068,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2413,
+ "bbox": [
+ 176,
+ 184,
+ 236,
+ 211
+ ],
+ "category_id": 16,
+ "id": 4776,
+ "area": 49796,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2414,
+ "bbox": [
+ 112,
+ 245,
+ 160,
+ 114
+ ],
+ "category_id": 15,
+ "id": 4777,
+ "area": 18240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2414,
+ "bbox": [
+ 282,
+ 328,
+ 158,
+ 114
+ ],
+ "category_id": 15,
+ "id": 4778,
+ "area": 18012,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2415,
+ "bbox": [
+ 134,
+ 368,
+ 35,
+ 38
+ ],
+ "category_id": 4,
+ "id": 4779,
+ "area": 1330,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2415,
+ "bbox": [
+ 408,
+ 152,
+ 31,
+ 36
+ ],
+ "category_id": 4,
+ "id": 4780,
+ "area": 1116,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2416,
+ "bbox": [
+ 8,
+ 204,
+ 488,
+ 163
+ ],
+ "category_id": 19,
+ "id": 4781,
+ "area": 79544,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2417,
+ "bbox": [
+ 479,
+ 87,
+ 25,
+ 17
+ ],
+ "category_id": 4,
+ "id": 4782,
+ "area": 425,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2417,
+ "bbox": [
+ 287,
+ 113,
+ 17,
+ 18
+ ],
+ "category_id": 4,
+ "id": 4783,
+ "area": 306,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2417,
+ "bbox": [
+ 169,
+ 165,
+ 21,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4784,
+ "area": 231,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2417,
+ "bbox": [
+ 34,
+ 233,
+ 21,
+ 8
+ ],
+ "category_id": 4,
+ "id": 4785,
+ "area": 168,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2417,
+ "bbox": [
+ 129,
+ 403,
+ 20,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4786,
+ "area": 280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2417,
+ "bbox": [
+ 432,
+ 428,
+ 25,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4787,
+ "area": 400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2418,
+ "bbox": [
+ 152,
+ 129,
+ 319,
+ 246
+ ],
+ "category_id": 16,
+ "id": 4788,
+ "area": 78474,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2419,
+ "bbox": [
+ 362,
+ 215,
+ 78,
+ 104
+ ],
+ "category_id": 15,
+ "id": 4789,
+ "area": 8112,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2419,
+ "bbox": [
+ 330,
+ 319,
+ 92,
+ 103
+ ],
+ "category_id": 15,
+ "id": 4790,
+ "area": 9476,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2420,
+ "bbox": [
+ 34,
+ 168,
+ 400,
+ 246
+ ],
+ "category_id": 10,
+ "id": 4791,
+ "area": 98400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2421,
+ "bbox": [
+ 207,
+ 141,
+ 165,
+ 160
+ ],
+ "category_id": 16,
+ "id": 4792,
+ "area": 26400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2422,
+ "bbox": [
+ 130,
+ 65,
+ 201,
+ 306
+ ],
+ "category_id": 16,
+ "id": 4793,
+ "area": 61506,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2423,
+ "bbox": [
+ 154,
+ 226,
+ 162,
+ 53
+ ],
+ "category_id": 19,
+ "id": 4794,
+ "area": 8586,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2424,
+ "bbox": [
+ 122,
+ 114,
+ 323,
+ 328
+ ],
+ "category_id": 10,
+ "id": 4795,
+ "area": 105944,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2425,
+ "bbox": [
+ 220,
+ 163,
+ 94,
+ 274
+ ],
+ "category_id": 16,
+ "id": 4796,
+ "area": 25756,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2426,
+ "bbox": [
+ 209,
+ 217,
+ 225,
+ 231
+ ],
+ "category_id": 16,
+ "id": 4797,
+ "area": 51975,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2427,
+ "bbox": [
+ 311,
+ 63,
+ 57,
+ 45
+ ],
+ "category_id": 4,
+ "id": 4798,
+ "area": 2565,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2427,
+ "bbox": [
+ 113,
+ 396,
+ 83,
+ 26
+ ],
+ "category_id": 4,
+ "id": 4799,
+ "area": 2158,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2428,
+ "bbox": [
+ 200,
+ 119,
+ 76,
+ 234
+ ],
+ "category_id": 19,
+ "id": 4800,
+ "area": 17784,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2429,
+ "bbox": [
+ 408,
+ 236,
+ 32,
+ 19
+ ],
+ "category_id": 4,
+ "id": 4801,
+ "area": 608,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2429,
+ "bbox": [
+ 471,
+ 460,
+ 27,
+ 18
+ ],
+ "category_id": 4,
+ "id": 4802,
+ "area": 486,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2430,
+ "bbox": [
+ 417,
+ 380,
+ 68,
+ 100
+ ],
+ "category_id": 19,
+ "id": 4803,
+ "area": 6800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2431,
+ "bbox": [
+ 127,
+ 188,
+ 233,
+ 196
+ ],
+ "category_id": 16,
+ "id": 4804,
+ "area": 45668,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2432,
+ "bbox": [
+ 154,
+ 97,
+ 54,
+ 313
+ ],
+ "category_id": 19,
+ "id": 4805,
+ "area": 16902,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2433,
+ "bbox": [
+ 152,
+ 420,
+ 346,
+ 51
+ ],
+ "category_id": 10,
+ "id": 4806,
+ "area": 17646,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2434,
+ "bbox": [
+ 49,
+ 72,
+ 96,
+ 86
+ ],
+ "category_id": 19,
+ "id": 4807,
+ "area": 8256,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2435,
+ "bbox": [
+ 78,
+ 238,
+ 359,
+ 59
+ ],
+ "category_id": 19,
+ "id": 4808,
+ "area": 21181,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2436,
+ "bbox": [
+ 92,
+ 126,
+ 190,
+ 209
+ ],
+ "category_id": 15,
+ "id": 4809,
+ "area": 39710,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2436,
+ "bbox": [
+ 355,
+ 195,
+ 68,
+ 114
+ ],
+ "category_id": 15,
+ "id": 4810,
+ "area": 7752,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2437,
+ "bbox": [
+ 143,
+ 226,
+ 192,
+ 67
+ ],
+ "category_id": 19,
+ "id": 4811,
+ "area": 12864,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2438,
+ "bbox": [
+ 173,
+ 129,
+ 144,
+ 84
+ ],
+ "category_id": 19,
+ "id": 4812,
+ "area": 12096,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2439,
+ "bbox": [
+ 131,
+ 136,
+ 196,
+ 189
+ ],
+ "category_id": 10,
+ "id": 4813,
+ "area": 37044,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2440,
+ "bbox": [
+ 67,
+ 191,
+ 400,
+ 144
+ ],
+ "category_id": 16,
+ "id": 4814,
+ "area": 57600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2441,
+ "bbox": [
+ 146,
+ 80,
+ 227,
+ 347
+ ],
+ "category_id": 16,
+ "id": 4815,
+ "area": 78769,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2442,
+ "bbox": [
+ 148,
+ 39,
+ 190,
+ 415
+ ],
+ "category_id": 10,
+ "id": 4816,
+ "area": 78850,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2443,
+ "bbox": [
+ 32,
+ 84,
+ 133,
+ 179
+ ],
+ "category_id": 15,
+ "id": 4817,
+ "area": 23807,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2443,
+ "bbox": [
+ 257,
+ 108,
+ 141,
+ 192
+ ],
+ "category_id": 15,
+ "id": 4818,
+ "area": 27072,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2444,
+ "bbox": [
+ 169,
+ 165,
+ 263,
+ 232
+ ],
+ "category_id": 16,
+ "id": 4819,
+ "area": 61016,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2445,
+ "bbox": [
+ 52,
+ 193,
+ 203,
+ 148
+ ],
+ "category_id": 16,
+ "id": 4820,
+ "area": 30044,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2445,
+ "bbox": [
+ 340,
+ 318,
+ 113,
+ 158
+ ],
+ "category_id": 16,
+ "id": 4821,
+ "area": 17854,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2446,
+ "bbox": [
+ 226,
+ 173,
+ 153,
+ 90
+ ],
+ "category_id": 19,
+ "id": 4822,
+ "area": 13770,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2447,
+ "bbox": [
+ 204,
+ 115,
+ 261,
+ 240
+ ],
+ "category_id": 10,
+ "id": 4823,
+ "area": 62640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2448,
+ "bbox": [
+ 140,
+ 280,
+ 245,
+ 66
+ ],
+ "category_id": 16,
+ "id": 4824,
+ "area": 16170,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2449,
+ "bbox": [
+ 136,
+ 300,
+ 80,
+ 114
+ ],
+ "category_id": 19,
+ "id": 4825,
+ "area": 9120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2450,
+ "bbox": [
+ 58,
+ 96,
+ 237,
+ 171
+ ],
+ "category_id": 16,
+ "id": 4826,
+ "area": 40527,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2451,
+ "bbox": [
+ 38,
+ 216,
+ 433,
+ 52
+ ],
+ "category_id": 10,
+ "id": 4827,
+ "area": 22516,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2452,
+ "bbox": [
+ 153,
+ 184,
+ 143,
+ 139
+ ],
+ "category_id": 15,
+ "id": 4828,
+ "area": 19877,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2453,
+ "bbox": [
+ 129,
+ 193,
+ 224,
+ 166
+ ],
+ "category_id": 16,
+ "id": 4829,
+ "area": 37184,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2454,
+ "bbox": [
+ 49,
+ 179,
+ 427,
+ 192
+ ],
+ "category_id": 10,
+ "id": 4830,
+ "area": 81984,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2455,
+ "bbox": [
+ 171,
+ 159,
+ 192,
+ 172
+ ],
+ "category_id": 19,
+ "id": 4831,
+ "area": 33024,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2456,
+ "bbox": [
+ 103,
+ 131,
+ 243,
+ 258
+ ],
+ "category_id": 16,
+ "id": 4832,
+ "area": 62694,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2457,
+ "bbox": [
+ 84,
+ 88,
+ 342,
+ 293
+ ],
+ "category_id": 10,
+ "id": 4833,
+ "area": 100206,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2458,
+ "bbox": [
+ 120,
+ 158,
+ 360,
+ 116
+ ],
+ "category_id": 10,
+ "id": 4834,
+ "area": 41760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2459,
+ "bbox": [
+ 194,
+ 63,
+ 93,
+ 372
+ ],
+ "category_id": 16,
+ "id": 4835,
+ "area": 34596,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2460,
+ "bbox": [
+ 168,
+ 154,
+ 259,
+ 128
+ ],
+ "category_id": 19,
+ "id": 4836,
+ "area": 33152,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2461,
+ "bbox": [
+ 394,
+ 69,
+ 13,
+ 25
+ ],
+ "category_id": 4,
+ "id": 4837,
+ "area": 325,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2461,
+ "bbox": [
+ 134,
+ 207,
+ 15,
+ 24
+ ],
+ "category_id": 4,
+ "id": 4838,
+ "area": 360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2461,
+ "bbox": [
+ 371,
+ 182,
+ 12,
+ 27
+ ],
+ "category_id": 4,
+ "id": 4839,
+ "area": 324,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2461,
+ "bbox": [
+ 323,
+ 324,
+ 14,
+ 25
+ ],
+ "category_id": 4,
+ "id": 4840,
+ "area": 350,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2461,
+ "bbox": [
+ 302,
+ 440,
+ 12,
+ 30
+ ],
+ "category_id": 4,
+ "id": 4841,
+ "area": 360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2461,
+ "bbox": [
+ 42,
+ 382,
+ 13,
+ 29
+ ],
+ "category_id": 4,
+ "id": 4842,
+ "area": 377,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2462,
+ "bbox": [
+ 84,
+ 342,
+ 51,
+ 37
+ ],
+ "category_id": 4,
+ "id": 4843,
+ "area": 1887,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2462,
+ "bbox": [
+ 321,
+ 81,
+ 60,
+ 50
+ ],
+ "category_id": 4,
+ "id": 4844,
+ "area": 3000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2463,
+ "bbox": [
+ 122,
+ 92,
+ 171,
+ 215
+ ],
+ "category_id": 16,
+ "id": 4845,
+ "area": 36765,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2464,
+ "bbox": [
+ 205,
+ 234,
+ 160,
+ 127
+ ],
+ "category_id": 15,
+ "id": 4846,
+ "area": 20320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2464,
+ "bbox": [
+ 25,
+ 216,
+ 155,
+ 125
+ ],
+ "category_id": 15,
+ "id": 4847,
+ "area": 19375,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2465,
+ "bbox": [
+ 63,
+ 48,
+ 7,
+ 17
+ ],
+ "category_id": 4,
+ "id": 4848,
+ "area": 119,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2465,
+ "bbox": [
+ 225,
+ 40,
+ 11,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4849,
+ "area": 176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2465,
+ "bbox": [
+ 459,
+ 106,
+ 13,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4850,
+ "area": 208,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2465,
+ "bbox": [
+ 435,
+ 425,
+ 9,
+ 17
+ ],
+ "category_id": 4,
+ "id": 4851,
+ "area": 153,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2465,
+ "bbox": [
+ 227,
+ 467,
+ 14,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4852,
+ "area": 224,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2466,
+ "bbox": [
+ 303,
+ 140,
+ 128,
+ 77
+ ],
+ "category_id": 15,
+ "id": 4853,
+ "area": 9856,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2467,
+ "bbox": [
+ 141,
+ 208,
+ 266,
+ 225
+ ],
+ "category_id": 10,
+ "id": 4854,
+ "area": 59850,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2468,
+ "bbox": [
+ 83,
+ 70,
+ 158,
+ 92
+ ],
+ "category_id": 10,
+ "id": 4855,
+ "area": 14536,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2469,
+ "bbox": [
+ 186,
+ 133,
+ 146,
+ 144
+ ],
+ "category_id": 19,
+ "id": 4856,
+ "area": 21024,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2470,
+ "bbox": [
+ 181,
+ 96,
+ 258,
+ 292
+ ],
+ "category_id": 10,
+ "id": 4857,
+ "area": 75336,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2471,
+ "bbox": [
+ 166,
+ 51,
+ 17,
+ 16
+ ],
+ "category_id": 4,
+ "id": 4858,
+ "area": 272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2471,
+ "bbox": [
+ 65,
+ 251,
+ 19,
+ 17
+ ],
+ "category_id": 4,
+ "id": 4859,
+ "area": 323,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2471,
+ "bbox": [
+ 348,
+ 133,
+ 17,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4860,
+ "area": 238,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2471,
+ "bbox": [
+ 343,
+ 371,
+ 15,
+ 15
+ ],
+ "category_id": 4,
+ "id": 4861,
+ "area": 225,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2471,
+ "bbox": [
+ 115,
+ 480,
+ 22,
+ 19
+ ],
+ "category_id": 4,
+ "id": 4862,
+ "area": 418,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2472,
+ "bbox": [
+ 172,
+ 190,
+ 140,
+ 42
+ ],
+ "category_id": 4,
+ "id": 4863,
+ "area": 5880,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2473,
+ "bbox": [
+ 160,
+ 201,
+ 140,
+ 234
+ ],
+ "category_id": 16,
+ "id": 4864,
+ "area": 32760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2474,
+ "bbox": [
+ 278,
+ 388,
+ 166,
+ 101
+ ],
+ "category_id": 10,
+ "id": 4865,
+ "area": 16766,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2475,
+ "bbox": [
+ 303,
+ 170,
+ 94,
+ 108
+ ],
+ "category_id": 4,
+ "id": 4866,
+ "area": 10152,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2476,
+ "bbox": [
+ 200,
+ 177,
+ 112,
+ 183
+ ],
+ "category_id": 19,
+ "id": 4867,
+ "area": 20496,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2477,
+ "bbox": [
+ 64,
+ 90,
+ 378,
+ 278
+ ],
+ "category_id": 16,
+ "id": 4868,
+ "area": 105084,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2478,
+ "bbox": [
+ 268,
+ 122,
+ 132,
+ 126
+ ],
+ "category_id": 15,
+ "id": 4869,
+ "area": 16632,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2478,
+ "bbox": [
+ 393,
+ 239,
+ 70,
+ 63
+ ],
+ "category_id": 15,
+ "id": 4870,
+ "area": 4410,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2479,
+ "bbox": [
+ 200,
+ 276,
+ 65,
+ 152
+ ],
+ "category_id": 16,
+ "id": 4871,
+ "area": 9880,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2480,
+ "bbox": [
+ 76,
+ 163,
+ 375,
+ 221
+ ],
+ "category_id": 10,
+ "id": 4872,
+ "area": 82875,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2481,
+ "bbox": [
+ 173,
+ 239,
+ 163,
+ 106
+ ],
+ "category_id": 16,
+ "id": 4873,
+ "area": 17278,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2482,
+ "bbox": [
+ 65,
+ 343,
+ 32,
+ 80
+ ],
+ "category_id": 4,
+ "id": 4874,
+ "area": 2560,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2482,
+ "bbox": [
+ 392,
+ 161,
+ 22,
+ 78
+ ],
+ "category_id": 4,
+ "id": 4875,
+ "area": 1716,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2483,
+ "bbox": [
+ 41,
+ 76,
+ 111,
+ 322
+ ],
+ "category_id": 10,
+ "id": 4876,
+ "area": 35742,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2484,
+ "bbox": [
+ 180,
+ 84,
+ 64,
+ 298
+ ],
+ "category_id": 19,
+ "id": 4877,
+ "area": 19072,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2485,
+ "bbox": [
+ 199,
+ 176,
+ 67,
+ 195
+ ],
+ "category_id": 19,
+ "id": 4878,
+ "area": 13065,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2486,
+ "bbox": [
+ 96,
+ 7,
+ 227,
+ 249
+ ],
+ "category_id": 15,
+ "id": 4879,
+ "area": 56523,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2486,
+ "bbox": [
+ 144,
+ 252,
+ 227,
+ 247
+ ],
+ "category_id": 15,
+ "id": 4880,
+ "area": 56069,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2486,
+ "bbox": [
+ 419,
+ 0,
+ 64,
+ 69
+ ],
+ "category_id": 15,
+ "id": 4881,
+ "area": 4416,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2487,
+ "bbox": [
+ 296,
+ 39,
+ 126,
+ 410
+ ],
+ "category_id": 10,
+ "id": 4882,
+ "area": 51660,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2488,
+ "bbox": [
+ 44,
+ 193,
+ 389,
+ 32
+ ],
+ "category_id": 10,
+ "id": 4883,
+ "area": 12448,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2489,
+ "bbox": [
+ 12,
+ 370,
+ 43,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4884,
+ "area": 430,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2489,
+ "bbox": [
+ 202,
+ 321,
+ 34,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4885,
+ "area": 306,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2489,
+ "bbox": [
+ 343,
+ 239,
+ 37,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4886,
+ "area": 518,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2489,
+ "bbox": [
+ 444,
+ 135,
+ 34,
+ 12
+ ],
+ "category_id": 4,
+ "id": 4887,
+ "area": 408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2490,
+ "bbox": [
+ 44,
+ 264,
+ 412,
+ 152
+ ],
+ "category_id": 16,
+ "id": 4888,
+ "area": 62624,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2491,
+ "bbox": [
+ 80,
+ 195,
+ 412,
+ 183
+ ],
+ "category_id": 16,
+ "id": 4889,
+ "area": 75396,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2492,
+ "bbox": [
+ 117,
+ 209,
+ 315,
+ 209
+ ],
+ "category_id": 10,
+ "id": 4890,
+ "area": 65835,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2493,
+ "bbox": [
+ 163,
+ 180,
+ 259,
+ 123
+ ],
+ "category_id": 16,
+ "id": 4891,
+ "area": 31857,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2494,
+ "bbox": [
+ 179,
+ 179,
+ 126,
+ 91
+ ],
+ "category_id": 15,
+ "id": 4892,
+ "area": 11466,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2495,
+ "bbox": [
+ 152,
+ 186,
+ 123,
+ 124
+ ],
+ "category_id": 15,
+ "id": 4893,
+ "area": 15252,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2496,
+ "bbox": [
+ 96,
+ 357,
+ 101,
+ 87
+ ],
+ "category_id": 10,
+ "id": 4894,
+ "area": 8787,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2497,
+ "bbox": [
+ 307,
+ 240,
+ 93,
+ 92
+ ],
+ "category_id": 4,
+ "id": 4895,
+ "area": 8556,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2498,
+ "bbox": [
+ 296,
+ 32,
+ 12,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4896,
+ "area": 168,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2498,
+ "bbox": [
+ 101,
+ 154,
+ 13,
+ 13
+ ],
+ "category_id": 4,
+ "id": 4897,
+ "area": 169,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2498,
+ "bbox": [
+ 164,
+ 333,
+ 28,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4898,
+ "area": 392,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2498,
+ "bbox": [
+ 319,
+ 392,
+ 28,
+ 18
+ ],
+ "category_id": 4,
+ "id": 4899,
+ "area": 504,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2498,
+ "bbox": [
+ 399,
+ 460,
+ 27,
+ 17
+ ],
+ "category_id": 4,
+ "id": 4900,
+ "area": 459,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2499,
+ "bbox": [
+ 123,
+ 336,
+ 26,
+ 54
+ ],
+ "category_id": 4,
+ "id": 4901,
+ "area": 1404,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2499,
+ "bbox": [
+ 403,
+ 211,
+ 32,
+ 55
+ ],
+ "category_id": 4,
+ "id": 4902,
+ "area": 1760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2500,
+ "bbox": [
+ 224,
+ 95,
+ 54,
+ 291
+ ],
+ "category_id": 19,
+ "id": 4903,
+ "area": 15714,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2501,
+ "bbox": [
+ 201,
+ 214,
+ 7,
+ 7
+ ],
+ "category_id": 15,
+ "id": 4904,
+ "area": 49,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2501,
+ "bbox": [
+ 193,
+ 215,
+ 8,
+ 8
+ ],
+ "category_id": 15,
+ "id": 4905,
+ "area": 64,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2501,
+ "bbox": [
+ 186,
+ 215,
+ 7,
+ 9
+ ],
+ "category_id": 15,
+ "id": 4906,
+ "area": 63,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2502,
+ "bbox": [
+ 304,
+ 282,
+ 44,
+ 24
+ ],
+ "category_id": 15,
+ "id": 4907,
+ "area": 1056,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2502,
+ "bbox": [
+ 189,
+ 396,
+ 93,
+ 100
+ ],
+ "category_id": 15,
+ "id": 4908,
+ "area": 9300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2503,
+ "bbox": [
+ 167,
+ 175,
+ 95,
+ 75
+ ],
+ "category_id": 15,
+ "id": 4909,
+ "area": 7125,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2503,
+ "bbox": [
+ 271,
+ 184,
+ 95,
+ 77
+ ],
+ "category_id": 15,
+ "id": 4910,
+ "area": 7315,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2504,
+ "bbox": [
+ 227,
+ 236,
+ 158,
+ 80
+ ],
+ "category_id": 16,
+ "id": 4911,
+ "area": 12640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2505,
+ "bbox": [
+ 174,
+ 24,
+ 32,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4912,
+ "area": 352,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2505,
+ "bbox": [
+ 394,
+ 21,
+ 32,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4913,
+ "area": 288,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2505,
+ "bbox": [
+ 273,
+ 150,
+ 39,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4914,
+ "area": 390,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2505,
+ "bbox": [
+ 263,
+ 261,
+ 30,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4915,
+ "area": 330,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2505,
+ "bbox": [
+ 296,
+ 369,
+ 35,
+ 10
+ ],
+ "category_id": 4,
+ "id": 4916,
+ "area": 350,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2505,
+ "bbox": [
+ 343,
+ 481,
+ 36,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4917,
+ "area": 324,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2506,
+ "bbox": [
+ 122,
+ 259,
+ 272,
+ 176
+ ],
+ "category_id": 16,
+ "id": 4918,
+ "area": 47872,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2507,
+ "bbox": [
+ 348,
+ 62,
+ 55,
+ 50
+ ],
+ "category_id": 4,
+ "id": 4919,
+ "area": 2750,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2507,
+ "bbox": [
+ 158,
+ 369,
+ 63,
+ 41
+ ],
+ "category_id": 4,
+ "id": 4920,
+ "area": 2583,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2508,
+ "bbox": [
+ 48,
+ 152,
+ 138,
+ 107
+ ],
+ "category_id": 15,
+ "id": 4921,
+ "area": 14766,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2508,
+ "bbox": [
+ 209,
+ 161,
+ 139,
+ 107
+ ],
+ "category_id": 15,
+ "id": 4922,
+ "area": 14873,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2508,
+ "bbox": [
+ 342,
+ 152,
+ 141,
+ 105
+ ],
+ "category_id": 15,
+ "id": 4923,
+ "area": 14805,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2509,
+ "bbox": [
+ 132,
+ 312,
+ 246,
+ 72
+ ],
+ "category_id": 19,
+ "id": 4924,
+ "area": 17712,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2510,
+ "bbox": [
+ 352,
+ 33,
+ 21,
+ 15
+ ],
+ "category_id": 4,
+ "id": 4925,
+ "area": 315,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2510,
+ "bbox": [
+ 227,
+ 115,
+ 18,
+ 23
+ ],
+ "category_id": 4,
+ "id": 4926,
+ "area": 414,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2510,
+ "bbox": [
+ 49,
+ 160,
+ 19,
+ 18
+ ],
+ "category_id": 4,
+ "id": 4927,
+ "area": 342,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2510,
+ "bbox": [
+ 14,
+ 368,
+ 34,
+ 12
+ ],
+ "category_id": 4,
+ "id": 4928,
+ "area": 408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2510,
+ "bbox": [
+ 264,
+ 433,
+ 29,
+ 15
+ ],
+ "category_id": 4,
+ "id": 4929,
+ "area": 435,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2510,
+ "bbox": [
+ 337,
+ 325,
+ 27,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4930,
+ "area": 378,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2510,
+ "bbox": [
+ 469,
+ 196,
+ 24,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4931,
+ "area": 336,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2511,
+ "bbox": [
+ 147,
+ 247,
+ 256,
+ 65
+ ],
+ "category_id": 19,
+ "id": 4932,
+ "area": 16640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2512,
+ "bbox": [
+ 161,
+ 204,
+ 139,
+ 174
+ ],
+ "category_id": 19,
+ "id": 4933,
+ "area": 24186,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2513,
+ "bbox": [
+ 24,
+ 261,
+ 459,
+ 46
+ ],
+ "category_id": 10,
+ "id": 4934,
+ "area": 21114,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2514,
+ "bbox": [
+ 81,
+ 128,
+ 315,
+ 191
+ ],
+ "category_id": 16,
+ "id": 4935,
+ "area": 60165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2515,
+ "bbox": [
+ 44,
+ 42,
+ 57,
+ 240
+ ],
+ "category_id": 10,
+ "id": 4936,
+ "area": 13680,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2516,
+ "bbox": [
+ 182,
+ 252,
+ 100,
+ 179
+ ],
+ "category_id": 19,
+ "id": 4937,
+ "area": 17900,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2517,
+ "bbox": [
+ 161,
+ 210,
+ 224,
+ 54
+ ],
+ "category_id": 10,
+ "id": 4938,
+ "area": 12096,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2518,
+ "bbox": [
+ 197,
+ 54,
+ 243,
+ 256
+ ],
+ "category_id": 16,
+ "id": 4939,
+ "area": 62208,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2519,
+ "bbox": [
+ 107,
+ 143,
+ 122,
+ 129
+ ],
+ "category_id": 15,
+ "id": 4940,
+ "area": 15738,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2519,
+ "bbox": [
+ 115,
+ 309,
+ 105,
+ 98
+ ],
+ "category_id": 15,
+ "id": 4941,
+ "area": 10290,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2520,
+ "bbox": [
+ 162,
+ 380,
+ 127,
+ 116
+ ],
+ "category_id": 15,
+ "id": 4942,
+ "area": 14732,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2520,
+ "bbox": [
+ 200,
+ 258,
+ 80,
+ 73
+ ],
+ "category_id": 15,
+ "id": 4943,
+ "area": 5840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2520,
+ "bbox": [
+ 300,
+ 293,
+ 81,
+ 70
+ ],
+ "category_id": 15,
+ "id": 4944,
+ "area": 5670,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2521,
+ "bbox": [
+ 114,
+ 265,
+ 97,
+ 124
+ ],
+ "category_id": 15,
+ "id": 4945,
+ "area": 12028,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2521,
+ "bbox": [
+ 353,
+ 98,
+ 29,
+ 123
+ ],
+ "category_id": 15,
+ "id": 4946,
+ "area": 3567,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2522,
+ "bbox": [
+ 146,
+ 445,
+ 14,
+ 20
+ ],
+ "category_id": 16,
+ "id": 4947,
+ "area": 280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2523,
+ "bbox": [
+ 285,
+ 36,
+ 189,
+ 99
+ ],
+ "category_id": 10,
+ "id": 4948,
+ "area": 18711,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2524,
+ "bbox": [
+ 176,
+ 159,
+ 269,
+ 317
+ ],
+ "category_id": 16,
+ "id": 4949,
+ "area": 85273,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2525,
+ "bbox": [
+ 75,
+ 67,
+ 274,
+ 68
+ ],
+ "category_id": 19,
+ "id": 4950,
+ "area": 18632,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2526,
+ "bbox": [
+ 208,
+ 89,
+ 221,
+ 172
+ ],
+ "category_id": 15,
+ "id": 4951,
+ "area": 38012,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2526,
+ "bbox": [
+ 125,
+ 303,
+ 234,
+ 166
+ ],
+ "category_id": 15,
+ "id": 4952,
+ "area": 38844,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2527,
+ "bbox": [
+ 50,
+ 63,
+ 63,
+ 57
+ ],
+ "category_id": 4,
+ "id": 4953,
+ "area": 3591,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2527,
+ "bbox": [
+ 412,
+ 372,
+ 77,
+ 40
+ ],
+ "category_id": 4,
+ "id": 4954,
+ "area": 3080,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2528,
+ "bbox": [
+ 55,
+ 0,
+ 23,
+ 21
+ ],
+ "category_id": 4,
+ "id": 4955,
+ "area": 483,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2528,
+ "bbox": [
+ 104,
+ 152,
+ 22,
+ 20
+ ],
+ "category_id": 4,
+ "id": 4956,
+ "area": 440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2528,
+ "bbox": [
+ 189,
+ 323,
+ 38,
+ 11
+ ],
+ "category_id": 4,
+ "id": 4957,
+ "area": 418,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2528,
+ "bbox": [
+ 43,
+ 417,
+ 26,
+ 24
+ ],
+ "category_id": 4,
+ "id": 4958,
+ "area": 624,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2528,
+ "bbox": [
+ 394,
+ 248,
+ 34,
+ 18
+ ],
+ "category_id": 4,
+ "id": 4959,
+ "area": 612,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2528,
+ "bbox": [
+ 478,
+ 410,
+ 30,
+ 13
+ ],
+ "category_id": 4,
+ "id": 4960,
+ "area": 390,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2528,
+ "bbox": [
+ 423,
+ 487,
+ 21,
+ 18
+ ],
+ "category_id": 4,
+ "id": 4961,
+ "area": 378,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2529,
+ "bbox": [
+ 229,
+ 286,
+ 132,
+ 67
+ ],
+ "category_id": 4,
+ "id": 4962,
+ "area": 8844,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2530,
+ "bbox": [
+ 272,
+ 254,
+ 195,
+ 202
+ ],
+ "category_id": 10,
+ "id": 4963,
+ "area": 39390,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2531,
+ "bbox": [
+ 50,
+ 154,
+ 390,
+ 176
+ ],
+ "category_id": 10,
+ "id": 4964,
+ "area": 68640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2532,
+ "bbox": [
+ 212,
+ 23,
+ 21,
+ 14
+ ],
+ "category_id": 4,
+ "id": 4965,
+ "area": 294,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2532,
+ "bbox": [
+ 450,
+ 63,
+ 20,
+ 18
+ ],
+ "category_id": 4,
+ "id": 4966,
+ "area": 360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2532,
+ "bbox": [
+ 62,
+ 264,
+ 15,
+ 25
+ ],
+ "category_id": 4,
+ "id": 4967,
+ "area": 375,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2532,
+ "bbox": [
+ 41,
+ 439,
+ 10,
+ 26
+ ],
+ "category_id": 4,
+ "id": 4968,
+ "area": 260,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2532,
+ "bbox": [
+ 442,
+ 401,
+ 14,
+ 23
+ ],
+ "category_id": 4,
+ "id": 4969,
+ "area": 322,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2533,
+ "bbox": [
+ 81,
+ 208,
+ 101,
+ 85
+ ],
+ "category_id": 15,
+ "id": 4970,
+ "area": 8585,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2533,
+ "bbox": [
+ 144,
+ 129,
+ 103,
+ 88
+ ],
+ "category_id": 15,
+ "id": 4971,
+ "area": 9064,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2533,
+ "bbox": [
+ 345,
+ 197,
+ 93,
+ 79
+ ],
+ "category_id": 15,
+ "id": 4972,
+ "area": 7347,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2533,
+ "bbox": [
+ 109,
+ 314,
+ 106,
+ 91
+ ],
+ "category_id": 15,
+ "id": 4973,
+ "area": 9646,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2534,
+ "bbox": [
+ 33,
+ 68,
+ 227,
+ 381
+ ],
+ "category_id": 10,
+ "id": 4974,
+ "area": 86487,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2535,
+ "bbox": [
+ 130,
+ 176,
+ 234,
+ 296
+ ],
+ "category_id": 15,
+ "id": 4975,
+ "area": 69264,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2536,
+ "bbox": [
+ 227,
+ 189,
+ 44,
+ 111
+ ],
+ "category_id": 19,
+ "id": 4976,
+ "area": 4884,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2537,
+ "bbox": [
+ 205,
+ 145,
+ 42,
+ 41
+ ],
+ "category_id": 4,
+ "id": 4977,
+ "area": 1722,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2537,
+ "bbox": [
+ 358,
+ 369,
+ 45,
+ 44
+ ],
+ "category_id": 4,
+ "id": 4978,
+ "area": 1980,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2538,
+ "bbox": [
+ 300,
+ 35,
+ 43,
+ 385
+ ],
+ "category_id": 10,
+ "id": 4979,
+ "area": 16555,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2539,
+ "bbox": [
+ 176,
+ 152,
+ 101,
+ 308
+ ],
+ "category_id": 16,
+ "id": 4980,
+ "area": 31108,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2540,
+ "bbox": [
+ 127,
+ 293,
+ 57,
+ 53
+ ],
+ "category_id": 16,
+ "id": 4981,
+ "area": 3021,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2540,
+ "bbox": [
+ 236,
+ 295,
+ 57,
+ 62
+ ],
+ "category_id": 16,
+ "id": 4982,
+ "area": 3534,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2541,
+ "bbox": [
+ 293,
+ 188,
+ 185,
+ 264
+ ],
+ "category_id": 19,
+ "id": 4983,
+ "area": 48840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2542,
+ "bbox": [
+ 69,
+ 181,
+ 419,
+ 64
+ ],
+ "category_id": 10,
+ "id": 4984,
+ "area": 26816,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2543,
+ "bbox": [
+ 24,
+ 147,
+ 48,
+ 23
+ ],
+ "category_id": 4,
+ "id": 4985,
+ "area": 1104,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2543,
+ "bbox": [
+ 375,
+ 67,
+ 55,
+ 27
+ ],
+ "category_id": 4,
+ "id": 4986,
+ "area": 1485,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2543,
+ "bbox": [
+ 383,
+ 397,
+ 50,
+ 30
+ ],
+ "category_id": 4,
+ "id": 4987,
+ "area": 1500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2544,
+ "bbox": [
+ 86,
+ 157,
+ 111,
+ 134
+ ],
+ "category_id": 15,
+ "id": 4988,
+ "area": 14874,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2544,
+ "bbox": [
+ 189,
+ 254,
+ 113,
+ 135
+ ],
+ "category_id": 15,
+ "id": 4989,
+ "area": 15255,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2544,
+ "bbox": [
+ 288,
+ 148,
+ 112,
+ 134
+ ],
+ "category_id": 15,
+ "id": 4990,
+ "area": 15008,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2545,
+ "bbox": [
+ 331,
+ 147,
+ 172,
+ 110
+ ],
+ "category_id": 10,
+ "id": 4991,
+ "area": 18920,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2546,
+ "bbox": [
+ 277,
+ 98,
+ 28,
+ 17
+ ],
+ "category_id": 4,
+ "id": 4992,
+ "area": 476,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2546,
+ "bbox": [
+ 240,
+ 365,
+ 39,
+ 32
+ ],
+ "category_id": 4,
+ "id": 4993,
+ "area": 1248,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2547,
+ "bbox": [
+ 51,
+ 326,
+ 103,
+ 29
+ ],
+ "category_id": 15,
+ "id": 4994,
+ "area": 2987,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2547,
+ "bbox": [
+ 176,
+ 244,
+ 71,
+ 24
+ ],
+ "category_id": 15,
+ "id": 4995,
+ "area": 1704,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2547,
+ "bbox": [
+ 375,
+ 163,
+ 104,
+ 34
+ ],
+ "category_id": 15,
+ "id": 4996,
+ "area": 3536,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2548,
+ "bbox": [
+ 360,
+ 356,
+ 104,
+ 106
+ ],
+ "category_id": 10,
+ "id": 4997,
+ "area": 11024,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2549,
+ "bbox": [
+ 40,
+ 263,
+ 7,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4998,
+ "area": 63,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2549,
+ "bbox": [
+ 82,
+ 204,
+ 6,
+ 9
+ ],
+ "category_id": 4,
+ "id": 4999,
+ "area": 54,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2549,
+ "bbox": [
+ 71,
+ 248,
+ 3,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5000,
+ "area": 27,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2549,
+ "bbox": [
+ 69,
+ 354,
+ 5,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5001,
+ "area": 45,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2549,
+ "bbox": [
+ 223,
+ 246,
+ 7,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5002,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2549,
+ "bbox": [
+ 238,
+ 158,
+ 5,
+ 8
+ ],
+ "category_id": 4,
+ "id": 5003,
+ "area": 40,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2549,
+ "bbox": [
+ 359,
+ 170,
+ 5,
+ 7
+ ],
+ "category_id": 4,
+ "id": 5004,
+ "area": 35,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2549,
+ "bbox": [
+ 325,
+ 239,
+ 7,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5005,
+ "area": 77,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2549,
+ "bbox": [
+ 277,
+ 279,
+ 3,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5006,
+ "area": 27,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2549,
+ "bbox": [
+ 446,
+ 220,
+ 6,
+ 10
+ ],
+ "category_id": 4,
+ "id": 5007,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2549,
+ "bbox": [
+ 148,
+ 315,
+ 8,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5008,
+ "area": 72,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2550,
+ "bbox": [
+ 203,
+ 101,
+ 151,
+ 295
+ ],
+ "category_id": 19,
+ "id": 5009,
+ "area": 44545,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2551,
+ "bbox": [
+ 174,
+ 144,
+ 180,
+ 258
+ ],
+ "category_id": 16,
+ "id": 5010,
+ "area": 46440,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2552,
+ "bbox": [
+ 147,
+ 211,
+ 148,
+ 135
+ ],
+ "category_id": 15,
+ "id": 5011,
+ "area": 19980,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2553,
+ "bbox": [
+ 105,
+ 176,
+ 44,
+ 31
+ ],
+ "category_id": 4,
+ "id": 5012,
+ "area": 1364,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2553,
+ "bbox": [
+ 384,
+ 317,
+ 56,
+ 34
+ ],
+ "category_id": 4,
+ "id": 5013,
+ "area": 1904,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2554,
+ "bbox": [
+ 270,
+ 61,
+ 30,
+ 160
+ ],
+ "category_id": 16,
+ "id": 5014,
+ "area": 4800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2554,
+ "bbox": [
+ 231,
+ 308,
+ 74,
+ 153
+ ],
+ "category_id": 16,
+ "id": 5015,
+ "area": 11322,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2555,
+ "bbox": [
+ 116,
+ 346,
+ 35,
+ 18
+ ],
+ "category_id": 4,
+ "id": 5016,
+ "area": 630,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2555,
+ "bbox": [
+ 193,
+ 76,
+ 32,
+ 17
+ ],
+ "category_id": 4,
+ "id": 5017,
+ "area": 544,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2555,
+ "bbox": [
+ 424,
+ 279,
+ 28,
+ 22
+ ],
+ "category_id": 4,
+ "id": 5018,
+ "area": 616,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2555,
+ "bbox": [
+ 365,
+ 451,
+ 27,
+ 16
+ ],
+ "category_id": 4,
+ "id": 5019,
+ "area": 432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2556,
+ "bbox": [
+ 325,
+ 95,
+ 29,
+ 45
+ ],
+ "category_id": 4,
+ "id": 5020,
+ "area": 1305,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2556,
+ "bbox": [
+ 154,
+ 343,
+ 68,
+ 38
+ ],
+ "category_id": 4,
+ "id": 5021,
+ "area": 2584,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2557,
+ "bbox": [
+ 10,
+ 146,
+ 23,
+ 24
+ ],
+ "category_id": 4,
+ "id": 5022,
+ "area": 552,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2557,
+ "bbox": [
+ 60,
+ 263,
+ 22,
+ 26
+ ],
+ "category_id": 4,
+ "id": 5023,
+ "area": 572,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2557,
+ "bbox": [
+ 56,
+ 390,
+ 15,
+ 31
+ ],
+ "category_id": 4,
+ "id": 5024,
+ "area": 465,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2557,
+ "bbox": [
+ 8,
+ 489,
+ 16,
+ 23
+ ],
+ "category_id": 4,
+ "id": 5025,
+ "area": 368,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2557,
+ "bbox": [
+ 457,
+ 104,
+ 20,
+ 29
+ ],
+ "category_id": 4,
+ "id": 5026,
+ "area": 580,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2557,
+ "bbox": [
+ 487,
+ 309,
+ 18,
+ 23
+ ],
+ "category_id": 4,
+ "id": 5027,
+ "area": 414,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2557,
+ "bbox": [
+ 106,
+ 28,
+ 22,
+ 27
+ ],
+ "category_id": 4,
+ "id": 5028,
+ "area": 594,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2558,
+ "bbox": [
+ 295,
+ 171,
+ 71,
+ 72
+ ],
+ "category_id": 15,
+ "id": 5029,
+ "area": 5112,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2559,
+ "bbox": [
+ 278,
+ 42,
+ 65,
+ 403
+ ],
+ "category_id": 10,
+ "id": 5030,
+ "area": 26195,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2560,
+ "bbox": [
+ 211,
+ 127,
+ 203,
+ 207
+ ],
+ "category_id": 10,
+ "id": 5031,
+ "area": 42021,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2561,
+ "bbox": [
+ 83,
+ 102,
+ 124,
+ 59
+ ],
+ "category_id": 19,
+ "id": 5032,
+ "area": 7316,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2562,
+ "bbox": [
+ 273,
+ 21,
+ 37,
+ 19
+ ],
+ "category_id": 4,
+ "id": 5033,
+ "area": 703,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2562,
+ "bbox": [
+ 425,
+ 30,
+ 40,
+ 18
+ ],
+ "category_id": 4,
+ "id": 5034,
+ "area": 720,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2562,
+ "bbox": [
+ 181,
+ 101,
+ 34,
+ 18
+ ],
+ "category_id": 4,
+ "id": 5035,
+ "area": 612,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2562,
+ "bbox": [
+ 366,
+ 120,
+ 35,
+ 20
+ ],
+ "category_id": 4,
+ "id": 5036,
+ "area": 700,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2562,
+ "bbox": [
+ 181,
+ 248,
+ 35,
+ 20
+ ],
+ "category_id": 4,
+ "id": 5037,
+ "area": 700,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2562,
+ "bbox": [
+ 12,
+ 389,
+ 36,
+ 18
+ ],
+ "category_id": 4,
+ "id": 5038,
+ "area": 648,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2563,
+ "bbox": [
+ 200,
+ 49,
+ 11,
+ 27
+ ],
+ "category_id": 4,
+ "id": 5039,
+ "area": 297,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2563,
+ "bbox": [
+ 20,
+ 248,
+ 16,
+ 23
+ ],
+ "category_id": 4,
+ "id": 5040,
+ "area": 368,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2563,
+ "bbox": [
+ 346,
+ 277,
+ 14,
+ 19
+ ],
+ "category_id": 4,
+ "id": 5041,
+ "area": 266,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2563,
+ "bbox": [
+ 124,
+ 435,
+ 15,
+ 22
+ ],
+ "category_id": 4,
+ "id": 5042,
+ "area": 330,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2563,
+ "bbox": [
+ 492,
+ 76,
+ 14,
+ 21
+ ],
+ "category_id": 4,
+ "id": 5043,
+ "area": 294,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2564,
+ "bbox": [
+ 187,
+ 280,
+ 127,
+ 20
+ ],
+ "category_id": 19,
+ "id": 5044,
+ "area": 2540,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2565,
+ "bbox": [
+ 109,
+ 179,
+ 90,
+ 92
+ ],
+ "category_id": 15,
+ "id": 5045,
+ "area": 8280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2565,
+ "bbox": [
+ 412,
+ 394,
+ 21,
+ 46
+ ],
+ "category_id": 15,
+ "id": 5046,
+ "area": 966,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2566,
+ "bbox": [
+ 51,
+ 122,
+ 440,
+ 254
+ ],
+ "category_id": 10,
+ "id": 5047,
+ "area": 111760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2567,
+ "bbox": [
+ 182,
+ 86,
+ 27,
+ 53
+ ],
+ "category_id": 4,
+ "id": 5048,
+ "area": 1431,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2567,
+ "bbox": [
+ 201,
+ 401,
+ 32,
+ 48
+ ],
+ "category_id": 4,
+ "id": 5049,
+ "area": 1536,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2568,
+ "bbox": [
+ 126,
+ 332,
+ 85,
+ 45
+ ],
+ "category_id": 15,
+ "id": 5050,
+ "area": 3825,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2569,
+ "bbox": [
+ 110,
+ 172,
+ 289,
+ 148
+ ],
+ "category_id": 19,
+ "id": 5051,
+ "area": 42772,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2570,
+ "bbox": [
+ 158,
+ 157,
+ 117,
+ 211
+ ],
+ "category_id": 16,
+ "id": 5052,
+ "area": 24687,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2571,
+ "bbox": [
+ 71,
+ 101,
+ 190,
+ 53
+ ],
+ "category_id": 19,
+ "id": 5053,
+ "area": 10070,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2572,
+ "bbox": [
+ 236,
+ 114,
+ 196,
+ 236
+ ],
+ "category_id": 16,
+ "id": 5054,
+ "area": 46256,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2573,
+ "bbox": [
+ 115,
+ 144,
+ 343,
+ 220
+ ],
+ "category_id": 10,
+ "id": 5055,
+ "area": 75460,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2574,
+ "bbox": [
+ 8,
+ 29,
+ 26,
+ 56
+ ],
+ "category_id": 16,
+ "id": 5056,
+ "area": 1456,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2575,
+ "bbox": [
+ 267,
+ 164,
+ 39,
+ 204
+ ],
+ "category_id": 19,
+ "id": 5057,
+ "area": 7956,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2576,
+ "bbox": [
+ 110,
+ 104,
+ 51,
+ 25
+ ],
+ "category_id": 4,
+ "id": 5058,
+ "area": 1275,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2576,
+ "bbox": [
+ 305,
+ 322,
+ 38,
+ 30
+ ],
+ "category_id": 4,
+ "id": 5059,
+ "area": 1140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2576,
+ "bbox": [
+ 340,
+ 437,
+ 50,
+ 38
+ ],
+ "category_id": 4,
+ "id": 5060,
+ "area": 1900,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2577,
+ "bbox": [
+ 127,
+ 88,
+ 280,
+ 304
+ ],
+ "category_id": 16,
+ "id": 5061,
+ "area": 85120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2578,
+ "bbox": [
+ 159,
+ 298,
+ 114,
+ 72
+ ],
+ "category_id": 16,
+ "id": 5062,
+ "area": 8208,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2579,
+ "bbox": [
+ 117,
+ 115,
+ 253,
+ 313
+ ],
+ "category_id": 10,
+ "id": 5063,
+ "area": 79189,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2580,
+ "bbox": [
+ 248,
+ 167,
+ 79,
+ 114
+ ],
+ "category_id": 15,
+ "id": 5064,
+ "area": 9006,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2580,
+ "bbox": [
+ 244,
+ 281,
+ 79,
+ 111
+ ],
+ "category_id": 15,
+ "id": 5065,
+ "area": 8769,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2581,
+ "bbox": [
+ 207,
+ 204,
+ 113,
+ 202
+ ],
+ "category_id": 16,
+ "id": 5066,
+ "area": 22826,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2582,
+ "bbox": [
+ 79,
+ 96,
+ 131,
+ 154
+ ],
+ "category_id": 15,
+ "id": 5067,
+ "area": 20174,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2582,
+ "bbox": [
+ 83,
+ 286,
+ 112,
+ 133
+ ],
+ "category_id": 15,
+ "id": 5068,
+ "area": 14896,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2582,
+ "bbox": [
+ 263,
+ 311,
+ 112,
+ 135
+ ],
+ "category_id": 15,
+ "id": 5069,
+ "area": 15120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2582,
+ "bbox": [
+ 329,
+ 135,
+ 137,
+ 169
+ ],
+ "category_id": 15,
+ "id": 5070,
+ "area": 23153,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2583,
+ "bbox": [
+ 90,
+ 158,
+ 48,
+ 26
+ ],
+ "category_id": 4,
+ "id": 5071,
+ "area": 1248,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2583,
+ "bbox": [
+ 324,
+ 80,
+ 48,
+ 28
+ ],
+ "category_id": 4,
+ "id": 5072,
+ "area": 1344,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2583,
+ "bbox": [
+ 364,
+ 407,
+ 49,
+ 26
+ ],
+ "category_id": 4,
+ "id": 5073,
+ "area": 1274,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2584,
+ "bbox": [
+ 107,
+ 151,
+ 47,
+ 39
+ ],
+ "category_id": 4,
+ "id": 5074,
+ "area": 1833,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2584,
+ "bbox": [
+ 85,
+ 424,
+ 35,
+ 28
+ ],
+ "category_id": 4,
+ "id": 5075,
+ "area": 980,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2584,
+ "bbox": [
+ 407,
+ 368,
+ 36,
+ 26
+ ],
+ "category_id": 4,
+ "id": 5076,
+ "area": 936,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2584,
+ "bbox": [
+ 368,
+ 109,
+ 40,
+ 26
+ ],
+ "category_id": 4,
+ "id": 5077,
+ "area": 1040,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2585,
+ "bbox": [
+ 101,
+ 204,
+ 371,
+ 36
+ ],
+ "category_id": 16,
+ "id": 5078,
+ "area": 13356,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2586,
+ "bbox": [
+ 79,
+ 21,
+ 174,
+ 279
+ ],
+ "category_id": 19,
+ "id": 5079,
+ "area": 48546,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2587,
+ "bbox": [
+ 165,
+ 158,
+ 283,
+ 64
+ ],
+ "category_id": 10,
+ "id": 5080,
+ "area": 18112,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2587,
+ "bbox": [
+ 72,
+ 283,
+ 324,
+ 58
+ ],
+ "category_id": 10,
+ "id": 5081,
+ "area": 18792,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2588,
+ "bbox": [
+ 270,
+ 80,
+ 58,
+ 311
+ ],
+ "category_id": 19,
+ "id": 5082,
+ "area": 18038,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2589,
+ "bbox": [
+ 119,
+ 193,
+ 288,
+ 230
+ ],
+ "category_id": 10,
+ "id": 5083,
+ "area": 66240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2590,
+ "bbox": [
+ 220,
+ 80,
+ 156,
+ 302
+ ],
+ "category_id": 10,
+ "id": 5084,
+ "area": 47112,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2591,
+ "bbox": [
+ 192,
+ 90,
+ 177,
+ 336
+ ],
+ "category_id": 10,
+ "id": 5085,
+ "area": 59472,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2592,
+ "bbox": [
+ 118,
+ 159,
+ 325,
+ 190
+ ],
+ "category_id": 10,
+ "id": 5086,
+ "area": 61750,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2593,
+ "bbox": [
+ 53,
+ 79,
+ 372,
+ 343
+ ],
+ "category_id": 10,
+ "id": 5087,
+ "area": 127596,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2594,
+ "bbox": [
+ 256,
+ 1,
+ 32,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5088,
+ "area": 352,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2594,
+ "bbox": [
+ 467,
+ 33,
+ 29,
+ 15
+ ],
+ "category_id": 4,
+ "id": 5089,
+ "area": 435,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2594,
+ "bbox": [
+ 58,
+ 168,
+ 31,
+ 13
+ ],
+ "category_id": 4,
+ "id": 5090,
+ "area": 403,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2594,
+ "bbox": [
+ 362,
+ 211,
+ 25,
+ 20
+ ],
+ "category_id": 4,
+ "id": 5091,
+ "area": 500,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2594,
+ "bbox": [
+ 321,
+ 333,
+ 28,
+ 19
+ ],
+ "category_id": 4,
+ "id": 5092,
+ "area": 532,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2594,
+ "bbox": [
+ 73,
+ 388,
+ 30,
+ 16
+ ],
+ "category_id": 4,
+ "id": 5093,
+ "area": 480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2594,
+ "bbox": [
+ 373,
+ 462,
+ 26,
+ 18
+ ],
+ "category_id": 4,
+ "id": 5094,
+ "area": 468,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2594,
+ "bbox": [
+ 238,
+ 488,
+ 27,
+ 16
+ ],
+ "category_id": 4,
+ "id": 5095,
+ "area": 432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2595,
+ "bbox": [
+ 46,
+ 72,
+ 414,
+ 109
+ ],
+ "category_id": 10,
+ "id": 5096,
+ "area": 45126,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2596,
+ "bbox": [
+ 128,
+ 119,
+ 170,
+ 231
+ ],
+ "category_id": 10,
+ "id": 5097,
+ "area": 39270,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2597,
+ "bbox": [
+ 215,
+ 203,
+ 119,
+ 98
+ ],
+ "category_id": 4,
+ "id": 5098,
+ "area": 11662,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2598,
+ "bbox": [
+ 173,
+ 271,
+ 141,
+ 41
+ ],
+ "category_id": 19,
+ "id": 5099,
+ "area": 5781,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2599,
+ "bbox": [
+ 61,
+ 118,
+ 56,
+ 68
+ ],
+ "category_id": 4,
+ "id": 5100,
+ "area": 3808,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2599,
+ "bbox": [
+ 55,
+ 376,
+ 36,
+ 75
+ ],
+ "category_id": 4,
+ "id": 5101,
+ "area": 2700,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2599,
+ "bbox": [
+ 432,
+ 244,
+ 43,
+ 77
+ ],
+ "category_id": 4,
+ "id": 5102,
+ "area": 3311,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2600,
+ "bbox": [
+ 12,
+ 17,
+ 30,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5103,
+ "area": 330,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2600,
+ "bbox": [
+ 129,
+ 36,
+ 31,
+ 10
+ ],
+ "category_id": 4,
+ "id": 5104,
+ "area": 310,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2600,
+ "bbox": [
+ 232,
+ 134,
+ 31,
+ 17
+ ],
+ "category_id": 4,
+ "id": 5105,
+ "area": 527,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2600,
+ "bbox": [
+ 413,
+ 211,
+ 29,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5106,
+ "area": 319,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2601,
+ "bbox": [
+ 71,
+ 328,
+ 37,
+ 63
+ ],
+ "category_id": 16,
+ "id": 5107,
+ "area": 2331,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2602,
+ "bbox": [
+ 247,
+ 99,
+ 73,
+ 349
+ ],
+ "category_id": 10,
+ "id": 5108,
+ "area": 25477,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2603,
+ "bbox": [
+ 192,
+ 211,
+ 112,
+ 87
+ ],
+ "category_id": 15,
+ "id": 5109,
+ "area": 9744,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2603,
+ "bbox": [
+ 193,
+ 96,
+ 112,
+ 89
+ ],
+ "category_id": 15,
+ "id": 5110,
+ "area": 9968,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2603,
+ "bbox": [
+ 384,
+ 3,
+ 107,
+ 17
+ ],
+ "category_id": 15,
+ "id": 5111,
+ "area": 1819,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2604,
+ "bbox": [
+ 129,
+ 103,
+ 34,
+ 20
+ ],
+ "category_id": 4,
+ "id": 5112,
+ "area": 680,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2604,
+ "bbox": [
+ 373,
+ 141,
+ 24,
+ 10
+ ],
+ "category_id": 4,
+ "id": 5113,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2604,
+ "bbox": [
+ 144,
+ 298,
+ 22,
+ 23
+ ],
+ "category_id": 4,
+ "id": 5114,
+ "area": 506,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2605,
+ "bbox": [
+ 198,
+ 225,
+ 118,
+ 89
+ ],
+ "category_id": 4,
+ "id": 5115,
+ "area": 10502,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2606,
+ "bbox": [
+ 164,
+ 108,
+ 27,
+ 30
+ ],
+ "category_id": 4,
+ "id": 5116,
+ "area": 810,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2606,
+ "bbox": [
+ 227,
+ 439,
+ 41,
+ 29
+ ],
+ "category_id": 4,
+ "id": 5117,
+ "area": 1189,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2606,
+ "bbox": [
+ 414,
+ 257,
+ 25,
+ 32
+ ],
+ "category_id": 4,
+ "id": 5118,
+ "area": 800,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2607,
+ "bbox": [
+ 318,
+ 58,
+ 76,
+ 411
+ ],
+ "category_id": 10,
+ "id": 5119,
+ "area": 31236,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2608,
+ "bbox": [
+ 112,
+ 201,
+ 323,
+ 193
+ ],
+ "category_id": 10,
+ "id": 5120,
+ "area": 62339,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2609,
+ "bbox": [
+ 171,
+ 218,
+ 24,
+ 27
+ ],
+ "category_id": 4,
+ "id": 5121,
+ "area": 648,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2609,
+ "bbox": [
+ 201,
+ 356,
+ 35,
+ 27
+ ],
+ "category_id": 4,
+ "id": 5122,
+ "area": 945,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2609,
+ "bbox": [
+ 362,
+ 437,
+ 28,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5123,
+ "area": 308,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2609,
+ "bbox": [
+ 128,
+ 437,
+ 18,
+ 19
+ ],
+ "category_id": 4,
+ "id": 5124,
+ "area": 342,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2610,
+ "bbox": [
+ 325,
+ 42,
+ 12,
+ 13
+ ],
+ "category_id": 4,
+ "id": 5125,
+ "area": 156,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2610,
+ "bbox": [
+ 118,
+ 245,
+ 11,
+ 16
+ ],
+ "category_id": 4,
+ "id": 5126,
+ "area": 176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2610,
+ "bbox": [
+ 334,
+ 192,
+ 12,
+ 14
+ ],
+ "category_id": 4,
+ "id": 5127,
+ "area": 168,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2610,
+ "bbox": [
+ 355,
+ 321,
+ 10,
+ 14
+ ],
+ "category_id": 4,
+ "id": 5128,
+ "area": 140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2611,
+ "bbox": [
+ 288,
+ 428,
+ 178,
+ 37
+ ],
+ "category_id": 10,
+ "id": 5129,
+ "area": 6586,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2612,
+ "bbox": [
+ 147,
+ 144,
+ 274,
+ 271
+ ],
+ "category_id": 16,
+ "id": 5130,
+ "area": 74254,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2613,
+ "bbox": [
+ 110,
+ 46,
+ 48,
+ 27
+ ],
+ "category_id": 4,
+ "id": 5131,
+ "area": 1296,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2613,
+ "bbox": [
+ 309,
+ 395,
+ 59,
+ 30
+ ],
+ "category_id": 4,
+ "id": 5132,
+ "area": 1770,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2613,
+ "bbox": [
+ 443,
+ 143,
+ 60,
+ 27
+ ],
+ "category_id": 4,
+ "id": 5133,
+ "area": 1620,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2614,
+ "bbox": [
+ 185,
+ 105,
+ 102,
+ 112
+ ],
+ "category_id": 15,
+ "id": 5134,
+ "area": 11424,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2614,
+ "bbox": [
+ 323,
+ 165,
+ 93,
+ 107
+ ],
+ "category_id": 15,
+ "id": 5135,
+ "area": 9951,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2614,
+ "bbox": [
+ 410,
+ 280,
+ 50,
+ 52
+ ],
+ "category_id": 15,
+ "id": 5136,
+ "area": 2600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2615,
+ "bbox": [
+ 186,
+ 240,
+ 131,
+ 68
+ ],
+ "category_id": 4,
+ "id": 5137,
+ "area": 8908,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2616,
+ "bbox": [
+ 92,
+ 90,
+ 148,
+ 147
+ ],
+ "category_id": 15,
+ "id": 5138,
+ "area": 21756,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2616,
+ "bbox": [
+ 118,
+ 290,
+ 147,
+ 142
+ ],
+ "category_id": 15,
+ "id": 5139,
+ "area": 20874,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2617,
+ "bbox": [
+ 101,
+ 343,
+ 77,
+ 45
+ ],
+ "category_id": 4,
+ "id": 5140,
+ "area": 3465,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2617,
+ "bbox": [
+ 299,
+ 131,
+ 83,
+ 28
+ ],
+ "category_id": 4,
+ "id": 5141,
+ "area": 2324,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2618,
+ "bbox": [
+ 69,
+ 237,
+ 152,
+ 127
+ ],
+ "category_id": 15,
+ "id": 5142,
+ "area": 19304,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2618,
+ "bbox": [
+ 240,
+ 220,
+ 156,
+ 119
+ ],
+ "category_id": 15,
+ "id": 5143,
+ "area": 18564,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2619,
+ "bbox": [
+ 16,
+ 385,
+ 42,
+ 25
+ ],
+ "category_id": 4,
+ "id": 5144,
+ "area": 1050,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2619,
+ "bbox": [
+ 177,
+ 81,
+ 38,
+ 26
+ ],
+ "category_id": 4,
+ "id": 5145,
+ "area": 988,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2619,
+ "bbox": [
+ 472,
+ 247,
+ 37,
+ 28
+ ],
+ "category_id": 4,
+ "id": 5146,
+ "area": 1036,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2620,
+ "bbox": [
+ 151,
+ 58,
+ 12,
+ 21
+ ],
+ "category_id": 4,
+ "id": 5147,
+ "area": 252,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2620,
+ "bbox": [
+ 107,
+ 360,
+ 12,
+ 19
+ ],
+ "category_id": 4,
+ "id": 5148,
+ "area": 228,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2620,
+ "bbox": [
+ 275,
+ 181,
+ 13,
+ 25
+ ],
+ "category_id": 4,
+ "id": 5149,
+ "area": 325,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2620,
+ "bbox": [
+ 450,
+ 132,
+ 11,
+ 22
+ ],
+ "category_id": 4,
+ "id": 5150,
+ "area": 242,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2620,
+ "bbox": [
+ 435,
+ 284,
+ 11,
+ 20
+ ],
+ "category_id": 4,
+ "id": 5151,
+ "area": 220,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2621,
+ "bbox": [
+ 139,
+ 246,
+ 228,
+ 62
+ ],
+ "category_id": 19,
+ "id": 5152,
+ "area": 14136,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2622,
+ "bbox": [
+ 188,
+ 209,
+ 185,
+ 95
+ ],
+ "category_id": 19,
+ "id": 5153,
+ "area": 17575,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2623,
+ "bbox": [
+ 61,
+ 79,
+ 31,
+ 88
+ ],
+ "category_id": 19,
+ "id": 5154,
+ "area": 2728,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2624,
+ "bbox": [
+ 255,
+ 217,
+ 111,
+ 112
+ ],
+ "category_id": 16,
+ "id": 5155,
+ "area": 12432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2625,
+ "bbox": [
+ 124,
+ 220,
+ 204,
+ 143
+ ],
+ "category_id": 19,
+ "id": 5156,
+ "area": 29172,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2626,
+ "bbox": [
+ 174,
+ 146,
+ 187,
+ 298
+ ],
+ "category_id": 10,
+ "id": 5157,
+ "area": 55726,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2627,
+ "bbox": [
+ 38,
+ 11,
+ 109,
+ 428
+ ],
+ "category_id": 16,
+ "id": 5158,
+ "area": 46652,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2628,
+ "bbox": [
+ 193,
+ 81,
+ 47,
+ 347
+ ],
+ "category_id": 19,
+ "id": 5159,
+ "area": 16309,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2629,
+ "bbox": [
+ 215,
+ 120,
+ 62,
+ 271
+ ],
+ "category_id": 19,
+ "id": 5160,
+ "area": 16802,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2630,
+ "bbox": [
+ 30,
+ 228,
+ 28,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5161,
+ "area": 336,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2630,
+ "bbox": [
+ 238,
+ 180,
+ 27,
+ 10
+ ],
+ "category_id": 4,
+ "id": 5162,
+ "area": 270,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2630,
+ "bbox": [
+ 347,
+ 284,
+ 21,
+ 15
+ ],
+ "category_id": 4,
+ "id": 5163,
+ "area": 315,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2630,
+ "bbox": [
+ 484,
+ 350,
+ 12,
+ 24
+ ],
+ "category_id": 4,
+ "id": 5164,
+ "area": 288,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2631,
+ "bbox": [
+ 249,
+ 155,
+ 190,
+ 175
+ ],
+ "category_id": 10,
+ "id": 5165,
+ "area": 33250,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2632,
+ "bbox": [
+ 62,
+ 331,
+ 33,
+ 148
+ ],
+ "category_id": 10,
+ "id": 5166,
+ "area": 4884,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2633,
+ "bbox": [
+ 22,
+ 147,
+ 158,
+ 36
+ ],
+ "category_id": 16,
+ "id": 5167,
+ "area": 5688,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2634,
+ "bbox": [
+ 208,
+ 272,
+ 131,
+ 151
+ ],
+ "category_id": 16,
+ "id": 5168,
+ "area": 19781,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2635,
+ "bbox": [
+ 175,
+ 95,
+ 43,
+ 263
+ ],
+ "category_id": 19,
+ "id": 5169,
+ "area": 11309,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2636,
+ "bbox": [
+ 159,
+ 65,
+ 196,
+ 420
+ ],
+ "category_id": 10,
+ "id": 5170,
+ "area": 82320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2637,
+ "bbox": [
+ 113,
+ 38,
+ 172,
+ 238
+ ],
+ "category_id": 15,
+ "id": 5171,
+ "area": 40936,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2638,
+ "bbox": [
+ 390,
+ 195,
+ 15,
+ 57
+ ],
+ "category_id": 4,
+ "id": 5172,
+ "area": 855,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2638,
+ "bbox": [
+ 74,
+ 464,
+ 34,
+ 40
+ ],
+ "category_id": 4,
+ "id": 5173,
+ "area": 1360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2638,
+ "bbox": [
+ 442,
+ 450,
+ 22,
+ 38
+ ],
+ "category_id": 4,
+ "id": 5174,
+ "area": 836,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2639,
+ "bbox": [
+ 15,
+ 431,
+ 190,
+ 36
+ ],
+ "category_id": 10,
+ "id": 5175,
+ "area": 6840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2640,
+ "bbox": [
+ 198,
+ 110,
+ 241,
+ 325
+ ],
+ "category_id": 10,
+ "id": 5176,
+ "area": 78325,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2641,
+ "bbox": [
+ 268,
+ 248,
+ 121,
+ 38
+ ],
+ "category_id": 4,
+ "id": 5177,
+ "area": 4598,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2642,
+ "bbox": [
+ 34,
+ 45,
+ 61,
+ 33
+ ],
+ "category_id": 4,
+ "id": 5178,
+ "area": 2013,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2642,
+ "bbox": [
+ 170,
+ 421,
+ 59,
+ 32
+ ],
+ "category_id": 4,
+ "id": 5179,
+ "area": 1888,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2642,
+ "bbox": [
+ 388,
+ 85,
+ 48,
+ 19
+ ],
+ "category_id": 4,
+ "id": 5180,
+ "area": 912,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2643,
+ "bbox": [
+ 428,
+ 200,
+ 66,
+ 188
+ ],
+ "category_id": 19,
+ "id": 5181,
+ "area": 12408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2644,
+ "bbox": [
+ 295,
+ 132,
+ 123,
+ 70
+ ],
+ "category_id": 10,
+ "id": 5182,
+ "area": 8610,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2644,
+ "bbox": [
+ 204,
+ 353,
+ 133,
+ 32
+ ],
+ "category_id": 10,
+ "id": 5183,
+ "area": 4256,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2645,
+ "bbox": [
+ 218,
+ 208,
+ 107,
+ 133
+ ],
+ "category_id": 15,
+ "id": 5184,
+ "area": 14231,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2645,
+ "bbox": [
+ 160,
+ 339,
+ 107,
+ 133
+ ],
+ "category_id": 15,
+ "id": 5185,
+ "area": 14231,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2645,
+ "bbox": [
+ 467,
+ 118,
+ 37,
+ 123
+ ],
+ "category_id": 15,
+ "id": 5186,
+ "area": 4551,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2646,
+ "bbox": [
+ 43,
+ 209,
+ 424,
+ 103
+ ],
+ "category_id": 10,
+ "id": 5187,
+ "area": 43672,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2647,
+ "bbox": [
+ 55,
+ 227,
+ 68,
+ 41
+ ],
+ "category_id": 4,
+ "id": 5188,
+ "area": 2788,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2647,
+ "bbox": [
+ 451,
+ 225,
+ 47,
+ 36
+ ],
+ "category_id": 4,
+ "id": 5189,
+ "area": 1692,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2648,
+ "bbox": [
+ 231,
+ 201,
+ 83,
+ 99
+ ],
+ "category_id": 15,
+ "id": 5190,
+ "area": 8217,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2648,
+ "bbox": [
+ 290,
+ 300,
+ 80,
+ 99
+ ],
+ "category_id": 15,
+ "id": 5191,
+ "area": 7920,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2648,
+ "bbox": [
+ 386,
+ 0,
+ 39,
+ 96
+ ],
+ "category_id": 15,
+ "id": 5192,
+ "area": 3744,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2649,
+ "bbox": [
+ 231,
+ 184,
+ 101,
+ 89
+ ],
+ "category_id": 4,
+ "id": 5193,
+ "area": 8989,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2650,
+ "bbox": [
+ 11,
+ 154,
+ 19,
+ 17
+ ],
+ "category_id": 4,
+ "id": 5194,
+ "area": 323,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2650,
+ "bbox": [
+ 245,
+ 180,
+ 21,
+ 28
+ ],
+ "category_id": 4,
+ "id": 5195,
+ "area": 588,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2650,
+ "bbox": [
+ 84,
+ 341,
+ 25,
+ 13
+ ],
+ "category_id": 4,
+ "id": 5196,
+ "area": 325,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2650,
+ "bbox": [
+ 372,
+ 320,
+ 24,
+ 17
+ ],
+ "category_id": 4,
+ "id": 5197,
+ "area": 408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2650,
+ "bbox": [
+ 478,
+ 123,
+ 22,
+ 16
+ ],
+ "category_id": 4,
+ "id": 5198,
+ "area": 352,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2651,
+ "bbox": [
+ 106,
+ 382,
+ 98,
+ 99
+ ],
+ "category_id": 19,
+ "id": 5199,
+ "area": 9702,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2652,
+ "bbox": [
+ 125,
+ 216,
+ 206,
+ 130
+ ],
+ "category_id": 19,
+ "id": 5200,
+ "area": 26780,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2653,
+ "bbox": [
+ 131,
+ 366,
+ 35,
+ 86
+ ],
+ "category_id": 4,
+ "id": 5201,
+ "area": 3010,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2653,
+ "bbox": [
+ 325,
+ 144,
+ 67,
+ 24
+ ],
+ "category_id": 4,
+ "id": 5202,
+ "area": 1608,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2654,
+ "bbox": [
+ 325,
+ 270,
+ 42,
+ 79
+ ],
+ "category_id": 4,
+ "id": 5203,
+ "area": 3318,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2655,
+ "bbox": [
+ 79,
+ 268,
+ 316,
+ 160
+ ],
+ "category_id": 19,
+ "id": 5204,
+ "area": 50560,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2656,
+ "bbox": [
+ 143,
+ 226,
+ 147,
+ 50
+ ],
+ "category_id": 19,
+ "id": 5205,
+ "area": 7350,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2657,
+ "bbox": [
+ 78,
+ 109,
+ 354,
+ 275
+ ],
+ "category_id": 16,
+ "id": 5206,
+ "area": 97350,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2658,
+ "bbox": [
+ 204,
+ 410,
+ 41,
+ 58
+ ],
+ "category_id": 4,
+ "id": 5207,
+ "area": 2378,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2658,
+ "bbox": [
+ 269,
+ 71,
+ 43,
+ 58
+ ],
+ "category_id": 4,
+ "id": 5208,
+ "area": 2494,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2659,
+ "bbox": [
+ 62,
+ 289,
+ 150,
+ 54
+ ],
+ "category_id": 10,
+ "id": 5209,
+ "area": 8100,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2660,
+ "bbox": [
+ 202,
+ 161,
+ 17,
+ 43
+ ],
+ "category_id": 4,
+ "id": 5210,
+ "area": 731,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2660,
+ "bbox": [
+ 369,
+ 65,
+ 24,
+ 47
+ ],
+ "category_id": 4,
+ "id": 5211,
+ "area": 1128,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2660,
+ "bbox": [
+ 257,
+ 446,
+ 20,
+ 46
+ ],
+ "category_id": 4,
+ "id": 5212,
+ "area": 920,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2661,
+ "bbox": [
+ 154,
+ 344,
+ 57,
+ 66
+ ],
+ "category_id": 16,
+ "id": 5213,
+ "area": 3762,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2662,
+ "bbox": [
+ 245,
+ 101,
+ 60,
+ 301
+ ],
+ "category_id": 16,
+ "id": 5214,
+ "area": 18060,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2663,
+ "bbox": [
+ 151,
+ 209,
+ 174,
+ 92
+ ],
+ "category_id": 19,
+ "id": 5215,
+ "area": 16008,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2664,
+ "bbox": [
+ 62,
+ 114,
+ 426,
+ 310
+ ],
+ "category_id": 10,
+ "id": 5216,
+ "area": 132060,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2665,
+ "bbox": [
+ 58,
+ 64,
+ 343,
+ 348
+ ],
+ "category_id": 10,
+ "id": 5217,
+ "area": 119364,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2666,
+ "bbox": [
+ 68,
+ 355,
+ 264,
+ 87
+ ],
+ "category_id": 10,
+ "id": 5218,
+ "area": 22968,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2667,
+ "bbox": [
+ 143,
+ 11,
+ 180,
+ 184
+ ],
+ "category_id": 15,
+ "id": 5219,
+ "area": 33120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2667,
+ "bbox": [
+ 141,
+ 299,
+ 185,
+ 180
+ ],
+ "category_id": 15,
+ "id": 5220,
+ "area": 33300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2668,
+ "bbox": [
+ 344,
+ 69,
+ 26,
+ 30
+ ],
+ "category_id": 4,
+ "id": 5221,
+ "area": 780,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2668,
+ "bbox": [
+ 479,
+ 177,
+ 24,
+ 24
+ ],
+ "category_id": 4,
+ "id": 5222,
+ "area": 576,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2668,
+ "bbox": [
+ 47,
+ 246,
+ 25,
+ 33
+ ],
+ "category_id": 4,
+ "id": 5223,
+ "area": 825,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2668,
+ "bbox": [
+ 19,
+ 394,
+ 26,
+ 25
+ ],
+ "category_id": 4,
+ "id": 5224,
+ "area": 650,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2669,
+ "bbox": [
+ 193,
+ 172,
+ 256,
+ 167
+ ],
+ "category_id": 16,
+ "id": 5225,
+ "area": 42752,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2670,
+ "bbox": [
+ 49,
+ 210,
+ 422,
+ 29
+ ],
+ "category_id": 10,
+ "id": 5226,
+ "area": 12238,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2671,
+ "bbox": [
+ 192,
+ 297,
+ 48,
+ 55
+ ],
+ "category_id": 16,
+ "id": 5227,
+ "area": 2640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2672,
+ "bbox": [
+ 124,
+ 96,
+ 327,
+ 337
+ ],
+ "category_id": 16,
+ "id": 5228,
+ "area": 110199,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2673,
+ "bbox": [
+ 480,
+ 91,
+ 29,
+ 22
+ ],
+ "category_id": 4,
+ "id": 5229,
+ "area": 638,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2673,
+ "bbox": [
+ 250,
+ 8,
+ 25,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5230,
+ "area": 275,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2673,
+ "bbox": [
+ 284,
+ 254,
+ 22,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5231,
+ "area": 264,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2673,
+ "bbox": [
+ 0,
+ 344,
+ 13,
+ 15
+ ],
+ "category_id": 4,
+ "id": 5232,
+ "area": 195,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2673,
+ "bbox": [
+ 281,
+ 502,
+ 19,
+ 10
+ ],
+ "category_id": 4,
+ "id": 5233,
+ "area": 190,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2674,
+ "bbox": [
+ 83,
+ 101,
+ 404,
+ 314
+ ],
+ "category_id": 15,
+ "id": 5234,
+ "area": 126856,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2675,
+ "bbox": [
+ 12,
+ 55,
+ 157,
+ 174
+ ],
+ "category_id": 10,
+ "id": 5235,
+ "area": 27318,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2676,
+ "bbox": [
+ 170,
+ 83,
+ 41,
+ 327
+ ],
+ "category_id": 19,
+ "id": 5236,
+ "area": 13407,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2677,
+ "bbox": [
+ 233,
+ 152,
+ 49,
+ 167
+ ],
+ "category_id": 19,
+ "id": 5237,
+ "area": 8183,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2678,
+ "bbox": [
+ 60,
+ 375,
+ 12,
+ 52
+ ],
+ "category_id": 4,
+ "id": 5238,
+ "area": 624,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2678,
+ "bbox": [
+ 76,
+ 93,
+ 16,
+ 48
+ ],
+ "category_id": 4,
+ "id": 5239,
+ "area": 768,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2678,
+ "bbox": [
+ 443,
+ 256,
+ 21,
+ 51
+ ],
+ "category_id": 4,
+ "id": 5240,
+ "area": 1071,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2679,
+ "bbox": [
+ 91,
+ 225,
+ 415,
+ 68
+ ],
+ "category_id": 10,
+ "id": 5241,
+ "area": 28220,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2680,
+ "bbox": [
+ 49,
+ 309,
+ 51,
+ 48
+ ],
+ "category_id": 4,
+ "id": 5242,
+ "area": 2448,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2680,
+ "bbox": [
+ 414,
+ 171,
+ 57,
+ 53
+ ],
+ "category_id": 4,
+ "id": 5243,
+ "area": 3021,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2681,
+ "bbox": [
+ 342,
+ 23,
+ 14,
+ 22
+ ],
+ "category_id": 4,
+ "id": 5244,
+ "area": 308,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2681,
+ "bbox": [
+ 234,
+ 126,
+ 16,
+ 25
+ ],
+ "category_id": 4,
+ "id": 5245,
+ "area": 400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2681,
+ "bbox": [
+ 76,
+ 283,
+ 23,
+ 18
+ ],
+ "category_id": 4,
+ "id": 5246,
+ "area": 414,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2681,
+ "bbox": [
+ 74,
+ 411,
+ 21,
+ 17
+ ],
+ "category_id": 4,
+ "id": 5247,
+ "area": 357,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2681,
+ "bbox": [
+ 206,
+ 488,
+ 19,
+ 18
+ ],
+ "category_id": 4,
+ "id": 5248,
+ "area": 342,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2681,
+ "bbox": [
+ 423,
+ 461,
+ 13,
+ 23
+ ],
+ "category_id": 4,
+ "id": 5249,
+ "area": 299,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2682,
+ "bbox": [
+ 272,
+ 34,
+ 16,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5250,
+ "area": 192,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2682,
+ "bbox": [
+ 178,
+ 174,
+ 15,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5251,
+ "area": 135,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2682,
+ "bbox": [
+ 52,
+ 325,
+ 19,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5252,
+ "area": 171,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2682,
+ "bbox": [
+ 261,
+ 297,
+ 18,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5253,
+ "area": 162,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2682,
+ "bbox": [
+ 65,
+ 493,
+ 18,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5254,
+ "area": 162,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2682,
+ "bbox": [
+ 412,
+ 466,
+ 23,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5255,
+ "area": 276,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2683,
+ "bbox": [
+ 129,
+ 19,
+ 20,
+ 13
+ ],
+ "category_id": 4,
+ "id": 5256,
+ "area": 260,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2683,
+ "bbox": [
+ 56,
+ 51,
+ 18,
+ 13
+ ],
+ "category_id": 4,
+ "id": 5257,
+ "area": 234,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2683,
+ "bbox": [
+ 14,
+ 165,
+ 19,
+ 13
+ ],
+ "category_id": 4,
+ "id": 5258,
+ "area": 247,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2683,
+ "bbox": [
+ 412,
+ 58,
+ 14,
+ 10
+ ],
+ "category_id": 4,
+ "id": 5259,
+ "area": 140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2683,
+ "bbox": [
+ 392,
+ 131,
+ 19,
+ 16
+ ],
+ "category_id": 4,
+ "id": 5260,
+ "area": 304,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2683,
+ "bbox": [
+ 462,
+ 215,
+ 34,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5261,
+ "area": 408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2683,
+ "bbox": [
+ 368,
+ 298,
+ 33,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5262,
+ "area": 363,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2683,
+ "bbox": [
+ 186,
+ 374,
+ 32,
+ 13
+ ],
+ "category_id": 4,
+ "id": 5263,
+ "area": 416,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2684,
+ "bbox": [
+ 83,
+ 243,
+ 75,
+ 84
+ ],
+ "category_id": 4,
+ "id": 5264,
+ "area": 6300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2684,
+ "bbox": [
+ 420,
+ 263,
+ 69,
+ 59
+ ],
+ "category_id": 4,
+ "id": 5265,
+ "area": 4071,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2685,
+ "bbox": [
+ 22,
+ 340,
+ 65,
+ 55
+ ],
+ "category_id": 4,
+ "id": 5266,
+ "area": 3575,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2685,
+ "bbox": [
+ 406,
+ 243,
+ 61,
+ 48
+ ],
+ "category_id": 4,
+ "id": 5267,
+ "area": 2928,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2686,
+ "bbox": [
+ 155,
+ 45,
+ 93,
+ 376
+ ],
+ "category_id": 19,
+ "id": 5268,
+ "area": 34968,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2687,
+ "bbox": [
+ 184,
+ 286,
+ 212,
+ 62
+ ],
+ "category_id": 10,
+ "id": 5269,
+ "area": 13144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2688,
+ "bbox": [
+ 204,
+ 116,
+ 52,
+ 48
+ ],
+ "category_id": 4,
+ "id": 5270,
+ "area": 2496,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2688,
+ "bbox": [
+ 338,
+ 428,
+ 48,
+ 48
+ ],
+ "category_id": 4,
+ "id": 5271,
+ "area": 2304,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2689,
+ "bbox": [
+ 145,
+ 317,
+ 179,
+ 92
+ ],
+ "category_id": 19,
+ "id": 5272,
+ "area": 16468,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2690,
+ "bbox": [
+ 36,
+ 14,
+ 22,
+ 13
+ ],
+ "category_id": 4,
+ "id": 5273,
+ "area": 286,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2690,
+ "bbox": [
+ 62,
+ 211,
+ 25,
+ 14
+ ],
+ "category_id": 4,
+ "id": 5274,
+ "area": 350,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2690,
+ "bbox": [
+ 240,
+ 184,
+ 24,
+ 20
+ ],
+ "category_id": 4,
+ "id": 5275,
+ "area": 480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2690,
+ "bbox": [
+ 435,
+ 40,
+ 22,
+ 13
+ ],
+ "category_id": 4,
+ "id": 5276,
+ "area": 286,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2690,
+ "bbox": [
+ 268,
+ 322,
+ 22,
+ 15
+ ],
+ "category_id": 4,
+ "id": 5277,
+ "area": 330,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2690,
+ "bbox": [
+ 293,
+ 449,
+ 25,
+ 18
+ ],
+ "category_id": 4,
+ "id": 5278,
+ "area": 450,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2691,
+ "bbox": [
+ 321,
+ 170,
+ 68,
+ 176
+ ],
+ "category_id": 10,
+ "id": 5279,
+ "area": 11968,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2692,
+ "bbox": [
+ 65,
+ 211,
+ 140,
+ 155
+ ],
+ "category_id": 10,
+ "id": 5280,
+ "area": 21700,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2693,
+ "bbox": [
+ 72,
+ 73,
+ 48,
+ 56
+ ],
+ "category_id": 4,
+ "id": 5281,
+ "area": 2688,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2693,
+ "bbox": [
+ 269,
+ 425,
+ 38,
+ 45
+ ],
+ "category_id": 4,
+ "id": 5282,
+ "area": 1710,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2693,
+ "bbox": [
+ 397,
+ 194,
+ 49,
+ 64
+ ],
+ "category_id": 4,
+ "id": 5283,
+ "area": 3136,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2694,
+ "bbox": [
+ 300,
+ 103,
+ 159,
+ 114
+ ],
+ "category_id": 15,
+ "id": 5284,
+ "area": 18126,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2694,
+ "bbox": [
+ 45,
+ 222,
+ 216,
+ 165
+ ],
+ "category_id": 15,
+ "id": 5285,
+ "area": 35640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2694,
+ "bbox": [
+ 271,
+ 281,
+ 162,
+ 122
+ ],
+ "category_id": 15,
+ "id": 5286,
+ "area": 19764,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2695,
+ "bbox": [
+ 68,
+ 44,
+ 24,
+ 31
+ ],
+ "category_id": 4,
+ "id": 5287,
+ "area": 744,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2695,
+ "bbox": [
+ 60,
+ 149,
+ 27,
+ 28
+ ],
+ "category_id": 4,
+ "id": 5288,
+ "area": 756,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2695,
+ "bbox": [
+ 142,
+ 233,
+ 27,
+ 26
+ ],
+ "category_id": 4,
+ "id": 5289,
+ "area": 702,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2695,
+ "bbox": [
+ 246,
+ 265,
+ 29,
+ 31
+ ],
+ "category_id": 4,
+ "id": 5290,
+ "area": 899,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2695,
+ "bbox": [
+ 434,
+ 212,
+ 28,
+ 29
+ ],
+ "category_id": 4,
+ "id": 5291,
+ "area": 812,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2695,
+ "bbox": [
+ 346,
+ 357,
+ 28,
+ 27
+ ],
+ "category_id": 4,
+ "id": 5292,
+ "area": 756,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2695,
+ "bbox": [
+ 346,
+ 463,
+ 28,
+ 28
+ ],
+ "category_id": 4,
+ "id": 5293,
+ "area": 784,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2696,
+ "bbox": [
+ 115,
+ 128,
+ 36,
+ 27
+ ],
+ "category_id": 4,
+ "id": 5294,
+ "area": 972,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2696,
+ "bbox": [
+ 143,
+ 332,
+ 29,
+ 32
+ ],
+ "category_id": 4,
+ "id": 5295,
+ "area": 928,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2696,
+ "bbox": [
+ 147,
+ 426,
+ 45,
+ 14
+ ],
+ "category_id": 4,
+ "id": 5296,
+ "area": 630,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2696,
+ "bbox": [
+ 336,
+ 207,
+ 27,
+ 27
+ ],
+ "category_id": 4,
+ "id": 5297,
+ "area": 729,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2696,
+ "bbox": [
+ 430,
+ 312,
+ 46,
+ 22
+ ],
+ "category_id": 4,
+ "id": 5298,
+ "area": 1012,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2696,
+ "bbox": [
+ 312,
+ 24,
+ 40,
+ 25
+ ],
+ "category_id": 4,
+ "id": 5299,
+ "area": 1000,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2697,
+ "bbox": [
+ 156,
+ 232,
+ 66,
+ 57
+ ],
+ "category_id": 4,
+ "id": 5300,
+ "area": 3762,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2697,
+ "bbox": [
+ 410,
+ 285,
+ 71,
+ 64
+ ],
+ "category_id": 4,
+ "id": 5301,
+ "area": 4544,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2698,
+ "bbox": [
+ 212,
+ 250,
+ 91,
+ 137
+ ],
+ "category_id": 19,
+ "id": 5302,
+ "area": 12467,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2699,
+ "bbox": [
+ 321,
+ 80,
+ 43,
+ 46
+ ],
+ "category_id": 4,
+ "id": 5303,
+ "area": 1978,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2699,
+ "bbox": [
+ 223,
+ 426,
+ 34,
+ 45
+ ],
+ "category_id": 4,
+ "id": 5304,
+ "area": 1530,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2700,
+ "bbox": [
+ 23,
+ 357,
+ 21,
+ 50
+ ],
+ "category_id": 4,
+ "id": 5305,
+ "area": 1050,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2700,
+ "bbox": [
+ 233,
+ 236,
+ 21,
+ 52
+ ],
+ "category_id": 4,
+ "id": 5306,
+ "area": 1092,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2700,
+ "bbox": [
+ 469,
+ 321,
+ 27,
+ 53
+ ],
+ "category_id": 4,
+ "id": 5307,
+ "area": 1431,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2701,
+ "bbox": [
+ 152,
+ 0,
+ 121,
+ 98
+ ],
+ "category_id": 15,
+ "id": 5308,
+ "area": 11858,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2701,
+ "bbox": [
+ 163,
+ 154,
+ 226,
+ 211
+ ],
+ "category_id": 15,
+ "id": 5309,
+ "area": 47686,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2702,
+ "bbox": [
+ 54,
+ 190,
+ 8,
+ 8
+ ],
+ "category_id": 15,
+ "id": 5310,
+ "area": 64,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2702,
+ "bbox": [
+ 49,
+ 202,
+ 7,
+ 7
+ ],
+ "category_id": 15,
+ "id": 5311,
+ "area": 49,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2702,
+ "bbox": [
+ 42,
+ 219,
+ 12,
+ 12
+ ],
+ "category_id": 15,
+ "id": 5312,
+ "area": 144,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2703,
+ "bbox": [
+ 119,
+ 147,
+ 140,
+ 121
+ ],
+ "category_id": 15,
+ "id": 5313,
+ "area": 16940,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2704,
+ "bbox": [
+ 169,
+ 58,
+ 199,
+ 198
+ ],
+ "category_id": 15,
+ "id": 5314,
+ "area": 39402,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2705,
+ "bbox": [
+ 161,
+ 28,
+ 32,
+ 32
+ ],
+ "category_id": 15,
+ "id": 5315,
+ "area": 1024,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2706,
+ "bbox": [
+ 191,
+ 193,
+ 134,
+ 135
+ ],
+ "category_id": 10,
+ "id": 5316,
+ "area": 18090,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2707,
+ "bbox": [
+ 190,
+ 39,
+ 136,
+ 337
+ ],
+ "category_id": 19,
+ "id": 5317,
+ "area": 45832,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2708,
+ "bbox": [
+ 222,
+ 238,
+ 131,
+ 80
+ ],
+ "category_id": 4,
+ "id": 5318,
+ "area": 10480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2709,
+ "bbox": [
+ 78,
+ 406,
+ 80,
+ 33
+ ],
+ "category_id": 4,
+ "id": 5319,
+ "area": 2640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2709,
+ "bbox": [
+ 301,
+ 76,
+ 75,
+ 34
+ ],
+ "category_id": 4,
+ "id": 5320,
+ "area": 2550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2710,
+ "bbox": [
+ 39,
+ 195,
+ 388,
+ 301
+ ],
+ "category_id": 16,
+ "id": 5321,
+ "area": 116788,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2711,
+ "bbox": [
+ 137,
+ 168,
+ 258,
+ 149
+ ],
+ "category_id": 16,
+ "id": 5322,
+ "area": 38442,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2712,
+ "bbox": [
+ 129,
+ 103,
+ 151,
+ 102
+ ],
+ "category_id": 15,
+ "id": 5323,
+ "area": 15402,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2712,
+ "bbox": [
+ 291,
+ 98,
+ 153,
+ 99
+ ],
+ "category_id": 15,
+ "id": 5324,
+ "area": 15147,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2712,
+ "bbox": [
+ 284,
+ 234,
+ 152,
+ 106
+ ],
+ "category_id": 15,
+ "id": 5325,
+ "area": 16112,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2713,
+ "bbox": [
+ 80,
+ 87,
+ 49,
+ 46
+ ],
+ "category_id": 4,
+ "id": 5326,
+ "area": 2254,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2713,
+ "bbox": [
+ 433,
+ 331,
+ 63,
+ 38
+ ],
+ "category_id": 4,
+ "id": 5327,
+ "area": 2394,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2714,
+ "bbox": [
+ 286,
+ 96,
+ 67,
+ 48
+ ],
+ "category_id": 19,
+ "id": 5328,
+ "area": 3216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2715,
+ "bbox": [
+ 189,
+ 139,
+ 58,
+ 232
+ ],
+ "category_id": 19,
+ "id": 5329,
+ "area": 13456,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2716,
+ "bbox": [
+ 95,
+ 144,
+ 264,
+ 170
+ ],
+ "category_id": 16,
+ "id": 5330,
+ "area": 44880,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2717,
+ "bbox": [
+ 257,
+ 78,
+ 110,
+ 318
+ ],
+ "category_id": 10,
+ "id": 5331,
+ "area": 34980,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2718,
+ "bbox": [
+ 174,
+ 145,
+ 161,
+ 262
+ ],
+ "category_id": 16,
+ "id": 5332,
+ "area": 42182,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2719,
+ "bbox": [
+ 104,
+ 172,
+ 149,
+ 43
+ ],
+ "category_id": 15,
+ "id": 5333,
+ "area": 6407,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2719,
+ "bbox": [
+ 117,
+ 304,
+ 162,
+ 44
+ ],
+ "category_id": 15,
+ "id": 5334,
+ "area": 7128,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2720,
+ "bbox": [
+ 204,
+ 122,
+ 75,
+ 19
+ ],
+ "category_id": 4,
+ "id": 5335,
+ "area": 1425,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2720,
+ "bbox": [
+ 225,
+ 406,
+ 66,
+ 25
+ ],
+ "category_id": 4,
+ "id": 5336,
+ "area": 1650,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2721,
+ "bbox": [
+ 112,
+ 250,
+ 80,
+ 89
+ ],
+ "category_id": 15,
+ "id": 5337,
+ "area": 7120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2721,
+ "bbox": [
+ 216,
+ 214,
+ 93,
+ 107
+ ],
+ "category_id": 15,
+ "id": 5338,
+ "area": 9951,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2722,
+ "bbox": [
+ 323,
+ 53,
+ 109,
+ 144
+ ],
+ "category_id": 15,
+ "id": 5339,
+ "area": 15696,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2722,
+ "bbox": [
+ 161,
+ 177,
+ 94,
+ 131
+ ],
+ "category_id": 15,
+ "id": 5340,
+ "area": 12314,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2722,
+ "bbox": [
+ 177,
+ 299,
+ 98,
+ 136
+ ],
+ "category_id": 15,
+ "id": 5341,
+ "area": 13328,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2722,
+ "bbox": [
+ 391,
+ 282,
+ 33,
+ 108
+ ],
+ "category_id": 15,
+ "id": 5342,
+ "area": 3564,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2723,
+ "bbox": [
+ 216,
+ 51,
+ 144,
+ 172
+ ],
+ "category_id": 15,
+ "id": 5343,
+ "area": 24768,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2723,
+ "bbox": [
+ 220,
+ 255,
+ 148,
+ 180
+ ],
+ "category_id": 15,
+ "id": 5344,
+ "area": 26640,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2724,
+ "bbox": [
+ 215,
+ 115,
+ 119,
+ 119
+ ],
+ "category_id": 15,
+ "id": 5345,
+ "area": 14161,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2724,
+ "bbox": [
+ 222,
+ 288,
+ 126,
+ 124
+ ],
+ "category_id": 15,
+ "id": 5346,
+ "area": 15624,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2725,
+ "bbox": [
+ 27,
+ 112,
+ 367,
+ 332
+ ],
+ "category_id": 10,
+ "id": 5347,
+ "area": 121844,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2726,
+ "bbox": [
+ 330,
+ 122,
+ 141,
+ 66
+ ],
+ "category_id": 10,
+ "id": 5348,
+ "area": 9306,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2727,
+ "bbox": [
+ 256,
+ 31,
+ 17,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5349,
+ "area": 187,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2727,
+ "bbox": [
+ 8,
+ 206,
+ 22,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5350,
+ "area": 242,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2727,
+ "bbox": [
+ 474,
+ 254,
+ 23,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5351,
+ "area": 276,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2727,
+ "bbox": [
+ 120,
+ 426,
+ 20,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5352,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2728,
+ "bbox": [
+ 210,
+ 230,
+ 189,
+ 164
+ ],
+ "category_id": 19,
+ "id": 5353,
+ "area": 30996,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2729,
+ "bbox": [
+ 225,
+ 221,
+ 185,
+ 198
+ ],
+ "category_id": 16,
+ "id": 5354,
+ "area": 36630,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2730,
+ "bbox": [
+ 170,
+ 330,
+ 142,
+ 41
+ ],
+ "category_id": 19,
+ "id": 5355,
+ "area": 5822,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2731,
+ "bbox": [
+ 190,
+ 120,
+ 139,
+ 141
+ ],
+ "category_id": 15,
+ "id": 5356,
+ "area": 19599,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2731,
+ "bbox": [
+ 184,
+ 301,
+ 139,
+ 147
+ ],
+ "category_id": 15,
+ "id": 5357,
+ "area": 20433,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2732,
+ "bbox": [
+ 122,
+ 122,
+ 49,
+ 34
+ ],
+ "category_id": 4,
+ "id": 5358,
+ "area": 1666,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2732,
+ "bbox": [
+ 392,
+ 328,
+ 56,
+ 34
+ ],
+ "category_id": 4,
+ "id": 5359,
+ "area": 1904,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2733,
+ "bbox": [
+ 46,
+ 378,
+ 16,
+ 43
+ ],
+ "category_id": 4,
+ "id": 5360,
+ "area": 688,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2733,
+ "bbox": [
+ 224,
+ 87,
+ 17,
+ 41
+ ],
+ "category_id": 4,
+ "id": 5361,
+ "area": 697,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2733,
+ "bbox": [
+ 464,
+ 378,
+ 18,
+ 40
+ ],
+ "category_id": 4,
+ "id": 5362,
+ "area": 720,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2734,
+ "bbox": [
+ 19,
+ 23,
+ 397,
+ 96
+ ],
+ "category_id": 19,
+ "id": 5363,
+ "area": 38112,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2735,
+ "bbox": [
+ 124,
+ 116,
+ 62,
+ 46
+ ],
+ "category_id": 4,
+ "id": 5364,
+ "area": 2852,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2735,
+ "bbox": [
+ 126,
+ 315,
+ 49,
+ 63
+ ],
+ "category_id": 4,
+ "id": 5365,
+ "area": 3087,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2735,
+ "bbox": [
+ 440,
+ 320,
+ 48,
+ 56
+ ],
+ "category_id": 4,
+ "id": 5366,
+ "area": 2688,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2736,
+ "bbox": [
+ 442,
+ 63,
+ 17,
+ 18
+ ],
+ "category_id": 4,
+ "id": 5367,
+ "area": 306,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2736,
+ "bbox": [
+ 434,
+ 152,
+ 15,
+ 27
+ ],
+ "category_id": 4,
+ "id": 5368,
+ "area": 405,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2736,
+ "bbox": [
+ 414,
+ 267,
+ 19,
+ 26
+ ],
+ "category_id": 4,
+ "id": 5369,
+ "area": 494,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2736,
+ "bbox": [
+ 389,
+ 366,
+ 14,
+ 22
+ ],
+ "category_id": 4,
+ "id": 5370,
+ "area": 308,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2736,
+ "bbox": [
+ 411,
+ 474,
+ 16,
+ 16
+ ],
+ "category_id": 4,
+ "id": 5371,
+ "area": 256,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2736,
+ "bbox": [
+ 32,
+ 61,
+ 18,
+ 20
+ ],
+ "category_id": 4,
+ "id": 5372,
+ "area": 360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2737,
+ "bbox": [
+ 154,
+ 133,
+ 32,
+ 31
+ ],
+ "category_id": 4,
+ "id": 5373,
+ "area": 992,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2737,
+ "bbox": [
+ 225,
+ 453,
+ 32,
+ 34
+ ],
+ "category_id": 4,
+ "id": 5374,
+ "area": 1088,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2737,
+ "bbox": [
+ 425,
+ 253,
+ 31,
+ 32
+ ],
+ "category_id": 4,
+ "id": 5375,
+ "area": 992,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2738,
+ "bbox": [
+ 60,
+ 122,
+ 416,
+ 314
+ ],
+ "category_id": 10,
+ "id": 5376,
+ "area": 130624,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2739,
+ "bbox": [
+ 204,
+ 110,
+ 89,
+ 294
+ ],
+ "category_id": 19,
+ "id": 5377,
+ "area": 26166,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2740,
+ "bbox": [
+ 6,
+ 60,
+ 19,
+ 10
+ ],
+ "category_id": 4,
+ "id": 5378,
+ "area": 190,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2740,
+ "bbox": [
+ 416,
+ 48,
+ 23,
+ 13
+ ],
+ "category_id": 4,
+ "id": 5379,
+ "area": 299,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2740,
+ "bbox": [
+ 134,
+ 220,
+ 28,
+ 15
+ ],
+ "category_id": 4,
+ "id": 5380,
+ "area": 420,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2740,
+ "bbox": [
+ 382,
+ 205,
+ 19,
+ 16
+ ],
+ "category_id": 4,
+ "id": 5381,
+ "area": 304,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2740,
+ "bbox": [
+ 250,
+ 389,
+ 20,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5382,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2740,
+ "bbox": [
+ 485,
+ 451,
+ 18,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5383,
+ "area": 162,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2741,
+ "bbox": [
+ 162,
+ 142,
+ 209,
+ 244
+ ],
+ "category_id": 16,
+ "id": 5384,
+ "area": 50996,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2742,
+ "bbox": [
+ 49,
+ 25,
+ 119,
+ 173
+ ],
+ "category_id": 10,
+ "id": 5385,
+ "area": 20587,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2743,
+ "bbox": [
+ 92,
+ 76,
+ 175,
+ 347
+ ],
+ "category_id": 16,
+ "id": 5386,
+ "area": 60725,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2744,
+ "bbox": [
+ 232,
+ 254,
+ 75,
+ 151
+ ],
+ "category_id": 19,
+ "id": 5387,
+ "area": 11325,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2745,
+ "bbox": [
+ 28,
+ 142,
+ 40,
+ 35
+ ],
+ "category_id": 4,
+ "id": 5388,
+ "area": 1400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2745,
+ "bbox": [
+ 451,
+ 273,
+ 43,
+ 38
+ ],
+ "category_id": 4,
+ "id": 5389,
+ "area": 1634,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2746,
+ "bbox": [
+ 152,
+ 190,
+ 192,
+ 217
+ ],
+ "category_id": 16,
+ "id": 5390,
+ "area": 41664,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2747,
+ "bbox": [
+ 109,
+ 263,
+ 161,
+ 27
+ ],
+ "category_id": 19,
+ "id": 5391,
+ "area": 4347,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2748,
+ "bbox": [
+ 75,
+ 304,
+ 200,
+ 192
+ ],
+ "category_id": 10,
+ "id": 5392,
+ "area": 38400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2749,
+ "bbox": [
+ 279,
+ 243,
+ 110,
+ 147
+ ],
+ "category_id": 4,
+ "id": 5393,
+ "area": 16170,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2750,
+ "bbox": [
+ 23,
+ 34,
+ 6,
+ 6
+ ],
+ "category_id": 4,
+ "id": 5394,
+ "area": 36,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2750,
+ "bbox": [
+ 93,
+ 216,
+ 5,
+ 7
+ ],
+ "category_id": 4,
+ "id": 5395,
+ "area": 35,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2750,
+ "bbox": [
+ 58,
+ 302,
+ 11,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5396,
+ "area": 99,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2750,
+ "bbox": [
+ 71,
+ 363,
+ 9,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5397,
+ "area": 108,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2750,
+ "bbox": [
+ 125,
+ 410,
+ 8,
+ 6
+ ],
+ "category_id": 4,
+ "id": 5398,
+ "area": 48,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2750,
+ "bbox": [
+ 199,
+ 78,
+ 8,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5399,
+ "area": 72,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2750,
+ "bbox": [
+ 262,
+ 86,
+ 9,
+ 6
+ ],
+ "category_id": 4,
+ "id": 5400,
+ "area": 54,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2750,
+ "bbox": [
+ 421,
+ 53,
+ 8,
+ 5
+ ],
+ "category_id": 4,
+ "id": 5401,
+ "area": 40,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2750,
+ "bbox": [
+ 476,
+ 265,
+ 10,
+ 8
+ ],
+ "category_id": 4,
+ "id": 5402,
+ "area": 80,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2750,
+ "bbox": [
+ 309,
+ 265,
+ 8,
+ 7
+ ],
+ "category_id": 4,
+ "id": 5403,
+ "area": 56,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2751,
+ "bbox": [
+ 22,
+ 324,
+ 91,
+ 136
+ ],
+ "category_id": 10,
+ "id": 5404,
+ "area": 12376,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2752,
+ "bbox": [
+ 48,
+ 195,
+ 402,
+ 181
+ ],
+ "category_id": 16,
+ "id": 5405,
+ "area": 72762,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2753,
+ "bbox": [
+ 75,
+ 48,
+ 14,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5406,
+ "area": 168,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2753,
+ "bbox": [
+ 56,
+ 129,
+ 15,
+ 15
+ ],
+ "category_id": 4,
+ "id": 5407,
+ "area": 225,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2753,
+ "bbox": [
+ 2,
+ 272,
+ 16,
+ 10
+ ],
+ "category_id": 4,
+ "id": 5408,
+ "area": 160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2753,
+ "bbox": [
+ 31,
+ 367,
+ 17,
+ 13
+ ],
+ "category_id": 4,
+ "id": 5409,
+ "area": 221,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2753,
+ "bbox": [
+ 0,
+ 472,
+ 9,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5410,
+ "area": 81,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2753,
+ "bbox": [
+ 320,
+ 21,
+ 17,
+ 7
+ ],
+ "category_id": 4,
+ "id": 5411,
+ "area": 119,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2753,
+ "bbox": [
+ 291,
+ 94,
+ 13,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5412,
+ "area": 156,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2753,
+ "bbox": [
+ 261,
+ 177,
+ 12,
+ 16
+ ],
+ "category_id": 4,
+ "id": 5413,
+ "area": 192,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2753,
+ "bbox": [
+ 213,
+ 245,
+ 17,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5414,
+ "area": 153,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2753,
+ "bbox": [
+ 132,
+ 290,
+ 15,
+ 10
+ ],
+ "category_id": 4,
+ "id": 5415,
+ "area": 150,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2753,
+ "bbox": [
+ 494,
+ 44,
+ 15,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5416,
+ "area": 165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2753,
+ "bbox": [
+ 462,
+ 175,
+ 15,
+ 14
+ ],
+ "category_id": 4,
+ "id": 5417,
+ "area": 210,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2753,
+ "bbox": [
+ 490,
+ 337,
+ 12,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5418,
+ "area": 132,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2753,
+ "bbox": [
+ 419,
+ 448,
+ 14,
+ 14
+ ],
+ "category_id": 4,
+ "id": 5419,
+ "area": 196,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2753,
+ "bbox": [
+ 322,
+ 481,
+ 17,
+ 16
+ ],
+ "category_id": 4,
+ "id": 5420,
+ "area": 272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2754,
+ "bbox": [
+ 209,
+ 168,
+ 97,
+ 128
+ ],
+ "category_id": 15,
+ "id": 5421,
+ "area": 12416,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2754,
+ "bbox": [
+ 194,
+ 304,
+ 97,
+ 135
+ ],
+ "category_id": 15,
+ "id": 5422,
+ "area": 13095,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2755,
+ "bbox": [
+ 264,
+ 88,
+ 57,
+ 34
+ ],
+ "category_id": 4,
+ "id": 5423,
+ "area": 1938,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2755,
+ "bbox": [
+ 267,
+ 373,
+ 40,
+ 46
+ ],
+ "category_id": 4,
+ "id": 5424,
+ "area": 1840,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2756,
+ "bbox": [
+ 203,
+ 105,
+ 139,
+ 163
+ ],
+ "category_id": 15,
+ "id": 5425,
+ "area": 22657,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2756,
+ "bbox": [
+ 165,
+ 282,
+ 136,
+ 168
+ ],
+ "category_id": 15,
+ "id": 5426,
+ "area": 22848,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2757,
+ "bbox": [
+ 234,
+ 58,
+ 153,
+ 419
+ ],
+ "category_id": 10,
+ "id": 5427,
+ "area": 64107,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2758,
+ "bbox": [
+ 34,
+ 100,
+ 22,
+ 33
+ ],
+ "category_id": 4,
+ "id": 5428,
+ "area": 726,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2758,
+ "bbox": [
+ 34,
+ 290,
+ 26,
+ 31
+ ],
+ "category_id": 4,
+ "id": 5429,
+ "area": 806,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2758,
+ "bbox": [
+ 17,
+ 478,
+ 25,
+ 28
+ ],
+ "category_id": 4,
+ "id": 5430,
+ "area": 700,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2758,
+ "bbox": [
+ 453,
+ 17,
+ 25,
+ 29
+ ],
+ "category_id": 4,
+ "id": 5431,
+ "area": 725,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2758,
+ "bbox": [
+ 443,
+ 202,
+ 22,
+ 25
+ ],
+ "category_id": 4,
+ "id": 5432,
+ "area": 550,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2758,
+ "bbox": [
+ 460,
+ 401,
+ 27,
+ 22
+ ],
+ "category_id": 4,
+ "id": 5433,
+ "area": 594,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2759,
+ "bbox": [
+ 104,
+ 157,
+ 368,
+ 162
+ ],
+ "category_id": 10,
+ "id": 5434,
+ "area": 59616,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2760,
+ "bbox": [
+ 354,
+ 133,
+ 90,
+ 78
+ ],
+ "category_id": 15,
+ "id": 5435,
+ "area": 7020,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2761,
+ "bbox": [
+ 243,
+ 183,
+ 74,
+ 158
+ ],
+ "category_id": 19,
+ "id": 5436,
+ "area": 11692,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2762,
+ "bbox": [
+ 136,
+ 62,
+ 58,
+ 26
+ ],
+ "category_id": 4,
+ "id": 5437,
+ "area": 1508,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2762,
+ "bbox": [
+ 338,
+ 387,
+ 31,
+ 46
+ ],
+ "category_id": 4,
+ "id": 5438,
+ "area": 1426,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2763,
+ "bbox": [
+ 361,
+ 31,
+ 7,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5439,
+ "area": 63,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2763,
+ "bbox": [
+ 66,
+ 90,
+ 6,
+ 7
+ ],
+ "category_id": 4,
+ "id": 5440,
+ "area": 42,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2763,
+ "bbox": [
+ 139,
+ 65,
+ 7,
+ 6
+ ],
+ "category_id": 4,
+ "id": 5441,
+ "area": 42,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2763,
+ "bbox": [
+ 200,
+ 90,
+ 4,
+ 8
+ ],
+ "category_id": 4,
+ "id": 5442,
+ "area": 32,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2763,
+ "bbox": [
+ 287,
+ 108,
+ 6,
+ 7
+ ],
+ "category_id": 4,
+ "id": 5443,
+ "area": 42,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2763,
+ "bbox": [
+ 341,
+ 71,
+ 5,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5444,
+ "area": 45,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2763,
+ "bbox": [
+ 343,
+ 144,
+ 4,
+ 7
+ ],
+ "category_id": 4,
+ "id": 5445,
+ "area": 28,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2763,
+ "bbox": [
+ 471,
+ 223,
+ 8,
+ 8
+ ],
+ "category_id": 4,
+ "id": 5446,
+ "area": 64,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2763,
+ "bbox": [
+ 446,
+ 316,
+ 4,
+ 5
+ ],
+ "category_id": 4,
+ "id": 5447,
+ "area": 20,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2763,
+ "bbox": [
+ 424,
+ 368,
+ 6,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5448,
+ "area": 66,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2763,
+ "bbox": [
+ 406,
+ 464,
+ 6,
+ 10
+ ],
+ "category_id": 4,
+ "id": 5449,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2763,
+ "bbox": [
+ 320,
+ 474,
+ 5,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5450,
+ "area": 55,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2763,
+ "bbox": [
+ 119,
+ 458,
+ 7,
+ 7
+ ],
+ "category_id": 4,
+ "id": 5451,
+ "area": 49,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2763,
+ "bbox": [
+ 251,
+ 449,
+ 4,
+ 8
+ ],
+ "category_id": 4,
+ "id": 5452,
+ "area": 32,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2764,
+ "bbox": [
+ 23,
+ 296,
+ 60,
+ 51
+ ],
+ "category_id": 4,
+ "id": 5453,
+ "area": 3060,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2764,
+ "bbox": [
+ 392,
+ 172,
+ 54,
+ 72
+ ],
+ "category_id": 4,
+ "id": 5454,
+ "area": 3888,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2765,
+ "bbox": [
+ 165,
+ 122,
+ 164,
+ 301
+ ],
+ "category_id": 16,
+ "id": 5455,
+ "area": 49364,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2766,
+ "bbox": [
+ 96,
+ 82,
+ 143,
+ 144
+ ],
+ "category_id": 15,
+ "id": 5456,
+ "area": 20592,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2766,
+ "bbox": [
+ 99,
+ 297,
+ 145,
+ 144
+ ],
+ "category_id": 15,
+ "id": 5457,
+ "area": 20880,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2767,
+ "bbox": [
+ 395,
+ 58,
+ 15,
+ 10
+ ],
+ "category_id": 4,
+ "id": 5458,
+ "area": 150,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2767,
+ "bbox": [
+ 378,
+ 126,
+ 19,
+ 14
+ ],
+ "category_id": 4,
+ "id": 5459,
+ "area": 266,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2767,
+ "bbox": [
+ 368,
+ 206,
+ 17,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5460,
+ "area": 153,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2767,
+ "bbox": [
+ 327,
+ 290,
+ 16,
+ 10
+ ],
+ "category_id": 4,
+ "id": 5461,
+ "area": 160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2767,
+ "bbox": [
+ 211,
+ 323,
+ 13,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5462,
+ "area": 143,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2767,
+ "bbox": [
+ 40,
+ 327,
+ 15,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5463,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2767,
+ "bbox": [
+ 79,
+ 396,
+ 18,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5464,
+ "area": 162,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2767,
+ "bbox": [
+ 480,
+ 354,
+ 16,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5465,
+ "area": 176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2768,
+ "bbox": [
+ 78,
+ 91,
+ 20,
+ 49
+ ],
+ "category_id": 4,
+ "id": 5466,
+ "area": 980,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2768,
+ "bbox": [
+ 126,
+ 389,
+ 25,
+ 53
+ ],
+ "category_id": 4,
+ "id": 5467,
+ "area": 1325,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2768,
+ "bbox": [
+ 413,
+ 71,
+ 34,
+ 60
+ ],
+ "category_id": 4,
+ "id": 5468,
+ "area": 2040,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2769,
+ "bbox": [
+ 24,
+ 124,
+ 455,
+ 259
+ ],
+ "category_id": 19,
+ "id": 5469,
+ "area": 117845,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2770,
+ "bbox": [
+ 181,
+ 420,
+ 25,
+ 67
+ ],
+ "category_id": 4,
+ "id": 5470,
+ "area": 1675,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2770,
+ "bbox": [
+ 346,
+ 161,
+ 36,
+ 105
+ ],
+ "category_id": 4,
+ "id": 5471,
+ "area": 3780,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2771,
+ "bbox": [
+ 147,
+ 382,
+ 69,
+ 72
+ ],
+ "category_id": 16,
+ "id": 5472,
+ "area": 4968,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2771,
+ "bbox": [
+ 302,
+ 106,
+ 153,
+ 227
+ ],
+ "category_id": 16,
+ "id": 5473,
+ "area": 34731,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2772,
+ "bbox": [
+ 76,
+ 128,
+ 384,
+ 187
+ ],
+ "category_id": 10,
+ "id": 5474,
+ "area": 71808,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2773,
+ "bbox": [
+ 44,
+ 491,
+ 11,
+ 8
+ ],
+ "category_id": 4,
+ "id": 5475,
+ "area": 88,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2773,
+ "bbox": [
+ 123,
+ 456,
+ 12,
+ 7
+ ],
+ "category_id": 4,
+ "id": 5476,
+ "area": 84,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2773,
+ "bbox": [
+ 212,
+ 436,
+ 10,
+ 5
+ ],
+ "category_id": 4,
+ "id": 5477,
+ "area": 50,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2773,
+ "bbox": [
+ 199,
+ 172,
+ 13,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5478,
+ "area": 156,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2773,
+ "bbox": [
+ 280,
+ 157,
+ 13,
+ 10
+ ],
+ "category_id": 4,
+ "id": 5479,
+ "area": 130,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2773,
+ "bbox": [
+ 371,
+ 160,
+ 12,
+ 8
+ ],
+ "category_id": 4,
+ "id": 5480,
+ "area": 96,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2773,
+ "bbox": [
+ 193,
+ 259,
+ 9,
+ 6
+ ],
+ "category_id": 4,
+ "id": 5481,
+ "area": 54,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2773,
+ "bbox": [
+ 296,
+ 245,
+ 8,
+ 3
+ ],
+ "category_id": 4,
+ "id": 5482,
+ "area": 24,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2773,
+ "bbox": [
+ 382,
+ 256,
+ 11,
+ 5
+ ],
+ "category_id": 4,
+ "id": 5483,
+ "area": 55,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2773,
+ "bbox": [
+ 486,
+ 267,
+ 10,
+ 6
+ ],
+ "category_id": 4,
+ "id": 5484,
+ "area": 60,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2774,
+ "bbox": [
+ 164,
+ 248,
+ 223,
+ 57
+ ],
+ "category_id": 19,
+ "id": 5485,
+ "area": 12711,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2775,
+ "bbox": [
+ 239,
+ 86,
+ 29,
+ 55
+ ],
+ "category_id": 4,
+ "id": 5486,
+ "area": 1595,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2775,
+ "bbox": [
+ 220,
+ 427,
+ 24,
+ 56
+ ],
+ "category_id": 4,
+ "id": 5487,
+ "area": 1344,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2776,
+ "bbox": [
+ 206,
+ 65,
+ 57,
+ 263
+ ],
+ "category_id": 10,
+ "id": 5488,
+ "area": 14991,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2777,
+ "bbox": [
+ 253,
+ 126,
+ 26,
+ 26
+ ],
+ "category_id": 4,
+ "id": 5489,
+ "area": 676,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2777,
+ "bbox": [
+ 93,
+ 332,
+ 37,
+ 28
+ ],
+ "category_id": 4,
+ "id": 5490,
+ "area": 1036,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2777,
+ "bbox": [
+ 471,
+ 330,
+ 29,
+ 29
+ ],
+ "category_id": 4,
+ "id": 5491,
+ "area": 841,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2778,
+ "bbox": [
+ 88,
+ 95,
+ 47,
+ 121
+ ],
+ "category_id": 19,
+ "id": 5492,
+ "area": 5687,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2779,
+ "bbox": [
+ 208,
+ 254,
+ 87,
+ 90
+ ],
+ "category_id": 15,
+ "id": 5493,
+ "area": 7830,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2779,
+ "bbox": [
+ 0,
+ 298,
+ 46,
+ 74
+ ],
+ "category_id": 15,
+ "id": 5494,
+ "area": 3404,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2779,
+ "bbox": [
+ 5,
+ 268,
+ 43,
+ 57
+ ],
+ "category_id": 15,
+ "id": 5495,
+ "area": 2451,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2779,
+ "bbox": [
+ 40,
+ 272,
+ 47,
+ 54
+ ],
+ "category_id": 15,
+ "id": 5496,
+ "area": 2538,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2780,
+ "bbox": [
+ 28,
+ 101,
+ 219,
+ 233
+ ],
+ "category_id": 10,
+ "id": 5497,
+ "area": 51027,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2781,
+ "bbox": [
+ 99,
+ 223,
+ 132,
+ 97
+ ],
+ "category_id": 15,
+ "id": 5498,
+ "area": 12804,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2782,
+ "bbox": [
+ 40,
+ 136,
+ 67,
+ 71
+ ],
+ "category_id": 4,
+ "id": 5499,
+ "area": 4757,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2782,
+ "bbox": [
+ 413,
+ 337,
+ 57,
+ 93
+ ],
+ "category_id": 4,
+ "id": 5500,
+ "area": 5301,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2783,
+ "bbox": [
+ 87,
+ 196,
+ 180,
+ 92
+ ],
+ "category_id": 19,
+ "id": 5501,
+ "area": 16560,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2784,
+ "bbox": [
+ 288,
+ 428,
+ 35,
+ 19
+ ],
+ "category_id": 16,
+ "id": 5502,
+ "area": 665,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2785,
+ "bbox": [
+ 214,
+ 337,
+ 36,
+ 74
+ ],
+ "category_id": 4,
+ "id": 5503,
+ "area": 2664,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2786,
+ "bbox": [
+ 181,
+ 44,
+ 279,
+ 367
+ ],
+ "category_id": 10,
+ "id": 5504,
+ "area": 102393,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2787,
+ "bbox": [
+ 38,
+ 367,
+ 17,
+ 47
+ ],
+ "category_id": 4,
+ "id": 5505,
+ "area": 799,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2787,
+ "bbox": [
+ 149,
+ 201,
+ 25,
+ 47
+ ],
+ "category_id": 4,
+ "id": 5506,
+ "area": 1175,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2787,
+ "bbox": [
+ 438,
+ 184,
+ 28,
+ 55
+ ],
+ "category_id": 4,
+ "id": 5507,
+ "area": 1540,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2788,
+ "bbox": [
+ 87,
+ 34,
+ 26,
+ 17
+ ],
+ "category_id": 4,
+ "id": 5508,
+ "area": 442,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2788,
+ "bbox": [
+ 289,
+ 135,
+ 16,
+ 17
+ ],
+ "category_id": 4,
+ "id": 5509,
+ "area": 272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2788,
+ "bbox": [
+ 469,
+ 65,
+ 19,
+ 18
+ ],
+ "category_id": 4,
+ "id": 5510,
+ "area": 342,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2788,
+ "bbox": [
+ 65,
+ 146,
+ 18,
+ 18
+ ],
+ "category_id": 4,
+ "id": 5511,
+ "area": 324,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2788,
+ "bbox": [
+ 52,
+ 251,
+ 16,
+ 19
+ ],
+ "category_id": 4,
+ "id": 5512,
+ "area": 304,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2788,
+ "bbox": [
+ 33,
+ 343,
+ 20,
+ 15
+ ],
+ "category_id": 4,
+ "id": 5513,
+ "area": 300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2788,
+ "bbox": [
+ 275,
+ 462,
+ 18,
+ 18
+ ],
+ "category_id": 4,
+ "id": 5514,
+ "area": 324,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2789,
+ "bbox": [
+ 122,
+ 288,
+ 38,
+ 34
+ ],
+ "category_id": 15,
+ "id": 5515,
+ "area": 1292,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2789,
+ "bbox": [
+ 287,
+ 271,
+ 73,
+ 67
+ ],
+ "category_id": 15,
+ "id": 5516,
+ "area": 4891,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2789,
+ "bbox": [
+ 284,
+ 180,
+ 75,
+ 69
+ ],
+ "category_id": 15,
+ "id": 5517,
+ "area": 5175,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2789,
+ "bbox": [
+ 129,
+ 403,
+ 40,
+ 32
+ ],
+ "category_id": 15,
+ "id": 5518,
+ "area": 1280,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2790,
+ "bbox": [
+ 136,
+ 133,
+ 302,
+ 238
+ ],
+ "category_id": 16,
+ "id": 5519,
+ "area": 71876,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2791,
+ "bbox": [
+ 149,
+ 31,
+ 171,
+ 137
+ ],
+ "category_id": 10,
+ "id": 5520,
+ "area": 23427,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2792,
+ "bbox": [
+ 352,
+ 389,
+ 42,
+ 35
+ ],
+ "category_id": 15,
+ "id": 5521,
+ "area": 1470,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2792,
+ "bbox": [
+ 250,
+ 318,
+ 37,
+ 30
+ ],
+ "category_id": 15,
+ "id": 5522,
+ "area": 1110,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2792,
+ "bbox": [
+ 193,
+ 272,
+ 42,
+ 33
+ ],
+ "category_id": 15,
+ "id": 5523,
+ "area": 1386,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2792,
+ "bbox": [
+ 105,
+ 277,
+ 44,
+ 32
+ ],
+ "category_id": 15,
+ "id": 5524,
+ "area": 1408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2792,
+ "bbox": [
+ 120,
+ 287,
+ 46,
+ 33
+ ],
+ "category_id": 15,
+ "id": 5525,
+ "area": 1518,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2792,
+ "bbox": [
+ 58,
+ 358,
+ 43,
+ 35
+ ],
+ "category_id": 15,
+ "id": 5526,
+ "area": 1505,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2793,
+ "bbox": [
+ 35,
+ 284,
+ 190,
+ 157
+ ],
+ "category_id": 15,
+ "id": 5527,
+ "area": 29830,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2793,
+ "bbox": [
+ 275,
+ 295,
+ 180,
+ 163
+ ],
+ "category_id": 15,
+ "id": 5528,
+ "area": 29340,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2794,
+ "bbox": [
+ 170,
+ 195,
+ 110,
+ 70
+ ],
+ "category_id": 15,
+ "id": 5529,
+ "area": 7700,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2794,
+ "bbox": [
+ 280,
+ 182,
+ 109,
+ 74
+ ],
+ "category_id": 15,
+ "id": 5530,
+ "area": 8066,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2795,
+ "bbox": [
+ 167,
+ 122,
+ 44,
+ 269
+ ],
+ "category_id": 10,
+ "id": 5531,
+ "area": 11836,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2796,
+ "bbox": [
+ 411,
+ 26,
+ 45,
+ 59
+ ],
+ "category_id": 10,
+ "id": 5532,
+ "area": 2655,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2797,
+ "bbox": [
+ 116,
+ 197,
+ 41,
+ 45
+ ],
+ "category_id": 16,
+ "id": 5533,
+ "area": 1845,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2797,
+ "bbox": [
+ 102,
+ 91,
+ 43,
+ 45
+ ],
+ "category_id": 16,
+ "id": 5534,
+ "area": 1935,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2798,
+ "bbox": [
+ 256,
+ 363,
+ 88,
+ 114
+ ],
+ "category_id": 15,
+ "id": 5535,
+ "area": 10032,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2798,
+ "bbox": [
+ 236,
+ 242,
+ 90,
+ 112
+ ],
+ "category_id": 15,
+ "id": 5536,
+ "area": 10080,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2799,
+ "bbox": [
+ 154,
+ 28,
+ 121,
+ 109
+ ],
+ "category_id": 15,
+ "id": 5537,
+ "area": 13189,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2799,
+ "bbox": [
+ 199,
+ 179,
+ 119,
+ 114
+ ],
+ "category_id": 15,
+ "id": 5538,
+ "area": 13566,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2799,
+ "bbox": [
+ 239,
+ 333,
+ 128,
+ 114
+ ],
+ "category_id": 15,
+ "id": 5539,
+ "area": 14592,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2800,
+ "bbox": [
+ 122,
+ 191,
+ 161,
+ 98
+ ],
+ "category_id": 19,
+ "id": 5540,
+ "area": 15778,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2801,
+ "bbox": [
+ 97,
+ 172,
+ 292,
+ 113
+ ],
+ "category_id": 19,
+ "id": 5541,
+ "area": 32996,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2802,
+ "bbox": [
+ 223,
+ 55,
+ 50,
+ 364
+ ],
+ "category_id": 10,
+ "id": 5542,
+ "area": 18200,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2803,
+ "bbox": [
+ 54,
+ 49,
+ 102,
+ 148
+ ],
+ "category_id": 10,
+ "id": 5543,
+ "area": 15096,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2804,
+ "bbox": [
+ 162,
+ 170,
+ 139,
+ 182
+ ],
+ "category_id": 16,
+ "id": 5544,
+ "area": 25298,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2804,
+ "bbox": [
+ 375,
+ 329,
+ 38,
+ 28
+ ],
+ "category_id": 16,
+ "id": 5545,
+ "area": 1064,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2805,
+ "bbox": [
+ 89,
+ 113,
+ 330,
+ 303
+ ],
+ "category_id": 10,
+ "id": 5546,
+ "area": 99990,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2806,
+ "bbox": [
+ 204,
+ 116,
+ 181,
+ 300
+ ],
+ "category_id": 16,
+ "id": 5547,
+ "area": 54300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2807,
+ "bbox": [
+ 64,
+ 140,
+ 413,
+ 200
+ ],
+ "category_id": 10,
+ "id": 5548,
+ "area": 82600,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2808,
+ "bbox": [
+ 149,
+ 138,
+ 127,
+ 133
+ ],
+ "category_id": 15,
+ "id": 5549,
+ "area": 16891,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2808,
+ "bbox": [
+ 268,
+ 239,
+ 114,
+ 118
+ ],
+ "category_id": 15,
+ "id": 5550,
+ "area": 13452,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2809,
+ "bbox": [
+ 231,
+ 147,
+ 37,
+ 202
+ ],
+ "category_id": 19,
+ "id": 5551,
+ "area": 7474,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2810,
+ "bbox": [
+ 46,
+ 155,
+ 447,
+ 97
+ ],
+ "category_id": 10,
+ "id": 5552,
+ "area": 43359,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2811,
+ "bbox": [
+ 243,
+ 216,
+ 78,
+ 104
+ ],
+ "category_id": 15,
+ "id": 5553,
+ "area": 8112,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2811,
+ "bbox": [
+ 0,
+ 304,
+ 42,
+ 124
+ ],
+ "category_id": 15,
+ "id": 5554,
+ "area": 5208,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2812,
+ "bbox": [
+ 47,
+ 118,
+ 77,
+ 80
+ ],
+ "category_id": 15,
+ "id": 5555,
+ "area": 6160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2812,
+ "bbox": [
+ 133,
+ 118,
+ 82,
+ 93
+ ],
+ "category_id": 15,
+ "id": 5556,
+ "area": 7626,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2813,
+ "bbox": [
+ 35,
+ 160,
+ 56,
+ 48
+ ],
+ "category_id": 4,
+ "id": 5557,
+ "area": 2688,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2813,
+ "bbox": [
+ 176,
+ 379,
+ 55,
+ 52
+ ],
+ "category_id": 4,
+ "id": 5558,
+ "area": 2860,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2813,
+ "bbox": [
+ 382,
+ 339,
+ 56,
+ 52
+ ],
+ "category_id": 4,
+ "id": 5559,
+ "area": 2912,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2814,
+ "bbox": [
+ 14,
+ 149,
+ 485,
+ 188
+ ],
+ "category_id": 10,
+ "id": 5560,
+ "area": 91180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2815,
+ "bbox": [
+ 240,
+ 179,
+ 60,
+ 284
+ ],
+ "category_id": 16,
+ "id": 5561,
+ "area": 17040,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2816,
+ "bbox": [
+ 257,
+ 238,
+ 36,
+ 69
+ ],
+ "category_id": 4,
+ "id": 5562,
+ "area": 2484,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2817,
+ "bbox": [
+ 74,
+ 261,
+ 365,
+ 146
+ ],
+ "category_id": 10,
+ "id": 5563,
+ "area": 53290,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2818,
+ "bbox": [
+ 62,
+ 140,
+ 373,
+ 254
+ ],
+ "category_id": 10,
+ "id": 5564,
+ "area": 94742,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2819,
+ "bbox": [
+ 125,
+ 193,
+ 245,
+ 126
+ ],
+ "category_id": 19,
+ "id": 5565,
+ "area": 30870,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2820,
+ "bbox": [
+ 296,
+ 131,
+ 118,
+ 206
+ ],
+ "category_id": 16,
+ "id": 5566,
+ "area": 24308,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2821,
+ "bbox": [
+ 40,
+ 196,
+ 58,
+ 35
+ ],
+ "category_id": 4,
+ "id": 5567,
+ "area": 2030,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2821,
+ "bbox": [
+ 312,
+ 407,
+ 59,
+ 40
+ ],
+ "category_id": 4,
+ "id": 5568,
+ "area": 2360,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2821,
+ "bbox": [
+ 424,
+ 90,
+ 58,
+ 36
+ ],
+ "category_id": 4,
+ "id": 5569,
+ "area": 2088,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2822,
+ "bbox": [
+ 94,
+ 148,
+ 298,
+ 195
+ ],
+ "category_id": 10,
+ "id": 5570,
+ "area": 58110,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2823,
+ "bbox": [
+ 267,
+ 73,
+ 96,
+ 231
+ ],
+ "category_id": 10,
+ "id": 5571,
+ "area": 22176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2824,
+ "bbox": [
+ 193,
+ 204,
+ 130,
+ 141
+ ],
+ "category_id": 16,
+ "id": 5572,
+ "area": 18330,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2825,
+ "bbox": [
+ 111,
+ 163,
+ 125,
+ 131
+ ],
+ "category_id": 15,
+ "id": 5573,
+ "area": 16375,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2825,
+ "bbox": [
+ 281,
+ 249,
+ 134,
+ 132
+ ],
+ "category_id": 15,
+ "id": 5574,
+ "area": 17688,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2826,
+ "bbox": [
+ 144,
+ 22,
+ 28,
+ 17
+ ],
+ "category_id": 4,
+ "id": 5575,
+ "area": 476,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2826,
+ "bbox": [
+ 22,
+ 213,
+ 27,
+ 17
+ ],
+ "category_id": 4,
+ "id": 5576,
+ "area": 459,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2826,
+ "bbox": [
+ 227,
+ 134,
+ 30,
+ 19
+ ],
+ "category_id": 4,
+ "id": 5577,
+ "area": 570,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2826,
+ "bbox": [
+ 401,
+ 97,
+ 29,
+ 17
+ ],
+ "category_id": 4,
+ "id": 5578,
+ "area": 493,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2826,
+ "bbox": [
+ 261,
+ 392,
+ 29,
+ 13
+ ],
+ "category_id": 4,
+ "id": 5579,
+ "area": 377,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2826,
+ "bbox": [
+ 447,
+ 346,
+ 32,
+ 18
+ ],
+ "category_id": 4,
+ "id": 5580,
+ "area": 576,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2827,
+ "bbox": [
+ 332,
+ 62,
+ 81,
+ 69
+ ],
+ "category_id": 16,
+ "id": 5581,
+ "area": 5589,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2828,
+ "bbox": [
+ 68,
+ 131,
+ 138,
+ 142
+ ],
+ "category_id": 19,
+ "id": 5582,
+ "area": 19596,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2829,
+ "bbox": [
+ 344,
+ 352,
+ 121,
+ 104
+ ],
+ "category_id": 10,
+ "id": 5583,
+ "area": 12584,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2830,
+ "bbox": [
+ 30,
+ 257,
+ 98,
+ 217
+ ],
+ "category_id": 10,
+ "id": 5584,
+ "area": 21266,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2831,
+ "bbox": [
+ 79,
+ 298,
+ 188,
+ 98
+ ],
+ "category_id": 16,
+ "id": 5585,
+ "area": 18424,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2832,
+ "bbox": [
+ 44,
+ 28,
+ 30,
+ 16
+ ],
+ "category_id": 4,
+ "id": 5586,
+ "area": 480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2832,
+ "bbox": [
+ 108,
+ 162,
+ 32,
+ 14
+ ],
+ "category_id": 4,
+ "id": 5587,
+ "area": 448,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2832,
+ "bbox": [
+ 274,
+ 53,
+ 29,
+ 13
+ ],
+ "category_id": 4,
+ "id": 5588,
+ "area": 377,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2832,
+ "bbox": [
+ 404,
+ 192,
+ 31,
+ 15
+ ],
+ "category_id": 4,
+ "id": 5589,
+ "area": 465,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2832,
+ "bbox": [
+ 232,
+ 318,
+ 32,
+ 15
+ ],
+ "category_id": 4,
+ "id": 5590,
+ "area": 480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2832,
+ "bbox": [
+ 349,
+ 453,
+ 27,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5591,
+ "area": 324,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2833,
+ "bbox": [
+ 64,
+ 135,
+ 391,
+ 188
+ ],
+ "category_id": 10,
+ "id": 5592,
+ "area": 73508,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2834,
+ "bbox": [
+ 204,
+ 102,
+ 86,
+ 119
+ ],
+ "category_id": 15,
+ "id": 5593,
+ "area": 10234,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2834,
+ "bbox": [
+ 156,
+ 224,
+ 71,
+ 98
+ ],
+ "category_id": 15,
+ "id": 5594,
+ "area": 6958,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2835,
+ "bbox": [
+ 154,
+ 118,
+ 121,
+ 25
+ ],
+ "category_id": 15,
+ "id": 5595,
+ "area": 3025,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2835,
+ "bbox": [
+ 359,
+ 144,
+ 121,
+ 24
+ ],
+ "category_id": 15,
+ "id": 5596,
+ "area": 2904,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2835,
+ "bbox": [
+ 28,
+ 201,
+ 112,
+ 76
+ ],
+ "category_id": 15,
+ "id": 5597,
+ "area": 8512,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2835,
+ "bbox": [
+ 140,
+ 215,
+ 108,
+ 79
+ ],
+ "category_id": 15,
+ "id": 5598,
+ "area": 8532,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2835,
+ "bbox": [
+ 256,
+ 232,
+ 112,
+ 80
+ ],
+ "category_id": 15,
+ "id": 5599,
+ "area": 8960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2835,
+ "bbox": [
+ 371,
+ 247,
+ 109,
+ 78
+ ],
+ "category_id": 15,
+ "id": 5600,
+ "area": 8502,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2836,
+ "bbox": [
+ 254,
+ 282,
+ 93,
+ 72
+ ],
+ "category_id": 4,
+ "id": 5601,
+ "area": 6696,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2837,
+ "bbox": [
+ 247,
+ 258,
+ 99,
+ 63
+ ],
+ "category_id": 4,
+ "id": 5602,
+ "area": 6237,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2838,
+ "bbox": [
+ 87,
+ 101,
+ 148,
+ 306
+ ],
+ "category_id": 19,
+ "id": 5603,
+ "area": 45288,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2839,
+ "bbox": [
+ 184,
+ 169,
+ 123,
+ 68
+ ],
+ "category_id": 15,
+ "id": 5604,
+ "area": 8364,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2839,
+ "bbox": [
+ 238,
+ 250,
+ 124,
+ 72
+ ],
+ "category_id": 15,
+ "id": 5605,
+ "area": 8928,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2840,
+ "bbox": [
+ 112,
+ 73,
+ 28,
+ 24
+ ],
+ "category_id": 4,
+ "id": 5606,
+ "area": 672,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2840,
+ "bbox": [
+ 179,
+ 423,
+ 32,
+ 29
+ ],
+ "category_id": 4,
+ "id": 5607,
+ "area": 928,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2840,
+ "bbox": [
+ 407,
+ 272,
+ 39,
+ 33
+ ],
+ "category_id": 4,
+ "id": 5608,
+ "area": 1287,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2841,
+ "bbox": [
+ 112,
+ 188,
+ 236,
+ 231
+ ],
+ "category_id": 16,
+ "id": 5609,
+ "area": 54516,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2842,
+ "bbox": [
+ 105,
+ 60,
+ 17,
+ 48
+ ],
+ "category_id": 4,
+ "id": 5610,
+ "area": 816,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2842,
+ "bbox": [
+ 128,
+ 401,
+ 17,
+ 51
+ ],
+ "category_id": 4,
+ "id": 5611,
+ "area": 867,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2842,
+ "bbox": [
+ 433,
+ 221,
+ 15,
+ 58
+ ],
+ "category_id": 4,
+ "id": 5612,
+ "area": 870,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2843,
+ "bbox": [
+ 42,
+ 186,
+ 427,
+ 107
+ ],
+ "category_id": 10,
+ "id": 5613,
+ "area": 45689,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2844,
+ "bbox": [
+ 221,
+ 180,
+ 33,
+ 106
+ ],
+ "category_id": 4,
+ "id": 5614,
+ "area": 3498,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2845,
+ "bbox": [
+ 45,
+ 231,
+ 426,
+ 37
+ ],
+ "category_id": 19,
+ "id": 5615,
+ "area": 15762,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2846,
+ "bbox": [
+ 117,
+ 112,
+ 150,
+ 163
+ ],
+ "category_id": 15,
+ "id": 5616,
+ "area": 24450,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2846,
+ "bbox": [
+ 267,
+ 147,
+ 154,
+ 167
+ ],
+ "category_id": 15,
+ "id": 5617,
+ "area": 25718,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2847,
+ "bbox": [
+ 38,
+ 96,
+ 180,
+ 183
+ ],
+ "category_id": 15,
+ "id": 5618,
+ "area": 32940,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2847,
+ "bbox": [
+ 275,
+ 184,
+ 178,
+ 192
+ ],
+ "category_id": 15,
+ "id": 5619,
+ "area": 34176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2848,
+ "bbox": [
+ 232,
+ 245,
+ 132,
+ 44
+ ],
+ "category_id": 19,
+ "id": 5620,
+ "area": 5808,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2849,
+ "bbox": [
+ 63,
+ 218,
+ 141,
+ 158
+ ],
+ "category_id": 15,
+ "id": 5621,
+ "area": 22278,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2849,
+ "bbox": [
+ 262,
+ 222,
+ 141,
+ 158
+ ],
+ "category_id": 15,
+ "id": 5622,
+ "area": 22278,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2850,
+ "bbox": [
+ 140,
+ 65,
+ 18,
+ 21
+ ],
+ "category_id": 4,
+ "id": 5623,
+ "area": 378,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2850,
+ "bbox": [
+ 103,
+ 190,
+ 22,
+ 28
+ ],
+ "category_id": 4,
+ "id": 5624,
+ "area": 616,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2850,
+ "bbox": [
+ 121,
+ 304,
+ 28,
+ 26
+ ],
+ "category_id": 4,
+ "id": 5625,
+ "area": 728,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2850,
+ "bbox": [
+ 56,
+ 474,
+ 19,
+ 22
+ ],
+ "category_id": 4,
+ "id": 5626,
+ "area": 418,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2850,
+ "bbox": [
+ 417,
+ 10,
+ 25,
+ 18
+ ],
+ "category_id": 4,
+ "id": 5627,
+ "area": 450,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2850,
+ "bbox": [
+ 440,
+ 140,
+ 24,
+ 20
+ ],
+ "category_id": 4,
+ "id": 5628,
+ "area": 480,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2850,
+ "bbox": [
+ 448,
+ 316,
+ 23,
+ 23
+ ],
+ "category_id": 4,
+ "id": 5629,
+ "area": 529,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2850,
+ "bbox": [
+ 273,
+ 410,
+ 23,
+ 21
+ ],
+ "category_id": 4,
+ "id": 5630,
+ "area": 483,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2851,
+ "bbox": [
+ 83,
+ 262,
+ 249,
+ 41
+ ],
+ "category_id": 19,
+ "id": 5631,
+ "area": 10209,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2852,
+ "bbox": [
+ 108,
+ 88,
+ 11,
+ 16
+ ],
+ "category_id": 4,
+ "id": 5632,
+ "area": 176,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2852,
+ "bbox": [
+ 67,
+ 200,
+ 10,
+ 14
+ ],
+ "category_id": 4,
+ "id": 5633,
+ "area": 140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2852,
+ "bbox": [
+ 43,
+ 302,
+ 9,
+ 13
+ ],
+ "category_id": 4,
+ "id": 5634,
+ "area": 117,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2852,
+ "bbox": [
+ 476,
+ 304,
+ 12,
+ 17
+ ],
+ "category_id": 4,
+ "id": 5635,
+ "area": 204,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2852,
+ "bbox": [
+ 453,
+ 364,
+ 12,
+ 17
+ ],
+ "category_id": 4,
+ "id": 5636,
+ "area": 204,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2853,
+ "bbox": [
+ 115,
+ 252,
+ 258,
+ 99
+ ],
+ "category_id": 16,
+ "id": 5637,
+ "area": 25542,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2854,
+ "bbox": [
+ 207,
+ 185,
+ 130,
+ 233
+ ],
+ "category_id": 16,
+ "id": 5638,
+ "area": 30290,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2855,
+ "bbox": [
+ 134,
+ 108,
+ 356,
+ 304
+ ],
+ "category_id": 16,
+ "id": 5639,
+ "area": 108224,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2856,
+ "bbox": [
+ 323,
+ 42,
+ 23,
+ 10
+ ],
+ "category_id": 4,
+ "id": 5640,
+ "area": 230,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2856,
+ "bbox": [
+ 340,
+ 149,
+ 34,
+ 18
+ ],
+ "category_id": 4,
+ "id": 5641,
+ "area": 612,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2856,
+ "bbox": [
+ 129,
+ 213,
+ 32,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5642,
+ "area": 352,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2856,
+ "bbox": [
+ 158,
+ 347,
+ 32,
+ 17
+ ],
+ "category_id": 4,
+ "id": 5643,
+ "area": 544,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2856,
+ "bbox": [
+ 385,
+ 382,
+ 20,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5644,
+ "area": 240,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2856,
+ "bbox": [
+ 156,
+ 480,
+ 27,
+ 18
+ ],
+ "category_id": 4,
+ "id": 5645,
+ "area": 486,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2857,
+ "bbox": [
+ 81,
+ 151,
+ 347,
+ 83
+ ],
+ "category_id": 19,
+ "id": 5646,
+ "area": 28801,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2858,
+ "bbox": [
+ 215,
+ 149,
+ 63,
+ 57
+ ],
+ "category_id": 4,
+ "id": 5647,
+ "area": 3591,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2858,
+ "bbox": [
+ 83,
+ 425,
+ 61,
+ 51
+ ],
+ "category_id": 4,
+ "id": 5648,
+ "area": 3111,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2859,
+ "bbox": [
+ 69,
+ 203,
+ 102,
+ 88
+ ],
+ "category_id": 15,
+ "id": 5649,
+ "area": 8976,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2859,
+ "bbox": [
+ 195,
+ 211,
+ 106,
+ 93
+ ],
+ "category_id": 15,
+ "id": 5650,
+ "area": 9858,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2860,
+ "bbox": [
+ 150,
+ 108,
+ 83,
+ 48
+ ],
+ "category_id": 4,
+ "id": 5651,
+ "area": 3984,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2860,
+ "bbox": [
+ 321,
+ 411,
+ 90,
+ 39
+ ],
+ "category_id": 4,
+ "id": 5652,
+ "area": 3510,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2861,
+ "bbox": [
+ 175,
+ 126,
+ 225,
+ 259
+ ],
+ "category_id": 10,
+ "id": 5653,
+ "area": 58275,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2862,
+ "bbox": [
+ 192,
+ 193,
+ 108,
+ 137
+ ],
+ "category_id": 15,
+ "id": 5654,
+ "area": 14796,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2862,
+ "bbox": [
+ 211,
+ 344,
+ 110,
+ 136
+ ],
+ "category_id": 15,
+ "id": 5655,
+ "area": 14960,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2863,
+ "bbox": [
+ 38,
+ 238,
+ 446,
+ 22
+ ],
+ "category_id": 10,
+ "id": 5656,
+ "area": 9812,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2864,
+ "bbox": [
+ 102,
+ 187,
+ 302,
+ 95
+ ],
+ "category_id": 19,
+ "id": 5657,
+ "area": 28690,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2865,
+ "bbox": [
+ 33,
+ 33,
+ 291,
+ 105
+ ],
+ "category_id": 10,
+ "id": 5658,
+ "area": 30555,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2866,
+ "bbox": [
+ 284,
+ 261,
+ 109,
+ 98
+ ],
+ "category_id": 4,
+ "id": 5659,
+ "area": 10682,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2867,
+ "bbox": [
+ 69,
+ 229,
+ 411,
+ 217
+ ],
+ "category_id": 16,
+ "id": 5660,
+ "area": 89187,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2868,
+ "bbox": [
+ 28,
+ 97,
+ 20,
+ 20
+ ],
+ "category_id": 4,
+ "id": 5661,
+ "area": 400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2868,
+ "bbox": [
+ 36,
+ 187,
+ 16,
+ 21
+ ],
+ "category_id": 4,
+ "id": 5662,
+ "area": 336,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2868,
+ "bbox": [
+ 199,
+ 334,
+ 24,
+ 18
+ ],
+ "category_id": 4,
+ "id": 5663,
+ "area": 432,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2868,
+ "bbox": [
+ 426,
+ 215,
+ 18,
+ 17
+ ],
+ "category_id": 4,
+ "id": 5664,
+ "area": 306,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2868,
+ "bbox": [
+ 466,
+ 288,
+ 24,
+ 17
+ ],
+ "category_id": 4,
+ "id": 5665,
+ "area": 408,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2869,
+ "bbox": [
+ 446,
+ 396,
+ 14,
+ 14
+ ],
+ "category_id": 15,
+ "id": 5666,
+ "area": 196,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2869,
+ "bbox": [
+ 438,
+ 414,
+ 14,
+ 15
+ ],
+ "category_id": 15,
+ "id": 5667,
+ "area": 210,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2869,
+ "bbox": [
+ 463,
+ 403,
+ 17,
+ 16
+ ],
+ "category_id": 15,
+ "id": 5668,
+ "area": 272,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2869,
+ "bbox": [
+ 455,
+ 422,
+ 17,
+ 14
+ ],
+ "category_id": 15,
+ "id": 5669,
+ "area": 238,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2870,
+ "bbox": [
+ 267,
+ 266,
+ 66,
+ 74
+ ],
+ "category_id": 15,
+ "id": 5670,
+ "area": 4884,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2870,
+ "bbox": [
+ 197,
+ 323,
+ 71,
+ 74
+ ],
+ "category_id": 15,
+ "id": 5671,
+ "area": 5254,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2871,
+ "bbox": [
+ 183,
+ 72,
+ 50,
+ 261
+ ],
+ "category_id": 19,
+ "id": 5672,
+ "area": 13050,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2872,
+ "bbox": [
+ 107,
+ 190,
+ 315,
+ 101
+ ],
+ "category_id": 16,
+ "id": 5673,
+ "area": 31815,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2873,
+ "bbox": [
+ 192,
+ 110,
+ 118,
+ 313
+ ],
+ "category_id": 10,
+ "id": 5674,
+ "area": 36934,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2874,
+ "bbox": [
+ 108,
+ 185,
+ 334,
+ 60
+ ],
+ "category_id": 16,
+ "id": 5675,
+ "area": 20040,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2875,
+ "bbox": [
+ 45,
+ 126,
+ 323,
+ 332
+ ],
+ "category_id": 16,
+ "id": 5676,
+ "area": 107236,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2876,
+ "bbox": [
+ 22,
+ 10,
+ 17,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5677,
+ "area": 204,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2876,
+ "bbox": [
+ 49,
+ 172,
+ 16,
+ 14
+ ],
+ "category_id": 4,
+ "id": 5678,
+ "area": 224,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2876,
+ "bbox": [
+ 143,
+ 131,
+ 14,
+ 10
+ ],
+ "category_id": 4,
+ "id": 5679,
+ "area": 140,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2876,
+ "bbox": [
+ 316,
+ 115,
+ 15,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5680,
+ "area": 165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2876,
+ "bbox": [
+ 404,
+ 58,
+ 24,
+ 13
+ ],
+ "category_id": 4,
+ "id": 5681,
+ "area": 312,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2876,
+ "bbox": [
+ 467,
+ 152,
+ 17,
+ 13
+ ],
+ "category_id": 4,
+ "id": 5682,
+ "area": 221,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2876,
+ "bbox": [
+ 306,
+ 186,
+ 16,
+ 13
+ ],
+ "category_id": 4,
+ "id": 5683,
+ "area": 208,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2876,
+ "bbox": [
+ 404,
+ 227,
+ 21,
+ 14
+ ],
+ "category_id": 4,
+ "id": 5684,
+ "area": 294,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2876,
+ "bbox": [
+ 376,
+ 298,
+ 21,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5685,
+ "area": 252,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2876,
+ "bbox": [
+ 225,
+ 321,
+ 15,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5686,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2876,
+ "bbox": [
+ 92,
+ 293,
+ 18,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5687,
+ "area": 216,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2876,
+ "bbox": [
+ 49,
+ 423,
+ 18,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5688,
+ "area": 162,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2876,
+ "bbox": [
+ 193,
+ 392,
+ 15,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5689,
+ "area": 180,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2877,
+ "bbox": [
+ 42,
+ 426,
+ 148,
+ 50
+ ],
+ "category_id": 10,
+ "id": 5690,
+ "area": 7400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2878,
+ "bbox": [
+ 197,
+ 229,
+ 64,
+ 80
+ ],
+ "category_id": 10,
+ "id": 5691,
+ "area": 5120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2879,
+ "bbox": [
+ 11,
+ 224,
+ 369,
+ 259
+ ],
+ "category_id": 15,
+ "id": 5692,
+ "area": 95571,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2879,
+ "bbox": [
+ 259,
+ 60,
+ 80,
+ 39
+ ],
+ "category_id": 15,
+ "id": 5693,
+ "area": 3120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2880,
+ "bbox": [
+ 39,
+ 17,
+ 181,
+ 172
+ ],
+ "category_id": 10,
+ "id": 5694,
+ "area": 31132,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2881,
+ "bbox": [
+ 174,
+ 82,
+ 19,
+ 8
+ ],
+ "category_id": 4,
+ "id": 5695,
+ "area": 152,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2881,
+ "bbox": [
+ 27,
+ 166,
+ 19,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5696,
+ "area": 209,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2881,
+ "bbox": [
+ 276,
+ 199,
+ 22,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5697,
+ "area": 198,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2881,
+ "bbox": [
+ 456,
+ 128,
+ 25,
+ 12
+ ],
+ "category_id": 4,
+ "id": 5698,
+ "area": 300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2882,
+ "bbox": [
+ 433,
+ 236,
+ 13,
+ 39
+ ],
+ "category_id": 15,
+ "id": 5699,
+ "area": 507,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2883,
+ "bbox": [
+ 48,
+ 42,
+ 37,
+ 9
+ ],
+ "category_id": 4,
+ "id": 5700,
+ "area": 333,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2883,
+ "bbox": [
+ 21,
+ 289,
+ 25,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5701,
+ "area": 275,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2883,
+ "bbox": [
+ 198,
+ 412,
+ 28,
+ 22
+ ],
+ "category_id": 4,
+ "id": 5702,
+ "area": 616,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2883,
+ "bbox": [
+ 320,
+ 483,
+ 29,
+ 22
+ ],
+ "category_id": 4,
+ "id": 5703,
+ "area": 638,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2884,
+ "bbox": [
+ 119,
+ 182,
+ 306,
+ 137
+ ],
+ "category_id": 10,
+ "id": 5704,
+ "area": 41922,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2885,
+ "bbox": [
+ 140,
+ 218,
+ 199,
+ 90
+ ],
+ "category_id": 19,
+ "id": 5705,
+ "area": 17910,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2886,
+ "bbox": [
+ 211,
+ 401,
+ 59,
+ 52
+ ],
+ "category_id": 4,
+ "id": 5706,
+ "area": 3068,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2886,
+ "bbox": [
+ 262,
+ 59,
+ 54,
+ 53
+ ],
+ "category_id": 4,
+ "id": 5707,
+ "area": 2862,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2887,
+ "bbox": [
+ 102,
+ 165,
+ 237,
+ 296
+ ],
+ "category_id": 16,
+ "id": 5708,
+ "area": 70152,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2888,
+ "bbox": [
+ 275,
+ 60,
+ 118,
+ 343
+ ],
+ "category_id": 10,
+ "id": 5709,
+ "area": 40474,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2889,
+ "bbox": [
+ 195,
+ 133,
+ 70,
+ 172
+ ],
+ "category_id": 19,
+ "id": 5710,
+ "area": 12040,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2890,
+ "bbox": [
+ 113,
+ 346,
+ 48,
+ 86
+ ],
+ "category_id": 10,
+ "id": 5711,
+ "area": 4128,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2891,
+ "bbox": [
+ 154,
+ 244,
+ 216,
+ 73
+ ],
+ "category_id": 19,
+ "id": 5712,
+ "area": 15768,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2892,
+ "bbox": [
+ 11,
+ 103,
+ 394,
+ 224
+ ],
+ "category_id": 10,
+ "id": 5713,
+ "area": 88256,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2893,
+ "bbox": [
+ 152,
+ 110,
+ 235,
+ 277
+ ],
+ "category_id": 16,
+ "id": 5714,
+ "area": 65095,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2894,
+ "bbox": [
+ 134,
+ 245,
+ 321,
+ 91
+ ],
+ "category_id": 16,
+ "id": 5715,
+ "area": 29211,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2895,
+ "bbox": [
+ 308,
+ 56,
+ 58,
+ 398
+ ],
+ "category_id": 10,
+ "id": 5716,
+ "area": 23084,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2896,
+ "bbox": [
+ 55,
+ 126,
+ 435,
+ 272
+ ],
+ "category_id": 10,
+ "id": 5717,
+ "area": 118320,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2897,
+ "bbox": [
+ 239,
+ 117,
+ 135,
+ 133
+ ],
+ "category_id": 15,
+ "id": 5718,
+ "area": 17955,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2897,
+ "bbox": [
+ 135,
+ 250,
+ 140,
+ 134
+ ],
+ "category_id": 15,
+ "id": 5719,
+ "area": 18760,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2897,
+ "bbox": [
+ 58,
+ 122,
+ 83,
+ 76
+ ],
+ "category_id": 15,
+ "id": 5720,
+ "area": 6308,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2898,
+ "bbox": [
+ 120,
+ 106,
+ 59,
+ 108
+ ],
+ "category_id": 16,
+ "id": 5721,
+ "area": 6372,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2899,
+ "bbox": [
+ 58,
+ 55,
+ 15,
+ 11
+ ],
+ "category_id": 4,
+ "id": 5722,
+ "area": 165,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2899,
+ "bbox": [
+ 241,
+ 131,
+ 16,
+ 10
+ ],
+ "category_id": 4,
+ "id": 5723,
+ "area": 160,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2899,
+ "bbox": [
+ 140,
+ 172,
+ 15,
+ 8
+ ],
+ "category_id": 4,
+ "id": 5724,
+ "area": 120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2899,
+ "bbox": [
+ 180,
+ 288,
+ 17,
+ 10
+ ],
+ "category_id": 4,
+ "id": 5725,
+ "area": 170,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2899,
+ "bbox": [
+ 149,
+ 397,
+ 15,
+ 8
+ ],
+ "category_id": 4,
+ "id": 5726,
+ "area": 120,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2899,
+ "bbox": [
+ 39,
+ 423,
+ 15,
+ 7
+ ],
+ "category_id": 4,
+ "id": 5727,
+ "area": 105,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2899,
+ "bbox": [
+ 393,
+ 487,
+ 26,
+ 15
+ ],
+ "category_id": 4,
+ "id": 5728,
+ "area": 390,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2900,
+ "bbox": [
+ 200,
+ 84,
+ 64,
+ 31
+ ],
+ "category_id": 4,
+ "id": 5729,
+ "area": 1984,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2900,
+ "bbox": [
+ 325,
+ 412,
+ 64,
+ 33
+ ],
+ "category_id": 4,
+ "id": 5730,
+ "area": 2112,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2901,
+ "bbox": [
+ 106,
+ 128,
+ 37,
+ 78
+ ],
+ "category_id": 16,
+ "id": 5731,
+ "area": 2886,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2902,
+ "bbox": [
+ 359,
+ 106,
+ 63,
+ 48
+ ],
+ "category_id": 4,
+ "id": 5732,
+ "area": 3024,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2902,
+ "bbox": [
+ 137,
+ 418,
+ 56,
+ 51
+ ],
+ "category_id": 4,
+ "id": 5733,
+ "area": 2856,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2903,
+ "bbox": [
+ 111,
+ 62,
+ 50,
+ 26
+ ],
+ "category_id": 4,
+ "id": 5734,
+ "area": 1300,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2903,
+ "bbox": [
+ 276,
+ 395,
+ 51,
+ 28
+ ],
+ "category_id": 4,
+ "id": 5735,
+ "area": 1428,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2904,
+ "bbox": [
+ 68,
+ 200,
+ 403,
+ 115
+ ],
+ "category_id": 10,
+ "id": 5736,
+ "area": 46345,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2905,
+ "bbox": [
+ 176,
+ 232,
+ 185,
+ 185
+ ],
+ "category_id": 19,
+ "id": 5737,
+ "area": 34225,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2906,
+ "bbox": [
+ 172,
+ 183,
+ 226,
+ 200
+ ],
+ "category_id": 10,
+ "id": 5738,
+ "area": 45200,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2907,
+ "bbox": [
+ 263,
+ 277,
+ 214,
+ 32
+ ],
+ "category_id": 10,
+ "id": 5739,
+ "area": 6848,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2908,
+ "bbox": [
+ 392,
+ 177,
+ 20,
+ 20
+ ],
+ "category_id": 15,
+ "id": 5740,
+ "area": 400,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2909,
+ "bbox": [
+ 236,
+ 88,
+ 117,
+ 151
+ ],
+ "category_id": 15,
+ "id": 5741,
+ "area": 17667,
+ "iscrowd": 0
+ },
+ {
+ "image_id": 2909,
+ "bbox": [
+ 307,
+ 227,
+ 25,
+ 115
+ ],
+ "category_id": 15,
+ "id": 5742,
+ "area": 2875,
+ "iscrowd": 0
+ }
+ ],
+ "categories": [
+ {
+ "id": 1,
+ "name": "vehicle"
+ },
+ {
+ "id": 2,
+ "name": "baseballfield"
+ },
+ {
+ "id": 3,
+ "name": "groundtrackfield"
+ },
+ {
+ "id": 4,
+ "name": "windmill"
+ },
+ {
+ "id": 5,
+ "name": "bridge"
+ },
+ {
+ "id": 6,
+ "name": "overpass"
+ },
+ {
+ "id": 7,
+ "name": "ship"
+ },
+ {
+ "id": 8,
+ "name": "airplane"
+ },
+ {
+ "id": 9,
+ "name": "tenniscourt"
+ },
+ {
+ "id": 10,
+ "name": "airport"
+ },
+ {
+ "id": 11,
+ "name": "expressway-service-area"
+ },
+ {
+ "id": 12,
+ "name": "basketballcourt"
+ },
+ {
+ "id": 13,
+ "name": "stadium"
+ },
+ {
+ "id": 14,
+ "name": "storagetank"
+ },
+ {
+ "id": 15,
+ "name": "chimney"
+ },
+ {
+ "id": 16,
+ "name": "dam"
+ },
+ {
+ "id": 17,
+ "name": "expressway-toll-station"
+ },
+ {
+ "id": 18,
+ "name": "golffield"
+ },
+ {
+ "id": 19,
+ "name": "trainstation"
+ },
+ {
+ "id": 20,
+ "name": "harbor"
+ }
+ ],
+ "info": {
+ "description": "",
+ "url": "",
+ "version": "",
+ "year": 2022,
+ "contributor": "\u7eaf\u7cb9ss",
+ "date_created": "2022-07-8"
+ },
+ "licenses": [
+ {
+ "id": 1,
+ "name": null,
+ "url": null
+ }
+ ]
+}
\ No newline at end of file
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/gt_jsons/exdark/instances_val_test_novel.json b/scripts/evaluation/FasterRCNN_score-mmdet/gt_jsons/exdark/instances_val_test_novel.json
new file mode 100644
index 0000000000000000000000000000000000000000..c90e31ed12a7347c359acf1adfa2be762c3eef5d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/gt_jsons/exdark/instances_val_test_novel.json
@@ -0,0 +1,43190 @@
+{
+ "info": {
+ "description": "ExDark Dataset in COCO Format",
+ "year": 2024,
+ "contributor": "Converted by Script"
+ },
+ "licenses": [],
+ "images": [
+ {
+ "id": 1,
+ "file_name": "2015_00267.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2,
+ "file_name": "2015_00273.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3,
+ "file_name": "2015_00288.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4,
+ "file_name": "2015_00292.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 5,
+ "file_name": "2015_00294.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 6,
+ "file_name": "2015_00295.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 7,
+ "file_name": "2015_00310.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 8,
+ "file_name": "2015_00325.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 9,
+ "file_name": "2015_00332.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 10,
+ "file_name": "2015_00339.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 11,
+ "file_name": "2015_00372.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 12,
+ "file_name": "2015_00383.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 13,
+ "file_name": "2015_00385.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 14,
+ "file_name": "2015_00387.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 15,
+ "file_name": "2015_00393.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 16,
+ "file_name": "2015_00394.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 17,
+ "file_name": "2015_00396.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 18,
+ "file_name": "2015_00398.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 19,
+ "file_name": "2015_00400.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 20,
+ "file_name": "2015_00414.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 21,
+ "file_name": "2015_00415.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 22,
+ "file_name": "2015_00424.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 23,
+ "file_name": "2015_00445.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 24,
+ "file_name": "2015_00457.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 25,
+ "file_name": "2015_00491.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 26,
+ "file_name": "2015_00531.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 27,
+ "file_name": "2015_00536.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 28,
+ "file_name": "2015_00575.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 29,
+ "file_name": "2015_00587.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 30,
+ "file_name": "2015_00588.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 31,
+ "file_name": "2015_00601.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 32,
+ "file_name": "2015_00604.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 33,
+ "file_name": "2015_00617.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 34,
+ "file_name": "2015_00627.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 35,
+ "file_name": "2015_00630.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 36,
+ "file_name": "2015_00990.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 37,
+ "file_name": "2015_01022.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 38,
+ "file_name": "2015_01070.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 39,
+ "file_name": "2015_01583.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 40,
+ "file_name": "2015_01604.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 41,
+ "file_name": "2015_01619.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 42,
+ "file_name": "2015_01620.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 43,
+ "file_name": "2015_01627.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 44,
+ "file_name": "2015_01629.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 45,
+ "file_name": "2015_01630.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 46,
+ "file_name": "2015_01639.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 47,
+ "file_name": "2015_01640.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 48,
+ "file_name": "2015_01642.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 49,
+ "file_name": "2015_01643.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 50,
+ "file_name": "2015_01645.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 51,
+ "file_name": "2015_01646.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 52,
+ "file_name": "2015_01647.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 53,
+ "file_name": "2015_01649.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 54,
+ "file_name": "2015_01650.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 55,
+ "file_name": "2015_01651.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 56,
+ "file_name": "2015_01652.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 57,
+ "file_name": "2015_01653.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 58,
+ "file_name": "2015_01654.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 59,
+ "file_name": "2015_01657.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 60,
+ "file_name": "2015_01661.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 61,
+ "file_name": "2015_01664.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 62,
+ "file_name": "2015_01665.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 63,
+ "file_name": "2015_01667.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 64,
+ "file_name": "2015_01668.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 65,
+ "file_name": "2015_01669.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 66,
+ "file_name": "2015_01671.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 67,
+ "file_name": "2015_01672.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 68,
+ "file_name": "2015_01673.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 69,
+ "file_name": "2015_01677.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 70,
+ "file_name": "2015_01678.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 71,
+ "file_name": "2015_01680.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 72,
+ "file_name": "2015_01683.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 73,
+ "file_name": "2015_01684.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 74,
+ "file_name": "2015_01685.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 75,
+ "file_name": "2015_01687.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 76,
+ "file_name": "2015_01691.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 77,
+ "file_name": "2015_01693.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 78,
+ "file_name": "2015_01700.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 79,
+ "file_name": "2015_01702.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 80,
+ "file_name": "2015_01703.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 81,
+ "file_name": "2015_01704.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 82,
+ "file_name": "2015_01706.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 83,
+ "file_name": "2015_01707.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 84,
+ "file_name": "2015_01708.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 85,
+ "file_name": "2015_01711.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 86,
+ "file_name": "2015_01712.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 87,
+ "file_name": "2015_01713.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 88,
+ "file_name": "2015_01714.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 89,
+ "file_name": "2015_01727.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 90,
+ "file_name": "2015_01728.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 91,
+ "file_name": "2015_01731.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 92,
+ "file_name": "2015_01737.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 93,
+ "file_name": "2015_01741.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 94,
+ "file_name": "2015_01744.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 95,
+ "file_name": "2015_01746.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 96,
+ "file_name": "2015_01748.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 97,
+ "file_name": "2015_01751.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 98,
+ "file_name": "2015_01755.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 99,
+ "file_name": "2015_01759.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 100,
+ "file_name": "2015_01765.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 101,
+ "file_name": "2015_01766.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 102,
+ "file_name": "2015_01771.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 103,
+ "file_name": "2015_01772.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 104,
+ "file_name": "2015_01773.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 105,
+ "file_name": "2015_01782.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 106,
+ "file_name": "2015_01783.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 107,
+ "file_name": "2015_01787.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 108,
+ "file_name": "2015_01790.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 109,
+ "file_name": "2015_01791.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 110,
+ "file_name": "2015_01792.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 111,
+ "file_name": "2015_01793.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 112,
+ "file_name": "2015_01794.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 113,
+ "file_name": "2015_01795.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 114,
+ "file_name": "2015_01802.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 115,
+ "file_name": "2015_01803.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 116,
+ "file_name": "2015_01809.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 117,
+ "file_name": "2015_01824.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 118,
+ "file_name": "2015_01828.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 119,
+ "file_name": "2015_01829.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 120,
+ "file_name": "2015_01830.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 121,
+ "file_name": "2015_01832.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 122,
+ "file_name": "2015_01833.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 123,
+ "file_name": "2015_01839.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 124,
+ "file_name": "2015_01842.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 125,
+ "file_name": "2015_01844.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 126,
+ "file_name": "2015_01855.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 127,
+ "file_name": "2015_01857.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 128,
+ "file_name": "2015_01858.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 129,
+ "file_name": "2015_01860.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 130,
+ "file_name": "2015_01861.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 131,
+ "file_name": "2015_01862.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 132,
+ "file_name": "2015_01863.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 133,
+ "file_name": "2015_01865.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 134,
+ "file_name": "2015_01866.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 135,
+ "file_name": "2015_01868.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 136,
+ "file_name": "2015_01869.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 137,
+ "file_name": "2015_01870.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 138,
+ "file_name": "2015_01874.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 139,
+ "file_name": "2015_01875.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 140,
+ "file_name": "2015_01876.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 141,
+ "file_name": "2015_01877.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 142,
+ "file_name": "2015_02129.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 143,
+ "file_name": "2015_02130.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 144,
+ "file_name": "2015_02131.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 145,
+ "file_name": "2015_02132.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 146,
+ "file_name": "2015_02133.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 147,
+ "file_name": "2015_02134.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 148,
+ "file_name": "2015_02135.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 149,
+ "file_name": "2015_02136.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 150,
+ "file_name": "2015_02137.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 151,
+ "file_name": "2015_02138.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 152,
+ "file_name": "2015_02139.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 153,
+ "file_name": "2015_02140.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 154,
+ "file_name": "2015_02141.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 155,
+ "file_name": "2015_02142.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 156,
+ "file_name": "2015_02143.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 157,
+ "file_name": "2015_02144.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 158,
+ "file_name": "2015_02145.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 159,
+ "file_name": "2015_02146.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 160,
+ "file_name": "2015_02147.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 161,
+ "file_name": "2015_02148.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 162,
+ "file_name": "2015_02149.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 163,
+ "file_name": "2015_02150.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 164,
+ "file_name": "2015_02151.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 165,
+ "file_name": "2015_02152.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 166,
+ "file_name": "2015_02153.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 167,
+ "file_name": "2015_02154.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 168,
+ "file_name": "2015_02155.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 169,
+ "file_name": "2015_02156.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 170,
+ "file_name": "2015_02157.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 171,
+ "file_name": "2015_02158.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 172,
+ "file_name": "2015_02159.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 173,
+ "file_name": "2015_02160.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 174,
+ "file_name": "2015_02161.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 175,
+ "file_name": "2015_02162.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 176,
+ "file_name": "2015_02163.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 177,
+ "file_name": "2015_02164.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 178,
+ "file_name": "2015_02165.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 179,
+ "file_name": "2015_02166.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 180,
+ "file_name": "2015_02167.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 181,
+ "file_name": "2015_02168.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 182,
+ "file_name": "2015_02169.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 183,
+ "file_name": "2015_02170.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 184,
+ "file_name": "2015_02171.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 185,
+ "file_name": "2015_02172.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 186,
+ "file_name": "2015_02173.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 187,
+ "file_name": "2015_02174.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 188,
+ "file_name": "2015_02175.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 189,
+ "file_name": "2015_02176.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 190,
+ "file_name": "2015_02177.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 191,
+ "file_name": "2015_02178.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 192,
+ "file_name": "2015_02179.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 193,
+ "file_name": "2015_02180.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 194,
+ "file_name": "2015_02181.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 195,
+ "file_name": "2015_02182.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 196,
+ "file_name": "2015_02183.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 197,
+ "file_name": "2015_02184.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 198,
+ "file_name": "2015_02185.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 199,
+ "file_name": "2015_02186.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 200,
+ "file_name": "2015_02187.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 201,
+ "file_name": "2015_02188.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 202,
+ "file_name": "2015_02189.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 203,
+ "file_name": "2015_02190.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 204,
+ "file_name": "2015_02191.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 205,
+ "file_name": "2015_02192.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 206,
+ "file_name": "2015_02193.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 207,
+ "file_name": "2015_02194.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 208,
+ "file_name": "2015_02195.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 209,
+ "file_name": "2015_02196.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 210,
+ "file_name": "2015_02197.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 211,
+ "file_name": "2015_02198.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 212,
+ "file_name": "2015_02199.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 213,
+ "file_name": "2015_02200.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 214,
+ "file_name": "2015_02201.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 215,
+ "file_name": "2015_02202.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 216,
+ "file_name": "2015_02203.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 217,
+ "file_name": "2015_02204.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 218,
+ "file_name": "2015_02205.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 219,
+ "file_name": "2015_02206.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 220,
+ "file_name": "2015_02207.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 221,
+ "file_name": "2015_02208.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 222,
+ "file_name": "2015_02209.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 223,
+ "file_name": "2015_02210.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 224,
+ "file_name": "2015_02211.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 225,
+ "file_name": "2015_02212.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 226,
+ "file_name": "2015_02213.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 227,
+ "file_name": "2015_02214.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 228,
+ "file_name": "2015_02215.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 229,
+ "file_name": "2015_02216.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 230,
+ "file_name": "2015_02217.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 231,
+ "file_name": "2015_02218.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 232,
+ "file_name": "2015_02219.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 233,
+ "file_name": "2015_02220.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 234,
+ "file_name": "2015_02221.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 235,
+ "file_name": "2015_02222.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 236,
+ "file_name": "2015_02223.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 237,
+ "file_name": "2015_02224.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 238,
+ "file_name": "2015_02225.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 239,
+ "file_name": "2015_02226.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 240,
+ "file_name": "2015_02227.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 241,
+ "file_name": "2015_02228.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 242,
+ "file_name": "2015_02229.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 243,
+ "file_name": "2015_02230.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 244,
+ "file_name": "2015_02231.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 245,
+ "file_name": "2015_02232.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 246,
+ "file_name": "2015_02233.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 247,
+ "file_name": "2015_02234.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 248,
+ "file_name": "2015_02235.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 249,
+ "file_name": "2015_02236.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 250,
+ "file_name": "2015_02237.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 251,
+ "file_name": "2015_02238.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 252,
+ "file_name": "2015_02239.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 253,
+ "file_name": "2015_02240.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 254,
+ "file_name": "2015_02241.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 255,
+ "file_name": "2015_02242.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 256,
+ "file_name": "2015_02243.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 257,
+ "file_name": "2015_02244.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 258,
+ "file_name": "2015_02245.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 259,
+ "file_name": "2015_02246.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 260,
+ "file_name": "2015_02247.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 261,
+ "file_name": "2015_02248.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 262,
+ "file_name": "2015_02249.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 263,
+ "file_name": "2015_02250.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 264,
+ "file_name": "2015_02251.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 265,
+ "file_name": "2015_02252.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 266,
+ "file_name": "2015_02253.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 267,
+ "file_name": "2015_02254.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 268,
+ "file_name": "2015_02255.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 269,
+ "file_name": "2015_02256.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 270,
+ "file_name": "2015_02257.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 271,
+ "file_name": "2015_02258.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 272,
+ "file_name": "2015_02259.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 273,
+ "file_name": "2015_02260.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 274,
+ "file_name": "2015_02261.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 275,
+ "file_name": "2015_02262.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 276,
+ "file_name": "2015_02263.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 277,
+ "file_name": "2015_02264.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 278,
+ "file_name": "2015_02265.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 279,
+ "file_name": "2015_02266.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 280,
+ "file_name": "2015_02267.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 281,
+ "file_name": "2015_02268.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 282,
+ "file_name": "2015_02269.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 283,
+ "file_name": "2015_02270.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 284,
+ "file_name": "2015_02271.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 285,
+ "file_name": "2015_02272.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 286,
+ "file_name": "2015_02273.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 287,
+ "file_name": "2015_02274.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 288,
+ "file_name": "2015_02275.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 289,
+ "file_name": "2015_02276.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 290,
+ "file_name": "2015_02277.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 291,
+ "file_name": "2015_02278.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 292,
+ "file_name": "2015_02279.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 293,
+ "file_name": "2015_02280.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 294,
+ "file_name": "2015_02281.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 295,
+ "file_name": "2015_02282.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 296,
+ "file_name": "2015_02283.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 297,
+ "file_name": "2015_02284.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 298,
+ "file_name": "2015_02285.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 299,
+ "file_name": "2015_02286.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 300,
+ "file_name": "2015_02287.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 301,
+ "file_name": "2015_02288.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 302,
+ "file_name": "2015_02289.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 303,
+ "file_name": "2015_02290.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 304,
+ "file_name": "2015_02291.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 305,
+ "file_name": "2015_02292.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 306,
+ "file_name": "2015_02293.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 307,
+ "file_name": "2015_02294.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 308,
+ "file_name": "2015_02295.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 309,
+ "file_name": "2015_02296.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 310,
+ "file_name": "2015_02297.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 311,
+ "file_name": "2015_02298.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 312,
+ "file_name": "2015_02299.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 313,
+ "file_name": "2015_02300.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 314,
+ "file_name": "2015_02301.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 315,
+ "file_name": "2015_02302.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 316,
+ "file_name": "2015_02303.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 317,
+ "file_name": "2015_02304.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 318,
+ "file_name": "2015_02305.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 319,
+ "file_name": "2015_02306.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 320,
+ "file_name": "2015_02307.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 321,
+ "file_name": "2015_02308.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 322,
+ "file_name": "2015_02309.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 323,
+ "file_name": "2015_02310.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 324,
+ "file_name": "2015_02311.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 325,
+ "file_name": "2015_02312.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 326,
+ "file_name": "2015_02313.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 327,
+ "file_name": "2015_02314.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 328,
+ "file_name": "2015_02315.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 329,
+ "file_name": "2015_02316.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 330,
+ "file_name": "2015_02317.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 331,
+ "file_name": "2015_02318.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 332,
+ "file_name": "2015_02319.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 333,
+ "file_name": "2015_02320.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 334,
+ "file_name": "2015_02321.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 335,
+ "file_name": "2015_02322.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 336,
+ "file_name": "2015_02323.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 337,
+ "file_name": "2015_02324.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 338,
+ "file_name": "2015_02325.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 339,
+ "file_name": "2015_02326.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 340,
+ "file_name": "2015_02327.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 341,
+ "file_name": "2015_02328.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 342,
+ "file_name": "2015_02329.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 343,
+ "file_name": "2015_02330.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 344,
+ "file_name": "2015_02331.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 345,
+ "file_name": "2015_02332.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 346,
+ "file_name": "2015_02333.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 347,
+ "file_name": "2015_02334.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 348,
+ "file_name": "2015_02335.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 349,
+ "file_name": "2015_02336.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 350,
+ "file_name": "2015_02337.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 351,
+ "file_name": "2015_02338.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 352,
+ "file_name": "2015_02339.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 353,
+ "file_name": "2015_02340.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 354,
+ "file_name": "2015_02341.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 355,
+ "file_name": "2015_02342.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 356,
+ "file_name": "2015_02343.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 357,
+ "file_name": "2015_02344.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 358,
+ "file_name": "2015_02345.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 359,
+ "file_name": "2015_02346.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 360,
+ "file_name": "2015_02347.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 361,
+ "file_name": "2015_02348.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 362,
+ "file_name": "2015_02349.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 363,
+ "file_name": "2015_02350.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 364,
+ "file_name": "2015_02351.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 365,
+ "file_name": "2015_02352.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 366,
+ "file_name": "2015_02353.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 367,
+ "file_name": "2015_02354.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 368,
+ "file_name": "2015_02355.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 369,
+ "file_name": "2015_02356.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 370,
+ "file_name": "2015_02357.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 371,
+ "file_name": "2015_02358.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 372,
+ "file_name": "2015_02359.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 373,
+ "file_name": "2015_02360.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 374,
+ "file_name": "2015_02361.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 375,
+ "file_name": "2015_02362.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 376,
+ "file_name": "2015_02363.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 377,
+ "file_name": "2015_02364.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 378,
+ "file_name": "2015_02365.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 379,
+ "file_name": "2015_02366.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 380,
+ "file_name": "2015_02367.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 381,
+ "file_name": "2015_02368.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 382,
+ "file_name": "2015_02369.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 383,
+ "file_name": "2015_02370.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 384,
+ "file_name": "2015_02371.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 385,
+ "file_name": "2015_02372.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 386,
+ "file_name": "2015_02373.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 387,
+ "file_name": "2015_02374.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 388,
+ "file_name": "2015_02375.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 389,
+ "file_name": "2015_02376.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 390,
+ "file_name": "2015_02377.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 391,
+ "file_name": "2015_02378.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 392,
+ "file_name": "2015_02379.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 393,
+ "file_name": "2015_02380.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 394,
+ "file_name": "2015_02381.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 395,
+ "file_name": "2015_02382.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 396,
+ "file_name": "2015_02383.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 397,
+ "file_name": "2015_02384.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 398,
+ "file_name": "2015_02385.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 399,
+ "file_name": "2015_02386.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 400,
+ "file_name": "2015_02387.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 401,
+ "file_name": "2015_02388.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 402,
+ "file_name": "2015_02389.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 403,
+ "file_name": "2015_02390.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 404,
+ "file_name": "2015_02391.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 405,
+ "file_name": "2015_02392.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 406,
+ "file_name": "2015_02393.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 407,
+ "file_name": "2015_02394.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 408,
+ "file_name": "2015_02395.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 409,
+ "file_name": "2015_02396.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 410,
+ "file_name": "2015_02397.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 411,
+ "file_name": "2015_02398.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 412,
+ "file_name": "2015_02399.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 413,
+ "file_name": "2015_02400.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 414,
+ "file_name": "2015_02401.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 415,
+ "file_name": "2015_02402.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 416,
+ "file_name": "2015_02403.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 417,
+ "file_name": "2015_02404.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 418,
+ "file_name": "2015_02405.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 419,
+ "file_name": "2015_02773.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 420,
+ "file_name": "2015_02800.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 421,
+ "file_name": "2015_02861.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 422,
+ "file_name": "2015_02876.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 423,
+ "file_name": "2015_02913.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 424,
+ "file_name": "2015_02914.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 425,
+ "file_name": "2015_02915.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 426,
+ "file_name": "2015_02916.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 427,
+ "file_name": "2015_02923.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 428,
+ "file_name": "2015_02925.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 429,
+ "file_name": "2015_02927.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 430,
+ "file_name": "2015_03039.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 431,
+ "file_name": "2017_07360.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 432,
+ "file_name": "2015_03303.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 433,
+ "file_name": "2015_03309.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 434,
+ "file_name": "2015_03332.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 435,
+ "file_name": "2015_03333.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 436,
+ "file_name": "2015_03357.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 437,
+ "file_name": "2015_03390.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 438,
+ "file_name": "2015_03440.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 439,
+ "file_name": "2015_03732.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 440,
+ "file_name": "2015_03741.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 441,
+ "file_name": "2015_04029.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 442,
+ "file_name": "2015_04031.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 443,
+ "file_name": "2015_04041.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 444,
+ "file_name": "2015_04042.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 445,
+ "file_name": "2015_04051.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 446,
+ "file_name": "2015_04055.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 447,
+ "file_name": "2015_04058.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 448,
+ "file_name": "2015_04061.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 449,
+ "file_name": "2015_04063.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 450,
+ "file_name": "2015_04064.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 451,
+ "file_name": "2015_04076.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 452,
+ "file_name": "2015_04079.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 453,
+ "file_name": "2015_04081.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 454,
+ "file_name": "2015_04082.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 455,
+ "file_name": "2015_04087.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 456,
+ "file_name": "2015_04088.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 457,
+ "file_name": "2015_04089.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 458,
+ "file_name": "2015_04090.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 459,
+ "file_name": "2015_04091.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 460,
+ "file_name": "2015_04092.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 461,
+ "file_name": "2015_04094.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 462,
+ "file_name": "2015_04095.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 463,
+ "file_name": "2015_04100.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 464,
+ "file_name": "2015_04102.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 465,
+ "file_name": "2015_04106.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 466,
+ "file_name": "2015_04108.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 467,
+ "file_name": "2015_04111.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 468,
+ "file_name": "2015_04112.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 469,
+ "file_name": "2015_04113.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 470,
+ "file_name": "2015_04114.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 471,
+ "file_name": "2015_04117.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 472,
+ "file_name": "2015_04118.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 473,
+ "file_name": "2015_04120.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 474,
+ "file_name": "2015_04121.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 475,
+ "file_name": "2015_04122.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 476,
+ "file_name": "2015_04123.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 477,
+ "file_name": "2015_04126.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 478,
+ "file_name": "2015_04127.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 479,
+ "file_name": "2015_04129.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 480,
+ "file_name": "2015_04130.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 481,
+ "file_name": "2015_04131.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 482,
+ "file_name": "2015_04132.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 483,
+ "file_name": "2015_04133.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 484,
+ "file_name": "2015_04134.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 485,
+ "file_name": "2015_04137.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 486,
+ "file_name": "2015_04138.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 487,
+ "file_name": "2015_04140.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 488,
+ "file_name": "2015_04143.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 489,
+ "file_name": "2015_04144.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 490,
+ "file_name": "2015_04146.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 491,
+ "file_name": "2015_04147.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 492,
+ "file_name": "2015_04149.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 493,
+ "file_name": "2015_04150.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 494,
+ "file_name": "2015_04151.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 495,
+ "file_name": "2015_04154.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 496,
+ "file_name": "2015_04155.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 497,
+ "file_name": "2015_04156.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 498,
+ "file_name": "2015_04158.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 499,
+ "file_name": "2015_04159.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 500,
+ "file_name": "2015_04161.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 501,
+ "file_name": "2015_04162.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 502,
+ "file_name": "2015_04163.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 503,
+ "file_name": "2015_04164.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 504,
+ "file_name": "2015_04165.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 505,
+ "file_name": "2015_04166.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 506,
+ "file_name": "2015_04168.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 507,
+ "file_name": "2015_04169.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 508,
+ "file_name": "2015_04183.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 509,
+ "file_name": "2015_04184.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 510,
+ "file_name": "2015_04185.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 511,
+ "file_name": "2015_04204.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 512,
+ "file_name": "2015_04211.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 513,
+ "file_name": "2015_04216.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 514,
+ "file_name": "2015_04218.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 515,
+ "file_name": "2015_04221.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 516,
+ "file_name": "2015_04222.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 517,
+ "file_name": "2015_04240.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 518,
+ "file_name": "2015_04254.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 519,
+ "file_name": "2015_04255.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 520,
+ "file_name": "2015_04275.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 521,
+ "file_name": "2015_04277.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 522,
+ "file_name": "2015_04278.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 523,
+ "file_name": "2015_04288.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 524,
+ "file_name": "2015_04289.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 525,
+ "file_name": "2015_04290.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 526,
+ "file_name": "2015_04291.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 527,
+ "file_name": "2015_04292.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 528,
+ "file_name": "2015_04293.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 529,
+ "file_name": "2015_04294.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 530,
+ "file_name": "2015_04295.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 531,
+ "file_name": "2015_04296.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 532,
+ "file_name": "2015_04297.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 533,
+ "file_name": "2015_04298.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 534,
+ "file_name": "2015_04299.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 535,
+ "file_name": "2015_04300.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 536,
+ "file_name": "2015_04304.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 537,
+ "file_name": "2015_04331.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 538,
+ "file_name": "2015_04342.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 539,
+ "file_name": "2015_04344.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 540,
+ "file_name": "2015_04345.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 541,
+ "file_name": "2015_04350.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 542,
+ "file_name": "2015_04355.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 543,
+ "file_name": "2015_04358.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 544,
+ "file_name": "2015_04362.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 545,
+ "file_name": "2015_04365.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 546,
+ "file_name": "2015_04368.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 547,
+ "file_name": "2015_04370.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 548,
+ "file_name": "2015_04371.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 549,
+ "file_name": "2015_04374.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 550,
+ "file_name": "2015_04377.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 551,
+ "file_name": "2015_04382.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 552,
+ "file_name": "2015_04388.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 553,
+ "file_name": "2015_04389.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 554,
+ "file_name": "2015_04390.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 555,
+ "file_name": "2015_04392.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 556,
+ "file_name": "2015_04393.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 557,
+ "file_name": "2015_04397.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 558,
+ "file_name": "2015_04400.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 559,
+ "file_name": "2015_04403.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 560,
+ "file_name": "2015_04415.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 561,
+ "file_name": "2015_04416.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 562,
+ "file_name": "2015_04417.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 563,
+ "file_name": "2015_04678.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 564,
+ "file_name": "2015_04697.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 565,
+ "file_name": "2015_04702.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 566,
+ "file_name": "2015_04705.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 567,
+ "file_name": "2015_04764.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 568,
+ "file_name": "2015_04792.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 569,
+ "file_name": "2015_04795.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 570,
+ "file_name": "2015_04836.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 571,
+ "file_name": "2015_04837.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 572,
+ "file_name": "2015_04838.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 573,
+ "file_name": "2015_04873.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 574,
+ "file_name": "2015_04921.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 575,
+ "file_name": "2015_04927.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 576,
+ "file_name": "2015_04928.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 577,
+ "file_name": "2015_04932.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 578,
+ "file_name": "2015_04941.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 579,
+ "file_name": "2015_05194.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 580,
+ "file_name": "2015_05195.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 581,
+ "file_name": "2015_05196.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 582,
+ "file_name": "2015_05197.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 583,
+ "file_name": "2015_05198.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 584,
+ "file_name": "2015_05199.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 585,
+ "file_name": "2015_05200.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 586,
+ "file_name": "2015_05201.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 587,
+ "file_name": "2015_05202.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 588,
+ "file_name": "2015_05203.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 589,
+ "file_name": "2015_05204.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 590,
+ "file_name": "2015_05205.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 591,
+ "file_name": "2015_05206.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 592,
+ "file_name": "2015_05207.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 593,
+ "file_name": "2015_05208.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 594,
+ "file_name": "2015_05209.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 595,
+ "file_name": "2015_05210.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 596,
+ "file_name": "2015_05211.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 597,
+ "file_name": "2015_05212.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 598,
+ "file_name": "2015_05213.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 599,
+ "file_name": "2015_05214.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 600,
+ "file_name": "2015_05215.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 601,
+ "file_name": "2015_05216.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 602,
+ "file_name": "2015_05217.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 603,
+ "file_name": "2015_05218.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 604,
+ "file_name": "2015_05219.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 605,
+ "file_name": "2015_05220.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 606,
+ "file_name": "2015_05221.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 607,
+ "file_name": "2015_05222.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 608,
+ "file_name": "2015_05223.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 609,
+ "file_name": "2015_05224.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 610,
+ "file_name": "2015_05225.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 611,
+ "file_name": "2015_05226.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 612,
+ "file_name": "2015_05227.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 613,
+ "file_name": "2015_05228.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 614,
+ "file_name": "2015_05229.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 615,
+ "file_name": "2015_05230.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 616,
+ "file_name": "2015_05231.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 617,
+ "file_name": "2015_05232.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 618,
+ "file_name": "2015_05233.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 619,
+ "file_name": "2015_05234.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 620,
+ "file_name": "2015_05235.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 621,
+ "file_name": "2015_05236.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 622,
+ "file_name": "2015_05237.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 623,
+ "file_name": "2015_05238.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 624,
+ "file_name": "2015_05239.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 625,
+ "file_name": "2015_05240.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 626,
+ "file_name": "2015_05241.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 627,
+ "file_name": "2015_05242.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 628,
+ "file_name": "2015_05243.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 629,
+ "file_name": "2015_05244.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 630,
+ "file_name": "2015_05245.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 631,
+ "file_name": "2015_05246.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 632,
+ "file_name": "2015_05247.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 633,
+ "file_name": "2015_05248.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 634,
+ "file_name": "2015_05249.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 635,
+ "file_name": "2015_05250.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 636,
+ "file_name": "2015_05251.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 637,
+ "file_name": "2015_05252.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 638,
+ "file_name": "2015_05253.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 639,
+ "file_name": "2015_05254.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 640,
+ "file_name": "2015_05255.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 641,
+ "file_name": "2015_05256.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 642,
+ "file_name": "2015_05257.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 643,
+ "file_name": "2015_05258.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 644,
+ "file_name": "2015_05259.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 645,
+ "file_name": "2015_05260.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 646,
+ "file_name": "2015_05261.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 647,
+ "file_name": "2015_05262.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 648,
+ "file_name": "2015_05263.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 649,
+ "file_name": "2015_05264.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 650,
+ "file_name": "2015_05265.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 651,
+ "file_name": "2015_05266.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 652,
+ "file_name": "2015_05267.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 653,
+ "file_name": "2015_05268.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 654,
+ "file_name": "2015_05269.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 655,
+ "file_name": "2015_05270.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 656,
+ "file_name": "2015_05271.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 657,
+ "file_name": "2015_05272.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 658,
+ "file_name": "2015_05273.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 659,
+ "file_name": "2015_05274.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 660,
+ "file_name": "2015_05275.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 661,
+ "file_name": "2015_05276.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 662,
+ "file_name": "2015_05277.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 663,
+ "file_name": "2015_05278.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 664,
+ "file_name": "2015_05279.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 665,
+ "file_name": "2015_05280.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 666,
+ "file_name": "2015_05281.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 667,
+ "file_name": "2015_05282.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 668,
+ "file_name": "2015_05283.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 669,
+ "file_name": "2015_05284.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 670,
+ "file_name": "2015_05285.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 671,
+ "file_name": "2015_05286.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 672,
+ "file_name": "2015_05287.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 673,
+ "file_name": "2015_05288.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 674,
+ "file_name": "2015_05289.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 675,
+ "file_name": "2015_05290.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 676,
+ "file_name": "2015_05291.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 677,
+ "file_name": "2015_05292.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 678,
+ "file_name": "2015_05293.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 679,
+ "file_name": "2015_05294.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 680,
+ "file_name": "2015_05295.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 681,
+ "file_name": "2015_05296.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 682,
+ "file_name": "2015_05297.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 683,
+ "file_name": "2015_05298.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 684,
+ "file_name": "2015_05299.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 685,
+ "file_name": "2015_05300.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 686,
+ "file_name": "2015_05301.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 687,
+ "file_name": "2015_05302.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 688,
+ "file_name": "2015_05303.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 689,
+ "file_name": "2015_05304.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 690,
+ "file_name": "2015_05305.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 691,
+ "file_name": "2015_05306.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 692,
+ "file_name": "2015_05307.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 693,
+ "file_name": "2015_05308.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 694,
+ "file_name": "2015_05309.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 695,
+ "file_name": "2015_05310.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 696,
+ "file_name": "2015_05311.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 697,
+ "file_name": "2015_05312.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 698,
+ "file_name": "2015_05313.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 699,
+ "file_name": "2015_05314.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 700,
+ "file_name": "2015_05315.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 701,
+ "file_name": "2015_05316.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 702,
+ "file_name": "2015_05317.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 703,
+ "file_name": "2015_05318.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 704,
+ "file_name": "2015_05319.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 705,
+ "file_name": "2015_05320.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 706,
+ "file_name": "2015_05321.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 707,
+ "file_name": "2015_05322.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 708,
+ "file_name": "2015_05323.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 709,
+ "file_name": "2015_05324.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 710,
+ "file_name": "2015_05325.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 711,
+ "file_name": "2015_05326.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 712,
+ "file_name": "2015_05327.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 713,
+ "file_name": "2015_05328.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 714,
+ "file_name": "2015_05329.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 715,
+ "file_name": "2015_05330.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 716,
+ "file_name": "2015_05331.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 717,
+ "file_name": "2015_05332.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 718,
+ "file_name": "2015_05333.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 719,
+ "file_name": "2015_05334.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 720,
+ "file_name": "2015_05335.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 721,
+ "file_name": "2015_05336.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 722,
+ "file_name": "2015_05337.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 723,
+ "file_name": "2015_05338.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 724,
+ "file_name": "2015_05339.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 725,
+ "file_name": "2015_05340.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 726,
+ "file_name": "2015_05341.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 727,
+ "file_name": "2015_05342.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 728,
+ "file_name": "2015_05343.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 729,
+ "file_name": "2015_05344.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 730,
+ "file_name": "2015_05345.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 731,
+ "file_name": "2015_05346.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 732,
+ "file_name": "2015_05347.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 733,
+ "file_name": "2015_05348.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 734,
+ "file_name": "2015_05349.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 735,
+ "file_name": "2015_05350.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 736,
+ "file_name": "2015_05351.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 737,
+ "file_name": "2015_05352.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 738,
+ "file_name": "2015_05353.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 739,
+ "file_name": "2015_05354.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 740,
+ "file_name": "2015_05355.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 741,
+ "file_name": "2015_05356.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 742,
+ "file_name": "2015_05357.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 743,
+ "file_name": "2015_05358.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 744,
+ "file_name": "2015_05359.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 745,
+ "file_name": "2015_05360.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 746,
+ "file_name": "2015_05361.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 747,
+ "file_name": "2015_05362.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 748,
+ "file_name": "2015_05363.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 749,
+ "file_name": "2015_05364.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 750,
+ "file_name": "2015_05365.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 751,
+ "file_name": "2015_05366.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 752,
+ "file_name": "2015_05367.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 753,
+ "file_name": "2015_05368.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 754,
+ "file_name": "2015_05369.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 755,
+ "file_name": "2015_05370.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 756,
+ "file_name": "2015_05371.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 757,
+ "file_name": "2015_05372.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 758,
+ "file_name": "2015_05373.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 759,
+ "file_name": "2015_05374.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 760,
+ "file_name": "2015_05375.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 761,
+ "file_name": "2015_05376.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 762,
+ "file_name": "2015_05377.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 763,
+ "file_name": "2015_05378.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 764,
+ "file_name": "2015_05379.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 765,
+ "file_name": "2015_05380.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 766,
+ "file_name": "2015_05381.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 767,
+ "file_name": "2015_05382.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 768,
+ "file_name": "2015_05383.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 769,
+ "file_name": "2015_05384.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 770,
+ "file_name": "2015_05385.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 771,
+ "file_name": "2015_05386.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 772,
+ "file_name": "2015_05387.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 773,
+ "file_name": "2015_05388.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 774,
+ "file_name": "2015_05389.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 775,
+ "file_name": "2015_05390.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 776,
+ "file_name": "2015_05391.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 777,
+ "file_name": "2015_05392.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 778,
+ "file_name": "2015_05393.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 779,
+ "file_name": "2015_05394.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 780,
+ "file_name": "2015_05395.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 781,
+ "file_name": "2015_05396.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 782,
+ "file_name": "2015_05397.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 783,
+ "file_name": "2015_05398.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 784,
+ "file_name": "2015_05399.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 785,
+ "file_name": "2015_05400.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 786,
+ "file_name": "2015_05401.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 787,
+ "file_name": "2015_05402.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 788,
+ "file_name": "2015_05403.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 789,
+ "file_name": "2015_05404.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 790,
+ "file_name": "2015_05405.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 791,
+ "file_name": "2015_05406.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 792,
+ "file_name": "2015_05407.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 793,
+ "file_name": "2015_05408.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 794,
+ "file_name": "2015_05409.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 795,
+ "file_name": "2015_05410.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 796,
+ "file_name": "2015_05411.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 797,
+ "file_name": "2015_05412.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 798,
+ "file_name": "2015_05413.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 799,
+ "file_name": "2015_05414.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 800,
+ "file_name": "2015_05415.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 801,
+ "file_name": "2015_05416.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 802,
+ "file_name": "2015_05417.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 803,
+ "file_name": "2015_05418.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 804,
+ "file_name": "2015_05419.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 805,
+ "file_name": "2015_05420.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 806,
+ "file_name": "2015_05421.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 807,
+ "file_name": "2015_05422.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 808,
+ "file_name": "2015_05423.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 809,
+ "file_name": "2015_05424.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 810,
+ "file_name": "2015_05425.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 811,
+ "file_name": "2015_05426.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 812,
+ "file_name": "2015_05427.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 813,
+ "file_name": "2015_05428.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 814,
+ "file_name": "2015_05429.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 815,
+ "file_name": "2015_05430.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 816,
+ "file_name": "2015_05431.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 817,
+ "file_name": "2015_05432.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 818,
+ "file_name": "2015_05433.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 819,
+ "file_name": "2015_05434.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 820,
+ "file_name": "2015_05435.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 821,
+ "file_name": "2015_05436.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 822,
+ "file_name": "2015_05437.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 823,
+ "file_name": "2015_05438.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 824,
+ "file_name": "2015_05439.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 825,
+ "file_name": "2015_05440.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 826,
+ "file_name": "2015_05441.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 827,
+ "file_name": "2015_05442.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 828,
+ "file_name": "2015_05443.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 829,
+ "file_name": "2015_05444.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 830,
+ "file_name": "2015_05445.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 831,
+ "file_name": "2015_05446.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 832,
+ "file_name": "2015_05447.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 833,
+ "file_name": "2015_05448.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 834,
+ "file_name": "2015_05449.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 835,
+ "file_name": "2015_05450.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 836,
+ "file_name": "2015_05451.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 837,
+ "file_name": "2015_05452.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 838,
+ "file_name": "2015_05453.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 839,
+ "file_name": "2015_05454.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 840,
+ "file_name": "2015_05455.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 841,
+ "file_name": "2015_05456.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 842,
+ "file_name": "2015_05457.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 843,
+ "file_name": "2015_05458.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 844,
+ "file_name": "2015_05459.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 845,
+ "file_name": "2015_05460.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 846,
+ "file_name": "2015_05461.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 847,
+ "file_name": "2015_05462.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 848,
+ "file_name": "2015_05463.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 849,
+ "file_name": "2015_05464.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 850,
+ "file_name": "2015_05465.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 851,
+ "file_name": "2015_05466.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 852,
+ "file_name": "2015_05467.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 853,
+ "file_name": "2015_05468.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 854,
+ "file_name": "2015_05469.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 855,
+ "file_name": "2015_05470.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 856,
+ "file_name": "2015_05471.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 857,
+ "file_name": "2015_05472.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 858,
+ "file_name": "2015_05473.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 859,
+ "file_name": "2015_05474.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 860,
+ "file_name": "2015_05475.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 861,
+ "file_name": "2015_05476.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 862,
+ "file_name": "2015_05477.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 863,
+ "file_name": "2015_05478.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 864,
+ "file_name": "2015_05479.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 865,
+ "file_name": "2015_05480.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 866,
+ "file_name": "2015_05481.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 867,
+ "file_name": "2015_05482.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 868,
+ "file_name": "2015_05483.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 869,
+ "file_name": "2015_05484.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 870,
+ "file_name": "2015_05485.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 871,
+ "file_name": "2015_05486.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 872,
+ "file_name": "2015_05487.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 873,
+ "file_name": "2015_05488.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 874,
+ "file_name": "2015_05489.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 875,
+ "file_name": "2015_05490.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 876,
+ "file_name": "2015_05491.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 877,
+ "file_name": "2015_05492.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 878,
+ "file_name": "2015_05493.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 879,
+ "file_name": "2015_05494.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 880,
+ "file_name": "2015_05495.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 881,
+ "file_name": "2015_05496.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 882,
+ "file_name": "2015_05497.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 883,
+ "file_name": "2015_05498.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 884,
+ "file_name": "2015_05499.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 885,
+ "file_name": "2015_05500.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 886,
+ "file_name": "2015_05501.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 887,
+ "file_name": "2015_05502.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 888,
+ "file_name": "2015_05503.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 889,
+ "file_name": "2015_05504.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 890,
+ "file_name": "2015_05505.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 891,
+ "file_name": "2015_05506.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 892,
+ "file_name": "2015_05507.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 893,
+ "file_name": "2015_05508.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 894,
+ "file_name": "2015_05509.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 895,
+ "file_name": "2015_05510.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 896,
+ "file_name": "2015_05511.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 897,
+ "file_name": "2015_05512.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 898,
+ "file_name": "2015_05513.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 899,
+ "file_name": "2015_05514.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 900,
+ "file_name": "2015_05515.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 901,
+ "file_name": "2015_05516.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 902,
+ "file_name": "2015_05517.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 903,
+ "file_name": "2015_05518.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 904,
+ "file_name": "2015_05519.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 905,
+ "file_name": "2015_05520.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 906,
+ "file_name": "2015_05521.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 907,
+ "file_name": "2015_05522.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 908,
+ "file_name": "2015_05523.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 909,
+ "file_name": "2015_05524.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 910,
+ "file_name": "2015_05525.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 911,
+ "file_name": "2015_05526.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 912,
+ "file_name": "2015_05527.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 913,
+ "file_name": "2015_05528.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 914,
+ "file_name": "2015_05529.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 915,
+ "file_name": "2015_05530.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 916,
+ "file_name": "2015_05531.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 917,
+ "file_name": "2015_05532.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 918,
+ "file_name": "2015_05533.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 919,
+ "file_name": "2015_05534.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 920,
+ "file_name": "2015_05535.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 921,
+ "file_name": "2015_05536.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 922,
+ "file_name": "2015_05537.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 923,
+ "file_name": "2015_05538.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 924,
+ "file_name": "2015_05539.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 925,
+ "file_name": "2015_05540.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 926,
+ "file_name": "2015_05541.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 927,
+ "file_name": "2015_05542.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 928,
+ "file_name": "2015_05543.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 929,
+ "file_name": "2015_05544.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 930,
+ "file_name": "2015_05545.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 931,
+ "file_name": "2015_05546.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 932,
+ "file_name": "2015_05547.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 933,
+ "file_name": "2015_05548.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 934,
+ "file_name": "2015_05549.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 935,
+ "file_name": "2015_05550.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 936,
+ "file_name": "2015_05551.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 937,
+ "file_name": "2015_05552.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 938,
+ "file_name": "2015_05553.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 939,
+ "file_name": "2015_05554.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 940,
+ "file_name": "2015_05555.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 941,
+ "file_name": "2015_05556.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 942,
+ "file_name": "2015_05557.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 943,
+ "file_name": "2015_05558.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 944,
+ "file_name": "2015_05559.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 945,
+ "file_name": "2015_05560.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 946,
+ "file_name": "2015_05561.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 947,
+ "file_name": "2015_05562.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 948,
+ "file_name": "2015_05563.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 949,
+ "file_name": "2015_05564.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 950,
+ "file_name": "2015_05565.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 951,
+ "file_name": "2015_05566.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 952,
+ "file_name": "2015_05567.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 953,
+ "file_name": "2015_05568.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 954,
+ "file_name": "2015_05569.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 955,
+ "file_name": "2015_05570.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 956,
+ "file_name": "2015_05571.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 957,
+ "file_name": "2015_05572.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 958,
+ "file_name": "2015_05573.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 959,
+ "file_name": "2015_05574.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 960,
+ "file_name": "2015_05575.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 961,
+ "file_name": "2015_05576.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 962,
+ "file_name": "2015_05577.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 963,
+ "file_name": "2015_05578.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 964,
+ "file_name": "2015_05579.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 965,
+ "file_name": "2015_05580.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 966,
+ "file_name": "2015_05581.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 967,
+ "file_name": "2015_05582.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 968,
+ "file_name": "2015_05583.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 969,
+ "file_name": "2015_05584.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 970,
+ "file_name": "2015_05585.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 971,
+ "file_name": "2015_05586.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 972,
+ "file_name": "2015_05587.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 973,
+ "file_name": "2015_05588.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 974,
+ "file_name": "2015_05589.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 975,
+ "file_name": "2015_05590.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 976,
+ "file_name": "2015_05591.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 977,
+ "file_name": "2015_05592.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 978,
+ "file_name": "2015_05593.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 979,
+ "file_name": "2015_05594.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 980,
+ "file_name": "2015_05595.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 981,
+ "file_name": "2015_05596.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 982,
+ "file_name": "2015_05597.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 983,
+ "file_name": "2015_05598.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 984,
+ "file_name": "2015_05599.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 985,
+ "file_name": "2015_05600.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 986,
+ "file_name": "2015_05601.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 987,
+ "file_name": "2015_05602.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 988,
+ "file_name": "2015_05603.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 989,
+ "file_name": "2015_05604.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 990,
+ "file_name": "2015_05605.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 991,
+ "file_name": "2015_05606.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 992,
+ "file_name": "2015_05607.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 993,
+ "file_name": "2015_05608.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 994,
+ "file_name": "2015_05609.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 995,
+ "file_name": "2015_05610.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 996,
+ "file_name": "2015_05611.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 997,
+ "file_name": "2015_05612.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 998,
+ "file_name": "2015_05613.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 999,
+ "file_name": "2015_05614.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1000,
+ "file_name": "2015_05615.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1001,
+ "file_name": "2015_05616.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1002,
+ "file_name": "2015_05617.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1003,
+ "file_name": "2015_05618.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1004,
+ "file_name": "2015_05619.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1005,
+ "file_name": "2015_05620.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1006,
+ "file_name": "2015_05621.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1007,
+ "file_name": "2015_05622.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1008,
+ "file_name": "2015_05623.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1009,
+ "file_name": "2015_05624.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1010,
+ "file_name": "2015_05625.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1011,
+ "file_name": "2015_05626.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1012,
+ "file_name": "2015_05627.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1013,
+ "file_name": "2015_05628.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1014,
+ "file_name": "2015_05629.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1015,
+ "file_name": "2015_05630.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1016,
+ "file_name": "2015_05631.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1017,
+ "file_name": "2015_05632.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1018,
+ "file_name": "2015_05633.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1019,
+ "file_name": "2015_05634.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1020,
+ "file_name": "2015_05635.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1021,
+ "file_name": "2015_05636.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1022,
+ "file_name": "2015_05637.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1023,
+ "file_name": "2015_05638.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1024,
+ "file_name": "2015_05639.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1025,
+ "file_name": "2015_05640.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1026,
+ "file_name": "2015_05641.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1027,
+ "file_name": "2015_05642.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1028,
+ "file_name": "2015_05643.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1029,
+ "file_name": "2015_05644.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1030,
+ "file_name": "2015_05645.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1031,
+ "file_name": "2015_05646.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1032,
+ "file_name": "2015_05647.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1033,
+ "file_name": "2015_05648.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1034,
+ "file_name": "2015_05649.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1035,
+ "file_name": "2015_05650.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1036,
+ "file_name": "2015_05651.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1037,
+ "file_name": "2015_05652.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1038,
+ "file_name": "2015_05653.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1039,
+ "file_name": "2015_05654.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1040,
+ "file_name": "2015_05655.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1041,
+ "file_name": "2015_05656.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1042,
+ "file_name": "2015_05657.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1043,
+ "file_name": "2015_05658.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1044,
+ "file_name": "2015_05659.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1045,
+ "file_name": "2015_05660.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1046,
+ "file_name": "2015_05661.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1047,
+ "file_name": "2015_05662.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1048,
+ "file_name": "2015_05663.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1049,
+ "file_name": "2015_05664.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1050,
+ "file_name": "2015_05665.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1051,
+ "file_name": "2015_05666.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1052,
+ "file_name": "2015_05667.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1053,
+ "file_name": "2015_05668.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1054,
+ "file_name": "2015_05669.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1055,
+ "file_name": "2015_05670.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1056,
+ "file_name": "2015_05671.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1057,
+ "file_name": "2015_05672.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1058,
+ "file_name": "2015_05673.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1059,
+ "file_name": "2015_05674.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1060,
+ "file_name": "2015_05675.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1061,
+ "file_name": "2015_05676.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1062,
+ "file_name": "2015_05677.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1063,
+ "file_name": "2015_05678.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1064,
+ "file_name": "2015_05679.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1065,
+ "file_name": "2015_05680.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1066,
+ "file_name": "2015_05681.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1067,
+ "file_name": "2015_05682.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1068,
+ "file_name": "2015_05683.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1069,
+ "file_name": "2015_05684.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1070,
+ "file_name": "2015_05685.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1071,
+ "file_name": "2015_05686.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1072,
+ "file_name": "2015_05687.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1073,
+ "file_name": "2015_05688.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1074,
+ "file_name": "2015_05689.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1075,
+ "file_name": "2015_05690.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1076,
+ "file_name": "2015_05691.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1077,
+ "file_name": "2015_05692.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1078,
+ "file_name": "2015_05693.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1079,
+ "file_name": "2015_05694.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1080,
+ "file_name": "2015_05695.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1081,
+ "file_name": "2015_05696.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1082,
+ "file_name": "2015_05697.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1083,
+ "file_name": "2015_05698.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1084,
+ "file_name": "2015_05699.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1085,
+ "file_name": "2015_05700.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1086,
+ "file_name": "2015_05701.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1087,
+ "file_name": "2015_05702.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1088,
+ "file_name": "2015_05703.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1089,
+ "file_name": "2015_05704.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1090,
+ "file_name": "2015_05705.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1091,
+ "file_name": "2015_05706.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1092,
+ "file_name": "2015_05707.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1093,
+ "file_name": "2015_05708.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1094,
+ "file_name": "2015_05709.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1095,
+ "file_name": "2015_05710.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1096,
+ "file_name": "2015_05711.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1097,
+ "file_name": "2015_05712.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1098,
+ "file_name": "2015_05713.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1099,
+ "file_name": "2015_05714.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1100,
+ "file_name": "2015_05715.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1101,
+ "file_name": "2015_05716.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1102,
+ "file_name": "2015_05717.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1103,
+ "file_name": "2015_05718.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1104,
+ "file_name": "2015_05719.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1105,
+ "file_name": "2015_05720.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1106,
+ "file_name": "2015_05721.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1107,
+ "file_name": "2015_05722.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1108,
+ "file_name": "2015_05723.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1109,
+ "file_name": "2015_05724.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1110,
+ "file_name": "2015_05725.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1111,
+ "file_name": "2015_05726.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1112,
+ "file_name": "2015_05727.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1113,
+ "file_name": "2015_05728.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1114,
+ "file_name": "2015_05729.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1115,
+ "file_name": "2015_05730.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1116,
+ "file_name": "2015_05731.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1117,
+ "file_name": "2015_05732.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1118,
+ "file_name": "2015_05733.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1119,
+ "file_name": "2015_05734.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1120,
+ "file_name": "2015_05735.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1121,
+ "file_name": "2015_05736.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1122,
+ "file_name": "2015_05737.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1123,
+ "file_name": "2015_05738.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1124,
+ "file_name": "2015_05739.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1125,
+ "file_name": "2015_05740.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1126,
+ "file_name": "2015_05741.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1127,
+ "file_name": "2015_05742.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1128,
+ "file_name": "2015_05743.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1129,
+ "file_name": "2015_05744.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1130,
+ "file_name": "2015_05995.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1131,
+ "file_name": "2015_05996.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1132,
+ "file_name": "2015_05997.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1133,
+ "file_name": "2015_05998.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1134,
+ "file_name": "2015_05999.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1135,
+ "file_name": "2015_06000.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1136,
+ "file_name": "2015_06001.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1137,
+ "file_name": "2015_06002.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1138,
+ "file_name": "2015_06003.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1139,
+ "file_name": "2015_06004.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1140,
+ "file_name": "2015_06005.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1141,
+ "file_name": "2015_06006.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1142,
+ "file_name": "2015_06007.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1143,
+ "file_name": "2015_06008.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1144,
+ "file_name": "2015_06009.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1145,
+ "file_name": "2015_06010.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1146,
+ "file_name": "2015_06011.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1147,
+ "file_name": "2015_06012.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1148,
+ "file_name": "2015_06013.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1149,
+ "file_name": "2015_06014.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1150,
+ "file_name": "2015_06015.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1151,
+ "file_name": "2015_06016.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1152,
+ "file_name": "2015_06017.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1153,
+ "file_name": "2015_06018.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1154,
+ "file_name": "2015_06019.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1155,
+ "file_name": "2015_06020.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1156,
+ "file_name": "2015_06021.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1157,
+ "file_name": "2015_06022.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1158,
+ "file_name": "2015_06023.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1159,
+ "file_name": "2015_06024.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1160,
+ "file_name": "2015_06025.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1161,
+ "file_name": "2015_06026.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1162,
+ "file_name": "2015_06027.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1163,
+ "file_name": "2015_06028.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1164,
+ "file_name": "2015_06029.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1165,
+ "file_name": "2015_06030.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1166,
+ "file_name": "2015_06031.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1167,
+ "file_name": "2015_06032.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1168,
+ "file_name": "2015_06033.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1169,
+ "file_name": "2015_06034.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1170,
+ "file_name": "2015_06035.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1171,
+ "file_name": "2015_06036.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1172,
+ "file_name": "2015_06037.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1173,
+ "file_name": "2015_06038.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1174,
+ "file_name": "2015_06039.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1175,
+ "file_name": "2015_06040.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1176,
+ "file_name": "2015_06041.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1177,
+ "file_name": "2015_06042.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1178,
+ "file_name": "2015_06043.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1179,
+ "file_name": "2015_06044.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1180,
+ "file_name": "2015_06045.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1181,
+ "file_name": "2015_06046.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1182,
+ "file_name": "2015_06047.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1183,
+ "file_name": "2015_06048.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1184,
+ "file_name": "2015_06049.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1185,
+ "file_name": "2015_06050.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1186,
+ "file_name": "2015_06051.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1187,
+ "file_name": "2015_06052.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1188,
+ "file_name": "2015_06053.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1189,
+ "file_name": "2015_06054.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1190,
+ "file_name": "2015_06055.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1191,
+ "file_name": "2015_06056.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1192,
+ "file_name": "2015_06057.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1193,
+ "file_name": "2015_06058.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1194,
+ "file_name": "2015_06059.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1195,
+ "file_name": "2015_06060.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1196,
+ "file_name": "2015_06061.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1197,
+ "file_name": "2015_06062.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1198,
+ "file_name": "2015_06063.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1199,
+ "file_name": "2015_06064.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1200,
+ "file_name": "2015_06065.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1201,
+ "file_name": "2015_06066.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1202,
+ "file_name": "2015_06067.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1203,
+ "file_name": "2015_06068.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1204,
+ "file_name": "2015_06069.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1205,
+ "file_name": "2015_06070.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1206,
+ "file_name": "2015_06071.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1207,
+ "file_name": "2015_06072.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1208,
+ "file_name": "2015_06073.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1209,
+ "file_name": "2015_06074.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1210,
+ "file_name": "2015_06075.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1211,
+ "file_name": "2015_06076.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1212,
+ "file_name": "2015_06077.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1213,
+ "file_name": "2015_06078.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1214,
+ "file_name": "2015_06079.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1215,
+ "file_name": "2015_06080.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1216,
+ "file_name": "2015_06081.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1217,
+ "file_name": "2015_06082.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1218,
+ "file_name": "2015_06083.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1219,
+ "file_name": "2015_06084.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1220,
+ "file_name": "2015_06085.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1221,
+ "file_name": "2015_06086.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1222,
+ "file_name": "2015_06087.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1223,
+ "file_name": "2015_06088.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1224,
+ "file_name": "2015_06089.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1225,
+ "file_name": "2015_06090.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1226,
+ "file_name": "2015_06091.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1227,
+ "file_name": "2015_06092.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1228,
+ "file_name": "2015_06093.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1229,
+ "file_name": "2015_06094.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1230,
+ "file_name": "2015_06095.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1231,
+ "file_name": "2015_06096.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1232,
+ "file_name": "2015_06097.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1233,
+ "file_name": "2015_06098.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1234,
+ "file_name": "2015_06099.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1235,
+ "file_name": "2015_06100.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1236,
+ "file_name": "2015_06101.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1237,
+ "file_name": "2015_06102.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1238,
+ "file_name": "2015_06103.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1239,
+ "file_name": "2015_06104.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1240,
+ "file_name": "2015_06105.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1241,
+ "file_name": "2015_06106.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1242,
+ "file_name": "2015_06107.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1243,
+ "file_name": "2015_06108.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1244,
+ "file_name": "2015_06109.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1245,
+ "file_name": "2015_06110.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1246,
+ "file_name": "2015_06111.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1247,
+ "file_name": "2015_06112.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1248,
+ "file_name": "2015_06113.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1249,
+ "file_name": "2015_06114.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1250,
+ "file_name": "2015_06115.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1251,
+ "file_name": "2015_06116.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1252,
+ "file_name": "2015_06117.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1253,
+ "file_name": "2015_06118.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1254,
+ "file_name": "2015_06119.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1255,
+ "file_name": "2015_06120.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1256,
+ "file_name": "2015_06121.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1257,
+ "file_name": "2015_06122.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1258,
+ "file_name": "2015_06123.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1259,
+ "file_name": "2015_06124.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1260,
+ "file_name": "2015_06125.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1261,
+ "file_name": "2015_06126.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1262,
+ "file_name": "2015_06127.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1263,
+ "file_name": "2015_06128.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1264,
+ "file_name": "2015_06129.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1265,
+ "file_name": "2015_06130.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1266,
+ "file_name": "2015_06131.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1267,
+ "file_name": "2015_06132.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1268,
+ "file_name": "2015_06133.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1269,
+ "file_name": "2015_06134.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1270,
+ "file_name": "2015_06135.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1271,
+ "file_name": "2015_06136.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1272,
+ "file_name": "2015_06137.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1273,
+ "file_name": "2015_06138.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1274,
+ "file_name": "2015_06139.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1275,
+ "file_name": "2015_06140.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1276,
+ "file_name": "2015_06141.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1277,
+ "file_name": "2015_06142.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1278,
+ "file_name": "2015_06143.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1279,
+ "file_name": "2015_06144.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1280,
+ "file_name": "2015_06145.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1281,
+ "file_name": "2015_06146.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1282,
+ "file_name": "2015_06147.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1283,
+ "file_name": "2015_06148.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1284,
+ "file_name": "2015_06149.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1285,
+ "file_name": "2015_06150.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1286,
+ "file_name": "2015_06151.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1287,
+ "file_name": "2015_06152.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1288,
+ "file_name": "2015_06153.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1289,
+ "file_name": "2015_06154.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1290,
+ "file_name": "2015_06155.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1291,
+ "file_name": "2015_06156.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1292,
+ "file_name": "2015_06157.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1293,
+ "file_name": "2015_06158.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1294,
+ "file_name": "2015_06159.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1295,
+ "file_name": "2015_06160.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1296,
+ "file_name": "2015_06161.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1297,
+ "file_name": "2015_06162.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1298,
+ "file_name": "2015_06163.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1299,
+ "file_name": "2015_06164.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1300,
+ "file_name": "2015_06165.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1301,
+ "file_name": "2015_06166.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1302,
+ "file_name": "2015_06167.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1303,
+ "file_name": "2015_06168.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1304,
+ "file_name": "2015_06169.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1305,
+ "file_name": "2015_06170.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1306,
+ "file_name": "2015_06171.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1307,
+ "file_name": "2015_06172.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1308,
+ "file_name": "2015_06173.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1309,
+ "file_name": "2015_06174.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1310,
+ "file_name": "2015_06175.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1311,
+ "file_name": "2015_06176.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1312,
+ "file_name": "2015_06177.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1313,
+ "file_name": "2015_06178.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1314,
+ "file_name": "2015_06179.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1315,
+ "file_name": "2015_06180.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1316,
+ "file_name": "2015_06181.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1317,
+ "file_name": "2015_06182.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1318,
+ "file_name": "2015_06183.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1319,
+ "file_name": "2015_06184.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1320,
+ "file_name": "2015_06185.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1321,
+ "file_name": "2015_06186.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1322,
+ "file_name": "2015_06187.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1323,
+ "file_name": "2015_06188.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1324,
+ "file_name": "2015_06189.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1325,
+ "file_name": "2015_06190.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1326,
+ "file_name": "2015_06191.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1327,
+ "file_name": "2015_06192.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1328,
+ "file_name": "2015_06193.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1329,
+ "file_name": "2015_06194.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1330,
+ "file_name": "2015_06195.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1331,
+ "file_name": "2015_06196.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1332,
+ "file_name": "2015_06197.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1333,
+ "file_name": "2015_06198.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1334,
+ "file_name": "2015_06199.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1335,
+ "file_name": "2015_06200.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1336,
+ "file_name": "2015_06201.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1337,
+ "file_name": "2015_06202.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1338,
+ "file_name": "2015_06203.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1339,
+ "file_name": "2015_06204.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1340,
+ "file_name": "2015_06205.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1341,
+ "file_name": "2015_06206.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1342,
+ "file_name": "2015_06207.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1343,
+ "file_name": "2015_06208.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1344,
+ "file_name": "2015_06209.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1345,
+ "file_name": "2015_06210.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1346,
+ "file_name": "2015_06211.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1347,
+ "file_name": "2015_06212.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1348,
+ "file_name": "2015_06213.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1349,
+ "file_name": "2015_06214.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1350,
+ "file_name": "2015_06215.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1351,
+ "file_name": "2015_06216.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1352,
+ "file_name": "2015_06217.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1353,
+ "file_name": "2015_06218.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1354,
+ "file_name": "2015_06219.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1355,
+ "file_name": "2015_06220.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1356,
+ "file_name": "2015_06221.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1357,
+ "file_name": "2015_06222.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1358,
+ "file_name": "2015_06223.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1359,
+ "file_name": "2015_06224.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1360,
+ "file_name": "2015_06225.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1361,
+ "file_name": "2015_06226.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1362,
+ "file_name": "2015_06227.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1363,
+ "file_name": "2015_06228.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1364,
+ "file_name": "2015_06229.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1365,
+ "file_name": "2015_06230.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1366,
+ "file_name": "2015_06231.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1367,
+ "file_name": "2015_06232.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1368,
+ "file_name": "2015_06233.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1369,
+ "file_name": "2015_06234.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1370,
+ "file_name": "2015_06235.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1371,
+ "file_name": "2015_06236.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1372,
+ "file_name": "2015_06237.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1373,
+ "file_name": "2015_06238.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1374,
+ "file_name": "2015_06239.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1375,
+ "file_name": "2015_06240.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1376,
+ "file_name": "2015_06241.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1377,
+ "file_name": "2015_06242.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1378,
+ "file_name": "2015_06243.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1379,
+ "file_name": "2015_06244.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1380,
+ "file_name": "2015_06245.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1381,
+ "file_name": "2016_07356.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1382,
+ "file_name": "2017_07361.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1383,
+ "file_name": "2015_06562.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1384,
+ "file_name": "2015_06567.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1385,
+ "file_name": "2015_06568.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1386,
+ "file_name": "2015_06569.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1387,
+ "file_name": "2015_06575.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1388,
+ "file_name": "2015_06581.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1389,
+ "file_name": "2015_06595.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1390,
+ "file_name": "2015_06606.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1391,
+ "file_name": "2015_06618.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1392,
+ "file_name": "2015_06619.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1393,
+ "file_name": "2015_06620.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1394,
+ "file_name": "2015_06624.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1395,
+ "file_name": "2015_06627.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1396,
+ "file_name": "2015_06632.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1397,
+ "file_name": "2015_06637.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1398,
+ "file_name": "2015_06644.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1399,
+ "file_name": "2015_06647.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1400,
+ "file_name": "2015_06655.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1401,
+ "file_name": "2015_06699.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1402,
+ "file_name": "2015_06704.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1403,
+ "file_name": "2015_06724.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1404,
+ "file_name": "2015_06725.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1405,
+ "file_name": "2015_06741.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1406,
+ "file_name": "2015_06743.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1407,
+ "file_name": "2015_06764.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1408,
+ "file_name": "2015_06775.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1409,
+ "file_name": "2015_06808.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1410,
+ "file_name": "2015_06809.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1411,
+ "file_name": "2017_07363.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1412,
+ "file_name": "2015_07102.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1413,
+ "file_name": "2015_07103.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1414,
+ "file_name": "2015_07104.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1415,
+ "file_name": "2015_07105.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1416,
+ "file_name": "2015_07106.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1417,
+ "file_name": "2015_07107.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1418,
+ "file_name": "2015_07108.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1419,
+ "file_name": "2015_07109.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1420,
+ "file_name": "2015_07110.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1421,
+ "file_name": "2015_07111.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1422,
+ "file_name": "2015_07112.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1423,
+ "file_name": "2015_07113.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1424,
+ "file_name": "2015_07114.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1425,
+ "file_name": "2015_07115.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1426,
+ "file_name": "2015_07117.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1427,
+ "file_name": "2015_07118.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1428,
+ "file_name": "2015_07119.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1429,
+ "file_name": "2015_07120.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1430,
+ "file_name": "2015_07121.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1431,
+ "file_name": "2015_07122.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1432,
+ "file_name": "2015_07123.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1433,
+ "file_name": "2015_07124.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1434,
+ "file_name": "2015_07125.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1435,
+ "file_name": "2015_07126.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1436,
+ "file_name": "2015_07127.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1437,
+ "file_name": "2015_07128.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1438,
+ "file_name": "2015_07129.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1439,
+ "file_name": "2015_07130.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1440,
+ "file_name": "2015_07131.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1441,
+ "file_name": "2015_07132.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1442,
+ "file_name": "2015_07133.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1443,
+ "file_name": "2015_07134.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1444,
+ "file_name": "2015_07135.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1445,
+ "file_name": "2015_07136.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1446,
+ "file_name": "2015_07137.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1447,
+ "file_name": "2015_07138.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1448,
+ "file_name": "2015_07139.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1449,
+ "file_name": "2015_07140.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1450,
+ "file_name": "2015_07141.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1451,
+ "file_name": "2015_07142.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1452,
+ "file_name": "2015_07143.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1453,
+ "file_name": "2015_07144.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1454,
+ "file_name": "2015_07145.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1455,
+ "file_name": "2015_07146.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1456,
+ "file_name": "2015_07147.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1457,
+ "file_name": "2015_07148.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1458,
+ "file_name": "2015_07149.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1459,
+ "file_name": "2015_07150.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1460,
+ "file_name": "2015_07151.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1461,
+ "file_name": "2015_07152.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1462,
+ "file_name": "2015_07153.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1463,
+ "file_name": "2015_07154.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1464,
+ "file_name": "2015_07155.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1465,
+ "file_name": "2015_07156.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1466,
+ "file_name": "2015_07157.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1467,
+ "file_name": "2015_07158.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1468,
+ "file_name": "2015_07159.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1469,
+ "file_name": "2015_07160.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1470,
+ "file_name": "2015_07161.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1471,
+ "file_name": "2015_07162.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1472,
+ "file_name": "2015_07163.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1473,
+ "file_name": "2015_07164.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1474,
+ "file_name": "2015_07165.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1475,
+ "file_name": "2015_07166.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1476,
+ "file_name": "2015_07167.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1477,
+ "file_name": "2015_07168.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1478,
+ "file_name": "2015_07169.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1479,
+ "file_name": "2015_07170.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1480,
+ "file_name": "2015_07171.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1481,
+ "file_name": "2015_07172.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1482,
+ "file_name": "2015_07173.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1483,
+ "file_name": "2015_07174.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1484,
+ "file_name": "2015_07175.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1485,
+ "file_name": "2015_07176.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1486,
+ "file_name": "2015_07177.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1487,
+ "file_name": "2015_07178.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1488,
+ "file_name": "2015_07179.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1489,
+ "file_name": "2015_07180.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1490,
+ "file_name": "2015_07181.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1491,
+ "file_name": "2015_07182.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1492,
+ "file_name": "2015_07183.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1493,
+ "file_name": "2015_07185.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1494,
+ "file_name": "2015_07186.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1495,
+ "file_name": "2015_07187.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1496,
+ "file_name": "2015_07188.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1497,
+ "file_name": "2015_07189.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1498,
+ "file_name": "2015_07190.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1499,
+ "file_name": "2015_07191.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1500,
+ "file_name": "2015_07192.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1501,
+ "file_name": "2015_07193.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1502,
+ "file_name": "2015_07194.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1503,
+ "file_name": "2015_07195.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1504,
+ "file_name": "2015_07196.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1505,
+ "file_name": "2015_07197.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1506,
+ "file_name": "2015_07198.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1507,
+ "file_name": "2015_07199.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1508,
+ "file_name": "2015_07200.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1509,
+ "file_name": "2015_07201.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1510,
+ "file_name": "2015_07202.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1511,
+ "file_name": "2015_07203.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1512,
+ "file_name": "2015_07204.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1513,
+ "file_name": "2015_07205.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1514,
+ "file_name": "2015_07206.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1515,
+ "file_name": "2015_07207.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1516,
+ "file_name": "2015_07208.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1517,
+ "file_name": "2015_07209.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1518,
+ "file_name": "2015_07210.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1519,
+ "file_name": "2015_07211.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1520,
+ "file_name": "2015_07212.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1521,
+ "file_name": "2015_07213.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1522,
+ "file_name": "2015_07214.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1523,
+ "file_name": "2015_07215.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1524,
+ "file_name": "2015_07216.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1525,
+ "file_name": "2015_07217.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1526,
+ "file_name": "2015_07218.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1527,
+ "file_name": "2015_07219.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1528,
+ "file_name": "2015_07220.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1529,
+ "file_name": "2015_07221.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1530,
+ "file_name": "2015_07222.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1531,
+ "file_name": "2015_07223.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1532,
+ "file_name": "2015_07224.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1533,
+ "file_name": "2015_07225.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1534,
+ "file_name": "2015_07226.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1535,
+ "file_name": "2015_07227.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1536,
+ "file_name": "2015_07228.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1537,
+ "file_name": "2015_07229.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1538,
+ "file_name": "2015_07230.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1539,
+ "file_name": "2015_07231.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1540,
+ "file_name": "2015_07232.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1541,
+ "file_name": "2015_07233.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1542,
+ "file_name": "2015_07234.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1543,
+ "file_name": "2015_07235.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1544,
+ "file_name": "2015_07236.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1545,
+ "file_name": "2015_07237.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1546,
+ "file_name": "2015_07238.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1547,
+ "file_name": "2015_07239.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1548,
+ "file_name": "2015_07240.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1549,
+ "file_name": "2015_07241.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1550,
+ "file_name": "2015_07242.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1551,
+ "file_name": "2015_07243.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1552,
+ "file_name": "2015_07244.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1553,
+ "file_name": "2015_07245.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1554,
+ "file_name": "2015_07246.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1555,
+ "file_name": "2015_07247.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1556,
+ "file_name": "2015_07248.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1557,
+ "file_name": "2015_07249.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1558,
+ "file_name": "2015_07250.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1559,
+ "file_name": "2015_07251.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1560,
+ "file_name": "2015_07252.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1561,
+ "file_name": "2015_07253.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1562,
+ "file_name": "2015_07254.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1563,
+ "file_name": "2015_07255.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1564,
+ "file_name": "2015_07256.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1565,
+ "file_name": "2015_07257.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1566,
+ "file_name": "2015_07258.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1567,
+ "file_name": "2015_07259.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1568,
+ "file_name": "2015_07260.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1569,
+ "file_name": "2015_07261.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1570,
+ "file_name": "2015_07262.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1571,
+ "file_name": "2015_07263.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1572,
+ "file_name": "2015_07264.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1573,
+ "file_name": "2015_07265.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1574,
+ "file_name": "2015_07266.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1575,
+ "file_name": "2015_07267.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1576,
+ "file_name": "2015_07268.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1577,
+ "file_name": "2015_07269.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1578,
+ "file_name": "2015_07270.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1579,
+ "file_name": "2015_07271.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1580,
+ "file_name": "2015_07272.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1581,
+ "file_name": "2015_07273.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1582,
+ "file_name": "2015_07274.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1583,
+ "file_name": "2015_07275.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1584,
+ "file_name": "2015_07276.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1585,
+ "file_name": "2015_07277.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1586,
+ "file_name": "2015_07278.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1587,
+ "file_name": "2015_07279.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1588,
+ "file_name": "2015_07280.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1589,
+ "file_name": "2015_07281.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1590,
+ "file_name": "2015_07282.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1591,
+ "file_name": "2015_07283.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1592,
+ "file_name": "2015_07284.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1593,
+ "file_name": "2015_07285.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1594,
+ "file_name": "2015_07286.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1595,
+ "file_name": "2015_07287.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1596,
+ "file_name": "2015_07288.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1597,
+ "file_name": "2015_07289.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1598,
+ "file_name": "2015_07290.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1599,
+ "file_name": "2015_07291.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1600,
+ "file_name": "2015_07292.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1601,
+ "file_name": "2015_07293.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1602,
+ "file_name": "2015_07294.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1603,
+ "file_name": "2015_07295.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1604,
+ "file_name": "2015_07296.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1605,
+ "file_name": "2015_07297.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1606,
+ "file_name": "2015_07298.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1607,
+ "file_name": "2015_07300.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1608,
+ "file_name": "2015_07301.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1609,
+ "file_name": "2015_07302.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1610,
+ "file_name": "2015_07303.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1611,
+ "file_name": "2015_07304.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1612,
+ "file_name": "2015_07305.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1613,
+ "file_name": "2015_07306.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1614,
+ "file_name": "2015_07307.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1615,
+ "file_name": "2015_07308.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1616,
+ "file_name": "2015_07309.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1617,
+ "file_name": "2015_07310.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1618,
+ "file_name": "2015_07311.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1619,
+ "file_name": "2015_07312.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1620,
+ "file_name": "2015_07313.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1621,
+ "file_name": "2015_07314.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1622,
+ "file_name": "2015_07315.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1623,
+ "file_name": "2015_07316.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1624,
+ "file_name": "2015_07317.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1625,
+ "file_name": "2015_07318.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1626,
+ "file_name": "2015_07319.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1627,
+ "file_name": "2015_07320.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1628,
+ "file_name": "2015_07321.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1629,
+ "file_name": "2015_07322.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1630,
+ "file_name": "2015_07323.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1631,
+ "file_name": "2015_07324.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1632,
+ "file_name": "2015_07325.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1633,
+ "file_name": "2015_07326.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1634,
+ "file_name": "2015_07327.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1635,
+ "file_name": "2015_07328.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1636,
+ "file_name": "2015_07329.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1637,
+ "file_name": "2015_07330.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1638,
+ "file_name": "2015_07331.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1639,
+ "file_name": "2015_07332.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1640,
+ "file_name": "2015_07333.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1641,
+ "file_name": "2015_07334.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1642,
+ "file_name": "2015_07335.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1643,
+ "file_name": "2015_07336.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1644,
+ "file_name": "2015_07337.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1645,
+ "file_name": "2015_07338.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1646,
+ "file_name": "2015_07339.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1647,
+ "file_name": "2015_07340.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1648,
+ "file_name": "2015_07341.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1649,
+ "file_name": "2015_07342.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1650,
+ "file_name": "2015_07343.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1651,
+ "file_name": "2015_07344.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1652,
+ "file_name": "2015_07345.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1653,
+ "file_name": "2015_07346.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1654,
+ "file_name": "2015_07347.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1655,
+ "file_name": "2015_07348.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1656,
+ "file_name": "2015_07349.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1657,
+ "file_name": "2015_07350.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1658,
+ "file_name": "2015_07351.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1659,
+ "file_name": "2015_07352.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1660,
+ "file_name": "2015_07353.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1661,
+ "file_name": "2015_07354.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1662,
+ "file_name": "2015_07355.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1663,
+ "file_name": "2017_07358.jpg",
+ "width": 512,
+ "height": 512
+ }
+ ],
+ "annotations": [
+ {
+ "id": 1,
+ "image_id": 1,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 445,
+ 150,
+ 65,
+ 43
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2,
+ "image_id": 2,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 340,
+ 262,
+ 40,
+ 71
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 3,
+ "image_id": 2,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 240,
+ 268,
+ 63,
+ 74
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 4,
+ "image_id": 3,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 145,
+ 262,
+ 56,
+ 57
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 5,
+ "image_id": 4,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 97,
+ 80,
+ 356,
+ 387
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 6,
+ "image_id": 5,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 8,
+ 176,
+ 141,
+ 194
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 7,
+ "image_id": 6,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 371,
+ 299,
+ 50,
+ 47
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 8,
+ "image_id": 7,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 110,
+ 182,
+ 24,
+ 57
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 9,
+ "image_id": 8,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 159,
+ 434,
+ 37,
+ 44
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 10,
+ "image_id": 8,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 393,
+ 431,
+ 36,
+ 47
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 11,
+ "image_id": 8,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 440,
+ 437,
+ 33,
+ 40
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 12,
+ "image_id": 8,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 104,
+ 475,
+ 56,
+ 20
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 13,
+ "image_id": 8,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 43,
+ 489,
+ 101,
+ 21
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 14,
+ "image_id": 9,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 375,
+ 79,
+ 81,
+ 64
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 15,
+ "image_id": 10,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 155,
+ 185,
+ 230,
+ 198
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 16,
+ "image_id": 11,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 290,
+ 453,
+ 45,
+ 44
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 17,
+ "image_id": 11,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 217,
+ 459,
+ 59,
+ 52
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 18,
+ "image_id": 11,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 460,
+ 56,
+ 49
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 19,
+ "image_id": 11,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 442,
+ 439,
+ 22,
+ 22
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 20,
+ "image_id": 12,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 399,
+ 119,
+ 91
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 21,
+ "image_id": 12,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 320,
+ 354,
+ 70,
+ 111
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 22,
+ "image_id": 12,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 416,
+ 332,
+ 85,
+ 145
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 23,
+ "image_id": 13,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 128,
+ 457,
+ 304
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 24,
+ "image_id": 14,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 381,
+ 127,
+ 81
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 25,
+ "image_id": 14,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 157,
+ 376,
+ 72,
+ 68
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 26,
+ "image_id": 15,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 132,
+ 185,
+ 112,
+ 186
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 27,
+ "image_id": 16,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 444,
+ 169,
+ 44,
+ 43
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 28,
+ "image_id": 17,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 169,
+ 368,
+ 326
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 29,
+ "image_id": 18,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 248,
+ 3,
+ 138,
+ 44
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 30,
+ "image_id": 19,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 177,
+ 102,
+ 71
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 31,
+ "image_id": 20,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 424,
+ 276,
+ 40,
+ 56
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 32,
+ "image_id": 20,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 141,
+ 241,
+ 109,
+ 47
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 33,
+ "image_id": 21,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 170,
+ 206,
+ 208,
+ 177
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 34,
+ "image_id": 22,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 113,
+ 256,
+ 395,
+ 236
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 35,
+ "image_id": 23,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 361,
+ 223,
+ 46,
+ 86
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 36,
+ "image_id": 23,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 380,
+ 233,
+ 123,
+ 170
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 37,
+ "image_id": 23,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 251,
+ 225,
+ 105,
+ 101
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 38,
+ "image_id": 24,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 178,
+ 128,
+ 29,
+ 27
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 39,
+ "image_id": 25,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 38,
+ 195,
+ 66,
+ 63
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 40,
+ "image_id": 25,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 117,
+ 221,
+ 53,
+ 76
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 41,
+ "image_id": 26,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 7,
+ 135,
+ 250
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 42,
+ "image_id": 27,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 133,
+ 234,
+ 37,
+ 81
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 43,
+ "image_id": 28,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 187,
+ 6,
+ 321,
+ 483
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 44,
+ "image_id": 29,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 53,
+ 161,
+ 399,
+ 226
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 45,
+ "image_id": 30,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 405,
+ 90,
+ 37,
+ 109
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 46,
+ "image_id": 30,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 266,
+ 187,
+ 183,
+ 274
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 47,
+ "image_id": 30,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 34,
+ 241,
+ 147,
+ 199
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 48,
+ "image_id": 31,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 354,
+ 319,
+ 22,
+ 33
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 49,
+ "image_id": 32,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 351,
+ 286,
+ 25,
+ 49
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 50,
+ "image_id": 32,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 484,
+ 282,
+ 27,
+ 60
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 51,
+ "image_id": 33,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 470,
+ 242,
+ 39,
+ 35
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 52,
+ "image_id": 34,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 248,
+ 72,
+ 118
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 53,
+ "image_id": 35,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 199,
+ 303,
+ 26,
+ 65
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 54,
+ "image_id": 36,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 352,
+ 65,
+ 48
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 55,
+ "image_id": 37,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 268,
+ 315,
+ 19,
+ 52
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 56,
+ "image_id": 38,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 466,
+ 215,
+ 33,
+ 32
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 57,
+ "image_id": 39,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 309,
+ 136,
+ 200,
+ 221
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 58,
+ "image_id": 40,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 245,
+ 500,
+ 255
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 59,
+ "image_id": 41,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 171,
+ 292,
+ 328,
+ 200
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 60,
+ "image_id": 42,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 160,
+ 505,
+ 346
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 61,
+ "image_id": 43,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 72,
+ 321,
+ 382,
+ 189
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 62,
+ "image_id": 44,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 294,
+ 402,
+ 118,
+ 106
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 63,
+ "image_id": 45,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 403,
+ 490,
+ 103
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 64,
+ "image_id": 46,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 331,
+ 497,
+ 175
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 65,
+ "image_id": 47,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 12,
+ 87,
+ 472,
+ 414
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 66,
+ "image_id": 48,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 140,
+ 280,
+ 279,
+ 186
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 67,
+ "image_id": 49,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 203,
+ 370,
+ 58,
+ 58
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 68,
+ "image_id": 50,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 21,
+ 149,
+ 484,
+ 354
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 69,
+ "image_id": 51,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 419,
+ 84,
+ 86
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 70,
+ "image_id": 52,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 144,
+ 310,
+ 224,
+ 194
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 71,
+ "image_id": 53,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 100,
+ 238,
+ 409,
+ 271
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 72,
+ "image_id": 54,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 1,
+ 292,
+ 368,
+ 152
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 73,
+ "image_id": 55,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 169,
+ 157,
+ 164
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 74,
+ "image_id": 55,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 301,
+ 502,
+ 208
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 75,
+ "image_id": 56,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 53,
+ 257,
+ 280,
+ 240
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 76,
+ "image_id": 57,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 7,
+ 423,
+ 420,
+ 83
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 77,
+ "image_id": 58,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 257,
+ 250,
+ 132,
+ 62
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 78,
+ "image_id": 58,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 45,
+ 304,
+ 342,
+ 191
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 79,
+ "image_id": 59,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 169,
+ 267,
+ 101,
+ 179
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 80,
+ "image_id": 59,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 8,
+ 280,
+ 66,
+ 218
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 81,
+ "image_id": 60,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 12,
+ 284,
+ 490,
+ 226
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 82,
+ "image_id": 61,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 12,
+ 315,
+ 223,
+ 186
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 83,
+ "image_id": 62,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 215,
+ 215,
+ 163
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 84,
+ "image_id": 63,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 252,
+ 507,
+ 258
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 85,
+ "image_id": 64,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 365,
+ 304,
+ 132
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 86,
+ "image_id": 65,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 32,
+ 216,
+ 76,
+ 127
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 87,
+ "image_id": 65,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 330,
+ 212,
+ 171
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 88,
+ "image_id": 66,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 281,
+ 261,
+ 227
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 89,
+ "image_id": 67,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 43,
+ 318,
+ 258,
+ 186
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 90,
+ "image_id": 68,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 19,
+ 378,
+ 288,
+ 120
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 91,
+ "image_id": 69,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 106,
+ 288,
+ 358,
+ 218
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 92,
+ "image_id": 70,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 268,
+ 503,
+ 240
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 93,
+ "image_id": 71,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 106,
+ 317,
+ 371,
+ 191
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 94,
+ "image_id": 72,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 109,
+ 273,
+ 387,
+ 236
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 95,
+ "image_id": 73,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 104,
+ 321,
+ 404,
+ 154
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 96,
+ "image_id": 74,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 321,
+ 242,
+ 46,
+ 53
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 97,
+ "image_id": 74,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 119,
+ 221,
+ 185,
+ 122
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 98,
+ "image_id": 75,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 307,
+ 167,
+ 135
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 99,
+ "image_id": 76,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 1,
+ 308,
+ 510,
+ 200
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 100,
+ "image_id": 77,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 18,
+ 304,
+ 492,
+ 203
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 101,
+ "image_id": 78,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 209,
+ 255,
+ 301,
+ 241
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 102,
+ "image_id": 79,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 101,
+ 307,
+ 166,
+ 149
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 103,
+ "image_id": 79,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 208,
+ 369,
+ 241,
+ 133
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 104,
+ "image_id": 80,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 290,
+ 324,
+ 199
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 105,
+ "image_id": 81,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 24,
+ 270,
+ 212,
+ 238
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 106,
+ "image_id": 82,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 143,
+ 335,
+ 351,
+ 167
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 107,
+ "image_id": 83,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 41,
+ 174,
+ 366,
+ 324
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 108,
+ "image_id": 84,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 440,
+ 12,
+ 71,
+ 62
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 109,
+ "image_id": 84,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 162,
+ 508,
+ 348
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 110,
+ "image_id": 85,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 51,
+ 502,
+ 459
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 111,
+ "image_id": 86,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 109,
+ 293,
+ 336,
+ 201
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 112,
+ "image_id": 87,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 147,
+ 265,
+ 361,
+ 235
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 113,
+ "image_id": 88,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 12,
+ 193,
+ 496,
+ 305
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 114,
+ "image_id": 89,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 8,
+ 77,
+ 496,
+ 432
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 115,
+ "image_id": 90,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 95,
+ 376,
+ 410
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 116,
+ "image_id": 91,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 75,
+ 239,
+ 430,
+ 266
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 117,
+ "image_id": 92,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 136,
+ 405,
+ 303,
+ 98
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 118,
+ "image_id": 93,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 57,
+ 341,
+ 436,
+ 159
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 119,
+ "image_id": 94,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 413,
+ 508,
+ 95
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 120,
+ "image_id": 95,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 300,
+ 422,
+ 130
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 121,
+ "image_id": 96,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 181,
+ 506,
+ 328
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 122,
+ "image_id": 97,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 11,
+ 177,
+ 496,
+ 331
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 123,
+ "image_id": 98,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 54,
+ 303,
+ 342,
+ 199
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 124,
+ "image_id": 99,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 130,
+ 240,
+ 204,
+ 205
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 125,
+ "image_id": 100,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 46,
+ 167,
+ 359,
+ 240
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 126,
+ "image_id": 101,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 87,
+ 94,
+ 403,
+ 410
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 127,
+ "image_id": 102,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 370,
+ 15,
+ 137,
+ 172
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 128,
+ "image_id": 103,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 234,
+ 273,
+ 238,
+ 234
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 129,
+ "image_id": 104,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 354,
+ 219,
+ 153
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 130,
+ "image_id": 105,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 27,
+ 416,
+ 460,
+ 66
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 131,
+ "image_id": 106,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 116,
+ 399,
+ 392,
+ 111
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 132,
+ "image_id": 107,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 276,
+ 471,
+ 226
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 133,
+ "image_id": 108,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 50,
+ 324,
+ 459,
+ 173
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 134,
+ "image_id": 109,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 177,
+ 164,
+ 332,
+ 341
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 135,
+ "image_id": 110,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 208,
+ 471,
+ 285
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 136,
+ "image_id": 111,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 179,
+ 293,
+ 102,
+ 207
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 137,
+ "image_id": 111,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 137,
+ 165,
+ 364
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 138,
+ "image_id": 112,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 157,
+ 267,
+ 140,
+ 138
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 139,
+ "image_id": 113,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 181,
+ 166,
+ 169,
+ 338
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 140,
+ "image_id": 114,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 20,
+ 183,
+ 485,
+ 324
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 141,
+ "image_id": 115,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 107,
+ 505,
+ 402
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 142,
+ "image_id": 116,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 115,
+ 49,
+ 393,
+ 458
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 143,
+ "image_id": 117,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 184,
+ 317,
+ 82,
+ 77
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 144,
+ "image_id": 117,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 376,
+ 506,
+ 133
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 145,
+ "image_id": 118,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 156,
+ 161,
+ 225,
+ 343
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 146,
+ "image_id": 119,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 403,
+ 46,
+ 107,
+ 169
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 147,
+ "image_id": 119,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 203,
+ 506,
+ 300
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 148,
+ "image_id": 120,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 11,
+ 240,
+ 498,
+ 269
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 149,
+ "image_id": 121,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 178,
+ 376,
+ 332,
+ 100
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 150,
+ "image_id": 122,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 141,
+ 502,
+ 338
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 151,
+ "image_id": 123,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 118,
+ 269,
+ 391,
+ 234
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 152,
+ "image_id": 124,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 1,
+ 125,
+ 506,
+ 380
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 153,
+ "image_id": 125,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 147,
+ 273,
+ 362,
+ 236
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 154,
+ "image_id": 126,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 259,
+ 107,
+ 121
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 155,
+ "image_id": 127,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 328,
+ 221,
+ 172
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 156,
+ "image_id": 127,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 211,
+ 221,
+ 100,
+ 124
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 157,
+ "image_id": 127,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 363,
+ 167,
+ 31,
+ 51
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 158,
+ "image_id": 127,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 396,
+ 296,
+ 96,
+ 201
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 159,
+ "image_id": 128,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 314,
+ 408,
+ 187
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 160,
+ "image_id": 129,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 214,
+ 373,
+ 159,
+ 124
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 161,
+ "image_id": 130,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 48,
+ 355,
+ 113,
+ 83
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 162,
+ "image_id": 131,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 208,
+ 352,
+ 157,
+ 149
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 163,
+ "image_id": 132,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 186,
+ 413,
+ 138,
+ 98
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 164,
+ "image_id": 133,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 26,
+ 274,
+ 154,
+ 214
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 165,
+ "image_id": 133,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 20,
+ 211,
+ 86,
+ 64
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 166,
+ "image_id": 134,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 413,
+ 191,
+ 87,
+ 61
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 167,
+ "image_id": 134,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 379,
+ 234,
+ 98,
+ 61
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 168,
+ "image_id": 134,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 372,
+ 296,
+ 139,
+ 201
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 169,
+ "image_id": 135,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 71,
+ 373,
+ 400,
+ 135
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 170,
+ "image_id": 136,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 201,
+ 337,
+ 70,
+ 101
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 171,
+ "image_id": 137,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 181,
+ 135,
+ 108
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 172,
+ "image_id": 137,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 320,
+ 172,
+ 118,
+ 69
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 173,
+ "image_id": 137,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 17,
+ 267,
+ 484,
+ 238
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 174,
+ "image_id": 137,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 195,
+ 164,
+ 22,
+ 65
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 175,
+ "image_id": 138,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 200,
+ 253,
+ 308,
+ 231
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 176,
+ "image_id": 139,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 410,
+ 482,
+ 96
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 177,
+ "image_id": 140,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 1,
+ 430,
+ 385,
+ 76
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 178,
+ "image_id": 141,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 321,
+ 412,
+ 183
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 179,
+ "image_id": 142,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 34,
+ 57,
+ 425,
+ 363
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 180,
+ "image_id": 143,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 238,
+ 124,
+ 232,
+ 239
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 181,
+ "image_id": 143,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 156,
+ 123,
+ 88,
+ 141
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 182,
+ "image_id": 144,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 224,
+ 134,
+ 108,
+ 274
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 183,
+ "image_id": 145,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 93,
+ 129,
+ 285,
+ 276
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 184,
+ "image_id": 146,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 187,
+ 18,
+ 252,
+ 491
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 185,
+ "image_id": 146,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 152,
+ 287,
+ 50,
+ 135
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 186,
+ "image_id": 146,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 423,
+ 282,
+ 45,
+ 153
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 187,
+ "image_id": 147,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 72,
+ 319,
+ 436,
+ 148
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 188,
+ "image_id": 148,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 14,
+ 40,
+ 413,
+ 449
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 189,
+ "image_id": 149,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 253,
+ 227,
+ 85,
+ 199
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 190,
+ "image_id": 149,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 258,
+ 146,
+ 26,
+ 53
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 191,
+ "image_id": 149,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 408,
+ 393,
+ 27,
+ 53
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 192,
+ "image_id": 149,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 446,
+ 399,
+ 33,
+ 63
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 193,
+ "image_id": 149,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 437,
+ 342,
+ 19,
+ 42
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 194,
+ "image_id": 149,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 464,
+ 351,
+ 17,
+ 40
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 195,
+ "image_id": 149,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 488,
+ 377,
+ 20,
+ 43
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 196,
+ "image_id": 149,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 453,
+ 309,
+ 18,
+ 36
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 197,
+ "image_id": 149,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 470,
+ 313,
+ 17,
+ 30
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 198,
+ "image_id": 149,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 487,
+ 328,
+ 16,
+ 19
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 199,
+ "image_id": 150,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 232,
+ 130,
+ 276,
+ 227
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 200,
+ "image_id": 150,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 64,
+ 144,
+ 186,
+ 198
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 201,
+ "image_id": 151,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 218,
+ 181,
+ 216,
+ 250
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 202,
+ "image_id": 152,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 131,
+ 98,
+ 180,
+ 207
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 203,
+ "image_id": 153,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 75,
+ 120,
+ 292,
+ 295
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 204,
+ "image_id": 154,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 36,
+ 52,
+ 422,
+ 385
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 205,
+ "image_id": 155,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 111,
+ 89,
+ 336,
+ 381
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 206,
+ "image_id": 156,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 102,
+ 65,
+ 361,
+ 386
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 207,
+ "image_id": 157,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 76,
+ 295,
+ 377
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 208,
+ "image_id": 158,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 99,
+ 99,
+ 351,
+ 412
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 209,
+ "image_id": 159,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 198,
+ 117,
+ 188,
+ 308
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 210,
+ "image_id": 160,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 85,
+ 15,
+ 342,
+ 487
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 211,
+ "image_id": 161,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 27,
+ 87,
+ 439,
+ 391
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 212,
+ "image_id": 162,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 72,
+ 221,
+ 272,
+ 215
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 213,
+ "image_id": 163,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 26,
+ 123,
+ 181,
+ 350
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 214,
+ "image_id": 164,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 11,
+ 226,
+ 471,
+ 252
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 215,
+ "image_id": 165,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 177,
+ 184,
+ 164,
+ 211
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 216,
+ "image_id": 166,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 240,
+ 104,
+ 165,
+ 198
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 217,
+ "image_id": 166,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 400,
+ 149,
+ 84,
+ 105
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 218,
+ "image_id": 167,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 12,
+ 249,
+ 476,
+ 239
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 219,
+ "image_id": 168,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 234,
+ 303,
+ 192,
+ 171
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 220,
+ "image_id": 169,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 25,
+ 47,
+ 481,
+ 401
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 221,
+ "image_id": 170,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 51,
+ 42,
+ 458,
+ 435
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 222,
+ "image_id": 171,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 24,
+ 94,
+ 187,
+ 190
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 223,
+ "image_id": 171,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 287,
+ 219,
+ 49,
+ 120
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 224,
+ "image_id": 172,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 79,
+ 278,
+ 138
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 225,
+ "image_id": 173,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 81,
+ 129,
+ 278,
+ 281
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 226,
+ "image_id": 174,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 38,
+ 141,
+ 388,
+ 220
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 227,
+ "image_id": 175,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 65,
+ 99,
+ 413,
+ 327
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 228,
+ "image_id": 175,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 263,
+ 373,
+ 61,
+ 86
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 229,
+ "image_id": 175,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 122,
+ 362,
+ 57,
+ 59
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 230,
+ "image_id": 176,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 25,
+ 203,
+ 252,
+ 275
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 231,
+ "image_id": 176,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 268,
+ 187,
+ 134,
+ 307
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 232,
+ "image_id": 177,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 40,
+ 408,
+ 393
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 233,
+ "image_id": 177,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 401,
+ 198,
+ 107,
+ 169
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 234,
+ "image_id": 178,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 39,
+ 80,
+ 336,
+ 297
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 235,
+ "image_id": 179,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 14,
+ 193,
+ 463,
+ 250
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 236,
+ "image_id": 180,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 24,
+ 99,
+ 245,
+ 289
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 237,
+ "image_id": 181,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 78,
+ 177,
+ 144,
+ 132
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 238,
+ "image_id": 182,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 65,
+ 33,
+ 417,
+ 385
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 239,
+ "image_id": 183,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 26,
+ 93,
+ 439,
+ 296
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 240,
+ "image_id": 184,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 66,
+ 140,
+ 274,
+ 252
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 241,
+ "image_id": 185,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 175,
+ 266,
+ 36,
+ 84
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 242,
+ "image_id": 185,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 222,
+ 233,
+ 22,
+ 56
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 243,
+ "image_id": 185,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 246,
+ 224,
+ 18,
+ 42
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 244,
+ "image_id": 185,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 277,
+ 218,
+ 24,
+ 34
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 245,
+ "image_id": 186,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 103,
+ 122,
+ 281,
+ 266
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 246,
+ "image_id": 187,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 171,
+ 240,
+ 241,
+ 132
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 247,
+ "image_id": 188,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 60,
+ 128,
+ 447,
+ 334
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 248,
+ "image_id": 189,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 49,
+ 117,
+ 339,
+ 268
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 249,
+ "image_id": 189,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 212,
+ 47,
+ 116
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 250,
+ "image_id": 190,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 212,
+ 169,
+ 146,
+ 136
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 251,
+ "image_id": 191,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 22,
+ 67,
+ 468,
+ 400
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 252,
+ "image_id": 192,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 248,
+ 241,
+ 254,
+ 208
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 253,
+ "image_id": 192,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 28,
+ 322,
+ 32,
+ 65
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 254,
+ "image_id": 193,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 44,
+ 142,
+ 413,
+ 310
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 255,
+ "image_id": 194,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 245,
+ 312,
+ 88,
+ 84
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 256,
+ "image_id": 194,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 168,
+ 393,
+ 56,
+ 93
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 257,
+ "image_id": 195,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 155,
+ 83,
+ 172,
+ 257
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 258,
+ "image_id": 196,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 247,
+ 333,
+ 202
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 259,
+ "image_id": 197,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 9,
+ 434,
+ 482
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 260,
+ "image_id": 198,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 128,
+ 7,
+ 350,
+ 458
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 261,
+ "image_id": 198,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 107,
+ 194,
+ 38,
+ 100
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 262,
+ "image_id": 199,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 18,
+ 195,
+ 178,
+ 231
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 263,
+ "image_id": 199,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 190,
+ 117,
+ 67,
+ 84
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 264,
+ "image_id": 199,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 440,
+ 153,
+ 14,
+ 24
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 265,
+ "image_id": 199,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 288,
+ 469,
+ 39,
+ 40
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 266,
+ "image_id": 199,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 258,
+ 84,
+ 42,
+ 57
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 267,
+ "image_id": 200,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 83,
+ 269,
+ 135,
+ 139
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 268,
+ "image_id": 201,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 33,
+ 38,
+ 446,
+ 411
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 269,
+ "image_id": 202,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 106,
+ 109,
+ 171,
+ 267
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 270,
+ "image_id": 202,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 333,
+ 238,
+ 176,
+ 121
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 271,
+ "image_id": 203,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 356,
+ 159,
+ 90,
+ 173
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 272,
+ "image_id": 204,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 42,
+ 8,
+ 437,
+ 478
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 273,
+ "image_id": 205,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 26,
+ 75,
+ 476,
+ 390
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 274,
+ "image_id": 206,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 145,
+ 223,
+ 192,
+ 209
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 275,
+ "image_id": 207,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 102,
+ 4,
+ 404,
+ 504
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 276,
+ "image_id": 208,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 174,
+ 53,
+ 332,
+ 380
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 277,
+ "image_id": 209,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 171,
+ 46,
+ 202,
+ 415
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 278,
+ "image_id": 210,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 11,
+ 113,
+ 458,
+ 378
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 279,
+ "image_id": 211,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 126,
+ 393,
+ 173,
+ 110
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 280,
+ "image_id": 212,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 49,
+ 246,
+ 108,
+ 106
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 281,
+ "image_id": 213,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 51,
+ 21,
+ 351,
+ 459
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 282,
+ "image_id": 214,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 180,
+ 234,
+ 97,
+ 62
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 283,
+ "image_id": 214,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 371,
+ 226,
+ 137,
+ 80
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 284,
+ "image_id": 215,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 259,
+ 121,
+ 115,
+ 261
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 285,
+ "image_id": 216,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 114,
+ 307,
+ 364,
+ 200
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 286,
+ "image_id": 217,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 160,
+ 256,
+ 134
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 287,
+ "image_id": 218,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 132,
+ 143,
+ 163,
+ 249
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 288,
+ "image_id": 218,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 29,
+ 171,
+ 111,
+ 235
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 289,
+ "image_id": 219,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 308,
+ 37,
+ 172,
+ 262
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 290,
+ "image_id": 220,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 170,
+ 198,
+ 264,
+ 207
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 291,
+ "image_id": 221,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 30,
+ 58,
+ 432,
+ 419
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 292,
+ "image_id": 222,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 28,
+ 106,
+ 444,
+ 335
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 293,
+ "image_id": 223,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 103,
+ 8,
+ 404,
+ 468
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 294,
+ "image_id": 224,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 99,
+ 109,
+ 376,
+ 310
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 295,
+ "image_id": 225,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 248,
+ 150,
+ 174,
+ 185
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 296,
+ "image_id": 225,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 4,
+ 268,
+ 441
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 297,
+ "image_id": 225,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 417,
+ 238,
+ 40,
+ 71
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 298,
+ "image_id": 225,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 448,
+ 254,
+ 25,
+ 51
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 299,
+ "image_id": 226,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 360,
+ 340,
+ 101,
+ 120
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 300,
+ "image_id": 227,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 228,
+ 299,
+ 204,
+ 154
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 301,
+ "image_id": 227,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 307,
+ 192,
+ 135
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 302,
+ "image_id": 228,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 386,
+ 227,
+ 122,
+ 109
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 303,
+ "image_id": 229,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 118,
+ 128,
+ 226,
+ 284
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 304,
+ "image_id": 230,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 138,
+ 40,
+ 321,
+ 400
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 305,
+ "image_id": 231,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 166,
+ 135,
+ 203,
+ 218
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 306,
+ "image_id": 231,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 382,
+ 131,
+ 126,
+ 218
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 307,
+ "image_id": 232,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 87,
+ 128,
+ 288,
+ 271
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 308,
+ "image_id": 233,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 164,
+ 64,
+ 236,
+ 375
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 309,
+ "image_id": 234,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 72,
+ 57,
+ 329,
+ 384
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 310,
+ "image_id": 235,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 101,
+ 407,
+ 400
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 311,
+ "image_id": 236,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 130,
+ 216,
+ 343
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 312,
+ "image_id": 237,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 363,
+ 345,
+ 101,
+ 107
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 313,
+ "image_id": 237,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 455,
+ 322,
+ 53,
+ 70
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 314,
+ "image_id": 237,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 320,
+ 301,
+ 63,
+ 38
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 315,
+ "image_id": 238,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 342,
+ 245,
+ 98,
+ 157
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 316,
+ "image_id": 238,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 291,
+ 326,
+ 23,
+ 41
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 317,
+ "image_id": 238,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 151,
+ 294,
+ 41,
+ 49
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 318,
+ "image_id": 239,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 301,
+ 143,
+ 53,
+ 39
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 319,
+ "image_id": 240,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 144,
+ 244,
+ 224
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 320,
+ "image_id": 241,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 87,
+ 91,
+ 375,
+ 363
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 321,
+ "image_id": 242,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 281,
+ 256,
+ 163,
+ 182
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 322,
+ "image_id": 242,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 310,
+ 57,
+ 97
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 323,
+ "image_id": 243,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 45,
+ 214,
+ 421,
+ 152
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 324,
+ "image_id": 244,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 81,
+ 54,
+ 420,
+ 412
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 325,
+ "image_id": 245,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 172,
+ 353,
+ 242
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 326,
+ "image_id": 246,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 156,
+ 39,
+ 346,
+ 316
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 327,
+ "image_id": 247,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 239,
+ 219,
+ 216,
+ 187
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 328,
+ "image_id": 248,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 132,
+ 51,
+ 279,
+ 319
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 329,
+ "image_id": 249,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 70,
+ 69,
+ 429,
+ 381
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 330,
+ "image_id": 250,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 46,
+ 281,
+ 415
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 331,
+ "image_id": 251,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 133,
+ 137,
+ 237,
+ 259
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 332,
+ "image_id": 252,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 156,
+ 148,
+ 270,
+ 289
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 333,
+ "image_id": 253,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 63,
+ 118,
+ 447,
+ 368
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 334,
+ "image_id": 254,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 244,
+ 120,
+ 262,
+ 259
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 335,
+ "image_id": 255,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 68,
+ 61,
+ 325,
+ 344
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 336,
+ "image_id": 256,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 102,
+ 51,
+ 280,
+ 291
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 337,
+ "image_id": 257,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 16,
+ 42,
+ 488,
+ 407
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 338,
+ "image_id": 258,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 192,
+ 174,
+ 212,
+ 232
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 339,
+ "image_id": 259,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 201,
+ 146,
+ 233,
+ 152
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 340,
+ "image_id": 260,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 124,
+ 117,
+ 283,
+ 277
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 341,
+ "image_id": 261,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 68,
+ 128,
+ 392,
+ 226
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 342,
+ "image_id": 262,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 14,
+ 123,
+ 493,
+ 275
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 343,
+ "image_id": 263,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 93,
+ 233,
+ 252,
+ 121
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 344,
+ "image_id": 264,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 45,
+ 114,
+ 270,
+ 320
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 345,
+ "image_id": 265,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 167,
+ 228,
+ 67,
+ 77
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 346,
+ "image_id": 266,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 53,
+ 494,
+ 402
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 347,
+ "image_id": 267,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 188,
+ 150,
+ 170,
+ 198
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 348,
+ "image_id": 268,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 12,
+ 52,
+ 492,
+ 375
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 349,
+ "image_id": 269,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 10,
+ 501,
+ 407
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 350,
+ "image_id": 270,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 203,
+ 94,
+ 306,
+ 324
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 351,
+ "image_id": 271,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 270,
+ 151,
+ 72,
+ 81
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 352,
+ "image_id": 272,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 161,
+ 276,
+ 221
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 353,
+ "image_id": 273,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 76,
+ 109,
+ 423,
+ 325
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 354,
+ "image_id": 274,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 148,
+ 429,
+ 309
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 355,
+ "image_id": 275,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 165,
+ 152,
+ 150,
+ 154
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 356,
+ "image_id": 276,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 96,
+ 339,
+ 247
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 357,
+ "image_id": 277,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 2,
+ 440,
+ 399
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 358,
+ "image_id": 278,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 11,
+ 5,
+ 425,
+ 440
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 359,
+ "image_id": 279,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 8,
+ 420,
+ 440
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 360,
+ "image_id": 280,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 86,
+ 32,
+ 378,
+ 396
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 361,
+ "image_id": 281,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 1,
+ 19,
+ 213,
+ 330
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 362,
+ "image_id": 281,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 211,
+ 19,
+ 295,
+ 358
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 363,
+ "image_id": 282,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 94,
+ 100,
+ 345,
+ 336
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 364,
+ "image_id": 283,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 36,
+ 87,
+ 448,
+ 385
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 365,
+ "image_id": 284,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 91,
+ 53,
+ 380,
+ 434
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 366,
+ "image_id": 285,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 115,
+ 107,
+ 253,
+ 291
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 367,
+ "image_id": 286,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 38,
+ 486,
+ 420
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 368,
+ "image_id": 287,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 391,
+ 158,
+ 95,
+ 182
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 369,
+ "image_id": 287,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 74,
+ 193,
+ 169,
+ 151
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 370,
+ "image_id": 287,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 345,
+ 231,
+ 37,
+ 73
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 371,
+ "image_id": 287,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 158,
+ 110,
+ 166
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 372,
+ "image_id": 288,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 19,
+ 35,
+ 434,
+ 469
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 373,
+ "image_id": 289,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 383,
+ 105,
+ 98
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 374,
+ "image_id": 289,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 90,
+ 373,
+ 228,
+ 125
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 375,
+ "image_id": 289,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 327,
+ 392,
+ 109,
+ 82
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 376,
+ "image_id": 289,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 411,
+ 397,
+ 97,
+ 68
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 377,
+ "image_id": 290,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 75,
+ 406,
+ 304
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 378,
+ "image_id": 291,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 85,
+ 43,
+ 422,
+ 398
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 379,
+ "image_id": 292,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 33,
+ 51,
+ 187,
+ 326
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 380,
+ "image_id": 293,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 54,
+ 100,
+ 320,
+ 261
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 381,
+ "image_id": 293,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 401,
+ 131,
+ 108,
+ 192
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 382,
+ "image_id": 294,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 51,
+ 27,
+ 397,
+ 451
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 383,
+ "image_id": 295,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 97,
+ 172,
+ 202,
+ 305
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 384,
+ "image_id": 295,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 285,
+ 224,
+ 136,
+ 141
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 385,
+ "image_id": 296,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 212,
+ 212,
+ 203,
+ 169
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 386,
+ "image_id": 297,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 96,
+ 149,
+ 406,
+ 327
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 387,
+ "image_id": 298,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 216,
+ 355,
+ 291,
+ 129
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 388,
+ "image_id": 299,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 238,
+ 10,
+ 266,
+ 477
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 389,
+ "image_id": 299,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 15,
+ 200,
+ 484
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 390,
+ "image_id": 300,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 322,
+ 265,
+ 164,
+ 105
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 391,
+ "image_id": 301,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 85,
+ 71,
+ 351,
+ 301
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 392,
+ "image_id": 302,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 174,
+ 139,
+ 169,
+ 253
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 393,
+ "image_id": 303,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 17,
+ 60,
+ 486,
+ 352
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 394,
+ "image_id": 304,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 71,
+ 117,
+ 263,
+ 172
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 395,
+ "image_id": 305,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 326,
+ 126,
+ 133,
+ 195
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 396,
+ "image_id": 306,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 108,
+ 76,
+ 339,
+ 359
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 397,
+ "image_id": 307,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 7,
+ 80,
+ 214,
+ 314
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 398,
+ "image_id": 308,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 8,
+ 150,
+ 301,
+ 307
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 399,
+ "image_id": 309,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 316,
+ 132,
+ 139,
+ 246
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 400,
+ "image_id": 310,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 80,
+ 159,
+ 405,
+ 213
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 401,
+ "image_id": 311,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 26,
+ 91,
+ 474,
+ 337
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 402,
+ "image_id": 312,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 50,
+ 492,
+ 386
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 403,
+ "image_id": 313,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 116,
+ 153,
+ 376,
+ 264
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 404,
+ "image_id": 314,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 251,
+ 89,
+ 254,
+ 249
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 405,
+ "image_id": 314,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 36,
+ 46,
+ 196,
+ 308
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 406,
+ "image_id": 315,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 168,
+ 478,
+ 242
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 407,
+ "image_id": 316,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 119,
+ 326,
+ 89,
+ 74
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 408,
+ "image_id": 316,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 351,
+ 96,
+ 81
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 409,
+ "image_id": 316,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 455,
+ 266,
+ 37,
+ 38
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 410,
+ "image_id": 317,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 333,
+ 123,
+ 98
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 411,
+ "image_id": 317,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 250,
+ 263,
+ 110,
+ 161
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 412,
+ "image_id": 317,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 410,
+ 331,
+ 30,
+ 35
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 413,
+ "image_id": 318,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 179,
+ 242,
+ 331,
+ 264
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 414,
+ "image_id": 319,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 135,
+ 234,
+ 79,
+ 90
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 415,
+ "image_id": 320,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 113,
+ 403,
+ 50,
+ 63
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 416,
+ "image_id": 321,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 34,
+ 173,
+ 362,
+ 94
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 417,
+ "image_id": 322,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 204,
+ 7,
+ 303,
+ 444
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 418,
+ "image_id": 322,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 80,
+ 27,
+ 131,
+ 232
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 419,
+ "image_id": 323,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 155,
+ 175,
+ 205,
+ 263
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 420,
+ "image_id": 323,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 63,
+ 239,
+ 62,
+ 79
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 421,
+ "image_id": 324,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 7,
+ 133,
+ 332,
+ 281
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 422,
+ "image_id": 325,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 234,
+ 258,
+ 159,
+ 103
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 423,
+ "image_id": 326,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 72,
+ 269,
+ 65,
+ 106
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 424,
+ "image_id": 327,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 68,
+ 208,
+ 299,
+ 280
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 425,
+ "image_id": 328,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 305,
+ 233,
+ 94,
+ 88
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 426,
+ "image_id": 329,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 74,
+ 208,
+ 424,
+ 281
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 427,
+ "image_id": 330,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 75,
+ 83,
+ 216,
+ 188
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 428,
+ "image_id": 331,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 30,
+ 47,
+ 322,
+ 407
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 429,
+ "image_id": 332,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 102,
+ 107,
+ 326,
+ 197
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 430,
+ "image_id": 333,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 297,
+ 247,
+ 92,
+ 132
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 431,
+ "image_id": 334,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 151,
+ 382,
+ 318
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 432,
+ "image_id": 335,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 7,
+ 59,
+ 490,
+ 424
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 433,
+ "image_id": 336,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 18,
+ 105,
+ 489,
+ 292
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 434,
+ "image_id": 337,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 295,
+ 364,
+ 64,
+ 38
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 435,
+ "image_id": 338,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 150,
+ 215,
+ 220,
+ 243
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 436,
+ "image_id": 339,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 29,
+ 485,
+ 397
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 437,
+ "image_id": 340,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 100,
+ 263,
+ 92,
+ 228
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 438,
+ "image_id": 340,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 200,
+ 304,
+ 96,
+ 166
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 439,
+ "image_id": 341,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 15,
+ 27,
+ 489,
+ 449
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 440,
+ "image_id": 342,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 14,
+ 49,
+ 421,
+ 402
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 441,
+ "image_id": 343,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 102,
+ 22,
+ 402,
+ 485
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 442,
+ "image_id": 344,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 55,
+ 66,
+ 367,
+ 301
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 443,
+ "image_id": 345,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 229,
+ 136,
+ 224,
+ 182
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 444,
+ "image_id": 346,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 202,
+ 119,
+ 285,
+ 272
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 445,
+ "image_id": 347,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 75,
+ 108,
+ 372,
+ 214
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 446,
+ "image_id": 348,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 92,
+ 312,
+ 187,
+ 162
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 447,
+ "image_id": 348,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 330,
+ 370,
+ 61,
+ 95
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 448,
+ "image_id": 348,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 1,
+ 259,
+ 85,
+ 220
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 449,
+ "image_id": 349,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 305,
+ 98,
+ 126,
+ 158
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 450,
+ "image_id": 349,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 243,
+ 132,
+ 68,
+ 73
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 451,
+ "image_id": 350,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 200,
+ 192,
+ 308,
+ 278
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 452,
+ "image_id": 351,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 49,
+ 74,
+ 328,
+ 343
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 453,
+ "image_id": 352,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 59,
+ 108,
+ 256,
+ 374
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 454,
+ "image_id": 353,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 105,
+ 160,
+ 282,
+ 246
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 455,
+ "image_id": 353,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 40,
+ 57,
+ 349
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 456,
+ "image_id": 354,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 114,
+ 78,
+ 168,
+ 372
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 457,
+ "image_id": 355,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 22,
+ 158,
+ 475,
+ 232
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 458,
+ "image_id": 356,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 93,
+ 97,
+ 196,
+ 284
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 459,
+ "image_id": 357,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 52,
+ 112,
+ 194,
+ 286
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 460,
+ "image_id": 358,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 111,
+ 239,
+ 168,
+ 116
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 461,
+ "image_id": 359,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 11,
+ 324,
+ 196,
+ 159
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 462,
+ "image_id": 359,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 363,
+ 284,
+ 17,
+ 18
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 463,
+ "image_id": 360,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 216,
+ 133,
+ 33,
+ 145
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 464,
+ "image_id": 360,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 380,
+ 56,
+ 64,
+ 271
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 465,
+ "image_id": 361,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 175,
+ 337,
+ 201
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 466,
+ "image_id": 362,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 97,
+ 73,
+ 330,
+ 329
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 467,
+ "image_id": 363,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 98,
+ 113,
+ 377,
+ 346
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 468,
+ "image_id": 364,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 83,
+ 46,
+ 295,
+ 391
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 469,
+ "image_id": 365,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 119,
+ 174,
+ 332,
+ 295
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 470,
+ "image_id": 366,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 20,
+ 105,
+ 487,
+ 208
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 471,
+ "image_id": 367,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 52,
+ 206,
+ 339,
+ 223
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 472,
+ "image_id": 368,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 50,
+ 117,
+ 429,
+ 389
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 473,
+ "image_id": 369,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 40,
+ 113,
+ 287,
+ 363
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 474,
+ "image_id": 370,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 52,
+ 128,
+ 404,
+ 334
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 475,
+ "image_id": 371,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 90,
+ 106,
+ 413,
+ 359
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 476,
+ "image_id": 372,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 165,
+ 120,
+ 193,
+ 301
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 477,
+ "image_id": 373,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 82,
+ 283,
+ 150,
+ 154
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 478,
+ "image_id": 374,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 233,
+ 286,
+ 100,
+ 148
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 479,
+ "image_id": 375,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 96,
+ 203,
+ 258,
+ 213
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 480,
+ "image_id": 376,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 215,
+ 203,
+ 293,
+ 243
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 481,
+ "image_id": 377,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 29,
+ 477,
+ 422
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 482,
+ "image_id": 378,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 248,
+ 173,
+ 82,
+ 128
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 483,
+ "image_id": 379,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 7,
+ 171,
+ 495,
+ 226
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 484,
+ "image_id": 380,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 338,
+ 277,
+ 95,
+ 107
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 485,
+ "image_id": 381,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 46,
+ 243,
+ 197,
+ 138
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 486,
+ "image_id": 382,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 15,
+ 103,
+ 467,
+ 292
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 487,
+ "image_id": 383,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 211,
+ 165,
+ 220,
+ 129
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 488,
+ "image_id": 384,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 246,
+ 188,
+ 230,
+ 228
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 489,
+ "image_id": 385,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 128,
+ 33,
+ 374,
+ 420
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 490,
+ "image_id": 386,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 163,
+ 178,
+ 110,
+ 142
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 491,
+ "image_id": 387,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 138,
+ 274,
+ 281,
+ 91
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 492,
+ "image_id": 387,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 14,
+ 328,
+ 397,
+ 166
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 493,
+ "image_id": 388,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 1,
+ 239,
+ 119,
+ 123
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 494,
+ "image_id": 388,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 108,
+ 254,
+ 63,
+ 84
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 495,
+ "image_id": 389,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 320,
+ 63,
+ 188,
+ 226
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 496,
+ "image_id": 389,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 236,
+ 242,
+ 64,
+ 139
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 497,
+ "image_id": 390,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 29,
+ 128,
+ 227,
+ 318
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 498,
+ "image_id": 391,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 30,
+ 126,
+ 458,
+ 273
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 499,
+ "image_id": 392,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 176,
+ 186,
+ 59,
+ 111
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 500,
+ "image_id": 392,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 266,
+ 176,
+ 61,
+ 108
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 501,
+ "image_id": 392,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 330,
+ 211,
+ 34,
+ 42
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 502,
+ "image_id": 393,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 60,
+ 249,
+ 316,
+ 215
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 503,
+ "image_id": 394,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 7,
+ 362,
+ 206,
+ 120
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 504,
+ "image_id": 395,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 140,
+ 359,
+ 107,
+ 58
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 505,
+ "image_id": 396,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 233,
+ 88,
+ 213,
+ 373
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 506,
+ "image_id": 397,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 21,
+ 36,
+ 459,
+ 353
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 507,
+ "image_id": 398,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 199,
+ 27,
+ 188,
+ 276
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 508,
+ "image_id": 399,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 126,
+ 494,
+ 269
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 509,
+ "image_id": 400,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 286,
+ 174,
+ 137,
+ 163
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 510,
+ "image_id": 401,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 151,
+ 41,
+ 245,
+ 351
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 511,
+ "image_id": 402,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 225,
+ 84,
+ 91,
+ 141
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 512,
+ "image_id": 403,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 79,
+ 83,
+ 137,
+ 378
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 513,
+ "image_id": 403,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 267,
+ 83,
+ 184,
+ 347
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 514,
+ "image_id": 404,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 126,
+ 116,
+ 167,
+ 343
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 515,
+ "image_id": 405,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 13,
+ 18,
+ 399,
+ 420
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 516,
+ "image_id": 406,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 9,
+ 456,
+ 434
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 517,
+ "image_id": 407,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 208,
+ 285,
+ 109,
+ 176
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 518,
+ "image_id": 408,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 247,
+ 108,
+ 133,
+ 283
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 519,
+ "image_id": 408,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 368,
+ 257,
+ 37,
+ 59
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 520,
+ "image_id": 409,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 89,
+ 170,
+ 76,
+ 155
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 521,
+ "image_id": 409,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 151,
+ 115,
+ 172,
+ 225
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 522,
+ "image_id": 410,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 187,
+ 123,
+ 259,
+ 272
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 523,
+ "image_id": 411,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 37,
+ 34,
+ 407,
+ 404
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 524,
+ "image_id": 412,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 211,
+ 327,
+ 111,
+ 82
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 525,
+ "image_id": 412,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 66,
+ 328,
+ 79,
+ 79
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 526,
+ "image_id": 412,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 266,
+ 170,
+ 16,
+ 50
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 527,
+ "image_id": 412,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 449,
+ 277,
+ 42,
+ 82
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 528,
+ "image_id": 413,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 168,
+ 5,
+ 242,
+ 422
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 529,
+ "image_id": 414,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 44,
+ 492,
+ 419
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 530,
+ "image_id": 415,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 26,
+ 7,
+ 425,
+ 416
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 531,
+ "image_id": 416,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 172,
+ 7,
+ 318,
+ 466
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 532,
+ "image_id": 417,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 210,
+ 170,
+ 152,
+ 176
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 533,
+ "image_id": 417,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 140,
+ 227,
+ 57,
+ 83
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 534,
+ "image_id": 418,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 54,
+ 131,
+ 301,
+ 231
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 535,
+ "image_id": 419,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 248,
+ 269,
+ 29,
+ 57
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 536,
+ "image_id": 419,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 280,
+ 271,
+ 11,
+ 32
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 537,
+ "image_id": 420,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 156,
+ 413,
+ 23,
+ 41
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 538,
+ "image_id": 420,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 286,
+ 403,
+ 16,
+ 24
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 539,
+ "image_id": 420,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 369,
+ 402,
+ 12,
+ 28
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 540,
+ "image_id": 420,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 64,
+ 399,
+ 28,
+ 26
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 541,
+ "image_id": 421,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 29,
+ 220,
+ 18,
+ 23
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 542,
+ "image_id": 421,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 52,
+ 220,
+ 24,
+ 25
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 543,
+ "image_id": 421,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 113,
+ 220,
+ 16,
+ 25
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 544,
+ "image_id": 422,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 30,
+ 384,
+ 95,
+ 117
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 545,
+ "image_id": 423,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 98,
+ 450,
+ 16,
+ 23
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 546,
+ "image_id": 423,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 427,
+ 464,
+ 19,
+ 29
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 547,
+ "image_id": 424,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 420,
+ 386,
+ 90,
+ 115
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 548,
+ "image_id": 425,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 331,
+ 33,
+ 50
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 549,
+ "image_id": 425,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 406,
+ 338,
+ 16,
+ 27
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 550,
+ "image_id": 426,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 420,
+ 337,
+ 77,
+ 89
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 551,
+ "image_id": 427,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 1,
+ 345,
+ 92,
+ 63
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 552,
+ "image_id": 428,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 210,
+ 280,
+ 21,
+ 21
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 553,
+ "image_id": 429,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 195,
+ 392,
+ 26,
+ 37
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 554,
+ "image_id": 430,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 444,
+ 195,
+ 30,
+ 52
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 555,
+ "image_id": 431,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 250,
+ 361,
+ 15,
+ 25
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 556,
+ "image_id": 431,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 273,
+ 365,
+ 10,
+ 17
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 557,
+ "image_id": 432,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 149,
+ 303,
+ 276,
+ 194
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 558,
+ "image_id": 433,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 14,
+ 501,
+ 493
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 559,
+ "image_id": 434,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 253,
+ 125,
+ 159,
+ 72
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 560,
+ "image_id": 435,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 8,
+ 497,
+ 456
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 561,
+ "image_id": 436,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 176,
+ 264,
+ 254,
+ 206
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 562,
+ "image_id": 437,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 277,
+ 175,
+ 122,
+ 90
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 563,
+ "image_id": 438,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 235,
+ 59,
+ 125
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 564,
+ "image_id": 439,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 77,
+ 190,
+ 194,
+ 238
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 565,
+ "image_id": 440,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 87,
+ 281,
+ 58,
+ 113
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 566,
+ "image_id": 441,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 12,
+ 238,
+ 274,
+ 274
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 567,
+ "image_id": 442,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 136,
+ 196,
+ 260,
+ 308
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 568,
+ "image_id": 443,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 236,
+ 281,
+ 24,
+ 69
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 569,
+ "image_id": 443,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 311,
+ 332,
+ 141,
+ 33
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 570,
+ "image_id": 444,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 126,
+ 186,
+ 211,
+ 260
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 571,
+ "image_id": 445,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 61,
+ 351,
+ 311,
+ 157
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 572,
+ "image_id": 446,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 293,
+ 347,
+ 118,
+ 106
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 573,
+ "image_id": 447,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 123,
+ 187,
+ 205,
+ 244
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 574,
+ "image_id": 448,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 57,
+ 288,
+ 334,
+ 219
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 575,
+ "image_id": 449,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 227,
+ 363,
+ 126
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 576,
+ "image_id": 450,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 166,
+ 290,
+ 122,
+ 190
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 577,
+ "image_id": 451,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 219,
+ 292,
+ 249
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 578,
+ "image_id": 452,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 258,
+ 353,
+ 161,
+ 144
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 579,
+ "image_id": 453,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 287,
+ 138,
+ 182
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 580,
+ "image_id": 454,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 206,
+ 225,
+ 205,
+ 266
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 581,
+ "image_id": 455,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 174,
+ 296,
+ 192,
+ 199
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 582,
+ "image_id": 456,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 164,
+ 157,
+ 343,
+ 338
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 583,
+ "image_id": 457,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 233,
+ 248,
+ 78,
+ 113
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 584,
+ "image_id": 458,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 240,
+ 347,
+ 114,
+ 141
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 585,
+ "image_id": 459,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 138,
+ 385,
+ 188,
+ 123
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 586,
+ "image_id": 460,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 417,
+ 249,
+ 92,
+ 235
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 587,
+ "image_id": 461,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 249,
+ 301,
+ 118,
+ 61
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 588,
+ "image_id": 461,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 447,
+ 293,
+ 64,
+ 66
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 589,
+ "image_id": 461,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 424,
+ 358,
+ 84,
+ 113
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 590,
+ "image_id": 462,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 84,
+ 276,
+ 84,
+ 44
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 591,
+ "image_id": 462,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 321,
+ 93,
+ 67
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 592,
+ "image_id": 463,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 70,
+ 275,
+ 176,
+ 228
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 593,
+ "image_id": 463,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 293,
+ 354,
+ 83,
+ 145
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 594,
+ "image_id": 464,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 135,
+ 111,
+ 361,
+ 392
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 595,
+ "image_id": 464,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 1,
+ 118,
+ 184,
+ 385
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 596,
+ "image_id": 465,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 208,
+ 223,
+ 284
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 597,
+ "image_id": 466,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 43,
+ 342,
+ 198,
+ 161
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 598,
+ "image_id": 467,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 120,
+ 368,
+ 88,
+ 136
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 599,
+ "image_id": 468,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 74,
+ 219,
+ 431,
+ 135
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 600,
+ "image_id": 469,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 145,
+ 270,
+ 301,
+ 72
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 601,
+ "image_id": 470,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 126,
+ 291,
+ 152,
+ 70
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 602,
+ "image_id": 471,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 69,
+ 265,
+ 98,
+ 138
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 603,
+ "image_id": 472,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 141,
+ 286,
+ 291,
+ 112
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 604,
+ "image_id": 473,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 78,
+ 265,
+ 230,
+ 199
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 605,
+ "image_id": 474,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 162,
+ 332,
+ 244,
+ 160
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 606,
+ "image_id": 475,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 135,
+ 197,
+ 371,
+ 278
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 607,
+ "image_id": 476,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 288,
+ 202,
+ 67,
+ 113
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 608,
+ "image_id": 476,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 187,
+ 217,
+ 52,
+ 67
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 609,
+ "image_id": 477,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 133,
+ 144,
+ 305,
+ 197
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 610,
+ "image_id": 478,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 276,
+ 320,
+ 215
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 611,
+ "image_id": 479,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 75,
+ 216,
+ 370,
+ 237
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 612,
+ "image_id": 480,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 264,
+ 351,
+ 116,
+ 155
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 613,
+ "image_id": 481,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 224,
+ 288,
+ 126,
+ 86
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 614,
+ "image_id": 482,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 292,
+ 67,
+ 181
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 615,
+ "image_id": 482,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 359,
+ 256,
+ 105,
+ 181
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 616,
+ "image_id": 483,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 111,
+ 197,
+ 137,
+ 214
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 617,
+ "image_id": 484,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 118,
+ 299,
+ 140,
+ 202
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 618,
+ "image_id": 485,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 144,
+ 155,
+ 155,
+ 103
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 619,
+ "image_id": 485,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 311,
+ 125,
+ 92,
+ 93
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 620,
+ "image_id": 486,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 248,
+ 233,
+ 119,
+ 276
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 621,
+ "image_id": 487,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 172,
+ 199,
+ 72,
+ 75
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 622,
+ "image_id": 488,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 246,
+ 231,
+ 36,
+ 27
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 623,
+ "image_id": 488,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 380,
+ 280,
+ 91,
+ 150
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 624,
+ "image_id": 489,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 186,
+ 249,
+ 160,
+ 252
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 625,
+ "image_id": 490,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 211,
+ 165,
+ 295,
+ 337
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 626,
+ "image_id": 491,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 330,
+ 230,
+ 102,
+ 155
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 627,
+ "image_id": 492,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 244,
+ 126,
+ 262,
+ 229
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 628,
+ "image_id": 493,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 262,
+ 202,
+ 244,
+ 285
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 629,
+ "image_id": 494,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 289,
+ 284,
+ 62,
+ 89
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 630,
+ "image_id": 494,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 142,
+ 291,
+ 100,
+ 170
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 631,
+ "image_id": 495,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 34,
+ 305,
+ 276,
+ 200
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 632,
+ "image_id": 496,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 327,
+ 269,
+ 173,
+ 222
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 633,
+ "image_id": 497,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 34,
+ 330,
+ 97,
+ 163
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 634,
+ "image_id": 498,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 182,
+ 202,
+ 316
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 635,
+ "image_id": 499,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 146,
+ 289,
+ 79,
+ 36
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 636,
+ "image_id": 499,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 387,
+ 292,
+ 124,
+ 55
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 637,
+ "image_id": 499,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 267,
+ 294,
+ 84,
+ 195
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 638,
+ "image_id": 499,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 382,
+ 371,
+ 124,
+ 137
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 639,
+ "image_id": 500,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 122,
+ 280,
+ 103,
+ 46
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 640,
+ "image_id": 501,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 172,
+ 502,
+ 182
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 641,
+ "image_id": 502,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 251,
+ 119,
+ 177
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 642,
+ "image_id": 502,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 365,
+ 249,
+ 126,
+ 163
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 643,
+ "image_id": 503,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 173,
+ 15,
+ 133,
+ 114
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 644,
+ "image_id": 503,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 120,
+ 39,
+ 62,
+ 60
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 645,
+ "image_id": 504,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 144,
+ 156,
+ 362,
+ 280
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 646,
+ "image_id": 505,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 25,
+ 271,
+ 460,
+ 94
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 647,
+ "image_id": 506,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 320,
+ 330,
+ 79,
+ 115
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 648,
+ "image_id": 507,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 66,
+ 80,
+ 303,
+ 208
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 649,
+ "image_id": 508,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 11,
+ 23,
+ 338,
+ 411
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 650,
+ "image_id": 509,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 74,
+ 213,
+ 378
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 651,
+ "image_id": 510,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 252,
+ 45,
+ 250,
+ 429
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 652,
+ "image_id": 511,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 177,
+ 194,
+ 158,
+ 278
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 653,
+ "image_id": 512,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 307,
+ 229,
+ 200,
+ 270
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 654,
+ "image_id": 513,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 47,
+ 267,
+ 217,
+ 170
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 655,
+ "image_id": 514,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 326,
+ 140,
+ 182,
+ 335
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 656,
+ "image_id": 515,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 262,
+ 220,
+ 89,
+ 102
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 657,
+ "image_id": 516,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 307,
+ 74,
+ 147
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 658,
+ "image_id": 517,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 293,
+ 156,
+ 215
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 659,
+ "image_id": 518,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 138,
+ 285,
+ 266,
+ 210
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 660,
+ "image_id": 518,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 168,
+ 230,
+ 257,
+ 82
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 661,
+ "image_id": 519,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 225,
+ 146,
+ 93,
+ 194
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 662,
+ "image_id": 520,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 44,
+ 314,
+ 166,
+ 176
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 663,
+ "image_id": 521,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 210,
+ 279,
+ 64,
+ 137
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 664,
+ "image_id": 522,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 185,
+ 245,
+ 192,
+ 241
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 665,
+ "image_id": 523,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 15,
+ 260,
+ 182,
+ 201
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 666,
+ "image_id": 524,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 60,
+ 245,
+ 336,
+ 220
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 667,
+ "image_id": 525,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 81,
+ 272,
+ 400,
+ 185
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 668,
+ "image_id": 526,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 216,
+ 131,
+ 174,
+ 371
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 669,
+ "image_id": 527,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 106,
+ 150,
+ 288,
+ 356
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 670,
+ "image_id": 527,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 259,
+ 122,
+ 250,
+ 315
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 671,
+ "image_id": 528,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 98,
+ 36,
+ 301,
+ 469
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 672,
+ "image_id": 529,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 129,
+ 115,
+ 116,
+ 315
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 673,
+ "image_id": 530,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 72,
+ 120,
+ 345,
+ 304
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 674,
+ "image_id": 531,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 174,
+ 180,
+ 166,
+ 240
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 675,
+ "image_id": 531,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 292,
+ 86,
+ 84,
+ 121
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 676,
+ "image_id": 532,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 126,
+ 15,
+ 237,
+ 490
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 677,
+ "image_id": 533,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 72,
+ 207,
+ 383,
+ 291
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 678,
+ "image_id": 534,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 65,
+ 176,
+ 237,
+ 322
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 679,
+ "image_id": 535,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 147,
+ 159,
+ 129,
+ 245
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 680,
+ "image_id": 536,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 257,
+ 236,
+ 106,
+ 103
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 681,
+ "image_id": 536,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 164,
+ 232,
+ 81,
+ 94
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 682,
+ "image_id": 537,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 178,
+ 342,
+ 199,
+ 163
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 683,
+ "image_id": 538,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 213,
+ 268,
+ 203,
+ 144
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 684,
+ "image_id": 538,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 265,
+ 231,
+ 102
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 685,
+ "image_id": 539,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 144,
+ 356,
+ 249,
+ 66
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 686,
+ "image_id": 540,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 290,
+ 116,
+ 119,
+ 99
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 687,
+ "image_id": 541,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 193,
+ 7,
+ 253,
+ 455
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 688,
+ "image_id": 542,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 133,
+ 21,
+ 374,
+ 135
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 689,
+ "image_id": 543,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 355,
+ 270,
+ 120,
+ 182
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 690,
+ "image_id": 544,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 58,
+ 55,
+ 66,
+ 158
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 691,
+ "image_id": 544,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 15,
+ 25,
+ 65,
+ 72
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 692,
+ "image_id": 544,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 186,
+ 142,
+ 125,
+ 305
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 693,
+ "image_id": 545,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 330,
+ 311,
+ 75,
+ 190
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 694,
+ "image_id": 546,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 395,
+ 307,
+ 116,
+ 197
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 695,
+ "image_id": 547,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 278,
+ 227,
+ 153,
+ 263
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 696,
+ "image_id": 548,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 310,
+ 107,
+ 127
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 697,
+ "image_id": 549,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 228,
+ 334,
+ 220,
+ 137
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 698,
+ "image_id": 549,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 253,
+ 221,
+ 112,
+ 66
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 699,
+ "image_id": 550,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 165,
+ 334,
+ 195,
+ 148
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 700,
+ "image_id": 551,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 116,
+ 227,
+ 47,
+ 103
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 701,
+ "image_id": 552,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 7,
+ 161,
+ 210,
+ 110
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 702,
+ "image_id": 553,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 210,
+ 287,
+ 109,
+ 186
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 703,
+ "image_id": 554,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 69,
+ 351,
+ 352,
+ 152
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 704,
+ "image_id": 554,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 12,
+ 300,
+ 111,
+ 155
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 705,
+ "image_id": 555,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 88,
+ 267,
+ 124,
+ 132
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 706,
+ "image_id": 556,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 184,
+ 282,
+ 111,
+ 177
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 707,
+ "image_id": 557,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 149,
+ 325,
+ 152,
+ 146
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 708,
+ "image_id": 557,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 17,
+ 309,
+ 117,
+ 108
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 709,
+ "image_id": 558,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 303,
+ 120,
+ 199
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 710,
+ "image_id": 558,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 389,
+ 335,
+ 122,
+ 174
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 711,
+ "image_id": 559,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 326,
+ 217,
+ 121,
+ 181
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 712,
+ "image_id": 560,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 180,
+ 256,
+ 263,
+ 205
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 713,
+ "image_id": 561,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 192,
+ 350,
+ 174,
+ 151
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 714,
+ "image_id": 561,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 243,
+ 167,
+ 142,
+ 70
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 715,
+ "image_id": 562,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 147,
+ 190,
+ 120,
+ 264
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 716,
+ "image_id": 563,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 36,
+ 250,
+ 453,
+ 247
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 717,
+ "image_id": 564,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 394,
+ 148,
+ 115,
+ 103
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 718,
+ "image_id": 565,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 98,
+ 358,
+ 382,
+ 144
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 719,
+ "image_id": 566,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 17,
+ 330,
+ 489,
+ 176
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 720,
+ "image_id": 567,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 21,
+ 38,
+ 483,
+ 462
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 721,
+ "image_id": 568,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 306,
+ 184,
+ 200,
+ 322
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 722,
+ "image_id": 569,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 42,
+ 213,
+ 466,
+ 271
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 723,
+ "image_id": 570,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 145,
+ 85,
+ 260,
+ 59
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 724,
+ "image_id": 570,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 60,
+ 101,
+ 428,
+ 330
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 725,
+ "image_id": 571,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 92,
+ 425,
+ 410
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 726,
+ "image_id": 572,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 171,
+ 215,
+ 79,
+ 77
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 727,
+ "image_id": 572,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 194,
+ 390,
+ 228,
+ 113
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 728,
+ "image_id": 573,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 36,
+ 203,
+ 442,
+ 296
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 729,
+ "image_id": 574,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 220,
+ 144,
+ 185
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 730,
+ "image_id": 575,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 142,
+ 280,
+ 179,
+ 146
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 731,
+ "image_id": 576,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 162,
+ 256,
+ 219,
+ 225
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 732,
+ "image_id": 577,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 75,
+ 302,
+ 77,
+ 192
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 733,
+ "image_id": 578,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 58,
+ 189,
+ 449,
+ 264
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 734,
+ "image_id": 579,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 166,
+ 287,
+ 88,
+ 223
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 735,
+ "image_id": 580,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 244,
+ 297,
+ 46,
+ 161
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 736,
+ "image_id": 581,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 15,
+ 166,
+ 464,
+ 333
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 737,
+ "image_id": 582,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 260,
+ 145,
+ 85,
+ 183
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 738,
+ "image_id": 583,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 264,
+ 351,
+ 103,
+ 157
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 739,
+ "image_id": 583,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 201,
+ 339,
+ 86,
+ 152
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 740,
+ "image_id": 583,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 124,
+ 406,
+ 41,
+ 97
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 741,
+ "image_id": 584,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 260,
+ 189,
+ 244,
+ 318
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 742,
+ "image_id": 585,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 29,
+ 177,
+ 304,
+ 327
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 743,
+ "image_id": 586,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 203,
+ 133,
+ 48,
+ 74
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 744,
+ "image_id": 587,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 224,
+ 102,
+ 174,
+ 250
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 745,
+ "image_id": 588,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 270,
+ 65,
+ 178,
+ 299
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 746,
+ "image_id": 588,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 82,
+ 184,
+ 209,
+ 310
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 747,
+ "image_id": 589,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 157,
+ 364,
+ 70,
+ 53
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 748,
+ "image_id": 590,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 304,
+ 97,
+ 183,
+ 334
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 749,
+ "image_id": 591,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 32,
+ 110,
+ 409,
+ 354
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 750,
+ "image_id": 592,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 73,
+ 355,
+ 183,
+ 149
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 751,
+ "image_id": 593,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 95,
+ 265,
+ 82,
+ 226
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 752,
+ "image_id": 594,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 137,
+ 376,
+ 123,
+ 105
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 753,
+ "image_id": 595,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 248,
+ 123,
+ 225,
+ 379
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 754,
+ "image_id": 596,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 18,
+ 347,
+ 163,
+ 158
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 755,
+ "image_id": 597,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 217,
+ 8,
+ 202,
+ 278
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 756,
+ "image_id": 598,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 33,
+ 361,
+ 65,
+ 30
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 757,
+ "image_id": 598,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 279,
+ 268,
+ 228,
+ 179
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 758,
+ "image_id": 599,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 289,
+ 314,
+ 80,
+ 106
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 759,
+ "image_id": 600,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 124,
+ 379,
+ 54,
+ 98
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 760,
+ "image_id": 601,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 24,
+ 26,
+ 331,
+ 428
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 761,
+ "image_id": 602,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 166,
+ 396,
+ 123,
+ 79
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 762,
+ "image_id": 602,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 108,
+ 269,
+ 85,
+ 173
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 763,
+ "image_id": 603,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 99,
+ 337,
+ 160,
+ 136
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 764,
+ "image_id": 604,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 218,
+ 178,
+ 93,
+ 122
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 765,
+ "image_id": 605,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 279,
+ 247,
+ 140,
+ 180
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 766,
+ "image_id": 606,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 225,
+ 320,
+ 32,
+ 101
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 767,
+ "image_id": 606,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 169,
+ 330,
+ 32,
+ 94
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 768,
+ "image_id": 607,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 68,
+ 27,
+ 393,
+ 450
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 769,
+ "image_id": 608,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 183,
+ 126,
+ 251,
+ 280
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 770,
+ "image_id": 609,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 188,
+ 288,
+ 200,
+ 161
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 771,
+ "image_id": 610,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 349,
+ 169,
+ 156,
+ 336
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 772,
+ "image_id": 611,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 116,
+ 160,
+ 249,
+ 297
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 773,
+ "image_id": 612,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 48,
+ 44,
+ 432,
+ 308
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 774,
+ "image_id": 613,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 85,
+ 296,
+ 112,
+ 171
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 775,
+ "image_id": 614,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 279,
+ 151,
+ 182,
+ 189
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 776,
+ "image_id": 615,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 116,
+ 99,
+ 90,
+ 213
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 777,
+ "image_id": 615,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 222,
+ 66,
+ 63,
+ 91
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 778,
+ "image_id": 615,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 105,
+ 32,
+ 54,
+ 66
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 779,
+ "image_id": 615,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 203,
+ 86,
+ 110,
+ 278
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 780,
+ "image_id": 615,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 304,
+ 96,
+ 68,
+ 120
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 781,
+ "image_id": 615,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 429,
+ 6,
+ 68,
+ 64
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 782,
+ "image_id": 616,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 20,
+ 413,
+ 193,
+ 92
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 783,
+ "image_id": 617,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 350,
+ 304,
+ 146,
+ 199
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 784,
+ "image_id": 618,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 132,
+ 185,
+ 233,
+ 168
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 785,
+ "image_id": 619,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 177,
+ 66,
+ 101,
+ 80
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 786,
+ "image_id": 619,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 183,
+ 15,
+ 68,
+ 41
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 787,
+ "image_id": 620,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 196,
+ 311,
+ 197,
+ 178
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 788,
+ "image_id": 621,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 112,
+ 260,
+ 277,
+ 233
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 789,
+ "image_id": 622,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 417,
+ 275,
+ 76,
+ 153
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 790,
+ "image_id": 623,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 254,
+ 195,
+ 132,
+ 301
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 791,
+ "image_id": 623,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 69,
+ 46,
+ 82,
+ 186
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 792,
+ "image_id": 624,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 68,
+ 389,
+ 70,
+ 81
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 793,
+ "image_id": 625,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 233,
+ 177,
+ 208,
+ 179
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 794,
+ "image_id": 625,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 288,
+ 92,
+ 173,
+ 187
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 795,
+ "image_id": 626,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 232,
+ 213,
+ 72,
+ 86
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 796,
+ "image_id": 627,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 152,
+ 245,
+ 42,
+ 44
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 797,
+ "image_id": 627,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 239,
+ 287,
+ 63,
+ 40
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 798,
+ "image_id": 628,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 69,
+ 309,
+ 108,
+ 133
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 799,
+ "image_id": 628,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 250,
+ 76,
+ 102,
+ 110
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 800,
+ "image_id": 629,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 161,
+ 154,
+ 104,
+ 235
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 801,
+ "image_id": 630,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 236,
+ 153,
+ 88,
+ 85
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 802,
+ "image_id": 631,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 143,
+ 172,
+ 317,
+ 208
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 803,
+ "image_id": 632,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 203,
+ 328,
+ 33,
+ 60
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 804,
+ "image_id": 633,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 244,
+ 261,
+ 140,
+ 205
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 805,
+ "image_id": 634,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 242,
+ 59,
+ 121,
+ 368
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 806,
+ "image_id": 635,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 96,
+ 257,
+ 32,
+ 83
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 807,
+ "image_id": 636,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 185,
+ 279,
+ 135,
+ 217
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 808,
+ "image_id": 637,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 181,
+ 346,
+ 96,
+ 85
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 809,
+ "image_id": 638,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 124,
+ 102,
+ 149,
+ 163
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 810,
+ "image_id": 639,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 64,
+ 231,
+ 438,
+ 265
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 811,
+ "image_id": 639,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 218,
+ 366,
+ 136,
+ 138
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 812,
+ "image_id": 640,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 227,
+ 252,
+ 219,
+ 181
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 813,
+ "image_id": 641,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 167,
+ 319,
+ 162,
+ 171
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 814,
+ "image_id": 642,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 97,
+ 129,
+ 409,
+ 372
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 815,
+ "image_id": 643,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 176,
+ 225,
+ 216,
+ 192
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 816,
+ "image_id": 644,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 205,
+ 285,
+ 66,
+ 58
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 817,
+ "image_id": 645,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 149,
+ 215,
+ 131,
+ 233
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 818,
+ "image_id": 646,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 66,
+ 214,
+ 120,
+ 136
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 819,
+ "image_id": 646,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 193,
+ 226,
+ 120,
+ 117
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 820,
+ "image_id": 646,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 310,
+ 199,
+ 120,
+ 130
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 821,
+ "image_id": 647,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 245,
+ 298,
+ 40,
+ 112
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 822,
+ "image_id": 648,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 228,
+ 213,
+ 112,
+ 172
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 823,
+ "image_id": 648,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 73,
+ 227,
+ 162,
+ 273
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 824,
+ "image_id": 649,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 99,
+ 408,
+ 165,
+ 91
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 825,
+ "image_id": 650,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 95,
+ 177,
+ 254,
+ 259
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 826,
+ "image_id": 651,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 204,
+ 260,
+ 52,
+ 61
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 827,
+ "image_id": 652,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 254,
+ 184,
+ 127,
+ 215
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 828,
+ "image_id": 652,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 176,
+ 228,
+ 56,
+ 146
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 829,
+ "image_id": 652,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 86,
+ 225,
+ 55,
+ 123
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 830,
+ "image_id": 653,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 24,
+ 373,
+ 36,
+ 74
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 831,
+ "image_id": 654,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 164,
+ 290,
+ 20,
+ 40
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 832,
+ "image_id": 655,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 191,
+ 406,
+ 44,
+ 93
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 833,
+ "image_id": 656,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 228,
+ 229,
+ 46,
+ 44
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 834,
+ "image_id": 657,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 290,
+ 424,
+ 28,
+ 43
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 835,
+ "image_id": 657,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 108,
+ 405,
+ 47,
+ 62
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 836,
+ "image_id": 658,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 272,
+ 330,
+ 16,
+ 41
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 837,
+ "image_id": 659,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 261,
+ 251,
+ 120,
+ 237
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 838,
+ "image_id": 660,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 161,
+ 324,
+ 34,
+ 85
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 839,
+ "image_id": 661,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 132,
+ 343,
+ 104,
+ 96
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 840,
+ "image_id": 661,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 417,
+ 230,
+ 57,
+ 59
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 841,
+ "image_id": 662,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 240,
+ 305,
+ 45,
+ 49
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 842,
+ "image_id": 663,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 146,
+ 145,
+ 168,
+ 233
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 843,
+ "image_id": 664,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 34,
+ 212,
+ 77,
+ 125
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 844,
+ "image_id": 664,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 409,
+ 250,
+ 58,
+ 102
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 845,
+ "image_id": 665,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 31,
+ 238,
+ 428,
+ 252
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 846,
+ "image_id": 666,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 355,
+ 287,
+ 48,
+ 39
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 847,
+ "image_id": 667,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 156,
+ 181,
+ 45,
+ 113
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 848,
+ "image_id": 667,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 208,
+ 155,
+ 25,
+ 79
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 849,
+ "image_id": 667,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 287,
+ 184,
+ 58,
+ 158
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 850,
+ "image_id": 667,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 358,
+ 262,
+ 72,
+ 182
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 851,
+ "image_id": 668,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 238,
+ 141,
+ 245,
+ 177
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 852,
+ "image_id": 668,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 28,
+ 157,
+ 283,
+ 285
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 853,
+ "image_id": 669,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 246,
+ 352,
+ 78,
+ 63
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 854,
+ "image_id": 670,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 17,
+ 28,
+ 268,
+ 468
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 855,
+ "image_id": 670,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 242,
+ 234,
+ 209,
+ 271
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 856,
+ "image_id": 671,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 79,
+ 369,
+ 101,
+ 58
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 857,
+ "image_id": 672,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 126,
+ 203,
+ 196,
+ 172
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 858,
+ "image_id": 673,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 4,
+ 498,
+ 497
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 859,
+ "image_id": 674,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 251,
+ 317,
+ 160,
+ 158
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 860,
+ "image_id": 675,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 35,
+ 116,
+ 455,
+ 339
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 861,
+ "image_id": 676,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 148,
+ 91,
+ 228,
+ 317
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 862,
+ "image_id": 677,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 288,
+ 320,
+ 92,
+ 156
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 863,
+ "image_id": 678,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 58,
+ 408,
+ 136,
+ 72
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 864,
+ "image_id": 679,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 137,
+ 277,
+ 56,
+ 110
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 865,
+ "image_id": 680,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 223,
+ 192,
+ 236,
+ 247
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 866,
+ "image_id": 681,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 217,
+ 49,
+ 263,
+ 455
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 867,
+ "image_id": 681,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 106,
+ 108,
+ 174,
+ 328
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 868,
+ "image_id": 682,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 106,
+ 174,
+ 86,
+ 93
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 869,
+ "image_id": 682,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 389,
+ 118,
+ 67,
+ 73
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 870,
+ "image_id": 682,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 325,
+ 118,
+ 69,
+ 70
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 871,
+ "image_id": 682,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 119,
+ 100,
+ 30,
+ 52
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 872,
+ "image_id": 683,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 109,
+ 282,
+ 163,
+ 199
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 873,
+ "image_id": 684,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 214,
+ 153,
+ 144,
+ 273
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 874,
+ "image_id": 684,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 316,
+ 162,
+ 193
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 875,
+ "image_id": 685,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 203,
+ 266,
+ 236,
+ 118
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 876,
+ "image_id": 685,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 8,
+ 281,
+ 188,
+ 105
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 877,
+ "image_id": 686,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 215,
+ 175,
+ 143,
+ 305
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 878,
+ "image_id": 687,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 73,
+ 258,
+ 263,
+ 229
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 879,
+ "image_id": 688,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 256,
+ 293,
+ 180,
+ 211
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 880,
+ "image_id": 689,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 220,
+ 208,
+ 100,
+ 91
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 881,
+ "image_id": 689,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 130,
+ 226,
+ 93,
+ 90
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 882,
+ "image_id": 690,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 115,
+ 32,
+ 224,
+ 248
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 883,
+ "image_id": 690,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 14,
+ 70,
+ 489,
+ 436
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 884,
+ "image_id": 690,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 386,
+ 5,
+ 109,
+ 100
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 885,
+ "image_id": 690,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 125,
+ 5,
+ 129,
+ 72
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 886,
+ "image_id": 691,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 224,
+ 226,
+ 87,
+ 142
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 887,
+ "image_id": 692,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 155,
+ 332,
+ 192,
+ 106
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 888,
+ "image_id": 692,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 371,
+ 321,
+ 78,
+ 107
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 889,
+ "image_id": 693,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 176,
+ 140,
+ 321
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 890,
+ "image_id": 693,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 176,
+ 35,
+ 311,
+ 331
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 891,
+ "image_id": 694,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 171,
+ 179,
+ 167,
+ 160
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 892,
+ "image_id": 695,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 167,
+ 205,
+ 126,
+ 122
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 893,
+ "image_id": 696,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 204,
+ 202,
+ 213,
+ 205
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 894,
+ "image_id": 697,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 50,
+ 275,
+ 126,
+ 230
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 895,
+ "image_id": 697,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 72,
+ 261,
+ 339,
+ 245
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 896,
+ "image_id": 698,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 193,
+ 229,
+ 212,
+ 215
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 897,
+ "image_id": 698,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 53,
+ 278,
+ 73,
+ 168
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 898,
+ "image_id": 699,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 359,
+ 295,
+ 136,
+ 154
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 899,
+ "image_id": 700,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 34,
+ 347,
+ 76,
+ 103
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 900,
+ "image_id": 701,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 183,
+ 212,
+ 211,
+ 64
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 901,
+ "image_id": 701,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 21,
+ 307,
+ 277,
+ 138
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 902,
+ "image_id": 702,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 210,
+ 331,
+ 94,
+ 142
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 903,
+ "image_id": 703,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 110,
+ 220,
+ 83,
+ 262
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 904,
+ "image_id": 703,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 199,
+ 232,
+ 165,
+ 163
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 905,
+ "image_id": 704,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 19,
+ 363,
+ 468,
+ 142
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 906,
+ "image_id": 704,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 232,
+ 226,
+ 122,
+ 148
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 907,
+ "image_id": 705,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 197,
+ 217,
+ 98,
+ 122
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 908,
+ "image_id": 706,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 100,
+ 156,
+ 42,
+ 140
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 909,
+ "image_id": 706,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 306,
+ 167,
+ 25,
+ 69
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 910,
+ "image_id": 707,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 297,
+ 186,
+ 66,
+ 93
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 911,
+ "image_id": 708,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 200,
+ 337,
+ 214,
+ 118
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 912,
+ "image_id": 709,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 17,
+ 83,
+ 454,
+ 296
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 913,
+ "image_id": 710,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 144,
+ 114,
+ 261,
+ 290
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 914,
+ "image_id": 711,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 166,
+ 175,
+ 119,
+ 151
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 915,
+ "image_id": 711,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 342,
+ 200,
+ 139,
+ 119
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 916,
+ "image_id": 712,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 17,
+ 100,
+ 427,
+ 350
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 917,
+ "image_id": 713,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 99,
+ 125,
+ 254,
+ 316
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 918,
+ "image_id": 714,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 20,
+ 56,
+ 479,
+ 388
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 919,
+ "image_id": 715,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 176,
+ 346,
+ 296
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 920,
+ "image_id": 716,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 133,
+ 120,
+ 341,
+ 369
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 921,
+ "image_id": 717,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 126,
+ 33,
+ 372,
+ 427
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 922,
+ "image_id": 717,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 30,
+ 127,
+ 224,
+ 309
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 923,
+ "image_id": 718,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 132,
+ 100,
+ 203,
+ 381
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 924,
+ "image_id": 719,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 24,
+ 13,
+ 259,
+ 469
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 925,
+ "image_id": 720,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 190,
+ 313,
+ 21,
+ 53
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 926,
+ "image_id": 721,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 277,
+ 23,
+ 195,
+ 475
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 927,
+ "image_id": 722,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 39,
+ 80,
+ 412,
+ 414
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 928,
+ "image_id": 723,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 284,
+ 258,
+ 104,
+ 49
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 929,
+ "image_id": 723,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 218,
+ 258,
+ 62,
+ 57
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 930,
+ "image_id": 724,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 231,
+ 192,
+ 107,
+ 122
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 931,
+ "image_id": 725,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 276,
+ 217,
+ 56,
+ 61
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 932,
+ "image_id": 726,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 215,
+ 239,
+ 104,
+ 108
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 933,
+ "image_id": 726,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 244,
+ 153,
+ 100,
+ 101
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 934,
+ "image_id": 727,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 189,
+ 142,
+ 197,
+ 249
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 935,
+ "image_id": 728,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 72,
+ 361,
+ 98,
+ 91
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 936,
+ "image_id": 729,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 149,
+ 173,
+ 154,
+ 332
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 937,
+ "image_id": 730,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 263,
+ 238,
+ 58,
+ 84
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 938,
+ "image_id": 731,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 41,
+ 295,
+ 157,
+ 189
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 939,
+ "image_id": 732,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 101,
+ 255,
+ 124,
+ 176
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 940,
+ "image_id": 733,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 41,
+ 22,
+ 390,
+ 484
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 941,
+ "image_id": 734,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 283,
+ 318,
+ 36,
+ 104
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 942,
+ "image_id": 735,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 19,
+ 9,
+ 420,
+ 472
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 943,
+ "image_id": 736,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 64,
+ 413,
+ 86,
+ 74
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 944,
+ "image_id": 737,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 19,
+ 8,
+ 382,
+ 316
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 945,
+ "image_id": 738,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 309,
+ 259,
+ 76,
+ 188
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 946,
+ "image_id": 739,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 33,
+ 69,
+ 336,
+ 430
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 947,
+ "image_id": 740,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 74,
+ 139,
+ 300,
+ 333
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 948,
+ "image_id": 741,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 245,
+ 259,
+ 213,
+ 165
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 949,
+ "image_id": 742,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 117,
+ 19,
+ 295,
+ 288
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 950,
+ "image_id": 743,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 51,
+ 91,
+ 448,
+ 282
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 951,
+ "image_id": 744,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 286,
+ 114,
+ 221,
+ 334
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 952,
+ "image_id": 745,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 176,
+ 94,
+ 325,
+ 291
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 953,
+ "image_id": 746,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 70,
+ 219,
+ 146,
+ 92
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 954,
+ "image_id": 746,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 81,
+ 298,
+ 204,
+ 96
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 955,
+ "image_id": 747,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 37,
+ 287,
+ 462
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 956,
+ "image_id": 748,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 159,
+ 177,
+ 248,
+ 128
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 957,
+ "image_id": 749,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 69,
+ 57,
+ 428,
+ 395
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 958,
+ "image_id": 750,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 162,
+ 201,
+ 179,
+ 283
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 959,
+ "image_id": 751,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 286,
+ 389,
+ 69,
+ 74
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 960,
+ "image_id": 752,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 74,
+ 158,
+ 192,
+ 332
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 961,
+ "image_id": 752,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 208,
+ 169,
+ 282,
+ 335
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 962,
+ "image_id": 753,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 204,
+ 349,
+ 193,
+ 146
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 963,
+ "image_id": 753,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 12,
+ 45,
+ 215,
+ 427
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 964,
+ "image_id": 754,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 216,
+ 182,
+ 181,
+ 324
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 965,
+ "image_id": 755,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 122,
+ 371,
+ 349
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 966,
+ "image_id": 756,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 143,
+ 133,
+ 352,
+ 139
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 967,
+ "image_id": 757,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 96,
+ 67,
+ 291,
+ 433
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 968,
+ "image_id": 758,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 183,
+ 275,
+ 160,
+ 163
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 969,
+ "image_id": 759,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 256,
+ 285,
+ 229,
+ 177
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 970,
+ "image_id": 760,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 39,
+ 103,
+ 366,
+ 387
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 971,
+ "image_id": 761,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 150,
+ 16,
+ 314,
+ 487
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 972,
+ "image_id": 762,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 49,
+ 35,
+ 452,
+ 461
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 973,
+ "image_id": 763,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 56,
+ 12,
+ 386,
+ 483
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 974,
+ "image_id": 764,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 217,
+ 133,
+ 93,
+ 265
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 975,
+ "image_id": 764,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 300,
+ 207,
+ 129,
+ 298
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 976,
+ "image_id": 765,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 129,
+ 320,
+ 69,
+ 116
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 977,
+ "image_id": 766,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 56,
+ 119,
+ 445,
+ 374
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 978,
+ "image_id": 767,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 8,
+ 8,
+ 432,
+ 453
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 979,
+ "image_id": 768,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 22,
+ 12,
+ 306,
+ 489
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 980,
+ "image_id": 769,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 11,
+ 509,
+ 469
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 981,
+ "image_id": 770,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 189,
+ 131,
+ 267,
+ 373
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 982,
+ "image_id": 770,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 314,
+ 36,
+ 134,
+ 217
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 983,
+ "image_id": 770,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 30,
+ 46,
+ 265,
+ 222
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 984,
+ "image_id": 771,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 24,
+ 32,
+ 318,
+ 382
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 985,
+ "image_id": 772,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 167,
+ 24,
+ 157,
+ 484
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 986,
+ "image_id": 772,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 182,
+ 211,
+ 210
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 987,
+ "image_id": 773,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 81,
+ 23,
+ 232,
+ 478
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 988,
+ "image_id": 774,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 11,
+ 25,
+ 482,
+ 399
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 989,
+ "image_id": 775,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 27,
+ 42,
+ 302,
+ 443
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 990,
+ "image_id": 776,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 91,
+ 184,
+ 231,
+ 319
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 991,
+ "image_id": 777,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 107,
+ 17,
+ 341,
+ 470
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 992,
+ "image_id": 778,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 232,
+ 135,
+ 264,
+ 371
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 993,
+ "image_id": 779,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 21,
+ 45,
+ 411,
+ 424
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 994,
+ "image_id": 780,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 193,
+ 54,
+ 305,
+ 414
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 995,
+ "image_id": 781,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 113,
+ 25,
+ 392,
+ 391
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 996,
+ "image_id": 782,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 15,
+ 160,
+ 384,
+ 332
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 997,
+ "image_id": 783,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 93,
+ 109,
+ 160,
+ 383
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 998,
+ "image_id": 784,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 294,
+ 138,
+ 114,
+ 256
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 999,
+ "image_id": 785,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 144,
+ 45,
+ 296,
+ 243
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1000,
+ "image_id": 786,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 248,
+ 215,
+ 97,
+ 137
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1001,
+ "image_id": 787,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 333,
+ 342,
+ 38,
+ 35
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1002,
+ "image_id": 788,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 133,
+ 44,
+ 291,
+ 439
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1003,
+ "image_id": 789,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 182,
+ 34,
+ 300,
+ 461
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1004,
+ "image_id": 790,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 92,
+ 47,
+ 382,
+ 429
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1005,
+ "image_id": 791,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 220,
+ 129,
+ 84,
+ 199
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1006,
+ "image_id": 792,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 47,
+ 41,
+ 434,
+ 303
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1007,
+ "image_id": 793,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 227,
+ 101,
+ 78,
+ 171
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1008,
+ "image_id": 794,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 155,
+ 97,
+ 349,
+ 409
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1009,
+ "image_id": 795,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 161,
+ 228,
+ 52,
+ 67
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1010,
+ "image_id": 796,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 349,
+ 163,
+ 50,
+ 154
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1011,
+ "image_id": 796,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 429,
+ 194,
+ 39,
+ 130
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1012,
+ "image_id": 796,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 300,
+ 137,
+ 30,
+ 87
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1013,
+ "image_id": 796,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 257,
+ 162,
+ 71,
+ 133
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1014,
+ "image_id": 796,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 181,
+ 160,
+ 66,
+ 137
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1015,
+ "image_id": 796,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 61,
+ 136,
+ 40,
+ 110
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1016,
+ "image_id": 796,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 122,
+ 84,
+ 61,
+ 128
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1017,
+ "image_id": 796,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 187,
+ 100,
+ 75,
+ 78
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1018,
+ "image_id": 797,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 210,
+ 193,
+ 235,
+ 284
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1019,
+ "image_id": 798,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 254,
+ 230,
+ 168,
+ 77
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1020,
+ "image_id": 799,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 142,
+ 132,
+ 221,
+ 293
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1021,
+ "image_id": 800,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 233,
+ 203,
+ 159,
+ 138
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1022,
+ "image_id": 801,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 346,
+ 264,
+ 133,
+ 222
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1023,
+ "image_id": 802,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 335,
+ 178,
+ 88,
+ 161
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1024,
+ "image_id": 803,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 184,
+ 57,
+ 54,
+ 130
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1025,
+ "image_id": 803,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 195,
+ 198,
+ 87,
+ 152
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1026,
+ "image_id": 804,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 206,
+ 236,
+ 73,
+ 124
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1027,
+ "image_id": 805,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 155,
+ 122,
+ 177,
+ 340
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1028,
+ "image_id": 806,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 34,
+ 116,
+ 436,
+ 272
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1029,
+ "image_id": 807,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 75,
+ 128,
+ 335,
+ 338
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1030,
+ "image_id": 808,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 67,
+ 15,
+ 346,
+ 471
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1031,
+ "image_id": 809,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 181,
+ 304,
+ 168,
+ 142
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1032,
+ "image_id": 810,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 246,
+ 184,
+ 206,
+ 203
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1033,
+ "image_id": 811,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 231,
+ 138,
+ 219,
+ 210
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1034,
+ "image_id": 812,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 175,
+ 108,
+ 183,
+ 365
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1035,
+ "image_id": 813,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 241,
+ 175,
+ 173,
+ 240
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1036,
+ "image_id": 813,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 101,
+ 134,
+ 260,
+ 286
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1037,
+ "image_id": 814,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 298,
+ 219,
+ 206,
+ 284
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1038,
+ "image_id": 814,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 170,
+ 129,
+ 241,
+ 334
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1039,
+ "image_id": 815,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 234,
+ 329,
+ 127,
+ 154
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1040,
+ "image_id": 816,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 58,
+ 123,
+ 134,
+ 357
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1041,
+ "image_id": 817,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 104,
+ 395,
+ 73,
+ 104
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1042,
+ "image_id": 818,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 85,
+ 156,
+ 397,
+ 325
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1043,
+ "image_id": 819,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 112,
+ 57,
+ 316,
+ 426
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1044,
+ "image_id": 820,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 186,
+ 232,
+ 316,
+ 249
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1045,
+ "image_id": 821,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 159,
+ 164,
+ 158,
+ 179
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1046,
+ "image_id": 822,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 162,
+ 139,
+ 220,
+ 219
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1047,
+ "image_id": 823,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 52,
+ 288,
+ 244,
+ 180
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1048,
+ "image_id": 824,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 192,
+ 70,
+ 147,
+ 375
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1049,
+ "image_id": 825,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 175,
+ 210,
+ 190,
+ 282
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1050,
+ "image_id": 826,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 7,
+ 106,
+ 343,
+ 377
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1051,
+ "image_id": 827,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 134,
+ 68,
+ 309,
+ 329
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1052,
+ "image_id": 828,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 96,
+ 73,
+ 217,
+ 423
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1053,
+ "image_id": 829,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 193,
+ 92,
+ 298,
+ 339
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1054,
+ "image_id": 830,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 287,
+ 237,
+ 39,
+ 55
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1055,
+ "image_id": 831,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 179,
+ 72,
+ 193,
+ 403
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1056,
+ "image_id": 832,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 195,
+ 109,
+ 194,
+ 204
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1057,
+ "image_id": 833,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 331,
+ 88,
+ 125,
+ 320
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1058,
+ "image_id": 833,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 13,
+ 75,
+ 348,
+ 419
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1059,
+ "image_id": 834,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 107,
+ 184,
+ 310,
+ 166
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1060,
+ "image_id": 835,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 87,
+ 249,
+ 362,
+ 244
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1061,
+ "image_id": 836,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 75,
+ 28,
+ 266,
+ 466
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1062,
+ "image_id": 837,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 89,
+ 38,
+ 335,
+ 450
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1063,
+ "image_id": 837,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 256,
+ 500,
+ 247
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1064,
+ "image_id": 838,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 52,
+ 85,
+ 265,
+ 315
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1065,
+ "image_id": 839,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 73,
+ 16,
+ 304,
+ 468
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1066,
+ "image_id": 839,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 314,
+ 6,
+ 183,
+ 373
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1067,
+ "image_id": 840,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 277,
+ 158,
+ 103,
+ 146
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1068,
+ "image_id": 840,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 83,
+ 74,
+ 163,
+ 252
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1069,
+ "image_id": 841,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 185,
+ 165,
+ 112,
+ 142
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1070,
+ "image_id": 842,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 40,
+ 229,
+ 359,
+ 275
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1071,
+ "image_id": 843,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 43,
+ 96,
+ 261,
+ 326
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1072,
+ "image_id": 844,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 293,
+ 308,
+ 48,
+ 108
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1073,
+ "image_id": 845,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 110,
+ 70,
+ 236,
+ 310
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1074,
+ "image_id": 846,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 38,
+ 92,
+ 366,
+ 366
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1075,
+ "image_id": 847,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 187,
+ 49,
+ 173,
+ 384
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1076,
+ "image_id": 848,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 143,
+ 66,
+ 349,
+ 432
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1077,
+ "image_id": 849,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 53,
+ 10,
+ 418,
+ 486
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1078,
+ "image_id": 850,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 268,
+ 40,
+ 225,
+ 447
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1079,
+ "image_id": 850,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 13,
+ 192,
+ 458
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1080,
+ "image_id": 851,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 25,
+ 55,
+ 423,
+ 446
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1081,
+ "image_id": 852,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 218,
+ 198,
+ 92,
+ 163
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1082,
+ "image_id": 853,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 144,
+ 195,
+ 181,
+ 197
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1083,
+ "image_id": 854,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 154,
+ 147,
+ 153,
+ 272
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1084,
+ "image_id": 855,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 22,
+ 10,
+ 464,
+ 459
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1085,
+ "image_id": 856,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 158,
+ 149,
+ 154,
+ 308
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1086,
+ "image_id": 857,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 282,
+ 330,
+ 47,
+ 98
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1087,
+ "image_id": 858,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 188,
+ 113,
+ 124,
+ 171
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1088,
+ "image_id": 859,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 87,
+ 109,
+ 290,
+ 333
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1089,
+ "image_id": 860,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 24,
+ 63,
+ 465,
+ 434
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1090,
+ "image_id": 861,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 75,
+ 293,
+ 140,
+ 75
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1091,
+ "image_id": 862,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 12,
+ 288,
+ 250,
+ 207
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1092,
+ "image_id": 862,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 147,
+ 142,
+ 307,
+ 129
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1093,
+ "image_id": 863,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 37,
+ 67,
+ 386,
+ 431
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1094,
+ "image_id": 864,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 209,
+ 15,
+ 258,
+ 487
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1095,
+ "image_id": 865,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 93,
+ 114,
+ 145,
+ 373
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1096,
+ "image_id": 865,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 273,
+ 121,
+ 134,
+ 329
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1097,
+ "image_id": 866,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 171,
+ 155,
+ 149,
+ 305
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1098,
+ "image_id": 867,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 142,
+ 104,
+ 314,
+ 376
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1099,
+ "image_id": 868,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 155,
+ 102,
+ 267,
+ 331
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1100,
+ "image_id": 869,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 85,
+ 82,
+ 372,
+ 411
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1101,
+ "image_id": 870,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 150,
+ 103,
+ 196,
+ 372
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1102,
+ "image_id": 871,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 80,
+ 124,
+ 349,
+ 329
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1103,
+ "image_id": 872,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 47,
+ 25,
+ 364,
+ 461
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1104,
+ "image_id": 873,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 208,
+ 145,
+ 208,
+ 358
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1105,
+ "image_id": 874,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 122,
+ 58,
+ 213,
+ 401
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1106,
+ "image_id": 875,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 59,
+ 18,
+ 340,
+ 455
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1107,
+ "image_id": 876,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 33,
+ 183,
+ 430
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1108,
+ "image_id": 876,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 401,
+ 164,
+ 97,
+ 112
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1109,
+ "image_id": 876,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 357,
+ 223,
+ 152,
+ 148
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1110,
+ "image_id": 877,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 19,
+ 72,
+ 441,
+ 387
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1111,
+ "image_id": 878,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 28,
+ 296,
+ 259
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1112,
+ "image_id": 878,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 117,
+ 192,
+ 386,
+ 304
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1113,
+ "image_id": 879,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 137,
+ 55,
+ 209,
+ 383
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1114,
+ "image_id": 880,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 14,
+ 98,
+ 217,
+ 389
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1115,
+ "image_id": 880,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 319,
+ 210,
+ 182,
+ 292
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1116,
+ "image_id": 881,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 282,
+ 245,
+ 159,
+ 147
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1117,
+ "image_id": 882,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 17,
+ 344,
+ 462
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1118,
+ "image_id": 883,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 13,
+ 154,
+ 268,
+ 284
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1119,
+ "image_id": 884,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 176,
+ 97,
+ 175,
+ 336
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1120,
+ "image_id": 885,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 247,
+ 213,
+ 228
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1121,
+ "image_id": 886,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 201,
+ 129,
+ 306,
+ 275
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1122,
+ "image_id": 887,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 65,
+ 233,
+ 204,
+ 191
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1123,
+ "image_id": 888,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 50,
+ 197,
+ 106,
+ 210
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1124,
+ "image_id": 888,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 180,
+ 191,
+ 123,
+ 150
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1125,
+ "image_id": 889,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 13,
+ 23,
+ 484,
+ 374
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1126,
+ "image_id": 890,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 120,
+ 143,
+ 367,
+ 350
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1127,
+ "image_id": 891,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 18,
+ 72,
+ 444,
+ 393
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1128,
+ "image_id": 892,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 68,
+ 4,
+ 304,
+ 502
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1129,
+ "image_id": 893,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 143,
+ 17,
+ 330,
+ 462
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1130,
+ "image_id": 894,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 45,
+ 13,
+ 421,
+ 468
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1131,
+ "image_id": 895,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 159,
+ 102,
+ 151,
+ 335
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1132,
+ "image_id": 896,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 272,
+ 191,
+ 226,
+ 301
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1133,
+ "image_id": 896,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 39,
+ 118,
+ 223,
+ 383
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1134,
+ "image_id": 897,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 141,
+ 13,
+ 207,
+ 450
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1135,
+ "image_id": 898,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 187,
+ 155,
+ 126,
+ 218
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1136,
+ "image_id": 898,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 36,
+ 43,
+ 317,
+ 165
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1137,
+ "image_id": 899,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 22,
+ 15,
+ 429,
+ 484
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1138,
+ "image_id": 900,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 138,
+ 6,
+ 352,
+ 484
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1139,
+ "image_id": 901,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 22,
+ 17,
+ 423,
+ 475
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1140,
+ "image_id": 902,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 29,
+ 237,
+ 355,
+ 255
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1141,
+ "image_id": 903,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 189,
+ 133,
+ 302,
+ 360
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1142,
+ "image_id": 904,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 106,
+ 219,
+ 237,
+ 157
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1143,
+ "image_id": 905,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 51,
+ 363,
+ 449,
+ 92
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1144,
+ "image_id": 906,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 130,
+ 27,
+ 372,
+ 465
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1145,
+ "image_id": 907,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 176,
+ 206,
+ 294,
+ 248
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1146,
+ "image_id": 908,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 219,
+ 229,
+ 128,
+ 248
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1147,
+ "image_id": 908,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 143,
+ 125,
+ 155,
+ 229
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1148,
+ "image_id": 909,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 11,
+ 86,
+ 430,
+ 361
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1149,
+ "image_id": 910,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 199,
+ 289,
+ 283,
+ 214
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1150,
+ "image_id": 910,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 258,
+ 181,
+ 106,
+ 126
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1151,
+ "image_id": 910,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 348,
+ 218,
+ 161,
+ 159
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1152,
+ "image_id": 910,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 15,
+ 273,
+ 207,
+ 208
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1153,
+ "image_id": 910,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 27,
+ 226,
+ 166,
+ 84
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1154,
+ "image_id": 910,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 152,
+ 214,
+ 122,
+ 128
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1155,
+ "image_id": 911,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 78,
+ 86,
+ 288,
+ 409
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1156,
+ "image_id": 912,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 26,
+ 170,
+ 204,
+ 265
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1157,
+ "image_id": 913,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 48,
+ 6,
+ 442,
+ 499
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1158,
+ "image_id": 914,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 70,
+ 39,
+ 393,
+ 454
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1159,
+ "image_id": 915,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 180,
+ 212,
+ 87,
+ 76
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1160,
+ "image_id": 916,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 70,
+ 378,
+ 279
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1161,
+ "image_id": 917,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 13,
+ 23,
+ 357,
+ 420
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1162,
+ "image_id": 918,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 19,
+ 9,
+ 195,
+ 477
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1163,
+ "image_id": 919,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 145,
+ 126,
+ 205,
+ 359
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1164,
+ "image_id": 920,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 218,
+ 253,
+ 67,
+ 175
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1165,
+ "image_id": 921,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 16,
+ 111,
+ 153,
+ 96
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1166,
+ "image_id": 921,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 148,
+ 15,
+ 324,
+ 470
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1167,
+ "image_id": 922,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 27,
+ 84,
+ 344,
+ 401
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1168,
+ "image_id": 923,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 119,
+ 140,
+ 109,
+ 185
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1169,
+ "image_id": 923,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 236,
+ 157,
+ 217,
+ 278
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1170,
+ "image_id": 924,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 119,
+ 277,
+ 316
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1171,
+ "image_id": 924,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 270,
+ 180,
+ 234,
+ 210
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1172,
+ "image_id": 925,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 146,
+ 117,
+ 202,
+ 218
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1173,
+ "image_id": 926,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 62,
+ 84,
+ 432,
+ 404
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1174,
+ "image_id": 927,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 256,
+ 109,
+ 195,
+ 395
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1175,
+ "image_id": 927,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 13,
+ 8,
+ 220,
+ 487
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1176,
+ "image_id": 928,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 137,
+ 138,
+ 267,
+ 187
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1177,
+ "image_id": 929,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 200,
+ 153,
+ 297,
+ 309
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1178,
+ "image_id": 929,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 174,
+ 177,
+ 132
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1179,
+ "image_id": 930,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 214,
+ 147,
+ 189,
+ 305
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1180,
+ "image_id": 931,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 88,
+ 81,
+ 315,
+ 360
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1181,
+ "image_id": 932,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 47,
+ 50,
+ 434,
+ 443
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1182,
+ "image_id": 933,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 310,
+ 211,
+ 88,
+ 174
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1183,
+ "image_id": 933,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 197,
+ 84,
+ 210,
+ 208
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1184,
+ "image_id": 934,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 78,
+ 74,
+ 359,
+ 398
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1185,
+ "image_id": 935,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 75,
+ 116,
+ 298,
+ 308
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1186,
+ "image_id": 935,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 375,
+ 118,
+ 91,
+ 124
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1187,
+ "image_id": 936,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 19,
+ 221,
+ 379,
+ 270
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1188,
+ "image_id": 937,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 33,
+ 25,
+ 387,
+ 443
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1189,
+ "image_id": 938,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 68,
+ 104,
+ 297,
+ 281
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1190,
+ "image_id": 939,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 25,
+ 19,
+ 423,
+ 443
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1191,
+ "image_id": 940,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 194,
+ 30,
+ 220,
+ 416
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1192,
+ "image_id": 941,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 57,
+ 195,
+ 295,
+ 307
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1193,
+ "image_id": 942,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 14,
+ 79,
+ 371,
+ 382
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1194,
+ "image_id": 943,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 26,
+ 180,
+ 288,
+ 324
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1195,
+ "image_id": 944,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 132,
+ 163,
+ 224,
+ 193
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1196,
+ "image_id": 945,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 38,
+ 4,
+ 418,
+ 501
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1197,
+ "image_id": 946,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 171,
+ 155,
+ 316,
+ 286
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1198,
+ "image_id": 947,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 162,
+ 122,
+ 343,
+ 264
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1199,
+ "image_id": 948,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 118,
+ 126,
+ 373,
+ 271
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1200,
+ "image_id": 949,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 107,
+ 36,
+ 327,
+ 442
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1201,
+ "image_id": 950,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 18,
+ 123,
+ 487,
+ 376
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1202,
+ "image_id": 951,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 79,
+ 21,
+ 375,
+ 475
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1203,
+ "image_id": 952,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 171,
+ 182,
+ 129,
+ 167
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1204,
+ "image_id": 953,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 201,
+ 25,
+ 89,
+ 247
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1205,
+ "image_id": 954,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 37,
+ 236,
+ 270,
+ 256
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1206,
+ "image_id": 955,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 66,
+ 159,
+ 167,
+ 196
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1207,
+ "image_id": 956,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 177,
+ 301,
+ 95,
+ 46
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1208,
+ "image_id": 957,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 147,
+ 217,
+ 176,
+ 278
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1209,
+ "image_id": 958,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 119,
+ 158,
+ 186,
+ 210
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1210,
+ "image_id": 959,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 33,
+ 25,
+ 424,
+ 436
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1211,
+ "image_id": 960,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 208,
+ 314,
+ 52,
+ 124
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1212,
+ "image_id": 961,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 25,
+ 40,
+ 410,
+ 450
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1213,
+ "image_id": 962,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 230,
+ 245,
+ 173,
+ 101
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1214,
+ "image_id": 962,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 198,
+ 191,
+ 166,
+ 79
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1215,
+ "image_id": 963,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 125,
+ 53,
+ 246,
+ 412
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1216,
+ "image_id": 964,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 306,
+ 185,
+ 44,
+ 109
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1217,
+ "image_id": 965,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 345,
+ 386,
+ 26,
+ 40
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1218,
+ "image_id": 966,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 157,
+ 281,
+ 124,
+ 150
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1219,
+ "image_id": 967,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 143,
+ 111,
+ 227,
+ 239
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1220,
+ "image_id": 968,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 126,
+ 4,
+ 367,
+ 465
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1221,
+ "image_id": 969,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 39,
+ 144,
+ 324,
+ 340
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1222,
+ "image_id": 970,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 172,
+ 97,
+ 149,
+ 248
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1223,
+ "image_id": 971,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 137,
+ 132,
+ 219,
+ 337
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1224,
+ "image_id": 972,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 50,
+ 138,
+ 344,
+ 345
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1225,
+ "image_id": 973,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 126,
+ 225,
+ 245,
+ 231
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1226,
+ "image_id": 974,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 36,
+ 398,
+ 459
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1227,
+ "image_id": 975,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 341,
+ 229,
+ 121,
+ 236
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1228,
+ "image_id": 976,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 69,
+ 53,
+ 399,
+ 385
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1229,
+ "image_id": 977,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 73,
+ 224,
+ 306,
+ 217
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1230,
+ "image_id": 978,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 68,
+ 77,
+ 310,
+ 139
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1231,
+ "image_id": 979,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 332,
+ 376,
+ 37,
+ 42
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1232,
+ "image_id": 980,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 337,
+ 342,
+ 65,
+ 82
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1233,
+ "image_id": 981,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 190,
+ 158,
+ 190,
+ 172
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1234,
+ "image_id": 982,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 18,
+ 207,
+ 486,
+ 223
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1235,
+ "image_id": 983,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 130,
+ 73,
+ 361,
+ 409
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1236,
+ "image_id": 984,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 15,
+ 55,
+ 403,
+ 440
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1237,
+ "image_id": 985,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 190,
+ 112,
+ 281,
+ 193
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1238,
+ "image_id": 986,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 266,
+ 109,
+ 187,
+ 392
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1239,
+ "image_id": 987,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 116,
+ 182,
+ 141,
+ 239
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1240,
+ "image_id": 988,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 210,
+ 221,
+ 121,
+ 199
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1241,
+ "image_id": 988,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 107,
+ 317,
+ 121,
+ 162
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1242,
+ "image_id": 989,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 102,
+ 93,
+ 353,
+ 400
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1243,
+ "image_id": 990,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 176,
+ 301,
+ 84,
+ 157
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1244,
+ "image_id": 990,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 119,
+ 355,
+ 55,
+ 67
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1245,
+ "image_id": 991,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 189,
+ 212,
+ 171,
+ 181
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1246,
+ "image_id": 992,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 42,
+ 77,
+ 413,
+ 369
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1247,
+ "image_id": 993,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 209,
+ 167,
+ 67,
+ 79
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1248,
+ "image_id": 994,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 296,
+ 188,
+ 98,
+ 172
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1249,
+ "image_id": 995,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 170,
+ 106,
+ 167,
+ 337
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1250,
+ "image_id": 996,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 43,
+ 345,
+ 210,
+ 95
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1251,
+ "image_id": 997,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 213,
+ 225,
+ 138,
+ 201
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1252,
+ "image_id": 998,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 124,
+ 331,
+ 163,
+ 109
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1253,
+ "image_id": 999,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 17,
+ 72,
+ 472,
+ 392
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1254,
+ "image_id": 1000,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 50,
+ 25,
+ 294,
+ 472
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1255,
+ "image_id": 1001,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 83,
+ 119,
+ 277,
+ 375
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1256,
+ "image_id": 1002,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 277,
+ 273,
+ 100,
+ 226
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1257,
+ "image_id": 1003,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 125,
+ 30,
+ 369,
+ 449
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1258,
+ "image_id": 1004,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 171,
+ 72,
+ 175,
+ 418
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1259,
+ "image_id": 1005,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 100,
+ 254,
+ 134,
+ 184
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1260,
+ "image_id": 1006,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 169,
+ 142,
+ 45,
+ 82
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1261,
+ "image_id": 1007,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 298,
+ 107,
+ 128,
+ 276
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1262,
+ "image_id": 1007,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 217,
+ 249,
+ 94,
+ 121
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1263,
+ "image_id": 1008,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 118,
+ 347,
+ 218,
+ 105
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1264,
+ "image_id": 1009,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 115,
+ 293,
+ 215,
+ 173
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1265,
+ "image_id": 1009,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 159,
+ 11,
+ 197,
+ 210
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1266,
+ "image_id": 1010,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 211,
+ 356,
+ 78,
+ 53
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1267,
+ "image_id": 1011,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 209,
+ 275,
+ 94,
+ 126
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1268,
+ "image_id": 1011,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 119,
+ 191,
+ 36,
+ 62
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1269,
+ "image_id": 1012,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 302,
+ 219,
+ 189,
+ 210
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1270,
+ "image_id": 1013,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 187,
+ 53,
+ 292,
+ 444
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1271,
+ "image_id": 1014,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 24,
+ 212,
+ 161,
+ 259
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1272,
+ "image_id": 1015,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 145,
+ 314,
+ 196,
+ 130
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1273,
+ "image_id": 1016,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 140,
+ 72,
+ 261,
+ 352
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1274,
+ "image_id": 1017,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 277,
+ 159,
+ 92,
+ 231
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1275,
+ "image_id": 1018,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 54,
+ 56,
+ 393,
+ 245
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1276,
+ "image_id": 1019,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 120,
+ 70,
+ 362,
+ 408
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1277,
+ "image_id": 1020,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 36,
+ 42,
+ 418,
+ 308
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1278,
+ "image_id": 1021,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 229,
+ 261,
+ 74,
+ 132
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1279,
+ "image_id": 1022,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 320,
+ 354,
+ 105,
+ 69
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1280,
+ "image_id": 1023,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 41,
+ 124,
+ 449,
+ 350
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1281,
+ "image_id": 1024,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 166,
+ 229,
+ 62,
+ 88
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1282,
+ "image_id": 1024,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 271,
+ 176,
+ 36,
+ 31
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1283,
+ "image_id": 1025,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 109,
+ 144,
+ 294,
+ 355
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1284,
+ "image_id": 1026,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 99,
+ 176,
+ 102,
+ 66
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1285,
+ "image_id": 1027,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 54,
+ 137,
+ 360,
+ 292
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1286,
+ "image_id": 1028,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 117,
+ 140,
+ 179,
+ 171
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1287,
+ "image_id": 1029,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 221,
+ 129,
+ 280,
+ 233
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1288,
+ "image_id": 1030,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 209,
+ 164,
+ 108,
+ 163
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1289,
+ "image_id": 1030,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 162,
+ 218,
+ 57,
+ 73
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1290,
+ "image_id": 1031,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 27,
+ 100,
+ 469,
+ 394
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1291,
+ "image_id": 1032,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 71,
+ 128,
+ 282,
+ 337
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1292,
+ "image_id": 1033,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 38,
+ 87,
+ 432,
+ 385
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1293,
+ "image_id": 1034,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 29,
+ 27,
+ 477,
+ 460
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1294,
+ "image_id": 1035,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 110,
+ 285,
+ 76,
+ 102
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1295,
+ "image_id": 1036,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 27,
+ 21,
+ 460,
+ 341
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1296,
+ "image_id": 1037,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 167,
+ 103,
+ 330,
+ 375
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1297,
+ "image_id": 1038,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 232,
+ 268,
+ 76,
+ 194
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1298,
+ "image_id": 1039,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 75,
+ 328,
+ 41,
+ 159
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1299,
+ "image_id": 1040,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 144,
+ 200,
+ 53,
+ 88
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1300,
+ "image_id": 1041,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 324,
+ 141,
+ 100,
+ 259
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1301,
+ "image_id": 1042,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 212,
+ 44,
+ 281,
+ 446
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1302,
+ "image_id": 1043,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 151,
+ 87,
+ 333,
+ 408
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1303,
+ "image_id": 1044,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 190,
+ 30,
+ 268,
+ 451
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1304,
+ "image_id": 1045,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 182,
+ 195,
+ 268,
+ 268
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1305,
+ "image_id": 1046,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 40,
+ 140,
+ 184,
+ 352
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1306,
+ "image_id": 1047,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 330,
+ 252,
+ 38,
+ 38
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1307,
+ "image_id": 1047,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 60,
+ 253,
+ 28,
+ 37
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1308,
+ "image_id": 1048,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 85,
+ 266,
+ 35,
+ 36
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1309,
+ "image_id": 1048,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 307,
+ 269,
+ 37,
+ 34
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1310,
+ "image_id": 1049,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 125,
+ 306,
+ 99,
+ 72
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1311,
+ "image_id": 1050,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 19,
+ 66,
+ 419,
+ 404
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1312,
+ "image_id": 1051,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 93,
+ 226,
+ 75,
+ 162
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1313,
+ "image_id": 1052,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 102,
+ 298,
+ 63,
+ 104
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1314,
+ "image_id": 1053,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 406,
+ 322,
+ 39,
+ 59
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1315,
+ "image_id": 1054,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 127,
+ 192,
+ 83,
+ 307
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1316,
+ "image_id": 1055,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 23,
+ 17,
+ 390,
+ 458
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1317,
+ "image_id": 1056,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 174,
+ 256,
+ 45,
+ 115
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1318,
+ "image_id": 1057,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 231,
+ 352,
+ 28,
+ 138
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1319,
+ "image_id": 1057,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 273,
+ 371,
+ 22,
+ 99
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1320,
+ "image_id": 1057,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 320,
+ 362,
+ 43,
+ 120
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1321,
+ "image_id": 1058,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 158,
+ 39,
+ 60,
+ 243
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1322,
+ "image_id": 1059,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 95,
+ 52,
+ 393,
+ 444
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1323,
+ "image_id": 1060,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 76,
+ 150,
+ 80,
+ 214
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1324,
+ "image_id": 1060,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 276,
+ 125,
+ 184,
+ 312
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1325,
+ "image_id": 1061,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 180,
+ 57,
+ 220,
+ 312
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1326,
+ "image_id": 1062,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 254,
+ 55,
+ 138,
+ 426
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1327,
+ "image_id": 1063,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 183,
+ 39,
+ 143,
+ 436
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1328,
+ "image_id": 1064,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 199,
+ 101,
+ 261,
+ 344
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1329,
+ "image_id": 1064,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 108,
+ 88,
+ 122,
+ 300
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1330,
+ "image_id": 1065,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 290,
+ 306,
+ 136,
+ 195
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1331,
+ "image_id": 1066,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 99,
+ 221,
+ 132,
+ 278
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1332,
+ "image_id": 1067,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 276,
+ 145,
+ 131,
+ 348
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1333,
+ "image_id": 1068,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 114,
+ 152,
+ 133,
+ 294
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1334,
+ "image_id": 1069,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 85,
+ 228,
+ 89,
+ 222
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1335,
+ "image_id": 1070,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 258,
+ 81,
+ 170,
+ 412
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1336,
+ "image_id": 1071,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 159,
+ 72,
+ 292,
+ 374
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1337,
+ "image_id": 1072,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 107,
+ 20,
+ 260,
+ 484
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1338,
+ "image_id": 1073,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 212,
+ 62,
+ 143,
+ 448
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1339,
+ "image_id": 1074,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 87,
+ 176,
+ 291,
+ 253
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1340,
+ "image_id": 1075,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 263,
+ 107,
+ 172,
+ 396
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1341,
+ "image_id": 1076,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 64,
+ 184,
+ 142,
+ 206
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1342,
+ "image_id": 1076,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 296,
+ 268,
+ 153,
+ 153
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1343,
+ "image_id": 1077,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 238,
+ 45,
+ 157,
+ 450
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1344,
+ "image_id": 1078,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 389,
+ 338,
+ 84,
+ 128
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1345,
+ "image_id": 1079,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 389,
+ 362,
+ 92,
+ 125
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1346,
+ "image_id": 1080,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 164,
+ 88,
+ 327,
+ 399
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1347,
+ "image_id": 1081,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 94,
+ 352,
+ 114,
+ 157
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1348,
+ "image_id": 1081,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 229,
+ 115,
+ 47,
+ 166
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1349,
+ "image_id": 1081,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 351,
+ 180,
+ 66,
+ 172
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1350,
+ "image_id": 1082,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 42,
+ 78,
+ 241,
+ 430
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1351,
+ "image_id": 1083,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 133,
+ 186,
+ 83,
+ 103
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1352,
+ "image_id": 1083,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 309,
+ 188,
+ 106,
+ 96
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1353,
+ "image_id": 1084,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 225,
+ 241,
+ 15,
+ 48
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1354,
+ "image_id": 1084,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 233,
+ 279,
+ 22,
+ 80
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1355,
+ "image_id": 1084,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 296,
+ 306,
+ 26,
+ 125
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1356,
+ "image_id": 1085,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 134,
+ 267,
+ 83,
+ 99
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1357,
+ "image_id": 1085,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 298,
+ 260,
+ 91,
+ 107
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1358,
+ "image_id": 1086,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 178,
+ 281,
+ 130,
+ 115
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1359,
+ "image_id": 1087,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 132,
+ 225,
+ 273,
+ 190
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1360,
+ "image_id": 1088,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 140,
+ 206,
+ 267,
+ 210
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1361,
+ "image_id": 1089,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 237,
+ 350,
+ 29,
+ 140
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1362,
+ "image_id": 1089,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 282,
+ 358,
+ 35,
+ 134
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1363,
+ "image_id": 1090,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 119,
+ 10,
+ 157,
+ 491
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1364,
+ "image_id": 1091,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 226,
+ 134,
+ 23,
+ 61
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1365,
+ "image_id": 1092,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 311,
+ 253,
+ 47,
+ 169
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1366,
+ "image_id": 1093,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 188,
+ 341,
+ 70,
+ 104
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1367,
+ "image_id": 1094,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 76,
+ 282,
+ 52,
+ 137
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1368,
+ "image_id": 1095,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 102,
+ 227,
+ 58,
+ 172
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1369,
+ "image_id": 1096,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 273,
+ 207,
+ 92,
+ 148
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1370,
+ "image_id": 1097,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 286,
+ 85,
+ 160,
+ 409
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1371,
+ "image_id": 1098,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 232,
+ 73,
+ 118,
+ 243
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1372,
+ "image_id": 1099,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 269,
+ 209,
+ 36,
+ 153
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1373,
+ "image_id": 1100,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 227,
+ 363,
+ 55,
+ 75
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1374,
+ "image_id": 1101,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 353,
+ 261,
+ 108,
+ 218
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1375,
+ "image_id": 1102,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 307,
+ 160,
+ 38,
+ 91
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1376,
+ "image_id": 1103,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 233,
+ 244,
+ 84,
+ 218
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1377,
+ "image_id": 1104,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 292,
+ 221,
+ 134,
+ 276
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1378,
+ "image_id": 1105,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 282,
+ 130,
+ 138,
+ 359
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1379,
+ "image_id": 1106,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 229,
+ 198,
+ 42,
+ 138
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1380,
+ "image_id": 1107,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 223,
+ 105,
+ 56,
+ 228
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1381,
+ "image_id": 1108,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 110,
+ 240,
+ 194,
+ 192
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1382,
+ "image_id": 1109,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 90,
+ 193,
+ 212,
+ 202
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1383,
+ "image_id": 1110,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 161,
+ 134,
+ 112,
+ 350
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1384,
+ "image_id": 1111,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 79,
+ 89,
+ 363,
+ 358
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1385,
+ "image_id": 1112,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 38,
+ 369,
+ 171,
+ 133
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1386,
+ "image_id": 1113,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 50,
+ 375,
+ 110,
+ 65
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1387,
+ "image_id": 1114,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 178,
+ 84,
+ 121,
+ 238
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1388,
+ "image_id": 1115,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 230,
+ 148,
+ 131,
+ 343
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1389,
+ "image_id": 1116,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 156,
+ 362,
+ 44,
+ 52
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1390,
+ "image_id": 1117,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 131,
+ 412,
+ 74,
+ 76
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1391,
+ "image_id": 1118,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 152,
+ 319,
+ 130,
+ 74
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1392,
+ "image_id": 1119,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 102,
+ 26,
+ 335,
+ 421
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1393,
+ "image_id": 1120,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 209,
+ 212,
+ 119,
+ 130
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1394,
+ "image_id": 1121,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 196,
+ 15,
+ 234,
+ 439
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1395,
+ "image_id": 1122,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 238,
+ 194,
+ 120,
+ 220
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1396,
+ "image_id": 1123,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 140,
+ 116,
+ 49,
+ 103
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1397,
+ "image_id": 1124,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 353,
+ 295,
+ 32,
+ 94
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1398,
+ "image_id": 1125,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 144,
+ 294,
+ 61,
+ 83
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1399,
+ "image_id": 1126,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 333,
+ 72,
+ 160,
+ 436
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1400,
+ "image_id": 1127,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 64,
+ 15,
+ 440,
+ 487
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1401,
+ "image_id": 1128,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 44,
+ 238,
+ 299,
+ 267
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1402,
+ "image_id": 1128,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 218,
+ 81,
+ 232,
+ 224
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1403,
+ "image_id": 1129,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 141,
+ 129,
+ 201,
+ 256
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1404,
+ "image_id": 1130,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 14,
+ 76,
+ 480,
+ 392
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1405,
+ "image_id": 1130,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 230,
+ 29,
+ 207,
+ 268
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1406,
+ "image_id": 1131,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 105,
+ 224,
+ 119,
+ 171
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1407,
+ "image_id": 1131,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 199,
+ 264,
+ 98,
+ 130
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1408,
+ "image_id": 1131,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 272,
+ 265,
+ 133,
+ 142
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1409,
+ "image_id": 1131,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 365,
+ 259,
+ 101,
+ 132
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1410,
+ "image_id": 1132,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 80,
+ 181,
+ 354,
+ 323
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1411,
+ "image_id": 1133,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 197,
+ 232,
+ 269,
+ 164
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1412,
+ "image_id": 1134,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 340,
+ 13,
+ 113,
+ 152
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1413,
+ "image_id": 1134,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 184,
+ 61,
+ 114,
+ 177
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1414,
+ "image_id": 1134,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 264,
+ 48,
+ 100,
+ 95
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1415,
+ "image_id": 1134,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 276,
+ 47,
+ 134,
+ 250
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1416,
+ "image_id": 1134,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 12,
+ 180,
+ 493,
+ 325
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1417,
+ "image_id": 1135,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 126,
+ 228,
+ 148,
+ 236
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1418,
+ "image_id": 1136,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 203,
+ 269,
+ 121,
+ 209
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1419,
+ "image_id": 1136,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 328,
+ 59,
+ 168
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1420,
+ "image_id": 1136,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 414,
+ 248,
+ 19,
+ 46
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1421,
+ "image_id": 1137,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 81,
+ 503,
+ 413
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1422,
+ "image_id": 1138,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 388,
+ 230,
+ 92,
+ 63
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1423,
+ "image_id": 1138,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 328,
+ 232,
+ 77,
+ 79
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1424,
+ "image_id": 1138,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 282,
+ 230,
+ 64,
+ 91
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1425,
+ "image_id": 1138,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 188,
+ 247,
+ 111,
+ 91
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1426,
+ "image_id": 1138,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 62,
+ 272,
+ 101,
+ 89
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1427,
+ "image_id": 1139,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 12,
+ 315,
+ 153,
+ 88
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1428,
+ "image_id": 1140,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 108,
+ 208,
+ 275,
+ 288
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1429,
+ "image_id": 1141,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 215,
+ 205,
+ 232,
+ 76
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1430,
+ "image_id": 1142,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 73,
+ 201,
+ 192,
+ 193
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1431,
+ "image_id": 1143,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 88,
+ 161,
+ 134,
+ 226
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1432,
+ "image_id": 1143,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 200,
+ 175,
+ 239,
+ 262
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1433,
+ "image_id": 1143,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 375,
+ 160,
+ 94,
+ 125
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1434,
+ "image_id": 1143,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 414,
+ 143,
+ 50,
+ 71
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1435,
+ "image_id": 1143,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 203,
+ 133,
+ 272
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1436,
+ "image_id": 1144,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 415,
+ 251,
+ 94,
+ 119
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1437,
+ "image_id": 1144,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 263,
+ 256,
+ 240,
+ 251
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1438,
+ "image_id": 1145,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 40,
+ 59,
+ 459,
+ 328
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1439,
+ "image_id": 1146,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 42,
+ 99,
+ 124,
+ 156
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1440,
+ "image_id": 1147,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 190,
+ 154,
+ 290,
+ 316
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1441,
+ "image_id": 1147,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 120,
+ 131,
+ 296,
+ 377
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1442,
+ "image_id": 1148,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 268,
+ 235,
+ 37,
+ 45
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1443,
+ "image_id": 1148,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 297,
+ 228,
+ 36,
+ 48
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1444,
+ "image_id": 1148,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 404,
+ 224,
+ 16,
+ 38
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1445,
+ "image_id": 1148,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 435,
+ 225,
+ 16,
+ 45
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1446,
+ "image_id": 1148,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 376,
+ 231,
+ 31,
+ 85
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1447,
+ "image_id": 1148,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 495,
+ 225,
+ 16,
+ 34
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1448,
+ "image_id": 1149,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 20,
+ 103,
+ 468,
+ 375
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1449,
+ "image_id": 1150,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 439,
+ 58,
+ 44
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1450,
+ "image_id": 1150,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 52,
+ 439,
+ 69,
+ 45
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1451,
+ "image_id": 1150,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 116,
+ 438,
+ 72,
+ 46
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1452,
+ "image_id": 1150,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 190,
+ 441,
+ 65,
+ 41
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1453,
+ "image_id": 1150,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 230,
+ 433,
+ 57,
+ 46
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1454,
+ "image_id": 1151,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 56,
+ 350,
+ 155,
+ 107
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1455,
+ "image_id": 1151,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 212,
+ 351,
+ 145,
+ 86
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1456,
+ "image_id": 1152,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 180,
+ 199,
+ 184,
+ 308
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1457,
+ "image_id": 1153,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 353,
+ 237,
+ 48,
+ 118
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1458,
+ "image_id": 1153,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 217,
+ 254,
+ 61,
+ 145
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1459,
+ "image_id": 1153,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 250,
+ 298,
+ 116,
+ 208
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1460,
+ "image_id": 1153,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 12,
+ 273,
+ 181,
+ 225
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1461,
+ "image_id": 1153,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 419,
+ 226,
+ 90,
+ 139
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1462,
+ "image_id": 1154,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 74,
+ 243,
+ 56,
+ 138
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1463,
+ "image_id": 1154,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 273,
+ 88,
+ 225
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1464,
+ "image_id": 1154,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 140,
+ 303,
+ 245,
+ 206
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1465,
+ "image_id": 1154,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 348,
+ 254,
+ 105,
+ 138
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1466,
+ "image_id": 1154,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 464,
+ 278,
+ 44,
+ 207
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1467,
+ "image_id": 1155,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 227,
+ 110,
+ 19,
+ 52
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1468,
+ "image_id": 1156,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 293,
+ 346,
+ 66,
+ 58
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1469,
+ "image_id": 1157,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 269,
+ 227,
+ 83,
+ 203
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1470,
+ "image_id": 1158,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 148,
+ 216,
+ 148,
+ 134
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1471,
+ "image_id": 1158,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 221,
+ 166,
+ 157
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1472,
+ "image_id": 1159,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 108,
+ 251,
+ 172,
+ 189
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1473,
+ "image_id": 1160,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 43,
+ 174,
+ 31,
+ 75
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1474,
+ "image_id": 1160,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 187,
+ 40,
+ 75
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1475,
+ "image_id": 1160,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 85,
+ 198,
+ 64,
+ 108
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1476,
+ "image_id": 1160,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 65,
+ 184,
+ 32,
+ 88
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1477,
+ "image_id": 1160,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 128,
+ 159,
+ 69,
+ 103
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1478,
+ "image_id": 1160,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 189,
+ 216,
+ 92,
+ 139
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1479,
+ "image_id": 1160,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 329,
+ 205,
+ 84,
+ 136
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1480,
+ "image_id": 1160,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 417,
+ 259,
+ 84,
+ 139
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1481,
+ "image_id": 1160,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 225,
+ 279,
+ 194,
+ 199
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1482,
+ "image_id": 1161,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 163,
+ 316,
+ 71,
+ 74
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1483,
+ "image_id": 1161,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 73,
+ 304,
+ 85,
+ 85
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1484,
+ "image_id": 1162,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 8,
+ 11,
+ 439,
+ 494
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1485,
+ "image_id": 1163,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 272,
+ 401,
+ 31,
+ 80
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1486,
+ "image_id": 1163,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 400,
+ 243,
+ 107,
+ 164
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1487,
+ "image_id": 1164,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 256,
+ 333,
+ 41,
+ 97
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1488,
+ "image_id": 1164,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 359,
+ 330,
+ 50,
+ 104
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1489,
+ "image_id": 1165,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 58,
+ 255,
+ 320,
+ 245
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1490,
+ "image_id": 1165,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 236,
+ 247,
+ 272,
+ 247
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1491,
+ "image_id": 1166,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 107,
+ 272,
+ 292,
+ 232
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1492,
+ "image_id": 1166,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 293,
+ 57,
+ 163
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1493,
+ "image_id": 1166,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 26,
+ 278,
+ 86,
+ 136
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1494,
+ "image_id": 1167,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 39,
+ 162,
+ 352,
+ 247
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1495,
+ "image_id": 1168,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 8,
+ 148,
+ 494,
+ 263
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1496,
+ "image_id": 1169,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 38,
+ 102,
+ 394,
+ 331
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1497,
+ "image_id": 1170,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 310,
+ 238,
+ 104,
+ 99
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1498,
+ "image_id": 1171,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 42,
+ 117,
+ 73,
+ 86
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1499,
+ "image_id": 1171,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 63,
+ 116,
+ 106,
+ 140
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1500,
+ "image_id": 1171,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 111,
+ 107,
+ 112,
+ 189
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1501,
+ "image_id": 1171,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 174,
+ 146,
+ 125,
+ 165
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1502,
+ "image_id": 1171,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 269,
+ 152,
+ 140,
+ 212
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1503,
+ "image_id": 1171,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 380,
+ 139,
+ 127,
+ 354
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1504,
+ "image_id": 1172,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 200,
+ 338,
+ 10,
+ 23
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1505,
+ "image_id": 1172,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 225,
+ 334,
+ 18,
+ 48
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1506,
+ "image_id": 1172,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 387,
+ 330,
+ 105,
+ 141
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1507,
+ "image_id": 1173,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 122,
+ 22,
+ 382,
+ 442
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1508,
+ "image_id": 1173,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 32,
+ 132,
+ 74,
+ 66
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1509,
+ "image_id": 1174,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 16,
+ 70,
+ 473,
+ 417
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1510,
+ "image_id": 1175,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 50,
+ 242,
+ 45,
+ 129
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1511,
+ "image_id": 1175,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 105,
+ 254,
+ 73,
+ 192
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1512,
+ "image_id": 1175,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 184,
+ 243,
+ 75,
+ 151
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1513,
+ "image_id": 1176,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 183,
+ 230,
+ 248,
+ 210
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1514,
+ "image_id": 1177,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 120,
+ 38,
+ 349,
+ 432
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1515,
+ "image_id": 1178,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 191,
+ 279,
+ 248,
+ 152
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1516,
+ "image_id": 1178,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 331,
+ 269,
+ 141,
+ 45
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1517,
+ "image_id": 1179,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 182,
+ 214,
+ 40,
+ 89
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1518,
+ "image_id": 1180,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 245,
+ 247,
+ 97,
+ 156
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1519,
+ "image_id": 1181,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 202,
+ 245,
+ 87,
+ 84
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1520,
+ "image_id": 1182,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 49,
+ 125,
+ 248,
+ 206
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1521,
+ "image_id": 1183,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 160,
+ 192,
+ 87,
+ 142
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1522,
+ "image_id": 1184,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 173,
+ 236,
+ 82,
+ 117
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1523,
+ "image_id": 1185,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 216,
+ 94,
+ 147
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1524,
+ "image_id": 1185,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 249,
+ 182,
+ 91,
+ 175
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1525,
+ "image_id": 1186,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 263,
+ 185,
+ 36,
+ 75
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1526,
+ "image_id": 1186,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 230,
+ 192,
+ 36,
+ 66
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1527,
+ "image_id": 1186,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 180,
+ 188,
+ 34,
+ 72
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1528,
+ "image_id": 1187,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 234,
+ 231,
+ 34,
+ 66
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1529,
+ "image_id": 1187,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 213,
+ 236,
+ 15,
+ 43
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1530,
+ "image_id": 1188,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 21,
+ 32,
+ 58,
+ 58
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1531,
+ "image_id": 1188,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 114,
+ 67,
+ 231,
+ 247
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1532,
+ "image_id": 1189,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 301,
+ 12,
+ 96,
+ 117
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1533,
+ "image_id": 1189,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 81,
+ 19,
+ 285,
+ 473
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1534,
+ "image_id": 1190,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 10,
+ 411,
+ 490
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1535,
+ "image_id": 1191,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 208,
+ 61,
+ 177,
+ 275
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1536,
+ "image_id": 1192,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 195,
+ 191,
+ 15,
+ 33
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1537,
+ "image_id": 1193,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 235,
+ 235,
+ 63,
+ 126
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1538,
+ "image_id": 1194,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 216,
+ 152,
+ 51,
+ 96
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1539,
+ "image_id": 1195,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 65,
+ 180,
+ 74,
+ 117
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1540,
+ "image_id": 1196,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 293,
+ 266,
+ 40,
+ 71
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1541,
+ "image_id": 1196,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 344,
+ 256,
+ 41,
+ 76
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1542,
+ "image_id": 1196,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 140,
+ 258,
+ 19,
+ 35
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1543,
+ "image_id": 1196,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 177,
+ 262,
+ 21,
+ 27
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1544,
+ "image_id": 1197,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 213,
+ 241,
+ 43,
+ 60
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1545,
+ "image_id": 1197,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 258,
+ 235,
+ 22,
+ 43
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1546,
+ "image_id": 1198,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 113,
+ 210,
+ 322,
+ 232
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1547,
+ "image_id": 1199,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 179,
+ 202,
+ 127,
+ 105
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1548,
+ "image_id": 1200,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 189,
+ 205,
+ 125,
+ 104
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1549,
+ "image_id": 1201,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 25,
+ 211,
+ 93,
+ 99
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1550,
+ "image_id": 1201,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 90,
+ 200,
+ 113,
+ 128
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1551,
+ "image_id": 1201,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 138,
+ 211,
+ 232,
+ 138
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1552,
+ "image_id": 1202,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 194,
+ 164,
+ 39,
+ 61
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1553,
+ "image_id": 1203,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 134,
+ 191,
+ 170,
+ 145
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1554,
+ "image_id": 1203,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 128,
+ 148,
+ 100,
+ 93
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1555,
+ "image_id": 1204,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 162,
+ 160,
+ 179,
+ 253
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1556,
+ "image_id": 1205,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 208,
+ 232,
+ 158,
+ 154
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1557,
+ "image_id": 1206,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 88,
+ 140,
+ 370,
+ 368
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1558,
+ "image_id": 1207,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 312,
+ 214,
+ 195
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1559,
+ "image_id": 1208,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 185,
+ 112,
+ 217,
+ 395
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1560,
+ "image_id": 1209,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 104,
+ 24,
+ 371,
+ 483
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1561,
+ "image_id": 1210,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 344,
+ 136,
+ 133,
+ 318
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1562,
+ "image_id": 1211,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 69,
+ 297,
+ 139,
+ 181
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1563,
+ "image_id": 1212,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 108,
+ 164,
+ 55,
+ 56
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1564,
+ "image_id": 1212,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 271,
+ 188,
+ 172,
+ 108
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1565,
+ "image_id": 1213,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 6,
+ 483,
+ 500
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1566,
+ "image_id": 1214,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 268,
+ 149,
+ 93,
+ 134
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1567,
+ "image_id": 1214,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 189,
+ 128,
+ 99,
+ 148
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1568,
+ "image_id": 1214,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 39,
+ 103,
+ 174,
+ 178
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1569,
+ "image_id": 1215,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 132,
+ 261,
+ 83,
+ 136
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1570,
+ "image_id": 1216,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 184,
+ 160,
+ 324,
+ 317
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1571,
+ "image_id": 1217,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 205,
+ 251,
+ 101,
+ 188
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1572,
+ "image_id": 1217,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 84,
+ 290,
+ 57,
+ 126
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1573,
+ "image_id": 1218,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 62,
+ 134,
+ 390,
+ 322
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1574,
+ "image_id": 1219,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 19,
+ 58,
+ 478,
+ 408
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1575,
+ "image_id": 1220,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 130,
+ 142,
+ 187,
+ 275
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1576,
+ "image_id": 1221,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 86,
+ 283,
+ 72,
+ 171
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1577,
+ "image_id": 1221,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 184,
+ 267,
+ 80,
+ 202
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1578,
+ "image_id": 1221,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 304,
+ 260,
+ 101,
+ 244
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1579,
+ "image_id": 1222,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 283,
+ 31,
+ 216,
+ 465
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1580,
+ "image_id": 1223,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 264,
+ 229,
+ 243,
+ 82
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1581,
+ "image_id": 1223,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 47,
+ 268,
+ 182,
+ 193
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1582,
+ "image_id": 1224,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 99,
+ 209,
+ 396,
+ 292
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1583,
+ "image_id": 1225,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 25,
+ 248,
+ 271,
+ 243
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1584,
+ "image_id": 1226,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 44,
+ 96,
+ 404,
+ 366
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1585,
+ "image_id": 1227,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 55,
+ 131,
+ 148,
+ 245
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1586,
+ "image_id": 1227,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 215,
+ 107,
+ 247,
+ 269
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1587,
+ "image_id": 1228,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 248,
+ 314,
+ 178,
+ 134
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1588,
+ "image_id": 1229,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 14,
+ 481,
+ 489
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1589,
+ "image_id": 1230,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 35,
+ 137,
+ 162,
+ 366
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1590,
+ "image_id": 1231,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 240,
+ 208,
+ 51,
+ 112
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1591,
+ "image_id": 1232,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 69,
+ 5,
+ 368,
+ 484
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1592,
+ "image_id": 1233,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 184,
+ 177,
+ 162,
+ 317
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1593,
+ "image_id": 1234,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 308,
+ 194,
+ 30,
+ 62
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1594,
+ "image_id": 1234,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 284,
+ 197,
+ 32,
+ 49
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1595,
+ "image_id": 1234,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 222,
+ 191,
+ 75,
+ 93
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1596,
+ "image_id": 1234,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 91,
+ 152,
+ 141,
+ 285
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1597,
+ "image_id": 1235,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 106,
+ 198,
+ 279,
+ 289
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1598,
+ "image_id": 1236,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 61,
+ 86,
+ 284,
+ 209
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1599,
+ "image_id": 1237,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 160,
+ 324,
+ 74,
+ 82
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1600,
+ "image_id": 1237,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 146,
+ 321,
+ 35,
+ 48
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1601,
+ "image_id": 1237,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 318,
+ 319,
+ 26,
+ 43
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1602,
+ "image_id": 1237,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 389,
+ 409,
+ 113,
+ 93
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1603,
+ "image_id": 1238,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 7,
+ 339,
+ 105,
+ 128
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1604,
+ "image_id": 1239,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 341,
+ 74,
+ 96
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1605,
+ "image_id": 1239,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 53,
+ 320,
+ 100,
+ 130
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1606,
+ "image_id": 1240,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 361,
+ 63,
+ 84
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1607,
+ "image_id": 1240,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 50,
+ 368,
+ 70,
+ 84
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1608,
+ "image_id": 1240,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 85,
+ 360,
+ 76,
+ 97
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1609,
+ "image_id": 1241,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 270,
+ 331,
+ 65,
+ 69
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1610,
+ "image_id": 1242,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 136,
+ 368,
+ 52,
+ 60
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1611,
+ "image_id": 1243,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 220,
+ 331,
+ 42,
+ 74
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1612,
+ "image_id": 1243,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 245,
+ 318,
+ 36,
+ 55
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1613,
+ "image_id": 1244,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 109,
+ 378,
+ 14,
+ 26
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1614,
+ "image_id": 1244,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 137,
+ 375,
+ 30,
+ 28
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1615,
+ "image_id": 1244,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 175,
+ 369,
+ 20,
+ 40
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1616,
+ "image_id": 1244,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 210,
+ 370,
+ 26,
+ 35
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1617,
+ "image_id": 1244,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 232,
+ 370,
+ 22,
+ 34
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1618,
+ "image_id": 1244,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 340,
+ 368,
+ 61,
+ 101
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1619,
+ "image_id": 1244,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 276,
+ 388,
+ 70,
+ 94
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1620,
+ "image_id": 1245,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 292,
+ 364,
+ 43,
+ 58
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1621,
+ "image_id": 1246,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 226,
+ 351,
+ 42,
+ 100
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1622,
+ "image_id": 1246,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 192,
+ 351,
+ 41,
+ 91
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1623,
+ "image_id": 1246,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 133,
+ 350,
+ 37,
+ 78
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1624,
+ "image_id": 1246,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 104,
+ 336,
+ 13,
+ 28
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1625,
+ "image_id": 1246,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 119,
+ 334,
+ 11,
+ 30
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1626,
+ "image_id": 1247,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 235,
+ 343,
+ 65,
+ 152
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1627,
+ "image_id": 1248,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 230,
+ 343,
+ 58,
+ 147
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1628,
+ "image_id": 1249,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 134,
+ 374,
+ 46,
+ 54
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1629,
+ "image_id": 1250,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 116,
+ 394,
+ 35,
+ 58
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1630,
+ "image_id": 1250,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 161,
+ 394,
+ 32,
+ 64
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1631,
+ "image_id": 1250,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 246,
+ 395,
+ 31,
+ 68
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1632,
+ "image_id": 1250,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 282,
+ 391,
+ 25,
+ 73
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1633,
+ "image_id": 1250,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 312,
+ 390,
+ 49,
+ 74
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1634,
+ "image_id": 1250,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 484,
+ 390,
+ 26,
+ 67
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1635,
+ "image_id": 1250,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 428,
+ 386,
+ 61,
+ 79
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1636,
+ "image_id": 1251,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 29,
+ 27,
+ 384,
+ 405
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1637,
+ "image_id": 1252,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 1,
+ 348,
+ 32,
+ 72
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1638,
+ "image_id": 1252,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 102,
+ 341,
+ 44,
+ 62
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1639,
+ "image_id": 1252,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 152,
+ 344,
+ 48,
+ 78
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1640,
+ "image_id": 1252,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 196,
+ 329,
+ 22,
+ 39
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1641,
+ "image_id": 1252,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 279,
+ 348,
+ 38,
+ 55
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1642,
+ "image_id": 1252,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 279,
+ 364,
+ 100,
+ 97
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1643,
+ "image_id": 1253,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 56,
+ 79,
+ 399,
+ 411
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1644,
+ "image_id": 1254,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 122,
+ 310,
+ 94,
+ 114
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1645,
+ "image_id": 1254,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 248,
+ 240,
+ 169,
+ 191
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1646,
+ "image_id": 1255,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 107,
+ 169,
+ 126,
+ 153
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1647,
+ "image_id": 1255,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 163,
+ 179,
+ 325,
+ 307
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1648,
+ "image_id": 1256,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 164,
+ 120,
+ 248
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1649,
+ "image_id": 1257,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 251,
+ 254,
+ 48,
+ 175
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1650,
+ "image_id": 1258,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 116,
+ 283,
+ 70,
+ 85
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1651,
+ "image_id": 1258,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 156,
+ 321,
+ 153,
+ 173
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1652,
+ "image_id": 1259,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 51,
+ 13,
+ 340,
+ 476
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1653,
+ "image_id": 1260,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 200,
+ 313,
+ 55,
+ 121
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1654,
+ "image_id": 1261,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 261,
+ 247,
+ 112,
+ 255
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1655,
+ "image_id": 1262,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 15,
+ 23,
+ 431,
+ 455
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1656,
+ "image_id": 1263,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 216,
+ 204,
+ 63,
+ 159
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1657,
+ "image_id": 1263,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 362,
+ 191,
+ 70,
+ 136
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1658,
+ "image_id": 1264,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 32,
+ 17,
+ 401,
+ 489
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1659,
+ "image_id": 1265,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 215,
+ 16,
+ 279,
+ 455
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1660,
+ "image_id": 1266,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 228,
+ 278,
+ 121,
+ 131
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1661,
+ "image_id": 1267,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 286,
+ 210,
+ 98,
+ 171
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1662,
+ "image_id": 1267,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 58,
+ 229,
+ 446
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1663,
+ "image_id": 1268,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 228,
+ 252,
+ 26,
+ 51
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1664,
+ "image_id": 1268,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 102,
+ 244,
+ 19,
+ 40
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1665,
+ "image_id": 1269,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 96,
+ 170,
+ 90,
+ 263
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1666,
+ "image_id": 1269,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 178,
+ 123,
+ 47,
+ 177
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1667,
+ "image_id": 1269,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 242,
+ 142,
+ 84,
+ 267
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1668,
+ "image_id": 1269,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 322,
+ 140,
+ 65,
+ 219
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1669,
+ "image_id": 1270,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 218,
+ 188,
+ 117,
+ 97
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1670,
+ "image_id": 1270,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 52,
+ 187,
+ 302,
+ 303
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1671,
+ "image_id": 1271,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 75,
+ 13,
+ 317,
+ 466
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1672,
+ "image_id": 1272,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 219,
+ 364,
+ 276,
+ 145
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1673,
+ "image_id": 1273,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 49,
+ 60,
+ 363,
+ 354
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1674,
+ "image_id": 1274,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 133,
+ 162,
+ 107,
+ 221
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1675,
+ "image_id": 1274,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 361,
+ 165,
+ 109,
+ 222
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1676,
+ "image_id": 1275,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 25,
+ 139,
+ 76,
+ 109
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1677,
+ "image_id": 1275,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 75,
+ 154,
+ 148,
+ 177
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1678,
+ "image_id": 1275,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 150,
+ 210,
+ 161,
+ 158
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1679,
+ "image_id": 1275,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 211,
+ 221,
+ 171,
+ 177
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1680,
+ "image_id": 1275,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 251,
+ 260,
+ 203,
+ 193
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1681,
+ "image_id": 1275,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 254,
+ 140,
+ 79,
+ 73
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1682,
+ "image_id": 1276,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 36,
+ 289,
+ 61,
+ 136
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1683,
+ "image_id": 1276,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 150,
+ 206,
+ 105,
+ 191
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1684,
+ "image_id": 1276,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 243,
+ 214,
+ 98,
+ 166
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1685,
+ "image_id": 1276,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 409,
+ 206,
+ 98,
+ 162
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1686,
+ "image_id": 1277,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 8,
+ 184,
+ 64,
+ 61
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1687,
+ "image_id": 1277,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 53,
+ 197,
+ 69,
+ 73
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1688,
+ "image_id": 1277,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 106,
+ 206,
+ 99,
+ 99
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1689,
+ "image_id": 1277,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 180,
+ 196,
+ 100,
+ 124
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1690,
+ "image_id": 1277,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 250,
+ 212,
+ 64,
+ 136
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1691,
+ "image_id": 1277,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 299,
+ 218,
+ 79,
+ 147
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1692,
+ "image_id": 1277,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 359,
+ 182,
+ 70,
+ 199
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1693,
+ "image_id": 1277,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 390,
+ 147,
+ 117,
+ 293
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1694,
+ "image_id": 1278,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 46,
+ 99,
+ 130,
+ 121
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1695,
+ "image_id": 1278,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 90,
+ 135,
+ 159,
+ 207
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1696,
+ "image_id": 1278,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 105,
+ 203,
+ 243,
+ 219
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1697,
+ "image_id": 1278,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 142,
+ 225,
+ 313,
+ 278
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1698,
+ "image_id": 1278,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 274,
+ 99,
+ 84,
+ 27
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1699,
+ "image_id": 1278,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 424,
+ 131,
+ 83,
+ 61
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1700,
+ "image_id": 1279,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 162,
+ 263,
+ 342,
+ 237
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1701,
+ "image_id": 1279,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 207,
+ 177,
+ 126,
+ 129
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1702,
+ "image_id": 1280,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 308,
+ 147,
+ 37,
+ 73
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1703,
+ "image_id": 1280,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 128,
+ 150,
+ 104,
+ 65
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1704,
+ "image_id": 1280,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 134,
+ 163,
+ 169,
+ 121
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1705,
+ "image_id": 1280,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 260,
+ 222,
+ 186,
+ 259
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1706,
+ "image_id": 1281,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 84,
+ 159,
+ 148,
+ 141
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1707,
+ "image_id": 1281,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 129,
+ 196,
+ 175,
+ 196
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1708,
+ "image_id": 1281,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 155,
+ 212,
+ 228,
+ 255
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1709,
+ "image_id": 1281,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 198,
+ 275,
+ 293,
+ 228
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1710,
+ "image_id": 1282,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 180,
+ 158,
+ 87,
+ 87
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1711,
+ "image_id": 1282,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 172,
+ 192,
+ 184,
+ 181
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1712,
+ "image_id": 1282,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 158,
+ 275,
+ 288,
+ 214
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1713,
+ "image_id": 1282,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 191,
+ 299,
+ 316,
+ 211
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1714,
+ "image_id": 1283,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 158,
+ 176,
+ 107,
+ 105
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1715,
+ "image_id": 1283,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 262,
+ 187,
+ 167,
+ 105
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1716,
+ "image_id": 1283,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 257,
+ 253,
+ 237,
+ 167
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1717,
+ "image_id": 1283,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 210,
+ 266,
+ 286,
+ 243
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1718,
+ "image_id": 1284,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 14,
+ 495,
+ 456
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1719,
+ "image_id": 1285,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 251,
+ 161,
+ 90,
+ 76
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1720,
+ "image_id": 1285,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 443,
+ 180,
+ 68,
+ 75
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1721,
+ "image_id": 1285,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 19,
+ 159,
+ 155,
+ 118
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1722,
+ "image_id": 1285,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 75,
+ 189,
+ 181,
+ 163
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1723,
+ "image_id": 1285,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 145,
+ 210,
+ 175,
+ 178
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1724,
+ "image_id": 1285,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 217,
+ 226,
+ 167,
+ 203
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1725,
+ "image_id": 1285,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 258,
+ 267,
+ 201,
+ 215
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1726,
+ "image_id": 1286,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 163,
+ 224,
+ 199,
+ 97
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1727,
+ "image_id": 1287,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 7,
+ 296,
+ 110,
+ 203
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1728,
+ "image_id": 1287,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 19,
+ 249,
+ 89,
+ 118
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1729,
+ "image_id": 1287,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 212,
+ 39,
+ 107
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1730,
+ "image_id": 1287,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 178,
+ 214,
+ 89,
+ 95
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1731,
+ "image_id": 1287,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 270,
+ 228,
+ 59,
+ 118
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1732,
+ "image_id": 1287,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 358,
+ 210,
+ 69,
+ 77
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1733,
+ "image_id": 1288,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 267,
+ 167,
+ 84,
+ 192
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1734,
+ "image_id": 1288,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 292,
+ 237,
+ 152,
+ 262
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1735,
+ "image_id": 1288,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 7,
+ 252,
+ 58,
+ 232
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1736,
+ "image_id": 1288,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 82,
+ 218,
+ 129,
+ 212
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1737,
+ "image_id": 1288,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 181,
+ 159,
+ 59,
+ 98
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1738,
+ "image_id": 1289,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 241,
+ 176,
+ 153,
+ 140
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1739,
+ "image_id": 1289,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 49,
+ 170,
+ 436,
+ 338
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1740,
+ "image_id": 1290,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 110,
+ 68,
+ 373,
+ 428
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1741,
+ "image_id": 1291,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 77,
+ 154,
+ 72,
+ 129
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1742,
+ "image_id": 1291,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 18,
+ 170,
+ 55,
+ 169
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1743,
+ "image_id": 1291,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 201,
+ 117,
+ 135,
+ 139
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1744,
+ "image_id": 1291,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 290,
+ 177,
+ 155,
+ 158
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1745,
+ "image_id": 1291,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 101,
+ 184,
+ 162,
+ 203
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1746,
+ "image_id": 1291,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 40,
+ 279,
+ 263,
+ 230
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1747,
+ "image_id": 1291,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 347,
+ 206,
+ 158,
+ 256
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1748,
+ "image_id": 1292,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 58,
+ 32,
+ 283,
+ 460
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1749,
+ "image_id": 1293,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 216,
+ 176,
+ 165,
+ 292
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1750,
+ "image_id": 1294,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 20,
+ 287,
+ 483
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1751,
+ "image_id": 1294,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 211,
+ 77,
+ 297,
+ 417
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1752,
+ "image_id": 1295,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 143,
+ 279,
+ 180,
+ 177
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1753,
+ "image_id": 1296,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 142,
+ 200,
+ 271,
+ 269
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1754,
+ "image_id": 1297,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 367,
+ 331,
+ 56,
+ 106
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1755,
+ "image_id": 1298,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 56,
+ 295,
+ 260,
+ 214
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1756,
+ "image_id": 1299,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 203,
+ 325,
+ 117,
+ 112
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1757,
+ "image_id": 1300,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 41,
+ 281,
+ 214,
+ 206
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1758,
+ "image_id": 1301,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 17,
+ 156,
+ 137,
+ 242
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1759,
+ "image_id": 1302,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 39,
+ 83,
+ 307,
+ 330
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1760,
+ "image_id": 1302,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 101,
+ 56,
+ 396,
+ 397
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1761,
+ "image_id": 1303,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 104,
+ 73,
+ 277,
+ 318
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1762,
+ "image_id": 1303,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 217,
+ 98,
+ 164,
+ 385
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1763,
+ "image_id": 1304,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 337,
+ 48,
+ 116
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1764,
+ "image_id": 1304,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 35,
+ 324,
+ 69,
+ 139
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1765,
+ "image_id": 1304,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 79,
+ 325,
+ 86,
+ 146
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1766,
+ "image_id": 1304,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 168,
+ 344,
+ 78,
+ 131
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1767,
+ "image_id": 1304,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 303,
+ 323,
+ 37,
+ 34
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1768,
+ "image_id": 1304,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 340,
+ 310,
+ 20,
+ 43
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1769,
+ "image_id": 1304,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 241,
+ 342,
+ 94,
+ 148
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1770,
+ "image_id": 1304,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 348,
+ 316,
+ 101,
+ 182
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1771,
+ "image_id": 1304,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 458,
+ 323,
+ 52,
+ 40
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1772,
+ "image_id": 1305,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 188,
+ 298,
+ 117,
+ 194
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1773,
+ "image_id": 1306,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 130,
+ 174,
+ 153,
+ 202
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1774,
+ "image_id": 1307,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 42,
+ 447,
+ 29,
+ 47
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1775,
+ "image_id": 1307,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 69,
+ 445,
+ 25,
+ 47
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1776,
+ "image_id": 1307,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 128,
+ 439,
+ 30,
+ 46
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1777,
+ "image_id": 1308,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 464,
+ 272,
+ 45,
+ 141
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1778,
+ "image_id": 1308,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 113,
+ 352,
+ 42,
+ 66
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1779,
+ "image_id": 1308,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 67,
+ 365,
+ 29,
+ 90
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1780,
+ "image_id": 1308,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 346,
+ 64,
+ 112
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1781,
+ "image_id": 1309,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 186,
+ 87,
+ 241,
+ 414
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1782,
+ "image_id": 1310,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 461,
+ 119,
+ 49,
+ 88
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1783,
+ "image_id": 1310,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 158,
+ 125,
+ 202,
+ 343
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1784,
+ "image_id": 1311,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 44,
+ 291,
+ 162,
+ 170
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1785,
+ "image_id": 1312,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 36,
+ 140,
+ 190,
+ 355
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1786,
+ "image_id": 1313,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 375,
+ 192,
+ 112,
+ 215
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1787,
+ "image_id": 1313,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 335,
+ 200,
+ 49,
+ 106
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1788,
+ "image_id": 1313,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 160,
+ 182,
+ 64,
+ 117
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1789,
+ "image_id": 1313,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 18,
+ 153,
+ 137,
+ 205
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1790,
+ "image_id": 1313,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 203,
+ 174,
+ 108,
+ 229
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1791,
+ "image_id": 1314,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 117,
+ 231,
+ 99,
+ 44
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1792,
+ "image_id": 1314,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 134,
+ 235,
+ 161,
+ 84
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1793,
+ "image_id": 1314,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 161,
+ 234,
+ 209,
+ 148
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1794,
+ "image_id": 1315,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 285,
+ 228,
+ 183,
+ 240
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1795,
+ "image_id": 1316,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 95,
+ 169,
+ 264,
+ 335
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1796,
+ "image_id": 1317,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 211,
+ 203,
+ 169,
+ 140
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1797,
+ "image_id": 1318,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 34,
+ 237,
+ 156,
+ 215
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1798,
+ "image_id": 1319,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 106,
+ 303,
+ 140,
+ 171
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1799,
+ "image_id": 1320,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 259,
+ 413,
+ 41,
+ 36
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1800,
+ "image_id": 1320,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 230,
+ 417,
+ 31,
+ 47
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1801,
+ "image_id": 1320,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 173,
+ 410,
+ 54,
+ 57
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1802,
+ "image_id": 1321,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 61,
+ 60,
+ 397,
+ 414
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1803,
+ "image_id": 1322,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 142,
+ 235,
+ 83,
+ 232
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1804,
+ "image_id": 1323,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 188,
+ 120,
+ 142,
+ 385
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1805,
+ "image_id": 1324,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 195,
+ 354,
+ 122,
+ 109
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1806,
+ "image_id": 1325,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 29,
+ 261,
+ 155,
+ 164
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1807,
+ "image_id": 1325,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 167,
+ 252,
+ 178,
+ 162
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1808,
+ "image_id": 1325,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 321,
+ 243,
+ 183,
+ 165
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1809,
+ "image_id": 1326,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 25,
+ 70,
+ 274,
+ 328
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1810,
+ "image_id": 1327,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 262,
+ 152,
+ 183,
+ 324
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1811,
+ "image_id": 1328,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 57,
+ 225,
+ 107,
+ 275
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1812,
+ "image_id": 1329,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 93,
+ 51,
+ 325,
+ 422
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1813,
+ "image_id": 1330,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 86,
+ 274,
+ 87,
+ 144
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1814,
+ "image_id": 1330,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 253,
+ 178,
+ 244
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1815,
+ "image_id": 1330,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 149,
+ 267,
+ 101,
+ 186
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1816,
+ "image_id": 1330,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 213,
+ 257,
+ 122,
+ 227
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1817,
+ "image_id": 1330,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 303,
+ 216,
+ 136,
+ 287
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1818,
+ "image_id": 1330,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 424,
+ 225,
+ 81,
+ 247
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1819,
+ "image_id": 1330,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 236,
+ 168,
+ 37,
+ 88
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1820,
+ "image_id": 1331,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 125,
+ 126,
+ 381
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1821,
+ "image_id": 1331,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 108,
+ 187,
+ 172,
+ 317
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1822,
+ "image_id": 1331,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 372,
+ 152,
+ 98,
+ 355
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1823,
+ "image_id": 1331,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 275,
+ 196,
+ 163,
+ 313
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1824,
+ "image_id": 1332,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 104,
+ 12,
+ 337,
+ 491
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1825,
+ "image_id": 1333,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 266,
+ 243,
+ 159,
+ 241
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1826,
+ "image_id": 1334,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 253,
+ 135,
+ 19,
+ 74
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1827,
+ "image_id": 1334,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 292,
+ 116,
+ 63,
+ 72
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1828,
+ "image_id": 1334,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 283,
+ 156,
+ 56,
+ 60
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1829,
+ "image_id": 1334,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 253,
+ 211,
+ 84,
+ 72
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1830,
+ "image_id": 1334,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 234,
+ 253,
+ 87,
+ 90
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1831,
+ "image_id": 1334,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 217,
+ 315,
+ 100,
+ 89
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1832,
+ "image_id": 1334,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 66,
+ 285,
+ 70,
+ 104
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1833,
+ "image_id": 1334,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 162,
+ 164,
+ 57,
+ 128
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1834,
+ "image_id": 1334,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 221,
+ 107,
+ 36,
+ 78
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1835,
+ "image_id": 1335,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 305,
+ 11,
+ 68,
+ 86
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1836,
+ "image_id": 1335,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 290,
+ 55,
+ 67,
+ 64
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1837,
+ "image_id": 1335,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 265,
+ 110,
+ 86,
+ 77
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1838,
+ "image_id": 1335,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 247,
+ 166,
+ 87,
+ 81
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1839,
+ "image_id": 1335,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 227,
+ 216,
+ 100,
+ 90
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1840,
+ "image_id": 1335,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 75,
+ 288,
+ 96,
+ 136
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1841,
+ "image_id": 1335,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 233,
+ 11,
+ 35,
+ 73
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1842,
+ "image_id": 1335,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 177,
+ 69,
+ 56,
+ 120
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1843,
+ "image_id": 1335,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 88,
+ 181,
+ 65,
+ 99
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1844,
+ "image_id": 1336,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 268,
+ 268,
+ 121,
+ 176
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1845,
+ "image_id": 1336,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 369,
+ 260,
+ 94,
+ 171
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1846,
+ "image_id": 1336,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 456,
+ 260,
+ 52,
+ 141
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1847,
+ "image_id": 1337,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 229,
+ 177,
+ 99,
+ 84
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1848,
+ "image_id": 1337,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 249,
+ 131,
+ 82,
+ 79
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1849,
+ "image_id": 1337,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 271,
+ 72,
+ 76,
+ 80
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1850,
+ "image_id": 1337,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 292,
+ 23,
+ 60,
+ 57
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1851,
+ "image_id": 1337,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 309,
+ 6,
+ 58,
+ 45
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1852,
+ "image_id": 1337,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 74,
+ 253,
+ 108,
+ 124
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1853,
+ "image_id": 1337,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 82,
+ 143,
+ 66,
+ 102
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1854,
+ "image_id": 1337,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 173,
+ 30,
+ 56,
+ 110
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1855,
+ "image_id": 1338,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 306,
+ 224,
+ 52,
+ 107
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1856,
+ "image_id": 1339,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 151,
+ 163,
+ 268,
+ 264
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1857,
+ "image_id": 1339,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 159,
+ 66,
+ 105,
+ 241
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1858,
+ "image_id": 1340,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 92,
+ 208,
+ 300,
+ 298
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1859,
+ "image_id": 1341,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 82,
+ 175,
+ 239,
+ 309
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1860,
+ "image_id": 1342,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 220,
+ 338,
+ 58,
+ 169
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1861,
+ "image_id": 1343,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 150,
+ 246,
+ 76,
+ 99
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1862,
+ "image_id": 1343,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 279,
+ 246,
+ 67,
+ 97
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1863,
+ "image_id": 1344,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 217,
+ 324,
+ 34,
+ 149
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1864,
+ "image_id": 1344,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 256,
+ 390,
+ 44,
+ 115
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1865,
+ "image_id": 1344,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 246,
+ 266,
+ 26,
+ 71
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1866,
+ "image_id": 1345,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 151,
+ 252,
+ 110,
+ 178
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1867,
+ "image_id": 1346,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 112,
+ 221,
+ 208,
+ 282
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1868,
+ "image_id": 1347,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 241,
+ 240,
+ 106,
+ 151
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1869,
+ "image_id": 1348,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 225,
+ 320,
+ 228,
+ 182
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1870,
+ "image_id": 1349,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 22,
+ 194,
+ 487
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1871,
+ "image_id": 1349,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 186,
+ 115,
+ 213,
+ 288
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1872,
+ "image_id": 1350,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 111,
+ 268,
+ 242,
+ 238
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1873,
+ "image_id": 1351,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 116,
+ 148,
+ 200,
+ 321
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1874,
+ "image_id": 1352,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 28,
+ 39,
+ 124,
+ 237
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1875,
+ "image_id": 1352,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 104,
+ 148,
+ 148,
+ 223
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1876,
+ "image_id": 1353,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 349,
+ 150,
+ 155,
+ 280
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1877,
+ "image_id": 1353,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 169,
+ 116,
+ 182
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1878,
+ "image_id": 1354,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 279,
+ 231,
+ 104,
+ 165
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1879,
+ "image_id": 1354,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 151,
+ 62,
+ 68,
+ 117
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1880,
+ "image_id": 1355,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 206,
+ 176,
+ 71,
+ 110
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1881,
+ "image_id": 1356,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 16,
+ 165,
+ 157,
+ 181
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1882,
+ "image_id": 1356,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 86,
+ 114,
+ 79,
+ 86
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1883,
+ "image_id": 1356,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 234,
+ 139,
+ 40,
+ 117
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1884,
+ "image_id": 1357,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 128,
+ 163,
+ 219,
+ 331
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1885,
+ "image_id": 1357,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 208,
+ 90,
+ 250
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1886,
+ "image_id": 1358,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 14,
+ 238,
+ 290,
+ 260
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1887,
+ "image_id": 1359,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 406,
+ 223,
+ 35,
+ 70
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1888,
+ "image_id": 1359,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 302,
+ 227,
+ 158,
+ 271
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1889,
+ "image_id": 1360,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 14,
+ 187,
+ 187,
+ 238
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1890,
+ "image_id": 1360,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 175,
+ 169,
+ 170,
+ 308
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1891,
+ "image_id": 1361,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 14,
+ 183,
+ 280,
+ 295
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1892,
+ "image_id": 1361,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 255,
+ 203,
+ 126,
+ 227
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1893,
+ "image_id": 1362,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 232,
+ 219,
+ 102,
+ 130
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1894,
+ "image_id": 1362,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 324,
+ 227,
+ 85,
+ 181
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1895,
+ "image_id": 1363,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 166,
+ 209,
+ 170,
+ 251
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1896,
+ "image_id": 1364,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 199,
+ 268,
+ 69,
+ 193
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1897,
+ "image_id": 1365,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 320,
+ 332,
+ 82,
+ 93
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1898,
+ "image_id": 1366,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 234,
+ 263,
+ 51,
+ 121
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1899,
+ "image_id": 1367,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 242,
+ 105,
+ 130,
+ 260
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1900,
+ "image_id": 1368,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 108,
+ 99,
+ 67,
+ 166
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1901,
+ "image_id": 1368,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 28,
+ 65,
+ 47,
+ 263
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1902,
+ "image_id": 1369,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 173,
+ 160,
+ 48,
+ 199
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1903,
+ "image_id": 1370,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 421,
+ 128,
+ 34,
+ 112
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1904,
+ "image_id": 1370,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 206,
+ 148,
+ 44,
+ 116
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1905,
+ "image_id": 1371,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 328,
+ 357,
+ 32,
+ 102
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1906,
+ "image_id": 1372,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 194,
+ 284,
+ 219,
+ 224
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1907,
+ "image_id": 1373,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 195,
+ 178,
+ 132,
+ 327
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1908,
+ "image_id": 1374,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 202,
+ 253,
+ 64,
+ 76
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1909,
+ "image_id": 1375,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 236,
+ 243,
+ 83,
+ 96
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1910,
+ "image_id": 1376,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 198,
+ 196,
+ 191,
+ 182
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1911,
+ "image_id": 1377,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 164,
+ 394,
+ 344
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1912,
+ "image_id": 1378,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 14,
+ 62,
+ 47,
+ 34
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1913,
+ "image_id": 1378,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 91,
+ 100,
+ 110
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1914,
+ "image_id": 1378,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 19,
+ 177,
+ 184,
+ 142
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1915,
+ "image_id": 1378,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 208,
+ 183,
+ 152,
+ 108
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1916,
+ "image_id": 1378,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 101,
+ 284,
+ 351,
+ 194
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1917,
+ "image_id": 1379,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 232,
+ 279,
+ 266,
+ 210
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1918,
+ "image_id": 1380,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 51,
+ 236,
+ 266,
+ 266
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1919,
+ "image_id": 1381,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 196,
+ 117,
+ 150,
+ 204
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1920,
+ "image_id": 1381,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 233,
+ 116,
+ 268,
+ 361
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1921,
+ "image_id": 1382,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 270,
+ 295,
+ 29,
+ 31
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1922,
+ "image_id": 1382,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 193,
+ 376,
+ 112,
+ 96
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1923,
+ "image_id": 1382,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 149,
+ 149,
+ 18,
+ 16
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1924,
+ "image_id": 1382,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 178,
+ 282,
+ 59,
+ 21
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1925,
+ "image_id": 1383,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 181,
+ 103,
+ 66
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1926,
+ "image_id": 1384,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 399,
+ 424,
+ 109,
+ 82
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1927,
+ "image_id": 1385,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 196,
+ 405,
+ 136,
+ 102
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1928,
+ "image_id": 1386,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 221,
+ 402,
+ 39,
+ 103
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1929,
+ "image_id": 1387,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 228,
+ 342,
+ 52,
+ 44
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1930,
+ "image_id": 1387,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 301,
+ 346,
+ 99,
+ 28
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1931,
+ "image_id": 1387,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 397,
+ 306,
+ 108,
+ 67
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1932,
+ "image_id": 1388,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 56,
+ 148,
+ 78,
+ 146
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1933,
+ "image_id": 1388,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 153,
+ 215,
+ 68,
+ 54
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1934,
+ "image_id": 1389,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 237,
+ 53,
+ 152,
+ 99
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1935,
+ "image_id": 1389,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 405,
+ 19,
+ 103,
+ 194
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1936,
+ "image_id": 1390,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 243,
+ 157,
+ 116
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1937,
+ "image_id": 1391,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 121,
+ 196,
+ 96,
+ 152
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1938,
+ "image_id": 1392,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 130,
+ 228,
+ 100,
+ 156
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1939,
+ "image_id": 1393,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 125,
+ 381,
+ 182,
+ 126
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1940,
+ "image_id": 1394,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 415,
+ 217,
+ 21,
+ 47
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1941,
+ "image_id": 1395,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 184,
+ 367,
+ 266
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1942,
+ "image_id": 1396,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 51,
+ 314,
+ 387,
+ 167
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1943,
+ "image_id": 1397,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 258,
+ 493,
+ 247
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1944,
+ "image_id": 1397,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 310,
+ 225,
+ 162,
+ 57
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1945,
+ "image_id": 1398,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 53,
+ 278,
+ 318,
+ 219
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1946,
+ "image_id": 1399,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 210,
+ 210,
+ 138,
+ 87
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1947,
+ "image_id": 1400,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 18,
+ 209,
+ 123,
+ 83
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1948,
+ "image_id": 1400,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 140,
+ 208,
+ 112,
+ 74
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1949,
+ "image_id": 1400,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 240,
+ 206,
+ 112,
+ 72
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1950,
+ "image_id": 1400,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 339,
+ 210,
+ 109,
+ 73
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1951,
+ "image_id": 1400,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 431,
+ 221,
+ 78,
+ 80
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1952,
+ "image_id": 1401,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 225,
+ 252,
+ 128,
+ 147
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1953,
+ "image_id": 1401,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 68,
+ 371,
+ 349,
+ 136
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1954,
+ "image_id": 1402,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 183,
+ 298,
+ 52,
+ 75
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1955,
+ "image_id": 1403,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 315,
+ 312,
+ 53,
+ 36
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1956,
+ "image_id": 1403,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 392,
+ 329,
+ 58,
+ 35
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1957,
+ "image_id": 1404,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 133,
+ 324,
+ 21,
+ 39
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1958,
+ "image_id": 1405,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 25,
+ 115,
+ 109,
+ 88
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1959,
+ "image_id": 1405,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 105,
+ 125,
+ 315,
+ 290
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1960,
+ "image_id": 1406,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 404,
+ 122,
+ 89,
+ 118
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1961,
+ "image_id": 1407,
+ "category_id": 4,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 345,
+ 306,
+ 163,
+ 128
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1962,
+ "image_id": 1408,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 329,
+ 366,
+ 25,
+ 27
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1963,
+ "image_id": 1409,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 311,
+ 374,
+ 196,
+ 130
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1964,
+ "image_id": 1410,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 241,
+ 322,
+ 26,
+ 98
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1965,
+ "image_id": 1411,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 469,
+ 216,
+ 39,
+ 66
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1966,
+ "image_id": 1412,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 62,
+ 263,
+ 217,
+ 131
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1967,
+ "image_id": 1412,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 11,
+ 383,
+ 494,
+ 125
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1968,
+ "image_id": 1413,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 231,
+ 399,
+ 156,
+ 105
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1969,
+ "image_id": 1414,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 141,
+ 284,
+ 152,
+ 80
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1970,
+ "image_id": 1414,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 77,
+ 327,
+ 425,
+ 172
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1971,
+ "image_id": 1415,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 219,
+ 303,
+ 109,
+ 145
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1972,
+ "image_id": 1416,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 38,
+ 187,
+ 302,
+ 243
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1973,
+ "image_id": 1416,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 307,
+ 214,
+ 204,
+ 247
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1974,
+ "image_id": 1417,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 279,
+ 303,
+ 93,
+ 143
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1975,
+ "image_id": 1417,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 438,
+ 375,
+ 55,
+ 131
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1976,
+ "image_id": 1418,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 192,
+ 387,
+ 112,
+ 118
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1977,
+ "image_id": 1418,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 304,
+ 322,
+ 104,
+ 130
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1978,
+ "image_id": 1419,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 77,
+ 377,
+ 76,
+ 56
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1979,
+ "image_id": 1419,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 28,
+ 421,
+ 480,
+ 86
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1980,
+ "image_id": 1420,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 241,
+ 274,
+ 65,
+ 85
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1981,
+ "image_id": 1420,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 177,
+ 350,
+ 196,
+ 154
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1982,
+ "image_id": 1421,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 11,
+ 334,
+ 200,
+ 170
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1983,
+ "image_id": 1422,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 211,
+ 172,
+ 60,
+ 87
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1984,
+ "image_id": 1422,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 286,
+ 225,
+ 84,
+ 113
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1985,
+ "image_id": 1422,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 446,
+ 252,
+ 66,
+ 101
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1986,
+ "image_id": 1423,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 150,
+ 194,
+ 124
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1987,
+ "image_id": 1423,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 144,
+ 314,
+ 362,
+ 193
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1988,
+ "image_id": 1423,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 241,
+ 135,
+ 82,
+ 109
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1989,
+ "image_id": 1424,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 199,
+ 433,
+ 309
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1990,
+ "image_id": 1425,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 45,
+ 219,
+ 141,
+ 98
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1991,
+ "image_id": 1425,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 359,
+ 291,
+ 124,
+ 213
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1992,
+ "image_id": 1426,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 198,
+ 407,
+ 197,
+ 96
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1993,
+ "image_id": 1427,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 102,
+ 295,
+ 406,
+ 193
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1994,
+ "image_id": 1428,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 252,
+ 185,
+ 116,
+ 78
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1995,
+ "image_id": 1428,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 150,
+ 161,
+ 111,
+ 62
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1996,
+ "image_id": 1428,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 88,
+ 215,
+ 308,
+ 294
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1997,
+ "image_id": 1429,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 352,
+ 252,
+ 96,
+ 144
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1998,
+ "image_id": 1430,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 240,
+ 299,
+ 116,
+ 132
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 1999,
+ "image_id": 1430,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 16,
+ 395,
+ 184,
+ 107
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2000,
+ "image_id": 1431,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 202,
+ 303,
+ 300,
+ 142
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2001,
+ "image_id": 1431,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 248,
+ 283,
+ 124
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2002,
+ "image_id": 1432,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 134,
+ 256,
+ 228,
+ 186
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2003,
+ "image_id": 1432,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 47,
+ 422,
+ 407,
+ 82
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2004,
+ "image_id": 1432,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 435,
+ 242,
+ 75,
+ 166
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2005,
+ "image_id": 1433,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 116,
+ 249,
+ 156,
+ 168
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2006,
+ "image_id": 1433,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 239,
+ 423,
+ 184,
+ 81
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2007,
+ "image_id": 1434,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 390,
+ 201,
+ 116,
+ 185
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2008,
+ "image_id": 1434,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 116,
+ 286,
+ 52,
+ 76
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2009,
+ "image_id": 1435,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 271,
+ 332,
+ 131,
+ 143
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2010,
+ "image_id": 1436,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 288,
+ 296,
+ 133,
+ 61
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2011,
+ "image_id": 1437,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 203,
+ 338,
+ 284,
+ 148
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2012,
+ "image_id": 1438,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 87,
+ 165,
+ 352,
+ 322
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2013,
+ "image_id": 1439,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 312,
+ 81,
+ 189
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2014,
+ "image_id": 1439,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 77,
+ 208,
+ 111,
+ 233
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2015,
+ "image_id": 1440,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 141,
+ 221,
+ 265,
+ 284
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2016,
+ "image_id": 1441,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 203,
+ 285,
+ 40,
+ 46
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2017,
+ "image_id": 1441,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 188,
+ 328,
+ 152,
+ 106
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2018,
+ "image_id": 1442,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 157,
+ 327,
+ 117,
+ 154
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2019,
+ "image_id": 1443,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 316,
+ 314,
+ 85,
+ 105
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2020,
+ "image_id": 1444,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 151,
+ 362,
+ 231,
+ 143
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2021,
+ "image_id": 1444,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 312,
+ 114,
+ 161
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2022,
+ "image_id": 1445,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 136,
+ 114,
+ 153,
+ 230
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2023,
+ "image_id": 1445,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 308,
+ 292,
+ 201
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2024,
+ "image_id": 1445,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 461,
+ 126,
+ 47,
+ 116
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2025,
+ "image_id": 1446,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 65,
+ 256,
+ 49,
+ 85
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2026,
+ "image_id": 1446,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 121,
+ 278,
+ 58,
+ 103
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2027,
+ "image_id": 1446,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 201,
+ 296,
+ 125,
+ 131
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2028,
+ "image_id": 1447,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 312,
+ 388,
+ 79,
+ 113
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2029,
+ "image_id": 1448,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 85,
+ 267,
+ 84,
+ 56
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2030,
+ "image_id": 1448,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 147,
+ 260,
+ 84,
+ 67
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2031,
+ "image_id": 1448,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 312,
+ 291,
+ 151,
+ 123
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2032,
+ "image_id": 1449,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 214,
+ 240,
+ 48,
+ 64
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2033,
+ "image_id": 1449,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 254,
+ 282,
+ 68,
+ 113
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2034,
+ "image_id": 1450,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 122,
+ 161,
+ 140,
+ 165
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2035,
+ "image_id": 1450,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 7,
+ 292,
+ 272,
+ 216
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2036,
+ "image_id": 1450,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 413,
+ 177,
+ 95,
+ 149
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2037,
+ "image_id": 1451,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 205,
+ 367,
+ 58,
+ 35
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2038,
+ "image_id": 1452,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 205,
+ 316,
+ 215,
+ 182
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2039,
+ "image_id": 1453,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 83,
+ 230,
+ 171,
+ 138
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2040,
+ "image_id": 1454,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 28,
+ 246,
+ 181,
+ 216
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2041,
+ "image_id": 1455,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 117,
+ 376,
+ 182,
+ 129
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2042,
+ "image_id": 1456,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 37,
+ 238,
+ 346,
+ 79
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2043,
+ "image_id": 1457,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 94,
+ 160,
+ 100,
+ 161
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2044,
+ "image_id": 1458,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 150,
+ 280,
+ 354,
+ 163
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2045,
+ "image_id": 1459,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 242,
+ 357,
+ 80,
+ 98
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2046,
+ "image_id": 1459,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 315,
+ 363,
+ 131,
+ 99
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2047,
+ "image_id": 1459,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 467,
+ 331,
+ 41,
+ 107
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2048,
+ "image_id": 1459,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 94,
+ 348,
+ 131,
+ 103
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2049,
+ "image_id": 1460,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 81,
+ 322,
+ 345,
+ 189
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2050,
+ "image_id": 1461,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 309,
+ 244,
+ 117,
+ 100
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2051,
+ "image_id": 1461,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 305,
+ 321,
+ 118,
+ 116
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2052,
+ "image_id": 1461,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 331,
+ 335,
+ 141,
+ 140
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2053,
+ "image_id": 1461,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 416,
+ 324,
+ 94,
+ 116
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2054,
+ "image_id": 1461,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 417,
+ 248,
+ 91,
+ 105
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2055,
+ "image_id": 1461,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 251,
+ 294,
+ 50,
+ 15
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2056,
+ "image_id": 1462,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 96,
+ 374,
+ 143,
+ 131
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2057,
+ "image_id": 1463,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 32,
+ 340,
+ 247,
+ 112
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2058,
+ "image_id": 1464,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 161,
+ 284,
+ 101,
+ 20
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2059,
+ "image_id": 1464,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 283,
+ 305,
+ 199,
+ 55
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2060,
+ "image_id": 1465,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 277,
+ 383,
+ 212,
+ 122
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2061,
+ "image_id": 1465,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 221,
+ 299,
+ 211,
+ 207
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2062,
+ "image_id": 1466,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 73,
+ 182,
+ 142,
+ 231
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2063,
+ "image_id": 1466,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 368,
+ 84,
+ 133,
+ 142
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2064,
+ "image_id": 1466,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 403,
+ 164,
+ 107,
+ 316
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2065,
+ "image_id": 1466,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 276,
+ 57,
+ 37,
+ 106
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2066,
+ "image_id": 1467,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 166,
+ 286,
+ 134,
+ 160
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2067,
+ "image_id": 1468,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 112,
+ 285,
+ 399,
+ 221
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2068,
+ "image_id": 1468,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 360,
+ 132,
+ 146,
+ 187
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2069,
+ "image_id": 1469,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 321,
+ 315,
+ 189,
+ 80
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2070,
+ "image_id": 1469,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 192,
+ 374,
+ 294,
+ 125
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2071,
+ "image_id": 1470,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 156,
+ 292,
+ 237,
+ 209
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2072,
+ "image_id": 1471,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 203,
+ 307,
+ 189,
+ 203
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2073,
+ "image_id": 1472,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 32,
+ 340,
+ 98,
+ 87
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2074,
+ "image_id": 1472,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 178,
+ 385,
+ 323,
+ 116
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2075,
+ "image_id": 1473,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 69,
+ 293,
+ 290,
+ 214
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2076,
+ "image_id": 1474,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 311,
+ 272,
+ 133,
+ 55
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2077,
+ "image_id": 1474,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 58,
+ 327,
+ 123,
+ 138
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2078,
+ "image_id": 1475,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 44,
+ 38,
+ 456,
+ 265
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2079,
+ "image_id": 1476,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 112,
+ 49,
+ 130,
+ 71
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2080,
+ "image_id": 1476,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 79,
+ 140,
+ 309,
+ 186
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2081,
+ "image_id": 1476,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 298,
+ 10,
+ 83,
+ 77
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2082,
+ "image_id": 1477,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 134,
+ 100,
+ 331,
+ 234
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2083,
+ "image_id": 1477,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 416,
+ 45,
+ 90,
+ 65
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2084,
+ "image_id": 1478,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 130,
+ 113,
+ 234,
+ 188
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2085,
+ "image_id": 1478,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 84,
+ 45,
+ 103,
+ 85
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2086,
+ "image_id": 1479,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 138,
+ 84,
+ 264,
+ 367
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2087,
+ "image_id": 1480,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 18,
+ 107,
+ 475,
+ 396
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2088,
+ "image_id": 1481,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 69,
+ 489,
+ 427
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2089,
+ "image_id": 1482,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 20,
+ 78,
+ 464,
+ 304
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2090,
+ "image_id": 1483,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 28,
+ 70,
+ 473,
+ 414
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2091,
+ "image_id": 1484,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 140,
+ 495,
+ 333
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2092,
+ "image_id": 1485,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 126,
+ 161,
+ 275,
+ 253
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2093,
+ "image_id": 1486,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 47,
+ 137,
+ 381,
+ 244
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2094,
+ "image_id": 1487,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 13,
+ 69,
+ 487,
+ 399
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2095,
+ "image_id": 1488,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 23,
+ 102,
+ 469,
+ 346
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2096,
+ "image_id": 1489,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 65,
+ 125,
+ 408,
+ 298
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2097,
+ "image_id": 1490,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 8,
+ 87,
+ 498,
+ 380
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2098,
+ "image_id": 1491,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 35,
+ 41,
+ 130,
+ 81
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2099,
+ "image_id": 1491,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 39,
+ 118,
+ 383,
+ 252
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2100,
+ "image_id": 1492,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 26,
+ 43,
+ 141,
+ 85
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2101,
+ "image_id": 1492,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 51,
+ 113,
+ 364,
+ 255
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2102,
+ "image_id": 1493,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 69,
+ 101,
+ 348,
+ 247
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2103,
+ "image_id": 1494,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 174,
+ 155,
+ 191,
+ 118
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2104,
+ "image_id": 1494,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 7,
+ 271,
+ 339,
+ 229
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2105,
+ "image_id": 1495,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 167,
+ 496,
+ 208
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2106,
+ "image_id": 1496,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 44,
+ 113,
+ 351,
+ 317
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2107,
+ "image_id": 1497,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 183,
+ 204,
+ 141,
+ 123
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2108,
+ "image_id": 1498,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 181,
+ 207,
+ 132,
+ 131
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2109,
+ "image_id": 1499,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 143,
+ 231,
+ 175,
+ 156
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2110,
+ "image_id": 1500,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 213,
+ 175,
+ 117,
+ 151
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2111,
+ "image_id": 1501,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 151,
+ 168,
+ 168,
+ 143
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2112,
+ "image_id": 1502,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 96,
+ 137,
+ 382,
+ 322
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2113,
+ "image_id": 1503,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 30,
+ 103,
+ 457,
+ 388
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2114,
+ "image_id": 1504,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 163,
+ 378,
+ 120,
+ 130
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2115,
+ "image_id": 1504,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 280,
+ 43,
+ 62
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2116,
+ "image_id": 1505,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 195,
+ 255,
+ 132,
+ 90
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2117,
+ "image_id": 1505,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 423,
+ 277,
+ 85,
+ 65
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2118,
+ "image_id": 1506,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 200,
+ 328,
+ 80,
+ 44
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2119,
+ "image_id": 1506,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 407,
+ 338,
+ 90,
+ 117
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2120,
+ "image_id": 1507,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 56,
+ 269,
+ 377,
+ 230
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2121,
+ "image_id": 1508,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 353,
+ 237,
+ 84,
+ 72
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2122,
+ "image_id": 1508,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 210,
+ 289,
+ 277,
+ 87
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2123,
+ "image_id": 1508,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 107,
+ 335,
+ 399,
+ 169
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2124,
+ "image_id": 1508,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 173,
+ 256,
+ 39,
+ 63
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2125,
+ "image_id": 1509,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 198,
+ 265,
+ 172,
+ 200
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2126,
+ "image_id": 1510,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 291,
+ 124,
+ 215,
+ 381
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2127,
+ "image_id": 1511,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 151,
+ 205,
+ 219,
+ 177
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2128,
+ "image_id": 1512,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 22,
+ 7,
+ 465,
+ 481
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2129,
+ "image_id": 1513,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 234,
+ 317,
+ 190,
+ 182
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2130,
+ "image_id": 1514,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 75,
+ 211,
+ 344,
+ 288
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2131,
+ "image_id": 1515,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 270,
+ 197,
+ 120,
+ 39
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2132,
+ "image_id": 1515,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 111,
+ 243,
+ 259,
+ 162
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2133,
+ "image_id": 1516,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 151,
+ 257,
+ 293,
+ 233
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2134,
+ "image_id": 1517,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 241,
+ 298,
+ 200,
+ 161
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2135,
+ "image_id": 1518,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 143,
+ 488,
+ 357
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2136,
+ "image_id": 1519,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 7,
+ 328,
+ 212,
+ 136
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2137,
+ "image_id": 1520,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 59,
+ 443,
+ 448
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2138,
+ "image_id": 1521,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 69,
+ 212,
+ 317,
+ 289
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2139,
+ "image_id": 1522,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 127,
+ 307,
+ 178,
+ 92
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2140,
+ "image_id": 1522,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 254,
+ 317,
+ 154,
+ 173
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2141,
+ "image_id": 1523,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 26,
+ 294,
+ 135,
+ 96
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2142,
+ "image_id": 1523,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 386,
+ 162,
+ 117
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2143,
+ "image_id": 1523,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 237,
+ 284,
+ 145,
+ 223
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2144,
+ "image_id": 1523,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 375,
+ 359,
+ 119,
+ 137
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2145,
+ "image_id": 1524,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 142,
+ 395,
+ 360
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2146,
+ "image_id": 1525,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 169,
+ 186,
+ 133,
+ 218
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2147,
+ "image_id": 1526,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 310,
+ 487,
+ 187
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2148,
+ "image_id": 1526,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 102,
+ 199,
+ 123,
+ 114
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2149,
+ "image_id": 1527,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 126,
+ 282,
+ 200,
+ 228
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2150,
+ "image_id": 1528,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 15,
+ 295,
+ 103,
+ 75
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2151,
+ "image_id": 1529,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 75,
+ 296,
+ 342,
+ 32
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2152,
+ "image_id": 1529,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 45,
+ 317,
+ 416,
+ 28
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2153,
+ "image_id": 1529,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 9,
+ 339,
+ 488,
+ 64
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2154,
+ "image_id": 1529,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 466,
+ 296,
+ 34,
+ 43
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2155,
+ "image_id": 1529,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 17,
+ 390,
+ 482,
+ 108
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2156,
+ "image_id": 1530,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 179,
+ 235,
+ 183,
+ 153
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2157,
+ "image_id": 1531,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 180,
+ 278,
+ 160,
+ 169
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2158,
+ "image_id": 1532,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 126,
+ 68,
+ 321,
+ 381
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2159,
+ "image_id": 1533,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 109,
+ 218,
+ 310,
+ 142
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2160,
+ "image_id": 1534,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 57,
+ 76,
+ 344,
+ 314
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2161,
+ "image_id": 1535,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 251,
+ 338,
+ 176,
+ 164
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2162,
+ "image_id": 1536,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 22,
+ 347,
+ 449,
+ 109
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2163,
+ "image_id": 1537,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 173,
+ 446,
+ 331
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2164,
+ "image_id": 1538,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 225,
+ 325,
+ 281
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2165,
+ "image_id": 1539,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 86,
+ 374,
+ 384,
+ 129
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2166,
+ "image_id": 1540,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 125,
+ 222,
+ 309,
+ 274
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2167,
+ "image_id": 1541,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 138,
+ 347,
+ 108,
+ 155
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2168,
+ "image_id": 1541,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 256,
+ 309,
+ 64,
+ 89
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2169,
+ "image_id": 1541,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 421,
+ 358,
+ 87,
+ 147
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2170,
+ "image_id": 1541,
+ "category_id": 10,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 406,
+ 281,
+ 41,
+ 47
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2171,
+ "image_id": 1542,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 86,
+ 139,
+ 309,
+ 334
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2172,
+ "image_id": 1543,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 46,
+ 208,
+ 461,
+ 271
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2173,
+ "image_id": 1544,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 159,
+ 315,
+ 201,
+ 185
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2174,
+ "image_id": 1544,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 5,
+ 225,
+ 170,
+ 143
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2175,
+ "image_id": 1544,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 277,
+ 260,
+ 56,
+ 70
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2176,
+ "image_id": 1545,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 261,
+ 261,
+ 157,
+ 86
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2177,
+ "image_id": 1545,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 75,
+ 293,
+ 235,
+ 153
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2178,
+ "image_id": 1546,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 193,
+ 327,
+ 275,
+ 176
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2179,
+ "image_id": 1547,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 254,
+ 133,
+ 201
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2180,
+ "image_id": 1547,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 256,
+ 291,
+ 251,
+ 199
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2181,
+ "image_id": 1548,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 15,
+ 173,
+ 395,
+ 336
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2182,
+ "image_id": 1549,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 24,
+ 326,
+ 479,
+ 179
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2183,
+ "image_id": 1550,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 6,
+ 161,
+ 294,
+ 253
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2184,
+ "image_id": 1551,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 259,
+ 253,
+ 54,
+ 30
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2185,
+ "image_id": 1551,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 179,
+ 243,
+ 43,
+ 44
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2186,
+ "image_id": 1551,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 80,
+ 280,
+ 76,
+ 73
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2187,
+ "image_id": 1551,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 127,
+ 365,
+ 71,
+ 92
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2188,
+ "image_id": 1551,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 164,
+ 277,
+ 121,
+ 105
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2189,
+ "image_id": 1551,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 354,
+ 289,
+ 99,
+ 46
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2190,
+ "image_id": 1551,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 339,
+ 315,
+ 66,
+ 187
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2191,
+ "image_id": 1552,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 12,
+ 311,
+ 287,
+ 193
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2192,
+ "image_id": 1553,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 204,
+ 251,
+ 134,
+ 240
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2193,
+ "image_id": 1554,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 274,
+ 151,
+ 176
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2194,
+ "image_id": 1554,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 99,
+ 247,
+ 41,
+ 35
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2195,
+ "image_id": 1554,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 163,
+ 241,
+ 52,
+ 68
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2196,
+ "image_id": 1554,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 277,
+ 251,
+ 88,
+ 69
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2197,
+ "image_id": 1554,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 311,
+ 248,
+ 80,
+ 38
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2198,
+ "image_id": 1554,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 412,
+ 247,
+ 92,
+ 59
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2199,
+ "image_id": 1554,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 414,
+ 275,
+ 95,
+ 121
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2200,
+ "image_id": 1554,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 279,
+ 302,
+ 229,
+ 206
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2201,
+ "image_id": 1555,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 141,
+ 156,
+ 223,
+ 218
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2202,
+ "image_id": 1556,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 103,
+ 67,
+ 311,
+ 438
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2203,
+ "image_id": 1556,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 413,
+ 103,
+ 97,
+ 55
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2204,
+ "image_id": 1556,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 372,
+ 81,
+ 48,
+ 21
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2205,
+ "image_id": 1557,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 155,
+ 304,
+ 162,
+ 190
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2206,
+ "image_id": 1558,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 7,
+ 316,
+ 284,
+ 187
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2207,
+ "image_id": 1559,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 331,
+ 333,
+ 123,
+ 162
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2208,
+ "image_id": 1559,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 103,
+ 340,
+ 119,
+ 165
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2209,
+ "image_id": 1559,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 181,
+ 304,
+ 53,
+ 78
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2210,
+ "image_id": 1560,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 179,
+ 496,
+ 284
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2211,
+ "image_id": 1561,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 197,
+ 239,
+ 254,
+ 231
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2212,
+ "image_id": 1562,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 113,
+ 307,
+ 128,
+ 185
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2213,
+ "image_id": 1563,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 16,
+ 127,
+ 131,
+ 76
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2214,
+ "image_id": 1563,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 246,
+ 53,
+ 66,
+ 91
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2215,
+ "image_id": 1563,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 34,
+ 86,
+ 75,
+ 48
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2216,
+ "image_id": 1563,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 152,
+ 236,
+ 344,
+ 266
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2217,
+ "image_id": 1564,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 119,
+ 298,
+ 286,
+ 188
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2218,
+ "image_id": 1565,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 64,
+ 392,
+ 345,
+ 95
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2219,
+ "image_id": 1565,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 349,
+ 329,
+ 145,
+ 81
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2220,
+ "image_id": 1566,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 27,
+ 220,
+ 217,
+ 125
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2221,
+ "image_id": 1566,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 315,
+ 350,
+ 187,
+ 154
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2222,
+ "image_id": 1567,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 204,
+ 261,
+ 139,
+ 26
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2223,
+ "image_id": 1567,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 78,
+ 315,
+ 47,
+ 128
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2224,
+ "image_id": 1567,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 219,
+ 358,
+ 97,
+ 115
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2225,
+ "image_id": 1568,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 31,
+ 271,
+ 75,
+ 116
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2226,
+ "image_id": 1568,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 166,
+ 248,
+ 53,
+ 66
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2227,
+ "image_id": 1569,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 52,
+ 197,
+ 280,
+ 307
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2228,
+ "image_id": 1570,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 171,
+ 193,
+ 207,
+ 193
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2229,
+ "image_id": 1571,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 167,
+ 283,
+ 87,
+ 61
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2230,
+ "image_id": 1571,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 31,
+ 312,
+ 107,
+ 33
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2231,
+ "image_id": 1571,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 14,
+ 338,
+ 273,
+ 163
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2232,
+ "image_id": 1572,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 102,
+ 119,
+ 108,
+ 230
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2233,
+ "image_id": 1573,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 104,
+ 114,
+ 260,
+ 321
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2234,
+ "image_id": 1574,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 216,
+ 104,
+ 113,
+ 28
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2235,
+ "image_id": 1574,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 201,
+ 132,
+ 149,
+ 52
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2236,
+ "image_id": 1574,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 162,
+ 175,
+ 231,
+ 91
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2237,
+ "image_id": 1574,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 10,
+ 355,
+ 498,
+ 152
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2238,
+ "image_id": 1575,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 103,
+ 320,
+ 259,
+ 167
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2239,
+ "image_id": 1576,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 107,
+ 294,
+ 351,
+ 214
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2240,
+ "image_id": 1577,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 68,
+ 364,
+ 178,
+ 94
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2241,
+ "image_id": 1578,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 212,
+ 313,
+ 275,
+ 96
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2242,
+ "image_id": 1579,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 105,
+ 57,
+ 314,
+ 448
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2243,
+ "image_id": 1580,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 196,
+ 191,
+ 281,
+ 317
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2244,
+ "image_id": 1581,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 299,
+ 239,
+ 173,
+ 183
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2245,
+ "image_id": 1581,
+ "category_id": 9,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 74,
+ 47,
+ 236,
+ 459
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2246,
+ "image_id": 1582,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 60,
+ 125,
+ 441,
+ 378
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2247,
+ "image_id": 1583,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 410,
+ 275,
+ 96,
+ 81
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2248,
+ "image_id": 1583,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 46,
+ 254,
+ 329,
+ 180
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2249,
+ "image_id": 1584,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 109,
+ 386,
+ 245,
+ 118
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2250,
+ "image_id": 1584,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 72,
+ 289,
+ 138,
+ 53
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2251,
+ "image_id": 1584,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 286,
+ 316,
+ 220,
+ 124
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2252,
+ "image_id": 1585,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 199,
+ 298,
+ 286,
+ 207
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2253,
+ "image_id": 1586,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 134,
+ 92,
+ 184,
+ 291
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2254,
+ "image_id": 1587,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 158,
+ 298,
+ 90,
+ 38
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2255,
+ "image_id": 1587,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 116,
+ 329,
+ 151,
+ 164
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2256,
+ "image_id": 1587,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 367,
+ 324,
+ 129,
+ 168
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2257,
+ "image_id": 1588,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 183,
+ 333,
+ 48,
+ 79
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2258,
+ "image_id": 1588,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 145,
+ 347,
+ 57,
+ 96
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2259,
+ "image_id": 1589,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 97,
+ 316,
+ 168,
+ 139
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2260,
+ "image_id": 1589,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 326,
+ 328,
+ 109,
+ 89
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2261,
+ "image_id": 1590,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 116,
+ 116,
+ 174,
+ 227
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2262,
+ "image_id": 1591,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 82,
+ 339,
+ 156,
+ 170
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2263,
+ "image_id": 1592,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 24,
+ 209,
+ 376,
+ 253
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2264,
+ "image_id": 1593,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 137,
+ 325,
+ 210,
+ 175
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2265,
+ "image_id": 1594,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 138,
+ 244,
+ 126,
+ 60
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2266,
+ "image_id": 1594,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 321,
+ 237,
+ 84,
+ 24
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2267,
+ "image_id": 1594,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 268,
+ 251,
+ 97,
+ 25
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2268,
+ "image_id": 1594,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 251,
+ 269,
+ 133,
+ 40
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2269,
+ "image_id": 1594,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 195,
+ 298,
+ 222,
+ 101
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2270,
+ "image_id": 1594,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 313,
+ 209,
+ 102
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2271,
+ "image_id": 1595,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 88,
+ 264,
+ 243,
+ 180
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2272,
+ "image_id": 1595,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 440,
+ 265,
+ 68,
+ 115
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2273,
+ "image_id": 1596,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 68,
+ 279,
+ 55,
+ 50
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2274,
+ "image_id": 1596,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 153,
+ 264,
+ 67,
+ 65
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2275,
+ "image_id": 1596,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 342,
+ 270,
+ 59,
+ 52
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2276,
+ "image_id": 1597,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 266,
+ 214,
+ 203,
+ 249
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2277,
+ "image_id": 1598,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 127,
+ 178,
+ 297,
+ 239
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2278,
+ "image_id": 1599,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 162,
+ 434,
+ 341
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2279,
+ "image_id": 1600,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 7,
+ 312,
+ 153,
+ 194
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2280,
+ "image_id": 1601,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 110,
+ 289,
+ 240,
+ 202
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2281,
+ "image_id": 1602,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 384,
+ 300,
+ 67,
+ 77
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2282,
+ "image_id": 1603,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 183,
+ 286,
+ 174,
+ 83
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2283,
+ "image_id": 1604,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 335,
+ 318,
+ 55,
+ 60
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2284,
+ "image_id": 1604,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 388,
+ 305,
+ 56,
+ 51
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2285,
+ "image_id": 1605,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 88,
+ 172,
+ 243
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2286,
+ "image_id": 1606,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 70,
+ 306,
+ 356,
+ 200
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2287,
+ "image_id": 1607,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 397,
+ 274,
+ 111,
+ 70
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2288,
+ "image_id": 1607,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 90,
+ 296,
+ 292,
+ 138
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2289,
+ "image_id": 1608,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 76,
+ 263,
+ 368,
+ 240
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2290,
+ "image_id": 1609,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 229,
+ 80,
+ 115,
+ 125
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2291,
+ "image_id": 1609,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 112,
+ 234,
+ 368,
+ 262
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2292,
+ "image_id": 1610,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 52,
+ 311,
+ 408,
+ 193
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2293,
+ "image_id": 1611,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 319,
+ 413,
+ 189
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2294,
+ "image_id": 1612,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 37,
+ 290,
+ 447,
+ 217
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2295,
+ "image_id": 1613,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 133,
+ 291,
+ 311,
+ 162
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2296,
+ "image_id": 1614,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 253,
+ 189,
+ 241,
+ 320
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2297,
+ "image_id": 1615,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 105,
+ 337,
+ 272,
+ 156
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2298,
+ "image_id": 1616,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 39,
+ 321,
+ 448,
+ 175
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2299,
+ "image_id": 1617,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 8,
+ 314,
+ 488,
+ 183
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2300,
+ "image_id": 1618,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 220,
+ 237,
+ 68,
+ 109
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2301,
+ "image_id": 1618,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 150,
+ 123,
+ 106,
+ 68
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2302,
+ "image_id": 1619,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 29,
+ 254,
+ 228,
+ 164
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2303,
+ "image_id": 1620,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 252,
+ 485,
+ 256
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2304,
+ "image_id": 1621,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 215,
+ 147,
+ 212,
+ 197
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2305,
+ "image_id": 1622,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 42,
+ 172,
+ 256,
+ 332
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2306,
+ "image_id": 1623,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 150,
+ 361,
+ 141,
+ 145
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2307,
+ "image_id": 1623,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 396,
+ 217,
+ 82,
+ 126
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2308,
+ "image_id": 1623,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 340,
+ 171,
+ 51,
+ 86
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2309,
+ "image_id": 1624,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 70,
+ 249,
+ 157,
+ 177
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2310,
+ "image_id": 1625,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 3,
+ 304,
+ 96,
+ 102
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2311,
+ "image_id": 1625,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 145,
+ 209,
+ 97,
+ 133
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2312,
+ "image_id": 1625,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 293,
+ 275,
+ 106,
+ 233
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2313,
+ "image_id": 1626,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 197,
+ 224,
+ 94,
+ 57
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2314,
+ "image_id": 1626,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 158,
+ 272,
+ 194,
+ 179
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2315,
+ "image_id": 1627,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 230,
+ 296,
+ 160,
+ 207
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2316,
+ "image_id": 1628,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 16,
+ 348,
+ 222,
+ 154
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2317,
+ "image_id": 1629,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 91,
+ 253,
+ 328,
+ 243
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2318,
+ "image_id": 1630,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 191,
+ 167,
+ 51,
+ 106
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2319,
+ "image_id": 1631,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 419,
+ 242,
+ 61,
+ 91
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2320,
+ "image_id": 1632,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 352,
+ 227,
+ 123,
+ 135
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2321,
+ "image_id": 1632,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 98,
+ 289,
+ 321,
+ 144
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2322,
+ "image_id": 1632,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 1,
+ 248,
+ 92,
+ 88
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2323,
+ "image_id": 1633,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 414,
+ 331,
+ 98,
+ 174
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2324,
+ "image_id": 1634,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 198,
+ 271,
+ 221,
+ 235
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2325,
+ "image_id": 1635,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 297,
+ 320,
+ 150,
+ 164
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2326,
+ "image_id": 1635,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 158,
+ 311,
+ 136,
+ 81
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2327,
+ "image_id": 1636,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 155,
+ 350,
+ 124,
+ 156
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2328,
+ "image_id": 1637,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 240,
+ 157,
+ 187,
+ 349
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2329,
+ "image_id": 1638,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 354,
+ 266,
+ 156,
+ 230
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2330,
+ "image_id": 1638,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 308,
+ 231,
+ 89,
+ 114
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2331,
+ "image_id": 1639,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 2,
+ 292,
+ 282,
+ 204
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2332,
+ "image_id": 1640,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 130,
+ 298,
+ 276,
+ 206
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2333,
+ "image_id": 1640,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 424,
+ 288,
+ 85,
+ 203
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2334,
+ "image_id": 1640,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 65,
+ 286,
+ 80,
+ 127
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2335,
+ "image_id": 1641,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 344,
+ 204,
+ 160,
+ 253
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2336,
+ "image_id": 1642,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 50,
+ 299,
+ 103,
+ 111
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2337,
+ "image_id": 1642,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 184,
+ 272,
+ 46,
+ 58
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2338,
+ "image_id": 1643,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 8,
+ 192,
+ 234,
+ 309
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2339,
+ "image_id": 1644,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 259,
+ 264,
+ 235
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2340,
+ "image_id": 1645,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 114,
+ 307,
+ 105,
+ 97
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2341,
+ "image_id": 1645,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 353,
+ 277,
+ 44,
+ 65
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2342,
+ "image_id": 1645,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 360,
+ 306,
+ 144,
+ 197
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2343,
+ "image_id": 1646,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 246,
+ 348,
+ 125,
+ 155
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2344,
+ "image_id": 1647,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 150,
+ 313,
+ 210,
+ 196
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2345,
+ "image_id": 1648,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 327,
+ 280,
+ 172,
+ 222
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2346,
+ "image_id": 1649,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 104,
+ 335,
+ 119,
+ 163
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2347,
+ "image_id": 1650,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 215,
+ 287,
+ 88,
+ 158
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2348,
+ "image_id": 1651,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 235,
+ 324,
+ 99,
+ 150
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2349,
+ "image_id": 1652,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 182,
+ 311,
+ 186,
+ 179
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2350,
+ "image_id": 1652,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 1,
+ 204,
+ 147,
+ 213
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2351,
+ "image_id": 1653,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 74,
+ 380,
+ 244,
+ 127
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2352,
+ "image_id": 1654,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 110,
+ 283,
+ 65,
+ 85
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2353,
+ "image_id": 1654,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 400,
+ 340,
+ 109,
+ 166
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2354,
+ "image_id": 1654,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 254,
+ 364,
+ 198,
+ 142
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2355,
+ "image_id": 1655,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 113,
+ 355,
+ 147,
+ 150
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2356,
+ "image_id": 1655,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 260,
+ 160,
+ 102,
+ 74
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2357,
+ "image_id": 1655,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 200,
+ 145,
+ 26,
+ 88
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2358,
+ "image_id": 1656,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 286,
+ 388,
+ 67,
+ 120
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2359,
+ "image_id": 1657,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 62,
+ 243,
+ 377,
+ 261
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2360,
+ "image_id": 1658,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 68,
+ 228,
+ 347,
+ 275
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2361,
+ "image_id": 1659,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 296,
+ 364,
+ 191,
+ 142
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2362,
+ "image_id": 1660,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 74,
+ 229,
+ 180,
+ 150
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2363,
+ "image_id": 1660,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 39,
+ 205,
+ 34,
+ 130
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2364,
+ "image_id": 1661,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 260,
+ 364,
+ 202,
+ 146
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2365,
+ "image_id": 1662,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 69,
+ 179,
+ 170,
+ 140
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2366,
+ "image_id": 1662,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 4,
+ 334,
+ 295,
+ 169
+ ],
+ "iscrowd": 0
+ },
+ {
+ "id": 2367,
+ "image_id": 1663,
+ "category_id": 12,
+ "segmentation": [],
+ "area": 262144,
+ "bbox": [
+ 317,
+ 122,
+ 177,
+ 251
+ ],
+ "iscrowd": 0
+ }
+ ],
+ "categories": [
+ {
+ "id": 1,
+ "name": "bicycle",
+ "supercategory": "object"
+ },
+ {
+ "id": 2,
+ "name": "boat",
+ "supercategory": "object"
+ },
+ {
+ "id": 3,
+ "name": "bottle",
+ "supercategory": "object"
+ },
+ {
+ "id": 4,
+ "name": "bus",
+ "supercategory": "object"
+ },
+ {
+ "id": 5,
+ "name": "car",
+ "supercategory": "object"
+ },
+ {
+ "id": 6,
+ "name": "cat",
+ "supercategory": "object"
+ },
+ {
+ "id": 7,
+ "name": "chair",
+ "supercategory": "object"
+ },
+ {
+ "id": 8,
+ "name": "cup",
+ "supercategory": "object"
+ },
+ {
+ "id": 9,
+ "name": "dog",
+ "supercategory": "object"
+ },
+ {
+ "id": 10,
+ "name": "motorbike",
+ "supercategory": "object"
+ },
+ {
+ "id": 11,
+ "name": "people",
+ "supercategory": "object"
+ },
+ {
+ "id": 12,
+ "name": "table",
+ "supercategory": "object"
+ }
+ ]
+}
\ No newline at end of file
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/gt_jsons/ruod/instance_test_novel.json b/scripts/evaluation/FasterRCNN_score-mmdet/gt_jsons/ruod/instance_test_novel.json
new file mode 100644
index 0000000000000000000000000000000000000000..7c3588f6cae39143a25d31a9404b800098bcbd20
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/gt_jsons/ruod/instance_test_novel.json
@@ -0,0 +1,100892 @@
+{
+ "info": {
+ "description": "",
+ "url": "",
+ "version": "",
+ "year": 2022,
+ "contributor": "\u7eaf\u7cb9",
+ "date_created": "2022-07-8"
+ },
+ "licenses": [
+ {
+ "id": 1,
+ "name": null,
+ "url": null
+ }
+ ],
+ "images": [
+ {
+ "id": 1,
+ "file_name": "007937.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2,
+ "file_name": "008551.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 7,
+ "file_name": "012163.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 8,
+ "file_name": "012230.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 9,
+ "file_name": "009746.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 11,
+ "file_name": "005032.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 16,
+ "file_name": "013776.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 17,
+ "file_name": "011997.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 19,
+ "file_name": "009127.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 22,
+ "file_name": "010980.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 23,
+ "file_name": "006969.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 24,
+ "file_name": "004553.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 26,
+ "file_name": "004998.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 27,
+ "file_name": "012266.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 29,
+ "file_name": "007747.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 30,
+ "file_name": "009815.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 33,
+ "file_name": "005741.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 34,
+ "file_name": "012390.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 37,
+ "file_name": "007102.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 39,
+ "file_name": "013280.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 42,
+ "file_name": "010600.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 44,
+ "file_name": "004184.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 46,
+ "file_name": "010697.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 47,
+ "file_name": "013604.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 49,
+ "file_name": "005250.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 51,
+ "file_name": "010624.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 52,
+ "file_name": "010986.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 57,
+ "file_name": "010851.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 58,
+ "file_name": "007748.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 59,
+ "file_name": "006704.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 60,
+ "file_name": "012759.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 62,
+ "file_name": "007467.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 63,
+ "file_name": "008421.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 64,
+ "file_name": "009125.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 69,
+ "file_name": "007177.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 71,
+ "file_name": "013090.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 72,
+ "file_name": "010324.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 73,
+ "file_name": "008540.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 75,
+ "file_name": "011066.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 76,
+ "file_name": "008614.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 77,
+ "file_name": "006851.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 78,
+ "file_name": "006918.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 79,
+ "file_name": "010763.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 80,
+ "file_name": "008123.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 81,
+ "file_name": "008588.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 82,
+ "file_name": "009507.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 86,
+ "file_name": "010719.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 91,
+ "file_name": "007810.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 92,
+ "file_name": "007159.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 93,
+ "file_name": "013841.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 95,
+ "file_name": "007613.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 97,
+ "file_name": "003934.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 98,
+ "file_name": "008507.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 99,
+ "file_name": "011506.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 100,
+ "file_name": "012156.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 102,
+ "file_name": "005083.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 104,
+ "file_name": "010642.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 109,
+ "file_name": "010723.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 110,
+ "file_name": "008059.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 111,
+ "file_name": "010399.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 112,
+ "file_name": "012008.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 113,
+ "file_name": "005026.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 114,
+ "file_name": "004443.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 118,
+ "file_name": "009925.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 119,
+ "file_name": "011024.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 120,
+ "file_name": "009537.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 121,
+ "file_name": "012098.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 122,
+ "file_name": "013736.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 123,
+ "file_name": "013591.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 128,
+ "file_name": "007956.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 129,
+ "file_name": "010413.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 130,
+ "file_name": "011520.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 134,
+ "file_name": "005112.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 135,
+ "file_name": "008187.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 136,
+ "file_name": "012054.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 138,
+ "file_name": "010622.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 139,
+ "file_name": "004794.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 140,
+ "file_name": "011406.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 141,
+ "file_name": "011113.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 143,
+ "file_name": "010475.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 148,
+ "file_name": "009121.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 150,
+ "file_name": "006191.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 152,
+ "file_name": "012784.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 155,
+ "file_name": "009497.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 159,
+ "file_name": "010728.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 161,
+ "file_name": "005781.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 162,
+ "file_name": "009120.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 164,
+ "file_name": "012673.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 165,
+ "file_name": "005151.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 168,
+ "file_name": "010008.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 169,
+ "file_name": "010979.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 170,
+ "file_name": "010582.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 171,
+ "file_name": "013996.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 175,
+ "file_name": "007577.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 179,
+ "file_name": "008952.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 181,
+ "file_name": "008784.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 187,
+ "file_name": "008858.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 188,
+ "file_name": "007355.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 189,
+ "file_name": "010303.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 190,
+ "file_name": "011160.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 196,
+ "file_name": "012354.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 198,
+ "file_name": "012265.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 200,
+ "file_name": "008217.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 202,
+ "file_name": "011191.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 203,
+ "file_name": "008595.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 204,
+ "file_name": "007749.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 209,
+ "file_name": "007073.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 210,
+ "file_name": "004920.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 211,
+ "file_name": "009978.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 212,
+ "file_name": "007707.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 214,
+ "file_name": "007686.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 217,
+ "file_name": "009760.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 220,
+ "file_name": "009462.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 221,
+ "file_name": "009088.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 222,
+ "file_name": "008745.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 228,
+ "file_name": "011404.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 231,
+ "file_name": "011262.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 232,
+ "file_name": "005190.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 233,
+ "file_name": "009960.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 235,
+ "file_name": "013888.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 236,
+ "file_name": "005181.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 238,
+ "file_name": "006853.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 239,
+ "file_name": "008872.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 240,
+ "file_name": "008996.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 241,
+ "file_name": "006783.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 242,
+ "file_name": "009472.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 243,
+ "file_name": "006616.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 246,
+ "file_name": "007217.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 247,
+ "file_name": "004781.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 251,
+ "file_name": "004758.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 254,
+ "file_name": "010612.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 256,
+ "file_name": "006171.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 257,
+ "file_name": "008364.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 260,
+ "file_name": "011008.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 261,
+ "file_name": "012756.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 263,
+ "file_name": "011140.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 264,
+ "file_name": "011533.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 268,
+ "file_name": "008797.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 270,
+ "file_name": "012923.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 271,
+ "file_name": "007547.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 274,
+ "file_name": "007676.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 275,
+ "file_name": "006964.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 278,
+ "file_name": "012094.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 279,
+ "file_name": "010431.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 280,
+ "file_name": "009086.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 281,
+ "file_name": "008779.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 282,
+ "file_name": "013758.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 283,
+ "file_name": "003931.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 284,
+ "file_name": "011182.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 287,
+ "file_name": "011178.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 288,
+ "file_name": "011521.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 289,
+ "file_name": "010383.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 294,
+ "file_name": "006496.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 296,
+ "file_name": "007046.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 297,
+ "file_name": "006840.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 298,
+ "file_name": "013409.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 299,
+ "file_name": "008241.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 306,
+ "file_name": "008447.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 307,
+ "file_name": "012707.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 310,
+ "file_name": "013470.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 311,
+ "file_name": "004929.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 313,
+ "file_name": "009485.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 315,
+ "file_name": "011469.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 316,
+ "file_name": "005688.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 317,
+ "file_name": "004873.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 318,
+ "file_name": "008319.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 319,
+ "file_name": "005252.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 321,
+ "file_name": "004469.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 322,
+ "file_name": "006587.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 323,
+ "file_name": "006504.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 324,
+ "file_name": "011121.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 325,
+ "file_name": "013800.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 326,
+ "file_name": "005145.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 328,
+ "file_name": "011311.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 330,
+ "file_name": "009729.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 331,
+ "file_name": "007733.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 332,
+ "file_name": "013571.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 333,
+ "file_name": "013051.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 334,
+ "file_name": "011729.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 335,
+ "file_name": "012853.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 337,
+ "file_name": "009916.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 338,
+ "file_name": "013437.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 339,
+ "file_name": "007644.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 346,
+ "file_name": "007756.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 349,
+ "file_name": "008436.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 350,
+ "file_name": "004151.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 352,
+ "file_name": "012717.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 353,
+ "file_name": "004084.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 355,
+ "file_name": "006693.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 356,
+ "file_name": "007609.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 357,
+ "file_name": "007023.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 358,
+ "file_name": "010726.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 360,
+ "file_name": "013964.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 362,
+ "file_name": "009039.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 364,
+ "file_name": "007961.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 365,
+ "file_name": "004824.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 366,
+ "file_name": "008394.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 367,
+ "file_name": "012365.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 369,
+ "file_name": "009446.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 370,
+ "file_name": "009093.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 371,
+ "file_name": "004213.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 374,
+ "file_name": "005786.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 375,
+ "file_name": "011338.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 376,
+ "file_name": "007540.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 378,
+ "file_name": "005788.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 379,
+ "file_name": "007054.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 380,
+ "file_name": "010934.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 381,
+ "file_name": "012761.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 385,
+ "file_name": "008825.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 386,
+ "file_name": "008069.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 387,
+ "file_name": "008998.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 388,
+ "file_name": "006824.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 390,
+ "file_name": "006697.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 391,
+ "file_name": "006389.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 392,
+ "file_name": "010415.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 393,
+ "file_name": "008965.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 396,
+ "file_name": "012912.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 397,
+ "file_name": "013045.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 399,
+ "file_name": "007460.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 402,
+ "file_name": "013886.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 403,
+ "file_name": "013465.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 404,
+ "file_name": "012825.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 405,
+ "file_name": "007347.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 407,
+ "file_name": "008038.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 410,
+ "file_name": "013170.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 411,
+ "file_name": "004660.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 412,
+ "file_name": "009949.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 413,
+ "file_name": "008283.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 414,
+ "file_name": "013755.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 415,
+ "file_name": "004206.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 417,
+ "file_name": "006677.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 419,
+ "file_name": "006750.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 421,
+ "file_name": "008692.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 426,
+ "file_name": "006176.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 427,
+ "file_name": "007963.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 429,
+ "file_name": "013244.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 430,
+ "file_name": "012203.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 433,
+ "file_name": "013397.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 434,
+ "file_name": "008798.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 435,
+ "file_name": "004444.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 436,
+ "file_name": "008820.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 437,
+ "file_name": "010544.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 438,
+ "file_name": "010307.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 439,
+ "file_name": "011234.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 440,
+ "file_name": "004533.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 443,
+ "file_name": "011196.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 446,
+ "file_name": "006406.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 447,
+ "file_name": "008066.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 450,
+ "file_name": "006561.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 451,
+ "file_name": "013628.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 452,
+ "file_name": "012282.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 453,
+ "file_name": "007583.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 454,
+ "file_name": "013920.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 455,
+ "file_name": "013652.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 456,
+ "file_name": "010411.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 461,
+ "file_name": "008004.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 462,
+ "file_name": "007730.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 463,
+ "file_name": "012369.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 465,
+ "file_name": "007057.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 466,
+ "file_name": "004904.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 470,
+ "file_name": "009743.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 471,
+ "file_name": "008284.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 472,
+ "file_name": "011576.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 474,
+ "file_name": "013478.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 475,
+ "file_name": "013794.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 477,
+ "file_name": "008427.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 478,
+ "file_name": "009991.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 481,
+ "file_name": "008964.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 482,
+ "file_name": "008378.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 484,
+ "file_name": "006468.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 485,
+ "file_name": "013109.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 486,
+ "file_name": "009779.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 489,
+ "file_name": "012993.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 490,
+ "file_name": "013319.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 492,
+ "file_name": "013055.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 494,
+ "file_name": "004943.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 496,
+ "file_name": "011272.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 497,
+ "file_name": "010757.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 498,
+ "file_name": "010985.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 499,
+ "file_name": "011998.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 501,
+ "file_name": "007433.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 505,
+ "file_name": "008845.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 507,
+ "file_name": "009763.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 508,
+ "file_name": "007248.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 510,
+ "file_name": "012031.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 512,
+ "file_name": "007536.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 514,
+ "file_name": "006640.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 515,
+ "file_name": "005213.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 517,
+ "file_name": "008333.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 518,
+ "file_name": "013436.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 519,
+ "file_name": "012564.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 523,
+ "file_name": "008294.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 524,
+ "file_name": "008457.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 525,
+ "file_name": "006882.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 526,
+ "file_name": "009974.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 530,
+ "file_name": "007604.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 531,
+ "file_name": "006369.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 534,
+ "file_name": "006396.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 536,
+ "file_name": "007050.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 537,
+ "file_name": "012844.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 539,
+ "file_name": "007706.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 540,
+ "file_name": "011438.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 541,
+ "file_name": "013751.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 546,
+ "file_name": "011277.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 547,
+ "file_name": "013168.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 550,
+ "file_name": "008212.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 551,
+ "file_name": "008213.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 553,
+ "file_name": "008068.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 556,
+ "file_name": "009057.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 558,
+ "file_name": "010158.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 559,
+ "file_name": "008554.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 561,
+ "file_name": "008502.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 562,
+ "file_name": "011021.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 563,
+ "file_name": "007636.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 564,
+ "file_name": "008393.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 566,
+ "file_name": "010793.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 572,
+ "file_name": "013208.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 573,
+ "file_name": "004993.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 574,
+ "file_name": "008387.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 577,
+ "file_name": "006522.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 579,
+ "file_name": "013242.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 581,
+ "file_name": "006250.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 584,
+ "file_name": "010503.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 586,
+ "file_name": "012957.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 587,
+ "file_name": "007056.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 588,
+ "file_name": "004964.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 589,
+ "file_name": "011174.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 590,
+ "file_name": "005588.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 593,
+ "file_name": "005011.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 594,
+ "file_name": "012335.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 596,
+ "file_name": "011540.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 597,
+ "file_name": "012491.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 598,
+ "file_name": "012313.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 601,
+ "file_name": "011305.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 602,
+ "file_name": "007907.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 603,
+ "file_name": "009586.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 604,
+ "file_name": "009301.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 607,
+ "file_name": "004108.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 608,
+ "file_name": "007896.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 609,
+ "file_name": "007220.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 610,
+ "file_name": "013258.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 616,
+ "file_name": "007332.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 617,
+ "file_name": "013889.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 619,
+ "file_name": "011327.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 623,
+ "file_name": "013549.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 624,
+ "file_name": "004975.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 626,
+ "file_name": "009523.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 629,
+ "file_name": "007843.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 630,
+ "file_name": "012959.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 632,
+ "file_name": "011680.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 638,
+ "file_name": "011440.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 640,
+ "file_name": "009481.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 642,
+ "file_name": "008528.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 645,
+ "file_name": "006958.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 649,
+ "file_name": "006339.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 652,
+ "file_name": "013349.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 654,
+ "file_name": "006706.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 655,
+ "file_name": "012699.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 656,
+ "file_name": "010493.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 657,
+ "file_name": "010778.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 658,
+ "file_name": "012106.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 660,
+ "file_name": "008953.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 661,
+ "file_name": "012349.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 662,
+ "file_name": "008096.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 665,
+ "file_name": "007246.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 666,
+ "file_name": "006211.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 667,
+ "file_name": "010504.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 670,
+ "file_name": "008041.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 672,
+ "file_name": "004318.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 674,
+ "file_name": "006965.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 677,
+ "file_name": "011348.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 678,
+ "file_name": "005039.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 681,
+ "file_name": "011732.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 683,
+ "file_name": "008028.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 689,
+ "file_name": "006736.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 690,
+ "file_name": "004146.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 693,
+ "file_name": "008050.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 695,
+ "file_name": "013613.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 696,
+ "file_name": "011377.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 697,
+ "file_name": "004118.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 698,
+ "file_name": "009512.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 699,
+ "file_name": "008814.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 700,
+ "file_name": "004982.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 703,
+ "file_name": "007658.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 706,
+ "file_name": "012776.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 707,
+ "file_name": "011557.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 708,
+ "file_name": "007297.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 709,
+ "file_name": "008647.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 713,
+ "file_name": "009753.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 717,
+ "file_name": "010627.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 718,
+ "file_name": "007838.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 719,
+ "file_name": "006748.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 722,
+ "file_name": "012605.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 723,
+ "file_name": "013498.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 725,
+ "file_name": "013621.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 726,
+ "file_name": "013932.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 729,
+ "file_name": "013780.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 730,
+ "file_name": "004116.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 731,
+ "file_name": "011026.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 733,
+ "file_name": "013667.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 734,
+ "file_name": "011556.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 735,
+ "file_name": "004768.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 736,
+ "file_name": "007086.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 737,
+ "file_name": "008229.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 742,
+ "file_name": "011499.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 743,
+ "file_name": "012617.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 744,
+ "file_name": "013926.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 747,
+ "file_name": "007880.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 748,
+ "file_name": "007860.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 755,
+ "file_name": "008397.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 759,
+ "file_name": "007020.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 760,
+ "file_name": "008685.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 761,
+ "file_name": "003868.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 763,
+ "file_name": "011401.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 765,
+ "file_name": "009570.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 767,
+ "file_name": "005765.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 768,
+ "file_name": "013919.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 770,
+ "file_name": "008742.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 771,
+ "file_name": "013489.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 774,
+ "file_name": "006859.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 775,
+ "file_name": "011989.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 776,
+ "file_name": "011523.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 777,
+ "file_name": "012486.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 778,
+ "file_name": "005107.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 781,
+ "file_name": "009872.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 782,
+ "file_name": "010925.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 783,
+ "file_name": "005837.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 795,
+ "file_name": "012247.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 797,
+ "file_name": "009979.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 799,
+ "file_name": "003912.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 800,
+ "file_name": "012255.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 802,
+ "file_name": "013754.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 803,
+ "file_name": "013824.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 804,
+ "file_name": "008802.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 805,
+ "file_name": "010569.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 806,
+ "file_name": "010691.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 808,
+ "file_name": "009563.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 809,
+ "file_name": "013375.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 812,
+ "file_name": "006690.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 813,
+ "file_name": "010391.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 815,
+ "file_name": "012746.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 817,
+ "file_name": "012070.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 820,
+ "file_name": "010585.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 821,
+ "file_name": "003965.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 823,
+ "file_name": "006641.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 824,
+ "file_name": "007018.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 825,
+ "file_name": "011076.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 826,
+ "file_name": "007682.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 828,
+ "file_name": "008331.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 829,
+ "file_name": "006916.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 830,
+ "file_name": "013647.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 831,
+ "file_name": "009334.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 833,
+ "file_name": "007040.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 834,
+ "file_name": "009383.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 835,
+ "file_name": "013339.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 837,
+ "file_name": "008133.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 840,
+ "file_name": "013199.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 841,
+ "file_name": "004503.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 843,
+ "file_name": "010711.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 845,
+ "file_name": "011148.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 850,
+ "file_name": "004481.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 851,
+ "file_name": "013125.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 853,
+ "file_name": "008617.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 855,
+ "file_name": "013419.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 858,
+ "file_name": "006667.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 860,
+ "file_name": "008120.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 865,
+ "file_name": "004150.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 867,
+ "file_name": "010565.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 868,
+ "file_name": "012801.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 869,
+ "file_name": "009065.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 870,
+ "file_name": "008543.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 872,
+ "file_name": "006243.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 873,
+ "file_name": "013481.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 875,
+ "file_name": "006324.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 878,
+ "file_name": "013972.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 880,
+ "file_name": "010396.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 882,
+ "file_name": "012172.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 885,
+ "file_name": "008539.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 888,
+ "file_name": "010854.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 891,
+ "file_name": "005070.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 892,
+ "file_name": "007490.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 893,
+ "file_name": "011227.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 895,
+ "file_name": "006665.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 897,
+ "file_name": "010795.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 898,
+ "file_name": "010672.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 900,
+ "file_name": "003952.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 902,
+ "file_name": "013543.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 903,
+ "file_name": "007235.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 906,
+ "file_name": "011028.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 907,
+ "file_name": "007133.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 909,
+ "file_name": "007899.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 911,
+ "file_name": "006644.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 913,
+ "file_name": "012153.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 916,
+ "file_name": "009259.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 917,
+ "file_name": "008602.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 918,
+ "file_name": "007878.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 919,
+ "file_name": "012101.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 922,
+ "file_name": "010203.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 923,
+ "file_name": "009905.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 925,
+ "file_name": "008701.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 928,
+ "file_name": "013497.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 931,
+ "file_name": "005005.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 932,
+ "file_name": "013278.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 935,
+ "file_name": "012587.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 938,
+ "file_name": "007165.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 943,
+ "file_name": "009804.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 946,
+ "file_name": "004446.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 950,
+ "file_name": "007148.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 951,
+ "file_name": "013614.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 952,
+ "file_name": "010445.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 953,
+ "file_name": "007809.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 954,
+ "file_name": "004199.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 955,
+ "file_name": "011419.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 956,
+ "file_name": "006578.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 957,
+ "file_name": "004159.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 960,
+ "file_name": "012870.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 962,
+ "file_name": "011962.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 964,
+ "file_name": "007329.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 966,
+ "file_name": "008976.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 967,
+ "file_name": "003980.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 970,
+ "file_name": "011349.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 973,
+ "file_name": "006570.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 974,
+ "file_name": "005016.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 980,
+ "file_name": "007928.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 983,
+ "file_name": "006453.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 984,
+ "file_name": "013702.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 986,
+ "file_name": "012312.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 987,
+ "file_name": "011720.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 988,
+ "file_name": "007570.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 989,
+ "file_name": "010908.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 991,
+ "file_name": "011975.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 992,
+ "file_name": "011162.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 993,
+ "file_name": "004552.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 995,
+ "file_name": "008661.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1001,
+ "file_name": "012680.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1005,
+ "file_name": "007455.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1006,
+ "file_name": "013730.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1010,
+ "file_name": "012185.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1012,
+ "file_name": "007697.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1015,
+ "file_name": "009562.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1016,
+ "file_name": "007949.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1018,
+ "file_name": "013202.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1019,
+ "file_name": "008959.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1020,
+ "file_name": "010755.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1021,
+ "file_name": "006708.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1023,
+ "file_name": "013813.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1027,
+ "file_name": "007141.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1028,
+ "file_name": "005244.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1030,
+ "file_name": "011293.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1031,
+ "file_name": "012582.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1034,
+ "file_name": "004495.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1036,
+ "file_name": "007807.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1037,
+ "file_name": "005744.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1039,
+ "file_name": "012730.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1040,
+ "file_name": "013383.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1042,
+ "file_name": "011205.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1043,
+ "file_name": "013814.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1046,
+ "file_name": "010637.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1047,
+ "file_name": "013080.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1051,
+ "file_name": "010706.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1053,
+ "file_name": "013312.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1054,
+ "file_name": "013182.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1056,
+ "file_name": "013781.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1057,
+ "file_name": "008775.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1059,
+ "file_name": "012173.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1060,
+ "file_name": "011151.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1061,
+ "file_name": "013148.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1062,
+ "file_name": "009352.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1063,
+ "file_name": "013333.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1064,
+ "file_name": "007155.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1069,
+ "file_name": "008381.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1070,
+ "file_name": "013640.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1071,
+ "file_name": "004973.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1073,
+ "file_name": "013279.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1075,
+ "file_name": "008448.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1077,
+ "file_name": "008520.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1078,
+ "file_name": "013359.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1080,
+ "file_name": "009805.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1081,
+ "file_name": "013248.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1083,
+ "file_name": "013846.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1086,
+ "file_name": "007537.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1087,
+ "file_name": "011126.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1088,
+ "file_name": "008556.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1091,
+ "file_name": "013527.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1092,
+ "file_name": "004559.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1096,
+ "file_name": "005391.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1097,
+ "file_name": "010167.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1099,
+ "file_name": "011953.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1101,
+ "file_name": "012724.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1102,
+ "file_name": "011250.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1103,
+ "file_name": "007374.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1106,
+ "file_name": "007259.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1109,
+ "file_name": "006937.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1111,
+ "file_name": "011498.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1112,
+ "file_name": "008222.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1113,
+ "file_name": "005835.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1115,
+ "file_name": "011450.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1118,
+ "file_name": "009417.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1121,
+ "file_name": "010940.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1122,
+ "file_name": "006348.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1123,
+ "file_name": "009571.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1124,
+ "file_name": "013057.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1125,
+ "file_name": "013224.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1126,
+ "file_name": "010976.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1127,
+ "file_name": "012718.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1128,
+ "file_name": "007367.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1131,
+ "file_name": "008824.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1132,
+ "file_name": "008773.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1133,
+ "file_name": "008029.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1134,
+ "file_name": "005468.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1135,
+ "file_name": "009808.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1136,
+ "file_name": "009969.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1137,
+ "file_name": "004041.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1138,
+ "file_name": "013004.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1140,
+ "file_name": "007463.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1141,
+ "file_name": "004111.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1143,
+ "file_name": "012210.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1145,
+ "file_name": "006689.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1146,
+ "file_name": "006466.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1147,
+ "file_name": "010200.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1149,
+ "file_name": "007223.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1150,
+ "file_name": "013387.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1151,
+ "file_name": "005572.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1152,
+ "file_name": "008841.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1153,
+ "file_name": "006000.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1155,
+ "file_name": "007527.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1156,
+ "file_name": "005810.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1157,
+ "file_name": "006805.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1158,
+ "file_name": "004615.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1161,
+ "file_name": "007967.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1162,
+ "file_name": "013673.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1163,
+ "file_name": "013286.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1164,
+ "file_name": "012148.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1165,
+ "file_name": "007600.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1170,
+ "file_name": "007488.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1173,
+ "file_name": "012762.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1175,
+ "file_name": "006008.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1179,
+ "file_name": "008673.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1180,
+ "file_name": "004956.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1182,
+ "file_name": "005226.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1183,
+ "file_name": "006385.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1184,
+ "file_name": "013679.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1192,
+ "file_name": "007181.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1194,
+ "file_name": "004236.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1195,
+ "file_name": "007224.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1197,
+ "file_name": "005051.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1198,
+ "file_name": "006613.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1199,
+ "file_name": "012146.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1200,
+ "file_name": "007830.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1201,
+ "file_name": "010146.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1206,
+ "file_name": "007360.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1207,
+ "file_name": "004175.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1209,
+ "file_name": "004643.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1210,
+ "file_name": "008402.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1213,
+ "file_name": "006381.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1214,
+ "file_name": "013306.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1216,
+ "file_name": "009747.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1217,
+ "file_name": "008530.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1219,
+ "file_name": "005046.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1228,
+ "file_name": "008829.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1230,
+ "file_name": "008424.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1232,
+ "file_name": "007780.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1234,
+ "file_name": "012553.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1235,
+ "file_name": "007402.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1238,
+ "file_name": "013517.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1240,
+ "file_name": "012123.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1241,
+ "file_name": "012019.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1242,
+ "file_name": "008277.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1243,
+ "file_name": "010759.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1245,
+ "file_name": "006556.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1247,
+ "file_name": "013773.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1249,
+ "file_name": "011475.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1251,
+ "file_name": "004763.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1252,
+ "file_name": "012326.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1253,
+ "file_name": "006901.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1255,
+ "file_name": "009360.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1256,
+ "file_name": "007699.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1257,
+ "file_name": "004062.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1259,
+ "file_name": "004242.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1260,
+ "file_name": "007811.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1263,
+ "file_name": "010860.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1264,
+ "file_name": "011383.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1265,
+ "file_name": "007228.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1267,
+ "file_name": "012074.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1269,
+ "file_name": "010590.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1271,
+ "file_name": "013865.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1272,
+ "file_name": "013635.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1274,
+ "file_name": "004614.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1278,
+ "file_name": "009100.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1280,
+ "file_name": "013210.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1281,
+ "file_name": "013029.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1283,
+ "file_name": "007183.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1287,
+ "file_name": "010930.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1288,
+ "file_name": "006618.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1289,
+ "file_name": "007112.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1290,
+ "file_name": "012300.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1291,
+ "file_name": "011934.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1292,
+ "file_name": "013330.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1293,
+ "file_name": "012601.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1296,
+ "file_name": "009591.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1297,
+ "file_name": "010206.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1299,
+ "file_name": "012413.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1300,
+ "file_name": "013354.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1301,
+ "file_name": "007188.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1303,
+ "file_name": "006617.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1305,
+ "file_name": "012485.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1306,
+ "file_name": "011446.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1308,
+ "file_name": "008020.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1309,
+ "file_name": "004122.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1310,
+ "file_name": "008956.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1311,
+ "file_name": "005099.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1313,
+ "file_name": "008062.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1314,
+ "file_name": "013400.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1317,
+ "file_name": "004166.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1320,
+ "file_name": "013660.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1323,
+ "file_name": "011982.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1324,
+ "file_name": "008245.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1325,
+ "file_name": "004072.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1327,
+ "file_name": "006932.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1328,
+ "file_name": "013822.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1329,
+ "file_name": "007061.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1331,
+ "file_name": "007932.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1332,
+ "file_name": "010714.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1333,
+ "file_name": "005161.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1337,
+ "file_name": "009062.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1338,
+ "file_name": "004590.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1340,
+ "file_name": "004999.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1341,
+ "file_name": "007588.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1342,
+ "file_name": "009106.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1344,
+ "file_name": "008852.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1345,
+ "file_name": "012165.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1347,
+ "file_name": "013705.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1348,
+ "file_name": "012143.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1350,
+ "file_name": "013884.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1352,
+ "file_name": "004195.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1355,
+ "file_name": "003908.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1356,
+ "file_name": "012548.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1357,
+ "file_name": "008257.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1359,
+ "file_name": "010359.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1360,
+ "file_name": "004343.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1361,
+ "file_name": "009032.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1362,
+ "file_name": "011676.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1366,
+ "file_name": "008159.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1368,
+ "file_name": "013748.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1370,
+ "file_name": "006944.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1371,
+ "file_name": "004981.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1372,
+ "file_name": "005128.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1377,
+ "file_name": "010932.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1380,
+ "file_name": "012979.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1383,
+ "file_name": "010480.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1384,
+ "file_name": "013732.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1385,
+ "file_name": "012531.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1387,
+ "file_name": "009921.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1389,
+ "file_name": "009018.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1391,
+ "file_name": "008945.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1392,
+ "file_name": "013099.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1393,
+ "file_name": "013198.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1396,
+ "file_name": "007767.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1397,
+ "file_name": "008250.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1398,
+ "file_name": "013114.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1401,
+ "file_name": "009732.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1403,
+ "file_name": "007168.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1404,
+ "file_name": "009770.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1406,
+ "file_name": "006375.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1409,
+ "file_name": "007178.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1414,
+ "file_name": "007715.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1415,
+ "file_name": "007977.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1417,
+ "file_name": "011319.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1418,
+ "file_name": "013906.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1423,
+ "file_name": "007152.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1426,
+ "file_name": "009509.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1427,
+ "file_name": "013046.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1429,
+ "file_name": "009879.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1431,
+ "file_name": "012177.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1432,
+ "file_name": "009063.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1433,
+ "file_name": "013142.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1434,
+ "file_name": "008230.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1435,
+ "file_name": "008163.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1437,
+ "file_name": "010870.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1440,
+ "file_name": "012500.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1441,
+ "file_name": "012041.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1443,
+ "file_name": "007257.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1444,
+ "file_name": "007816.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1445,
+ "file_name": "007492.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1446,
+ "file_name": "010646.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1447,
+ "file_name": "004137.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1448,
+ "file_name": "005831.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1449,
+ "file_name": "005094.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1450,
+ "file_name": "013587.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1453,
+ "file_name": "007786.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1454,
+ "file_name": "012011.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1455,
+ "file_name": "013238.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1456,
+ "file_name": "013799.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1461,
+ "file_name": "012961.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1462,
+ "file_name": "010653.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1463,
+ "file_name": "010608.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1464,
+ "file_name": "008435.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1465,
+ "file_name": "008545.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1466,
+ "file_name": "008371.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1468,
+ "file_name": "007411.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1470,
+ "file_name": "012797.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1472,
+ "file_name": "013134.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1474,
+ "file_name": "008562.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1478,
+ "file_name": "012817.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1479,
+ "file_name": "013157.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1480,
+ "file_name": "013802.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1481,
+ "file_name": "011068.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1482,
+ "file_name": "003852.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1483,
+ "file_name": "006358.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1486,
+ "file_name": "005172.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1487,
+ "file_name": "003895.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1488,
+ "file_name": "004934.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1490,
+ "file_name": "008769.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1491,
+ "file_name": "008105.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1495,
+ "file_name": "006629.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1496,
+ "file_name": "003915.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1497,
+ "file_name": "012710.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1498,
+ "file_name": "010961.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1499,
+ "file_name": "013564.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1500,
+ "file_name": "004065.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1502,
+ "file_name": "013361.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1504,
+ "file_name": "008812.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1506,
+ "file_name": "007950.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1507,
+ "file_name": "007596.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1509,
+ "file_name": "007368.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1510,
+ "file_name": "010777.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1511,
+ "file_name": "010972.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1514,
+ "file_name": "004028.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1515,
+ "file_name": "009474.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1518,
+ "file_name": "007111.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1520,
+ "file_name": "008835.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1524,
+ "file_name": "013050.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1526,
+ "file_name": "012687.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1529,
+ "file_name": "009965.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1530,
+ "file_name": "007298.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1532,
+ "file_name": "004508.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1535,
+ "file_name": "004928.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1539,
+ "file_name": "013282.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1540,
+ "file_name": "012971.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1542,
+ "file_name": "013542.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1543,
+ "file_name": "011188.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1546,
+ "file_name": "008019.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1547,
+ "file_name": "012135.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1548,
+ "file_name": "009580.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1549,
+ "file_name": "013661.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1551,
+ "file_name": "013874.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1552,
+ "file_name": "008030.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1553,
+ "file_name": "005040.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1555,
+ "file_name": "013616.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1557,
+ "file_name": "008286.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1558,
+ "file_name": "007346.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1559,
+ "file_name": "004256.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1560,
+ "file_name": "010401.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1561,
+ "file_name": "006456.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1562,
+ "file_name": "008034.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1564,
+ "file_name": "012546.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1565,
+ "file_name": "004704.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1566,
+ "file_name": "006896.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1569,
+ "file_name": "008299.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1571,
+ "file_name": "012911.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1573,
+ "file_name": "013854.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1574,
+ "file_name": "009924.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1575,
+ "file_name": "010667.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1576,
+ "file_name": "010522.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1578,
+ "file_name": "013033.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1580,
+ "file_name": "007256.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1581,
+ "file_name": "013703.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1582,
+ "file_name": "010746.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1584,
+ "file_name": "013931.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1585,
+ "file_name": "008788.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1586,
+ "file_name": "007607.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1589,
+ "file_name": "008630.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1590,
+ "file_name": "005735.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1591,
+ "file_name": "007976.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1592,
+ "file_name": "007701.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1593,
+ "file_name": "012863.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1594,
+ "file_name": "013427.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1595,
+ "file_name": "006767.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1596,
+ "file_name": "007379.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1598,
+ "file_name": "010896.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1599,
+ "file_name": "010733.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1600,
+ "file_name": "008434.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1602,
+ "file_name": "011478.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1606,
+ "file_name": "007504.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1607,
+ "file_name": "008428.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1611,
+ "file_name": "012060.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1613,
+ "file_name": "013682.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1615,
+ "file_name": "004190.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1616,
+ "file_name": "008689.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1618,
+ "file_name": "011403.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1621,
+ "file_name": "012742.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1623,
+ "file_name": "006370.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1625,
+ "file_name": "012690.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1626,
+ "file_name": "006387.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1628,
+ "file_name": "006656.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1631,
+ "file_name": "008892.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1632,
+ "file_name": "008390.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1633,
+ "file_name": "011099.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1635,
+ "file_name": "005693.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1637,
+ "file_name": "009124.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1639,
+ "file_name": "008195.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1640,
+ "file_name": "010421.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1641,
+ "file_name": "012284.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1642,
+ "file_name": "011138.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1643,
+ "file_name": "012965.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1644,
+ "file_name": "009759.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1645,
+ "file_name": "010949.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1646,
+ "file_name": "009351.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1648,
+ "file_name": "011468.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1650,
+ "file_name": "013399.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1652,
+ "file_name": "005205.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1654,
+ "file_name": "004833.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1657,
+ "file_name": "007154.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1660,
+ "file_name": "013172.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1661,
+ "file_name": "006338.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1663,
+ "file_name": "010947.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1664,
+ "file_name": "006558.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1666,
+ "file_name": "006952.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1667,
+ "file_name": "003858.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1668,
+ "file_name": "004459.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1669,
+ "file_name": "008154.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1670,
+ "file_name": "005034.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1674,
+ "file_name": "011112.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1675,
+ "file_name": "007746.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1676,
+ "file_name": "010883.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1677,
+ "file_name": "008007.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1678,
+ "file_name": "013304.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1681,
+ "file_name": "009640.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1683,
+ "file_name": "013149.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1684,
+ "file_name": "007859.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1685,
+ "file_name": "012621.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1686,
+ "file_name": "008796.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1687,
+ "file_name": "006725.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1691,
+ "file_name": "009410.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1695,
+ "file_name": "005300.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1697,
+ "file_name": "012771.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1698,
+ "file_name": "005207.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1701,
+ "file_name": "007344.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1702,
+ "file_name": "006635.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1703,
+ "file_name": "009340.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1704,
+ "file_name": "006903.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1708,
+ "file_name": "004091.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1712,
+ "file_name": "012984.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1715,
+ "file_name": "008075.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1720,
+ "file_name": "013444.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1722,
+ "file_name": "013100.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1723,
+ "file_name": "005574.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1724,
+ "file_name": "010378.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1727,
+ "file_name": "005069.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1728,
+ "file_name": "012175.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1729,
+ "file_name": "009780.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1730,
+ "file_name": "011436.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1732,
+ "file_name": "006876.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1734,
+ "file_name": "010613.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1735,
+ "file_name": "010308.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1737,
+ "file_name": "012847.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1743,
+ "file_name": "008072.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1744,
+ "file_name": "013207.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1745,
+ "file_name": "013759.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1746,
+ "file_name": "006915.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1748,
+ "file_name": "007693.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1749,
+ "file_name": "012918.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1750,
+ "file_name": "008193.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1751,
+ "file_name": "009794.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1753,
+ "file_name": "009307.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1755,
+ "file_name": "011420.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1756,
+ "file_name": "008481.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1757,
+ "file_name": "006590.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1758,
+ "file_name": "007864.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1759,
+ "file_name": "006383.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1760,
+ "file_name": "007508.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1762,
+ "file_name": "005329.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1764,
+ "file_name": "007667.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1766,
+ "file_name": "009369.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1767,
+ "file_name": "012800.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1772,
+ "file_name": "011960.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1773,
+ "file_name": "008021.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1776,
+ "file_name": "013241.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1777,
+ "file_name": "007944.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1780,
+ "file_name": "013954.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1783,
+ "file_name": "013945.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1784,
+ "file_name": "006940.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1785,
+ "file_name": "011696.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1786,
+ "file_name": "007824.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1787,
+ "file_name": "008497.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1791,
+ "file_name": "012939.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1792,
+ "file_name": "010802.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1797,
+ "file_name": "012498.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1798,
+ "file_name": "009173.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1799,
+ "file_name": "010181.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1800,
+ "file_name": "006438.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1803,
+ "file_name": "005704.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1804,
+ "file_name": "009108.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1805,
+ "file_name": "004624.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1807,
+ "file_name": "013358.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1808,
+ "file_name": "011306.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1809,
+ "file_name": "012962.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1810,
+ "file_name": "007214.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1814,
+ "file_name": "007741.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1816,
+ "file_name": "013315.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1820,
+ "file_name": "008175.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1821,
+ "file_name": "008309.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1823,
+ "file_name": "007910.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1824,
+ "file_name": "009865.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1825,
+ "file_name": "008419.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1826,
+ "file_name": "013188.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1829,
+ "file_name": "010712.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1830,
+ "file_name": "009142.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1831,
+ "file_name": "013714.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1833,
+ "file_name": "006482.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1834,
+ "file_name": "010891.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1835,
+ "file_name": "013594.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1836,
+ "file_name": "004983.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1837,
+ "file_name": "008790.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1838,
+ "file_name": "009994.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1839,
+ "file_name": "013022.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1840,
+ "file_name": "008467.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1841,
+ "file_name": "008764.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1842,
+ "file_name": "013346.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1845,
+ "file_name": "009337.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1850,
+ "file_name": "008281.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1851,
+ "file_name": "004910.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1853,
+ "file_name": "008328.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1854,
+ "file_name": "008119.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1856,
+ "file_name": "006938.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1857,
+ "file_name": "008224.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1859,
+ "file_name": "011201.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1860,
+ "file_name": "012867.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1861,
+ "file_name": "007191.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1862,
+ "file_name": "011090.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1863,
+ "file_name": "008267.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1866,
+ "file_name": "012221.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1873,
+ "file_name": "011074.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1874,
+ "file_name": "006397.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1875,
+ "file_name": "010788.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1876,
+ "file_name": "004762.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1877,
+ "file_name": "011039.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1881,
+ "file_name": "006681.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1884,
+ "file_name": "005309.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1885,
+ "file_name": "008045.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1886,
+ "file_name": "007889.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1887,
+ "file_name": "008583.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1890,
+ "file_name": "004203.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1894,
+ "file_name": "010521.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1895,
+ "file_name": "006462.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1897,
+ "file_name": "012978.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1898,
+ "file_name": "009765.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1899,
+ "file_name": "004902.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1900,
+ "file_name": "009931.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1901,
+ "file_name": "005109.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1903,
+ "file_name": "008370.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1904,
+ "file_name": "012086.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1908,
+ "file_name": "006546.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1909,
+ "file_name": "009130.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1910,
+ "file_name": "012689.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1911,
+ "file_name": "013523.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1912,
+ "file_name": "009940.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1913,
+ "file_name": "006388.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1915,
+ "file_name": "007128.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1917,
+ "file_name": "011492.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1918,
+ "file_name": "004783.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1919,
+ "file_name": "013062.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1921,
+ "file_name": "010722.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1923,
+ "file_name": "011308.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1924,
+ "file_name": "012741.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1925,
+ "file_name": "008646.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1927,
+ "file_name": "013536.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1928,
+ "file_name": "009736.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1931,
+ "file_name": "006520.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1932,
+ "file_name": "006521.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1933,
+ "file_name": "012488.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1934,
+ "file_name": "004098.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1936,
+ "file_name": "009380.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1937,
+ "file_name": "009508.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1940,
+ "file_name": "013807.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1942,
+ "file_name": "013418.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1943,
+ "file_name": "004992.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1946,
+ "file_name": "009752.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1950,
+ "file_name": "008855.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1951,
+ "file_name": "004289.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1953,
+ "file_name": "006801.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1955,
+ "file_name": "013519.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1957,
+ "file_name": "010771.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1960,
+ "file_name": "011078.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1963,
+ "file_name": "012136.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1964,
+ "file_name": "008569.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1969,
+ "file_name": "004243.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1970,
+ "file_name": "007428.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1971,
+ "file_name": "012807.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1972,
+ "file_name": "004507.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1973,
+ "file_name": "013569.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1974,
+ "file_name": "013662.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1975,
+ "file_name": "011135.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1978,
+ "file_name": "011101.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1980,
+ "file_name": "007587.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1984,
+ "file_name": "009134.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1985,
+ "file_name": "005114.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1987,
+ "file_name": "009927.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1988,
+ "file_name": "006533.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1990,
+ "file_name": "009171.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1993,
+ "file_name": "010116.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1995,
+ "file_name": "013580.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1997,
+ "file_name": "008313.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 1998,
+ "file_name": "013001.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2000,
+ "file_name": "012968.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2001,
+ "file_name": "006647.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2002,
+ "file_name": "008321.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2003,
+ "file_name": "013291.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2004,
+ "file_name": "005843.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2006,
+ "file_name": "008517.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2008,
+ "file_name": "012503.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2010,
+ "file_name": "003901.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2012,
+ "file_name": "012052.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2014,
+ "file_name": "006790.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2017,
+ "file_name": "004488.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2018,
+ "file_name": "011144.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2019,
+ "file_name": "003910.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2022,
+ "file_name": "011411.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2025,
+ "file_name": "008992.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2027,
+ "file_name": "006740.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2028,
+ "file_name": "011395.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2030,
+ "file_name": "008167.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2032,
+ "file_name": "013048.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2034,
+ "file_name": "010984.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2035,
+ "file_name": "013767.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2036,
+ "file_name": "003903.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2038,
+ "file_name": "009431.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2042,
+ "file_name": "009559.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2045,
+ "file_name": "012356.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2050,
+ "file_name": "012934.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2051,
+ "file_name": "006813.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2052,
+ "file_name": "010542.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2053,
+ "file_name": "007282.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2055,
+ "file_name": "008867.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2056,
+ "file_name": "006469.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2058,
+ "file_name": "006924.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2059,
+ "file_name": "008200.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2060,
+ "file_name": "012846.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2061,
+ "file_name": "007238.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2062,
+ "file_name": "007873.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2063,
+ "file_name": "009526.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2064,
+ "file_name": "009418.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2066,
+ "file_name": "007406.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2068,
+ "file_name": "006579.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2069,
+ "file_name": "013791.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2070,
+ "file_name": "012674.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2072,
+ "file_name": "007149.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2074,
+ "file_name": "006973.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2077,
+ "file_name": "007315.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2078,
+ "file_name": "013905.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2079,
+ "file_name": "009531.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2081,
+ "file_name": "013296.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2083,
+ "file_name": "011213.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2084,
+ "file_name": "003922.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2087,
+ "file_name": "006420.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2092,
+ "file_name": "009265.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2093,
+ "file_name": "007522.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2094,
+ "file_name": "007597.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2095,
+ "file_name": "010887.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2099,
+ "file_name": "012206.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2101,
+ "file_name": "009098.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2102,
+ "file_name": "012420.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2105,
+ "file_name": "010902.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2108,
+ "file_name": "007900.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2109,
+ "file_name": "013572.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2110,
+ "file_name": "004070.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2111,
+ "file_name": "012782.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2112,
+ "file_name": "006398.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2114,
+ "file_name": "004032.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2115,
+ "file_name": "012329.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2117,
+ "file_name": "004338.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2118,
+ "file_name": "007241.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2119,
+ "file_name": "009332.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2121,
+ "file_name": "007230.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2122,
+ "file_name": "012216.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2123,
+ "file_name": "013479.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2127,
+ "file_name": "012676.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2129,
+ "file_name": "012456.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2130,
+ "file_name": "006415.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2131,
+ "file_name": "008183.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2135,
+ "file_name": "004880.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2136,
+ "file_name": "010437.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2137,
+ "file_name": "013957.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2139,
+ "file_name": "010184.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2141,
+ "file_name": "012805.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2143,
+ "file_name": "006928.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2144,
+ "file_name": "012977.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2146,
+ "file_name": "013688.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2147,
+ "file_name": "010650.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2149,
+ "file_name": "013715.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2154,
+ "file_name": "007823.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2155,
+ "file_name": "007650.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2156,
+ "file_name": "013632.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2159,
+ "file_name": "006642.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2161,
+ "file_name": "010205.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2162,
+ "file_name": "012963.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2166,
+ "file_name": "013830.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2167,
+ "file_name": "009438.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2168,
+ "file_name": "010394.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2169,
+ "file_name": "009001.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2170,
+ "file_name": "008302.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2172,
+ "file_name": "008651.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2176,
+ "file_name": "008087.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2179,
+ "file_name": "007681.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2182,
+ "file_name": "006835.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2185,
+ "file_name": "012161.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2186,
+ "file_name": "004110.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2188,
+ "file_name": "009464.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2189,
+ "file_name": "010734.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2191,
+ "file_name": "007357.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2192,
+ "file_name": "004845.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2195,
+ "file_name": "004320.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2196,
+ "file_name": "006471.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2197,
+ "file_name": "008307.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2198,
+ "file_name": "012697.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2199,
+ "file_name": "007247.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2202,
+ "file_name": "008114.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2203,
+ "file_name": "013743.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2204,
+ "file_name": "007520.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2205,
+ "file_name": "004173.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2207,
+ "file_name": "010329.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2210,
+ "file_name": "010909.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2213,
+ "file_name": "012089.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2214,
+ "file_name": "009788.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2217,
+ "file_name": "009818.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2218,
+ "file_name": "011632.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2222,
+ "file_name": "012040.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2223,
+ "file_name": "006580.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2225,
+ "file_name": "010283.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2226,
+ "file_name": "004273.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2230,
+ "file_name": "009795.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2231,
+ "file_name": "012644.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2232,
+ "file_name": "011744.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2235,
+ "file_name": "013439.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2236,
+ "file_name": "011933.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2237,
+ "file_name": "010311.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2238,
+ "file_name": "011220.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2240,
+ "file_name": "010488.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2241,
+ "file_name": "009863.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2242,
+ "file_name": "006574.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2243,
+ "file_name": "009362.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2245,
+ "file_name": "010377.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2247,
+ "file_name": "009814.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2248,
+ "file_name": "010175.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2252,
+ "file_name": "007987.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2254,
+ "file_name": "008117.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2256,
+ "file_name": "006437.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2260,
+ "file_name": "013052.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2261,
+ "file_name": "007688.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2262,
+ "file_name": "008785.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2263,
+ "file_name": "013464.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2265,
+ "file_name": "006905.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2266,
+ "file_name": "013739.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2267,
+ "file_name": "009465.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2273,
+ "file_name": "004315.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2274,
+ "file_name": "013164.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2276,
+ "file_name": "008215.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2280,
+ "file_name": "012700.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2281,
+ "file_name": "010657.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2282,
+ "file_name": "009518.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2283,
+ "file_name": "013316.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2284,
+ "file_name": "013348.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2288,
+ "file_name": "004253.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2289,
+ "file_name": "012315.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2291,
+ "file_name": "011195.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2294,
+ "file_name": "004457.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2295,
+ "file_name": "013870.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2296,
+ "file_name": "011513.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2298,
+ "file_name": "008893.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2299,
+ "file_name": "009011.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2305,
+ "file_name": "008118.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2307,
+ "file_name": "008413.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2308,
+ "file_name": "011181.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2309,
+ "file_name": "005062.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2310,
+ "file_name": "013953.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2311,
+ "file_name": "007703.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2312,
+ "file_name": "007576.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2313,
+ "file_name": "012705.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2315,
+ "file_name": "011509.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2316,
+ "file_name": "007775.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2317,
+ "file_name": "012115.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2318,
+ "file_name": "010583.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2320,
+ "file_name": "013123.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2321,
+ "file_name": "006596.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2322,
+ "file_name": "005048.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2323,
+ "file_name": "006941.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2324,
+ "file_name": "005390.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2330,
+ "file_name": "010418.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2331,
+ "file_name": "006562.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2333,
+ "file_name": "009582.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2334,
+ "file_name": "005526.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2336,
+ "file_name": "005811.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2337,
+ "file_name": "008373.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2338,
+ "file_name": "008590.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2339,
+ "file_name": "013601.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2340,
+ "file_name": "007876.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2341,
+ "file_name": "004124.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2343,
+ "file_name": "007628.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2349,
+ "file_name": "012411.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2351,
+ "file_name": "005455.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2352,
+ "file_name": "010889.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2355,
+ "file_name": "005552.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2356,
+ "file_name": "012654.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2357,
+ "file_name": "005317.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2358,
+ "file_name": "013981.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2359,
+ "file_name": "007202.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2360,
+ "file_name": "013139.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2361,
+ "file_name": "006730.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2363,
+ "file_name": "006513.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2364,
+ "file_name": "009331.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2366,
+ "file_name": "013408.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2367,
+ "file_name": "005036.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2368,
+ "file_name": "010855.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2372,
+ "file_name": "007842.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2377,
+ "file_name": "007782.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2379,
+ "file_name": "004918.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2383,
+ "file_name": "013929.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2384,
+ "file_name": "012320.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2385,
+ "file_name": "011164.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2387,
+ "file_name": "004265.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2388,
+ "file_name": "010462.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2390,
+ "file_name": "013828.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2393,
+ "file_name": "008677.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2395,
+ "file_name": "011631.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2396,
+ "file_name": "007661.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2397,
+ "file_name": "006839.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2398,
+ "file_name": "008960.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2403,
+ "file_name": "004069.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2404,
+ "file_name": "012090.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2409,
+ "file_name": "012581.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2410,
+ "file_name": "004866.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2414,
+ "file_name": "006912.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2415,
+ "file_name": "010946.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2417,
+ "file_name": "011408.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2419,
+ "file_name": "013825.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2420,
+ "file_name": "011428.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2421,
+ "file_name": "011434.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2422,
+ "file_name": "013914.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2424,
+ "file_name": "005330.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2428,
+ "file_name": "004188.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2430,
+ "file_name": "007710.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2433,
+ "file_name": "008015.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2434,
+ "file_name": "004675.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2435,
+ "file_name": "006416.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2436,
+ "file_name": "013128.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2438,
+ "file_name": "004807.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2439,
+ "file_name": "007904.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2440,
+ "file_name": "013350.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2441,
+ "file_name": "009437.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2442,
+ "file_name": "011369.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2445,
+ "file_name": "004502.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2446,
+ "file_name": "009267.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2449,
+ "file_name": "011580.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2451,
+ "file_name": "012831.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2452,
+ "file_name": "005031.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2453,
+ "file_name": "004029.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2454,
+ "file_name": "006816.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2458,
+ "file_name": "007947.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2459,
+ "file_name": "006710.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2460,
+ "file_name": "007164.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2461,
+ "file_name": "004862.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2462,
+ "file_name": "012858.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2465,
+ "file_name": "008162.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2466,
+ "file_name": "012309.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2469,
+ "file_name": "013337.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2470,
+ "file_name": "013686.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2472,
+ "file_name": "013617.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2474,
+ "file_name": "010501.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2478,
+ "file_name": "008324.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2480,
+ "file_name": "012341.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2484,
+ "file_name": "010919.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2485,
+ "file_name": "011544.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2486,
+ "file_name": "011371.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2489,
+ "file_name": "007048.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2492,
+ "file_name": "012387.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2494,
+ "file_name": "004820.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2495,
+ "file_name": "009817.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2496,
+ "file_name": "007470.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2498,
+ "file_name": "011526.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2499,
+ "file_name": "013768.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2500,
+ "file_name": "013472.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2501,
+ "file_name": "004492.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2502,
+ "file_name": "009109.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2504,
+ "file_name": "008572.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2506,
+ "file_name": "011007.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2507,
+ "file_name": "012245.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2509,
+ "file_name": "010314.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2510,
+ "file_name": "013234.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2511,
+ "file_name": "007239.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2514,
+ "file_name": "009792.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2515,
+ "file_name": "004046.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2516,
+ "file_name": "004916.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2519,
+ "file_name": "012263.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2522,
+ "file_name": "011073.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2523,
+ "file_name": "008367.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2524,
+ "file_name": "010575.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2526,
+ "file_name": "009950.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2528,
+ "file_name": "012856.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2531,
+ "file_name": "011723.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2534,
+ "file_name": "006865.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2535,
+ "file_name": "006410.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2538,
+ "file_name": "007560.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2541,
+ "file_name": "004212.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2542,
+ "file_name": "013538.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2544,
+ "file_name": "008508.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2547,
+ "file_name": "007478.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2548,
+ "file_name": "004191.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2549,
+ "file_name": "008770.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2550,
+ "file_name": "012528.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2551,
+ "file_name": "007631.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2554,
+ "file_name": "010587.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2555,
+ "file_name": "006605.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2556,
+ "file_name": "009593.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2561,
+ "file_name": "010886.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2562,
+ "file_name": "008643.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2565,
+ "file_name": "010599.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2566,
+ "file_name": "011340.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2567,
+ "file_name": "012515.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2569,
+ "file_name": "007751.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2572,
+ "file_name": "012032.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2574,
+ "file_name": "006685.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2576,
+ "file_name": "013826.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2577,
+ "file_name": "013856.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2579,
+ "file_name": "005357.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2582,
+ "file_name": "008498.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2583,
+ "file_name": "008524.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2589,
+ "file_name": "012726.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2591,
+ "file_name": "011022.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2592,
+ "file_name": "009238.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2593,
+ "file_name": "010271.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2596,
+ "file_name": "008606.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2602,
+ "file_name": "006501.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2603,
+ "file_name": "007472.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2605,
+ "file_name": "011232.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2607,
+ "file_name": "008926.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2611,
+ "file_name": "004591.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2615,
+ "file_name": "013335.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2617,
+ "file_name": "008383.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2620,
+ "file_name": "007399.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2621,
+ "file_name": "007219.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2624,
+ "file_name": "008237.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2627,
+ "file_name": "009789.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2629,
+ "file_name": "008003.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2631,
+ "file_name": "011991.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2632,
+ "file_name": "010648.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2633,
+ "file_name": "013068.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2634,
+ "file_name": "007143.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2638,
+ "file_name": "004235.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2640,
+ "file_name": "011939.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2642,
+ "file_name": "012235.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2644,
+ "file_name": "011455.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2645,
+ "file_name": "013137.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2648,
+ "file_name": "008008.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2649,
+ "file_name": "012289.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2652,
+ "file_name": "012343.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2653,
+ "file_name": "009213.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2654,
+ "file_name": "007464.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2656,
+ "file_name": "013726.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2657,
+ "file_name": "007739.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2658,
+ "file_name": "009491.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2659,
+ "file_name": "009944.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2661,
+ "file_name": "007991.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2665,
+ "file_name": "006544.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2669,
+ "file_name": "013396.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2671,
+ "file_name": "012184.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2672,
+ "file_name": "006573.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2673,
+ "file_name": "010516.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2674,
+ "file_name": "013816.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2675,
+ "file_name": "009970.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2678,
+ "file_name": "008815.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2680,
+ "file_name": "013843.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2681,
+ "file_name": "008736.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2683,
+ "file_name": "006712.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2684,
+ "file_name": "006622.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2685,
+ "file_name": "008627.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2693,
+ "file_name": "010607.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2697,
+ "file_name": "006627.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2698,
+ "file_name": "013821.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2699,
+ "file_name": "007371.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2701,
+ "file_name": "005042.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2702,
+ "file_name": "006674.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2704,
+ "file_name": "012487.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2710,
+ "file_name": "012901.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2712,
+ "file_name": "010649.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2714,
+ "file_name": "008984.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2716,
+ "file_name": "011980.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2717,
+ "file_name": "006357.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2718,
+ "file_name": "008951.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2721,
+ "file_name": "006670.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2722,
+ "file_name": "008568.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2723,
+ "file_name": "007471.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2725,
+ "file_name": "011405.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2726,
+ "file_name": "012899.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2727,
+ "file_name": "006377.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2729,
+ "file_name": "009917.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2730,
+ "file_name": "010685.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2732,
+ "file_name": "011948.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2734,
+ "file_name": "013965.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2735,
+ "file_name": "004077.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2736,
+ "file_name": "007879.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2737,
+ "file_name": "009227.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2738,
+ "file_name": "004835.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2739,
+ "file_name": "012344.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2740,
+ "file_name": "011480.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2742,
+ "file_name": "012285.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2743,
+ "file_name": "013570.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2744,
+ "file_name": "012834.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2746,
+ "file_name": "012958.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2747,
+ "file_name": "006929.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2750,
+ "file_name": "007581.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2751,
+ "file_name": "007627.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2754,
+ "file_name": "007319.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2756,
+ "file_name": "013212.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2758,
+ "file_name": "010523.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2759,
+ "file_name": "008763.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2760,
+ "file_name": "005569.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2762,
+ "file_name": "009002.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2766,
+ "file_name": "006856.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2768,
+ "file_name": "012009.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2769,
+ "file_name": "007245.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2770,
+ "file_name": "008327.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2773,
+ "file_name": "007287.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2774,
+ "file_name": "010835.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2775,
+ "file_name": "005173.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2776,
+ "file_name": "007535.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2779,
+ "file_name": "013194.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2780,
+ "file_name": "004829.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2781,
+ "file_name": "006836.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2782,
+ "file_name": "012594.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2784,
+ "file_name": "010422.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2786,
+ "file_name": "013395.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2788,
+ "file_name": "004506.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2789,
+ "file_name": "013598.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2790,
+ "file_name": "005092.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2791,
+ "file_name": "006132.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2793,
+ "file_name": "012544.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2795,
+ "file_name": "008967.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2796,
+ "file_name": "010510.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2798,
+ "file_name": "007856.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2799,
+ "file_name": "012093.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2803,
+ "file_name": "010731.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2804,
+ "file_name": "013880.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2807,
+ "file_name": "004744.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2810,
+ "file_name": "013575.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2815,
+ "file_name": "013458.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2816,
+ "file_name": "008978.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2817,
+ "file_name": "013328.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2818,
+ "file_name": "007454.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2819,
+ "file_name": "010910.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2821,
+ "file_name": "005341.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2823,
+ "file_name": "009945.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2824,
+ "file_name": "007617.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2825,
+ "file_name": "005038.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2828,
+ "file_name": "009027.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2830,
+ "file_name": "011963.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2833,
+ "file_name": "012950.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2834,
+ "file_name": "007336.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2837,
+ "file_name": "008948.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2838,
+ "file_name": "004478.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2839,
+ "file_name": "009068.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2840,
+ "file_name": "007793.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2842,
+ "file_name": "011579.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2843,
+ "file_name": "012085.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2845,
+ "file_name": "004463.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2846,
+ "file_name": "013902.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2847,
+ "file_name": "008981.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2848,
+ "file_name": "005417.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2849,
+ "file_name": "008470.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2850,
+ "file_name": "006908.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2853,
+ "file_name": "013551.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2854,
+ "file_name": "004695.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2856,
+ "file_name": "013993.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2858,
+ "file_name": "011548.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2860,
+ "file_name": "005397.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2864,
+ "file_name": "006255.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2866,
+ "file_name": "013521.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2868,
+ "file_name": "004855.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2869,
+ "file_name": "011466.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2870,
+ "file_name": "012907.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2872,
+ "file_name": "007649.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2873,
+ "file_name": "005100.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2877,
+ "file_name": "003925.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2880,
+ "file_name": "011515.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2881,
+ "file_name": "012117.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2882,
+ "file_name": "012728.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2883,
+ "file_name": "013928.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2884,
+ "file_name": "013492.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2887,
+ "file_name": "011987.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2888,
+ "file_name": "008347.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2890,
+ "file_name": "011067.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2891,
+ "file_name": "006800.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2892,
+ "file_name": "004935.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2894,
+ "file_name": "012501.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2896,
+ "file_name": "013025.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2899,
+ "file_name": "005057.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2900,
+ "file_name": "010776.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2904,
+ "file_name": "009783.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2905,
+ "file_name": "005770.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2906,
+ "file_name": "012368.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2907,
+ "file_name": "008354.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2909,
+ "file_name": "012848.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2911,
+ "file_name": "006571.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2913,
+ "file_name": "011285.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2914,
+ "file_name": "004465.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2915,
+ "file_name": "013185.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2916,
+ "file_name": "009578.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2917,
+ "file_name": "013031.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2918,
+ "file_name": "009869.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2919,
+ "file_name": "009074.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2920,
+ "file_name": "013939.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2924,
+ "file_name": "009070.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2925,
+ "file_name": "007642.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2926,
+ "file_name": "011583.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2927,
+ "file_name": "013771.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2928,
+ "file_name": "007137.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2929,
+ "file_name": "008601.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2930,
+ "file_name": "008979.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2931,
+ "file_name": "007968.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2933,
+ "file_name": "004895.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2935,
+ "file_name": "013084.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2936,
+ "file_name": "012192.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2938,
+ "file_name": "011023.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2939,
+ "file_name": "007278.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2940,
+ "file_name": "009090.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2942,
+ "file_name": "012358.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2943,
+ "file_name": "013292.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2944,
+ "file_name": "009588.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2945,
+ "file_name": "010364.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2946,
+ "file_name": "011149.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2949,
+ "file_name": "006794.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2950,
+ "file_name": "012132.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2951,
+ "file_name": "007550.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2952,
+ "file_name": "010588.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2954,
+ "file_name": "008592.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2959,
+ "file_name": "004068.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2962,
+ "file_name": "012892.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2963,
+ "file_name": "012027.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2965,
+ "file_name": "007161.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2966,
+ "file_name": "005380.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2968,
+ "file_name": "006897.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2969,
+ "file_name": "004498.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2970,
+ "file_name": "008628.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2971,
+ "file_name": "008830.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2973,
+ "file_name": "013913.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2974,
+ "file_name": "007869.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2975,
+ "file_name": "012405.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2977,
+ "file_name": "011035.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2980,
+ "file_name": "006981.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2981,
+ "file_name": "010660.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2985,
+ "file_name": "011179.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2986,
+ "file_name": "005210.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2987,
+ "file_name": "004450.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2989,
+ "file_name": "004340.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2990,
+ "file_name": "012071.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2992,
+ "file_name": "008514.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2993,
+ "file_name": "007039.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2994,
+ "file_name": "011042.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2997,
+ "file_name": "013018.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 2998,
+ "file_name": "009748.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3001,
+ "file_name": "009576.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3002,
+ "file_name": "013145.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3007,
+ "file_name": "013712.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3009,
+ "file_name": "006204.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3011,
+ "file_name": "004329.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3012,
+ "file_name": "009132.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3013,
+ "file_name": "011742.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3015,
+ "file_name": "005202.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3016,
+ "file_name": "012774.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3017,
+ "file_name": "007275.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3020,
+ "file_name": "011374.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3022,
+ "file_name": "005082.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3023,
+ "file_name": "004599.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3025,
+ "file_name": "003863.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3027,
+ "file_name": "013692.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3032,
+ "file_name": "010433.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3033,
+ "file_name": "006845.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3035,
+ "file_name": "013071.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3036,
+ "file_name": "005084.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3037,
+ "file_name": "008504.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3040,
+ "file_name": "013473.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3041,
+ "file_name": "009803.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3042,
+ "file_name": "008534.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3045,
+ "file_name": "004957.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3047,
+ "file_name": "007416.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3049,
+ "file_name": "010841.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3050,
+ "file_name": "008139.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3051,
+ "file_name": "007705.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3052,
+ "file_name": "013861.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3054,
+ "file_name": "007077.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3055,
+ "file_name": "004303.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3059,
+ "file_name": "009381.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3061,
+ "file_name": "007476.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3063,
+ "file_name": "008337.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3065,
+ "file_name": "012231.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3070,
+ "file_name": "003971.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3071,
+ "file_name": "006842.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3072,
+ "file_name": "011430.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3073,
+ "file_name": "004154.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3077,
+ "file_name": "007901.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3078,
+ "file_name": "013956.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3079,
+ "file_name": "006849.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3081,
+ "file_name": "007151.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3087,
+ "file_name": "011061.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3088,
+ "file_name": "006575.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3089,
+ "file_name": "006863.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3091,
+ "file_name": "004114.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3092,
+ "file_name": "011453.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3094,
+ "file_name": "008911.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3096,
+ "file_name": "007763.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3098,
+ "file_name": "006606.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3099,
+ "file_name": "007524.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3100,
+ "file_name": "006695.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3101,
+ "file_name": "004267.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3102,
+ "file_name": "010531.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3106,
+ "file_name": "010965.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3108,
+ "file_name": "009514.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3109,
+ "file_name": "004930.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3111,
+ "file_name": "005091.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3113,
+ "file_name": "010176.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3114,
+ "file_name": "013806.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3116,
+ "file_name": "005323.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3117,
+ "file_name": "004075.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3118,
+ "file_name": "006866.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3120,
+ "file_name": "006888.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3121,
+ "file_name": "008466.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3124,
+ "file_name": "010414.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3127,
+ "file_name": "013225.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3129,
+ "file_name": "011964.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3131,
+ "file_name": "006956.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3132,
+ "file_name": "012283.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3134,
+ "file_name": "010408.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3137,
+ "file_name": "012331.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3138,
+ "file_name": "008268.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3139,
+ "file_name": "010318.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3140,
+ "file_name": "013762.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3142,
+ "file_name": "013434.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3144,
+ "file_name": "013097.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3145,
+ "file_name": "012602.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3146,
+ "file_name": "010866.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3147,
+ "file_name": "010497.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3148,
+ "file_name": "010419.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3149,
+ "file_name": "008017.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3151,
+ "file_name": "008366.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3154,
+ "file_name": "005056.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3156,
+ "file_name": "005834.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3157,
+ "file_name": "005171.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3158,
+ "file_name": "013974.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3160,
+ "file_name": "011378.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3161,
+ "file_name": "007507.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3166,
+ "file_name": "013072.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3167,
+ "file_name": "010850.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3173,
+ "file_name": "007554.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3176,
+ "file_name": "010957.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3178,
+ "file_name": "007680.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3180,
+ "file_name": "013370.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3181,
+ "file_name": "013851.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3182,
+ "file_name": "011946.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3184,
+ "file_name": "011582.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3186,
+ "file_name": "008931.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3187,
+ "file_name": "005122.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3194,
+ "file_name": "010560.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3196,
+ "file_name": "004816.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3197,
+ "file_name": "010557.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3198,
+ "file_name": "012849.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3199,
+ "file_name": "010765.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3200,
+ "file_name": "011204.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3202,
+ "file_name": "008816.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3203,
+ "file_name": "007276.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3205,
+ "file_name": "011581.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3209,
+ "file_name": "010645.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3210,
+ "file_name": "004237.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3211,
+ "file_name": "003938.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3212,
+ "file_name": "004037.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3215,
+ "file_name": "006623.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3216,
+ "file_name": "013559.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3219,
+ "file_name": "009094.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3220,
+ "file_name": "007497.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3222,
+ "file_name": "004030.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3223,
+ "file_name": "010153.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3225,
+ "file_name": "012694.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3226,
+ "file_name": "010316.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3227,
+ "file_name": "011502.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3230,
+ "file_name": "006213.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3231,
+ "file_name": "009733.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3233,
+ "file_name": "004892.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3234,
+ "file_name": "007914.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3235,
+ "file_name": "006459.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3239,
+ "file_name": "007366.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3242,
+ "file_name": "006249.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3247,
+ "file_name": "009047.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3254,
+ "file_name": "004099.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3256,
+ "file_name": "013091.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3257,
+ "file_name": "010436.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3261,
+ "file_name": "008088.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3262,
+ "file_name": "013281.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3264,
+ "file_name": "004455.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3265,
+ "file_name": "012001.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3266,
+ "file_name": "013506.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3269,
+ "file_name": "013305.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3271,
+ "file_name": "009373.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3272,
+ "file_name": "012720.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3273,
+ "file_name": "004581.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3275,
+ "file_name": "006779.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3276,
+ "file_name": "013948.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3278,
+ "file_name": "013415.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3280,
+ "file_name": "009128.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3281,
+ "file_name": "012202.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3284,
+ "file_name": "006446.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3285,
+ "file_name": "009574.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3286,
+ "file_name": "007034.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3287,
+ "file_name": "007983.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3289,
+ "file_name": "007654.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3290,
+ "file_name": "008208.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3291,
+ "file_name": "011158.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3292,
+ "file_name": "007833.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3293,
+ "file_name": "007834.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3294,
+ "file_name": "008168.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3296,
+ "file_name": "007795.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3297,
+ "file_name": "004442.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3298,
+ "file_name": "004268.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3299,
+ "file_name": "005030.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3300,
+ "file_name": "004965.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3301,
+ "file_name": "012841.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3303,
+ "file_name": "013716.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3304,
+ "file_name": "012740.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3305,
+ "file_name": "008726.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3306,
+ "file_name": "005322.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3308,
+ "file_name": "011168.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3310,
+ "file_name": "013466.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3311,
+ "file_name": "006492.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3312,
+ "file_name": "005146.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3314,
+ "file_name": "010798.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3315,
+ "file_name": "011999.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3316,
+ "file_name": "013727.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3317,
+ "file_name": "008631.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3318,
+ "file_name": "006394.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3319,
+ "file_name": "007678.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3321,
+ "file_name": "007254.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3322,
+ "file_name": "006208.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3326,
+ "file_name": "009480.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3332,
+ "file_name": "012364.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3333,
+ "file_name": "006749.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3336,
+ "file_name": "009500.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3338,
+ "file_name": "010967.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3339,
+ "file_name": "011258.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3342,
+ "file_name": "013573.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3343,
+ "file_name": "008271.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3344,
+ "file_name": "006470.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3345,
+ "file_name": "010309.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3348,
+ "file_name": "013390.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3349,
+ "file_name": "011016.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3353,
+ "file_name": "009740.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3354,
+ "file_name": "004220.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3355,
+ "file_name": "010664.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3356,
+ "file_name": "005192.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3358,
+ "file_name": "004302.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3359,
+ "file_name": "007204.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3361,
+ "file_name": "008854.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3366,
+ "file_name": "008923.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3367,
+ "file_name": "007801.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3368,
+ "file_name": "004167.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3369,
+ "file_name": "004328.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3371,
+ "file_name": "008733.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3374,
+ "file_name": "007213.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3376,
+ "file_name": "006371.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3378,
+ "file_name": "005790.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3379,
+ "file_name": "012789.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3381,
+ "file_name": "009527.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3386,
+ "file_name": "007029.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3388,
+ "file_name": "008557.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3391,
+ "file_name": "010555.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3395,
+ "file_name": "005789.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3396,
+ "file_name": "006668.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3398,
+ "file_name": "013643.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3399,
+ "file_name": "012763.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3401,
+ "file_name": "010676.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3402,
+ "file_name": "007815.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3404,
+ "file_name": "010507.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3409,
+ "file_name": "009382.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3410,
+ "file_name": "010663.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3416,
+ "file_name": "012345.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3419,
+ "file_name": "012256.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3420,
+ "file_name": "010836.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3421,
+ "file_name": "013160.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3422,
+ "file_name": "009928.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3423,
+ "file_name": "005551.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3424,
+ "file_name": "012824.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3428,
+ "file_name": "008083.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3430,
+ "file_name": "005101.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3431,
+ "file_name": "011062.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3432,
+ "file_name": "009114.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3433,
+ "file_name": "004840.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3434,
+ "file_name": "011699.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3438,
+ "file_name": "010154.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3439,
+ "file_name": "006400.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3440,
+ "file_name": "007677.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3441,
+ "file_name": "012607.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3445,
+ "file_name": "012478.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3450,
+ "file_name": "012625.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3451,
+ "file_name": "013131.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3452,
+ "file_name": "012966.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3453,
+ "file_name": "012733.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3454,
+ "file_name": "012875.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3455,
+ "file_name": "013927.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3456,
+ "file_name": "004932.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3458,
+ "file_name": "007253.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3459,
+ "file_name": "008653.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3460,
+ "file_name": "009033.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3461,
+ "file_name": "013135.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3462,
+ "file_name": "012743.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3463,
+ "file_name": "013104.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3464,
+ "file_name": "013030.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3465,
+ "file_name": "008295.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3466,
+ "file_name": "009071.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3467,
+ "file_name": "005793.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3469,
+ "file_name": "006626.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3471,
+ "file_name": "008809.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3475,
+ "file_name": "012510.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3476,
+ "file_name": "005200.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3478,
+ "file_name": "004881.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3483,
+ "file_name": "011718.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3487,
+ "file_name": "011089.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3498,
+ "file_name": "008518.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3500,
+ "file_name": "012736.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3501,
+ "file_name": "011441.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3503,
+ "file_name": "008783.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3504,
+ "file_name": "011477.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3505,
+ "file_name": "006402.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3507,
+ "file_name": "005415.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3508,
+ "file_name": "012385.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3509,
+ "file_name": "011088.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3510,
+ "file_name": "010864.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3511,
+ "file_name": "008282.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3514,
+ "file_name": "011038.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3516,
+ "file_name": "012552.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3517,
+ "file_name": "008426.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3518,
+ "file_name": "006017.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3519,
+ "file_name": "013300.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3522,
+ "file_name": "010159.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3525,
+ "file_name": "009015.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3529,
+ "file_name": "005144.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3530,
+ "file_name": "012686.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3533,
+ "file_name": "008464.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3534,
+ "file_name": "004484.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3539,
+ "file_name": "008014.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3540,
+ "file_name": "013597.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3542,
+ "file_name": "012932.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3547,
+ "file_name": "012836.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3548,
+ "file_name": "004809.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3549,
+ "file_name": "006971.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3552,
+ "file_name": "007789.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3554,
+ "file_name": "008582.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3555,
+ "file_name": "008263.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3561,
+ "file_name": "009483.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3566,
+ "file_name": "009470.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3568,
+ "file_name": "007884.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3569,
+ "file_name": "012468.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3570,
+ "file_name": "007358.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3571,
+ "file_name": "013407.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3572,
+ "file_name": "013127.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3573,
+ "file_name": "013119.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3575,
+ "file_name": "010017.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3577,
+ "file_name": "012288.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3580,
+ "file_name": "008318.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3582,
+ "file_name": "010922.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3585,
+ "file_name": "007624.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3586,
+ "file_name": "011011.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3589,
+ "file_name": "006350.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3591,
+ "file_name": "010373.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3592,
+ "file_name": "005305.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3593,
+ "file_name": "006479.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3594,
+ "file_name": "008596.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3601,
+ "file_name": "012982.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3603,
+ "file_name": "009390.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3604,
+ "file_name": "011940.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3606,
+ "file_name": "013946.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3608,
+ "file_name": "012570.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3609,
+ "file_name": "004489.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3611,
+ "file_name": "005155.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3612,
+ "file_name": "010473.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3613,
+ "file_name": "007806.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3614,
+ "file_name": "007938.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3616,
+ "file_name": "005003.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3619,
+ "file_name": "007957.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3620,
+ "file_name": "008422.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3621,
+ "file_name": "013133.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3623,
+ "file_name": "004066.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3626,
+ "file_name": "003899.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3627,
+ "file_name": "004995.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3628,
+ "file_name": "007759.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3629,
+ "file_name": "008340.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3630,
+ "file_name": "010745.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3633,
+ "file_name": "009006.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3636,
+ "file_name": "013943.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3637,
+ "file_name": "013509.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3638,
+ "file_name": "011251.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3639,
+ "file_name": "008001.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3640,
+ "file_name": "010306.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3641,
+ "file_name": "011360.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3642,
+ "file_name": "012250.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3643,
+ "file_name": "013331.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3646,
+ "file_name": "011315.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3647,
+ "file_name": "011388.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3649,
+ "file_name": "011375.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3650,
+ "file_name": "007883.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3651,
+ "file_name": "013747.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3653,
+ "file_name": "004670.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3655,
+ "file_name": "009971.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3656,
+ "file_name": "012181.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3658,
+ "file_name": "005456.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3659,
+ "file_name": "008799.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3662,
+ "file_name": "013871.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3663,
+ "file_name": "006871.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3666,
+ "file_name": "007558.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3667,
+ "file_name": "007328.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3669,
+ "file_name": "010152.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3672,
+ "file_name": "009990.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3673,
+ "file_name": "012408.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3674,
+ "file_name": "009035.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3675,
+ "file_name": "011155.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3676,
+ "file_name": "003928.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3680,
+ "file_name": "013501.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3684,
+ "file_name": "009089.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3685,
+ "file_name": "005149.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3686,
+ "file_name": "011659.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3687,
+ "file_name": "011020.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3688,
+ "file_name": "007422.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3690,
+ "file_name": "009012.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3691,
+ "file_name": "013684.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3693,
+ "file_name": "003916.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3694,
+ "file_name": "008765.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3695,
+ "file_name": "003946.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3696,
+ "file_name": "013775.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3697,
+ "file_name": "010797.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3698,
+ "file_name": "004814.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3701,
+ "file_name": "012457.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3702,
+ "file_name": "006925.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3703,
+ "file_name": "005004.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3706,
+ "file_name": "007335.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3707,
+ "file_name": "009407.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3708,
+ "file_name": "005436.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3709,
+ "file_name": "005687.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3711,
+ "file_name": "013118.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3713,
+ "file_name": "009136.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3714,
+ "file_name": "008032.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3715,
+ "file_name": "003906.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3717,
+ "file_name": "009941.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3720,
+ "file_name": "012920.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3721,
+ "file_name": "008618.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3722,
+ "file_name": "011197.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3723,
+ "file_name": "006353.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3724,
+ "file_name": "008396.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3725,
+ "file_name": "005187.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3726,
+ "file_name": "009328.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3728,
+ "file_name": "010520.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3730,
+ "file_name": "007166.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3731,
+ "file_name": "011170.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3732,
+ "file_name": "010767.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3734,
+ "file_name": "011036.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3735,
+ "file_name": "010867.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3738,
+ "file_name": "006531.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3739,
+ "file_name": "009566.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3741,
+ "file_name": "009922.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3743,
+ "file_name": "010212.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3744,
+ "file_name": "010638.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3746,
+ "file_name": "004905.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3748,
+ "file_name": "012021.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3749,
+ "file_name": "013533.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3750,
+ "file_name": "010319.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3751,
+ "file_name": "006920.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3753,
+ "file_name": "010944.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3754,
+ "file_name": "009519.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3756,
+ "file_name": "004217.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3757,
+ "file_name": "013576.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3758,
+ "file_name": "006560.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3761,
+ "file_name": "011163.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3762,
+ "file_name": "007280.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3763,
+ "file_name": "007857.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3765,
+ "file_name": "009630.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3767,
+ "file_name": "004460.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3768,
+ "file_name": "012278.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3769,
+ "file_name": "007216.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3770,
+ "file_name": "006245.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3775,
+ "file_name": "010615.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3778,
+ "file_name": "003943.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3779,
+ "file_name": "012188.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3780,
+ "file_name": "004447.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3781,
+ "file_name": "006330.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3783,
+ "file_name": "004953.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3787,
+ "file_name": "008625.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3789,
+ "file_name": "013426.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3790,
+ "file_name": "007978.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3791,
+ "file_name": "004205.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3792,
+ "file_name": "004337.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3794,
+ "file_name": "004514.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3796,
+ "file_name": "007675.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3798,
+ "file_name": "013392.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3800,
+ "file_name": "013904.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3808,
+ "file_name": "007505.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3810,
+ "file_name": "008874.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3811,
+ "file_name": "012861.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3812,
+ "file_name": "013301.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3814,
+ "file_name": "013832.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3815,
+ "file_name": "009643.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3817,
+ "file_name": "010880.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3818,
+ "file_name": "007685.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3821,
+ "file_name": "006772.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3822,
+ "file_name": "011971.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3823,
+ "file_name": "006947.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3824,
+ "file_name": "013085.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3825,
+ "file_name": "008821.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3827,
+ "file_name": "008620.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3833,
+ "file_name": "012574.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3835,
+ "file_name": "007413.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3837,
+ "file_name": "010602.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3839,
+ "file_name": "013181.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3841,
+ "file_name": "012729.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3843,
+ "file_name": "013845.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3845,
+ "file_name": "010161.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3846,
+ "file_name": "013721.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3848,
+ "file_name": "007330.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3849,
+ "file_name": "009164.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3853,
+ "file_name": "013812.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3854,
+ "file_name": "007231.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3858,
+ "file_name": "011414.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3863,
+ "file_name": "011558.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3864,
+ "file_name": "013774.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3866,
+ "file_name": "010920.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3867,
+ "file_name": "010786.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3870,
+ "file_name": "012758.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3872,
+ "file_name": "008711.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3873,
+ "file_name": "012883.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3875,
+ "file_name": "013655.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3876,
+ "file_name": "007418.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3877,
+ "file_name": "013338.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3879,
+ "file_name": "005251.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3880,
+ "file_name": "006480.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3881,
+ "file_name": "008684.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3882,
+ "file_name": "013734.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3884,
+ "file_name": "007534.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3885,
+ "file_name": "007203.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3886,
+ "file_name": "011660.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3887,
+ "file_name": "003926.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3888,
+ "file_name": "004472.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3889,
+ "file_name": "005347.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3890,
+ "file_name": "012770.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3891,
+ "file_name": "007929.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3892,
+ "file_name": "006582.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3894,
+ "file_name": "005060.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3896,
+ "file_name": "007870.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3897,
+ "file_name": "008349.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3899,
+ "file_name": "012681.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3900,
+ "file_name": "007718.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3901,
+ "file_name": "009522.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3904,
+ "file_name": "007844.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3905,
+ "file_name": "007136.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3906,
+ "file_name": "005768.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3907,
+ "file_name": "011925.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3909,
+ "file_name": "007258.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3910,
+ "file_name": "009017.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3912,
+ "file_name": "006877.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3913,
+ "file_name": "010782.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3914,
+ "file_name": "011015.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3915,
+ "file_name": "012769.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3916,
+ "file_name": "013999.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3919,
+ "file_name": "011504.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3921,
+ "file_name": "008356.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3922,
+ "file_name": "006563.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3923,
+ "file_name": "008192.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3924,
+ "file_name": "006401.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3925,
+ "file_name": "012294.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3928,
+ "file_name": "011352.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3930,
+ "file_name": "008343.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3931,
+ "file_name": "006874.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3933,
+ "file_name": "008233.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3934,
+ "file_name": "009956.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3935,
+ "file_name": "009420.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3939,
+ "file_name": "004101.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3941,
+ "file_name": "007865.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3942,
+ "file_name": "011622.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3945,
+ "file_name": "010012.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3947,
+ "file_name": "006344.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3949,
+ "file_name": "007281.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3951,
+ "file_name": "011064.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3952,
+ "file_name": "008013.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3953,
+ "file_name": "012872.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3954,
+ "file_name": "012223.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3955,
+ "file_name": "013295.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3956,
+ "file_name": "007171.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3957,
+ "file_name": "011663.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3958,
+ "file_name": "009962.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3959,
+ "file_name": "010690.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3964,
+ "file_name": "004703.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3966,
+ "file_name": "012751.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3969,
+ "file_name": "007227.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3972,
+ "file_name": "012269.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3974,
+ "file_name": "005764.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3975,
+ "file_name": "007339.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3976,
+ "file_name": "005154.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3979,
+ "file_name": "007024.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3980,
+ "file_name": "007401.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3981,
+ "file_name": "010890.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3982,
+ "file_name": "011679.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3984,
+ "file_name": "009376.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3985,
+ "file_name": "013624.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3986,
+ "file_name": "010387.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3987,
+ "file_name": "013167.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3988,
+ "file_name": "013366.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3991,
+ "file_name": "010948.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3992,
+ "file_name": "007303.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3994,
+ "file_name": "008429.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 3996,
+ "file_name": "009492.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4000,
+ "file_name": "008738.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4002,
+ "file_name": "011384.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4003,
+ "file_name": "010002.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4005,
+ "file_name": "009165.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4008,
+ "file_name": "007799.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4009,
+ "file_name": "008666.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4012,
+ "file_name": "004925.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4017,
+ "file_name": "013445.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4021,
+ "file_name": "013663.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4023,
+ "file_name": "012002.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4024,
+ "file_name": "005191.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4027,
+ "file_name": "012790.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4028,
+ "file_name": "008578.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4029,
+ "file_name": "006584.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4030,
+ "file_name": "008698.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4031,
+ "file_name": "011692.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4032,
+ "file_name": "009957.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4036,
+ "file_name": "006512.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4037,
+ "file_name": "009556.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4038,
+ "file_name": "010198.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4042,
+ "file_name": "006488.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4043,
+ "file_name": "007934.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4044,
+ "file_name": "008768.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4045,
+ "file_name": "011717.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4046,
+ "file_name": "005147.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4047,
+ "file_name": "008748.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4048,
+ "file_name": "009345.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4049,
+ "file_name": "010758.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4050,
+ "file_name": "006123.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4051,
+ "file_name": "012791.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4052,
+ "file_name": "009787.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4053,
+ "file_name": "007847.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4059,
+ "file_name": "011571.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4060,
+ "file_name": "010192.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4061,
+ "file_name": "005024.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4062,
+ "file_name": "008016.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4064,
+ "file_name": "012779.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4069,
+ "file_name": "012951.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4070,
+ "file_name": "012145.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4071,
+ "file_name": "010367.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4072,
+ "file_name": "013054.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4074,
+ "file_name": "008613.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4079,
+ "file_name": "003850.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4080,
+ "file_name": "008379.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4081,
+ "file_name": "010019.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4085,
+ "file_name": "013557.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4087,
+ "file_name": "009636.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4088,
+ "file_name": "013909.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4089,
+ "file_name": "011983.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4092,
+ "file_name": "013233.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4093,
+ "file_name": "008415.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4094,
+ "file_name": "011054.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4097,
+ "file_name": "006608.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4101,
+ "file_name": "007138.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4107,
+ "file_name": "013193.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4108,
+ "file_name": "008461.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4109,
+ "file_name": "007170.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4113,
+ "file_name": "006742.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4115,
+ "file_name": "013422.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4117,
+ "file_name": "013493.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4118,
+ "file_name": "007532.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4119,
+ "file_name": "010160.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4121,
+ "file_name": "011973.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4122,
+ "file_name": "008160.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4123,
+ "file_name": "011464.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4124,
+ "file_name": "010906.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4130,
+ "file_name": "005152.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4132,
+ "file_name": "013412.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4133,
+ "file_name": "010490.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4135,
+ "file_name": "012360.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4136,
+ "file_name": "006609.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4137,
+ "file_name": "013026.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4138,
+ "file_name": "011056.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4140,
+ "file_name": "009049.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4142,
+ "file_name": "008275.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4147,
+ "file_name": "005217.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4149,
+ "file_name": "013903.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4151,
+ "file_name": "011157.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4152,
+ "file_name": "008206.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4154,
+ "file_name": "009263.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4155,
+ "file_name": "004204.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4158,
+ "file_name": "004554.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4159,
+ "file_name": "007430.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4162,
+ "file_name": "005466.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4167,
+ "file_name": "013206.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4168,
+ "file_name": "013313.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4170,
+ "file_name": "008161.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4171,
+ "file_name": "013180.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4176,
+ "file_name": "008314.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4179,
+ "file_name": "009416.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4180,
+ "file_name": "011260.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4181,
+ "file_name": "011335.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4183,
+ "file_name": "008659.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4184,
+ "file_name": "008171.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4185,
+ "file_name": "010357.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4187,
+ "file_name": "011695.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4189,
+ "file_name": "010737.jpg",
+ "width": 512,
+ "height": 512
+ },
+ {
+ "id": 4200,
+ "file_name": "009981.jpg",
+ "width": 512,
+ "height": 512
+ }
+ ],
+ "type": "instances",
+ "annotations": [
+ {
+ "id": 1,
+ "image_id": 1,
+ "bbox": [
+ 2,
+ 40,
+ 95,
+ 122
+ ],
+ "category_id": 6,
+ "area": 142572,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2,
+ "image_id": 1,
+ "bbox": [
+ 11,
+ 27,
+ 86,
+ 54
+ ],
+ "category_id": 6,
+ "area": 56935,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3,
+ "image_id": 1,
+ "bbox": [
+ 44,
+ 110,
+ 134,
+ 153
+ ],
+ "category_id": 6,
+ "area": 251328,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4,
+ "image_id": 1,
+ "bbox": [
+ 54,
+ 260,
+ 152,
+ 96
+ ],
+ "category_id": 6,
+ "area": 179568,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5,
+ "image_id": 1,
+ "bbox": [
+ 7,
+ 357,
+ 141,
+ 92
+ ],
+ "category_id": 6,
+ "area": 158595,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6,
+ "image_id": 1,
+ "bbox": [
+ 106,
+ 32,
+ 121,
+ 87
+ ],
+ "category_id": 6,
+ "area": 129065,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7,
+ "image_id": 1,
+ "bbox": [
+ 196,
+ 41,
+ 137,
+ 182
+ ],
+ "category_id": 6,
+ "area": 305856,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8,
+ "image_id": 1,
+ "bbox": [
+ 197,
+ 217,
+ 113,
+ 73
+ ],
+ "category_id": 6,
+ "area": 101140,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9,
+ "image_id": 1,
+ "bbox": [
+ 224,
+ 286,
+ 170,
+ 85
+ ],
+ "category_id": 6,
+ "area": 177815,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10,
+ "image_id": 1,
+ "bbox": [
+ 310,
+ 2,
+ 113,
+ 98
+ ],
+ "category_id": 6,
+ "area": 136150,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11,
+ "image_id": 1,
+ "bbox": [
+ 334,
+ 63,
+ 40,
+ 71
+ ],
+ "category_id": 6,
+ "area": 35560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12,
+ "image_id": 1,
+ "bbox": [
+ 354,
+ 111,
+ 74,
+ 66
+ ],
+ "category_id": 6,
+ "area": 60416,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13,
+ "image_id": 1,
+ "bbox": [
+ 441,
+ 113,
+ 68,
+ 85
+ ],
+ "category_id": 6,
+ "area": 71675,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14,
+ "image_id": 1,
+ "bbox": [
+ 486,
+ 304,
+ 25,
+ 68
+ ],
+ "category_id": 6,
+ "area": 21472,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18,
+ "image_id": 2,
+ "bbox": [
+ 76,
+ 130,
+ 284,
+ 236
+ ],
+ "category_id": 8,
+ "area": 531435,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19,
+ "image_id": 2,
+ "bbox": [
+ 260,
+ 140,
+ 250,
+ 362
+ ],
+ "category_id": 6,
+ "area": 719100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 41,
+ "image_id": 7,
+ "bbox": [
+ 154,
+ 300,
+ 102,
+ 110
+ ],
+ "category_id": 9,
+ "area": 17069,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 42,
+ "image_id": 7,
+ "bbox": [
+ 260,
+ 410,
+ 68,
+ 97
+ ],
+ "category_id": 6,
+ "area": 10057,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 43,
+ "image_id": 7,
+ "bbox": [
+ 232,
+ 418,
+ 29,
+ 33
+ ],
+ "category_id": 6,
+ "area": 1519,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 44,
+ "image_id": 7,
+ "bbox": [
+ 231,
+ 456,
+ 30,
+ 31
+ ],
+ "category_id": 6,
+ "area": 1479,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 45,
+ "image_id": 8,
+ "bbox": [
+ 7,
+ 117,
+ 60,
+ 62
+ ],
+ "category_id": 8,
+ "area": 5000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 46,
+ "image_id": 8,
+ "bbox": [
+ 0,
+ 266,
+ 153,
+ 136
+ ],
+ "category_id": 8,
+ "area": 27686,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 47,
+ "image_id": 8,
+ "bbox": [
+ 192,
+ 230,
+ 92,
+ 281
+ ],
+ "category_id": 8,
+ "area": 34272,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 48,
+ "image_id": 8,
+ "bbox": [
+ 316,
+ 154,
+ 86,
+ 209
+ ],
+ "category_id": 8,
+ "area": 23881,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 49,
+ "image_id": 8,
+ "bbox": [
+ 135,
+ 188,
+ 45,
+ 70
+ ],
+ "category_id": 8,
+ "area": 4200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 50,
+ "image_id": 8,
+ "bbox": [
+ 267,
+ 164,
+ 34,
+ 43
+ ],
+ "category_id": 8,
+ "area": 1995,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 51,
+ "image_id": 8,
+ "bbox": [
+ 489,
+ 321,
+ 21,
+ 77
+ ],
+ "category_id": 6,
+ "area": 2232,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 52,
+ "image_id": 9,
+ "bbox": [
+ 40,
+ 154,
+ 441,
+ 355
+ ],
+ "category_id": 6,
+ "area": 8877722,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 54,
+ "image_id": 11,
+ "bbox": [
+ 98,
+ 132,
+ 303,
+ 251
+ ],
+ "category_id": 9,
+ "area": 363912,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 99,
+ "image_id": 16,
+ "bbox": [
+ 280,
+ 59,
+ 83,
+ 123
+ ],
+ "category_id": 9,
+ "area": 36366,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 100,
+ "image_id": 16,
+ "bbox": [
+ 248,
+ 136,
+ 90,
+ 184
+ ],
+ "category_id": 9,
+ "area": 58534,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 101,
+ "image_id": 16,
+ "bbox": [
+ 212,
+ 469,
+ 66,
+ 36
+ ],
+ "category_id": 9,
+ "area": 8466,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 102,
+ "image_id": 16,
+ "bbox": [
+ 275,
+ 66,
+ 41,
+ 105
+ ],
+ "category_id": 9,
+ "area": 15244,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 103,
+ "image_id": 17,
+ "bbox": [
+ 18,
+ 13,
+ 466,
+ 492
+ ],
+ "category_id": 8,
+ "area": 213850,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 124,
+ "image_id": 19,
+ "bbox": [
+ 40,
+ 56,
+ 336,
+ 454
+ ],
+ "category_id": 8,
+ "area": 537399,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 142,
+ "image_id": 22,
+ "bbox": [
+ 122,
+ 101,
+ 205,
+ 111
+ ],
+ "category_id": 8,
+ "area": 181484,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 145,
+ "image_id": 23,
+ "bbox": [
+ 42,
+ 128,
+ 335,
+ 362
+ ],
+ "category_id": 9,
+ "area": 365024,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 148,
+ "image_id": 24,
+ "bbox": [
+ 229,
+ 8,
+ 130,
+ 156
+ ],
+ "category_id": 6,
+ "area": 15120,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 149,
+ "image_id": 24,
+ "bbox": [
+ 384,
+ 17,
+ 122,
+ 272
+ ],
+ "category_id": 6,
+ "area": 24816,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 150,
+ "image_id": 24,
+ "bbox": [
+ 241,
+ 372,
+ 41,
+ 110
+ ],
+ "category_id": 6,
+ "area": 3420,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 151,
+ "image_id": 24,
+ "bbox": [
+ 274,
+ 349,
+ 41,
+ 110
+ ],
+ "category_id": 6,
+ "area": 3420,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 152,
+ "image_id": 24,
+ "bbox": [
+ 408,
+ 284,
+ 65,
+ 114
+ ],
+ "category_id": 6,
+ "area": 5530,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 162,
+ "image_id": 26,
+ "bbox": [
+ 67,
+ 245,
+ 373,
+ 192
+ ],
+ "category_id": 9,
+ "area": 148482,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 163,
+ "image_id": 26,
+ "bbox": [
+ 91,
+ 324,
+ 419,
+ 187
+ ],
+ "category_id": 6,
+ "area": 161868,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 164,
+ "image_id": 26,
+ "bbox": [
+ 0,
+ 63,
+ 512,
+ 247
+ ],
+ "category_id": 6,
+ "area": 261000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 165,
+ "image_id": 27,
+ "bbox": [
+ 279,
+ 260,
+ 232,
+ 249
+ ],
+ "category_id": 6,
+ "area": 74844,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 166,
+ "image_id": 27,
+ "bbox": [
+ 292,
+ 215,
+ 168,
+ 221
+ ],
+ "category_id": 6,
+ "area": 48224,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 167,
+ "image_id": 27,
+ "bbox": [
+ 0,
+ 344,
+ 230,
+ 167
+ ],
+ "category_id": 8,
+ "area": 49875,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 168,
+ "image_id": 27,
+ "bbox": [
+ 17,
+ 15,
+ 303,
+ 173
+ ],
+ "category_id": 8,
+ "area": 68034,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 172,
+ "image_id": 29,
+ "bbox": [
+ 0,
+ 115,
+ 300,
+ 396
+ ],
+ "category_id": 6,
+ "area": 382228,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 173,
+ "image_id": 29,
+ "bbox": [
+ 349,
+ 432,
+ 113,
+ 79
+ ],
+ "category_id": 6,
+ "area": 28764,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 174,
+ "image_id": 29,
+ "bbox": [
+ 180,
+ 142,
+ 46,
+ 108
+ ],
+ "category_id": 8,
+ "area": 16240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 175,
+ "image_id": 29,
+ "bbox": [
+ 271,
+ 279,
+ 109,
+ 151
+ ],
+ "category_id": 8,
+ "area": 53235,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 176,
+ "image_id": 30,
+ "bbox": [
+ 1,
+ 110,
+ 222,
+ 135
+ ],
+ "category_id": 8,
+ "area": 75012,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 178,
+ "image_id": 30,
+ "bbox": [
+ 324,
+ 213,
+ 170,
+ 191
+ ],
+ "category_id": 8,
+ "area": 80825,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 191,
+ "image_id": 33,
+ "bbox": [
+ 32,
+ 160,
+ 60,
+ 60
+ ],
+ "category_id": 6,
+ "area": 4275,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 192,
+ "image_id": 33,
+ "bbox": [
+ 94,
+ 102,
+ 71,
+ 54
+ ],
+ "category_id": 6,
+ "area": 4539,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 193,
+ "image_id": 33,
+ "bbox": [
+ 136,
+ 401,
+ 88,
+ 103
+ ],
+ "category_id": 6,
+ "area": 10767,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 194,
+ "image_id": 33,
+ "bbox": [
+ 249,
+ 263,
+ 52,
+ 70
+ ],
+ "category_id": 6,
+ "area": 4356,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 195,
+ "image_id": 34,
+ "bbox": [
+ 129,
+ 241,
+ 313,
+ 241
+ ],
+ "category_id": 8,
+ "area": 110973,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 196,
+ "image_id": 34,
+ "bbox": [
+ 132,
+ 0,
+ 379,
+ 381
+ ],
+ "category_id": 8,
+ "area": 212016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 205,
+ "image_id": 37,
+ "bbox": [
+ 328,
+ 255,
+ 80,
+ 192
+ ],
+ "category_id": 6,
+ "area": 23364,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 206,
+ "image_id": 37,
+ "bbox": [
+ 81,
+ 368,
+ 85,
+ 143
+ ],
+ "category_id": 6,
+ "area": 18480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 207,
+ "image_id": 37,
+ "bbox": [
+ 0,
+ 122,
+ 195,
+ 355
+ ],
+ "category_id": 6,
+ "area": 104640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 208,
+ "image_id": 37,
+ "bbox": [
+ 122,
+ 76,
+ 336,
+ 434
+ ],
+ "category_id": 9,
+ "area": 220000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 210,
+ "image_id": 39,
+ "bbox": [
+ 96,
+ 80,
+ 328,
+ 350
+ ],
+ "category_id": 9,
+ "area": 296652,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 228,
+ "image_id": 42,
+ "bbox": [
+ 148,
+ 237,
+ 175,
+ 105
+ ],
+ "category_id": 8,
+ "area": 64824,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 229,
+ "image_id": 42,
+ "bbox": [
+ 272,
+ 303,
+ 166,
+ 188
+ ],
+ "category_id": 8,
+ "area": 109560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 244,
+ "image_id": 44,
+ "bbox": [
+ 47,
+ 0,
+ 99,
+ 224
+ ],
+ "category_id": 10,
+ "area": 36676,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 245,
+ "image_id": 44,
+ "bbox": [
+ 138,
+ 117,
+ 93,
+ 238
+ ],
+ "category_id": 10,
+ "area": 36675,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 246,
+ "image_id": 44,
+ "bbox": [
+ 228,
+ 106,
+ 154,
+ 356
+ ],
+ "category_id": 10,
+ "area": 90653,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 247,
+ "image_id": 44,
+ "bbox": [
+ 311,
+ 419,
+ 69,
+ 92
+ ],
+ "category_id": 10,
+ "area": 10527,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 248,
+ "image_id": 44,
+ "bbox": [
+ 467,
+ 71,
+ 44,
+ 187
+ ],
+ "category_id": 10,
+ "area": 13806,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 264,
+ "image_id": 46,
+ "bbox": [
+ 128,
+ 277,
+ 228,
+ 234
+ ],
+ "category_id": 8,
+ "area": 187288,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 265,
+ "image_id": 46,
+ "bbox": [
+ 222,
+ 194,
+ 168,
+ 222
+ ],
+ "category_id": 8,
+ "area": 131242,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 266,
+ "image_id": 46,
+ "bbox": [
+ 292,
+ 117,
+ 66,
+ 116
+ ],
+ "category_id": 8,
+ "area": 26895,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 267,
+ "image_id": 47,
+ "bbox": [
+ 145,
+ 86,
+ 279,
+ 269
+ ],
+ "category_id": 9,
+ "area": 264921,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 287,
+ "image_id": 49,
+ "bbox": [
+ 99,
+ 117,
+ 411,
+ 345
+ ],
+ "category_id": 9,
+ "area": 131970,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 288,
+ "image_id": 49,
+ "bbox": [
+ 86,
+ 308,
+ 58,
+ 181
+ ],
+ "category_id": 6,
+ "area": 9869,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 290,
+ "image_id": 51,
+ "bbox": [
+ 155,
+ 147,
+ 232,
+ 135
+ ],
+ "category_id": 8,
+ "area": 110200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 291,
+ "image_id": 51,
+ "bbox": [
+ 92,
+ 209,
+ 348,
+ 281
+ ],
+ "category_id": 8,
+ "area": 343174,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 292,
+ "image_id": 51,
+ "bbox": [
+ 415,
+ 284,
+ 28,
+ 38
+ ],
+ "category_id": 8,
+ "area": 3888,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 293,
+ "image_id": 52,
+ "bbox": [
+ 164,
+ 145,
+ 227,
+ 170
+ ],
+ "category_id": 8,
+ "area": 305868,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 294,
+ "image_id": 52,
+ "bbox": [
+ 0,
+ 190,
+ 202,
+ 146
+ ],
+ "category_id": 8,
+ "area": 234080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 320,
+ "image_id": 57,
+ "bbox": [
+ 241,
+ 272,
+ 110,
+ 73
+ ],
+ "category_id": 8,
+ "area": 64428,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 321,
+ "image_id": 57,
+ "bbox": [
+ 232,
+ 241,
+ 54,
+ 30
+ ],
+ "category_id": 8,
+ "area": 13184,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 322,
+ "image_id": 57,
+ "bbox": [
+ 185,
+ 271,
+ 76,
+ 96
+ ],
+ "category_id": 8,
+ "area": 58752,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 323,
+ "image_id": 57,
+ "bbox": [
+ 292,
+ 209,
+ 32,
+ 52
+ ],
+ "category_id": 6,
+ "area": 13431,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 324,
+ "image_id": 58,
+ "bbox": [
+ 0,
+ 207,
+ 511,
+ 304
+ ],
+ "category_id": 6,
+ "area": 498624,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 325,
+ "image_id": 58,
+ "bbox": [
+ 265,
+ 193,
+ 165,
+ 127
+ ],
+ "category_id": 8,
+ "area": 67568,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 326,
+ "image_id": 59,
+ "bbox": [
+ 152,
+ 71,
+ 266,
+ 375
+ ],
+ "category_id": 10,
+ "area": 255987,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 328,
+ "image_id": 60,
+ "bbox": [
+ 158,
+ 325,
+ 230,
+ 184
+ ],
+ "category_id": 10,
+ "area": 67100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 334,
+ "image_id": 62,
+ "bbox": [
+ 92,
+ 119,
+ 364,
+ 307
+ ],
+ "category_id": 8,
+ "area": 393120,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 337,
+ "image_id": 63,
+ "bbox": [
+ 214,
+ 87,
+ 214,
+ 309
+ ],
+ "category_id": 8,
+ "area": 87290,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 338,
+ "image_id": 64,
+ "bbox": [
+ 34,
+ 46,
+ 348,
+ 464
+ ],
+ "category_id": 8,
+ "area": 568763,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 351,
+ "image_id": 69,
+ "bbox": [
+ 43,
+ 51,
+ 270,
+ 353
+ ],
+ "category_id": 9,
+ "area": 199356,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 352,
+ "image_id": 69,
+ "bbox": [
+ 131,
+ 51,
+ 295,
+ 442
+ ],
+ "category_id": 9,
+ "area": 272008,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 360,
+ "image_id": 71,
+ "bbox": [
+ 186,
+ 163,
+ 40,
+ 82
+ ],
+ "category_id": 10,
+ "area": 9520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 361,
+ "image_id": 71,
+ "bbox": [
+ 71,
+ 133,
+ 282,
+ 378
+ ],
+ "category_id": 9,
+ "area": 305270,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 364,
+ "image_id": 72,
+ "bbox": [
+ 33,
+ 216,
+ 33,
+ 33
+ ],
+ "category_id": 6,
+ "area": 7952,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 365,
+ "image_id": 72,
+ "bbox": [
+ 1,
+ 222,
+ 38,
+ 57
+ ],
+ "category_id": 6,
+ "area": 15730,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 366,
+ "image_id": 72,
+ "bbox": [
+ 0,
+ 273,
+ 27,
+ 39
+ ],
+ "category_id": 6,
+ "area": 7644,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 370,
+ "image_id": 72,
+ "bbox": [
+ 18,
+ 76,
+ 14,
+ 24
+ ],
+ "category_id": 6,
+ "area": 2444,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 371,
+ "image_id": 73,
+ "bbox": [
+ 89,
+ 204,
+ 144,
+ 106
+ ],
+ "category_id": 8,
+ "area": 121184,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 372,
+ "image_id": 73,
+ "bbox": [
+ 195,
+ 194,
+ 127,
+ 81
+ ],
+ "category_id": 8,
+ "area": 81567,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 376,
+ "image_id": 75,
+ "bbox": [
+ 0,
+ 0,
+ 119,
+ 175
+ ],
+ "category_id": 6,
+ "area": 48672,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 377,
+ "image_id": 75,
+ "bbox": [
+ 260,
+ 198,
+ 215,
+ 252
+ ],
+ "category_id": 9,
+ "area": 125664,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 378,
+ "image_id": 76,
+ "bbox": [
+ 198,
+ 107,
+ 71,
+ 61
+ ],
+ "category_id": 8,
+ "area": 34840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 379,
+ "image_id": 76,
+ "bbox": [
+ 245,
+ 12,
+ 63,
+ 91
+ ],
+ "category_id": 6,
+ "area": 45504,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 380,
+ "image_id": 77,
+ "bbox": [
+ 58,
+ 118,
+ 180,
+ 239
+ ],
+ "category_id": 10,
+ "area": 40916,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 382,
+ "image_id": 77,
+ "bbox": [
+ 5,
+ 330,
+ 46,
+ 65
+ ],
+ "category_id": 10,
+ "area": 2862,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 383,
+ "image_id": 78,
+ "bbox": [
+ 227,
+ 151,
+ 225,
+ 288
+ ],
+ "category_id": 9,
+ "area": 43571,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 384,
+ "image_id": 78,
+ "bbox": [
+ 12,
+ 67,
+ 225,
+ 293
+ ],
+ "category_id": 9,
+ "area": 44270,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 388,
+ "image_id": 79,
+ "bbox": [
+ 136,
+ 101,
+ 30,
+ 34
+ ],
+ "category_id": 6,
+ "area": 3724,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 389,
+ "image_id": 79,
+ "bbox": [
+ 173,
+ 145,
+ 31,
+ 44
+ ],
+ "category_id": 6,
+ "area": 4836,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 390,
+ "image_id": 79,
+ "bbox": [
+ 159,
+ 192,
+ 41,
+ 47
+ ],
+ "category_id": 6,
+ "area": 6968,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 391,
+ "image_id": 79,
+ "bbox": [
+ 208,
+ 183,
+ 28,
+ 23
+ ],
+ "category_id": 6,
+ "area": 2376,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 392,
+ "image_id": 79,
+ "bbox": [
+ 114,
+ 462,
+ 48,
+ 49
+ ],
+ "category_id": 6,
+ "area": 8349,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 393,
+ "image_id": 79,
+ "bbox": [
+ 16,
+ 200,
+ 263,
+ 235
+ ],
+ "category_id": 8,
+ "area": 217140,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 394,
+ "image_id": 79,
+ "bbox": [
+ 194,
+ 201,
+ 281,
+ 250
+ ],
+ "category_id": 8,
+ "area": 246753,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 395,
+ "image_id": 80,
+ "bbox": [
+ 288,
+ 0,
+ 163,
+ 163
+ ],
+ "category_id": 6,
+ "area": 131097,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 396,
+ "image_id": 80,
+ "bbox": [
+ 173,
+ 182,
+ 205,
+ 281
+ ],
+ "category_id": 8,
+ "area": 281826,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 397,
+ "image_id": 81,
+ "bbox": [
+ 143,
+ 39,
+ 266,
+ 370
+ ],
+ "category_id": 8,
+ "area": 925925,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 399,
+ "image_id": 82,
+ "bbox": [
+ 180,
+ 46,
+ 319,
+ 337
+ ],
+ "category_id": 9,
+ "area": 236458,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 418,
+ "image_id": 86,
+ "bbox": [
+ 176,
+ 204,
+ 145,
+ 114
+ ],
+ "category_id": 8,
+ "area": 58443,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 419,
+ "image_id": 86,
+ "bbox": [
+ 289,
+ 230,
+ 118,
+ 89
+ ],
+ "category_id": 8,
+ "area": 37125,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 455,
+ "image_id": 91,
+ "bbox": [
+ 108,
+ 127,
+ 85,
+ 196
+ ],
+ "category_id": 8,
+ "area": 53848,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 456,
+ "image_id": 91,
+ "bbox": [
+ 308,
+ 262,
+ 84,
+ 105
+ ],
+ "category_id": 8,
+ "area": 28907,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 459,
+ "image_id": 92,
+ "bbox": [
+ 171,
+ 159,
+ 183,
+ 178
+ ],
+ "category_id": 9,
+ "area": 55290,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 460,
+ "image_id": 92,
+ "bbox": [
+ 182,
+ 137,
+ 136,
+ 92
+ ],
+ "category_id": 9,
+ "area": 21483,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 461,
+ "image_id": 93,
+ "bbox": [
+ 68,
+ 113,
+ 407,
+ 264
+ ],
+ "category_id": 9,
+ "area": 378696,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 462,
+ "image_id": 93,
+ "bbox": [
+ 232,
+ 70,
+ 242,
+ 307
+ ],
+ "category_id": 9,
+ "area": 262831,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 468,
+ "image_id": 95,
+ "bbox": [
+ 214,
+ 208,
+ 206,
+ 177
+ ],
+ "category_id": 8,
+ "area": 514932,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 483,
+ "image_id": 97,
+ "bbox": [
+ 0,
+ 176,
+ 321,
+ 173
+ ],
+ "category_id": 9,
+ "area": 77285,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 484,
+ "image_id": 98,
+ "bbox": [
+ 51,
+ 91,
+ 415,
+ 405
+ ],
+ "category_id": 8,
+ "area": 2019776,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 485,
+ "image_id": 99,
+ "bbox": [
+ 0,
+ 205,
+ 154,
+ 198
+ ],
+ "category_id": 8,
+ "area": 242182,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 486,
+ "image_id": 99,
+ "bbox": [
+ 70,
+ 23,
+ 149,
+ 106
+ ],
+ "category_id": 8,
+ "area": 126450,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 487,
+ "image_id": 99,
+ "bbox": [
+ 170,
+ 0,
+ 199,
+ 348
+ ],
+ "category_id": 8,
+ "area": 549780,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 488,
+ "image_id": 99,
+ "bbox": [
+ 293,
+ 356,
+ 101,
+ 122
+ ],
+ "category_id": 6,
+ "area": 98420,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 489,
+ "image_id": 99,
+ "bbox": [
+ 466,
+ 216,
+ 45,
+ 58
+ ],
+ "category_id": 6,
+ "area": 21204,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 495,
+ "image_id": 100,
+ "bbox": [
+ 131,
+ 298,
+ 104,
+ 77
+ ],
+ "category_id": 9,
+ "area": 12283,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 496,
+ "image_id": 100,
+ "bbox": [
+ 73,
+ 451,
+ 52,
+ 60
+ ],
+ "category_id": 6,
+ "area": 4785,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 497,
+ "image_id": 100,
+ "bbox": [
+ 136,
+ 362,
+ 33,
+ 44
+ ],
+ "category_id": 6,
+ "area": 2255,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 498,
+ "image_id": 100,
+ "bbox": [
+ 164,
+ 459,
+ 41,
+ 49
+ ],
+ "category_id": 6,
+ "area": 3060,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 499,
+ "image_id": 100,
+ "bbox": [
+ 202,
+ 446,
+ 44,
+ 53
+ ],
+ "category_id": 6,
+ "area": 3626,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 503,
+ "image_id": 102,
+ "bbox": [
+ 95,
+ 195,
+ 310,
+ 219
+ ],
+ "category_id": 9,
+ "area": 78897,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 512,
+ "image_id": 104,
+ "bbox": [
+ 41,
+ 123,
+ 64,
+ 91
+ ],
+ "category_id": 8,
+ "area": 20480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 513,
+ "image_id": 104,
+ "bbox": [
+ 202,
+ 262,
+ 109,
+ 84
+ ],
+ "category_id": 8,
+ "area": 32606,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 514,
+ "image_id": 104,
+ "bbox": [
+ 289,
+ 254,
+ 51,
+ 107
+ ],
+ "category_id": 8,
+ "area": 19479,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 539,
+ "image_id": 109,
+ "bbox": [
+ 176,
+ 284,
+ 84,
+ 104
+ ],
+ "category_id": 8,
+ "area": 30806,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 540,
+ "image_id": 109,
+ "bbox": [
+ 243,
+ 308,
+ 88,
+ 64
+ ],
+ "category_id": 8,
+ "area": 19800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 544,
+ "image_id": 110,
+ "bbox": [
+ 0,
+ 209,
+ 340,
+ 302
+ ],
+ "category_id": 6,
+ "area": 815364,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 545,
+ "image_id": 110,
+ "bbox": [
+ 189,
+ 378,
+ 127,
+ 133
+ ],
+ "category_id": 6,
+ "area": 134796,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 546,
+ "image_id": 110,
+ "bbox": [
+ 421,
+ 412,
+ 79,
+ 99
+ ],
+ "category_id": 6,
+ "area": 62580,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 548,
+ "image_id": 110,
+ "bbox": [
+ 106,
+ 100,
+ 227,
+ 265
+ ],
+ "category_id": 8,
+ "area": 476268,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 552,
+ "image_id": 111,
+ "bbox": [
+ 337,
+ 284,
+ 24,
+ 59
+ ],
+ "category_id": 6,
+ "area": 9954,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 553,
+ "image_id": 111,
+ "bbox": [
+ 364,
+ 305,
+ 14,
+ 47
+ ],
+ "category_id": 6,
+ "area": 4900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 554,
+ "image_id": 111,
+ "bbox": [
+ 322,
+ 356,
+ 21,
+ 40
+ ],
+ "category_id": 6,
+ "area": 6035,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 555,
+ "image_id": 112,
+ "bbox": [
+ 168,
+ 83,
+ 342,
+ 234
+ ],
+ "category_id": 8,
+ "area": 76824,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 556,
+ "image_id": 112,
+ "bbox": [
+ 0,
+ 85,
+ 310,
+ 348
+ ],
+ "category_id": 8,
+ "area": 103840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 557,
+ "image_id": 113,
+ "bbox": [
+ 105,
+ 76,
+ 217,
+ 377
+ ],
+ "category_id": 9,
+ "area": 387353,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 559,
+ "image_id": 114,
+ "bbox": [
+ 183,
+ 72,
+ 315,
+ 351
+ ],
+ "category_id": 9,
+ "area": 384426,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 632,
+ "image_id": 118,
+ "bbox": [
+ 0,
+ 0,
+ 511,
+ 249
+ ],
+ "category_id": 6,
+ "area": 304320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 633,
+ "image_id": 119,
+ "bbox": [
+ 0,
+ 230,
+ 18,
+ 141
+ ],
+ "category_id": 6,
+ "area": 5760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 634,
+ "image_id": 119,
+ "bbox": [
+ 197,
+ 300,
+ 307,
+ 211
+ ],
+ "category_id": 9,
+ "area": 145770,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 637,
+ "image_id": 120,
+ "bbox": [
+ 22,
+ 444,
+ 111,
+ 67
+ ],
+ "category_id": 6,
+ "area": 23842,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 638,
+ "image_id": 120,
+ "bbox": [
+ 178,
+ 427,
+ 131,
+ 84
+ ],
+ "category_id": 6,
+ "area": 35226,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 639,
+ "image_id": 120,
+ "bbox": [
+ 263,
+ 422,
+ 64,
+ 89
+ ],
+ "category_id": 6,
+ "area": 18271,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 655,
+ "image_id": 121,
+ "bbox": [
+ 166,
+ 154,
+ 289,
+ 317
+ ],
+ "category_id": 6,
+ "area": 382467,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 656,
+ "image_id": 121,
+ "bbox": [
+ 89,
+ 192,
+ 100,
+ 139
+ ],
+ "category_id": 6,
+ "area": 58000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 657,
+ "image_id": 121,
+ "bbox": [
+ 323,
+ 102,
+ 132,
+ 130
+ ],
+ "category_id": 6,
+ "area": 71827,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 658,
+ "image_id": 121,
+ "bbox": [
+ 16,
+ 217,
+ 86,
+ 120
+ ],
+ "category_id": 6,
+ "area": 43416,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 659,
+ "image_id": 122,
+ "bbox": [
+ 2,
+ 113,
+ 281,
+ 203
+ ],
+ "category_id": 9,
+ "area": 201058,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 660,
+ "image_id": 122,
+ "bbox": [
+ 247,
+ 196,
+ 111,
+ 144
+ ],
+ "category_id": 9,
+ "area": 56434,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 661,
+ "image_id": 122,
+ "bbox": [
+ 212,
+ 305,
+ 154,
+ 187
+ ],
+ "category_id": 9,
+ "area": 101255,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 662,
+ "image_id": 123,
+ "bbox": [
+ 2,
+ 37,
+ 506,
+ 469
+ ],
+ "category_id": 9,
+ "area": 835560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 676,
+ "image_id": 128,
+ "bbox": [
+ 158,
+ 116,
+ 101,
+ 211
+ ],
+ "category_id": 10,
+ "area": 27720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 677,
+ "image_id": 128,
+ "bbox": [
+ 210,
+ 154,
+ 243,
+ 352
+ ],
+ "category_id": 9,
+ "area": 110880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 678,
+ "image_id": 129,
+ "bbox": [
+ 316,
+ 336,
+ 194,
+ 175
+ ],
+ "category_id": 8,
+ "area": 119310,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 679,
+ "image_id": 129,
+ "bbox": [
+ 0,
+ 2,
+ 319,
+ 507
+ ],
+ "category_id": 8,
+ "area": 568974,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 680,
+ "image_id": 129,
+ "bbox": [
+ 212,
+ 44,
+ 299,
+ 336
+ ],
+ "category_id": 8,
+ "area": 353056,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 681,
+ "image_id": 130,
+ "bbox": [
+ 0,
+ 0,
+ 500,
+ 325
+ ],
+ "category_id": 8,
+ "area": 1256112,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 682,
+ "image_id": 130,
+ "bbox": [
+ 66,
+ 64,
+ 134,
+ 185
+ ],
+ "category_id": 8,
+ "area": 192528,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 683,
+ "image_id": 130,
+ "bbox": [
+ 101,
+ 252,
+ 307,
+ 217
+ ],
+ "category_id": 8,
+ "area": 517248,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 699,
+ "image_id": 134,
+ "bbox": [
+ 3,
+ 59,
+ 310,
+ 232
+ ],
+ "category_id": 9,
+ "area": 192045,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 701,
+ "image_id": 135,
+ "bbox": [
+ 96,
+ 267,
+ 188,
+ 114
+ ],
+ "category_id": 8,
+ "area": 144818,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 702,
+ "image_id": 135,
+ "bbox": [
+ 177,
+ 34,
+ 204,
+ 415
+ ],
+ "category_id": 8,
+ "area": 572236,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 703,
+ "image_id": 135,
+ "bbox": [
+ 263,
+ 26,
+ 162,
+ 256
+ ],
+ "category_id": 8,
+ "area": 280578,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 704,
+ "image_id": 136,
+ "bbox": [
+ 52,
+ 56,
+ 438,
+ 394
+ ],
+ "category_id": 8,
+ "area": 666523,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 706,
+ "image_id": 138,
+ "bbox": [
+ 50,
+ 131,
+ 255,
+ 172
+ ],
+ "category_id": 8,
+ "area": 153758,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 707,
+ "image_id": 138,
+ "bbox": [
+ 31,
+ 175,
+ 325,
+ 282
+ ],
+ "category_id": 8,
+ "area": 321552,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 708,
+ "image_id": 138,
+ "bbox": [
+ 356,
+ 265,
+ 56,
+ 71
+ ],
+ "category_id": 8,
+ "area": 14100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 709,
+ "image_id": 139,
+ "bbox": [
+ 0,
+ 1,
+ 444,
+ 510
+ ],
+ "category_id": 6,
+ "area": 2190438,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 711,
+ "image_id": 140,
+ "bbox": [
+ 144,
+ 147,
+ 233,
+ 172
+ ],
+ "category_id": 8,
+ "area": 141328,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 712,
+ "image_id": 140,
+ "bbox": [
+ 254,
+ 12,
+ 257,
+ 203
+ ],
+ "category_id": 8,
+ "area": 183255,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 714,
+ "image_id": 141,
+ "bbox": [
+ 197,
+ 311,
+ 294,
+ 199
+ ],
+ "category_id": 9,
+ "area": 77572,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 720,
+ "image_id": 143,
+ "bbox": [
+ 0,
+ 236,
+ 154,
+ 182
+ ],
+ "category_id": 8,
+ "area": 98175,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 721,
+ "image_id": 143,
+ "bbox": [
+ 260,
+ 255,
+ 173,
+ 107
+ ],
+ "category_id": 8,
+ "area": 65534,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 759,
+ "image_id": 148,
+ "bbox": [
+ 25,
+ 29,
+ 361,
+ 482
+ ],
+ "category_id": 8,
+ "area": 612234,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 774,
+ "image_id": 150,
+ "bbox": [
+ 0,
+ 171,
+ 213,
+ 299
+ ],
+ "category_id": 6,
+ "area": 471600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 775,
+ "image_id": 150,
+ "bbox": [
+ 149,
+ 18,
+ 166,
+ 489
+ ],
+ "category_id": 6,
+ "area": 600740,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 789,
+ "image_id": 152,
+ "bbox": [
+ 172,
+ 287,
+ 237,
+ 177
+ ],
+ "category_id": 10,
+ "area": 66780,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 796,
+ "image_id": 155,
+ "bbox": [
+ 197,
+ 232,
+ 240,
+ 189
+ ],
+ "category_id": 9,
+ "area": 116334,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 837,
+ "image_id": 159,
+ "bbox": [
+ 222,
+ 224,
+ 131,
+ 137
+ ],
+ "category_id": 8,
+ "area": 63497,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 858,
+ "image_id": 161,
+ "bbox": [
+ 79,
+ 0,
+ 73,
+ 100
+ ],
+ "category_id": 6,
+ "area": 8648,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 859,
+ "image_id": 162,
+ "bbox": [
+ 28,
+ 34,
+ 359,
+ 475
+ ],
+ "category_id": 8,
+ "area": 600762,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 870,
+ "image_id": 164,
+ "bbox": [
+ 95,
+ 322,
+ 220,
+ 188
+ ],
+ "category_id": 9,
+ "area": 83430,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 872,
+ "image_id": 165,
+ "bbox": [
+ 144,
+ 21,
+ 367,
+ 266
+ ],
+ "category_id": 9,
+ "area": 70005,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 896,
+ "image_id": 168,
+ "bbox": [
+ 0,
+ 6,
+ 285,
+ 301
+ ],
+ "category_id": 6,
+ "area": 216800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 897,
+ "image_id": 168,
+ "bbox": [
+ 335,
+ 229,
+ 76,
+ 78
+ ],
+ "category_id": 6,
+ "area": 15184,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 898,
+ "image_id": 168,
+ "bbox": [
+ 423,
+ 364,
+ 45,
+ 39
+ ],
+ "category_id": 6,
+ "area": 4558,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 901,
+ "image_id": 169,
+ "bbox": [
+ 169,
+ 131,
+ 157,
+ 90
+ ],
+ "category_id": 8,
+ "area": 112690,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 902,
+ "image_id": 169,
+ "bbox": [
+ 158,
+ 165,
+ 64,
+ 48
+ ],
+ "category_id": 8,
+ "area": 24786,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 903,
+ "image_id": 170,
+ "bbox": [
+ 153,
+ 207,
+ 89,
+ 114
+ ],
+ "category_id": 8,
+ "area": 35903,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 904,
+ "image_id": 170,
+ "bbox": [
+ 364,
+ 152,
+ 110,
+ 120
+ ],
+ "category_id": 8,
+ "area": 46475,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 905,
+ "image_id": 170,
+ "bbox": [
+ 464,
+ 344,
+ 33,
+ 32
+ ],
+ "category_id": 6,
+ "area": 3735,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 906,
+ "image_id": 171,
+ "bbox": [
+ 195,
+ 192,
+ 313,
+ 197
+ ],
+ "category_id": 9,
+ "area": 217952,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 924,
+ "image_id": 175,
+ "bbox": [
+ 0,
+ 101,
+ 169,
+ 110
+ ],
+ "category_id": 8,
+ "area": 264264,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 925,
+ "image_id": 175,
+ "bbox": [
+ 188,
+ 67,
+ 252,
+ 153
+ ],
+ "category_id": 8,
+ "area": 544320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 963,
+ "image_id": 179,
+ "bbox": [
+ 160,
+ 262,
+ 151,
+ 206
+ ],
+ "category_id": 9,
+ "area": 46930,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 988,
+ "image_id": 181,
+ "bbox": [
+ 91,
+ 69,
+ 108,
+ 86
+ ],
+ "category_id": 8,
+ "area": 72495,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 989,
+ "image_id": 181,
+ "bbox": [
+ 0,
+ 7,
+ 498,
+ 325
+ ],
+ "category_id": 8,
+ "area": 1251415,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 990,
+ "image_id": 181,
+ "bbox": [
+ 103,
+ 254,
+ 304,
+ 215
+ ],
+ "category_id": 8,
+ "area": 506160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1038,
+ "image_id": 187,
+ "bbox": [
+ 137,
+ 0,
+ 288,
+ 512
+ ],
+ "category_id": 6,
+ "area": 177630,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1039,
+ "image_id": 188,
+ "bbox": [
+ 65,
+ 82,
+ 222,
+ 281
+ ],
+ "category_id": 8,
+ "area": 101324,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1040,
+ "image_id": 188,
+ "bbox": [
+ 72,
+ 130,
+ 326,
+ 355
+ ],
+ "category_id": 8,
+ "area": 188190,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1041,
+ "image_id": 188,
+ "bbox": [
+ 336,
+ 193,
+ 95,
+ 100
+ ],
+ "category_id": 8,
+ "area": 15496,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1044,
+ "image_id": 189,
+ "bbox": [
+ 170,
+ 176,
+ 32,
+ 80
+ ],
+ "category_id": 6,
+ "area": 17928,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1045,
+ "image_id": 189,
+ "bbox": [
+ 249,
+ 220,
+ 21,
+ 60
+ ],
+ "category_id": 6,
+ "area": 8804,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1046,
+ "image_id": 189,
+ "bbox": [
+ 265,
+ 238,
+ 14,
+ 41
+ ],
+ "category_id": 6,
+ "area": 4128,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1047,
+ "image_id": 189,
+ "bbox": [
+ 257,
+ 424,
+ 34,
+ 66
+ ],
+ "category_id": 6,
+ "area": 15618,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1048,
+ "image_id": 189,
+ "bbox": [
+ 296,
+ 379,
+ 26,
+ 39
+ ],
+ "category_id": 6,
+ "area": 7298,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1049,
+ "image_id": 189,
+ "bbox": [
+ 327,
+ 401,
+ 18,
+ 32
+ ],
+ "category_id": 6,
+ "area": 4087,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1050,
+ "image_id": 189,
+ "bbox": [
+ 41,
+ 28,
+ 19,
+ 52
+ ],
+ "category_id": 6,
+ "area": 7085,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1051,
+ "image_id": 189,
+ "bbox": [
+ 96,
+ 21,
+ 19,
+ 40
+ ],
+ "category_id": 6,
+ "area": 5395,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1052,
+ "image_id": 189,
+ "bbox": [
+ 329,
+ 360,
+ 16,
+ 36
+ ],
+ "category_id": 6,
+ "area": 4125,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1053,
+ "image_id": 189,
+ "bbox": [
+ 347,
+ 314,
+ 22,
+ 57
+ ],
+ "category_id": 6,
+ "area": 8614,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1054,
+ "image_id": 189,
+ "bbox": [
+ 365,
+ 365,
+ 14,
+ 29
+ ],
+ "category_id": 6,
+ "area": 2867,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1055,
+ "image_id": 189,
+ "bbox": [
+ 329,
+ 476,
+ 26,
+ 35
+ ],
+ "category_id": 6,
+ "area": 6586,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1056,
+ "image_id": 190,
+ "bbox": [
+ 47,
+ 131,
+ 9,
+ 53
+ ],
+ "category_id": 6,
+ "area": 1311,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1057,
+ "image_id": 190,
+ "bbox": [
+ 118,
+ 124,
+ 330,
+ 316
+ ],
+ "category_id": 9,
+ "area": 260760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1058,
+ "image_id": 190,
+ "bbox": [
+ 426,
+ 68,
+ 85,
+ 105
+ ],
+ "category_id": 6,
+ "area": 22440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1059,
+ "image_id": 190,
+ "bbox": [
+ 252,
+ 65,
+ 117,
+ 111
+ ],
+ "category_id": 6,
+ "area": 32770,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1099,
+ "image_id": 196,
+ "bbox": [
+ 150,
+ 179,
+ 333,
+ 332
+ ],
+ "category_id": 8,
+ "area": 128650,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1105,
+ "image_id": 198,
+ "bbox": [
+ 76,
+ 342,
+ 230,
+ 119
+ ],
+ "category_id": 8,
+ "area": 35625,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1106,
+ "image_id": 198,
+ "bbox": [
+ 0,
+ 6,
+ 298,
+ 182
+ ],
+ "category_id": 8,
+ "area": 70325,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1107,
+ "image_id": 198,
+ "bbox": [
+ 285,
+ 271,
+ 226,
+ 236
+ ],
+ "category_id": 6,
+ "area": 69372,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1108,
+ "image_id": 198,
+ "bbox": [
+ 293,
+ 215,
+ 168,
+ 206
+ ],
+ "category_id": 6,
+ "area": 44936,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1125,
+ "image_id": 200,
+ "bbox": [
+ 144,
+ 101,
+ 197,
+ 256
+ ],
+ "category_id": 8,
+ "area": 342620,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1126,
+ "image_id": 200,
+ "bbox": [
+ 430,
+ 308,
+ 80,
+ 89
+ ],
+ "category_id": 8,
+ "area": 48300,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1129,
+ "image_id": 202,
+ "bbox": [
+ 234,
+ 65,
+ 174,
+ 218
+ ],
+ "category_id": 8,
+ "area": 44268,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1130,
+ "image_id": 203,
+ "bbox": [
+ 143,
+ 50,
+ 313,
+ 439
+ ],
+ "category_id": 8,
+ "area": 203320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1131,
+ "image_id": 203,
+ "bbox": [
+ 0,
+ 43,
+ 76,
+ 277
+ ],
+ "category_id": 6,
+ "area": 31369,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1132,
+ "image_id": 203,
+ "bbox": [
+ 0,
+ 292,
+ 78,
+ 178
+ ],
+ "category_id": 6,
+ "area": 20829,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1133,
+ "image_id": 203,
+ "bbox": [
+ 37,
+ 60,
+ 112,
+ 276
+ ],
+ "category_id": 6,
+ "area": 45756,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1134,
+ "image_id": 203,
+ "bbox": [
+ 297,
+ 0,
+ 129,
+ 255
+ ],
+ "category_id": 6,
+ "area": 48805,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1135,
+ "image_id": 203,
+ "bbox": [
+ 411,
+ 2,
+ 90,
+ 230
+ ],
+ "category_id": 6,
+ "area": 30955,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1136,
+ "image_id": 204,
+ "bbox": [
+ 0,
+ 134,
+ 363,
+ 377
+ ],
+ "category_id": 6,
+ "area": 441640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1137,
+ "image_id": 204,
+ "bbox": [
+ 308,
+ 443,
+ 199,
+ 68
+ ],
+ "category_id": 6,
+ "area": 44233,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1138,
+ "image_id": 204,
+ "bbox": [
+ 164,
+ 268,
+ 50,
+ 169
+ ],
+ "category_id": 8,
+ "area": 27375,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1157,
+ "image_id": 209,
+ "bbox": [
+ 432,
+ 148,
+ 79,
+ 115
+ ],
+ "category_id": 6,
+ "area": 12495,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1158,
+ "image_id": 209,
+ "bbox": [
+ 100,
+ 106,
+ 341,
+ 404
+ ],
+ "category_id": 9,
+ "area": 189666,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1159,
+ "image_id": 209,
+ "bbox": [
+ 125,
+ 260,
+ 42,
+ 65
+ ],
+ "category_id": 6,
+ "area": 3840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1160,
+ "image_id": 209,
+ "bbox": [
+ 342,
+ 241,
+ 169,
+ 220
+ ],
+ "category_id": 6,
+ "area": 51255,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1161,
+ "image_id": 210,
+ "bbox": [
+ 105,
+ 115,
+ 236,
+ 332
+ ],
+ "category_id": 9,
+ "area": 237446,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1165,
+ "image_id": 211,
+ "bbox": [
+ 0,
+ 102,
+ 369,
+ 250
+ ],
+ "category_id": 6,
+ "area": 224900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1166,
+ "image_id": 212,
+ "bbox": [
+ 165,
+ 218,
+ 93,
+ 108
+ ],
+ "category_id": 8,
+ "area": 35568,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1167,
+ "image_id": 212,
+ "bbox": [
+ 370,
+ 175,
+ 104,
+ 124
+ ],
+ "category_id": 8,
+ "area": 45850,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1174,
+ "image_id": 214,
+ "bbox": [
+ 114,
+ 235,
+ 156,
+ 159
+ ],
+ "category_id": 8,
+ "area": 87584,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1175,
+ "image_id": 214,
+ "bbox": [
+ 221,
+ 231,
+ 194,
+ 160
+ ],
+ "category_id": 8,
+ "area": 109350,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1184,
+ "image_id": 217,
+ "bbox": [
+ 0,
+ 341,
+ 101,
+ 165
+ ],
+ "category_id": 6,
+ "area": 88682,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1185,
+ "image_id": 217,
+ "bbox": [
+ 306,
+ 298,
+ 96,
+ 102
+ ],
+ "category_id": 6,
+ "area": 52116,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1186,
+ "image_id": 217,
+ "bbox": [
+ 380,
+ 0,
+ 130,
+ 383
+ ],
+ "category_id": 6,
+ "area": 265095,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1187,
+ "image_id": 217,
+ "bbox": [
+ 50,
+ 113,
+ 326,
+ 214
+ ],
+ "category_id": 8,
+ "area": 371108,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1196,
+ "image_id": 220,
+ "bbox": [
+ 321,
+ 316,
+ 188,
+ 113
+ ],
+ "category_id": 9,
+ "area": 25839,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1197,
+ "image_id": 221,
+ "bbox": [
+ 84,
+ 158,
+ 90,
+ 93
+ ],
+ "category_id": 6,
+ "area": 29606,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1198,
+ "image_id": 221,
+ "bbox": [
+ 84,
+ 219,
+ 91,
+ 108
+ ],
+ "category_id": 6,
+ "area": 35037,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1199,
+ "image_id": 221,
+ "bbox": [
+ 91,
+ 328,
+ 46,
+ 63
+ ],
+ "category_id": 6,
+ "area": 10413,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1200,
+ "image_id": 221,
+ "bbox": [
+ 285,
+ 361,
+ 83,
+ 83
+ ],
+ "category_id": 6,
+ "area": 24544,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1201,
+ "image_id": 221,
+ "bbox": [
+ 177,
+ 448,
+ 99,
+ 63
+ ],
+ "category_id": 6,
+ "area": 22161,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1202,
+ "image_id": 221,
+ "bbox": [
+ 276,
+ 412,
+ 108,
+ 98
+ ],
+ "category_id": 6,
+ "area": 37808,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1203,
+ "image_id": 221,
+ "bbox": [
+ 366,
+ 377,
+ 145,
+ 133
+ ],
+ "category_id": 6,
+ "area": 68432,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1204,
+ "image_id": 221,
+ "bbox": [
+ 210,
+ 297,
+ 39,
+ 56
+ ],
+ "category_id": 6,
+ "area": 7821,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1206,
+ "image_id": 221,
+ "bbox": [
+ 212,
+ 192,
+ 114,
+ 89
+ ],
+ "category_id": 8,
+ "area": 36036,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1207,
+ "image_id": 221,
+ "bbox": [
+ 368,
+ 243,
+ 40,
+ 59
+ ],
+ "category_id": 6,
+ "area": 8383,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1210,
+ "image_id": 222,
+ "bbox": [
+ 226,
+ 154,
+ 158,
+ 241
+ ],
+ "category_id": 8,
+ "area": 133510,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1273,
+ "image_id": 228,
+ "bbox": [
+ 33,
+ 196,
+ 324,
+ 274
+ ],
+ "category_id": 8,
+ "area": 311850,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1274,
+ "image_id": 228,
+ "bbox": [
+ 40,
+ 157,
+ 283,
+ 231
+ ],
+ "category_id": 8,
+ "area": 229716,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1275,
+ "image_id": 228,
+ "bbox": [
+ 360,
+ 277,
+ 57,
+ 79
+ ],
+ "category_id": 8,
+ "area": 15984,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1283,
+ "image_id": 231,
+ "bbox": [
+ 151,
+ 155,
+ 226,
+ 161
+ ],
+ "category_id": 8,
+ "area": 88032,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1285,
+ "image_id": 231,
+ "bbox": [
+ 286,
+ 180,
+ 220,
+ 108
+ ],
+ "category_id": 8,
+ "area": 57450,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1286,
+ "image_id": 232,
+ "bbox": [
+ 129,
+ 148,
+ 277,
+ 185
+ ],
+ "category_id": 9,
+ "area": 47450,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1287,
+ "image_id": 233,
+ "bbox": [
+ 0,
+ 9,
+ 512,
+ 288
+ ],
+ "category_id": 6,
+ "area": 369608,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1295,
+ "image_id": 235,
+ "bbox": [
+ 3,
+ 83,
+ 506,
+ 425
+ ],
+ "category_id": 9,
+ "area": 757068,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1296,
+ "image_id": 235,
+ "bbox": [
+ 28,
+ 2,
+ 450,
+ 373
+ ],
+ "category_id": 9,
+ "area": 590625,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1297,
+ "image_id": 235,
+ "bbox": [
+ 14,
+ 130,
+ 449,
+ 378
+ ],
+ "category_id": 9,
+ "area": 597436,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1298,
+ "image_id": 236,
+ "bbox": [
+ 9,
+ 204,
+ 232,
+ 260
+ ],
+ "category_id": 9,
+ "area": 35866,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1328,
+ "image_id": 238,
+ "bbox": [
+ 252,
+ 218,
+ 208,
+ 275
+ ],
+ "category_id": 10,
+ "area": 28985,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1336,
+ "image_id": 239,
+ "bbox": [
+ 183,
+ 76,
+ 37,
+ 64
+ ],
+ "category_id": 6,
+ "area": 2928,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1337,
+ "image_id": 239,
+ "bbox": [
+ 63,
+ 127,
+ 40,
+ 115
+ ],
+ "category_id": 6,
+ "area": 5676,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1338,
+ "image_id": 239,
+ "bbox": [
+ 63,
+ 247,
+ 88,
+ 230
+ ],
+ "category_id": 6,
+ "area": 24424,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1339,
+ "image_id": 239,
+ "bbox": [
+ 255,
+ 92,
+ 36,
+ 48
+ ],
+ "category_id": 6,
+ "area": 2124,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1340,
+ "image_id": 239,
+ "bbox": [
+ 232,
+ 274,
+ 152,
+ 237
+ ],
+ "category_id": 6,
+ "area": 43365,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1341,
+ "image_id": 240,
+ "bbox": [
+ 38,
+ 169,
+ 432,
+ 200
+ ],
+ "category_id": 8,
+ "area": 317382,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1342,
+ "image_id": 240,
+ "bbox": [
+ 266,
+ 435,
+ 159,
+ 76
+ ],
+ "category_id": 6,
+ "area": 44506,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1343,
+ "image_id": 240,
+ "bbox": [
+ 200,
+ 329,
+ 250,
+ 180
+ ],
+ "category_id": 6,
+ "area": 165816,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1344,
+ "image_id": 240,
+ "bbox": [
+ 47,
+ 456,
+ 90,
+ 55
+ ],
+ "category_id": 6,
+ "area": 18146,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1345,
+ "image_id": 240,
+ "bbox": [
+ 0,
+ 267,
+ 27,
+ 90
+ ],
+ "category_id": 6,
+ "area": 9024,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1346,
+ "image_id": 241,
+ "bbox": [
+ 163,
+ 255,
+ 116,
+ 213
+ ],
+ "category_id": 10,
+ "area": 43416,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1347,
+ "image_id": 241,
+ "bbox": [
+ 288,
+ 0,
+ 120,
+ 205
+ ],
+ "category_id": 10,
+ "area": 43232,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1348,
+ "image_id": 241,
+ "bbox": [
+ 34,
+ 279,
+ 40,
+ 59
+ ],
+ "category_id": 10,
+ "area": 4200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1349,
+ "image_id": 241,
+ "bbox": [
+ 18,
+ 0,
+ 75,
+ 62
+ ],
+ "category_id": 10,
+ "area": 8260,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1350,
+ "image_id": 242,
+ "bbox": [
+ 203,
+ 198,
+ 154,
+ 204
+ ],
+ "category_id": 6,
+ "area": 45400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1351,
+ "image_id": 242,
+ "bbox": [
+ 217,
+ 387,
+ 213,
+ 123
+ ],
+ "category_id": 9,
+ "area": 37994,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1353,
+ "image_id": 243,
+ "bbox": [
+ 221,
+ 172,
+ 121,
+ 218
+ ],
+ "category_id": 9,
+ "area": 79376,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1359,
+ "image_id": 246,
+ "bbox": [
+ 0,
+ 154,
+ 499,
+ 357
+ ],
+ "category_id": 9,
+ "area": 325268,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1360,
+ "image_id": 246,
+ "bbox": [
+ 165,
+ 9,
+ 346,
+ 500
+ ],
+ "category_id": 9,
+ "area": 315568,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1367,
+ "image_id": 247,
+ "bbox": [
+ 142,
+ 139,
+ 272,
+ 335
+ ],
+ "category_id": 6,
+ "area": 167418,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1378,
+ "image_id": 251,
+ "bbox": [
+ 158,
+ 332,
+ 162,
+ 179
+ ],
+ "category_id": 6,
+ "area": 18297,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1379,
+ "image_id": 251,
+ "bbox": [
+ 0,
+ 120,
+ 256,
+ 391
+ ],
+ "category_id": 6,
+ "area": 63180,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1406,
+ "image_id": 254,
+ "bbox": [
+ 0,
+ 100,
+ 322,
+ 411
+ ],
+ "category_id": 8,
+ "area": 464712,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1407,
+ "image_id": 254,
+ "bbox": [
+ 183,
+ 112,
+ 328,
+ 366
+ ],
+ "category_id": 8,
+ "area": 421480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1415,
+ "image_id": 256,
+ "bbox": [
+ 146,
+ 186,
+ 80,
+ 143
+ ],
+ "category_id": 6,
+ "area": 84600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1416,
+ "image_id": 257,
+ "bbox": [
+ 180,
+ 93,
+ 105,
+ 284
+ ],
+ "category_id": 8,
+ "area": 237395,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1417,
+ "image_id": 257,
+ "bbox": [
+ 293,
+ 15,
+ 218,
+ 179
+ ],
+ "category_id": 8,
+ "area": 309204,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1421,
+ "image_id": 260,
+ "bbox": [
+ 0,
+ 88,
+ 64,
+ 287
+ ],
+ "category_id": 6,
+ "area": 41899,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1422,
+ "image_id": 260,
+ "bbox": [
+ 0,
+ 385,
+ 14,
+ 56
+ ],
+ "category_id": 6,
+ "area": 1881,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1423,
+ "image_id": 260,
+ "bbox": [
+ 190,
+ 284,
+ 285,
+ 227
+ ],
+ "category_id": 9,
+ "area": 145928,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1426,
+ "image_id": 261,
+ "bbox": [
+ 243,
+ 262,
+ 125,
+ 194
+ ],
+ "category_id": 10,
+ "area": 49486,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1427,
+ "image_id": 261,
+ "bbox": [
+ 291,
+ 95,
+ 184,
+ 239
+ ],
+ "category_id": 9,
+ "area": 89512,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1428,
+ "image_id": 261,
+ "bbox": [
+ 28,
+ 40,
+ 483,
+ 470
+ ],
+ "category_id": 6,
+ "area": 460250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1436,
+ "image_id": 263,
+ "bbox": [
+ 0,
+ 62,
+ 319,
+ 384
+ ],
+ "category_id": 9,
+ "area": 305772,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1437,
+ "image_id": 263,
+ "bbox": [
+ 317,
+ 129,
+ 16,
+ 87
+ ],
+ "category_id": 6,
+ "area": 3616,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1438,
+ "image_id": 263,
+ "bbox": [
+ 445,
+ 284,
+ 66,
+ 203
+ ],
+ "category_id": 6,
+ "area": 33664,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1439,
+ "image_id": 264,
+ "bbox": [
+ 113,
+ 245,
+ 174,
+ 232
+ ],
+ "category_id": 8,
+ "area": 141050,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1440,
+ "image_id": 264,
+ "bbox": [
+ 266,
+ 231,
+ 135,
+ 136
+ ],
+ "category_id": 8,
+ "area": 64558,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1475,
+ "image_id": 268,
+ "bbox": [
+ 44,
+ 123,
+ 467,
+ 335
+ ],
+ "category_id": 8,
+ "area": 1210308,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1476,
+ "image_id": 268,
+ "bbox": [
+ 231,
+ 0,
+ 253,
+ 118
+ ],
+ "category_id": 8,
+ "area": 231800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1484,
+ "image_id": 270,
+ "bbox": [
+ 2,
+ 5,
+ 355,
+ 405
+ ],
+ "category_id": 9,
+ "area": 335520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1485,
+ "image_id": 270,
+ "bbox": [
+ 225,
+ 4,
+ 243,
+ 173
+ ],
+ "category_id": 6,
+ "area": 98400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1486,
+ "image_id": 270,
+ "bbox": [
+ 310,
+ 126,
+ 106,
+ 271
+ ],
+ "category_id": 6,
+ "area": 67410,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1487,
+ "image_id": 271,
+ "bbox": [
+ 48,
+ 118,
+ 361,
+ 389
+ ],
+ "category_id": 8,
+ "area": 867542,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1499,
+ "image_id": 274,
+ "bbox": [
+ 61,
+ 122,
+ 198,
+ 187
+ ],
+ "category_id": 8,
+ "area": 130185,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1500,
+ "image_id": 274,
+ "bbox": [
+ 246,
+ 111,
+ 208,
+ 227
+ ],
+ "category_id": 8,
+ "area": 165360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1501,
+ "image_id": 275,
+ "bbox": [
+ 104,
+ 301,
+ 380,
+ 188
+ ],
+ "category_id": 9,
+ "area": 76398,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1502,
+ "image_id": 275,
+ "bbox": [
+ 91,
+ 109,
+ 91,
+ 62
+ ],
+ "category_id": 9,
+ "area": 6083,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1520,
+ "image_id": 278,
+ "bbox": [
+ 3,
+ 28,
+ 155,
+ 162
+ ],
+ "category_id": 6,
+ "area": 1180374,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1521,
+ "image_id": 278,
+ "bbox": [
+ 159,
+ 1,
+ 136,
+ 159
+ ],
+ "category_id": 6,
+ "area": 1020464,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1522,
+ "image_id": 278,
+ "bbox": [
+ 307,
+ 60,
+ 73,
+ 82
+ ],
+ "category_id": 6,
+ "area": 285762,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1523,
+ "image_id": 278,
+ "bbox": [
+ 156,
+ 269,
+ 169,
+ 165
+ ],
+ "category_id": 6,
+ "area": 1315254,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1524,
+ "image_id": 278,
+ "bbox": [
+ 0,
+ 279,
+ 66,
+ 193
+ ],
+ "category_id": 6,
+ "area": 600300,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1525,
+ "image_id": 278,
+ "bbox": [
+ 52,
+ 279,
+ 94,
+ 179
+ ],
+ "category_id": 6,
+ "area": 793104,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1526,
+ "image_id": 278,
+ "bbox": [
+ 177,
+ 167,
+ 91,
+ 101
+ ],
+ "category_id": 6,
+ "area": 435366,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1527,
+ "image_id": 278,
+ "bbox": [
+ 281,
+ 345,
+ 139,
+ 153
+ ],
+ "category_id": 6,
+ "area": 1002627,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1528,
+ "image_id": 278,
+ "bbox": [
+ 355,
+ 222,
+ 155,
+ 178
+ ],
+ "category_id": 6,
+ "area": 1296768,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1529,
+ "image_id": 278,
+ "bbox": [
+ 289,
+ 200,
+ 86,
+ 135
+ ],
+ "category_id": 6,
+ "area": 548449,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1530,
+ "image_id": 278,
+ "bbox": [
+ 351,
+ 64,
+ 160,
+ 131
+ ],
+ "category_id": 6,
+ "area": 991870,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1531,
+ "image_id": 279,
+ "bbox": [
+ 209,
+ 99,
+ 128,
+ 144
+ ],
+ "category_id": 8,
+ "area": 65163,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1532,
+ "image_id": 279,
+ "bbox": [
+ 130,
+ 219,
+ 244,
+ 227
+ ],
+ "category_id": 8,
+ "area": 195228,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1533,
+ "image_id": 279,
+ "bbox": [
+ 452,
+ 56,
+ 44,
+ 64
+ ],
+ "category_id": 8,
+ "area": 10101,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1534,
+ "image_id": 279,
+ "bbox": [
+ 370,
+ 403,
+ 50,
+ 64
+ ],
+ "category_id": 8,
+ "area": 11250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1535,
+ "image_id": 280,
+ "bbox": [
+ 84,
+ 162,
+ 88,
+ 94
+ ],
+ "category_id": 6,
+ "area": 29526,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1536,
+ "image_id": 280,
+ "bbox": [
+ 83,
+ 218,
+ 88,
+ 110
+ ],
+ "category_id": 6,
+ "area": 34100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1537,
+ "image_id": 280,
+ "bbox": [
+ 91,
+ 328,
+ 46,
+ 57
+ ],
+ "category_id": 6,
+ "area": 9477,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1538,
+ "image_id": 280,
+ "bbox": [
+ 283,
+ 361,
+ 80,
+ 84
+ ],
+ "category_id": 6,
+ "area": 23800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1539,
+ "image_id": 280,
+ "bbox": [
+ 209,
+ 299,
+ 39,
+ 57
+ ],
+ "category_id": 6,
+ "area": 8019,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1540,
+ "image_id": 280,
+ "bbox": [
+ 368,
+ 243,
+ 40,
+ 61
+ ],
+ "category_id": 6,
+ "area": 8874,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1541,
+ "image_id": 280,
+ "bbox": [
+ 176,
+ 447,
+ 98,
+ 64
+ ],
+ "category_id": 6,
+ "area": 22295,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1542,
+ "image_id": 280,
+ "bbox": [
+ 281,
+ 413,
+ 113,
+ 98
+ ],
+ "category_id": 6,
+ "area": 39337,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1543,
+ "image_id": 280,
+ "bbox": [
+ 366,
+ 377,
+ 145,
+ 134
+ ],
+ "category_id": 6,
+ "area": 68607,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1544,
+ "image_id": 280,
+ "bbox": [
+ 214,
+ 194,
+ 113,
+ 88
+ ],
+ "category_id": 8,
+ "area": 35500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1546,
+ "image_id": 280,
+ "bbox": [
+ 473,
+ 326,
+ 37,
+ 108
+ ],
+ "category_id": 6,
+ "area": 14288,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1548,
+ "image_id": 281,
+ "bbox": [
+ 26,
+ 0,
+ 115,
+ 160
+ ],
+ "category_id": 8,
+ "area": 142890,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1549,
+ "image_id": 281,
+ "bbox": [
+ 4,
+ 34,
+ 451,
+ 284
+ ],
+ "category_id": 8,
+ "area": 989168,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1550,
+ "image_id": 281,
+ "bbox": [
+ 131,
+ 198,
+ 299,
+ 216
+ ],
+ "category_id": 8,
+ "area": 501534,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1551,
+ "image_id": 282,
+ "bbox": [
+ 144,
+ 90,
+ 365,
+ 416
+ ],
+ "category_id": 9,
+ "area": 535018,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1552,
+ "image_id": 282,
+ "bbox": [
+ 0,
+ 98,
+ 318,
+ 372
+ ],
+ "category_id": 9,
+ "area": 416580,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1553,
+ "image_id": 282,
+ "bbox": [
+ 115,
+ 0,
+ 392,
+ 282
+ ],
+ "category_id": 9,
+ "area": 389457,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1559,
+ "image_id": 283,
+ "bbox": [
+ 141,
+ 125,
+ 256,
+ 386
+ ],
+ "category_id": 9,
+ "area": 103040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1567,
+ "image_id": 284,
+ "bbox": [
+ 275,
+ 159,
+ 148,
+ 118
+ ],
+ "category_id": 8,
+ "area": 20535,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1571,
+ "image_id": 287,
+ "bbox": [
+ 300,
+ 136,
+ 52,
+ 74
+ ],
+ "category_id": 8,
+ "area": 4550,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1573,
+ "image_id": 288,
+ "bbox": [
+ 76,
+ 276,
+ 322,
+ 201
+ ],
+ "category_id": 8,
+ "area": 501280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1574,
+ "image_id": 288,
+ "bbox": [
+ 78,
+ 87,
+ 130,
+ 165
+ ],
+ "category_id": 8,
+ "area": 166749,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1575,
+ "image_id": 288,
+ "bbox": [
+ 0,
+ 0,
+ 512,
+ 339
+ ],
+ "category_id": 8,
+ "area": 1341200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1576,
+ "image_id": 289,
+ "bbox": [
+ 0,
+ 59,
+ 32,
+ 50
+ ],
+ "category_id": 6,
+ "area": 11130,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1577,
+ "image_id": 289,
+ "bbox": [
+ 1,
+ 113,
+ 40,
+ 54
+ ],
+ "category_id": 6,
+ "area": 15065,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1578,
+ "image_id": 289,
+ "bbox": [
+ 42,
+ 167,
+ 43,
+ 58
+ ],
+ "category_id": 6,
+ "area": 17466,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1579,
+ "image_id": 289,
+ "bbox": [
+ 88,
+ 218,
+ 52,
+ 61
+ ],
+ "category_id": 6,
+ "area": 22059,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1580,
+ "image_id": 289,
+ "bbox": [
+ 97,
+ 231,
+ 33,
+ 59
+ ],
+ "category_id": 6,
+ "area": 13860,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1605,
+ "image_id": 294,
+ "bbox": [
+ 0,
+ 106,
+ 512,
+ 405
+ ],
+ "category_id": 6,
+ "area": 8802760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1612,
+ "image_id": 296,
+ "bbox": [
+ 340,
+ 0,
+ 150,
+ 178
+ ],
+ "category_id": 6,
+ "area": 40504,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1613,
+ "image_id": 296,
+ "bbox": [
+ 351,
+ 389,
+ 62,
+ 78
+ ],
+ "category_id": 6,
+ "area": 7373,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1614,
+ "image_id": 296,
+ "bbox": [
+ 20,
+ 466,
+ 92,
+ 44
+ ],
+ "category_id": 6,
+ "area": 6150,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1615,
+ "image_id": 296,
+ "bbox": [
+ 395,
+ 392,
+ 81,
+ 117
+ ],
+ "category_id": 6,
+ "area": 14279,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1616,
+ "image_id": 296,
+ "bbox": [
+ 471,
+ 457,
+ 39,
+ 52
+ ],
+ "category_id": 6,
+ "area": 3136,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1617,
+ "image_id": 296,
+ "bbox": [
+ 116,
+ 188,
+ 343,
+ 321
+ ],
+ "category_id": 9,
+ "area": 165390,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1618,
+ "image_id": 296,
+ "bbox": [
+ 0,
+ 0,
+ 264,
+ 458
+ ],
+ "category_id": 6,
+ "area": 181900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1620,
+ "image_id": 297,
+ "bbox": [
+ 417,
+ 206,
+ 48,
+ 77
+ ],
+ "category_id": 10,
+ "area": 286029,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1621,
+ "image_id": 297,
+ "bbox": [
+ 85,
+ 320,
+ 207,
+ 118
+ ],
+ "category_id": 9,
+ "area": 1876745,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1622,
+ "image_id": 298,
+ "bbox": [
+ 135,
+ 128,
+ 198,
+ 329
+ ],
+ "category_id": 10,
+ "area": 516385,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1624,
+ "image_id": 299,
+ "bbox": [
+ 272,
+ 196,
+ 231,
+ 312
+ ],
+ "category_id": 8,
+ "area": 95518,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1692,
+ "image_id": 306,
+ "bbox": [
+ 21,
+ 252,
+ 220,
+ 195
+ ],
+ "category_id": 8,
+ "area": 85078,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1693,
+ "image_id": 306,
+ "bbox": [
+ 274,
+ 361,
+ 176,
+ 149
+ ],
+ "category_id": 8,
+ "area": 52140,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1694,
+ "image_id": 307,
+ "bbox": [
+ 154,
+ 41,
+ 211,
+ 396
+ ],
+ "category_id": 10,
+ "area": 664578,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1702,
+ "image_id": 310,
+ "bbox": [
+ 162,
+ 71,
+ 242,
+ 377
+ ],
+ "category_id": 9,
+ "area": 321786,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1704,
+ "image_id": 311,
+ "bbox": [
+ 36,
+ 190,
+ 299,
+ 242
+ ],
+ "category_id": 9,
+ "area": 302596,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1718,
+ "image_id": 313,
+ "bbox": [
+ 22,
+ 234,
+ 172,
+ 277
+ ],
+ "category_id": 6,
+ "area": 105576,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1719,
+ "image_id": 313,
+ "bbox": [
+ 43,
+ 79,
+ 360,
+ 320
+ ],
+ "category_id": 9,
+ "area": 254331,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1723,
+ "image_id": 315,
+ "bbox": [
+ 385,
+ 432,
+ 125,
+ 79
+ ],
+ "category_id": 6,
+ "area": 34743,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1724,
+ "image_id": 315,
+ "bbox": [
+ 217,
+ 128,
+ 114,
+ 100
+ ],
+ "category_id": 8,
+ "area": 40185,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1725,
+ "image_id": 315,
+ "bbox": [
+ 175,
+ 188,
+ 143,
+ 92
+ ],
+ "category_id": 8,
+ "area": 46410,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1727,
+ "image_id": 316,
+ "bbox": [
+ 220,
+ 121,
+ 291,
+ 390
+ ],
+ "category_id": 6,
+ "area": 494525,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1728,
+ "image_id": 317,
+ "bbox": [
+ 65,
+ 192,
+ 322,
+ 256
+ ],
+ "category_id": 9,
+ "area": 252504,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1737,
+ "image_id": 318,
+ "bbox": [
+ 26,
+ 252,
+ 76,
+ 61
+ ],
+ "category_id": 8,
+ "area": 36894,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1738,
+ "image_id": 318,
+ "bbox": [
+ 219,
+ 220,
+ 41,
+ 86
+ ],
+ "category_id": 8,
+ "area": 28548,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1739,
+ "image_id": 318,
+ "bbox": [
+ 273,
+ 155,
+ 28,
+ 18
+ ],
+ "category_id": 8,
+ "area": 4320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1740,
+ "image_id": 318,
+ "bbox": [
+ 243,
+ 166,
+ 17,
+ 13
+ ],
+ "category_id": 8,
+ "area": 1876,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1741,
+ "image_id": 318,
+ "bbox": [
+ 328,
+ 172,
+ 49,
+ 31
+ ],
+ "category_id": 8,
+ "area": 12462,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1742,
+ "image_id": 318,
+ "bbox": [
+ 370,
+ 244,
+ 78,
+ 43
+ ],
+ "category_id": 8,
+ "area": 27232,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1743,
+ "image_id": 318,
+ "bbox": [
+ 434,
+ 213,
+ 77,
+ 80
+ ],
+ "category_id": 8,
+ "area": 48841,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1744,
+ "image_id": 319,
+ "bbox": [
+ 64,
+ 103,
+ 350,
+ 408
+ ],
+ "category_id": 9,
+ "area": 380100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1748,
+ "image_id": 321,
+ "bbox": [
+ 243,
+ 221,
+ 133,
+ 149
+ ],
+ "category_id": 9,
+ "area": 41125,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1759,
+ "image_id": 322,
+ "bbox": [
+ 2,
+ 251,
+ 136,
+ 142
+ ],
+ "category_id": 6,
+ "area": 139776,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1760,
+ "image_id": 322,
+ "bbox": [
+ 124,
+ 283,
+ 110,
+ 226
+ ],
+ "category_id": 6,
+ "area": 179190,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1763,
+ "image_id": 323,
+ "bbox": [
+ 56,
+ 234,
+ 387,
+ 264
+ ],
+ "category_id": 9,
+ "area": 307278,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1766,
+ "image_id": 324,
+ "bbox": [
+ 191,
+ 344,
+ 285,
+ 167
+ ],
+ "category_id": 9,
+ "area": 62928,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1767,
+ "image_id": 325,
+ "bbox": [
+ 141,
+ 84,
+ 273,
+ 423
+ ],
+ "category_id": 9,
+ "area": 407068,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1768,
+ "image_id": 325,
+ "bbox": [
+ 315,
+ 0,
+ 194,
+ 450
+ ],
+ "category_id": 9,
+ "area": 307490,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1769,
+ "image_id": 326,
+ "bbox": [
+ 242,
+ 209,
+ 244,
+ 242
+ ],
+ "category_id": 9,
+ "area": 33699,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1772,
+ "image_id": 328,
+ "bbox": [
+ 178,
+ 223,
+ 131,
+ 112
+ ],
+ "category_id": 8,
+ "area": 30371,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1786,
+ "image_id": 330,
+ "bbox": [
+ 73,
+ 169,
+ 155,
+ 222
+ ],
+ "category_id": 6,
+ "area": 142464,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1788,
+ "image_id": 331,
+ "bbox": [
+ 267,
+ 297,
+ 59,
+ 90
+ ],
+ "category_id": 9,
+ "area": 18648,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1789,
+ "image_id": 332,
+ "bbox": [
+ 3,
+ 0,
+ 505,
+ 506
+ ],
+ "category_id": 9,
+ "area": 899968,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1790,
+ "image_id": 333,
+ "bbox": [
+ 166,
+ 0,
+ 230,
+ 242
+ ],
+ "category_id": 9,
+ "area": 141198,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1791,
+ "image_id": 333,
+ "bbox": [
+ 167,
+ 184,
+ 213,
+ 307
+ ],
+ "category_id": 10,
+ "area": 166320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1793,
+ "image_id": 334,
+ "bbox": [
+ 12,
+ 140,
+ 499,
+ 371
+ ],
+ "category_id": 9,
+ "area": 157785,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1794,
+ "image_id": 335,
+ "bbox": [
+ 47,
+ 288,
+ 80,
+ 73
+ ],
+ "category_id": 10,
+ "area": 14112,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1796,
+ "image_id": 335,
+ "bbox": [
+ 234,
+ 136,
+ 277,
+ 163
+ ],
+ "category_id": 9,
+ "area": 107991,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1808,
+ "image_id": 337,
+ "bbox": [
+ 70,
+ 0,
+ 440,
+ 182
+ ],
+ "category_id": 6,
+ "area": 187712,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1809,
+ "image_id": 338,
+ "bbox": [
+ 72,
+ 105,
+ 399,
+ 374
+ ],
+ "category_id": 9,
+ "area": 524948,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1811,
+ "image_id": 339,
+ "bbox": [
+ 96,
+ 210,
+ 204,
+ 120
+ ],
+ "category_id": 8,
+ "area": 86359,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1812,
+ "image_id": 339,
+ "bbox": [
+ 312,
+ 259,
+ 140,
+ 95
+ ],
+ "category_id": 8,
+ "area": 47034,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1850,
+ "image_id": 346,
+ "bbox": [
+ 366,
+ 258,
+ 69,
+ 57
+ ],
+ "category_id": 8,
+ "area": 12975,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1851,
+ "image_id": 346,
+ "bbox": [
+ 127,
+ 144,
+ 69,
+ 67
+ ],
+ "category_id": 8,
+ "area": 15138,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1888,
+ "image_id": 349,
+ "bbox": [
+ 51,
+ 48,
+ 159,
+ 95
+ ],
+ "category_id": 8,
+ "area": 99400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1889,
+ "image_id": 349,
+ "bbox": [
+ 133,
+ 131,
+ 228,
+ 180
+ ],
+ "category_id": 8,
+ "area": 268840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1890,
+ "image_id": 349,
+ "bbox": [
+ 203,
+ 272,
+ 305,
+ 210
+ ],
+ "category_id": 8,
+ "area": 418290,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1893,
+ "image_id": 350,
+ "bbox": [
+ 3,
+ 28,
+ 507,
+ 483
+ ],
+ "category_id": 6,
+ "area": 697910,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1895,
+ "image_id": 352,
+ "bbox": [
+ 55,
+ 253,
+ 232,
+ 210
+ ],
+ "category_id": 10,
+ "area": 77308,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1897,
+ "image_id": 353,
+ "bbox": [
+ 324,
+ 134,
+ 164,
+ 213
+ ],
+ "category_id": 10,
+ "area": 57486,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1898,
+ "image_id": 353,
+ "bbox": [
+ 362,
+ 348,
+ 61,
+ 95
+ ],
+ "category_id": 10,
+ "area": 9630,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1899,
+ "image_id": 353,
+ "bbox": [
+ 277,
+ 294,
+ 56,
+ 110
+ ],
+ "category_id": 10,
+ "area": 10296,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1900,
+ "image_id": 353,
+ "bbox": [
+ 202,
+ 362,
+ 51,
+ 118
+ ],
+ "category_id": 10,
+ "area": 10080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1901,
+ "image_id": 353,
+ "bbox": [
+ 0,
+ 195,
+ 168,
+ 270
+ ],
+ "category_id": 10,
+ "area": 74970,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1919,
+ "image_id": 355,
+ "bbox": [
+ 79,
+ 253,
+ 360,
+ 215
+ ],
+ "category_id": 9,
+ "area": 791296,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1920,
+ "image_id": 356,
+ "bbox": [
+ 166,
+ 252,
+ 234,
+ 227
+ ],
+ "category_id": 8,
+ "area": 750186,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1921,
+ "image_id": 357,
+ "bbox": [
+ 376,
+ 345,
+ 134,
+ 165
+ ],
+ "category_id": 9,
+ "area": 57190,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1922,
+ "image_id": 357,
+ "bbox": [
+ 0,
+ 134,
+ 141,
+ 249
+ ],
+ "category_id": 9,
+ "area": 90396,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1923,
+ "image_id": 357,
+ "bbox": [
+ 133,
+ 272,
+ 104,
+ 205
+ ],
+ "category_id": 9,
+ "area": 55269,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1924,
+ "image_id": 357,
+ "bbox": [
+ 214,
+ 239,
+ 207,
+ 270
+ ],
+ "category_id": 9,
+ "area": 143910,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1925,
+ "image_id": 357,
+ "bbox": [
+ 251,
+ 92,
+ 177,
+ 282
+ ],
+ "category_id": 9,
+ "area": 128100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1926,
+ "image_id": 357,
+ "bbox": [
+ 143,
+ 0,
+ 239,
+ 315
+ ],
+ "category_id": 9,
+ "area": 193457,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1927,
+ "image_id": 358,
+ "bbox": [
+ 198,
+ 279,
+ 88,
+ 141
+ ],
+ "category_id": 8,
+ "area": 43956,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1928,
+ "image_id": 358,
+ "bbox": [
+ 265,
+ 296,
+ 82,
+ 59
+ ],
+ "category_id": 8,
+ "area": 17098,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1936,
+ "image_id": 360,
+ "bbox": [
+ 144,
+ 1,
+ 246,
+ 332
+ ],
+ "category_id": 9,
+ "area": 288288,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1937,
+ "image_id": 360,
+ "bbox": [
+ 172,
+ 144,
+ 186,
+ 246
+ ],
+ "category_id": 9,
+ "area": 160890,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1957,
+ "image_id": 362,
+ "bbox": [
+ 0,
+ 88,
+ 166,
+ 253
+ ],
+ "category_id": 6,
+ "area": 147740,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1958,
+ "image_id": 362,
+ "bbox": [
+ 132,
+ 218,
+ 128,
+ 202
+ ],
+ "category_id": 6,
+ "area": 91485,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1959,
+ "image_id": 362,
+ "bbox": [
+ 157,
+ 112,
+ 86,
+ 106
+ ],
+ "category_id": 6,
+ "area": 32400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1960,
+ "image_id": 362,
+ "bbox": [
+ 203,
+ 196,
+ 134,
+ 103
+ ],
+ "category_id": 6,
+ "area": 48865,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1981,
+ "image_id": 364,
+ "bbox": [
+ 3,
+ 21,
+ 171,
+ 145
+ ],
+ "category_id": 6,
+ "area": 344052,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1982,
+ "image_id": 364,
+ "bbox": [
+ 3,
+ 217,
+ 33,
+ 28
+ ],
+ "category_id": 6,
+ "area": 13266,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1983,
+ "image_id": 364,
+ "bbox": [
+ 0,
+ 398,
+ 38,
+ 109
+ ],
+ "category_id": 6,
+ "area": 58140,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1984,
+ "image_id": 364,
+ "bbox": [
+ 100,
+ 21,
+ 73,
+ 52
+ ],
+ "category_id": 6,
+ "area": 53619,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1985,
+ "image_id": 364,
+ "bbox": [
+ 183,
+ 28,
+ 100,
+ 83
+ ],
+ "category_id": 6,
+ "area": 116691,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1986,
+ "image_id": 364,
+ "bbox": [
+ 129,
+ 103,
+ 115,
+ 153
+ ],
+ "category_id": 6,
+ "area": 245252,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1987,
+ "image_id": 364,
+ "bbox": [
+ 137,
+ 258,
+ 129,
+ 91
+ ],
+ "category_id": 6,
+ "area": 164088,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1988,
+ "image_id": 364,
+ "bbox": [
+ 88,
+ 356,
+ 125,
+ 98
+ ],
+ "category_id": 6,
+ "area": 171523,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1989,
+ "image_id": 364,
+ "bbox": [
+ 261,
+ 34,
+ 114,
+ 191
+ ],
+ "category_id": 6,
+ "area": 302991,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1990,
+ "image_id": 364,
+ "bbox": [
+ 256,
+ 215,
+ 99,
+ 76
+ ],
+ "category_id": 6,
+ "area": 106134,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1991,
+ "image_id": 364,
+ "bbox": [
+ 278,
+ 288,
+ 149,
+ 86
+ ],
+ "category_id": 6,
+ "area": 180299,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1992,
+ "image_id": 364,
+ "bbox": [
+ 378,
+ 60,
+ 33,
+ 70
+ ],
+ "category_id": 6,
+ "area": 32964,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1993,
+ "image_id": 364,
+ "bbox": [
+ 357,
+ 1,
+ 95,
+ 85
+ ],
+ "category_id": 6,
+ "area": 112100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1994,
+ "image_id": 364,
+ "bbox": [
+ 393,
+ 108,
+ 74,
+ 69
+ ],
+ "category_id": 6,
+ "area": 71280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 1995,
+ "image_id": 364,
+ "bbox": [
+ 471,
+ 109,
+ 39,
+ 88
+ ],
+ "category_id": 6,
+ "area": 48506,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2002,
+ "image_id": 365,
+ "bbox": [
+ 216,
+ 256,
+ 41,
+ 51
+ ],
+ "category_id": 6,
+ "area": 5427,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2003,
+ "image_id": 365,
+ "bbox": [
+ 200,
+ 300,
+ 54,
+ 76
+ ],
+ "category_id": 6,
+ "area": 10593,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2004,
+ "image_id": 366,
+ "bbox": [
+ 115,
+ 253,
+ 173,
+ 120
+ ],
+ "category_id": 8,
+ "area": 165354,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2005,
+ "image_id": 366,
+ "bbox": [
+ 6,
+ 75,
+ 228,
+ 193
+ ],
+ "category_id": 6,
+ "area": 349248,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2006,
+ "image_id": 366,
+ "bbox": [
+ 157,
+ 196,
+ 65,
+ 48
+ ],
+ "category_id": 6,
+ "area": 25092,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2007,
+ "image_id": 366,
+ "bbox": [
+ 273,
+ 103,
+ 40,
+ 49
+ ],
+ "category_id": 6,
+ "area": 15704,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2008,
+ "image_id": 366,
+ "bbox": [
+ 241,
+ 168,
+ 44,
+ 68
+ ],
+ "category_id": 6,
+ "area": 23760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2009,
+ "image_id": 366,
+ "bbox": [
+ 235,
+ 83,
+ 35,
+ 69
+ ],
+ "category_id": 6,
+ "area": 19564,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2010,
+ "image_id": 367,
+ "bbox": [
+ 189,
+ 52,
+ 118,
+ 100
+ ],
+ "category_id": 8,
+ "area": 94359,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2011,
+ "image_id": 367,
+ "bbox": [
+ 177,
+ 206,
+ 221,
+ 212
+ ],
+ "category_id": 8,
+ "area": 372736,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2015,
+ "image_id": 369,
+ "bbox": [
+ 199,
+ 142,
+ 141,
+ 240
+ ],
+ "category_id": 6,
+ "area": 74347,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2016,
+ "image_id": 369,
+ "bbox": [
+ 203,
+ 366,
+ 232,
+ 145
+ ],
+ "category_id": 9,
+ "area": 73904,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2017,
+ "image_id": 370,
+ "bbox": [
+ 86,
+ 155,
+ 92,
+ 92
+ ],
+ "category_id": 6,
+ "area": 30160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2018,
+ "image_id": 370,
+ "bbox": [
+ 89,
+ 216,
+ 86,
+ 110
+ ],
+ "category_id": 6,
+ "area": 33696,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2019,
+ "image_id": 370,
+ "bbox": [
+ 96,
+ 327,
+ 44,
+ 60
+ ],
+ "category_id": 6,
+ "area": 9435,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2020,
+ "image_id": 370,
+ "bbox": [
+ 212,
+ 299,
+ 39,
+ 51
+ ],
+ "category_id": 6,
+ "area": 7227,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2021,
+ "image_id": 370,
+ "bbox": [
+ 288,
+ 363,
+ 83,
+ 78
+ ],
+ "category_id": 6,
+ "area": 23199,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2022,
+ "image_id": 370,
+ "bbox": [
+ 179,
+ 446,
+ 96,
+ 65
+ ],
+ "category_id": 6,
+ "area": 22172,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2023,
+ "image_id": 370,
+ "bbox": [
+ 276,
+ 413,
+ 106,
+ 98
+ ],
+ "category_id": 6,
+ "area": 36846,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2024,
+ "image_id": 370,
+ "bbox": [
+ 372,
+ 379,
+ 139,
+ 131
+ ],
+ "category_id": 6,
+ "area": 64380,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2026,
+ "image_id": 370,
+ "bbox": [
+ 217,
+ 192,
+ 110,
+ 91
+ ],
+ "category_id": 8,
+ "area": 35200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2028,
+ "image_id": 370,
+ "bbox": [
+ 361,
+ 242,
+ 52,
+ 62
+ ],
+ "category_id": 6,
+ "area": 11528,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2029,
+ "image_id": 371,
+ "bbox": [
+ 0,
+ 34,
+ 511,
+ 477
+ ],
+ "category_id": 10,
+ "area": 446472,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2041,
+ "image_id": 374,
+ "bbox": [
+ 147,
+ 374,
+ 134,
+ 117
+ ],
+ "category_id": 6,
+ "area": 18480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2042,
+ "image_id": 374,
+ "bbox": [
+ 290,
+ 359,
+ 220,
+ 151
+ ],
+ "category_id": 6,
+ "area": 39192,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2043,
+ "image_id": 375,
+ "bbox": [
+ 69,
+ 244,
+ 237,
+ 151
+ ],
+ "category_id": 8,
+ "area": 42411,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2046,
+ "image_id": 376,
+ "bbox": [
+ 0,
+ 242,
+ 300,
+ 267
+ ],
+ "category_id": 8,
+ "area": 542564,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2050,
+ "image_id": 378,
+ "bbox": [
+ 104,
+ 75,
+ 84,
+ 62
+ ],
+ "category_id": 6,
+ "area": 6195,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2051,
+ "image_id": 378,
+ "bbox": [
+ 350,
+ 326,
+ 67,
+ 72
+ ],
+ "category_id": 6,
+ "area": 5712,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2052,
+ "image_id": 378,
+ "bbox": [
+ 232,
+ 404,
+ 57,
+ 103
+ ],
+ "category_id": 6,
+ "area": 6984,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2053,
+ "image_id": 379,
+ "bbox": [
+ 0,
+ 199,
+ 509,
+ 311
+ ],
+ "category_id": 9,
+ "area": 238136,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2054,
+ "image_id": 379,
+ "bbox": [
+ 29,
+ 223,
+ 189,
+ 273
+ ],
+ "category_id": 6,
+ "area": 77978,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2057,
+ "image_id": 380,
+ "bbox": [
+ 312,
+ 173,
+ 154,
+ 176
+ ],
+ "category_id": 8,
+ "area": 215016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2058,
+ "image_id": 380,
+ "bbox": [
+ 185,
+ 152,
+ 153,
+ 178
+ ],
+ "category_id": 8,
+ "area": 217529,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2059,
+ "image_id": 380,
+ "bbox": [
+ 168,
+ 211,
+ 85,
+ 90
+ ],
+ "category_id": 8,
+ "area": 60929,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2060,
+ "image_id": 380,
+ "bbox": [
+ 0,
+ 174,
+ 113,
+ 136
+ ],
+ "category_id": 8,
+ "area": 121688,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2062,
+ "image_id": 381,
+ "bbox": [
+ 173,
+ 428,
+ 338,
+ 82
+ ],
+ "category_id": 9,
+ "area": 43524,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2090,
+ "image_id": 385,
+ "bbox": [
+ 237,
+ 151,
+ 170,
+ 320
+ ],
+ "category_id": 9,
+ "area": 108836,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2091,
+ "image_id": 386,
+ "bbox": [
+ 62,
+ 102,
+ 278,
+ 342
+ ],
+ "category_id": 8,
+ "area": 754812,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2092,
+ "image_id": 387,
+ "bbox": [
+ 83,
+ 149,
+ 112,
+ 136
+ ],
+ "category_id": 8,
+ "area": 88708,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2093,
+ "image_id": 387,
+ "bbox": [
+ 140,
+ 131,
+ 236,
+ 312
+ ],
+ "category_id": 8,
+ "area": 424116,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2095,
+ "image_id": 388,
+ "bbox": [
+ 92,
+ 218,
+ 305,
+ 151
+ ],
+ "category_id": 9,
+ "area": 220384,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2110,
+ "image_id": 390,
+ "bbox": [
+ 88,
+ 18,
+ 315,
+ 469
+ ],
+ "category_id": 10,
+ "area": 259012,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2111,
+ "image_id": 390,
+ "bbox": [
+ 337,
+ 0,
+ 170,
+ 213
+ ],
+ "category_id": 10,
+ "area": 63516,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2112,
+ "image_id": 391,
+ "bbox": [
+ 49,
+ 89,
+ 310,
+ 248
+ ],
+ "category_id": 9,
+ "area": 70616,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2113,
+ "image_id": 391,
+ "bbox": [
+ 27,
+ 325,
+ 449,
+ 175
+ ],
+ "category_id": 6,
+ "area": 72199,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2114,
+ "image_id": 392,
+ "bbox": [
+ 68,
+ 175,
+ 319,
+ 260
+ ],
+ "category_id": 8,
+ "area": 292068,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2115,
+ "image_id": 392,
+ "bbox": [
+ 301,
+ 324,
+ 103,
+ 123
+ ],
+ "category_id": 8,
+ "area": 44634,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2116,
+ "image_id": 393,
+ "bbox": [
+ 102,
+ 338,
+ 302,
+ 126
+ ],
+ "category_id": 9,
+ "area": 81328,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2130,
+ "image_id": 396,
+ "bbox": [
+ 0,
+ 154,
+ 286,
+ 357
+ ],
+ "category_id": 9,
+ "area": 225150,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2131,
+ "image_id": 396,
+ "bbox": [
+ 262,
+ 1,
+ 249,
+ 277
+ ],
+ "category_id": 6,
+ "area": 151984,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2132,
+ "image_id": 397,
+ "bbox": [
+ 144,
+ 129,
+ 326,
+ 381
+ ],
+ "category_id": 9,
+ "area": 295120,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2135,
+ "image_id": 397,
+ "bbox": [
+ 271,
+ 135,
+ 36,
+ 74
+ ],
+ "category_id": 10,
+ "area": 6499,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2144,
+ "image_id": 399,
+ "bbox": [
+ 151,
+ 163,
+ 158,
+ 316
+ ],
+ "category_id": 9,
+ "area": 387288,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2157,
+ "image_id": 402,
+ "bbox": [
+ 2,
+ 178,
+ 508,
+ 327
+ ],
+ "category_id": 9,
+ "area": 584200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2158,
+ "image_id": 402,
+ "bbox": [
+ 2,
+ 185,
+ 506,
+ 260
+ ],
+ "category_id": 9,
+ "area": 464622,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2159,
+ "image_id": 402,
+ "bbox": [
+ 14,
+ 7,
+ 446,
+ 443
+ ],
+ "category_id": 9,
+ "area": 694645,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2160,
+ "image_id": 403,
+ "bbox": [
+ 158,
+ 48,
+ 220,
+ 430
+ ],
+ "category_id": 9,
+ "area": 333906,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2162,
+ "image_id": 403,
+ "bbox": [
+ 2,
+ 130,
+ 123,
+ 114
+ ],
+ "category_id": 6,
+ "area": 49749,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2163,
+ "image_id": 403,
+ "bbox": [
+ 141,
+ 307,
+ 81,
+ 110
+ ],
+ "category_id": 6,
+ "area": 31824,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2164,
+ "image_id": 403,
+ "bbox": [
+ 227,
+ 0,
+ 264,
+ 507
+ ],
+ "category_id": 6,
+ "area": 471240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2165,
+ "image_id": 404,
+ "bbox": [
+ 145,
+ 114,
+ 215,
+ 397
+ ],
+ "category_id": 10,
+ "area": 677942,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2166,
+ "image_id": 405,
+ "bbox": [
+ 0,
+ 67,
+ 268,
+ 437
+ ],
+ "category_id": 8,
+ "area": 211422,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2167,
+ "image_id": 405,
+ "bbox": [
+ 258,
+ 0,
+ 253,
+ 510
+ ],
+ "category_id": 8,
+ "area": 232696,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2180,
+ "image_id": 407,
+ "bbox": [
+ 333,
+ 167,
+ 178,
+ 289
+ ],
+ "category_id": 8,
+ "area": 59940,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2196,
+ "image_id": 410,
+ "bbox": [
+ 160,
+ 151,
+ 183,
+ 77
+ ],
+ "category_id": 9,
+ "area": 46280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2197,
+ "image_id": 410,
+ "bbox": [
+ 160,
+ 185,
+ 138,
+ 211
+ ],
+ "category_id": 10,
+ "area": 95371,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2199,
+ "image_id": 411,
+ "bbox": [
+ 306,
+ 272,
+ 125,
+ 178
+ ],
+ "category_id": 9,
+ "area": 50451,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2200,
+ "image_id": 411,
+ "bbox": [
+ 212,
+ 101,
+ 63,
+ 104
+ ],
+ "category_id": 6,
+ "area": 14986,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2201,
+ "image_id": 411,
+ "bbox": [
+ 105,
+ 199,
+ 45,
+ 85
+ ],
+ "category_id": 6,
+ "area": 8640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2202,
+ "image_id": 411,
+ "bbox": [
+ 140,
+ 234,
+ 47,
+ 63
+ ],
+ "category_id": 6,
+ "area": 6745,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2203,
+ "image_id": 411,
+ "bbox": [
+ 161,
+ 200,
+ 47,
+ 63
+ ],
+ "category_id": 6,
+ "area": 6745,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2204,
+ "image_id": 411,
+ "bbox": [
+ 195,
+ 208,
+ 47,
+ 63
+ ],
+ "category_id": 6,
+ "area": 6745,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2205,
+ "image_id": 411,
+ "bbox": [
+ 237,
+ 199,
+ 47,
+ 63
+ ],
+ "category_id": 6,
+ "area": 6745,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2206,
+ "image_id": 411,
+ "bbox": [
+ 404,
+ 441,
+ 106,
+ 70
+ ],
+ "category_id": 6,
+ "area": 16827,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2207,
+ "image_id": 411,
+ "bbox": [
+ 356,
+ 443,
+ 79,
+ 68
+ ],
+ "category_id": 6,
+ "area": 12243,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2208,
+ "image_id": 411,
+ "bbox": [
+ 306,
+ 460,
+ 57,
+ 51
+ ],
+ "category_id": 6,
+ "area": 6670,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2209,
+ "image_id": 411,
+ "bbox": [
+ 227,
+ 448,
+ 84,
+ 64
+ ],
+ "category_id": 6,
+ "area": 12096,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2210,
+ "image_id": 411,
+ "bbox": [
+ 297,
+ 366,
+ 50,
+ 84
+ ],
+ "category_id": 6,
+ "area": 9500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2211,
+ "image_id": 411,
+ "bbox": [
+ 59,
+ 264,
+ 44,
+ 47
+ ],
+ "category_id": 6,
+ "area": 4664,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2212,
+ "image_id": 411,
+ "bbox": [
+ 71,
+ 354,
+ 45,
+ 66
+ ],
+ "category_id": 6,
+ "area": 6750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2213,
+ "image_id": 411,
+ "bbox": [
+ 350,
+ 180,
+ 80,
+ 89
+ ],
+ "category_id": 6,
+ "area": 16261,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2221,
+ "image_id": 412,
+ "bbox": [
+ 37,
+ 4,
+ 474,
+ 289
+ ],
+ "category_id": 6,
+ "area": 330990,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2222,
+ "image_id": 413,
+ "bbox": [
+ 0,
+ 111,
+ 510,
+ 393
+ ],
+ "category_id": 8,
+ "area": 705628,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2223,
+ "image_id": 413,
+ "bbox": [
+ 346,
+ 62,
+ 143,
+ 245
+ ],
+ "category_id": 8,
+ "area": 123510,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2224,
+ "image_id": 414,
+ "bbox": [
+ 2,
+ 38,
+ 506,
+ 468
+ ],
+ "category_id": 9,
+ "area": 834953,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2225,
+ "image_id": 414,
+ "bbox": [
+ 61,
+ 54,
+ 120,
+ 440
+ ],
+ "category_id": 9,
+ "area": 186938,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2226,
+ "image_id": 414,
+ "bbox": [
+ 184,
+ 2,
+ 324,
+ 88
+ ],
+ "category_id": 9,
+ "area": 100564,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2227,
+ "image_id": 415,
+ "bbox": [
+ 63,
+ 47,
+ 229,
+ 377
+ ],
+ "category_id": 10,
+ "area": 123520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2236,
+ "image_id": 417,
+ "bbox": [
+ 46,
+ 131,
+ 192,
+ 380
+ ],
+ "category_id": 10,
+ "area": 225450,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2237,
+ "image_id": 417,
+ "bbox": [
+ 210,
+ 0,
+ 154,
+ 332
+ ],
+ "category_id": 10,
+ "area": 158556,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2238,
+ "image_id": 417,
+ "bbox": [
+ 330,
+ 0,
+ 85,
+ 198
+ ],
+ "category_id": 10,
+ "area": 52400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2241,
+ "image_id": 419,
+ "bbox": [
+ 74,
+ 209,
+ 231,
+ 198
+ ],
+ "category_id": 9,
+ "area": 161541,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2242,
+ "image_id": 419,
+ "bbox": [
+ 277,
+ 86,
+ 203,
+ 283
+ ],
+ "category_id": 10,
+ "area": 202692,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2245,
+ "image_id": 421,
+ "bbox": [
+ 131,
+ 218,
+ 164,
+ 133
+ ],
+ "category_id": 8,
+ "area": 76074,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2246,
+ "image_id": 421,
+ "bbox": [
+ 159,
+ 5,
+ 266,
+ 326
+ ],
+ "category_id": 8,
+ "area": 302328,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2247,
+ "image_id": 421,
+ "bbox": [
+ 15,
+ 0,
+ 494,
+ 504
+ ],
+ "category_id": 6,
+ "area": 867855,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2264,
+ "image_id": 426,
+ "bbox": [
+ 236,
+ 128,
+ 121,
+ 177
+ ],
+ "category_id": 6,
+ "area": 157850,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2267,
+ "image_id": 427,
+ "bbox": [
+ 2,
+ 21,
+ 172,
+ 144
+ ],
+ "category_id": 6,
+ "area": 345878,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2268,
+ "image_id": 427,
+ "bbox": [
+ 0,
+ 215,
+ 41,
+ 29
+ ],
+ "category_id": 6,
+ "area": 16932,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2269,
+ "image_id": 427,
+ "bbox": [
+ 2,
+ 398,
+ 35,
+ 104
+ ],
+ "category_id": 6,
+ "area": 51909,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2270,
+ "image_id": 427,
+ "bbox": [
+ 99,
+ 19,
+ 75,
+ 57
+ ],
+ "category_id": 6,
+ "area": 59598,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2271,
+ "image_id": 427,
+ "bbox": [
+ 183,
+ 18,
+ 101,
+ 102
+ ],
+ "category_id": 6,
+ "area": 144078,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2272,
+ "image_id": 427,
+ "bbox": [
+ 127,
+ 105,
+ 117,
+ 155
+ ],
+ "category_id": 6,
+ "area": 253188,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2273,
+ "image_id": 427,
+ "bbox": [
+ 133,
+ 261,
+ 129,
+ 97
+ ],
+ "category_id": 6,
+ "area": 175422,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2274,
+ "image_id": 427,
+ "bbox": [
+ 86,
+ 357,
+ 131,
+ 98
+ ],
+ "category_id": 6,
+ "area": 179366,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2275,
+ "image_id": 427,
+ "bbox": [
+ 257,
+ 35,
+ 119,
+ 190
+ ],
+ "category_id": 6,
+ "area": 315480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2276,
+ "image_id": 427,
+ "bbox": [
+ 258,
+ 217,
+ 96,
+ 76
+ ],
+ "category_id": 6,
+ "area": 103329,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2277,
+ "image_id": 427,
+ "bbox": [
+ 280,
+ 288,
+ 152,
+ 80
+ ],
+ "category_id": 6,
+ "area": 170520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2278,
+ "image_id": 427,
+ "bbox": [
+ 357,
+ 0,
+ 95,
+ 100
+ ],
+ "category_id": 6,
+ "area": 133000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2279,
+ "image_id": 427,
+ "bbox": [
+ 375,
+ 56,
+ 36,
+ 75
+ ],
+ "category_id": 6,
+ "area": 38661,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2280,
+ "image_id": 427,
+ "bbox": [
+ 394,
+ 107,
+ 75,
+ 62
+ ],
+ "category_id": 6,
+ "area": 65016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2281,
+ "image_id": 427,
+ "bbox": [
+ 470,
+ 108,
+ 42,
+ 95
+ ],
+ "category_id": 6,
+ "area": 55608,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2295,
+ "image_id": 429,
+ "bbox": [
+ 0,
+ 30,
+ 413,
+ 480
+ ],
+ "category_id": 9,
+ "area": 201609,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2298,
+ "image_id": 430,
+ "bbox": [
+ 171,
+ 150,
+ 126,
+ 263
+ ],
+ "category_id": 10,
+ "area": 43890,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2303,
+ "image_id": 430,
+ "bbox": [
+ 441,
+ 317,
+ 69,
+ 194
+ ],
+ "category_id": 6,
+ "area": 17825,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2304,
+ "image_id": 430,
+ "bbox": [
+ 0,
+ 338,
+ 53,
+ 124
+ ],
+ "category_id": 6,
+ "area": 8811,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2305,
+ "image_id": 430,
+ "bbox": [
+ 86,
+ 386,
+ 91,
+ 121
+ ],
+ "category_id": 6,
+ "area": 14647,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2311,
+ "image_id": 433,
+ "bbox": [
+ 9,
+ 35,
+ 322,
+ 476
+ ],
+ "category_id": 10,
+ "area": 1216254,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2312,
+ "image_id": 434,
+ "bbox": [
+ 159,
+ 0,
+ 201,
+ 141
+ ],
+ "category_id": 8,
+ "area": 219705,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2313,
+ "image_id": 434,
+ "bbox": [
+ 24,
+ 112,
+ 486,
+ 399
+ ],
+ "category_id": 8,
+ "area": 1500504,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2314,
+ "image_id": 435,
+ "bbox": [
+ 71,
+ 141,
+ 440,
+ 370
+ ],
+ "category_id": 9,
+ "area": 654198,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2315,
+ "image_id": 436,
+ "bbox": [
+ 149,
+ 47,
+ 326,
+ 384
+ ],
+ "category_id": 9,
+ "area": 366624,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2316,
+ "image_id": 437,
+ "bbox": [
+ 44,
+ 200,
+ 209,
+ 264
+ ],
+ "category_id": 8,
+ "area": 194033,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2317,
+ "image_id": 437,
+ "bbox": [
+ 184,
+ 214,
+ 327,
+ 252
+ ],
+ "category_id": 8,
+ "area": 289107,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2318,
+ "image_id": 438,
+ "bbox": [
+ 44,
+ 102,
+ 24,
+ 32
+ ],
+ "category_id": 6,
+ "area": 5658,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2319,
+ "image_id": 438,
+ "bbox": [
+ 61,
+ 326,
+ 16,
+ 52
+ ],
+ "category_id": 6,
+ "area": 5994,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2320,
+ "image_id": 438,
+ "bbox": [
+ 319,
+ 412,
+ 33,
+ 55
+ ],
+ "category_id": 6,
+ "area": 12992,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2321,
+ "image_id": 438,
+ "bbox": [
+ 128,
+ 301,
+ 26,
+ 28
+ ],
+ "category_id": 6,
+ "area": 5251,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2322,
+ "image_id": 438,
+ "bbox": [
+ 109,
+ 307,
+ 23,
+ 30
+ ],
+ "category_id": 6,
+ "area": 4928,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2323,
+ "image_id": 438,
+ "bbox": [
+ 349,
+ 452,
+ 55,
+ 59
+ ],
+ "category_id": 6,
+ "area": 22692,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2324,
+ "image_id": 438,
+ "bbox": [
+ 339,
+ 276,
+ 12,
+ 23
+ ],
+ "category_id": 6,
+ "area": 2107,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2325,
+ "image_id": 438,
+ "bbox": [
+ 347,
+ 380,
+ 17,
+ 22
+ ],
+ "category_id": 6,
+ "area": 2679,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2326,
+ "image_id": 438,
+ "bbox": [
+ 346,
+ 344,
+ 18,
+ 36
+ ],
+ "category_id": 6,
+ "area": 4636,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2329,
+ "image_id": 439,
+ "bbox": [
+ 75,
+ 211,
+ 132,
+ 112
+ ],
+ "category_id": 8,
+ "area": 38710,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2330,
+ "image_id": 439,
+ "bbox": [
+ 345,
+ 161,
+ 166,
+ 148
+ ],
+ "category_id": 8,
+ "area": 63856,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2332,
+ "image_id": 440,
+ "bbox": [
+ 244,
+ 347,
+ 91,
+ 92
+ ],
+ "category_id": 6,
+ "area": 22263,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2333,
+ "image_id": 440,
+ "bbox": [
+ 161,
+ 377,
+ 88,
+ 71
+ ],
+ "category_id": 6,
+ "area": 16800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2334,
+ "image_id": 440,
+ "bbox": [
+ 422,
+ 405,
+ 23,
+ 43
+ ],
+ "category_id": 6,
+ "area": 2726,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2335,
+ "image_id": 440,
+ "bbox": [
+ 298,
+ 431,
+ 37,
+ 41
+ ],
+ "category_id": 6,
+ "area": 4070,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2336,
+ "image_id": 440,
+ "bbox": [
+ 322,
+ 400,
+ 37,
+ 53
+ ],
+ "category_id": 6,
+ "area": 5254,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2337,
+ "image_id": 440,
+ "bbox": [
+ 365,
+ 415,
+ 47,
+ 49
+ ],
+ "category_id": 6,
+ "area": 6204,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2338,
+ "image_id": 440,
+ "bbox": [
+ 4,
+ 419,
+ 36,
+ 52
+ ],
+ "category_id": 6,
+ "area": 5040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2339,
+ "image_id": 440,
+ "bbox": [
+ 207,
+ 420,
+ 58,
+ 68
+ ],
+ "category_id": 6,
+ "area": 10647,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2340,
+ "image_id": 440,
+ "bbox": [
+ 205,
+ 286,
+ 33,
+ 32
+ ],
+ "category_id": 6,
+ "area": 2881,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2341,
+ "image_id": 440,
+ "bbox": [
+ 288,
+ 291,
+ 49,
+ 51
+ ],
+ "category_id": 6,
+ "area": 6762,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2342,
+ "image_id": 440,
+ "bbox": [
+ 322,
+ 452,
+ 30,
+ 50
+ ],
+ "category_id": 6,
+ "area": 4087,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2343,
+ "image_id": 440,
+ "bbox": [
+ 397,
+ 449,
+ 46,
+ 42
+ ],
+ "category_id": 6,
+ "area": 5244,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2365,
+ "image_id": 443,
+ "bbox": [
+ 178,
+ 212,
+ 217,
+ 158
+ ],
+ "category_id": 8,
+ "area": 40108,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2382,
+ "image_id": 446,
+ "bbox": [
+ 0,
+ 106,
+ 247,
+ 394
+ ],
+ "category_id": 9,
+ "area": 357315,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2384,
+ "image_id": 447,
+ "bbox": [
+ 172,
+ 177,
+ 285,
+ 320
+ ],
+ "category_id": 8,
+ "area": 106444,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2404,
+ "image_id": 450,
+ "bbox": [
+ 0,
+ 272,
+ 93,
+ 145
+ ],
+ "category_id": 10,
+ "area": 34770,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2405,
+ "image_id": 450,
+ "bbox": [
+ 59,
+ 142,
+ 120,
+ 163
+ ],
+ "category_id": 10,
+ "area": 50268,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2406,
+ "image_id": 450,
+ "bbox": [
+ 116,
+ 418,
+ 61,
+ 71
+ ],
+ "category_id": 10,
+ "area": 11160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2407,
+ "image_id": 450,
+ "bbox": [
+ 206,
+ 318,
+ 37,
+ 63
+ ],
+ "category_id": 10,
+ "area": 6059,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2408,
+ "image_id": 450,
+ "bbox": [
+ 147,
+ 296,
+ 33,
+ 51
+ ],
+ "category_id": 10,
+ "area": 4355,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2409,
+ "image_id": 450,
+ "bbox": [
+ 369,
+ 363,
+ 48,
+ 68
+ ],
+ "category_id": 10,
+ "area": 8455,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2410,
+ "image_id": 450,
+ "bbox": [
+ 339,
+ 416,
+ 51,
+ 81
+ ],
+ "category_id": 10,
+ "area": 10706,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2411,
+ "image_id": 450,
+ "bbox": [
+ 338,
+ 185,
+ 65,
+ 92
+ ],
+ "category_id": 10,
+ "area": 15367,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2412,
+ "image_id": 450,
+ "bbox": [
+ 349,
+ 94,
+ 82,
+ 101
+ ],
+ "category_id": 10,
+ "area": 21384,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2413,
+ "image_id": 450,
+ "bbox": [
+ 423,
+ 272,
+ 48,
+ 47
+ ],
+ "category_id": 10,
+ "area": 5890,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2414,
+ "image_id": 450,
+ "bbox": [
+ 425,
+ 188,
+ 85,
+ 101
+ ],
+ "category_id": 10,
+ "area": 22044,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2415,
+ "image_id": 450,
+ "bbox": [
+ 443,
+ 129,
+ 68,
+ 69
+ ],
+ "category_id": 10,
+ "area": 12103,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2416,
+ "image_id": 450,
+ "bbox": [
+ 424,
+ 34,
+ 54,
+ 72
+ ],
+ "category_id": 10,
+ "area": 10165,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2417,
+ "image_id": 450,
+ "bbox": [
+ 301,
+ 98,
+ 47,
+ 64
+ ],
+ "category_id": 10,
+ "area": 7728,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2418,
+ "image_id": 450,
+ "bbox": [
+ 350,
+ 65,
+ 29,
+ 52
+ ],
+ "category_id": 10,
+ "area": 3944,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2426,
+ "image_id": 451,
+ "bbox": [
+ 16,
+ 304,
+ 13,
+ 19
+ ],
+ "category_id": 9,
+ "area": 924,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2427,
+ "image_id": 451,
+ "bbox": [
+ 55,
+ 227,
+ 15,
+ 29
+ ],
+ "category_id": 9,
+ "area": 1638,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2428,
+ "image_id": 451,
+ "bbox": [
+ 156,
+ 209,
+ 167,
+ 137
+ ],
+ "category_id": 9,
+ "area": 81286,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2429,
+ "image_id": 451,
+ "bbox": [
+ 419,
+ 124,
+ 50,
+ 199
+ ],
+ "category_id": 9,
+ "area": 35000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2430,
+ "image_id": 452,
+ "bbox": [
+ 64,
+ 135,
+ 381,
+ 299
+ ],
+ "category_id": 8,
+ "area": 420180,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2431,
+ "image_id": 453,
+ "bbox": [
+ 301,
+ 167,
+ 138,
+ 76
+ ],
+ "category_id": 8,
+ "area": 149210,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2432,
+ "image_id": 453,
+ "bbox": [
+ 312,
+ 124,
+ 160,
+ 57
+ ],
+ "category_id": 8,
+ "area": 129600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2433,
+ "image_id": 454,
+ "bbox": [
+ 1,
+ 105,
+ 433,
+ 403
+ ],
+ "category_id": 9,
+ "area": 614628,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2434,
+ "image_id": 454,
+ "bbox": [
+ 3,
+ 18,
+ 506,
+ 487
+ ],
+ "category_id": 9,
+ "area": 869162,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2435,
+ "image_id": 454,
+ "bbox": [
+ 3,
+ 0,
+ 396,
+ 420
+ ],
+ "category_id": 9,
+ "area": 586672,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2436,
+ "image_id": 455,
+ "bbox": [
+ 151,
+ 112,
+ 264,
+ 351
+ ],
+ "category_id": 9,
+ "area": 326040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2437,
+ "image_id": 455,
+ "bbox": [
+ 221,
+ 30,
+ 215,
+ 401
+ ],
+ "category_id": 9,
+ "area": 303432,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2438,
+ "image_id": 456,
+ "bbox": [
+ 31,
+ 118,
+ 297,
+ 393
+ ],
+ "category_id": 8,
+ "area": 410326,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2439,
+ "image_id": 456,
+ "bbox": [
+ 71,
+ 0,
+ 416,
+ 471
+ ],
+ "category_id": 8,
+ "area": 687818,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2467,
+ "image_id": 461,
+ "bbox": [
+ 100,
+ 129,
+ 63,
+ 158
+ ],
+ "category_id": 6,
+ "area": 12036,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2468,
+ "image_id": 461,
+ "bbox": [
+ 115,
+ 4,
+ 68,
+ 127
+ ],
+ "category_id": 6,
+ "area": 10450,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2469,
+ "image_id": 462,
+ "bbox": [
+ 181,
+ 43,
+ 84,
+ 93
+ ],
+ "category_id": 9,
+ "area": 27510,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2470,
+ "image_id": 463,
+ "bbox": [
+ 162,
+ 73,
+ 62,
+ 133
+ ],
+ "category_id": 8,
+ "area": 65754,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2471,
+ "image_id": 463,
+ "bbox": [
+ 247,
+ 145,
+ 85,
+ 283
+ ],
+ "category_id": 8,
+ "area": 191081,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2472,
+ "image_id": 463,
+ "bbox": [
+ 297,
+ 343,
+ 166,
+ 168
+ ],
+ "category_id": 8,
+ "area": 221788,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2475,
+ "image_id": 463,
+ "bbox": [
+ 3,
+ 228,
+ 44,
+ 80
+ ],
+ "category_id": 6,
+ "area": 28392,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2476,
+ "image_id": 463,
+ "bbox": [
+ 69,
+ 144,
+ 25,
+ 100
+ ],
+ "category_id": 6,
+ "area": 20448,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2478,
+ "image_id": 465,
+ "bbox": [
+ 86,
+ 171,
+ 370,
+ 339
+ ],
+ "category_id": 9,
+ "area": 176176,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2479,
+ "image_id": 465,
+ "bbox": [
+ 1,
+ 317,
+ 189,
+ 193
+ ],
+ "category_id": 6,
+ "area": 51100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2481,
+ "image_id": 466,
+ "bbox": [
+ 8,
+ 101,
+ 405,
+ 254
+ ],
+ "category_id": 9,
+ "area": 309420,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2499,
+ "image_id": 470,
+ "bbox": [
+ 171,
+ 395,
+ 112,
+ 113
+ ],
+ "category_id": 6,
+ "area": 36023,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2500,
+ "image_id": 470,
+ "bbox": [
+ 39,
+ 424,
+ 95,
+ 87
+ ],
+ "category_id": 6,
+ "area": 23688,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2501,
+ "image_id": 470,
+ "bbox": [
+ 0,
+ 286,
+ 173,
+ 134
+ ],
+ "category_id": 6,
+ "area": 66348,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2502,
+ "image_id": 470,
+ "bbox": [
+ 10,
+ 158,
+ 124,
+ 123
+ ],
+ "category_id": 6,
+ "area": 43365,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2503,
+ "image_id": 470,
+ "bbox": [
+ 107,
+ 219,
+ 136,
+ 132
+ ],
+ "category_id": 6,
+ "area": 50920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2504,
+ "image_id": 470,
+ "bbox": [
+ 207,
+ 253,
+ 143,
+ 157
+ ],
+ "category_id": 6,
+ "area": 64014,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2505,
+ "image_id": 470,
+ "bbox": [
+ 371,
+ 327,
+ 86,
+ 104
+ ],
+ "category_id": 6,
+ "area": 25500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2506,
+ "image_id": 470,
+ "bbox": [
+ 217,
+ 114,
+ 139,
+ 171
+ ],
+ "category_id": 6,
+ "area": 67678,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2507,
+ "image_id": 470,
+ "bbox": [
+ 349,
+ 141,
+ 147,
+ 128
+ ],
+ "category_id": 6,
+ "area": 53360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2508,
+ "image_id": 470,
+ "bbox": [
+ 387,
+ 0,
+ 123,
+ 107
+ ],
+ "category_id": 6,
+ "area": 37820,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2509,
+ "image_id": 471,
+ "bbox": [
+ 1,
+ 105,
+ 428,
+ 402
+ ],
+ "category_id": 8,
+ "area": 605620,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2510,
+ "image_id": 471,
+ "bbox": [
+ 364,
+ 79,
+ 128,
+ 243
+ ],
+ "category_id": 8,
+ "area": 109760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2515,
+ "image_id": 472,
+ "bbox": [
+ 1,
+ 212,
+ 249,
+ 285
+ ],
+ "category_id": 6,
+ "area": 64911,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2516,
+ "image_id": 472,
+ "bbox": [
+ 184,
+ 169,
+ 76,
+ 111
+ ],
+ "category_id": 6,
+ "area": 7740,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2517,
+ "image_id": 472,
+ "bbox": [
+ 254,
+ 140,
+ 153,
+ 127
+ ],
+ "category_id": 6,
+ "area": 17819,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2518,
+ "image_id": 472,
+ "bbox": [
+ 355,
+ 197,
+ 90,
+ 143
+ ],
+ "category_id": 6,
+ "area": 11832,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2522,
+ "image_id": 474,
+ "bbox": [
+ 114,
+ 19,
+ 394,
+ 487
+ ],
+ "category_id": 9,
+ "area": 677082,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2523,
+ "image_id": 475,
+ "bbox": [
+ 96,
+ 0,
+ 292,
+ 505
+ ],
+ "category_id": 9,
+ "area": 519030,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2524,
+ "image_id": 475,
+ "bbox": [
+ 286,
+ 0,
+ 180,
+ 221
+ ],
+ "category_id": 9,
+ "area": 140261,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2533,
+ "image_id": 477,
+ "bbox": [
+ 58,
+ 0,
+ 329,
+ 384
+ ],
+ "category_id": 9,
+ "area": 234112,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2534,
+ "image_id": 477,
+ "bbox": [
+ 67,
+ 74,
+ 249,
+ 376
+ ],
+ "category_id": 9,
+ "area": 173625,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2535,
+ "image_id": 478,
+ "bbox": [
+ 0,
+ 0,
+ 407,
+ 310
+ ],
+ "category_id": 6,
+ "area": 307278,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2536,
+ "image_id": 478,
+ "bbox": [
+ 468,
+ 151,
+ 43,
+ 87
+ ],
+ "category_id": 6,
+ "area": 9184,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2541,
+ "image_id": 478,
+ "bbox": [
+ 1,
+ 0,
+ 403,
+ 285
+ ],
+ "category_id": 6,
+ "area": 279590,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2542,
+ "image_id": 478,
+ "bbox": [
+ 469,
+ 151,
+ 42,
+ 86
+ ],
+ "category_id": 6,
+ "area": 8991,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2571,
+ "image_id": 481,
+ "bbox": [
+ 55,
+ 177,
+ 254,
+ 125
+ ],
+ "category_id": 9,
+ "area": 75990,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2572,
+ "image_id": 481,
+ "bbox": [
+ 250,
+ 119,
+ 226,
+ 102
+ ],
+ "category_id": 9,
+ "area": 55860,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2573,
+ "image_id": 482,
+ "bbox": [
+ 7,
+ 129,
+ 68,
+ 54
+ ],
+ "category_id": 8,
+ "area": 29580,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2574,
+ "image_id": 482,
+ "bbox": [
+ 1,
+ 188,
+ 81,
+ 68
+ ],
+ "category_id": 8,
+ "area": 44208,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2575,
+ "image_id": 482,
+ "bbox": [
+ 71,
+ 210,
+ 373,
+ 226
+ ],
+ "category_id": 8,
+ "area": 669200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2576,
+ "image_id": 482,
+ "bbox": [
+ 0,
+ 389,
+ 291,
+ 122
+ ],
+ "category_id": 8,
+ "area": 282252,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2581,
+ "image_id": 484,
+ "bbox": [
+ 4,
+ 48,
+ 475,
+ 416
+ ],
+ "category_id": 6,
+ "area": 232366,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2582,
+ "image_id": 485,
+ "bbox": [
+ 25,
+ 126,
+ 330,
+ 385
+ ],
+ "category_id": 9,
+ "area": 364490,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2585,
+ "image_id": 485,
+ "bbox": [
+ 183,
+ 128,
+ 73,
+ 64
+ ],
+ "category_id": 10,
+ "area": 13589,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2587,
+ "image_id": 486,
+ "bbox": [
+ 158,
+ 97,
+ 323,
+ 328
+ ],
+ "category_id": 8,
+ "area": 480782,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2595,
+ "image_id": 489,
+ "bbox": [
+ 98,
+ 0,
+ 394,
+ 436
+ ],
+ "category_id": 10,
+ "area": 1362159,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2596,
+ "image_id": 490,
+ "bbox": [
+ 21,
+ 0,
+ 401,
+ 479
+ ],
+ "category_id": 10,
+ "area": 1523060,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2597,
+ "image_id": 490,
+ "bbox": [
+ 356,
+ 448,
+ 34,
+ 62
+ ],
+ "category_id": 10,
+ "area": 17028,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2598,
+ "image_id": 490,
+ "bbox": [
+ 33,
+ 325,
+ 39,
+ 91
+ ],
+ "category_id": 10,
+ "area": 28757,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2603,
+ "image_id": 492,
+ "bbox": [
+ 191,
+ 145,
+ 237,
+ 366
+ ],
+ "category_id": 9,
+ "area": 206541,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2605,
+ "image_id": 492,
+ "bbox": [
+ 276,
+ 142,
+ 53,
+ 100
+ ],
+ "category_id": 10,
+ "area": 12707,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2613,
+ "image_id": 494,
+ "bbox": [
+ 141,
+ 148,
+ 255,
+ 163
+ ],
+ "category_id": 9,
+ "area": 125195,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2622,
+ "image_id": 496,
+ "bbox": [
+ 270,
+ 181,
+ 213,
+ 168
+ ],
+ "category_id": 8,
+ "area": 92169,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2623,
+ "image_id": 496,
+ "bbox": [
+ 442,
+ 212,
+ 69,
+ 101
+ ],
+ "category_id": 8,
+ "area": 18209,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2624,
+ "image_id": 497,
+ "bbox": [
+ 101,
+ 210,
+ 154,
+ 167
+ ],
+ "category_id": 8,
+ "area": 90090,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2625,
+ "image_id": 497,
+ "bbox": [
+ 208,
+ 213,
+ 201,
+ 162
+ ],
+ "category_id": 8,
+ "area": 114684,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2626,
+ "image_id": 497,
+ "bbox": [
+ 94,
+ 126,
+ 23,
+ 24
+ ],
+ "category_id": 6,
+ "area": 2030,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2627,
+ "image_id": 497,
+ "bbox": [
+ 126,
+ 163,
+ 28,
+ 34
+ ],
+ "category_id": 6,
+ "area": 3360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2628,
+ "image_id": 497,
+ "bbox": [
+ 136,
+ 209,
+ 29,
+ 27
+ ],
+ "category_id": 6,
+ "area": 2847,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2629,
+ "image_id": 497,
+ "bbox": [
+ 156,
+ 194,
+ 18,
+ 17
+ ],
+ "category_id": 6,
+ "area": 1150,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2630,
+ "image_id": 497,
+ "bbox": [
+ 136,
+ 399,
+ 36,
+ 38
+ ],
+ "category_id": 6,
+ "area": 4860,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2632,
+ "image_id": 498,
+ "bbox": [
+ 154,
+ 173,
+ 237,
+ 169
+ ],
+ "category_id": 8,
+ "area": 318262,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2633,
+ "image_id": 498,
+ "bbox": [
+ 0,
+ 192,
+ 223,
+ 181
+ ],
+ "category_id": 8,
+ "area": 320116,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2634,
+ "image_id": 498,
+ "bbox": [
+ 425,
+ 303,
+ 37,
+ 55
+ ],
+ "category_id": 8,
+ "area": 16756,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2635,
+ "image_id": 498,
+ "bbox": [
+ 421,
+ 318,
+ 49,
+ 37
+ ],
+ "category_id": 8,
+ "area": 14960,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2636,
+ "image_id": 499,
+ "bbox": [
+ 6,
+ 78,
+ 379,
+ 346
+ ],
+ "category_id": 8,
+ "area": 331968,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2637,
+ "image_id": 499,
+ "bbox": [
+ 356,
+ 371,
+ 154,
+ 138
+ ],
+ "category_id": 6,
+ "area": 53879,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2640,
+ "image_id": 501,
+ "bbox": [
+ 0,
+ 94,
+ 223,
+ 177
+ ],
+ "category_id": 8,
+ "area": 104370,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2641,
+ "image_id": 501,
+ "bbox": [
+ 162,
+ 218,
+ 195,
+ 148
+ ],
+ "category_id": 8,
+ "area": 76670,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2642,
+ "image_id": 501,
+ "bbox": [
+ 375,
+ 269,
+ 135,
+ 134
+ ],
+ "category_id": 8,
+ "area": 48174,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2686,
+ "image_id": 505,
+ "bbox": [
+ 33,
+ 139,
+ 96,
+ 102
+ ],
+ "category_id": 6,
+ "area": 39456,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2687,
+ "image_id": 505,
+ "bbox": [
+ 74,
+ 197,
+ 79,
+ 112
+ ],
+ "category_id": 6,
+ "area": 35639,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2688,
+ "image_id": 505,
+ "bbox": [
+ 149,
+ 212,
+ 60,
+ 91
+ ],
+ "category_id": 6,
+ "area": 22016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2689,
+ "image_id": 505,
+ "bbox": [
+ 0,
+ 274,
+ 55,
+ 217
+ ],
+ "category_id": 6,
+ "area": 48032,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2690,
+ "image_id": 505,
+ "bbox": [
+ 265,
+ 287,
+ 70,
+ 117
+ ],
+ "category_id": 6,
+ "area": 32964,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2691,
+ "image_id": 505,
+ "bbox": [
+ 68,
+ 321,
+ 96,
+ 189
+ ],
+ "category_id": 6,
+ "area": 73416,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2692,
+ "image_id": 505,
+ "bbox": [
+ 141,
+ 313,
+ 79,
+ 164
+ ],
+ "category_id": 6,
+ "area": 52210,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2693,
+ "image_id": 505,
+ "bbox": [
+ 298,
+ 196,
+ 34,
+ 76
+ ],
+ "category_id": 6,
+ "area": 10593,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2694,
+ "image_id": 505,
+ "bbox": [
+ 295,
+ 107,
+ 37,
+ 75
+ ],
+ "category_id": 6,
+ "area": 11448,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2695,
+ "image_id": 505,
+ "bbox": [
+ 324,
+ 112,
+ 23,
+ 32
+ ],
+ "category_id": 6,
+ "area": 3082,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2696,
+ "image_id": 505,
+ "bbox": [
+ 348,
+ 87,
+ 53,
+ 92
+ ],
+ "category_id": 6,
+ "area": 19760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2697,
+ "image_id": 505,
+ "bbox": [
+ 378,
+ 107,
+ 62,
+ 144
+ ],
+ "category_id": 6,
+ "area": 35754,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2698,
+ "image_id": 505,
+ "bbox": [
+ 384,
+ 152,
+ 127,
+ 252
+ ],
+ "category_id": 6,
+ "area": 128148,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2699,
+ "image_id": 505,
+ "bbox": [
+ 397,
+ 396,
+ 72,
+ 115
+ ],
+ "category_id": 6,
+ "area": 33372,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2710,
+ "image_id": 507,
+ "bbox": [
+ 240,
+ 104,
+ 210,
+ 338
+ ],
+ "category_id": 8,
+ "area": 341960,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2712,
+ "image_id": 508,
+ "bbox": [
+ 20,
+ 297,
+ 444,
+ 211
+ ],
+ "category_id": 8,
+ "area": 249150,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2720,
+ "image_id": 510,
+ "bbox": [
+ 26,
+ 175,
+ 216,
+ 286
+ ],
+ "category_id": 8,
+ "area": 83780,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2721,
+ "image_id": 510,
+ "bbox": [
+ 233,
+ 297,
+ 146,
+ 110
+ ],
+ "category_id": 8,
+ "area": 21840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2722,
+ "image_id": 510,
+ "bbox": [
+ 368,
+ 321,
+ 99,
+ 70
+ ],
+ "category_id": 8,
+ "area": 9512,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2723,
+ "image_id": 510,
+ "bbox": [
+ 363,
+ 236,
+ 93,
+ 76
+ ],
+ "category_id": 8,
+ "area": 9639,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2728,
+ "image_id": 512,
+ "bbox": [
+ 64,
+ 85,
+ 304,
+ 349
+ ],
+ "category_id": 8,
+ "area": 716056,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2748,
+ "image_id": 514,
+ "bbox": [
+ 178,
+ 161,
+ 181,
+ 266
+ ],
+ "category_id": 9,
+ "area": 316944,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2749,
+ "image_id": 514,
+ "bbox": [
+ 0,
+ 163,
+ 79,
+ 139
+ ],
+ "category_id": 6,
+ "area": 72416,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2750,
+ "image_id": 514,
+ "bbox": [
+ 314,
+ 197,
+ 196,
+ 267
+ ],
+ "category_id": 6,
+ "area": 345630,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2751,
+ "image_id": 515,
+ "bbox": [
+ 82,
+ 12,
+ 419,
+ 431
+ ],
+ "category_id": 9,
+ "area": 181828,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2757,
+ "image_id": 517,
+ "bbox": [
+ 207,
+ 189,
+ 90,
+ 135
+ ],
+ "category_id": 8,
+ "area": 97185,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2758,
+ "image_id": 517,
+ "bbox": [
+ 308,
+ 211,
+ 140,
+ 74
+ ],
+ "category_id": 8,
+ "area": 82950,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2759,
+ "image_id": 517,
+ "bbox": [
+ 140,
+ 464,
+ 268,
+ 47
+ ],
+ "category_id": 8,
+ "area": 101505,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2760,
+ "image_id": 517,
+ "bbox": [
+ 297,
+ 258,
+ 17,
+ 71
+ ],
+ "category_id": 8,
+ "area": 9900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2761,
+ "image_id": 518,
+ "bbox": [
+ 127,
+ 133,
+ 242,
+ 377
+ ],
+ "category_id": 10,
+ "area": 726067,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2762,
+ "image_id": 519,
+ "bbox": [
+ 210,
+ 2,
+ 189,
+ 405
+ ],
+ "category_id": 9,
+ "area": 109531,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2763,
+ "image_id": 519,
+ "bbox": [
+ 1,
+ 284,
+ 105,
+ 220
+ ],
+ "category_id": 9,
+ "area": 32960,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2777,
+ "image_id": 523,
+ "bbox": [
+ 141,
+ 204,
+ 174,
+ 187
+ ],
+ "category_id": 8,
+ "area": 115368,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2778,
+ "image_id": 523,
+ "bbox": [
+ 284,
+ 414,
+ 174,
+ 94
+ ],
+ "category_id": 8,
+ "area": 57855,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2781,
+ "image_id": 524,
+ "bbox": [
+ 1,
+ 0,
+ 104,
+ 169
+ ],
+ "category_id": 6,
+ "area": 57204,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2782,
+ "image_id": 524,
+ "bbox": [
+ 1,
+ 157,
+ 51,
+ 67
+ ],
+ "category_id": 6,
+ "area": 11160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2783,
+ "image_id": 524,
+ "bbox": [
+ 95,
+ 0,
+ 151,
+ 143
+ ],
+ "category_id": 6,
+ "area": 70272,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2784,
+ "image_id": 524,
+ "bbox": [
+ 0,
+ 325,
+ 173,
+ 179
+ ],
+ "category_id": 6,
+ "area": 100800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2785,
+ "image_id": 525,
+ "bbox": [
+ 8,
+ 67,
+ 307,
+ 257
+ ],
+ "category_id": 10,
+ "area": 153825,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2789,
+ "image_id": 526,
+ "bbox": [
+ 0,
+ 114,
+ 393,
+ 279
+ ],
+ "category_id": 6,
+ "area": 271440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2808,
+ "image_id": 530,
+ "bbox": [
+ 172,
+ 59,
+ 338,
+ 450
+ ],
+ "category_id": 8,
+ "area": 4832342,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2809,
+ "image_id": 531,
+ "bbox": [
+ 83,
+ 4,
+ 323,
+ 420
+ ],
+ "category_id": 9,
+ "area": 345788,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2824,
+ "image_id": 534,
+ "bbox": [
+ 141,
+ 99,
+ 256,
+ 157
+ ],
+ "category_id": 9,
+ "area": 153807,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2825,
+ "image_id": 534,
+ "bbox": [
+ 20,
+ 253,
+ 440,
+ 240
+ ],
+ "category_id": 6,
+ "area": 403340,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2838,
+ "image_id": 536,
+ "bbox": [
+ 0,
+ 191,
+ 511,
+ 319
+ ],
+ "category_id": 9,
+ "area": 244792,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2839,
+ "image_id": 536,
+ "bbox": [
+ 0,
+ 0,
+ 329,
+ 490
+ ],
+ "category_id": 6,
+ "area": 242515,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2840,
+ "image_id": 537,
+ "bbox": [
+ 130,
+ 5,
+ 179,
+ 191
+ ],
+ "category_id": 9,
+ "area": 98800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2841,
+ "image_id": 537,
+ "bbox": [
+ 235,
+ 53,
+ 137,
+ 219
+ ],
+ "category_id": 10,
+ "area": 86598,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2859,
+ "image_id": 539,
+ "bbox": [
+ 164,
+ 179,
+ 87,
+ 109
+ ],
+ "category_id": 8,
+ "area": 33354,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2860,
+ "image_id": 539,
+ "bbox": [
+ 354,
+ 155,
+ 110,
+ 108
+ ],
+ "category_id": 8,
+ "area": 41800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2861,
+ "image_id": 540,
+ "bbox": [
+ 0,
+ 141,
+ 316,
+ 369
+ ],
+ "category_id": 6,
+ "area": 408430,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2862,
+ "image_id": 540,
+ "bbox": [
+ 202,
+ 114,
+ 215,
+ 167
+ ],
+ "category_id": 8,
+ "area": 126430,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2863,
+ "image_id": 540,
+ "bbox": [
+ 315,
+ 125,
+ 194,
+ 275
+ ],
+ "category_id": 8,
+ "area": 187982,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2864,
+ "image_id": 541,
+ "bbox": [
+ 290,
+ 14,
+ 219,
+ 152
+ ],
+ "category_id": 9,
+ "area": 117272,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2865,
+ "image_id": 541,
+ "bbox": [
+ 151,
+ 147,
+ 142,
+ 177
+ ],
+ "category_id": 9,
+ "area": 88750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2866,
+ "image_id": 541,
+ "bbox": [
+ 2,
+ 174,
+ 75,
+ 161
+ ],
+ "category_id": 9,
+ "area": 42676,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2867,
+ "image_id": 541,
+ "bbox": [
+ 15,
+ 128,
+ 494,
+ 376
+ ],
+ "category_id": 9,
+ "area": 655080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2902,
+ "image_id": 546,
+ "bbox": [
+ 0,
+ 209,
+ 207,
+ 192
+ ],
+ "category_id": 8,
+ "area": 107334,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2903,
+ "image_id": 546,
+ "bbox": [
+ 162,
+ 150,
+ 349,
+ 228
+ ],
+ "category_id": 8,
+ "area": 215604,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2905,
+ "image_id": 547,
+ "bbox": [
+ 383,
+ 129,
+ 128,
+ 190
+ ],
+ "category_id": 6,
+ "area": 41193,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2906,
+ "image_id": 547,
+ "bbox": [
+ 89,
+ 300,
+ 420,
+ 210
+ ],
+ "category_id": 6,
+ "area": 149308,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2907,
+ "image_id": 547,
+ "bbox": [
+ 132,
+ 194,
+ 379,
+ 255
+ ],
+ "category_id": 9,
+ "area": 162876,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2914,
+ "image_id": 550,
+ "bbox": [
+ 103,
+ 119,
+ 408,
+ 389
+ ],
+ "category_id": 8,
+ "area": 1096334,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2915,
+ "image_id": 550,
+ "bbox": [
+ 261,
+ 113,
+ 115,
+ 219
+ ],
+ "category_id": 8,
+ "area": 173635,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2916,
+ "image_id": 550,
+ "bbox": [
+ 142,
+ 474,
+ 77,
+ 34
+ ],
+ "category_id": 8,
+ "area": 18576,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2917,
+ "image_id": 551,
+ "bbox": [
+ 76,
+ 115,
+ 405,
+ 391
+ ],
+ "category_id": 8,
+ "area": 1092936,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2918,
+ "image_id": 551,
+ "bbox": [
+ 266,
+ 129,
+ 107,
+ 220
+ ],
+ "category_id": 8,
+ "area": 163080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2919,
+ "image_id": 551,
+ "bbox": [
+ 138,
+ 442,
+ 71,
+ 67
+ ],
+ "category_id": 8,
+ "area": 32982,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2923,
+ "image_id": 553,
+ "bbox": [
+ 189,
+ 156,
+ 269,
+ 351
+ ],
+ "category_id": 8,
+ "area": 110208,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2937,
+ "image_id": 556,
+ "bbox": [
+ 62,
+ 302,
+ 50,
+ 65
+ ],
+ "category_id": 6,
+ "area": 11684,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2938,
+ "image_id": 556,
+ "bbox": [
+ 22,
+ 283,
+ 37,
+ 49
+ ],
+ "category_id": 6,
+ "area": 6580,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2939,
+ "image_id": 556,
+ "bbox": [
+ 49,
+ 187,
+ 103,
+ 115
+ ],
+ "category_id": 6,
+ "area": 41796,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2940,
+ "image_id": 556,
+ "bbox": [
+ 43,
+ 136,
+ 95,
+ 89
+ ],
+ "category_id": 6,
+ "area": 30114,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2942,
+ "image_id": 556,
+ "bbox": [
+ 252,
+ 120,
+ 131,
+ 228
+ ],
+ "category_id": 8,
+ "area": 105616,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2949,
+ "image_id": 556,
+ "bbox": [
+ 442,
+ 288,
+ 69,
+ 120
+ ],
+ "category_id": 6,
+ "area": 29580,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2950,
+ "image_id": 556,
+ "bbox": [
+ 165,
+ 338,
+ 346,
+ 172
+ ],
+ "category_id": 6,
+ "area": 210438,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2958,
+ "image_id": 558,
+ "bbox": [
+ 0,
+ 2,
+ 117,
+ 120
+ ],
+ "category_id": 6,
+ "area": 122475,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2959,
+ "image_id": 558,
+ "bbox": [
+ 0,
+ 114,
+ 45,
+ 96
+ ],
+ "category_id": 6,
+ "area": 38088,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2960,
+ "image_id": 558,
+ "bbox": [
+ 103,
+ 0,
+ 84,
+ 153
+ ],
+ "category_id": 6,
+ "area": 111945,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2961,
+ "image_id": 558,
+ "bbox": [
+ 93,
+ 146,
+ 72,
+ 108
+ ],
+ "category_id": 6,
+ "area": 68510,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2962,
+ "image_id": 558,
+ "bbox": [
+ 193,
+ 342,
+ 99,
+ 98
+ ],
+ "category_id": 6,
+ "area": 84280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2963,
+ "image_id": 558,
+ "bbox": [
+ 178,
+ 20,
+ 121,
+ 119
+ ],
+ "category_id": 6,
+ "area": 125829,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2964,
+ "image_id": 558,
+ "bbox": [
+ 253,
+ 0,
+ 95,
+ 106
+ ],
+ "category_id": 6,
+ "area": 87856,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2965,
+ "image_id": 558,
+ "bbox": [
+ 318,
+ 291,
+ 160,
+ 164
+ ],
+ "category_id": 6,
+ "area": 227934,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2971,
+ "image_id": 558,
+ "bbox": [
+ 387,
+ 88,
+ 87,
+ 68
+ ],
+ "category_id": 6,
+ "area": 51216,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2972,
+ "image_id": 558,
+ "bbox": [
+ 58,
+ 108,
+ 73,
+ 54
+ ],
+ "category_id": 6,
+ "area": 34720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2973,
+ "image_id": 559,
+ "bbox": [
+ 30,
+ 50,
+ 345,
+ 406
+ ],
+ "category_id": 8,
+ "area": 232389,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2979,
+ "image_id": 561,
+ "bbox": [
+ 92,
+ 25,
+ 271,
+ 437
+ ],
+ "category_id": 8,
+ "area": 138990,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2980,
+ "image_id": 561,
+ "bbox": [
+ 55,
+ 2,
+ 108,
+ 171
+ ],
+ "category_id": 6,
+ "area": 21735,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2981,
+ "image_id": 561,
+ "bbox": [
+ 68,
+ 98,
+ 98,
+ 113
+ ],
+ "category_id": 6,
+ "area": 13038,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2982,
+ "image_id": 561,
+ "bbox": [
+ 12,
+ 163,
+ 158,
+ 229
+ ],
+ "category_id": 6,
+ "area": 42570,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2983,
+ "image_id": 561,
+ "bbox": [
+ 120,
+ 376,
+ 100,
+ 133
+ ],
+ "category_id": 6,
+ "area": 15750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2988,
+ "image_id": 562,
+ "bbox": [
+ 356,
+ 377,
+ 146,
+ 66
+ ],
+ "category_id": 9,
+ "area": 8791,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2989,
+ "image_id": 563,
+ "bbox": [
+ 159,
+ 119,
+ 63,
+ 132
+ ],
+ "category_id": 8,
+ "area": 29415,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2990,
+ "image_id": 563,
+ "bbox": [
+ 242,
+ 234,
+ 114,
+ 249
+ ],
+ "category_id": 8,
+ "area": 100163,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2991,
+ "image_id": 564,
+ "bbox": [
+ 124,
+ 297,
+ 188,
+ 124
+ ],
+ "category_id": 8,
+ "area": 184710,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2992,
+ "image_id": 564,
+ "bbox": [
+ 217,
+ 110,
+ 68,
+ 105
+ ],
+ "category_id": 6,
+ "area": 56865,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2993,
+ "image_id": 564,
+ "bbox": [
+ 231,
+ 195,
+ 45,
+ 64
+ ],
+ "category_id": 6,
+ "area": 22984,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2994,
+ "image_id": 564,
+ "bbox": [
+ 0,
+ 98,
+ 226,
+ 169
+ ],
+ "category_id": 6,
+ "area": 303584,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 2995,
+ "image_id": 564,
+ "bbox": [
+ 146,
+ 224,
+ 70,
+ 47
+ ],
+ "category_id": 6,
+ "area": 26664,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3006,
+ "image_id": 566,
+ "bbox": [
+ 196,
+ 120,
+ 216,
+ 165
+ ],
+ "category_id": 8,
+ "area": 125744,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3007,
+ "image_id": 566,
+ "bbox": [
+ 219,
+ 235,
+ 196,
+ 161
+ ],
+ "category_id": 8,
+ "area": 110740,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3030,
+ "image_id": 572,
+ "bbox": [
+ 10,
+ 84,
+ 387,
+ 426
+ ],
+ "category_id": 9,
+ "area": 471744,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3041,
+ "image_id": 573,
+ "bbox": [
+ 94,
+ 130,
+ 304,
+ 381
+ ],
+ "category_id": 9,
+ "area": 174956,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3042,
+ "image_id": 574,
+ "bbox": [
+ 83,
+ 178,
+ 60,
+ 117
+ ],
+ "category_id": 8,
+ "area": 55575,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3043,
+ "image_id": 574,
+ "bbox": [
+ 193,
+ 186,
+ 44,
+ 30
+ ],
+ "category_id": 8,
+ "area": 10624,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3044,
+ "image_id": 574,
+ "bbox": [
+ 238,
+ 186,
+ 69,
+ 42
+ ],
+ "category_id": 8,
+ "area": 23229,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3045,
+ "image_id": 574,
+ "bbox": [
+ 296,
+ 261,
+ 87,
+ 97
+ ],
+ "category_id": 8,
+ "area": 67240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3046,
+ "image_id": 574,
+ "bbox": [
+ 388,
+ 230,
+ 75,
+ 70
+ ],
+ "category_id": 8,
+ "area": 41736,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3047,
+ "image_id": 574,
+ "bbox": [
+ 436,
+ 203,
+ 42,
+ 38
+ ],
+ "category_id": 8,
+ "area": 12956,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3048,
+ "image_id": 574,
+ "bbox": [
+ 469,
+ 219,
+ 42,
+ 50
+ ],
+ "category_id": 8,
+ "area": 17227,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3055,
+ "image_id": 577,
+ "bbox": [
+ 4,
+ 173,
+ 495,
+ 254
+ ],
+ "category_id": 9,
+ "area": 1929252,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3069,
+ "image_id": 579,
+ "bbox": [
+ 192,
+ 5,
+ 73,
+ 190
+ ],
+ "category_id": 6,
+ "area": 23598,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3070,
+ "image_id": 579,
+ "bbox": [
+ 238,
+ 76,
+ 195,
+ 159
+ ],
+ "category_id": 6,
+ "area": 52419,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3071,
+ "image_id": 579,
+ "bbox": [
+ 34,
+ 228,
+ 476,
+ 281
+ ],
+ "category_id": 6,
+ "area": 226134,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3072,
+ "image_id": 579,
+ "bbox": [
+ 127,
+ 199,
+ 384,
+ 253
+ ],
+ "category_id": 9,
+ "area": 164175,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3080,
+ "image_id": 581,
+ "bbox": [
+ 87,
+ 120,
+ 107,
+ 154
+ ],
+ "category_id": 10,
+ "area": 68272,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3118,
+ "image_id": 584,
+ "bbox": [
+ 49,
+ 242,
+ 201,
+ 202
+ ],
+ "category_id": 8,
+ "area": 142349,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3119,
+ "image_id": 584,
+ "bbox": [
+ 166,
+ 207,
+ 214,
+ 269
+ ],
+ "category_id": 8,
+ "area": 202072,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3120,
+ "image_id": 584,
+ "bbox": [
+ 307,
+ 57,
+ 110,
+ 110
+ ],
+ "category_id": 8,
+ "area": 42625,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3123,
+ "image_id": 586,
+ "bbox": [
+ 65,
+ 0,
+ 333,
+ 506
+ ],
+ "category_id": 9,
+ "area": 461472,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3124,
+ "image_id": 586,
+ "bbox": [
+ 320,
+ 0,
+ 191,
+ 172
+ ],
+ "category_id": 6,
+ "area": 90522,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3125,
+ "image_id": 587,
+ "bbox": [
+ 0,
+ 162,
+ 512,
+ 347
+ ],
+ "category_id": 9,
+ "area": 248850,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3126,
+ "image_id": 587,
+ "bbox": [
+ 0,
+ 163,
+ 219,
+ 346
+ ],
+ "category_id": 6,
+ "area": 106132,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3127,
+ "image_id": 587,
+ "bbox": [
+ 328,
+ 346,
+ 149,
+ 151
+ ],
+ "category_id": 6,
+ "area": 31510,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3137,
+ "image_id": 588,
+ "bbox": [
+ 76,
+ 147,
+ 347,
+ 274
+ ],
+ "category_id": 9,
+ "area": 592596,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3138,
+ "image_id": 588,
+ "bbox": [
+ 366,
+ 228,
+ 145,
+ 45
+ ],
+ "category_id": 9,
+ "area": 41283,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3139,
+ "image_id": 589,
+ "bbox": [
+ 145,
+ 207,
+ 310,
+ 216
+ ],
+ "category_id": 9,
+ "area": 172870,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3140,
+ "image_id": 589,
+ "bbox": [
+ 140,
+ 115,
+ 89,
+ 128
+ ],
+ "category_id": 6,
+ "area": 29754,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3141,
+ "image_id": 589,
+ "bbox": [
+ 255,
+ 0,
+ 157,
+ 152
+ ],
+ "category_id": 6,
+ "area": 62100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3142,
+ "image_id": 589,
+ "bbox": [
+ 392,
+ 0,
+ 92,
+ 118
+ ],
+ "category_id": 6,
+ "area": 28175,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3143,
+ "image_id": 590,
+ "bbox": [
+ 201,
+ 320,
+ 53,
+ 76
+ ],
+ "category_id": 6,
+ "area": 21574,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3144,
+ "image_id": 590,
+ "bbox": [
+ 257,
+ 254,
+ 112,
+ 95
+ ],
+ "category_id": 6,
+ "area": 57456,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3145,
+ "image_id": 590,
+ "bbox": [
+ 377,
+ 317,
+ 49,
+ 65
+ ],
+ "category_id": 6,
+ "area": 17250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3159,
+ "image_id": 593,
+ "bbox": [
+ 9,
+ 170,
+ 463,
+ 302
+ ],
+ "category_id": 9,
+ "area": 356570,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3160,
+ "image_id": 594,
+ "bbox": [
+ 75,
+ 122,
+ 280,
+ 234
+ ],
+ "category_id": 8,
+ "area": 520182,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3163,
+ "image_id": 594,
+ "bbox": [
+ 27,
+ 426,
+ 29,
+ 59
+ ],
+ "category_id": 6,
+ "area": 13750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3164,
+ "image_id": 594,
+ "bbox": [
+ 20,
+ 161,
+ 64,
+ 121
+ ],
+ "category_id": 6,
+ "area": 61952,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3165,
+ "image_id": 594,
+ "bbox": [
+ 265,
+ 146,
+ 246,
+ 362
+ ],
+ "category_id": 6,
+ "area": 706700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3190,
+ "image_id": 596,
+ "bbox": [
+ 189,
+ 169,
+ 47,
+ 60
+ ],
+ "category_id": 8,
+ "area": 10115,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3191,
+ "image_id": 596,
+ "bbox": [
+ 267,
+ 248,
+ 52,
+ 115
+ ],
+ "category_id": 8,
+ "area": 21060,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3192,
+ "image_id": 597,
+ "bbox": [
+ 182,
+ 172,
+ 230,
+ 151
+ ],
+ "category_id": 6,
+ "area": 239470,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3193,
+ "image_id": 597,
+ "bbox": [
+ 104,
+ 278,
+ 254,
+ 221
+ ],
+ "category_id": 8,
+ "area": 388907,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3194,
+ "image_id": 597,
+ "bbox": [
+ 220,
+ 34,
+ 291,
+ 307
+ ],
+ "category_id": 8,
+ "area": 615568,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3195,
+ "image_id": 598,
+ "bbox": [
+ 44,
+ 161,
+ 270,
+ 324
+ ],
+ "category_id": 8,
+ "area": 221446,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3216,
+ "image_id": 601,
+ "bbox": [
+ 87,
+ 249,
+ 358,
+ 261
+ ],
+ "category_id": 8,
+ "area": 58300,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3218,
+ "image_id": 602,
+ "bbox": [
+ 2,
+ 41,
+ 67,
+ 52
+ ],
+ "category_id": 6,
+ "area": 45140,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3219,
+ "image_id": 602,
+ "bbox": [
+ 1,
+ 60,
+ 67,
+ 123
+ ],
+ "category_id": 6,
+ "area": 105754,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3220,
+ "image_id": 602,
+ "bbox": [
+ 79,
+ 53,
+ 113,
+ 83
+ ],
+ "category_id": 6,
+ "area": 120360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3221,
+ "image_id": 602,
+ "bbox": [
+ 165,
+ 57,
+ 129,
+ 180
+ ],
+ "category_id": 6,
+ "area": 296376,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3222,
+ "image_id": 602,
+ "bbox": [
+ 164,
+ 232,
+ 106,
+ 78
+ ],
+ "category_id": 6,
+ "area": 106474,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3223,
+ "image_id": 602,
+ "bbox": [
+ 25,
+ 274,
+ 146,
+ 96
+ ],
+ "category_id": 6,
+ "area": 178992,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3224,
+ "image_id": 602,
+ "bbox": [
+ 2,
+ 377,
+ 113,
+ 83
+ ],
+ "category_id": 6,
+ "area": 119837,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3225,
+ "image_id": 602,
+ "bbox": [
+ 274,
+ 1,
+ 107,
+ 98
+ ],
+ "category_id": 6,
+ "area": 133980,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3226,
+ "image_id": 602,
+ "bbox": [
+ 188,
+ 51,
+ 59,
+ 53
+ ],
+ "category_id": 6,
+ "area": 40257,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3227,
+ "image_id": 602,
+ "bbox": [
+ 292,
+ 80,
+ 40,
+ 77
+ ],
+ "category_id": 6,
+ "area": 39712,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3228,
+ "image_id": 602,
+ "bbox": [
+ 311,
+ 127,
+ 74,
+ 65
+ ],
+ "category_id": 6,
+ "area": 61640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3229,
+ "image_id": 602,
+ "bbox": [
+ 392,
+ 47,
+ 36,
+ 49
+ ],
+ "category_id": 6,
+ "area": 22794,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3230,
+ "image_id": 602,
+ "bbox": [
+ 354,
+ 81,
+ 76,
+ 86
+ ],
+ "category_id": 6,
+ "area": 83875,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3231,
+ "image_id": 602,
+ "bbox": [
+ 399,
+ 120,
+ 110,
+ 99
+ ],
+ "category_id": 6,
+ "area": 138950,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3232,
+ "image_id": 602,
+ "bbox": [
+ 189,
+ 304,
+ 168,
+ 85
+ ],
+ "category_id": 6,
+ "area": 182707,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3243,
+ "image_id": 603,
+ "bbox": [
+ 14,
+ 445,
+ 112,
+ 65
+ ],
+ "category_id": 6,
+ "area": 20564,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3244,
+ "image_id": 603,
+ "bbox": [
+ 98,
+ 439,
+ 129,
+ 72
+ ],
+ "category_id": 6,
+ "area": 26215,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3247,
+ "image_id": 604,
+ "bbox": [
+ 0,
+ 230,
+ 159,
+ 240
+ ],
+ "category_id": 6,
+ "area": 204303,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3262,
+ "image_id": 607,
+ "bbox": [
+ 157,
+ 42,
+ 79,
+ 262
+ ],
+ "category_id": 10,
+ "area": 36852,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3263,
+ "image_id": 607,
+ "bbox": [
+ 221,
+ 0,
+ 91,
+ 86
+ ],
+ "category_id": 10,
+ "area": 13940,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3264,
+ "image_id": 607,
+ "bbox": [
+ 57,
+ 0,
+ 81,
+ 56
+ ],
+ "category_id": 10,
+ "area": 8100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3265,
+ "image_id": 607,
+ "bbox": [
+ 368,
+ 2,
+ 143,
+ 475
+ ],
+ "category_id": 10,
+ "area": 119515,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3266,
+ "image_id": 608,
+ "bbox": [
+ 2,
+ 15,
+ 111,
+ 141
+ ],
+ "category_id": 6,
+ "area": 191590,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3267,
+ "image_id": 608,
+ "bbox": [
+ 60,
+ 96,
+ 135,
+ 162
+ ],
+ "category_id": 6,
+ "area": 268375,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3268,
+ "image_id": 608,
+ "bbox": [
+ 125,
+ 17,
+ 116,
+ 94
+ ],
+ "category_id": 6,
+ "area": 135630,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3269,
+ "image_id": 608,
+ "bbox": [
+ 210,
+ 27,
+ 135,
+ 189
+ ],
+ "category_id": 6,
+ "area": 315480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3270,
+ "image_id": 608,
+ "bbox": [
+ 211,
+ 204,
+ 108,
+ 77
+ ],
+ "category_id": 6,
+ "area": 101840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3271,
+ "image_id": 608,
+ "bbox": [
+ 71,
+ 248,
+ 145,
+ 97
+ ],
+ "category_id": 6,
+ "area": 173568,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3272,
+ "image_id": 608,
+ "bbox": [
+ 13,
+ 346,
+ 143,
+ 95
+ ],
+ "category_id": 6,
+ "area": 166996,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3273,
+ "image_id": 608,
+ "bbox": [
+ 237,
+ 274,
+ 166,
+ 100
+ ],
+ "category_id": 6,
+ "area": 204165,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3274,
+ "image_id": 608,
+ "bbox": [
+ 366,
+ 98,
+ 67,
+ 72
+ ],
+ "category_id": 6,
+ "area": 59708,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3275,
+ "image_id": 608,
+ "bbox": [
+ 452,
+ 99,
+ 59,
+ 90
+ ],
+ "category_id": 6,
+ "area": 65728,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3276,
+ "image_id": 608,
+ "bbox": [
+ 408,
+ 51,
+ 75,
+ 96
+ ],
+ "category_id": 6,
+ "area": 89712,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3280,
+ "image_id": 609,
+ "bbox": [
+ 162,
+ 46,
+ 343,
+ 303
+ ],
+ "category_id": 9,
+ "area": 226730,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3281,
+ "image_id": 609,
+ "bbox": [
+ 103,
+ 177,
+ 152,
+ 259
+ ],
+ "category_id": 9,
+ "area": 86190,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3282,
+ "image_id": 609,
+ "bbox": [
+ 203,
+ 210,
+ 308,
+ 300
+ ],
+ "category_id": 9,
+ "area": 201365,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3283,
+ "image_id": 610,
+ "bbox": [
+ 0,
+ 26,
+ 353,
+ 483
+ ],
+ "category_id": 9,
+ "area": 173448,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3309,
+ "image_id": 616,
+ "bbox": [
+ 30,
+ 0,
+ 166,
+ 207
+ ],
+ "category_id": 8,
+ "area": 81796,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3311,
+ "image_id": 617,
+ "bbox": [
+ 1,
+ 17,
+ 506,
+ 489
+ ],
+ "category_id": 9,
+ "area": 872963,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3312,
+ "image_id": 617,
+ "bbox": [
+ 38,
+ 0,
+ 461,
+ 464
+ ],
+ "category_id": 9,
+ "area": 752909,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3313,
+ "image_id": 617,
+ "bbox": [
+ 32,
+ 91,
+ 450,
+ 363
+ ],
+ "category_id": 9,
+ "area": 575386,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3317,
+ "image_id": 619,
+ "bbox": [
+ 27,
+ 241,
+ 125,
+ 162
+ ],
+ "category_id": 8,
+ "area": 52974,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3318,
+ "image_id": 619,
+ "bbox": [
+ 104,
+ 230,
+ 315,
+ 280
+ ],
+ "category_id": 8,
+ "area": 229875,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3327,
+ "image_id": 623,
+ "bbox": [
+ 225,
+ 168,
+ 186,
+ 315
+ ],
+ "category_id": 10,
+ "area": 466165,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3328,
+ "image_id": 623,
+ "bbox": [
+ 84,
+ 125,
+ 28,
+ 52
+ ],
+ "category_id": 10,
+ "area": 11660,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3329,
+ "image_id": 623,
+ "bbox": [
+ 128,
+ 281,
+ 37,
+ 64
+ ],
+ "category_id": 10,
+ "area": 19312,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3330,
+ "image_id": 623,
+ "bbox": [
+ 326,
+ 143,
+ 26,
+ 46
+ ],
+ "category_id": 10,
+ "area": 9800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3331,
+ "image_id": 623,
+ "bbox": [
+ 98,
+ 353,
+ 31,
+ 56
+ ],
+ "category_id": 10,
+ "area": 14161,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3332,
+ "image_id": 623,
+ "bbox": [
+ 159,
+ 229,
+ 20,
+ 36
+ ],
+ "category_id": 10,
+ "area": 5700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3333,
+ "image_id": 623,
+ "bbox": [
+ 266,
+ 28,
+ 17,
+ 44
+ ],
+ "category_id": 10,
+ "area": 6138,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3334,
+ "image_id": 623,
+ "bbox": [
+ 20,
+ 374,
+ 28,
+ 55
+ ],
+ "category_id": 10,
+ "area": 12508,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3335,
+ "image_id": 624,
+ "bbox": [
+ 26,
+ 54,
+ 311,
+ 406
+ ],
+ "category_id": 9,
+ "area": 1112888,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3350,
+ "image_id": 626,
+ "bbox": [
+ 271,
+ 328,
+ 118,
+ 183
+ ],
+ "category_id": 6,
+ "area": 41790,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3351,
+ "image_id": 626,
+ "bbox": [
+ 119,
+ 61,
+ 273,
+ 333
+ ],
+ "category_id": 9,
+ "area": 175338,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3354,
+ "image_id": 629,
+ "bbox": [
+ 374,
+ 26,
+ 90,
+ 87
+ ],
+ "category_id": 8,
+ "area": 18683,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3355,
+ "image_id": 629,
+ "bbox": [
+ 169,
+ 0,
+ 215,
+ 128
+ ],
+ "category_id": 8,
+ "area": 65800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3356,
+ "image_id": 629,
+ "bbox": [
+ 5,
+ 31,
+ 506,
+ 457
+ ],
+ "category_id": 8,
+ "area": 549486,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3357,
+ "image_id": 629,
+ "bbox": [
+ 0,
+ 185,
+ 407,
+ 325
+ ],
+ "category_id": 8,
+ "area": 314530,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3359,
+ "image_id": 630,
+ "bbox": [
+ 175,
+ 153,
+ 111,
+ 228
+ ],
+ "category_id": 10,
+ "area": 85722,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3360,
+ "image_id": 630,
+ "bbox": [
+ 238,
+ 86,
+ 200,
+ 424
+ ],
+ "category_id": 9,
+ "area": 286836,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3363,
+ "image_id": 632,
+ "bbox": [
+ 0,
+ 50,
+ 194,
+ 101
+ ],
+ "category_id": 9,
+ "area": 21762,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3364,
+ "image_id": 632,
+ "bbox": [
+ 245,
+ 167,
+ 266,
+ 139
+ ],
+ "category_id": 9,
+ "area": 40767,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3400,
+ "image_id": 638,
+ "bbox": [
+ 111,
+ 99,
+ 267,
+ 160
+ ],
+ "category_id": 8,
+ "area": 150075,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3401,
+ "image_id": 638,
+ "bbox": [
+ 194,
+ 199,
+ 248,
+ 219
+ ],
+ "category_id": 8,
+ "area": 191268,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3444,
+ "image_id": 640,
+ "bbox": [
+ 211,
+ 228,
+ 93,
+ 194
+ ],
+ "category_id": 9,
+ "area": 43584,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3445,
+ "image_id": 640,
+ "bbox": [
+ 87,
+ 0,
+ 424,
+ 439
+ ],
+ "category_id": 6,
+ "area": 448000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3448,
+ "image_id": 642,
+ "bbox": [
+ 161,
+ 78,
+ 72,
+ 141
+ ],
+ "category_id": 8,
+ "area": 80730,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3449,
+ "image_id": 642,
+ "bbox": [
+ 252,
+ 109,
+ 77,
+ 283
+ ],
+ "category_id": 8,
+ "area": 173130,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3450,
+ "image_id": 642,
+ "bbox": [
+ 320,
+ 335,
+ 172,
+ 161
+ ],
+ "category_id": 8,
+ "area": 220968,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3475,
+ "image_id": 645,
+ "bbox": [
+ 116,
+ 75,
+ 295,
+ 410
+ ],
+ "category_id": 9,
+ "area": 160362,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3476,
+ "image_id": 645,
+ "bbox": [
+ 368,
+ 196,
+ 79,
+ 70
+ ],
+ "category_id": 9,
+ "area": 7442,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3514,
+ "image_id": 649,
+ "bbox": [
+ 55,
+ 138,
+ 302,
+ 230
+ ],
+ "category_id": 9,
+ "area": 184525,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3515,
+ "image_id": 649,
+ "bbox": [
+ 284,
+ 54,
+ 156,
+ 331
+ ],
+ "category_id": 10,
+ "area": 136968,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3525,
+ "image_id": 652,
+ "bbox": [
+ 69,
+ 48,
+ 69,
+ 107
+ ],
+ "category_id": 10,
+ "area": 58760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3563,
+ "image_id": 654,
+ "bbox": [
+ 360,
+ 6,
+ 94,
+ 172
+ ],
+ "category_id": 10,
+ "area": 29040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3564,
+ "image_id": 654,
+ "bbox": [
+ 1,
+ 423,
+ 37,
+ 59
+ ],
+ "category_id": 10,
+ "area": 3933,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3565,
+ "image_id": 654,
+ "bbox": [
+ 48,
+ 282,
+ 172,
+ 228
+ ],
+ "category_id": 10,
+ "area": 69861,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3566,
+ "image_id": 655,
+ "bbox": [
+ 64,
+ 166,
+ 235,
+ 182
+ ],
+ "category_id": 10,
+ "area": 68016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3568,
+ "image_id": 656,
+ "bbox": [
+ 174,
+ 253,
+ 139,
+ 159
+ ],
+ "category_id": 8,
+ "area": 77952,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3569,
+ "image_id": 656,
+ "bbox": [
+ 293,
+ 264,
+ 43,
+ 88
+ ],
+ "category_id": 8,
+ "area": 13392,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3572,
+ "image_id": 657,
+ "bbox": [
+ 164,
+ 179,
+ 87,
+ 110
+ ],
+ "category_id": 8,
+ "area": 33790,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3573,
+ "image_id": 657,
+ "bbox": [
+ 354,
+ 154,
+ 108,
+ 109
+ ],
+ "category_id": 8,
+ "area": 41888,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3574,
+ "image_id": 657,
+ "bbox": [
+ 449,
+ 288,
+ 33,
+ 35
+ ],
+ "category_id": 6,
+ "area": 4150,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3575,
+ "image_id": 657,
+ "bbox": [
+ 491,
+ 334,
+ 18,
+ 31
+ ],
+ "category_id": 6,
+ "area": 2068,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3576,
+ "image_id": 658,
+ "bbox": [
+ 1,
+ 325,
+ 148,
+ 151
+ ],
+ "category_id": 6,
+ "area": 62835,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3577,
+ "image_id": 658,
+ "bbox": [
+ 61,
+ 269,
+ 133,
+ 88
+ ],
+ "category_id": 6,
+ "area": 32860,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3578,
+ "image_id": 658,
+ "bbox": [
+ 403,
+ 26,
+ 108,
+ 206
+ ],
+ "category_id": 6,
+ "area": 62350,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3579,
+ "image_id": 658,
+ "bbox": [
+ 344,
+ 220,
+ 167,
+ 130
+ ],
+ "category_id": 6,
+ "area": 61456,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3677,
+ "image_id": 660,
+ "bbox": [
+ 0,
+ 325,
+ 314,
+ 186
+ ],
+ "category_id": 9,
+ "area": 87723,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3682,
+ "image_id": 661,
+ "bbox": [
+ 105,
+ 103,
+ 224,
+ 266
+ ],
+ "category_id": 8,
+ "area": 474609,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3684,
+ "image_id": 661,
+ "bbox": [
+ 0,
+ 201,
+ 191,
+ 309
+ ],
+ "category_id": 6,
+ "area": 468136,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3685,
+ "image_id": 661,
+ "bbox": [
+ 239,
+ 325,
+ 92,
+ 155
+ ],
+ "category_id": 6,
+ "area": 114163,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3686,
+ "image_id": 661,
+ "bbox": [
+ 405,
+ 404,
+ 83,
+ 105
+ ],
+ "category_id": 6,
+ "area": 69264,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3687,
+ "image_id": 661,
+ "bbox": [
+ 88,
+ 316,
+ 137,
+ 159
+ ],
+ "category_id": 6,
+ "area": 174229,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3688,
+ "image_id": 662,
+ "bbox": [
+ 132,
+ 6,
+ 379,
+ 372
+ ],
+ "category_id": 8,
+ "area": 206968,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3689,
+ "image_id": 662,
+ "bbox": [
+ 129,
+ 247,
+ 311,
+ 239
+ ],
+ "category_id": 8,
+ "area": 109509,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3706,
+ "image_id": 665,
+ "bbox": [
+ 12,
+ 268,
+ 379,
+ 143
+ ],
+ "category_id": 8,
+ "area": 49840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3709,
+ "image_id": 666,
+ "bbox": [
+ 89,
+ 454,
+ 92,
+ 54
+ ],
+ "category_id": 6,
+ "area": 37605,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3710,
+ "image_id": 666,
+ "bbox": [
+ 0,
+ 318,
+ 102,
+ 113
+ ],
+ "category_id": 6,
+ "area": 86714,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3713,
+ "image_id": 667,
+ "bbox": [
+ 70,
+ 224,
+ 183,
+ 210
+ ],
+ "category_id": 8,
+ "area": 135405,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3714,
+ "image_id": 667,
+ "bbox": [
+ 181,
+ 210,
+ 216,
+ 275
+ ],
+ "category_id": 8,
+ "area": 209212,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3715,
+ "image_id": 667,
+ "bbox": [
+ 337,
+ 52,
+ 110,
+ 110
+ ],
+ "category_id": 8,
+ "area": 42780,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3722,
+ "image_id": 670,
+ "bbox": [
+ 73,
+ 195,
+ 161,
+ 108
+ ],
+ "category_id": 8,
+ "area": 138396,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3723,
+ "image_id": 670,
+ "bbox": [
+ 262,
+ 215,
+ 176,
+ 97
+ ],
+ "category_id": 8,
+ "area": 136372,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3726,
+ "image_id": 672,
+ "bbox": [
+ 72,
+ 41,
+ 329,
+ 470
+ ],
+ "category_id": 10,
+ "area": 193050,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3740,
+ "image_id": 674,
+ "bbox": [
+ 89,
+ 249,
+ 259,
+ 229
+ ],
+ "category_id": 9,
+ "area": 34428,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3773,
+ "image_id": 677,
+ "bbox": [
+ 93,
+ 198,
+ 177,
+ 207
+ ],
+ "category_id": 8,
+ "area": 25785,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3774,
+ "image_id": 677,
+ "bbox": [
+ 227,
+ 152,
+ 161,
+ 90
+ ],
+ "category_id": 8,
+ "area": 10266,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3783,
+ "image_id": 678,
+ "bbox": [
+ 259,
+ 301,
+ 252,
+ 210
+ ],
+ "category_id": 6,
+ "area": 388977,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3784,
+ "image_id": 678,
+ "bbox": [
+ 0,
+ 231,
+ 305,
+ 280
+ ],
+ "category_id": 6,
+ "area": 627435,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3785,
+ "image_id": 678,
+ "bbox": [
+ 193,
+ 0,
+ 312,
+ 246
+ ],
+ "category_id": 9,
+ "area": 563152,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3797,
+ "image_id": 681,
+ "bbox": [
+ 7,
+ 140,
+ 503,
+ 371
+ ],
+ "category_id": 9,
+ "area": 159125,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3806,
+ "image_id": 683,
+ "bbox": [
+ 264,
+ 207,
+ 247,
+ 288
+ ],
+ "category_id": 8,
+ "area": 82852,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3823,
+ "image_id": 689,
+ "bbox": [
+ 79,
+ 70,
+ 326,
+ 440
+ ],
+ "category_id": 10,
+ "area": 505104,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3825,
+ "image_id": 690,
+ "bbox": [
+ 89,
+ 0,
+ 375,
+ 477
+ ],
+ "category_id": 10,
+ "area": 138684,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3831,
+ "image_id": 693,
+ "bbox": [
+ 20,
+ 123,
+ 340,
+ 329
+ ],
+ "category_id": 8,
+ "area": 130592,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3842,
+ "image_id": 695,
+ "bbox": [
+ 131,
+ 68,
+ 242,
+ 315
+ ],
+ "category_id": 9,
+ "area": 268901,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3844,
+ "image_id": 696,
+ "bbox": [
+ 38,
+ 125,
+ 401,
+ 385
+ ],
+ "category_id": 8,
+ "area": 417105,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3845,
+ "image_id": 696,
+ "bbox": [
+ 215,
+ 11,
+ 295,
+ 383
+ ],
+ "category_id": 8,
+ "area": 305025,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3847,
+ "image_id": 697,
+ "bbox": [
+ 212,
+ 28,
+ 143,
+ 226
+ ],
+ "category_id": 6,
+ "area": 268926,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3848,
+ "image_id": 697,
+ "bbox": [
+ 271,
+ 346,
+ 181,
+ 82
+ ],
+ "category_id": 6,
+ "area": 123540,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3849,
+ "image_id": 698,
+ "bbox": [
+ 8,
+ 5,
+ 176,
+ 251
+ ],
+ "category_id": 6,
+ "area": 59648,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3850,
+ "image_id": 698,
+ "bbox": [
+ 0,
+ 246,
+ 209,
+ 258
+ ],
+ "category_id": 9,
+ "area": 72960,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3851,
+ "image_id": 699,
+ "bbox": [
+ 84,
+ 101,
+ 197,
+ 321
+ ],
+ "category_id": 6,
+ "area": 220908,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3859,
+ "image_id": 700,
+ "bbox": [
+ 94,
+ 140,
+ 342,
+ 300
+ ],
+ "category_id": 9,
+ "area": 128232,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3889,
+ "image_id": 703,
+ "bbox": [
+ 248,
+ 222,
+ 177,
+ 155
+ ],
+ "category_id": 8,
+ "area": 96574,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3890,
+ "image_id": 703,
+ "bbox": [
+ 218,
+ 222,
+ 64,
+ 135
+ ],
+ "category_id": 8,
+ "area": 30780,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3907,
+ "image_id": 706,
+ "bbox": [
+ 267,
+ 230,
+ 129,
+ 186
+ ],
+ "category_id": 10,
+ "area": 49115,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3908,
+ "image_id": 706,
+ "bbox": [
+ 202,
+ 33,
+ 309,
+ 245
+ ],
+ "category_id": 9,
+ "area": 153725,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3909,
+ "image_id": 706,
+ "bbox": [
+ 2,
+ 12,
+ 509,
+ 495
+ ],
+ "category_id": 6,
+ "area": 511710,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3910,
+ "image_id": 707,
+ "bbox": [
+ 176,
+ 364,
+ 44,
+ 130
+ ],
+ "category_id": 8,
+ "area": 18260,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3911,
+ "image_id": 707,
+ "bbox": [
+ 231,
+ 119,
+ 70,
+ 109
+ ],
+ "category_id": 8,
+ "area": 24640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3913,
+ "image_id": 708,
+ "bbox": [
+ 195,
+ 56,
+ 257,
+ 365
+ ],
+ "category_id": 8,
+ "area": 110446,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3915,
+ "image_id": 709,
+ "bbox": [
+ 143,
+ 130,
+ 368,
+ 304
+ ],
+ "category_id": 8,
+ "area": 393267,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3916,
+ "image_id": 709,
+ "bbox": [
+ 285,
+ 16,
+ 226,
+ 247
+ ],
+ "category_id": 8,
+ "area": 196402,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3917,
+ "image_id": 709,
+ "bbox": [
+ 0,
+ 2,
+ 240,
+ 505
+ ],
+ "category_id": 6,
+ "area": 426216,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3948,
+ "image_id": 713,
+ "bbox": [
+ 202,
+ 163,
+ 163,
+ 168
+ ],
+ "category_id": 6,
+ "area": 108808,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3949,
+ "image_id": 713,
+ "bbox": [
+ 272,
+ 292,
+ 142,
+ 217
+ ],
+ "category_id": 6,
+ "area": 123176,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3950,
+ "image_id": 713,
+ "bbox": [
+ 414,
+ 350,
+ 56,
+ 60
+ ],
+ "category_id": 6,
+ "area": 13440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3951,
+ "image_id": 713,
+ "bbox": [
+ 167,
+ 437,
+ 117,
+ 72
+ ],
+ "category_id": 6,
+ "area": 33872,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3952,
+ "image_id": 713,
+ "bbox": [
+ 415,
+ 460,
+ 95,
+ 48
+ ],
+ "category_id": 6,
+ "area": 18642,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3953,
+ "image_id": 713,
+ "bbox": [
+ 0,
+ 451,
+ 78,
+ 59
+ ],
+ "category_id": 6,
+ "area": 18620,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3972,
+ "image_id": 713,
+ "bbox": [
+ 104,
+ 357,
+ 73,
+ 85
+ ],
+ "category_id": 6,
+ "area": 24888,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3973,
+ "image_id": 713,
+ "bbox": [
+ 108,
+ 307,
+ 71,
+ 57
+ ],
+ "category_id": 6,
+ "area": 16289,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3974,
+ "image_id": 713,
+ "bbox": [
+ 63,
+ 239,
+ 87,
+ 68
+ ],
+ "category_id": 6,
+ "area": 23653,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3993,
+ "image_id": 717,
+ "bbox": [
+ 225,
+ 42,
+ 164,
+ 140
+ ],
+ "category_id": 8,
+ "area": 80967,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3994,
+ "image_id": 717,
+ "bbox": [
+ 70,
+ 159,
+ 343,
+ 306
+ ],
+ "category_id": 8,
+ "area": 368082,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3995,
+ "image_id": 717,
+ "bbox": [
+ 374,
+ 280,
+ 37,
+ 64
+ ],
+ "category_id": 8,
+ "area": 8554,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3996,
+ "image_id": 718,
+ "bbox": [
+ 203,
+ 150,
+ 268,
+ 263
+ ],
+ "category_id": 8,
+ "area": 558330,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3998,
+ "image_id": 719,
+ "bbox": [
+ 156,
+ 201,
+ 154,
+ 303
+ ],
+ "category_id": 10,
+ "area": 165249,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 3999,
+ "image_id": 719,
+ "bbox": [
+ 222,
+ 0,
+ 288,
+ 332
+ ],
+ "category_id": 9,
+ "area": 336707,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4010,
+ "image_id": 722,
+ "bbox": [
+ 152,
+ 188,
+ 222,
+ 126
+ ],
+ "category_id": 10,
+ "area": 44394,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4012,
+ "image_id": 723,
+ "bbox": [
+ 0,
+ 0,
+ 377,
+ 508
+ ],
+ "category_id": 10,
+ "area": 1518295,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4029,
+ "image_id": 725,
+ "bbox": [
+ 230,
+ 169,
+ 124,
+ 86
+ ],
+ "category_id": 9,
+ "area": 37942,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4030,
+ "image_id": 725,
+ "bbox": [
+ 258,
+ 114,
+ 143,
+ 145
+ ],
+ "category_id": 9,
+ "area": 73032,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4032,
+ "image_id": 726,
+ "bbox": [
+ 2,
+ 81,
+ 507,
+ 430
+ ],
+ "category_id": 9,
+ "area": 767140,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4033,
+ "image_id": 726,
+ "bbox": [
+ 252,
+ 1,
+ 260,
+ 506
+ ],
+ "category_id": 9,
+ "area": 462800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4043,
+ "image_id": 729,
+ "bbox": [
+ 246,
+ 61,
+ 106,
+ 128
+ ],
+ "category_id": 9,
+ "area": 47880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4044,
+ "image_id": 729,
+ "bbox": [
+ 240,
+ 61,
+ 62,
+ 126
+ ],
+ "category_id": 9,
+ "area": 27768,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4045,
+ "image_id": 729,
+ "bbox": [
+ 208,
+ 219,
+ 135,
+ 185
+ ],
+ "category_id": 9,
+ "area": 88218,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4046,
+ "image_id": 730,
+ "bbox": [
+ 112,
+ 157,
+ 307,
+ 282
+ ],
+ "category_id": 10,
+ "area": 231478,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4047,
+ "image_id": 731,
+ "bbox": [
+ 0,
+ 223,
+ 16,
+ 130
+ ],
+ "category_id": 6,
+ "area": 4788,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4048,
+ "image_id": 731,
+ "bbox": [
+ 199,
+ 294,
+ 311,
+ 214
+ ],
+ "category_id": 9,
+ "area": 149548,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4053,
+ "image_id": 733,
+ "bbox": [
+ 325,
+ 142,
+ 184,
+ 364
+ ],
+ "category_id": 9,
+ "area": 235980,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4054,
+ "image_id": 733,
+ "bbox": [
+ 29,
+ 90,
+ 394,
+ 418
+ ],
+ "category_id": 9,
+ "area": 579180,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4055,
+ "image_id": 734,
+ "bbox": [
+ 154,
+ 302,
+ 35,
+ 112
+ ],
+ "category_id": 8,
+ "area": 12672,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4056,
+ "image_id": 734,
+ "bbox": [
+ 199,
+ 136,
+ 74,
+ 112
+ ],
+ "category_id": 8,
+ "area": 26825,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4058,
+ "image_id": 735,
+ "bbox": [
+ 53,
+ 123,
+ 242,
+ 342
+ ],
+ "category_id": 6,
+ "area": 389052,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4059,
+ "image_id": 735,
+ "bbox": [
+ 140,
+ 258,
+ 201,
+ 166
+ ],
+ "category_id": 6,
+ "area": 157248,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4061,
+ "image_id": 736,
+ "bbox": [
+ 0,
+ 147,
+ 512,
+ 363
+ ],
+ "category_id": 9,
+ "area": 274992,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4062,
+ "image_id": 736,
+ "bbox": [
+ 313,
+ 57,
+ 198,
+ 223
+ ],
+ "category_id": 6,
+ "area": 65619,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4063,
+ "image_id": 737,
+ "bbox": [
+ 115,
+ 99,
+ 200,
+ 382
+ ],
+ "category_id": 8,
+ "area": 101238,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4082,
+ "image_id": 742,
+ "bbox": [
+ 156,
+ 95,
+ 181,
+ 323
+ ],
+ "category_id": 8,
+ "area": 204756,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4083,
+ "image_id": 742,
+ "bbox": [
+ 317,
+ 371,
+ 74,
+ 110
+ ],
+ "category_id": 6,
+ "area": 28520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4084,
+ "image_id": 742,
+ "bbox": [
+ 210,
+ 344,
+ 48,
+ 120
+ ],
+ "category_id": 6,
+ "area": 20449,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4085,
+ "image_id": 743,
+ "bbox": [
+ 109,
+ 51,
+ 297,
+ 457
+ ],
+ "category_id": 10,
+ "area": 1074860,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4086,
+ "image_id": 744,
+ "bbox": [
+ 0,
+ 33,
+ 273,
+ 477
+ ],
+ "category_id": 9,
+ "area": 459648,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4087,
+ "image_id": 744,
+ "bbox": [
+ 179,
+ 0,
+ 330,
+ 509
+ ],
+ "category_id": 9,
+ "area": 591525,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4089,
+ "image_id": 744,
+ "bbox": [
+ 49,
+ 0,
+ 322,
+ 334
+ ],
+ "category_id": 9,
+ "area": 380097,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4104,
+ "image_id": 747,
+ "bbox": [
+ 2,
+ 0,
+ 122,
+ 164
+ ],
+ "category_id": 6,
+ "area": 197722,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4105,
+ "image_id": 747,
+ "bbox": [
+ 150,
+ 0,
+ 137,
+ 120
+ ],
+ "category_id": 6,
+ "area": 162792,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4106,
+ "image_id": 747,
+ "bbox": [
+ 3,
+ 158,
+ 165,
+ 149
+ ],
+ "category_id": 6,
+ "area": 241878,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4107,
+ "image_id": 747,
+ "bbox": [
+ 1,
+ 273,
+ 91,
+ 75
+ ],
+ "category_id": 6,
+ "area": 67872,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4108,
+ "image_id": 747,
+ "bbox": [
+ 172,
+ 190,
+ 179,
+ 101
+ ],
+ "category_id": 6,
+ "area": 177900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4109,
+ "image_id": 747,
+ "bbox": [
+ 306,
+ 273,
+ 90,
+ 108
+ ],
+ "category_id": 6,
+ "area": 95956,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4110,
+ "image_id": 747,
+ "bbox": [
+ 300,
+ 0,
+ 96,
+ 54
+ ],
+ "category_id": 6,
+ "area": 51359,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4111,
+ "image_id": 747,
+ "bbox": [
+ 400,
+ 1,
+ 111,
+ 99
+ ],
+ "category_id": 6,
+ "area": 108265,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4112,
+ "image_id": 747,
+ "bbox": [
+ 443,
+ 200,
+ 68,
+ 117
+ ],
+ "category_id": 6,
+ "area": 78996,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4116,
+ "image_id": 748,
+ "bbox": [
+ 251,
+ 243,
+ 43,
+ 73
+ ],
+ "category_id": 8,
+ "area": 25272,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4130,
+ "image_id": 755,
+ "bbox": [
+ 189,
+ 103,
+ 130,
+ 141
+ ],
+ "category_id": 8,
+ "area": 145722,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4131,
+ "image_id": 755,
+ "bbox": [
+ 15,
+ 141,
+ 137,
+ 136
+ ],
+ "category_id": 6,
+ "area": 148896,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4132,
+ "image_id": 755,
+ "bbox": [
+ 353,
+ 169,
+ 66,
+ 90
+ ],
+ "category_id": 6,
+ "area": 47750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4133,
+ "image_id": 755,
+ "bbox": [
+ 174,
+ 401,
+ 112,
+ 109
+ ],
+ "category_id": 6,
+ "area": 97251,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4134,
+ "image_id": 755,
+ "bbox": [
+ 307,
+ 354,
+ 41,
+ 110
+ ],
+ "category_id": 6,
+ "area": 36115,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4155,
+ "image_id": 759,
+ "bbox": [
+ 411,
+ 373,
+ 76,
+ 137
+ ],
+ "category_id": 9,
+ "area": 26878,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4156,
+ "image_id": 759,
+ "bbox": [
+ 31,
+ 161,
+ 95,
+ 248
+ ],
+ "category_id": 9,
+ "area": 60536,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4157,
+ "image_id": 759,
+ "bbox": [
+ 155,
+ 287,
+ 85,
+ 214
+ ],
+ "category_id": 9,
+ "area": 46704,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4158,
+ "image_id": 759,
+ "bbox": [
+ 172,
+ 213,
+ 246,
+ 296
+ ],
+ "category_id": 9,
+ "area": 186624,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4159,
+ "image_id": 759,
+ "bbox": [
+ 241,
+ 94,
+ 170,
+ 275
+ ],
+ "category_id": 9,
+ "area": 120309,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4160,
+ "image_id": 759,
+ "bbox": [
+ 190,
+ 39,
+ 167,
+ 273
+ ],
+ "category_id": 9,
+ "area": 117505,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4161,
+ "image_id": 760,
+ "bbox": [
+ 116,
+ 195,
+ 195,
+ 189
+ ],
+ "category_id": 8,
+ "area": 129055,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4162,
+ "image_id": 760,
+ "bbox": [
+ 168,
+ 19,
+ 343,
+ 469
+ ],
+ "category_id": 8,
+ "area": 560880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4163,
+ "image_id": 760,
+ "bbox": [
+ 0,
+ 2,
+ 331,
+ 509
+ ],
+ "category_id": 6,
+ "area": 586688,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4164,
+ "image_id": 761,
+ "bbox": [
+ 184,
+ 176,
+ 177,
+ 159
+ ],
+ "category_id": 9,
+ "area": 43955,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4174,
+ "image_id": 763,
+ "bbox": [
+ 39,
+ 222,
+ 49,
+ 69
+ ],
+ "category_id": 8,
+ "area": 6794,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4175,
+ "image_id": 763,
+ "bbox": [
+ 116,
+ 294,
+ 239,
+ 160
+ ],
+ "category_id": 8,
+ "area": 76258,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4176,
+ "image_id": 763,
+ "bbox": [
+ 178,
+ 219,
+ 125,
+ 118
+ ],
+ "category_id": 8,
+ "area": 29565,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4190,
+ "image_id": 765,
+ "bbox": [
+ 29,
+ 405,
+ 96,
+ 98
+ ],
+ "category_id": 6,
+ "area": 26535,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4191,
+ "image_id": 765,
+ "bbox": [
+ 107,
+ 403,
+ 120,
+ 108
+ ],
+ "category_id": 6,
+ "area": 36411,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4205,
+ "image_id": 767,
+ "bbox": [
+ 76,
+ 308,
+ 180,
+ 199
+ ],
+ "category_id": 6,
+ "area": 42262,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4206,
+ "image_id": 767,
+ "bbox": [
+ 200,
+ 118,
+ 93,
+ 96
+ ],
+ "category_id": 6,
+ "area": 10530,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4207,
+ "image_id": 767,
+ "bbox": [
+ 325,
+ 158,
+ 40,
+ 73
+ ],
+ "category_id": 6,
+ "area": 3519,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4208,
+ "image_id": 767,
+ "bbox": [
+ 93,
+ 20,
+ 52,
+ 66
+ ],
+ "category_id": 6,
+ "area": 4092,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4209,
+ "image_id": 768,
+ "bbox": [
+ 0,
+ 93,
+ 484,
+ 413
+ ],
+ "category_id": 9,
+ "area": 704220,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4210,
+ "image_id": 768,
+ "bbox": [
+ 1,
+ 0,
+ 395,
+ 384
+ ],
+ "category_id": 9,
+ "area": 534508,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4211,
+ "image_id": 768,
+ "bbox": [
+ 0,
+ 14,
+ 509,
+ 493
+ ],
+ "category_id": 9,
+ "area": 883462,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4216,
+ "image_id": 770,
+ "bbox": [
+ 115,
+ 218,
+ 86,
+ 107
+ ],
+ "category_id": 8,
+ "area": 32465,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4222,
+ "image_id": 770,
+ "bbox": [
+ 226,
+ 311,
+ 218,
+ 192
+ ],
+ "category_id": 6,
+ "area": 146067,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4223,
+ "image_id": 771,
+ "bbox": [
+ 124,
+ 267,
+ 138,
+ 81
+ ],
+ "category_id": 9,
+ "area": 39675,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4234,
+ "image_id": 774,
+ "bbox": [
+ 18,
+ 0,
+ 492,
+ 510
+ ],
+ "category_id": 6,
+ "area": 197046,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4235,
+ "image_id": 775,
+ "bbox": [
+ 43,
+ 191,
+ 467,
+ 319
+ ],
+ "category_id": 8,
+ "area": 183893,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4237,
+ "image_id": 776,
+ "bbox": [
+ 45,
+ 285,
+ 329,
+ 185
+ ],
+ "category_id": 8,
+ "area": 472622,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4238,
+ "image_id": 776,
+ "bbox": [
+ 0,
+ 0,
+ 511,
+ 355
+ ],
+ "category_id": 8,
+ "area": 1402962,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4239,
+ "image_id": 776,
+ "bbox": [
+ 83,
+ 104,
+ 123,
+ 169
+ ],
+ "category_id": 8,
+ "area": 161700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4240,
+ "image_id": 777,
+ "bbox": [
+ 107,
+ 237,
+ 288,
+ 219
+ ],
+ "category_id": 8,
+ "area": 427966,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4241,
+ "image_id": 777,
+ "bbox": [
+ 173,
+ 17,
+ 196,
+ 392
+ ],
+ "category_id": 8,
+ "area": 521115,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4242,
+ "image_id": 777,
+ "bbox": [
+ 187,
+ 27,
+ 253,
+ 235
+ ],
+ "category_id": 8,
+ "area": 402900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4243,
+ "image_id": 778,
+ "bbox": [
+ 113,
+ 126,
+ 176,
+ 150
+ ],
+ "category_id": 9,
+ "area": 18920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4255,
+ "image_id": 781,
+ "bbox": [
+ 0,
+ 105,
+ 200,
+ 249
+ ],
+ "category_id": 6,
+ "area": 116656,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4258,
+ "image_id": 782,
+ "bbox": [
+ 228,
+ 135,
+ 183,
+ 173
+ ],
+ "category_id": 8,
+ "area": 250755,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4259,
+ "image_id": 782,
+ "bbox": [
+ 259,
+ 149,
+ 147,
+ 216
+ ],
+ "category_id": 8,
+ "area": 252624,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4260,
+ "image_id": 782,
+ "bbox": [
+ 175,
+ 192,
+ 99,
+ 99
+ ],
+ "category_id": 8,
+ "area": 77957,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4261,
+ "image_id": 783,
+ "bbox": [
+ 304,
+ 56,
+ 75,
+ 278
+ ],
+ "category_id": 6,
+ "area": 24534,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4262,
+ "image_id": 783,
+ "bbox": [
+ 64,
+ 5,
+ 97,
+ 110
+ ],
+ "category_id": 6,
+ "area": 12688,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4263,
+ "image_id": 783,
+ "bbox": [
+ 247,
+ 260,
+ 23,
+ 53
+ ],
+ "category_id": 6,
+ "area": 1450,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4344,
+ "image_id": 795,
+ "bbox": [
+ 22,
+ 315,
+ 239,
+ 185
+ ],
+ "category_id": 6,
+ "area": 49928,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4345,
+ "image_id": 795,
+ "bbox": [
+ 0,
+ 264,
+ 73,
+ 152
+ ],
+ "category_id": 6,
+ "area": 12610,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4346,
+ "image_id": 795,
+ "bbox": [
+ 288,
+ 380,
+ 67,
+ 123
+ ],
+ "category_id": 6,
+ "area": 9345,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4347,
+ "image_id": 795,
+ "bbox": [
+ 356,
+ 409,
+ 44,
+ 90
+ ],
+ "category_id": 6,
+ "area": 4543,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4348,
+ "image_id": 795,
+ "bbox": [
+ 395,
+ 373,
+ 40,
+ 71
+ ],
+ "category_id": 6,
+ "area": 3294,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4349,
+ "image_id": 795,
+ "bbox": [
+ 395,
+ 442,
+ 31,
+ 65
+ ],
+ "category_id": 6,
+ "area": 2296,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4352,
+ "image_id": 795,
+ "bbox": [
+ 251,
+ 326,
+ 91,
+ 77
+ ],
+ "category_id": 9,
+ "area": 7986,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4353,
+ "image_id": 795,
+ "bbox": [
+ 340,
+ 321,
+ 100,
+ 56
+ ],
+ "category_id": 6,
+ "area": 6336,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4357,
+ "image_id": 797,
+ "bbox": [
+ 0,
+ 5,
+ 435,
+ 280
+ ],
+ "category_id": 6,
+ "area": 307746,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4367,
+ "image_id": 799,
+ "bbox": [
+ 104,
+ 130,
+ 380,
+ 353
+ ],
+ "category_id": 9,
+ "area": 157556,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4368,
+ "image_id": 800,
+ "bbox": [
+ 187,
+ 341,
+ 142,
+ 140
+ ],
+ "category_id": 9,
+ "area": 26320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4372,
+ "image_id": 802,
+ "bbox": [
+ 226,
+ 0,
+ 282,
+ 86
+ ],
+ "category_id": 9,
+ "area": 85426,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4373,
+ "image_id": 802,
+ "bbox": [
+ 176,
+ 49,
+ 54,
+ 456
+ ],
+ "category_id": 9,
+ "area": 87312,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4374,
+ "image_id": 802,
+ "bbox": [
+ 2,
+ 49,
+ 505,
+ 457
+ ],
+ "category_id": 9,
+ "area": 813372,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4375,
+ "image_id": 803,
+ "bbox": [
+ 259,
+ 189,
+ 250,
+ 318
+ ],
+ "category_id": 9,
+ "area": 280896,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4376,
+ "image_id": 803,
+ "bbox": [
+ 182,
+ 0,
+ 327,
+ 433
+ ],
+ "category_id": 9,
+ "area": 498162,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4377,
+ "image_id": 803,
+ "bbox": [
+ 210,
+ 0,
+ 299,
+ 430
+ ],
+ "category_id": 9,
+ "area": 452540,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4383,
+ "image_id": 804,
+ "bbox": [
+ 177,
+ 208,
+ 61,
+ 62
+ ],
+ "category_id": 6,
+ "area": 13398,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4384,
+ "image_id": 804,
+ "bbox": [
+ 100,
+ 227,
+ 69,
+ 91
+ ],
+ "category_id": 6,
+ "area": 22144,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4385,
+ "image_id": 804,
+ "bbox": [
+ 215,
+ 304,
+ 43,
+ 124
+ ],
+ "category_id": 6,
+ "area": 18618,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4386,
+ "image_id": 804,
+ "bbox": [
+ 255,
+ 339,
+ 71,
+ 130
+ ],
+ "category_id": 6,
+ "area": 32578,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4387,
+ "image_id": 804,
+ "bbox": [
+ 239,
+ 169,
+ 68,
+ 105
+ ],
+ "category_id": 6,
+ "area": 25137,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4388,
+ "image_id": 805,
+ "bbox": [
+ 111,
+ 205,
+ 162,
+ 194
+ ],
+ "category_id": 8,
+ "area": 110432,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4389,
+ "image_id": 805,
+ "bbox": [
+ 268,
+ 159,
+ 242,
+ 234
+ ],
+ "category_id": 8,
+ "area": 199045,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4391,
+ "image_id": 805,
+ "bbox": [
+ 369,
+ 124,
+ 30,
+ 40
+ ],
+ "category_id": 6,
+ "area": 4389,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4392,
+ "image_id": 805,
+ "bbox": [
+ 406,
+ 170,
+ 29,
+ 36
+ ],
+ "category_id": 6,
+ "area": 3723,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4393,
+ "image_id": 805,
+ "bbox": [
+ 447,
+ 212,
+ 23,
+ 27
+ ],
+ "category_id": 6,
+ "area": 2262,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4394,
+ "image_id": 806,
+ "bbox": [
+ 206,
+ 186,
+ 147,
+ 169
+ ],
+ "category_id": 8,
+ "area": 87216,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4395,
+ "image_id": 806,
+ "bbox": [
+ 323,
+ 218,
+ 74,
+ 99
+ ],
+ "category_id": 8,
+ "area": 25993,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4400,
+ "image_id": 808,
+ "bbox": [
+ 32,
+ 397,
+ 98,
+ 93
+ ],
+ "category_id": 6,
+ "area": 25668,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4401,
+ "image_id": 808,
+ "bbox": [
+ 103,
+ 391,
+ 124,
+ 120
+ ],
+ "category_id": 6,
+ "area": 41772,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4402,
+ "image_id": 809,
+ "bbox": [
+ 1,
+ 113,
+ 337,
+ 310
+ ],
+ "category_id": 9,
+ "area": 260445,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4422,
+ "image_id": 812,
+ "bbox": [
+ 195,
+ 154,
+ 73,
+ 166
+ ],
+ "category_id": 10,
+ "area": 37840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4423,
+ "image_id": 813,
+ "bbox": [
+ 10,
+ 276,
+ 27,
+ 73
+ ],
+ "category_id": 6,
+ "area": 13860,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4424,
+ "image_id": 813,
+ "bbox": [
+ 479,
+ 436,
+ 21,
+ 75
+ ],
+ "category_id": 6,
+ "area": 11200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4425,
+ "image_id": 813,
+ "bbox": [
+ 338,
+ 308,
+ 27,
+ 45
+ ],
+ "category_id": 6,
+ "area": 8640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4426,
+ "image_id": 813,
+ "bbox": [
+ 488,
+ 241,
+ 16,
+ 35
+ ],
+ "category_id": 6,
+ "area": 4050,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4427,
+ "image_id": 813,
+ "bbox": [
+ 498,
+ 217,
+ 13,
+ 37
+ ],
+ "category_id": 6,
+ "area": 3600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4434,
+ "image_id": 815,
+ "bbox": [
+ 0,
+ 228,
+ 349,
+ 283
+ ],
+ "category_id": 6,
+ "area": 248976,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4435,
+ "image_id": 815,
+ "bbox": [
+ 116,
+ 0,
+ 351,
+ 490
+ ],
+ "category_id": 6,
+ "area": 434172,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4437,
+ "image_id": 815,
+ "bbox": [
+ 295,
+ 376,
+ 216,
+ 132
+ ],
+ "category_id": 9,
+ "area": 72063,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4443,
+ "image_id": 817,
+ "bbox": [
+ 81,
+ 199,
+ 284,
+ 218
+ ],
+ "category_id": 9,
+ "area": 56943,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4452,
+ "image_id": 820,
+ "bbox": [
+ 211,
+ 249,
+ 86,
+ 68
+ ],
+ "category_id": 8,
+ "area": 20736,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4453,
+ "image_id": 820,
+ "bbox": [
+ 315,
+ 189,
+ 84,
+ 134
+ ],
+ "category_id": 8,
+ "area": 40068,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4454,
+ "image_id": 821,
+ "bbox": [
+ 0,
+ 110,
+ 442,
+ 398
+ ],
+ "category_id": 9,
+ "area": 529230,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4458,
+ "image_id": 823,
+ "bbox": [
+ 296,
+ 251,
+ 110,
+ 117
+ ],
+ "category_id": 10,
+ "area": 13287,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4459,
+ "image_id": 823,
+ "bbox": [
+ 188,
+ 142,
+ 267,
+ 128
+ ],
+ "category_id": 9,
+ "area": 35369,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4460,
+ "image_id": 823,
+ "bbox": [
+ 75,
+ 316,
+ 249,
+ 195
+ ],
+ "category_id": 9,
+ "area": 50224,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4461,
+ "image_id": 824,
+ "bbox": [
+ 380,
+ 333,
+ 86,
+ 177
+ ],
+ "category_id": 9,
+ "area": 39330,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4462,
+ "image_id": 824,
+ "bbox": [
+ 8,
+ 153,
+ 142,
+ 257
+ ],
+ "category_id": 9,
+ "area": 93854,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4463,
+ "image_id": 824,
+ "bbox": [
+ 147,
+ 261,
+ 78,
+ 238
+ ],
+ "category_id": 9,
+ "area": 47586,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4464,
+ "image_id": 824,
+ "bbox": [
+ 196,
+ 189,
+ 183,
+ 295
+ ],
+ "category_id": 9,
+ "area": 138646,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4465,
+ "image_id": 824,
+ "bbox": [
+ 238,
+ 85,
+ 157,
+ 245
+ ],
+ "category_id": 9,
+ "area": 98580,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4466,
+ "image_id": 824,
+ "bbox": [
+ 188,
+ 24,
+ 152,
+ 282
+ ],
+ "category_id": 9,
+ "area": 110100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4467,
+ "image_id": 825,
+ "bbox": [
+ 188,
+ 0,
+ 277,
+ 199
+ ],
+ "category_id": 6,
+ "area": 155344,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4468,
+ "image_id": 825,
+ "bbox": [
+ 253,
+ 144,
+ 249,
+ 357
+ ],
+ "category_id": 9,
+ "area": 249375,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4469,
+ "image_id": 826,
+ "bbox": [
+ 63,
+ 195,
+ 196,
+ 183
+ ],
+ "category_id": 8,
+ "area": 125930,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4470,
+ "image_id": 826,
+ "bbox": [
+ 191,
+ 194,
+ 241,
+ 208
+ ],
+ "category_id": 8,
+ "area": 176076,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4493,
+ "image_id": 828,
+ "bbox": [
+ 126,
+ 248,
+ 328,
+ 206
+ ],
+ "category_id": 8,
+ "area": 537588,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4494,
+ "image_id": 828,
+ "bbox": [
+ 245,
+ 139,
+ 100,
+ 66
+ ],
+ "category_id": 8,
+ "area": 53298,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4495,
+ "image_id": 828,
+ "bbox": [
+ 196,
+ 149,
+ 39,
+ 48
+ ],
+ "category_id": 8,
+ "area": 15244,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4496,
+ "image_id": 829,
+ "bbox": [
+ 213,
+ 145,
+ 280,
+ 202
+ ],
+ "category_id": 9,
+ "area": 88803,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4497,
+ "image_id": 829,
+ "bbox": [
+ 50,
+ 224,
+ 239,
+ 109
+ ],
+ "category_id": 9,
+ "area": 40963,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4498,
+ "image_id": 829,
+ "bbox": [
+ 30,
+ 329,
+ 192,
+ 151
+ ],
+ "category_id": 9,
+ "area": 45549,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4499,
+ "image_id": 829,
+ "bbox": [
+ 413,
+ 368,
+ 77,
+ 38
+ ],
+ "category_id": 9,
+ "area": 4656,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4500,
+ "image_id": 830,
+ "bbox": [
+ 228,
+ 40,
+ 213,
+ 387
+ ],
+ "category_id": 9,
+ "area": 290485,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4501,
+ "image_id": 830,
+ "bbox": [
+ 171,
+ 108,
+ 247,
+ 319
+ ],
+ "category_id": 9,
+ "area": 277482,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4502,
+ "image_id": 831,
+ "bbox": [
+ 0,
+ 0,
+ 280,
+ 510
+ ],
+ "category_id": 6,
+ "area": 1133028,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4518,
+ "image_id": 833,
+ "bbox": [
+ 0,
+ 202,
+ 158,
+ 246
+ ],
+ "category_id": 6,
+ "area": 58624,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4519,
+ "image_id": 833,
+ "bbox": [
+ 0,
+ 148,
+ 511,
+ 361
+ ],
+ "category_id": 9,
+ "area": 277045,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4520,
+ "image_id": 833,
+ "bbox": [
+ 223,
+ 0,
+ 89,
+ 81
+ ],
+ "category_id": 6,
+ "area": 11020,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4522,
+ "image_id": 834,
+ "bbox": [
+ 215,
+ 250,
+ 206,
+ 148
+ ],
+ "category_id": 9,
+ "area": 49248,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4523,
+ "image_id": 835,
+ "bbox": [
+ 153,
+ 128,
+ 238,
+ 222
+ ],
+ "category_id": 9,
+ "area": 116772,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4531,
+ "image_id": 837,
+ "bbox": [
+ 136,
+ 33,
+ 223,
+ 358
+ ],
+ "category_id": 8,
+ "area": 548052,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4536,
+ "image_id": 840,
+ "bbox": [
+ 226,
+ 74,
+ 183,
+ 401
+ ],
+ "category_id": 9,
+ "area": 239502,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4537,
+ "image_id": 840,
+ "bbox": [
+ 201,
+ 235,
+ 110,
+ 184
+ ],
+ "category_id": 10,
+ "area": 66443,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4538,
+ "image_id": 841,
+ "bbox": [
+ 193,
+ 71,
+ 300,
+ 435
+ ],
+ "category_id": 9,
+ "area": 523340,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4540,
+ "image_id": 843,
+ "bbox": [
+ 138,
+ 67,
+ 52,
+ 82
+ ],
+ "category_id": 8,
+ "area": 14950,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4541,
+ "image_id": 843,
+ "bbox": [
+ 257,
+ 156,
+ 83,
+ 181
+ ],
+ "category_id": 8,
+ "area": 52832,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4554,
+ "image_id": 845,
+ "bbox": [
+ 0,
+ 84,
+ 379,
+ 383
+ ],
+ "category_id": 9,
+ "area": 362313,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4555,
+ "image_id": 845,
+ "bbox": [
+ 408,
+ 198,
+ 103,
+ 280
+ ],
+ "category_id": 6,
+ "area": 71874,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4556,
+ "image_id": 845,
+ "bbox": [
+ 224,
+ 132,
+ 10,
+ 48
+ ],
+ "category_id": 6,
+ "area": 1323,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4568,
+ "image_id": 850,
+ "bbox": [
+ 165,
+ 20,
+ 248,
+ 349
+ ],
+ "category_id": 9,
+ "area": 369810,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4569,
+ "image_id": 851,
+ "bbox": [
+ 157,
+ 65,
+ 136,
+ 264
+ ],
+ "category_id": 10,
+ "area": 117196,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4570,
+ "image_id": 851,
+ "bbox": [
+ 174,
+ 0,
+ 271,
+ 454
+ ],
+ "category_id": 9,
+ "area": 399406,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4577,
+ "image_id": 853,
+ "bbox": [
+ 232,
+ 109,
+ 68,
+ 113
+ ],
+ "category_id": 8,
+ "area": 61920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4578,
+ "image_id": 853,
+ "bbox": [
+ 317,
+ 21,
+ 84,
+ 112
+ ],
+ "category_id": 6,
+ "area": 75208,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4586,
+ "image_id": 855,
+ "bbox": [
+ 111,
+ 162,
+ 277,
+ 349
+ ],
+ "category_id": 10,
+ "area": 768258,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4598,
+ "image_id": 858,
+ "bbox": [
+ 82,
+ 113,
+ 381,
+ 398
+ ],
+ "category_id": 10,
+ "area": 6270990,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4602,
+ "image_id": 860,
+ "bbox": [
+ 97,
+ 140,
+ 123,
+ 152
+ ],
+ "category_id": 6,
+ "area": 92379,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4603,
+ "image_id": 860,
+ "bbox": [
+ 313,
+ 98,
+ 195,
+ 207
+ ],
+ "category_id": 6,
+ "area": 197730,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4644,
+ "image_id": 865,
+ "bbox": [
+ 23,
+ 11,
+ 426,
+ 415
+ ],
+ "category_id": 6,
+ "area": 530370,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4660,
+ "image_id": 867,
+ "bbox": [
+ 140,
+ 85,
+ 26,
+ 27
+ ],
+ "category_id": 6,
+ "area": 2613,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4661,
+ "image_id": 867,
+ "bbox": [
+ 175,
+ 132,
+ 31,
+ 46
+ ],
+ "category_id": 6,
+ "area": 5070,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4662,
+ "image_id": 867,
+ "bbox": [
+ 156,
+ 177,
+ 38,
+ 44
+ ],
+ "category_id": 6,
+ "area": 6014,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4663,
+ "image_id": 867,
+ "bbox": [
+ 210,
+ 169,
+ 20,
+ 24
+ ],
+ "category_id": 6,
+ "area": 1700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4664,
+ "image_id": 867,
+ "bbox": [
+ 86,
+ 442,
+ 54,
+ 57
+ ],
+ "category_id": 6,
+ "area": 10960,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4665,
+ "image_id": 867,
+ "bbox": [
+ 12,
+ 213,
+ 269,
+ 241
+ ],
+ "category_id": 8,
+ "area": 227812,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4666,
+ "image_id": 867,
+ "bbox": [
+ 198,
+ 199,
+ 231,
+ 257
+ ],
+ "category_id": 8,
+ "area": 208080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4668,
+ "image_id": 868,
+ "bbox": [
+ 136,
+ 47,
+ 324,
+ 436
+ ],
+ "category_id": 10,
+ "area": 1117800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4676,
+ "image_id": 869,
+ "bbox": [
+ 239,
+ 144,
+ 126,
+ 189
+ ],
+ "category_id": 8,
+ "area": 84372,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4677,
+ "image_id": 869,
+ "bbox": [
+ 46,
+ 139,
+ 97,
+ 91
+ ],
+ "category_id": 6,
+ "area": 31347,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4678,
+ "image_id": 869,
+ "bbox": [
+ 50,
+ 192,
+ 90,
+ 116
+ ],
+ "category_id": 6,
+ "area": 36900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4679,
+ "image_id": 869,
+ "bbox": [
+ 65,
+ 299,
+ 48,
+ 76
+ ],
+ "category_id": 6,
+ "area": 13054,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4680,
+ "image_id": 869,
+ "bbox": [
+ 266,
+ 338,
+ 90,
+ 91
+ ],
+ "category_id": 6,
+ "area": 29154,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4681,
+ "image_id": 869,
+ "bbox": [
+ 371,
+ 337,
+ 140,
+ 174
+ ],
+ "category_id": 6,
+ "area": 85995,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4682,
+ "image_id": 869,
+ "bbox": [
+ 282,
+ 374,
+ 133,
+ 137
+ ],
+ "category_id": 6,
+ "area": 64462,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4683,
+ "image_id": 869,
+ "bbox": [
+ 156,
+ 408,
+ 148,
+ 103
+ ],
+ "category_id": 6,
+ "area": 54312,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4684,
+ "image_id": 869,
+ "bbox": [
+ 180,
+ 282,
+ 42,
+ 52
+ ],
+ "category_id": 6,
+ "area": 7844,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4685,
+ "image_id": 869,
+ "bbox": [
+ 442,
+ 281,
+ 69,
+ 120
+ ],
+ "category_id": 6,
+ "area": 29410,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4686,
+ "image_id": 870,
+ "bbox": [
+ 82,
+ 190,
+ 144,
+ 134
+ ],
+ "category_id": 8,
+ "area": 153669,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4687,
+ "image_id": 870,
+ "bbox": [
+ 187,
+ 193,
+ 120,
+ 91
+ ],
+ "category_id": 8,
+ "area": 87882,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4690,
+ "image_id": 872,
+ "bbox": [
+ 0,
+ 119,
+ 510,
+ 389
+ ],
+ "category_id": 6,
+ "area": 301490,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4691,
+ "image_id": 873,
+ "bbox": [
+ 112,
+ 24,
+ 396,
+ 482
+ ],
+ "category_id": 9,
+ "area": 671898,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4694,
+ "image_id": 875,
+ "bbox": [
+ 23,
+ 29,
+ 487,
+ 481
+ ],
+ "category_id": 6,
+ "area": 848232,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4732,
+ "image_id": 878,
+ "bbox": [
+ 261,
+ 147,
+ 79,
+ 143
+ ],
+ "category_id": 9,
+ "area": 40198,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4733,
+ "image_id": 878,
+ "bbox": [
+ 253,
+ 205,
+ 70,
+ 86
+ ],
+ "category_id": 9,
+ "area": 21417,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4736,
+ "image_id": 880,
+ "bbox": [
+ 324,
+ 277,
+ 27,
+ 70
+ ],
+ "category_id": 6,
+ "area": 13764,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4737,
+ "image_id": 880,
+ "bbox": [
+ 307,
+ 354,
+ 19,
+ 39
+ ],
+ "category_id": 6,
+ "area": 5376,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4745,
+ "image_id": 882,
+ "bbox": [
+ 121,
+ 453,
+ 114,
+ 57
+ ],
+ "category_id": 6,
+ "area": 10017,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4746,
+ "image_id": 882,
+ "bbox": [
+ 205,
+ 430,
+ 67,
+ 70
+ ],
+ "category_id": 6,
+ "area": 7215,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4747,
+ "image_id": 882,
+ "bbox": [
+ 53,
+ 470,
+ 65,
+ 40
+ ],
+ "category_id": 6,
+ "area": 3996,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4748,
+ "image_id": 882,
+ "bbox": [
+ 203,
+ 394,
+ 66,
+ 52
+ ],
+ "category_id": 6,
+ "area": 5232,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4749,
+ "image_id": 882,
+ "bbox": [
+ 276,
+ 425,
+ 51,
+ 58
+ ],
+ "category_id": 6,
+ "area": 4590,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4750,
+ "image_id": 882,
+ "bbox": [
+ 382,
+ 472,
+ 86,
+ 36
+ ],
+ "category_id": 6,
+ "area": 4719,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4755,
+ "image_id": 885,
+ "bbox": [
+ 106,
+ 51,
+ 257,
+ 206
+ ],
+ "category_id": 8,
+ "area": 419775,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4756,
+ "image_id": 885,
+ "bbox": [
+ 118,
+ 166,
+ 236,
+ 191
+ ],
+ "category_id": 8,
+ "area": 357058,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4771,
+ "image_id": 888,
+ "bbox": [
+ 234,
+ 267,
+ 121,
+ 73
+ ],
+ "category_id": 8,
+ "area": 70525,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4772,
+ "image_id": 888,
+ "bbox": [
+ 181,
+ 246,
+ 78,
+ 114
+ ],
+ "category_id": 8,
+ "area": 71148,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4773,
+ "image_id": 888,
+ "bbox": [
+ 232,
+ 232,
+ 52,
+ 26
+ ],
+ "category_id": 8,
+ "area": 11088,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4774,
+ "image_id": 888,
+ "bbox": [
+ 334,
+ 199,
+ 34,
+ 53
+ ],
+ "category_id": 6,
+ "area": 14448,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4780,
+ "image_id": 891,
+ "bbox": [
+ 153,
+ 17,
+ 299,
+ 368
+ ],
+ "category_id": 9,
+ "area": 185399,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4788,
+ "image_id": 892,
+ "bbox": [
+ 10,
+ 167,
+ 219,
+ 204
+ ],
+ "category_id": 8,
+ "area": 157563,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4789,
+ "image_id": 892,
+ "bbox": [
+ 201,
+ 107,
+ 310,
+ 264
+ ],
+ "category_id": 8,
+ "area": 288300,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4790,
+ "image_id": 893,
+ "bbox": [
+ 76,
+ 195,
+ 242,
+ 203
+ ],
+ "category_id": 8,
+ "area": 127916,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4791,
+ "image_id": 893,
+ "bbox": [
+ 260,
+ 148,
+ 250,
+ 233
+ ],
+ "category_id": 8,
+ "area": 151308,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4799,
+ "image_id": 895,
+ "bbox": [
+ 170,
+ 275,
+ 169,
+ 215
+ ],
+ "category_id": 10,
+ "area": 30135,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4800,
+ "image_id": 895,
+ "bbox": [
+ 232,
+ 64,
+ 276,
+ 249
+ ],
+ "category_id": 9,
+ "area": 56950,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4807,
+ "image_id": 897,
+ "bbox": [
+ 190,
+ 213,
+ 198,
+ 177
+ ],
+ "category_id": 8,
+ "area": 123753,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4808,
+ "image_id": 897,
+ "bbox": [
+ 305,
+ 104,
+ 204,
+ 129
+ ],
+ "category_id": 8,
+ "area": 92491,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4809,
+ "image_id": 898,
+ "bbox": [
+ 214,
+ 269,
+ 150,
+ 117
+ ],
+ "category_id": 8,
+ "area": 61828,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4813,
+ "image_id": 900,
+ "bbox": [
+ 182,
+ 81,
+ 114,
+ 247
+ ],
+ "category_id": 9,
+ "area": 40414,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4814,
+ "image_id": 900,
+ "bbox": [
+ 359,
+ 374,
+ 68,
+ 82
+ ],
+ "category_id": 6,
+ "area": 8181,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4825,
+ "image_id": 902,
+ "bbox": [
+ 1,
+ 306,
+ 47,
+ 75
+ ],
+ "category_id": 10,
+ "area": 28640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4826,
+ "image_id": 902,
+ "bbox": [
+ 125,
+ 368,
+ 37,
+ 65
+ ],
+ "category_id": 10,
+ "area": 19182,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4827,
+ "image_id": 902,
+ "bbox": [
+ 47,
+ 426,
+ 28,
+ 47
+ ],
+ "category_id": 10,
+ "area": 10908,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4828,
+ "image_id": 902,
+ "bbox": [
+ 151,
+ 209,
+ 33,
+ 50
+ ],
+ "category_id": 10,
+ "area": 13589,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4829,
+ "image_id": 902,
+ "bbox": [
+ 398,
+ 328,
+ 40,
+ 73
+ ],
+ "category_id": 10,
+ "area": 23560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4830,
+ "image_id": 902,
+ "bbox": [
+ 242,
+ 350,
+ 77,
+ 127
+ ],
+ "category_id": 10,
+ "area": 78279,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4831,
+ "image_id": 902,
+ "bbox": [
+ 221,
+ 133,
+ 130,
+ 232
+ ],
+ "category_id": 10,
+ "area": 239608,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4832,
+ "image_id": 902,
+ "bbox": [
+ 194,
+ 66,
+ 22,
+ 38
+ ],
+ "category_id": 10,
+ "area": 6970,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4833,
+ "image_id": 902,
+ "bbox": [
+ 7,
+ 208,
+ 22,
+ 36
+ ],
+ "category_id": 10,
+ "area": 6308,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4834,
+ "image_id": 903,
+ "bbox": [
+ 396,
+ 360,
+ 104,
+ 151
+ ],
+ "category_id": 9,
+ "area": 40376,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4835,
+ "image_id": 903,
+ "bbox": [
+ 23,
+ 153,
+ 109,
+ 237
+ ],
+ "category_id": 9,
+ "area": 66528,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4836,
+ "image_id": 903,
+ "bbox": [
+ 144,
+ 276,
+ 94,
+ 208
+ ],
+ "category_id": 9,
+ "area": 50406,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4837,
+ "image_id": 903,
+ "bbox": [
+ 185,
+ 0,
+ 199,
+ 308
+ ],
+ "category_id": 9,
+ "area": 157200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4838,
+ "image_id": 903,
+ "bbox": [
+ 243,
+ 160,
+ 177,
+ 207
+ ],
+ "category_id": 9,
+ "area": 94419,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4839,
+ "image_id": 903,
+ "bbox": [
+ 202,
+ 205,
+ 155,
+ 306
+ ],
+ "category_id": 9,
+ "area": 121788,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4855,
+ "image_id": 906,
+ "bbox": [
+ 0,
+ 201,
+ 13,
+ 113
+ ],
+ "category_id": 6,
+ "area": 3480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4856,
+ "image_id": 906,
+ "bbox": [
+ 201,
+ 271,
+ 310,
+ 239
+ ],
+ "category_id": 9,
+ "area": 166896,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4863,
+ "image_id": 907,
+ "bbox": [
+ 141,
+ 269,
+ 120,
+ 71
+ ],
+ "category_id": 9,
+ "area": 17064,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4872,
+ "image_id": 909,
+ "bbox": [
+ 2,
+ 40,
+ 131,
+ 101
+ ],
+ "category_id": 6,
+ "area": 172914,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4873,
+ "image_id": 909,
+ "bbox": [
+ 7,
+ 110,
+ 70,
+ 139
+ ],
+ "category_id": 6,
+ "area": 127194,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4874,
+ "image_id": 909,
+ "bbox": [
+ 40,
+ 24,
+ 114,
+ 74
+ ],
+ "category_id": 6,
+ "area": 110880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4875,
+ "image_id": 909,
+ "bbox": [
+ 2,
+ 242,
+ 59,
+ 89
+ ],
+ "category_id": 6,
+ "area": 68888,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4876,
+ "image_id": 909,
+ "bbox": [
+ 75,
+ 268,
+ 67,
+ 103
+ ],
+ "category_id": 6,
+ "area": 90520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4877,
+ "image_id": 909,
+ "bbox": [
+ 80,
+ 402,
+ 85,
+ 83
+ ],
+ "category_id": 6,
+ "area": 91709,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4878,
+ "image_id": 909,
+ "bbox": [
+ 133,
+ 127,
+ 78,
+ 80
+ ],
+ "category_id": 6,
+ "area": 81221,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4879,
+ "image_id": 909,
+ "bbox": [
+ 214,
+ 7,
+ 227,
+ 143
+ ],
+ "category_id": 6,
+ "area": 422180,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4880,
+ "image_id": 909,
+ "bbox": [
+ 195,
+ 340,
+ 121,
+ 158
+ ],
+ "category_id": 6,
+ "area": 249760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4881,
+ "image_id": 909,
+ "bbox": [
+ 474,
+ 26,
+ 37,
+ 73
+ ],
+ "category_id": 6,
+ "area": 35880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4882,
+ "image_id": 909,
+ "bbox": [
+ 412,
+ 119,
+ 99,
+ 104
+ ],
+ "category_id": 6,
+ "area": 135420,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4883,
+ "image_id": 909,
+ "bbox": [
+ 242,
+ 200,
+ 74,
+ 49
+ ],
+ "category_id": 6,
+ "area": 47056,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4884,
+ "image_id": 909,
+ "bbox": [
+ 81,
+ 146,
+ 89,
+ 87
+ ],
+ "category_id": 6,
+ "area": 101970,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4885,
+ "image_id": 909,
+ "bbox": [
+ 387,
+ 6,
+ 84,
+ 58
+ ],
+ "category_id": 6,
+ "area": 63345,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4890,
+ "image_id": 911,
+ "bbox": [
+ 83,
+ 8,
+ 123,
+ 191
+ ],
+ "category_id": 6,
+ "area": 62496,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4891,
+ "image_id": 911,
+ "bbox": [
+ 161,
+ 87,
+ 350,
+ 424
+ ],
+ "category_id": 6,
+ "area": 392977,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4897,
+ "image_id": 913,
+ "bbox": [
+ 108,
+ 173,
+ 65,
+ 117
+ ],
+ "category_id": 9,
+ "area": 11664,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4945,
+ "image_id": 916,
+ "bbox": [
+ 321,
+ 256,
+ 56,
+ 255
+ ],
+ "category_id": 6,
+ "area": 98010,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4947,
+ "image_id": 917,
+ "bbox": [
+ 169,
+ 71,
+ 342,
+ 433
+ ],
+ "category_id": 8,
+ "area": 522160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4948,
+ "image_id": 917,
+ "bbox": [
+ 1,
+ 0,
+ 442,
+ 511
+ ],
+ "category_id": 6,
+ "area": 795214,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4949,
+ "image_id": 918,
+ "bbox": [
+ 2,
+ 30,
+ 24,
+ 82
+ ],
+ "category_id": 6,
+ "area": 38645,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4950,
+ "image_id": 918,
+ "bbox": [
+ 3,
+ 110,
+ 57,
+ 58
+ ],
+ "category_id": 6,
+ "area": 64581,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4951,
+ "image_id": 918,
+ "bbox": [
+ 1,
+ 404,
+ 31,
+ 78
+ ],
+ "category_id": 6,
+ "area": 47940,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4952,
+ "image_id": 918,
+ "bbox": [
+ 53,
+ 361,
+ 88,
+ 130
+ ],
+ "category_id": 6,
+ "area": 222300,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4953,
+ "image_id": 918,
+ "bbox": [
+ 86,
+ 202,
+ 49,
+ 59
+ ],
+ "category_id": 6,
+ "area": 57352,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4954,
+ "image_id": 918,
+ "bbox": [
+ 37,
+ 24,
+ 201,
+ 126
+ ],
+ "category_id": 6,
+ "area": 490872,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4955,
+ "image_id": 918,
+ "bbox": [
+ 199,
+ 12,
+ 39,
+ 61
+ ],
+ "category_id": 6,
+ "area": 47073,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4956,
+ "image_id": 918,
+ "bbox": [
+ 246,
+ 17,
+ 74,
+ 80
+ ],
+ "category_id": 6,
+ "area": 115661,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4957,
+ "image_id": 918,
+ "bbox": [
+ 296,
+ 74,
+ 91,
+ 127
+ ],
+ "category_id": 6,
+ "area": 226176,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4958,
+ "image_id": 918,
+ "bbox": [
+ 202,
+ 94,
+ 85,
+ 151
+ ],
+ "category_id": 6,
+ "area": 251409,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4959,
+ "image_id": 918,
+ "bbox": [
+ 174,
+ 241,
+ 144,
+ 192
+ ],
+ "category_id": 6,
+ "area": 538289,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4960,
+ "image_id": 918,
+ "bbox": [
+ 316,
+ 270,
+ 108,
+ 81
+ ],
+ "category_id": 6,
+ "area": 171698,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4961,
+ "image_id": 918,
+ "bbox": [
+ 402,
+ 96,
+ 55,
+ 65
+ ],
+ "category_id": 6,
+ "area": 70265,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4962,
+ "image_id": 918,
+ "bbox": [
+ 461,
+ 97,
+ 50,
+ 85
+ ],
+ "category_id": 6,
+ "area": 83811,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4968,
+ "image_id": 919,
+ "bbox": [
+ 96,
+ 290,
+ 121,
+ 66
+ ],
+ "category_id": 6,
+ "area": 490958,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4969,
+ "image_id": 919,
+ "bbox": [
+ 401,
+ 110,
+ 110,
+ 194
+ ],
+ "category_id": 6,
+ "area": 1301515,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4970,
+ "image_id": 919,
+ "bbox": [
+ 47,
+ 288,
+ 464,
+ 194
+ ],
+ "category_id": 6,
+ "area": 5477202,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4981,
+ "image_id": 922,
+ "bbox": [
+ 1,
+ 138,
+ 62,
+ 125
+ ],
+ "category_id": 6,
+ "area": 63388,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4982,
+ "image_id": 922,
+ "bbox": [
+ 1,
+ 343,
+ 85,
+ 117
+ ],
+ "category_id": 6,
+ "area": 81189,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4983,
+ "image_id": 922,
+ "bbox": [
+ 108,
+ 269,
+ 147,
+ 212
+ ],
+ "category_id": 6,
+ "area": 253005,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4984,
+ "image_id": 922,
+ "bbox": [
+ 171,
+ 37,
+ 78,
+ 79
+ ],
+ "category_id": 6,
+ "area": 50384,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4990,
+ "image_id": 923,
+ "bbox": [
+ 0,
+ 0,
+ 512,
+ 442
+ ],
+ "category_id": 6,
+ "area": 557770,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4992,
+ "image_id": 925,
+ "bbox": [
+ 175,
+ 190,
+ 145,
+ 90
+ ],
+ "category_id": 8,
+ "area": 45847,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4993,
+ "image_id": 925,
+ "bbox": [
+ 216,
+ 128,
+ 114,
+ 111
+ ],
+ "category_id": 8,
+ "area": 44616,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 4994,
+ "image_id": 925,
+ "bbox": [
+ 204,
+ 186,
+ 307,
+ 322
+ ],
+ "category_id": 6,
+ "area": 345015,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5011,
+ "image_id": 928,
+ "bbox": [
+ 119,
+ 243,
+ 136,
+ 112
+ ],
+ "category_id": 9,
+ "area": 53720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5017,
+ "image_id": 931,
+ "bbox": [
+ 172,
+ 166,
+ 248,
+ 171
+ ],
+ "category_id": 9,
+ "area": 201289,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5018,
+ "image_id": 932,
+ "bbox": [
+ 44,
+ 117,
+ 460,
+ 275
+ ],
+ "category_id": 9,
+ "area": 100580,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5025,
+ "image_id": 935,
+ "bbox": [
+ 1,
+ 95,
+ 510,
+ 338
+ ],
+ "category_id": 9,
+ "area": 245532,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5026,
+ "image_id": 935,
+ "bbox": [
+ 105,
+ 54,
+ 61,
+ 97
+ ],
+ "category_id": 9,
+ "area": 8554,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5040,
+ "image_id": 938,
+ "bbox": [
+ 30,
+ 260,
+ 239,
+ 154
+ ],
+ "category_id": 9,
+ "area": 62865,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5041,
+ "image_id": 938,
+ "bbox": [
+ 389,
+ 157,
+ 81,
+ 242
+ ],
+ "category_id": 9,
+ "area": 33670,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5074,
+ "image_id": 943,
+ "bbox": [
+ 277,
+ 22,
+ 172,
+ 228
+ ],
+ "category_id": 8,
+ "area": 171160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5075,
+ "image_id": 943,
+ "bbox": [
+ 147,
+ 131,
+ 190,
+ 379
+ ],
+ "category_id": 8,
+ "area": 314442,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5076,
+ "image_id": 943,
+ "bbox": [
+ 325,
+ 302,
+ 169,
+ 208
+ ],
+ "category_id": 8,
+ "area": 153792,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5085,
+ "image_id": 946,
+ "bbox": [
+ 180,
+ 153,
+ 201,
+ 335
+ ],
+ "category_id": 9,
+ "area": 171784,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5095,
+ "image_id": 950,
+ "bbox": [
+ 57,
+ 101,
+ 353,
+ 248
+ ],
+ "category_id": 9,
+ "area": 184314,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5098,
+ "image_id": 951,
+ "bbox": [
+ 261,
+ 56,
+ 94,
+ 96
+ ],
+ "category_id": 9,
+ "area": 32232,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5099,
+ "image_id": 951,
+ "bbox": [
+ 271,
+ 92,
+ 120,
+ 169
+ ],
+ "category_id": 9,
+ "area": 71638,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5102,
+ "image_id": 952,
+ "bbox": [
+ 0,
+ 121,
+ 50,
+ 114
+ ],
+ "category_id": 8,
+ "area": 20320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5103,
+ "image_id": 952,
+ "bbox": [
+ 163,
+ 262,
+ 154,
+ 127
+ ],
+ "category_id": 8,
+ "area": 69094,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5104,
+ "image_id": 952,
+ "bbox": [
+ 272,
+ 311,
+ 110,
+ 97
+ ],
+ "category_id": 8,
+ "area": 37536,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5105,
+ "image_id": 953,
+ "bbox": [
+ 113,
+ 127,
+ 85,
+ 200
+ ],
+ "category_id": 8,
+ "area": 54954,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5106,
+ "image_id": 953,
+ "bbox": [
+ 308,
+ 250,
+ 84,
+ 105
+ ],
+ "category_id": 8,
+ "area": 28560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5108,
+ "image_id": 954,
+ "bbox": [
+ 42,
+ 173,
+ 275,
+ 338
+ ],
+ "category_id": 10,
+ "area": 165240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5109,
+ "image_id": 954,
+ "bbox": [
+ 184,
+ 47,
+ 28,
+ 49
+ ],
+ "category_id": 10,
+ "area": 2491,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5110,
+ "image_id": 955,
+ "bbox": [
+ 0,
+ 0,
+ 264,
+ 511
+ ],
+ "category_id": 6,
+ "area": 472560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5111,
+ "image_id": 955,
+ "bbox": [
+ 159,
+ 93,
+ 352,
+ 283
+ ],
+ "category_id": 8,
+ "area": 349360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5112,
+ "image_id": 955,
+ "bbox": [
+ 286,
+ 8,
+ 225,
+ 407
+ ],
+ "category_id": 8,
+ "area": 320910,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5113,
+ "image_id": 956,
+ "bbox": [
+ 21,
+ 82,
+ 48,
+ 87
+ ],
+ "category_id": 10,
+ "area": 12576,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5114,
+ "image_id": 956,
+ "bbox": [
+ 216,
+ 54,
+ 30,
+ 42
+ ],
+ "category_id": 10,
+ "area": 3843,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5115,
+ "image_id": 956,
+ "bbox": [
+ 127,
+ 118,
+ 49,
+ 66
+ ],
+ "category_id": 10,
+ "area": 9801,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5116,
+ "image_id": 956,
+ "bbox": [
+ 348,
+ 205,
+ 60,
+ 54
+ ],
+ "category_id": 10,
+ "area": 9922,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5117,
+ "image_id": 956,
+ "bbox": [
+ 311,
+ 190,
+ 36,
+ 65
+ ],
+ "category_id": 10,
+ "area": 7154,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5118,
+ "image_id": 956,
+ "bbox": [
+ 235,
+ 8,
+ 21,
+ 33
+ ],
+ "category_id": 10,
+ "area": 2100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5119,
+ "image_id": 956,
+ "bbox": [
+ 455,
+ 455,
+ 43,
+ 56
+ ],
+ "category_id": 10,
+ "area": 7224,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5120,
+ "image_id": 956,
+ "bbox": [
+ 316,
+ 322,
+ 25,
+ 30
+ ],
+ "category_id": 10,
+ "area": 2250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5121,
+ "image_id": 956,
+ "bbox": [
+ 395,
+ 78,
+ 19,
+ 29
+ ],
+ "category_id": 10,
+ "area": 1716,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5126,
+ "image_id": 956,
+ "bbox": [
+ 302,
+ 375,
+ 22,
+ 30
+ ],
+ "category_id": 10,
+ "area": 2025,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5127,
+ "image_id": 956,
+ "bbox": [
+ 311,
+ 463,
+ 29,
+ 46
+ ],
+ "category_id": 10,
+ "area": 4130,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5130,
+ "image_id": 957,
+ "bbox": [
+ 1,
+ 85,
+ 506,
+ 426
+ ],
+ "category_id": 6,
+ "area": 182160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5137,
+ "image_id": 960,
+ "bbox": [
+ 70,
+ 269,
+ 75,
+ 70
+ ],
+ "category_id": 10,
+ "area": 12696,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5138,
+ "image_id": 960,
+ "bbox": [
+ 217,
+ 132,
+ 294,
+ 168
+ ],
+ "category_id": 9,
+ "area": 117822,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5150,
+ "image_id": 962,
+ "bbox": [
+ 65,
+ 95,
+ 380,
+ 337
+ ],
+ "category_id": 8,
+ "area": 235620,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5154,
+ "image_id": 964,
+ "bbox": [
+ 0,
+ 101,
+ 300,
+ 308
+ ],
+ "category_id": 8,
+ "area": 239418,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5162,
+ "image_id": 966,
+ "bbox": [
+ 159,
+ 201,
+ 106,
+ 299
+ ],
+ "category_id": 8,
+ "area": 253031,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5163,
+ "image_id": 966,
+ "bbox": [
+ 296,
+ 54,
+ 215,
+ 261
+ ],
+ "category_id": 6,
+ "area": 445464,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5164,
+ "image_id": 966,
+ "bbox": [
+ 0,
+ 301,
+ 57,
+ 81
+ ],
+ "category_id": 6,
+ "area": 37107,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5167,
+ "image_id": 967,
+ "bbox": [
+ 0,
+ 156,
+ 324,
+ 330
+ ],
+ "category_id": 6,
+ "area": 598500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5189,
+ "image_id": 970,
+ "bbox": [
+ 56,
+ 105,
+ 310,
+ 389
+ ],
+ "category_id": 8,
+ "area": 202007,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5190,
+ "image_id": 970,
+ "bbox": [
+ 287,
+ 40,
+ 149,
+ 269
+ ],
+ "category_id": 8,
+ "area": 67158,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5203,
+ "image_id": 973,
+ "bbox": [
+ 96,
+ 325,
+ 158,
+ 157
+ ],
+ "category_id": 10,
+ "area": 80080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5206,
+ "image_id": 974,
+ "bbox": [
+ 0,
+ 160,
+ 412,
+ 317
+ ],
+ "category_id": 9,
+ "area": 61560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5228,
+ "image_id": 980,
+ "bbox": [
+ 0,
+ 49,
+ 7,
+ 83
+ ],
+ "category_id": 6,
+ "area": 16250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5229,
+ "image_id": 980,
+ "bbox": [
+ 9,
+ 74,
+ 79,
+ 77
+ ],
+ "category_id": 6,
+ "area": 154413,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5230,
+ "image_id": 980,
+ "bbox": [
+ 1,
+ 155,
+ 35,
+ 94
+ ],
+ "category_id": 6,
+ "area": 82942,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5231,
+ "image_id": 980,
+ "bbox": [
+ 26,
+ 138,
+ 43,
+ 133
+ ],
+ "category_id": 6,
+ "area": 144040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5232,
+ "image_id": 980,
+ "bbox": [
+ 7,
+ 217,
+ 35,
+ 73
+ ],
+ "category_id": 6,
+ "area": 65780,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5233,
+ "image_id": 980,
+ "bbox": [
+ 22,
+ 262,
+ 36,
+ 75
+ ],
+ "category_id": 6,
+ "area": 68968,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5234,
+ "image_id": 980,
+ "bbox": [
+ 48,
+ 63,
+ 62,
+ 73
+ ],
+ "category_id": 6,
+ "area": 114686,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5235,
+ "image_id": 980,
+ "bbox": [
+ 65,
+ 275,
+ 39,
+ 84
+ ],
+ "category_id": 6,
+ "area": 84150,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5236,
+ "image_id": 980,
+ "bbox": [
+ 67,
+ 404,
+ 48,
+ 66
+ ],
+ "category_id": 6,
+ "area": 81067,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5237,
+ "image_id": 980,
+ "bbox": [
+ 122,
+ 58,
+ 57,
+ 60
+ ],
+ "category_id": 6,
+ "area": 87453,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5238,
+ "image_id": 980,
+ "bbox": [
+ 148,
+ 48,
+ 142,
+ 129
+ ],
+ "category_id": 6,
+ "area": 461978,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5239,
+ "image_id": 980,
+ "bbox": [
+ 162,
+ 219,
+ 42,
+ 45
+ ],
+ "category_id": 6,
+ "area": 49225,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5240,
+ "image_id": 980,
+ "bbox": [
+ 131,
+ 347,
+ 74,
+ 138
+ ],
+ "category_id": 6,
+ "area": 258660,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5241,
+ "image_id": 980,
+ "bbox": [
+ 243,
+ 41,
+ 48,
+ 54
+ ],
+ "category_id": 6,
+ "area": 65199,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5242,
+ "image_id": 980,
+ "bbox": [
+ 296,
+ 40,
+ 62,
+ 92
+ ],
+ "category_id": 6,
+ "area": 145162,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5266,
+ "image_id": 983,
+ "bbox": [
+ 0,
+ 105,
+ 511,
+ 351
+ ],
+ "category_id": 10,
+ "area": 395601,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5267,
+ "image_id": 983,
+ "bbox": [
+ 65,
+ 5,
+ 225,
+ 124
+ ],
+ "category_id": 10,
+ "area": 61588,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5268,
+ "image_id": 984,
+ "bbox": [
+ 20,
+ 307,
+ 56,
+ 109
+ ],
+ "category_id": 9,
+ "area": 21560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5269,
+ "image_id": 984,
+ "bbox": [
+ 216,
+ 165,
+ 118,
+ 165
+ ],
+ "category_id": 9,
+ "area": 69201,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5278,
+ "image_id": 986,
+ "bbox": [
+ 163,
+ 302,
+ 168,
+ 164
+ ],
+ "category_id": 8,
+ "area": 70192,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5279,
+ "image_id": 986,
+ "bbox": [
+ 284,
+ 266,
+ 35,
+ 34
+ ],
+ "category_id": 8,
+ "area": 3105,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5280,
+ "image_id": 986,
+ "bbox": [
+ 377,
+ 259,
+ 49,
+ 16
+ ],
+ "category_id": 8,
+ "area": 2016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5282,
+ "image_id": 987,
+ "bbox": [
+ 156,
+ 153,
+ 158,
+ 78
+ ],
+ "category_id": 9,
+ "area": 12528,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5287,
+ "image_id": 988,
+ "bbox": [
+ 0,
+ 211,
+ 213,
+ 181
+ ],
+ "category_id": 8,
+ "area": 546259,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5288,
+ "image_id": 988,
+ "bbox": [
+ 210,
+ 176,
+ 283,
+ 165
+ ],
+ "category_id": 8,
+ "area": 659856,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5289,
+ "image_id": 989,
+ "bbox": [
+ 184,
+ 125,
+ 145,
+ 118
+ ],
+ "category_id": 8,
+ "area": 136750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5290,
+ "image_id": 989,
+ "bbox": [
+ 159,
+ 58,
+ 110,
+ 189
+ ],
+ "category_id": 8,
+ "area": 164787,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5291,
+ "image_id": 989,
+ "bbox": [
+ 0,
+ 49,
+ 169,
+ 161
+ ],
+ "category_id": 8,
+ "area": 215900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5298,
+ "image_id": 991,
+ "bbox": [
+ 0,
+ 58,
+ 479,
+ 440
+ ],
+ "category_id": 9,
+ "area": 217494,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5299,
+ "image_id": 991,
+ "bbox": [
+ 0,
+ 415,
+ 203,
+ 96
+ ],
+ "category_id": 6,
+ "area": 20315,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5300,
+ "image_id": 991,
+ "bbox": [
+ 209,
+ 380,
+ 299,
+ 131
+ ],
+ "category_id": 6,
+ "area": 40716,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5301,
+ "image_id": 992,
+ "bbox": [
+ 115,
+ 125,
+ 331,
+ 293
+ ],
+ "category_id": 9,
+ "area": 242060,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5302,
+ "image_id": 992,
+ "bbox": [
+ 230,
+ 61,
+ 119,
+ 107
+ ],
+ "category_id": 6,
+ "area": 31831,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5303,
+ "image_id": 992,
+ "bbox": [
+ 21,
+ 134,
+ 10,
+ 56
+ ],
+ "category_id": 6,
+ "area": 1533,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5304,
+ "image_id": 992,
+ "bbox": [
+ 402,
+ 9,
+ 109,
+ 152
+ ],
+ "category_id": 6,
+ "area": 41778,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5306,
+ "image_id": 993,
+ "bbox": [
+ 18,
+ 105,
+ 68,
+ 74
+ ],
+ "category_id": 6,
+ "area": 6142,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5307,
+ "image_id": 993,
+ "bbox": [
+ 63,
+ 41,
+ 87,
+ 54
+ ],
+ "category_id": 6,
+ "area": 5640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5308,
+ "image_id": 993,
+ "bbox": [
+ 218,
+ 29,
+ 87,
+ 57
+ ],
+ "category_id": 6,
+ "area": 6016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5309,
+ "image_id": 993,
+ "bbox": [
+ 176,
+ 76,
+ 87,
+ 57
+ ],
+ "category_id": 6,
+ "area": 6016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5310,
+ "image_id": 993,
+ "bbox": [
+ 64,
+ 174,
+ 100,
+ 65
+ ],
+ "category_id": 6,
+ "area": 7776,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5311,
+ "image_id": 993,
+ "bbox": [
+ 81,
+ 112,
+ 65,
+ 63
+ ],
+ "category_id": 6,
+ "area": 4900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5312,
+ "image_id": 993,
+ "bbox": [
+ 137,
+ 129,
+ 71,
+ 60
+ ],
+ "category_id": 6,
+ "area": 5159,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5313,
+ "image_id": 993,
+ "bbox": [
+ 143,
+ 252,
+ 100,
+ 93
+ ],
+ "category_id": 6,
+ "area": 11232,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5314,
+ "image_id": 993,
+ "bbox": [
+ 0,
+ 269,
+ 142,
+ 235
+ ],
+ "category_id": 6,
+ "area": 39933,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5315,
+ "image_id": 993,
+ "bbox": [
+ 108,
+ 359,
+ 173,
+ 105
+ ],
+ "category_id": 6,
+ "area": 21762,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5316,
+ "image_id": 993,
+ "bbox": [
+ 159,
+ 438,
+ 92,
+ 67
+ ],
+ "category_id": 6,
+ "area": 7425,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5317,
+ "image_id": 993,
+ "bbox": [
+ 232,
+ 269,
+ 122,
+ 84
+ ],
+ "category_id": 6,
+ "area": 12408,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5341,
+ "image_id": 995,
+ "bbox": [
+ 0,
+ 97,
+ 247,
+ 410
+ ],
+ "category_id": 6,
+ "area": 355925,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5342,
+ "image_id": 995,
+ "bbox": [
+ 134,
+ 271,
+ 201,
+ 232
+ ],
+ "category_id": 8,
+ "area": 163475,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5343,
+ "image_id": 995,
+ "bbox": [
+ 297,
+ 89,
+ 212,
+ 363
+ ],
+ "category_id": 8,
+ "area": 270788,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5382,
+ "image_id": 1001,
+ "bbox": [
+ 176,
+ 165,
+ 334,
+ 279
+ ],
+ "category_id": 9,
+ "area": 219447,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5383,
+ "image_id": 1001,
+ "bbox": [
+ 4,
+ 0,
+ 53,
+ 63
+ ],
+ "category_id": 6,
+ "area": 7980,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5384,
+ "image_id": 1001,
+ "bbox": [
+ 0,
+ 233,
+ 58,
+ 106
+ ],
+ "category_id": 6,
+ "area": 14605,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5416,
+ "image_id": 1005,
+ "bbox": [
+ 199,
+ 96,
+ 187,
+ 286
+ ],
+ "category_id": 9,
+ "area": 414067,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5417,
+ "image_id": 1006,
+ "bbox": [
+ 206,
+ 316,
+ 150,
+ 189
+ ],
+ "category_id": 9,
+ "area": 100392,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5418,
+ "image_id": 1006,
+ "bbox": [
+ 240,
+ 216,
+ 122,
+ 112
+ ],
+ "category_id": 9,
+ "area": 48348,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5419,
+ "image_id": 1006,
+ "bbox": [
+ 0,
+ 248,
+ 289,
+ 199
+ ],
+ "category_id": 9,
+ "area": 202440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5454,
+ "image_id": 1010,
+ "bbox": [
+ 132,
+ 163,
+ 57,
+ 91
+ ],
+ "category_id": 8,
+ "area": 7008,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5455,
+ "image_id": 1010,
+ "bbox": [
+ 188,
+ 37,
+ 85,
+ 179
+ ],
+ "category_id": 8,
+ "area": 20306,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5456,
+ "image_id": 1010,
+ "bbox": [
+ 223,
+ 243,
+ 106,
+ 230
+ ],
+ "category_id": 6,
+ "area": 32568,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5457,
+ "image_id": 1010,
+ "bbox": [
+ 140,
+ 386,
+ 148,
+ 121
+ ],
+ "category_id": 6,
+ "area": 23959,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5458,
+ "image_id": 1010,
+ "bbox": [
+ 164,
+ 258,
+ 69,
+ 145
+ ],
+ "category_id": 6,
+ "area": 13340,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5478,
+ "image_id": 1012,
+ "bbox": [
+ 108,
+ 209,
+ 168,
+ 192
+ ],
+ "category_id": 8,
+ "area": 112980,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5479,
+ "image_id": 1012,
+ "bbox": [
+ 265,
+ 160,
+ 245,
+ 233
+ ],
+ "category_id": 8,
+ "area": 200778,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5508,
+ "image_id": 1015,
+ "bbox": [
+ 35,
+ 393,
+ 94,
+ 99
+ ],
+ "category_id": 6,
+ "area": 26166,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5509,
+ "image_id": 1015,
+ "bbox": [
+ 104,
+ 390,
+ 124,
+ 121
+ ],
+ "category_id": 6,
+ "area": 42065,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5512,
+ "image_id": 1016,
+ "bbox": [
+ 1,
+ 57,
+ 169,
+ 126
+ ],
+ "category_id": 6,
+ "area": 379739,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5513,
+ "image_id": 1016,
+ "bbox": [
+ 104,
+ 52,
+ 65,
+ 53
+ ],
+ "category_id": 6,
+ "area": 61904,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5514,
+ "image_id": 1016,
+ "bbox": [
+ 177,
+ 61,
+ 87,
+ 69
+ ],
+ "category_id": 6,
+ "area": 107562,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5515,
+ "image_id": 1016,
+ "bbox": [
+ 245,
+ 67,
+ 103,
+ 164
+ ],
+ "category_id": 6,
+ "area": 302434,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5516,
+ "image_id": 1016,
+ "bbox": [
+ 333,
+ 6,
+ 86,
+ 95
+ ],
+ "category_id": 6,
+ "area": 146250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5517,
+ "image_id": 1016,
+ "bbox": [
+ 366,
+ 146,
+ 67,
+ 45
+ ],
+ "category_id": 6,
+ "area": 54481,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5518,
+ "image_id": 1016,
+ "bbox": [
+ 439,
+ 127,
+ 72,
+ 84
+ ],
+ "category_id": 6,
+ "area": 109545,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5519,
+ "image_id": 1016,
+ "bbox": [
+ 127,
+ 130,
+ 105,
+ 135
+ ],
+ "category_id": 6,
+ "area": 253175,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5520,
+ "image_id": 1016,
+ "bbox": [
+ 8,
+ 228,
+ 40,
+ 30
+ ],
+ "category_id": 6,
+ "area": 22080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5521,
+ "image_id": 1016,
+ "bbox": [
+ 0,
+ 373,
+ 54,
+ 119
+ ],
+ "category_id": 6,
+ "area": 114453,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5522,
+ "image_id": 1016,
+ "bbox": [
+ 134,
+ 267,
+ 118,
+ 81
+ ],
+ "category_id": 6,
+ "area": 171836,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5523,
+ "image_id": 1016,
+ "bbox": [
+ 91,
+ 352,
+ 109,
+ 86
+ ],
+ "category_id": 6,
+ "area": 167280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5524,
+ "image_id": 1016,
+ "bbox": [
+ 246,
+ 227,
+ 85,
+ 67
+ ],
+ "category_id": 6,
+ "area": 102410,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5525,
+ "image_id": 1016,
+ "bbox": [
+ 265,
+ 289,
+ 130,
+ 78
+ ],
+ "category_id": 6,
+ "area": 180488,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5526,
+ "image_id": 1016,
+ "bbox": [
+ 468,
+ 297,
+ 41,
+ 82
+ ],
+ "category_id": 6,
+ "area": 60822,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5535,
+ "image_id": 1018,
+ "bbox": [
+ 424,
+ 104,
+ 86,
+ 145
+ ],
+ "category_id": 6,
+ "area": 21172,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5536,
+ "image_id": 1018,
+ "bbox": [
+ 144,
+ 285,
+ 366,
+ 225
+ ],
+ "category_id": 6,
+ "area": 139405,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5537,
+ "image_id": 1018,
+ "bbox": [
+ 179,
+ 174,
+ 330,
+ 303
+ ],
+ "category_id": 9,
+ "area": 169290,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5540,
+ "image_id": 1019,
+ "bbox": [
+ 0,
+ 18,
+ 139,
+ 333
+ ],
+ "category_id": 6,
+ "area": 99876,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5541,
+ "image_id": 1019,
+ "bbox": [
+ 384,
+ 1,
+ 127,
+ 157
+ ],
+ "category_id": 6,
+ "area": 43008,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5542,
+ "image_id": 1019,
+ "bbox": [
+ 40,
+ 54,
+ 451,
+ 398
+ ],
+ "category_id": 9,
+ "area": 385884,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5543,
+ "image_id": 1020,
+ "bbox": [
+ 87,
+ 183,
+ 183,
+ 176
+ ],
+ "category_id": 8,
+ "area": 113126,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5544,
+ "image_id": 1020,
+ "bbox": [
+ 216,
+ 183,
+ 217,
+ 197
+ ],
+ "category_id": 8,
+ "area": 149868,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5545,
+ "image_id": 1020,
+ "bbox": [
+ 109,
+ 109,
+ 23,
+ 22
+ ],
+ "category_id": 6,
+ "area": 1856,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5546,
+ "image_id": 1020,
+ "bbox": [
+ 141,
+ 147,
+ 26,
+ 36
+ ],
+ "category_id": 6,
+ "area": 3315,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5547,
+ "image_id": 1020,
+ "bbox": [
+ 154,
+ 189,
+ 26,
+ 39
+ ],
+ "category_id": 6,
+ "area": 3640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5549,
+ "image_id": 1020,
+ "bbox": [
+ 158,
+ 375,
+ 38,
+ 39
+ ],
+ "category_id": 6,
+ "area": 5225,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5550,
+ "image_id": 1021,
+ "bbox": [
+ 119,
+ 67,
+ 122,
+ 144
+ ],
+ "category_id": 10,
+ "area": 19110,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5551,
+ "image_id": 1021,
+ "bbox": [
+ 224,
+ 200,
+ 200,
+ 310
+ ],
+ "category_id": 9,
+ "area": 67239,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5562,
+ "image_id": 1023,
+ "bbox": [
+ 160,
+ 182,
+ 181,
+ 191
+ ],
+ "category_id": 9,
+ "area": 122126,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5563,
+ "image_id": 1023,
+ "bbox": [
+ 251,
+ 46,
+ 256,
+ 307
+ ],
+ "category_id": 9,
+ "area": 277344,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5564,
+ "image_id": 1023,
+ "bbox": [
+ 275,
+ 209,
+ 233,
+ 296
+ ],
+ "category_id": 9,
+ "area": 243528,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5584,
+ "image_id": 1027,
+ "bbox": [
+ 43,
+ 246,
+ 167,
+ 126
+ ],
+ "category_id": 9,
+ "area": 46690,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5587,
+ "image_id": 1028,
+ "bbox": [
+ 0,
+ 180,
+ 349,
+ 231
+ ],
+ "category_id": 9,
+ "area": 141155,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5592,
+ "image_id": 1030,
+ "bbox": [
+ 169,
+ 275,
+ 342,
+ 230
+ ],
+ "category_id": 8,
+ "area": 174303,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5593,
+ "image_id": 1030,
+ "bbox": [
+ 191,
+ 28,
+ 320,
+ 290
+ ],
+ "category_id": 8,
+ "area": 204828,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5594,
+ "image_id": 1031,
+ "bbox": [
+ 52,
+ 21,
+ 70,
+ 102
+ ],
+ "category_id": 9,
+ "area": 10272,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5595,
+ "image_id": 1031,
+ "bbox": [
+ 130,
+ 77,
+ 134,
+ 88
+ ],
+ "category_id": 9,
+ "area": 17015,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5596,
+ "image_id": 1031,
+ "bbox": [
+ 0,
+ 125,
+ 512,
+ 383
+ ],
+ "category_id": 9,
+ "area": 279240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5616,
+ "image_id": 1034,
+ "bbox": [
+ 0,
+ 85,
+ 276,
+ 223
+ ],
+ "category_id": 9,
+ "area": 190512,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5617,
+ "image_id": 1034,
+ "bbox": [
+ 24,
+ 398,
+ 122,
+ 108
+ ],
+ "category_id": 6,
+ "area": 41041,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5618,
+ "image_id": 1034,
+ "bbox": [
+ 131,
+ 362,
+ 139,
+ 144
+ ],
+ "category_id": 6,
+ "area": 62130,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5619,
+ "image_id": 1034,
+ "bbox": [
+ 244,
+ 434,
+ 116,
+ 77
+ ],
+ "category_id": 6,
+ "area": 27948,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5620,
+ "image_id": 1034,
+ "bbox": [
+ 284,
+ 345,
+ 148,
+ 117
+ ],
+ "category_id": 6,
+ "area": 53438,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5621,
+ "image_id": 1034,
+ "bbox": [
+ 383,
+ 244,
+ 128,
+ 263
+ ],
+ "category_id": 6,
+ "area": 104794,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5628,
+ "image_id": 1036,
+ "bbox": [
+ 88,
+ 328,
+ 141,
+ 181
+ ],
+ "category_id": 8,
+ "area": 82249,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5629,
+ "image_id": 1036,
+ "bbox": [
+ 331,
+ 267,
+ 136,
+ 103
+ ],
+ "category_id": 8,
+ "area": 45087,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5631,
+ "image_id": 1037,
+ "bbox": [
+ 371,
+ 250,
+ 72,
+ 37
+ ],
+ "category_id": 6,
+ "area": 3150,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5632,
+ "image_id": 1037,
+ "bbox": [
+ 259,
+ 346,
+ 127,
+ 75
+ ],
+ "category_id": 6,
+ "area": 11289,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5636,
+ "image_id": 1039,
+ "bbox": [
+ 33,
+ 85,
+ 478,
+ 425
+ ],
+ "category_id": 9,
+ "area": 490776,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5637,
+ "image_id": 1039,
+ "bbox": [
+ 0,
+ 36,
+ 257,
+ 472
+ ],
+ "category_id": 6,
+ "area": 293210,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5638,
+ "image_id": 1039,
+ "bbox": [
+ 0,
+ 0,
+ 510,
+ 510
+ ],
+ "category_id": 6,
+ "area": 628642,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5641,
+ "image_id": 1040,
+ "bbox": [
+ 31,
+ 251,
+ 340,
+ 259
+ ],
+ "category_id": 9,
+ "area": 219510,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5648,
+ "image_id": 1042,
+ "bbox": [
+ 0,
+ 236,
+ 238,
+ 179
+ ],
+ "category_id": 8,
+ "area": 109864,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5649,
+ "image_id": 1042,
+ "bbox": [
+ 300,
+ 249,
+ 211,
+ 115
+ ],
+ "category_id": 8,
+ "area": 62880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5650,
+ "image_id": 1043,
+ "bbox": [
+ 191,
+ 166,
+ 133,
+ 228
+ ],
+ "category_id": 9,
+ "area": 107226,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5651,
+ "image_id": 1043,
+ "bbox": [
+ 241,
+ 33,
+ 253,
+ 302
+ ],
+ "category_id": 9,
+ "area": 269450,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5652,
+ "image_id": 1043,
+ "bbox": [
+ 264,
+ 199,
+ 245,
+ 305
+ ],
+ "category_id": 9,
+ "area": 262977,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5669,
+ "image_id": 1046,
+ "bbox": [
+ 271,
+ 256,
+ 51,
+ 48
+ ],
+ "category_id": 8,
+ "area": 8772,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5670,
+ "image_id": 1046,
+ "bbox": [
+ 324,
+ 284,
+ 25,
+ 32
+ ],
+ "category_id": 8,
+ "area": 2898,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5672,
+ "image_id": 1047,
+ "bbox": [
+ 220,
+ 202,
+ 40,
+ 98
+ ],
+ "category_id": 10,
+ "area": 11340,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5673,
+ "image_id": 1047,
+ "bbox": [
+ 80,
+ 185,
+ 283,
+ 326
+ ],
+ "category_id": 9,
+ "area": 264158,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5701,
+ "image_id": 1051,
+ "bbox": [
+ 68,
+ 185,
+ 116,
+ 202
+ ],
+ "category_id": 8,
+ "area": 82070,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5702,
+ "image_id": 1051,
+ "bbox": [
+ 162,
+ 245,
+ 208,
+ 265
+ ],
+ "category_id": 8,
+ "area": 193812,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5703,
+ "image_id": 1051,
+ "bbox": [
+ 409,
+ 70,
+ 101,
+ 117
+ ],
+ "category_id": 8,
+ "area": 41656,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5712,
+ "image_id": 1053,
+ "bbox": [
+ 89,
+ 187,
+ 280,
+ 322
+ ],
+ "category_id": 9,
+ "area": 224841,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5715,
+ "image_id": 1054,
+ "bbox": [
+ 428,
+ 88,
+ 83,
+ 119
+ ],
+ "category_id": 6,
+ "area": 16770,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5716,
+ "image_id": 1054,
+ "bbox": [
+ 145,
+ 243,
+ 366,
+ 267
+ ],
+ "category_id": 6,
+ "area": 165288,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5717,
+ "image_id": 1054,
+ "bbox": [
+ 186,
+ 160,
+ 324,
+ 331
+ ],
+ "category_id": 9,
+ "area": 181440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5742,
+ "image_id": 1056,
+ "bbox": [
+ 249,
+ 41,
+ 106,
+ 122
+ ],
+ "category_id": 9,
+ "area": 45924,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5743,
+ "image_id": 1056,
+ "bbox": [
+ 211,
+ 183,
+ 140,
+ 211
+ ],
+ "category_id": 9,
+ "area": 104247,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5744,
+ "image_id": 1056,
+ "bbox": [
+ 301,
+ 469,
+ 14,
+ 36
+ ],
+ "category_id": 9,
+ "area": 1924,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5745,
+ "image_id": 1057,
+ "bbox": [
+ 69,
+ 92,
+ 442,
+ 314
+ ],
+ "category_id": 8,
+ "area": 1074744,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5749,
+ "image_id": 1059,
+ "bbox": [
+ 57,
+ 244,
+ 109,
+ 70
+ ],
+ "category_id": 9,
+ "area": 11700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5750,
+ "image_id": 1060,
+ "bbox": [
+ 198,
+ 319,
+ 164,
+ 98
+ ],
+ "category_id": 9,
+ "area": 26956,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5752,
+ "image_id": 1061,
+ "bbox": [
+ 0,
+ 134,
+ 277,
+ 162
+ ],
+ "category_id": 9,
+ "area": 75856,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5753,
+ "image_id": 1061,
+ "bbox": [
+ 0,
+ 211,
+ 184,
+ 299
+ ],
+ "category_id": 6,
+ "area": 92950,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5754,
+ "image_id": 1062,
+ "bbox": [
+ 22,
+ 173,
+ 335,
+ 338
+ ],
+ "category_id": 9,
+ "area": 280136,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5755,
+ "image_id": 1062,
+ "bbox": [
+ 342,
+ 250,
+ 72,
+ 157
+ ],
+ "category_id": 6,
+ "area": 28080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5756,
+ "image_id": 1063,
+ "bbox": [
+ 3,
+ 92,
+ 446,
+ 417
+ ],
+ "category_id": 9,
+ "area": 156784,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5759,
+ "image_id": 1064,
+ "bbox": [
+ 178,
+ 33,
+ 145,
+ 136
+ ],
+ "category_id": 9,
+ "area": 26750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5760,
+ "image_id": 1064,
+ "bbox": [
+ 199,
+ 101,
+ 178,
+ 228
+ ],
+ "category_id": 9,
+ "area": 55020,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5787,
+ "image_id": 1069,
+ "bbox": [
+ 4,
+ 165,
+ 102,
+ 80
+ ],
+ "category_id": 8,
+ "area": 65110,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5788,
+ "image_id": 1069,
+ "bbox": [
+ 52,
+ 113,
+ 60,
+ 48
+ ],
+ "category_id": 8,
+ "area": 23052,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5789,
+ "image_id": 1069,
+ "bbox": [
+ 164,
+ 86,
+ 39,
+ 42
+ ],
+ "category_id": 8,
+ "area": 13410,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5790,
+ "image_id": 1069,
+ "bbox": [
+ 84,
+ 179,
+ 313,
+ 265
+ ],
+ "category_id": 8,
+ "area": 659120,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5791,
+ "image_id": 1069,
+ "bbox": [
+ 0,
+ 386,
+ 146,
+ 120
+ ],
+ "category_id": 8,
+ "area": 139446,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5792,
+ "image_id": 1070,
+ "bbox": [
+ 352,
+ 6,
+ 78,
+ 76
+ ],
+ "category_id": 10,
+ "area": 21060,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5793,
+ "image_id": 1070,
+ "bbox": [
+ 285,
+ 363,
+ 47,
+ 67
+ ],
+ "category_id": 10,
+ "area": 11305,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5794,
+ "image_id": 1070,
+ "bbox": [
+ 244,
+ 314,
+ 27,
+ 49
+ ],
+ "category_id": 10,
+ "area": 4760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5795,
+ "image_id": 1070,
+ "bbox": [
+ 298,
+ 339,
+ 23,
+ 32
+ ],
+ "category_id": 10,
+ "area": 2714,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5796,
+ "image_id": 1070,
+ "bbox": [
+ 434,
+ 290,
+ 36,
+ 64
+ ],
+ "category_id": 10,
+ "area": 8280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5797,
+ "image_id": 1070,
+ "bbox": [
+ 472,
+ 307,
+ 22,
+ 48
+ ],
+ "category_id": 10,
+ "area": 3808,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5798,
+ "image_id": 1070,
+ "bbox": [
+ 135,
+ 95,
+ 20,
+ 41
+ ],
+ "category_id": 10,
+ "area": 2958,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5799,
+ "image_id": 1070,
+ "bbox": [
+ 122,
+ 249,
+ 27,
+ 51
+ ],
+ "category_id": 10,
+ "area": 4964,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5800,
+ "image_id": 1070,
+ "bbox": [
+ 51,
+ 450,
+ 16,
+ 39
+ ],
+ "category_id": 10,
+ "area": 2200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5801,
+ "image_id": 1070,
+ "bbox": [
+ 6,
+ 427,
+ 19,
+ 44
+ ],
+ "category_id": 10,
+ "area": 3087,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5802,
+ "image_id": 1070,
+ "bbox": [
+ 417,
+ 327,
+ 16,
+ 38
+ ],
+ "category_id": 10,
+ "area": 2214,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5803,
+ "image_id": 1071,
+ "bbox": [
+ 22,
+ 54,
+ 435,
+ 356
+ ],
+ "category_id": 9,
+ "area": 1228768,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5808,
+ "image_id": 1073,
+ "bbox": [
+ 138,
+ 0,
+ 322,
+ 345
+ ],
+ "category_id": 10,
+ "area": 880632,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5809,
+ "image_id": 1073,
+ "bbox": [
+ 0,
+ 10,
+ 325,
+ 499
+ ],
+ "category_id": 10,
+ "area": 1284660,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5813,
+ "image_id": 1075,
+ "bbox": [
+ 130,
+ 125,
+ 276,
+ 284
+ ],
+ "category_id": 8,
+ "area": 235578,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5814,
+ "image_id": 1075,
+ "bbox": [
+ 1,
+ 377,
+ 258,
+ 132
+ ],
+ "category_id": 6,
+ "area": 102684,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5815,
+ "image_id": 1075,
+ "bbox": [
+ 1,
+ 45,
+ 104,
+ 144
+ ],
+ "category_id": 6,
+ "area": 45144,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5816,
+ "image_id": 1075,
+ "bbox": [
+ 7,
+ 143,
+ 166,
+ 161
+ ],
+ "category_id": 6,
+ "area": 80344,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5817,
+ "image_id": 1075,
+ "bbox": [
+ 174,
+ 34,
+ 141,
+ 98
+ ],
+ "category_id": 6,
+ "area": 41736,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5818,
+ "image_id": 1075,
+ "bbox": [
+ 266,
+ 320,
+ 244,
+ 188
+ ],
+ "category_id": 6,
+ "area": 138387,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5819,
+ "image_id": 1075,
+ "bbox": [
+ 392,
+ 130,
+ 75,
+ 102
+ ],
+ "category_id": 6,
+ "area": 22950,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5820,
+ "image_id": 1075,
+ "bbox": [
+ 405,
+ 229,
+ 52,
+ 32
+ ],
+ "category_id": 6,
+ "area": 5145,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5824,
+ "image_id": 1077,
+ "bbox": [
+ 148,
+ 89,
+ 184,
+ 288
+ ],
+ "category_id": 8,
+ "area": 46134,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5825,
+ "image_id": 1077,
+ "bbox": [
+ 35,
+ 27,
+ 108,
+ 65
+ ],
+ "category_id": 6,
+ "area": 6201,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5826,
+ "image_id": 1077,
+ "bbox": [
+ 159,
+ 13,
+ 149,
+ 83
+ ],
+ "category_id": 6,
+ "area": 10787,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5827,
+ "image_id": 1077,
+ "bbox": [
+ 326,
+ 66,
+ 85,
+ 90
+ ],
+ "category_id": 6,
+ "area": 6716,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5828,
+ "image_id": 1078,
+ "bbox": [
+ 30,
+ 24,
+ 310,
+ 485
+ ],
+ "category_id": 10,
+ "area": 1192075,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5840,
+ "image_id": 1080,
+ "bbox": [
+ 101,
+ 139,
+ 305,
+ 316
+ ],
+ "category_id": 8,
+ "area": 134838,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5841,
+ "image_id": 1080,
+ "bbox": [
+ 0,
+ 193,
+ 45,
+ 163
+ ],
+ "category_id": 6,
+ "area": 10251,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5842,
+ "image_id": 1081,
+ "bbox": [
+ 0,
+ 25,
+ 416,
+ 483
+ ],
+ "category_id": 9,
+ "area": 204732,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5850,
+ "image_id": 1083,
+ "bbox": [
+ 251,
+ 0,
+ 257,
+ 479
+ ],
+ "category_id": 9,
+ "area": 434056,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5851,
+ "image_id": 1083,
+ "bbox": [
+ 213,
+ 225,
+ 296,
+ 280
+ ],
+ "category_id": 9,
+ "area": 292695,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5861,
+ "image_id": 1086,
+ "bbox": [
+ 129,
+ 106,
+ 107,
+ 190
+ ],
+ "category_id": 8,
+ "area": 138345,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5862,
+ "image_id": 1087,
+ "bbox": [
+ 65,
+ 100,
+ 347,
+ 411
+ ],
+ "category_id": 9,
+ "area": 275793,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5864,
+ "image_id": 1087,
+ "bbox": [
+ 103,
+ 3,
+ 99,
+ 120
+ ],
+ "category_id": 6,
+ "area": 22983,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5865,
+ "image_id": 1088,
+ "bbox": [
+ 124,
+ 184,
+ 241,
+ 283
+ ],
+ "category_id": 8,
+ "area": 117528,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5866,
+ "image_id": 1088,
+ "bbox": [
+ 51,
+ 205,
+ 87,
+ 192
+ ],
+ "category_id": 6,
+ "area": 28730,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5867,
+ "image_id": 1088,
+ "bbox": [
+ 368,
+ 224,
+ 140,
+ 276
+ ],
+ "category_id": 6,
+ "area": 66825,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5868,
+ "image_id": 1088,
+ "bbox": [
+ 2,
+ 316,
+ 67,
+ 146
+ ],
+ "category_id": 6,
+ "area": 16899,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5883,
+ "image_id": 1091,
+ "bbox": [
+ 110,
+ 90,
+ 361,
+ 282
+ ],
+ "category_id": 9,
+ "area": 358888,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5885,
+ "image_id": 1092,
+ "bbox": [
+ 0,
+ 400,
+ 256,
+ 108
+ ],
+ "category_id": 6,
+ "area": 30690,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5911,
+ "image_id": 1096,
+ "bbox": [
+ 391,
+ 109,
+ 70,
+ 78
+ ],
+ "category_id": 6,
+ "area": 36704,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5912,
+ "image_id": 1097,
+ "bbox": [
+ 0,
+ 86,
+ 56,
+ 109
+ ],
+ "category_id": 6,
+ "area": 67671,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5913,
+ "image_id": 1097,
+ "bbox": [
+ 76,
+ 335,
+ 116,
+ 174
+ ],
+ "category_id": 6,
+ "area": 222794,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5914,
+ "image_id": 1097,
+ "bbox": [
+ 0,
+ 400,
+ 51,
+ 101
+ ],
+ "category_id": 6,
+ "area": 57600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5915,
+ "image_id": 1097,
+ "bbox": [
+ 113,
+ 0,
+ 219,
+ 78
+ ],
+ "category_id": 6,
+ "area": 188955,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5916,
+ "image_id": 1097,
+ "bbox": [
+ 18,
+ 26,
+ 74,
+ 132
+ ],
+ "category_id": 6,
+ "area": 109040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5917,
+ "image_id": 1097,
+ "bbox": [
+ 285,
+ 9,
+ 120,
+ 193
+ ],
+ "category_id": 6,
+ "area": 256543,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5918,
+ "image_id": 1097,
+ "bbox": [
+ 420,
+ 1,
+ 90,
+ 153
+ ],
+ "category_id": 6,
+ "area": 153202,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5919,
+ "image_id": 1097,
+ "bbox": [
+ 421,
+ 144,
+ 87,
+ 101
+ ],
+ "category_id": 6,
+ "area": 97580,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5924,
+ "image_id": 1097,
+ "bbox": [
+ 152,
+ 144,
+ 39,
+ 39
+ ],
+ "category_id": 6,
+ "area": 17094,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5931,
+ "image_id": 1099,
+ "bbox": [
+ 271,
+ 75,
+ 94,
+ 101
+ ],
+ "category_id": 8,
+ "area": 660480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5932,
+ "image_id": 1099,
+ "bbox": [
+ 215,
+ 296,
+ 239,
+ 99
+ ],
+ "category_id": 8,
+ "area": 1627430,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5933,
+ "image_id": 1099,
+ "bbox": [
+ 129,
+ 175,
+ 190,
+ 256
+ ],
+ "category_id": 8,
+ "area": 3333710,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5934,
+ "image_id": 1099,
+ "bbox": [
+ 100,
+ 249,
+ 140,
+ 209
+ ],
+ "category_id": 8,
+ "area": 2014950,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5935,
+ "image_id": 1099,
+ "bbox": [
+ 0,
+ 44,
+ 417,
+ 463
+ ],
+ "category_id": 6,
+ "area": 13217829,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5939,
+ "image_id": 1101,
+ "bbox": [
+ 58,
+ 291,
+ 227,
+ 219
+ ],
+ "category_id": 10,
+ "area": 78561,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5941,
+ "image_id": 1102,
+ "bbox": [
+ 96,
+ 164,
+ 271,
+ 199
+ ],
+ "category_id": 8,
+ "area": 131928,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5942,
+ "image_id": 1102,
+ "bbox": [
+ 297,
+ 101,
+ 213,
+ 331
+ ],
+ "category_id": 8,
+ "area": 172125,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5943,
+ "image_id": 1102,
+ "bbox": [
+ 0,
+ 12,
+ 151,
+ 498
+ ],
+ "category_id": 6,
+ "area": 183540,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5944,
+ "image_id": 1103,
+ "bbox": [
+ 10,
+ 159,
+ 307,
+ 266
+ ],
+ "category_id": 8,
+ "area": 225610,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5945,
+ "image_id": 1103,
+ "bbox": [
+ 377,
+ 32,
+ 132,
+ 232
+ ],
+ "category_id": 8,
+ "area": 85008,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5946,
+ "image_id": 1103,
+ "bbox": [
+ 0,
+ 38,
+ 226,
+ 141
+ ],
+ "category_id": 6,
+ "area": 87924,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5947,
+ "image_id": 1103,
+ "bbox": [
+ 283,
+ 41,
+ 101,
+ 128
+ ],
+ "category_id": 6,
+ "area": 35705,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5948,
+ "image_id": 1103,
+ "bbox": [
+ 171,
+ 337,
+ 153,
+ 146
+ ],
+ "category_id": 6,
+ "area": 61823,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5960,
+ "image_id": 1106,
+ "bbox": [
+ 65,
+ 149,
+ 416,
+ 360
+ ],
+ "category_id": 8,
+ "area": 547149,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5967,
+ "image_id": 1109,
+ "bbox": [
+ 17,
+ 63,
+ 348,
+ 290
+ ],
+ "category_id": 9,
+ "area": 69360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5970,
+ "image_id": 1109,
+ "bbox": [
+ 336,
+ 54,
+ 137,
+ 180
+ ],
+ "category_id": 9,
+ "area": 17066,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5974,
+ "image_id": 1109,
+ "bbox": [
+ 143,
+ 394,
+ 166,
+ 117
+ ],
+ "category_id": 9,
+ "area": 13455,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5980,
+ "image_id": 1111,
+ "bbox": [
+ 226,
+ 153,
+ 154,
+ 246
+ ],
+ "category_id": 8,
+ "area": 132135,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5986,
+ "image_id": 1112,
+ "bbox": [
+ 159,
+ 88,
+ 153,
+ 293
+ ],
+ "category_id": 8,
+ "area": 304750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5987,
+ "image_id": 1112,
+ "bbox": [
+ 349,
+ 237,
+ 99,
+ 118
+ ],
+ "category_id": 8,
+ "area": 79822,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5988,
+ "image_id": 1113,
+ "bbox": [
+ 113,
+ 38,
+ 109,
+ 182
+ ],
+ "category_id": 6,
+ "area": 23427,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5989,
+ "image_id": 1113,
+ "bbox": [
+ 57,
+ 104,
+ 49,
+ 69
+ ],
+ "category_id": 6,
+ "area": 4030,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 5990,
+ "image_id": 1113,
+ "bbox": [
+ 316,
+ 283,
+ 38,
+ 68
+ ],
+ "category_id": 6,
+ "area": 3072,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6002,
+ "image_id": 1115,
+ "bbox": [
+ 146,
+ 89,
+ 365,
+ 421
+ ],
+ "category_id": 8,
+ "area": 536310,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6003,
+ "image_id": 1115,
+ "bbox": [
+ 101,
+ 283,
+ 207,
+ 149
+ ],
+ "category_id": 8,
+ "area": 107844,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6007,
+ "image_id": 1118,
+ "bbox": [
+ 245,
+ 254,
+ 192,
+ 233
+ ],
+ "category_id": 9,
+ "area": 91520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6008,
+ "image_id": 1118,
+ "bbox": [
+ 68,
+ 0,
+ 136,
+ 161
+ ],
+ "category_id": 6,
+ "area": 45000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6009,
+ "image_id": 1118,
+ "bbox": [
+ 0,
+ 0,
+ 90,
+ 182
+ ],
+ "category_id": 6,
+ "area": 33698,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6010,
+ "image_id": 1118,
+ "bbox": [
+ 8,
+ 123,
+ 238,
+ 388
+ ],
+ "category_id": 6,
+ "area": 188788,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6032,
+ "image_id": 1121,
+ "bbox": [
+ 168,
+ 232,
+ 89,
+ 82
+ ],
+ "category_id": 8,
+ "area": 58116,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6033,
+ "image_id": 1121,
+ "bbox": [
+ 56,
+ 265,
+ 172,
+ 110
+ ],
+ "category_id": 8,
+ "area": 150930,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6034,
+ "image_id": 1121,
+ "bbox": [
+ 379,
+ 132,
+ 132,
+ 291
+ ],
+ "category_id": 8,
+ "area": 305655,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6035,
+ "image_id": 1122,
+ "bbox": [
+ 272,
+ 330,
+ 141,
+ 68
+ ],
+ "category_id": 10,
+ "area": 15476,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6036,
+ "image_id": 1122,
+ "bbox": [
+ 0,
+ 25,
+ 478,
+ 355
+ ],
+ "category_id": 9,
+ "area": 272745,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6039,
+ "image_id": 1123,
+ "bbox": [
+ 28,
+ 405,
+ 96,
+ 98
+ ],
+ "category_id": 6,
+ "area": 26535,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6040,
+ "image_id": 1123,
+ "bbox": [
+ 101,
+ 403,
+ 127,
+ 108
+ ],
+ "category_id": 6,
+ "area": 38319,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6044,
+ "image_id": 1124,
+ "bbox": [
+ 152,
+ 36,
+ 206,
+ 423
+ ],
+ "category_id": 10,
+ "area": 691956,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6047,
+ "image_id": 1125,
+ "bbox": [
+ 0,
+ 294,
+ 512,
+ 217
+ ],
+ "category_id": 6,
+ "area": 113030,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6048,
+ "image_id": 1125,
+ "bbox": [
+ 0,
+ 145,
+ 114,
+ 217
+ ],
+ "category_id": 6,
+ "area": 25276,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6049,
+ "image_id": 1125,
+ "bbox": [
+ 0,
+ 284,
+ 145,
+ 140
+ ],
+ "category_id": 6,
+ "area": 20815,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6050,
+ "image_id": 1126,
+ "bbox": [
+ 224,
+ 111,
+ 102,
+ 126
+ ],
+ "category_id": 8,
+ "area": 103062,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6051,
+ "image_id": 1126,
+ "bbox": [
+ 319,
+ 107,
+ 72,
+ 78
+ ],
+ "category_id": 8,
+ "area": 44820,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6053,
+ "image_id": 1127,
+ "bbox": [
+ 284,
+ 170,
+ 88,
+ 82
+ ],
+ "category_id": 10,
+ "area": 13803,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6054,
+ "image_id": 1127,
+ "bbox": [
+ 235,
+ 106,
+ 171,
+ 62
+ ],
+ "category_id": 9,
+ "area": 20250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6060,
+ "image_id": 1128,
+ "bbox": [
+ 163,
+ 169,
+ 331,
+ 241
+ ],
+ "category_id": 8,
+ "area": 82256,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6074,
+ "image_id": 1131,
+ "bbox": [
+ 124,
+ 102,
+ 377,
+ 404
+ ],
+ "category_id": 6,
+ "area": 404233,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6075,
+ "image_id": 1131,
+ "bbox": [
+ 2,
+ 277,
+ 137,
+ 219
+ ],
+ "category_id": 6,
+ "area": 79893,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6105,
+ "image_id": 1132,
+ "bbox": [
+ 105,
+ 126,
+ 305,
+ 377
+ ],
+ "category_id": 8,
+ "area": 914159,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6106,
+ "image_id": 1132,
+ "bbox": [
+ 354,
+ 144,
+ 157,
+ 164
+ ],
+ "category_id": 8,
+ "area": 204730,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6107,
+ "image_id": 1133,
+ "bbox": [
+ 250,
+ 40,
+ 79,
+ 180
+ ],
+ "category_id": 8,
+ "area": 113620,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6108,
+ "image_id": 1133,
+ "bbox": [
+ 281,
+ 240,
+ 100,
+ 270
+ ],
+ "category_id": 8,
+ "area": 213750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6109,
+ "image_id": 1133,
+ "bbox": [
+ 249,
+ 22,
+ 37,
+ 83
+ ],
+ "category_id": 8,
+ "area": 24957,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6110,
+ "image_id": 1133,
+ "bbox": [
+ 141,
+ 107,
+ 20,
+ 66
+ ],
+ "category_id": 6,
+ "area": 10716,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6111,
+ "image_id": 1133,
+ "bbox": [
+ 176,
+ 63,
+ 22,
+ 74
+ ],
+ "category_id": 6,
+ "area": 13114,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6112,
+ "image_id": 1134,
+ "bbox": [
+ 174,
+ 292,
+ 115,
+ 81
+ ],
+ "category_id": 6,
+ "area": 5989,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6113,
+ "image_id": 1134,
+ "bbox": [
+ 0,
+ 425,
+ 129,
+ 86
+ ],
+ "category_id": 6,
+ "area": 7056,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6115,
+ "image_id": 1134,
+ "bbox": [
+ 311,
+ 322,
+ 200,
+ 189
+ ],
+ "category_id": 6,
+ "area": 24108,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6116,
+ "image_id": 1135,
+ "bbox": [
+ 0,
+ 301,
+ 128,
+ 210
+ ],
+ "category_id": 6,
+ "area": 27935,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6117,
+ "image_id": 1135,
+ "bbox": [
+ 0,
+ 0,
+ 512,
+ 512
+ ],
+ "category_id": 6,
+ "area": 270000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6121,
+ "image_id": 1136,
+ "bbox": [
+ 0,
+ 0,
+ 512,
+ 264
+ ],
+ "category_id": 6,
+ "area": 330714,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6122,
+ "image_id": 1137,
+ "bbox": [
+ 1,
+ 98,
+ 510,
+ 411
+ ],
+ "category_id": 6,
+ "area": 200598,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6123,
+ "image_id": 1138,
+ "bbox": [
+ 58,
+ 313,
+ 208,
+ 144
+ ],
+ "category_id": 9,
+ "area": 50868,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6124,
+ "image_id": 1138,
+ "bbox": [
+ 14,
+ 323,
+ 275,
+ 188
+ ],
+ "category_id": 6,
+ "area": 87740,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6130,
+ "image_id": 1140,
+ "bbox": [
+ 138,
+ 238,
+ 130,
+ 273
+ ],
+ "category_id": 9,
+ "area": 273694,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6132,
+ "image_id": 1141,
+ "bbox": [
+ 116,
+ 67,
+ 389,
+ 444
+ ],
+ "category_id": 10,
+ "area": 270838,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6133,
+ "image_id": 1141,
+ "bbox": [
+ 278,
+ 0,
+ 148,
+ 161
+ ],
+ "category_id": 10,
+ "area": 37530,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6134,
+ "image_id": 1141,
+ "bbox": [
+ 109,
+ 0,
+ 108,
+ 91
+ ],
+ "category_id": 10,
+ "area": 15642,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6145,
+ "image_id": 1143,
+ "bbox": [
+ 186,
+ 77,
+ 120,
+ 296
+ ],
+ "category_id": 10,
+ "area": 46964,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6150,
+ "image_id": 1143,
+ "bbox": [
+ 311,
+ 460,
+ 86,
+ 50
+ ],
+ "category_id": 6,
+ "area": 5720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6170,
+ "image_id": 1145,
+ "bbox": [
+ 299,
+ 179,
+ 113,
+ 160
+ ],
+ "category_id": 10,
+ "area": 192100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6171,
+ "image_id": 1145,
+ "bbox": [
+ 325,
+ 454,
+ 69,
+ 57
+ ],
+ "category_id": 10,
+ "area": 41958,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6173,
+ "image_id": 1146,
+ "bbox": [
+ 0,
+ 203,
+ 61,
+ 306
+ ],
+ "category_id": 6,
+ "area": 12551,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6177,
+ "image_id": 1147,
+ "bbox": [
+ 0,
+ 11,
+ 187,
+ 131
+ ],
+ "category_id": 6,
+ "area": 434010,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6178,
+ "image_id": 1147,
+ "bbox": [
+ 0,
+ 333,
+ 69,
+ 124
+ ],
+ "category_id": 6,
+ "area": 152312,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6179,
+ "image_id": 1147,
+ "bbox": [
+ 123,
+ 9,
+ 80,
+ 57
+ ],
+ "category_id": 6,
+ "area": 81474,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6180,
+ "image_id": 1147,
+ "bbox": [
+ 192,
+ 15,
+ 93,
+ 94
+ ],
+ "category_id": 6,
+ "area": 155490,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6181,
+ "image_id": 1147,
+ "bbox": [
+ 148,
+ 87,
+ 100,
+ 140
+ ],
+ "category_id": 6,
+ "area": 248151,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6182,
+ "image_id": 1147,
+ "bbox": [
+ 267,
+ 23,
+ 98,
+ 167
+ ],
+ "category_id": 6,
+ "area": 289209,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6183,
+ "image_id": 1147,
+ "bbox": [
+ 262,
+ 184,
+ 84,
+ 71
+ ],
+ "category_id": 6,
+ "area": 105984,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6184,
+ "image_id": 1147,
+ "bbox": [
+ 367,
+ 44,
+ 28,
+ 80
+ ],
+ "category_id": 6,
+ "area": 39680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6185,
+ "image_id": 1147,
+ "bbox": [
+ 483,
+ 261,
+ 28,
+ 82
+ ],
+ "category_id": 6,
+ "area": 41470,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6186,
+ "image_id": 1147,
+ "bbox": [
+ 382,
+ 88,
+ 63,
+ 71
+ ],
+ "category_id": 6,
+ "area": 80064,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6187,
+ "image_id": 1147,
+ "bbox": [
+ 451,
+ 89,
+ 60,
+ 77
+ ],
+ "category_id": 6,
+ "area": 83076,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6193,
+ "image_id": 1149,
+ "bbox": [
+ 10,
+ 82,
+ 160,
+ 277
+ ],
+ "category_id": 9,
+ "area": 104229,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6194,
+ "image_id": 1149,
+ "bbox": [
+ 368,
+ 411,
+ 132,
+ 100
+ ],
+ "category_id": 9,
+ "area": 31080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6195,
+ "image_id": 1149,
+ "bbox": [
+ 140,
+ 355,
+ 157,
+ 156
+ ],
+ "category_id": 9,
+ "area": 57904,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6196,
+ "image_id": 1149,
+ "bbox": [
+ 172,
+ 209,
+ 112,
+ 209
+ ],
+ "category_id": 9,
+ "area": 55220,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6197,
+ "image_id": 1149,
+ "bbox": [
+ 251,
+ 60,
+ 120,
+ 146
+ ],
+ "category_id": 9,
+ "area": 41536,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6198,
+ "image_id": 1149,
+ "bbox": [
+ 244,
+ 94,
+ 113,
+ 181
+ ],
+ "category_id": 9,
+ "area": 48396,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6199,
+ "image_id": 1150,
+ "bbox": [
+ 191,
+ 32,
+ 85,
+ 476
+ ],
+ "category_id": 10,
+ "area": 323610,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6200,
+ "image_id": 1151,
+ "bbox": [
+ 75,
+ 323,
+ 434,
+ 185
+ ],
+ "category_id": 6,
+ "area": 362202,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6204,
+ "image_id": 1152,
+ "bbox": [
+ 10,
+ 144,
+ 225,
+ 288
+ ],
+ "category_id": 8,
+ "area": 94688,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6205,
+ "image_id": 1152,
+ "bbox": [
+ 253,
+ 81,
+ 250,
+ 186
+ ],
+ "category_id": 8,
+ "area": 68208,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6207,
+ "image_id": 1153,
+ "bbox": [
+ 171,
+ 233,
+ 320,
+ 215
+ ],
+ "category_id": 6,
+ "area": 31560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6214,
+ "image_id": 1155,
+ "bbox": [
+ 48,
+ 0,
+ 166,
+ 179
+ ],
+ "category_id": 6,
+ "area": 199104,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6215,
+ "image_id": 1155,
+ "bbox": [
+ 186,
+ 0,
+ 152,
+ 254
+ ],
+ "category_id": 6,
+ "area": 257943,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6216,
+ "image_id": 1155,
+ "bbox": [
+ 333,
+ 5,
+ 117,
+ 259
+ ],
+ "category_id": 6,
+ "area": 203520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6217,
+ "image_id": 1156,
+ "bbox": [
+ 311,
+ 293,
+ 19,
+ 33
+ ],
+ "category_id": 6,
+ "area": 744,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6218,
+ "image_id": 1156,
+ "bbox": [
+ 332,
+ 294,
+ 24,
+ 43
+ ],
+ "category_id": 6,
+ "area": 1230,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6219,
+ "image_id": 1156,
+ "bbox": [
+ 268,
+ 331,
+ 24,
+ 30
+ ],
+ "category_id": 6,
+ "area": 899,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6220,
+ "image_id": 1157,
+ "bbox": [
+ 146,
+ 118,
+ 257,
+ 252
+ ],
+ "category_id": 10,
+ "area": 150516,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6225,
+ "image_id": 1158,
+ "bbox": [
+ 398,
+ 406,
+ 110,
+ 96
+ ],
+ "category_id": 6,
+ "area": 34600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6226,
+ "image_id": 1158,
+ "bbox": [
+ 47,
+ 362,
+ 164,
+ 77
+ ],
+ "category_id": 6,
+ "area": 41377,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6227,
+ "image_id": 1158,
+ "bbox": [
+ 200,
+ 421,
+ 91,
+ 77
+ ],
+ "category_id": 6,
+ "area": 23023,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6228,
+ "image_id": 1158,
+ "bbox": [
+ 0,
+ 441,
+ 167,
+ 70
+ ],
+ "category_id": 6,
+ "area": 37990,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6233,
+ "image_id": 1161,
+ "bbox": [
+ 0,
+ 20,
+ 10,
+ 71
+ ],
+ "category_id": 6,
+ "area": 7632,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6234,
+ "image_id": 1161,
+ "bbox": [
+ 37,
+ 0,
+ 23,
+ 100
+ ],
+ "category_id": 6,
+ "area": 23621,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6235,
+ "image_id": 1161,
+ "bbox": [
+ 0,
+ 87,
+ 49,
+ 59
+ ],
+ "category_id": 6,
+ "area": 29040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6236,
+ "image_id": 1161,
+ "bbox": [
+ 5,
+ 129,
+ 103,
+ 110
+ ],
+ "category_id": 6,
+ "area": 114821,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6237,
+ "image_id": 1161,
+ "bbox": [
+ 3,
+ 234,
+ 51,
+ 98
+ ],
+ "category_id": 6,
+ "area": 50689,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6238,
+ "image_id": 1161,
+ "bbox": [
+ 49,
+ 21,
+ 169,
+ 107
+ ],
+ "category_id": 6,
+ "area": 182720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6239,
+ "image_id": 1161,
+ "bbox": [
+ 95,
+ 102,
+ 78,
+ 174
+ ],
+ "category_id": 6,
+ "area": 138065,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6240,
+ "image_id": 1161,
+ "bbox": [
+ 59,
+ 204,
+ 64,
+ 98
+ ],
+ "category_id": 6,
+ "area": 63288,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6241,
+ "image_id": 1161,
+ "bbox": [
+ 86,
+ 264,
+ 70,
+ 102
+ ],
+ "category_id": 6,
+ "area": 72522,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6242,
+ "image_id": 1161,
+ "bbox": [
+ 132,
+ 8,
+ 124,
+ 80
+ ],
+ "category_id": 6,
+ "area": 99902,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6243,
+ "image_id": 1161,
+ "bbox": [
+ 167,
+ 142,
+ 103,
+ 119
+ ],
+ "category_id": 6,
+ "area": 123895,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6244,
+ "image_id": 1161,
+ "bbox": [
+ 231,
+ 95,
+ 86,
+ 121
+ ],
+ "category_id": 6,
+ "area": 105342,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6245,
+ "image_id": 1161,
+ "bbox": [
+ 220,
+ 80,
+ 93,
+ 82
+ ],
+ "category_id": 6,
+ "area": 76930,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6246,
+ "image_id": 1161,
+ "bbox": [
+ 275,
+ 1,
+ 104,
+ 61
+ ],
+ "category_id": 6,
+ "area": 63882,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6247,
+ "image_id": 1161,
+ "bbox": [
+ 324,
+ 6,
+ 187,
+ 149
+ ],
+ "category_id": 6,
+ "area": 281610,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6253,
+ "image_id": 1162,
+ "bbox": [
+ 133,
+ 0,
+ 243,
+ 506
+ ],
+ "category_id": 9,
+ "area": 433608,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6254,
+ "image_id": 1162,
+ "bbox": [
+ 288,
+ 0,
+ 125,
+ 144
+ ],
+ "category_id": 9,
+ "area": 63742,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6256,
+ "image_id": 1163,
+ "bbox": [
+ 81,
+ 43,
+ 324,
+ 399
+ ],
+ "category_id": 9,
+ "area": 334122,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6258,
+ "image_id": 1164,
+ "bbox": [
+ 367,
+ 156,
+ 129,
+ 115
+ ],
+ "category_id": 9,
+ "area": 22578,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6260,
+ "image_id": 1165,
+ "bbox": [
+ 103,
+ 50,
+ 391,
+ 459
+ ],
+ "category_id": 8,
+ "area": 5685095,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6292,
+ "image_id": 1170,
+ "bbox": [
+ 331,
+ 239,
+ 34,
+ 113
+ ],
+ "category_id": 8,
+ "area": 13674,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6349,
+ "image_id": 1173,
+ "bbox": [
+ 173,
+ 424,
+ 333,
+ 87
+ ],
+ "category_id": 9,
+ "area": 45639,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6354,
+ "image_id": 1175,
+ "bbox": [
+ 125,
+ 284,
+ 100,
+ 111
+ ],
+ "category_id": 6,
+ "area": 15368,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6355,
+ "image_id": 1175,
+ "bbox": [
+ 182,
+ 131,
+ 83,
+ 111
+ ],
+ "category_id": 6,
+ "area": 12656,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6356,
+ "image_id": 1175,
+ "bbox": [
+ 302,
+ 0,
+ 87,
+ 159
+ ],
+ "category_id": 6,
+ "area": 18998,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6357,
+ "image_id": 1175,
+ "bbox": [
+ 30,
+ 2,
+ 53,
+ 58
+ ],
+ "category_id": 6,
+ "area": 4248,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6358,
+ "image_id": 1175,
+ "bbox": [
+ 33,
+ 108,
+ 140,
+ 93
+ ],
+ "category_id": 6,
+ "area": 17860,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6359,
+ "image_id": 1175,
+ "bbox": [
+ 207,
+ 46,
+ 81,
+ 85
+ ],
+ "category_id": 6,
+ "area": 9460,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6360,
+ "image_id": 1175,
+ "bbox": [
+ 86,
+ 0,
+ 60,
+ 68
+ ],
+ "category_id": 6,
+ "area": 5658,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6361,
+ "image_id": 1175,
+ "bbox": [
+ 338,
+ 167,
+ 115,
+ 105
+ ],
+ "category_id": 6,
+ "area": 16585,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6374,
+ "image_id": 1179,
+ "bbox": [
+ 133,
+ 235,
+ 219,
+ 256
+ ],
+ "category_id": 8,
+ "area": 197091,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6375,
+ "image_id": 1179,
+ "bbox": [
+ 230,
+ 143,
+ 254,
+ 170
+ ],
+ "category_id": 8,
+ "area": 152243,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6376,
+ "image_id": 1180,
+ "bbox": [
+ 38,
+ 187,
+ 271,
+ 236
+ ],
+ "category_id": 9,
+ "area": 235690,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6393,
+ "image_id": 1182,
+ "bbox": [
+ 53,
+ 53,
+ 275,
+ 408
+ ],
+ "category_id": 9,
+ "area": 682370,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6397,
+ "image_id": 1183,
+ "bbox": [
+ 193,
+ 225,
+ 120,
+ 181
+ ],
+ "category_id": 10,
+ "area": 173499,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6398,
+ "image_id": 1183,
+ "bbox": [
+ 101,
+ 149,
+ 210,
+ 215
+ ],
+ "category_id": 9,
+ "area": 358660,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6403,
+ "image_id": 1184,
+ "bbox": [
+ 206,
+ 67,
+ 118,
+ 168
+ ],
+ "category_id": 9,
+ "area": 70389,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6404,
+ "image_id": 1184,
+ "bbox": [
+ 217,
+ 161,
+ 144,
+ 155
+ ],
+ "category_id": 9,
+ "area": 79278,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6449,
+ "image_id": 1192,
+ "bbox": [
+ 26,
+ 173,
+ 310,
+ 246
+ ],
+ "category_id": 9,
+ "area": 160140,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6450,
+ "image_id": 1192,
+ "bbox": [
+ 146,
+ 79,
+ 306,
+ 432
+ ],
+ "category_id": 9,
+ "area": 276650,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6456,
+ "image_id": 1194,
+ "bbox": [
+ 440,
+ 47,
+ 33,
+ 24
+ ],
+ "category_id": 10,
+ "area": 1320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6457,
+ "image_id": 1194,
+ "bbox": [
+ 425,
+ 69,
+ 23,
+ 24
+ ],
+ "category_id": 10,
+ "area": 912,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6458,
+ "image_id": 1194,
+ "bbox": [
+ 367,
+ 97,
+ 38,
+ 25
+ ],
+ "category_id": 10,
+ "area": 1575,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6459,
+ "image_id": 1194,
+ "bbox": [
+ 355,
+ 158,
+ 51,
+ 40
+ ],
+ "category_id": 10,
+ "area": 3360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6460,
+ "image_id": 1194,
+ "bbox": [
+ 408,
+ 119,
+ 53,
+ 38
+ ],
+ "category_id": 10,
+ "area": 3306,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6461,
+ "image_id": 1194,
+ "bbox": [
+ 421,
+ 195,
+ 47,
+ 34
+ ],
+ "category_id": 10,
+ "area": 2618,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6462,
+ "image_id": 1194,
+ "bbox": [
+ 337,
+ 226,
+ 62,
+ 68
+ ],
+ "category_id": 10,
+ "area": 6767,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6463,
+ "image_id": 1194,
+ "bbox": [
+ 254,
+ 22,
+ 24,
+ 49
+ ],
+ "category_id": 10,
+ "area": 1920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6464,
+ "image_id": 1194,
+ "bbox": [
+ 139,
+ 0,
+ 77,
+ 30
+ ],
+ "category_id": 10,
+ "area": 3750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6465,
+ "image_id": 1194,
+ "bbox": [
+ 27,
+ 0,
+ 99,
+ 35
+ ],
+ "category_id": 10,
+ "area": 5670,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6466,
+ "image_id": 1194,
+ "bbox": [
+ 293,
+ 66,
+ 42,
+ 31
+ ],
+ "category_id": 10,
+ "area": 2139,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6467,
+ "image_id": 1194,
+ "bbox": [
+ 276,
+ 0,
+ 62,
+ 38
+ ],
+ "category_id": 10,
+ "area": 3838,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6468,
+ "image_id": 1194,
+ "bbox": [
+ 313,
+ 0,
+ 62,
+ 71
+ ],
+ "category_id": 10,
+ "area": 7140,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6469,
+ "image_id": 1194,
+ "bbox": [
+ 97,
+ 132,
+ 37,
+ 33
+ ],
+ "category_id": 10,
+ "area": 1980,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6470,
+ "image_id": 1194,
+ "bbox": [
+ 0,
+ 129,
+ 48,
+ 105
+ ],
+ "category_id": 10,
+ "area": 8034,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6544,
+ "image_id": 1195,
+ "bbox": [
+ 11,
+ 77,
+ 158,
+ 277
+ ],
+ "category_id": 9,
+ "area": 102897,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6545,
+ "image_id": 1195,
+ "bbox": [
+ 357,
+ 391,
+ 135,
+ 118
+ ],
+ "category_id": 9,
+ "area": 37630,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6546,
+ "image_id": 1195,
+ "bbox": [
+ 143,
+ 345,
+ 135,
+ 166
+ ],
+ "category_id": 9,
+ "area": 53000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6547,
+ "image_id": 1195,
+ "bbox": [
+ 166,
+ 208,
+ 114,
+ 197
+ ],
+ "category_id": 9,
+ "area": 53088,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6548,
+ "image_id": 1195,
+ "bbox": [
+ 244,
+ 72,
+ 124,
+ 165
+ ],
+ "category_id": 9,
+ "area": 48312,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6549,
+ "image_id": 1195,
+ "bbox": [
+ 244,
+ 105,
+ 112,
+ 171
+ ],
+ "category_id": 9,
+ "area": 45320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6572,
+ "image_id": 1197,
+ "bbox": [
+ 98,
+ 146,
+ 331,
+ 325
+ ],
+ "category_id": 9,
+ "area": 332258,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6592,
+ "image_id": 1198,
+ "bbox": [
+ 174,
+ 152,
+ 128,
+ 124
+ ],
+ "category_id": 10,
+ "area": 48100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6593,
+ "image_id": 1199,
+ "bbox": [
+ 93,
+ 291,
+ 99,
+ 84
+ ],
+ "category_id": 9,
+ "area": 12714,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6596,
+ "image_id": 1199,
+ "bbox": [
+ 70,
+ 387,
+ 45,
+ 72
+ ],
+ "category_id": 6,
+ "area": 4958,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6597,
+ "image_id": 1200,
+ "bbox": [
+ 68,
+ 160,
+ 182,
+ 336
+ ],
+ "category_id": 8,
+ "area": 105915,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6598,
+ "image_id": 1200,
+ "bbox": [
+ 198,
+ 40,
+ 296,
+ 311
+ ],
+ "category_id": 8,
+ "area": 159324,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6599,
+ "image_id": 1201,
+ "bbox": [
+ 447,
+ 55,
+ 64,
+ 105
+ ],
+ "category_id": 6,
+ "area": 116865,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6600,
+ "image_id": 1201,
+ "bbox": [
+ 257,
+ 50,
+ 185,
+ 144
+ ],
+ "category_id": 6,
+ "area": 459510,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6601,
+ "image_id": 1201,
+ "bbox": [
+ 190,
+ 146,
+ 59,
+ 110
+ ],
+ "category_id": 6,
+ "area": 111643,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6602,
+ "image_id": 1201,
+ "bbox": [
+ 218,
+ 62,
+ 76,
+ 62
+ ],
+ "category_id": 6,
+ "area": 81400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6603,
+ "image_id": 1201,
+ "bbox": [
+ 176,
+ 133,
+ 72,
+ 72
+ ],
+ "category_id": 6,
+ "area": 90624,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6604,
+ "image_id": 1201,
+ "bbox": [
+ 0,
+ 36,
+ 71,
+ 137
+ ],
+ "category_id": 6,
+ "area": 169400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6605,
+ "image_id": 1201,
+ "bbox": [
+ 1,
+ 174,
+ 92,
+ 97
+ ],
+ "category_id": 6,
+ "area": 155488,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6606,
+ "image_id": 1201,
+ "bbox": [
+ 145,
+ 189,
+ 74,
+ 93
+ ],
+ "category_id": 6,
+ "area": 119427,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6607,
+ "image_id": 1201,
+ "bbox": [
+ 0,
+ 263,
+ 62,
+ 91
+ ],
+ "category_id": 6,
+ "area": 98532,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6608,
+ "image_id": 1201,
+ "bbox": [
+ 90,
+ 290,
+ 45,
+ 86
+ ],
+ "category_id": 6,
+ "area": 66660,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6609,
+ "image_id": 1201,
+ "bbox": [
+ 144,
+ 310,
+ 54,
+ 83
+ ],
+ "category_id": 6,
+ "area": 77645,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6610,
+ "image_id": 1201,
+ "bbox": [
+ 230,
+ 406,
+ 98,
+ 105
+ ],
+ "category_id": 6,
+ "area": 178080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6611,
+ "image_id": 1201,
+ "bbox": [
+ 122,
+ 67,
+ 84,
+ 74
+ ],
+ "category_id": 6,
+ "area": 107830,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6612,
+ "image_id": 1201,
+ "bbox": [
+ 403,
+ 132,
+ 96,
+ 158
+ ],
+ "category_id": 6,
+ "area": 260119,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6617,
+ "image_id": 1201,
+ "bbox": [
+ 68,
+ 241,
+ 46,
+ 58
+ ],
+ "category_id": 6,
+ "area": 46556,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6637,
+ "image_id": 1206,
+ "bbox": [
+ 99,
+ 14,
+ 59,
+ 174
+ ],
+ "category_id": 8,
+ "area": 34992,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6638,
+ "image_id": 1206,
+ "bbox": [
+ 164,
+ 260,
+ 223,
+ 195
+ ],
+ "category_id": 8,
+ "area": 148104,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6639,
+ "image_id": 1206,
+ "bbox": [
+ 331,
+ 260,
+ 157,
+ 206
+ ],
+ "category_id": 8,
+ "area": 110080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6640,
+ "image_id": 1206,
+ "bbox": [
+ 81,
+ 201,
+ 143,
+ 114
+ ],
+ "category_id": 8,
+ "area": 55806,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6641,
+ "image_id": 1206,
+ "bbox": [
+ 204,
+ 148,
+ 58,
+ 113
+ ],
+ "category_id": 8,
+ "area": 22400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6642,
+ "image_id": 1207,
+ "bbox": [
+ 151,
+ 137,
+ 276,
+ 374
+ ],
+ "category_id": 10,
+ "area": 171360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6649,
+ "image_id": 1209,
+ "bbox": [
+ 50,
+ 283,
+ 148,
+ 160
+ ],
+ "category_id": 9,
+ "area": 91866,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6650,
+ "image_id": 1209,
+ "bbox": [
+ 0,
+ 278,
+ 59,
+ 81
+ ],
+ "category_id": 6,
+ "area": 18944,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6651,
+ "image_id": 1209,
+ "bbox": [
+ 36,
+ 207,
+ 50,
+ 63
+ ],
+ "category_id": 6,
+ "area": 12276,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6652,
+ "image_id": 1209,
+ "bbox": [
+ 165,
+ 171,
+ 27,
+ 52
+ ],
+ "category_id": 6,
+ "area": 5658,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6653,
+ "image_id": 1209,
+ "bbox": [
+ 237,
+ 199,
+ 19,
+ 23
+ ],
+ "category_id": 6,
+ "area": 1739,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6654,
+ "image_id": 1210,
+ "bbox": [
+ 126,
+ 144,
+ 185,
+ 238
+ ],
+ "category_id": 8,
+ "area": 349776,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6655,
+ "image_id": 1210,
+ "bbox": [
+ 23,
+ 232,
+ 105,
+ 150
+ ],
+ "category_id": 6,
+ "area": 125292,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6656,
+ "image_id": 1210,
+ "bbox": [
+ 90,
+ 24,
+ 77,
+ 91
+ ],
+ "category_id": 6,
+ "area": 56356,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6657,
+ "image_id": 1210,
+ "bbox": [
+ 355,
+ 414,
+ 57,
+ 92
+ ],
+ "category_id": 6,
+ "area": 41925,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6658,
+ "image_id": 1210,
+ "bbox": [
+ 374,
+ 222,
+ 72,
+ 92
+ ],
+ "category_id": 6,
+ "area": 53235,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6671,
+ "image_id": 1213,
+ "bbox": [
+ 133,
+ 157,
+ 355,
+ 246
+ ],
+ "category_id": 9,
+ "area": 307594,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6672,
+ "image_id": 1213,
+ "bbox": [
+ 148,
+ 146,
+ 76,
+ 174
+ ],
+ "category_id": 10,
+ "area": 46550,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6674,
+ "image_id": 1214,
+ "bbox": [
+ 107,
+ 153,
+ 295,
+ 199
+ ],
+ "category_id": 9,
+ "area": 151765,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6676,
+ "image_id": 1216,
+ "bbox": [
+ 104,
+ 217,
+ 42,
+ 65
+ ],
+ "category_id": 6,
+ "area": 22080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6677,
+ "image_id": 1216,
+ "bbox": [
+ 265,
+ 309,
+ 149,
+ 170
+ ],
+ "category_id": 6,
+ "area": 201600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6709,
+ "image_id": 1217,
+ "bbox": [
+ 175,
+ 24,
+ 70,
+ 128
+ ],
+ "category_id": 8,
+ "area": 71544,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6710,
+ "image_id": 1217,
+ "bbox": [
+ 188,
+ 111,
+ 166,
+ 302
+ ],
+ "category_id": 8,
+ "area": 399375,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6711,
+ "image_id": 1217,
+ "bbox": [
+ 237,
+ 293,
+ 150,
+ 217
+ ],
+ "category_id": 8,
+ "area": 259794,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6716,
+ "image_id": 1219,
+ "bbox": [
+ 220,
+ 93,
+ 281,
+ 418
+ ],
+ "category_id": 9,
+ "area": 94416,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6782,
+ "image_id": 1228,
+ "bbox": [
+ 88,
+ 39,
+ 134,
+ 128
+ ],
+ "category_id": 8,
+ "area": 42174,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6783,
+ "image_id": 1228,
+ "bbox": [
+ 67,
+ 101,
+ 379,
+ 356
+ ],
+ "category_id": 8,
+ "area": 329810,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6787,
+ "image_id": 1230,
+ "bbox": [
+ 213,
+ 170,
+ 187,
+ 187
+ ],
+ "category_id": 8,
+ "area": 46288,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6792,
+ "image_id": 1232,
+ "bbox": [
+ 141,
+ 245,
+ 145,
+ 150
+ ],
+ "category_id": 8,
+ "area": 70059,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6793,
+ "image_id": 1232,
+ "bbox": [
+ 318,
+ 158,
+ 89,
+ 115
+ ],
+ "category_id": 8,
+ "area": 33004,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6796,
+ "image_id": 1234,
+ "bbox": [
+ 72,
+ 56,
+ 428,
+ 450
+ ],
+ "category_id": 9,
+ "area": 679648,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6797,
+ "image_id": 1235,
+ "bbox": [
+ 56,
+ 40,
+ 157,
+ 127
+ ],
+ "category_id": 8,
+ "area": 50328,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6798,
+ "image_id": 1235,
+ "bbox": [
+ 32,
+ 102,
+ 449,
+ 357
+ ],
+ "category_id": 8,
+ "area": 399714,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6803,
+ "image_id": 1238,
+ "bbox": [
+ 126,
+ 114,
+ 351,
+ 251
+ ],
+ "category_id": 9,
+ "area": 311166,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6809,
+ "image_id": 1240,
+ "bbox": [
+ 117,
+ 322,
+ 147,
+ 187
+ ],
+ "category_id": 9,
+ "area": 41452,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6810,
+ "image_id": 1241,
+ "bbox": [
+ 26,
+ 59,
+ 481,
+ 450
+ ],
+ "category_id": 6,
+ "area": 592245,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6811,
+ "image_id": 1241,
+ "bbox": [
+ 226,
+ 109,
+ 98,
+ 118
+ ],
+ "category_id": 8,
+ "area": 31752,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6812,
+ "image_id": 1242,
+ "bbox": [
+ 2,
+ 44,
+ 300,
+ 305
+ ],
+ "category_id": 8,
+ "area": 322500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6813,
+ "image_id": 1242,
+ "bbox": [
+ 95,
+ 231,
+ 390,
+ 205
+ ],
+ "category_id": 8,
+ "area": 281775,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6814,
+ "image_id": 1243,
+ "bbox": [
+ 131,
+ 212,
+ 156,
+ 158
+ ],
+ "category_id": 8,
+ "area": 86580,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6815,
+ "image_id": 1243,
+ "bbox": [
+ 240,
+ 209,
+ 202,
+ 154
+ ],
+ "category_id": 8,
+ "area": 109512,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6817,
+ "image_id": 1243,
+ "bbox": [
+ 160,
+ 374,
+ 35,
+ 37
+ ],
+ "category_id": 6,
+ "area": 4628,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6818,
+ "image_id": 1243,
+ "bbox": [
+ 116,
+ 100,
+ 27,
+ 29
+ ],
+ "category_id": 6,
+ "area": 2898,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6819,
+ "image_id": 1243,
+ "bbox": [
+ 147,
+ 134,
+ 26,
+ 31
+ ],
+ "category_id": 6,
+ "area": 2860,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6820,
+ "image_id": 1243,
+ "bbox": [
+ 156,
+ 183,
+ 31,
+ 31
+ ],
+ "category_id": 6,
+ "area": 3476,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6821,
+ "image_id": 1243,
+ "bbox": [
+ 179,
+ 164,
+ 18,
+ 16
+ ],
+ "category_id": 6,
+ "area": 1035,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6830,
+ "image_id": 1245,
+ "bbox": [
+ 64,
+ 124,
+ 228,
+ 269
+ ],
+ "category_id": 9,
+ "area": 486208,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6834,
+ "image_id": 1247,
+ "bbox": [
+ 2,
+ 0,
+ 375,
+ 374
+ ],
+ "category_id": 9,
+ "area": 494326,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6835,
+ "image_id": 1247,
+ "bbox": [
+ 2,
+ 85,
+ 304,
+ 421
+ ],
+ "category_id": 9,
+ "area": 450680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6836,
+ "image_id": 1247,
+ "bbox": [
+ 269,
+ 0,
+ 239,
+ 506
+ ],
+ "category_id": 9,
+ "area": 425776,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6838,
+ "image_id": 1249,
+ "bbox": [
+ 159,
+ 155,
+ 151,
+ 214
+ ],
+ "category_id": 8,
+ "area": 250172,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6839,
+ "image_id": 1249,
+ "bbox": [
+ 129,
+ 0,
+ 90,
+ 85
+ ],
+ "category_id": 6,
+ "area": 59826,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6840,
+ "image_id": 1249,
+ "bbox": [
+ 230,
+ 0,
+ 126,
+ 129
+ ],
+ "category_id": 6,
+ "area": 126291,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6841,
+ "image_id": 1249,
+ "bbox": [
+ 317,
+ 0,
+ 174,
+ 286
+ ],
+ "category_id": 6,
+ "area": 386514,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6851,
+ "image_id": 1251,
+ "bbox": [
+ 24,
+ 140,
+ 487,
+ 371
+ ],
+ "category_id": 6,
+ "area": 110044,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6852,
+ "image_id": 1252,
+ "bbox": [
+ 300,
+ 159,
+ 211,
+ 283
+ ],
+ "category_id": 8,
+ "area": 69960,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6855,
+ "image_id": 1253,
+ "bbox": [
+ 214,
+ 32,
+ 169,
+ 345
+ ],
+ "category_id": 10,
+ "area": 102664,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6856,
+ "image_id": 1253,
+ "bbox": [
+ 59,
+ 311,
+ 84,
+ 168
+ ],
+ "category_id": 10,
+ "area": 24960,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6857,
+ "image_id": 1253,
+ "bbox": [
+ 341,
+ 365,
+ 58,
+ 122
+ ],
+ "category_id": 10,
+ "area": 12528,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6858,
+ "image_id": 1253,
+ "bbox": [
+ 306,
+ 361,
+ 42,
+ 71
+ ],
+ "category_id": 10,
+ "area": 5372,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6863,
+ "image_id": 1255,
+ "bbox": [
+ 71,
+ 223,
+ 347,
+ 288
+ ],
+ "category_id": 9,
+ "area": 247919,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6864,
+ "image_id": 1255,
+ "bbox": [
+ 411,
+ 292,
+ 77,
+ 134
+ ],
+ "category_id": 6,
+ "area": 25718,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6866,
+ "image_id": 1256,
+ "bbox": [
+ 137,
+ 189,
+ 148,
+ 159
+ ],
+ "category_id": 8,
+ "area": 82880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6867,
+ "image_id": 1256,
+ "bbox": [
+ 324,
+ 146,
+ 176,
+ 254
+ ],
+ "category_id": 8,
+ "area": 157352,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6869,
+ "image_id": 1257,
+ "bbox": [
+ 13,
+ 241,
+ 483,
+ 270
+ ],
+ "category_id": 10,
+ "area": 176715,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6877,
+ "image_id": 1259,
+ "bbox": [
+ 133,
+ 0,
+ 378,
+ 431
+ ],
+ "category_id": 10,
+ "area": 368532,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6879,
+ "image_id": 1260,
+ "bbox": [
+ 98,
+ 124,
+ 96,
+ 184
+ ],
+ "category_id": 8,
+ "area": 57121,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6880,
+ "image_id": 1260,
+ "bbox": [
+ 312,
+ 266,
+ 97,
+ 106
+ ],
+ "category_id": 8,
+ "area": 33396,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6913,
+ "image_id": 1263,
+ "bbox": [
+ 179,
+ 200,
+ 70,
+ 30
+ ],
+ "category_id": 8,
+ "area": 16960,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6914,
+ "image_id": 1263,
+ "bbox": [
+ 138,
+ 222,
+ 114,
+ 118
+ ],
+ "category_id": 8,
+ "area": 107679,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6915,
+ "image_id": 1263,
+ "bbox": [
+ 145,
+ 260,
+ 129,
+ 54
+ ],
+ "category_id": 8,
+ "area": 56260,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6916,
+ "image_id": 1263,
+ "bbox": [
+ 339,
+ 152,
+ 42,
+ 60
+ ],
+ "category_id": 6,
+ "area": 20066,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6917,
+ "image_id": 1264,
+ "bbox": [
+ 57,
+ 0,
+ 272,
+ 128
+ ],
+ "category_id": 8,
+ "area": 42427,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6919,
+ "image_id": 1265,
+ "bbox": [
+ 47,
+ 79,
+ 118,
+ 238
+ ],
+ "category_id": 9,
+ "area": 66066,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6920,
+ "image_id": 1265,
+ "bbox": [
+ 384,
+ 417,
+ 91,
+ 92
+ ],
+ "category_id": 9,
+ "area": 19758,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6921,
+ "image_id": 1265,
+ "bbox": [
+ 120,
+ 243,
+ 153,
+ 268
+ ],
+ "category_id": 9,
+ "area": 96600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6922,
+ "image_id": 1265,
+ "bbox": [
+ 166,
+ 174,
+ 117,
+ 204
+ ],
+ "category_id": 9,
+ "area": 56105,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6923,
+ "image_id": 1265,
+ "bbox": [
+ 247,
+ 62,
+ 103,
+ 179
+ ],
+ "category_id": 9,
+ "area": 43645,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6924,
+ "image_id": 1265,
+ "bbox": [
+ 265,
+ 100,
+ 112,
+ 180
+ ],
+ "category_id": 9,
+ "area": 47740,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6927,
+ "image_id": 1267,
+ "bbox": [
+ 0,
+ 155,
+ 391,
+ 281
+ ],
+ "category_id": 9,
+ "area": 1520386,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6928,
+ "image_id": 1267,
+ "bbox": [
+ 277,
+ 130,
+ 153,
+ 232
+ ],
+ "category_id": 8,
+ "area": 492480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6933,
+ "image_id": 1269,
+ "bbox": [
+ 166,
+ 189,
+ 124,
+ 182
+ ],
+ "category_id": 8,
+ "area": 79305,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6934,
+ "image_id": 1269,
+ "bbox": [
+ 252,
+ 147,
+ 150,
+ 167
+ ],
+ "category_id": 8,
+ "area": 88125,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6936,
+ "image_id": 1271,
+ "bbox": [
+ 10,
+ 141,
+ 104,
+ 258
+ ],
+ "category_id": 9,
+ "area": 94380,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6937,
+ "image_id": 1271,
+ "bbox": [
+ 342,
+ 295,
+ 81,
+ 214
+ ],
+ "category_id": 9,
+ "area": 61103,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6938,
+ "image_id": 1271,
+ "bbox": [
+ 119,
+ 250,
+ 100,
+ 241
+ ],
+ "category_id": 9,
+ "area": 85000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6939,
+ "image_id": 1271,
+ "bbox": [
+ 189,
+ 174,
+ 130,
+ 286
+ ],
+ "category_id": 9,
+ "area": 130975,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6940,
+ "image_id": 1271,
+ "bbox": [
+ 234,
+ 84,
+ 126,
+ 296
+ ],
+ "category_id": 9,
+ "area": 132189,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6941,
+ "image_id": 1271,
+ "bbox": [
+ 145,
+ 0,
+ 178,
+ 295
+ ],
+ "category_id": 9,
+ "area": 185505,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6942,
+ "image_id": 1272,
+ "bbox": [
+ 166,
+ 53,
+ 49,
+ 93
+ ],
+ "category_id": 10,
+ "area": 16244,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6943,
+ "image_id": 1272,
+ "bbox": [
+ 6,
+ 41,
+ 29,
+ 105
+ ],
+ "category_id": 10,
+ "area": 11026,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6944,
+ "image_id": 1272,
+ "bbox": [
+ 82,
+ 219,
+ 22,
+ 50
+ ],
+ "category_id": 10,
+ "area": 3976,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6945,
+ "image_id": 1272,
+ "bbox": [
+ 92,
+ 442,
+ 15,
+ 30
+ ],
+ "category_id": 10,
+ "area": 1634,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6946,
+ "image_id": 1272,
+ "bbox": [
+ 68,
+ 398,
+ 14,
+ 32
+ ],
+ "category_id": 10,
+ "area": 1575,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6947,
+ "image_id": 1272,
+ "bbox": [
+ 370,
+ 409,
+ 44,
+ 64
+ ],
+ "category_id": 10,
+ "area": 9900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6948,
+ "image_id": 1272,
+ "bbox": [
+ 346,
+ 4,
+ 32,
+ 27
+ ],
+ "category_id": 10,
+ "area": 3198,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6949,
+ "image_id": 1272,
+ "bbox": [
+ 400,
+ 7,
+ 20,
+ 27
+ ],
+ "category_id": 10,
+ "area": 2028,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6950,
+ "image_id": 1272,
+ "bbox": [
+ 457,
+ 41,
+ 22,
+ 44
+ ],
+ "category_id": 10,
+ "area": 3591,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6951,
+ "image_id": 1272,
+ "bbox": [
+ 468,
+ 89,
+ 22,
+ 42
+ ],
+ "category_id": 10,
+ "area": 3300,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6952,
+ "image_id": 1272,
+ "bbox": [
+ 438,
+ 166,
+ 17,
+ 44
+ ],
+ "category_id": 10,
+ "area": 2666,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6953,
+ "image_id": 1272,
+ "bbox": [
+ 450,
+ 226,
+ 18,
+ 33
+ ],
+ "category_id": 10,
+ "area": 2162,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6954,
+ "image_id": 1272,
+ "bbox": [
+ 361,
+ 265,
+ 24,
+ 40
+ ],
+ "category_id": 10,
+ "area": 3477,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6955,
+ "image_id": 1272,
+ "bbox": [
+ 421,
+ 272,
+ 20,
+ 32
+ ],
+ "category_id": 10,
+ "area": 2295,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6961,
+ "image_id": 1274,
+ "bbox": [
+ 42,
+ 242,
+ 162,
+ 128
+ ],
+ "category_id": 9,
+ "area": 18025,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6962,
+ "image_id": 1274,
+ "bbox": [
+ 108,
+ 408,
+ 46,
+ 70
+ ],
+ "category_id": 6,
+ "area": 2850,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6963,
+ "image_id": 1274,
+ "bbox": [
+ 18,
+ 401,
+ 87,
+ 57
+ ],
+ "category_id": 6,
+ "area": 4324,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6964,
+ "image_id": 1274,
+ "bbox": [
+ 0,
+ 339,
+ 108,
+ 65
+ ],
+ "category_id": 6,
+ "area": 6201,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6969,
+ "image_id": 1278,
+ "bbox": [
+ 93,
+ 160,
+ 90,
+ 86
+ ],
+ "category_id": 6,
+ "area": 27572,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6970,
+ "image_id": 1278,
+ "bbox": [
+ 95,
+ 215,
+ 68,
+ 109
+ ],
+ "category_id": 6,
+ "area": 26334,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6971,
+ "image_id": 1278,
+ "bbox": [
+ 288,
+ 359,
+ 87,
+ 89
+ ],
+ "category_id": 6,
+ "area": 27594,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6972,
+ "image_id": 1278,
+ "bbox": [
+ 176,
+ 443,
+ 99,
+ 68
+ ],
+ "category_id": 6,
+ "area": 24153,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6973,
+ "image_id": 1278,
+ "bbox": [
+ 280,
+ 406,
+ 123,
+ 105
+ ],
+ "category_id": 6,
+ "area": 45584,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6974,
+ "image_id": 1278,
+ "bbox": [
+ 374,
+ 374,
+ 137,
+ 137
+ ],
+ "category_id": 6,
+ "area": 66736,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6975,
+ "image_id": 1278,
+ "bbox": [
+ 216,
+ 297,
+ 37,
+ 48
+ ],
+ "category_id": 6,
+ "area": 6324,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6976,
+ "image_id": 1278,
+ "bbox": [
+ 3,
+ 283,
+ 48,
+ 130
+ ],
+ "category_id": 6,
+ "area": 22448,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6977,
+ "image_id": 1278,
+ "bbox": [
+ 103,
+ 324,
+ 44,
+ 56
+ ],
+ "category_id": 6,
+ "area": 8690,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6978,
+ "image_id": 1278,
+ "bbox": [
+ 378,
+ 244,
+ 37,
+ 54
+ ],
+ "category_id": 6,
+ "area": 7068,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6979,
+ "image_id": 1278,
+ "bbox": [
+ 217,
+ 192,
+ 106,
+ 86
+ ],
+ "category_id": 8,
+ "area": 32452,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6994,
+ "image_id": 1280,
+ "bbox": [
+ 246,
+ 148,
+ 117,
+ 309
+ ],
+ "category_id": 10,
+ "area": 286667,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6995,
+ "image_id": 1281,
+ "bbox": [
+ 90,
+ 265,
+ 232,
+ 209
+ ],
+ "category_id": 9,
+ "area": 81947,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 6996,
+ "image_id": 1281,
+ "bbox": [
+ 0,
+ 307,
+ 243,
+ 204
+ ],
+ "category_id": 6,
+ "area": 83694,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7000,
+ "image_id": 1283,
+ "bbox": [
+ 90,
+ 143,
+ 219,
+ 269
+ ],
+ "category_id": 9,
+ "area": 123480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7001,
+ "image_id": 1283,
+ "bbox": [
+ 131,
+ 61,
+ 295,
+ 450
+ ],
+ "category_id": 9,
+ "area": 277905,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7041,
+ "image_id": 1287,
+ "bbox": [
+ 134,
+ 215,
+ 107,
+ 91
+ ],
+ "category_id": 8,
+ "area": 77586,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7042,
+ "image_id": 1287,
+ "bbox": [
+ 180,
+ 146,
+ 145,
+ 169
+ ],
+ "category_id": 8,
+ "area": 194752,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7043,
+ "image_id": 1287,
+ "bbox": [
+ 305,
+ 153,
+ 148,
+ 196
+ ],
+ "category_id": 8,
+ "area": 231155,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7044,
+ "image_id": 1288,
+ "bbox": [
+ 55,
+ 55,
+ 224,
+ 202
+ ],
+ "category_id": 9,
+ "area": 129944,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7050,
+ "image_id": 1289,
+ "bbox": [
+ 149,
+ 94,
+ 362,
+ 405
+ ],
+ "category_id": 9,
+ "area": 448610,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7051,
+ "image_id": 1289,
+ "bbox": [
+ 0,
+ 112,
+ 171,
+ 138
+ ],
+ "category_id": 6,
+ "area": 72568,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7052,
+ "image_id": 1289,
+ "bbox": [
+ 170,
+ 149,
+ 90,
+ 73
+ ],
+ "category_id": 6,
+ "area": 20196,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7053,
+ "image_id": 1289,
+ "bbox": [
+ 0,
+ 247,
+ 291,
+ 263
+ ],
+ "category_id": 6,
+ "area": 234880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7054,
+ "image_id": 1289,
+ "bbox": [
+ 283,
+ 352,
+ 124,
+ 157
+ ],
+ "category_id": 6,
+ "area": 60280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7055,
+ "image_id": 1289,
+ "bbox": [
+ 353,
+ 311,
+ 75,
+ 112
+ ],
+ "category_id": 6,
+ "area": 25905,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7056,
+ "image_id": 1289,
+ "bbox": [
+ 370,
+ 300,
+ 72,
+ 70
+ ],
+ "category_id": 6,
+ "area": 15840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7057,
+ "image_id": 1289,
+ "bbox": [
+ 426,
+ 333,
+ 85,
+ 176
+ ],
+ "category_id": 6,
+ "area": 46248,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7058,
+ "image_id": 1290,
+ "bbox": [
+ 57,
+ 242,
+ 407,
+ 140
+ ],
+ "category_id": 8,
+ "area": 61362,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7062,
+ "image_id": 1291,
+ "bbox": [
+ 23,
+ 105,
+ 483,
+ 403
+ ],
+ "category_id": 9,
+ "area": 674610,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7064,
+ "image_id": 1292,
+ "bbox": [
+ 33,
+ 199,
+ 263,
+ 179
+ ],
+ "category_id": 9,
+ "area": 111639,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7066,
+ "image_id": 1293,
+ "bbox": [
+ 4,
+ 6,
+ 472,
+ 392
+ ],
+ "category_id": 9,
+ "area": 651360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7083,
+ "image_id": 1296,
+ "bbox": [
+ 143,
+ 480,
+ 87,
+ 31
+ ],
+ "category_id": 6,
+ "area": 7755,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7085,
+ "image_id": 1297,
+ "bbox": [
+ 0,
+ 49,
+ 99,
+ 159
+ ],
+ "category_id": 6,
+ "area": 184041,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7086,
+ "image_id": 1297,
+ "bbox": [
+ 17,
+ 164,
+ 68,
+ 50
+ ],
+ "category_id": 6,
+ "area": 40000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7087,
+ "image_id": 1297,
+ "bbox": [
+ 0,
+ 201,
+ 59,
+ 105
+ ],
+ "category_id": 6,
+ "area": 72695,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7088,
+ "image_id": 1297,
+ "bbox": [
+ 0,
+ 304,
+ 81,
+ 92
+ ],
+ "category_id": 6,
+ "area": 87320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7089,
+ "image_id": 1297,
+ "bbox": [
+ 86,
+ 98,
+ 170,
+ 120
+ ],
+ "category_id": 6,
+ "area": 237604,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7090,
+ "image_id": 1297,
+ "bbox": [
+ 96,
+ 306,
+ 58,
+ 30
+ ],
+ "category_id": 6,
+ "area": 20448,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7091,
+ "image_id": 1297,
+ "bbox": [
+ 129,
+ 179,
+ 72,
+ 168
+ ],
+ "category_id": 6,
+ "area": 140968,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7092,
+ "image_id": 1297,
+ "bbox": [
+ 123,
+ 331,
+ 59,
+ 95
+ ],
+ "category_id": 6,
+ "area": 65751,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7093,
+ "image_id": 1297,
+ "bbox": [
+ 199,
+ 353,
+ 66,
+ 94
+ ],
+ "category_id": 6,
+ "area": 73444,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7094,
+ "image_id": 1297,
+ "bbox": [
+ 195,
+ 218,
+ 92,
+ 106
+ ],
+ "category_id": 6,
+ "area": 114580,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7095,
+ "image_id": 1297,
+ "bbox": [
+ 237,
+ 157,
+ 92,
+ 81
+ ],
+ "category_id": 6,
+ "area": 86946,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7096,
+ "image_id": 1297,
+ "bbox": [
+ 343,
+ 66,
+ 167,
+ 159
+ ],
+ "category_id": 6,
+ "area": 308154,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7097,
+ "image_id": 1297,
+ "bbox": [
+ 259,
+ 172,
+ 76,
+ 119
+ ],
+ "category_id": 6,
+ "area": 106400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7098,
+ "image_id": 1297,
+ "bbox": [
+ 368,
+ 278,
+ 73,
+ 60
+ ],
+ "category_id": 6,
+ "area": 51456,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7099,
+ "image_id": 1297,
+ "bbox": [
+ 315,
+ 448,
+ 115,
+ 61
+ ],
+ "category_id": 6,
+ "area": 82740,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7112,
+ "image_id": 1299,
+ "bbox": [
+ 148,
+ 105,
+ 119,
+ 291
+ ],
+ "category_id": 8,
+ "area": 100426,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7113,
+ "image_id": 1299,
+ "bbox": [
+ 132,
+ 214,
+ 94,
+ 121
+ ],
+ "category_id": 6,
+ "area": 33135,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7117,
+ "image_id": 1300,
+ "bbox": [
+ 61,
+ 165,
+ 311,
+ 334
+ ],
+ "category_id": 9,
+ "area": 258912,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7118,
+ "image_id": 1301,
+ "bbox": [
+ 0,
+ 57,
+ 322,
+ 452
+ ],
+ "category_id": 9,
+ "area": 304968,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7119,
+ "image_id": 1301,
+ "bbox": [
+ 257,
+ 97,
+ 115,
+ 413
+ ],
+ "category_id": 9,
+ "area": 99484,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7122,
+ "image_id": 1303,
+ "bbox": [
+ 124,
+ 48,
+ 126,
+ 242
+ ],
+ "category_id": 6,
+ "area": 243236,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7123,
+ "image_id": 1303,
+ "bbox": [
+ 348,
+ 59,
+ 110,
+ 202
+ ],
+ "category_id": 6,
+ "area": 177632,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7124,
+ "image_id": 1303,
+ "bbox": [
+ 416,
+ 278,
+ 94,
+ 137
+ ],
+ "category_id": 6,
+ "area": 103240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7147,
+ "image_id": 1305,
+ "bbox": [
+ 96,
+ 272,
+ 284,
+ 177
+ ],
+ "category_id": 8,
+ "area": 341544,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7148,
+ "image_id": 1305,
+ "bbox": [
+ 185,
+ 34,
+ 179,
+ 392
+ ],
+ "category_id": 8,
+ "area": 476448,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7149,
+ "image_id": 1305,
+ "bbox": [
+ 178,
+ 27,
+ 247,
+ 253
+ ],
+ "category_id": 8,
+ "area": 423182,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7150,
+ "image_id": 1306,
+ "bbox": [
+ 0,
+ 39,
+ 198,
+ 247
+ ],
+ "category_id": 8,
+ "area": 170924,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7151,
+ "image_id": 1306,
+ "bbox": [
+ 160,
+ 214,
+ 217,
+ 158
+ ],
+ "category_id": 8,
+ "area": 119340,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7154,
+ "image_id": 1308,
+ "bbox": [
+ 65,
+ 41,
+ 175,
+ 112
+ ],
+ "category_id": 8,
+ "area": 25839,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7155,
+ "image_id": 1308,
+ "bbox": [
+ 114,
+ 212,
+ 303,
+ 229
+ ],
+ "category_id": 8,
+ "area": 91553,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7157,
+ "image_id": 1309,
+ "bbox": [
+ 68,
+ 327,
+ 187,
+ 156
+ ],
+ "category_id": 6,
+ "area": 33456,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7158,
+ "image_id": 1309,
+ "bbox": [
+ 77,
+ 137,
+ 38,
+ 52
+ ],
+ "category_id": 6,
+ "area": 2346,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7159,
+ "image_id": 1309,
+ "bbox": [
+ 121,
+ 156,
+ 40,
+ 42
+ ],
+ "category_id": 6,
+ "area": 1961,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7160,
+ "image_id": 1309,
+ "bbox": [
+ 110,
+ 250,
+ 52,
+ 76
+ ],
+ "category_id": 6,
+ "area": 4623,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7161,
+ "image_id": 1309,
+ "bbox": [
+ 163,
+ 239,
+ 75,
+ 94
+ ],
+ "category_id": 6,
+ "area": 8118,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7164,
+ "image_id": 1310,
+ "bbox": [
+ 177,
+ 86,
+ 312,
+ 425
+ ],
+ "category_id": 9,
+ "area": 199019,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7165,
+ "image_id": 1311,
+ "bbox": [
+ 137,
+ 0,
+ 249,
+ 512
+ ],
+ "category_id": 9,
+ "area": 84180,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7169,
+ "image_id": 1313,
+ "bbox": [
+ 151,
+ 178,
+ 327,
+ 332
+ ],
+ "category_id": 8,
+ "area": 126480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7171,
+ "image_id": 1314,
+ "bbox": [
+ 21,
+ 50,
+ 325,
+ 460
+ ],
+ "category_id": 10,
+ "area": 1184620,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7199,
+ "image_id": 1317,
+ "bbox": [
+ 161,
+ 46,
+ 220,
+ 443
+ ],
+ "category_id": 10,
+ "area": 173840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7251,
+ "image_id": 1320,
+ "bbox": [
+ 154,
+ 172,
+ 297,
+ 334
+ ],
+ "category_id": 9,
+ "area": 349953,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7252,
+ "image_id": 1320,
+ "bbox": [
+ 244,
+ 91,
+ 230,
+ 385
+ ],
+ "category_id": 9,
+ "area": 312734,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7259,
+ "image_id": 1323,
+ "bbox": [
+ 12,
+ 37,
+ 475,
+ 421
+ ],
+ "category_id": 9,
+ "area": 366206,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7262,
+ "image_id": 1324,
+ "bbox": [
+ 338,
+ 170,
+ 167,
+ 292
+ ],
+ "category_id": 8,
+ "area": 64664,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7263,
+ "image_id": 1325,
+ "bbox": [
+ 213,
+ 110,
+ 191,
+ 401
+ ],
+ "category_id": 10,
+ "area": 128305,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7264,
+ "image_id": 1325,
+ "bbox": [
+ 98,
+ 116,
+ 83,
+ 118
+ ],
+ "category_id": 10,
+ "area": 16498,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7269,
+ "image_id": 1327,
+ "bbox": [
+ 91,
+ 301,
+ 327,
+ 130
+ ],
+ "category_id": 9,
+ "area": 120384,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7270,
+ "image_id": 1328,
+ "bbox": [
+ 154,
+ 51,
+ 278,
+ 310
+ ],
+ "category_id": 9,
+ "area": 303456,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7271,
+ "image_id": 1328,
+ "bbox": [
+ 215,
+ 0,
+ 294,
+ 388
+ ],
+ "category_id": 9,
+ "area": 401310,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7272,
+ "image_id": 1328,
+ "bbox": [
+ 206,
+ 231,
+ 303,
+ 275
+ ],
+ "category_id": 9,
+ "area": 294104,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7274,
+ "image_id": 1329,
+ "bbox": [
+ 95,
+ 169,
+ 366,
+ 341
+ ],
+ "category_id": 9,
+ "area": 162534,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7291,
+ "image_id": 1331,
+ "bbox": [
+ 2,
+ 3,
+ 42,
+ 74
+ ],
+ "category_id": 6,
+ "area": 61776,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7292,
+ "image_id": 1331,
+ "bbox": [
+ 1,
+ 52,
+ 28,
+ 68
+ ],
+ "category_id": 6,
+ "area": 38151,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7293,
+ "image_id": 1331,
+ "bbox": [
+ 18,
+ 71,
+ 58,
+ 74
+ ],
+ "category_id": 6,
+ "area": 86920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7294,
+ "image_id": 1331,
+ "bbox": [
+ 2,
+ 122,
+ 51,
+ 93
+ ],
+ "category_id": 6,
+ "area": 95858,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7295,
+ "image_id": 1331,
+ "bbox": [
+ 28,
+ 82,
+ 52,
+ 96
+ ],
+ "category_id": 6,
+ "area": 101775,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7296,
+ "image_id": 1331,
+ "bbox": [
+ 1,
+ 243,
+ 34,
+ 88
+ ],
+ "category_id": 6,
+ "area": 60356,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7297,
+ "image_id": 1331,
+ "bbox": [
+ 52,
+ 2,
+ 68,
+ 52
+ ],
+ "category_id": 6,
+ "area": 70680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7298,
+ "image_id": 1331,
+ "bbox": [
+ 86,
+ 2,
+ 162,
+ 129
+ ],
+ "category_id": 6,
+ "area": 419941,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7299,
+ "image_id": 1331,
+ "bbox": [
+ 103,
+ 176,
+ 46,
+ 45
+ ],
+ "category_id": 6,
+ "area": 41377,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7300,
+ "image_id": 1331,
+ "bbox": [
+ 3,
+ 380,
+ 47,
+ 76
+ ],
+ "category_id": 6,
+ "area": 72336,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7301,
+ "image_id": 1331,
+ "bbox": [
+ 70,
+ 337,
+ 77,
+ 134
+ ],
+ "category_id": 6,
+ "area": 206974,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7302,
+ "image_id": 1331,
+ "bbox": [
+ 198,
+ 1,
+ 49,
+ 37
+ ],
+ "category_id": 6,
+ "area": 37252,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7303,
+ "image_id": 1331,
+ "bbox": [
+ 253,
+ 2,
+ 72,
+ 77
+ ],
+ "category_id": 6,
+ "area": 112590,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7304,
+ "image_id": 1331,
+ "bbox": [
+ 214,
+ 70,
+ 83,
+ 151
+ ],
+ "category_id": 6,
+ "area": 251640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7305,
+ "image_id": 1331,
+ "bbox": [
+ 220,
+ 216,
+ 95,
+ 98
+ ],
+ "category_id": 6,
+ "area": 186030,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7318,
+ "image_id": 1332,
+ "bbox": [
+ 209,
+ 230,
+ 70,
+ 158
+ ],
+ "category_id": 8,
+ "area": 39294,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7319,
+ "image_id": 1333,
+ "bbox": [
+ 0,
+ 209,
+ 496,
+ 301
+ ],
+ "category_id": 9,
+ "area": 102820,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7356,
+ "image_id": 1337,
+ "bbox": [
+ 253,
+ 147,
+ 110,
+ 189
+ ],
+ "category_id": 8,
+ "area": 73416,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7362,
+ "image_id": 1337,
+ "bbox": [
+ 46,
+ 137,
+ 95,
+ 89
+ ],
+ "category_id": 6,
+ "area": 29988,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7363,
+ "image_id": 1337,
+ "bbox": [
+ 50,
+ 192,
+ 102,
+ 111
+ ],
+ "category_id": 6,
+ "area": 40349,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7364,
+ "image_id": 1337,
+ "bbox": [
+ 180,
+ 284,
+ 39,
+ 46
+ ],
+ "category_id": 6,
+ "area": 6468,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7365,
+ "image_id": 1337,
+ "bbox": [
+ 167,
+ 187,
+ 33,
+ 49
+ ],
+ "category_id": 6,
+ "area": 5810,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7366,
+ "image_id": 1337,
+ "bbox": [
+ 61,
+ 302,
+ 52,
+ 68
+ ],
+ "category_id": 6,
+ "area": 12707,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7367,
+ "image_id": 1337,
+ "bbox": [
+ 264,
+ 331,
+ 90,
+ 100
+ ],
+ "category_id": 6,
+ "area": 32234,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7368,
+ "image_id": 1337,
+ "bbox": [
+ 446,
+ 286,
+ 65,
+ 105
+ ],
+ "category_id": 6,
+ "area": 24124,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7369,
+ "image_id": 1337,
+ "bbox": [
+ 368,
+ 334,
+ 142,
+ 177
+ ],
+ "category_id": 6,
+ "area": 89250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7370,
+ "image_id": 1337,
+ "bbox": [
+ 272,
+ 373,
+ 146,
+ 138
+ ],
+ "category_id": 6,
+ "area": 71175,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7371,
+ "image_id": 1337,
+ "bbox": [
+ 158,
+ 403,
+ 151,
+ 108
+ ],
+ "category_id": 6,
+ "area": 57456,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7373,
+ "image_id": 1338,
+ "bbox": [
+ 306,
+ 282,
+ 74,
+ 72
+ ],
+ "category_id": 6,
+ "area": 5568,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7374,
+ "image_id": 1338,
+ "bbox": [
+ 218,
+ 278,
+ 54,
+ 48
+ ],
+ "category_id": 6,
+ "area": 2752,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7375,
+ "image_id": 1338,
+ "bbox": [
+ 457,
+ 463,
+ 54,
+ 48
+ ],
+ "category_id": 6,
+ "area": 2752,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7376,
+ "image_id": 1338,
+ "bbox": [
+ 401,
+ 449,
+ 52,
+ 37
+ ],
+ "category_id": 6,
+ "area": 2013,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7377,
+ "image_id": 1338,
+ "bbox": [
+ 407,
+ 480,
+ 40,
+ 29
+ ],
+ "category_id": 6,
+ "area": 1222,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7378,
+ "image_id": 1338,
+ "bbox": [
+ 0,
+ 244,
+ 50,
+ 75
+ ],
+ "category_id": 6,
+ "area": 3894,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7379,
+ "image_id": 1338,
+ "bbox": [
+ 458,
+ 407,
+ 53,
+ 56
+ ],
+ "category_id": 6,
+ "area": 3150,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7380,
+ "image_id": 1338,
+ "bbox": [
+ 85,
+ 401,
+ 74,
+ 60
+ ],
+ "category_id": 6,
+ "area": 4611,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7381,
+ "image_id": 1338,
+ "bbox": [
+ 442,
+ 336,
+ 53,
+ 56
+ ],
+ "category_id": 6,
+ "area": 3150,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7382,
+ "image_id": 1338,
+ "bbox": [
+ 387,
+ 381,
+ 40,
+ 58
+ ],
+ "category_id": 6,
+ "area": 2448,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7383,
+ "image_id": 1338,
+ "bbox": [
+ 347,
+ 374,
+ 49,
+ 56
+ ],
+ "category_id": 6,
+ "area": 2900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7384,
+ "image_id": 1338,
+ "bbox": [
+ 273,
+ 325,
+ 72,
+ 85
+ ],
+ "category_id": 6,
+ "area": 6375,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7387,
+ "image_id": 1340,
+ "bbox": [
+ 66,
+ 289,
+ 109,
+ 65
+ ],
+ "category_id": 6,
+ "area": 18161,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7388,
+ "image_id": 1340,
+ "bbox": [
+ 39,
+ 164,
+ 392,
+ 266
+ ],
+ "category_id": 9,
+ "area": 265720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7389,
+ "image_id": 1341,
+ "bbox": [
+ 0,
+ 0,
+ 511,
+ 509
+ ],
+ "category_id": 6,
+ "area": 8255538,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7390,
+ "image_id": 1342,
+ "bbox": [
+ 132,
+ 407,
+ 181,
+ 102
+ ],
+ "category_id": 6,
+ "area": 65232,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7391,
+ "image_id": 1342,
+ "bbox": [
+ 357,
+ 228,
+ 154,
+ 283
+ ],
+ "category_id": 6,
+ "area": 154026,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7392,
+ "image_id": 1342,
+ "bbox": [
+ 108,
+ 233,
+ 257,
+ 197
+ ],
+ "category_id": 6,
+ "area": 178754,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7393,
+ "image_id": 1342,
+ "bbox": [
+ 0,
+ 216,
+ 147,
+ 256
+ ],
+ "category_id": 8,
+ "area": 132840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7402,
+ "image_id": 1344,
+ "bbox": [
+ 221,
+ 0,
+ 202,
+ 299
+ ],
+ "category_id": 6,
+ "area": 81983,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7403,
+ "image_id": 1344,
+ "bbox": [
+ 136,
+ 15,
+ 56,
+ 86
+ ],
+ "category_id": 6,
+ "area": 6586,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7404,
+ "image_id": 1345,
+ "bbox": [
+ 349,
+ 141,
+ 55,
+ 114
+ ],
+ "category_id": 6,
+ "area": 9555,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7405,
+ "image_id": 1345,
+ "bbox": [
+ 414,
+ 118,
+ 96,
+ 316
+ ],
+ "category_id": 6,
+ "area": 46110,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7406,
+ "image_id": 1345,
+ "bbox": [
+ 429,
+ 93,
+ 81,
+ 110
+ ],
+ "category_id": 6,
+ "area": 13635,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7407,
+ "image_id": 1345,
+ "bbox": [
+ 289,
+ 421,
+ 46,
+ 56
+ ],
+ "category_id": 6,
+ "area": 3952,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7408,
+ "image_id": 1345,
+ "bbox": [
+ 343,
+ 480,
+ 43,
+ 30
+ ],
+ "category_id": 6,
+ "area": 2016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7413,
+ "image_id": 1347,
+ "bbox": [
+ 169,
+ 151,
+ 148,
+ 236
+ ],
+ "category_id": 9,
+ "area": 123172,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7414,
+ "image_id": 1347,
+ "bbox": [
+ 2,
+ 366,
+ 28,
+ 83
+ ],
+ "category_id": 9,
+ "area": 8496,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7416,
+ "image_id": 1348,
+ "bbox": [
+ 60,
+ 208,
+ 149,
+ 183
+ ],
+ "category_id": 9,
+ "area": 40992,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7417,
+ "image_id": 1348,
+ "bbox": [
+ 22,
+ 226,
+ 51,
+ 84
+ ],
+ "category_id": 6,
+ "area": 6552,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7420,
+ "image_id": 1350,
+ "bbox": [
+ 1,
+ 199,
+ 388,
+ 307
+ ],
+ "category_id": 9,
+ "area": 419040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7421,
+ "image_id": 1350,
+ "bbox": [
+ 2,
+ 12,
+ 414,
+ 424
+ ],
+ "category_id": 9,
+ "area": 618492,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7422,
+ "image_id": 1350,
+ "bbox": [
+ 0,
+ 177,
+ 421,
+ 318
+ ],
+ "category_id": 9,
+ "area": 471744,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7431,
+ "image_id": 1352,
+ "bbox": [
+ 120,
+ 254,
+ 40,
+ 65
+ ],
+ "category_id": 10,
+ "area": 1480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7432,
+ "image_id": 1352,
+ "bbox": [
+ 49,
+ 225,
+ 63,
+ 104
+ ],
+ "category_id": 10,
+ "area": 3717,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7433,
+ "image_id": 1352,
+ "bbox": [
+ 9,
+ 394,
+ 38,
+ 69
+ ],
+ "category_id": 10,
+ "area": 1482,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7434,
+ "image_id": 1352,
+ "bbox": [
+ 252,
+ 202,
+ 32,
+ 58
+ ],
+ "category_id": 10,
+ "area": 1056,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7435,
+ "image_id": 1352,
+ "bbox": [
+ 282,
+ 0,
+ 137,
+ 195
+ ],
+ "category_id": 10,
+ "area": 15070,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7436,
+ "image_id": 1352,
+ "bbox": [
+ 124,
+ 0,
+ 219,
+ 115
+ ],
+ "category_id": 10,
+ "area": 14235,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7437,
+ "image_id": 1352,
+ "bbox": [
+ 401,
+ 341,
+ 27,
+ 46
+ ],
+ "category_id": 10,
+ "area": 702,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7438,
+ "image_id": 1352,
+ "bbox": [
+ 411,
+ 272,
+ 52,
+ 81
+ ],
+ "category_id": 10,
+ "area": 2392,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7439,
+ "image_id": 1352,
+ "bbox": [
+ 423,
+ 209,
+ 45,
+ 69
+ ],
+ "category_id": 10,
+ "area": 1755,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7440,
+ "image_id": 1352,
+ "bbox": [
+ 488,
+ 280,
+ 24,
+ 60
+ ],
+ "category_id": 10,
+ "area": 816,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7441,
+ "image_id": 1352,
+ "bbox": [
+ 276,
+ 465,
+ 25,
+ 42
+ ],
+ "category_id": 10,
+ "area": 600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7442,
+ "image_id": 1352,
+ "bbox": [
+ 74,
+ 394,
+ 29,
+ 48
+ ],
+ "category_id": 10,
+ "area": 783,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7443,
+ "image_id": 1352,
+ "bbox": [
+ 287,
+ 188,
+ 14,
+ 23
+ ],
+ "category_id": 10,
+ "area": 182,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7445,
+ "image_id": 1352,
+ "bbox": [
+ 422,
+ 145,
+ 23,
+ 42
+ ],
+ "category_id": 10,
+ "area": 552,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7446,
+ "image_id": 1352,
+ "bbox": [
+ 459,
+ 389,
+ 24,
+ 37
+ ],
+ "category_id": 10,
+ "area": 504,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7483,
+ "image_id": 1355,
+ "bbox": [
+ 0,
+ 60,
+ 435,
+ 408
+ ],
+ "category_id": 9,
+ "area": 288744,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7484,
+ "image_id": 1356,
+ "bbox": [
+ 68,
+ 110,
+ 439,
+ 398
+ ],
+ "category_id": 9,
+ "area": 614880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7487,
+ "image_id": 1357,
+ "bbox": [
+ 6,
+ 289,
+ 131,
+ 221
+ ],
+ "category_id": 8,
+ "area": 38480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7494,
+ "image_id": 1359,
+ "bbox": [
+ 271,
+ 90,
+ 32,
+ 44
+ ],
+ "category_id": 6,
+ "area": 9400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7495,
+ "image_id": 1359,
+ "bbox": [
+ 334,
+ 224,
+ 29,
+ 37
+ ],
+ "category_id": 6,
+ "area": 6942,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7496,
+ "image_id": 1359,
+ "bbox": [
+ 437,
+ 175,
+ 33,
+ 45
+ ],
+ "category_id": 6,
+ "area": 9785,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7685,
+ "image_id": 1360,
+ "bbox": [
+ 290,
+ 0,
+ 221,
+ 508
+ ],
+ "category_id": 6,
+ "area": 407540,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7686,
+ "image_id": 1360,
+ "bbox": [
+ 3,
+ 0,
+ 289,
+ 509
+ ],
+ "category_id": 6,
+ "area": 531330,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7692,
+ "image_id": 1361,
+ "bbox": [
+ 171,
+ 112,
+ 278,
+ 259
+ ],
+ "category_id": 8,
+ "area": 116580,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7693,
+ "image_id": 1361,
+ "bbox": [
+ 104,
+ 142,
+ 223,
+ 172
+ ],
+ "category_id": 8,
+ "area": 62122,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7694,
+ "image_id": 1362,
+ "bbox": [
+ 12,
+ 69,
+ 186,
+ 149
+ ],
+ "category_id": 9,
+ "area": 30464,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7695,
+ "image_id": 1362,
+ "bbox": [
+ 197,
+ 192,
+ 305,
+ 151
+ ],
+ "category_id": 9,
+ "area": 50784,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7718,
+ "image_id": 1366,
+ "bbox": [
+ 108,
+ 140,
+ 87,
+ 65
+ ],
+ "category_id": 8,
+ "area": 38794,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7719,
+ "image_id": 1366,
+ "bbox": [
+ 123,
+ 83,
+ 49,
+ 67
+ ],
+ "category_id": 8,
+ "area": 22755,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7720,
+ "image_id": 1366,
+ "bbox": [
+ 206,
+ 78,
+ 24,
+ 36
+ ],
+ "category_id": 8,
+ "area": 6164,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7721,
+ "image_id": 1366,
+ "bbox": [
+ 237,
+ 201,
+ 55,
+ 123
+ ],
+ "category_id": 8,
+ "area": 46800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7722,
+ "image_id": 1366,
+ "bbox": [
+ 321,
+ 116,
+ 48,
+ 81
+ ],
+ "category_id": 8,
+ "area": 26640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7726,
+ "image_id": 1368,
+ "bbox": [
+ 239,
+ 162,
+ 169,
+ 344
+ ],
+ "category_id": 9,
+ "area": 205155,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7727,
+ "image_id": 1368,
+ "bbox": [
+ 12,
+ 123,
+ 252,
+ 320
+ ],
+ "category_id": 9,
+ "area": 284130,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7728,
+ "image_id": 1368,
+ "bbox": [
+ 344,
+ 39,
+ 164,
+ 96
+ ],
+ "category_id": 9,
+ "area": 55760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7729,
+ "image_id": 1368,
+ "bbox": [
+ 326,
+ 115,
+ 183,
+ 390
+ ],
+ "category_id": 9,
+ "area": 251442,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7737,
+ "image_id": 1370,
+ "bbox": [
+ 7,
+ 41,
+ 287,
+ 466
+ ],
+ "category_id": 9,
+ "area": 116220,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7738,
+ "image_id": 1370,
+ "bbox": [
+ 208,
+ 9,
+ 295,
+ 308
+ ],
+ "category_id": 9,
+ "area": 78948,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7739,
+ "image_id": 1371,
+ "bbox": [
+ 161,
+ 197,
+ 309,
+ 216
+ ],
+ "category_id": 9,
+ "area": 309890,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7740,
+ "image_id": 1372,
+ "bbox": [
+ 13,
+ 279,
+ 435,
+ 173
+ ],
+ "category_id": 9,
+ "area": 48025,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7763,
+ "image_id": 1377,
+ "bbox": [
+ 0,
+ 159,
+ 61,
+ 124
+ ],
+ "category_id": 8,
+ "area": 60784,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7764,
+ "image_id": 1377,
+ "bbox": [
+ 181,
+ 148,
+ 140,
+ 162
+ ],
+ "category_id": 8,
+ "area": 179550,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7765,
+ "image_id": 1377,
+ "bbox": [
+ 310,
+ 156,
+ 142,
+ 198
+ ],
+ "category_id": 8,
+ "area": 223746,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7766,
+ "image_id": 1377,
+ "bbox": [
+ 140,
+ 219,
+ 98,
+ 73
+ ],
+ "category_id": 8,
+ "area": 57040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7777,
+ "image_id": 1380,
+ "bbox": [
+ 65,
+ 0,
+ 328,
+ 512
+ ],
+ "category_id": 9,
+ "area": 460284,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7778,
+ "image_id": 1380,
+ "bbox": [
+ 323,
+ 0,
+ 188,
+ 244
+ ],
+ "category_id": 6,
+ "area": 126048,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7795,
+ "image_id": 1383,
+ "bbox": [
+ 81,
+ 262,
+ 140,
+ 237
+ ],
+ "category_id": 8,
+ "area": 116200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7796,
+ "image_id": 1383,
+ "bbox": [
+ 385,
+ 268,
+ 126,
+ 152
+ ],
+ "category_id": 8,
+ "area": 67095,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7799,
+ "image_id": 1384,
+ "bbox": [
+ 255,
+ 165,
+ 108,
+ 126
+ ],
+ "category_id": 9,
+ "area": 48238,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7800,
+ "image_id": 1384,
+ "bbox": [
+ 250,
+ 256,
+ 112,
+ 145
+ ],
+ "category_id": 9,
+ "area": 57528,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7801,
+ "image_id": 1384,
+ "bbox": [
+ 6,
+ 191,
+ 337,
+ 192
+ ],
+ "category_id": 9,
+ "area": 228453,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7802,
+ "image_id": 1385,
+ "bbox": [
+ 127,
+ 425,
+ 214,
+ 84
+ ],
+ "category_id": 10,
+ "area": 28684,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7808,
+ "image_id": 1387,
+ "bbox": [
+ 0,
+ 0,
+ 512,
+ 252
+ ],
+ "category_id": 6,
+ "area": 320128,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7810,
+ "image_id": 1389,
+ "bbox": [
+ 26,
+ 124,
+ 408,
+ 290
+ ],
+ "category_id": 8,
+ "area": 416976,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7811,
+ "image_id": 1389,
+ "bbox": [
+ 130,
+ 6,
+ 70,
+ 93
+ ],
+ "category_id": 6,
+ "area": 22925,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7812,
+ "image_id": 1389,
+ "bbox": [
+ 196,
+ 0,
+ 173,
+ 130
+ ],
+ "category_id": 6,
+ "area": 79672,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7813,
+ "image_id": 1389,
+ "bbox": [
+ 232,
+ 161,
+ 137,
+ 350
+ ],
+ "category_id": 6,
+ "area": 169099,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7814,
+ "image_id": 1389,
+ "bbox": [
+ 441,
+ 161,
+ 36,
+ 62
+ ],
+ "category_id": 6,
+ "area": 7920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7815,
+ "image_id": 1389,
+ "bbox": [
+ 454,
+ 216,
+ 56,
+ 54
+ ],
+ "category_id": 6,
+ "area": 10640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7828,
+ "image_id": 1391,
+ "bbox": [
+ 65,
+ 19,
+ 60,
+ 83
+ ],
+ "category_id": 9,
+ "area": 7546,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7830,
+ "image_id": 1392,
+ "bbox": [
+ 1,
+ 105,
+ 317,
+ 162
+ ],
+ "category_id": 9,
+ "area": 87084,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7831,
+ "image_id": 1392,
+ "bbox": [
+ 0,
+ 190,
+ 199,
+ 320
+ ],
+ "category_id": 6,
+ "area": 107880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7832,
+ "image_id": 1393,
+ "bbox": [
+ 0,
+ 91,
+ 363,
+ 419
+ ],
+ "category_id": 9,
+ "area": 435390,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7841,
+ "image_id": 1396,
+ "bbox": [
+ 183,
+ 233,
+ 93,
+ 103
+ ],
+ "category_id": 8,
+ "area": 30989,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7842,
+ "image_id": 1396,
+ "bbox": [
+ 236,
+ 293,
+ 91,
+ 71
+ ],
+ "category_id": 8,
+ "area": 20884,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7845,
+ "image_id": 1397,
+ "bbox": [
+ 211,
+ 101,
+ 293,
+ 251
+ ],
+ "category_id": 8,
+ "area": 97468,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7849,
+ "image_id": 1398,
+ "bbox": [
+ 181,
+ 119,
+ 76,
+ 63
+ ],
+ "category_id": 10,
+ "area": 13965,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7850,
+ "image_id": 1398,
+ "bbox": [
+ 28,
+ 116,
+ 336,
+ 394
+ ],
+ "category_id": 9,
+ "area": 379600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7867,
+ "image_id": 1401,
+ "bbox": [
+ 190,
+ 313,
+ 77,
+ 120
+ ],
+ "category_id": 6,
+ "area": 55626,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7900,
+ "image_id": 1403,
+ "bbox": [
+ 122,
+ 31,
+ 304,
+ 439
+ ],
+ "category_id": 9,
+ "area": 241332,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7901,
+ "image_id": 1403,
+ "bbox": [
+ 35,
+ 99,
+ 280,
+ 298
+ ],
+ "category_id": 9,
+ "area": 150672,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7902,
+ "image_id": 1404,
+ "bbox": [
+ 233,
+ 239,
+ 92,
+ 272
+ ],
+ "category_id": 8,
+ "area": 95436,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7903,
+ "image_id": 1404,
+ "bbox": [
+ 233,
+ 112,
+ 76,
+ 209
+ ],
+ "category_id": 8,
+ "area": 60390,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7904,
+ "image_id": 1404,
+ "bbox": [
+ 319,
+ 137,
+ 70,
+ 181
+ ],
+ "category_id": 8,
+ "area": 48048,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7916,
+ "image_id": 1406,
+ "bbox": [
+ 98,
+ 126,
+ 377,
+ 277
+ ],
+ "category_id": 9,
+ "area": 266057,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7929,
+ "image_id": 1409,
+ "bbox": [
+ 120,
+ 44,
+ 204,
+ 368
+ ],
+ "category_id": 9,
+ "area": 157248,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7930,
+ "image_id": 1409,
+ "bbox": [
+ 140,
+ 55,
+ 298,
+ 447
+ ],
+ "category_id": 9,
+ "area": 278810,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7947,
+ "image_id": 1414,
+ "bbox": [
+ 178,
+ 267,
+ 102,
+ 67
+ ],
+ "category_id": 8,
+ "area": 24415,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7948,
+ "image_id": 1414,
+ "bbox": [
+ 279,
+ 184,
+ 103,
+ 149
+ ],
+ "category_id": 8,
+ "area": 54390,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7949,
+ "image_id": 1415,
+ "bbox": [
+ 51,
+ 2,
+ 117,
+ 80
+ ],
+ "category_id": 6,
+ "area": 96924,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7950,
+ "image_id": 1415,
+ "bbox": [
+ 5,
+ 69,
+ 118,
+ 181
+ ],
+ "category_id": 6,
+ "area": 219938,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7951,
+ "image_id": 1415,
+ "bbox": [
+ 4,
+ 243,
+ 147,
+ 113
+ ],
+ "category_id": 6,
+ "area": 171418,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7952,
+ "image_id": 1415,
+ "bbox": [
+ 4,
+ 361,
+ 84,
+ 68
+ ],
+ "category_id": 6,
+ "area": 59924,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7953,
+ "image_id": 1415,
+ "bbox": [
+ 141,
+ 4,
+ 141,
+ 201
+ ],
+ "category_id": 6,
+ "area": 290752,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7954,
+ "image_id": 1415,
+ "bbox": [
+ 141,
+ 194,
+ 116,
+ 88
+ ],
+ "category_id": 6,
+ "area": 105961,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7955,
+ "image_id": 1415,
+ "bbox": [
+ 169,
+ 275,
+ 178,
+ 94
+ ],
+ "category_id": 6,
+ "area": 172822,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7956,
+ "image_id": 1415,
+ "bbox": [
+ 282,
+ 16,
+ 42,
+ 88
+ ],
+ "category_id": 6,
+ "area": 38211,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7957,
+ "image_id": 1415,
+ "bbox": [
+ 302,
+ 73,
+ 92,
+ 72
+ ],
+ "category_id": 6,
+ "area": 68068,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7958,
+ "image_id": 1415,
+ "bbox": [
+ 395,
+ 66,
+ 115,
+ 105
+ ],
+ "category_id": 6,
+ "area": 124678,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7959,
+ "image_id": 1415,
+ "bbox": [
+ 448,
+ 283,
+ 63,
+ 68
+ ],
+ "category_id": 6,
+ "area": 44730,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7966,
+ "image_id": 1417,
+ "bbox": [
+ 184,
+ 202,
+ 129,
+ 118
+ ],
+ "category_id": 8,
+ "area": 39435,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7967,
+ "image_id": 1417,
+ "bbox": [
+ 219,
+ 333,
+ 97,
+ 105
+ ],
+ "category_id": 8,
+ "area": 26460,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7972,
+ "image_id": 1418,
+ "bbox": [
+ 2,
+ 172,
+ 404,
+ 334
+ ],
+ "category_id": 9,
+ "area": 475710,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7973,
+ "image_id": 1418,
+ "bbox": [
+ 161,
+ 0,
+ 348,
+ 179
+ ],
+ "category_id": 9,
+ "area": 219744,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 7974,
+ "image_id": 1418,
+ "bbox": [
+ 171,
+ 102,
+ 340,
+ 307
+ ],
+ "category_id": 9,
+ "area": 368916,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8020,
+ "image_id": 1423,
+ "bbox": [
+ 38,
+ 23,
+ 260,
+ 335
+ ],
+ "category_id": 9,
+ "area": 175205,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8043,
+ "image_id": 1426,
+ "bbox": [
+ 55,
+ 0,
+ 372,
+ 428
+ ],
+ "category_id": 9,
+ "area": 304848,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8045,
+ "image_id": 1427,
+ "bbox": [
+ 152,
+ 35,
+ 272,
+ 278
+ ],
+ "category_id": 9,
+ "area": 191950,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8046,
+ "image_id": 1427,
+ "bbox": [
+ 160,
+ 153,
+ 220,
+ 341
+ ],
+ "category_id": 10,
+ "area": 190442,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8050,
+ "image_id": 1429,
+ "bbox": [
+ 0,
+ 0,
+ 511,
+ 163
+ ],
+ "category_id": 6,
+ "area": 210231,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8060,
+ "image_id": 1431,
+ "bbox": [
+ 0,
+ 260,
+ 247,
+ 249
+ ],
+ "category_id": 9,
+ "area": 93432,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8067,
+ "image_id": 1432,
+ "bbox": [
+ 46,
+ 135,
+ 96,
+ 92
+ ],
+ "category_id": 6,
+ "area": 31460,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8068,
+ "image_id": 1432,
+ "bbox": [
+ 50,
+ 190,
+ 100,
+ 113
+ ],
+ "category_id": 6,
+ "area": 40320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8069,
+ "image_id": 1432,
+ "bbox": [
+ 60,
+ 305,
+ 54,
+ 73
+ ],
+ "category_id": 6,
+ "area": 14040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8070,
+ "image_id": 1432,
+ "bbox": [
+ 250,
+ 145,
+ 114,
+ 189
+ ],
+ "category_id": 8,
+ "area": 75810,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8074,
+ "image_id": 1432,
+ "bbox": [
+ 446,
+ 281,
+ 65,
+ 99
+ ],
+ "category_id": 6,
+ "area": 22820,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8075,
+ "image_id": 1432,
+ "bbox": [
+ 370,
+ 337,
+ 141,
+ 173
+ ],
+ "category_id": 6,
+ "area": 86132,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8076,
+ "image_id": 1432,
+ "bbox": [
+ 284,
+ 373,
+ 134,
+ 138
+ ],
+ "category_id": 6,
+ "area": 65715,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8077,
+ "image_id": 1432,
+ "bbox": [
+ 156,
+ 406,
+ 151,
+ 105
+ ],
+ "category_id": 6,
+ "area": 56322,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8078,
+ "image_id": 1432,
+ "bbox": [
+ 182,
+ 284,
+ 38,
+ 50
+ ],
+ "category_id": 6,
+ "area": 6887,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8079,
+ "image_id": 1433,
+ "bbox": [
+ 0,
+ 123,
+ 280,
+ 164
+ ],
+ "category_id": 9,
+ "area": 77865,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8080,
+ "image_id": 1433,
+ "bbox": [
+ 0,
+ 204,
+ 183,
+ 305
+ ],
+ "category_id": 6,
+ "area": 94620,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8082,
+ "image_id": 1434,
+ "bbox": [
+ 170,
+ 269,
+ 66,
+ 61
+ ],
+ "category_id": 8,
+ "area": 32500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8083,
+ "image_id": 1435,
+ "bbox": [
+ 79,
+ 93,
+ 54,
+ 43
+ ],
+ "category_id": 8,
+ "area": 16116,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8084,
+ "image_id": 1435,
+ "bbox": [
+ 105,
+ 159,
+ 95,
+ 55
+ ],
+ "category_id": 8,
+ "area": 36259,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8085,
+ "image_id": 1435,
+ "bbox": [
+ 202,
+ 83,
+ 36,
+ 41
+ ],
+ "category_id": 8,
+ "area": 10275,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8086,
+ "image_id": 1435,
+ "bbox": [
+ 215,
+ 172,
+ 51,
+ 139
+ ],
+ "category_id": 8,
+ "area": 49082,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8087,
+ "image_id": 1435,
+ "bbox": [
+ 311,
+ 131,
+ 59,
+ 80
+ ],
+ "category_id": 8,
+ "area": 32634,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8092,
+ "image_id": 1437,
+ "bbox": [
+ 208,
+ 289,
+ 100,
+ 53
+ ],
+ "category_id": 8,
+ "area": 42714,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8093,
+ "image_id": 1437,
+ "bbox": [
+ 200,
+ 352,
+ 166,
+ 91
+ ],
+ "category_id": 8,
+ "area": 120862,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8094,
+ "image_id": 1437,
+ "bbox": [
+ 28,
+ 332,
+ 185,
+ 159
+ ],
+ "category_id": 8,
+ "area": 233878,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8095,
+ "image_id": 1437,
+ "bbox": [
+ 485,
+ 210,
+ 25,
+ 73
+ ],
+ "category_id": 6,
+ "area": 14476,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8110,
+ "image_id": 1440,
+ "bbox": [
+ 0,
+ 277,
+ 66,
+ 234
+ ],
+ "category_id": 8,
+ "area": 107226,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8111,
+ "image_id": 1440,
+ "bbox": [
+ 67,
+ 185,
+ 351,
+ 259
+ ],
+ "category_id": 8,
+ "area": 628518,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8112,
+ "image_id": 1440,
+ "bbox": [
+ 205,
+ 174,
+ 208,
+ 173
+ ],
+ "category_id": 8,
+ "area": 250242,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8114,
+ "image_id": 1441,
+ "bbox": [
+ 246,
+ 23,
+ 265,
+ 298
+ ],
+ "category_id": 6,
+ "area": 356025,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8115,
+ "image_id": 1441,
+ "bbox": [
+ 72,
+ 188,
+ 125,
+ 306
+ ],
+ "category_id": 8,
+ "area": 172161,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8118,
+ "image_id": 1443,
+ "bbox": [
+ 49,
+ 346,
+ 151,
+ 154
+ ],
+ "category_id": 8,
+ "area": 46610,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8119,
+ "image_id": 1443,
+ "bbox": [
+ 208,
+ 282,
+ 241,
+ 228
+ ],
+ "category_id": 8,
+ "area": 109277,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8121,
+ "image_id": 1444,
+ "bbox": [
+ 250,
+ 270,
+ 125,
+ 103
+ ],
+ "category_id": 8,
+ "area": 41184,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8122,
+ "image_id": 1444,
+ "bbox": [
+ 360,
+ 223,
+ 62,
+ 90
+ ],
+ "category_id": 8,
+ "area": 17825,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8123,
+ "image_id": 1445,
+ "bbox": [
+ 10,
+ 200,
+ 165,
+ 158
+ ],
+ "category_id": 8,
+ "area": 92322,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8124,
+ "image_id": 1445,
+ "bbox": [
+ 80,
+ 20,
+ 365,
+ 262
+ ],
+ "category_id": 8,
+ "area": 335984,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8125,
+ "image_id": 1446,
+ "bbox": [
+ 176,
+ 262,
+ 136,
+ 147
+ ],
+ "category_id": 8,
+ "area": 70246,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8126,
+ "image_id": 1446,
+ "bbox": [
+ 262,
+ 321,
+ 126,
+ 102
+ ],
+ "category_id": 8,
+ "area": 45360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8129,
+ "image_id": 1447,
+ "bbox": [
+ 41,
+ 144,
+ 274,
+ 233
+ ],
+ "category_id": 6,
+ "area": 159432,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8136,
+ "image_id": 1448,
+ "bbox": [
+ 40,
+ 310,
+ 50,
+ 58
+ ],
+ "category_id": 6,
+ "area": 3465,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8137,
+ "image_id": 1448,
+ "bbox": [
+ 234,
+ 358,
+ 51,
+ 69
+ ],
+ "category_id": 6,
+ "area": 4160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8140,
+ "image_id": 1449,
+ "bbox": [
+ 124,
+ 139,
+ 240,
+ 367
+ ],
+ "category_id": 9,
+ "area": 63215,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8141,
+ "image_id": 1450,
+ "bbox": [
+ 2,
+ 23,
+ 506,
+ 483
+ ],
+ "category_id": 9,
+ "area": 860200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8149,
+ "image_id": 1453,
+ "bbox": [
+ 91,
+ 302,
+ 139,
+ 125
+ ],
+ "category_id": 8,
+ "area": 55173,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8150,
+ "image_id": 1453,
+ "bbox": [
+ 281,
+ 207,
+ 109,
+ 72
+ ],
+ "category_id": 8,
+ "area": 25116,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8151,
+ "image_id": 1454,
+ "bbox": [
+ 51,
+ 108,
+ 373,
+ 273
+ ],
+ "category_id": 8,
+ "area": 149504,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8152,
+ "image_id": 1454,
+ "bbox": [
+ 42,
+ 242,
+ 348,
+ 243
+ ],
+ "category_id": 8,
+ "area": 124260,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8153,
+ "image_id": 1455,
+ "bbox": [
+ 81,
+ 168,
+ 430,
+ 341
+ ],
+ "category_id": 9,
+ "area": 420561,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8157,
+ "image_id": 1456,
+ "bbox": [
+ 117,
+ 76,
+ 312,
+ 430
+ ],
+ "category_id": 9,
+ "area": 473110,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8158,
+ "image_id": 1456,
+ "bbox": [
+ 251,
+ 0,
+ 257,
+ 425
+ ],
+ "category_id": 9,
+ "area": 385157,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8199,
+ "image_id": 1461,
+ "bbox": [
+ 67,
+ 0,
+ 336,
+ 512
+ ],
+ "category_id": 9,
+ "area": 471062,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8200,
+ "image_id": 1461,
+ "bbox": [
+ 323,
+ 0,
+ 188,
+ 205
+ ],
+ "category_id": 6,
+ "area": 105918,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8201,
+ "image_id": 1462,
+ "bbox": [
+ 146,
+ 272,
+ 154,
+ 164
+ ],
+ "category_id": 8,
+ "area": 89397,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8202,
+ "image_id": 1462,
+ "bbox": [
+ 238,
+ 323,
+ 144,
+ 132
+ ],
+ "category_id": 8,
+ "area": 67332,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8206,
+ "image_id": 1463,
+ "bbox": [
+ 1,
+ 112,
+ 366,
+ 399
+ ],
+ "category_id": 8,
+ "area": 512754,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8207,
+ "image_id": 1463,
+ "bbox": [
+ 170,
+ 36,
+ 251,
+ 370
+ ],
+ "category_id": 8,
+ "area": 327080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8208,
+ "image_id": 1464,
+ "bbox": [
+ 71,
+ 150,
+ 304,
+ 145
+ ],
+ "category_id": 8,
+ "area": 144599,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8209,
+ "image_id": 1464,
+ "bbox": [
+ 24,
+ 215,
+ 111,
+ 108
+ ],
+ "category_id": 6,
+ "area": 39440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8210,
+ "image_id": 1464,
+ "bbox": [
+ 0,
+ 330,
+ 103,
+ 181
+ ],
+ "category_id": 6,
+ "area": 61128,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8211,
+ "image_id": 1464,
+ "bbox": [
+ 100,
+ 331,
+ 185,
+ 179
+ ],
+ "category_id": 6,
+ "area": 109028,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8212,
+ "image_id": 1464,
+ "bbox": [
+ 276,
+ 106,
+ 233,
+ 402
+ ],
+ "category_id": 6,
+ "area": 306323,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8213,
+ "image_id": 1464,
+ "bbox": [
+ 153,
+ 280,
+ 61,
+ 55
+ ],
+ "category_id": 6,
+ "area": 11008,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8214,
+ "image_id": 1465,
+ "bbox": [
+ 77,
+ 192,
+ 168,
+ 115
+ ],
+ "category_id": 8,
+ "area": 153576,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8215,
+ "image_id": 1465,
+ "bbox": [
+ 333,
+ 219,
+ 159,
+ 106
+ ],
+ "category_id": 8,
+ "area": 133952,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8216,
+ "image_id": 1465,
+ "bbox": [
+ 145,
+ 313,
+ 37,
+ 55
+ ],
+ "category_id": 6,
+ "area": 16520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8217,
+ "image_id": 1466,
+ "bbox": [
+ 66,
+ 232,
+ 67,
+ 55
+ ],
+ "category_id": 8,
+ "area": 29972,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8218,
+ "image_id": 1466,
+ "bbox": [
+ 119,
+ 239,
+ 196,
+ 271
+ ],
+ "category_id": 8,
+ "area": 421564,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8219,
+ "image_id": 1466,
+ "bbox": [
+ 83,
+ 463,
+ 169,
+ 47
+ ],
+ "category_id": 8,
+ "area": 63700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8228,
+ "image_id": 1468,
+ "bbox": [
+ 0,
+ 200,
+ 427,
+ 239
+ ],
+ "category_id": 8,
+ "area": 196877,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8229,
+ "image_id": 1468,
+ "bbox": [
+ 252,
+ 109,
+ 205,
+ 103
+ ],
+ "category_id": 8,
+ "area": 40905,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8236,
+ "image_id": 1470,
+ "bbox": [
+ 197,
+ 282,
+ 98,
+ 100
+ ],
+ "category_id": 9,
+ "area": 23520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8253,
+ "image_id": 1472,
+ "bbox": [
+ 15,
+ 156,
+ 358,
+ 355
+ ],
+ "category_id": 9,
+ "area": 363285,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8255,
+ "image_id": 1472,
+ "bbox": [
+ 181,
+ 166,
+ 34,
+ 52
+ ],
+ "category_id": 10,
+ "area": 5074,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8261,
+ "image_id": 1474,
+ "bbox": [
+ 192,
+ 119,
+ 100,
+ 200
+ ],
+ "category_id": 8,
+ "area": 44548,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8276,
+ "image_id": 1478,
+ "bbox": [
+ 3,
+ 1,
+ 507,
+ 510
+ ],
+ "category_id": 6,
+ "area": 555282,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8277,
+ "image_id": 1478,
+ "bbox": [
+ 25,
+ 272,
+ 443,
+ 240
+ ],
+ "category_id": 9,
+ "area": 228360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8278,
+ "image_id": 1479,
+ "bbox": [
+ 19,
+ 135,
+ 274,
+ 171
+ ],
+ "category_id": 9,
+ "area": 79050,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8279,
+ "image_id": 1479,
+ "bbox": [
+ 0,
+ 212,
+ 197,
+ 219
+ ],
+ "category_id": 6,
+ "area": 72828,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8280,
+ "image_id": 1480,
+ "bbox": [
+ 161,
+ 134,
+ 336,
+ 373
+ ],
+ "category_id": 9,
+ "area": 441525,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8281,
+ "image_id": 1480,
+ "bbox": [
+ 292,
+ 0,
+ 218,
+ 508
+ ],
+ "category_id": 9,
+ "area": 389675,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8282,
+ "image_id": 1481,
+ "bbox": [
+ 0,
+ 0,
+ 177,
+ 193
+ ],
+ "category_id": 6,
+ "area": 79464,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8283,
+ "image_id": 1481,
+ "bbox": [
+ 280,
+ 197,
+ 231,
+ 254
+ ],
+ "category_id": 9,
+ "area": 136278,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8284,
+ "image_id": 1482,
+ "bbox": [
+ 132,
+ 152,
+ 178,
+ 322
+ ],
+ "category_id": 9,
+ "area": 84100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8285,
+ "image_id": 1482,
+ "bbox": [
+ 344,
+ 149,
+ 94,
+ 119
+ ],
+ "category_id": 6,
+ "area": 16478,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8286,
+ "image_id": 1483,
+ "bbox": [
+ 121,
+ 67,
+ 125,
+ 154
+ ],
+ "category_id": 10,
+ "area": 22165,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8287,
+ "image_id": 1483,
+ "bbox": [
+ 225,
+ 201,
+ 206,
+ 305
+ ],
+ "category_id": 9,
+ "area": 72448,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8339,
+ "image_id": 1486,
+ "bbox": [
+ 146,
+ 47,
+ 263,
+ 346
+ ],
+ "category_id": 9,
+ "area": 64250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8340,
+ "image_id": 1486,
+ "bbox": [
+ 26,
+ 190,
+ 375,
+ 292
+ ],
+ "category_id": 9,
+ "area": 77437,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8341,
+ "image_id": 1487,
+ "bbox": [
+ 168,
+ 47,
+ 332,
+ 464
+ ],
+ "category_id": 9,
+ "area": 282336,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8343,
+ "image_id": 1488,
+ "bbox": [
+ 67,
+ 280,
+ 360,
+ 134
+ ],
+ "category_id": 9,
+ "area": 145529,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8345,
+ "image_id": 1488,
+ "bbox": [
+ 179,
+ 30,
+ 139,
+ 240
+ ],
+ "category_id": 6,
+ "area": 100529,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8346,
+ "image_id": 1488,
+ "bbox": [
+ 0,
+ 0,
+ 163,
+ 134
+ ],
+ "category_id": 6,
+ "area": 65905,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8347,
+ "image_id": 1488,
+ "bbox": [
+ 0,
+ 204,
+ 85,
+ 168
+ ],
+ "category_id": 6,
+ "area": 43008,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8363,
+ "image_id": 1490,
+ "bbox": [
+ 2,
+ 5,
+ 509,
+ 320
+ ],
+ "category_id": 8,
+ "area": 1292512,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8364,
+ "image_id": 1490,
+ "bbox": [
+ 4,
+ 283,
+ 347,
+ 184
+ ],
+ "category_id": 8,
+ "area": 508170,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8365,
+ "image_id": 1491,
+ "bbox": [
+ 0,
+ 107,
+ 59,
+ 262
+ ],
+ "category_id": 6,
+ "area": 44992,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8366,
+ "image_id": 1491,
+ "bbox": [
+ 60,
+ 21,
+ 267,
+ 395
+ ],
+ "category_id": 6,
+ "area": 303905,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8367,
+ "image_id": 1491,
+ "bbox": [
+ 299,
+ 4,
+ 208,
+ 346
+ ],
+ "category_id": 6,
+ "area": 208520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8381,
+ "image_id": 1495,
+ "bbox": [
+ 131,
+ 185,
+ 152,
+ 279
+ ],
+ "category_id": 9,
+ "area": 289656,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8382,
+ "image_id": 1496,
+ "bbox": [
+ 94,
+ 163,
+ 82,
+ 154
+ ],
+ "category_id": 10,
+ "area": 46948,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8383,
+ "image_id": 1496,
+ "bbox": [
+ 153,
+ 117,
+ 160,
+ 171
+ ],
+ "category_id": 9,
+ "area": 100500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8384,
+ "image_id": 1496,
+ "bbox": [
+ 229,
+ 0,
+ 282,
+ 511
+ ],
+ "category_id": 9,
+ "area": 528938,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8385,
+ "image_id": 1497,
+ "bbox": [
+ 325,
+ 163,
+ 80,
+ 75
+ ],
+ "category_id": 10,
+ "area": 11583,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8386,
+ "image_id": 1497,
+ "bbox": [
+ 302,
+ 105,
+ 138,
+ 56
+ ],
+ "category_id": 9,
+ "area": 14948,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8389,
+ "image_id": 1498,
+ "bbox": [
+ 0,
+ 67,
+ 158,
+ 83
+ ],
+ "category_id": 8,
+ "area": 104368,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8390,
+ "image_id": 1498,
+ "bbox": [
+ 206,
+ 127,
+ 109,
+ 117
+ ],
+ "category_id": 8,
+ "area": 101764,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8391,
+ "image_id": 1498,
+ "bbox": [
+ 126,
+ 12,
+ 84,
+ 119
+ ],
+ "category_id": 8,
+ "area": 79948,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8392,
+ "image_id": 1498,
+ "bbox": [
+ 41,
+ 64,
+ 85,
+ 43
+ ],
+ "category_id": 8,
+ "area": 29532,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8394,
+ "image_id": 1499,
+ "bbox": [
+ 129,
+ 59,
+ 382,
+ 451
+ ],
+ "category_id": 10,
+ "area": 1365649,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8397,
+ "image_id": 1500,
+ "bbox": [
+ 201,
+ 78,
+ 93,
+ 206
+ ],
+ "category_id": 10,
+ "area": 30070,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8410,
+ "image_id": 1502,
+ "bbox": [
+ 0,
+ 159,
+ 357,
+ 223
+ ],
+ "category_id": 9,
+ "area": 165690,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8427,
+ "image_id": 1504,
+ "bbox": [
+ 156,
+ 97,
+ 180,
+ 322
+ ],
+ "category_id": 8,
+ "area": 202499,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8438,
+ "image_id": 1506,
+ "bbox": [
+ 185,
+ 57,
+ 121,
+ 298
+ ],
+ "category_id": 10,
+ "area": 46926,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8443,
+ "image_id": 1507,
+ "bbox": [
+ 166,
+ 62,
+ 266,
+ 447
+ ],
+ "category_id": 8,
+ "area": 3774222,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8448,
+ "image_id": 1509,
+ "bbox": [
+ 58,
+ 80,
+ 272,
+ 94
+ ],
+ "category_id": 8,
+ "area": 47212,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8449,
+ "image_id": 1509,
+ "bbox": [
+ 0,
+ 257,
+ 224,
+ 192
+ ],
+ "category_id": 6,
+ "area": 79163,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8452,
+ "image_id": 1510,
+ "bbox": [
+ 155,
+ 171,
+ 86,
+ 114
+ ],
+ "category_id": 8,
+ "area": 34560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8453,
+ "image_id": 1510,
+ "bbox": [
+ 338,
+ 149,
+ 115,
+ 126
+ ],
+ "category_id": 8,
+ "area": 50976,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8454,
+ "image_id": 1510,
+ "bbox": [
+ 433,
+ 277,
+ 30,
+ 32
+ ],
+ "category_id": 6,
+ "area": 3542,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8455,
+ "image_id": 1510,
+ "bbox": [
+ 472,
+ 322,
+ 36,
+ 47
+ ],
+ "category_id": 6,
+ "area": 6072,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8457,
+ "image_id": 1511,
+ "bbox": [
+ 352,
+ 69,
+ 92,
+ 103
+ ],
+ "category_id": 8,
+ "area": 75646,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8458,
+ "image_id": 1511,
+ "bbox": [
+ 250,
+ 139,
+ 73,
+ 121
+ ],
+ "category_id": 8,
+ "area": 70932,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8459,
+ "image_id": 1511,
+ "bbox": [
+ 30,
+ 152,
+ 91,
+ 41
+ ],
+ "category_id": 8,
+ "area": 29841,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8462,
+ "image_id": 1514,
+ "bbox": [
+ 64,
+ 114,
+ 420,
+ 389
+ ],
+ "category_id": 6,
+ "area": 235092,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8463,
+ "image_id": 1514,
+ "bbox": [
+ 344,
+ 34,
+ 61,
+ 64
+ ],
+ "category_id": 6,
+ "area": 5640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8464,
+ "image_id": 1514,
+ "bbox": [
+ 408,
+ 20,
+ 98,
+ 128
+ ],
+ "category_id": 6,
+ "area": 18048,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8465,
+ "image_id": 1515,
+ "bbox": [
+ 209,
+ 209,
+ 158,
+ 208
+ ],
+ "category_id": 6,
+ "area": 47736,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8466,
+ "image_id": 1515,
+ "bbox": [
+ 223,
+ 402,
+ 205,
+ 109
+ ],
+ "category_id": 6,
+ "area": 32421,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8502,
+ "image_id": 1518,
+ "bbox": [
+ 0,
+ 111,
+ 176,
+ 162
+ ],
+ "category_id": 6,
+ "area": 87462,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8503,
+ "image_id": 1518,
+ "bbox": [
+ 33,
+ 243,
+ 227,
+ 268
+ ],
+ "category_id": 6,
+ "area": 186252,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8504,
+ "image_id": 1518,
+ "bbox": [
+ 322,
+ 338,
+ 88,
+ 156
+ ],
+ "category_id": 6,
+ "area": 42510,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8505,
+ "image_id": 1518,
+ "bbox": [
+ 328,
+ 318,
+ 99,
+ 94
+ ],
+ "category_id": 6,
+ "area": 28908,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8506,
+ "image_id": 1518,
+ "bbox": [
+ 381,
+ 308,
+ 63,
+ 58
+ ],
+ "category_id": 6,
+ "area": 11480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8507,
+ "image_id": 1518,
+ "bbox": [
+ 426,
+ 322,
+ 84,
+ 187
+ ],
+ "category_id": 6,
+ "area": 48470,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8508,
+ "image_id": 1518,
+ "bbox": [
+ 175,
+ 152,
+ 61,
+ 58
+ ],
+ "category_id": 6,
+ "area": 10988,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8509,
+ "image_id": 1518,
+ "bbox": [
+ 114,
+ 48,
+ 397,
+ 462
+ ],
+ "category_id": 9,
+ "area": 561795,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8512,
+ "image_id": 1520,
+ "bbox": [
+ 6,
+ 53,
+ 486,
+ 458
+ ],
+ "category_id": 9,
+ "area": 177080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8550,
+ "image_id": 1524,
+ "bbox": [
+ 274,
+ 146,
+ 37,
+ 93
+ ],
+ "category_id": 10,
+ "area": 8296,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8551,
+ "image_id": 1524,
+ "bbox": [
+ 200,
+ 143,
+ 221,
+ 368
+ ],
+ "category_id": 9,
+ "area": 193516,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8559,
+ "image_id": 1526,
+ "bbox": [
+ 104,
+ 135,
+ 233,
+ 168
+ ],
+ "category_id": 10,
+ "area": 62109,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8584,
+ "image_id": 1529,
+ "bbox": [
+ 1,
+ 0,
+ 510,
+ 280
+ ],
+ "category_id": 6,
+ "area": 368520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8585,
+ "image_id": 1530,
+ "bbox": [
+ 212,
+ 170,
+ 187,
+ 190
+ ],
+ "category_id": 8,
+ "area": 41886,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8591,
+ "image_id": 1532,
+ "bbox": [
+ 174,
+ 152,
+ 167,
+ 192
+ ],
+ "category_id": 9,
+ "area": 180677,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8592,
+ "image_id": 1532,
+ "bbox": [
+ 47,
+ 416,
+ 112,
+ 94
+ ],
+ "category_id": 9,
+ "area": 60207,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8608,
+ "image_id": 1535,
+ "bbox": [
+ 141,
+ 40,
+ 284,
+ 443
+ ],
+ "category_id": 9,
+ "area": 438075,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8623,
+ "image_id": 1539,
+ "bbox": [
+ 111,
+ 53,
+ 275,
+ 384
+ ],
+ "category_id": 9,
+ "area": 263400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8627,
+ "image_id": 1540,
+ "bbox": [
+ 212,
+ 136,
+ 40,
+ 142
+ ],
+ "category_id": 10,
+ "area": 13505,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8628,
+ "image_id": 1540,
+ "bbox": [
+ 206,
+ 158,
+ 294,
+ 285
+ ],
+ "category_id": 9,
+ "area": 199764,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8645,
+ "image_id": 1542,
+ "bbox": [
+ 136,
+ 0,
+ 372,
+ 506
+ ],
+ "category_id": 9,
+ "area": 663584,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8647,
+ "image_id": 1543,
+ "bbox": [
+ 264,
+ 143,
+ 157,
+ 176
+ ],
+ "category_id": 8,
+ "area": 32340,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8668,
+ "image_id": 1546,
+ "bbox": [
+ 228,
+ 54,
+ 227,
+ 306
+ ],
+ "category_id": 8,
+ "area": 89856,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8669,
+ "image_id": 1546,
+ "bbox": [
+ 113,
+ 118,
+ 190,
+ 239
+ ],
+ "category_id": 8,
+ "area": 58950,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8672,
+ "image_id": 1547,
+ "bbox": [
+ 90,
+ 381,
+ 118,
+ 87
+ ],
+ "category_id": 9,
+ "area": 15440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8676,
+ "image_id": 1548,
+ "bbox": [
+ 21,
+ 405,
+ 109,
+ 99
+ ],
+ "category_id": 6,
+ "area": 30368,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8677,
+ "image_id": 1548,
+ "bbox": [
+ 101,
+ 404,
+ 131,
+ 107
+ ],
+ "category_id": 6,
+ "area": 39342,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8678,
+ "image_id": 1549,
+ "bbox": [
+ 198,
+ 167,
+ 247,
+ 333
+ ],
+ "category_id": 9,
+ "area": 289842,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8679,
+ "image_id": 1549,
+ "bbox": [
+ 240,
+ 86,
+ 227,
+ 384
+ ],
+ "category_id": 9,
+ "area": 307829,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8692,
+ "image_id": 1551,
+ "bbox": [
+ 66,
+ 26,
+ 82,
+ 233
+ ],
+ "category_id": 9,
+ "area": 67896,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8693,
+ "image_id": 1551,
+ "bbox": [
+ 202,
+ 139,
+ 167,
+ 307
+ ],
+ "category_id": 9,
+ "area": 180994,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8694,
+ "image_id": 1551,
+ "bbox": [
+ 246,
+ 12,
+ 136,
+ 258
+ ],
+ "category_id": 9,
+ "area": 124488,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8695,
+ "image_id": 1551,
+ "bbox": [
+ 204,
+ 2,
+ 148,
+ 299
+ ],
+ "category_id": 9,
+ "area": 156191,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8696,
+ "image_id": 1551,
+ "bbox": [
+ 355,
+ 263,
+ 102,
+ 226
+ ],
+ "category_id": 9,
+ "area": 81090,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8698,
+ "image_id": 1552,
+ "bbox": [
+ 269,
+ 192,
+ 241,
+ 284
+ ],
+ "category_id": 8,
+ "area": 80066,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8699,
+ "image_id": 1553,
+ "bbox": [
+ 0,
+ 0,
+ 512,
+ 468
+ ],
+ "category_id": 9,
+ "area": 1267500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8703,
+ "image_id": 1555,
+ "bbox": [
+ 269,
+ 103,
+ 92,
+ 79
+ ],
+ "category_id": 9,
+ "area": 25872,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8704,
+ "image_id": 1555,
+ "bbox": [
+ 258,
+ 123,
+ 130,
+ 147
+ ],
+ "category_id": 9,
+ "area": 67482,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8708,
+ "image_id": 1557,
+ "bbox": [
+ 9,
+ 57,
+ 414,
+ 364
+ ],
+ "category_id": 8,
+ "area": 530944,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8709,
+ "image_id": 1557,
+ "bbox": [
+ 377,
+ 107,
+ 134,
+ 295
+ ],
+ "category_id": 8,
+ "area": 139360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8710,
+ "image_id": 1558,
+ "bbox": [
+ 73,
+ 129,
+ 251,
+ 214
+ ],
+ "category_id": 8,
+ "area": 98686,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8714,
+ "image_id": 1559,
+ "bbox": [
+ 209,
+ 204,
+ 253,
+ 305
+ ],
+ "category_id": 10,
+ "area": 183997,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8715,
+ "image_id": 1559,
+ "bbox": [
+ 91,
+ 159,
+ 222,
+ 133
+ ],
+ "category_id": 9,
+ "area": 70550,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8719,
+ "image_id": 1560,
+ "bbox": [
+ 168,
+ 362,
+ 19,
+ 39
+ ],
+ "category_id": 6,
+ "area": 5248,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8720,
+ "image_id": 1560,
+ "bbox": [
+ 60,
+ 264,
+ 38,
+ 55
+ ],
+ "category_id": 6,
+ "area": 14478,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8721,
+ "image_id": 1561,
+ "bbox": [
+ 7,
+ 50,
+ 504,
+ 460
+ ],
+ "category_id": 10,
+ "area": 533520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8724,
+ "image_id": 1562,
+ "bbox": [
+ 300,
+ 160,
+ 210,
+ 282
+ ],
+ "category_id": 8,
+ "area": 69168,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8727,
+ "image_id": 1564,
+ "bbox": [
+ 103,
+ 207,
+ 292,
+ 269
+ ],
+ "category_id": 9,
+ "area": 139836,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8731,
+ "image_id": 1565,
+ "bbox": [
+ 60,
+ 210,
+ 237,
+ 279
+ ],
+ "category_id": 9,
+ "area": 2524464,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8732,
+ "image_id": 1565,
+ "bbox": [
+ 212,
+ 222,
+ 169,
+ 141
+ ],
+ "category_id": 6,
+ "area": 910078,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8733,
+ "image_id": 1565,
+ "bbox": [
+ 0,
+ 163,
+ 44,
+ 64
+ ],
+ "category_id": 6,
+ "area": 107388,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8734,
+ "image_id": 1565,
+ "bbox": [
+ 33,
+ 216,
+ 52,
+ 81
+ ],
+ "category_id": 6,
+ "area": 162564,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8735,
+ "image_id": 1566,
+ "bbox": [
+ 237,
+ 160,
+ 59,
+ 136
+ ],
+ "category_id": 10,
+ "area": 13530,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8736,
+ "image_id": 1566,
+ "bbox": [
+ 314,
+ 128,
+ 71,
+ 155
+ ],
+ "category_id": 10,
+ "area": 18480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8737,
+ "image_id": 1566,
+ "bbox": [
+ 12,
+ 316,
+ 101,
+ 181
+ ],
+ "category_id": 10,
+ "area": 30668,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8738,
+ "image_id": 1566,
+ "bbox": [
+ 172,
+ 287,
+ 116,
+ 149
+ ],
+ "category_id": 10,
+ "area": 29025,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8739,
+ "image_id": 1566,
+ "bbox": [
+ 311,
+ 299,
+ 114,
+ 178
+ ],
+ "category_id": 10,
+ "area": 34132,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8740,
+ "image_id": 1566,
+ "bbox": [
+ 35,
+ 24,
+ 111,
+ 156
+ ],
+ "category_id": 10,
+ "area": 28905,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8741,
+ "image_id": 1566,
+ "bbox": [
+ 380,
+ 29,
+ 120,
+ 260
+ ],
+ "category_id": 10,
+ "area": 52405,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8747,
+ "image_id": 1569,
+ "bbox": [
+ 16,
+ 108,
+ 321,
+ 397
+ ],
+ "category_id": 8,
+ "area": 448877,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8748,
+ "image_id": 1569,
+ "bbox": [
+ 85,
+ 0,
+ 408,
+ 467
+ ],
+ "category_id": 8,
+ "area": 671160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8758,
+ "image_id": 1571,
+ "bbox": [
+ 200,
+ 0,
+ 311,
+ 244
+ ],
+ "category_id": 9,
+ "area": 176904,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8759,
+ "image_id": 1571,
+ "bbox": [
+ 304,
+ 138,
+ 171,
+ 188
+ ],
+ "category_id": 10,
+ "area": 74640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8765,
+ "image_id": 1573,
+ "bbox": [
+ 72,
+ 96,
+ 134,
+ 259
+ ],
+ "category_id": 9,
+ "area": 123005,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8766,
+ "image_id": 1573,
+ "bbox": [
+ 161,
+ 329,
+ 126,
+ 171
+ ],
+ "category_id": 9,
+ "area": 76397,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8767,
+ "image_id": 1573,
+ "bbox": [
+ 374,
+ 380,
+ 108,
+ 125
+ ],
+ "category_id": 9,
+ "area": 47967,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8768,
+ "image_id": 1573,
+ "bbox": [
+ 202,
+ 216,
+ 100,
+ 181
+ ],
+ "category_id": 9,
+ "area": 63750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8769,
+ "image_id": 1573,
+ "bbox": [
+ 252,
+ 93,
+ 129,
+ 153
+ ],
+ "category_id": 9,
+ "area": 69768,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8770,
+ "image_id": 1573,
+ "bbox": [
+ 278,
+ 130,
+ 91,
+ 160
+ ],
+ "category_id": 9,
+ "area": 51525,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8774,
+ "image_id": 1574,
+ "bbox": [
+ 0,
+ 0,
+ 511,
+ 265
+ ],
+ "category_id": 6,
+ "area": 328536,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8775,
+ "image_id": 1575,
+ "bbox": [
+ 122,
+ 233,
+ 142,
+ 197
+ ],
+ "category_id": 8,
+ "area": 98256,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8776,
+ "image_id": 1575,
+ "bbox": [
+ 220,
+ 237,
+ 116,
+ 191
+ ],
+ "category_id": 8,
+ "area": 78256,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8777,
+ "image_id": 1575,
+ "bbox": [
+ 283,
+ 347,
+ 122,
+ 163
+ ],
+ "category_id": 8,
+ "area": 70074,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8778,
+ "image_id": 1576,
+ "bbox": [
+ 158,
+ 254,
+ 84,
+ 93
+ ],
+ "category_id": 8,
+ "area": 27772,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8779,
+ "image_id": 1576,
+ "bbox": [
+ 226,
+ 269,
+ 93,
+ 74
+ ],
+ "category_id": 8,
+ "area": 24465,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8802,
+ "image_id": 1578,
+ "bbox": [
+ 72,
+ 251,
+ 303,
+ 210
+ ],
+ "category_id": 9,
+ "area": 107859,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8803,
+ "image_id": 1578,
+ "bbox": [
+ 0,
+ 306,
+ 252,
+ 205
+ ],
+ "category_id": 6,
+ "area": 87416,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8811,
+ "image_id": 1580,
+ "bbox": [
+ 83,
+ 270,
+ 167,
+ 113
+ ],
+ "category_id": 8,
+ "area": 48396,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8813,
+ "image_id": 1580,
+ "bbox": [
+ 185,
+ 380,
+ 83,
+ 82
+ ],
+ "category_id": 6,
+ "area": 17441,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8814,
+ "image_id": 1581,
+ "bbox": [
+ 6,
+ 321,
+ 60,
+ 96
+ ],
+ "category_id": 9,
+ "area": 20385,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8815,
+ "image_id": 1581,
+ "bbox": [
+ 204,
+ 153,
+ 124,
+ 135
+ ],
+ "category_id": 9,
+ "area": 59210,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8818,
+ "image_id": 1582,
+ "bbox": [
+ 25,
+ 203,
+ 210,
+ 262
+ ],
+ "category_id": 8,
+ "area": 193200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8819,
+ "image_id": 1582,
+ "bbox": [
+ 168,
+ 227,
+ 338,
+ 216
+ ],
+ "category_id": 8,
+ "area": 256035,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8834,
+ "image_id": 1584,
+ "bbox": [
+ 2,
+ 104,
+ 508,
+ 404
+ ],
+ "category_id": 9,
+ "area": 723199,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8835,
+ "image_id": 1584,
+ "bbox": [
+ 202,
+ 2,
+ 309,
+ 504
+ ],
+ "category_id": 9,
+ "area": 548830,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8836,
+ "image_id": 1585,
+ "bbox": [
+ 1,
+ 0,
+ 510,
+ 331
+ ],
+ "category_id": 8,
+ "area": 1305756,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8837,
+ "image_id": 1585,
+ "bbox": [
+ 148,
+ 110,
+ 62,
+ 135
+ ],
+ "category_id": 8,
+ "area": 65286,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8838,
+ "image_id": 1585,
+ "bbox": [
+ 8,
+ 289,
+ 348,
+ 187
+ ],
+ "category_id": 8,
+ "area": 503344,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8839,
+ "image_id": 1586,
+ "bbox": [
+ 166,
+ 232,
+ 189,
+ 182
+ ],
+ "category_id": 8,
+ "area": 487272,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8846,
+ "image_id": 1589,
+ "bbox": [
+ 127,
+ 128,
+ 125,
+ 100
+ ],
+ "category_id": 8,
+ "area": 44133,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8847,
+ "image_id": 1589,
+ "bbox": [
+ 335,
+ 190,
+ 138,
+ 166
+ ],
+ "category_id": 8,
+ "area": 80851,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8849,
+ "image_id": 1590,
+ "bbox": [
+ 172,
+ 266,
+ 302,
+ 180
+ ],
+ "category_id": 6,
+ "area": 226500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8850,
+ "image_id": 1591,
+ "bbox": [
+ 153,
+ 236,
+ 144,
+ 252
+ ],
+ "category_id": 10,
+ "area": 47235,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8851,
+ "image_id": 1591,
+ "bbox": [
+ 113,
+ 28,
+ 316,
+ 324
+ ],
+ "category_id": 9,
+ "area": 132870,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8853,
+ "image_id": 1592,
+ "bbox": [
+ 148,
+ 180,
+ 129,
+ 125
+ ],
+ "category_id": 8,
+ "area": 56848,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8854,
+ "image_id": 1592,
+ "bbox": [
+ 325,
+ 155,
+ 177,
+ 259
+ ],
+ "category_id": 8,
+ "area": 161252,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8855,
+ "image_id": 1593,
+ "bbox": [
+ 230,
+ 1,
+ 133,
+ 176
+ ],
+ "category_id": 10,
+ "area": 67716,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8857,
+ "image_id": 1594,
+ "bbox": [
+ 120,
+ 153,
+ 306,
+ 233
+ ],
+ "category_id": 9,
+ "area": 251685,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8859,
+ "image_id": 1595,
+ "bbox": [
+ 192,
+ 245,
+ 318,
+ 264
+ ],
+ "category_id": 10,
+ "area": 149017,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8860,
+ "image_id": 1596,
+ "bbox": [
+ 73,
+ 75,
+ 328,
+ 189
+ ],
+ "category_id": 8,
+ "area": 106020,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8882,
+ "image_id": 1598,
+ "bbox": [
+ 266,
+ 182,
+ 129,
+ 156
+ ],
+ "category_id": 8,
+ "area": 160380,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8883,
+ "image_id": 1598,
+ "bbox": [
+ 137,
+ 215,
+ 156,
+ 132
+ ],
+ "category_id": 8,
+ "area": 164052,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8884,
+ "image_id": 1598,
+ "bbox": [
+ 140,
+ 270,
+ 121,
+ 116
+ ],
+ "category_id": 8,
+ "area": 111230,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8886,
+ "image_id": 1599,
+ "bbox": [
+ 214,
+ 202,
+ 190,
+ 181
+ ],
+ "category_id": 8,
+ "area": 120904,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8887,
+ "image_id": 1599,
+ "bbox": [
+ 140,
+ 214,
+ 76,
+ 149
+ ],
+ "category_id": 8,
+ "area": 39919,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8888,
+ "image_id": 1600,
+ "bbox": [
+ 113,
+ 91,
+ 316,
+ 289
+ ],
+ "category_id": 8,
+ "area": 171765,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8889,
+ "image_id": 1600,
+ "bbox": [
+ 308,
+ 239,
+ 202,
+ 268
+ ],
+ "category_id": 6,
+ "area": 101752,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8890,
+ "image_id": 1600,
+ "bbox": [
+ 167,
+ 377,
+ 150,
+ 130
+ ],
+ "category_id": 6,
+ "area": 36660,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8891,
+ "image_id": 1600,
+ "bbox": [
+ 3,
+ 426,
+ 159,
+ 84
+ ],
+ "category_id": 6,
+ "area": 25149,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8892,
+ "image_id": 1600,
+ "bbox": [
+ 92,
+ 346,
+ 94,
+ 75
+ ],
+ "category_id": 6,
+ "area": 13320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8893,
+ "image_id": 1600,
+ "bbox": [
+ 0,
+ 278,
+ 55,
+ 119
+ ],
+ "category_id": 6,
+ "area": 12298,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8909,
+ "image_id": 1602,
+ "bbox": [
+ 0,
+ 271,
+ 131,
+ 96
+ ],
+ "category_id": 6,
+ "area": 97812,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8910,
+ "image_id": 1602,
+ "bbox": [
+ 128,
+ 429,
+ 36,
+ 54
+ ],
+ "category_id": 6,
+ "area": 15368,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8911,
+ "image_id": 1602,
+ "bbox": [
+ 37,
+ 119,
+ 33,
+ 61
+ ],
+ "category_id": 6,
+ "area": 16002,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8912,
+ "image_id": 1602,
+ "bbox": [
+ 100,
+ 154,
+ 38,
+ 64
+ ],
+ "category_id": 6,
+ "area": 19285,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8913,
+ "image_id": 1602,
+ "bbox": [
+ 120,
+ 113,
+ 38,
+ 104
+ ],
+ "category_id": 6,
+ "area": 31390,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8914,
+ "image_id": 1602,
+ "bbox": [
+ 158,
+ 115,
+ 30,
+ 66
+ ],
+ "category_id": 6,
+ "area": 15732,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8915,
+ "image_id": 1602,
+ "bbox": [
+ 68,
+ 38,
+ 31,
+ 52
+ ],
+ "category_id": 6,
+ "area": 12971,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8916,
+ "image_id": 1602,
+ "bbox": [
+ 89,
+ 66,
+ 39,
+ 74
+ ],
+ "category_id": 6,
+ "area": 22644,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8917,
+ "image_id": 1602,
+ "bbox": [
+ 288,
+ 99,
+ 25,
+ 58
+ ],
+ "category_id": 6,
+ "area": 11374,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8918,
+ "image_id": 1602,
+ "bbox": [
+ 285,
+ 161,
+ 53,
+ 76
+ ],
+ "category_id": 6,
+ "area": 31442,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8919,
+ "image_id": 1602,
+ "bbox": [
+ 404,
+ 103,
+ 43,
+ 82
+ ],
+ "category_id": 6,
+ "area": 27540,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8920,
+ "image_id": 1602,
+ "bbox": [
+ 409,
+ 0,
+ 22,
+ 37
+ ],
+ "category_id": 6,
+ "area": 6545,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8921,
+ "image_id": 1602,
+ "bbox": [
+ 313,
+ 112,
+ 25,
+ 48
+ ],
+ "category_id": 6,
+ "area": 9306,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8922,
+ "image_id": 1602,
+ "bbox": [
+ 387,
+ 43,
+ 13,
+ 52
+ ],
+ "category_id": 6,
+ "area": 5616,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8923,
+ "image_id": 1602,
+ "bbox": [
+ 343,
+ 54,
+ 20,
+ 21
+ ],
+ "category_id": 6,
+ "area": 3344,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8928,
+ "image_id": 1606,
+ "bbox": [
+ 243,
+ 300,
+ 263,
+ 210
+ ],
+ "category_id": 8,
+ "area": 194405,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8929,
+ "image_id": 1606,
+ "bbox": [
+ 95,
+ 32,
+ 155,
+ 231
+ ],
+ "category_id": 8,
+ "area": 126100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8930,
+ "image_id": 1606,
+ "bbox": [
+ 145,
+ 140,
+ 237,
+ 311
+ ],
+ "category_id": 8,
+ "area": 260172,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8931,
+ "image_id": 1607,
+ "bbox": [
+ 142,
+ 177,
+ 300,
+ 226
+ ],
+ "category_id": 8,
+ "area": 173752,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8932,
+ "image_id": 1607,
+ "bbox": [
+ 0,
+ 101,
+ 136,
+ 229
+ ],
+ "category_id": 6,
+ "area": 79833,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8933,
+ "image_id": 1607,
+ "bbox": [
+ 42,
+ 145,
+ 201,
+ 177
+ ],
+ "category_id": 6,
+ "area": 90783,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8934,
+ "image_id": 1607,
+ "bbox": [
+ 94,
+ 294,
+ 60,
+ 144
+ ],
+ "category_id": 6,
+ "area": 22184,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8935,
+ "image_id": 1607,
+ "bbox": [
+ 36,
+ 329,
+ 69,
+ 101
+ ],
+ "category_id": 6,
+ "area": 18088,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8936,
+ "image_id": 1607,
+ "bbox": [
+ 209,
+ 68,
+ 126,
+ 130
+ ],
+ "category_id": 6,
+ "area": 42160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8937,
+ "image_id": 1607,
+ "bbox": [
+ 334,
+ 111,
+ 56,
+ 68
+ ],
+ "category_id": 6,
+ "area": 9879,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8938,
+ "image_id": 1607,
+ "bbox": [
+ 412,
+ 338,
+ 99,
+ 170
+ ],
+ "category_id": 6,
+ "area": 43262,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8939,
+ "image_id": 1607,
+ "bbox": [
+ 439,
+ 210,
+ 69,
+ 173
+ ],
+ "category_id": 6,
+ "area": 30510,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8956,
+ "image_id": 1611,
+ "bbox": [
+ 43,
+ 250,
+ 261,
+ 246
+ ],
+ "category_id": 9,
+ "area": 58637,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8961,
+ "image_id": 1613,
+ "bbox": [
+ 1,
+ 64,
+ 352,
+ 447
+ ],
+ "category_id": 9,
+ "area": 554149,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8962,
+ "image_id": 1613,
+ "bbox": [
+ 147,
+ 260,
+ 276,
+ 248
+ ],
+ "category_id": 9,
+ "area": 241159,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8970,
+ "image_id": 1615,
+ "bbox": [
+ 4,
+ 1,
+ 388,
+ 426
+ ],
+ "category_id": 10,
+ "area": 276750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8971,
+ "image_id": 1616,
+ "bbox": [
+ 5,
+ 0,
+ 506,
+ 511
+ ],
+ "category_id": 6,
+ "area": 901615,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8972,
+ "image_id": 1616,
+ "bbox": [
+ 200,
+ 78,
+ 262,
+ 378
+ ],
+ "category_id": 8,
+ "area": 345560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8973,
+ "image_id": 1616,
+ "bbox": [
+ 180,
+ 279,
+ 120,
+ 137
+ ],
+ "category_id": 8,
+ "area": 57600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8976,
+ "image_id": 1618,
+ "bbox": [
+ 383,
+ 267,
+ 66,
+ 77
+ ],
+ "category_id": 8,
+ "area": 18203,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8977,
+ "image_id": 1618,
+ "bbox": [
+ 71,
+ 148,
+ 323,
+ 254
+ ],
+ "category_id": 8,
+ "area": 288099,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8978,
+ "image_id": 1618,
+ "bbox": [
+ 60,
+ 112,
+ 292,
+ 247
+ ],
+ "category_id": 8,
+ "area": 253657,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 8999,
+ "image_id": 1621,
+ "bbox": [
+ 107,
+ 0,
+ 345,
+ 442
+ ],
+ "category_id": 6,
+ "area": 384825,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9002,
+ "image_id": 1621,
+ "bbox": [
+ 286,
+ 333,
+ 225,
+ 178
+ ],
+ "category_id": 9,
+ "area": 101336,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9003,
+ "image_id": 1621,
+ "bbox": [
+ 0,
+ 246,
+ 344,
+ 265
+ ],
+ "category_id": 6,
+ "area": 229950,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9008,
+ "image_id": 1623,
+ "bbox": [
+ 83,
+ 139,
+ 363,
+ 372
+ ],
+ "category_id": 9,
+ "area": 5148750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9009,
+ "image_id": 1623,
+ "bbox": [
+ 268,
+ 42,
+ 113,
+ 196
+ ],
+ "category_id": 10,
+ "area": 849420,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9011,
+ "image_id": 1625,
+ "bbox": [
+ 89,
+ 145,
+ 236,
+ 168
+ ],
+ "category_id": 10,
+ "area": 62913,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9013,
+ "image_id": 1626,
+ "bbox": [
+ 39,
+ 72,
+ 191,
+ 188
+ ],
+ "category_id": 9,
+ "area": 74140,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9014,
+ "image_id": 1626,
+ "bbox": [
+ 164,
+ 9,
+ 315,
+ 445
+ ],
+ "category_id": 6,
+ "area": 287526,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9018,
+ "image_id": 1628,
+ "bbox": [
+ 43,
+ 465,
+ 33,
+ 45
+ ],
+ "category_id": 10,
+ "area": 8633,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9019,
+ "image_id": 1628,
+ "bbox": [
+ 176,
+ 474,
+ 31,
+ 37
+ ],
+ "category_id": 10,
+ "area": 6643,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9020,
+ "image_id": 1628,
+ "bbox": [
+ 133,
+ 370,
+ 37,
+ 53
+ ],
+ "category_id": 10,
+ "area": 11655,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9021,
+ "image_id": 1628,
+ "bbox": [
+ 295,
+ 443,
+ 20,
+ 32
+ ],
+ "category_id": 10,
+ "area": 3840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9022,
+ "image_id": 1628,
+ "bbox": [
+ 291,
+ 477,
+ 22,
+ 33
+ ],
+ "category_id": 10,
+ "area": 4355,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9023,
+ "image_id": 1628,
+ "bbox": [
+ 446,
+ 473,
+ 32,
+ 37
+ ],
+ "category_id": 10,
+ "area": 7104,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9024,
+ "image_id": 1628,
+ "bbox": [
+ 492,
+ 420,
+ 18,
+ 46
+ ],
+ "category_id": 10,
+ "area": 4914,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9025,
+ "image_id": 1628,
+ "bbox": [
+ 467,
+ 196,
+ 30,
+ 58
+ ],
+ "category_id": 10,
+ "area": 10146,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9026,
+ "image_id": 1628,
+ "bbox": [
+ 467,
+ 322,
+ 18,
+ 31
+ ],
+ "category_id": 10,
+ "area": 3286,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9027,
+ "image_id": 1628,
+ "bbox": [
+ 405,
+ 297,
+ 17,
+ 30
+ ],
+ "category_id": 10,
+ "area": 2950,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9028,
+ "image_id": 1628,
+ "bbox": [
+ 392,
+ 220,
+ 27,
+ 46
+ ],
+ "category_id": 10,
+ "area": 7280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9029,
+ "image_id": 1628,
+ "bbox": [
+ 107,
+ 301,
+ 16,
+ 30
+ ],
+ "category_id": 10,
+ "area": 2940,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9030,
+ "image_id": 1628,
+ "bbox": [
+ 22,
+ 305,
+ 14,
+ 25
+ ],
+ "category_id": 10,
+ "area": 2107,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9031,
+ "image_id": 1628,
+ "bbox": [
+ 98,
+ 229,
+ 17,
+ 24
+ ],
+ "category_id": 10,
+ "area": 2496,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9032,
+ "image_id": 1628,
+ "bbox": [
+ 184,
+ 115,
+ 35,
+ 52
+ ],
+ "category_id": 10,
+ "area": 10815,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9050,
+ "image_id": 1631,
+ "bbox": [
+ 134,
+ 41,
+ 58,
+ 83
+ ],
+ "category_id": 6,
+ "area": 6144,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9051,
+ "image_id": 1631,
+ "bbox": [
+ 222,
+ 19,
+ 77,
+ 259
+ ],
+ "category_id": 6,
+ "area": 25200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9052,
+ "image_id": 1631,
+ "bbox": [
+ 292,
+ 0,
+ 130,
+ 479
+ ],
+ "category_id": 6,
+ "area": 78597,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9056,
+ "image_id": 1632,
+ "bbox": [
+ 189,
+ 226,
+ 43,
+ 108
+ ],
+ "category_id": 8,
+ "area": 37327,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9057,
+ "image_id": 1632,
+ "bbox": [
+ 449,
+ 277,
+ 61,
+ 41
+ ],
+ "category_id": 8,
+ "area": 20097,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9058,
+ "image_id": 1632,
+ "bbox": [
+ 363,
+ 316,
+ 92,
+ 63
+ ],
+ "category_id": 8,
+ "area": 46632,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9059,
+ "image_id": 1632,
+ "bbox": [
+ 307,
+ 217,
+ 65,
+ 39
+ ],
+ "category_id": 8,
+ "area": 20418,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9060,
+ "image_id": 1632,
+ "bbox": [
+ 258,
+ 201,
+ 32,
+ 21
+ ],
+ "category_id": 8,
+ "area": 5612,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9062,
+ "image_id": 1633,
+ "bbox": [
+ 212,
+ 296,
+ 232,
+ 174
+ ],
+ "category_id": 9,
+ "area": 53400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9066,
+ "image_id": 1635,
+ "bbox": [
+ 256,
+ 58,
+ 254,
+ 451
+ ],
+ "category_id": 6,
+ "area": 805779,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9073,
+ "image_id": 1637,
+ "bbox": [
+ 35,
+ 39,
+ 348,
+ 472
+ ],
+ "category_id": 8,
+ "area": 579215,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9088,
+ "image_id": 1639,
+ "bbox": [
+ 112,
+ 288,
+ 246,
+ 210
+ ],
+ "category_id": 8,
+ "area": 358484,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9089,
+ "image_id": 1639,
+ "bbox": [
+ 222,
+ 22,
+ 288,
+ 264
+ ],
+ "category_id": 8,
+ "area": 524960,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9090,
+ "image_id": 1640,
+ "bbox": [
+ 0,
+ 107,
+ 194,
+ 189
+ ],
+ "category_id": 8,
+ "area": 129276,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9091,
+ "image_id": 1640,
+ "bbox": [
+ 69,
+ 154,
+ 304,
+ 249
+ ],
+ "category_id": 8,
+ "area": 266000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9092,
+ "image_id": 1640,
+ "bbox": [
+ 347,
+ 291,
+ 75,
+ 91
+ ],
+ "category_id": 8,
+ "area": 24064,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9093,
+ "image_id": 1641,
+ "bbox": [
+ 108,
+ 12,
+ 403,
+ 457
+ ],
+ "category_id": 9,
+ "area": 691488,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9094,
+ "image_id": 1641,
+ "bbox": [
+ 34,
+ 115,
+ 43,
+ 95
+ ],
+ "category_id": 6,
+ "area": 15444,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9095,
+ "image_id": 1642,
+ "bbox": [
+ 0,
+ 69,
+ 310,
+ 390
+ ],
+ "category_id": 9,
+ "area": 302082,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9096,
+ "image_id": 1642,
+ "bbox": [
+ 335,
+ 139,
+ 20,
+ 88
+ ],
+ "category_id": 6,
+ "area": 4446,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9097,
+ "image_id": 1642,
+ "bbox": [
+ 453,
+ 301,
+ 58,
+ 209
+ ],
+ "category_id": 6,
+ "area": 30623,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9098,
+ "image_id": 1643,
+ "bbox": [
+ 62,
+ 0,
+ 352,
+ 512
+ ],
+ "category_id": 9,
+ "area": 493252,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9099,
+ "image_id": 1643,
+ "bbox": [
+ 330,
+ 0,
+ 181,
+ 228
+ ],
+ "category_id": 6,
+ "area": 113200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9100,
+ "image_id": 1644,
+ "bbox": [
+ 91,
+ 101,
+ 291,
+ 238
+ ],
+ "category_id": 8,
+ "area": 377460,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9101,
+ "image_id": 1644,
+ "bbox": [
+ 463,
+ 303,
+ 47,
+ 185
+ ],
+ "category_id": 6,
+ "area": 48195,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9102,
+ "image_id": 1644,
+ "bbox": [
+ 236,
+ 296,
+ 206,
+ 213
+ ],
+ "category_id": 6,
+ "area": 239644,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9103,
+ "image_id": 1644,
+ "bbox": [
+ 0,
+ 160,
+ 98,
+ 121
+ ],
+ "category_id": 6,
+ "area": 64998,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9104,
+ "image_id": 1645,
+ "bbox": [
+ 126,
+ 177,
+ 46,
+ 33
+ ],
+ "category_id": 8,
+ "area": 12250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9105,
+ "image_id": 1645,
+ "bbox": [
+ 140,
+ 210,
+ 64,
+ 61
+ ],
+ "category_id": 8,
+ "area": 31330,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9106,
+ "image_id": 1645,
+ "bbox": [
+ 192,
+ 169,
+ 62,
+ 87
+ ],
+ "category_id": 8,
+ "area": 43056,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9107,
+ "image_id": 1645,
+ "bbox": [
+ 253,
+ 205,
+ 71,
+ 83
+ ],
+ "category_id": 8,
+ "area": 47344,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9109,
+ "image_id": 1646,
+ "bbox": [
+ 38,
+ 266,
+ 243,
+ 171
+ ],
+ "category_id": 9,
+ "area": 107416,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9110,
+ "image_id": 1646,
+ "bbox": [
+ 134,
+ 310,
+ 263,
+ 201
+ ],
+ "category_id": 6,
+ "area": 136773,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9114,
+ "image_id": 1648,
+ "bbox": [
+ 183,
+ 212,
+ 124,
+ 143
+ ],
+ "category_id": 8,
+ "area": 62511,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9115,
+ "image_id": 1648,
+ "bbox": [
+ 143,
+ 268,
+ 154,
+ 107
+ ],
+ "category_id": 8,
+ "area": 58135,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9117,
+ "image_id": 1648,
+ "bbox": [
+ 0,
+ 270,
+ 312,
+ 236
+ ],
+ "category_id": 6,
+ "area": 257187,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9128,
+ "image_id": 1650,
+ "bbox": [
+ 10,
+ 51,
+ 332,
+ 460
+ ],
+ "category_id": 10,
+ "area": 1212084,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9131,
+ "image_id": 1652,
+ "bbox": [
+ 188,
+ 10,
+ 322,
+ 407
+ ],
+ "category_id": 9,
+ "area": 363800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9137,
+ "image_id": 1654,
+ "bbox": [
+ 12,
+ 2,
+ 150,
+ 309
+ ],
+ "category_id": 6,
+ "area": 489810,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9138,
+ "image_id": 1654,
+ "bbox": [
+ 263,
+ 221,
+ 88,
+ 290
+ ],
+ "category_id": 6,
+ "area": 271576,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9139,
+ "image_id": 1654,
+ "bbox": [
+ 336,
+ 50,
+ 175,
+ 427
+ ],
+ "category_id": 6,
+ "area": 790371,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9168,
+ "image_id": 1657,
+ "bbox": [
+ 25,
+ 26,
+ 310,
+ 390
+ ],
+ "category_id": 9,
+ "area": 217616,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9180,
+ "image_id": 1660,
+ "bbox": [
+ 146,
+ 178,
+ 365,
+ 239
+ ],
+ "category_id": 9,
+ "area": 147420,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9181,
+ "image_id": 1660,
+ "bbox": [
+ 117,
+ 269,
+ 393,
+ 241
+ ],
+ "category_id": 6,
+ "area": 159820,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9182,
+ "image_id": 1660,
+ "bbox": [
+ 404,
+ 96,
+ 107,
+ 163
+ ],
+ "category_id": 6,
+ "area": 29548,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9185,
+ "image_id": 1661,
+ "bbox": [
+ 214,
+ 232,
+ 147,
+ 171
+ ],
+ "category_id": 9,
+ "area": 59064,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9186,
+ "image_id": 1661,
+ "bbox": [
+ 344,
+ 184,
+ 114,
+ 179
+ ],
+ "category_id": 9,
+ "area": 47936,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9187,
+ "image_id": 1661,
+ "bbox": [
+ 313,
+ 264,
+ 116,
+ 188
+ ],
+ "category_id": 10,
+ "area": 51465,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9195,
+ "image_id": 1663,
+ "bbox": [
+ 215,
+ 146,
+ 73,
+ 139
+ ],
+ "category_id": 8,
+ "area": 81144,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9196,
+ "image_id": 1663,
+ "bbox": [
+ 273,
+ 202,
+ 74,
+ 85
+ ],
+ "category_id": 8,
+ "area": 50400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9197,
+ "image_id": 1663,
+ "bbox": [
+ 162,
+ 218,
+ 104,
+ 55
+ ],
+ "category_id": 8,
+ "area": 46256,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9198,
+ "image_id": 1663,
+ "bbox": [
+ 148,
+ 169,
+ 59,
+ 33
+ ],
+ "category_id": 8,
+ "area": 15762,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9199,
+ "image_id": 1664,
+ "bbox": [
+ 96,
+ 0,
+ 312,
+ 443
+ ],
+ "category_id": 9,
+ "area": 1095120,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9200,
+ "image_id": 1664,
+ "bbox": [
+ 38,
+ 394,
+ 197,
+ 116
+ ],
+ "category_id": 6,
+ "area": 182040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9212,
+ "image_id": 1666,
+ "bbox": [
+ 71,
+ 16,
+ 302,
+ 261
+ ],
+ "category_id": 9,
+ "area": 81396,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9213,
+ "image_id": 1666,
+ "bbox": [
+ 137,
+ 250,
+ 326,
+ 174
+ ],
+ "category_id": 6,
+ "area": 58548,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9214,
+ "image_id": 1667,
+ "bbox": [
+ 94,
+ 132,
+ 417,
+ 379
+ ],
+ "category_id": 9,
+ "area": 164304,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9216,
+ "image_id": 1668,
+ "bbox": [
+ 115,
+ 28,
+ 214,
+ 252
+ ],
+ "category_id": 9,
+ "area": 285692,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9217,
+ "image_id": 1668,
+ "bbox": [
+ 56,
+ 160,
+ 47,
+ 63
+ ],
+ "category_id": 9,
+ "area": 15827,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9218,
+ "image_id": 1669,
+ "bbox": [
+ 91,
+ 130,
+ 82,
+ 53
+ ],
+ "category_id": 8,
+ "area": 29876,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9219,
+ "image_id": 1669,
+ "bbox": [
+ 176,
+ 109,
+ 52,
+ 56
+ ],
+ "category_id": 8,
+ "area": 20188,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9220,
+ "image_id": 1669,
+ "bbox": [
+ 200,
+ 77,
+ 26,
+ 37
+ ],
+ "category_id": 8,
+ "area": 6762,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9221,
+ "image_id": 1669,
+ "bbox": [
+ 295,
+ 88,
+ 53,
+ 66
+ ],
+ "category_id": 8,
+ "area": 24120,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9222,
+ "image_id": 1669,
+ "bbox": [
+ 249,
+ 202,
+ 43,
+ 99
+ ],
+ "category_id": 8,
+ "area": 29684,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9223,
+ "image_id": 1670,
+ "bbox": [
+ 24,
+ 291,
+ 487,
+ 170
+ ],
+ "category_id": 9,
+ "area": 171626,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9235,
+ "image_id": 1674,
+ "bbox": [
+ 29,
+ 68,
+ 354,
+ 443
+ ],
+ "category_id": 9,
+ "area": 308864,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9236,
+ "image_id": 1674,
+ "bbox": [
+ 240,
+ 0,
+ 86,
+ 113
+ ],
+ "category_id": 6,
+ "area": 19370,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9237,
+ "image_id": 1675,
+ "bbox": [
+ 185,
+ 161,
+ 39,
+ 80
+ ],
+ "category_id": 8,
+ "area": 10197,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9238,
+ "image_id": 1675,
+ "bbox": [
+ 280,
+ 269,
+ 90,
+ 123
+ ],
+ "category_id": 8,
+ "area": 35550,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9239,
+ "image_id": 1675,
+ "bbox": [
+ 0,
+ 110,
+ 300,
+ 401
+ ],
+ "category_id": 6,
+ "area": 383958,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9240,
+ "image_id": 1675,
+ "bbox": [
+ 337,
+ 392,
+ 113,
+ 119
+ ],
+ "category_id": 6,
+ "area": 42993,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9241,
+ "image_id": 1676,
+ "bbox": [
+ 89,
+ 249,
+ 158,
+ 107
+ ],
+ "category_id": 8,
+ "area": 134244,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9242,
+ "image_id": 1676,
+ "bbox": [
+ 158,
+ 146,
+ 91,
+ 169
+ ],
+ "category_id": 8,
+ "area": 122794,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9243,
+ "image_id": 1676,
+ "bbox": [
+ 233,
+ 140,
+ 125,
+ 140
+ ],
+ "category_id": 8,
+ "area": 138824,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9244,
+ "image_id": 1676,
+ "bbox": [
+ 322,
+ 154,
+ 115,
+ 71
+ ],
+ "category_id": 8,
+ "area": 65534,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9245,
+ "image_id": 1677,
+ "bbox": [
+ 1,
+ 1,
+ 120,
+ 58
+ ],
+ "category_id": 6,
+ "area": 66744,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9246,
+ "image_id": 1677,
+ "bbox": [
+ 96,
+ 2,
+ 137,
+ 178
+ ],
+ "category_id": 6,
+ "area": 231240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9247,
+ "image_id": 1677,
+ "bbox": [
+ 260,
+ 2,
+ 130,
+ 134
+ ],
+ "category_id": 6,
+ "area": 165020,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9248,
+ "image_id": 1677,
+ "bbox": [
+ 252,
+ 120,
+ 114,
+ 99
+ ],
+ "category_id": 6,
+ "area": 106586,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9249,
+ "image_id": 1677,
+ "bbox": [
+ 109,
+ 175,
+ 151,
+ 122
+ ],
+ "category_id": 6,
+ "area": 174408,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9250,
+ "image_id": 1677,
+ "bbox": [
+ 51,
+ 301,
+ 154,
+ 122
+ ],
+ "category_id": 6,
+ "area": 177936,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9251,
+ "image_id": 1677,
+ "bbox": [
+ 413,
+ 1,
+ 74,
+ 63
+ ],
+ "category_id": 6,
+ "area": 44704,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9252,
+ "image_id": 1677,
+ "bbox": [
+ 304,
+ 212,
+ 138,
+ 103
+ ],
+ "category_id": 6,
+ "area": 134235,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9257,
+ "image_id": 1678,
+ "bbox": [
+ 187,
+ 34,
+ 48,
+ 93
+ ],
+ "category_id": 10,
+ "area": 35854,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9258,
+ "image_id": 1678,
+ "bbox": [
+ 410,
+ 180,
+ 54,
+ 98
+ ],
+ "category_id": 10,
+ "area": 42642,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9259,
+ "image_id": 1678,
+ "bbox": [
+ 201,
+ 312,
+ 42,
+ 75
+ ],
+ "category_id": 10,
+ "area": 25440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9260,
+ "image_id": 1678,
+ "bbox": [
+ 357,
+ 322,
+ 27,
+ 55
+ ],
+ "category_id": 10,
+ "area": 12036,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9261,
+ "image_id": 1678,
+ "bbox": [
+ 432,
+ 328,
+ 50,
+ 88
+ ],
+ "category_id": 10,
+ "area": 35530,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9262,
+ "image_id": 1678,
+ "bbox": [
+ 252,
+ 430,
+ 180,
+ 79
+ ],
+ "category_id": 10,
+ "area": 112725,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9263,
+ "image_id": 1678,
+ "bbox": [
+ 89,
+ 281,
+ 24,
+ 48
+ ],
+ "category_id": 10,
+ "area": 9579,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9264,
+ "image_id": 1678,
+ "bbox": [
+ 121,
+ 102,
+ 28,
+ 48
+ ],
+ "category_id": 10,
+ "area": 10918,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9269,
+ "image_id": 1681,
+ "bbox": [
+ 129,
+ 81,
+ 122,
+ 35
+ ],
+ "category_id": 8,
+ "area": 8688,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9270,
+ "image_id": 1681,
+ "bbox": [
+ 203,
+ 254,
+ 200,
+ 96
+ ],
+ "category_id": 8,
+ "area": 38480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9280,
+ "image_id": 1683,
+ "bbox": [
+ 181,
+ 96,
+ 31,
+ 42
+ ],
+ "category_id": 10,
+ "area": 3850,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9281,
+ "image_id": 1683,
+ "bbox": [
+ 17,
+ 92,
+ 401,
+ 419
+ ],
+ "category_id": 9,
+ "area": 480936,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9284,
+ "image_id": 1684,
+ "bbox": [
+ 232,
+ 289,
+ 64,
+ 42
+ ],
+ "category_id": 8,
+ "area": 21870,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9286,
+ "image_id": 1685,
+ "bbox": [
+ 0,
+ 251,
+ 72,
+ 91
+ ],
+ "category_id": 9,
+ "area": 9435,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9287,
+ "image_id": 1685,
+ "bbox": [
+ 1,
+ 80,
+ 394,
+ 263
+ ],
+ "category_id": 9,
+ "area": 147846,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9288,
+ "image_id": 1686,
+ "bbox": [
+ 66,
+ 119,
+ 444,
+ 315
+ ],
+ "category_id": 8,
+ "area": 1081962,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9289,
+ "image_id": 1686,
+ "bbox": [
+ 261,
+ 0,
+ 249,
+ 143
+ ],
+ "category_id": 8,
+ "area": 274940,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9290,
+ "image_id": 1687,
+ "bbox": [
+ 163,
+ 144,
+ 122,
+ 115
+ ],
+ "category_id": 10,
+ "area": 49572,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9291,
+ "image_id": 1687,
+ "bbox": [
+ 103,
+ 317,
+ 63,
+ 137
+ ],
+ "category_id": 10,
+ "area": 30846,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9292,
+ "image_id": 1687,
+ "bbox": [
+ 221,
+ 210,
+ 203,
+ 301
+ ],
+ "category_id": 10,
+ "area": 215816,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9303,
+ "image_id": 1691,
+ "bbox": [
+ 39,
+ 0,
+ 192,
+ 256
+ ],
+ "category_id": 6,
+ "area": 61411,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9304,
+ "image_id": 1691,
+ "bbox": [
+ 87,
+ 229,
+ 327,
+ 282
+ ],
+ "category_id": 9,
+ "area": 114959,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9325,
+ "image_id": 1695,
+ "bbox": [
+ 44,
+ 21,
+ 465,
+ 489
+ ],
+ "category_id": 6,
+ "area": 293997,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9334,
+ "image_id": 1697,
+ "bbox": [
+ 170,
+ 287,
+ 225,
+ 207
+ ],
+ "category_id": 10,
+ "area": 73606,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9336,
+ "image_id": 1698,
+ "bbox": [
+ 105,
+ 357,
+ 194,
+ 138
+ ],
+ "category_id": 9,
+ "area": 25758,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9345,
+ "image_id": 1701,
+ "bbox": [
+ 50,
+ 0,
+ 224,
+ 511
+ ],
+ "category_id": 8,
+ "area": 367463,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9346,
+ "image_id": 1701,
+ "bbox": [
+ 318,
+ 158,
+ 193,
+ 352
+ ],
+ "category_id": 8,
+ "area": 217440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9347,
+ "image_id": 1702,
+ "bbox": [
+ 83,
+ 140,
+ 362,
+ 371
+ ],
+ "category_id": 9,
+ "area": 5107112,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9348,
+ "image_id": 1702,
+ "bbox": [
+ 264,
+ 43,
+ 116,
+ 189
+ ],
+ "category_id": 10,
+ "area": 838490,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9350,
+ "image_id": 1703,
+ "bbox": [
+ 146,
+ 208,
+ 294,
+ 215
+ ],
+ "category_id": 9,
+ "area": 163520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9351,
+ "image_id": 1703,
+ "bbox": [
+ 129,
+ 107,
+ 82,
+ 133
+ ],
+ "category_id": 6,
+ "area": 28236,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9352,
+ "image_id": 1703,
+ "bbox": [
+ 234,
+ 64,
+ 61,
+ 76
+ ],
+ "category_id": 6,
+ "area": 12168,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9353,
+ "image_id": 1703,
+ "bbox": [
+ 278,
+ 19,
+ 80,
+ 115
+ ],
+ "category_id": 6,
+ "area": 23868,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9354,
+ "image_id": 1703,
+ "bbox": [
+ 318,
+ 0,
+ 126,
+ 111
+ ],
+ "category_id": 6,
+ "area": 36240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9355,
+ "image_id": 1704,
+ "bbox": [
+ 236,
+ 173,
+ 92,
+ 188
+ ],
+ "category_id": 10,
+ "area": 28819,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9366,
+ "image_id": 1708,
+ "bbox": [
+ 211,
+ 242,
+ 151,
+ 188
+ ],
+ "category_id": 9,
+ "area": 62604,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9367,
+ "image_id": 1708,
+ "bbox": [
+ 346,
+ 196,
+ 116,
+ 195
+ ],
+ "category_id": 9,
+ "area": 50140,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9392,
+ "image_id": 1712,
+ "bbox": [
+ 70,
+ 0,
+ 325,
+ 512
+ ],
+ "category_id": 9,
+ "area": 455846,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9393,
+ "image_id": 1712,
+ "bbox": [
+ 318,
+ 0,
+ 193,
+ 226
+ ],
+ "category_id": 6,
+ "area": 119840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9404,
+ "image_id": 1715,
+ "bbox": [
+ 164,
+ 83,
+ 65,
+ 134
+ ],
+ "category_id": 8,
+ "area": 69580,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9405,
+ "image_id": 1715,
+ "bbox": [
+ 254,
+ 151,
+ 76,
+ 250
+ ],
+ "category_id": 8,
+ "area": 151294,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9406,
+ "image_id": 1715,
+ "bbox": [
+ 333,
+ 339,
+ 159,
+ 167
+ ],
+ "category_id": 8,
+ "area": 210741,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9418,
+ "image_id": 1720,
+ "bbox": [
+ 119,
+ 90,
+ 277,
+ 420
+ ],
+ "category_id": 10,
+ "area": 921440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9424,
+ "image_id": 1722,
+ "bbox": [
+ 179,
+ 186,
+ 43,
+ 67
+ ],
+ "category_id": 10,
+ "area": 8436,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9426,
+ "image_id": 1722,
+ "bbox": [
+ 73,
+ 155,
+ 283,
+ 356
+ ],
+ "category_id": 9,
+ "area": 288804,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9430,
+ "image_id": 1723,
+ "bbox": [
+ 233,
+ 97,
+ 278,
+ 414
+ ],
+ "category_id": 6,
+ "area": 518300,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9433,
+ "image_id": 1724,
+ "bbox": [
+ 195,
+ 375,
+ 25,
+ 35
+ ],
+ "category_id": 6,
+ "area": 6225,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9434,
+ "image_id": 1724,
+ "bbox": [
+ 247,
+ 364,
+ 33,
+ 29
+ ],
+ "category_id": 6,
+ "area": 6867,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9435,
+ "image_id": 1724,
+ "bbox": [
+ 293,
+ 323,
+ 42,
+ 45
+ ],
+ "category_id": 6,
+ "area": 13110,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9436,
+ "image_id": 1724,
+ "bbox": [
+ 318,
+ 300,
+ 34,
+ 62
+ ],
+ "category_id": 6,
+ "area": 14803,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9441,
+ "image_id": 1727,
+ "bbox": [
+ 208,
+ 196,
+ 152,
+ 104
+ ],
+ "category_id": 9,
+ "area": 47728,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9443,
+ "image_id": 1728,
+ "bbox": [
+ 0,
+ 316,
+ 225,
+ 194
+ ],
+ "category_id": 9,
+ "area": 66038,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9444,
+ "image_id": 1729,
+ "bbox": [
+ 208,
+ 283,
+ 221,
+ 159
+ ],
+ "category_id": 8,
+ "area": 159016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9445,
+ "image_id": 1730,
+ "bbox": [
+ 0,
+ 151,
+ 410,
+ 359
+ ],
+ "category_id": 6,
+ "area": 517104,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9446,
+ "image_id": 1730,
+ "bbox": [
+ 190,
+ 212,
+ 242,
+ 184
+ ],
+ "category_id": 8,
+ "area": 156090,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9447,
+ "image_id": 1730,
+ "bbox": [
+ 460,
+ 93,
+ 50,
+ 274
+ ],
+ "category_id": 8,
+ "area": 48895,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9451,
+ "image_id": 1732,
+ "bbox": [
+ 44,
+ 163,
+ 133,
+ 177
+ ],
+ "category_id": 6,
+ "area": 55471,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9452,
+ "image_id": 1732,
+ "bbox": [
+ 286,
+ 240,
+ 107,
+ 192
+ ],
+ "category_id": 6,
+ "area": 48278,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9453,
+ "image_id": 1732,
+ "bbox": [
+ 242,
+ 140,
+ 102,
+ 195
+ ],
+ "category_id": 6,
+ "area": 46899,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9454,
+ "image_id": 1732,
+ "bbox": [
+ 222,
+ 384,
+ 179,
+ 127
+ ],
+ "category_id": 6,
+ "area": 53246,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9455,
+ "image_id": 1732,
+ "bbox": [
+ 170,
+ 315,
+ 105,
+ 141
+ ],
+ "category_id": 6,
+ "area": 34672,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9456,
+ "image_id": 1732,
+ "bbox": [
+ 130,
+ 150,
+ 105,
+ 136
+ ],
+ "category_id": 6,
+ "area": 33293,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9457,
+ "image_id": 1732,
+ "bbox": [
+ 353,
+ 212,
+ 134,
+ 169
+ ],
+ "category_id": 6,
+ "area": 53130,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9482,
+ "image_id": 1734,
+ "bbox": [
+ 0,
+ 2,
+ 344,
+ 509
+ ],
+ "category_id": 8,
+ "area": 614900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9483,
+ "image_id": 1734,
+ "bbox": [
+ 211,
+ 44,
+ 300,
+ 344
+ ],
+ "category_id": 8,
+ "area": 363484,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9484,
+ "image_id": 1735,
+ "bbox": [
+ 35,
+ 93,
+ 25,
+ 32
+ ],
+ "category_id": 6,
+ "area": 5796,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9485,
+ "image_id": 1735,
+ "bbox": [
+ 50,
+ 331,
+ 19,
+ 51
+ ],
+ "category_id": 6,
+ "area": 7020,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9486,
+ "image_id": 1735,
+ "bbox": [
+ 101,
+ 304,
+ 27,
+ 34
+ ],
+ "category_id": 6,
+ "area": 6570,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9487,
+ "image_id": 1735,
+ "bbox": [
+ 123,
+ 298,
+ 27,
+ 33
+ ],
+ "category_id": 6,
+ "area": 6532,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9488,
+ "image_id": 1735,
+ "bbox": [
+ 338,
+ 436,
+ 34,
+ 66
+ ],
+ "category_id": 6,
+ "area": 15707,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9489,
+ "image_id": 1735,
+ "bbox": [
+ 341,
+ 275,
+ 10,
+ 25
+ ],
+ "category_id": 6,
+ "area": 1855,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9490,
+ "image_id": 1735,
+ "bbox": [
+ 353,
+ 346,
+ 16,
+ 21
+ ],
+ "category_id": 6,
+ "area": 2530,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9491,
+ "image_id": 1735,
+ "bbox": [
+ 363,
+ 469,
+ 53,
+ 42
+ ],
+ "category_id": 6,
+ "area": 15930,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9497,
+ "image_id": 1737,
+ "bbox": [
+ 0,
+ 122,
+ 195,
+ 353
+ ],
+ "category_id": 9,
+ "area": 172584,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9528,
+ "image_id": 1743,
+ "bbox": [
+ 198,
+ 152,
+ 279,
+ 343
+ ],
+ "category_id": 8,
+ "area": 111708,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9530,
+ "image_id": 1744,
+ "bbox": [
+ 382,
+ 95,
+ 129,
+ 177
+ ],
+ "category_id": 6,
+ "area": 38793,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9531,
+ "image_id": 1744,
+ "bbox": [
+ 98,
+ 281,
+ 412,
+ 229
+ ],
+ "category_id": 6,
+ "area": 159360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9532,
+ "image_id": 1744,
+ "bbox": [
+ 102,
+ 180,
+ 408,
+ 270
+ ],
+ "category_id": 9,
+ "area": 186396,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9533,
+ "image_id": 1745,
+ "bbox": [
+ 154,
+ 0,
+ 353,
+ 390
+ ],
+ "category_id": 9,
+ "area": 485316,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9534,
+ "image_id": 1745,
+ "bbox": [
+ 2,
+ 117,
+ 336,
+ 394
+ ],
+ "category_id": 9,
+ "area": 467310,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9535,
+ "image_id": 1745,
+ "bbox": [
+ 176,
+ 164,
+ 332,
+ 343
+ ],
+ "category_id": 9,
+ "area": 401856,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9537,
+ "image_id": 1746,
+ "bbox": [
+ 142,
+ 84,
+ 188,
+ 112
+ ],
+ "category_id": 9,
+ "area": 30150,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9538,
+ "image_id": 1746,
+ "bbox": [
+ 141,
+ 152,
+ 300,
+ 225
+ ],
+ "category_id": 9,
+ "area": 96571,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9539,
+ "image_id": 1746,
+ "bbox": [
+ 45,
+ 276,
+ 433,
+ 229
+ ],
+ "category_id": 9,
+ "area": 141932,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9543,
+ "image_id": 1748,
+ "bbox": [
+ 11,
+ 268,
+ 285,
+ 182
+ ],
+ "category_id": 8,
+ "area": 182272,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9544,
+ "image_id": 1748,
+ "bbox": [
+ 204,
+ 196,
+ 225,
+ 175
+ ],
+ "category_id": 8,
+ "area": 138498,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9545,
+ "image_id": 1749,
+ "bbox": [
+ 0,
+ 144,
+ 266,
+ 367
+ ],
+ "category_id": 9,
+ "area": 215254,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9546,
+ "image_id": 1749,
+ "bbox": [
+ 259,
+ 0,
+ 252,
+ 286
+ ],
+ "category_id": 6,
+ "area": 158840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9549,
+ "image_id": 1750,
+ "bbox": [
+ 112,
+ 278,
+ 246,
+ 223
+ ],
+ "category_id": 8,
+ "area": 379500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9550,
+ "image_id": 1750,
+ "bbox": [
+ 219,
+ 36,
+ 291,
+ 300
+ ],
+ "category_id": 8,
+ "area": 603525,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9551,
+ "image_id": 1751,
+ "bbox": [
+ 66,
+ 139,
+ 281,
+ 371
+ ],
+ "category_id": 8,
+ "area": 518056,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9552,
+ "image_id": 1751,
+ "bbox": [
+ 84,
+ 0,
+ 427,
+ 465
+ ],
+ "category_id": 6,
+ "area": 984300,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9557,
+ "image_id": 1753,
+ "bbox": [
+ 211,
+ 358,
+ 125,
+ 111
+ ],
+ "category_id": 6,
+ "area": 49536,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9558,
+ "image_id": 1753,
+ "bbox": [
+ 311,
+ 337,
+ 103,
+ 105
+ ],
+ "category_id": 6,
+ "area": 38584,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9559,
+ "image_id": 1753,
+ "bbox": [
+ 178,
+ 281,
+ 181,
+ 65
+ ],
+ "category_id": 6,
+ "area": 42149,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9562,
+ "image_id": 1755,
+ "bbox": [
+ 0,
+ 0,
+ 202,
+ 509
+ ],
+ "category_id": 6,
+ "area": 361284,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9563,
+ "image_id": 1755,
+ "bbox": [
+ 134,
+ 120,
+ 297,
+ 241
+ ],
+ "category_id": 8,
+ "area": 250796,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9564,
+ "image_id": 1755,
+ "bbox": [
+ 236,
+ 52,
+ 275,
+ 399
+ ],
+ "category_id": 8,
+ "area": 384033,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9568,
+ "image_id": 1756,
+ "bbox": [
+ 1,
+ 178,
+ 96,
+ 127
+ ],
+ "category_id": 6,
+ "area": 35432,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9569,
+ "image_id": 1756,
+ "bbox": [
+ 3,
+ 311,
+ 51,
+ 57
+ ],
+ "category_id": 6,
+ "area": 8547,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9570,
+ "image_id": 1756,
+ "bbox": [
+ 27,
+ 2,
+ 146,
+ 169
+ ],
+ "category_id": 6,
+ "area": 71906,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9571,
+ "image_id": 1757,
+ "bbox": [
+ 188,
+ 58,
+ 289,
+ 453
+ ],
+ "category_id": 10,
+ "area": 234768,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9572,
+ "image_id": 1757,
+ "bbox": [
+ 41,
+ 191,
+ 233,
+ 300
+ ],
+ "category_id": 10,
+ "area": 125280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9573,
+ "image_id": 1758,
+ "bbox": [
+ 185,
+ 163,
+ 90,
+ 85
+ ],
+ "category_id": 8,
+ "area": 61020,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9575,
+ "image_id": 1759,
+ "bbox": [
+ 201,
+ 125,
+ 298,
+ 256
+ ],
+ "category_id": 9,
+ "area": 268945,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9576,
+ "image_id": 1760,
+ "bbox": [
+ 150,
+ 225,
+ 281,
+ 286
+ ],
+ "category_id": 8,
+ "area": 282204,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9577,
+ "image_id": 1760,
+ "bbox": [
+ 220,
+ 64,
+ 290,
+ 232
+ ],
+ "category_id": 8,
+ "area": 236350,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9578,
+ "image_id": 1760,
+ "bbox": [
+ 365,
+ 9,
+ 131,
+ 187
+ ],
+ "category_id": 8,
+ "area": 86856,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9590,
+ "image_id": 1762,
+ "bbox": [
+ 433,
+ 305,
+ 78,
+ 43
+ ],
+ "category_id": 6,
+ "area": 3850,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9591,
+ "image_id": 1762,
+ "bbox": [
+ 101,
+ 322,
+ 134,
+ 107
+ ],
+ "category_id": 6,
+ "area": 16113,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9592,
+ "image_id": 1762,
+ "bbox": [
+ 188,
+ 170,
+ 171,
+ 116
+ ],
+ "category_id": 6,
+ "area": 22211,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9593,
+ "image_id": 1762,
+ "bbox": [
+ 391,
+ 200,
+ 99,
+ 69
+ ],
+ "category_id": 6,
+ "area": 7663,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9594,
+ "image_id": 1762,
+ "bbox": [
+ 369,
+ 246,
+ 99,
+ 76
+ ],
+ "category_id": 6,
+ "area": 8439,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9595,
+ "image_id": 1762,
+ "bbox": [
+ 0,
+ 184,
+ 117,
+ 84
+ ],
+ "category_id": 6,
+ "area": 11040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9596,
+ "image_id": 1762,
+ "bbox": [
+ 0,
+ 297,
+ 84,
+ 80
+ ],
+ "category_id": 6,
+ "area": 7636,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9597,
+ "image_id": 1762,
+ "bbox": [
+ 212,
+ 306,
+ 94,
+ 85
+ ],
+ "category_id": 6,
+ "area": 8924,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9598,
+ "image_id": 1762,
+ "bbox": [
+ 337,
+ 344,
+ 174,
+ 167
+ ],
+ "category_id": 6,
+ "area": 32470,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9605,
+ "image_id": 1764,
+ "bbox": [
+ 0,
+ 197,
+ 258,
+ 297
+ ],
+ "category_id": 8,
+ "area": 268320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9606,
+ "image_id": 1764,
+ "bbox": [
+ 257,
+ 217,
+ 253,
+ 275
+ ],
+ "category_id": 8,
+ "area": 244724,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9608,
+ "image_id": 1766,
+ "bbox": [
+ 86,
+ 199,
+ 144,
+ 102
+ ],
+ "category_id": 9,
+ "area": 38086,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9609,
+ "image_id": 1766,
+ "bbox": [
+ 243,
+ 155,
+ 134,
+ 243
+ ],
+ "category_id": 6,
+ "area": 84480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9610,
+ "image_id": 1766,
+ "bbox": [
+ 386,
+ 316,
+ 75,
+ 127
+ ],
+ "category_id": 6,
+ "area": 24739,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9611,
+ "image_id": 1767,
+ "bbox": [
+ 0,
+ 61,
+ 305,
+ 203
+ ],
+ "category_id": 9,
+ "area": 178840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9612,
+ "image_id": 1767,
+ "bbox": [
+ 210,
+ 74,
+ 124,
+ 195
+ ],
+ "category_id": 10,
+ "area": 69804,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9633,
+ "image_id": 1772,
+ "bbox": [
+ 106,
+ 212,
+ 209,
+ 227
+ ],
+ "category_id": 9,
+ "area": 62622,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9634,
+ "image_id": 1773,
+ "bbox": [
+ 248,
+ 363,
+ 126,
+ 100
+ ],
+ "category_id": 8,
+ "area": 23180,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9635,
+ "image_id": 1773,
+ "bbox": [
+ 39,
+ 206,
+ 81,
+ 78
+ ],
+ "category_id": 8,
+ "area": 11618,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9636,
+ "image_id": 1773,
+ "bbox": [
+ 0,
+ 35,
+ 36,
+ 44
+ ],
+ "category_id": 8,
+ "area": 2940,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9637,
+ "image_id": 1773,
+ "bbox": [
+ 434,
+ 162,
+ 46,
+ 78
+ ],
+ "category_id": 8,
+ "area": 6586,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9638,
+ "image_id": 1773,
+ "bbox": [
+ 43,
+ 172,
+ 27,
+ 50
+ ],
+ "category_id": 8,
+ "area": 2592,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9639,
+ "image_id": 1773,
+ "bbox": [
+ 374,
+ 90,
+ 34,
+ 42
+ ],
+ "category_id": 8,
+ "area": 2680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9640,
+ "image_id": 1773,
+ "bbox": [
+ 173,
+ 144,
+ 41,
+ 66
+ ],
+ "category_id": 8,
+ "area": 5040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9641,
+ "image_id": 1773,
+ "bbox": [
+ 428,
+ 130,
+ 40,
+ 33
+ ],
+ "category_id": 8,
+ "area": 2496,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9642,
+ "image_id": 1773,
+ "bbox": [
+ 477,
+ 177,
+ 33,
+ 39
+ ],
+ "category_id": 8,
+ "area": 2405,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9694,
+ "image_id": 1776,
+ "bbox": [
+ 125,
+ 197,
+ 386,
+ 253
+ ],
+ "category_id": 9,
+ "area": 164725,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9695,
+ "image_id": 1776,
+ "bbox": [
+ 35,
+ 229,
+ 475,
+ 280
+ ],
+ "category_id": 6,
+ "area": 225090,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9696,
+ "image_id": 1776,
+ "bbox": [
+ 187,
+ 8,
+ 79,
+ 189
+ ],
+ "category_id": 6,
+ "area": 25544,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9697,
+ "image_id": 1776,
+ "bbox": [
+ 239,
+ 75,
+ 184,
+ 170
+ ],
+ "category_id": 6,
+ "area": 52910,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9698,
+ "image_id": 1777,
+ "bbox": [
+ 0,
+ 0,
+ 83,
+ 114
+ ],
+ "category_id": 6,
+ "area": 57840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9699,
+ "image_id": 1777,
+ "bbox": [
+ 103,
+ 24,
+ 206,
+ 140
+ ],
+ "category_id": 6,
+ "area": 176106,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9700,
+ "image_id": 1777,
+ "bbox": [
+ 417,
+ 39,
+ 94,
+ 158
+ ],
+ "category_id": 6,
+ "area": 90636,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9710,
+ "image_id": 1780,
+ "bbox": [
+ 202,
+ 90,
+ 188,
+ 210
+ ],
+ "category_id": 9,
+ "area": 139120,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9711,
+ "image_id": 1780,
+ "bbox": [
+ 254,
+ 75,
+ 175,
+ 225
+ ],
+ "category_id": 9,
+ "area": 139163,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9720,
+ "image_id": 1783,
+ "bbox": [
+ 460,
+ 185,
+ 48,
+ 93
+ ],
+ "category_id": 9,
+ "area": 16104,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9721,
+ "image_id": 1783,
+ "bbox": [
+ 250,
+ 95,
+ 88,
+ 127
+ ],
+ "category_id": 9,
+ "area": 39380,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9722,
+ "image_id": 1783,
+ "bbox": [
+ 332,
+ 115,
+ 61,
+ 78
+ ],
+ "category_id": 9,
+ "area": 16983,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9723,
+ "image_id": 1783,
+ "bbox": [
+ 364,
+ 160,
+ 111,
+ 123
+ ],
+ "category_id": 9,
+ "area": 48546,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9724,
+ "image_id": 1784,
+ "bbox": [
+ 6,
+ 70,
+ 238,
+ 433
+ ],
+ "category_id": 9,
+ "area": 69407,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9725,
+ "image_id": 1784,
+ "bbox": [
+ 204,
+ 103,
+ 297,
+ 320
+ ],
+ "category_id": 9,
+ "area": 64064,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9727,
+ "image_id": 1785,
+ "bbox": [
+ 260,
+ 132,
+ 251,
+ 264
+ ],
+ "category_id": 9,
+ "area": 73023,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9728,
+ "image_id": 1786,
+ "bbox": [
+ 219,
+ 108,
+ 291,
+ 288
+ ],
+ "category_id": 8,
+ "area": 97647,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9730,
+ "image_id": 1787,
+ "bbox": [
+ 107,
+ 108,
+ 148,
+ 224
+ ],
+ "category_id": 10,
+ "area": 95948,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9731,
+ "image_id": 1787,
+ "bbox": [
+ 239,
+ 104,
+ 54,
+ 94
+ ],
+ "category_id": 10,
+ "area": 14700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9732,
+ "image_id": 1787,
+ "bbox": [
+ 270,
+ 146,
+ 50,
+ 76
+ ],
+ "category_id": 10,
+ "area": 11187,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9733,
+ "image_id": 1787,
+ "bbox": [
+ 224,
+ 368,
+ 106,
+ 122
+ ],
+ "category_id": 10,
+ "area": 37286,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9734,
+ "image_id": 1787,
+ "bbox": [
+ 249,
+ 395,
+ 101,
+ 111
+ ],
+ "category_id": 10,
+ "area": 32670,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9756,
+ "image_id": 1791,
+ "bbox": [
+ 169,
+ 2,
+ 236,
+ 292
+ ],
+ "category_id": 9,
+ "area": 232580,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9757,
+ "image_id": 1791,
+ "bbox": [
+ 148,
+ 128,
+ 136,
+ 216
+ ],
+ "category_id": 10,
+ "area": 99792,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9761,
+ "image_id": 1792,
+ "bbox": [
+ 75,
+ 189,
+ 158,
+ 102
+ ],
+ "category_id": 8,
+ "area": 57024,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9762,
+ "image_id": 1792,
+ "bbox": [
+ 216,
+ 250,
+ 148,
+ 173
+ ],
+ "category_id": 8,
+ "area": 90396,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9787,
+ "image_id": 1797,
+ "bbox": [
+ 2,
+ 221,
+ 329,
+ 285
+ ],
+ "category_id": 8,
+ "area": 647388,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9788,
+ "image_id": 1797,
+ "bbox": [
+ 130,
+ 206,
+ 185,
+ 122
+ ],
+ "category_id": 8,
+ "area": 156240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9795,
+ "image_id": 1798,
+ "bbox": [
+ 165,
+ 384,
+ 33,
+ 56
+ ],
+ "category_id": 6,
+ "area": 13090,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9796,
+ "image_id": 1799,
+ "bbox": [
+ 62,
+ 0,
+ 103,
+ 81
+ ],
+ "category_id": 6,
+ "area": 75072,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9803,
+ "image_id": 1799,
+ "bbox": [
+ 464,
+ 130,
+ 47,
+ 129
+ ],
+ "category_id": 6,
+ "area": 54587,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9804,
+ "image_id": 1800,
+ "bbox": [
+ 246,
+ 184,
+ 192,
+ 266
+ ],
+ "category_id": 10,
+ "area": 136704,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9813,
+ "image_id": 1803,
+ "bbox": [
+ 172,
+ 197,
+ 114,
+ 153
+ ],
+ "category_id": 9,
+ "area": 118440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9815,
+ "image_id": 1803,
+ "bbox": [
+ 428,
+ 388,
+ 83,
+ 123
+ ],
+ "category_id": 6,
+ "area": 69235,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9820,
+ "image_id": 1804,
+ "bbox": [
+ 148,
+ 407,
+ 165,
+ 102
+ ],
+ "category_id": 6,
+ "area": 59616,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9821,
+ "image_id": 1804,
+ "bbox": [
+ 356,
+ 228,
+ 155,
+ 280
+ ],
+ "category_id": 6,
+ "area": 153260,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9822,
+ "image_id": 1804,
+ "bbox": [
+ 112,
+ 232,
+ 253,
+ 198
+ ],
+ "category_id": 6,
+ "area": 176886,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9823,
+ "image_id": 1804,
+ "bbox": [
+ 0,
+ 219,
+ 160,
+ 273
+ ],
+ "category_id": 8,
+ "area": 153984,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9842,
+ "image_id": 1805,
+ "bbox": [
+ 107,
+ 423,
+ 159,
+ 88
+ ],
+ "category_id": 6,
+ "area": 99575,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9843,
+ "image_id": 1805,
+ "bbox": [
+ 179,
+ 401,
+ 270,
+ 110
+ ],
+ "category_id": 6,
+ "area": 210370,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9849,
+ "image_id": 1807,
+ "bbox": [
+ 60,
+ 161,
+ 317,
+ 341
+ ],
+ "category_id": 9,
+ "area": 269698,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9854,
+ "image_id": 1808,
+ "bbox": [
+ 276,
+ 177,
+ 235,
+ 287
+ ],
+ "category_id": 8,
+ "area": 44642,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9855,
+ "image_id": 1809,
+ "bbox": [
+ 219,
+ 142,
+ 33,
+ 141
+ ],
+ "category_id": 10,
+ "area": 11224,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9857,
+ "image_id": 1809,
+ "bbox": [
+ 226,
+ 155,
+ 260,
+ 299
+ ],
+ "category_id": 9,
+ "area": 184775,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9859,
+ "image_id": 1810,
+ "bbox": [
+ 190,
+ 69,
+ 134,
+ 128
+ ],
+ "category_id": 9,
+ "area": 32000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9860,
+ "image_id": 1810,
+ "bbox": [
+ 148,
+ 150,
+ 134,
+ 191
+ ],
+ "category_id": 9,
+ "area": 47362,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9886,
+ "image_id": 1814,
+ "bbox": [
+ 165,
+ 209,
+ 88,
+ 140
+ ],
+ "category_id": 8,
+ "area": 40040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9887,
+ "image_id": 1814,
+ "bbox": [
+ 308,
+ 311,
+ 42,
+ 67
+ ],
+ "category_id": 8,
+ "area": 9240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9894,
+ "image_id": 1816,
+ "bbox": [
+ 98,
+ 168,
+ 291,
+ 276
+ ],
+ "category_id": 9,
+ "area": 207868,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9906,
+ "image_id": 1820,
+ "bbox": [
+ 39,
+ 284,
+ 453,
+ 222
+ ],
+ "category_id": 8,
+ "area": 689016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9908,
+ "image_id": 1821,
+ "bbox": [
+ 24,
+ 107,
+ 255,
+ 219
+ ],
+ "category_id": 8,
+ "area": 196504,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9909,
+ "image_id": 1821,
+ "bbox": [
+ 378,
+ 0,
+ 53,
+ 135
+ ],
+ "category_id": 6,
+ "area": 25460,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9910,
+ "image_id": 1821,
+ "bbox": [
+ 464,
+ 215,
+ 37,
+ 147
+ ],
+ "category_id": 6,
+ "area": 19344,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9924,
+ "image_id": 1823,
+ "bbox": [
+ 13,
+ 3,
+ 126,
+ 159
+ ],
+ "category_id": 6,
+ "area": 216534,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9925,
+ "image_id": 1823,
+ "bbox": [
+ 18,
+ 162,
+ 146,
+ 116
+ ],
+ "category_id": 6,
+ "area": 183050,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9926,
+ "image_id": 1823,
+ "bbox": [
+ 3,
+ 279,
+ 107,
+ 98
+ ],
+ "category_id": 6,
+ "area": 113870,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9927,
+ "image_id": 1823,
+ "bbox": [
+ 178,
+ 195,
+ 171,
+ 100
+ ],
+ "category_id": 6,
+ "area": 184513,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9928,
+ "image_id": 1823,
+ "bbox": [
+ 162,
+ 0,
+ 127,
+ 125
+ ],
+ "category_id": 6,
+ "area": 171990,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9929,
+ "image_id": 1823,
+ "bbox": [
+ 156,
+ 113,
+ 108,
+ 92
+ ],
+ "category_id": 6,
+ "area": 108531,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9930,
+ "image_id": 1823,
+ "bbox": [
+ 308,
+ 1,
+ 87,
+ 65
+ ],
+ "category_id": 6,
+ "area": 61152,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9931,
+ "image_id": 1823,
+ "bbox": [
+ 396,
+ 1,
+ 113,
+ 88
+ ],
+ "category_id": 6,
+ "area": 107060,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9932,
+ "image_id": 1823,
+ "bbox": [
+ 436,
+ 217,
+ 73,
+ 100
+ ],
+ "category_id": 6,
+ "area": 79083,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9936,
+ "image_id": 1824,
+ "bbox": [
+ 171,
+ 188,
+ 154,
+ 140
+ ],
+ "category_id": 6,
+ "area": 51392,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9942,
+ "image_id": 1825,
+ "bbox": [
+ 270,
+ 157,
+ 150,
+ 154
+ ],
+ "category_id": 8,
+ "area": 30595,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9943,
+ "image_id": 1826,
+ "bbox": [
+ 33,
+ 105,
+ 329,
+ 404
+ ],
+ "category_id": 9,
+ "area": 380952,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9976,
+ "image_id": 1829,
+ "bbox": [
+ 103,
+ 104,
+ 61,
+ 96
+ ],
+ "category_id": 8,
+ "area": 20655,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9977,
+ "image_id": 1829,
+ "bbox": [
+ 245,
+ 183,
+ 85,
+ 150
+ ],
+ "category_id": 8,
+ "area": 44943,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9979,
+ "image_id": 1830,
+ "bbox": [
+ 45,
+ 110,
+ 339,
+ 401
+ ],
+ "category_id": 8,
+ "area": 479685,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9980,
+ "image_id": 1831,
+ "bbox": [
+ 112,
+ 64,
+ 396,
+ 440
+ ],
+ "category_id": 9,
+ "area": 613800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9982,
+ "image_id": 1833,
+ "bbox": [
+ 38,
+ 182,
+ 167,
+ 289
+ ],
+ "category_id": 10,
+ "area": 86180,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9983,
+ "image_id": 1833,
+ "bbox": [
+ 187,
+ 157,
+ 128,
+ 168
+ ],
+ "category_id": 10,
+ "area": 38556,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9984,
+ "image_id": 1833,
+ "bbox": [
+ 186,
+ 11,
+ 106,
+ 162
+ ],
+ "category_id": 10,
+ "area": 30732,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9985,
+ "image_id": 1833,
+ "bbox": [
+ 327,
+ 52,
+ 101,
+ 165
+ ],
+ "category_id": 10,
+ "area": 29892,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9986,
+ "image_id": 1833,
+ "bbox": [
+ 417,
+ 80,
+ 94,
+ 187
+ ],
+ "category_id": 10,
+ "area": 31500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9987,
+ "image_id": 1833,
+ "bbox": [
+ 388,
+ 209,
+ 103,
+ 223
+ ],
+ "category_id": 10,
+ "area": 41088,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9990,
+ "image_id": 1834,
+ "bbox": [
+ 146,
+ 253,
+ 139,
+ 112
+ ],
+ "category_id": 8,
+ "area": 124188,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9991,
+ "image_id": 1834,
+ "bbox": [
+ 161,
+ 196,
+ 167,
+ 137
+ ],
+ "category_id": 8,
+ "area": 182748,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9992,
+ "image_id": 1834,
+ "bbox": [
+ 281,
+ 148,
+ 125,
+ 136
+ ],
+ "category_id": 8,
+ "area": 135177,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9993,
+ "image_id": 1835,
+ "bbox": [
+ 182,
+ 121,
+ 329,
+ 281
+ ],
+ "category_id": 10,
+ "area": 732355,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9994,
+ "image_id": 1836,
+ "bbox": [
+ 23,
+ 20,
+ 462,
+ 451
+ ],
+ "category_id": 9,
+ "area": 1353274,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9995,
+ "image_id": 1837,
+ "bbox": [
+ 149,
+ 2,
+ 230,
+ 307
+ ],
+ "category_id": 8,
+ "area": 546508,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9996,
+ "image_id": 1837,
+ "bbox": [
+ 1,
+ 0,
+ 147,
+ 85
+ ],
+ "category_id": 8,
+ "area": 97881,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9997,
+ "image_id": 1837,
+ "bbox": [
+ 0,
+ 223,
+ 100,
+ 164
+ ],
+ "category_id": 8,
+ "area": 127125,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9998,
+ "image_id": 1838,
+ "bbox": [
+ 0,
+ 0,
+ 316,
+ 335
+ ],
+ "category_id": 6,
+ "area": 267129,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 9999,
+ "image_id": 1838,
+ "bbox": [
+ 374,
+ 198,
+ 93,
+ 85
+ ],
+ "category_id": 6,
+ "area": 20001,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10000,
+ "image_id": 1838,
+ "bbox": [
+ 417,
+ 290,
+ 42,
+ 37
+ ],
+ "category_id": 6,
+ "area": 3920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10005,
+ "image_id": 1839,
+ "bbox": [
+ 161,
+ 76,
+ 206,
+ 361
+ ],
+ "category_id": 10,
+ "area": 591312,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10013,
+ "image_id": 1840,
+ "bbox": [
+ 431,
+ 119,
+ 76,
+ 162
+ ],
+ "category_id": 6,
+ "area": 38016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10014,
+ "image_id": 1841,
+ "bbox": [
+ 102,
+ 132,
+ 336,
+ 378
+ ],
+ "category_id": 8,
+ "area": 1005480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10015,
+ "image_id": 1841,
+ "bbox": [
+ 304,
+ 141,
+ 208,
+ 149
+ ],
+ "category_id": 8,
+ "area": 246480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10016,
+ "image_id": 1842,
+ "bbox": [
+ 0,
+ 126,
+ 312,
+ 193
+ ],
+ "category_id": 9,
+ "area": 125077,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10026,
+ "image_id": 1845,
+ "bbox": [
+ 77,
+ 213,
+ 204,
+ 297
+ ],
+ "category_id": 6,
+ "area": 482304,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10061,
+ "image_id": 1850,
+ "bbox": [
+ 78,
+ 30,
+ 285,
+ 283
+ ],
+ "category_id": 8,
+ "area": 284172,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10062,
+ "image_id": 1850,
+ "bbox": [
+ 86,
+ 155,
+ 410,
+ 315
+ ],
+ "category_id": 8,
+ "area": 454075,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10063,
+ "image_id": 1851,
+ "bbox": [
+ 285,
+ 6,
+ 46,
+ 59
+ ],
+ "category_id": 6,
+ "area": 8610,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10064,
+ "image_id": 1851,
+ "bbox": [
+ 362,
+ 25,
+ 40,
+ 43
+ ],
+ "category_id": 6,
+ "area": 5544,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10065,
+ "image_id": 1851,
+ "bbox": [
+ 327,
+ 70,
+ 85,
+ 63
+ ],
+ "category_id": 6,
+ "area": 17063,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10066,
+ "image_id": 1851,
+ "bbox": [
+ 0,
+ 370,
+ 132,
+ 141
+ ],
+ "category_id": 6,
+ "area": 58750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10067,
+ "image_id": 1851,
+ "bbox": [
+ 356,
+ 430,
+ 155,
+ 81
+ ],
+ "category_id": 6,
+ "area": 39456,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10068,
+ "image_id": 1851,
+ "bbox": [
+ 196,
+ 392,
+ 122,
+ 55
+ ],
+ "category_id": 6,
+ "area": 21266,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10069,
+ "image_id": 1851,
+ "bbox": [
+ 95,
+ 270,
+ 373,
+ 204
+ ],
+ "category_id": 9,
+ "area": 239282,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10073,
+ "image_id": 1853,
+ "bbox": [
+ 126,
+ 253,
+ 276,
+ 194
+ ],
+ "category_id": 8,
+ "area": 425170,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10074,
+ "image_id": 1853,
+ "bbox": [
+ 116,
+ 186,
+ 105,
+ 88
+ ],
+ "category_id": 8,
+ "area": 73656,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10075,
+ "image_id": 1853,
+ "bbox": [
+ 396,
+ 226,
+ 82,
+ 50
+ ],
+ "category_id": 8,
+ "area": 32956,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10076,
+ "image_id": 1853,
+ "bbox": [
+ 422,
+ 192,
+ 50,
+ 28
+ ],
+ "category_id": 8,
+ "area": 11651,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10078,
+ "image_id": 1854,
+ "bbox": [
+ 96,
+ 142,
+ 125,
+ 161
+ ],
+ "category_id": 6,
+ "area": 99264,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10079,
+ "image_id": 1854,
+ "bbox": [
+ 310,
+ 102,
+ 199,
+ 172
+ ],
+ "category_id": 6,
+ "area": 168038,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10105,
+ "image_id": 1856,
+ "bbox": [
+ 6,
+ 16,
+ 426,
+ 356
+ ],
+ "category_id": 9,
+ "area": 113152,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10106,
+ "image_id": 1856,
+ "bbox": [
+ 129,
+ 96,
+ 373,
+ 358
+ ],
+ "category_id": 9,
+ "area": 99846,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10107,
+ "image_id": 1857,
+ "bbox": [
+ 152,
+ 98,
+ 148,
+ 304
+ ],
+ "category_id": 8,
+ "area": 304695,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10108,
+ "image_id": 1857,
+ "bbox": [
+ 338,
+ 238,
+ 109,
+ 127
+ ],
+ "category_id": 8,
+ "area": 93432,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10111,
+ "image_id": 1859,
+ "bbox": [
+ 114,
+ 94,
+ 198,
+ 385
+ ],
+ "category_id": 8,
+ "area": 88920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10113,
+ "image_id": 1860,
+ "bbox": [
+ 201,
+ 0,
+ 128,
+ 186
+ ],
+ "category_id": 10,
+ "area": 68926,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10116,
+ "image_id": 1861,
+ "bbox": [
+ 0,
+ 45,
+ 301,
+ 464
+ ],
+ "category_id": 9,
+ "area": 262314,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10117,
+ "image_id": 1861,
+ "bbox": [
+ 229,
+ 107,
+ 282,
+ 402
+ ],
+ "category_id": 9,
+ "area": 212980,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10118,
+ "image_id": 1862,
+ "bbox": [
+ 0,
+ 20,
+ 96,
+ 158
+ ],
+ "category_id": 6,
+ "area": 33300,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10119,
+ "image_id": 1862,
+ "bbox": [
+ 145,
+ 153,
+ 326,
+ 358
+ ],
+ "category_id": 9,
+ "area": 254144,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10120,
+ "image_id": 1863,
+ "bbox": [
+ 58,
+ 77,
+ 81,
+ 136
+ ],
+ "category_id": 6,
+ "area": 17408,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10121,
+ "image_id": 1863,
+ "bbox": [
+ 174,
+ 174,
+ 128,
+ 130
+ ],
+ "category_id": 8,
+ "area": 26230,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10128,
+ "image_id": 1866,
+ "bbox": [
+ 229,
+ 134,
+ 232,
+ 377
+ ],
+ "category_id": 9,
+ "area": 116186,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10129,
+ "image_id": 1866,
+ "bbox": [
+ 48,
+ 135,
+ 110,
+ 197
+ ],
+ "category_id": 10,
+ "area": 28731,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10175,
+ "image_id": 1873,
+ "bbox": [
+ 90,
+ 0,
+ 250,
+ 178
+ ],
+ "category_id": 6,
+ "area": 125426,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10176,
+ "image_id": 1873,
+ "bbox": [
+ 224,
+ 170,
+ 256,
+ 318
+ ],
+ "category_id": 9,
+ "area": 229384,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10177,
+ "image_id": 1874,
+ "bbox": [
+ 49,
+ 135,
+ 341,
+ 223
+ ],
+ "category_id": 9,
+ "area": 1214304,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10178,
+ "image_id": 1875,
+ "bbox": [
+ 371,
+ 417,
+ 48,
+ 60
+ ],
+ "category_id": 6,
+ "area": 10370,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10179,
+ "image_id": 1875,
+ "bbox": [
+ 297,
+ 185,
+ 90,
+ 114
+ ],
+ "category_id": 8,
+ "area": 36160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10180,
+ "image_id": 1875,
+ "bbox": [
+ 187,
+ 254,
+ 90,
+ 73
+ ],
+ "category_id": 8,
+ "area": 23381,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10181,
+ "image_id": 1876,
+ "bbox": [
+ 61,
+ 229,
+ 189,
+ 201
+ ],
+ "category_id": 6,
+ "area": 24500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10183,
+ "image_id": 1877,
+ "bbox": [
+ 8,
+ 225,
+ 407,
+ 286
+ ],
+ "category_id": 9,
+ "area": 236628,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10184,
+ "image_id": 1877,
+ "bbox": [
+ 413,
+ 18,
+ 98,
+ 178
+ ],
+ "category_id": 6,
+ "area": 35685,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10216,
+ "image_id": 1881,
+ "bbox": [
+ 207,
+ 90,
+ 92,
+ 166
+ ],
+ "category_id": 10,
+ "area": 21717,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10247,
+ "image_id": 1884,
+ "bbox": [
+ 0,
+ 0,
+ 390,
+ 511
+ ],
+ "category_id": 6,
+ "area": 281940,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10249,
+ "image_id": 1885,
+ "bbox": [
+ 76,
+ 136,
+ 282,
+ 232
+ ],
+ "category_id": 8,
+ "area": 519400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10251,
+ "image_id": 1885,
+ "bbox": [
+ 254,
+ 146,
+ 257,
+ 363
+ ],
+ "category_id": 6,
+ "area": 740922,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10252,
+ "image_id": 1886,
+ "bbox": [
+ 0,
+ 68,
+ 58,
+ 73
+ ],
+ "category_id": 6,
+ "area": 90880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10253,
+ "image_id": 1886,
+ "bbox": [
+ 0,
+ 114,
+ 42,
+ 66
+ ],
+ "category_id": 6,
+ "area": 59881,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10254,
+ "image_id": 1886,
+ "bbox": [
+ 1,
+ 179,
+ 65,
+ 91
+ ],
+ "category_id": 6,
+ "area": 127050,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10255,
+ "image_id": 1886,
+ "bbox": [
+ 3,
+ 290,
+ 47,
+ 88
+ ],
+ "category_id": 6,
+ "area": 89342,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10256,
+ "image_id": 1886,
+ "bbox": [
+ 8,
+ 417,
+ 55,
+ 74
+ ],
+ "category_id": 6,
+ "area": 88704,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10257,
+ "image_id": 1886,
+ "bbox": [
+ 85,
+ 385,
+ 83,
+ 114
+ ],
+ "category_id": 6,
+ "area": 201062,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10258,
+ "image_id": 1886,
+ "bbox": [
+ 115,
+ 230,
+ 49,
+ 46
+ ],
+ "category_id": 6,
+ "area": 48780,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10259,
+ "image_id": 1886,
+ "bbox": [
+ 41,
+ 142,
+ 52,
+ 99
+ ],
+ "category_id": 6,
+ "area": 110016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10260,
+ "image_id": 1886,
+ "bbox": [
+ 67,
+ 66,
+ 68,
+ 66
+ ],
+ "category_id": 6,
+ "area": 97146,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10261,
+ "image_id": 1886,
+ "bbox": [
+ 100,
+ 55,
+ 164,
+ 135
+ ],
+ "category_id": 6,
+ "area": 474498,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10262,
+ "image_id": 1886,
+ "bbox": [
+ 270,
+ 58,
+ 74,
+ 80
+ ],
+ "category_id": 6,
+ "area": 125972,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10263,
+ "image_id": 1886,
+ "bbox": [
+ 325,
+ 65,
+ 86,
+ 176
+ ],
+ "category_id": 6,
+ "area": 322728,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10264,
+ "image_id": 1886,
+ "bbox": [
+ 324,
+ 229,
+ 67,
+ 74
+ ],
+ "category_id": 6,
+ "area": 106678,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10265,
+ "image_id": 1886,
+ "bbox": [
+ 397,
+ 0,
+ 69,
+ 119
+ ],
+ "category_id": 6,
+ "area": 176788,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10266,
+ "image_id": 1886,
+ "bbox": [
+ 412,
+ 96,
+ 24,
+ 66
+ ],
+ "category_id": 6,
+ "area": 35072,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10275,
+ "image_id": 1887,
+ "bbox": [
+ 3,
+ 116,
+ 250,
+ 134
+ ],
+ "category_id": 8,
+ "area": 85575,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10276,
+ "image_id": 1887,
+ "bbox": [
+ 88,
+ 128,
+ 309,
+ 345
+ ],
+ "category_id": 8,
+ "area": 272250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10277,
+ "image_id": 1887,
+ "bbox": [
+ 323,
+ 251,
+ 167,
+ 186
+ ],
+ "category_id": 8,
+ "area": 79461,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10296,
+ "image_id": 1890,
+ "bbox": [
+ 241,
+ 0,
+ 88,
+ 89
+ ],
+ "category_id": 10,
+ "area": 13612,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10297,
+ "image_id": 1890,
+ "bbox": [
+ 295,
+ 19,
+ 82,
+ 208
+ ],
+ "category_id": 10,
+ "area": 29529,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10298,
+ "image_id": 1890,
+ "bbox": [
+ 389,
+ 294,
+ 93,
+ 202
+ ],
+ "category_id": 10,
+ "area": 32712,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10299,
+ "image_id": 1890,
+ "bbox": [
+ 132,
+ 183,
+ 100,
+ 230
+ ],
+ "category_id": 10,
+ "area": 40018,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10300,
+ "image_id": 1890,
+ "bbox": [
+ 149,
+ 435,
+ 77,
+ 76
+ ],
+ "category_id": 10,
+ "area": 10224,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10312,
+ "image_id": 1894,
+ "bbox": [
+ 194,
+ 219,
+ 94,
+ 97
+ ],
+ "category_id": 8,
+ "area": 32195,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10313,
+ "image_id": 1894,
+ "bbox": [
+ 269,
+ 238,
+ 82,
+ 71
+ ],
+ "category_id": 8,
+ "area": 20600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10315,
+ "image_id": 1895,
+ "bbox": [
+ 164,
+ 43,
+ 175,
+ 261
+ ],
+ "category_id": 10,
+ "area": 115596,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10316,
+ "image_id": 1895,
+ "bbox": [
+ 25,
+ 224,
+ 92,
+ 127
+ ],
+ "category_id": 10,
+ "area": 29684,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10317,
+ "image_id": 1895,
+ "bbox": [
+ 79,
+ 431,
+ 68,
+ 79
+ ],
+ "category_id": 10,
+ "area": 13699,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10318,
+ "image_id": 1895,
+ "bbox": [
+ 307,
+ 279,
+ 53,
+ 84
+ ],
+ "category_id": 10,
+ "area": 11445,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10319,
+ "image_id": 1895,
+ "bbox": [
+ 211,
+ 347,
+ 52,
+ 72
+ ],
+ "category_id": 10,
+ "area": 9588,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10320,
+ "image_id": 1895,
+ "bbox": [
+ 419,
+ 103,
+ 42,
+ 53
+ ],
+ "category_id": 10,
+ "area": 5658,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10321,
+ "image_id": 1895,
+ "bbox": [
+ 405,
+ 180,
+ 36,
+ 48
+ ],
+ "category_id": 10,
+ "area": 4402,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10322,
+ "image_id": 1895,
+ "bbox": [
+ 51,
+ 114,
+ 40,
+ 44
+ ],
+ "category_id": 10,
+ "area": 4446,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10323,
+ "image_id": 1895,
+ "bbox": [
+ 56,
+ 0,
+ 40,
+ 49
+ ],
+ "category_id": 10,
+ "area": 4992,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10324,
+ "image_id": 1895,
+ "bbox": [
+ 375,
+ 0,
+ 53,
+ 52
+ ],
+ "category_id": 10,
+ "area": 7072,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10325,
+ "image_id": 1895,
+ "bbox": [
+ 460,
+ 4,
+ 38,
+ 48
+ ],
+ "category_id": 10,
+ "area": 4650,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10326,
+ "image_id": 1895,
+ "bbox": [
+ 296,
+ 388,
+ 43,
+ 70
+ ],
+ "category_id": 10,
+ "area": 7735,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10327,
+ "image_id": 1895,
+ "bbox": [
+ 172,
+ 444,
+ 62,
+ 54
+ ],
+ "category_id": 10,
+ "area": 8540,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10328,
+ "image_id": 1895,
+ "bbox": [
+ 469,
+ 366,
+ 40,
+ 58
+ ],
+ "category_id": 10,
+ "area": 6004,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10374,
+ "image_id": 1897,
+ "bbox": [
+ 73,
+ 0,
+ 397,
+ 471
+ ],
+ "category_id": 10,
+ "area": 1483048,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10375,
+ "image_id": 1898,
+ "bbox": [
+ 159,
+ 177,
+ 325,
+ 333
+ ],
+ "category_id": 8,
+ "area": 478969,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10376,
+ "image_id": 1898,
+ "bbox": [
+ 106,
+ 297,
+ 145,
+ 213
+ ],
+ "category_id": 8,
+ "area": 136920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10377,
+ "image_id": 1899,
+ "bbox": [
+ 83,
+ 217,
+ 134,
+ 290
+ ],
+ "category_id": 9,
+ "area": 106260,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10378,
+ "image_id": 1899,
+ "bbox": [
+ 303,
+ 159,
+ 128,
+ 145
+ ],
+ "category_id": 10,
+ "area": 51062,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10384,
+ "image_id": 1900,
+ "bbox": [
+ 0,
+ 0,
+ 294,
+ 174
+ ],
+ "category_id": 6,
+ "area": 127908,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10385,
+ "image_id": 1900,
+ "bbox": [
+ 476,
+ 210,
+ 31,
+ 55
+ ],
+ "category_id": 6,
+ "area": 4320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10386,
+ "image_id": 1900,
+ "bbox": [
+ 432,
+ 0,
+ 79,
+ 111
+ ],
+ "category_id": 6,
+ "area": 21895,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10387,
+ "image_id": 1900,
+ "bbox": [
+ 484,
+ 132,
+ 23,
+ 39
+ ],
+ "category_id": 6,
+ "area": 2295,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10388,
+ "image_id": 1901,
+ "bbox": [
+ 1,
+ 213,
+ 509,
+ 271
+ ],
+ "category_id": 9,
+ "area": 89142,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10391,
+ "image_id": 1903,
+ "bbox": [
+ 76,
+ 222,
+ 66,
+ 57
+ ],
+ "category_id": 8,
+ "area": 30622,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10392,
+ "image_id": 1903,
+ "bbox": [
+ 155,
+ 230,
+ 177,
+ 281
+ ],
+ "category_id": 8,
+ "area": 394938,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10393,
+ "image_id": 1903,
+ "bbox": [
+ 74,
+ 453,
+ 191,
+ 58
+ ],
+ "category_id": 8,
+ "area": 89032,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10394,
+ "image_id": 1904,
+ "bbox": [
+ 39,
+ 88,
+ 395,
+ 355
+ ],
+ "category_id": 9,
+ "area": 3021957,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10411,
+ "image_id": 1908,
+ "bbox": [
+ 217,
+ 219,
+ 294,
+ 218
+ ],
+ "category_id": 10,
+ "area": 95872,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10412,
+ "image_id": 1909,
+ "bbox": [
+ 48,
+ 75,
+ 323,
+ 435
+ ],
+ "category_id": 8,
+ "area": 495304,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10413,
+ "image_id": 1910,
+ "bbox": [
+ 139,
+ 142,
+ 371,
+ 276
+ ],
+ "category_id": 9,
+ "area": 240368,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10414,
+ "image_id": 1910,
+ "bbox": [
+ 0,
+ 0,
+ 76,
+ 87
+ ],
+ "category_id": 6,
+ "area": 15645,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10415,
+ "image_id": 1910,
+ "bbox": [
+ 0,
+ 396,
+ 151,
+ 115
+ ],
+ "category_id": 6,
+ "area": 40710,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10416,
+ "image_id": 1911,
+ "bbox": [
+ 106,
+ 96,
+ 365,
+ 273
+ ],
+ "category_id": 9,
+ "area": 351505,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10421,
+ "image_id": 1912,
+ "bbox": [
+ 59,
+ 0,
+ 452,
+ 245
+ ],
+ "category_id": 6,
+ "area": 275842,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10422,
+ "image_id": 1913,
+ "bbox": [
+ 65,
+ 138,
+ 245,
+ 276
+ ],
+ "category_id": 9,
+ "area": 189880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10423,
+ "image_id": 1913,
+ "bbox": [
+ 216,
+ 247,
+ 249,
+ 244
+ ],
+ "category_id": 6,
+ "area": 170646,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10426,
+ "image_id": 1915,
+ "bbox": [
+ 155,
+ 107,
+ 253,
+ 356
+ ],
+ "category_id": 9,
+ "area": 263840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10441,
+ "image_id": 1917,
+ "bbox": [
+ 14,
+ 156,
+ 199,
+ 284
+ ],
+ "category_id": 6,
+ "area": 438165,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10442,
+ "image_id": 1917,
+ "bbox": [
+ 379,
+ 78,
+ 71,
+ 125
+ ],
+ "category_id": 6,
+ "area": 68886,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10443,
+ "image_id": 1917,
+ "bbox": [
+ 443,
+ 271,
+ 47,
+ 86
+ ],
+ "category_id": 6,
+ "area": 31506,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10444,
+ "image_id": 1917,
+ "bbox": [
+ 105,
+ 84,
+ 183,
+ 170
+ ],
+ "category_id": 6,
+ "area": 241824,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10445,
+ "image_id": 1917,
+ "bbox": [
+ 207,
+ 50,
+ 304,
+ 461
+ ],
+ "category_id": 6,
+ "area": 1083000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10446,
+ "image_id": 1917,
+ "bbox": [
+ 335,
+ 21,
+ 176,
+ 134
+ ],
+ "category_id": 6,
+ "area": 182988,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10466,
+ "image_id": 1918,
+ "bbox": [
+ 228,
+ 327,
+ 52,
+ 76
+ ],
+ "category_id": 6,
+ "area": 6003,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10467,
+ "image_id": 1918,
+ "bbox": [
+ 84,
+ 435,
+ 78,
+ 76
+ ],
+ "category_id": 6,
+ "area": 8970,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10468,
+ "image_id": 1918,
+ "bbox": [
+ 0,
+ 422,
+ 66,
+ 89
+ ],
+ "category_id": 6,
+ "area": 8910,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10469,
+ "image_id": 1918,
+ "bbox": [
+ 348,
+ 216,
+ 94,
+ 121
+ ],
+ "category_id": 6,
+ "area": 17160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10470,
+ "image_id": 1918,
+ "bbox": [
+ 449,
+ 345,
+ 62,
+ 134
+ ],
+ "category_id": 6,
+ "area": 12566,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10471,
+ "image_id": 1918,
+ "bbox": [
+ 419,
+ 220,
+ 92,
+ 130
+ ],
+ "category_id": 6,
+ "area": 18054,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10472,
+ "image_id": 1918,
+ "bbox": [
+ 276,
+ 149,
+ 61,
+ 86
+ ],
+ "category_id": 6,
+ "area": 7878,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10473,
+ "image_id": 1918,
+ "bbox": [
+ 204,
+ 259,
+ 55,
+ 89
+ ],
+ "category_id": 6,
+ "area": 7452,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10477,
+ "image_id": 1919,
+ "bbox": [
+ 144,
+ 47,
+ 227,
+ 405
+ ],
+ "category_id": 10,
+ "area": 730168,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10481,
+ "image_id": 1921,
+ "bbox": [
+ 158,
+ 254,
+ 84,
+ 93
+ ],
+ "category_id": 8,
+ "area": 27772,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10482,
+ "image_id": 1921,
+ "bbox": [
+ 226,
+ 269,
+ 93,
+ 74
+ ],
+ "category_id": 8,
+ "area": 24465,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10488,
+ "image_id": 1923,
+ "bbox": [
+ 0,
+ 59,
+ 459,
+ 295
+ ],
+ "category_id": 8,
+ "area": 437326,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10489,
+ "image_id": 1923,
+ "bbox": [
+ 358,
+ 4,
+ 153,
+ 154
+ ],
+ "category_id": 8,
+ "area": 76245,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10493,
+ "image_id": 1924,
+ "bbox": [
+ 107,
+ 0,
+ 344,
+ 431
+ ],
+ "category_id": 6,
+ "area": 373760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10494,
+ "image_id": 1924,
+ "bbox": [
+ 285,
+ 334,
+ 226,
+ 177
+ ],
+ "category_id": 9,
+ "area": 100800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10495,
+ "image_id": 1924,
+ "bbox": [
+ 0,
+ 249,
+ 343,
+ 262
+ ],
+ "category_id": 6,
+ "area": 226408,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10496,
+ "image_id": 1925,
+ "bbox": [
+ 167,
+ 206,
+ 344,
+ 273
+ ],
+ "category_id": 8,
+ "area": 329763,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10497,
+ "image_id": 1925,
+ "bbox": [
+ 0,
+ 0,
+ 233,
+ 510
+ ],
+ "category_id": 6,
+ "area": 417560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10508,
+ "image_id": 1927,
+ "bbox": [
+ 203,
+ 0,
+ 304,
+ 505
+ ],
+ "category_id": 9,
+ "area": 541782,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10509,
+ "image_id": 1928,
+ "bbox": [
+ 20,
+ 33,
+ 404,
+ 473
+ ],
+ "category_id": 8,
+ "area": 1401776,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10534,
+ "image_id": 1931,
+ "bbox": [
+ 60,
+ 3,
+ 110,
+ 107
+ ],
+ "category_id": 6,
+ "area": 22644,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10535,
+ "image_id": 1931,
+ "bbox": [
+ 63,
+ 102,
+ 125,
+ 159
+ ],
+ "category_id": 6,
+ "area": 38280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10536,
+ "image_id": 1931,
+ "bbox": [
+ 178,
+ 299,
+ 146,
+ 208
+ ],
+ "category_id": 6,
+ "area": 58320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10537,
+ "image_id": 1931,
+ "bbox": [
+ 424,
+ 334,
+ 87,
+ 174
+ ],
+ "category_id": 6,
+ "area": 29141,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10538,
+ "image_id": 1932,
+ "bbox": [
+ 19,
+ 15,
+ 471,
+ 469
+ ],
+ "category_id": 6,
+ "area": 164864,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10539,
+ "image_id": 1933,
+ "bbox": [
+ 89,
+ 255,
+ 293,
+ 202
+ ],
+ "category_id": 8,
+ "area": 409077,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10540,
+ "image_id": 1933,
+ "bbox": [
+ 214,
+ 35,
+ 296,
+ 384
+ ],
+ "category_id": 8,
+ "area": 785463,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10541,
+ "image_id": 1933,
+ "bbox": [
+ 178,
+ 158,
+ 201,
+ 162
+ ],
+ "category_id": 6,
+ "area": 225455,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10542,
+ "image_id": 1934,
+ "bbox": [
+ 94,
+ 87,
+ 245,
+ 424
+ ],
+ "category_id": 10,
+ "area": 173745,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10552,
+ "image_id": 1936,
+ "bbox": [
+ 32,
+ 118,
+ 429,
+ 393
+ ],
+ "category_id": 9,
+ "area": 417175,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10553,
+ "image_id": 1936,
+ "bbox": [
+ 381,
+ 151,
+ 79,
+ 141
+ ],
+ "category_id": 6,
+ "area": 27702,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10555,
+ "image_id": 1936,
+ "bbox": [
+ 387,
+ 351,
+ 39,
+ 67
+ ],
+ "category_id": 6,
+ "area": 6545,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10556,
+ "image_id": 1937,
+ "bbox": [
+ 262,
+ 198,
+ 165,
+ 193
+ ],
+ "category_id": 6,
+ "area": 46116,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10557,
+ "image_id": 1937,
+ "bbox": [
+ 256,
+ 367,
+ 209,
+ 144
+ ],
+ "category_id": 9,
+ "area": 43428,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10563,
+ "image_id": 1940,
+ "bbox": [
+ 121,
+ 139,
+ 388,
+ 368
+ ],
+ "category_id": 9,
+ "area": 502978,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10564,
+ "image_id": 1940,
+ "bbox": [
+ 233,
+ 32,
+ 112,
+ 177
+ ],
+ "category_id": 9,
+ "area": 69969,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10565,
+ "image_id": 1940,
+ "bbox": [
+ 288,
+ 7,
+ 222,
+ 500
+ ],
+ "category_id": 9,
+ "area": 392128,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10615,
+ "image_id": 1942,
+ "bbox": [
+ 168,
+ 95,
+ 342,
+ 415
+ ],
+ "category_id": 10,
+ "area": 1126536,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10616,
+ "image_id": 1943,
+ "bbox": [
+ 100,
+ 52,
+ 387,
+ 314
+ ],
+ "category_id": 9,
+ "area": 210912,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10653,
+ "image_id": 1946,
+ "bbox": [
+ 309,
+ 257,
+ 105,
+ 98
+ ],
+ "category_id": 6,
+ "area": 31017,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10654,
+ "image_id": 1946,
+ "bbox": [
+ 433,
+ 218,
+ 78,
+ 152
+ ],
+ "category_id": 6,
+ "area": 35953,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10655,
+ "image_id": 1946,
+ "bbox": [
+ 161,
+ 60,
+ 193,
+ 80
+ ],
+ "category_id": 6,
+ "area": 46827,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10656,
+ "image_id": 1946,
+ "bbox": [
+ 0,
+ 72,
+ 197,
+ 173
+ ],
+ "category_id": 6,
+ "area": 102440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10657,
+ "image_id": 1946,
+ "bbox": [
+ 263,
+ 104,
+ 248,
+ 153
+ ],
+ "category_id": 6,
+ "area": 114080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10658,
+ "image_id": 1946,
+ "bbox": [
+ 0,
+ 377,
+ 120,
+ 133
+ ],
+ "category_id": 6,
+ "area": 48200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10659,
+ "image_id": 1946,
+ "bbox": [
+ 8,
+ 280,
+ 74,
+ 112
+ ],
+ "category_id": 6,
+ "area": 24864,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10660,
+ "image_id": 1946,
+ "bbox": [
+ 77,
+ 229,
+ 62,
+ 78
+ ],
+ "category_id": 6,
+ "area": 14508,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10661,
+ "image_id": 1946,
+ "bbox": [
+ 73,
+ 326,
+ 69,
+ 63
+ ],
+ "category_id": 6,
+ "area": 13110,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10694,
+ "image_id": 1950,
+ "bbox": [
+ 44,
+ 251,
+ 260,
+ 245
+ ],
+ "category_id": 9,
+ "area": 130988,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10695,
+ "image_id": 1951,
+ "bbox": [
+ 185,
+ 0,
+ 156,
+ 208
+ ],
+ "category_id": 10,
+ "area": 58200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10696,
+ "image_id": 1951,
+ "bbox": [
+ 21,
+ 0,
+ 146,
+ 174
+ ],
+ "category_id": 10,
+ "area": 45696,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10697,
+ "image_id": 1951,
+ "bbox": [
+ 77,
+ 107,
+ 172,
+ 331
+ ],
+ "category_id": 10,
+ "area": 102399,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10698,
+ "image_id": 1951,
+ "bbox": [
+ 198,
+ 240,
+ 175,
+ 271
+ ],
+ "category_id": 10,
+ "area": 85086,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10699,
+ "image_id": 1951,
+ "bbox": [
+ 320,
+ 135,
+ 171,
+ 306
+ ],
+ "category_id": 10,
+ "area": 94105,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10700,
+ "image_id": 1951,
+ "bbox": [
+ 460,
+ 282,
+ 49,
+ 208
+ ],
+ "category_id": 10,
+ "area": 18400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10701,
+ "image_id": 1951,
+ "bbox": [
+ 0,
+ 137,
+ 69,
+ 289
+ ],
+ "category_id": 10,
+ "area": 36140,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10707,
+ "image_id": 1953,
+ "bbox": [
+ 0,
+ 222,
+ 256,
+ 287
+ ],
+ "category_id": 9,
+ "area": 307680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10727,
+ "image_id": 1955,
+ "bbox": [
+ 122,
+ 113,
+ 355,
+ 244
+ ],
+ "category_id": 9,
+ "area": 305472,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10745,
+ "image_id": 1957,
+ "bbox": [
+ 138,
+ 192,
+ 144,
+ 154
+ ],
+ "category_id": 8,
+ "area": 77976,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10746,
+ "image_id": 1957,
+ "bbox": [
+ 324,
+ 145,
+ 174,
+ 264
+ ],
+ "category_id": 8,
+ "area": 161320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10747,
+ "image_id": 1957,
+ "bbox": [
+ 405,
+ 119,
+ 31,
+ 38
+ ],
+ "category_id": 6,
+ "area": 4266,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10748,
+ "image_id": 1957,
+ "bbox": [
+ 441,
+ 163,
+ 39,
+ 47
+ ],
+ "category_id": 6,
+ "area": 6534,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10749,
+ "image_id": 1957,
+ "bbox": [
+ 481,
+ 202,
+ 28,
+ 29
+ ],
+ "category_id": 6,
+ "area": 2870,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10785,
+ "image_id": 1960,
+ "bbox": [
+ 246,
+ 0,
+ 265,
+ 235
+ ],
+ "category_id": 6,
+ "area": 202880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10786,
+ "image_id": 1960,
+ "bbox": [
+ 215,
+ 166,
+ 296,
+ 325
+ ],
+ "category_id": 9,
+ "area": 312494,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10797,
+ "image_id": 1963,
+ "bbox": [
+ 168,
+ 206,
+ 166,
+ 228
+ ],
+ "category_id": 9,
+ "area": 57120,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10798,
+ "image_id": 1963,
+ "bbox": [
+ 95,
+ 210,
+ 59,
+ 96
+ ],
+ "category_id": 6,
+ "area": 8633,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10799,
+ "image_id": 1964,
+ "bbox": [
+ 30,
+ 198,
+ 278,
+ 289
+ ],
+ "category_id": 8,
+ "area": 205088,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10800,
+ "image_id": 1964,
+ "bbox": [
+ 239,
+ 36,
+ 212,
+ 261
+ ],
+ "category_id": 8,
+ "area": 141100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10812,
+ "image_id": 1969,
+ "bbox": [
+ 102,
+ 182,
+ 256,
+ 329
+ ],
+ "category_id": 10,
+ "area": 257580,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10813,
+ "image_id": 1969,
+ "bbox": [
+ 262,
+ 115,
+ 37,
+ 61
+ ],
+ "category_id": 10,
+ "area": 7120,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10814,
+ "image_id": 1970,
+ "bbox": [
+ 2,
+ 133,
+ 218,
+ 171
+ ],
+ "category_id": 8,
+ "area": 92984,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10815,
+ "image_id": 1970,
+ "bbox": [
+ 170,
+ 165,
+ 146,
+ 211
+ ],
+ "category_id": 8,
+ "area": 76824,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10816,
+ "image_id": 1970,
+ "bbox": [
+ 263,
+ 214,
+ 204,
+ 186
+ ],
+ "category_id": 8,
+ "area": 94208,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10817,
+ "image_id": 1971,
+ "bbox": [
+ 28,
+ 276,
+ 475,
+ 235
+ ],
+ "category_id": 9,
+ "area": 240084,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10818,
+ "image_id": 1971,
+ "bbox": [
+ 34,
+ 3,
+ 476,
+ 508
+ ],
+ "category_id": 6,
+ "area": 519357,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10819,
+ "image_id": 1971,
+ "bbox": [
+ 0,
+ 367,
+ 64,
+ 144
+ ],
+ "category_id": 6,
+ "area": 20099,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10820,
+ "image_id": 1972,
+ "bbox": [
+ 188,
+ 193,
+ 275,
+ 202
+ ],
+ "category_id": 9,
+ "area": 351714,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10821,
+ "image_id": 1973,
+ "bbox": [
+ 3,
+ 0,
+ 505,
+ 507
+ ],
+ "category_id": 9,
+ "area": 901232,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10822,
+ "image_id": 1974,
+ "bbox": [
+ 167,
+ 112,
+ 275,
+ 394
+ ],
+ "category_id": 9,
+ "area": 382395,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10823,
+ "image_id": 1974,
+ "bbox": [
+ 235,
+ 93,
+ 228,
+ 384
+ ],
+ "category_id": 9,
+ "area": 308880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10825,
+ "image_id": 1975,
+ "bbox": [
+ 259,
+ 236,
+ 229,
+ 275
+ ],
+ "category_id": 9,
+ "area": 83424,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10833,
+ "image_id": 1978,
+ "bbox": [
+ 205,
+ 220,
+ 261,
+ 271
+ ],
+ "category_id": 9,
+ "area": 93834,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10837,
+ "image_id": 1980,
+ "bbox": [
+ 121,
+ 108,
+ 191,
+ 287
+ ],
+ "category_id": 8,
+ "area": 1744518,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10849,
+ "image_id": 1984,
+ "bbox": [
+ 53,
+ 93,
+ 320,
+ 418
+ ],
+ "category_id": 8,
+ "area": 471200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10852,
+ "image_id": 1985,
+ "bbox": [
+ 251,
+ 266,
+ 169,
+ 202
+ ],
+ "category_id": 9,
+ "area": 20295,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10870,
+ "image_id": 1987,
+ "bbox": [
+ 2,
+ 0,
+ 509,
+ 255
+ ],
+ "category_id": 6,
+ "area": 326256,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10871,
+ "image_id": 1988,
+ "bbox": [
+ 90,
+ 20,
+ 384,
+ 455
+ ],
+ "category_id": 6,
+ "area": 1282668,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10878,
+ "image_id": 1990,
+ "bbox": [
+ 211,
+ 399,
+ 35,
+ 52
+ ],
+ "category_id": 6,
+ "area": 12321,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10896,
+ "image_id": 1993,
+ "bbox": [
+ 0,
+ 290,
+ 511,
+ 220
+ ],
+ "category_id": 6,
+ "area": 223020,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10897,
+ "image_id": 1993,
+ "bbox": [
+ 0,
+ 1,
+ 511,
+ 320
+ ],
+ "category_id": 6,
+ "area": 323910,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10899,
+ "image_id": 1995,
+ "bbox": [
+ 171,
+ 204,
+ 340,
+ 306
+ ],
+ "category_id": 10,
+ "area": 825572,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10901,
+ "image_id": 1997,
+ "bbox": [
+ 136,
+ 98,
+ 111,
+ 364
+ ],
+ "category_id": 8,
+ "area": 142336,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10902,
+ "image_id": 1997,
+ "bbox": [
+ 227,
+ 114,
+ 148,
+ 213
+ ],
+ "category_id": 8,
+ "area": 111600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10905,
+ "image_id": 1998,
+ "bbox": [
+ 241,
+ 165,
+ 39,
+ 92
+ ],
+ "category_id": 10,
+ "area": 8640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10910,
+ "image_id": 2000,
+ "bbox": [
+ 0,
+ 0,
+ 512,
+ 400
+ ],
+ "category_id": 10,
+ "area": 1620480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10911,
+ "image_id": 2001,
+ "bbox": [
+ 65,
+ 55,
+ 310,
+ 356
+ ],
+ "category_id": 10,
+ "area": 164628,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10912,
+ "image_id": 2002,
+ "bbox": [
+ 20,
+ 276,
+ 49,
+ 29
+ ],
+ "category_id": 8,
+ "area": 11718,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10913,
+ "image_id": 2002,
+ "bbox": [
+ 164,
+ 259,
+ 130,
+ 101
+ ],
+ "category_id": 8,
+ "area": 105350,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10914,
+ "image_id": 2003,
+ "bbox": [
+ 288,
+ 233,
+ 152,
+ 278
+ ],
+ "category_id": 6,
+ "area": 109557,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10915,
+ "image_id": 2003,
+ "bbox": [
+ 98,
+ 82,
+ 321,
+ 302
+ ],
+ "category_id": 9,
+ "area": 250173,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10917,
+ "image_id": 2004,
+ "bbox": [
+ 312,
+ 126,
+ 199,
+ 343
+ ],
+ "category_id": 6,
+ "area": 80178,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10918,
+ "image_id": 2004,
+ "bbox": [
+ 144,
+ 258,
+ 206,
+ 252
+ ],
+ "category_id": 6,
+ "area": 61146,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10936,
+ "image_id": 2006,
+ "bbox": [
+ 149,
+ 10,
+ 261,
+ 453
+ ],
+ "category_id": 8,
+ "area": 317730,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10948,
+ "image_id": 2008,
+ "bbox": [
+ 53,
+ 150,
+ 458,
+ 360
+ ],
+ "category_id": 8,
+ "area": 1140505,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10949,
+ "image_id": 2008,
+ "bbox": [
+ 138,
+ 404,
+ 229,
+ 107
+ ],
+ "category_id": 8,
+ "area": 169728,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10950,
+ "image_id": 2008,
+ "bbox": [
+ 244,
+ 34,
+ 242,
+ 359
+ ],
+ "category_id": 8,
+ "area": 600210,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10991,
+ "image_id": 2010,
+ "bbox": [
+ 243,
+ 194,
+ 171,
+ 176
+ ],
+ "category_id": 9,
+ "area": 42354,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10995,
+ "image_id": 2012,
+ "bbox": [
+ 86,
+ 39,
+ 138,
+ 128
+ ],
+ "category_id": 8,
+ "area": 43248,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 10996,
+ "image_id": 2012,
+ "bbox": [
+ 67,
+ 101,
+ 378,
+ 356
+ ],
+ "category_id": 8,
+ "area": 329187,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11000,
+ "image_id": 2014,
+ "bbox": [
+ 81,
+ 0,
+ 430,
+ 417
+ ],
+ "category_id": 10,
+ "area": 295812,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11013,
+ "image_id": 2017,
+ "bbox": [
+ 117,
+ 124,
+ 301,
+ 236
+ ],
+ "category_id": 9,
+ "area": 751450,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11014,
+ "image_id": 2018,
+ "bbox": [
+ 0,
+ 77,
+ 349,
+ 376
+ ],
+ "category_id": 9,
+ "area": 327264,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11015,
+ "image_id": 2018,
+ "bbox": [
+ 294,
+ 129,
+ 14,
+ 90
+ ],
+ "category_id": 6,
+ "area": 3159,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11016,
+ "image_id": 2018,
+ "bbox": [
+ 442,
+ 261,
+ 68,
+ 250
+ ],
+ "category_id": 6,
+ "area": 42768,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11017,
+ "image_id": 2019,
+ "bbox": [
+ 0,
+ 59,
+ 210,
+ 167
+ ],
+ "category_id": 9,
+ "area": 123900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11018,
+ "image_id": 2019,
+ "bbox": [
+ 198,
+ 258,
+ 113,
+ 224
+ ],
+ "category_id": 10,
+ "area": 89744,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11019,
+ "image_id": 2019,
+ "bbox": [
+ 266,
+ 176,
+ 170,
+ 138
+ ],
+ "category_id": 9,
+ "area": 83265,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11057,
+ "image_id": 2022,
+ "bbox": [
+ 0,
+ 0,
+ 114,
+ 510
+ ],
+ "category_id": 6,
+ "area": 205205,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11058,
+ "image_id": 2022,
+ "bbox": [
+ 182,
+ 241,
+ 262,
+ 156
+ ],
+ "category_id": 8,
+ "area": 143664,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11066,
+ "image_id": 2025,
+ "bbox": [
+ 74,
+ 47,
+ 375,
+ 448
+ ],
+ "category_id": 8,
+ "area": 237839,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11070,
+ "image_id": 2027,
+ "bbox": [
+ 100,
+ 324,
+ 96,
+ 133
+ ],
+ "category_id": 10,
+ "area": 45120,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11072,
+ "image_id": 2027,
+ "bbox": [
+ 310,
+ 0,
+ 91,
+ 113
+ ],
+ "category_id": 10,
+ "area": 36480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11074,
+ "image_id": 2028,
+ "bbox": [
+ 59,
+ 0,
+ 270,
+ 130
+ ],
+ "category_id": 8,
+ "area": 29832,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11078,
+ "image_id": 2030,
+ "bbox": [
+ 26,
+ 94,
+ 66,
+ 87
+ ],
+ "category_id": 8,
+ "area": 39432,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11079,
+ "image_id": 2030,
+ "bbox": [
+ 114,
+ 222,
+ 108,
+ 123
+ ],
+ "category_id": 8,
+ "area": 90944,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11080,
+ "image_id": 2030,
+ "bbox": [
+ 249,
+ 89,
+ 20,
+ 45
+ ],
+ "category_id": 8,
+ "area": 6308,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11081,
+ "image_id": 2030,
+ "bbox": [
+ 198,
+ 194,
+ 60,
+ 170
+ ],
+ "category_id": 8,
+ "area": 70452,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11082,
+ "image_id": 2030,
+ "bbox": [
+ 315,
+ 192,
+ 86,
+ 123
+ ],
+ "category_id": 8,
+ "area": 72450,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11083,
+ "image_id": 2030,
+ "bbox": [
+ 394,
+ 193,
+ 28,
+ 70
+ ],
+ "category_id": 8,
+ "area": 13824,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11096,
+ "image_id": 2032,
+ "bbox": [
+ 61,
+ 279,
+ 412,
+ 232
+ ],
+ "category_id": 9,
+ "area": 161028,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11097,
+ "image_id": 2032,
+ "bbox": [
+ 0,
+ 304,
+ 195,
+ 207
+ ],
+ "category_id": 6,
+ "area": 68175,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11103,
+ "image_id": 2034,
+ "bbox": [
+ 0,
+ 185,
+ 285,
+ 210
+ ],
+ "category_id": 8,
+ "area": 477040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11104,
+ "image_id": 2034,
+ "bbox": [
+ 150,
+ 136,
+ 278,
+ 226
+ ],
+ "category_id": 8,
+ "area": 499988,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11105,
+ "image_id": 2034,
+ "bbox": [
+ 364,
+ 328,
+ 40,
+ 59
+ ],
+ "category_id": 8,
+ "area": 19000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11106,
+ "image_id": 2034,
+ "bbox": [
+ 388,
+ 341,
+ 21,
+ 21
+ ],
+ "category_id": 8,
+ "area": 3690,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11107,
+ "image_id": 2035,
+ "bbox": [
+ 59,
+ 0,
+ 409,
+ 389
+ ],
+ "category_id": 9,
+ "area": 560604,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11108,
+ "image_id": 2035,
+ "bbox": [
+ 273,
+ 90,
+ 235,
+ 418
+ ],
+ "category_id": 9,
+ "area": 345744,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11109,
+ "image_id": 2035,
+ "bbox": [
+ 2,
+ 7,
+ 412,
+ 497
+ ],
+ "category_id": 9,
+ "area": 722400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11112,
+ "image_id": 2036,
+ "bbox": [
+ 67,
+ 312,
+ 275,
+ 169
+ ],
+ "category_id": 9,
+ "area": 304794,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11116,
+ "image_id": 2038,
+ "bbox": [
+ 51,
+ 0,
+ 172,
+ 256
+ ],
+ "category_id": 6,
+ "area": 101598,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11117,
+ "image_id": 2038,
+ "bbox": [
+ 0,
+ 201,
+ 40,
+ 147
+ ],
+ "category_id": 6,
+ "area": 13695,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11118,
+ "image_id": 2038,
+ "bbox": [
+ 242,
+ 371,
+ 81,
+ 128
+ ],
+ "category_id": 6,
+ "area": 24024,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11119,
+ "image_id": 2038,
+ "bbox": [
+ 351,
+ 0,
+ 160,
+ 475
+ ],
+ "category_id": 6,
+ "area": 174699,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11120,
+ "image_id": 2038,
+ "bbox": [
+ 259,
+ 13,
+ 100,
+ 147
+ ],
+ "category_id": 6,
+ "area": 34155,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11121,
+ "image_id": 2038,
+ "bbox": [
+ 108,
+ 183,
+ 163,
+ 288
+ ],
+ "category_id": 9,
+ "area": 107870,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11122,
+ "image_id": 2038,
+ "bbox": [
+ 130,
+ 362,
+ 110,
+ 149
+ ],
+ "category_id": 6,
+ "area": 37909,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11135,
+ "image_id": 2042,
+ "bbox": [
+ 46,
+ 379,
+ 90,
+ 93
+ ],
+ "category_id": 6,
+ "area": 23427,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11136,
+ "image_id": 2042,
+ "bbox": [
+ 110,
+ 373,
+ 120,
+ 138
+ ],
+ "category_id": 6,
+ "area": 46487,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11164,
+ "image_id": 2045,
+ "bbox": [
+ 166,
+ 181,
+ 302,
+ 330
+ ],
+ "category_id": 8,
+ "area": 116493,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11213,
+ "image_id": 2050,
+ "bbox": [
+ 138,
+ 222,
+ 67,
+ 98
+ ],
+ "category_id": 10,
+ "area": 15872,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11215,
+ "image_id": 2050,
+ "bbox": [
+ 204,
+ 212,
+ 287,
+ 223
+ ],
+ "category_id": 9,
+ "area": 152775,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11217,
+ "image_id": 2051,
+ "bbox": [
+ 8,
+ 0,
+ 475,
+ 512
+ ],
+ "category_id": 6,
+ "area": 260000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11218,
+ "image_id": 2052,
+ "bbox": [
+ 26,
+ 219,
+ 239,
+ 252
+ ],
+ "category_id": 8,
+ "area": 211447,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11219,
+ "image_id": 2052,
+ "bbox": [
+ 182,
+ 231,
+ 309,
+ 255
+ ],
+ "category_id": 8,
+ "area": 276376,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11221,
+ "image_id": 2053,
+ "bbox": [
+ 0,
+ 188,
+ 510,
+ 321
+ ],
+ "category_id": 8,
+ "area": 607620,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11225,
+ "image_id": 2055,
+ "bbox": [
+ 74,
+ 0,
+ 406,
+ 411
+ ],
+ "category_id": 9,
+ "area": 445788,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11226,
+ "image_id": 2056,
+ "bbox": [
+ 29,
+ 42,
+ 95,
+ 174
+ ],
+ "category_id": 10,
+ "area": 22499,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11227,
+ "image_id": 2056,
+ "bbox": [
+ 201,
+ 178,
+ 104,
+ 160
+ ],
+ "category_id": 10,
+ "area": 22657,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11228,
+ "image_id": 2056,
+ "bbox": [
+ 386,
+ 109,
+ 83,
+ 147
+ ],
+ "category_id": 10,
+ "area": 16768,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11229,
+ "image_id": 2056,
+ "bbox": [
+ 123,
+ 285,
+ 88,
+ 139
+ ],
+ "category_id": 10,
+ "area": 16819,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11230,
+ "image_id": 2056,
+ "bbox": [
+ 391,
+ 321,
+ 94,
+ 122
+ ],
+ "category_id": 10,
+ "area": 15688,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11231,
+ "image_id": 2056,
+ "bbox": [
+ 211,
+ 398,
+ 59,
+ 113
+ ],
+ "category_id": 10,
+ "area": 9114,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11232,
+ "image_id": 2056,
+ "bbox": [
+ 139,
+ 57,
+ 99,
+ 81
+ ],
+ "category_id": 10,
+ "area": 11005,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11233,
+ "image_id": 2056,
+ "bbox": [
+ 224,
+ 23,
+ 74,
+ 89
+ ],
+ "category_id": 10,
+ "area": 9126,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11234,
+ "image_id": 2056,
+ "bbox": [
+ 123,
+ 6,
+ 67,
+ 66
+ ],
+ "category_id": 10,
+ "area": 6090,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11235,
+ "image_id": 2056,
+ "bbox": [
+ 296,
+ 220,
+ 49,
+ 147
+ ],
+ "category_id": 10,
+ "area": 9984,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11236,
+ "image_id": 2056,
+ "bbox": [
+ 15,
+ 63,
+ 108,
+ 106
+ ],
+ "category_id": 10,
+ "area": 15548,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11237,
+ "image_id": 2056,
+ "bbox": [
+ 57,
+ 190,
+ 80,
+ 136
+ ],
+ "category_id": 10,
+ "area": 14750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11238,
+ "image_id": 2056,
+ "bbox": [
+ 342,
+ 80,
+ 79,
+ 187
+ ],
+ "category_id": 10,
+ "area": 20212,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11239,
+ "image_id": 2056,
+ "bbox": [
+ 339,
+ 416,
+ 101,
+ 94
+ ],
+ "category_id": 10,
+ "area": 13038,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11242,
+ "image_id": 2058,
+ "bbox": [
+ 4,
+ 203,
+ 338,
+ 217
+ ],
+ "category_id": 9,
+ "area": 518612,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11244,
+ "image_id": 2059,
+ "bbox": [
+ 3,
+ 205,
+ 328,
+ 300
+ ],
+ "category_id": 8,
+ "area": 679662,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11246,
+ "image_id": 2060,
+ "bbox": [
+ 0,
+ 122,
+ 194,
+ 356
+ ],
+ "category_id": 9,
+ "area": 173040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11248,
+ "image_id": 2061,
+ "bbox": [
+ 107,
+ 337,
+ 191,
+ 106
+ ],
+ "category_id": 8,
+ "area": 15150,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11250,
+ "image_id": 2062,
+ "bbox": [
+ 6,
+ 39,
+ 73,
+ 148
+ ],
+ "category_id": 6,
+ "area": 154762,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11251,
+ "image_id": 2062,
+ "bbox": [
+ 0,
+ 157,
+ 22,
+ 50
+ ],
+ "category_id": 6,
+ "area": 15808,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11252,
+ "image_id": 2062,
+ "bbox": [
+ 16,
+ 161,
+ 54,
+ 56
+ ],
+ "category_id": 6,
+ "area": 43605,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11253,
+ "image_id": 2062,
+ "bbox": [
+ 5,
+ 201,
+ 107,
+ 73
+ ],
+ "category_id": 6,
+ "area": 112554,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11254,
+ "image_id": 2062,
+ "bbox": [
+ 1,
+ 294,
+ 75,
+ 113
+ ],
+ "category_id": 6,
+ "area": 120700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11255,
+ "image_id": 2062,
+ "bbox": [
+ 77,
+ 303,
+ 46,
+ 46
+ ],
+ "category_id": 6,
+ "area": 31020,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11256,
+ "image_id": 2062,
+ "bbox": [
+ 99,
+ 334,
+ 48,
+ 111
+ ],
+ "category_id": 6,
+ "area": 76944,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11257,
+ "image_id": 2062,
+ "bbox": [
+ 155,
+ 358,
+ 56,
+ 100
+ ],
+ "category_id": 6,
+ "area": 80030,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11258,
+ "image_id": 2062,
+ "bbox": [
+ 71,
+ 96,
+ 131,
+ 122
+ ],
+ "category_id": 6,
+ "area": 227424,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11259,
+ "image_id": 2062,
+ "bbox": [
+ 105,
+ 177,
+ 55,
+ 174
+ ],
+ "category_id": 6,
+ "area": 135716,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11260,
+ "image_id": 2062,
+ "bbox": [
+ 155,
+ 216,
+ 74,
+ 110
+ ],
+ "category_id": 6,
+ "area": 116512,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11261,
+ "image_id": 2062,
+ "bbox": [
+ 187,
+ 153,
+ 62,
+ 79
+ ],
+ "category_id": 6,
+ "area": 70210,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11262,
+ "image_id": 2062,
+ "bbox": [
+ 257,
+ 466,
+ 68,
+ 45
+ ],
+ "category_id": 6,
+ "area": 43520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11263,
+ "image_id": 2062,
+ "bbox": [
+ 285,
+ 279,
+ 60,
+ 56
+ ],
+ "category_id": 6,
+ "area": 48222,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11264,
+ "image_id": 2062,
+ "bbox": [
+ 231,
+ 66,
+ 78,
+ 68
+ ],
+ "category_id": 6,
+ "area": 76055,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11275,
+ "image_id": 2063,
+ "bbox": [
+ 288,
+ 319,
+ 169,
+ 108
+ ],
+ "category_id": 9,
+ "area": 36366,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11276,
+ "image_id": 2064,
+ "bbox": [
+ 48,
+ 0,
+ 204,
+ 245
+ ],
+ "category_id": 6,
+ "area": 62608,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11277,
+ "image_id": 2064,
+ "bbox": [
+ 91,
+ 202,
+ 361,
+ 281
+ ],
+ "category_id": 9,
+ "area": 126616,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11283,
+ "image_id": 2066,
+ "bbox": [
+ 130,
+ 93,
+ 294,
+ 287
+ ],
+ "category_id": 8,
+ "area": 67176,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11286,
+ "image_id": 2068,
+ "bbox": [
+ 0,
+ 1,
+ 172,
+ 201
+ ],
+ "category_id": 10,
+ "area": 61886,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11287,
+ "image_id": 2068,
+ "bbox": [
+ 371,
+ 25,
+ 91,
+ 241
+ ],
+ "category_id": 10,
+ "area": 39377,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11288,
+ "image_id": 2068,
+ "bbox": [
+ 232,
+ 259,
+ 113,
+ 245
+ ],
+ "category_id": 10,
+ "area": 49796,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11289,
+ "image_id": 2068,
+ "bbox": [
+ 322,
+ 181,
+ 101,
+ 328
+ ],
+ "category_id": 10,
+ "area": 59408,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11290,
+ "image_id": 2069,
+ "bbox": [
+ 1,
+ 87,
+ 463,
+ 419
+ ],
+ "category_id": 9,
+ "area": 683220,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11291,
+ "image_id": 2069,
+ "bbox": [
+ 216,
+ 0,
+ 160,
+ 168
+ ],
+ "category_id": 9,
+ "area": 95037,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11292,
+ "image_id": 2070,
+ "bbox": [
+ 136,
+ 146,
+ 217,
+ 163
+ ],
+ "category_id": 10,
+ "area": 56160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11300,
+ "image_id": 2072,
+ "bbox": [
+ 91,
+ 115,
+ 317,
+ 238
+ ],
+ "category_id": 9,
+ "area": 158004,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11309,
+ "image_id": 2074,
+ "bbox": [
+ 58,
+ 121,
+ 189,
+ 305
+ ],
+ "category_id": 9,
+ "area": 56282,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11315,
+ "image_id": 2077,
+ "bbox": [
+ 38,
+ 67,
+ 416,
+ 442
+ ],
+ "category_id": 8,
+ "area": 215800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11317,
+ "image_id": 2078,
+ "bbox": [
+ 1,
+ 164,
+ 388,
+ 343
+ ],
+ "category_id": 9,
+ "area": 468993,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11318,
+ "image_id": 2078,
+ "bbox": [
+ 168,
+ 0,
+ 344,
+ 180
+ ],
+ "category_id": 9,
+ "area": 218440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11319,
+ "image_id": 2078,
+ "bbox": [
+ 174,
+ 73,
+ 338,
+ 315
+ ],
+ "category_id": 9,
+ "area": 374335,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11322,
+ "image_id": 2079,
+ "bbox": [
+ 45,
+ 440,
+ 101,
+ 71
+ ],
+ "category_id": 6,
+ "area": 20301,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11323,
+ "image_id": 2079,
+ "bbox": [
+ 128,
+ 435,
+ 108,
+ 75
+ ],
+ "category_id": 6,
+ "area": 22684,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11326,
+ "image_id": 2081,
+ "bbox": [
+ 349,
+ 295,
+ 162,
+ 214
+ ],
+ "category_id": 6,
+ "area": 89600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11327,
+ "image_id": 2081,
+ "bbox": [
+ 88,
+ 130,
+ 318,
+ 211
+ ],
+ "category_id": 9,
+ "area": 173558,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11333,
+ "image_id": 2083,
+ "bbox": [
+ 127,
+ 181,
+ 222,
+ 150
+ ],
+ "category_id": 8,
+ "area": 89661,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11334,
+ "image_id": 2083,
+ "bbox": [
+ 69,
+ 147,
+ 150,
+ 106
+ ],
+ "category_id": 8,
+ "area": 42777,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11337,
+ "image_id": 2084,
+ "bbox": [
+ 33,
+ 163,
+ 367,
+ 215
+ ],
+ "category_id": 9,
+ "area": 422050,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11338,
+ "image_id": 2084,
+ "bbox": [
+ 29,
+ 0,
+ 483,
+ 187
+ ],
+ "category_id": 6,
+ "area": 483966,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11358,
+ "image_id": 2087,
+ "bbox": [
+ 327,
+ 199,
+ 43,
+ 63
+ ],
+ "category_id": 9,
+ "area": 11770,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11380,
+ "image_id": 2092,
+ "bbox": [
+ 204,
+ 243,
+ 202,
+ 127
+ ],
+ "category_id": 8,
+ "area": 142438,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11392,
+ "image_id": 2093,
+ "bbox": [
+ 167,
+ 6,
+ 164,
+ 238
+ ],
+ "category_id": 6,
+ "area": 261468,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11393,
+ "image_id": 2093,
+ "bbox": [
+ 0,
+ 119,
+ 123,
+ 292
+ ],
+ "category_id": 6,
+ "area": 240994,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11395,
+ "image_id": 2094,
+ "bbox": [
+ 180,
+ 76,
+ 299,
+ 435
+ ],
+ "category_id": 8,
+ "area": 4122228,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11396,
+ "image_id": 2095,
+ "bbox": [
+ 126,
+ 271,
+ 171,
+ 110
+ ],
+ "category_id": 8,
+ "area": 149819,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11397,
+ "image_id": 2095,
+ "bbox": [
+ 179,
+ 215,
+ 116,
+ 124
+ ],
+ "category_id": 8,
+ "area": 115194,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11398,
+ "image_id": 2095,
+ "bbox": [
+ 289,
+ 173,
+ 140,
+ 107
+ ],
+ "category_id": 8,
+ "area": 119402,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11399,
+ "image_id": 2095,
+ "bbox": [
+ 440,
+ 195,
+ 71,
+ 75
+ ],
+ "category_id": 8,
+ "area": 42880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11427,
+ "image_id": 2099,
+ "bbox": [
+ 183,
+ 102,
+ 134,
+ 309
+ ],
+ "category_id": 10,
+ "area": 55081,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11428,
+ "image_id": 2099,
+ "bbox": [
+ 4,
+ 407,
+ 109,
+ 104
+ ],
+ "category_id": 6,
+ "area": 15106,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11429,
+ "image_id": 2099,
+ "bbox": [
+ 297,
+ 412,
+ 80,
+ 99
+ ],
+ "category_id": 6,
+ "area": 10586,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11433,
+ "image_id": 2101,
+ "bbox": [
+ 94,
+ 160,
+ 88,
+ 93
+ ],
+ "category_id": 6,
+ "area": 29082,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11434,
+ "image_id": 2101,
+ "bbox": [
+ 92,
+ 216,
+ 93,
+ 112
+ ],
+ "category_id": 6,
+ "area": 36814,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11435,
+ "image_id": 2101,
+ "bbox": [
+ 100,
+ 324,
+ 44,
+ 64
+ ],
+ "category_id": 6,
+ "area": 9990,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11436,
+ "image_id": 2101,
+ "bbox": [
+ 4,
+ 280,
+ 46,
+ 140
+ ],
+ "category_id": 6,
+ "area": 22852,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11437,
+ "image_id": 2101,
+ "bbox": [
+ 176,
+ 445,
+ 99,
+ 66
+ ],
+ "category_id": 6,
+ "area": 23064,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11438,
+ "image_id": 2101,
+ "bbox": [
+ 278,
+ 409,
+ 112,
+ 100
+ ],
+ "category_id": 6,
+ "area": 40044,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11439,
+ "image_id": 2101,
+ "bbox": [
+ 370,
+ 378,
+ 141,
+ 133
+ ],
+ "category_id": 6,
+ "area": 66364,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11440,
+ "image_id": 2101,
+ "bbox": [
+ 289,
+ 361,
+ 82,
+ 77
+ ],
+ "category_id": 6,
+ "area": 22563,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11441,
+ "image_id": 2101,
+ "bbox": [
+ 215,
+ 296,
+ 40,
+ 54
+ ],
+ "category_id": 6,
+ "area": 7854,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11442,
+ "image_id": 2101,
+ "bbox": [
+ 409,
+ 327,
+ 6,
+ 47
+ ],
+ "category_id": 6,
+ "area": 1139,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11443,
+ "image_id": 2101,
+ "bbox": [
+ 216,
+ 192,
+ 109,
+ 88
+ ],
+ "category_id": 8,
+ "area": 34250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11444,
+ "image_id": 2101,
+ "bbox": [
+ 376,
+ 243,
+ 39,
+ 59
+ ],
+ "category_id": 6,
+ "area": 8316,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11446,
+ "image_id": 2102,
+ "bbox": [
+ 162,
+ 0,
+ 106,
+ 102
+ ],
+ "category_id": 6,
+ "area": 53440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11447,
+ "image_id": 2102,
+ "bbox": [
+ 200,
+ 196,
+ 184,
+ 263
+ ],
+ "category_id": 8,
+ "area": 236930,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11449,
+ "image_id": 2102,
+ "bbox": [
+ 282,
+ 0,
+ 151,
+ 163
+ ],
+ "category_id": 6,
+ "area": 121030,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11460,
+ "image_id": 2105,
+ "bbox": [
+ 99,
+ 10,
+ 144,
+ 166
+ ],
+ "category_id": 8,
+ "area": 189540,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11461,
+ "image_id": 2105,
+ "bbox": [
+ 154,
+ 56,
+ 142,
+ 121
+ ],
+ "category_id": 8,
+ "area": 137238,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11462,
+ "image_id": 2105,
+ "bbox": [
+ 0,
+ 19,
+ 108,
+ 141
+ ],
+ "category_id": 8,
+ "area": 120988,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11470,
+ "image_id": 2108,
+ "bbox": [
+ 3,
+ 42,
+ 128,
+ 99
+ ],
+ "category_id": 6,
+ "area": 166144,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11471,
+ "image_id": 2108,
+ "bbox": [
+ 4,
+ 111,
+ 76,
+ 146
+ ],
+ "category_id": 6,
+ "area": 144480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11472,
+ "image_id": 2108,
+ "bbox": [
+ 3,
+ 243,
+ 58,
+ 90
+ ],
+ "category_id": 6,
+ "area": 68585,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11473,
+ "image_id": 2108,
+ "bbox": [
+ 74,
+ 266,
+ 71,
+ 87
+ ],
+ "category_id": 6,
+ "area": 80434,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11474,
+ "image_id": 2108,
+ "bbox": [
+ 83,
+ 404,
+ 81,
+ 79
+ ],
+ "category_id": 6,
+ "area": 84882,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11475,
+ "image_id": 2108,
+ "bbox": [
+ 4,
+ 114,
+ 70,
+ 131
+ ],
+ "category_id": 6,
+ "area": 119970,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11476,
+ "image_id": 2108,
+ "bbox": [
+ 44,
+ 25,
+ 111,
+ 72
+ ],
+ "category_id": 6,
+ "area": 104704,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11477,
+ "image_id": 2108,
+ "bbox": [
+ 103,
+ 105,
+ 105,
+ 62
+ ],
+ "category_id": 6,
+ "area": 84753,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11478,
+ "image_id": 2108,
+ "bbox": [
+ 78,
+ 145,
+ 92,
+ 96
+ ],
+ "category_id": 6,
+ "area": 115258,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11479,
+ "image_id": 2108,
+ "bbox": [
+ 207,
+ 340,
+ 113,
+ 154
+ ],
+ "category_id": 6,
+ "area": 226590,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11480,
+ "image_id": 2108,
+ "bbox": [
+ 218,
+ 7,
+ 247,
+ 148
+ ],
+ "category_id": 6,
+ "area": 477750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11481,
+ "image_id": 2108,
+ "bbox": [
+ 238,
+ 198,
+ 79,
+ 51
+ ],
+ "category_id": 6,
+ "area": 52671,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11482,
+ "image_id": 2108,
+ "bbox": [
+ 470,
+ 31,
+ 39,
+ 67
+ ],
+ "category_id": 6,
+ "area": 34894,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11483,
+ "image_id": 2108,
+ "bbox": [
+ 413,
+ 116,
+ 98,
+ 110
+ ],
+ "category_id": 6,
+ "area": 140400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11487,
+ "image_id": 2109,
+ "bbox": [
+ 140,
+ 198,
+ 370,
+ 310
+ ],
+ "category_id": 10,
+ "area": 912496,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11493,
+ "image_id": 2110,
+ "bbox": [
+ 160,
+ 215,
+ 322,
+ 273
+ ],
+ "category_id": 6,
+ "area": 15532641,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11494,
+ "image_id": 2111,
+ "bbox": [
+ 0,
+ 86,
+ 353,
+ 280
+ ],
+ "category_id": 9,
+ "area": 169600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11495,
+ "image_id": 2111,
+ "bbox": [
+ 389,
+ 308,
+ 122,
+ 74
+ ],
+ "category_id": 9,
+ "area": 15640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11496,
+ "image_id": 2112,
+ "bbox": [
+ 97,
+ 9,
+ 198,
+ 244
+ ],
+ "category_id": 9,
+ "area": 111356,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11497,
+ "image_id": 2112,
+ "bbox": [
+ 8,
+ 100,
+ 271,
+ 372
+ ],
+ "category_id": 9,
+ "area": 231610,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11498,
+ "image_id": 2112,
+ "bbox": [
+ 182,
+ 267,
+ 219,
+ 167
+ ],
+ "category_id": 9,
+ "area": 83888,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11499,
+ "image_id": 2112,
+ "bbox": [
+ 295,
+ 60,
+ 216,
+ 339
+ ],
+ "category_id": 9,
+ "area": 167956,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11514,
+ "image_id": 2114,
+ "bbox": [
+ 3,
+ 102,
+ 508,
+ 404
+ ],
+ "category_id": 6,
+ "area": 261919,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11515,
+ "image_id": 2115,
+ "bbox": [
+ 129,
+ 161,
+ 244,
+ 201
+ ],
+ "category_id": 8,
+ "area": 391068,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11516,
+ "image_id": 2115,
+ "bbox": [
+ 108,
+ 48,
+ 257,
+ 219
+ ],
+ "category_id": 8,
+ "area": 447258,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11521,
+ "image_id": 2117,
+ "bbox": [
+ 142,
+ 206,
+ 213,
+ 261
+ ],
+ "category_id": 10,
+ "area": 125685,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11523,
+ "image_id": 2118,
+ "bbox": [
+ 59,
+ 0,
+ 348,
+ 496
+ ],
+ "category_id": 8,
+ "area": 139425,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11525,
+ "image_id": 2119,
+ "bbox": [
+ 1,
+ 0,
+ 340,
+ 511
+ ],
+ "category_id": 6,
+ "area": 1375725,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11531,
+ "image_id": 2121,
+ "bbox": [
+ 364,
+ 417,
+ 106,
+ 93
+ ],
+ "category_id": 9,
+ "area": 23296,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11532,
+ "image_id": 2121,
+ "bbox": [
+ 40,
+ 87,
+ 115,
+ 235
+ ],
+ "category_id": 9,
+ "area": 63958,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11533,
+ "image_id": 2121,
+ "bbox": [
+ 124,
+ 289,
+ 116,
+ 221
+ ],
+ "category_id": 9,
+ "area": 60382,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11534,
+ "image_id": 2121,
+ "bbox": [
+ 177,
+ 230,
+ 109,
+ 156
+ ],
+ "category_id": 9,
+ "area": 40232,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11535,
+ "image_id": 2121,
+ "bbox": [
+ 248,
+ 54,
+ 136,
+ 203
+ ],
+ "category_id": 9,
+ "area": 65148,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11536,
+ "image_id": 2121,
+ "bbox": [
+ 273,
+ 114,
+ 110,
+ 183
+ ],
+ "category_id": 9,
+ "area": 47520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11537,
+ "image_id": 2122,
+ "bbox": [
+ 96,
+ 129,
+ 104,
+ 240
+ ],
+ "category_id": 10,
+ "area": 33216,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11538,
+ "image_id": 2122,
+ "bbox": [
+ 221,
+ 228,
+ 220,
+ 283
+ ],
+ "category_id": 9,
+ "area": 82716,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11539,
+ "image_id": 2123,
+ "bbox": [
+ 114,
+ 22,
+ 394,
+ 485
+ ],
+ "category_id": 9,
+ "area": 673438,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11561,
+ "image_id": 2127,
+ "bbox": [
+ 95,
+ 63,
+ 311,
+ 261
+ ],
+ "category_id": 9,
+ "area": 115656,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11562,
+ "image_id": 2127,
+ "bbox": [
+ 340,
+ 164,
+ 44,
+ 81
+ ],
+ "category_id": 9,
+ "area": 5168,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11567,
+ "image_id": 2129,
+ "bbox": [
+ 245,
+ 205,
+ 54,
+ 103
+ ],
+ "category_id": 8,
+ "area": 38164,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11568,
+ "image_id": 2129,
+ "bbox": [
+ 99,
+ 138,
+ 87,
+ 53
+ ],
+ "category_id": 8,
+ "area": 31622,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11569,
+ "image_id": 2129,
+ "bbox": [
+ 149,
+ 83,
+ 49,
+ 77
+ ],
+ "category_id": 8,
+ "area": 25900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11570,
+ "image_id": 2129,
+ "bbox": [
+ 203,
+ 70,
+ 23,
+ 34
+ ],
+ "category_id": 8,
+ "area": 5607,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11571,
+ "image_id": 2129,
+ "bbox": [
+ 316,
+ 94,
+ 41,
+ 77
+ ],
+ "category_id": 8,
+ "area": 21840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11572,
+ "image_id": 2130,
+ "bbox": [
+ 256,
+ 219,
+ 254,
+ 209
+ ],
+ "category_id": 10,
+ "area": 1173420,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11574,
+ "image_id": 2131,
+ "bbox": [
+ 41,
+ 3,
+ 367,
+ 354
+ ],
+ "category_id": 8,
+ "area": 841347,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11575,
+ "image_id": 2131,
+ "bbox": [
+ 126,
+ 153,
+ 341,
+ 358
+ ],
+ "category_id": 8,
+ "area": 789186,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11576,
+ "image_id": 2131,
+ "bbox": [
+ 408,
+ 70,
+ 75,
+ 67
+ ],
+ "category_id": 8,
+ "area": 32994,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11597,
+ "image_id": 2135,
+ "bbox": [
+ 29,
+ 59,
+ 415,
+ 435
+ ],
+ "category_id": 9,
+ "area": 702180,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11598,
+ "image_id": 2136,
+ "bbox": [
+ 271,
+ 256,
+ 51,
+ 48
+ ],
+ "category_id": 8,
+ "area": 8772,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11599,
+ "image_id": 2136,
+ "bbox": [
+ 324,
+ 284,
+ 25,
+ 32
+ ],
+ "category_id": 8,
+ "area": 2898,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11600,
+ "image_id": 2137,
+ "bbox": [
+ 174,
+ 45,
+ 253,
+ 236
+ ],
+ "category_id": 9,
+ "area": 211122,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11601,
+ "image_id": 2137,
+ "bbox": [
+ 205,
+ 135,
+ 222,
+ 271
+ ],
+ "category_id": 9,
+ "area": 212392,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11606,
+ "image_id": 2139,
+ "bbox": [
+ 21,
+ 34,
+ 259,
+ 172
+ ],
+ "category_id": 6,
+ "area": 467343,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11607,
+ "image_id": 2139,
+ "bbox": [
+ 282,
+ 36,
+ 124,
+ 126
+ ],
+ "category_id": 6,
+ "area": 163500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11608,
+ "image_id": 2139,
+ "bbox": [
+ 377,
+ 48,
+ 134,
+ 224
+ ],
+ "category_id": 6,
+ "area": 314157,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11609,
+ "image_id": 2139,
+ "bbox": [
+ 226,
+ 130,
+ 135,
+ 187
+ ],
+ "category_id": 6,
+ "area": 263544,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11610,
+ "image_id": 2139,
+ "bbox": [
+ 0,
+ 434,
+ 113,
+ 76
+ ],
+ "category_id": 6,
+ "area": 90972,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11611,
+ "image_id": 2139,
+ "bbox": [
+ 375,
+ 257,
+ 114,
+ 97
+ ],
+ "category_id": 6,
+ "area": 116290,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11622,
+ "image_id": 2141,
+ "bbox": [
+ 0,
+ 92,
+ 306,
+ 217
+ ],
+ "category_id": 9,
+ "area": 191923,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11623,
+ "image_id": 2141,
+ "bbox": [
+ 220,
+ 80,
+ 147,
+ 217
+ ],
+ "category_id": 10,
+ "area": 92168,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11627,
+ "image_id": 2143,
+ "bbox": [
+ 88,
+ 96,
+ 218,
+ 187
+ ],
+ "category_id": 9,
+ "area": 66836,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11644,
+ "image_id": 2144,
+ "bbox": [
+ 156,
+ 83,
+ 285,
+ 428
+ ],
+ "category_id": 9,
+ "area": 411600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11645,
+ "image_id": 2144,
+ "bbox": [
+ 222,
+ 104,
+ 113,
+ 231
+ ],
+ "category_id": 10,
+ "area": 88722,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11647,
+ "image_id": 2146,
+ "bbox": [
+ 3,
+ 5,
+ 354,
+ 498
+ ],
+ "category_id": 9,
+ "area": 620385,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11648,
+ "image_id": 2146,
+ "bbox": [
+ 144,
+ 201,
+ 332,
+ 300
+ ],
+ "category_id": 9,
+ "area": 351513,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11649,
+ "image_id": 2147,
+ "bbox": [
+ 118,
+ 241,
+ 198,
+ 164
+ ],
+ "category_id": 8,
+ "area": 114576,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11650,
+ "image_id": 2147,
+ "bbox": [
+ 247,
+ 297,
+ 140,
+ 139
+ ],
+ "category_id": 8,
+ "area": 68796,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11678,
+ "image_id": 2149,
+ "bbox": [
+ 168,
+ 1,
+ 147,
+ 162
+ ],
+ "category_id": 9,
+ "area": 84501,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11679,
+ "image_id": 2149,
+ "bbox": [
+ 151,
+ 93,
+ 147,
+ 122
+ ],
+ "category_id": 9,
+ "area": 63468,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11707,
+ "image_id": 2154,
+ "bbox": [
+ 79,
+ 210,
+ 180,
+ 192
+ ],
+ "category_id": 8,
+ "area": 110454,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11708,
+ "image_id": 2154,
+ "bbox": [
+ 231,
+ 167,
+ 186,
+ 232
+ ],
+ "category_id": 8,
+ "area": 137344,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11709,
+ "image_id": 2155,
+ "bbox": [
+ 158,
+ 254,
+ 86,
+ 92
+ ],
+ "category_id": 8,
+ "area": 28210,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11710,
+ "image_id": 2155,
+ "bbox": [
+ 217,
+ 266,
+ 103,
+ 79
+ ],
+ "category_id": 8,
+ "area": 29008,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11712,
+ "image_id": 2156,
+ "bbox": [
+ 99,
+ 221,
+ 220,
+ 147
+ ],
+ "category_id": 9,
+ "area": 114057,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11713,
+ "image_id": 2156,
+ "bbox": [
+ 383,
+ 148,
+ 59,
+ 182
+ ],
+ "category_id": 9,
+ "area": 38036,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11773,
+ "image_id": 2159,
+ "bbox": [
+ 69,
+ 60,
+ 322,
+ 436
+ ],
+ "category_id": 9,
+ "area": 1613350,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11778,
+ "image_id": 2161,
+ "bbox": [
+ 40,
+ 1,
+ 269,
+ 79
+ ],
+ "category_id": 6,
+ "area": 174811,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11779,
+ "image_id": 2161,
+ "bbox": [
+ 72,
+ 142,
+ 87,
+ 66
+ ],
+ "category_id": 6,
+ "area": 47288,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11780,
+ "image_id": 2161,
+ "bbox": [
+ 286,
+ 5,
+ 154,
+ 195
+ ],
+ "category_id": 6,
+ "area": 245700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11781,
+ "image_id": 2161,
+ "bbox": [
+ 461,
+ 138,
+ 50,
+ 84
+ ],
+ "category_id": 6,
+ "area": 34866,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11782,
+ "image_id": 2161,
+ "bbox": [
+ 10,
+ 343,
+ 157,
+ 168
+ ],
+ "category_id": 6,
+ "area": 215760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11783,
+ "image_id": 2161,
+ "bbox": [
+ 461,
+ 0,
+ 50,
+ 104
+ ],
+ "category_id": 6,
+ "area": 43210,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11788,
+ "image_id": 2162,
+ "bbox": [
+ 184,
+ 159,
+ 89,
+ 227
+ ],
+ "category_id": 10,
+ "area": 68640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11789,
+ "image_id": 2162,
+ "bbox": [
+ 219,
+ 93,
+ 224,
+ 418
+ ],
+ "category_id": 9,
+ "area": 315700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11819,
+ "image_id": 2166,
+ "bbox": [
+ 110,
+ 150,
+ 383,
+ 357
+ ],
+ "category_id": 9,
+ "area": 482377,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11820,
+ "image_id": 2166,
+ "bbox": [
+ 150,
+ 39,
+ 199,
+ 232
+ ],
+ "category_id": 9,
+ "area": 163173,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11821,
+ "image_id": 2166,
+ "bbox": [
+ 238,
+ 0,
+ 256,
+ 423
+ ],
+ "category_id": 9,
+ "area": 382036,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11822,
+ "image_id": 2167,
+ "bbox": [
+ 199,
+ 130,
+ 136,
+ 242
+ ],
+ "category_id": 6,
+ "area": 72459,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11823,
+ "image_id": 2167,
+ "bbox": [
+ 204,
+ 343,
+ 187,
+ 168
+ ],
+ "category_id": 9,
+ "area": 69027,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11824,
+ "image_id": 2168,
+ "bbox": [
+ 497,
+ 219,
+ 13,
+ 35
+ ],
+ "category_id": 6,
+ "area": 3375,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11825,
+ "image_id": 2168,
+ "bbox": [
+ 326,
+ 290,
+ 28,
+ 69
+ ],
+ "category_id": 6,
+ "area": 13870,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11826,
+ "image_id": 2168,
+ "bbox": [
+ 308,
+ 363,
+ 20,
+ 34
+ ],
+ "category_id": 6,
+ "area": 4891,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11830,
+ "image_id": 2169,
+ "bbox": [
+ 44,
+ 161,
+ 268,
+ 321
+ ],
+ "category_id": 8,
+ "area": 218614,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11832,
+ "image_id": 2170,
+ "bbox": [
+ 69,
+ 211,
+ 308,
+ 222
+ ],
+ "category_id": 8,
+ "area": 241323,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11833,
+ "image_id": 2170,
+ "bbox": [
+ 332,
+ 275,
+ 117,
+ 177
+ ],
+ "category_id": 8,
+ "area": 73206,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11852,
+ "image_id": 2172,
+ "bbox": [
+ 148,
+ 244,
+ 261,
+ 214
+ ],
+ "category_id": 8,
+ "area": 196200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11853,
+ "image_id": 2172,
+ "bbox": [
+ 236,
+ 88,
+ 230,
+ 210
+ ],
+ "category_id": 8,
+ "area": 169920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11854,
+ "image_id": 2172,
+ "bbox": [
+ 1,
+ 14,
+ 244,
+ 493
+ ],
+ "category_id": 6,
+ "area": 422892,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11874,
+ "image_id": 2176,
+ "bbox": [
+ 1,
+ 45,
+ 121,
+ 194
+ ],
+ "category_id": 8,
+ "area": 34138,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11875,
+ "image_id": 2176,
+ "bbox": [
+ 95,
+ 40,
+ 154,
+ 278
+ ],
+ "category_id": 8,
+ "area": 62451,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11876,
+ "image_id": 2176,
+ "bbox": [
+ 210,
+ 25,
+ 197,
+ 258
+ ],
+ "category_id": 8,
+ "area": 74025,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11877,
+ "image_id": 2176,
+ "bbox": [
+ 428,
+ 43,
+ 83,
+ 208
+ ],
+ "category_id": 8,
+ "area": 25298,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11884,
+ "image_id": 2179,
+ "bbox": [
+ 309,
+ 44,
+ 48,
+ 44
+ ],
+ "category_id": 8,
+ "area": 7560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11885,
+ "image_id": 2179,
+ "bbox": [
+ 75,
+ 212,
+ 188,
+ 173
+ ],
+ "category_id": 8,
+ "area": 114210,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11886,
+ "image_id": 2179,
+ "bbox": [
+ 208,
+ 206,
+ 229,
+ 177
+ ],
+ "category_id": 8,
+ "area": 142926,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11909,
+ "image_id": 2182,
+ "bbox": [
+ 186,
+ 58,
+ 276,
+ 349
+ ],
+ "category_id": 9,
+ "area": 83754,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11910,
+ "image_id": 2182,
+ "bbox": [
+ 120,
+ 177,
+ 107,
+ 177
+ ],
+ "category_id": 10,
+ "area": 16588,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11920,
+ "image_id": 2185,
+ "bbox": [
+ 64,
+ 312,
+ 142,
+ 87
+ ],
+ "category_id": 9,
+ "area": 18800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11921,
+ "image_id": 2185,
+ "bbox": [
+ 168,
+ 421,
+ 60,
+ 85
+ ],
+ "category_id": 6,
+ "area": 7722,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11922,
+ "image_id": 2185,
+ "bbox": [
+ 126,
+ 442,
+ 30,
+ 42
+ ],
+ "category_id": 6,
+ "area": 1950,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11923,
+ "image_id": 2185,
+ "bbox": [
+ 142,
+ 480,
+ 30,
+ 30
+ ],
+ "category_id": 6,
+ "area": 1428,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11925,
+ "image_id": 2186,
+ "bbox": [
+ 0,
+ 17,
+ 222,
+ 200
+ ],
+ "category_id": 9,
+ "area": 58305,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11926,
+ "image_id": 2186,
+ "bbox": [
+ 317,
+ 180,
+ 194,
+ 180
+ ],
+ "category_id": 9,
+ "area": 45850,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11927,
+ "image_id": 2186,
+ "bbox": [
+ 194,
+ 280,
+ 203,
+ 231
+ ],
+ "category_id": 10,
+ "area": 61425,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11947,
+ "image_id": 2188,
+ "bbox": [
+ 263,
+ 292,
+ 164,
+ 165
+ ],
+ "category_id": 9,
+ "area": 34240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11953,
+ "image_id": 2189,
+ "bbox": [
+ 224,
+ 247,
+ 193,
+ 195
+ ],
+ "category_id": 8,
+ "area": 132616,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11954,
+ "image_id": 2189,
+ "bbox": [
+ 152,
+ 249,
+ 72,
+ 159
+ ],
+ "category_id": 8,
+ "area": 40586,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11959,
+ "image_id": 2191,
+ "bbox": [
+ 37,
+ 208,
+ 195,
+ 217
+ ],
+ "category_id": 8,
+ "area": 54060,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11960,
+ "image_id": 2191,
+ "bbox": [
+ 224,
+ 265,
+ 126,
+ 123
+ ],
+ "category_id": 8,
+ "area": 19965,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11961,
+ "image_id": 2191,
+ "bbox": [
+ 161,
+ 13,
+ 333,
+ 243
+ ],
+ "category_id": 8,
+ "area": 103292,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11970,
+ "image_id": 2192,
+ "bbox": [
+ 136,
+ 266,
+ 239,
+ 213
+ ],
+ "category_id": 6,
+ "area": 157590,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11995,
+ "image_id": 2195,
+ "bbox": [
+ 51,
+ 119,
+ 345,
+ 276
+ ],
+ "category_id": 10,
+ "area": 209628,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11996,
+ "image_id": 2196,
+ "bbox": [
+ 59,
+ 75,
+ 439,
+ 347
+ ],
+ "category_id": 10,
+ "area": 178808,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11997,
+ "image_id": 2197,
+ "bbox": [
+ 92,
+ 162,
+ 61,
+ 67
+ ],
+ "category_id": 8,
+ "area": 14630,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11998,
+ "image_id": 2197,
+ "bbox": [
+ 185,
+ 139,
+ 196,
+ 372
+ ],
+ "category_id": 8,
+ "area": 257284,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 11999,
+ "image_id": 2198,
+ "bbox": [
+ 201,
+ 180,
+ 128,
+ 168
+ ],
+ "category_id": 9,
+ "area": 24252,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12000,
+ "image_id": 2198,
+ "bbox": [
+ 349,
+ 72,
+ 108,
+ 183
+ ],
+ "category_id": 9,
+ "area": 22185,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12002,
+ "image_id": 2199,
+ "bbox": [
+ 69,
+ 183,
+ 320,
+ 215
+ ],
+ "category_id": 8,
+ "area": 73980,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12014,
+ "image_id": 2202,
+ "bbox": [
+ 122,
+ 224,
+ 222,
+ 225
+ ],
+ "category_id": 8,
+ "area": 144333,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12015,
+ "image_id": 2203,
+ "bbox": [
+ 38,
+ 33,
+ 326,
+ 392
+ ],
+ "category_id": 9,
+ "area": 450432,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12027,
+ "image_id": 2204,
+ "bbox": [
+ 159,
+ 2,
+ 166,
+ 240
+ ],
+ "category_id": 6,
+ "area": 267050,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12028,
+ "image_id": 2204,
+ "bbox": [
+ 0,
+ 122,
+ 112,
+ 282
+ ],
+ "category_id": 6,
+ "area": 212336,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12029,
+ "image_id": 2204,
+ "bbox": [
+ 57,
+ 385,
+ 57,
+ 125
+ ],
+ "category_id": 6,
+ "area": 48384,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12030,
+ "image_id": 2205,
+ "bbox": [
+ 231,
+ 67,
+ 117,
+ 344
+ ],
+ "category_id": 10,
+ "area": 66830,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12034,
+ "image_id": 2207,
+ "bbox": [
+ 14,
+ 78,
+ 56,
+ 50
+ ],
+ "category_id": 6,
+ "area": 19530,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12035,
+ "image_id": 2207,
+ "bbox": [
+ 27,
+ 27,
+ 70,
+ 72
+ ],
+ "category_id": 6,
+ "area": 35264,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12036,
+ "image_id": 2207,
+ "bbox": [
+ 108,
+ 39,
+ 28,
+ 31
+ ],
+ "category_id": 6,
+ "area": 6298,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12061,
+ "image_id": 2210,
+ "bbox": [
+ 211,
+ 130,
+ 120,
+ 112
+ ],
+ "category_id": 8,
+ "area": 106887,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12062,
+ "image_id": 2210,
+ "bbox": [
+ 166,
+ 51,
+ 113,
+ 200
+ ],
+ "category_id": 8,
+ "area": 180621,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12063,
+ "image_id": 2210,
+ "bbox": [
+ 0,
+ 41,
+ 178,
+ 162
+ ],
+ "category_id": 8,
+ "area": 228798,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12077,
+ "image_id": 2213,
+ "bbox": [
+ 80,
+ 157,
+ 404,
+ 255
+ ],
+ "category_id": 8,
+ "area": 136427,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12078,
+ "image_id": 2214,
+ "bbox": [
+ 84,
+ 98,
+ 112,
+ 127
+ ],
+ "category_id": 8,
+ "area": 68400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12079,
+ "image_id": 2214,
+ "bbox": [
+ 128,
+ 138,
+ 146,
+ 139
+ ],
+ "category_id": 8,
+ "area": 97344,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12080,
+ "image_id": 2214,
+ "bbox": [
+ 204,
+ 197,
+ 111,
+ 270
+ ],
+ "category_id": 8,
+ "area": 143871,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12081,
+ "image_id": 2214,
+ "bbox": [
+ 315,
+ 0,
+ 196,
+ 373
+ ],
+ "category_id": 8,
+ "area": 350424,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12082,
+ "image_id": 2214,
+ "bbox": [
+ 310,
+ 34,
+ 62,
+ 224
+ ],
+ "category_id": 8,
+ "area": 67000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12086,
+ "image_id": 2217,
+ "bbox": [
+ 0,
+ 134,
+ 112,
+ 257
+ ],
+ "category_id": 8,
+ "area": 67482,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12087,
+ "image_id": 2217,
+ "bbox": [
+ 317,
+ 178,
+ 194,
+ 134
+ ],
+ "category_id": 8,
+ "area": 61218,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12094,
+ "image_id": 2218,
+ "bbox": [
+ 166,
+ 186,
+ 264,
+ 167
+ ],
+ "category_id": 9,
+ "area": 50616,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12127,
+ "image_id": 2222,
+ "bbox": [
+ 101,
+ 164,
+ 109,
+ 273
+ ],
+ "category_id": 8,
+ "area": 130242,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12128,
+ "image_id": 2222,
+ "bbox": [
+ 258,
+ 0,
+ 245,
+ 261
+ ],
+ "category_id": 6,
+ "area": 278334,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12129,
+ "image_id": 2223,
+ "bbox": [
+ 299,
+ 268,
+ 212,
+ 242
+ ],
+ "category_id": 6,
+ "area": 63800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12130,
+ "image_id": 2223,
+ "bbox": [
+ 0,
+ 0,
+ 323,
+ 407
+ ],
+ "category_id": 6,
+ "area": 163540,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12145,
+ "image_id": 2225,
+ "bbox": [
+ 240,
+ 130,
+ 41,
+ 62
+ ],
+ "category_id": 6,
+ "area": 20567,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12146,
+ "image_id": 2226,
+ "bbox": [
+ 88,
+ 275,
+ 314,
+ 113
+ ],
+ "category_id": 10,
+ "area": 48276,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12147,
+ "image_id": 2226,
+ "bbox": [
+ 214,
+ 354,
+ 289,
+ 113
+ ],
+ "category_id": 10,
+ "area": 44388,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12148,
+ "image_id": 2226,
+ "bbox": [
+ 415,
+ 361,
+ 93,
+ 46
+ ],
+ "category_id": 10,
+ "area": 5963,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12149,
+ "image_id": 2226,
+ "bbox": [
+ 374,
+ 238,
+ 135,
+ 88
+ ],
+ "category_id": 10,
+ "area": 16256,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12150,
+ "image_id": 2226,
+ "bbox": [
+ 445,
+ 196,
+ 65,
+ 64
+ ],
+ "category_id": 10,
+ "area": 5704,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12151,
+ "image_id": 2226,
+ "bbox": [
+ 364,
+ 35,
+ 103,
+ 95
+ ],
+ "category_id": 10,
+ "area": 13426,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12152,
+ "image_id": 2226,
+ "bbox": [
+ 390,
+ 134,
+ 106,
+ 101
+ ],
+ "category_id": 10,
+ "area": 14645,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12153,
+ "image_id": 2226,
+ "bbox": [
+ 85,
+ 20,
+ 346,
+ 235
+ ],
+ "category_id": 10,
+ "area": 110536,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12154,
+ "image_id": 2226,
+ "bbox": [
+ 97,
+ 199,
+ 111,
+ 54
+ ],
+ "category_id": 10,
+ "area": 8268,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12155,
+ "image_id": 2226,
+ "bbox": [
+ 0,
+ 87,
+ 99,
+ 90
+ ],
+ "category_id": 10,
+ "area": 12126,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12156,
+ "image_id": 2226,
+ "bbox": [
+ 32,
+ 28,
+ 64,
+ 78
+ ],
+ "category_id": 10,
+ "area": 6832,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12157,
+ "image_id": 2226,
+ "bbox": [
+ 68,
+ 13,
+ 126,
+ 70
+ ],
+ "category_id": 10,
+ "area": 12120,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12158,
+ "image_id": 2226,
+ "bbox": [
+ 29,
+ 0,
+ 166,
+ 20
+ ],
+ "category_id": 10,
+ "area": 4740,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12177,
+ "image_id": 2230,
+ "bbox": [
+ 0,
+ 143,
+ 511,
+ 366
+ ],
+ "category_id": 6,
+ "area": 882233,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12178,
+ "image_id": 2230,
+ "bbox": [
+ 0,
+ 40,
+ 376,
+ 394
+ ],
+ "category_id": 8,
+ "area": 698740,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12179,
+ "image_id": 2231,
+ "bbox": [
+ 186,
+ 125,
+ 213,
+ 148
+ ],
+ "category_id": 10,
+ "area": 50091,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12182,
+ "image_id": 2232,
+ "bbox": [
+ 149,
+ 296,
+ 164,
+ 164
+ ],
+ "category_id": 9,
+ "area": 26296,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12196,
+ "image_id": 2235,
+ "bbox": [
+ 56,
+ 88,
+ 427,
+ 413
+ ],
+ "category_id": 9,
+ "area": 620508,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12198,
+ "image_id": 2236,
+ "bbox": [
+ 81,
+ 178,
+ 317,
+ 280
+ ],
+ "category_id": 8,
+ "area": 119082,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12200,
+ "image_id": 2237,
+ "bbox": [
+ 139,
+ 32,
+ 17,
+ 31
+ ],
+ "category_id": 6,
+ "area": 3640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12201,
+ "image_id": 2237,
+ "bbox": [
+ 246,
+ 379,
+ 14,
+ 30
+ ],
+ "category_id": 6,
+ "area": 2970,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12202,
+ "image_id": 2237,
+ "bbox": [
+ 294,
+ 396,
+ 35,
+ 44
+ ],
+ "category_id": 6,
+ "area": 10507,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12203,
+ "image_id": 2237,
+ "bbox": [
+ 271,
+ 437,
+ 45,
+ 74
+ ],
+ "category_id": 6,
+ "area": 22914,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12204,
+ "image_id": 2237,
+ "bbox": [
+ 413,
+ 384,
+ 37,
+ 104
+ ],
+ "category_id": 6,
+ "area": 26180,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12209,
+ "image_id": 2238,
+ "bbox": [
+ 171,
+ 216,
+ 282,
+ 171
+ ],
+ "category_id": 8,
+ "area": 120159,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12210,
+ "image_id": 2238,
+ "bbox": [
+ 124,
+ 228,
+ 110,
+ 155
+ ],
+ "category_id": 8,
+ "area": 42570,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12212,
+ "image_id": 2240,
+ "bbox": [
+ 125,
+ 185,
+ 120,
+ 72
+ ],
+ "category_id": 8,
+ "area": 30401,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12213,
+ "image_id": 2240,
+ "bbox": [
+ 254,
+ 239,
+ 129,
+ 149
+ ],
+ "category_id": 8,
+ "area": 67830,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12214,
+ "image_id": 2241,
+ "bbox": [
+ 141,
+ 221,
+ 185,
+ 115
+ ],
+ "category_id": 6,
+ "area": 52800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12219,
+ "image_id": 2242,
+ "bbox": [
+ 123,
+ 113,
+ 246,
+ 157
+ ],
+ "category_id": 10,
+ "area": 253402,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12220,
+ "image_id": 2243,
+ "bbox": [
+ 67,
+ 191,
+ 373,
+ 320
+ ],
+ "category_id": 9,
+ "area": 295504,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12221,
+ "image_id": 2243,
+ "bbox": [
+ 421,
+ 233,
+ 79,
+ 138
+ ],
+ "category_id": 6,
+ "area": 27348,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12222,
+ "image_id": 2243,
+ "bbox": [
+ 422,
+ 424,
+ 43,
+ 67
+ ],
+ "category_id": 6,
+ "area": 7161,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12226,
+ "image_id": 2245,
+ "bbox": [
+ 196,
+ 401,
+ 23,
+ 34
+ ],
+ "category_id": 6,
+ "area": 5840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12227,
+ "image_id": 2245,
+ "bbox": [
+ 248,
+ 393,
+ 25,
+ 28
+ ],
+ "category_id": 6,
+ "area": 5015,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12228,
+ "image_id": 2245,
+ "bbox": [
+ 291,
+ 355,
+ 35,
+ 50
+ ],
+ "category_id": 6,
+ "area": 12720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12229,
+ "image_id": 2245,
+ "bbox": [
+ 313,
+ 340,
+ 31,
+ 51
+ ],
+ "category_id": 6,
+ "area": 11232,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12254,
+ "image_id": 2247,
+ "bbox": [
+ 54,
+ 179,
+ 351,
+ 332
+ ],
+ "category_id": 8,
+ "area": 486864,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12255,
+ "image_id": 2247,
+ "bbox": [
+ 171,
+ 9,
+ 196,
+ 194
+ ],
+ "category_id": 6,
+ "area": 159239,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12256,
+ "image_id": 2248,
+ "bbox": [
+ 1,
+ 1,
+ 91,
+ 99
+ ],
+ "category_id": 6,
+ "area": 85215,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12257,
+ "image_id": 2248,
+ "bbox": [
+ 0,
+ 57,
+ 79,
+ 76
+ ],
+ "category_id": 6,
+ "area": 57159,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12258,
+ "image_id": 2248,
+ "bbox": [
+ 79,
+ 0,
+ 188,
+ 112
+ ],
+ "category_id": 6,
+ "area": 200232,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12259,
+ "image_id": 2248,
+ "bbox": [
+ 2,
+ 96,
+ 137,
+ 123
+ ],
+ "category_id": 6,
+ "area": 159750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12260,
+ "image_id": 2248,
+ "bbox": [
+ 0,
+ 206,
+ 73,
+ 99
+ ],
+ "category_id": 6,
+ "area": 69212,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12261,
+ "image_id": 2248,
+ "bbox": [
+ 90,
+ 177,
+ 66,
+ 74
+ ],
+ "category_id": 6,
+ "area": 46440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12262,
+ "image_id": 2248,
+ "bbox": [
+ 120,
+ 240,
+ 69,
+ 107
+ ],
+ "category_id": 6,
+ "area": 69834,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12263,
+ "image_id": 2248,
+ "bbox": [
+ 202,
+ 263,
+ 77,
+ 99
+ ],
+ "category_id": 6,
+ "area": 72898,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12264,
+ "image_id": 2248,
+ "bbox": [
+ 128,
+ 74,
+ 79,
+ 186
+ ],
+ "category_id": 6,
+ "area": 139360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12265,
+ "image_id": 2248,
+ "bbox": [
+ 168,
+ 0,
+ 124,
+ 69
+ ],
+ "category_id": 6,
+ "area": 81807,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12266,
+ "image_id": 2248,
+ "bbox": [
+ 366,
+ 2,
+ 145,
+ 122
+ ],
+ "category_id": 6,
+ "area": 167904,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12267,
+ "image_id": 2248,
+ "bbox": [
+ 200,
+ 123,
+ 108,
+ 107
+ ],
+ "category_id": 6,
+ "area": 109695,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12268,
+ "image_id": 2248,
+ "bbox": [
+ 215,
+ 434,
+ 88,
+ 77
+ ],
+ "category_id": 6,
+ "area": 64512,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12269,
+ "image_id": 2248,
+ "bbox": [
+ 335,
+ 372,
+ 143,
+ 139
+ ],
+ "category_id": 6,
+ "area": 188000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12270,
+ "image_id": 2248,
+ "bbox": [
+ 394,
+ 180,
+ 81,
+ 75
+ ],
+ "category_id": 6,
+ "area": 57988,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12301,
+ "image_id": 2252,
+ "bbox": [
+ 169,
+ 315,
+ 129,
+ 191
+ ],
+ "category_id": 10,
+ "area": 31920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12302,
+ "image_id": 2252,
+ "bbox": [
+ 216,
+ 67,
+ 236,
+ 291
+ ],
+ "category_id": 9,
+ "area": 89088,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12311,
+ "image_id": 2254,
+ "bbox": [
+ 71,
+ 337,
+ 439,
+ 169
+ ],
+ "category_id": 6,
+ "area": 364532,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12312,
+ "image_id": 2254,
+ "bbox": [
+ 313,
+ 24,
+ 198,
+ 312
+ ],
+ "category_id": 6,
+ "area": 302940,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12315,
+ "image_id": 2256,
+ "bbox": [
+ 316,
+ 250,
+ 76,
+ 127
+ ],
+ "category_id": 10,
+ "area": 9696,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12316,
+ "image_id": 2256,
+ "bbox": [
+ 124,
+ 113,
+ 260,
+ 203
+ ],
+ "category_id": 9,
+ "area": 52650,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12332,
+ "image_id": 2260,
+ "bbox": [
+ 149,
+ 36,
+ 232,
+ 411
+ ],
+ "category_id": 10,
+ "area": 757764,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12335,
+ "image_id": 2261,
+ "bbox": [
+ 78,
+ 235,
+ 172,
+ 160
+ ],
+ "category_id": 8,
+ "area": 97200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12336,
+ "image_id": 2261,
+ "bbox": [
+ 207,
+ 229,
+ 206,
+ 178
+ ],
+ "category_id": 8,
+ "area": 129000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12337,
+ "image_id": 2262,
+ "bbox": [
+ 0,
+ 0,
+ 507,
+ 340
+ ],
+ "category_id": 8,
+ "area": 1332601,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12338,
+ "image_id": 2262,
+ "bbox": [
+ 78,
+ 89,
+ 130,
+ 146
+ ],
+ "category_id": 8,
+ "area": 146888,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12339,
+ "image_id": 2262,
+ "bbox": [
+ 77,
+ 277,
+ 320,
+ 200
+ ],
+ "category_id": 8,
+ "area": 495600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12340,
+ "image_id": 2263,
+ "bbox": [
+ 157,
+ 49,
+ 222,
+ 430
+ ],
+ "category_id": 9,
+ "area": 336936,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12342,
+ "image_id": 2263,
+ "bbox": [
+ 3,
+ 129,
+ 121,
+ 119
+ ],
+ "category_id": 6,
+ "area": 50904,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12343,
+ "image_id": 2263,
+ "bbox": [
+ 139,
+ 308,
+ 82,
+ 117
+ ],
+ "category_id": 6,
+ "area": 34155,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12344,
+ "image_id": 2263,
+ "bbox": [
+ 222,
+ 0,
+ 268,
+ 509
+ ],
+ "category_id": 6,
+ "area": 481824,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12349,
+ "image_id": 2265,
+ "bbox": [
+ 126,
+ 54,
+ 250,
+ 437
+ ],
+ "category_id": 10,
+ "area": 190756,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12350,
+ "image_id": 2265,
+ "bbox": [
+ 83,
+ 147,
+ 37,
+ 54
+ ],
+ "category_id": 10,
+ "area": 3519,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12351,
+ "image_id": 2265,
+ "bbox": [
+ 395,
+ 318,
+ 28,
+ 53
+ ],
+ "category_id": 10,
+ "area": 2600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12352,
+ "image_id": 2265,
+ "bbox": [
+ 468,
+ 311,
+ 35,
+ 54
+ ],
+ "category_id": 10,
+ "area": 3366,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12353,
+ "image_id": 2265,
+ "bbox": [
+ 102,
+ 20,
+ 42,
+ 31
+ ],
+ "category_id": 10,
+ "area": 2340,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12354,
+ "image_id": 2266,
+ "bbox": [
+ 69,
+ 28,
+ 287,
+ 250
+ ],
+ "category_id": 9,
+ "area": 253088,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12355,
+ "image_id": 2266,
+ "bbox": [
+ 205,
+ 205,
+ 152,
+ 204
+ ],
+ "category_id": 9,
+ "area": 109440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12356,
+ "image_id": 2266,
+ "bbox": [
+ 3,
+ 355,
+ 136,
+ 152
+ ],
+ "category_id": 9,
+ "area": 72974,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12357,
+ "image_id": 2267,
+ "bbox": [
+ 223,
+ 196,
+ 75,
+ 135
+ ],
+ "category_id": 9,
+ "area": 15594,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12358,
+ "image_id": 2267,
+ "bbox": [
+ 95,
+ 17,
+ 390,
+ 307
+ ],
+ "category_id": 6,
+ "area": 182479,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12395,
+ "image_id": 2273,
+ "bbox": [
+ 0,
+ 0,
+ 369,
+ 512
+ ],
+ "category_id": 6,
+ "area": 389754,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12396,
+ "image_id": 2274,
+ "bbox": [
+ 49,
+ 80,
+ 261,
+ 431
+ ],
+ "category_id": 9,
+ "area": 322083,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12400,
+ "image_id": 2274,
+ "bbox": [
+ 160,
+ 92,
+ 24,
+ 63
+ ],
+ "category_id": 10,
+ "area": 4472,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12410,
+ "image_id": 2276,
+ "bbox": [
+ 68,
+ 102,
+ 419,
+ 406
+ ],
+ "category_id": 8,
+ "area": 1172908,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12411,
+ "image_id": 2276,
+ "bbox": [
+ 294,
+ 178,
+ 103,
+ 237
+ ],
+ "category_id": 8,
+ "area": 169336,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12420,
+ "image_id": 2280,
+ "bbox": [
+ 309,
+ 174,
+ 53,
+ 78
+ ],
+ "category_id": 10,
+ "area": 6942,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12421,
+ "image_id": 2280,
+ "bbox": [
+ 287,
+ 121,
+ 86,
+ 69
+ ],
+ "category_id": 9,
+ "area": 10005,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12425,
+ "image_id": 2281,
+ "bbox": [
+ 99,
+ 271,
+ 192,
+ 199
+ ],
+ "category_id": 8,
+ "area": 134680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12426,
+ "image_id": 2281,
+ "bbox": [
+ 177,
+ 329,
+ 226,
+ 171
+ ],
+ "category_id": 8,
+ "area": 135840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12433,
+ "image_id": 2281,
+ "bbox": [
+ 385,
+ 1,
+ 30,
+ 32
+ ],
+ "category_id": 9,
+ "area": 3465,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12434,
+ "image_id": 2282,
+ "bbox": [
+ 0,
+ 31,
+ 178,
+ 253
+ ],
+ "category_id": 6,
+ "area": 60865,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12435,
+ "image_id": 2282,
+ "bbox": [
+ 0,
+ 259,
+ 212,
+ 252
+ ],
+ "category_id": 9,
+ "area": 72072,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12437,
+ "image_id": 2283,
+ "bbox": [
+ 98,
+ 166,
+ 291,
+ 279
+ ],
+ "category_id": 9,
+ "area": 209752,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12440,
+ "image_id": 2284,
+ "bbox": [
+ 80,
+ 154,
+ 297,
+ 310
+ ],
+ "category_id": 9,
+ "area": 229416,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12482,
+ "image_id": 2288,
+ "bbox": [
+ 288,
+ 75,
+ 158,
+ 163
+ ],
+ "category_id": 10,
+ "area": 62456,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12483,
+ "image_id": 2288,
+ "bbox": [
+ 39,
+ 199,
+ 216,
+ 284
+ ],
+ "category_id": 9,
+ "area": 147498,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12484,
+ "image_id": 2289,
+ "bbox": [
+ 132,
+ 129,
+ 321,
+ 324
+ ],
+ "category_id": 9,
+ "area": 176098,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12485,
+ "image_id": 2289,
+ "bbox": [
+ 0,
+ 284,
+ 146,
+ 66
+ ],
+ "category_id": 6,
+ "area": 16422,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12486,
+ "image_id": 2289,
+ "bbox": [
+ 0,
+ 437,
+ 252,
+ 74
+ ],
+ "category_id": 6,
+ "area": 31980,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12490,
+ "image_id": 2291,
+ "bbox": [
+ 219,
+ 77,
+ 215,
+ 279
+ ],
+ "category_id": 8,
+ "area": 70209,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12510,
+ "image_id": 2294,
+ "bbox": [
+ 2,
+ 161,
+ 416,
+ 346
+ ],
+ "category_id": 9,
+ "area": 338954,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12511,
+ "image_id": 2295,
+ "bbox": [
+ 10,
+ 108,
+ 141,
+ 251
+ ],
+ "category_id": 9,
+ "area": 124609,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12512,
+ "image_id": 2295,
+ "bbox": [
+ 153,
+ 0,
+ 210,
+ 301
+ ],
+ "category_id": 9,
+ "area": 223448,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12513,
+ "image_id": 2295,
+ "bbox": [
+ 239,
+ 64,
+ 162,
+ 295
+ ],
+ "category_id": 9,
+ "area": 168480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12514,
+ "image_id": 2295,
+ "bbox": [
+ 164,
+ 229,
+ 74,
+ 219
+ ],
+ "category_id": 9,
+ "area": 57165,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12515,
+ "image_id": 2295,
+ "bbox": [
+ 215,
+ 226,
+ 182,
+ 278
+ ],
+ "category_id": 9,
+ "area": 178752,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12516,
+ "image_id": 2295,
+ "bbox": [
+ 368,
+ 321,
+ 108,
+ 185
+ ],
+ "category_id": 9,
+ "area": 70731,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12517,
+ "image_id": 2296,
+ "bbox": [
+ 3,
+ 38,
+ 508,
+ 469
+ ],
+ "category_id": 8,
+ "area": 1842135,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12518,
+ "image_id": 2296,
+ "bbox": [
+ 368,
+ 183,
+ 143,
+ 219
+ ],
+ "category_id": 8,
+ "area": 244167,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12527,
+ "image_id": 2298,
+ "bbox": [
+ 145,
+ 32,
+ 53,
+ 88
+ ],
+ "category_id": 6,
+ "area": 5984,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12528,
+ "image_id": 2298,
+ "bbox": [
+ 226,
+ 25,
+ 67,
+ 253
+ ],
+ "category_id": 6,
+ "area": 21450,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12529,
+ "image_id": 2298,
+ "bbox": [
+ 279,
+ 0,
+ 143,
+ 495
+ ],
+ "category_id": 6,
+ "area": 89535,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12530,
+ "image_id": 2299,
+ "bbox": [
+ 69,
+ 178,
+ 400,
+ 241
+ ],
+ "category_id": 8,
+ "area": 340680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12531,
+ "image_id": 2299,
+ "bbox": [
+ 26,
+ 142,
+ 92,
+ 135
+ ],
+ "category_id": 6,
+ "area": 43890,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12532,
+ "image_id": 2299,
+ "bbox": [
+ 127,
+ 317,
+ 74,
+ 93
+ ],
+ "category_id": 6,
+ "area": 24366,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12533,
+ "image_id": 2299,
+ "bbox": [
+ 0,
+ 288,
+ 36,
+ 112
+ ],
+ "category_id": 6,
+ "area": 14220,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12534,
+ "image_id": 2299,
+ "bbox": [
+ 363,
+ 69,
+ 72,
+ 95
+ ],
+ "category_id": 6,
+ "area": 24254,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12535,
+ "image_id": 2299,
+ "bbox": [
+ 414,
+ 35,
+ 97,
+ 162
+ ],
+ "category_id": 6,
+ "area": 55647,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12536,
+ "image_id": 2299,
+ "bbox": [
+ 0,
+ 10,
+ 71,
+ 132
+ ],
+ "category_id": 6,
+ "area": 33108,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12537,
+ "image_id": 2299,
+ "bbox": [
+ 399,
+ 223,
+ 63,
+ 130
+ ],
+ "category_id": 6,
+ "area": 28914,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12538,
+ "image_id": 2299,
+ "bbox": [
+ 221,
+ 125,
+ 30,
+ 32
+ ],
+ "category_id": 6,
+ "area": 3542,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12539,
+ "image_id": 2299,
+ "bbox": [
+ 145,
+ 17,
+ 20,
+ 25
+ ],
+ "category_id": 6,
+ "area": 1872,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12562,
+ "image_id": 2305,
+ "bbox": [
+ 96,
+ 144,
+ 126,
+ 191
+ ],
+ "category_id": 6,
+ "area": 118248,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12563,
+ "image_id": 2305,
+ "bbox": [
+ 304,
+ 103,
+ 205,
+ 180
+ ],
+ "category_id": 6,
+ "area": 181130,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12570,
+ "image_id": 2307,
+ "bbox": [
+ 188,
+ 45,
+ 161,
+ 326
+ ],
+ "category_id": 8,
+ "area": 416240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12571,
+ "image_id": 2307,
+ "bbox": [
+ 50,
+ 120,
+ 130,
+ 306
+ ],
+ "category_id": 6,
+ "area": 317030,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12572,
+ "image_id": 2307,
+ "bbox": [
+ 240,
+ 443,
+ 43,
+ 65
+ ],
+ "category_id": 6,
+ "area": 22632,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12573,
+ "image_id": 2307,
+ "bbox": [
+ 312,
+ 289,
+ 37,
+ 110
+ ],
+ "category_id": 6,
+ "area": 32526,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12574,
+ "image_id": 2307,
+ "bbox": [
+ 364,
+ 364,
+ 61,
+ 62
+ ],
+ "category_id": 6,
+ "area": 30130,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12575,
+ "image_id": 2307,
+ "bbox": [
+ 341,
+ 61,
+ 151,
+ 92
+ ],
+ "category_id": 6,
+ "area": 110760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12577,
+ "image_id": 2308,
+ "bbox": [
+ 284,
+ 156,
+ 138,
+ 95
+ ],
+ "category_id": 8,
+ "area": 15397,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12578,
+ "image_id": 2309,
+ "bbox": [
+ 213,
+ 72,
+ 131,
+ 190
+ ],
+ "category_id": 9,
+ "area": 88172,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12579,
+ "image_id": 2310,
+ "bbox": [
+ 198,
+ 113,
+ 166,
+ 213
+ ],
+ "category_id": 9,
+ "area": 124500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12580,
+ "image_id": 2310,
+ "bbox": [
+ 252,
+ 117,
+ 148,
+ 162
+ ],
+ "category_id": 9,
+ "area": 85188,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12581,
+ "image_id": 2311,
+ "bbox": [
+ 155,
+ 217,
+ 101,
+ 124
+ ],
+ "category_id": 8,
+ "area": 44450,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12582,
+ "image_id": 2311,
+ "bbox": [
+ 300,
+ 211,
+ 185,
+ 157
+ ],
+ "category_id": 8,
+ "area": 102080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12583,
+ "image_id": 2312,
+ "bbox": [
+ 199,
+ 124,
+ 278,
+ 171
+ ],
+ "category_id": 8,
+ "area": 669071,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12584,
+ "image_id": 2313,
+ "bbox": [
+ 56,
+ 200,
+ 232,
+ 195
+ ],
+ "category_id": 10,
+ "area": 71764,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12594,
+ "image_id": 2315,
+ "bbox": [
+ 112,
+ 82,
+ 399,
+ 425
+ ],
+ "category_id": 8,
+ "area": 1309620,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12595,
+ "image_id": 2316,
+ "bbox": [
+ 182,
+ 256,
+ 57,
+ 52
+ ],
+ "category_id": 8,
+ "area": 9648,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12596,
+ "image_id": 2316,
+ "bbox": [
+ 249,
+ 266,
+ 80,
+ 57
+ ],
+ "category_id": 8,
+ "area": 14527,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12597,
+ "image_id": 2317,
+ "bbox": [
+ 3,
+ 41,
+ 344,
+ 382
+ ],
+ "category_id": 6,
+ "area": 1226450,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12598,
+ "image_id": 2317,
+ "bbox": [
+ 263,
+ 293,
+ 100,
+ 211
+ ],
+ "category_id": 6,
+ "area": 197400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12599,
+ "image_id": 2317,
+ "bbox": [
+ 297,
+ 24,
+ 134,
+ 176
+ ],
+ "category_id": 6,
+ "area": 221256,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12600,
+ "image_id": 2317,
+ "bbox": [
+ 376,
+ 186,
+ 135,
+ 323
+ ],
+ "category_id": 6,
+ "area": 407416,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12601,
+ "image_id": 2317,
+ "bbox": [
+ 455,
+ 36,
+ 56,
+ 120
+ ],
+ "category_id": 6,
+ "area": 63600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12602,
+ "image_id": 2317,
+ "bbox": [
+ 0,
+ 1,
+ 86,
+ 178
+ ],
+ "category_id": 6,
+ "area": 143975,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12603,
+ "image_id": 2318,
+ "bbox": [
+ 236,
+ 231,
+ 70,
+ 88
+ ],
+ "category_id": 8,
+ "area": 21948,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12604,
+ "image_id": 2318,
+ "bbox": [
+ 397,
+ 156,
+ 93,
+ 172
+ ],
+ "category_id": 8,
+ "area": 56386,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12614,
+ "image_id": 2320,
+ "bbox": [
+ 0,
+ 122,
+ 281,
+ 160
+ ],
+ "category_id": 9,
+ "area": 76038,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12615,
+ "image_id": 2320,
+ "bbox": [
+ 0,
+ 223,
+ 185,
+ 287
+ ],
+ "category_id": 6,
+ "area": 89544,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12616,
+ "image_id": 2321,
+ "bbox": [
+ 164,
+ 208,
+ 151,
+ 200
+ ],
+ "category_id": 10,
+ "area": 50160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12618,
+ "image_id": 2322,
+ "bbox": [
+ 88,
+ 9,
+ 404,
+ 501
+ ],
+ "category_id": 9,
+ "area": 1602870,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12621,
+ "image_id": 2323,
+ "bbox": [
+ 8,
+ 70,
+ 267,
+ 246
+ ],
+ "category_id": 9,
+ "area": 45428,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12622,
+ "image_id": 2323,
+ "bbox": [
+ 273,
+ 69,
+ 228,
+ 433
+ ],
+ "category_id": 9,
+ "area": 68493,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12623,
+ "image_id": 2324,
+ "bbox": [
+ 60,
+ 334,
+ 239,
+ 177
+ ],
+ "category_id": 9,
+ "area": 76824,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12649,
+ "image_id": 2330,
+ "bbox": [
+ 115,
+ 121,
+ 135,
+ 376
+ ],
+ "category_id": 8,
+ "area": 178273,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12650,
+ "image_id": 2330,
+ "bbox": [
+ 242,
+ 140,
+ 149,
+ 212
+ ],
+ "category_id": 8,
+ "area": 111228,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12652,
+ "image_id": 2331,
+ "bbox": [
+ 19,
+ 235,
+ 492,
+ 136
+ ],
+ "category_id": 9,
+ "area": 467075,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12675,
+ "image_id": 2333,
+ "bbox": [
+ 30,
+ 406,
+ 104,
+ 99
+ ],
+ "category_id": 6,
+ "area": 28762,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12676,
+ "image_id": 2333,
+ "bbox": [
+ 109,
+ 404,
+ 124,
+ 107
+ ],
+ "category_id": 6,
+ "area": 37130,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12681,
+ "image_id": 2334,
+ "bbox": [
+ 421,
+ 353,
+ 84,
+ 113
+ ],
+ "category_id": 6,
+ "area": 69930,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12686,
+ "image_id": 2336,
+ "bbox": [
+ 311,
+ 210,
+ 24,
+ 60
+ ],
+ "category_id": 6,
+ "area": 1767,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12687,
+ "image_id": 2336,
+ "bbox": [
+ 96,
+ 129,
+ 52,
+ 67
+ ],
+ "category_id": 6,
+ "area": 4095,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12688,
+ "image_id": 2336,
+ "bbox": [
+ 290,
+ 84,
+ 52,
+ 69
+ ],
+ "category_id": 6,
+ "area": 4290,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12689,
+ "image_id": 2336,
+ "bbox": [
+ 376,
+ 105,
+ 36,
+ 58
+ ],
+ "category_id": 6,
+ "area": 2475,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12690,
+ "image_id": 2336,
+ "bbox": [
+ 437,
+ 353,
+ 66,
+ 69
+ ],
+ "category_id": 6,
+ "area": 5395,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12691,
+ "image_id": 2336,
+ "bbox": [
+ 396,
+ 408,
+ 48,
+ 68
+ ],
+ "category_id": 6,
+ "area": 3840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12692,
+ "image_id": 2337,
+ "bbox": [
+ 62,
+ 220,
+ 70,
+ 61
+ ],
+ "category_id": 8,
+ "area": 34580,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12693,
+ "image_id": 2337,
+ "bbox": [
+ 64,
+ 271,
+ 262,
+ 228
+ ],
+ "category_id": 8,
+ "area": 474266,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12694,
+ "image_id": 2338,
+ "bbox": [
+ 33,
+ 81,
+ 333,
+ 396
+ ],
+ "category_id": 8,
+ "area": 968018,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12695,
+ "image_id": 2339,
+ "bbox": [
+ 72,
+ 187,
+ 436,
+ 320
+ ],
+ "category_id": 9,
+ "area": 491400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12696,
+ "image_id": 2340,
+ "bbox": [
+ 4,
+ 35,
+ 77,
+ 188
+ ],
+ "category_id": 6,
+ "area": 205254,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12697,
+ "image_id": 2340,
+ "bbox": [
+ 0,
+ 158,
+ 23,
+ 58
+ ],
+ "category_id": 6,
+ "area": 19184,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12698,
+ "image_id": 2340,
+ "bbox": [
+ 15,
+ 160,
+ 56,
+ 56
+ ],
+ "category_id": 6,
+ "area": 45144,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12699,
+ "image_id": 2340,
+ "bbox": [
+ 2,
+ 198,
+ 111,
+ 117
+ ],
+ "category_id": 6,
+ "area": 185142,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12700,
+ "image_id": 2340,
+ "bbox": [
+ 0,
+ 303,
+ 76,
+ 98
+ ],
+ "category_id": 6,
+ "area": 106326,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12701,
+ "image_id": 2340,
+ "bbox": [
+ 75,
+ 277,
+ 49,
+ 91
+ ],
+ "category_id": 6,
+ "area": 64584,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12702,
+ "image_id": 2340,
+ "bbox": [
+ 98,
+ 335,
+ 49,
+ 109
+ ],
+ "category_id": 6,
+ "area": 76657,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12703,
+ "image_id": 2340,
+ "bbox": [
+ 155,
+ 361,
+ 55,
+ 123
+ ],
+ "category_id": 6,
+ "area": 97464,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12704,
+ "image_id": 2340,
+ "bbox": [
+ 69,
+ 93,
+ 132,
+ 135
+ ],
+ "category_id": 6,
+ "area": 253776,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12705,
+ "image_id": 2340,
+ "bbox": [
+ 102,
+ 175,
+ 57,
+ 180
+ ],
+ "category_id": 6,
+ "area": 146611,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12706,
+ "image_id": 2340,
+ "bbox": [
+ 185,
+ 154,
+ 74,
+ 80
+ ],
+ "category_id": 6,
+ "area": 85293,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12707,
+ "image_id": 2340,
+ "bbox": [
+ 232,
+ 74,
+ 37,
+ 52
+ ],
+ "category_id": 6,
+ "area": 27475,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12708,
+ "image_id": 2340,
+ "bbox": [
+ 268,
+ 42,
+ 194,
+ 180
+ ],
+ "category_id": 6,
+ "area": 496845,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12709,
+ "image_id": 2340,
+ "bbox": [
+ 401,
+ 52,
+ 61,
+ 71
+ ],
+ "category_id": 6,
+ "area": 62060,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12710,
+ "image_id": 2340,
+ "bbox": [
+ 470,
+ 53,
+ 39,
+ 117
+ ],
+ "category_id": 6,
+ "area": 66198,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12720,
+ "image_id": 2341,
+ "bbox": [
+ 215,
+ 51,
+ 295,
+ 458
+ ],
+ "category_id": 9,
+ "area": 143541,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12721,
+ "image_id": 2341,
+ "bbox": [
+ 31,
+ 0,
+ 194,
+ 319
+ ],
+ "category_id": 10,
+ "area": 65853,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12723,
+ "image_id": 2343,
+ "bbox": [
+ 80,
+ 234,
+ 228,
+ 233
+ ],
+ "category_id": 8,
+ "area": 186717,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12724,
+ "image_id": 2343,
+ "bbox": [
+ 202,
+ 179,
+ 206,
+ 274
+ ],
+ "category_id": 8,
+ "area": 199045,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12725,
+ "image_id": 2343,
+ "bbox": [
+ 295,
+ 55,
+ 101,
+ 124
+ ],
+ "category_id": 8,
+ "area": 44196,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12737,
+ "image_id": 2349,
+ "bbox": [
+ 132,
+ 242,
+ 202,
+ 187
+ ],
+ "category_id": 8,
+ "area": 109368,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12741,
+ "image_id": 2351,
+ "bbox": [
+ 0,
+ 206,
+ 231,
+ 305
+ ],
+ "category_id": 6,
+ "area": 44974,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12742,
+ "image_id": 2351,
+ "bbox": [
+ 159,
+ 106,
+ 352,
+ 355
+ ],
+ "category_id": 6,
+ "area": 79464,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12743,
+ "image_id": 2352,
+ "bbox": [
+ 125,
+ 265,
+ 147,
+ 107
+ ],
+ "category_id": 8,
+ "area": 124978,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12744,
+ "image_id": 2352,
+ "bbox": [
+ 148,
+ 199,
+ 162,
+ 149
+ ],
+ "category_id": 8,
+ "area": 193076,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12745,
+ "image_id": 2352,
+ "bbox": [
+ 265,
+ 151,
+ 121,
+ 127
+ ],
+ "category_id": 8,
+ "area": 122395,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12755,
+ "image_id": 2355,
+ "bbox": [
+ 0,
+ 0,
+ 333,
+ 512
+ ],
+ "category_id": 6,
+ "area": 1242952,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12759,
+ "image_id": 2356,
+ "bbox": [
+ 164,
+ 128,
+ 216,
+ 154
+ ],
+ "category_id": 10,
+ "area": 52808,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12763,
+ "image_id": 2357,
+ "bbox": [
+ 0,
+ 0,
+ 102,
+ 163
+ ],
+ "category_id": 6,
+ "area": 21300,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12764,
+ "image_id": 2357,
+ "bbox": [
+ 100,
+ 118,
+ 94,
+ 107
+ ],
+ "category_id": 6,
+ "area": 12880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12765,
+ "image_id": 2358,
+ "bbox": [
+ 71,
+ 130,
+ 393,
+ 285
+ ],
+ "category_id": 9,
+ "area": 395166,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12766,
+ "image_id": 2359,
+ "bbox": [
+ 0,
+ 100,
+ 220,
+ 398
+ ],
+ "category_id": 9,
+ "area": 227766,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12767,
+ "image_id": 2359,
+ "bbox": [
+ 230,
+ 56,
+ 164,
+ 301
+ ],
+ "category_id": 9,
+ "area": 128800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12768,
+ "image_id": 2359,
+ "bbox": [
+ 312,
+ 231,
+ 198,
+ 279
+ ],
+ "category_id": 9,
+ "area": 143856,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12771,
+ "image_id": 2360,
+ "bbox": [
+ 10,
+ 144,
+ 358,
+ 367
+ ],
+ "category_id": 9,
+ "area": 375705,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12772,
+ "image_id": 2360,
+ "bbox": [
+ 178,
+ 163,
+ 44,
+ 40
+ ],
+ "category_id": 10,
+ "area": 5148,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12776,
+ "image_id": 2361,
+ "bbox": [
+ 165,
+ 7,
+ 179,
+ 504
+ ],
+ "category_id": 10,
+ "area": 318080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12785,
+ "image_id": 2363,
+ "bbox": [
+ 280,
+ 81,
+ 186,
+ 163
+ ],
+ "category_id": 9,
+ "area": 223672,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12786,
+ "image_id": 2363,
+ "bbox": [
+ 70,
+ 200,
+ 219,
+ 295
+ ],
+ "category_id": 9,
+ "area": 476091,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12787,
+ "image_id": 2363,
+ "bbox": [
+ 233,
+ 176,
+ 278,
+ 279
+ ],
+ "category_id": 6,
+ "area": 569634,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12794,
+ "image_id": 2363,
+ "bbox": [
+ 63,
+ 350,
+ 26,
+ 92
+ ],
+ "category_id": 6,
+ "area": 18144,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12795,
+ "image_id": 2363,
+ "bbox": [
+ 28,
+ 280,
+ 41,
+ 54
+ ],
+ "category_id": 6,
+ "area": 16512,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12808,
+ "image_id": 2364,
+ "bbox": [
+ 1,
+ 0,
+ 364,
+ 510
+ ],
+ "category_id": 6,
+ "area": 1469816,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12826,
+ "image_id": 2366,
+ "bbox": [
+ 135,
+ 82,
+ 180,
+ 354
+ ],
+ "category_id": 10,
+ "area": 505648,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12827,
+ "image_id": 2367,
+ "bbox": [
+ 42,
+ 196,
+ 417,
+ 254
+ ],
+ "category_id": 9,
+ "area": 117284,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12828,
+ "image_id": 2368,
+ "bbox": [
+ 226,
+ 230,
+ 57,
+ 29
+ ],
+ "category_id": 8,
+ "area": 13482,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12829,
+ "image_id": 2368,
+ "bbox": [
+ 219,
+ 264,
+ 123,
+ 73
+ ],
+ "category_id": 8,
+ "area": 71610,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12830,
+ "image_id": 2368,
+ "bbox": [
+ 172,
+ 242,
+ 84,
+ 118
+ ],
+ "category_id": 8,
+ "area": 79065,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12831,
+ "image_id": 2368,
+ "bbox": [
+ 339,
+ 196,
+ 33,
+ 51
+ ],
+ "category_id": 6,
+ "area": 13734,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12854,
+ "image_id": 2372,
+ "bbox": [
+ 138,
+ 217,
+ 253,
+ 239
+ ],
+ "category_id": 8,
+ "area": 480760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12864,
+ "image_id": 2377,
+ "bbox": [
+ 173,
+ 276,
+ 159,
+ 132
+ ],
+ "category_id": 8,
+ "area": 66924,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12865,
+ "image_id": 2377,
+ "bbox": [
+ 313,
+ 191,
+ 43,
+ 97
+ ],
+ "category_id": 8,
+ "area": 13268,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12877,
+ "image_id": 2379,
+ "bbox": [
+ 41,
+ 23,
+ 416,
+ 488
+ ],
+ "category_id": 9,
+ "area": 3880560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12885,
+ "image_id": 2383,
+ "bbox": [
+ 2,
+ 144,
+ 507,
+ 364
+ ],
+ "category_id": 9,
+ "area": 649728,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12886,
+ "image_id": 2383,
+ "bbox": [
+ 114,
+ 16,
+ 218,
+ 458
+ ],
+ "category_id": 9,
+ "area": 351525,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12887,
+ "image_id": 2384,
+ "bbox": [
+ 265,
+ 208,
+ 246,
+ 288
+ ],
+ "category_id": 8,
+ "area": 82583,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12890,
+ "image_id": 2385,
+ "bbox": [
+ 121,
+ 125,
+ 336,
+ 244
+ ],
+ "category_id": 9,
+ "area": 204136,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12891,
+ "image_id": 2385,
+ "bbox": [
+ 365,
+ 0,
+ 146,
+ 130
+ ],
+ "category_id": 6,
+ "area": 47489,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12892,
+ "image_id": 2385,
+ "bbox": [
+ 207,
+ 37,
+ 92,
+ 112
+ ],
+ "category_id": 6,
+ "area": 25988,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12895,
+ "image_id": 2387,
+ "bbox": [
+ 74,
+ 180,
+ 191,
+ 281
+ ],
+ "category_id": 10,
+ "area": 49980,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12896,
+ "image_id": 2387,
+ "bbox": [
+ 128,
+ 16,
+ 270,
+ 331
+ ],
+ "category_id": 9,
+ "area": 82992,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12897,
+ "image_id": 2387,
+ "bbox": [
+ 256,
+ 147,
+ 185,
+ 135
+ ],
+ "category_id": 9,
+ "area": 23331,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12898,
+ "image_id": 2388,
+ "bbox": [
+ 0,
+ 168,
+ 200,
+ 194
+ ],
+ "category_id": 8,
+ "area": 136000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12899,
+ "image_id": 2388,
+ "bbox": [
+ 188,
+ 267,
+ 164,
+ 241
+ ],
+ "category_id": 8,
+ "area": 138918,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12900,
+ "image_id": 2388,
+ "bbox": [
+ 278,
+ 362,
+ 161,
+ 138
+ ],
+ "category_id": 8,
+ "area": 78182,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12906,
+ "image_id": 2390,
+ "bbox": [
+ 188,
+ 1,
+ 268,
+ 505
+ ],
+ "category_id": 9,
+ "area": 477081,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12907,
+ "image_id": 2390,
+ "bbox": [
+ 142,
+ 1,
+ 349,
+ 420
+ ],
+ "category_id": 9,
+ "area": 516816,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12908,
+ "image_id": 2390,
+ "bbox": [
+ 206,
+ 3,
+ 300,
+ 335
+ ],
+ "category_id": 9,
+ "area": 354000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12909,
+ "image_id": 2390,
+ "bbox": [
+ 445,
+ 352,
+ 66,
+ 154
+ ],
+ "category_id": 9,
+ "area": 35805,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12915,
+ "image_id": 2393,
+ "bbox": [
+ 3,
+ 24,
+ 314,
+ 487
+ ],
+ "category_id": 8,
+ "area": 534006,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12916,
+ "image_id": 2393,
+ "bbox": [
+ 296,
+ 199,
+ 162,
+ 201
+ ],
+ "category_id": 8,
+ "area": 114210,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12917,
+ "image_id": 2393,
+ "bbox": [
+ 24,
+ 0,
+ 487,
+ 505
+ ],
+ "category_id": 6,
+ "area": 857591,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12920,
+ "image_id": 2395,
+ "bbox": [
+ 37,
+ 243,
+ 301,
+ 238
+ ],
+ "category_id": 9,
+ "area": 82243,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12923,
+ "image_id": 2396,
+ "bbox": [
+ 139,
+ 214,
+ 126,
+ 148
+ ],
+ "category_id": 8,
+ "area": 65728,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12924,
+ "image_id": 2396,
+ "bbox": [
+ 227,
+ 202,
+ 177,
+ 184
+ ],
+ "category_id": 8,
+ "area": 114996,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12925,
+ "image_id": 2397,
+ "bbox": [
+ 122,
+ 83,
+ 265,
+ 425
+ ],
+ "category_id": 10,
+ "area": 201637,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12927,
+ "image_id": 2398,
+ "bbox": [
+ 92,
+ 119,
+ 331,
+ 217
+ ],
+ "category_id": 8,
+ "area": 81430,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12948,
+ "image_id": 2403,
+ "bbox": [
+ 103,
+ 17,
+ 393,
+ 322
+ ],
+ "category_id": 9,
+ "area": 111684,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12949,
+ "image_id": 2403,
+ "bbox": [
+ 174,
+ 219,
+ 152,
+ 271
+ ],
+ "category_id": 10,
+ "area": 36481,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12950,
+ "image_id": 2404,
+ "bbox": [
+ 66,
+ 1,
+ 347,
+ 467
+ ],
+ "category_id": 8,
+ "area": 595034,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12981,
+ "image_id": 2409,
+ "bbox": [
+ 31,
+ 132,
+ 472,
+ 379
+ ],
+ "category_id": 9,
+ "area": 628940,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12982,
+ "image_id": 2410,
+ "bbox": [
+ 176,
+ 184,
+ 156,
+ 98
+ ],
+ "category_id": 9,
+ "area": 68640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12992,
+ "image_id": 2414,
+ "bbox": [
+ 62,
+ 68,
+ 359,
+ 385
+ ],
+ "category_id": 9,
+ "area": 420042,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12993,
+ "image_id": 2415,
+ "bbox": [
+ 283,
+ 0,
+ 78,
+ 129
+ ],
+ "category_id": 8,
+ "area": 80556,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12994,
+ "image_id": 2415,
+ "bbox": [
+ 202,
+ 117,
+ 82,
+ 147
+ ],
+ "category_id": 8,
+ "area": 96099,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12995,
+ "image_id": 2415,
+ "bbox": [
+ 291,
+ 194,
+ 82,
+ 90
+ ],
+ "category_id": 8,
+ "area": 58828,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12996,
+ "image_id": 2415,
+ "bbox": [
+ 179,
+ 212,
+ 113,
+ 63
+ ],
+ "category_id": 8,
+ "area": 56392,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 12997,
+ "image_id": 2415,
+ "bbox": [
+ 152,
+ 164,
+ 56,
+ 38
+ ],
+ "category_id": 8,
+ "area": 17172,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13003,
+ "image_id": 2417,
+ "bbox": [
+ 106,
+ 216,
+ 230,
+ 144
+ ],
+ "category_id": 8,
+ "area": 117131,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13004,
+ "image_id": 2417,
+ "bbox": [
+ 208,
+ 29,
+ 279,
+ 216
+ ],
+ "category_id": 8,
+ "area": 211191,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13008,
+ "image_id": 2419,
+ "bbox": [
+ 58,
+ 128,
+ 448,
+ 378
+ ],
+ "category_id": 9,
+ "area": 595840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13009,
+ "image_id": 2419,
+ "bbox": [
+ 442,
+ 299,
+ 69,
+ 195
+ ],
+ "category_id": 9,
+ "area": 47575,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13010,
+ "image_id": 2419,
+ "bbox": [
+ 160,
+ 0,
+ 307,
+ 344
+ ],
+ "category_id": 9,
+ "area": 372965,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13011,
+ "image_id": 2419,
+ "bbox": [
+ 213,
+ 0,
+ 278,
+ 298
+ ],
+ "category_id": 9,
+ "area": 291900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13012,
+ "image_id": 2420,
+ "bbox": [
+ 0,
+ 6,
+ 244,
+ 504
+ ],
+ "category_id": 6,
+ "area": 432684,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13013,
+ "image_id": 2420,
+ "bbox": [
+ 105,
+ 182,
+ 238,
+ 257
+ ],
+ "category_id": 8,
+ "area": 214920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13014,
+ "image_id": 2420,
+ "bbox": [
+ 230,
+ 75,
+ 280,
+ 277
+ ],
+ "category_id": 8,
+ "area": 271600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13018,
+ "image_id": 2421,
+ "bbox": [
+ 0,
+ 196,
+ 260,
+ 314
+ ],
+ "category_id": 6,
+ "area": 286440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13019,
+ "image_id": 2421,
+ "bbox": [
+ 168,
+ 212,
+ 335,
+ 236
+ ],
+ "category_id": 8,
+ "area": 277709,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13020,
+ "image_id": 2422,
+ "bbox": [
+ 2,
+ 123,
+ 506,
+ 384
+ ],
+ "category_id": 9,
+ "area": 684180,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13021,
+ "image_id": 2422,
+ "bbox": [
+ 2,
+ 0,
+ 506,
+ 345
+ ],
+ "category_id": 9,
+ "area": 615762,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13022,
+ "image_id": 2422,
+ "bbox": [
+ 1,
+ 0,
+ 376,
+ 211
+ ],
+ "category_id": 9,
+ "area": 280716,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13023,
+ "image_id": 2422,
+ "bbox": [
+ 360,
+ 64,
+ 149,
+ 246
+ ],
+ "category_id": 9,
+ "area": 129058,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13029,
+ "image_id": 2424,
+ "bbox": [
+ 118,
+ 33,
+ 216,
+ 283
+ ],
+ "category_id": 10,
+ "area": 46420,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13053,
+ "image_id": 2428,
+ "bbox": [
+ 0,
+ 0,
+ 88,
+ 240
+ ],
+ "category_id": 10,
+ "area": 34958,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13054,
+ "image_id": 2428,
+ "bbox": [
+ 130,
+ 1,
+ 165,
+ 228
+ ],
+ "category_id": 10,
+ "area": 62208,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13055,
+ "image_id": 2428,
+ "bbox": [
+ 195,
+ 0,
+ 260,
+ 406
+ ],
+ "category_id": 10,
+ "area": 173952,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13056,
+ "image_id": 2428,
+ "bbox": [
+ 394,
+ 414,
+ 117,
+ 96
+ ],
+ "category_id": 10,
+ "area": 18564,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13057,
+ "image_id": 2428,
+ "bbox": [
+ 440,
+ 3,
+ 69,
+ 100
+ ],
+ "category_id": 10,
+ "area": 11495,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13058,
+ "image_id": 2428,
+ "bbox": [
+ 8,
+ 299,
+ 46,
+ 88
+ ],
+ "category_id": 10,
+ "area": 6804,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13059,
+ "image_id": 2428,
+ "bbox": [
+ 56,
+ 280,
+ 64,
+ 117
+ ],
+ "category_id": 10,
+ "area": 12543,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13065,
+ "image_id": 2430,
+ "bbox": [
+ 152,
+ 207,
+ 90,
+ 112
+ ],
+ "category_id": 8,
+ "area": 35708,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13066,
+ "image_id": 2430,
+ "bbox": [
+ 363,
+ 152,
+ 111,
+ 119
+ ],
+ "category_id": 8,
+ "area": 46593,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13085,
+ "image_id": 2433,
+ "bbox": [
+ 1,
+ 2,
+ 135,
+ 127
+ ],
+ "category_id": 6,
+ "area": 277200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13086,
+ "image_id": 2433,
+ "bbox": [
+ 113,
+ 84,
+ 98,
+ 136
+ ],
+ "category_id": 6,
+ "area": 214776,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13087,
+ "image_id": 2433,
+ "bbox": [
+ 122,
+ 226,
+ 114,
+ 99
+ ],
+ "category_id": 6,
+ "area": 183885,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13088,
+ "image_id": 2433,
+ "bbox": [
+ 78,
+ 328,
+ 113,
+ 95
+ ],
+ "category_id": 6,
+ "area": 174437,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13089,
+ "image_id": 2433,
+ "bbox": [
+ 0,
+ 375,
+ 40,
+ 102
+ ],
+ "category_id": 6,
+ "area": 67260,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13090,
+ "image_id": 2433,
+ "bbox": [
+ 229,
+ 2,
+ 101,
+ 189
+ ],
+ "category_id": 6,
+ "area": 306726,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13091,
+ "image_id": 2433,
+ "bbox": [
+ 227,
+ 181,
+ 84,
+ 79
+ ],
+ "category_id": 6,
+ "area": 107640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13092,
+ "image_id": 2433,
+ "bbox": [
+ 250,
+ 255,
+ 127,
+ 82
+ ],
+ "category_id": 6,
+ "area": 168740,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13093,
+ "image_id": 2433,
+ "bbox": [
+ 329,
+ 24,
+ 28,
+ 80
+ ],
+ "category_id": 6,
+ "area": 36841,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13094,
+ "image_id": 2433,
+ "bbox": [
+ 345,
+ 75,
+ 65,
+ 70
+ ],
+ "category_id": 6,
+ "area": 73326,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13095,
+ "image_id": 2433,
+ "bbox": [
+ 412,
+ 66,
+ 95,
+ 96
+ ],
+ "category_id": 6,
+ "area": 147294,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13096,
+ "image_id": 2433,
+ "bbox": [
+ 441,
+ 254,
+ 70,
+ 103
+ ],
+ "category_id": 6,
+ "area": 116025,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13103,
+ "image_id": 2434,
+ "bbox": [
+ 0,
+ 267,
+ 343,
+ 160
+ ],
+ "category_id": 9,
+ "area": 176000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13104,
+ "image_id": 2435,
+ "bbox": [
+ 56,
+ 179,
+ 210,
+ 327
+ ],
+ "category_id": 9,
+ "area": 71536,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13105,
+ "image_id": 2435,
+ "bbox": [
+ 144,
+ 106,
+ 330,
+ 324
+ ],
+ "category_id": 10,
+ "area": 111097,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13106,
+ "image_id": 2436,
+ "bbox": [
+ 0,
+ 115,
+ 287,
+ 159
+ ],
+ "category_id": 9,
+ "area": 77158,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13107,
+ "image_id": 2436,
+ "bbox": [
+ 0,
+ 208,
+ 190,
+ 302
+ ],
+ "category_id": 6,
+ "area": 97088,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13110,
+ "image_id": 2438,
+ "bbox": [
+ 30,
+ 309,
+ 211,
+ 199
+ ],
+ "category_id": 6,
+ "area": 792442,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13145,
+ "image_id": 2439,
+ "bbox": [
+ 2,
+ 1,
+ 63,
+ 107
+ ],
+ "category_id": 6,
+ "area": 101871,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13146,
+ "image_id": 2439,
+ "bbox": [
+ 2,
+ 98,
+ 54,
+ 48
+ ],
+ "category_id": 6,
+ "area": 39680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13147,
+ "image_id": 2439,
+ "bbox": [
+ 4,
+ 136,
+ 96,
+ 107
+ ],
+ "category_id": 6,
+ "area": 154693,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13148,
+ "image_id": 2439,
+ "bbox": [
+ 4,
+ 235,
+ 47,
+ 90
+ ],
+ "category_id": 6,
+ "area": 64447,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13149,
+ "image_id": 2439,
+ "bbox": [
+ 62,
+ 203,
+ 47,
+ 91
+ ],
+ "category_id": 6,
+ "area": 65408,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13150,
+ "image_id": 2439,
+ "bbox": [
+ 83,
+ 261,
+ 48,
+ 96
+ ],
+ "category_id": 6,
+ "area": 69916,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13151,
+ "image_id": 2439,
+ "bbox": [
+ 142,
+ 285,
+ 54,
+ 86
+ ],
+ "category_id": 6,
+ "area": 70612,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13152,
+ "image_id": 2439,
+ "bbox": [
+ 58,
+ 62,
+ 119,
+ 86
+ ],
+ "category_id": 6,
+ "area": 155402,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13153,
+ "image_id": 2439,
+ "bbox": [
+ 88,
+ 113,
+ 58,
+ 154
+ ],
+ "category_id": 6,
+ "area": 135904,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13154,
+ "image_id": 2439,
+ "bbox": [
+ 144,
+ 153,
+ 73,
+ 103
+ ],
+ "category_id": 6,
+ "area": 113533,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13155,
+ "image_id": 2439,
+ "bbox": [
+ 145,
+ 439,
+ 67,
+ 70
+ ],
+ "category_id": 6,
+ "area": 71959,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13156,
+ "image_id": 2439,
+ "bbox": [
+ 218,
+ 11,
+ 76,
+ 73
+ ],
+ "category_id": 6,
+ "area": 84609,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13157,
+ "image_id": 2439,
+ "bbox": [
+ 256,
+ 4,
+ 194,
+ 156
+ ],
+ "category_id": 6,
+ "area": 454908,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13158,
+ "image_id": 2439,
+ "bbox": [
+ 188,
+ 107,
+ 61,
+ 115
+ ],
+ "category_id": 6,
+ "area": 106848,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13159,
+ "image_id": 2439,
+ "bbox": [
+ 240,
+ 385,
+ 93,
+ 122
+ ],
+ "category_id": 6,
+ "area": 172088,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13172,
+ "image_id": 2440,
+ "bbox": [
+ 79,
+ 155,
+ 289,
+ 308
+ ],
+ "category_id": 9,
+ "area": 222202,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13173,
+ "image_id": 2441,
+ "bbox": [
+ 112,
+ 229,
+ 300,
+ 282
+ ],
+ "category_id": 9,
+ "area": 133095,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13174,
+ "image_id": 2441,
+ "bbox": [
+ 165,
+ 0,
+ 279,
+ 329
+ ],
+ "category_id": 6,
+ "area": 144522,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13175,
+ "image_id": 2441,
+ "bbox": [
+ 74,
+ 296,
+ 66,
+ 175
+ ],
+ "category_id": 6,
+ "area": 18408,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13176,
+ "image_id": 2441,
+ "bbox": [
+ 385,
+ 281,
+ 56,
+ 229
+ ],
+ "category_id": 6,
+ "area": 20416,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13177,
+ "image_id": 2441,
+ "bbox": [
+ 425,
+ 133,
+ 86,
+ 183
+ ],
+ "category_id": 6,
+ "area": 24975,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13178,
+ "image_id": 2442,
+ "bbox": [
+ 69,
+ 214,
+ 215,
+ 155
+ ],
+ "category_id": 8,
+ "area": 103145,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13179,
+ "image_id": 2442,
+ "bbox": [
+ 208,
+ 122,
+ 217,
+ 146
+ ],
+ "category_id": 8,
+ "area": 98406,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13209,
+ "image_id": 2445,
+ "bbox": [
+ 146,
+ 74,
+ 314,
+ 201
+ ],
+ "category_id": 9,
+ "area": 231840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13218,
+ "image_id": 2445,
+ "bbox": [
+ 236,
+ 394,
+ 31,
+ 40
+ ],
+ "category_id": 6,
+ "area": 4672,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13219,
+ "image_id": 2445,
+ "bbox": [
+ 194,
+ 391,
+ 39,
+ 85
+ ],
+ "category_id": 6,
+ "area": 12236,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13220,
+ "image_id": 2445,
+ "bbox": [
+ 144,
+ 327,
+ 29,
+ 66
+ ],
+ "category_id": 6,
+ "area": 7072,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13221,
+ "image_id": 2445,
+ "bbox": [
+ 168,
+ 320,
+ 45,
+ 91
+ ],
+ "category_id": 6,
+ "area": 15158,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13222,
+ "image_id": 2445,
+ "bbox": [
+ 238,
+ 329,
+ 46,
+ 60
+ ],
+ "category_id": 6,
+ "area": 10340,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13223,
+ "image_id": 2445,
+ "bbox": [
+ 346,
+ 398,
+ 59,
+ 75
+ ],
+ "category_id": 6,
+ "area": 16402,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13224,
+ "image_id": 2445,
+ "bbox": [
+ 261,
+ 396,
+ 59,
+ 75
+ ],
+ "category_id": 6,
+ "area": 16402,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13225,
+ "image_id": 2445,
+ "bbox": [
+ 392,
+ 471,
+ 58,
+ 40
+ ],
+ "category_id": 6,
+ "area": 8832,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13226,
+ "image_id": 2445,
+ "bbox": [
+ 440,
+ 460,
+ 31,
+ 40
+ ],
+ "category_id": 6,
+ "area": 4672,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13227,
+ "image_id": 2445,
+ "bbox": [
+ 369,
+ 314,
+ 86,
+ 90
+ ],
+ "category_id": 6,
+ "area": 28826,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13228,
+ "image_id": 2445,
+ "bbox": [
+ 404,
+ 282,
+ 59,
+ 75
+ ],
+ "category_id": 6,
+ "area": 16402,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13229,
+ "image_id": 2445,
+ "bbox": [
+ 347,
+ 281,
+ 59,
+ 75
+ ],
+ "category_id": 6,
+ "area": 16402,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13230,
+ "image_id": 2445,
+ "bbox": [
+ 279,
+ 300,
+ 59,
+ 75
+ ],
+ "category_id": 6,
+ "area": 16402,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13231,
+ "image_id": 2445,
+ "bbox": [
+ 56,
+ 344,
+ 30,
+ 58
+ ],
+ "category_id": 6,
+ "area": 6624,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13250,
+ "image_id": 2446,
+ "bbox": [
+ 270,
+ 221,
+ 125,
+ 243
+ ],
+ "category_id": 6,
+ "area": 167860,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13251,
+ "image_id": 2446,
+ "bbox": [
+ 207,
+ 297,
+ 73,
+ 214
+ ],
+ "category_id": 6,
+ "area": 87168,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13252,
+ "image_id": 2446,
+ "bbox": [
+ 7,
+ 13,
+ 172,
+ 365
+ ],
+ "category_id": 6,
+ "area": 348336,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13253,
+ "image_id": 2446,
+ "bbox": [
+ 140,
+ 4,
+ 77,
+ 205
+ ],
+ "category_id": 6,
+ "area": 87584,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13254,
+ "image_id": 2446,
+ "bbox": [
+ 150,
+ 305,
+ 132,
+ 126
+ ],
+ "category_id": 6,
+ "area": 92843,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13259,
+ "image_id": 2449,
+ "bbox": [
+ 15,
+ 218,
+ 236,
+ 288
+ ],
+ "category_id": 6,
+ "area": 62211,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13260,
+ "image_id": 2449,
+ "bbox": [
+ 252,
+ 144,
+ 129,
+ 147
+ ],
+ "category_id": 6,
+ "area": 17374,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13261,
+ "image_id": 2449,
+ "bbox": [
+ 352,
+ 196,
+ 79,
+ 133
+ ],
+ "category_id": 6,
+ "area": 9720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13262,
+ "image_id": 2449,
+ "bbox": [
+ 173,
+ 173,
+ 78,
+ 123
+ ],
+ "category_id": 6,
+ "area": 8800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13268,
+ "image_id": 2451,
+ "bbox": [
+ 96,
+ 398,
+ 182,
+ 93
+ ],
+ "category_id": 9,
+ "area": 26412,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13270,
+ "image_id": 2452,
+ "bbox": [
+ 22,
+ 78,
+ 455,
+ 424
+ ],
+ "category_id": 9,
+ "area": 88452,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13273,
+ "image_id": 2453,
+ "bbox": [
+ 325,
+ 283,
+ 186,
+ 187
+ ],
+ "category_id": 6,
+ "area": 39494,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13274,
+ "image_id": 2453,
+ "bbox": [
+ 406,
+ 85,
+ 105,
+ 175
+ ],
+ "category_id": 6,
+ "area": 20909,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13275,
+ "image_id": 2454,
+ "bbox": [
+ 52,
+ 8,
+ 419,
+ 449
+ ],
+ "category_id": 6,
+ "area": 1232340,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13310,
+ "image_id": 2458,
+ "bbox": [
+ 0,
+ 57,
+ 170,
+ 117
+ ],
+ "category_id": 6,
+ "area": 353430,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13311,
+ "image_id": 2458,
+ "bbox": [
+ 104,
+ 52,
+ 65,
+ 57
+ ],
+ "category_id": 6,
+ "area": 65700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13312,
+ "image_id": 2458,
+ "bbox": [
+ 175,
+ 58,
+ 92,
+ 72
+ ],
+ "category_id": 6,
+ "area": 118976,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13313,
+ "image_id": 2458,
+ "bbox": [
+ 244,
+ 68,
+ 98,
+ 170
+ ],
+ "category_id": 6,
+ "area": 297466,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13314,
+ "image_id": 2458,
+ "bbox": [
+ 331,
+ 4,
+ 86,
+ 105
+ ],
+ "category_id": 6,
+ "area": 161408,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13315,
+ "image_id": 2458,
+ "bbox": [
+ 350,
+ 88,
+ 31,
+ 74
+ ],
+ "category_id": 6,
+ "area": 41748,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13316,
+ "image_id": 2458,
+ "bbox": [
+ 0,
+ 224,
+ 49,
+ 43
+ ],
+ "category_id": 6,
+ "area": 37570,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13317,
+ "image_id": 2458,
+ "bbox": [
+ 0,
+ 374,
+ 52,
+ 119
+ ],
+ "category_id": 6,
+ "area": 111156,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13318,
+ "image_id": 2458,
+ "bbox": [
+ 126,
+ 127,
+ 107,
+ 135
+ ],
+ "category_id": 6,
+ "area": 257816,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13319,
+ "image_id": 2458,
+ "bbox": [
+ 132,
+ 262,
+ 120,
+ 90
+ ],
+ "category_id": 6,
+ "area": 193320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13320,
+ "image_id": 2458,
+ "bbox": [
+ 268,
+ 291,
+ 132,
+ 75
+ ],
+ "category_id": 6,
+ "area": 177608,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13321,
+ "image_id": 2458,
+ "bbox": [
+ 333,
+ 128,
+ 80,
+ 61
+ ],
+ "category_id": 6,
+ "area": 86760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13322,
+ "image_id": 2458,
+ "bbox": [
+ 467,
+ 299,
+ 42,
+ 80
+ ],
+ "category_id": 6,
+ "area": 60610,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13340,
+ "image_id": 2459,
+ "bbox": [
+ 166,
+ 153,
+ 290,
+ 312
+ ],
+ "category_id": 6,
+ "area": 378246,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13341,
+ "image_id": 2460,
+ "bbox": [
+ 40,
+ 256,
+ 234,
+ 143
+ ],
+ "category_id": 9,
+ "area": 57069,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13342,
+ "image_id": 2460,
+ "bbox": [
+ 374,
+ 142,
+ 99,
+ 252
+ ],
+ "category_id": 9,
+ "area": 42502,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13343,
+ "image_id": 2461,
+ "bbox": [
+ 140,
+ 144,
+ 240,
+ 206
+ ],
+ "category_id": 9,
+ "area": 221769,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13344,
+ "image_id": 2462,
+ "bbox": [
+ 53,
+ 288,
+ 81,
+ 72
+ ],
+ "category_id": 10,
+ "area": 14006,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13346,
+ "image_id": 2462,
+ "bbox": [
+ 228,
+ 136,
+ 283,
+ 162
+ ],
+ "category_id": 9,
+ "area": 109604,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13352,
+ "image_id": 2465,
+ "bbox": [
+ 84,
+ 89,
+ 55,
+ 45
+ ],
+ "category_id": 8,
+ "area": 16974,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13353,
+ "image_id": 2465,
+ "bbox": [
+ 107,
+ 155,
+ 89,
+ 61
+ ],
+ "category_id": 8,
+ "area": 37296,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13354,
+ "image_id": 2465,
+ "bbox": [
+ 201,
+ 84,
+ 33,
+ 37
+ ],
+ "category_id": 8,
+ "area": 8694,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13355,
+ "image_id": 2465,
+ "bbox": [
+ 216,
+ 180,
+ 55,
+ 130
+ ],
+ "category_id": 8,
+ "area": 48822,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13356,
+ "image_id": 2465,
+ "bbox": [
+ 312,
+ 124,
+ 58,
+ 89
+ ],
+ "category_id": 8,
+ "area": 35478,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13357,
+ "image_id": 2466,
+ "bbox": [
+ 77,
+ 81,
+ 371,
+ 367
+ ],
+ "category_id": 8,
+ "area": 444276,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13384,
+ "image_id": 2469,
+ "bbox": [
+ 24,
+ 249,
+ 291,
+ 162
+ ],
+ "category_id": 9,
+ "area": 108900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13386,
+ "image_id": 2470,
+ "bbox": [
+ 1,
+ 49,
+ 348,
+ 459
+ ],
+ "category_id": 9,
+ "area": 562020,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13387,
+ "image_id": 2470,
+ "bbox": [
+ 149,
+ 172,
+ 360,
+ 333
+ ],
+ "category_id": 9,
+ "area": 422100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13395,
+ "image_id": 2472,
+ "bbox": [
+ 266,
+ 135,
+ 111,
+ 54
+ ],
+ "category_id": 9,
+ "area": 21483,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13396,
+ "image_id": 2472,
+ "bbox": [
+ 255,
+ 125,
+ 134,
+ 148
+ ],
+ "category_id": 9,
+ "area": 70015,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13399,
+ "image_id": 2474,
+ "bbox": [
+ 76,
+ 204,
+ 222,
+ 216
+ ],
+ "category_id": 8,
+ "area": 168165,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13400,
+ "image_id": 2474,
+ "bbox": [
+ 191,
+ 150,
+ 215,
+ 258
+ ],
+ "category_id": 8,
+ "area": 194756,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13401,
+ "image_id": 2474,
+ "bbox": [
+ 303,
+ 12,
+ 106,
+ 122
+ ],
+ "category_id": 8,
+ "area": 45657,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13442,
+ "image_id": 2478,
+ "bbox": [
+ 66,
+ 163,
+ 119,
+ 80
+ ],
+ "category_id": 8,
+ "area": 75543,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13443,
+ "image_id": 2478,
+ "bbox": [
+ 185,
+ 212,
+ 175,
+ 237
+ ],
+ "category_id": 8,
+ "area": 329658,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13444,
+ "image_id": 2478,
+ "bbox": [
+ 373,
+ 181,
+ 41,
+ 29
+ ],
+ "category_id": 8,
+ "area": 9702,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13445,
+ "image_id": 2478,
+ "bbox": [
+ 338,
+ 212,
+ 92,
+ 58
+ ],
+ "category_id": 8,
+ "area": 42681,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13446,
+ "image_id": 2478,
+ "bbox": [
+ 417,
+ 234,
+ 55,
+ 34
+ ],
+ "category_id": 8,
+ "area": 15184,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13447,
+ "image_id": 2478,
+ "bbox": [
+ 450,
+ 210,
+ 60,
+ 27
+ ],
+ "category_id": 8,
+ "area": 13452,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13456,
+ "image_id": 2480,
+ "bbox": [
+ 175,
+ 99,
+ 238,
+ 392
+ ],
+ "category_id": 8,
+ "area": 741060,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13457,
+ "image_id": 2480,
+ "bbox": [
+ 3,
+ 296,
+ 248,
+ 215
+ ],
+ "category_id": 6,
+ "area": 422220,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13490,
+ "image_id": 2484,
+ "bbox": [
+ 57,
+ 64,
+ 244,
+ 256
+ ],
+ "category_id": 8,
+ "area": 497014,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13491,
+ "image_id": 2484,
+ "bbox": [
+ 262,
+ 82,
+ 126,
+ 212
+ ],
+ "category_id": 8,
+ "area": 212826,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13492,
+ "image_id": 2485,
+ "bbox": [
+ 164,
+ 207,
+ 89,
+ 142
+ ],
+ "category_id": 8,
+ "area": 41070,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13493,
+ "image_id": 2485,
+ "bbox": [
+ 306,
+ 311,
+ 47,
+ 68
+ ],
+ "category_id": 8,
+ "area": 10413,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13494,
+ "image_id": 2486,
+ "bbox": [
+ 1,
+ 68,
+ 134,
+ 238
+ ],
+ "category_id": 8,
+ "area": 73395,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13495,
+ "image_id": 2486,
+ "bbox": [
+ 200,
+ 36,
+ 311,
+ 372
+ ],
+ "category_id": 8,
+ "area": 265188,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13511,
+ "image_id": 2489,
+ "bbox": [
+ 0,
+ 186,
+ 510,
+ 324
+ ],
+ "category_id": 9,
+ "area": 248325,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13512,
+ "image_id": 2489,
+ "bbox": [
+ 375,
+ 57,
+ 134,
+ 212
+ ],
+ "category_id": 6,
+ "area": 42946,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13513,
+ "image_id": 2489,
+ "bbox": [
+ 0,
+ 43,
+ 287,
+ 414
+ ],
+ "category_id": 6,
+ "area": 179025,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13521,
+ "image_id": 2492,
+ "bbox": [
+ 0,
+ 221,
+ 398,
+ 252
+ ],
+ "category_id": 8,
+ "area": 147186,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13522,
+ "image_id": 2492,
+ "bbox": [
+ 54,
+ 153,
+ 256,
+ 162
+ ],
+ "category_id": 8,
+ "area": 60918,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13523,
+ "image_id": 2492,
+ "bbox": [
+ 0,
+ 97,
+ 239,
+ 170
+ ],
+ "category_id": 8,
+ "area": 59850,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13524,
+ "image_id": 2492,
+ "bbox": [
+ 118,
+ 152,
+ 221,
+ 189
+ ],
+ "category_id": 8,
+ "area": 61623,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13562,
+ "image_id": 2494,
+ "bbox": [
+ 0,
+ 397,
+ 84,
+ 113
+ ],
+ "category_id": 6,
+ "area": 39480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13563,
+ "image_id": 2494,
+ "bbox": [
+ 90,
+ 413,
+ 85,
+ 89
+ ],
+ "category_id": 6,
+ "area": 31672,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13564,
+ "image_id": 2494,
+ "bbox": [
+ 114,
+ 284,
+ 65,
+ 60
+ ],
+ "category_id": 6,
+ "area": 16463,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13565,
+ "image_id": 2494,
+ "bbox": [
+ 348,
+ 419,
+ 114,
+ 92
+ ],
+ "category_id": 6,
+ "area": 43911,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13566,
+ "image_id": 2494,
+ "bbox": [
+ 246,
+ 313,
+ 40,
+ 81
+ ],
+ "category_id": 6,
+ "area": 13500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13567,
+ "image_id": 2494,
+ "bbox": [
+ 220,
+ 260,
+ 62,
+ 73
+ ],
+ "category_id": 6,
+ "area": 18876,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13568,
+ "image_id": 2494,
+ "bbox": [
+ 363,
+ 227,
+ 102,
+ 131
+ ],
+ "category_id": 6,
+ "area": 56026,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13569,
+ "image_id": 2494,
+ "bbox": [
+ 438,
+ 232,
+ 72,
+ 91
+ ],
+ "category_id": 6,
+ "area": 27664,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13570,
+ "image_id": 2494,
+ "bbox": [
+ 295,
+ 163,
+ 67,
+ 70
+ ],
+ "category_id": 6,
+ "area": 19488,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13571,
+ "image_id": 2495,
+ "bbox": [
+ 23,
+ 115,
+ 117,
+ 233
+ ],
+ "category_id": 8,
+ "area": 68930,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13572,
+ "image_id": 2495,
+ "bbox": [
+ 282,
+ 166,
+ 228,
+ 153
+ ],
+ "category_id": 8,
+ "area": 88000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13573,
+ "image_id": 2495,
+ "bbox": [
+ 444,
+ 352,
+ 67,
+ 143
+ ],
+ "category_id": 6,
+ "area": 24497,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13575,
+ "image_id": 2496,
+ "bbox": [
+ 64,
+ 207,
+ 415,
+ 245
+ ],
+ "category_id": 8,
+ "area": 358455,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13584,
+ "image_id": 2498,
+ "bbox": [
+ 0,
+ 0,
+ 153,
+ 88
+ ],
+ "category_id": 8,
+ "area": 104859,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13585,
+ "image_id": 2498,
+ "bbox": [
+ 0,
+ 224,
+ 101,
+ 170
+ ],
+ "category_id": 8,
+ "area": 133731,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13586,
+ "image_id": 2498,
+ "bbox": [
+ 149,
+ 0,
+ 231,
+ 314
+ ],
+ "category_id": 8,
+ "area": 561385,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13587,
+ "image_id": 2498,
+ "bbox": [
+ 282,
+ 362,
+ 99,
+ 125
+ ],
+ "category_id": 6,
+ "area": 96607,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13589,
+ "image_id": 2499,
+ "bbox": [
+ 292,
+ 92,
+ 216,
+ 414
+ ],
+ "category_id": 9,
+ "area": 315986,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13590,
+ "image_id": 2499,
+ "bbox": [
+ 0,
+ 0,
+ 427,
+ 427
+ ],
+ "category_id": 9,
+ "area": 641868,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13591,
+ "image_id": 2499,
+ "bbox": [
+ 2,
+ 42,
+ 317,
+ 445
+ ],
+ "category_id": 9,
+ "area": 497044,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13592,
+ "image_id": 2500,
+ "bbox": [
+ 162,
+ 71,
+ 242,
+ 374
+ ],
+ "category_id": 9,
+ "area": 319362,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13596,
+ "image_id": 2501,
+ "bbox": [
+ 71,
+ 9,
+ 401,
+ 433
+ ],
+ "category_id": 9,
+ "area": 1121076,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13598,
+ "image_id": 2502,
+ "bbox": [
+ 150,
+ 413,
+ 162,
+ 97
+ ],
+ "category_id": 6,
+ "area": 55759,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13599,
+ "image_id": 2502,
+ "bbox": [
+ 358,
+ 226,
+ 152,
+ 285
+ ],
+ "category_id": 6,
+ "area": 152781,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13600,
+ "image_id": 2502,
+ "bbox": [
+ 114,
+ 232,
+ 251,
+ 197
+ ],
+ "category_id": 6,
+ "area": 174584,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13601,
+ "image_id": 2502,
+ "bbox": [
+ 0,
+ 219,
+ 163,
+ 278
+ ],
+ "category_id": 8,
+ "area": 159919,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13606,
+ "image_id": 2504,
+ "bbox": [
+ 134,
+ 115,
+ 197,
+ 333
+ ],
+ "category_id": 8,
+ "area": 70782,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13607,
+ "image_id": 2504,
+ "bbox": [
+ 3,
+ 209,
+ 229,
+ 302
+ ],
+ "category_id": 6,
+ "area": 74496,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13608,
+ "image_id": 2504,
+ "bbox": [
+ 200,
+ 380,
+ 230,
+ 124
+ ],
+ "category_id": 6,
+ "area": 30765,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13614,
+ "image_id": 2506,
+ "bbox": [
+ 209,
+ 189,
+ 117,
+ 129
+ ],
+ "category_id": 9,
+ "area": 20383,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13616,
+ "image_id": 2507,
+ "bbox": [
+ 101,
+ 188,
+ 208,
+ 154
+ ],
+ "category_id": 8,
+ "area": 42558,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13617,
+ "image_id": 2507,
+ "bbox": [
+ 274,
+ 200,
+ 237,
+ 224
+ ],
+ "category_id": 8,
+ "area": 70347,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13618,
+ "image_id": 2507,
+ "bbox": [
+ 0,
+ 289,
+ 75,
+ 217
+ ],
+ "category_id": 6,
+ "area": 21625,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13621,
+ "image_id": 2509,
+ "bbox": [
+ 28,
+ 300,
+ 69,
+ 75
+ ],
+ "category_id": 6,
+ "area": 36640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13622,
+ "image_id": 2509,
+ "bbox": [
+ 384,
+ 29,
+ 17,
+ 40
+ ],
+ "category_id": 6,
+ "area": 5015,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13623,
+ "image_id": 2509,
+ "bbox": [
+ 446,
+ 23,
+ 10,
+ 22
+ ],
+ "category_id": 6,
+ "area": 1680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13624,
+ "image_id": 2509,
+ "bbox": [
+ 415,
+ 44,
+ 14,
+ 21
+ ],
+ "category_id": 6,
+ "area": 2208,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13625,
+ "image_id": 2509,
+ "bbox": [
+ 458,
+ 52,
+ 9,
+ 39
+ ],
+ "category_id": 6,
+ "area": 2688,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13626,
+ "image_id": 2509,
+ "bbox": [
+ 465,
+ 1,
+ 15,
+ 27
+ ],
+ "category_id": 6,
+ "area": 2900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13627,
+ "image_id": 2509,
+ "bbox": [
+ 0,
+ 251,
+ 37,
+ 98
+ ],
+ "category_id": 6,
+ "area": 25584,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13630,
+ "image_id": 2510,
+ "bbox": [
+ 0,
+ 250,
+ 338,
+ 175
+ ],
+ "category_id": 9,
+ "area": 60480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13632,
+ "image_id": 2511,
+ "bbox": [
+ 15,
+ 79,
+ 298,
+ 298
+ ],
+ "category_id": 8,
+ "area": 110976,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13633,
+ "image_id": 2511,
+ "bbox": [
+ 317,
+ 230,
+ 184,
+ 170
+ ],
+ "category_id": 8,
+ "area": 39060,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13648,
+ "image_id": 2514,
+ "bbox": [
+ 103,
+ 186,
+ 237,
+ 289
+ ],
+ "category_id": 8,
+ "area": 361089,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13649,
+ "image_id": 2514,
+ "bbox": [
+ 255,
+ 243,
+ 256,
+ 266
+ ],
+ "category_id": 6,
+ "area": 358600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13650,
+ "image_id": 2514,
+ "bbox": [
+ 0,
+ 0,
+ 287,
+ 509
+ ],
+ "category_id": 6,
+ "area": 767760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13651,
+ "image_id": 2515,
+ "bbox": [
+ 0,
+ 0,
+ 321,
+ 394
+ ],
+ "category_id": 6,
+ "area": 161396,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13653,
+ "image_id": 2516,
+ "bbox": [
+ 69,
+ 70,
+ 279,
+ 250
+ ],
+ "category_id": 9,
+ "area": 170515,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13665,
+ "image_id": 2519,
+ "bbox": [
+ 114,
+ 171,
+ 39,
+ 67
+ ],
+ "category_id": 8,
+ "area": 3456,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13666,
+ "image_id": 2519,
+ "bbox": [
+ 193,
+ 171,
+ 46,
+ 89
+ ],
+ "category_id": 8,
+ "area": 5396,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13667,
+ "image_id": 2519,
+ "bbox": [
+ 304,
+ 218,
+ 63,
+ 114
+ ],
+ "category_id": 8,
+ "area": 9373,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13668,
+ "image_id": 2519,
+ "bbox": [
+ 221,
+ 315,
+ 103,
+ 194
+ ],
+ "category_id": 6,
+ "area": 26040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13669,
+ "image_id": 2519,
+ "bbox": [
+ 153,
+ 464,
+ 75,
+ 44
+ ],
+ "category_id": 6,
+ "area": 4305,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13670,
+ "image_id": 2519,
+ "bbox": [
+ 162,
+ 339,
+ 64,
+ 138
+ ],
+ "category_id": 6,
+ "area": 11550,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13679,
+ "image_id": 2522,
+ "bbox": [
+ 222,
+ 166,
+ 249,
+ 324
+ ],
+ "category_id": 9,
+ "area": 226275,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13683,
+ "image_id": 2522,
+ "bbox": [
+ 67,
+ 0,
+ 228,
+ 184
+ ],
+ "category_id": 6,
+ "area": 118326,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13684,
+ "image_id": 2523,
+ "bbox": [
+ 106,
+ 200,
+ 56,
+ 58
+ ],
+ "category_id": 8,
+ "area": 26164,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13685,
+ "image_id": 2523,
+ "bbox": [
+ 139,
+ 205,
+ 262,
+ 284
+ ],
+ "category_id": 8,
+ "area": 590400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13686,
+ "image_id": 2523,
+ "bbox": [
+ 70,
+ 410,
+ 273,
+ 101
+ ],
+ "category_id": 8,
+ "area": 219564,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13687,
+ "image_id": 2524,
+ "bbox": [
+ 155,
+ 217,
+ 99,
+ 125
+ ],
+ "category_id": 8,
+ "area": 43824,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13688,
+ "image_id": 2524,
+ "bbox": [
+ 303,
+ 211,
+ 181,
+ 148
+ ],
+ "category_id": 8,
+ "area": 94432,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13689,
+ "image_id": 2524,
+ "bbox": [
+ 478,
+ 308,
+ 32,
+ 49
+ ],
+ "category_id": 6,
+ "area": 5670,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13694,
+ "image_id": 2526,
+ "bbox": [
+ 0,
+ 0,
+ 511,
+ 298
+ ],
+ "category_id": 6,
+ "area": 374784,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13711,
+ "image_id": 2528,
+ "bbox": [
+ 3,
+ 0,
+ 508,
+ 408
+ ],
+ "category_id": 9,
+ "area": 419015,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13734,
+ "image_id": 2531,
+ "bbox": [
+ 182,
+ 151,
+ 164,
+ 68
+ ],
+ "category_id": 9,
+ "area": 11403,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13748,
+ "image_id": 2534,
+ "bbox": [
+ 86,
+ 110,
+ 345,
+ 357
+ ],
+ "category_id": 10,
+ "area": 977725,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13767,
+ "image_id": 2535,
+ "bbox": [
+ 83,
+ 301,
+ 288,
+ 155
+ ],
+ "category_id": 9,
+ "area": 179423,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13783,
+ "image_id": 2538,
+ "bbox": [
+ 39,
+ 84,
+ 268,
+ 410
+ ],
+ "category_id": 8,
+ "area": 127922,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13804,
+ "image_id": 2541,
+ "bbox": [
+ 406,
+ 0,
+ 105,
+ 136
+ ],
+ "category_id": 10,
+ "area": 25284,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13805,
+ "image_id": 2541,
+ "bbox": [
+ 141,
+ 0,
+ 226,
+ 509
+ ],
+ "category_id": 10,
+ "area": 202377,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13806,
+ "image_id": 2542,
+ "bbox": [
+ 180,
+ 0,
+ 328,
+ 507
+ ],
+ "category_id": 9,
+ "area": 585373,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13809,
+ "image_id": 2544,
+ "bbox": [
+ 204,
+ 155,
+ 149,
+ 112
+ ],
+ "category_id": 8,
+ "area": 34584,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13810,
+ "image_id": 2544,
+ "bbox": [
+ 306,
+ 143,
+ 123,
+ 71
+ ],
+ "category_id": 8,
+ "area": 18228,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13811,
+ "image_id": 2544,
+ "bbox": [
+ 2,
+ 147,
+ 124,
+ 292
+ ],
+ "category_id": 6,
+ "area": 74774,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13812,
+ "image_id": 2544,
+ "bbox": [
+ 110,
+ 362,
+ 158,
+ 146
+ ],
+ "category_id": 6,
+ "area": 47816,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13828,
+ "image_id": 2547,
+ "bbox": [
+ 140,
+ 120,
+ 206,
+ 176
+ ],
+ "category_id": 8,
+ "area": 128216,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13829,
+ "image_id": 2548,
+ "bbox": [
+ 28,
+ 41,
+ 92,
+ 212
+ ],
+ "category_id": 10,
+ "area": 34916,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13830,
+ "image_id": 2548,
+ "bbox": [
+ 133,
+ 5,
+ 100,
+ 219
+ ],
+ "category_id": 10,
+ "area": 39270,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13831,
+ "image_id": 2548,
+ "bbox": [
+ 314,
+ 37,
+ 97,
+ 233
+ ],
+ "category_id": 10,
+ "area": 40363,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13832,
+ "image_id": 2548,
+ "bbox": [
+ 447,
+ 211,
+ 63,
+ 174
+ ],
+ "category_id": 10,
+ "area": 19539,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13833,
+ "image_id": 2548,
+ "bbox": [
+ 417,
+ 452,
+ 60,
+ 59
+ ],
+ "category_id": 10,
+ "area": 6384,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13834,
+ "image_id": 2549,
+ "bbox": [
+ 149,
+ 67,
+ 97,
+ 100
+ ],
+ "category_id": 8,
+ "area": 77958,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13835,
+ "image_id": 2549,
+ "bbox": [
+ 142,
+ 3,
+ 306,
+ 389
+ ],
+ "category_id": 8,
+ "area": 943656,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13836,
+ "image_id": 2549,
+ "bbox": [
+ 2,
+ 174,
+ 205,
+ 230
+ ],
+ "category_id": 8,
+ "area": 374503,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13837,
+ "image_id": 2550,
+ "bbox": [
+ 42,
+ 93,
+ 444,
+ 414
+ ],
+ "category_id": 9,
+ "area": 647130,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13838,
+ "image_id": 2551,
+ "bbox": [
+ 306,
+ 57,
+ 112,
+ 112
+ ],
+ "category_id": 8,
+ "area": 44274,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13839,
+ "image_id": 2551,
+ "bbox": [
+ 48,
+ 244,
+ 202,
+ 200
+ ],
+ "category_id": 8,
+ "area": 142186,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13840,
+ "image_id": 2551,
+ "bbox": [
+ 165,
+ 207,
+ 214,
+ 269
+ ],
+ "category_id": 8,
+ "area": 201695,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13848,
+ "image_id": 2554,
+ "bbox": [
+ 178,
+ 267,
+ 102,
+ 65
+ ],
+ "category_id": 8,
+ "area": 23552,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13849,
+ "image_id": 2554,
+ "bbox": [
+ 279,
+ 184,
+ 102,
+ 147
+ ],
+ "category_id": 8,
+ "area": 52736,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13850,
+ "image_id": 2554,
+ "bbox": [
+ 406,
+ 415,
+ 46,
+ 48
+ ],
+ "category_id": 6,
+ "area": 7820,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13851,
+ "image_id": 2554,
+ "bbox": [
+ 465,
+ 484,
+ 44,
+ 27
+ ],
+ "category_id": 6,
+ "area": 4368,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13852,
+ "image_id": 2555,
+ "bbox": [
+ 160,
+ 132,
+ 222,
+ 377
+ ],
+ "category_id": 10,
+ "area": 54390,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13856,
+ "image_id": 2556,
+ "bbox": [
+ 183,
+ 495,
+ 53,
+ 16
+ ],
+ "category_id": 6,
+ "area": 2424,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13889,
+ "image_id": 2561,
+ "bbox": [
+ 124,
+ 266,
+ 169,
+ 110
+ ],
+ "category_id": 8,
+ "area": 147722,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13890,
+ "image_id": 2561,
+ "bbox": [
+ 175,
+ 211,
+ 112,
+ 110
+ ],
+ "category_id": 8,
+ "area": 98280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13891,
+ "image_id": 2561,
+ "bbox": [
+ 292,
+ 169,
+ 128,
+ 108
+ ],
+ "category_id": 8,
+ "area": 110149,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13892,
+ "image_id": 2561,
+ "bbox": [
+ 429,
+ 187,
+ 82,
+ 75
+ ],
+ "category_id": 8,
+ "area": 49449,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13893,
+ "image_id": 2562,
+ "bbox": [
+ 0,
+ 0,
+ 256,
+ 511
+ ],
+ "category_id": 6,
+ "area": 458240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13894,
+ "image_id": 2562,
+ "bbox": [
+ 220,
+ 186,
+ 194,
+ 199
+ ],
+ "category_id": 8,
+ "area": 135315,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13895,
+ "image_id": 2562,
+ "bbox": [
+ 364,
+ 122,
+ 146,
+ 334
+ ],
+ "category_id": 8,
+ "area": 171756,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13904,
+ "image_id": 2565,
+ "bbox": [
+ 163,
+ 147,
+ 178,
+ 114
+ ],
+ "category_id": 8,
+ "area": 71520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13905,
+ "image_id": 2565,
+ "bbox": [
+ 219,
+ 254,
+ 158,
+ 169
+ ],
+ "category_id": 8,
+ "area": 93615,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13906,
+ "image_id": 2566,
+ "bbox": [
+ 77,
+ 306,
+ 177,
+ 204
+ ],
+ "category_id": 8,
+ "area": 42919,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13907,
+ "image_id": 2566,
+ "bbox": [
+ 247,
+ 396,
+ 114,
+ 71
+ ],
+ "category_id": 8,
+ "area": 9570,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13908,
+ "image_id": 2566,
+ "bbox": [
+ 341,
+ 405,
+ 99,
+ 61
+ ],
+ "category_id": 8,
+ "area": 7200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13909,
+ "image_id": 2566,
+ "bbox": [
+ 355,
+ 352,
+ 76,
+ 55
+ ],
+ "category_id": 8,
+ "area": 4995,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13911,
+ "image_id": 2567,
+ "bbox": [
+ 183,
+ 0,
+ 53,
+ 64
+ ],
+ "category_id": 8,
+ "area": 23807,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13912,
+ "image_id": 2567,
+ "bbox": [
+ 212,
+ 55,
+ 147,
+ 85
+ ],
+ "category_id": 8,
+ "area": 86592,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13913,
+ "image_id": 2567,
+ "bbox": [
+ 225,
+ 97,
+ 166,
+ 239
+ ],
+ "category_id": 8,
+ "area": 274664,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13914,
+ "image_id": 2567,
+ "bbox": [
+ 190,
+ 194,
+ 219,
+ 148
+ ],
+ "category_id": 8,
+ "area": 224604,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13915,
+ "image_id": 2567,
+ "bbox": [
+ 234,
+ 328,
+ 114,
+ 128
+ ],
+ "category_id": 8,
+ "area": 100584,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13936,
+ "image_id": 2569,
+ "bbox": [
+ 157,
+ 160,
+ 279,
+ 254
+ ],
+ "category_id": 8,
+ "area": 230376,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13955,
+ "image_id": 2572,
+ "bbox": [
+ 116,
+ 285,
+ 239,
+ 157
+ ],
+ "category_id": 8,
+ "area": 77096,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13956,
+ "image_id": 2572,
+ "bbox": [
+ 53,
+ 215,
+ 36,
+ 65
+ ],
+ "category_id": 8,
+ "area": 4928,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13957,
+ "image_id": 2572,
+ "bbox": [
+ 178,
+ 210,
+ 126,
+ 121
+ ],
+ "category_id": 8,
+ "area": 31524,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13963,
+ "image_id": 2572,
+ "bbox": [
+ 362,
+ 205,
+ 149,
+ 289
+ ],
+ "category_id": 6,
+ "area": 88818,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13969,
+ "image_id": 2574,
+ "bbox": [
+ 76,
+ 84,
+ 259,
+ 389
+ ],
+ "category_id": 10,
+ "area": 3319635,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13970,
+ "image_id": 2574,
+ "bbox": [
+ 326,
+ 116,
+ 122,
+ 285
+ ],
+ "category_id": 9,
+ "area": 1150920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13972,
+ "image_id": 2576,
+ "bbox": [
+ 25,
+ 100,
+ 460,
+ 406
+ ],
+ "category_id": 9,
+ "area": 658944,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13973,
+ "image_id": 2576,
+ "bbox": [
+ 428,
+ 305,
+ 81,
+ 204
+ ],
+ "category_id": 9,
+ "area": 58752,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13974,
+ "image_id": 2576,
+ "bbox": [
+ 132,
+ 0,
+ 308,
+ 354
+ ],
+ "category_id": 9,
+ "area": 384729,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13975,
+ "image_id": 2576,
+ "bbox": [
+ 193,
+ 0,
+ 273,
+ 322
+ ],
+ "category_id": 9,
+ "area": 309399,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13976,
+ "image_id": 2577,
+ "bbox": [
+ 100,
+ 96,
+ 103,
+ 225
+ ],
+ "category_id": 9,
+ "area": 81786,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13977,
+ "image_id": 2577,
+ "bbox": [
+ 164,
+ 253,
+ 133,
+ 255
+ ],
+ "category_id": 9,
+ "area": 119547,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13978,
+ "image_id": 2577,
+ "bbox": [
+ 396,
+ 410,
+ 80,
+ 96
+ ],
+ "category_id": 9,
+ "area": 27200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13979,
+ "image_id": 2577,
+ "bbox": [
+ 203,
+ 182,
+ 103,
+ 192
+ ],
+ "category_id": 9,
+ "area": 69660,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13980,
+ "image_id": 2577,
+ "bbox": [
+ 292,
+ 117,
+ 98,
+ 169
+ ],
+ "category_id": 9,
+ "area": 58555,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13981,
+ "image_id": 2577,
+ "bbox": [
+ 277,
+ 78,
+ 96,
+ 180
+ ],
+ "category_id": 9,
+ "area": 60960,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13991,
+ "image_id": 2579,
+ "bbox": [
+ 56,
+ 423,
+ 313,
+ 88
+ ],
+ "category_id": 6,
+ "area": 256533,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 13992,
+ "image_id": 2579,
+ "bbox": [
+ 287,
+ 237,
+ 134,
+ 86
+ ],
+ "category_id": 6,
+ "area": 107844,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14002,
+ "image_id": 2582,
+ "bbox": [
+ 331,
+ 93,
+ 103,
+ 153
+ ],
+ "category_id": 10,
+ "area": 39294,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14003,
+ "image_id": 2583,
+ "bbox": [
+ 91,
+ 106,
+ 275,
+ 356
+ ],
+ "category_id": 8,
+ "area": 775783,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14004,
+ "image_id": 2583,
+ "bbox": [
+ 397,
+ 275,
+ 49,
+ 55
+ ],
+ "category_id": 6,
+ "area": 21712,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14037,
+ "image_id": 2589,
+ "bbox": [
+ 3,
+ 94,
+ 507,
+ 407
+ ],
+ "category_id": 9,
+ "area": 530105,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14048,
+ "image_id": 2591,
+ "bbox": [
+ 0,
+ 223,
+ 32,
+ 155
+ ],
+ "category_id": 6,
+ "area": 11218,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14049,
+ "image_id": 2591,
+ "bbox": [
+ 204,
+ 293,
+ 302,
+ 218
+ ],
+ "category_id": 9,
+ "area": 148296,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14050,
+ "image_id": 2592,
+ "bbox": [
+ 71,
+ 133,
+ 439,
+ 377
+ ],
+ "category_id": 6,
+ "area": 955647,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14057,
+ "image_id": 2593,
+ "bbox": [
+ 0,
+ 0,
+ 341,
+ 511
+ ],
+ "category_id": 6,
+ "area": 1378762,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14079,
+ "image_id": 2596,
+ "bbox": [
+ 0,
+ 0,
+ 251,
+ 511
+ ],
+ "category_id": 6,
+ "area": 451532,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14080,
+ "image_id": 2596,
+ "bbox": [
+ 172,
+ 70,
+ 339,
+ 433
+ ],
+ "category_id": 8,
+ "area": 516432,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14126,
+ "image_id": 2602,
+ "bbox": [
+ 78,
+ 106,
+ 236,
+ 240
+ ],
+ "category_id": 9,
+ "area": 681120,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14127,
+ "image_id": 2602,
+ "bbox": [
+ 351,
+ 5,
+ 159,
+ 181
+ ],
+ "category_id": 6,
+ "area": 348255,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14130,
+ "image_id": 2603,
+ "bbox": [
+ 4,
+ 128,
+ 435,
+ 282
+ ],
+ "category_id": 8,
+ "area": 432333,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14134,
+ "image_id": 2605,
+ "bbox": [
+ 117,
+ 185,
+ 131,
+ 130
+ ],
+ "category_id": 8,
+ "area": 43498,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14135,
+ "image_id": 2605,
+ "bbox": [
+ 317,
+ 174,
+ 194,
+ 161
+ ],
+ "category_id": 8,
+ "area": 79875,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14145,
+ "image_id": 2607,
+ "bbox": [
+ 151,
+ 187,
+ 349,
+ 259
+ ],
+ "category_id": 8,
+ "area": 76557,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14169,
+ "image_id": 2611,
+ "bbox": [
+ 347,
+ 275,
+ 164,
+ 233
+ ],
+ "category_id": 6,
+ "area": 41989,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14182,
+ "image_id": 2615,
+ "bbox": [
+ 62,
+ 233,
+ 259,
+ 205
+ ],
+ "category_id": 9,
+ "area": 125856,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14190,
+ "image_id": 2617,
+ "bbox": [
+ 173,
+ 218,
+ 166,
+ 139
+ ],
+ "category_id": 8,
+ "area": 184375,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14191,
+ "image_id": 2617,
+ "bbox": [
+ 174,
+ 350,
+ 124,
+ 102
+ ],
+ "category_id": 8,
+ "area": 100905,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14202,
+ "image_id": 2620,
+ "bbox": [
+ 146,
+ 211,
+ 123,
+ 257
+ ],
+ "category_id": 8,
+ "area": 117114,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14204,
+ "image_id": 2621,
+ "bbox": [
+ 170,
+ 59,
+ 341,
+ 300
+ ],
+ "category_id": 9,
+ "area": 222870,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14205,
+ "image_id": 2621,
+ "bbox": [
+ 200,
+ 218,
+ 310,
+ 292
+ ],
+ "category_id": 9,
+ "area": 197220,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14206,
+ "image_id": 2621,
+ "bbox": [
+ 46,
+ 194,
+ 221,
+ 175
+ ],
+ "category_id": 9,
+ "area": 84360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14217,
+ "image_id": 2624,
+ "bbox": [
+ 280,
+ 154,
+ 69,
+ 85
+ ],
+ "category_id": 8,
+ "area": 7840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14226,
+ "image_id": 2627,
+ "bbox": [
+ 89,
+ 106,
+ 122,
+ 144
+ ],
+ "category_id": 8,
+ "area": 85968,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14227,
+ "image_id": 2627,
+ "bbox": [
+ 131,
+ 146,
+ 145,
+ 146
+ ],
+ "category_id": 8,
+ "area": 102896,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14228,
+ "image_id": 2627,
+ "bbox": [
+ 319,
+ 28,
+ 66,
+ 281
+ ],
+ "category_id": 8,
+ "area": 91140,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14229,
+ "image_id": 2627,
+ "bbox": [
+ 370,
+ 0,
+ 141,
+ 256
+ ],
+ "category_id": 8,
+ "area": 175720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14230,
+ "image_id": 2627,
+ "bbox": [
+ 215,
+ 195,
+ 146,
+ 279
+ ],
+ "category_id": 8,
+ "area": 199326,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14242,
+ "image_id": 2629,
+ "bbox": [
+ 0,
+ 297,
+ 29,
+ 135
+ ],
+ "category_id": 6,
+ "area": 38517,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14243,
+ "image_id": 2629,
+ "bbox": [
+ 146,
+ 5,
+ 94,
+ 92
+ ],
+ "category_id": 6,
+ "area": 85432,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14244,
+ "image_id": 2629,
+ "bbox": [
+ 126,
+ 88,
+ 114,
+ 135
+ ],
+ "category_id": 6,
+ "area": 152333,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14245,
+ "image_id": 2629,
+ "bbox": [
+ 78,
+ 225,
+ 80,
+ 133
+ ],
+ "category_id": 6,
+ "area": 105336,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14246,
+ "image_id": 2629,
+ "bbox": [
+ 298,
+ 0,
+ 30,
+ 49
+ ],
+ "category_id": 6,
+ "area": 14490,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14247,
+ "image_id": 2629,
+ "bbox": [
+ 256,
+ 33,
+ 101,
+ 103
+ ],
+ "category_id": 6,
+ "area": 102432,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14248,
+ "image_id": 2629,
+ "bbox": [
+ 292,
+ 131,
+ 147,
+ 113
+ ],
+ "category_id": 6,
+ "area": 162980,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14262,
+ "image_id": 2631,
+ "bbox": [
+ 94,
+ 180,
+ 325,
+ 206
+ ],
+ "category_id": 8,
+ "area": 200850,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14263,
+ "image_id": 2632,
+ "bbox": [
+ 148,
+ 274,
+ 146,
+ 161
+ ],
+ "category_id": 8,
+ "area": 82490,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14264,
+ "image_id": 2632,
+ "bbox": [
+ 223,
+ 331,
+ 155,
+ 119
+ ],
+ "category_id": 8,
+ "area": 64796,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14265,
+ "image_id": 2633,
+ "bbox": [
+ 0,
+ 116,
+ 281,
+ 176
+ ],
+ "category_id": 9,
+ "area": 83904,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14266,
+ "image_id": 2633,
+ "bbox": [
+ 0,
+ 214,
+ 308,
+ 296
+ ],
+ "category_id": 6,
+ "area": 154238,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14272,
+ "image_id": 2634,
+ "bbox": [
+ 36,
+ 281,
+ 181,
+ 89
+ ],
+ "category_id": 9,
+ "area": 35796,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14289,
+ "image_id": 2638,
+ "bbox": [
+ 239,
+ 25,
+ 112,
+ 342
+ ],
+ "category_id": 10,
+ "area": 63434,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14299,
+ "image_id": 2640,
+ "bbox": [
+ 106,
+ 176,
+ 326,
+ 211
+ ],
+ "category_id": 9,
+ "area": 206684,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14312,
+ "image_id": 2642,
+ "bbox": [
+ 489,
+ 340,
+ 21,
+ 80
+ ],
+ "category_id": 6,
+ "area": 2304,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14313,
+ "image_id": 2642,
+ "bbox": [
+ 58,
+ 239,
+ 62,
+ 96
+ ],
+ "category_id": 6,
+ "area": 8008,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14314,
+ "image_id": 2642,
+ "bbox": [
+ 112,
+ 230,
+ 47,
+ 48
+ ],
+ "category_id": 6,
+ "area": 3042,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14315,
+ "image_id": 2642,
+ "bbox": [
+ 190,
+ 283,
+ 78,
+ 228
+ ],
+ "category_id": 6,
+ "area": 23660,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14316,
+ "image_id": 2642,
+ "bbox": [
+ 264,
+ 125,
+ 66,
+ 131
+ ],
+ "category_id": 6,
+ "area": 11655,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14317,
+ "image_id": 2642,
+ "bbox": [
+ 334,
+ 233,
+ 28,
+ 40
+ ],
+ "category_id": 6,
+ "area": 1536,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14318,
+ "image_id": 2642,
+ "bbox": [
+ 411,
+ 259,
+ 79,
+ 51
+ ],
+ "category_id": 6,
+ "area": 5412,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14319,
+ "image_id": 2642,
+ "bbox": [
+ 295,
+ 249,
+ 37,
+ 72
+ ],
+ "category_id": 6,
+ "area": 3596,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14344,
+ "image_id": 2644,
+ "bbox": [
+ 139,
+ 52,
+ 341,
+ 459
+ ],
+ "category_id": 8,
+ "area": 546550,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14345,
+ "image_id": 2644,
+ "bbox": [
+ 98,
+ 217,
+ 206,
+ 210
+ ],
+ "category_id": 8,
+ "area": 151116,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14349,
+ "image_id": 2645,
+ "bbox": [
+ 0,
+ 114,
+ 273,
+ 155
+ ],
+ "category_id": 9,
+ "area": 71656,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14350,
+ "image_id": 2645,
+ "bbox": [
+ 0,
+ 197,
+ 180,
+ 313
+ ],
+ "category_id": 6,
+ "area": 95200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14363,
+ "image_id": 2648,
+ "bbox": [
+ 0,
+ 0,
+ 121,
+ 58
+ ],
+ "category_id": 6,
+ "area": 66976,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14364,
+ "image_id": 2648,
+ "bbox": [
+ 97,
+ 2,
+ 136,
+ 180
+ ],
+ "category_id": 6,
+ "area": 231570,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14365,
+ "image_id": 2648,
+ "bbox": [
+ 105,
+ 173,
+ 158,
+ 125
+ ],
+ "category_id": 6,
+ "area": 186300,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14366,
+ "image_id": 2648,
+ "bbox": [
+ 51,
+ 300,
+ 152,
+ 121
+ ],
+ "category_id": 6,
+ "area": 174348,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14367,
+ "image_id": 2648,
+ "bbox": [
+ 260,
+ 0,
+ 131,
+ 138
+ ],
+ "category_id": 6,
+ "area": 170754,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14368,
+ "image_id": 2648,
+ "bbox": [
+ 253,
+ 119,
+ 110,
+ 101
+ ],
+ "category_id": 6,
+ "area": 105840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14369,
+ "image_id": 2648,
+ "bbox": [
+ 299,
+ 211,
+ 151,
+ 106
+ ],
+ "category_id": 6,
+ "area": 151188,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14370,
+ "image_id": 2648,
+ "bbox": [
+ 411,
+ 0,
+ 74,
+ 66
+ ],
+ "category_id": 6,
+ "area": 46552,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14375,
+ "image_id": 2649,
+ "bbox": [
+ 111,
+ 254,
+ 98,
+ 79
+ ],
+ "category_id": 8,
+ "area": 11786,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14376,
+ "image_id": 2649,
+ "bbox": [
+ 191,
+ 176,
+ 62,
+ 130
+ ],
+ "category_id": 8,
+ "area": 12285,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14377,
+ "image_id": 2649,
+ "bbox": [
+ 234,
+ 175,
+ 104,
+ 172
+ ],
+ "category_id": 8,
+ "area": 27280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14393,
+ "image_id": 2652,
+ "bbox": [
+ 74,
+ 118,
+ 270,
+ 277
+ ],
+ "category_id": 8,
+ "area": 593618,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14394,
+ "image_id": 2652,
+ "bbox": [
+ 328,
+ 161,
+ 183,
+ 349
+ ],
+ "category_id": 6,
+ "area": 507006,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14395,
+ "image_id": 2652,
+ "bbox": [
+ 210,
+ 60,
+ 142,
+ 365
+ ],
+ "category_id": 6,
+ "area": 411180,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14396,
+ "image_id": 2652,
+ "bbox": [
+ 158,
+ 311,
+ 29,
+ 199
+ ],
+ "category_id": 6,
+ "area": 46310,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14397,
+ "image_id": 2652,
+ "bbox": [
+ 66,
+ 391,
+ 33,
+ 120
+ ],
+ "category_id": 6,
+ "area": 31875,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14398,
+ "image_id": 2653,
+ "bbox": [
+ 224,
+ 95,
+ 95,
+ 38
+ ],
+ "category_id": 8,
+ "area": 7876,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14399,
+ "image_id": 2653,
+ "bbox": [
+ 266,
+ 318,
+ 153,
+ 119
+ ],
+ "category_id": 8,
+ "area": 39319,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14400,
+ "image_id": 2654,
+ "bbox": [
+ 115,
+ 186,
+ 125,
+ 235
+ ],
+ "category_id": 9,
+ "area": 227465,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14405,
+ "image_id": 2656,
+ "bbox": [
+ 2,
+ 232,
+ 309,
+ 219
+ ],
+ "category_id": 9,
+ "area": 238392,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14406,
+ "image_id": 2656,
+ "bbox": [
+ 250,
+ 176,
+ 134,
+ 147
+ ],
+ "category_id": 9,
+ "area": 69680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14407,
+ "image_id": 2656,
+ "bbox": [
+ 290,
+ 367,
+ 219,
+ 137
+ ],
+ "category_id": 9,
+ "area": 105957,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14408,
+ "image_id": 2657,
+ "bbox": [
+ 170,
+ 208,
+ 96,
+ 124
+ ],
+ "category_id": 8,
+ "area": 41586,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14409,
+ "image_id": 2657,
+ "bbox": [
+ 270,
+ 281,
+ 78,
+ 70
+ ],
+ "category_id": 8,
+ "area": 19110,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14410,
+ "image_id": 2658,
+ "bbox": [
+ 173,
+ 170,
+ 272,
+ 296
+ ],
+ "category_id": 9,
+ "area": 206064,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14417,
+ "image_id": 2659,
+ "bbox": [
+ 38,
+ 0,
+ 473,
+ 286
+ ],
+ "category_id": 6,
+ "area": 329230,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14432,
+ "image_id": 2661,
+ "bbox": [
+ 166,
+ 265,
+ 135,
+ 241
+ ],
+ "category_id": 10,
+ "area": 42240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14433,
+ "image_id": 2661,
+ "bbox": [
+ 216,
+ 59,
+ 270,
+ 245
+ ],
+ "category_id": 9,
+ "area": 85800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14444,
+ "image_id": 2665,
+ "bbox": [
+ 398,
+ 314,
+ 112,
+ 169
+ ],
+ "category_id": 10,
+ "area": 17688,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14445,
+ "image_id": 2665,
+ "bbox": [
+ 174,
+ 168,
+ 58,
+ 67
+ ],
+ "category_id": 10,
+ "area": 3604,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14446,
+ "image_id": 2665,
+ "bbox": [
+ 40,
+ 275,
+ 78,
+ 102
+ ],
+ "category_id": 10,
+ "area": 7452,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14447,
+ "image_id": 2665,
+ "bbox": [
+ 360,
+ 408,
+ 52,
+ 73
+ ],
+ "category_id": 10,
+ "area": 3596,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14448,
+ "image_id": 2665,
+ "bbox": [
+ 0,
+ 394,
+ 52,
+ 74
+ ],
+ "category_id": 10,
+ "area": 3599,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14449,
+ "image_id": 2665,
+ "bbox": [
+ 351,
+ 280,
+ 27,
+ 34
+ ],
+ "category_id": 10,
+ "area": 864,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14450,
+ "image_id": 2665,
+ "bbox": [
+ 361,
+ 318,
+ 26,
+ 32
+ ],
+ "category_id": 10,
+ "area": 806,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14451,
+ "image_id": 2665,
+ "bbox": [
+ 383,
+ 195,
+ 25,
+ 32
+ ],
+ "category_id": 10,
+ "area": 780,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14452,
+ "image_id": 2665,
+ "bbox": [
+ 0,
+ 261,
+ 29,
+ 46
+ ],
+ "category_id": 10,
+ "area": 1295,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14453,
+ "image_id": 2665,
+ "bbox": [
+ 115,
+ 206,
+ 29,
+ 39
+ ],
+ "category_id": 10,
+ "area": 1054,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14454,
+ "image_id": 2665,
+ "bbox": [
+ 137,
+ 359,
+ 29,
+ 37
+ ],
+ "category_id": 10,
+ "area": 1050,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14455,
+ "image_id": 2665,
+ "bbox": [
+ 155,
+ 321,
+ 29,
+ 48
+ ],
+ "category_id": 10,
+ "area": 1330,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14456,
+ "image_id": 2665,
+ "bbox": [
+ 250,
+ 149,
+ 42,
+ 56
+ ],
+ "category_id": 10,
+ "area": 2250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14477,
+ "image_id": 2669,
+ "bbox": [
+ 7,
+ 54,
+ 323,
+ 457
+ ],
+ "category_id": 10,
+ "area": 1170792,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14483,
+ "image_id": 2671,
+ "bbox": [
+ 117,
+ 170,
+ 36,
+ 71
+ ],
+ "category_id": 8,
+ "area": 3420,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14484,
+ "image_id": 2671,
+ "bbox": [
+ 194,
+ 169,
+ 44,
+ 91
+ ],
+ "category_id": 8,
+ "area": 5402,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14485,
+ "image_id": 2671,
+ "bbox": [
+ 303,
+ 218,
+ 62,
+ 116
+ ],
+ "category_id": 8,
+ "area": 9672,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14486,
+ "image_id": 2671,
+ "bbox": [
+ 219,
+ 313,
+ 102,
+ 194
+ ],
+ "category_id": 6,
+ "area": 26350,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14487,
+ "image_id": 2671,
+ "bbox": [
+ 151,
+ 336,
+ 71,
+ 146
+ ],
+ "category_id": 6,
+ "area": 13923,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14488,
+ "image_id": 2671,
+ "bbox": [
+ 153,
+ 468,
+ 71,
+ 41
+ ],
+ "category_id": 6,
+ "area": 3894,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14490,
+ "image_id": 2672,
+ "bbox": [
+ 66,
+ 170,
+ 296,
+ 331
+ ],
+ "category_id": 9,
+ "area": 301774,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14491,
+ "image_id": 2673,
+ "bbox": [
+ 97,
+ 211,
+ 202,
+ 117
+ ],
+ "category_id": 8,
+ "area": 83148,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14492,
+ "image_id": 2673,
+ "bbox": [
+ 312,
+ 261,
+ 134,
+ 92
+ ],
+ "category_id": 8,
+ "area": 43550,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14493,
+ "image_id": 2674,
+ "bbox": [
+ 236,
+ 66,
+ 273,
+ 306
+ ],
+ "category_id": 9,
+ "area": 294804,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14494,
+ "image_id": 2674,
+ "bbox": [
+ 257,
+ 238,
+ 252,
+ 267
+ ],
+ "category_id": 9,
+ "area": 237632,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14495,
+ "image_id": 2674,
+ "bbox": [
+ 155,
+ 121,
+ 177,
+ 251
+ ],
+ "category_id": 9,
+ "area": 157176,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14500,
+ "image_id": 2675,
+ "bbox": [
+ 12,
+ 0,
+ 499,
+ 261
+ ],
+ "category_id": 6,
+ "area": 315400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14517,
+ "image_id": 2678,
+ "bbox": [
+ 115,
+ 258,
+ 134,
+ 215
+ ],
+ "category_id": 8,
+ "area": 100835,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14518,
+ "image_id": 2678,
+ "bbox": [
+ 263,
+ 228,
+ 137,
+ 130
+ ],
+ "category_id": 8,
+ "area": 62426,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14519,
+ "image_id": 2678,
+ "bbox": [
+ 6,
+ 188,
+ 287,
+ 319
+ ],
+ "category_id": 6,
+ "area": 320052,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14520,
+ "image_id": 2678,
+ "bbox": [
+ 109,
+ 39,
+ 141,
+ 161
+ ],
+ "category_id": 6,
+ "area": 79326,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14521,
+ "image_id": 2678,
+ "bbox": [
+ 256,
+ 22,
+ 190,
+ 200
+ ],
+ "category_id": 6,
+ "area": 132720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14522,
+ "image_id": 2678,
+ "bbox": [
+ 358,
+ 209,
+ 149,
+ 235
+ ],
+ "category_id": 6,
+ "area": 122059,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14537,
+ "image_id": 2680,
+ "bbox": [
+ 146,
+ 146,
+ 365,
+ 346
+ ],
+ "category_id": 9,
+ "area": 445118,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14538,
+ "image_id": 2680,
+ "bbox": [
+ 252,
+ 39,
+ 260,
+ 356
+ ],
+ "category_id": 9,
+ "area": 325650,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14544,
+ "image_id": 2681,
+ "bbox": [
+ 257,
+ 342,
+ 75,
+ 129
+ ],
+ "category_id": 6,
+ "area": 33847,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14545,
+ "image_id": 2681,
+ "bbox": [
+ 102,
+ 228,
+ 69,
+ 69
+ ],
+ "category_id": 6,
+ "area": 16781,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14546,
+ "image_id": 2681,
+ "bbox": [
+ 177,
+ 208,
+ 62,
+ 60
+ ],
+ "category_id": 6,
+ "area": 13175,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14547,
+ "image_id": 2681,
+ "bbox": [
+ 133,
+ 37,
+ 113,
+ 113
+ ],
+ "category_id": 6,
+ "area": 44838,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14548,
+ "image_id": 2681,
+ "bbox": [
+ 241,
+ 15,
+ 36,
+ 80
+ ],
+ "category_id": 6,
+ "area": 10283,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14549,
+ "image_id": 2681,
+ "bbox": [
+ 238,
+ 175,
+ 71,
+ 93
+ ],
+ "category_id": 6,
+ "area": 23187,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14550,
+ "image_id": 2681,
+ "bbox": [
+ 34,
+ 266,
+ 100,
+ 155
+ ],
+ "category_id": 6,
+ "area": 54500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14551,
+ "image_id": 2681,
+ "bbox": [
+ 216,
+ 303,
+ 41,
+ 125
+ ],
+ "category_id": 6,
+ "area": 17952,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14552,
+ "image_id": 2681,
+ "bbox": [
+ 349,
+ 357,
+ 50,
+ 59
+ ],
+ "category_id": 6,
+ "area": 10375,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14558,
+ "image_id": 2683,
+ "bbox": [
+ 37,
+ 177,
+ 423,
+ 274
+ ],
+ "category_id": 6,
+ "area": 851368,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14559,
+ "image_id": 2683,
+ "bbox": [
+ 174,
+ 337,
+ 63,
+ 107
+ ],
+ "category_id": 6,
+ "area": 49447,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14560,
+ "image_id": 2683,
+ "bbox": [
+ 348,
+ 324,
+ 43,
+ 77
+ ],
+ "category_id": 6,
+ "area": 24797,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14561,
+ "image_id": 2683,
+ "bbox": [
+ 250,
+ 440,
+ 82,
+ 54
+ ],
+ "category_id": 6,
+ "area": 32766,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14562,
+ "image_id": 2684,
+ "bbox": [
+ 0,
+ 76,
+ 393,
+ 419
+ ],
+ "category_id": 10,
+ "area": 678592,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14564,
+ "image_id": 2685,
+ "bbox": [
+ 151,
+ 235,
+ 169,
+ 108
+ ],
+ "category_id": 8,
+ "area": 64296,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14565,
+ "image_id": 2685,
+ "bbox": [
+ 275,
+ 299,
+ 172,
+ 212
+ ],
+ "category_id": 8,
+ "area": 127710,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14611,
+ "image_id": 2693,
+ "bbox": [
+ 2,
+ 121,
+ 397,
+ 390
+ ],
+ "category_id": 8,
+ "area": 543616,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14612,
+ "image_id": 2693,
+ "bbox": [
+ 214,
+ 29,
+ 205,
+ 317
+ ],
+ "category_id": 8,
+ "area": 229244,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14629,
+ "image_id": 2697,
+ "bbox": [
+ 174,
+ 379,
+ 76,
+ 63
+ ],
+ "category_id": 9,
+ "area": 217005,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14630,
+ "image_id": 2697,
+ "bbox": [
+ 231,
+ 411,
+ 120,
+ 88
+ ],
+ "category_id": 9,
+ "area": 477180,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14632,
+ "image_id": 2698,
+ "bbox": [
+ 160,
+ 276,
+ 349,
+ 231
+ ],
+ "category_id": 9,
+ "area": 284050,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14633,
+ "image_id": 2698,
+ "bbox": [
+ 212,
+ 0,
+ 292,
+ 364
+ ],
+ "category_id": 9,
+ "area": 374784,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14634,
+ "image_id": 2698,
+ "bbox": [
+ 117,
+ 19,
+ 287,
+ 361
+ ],
+ "category_id": 9,
+ "area": 364744,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14636,
+ "image_id": 2699,
+ "bbox": [
+ 127,
+ 205,
+ 66,
+ 105
+ ],
+ "category_id": 8,
+ "area": 10340,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14637,
+ "image_id": 2699,
+ "bbox": [
+ 168,
+ 121,
+ 259,
+ 390
+ ],
+ "category_id": 8,
+ "area": 149210,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14640,
+ "image_id": 2701,
+ "bbox": [
+ 41,
+ 139,
+ 444,
+ 314
+ ],
+ "category_id": 9,
+ "area": 26775,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14643,
+ "image_id": 2702,
+ "bbox": [
+ 95,
+ 7,
+ 413,
+ 453
+ ],
+ "category_id": 10,
+ "area": 108678,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14657,
+ "image_id": 2704,
+ "bbox": [
+ 69,
+ 253,
+ 325,
+ 200
+ ],
+ "category_id": 8,
+ "area": 438615,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14658,
+ "image_id": 2704,
+ "bbox": [
+ 156,
+ 29,
+ 183,
+ 434
+ ],
+ "category_id": 8,
+ "area": 536940,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14659,
+ "image_id": 2704,
+ "bbox": [
+ 175,
+ 21,
+ 256,
+ 295
+ ],
+ "category_id": 8,
+ "area": 512640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14680,
+ "image_id": 2710,
+ "bbox": [
+ 278,
+ 130,
+ 144,
+ 189
+ ],
+ "category_id": 10,
+ "area": 77778,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14681,
+ "image_id": 2710,
+ "bbox": [
+ 202,
+ 16,
+ 300,
+ 196
+ ],
+ "category_id": 9,
+ "area": 168096,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14690,
+ "image_id": 2712,
+ "bbox": [
+ 127,
+ 242,
+ 159,
+ 169
+ ],
+ "category_id": 8,
+ "area": 94563,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14691,
+ "image_id": 2712,
+ "bbox": [
+ 216,
+ 305,
+ 156,
+ 120
+ ],
+ "category_id": 8,
+ "area": 65910,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14713,
+ "image_id": 2714,
+ "bbox": [
+ 35,
+ 199,
+ 367,
+ 246
+ ],
+ "category_id": 8,
+ "area": 271215,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14714,
+ "image_id": 2714,
+ "bbox": [
+ 324,
+ 167,
+ 126,
+ 165
+ ],
+ "category_id": 8,
+ "area": 62496,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14716,
+ "image_id": 2716,
+ "bbox": [
+ 52,
+ 2,
+ 458,
+ 507
+ ],
+ "category_id": 9,
+ "area": 697076,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14717,
+ "image_id": 2716,
+ "bbox": [
+ 25,
+ 1,
+ 376,
+ 196
+ ],
+ "category_id": 6,
+ "area": 221840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14718,
+ "image_id": 2716,
+ "bbox": [
+ 0,
+ 301,
+ 113,
+ 208
+ ],
+ "category_id": 6,
+ "area": 70512,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14719,
+ "image_id": 2717,
+ "bbox": [
+ 227,
+ 116,
+ 107,
+ 219
+ ],
+ "category_id": 10,
+ "area": 23541,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14720,
+ "image_id": 2717,
+ "bbox": [
+ 197,
+ 63,
+ 190,
+ 330
+ ],
+ "category_id": 9,
+ "area": 62776,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14722,
+ "image_id": 2718,
+ "bbox": [
+ 0,
+ 0,
+ 512,
+ 145
+ ],
+ "category_id": 6,
+ "area": 111890,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14729,
+ "image_id": 2721,
+ "bbox": [
+ 203,
+ 232,
+ 55,
+ 91
+ ],
+ "category_id": 10,
+ "area": 135552,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14730,
+ "image_id": 2721,
+ "bbox": [
+ 347,
+ 376,
+ 69,
+ 91
+ ],
+ "category_id": 10,
+ "area": 168192,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14731,
+ "image_id": 2721,
+ "bbox": [
+ 463,
+ 230,
+ 38,
+ 57
+ ],
+ "category_id": 10,
+ "area": 58564,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14732,
+ "image_id": 2721,
+ "bbox": [
+ 119,
+ 393,
+ 23,
+ 41
+ ],
+ "category_id": 10,
+ "area": 26425,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14734,
+ "image_id": 2721,
+ "bbox": [
+ 403,
+ 343,
+ 65,
+ 111
+ ],
+ "category_id": 10,
+ "area": 195104,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14735,
+ "image_id": 2722,
+ "bbox": [
+ 7,
+ 229,
+ 293,
+ 116
+ ],
+ "category_id": 8,
+ "area": 35599,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14736,
+ "image_id": 2722,
+ "bbox": [
+ 352,
+ 255,
+ 87,
+ 135
+ ],
+ "category_id": 8,
+ "area": 12317,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14737,
+ "image_id": 2722,
+ "bbox": [
+ 432,
+ 309,
+ 64,
+ 95
+ ],
+ "category_id": 8,
+ "area": 6480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14738,
+ "image_id": 2722,
+ "bbox": [
+ 168,
+ 100,
+ 96,
+ 148
+ ],
+ "category_id": 6,
+ "area": 14880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14739,
+ "image_id": 2722,
+ "bbox": [
+ 0,
+ 324,
+ 120,
+ 152
+ ],
+ "category_id": 6,
+ "area": 19177,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14740,
+ "image_id": 2722,
+ "bbox": [
+ 122,
+ 360,
+ 56,
+ 97
+ ],
+ "category_id": 6,
+ "area": 5751,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14743,
+ "image_id": 2723,
+ "bbox": [
+ 63,
+ 185,
+ 397,
+ 256
+ ],
+ "category_id": 8,
+ "area": 357480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14747,
+ "image_id": 2725,
+ "bbox": [
+ 0,
+ 0,
+ 199,
+ 510
+ ],
+ "category_id": 6,
+ "area": 356070,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14748,
+ "image_id": 2725,
+ "bbox": [
+ 205,
+ 281,
+ 166,
+ 113
+ ],
+ "category_id": 8,
+ "area": 65985,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14749,
+ "image_id": 2725,
+ "bbox": [
+ 333,
+ 179,
+ 173,
+ 180
+ ],
+ "category_id": 8,
+ "area": 109549,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14750,
+ "image_id": 2726,
+ "bbox": [
+ 123,
+ 136,
+ 354,
+ 375
+ ],
+ "category_id": 9,
+ "area": 292326,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14751,
+ "image_id": 2726,
+ "bbox": [
+ 16,
+ 438,
+ 109,
+ 73
+ ],
+ "category_id": 6,
+ "area": 17738,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14752,
+ "image_id": 2726,
+ "bbox": [
+ 412,
+ 8,
+ 99,
+ 297
+ ],
+ "category_id": 6,
+ "area": 64616,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14753,
+ "image_id": 2726,
+ "bbox": [
+ 1,
+ 147,
+ 180,
+ 96
+ ],
+ "category_id": 6,
+ "area": 38272,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14754,
+ "image_id": 2727,
+ "bbox": [
+ 195,
+ 0,
+ 310,
+ 244
+ ],
+ "category_id": 9,
+ "area": 555237,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14762,
+ "image_id": 2727,
+ "bbox": [
+ 259,
+ 300,
+ 252,
+ 208
+ ],
+ "category_id": 6,
+ "area": 384544,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14763,
+ "image_id": 2727,
+ "bbox": [
+ 63,
+ 345,
+ 242,
+ 164
+ ],
+ "category_id": 6,
+ "area": 292202,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14780,
+ "image_id": 2729,
+ "bbox": [
+ 47,
+ 2,
+ 464,
+ 198
+ ],
+ "category_id": 6,
+ "area": 225147,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14781,
+ "image_id": 2730,
+ "bbox": [
+ 95,
+ 135,
+ 127,
+ 92
+ ],
+ "category_id": 8,
+ "area": 41340,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14782,
+ "image_id": 2730,
+ "bbox": [
+ 155,
+ 185,
+ 143,
+ 127
+ ],
+ "category_id": 8,
+ "area": 64082,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14785,
+ "image_id": 2732,
+ "bbox": [
+ 55,
+ 70,
+ 398,
+ 392
+ ],
+ "category_id": 8,
+ "area": 162932,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14790,
+ "image_id": 2734,
+ "bbox": [
+ 147,
+ 0,
+ 234,
+ 273
+ ],
+ "category_id": 9,
+ "area": 225408,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14791,
+ "image_id": 2734,
+ "bbox": [
+ 154,
+ 88,
+ 248,
+ 230
+ ],
+ "category_id": 9,
+ "area": 201528,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14794,
+ "image_id": 2735,
+ "bbox": [
+ 265,
+ 335,
+ 60,
+ 78
+ ],
+ "category_id": 6,
+ "area": 11520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14795,
+ "image_id": 2736,
+ "bbox": [
+ 0,
+ 30,
+ 27,
+ 75
+ ],
+ "category_id": 6,
+ "area": 39712,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14796,
+ "image_id": 2736,
+ "bbox": [
+ 2,
+ 98,
+ 62,
+ 84
+ ],
+ "category_id": 6,
+ "area": 102414,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14797,
+ "image_id": 2736,
+ "bbox": [
+ 1,
+ 155,
+ 34,
+ 85
+ ],
+ "category_id": 6,
+ "area": 56488,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14798,
+ "image_id": 2736,
+ "bbox": [
+ 3,
+ 402,
+ 30,
+ 73
+ ],
+ "category_id": 6,
+ "area": 42444,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14799,
+ "image_id": 2736,
+ "bbox": [
+ 52,
+ 362,
+ 86,
+ 138
+ ],
+ "category_id": 6,
+ "area": 231165,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14800,
+ "image_id": 2736,
+ "bbox": [
+ 86,
+ 200,
+ 50,
+ 60
+ ],
+ "category_id": 6,
+ "area": 59024,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14801,
+ "image_id": 2736,
+ "bbox": [
+ 38,
+ 13,
+ 200,
+ 139
+ ],
+ "category_id": 6,
+ "area": 541080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14802,
+ "image_id": 2736,
+ "bbox": [
+ 244,
+ 11,
+ 75,
+ 101
+ ],
+ "category_id": 6,
+ "area": 147334,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14803,
+ "image_id": 2736,
+ "bbox": [
+ 208,
+ 94,
+ 81,
+ 149
+ ],
+ "category_id": 6,
+ "area": 235206,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14804,
+ "image_id": 2736,
+ "bbox": [
+ 300,
+ 73,
+ 87,
+ 199
+ ],
+ "category_id": 6,
+ "area": 337480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14805,
+ "image_id": 2736,
+ "bbox": [
+ 402,
+ 95,
+ 53,
+ 67
+ ],
+ "category_id": 6,
+ "area": 70227,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14806,
+ "image_id": 2736,
+ "bbox": [
+ 459,
+ 97,
+ 52,
+ 83
+ ],
+ "category_id": 6,
+ "area": 84300,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14807,
+ "image_id": 2736,
+ "bbox": [
+ 175,
+ 241,
+ 133,
+ 193
+ ],
+ "category_id": 6,
+ "area": 499653,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14808,
+ "image_id": 2736,
+ "bbox": [
+ 315,
+ 269,
+ 114,
+ 79
+ ],
+ "category_id": 6,
+ "area": 175177,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14809,
+ "image_id": 2736,
+ "bbox": [
+ 440,
+ 49,
+ 40,
+ 78
+ ],
+ "category_id": 6,
+ "area": 62101,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14818,
+ "image_id": 2737,
+ "bbox": [
+ 198,
+ 223,
+ 73,
+ 84
+ ],
+ "category_id": 6,
+ "area": 21942,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14823,
+ "image_id": 2738,
+ "bbox": [
+ 117,
+ 27,
+ 206,
+ 238
+ ],
+ "category_id": 6,
+ "area": 43554,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14824,
+ "image_id": 2738,
+ "bbox": [
+ 0,
+ 110,
+ 144,
+ 209
+ ],
+ "category_id": 6,
+ "area": 26887,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14825,
+ "image_id": 2739,
+ "bbox": [
+ 28,
+ 152,
+ 321,
+ 327
+ ],
+ "category_id": 8,
+ "area": 122706,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14832,
+ "image_id": 2740,
+ "bbox": [
+ 431,
+ 4,
+ 80,
+ 115
+ ],
+ "category_id": 6,
+ "area": 71939,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14833,
+ "image_id": 2740,
+ "bbox": [
+ 309,
+ 284,
+ 31,
+ 73
+ ],
+ "category_id": 6,
+ "area": 18088,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14834,
+ "image_id": 2740,
+ "bbox": [
+ 373,
+ 234,
+ 49,
+ 73
+ ],
+ "category_id": 6,
+ "area": 27968,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14836,
+ "image_id": 2740,
+ "bbox": [
+ 341,
+ 142,
+ 28,
+ 56
+ ],
+ "category_id": 6,
+ "area": 12412,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14839,
+ "image_id": 2742,
+ "bbox": [
+ 106,
+ 105,
+ 351,
+ 254
+ ],
+ "category_id": 9,
+ "area": 424557,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14841,
+ "image_id": 2743,
+ "bbox": [
+ 172,
+ 178,
+ 339,
+ 332
+ ],
+ "category_id": 10,
+ "area": 891672,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14842,
+ "image_id": 2744,
+ "bbox": [
+ 88,
+ 52,
+ 205,
+ 204
+ ],
+ "category_id": 9,
+ "area": 120648,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14844,
+ "image_id": 2744,
+ "bbox": [
+ 266,
+ 89,
+ 121,
+ 201
+ ],
+ "category_id": 10,
+ "area": 70460,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14852,
+ "image_id": 2746,
+ "bbox": [
+ 212,
+ 193,
+ 36,
+ 132
+ ],
+ "category_id": 10,
+ "area": 11591,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14853,
+ "image_id": 2746,
+ "bbox": [
+ 220,
+ 209,
+ 268,
+ 275
+ ],
+ "category_id": 9,
+ "area": 175062,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14856,
+ "image_id": 2747,
+ "bbox": [
+ 190,
+ 162,
+ 71,
+ 73
+ ],
+ "category_id": 9,
+ "area": 7776,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14865,
+ "image_id": 2750,
+ "bbox": [
+ 341,
+ 75,
+ 153,
+ 90
+ ],
+ "category_id": 8,
+ "area": 194818,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14867,
+ "image_id": 2751,
+ "bbox": [
+ 97,
+ 234,
+ 227,
+ 245
+ ],
+ "category_id": 8,
+ "area": 195736,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14868,
+ "image_id": 2751,
+ "bbox": [
+ 299,
+ 74,
+ 96,
+ 122
+ ],
+ "category_id": 8,
+ "area": 41211,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14869,
+ "image_id": 2751,
+ "bbox": [
+ 222,
+ 173,
+ 192,
+ 307
+ ],
+ "category_id": 8,
+ "area": 206830,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14936,
+ "image_id": 2754,
+ "bbox": [
+ 37,
+ 217,
+ 94,
+ 237
+ ],
+ "category_id": 8,
+ "area": 59600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14937,
+ "image_id": 2754,
+ "bbox": [
+ 98,
+ 141,
+ 77,
+ 159
+ ],
+ "category_id": 8,
+ "area": 32800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14938,
+ "image_id": 2754,
+ "bbox": [
+ 226,
+ 157,
+ 88,
+ 68
+ ],
+ "category_id": 8,
+ "area": 15996,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14939,
+ "image_id": 2754,
+ "bbox": [
+ 242,
+ 194,
+ 73,
+ 56
+ ],
+ "category_id": 8,
+ "area": 11076,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14940,
+ "image_id": 2754,
+ "bbox": [
+ 242,
+ 390,
+ 139,
+ 99
+ ],
+ "category_id": 8,
+ "area": 36750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14941,
+ "image_id": 2754,
+ "bbox": [
+ 284,
+ 241,
+ 138,
+ 165
+ ],
+ "category_id": 8,
+ "area": 60651,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14942,
+ "image_id": 2754,
+ "bbox": [
+ 383,
+ 253,
+ 110,
+ 132
+ ],
+ "category_id": 8,
+ "area": 38678,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14947,
+ "image_id": 2756,
+ "bbox": [
+ 338,
+ 90,
+ 173,
+ 216
+ ],
+ "category_id": 6,
+ "area": 63215,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14948,
+ "image_id": 2756,
+ "bbox": [
+ 69,
+ 301,
+ 434,
+ 209
+ ],
+ "category_id": 6,
+ "area": 152998,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14949,
+ "image_id": 2756,
+ "bbox": [
+ 116,
+ 173,
+ 394,
+ 301
+ ],
+ "category_id": 9,
+ "area": 200124,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14956,
+ "image_id": 2758,
+ "bbox": [
+ 176,
+ 284,
+ 84,
+ 104
+ ],
+ "category_id": 8,
+ "area": 30806,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14957,
+ "image_id": 2758,
+ "bbox": [
+ 243,
+ 308,
+ 88,
+ 64
+ ],
+ "category_id": 8,
+ "area": 19800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14960,
+ "image_id": 2759,
+ "bbox": [
+ 191,
+ 168,
+ 45,
+ 60
+ ],
+ "category_id": 8,
+ "area": 9605,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14961,
+ "image_id": 2759,
+ "bbox": [
+ 267,
+ 249,
+ 51,
+ 107
+ ],
+ "category_id": 8,
+ "area": 19050,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14964,
+ "image_id": 2760,
+ "bbox": [
+ 170,
+ 0,
+ 120,
+ 247
+ ],
+ "category_id": 6,
+ "area": 98010,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14965,
+ "image_id": 2760,
+ "bbox": [
+ 129,
+ 1,
+ 51,
+ 58
+ ],
+ "category_id": 6,
+ "area": 9906,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14966,
+ "image_id": 2760,
+ "bbox": [
+ 241,
+ 96,
+ 207,
+ 284
+ ],
+ "category_id": 6,
+ "area": 194427,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14967,
+ "image_id": 2760,
+ "bbox": [
+ 452,
+ 0,
+ 59,
+ 91
+ ],
+ "category_id": 6,
+ "area": 17812,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 14972,
+ "image_id": 2762,
+ "bbox": [
+ 135,
+ 99,
+ 337,
+ 187
+ ],
+ "category_id": 8,
+ "area": 135993,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15000,
+ "image_id": 2766,
+ "bbox": [
+ 76,
+ 108,
+ 388,
+ 401
+ ],
+ "category_id": 9,
+ "area": 892545,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15006,
+ "image_id": 2768,
+ "bbox": [
+ 34,
+ 21,
+ 446,
+ 474
+ ],
+ "category_id": 8,
+ "area": 199786,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15009,
+ "image_id": 2769,
+ "bbox": [
+ 112,
+ 316,
+ 291,
+ 176
+ ],
+ "category_id": 8,
+ "area": 34210,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15010,
+ "image_id": 2770,
+ "bbox": [
+ 111,
+ 171,
+ 104,
+ 84
+ ],
+ "category_id": 8,
+ "area": 70347,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15011,
+ "image_id": 2770,
+ "bbox": [
+ 159,
+ 224,
+ 247,
+ 204
+ ],
+ "category_id": 8,
+ "area": 400464,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15012,
+ "image_id": 2770,
+ "bbox": [
+ 381,
+ 220,
+ 107,
+ 74
+ ],
+ "category_id": 8,
+ "area": 63516,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15013,
+ "image_id": 2770,
+ "bbox": [
+ 482,
+ 257,
+ 29,
+ 33
+ ],
+ "category_id": 8,
+ "area": 7840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15014,
+ "image_id": 2770,
+ "bbox": [
+ 414,
+ 186,
+ 50,
+ 28
+ ],
+ "category_id": 8,
+ "area": 11468,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15025,
+ "image_id": 2773,
+ "bbox": [
+ 0,
+ 232,
+ 426,
+ 166
+ ],
+ "category_id": 8,
+ "area": 212397,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15026,
+ "image_id": 2773,
+ "bbox": [
+ 226,
+ 96,
+ 285,
+ 226
+ ],
+ "category_id": 8,
+ "area": 193230,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15029,
+ "image_id": 2774,
+ "bbox": [
+ 452,
+ 120,
+ 57,
+ 171
+ ],
+ "category_id": 6,
+ "area": 77976,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15030,
+ "image_id": 2774,
+ "bbox": [
+ 391,
+ 191,
+ 65,
+ 124
+ ],
+ "category_id": 6,
+ "area": 64961,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15031,
+ "image_id": 2774,
+ "bbox": [
+ 429,
+ 289,
+ 69,
+ 154
+ ],
+ "category_id": 6,
+ "area": 85412,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15032,
+ "image_id": 2774,
+ "bbox": [
+ 349,
+ 321,
+ 76,
+ 151
+ ],
+ "category_id": 6,
+ "area": 91234,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15033,
+ "image_id": 2774,
+ "bbox": [
+ 326,
+ 208,
+ 66,
+ 163
+ ],
+ "category_id": 6,
+ "area": 86344,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15034,
+ "image_id": 2774,
+ "bbox": [
+ 315,
+ 340,
+ 53,
+ 111
+ ],
+ "category_id": 6,
+ "area": 47436,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15035,
+ "image_id": 2774,
+ "bbox": [
+ 206,
+ 347,
+ 144,
+ 164
+ ],
+ "category_id": 6,
+ "area": 188268,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15036,
+ "image_id": 2774,
+ "bbox": [
+ 36,
+ 377,
+ 44,
+ 70
+ ],
+ "category_id": 6,
+ "area": 24585,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15037,
+ "image_id": 2774,
+ "bbox": [
+ 134,
+ 144,
+ 82,
+ 110
+ ],
+ "category_id": 6,
+ "area": 72230,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15038,
+ "image_id": 2774,
+ "bbox": [
+ 287,
+ 240,
+ 47,
+ 69
+ ],
+ "category_id": 6,
+ "area": 26313,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15039,
+ "image_id": 2774,
+ "bbox": [
+ 185,
+ 198,
+ 63,
+ 65
+ ],
+ "category_id": 6,
+ "area": 32706,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15040,
+ "image_id": 2775,
+ "bbox": [
+ 111,
+ 208,
+ 264,
+ 221
+ ],
+ "category_id": 9,
+ "area": 41796,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15044,
+ "image_id": 2776,
+ "bbox": [
+ 79,
+ 178,
+ 78,
+ 97
+ ],
+ "category_id": 6,
+ "area": 51861,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15045,
+ "image_id": 2776,
+ "bbox": [
+ 61,
+ 237,
+ 102,
+ 124
+ ],
+ "category_id": 6,
+ "area": 86558,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15046,
+ "image_id": 2776,
+ "bbox": [
+ 158,
+ 394,
+ 322,
+ 116
+ ],
+ "category_id": 6,
+ "area": 252630,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15047,
+ "image_id": 2776,
+ "bbox": [
+ 186,
+ 217,
+ 89,
+ 85
+ ],
+ "category_id": 8,
+ "area": 51925,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15064,
+ "image_id": 2779,
+ "bbox": [
+ 224,
+ 80,
+ 139,
+ 350
+ ],
+ "category_id": 9,
+ "area": 158652,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15065,
+ "image_id": 2779,
+ "bbox": [
+ 178,
+ 178,
+ 116,
+ 204
+ ],
+ "category_id": 10,
+ "area": 77259,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15074,
+ "image_id": 2780,
+ "bbox": [
+ 181,
+ 329,
+ 97,
+ 145
+ ],
+ "category_id": 6,
+ "area": 221130,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15075,
+ "image_id": 2780,
+ "bbox": [
+ 328,
+ 421,
+ 151,
+ 90
+ ],
+ "category_id": 6,
+ "area": 214420,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15076,
+ "image_id": 2780,
+ "bbox": [
+ 313,
+ 224,
+ 65,
+ 103
+ ],
+ "category_id": 6,
+ "area": 104975,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15077,
+ "image_id": 2780,
+ "bbox": [
+ 292,
+ 161,
+ 34,
+ 69
+ ],
+ "category_id": 6,
+ "area": 37107,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15078,
+ "image_id": 2780,
+ "bbox": [
+ 329,
+ 144,
+ 76,
+ 82
+ ],
+ "category_id": 6,
+ "area": 98556,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15079,
+ "image_id": 2780,
+ "bbox": [
+ 343,
+ 345,
+ 88,
+ 119
+ ],
+ "category_id": 6,
+ "area": 166056,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15080,
+ "image_id": 2780,
+ "bbox": [
+ 246,
+ 88,
+ 36,
+ 55
+ ],
+ "category_id": 6,
+ "area": 31320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15081,
+ "image_id": 2780,
+ "bbox": [
+ 413,
+ 44,
+ 51,
+ 38
+ ],
+ "category_id": 6,
+ "area": 30345,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15082,
+ "image_id": 2780,
+ "bbox": [
+ 438,
+ 139,
+ 53,
+ 43
+ ],
+ "category_id": 6,
+ "area": 36716,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15083,
+ "image_id": 2780,
+ "bbox": [
+ 388,
+ 200,
+ 49,
+ 59
+ ],
+ "category_id": 6,
+ "area": 45756,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15084,
+ "image_id": 2780,
+ "bbox": [
+ 448,
+ 2,
+ 54,
+ 48
+ ],
+ "category_id": 6,
+ "area": 41769,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15096,
+ "image_id": 2781,
+ "bbox": [
+ 96,
+ 27,
+ 128,
+ 257
+ ],
+ "category_id": 10,
+ "area": 53016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15097,
+ "image_id": 2781,
+ "bbox": [
+ 91,
+ 122,
+ 418,
+ 389
+ ],
+ "category_id": 9,
+ "area": 261751,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15098,
+ "image_id": 2782,
+ "bbox": [
+ 155,
+ 280,
+ 283,
+ 199
+ ],
+ "category_id": 9,
+ "area": 109307,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15107,
+ "image_id": 2784,
+ "bbox": [
+ 50,
+ 131,
+ 255,
+ 172
+ ],
+ "category_id": 8,
+ "area": 153758,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15108,
+ "image_id": 2784,
+ "bbox": [
+ 31,
+ 175,
+ 309,
+ 277
+ ],
+ "category_id": 8,
+ "area": 300308,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15109,
+ "image_id": 2784,
+ "bbox": [
+ 356,
+ 265,
+ 56,
+ 71
+ ],
+ "category_id": 8,
+ "area": 14100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15120,
+ "image_id": 2786,
+ "bbox": [
+ 13,
+ 47,
+ 324,
+ 464
+ ],
+ "category_id": 10,
+ "area": 1189485,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15123,
+ "image_id": 2788,
+ "bbox": [
+ 132,
+ 102,
+ 264,
+ 150
+ ],
+ "category_id": 9,
+ "area": 223686,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15124,
+ "image_id": 2788,
+ "bbox": [
+ 184,
+ 316,
+ 47,
+ 57
+ ],
+ "category_id": 6,
+ "area": 15400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15125,
+ "image_id": 2788,
+ "bbox": [
+ 0,
+ 322,
+ 94,
+ 189
+ ],
+ "category_id": 6,
+ "area": 100464,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15126,
+ "image_id": 2788,
+ "bbox": [
+ 77,
+ 351,
+ 84,
+ 155
+ ],
+ "category_id": 6,
+ "area": 73800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15127,
+ "image_id": 2788,
+ "bbox": [
+ 196,
+ 350,
+ 87,
+ 148
+ ],
+ "category_id": 6,
+ "area": 73502,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15128,
+ "image_id": 2788,
+ "bbox": [
+ 291,
+ 286,
+ 129,
+ 147
+ ],
+ "category_id": 6,
+ "area": 107540,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15129,
+ "image_id": 2788,
+ "bbox": [
+ 3,
+ 255,
+ 90,
+ 90
+ ],
+ "category_id": 6,
+ "area": 46200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15130,
+ "image_id": 2788,
+ "bbox": [
+ 388,
+ 387,
+ 123,
+ 124
+ ],
+ "category_id": 6,
+ "area": 86040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15131,
+ "image_id": 2788,
+ "bbox": [
+ 348,
+ 477,
+ 29,
+ 31
+ ],
+ "category_id": 6,
+ "area": 5307,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15132,
+ "image_id": 2788,
+ "bbox": [
+ 348,
+ 448,
+ 29,
+ 31
+ ],
+ "category_id": 6,
+ "area": 5307,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15133,
+ "image_id": 2788,
+ "bbox": [
+ 439,
+ 275,
+ 72,
+ 125
+ ],
+ "category_id": 6,
+ "area": 51304,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15134,
+ "image_id": 2788,
+ "bbox": [
+ 238,
+ 255,
+ 75,
+ 98
+ ],
+ "category_id": 6,
+ "area": 41958,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15135,
+ "image_id": 2788,
+ "bbox": [
+ 300,
+ 258,
+ 48,
+ 44
+ ],
+ "category_id": 6,
+ "area": 12298,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15136,
+ "image_id": 2788,
+ "bbox": [
+ 145,
+ 265,
+ 37,
+ 45
+ ],
+ "category_id": 6,
+ "area": 9570,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15137,
+ "image_id": 2788,
+ "bbox": [
+ 183,
+ 259,
+ 68,
+ 67
+ ],
+ "category_id": 6,
+ "area": 25929,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15144,
+ "image_id": 2789,
+ "bbox": [
+ 145,
+ 46,
+ 364,
+ 460
+ ],
+ "category_id": 9,
+ "area": 590328,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15146,
+ "image_id": 2790,
+ "bbox": [
+ 71,
+ 109,
+ 396,
+ 327
+ ],
+ "category_id": 9,
+ "area": 71410,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15160,
+ "image_id": 2791,
+ "bbox": [
+ 186,
+ 399,
+ 252,
+ 112
+ ],
+ "category_id": 6,
+ "area": 144400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15161,
+ "image_id": 2791,
+ "bbox": [
+ 0,
+ 320,
+ 105,
+ 159
+ ],
+ "category_id": 6,
+ "area": 85492,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15166,
+ "image_id": 2793,
+ "bbox": [
+ 3,
+ 193,
+ 166,
+ 205
+ ],
+ "category_id": 9,
+ "area": 48768,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15167,
+ "image_id": 2793,
+ "bbox": [
+ 200,
+ 1,
+ 263,
+ 300
+ ],
+ "category_id": 9,
+ "area": 112681,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15168,
+ "image_id": 2793,
+ "bbox": [
+ 95,
+ 236,
+ 280,
+ 273
+ ],
+ "category_id": 9,
+ "area": 109140,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15188,
+ "image_id": 2795,
+ "bbox": [
+ 193,
+ 37,
+ 224,
+ 474
+ ],
+ "category_id": 8,
+ "area": 203863,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15189,
+ "image_id": 2795,
+ "bbox": [
+ 314,
+ 156,
+ 163,
+ 204
+ ],
+ "category_id": 8,
+ "area": 64170,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15190,
+ "image_id": 2796,
+ "bbox": [
+ 141,
+ 104,
+ 59,
+ 106
+ ],
+ "category_id": 8,
+ "area": 22201,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15191,
+ "image_id": 2796,
+ "bbox": [
+ 236,
+ 198,
+ 85,
+ 198
+ ],
+ "category_id": 8,
+ "area": 59214,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15196,
+ "image_id": 2798,
+ "bbox": [
+ 187,
+ 324,
+ 35,
+ 54
+ ],
+ "category_id": 8,
+ "area": 15312,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15198,
+ "image_id": 2799,
+ "bbox": [
+ 75,
+ 304,
+ 122,
+ 80
+ ],
+ "category_id": 6,
+ "area": 40754,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15199,
+ "image_id": 2799,
+ "bbox": [
+ 101,
+ 1,
+ 182,
+ 116
+ ],
+ "category_id": 6,
+ "area": 87535,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15200,
+ "image_id": 2799,
+ "bbox": [
+ 209,
+ 79,
+ 127,
+ 79
+ ],
+ "category_id": 6,
+ "area": 41561,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15201,
+ "image_id": 2799,
+ "bbox": [
+ 189,
+ 288,
+ 154,
+ 108
+ ],
+ "category_id": 6,
+ "area": 68780,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15202,
+ "image_id": 2799,
+ "bbox": [
+ 30,
+ 182,
+ 92,
+ 98
+ ],
+ "category_id": 6,
+ "area": 37368,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15203,
+ "image_id": 2799,
+ "bbox": [
+ 317,
+ 304,
+ 81,
+ 124
+ ],
+ "category_id": 6,
+ "area": 41638,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15204,
+ "image_id": 2799,
+ "bbox": [
+ 426,
+ 0,
+ 85,
+ 290
+ ],
+ "category_id": 6,
+ "area": 102000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15205,
+ "image_id": 2799,
+ "bbox": [
+ 0,
+ 286,
+ 72,
+ 105
+ ],
+ "category_id": 6,
+ "area": 31620,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15206,
+ "image_id": 2799,
+ "bbox": [
+ 337,
+ 179,
+ 90,
+ 104
+ ],
+ "category_id": 6,
+ "area": 39192,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15207,
+ "image_id": 2799,
+ "bbox": [
+ 66,
+ 98,
+ 119,
+ 78
+ ],
+ "category_id": 6,
+ "area": 38640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15208,
+ "image_id": 2799,
+ "bbox": [
+ 105,
+ 180,
+ 66,
+ 81
+ ],
+ "category_id": 6,
+ "area": 22320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15209,
+ "image_id": 2799,
+ "bbox": [
+ 313,
+ 31,
+ 107,
+ 145
+ ],
+ "category_id": 6,
+ "area": 64512,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15210,
+ "image_id": 2799,
+ "bbox": [
+ 238,
+ 393,
+ 111,
+ 118
+ ],
+ "category_id": 6,
+ "area": 54549,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15211,
+ "image_id": 2799,
+ "bbox": [
+ 357,
+ 454,
+ 122,
+ 56
+ ],
+ "category_id": 6,
+ "area": 28600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15260,
+ "image_id": 2803,
+ "bbox": [
+ 268,
+ 210,
+ 206,
+ 163
+ ],
+ "category_id": 8,
+ "area": 117935,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15262,
+ "image_id": 2803,
+ "bbox": [
+ 226,
+ 215,
+ 87,
+ 154
+ ],
+ "category_id": 8,
+ "area": 47523,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15263,
+ "image_id": 2804,
+ "bbox": [
+ 2,
+ 176,
+ 328,
+ 330
+ ],
+ "category_id": 9,
+ "area": 381765,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15264,
+ "image_id": 2804,
+ "bbox": [
+ 2,
+ 0,
+ 428,
+ 401
+ ],
+ "category_id": 9,
+ "area": 605115,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15265,
+ "image_id": 2804,
+ "bbox": [
+ 2,
+ 20,
+ 390,
+ 487
+ ],
+ "category_id": 9,
+ "area": 668560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15272,
+ "image_id": 2807,
+ "bbox": [
+ 0,
+ 0,
+ 511,
+ 510
+ ],
+ "category_id": 6,
+ "area": 783618,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15290,
+ "image_id": 2810,
+ "bbox": [
+ 2,
+ 0,
+ 506,
+ 506
+ ],
+ "category_id": 9,
+ "area": 902104,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15315,
+ "image_id": 2815,
+ "bbox": [
+ 159,
+ 38,
+ 220,
+ 440
+ ],
+ "category_id": 9,
+ "area": 341688,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15316,
+ "image_id": 2815,
+ "bbox": [
+ 230,
+ 0,
+ 257,
+ 506
+ ],
+ "category_id": 6,
+ "area": 457816,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15317,
+ "image_id": 2815,
+ "bbox": [
+ 3,
+ 121,
+ 124,
+ 124
+ ],
+ "category_id": 6,
+ "area": 54425,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15318,
+ "image_id": 2815,
+ "bbox": [
+ 142,
+ 302,
+ 82,
+ 115
+ ],
+ "category_id": 6,
+ "area": 33415,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15319,
+ "image_id": 2816,
+ "bbox": [
+ 209,
+ 66,
+ 129,
+ 348
+ ],
+ "category_id": 8,
+ "area": 357696,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15320,
+ "image_id": 2816,
+ "bbox": [
+ 362,
+ 0,
+ 149,
+ 193
+ ],
+ "category_id": 6,
+ "area": 229858,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15321,
+ "image_id": 2816,
+ "bbox": [
+ 34,
+ 219,
+ 66,
+ 78
+ ],
+ "category_id": 6,
+ "area": 40920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15322,
+ "image_id": 2817,
+ "bbox": [
+ 12,
+ 149,
+ 388,
+ 224
+ ],
+ "category_id": 9,
+ "area": 73216,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15323,
+ "image_id": 2818,
+ "bbox": [
+ 197,
+ 107,
+ 216,
+ 308
+ ],
+ "category_id": 9,
+ "area": 514524,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15324,
+ "image_id": 2819,
+ "bbox": [
+ 206,
+ 140,
+ 128,
+ 116
+ ],
+ "category_id": 8,
+ "area": 118090,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15325,
+ "image_id": 2819,
+ "bbox": [
+ 178,
+ 61,
+ 109,
+ 201
+ ],
+ "category_id": 8,
+ "area": 173840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15326,
+ "image_id": 2819,
+ "bbox": [
+ 0,
+ 45,
+ 193,
+ 172
+ ],
+ "category_id": 8,
+ "area": 262812,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15332,
+ "image_id": 2821,
+ "bbox": [
+ 141,
+ 169,
+ 254,
+ 297
+ ],
+ "category_id": 10,
+ "area": 308360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15338,
+ "image_id": 2823,
+ "bbox": [
+ 59,
+ 0,
+ 452,
+ 292
+ ],
+ "category_id": 6,
+ "area": 333899,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15339,
+ "image_id": 2824,
+ "bbox": [
+ 0,
+ 0,
+ 410,
+ 510
+ ],
+ "category_id": 6,
+ "area": 2946672,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15340,
+ "image_id": 2824,
+ "bbox": [
+ 176,
+ 200,
+ 173,
+ 143
+ ],
+ "category_id": 8,
+ "area": 348595,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15341,
+ "image_id": 2825,
+ "bbox": [
+ 108,
+ 196,
+ 90,
+ 102
+ ],
+ "category_id": 6,
+ "area": 11583,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15342,
+ "image_id": 2825,
+ "bbox": [
+ 285,
+ 243,
+ 99,
+ 106
+ ],
+ "category_id": 6,
+ "area": 13184,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15344,
+ "image_id": 2825,
+ "bbox": [
+ 268,
+ 34,
+ 243,
+ 204
+ ],
+ "category_id": 9,
+ "area": 62172,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15400,
+ "image_id": 2828,
+ "bbox": [
+ 46,
+ 90,
+ 241,
+ 158
+ ],
+ "category_id": 8,
+ "area": 106920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15402,
+ "image_id": 2828,
+ "bbox": [
+ 0,
+ 61,
+ 124,
+ 346
+ ],
+ "category_id": 6,
+ "area": 121088,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15403,
+ "image_id": 2828,
+ "bbox": [
+ 129,
+ 196,
+ 67,
+ 51
+ ],
+ "category_id": 6,
+ "area": 9660,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15404,
+ "image_id": 2828,
+ "bbox": [
+ 223,
+ 84,
+ 287,
+ 427
+ ],
+ "category_id": 6,
+ "area": 343970,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15411,
+ "image_id": 2830,
+ "bbox": [
+ 115,
+ 110,
+ 324,
+ 341
+ ],
+ "category_id": 9,
+ "area": 497475,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15412,
+ "image_id": 2830,
+ "bbox": [
+ 154,
+ 284,
+ 337,
+ 179
+ ],
+ "category_id": 9,
+ "area": 272526,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15418,
+ "image_id": 2833,
+ "bbox": [
+ 202,
+ 209,
+ 47,
+ 140
+ ],
+ "category_id": 10,
+ "area": 15921,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15420,
+ "image_id": 2833,
+ "bbox": [
+ 182,
+ 230,
+ 299,
+ 274
+ ],
+ "category_id": 9,
+ "area": 194922,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15422,
+ "image_id": 2834,
+ "bbox": [
+ 84,
+ 117,
+ 268,
+ 140
+ ],
+ "category_id": 8,
+ "area": 94575,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15430,
+ "image_id": 2837,
+ "bbox": [
+ 192,
+ 132,
+ 319,
+ 360
+ ],
+ "category_id": 9,
+ "area": 172451,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15431,
+ "image_id": 2838,
+ "bbox": [
+ 73,
+ 122,
+ 270,
+ 247
+ ],
+ "category_id": 9,
+ "area": 61181,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15432,
+ "image_id": 2839,
+ "bbox": [
+ 241,
+ 144,
+ 128,
+ 159
+ ],
+ "category_id": 8,
+ "area": 71680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15434,
+ "image_id": 2839,
+ "bbox": [
+ 46,
+ 142,
+ 98,
+ 83
+ ],
+ "category_id": 6,
+ "area": 28910,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15435,
+ "image_id": 2839,
+ "bbox": [
+ 51,
+ 196,
+ 100,
+ 112
+ ],
+ "category_id": 6,
+ "area": 39500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15436,
+ "image_id": 2839,
+ "bbox": [
+ 66,
+ 312,
+ 49,
+ 64
+ ],
+ "category_id": 6,
+ "area": 11160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15437,
+ "image_id": 2839,
+ "bbox": [
+ 181,
+ 285,
+ 44,
+ 54
+ ],
+ "category_id": 6,
+ "area": 8470,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15438,
+ "image_id": 2839,
+ "bbox": [
+ 451,
+ 293,
+ 60,
+ 98
+ ],
+ "category_id": 6,
+ "area": 20976,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15439,
+ "image_id": 2839,
+ "bbox": [
+ 371,
+ 339,
+ 140,
+ 172
+ ],
+ "category_id": 6,
+ "area": 85536,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15440,
+ "image_id": 2839,
+ "bbox": [
+ 268,
+ 375,
+ 144,
+ 136
+ ],
+ "category_id": 6,
+ "area": 69312,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15441,
+ "image_id": 2839,
+ "bbox": [
+ 181,
+ 414,
+ 122,
+ 97
+ ],
+ "category_id": 6,
+ "area": 41922,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15442,
+ "image_id": 2839,
+ "bbox": [
+ 267,
+ 342,
+ 89,
+ 73
+ ],
+ "category_id": 6,
+ "area": 23296,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15448,
+ "image_id": 2840,
+ "bbox": [
+ 51,
+ 62,
+ 71,
+ 51
+ ],
+ "category_id": 8,
+ "area": 11814,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15449,
+ "image_id": 2840,
+ "bbox": [
+ 195,
+ 157,
+ 250,
+ 197
+ ],
+ "category_id": 8,
+ "area": 158610,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15456,
+ "image_id": 2842,
+ "bbox": [
+ 255,
+ 148,
+ 130,
+ 124
+ ],
+ "category_id": 6,
+ "area": 14847,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15457,
+ "image_id": 2842,
+ "bbox": [
+ 350,
+ 202,
+ 85,
+ 133
+ ],
+ "category_id": 6,
+ "area": 10368,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15458,
+ "image_id": 2842,
+ "bbox": [
+ 181,
+ 170,
+ 71,
+ 117
+ ],
+ "category_id": 6,
+ "area": 7695,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15459,
+ "image_id": 2842,
+ "bbox": [
+ 3,
+ 217,
+ 245,
+ 286
+ ],
+ "category_id": 6,
+ "area": 64264,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15460,
+ "image_id": 2843,
+ "bbox": [
+ 259,
+ 344,
+ 120,
+ 166
+ ],
+ "category_id": 8,
+ "area": 281736,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15461,
+ "image_id": 2843,
+ "bbox": [
+ 123,
+ 228,
+ 244,
+ 221
+ ],
+ "category_id": 8,
+ "area": 760683,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15462,
+ "image_id": 2843,
+ "bbox": [
+ 101,
+ 211,
+ 227,
+ 227
+ ],
+ "category_id": 8,
+ "area": 727680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15470,
+ "image_id": 2845,
+ "bbox": [
+ 207,
+ 164,
+ 171,
+ 223
+ ],
+ "category_id": 9,
+ "area": 736044,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15472,
+ "image_id": 2846,
+ "bbox": [
+ 0,
+ 105,
+ 402,
+ 404
+ ],
+ "category_id": 9,
+ "area": 572983,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15473,
+ "image_id": 2846,
+ "bbox": [
+ 197,
+ 2,
+ 312,
+ 197
+ ],
+ "category_id": 9,
+ "area": 217396,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15474,
+ "image_id": 2846,
+ "bbox": [
+ 193,
+ 56,
+ 317,
+ 347
+ ],
+ "category_id": 9,
+ "area": 386984,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15475,
+ "image_id": 2847,
+ "bbox": [
+ 0,
+ 0,
+ 511,
+ 184
+ ],
+ "category_id": 6,
+ "area": 359640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15476,
+ "image_id": 2847,
+ "bbox": [
+ 40,
+ 224,
+ 212,
+ 146
+ ],
+ "category_id": 8,
+ "area": 118690,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15477,
+ "image_id": 2847,
+ "bbox": [
+ 248,
+ 235,
+ 116,
+ 97
+ ],
+ "category_id": 8,
+ "area": 43548,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15478,
+ "image_id": 2847,
+ "bbox": [
+ 314,
+ 311,
+ 128,
+ 119
+ ],
+ "category_id": 8,
+ "area": 58734,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15479,
+ "image_id": 2848,
+ "bbox": [
+ 0,
+ 37,
+ 512,
+ 474
+ ],
+ "category_id": 6,
+ "area": 2161800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15497,
+ "image_id": 2849,
+ "bbox": [
+ 120,
+ 151,
+ 67,
+ 73
+ ],
+ "category_id": 6,
+ "area": 12688,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15498,
+ "image_id": 2849,
+ "bbox": [
+ 1,
+ 338,
+ 62,
+ 82
+ ],
+ "category_id": 6,
+ "area": 12992,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15499,
+ "image_id": 2849,
+ "bbox": [
+ 232,
+ 0,
+ 37,
+ 253
+ ],
+ "category_id": 6,
+ "area": 24344,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15500,
+ "image_id": 2849,
+ "bbox": [
+ 269,
+ 41,
+ 242,
+ 190
+ ],
+ "category_id": 6,
+ "area": 117384,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15501,
+ "image_id": 2849,
+ "bbox": [
+ 98,
+ 224,
+ 178,
+ 144
+ ],
+ "category_id": 6,
+ "area": 65688,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15502,
+ "image_id": 2850,
+ "bbox": [
+ 135,
+ 129,
+ 368,
+ 332
+ ],
+ "category_id": 9,
+ "area": 85344,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15503,
+ "image_id": 2850,
+ "bbox": [
+ 27,
+ 71,
+ 333,
+ 378
+ ],
+ "category_id": 9,
+ "area": 87975,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15513,
+ "image_id": 2853,
+ "bbox": [
+ 47,
+ 7,
+ 461,
+ 499
+ ],
+ "category_id": 9,
+ "area": 811262,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15515,
+ "image_id": 2854,
+ "bbox": [
+ 0,
+ 59,
+ 137,
+ 120
+ ],
+ "category_id": 6,
+ "area": 22869,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15516,
+ "image_id": 2854,
+ "bbox": [
+ 5,
+ 213,
+ 80,
+ 91
+ ],
+ "category_id": 6,
+ "area": 10200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15517,
+ "image_id": 2854,
+ "bbox": [
+ 69,
+ 276,
+ 70,
+ 135
+ ],
+ "category_id": 6,
+ "area": 13209,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15518,
+ "image_id": 2854,
+ "bbox": [
+ 145,
+ 332,
+ 69,
+ 73
+ ],
+ "category_id": 6,
+ "area": 7020,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15519,
+ "image_id": 2854,
+ "bbox": [
+ 126,
+ 446,
+ 81,
+ 65
+ ],
+ "category_id": 6,
+ "area": 7398,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15520,
+ "image_id": 2854,
+ "bbox": [
+ 187,
+ 448,
+ 79,
+ 63
+ ],
+ "category_id": 6,
+ "area": 6916,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15521,
+ "image_id": 2854,
+ "bbox": [
+ 160,
+ 264,
+ 59,
+ 75
+ ],
+ "category_id": 6,
+ "area": 6200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15522,
+ "image_id": 2854,
+ "bbox": [
+ 119,
+ 229,
+ 59,
+ 75
+ ],
+ "category_id": 6,
+ "area": 6200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15523,
+ "image_id": 2854,
+ "bbox": [
+ 73,
+ 198,
+ 50,
+ 78
+ ],
+ "category_id": 6,
+ "area": 5440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15540,
+ "image_id": 2856,
+ "bbox": [
+ 58,
+ 126,
+ 450,
+ 248
+ ],
+ "category_id": 9,
+ "area": 393323,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15548,
+ "image_id": 2858,
+ "bbox": [
+ 0,
+ 101,
+ 298,
+ 409
+ ],
+ "category_id": 6,
+ "area": 390818,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15549,
+ "image_id": 2858,
+ "bbox": [
+ 299,
+ 389,
+ 188,
+ 120
+ ],
+ "category_id": 6,
+ "area": 72540,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15550,
+ "image_id": 2858,
+ "bbox": [
+ 296,
+ 280,
+ 87,
+ 108
+ ],
+ "category_id": 8,
+ "area": 30520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15551,
+ "image_id": 2858,
+ "bbox": [
+ 183,
+ 168,
+ 35,
+ 84
+ ],
+ "category_id": 8,
+ "area": 9612,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15557,
+ "image_id": 2860,
+ "bbox": [
+ 129,
+ 260,
+ 124,
+ 141
+ ],
+ "category_id": 6,
+ "area": 156883,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15558,
+ "image_id": 2860,
+ "bbox": [
+ 220,
+ 407,
+ 90,
+ 104
+ ],
+ "category_id": 6,
+ "area": 84270,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15559,
+ "image_id": 2860,
+ "bbox": [
+ 73,
+ 385,
+ 51,
+ 84
+ ],
+ "category_id": 6,
+ "area": 38948,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15560,
+ "image_id": 2860,
+ "bbox": [
+ 0,
+ 90,
+ 126,
+ 93
+ ],
+ "category_id": 6,
+ "area": 105228,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15561,
+ "image_id": 2860,
+ "bbox": [
+ 132,
+ 98,
+ 57,
+ 104
+ ],
+ "category_id": 6,
+ "area": 53265,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15562,
+ "image_id": 2860,
+ "bbox": [
+ 257,
+ 105,
+ 120,
+ 174
+ ],
+ "category_id": 6,
+ "area": 186984,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15563,
+ "image_id": 2860,
+ "bbox": [
+ 277,
+ 304,
+ 68,
+ 66
+ ],
+ "category_id": 6,
+ "area": 40560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15564,
+ "image_id": 2860,
+ "bbox": [
+ 301,
+ 339,
+ 129,
+ 172
+ ],
+ "category_id": 6,
+ "area": 198380,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15610,
+ "image_id": 2864,
+ "bbox": [
+ 230,
+ 343,
+ 174,
+ 168
+ ],
+ "category_id": 6,
+ "area": 65961,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15619,
+ "image_id": 2866,
+ "bbox": [
+ 112,
+ 105,
+ 362,
+ 253
+ ],
+ "category_id": 9,
+ "area": 322892,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15626,
+ "image_id": 2868,
+ "bbox": [
+ 0,
+ 1,
+ 403,
+ 465
+ ],
+ "category_id": 9,
+ "area": 186238,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15628,
+ "image_id": 2869,
+ "bbox": [
+ 136,
+ 255,
+ 237,
+ 188
+ ],
+ "category_id": 8,
+ "area": 155696,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15629,
+ "image_id": 2869,
+ "bbox": [
+ 76,
+ 284,
+ 211,
+ 150
+ ],
+ "category_id": 8,
+ "area": 110670,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15630,
+ "image_id": 2870,
+ "bbox": [
+ 196,
+ 1,
+ 138,
+ 416
+ ],
+ "category_id": 10,
+ "area": 457959,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15640,
+ "image_id": 2872,
+ "bbox": [
+ 194,
+ 219,
+ 96,
+ 98
+ ],
+ "category_id": 8,
+ "area": 33120,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15641,
+ "image_id": 2872,
+ "bbox": [
+ 264,
+ 240,
+ 88,
+ 69
+ ],
+ "category_id": 8,
+ "area": 21340,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15642,
+ "image_id": 2873,
+ "bbox": [
+ 73,
+ 175,
+ 305,
+ 218
+ ],
+ "category_id": 9,
+ "area": 42316,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15677,
+ "image_id": 2877,
+ "bbox": [
+ 0,
+ 73,
+ 457,
+ 366
+ ],
+ "category_id": 9,
+ "area": 488990,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15681,
+ "image_id": 2880,
+ "bbox": [
+ 24,
+ 0,
+ 116,
+ 247
+ ],
+ "category_id": 8,
+ "area": 222870,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15682,
+ "image_id": 2880,
+ "bbox": [
+ 5,
+ 32,
+ 451,
+ 287
+ ],
+ "category_id": 8,
+ "area": 1000984,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15683,
+ "image_id": 2880,
+ "bbox": [
+ 129,
+ 200,
+ 300,
+ 224
+ ],
+ "category_id": 8,
+ "area": 518826,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15684,
+ "image_id": 2881,
+ "bbox": [
+ 53,
+ 0,
+ 122,
+ 169
+ ],
+ "category_id": 6,
+ "area": 30360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15685,
+ "image_id": 2881,
+ "bbox": [
+ 290,
+ 0,
+ 221,
+ 199
+ ],
+ "category_id": 6,
+ "area": 64152,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15686,
+ "image_id": 2881,
+ "bbox": [
+ 207,
+ 6,
+ 50,
+ 149
+ ],
+ "category_id": 6,
+ "area": 10927,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15687,
+ "image_id": 2881,
+ "bbox": [
+ 270,
+ 370,
+ 146,
+ 131
+ ],
+ "category_id": 6,
+ "area": 27885,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15688,
+ "image_id": 2881,
+ "bbox": [
+ 76,
+ 196,
+ 315,
+ 185
+ ],
+ "category_id": 8,
+ "area": 85008,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15690,
+ "image_id": 2882,
+ "bbox": [
+ 326,
+ 198,
+ 98,
+ 106
+ ],
+ "category_id": 10,
+ "area": 19877,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15691,
+ "image_id": 2882,
+ "bbox": [
+ 307,
+ 127,
+ 175,
+ 101
+ ],
+ "category_id": 9,
+ "area": 33792,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15693,
+ "image_id": 2883,
+ "bbox": [
+ 1,
+ 147,
+ 322,
+ 359
+ ],
+ "category_id": 9,
+ "area": 408342,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15694,
+ "image_id": 2883,
+ "bbox": [
+ 246,
+ 1,
+ 262,
+ 503
+ ],
+ "category_id": 9,
+ "area": 463740,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15695,
+ "image_id": 2883,
+ "bbox": [
+ 129,
+ 2,
+ 249,
+ 414
+ ],
+ "category_id": 9,
+ "area": 363792,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15697,
+ "image_id": 2884,
+ "bbox": [
+ 120,
+ 201,
+ 123,
+ 135
+ ],
+ "category_id": 9,
+ "area": 58710,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15734,
+ "image_id": 2887,
+ "bbox": [
+ 177,
+ 333,
+ 133,
+ 178
+ ],
+ "category_id": 6,
+ "area": 262546,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15735,
+ "image_id": 2887,
+ "bbox": [
+ 0,
+ 0,
+ 140,
+ 510
+ ],
+ "category_id": 6,
+ "area": 792338,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15736,
+ "image_id": 2887,
+ "bbox": [
+ 264,
+ 174,
+ 113,
+ 195
+ ],
+ "category_id": 8,
+ "area": 245195,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15737,
+ "image_id": 2887,
+ "bbox": [
+ 108,
+ 99,
+ 196,
+ 355
+ ],
+ "category_id": 8,
+ "area": 770770,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15738,
+ "image_id": 2888,
+ "bbox": [
+ 0,
+ 256,
+ 287,
+ 255
+ ],
+ "category_id": 8,
+ "area": 579964,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15739,
+ "image_id": 2888,
+ "bbox": [
+ 185,
+ 225,
+ 138,
+ 136
+ ],
+ "category_id": 8,
+ "area": 149472,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15753,
+ "image_id": 2890,
+ "bbox": [
+ 0,
+ 0,
+ 145,
+ 181
+ ],
+ "category_id": 6,
+ "area": 61226,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15754,
+ "image_id": 2890,
+ "bbox": [
+ 266,
+ 191,
+ 199,
+ 250
+ ],
+ "category_id": 9,
+ "area": 115564,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15755,
+ "image_id": 2891,
+ "bbox": [
+ 74,
+ 269,
+ 144,
+ 236
+ ],
+ "category_id": 6,
+ "area": 110175,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15765,
+ "image_id": 2892,
+ "bbox": [
+ 214,
+ 141,
+ 244,
+ 197
+ ],
+ "category_id": 9,
+ "area": 122589,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15772,
+ "image_id": 2894,
+ "bbox": [
+ 0,
+ 259,
+ 88,
+ 251
+ ],
+ "category_id": 8,
+ "area": 153032,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15773,
+ "image_id": 2894,
+ "bbox": [
+ 114,
+ 166,
+ 307,
+ 210
+ ],
+ "category_id": 8,
+ "area": 445124,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15774,
+ "image_id": 2894,
+ "bbox": [
+ 204,
+ 161,
+ 219,
+ 112
+ ],
+ "category_id": 8,
+ "area": 169554,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15778,
+ "image_id": 2896,
+ "bbox": [
+ 242,
+ 152,
+ 52,
+ 94
+ ],
+ "category_id": 10,
+ "area": 11685,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15779,
+ "image_id": 2896,
+ "bbox": [
+ 206,
+ 145,
+ 253,
+ 365
+ ],
+ "category_id": 9,
+ "area": 220388,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15793,
+ "image_id": 2899,
+ "bbox": [
+ 5,
+ 158,
+ 369,
+ 349
+ ],
+ "category_id": 9,
+ "area": 453193,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15796,
+ "image_id": 2900,
+ "bbox": [
+ 151,
+ 175,
+ 90,
+ 116
+ ],
+ "category_id": 8,
+ "area": 37001,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15797,
+ "image_id": 2900,
+ "bbox": [
+ 308,
+ 165,
+ 150,
+ 109
+ ],
+ "category_id": 8,
+ "area": 57750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15798,
+ "image_id": 2900,
+ "bbox": [
+ 428,
+ 267,
+ 30,
+ 34
+ ],
+ "category_id": 6,
+ "area": 3675,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15799,
+ "image_id": 2900,
+ "bbox": [
+ 468,
+ 314,
+ 36,
+ 44
+ ],
+ "category_id": 6,
+ "area": 5642,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15800,
+ "image_id": 2900,
+ "bbox": [
+ 489,
+ 388,
+ 22,
+ 55
+ ],
+ "category_id": 6,
+ "area": 4290,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15828,
+ "image_id": 2904,
+ "bbox": [
+ 105,
+ 216,
+ 382,
+ 230
+ ],
+ "category_id": 8,
+ "area": 420894,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15829,
+ "image_id": 2904,
+ "bbox": [
+ 334,
+ 44,
+ 158,
+ 128
+ ],
+ "category_id": 8,
+ "area": 97000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15830,
+ "image_id": 2905,
+ "bbox": [
+ 159,
+ 128,
+ 125,
+ 145
+ ],
+ "category_id": 6,
+ "area": 21352,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15831,
+ "image_id": 2905,
+ "bbox": [
+ 316,
+ 151,
+ 31,
+ 27
+ ],
+ "category_id": 6,
+ "area": 1014,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15832,
+ "image_id": 2906,
+ "bbox": [
+ 211,
+ 148,
+ 296,
+ 348
+ ],
+ "category_id": 8,
+ "area": 120250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15834,
+ "image_id": 2907,
+ "bbox": [
+ 224,
+ 142,
+ 197,
+ 237
+ ],
+ "category_id": 8,
+ "area": 371480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15835,
+ "image_id": 2907,
+ "bbox": [
+ 325,
+ 107,
+ 157,
+ 137
+ ],
+ "category_id": 8,
+ "area": 171680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15838,
+ "image_id": 2909,
+ "bbox": [
+ 42,
+ 288,
+ 83,
+ 72
+ ],
+ "category_id": 10,
+ "area": 14382,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15840,
+ "image_id": 2909,
+ "bbox": [
+ 237,
+ 130,
+ 274,
+ 163
+ ],
+ "category_id": 9,
+ "area": 106500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15846,
+ "image_id": 2911,
+ "bbox": [
+ 171,
+ 135,
+ 123,
+ 229
+ ],
+ "category_id": 10,
+ "area": 75192,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15852,
+ "image_id": 2913,
+ "bbox": [
+ 0,
+ 258,
+ 144,
+ 77
+ ],
+ "category_id": 8,
+ "area": 24075,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15853,
+ "image_id": 2913,
+ "bbox": [
+ 149,
+ 196,
+ 147,
+ 152
+ ],
+ "category_id": 8,
+ "area": 48300,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15855,
+ "image_id": 2914,
+ "bbox": [
+ 102,
+ 186,
+ 250,
+ 280
+ ],
+ "category_id": 9,
+ "area": 308416,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15856,
+ "image_id": 2915,
+ "bbox": [
+ 0,
+ 0,
+ 511,
+ 511
+ ],
+ "category_id": 10,
+ "area": 2067604,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15859,
+ "image_id": 2916,
+ "bbox": [
+ 30,
+ 405,
+ 101,
+ 99
+ ],
+ "category_id": 6,
+ "area": 28371,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15860,
+ "image_id": 2916,
+ "bbox": [
+ 109,
+ 405,
+ 120,
+ 106
+ ],
+ "category_id": 6,
+ "area": 35796,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15861,
+ "image_id": 2917,
+ "bbox": [
+ 122,
+ 57,
+ 237,
+ 420
+ ],
+ "category_id": 9,
+ "area": 253487,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15862,
+ "image_id": 2917,
+ "bbox": [
+ 240,
+ 155,
+ 198,
+ 302
+ ],
+ "category_id": 10,
+ "area": 152358,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15864,
+ "image_id": 2918,
+ "bbox": [
+ 0,
+ 130,
+ 291,
+ 241
+ ],
+ "category_id": 6,
+ "area": 148470,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15865,
+ "image_id": 2918,
+ "bbox": [
+ 407,
+ 256,
+ 102,
+ 173
+ ],
+ "category_id": 6,
+ "area": 37541,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15866,
+ "image_id": 2918,
+ "bbox": [
+ 474,
+ 425,
+ 37,
+ 37
+ ],
+ "category_id": 6,
+ "area": 2961,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15867,
+ "image_id": 2918,
+ "bbox": [
+ 431,
+ 427,
+ 52,
+ 66
+ ],
+ "category_id": 6,
+ "area": 7476,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15869,
+ "image_id": 2919,
+ "bbox": [
+ 54,
+ 150,
+ 94,
+ 85
+ ],
+ "category_id": 6,
+ "area": 28200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15870,
+ "image_id": 2919,
+ "bbox": [
+ 58,
+ 206,
+ 98,
+ 111
+ ],
+ "category_id": 6,
+ "area": 38622,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15871,
+ "image_id": 2919,
+ "bbox": [
+ 70,
+ 318,
+ 47,
+ 66
+ ],
+ "category_id": 6,
+ "area": 11186,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15872,
+ "image_id": 2919,
+ "bbox": [
+ 185,
+ 292,
+ 44,
+ 51
+ ],
+ "category_id": 6,
+ "area": 8030,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15873,
+ "image_id": 2919,
+ "bbox": [
+ 268,
+ 347,
+ 90,
+ 92
+ ],
+ "category_id": 6,
+ "area": 29380,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15875,
+ "image_id": 2919,
+ "bbox": [
+ 178,
+ 425,
+ 102,
+ 86
+ ],
+ "category_id": 6,
+ "area": 31232,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15876,
+ "image_id": 2919,
+ "bbox": [
+ 260,
+ 388,
+ 146,
+ 123
+ ],
+ "category_id": 6,
+ "area": 63510,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15877,
+ "image_id": 2919,
+ "bbox": [
+ 373,
+ 348,
+ 138,
+ 162
+ ],
+ "category_id": 6,
+ "area": 79234,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15879,
+ "image_id": 2919,
+ "bbox": [
+ 238,
+ 169,
+ 136,
+ 91
+ ],
+ "category_id": 8,
+ "area": 43989,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15881,
+ "image_id": 2919,
+ "bbox": [
+ 447,
+ 313,
+ 64,
+ 84
+ ],
+ "category_id": 6,
+ "area": 19159,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15885,
+ "image_id": 2920,
+ "bbox": [
+ 242,
+ 105,
+ 92,
+ 81
+ ],
+ "category_id": 9,
+ "area": 26680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15886,
+ "image_id": 2920,
+ "bbox": [
+ 333,
+ 101,
+ 68,
+ 77
+ ],
+ "category_id": 9,
+ "area": 18639,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15887,
+ "image_id": 2920,
+ "bbox": [
+ 298,
+ 158,
+ 113,
+ 83
+ ],
+ "category_id": 9,
+ "area": 33228,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15888,
+ "image_id": 2920,
+ "bbox": [
+ 390,
+ 199,
+ 118,
+ 75
+ ],
+ "category_id": 9,
+ "area": 31376,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15940,
+ "image_id": 2924,
+ "bbox": [
+ 51,
+ 145,
+ 95,
+ 89
+ ],
+ "category_id": 6,
+ "area": 30114,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15941,
+ "image_id": 2924,
+ "bbox": [
+ 54,
+ 198,
+ 100,
+ 114
+ ],
+ "category_id": 6,
+ "area": 40250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15942,
+ "image_id": 2924,
+ "bbox": [
+ 66,
+ 312,
+ 50,
+ 63
+ ],
+ "category_id": 6,
+ "area": 11125,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15944,
+ "image_id": 2924,
+ "bbox": [
+ 184,
+ 288,
+ 42,
+ 49
+ ],
+ "category_id": 6,
+ "area": 7383,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15945,
+ "image_id": 2924,
+ "bbox": [
+ 268,
+ 339,
+ 88,
+ 92
+ ],
+ "category_id": 6,
+ "area": 28860,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15947,
+ "image_id": 2924,
+ "bbox": [
+ 181,
+ 418,
+ 118,
+ 93
+ ],
+ "category_id": 6,
+ "area": 38776,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15948,
+ "image_id": 2924,
+ "bbox": [
+ 281,
+ 380,
+ 128,
+ 131
+ ],
+ "category_id": 6,
+ "area": 59570,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15949,
+ "image_id": 2924,
+ "bbox": [
+ 369,
+ 340,
+ 142,
+ 169
+ ],
+ "category_id": 6,
+ "area": 85084,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15950,
+ "image_id": 2924,
+ "bbox": [
+ 241,
+ 167,
+ 129,
+ 133
+ ],
+ "category_id": 8,
+ "area": 60912,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15953,
+ "image_id": 2924,
+ "bbox": [
+ 452,
+ 306,
+ 59,
+ 78
+ ],
+ "category_id": 6,
+ "area": 16280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15955,
+ "image_id": 2925,
+ "bbox": [
+ 208,
+ 229,
+ 72,
+ 159
+ ],
+ "category_id": 8,
+ "area": 40363,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15956,
+ "image_id": 2926,
+ "bbox": [
+ 51,
+ 41,
+ 295,
+ 239
+ ],
+ "category_id": 6,
+ "area": 76755,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15957,
+ "image_id": 2926,
+ "bbox": [
+ 256,
+ 163,
+ 255,
+ 263
+ ],
+ "category_id": 6,
+ "area": 72924,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15967,
+ "image_id": 2927,
+ "bbox": [
+ 241,
+ 34,
+ 268,
+ 473
+ ],
+ "category_id": 9,
+ "area": 446220,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15968,
+ "image_id": 2927,
+ "bbox": [
+ 0,
+ 0,
+ 462,
+ 411
+ ],
+ "category_id": 9,
+ "area": 668745,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15969,
+ "image_id": 2927,
+ "bbox": [
+ 2,
+ 78,
+ 310,
+ 429
+ ],
+ "category_id": 9,
+ "area": 468704,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15971,
+ "image_id": 2928,
+ "bbox": [
+ 87,
+ 175,
+ 151,
+ 146
+ ],
+ "category_id": 9,
+ "area": 45045,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15977,
+ "image_id": 2929,
+ "bbox": [
+ 55,
+ 13,
+ 374,
+ 471
+ ],
+ "category_id": 8,
+ "area": 621231,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15978,
+ "image_id": 2929,
+ "bbox": [
+ 398,
+ 7,
+ 109,
+ 194
+ ],
+ "category_id": 8,
+ "area": 74802,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15979,
+ "image_id": 2930,
+ "bbox": [
+ 161,
+ 97,
+ 130,
+ 364
+ ],
+ "category_id": 8,
+ "area": 374784,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15980,
+ "image_id": 2930,
+ "bbox": [
+ 321,
+ 0,
+ 190,
+ 232
+ ],
+ "category_id": 6,
+ "area": 351556,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15981,
+ "image_id": 2930,
+ "bbox": [
+ 17,
+ 227,
+ 54,
+ 82
+ ],
+ "category_id": 6,
+ "area": 35844,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15982,
+ "image_id": 2931,
+ "bbox": [
+ 82,
+ 227,
+ 125,
+ 211
+ ],
+ "category_id": 10,
+ "area": 34272,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15983,
+ "image_id": 2931,
+ "bbox": [
+ 176,
+ 111,
+ 218,
+ 391
+ ],
+ "category_id": 9,
+ "area": 110716,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15986,
+ "image_id": 2933,
+ "bbox": [
+ 37,
+ 99,
+ 318,
+ 299
+ ],
+ "category_id": 9,
+ "area": 349128,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 15987,
+ "image_id": 2933,
+ "bbox": [
+ 322,
+ 8,
+ 189,
+ 330
+ ],
+ "category_id": 6,
+ "area": 229031,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16000,
+ "image_id": 2935,
+ "bbox": [
+ 0,
+ 121,
+ 321,
+ 162
+ ],
+ "category_id": 9,
+ "area": 87824,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16001,
+ "image_id": 2935,
+ "bbox": [
+ 0,
+ 226,
+ 239,
+ 284
+ ],
+ "category_id": 6,
+ "area": 114948,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16002,
+ "image_id": 2936,
+ "bbox": [
+ 75,
+ 104,
+ 89,
+ 183
+ ],
+ "category_id": 8,
+ "area": 21754,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16006,
+ "image_id": 2938,
+ "bbox": [
+ 79,
+ 132,
+ 221,
+ 173
+ ],
+ "category_id": 9,
+ "area": 54450,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16008,
+ "image_id": 2939,
+ "bbox": [
+ 53,
+ 119,
+ 299,
+ 332
+ ],
+ "category_id": 8,
+ "area": 107590,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16010,
+ "image_id": 2940,
+ "bbox": [
+ 83,
+ 159,
+ 93,
+ 93
+ ],
+ "category_id": 6,
+ "area": 30756,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16011,
+ "image_id": 2940,
+ "bbox": [
+ 86,
+ 217,
+ 72,
+ 109
+ ],
+ "category_id": 6,
+ "area": 27720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16012,
+ "image_id": 2940,
+ "bbox": [
+ 90,
+ 327,
+ 49,
+ 67
+ ],
+ "category_id": 6,
+ "area": 11780,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16013,
+ "image_id": 2940,
+ "bbox": [
+ 209,
+ 298,
+ 40,
+ 54
+ ],
+ "category_id": 6,
+ "area": 7777,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16014,
+ "image_id": 2940,
+ "bbox": [
+ 284,
+ 361,
+ 85,
+ 84
+ ],
+ "category_id": 6,
+ "area": 25347,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16015,
+ "image_id": 2940,
+ "bbox": [
+ 179,
+ 448,
+ 92,
+ 64
+ ],
+ "category_id": 6,
+ "area": 20880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16016,
+ "image_id": 2940,
+ "bbox": [
+ 284,
+ 411,
+ 108,
+ 99
+ ],
+ "category_id": 6,
+ "area": 37800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16017,
+ "image_id": 2940,
+ "bbox": [
+ 367,
+ 384,
+ 144,
+ 125
+ ],
+ "category_id": 6,
+ "area": 63536,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16019,
+ "image_id": 2940,
+ "bbox": [
+ 212,
+ 192,
+ 114,
+ 88
+ ],
+ "category_id": 8,
+ "area": 35750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16020,
+ "image_id": 2940,
+ "bbox": [
+ 370,
+ 241,
+ 38,
+ 62
+ ],
+ "category_id": 6,
+ "area": 8448,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16021,
+ "image_id": 2940,
+ "bbox": [
+ 2,
+ 288,
+ 40,
+ 89
+ ],
+ "category_id": 6,
+ "area": 12852,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16027,
+ "image_id": 2942,
+ "bbox": [
+ 170,
+ 177,
+ 289,
+ 306
+ ],
+ "category_id": 8,
+ "area": 103246,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16029,
+ "image_id": 2943,
+ "bbox": [
+ 99,
+ 39,
+ 312,
+ 394
+ ],
+ "category_id": 9,
+ "area": 306152,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16034,
+ "image_id": 2944,
+ "bbox": [
+ 17,
+ 452,
+ 108,
+ 59
+ ],
+ "category_id": 6,
+ "area": 17922,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16035,
+ "image_id": 2944,
+ "bbox": [
+ 117,
+ 444,
+ 96,
+ 67
+ ],
+ "category_id": 6,
+ "area": 18117,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16036,
+ "image_id": 2945,
+ "bbox": [
+ 80,
+ 329,
+ 11,
+ 31
+ ],
+ "category_id": 6,
+ "area": 1980,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16037,
+ "image_id": 2945,
+ "bbox": [
+ 108,
+ 335,
+ 20,
+ 32
+ ],
+ "category_id": 6,
+ "area": 3528,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16038,
+ "image_id": 2945,
+ "bbox": [
+ 293,
+ 287,
+ 16,
+ 32
+ ],
+ "category_id": 6,
+ "area": 2912,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16039,
+ "image_id": 2945,
+ "bbox": [
+ 3,
+ 0,
+ 35,
+ 83
+ ],
+ "category_id": 6,
+ "area": 15696,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16044,
+ "image_id": 2946,
+ "bbox": [
+ 209,
+ 316,
+ 171,
+ 150
+ ],
+ "category_id": 9,
+ "area": 42700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16053,
+ "image_id": 2949,
+ "bbox": [
+ 4,
+ 52,
+ 413,
+ 408
+ ],
+ "category_id": 9,
+ "area": 10269825,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16059,
+ "image_id": 2950,
+ "bbox": [
+ 168,
+ 354,
+ 109,
+ 148
+ ],
+ "category_id": 9,
+ "area": 24344,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16062,
+ "image_id": 2951,
+ "bbox": [
+ 109,
+ 335,
+ 135,
+ 153
+ ],
+ "category_id": 8,
+ "area": 139380,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16063,
+ "image_id": 2952,
+ "bbox": [
+ 371,
+ 417,
+ 48,
+ 60
+ ],
+ "category_id": 6,
+ "area": 10370,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16064,
+ "image_id": 2952,
+ "bbox": [
+ 297,
+ 185,
+ 90,
+ 114
+ ],
+ "category_id": 8,
+ "area": 36160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16065,
+ "image_id": 2952,
+ "bbox": [
+ 187,
+ 254,
+ 90,
+ 73
+ ],
+ "category_id": 8,
+ "area": 23381,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16068,
+ "image_id": 2954,
+ "bbox": [
+ 120,
+ 174,
+ 291,
+ 209
+ ],
+ "category_id": 8,
+ "area": 1119053,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16069,
+ "image_id": 2954,
+ "bbox": [
+ 22,
+ 124,
+ 135,
+ 65
+ ],
+ "category_id": 6,
+ "area": 162281,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16070,
+ "image_id": 2954,
+ "bbox": [
+ 1,
+ 1,
+ 141,
+ 95
+ ],
+ "category_id": 6,
+ "area": 247441,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16071,
+ "image_id": 2954,
+ "bbox": [
+ 71,
+ 11,
+ 103,
+ 58
+ ],
+ "category_id": 6,
+ "area": 110547,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16072,
+ "image_id": 2954,
+ "bbox": [
+ 9,
+ 191,
+ 148,
+ 96
+ ],
+ "category_id": 6,
+ "area": 262279,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16073,
+ "image_id": 2954,
+ "bbox": [
+ 3,
+ 232,
+ 69,
+ 97
+ ],
+ "category_id": 6,
+ "area": 123895,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16074,
+ "image_id": 2954,
+ "bbox": [
+ 0,
+ 310,
+ 181,
+ 191
+ ],
+ "category_id": 6,
+ "area": 634692,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16075,
+ "image_id": 2954,
+ "bbox": [
+ 141,
+ 366,
+ 154,
+ 144
+ ],
+ "category_id": 6,
+ "area": 407859,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16076,
+ "image_id": 2954,
+ "bbox": [
+ 322,
+ 121,
+ 186,
+ 385
+ ],
+ "category_id": 6,
+ "area": 1314597,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16099,
+ "image_id": 2959,
+ "bbox": [
+ 55,
+ 64,
+ 292,
+ 419
+ ],
+ "category_id": 9,
+ "area": 198470,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16100,
+ "image_id": 2959,
+ "bbox": [
+ 188,
+ 69,
+ 292,
+ 374
+ ],
+ "category_id": 10,
+ "area": 176665,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16117,
+ "image_id": 2962,
+ "bbox": [
+ 129,
+ 0,
+ 173,
+ 345
+ ],
+ "category_id": 10,
+ "area": 473121,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16118,
+ "image_id": 2963,
+ "bbox": [
+ 39,
+ 6,
+ 343,
+ 459
+ ],
+ "category_id": 8,
+ "area": 165910,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16121,
+ "image_id": 2965,
+ "bbox": [
+ 61,
+ 244,
+ 228,
+ 106
+ ],
+ "category_id": 9,
+ "area": 41496,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16122,
+ "image_id": 2965,
+ "bbox": [
+ 276,
+ 153,
+ 175,
+ 261
+ ],
+ "category_id": 9,
+ "area": 77841,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16123,
+ "image_id": 2966,
+ "bbox": [
+ 44,
+ 337,
+ 224,
+ 174
+ ],
+ "category_id": 9,
+ "area": 127680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16147,
+ "image_id": 2968,
+ "bbox": [
+ 116,
+ 69,
+ 206,
+ 382
+ ],
+ "category_id": 10,
+ "area": 131019,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16149,
+ "image_id": 2969,
+ "bbox": [
+ 195,
+ 187,
+ 144,
+ 146
+ ],
+ "category_id": 9,
+ "area": 47685,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16178,
+ "image_id": 2970,
+ "bbox": [
+ 76,
+ 189,
+ 160,
+ 103
+ ],
+ "category_id": 8,
+ "area": 58145,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16179,
+ "image_id": 2970,
+ "bbox": [
+ 218,
+ 251,
+ 146,
+ 172
+ ],
+ "category_id": 8,
+ "area": 88447,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16181,
+ "image_id": 2971,
+ "bbox": [
+ 0,
+ 142,
+ 131,
+ 341
+ ],
+ "category_id": 6,
+ "area": 56760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16195,
+ "image_id": 2973,
+ "bbox": [
+ 0,
+ 102,
+ 394,
+ 407
+ ],
+ "category_id": 9,
+ "area": 564405,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16196,
+ "image_id": 2973,
+ "bbox": [
+ 1,
+ 0,
+ 465,
+ 507
+ ],
+ "category_id": 9,
+ "area": 831096,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16197,
+ "image_id": 2973,
+ "bbox": [
+ 2,
+ 0,
+ 396,
+ 168
+ ],
+ "category_id": 9,
+ "area": 235104,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16198,
+ "image_id": 2973,
+ "bbox": [
+ 392,
+ 0,
+ 116,
+ 310
+ ],
+ "category_id": 9,
+ "area": 126876,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16202,
+ "image_id": 2974,
+ "bbox": [
+ 1,
+ 30,
+ 230,
+ 143
+ ],
+ "category_id": 6,
+ "area": 411651,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16203,
+ "image_id": 2974,
+ "bbox": [
+ 237,
+ 21,
+ 110,
+ 88
+ ],
+ "category_id": 6,
+ "area": 122720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16204,
+ "image_id": 2974,
+ "bbox": [
+ 346,
+ 26,
+ 53,
+ 60
+ ],
+ "category_id": 6,
+ "area": 40200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16205,
+ "image_id": 2974,
+ "bbox": [
+ 320,
+ 34,
+ 126,
+ 197
+ ],
+ "category_id": 6,
+ "area": 309815,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16206,
+ "image_id": 2974,
+ "bbox": [
+ 465,
+ 106,
+ 43,
+ 69
+ ],
+ "category_id": 6,
+ "area": 38048,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16207,
+ "image_id": 2974,
+ "bbox": [
+ 445,
+ 55,
+ 35,
+ 86
+ ],
+ "category_id": 6,
+ "area": 38437,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16208,
+ "image_id": 2974,
+ "bbox": [
+ 180,
+ 107,
+ 124,
+ 163
+ ],
+ "category_id": 6,
+ "area": 253038,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16209,
+ "image_id": 2974,
+ "bbox": [
+ 4,
+ 212,
+ 142,
+ 123
+ ],
+ "category_id": 6,
+ "area": 219350,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16210,
+ "image_id": 2974,
+ "bbox": [
+ 140,
+ 264,
+ 188,
+ 200
+ ],
+ "category_id": 6,
+ "area": 470196,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16211,
+ "image_id": 2974,
+ "bbox": [
+ 345,
+ 292,
+ 156,
+ 86
+ ],
+ "category_id": 6,
+ "area": 168480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16212,
+ "image_id": 2974,
+ "bbox": [
+ 0,
+ 384,
+ 86,
+ 121
+ ],
+ "category_id": 6,
+ "area": 131625,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16213,
+ "image_id": 2974,
+ "bbox": [
+ 425,
+ 0,
+ 81,
+ 96
+ ],
+ "category_id": 6,
+ "area": 98547,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16214,
+ "image_id": 2975,
+ "bbox": [
+ 208,
+ 128,
+ 257,
+ 382
+ ],
+ "category_id": 8,
+ "area": 282880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16215,
+ "image_id": 2975,
+ "bbox": [
+ 6,
+ 291,
+ 66,
+ 106
+ ],
+ "category_id": 6,
+ "area": 20418,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16233,
+ "image_id": 2977,
+ "bbox": [
+ 0,
+ 0,
+ 83,
+ 101
+ ],
+ "category_id": 6,
+ "area": 17892,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16234,
+ "image_id": 2977,
+ "bbox": [
+ 265,
+ 218,
+ 246,
+ 190
+ ],
+ "category_id": 9,
+ "area": 98230,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16248,
+ "image_id": 2980,
+ "bbox": [
+ 0,
+ 276,
+ 312,
+ 234
+ ],
+ "category_id": 9,
+ "area": 84847,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16256,
+ "image_id": 2981,
+ "bbox": [
+ 116,
+ 285,
+ 235,
+ 222
+ ],
+ "category_id": 8,
+ "area": 182868,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16257,
+ "image_id": 2981,
+ "bbox": [
+ 233,
+ 354,
+ 199,
+ 157
+ ],
+ "category_id": 8,
+ "area": 109780,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16258,
+ "image_id": 2981,
+ "bbox": [
+ 120,
+ 58,
+ 21,
+ 26
+ ],
+ "category_id": 9,
+ "area": 1961,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16274,
+ "image_id": 2985,
+ "bbox": [
+ 303,
+ 144,
+ 55,
+ 76
+ ],
+ "category_id": 8,
+ "area": 4899,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16275,
+ "image_id": 2986,
+ "bbox": [
+ 75,
+ 2,
+ 436,
+ 406
+ ],
+ "category_id": 9,
+ "area": 264060,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16278,
+ "image_id": 2987,
+ "bbox": [
+ 61,
+ 114,
+ 350,
+ 394
+ ],
+ "category_id": 9,
+ "area": 63294,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16279,
+ "image_id": 2987,
+ "bbox": [
+ 302,
+ 136,
+ 61,
+ 92
+ ],
+ "category_id": 6,
+ "area": 2592,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16280,
+ "image_id": 2987,
+ "bbox": [
+ 273,
+ 334,
+ 152,
+ 174
+ ],
+ "category_id": 6,
+ "area": 12138,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16281,
+ "image_id": 2987,
+ "bbox": [
+ 364,
+ 46,
+ 145,
+ 262
+ ],
+ "category_id": 6,
+ "area": 17556,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16282,
+ "image_id": 2987,
+ "bbox": [
+ 413,
+ 314,
+ 96,
+ 192
+ ],
+ "category_id": 6,
+ "area": 8475,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16297,
+ "image_id": 2989,
+ "bbox": [
+ 212,
+ 146,
+ 103,
+ 76
+ ],
+ "category_id": 6,
+ "area": 20200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16298,
+ "image_id": 2989,
+ "bbox": [
+ 89,
+ 18,
+ 123,
+ 111
+ ],
+ "category_id": 6,
+ "area": 35090,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16299,
+ "image_id": 2989,
+ "bbox": [
+ 37,
+ 1,
+ 62,
+ 36
+ ],
+ "category_id": 6,
+ "area": 5856,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16300,
+ "image_id": 2989,
+ "bbox": [
+ 0,
+ 8,
+ 153,
+ 182
+ ],
+ "category_id": 6,
+ "area": 70863,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16301,
+ "image_id": 2989,
+ "bbox": [
+ 1,
+ 220,
+ 92,
+ 167
+ ],
+ "category_id": 6,
+ "area": 39458,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16302,
+ "image_id": 2989,
+ "bbox": [
+ 0,
+ 390,
+ 177,
+ 121
+ ],
+ "category_id": 6,
+ "area": 54826,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16303,
+ "image_id": 2989,
+ "bbox": [
+ 169,
+ 441,
+ 67,
+ 62
+ ],
+ "category_id": 6,
+ "area": 10692,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16304,
+ "image_id": 2989,
+ "bbox": [
+ 227,
+ 451,
+ 88,
+ 59
+ ],
+ "category_id": 6,
+ "area": 13321,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16305,
+ "image_id": 2989,
+ "bbox": [
+ 392,
+ 309,
+ 50,
+ 73
+ ],
+ "category_id": 6,
+ "area": 9504,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16306,
+ "image_id": 2989,
+ "bbox": [
+ 371,
+ 222,
+ 82,
+ 100
+ ],
+ "category_id": 10,
+ "area": 21091,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16307,
+ "image_id": 2989,
+ "bbox": [
+ 261,
+ 255,
+ 51,
+ 68
+ ],
+ "category_id": 6,
+ "area": 8989,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16308,
+ "image_id": 2990,
+ "bbox": [
+ 215,
+ 188,
+ 135,
+ 171
+ ],
+ "category_id": 9,
+ "area": 21306,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16311,
+ "image_id": 2992,
+ "bbox": [
+ 199,
+ 36,
+ 312,
+ 173
+ ],
+ "category_id": 8,
+ "area": 113274,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16313,
+ "image_id": 2992,
+ "bbox": [
+ 2,
+ 228,
+ 182,
+ 142
+ ],
+ "category_id": 6,
+ "area": 54112,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16314,
+ "image_id": 2992,
+ "bbox": [
+ 1,
+ 333,
+ 509,
+ 178
+ ],
+ "category_id": 6,
+ "area": 189773,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16316,
+ "image_id": 2993,
+ "bbox": [
+ 81,
+ 178,
+ 312,
+ 279
+ ],
+ "category_id": 8,
+ "area": 115018,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16317,
+ "image_id": 2994,
+ "bbox": [
+ 72,
+ 0,
+ 66,
+ 43
+ ],
+ "category_id": 6,
+ "area": 6728,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16318,
+ "image_id": 2994,
+ "bbox": [
+ 243,
+ 183,
+ 223,
+ 238
+ ],
+ "category_id": 9,
+ "area": 123384,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16325,
+ "image_id": 2997,
+ "bbox": [
+ 34,
+ 260,
+ 272,
+ 145
+ ],
+ "category_id": 9,
+ "area": 66676,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16326,
+ "image_id": 2997,
+ "bbox": [
+ 0,
+ 320,
+ 259,
+ 191
+ ],
+ "category_id": 6,
+ "area": 83616,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16327,
+ "image_id": 2998,
+ "bbox": [
+ 139,
+ 4,
+ 278,
+ 302
+ ],
+ "category_id": 6,
+ "area": 90624,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16342,
+ "image_id": 3001,
+ "bbox": [
+ 27,
+ 406,
+ 96,
+ 98
+ ],
+ "category_id": 6,
+ "area": 26535,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16343,
+ "image_id": 3001,
+ "bbox": [
+ 96,
+ 404,
+ 135,
+ 106
+ ],
+ "category_id": 6,
+ "area": 39936,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16344,
+ "image_id": 3002,
+ "bbox": [
+ 135,
+ 221,
+ 129,
+ 219
+ ],
+ "category_id": 10,
+ "area": 92002,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16345,
+ "image_id": 3002,
+ "bbox": [
+ 88,
+ 50,
+ 205,
+ 264
+ ],
+ "category_id": 9,
+ "area": 176646,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16383,
+ "image_id": 3007,
+ "bbox": [
+ 60,
+ 0,
+ 448,
+ 507
+ ],
+ "category_id": 9,
+ "area": 799680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16405,
+ "image_id": 3009,
+ "bbox": [
+ 98,
+ 278,
+ 112,
+ 80
+ ],
+ "category_id": 6,
+ "area": 68040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16406,
+ "image_id": 3009,
+ "bbox": [
+ 326,
+ 315,
+ 141,
+ 196
+ ],
+ "category_id": 6,
+ "area": 208165,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16429,
+ "image_id": 3011,
+ "bbox": [
+ 0,
+ 105,
+ 468,
+ 406
+ ],
+ "category_id": 10,
+ "area": 464355,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16431,
+ "image_id": 3012,
+ "bbox": [
+ 40,
+ 83,
+ 331,
+ 426
+ ],
+ "category_id": 8,
+ "area": 497400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16433,
+ "image_id": 3013,
+ "bbox": [
+ 139,
+ 276,
+ 176,
+ 177
+ ],
+ "category_id": 9,
+ "area": 30504,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16439,
+ "image_id": 3015,
+ "bbox": [
+ 105,
+ 69,
+ 221,
+ 260
+ ],
+ "category_id": 9,
+ "area": 51205,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16442,
+ "image_id": 3015,
+ "bbox": [
+ 68,
+ 297,
+ 358,
+ 214
+ ],
+ "category_id": 10,
+ "area": 68112,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16443,
+ "image_id": 3016,
+ "bbox": [
+ 76,
+ 333,
+ 237,
+ 138
+ ],
+ "category_id": 9,
+ "area": 49728,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16447,
+ "image_id": 3017,
+ "bbox": [
+ 140,
+ 298,
+ 100,
+ 159
+ ],
+ "category_id": 8,
+ "area": 40488,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16448,
+ "image_id": 3017,
+ "bbox": [
+ 231,
+ 287,
+ 136,
+ 126
+ ],
+ "category_id": 8,
+ "area": 43776,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16468,
+ "image_id": 3020,
+ "bbox": [
+ 0,
+ 35,
+ 242,
+ 329
+ ],
+ "category_id": 8,
+ "area": 267960,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16469,
+ "image_id": 3020,
+ "bbox": [
+ 321,
+ 0,
+ 190,
+ 294
+ ],
+ "category_id": 8,
+ "area": 187878,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16475,
+ "image_id": 3022,
+ "bbox": [
+ 71,
+ 256,
+ 378,
+ 190
+ ],
+ "category_id": 9,
+ "area": 83583,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16478,
+ "image_id": 3023,
+ "bbox": [
+ 143,
+ 300,
+ 138,
+ 208
+ ],
+ "category_id": 6,
+ "area": 104000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16479,
+ "image_id": 3023,
+ "bbox": [
+ 255,
+ 327,
+ 94,
+ 173
+ ],
+ "category_id": 6,
+ "area": 59274,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16480,
+ "image_id": 3023,
+ "bbox": [
+ 401,
+ 173,
+ 61,
+ 90
+ ],
+ "category_id": 6,
+ "area": 20016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16481,
+ "image_id": 3023,
+ "bbox": [
+ 345,
+ 237,
+ 129,
+ 166
+ ],
+ "category_id": 6,
+ "area": 77265,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16488,
+ "image_id": 3025,
+ "bbox": [
+ 4,
+ 164,
+ 487,
+ 347
+ ],
+ "category_id": 9,
+ "area": 453375,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16497,
+ "image_id": 3027,
+ "bbox": [
+ 2,
+ 19,
+ 348,
+ 487
+ ],
+ "category_id": 9,
+ "area": 597320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16498,
+ "image_id": 3027,
+ "bbox": [
+ 152,
+ 204,
+ 356,
+ 304
+ ],
+ "category_id": 9,
+ "area": 381348,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16554,
+ "image_id": 3032,
+ "bbox": [
+ 177,
+ 70,
+ 120,
+ 123
+ ],
+ "category_id": 8,
+ "area": 51900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16555,
+ "image_id": 3032,
+ "bbox": [
+ 240,
+ 208,
+ 134,
+ 181
+ ],
+ "category_id": 8,
+ "area": 85090,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16556,
+ "image_id": 3032,
+ "bbox": [
+ 379,
+ 397,
+ 51,
+ 57
+ ],
+ "category_id": 8,
+ "area": 10449,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16558,
+ "image_id": 3033,
+ "bbox": [
+ 246,
+ 98,
+ 265,
+ 411
+ ],
+ "category_id": 9,
+ "area": 535534,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16563,
+ "image_id": 3035,
+ "bbox": [
+ 7,
+ 159,
+ 174,
+ 342
+ ],
+ "category_id": 10,
+ "area": 151008,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16564,
+ "image_id": 3035,
+ "bbox": [
+ 92,
+ 0,
+ 325,
+ 357
+ ],
+ "category_id": 9,
+ "area": 293679,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16566,
+ "image_id": 3036,
+ "bbox": [
+ 0,
+ 104,
+ 494,
+ 405
+ ],
+ "category_id": 9,
+ "area": 127512,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16567,
+ "image_id": 3037,
+ "bbox": [
+ 37,
+ 270,
+ 292,
+ 157
+ ],
+ "category_id": 8,
+ "area": 2608966,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16568,
+ "image_id": 3037,
+ "bbox": [
+ 238,
+ 53,
+ 208,
+ 159
+ ],
+ "category_id": 8,
+ "area": 1883560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16569,
+ "image_id": 3037,
+ "bbox": [
+ 36,
+ 54,
+ 209,
+ 257
+ ],
+ "category_id": 6,
+ "area": 3060288,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16571,
+ "image_id": 3037,
+ "bbox": [
+ 191,
+ 190,
+ 130,
+ 110
+ ],
+ "category_id": 6,
+ "area": 819924,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16572,
+ "image_id": 3037,
+ "bbox": [
+ 316,
+ 172,
+ 119,
+ 267
+ ],
+ "category_id": 6,
+ "area": 1815515,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16573,
+ "image_id": 3037,
+ "bbox": [
+ 384,
+ 195,
+ 127,
+ 276
+ ],
+ "category_id": 6,
+ "area": 2001025,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16576,
+ "image_id": 3040,
+ "bbox": [
+ 161,
+ 71,
+ 244,
+ 376
+ ],
+ "category_id": 9,
+ "area": 323830,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16580,
+ "image_id": 3041,
+ "bbox": [
+ 275,
+ 108,
+ 199,
+ 171
+ ],
+ "category_id": 8,
+ "area": 165378,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16581,
+ "image_id": 3041,
+ "bbox": [
+ 2,
+ 267,
+ 209,
+ 212
+ ],
+ "category_id": 8,
+ "area": 216675,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16582,
+ "image_id": 3042,
+ "bbox": [
+ 238,
+ 122,
+ 81,
+ 198
+ ],
+ "category_id": 8,
+ "area": 127072,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16583,
+ "image_id": 3042,
+ "bbox": [
+ 282,
+ 242,
+ 86,
+ 202
+ ],
+ "category_id": 8,
+ "area": 138244,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16584,
+ "image_id": 3042,
+ "bbox": [
+ 229,
+ 120,
+ 35,
+ 72
+ ],
+ "category_id": 8,
+ "area": 20064,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16596,
+ "image_id": 3045,
+ "bbox": [
+ 114,
+ 150,
+ 244,
+ 210
+ ],
+ "category_id": 9,
+ "area": 183933,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16598,
+ "image_id": 3047,
+ "bbox": [
+ 0,
+ 34,
+ 340,
+ 476
+ ],
+ "category_id": 8,
+ "area": 441558,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16599,
+ "image_id": 3047,
+ "bbox": [
+ 318,
+ 205,
+ 192,
+ 202
+ ],
+ "category_id": 8,
+ "area": 105842,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16614,
+ "image_id": 3049,
+ "bbox": [
+ 233,
+ 208,
+ 72,
+ 96
+ ],
+ "category_id": 8,
+ "area": 55419,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16615,
+ "image_id": 3050,
+ "bbox": [
+ 197,
+ 171,
+ 96,
+ 233
+ ],
+ "category_id": 8,
+ "area": 152640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16616,
+ "image_id": 3051,
+ "bbox": [
+ 155,
+ 172,
+ 88,
+ 114
+ ],
+ "category_id": 8,
+ "area": 35360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16617,
+ "image_id": 3051,
+ "bbox": [
+ 339,
+ 150,
+ 114,
+ 118
+ ],
+ "category_id": 8,
+ "area": 47642,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16618,
+ "image_id": 3052,
+ "bbox": [
+ 44,
+ 152,
+ 144,
+ 224
+ ],
+ "category_id": 9,
+ "area": 113715,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16619,
+ "image_id": 3052,
+ "bbox": [
+ 158,
+ 287,
+ 150,
+ 219
+ ],
+ "category_id": 9,
+ "area": 116184,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16620,
+ "image_id": 3052,
+ "bbox": [
+ 169,
+ 252,
+ 85,
+ 157
+ ],
+ "category_id": 9,
+ "area": 47073,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16621,
+ "image_id": 3052,
+ "bbox": [
+ 358,
+ 408,
+ 81,
+ 96
+ ],
+ "category_id": 9,
+ "area": 27608,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16622,
+ "image_id": 3052,
+ "bbox": [
+ 266,
+ 147,
+ 115,
+ 214
+ ],
+ "category_id": 9,
+ "area": 86976,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16623,
+ "image_id": 3052,
+ "bbox": [
+ 228,
+ 93,
+ 128,
+ 225
+ ],
+ "category_id": 9,
+ "area": 101440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16632,
+ "image_id": 3054,
+ "bbox": [
+ 0,
+ 112,
+ 461,
+ 397
+ ],
+ "category_id": 9,
+ "area": 267531,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16633,
+ "image_id": 3054,
+ "bbox": [
+ 0,
+ 67,
+ 256,
+ 443
+ ],
+ "category_id": 6,
+ "area": 165645,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16635,
+ "image_id": 3055,
+ "bbox": [
+ 37,
+ 66,
+ 474,
+ 444
+ ],
+ "category_id": 6,
+ "area": 199525,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16657,
+ "image_id": 3059,
+ "bbox": [
+ 74,
+ 237,
+ 137,
+ 113
+ ],
+ "category_id": 9,
+ "area": 40086,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16670,
+ "image_id": 3061,
+ "bbox": [
+ 7,
+ 94,
+ 395,
+ 368
+ ],
+ "category_id": 8,
+ "area": 510796,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16679,
+ "image_id": 3063,
+ "bbox": [
+ 59,
+ 357,
+ 235,
+ 154
+ ],
+ "category_id": 8,
+ "area": 286650,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16680,
+ "image_id": 3063,
+ "bbox": [
+ 280,
+ 365,
+ 59,
+ 97
+ ],
+ "category_id": 8,
+ "area": 45715,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16681,
+ "image_id": 3063,
+ "bbox": [
+ 263,
+ 412,
+ 248,
+ 99
+ ],
+ "category_id": 8,
+ "area": 194788,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16703,
+ "image_id": 3065,
+ "bbox": [
+ 467,
+ 326,
+ 31,
+ 91
+ ],
+ "category_id": 6,
+ "area": 3796,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16704,
+ "image_id": 3065,
+ "bbox": [
+ 309,
+ 183,
+ 86,
+ 230
+ ],
+ "category_id": 8,
+ "area": 26496,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16705,
+ "image_id": 3065,
+ "bbox": [
+ 175,
+ 219,
+ 88,
+ 292
+ ],
+ "category_id": 8,
+ "area": 34251,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16706,
+ "image_id": 3065,
+ "bbox": [
+ 0,
+ 273,
+ 54,
+ 80
+ ],
+ "category_id": 8,
+ "area": 5824,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16707,
+ "image_id": 3065,
+ "bbox": [
+ 122,
+ 185,
+ 39,
+ 79
+ ],
+ "category_id": 8,
+ "area": 4158,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16708,
+ "image_id": 3065,
+ "bbox": [
+ 268,
+ 166,
+ 40,
+ 51
+ ],
+ "category_id": 8,
+ "area": 2747,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16734,
+ "image_id": 3070,
+ "bbox": [
+ 46,
+ 127,
+ 455,
+ 384
+ ],
+ "category_id": 9,
+ "area": 729444,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16735,
+ "image_id": 3071,
+ "bbox": [
+ 58,
+ 45,
+ 411,
+ 448
+ ],
+ "category_id": 6,
+ "area": 380824,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16736,
+ "image_id": 3072,
+ "bbox": [
+ 0,
+ 0,
+ 374,
+ 510
+ ],
+ "category_id": 6,
+ "area": 668525,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16737,
+ "image_id": 3072,
+ "bbox": [
+ 181,
+ 256,
+ 252,
+ 232
+ ],
+ "category_id": 8,
+ "area": 205075,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16738,
+ "image_id": 3072,
+ "bbox": [
+ 205,
+ 9,
+ 288,
+ 296
+ ],
+ "category_id": 8,
+ "area": 299215,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16741,
+ "image_id": 3073,
+ "bbox": [
+ 21,
+ 140,
+ 326,
+ 326
+ ],
+ "category_id": 9,
+ "area": 101761,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16742,
+ "image_id": 3073,
+ "bbox": [
+ 172,
+ 145,
+ 326,
+ 289
+ ],
+ "category_id": 10,
+ "area": 90277,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16755,
+ "image_id": 3077,
+ "bbox": [
+ 2,
+ 2,
+ 65,
+ 119
+ ],
+ "category_id": 6,
+ "area": 72488,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16756,
+ "image_id": 3077,
+ "bbox": [
+ 5,
+ 48,
+ 138,
+ 184
+ ],
+ "category_id": 6,
+ "area": 235330,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16757,
+ "image_id": 3077,
+ "bbox": [
+ 76,
+ 0,
+ 121,
+ 56
+ ],
+ "category_id": 6,
+ "area": 62832,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16758,
+ "image_id": 3077,
+ "bbox": [
+ 171,
+ 5,
+ 135,
+ 196
+ ],
+ "category_id": 6,
+ "area": 246323,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16759,
+ "image_id": 3077,
+ "bbox": [
+ 167,
+ 185,
+ 113,
+ 98
+ ],
+ "category_id": 6,
+ "area": 103027,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16760,
+ "image_id": 3077,
+ "bbox": [
+ 308,
+ 0,
+ 38,
+ 92
+ ],
+ "category_id": 6,
+ "area": 33020,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16761,
+ "image_id": 3077,
+ "bbox": [
+ 326,
+ 52,
+ 77,
+ 81
+ ],
+ "category_id": 6,
+ "area": 58016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16762,
+ "image_id": 3077,
+ "bbox": [
+ 423,
+ 43,
+ 87,
+ 129
+ ],
+ "category_id": 6,
+ "area": 104664,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16763,
+ "image_id": 3077,
+ "bbox": [
+ 462,
+ 294,
+ 49,
+ 110
+ ],
+ "category_id": 6,
+ "area": 49995,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16767,
+ "image_id": 3077,
+ "bbox": [
+ 24,
+ 239,
+ 151,
+ 114
+ ],
+ "category_id": 6,
+ "area": 159630,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16768,
+ "image_id": 3077,
+ "bbox": [
+ 2,
+ 367,
+ 114,
+ 108
+ ],
+ "category_id": 6,
+ "area": 114432,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16769,
+ "image_id": 3077,
+ "bbox": [
+ 196,
+ 277,
+ 176,
+ 103
+ ],
+ "category_id": 6,
+ "area": 168128,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16770,
+ "image_id": 3078,
+ "bbox": [
+ 144,
+ 69,
+ 239,
+ 169
+ ],
+ "category_id": 9,
+ "area": 143161,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16771,
+ "image_id": 3078,
+ "bbox": [
+ 185,
+ 140,
+ 270,
+ 172
+ ],
+ "category_id": 9,
+ "area": 163592,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16772,
+ "image_id": 3079,
+ "bbox": [
+ 92,
+ 96,
+ 418,
+ 361
+ ],
+ "category_id": 9,
+ "area": 817182,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16782,
+ "image_id": 3081,
+ "bbox": [
+ 100,
+ 47,
+ 235,
+ 316
+ ],
+ "category_id": 9,
+ "area": 142967,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16827,
+ "image_id": 3087,
+ "bbox": [
+ 173,
+ 0,
+ 154,
+ 148
+ ],
+ "category_id": 6,
+ "area": 46332,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16828,
+ "image_id": 3087,
+ "bbox": [
+ 0,
+ 185,
+ 366,
+ 324
+ ],
+ "category_id": 9,
+ "area": 240720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16829,
+ "image_id": 3088,
+ "bbox": [
+ 116,
+ 34,
+ 347,
+ 397
+ ],
+ "category_id": 6,
+ "area": 106959,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16833,
+ "image_id": 3089,
+ "bbox": [
+ 148,
+ 84,
+ 101,
+ 100
+ ],
+ "category_id": 6,
+ "area": 14162,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16834,
+ "image_id": 3089,
+ "bbox": [
+ 406,
+ 177,
+ 100,
+ 171
+ ],
+ "category_id": 6,
+ "area": 23925,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16835,
+ "image_id": 3089,
+ "bbox": [
+ 330,
+ 210,
+ 89,
+ 176
+ ],
+ "category_id": 6,
+ "area": 21930,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16843,
+ "image_id": 3091,
+ "bbox": [
+ 80,
+ 139,
+ 365,
+ 372
+ ],
+ "category_id": 9,
+ "area": 113424,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16845,
+ "image_id": 3091,
+ "bbox": [
+ 265,
+ 44,
+ 108,
+ 188
+ ],
+ "category_id": 10,
+ "area": 17061,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16846,
+ "image_id": 3092,
+ "bbox": [
+ 155,
+ 218,
+ 172,
+ 245
+ ],
+ "category_id": 8,
+ "area": 147147,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16847,
+ "image_id": 3092,
+ "bbox": [
+ 217,
+ 45,
+ 294,
+ 466
+ ],
+ "category_id": 8,
+ "area": 477996,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16866,
+ "image_id": 3094,
+ "bbox": [
+ 0,
+ 445,
+ 73,
+ 66
+ ],
+ "category_id": 6,
+ "area": 6426,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16867,
+ "image_id": 3094,
+ "bbox": [
+ 0,
+ 273,
+ 58,
+ 117
+ ],
+ "category_id": 6,
+ "area": 9120,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16868,
+ "image_id": 3094,
+ "bbox": [
+ 232,
+ 356,
+ 132,
+ 155
+ ],
+ "category_id": 6,
+ "area": 27305,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16869,
+ "image_id": 3094,
+ "bbox": [
+ 58,
+ 79,
+ 279,
+ 346
+ ],
+ "category_id": 9,
+ "area": 128199,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16895,
+ "image_id": 3096,
+ "bbox": [
+ 0,
+ 90,
+ 240,
+ 421
+ ],
+ "category_id": 8,
+ "area": 352811,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16898,
+ "image_id": 3098,
+ "bbox": [
+ 285,
+ 203,
+ 112,
+ 173
+ ],
+ "category_id": 10,
+ "area": 22983,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16899,
+ "image_id": 3098,
+ "bbox": [
+ 172,
+ 124,
+ 24,
+ 37
+ ],
+ "category_id": 10,
+ "area": 1050,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16900,
+ "image_id": 3098,
+ "bbox": [
+ 153,
+ 237,
+ 35,
+ 44
+ ],
+ "category_id": 10,
+ "area": 1848,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16901,
+ "image_id": 3098,
+ "bbox": [
+ 398,
+ 72,
+ 34,
+ 51
+ ],
+ "category_id": 10,
+ "area": 2064,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16902,
+ "image_id": 3098,
+ "bbox": [
+ 437,
+ 141,
+ 52,
+ 54
+ ],
+ "category_id": 10,
+ "area": 3366,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16903,
+ "image_id": 3098,
+ "bbox": [
+ 444,
+ 361,
+ 43,
+ 69
+ ],
+ "category_id": 10,
+ "area": 3510,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16904,
+ "image_id": 3098,
+ "bbox": [
+ 174,
+ 466,
+ 26,
+ 42
+ ],
+ "category_id": 10,
+ "area": 1320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16905,
+ "image_id": 3098,
+ "bbox": [
+ 382,
+ 0,
+ 45,
+ 59
+ ],
+ "category_id": 10,
+ "area": 3192,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16906,
+ "image_id": 3098,
+ "bbox": [
+ 324,
+ 345,
+ 76,
+ 83
+ ],
+ "category_id": 10,
+ "area": 7488,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16907,
+ "image_id": 3098,
+ "bbox": [
+ 287,
+ 170,
+ 39,
+ 54
+ ],
+ "category_id": 10,
+ "area": 2499,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16908,
+ "image_id": 3098,
+ "bbox": [
+ 311,
+ 152,
+ 37,
+ 55
+ ],
+ "category_id": 10,
+ "area": 2444,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16909,
+ "image_id": 3098,
+ "bbox": [
+ 359,
+ 452,
+ 33,
+ 44
+ ],
+ "category_id": 10,
+ "area": 1764,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16910,
+ "image_id": 3098,
+ "bbox": [
+ 255,
+ 450,
+ 33,
+ 44
+ ],
+ "category_id": 10,
+ "area": 1764,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16911,
+ "image_id": 3098,
+ "bbox": [
+ 40,
+ 474,
+ 31,
+ 35
+ ],
+ "category_id": 10,
+ "area": 1287,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16912,
+ "image_id": 3098,
+ "bbox": [
+ 446,
+ 189,
+ 33,
+ 32
+ ],
+ "category_id": 10,
+ "area": 1260,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16925,
+ "image_id": 3099,
+ "bbox": [
+ 0,
+ 115,
+ 124,
+ 282
+ ],
+ "category_id": 6,
+ "area": 235416,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16926,
+ "image_id": 3099,
+ "bbox": [
+ 167,
+ 8,
+ 164,
+ 238
+ ],
+ "category_id": 6,
+ "area": 261519,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16928,
+ "image_id": 3100,
+ "bbox": [
+ 32,
+ 284,
+ 162,
+ 158
+ ],
+ "category_id": 9,
+ "area": 99294,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16929,
+ "image_id": 3101,
+ "bbox": [
+ 82,
+ 75,
+ 356,
+ 409
+ ],
+ "category_id": 10,
+ "area": 145136,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16930,
+ "image_id": 3102,
+ "bbox": [
+ 268,
+ 210,
+ 206,
+ 163
+ ],
+ "category_id": 8,
+ "area": 117935,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16932,
+ "image_id": 3102,
+ "bbox": [
+ 226,
+ 215,
+ 87,
+ 154
+ ],
+ "category_id": 8,
+ "area": 47523,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16953,
+ "image_id": 3106,
+ "bbox": [
+ 173,
+ 139,
+ 170,
+ 104
+ ],
+ "category_id": 8,
+ "area": 141219,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16954,
+ "image_id": 3106,
+ "bbox": [
+ 0,
+ 20,
+ 86,
+ 89
+ ],
+ "category_id": 8,
+ "area": 61047,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16955,
+ "image_id": 3106,
+ "bbox": [
+ 27,
+ 38,
+ 92,
+ 59
+ ],
+ "category_id": 8,
+ "area": 43250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16956,
+ "image_id": 3106,
+ "bbox": [
+ 171,
+ 0,
+ 93,
+ 101
+ ],
+ "category_id": 8,
+ "area": 74686,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16961,
+ "image_id": 3108,
+ "bbox": [
+ 0,
+ 31,
+ 174,
+ 251
+ ],
+ "category_id": 6,
+ "area": 58949,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16962,
+ "image_id": 3108,
+ "bbox": [
+ 0,
+ 273,
+ 204,
+ 238
+ ],
+ "category_id": 9,
+ "area": 65637,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16971,
+ "image_id": 3109,
+ "bbox": [
+ 33,
+ 165,
+ 462,
+ 337
+ ],
+ "category_id": 9,
+ "area": 1302375,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16982,
+ "image_id": 3111,
+ "bbox": [
+ 0,
+ 103,
+ 512,
+ 398
+ ],
+ "category_id": 9,
+ "area": 113500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16984,
+ "image_id": 3113,
+ "bbox": [
+ 2,
+ 0,
+ 116,
+ 137
+ ],
+ "category_id": 6,
+ "area": 138294,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16985,
+ "image_id": 3113,
+ "bbox": [
+ 18,
+ 83,
+ 86,
+ 87
+ ],
+ "category_id": 6,
+ "area": 65000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16986,
+ "image_id": 3113,
+ "bbox": [
+ 0,
+ 126,
+ 170,
+ 121
+ ],
+ "category_id": 6,
+ "area": 178190,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16987,
+ "image_id": 3113,
+ "bbox": [
+ 0,
+ 236,
+ 103,
+ 111
+ ],
+ "category_id": 6,
+ "area": 99209,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16988,
+ "image_id": 3113,
+ "bbox": [
+ 114,
+ 206,
+ 72,
+ 66
+ ],
+ "category_id": 6,
+ "area": 41202,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16989,
+ "image_id": 3113,
+ "bbox": [
+ 102,
+ 15,
+ 207,
+ 132
+ ],
+ "category_id": 6,
+ "area": 235872,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16990,
+ "image_id": 3113,
+ "bbox": [
+ 199,
+ 1,
+ 134,
+ 74
+ ],
+ "category_id": 6,
+ "area": 86072,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16991,
+ "image_id": 3113,
+ "bbox": [
+ 289,
+ 76,
+ 113,
+ 87
+ ],
+ "category_id": 6,
+ "area": 85591,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16992,
+ "image_id": 3113,
+ "bbox": [
+ 309,
+ 93,
+ 95,
+ 132
+ ],
+ "category_id": 6,
+ "area": 108486,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16993,
+ "image_id": 3113,
+ "bbox": [
+ 415,
+ 4,
+ 95,
+ 149
+ ],
+ "category_id": 6,
+ "area": 122836,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16994,
+ "image_id": 3113,
+ "bbox": [
+ 442,
+ 209,
+ 69,
+ 72
+ ],
+ "category_id": 6,
+ "area": 42848,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16995,
+ "image_id": 3113,
+ "bbox": [
+ 384,
+ 415,
+ 127,
+ 96
+ ],
+ "category_id": 6,
+ "area": 105984,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 16996,
+ "image_id": 3113,
+ "bbox": [
+ 262,
+ 464,
+ 84,
+ 47
+ ],
+ "category_id": 6,
+ "area": 34408,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17001,
+ "image_id": 3113,
+ "bbox": [
+ 356,
+ 0,
+ 122,
+ 50
+ ],
+ "category_id": 6,
+ "area": 52624,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17002,
+ "image_id": 3114,
+ "bbox": [
+ 156,
+ 118,
+ 352,
+ 387
+ ],
+ "category_id": 9,
+ "area": 480690,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17003,
+ "image_id": 3114,
+ "bbox": [
+ 259,
+ 7,
+ 213,
+ 502
+ ],
+ "category_id": 9,
+ "area": 376831,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17004,
+ "image_id": 3114,
+ "bbox": [
+ 286,
+ 0,
+ 222,
+ 504
+ ],
+ "category_id": 9,
+ "area": 394204,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17010,
+ "image_id": 3116,
+ "bbox": [
+ 85,
+ 296,
+ 423,
+ 215
+ ],
+ "category_id": 6,
+ "area": 124836,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17011,
+ "image_id": 3116,
+ "bbox": [
+ 339,
+ 16,
+ 171,
+ 241
+ ],
+ "category_id": 6,
+ "area": 56613,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17019,
+ "image_id": 3117,
+ "bbox": [
+ 287,
+ 39,
+ 224,
+ 280
+ ],
+ "category_id": 6,
+ "area": 54080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17020,
+ "image_id": 3117,
+ "bbox": [
+ 0,
+ 66,
+ 495,
+ 445
+ ],
+ "category_id": 6,
+ "area": 189108,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17021,
+ "image_id": 3118,
+ "bbox": [
+ 118,
+ 254,
+ 266,
+ 88
+ ],
+ "category_id": 9,
+ "area": 85444,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17026,
+ "image_id": 3120,
+ "bbox": [
+ 27,
+ 63,
+ 250,
+ 295
+ ],
+ "category_id": 10,
+ "area": 79968,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17031,
+ "image_id": 3121,
+ "bbox": [
+ 413,
+ 160,
+ 76,
+ 132
+ ],
+ "category_id": 6,
+ "area": 33998,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17039,
+ "image_id": 3124,
+ "bbox": [
+ 127,
+ 25,
+ 384,
+ 348
+ ],
+ "category_id": 8,
+ "area": 469440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17040,
+ "image_id": 3124,
+ "bbox": [
+ 137,
+ 269,
+ 374,
+ 227
+ ],
+ "category_id": 8,
+ "area": 299520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17076,
+ "image_id": 3127,
+ "bbox": [
+ 257,
+ 107,
+ 86,
+ 303
+ ],
+ "category_id": 10,
+ "area": 208000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17091,
+ "image_id": 3129,
+ "bbox": [
+ 65,
+ 119,
+ 389,
+ 389
+ ],
+ "category_id": 9,
+ "area": 330170,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17109,
+ "image_id": 3131,
+ "bbox": [
+ 116,
+ 53,
+ 378,
+ 292
+ ],
+ "category_id": 9,
+ "area": 317031,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17111,
+ "image_id": 3132,
+ "bbox": [
+ 164,
+ 131,
+ 194,
+ 244
+ ],
+ "category_id": 9,
+ "area": 236676,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17117,
+ "image_id": 3134,
+ "bbox": [
+ 1,
+ 112,
+ 366,
+ 399
+ ],
+ "category_id": 8,
+ "area": 512754,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17118,
+ "image_id": 3134,
+ "bbox": [
+ 170,
+ 36,
+ 251,
+ 370
+ ],
+ "category_id": 8,
+ "area": 327080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17126,
+ "image_id": 3137,
+ "bbox": [
+ 183,
+ 202,
+ 123,
+ 91
+ ],
+ "category_id": 8,
+ "area": 89552,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17127,
+ "image_id": 3137,
+ "bbox": [
+ 85,
+ 206,
+ 150,
+ 127
+ ],
+ "category_id": 8,
+ "area": 151985,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17128,
+ "image_id": 3137,
+ "bbox": [
+ 8,
+ 150,
+ 82,
+ 99
+ ],
+ "category_id": 6,
+ "area": 65310,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17129,
+ "image_id": 3137,
+ "bbox": [
+ 12,
+ 292,
+ 52,
+ 100
+ ],
+ "category_id": 6,
+ "area": 41552,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17130,
+ "image_id": 3137,
+ "bbox": [
+ 51,
+ 256,
+ 50,
+ 91
+ ],
+ "category_id": 6,
+ "area": 36477,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17131,
+ "image_id": 3138,
+ "bbox": [
+ 43,
+ 77,
+ 84,
+ 144
+ ],
+ "category_id": 6,
+ "area": 19035,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17132,
+ "image_id": 3138,
+ "bbox": [
+ 181,
+ 180,
+ 121,
+ 131
+ ],
+ "category_id": 8,
+ "area": 24969,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17133,
+ "image_id": 3139,
+ "bbox": [
+ 0,
+ 235,
+ 28,
+ 49
+ ],
+ "category_id": 6,
+ "area": 9682,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17134,
+ "image_id": 3139,
+ "bbox": [
+ 0,
+ 286,
+ 44,
+ 73
+ ],
+ "category_id": 6,
+ "area": 22630,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17135,
+ "image_id": 3139,
+ "bbox": [
+ 416,
+ 21,
+ 20,
+ 36
+ ],
+ "category_id": 6,
+ "area": 5016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17136,
+ "image_id": 3139,
+ "bbox": [
+ 490,
+ 44,
+ 10,
+ 39
+ ],
+ "category_id": 6,
+ "area": 2739,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17140,
+ "image_id": 3140,
+ "bbox": [
+ 108,
+ 0,
+ 400,
+ 359
+ ],
+ "category_id": 9,
+ "area": 506010,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17141,
+ "image_id": 3140,
+ "bbox": [
+ 2,
+ 85,
+ 398,
+ 426
+ ],
+ "category_id": 9,
+ "area": 597000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17142,
+ "image_id": 3140,
+ "bbox": [
+ 166,
+ 46,
+ 343,
+ 461
+ ],
+ "category_id": 9,
+ "area": 556842,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17148,
+ "image_id": 3142,
+ "bbox": [
+ 147,
+ 92,
+ 193,
+ 418
+ ],
+ "category_id": 10,
+ "area": 639292,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17152,
+ "image_id": 3144,
+ "bbox": [
+ 78,
+ 224,
+ 87,
+ 188
+ ],
+ "category_id": 9,
+ "area": 130216,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17154,
+ "image_id": 3144,
+ "bbox": [
+ 169,
+ 95,
+ 150,
+ 329
+ ],
+ "category_id": 10,
+ "area": 392110,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17155,
+ "image_id": 3145,
+ "bbox": [
+ 38,
+ 103,
+ 364,
+ 329
+ ],
+ "category_id": 9,
+ "area": 170940,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17156,
+ "image_id": 3145,
+ "bbox": [
+ 0,
+ 284,
+ 80,
+ 84
+ ],
+ "category_id": 9,
+ "area": 9717,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17157,
+ "image_id": 3145,
+ "bbox": [
+ 218,
+ 256,
+ 34,
+ 77
+ ],
+ "category_id": 9,
+ "area": 3744,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17158,
+ "image_id": 3146,
+ "bbox": [
+ 404,
+ 242,
+ 39,
+ 59
+ ],
+ "category_id": 6,
+ "area": 18625,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17159,
+ "image_id": 3146,
+ "bbox": [
+ 167,
+ 348,
+ 125,
+ 81
+ ],
+ "category_id": 8,
+ "area": 80840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17160,
+ "image_id": 3146,
+ "bbox": [
+ 175,
+ 273,
+ 86,
+ 45
+ ],
+ "category_id": 8,
+ "area": 31331,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17161,
+ "image_id": 3146,
+ "bbox": [
+ 93,
+ 298,
+ 119,
+ 130
+ ],
+ "category_id": 8,
+ "area": 123200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17162,
+ "image_id": 3146,
+ "bbox": [
+ 99,
+ 347,
+ 82,
+ 70
+ ],
+ "category_id": 8,
+ "area": 46339,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17163,
+ "image_id": 3147,
+ "bbox": [
+ 128,
+ 277,
+ 228,
+ 234
+ ],
+ "category_id": 8,
+ "area": 187288,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17164,
+ "image_id": 3147,
+ "bbox": [
+ 222,
+ 194,
+ 168,
+ 222
+ ],
+ "category_id": 8,
+ "area": 131242,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17165,
+ "image_id": 3147,
+ "bbox": [
+ 292,
+ 117,
+ 66,
+ 116
+ ],
+ "category_id": 8,
+ "area": 26895,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17166,
+ "image_id": 3148,
+ "bbox": [
+ 142,
+ 111,
+ 122,
+ 386
+ ],
+ "category_id": 8,
+ "area": 166701,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17167,
+ "image_id": 3148,
+ "bbox": [
+ 249,
+ 128,
+ 149,
+ 207
+ ],
+ "category_id": 8,
+ "area": 108252,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17168,
+ "image_id": 3149,
+ "bbox": [
+ 132,
+ 179,
+ 181,
+ 138
+ ],
+ "category_id": 9,
+ "area": 29083,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17172,
+ "image_id": 3151,
+ "bbox": [
+ 161,
+ 243,
+ 39,
+ 51
+ ],
+ "category_id": 8,
+ "area": 16132,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17173,
+ "image_id": 3151,
+ "bbox": [
+ 92,
+ 426,
+ 186,
+ 84
+ ],
+ "category_id": 8,
+ "area": 124244,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17174,
+ "image_id": 3151,
+ "bbox": [
+ 95,
+ 285,
+ 344,
+ 226
+ ],
+ "category_id": 8,
+ "area": 615807,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17203,
+ "image_id": 3154,
+ "bbox": [
+ 91,
+ 15,
+ 263,
+ 496
+ ],
+ "category_id": 9,
+ "area": 459284,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17215,
+ "image_id": 3156,
+ "bbox": [
+ 86,
+ 189,
+ 424,
+ 321
+ ],
+ "category_id": 6,
+ "area": 264830,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17217,
+ "image_id": 3157,
+ "bbox": [
+ 72,
+ 117,
+ 305,
+ 391
+ ],
+ "category_id": 9,
+ "area": 66752,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17223,
+ "image_id": 3158,
+ "bbox": [
+ 301,
+ 130,
+ 72,
+ 118
+ ],
+ "category_id": 9,
+ "area": 30046,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17224,
+ "image_id": 3158,
+ "bbox": [
+ 295,
+ 196,
+ 56,
+ 76
+ ],
+ "category_id": 9,
+ "area": 15336,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17228,
+ "image_id": 3160,
+ "bbox": [
+ 163,
+ 119,
+ 336,
+ 391
+ ],
+ "category_id": 8,
+ "area": 385152,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17229,
+ "image_id": 3160,
+ "bbox": [
+ 208,
+ 6,
+ 303,
+ 330
+ ],
+ "category_id": 8,
+ "area": 294109,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17230,
+ "image_id": 3161,
+ "bbox": [
+ 182,
+ 267,
+ 283,
+ 242
+ ],
+ "category_id": 8,
+ "area": 241087,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17231,
+ "image_id": 3161,
+ "bbox": [
+ 144,
+ 101,
+ 338,
+ 254
+ ],
+ "category_id": 8,
+ "area": 302022,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17232,
+ "image_id": 3161,
+ "bbox": [
+ 178,
+ 22,
+ 199,
+ 202
+ ],
+ "category_id": 8,
+ "area": 142215,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17292,
+ "image_id": 3166,
+ "bbox": [
+ 161,
+ 58,
+ 217,
+ 383
+ ],
+ "category_id": 10,
+ "area": 660136,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17294,
+ "image_id": 3167,
+ "bbox": [
+ 232,
+ 255,
+ 103,
+ 79
+ ],
+ "category_id": 8,
+ "area": 64629,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17295,
+ "image_id": 3167,
+ "bbox": [
+ 226,
+ 227,
+ 55,
+ 30
+ ],
+ "category_id": 8,
+ "area": 13455,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17296,
+ "image_id": 3167,
+ "bbox": [
+ 176,
+ 257,
+ 74,
+ 88
+ ],
+ "category_id": 8,
+ "area": 52266,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17297,
+ "image_id": 3167,
+ "bbox": [
+ 269,
+ 191,
+ 39,
+ 56
+ ],
+ "category_id": 6,
+ "area": 17880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17312,
+ "image_id": 3173,
+ "bbox": [
+ 145,
+ 238,
+ 148,
+ 195
+ ],
+ "category_id": 9,
+ "area": 224068,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17322,
+ "image_id": 3176,
+ "bbox": [
+ 246,
+ 144,
+ 60,
+ 90
+ ],
+ "category_id": 8,
+ "area": 42940,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17323,
+ "image_id": 3176,
+ "bbox": [
+ 72,
+ 125,
+ 165,
+ 66
+ ],
+ "category_id": 8,
+ "area": 87279,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17324,
+ "image_id": 3176,
+ "bbox": [
+ 131,
+ 87,
+ 82,
+ 53
+ ],
+ "category_id": 8,
+ "area": 34608,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17325,
+ "image_id": 3176,
+ "bbox": [
+ 92,
+ 98,
+ 57,
+ 32
+ ],
+ "category_id": 8,
+ "area": 14973,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17330,
+ "image_id": 3178,
+ "bbox": [
+ 80,
+ 213,
+ 208,
+ 198
+ ],
+ "category_id": 8,
+ "area": 144560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17331,
+ "image_id": 3178,
+ "bbox": [
+ 220,
+ 217,
+ 226,
+ 199
+ ],
+ "category_id": 8,
+ "area": 158760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17339,
+ "image_id": 3180,
+ "bbox": [
+ 230,
+ 210,
+ 39,
+ 71
+ ],
+ "category_id": 10,
+ "area": 22350,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17340,
+ "image_id": 3180,
+ "bbox": [
+ 162,
+ 4,
+ 46,
+ 63
+ ],
+ "category_id": 10,
+ "area": 23142,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17341,
+ "image_id": 3180,
+ "bbox": [
+ 377,
+ 219,
+ 30,
+ 57
+ ],
+ "category_id": 10,
+ "area": 14152,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17342,
+ "image_id": 3181,
+ "bbox": [
+ 242,
+ 0,
+ 215,
+ 420
+ ],
+ "category_id": 9,
+ "area": 317958,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17343,
+ "image_id": 3181,
+ "bbox": [
+ 279,
+ 0,
+ 162,
+ 503
+ ],
+ "category_id": 9,
+ "area": 286740,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17344,
+ "image_id": 3182,
+ "bbox": [
+ 107,
+ 3,
+ 230,
+ 508
+ ],
+ "category_id": 8,
+ "area": 121764,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17367,
+ "image_id": 3184,
+ "bbox": [
+ 54,
+ 45,
+ 296,
+ 237
+ ],
+ "category_id": 6,
+ "area": 76254,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17368,
+ "image_id": 3184,
+ "bbox": [
+ 257,
+ 169,
+ 254,
+ 262
+ ],
+ "category_id": 6,
+ "area": 72380,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17378,
+ "image_id": 3186,
+ "bbox": [
+ 173,
+ 256,
+ 168,
+ 190
+ ],
+ "category_id": 8,
+ "area": 27156,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17381,
+ "image_id": 3187,
+ "bbox": [
+ 16,
+ 57,
+ 348,
+ 450
+ ],
+ "category_id": 9,
+ "area": 170016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17382,
+ "image_id": 3187,
+ "bbox": [
+ 131,
+ 124,
+ 379,
+ 387
+ ],
+ "category_id": 9,
+ "area": 159378,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17446,
+ "image_id": 3194,
+ "bbox": [
+ 79,
+ 236,
+ 167,
+ 162
+ ],
+ "category_id": 8,
+ "area": 95304,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17447,
+ "image_id": 3194,
+ "bbox": [
+ 203,
+ 229,
+ 208,
+ 195
+ ],
+ "category_id": 8,
+ "area": 142754,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17450,
+ "image_id": 3194,
+ "bbox": [
+ 70,
+ 171,
+ 23,
+ 24
+ ],
+ "category_id": 6,
+ "area": 2006,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17451,
+ "image_id": 3194,
+ "bbox": [
+ 102,
+ 204,
+ 24,
+ 34
+ ],
+ "category_id": 6,
+ "area": 3038,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17452,
+ "image_id": 3194,
+ "bbox": [
+ 112,
+ 246,
+ 30,
+ 37
+ ],
+ "category_id": 6,
+ "area": 4028,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17453,
+ "image_id": 3194,
+ "bbox": [
+ 110,
+ 447,
+ 41,
+ 40
+ ],
+ "category_id": 6,
+ "area": 5871,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17454,
+ "image_id": 3194,
+ "bbox": [
+ 134,
+ 234,
+ 20,
+ 16
+ ],
+ "category_id": 6,
+ "area": 1196,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17460,
+ "image_id": 3196,
+ "bbox": [
+ 302,
+ 0,
+ 209,
+ 511
+ ],
+ "category_id": 6,
+ "area": 267102,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17461,
+ "image_id": 3197,
+ "bbox": [
+ 101,
+ 210,
+ 154,
+ 167
+ ],
+ "category_id": 8,
+ "area": 90090,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17462,
+ "image_id": 3197,
+ "bbox": [
+ 208,
+ 213,
+ 201,
+ 162
+ ],
+ "category_id": 8,
+ "area": 114684,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17463,
+ "image_id": 3197,
+ "bbox": [
+ 94,
+ 126,
+ 23,
+ 24
+ ],
+ "category_id": 6,
+ "area": 2030,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17464,
+ "image_id": 3197,
+ "bbox": [
+ 126,
+ 163,
+ 28,
+ 34
+ ],
+ "category_id": 6,
+ "area": 3360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17465,
+ "image_id": 3197,
+ "bbox": [
+ 136,
+ 209,
+ 29,
+ 27
+ ],
+ "category_id": 6,
+ "area": 2847,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17466,
+ "image_id": 3197,
+ "bbox": [
+ 156,
+ 194,
+ 18,
+ 17
+ ],
+ "category_id": 6,
+ "area": 1150,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17467,
+ "image_id": 3197,
+ "bbox": [
+ 136,
+ 399,
+ 36,
+ 38
+ ],
+ "category_id": 6,
+ "area": 4860,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17469,
+ "image_id": 3198,
+ "bbox": [
+ 128,
+ 0,
+ 207,
+ 108
+ ],
+ "category_id": 9,
+ "area": 64680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17470,
+ "image_id": 3198,
+ "bbox": [
+ 226,
+ 37,
+ 129,
+ 217
+ ],
+ "category_id": 10,
+ "area": 80928,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17472,
+ "image_id": 3199,
+ "bbox": [
+ 140,
+ 85,
+ 26,
+ 27
+ ],
+ "category_id": 6,
+ "area": 2613,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17473,
+ "image_id": 3199,
+ "bbox": [
+ 175,
+ 132,
+ 31,
+ 46
+ ],
+ "category_id": 6,
+ "area": 5070,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17474,
+ "image_id": 3199,
+ "bbox": [
+ 156,
+ 177,
+ 38,
+ 44
+ ],
+ "category_id": 6,
+ "area": 6014,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17475,
+ "image_id": 3199,
+ "bbox": [
+ 210,
+ 169,
+ 20,
+ 24
+ ],
+ "category_id": 6,
+ "area": 1700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17476,
+ "image_id": 3199,
+ "bbox": [
+ 86,
+ 442,
+ 54,
+ 57
+ ],
+ "category_id": 6,
+ "area": 10960,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17477,
+ "image_id": 3199,
+ "bbox": [
+ 12,
+ 213,
+ 269,
+ 241
+ ],
+ "category_id": 8,
+ "area": 227812,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17478,
+ "image_id": 3199,
+ "bbox": [
+ 198,
+ 199,
+ 231,
+ 257
+ ],
+ "category_id": 8,
+ "area": 208080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17480,
+ "image_id": 3200,
+ "bbox": [
+ 63,
+ 93,
+ 233,
+ 378
+ ],
+ "category_id": 8,
+ "area": 102723,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17485,
+ "image_id": 3202,
+ "bbox": [
+ 57,
+ 2,
+ 117,
+ 95
+ ],
+ "category_id": 6,
+ "area": 37592,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17486,
+ "image_id": 3202,
+ "bbox": [
+ 91,
+ 83,
+ 79,
+ 136
+ ],
+ "category_id": 6,
+ "area": 36582,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17487,
+ "image_id": 3202,
+ "bbox": [
+ 7,
+ 333,
+ 96,
+ 173
+ ],
+ "category_id": 6,
+ "area": 56376,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17488,
+ "image_id": 3202,
+ "bbox": [
+ 97,
+ 311,
+ 151,
+ 197
+ ],
+ "category_id": 6,
+ "area": 100466,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17493,
+ "image_id": 3203,
+ "bbox": [
+ 65,
+ 129,
+ 344,
+ 312
+ ],
+ "category_id": 8,
+ "area": 193006,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17494,
+ "image_id": 3203,
+ "bbox": [
+ 173,
+ 73,
+ 263,
+ 191
+ ],
+ "category_id": 8,
+ "area": 90497,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17506,
+ "image_id": 3205,
+ "bbox": [
+ 246,
+ 140,
+ 126,
+ 137
+ ],
+ "category_id": 6,
+ "area": 15873,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17507,
+ "image_id": 3205,
+ "bbox": [
+ 344,
+ 195,
+ 78,
+ 136
+ ],
+ "category_id": 6,
+ "area": 9790,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17508,
+ "image_id": 3205,
+ "bbox": [
+ 169,
+ 169,
+ 76,
+ 123
+ ],
+ "category_id": 6,
+ "area": 8600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17509,
+ "image_id": 3205,
+ "bbox": [
+ 0,
+ 218,
+ 250,
+ 281
+ ],
+ "category_id": 6,
+ "area": 64296,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17516,
+ "image_id": 3209,
+ "bbox": [
+ 0,
+ 121,
+ 50,
+ 114
+ ],
+ "category_id": 8,
+ "area": 20320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17517,
+ "image_id": 3209,
+ "bbox": [
+ 163,
+ 262,
+ 154,
+ 127
+ ],
+ "category_id": 8,
+ "area": 69094,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17518,
+ "image_id": 3209,
+ "bbox": [
+ 272,
+ 311,
+ 110,
+ 97
+ ],
+ "category_id": 8,
+ "area": 37536,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17519,
+ "image_id": 3210,
+ "bbox": [
+ 144,
+ 113,
+ 235,
+ 271
+ ],
+ "category_id": 10,
+ "area": 1952192,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17520,
+ "image_id": 3211,
+ "bbox": [
+ 137,
+ 115,
+ 332,
+ 118
+ ],
+ "category_id": 9,
+ "area": 52156,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17528,
+ "image_id": 3212,
+ "bbox": [
+ 106,
+ 257,
+ 262,
+ 233
+ ],
+ "category_id": 6,
+ "area": 73216,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17529,
+ "image_id": 3212,
+ "bbox": [
+ 22,
+ 322,
+ 88,
+ 63
+ ],
+ "category_id": 6,
+ "area": 6708,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17540,
+ "image_id": 3215,
+ "bbox": [
+ 64,
+ 152,
+ 377,
+ 350
+ ],
+ "category_id": 10,
+ "area": 423436,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17541,
+ "image_id": 3216,
+ "bbox": [
+ 56,
+ 50,
+ 453,
+ 456
+ ],
+ "category_id": 9,
+ "area": 727386,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17557,
+ "image_id": 3219,
+ "bbox": [
+ 86,
+ 157,
+ 92,
+ 93
+ ],
+ "category_id": 6,
+ "area": 30624,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17558,
+ "image_id": 3219,
+ "bbox": [
+ 90,
+ 217,
+ 92,
+ 110
+ ],
+ "category_id": 6,
+ "area": 35880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17559,
+ "image_id": 3219,
+ "bbox": [
+ 95,
+ 322,
+ 47,
+ 67
+ ],
+ "category_id": 6,
+ "area": 11305,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17560,
+ "image_id": 3219,
+ "bbox": [
+ 173,
+ 457,
+ 88,
+ 54
+ ],
+ "category_id": 6,
+ "area": 16796,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17561,
+ "image_id": 3219,
+ "bbox": [
+ 274,
+ 411,
+ 108,
+ 99
+ ],
+ "category_id": 6,
+ "area": 38080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17562,
+ "image_id": 3219,
+ "bbox": [
+ 373,
+ 383,
+ 138,
+ 128
+ ],
+ "category_id": 6,
+ "area": 62460,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17564,
+ "image_id": 3219,
+ "bbox": [
+ 216,
+ 190,
+ 113,
+ 93
+ ],
+ "category_id": 8,
+ "area": 37073,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17565,
+ "image_id": 3219,
+ "bbox": [
+ 286,
+ 357,
+ 85,
+ 78
+ ],
+ "category_id": 6,
+ "area": 23643,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17566,
+ "image_id": 3219,
+ "bbox": [
+ 212,
+ 297,
+ 40,
+ 52
+ ],
+ "category_id": 6,
+ "area": 7548,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17567,
+ "image_id": 3219,
+ "bbox": [
+ 367,
+ 244,
+ 45,
+ 57
+ ],
+ "category_id": 6,
+ "area": 9234,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17569,
+ "image_id": 3220,
+ "bbox": [
+ 100,
+ 115,
+ 360,
+ 310
+ ],
+ "category_id": 8,
+ "area": 392400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17570,
+ "image_id": 3220,
+ "bbox": [
+ 32,
+ 201,
+ 231,
+ 265
+ ],
+ "category_id": 8,
+ "area": 215967,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17577,
+ "image_id": 3222,
+ "bbox": [
+ 0,
+ 5,
+ 162,
+ 385
+ ],
+ "category_id": 6,
+ "area": 59784,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17586,
+ "image_id": 3223,
+ "bbox": [
+ 33,
+ 8,
+ 75,
+ 90
+ ],
+ "category_id": 6,
+ "area": 48816,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17587,
+ "image_id": 3223,
+ "bbox": [
+ 68,
+ 82,
+ 73,
+ 132
+ ],
+ "category_id": 6,
+ "area": 70057,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17588,
+ "image_id": 3223,
+ "bbox": [
+ 156,
+ 114,
+ 85,
+ 110
+ ],
+ "category_id": 6,
+ "area": 67840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17589,
+ "image_id": 3223,
+ "bbox": [
+ 155,
+ 1,
+ 118,
+ 76
+ ],
+ "category_id": 6,
+ "area": 65148,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17590,
+ "image_id": 3223,
+ "bbox": [
+ 283,
+ 249,
+ 173,
+ 197
+ ],
+ "category_id": 6,
+ "area": 246954,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17602,
+ "image_id": 3225,
+ "bbox": [
+ 194,
+ 156,
+ 150,
+ 230
+ ],
+ "category_id": 9,
+ "area": 38986,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17603,
+ "image_id": 3225,
+ "bbox": [
+ 367,
+ 28,
+ 144,
+ 192
+ ],
+ "category_id": 9,
+ "area": 31234,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17604,
+ "image_id": 3226,
+ "bbox": [
+ 0,
+ 318,
+ 66,
+ 77
+ ],
+ "category_id": 6,
+ "area": 35208,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17605,
+ "image_id": 3226,
+ "bbox": [
+ 0,
+ 265,
+ 27,
+ 54
+ ],
+ "category_id": 6,
+ "area": 10146,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17606,
+ "image_id": 3226,
+ "bbox": [
+ 388,
+ 27,
+ 20,
+ 41
+ ],
+ "category_id": 6,
+ "area": 5720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17607,
+ "image_id": 3226,
+ "bbox": [
+ 417,
+ 35,
+ 19,
+ 35
+ ],
+ "category_id": 6,
+ "area": 4650,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17608,
+ "image_id": 3226,
+ "bbox": [
+ 465,
+ 49,
+ 8,
+ 45
+ ],
+ "category_id": 6,
+ "area": 2688,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17611,
+ "image_id": 3227,
+ "bbox": [
+ 190,
+ 166,
+ 46,
+ 65
+ ],
+ "category_id": 8,
+ "area": 10556,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17612,
+ "image_id": 3227,
+ "bbox": [
+ 265,
+ 246,
+ 54,
+ 112
+ ],
+ "category_id": 8,
+ "area": 21195,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17613,
+ "image_id": 3227,
+ "bbox": [
+ 0,
+ 346,
+ 70,
+ 136
+ ],
+ "category_id": 6,
+ "area": 33616,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17614,
+ "image_id": 3227,
+ "bbox": [
+ 62,
+ 291,
+ 39,
+ 64
+ ],
+ "category_id": 6,
+ "area": 8910,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17615,
+ "image_id": 3227,
+ "bbox": [
+ 101,
+ 329,
+ 42,
+ 47
+ ],
+ "category_id": 6,
+ "area": 6930,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17616,
+ "image_id": 3227,
+ "bbox": [
+ 0,
+ 290,
+ 50,
+ 63
+ ],
+ "category_id": 6,
+ "area": 11125,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17635,
+ "image_id": 3230,
+ "bbox": [
+ 173,
+ 382,
+ 67,
+ 99
+ ],
+ "category_id": 6,
+ "area": 50347,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17636,
+ "image_id": 3230,
+ "bbox": [
+ 248,
+ 351,
+ 66,
+ 75
+ ],
+ "category_id": 6,
+ "area": 37297,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17642,
+ "image_id": 3231,
+ "bbox": [
+ 361,
+ 323,
+ 145,
+ 170
+ ],
+ "category_id": 6,
+ "area": 57368,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17644,
+ "image_id": 3233,
+ "bbox": [
+ 219,
+ 145,
+ 268,
+ 313
+ ],
+ "category_id": 9,
+ "area": 507360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17670,
+ "image_id": 3234,
+ "bbox": [
+ 8,
+ 0,
+ 127,
+ 148
+ ],
+ "category_id": 6,
+ "area": 203376,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17671,
+ "image_id": 3234,
+ "bbox": [
+ 17,
+ 211,
+ 81,
+ 65
+ ],
+ "category_id": 6,
+ "area": 57036,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17672,
+ "image_id": 3234,
+ "bbox": [
+ 2,
+ 279,
+ 106,
+ 100
+ ],
+ "category_id": 6,
+ "area": 114079,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17673,
+ "image_id": 3234,
+ "bbox": [
+ 164,
+ 1,
+ 125,
+ 126
+ ],
+ "category_id": 6,
+ "area": 170550,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17674,
+ "image_id": 3234,
+ "bbox": [
+ 160,
+ 112,
+ 105,
+ 86
+ ],
+ "category_id": 6,
+ "area": 97760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17675,
+ "image_id": 3234,
+ "bbox": [
+ 179,
+ 191,
+ 169,
+ 100
+ ],
+ "category_id": 6,
+ "area": 183618,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17676,
+ "image_id": 3234,
+ "bbox": [
+ 308,
+ 1,
+ 85,
+ 80
+ ],
+ "category_id": 6,
+ "area": 74052,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17677,
+ "image_id": 3234,
+ "bbox": [
+ 395,
+ 2,
+ 115,
+ 88
+ ],
+ "category_id": 6,
+ "area": 109858,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17678,
+ "image_id": 3234,
+ "bbox": [
+ 438,
+ 202,
+ 73,
+ 109
+ ],
+ "category_id": 6,
+ "area": 86856,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17682,
+ "image_id": 3235,
+ "bbox": [
+ 92,
+ 125,
+ 419,
+ 384
+ ],
+ "category_id": 9,
+ "area": 425226,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17683,
+ "image_id": 3235,
+ "bbox": [
+ 97,
+ 27,
+ 128,
+ 256
+ ],
+ "category_id": 10,
+ "area": 87001,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17695,
+ "image_id": 3239,
+ "bbox": [
+ 225,
+ 37,
+ 183,
+ 292
+ ],
+ "category_id": 8,
+ "area": 142740,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17696,
+ "image_id": 3239,
+ "bbox": [
+ 158,
+ 243,
+ 265,
+ 221
+ ],
+ "category_id": 8,
+ "area": 156645,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17703,
+ "image_id": 3242,
+ "bbox": [
+ 79,
+ 172,
+ 127,
+ 151
+ ],
+ "category_id": 6,
+ "area": 39872,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17704,
+ "image_id": 3242,
+ "bbox": [
+ 65,
+ 76,
+ 157,
+ 157
+ ],
+ "category_id": 6,
+ "area": 50784,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17705,
+ "image_id": 3242,
+ "bbox": [
+ 254,
+ 5,
+ 164,
+ 207
+ ],
+ "category_id": 6,
+ "area": 70470,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17706,
+ "image_id": 3242,
+ "bbox": [
+ 232,
+ 216,
+ 278,
+ 279
+ ],
+ "category_id": 6,
+ "area": 160392,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17735,
+ "image_id": 3247,
+ "bbox": [
+ 12,
+ 438,
+ 86,
+ 71
+ ],
+ "category_id": 6,
+ "area": 21816,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17736,
+ "image_id": 3247,
+ "bbox": [
+ 400,
+ 256,
+ 110,
+ 254
+ ],
+ "category_id": 6,
+ "area": 98808,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17737,
+ "image_id": 3247,
+ "bbox": [
+ 0,
+ 270,
+ 122,
+ 184
+ ],
+ "category_id": 6,
+ "area": 79254,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17738,
+ "image_id": 3247,
+ "bbox": [
+ 0,
+ 147,
+ 192,
+ 258
+ ],
+ "category_id": 6,
+ "area": 174966,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17739,
+ "image_id": 3247,
+ "bbox": [
+ 131,
+ 352,
+ 231,
+ 158
+ ],
+ "category_id": 6,
+ "area": 129117,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17740,
+ "image_id": 3247,
+ "bbox": [
+ 247,
+ 292,
+ 82,
+ 144
+ ],
+ "category_id": 8,
+ "area": 41615,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17793,
+ "image_id": 3254,
+ "bbox": [
+ 195,
+ 47,
+ 107,
+ 240
+ ],
+ "category_id": 10,
+ "area": 45173,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17794,
+ "image_id": 3254,
+ "bbox": [
+ 343,
+ 305,
+ 73,
+ 157
+ ],
+ "category_id": 10,
+ "area": 20115,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17795,
+ "image_id": 3254,
+ "bbox": [
+ 435,
+ 220,
+ 76,
+ 265
+ ],
+ "category_id": 10,
+ "area": 35642,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17796,
+ "image_id": 3254,
+ "bbox": [
+ 90,
+ 399,
+ 79,
+ 112
+ ],
+ "category_id": 10,
+ "area": 15582,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17797,
+ "image_id": 3254,
+ "bbox": [
+ 8,
+ 326,
+ 77,
+ 155
+ ],
+ "category_id": 10,
+ "area": 21168,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17803,
+ "image_id": 3256,
+ "bbox": [
+ 0,
+ 159,
+ 123,
+ 233
+ ],
+ "category_id": 10,
+ "area": 93600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17804,
+ "image_id": 3256,
+ "bbox": [
+ 336,
+ 126,
+ 175,
+ 288
+ ],
+ "category_id": 9,
+ "area": 164822,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17805,
+ "image_id": 3257,
+ "bbox": [
+ 156,
+ 290,
+ 116,
+ 66
+ ],
+ "category_id": 8,
+ "area": 27156,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17806,
+ "image_id": 3257,
+ "bbox": [
+ 246,
+ 333,
+ 52,
+ 44
+ ],
+ "category_id": 8,
+ "area": 8316,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17829,
+ "image_id": 3261,
+ "bbox": [
+ 18,
+ 70,
+ 103,
+ 179
+ ],
+ "category_id": 8,
+ "area": 26832,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17830,
+ "image_id": 3261,
+ "bbox": [
+ 95,
+ 8,
+ 174,
+ 334
+ ],
+ "category_id": 8,
+ "area": 84681,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17831,
+ "image_id": 3261,
+ "bbox": [
+ 211,
+ 39,
+ 194,
+ 218
+ ],
+ "category_id": 8,
+ "area": 61560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17832,
+ "image_id": 3261,
+ "bbox": [
+ 426,
+ 40,
+ 83,
+ 199
+ ],
+ "category_id": 8,
+ "area": 24360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17834,
+ "image_id": 3262,
+ "bbox": [
+ 96,
+ 78,
+ 328,
+ 363
+ ],
+ "category_id": 9,
+ "area": 307272,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17839,
+ "image_id": 3264,
+ "bbox": [
+ 341,
+ 18,
+ 112,
+ 201
+ ],
+ "category_id": 9,
+ "area": 6534,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17840,
+ "image_id": 3264,
+ "bbox": [
+ 124,
+ 234,
+ 151,
+ 132
+ ],
+ "category_id": 9,
+ "area": 5785,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17852,
+ "image_id": 3264,
+ "bbox": [
+ 431,
+ 297,
+ 76,
+ 118
+ ],
+ "category_id": 6,
+ "area": 2610,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17853,
+ "image_id": 3264,
+ "bbox": [
+ 394,
+ 210,
+ 71,
+ 112
+ ],
+ "category_id": 6,
+ "area": 2310,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17854,
+ "image_id": 3264,
+ "bbox": [
+ 18,
+ 318,
+ 98,
+ 122
+ ],
+ "category_id": 6,
+ "area": 3480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17855,
+ "image_id": 3264,
+ "bbox": [
+ 150,
+ 59,
+ 175,
+ 175
+ ],
+ "category_id": 6,
+ "area": 8858,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17856,
+ "image_id": 3264,
+ "bbox": [
+ 356,
+ 297,
+ 64,
+ 65
+ ],
+ "category_id": 6,
+ "area": 1216,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17857,
+ "image_id": 3264,
+ "bbox": [
+ 312,
+ 379,
+ 97,
+ 75
+ ],
+ "category_id": 6,
+ "area": 2109,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17858,
+ "image_id": 3264,
+ "bbox": [
+ 64,
+ 161,
+ 83,
+ 83
+ ],
+ "category_id": 6,
+ "area": 2009,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17859,
+ "image_id": 3265,
+ "bbox": [
+ 30,
+ 34,
+ 353,
+ 473
+ ],
+ "category_id": 8,
+ "area": 159390,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17860,
+ "image_id": 3265,
+ "bbox": [
+ 22,
+ 193,
+ 485,
+ 315
+ ],
+ "category_id": 6,
+ "area": 145408,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17861,
+ "image_id": 3265,
+ "bbox": [
+ 217,
+ 0,
+ 291,
+ 226
+ ],
+ "category_id": 6,
+ "area": 62744,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17862,
+ "image_id": 3266,
+ "bbox": [
+ 75,
+ 0,
+ 380,
+ 368
+ ],
+ "category_id": 10,
+ "area": 1110984,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17871,
+ "image_id": 3269,
+ "bbox": [
+ 107,
+ 153,
+ 296,
+ 197
+ ],
+ "category_id": 9,
+ "area": 150568,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17875,
+ "image_id": 3271,
+ "bbox": [
+ 73,
+ 190,
+ 129,
+ 158
+ ],
+ "category_id": 9,
+ "area": 52858,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17876,
+ "image_id": 3271,
+ "bbox": [
+ 242,
+ 226,
+ 151,
+ 285
+ ],
+ "category_id": 6,
+ "area": 111069,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17877,
+ "image_id": 3271,
+ "bbox": [
+ 418,
+ 446,
+ 91,
+ 65
+ ],
+ "category_id": 6,
+ "area": 15397,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17878,
+ "image_id": 3272,
+ "bbox": [
+ 1,
+ 85,
+ 509,
+ 307
+ ],
+ "category_id": 9,
+ "area": 360774,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17879,
+ "image_id": 3273,
+ "bbox": [
+ 29,
+ 152,
+ 257,
+ 287
+ ],
+ "category_id": 9,
+ "area": 204058,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17881,
+ "image_id": 3273,
+ "bbox": [
+ 378,
+ 193,
+ 87,
+ 145
+ ],
+ "category_id": 6,
+ "area": 34974,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17882,
+ "image_id": 3273,
+ "bbox": [
+ 277,
+ 212,
+ 137,
+ 185
+ ],
+ "category_id": 6,
+ "area": 70144,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17883,
+ "image_id": 3273,
+ "bbox": [
+ 82,
+ 414,
+ 81,
+ 94
+ ],
+ "category_id": 6,
+ "area": 21190,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17884,
+ "image_id": 3273,
+ "bbox": [
+ 237,
+ 316,
+ 111,
+ 127
+ ],
+ "category_id": 6,
+ "area": 39072,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17885,
+ "image_id": 3273,
+ "bbox": [
+ 387,
+ 336,
+ 111,
+ 127
+ ],
+ "category_id": 6,
+ "area": 39072,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17894,
+ "image_id": 3275,
+ "bbox": [
+ 179,
+ 220,
+ 146,
+ 268
+ ],
+ "category_id": 10,
+ "area": 311300,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17895,
+ "image_id": 3275,
+ "bbox": [
+ 229,
+ 0,
+ 282,
+ 307
+ ],
+ "category_id": 9,
+ "area": 687528,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17896,
+ "image_id": 3276,
+ "bbox": [
+ 234,
+ 34,
+ 83,
+ 160
+ ],
+ "category_id": 9,
+ "area": 47234,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17897,
+ "image_id": 3276,
+ "bbox": [
+ 298,
+ 68,
+ 77,
+ 163
+ ],
+ "category_id": 9,
+ "area": 44390,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17898,
+ "image_id": 3276,
+ "bbox": [
+ 322,
+ 59,
+ 135,
+ 182
+ ],
+ "category_id": 9,
+ "area": 86866,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17899,
+ "image_id": 3276,
+ "bbox": [
+ 428,
+ 132,
+ 81,
+ 65
+ ],
+ "category_id": 9,
+ "area": 18676,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17908,
+ "image_id": 3278,
+ "bbox": [
+ 152,
+ 109,
+ 359,
+ 401
+ ],
+ "category_id": 10,
+ "area": 1140408,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17921,
+ "image_id": 3280,
+ "bbox": [
+ 52,
+ 66,
+ 321,
+ 445
+ ],
+ "category_id": 8,
+ "area": 503304,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17922,
+ "image_id": 3281,
+ "bbox": [
+ 429,
+ 307,
+ 80,
+ 204
+ ],
+ "category_id": 6,
+ "area": 21679,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17923,
+ "image_id": 3281,
+ "bbox": [
+ 0,
+ 342,
+ 48,
+ 124
+ ],
+ "category_id": 6,
+ "area": 8019,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17924,
+ "image_id": 3281,
+ "bbox": [
+ 66,
+ 381,
+ 93,
+ 130
+ ],
+ "category_id": 6,
+ "area": 16120,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17928,
+ "image_id": 3281,
+ "bbox": [
+ 190,
+ 179,
+ 115,
+ 276
+ ],
+ "category_id": 10,
+ "area": 42240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17943,
+ "image_id": 3284,
+ "bbox": [
+ 12,
+ 81,
+ 171,
+ 134
+ ],
+ "category_id": 10,
+ "area": 29700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17944,
+ "image_id": 3284,
+ "bbox": [
+ 86,
+ 33,
+ 55,
+ 49
+ ],
+ "category_id": 10,
+ "area": 3577,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17945,
+ "image_id": 3284,
+ "bbox": [
+ 48,
+ 379,
+ 62,
+ 104
+ ],
+ "category_id": 10,
+ "area": 8364,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17946,
+ "image_id": 3284,
+ "bbox": [
+ 41,
+ 265,
+ 62,
+ 88
+ ],
+ "category_id": 10,
+ "area": 7134,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17947,
+ "image_id": 3284,
+ "bbox": [
+ 8,
+ 353,
+ 55,
+ 93
+ ],
+ "category_id": 10,
+ "area": 6716,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17948,
+ "image_id": 3284,
+ "bbox": [
+ 129,
+ 240,
+ 99,
+ 119
+ ],
+ "category_id": 10,
+ "area": 15210,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17949,
+ "image_id": 3284,
+ "bbox": [
+ 209,
+ 246,
+ 59,
+ 98
+ ],
+ "category_id": 10,
+ "area": 7566,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17950,
+ "image_id": 3284,
+ "bbox": [
+ 239,
+ 194,
+ 51,
+ 80
+ ],
+ "category_id": 10,
+ "area": 5372,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17951,
+ "image_id": 3284,
+ "bbox": [
+ 141,
+ 0,
+ 62,
+ 55
+ ],
+ "category_id": 10,
+ "area": 4428,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17952,
+ "image_id": 3284,
+ "bbox": [
+ 350,
+ 201,
+ 68,
+ 103
+ ],
+ "category_id": 10,
+ "area": 9090,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17953,
+ "image_id": 3284,
+ "bbox": [
+ 357,
+ 373,
+ 77,
+ 77
+ ],
+ "category_id": 10,
+ "area": 7676,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17954,
+ "image_id": 3284,
+ "bbox": [
+ 393,
+ 317,
+ 66,
+ 65
+ ],
+ "category_id": 10,
+ "area": 5568,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17955,
+ "image_id": 3284,
+ "bbox": [
+ 428,
+ 182,
+ 68,
+ 82
+ ],
+ "category_id": 10,
+ "area": 7290,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17956,
+ "image_id": 3284,
+ "bbox": [
+ 440,
+ 91,
+ 71,
+ 88
+ ],
+ "category_id": 10,
+ "area": 8091,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17957,
+ "image_id": 3285,
+ "bbox": [
+ 20,
+ 407,
+ 101,
+ 97
+ ],
+ "category_id": 6,
+ "area": 27599,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17958,
+ "image_id": 3285,
+ "bbox": [
+ 105,
+ 405,
+ 125,
+ 105
+ ],
+ "category_id": 6,
+ "area": 36735,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17961,
+ "image_id": 3286,
+ "bbox": [
+ 310,
+ 13,
+ 175,
+ 223
+ ],
+ "category_id": 6,
+ "area": 310104,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17962,
+ "image_id": 3286,
+ "bbox": [
+ 216,
+ 87,
+ 104,
+ 233
+ ],
+ "category_id": 8,
+ "area": 193256,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17963,
+ "image_id": 3287,
+ "bbox": [
+ 175,
+ 315,
+ 129,
+ 192
+ ],
+ "category_id": 10,
+ "area": 32283,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17964,
+ "image_id": 3287,
+ "bbox": [
+ 195,
+ 100,
+ 223,
+ 254
+ ],
+ "category_id": 9,
+ "area": 73528,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17989,
+ "image_id": 3289,
+ "bbox": [
+ 198,
+ 278,
+ 90,
+ 144
+ ],
+ "category_id": 8,
+ "area": 45450,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17990,
+ "image_id": 3289,
+ "bbox": [
+ 263,
+ 296,
+ 88,
+ 59
+ ],
+ "category_id": 8,
+ "area": 18426,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17992,
+ "image_id": 3290,
+ "bbox": [
+ 113,
+ 457,
+ 115,
+ 50
+ ],
+ "category_id": 8,
+ "area": 39758,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17993,
+ "image_id": 3290,
+ "bbox": [
+ 4,
+ 186,
+ 503,
+ 315
+ ],
+ "category_id": 8,
+ "area": 1095250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17994,
+ "image_id": 3290,
+ "bbox": [
+ 242,
+ 8,
+ 228,
+ 358
+ ],
+ "category_id": 8,
+ "area": 564596,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17995,
+ "image_id": 3291,
+ "bbox": [
+ 27,
+ 132,
+ 23,
+ 113
+ ],
+ "category_id": 6,
+ "area": 6615,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17996,
+ "image_id": 3291,
+ "bbox": [
+ 296,
+ 78,
+ 132,
+ 100
+ ],
+ "category_id": 6,
+ "area": 33150,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17997,
+ "image_id": 3291,
+ "bbox": [
+ 85,
+ 106,
+ 373,
+ 359
+ ],
+ "category_id": 9,
+ "area": 333870,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 17998,
+ "image_id": 3292,
+ "bbox": [
+ 215,
+ 237,
+ 67,
+ 63
+ ],
+ "category_id": 8,
+ "area": 33782,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18000,
+ "image_id": 3293,
+ "bbox": [
+ 192,
+ 262,
+ 70,
+ 62
+ ],
+ "category_id": 8,
+ "area": 34846,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18002,
+ "image_id": 3294,
+ "bbox": [
+ 7,
+ 79,
+ 69,
+ 106
+ ],
+ "category_id": 8,
+ "area": 50180,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18003,
+ "image_id": 3294,
+ "bbox": [
+ 99,
+ 215,
+ 118,
+ 132
+ ],
+ "category_id": 8,
+ "area": 106763,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18004,
+ "image_id": 3294,
+ "bbox": [
+ 189,
+ 184,
+ 59,
+ 179
+ ],
+ "category_id": 8,
+ "area": 72698,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18005,
+ "image_id": 3294,
+ "bbox": [
+ 255,
+ 77,
+ 20,
+ 46
+ ],
+ "category_id": 8,
+ "area": 6468,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18006,
+ "image_id": 3294,
+ "bbox": [
+ 309,
+ 180,
+ 98,
+ 145
+ ],
+ "category_id": 8,
+ "area": 97520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18007,
+ "image_id": 3294,
+ "bbox": [
+ 396,
+ 181,
+ 32,
+ 67
+ ],
+ "category_id": 8,
+ "area": 15129,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18011,
+ "image_id": 3296,
+ "bbox": [
+ 168,
+ 197,
+ 49,
+ 47
+ ],
+ "category_id": 8,
+ "area": 7442,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18012,
+ "image_id": 3296,
+ "bbox": [
+ 207,
+ 219,
+ 103,
+ 64
+ ],
+ "category_id": 8,
+ "area": 21331,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18014,
+ "image_id": 3297,
+ "bbox": [
+ 88,
+ 107,
+ 338,
+ 194
+ ],
+ "category_id": 9,
+ "area": 289738,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18016,
+ "image_id": 3297,
+ "bbox": [
+ 0,
+ 360,
+ 512,
+ 151
+ ],
+ "category_id": 6,
+ "area": 340800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18017,
+ "image_id": 3298,
+ "bbox": [
+ 0,
+ 383,
+ 53,
+ 73
+ ],
+ "category_id": 10,
+ "area": 6732,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18018,
+ "image_id": 3298,
+ "bbox": [
+ 35,
+ 416,
+ 58,
+ 56
+ ],
+ "category_id": 10,
+ "area": 5668,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18019,
+ "image_id": 3298,
+ "bbox": [
+ 86,
+ 352,
+ 58,
+ 71
+ ],
+ "category_id": 10,
+ "area": 7128,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18020,
+ "image_id": 3298,
+ "bbox": [
+ 22,
+ 101,
+ 34,
+ 30
+ ],
+ "category_id": 10,
+ "area": 1792,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18021,
+ "image_id": 3298,
+ "bbox": [
+ 97,
+ 58,
+ 37,
+ 37
+ ],
+ "category_id": 10,
+ "area": 2450,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18022,
+ "image_id": 3298,
+ "bbox": [
+ 131,
+ 154,
+ 58,
+ 56
+ ],
+ "category_id": 10,
+ "area": 5616,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18023,
+ "image_id": 3298,
+ "bbox": [
+ 152,
+ 198,
+ 35,
+ 46
+ ],
+ "category_id": 10,
+ "area": 2795,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18024,
+ "image_id": 3298,
+ "bbox": [
+ 189,
+ 99,
+ 39,
+ 43
+ ],
+ "category_id": 10,
+ "area": 2960,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18025,
+ "image_id": 3298,
+ "bbox": [
+ 140,
+ 82,
+ 28,
+ 32
+ ],
+ "category_id": 10,
+ "area": 1590,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18026,
+ "image_id": 3298,
+ "bbox": [
+ 277,
+ 312,
+ 53,
+ 71
+ ],
+ "category_id": 10,
+ "area": 6534,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18027,
+ "image_id": 3298,
+ "bbox": [
+ 289,
+ 167,
+ 38,
+ 50
+ ],
+ "category_id": 10,
+ "area": 3384,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18028,
+ "image_id": 3298,
+ "bbox": [
+ 143,
+ 22,
+ 29,
+ 32
+ ],
+ "category_id": 10,
+ "area": 1620,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18029,
+ "image_id": 3298,
+ "bbox": [
+ 162,
+ 2,
+ 19,
+ 22
+ ],
+ "category_id": 10,
+ "area": 777,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18030,
+ "image_id": 3298,
+ "bbox": [
+ 224,
+ 15,
+ 33,
+ 49
+ ],
+ "category_id": 10,
+ "area": 2852,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18031,
+ "image_id": 3298,
+ "bbox": [
+ 259,
+ 49,
+ 63,
+ 43
+ ],
+ "category_id": 10,
+ "area": 4680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18052,
+ "image_id": 3299,
+ "bbox": [
+ 141,
+ 166,
+ 272,
+ 216
+ ],
+ "category_id": 9,
+ "area": 95850,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18053,
+ "image_id": 3299,
+ "bbox": [
+ 263,
+ 451,
+ 64,
+ 60
+ ],
+ "category_id": 6,
+ "area": 6363,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18054,
+ "image_id": 3299,
+ "bbox": [
+ 428,
+ 174,
+ 83,
+ 152
+ ],
+ "category_id": 6,
+ "area": 20670,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18055,
+ "image_id": 3300,
+ "bbox": [
+ 0,
+ 79,
+ 417,
+ 221
+ ],
+ "category_id": 9,
+ "area": 81770,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18058,
+ "image_id": 3301,
+ "bbox": [
+ 169,
+ 452,
+ 97,
+ 59
+ ],
+ "category_id": 9,
+ "area": 6916,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18060,
+ "image_id": 3303,
+ "bbox": [
+ 200,
+ 59,
+ 142,
+ 145
+ ],
+ "category_id": 9,
+ "area": 72624,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18061,
+ "image_id": 3303,
+ "bbox": [
+ 150,
+ 127,
+ 183,
+ 88
+ ],
+ "category_id": 9,
+ "area": 56792,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18063,
+ "image_id": 3304,
+ "bbox": [
+ 303,
+ 67,
+ 189,
+ 149
+ ],
+ "category_id": 9,
+ "area": 53820,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18064,
+ "image_id": 3304,
+ "bbox": [
+ 314,
+ 175,
+ 124,
+ 120
+ ],
+ "category_id": 10,
+ "area": 28417,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18065,
+ "image_id": 3304,
+ "bbox": [
+ 0,
+ 170,
+ 512,
+ 340
+ ],
+ "category_id": 6,
+ "area": 330780,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18070,
+ "image_id": 3305,
+ "bbox": [
+ 15,
+ 157,
+ 197,
+ 273
+ ],
+ "category_id": 6,
+ "area": 416232,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18071,
+ "image_id": 3305,
+ "bbox": [
+ 209,
+ 26,
+ 297,
+ 474
+ ],
+ "category_id": 6,
+ "area": 1090332,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18072,
+ "image_id": 3305,
+ "bbox": [
+ 381,
+ 83,
+ 65,
+ 121
+ ],
+ "category_id": 6,
+ "area": 61250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18073,
+ "image_id": 3305,
+ "bbox": [
+ 446,
+ 270,
+ 44,
+ 88
+ ],
+ "category_id": 6,
+ "area": 30195,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18103,
+ "image_id": 3306,
+ "bbox": [
+ 83,
+ 268,
+ 108,
+ 67
+ ],
+ "category_id": 6,
+ "area": 6996,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18104,
+ "image_id": 3306,
+ "bbox": [
+ 0,
+ 294,
+ 66,
+ 106
+ ],
+ "category_id": 6,
+ "area": 6760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18105,
+ "image_id": 3306,
+ "bbox": [
+ 0,
+ 458,
+ 115,
+ 53
+ ],
+ "category_id": 6,
+ "area": 5876,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18106,
+ "image_id": 3306,
+ "bbox": [
+ 395,
+ 76,
+ 114,
+ 177
+ ],
+ "category_id": 6,
+ "area": 19376,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18107,
+ "image_id": 3306,
+ "bbox": [
+ 425,
+ 336,
+ 86,
+ 159
+ ],
+ "category_id": 6,
+ "area": 13104,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18119,
+ "image_id": 3308,
+ "bbox": [
+ 117,
+ 187,
+ 338,
+ 222
+ ],
+ "category_id": 9,
+ "area": 193844,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18120,
+ "image_id": 3308,
+ "bbox": [
+ 169,
+ 93,
+ 102,
+ 121
+ ],
+ "category_id": 6,
+ "area": 32175,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18121,
+ "image_id": 3308,
+ "bbox": [
+ 307,
+ 0,
+ 195,
+ 169
+ ],
+ "category_id": 6,
+ "area": 85560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18126,
+ "image_id": 3310,
+ "bbox": [
+ 158,
+ 49,
+ 219,
+ 430
+ ],
+ "category_id": 9,
+ "area": 331540,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18130,
+ "image_id": 3310,
+ "bbox": [
+ 2,
+ 132,
+ 124,
+ 115
+ ],
+ "category_id": 6,
+ "area": 50693,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18131,
+ "image_id": 3310,
+ "bbox": [
+ 138,
+ 307,
+ 82,
+ 111
+ ],
+ "category_id": 6,
+ "area": 32499,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18132,
+ "image_id": 3310,
+ "bbox": [
+ 229,
+ 0,
+ 261,
+ 506
+ ],
+ "category_id": 6,
+ "area": 465648,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18134,
+ "image_id": 3311,
+ "bbox": [
+ 227,
+ 47,
+ 122,
+ 249
+ ],
+ "category_id": 6,
+ "area": 486987,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18135,
+ "image_id": 3311,
+ "bbox": [
+ 329,
+ 111,
+ 80,
+ 141
+ ],
+ "category_id": 6,
+ "area": 181566,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18136,
+ "image_id": 3311,
+ "bbox": [
+ 398,
+ 137,
+ 109,
+ 198
+ ],
+ "category_id": 6,
+ "area": 345610,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18137,
+ "image_id": 3312,
+ "bbox": [
+ 9,
+ 95,
+ 343,
+ 215
+ ],
+ "category_id": 9,
+ "area": 46900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18156,
+ "image_id": 3314,
+ "bbox": [
+ 188,
+ 135,
+ 174,
+ 126
+ ],
+ "category_id": 8,
+ "area": 77349,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18157,
+ "image_id": 3314,
+ "bbox": [
+ 192,
+ 252,
+ 161,
+ 186
+ ],
+ "category_id": 8,
+ "area": 105444,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18158,
+ "image_id": 3315,
+ "bbox": [
+ 17,
+ 272,
+ 204,
+ 164
+ ],
+ "category_id": 8,
+ "area": 49280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18159,
+ "image_id": 3315,
+ "bbox": [
+ 168,
+ 132,
+ 163,
+ 273
+ ],
+ "category_id": 8,
+ "area": 65280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18160,
+ "image_id": 3315,
+ "bbox": [
+ 240,
+ 129,
+ 237,
+ 318
+ ],
+ "category_id": 8,
+ "area": 110929,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18161,
+ "image_id": 3316,
+ "bbox": [
+ 0,
+ 178,
+ 269,
+ 251
+ ],
+ "category_id": 9,
+ "area": 238242,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18162,
+ "image_id": 3316,
+ "bbox": [
+ 248,
+ 132,
+ 134,
+ 179
+ ],
+ "category_id": 9,
+ "area": 85261,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18163,
+ "image_id": 3316,
+ "bbox": [
+ 248,
+ 230,
+ 127,
+ 156
+ ],
+ "category_id": 9,
+ "area": 70180,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18164,
+ "image_id": 3316,
+ "bbox": [
+ 253,
+ 376,
+ 256,
+ 132
+ ],
+ "category_id": 9,
+ "area": 119040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18165,
+ "image_id": 3317,
+ "bbox": [
+ 240,
+ 229,
+ 90,
+ 109
+ ],
+ "category_id": 8,
+ "area": 34958,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18166,
+ "image_id": 3318,
+ "bbox": [
+ 117,
+ 125,
+ 379,
+ 318
+ ],
+ "category_id": 9,
+ "area": 174720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18167,
+ "image_id": 3318,
+ "bbox": [
+ 62,
+ 144,
+ 96,
+ 73
+ ],
+ "category_id": 10,
+ "area": 10296,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18168,
+ "image_id": 3319,
+ "bbox": [
+ 68,
+ 217,
+ 203,
+ 174
+ ],
+ "category_id": 8,
+ "area": 124460,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18169,
+ "image_id": 3319,
+ "bbox": [
+ 193,
+ 206,
+ 228,
+ 210
+ ],
+ "category_id": 8,
+ "area": 168740,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18180,
+ "image_id": 3321,
+ "bbox": [
+ 19,
+ 197,
+ 101,
+ 186
+ ],
+ "category_id": 8,
+ "area": 43560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18181,
+ "image_id": 3321,
+ "bbox": [
+ 98,
+ 60,
+ 164,
+ 423
+ ],
+ "category_id": 8,
+ "area": 160000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18182,
+ "image_id": 3321,
+ "bbox": [
+ 263,
+ 284,
+ 244,
+ 207
+ ],
+ "category_id": 8,
+ "area": 116375,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18183,
+ "image_id": 3322,
+ "bbox": [
+ 216,
+ 409,
+ 108,
+ 102
+ ],
+ "category_id": 6,
+ "area": 83430,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18184,
+ "image_id": 3322,
+ "bbox": [
+ 326,
+ 400,
+ 177,
+ 111
+ ],
+ "category_id": 6,
+ "area": 147616,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18201,
+ "image_id": 3326,
+ "bbox": [
+ 215,
+ 226,
+ 165,
+ 205
+ ],
+ "category_id": 6,
+ "area": 48843,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18202,
+ "image_id": 3326,
+ "bbox": [
+ 243,
+ 419,
+ 209,
+ 92
+ ],
+ "category_id": 9,
+ "area": 27720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18236,
+ "image_id": 3332,
+ "bbox": [
+ 214,
+ 153,
+ 264,
+ 349
+ ],
+ "category_id": 8,
+ "area": 107580,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18238,
+ "image_id": 3333,
+ "bbox": [
+ 16,
+ 177,
+ 450,
+ 332
+ ],
+ "category_id": 9,
+ "area": 526968,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18239,
+ "image_id": 3333,
+ "bbox": [
+ 197,
+ 41,
+ 138,
+ 226
+ ],
+ "category_id": 10,
+ "area": 110055,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18244,
+ "image_id": 3336,
+ "bbox": [
+ 217,
+ 252,
+ 150,
+ 177
+ ],
+ "category_id": 6,
+ "area": 38233,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18245,
+ "image_id": 3336,
+ "bbox": [
+ 211,
+ 420,
+ 189,
+ 90
+ ],
+ "category_id": 9,
+ "area": 24552,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18250,
+ "image_id": 3338,
+ "bbox": [
+ 194,
+ 7,
+ 111,
+ 111
+ ],
+ "category_id": 8,
+ "area": 98884,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18251,
+ "image_id": 3338,
+ "bbox": [
+ 168,
+ 107,
+ 153,
+ 182
+ ],
+ "category_id": 8,
+ "area": 220990,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18252,
+ "image_id": 3338,
+ "bbox": [
+ 0,
+ 60,
+ 94,
+ 48
+ ],
+ "category_id": 8,
+ "area": 36668,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18253,
+ "image_id": 3339,
+ "bbox": [
+ 101,
+ 135,
+ 410,
+ 270
+ ],
+ "category_id": 8,
+ "area": 299466,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18254,
+ "image_id": 3339,
+ "bbox": [
+ 84,
+ 97,
+ 344,
+ 157
+ ],
+ "category_id": 8,
+ "area": 146520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18267,
+ "image_id": 3342,
+ "bbox": [
+ 2,
+ 0,
+ 506,
+ 507
+ ],
+ "category_id": 9,
+ "area": 904638,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18268,
+ "image_id": 3343,
+ "bbox": [
+ 139,
+ 325,
+ 372,
+ 185
+ ],
+ "category_id": 8,
+ "area": 243252,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18269,
+ "image_id": 3343,
+ "bbox": [
+ 146,
+ 134,
+ 356,
+ 172
+ ],
+ "category_id": 8,
+ "area": 215864,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18270,
+ "image_id": 3343,
+ "bbox": [
+ 155,
+ 27,
+ 281,
+ 167
+ ],
+ "category_id": 8,
+ "area": 165440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18271,
+ "image_id": 3344,
+ "bbox": [
+ 154,
+ 86,
+ 299,
+ 420
+ ],
+ "category_id": 10,
+ "area": 205016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18272,
+ "image_id": 3345,
+ "bbox": [
+ 57,
+ 352,
+ 15,
+ 63
+ ],
+ "category_id": 6,
+ "area": 7128,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18273,
+ "image_id": 3345,
+ "bbox": [
+ 110,
+ 318,
+ 29,
+ 42
+ ],
+ "category_id": 6,
+ "area": 8989,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18274,
+ "image_id": 3345,
+ "bbox": [
+ 133,
+ 305,
+ 21,
+ 44
+ ],
+ "category_id": 6,
+ "area": 6696,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18275,
+ "image_id": 3345,
+ "bbox": [
+ 35,
+ 89,
+ 29,
+ 36
+ ],
+ "category_id": 6,
+ "area": 7777,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18276,
+ "image_id": 3345,
+ "bbox": [
+ 392,
+ 456,
+ 45,
+ 55
+ ],
+ "category_id": 6,
+ "area": 17748,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18277,
+ "image_id": 3345,
+ "bbox": [
+ 355,
+ 273,
+ 7,
+ 23
+ ],
+ "category_id": 6,
+ "area": 1350,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18278,
+ "image_id": 3345,
+ "bbox": [
+ 373,
+ 351,
+ 11,
+ 22
+ ],
+ "category_id": 6,
+ "area": 1840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18301,
+ "image_id": 3348,
+ "bbox": [
+ 38,
+ 160,
+ 368,
+ 350
+ ],
+ "category_id": 9,
+ "area": 320542,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18302,
+ "image_id": 3349,
+ "bbox": [
+ 149,
+ 158,
+ 139,
+ 141
+ ],
+ "category_id": 9,
+ "area": 16068,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18319,
+ "image_id": 3353,
+ "bbox": [
+ 0,
+ 174,
+ 299,
+ 337
+ ],
+ "category_id": 6,
+ "area": 281200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18320,
+ "image_id": 3353,
+ "bbox": [
+ 1,
+ 240,
+ 507,
+ 271
+ ],
+ "category_id": 6,
+ "area": 383910,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18327,
+ "image_id": 3354,
+ "bbox": [
+ 0,
+ 236,
+ 512,
+ 275
+ ],
+ "category_id": 6,
+ "area": 219816,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18328,
+ "image_id": 3354,
+ "bbox": [
+ 169,
+ 309,
+ 117,
+ 202
+ ],
+ "category_id": 6,
+ "area": 37050,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18329,
+ "image_id": 3355,
+ "bbox": [
+ 62,
+ 217,
+ 193,
+ 225
+ ],
+ "category_id": 8,
+ "area": 152628,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18330,
+ "image_id": 3355,
+ "bbox": [
+ 221,
+ 272,
+ 137,
+ 239
+ ],
+ "category_id": 8,
+ "area": 115240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18331,
+ "image_id": 3355,
+ "bbox": [
+ 284,
+ 384,
+ 165,
+ 121
+ ],
+ "category_id": 8,
+ "area": 70380,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18332,
+ "image_id": 3356,
+ "bbox": [
+ 148,
+ 0,
+ 363,
+ 488
+ ],
+ "category_id": 9,
+ "area": 5609380,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18333,
+ "image_id": 3356,
+ "bbox": [
+ 0,
+ 0,
+ 189,
+ 383
+ ],
+ "category_id": 9,
+ "area": 2296336,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18334,
+ "image_id": 3356,
+ "bbox": [
+ 166,
+ 293,
+ 257,
+ 202
+ ],
+ "category_id": 9,
+ "area": 1653792,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18339,
+ "image_id": 3358,
+ "bbox": [
+ 82,
+ 198,
+ 201,
+ 157
+ ],
+ "category_id": 9,
+ "area": 250660,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18340,
+ "image_id": 3358,
+ "bbox": [
+ 184,
+ 207,
+ 207,
+ 189
+ ],
+ "category_id": 10,
+ "area": 311600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18342,
+ "image_id": 3359,
+ "bbox": [
+ 0,
+ 189,
+ 260,
+ 205
+ ],
+ "category_id": 9,
+ "area": 138992,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18343,
+ "image_id": 3359,
+ "bbox": [
+ 201,
+ 114,
+ 137,
+ 243
+ ],
+ "category_id": 9,
+ "area": 86881,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18344,
+ "image_id": 3359,
+ "bbox": [
+ 211,
+ 284,
+ 300,
+ 226
+ ],
+ "category_id": 9,
+ "area": 176736,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18351,
+ "image_id": 3361,
+ "bbox": [
+ 217,
+ 0,
+ 205,
+ 360
+ ],
+ "category_id": 6,
+ "area": 100750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18369,
+ "image_id": 3366,
+ "bbox": [
+ 177,
+ 142,
+ 276,
+ 287
+ ],
+ "category_id": 8,
+ "area": 67133,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18370,
+ "image_id": 3367,
+ "bbox": [
+ 242,
+ 233,
+ 127,
+ 209
+ ],
+ "category_id": 8,
+ "area": 86178,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18371,
+ "image_id": 3368,
+ "bbox": [
+ 139,
+ 80,
+ 276,
+ 431
+ ],
+ "category_id": 10,
+ "area": 198924,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18372,
+ "image_id": 3369,
+ "bbox": [
+ 81,
+ 30,
+ 358,
+ 469
+ ],
+ "category_id": 10,
+ "area": 1738286,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18384,
+ "image_id": 3371,
+ "bbox": [
+ 121,
+ 50,
+ 103,
+ 414
+ ],
+ "category_id": 6,
+ "area": 329258,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18398,
+ "image_id": 3374,
+ "bbox": [
+ 151,
+ 0,
+ 360,
+ 271
+ ],
+ "category_id": 9,
+ "area": 266057,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18399,
+ "image_id": 3374,
+ "bbox": [
+ 166,
+ 190,
+ 344,
+ 321
+ ],
+ "category_id": 9,
+ "area": 301035,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18402,
+ "image_id": 3376,
+ "bbox": [
+ 114,
+ 144,
+ 255,
+ 124
+ ],
+ "category_id": 9,
+ "area": 31411,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18403,
+ "image_id": 3376,
+ "bbox": [
+ 134,
+ 269,
+ 306,
+ 236
+ ],
+ "category_id": 6,
+ "area": 71616,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18427,
+ "image_id": 3378,
+ "bbox": [
+ 192,
+ 67,
+ 117,
+ 133
+ ],
+ "category_id": 6,
+ "area": 18375,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18428,
+ "image_id": 3378,
+ "bbox": [
+ 90,
+ 305,
+ 88,
+ 120
+ ],
+ "category_id": 6,
+ "area": 12543,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18430,
+ "image_id": 3379,
+ "bbox": [
+ 175,
+ 279,
+ 237,
+ 170
+ ],
+ "category_id": 10,
+ "area": 63945,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18438,
+ "image_id": 3381,
+ "bbox": [
+ 52,
+ 398,
+ 91,
+ 99
+ ],
+ "category_id": 6,
+ "area": 25340,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18439,
+ "image_id": 3381,
+ "bbox": [
+ 120,
+ 396,
+ 114,
+ 115
+ ],
+ "category_id": 6,
+ "area": 36838,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18467,
+ "image_id": 3386,
+ "bbox": [
+ 393,
+ 157,
+ 116,
+ 161
+ ],
+ "category_id": 9,
+ "area": 31688,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18468,
+ "image_id": 3386,
+ "bbox": [
+ 279,
+ 112,
+ 139,
+ 203
+ ],
+ "category_id": 9,
+ "area": 47988,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18469,
+ "image_id": 3386,
+ "bbox": [
+ 193,
+ 58,
+ 123,
+ 205
+ ],
+ "category_id": 9,
+ "area": 42731,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18470,
+ "image_id": 3386,
+ "bbox": [
+ 163,
+ 20,
+ 97,
+ 110
+ ],
+ "category_id": 9,
+ "area": 18042,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18478,
+ "image_id": 3388,
+ "bbox": [
+ 158,
+ 212,
+ 124,
+ 179
+ ],
+ "category_id": 8,
+ "area": 31122,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18479,
+ "image_id": 3388,
+ "bbox": [
+ 119,
+ 222,
+ 35,
+ 130
+ ],
+ "category_id": 6,
+ "area": 6402,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18480,
+ "image_id": 3388,
+ "bbox": [
+ 118,
+ 6,
+ 49,
+ 130
+ ],
+ "category_id": 6,
+ "area": 9021,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18481,
+ "image_id": 3388,
+ "bbox": [
+ 158,
+ 99,
+ 53,
+ 157
+ ],
+ "category_id": 6,
+ "area": 11700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18482,
+ "image_id": 3388,
+ "bbox": [
+ 204,
+ 66,
+ 21,
+ 165
+ ],
+ "category_id": 6,
+ "area": 5043,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18483,
+ "image_id": 3388,
+ "bbox": [
+ 225,
+ 37,
+ 35,
+ 134
+ ],
+ "category_id": 6,
+ "area": 6700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18484,
+ "image_id": 3388,
+ "bbox": [
+ 245,
+ 20,
+ 60,
+ 130
+ ],
+ "category_id": 6,
+ "area": 11058,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18485,
+ "image_id": 3388,
+ "bbox": [
+ 283,
+ 0,
+ 161,
+ 102
+ ],
+ "category_id": 6,
+ "area": 22952,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18490,
+ "image_id": 3391,
+ "bbox": [
+ 87,
+ 183,
+ 183,
+ 176
+ ],
+ "category_id": 8,
+ "area": 113126,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18491,
+ "image_id": 3391,
+ "bbox": [
+ 216,
+ 183,
+ 217,
+ 197
+ ],
+ "category_id": 8,
+ "area": 149868,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18492,
+ "image_id": 3391,
+ "bbox": [
+ 109,
+ 109,
+ 23,
+ 22
+ ],
+ "category_id": 6,
+ "area": 1856,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18493,
+ "image_id": 3391,
+ "bbox": [
+ 141,
+ 147,
+ 26,
+ 36
+ ],
+ "category_id": 6,
+ "area": 3315,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18494,
+ "image_id": 3391,
+ "bbox": [
+ 154,
+ 189,
+ 26,
+ 39
+ ],
+ "category_id": 6,
+ "area": 3640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18496,
+ "image_id": 3391,
+ "bbox": [
+ 158,
+ 375,
+ 38,
+ 39
+ ],
+ "category_id": 6,
+ "area": 5225,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18506,
+ "image_id": 3395,
+ "bbox": [
+ 0,
+ 398,
+ 160,
+ 110
+ ],
+ "category_id": 6,
+ "area": 20904,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18507,
+ "image_id": 3395,
+ "bbox": [
+ 460,
+ 457,
+ 46,
+ 52
+ ],
+ "category_id": 6,
+ "area": 2842,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18508,
+ "image_id": 3395,
+ "bbox": [
+ 404,
+ 391,
+ 52,
+ 116
+ ],
+ "category_id": 6,
+ "area": 7085,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18509,
+ "image_id": 3395,
+ "bbox": [
+ 348,
+ 370,
+ 84,
+ 128
+ ],
+ "category_id": 6,
+ "area": 12720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18510,
+ "image_id": 3395,
+ "bbox": [
+ 393,
+ 0,
+ 88,
+ 84
+ ],
+ "category_id": 6,
+ "area": 8769,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18513,
+ "image_id": 3396,
+ "bbox": [
+ 324,
+ 360,
+ 51,
+ 51
+ ],
+ "category_id": 10,
+ "area": 206460,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18514,
+ "image_id": 3396,
+ "bbox": [
+ 0,
+ 145,
+ 62,
+ 70
+ ],
+ "category_id": 10,
+ "area": 348492,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18515,
+ "image_id": 3396,
+ "bbox": [
+ 69,
+ 32,
+ 57,
+ 50
+ ],
+ "category_id": 10,
+ "area": 229482,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18516,
+ "image_id": 3396,
+ "bbox": [
+ 210,
+ 143,
+ 47,
+ 42
+ ],
+ "category_id": 10,
+ "area": 160160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18517,
+ "image_id": 3396,
+ "bbox": [
+ 259,
+ 394,
+ 29,
+ 41
+ ],
+ "category_id": 10,
+ "area": 96320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18518,
+ "image_id": 3396,
+ "bbox": [
+ 423,
+ 115,
+ 46,
+ 66
+ ],
+ "category_id": 10,
+ "area": 244881,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18524,
+ "image_id": 3398,
+ "bbox": [
+ 281,
+ 99,
+ 107,
+ 237
+ ],
+ "category_id": 10,
+ "area": 89846,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18526,
+ "image_id": 3399,
+ "bbox": [
+ 159,
+ 305,
+ 226,
+ 200
+ ],
+ "category_id": 10,
+ "area": 71700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18544,
+ "image_id": 3401,
+ "bbox": [
+ 0,
+ 240,
+ 223,
+ 149
+ ],
+ "category_id": 8,
+ "area": 117390,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18545,
+ "image_id": 3401,
+ "bbox": [
+ 282,
+ 249,
+ 193,
+ 112
+ ],
+ "category_id": 8,
+ "area": 76472,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18547,
+ "image_id": 3402,
+ "bbox": [
+ 171,
+ 234,
+ 135,
+ 106
+ ],
+ "category_id": 8,
+ "area": 46306,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18548,
+ "image_id": 3402,
+ "bbox": [
+ 368,
+ 194,
+ 60,
+ 85
+ ],
+ "category_id": 8,
+ "area": 16459,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18555,
+ "image_id": 3404,
+ "bbox": [
+ 150,
+ 130,
+ 64,
+ 161
+ ],
+ "category_id": 8,
+ "area": 36612,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18556,
+ "image_id": 3404,
+ "bbox": [
+ 230,
+ 234,
+ 142,
+ 249
+ ],
+ "category_id": 8,
+ "area": 124244,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18577,
+ "image_id": 3409,
+ "bbox": [
+ 18,
+ 141,
+ 383,
+ 370
+ ],
+ "category_id": 9,
+ "area": 350625,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18578,
+ "image_id": 3409,
+ "bbox": [
+ 394,
+ 222,
+ 91,
+ 153
+ ],
+ "category_id": 6,
+ "area": 34672,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18579,
+ "image_id": 3409,
+ "bbox": [
+ 400,
+ 431,
+ 43,
+ 68
+ ],
+ "category_id": 6,
+ "area": 7254,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18580,
+ "image_id": 3410,
+ "bbox": [
+ 42,
+ 197,
+ 205,
+ 229
+ ],
+ "category_id": 8,
+ "area": 164673,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18581,
+ "image_id": 3410,
+ "bbox": [
+ 215,
+ 272,
+ 150,
+ 227
+ ],
+ "category_id": 8,
+ "area": 119625,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18582,
+ "image_id": 3410,
+ "bbox": [
+ 271,
+ 378,
+ 178,
+ 132
+ ],
+ "category_id": 8,
+ "area": 82510,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18626,
+ "image_id": 3416,
+ "bbox": [
+ 83,
+ 106,
+ 219,
+ 255
+ ],
+ "category_id": 8,
+ "area": 443312,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18627,
+ "image_id": 3416,
+ "bbox": [
+ 0,
+ 202,
+ 181,
+ 307
+ ],
+ "category_id": 6,
+ "area": 441320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18628,
+ "image_id": 3416,
+ "bbox": [
+ 36,
+ 408,
+ 130,
+ 101
+ ],
+ "category_id": 6,
+ "area": 104646,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18629,
+ "image_id": 3416,
+ "bbox": [
+ 226,
+ 315,
+ 62,
+ 101
+ ],
+ "category_id": 6,
+ "area": 50740,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18630,
+ "image_id": 3416,
+ "bbox": [
+ 340,
+ 364,
+ 95,
+ 145
+ ],
+ "category_id": 6,
+ "area": 110213,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18631,
+ "image_id": 3416,
+ "bbox": [
+ 421,
+ 457,
+ 43,
+ 53
+ ],
+ "category_id": 6,
+ "area": 18368,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18640,
+ "image_id": 3419,
+ "bbox": [
+ 200,
+ 332,
+ 132,
+ 146
+ ],
+ "category_id": 9,
+ "area": 25506,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18641,
+ "image_id": 3420,
+ "bbox": [
+ 216,
+ 204,
+ 117,
+ 155
+ ],
+ "category_id": 8,
+ "area": 144534,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18642,
+ "image_id": 3421,
+ "bbox": [
+ 128,
+ 219,
+ 123,
+ 200
+ ],
+ "category_id": 10,
+ "area": 80668,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18643,
+ "image_id": 3421,
+ "bbox": [
+ 128,
+ 80,
+ 132,
+ 199
+ ],
+ "category_id": 9,
+ "area": 85652,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18648,
+ "image_id": 3422,
+ "bbox": [
+ 0,
+ 0,
+ 512,
+ 248
+ ],
+ "category_id": 6,
+ "area": 307173,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18653,
+ "image_id": 3423,
+ "bbox": [
+ 0,
+ 0,
+ 351,
+ 512
+ ],
+ "category_id": 6,
+ "area": 1313232,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18654,
+ "image_id": 3424,
+ "bbox": [
+ 13,
+ 9,
+ 262,
+ 242
+ ],
+ "category_id": 9,
+ "area": 182792,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18656,
+ "image_id": 3424,
+ "bbox": [
+ 262,
+ 2,
+ 127,
+ 209
+ ],
+ "category_id": 10,
+ "area": 76950,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18687,
+ "image_id": 3428,
+ "bbox": [
+ 1,
+ 19,
+ 69,
+ 154
+ ],
+ "category_id": 8,
+ "area": 15660,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18688,
+ "image_id": 3428,
+ "bbox": [
+ 1,
+ 80,
+ 222,
+ 244
+ ],
+ "category_id": 8,
+ "area": 79023,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18689,
+ "image_id": 3428,
+ "bbox": [
+ 215,
+ 21,
+ 195,
+ 285
+ ],
+ "category_id": 8,
+ "area": 81174,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18699,
+ "image_id": 3430,
+ "bbox": [
+ 0,
+ 103,
+ 512,
+ 397
+ ],
+ "category_id": 9,
+ "area": 109872,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18700,
+ "image_id": 3431,
+ "bbox": [
+ 0,
+ 0,
+ 106,
+ 162
+ ],
+ "category_id": 6,
+ "area": 39960,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18701,
+ "image_id": 3431,
+ "bbox": [
+ 213,
+ 183,
+ 220,
+ 295
+ ],
+ "category_id": 9,
+ "area": 150519,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18702,
+ "image_id": 3432,
+ "bbox": [
+ 169,
+ 420,
+ 142,
+ 91
+ ],
+ "category_id": 6,
+ "area": 45568,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18703,
+ "image_id": 3432,
+ "bbox": [
+ 356,
+ 224,
+ 154,
+ 285
+ ],
+ "category_id": 6,
+ "area": 155172,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18704,
+ "image_id": 3432,
+ "bbox": [
+ 128,
+ 231,
+ 234,
+ 192
+ ],
+ "category_id": 6,
+ "area": 157950,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18705,
+ "image_id": 3432,
+ "bbox": [
+ 0,
+ 215,
+ 193,
+ 295
+ ],
+ "category_id": 8,
+ "area": 200928,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18708,
+ "image_id": 3433,
+ "bbox": [
+ 0,
+ 113,
+ 75,
+ 113
+ ],
+ "category_id": 6,
+ "area": 39424,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18709,
+ "image_id": 3433,
+ "bbox": [
+ 0,
+ 151,
+ 313,
+ 327
+ ],
+ "category_id": 6,
+ "area": 470640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18710,
+ "image_id": 3433,
+ "bbox": [
+ 312,
+ 431,
+ 136,
+ 47
+ ],
+ "category_id": 6,
+ "area": 29568,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18711,
+ "image_id": 3433,
+ "bbox": [
+ 351,
+ 262,
+ 160,
+ 191
+ ],
+ "category_id": 6,
+ "area": 140649,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18723,
+ "image_id": 3434,
+ "bbox": [
+ 155,
+ 127,
+ 324,
+ 187
+ ],
+ "category_id": 9,
+ "area": 66690,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18738,
+ "image_id": 3438,
+ "bbox": [
+ 1,
+ 14,
+ 81,
+ 115
+ ],
+ "category_id": 6,
+ "area": 78030,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18739,
+ "image_id": 3438,
+ "bbox": [
+ 98,
+ 0,
+ 61,
+ 65
+ ],
+ "category_id": 6,
+ "area": 33418,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18740,
+ "image_id": 3438,
+ "bbox": [
+ 126,
+ 54,
+ 61,
+ 128
+ ],
+ "category_id": 6,
+ "area": 65618,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18741,
+ "image_id": 3438,
+ "bbox": [
+ 201,
+ 84,
+ 73,
+ 111
+ ],
+ "category_id": 6,
+ "area": 67858,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18742,
+ "image_id": 3438,
+ "bbox": [
+ 210,
+ 291,
+ 86,
+ 127
+ ],
+ "category_id": 6,
+ "area": 91195,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18744,
+ "image_id": 3438,
+ "bbox": [
+ 313,
+ 197,
+ 144,
+ 228
+ ],
+ "category_id": 6,
+ "area": 273408,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18748,
+ "image_id": 3439,
+ "bbox": [
+ 25,
+ 51,
+ 312,
+ 395
+ ],
+ "category_id": 9,
+ "area": 169824,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18749,
+ "image_id": 3440,
+ "bbox": [
+ 57,
+ 188,
+ 202,
+ 182
+ ],
+ "category_id": 8,
+ "area": 128775,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18750,
+ "image_id": 3440,
+ "bbox": [
+ 241,
+ 179,
+ 198,
+ 214
+ ],
+ "category_id": 8,
+ "area": 149296,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18752,
+ "image_id": 3441,
+ "bbox": [
+ 0,
+ 264,
+ 81,
+ 106
+ ],
+ "category_id": 9,
+ "area": 12276,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18753,
+ "image_id": 3441,
+ "bbox": [
+ 43,
+ 113,
+ 391,
+ 341
+ ],
+ "category_id": 9,
+ "area": 190124,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18754,
+ "image_id": 3441,
+ "bbox": [
+ 225,
+ 267,
+ 68,
+ 43
+ ],
+ "category_id": 9,
+ "area": 4264,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18761,
+ "image_id": 3445,
+ "bbox": [
+ 68,
+ 51,
+ 348,
+ 370
+ ],
+ "category_id": 8,
+ "area": 832617,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18762,
+ "image_id": 3445,
+ "bbox": [
+ 132,
+ 188,
+ 330,
+ 321
+ ],
+ "category_id": 8,
+ "area": 684061,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18763,
+ "image_id": 3445,
+ "bbox": [
+ 408,
+ 112,
+ 78,
+ 85
+ ],
+ "category_id": 8,
+ "area": 43365,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18788,
+ "image_id": 3450,
+ "bbox": [
+ 2,
+ 49,
+ 488,
+ 421
+ ],
+ "category_id": 9,
+ "area": 723460,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18789,
+ "image_id": 3451,
+ "bbox": [
+ 141,
+ 166,
+ 221,
+ 344
+ ],
+ "category_id": 10,
+ "area": 604864,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18790,
+ "image_id": 3452,
+ "bbox": [
+ 220,
+ 137,
+ 33,
+ 126
+ ],
+ "category_id": 10,
+ "area": 10230,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18791,
+ "image_id": 3452,
+ "bbox": [
+ 218,
+ 154,
+ 274,
+ 299
+ ],
+ "category_id": 9,
+ "area": 195000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18794,
+ "image_id": 3453,
+ "bbox": [
+ 80,
+ 85,
+ 429,
+ 426
+ ],
+ "category_id": 9,
+ "area": 442746,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18795,
+ "image_id": 3453,
+ "bbox": [
+ 42,
+ 0,
+ 133,
+ 494
+ ],
+ "category_id": 6,
+ "area": 158766,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18796,
+ "image_id": 3453,
+ "bbox": [
+ 0,
+ 0,
+ 509,
+ 512
+ ],
+ "category_id": 6,
+ "area": 629057,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18797,
+ "image_id": 3454,
+ "bbox": [
+ 190,
+ 0,
+ 226,
+ 148
+ ],
+ "category_id": 9,
+ "area": 96960,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18798,
+ "image_id": 3454,
+ "bbox": [
+ 198,
+ 69,
+ 117,
+ 185
+ ],
+ "category_id": 10,
+ "area": 62379,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18800,
+ "image_id": 3455,
+ "bbox": [
+ 0,
+ 54,
+ 268,
+ 456
+ ],
+ "category_id": 9,
+ "area": 430782,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18801,
+ "image_id": 3455,
+ "bbox": [
+ 176,
+ 0,
+ 332,
+ 505
+ ],
+ "category_id": 9,
+ "area": 590841,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18802,
+ "image_id": 3455,
+ "bbox": [
+ 84,
+ 0,
+ 295,
+ 320
+ ],
+ "category_id": 9,
+ "area": 332550,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18810,
+ "image_id": 3456,
+ "bbox": [
+ 102,
+ 199,
+ 348,
+ 238
+ ],
+ "category_id": 9,
+ "area": 140937,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18814,
+ "image_id": 3458,
+ "bbox": [
+ 146,
+ 86,
+ 226,
+ 256
+ ],
+ "category_id": 8,
+ "area": 171310,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18815,
+ "image_id": 3459,
+ "bbox": [
+ 138,
+ 247,
+ 248,
+ 207
+ ],
+ "category_id": 8,
+ "area": 180711,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18816,
+ "image_id": 3459,
+ "bbox": [
+ 230,
+ 135,
+ 229,
+ 175
+ ],
+ "category_id": 8,
+ "area": 141204,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18817,
+ "image_id": 3459,
+ "bbox": [
+ 0,
+ 74,
+ 233,
+ 434
+ ],
+ "category_id": 6,
+ "area": 355047,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18820,
+ "image_id": 3460,
+ "bbox": [
+ 175,
+ 125,
+ 72,
+ 125
+ ],
+ "category_id": 6,
+ "area": 31856,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18821,
+ "image_id": 3460,
+ "bbox": [
+ 241,
+ 194,
+ 44,
+ 66
+ ],
+ "category_id": 6,
+ "area": 10416,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18822,
+ "image_id": 3460,
+ "bbox": [
+ 234,
+ 270,
+ 110,
+ 187
+ ],
+ "category_id": 6,
+ "area": 72588,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18849,
+ "image_id": 3460,
+ "bbox": [
+ 217,
+ 125,
+ 42,
+ 60
+ ],
+ "category_id": 6,
+ "area": 9010,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18850,
+ "image_id": 3460,
+ "bbox": [
+ 241,
+ 192,
+ 93,
+ 75
+ ],
+ "category_id": 6,
+ "area": 24804,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18851,
+ "image_id": 3461,
+ "bbox": [
+ 79,
+ 17,
+ 198,
+ 397
+ ],
+ "category_id": 9,
+ "area": 255411,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18852,
+ "image_id": 3461,
+ "bbox": [
+ 137,
+ 232,
+ 128,
+ 222
+ ],
+ "category_id": 10,
+ "area": 92664,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18854,
+ "image_id": 3462,
+ "bbox": [
+ 72,
+ 322,
+ 224,
+ 185
+ ],
+ "category_id": 10,
+ "area": 65637,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18856,
+ "image_id": 3463,
+ "bbox": [
+ 77,
+ 145,
+ 289,
+ 366
+ ],
+ "category_id": 9,
+ "area": 303208,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18859,
+ "image_id": 3464,
+ "bbox": [
+ 219,
+ 155,
+ 72,
+ 130
+ ],
+ "category_id": 10,
+ "area": 22440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18863,
+ "image_id": 3464,
+ "bbox": [
+ 221,
+ 149,
+ 230,
+ 362
+ ],
+ "category_id": 9,
+ "area": 197820,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18864,
+ "image_id": 3465,
+ "bbox": [
+ 14,
+ 68,
+ 363,
+ 441
+ ],
+ "category_id": 8,
+ "area": 564489,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18865,
+ "image_id": 3465,
+ "bbox": [
+ 137,
+ 1,
+ 298,
+ 354
+ ],
+ "category_id": 8,
+ "area": 372254,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18866,
+ "image_id": 3466,
+ "bbox": [
+ 50,
+ 147,
+ 96,
+ 86
+ ],
+ "category_id": 6,
+ "area": 29280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18867,
+ "image_id": 3466,
+ "bbox": [
+ 57,
+ 199,
+ 97,
+ 115
+ ],
+ "category_id": 6,
+ "area": 39366,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18868,
+ "image_id": 3466,
+ "bbox": [
+ 66,
+ 312,
+ 51,
+ 69
+ ],
+ "category_id": 6,
+ "area": 12544,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18870,
+ "image_id": 3466,
+ "bbox": [
+ 184,
+ 286,
+ 41,
+ 49
+ ],
+ "category_id": 6,
+ "area": 7210,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18871,
+ "image_id": 3466,
+ "bbox": [
+ 180,
+ 418,
+ 117,
+ 91
+ ],
+ "category_id": 6,
+ "area": 37632,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18872,
+ "image_id": 3466,
+ "bbox": [
+ 281,
+ 381,
+ 127,
+ 130
+ ],
+ "category_id": 6,
+ "area": 58377,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18873,
+ "image_id": 3466,
+ "bbox": [
+ 368,
+ 344,
+ 144,
+ 166
+ ],
+ "category_id": 6,
+ "area": 84240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18874,
+ "image_id": 3466,
+ "bbox": [
+ 452,
+ 310,
+ 59,
+ 78
+ ],
+ "category_id": 6,
+ "area": 16390,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18875,
+ "image_id": 3466,
+ "bbox": [
+ 240,
+ 170,
+ 135,
+ 120
+ ],
+ "category_id": 8,
+ "area": 57630,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18877,
+ "image_id": 3466,
+ "bbox": [
+ 266,
+ 342,
+ 92,
+ 92
+ ],
+ "category_id": 6,
+ "area": 30030,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18883,
+ "image_id": 3467,
+ "bbox": [
+ 274,
+ 24,
+ 92,
+ 149
+ ],
+ "category_id": 6,
+ "area": 16240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18884,
+ "image_id": 3467,
+ "bbox": [
+ 429,
+ 119,
+ 81,
+ 193
+ ],
+ "category_id": 6,
+ "area": 18462,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18885,
+ "image_id": 3467,
+ "bbox": [
+ 99,
+ 196,
+ 181,
+ 221
+ ],
+ "category_id": 6,
+ "area": 47216,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18886,
+ "image_id": 3467,
+ "bbox": [
+ 0,
+ 0,
+ 184,
+ 156
+ ],
+ "category_id": 6,
+ "area": 33957,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18887,
+ "image_id": 3467,
+ "bbox": [
+ 0,
+ 421,
+ 56,
+ 89
+ ],
+ "category_id": 6,
+ "area": 5964,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18888,
+ "image_id": 3467,
+ "bbox": [
+ 56,
+ 434,
+ 118,
+ 75
+ ],
+ "category_id": 6,
+ "area": 10508,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18889,
+ "image_id": 3467,
+ "bbox": [
+ 252,
+ 424,
+ 76,
+ 86
+ ],
+ "category_id": 6,
+ "area": 7776,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18897,
+ "image_id": 3469,
+ "bbox": [
+ 104,
+ 0,
+ 406,
+ 512
+ ],
+ "category_id": 10,
+ "area": 203196,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18901,
+ "image_id": 3471,
+ "bbox": [
+ 90,
+ 0,
+ 201,
+ 423
+ ],
+ "category_id": 6,
+ "area": 271542,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18902,
+ "image_id": 3471,
+ "bbox": [
+ 365,
+ 0,
+ 88,
+ 97
+ ],
+ "category_id": 6,
+ "area": 27375,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18920,
+ "image_id": 3475,
+ "bbox": [
+ 103,
+ 119,
+ 407,
+ 392
+ ],
+ "category_id": 8,
+ "area": 1101555,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18921,
+ "image_id": 3475,
+ "bbox": [
+ 262,
+ 113,
+ 114,
+ 217
+ ],
+ "category_id": 8,
+ "area": 171136,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18922,
+ "image_id": 3476,
+ "bbox": [
+ 0,
+ 211,
+ 507,
+ 210
+ ],
+ "category_id": 9,
+ "area": 175525,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18924,
+ "image_id": 3478,
+ "bbox": [
+ 6,
+ 0,
+ 311,
+ 270
+ ],
+ "category_id": 9,
+ "area": 665190,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18926,
+ "image_id": 3478,
+ "bbox": [
+ 49,
+ 343,
+ 200,
+ 168
+ ],
+ "category_id": 6,
+ "area": 267356,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18927,
+ "image_id": 3478,
+ "bbox": [
+ 204,
+ 260,
+ 307,
+ 251
+ ],
+ "category_id": 6,
+ "area": 611712,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18963,
+ "image_id": 3483,
+ "bbox": [
+ 147,
+ 132,
+ 155,
+ 70
+ ],
+ "category_id": 9,
+ "area": 11115,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 18994,
+ "image_id": 3487,
+ "bbox": [
+ 172,
+ 164,
+ 270,
+ 300
+ ],
+ "category_id": 9,
+ "area": 151294,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19036,
+ "image_id": 3498,
+ "bbox": [
+ 173,
+ 187,
+ 180,
+ 117
+ ],
+ "category_id": 8,
+ "area": 34185,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19037,
+ "image_id": 3498,
+ "bbox": [
+ 0,
+ 360,
+ 483,
+ 148
+ ],
+ "category_id": 6,
+ "area": 115404,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19038,
+ "image_id": 3498,
+ "bbox": [
+ 410,
+ 386,
+ 97,
+ 105
+ ],
+ "category_id": 6,
+ "area": 16588,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19039,
+ "image_id": 3498,
+ "bbox": [
+ 135,
+ 94,
+ 207,
+ 109
+ ],
+ "category_id": 6,
+ "area": 36480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19040,
+ "image_id": 3498,
+ "bbox": [
+ 47,
+ 148,
+ 129,
+ 190
+ ],
+ "category_id": 6,
+ "area": 39690,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19041,
+ "image_id": 3498,
+ "bbox": [
+ 157,
+ 312,
+ 105,
+ 51
+ ],
+ "category_id": 6,
+ "area": 8835,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19042,
+ "image_id": 3498,
+ "bbox": [
+ 283,
+ 294,
+ 106,
+ 68
+ ],
+ "category_id": 6,
+ "area": 11700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19045,
+ "image_id": 3500,
+ "bbox": [
+ 224,
+ 102,
+ 246,
+ 133
+ ],
+ "category_id": 9,
+ "area": 62292,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19046,
+ "image_id": 3500,
+ "bbox": [
+ 280,
+ 194,
+ 127,
+ 108
+ ],
+ "category_id": 10,
+ "area": 26085,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19048,
+ "image_id": 3501,
+ "bbox": [
+ 127,
+ 120,
+ 275,
+ 177
+ ],
+ "category_id": 8,
+ "area": 170624,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19049,
+ "image_id": 3501,
+ "bbox": [
+ 180,
+ 262,
+ 240,
+ 203
+ ],
+ "category_id": 8,
+ "area": 171570,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19050,
+ "image_id": 3501,
+ "bbox": [
+ 130,
+ 433,
+ 51,
+ 36
+ ],
+ "category_id": 8,
+ "area": 6528,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19051,
+ "image_id": 3501,
+ "bbox": [
+ 258,
+ 420,
+ 45,
+ 39
+ ],
+ "category_id": 8,
+ "area": 6384,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19052,
+ "image_id": 3501,
+ "bbox": [
+ 337,
+ 475,
+ 50,
+ 31
+ ],
+ "category_id": 8,
+ "area": 5588,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19056,
+ "image_id": 3503,
+ "bbox": [
+ 66,
+ 43,
+ 116,
+ 109
+ ],
+ "category_id": 8,
+ "area": 97875,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19057,
+ "image_id": 3503,
+ "bbox": [
+ 0,
+ 34,
+ 512,
+ 288
+ ],
+ "category_id": 8,
+ "area": 1140020,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19058,
+ "image_id": 3503,
+ "bbox": [
+ 118,
+ 234,
+ 289,
+ 225
+ ],
+ "category_id": 8,
+ "area": 502048,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19065,
+ "image_id": 3504,
+ "bbox": [
+ 368,
+ 115,
+ 39,
+ 82
+ ],
+ "category_id": 6,
+ "area": 25181,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19066,
+ "image_id": 3504,
+ "bbox": [
+ 267,
+ 174,
+ 44,
+ 74
+ ],
+ "category_id": 6,
+ "area": 25718,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19067,
+ "image_id": 3504,
+ "bbox": [
+ 285,
+ 130,
+ 28,
+ 37
+ ],
+ "category_id": 6,
+ "area": 8239,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19068,
+ "image_id": 3504,
+ "bbox": [
+ 262,
+ 119,
+ 24,
+ 46
+ ],
+ "category_id": 6,
+ "area": 8640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19069,
+ "image_id": 3504,
+ "bbox": [
+ 76,
+ 168,
+ 36,
+ 62
+ ],
+ "category_id": 6,
+ "area": 17415,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19070,
+ "image_id": 3504,
+ "bbox": [
+ 93,
+ 122,
+ 64,
+ 102
+ ],
+ "category_id": 6,
+ "area": 51516,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19071,
+ "image_id": 3504,
+ "bbox": [
+ 0,
+ 294,
+ 109,
+ 87
+ ],
+ "category_id": 6,
+ "area": 73980,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19072,
+ "image_id": 3504,
+ "bbox": [
+ 106,
+ 441,
+ 35,
+ 60
+ ],
+ "category_id": 6,
+ "area": 16750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19073,
+ "image_id": 3504,
+ "bbox": [
+ 356,
+ 210,
+ 22,
+ 42
+ ],
+ "category_id": 6,
+ "area": 7308,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19074,
+ "image_id": 3504,
+ "bbox": [
+ 6,
+ 132,
+ 35,
+ 60
+ ],
+ "category_id": 6,
+ "area": 16368,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19075,
+ "image_id": 3504,
+ "bbox": [
+ 57,
+ 79,
+ 36,
+ 71
+ ],
+ "category_id": 6,
+ "area": 19845,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19076,
+ "image_id": 3504,
+ "bbox": [
+ 35,
+ 51,
+ 32,
+ 50
+ ],
+ "category_id": 6,
+ "area": 12480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19077,
+ "image_id": 3504,
+ "bbox": [
+ 18,
+ 446,
+ 35,
+ 65
+ ],
+ "category_id": 6,
+ "area": 17955,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19078,
+ "image_id": 3504,
+ "bbox": [
+ 351,
+ 59,
+ 13,
+ 42
+ ],
+ "category_id": 6,
+ "area": 4437,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19079,
+ "image_id": 3504,
+ "bbox": [
+ 371,
+ 2,
+ 20,
+ 50
+ ],
+ "category_id": 6,
+ "area": 8190,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19080,
+ "image_id": 3505,
+ "bbox": [
+ 239,
+ 131,
+ 171,
+ 357
+ ],
+ "category_id": 9,
+ "area": 2895600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19081,
+ "image_id": 3505,
+ "bbox": [
+ 137,
+ 196,
+ 181,
+ 272
+ ],
+ "category_id": 10,
+ "area": 2330584,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19116,
+ "image_id": 3507,
+ "bbox": [
+ 0,
+ 62,
+ 467,
+ 449
+ ],
+ "category_id": 6,
+ "area": 1946496,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19118,
+ "image_id": 3508,
+ "bbox": [
+ 38,
+ 21,
+ 386,
+ 390
+ ],
+ "category_id": 8,
+ "area": 218960,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19119,
+ "image_id": 3508,
+ "bbox": [
+ 126,
+ 78,
+ 385,
+ 314
+ ],
+ "category_id": 8,
+ "area": 176182,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19120,
+ "image_id": 3509,
+ "bbox": [
+ 0,
+ 24,
+ 121,
+ 153
+ ],
+ "category_id": 6,
+ "area": 40454,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19121,
+ "image_id": 3509,
+ "bbox": [
+ 144,
+ 164,
+ 328,
+ 347
+ ],
+ "category_id": 9,
+ "area": 247860,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19122,
+ "image_id": 3510,
+ "bbox": [
+ 383,
+ 168,
+ 36,
+ 57
+ ],
+ "category_id": 6,
+ "area": 16714,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19123,
+ "image_id": 3510,
+ "bbox": [
+ 199,
+ 203,
+ 67,
+ 35
+ ],
+ "category_id": 8,
+ "area": 18648,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19124,
+ "image_id": 3510,
+ "bbox": [
+ 153,
+ 230,
+ 119,
+ 117
+ ],
+ "category_id": 8,
+ "area": 110656,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19125,
+ "image_id": 3510,
+ "bbox": [
+ 154,
+ 268,
+ 132,
+ 63
+ ],
+ "category_id": 8,
+ "area": 66234,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19126,
+ "image_id": 3511,
+ "bbox": [
+ 82,
+ 90,
+ 295,
+ 237
+ ],
+ "category_id": 8,
+ "area": 246492,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19127,
+ "image_id": 3511,
+ "bbox": [
+ 113,
+ 182,
+ 380,
+ 322
+ ],
+ "category_id": 8,
+ "area": 431300,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19156,
+ "image_id": 3514,
+ "bbox": [
+ 14,
+ 234,
+ 403,
+ 276
+ ],
+ "category_id": 9,
+ "area": 225594,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19158,
+ "image_id": 3514,
+ "bbox": [
+ 434,
+ 32,
+ 76,
+ 177
+ ],
+ "category_id": 6,
+ "area": 27548,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19169,
+ "image_id": 3516,
+ "bbox": [
+ 65,
+ 355,
+ 248,
+ 94
+ ],
+ "category_id": 10,
+ "area": 37177,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19171,
+ "image_id": 3517,
+ "bbox": [
+ 67,
+ 0,
+ 306,
+ 317
+ ],
+ "category_id": 9,
+ "area": 195455,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19172,
+ "image_id": 3517,
+ "bbox": [
+ 97,
+ 33,
+ 199,
+ 456
+ ],
+ "category_id": 9,
+ "area": 182700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19176,
+ "image_id": 3518,
+ "bbox": [
+ 63,
+ 0,
+ 314,
+ 221
+ ],
+ "category_id": 6,
+ "area": 94976,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19177,
+ "image_id": 3518,
+ "bbox": [
+ 210,
+ 357,
+ 191,
+ 151
+ ],
+ "category_id": 6,
+ "area": 39474,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19178,
+ "image_id": 3519,
+ "bbox": [
+ 70,
+ 143,
+ 311,
+ 198
+ ],
+ "category_id": 9,
+ "area": 158790,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19180,
+ "image_id": 3519,
+ "bbox": [
+ 410,
+ 333,
+ 101,
+ 175
+ ],
+ "category_id": 6,
+ "area": 45990,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19215,
+ "image_id": 3522,
+ "bbox": [
+ 0,
+ 113,
+ 56,
+ 107
+ ],
+ "category_id": 6,
+ "area": 52326,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19216,
+ "image_id": 3522,
+ "bbox": [
+ 0,
+ 3,
+ 119,
+ 122
+ ],
+ "category_id": 6,
+ "area": 126324,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19217,
+ "image_id": 3522,
+ "bbox": [
+ 92,
+ 145,
+ 75,
+ 117
+ ],
+ "category_id": 6,
+ "area": 77280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19218,
+ "image_id": 3522,
+ "bbox": [
+ 192,
+ 340,
+ 100,
+ 101
+ ],
+ "category_id": 6,
+ "area": 88160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19219,
+ "image_id": 3522,
+ "bbox": [
+ 59,
+ 106,
+ 72,
+ 56
+ ],
+ "category_id": 6,
+ "area": 35802,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19220,
+ "image_id": 3522,
+ "bbox": [
+ 178,
+ 21,
+ 122,
+ 127
+ ],
+ "category_id": 6,
+ "area": 135036,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19221,
+ "image_id": 3522,
+ "bbox": [
+ 323,
+ 286,
+ 156,
+ 167
+ ],
+ "category_id": 6,
+ "area": 225621,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19222,
+ "image_id": 3522,
+ "bbox": [
+ 253,
+ 0,
+ 95,
+ 108
+ ],
+ "category_id": 6,
+ "area": 89590,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19223,
+ "image_id": 3522,
+ "bbox": [
+ 386,
+ 88,
+ 88,
+ 68
+ ],
+ "category_id": 6,
+ "area": 51992,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19245,
+ "image_id": 3525,
+ "bbox": [
+ 59,
+ 178,
+ 396,
+ 251
+ ],
+ "category_id": 8,
+ "area": 351168,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19246,
+ "image_id": 3525,
+ "bbox": [
+ 0,
+ 219,
+ 57,
+ 154
+ ],
+ "category_id": 6,
+ "area": 31248,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19247,
+ "image_id": 3525,
+ "bbox": [
+ 222,
+ 35,
+ 69,
+ 87
+ ],
+ "category_id": 6,
+ "area": 21279,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19248,
+ "image_id": 3525,
+ "bbox": [
+ 281,
+ 2,
+ 178,
+ 161
+ ],
+ "category_id": 6,
+ "area": 101469,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19251,
+ "image_id": 3525,
+ "bbox": [
+ 330,
+ 193,
+ 135,
+ 145
+ ],
+ "category_id": 6,
+ "area": 69156,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19252,
+ "image_id": 3525,
+ "bbox": [
+ 86,
+ 85,
+ 29,
+ 33
+ ],
+ "category_id": 6,
+ "area": 3431,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19263,
+ "image_id": 3529,
+ "bbox": [
+ 173,
+ 176,
+ 224,
+ 266
+ ],
+ "category_id": 9,
+ "area": 42705,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19264,
+ "image_id": 3529,
+ "bbox": [
+ 323,
+ 76,
+ 123,
+ 92
+ ],
+ "category_id": 9,
+ "area": 8228,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19267,
+ "image_id": 3530,
+ "bbox": [
+ 268,
+ 195,
+ 242,
+ 315
+ ],
+ "category_id": 9,
+ "area": 174741,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19268,
+ "image_id": 3530,
+ "bbox": [
+ 13,
+ 86,
+ 55,
+ 73
+ ],
+ "category_id": 6,
+ "area": 9288,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19269,
+ "image_id": 3530,
+ "bbox": [
+ 0,
+ 331,
+ 156,
+ 132
+ ],
+ "category_id": 6,
+ "area": 47424,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19270,
+ "image_id": 3530,
+ "bbox": [
+ 130,
+ 0,
+ 64,
+ 62
+ ],
+ "category_id": 6,
+ "area": 9324,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19289,
+ "image_id": 3533,
+ "bbox": [
+ 4,
+ 294,
+ 310,
+ 212
+ ],
+ "category_id": 6,
+ "area": 187664,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19290,
+ "image_id": 3534,
+ "bbox": [
+ 3,
+ 188,
+ 504,
+ 293
+ ],
+ "category_id": 9,
+ "area": 304956,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19314,
+ "image_id": 3539,
+ "bbox": [
+ 0,
+ 1,
+ 135,
+ 130
+ ],
+ "category_id": 6,
+ "area": 283404,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19315,
+ "image_id": 3539,
+ "bbox": [
+ 177,
+ 0,
+ 74,
+ 60
+ ],
+ "category_id": 6,
+ "area": 72105,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19316,
+ "image_id": 3539,
+ "bbox": [
+ 115,
+ 110,
+ 100,
+ 121
+ ],
+ "category_id": 6,
+ "area": 195300,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19317,
+ "image_id": 3539,
+ "bbox": [
+ 120,
+ 233,
+ 115,
+ 92
+ ],
+ "category_id": 6,
+ "area": 171840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19318,
+ "image_id": 3539,
+ "bbox": [
+ 79,
+ 326,
+ 111,
+ 103
+ ],
+ "category_id": 6,
+ "area": 184926,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19319,
+ "image_id": 3539,
+ "bbox": [
+ 1,
+ 380,
+ 35,
+ 93
+ ],
+ "category_id": 6,
+ "area": 52649,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19320,
+ "image_id": 3539,
+ "bbox": [
+ 229,
+ 0,
+ 99,
+ 194
+ ],
+ "category_id": 6,
+ "area": 310673,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19321,
+ "image_id": 3539,
+ "bbox": [
+ 227,
+ 182,
+ 84,
+ 76
+ ],
+ "category_id": 6,
+ "area": 102833,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19322,
+ "image_id": 3539,
+ "bbox": [
+ 244,
+ 255,
+ 132,
+ 85
+ ],
+ "category_id": 6,
+ "area": 182336,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19323,
+ "image_id": 3539,
+ "bbox": [
+ 329,
+ 24,
+ 30,
+ 83
+ ],
+ "category_id": 6,
+ "area": 41184,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19324,
+ "image_id": 3539,
+ "bbox": [
+ 342,
+ 74,
+ 67,
+ 70
+ ],
+ "category_id": 6,
+ "area": 76372,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19325,
+ "image_id": 3539,
+ "bbox": [
+ 412,
+ 66,
+ 95,
+ 100
+ ],
+ "category_id": 6,
+ "area": 154164,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19326,
+ "image_id": 3539,
+ "bbox": [
+ 444,
+ 258,
+ 67,
+ 101
+ ],
+ "category_id": 6,
+ "area": 109472,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19332,
+ "image_id": 3540,
+ "bbox": [
+ 1,
+ 55,
+ 507,
+ 452
+ ],
+ "category_id": 9,
+ "area": 808353,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19335,
+ "image_id": 3542,
+ "bbox": [
+ 132,
+ 308,
+ 343,
+ 203
+ ],
+ "category_id": 9,
+ "area": 123280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19353,
+ "image_id": 3547,
+ "bbox": [
+ 111,
+ 415,
+ 172,
+ 96
+ ],
+ "category_id": 9,
+ "area": 19824,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19358,
+ "image_id": 3548,
+ "bbox": [
+ 210,
+ 76,
+ 167,
+ 129
+ ],
+ "category_id": 6,
+ "area": 30996,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19359,
+ "image_id": 3548,
+ "bbox": [
+ 296,
+ 173,
+ 83,
+ 88
+ ],
+ "category_id": 6,
+ "area": 10578,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19360,
+ "image_id": 3549,
+ "bbox": [
+ 73,
+ 30,
+ 435,
+ 478
+ ],
+ "category_id": 9,
+ "area": 624507,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19367,
+ "image_id": 3552,
+ "bbox": [
+ 182,
+ 259,
+ 162,
+ 221
+ ],
+ "category_id": 8,
+ "area": 115140,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19368,
+ "image_id": 3552,
+ "bbox": [
+ 230,
+ 175,
+ 45,
+ 94
+ ],
+ "category_id": 8,
+ "area": 13908,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19373,
+ "image_id": 3554,
+ "bbox": [
+ 115,
+ 252,
+ 132,
+ 148
+ ],
+ "category_id": 8,
+ "area": 29097,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19375,
+ "image_id": 3554,
+ "bbox": [
+ 113,
+ 100,
+ 57,
+ 32
+ ],
+ "category_id": 6,
+ "area": 2765,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19376,
+ "image_id": 3554,
+ "bbox": [
+ 3,
+ 126,
+ 47,
+ 58
+ ],
+ "category_id": 6,
+ "area": 4095,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19377,
+ "image_id": 3554,
+ "bbox": [
+ 1,
+ 277,
+ 47,
+ 81
+ ],
+ "category_id": 6,
+ "area": 5720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19378,
+ "image_id": 3554,
+ "bbox": [
+ 137,
+ 116,
+ 92,
+ 34
+ ],
+ "category_id": 6,
+ "area": 4736,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19379,
+ "image_id": 3554,
+ "bbox": [
+ 81,
+ 143,
+ 32,
+ 34
+ ],
+ "category_id": 6,
+ "area": 1665,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19380,
+ "image_id": 3554,
+ "bbox": [
+ 287,
+ 180,
+ 224,
+ 124
+ ],
+ "category_id": 6,
+ "area": 41406,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19381,
+ "image_id": 3554,
+ "bbox": [
+ 225,
+ 258,
+ 202,
+ 139
+ ],
+ "category_id": 6,
+ "area": 41850,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19382,
+ "image_id": 3554,
+ "bbox": [
+ 178,
+ 373,
+ 229,
+ 134
+ ],
+ "category_id": 6,
+ "area": 45504,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19383,
+ "image_id": 3554,
+ "bbox": [
+ 23,
+ 336,
+ 81,
+ 68
+ ],
+ "category_id": 6,
+ "area": 8288,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19384,
+ "image_id": 3554,
+ "bbox": [
+ 0,
+ 396,
+ 63,
+ 66
+ ],
+ "category_id": 6,
+ "area": 6177,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19385,
+ "image_id": 3554,
+ "bbox": [
+ 422,
+ 285,
+ 84,
+ 134
+ ],
+ "category_id": 6,
+ "area": 16965,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19389,
+ "image_id": 3555,
+ "bbox": [
+ 105,
+ 104,
+ 378,
+ 397
+ ],
+ "category_id": 8,
+ "area": 198436,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19412,
+ "image_id": 3561,
+ "bbox": [
+ 193,
+ 139,
+ 299,
+ 199
+ ],
+ "category_id": 9,
+ "area": 157884,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19413,
+ "image_id": 3561,
+ "bbox": [
+ 12,
+ 309,
+ 157,
+ 202
+ ],
+ "category_id": 6,
+ "area": 83889,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19427,
+ "image_id": 3566,
+ "bbox": [
+ 200,
+ 197,
+ 150,
+ 205
+ ],
+ "category_id": 6,
+ "area": 44421,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19428,
+ "image_id": 3566,
+ "bbox": [
+ 213,
+ 382,
+ 221,
+ 129
+ ],
+ "category_id": 9,
+ "area": 41076,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19434,
+ "image_id": 3568,
+ "bbox": [
+ 2,
+ 97,
+ 64,
+ 97
+ ],
+ "category_id": 6,
+ "area": 104228,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19435,
+ "image_id": 3568,
+ "bbox": [
+ 32,
+ 19,
+ 84,
+ 58
+ ],
+ "category_id": 6,
+ "area": 81468,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19436,
+ "image_id": 3568,
+ "bbox": [
+ 74,
+ 7,
+ 204,
+ 136
+ ],
+ "category_id": 6,
+ "area": 462086,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19437,
+ "image_id": 3568,
+ "bbox": [
+ 213,
+ 5,
+ 67,
+ 54
+ ],
+ "category_id": 6,
+ "area": 60180,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19438,
+ "image_id": 3568,
+ "bbox": [
+ 286,
+ 9,
+ 93,
+ 74
+ ],
+ "category_id": 6,
+ "area": 115080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19439,
+ "image_id": 3568,
+ "bbox": [
+ 391,
+ 20,
+ 56,
+ 62
+ ],
+ "category_id": 6,
+ "area": 57798,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19440,
+ "image_id": 3568,
+ "bbox": [
+ 390,
+ 77,
+ 75,
+ 120
+ ],
+ "category_id": 6,
+ "area": 149160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19441,
+ "image_id": 3568,
+ "bbox": [
+ 358,
+ 183,
+ 86,
+ 77
+ ],
+ "category_id": 6,
+ "area": 110580,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19442,
+ "image_id": 3568,
+ "bbox": [
+ 462,
+ 41,
+ 31,
+ 71
+ ],
+ "category_id": 6,
+ "area": 37800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19443,
+ "image_id": 3568,
+ "bbox": [
+ 480,
+ 87,
+ 31,
+ 58
+ ],
+ "category_id": 6,
+ "area": 30580,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19444,
+ "image_id": 3568,
+ "bbox": [
+ 244,
+ 222,
+ 119,
+ 99
+ ],
+ "category_id": 6,
+ "area": 194556,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19445,
+ "image_id": 3568,
+ "bbox": [
+ 201,
+ 326,
+ 111,
+ 81
+ ],
+ "category_id": 6,
+ "area": 148840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19446,
+ "image_id": 3568,
+ "bbox": [
+ 373,
+ 253,
+ 132,
+ 88
+ ],
+ "category_id": 6,
+ "area": 193556,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19447,
+ "image_id": 3568,
+ "bbox": [
+ 57,
+ 333,
+ 96,
+ 131
+ ],
+ "category_id": 6,
+ "area": 207624,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19451,
+ "image_id": 3569,
+ "bbox": [
+ 0,
+ 43,
+ 44,
+ 101
+ ],
+ "category_id": 8,
+ "area": 30710,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19452,
+ "image_id": 3569,
+ "bbox": [
+ 88,
+ 164,
+ 105,
+ 167
+ ],
+ "category_id": 8,
+ "area": 120080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19453,
+ "image_id": 3569,
+ "bbox": [
+ 155,
+ 161,
+ 90,
+ 199
+ ],
+ "category_id": 8,
+ "area": 123057,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19454,
+ "image_id": 3569,
+ "bbox": [
+ 268,
+ 64,
+ 34,
+ 52
+ ],
+ "category_id": 8,
+ "area": 12288,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19455,
+ "image_id": 3569,
+ "bbox": [
+ 313,
+ 164,
+ 94,
+ 157
+ ],
+ "category_id": 8,
+ "area": 100958,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19456,
+ "image_id": 3569,
+ "bbox": [
+ 406,
+ 169,
+ 39,
+ 68
+ ],
+ "category_id": 8,
+ "area": 18250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19457,
+ "image_id": 3570,
+ "bbox": [
+ 114,
+ 0,
+ 378,
+ 281
+ ],
+ "category_id": 8,
+ "area": 152070,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19458,
+ "image_id": 3570,
+ "bbox": [
+ 0,
+ 152,
+ 321,
+ 356
+ ],
+ "category_id": 8,
+ "area": 163437,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19459,
+ "image_id": 3571,
+ "bbox": [
+ 150,
+ 56,
+ 194,
+ 389
+ ],
+ "category_id": 10,
+ "area": 599330,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19460,
+ "image_id": 3572,
+ "bbox": [
+ 0,
+ 114,
+ 286,
+ 162
+ ],
+ "category_id": 9,
+ "area": 78320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19461,
+ "image_id": 3572,
+ "bbox": [
+ 0,
+ 209,
+ 190,
+ 301
+ ],
+ "category_id": 6,
+ "area": 96465,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19462,
+ "image_id": 3573,
+ "bbox": [
+ 170,
+ 133,
+ 29,
+ 59
+ ],
+ "category_id": 10,
+ "area": 5096,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19464,
+ "image_id": 3573,
+ "bbox": [
+ 107,
+ 123,
+ 221,
+ 388
+ ],
+ "category_id": 9,
+ "area": 246400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19476,
+ "image_id": 3575,
+ "bbox": [
+ 0,
+ 1,
+ 512,
+ 252
+ ],
+ "category_id": 6,
+ "area": 320133,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19494,
+ "image_id": 3577,
+ "bbox": [
+ 3,
+ 76,
+ 507,
+ 397
+ ],
+ "category_id": 9,
+ "area": 234244,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19510,
+ "image_id": 3580,
+ "bbox": [
+ 56,
+ 56,
+ 130,
+ 261
+ ],
+ "category_id": 8,
+ "area": 120336,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19511,
+ "image_id": 3580,
+ "bbox": [
+ 45,
+ 212,
+ 221,
+ 216
+ ],
+ "category_id": 8,
+ "area": 168112,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19512,
+ "image_id": 3580,
+ "bbox": [
+ 256,
+ 134,
+ 255,
+ 171
+ ],
+ "category_id": 8,
+ "area": 153758,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19513,
+ "image_id": 3580,
+ "bbox": [
+ 403,
+ 0,
+ 108,
+ 169
+ ],
+ "category_id": 6,
+ "area": 65008,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19514,
+ "image_id": 3580,
+ "bbox": [
+ 440,
+ 304,
+ 70,
+ 109
+ ],
+ "category_id": 6,
+ "area": 26950,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19523,
+ "image_id": 3582,
+ "bbox": [
+ 162,
+ 105,
+ 184,
+ 258
+ ],
+ "category_id": 8,
+ "area": 376595,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19524,
+ "image_id": 3582,
+ "bbox": [
+ 330,
+ 153,
+ 71,
+ 158
+ ],
+ "category_id": 8,
+ "area": 90115,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19525,
+ "image_id": 3582,
+ "bbox": [
+ 174,
+ 155,
+ 78,
+ 99
+ ],
+ "category_id": 8,
+ "area": 61237,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19535,
+ "image_id": 3585,
+ "bbox": [
+ 164,
+ 263,
+ 172,
+ 162
+ ],
+ "category_id": 8,
+ "area": 394848,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19537,
+ "image_id": 3586,
+ "bbox": [
+ 37,
+ 0,
+ 86,
+ 69
+ ],
+ "category_id": 6,
+ "area": 13137,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19538,
+ "image_id": 3586,
+ "bbox": [
+ 234,
+ 237,
+ 220,
+ 209
+ ],
+ "category_id": 9,
+ "area": 100485,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19558,
+ "image_id": 3589,
+ "bbox": [
+ 33,
+ 106,
+ 475,
+ 272
+ ],
+ "category_id": 9,
+ "area": 173016,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19564,
+ "image_id": 3591,
+ "bbox": [
+ 150,
+ 369,
+ 21,
+ 25
+ ],
+ "category_id": 6,
+ "area": 3604,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19565,
+ "image_id": 3591,
+ "bbox": [
+ 193,
+ 361,
+ 18,
+ 22
+ ],
+ "category_id": 6,
+ "area": 2880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19566,
+ "image_id": 3591,
+ "bbox": [
+ 230,
+ 333,
+ 23,
+ 42
+ ],
+ "category_id": 6,
+ "area": 6660,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19567,
+ "image_id": 3591,
+ "bbox": [
+ 255,
+ 328,
+ 18,
+ 41
+ ],
+ "category_id": 6,
+ "area": 5133,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19568,
+ "image_id": 3591,
+ "bbox": [
+ 395,
+ 449,
+ 19,
+ 34
+ ],
+ "category_id": 6,
+ "area": 4453,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19569,
+ "image_id": 3591,
+ "bbox": [
+ 437,
+ 206,
+ 16,
+ 44
+ ],
+ "category_id": 6,
+ "area": 4982,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19570,
+ "image_id": 3591,
+ "bbox": [
+ 473,
+ 129,
+ 32,
+ 38
+ ],
+ "category_id": 6,
+ "area": 8610,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19571,
+ "image_id": 3591,
+ "bbox": [
+ 485,
+ 218,
+ 19,
+ 28
+ ],
+ "category_id": 6,
+ "area": 3904,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19572,
+ "image_id": 3591,
+ "bbox": [
+ 470,
+ 181,
+ 11,
+ 24
+ ],
+ "category_id": 6,
+ "area": 1887,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19576,
+ "image_id": 3592,
+ "bbox": [
+ 40,
+ 214,
+ 204,
+ 237
+ ],
+ "category_id": 6,
+ "area": 46400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19577,
+ "image_id": 3593,
+ "bbox": [
+ 223,
+ 115,
+ 110,
+ 218
+ ],
+ "category_id": 10,
+ "area": 24112,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19578,
+ "image_id": 3593,
+ "bbox": [
+ 202,
+ 63,
+ 188,
+ 326
+ ],
+ "category_id": 9,
+ "area": 61542,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19580,
+ "image_id": 3594,
+ "bbox": [
+ 119,
+ 132,
+ 285,
+ 319
+ ],
+ "category_id": 8,
+ "area": 130416,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19616,
+ "image_id": 3601,
+ "bbox": [
+ 86,
+ 112,
+ 287,
+ 399
+ ],
+ "category_id": 9,
+ "area": 270417,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19617,
+ "image_id": 3601,
+ "bbox": [
+ 117,
+ 60,
+ 177,
+ 287
+ ],
+ "category_id": 10,
+ "area": 119691,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19621,
+ "image_id": 3603,
+ "bbox": [
+ 180,
+ 204,
+ 124,
+ 218
+ ],
+ "category_id": 6,
+ "area": 67519,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19622,
+ "image_id": 3603,
+ "bbox": [
+ 217,
+ 407,
+ 168,
+ 104
+ ],
+ "category_id": 9,
+ "area": 43440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19624,
+ "image_id": 3604,
+ "bbox": [
+ 22,
+ 179,
+ 462,
+ 278
+ ],
+ "category_id": 9,
+ "area": 298452,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19632,
+ "image_id": 3606,
+ "bbox": [
+ 245,
+ 81,
+ 88,
+ 127
+ ],
+ "category_id": 9,
+ "area": 39738,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19633,
+ "image_id": 3606,
+ "bbox": [
+ 326,
+ 99,
+ 67,
+ 69
+ ],
+ "category_id": 9,
+ "area": 16562,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19634,
+ "image_id": 3606,
+ "bbox": [
+ 357,
+ 120,
+ 108,
+ 120
+ ],
+ "category_id": 9,
+ "area": 46070,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19635,
+ "image_id": 3606,
+ "bbox": [
+ 450,
+ 164,
+ 58,
+ 60
+ ],
+ "category_id": 9,
+ "area": 12495,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19649,
+ "image_id": 3608,
+ "bbox": [
+ 145,
+ 276,
+ 252,
+ 229
+ ],
+ "category_id": 9,
+ "area": 112065,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19650,
+ "image_id": 3609,
+ "bbox": [
+ 134,
+ 91,
+ 341,
+ 372
+ ],
+ "category_id": 9,
+ "area": 227484,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19661,
+ "image_id": 3611,
+ "bbox": [
+ 0,
+ 312,
+ 260,
+ 78
+ ],
+ "category_id": 9,
+ "area": 29210,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19663,
+ "image_id": 3612,
+ "bbox": [
+ 241,
+ 277,
+ 164,
+ 116
+ ],
+ "category_id": 8,
+ "area": 66830,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19664,
+ "image_id": 3613,
+ "bbox": [
+ 86,
+ 311,
+ 164,
+ 199
+ ],
+ "category_id": 8,
+ "area": 106080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19665,
+ "image_id": 3613,
+ "bbox": [
+ 304,
+ 267,
+ 139,
+ 107
+ ],
+ "category_id": 8,
+ "area": 48580,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19666,
+ "image_id": 3614,
+ "bbox": [
+ 0,
+ 39,
+ 97,
+ 131
+ ],
+ "category_id": 6,
+ "area": 155178,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19667,
+ "image_id": 3614,
+ "bbox": [
+ 12,
+ 26,
+ 84,
+ 52
+ ],
+ "category_id": 6,
+ "area": 53835,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19668,
+ "image_id": 3614,
+ "bbox": [
+ 109,
+ 32,
+ 118,
+ 90
+ ],
+ "category_id": 6,
+ "area": 129600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19669,
+ "image_id": 3614,
+ "bbox": [
+ 42,
+ 112,
+ 143,
+ 150
+ ],
+ "category_id": 6,
+ "area": 263176,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19670,
+ "image_id": 3614,
+ "bbox": [
+ 49,
+ 258,
+ 157,
+ 97
+ ],
+ "category_id": 6,
+ "area": 186300,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19671,
+ "image_id": 3614,
+ "bbox": [
+ 1,
+ 360,
+ 147,
+ 93
+ ],
+ "category_id": 6,
+ "area": 167817,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19672,
+ "image_id": 3614,
+ "bbox": [
+ 199,
+ 42,
+ 136,
+ 184
+ ],
+ "category_id": 6,
+ "area": 306726,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19673,
+ "image_id": 3614,
+ "bbox": [
+ 196,
+ 218,
+ 112,
+ 76
+ ],
+ "category_id": 6,
+ "area": 104720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19674,
+ "image_id": 3614,
+ "bbox": [
+ 220,
+ 288,
+ 163,
+ 82
+ ],
+ "category_id": 6,
+ "area": 164080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19675,
+ "image_id": 3614,
+ "bbox": [
+ 310,
+ 2,
+ 115,
+ 108
+ ],
+ "category_id": 6,
+ "area": 153639,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19676,
+ "image_id": 3614,
+ "bbox": [
+ 328,
+ 64,
+ 45,
+ 76
+ ],
+ "category_id": 6,
+ "area": 42120,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19677,
+ "image_id": 3614,
+ "bbox": [
+ 354,
+ 113,
+ 87,
+ 67
+ ],
+ "category_id": 6,
+ "area": 71162,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19678,
+ "image_id": 3614,
+ "bbox": [
+ 443,
+ 112,
+ 65,
+ 88
+ ],
+ "category_id": 6,
+ "area": 70468,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19679,
+ "image_id": 3614,
+ "bbox": [
+ 488,
+ 305,
+ 21,
+ 67
+ ],
+ "category_id": 6,
+ "area": 18000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19685,
+ "image_id": 3616,
+ "bbox": [
+ 263,
+ 207,
+ 188,
+ 304
+ ],
+ "category_id": 9,
+ "area": 104958,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19715,
+ "image_id": 3619,
+ "bbox": [
+ 1,
+ 54,
+ 167,
+ 127
+ ],
+ "category_id": 6,
+ "area": 378759,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19716,
+ "image_id": 3619,
+ "bbox": [
+ 3,
+ 225,
+ 46,
+ 37
+ ],
+ "category_id": 6,
+ "area": 30576,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19717,
+ "image_id": 3619,
+ "bbox": [
+ 2,
+ 371,
+ 49,
+ 117
+ ],
+ "category_id": 6,
+ "area": 103472,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19718,
+ "image_id": 3619,
+ "bbox": [
+ 105,
+ 56,
+ 63,
+ 46
+ ],
+ "category_id": 6,
+ "area": 52052,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19719,
+ "image_id": 3619,
+ "bbox": [
+ 178,
+ 58,
+ 66,
+ 55
+ ],
+ "category_id": 6,
+ "area": 65043,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19720,
+ "image_id": 3619,
+ "bbox": [
+ 128,
+ 151,
+ 103,
+ 113
+ ],
+ "category_id": 6,
+ "area": 207872,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19721,
+ "image_id": 3619,
+ "bbox": [
+ 250,
+ 67,
+ 98,
+ 166
+ ],
+ "category_id": 6,
+ "area": 290836,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19722,
+ "image_id": 3619,
+ "bbox": [
+ 242,
+ 226,
+ 89,
+ 68
+ ],
+ "category_id": 6,
+ "area": 108671,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19723,
+ "image_id": 3619,
+ "bbox": [
+ 134,
+ 265,
+ 116,
+ 85
+ ],
+ "category_id": 6,
+ "area": 176925,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19724,
+ "image_id": 3619,
+ "bbox": [
+ 91,
+ 352,
+ 115,
+ 84
+ ],
+ "category_id": 6,
+ "area": 173012,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19725,
+ "image_id": 3619,
+ "bbox": [
+ 331,
+ 4,
+ 86,
+ 104
+ ],
+ "category_id": 6,
+ "area": 160632,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19726,
+ "image_id": 3619,
+ "bbox": [
+ 349,
+ 89,
+ 32,
+ 68
+ ],
+ "category_id": 6,
+ "area": 38880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19727,
+ "image_id": 3619,
+ "bbox": [
+ 362,
+ 132,
+ 68,
+ 61
+ ],
+ "category_id": 6,
+ "area": 75152,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19728,
+ "image_id": 3619,
+ "bbox": [
+ 434,
+ 122,
+ 76,
+ 90
+ ],
+ "category_id": 6,
+ "area": 121752,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19729,
+ "image_id": 3619,
+ "bbox": [
+ 262,
+ 289,
+ 138,
+ 74
+ ],
+ "category_id": 6,
+ "area": 183195,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19739,
+ "image_id": 3620,
+ "bbox": [
+ 204,
+ 100,
+ 229,
+ 324
+ ],
+ "category_id": 8,
+ "area": 98192,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19740,
+ "image_id": 3621,
+ "bbox": [
+ 0,
+ 116,
+ 279,
+ 160
+ ],
+ "category_id": 9,
+ "area": 75342,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19741,
+ "image_id": 3621,
+ "bbox": [
+ 0,
+ 206,
+ 183,
+ 304
+ ],
+ "category_id": 6,
+ "area": 94335,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19752,
+ "image_id": 3623,
+ "bbox": [
+ 2,
+ 194,
+ 492,
+ 317
+ ],
+ "category_id": 6,
+ "area": 371037,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19767,
+ "image_id": 3626,
+ "bbox": [
+ 203,
+ 203,
+ 217,
+ 159
+ ],
+ "category_id": 9,
+ "area": 92442,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19768,
+ "image_id": 3627,
+ "bbox": [
+ 36,
+ 220,
+ 400,
+ 156
+ ],
+ "category_id": 9,
+ "area": 152500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19769,
+ "image_id": 3627,
+ "bbox": [
+ 431,
+ 337,
+ 80,
+ 108
+ ],
+ "category_id": 6,
+ "area": 21420,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19770,
+ "image_id": 3628,
+ "bbox": [
+ 133,
+ 227,
+ 37,
+ 65
+ ],
+ "category_id": 8,
+ "area": 7728,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19771,
+ "image_id": 3628,
+ "bbox": [
+ 243,
+ 261,
+ 85,
+ 54
+ ],
+ "category_id": 8,
+ "area": 14628,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19772,
+ "image_id": 3629,
+ "bbox": [
+ 2,
+ 184,
+ 279,
+ 261
+ ],
+ "category_id": 8,
+ "area": 578496,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19773,
+ "image_id": 3629,
+ "bbox": [
+ 380,
+ 186,
+ 131,
+ 317
+ ],
+ "category_id": 8,
+ "area": 330486,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19774,
+ "image_id": 3630,
+ "bbox": [
+ 48,
+ 192,
+ 195,
+ 254
+ ],
+ "category_id": 8,
+ "area": 173728,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19775,
+ "image_id": 3630,
+ "bbox": [
+ 180,
+ 209,
+ 329,
+ 242
+ ],
+ "category_id": 8,
+ "area": 279336,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19785,
+ "image_id": 3633,
+ "bbox": [
+ 128,
+ 102,
+ 284,
+ 351
+ ],
+ "category_id": 8,
+ "area": 351234,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19786,
+ "image_id": 3633,
+ "bbox": [
+ 44,
+ 24,
+ 95,
+ 130
+ ],
+ "category_id": 6,
+ "area": 43554,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19787,
+ "image_id": 3633,
+ "bbox": [
+ 6,
+ 355,
+ 93,
+ 91
+ ],
+ "category_id": 6,
+ "area": 29824,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19788,
+ "image_id": 3633,
+ "bbox": [
+ 420,
+ 214,
+ 91,
+ 194
+ ],
+ "category_id": 6,
+ "area": 62746,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19789,
+ "image_id": 3633,
+ "bbox": [
+ 438,
+ 71,
+ 74,
+ 105
+ ],
+ "category_id": 6,
+ "area": 27565,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19790,
+ "image_id": 3633,
+ "bbox": [
+ 42,
+ 148,
+ 48,
+ 96
+ ],
+ "category_id": 6,
+ "area": 16320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19803,
+ "image_id": 3636,
+ "bbox": [
+ 362,
+ 171,
+ 124,
+ 88
+ ],
+ "category_id": 9,
+ "area": 38564,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19804,
+ "image_id": 3636,
+ "bbox": [
+ 270,
+ 102,
+ 85,
+ 128
+ ],
+ "category_id": 9,
+ "area": 38340,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19805,
+ "image_id": 3636,
+ "bbox": [
+ 300,
+ 112,
+ 110,
+ 150
+ ],
+ "category_id": 9,
+ "area": 58447,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19807,
+ "image_id": 3637,
+ "bbox": [
+ 113,
+ 279,
+ 143,
+ 74
+ ],
+ "category_id": 9,
+ "area": 37590,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19811,
+ "image_id": 3638,
+ "bbox": [
+ 34,
+ 198,
+ 421,
+ 228
+ ],
+ "category_id": 8,
+ "area": 241424,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19812,
+ "image_id": 3638,
+ "bbox": [
+ 318,
+ 59,
+ 193,
+ 366
+ ],
+ "category_id": 8,
+ "area": 178308,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19818,
+ "image_id": 3639,
+ "bbox": [
+ 128,
+ 91,
+ 135,
+ 133
+ ],
+ "category_id": 6,
+ "area": 175956,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19819,
+ "image_id": 3639,
+ "bbox": [
+ 78,
+ 227,
+ 51,
+ 131
+ ],
+ "category_id": 6,
+ "area": 66192,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19820,
+ "image_id": 3639,
+ "bbox": [
+ 298,
+ 1,
+ 31,
+ 47
+ ],
+ "category_id": 6,
+ "area": 14640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19821,
+ "image_id": 3639,
+ "bbox": [
+ 257,
+ 33,
+ 98,
+ 100
+ ],
+ "category_id": 6,
+ "area": 97524,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19822,
+ "image_id": 3639,
+ "bbox": [
+ 279,
+ 131,
+ 157,
+ 116
+ ],
+ "category_id": 6,
+ "area": 178794,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19823,
+ "image_id": 3639,
+ "bbox": [
+ 0,
+ 300,
+ 30,
+ 131
+ ],
+ "category_id": 6,
+ "area": 39766,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19828,
+ "image_id": 3639,
+ "bbox": [
+ 146,
+ 1,
+ 94,
+ 95
+ ],
+ "category_id": 6,
+ "area": 88200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19829,
+ "image_id": 3640,
+ "bbox": [
+ 0,
+ 79,
+ 25,
+ 52
+ ],
+ "category_id": 6,
+ "area": 9546,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19830,
+ "image_id": 3640,
+ "bbox": [
+ 60,
+ 126,
+ 21,
+ 24
+ ],
+ "category_id": 6,
+ "area": 3744,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19831,
+ "image_id": 3640,
+ "bbox": [
+ 26,
+ 110,
+ 15,
+ 21
+ ],
+ "category_id": 6,
+ "area": 2340,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19832,
+ "image_id": 3640,
+ "bbox": [
+ 80,
+ 332,
+ 13,
+ 53
+ ],
+ "category_id": 6,
+ "area": 4928,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19833,
+ "image_id": 3640,
+ "bbox": [
+ 125,
+ 310,
+ 19,
+ 30
+ ],
+ "category_id": 6,
+ "area": 4096,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19834,
+ "image_id": 3640,
+ "bbox": [
+ 142,
+ 305,
+ 22,
+ 27
+ ],
+ "category_id": 6,
+ "area": 4292,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19835,
+ "image_id": 3640,
+ "bbox": [
+ 355,
+ 433,
+ 48,
+ 78
+ ],
+ "category_id": 6,
+ "area": 26560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19836,
+ "image_id": 3640,
+ "bbox": [
+ 317,
+ 374,
+ 34,
+ 66
+ ],
+ "category_id": 6,
+ "area": 15846,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19837,
+ "image_id": 3640,
+ "bbox": [
+ 351,
+ 336,
+ 16,
+ 40
+ ],
+ "category_id": 6,
+ "area": 4816,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19838,
+ "image_id": 3640,
+ "bbox": [
+ 353,
+ 405,
+ 12,
+ 34
+ ],
+ "category_id": 6,
+ "area": 2920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19839,
+ "image_id": 3640,
+ "bbox": [
+ 349,
+ 307,
+ 11,
+ 22
+ ],
+ "category_id": 6,
+ "area": 1739,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19840,
+ "image_id": 3640,
+ "bbox": [
+ 353,
+ 377,
+ 12,
+ 20
+ ],
+ "category_id": 6,
+ "area": 1763,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19841,
+ "image_id": 3640,
+ "bbox": [
+ 306,
+ 193,
+ 141,
+ 167
+ ],
+ "category_id": 6,
+ "area": 165440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19842,
+ "image_id": 3640,
+ "bbox": [
+ 345,
+ 166,
+ 153,
+ 136
+ ],
+ "category_id": 6,
+ "area": 145796,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19843,
+ "image_id": 3641,
+ "bbox": [
+ 289,
+ 30,
+ 210,
+ 427
+ ],
+ "category_id": 8,
+ "area": 207163,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19844,
+ "image_id": 3641,
+ "bbox": [
+ 32,
+ 0,
+ 345,
+ 509
+ ],
+ "category_id": 8,
+ "area": 404808,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19846,
+ "image_id": 3642,
+ "bbox": [
+ 239,
+ 277,
+ 82,
+ 52
+ ],
+ "category_id": 9,
+ "area": 6480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19847,
+ "image_id": 3643,
+ "bbox": [
+ 34,
+ 199,
+ 261,
+ 178
+ ],
+ "category_id": 9,
+ "area": 110088,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19854,
+ "image_id": 3646,
+ "bbox": [
+ 0,
+ 117,
+ 90,
+ 127
+ ],
+ "category_id": 8,
+ "area": 28512,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19864,
+ "image_id": 3647,
+ "bbox": [
+ 136,
+ 195,
+ 282,
+ 246
+ ],
+ "category_id": 8,
+ "area": 172131,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19865,
+ "image_id": 3647,
+ "bbox": [
+ 36,
+ 105,
+ 148,
+ 142
+ ],
+ "category_id": 8,
+ "area": 52338,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19866,
+ "image_id": 3647,
+ "bbox": [
+ 164,
+ 53,
+ 164,
+ 174
+ ],
+ "category_id": 8,
+ "area": 70560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19867,
+ "image_id": 3647,
+ "bbox": [
+ 403,
+ 82,
+ 87,
+ 98
+ ],
+ "category_id": 8,
+ "area": 21168,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19875,
+ "image_id": 3649,
+ "bbox": [
+ 0,
+ 151,
+ 363,
+ 359
+ ],
+ "category_id": 8,
+ "area": 383780,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19876,
+ "image_id": 3649,
+ "bbox": [
+ 297,
+ 58,
+ 214,
+ 330
+ ],
+ "category_id": 8,
+ "area": 208823,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19877,
+ "image_id": 3650,
+ "bbox": [
+ 3,
+ 2,
+ 122,
+ 164
+ ],
+ "category_id": 6,
+ "area": 197316,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19878,
+ "image_id": 3650,
+ "bbox": [
+ 1,
+ 158,
+ 150,
+ 91
+ ],
+ "category_id": 6,
+ "area": 135728,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19879,
+ "image_id": 3650,
+ "bbox": [
+ 167,
+ 190,
+ 184,
+ 99
+ ],
+ "category_id": 6,
+ "area": 180560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19880,
+ "image_id": 3650,
+ "bbox": [
+ 0,
+ 275,
+ 91,
+ 75
+ ],
+ "category_id": 6,
+ "area": 68400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19881,
+ "image_id": 3650,
+ "bbox": [
+ 443,
+ 200,
+ 65,
+ 117
+ ],
+ "category_id": 6,
+ "area": 75646,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19882,
+ "image_id": 3650,
+ "bbox": [
+ 148,
+ 0,
+ 138,
+ 119
+ ],
+ "category_id": 6,
+ "area": 161778,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19883,
+ "image_id": 3650,
+ "bbox": [
+ 142,
+ 107,
+ 119,
+ 92
+ ],
+ "category_id": 6,
+ "area": 108900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19884,
+ "image_id": 3650,
+ "bbox": [
+ 286,
+ 0,
+ 98,
+ 78
+ ],
+ "category_id": 6,
+ "area": 75725,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19885,
+ "image_id": 3650,
+ "bbox": [
+ 399,
+ 0,
+ 112,
+ 82
+ ],
+ "category_id": 6,
+ "area": 90639,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19889,
+ "image_id": 3651,
+ "bbox": [
+ 76,
+ 158,
+ 222,
+ 242
+ ],
+ "category_id": 9,
+ "area": 189255,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19890,
+ "image_id": 3651,
+ "bbox": [
+ 343,
+ 58,
+ 166,
+ 121
+ ],
+ "category_id": 9,
+ "area": 71136,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19891,
+ "image_id": 3651,
+ "bbox": [
+ 120,
+ 194,
+ 278,
+ 275
+ ],
+ "category_id": 9,
+ "area": 269352,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19892,
+ "image_id": 3651,
+ "bbox": [
+ 318,
+ 167,
+ 190,
+ 342
+ ],
+ "category_id": 9,
+ "area": 229437,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19896,
+ "image_id": 3653,
+ "bbox": [
+ 345,
+ 30,
+ 166,
+ 192
+ ],
+ "category_id": 9,
+ "area": 24674,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19900,
+ "image_id": 3655,
+ "bbox": [
+ 2,
+ 2,
+ 509,
+ 269
+ ],
+ "category_id": 6,
+ "area": 349860,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19905,
+ "image_id": 3656,
+ "bbox": [
+ 91,
+ 159,
+ 41,
+ 56
+ ],
+ "category_id": 8,
+ "area": 3060,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19906,
+ "image_id": 3656,
+ "bbox": [
+ 183,
+ 107,
+ 110,
+ 140
+ ],
+ "category_id": 8,
+ "area": 20608,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19907,
+ "image_id": 3656,
+ "bbox": [
+ 314,
+ 214,
+ 83,
+ 110
+ ],
+ "category_id": 8,
+ "area": 12232,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19908,
+ "image_id": 3656,
+ "bbox": [
+ 227,
+ 347,
+ 97,
+ 163
+ ],
+ "category_id": 6,
+ "area": 21060,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19909,
+ "image_id": 3656,
+ "bbox": [
+ 151,
+ 330,
+ 78,
+ 179
+ ],
+ "category_id": 6,
+ "area": 18590,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19910,
+ "image_id": 3656,
+ "bbox": [
+ 159,
+ 481,
+ 69,
+ 27
+ ],
+ "category_id": 6,
+ "area": 2530,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19917,
+ "image_id": 3658,
+ "bbox": [
+ 318,
+ 64,
+ 136,
+ 121
+ ],
+ "category_id": 6,
+ "area": 12789,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19918,
+ "image_id": 3658,
+ "bbox": [
+ 215,
+ 0,
+ 126,
+ 79
+ ],
+ "category_id": 6,
+ "area": 7752,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19919,
+ "image_id": 3659,
+ "bbox": [
+ 67,
+ 97,
+ 426,
+ 414
+ ],
+ "category_id": 8,
+ "area": 1362984,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19932,
+ "image_id": 3662,
+ "bbox": [
+ 8,
+ 74,
+ 140,
+ 265
+ ],
+ "category_id": 9,
+ "area": 131296,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19933,
+ "image_id": 3662,
+ "bbox": [
+ 172,
+ 234,
+ 79,
+ 173
+ ],
+ "category_id": 9,
+ "area": 48312,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19934,
+ "image_id": 3662,
+ "bbox": [
+ 213,
+ 221,
+ 168,
+ 282
+ ],
+ "category_id": 9,
+ "area": 166740,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19935,
+ "image_id": 3662,
+ "bbox": [
+ 225,
+ 8,
+ 141,
+ 270
+ ],
+ "category_id": 9,
+ "area": 134520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19936,
+ "image_id": 3662,
+ "bbox": [
+ 242,
+ 43,
+ 161,
+ 280
+ ],
+ "category_id": 9,
+ "area": 158782,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19937,
+ "image_id": 3662,
+ "bbox": [
+ 380,
+ 322,
+ 89,
+ 182
+ ],
+ "category_id": 9,
+ "area": 57311,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19938,
+ "image_id": 3663,
+ "bbox": [
+ 148,
+ 156,
+ 314,
+ 338
+ ],
+ "category_id": 9,
+ "area": 283679,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19939,
+ "image_id": 3663,
+ "bbox": [
+ 91,
+ 149,
+ 61,
+ 107
+ ],
+ "category_id": 10,
+ "area": 17446,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19956,
+ "image_id": 3666,
+ "bbox": [
+ 142,
+ 189,
+ 174,
+ 259
+ ],
+ "category_id": 9,
+ "area": 348168,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19957,
+ "image_id": 3667,
+ "bbox": [
+ 309,
+ 221,
+ 201,
+ 184
+ ],
+ "category_id": 8,
+ "area": 97155,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19963,
+ "image_id": 3669,
+ "bbox": [
+ 0,
+ 9,
+ 85,
+ 132
+ ],
+ "category_id": 6,
+ "area": 125268,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19964,
+ "image_id": 3669,
+ "bbox": [
+ 91,
+ 0,
+ 126,
+ 103
+ ],
+ "category_id": 6,
+ "area": 144584,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19966,
+ "image_id": 3669,
+ "bbox": [
+ 186,
+ 2,
+ 140,
+ 201
+ ],
+ "category_id": 6,
+ "area": 313880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19967,
+ "image_id": 3669,
+ "bbox": [
+ 185,
+ 192,
+ 116,
+ 81
+ ],
+ "category_id": 6,
+ "area": 104910,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19968,
+ "image_id": 3669,
+ "bbox": [
+ 327,
+ 27,
+ 35,
+ 91
+ ],
+ "category_id": 6,
+ "area": 36057,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19969,
+ "image_id": 3669,
+ "bbox": [
+ 30,
+ 75,
+ 137,
+ 167
+ ],
+ "category_id": 6,
+ "area": 254562,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19970,
+ "image_id": 3669,
+ "bbox": [
+ 348,
+ 77,
+ 75,
+ 83
+ ],
+ "category_id": 6,
+ "area": 69828,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 19971,
+ "image_id": 3669,
+ "bbox": [
+ 438,
+ 80,
+ 72,
+ 92
+ ],
+ "category_id": 6,
+ "area": 74420,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20005,
+ "image_id": 3672,
+ "bbox": [
+ 0,
+ 0,
+ 408,
+ 290
+ ],
+ "category_id": 6,
+ "area": 295974,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20006,
+ "image_id": 3672,
+ "bbox": [
+ 472,
+ 127,
+ 39,
+ 92
+ ],
+ "category_id": 6,
+ "area": 9000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20007,
+ "image_id": 3673,
+ "bbox": [
+ 124,
+ 103,
+ 144,
+ 273
+ ],
+ "category_id": 8,
+ "area": 113760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20009,
+ "image_id": 3673,
+ "bbox": [
+ 116,
+ 428,
+ 289,
+ 82
+ ],
+ "category_id": 6,
+ "area": 68400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20010,
+ "image_id": 3673,
+ "bbox": [
+ 63,
+ 134,
+ 94,
+ 129
+ ],
+ "category_id": 6,
+ "area": 35100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20011,
+ "image_id": 3673,
+ "bbox": [
+ 89,
+ 356,
+ 85,
+ 149
+ ],
+ "category_id": 6,
+ "area": 36849,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20012,
+ "image_id": 3674,
+ "bbox": [
+ 166,
+ 128,
+ 75,
+ 133
+ ],
+ "category_id": 6,
+ "area": 35344,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20013,
+ "image_id": 3674,
+ "bbox": [
+ 230,
+ 202,
+ 61,
+ 110
+ ],
+ "category_id": 6,
+ "area": 23870,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20014,
+ "image_id": 3674,
+ "bbox": [
+ 221,
+ 280,
+ 122,
+ 203
+ ],
+ "category_id": 6,
+ "area": 87802,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20015,
+ "image_id": 3674,
+ "bbox": [
+ 205,
+ 123,
+ 52,
+ 71
+ ],
+ "category_id": 6,
+ "area": 13000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20016,
+ "image_id": 3674,
+ "bbox": [
+ 236,
+ 192,
+ 105,
+ 86
+ ],
+ "category_id": 6,
+ "area": 31944,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20039,
+ "image_id": 3675,
+ "bbox": [
+ 67,
+ 329,
+ 161,
+ 143
+ ],
+ "category_id": 9,
+ "area": 36720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20040,
+ "image_id": 3676,
+ "bbox": [
+ 9,
+ 10,
+ 474,
+ 500
+ ],
+ "category_id": 9,
+ "area": 659555,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20069,
+ "image_id": 3680,
+ "bbox": [
+ 114,
+ 262,
+ 144,
+ 86
+ ],
+ "category_id": 9,
+ "area": 44164,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20122,
+ "image_id": 3684,
+ "bbox": [
+ 211,
+ 192,
+ 117,
+ 91
+ ],
+ "category_id": 8,
+ "area": 37504,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20124,
+ "image_id": 3684,
+ "bbox": [
+ 208,
+ 300,
+ 38,
+ 52
+ ],
+ "category_id": 6,
+ "area": 7178,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20125,
+ "image_id": 3684,
+ "bbox": [
+ 283,
+ 357,
+ 87,
+ 91
+ ],
+ "category_id": 6,
+ "area": 28032,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20126,
+ "image_id": 3684,
+ "bbox": [
+ 82,
+ 157,
+ 90,
+ 93
+ ],
+ "category_id": 6,
+ "area": 29475,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20127,
+ "image_id": 3684,
+ "bbox": [
+ 86,
+ 218,
+ 91,
+ 103
+ ],
+ "category_id": 6,
+ "area": 33288,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20128,
+ "image_id": 3684,
+ "bbox": [
+ 92,
+ 327,
+ 47,
+ 59
+ ],
+ "category_id": 6,
+ "area": 9794,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20129,
+ "image_id": 3684,
+ "bbox": [
+ 177,
+ 447,
+ 92,
+ 64
+ ],
+ "category_id": 6,
+ "area": 21112,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20130,
+ "image_id": 3684,
+ "bbox": [
+ 278,
+ 414,
+ 115,
+ 97
+ ],
+ "category_id": 6,
+ "area": 39593,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20131,
+ "image_id": 3684,
+ "bbox": [
+ 369,
+ 383,
+ 142,
+ 128
+ ],
+ "category_id": 6,
+ "area": 64255,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20132,
+ "image_id": 3685,
+ "bbox": [
+ 37,
+ 172,
+ 367,
+ 156
+ ],
+ "category_id": 9,
+ "area": 35182,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20139,
+ "image_id": 3686,
+ "bbox": [
+ 0,
+ 74,
+ 503,
+ 221
+ ],
+ "category_id": 9,
+ "area": 109200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20140,
+ "image_id": 3686,
+ "bbox": [
+ 234,
+ 247,
+ 124,
+ 59
+ ],
+ "category_id": 9,
+ "area": 7290,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20141,
+ "image_id": 3687,
+ "bbox": [
+ 244,
+ 260,
+ 199,
+ 178
+ ],
+ "category_id": 9,
+ "area": 76935,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20142,
+ "image_id": 3687,
+ "bbox": [
+ 0,
+ 0,
+ 86,
+ 74
+ ],
+ "category_id": 6,
+ "area": 13857,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20143,
+ "image_id": 3688,
+ "bbox": [
+ 6,
+ 85,
+ 229,
+ 160
+ ],
+ "category_id": 8,
+ "area": 96574,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20144,
+ "image_id": 3688,
+ "bbox": [
+ 423,
+ 248,
+ 88,
+ 83
+ ],
+ "category_id": 8,
+ "area": 19323,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20145,
+ "image_id": 3688,
+ "bbox": [
+ 22,
+ 121,
+ 412,
+ 279
+ ],
+ "category_id": 8,
+ "area": 301684,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20149,
+ "image_id": 3690,
+ "bbox": [
+ 0,
+ 169,
+ 95,
+ 130
+ ],
+ "category_id": 6,
+ "area": 43737,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20150,
+ "image_id": 3690,
+ "bbox": [
+ 111,
+ 320,
+ 52,
+ 85
+ ],
+ "category_id": 6,
+ "area": 15600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20151,
+ "image_id": 3690,
+ "bbox": [
+ 335,
+ 68,
+ 72,
+ 94
+ ],
+ "category_id": 6,
+ "area": 23940,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20152,
+ "image_id": 3690,
+ "bbox": [
+ 394,
+ 34,
+ 116,
+ 162
+ ],
+ "category_id": 6,
+ "area": 66639,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20153,
+ "image_id": 3690,
+ "bbox": [
+ 407,
+ 280,
+ 96,
+ 231
+ ],
+ "category_id": 6,
+ "area": 78892,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20154,
+ "image_id": 3690,
+ "bbox": [
+ 73,
+ 188,
+ 404,
+ 271
+ ],
+ "category_id": 8,
+ "area": 385820,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20155,
+ "image_id": 3690,
+ "bbox": [
+ 286,
+ 71,
+ 22,
+ 71
+ ],
+ "category_id": 6,
+ "area": 5656,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20156,
+ "image_id": 3691,
+ "bbox": [
+ 2,
+ 54,
+ 340,
+ 452
+ ],
+ "category_id": 9,
+ "area": 542087,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20157,
+ "image_id": 3691,
+ "bbox": [
+ 148,
+ 36,
+ 358,
+ 475
+ ],
+ "category_id": 9,
+ "area": 598755,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20166,
+ "image_id": 3693,
+ "bbox": [
+ 28,
+ 154,
+ 447,
+ 272
+ ],
+ "category_id": 9,
+ "area": 142800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20167,
+ "image_id": 3694,
+ "bbox": [
+ 85,
+ 118,
+ 361,
+ 389
+ ],
+ "category_id": 8,
+ "area": 1113810,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20168,
+ "image_id": 3694,
+ "bbox": [
+ 293,
+ 126,
+ 218,
+ 143
+ ],
+ "category_id": 8,
+ "area": 247338,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20169,
+ "image_id": 3695,
+ "bbox": [
+ 197,
+ 155,
+ 314,
+ 356
+ ],
+ "category_id": 9,
+ "area": 137484,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20170,
+ "image_id": 3696,
+ "bbox": [
+ 286,
+ 57,
+ 78,
+ 123
+ ],
+ "category_id": 9,
+ "area": 34278,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20171,
+ "image_id": 3696,
+ "bbox": [
+ 250,
+ 129,
+ 76,
+ 190
+ ],
+ "category_id": 9,
+ "area": 50920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20172,
+ "image_id": 3696,
+ "bbox": [
+ 218,
+ 464,
+ 51,
+ 42
+ ],
+ "category_id": 9,
+ "area": 7740,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20173,
+ "image_id": 3696,
+ "bbox": [
+ 262,
+ 61,
+ 29,
+ 61
+ ],
+ "category_id": 9,
+ "area": 6364,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20174,
+ "image_id": 3697,
+ "bbox": [
+ 174,
+ 220,
+ 174,
+ 173
+ ],
+ "category_id": 8,
+ "area": 105948,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20175,
+ "image_id": 3697,
+ "bbox": [
+ 277,
+ 102,
+ 182,
+ 128
+ ],
+ "category_id": 8,
+ "area": 81900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20191,
+ "image_id": 3698,
+ "bbox": [
+ 0,
+ 50,
+ 393,
+ 460
+ ],
+ "category_id": 6,
+ "area": 543817,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20197,
+ "image_id": 3701,
+ "bbox": [
+ 124,
+ 85,
+ 48,
+ 67
+ ],
+ "category_id": 8,
+ "area": 22204,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20198,
+ "image_id": 3701,
+ "bbox": [
+ 110,
+ 142,
+ 85,
+ 71
+ ],
+ "category_id": 8,
+ "area": 41470,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20199,
+ "image_id": 3701,
+ "bbox": [
+ 205,
+ 80,
+ 25,
+ 36
+ ],
+ "category_id": 8,
+ "area": 6365,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20200,
+ "image_id": 3701,
+ "bbox": [
+ 238,
+ 203,
+ 53,
+ 120
+ ],
+ "category_id": 8,
+ "area": 44238,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20201,
+ "image_id": 3701,
+ "bbox": [
+ 322,
+ 113,
+ 47,
+ 78
+ ],
+ "category_id": 8,
+ "area": 25311,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20202,
+ "image_id": 3702,
+ "bbox": [
+ 34,
+ 101,
+ 432,
+ 372
+ ],
+ "category_id": 9,
+ "area": 1636641,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20203,
+ "image_id": 3703,
+ "bbox": [
+ 31,
+ 167,
+ 378,
+ 256
+ ],
+ "category_id": 9,
+ "area": 404888,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20220,
+ "image_id": 3706,
+ "bbox": [
+ 28,
+ 166,
+ 309,
+ 118
+ ],
+ "category_id": 8,
+ "area": 92628,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20222,
+ "image_id": 3707,
+ "bbox": [
+ 81,
+ 0,
+ 148,
+ 170
+ ],
+ "category_id": 6,
+ "area": 46081,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20223,
+ "image_id": 3707,
+ "bbox": [
+ 16,
+ 110,
+ 219,
+ 281
+ ],
+ "category_id": 6,
+ "area": 112895,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20224,
+ "image_id": 3707,
+ "bbox": [
+ 211,
+ 255,
+ 193,
+ 142
+ ],
+ "category_id": 9,
+ "area": 50320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20225,
+ "image_id": 3707,
+ "bbox": [
+ 25,
+ 22,
+ 96,
+ 142
+ ],
+ "category_id": 6,
+ "area": 25160,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20229,
+ "image_id": 3708,
+ "bbox": [
+ 225,
+ 297,
+ 80,
+ 82
+ ],
+ "category_id": 6,
+ "area": 44304,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20230,
+ "image_id": 3708,
+ "bbox": [
+ 275,
+ 418,
+ 38,
+ 93
+ ],
+ "category_id": 6,
+ "area": 23760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20232,
+ "image_id": 3709,
+ "bbox": [
+ 207,
+ 133,
+ 304,
+ 378
+ ],
+ "category_id": 6,
+ "area": 500004,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20235,
+ "image_id": 3711,
+ "bbox": [
+ 0,
+ 127,
+ 279,
+ 157
+ ],
+ "category_id": 9,
+ "area": 74214,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20236,
+ "image_id": 3711,
+ "bbox": [
+ 0,
+ 211,
+ 179,
+ 299
+ ],
+ "category_id": 6,
+ "area": 90675,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20246,
+ "image_id": 3713,
+ "bbox": [
+ 43,
+ 103,
+ 332,
+ 408
+ ],
+ "category_id": 8,
+ "area": 477568,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20249,
+ "image_id": 3714,
+ "bbox": [
+ 280,
+ 172,
+ 231,
+ 283
+ ],
+ "category_id": 8,
+ "area": 76585,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20250,
+ "image_id": 3715,
+ "bbox": [
+ 114,
+ 80,
+ 202,
+ 365
+ ],
+ "category_id": 9,
+ "area": 780520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20251,
+ "image_id": 3715,
+ "bbox": [
+ 333,
+ 132,
+ 48,
+ 73
+ ],
+ "category_id": 6,
+ "area": 37698,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20252,
+ "image_id": 3715,
+ "bbox": [
+ 460,
+ 242,
+ 51,
+ 56
+ ],
+ "category_id": 6,
+ "area": 30846,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20267,
+ "image_id": 3717,
+ "bbox": [
+ 40,
+ 1,
+ 471,
+ 245
+ ],
+ "category_id": 6,
+ "area": 285505,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20298,
+ "image_id": 3720,
+ "bbox": [
+ 155,
+ 2,
+ 188,
+ 235
+ ],
+ "category_id": 9,
+ "area": 149872,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20299,
+ "image_id": 3720,
+ "bbox": [
+ 207,
+ 101,
+ 113,
+ 198
+ ],
+ "category_id": 10,
+ "area": 76167,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20300,
+ "image_id": 3721,
+ "bbox": [
+ 252,
+ 76,
+ 60,
+ 131
+ ],
+ "category_id": 8,
+ "area": 62602,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20301,
+ "image_id": 3721,
+ "bbox": [
+ 319,
+ 24,
+ 91,
+ 119
+ ],
+ "category_id": 6,
+ "area": 87032,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20302,
+ "image_id": 3722,
+ "bbox": [
+ 182,
+ 206,
+ 226,
+ 170
+ ],
+ "category_id": 8,
+ "area": 44838,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20303,
+ "image_id": 3723,
+ "bbox": [
+ 131,
+ 9,
+ 345,
+ 467
+ ],
+ "category_id": 9,
+ "area": 430493,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20304,
+ "image_id": 3723,
+ "bbox": [
+ 111,
+ 270,
+ 73,
+ 134
+ ],
+ "category_id": 10,
+ "area": 26313,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20305,
+ "image_id": 3724,
+ "bbox": [
+ 42,
+ 161,
+ 139,
+ 132
+ ],
+ "category_id": 6,
+ "area": 146720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20306,
+ "image_id": 3724,
+ "bbox": [
+ 121,
+ 1,
+ 76,
+ 54
+ ],
+ "category_id": 6,
+ "area": 33408,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20307,
+ "image_id": 3724,
+ "bbox": [
+ 383,
+ 206,
+ 64,
+ 91
+ ],
+ "category_id": 6,
+ "area": 46656,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20308,
+ "image_id": 3724,
+ "bbox": [
+ 223,
+ 122,
+ 122,
+ 141
+ ],
+ "category_id": 8,
+ "area": 137540,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20309,
+ "image_id": 3725,
+ "bbox": [
+ 57,
+ 180,
+ 303,
+ 247
+ ],
+ "category_id": 9,
+ "area": 51504,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20310,
+ "image_id": 3726,
+ "bbox": [
+ 161,
+ 101,
+ 186,
+ 405
+ ],
+ "category_id": 6,
+ "area": 596790,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20317,
+ "image_id": 3728,
+ "bbox": [
+ 210,
+ 237,
+ 100,
+ 107
+ ],
+ "category_id": 8,
+ "area": 37901,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20318,
+ "image_id": 3728,
+ "bbox": [
+ 271,
+ 253,
+ 99,
+ 69
+ ],
+ "category_id": 8,
+ "area": 24304,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20322,
+ "image_id": 3730,
+ "bbox": [
+ 33,
+ 205,
+ 287,
+ 189
+ ],
+ "category_id": 9,
+ "area": 101061,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20323,
+ "image_id": 3730,
+ "bbox": [
+ 365,
+ 128,
+ 98,
+ 171
+ ],
+ "category_id": 9,
+ "area": 31328,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20324,
+ "image_id": 3731,
+ "bbox": [
+ 147,
+ 91,
+ 105,
+ 127
+ ],
+ "category_id": 6,
+ "area": 34773,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20325,
+ "image_id": 3731,
+ "bbox": [
+ 277,
+ 0,
+ 234,
+ 160
+ ],
+ "category_id": 6,
+ "area": 96565,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20326,
+ "image_id": 3731,
+ "bbox": [
+ 114,
+ 193,
+ 332,
+ 214
+ ],
+ "category_id": 9,
+ "area": 182990,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20327,
+ "image_id": 3732,
+ "bbox": [
+ 78,
+ 223,
+ 194,
+ 204
+ ],
+ "category_id": 8,
+ "area": 139482,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20328,
+ "image_id": 3732,
+ "bbox": [
+ 198,
+ 197,
+ 299,
+ 242
+ ],
+ "category_id": 8,
+ "area": 253911,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20329,
+ "image_id": 3732,
+ "bbox": [
+ 363,
+ 89,
+ 32,
+ 43
+ ],
+ "category_id": 6,
+ "area": 4941,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20330,
+ "image_id": 3732,
+ "bbox": [
+ 399,
+ 137,
+ 35,
+ 46
+ ],
+ "category_id": 6,
+ "area": 5720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20331,
+ "image_id": 3732,
+ "bbox": [
+ 438,
+ 179,
+ 33,
+ 23
+ ],
+ "category_id": 6,
+ "area": 2739,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20332,
+ "image_id": 3732,
+ "bbox": [
+ 392,
+ 189,
+ 46,
+ 40
+ ],
+ "category_id": 6,
+ "area": 6612,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20333,
+ "image_id": 3732,
+ "bbox": [
+ 351,
+ 470,
+ 58,
+ 41
+ ],
+ "category_id": 6,
+ "area": 8410,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20341,
+ "image_id": 3734,
+ "bbox": [
+ 26,
+ 239,
+ 399,
+ 272
+ ],
+ "category_id": 9,
+ "area": 219780,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20343,
+ "image_id": 3734,
+ "bbox": [
+ 459,
+ 70,
+ 51,
+ 140
+ ],
+ "category_id": 6,
+ "area": 14688,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20344,
+ "image_id": 3735,
+ "bbox": [
+ 173,
+ 278,
+ 87,
+ 48
+ ],
+ "category_id": 8,
+ "area": 33784,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20345,
+ "image_id": 3735,
+ "bbox": [
+ 85,
+ 303,
+ 122,
+ 125
+ ],
+ "category_id": 8,
+ "area": 121370,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20346,
+ "image_id": 3735,
+ "bbox": [
+ 168,
+ 348,
+ 128,
+ 79
+ ],
+ "category_id": 8,
+ "area": 80327,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20347,
+ "image_id": 3735,
+ "bbox": [
+ 405,
+ 238,
+ 41,
+ 61
+ ],
+ "category_id": 6,
+ "area": 20280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20355,
+ "image_id": 3738,
+ "bbox": [
+ 2,
+ 0,
+ 329,
+ 439
+ ],
+ "category_id": 6,
+ "area": 178227,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20356,
+ "image_id": 3738,
+ "bbox": [
+ 310,
+ 97,
+ 200,
+ 391
+ ],
+ "category_id": 6,
+ "area": 96726,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20359,
+ "image_id": 3739,
+ "bbox": [
+ 34,
+ 404,
+ 94,
+ 91
+ ],
+ "category_id": 6,
+ "area": 24165,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20360,
+ "image_id": 3739,
+ "bbox": [
+ 103,
+ 398,
+ 126,
+ 113
+ ],
+ "category_id": 6,
+ "area": 40080,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20379,
+ "image_id": 3741,
+ "bbox": [
+ 1,
+ 0,
+ 510,
+ 259
+ ],
+ "category_id": 6,
+ "area": 325242,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20387,
+ "image_id": 3743,
+ "bbox": [
+ 0,
+ 0,
+ 101,
+ 99
+ ],
+ "category_id": 6,
+ "area": 77275,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20388,
+ "image_id": 3743,
+ "bbox": [
+ 0,
+ 93,
+ 42,
+ 196
+ ],
+ "category_id": 6,
+ "area": 64055,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20389,
+ "image_id": 3743,
+ "bbox": [
+ 65,
+ 8,
+ 172,
+ 234
+ ],
+ "category_id": 6,
+ "area": 308958,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20390,
+ "image_id": 3743,
+ "bbox": [
+ 237,
+ 36,
+ 49,
+ 106
+ ],
+ "category_id": 6,
+ "area": 40033,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20391,
+ "image_id": 3743,
+ "bbox": [
+ 265,
+ 98,
+ 92,
+ 87
+ ],
+ "category_id": 6,
+ "area": 61997,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20392,
+ "image_id": 3743,
+ "bbox": [
+ 380,
+ 91,
+ 130,
+ 117
+ ],
+ "category_id": 6,
+ "area": 117549,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20393,
+ "image_id": 3743,
+ "bbox": [
+ 434,
+ 326,
+ 77,
+ 116
+ ],
+ "category_id": 6,
+ "area": 68970,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20394,
+ "image_id": 3743,
+ "bbox": [
+ 61,
+ 227,
+ 143,
+ 105
+ ],
+ "category_id": 6,
+ "area": 115624,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20397,
+ "image_id": 3744,
+ "bbox": [
+ 239,
+ 254,
+ 106,
+ 62
+ ],
+ "category_id": 8,
+ "area": 23320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20398,
+ "image_id": 3744,
+ "bbox": [
+ 15,
+ 0,
+ 72,
+ 89
+ ],
+ "category_id": 8,
+ "area": 22500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20413,
+ "image_id": 3746,
+ "bbox": [
+ 120,
+ 104,
+ 309,
+ 407
+ ],
+ "category_id": 9,
+ "area": 249830,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20430,
+ "image_id": 3748,
+ "bbox": [
+ 214,
+ 195,
+ 225,
+ 225
+ ],
+ "category_id": 8,
+ "area": 180960,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20431,
+ "image_id": 3749,
+ "bbox": [
+ 34,
+ 38,
+ 53,
+ 116
+ ],
+ "category_id": 10,
+ "area": 48755,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20432,
+ "image_id": 3749,
+ "bbox": [
+ 207,
+ 0,
+ 140,
+ 194
+ ],
+ "category_id": 10,
+ "area": 216480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20433,
+ "image_id": 3749,
+ "bbox": [
+ 10,
+ 382,
+ 42,
+ 86
+ ],
+ "category_id": 10,
+ "area": 29280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20434,
+ "image_id": 3749,
+ "bbox": [
+ 43,
+ 279,
+ 55,
+ 90
+ ],
+ "category_id": 10,
+ "area": 39919,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20435,
+ "image_id": 3749,
+ "bbox": [
+ 139,
+ 278,
+ 71,
+ 127
+ ],
+ "category_id": 10,
+ "area": 72092,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20436,
+ "image_id": 3749,
+ "bbox": [
+ 374,
+ 319,
+ 67,
+ 95
+ ],
+ "category_id": 10,
+ "area": 50652,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20437,
+ "image_id": 3749,
+ "bbox": [
+ 331,
+ 338,
+ 40,
+ 78
+ ],
+ "category_id": 10,
+ "area": 24915,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20438,
+ "image_id": 3749,
+ "bbox": [
+ 269,
+ 427,
+ 35,
+ 60
+ ],
+ "category_id": 10,
+ "area": 16891,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20439,
+ "image_id": 3749,
+ "bbox": [
+ 193,
+ 140,
+ 56,
+ 102
+ ],
+ "category_id": 10,
+ "area": 45570,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20440,
+ "image_id": 3749,
+ "bbox": [
+ 162,
+ 135,
+ 32,
+ 56
+ ],
+ "category_id": 10,
+ "area": 14760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20444,
+ "image_id": 3750,
+ "bbox": [
+ 0,
+ 302,
+ 30,
+ 52
+ ],
+ "category_id": 6,
+ "area": 11110,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20445,
+ "image_id": 3750,
+ "bbox": [
+ 431,
+ 2,
+ 19,
+ 37
+ ],
+ "category_id": 6,
+ "area": 5214,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20446,
+ "image_id": 3750,
+ "bbox": [
+ 0,
+ 238,
+ 26,
+ 41
+ ],
+ "category_id": 6,
+ "area": 7832,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20447,
+ "image_id": 3751,
+ "bbox": [
+ 114,
+ 145,
+ 290,
+ 279
+ ],
+ "category_id": 9,
+ "area": 233583,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20452,
+ "image_id": 3753,
+ "bbox": [
+ 184,
+ 47,
+ 88,
+ 146
+ ],
+ "category_id": 8,
+ "area": 101970,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20453,
+ "image_id": 3753,
+ "bbox": [
+ 338,
+ 191,
+ 127,
+ 99
+ ],
+ "category_id": 8,
+ "area": 100170,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20454,
+ "image_id": 3753,
+ "bbox": [
+ 172,
+ 206,
+ 111,
+ 73
+ ],
+ "category_id": 8,
+ "area": 64372,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20455,
+ "image_id": 3753,
+ "bbox": [
+ 162,
+ 194,
+ 38,
+ 40
+ ],
+ "category_id": 8,
+ "area": 12240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20456,
+ "image_id": 3754,
+ "bbox": [
+ 133,
+ 88,
+ 377,
+ 246
+ ],
+ "category_id": 9,
+ "area": 143405,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20457,
+ "image_id": 3754,
+ "bbox": [
+ 40,
+ 36,
+ 87,
+ 131
+ ],
+ "category_id": 6,
+ "area": 17825,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20461,
+ "image_id": 3756,
+ "bbox": [
+ 15,
+ 196,
+ 85,
+ 156
+ ],
+ "category_id": 10,
+ "area": 20874,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20462,
+ "image_id": 3756,
+ "bbox": [
+ 234,
+ 110,
+ 128,
+ 151
+ ],
+ "category_id": 10,
+ "area": 30388,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20463,
+ "image_id": 3756,
+ "bbox": [
+ 272,
+ 247,
+ 84,
+ 116
+ ],
+ "category_id": 10,
+ "area": 15369,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20464,
+ "image_id": 3756,
+ "bbox": [
+ 402,
+ 71,
+ 83,
+ 120
+ ],
+ "category_id": 10,
+ "area": 15707,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20465,
+ "image_id": 3756,
+ "bbox": [
+ 435,
+ 326,
+ 76,
+ 170
+ ],
+ "category_id": 10,
+ "area": 20320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20466,
+ "image_id": 3757,
+ "bbox": [
+ 131,
+ 174,
+ 380,
+ 336
+ ],
+ "category_id": 10,
+ "area": 1011034,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20467,
+ "image_id": 3758,
+ "bbox": [
+ 36,
+ 225,
+ 463,
+ 286
+ ],
+ "category_id": 9,
+ "area": 830277,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20476,
+ "image_id": 3761,
+ "bbox": [
+ 19,
+ 287,
+ 122,
+ 145
+ ],
+ "category_id": 9,
+ "area": 22896,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20478,
+ "image_id": 3762,
+ "bbox": [
+ 3,
+ 130,
+ 302,
+ 380
+ ],
+ "category_id": 8,
+ "area": 286161,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20480,
+ "image_id": 3763,
+ "bbox": [
+ 158,
+ 238,
+ 107,
+ 140
+ ],
+ "category_id": 8,
+ "area": 119394,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20487,
+ "image_id": 3765,
+ "bbox": [
+ 139,
+ 53,
+ 121,
+ 37
+ ],
+ "category_id": 8,
+ "area": 9129,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20488,
+ "image_id": 3765,
+ "bbox": [
+ 169,
+ 266,
+ 209,
+ 103
+ ],
+ "category_id": 8,
+ "area": 43120,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20492,
+ "image_id": 3767,
+ "bbox": [
+ 77,
+ 138,
+ 350,
+ 324
+ ],
+ "category_id": 9,
+ "area": 340200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20493,
+ "image_id": 3768,
+ "bbox": [
+ 169,
+ 290,
+ 131,
+ 217
+ ],
+ "category_id": 10,
+ "area": 37022,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20494,
+ "image_id": 3768,
+ "bbox": [
+ 144,
+ 64,
+ 238,
+ 278
+ ],
+ "category_id": 9,
+ "area": 85748,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20495,
+ "image_id": 3769,
+ "bbox": [
+ 158,
+ 70,
+ 158,
+ 128
+ ],
+ "category_id": 9,
+ "area": 37524,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20496,
+ "image_id": 3769,
+ "bbox": [
+ 121,
+ 171,
+ 157,
+ 186
+ ],
+ "category_id": 9,
+ "area": 54288,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20499,
+ "image_id": 3770,
+ "bbox": [
+ 246,
+ 305,
+ 125,
+ 205
+ ],
+ "category_id": 8,
+ "area": 199294,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20528,
+ "image_id": 3775,
+ "bbox": [
+ 68,
+ 175,
+ 319,
+ 260
+ ],
+ "category_id": 8,
+ "area": 292068,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20529,
+ "image_id": 3775,
+ "bbox": [
+ 301,
+ 324,
+ 103,
+ 123
+ ],
+ "category_id": 8,
+ "area": 44634,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20541,
+ "image_id": 3778,
+ "bbox": [
+ 212,
+ 123,
+ 210,
+ 185
+ ],
+ "category_id": 9,
+ "area": 43921,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20542,
+ "image_id": 3778,
+ "bbox": [
+ 24,
+ 155,
+ 202,
+ 223
+ ],
+ "category_id": 6,
+ "area": 50853,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20543,
+ "image_id": 3778,
+ "bbox": [
+ 0,
+ 385,
+ 224,
+ 126
+ ],
+ "category_id": 6,
+ "area": 31920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20544,
+ "image_id": 3779,
+ "bbox": [
+ 174,
+ 247,
+ 182,
+ 163
+ ],
+ "category_id": 8,
+ "area": 39260,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20545,
+ "image_id": 3779,
+ "bbox": [
+ 198,
+ 121,
+ 75,
+ 107
+ ],
+ "category_id": 8,
+ "area": 10836,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20546,
+ "image_id": 3780,
+ "bbox": [
+ 62,
+ 153,
+ 259,
+ 267
+ ],
+ "category_id": 9,
+ "area": 1140942,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20551,
+ "image_id": 3781,
+ "bbox": [
+ 1,
+ 165,
+ 38,
+ 143
+ ],
+ "category_id": 6,
+ "area": 27261,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20552,
+ "image_id": 3781,
+ "bbox": [
+ 0,
+ 319,
+ 94,
+ 168
+ ],
+ "category_id": 6,
+ "area": 77805,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20553,
+ "image_id": 3781,
+ "bbox": [
+ 176,
+ 399,
+ 238,
+ 112
+ ],
+ "category_id": 6,
+ "area": 132126,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20554,
+ "image_id": 3781,
+ "bbox": [
+ 58,
+ 285,
+ 65,
+ 76
+ ],
+ "category_id": 6,
+ "area": 24676,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20558,
+ "image_id": 3783,
+ "bbox": [
+ 291,
+ 163,
+ 108,
+ 117
+ ],
+ "category_id": 9,
+ "area": 46736,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20559,
+ "image_id": 3783,
+ "bbox": [
+ 55,
+ 103,
+ 339,
+ 283
+ ],
+ "category_id": 9,
+ "area": 352185,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20600,
+ "image_id": 3787,
+ "bbox": [
+ 163,
+ 147,
+ 178,
+ 103
+ ],
+ "category_id": 8,
+ "area": 64525,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20601,
+ "image_id": 3787,
+ "bbox": [
+ 218,
+ 252,
+ 157,
+ 169
+ ],
+ "category_id": 8,
+ "area": 93772,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20627,
+ "image_id": 3789,
+ "bbox": [
+ 168,
+ 106,
+ 188,
+ 405
+ ],
+ "category_id": 10,
+ "area": 605340,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20628,
+ "image_id": 3789,
+ "bbox": [
+ 87,
+ 0,
+ 240,
+ 170
+ ],
+ "category_id": 10,
+ "area": 323459,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20629,
+ "image_id": 3790,
+ "bbox": [
+ 2,
+ 68,
+ 120,
+ 176
+ ],
+ "category_id": 6,
+ "area": 218564,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20630,
+ "image_id": 3790,
+ "bbox": [
+ 53,
+ 2,
+ 116,
+ 78
+ ],
+ "category_id": 6,
+ "area": 93210,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20631,
+ "image_id": 3790,
+ "bbox": [
+ 5,
+ 245,
+ 149,
+ 111
+ ],
+ "category_id": 6,
+ "area": 170000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20632,
+ "image_id": 3790,
+ "bbox": [
+ 4,
+ 362,
+ 92,
+ 67
+ ],
+ "category_id": 6,
+ "area": 64064,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20633,
+ "image_id": 3790,
+ "bbox": [
+ 140,
+ 2,
+ 142,
+ 204
+ ],
+ "category_id": 6,
+ "area": 298602,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20634,
+ "image_id": 3790,
+ "bbox": [
+ 141,
+ 193,
+ 117,
+ 89
+ ],
+ "category_id": 6,
+ "area": 107682,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20635,
+ "image_id": 3790,
+ "bbox": [
+ 173,
+ 276,
+ 176,
+ 97
+ ],
+ "category_id": 6,
+ "area": 175522,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20636,
+ "image_id": 3790,
+ "bbox": [
+ 283,
+ 16,
+ 41,
+ 92
+ ],
+ "category_id": 6,
+ "area": 39337,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20637,
+ "image_id": 3790,
+ "bbox": [
+ 305,
+ 75,
+ 88,
+ 70
+ ],
+ "category_id": 6,
+ "area": 63936,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20638,
+ "image_id": 3790,
+ "bbox": [
+ 397,
+ 59,
+ 114,
+ 114
+ ],
+ "category_id": 6,
+ "area": 133318,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20639,
+ "image_id": 3790,
+ "bbox": [
+ 448,
+ 274,
+ 59,
+ 78
+ ],
+ "category_id": 6,
+ "area": 47520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20643,
+ "image_id": 3791,
+ "bbox": [
+ 168,
+ 200,
+ 83,
+ 163
+ ],
+ "category_id": 10,
+ "area": 24336,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20645,
+ "image_id": 3792,
+ "bbox": [
+ 0,
+ 0,
+ 512,
+ 510
+ ],
+ "category_id": 6,
+ "area": 740086,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20648,
+ "image_id": 3794,
+ "bbox": [
+ 25,
+ 38,
+ 285,
+ 400
+ ],
+ "category_id": 9,
+ "area": 1534794,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20649,
+ "image_id": 3794,
+ "bbox": [
+ 1,
+ 309,
+ 111,
+ 146
+ ],
+ "category_id": 6,
+ "area": 220038,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20650,
+ "image_id": 3794,
+ "bbox": [
+ 96,
+ 396,
+ 95,
+ 115
+ ],
+ "category_id": 6,
+ "area": 147552,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20651,
+ "image_id": 3794,
+ "bbox": [
+ 61,
+ 231,
+ 59,
+ 75
+ ],
+ "category_id": 6,
+ "area": 60528,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20652,
+ "image_id": 3794,
+ "bbox": [
+ 121,
+ 227,
+ 34,
+ 51
+ ],
+ "category_id": 6,
+ "area": 23970,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20653,
+ "image_id": 3794,
+ "bbox": [
+ 99,
+ 202,
+ 34,
+ 51
+ ],
+ "category_id": 6,
+ "area": 23970,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20654,
+ "image_id": 3794,
+ "bbox": [
+ 8,
+ 224,
+ 52,
+ 52
+ ],
+ "category_id": 6,
+ "area": 36576,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20655,
+ "image_id": 3794,
+ "bbox": [
+ 117,
+ 311,
+ 43,
+ 79
+ ],
+ "category_id": 6,
+ "area": 46209,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20656,
+ "image_id": 3794,
+ "bbox": [
+ 169,
+ 334,
+ 47,
+ 68
+ ],
+ "category_id": 6,
+ "area": 43470,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20657,
+ "image_id": 3794,
+ "bbox": [
+ 237,
+ 466,
+ 44,
+ 45
+ ],
+ "category_id": 6,
+ "area": 27000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20659,
+ "image_id": 3796,
+ "bbox": [
+ 57,
+ 188,
+ 198,
+ 210
+ ],
+ "category_id": 8,
+ "area": 146320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20660,
+ "image_id": 3796,
+ "bbox": [
+ 238,
+ 191,
+ 234,
+ 249
+ ],
+ "category_id": 8,
+ "area": 204863,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20666,
+ "image_id": 3798,
+ "bbox": [
+ 201,
+ 6,
+ 84,
+ 497
+ ],
+ "category_id": 10,
+ "area": 333582,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20689,
+ "image_id": 3800,
+ "bbox": [
+ 1,
+ 162,
+ 370,
+ 344
+ ],
+ "category_id": 9,
+ "area": 448625,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20690,
+ "image_id": 3800,
+ "bbox": [
+ 172,
+ 2,
+ 336,
+ 140
+ ],
+ "category_id": 9,
+ "area": 165874,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20691,
+ "image_id": 3800,
+ "bbox": [
+ 174,
+ 64,
+ 334,
+ 324
+ ],
+ "category_id": 9,
+ "area": 382052,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20741,
+ "image_id": 3808,
+ "bbox": [
+ 136,
+ 6,
+ 159,
+ 232
+ ],
+ "category_id": 8,
+ "area": 130146,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20742,
+ "image_id": 3808,
+ "bbox": [
+ 146,
+ 106,
+ 281,
+ 285
+ ],
+ "category_id": 8,
+ "area": 281903,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20743,
+ "image_id": 3808,
+ "bbox": [
+ 233,
+ 268,
+ 273,
+ 242
+ ],
+ "category_id": 8,
+ "area": 232560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20748,
+ "image_id": 3810,
+ "bbox": [
+ 68,
+ 65,
+ 66,
+ 109
+ ],
+ "category_id": 6,
+ "area": 8774,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20749,
+ "image_id": 3810,
+ "bbox": [
+ 128,
+ 67,
+ 71,
+ 132
+ ],
+ "category_id": 6,
+ "area": 11484,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20750,
+ "image_id": 3810,
+ "bbox": [
+ 207,
+ 95,
+ 69,
+ 77
+ ],
+ "category_id": 6,
+ "area": 6496,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20751,
+ "image_id": 3810,
+ "bbox": [
+ 264,
+ 309,
+ 105,
+ 178
+ ],
+ "category_id": 6,
+ "area": 22610,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20752,
+ "image_id": 3810,
+ "bbox": [
+ 245,
+ 360,
+ 39,
+ 117
+ ],
+ "category_id": 6,
+ "area": 5544,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20753,
+ "image_id": 3810,
+ "bbox": [
+ 355,
+ 274,
+ 27,
+ 73
+ ],
+ "category_id": 6,
+ "area": 2475,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20754,
+ "image_id": 3810,
+ "bbox": [
+ 349,
+ 323,
+ 96,
+ 188
+ ],
+ "category_id": 6,
+ "area": 21996,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20755,
+ "image_id": 3810,
+ "bbox": [
+ 323,
+ 1,
+ 96,
+ 100
+ ],
+ "category_id": 6,
+ "area": 11625,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20756,
+ "image_id": 3810,
+ "bbox": [
+ 399,
+ 233,
+ 43,
+ 69
+ ],
+ "category_id": 6,
+ "area": 3640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20757,
+ "image_id": 3810,
+ "bbox": [
+ 442,
+ 221,
+ 48,
+ 104
+ ],
+ "category_id": 6,
+ "area": 6084,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20758,
+ "image_id": 3810,
+ "bbox": [
+ 411,
+ 37,
+ 70,
+ 217
+ ],
+ "category_id": 6,
+ "area": 18306,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20759,
+ "image_id": 3810,
+ "bbox": [
+ 484,
+ 0,
+ 27,
+ 260
+ ],
+ "category_id": 6,
+ "area": 8730,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20760,
+ "image_id": 3810,
+ "bbox": [
+ 0,
+ 396,
+ 29,
+ 93
+ ],
+ "category_id": 6,
+ "area": 3360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20761,
+ "image_id": 3811,
+ "bbox": [
+ 123,
+ 296,
+ 106,
+ 215
+ ],
+ "category_id": 6,
+ "area": 46360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20762,
+ "image_id": 3811,
+ "bbox": [
+ 200,
+ 218,
+ 311,
+ 292
+ ],
+ "category_id": 9,
+ "area": 183928,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20763,
+ "image_id": 3812,
+ "bbox": [
+ 71,
+ 145,
+ 310,
+ 193
+ ],
+ "category_id": 9,
+ "area": 154308,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20765,
+ "image_id": 3812,
+ "bbox": [
+ 410,
+ 331,
+ 101,
+ 179
+ ],
+ "category_id": 6,
+ "area": 46652,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20777,
+ "image_id": 3814,
+ "bbox": [
+ 149,
+ 31,
+ 173,
+ 216
+ ],
+ "category_id": 9,
+ "area": 131936,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20778,
+ "image_id": 3814,
+ "bbox": [
+ 243,
+ 91,
+ 250,
+ 418
+ ],
+ "category_id": 9,
+ "area": 368714,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20779,
+ "image_id": 3814,
+ "bbox": [
+ 139,
+ 187,
+ 318,
+ 319
+ ],
+ "category_id": 9,
+ "area": 356955,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20780,
+ "image_id": 3815,
+ "bbox": [
+ 2,
+ 2,
+ 226,
+ 362
+ ],
+ "category_id": 6,
+ "area": 162837,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20781,
+ "image_id": 3815,
+ "bbox": [
+ 277,
+ 0,
+ 233,
+ 295
+ ],
+ "category_id": 6,
+ "area": 137256,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20782,
+ "image_id": 3815,
+ "bbox": [
+ 230,
+ 315,
+ 200,
+ 108
+ ],
+ "category_id": 8,
+ "area": 43070,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20844,
+ "image_id": 3817,
+ "bbox": [
+ 88,
+ 256,
+ 154,
+ 100
+ ],
+ "category_id": 8,
+ "area": 123540,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20845,
+ "image_id": 3817,
+ "bbox": [
+ 154,
+ 193,
+ 109,
+ 106
+ ],
+ "category_id": 8,
+ "area": 92250,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20846,
+ "image_id": 3817,
+ "bbox": [
+ 233,
+ 144,
+ 127,
+ 142
+ ],
+ "category_id": 8,
+ "area": 143700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20847,
+ "image_id": 3817,
+ "bbox": [
+ 327,
+ 156,
+ 120,
+ 71
+ ],
+ "category_id": 8,
+ "area": 68403,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20848,
+ "image_id": 3818,
+ "bbox": [
+ 102,
+ 212,
+ 159,
+ 156
+ ],
+ "category_id": 8,
+ "area": 87162,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20849,
+ "image_id": 3818,
+ "bbox": [
+ 210,
+ 213,
+ 200,
+ 156
+ ],
+ "category_id": 8,
+ "area": 109500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20876,
+ "image_id": 3821,
+ "bbox": [
+ 206,
+ 164,
+ 170,
+ 274
+ ],
+ "category_id": 10,
+ "area": 54692,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20878,
+ "image_id": 3822,
+ "bbox": [
+ 28,
+ 91,
+ 459,
+ 311
+ ],
+ "category_id": 9,
+ "area": 123994,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20879,
+ "image_id": 3823,
+ "bbox": [
+ 34,
+ 254,
+ 125,
+ 101
+ ],
+ "category_id": 9,
+ "area": 21131,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20884,
+ "image_id": 3824,
+ "bbox": [
+ 76,
+ 154,
+ 275,
+ 357
+ ],
+ "category_id": 9,
+ "area": 281064,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20886,
+ "image_id": 3824,
+ "bbox": [
+ 200,
+ 178,
+ 34,
+ 92
+ ],
+ "category_id": 10,
+ "area": 9027,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20887,
+ "image_id": 3825,
+ "bbox": [
+ 17,
+ 253,
+ 100,
+ 142
+ ],
+ "category_id": 6,
+ "area": 13098,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20888,
+ "image_id": 3825,
+ "bbox": [
+ 126,
+ 257,
+ 120,
+ 158
+ ],
+ "category_id": 6,
+ "area": 17484,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20889,
+ "image_id": 3825,
+ "bbox": [
+ 311,
+ 243,
+ 70,
+ 144
+ ],
+ "category_id": 6,
+ "area": 9379,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20890,
+ "image_id": 3825,
+ "bbox": [
+ 360,
+ 281,
+ 139,
+ 217
+ ],
+ "category_id": 6,
+ "area": 27710,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20891,
+ "image_id": 3825,
+ "bbox": [
+ 187,
+ 354,
+ 124,
+ 157
+ ],
+ "category_id": 6,
+ "area": 17958,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20892,
+ "image_id": 3825,
+ "bbox": [
+ 197,
+ 188,
+ 102,
+ 162
+ ],
+ "category_id": 6,
+ "area": 15240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20895,
+ "image_id": 3827,
+ "bbox": [
+ 220,
+ 53,
+ 55,
+ 125
+ ],
+ "category_id": 8,
+ "area": 54855,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20896,
+ "image_id": 3827,
+ "bbox": [
+ 266,
+ 40,
+ 96,
+ 128
+ ],
+ "category_id": 6,
+ "area": 98192,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20965,
+ "image_id": 3833,
+ "bbox": [
+ 1,
+ 1,
+ 226,
+ 467
+ ],
+ "category_id": 10,
+ "area": 839937,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20966,
+ "image_id": 3833,
+ "bbox": [
+ 170,
+ 180,
+ 161,
+ 331
+ ],
+ "category_id": 10,
+ "area": 423594,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20986,
+ "image_id": 3835,
+ "bbox": [
+ 0,
+ 0,
+ 232,
+ 248
+ ],
+ "category_id": 8,
+ "area": 147560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20987,
+ "image_id": 3835,
+ "bbox": [
+ 254,
+ 222,
+ 257,
+ 166
+ ],
+ "category_id": 8,
+ "area": 109212,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20991,
+ "image_id": 3837,
+ "bbox": [
+ 75,
+ 189,
+ 158,
+ 102
+ ],
+ "category_id": 8,
+ "area": 57024,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20992,
+ "image_id": 3837,
+ "bbox": [
+ 216,
+ 250,
+ 148,
+ 173
+ ],
+ "category_id": 8,
+ "area": 90396,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20996,
+ "image_id": 3839,
+ "bbox": [
+ 186,
+ 158,
+ 325,
+ 329
+ ],
+ "category_id": 9,
+ "area": 180790,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20997,
+ "image_id": 3839,
+ "bbox": [
+ 145,
+ 242,
+ 366,
+ 268
+ ],
+ "category_id": 6,
+ "area": 165856,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 20998,
+ "image_id": 3839,
+ "bbox": [
+ 429,
+ 88,
+ 82,
+ 119
+ ],
+ "category_id": 6,
+ "area": 16640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21009,
+ "image_id": 3841,
+ "bbox": [
+ 33,
+ 83,
+ 478,
+ 425
+ ],
+ "category_id": 9,
+ "area": 491305,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21010,
+ "image_id": 3841,
+ "bbox": [
+ 0,
+ 36,
+ 255,
+ 470
+ ],
+ "category_id": 6,
+ "area": 290512,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21011,
+ "image_id": 3841,
+ "bbox": [
+ 0,
+ 0,
+ 508,
+ 511
+ ],
+ "category_id": 6,
+ "area": 627396,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21017,
+ "image_id": 3843,
+ "bbox": [
+ 267,
+ 248,
+ 244,
+ 260
+ ],
+ "category_id": 9,
+ "area": 223992,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21018,
+ "image_id": 3843,
+ "bbox": [
+ 282,
+ 27,
+ 229,
+ 477
+ ],
+ "category_id": 9,
+ "area": 385154,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21020,
+ "image_id": 3845,
+ "bbox": [
+ 0,
+ 2,
+ 117,
+ 123
+ ],
+ "category_id": 6,
+ "area": 125307,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21021,
+ "image_id": 3845,
+ "bbox": [
+ 0,
+ 116,
+ 42,
+ 95
+ ],
+ "category_id": 6,
+ "area": 35490,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21022,
+ "image_id": 3845,
+ "bbox": [
+ 60,
+ 86,
+ 71,
+ 75
+ ],
+ "category_id": 6,
+ "area": 46652,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21023,
+ "image_id": 3845,
+ "bbox": [
+ 180,
+ 20,
+ 120,
+ 120
+ ],
+ "category_id": 6,
+ "area": 125216,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21024,
+ "image_id": 3845,
+ "bbox": [
+ 252,
+ 0,
+ 95,
+ 108
+ ],
+ "category_id": 6,
+ "area": 89301,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21025,
+ "image_id": 3845,
+ "bbox": [
+ 103,
+ 147,
+ 60,
+ 109
+ ],
+ "category_id": 6,
+ "area": 57279,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21026,
+ "image_id": 3845,
+ "bbox": [
+ 180,
+ 171,
+ 85,
+ 106
+ ],
+ "category_id": 6,
+ "area": 78174,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21027,
+ "image_id": 3845,
+ "bbox": [
+ 192,
+ 341,
+ 99,
+ 98
+ ],
+ "category_id": 6,
+ "area": 84581,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21028,
+ "image_id": 3845,
+ "bbox": [
+ 313,
+ 285,
+ 163,
+ 169
+ ],
+ "category_id": 6,
+ "area": 240051,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21029,
+ "image_id": 3845,
+ "bbox": [
+ 386,
+ 88,
+ 89,
+ 67
+ ],
+ "category_id": 6,
+ "area": 52110,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21036,
+ "image_id": 3846,
+ "bbox": [
+ 298,
+ 69,
+ 135,
+ 238
+ ],
+ "category_id": 9,
+ "area": 113568,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21037,
+ "image_id": 3846,
+ "bbox": [
+ 281,
+ 146,
+ 114,
+ 196
+ ],
+ "category_id": 9,
+ "area": 79222,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21038,
+ "image_id": 3846,
+ "bbox": [
+ 400,
+ 183,
+ 108,
+ 173
+ ],
+ "category_id": 9,
+ "area": 66124,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21039,
+ "image_id": 3846,
+ "bbox": [
+ 1,
+ 177,
+ 114,
+ 327
+ ],
+ "category_id": 9,
+ "area": 131846,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21044,
+ "image_id": 3848,
+ "bbox": [
+ 142,
+ 225,
+ 369,
+ 237
+ ],
+ "category_id": 8,
+ "area": 204227,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21049,
+ "image_id": 3849,
+ "bbox": [
+ 92,
+ 333,
+ 31,
+ 55
+ ],
+ "category_id": 6,
+ "area": 11817,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21050,
+ "image_id": 3849,
+ "bbox": [
+ 388,
+ 407,
+ 31,
+ 58
+ ],
+ "category_id": 6,
+ "area": 12524,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21068,
+ "image_id": 3853,
+ "bbox": [
+ 266,
+ 204,
+ 242,
+ 301
+ ],
+ "category_id": 9,
+ "area": 257368,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21069,
+ "image_id": 3853,
+ "bbox": [
+ 255,
+ 49,
+ 255,
+ 329
+ ],
+ "category_id": 9,
+ "area": 295857,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21070,
+ "image_id": 3853,
+ "bbox": [
+ 207,
+ 74,
+ 146,
+ 251
+ ],
+ "category_id": 9,
+ "area": 129918,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21071,
+ "image_id": 3854,
+ "bbox": [
+ 377,
+ 406,
+ 100,
+ 104
+ ],
+ "category_id": 9,
+ "area": 24384,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21072,
+ "image_id": 3854,
+ "bbox": [
+ 20,
+ 131,
+ 164,
+ 248
+ ],
+ "category_id": 9,
+ "area": 94514,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21073,
+ "image_id": 3854,
+ "bbox": [
+ 150,
+ 309,
+ 182,
+ 202
+ ],
+ "category_id": 9,
+ "area": 85608,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21074,
+ "image_id": 3854,
+ "bbox": [
+ 162,
+ 238,
+ 92,
+ 159
+ ],
+ "category_id": 9,
+ "area": 34144,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21075,
+ "image_id": 3854,
+ "bbox": [
+ 229,
+ 73,
+ 148,
+ 234
+ ],
+ "category_id": 9,
+ "area": 80655,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21076,
+ "image_id": 3854,
+ "bbox": [
+ 275,
+ 131,
+ 134,
+ 225
+ ],
+ "category_id": 9,
+ "area": 70418,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21087,
+ "image_id": 3858,
+ "bbox": [
+ 0,
+ 0,
+ 137,
+ 510
+ ],
+ "category_id": 6,
+ "area": 245960,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21088,
+ "image_id": 3858,
+ "bbox": [
+ 118,
+ 260,
+ 244,
+ 132
+ ],
+ "category_id": 8,
+ "area": 113646,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21089,
+ "image_id": 3858,
+ "bbox": [
+ 321,
+ 113,
+ 189,
+ 330
+ ],
+ "category_id": 8,
+ "area": 219462,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21125,
+ "image_id": 3863,
+ "bbox": [
+ 190,
+ 373,
+ 66,
+ 138
+ ],
+ "category_id": 8,
+ "area": 29880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21126,
+ "image_id": 3863,
+ "bbox": [
+ 246,
+ 73,
+ 77,
+ 108
+ ],
+ "category_id": 8,
+ "area": 27213,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21127,
+ "image_id": 3864,
+ "bbox": [
+ 2,
+ 0,
+ 386,
+ 494
+ ],
+ "category_id": 9,
+ "area": 673032,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21128,
+ "image_id": 3864,
+ "bbox": [
+ 321,
+ 0,
+ 188,
+ 506
+ ],
+ "category_id": 9,
+ "area": 334640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21129,
+ "image_id": 3864,
+ "bbox": [
+ 3,
+ 142,
+ 337,
+ 365
+ ],
+ "category_id": 9,
+ "area": 433816,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21142,
+ "image_id": 3866,
+ "bbox": [
+ 115,
+ 105,
+ 212,
+ 253
+ ],
+ "category_id": 8,
+ "area": 426132,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21143,
+ "image_id": 3866,
+ "bbox": [
+ 279,
+ 106,
+ 123,
+ 213
+ ],
+ "category_id": 8,
+ "area": 208813,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21145,
+ "image_id": 3867,
+ "bbox": [
+ 186,
+ 269,
+ 95,
+ 67
+ ],
+ "category_id": 8,
+ "area": 22466,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21146,
+ "image_id": 3867,
+ "bbox": [
+ 283,
+ 201,
+ 93,
+ 157
+ ],
+ "category_id": 8,
+ "area": 51714,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21192,
+ "image_id": 3870,
+ "bbox": [
+ 247,
+ 351,
+ 84,
+ 80
+ ],
+ "category_id": 9,
+ "area": 11659,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21193,
+ "image_id": 3870,
+ "bbox": [
+ 415,
+ 307,
+ 96,
+ 68
+ ],
+ "category_id": 9,
+ "area": 11324,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21202,
+ "image_id": 3872,
+ "bbox": [
+ 30,
+ 277,
+ 99,
+ 86
+ ],
+ "category_id": 6,
+ "area": 66767,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21203,
+ "image_id": 3872,
+ "bbox": [
+ 101,
+ 115,
+ 57,
+ 101
+ ],
+ "category_id": 6,
+ "area": 44935,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21204,
+ "image_id": 3872,
+ "bbox": [
+ 287,
+ 159,
+ 49,
+ 81
+ ],
+ "category_id": 6,
+ "area": 30912,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21205,
+ "image_id": 3872,
+ "bbox": [
+ 402,
+ 105,
+ 42,
+ 75
+ ],
+ "category_id": 6,
+ "area": 24645,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21208,
+ "image_id": 3873,
+ "bbox": [
+ 248,
+ 45,
+ 204,
+ 220
+ ],
+ "category_id": 9,
+ "area": 105050,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21209,
+ "image_id": 3873,
+ "bbox": [
+ 247,
+ 134,
+ 141,
+ 166
+ ],
+ "category_id": 10,
+ "area": 54912,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21211,
+ "image_id": 3875,
+ "bbox": [
+ 235,
+ 56,
+ 223,
+ 398
+ ],
+ "category_id": 9,
+ "area": 313040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21212,
+ "image_id": 3875,
+ "bbox": [
+ 146,
+ 142,
+ 288,
+ 339
+ ],
+ "category_id": 9,
+ "area": 343917,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21213,
+ "image_id": 3876,
+ "bbox": [
+ 38,
+ 275,
+ 301,
+ 145
+ ],
+ "category_id": 8,
+ "area": 111153,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21214,
+ "image_id": 3876,
+ "bbox": [
+ 74,
+ 41,
+ 437,
+ 415
+ ],
+ "category_id": 8,
+ "area": 462528,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21215,
+ "image_id": 3877,
+ "bbox": [
+ 154,
+ 129,
+ 237,
+ 216
+ ],
+ "category_id": 9,
+ "area": 112965,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21221,
+ "image_id": 3879,
+ "bbox": [
+ 329,
+ 297,
+ 70,
+ 51
+ ],
+ "category_id": 9,
+ "area": 9620,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21222,
+ "image_id": 3880,
+ "bbox": [
+ 44,
+ 1,
+ 182,
+ 350
+ ],
+ "category_id": 10,
+ "area": 153244,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21223,
+ "image_id": 3880,
+ "bbox": [
+ 217,
+ 47,
+ 288,
+ 464
+ ],
+ "category_id": 9,
+ "area": 321966,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21224,
+ "image_id": 3881,
+ "bbox": [
+ 156,
+ 218,
+ 170,
+ 235
+ ],
+ "category_id": 8,
+ "area": 139496,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21225,
+ "image_id": 3881,
+ "bbox": [
+ 219,
+ 49,
+ 292,
+ 457
+ ],
+ "category_id": 8,
+ "area": 465280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21226,
+ "image_id": 3881,
+ "bbox": [
+ 2,
+ 0,
+ 365,
+ 510
+ ],
+ "category_id": 6,
+ "area": 649026,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21227,
+ "image_id": 3882,
+ "bbox": [
+ 6,
+ 191,
+ 257,
+ 255
+ ],
+ "category_id": 9,
+ "area": 230837,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21228,
+ "image_id": 3882,
+ "bbox": [
+ 256,
+ 168,
+ 101,
+ 115
+ ],
+ "category_id": 9,
+ "area": 41239,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21229,
+ "image_id": 3882,
+ "bbox": [
+ 255,
+ 240,
+ 104,
+ 218
+ ],
+ "category_id": 9,
+ "area": 80127,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21232,
+ "image_id": 3884,
+ "bbox": [
+ 184,
+ 211,
+ 103,
+ 92
+ ],
+ "category_id": 8,
+ "area": 64512,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21233,
+ "image_id": 3884,
+ "bbox": [
+ 65,
+ 180,
+ 79,
+ 97
+ ],
+ "category_id": 6,
+ "area": 52569,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21234,
+ "image_id": 3884,
+ "bbox": [
+ 45,
+ 244,
+ 114,
+ 133
+ ],
+ "category_id": 6,
+ "area": 103334,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21235,
+ "image_id": 3884,
+ "bbox": [
+ 149,
+ 393,
+ 336,
+ 114
+ ],
+ "category_id": 6,
+ "area": 259785,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21237,
+ "image_id": 3885,
+ "bbox": [
+ 255,
+ 253,
+ 256,
+ 256
+ ],
+ "category_id": 9,
+ "area": 170181,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21238,
+ "image_id": 3885,
+ "bbox": [
+ 0,
+ 120,
+ 266,
+ 277
+ ],
+ "category_id": 9,
+ "area": 192234,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21239,
+ "image_id": 3885,
+ "bbox": [
+ 236,
+ 89,
+ 155,
+ 262
+ ],
+ "category_id": 9,
+ "area": 105488,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21241,
+ "image_id": 3886,
+ "bbox": [
+ 0,
+ 200,
+ 404,
+ 269
+ ],
+ "category_id": 9,
+ "area": 106434,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21242,
+ "image_id": 3886,
+ "bbox": [
+ 302,
+ 395,
+ 112,
+ 66
+ ],
+ "category_id": 9,
+ "area": 7320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21243,
+ "image_id": 3886,
+ "bbox": [
+ 396,
+ 224,
+ 92,
+ 87
+ ],
+ "category_id": 9,
+ "area": 7900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21244,
+ "image_id": 3886,
+ "bbox": [
+ 0,
+ 416,
+ 48,
+ 78
+ ],
+ "category_id": 9,
+ "area": 3763,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21246,
+ "image_id": 3887,
+ "bbox": [
+ 35,
+ 88,
+ 420,
+ 320
+ ],
+ "category_id": 9,
+ "area": 157500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21247,
+ "image_id": 3888,
+ "bbox": [
+ 14,
+ 21,
+ 467,
+ 479
+ ],
+ "category_id": 9,
+ "area": 396262,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21248,
+ "image_id": 3889,
+ "bbox": [
+ 176,
+ 139,
+ 135,
+ 140
+ ],
+ "category_id": 9,
+ "area": 114048,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21249,
+ "image_id": 3890,
+ "bbox": [
+ 129,
+ 369,
+ 55,
+ 37
+ ],
+ "category_id": 9,
+ "area": 2695,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21250,
+ "image_id": 3890,
+ "bbox": [
+ 207,
+ 133,
+ 159,
+ 228
+ ],
+ "category_id": 9,
+ "area": 47952,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21251,
+ "image_id": 3891,
+ "bbox": [
+ 3,
+ 77,
+ 95,
+ 82
+ ],
+ "category_id": 6,
+ "area": 196773,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21252,
+ "image_id": 3891,
+ "bbox": [
+ 2,
+ 140,
+ 44,
+ 86
+ ],
+ "category_id": 6,
+ "area": 96719,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21253,
+ "image_id": 3891,
+ "bbox": [
+ 21,
+ 263,
+ 37,
+ 74
+ ],
+ "category_id": 6,
+ "area": 69258,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21254,
+ "image_id": 3891,
+ "bbox": [
+ 65,
+ 280,
+ 39,
+ 78
+ ],
+ "category_id": 6,
+ "area": 77112,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21255,
+ "image_id": 3891,
+ "bbox": [
+ 70,
+ 400,
+ 47,
+ 77
+ ],
+ "category_id": 6,
+ "area": 91203,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21256,
+ "image_id": 3891,
+ "bbox": [
+ 137,
+ 350,
+ 70,
+ 135
+ ],
+ "category_id": 6,
+ "area": 237600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21257,
+ "image_id": 3891,
+ "bbox": [
+ 47,
+ 62,
+ 63,
+ 69
+ ],
+ "category_id": 6,
+ "area": 111792,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21258,
+ "image_id": 3891,
+ "bbox": [
+ 123,
+ 56,
+ 55,
+ 48
+ ],
+ "category_id": 6,
+ "area": 67473,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21259,
+ "image_id": 3891,
+ "bbox": [
+ 100,
+ 118,
+ 40,
+ 57
+ ],
+ "category_id": 6,
+ "area": 58950,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21260,
+ "image_id": 3891,
+ "bbox": [
+ 103,
+ 131,
+ 40,
+ 91
+ ],
+ "category_id": 6,
+ "area": 93534,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21261,
+ "image_id": 3891,
+ "bbox": [
+ 148,
+ 44,
+ 139,
+ 131
+ ],
+ "category_id": 6,
+ "area": 458878,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21262,
+ "image_id": 3891,
+ "bbox": [
+ 161,
+ 215,
+ 40,
+ 45
+ ],
+ "category_id": 6,
+ "area": 46361,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21263,
+ "image_id": 3891,
+ "bbox": [
+ 246,
+ 41,
+ 41,
+ 45
+ ],
+ "category_id": 6,
+ "area": 47256,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21264,
+ "image_id": 3891,
+ "bbox": [
+ 293,
+ 46,
+ 65,
+ 78
+ ],
+ "category_id": 6,
+ "area": 127908,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21265,
+ "image_id": 3891,
+ "bbox": [
+ 260,
+ 116,
+ 72,
+ 138
+ ],
+ "category_id": 6,
+ "area": 252030,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21283,
+ "image_id": 3892,
+ "bbox": [
+ 94,
+ 165,
+ 83,
+ 152
+ ],
+ "category_id": 10,
+ "area": 46605,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21284,
+ "image_id": 3892,
+ "bbox": [
+ 148,
+ 117,
+ 165,
+ 168
+ ],
+ "category_id": 9,
+ "area": 102044,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21285,
+ "image_id": 3892,
+ "bbox": [
+ 241,
+ 0,
+ 269,
+ 511
+ ],
+ "category_id": 9,
+ "area": 504169,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21287,
+ "image_id": 3894,
+ "bbox": [
+ 0,
+ 0,
+ 365,
+ 511
+ ],
+ "category_id": 9,
+ "area": 657166,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21305,
+ "image_id": 3896,
+ "bbox": [
+ 40,
+ 75,
+ 23,
+ 112
+ ],
+ "category_id": 6,
+ "area": 24480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21306,
+ "image_id": 3896,
+ "bbox": [
+ 50,
+ 105,
+ 204,
+ 124
+ ],
+ "category_id": 6,
+ "area": 232368,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21307,
+ "image_id": 3896,
+ "bbox": [
+ 0,
+ 168,
+ 47,
+ 68
+ ],
+ "category_id": 6,
+ "area": 29536,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21308,
+ "image_id": 3896,
+ "bbox": [
+ 1,
+ 210,
+ 114,
+ 92
+ ],
+ "category_id": 6,
+ "area": 97226,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21309,
+ "image_id": 3896,
+ "bbox": [
+ 103,
+ 187,
+ 86,
+ 173
+ ],
+ "category_id": 6,
+ "area": 136764,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21310,
+ "image_id": 3896,
+ "bbox": [
+ 182,
+ 226,
+ 116,
+ 118
+ ],
+ "category_id": 6,
+ "area": 125307,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21311,
+ "image_id": 3896,
+ "bbox": [
+ 255,
+ 181,
+ 94,
+ 119
+ ],
+ "category_id": 6,
+ "area": 102885,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21312,
+ "image_id": 3896,
+ "bbox": [
+ 144,
+ 85,
+ 135,
+ 89
+ ],
+ "category_id": 6,
+ "area": 111248,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21313,
+ "image_id": 3896,
+ "bbox": [
+ 235,
+ 163,
+ 112,
+ 86
+ ],
+ "category_id": 6,
+ "area": 88894,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21314,
+ "image_id": 3896,
+ "bbox": [
+ 302,
+ 81,
+ 68,
+ 61
+ ],
+ "category_id": 6,
+ "area": 38316,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21315,
+ "image_id": 3896,
+ "bbox": [
+ 359,
+ 66,
+ 152,
+ 171
+ ],
+ "category_id": 6,
+ "area": 239720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21316,
+ "image_id": 3896,
+ "bbox": [
+ 0,
+ 316,
+ 57,
+ 94
+ ],
+ "category_id": 6,
+ "area": 49478,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21317,
+ "image_id": 3896,
+ "bbox": [
+ 93,
+ 343,
+ 74,
+ 115
+ ],
+ "category_id": 6,
+ "area": 78874,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21318,
+ "image_id": 3896,
+ "bbox": [
+ 180,
+ 369,
+ 82,
+ 104
+ ],
+ "category_id": 6,
+ "area": 79000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21319,
+ "image_id": 3896,
+ "bbox": [
+ 390,
+ 290,
+ 88,
+ 57
+ ],
+ "category_id": 6,
+ "area": 46284,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21321,
+ "image_id": 3897,
+ "bbox": [
+ 26,
+ 274,
+ 257,
+ 236
+ ],
+ "category_id": 8,
+ "area": 481068,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21322,
+ "image_id": 3897,
+ "bbox": [
+ 174,
+ 238,
+ 144,
+ 135
+ ],
+ "category_id": 8,
+ "area": 154470,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21327,
+ "image_id": 3899,
+ "bbox": [
+ 117,
+ 149,
+ 227,
+ 167
+ ],
+ "category_id": 10,
+ "area": 59899,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21329,
+ "image_id": 3900,
+ "bbox": [
+ 168,
+ 189,
+ 124,
+ 184
+ ],
+ "category_id": 8,
+ "area": 79980,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21330,
+ "image_id": 3900,
+ "bbox": [
+ 252,
+ 147,
+ 150,
+ 165
+ ],
+ "category_id": 8,
+ "area": 87464,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21331,
+ "image_id": 3901,
+ "bbox": [
+ 0,
+ 31,
+ 190,
+ 273
+ ],
+ "category_id": 6,
+ "area": 70104,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21332,
+ "image_id": 3901,
+ "bbox": [
+ 0,
+ 279,
+ 213,
+ 230
+ ],
+ "category_id": 9,
+ "area": 66340,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21369,
+ "image_id": 3904,
+ "bbox": [
+ 184,
+ 341,
+ 77,
+ 97
+ ],
+ "category_id": 8,
+ "area": 59740,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21380,
+ "image_id": 3905,
+ "bbox": [
+ 86,
+ 192,
+ 152,
+ 132
+ ],
+ "category_id": 9,
+ "area": 40826,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21382,
+ "image_id": 3906,
+ "bbox": [
+ 221,
+ 330,
+ 187,
+ 181
+ ],
+ "category_id": 6,
+ "area": 39780,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21383,
+ "image_id": 3906,
+ "bbox": [
+ 338,
+ 151,
+ 48,
+ 59
+ ],
+ "category_id": 6,
+ "area": 3416,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21384,
+ "image_id": 3906,
+ "bbox": [
+ 296,
+ 32,
+ 47,
+ 64
+ ],
+ "category_id": 6,
+ "area": 3540,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21388,
+ "image_id": 3907,
+ "bbox": [
+ 288,
+ 109,
+ 102,
+ 176
+ ],
+ "category_id": 6,
+ "area": 39208,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21389,
+ "image_id": 3907,
+ "bbox": [
+ 237,
+ 0,
+ 160,
+ 243
+ ],
+ "category_id": 6,
+ "area": 84800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21390,
+ "image_id": 3907,
+ "bbox": [
+ 431,
+ 67,
+ 80,
+ 114
+ ],
+ "category_id": 6,
+ "area": 20100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21391,
+ "image_id": 3907,
+ "bbox": [
+ 402,
+ 259,
+ 52,
+ 73
+ ],
+ "category_id": 6,
+ "area": 8439,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21392,
+ "image_id": 3907,
+ "bbox": [
+ 454,
+ 276,
+ 36,
+ 53
+ ],
+ "category_id": 6,
+ "area": 4200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21393,
+ "image_id": 3907,
+ "bbox": [
+ 434,
+ 221,
+ 35,
+ 49
+ ],
+ "category_id": 6,
+ "area": 3835,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21394,
+ "image_id": 3907,
+ "bbox": [
+ 369,
+ 182,
+ 66,
+ 103
+ ],
+ "category_id": 6,
+ "area": 14960,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21403,
+ "image_id": 3909,
+ "bbox": [
+ 47,
+ 292,
+ 363,
+ 177
+ ],
+ "category_id": 8,
+ "area": 138060,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21404,
+ "image_id": 3909,
+ "bbox": [
+ 363,
+ 288,
+ 111,
+ 61
+ ],
+ "category_id": 8,
+ "area": 14688,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21405,
+ "image_id": 3910,
+ "bbox": [
+ 2,
+ 130,
+ 434,
+ 283
+ ],
+ "category_id": 8,
+ "area": 432626,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21406,
+ "image_id": 3910,
+ "bbox": [
+ 153,
+ 9,
+ 72,
+ 80
+ ],
+ "category_id": 6,
+ "area": 20566,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21407,
+ "image_id": 3910,
+ "bbox": [
+ 222,
+ 0,
+ 168,
+ 133
+ ],
+ "category_id": 6,
+ "area": 79336,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21408,
+ "image_id": 3910,
+ "bbox": [
+ 256,
+ 165,
+ 132,
+ 344
+ ],
+ "category_id": 6,
+ "area": 160688,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21409,
+ "image_id": 3910,
+ "bbox": [
+ 405,
+ 382,
+ 78,
+ 86
+ ],
+ "category_id": 6,
+ "area": 23716,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21415,
+ "image_id": 3912,
+ "bbox": [
+ 145,
+ 35,
+ 266,
+ 390
+ ],
+ "category_id": 10,
+ "area": 428688,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21427,
+ "image_id": 3913,
+ "bbox": [
+ 153,
+ 207,
+ 89,
+ 114
+ ],
+ "category_id": 8,
+ "area": 35903,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21428,
+ "image_id": 3913,
+ "bbox": [
+ 364,
+ 152,
+ 110,
+ 120
+ ],
+ "category_id": 8,
+ "area": 46475,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21429,
+ "image_id": 3913,
+ "bbox": [
+ 464,
+ 344,
+ 33,
+ 32
+ ],
+ "category_id": 6,
+ "area": 3735,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21430,
+ "image_id": 3914,
+ "bbox": [
+ 0,
+ 165,
+ 50,
+ 215
+ ],
+ "category_id": 6,
+ "area": 24309,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21431,
+ "image_id": 3914,
+ "bbox": [
+ 202,
+ 293,
+ 288,
+ 218
+ ],
+ "category_id": 9,
+ "area": 141192,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21432,
+ "image_id": 3915,
+ "bbox": [
+ 208,
+ 132,
+ 158,
+ 230
+ ],
+ "category_id": 9,
+ "area": 47740,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21433,
+ "image_id": 3915,
+ "bbox": [
+ 129,
+ 368,
+ 54,
+ 39
+ ],
+ "category_id": 9,
+ "area": 2812,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21434,
+ "image_id": 3916,
+ "bbox": [
+ 287,
+ 254,
+ 222,
+ 189
+ ],
+ "category_id": 9,
+ "area": 147896,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21445,
+ "image_id": 3919,
+ "bbox": [
+ 0,
+ 0,
+ 512,
+ 332
+ ],
+ "category_id": 8,
+ "area": 1345920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21446,
+ "image_id": 3919,
+ "bbox": [
+ 2,
+ 285,
+ 350,
+ 185
+ ],
+ "category_id": 8,
+ "area": 513383,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21449,
+ "image_id": 3921,
+ "bbox": [
+ 1,
+ 205,
+ 386,
+ 274
+ ],
+ "category_id": 8,
+ "area": 838971,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21450,
+ "image_id": 3921,
+ "bbox": [
+ 168,
+ 176,
+ 47,
+ 36
+ ],
+ "category_id": 8,
+ "area": 13629,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21451,
+ "image_id": 3921,
+ "bbox": [
+ 359,
+ 134,
+ 43,
+ 52
+ ],
+ "category_id": 8,
+ "area": 17820,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21452,
+ "image_id": 3922,
+ "bbox": [
+ 189,
+ 85,
+ 165,
+ 380
+ ],
+ "category_id": 10,
+ "area": 112728,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21453,
+ "image_id": 3923,
+ "bbox": [
+ 114,
+ 273,
+ 247,
+ 214
+ ],
+ "category_id": 8,
+ "area": 364707,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21454,
+ "image_id": 3923,
+ "bbox": [
+ 220,
+ 40,
+ 289,
+ 329
+ ],
+ "category_id": 8,
+ "area": 657951,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21455,
+ "image_id": 3924,
+ "bbox": [
+ 1,
+ 3,
+ 248,
+ 450
+ ],
+ "category_id": 9,
+ "area": 109340,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21456,
+ "image_id": 3924,
+ "bbox": [
+ 19,
+ 92,
+ 427,
+ 419
+ ],
+ "category_id": 10,
+ "area": 175192,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21457,
+ "image_id": 3925,
+ "bbox": [
+ 16,
+ 81,
+ 467,
+ 369
+ ],
+ "category_id": 8,
+ "area": 252096,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21465,
+ "image_id": 3928,
+ "bbox": [
+ 14,
+ 29,
+ 299,
+ 355
+ ],
+ "category_id": 8,
+ "area": 179172,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21466,
+ "image_id": 3928,
+ "bbox": [
+ 112,
+ 248,
+ 391,
+ 256
+ ],
+ "category_id": 8,
+ "area": 168226,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21470,
+ "image_id": 3930,
+ "bbox": [
+ 4,
+ 188,
+ 339,
+ 310
+ ],
+ "category_id": 8,
+ "area": 835088,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21471,
+ "image_id": 3930,
+ "bbox": [
+ 239,
+ 156,
+ 162,
+ 145
+ ],
+ "category_id": 8,
+ "area": 187577,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21472,
+ "image_id": 3931,
+ "bbox": [
+ 113,
+ 48,
+ 252,
+ 336
+ ],
+ "category_id": 10,
+ "area": 141680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21473,
+ "image_id": 3931,
+ "bbox": [
+ 336,
+ 375,
+ 120,
+ 136
+ ],
+ "category_id": 10,
+ "area": 27510,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21474,
+ "image_id": 3931,
+ "bbox": [
+ 351,
+ 169,
+ 70,
+ 107
+ ],
+ "category_id": 10,
+ "area": 12669,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21478,
+ "image_id": 3933,
+ "bbox": [
+ 282,
+ 184,
+ 92,
+ 122
+ ],
+ "category_id": 8,
+ "area": 14950,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21479,
+ "image_id": 3934,
+ "bbox": [
+ 0,
+ 13,
+ 512,
+ 284
+ ],
+ "category_id": 6,
+ "area": 372992,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21483,
+ "image_id": 3935,
+ "bbox": [
+ 38,
+ 0,
+ 206,
+ 248
+ ],
+ "category_id": 6,
+ "area": 63630,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21484,
+ "image_id": 3935,
+ "bbox": [
+ 69,
+ 198,
+ 376,
+ 300
+ ],
+ "category_id": 9,
+ "area": 140462,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21501,
+ "image_id": 3939,
+ "bbox": [
+ 157,
+ 127,
+ 208,
+ 384
+ ],
+ "category_id": 10,
+ "area": 140504,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21502,
+ "image_id": 3939,
+ "bbox": [
+ 55,
+ 0,
+ 114,
+ 223
+ ],
+ "category_id": 10,
+ "area": 44944,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21506,
+ "image_id": 3941,
+ "bbox": [
+ 214,
+ 171,
+ 185,
+ 185
+ ],
+ "category_id": 8,
+ "area": 45414,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21511,
+ "image_id": 3942,
+ "bbox": [
+ 341,
+ 429,
+ 170,
+ 81
+ ],
+ "category_id": 6,
+ "area": 11937,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21512,
+ "image_id": 3942,
+ "bbox": [
+ 287,
+ 493,
+ 64,
+ 18
+ ],
+ "category_id": 6,
+ "area": 1040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21539,
+ "image_id": 3945,
+ "bbox": [
+ 0,
+ 0,
+ 267,
+ 271
+ ],
+ "category_id": 6,
+ "area": 179275,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21560,
+ "image_id": 3947,
+ "bbox": [
+ 26,
+ 72,
+ 326,
+ 202
+ ],
+ "category_id": 9,
+ "area": 168850,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21561,
+ "image_id": 3947,
+ "bbox": [
+ 205,
+ 161,
+ 268,
+ 308
+ ],
+ "category_id": 10,
+ "area": 210672,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21563,
+ "image_id": 3949,
+ "bbox": [
+ 0,
+ 7,
+ 201,
+ 263
+ ],
+ "category_id": 8,
+ "area": 55825,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21564,
+ "image_id": 3949,
+ "bbox": [
+ 0,
+ 211,
+ 238,
+ 300
+ ],
+ "category_id": 8,
+ "area": 75306,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21565,
+ "image_id": 3949,
+ "bbox": [
+ 176,
+ 304,
+ 152,
+ 103
+ ],
+ "category_id": 8,
+ "area": 16640,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21566,
+ "image_id": 3949,
+ "bbox": [
+ 250,
+ 213,
+ 160,
+ 152
+ ],
+ "category_id": 8,
+ "area": 25623,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21574,
+ "image_id": 3951,
+ "bbox": [
+ 0,
+ 0,
+ 95,
+ 169
+ ],
+ "category_id": 6,
+ "area": 37516,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21575,
+ "image_id": 3951,
+ "bbox": [
+ 208,
+ 154,
+ 214,
+ 333
+ ],
+ "category_id": 9,
+ "area": 165168,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21576,
+ "image_id": 3952,
+ "bbox": [
+ 0,
+ 0,
+ 154,
+ 134
+ ],
+ "category_id": 6,
+ "area": 333870,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21577,
+ "image_id": 3952,
+ "bbox": [
+ 1,
+ 185,
+ 35,
+ 31
+ ],
+ "category_id": 6,
+ "area": 17604,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21578,
+ "image_id": 3952,
+ "bbox": [
+ 2,
+ 378,
+ 33,
+ 108
+ ],
+ "category_id": 6,
+ "area": 57904,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21579,
+ "image_id": 3952,
+ "bbox": [
+ 161,
+ 5,
+ 77,
+ 59
+ ],
+ "category_id": 6,
+ "area": 74106,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21580,
+ "image_id": 3952,
+ "bbox": [
+ 113,
+ 99,
+ 103,
+ 133
+ ],
+ "category_id": 6,
+ "area": 221741,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21581,
+ "image_id": 3952,
+ "bbox": [
+ 121,
+ 226,
+ 113,
+ 97
+ ],
+ "category_id": 6,
+ "area": 177262,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21582,
+ "image_id": 3952,
+ "bbox": [
+ 79,
+ 327,
+ 111,
+ 89
+ ],
+ "category_id": 6,
+ "area": 161098,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21583,
+ "image_id": 3952,
+ "bbox": [
+ 239,
+ 2,
+ 91,
+ 190
+ ],
+ "category_id": 6,
+ "area": 279416,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21584,
+ "image_id": 3952,
+ "bbox": [
+ 228,
+ 182,
+ 82,
+ 79
+ ],
+ "category_id": 6,
+ "area": 104394,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21585,
+ "image_id": 3952,
+ "bbox": [
+ 249,
+ 256,
+ 125,
+ 78
+ ],
+ "category_id": 6,
+ "area": 158264,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21586,
+ "image_id": 3952,
+ "bbox": [
+ 324,
+ 25,
+ 36,
+ 74
+ ],
+ "category_id": 6,
+ "area": 42752,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21587,
+ "image_id": 3952,
+ "bbox": [
+ 342,
+ 75,
+ 68,
+ 71
+ ],
+ "category_id": 6,
+ "area": 77736,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21588,
+ "image_id": 3952,
+ "bbox": [
+ 411,
+ 63,
+ 98,
+ 100
+ ],
+ "category_id": 6,
+ "area": 158340,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21589,
+ "image_id": 3952,
+ "bbox": [
+ 443,
+ 257,
+ 68,
+ 102
+ ],
+ "category_id": 6,
+ "area": 112926,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21594,
+ "image_id": 3953,
+ "bbox": [
+ 35,
+ 352,
+ 158,
+ 159
+ ],
+ "category_id": 6,
+ "area": 51042,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21595,
+ "image_id": 3953,
+ "bbox": [
+ 215,
+ 267,
+ 296,
+ 243
+ ],
+ "category_id": 9,
+ "area": 145728,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21596,
+ "image_id": 3954,
+ "bbox": [
+ 171,
+ 164,
+ 247,
+ 345
+ ],
+ "category_id": 9,
+ "area": 112750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21597,
+ "image_id": 3954,
+ "bbox": [
+ 50,
+ 193,
+ 117,
+ 199
+ ],
+ "category_id": 10,
+ "area": 31005,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21598,
+ "image_id": 3955,
+ "bbox": [
+ 90,
+ 131,
+ 317,
+ 209
+ ],
+ "category_id": 9,
+ "area": 170750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21600,
+ "image_id": 3955,
+ "bbox": [
+ 352,
+ 296,
+ 159,
+ 213
+ ],
+ "category_id": 6,
+ "area": 87465,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21601,
+ "image_id": 3956,
+ "bbox": [
+ 87,
+ 55,
+ 251,
+ 341
+ ],
+ "category_id": 9,
+ "area": 170280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21602,
+ "image_id": 3956,
+ "bbox": [
+ 136,
+ 69,
+ 322,
+ 399
+ ],
+ "category_id": 9,
+ "area": 254430,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21603,
+ "image_id": 3957,
+ "bbox": [
+ 9,
+ 203,
+ 230,
+ 193
+ ],
+ "category_id": 9,
+ "area": 33276,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21604,
+ "image_id": 3957,
+ "bbox": [
+ 251,
+ 93,
+ 247,
+ 173
+ ],
+ "category_id": 9,
+ "area": 32118,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21605,
+ "image_id": 3958,
+ "bbox": [
+ 0,
+ 0,
+ 512,
+ 296
+ ],
+ "category_id": 6,
+ "area": 383616,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21609,
+ "image_id": 3959,
+ "bbox": [
+ 79,
+ 152,
+ 123,
+ 67
+ ],
+ "category_id": 8,
+ "area": 29260,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21610,
+ "image_id": 3959,
+ "bbox": [
+ 251,
+ 194,
+ 144,
+ 139
+ ],
+ "category_id": 8,
+ "area": 70590,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21652,
+ "image_id": 3964,
+ "bbox": [
+ 132,
+ 89,
+ 215,
+ 398
+ ],
+ "category_id": 6,
+ "area": 257140,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21661,
+ "image_id": 3966,
+ "bbox": [
+ 175,
+ 352,
+ 206,
+ 159
+ ],
+ "category_id": 10,
+ "area": 51870,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21670,
+ "image_id": 3969,
+ "bbox": [
+ 385,
+ 406,
+ 87,
+ 105
+ ],
+ "category_id": 9,
+ "area": 21546,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21671,
+ "image_id": 3969,
+ "bbox": [
+ 48,
+ 75,
+ 117,
+ 249
+ ],
+ "category_id": 9,
+ "area": 68770,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21672,
+ "image_id": 3969,
+ "bbox": [
+ 120,
+ 278,
+ 135,
+ 233
+ ],
+ "category_id": 9,
+ "area": 74200,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21673,
+ "image_id": 3969,
+ "bbox": [
+ 162,
+ 185,
+ 117,
+ 196
+ ],
+ "category_id": 9,
+ "area": 54280,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21674,
+ "image_id": 3969,
+ "bbox": [
+ 240,
+ 31,
+ 98,
+ 208
+ ],
+ "category_id": 9,
+ "area": 48000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21675,
+ "image_id": 3969,
+ "bbox": [
+ 257,
+ 95,
+ 112,
+ 185
+ ],
+ "category_id": 9,
+ "area": 48618,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21693,
+ "image_id": 3972,
+ "bbox": [
+ 114,
+ 133,
+ 46,
+ 95
+ ],
+ "category_id": 8,
+ "area": 5776,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21694,
+ "image_id": 3972,
+ "bbox": [
+ 173,
+ 0,
+ 102,
+ 191
+ ],
+ "category_id": 8,
+ "area": 25384,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21695,
+ "image_id": 3972,
+ "bbox": [
+ 227,
+ 266,
+ 102,
+ 233
+ ],
+ "category_id": 6,
+ "area": 30876,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21696,
+ "image_id": 3972,
+ "bbox": [
+ 144,
+ 417,
+ 137,
+ 94
+ ],
+ "category_id": 6,
+ "area": 16800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21697,
+ "image_id": 3972,
+ "bbox": [
+ 162,
+ 283,
+ 64,
+ 140
+ ],
+ "category_id": 6,
+ "area": 11760,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21700,
+ "image_id": 3974,
+ "bbox": [
+ 215,
+ 315,
+ 106,
+ 91
+ ],
+ "category_id": 6,
+ "area": 11438,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21701,
+ "image_id": 3974,
+ "bbox": [
+ 389,
+ 397,
+ 64,
+ 64
+ ],
+ "category_id": 6,
+ "area": 4800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21702,
+ "image_id": 3974,
+ "bbox": [
+ 176,
+ 201,
+ 60,
+ 33
+ ],
+ "category_id": 6,
+ "area": 2356,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21703,
+ "image_id": 3974,
+ "bbox": [
+ 124,
+ 211,
+ 48,
+ 53
+ ],
+ "category_id": 6,
+ "area": 3000,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21705,
+ "image_id": 3975,
+ "bbox": [
+ 181,
+ 196,
+ 226,
+ 280
+ ],
+ "category_id": 8,
+ "area": 158304,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21713,
+ "image_id": 3976,
+ "bbox": [
+ 0,
+ 67,
+ 512,
+ 444
+ ],
+ "category_id": 6,
+ "area": 144500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21719,
+ "image_id": 3979,
+ "bbox": [
+ 394,
+ 340,
+ 116,
+ 171
+ ],
+ "category_id": 9,
+ "area": 50838,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21720,
+ "image_id": 3979,
+ "bbox": [
+ 0,
+ 120,
+ 144,
+ 251
+ ],
+ "category_id": 9,
+ "area": 92910,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21721,
+ "image_id": 3979,
+ "bbox": [
+ 217,
+ 238,
+ 204,
+ 271
+ ],
+ "category_id": 9,
+ "area": 141856,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21722,
+ "image_id": 3979,
+ "bbox": [
+ 162,
+ 244,
+ 82,
+ 219
+ ],
+ "category_id": 9,
+ "area": 46170,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21723,
+ "image_id": 3979,
+ "bbox": [
+ 244,
+ 79,
+ 183,
+ 268
+ ],
+ "category_id": 9,
+ "area": 125976,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21724,
+ "image_id": 3979,
+ "bbox": [
+ 148,
+ 0,
+ 210,
+ 315
+ ],
+ "category_id": 9,
+ "area": 170144,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21725,
+ "image_id": 3980,
+ "bbox": [
+ 2,
+ 240,
+ 280,
+ 210
+ ],
+ "category_id": 8,
+ "area": 133893,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21726,
+ "image_id": 3980,
+ "bbox": [
+ 198,
+ 118,
+ 312,
+ 361
+ ],
+ "category_id": 8,
+ "area": 256256,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21727,
+ "image_id": 3981,
+ "bbox": [
+ 268,
+ 146,
+ 123,
+ 136
+ ],
+ "category_id": 8,
+ "area": 132594,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21728,
+ "image_id": 3981,
+ "bbox": [
+ 132,
+ 189,
+ 182,
+ 152
+ ],
+ "category_id": 8,
+ "area": 220248,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21729,
+ "image_id": 3981,
+ "bbox": [
+ 132,
+ 249,
+ 141,
+ 115
+ ],
+ "category_id": 8,
+ "area": 128547,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21732,
+ "image_id": 3982,
+ "bbox": [
+ 0,
+ 43,
+ 177,
+ 119
+ ],
+ "category_id": 9,
+ "area": 23326,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21733,
+ "image_id": 3982,
+ "bbox": [
+ 235,
+ 164,
+ 255,
+ 156
+ ],
+ "category_id": 9,
+ "area": 43901,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21739,
+ "image_id": 3984,
+ "bbox": [
+ 40,
+ 131,
+ 418,
+ 380
+ ],
+ "category_id": 9,
+ "area": 392836,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21740,
+ "image_id": 3984,
+ "bbox": [
+ 386,
+ 158,
+ 84,
+ 150
+ ],
+ "category_id": 6,
+ "area": 31486,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21741,
+ "image_id": 3984,
+ "bbox": [
+ 393,
+ 358,
+ 38,
+ 76
+ ],
+ "category_id": 6,
+ "area": 7216,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21747,
+ "image_id": 3985,
+ "bbox": [
+ 194,
+ 187,
+ 100,
+ 78
+ ],
+ "category_id": 9,
+ "area": 27720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21748,
+ "image_id": 3985,
+ "bbox": [
+ 263,
+ 183,
+ 152,
+ 136
+ ],
+ "category_id": 9,
+ "area": 73344,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21749,
+ "image_id": 3986,
+ "bbox": [
+ 118,
+ 274,
+ 22,
+ 52
+ ],
+ "category_id": 6,
+ "area": 7992,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21750,
+ "image_id": 3986,
+ "bbox": [
+ 477,
+ 374,
+ 17,
+ 69
+ ],
+ "category_id": 6,
+ "area": 8085,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21754,
+ "image_id": 3987,
+ "bbox": [
+ 134,
+ 194,
+ 376,
+ 253
+ ],
+ "category_id": 9,
+ "area": 160600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21755,
+ "image_id": 3987,
+ "bbox": [
+ 93,
+ 301,
+ 418,
+ 209
+ ],
+ "category_id": 6,
+ "area": 147972,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21756,
+ "image_id": 3987,
+ "bbox": [
+ 382,
+ 130,
+ 128,
+ 192
+ ],
+ "category_id": 6,
+ "area": 41800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21760,
+ "image_id": 3988,
+ "bbox": [
+ 13,
+ 167,
+ 459,
+ 299
+ ],
+ "category_id": 9,
+ "area": 342576,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21779,
+ "image_id": 3991,
+ "bbox": [
+ 139,
+ 164,
+ 52,
+ 35
+ ],
+ "category_id": 8,
+ "area": 14625,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21780,
+ "image_id": 3991,
+ "bbox": [
+ 149,
+ 207,
+ 102,
+ 55
+ ],
+ "category_id": 8,
+ "area": 44811,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21781,
+ "image_id": 3991,
+ "bbox": [
+ 213,
+ 164,
+ 52,
+ 71
+ ],
+ "category_id": 8,
+ "area": 29400,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21782,
+ "image_id": 3991,
+ "bbox": [
+ 266,
+ 198,
+ 69,
+ 83
+ ],
+ "category_id": 8,
+ "area": 46020,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21784,
+ "image_id": 3992,
+ "bbox": [
+ 85,
+ 154,
+ 424,
+ 356
+ ],
+ "category_id": 8,
+ "area": 177020,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21794,
+ "image_id": 3994,
+ "bbox": [
+ 64,
+ 136,
+ 92,
+ 121
+ ],
+ "category_id": 6,
+ "area": 25345,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21795,
+ "image_id": 3994,
+ "bbox": [
+ 62,
+ 246,
+ 109,
+ 126
+ ],
+ "category_id": 6,
+ "area": 30956,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21796,
+ "image_id": 3994,
+ "bbox": [
+ 135,
+ 387,
+ 29,
+ 72
+ ],
+ "category_id": 6,
+ "area": 4698,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21797,
+ "image_id": 3994,
+ "bbox": [
+ 150,
+ 408,
+ 70,
+ 104
+ ],
+ "category_id": 6,
+ "area": 16380,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21798,
+ "image_id": 3994,
+ "bbox": [
+ 335,
+ 258,
+ 82,
+ 129
+ ],
+ "category_id": 6,
+ "area": 23944,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21799,
+ "image_id": 3994,
+ "bbox": [
+ 376,
+ 316,
+ 51,
+ 96
+ ],
+ "category_id": 6,
+ "area": 11118,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21800,
+ "image_id": 3994,
+ "bbox": [
+ 275,
+ 94,
+ 65,
+ 57
+ ],
+ "category_id": 6,
+ "area": 8450,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21801,
+ "image_id": 3994,
+ "bbox": [
+ 434,
+ 354,
+ 40,
+ 87
+ ],
+ "category_id": 6,
+ "area": 7840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21803,
+ "image_id": 3994,
+ "bbox": [
+ 192,
+ 225,
+ 38,
+ 54
+ ],
+ "category_id": 6,
+ "area": 4636,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21810,
+ "image_id": 3996,
+ "bbox": [
+ 221,
+ 203,
+ 152,
+ 203
+ ],
+ "category_id": 6,
+ "area": 44576,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21811,
+ "image_id": 3996,
+ "bbox": [
+ 236,
+ 375,
+ 215,
+ 136
+ ],
+ "category_id": 9,
+ "area": 42294,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21832,
+ "image_id": 4000,
+ "bbox": [
+ 100,
+ 328,
+ 59,
+ 107
+ ],
+ "category_id": 6,
+ "area": 22050,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21833,
+ "image_id": 4000,
+ "bbox": [
+ 123,
+ 68,
+ 133,
+ 146
+ ],
+ "category_id": 6,
+ "area": 68060,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21834,
+ "image_id": 4000,
+ "bbox": [
+ 166,
+ 309,
+ 85,
+ 75
+ ],
+ "category_id": 6,
+ "area": 22472,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21835,
+ "image_id": 4000,
+ "bbox": [
+ 246,
+ 243,
+ 70,
+ 81
+ ],
+ "category_id": 6,
+ "area": 20064,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21849,
+ "image_id": 4002,
+ "bbox": [
+ 26,
+ 229,
+ 195,
+ 180
+ ],
+ "category_id": 8,
+ "area": 54900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21850,
+ "image_id": 4002,
+ "bbox": [
+ 203,
+ 139,
+ 203,
+ 237
+ ],
+ "category_id": 8,
+ "area": 75366,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21851,
+ "image_id": 4002,
+ "bbox": [
+ 329,
+ 141,
+ 173,
+ 144
+ ],
+ "category_id": 8,
+ "area": 38880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21852,
+ "image_id": 4003,
+ "bbox": [
+ 0,
+ 12,
+ 352,
+ 357
+ ],
+ "category_id": 6,
+ "area": 317772,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21853,
+ "image_id": 4003,
+ "bbox": [
+ 413,
+ 215,
+ 49,
+ 75
+ ],
+ "category_id": 6,
+ "area": 9405,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21867,
+ "image_id": 4005,
+ "bbox": [
+ 69,
+ 324,
+ 22,
+ 50
+ ],
+ "category_id": 6,
+ "area": 7844,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21868,
+ "image_id": 4005,
+ "bbox": [
+ 360,
+ 404,
+ 34,
+ 59
+ ],
+ "category_id": 6,
+ "area": 13888,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21887,
+ "image_id": 4008,
+ "bbox": [
+ 68,
+ 187,
+ 193,
+ 196
+ ],
+ "category_id": 8,
+ "area": 121464,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21888,
+ "image_id": 4008,
+ "bbox": [
+ 221,
+ 154,
+ 185,
+ 217
+ ],
+ "category_id": 8,
+ "area": 128619,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21889,
+ "image_id": 4009,
+ "bbox": [
+ 189,
+ 221,
+ 242,
+ 172
+ ],
+ "category_id": 8,
+ "area": 146287,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21890,
+ "image_id": 4009,
+ "bbox": [
+ 460,
+ 90,
+ 51,
+ 275
+ ],
+ "category_id": 8,
+ "area": 49794,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21891,
+ "image_id": 4009,
+ "bbox": [
+ 64,
+ 153,
+ 156,
+ 170
+ ],
+ "category_id": 6,
+ "area": 93210,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21892,
+ "image_id": 4009,
+ "bbox": [
+ 1,
+ 228,
+ 406,
+ 282
+ ],
+ "category_id": 6,
+ "area": 402336,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21904,
+ "image_id": 4012,
+ "bbox": [
+ 0,
+ 123,
+ 351,
+ 328
+ ],
+ "category_id": 9,
+ "area": 2619792,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21929,
+ "image_id": 4017,
+ "bbox": [
+ 32,
+ 96,
+ 476,
+ 409
+ ],
+ "category_id": 9,
+ "area": 685440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21961,
+ "image_id": 4021,
+ "bbox": [
+ 217,
+ 99,
+ 235,
+ 412
+ ],
+ "category_id": 9,
+ "area": 341620,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21962,
+ "image_id": 4021,
+ "bbox": [
+ 240,
+ 114,
+ 233,
+ 377
+ ],
+ "category_id": 9,
+ "area": 310104,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21966,
+ "image_id": 4023,
+ "bbox": [
+ 36,
+ 175,
+ 404,
+ 211
+ ],
+ "category_id": 8,
+ "area": 89056,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21967,
+ "image_id": 4023,
+ "bbox": [
+ 8,
+ 39,
+ 172,
+ 164
+ ],
+ "category_id": 8,
+ "area": 29592,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21968,
+ "image_id": 4023,
+ "bbox": [
+ 414,
+ 41,
+ 97,
+ 159
+ ],
+ "category_id": 8,
+ "area": 16226,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21969,
+ "image_id": 4024,
+ "bbox": [
+ 33,
+ 243,
+ 281,
+ 268
+ ],
+ "category_id": 9,
+ "area": 90870,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21986,
+ "image_id": 4027,
+ "bbox": [
+ 138,
+ 92,
+ 281,
+ 239
+ ],
+ "category_id": 9,
+ "area": 136680,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21987,
+ "image_id": 4027,
+ "bbox": [
+ 248,
+ 193,
+ 137,
+ 193
+ ],
+ "category_id": 10,
+ "area": 54033,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21989,
+ "image_id": 4028,
+ "bbox": [
+ 126,
+ 115,
+ 326,
+ 321
+ ],
+ "category_id": 8,
+ "area": 514185,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21990,
+ "image_id": 4029,
+ "bbox": [
+ 49,
+ 71,
+ 443,
+ 369
+ ],
+ "category_id": 6,
+ "area": 192192,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21991,
+ "image_id": 4030,
+ "bbox": [
+ 77,
+ 255,
+ 296,
+ 188
+ ],
+ "category_id": 8,
+ "area": 193831,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21992,
+ "image_id": 4030,
+ "bbox": [
+ 2,
+ 295,
+ 113,
+ 210
+ ],
+ "category_id": 6,
+ "area": 83485,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21994,
+ "image_id": 4031,
+ "bbox": [
+ 203,
+ 0,
+ 242,
+ 320
+ ],
+ "category_id": 9,
+ "area": 85264,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 21998,
+ "image_id": 4032,
+ "bbox": [
+ 0,
+ 3,
+ 512,
+ 298
+ ],
+ "category_id": 6,
+ "area": 385672,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22013,
+ "image_id": 4036,
+ "bbox": [
+ 202,
+ 180,
+ 197,
+ 322
+ ],
+ "category_id": 10,
+ "area": 223822,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22014,
+ "image_id": 4036,
+ "bbox": [
+ 67,
+ 78,
+ 412,
+ 223
+ ],
+ "category_id": 9,
+ "area": 324048,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22015,
+ "image_id": 4037,
+ "bbox": [
+ 0,
+ 0,
+ 364,
+ 434
+ ],
+ "category_id": 6,
+ "area": 206498,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22022,
+ "image_id": 4038,
+ "bbox": [
+ 0,
+ 30,
+ 84,
+ 182
+ ],
+ "category_id": 6,
+ "area": 150612,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22023,
+ "image_id": 4038,
+ "bbox": [
+ 11,
+ 12,
+ 93,
+ 94
+ ],
+ "category_id": 6,
+ "area": 86880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22024,
+ "image_id": 4038,
+ "bbox": [
+ 90,
+ 20,
+ 108,
+ 140
+ ],
+ "category_id": 6,
+ "area": 148869,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22025,
+ "image_id": 4038,
+ "bbox": [
+ 39,
+ 130,
+ 117,
+ 214
+ ],
+ "category_id": 6,
+ "area": 246522,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22026,
+ "image_id": 4038,
+ "bbox": [
+ 173,
+ 33,
+ 121,
+ 255
+ ],
+ "category_id": 6,
+ "area": 304560,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22027,
+ "image_id": 4038,
+ "bbox": [
+ 296,
+ 69,
+ 33,
+ 114
+ ],
+ "category_id": 6,
+ "area": 37990,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22028,
+ "image_id": 4038,
+ "bbox": [
+ 173,
+ 280,
+ 99,
+ 111
+ ],
+ "category_id": 6,
+ "area": 107724,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22029,
+ "image_id": 4038,
+ "bbox": [
+ 315,
+ 134,
+ 62,
+ 106
+ ],
+ "category_id": 6,
+ "area": 65853,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22030,
+ "image_id": 4038,
+ "bbox": [
+ 393,
+ 125,
+ 115,
+ 131
+ ],
+ "category_id": 6,
+ "area": 148851,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22031,
+ "image_id": 4038,
+ "bbox": [
+ 435,
+ 382,
+ 75,
+ 121
+ ],
+ "category_id": 6,
+ "area": 90244,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22049,
+ "image_id": 4042,
+ "bbox": [
+ 368,
+ 4,
+ 127,
+ 241
+ ],
+ "category_id": 10,
+ "area": 108460,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22050,
+ "image_id": 4043,
+ "bbox": [
+ 4,
+ 58,
+ 51,
+ 70
+ ],
+ "category_id": 6,
+ "area": 86130,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22051,
+ "image_id": 4043,
+ "bbox": [
+ 0,
+ 107,
+ 41,
+ 59
+ ],
+ "category_id": 6,
+ "area": 58539,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22052,
+ "image_id": 4043,
+ "bbox": [
+ 3,
+ 170,
+ 61,
+ 81
+ ],
+ "category_id": 6,
+ "area": 118255,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22053,
+ "image_id": 4043,
+ "bbox": [
+ 6,
+ 272,
+ 37,
+ 72
+ ],
+ "category_id": 6,
+ "area": 64883,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22054,
+ "image_id": 4043,
+ "bbox": [
+ 8,
+ 391,
+ 53,
+ 64
+ ],
+ "category_id": 6,
+ "area": 81312,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22055,
+ "image_id": 4043,
+ "bbox": [
+ 64,
+ 59,
+ 66,
+ 60
+ ],
+ "category_id": 6,
+ "area": 94122,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22056,
+ "image_id": 4043,
+ "bbox": [
+ 29,
+ 122,
+ 58,
+ 62
+ ],
+ "category_id": 6,
+ "area": 86095,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22057,
+ "image_id": 4043,
+ "bbox": [
+ 39,
+ 129,
+ 50,
+ 89
+ ],
+ "category_id": 6,
+ "area": 107748,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22058,
+ "image_id": 4043,
+ "bbox": [
+ 88,
+ 343,
+ 71,
+ 123
+ ],
+ "category_id": 6,
+ "area": 207264,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22059,
+ "image_id": 4043,
+ "bbox": [
+ 96,
+ 49,
+ 158,
+ 127
+ ],
+ "category_id": 6,
+ "area": 474220,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22060,
+ "image_id": 4043,
+ "bbox": [
+ 205,
+ 47,
+ 50,
+ 48
+ ],
+ "category_id": 6,
+ "area": 57420,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22061,
+ "image_id": 4043,
+ "bbox": [
+ 112,
+ 211,
+ 46,
+ 38
+ ],
+ "category_id": 6,
+ "area": 43040,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22062,
+ "image_id": 4043,
+ "bbox": [
+ 260,
+ 54,
+ 68,
+ 74
+ ],
+ "category_id": 6,
+ "area": 120736,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22063,
+ "image_id": 4043,
+ "bbox": [
+ 222,
+ 120,
+ 81,
+ 130
+ ],
+ "category_id": 6,
+ "area": 251853,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22064,
+ "image_id": 4043,
+ "bbox": [
+ 226,
+ 247,
+ 92,
+ 81
+ ],
+ "category_id": 6,
+ "area": 176880,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22081,
+ "image_id": 4044,
+ "bbox": [
+ 31,
+ 3,
+ 117,
+ 149
+ ],
+ "category_id": 8,
+ "area": 138285,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22082,
+ "image_id": 4044,
+ "bbox": [
+ 4,
+ 26,
+ 487,
+ 289
+ ],
+ "category_id": 8,
+ "area": 1116908,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22083,
+ "image_id": 4044,
+ "bbox": [
+ 129,
+ 203,
+ 295,
+ 210
+ ],
+ "category_id": 8,
+ "area": 492396,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22085,
+ "image_id": 4045,
+ "bbox": [
+ 40,
+ 53,
+ 164,
+ 85
+ ],
+ "category_id": 9,
+ "area": 14118,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22086,
+ "image_id": 4046,
+ "bbox": [
+ 257,
+ 181,
+ 108,
+ 184
+ ],
+ "category_id": 9,
+ "area": 22052,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22088,
+ "image_id": 4047,
+ "bbox": [
+ 83,
+ 102,
+ 200,
+ 324
+ ],
+ "category_id": 6,
+ "area": 226546,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22098,
+ "image_id": 4048,
+ "bbox": [
+ 96,
+ 274,
+ 218,
+ 113
+ ],
+ "category_id": 9,
+ "area": 63910,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22099,
+ "image_id": 4048,
+ "bbox": [
+ 133,
+ 319,
+ 244,
+ 190
+ ],
+ "category_id": 6,
+ "area": 119970,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22100,
+ "image_id": 4049,
+ "bbox": [
+ 114,
+ 234,
+ 156,
+ 161
+ ],
+ "category_id": 8,
+ "area": 88140,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22101,
+ "image_id": 4049,
+ "bbox": [
+ 222,
+ 232,
+ 193,
+ 166
+ ],
+ "category_id": 8,
+ "area": 112539,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22103,
+ "image_id": 4049,
+ "bbox": [
+ 101,
+ 132,
+ 24,
+ 29
+ ],
+ "category_id": 6,
+ "area": 2562,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22104,
+ "image_id": 4049,
+ "bbox": [
+ 136,
+ 169,
+ 21,
+ 25
+ ],
+ "category_id": 6,
+ "area": 1944,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22105,
+ "image_id": 4049,
+ "bbox": [
+ 143,
+ 218,
+ 28,
+ 34
+ ],
+ "category_id": 6,
+ "area": 3408,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22106,
+ "image_id": 4049,
+ "bbox": [
+ 163,
+ 201,
+ 18,
+ 20
+ ],
+ "category_id": 6,
+ "area": 1334,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22109,
+ "image_id": 4050,
+ "bbox": [
+ 199,
+ 201,
+ 76,
+ 154
+ ],
+ "category_id": 9,
+ "area": 48240,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22110,
+ "image_id": 4051,
+ "bbox": [
+ 146,
+ 86,
+ 364,
+ 396
+ ],
+ "category_id": 10,
+ "area": 1141976,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22111,
+ "image_id": 4052,
+ "bbox": [
+ 76,
+ 187,
+ 147,
+ 79
+ ],
+ "category_id": 8,
+ "area": 54516,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22112,
+ "image_id": 4052,
+ "bbox": [
+ 140,
+ 41,
+ 64,
+ 170
+ ],
+ "category_id": 8,
+ "area": 50853,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22113,
+ "image_id": 4052,
+ "bbox": [
+ 223,
+ 110,
+ 81,
+ 158
+ ],
+ "category_id": 8,
+ "area": 59944,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22114,
+ "image_id": 4052,
+ "bbox": [
+ 88,
+ 281,
+ 142,
+ 133
+ ],
+ "category_id": 8,
+ "area": 88704,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22115,
+ "image_id": 4052,
+ "bbox": [
+ 319,
+ 117,
+ 146,
+ 218
+ ],
+ "category_id": 8,
+ "area": 149175,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22116,
+ "image_id": 4052,
+ "bbox": [
+ 354,
+ 379,
+ 74,
+ 131
+ ],
+ "category_id": 8,
+ "area": 45825,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22117,
+ "image_id": 4052,
+ "bbox": [
+ 252,
+ 302,
+ 237,
+ 159
+ ],
+ "category_id": 8,
+ "area": 176328,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22118,
+ "image_id": 4052,
+ "bbox": [
+ 197,
+ 409,
+ 154,
+ 101
+ ],
+ "category_id": 8,
+ "area": 73084,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22119,
+ "image_id": 4053,
+ "bbox": [
+ 45,
+ 71,
+ 404,
+ 437
+ ],
+ "category_id": 8,
+ "area": 233290,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22160,
+ "image_id": 4059,
+ "bbox": [
+ 0,
+ 262,
+ 65,
+ 117
+ ],
+ "category_id": 6,
+ "area": 6804,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22161,
+ "image_id": 4059,
+ "bbox": [
+ 0,
+ 163,
+ 73,
+ 99
+ ],
+ "category_id": 6,
+ "area": 6461,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22162,
+ "image_id": 4059,
+ "bbox": [
+ 75,
+ 137,
+ 100,
+ 113
+ ],
+ "category_id": 6,
+ "area": 10088,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22163,
+ "image_id": 4059,
+ "bbox": [
+ 188,
+ 186,
+ 98,
+ 89
+ ],
+ "category_id": 6,
+ "area": 7790,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22164,
+ "image_id": 4059,
+ "bbox": [
+ 160,
+ 155,
+ 71,
+ 103
+ ],
+ "category_id": 6,
+ "area": 6555,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22168,
+ "image_id": 4060,
+ "bbox": [
+ 0,
+ 208,
+ 102,
+ 215
+ ],
+ "category_id": 6,
+ "area": 131920,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22169,
+ "image_id": 4061,
+ "bbox": [
+ 11,
+ 208,
+ 311,
+ 219
+ ],
+ "category_id": 9,
+ "area": 111036,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22170,
+ "image_id": 4062,
+ "bbox": [
+ 0,
+ 1,
+ 153,
+ 137
+ ],
+ "category_id": 6,
+ "area": 339864,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22171,
+ "image_id": 4062,
+ "bbox": [
+ 0,
+ 372,
+ 39,
+ 97
+ ],
+ "category_id": 6,
+ "area": 62008,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22172,
+ "image_id": 4062,
+ "bbox": [
+ 113,
+ 71,
+ 101,
+ 159
+ ],
+ "category_id": 6,
+ "area": 259992,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22173,
+ "image_id": 4062,
+ "bbox": [
+ 169,
+ 2,
+ 44,
+ 81
+ ],
+ "category_id": 6,
+ "area": 58092,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22174,
+ "image_id": 4062,
+ "bbox": [
+ 120,
+ 225,
+ 114,
+ 99
+ ],
+ "category_id": 6,
+ "area": 183008,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22175,
+ "image_id": 4062,
+ "bbox": [
+ 80,
+ 322,
+ 111,
+ 102
+ ],
+ "category_id": 6,
+ "area": 182664,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22176,
+ "image_id": 4062,
+ "bbox": [
+ 234,
+ 71,
+ 96,
+ 120
+ ],
+ "category_id": 6,
+ "area": 187264,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22177,
+ "image_id": 4062,
+ "bbox": [
+ 226,
+ 180,
+ 86,
+ 84
+ ],
+ "category_id": 6,
+ "area": 117384,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22178,
+ "image_id": 4062,
+ "bbox": [
+ 246,
+ 256,
+ 129,
+ 82
+ ],
+ "category_id": 6,
+ "area": 170684,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22179,
+ "image_id": 4062,
+ "bbox": [
+ 328,
+ 25,
+ 31,
+ 82
+ ],
+ "category_id": 6,
+ "area": 42032,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22180,
+ "image_id": 4062,
+ "bbox": [
+ 345,
+ 74,
+ 53,
+ 67
+ ],
+ "category_id": 6,
+ "area": 58750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22181,
+ "image_id": 4062,
+ "bbox": [
+ 410,
+ 66,
+ 98,
+ 97
+ ],
+ "category_id": 6,
+ "area": 154466,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22182,
+ "image_id": 4062,
+ "bbox": [
+ 443,
+ 256,
+ 68,
+ 106
+ ],
+ "category_id": 6,
+ "area": 117711,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22191,
+ "image_id": 4064,
+ "bbox": [
+ 173,
+ 287,
+ 233,
+ 194
+ ],
+ "category_id": 10,
+ "area": 71688,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22204,
+ "image_id": 4069,
+ "bbox": [
+ 191,
+ 75,
+ 104,
+ 225
+ ],
+ "category_id": 10,
+ "area": 79360,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22205,
+ "image_id": 4069,
+ "bbox": [
+ 285,
+ 39,
+ 189,
+ 331
+ ],
+ "category_id": 9,
+ "area": 212030,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22208,
+ "image_id": 4070,
+ "bbox": [
+ 90,
+ 262,
+ 107,
+ 106
+ ],
+ "category_id": 9,
+ "area": 17150,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22209,
+ "image_id": 4070,
+ "bbox": [
+ 53,
+ 354,
+ 48,
+ 71
+ ],
+ "category_id": 6,
+ "area": 5214,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22210,
+ "image_id": 4071,
+ "bbox": [
+ 46,
+ 156,
+ 15,
+ 40
+ ],
+ "category_id": 6,
+ "area": 4128,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22211,
+ "image_id": 4071,
+ "bbox": [
+ 143,
+ 292,
+ 11,
+ 25
+ ],
+ "category_id": 6,
+ "area": 1836,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22212,
+ "image_id": 4071,
+ "bbox": [
+ 173,
+ 301,
+ 19,
+ 30
+ ],
+ "category_id": 6,
+ "area": 3770,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22213,
+ "image_id": 4071,
+ "bbox": [
+ 32,
+ 35,
+ 94,
+ 105
+ ],
+ "category_id": 6,
+ "area": 63270,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22214,
+ "image_id": 4071,
+ "bbox": [
+ 358,
+ 270,
+ 15,
+ 29
+ ],
+ "category_id": 6,
+ "area": 2961,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22218,
+ "image_id": 4072,
+ "bbox": [
+ 54,
+ 272,
+ 397,
+ 238
+ ],
+ "category_id": 9,
+ "area": 159544,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22219,
+ "image_id": 4072,
+ "bbox": [
+ 0,
+ 284,
+ 124,
+ 227
+ ],
+ "category_id": 6,
+ "area": 47671,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22223,
+ "image_id": 4074,
+ "bbox": [
+ 214,
+ 105,
+ 56,
+ 59
+ ],
+ "category_id": 8,
+ "area": 26500,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22224,
+ "image_id": 4074,
+ "bbox": [
+ 241,
+ 27,
+ 70,
+ 86
+ ],
+ "category_id": 6,
+ "area": 48678,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22261,
+ "image_id": 4079,
+ "bbox": [
+ 76,
+ 239,
+ 293,
+ 118
+ ],
+ "category_id": 9,
+ "area": 47790,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22262,
+ "image_id": 4080,
+ "bbox": [
+ 22,
+ 123,
+ 61,
+ 51
+ ],
+ "category_id": 8,
+ "area": 24961,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22263,
+ "image_id": 4080,
+ "bbox": [
+ 8,
+ 182,
+ 79,
+ 66
+ ],
+ "category_id": 8,
+ "area": 41877,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22264,
+ "image_id": 4080,
+ "bbox": [
+ 59,
+ 198,
+ 352,
+ 244
+ ],
+ "category_id": 8,
+ "area": 682152,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22265,
+ "image_id": 4080,
+ "bbox": [
+ 1,
+ 410,
+ 159,
+ 101
+ ],
+ "category_id": 8,
+ "area": 128785,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22270,
+ "image_id": 4081,
+ "bbox": [
+ 386,
+ 332,
+ 124,
+ 174
+ ],
+ "category_id": 6,
+ "area": 50508,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22286,
+ "image_id": 4085,
+ "bbox": [
+ 103,
+ 184,
+ 405,
+ 323
+ ],
+ "category_id": 9,
+ "area": 461370,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22291,
+ "image_id": 4087,
+ "bbox": [
+ 138,
+ 74,
+ 118,
+ 33
+ ],
+ "category_id": 8,
+ "area": 7875,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22292,
+ "image_id": 4087,
+ "bbox": [
+ 194,
+ 257,
+ 203,
+ 104
+ ],
+ "category_id": 8,
+ "area": 42159,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22293,
+ "image_id": 4088,
+ "bbox": [
+ 2,
+ 185,
+ 324,
+ 323
+ ],
+ "category_id": 9,
+ "area": 369005,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22294,
+ "image_id": 4088,
+ "bbox": [
+ 6,
+ 0,
+ 470,
+ 367
+ ],
+ "category_id": 9,
+ "area": 607475,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22295,
+ "image_id": 4088,
+ "bbox": [
+ 52,
+ 1,
+ 456,
+ 347
+ ],
+ "category_id": 9,
+ "area": 558438,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22296,
+ "image_id": 4089,
+ "bbox": [
+ 11,
+ 117,
+ 436,
+ 394
+ ],
+ "category_id": 9,
+ "area": 137970,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22297,
+ "image_id": 4089,
+ "bbox": [
+ 348,
+ 30,
+ 162,
+ 397
+ ],
+ "category_id": 6,
+ "area": 51952,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22298,
+ "image_id": 4089,
+ "bbox": [
+ 205,
+ 415,
+ 303,
+ 96
+ ],
+ "category_id": 6,
+ "area": 23496,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22311,
+ "image_id": 4092,
+ "bbox": [
+ 85,
+ 163,
+ 426,
+ 346
+ ],
+ "category_id": 9,
+ "area": 421969,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22313,
+ "image_id": 4093,
+ "bbox": [
+ 281,
+ 173,
+ 133,
+ 87
+ ],
+ "category_id": 8,
+ "area": 15416,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22314,
+ "image_id": 4094,
+ "bbox": [
+ 57,
+ 0,
+ 92,
+ 54
+ ],
+ "category_id": 6,
+ "area": 11520,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22315,
+ "image_id": 4094,
+ "bbox": [
+ 254,
+ 174,
+ 166,
+ 218
+ ],
+ "category_id": 9,
+ "area": 84099,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22323,
+ "image_id": 4097,
+ "bbox": [
+ 0,
+ 62,
+ 489,
+ 377
+ ],
+ "category_id": 10,
+ "area": 186368,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22357,
+ "image_id": 4101,
+ "bbox": [
+ 43,
+ 206,
+ 153,
+ 126
+ ],
+ "category_id": 9,
+ "area": 42665,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22393,
+ "image_id": 4107,
+ "bbox": [
+ 13,
+ 102,
+ 352,
+ 407
+ ],
+ "category_id": 9,
+ "area": 409981,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22398,
+ "image_id": 4108,
+ "bbox": [
+ 1,
+ 0,
+ 242,
+ 398
+ ],
+ "category_id": 6,
+ "area": 262171,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22399,
+ "image_id": 4108,
+ "bbox": [
+ 9,
+ 300,
+ 270,
+ 211
+ ],
+ "category_id": 6,
+ "area": 155628,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22400,
+ "image_id": 4108,
+ "bbox": [
+ 205,
+ 250,
+ 306,
+ 255
+ ],
+ "category_id": 6,
+ "area": 212652,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22401,
+ "image_id": 4108,
+ "bbox": [
+ 187,
+ 112,
+ 324,
+ 199
+ ],
+ "category_id": 6,
+ "area": 175840,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22402,
+ "image_id": 4109,
+ "bbox": [
+ 74,
+ 20,
+ 257,
+ 374
+ ],
+ "category_id": 9,
+ "area": 174096,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22403,
+ "image_id": 4109,
+ "bbox": [
+ 132,
+ 23,
+ 312,
+ 447
+ ],
+ "category_id": 9,
+ "area": 251808,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22411,
+ "image_id": 4113,
+ "bbox": [
+ 164,
+ 211,
+ 154,
+ 201
+ ],
+ "category_id": 10,
+ "area": 109908,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22412,
+ "image_id": 4113,
+ "bbox": [
+ 207,
+ 69,
+ 179,
+ 157
+ ],
+ "category_id": 9,
+ "area": 99456,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22413,
+ "image_id": 4113,
+ "bbox": [
+ 248,
+ 175,
+ 211,
+ 253
+ ],
+ "category_id": 9,
+ "area": 187968,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22420,
+ "image_id": 4115,
+ "bbox": [
+ 128,
+ 131,
+ 235,
+ 379
+ ],
+ "category_id": 10,
+ "area": 707283,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22421,
+ "image_id": 4115,
+ "bbox": [
+ 131,
+ 0,
+ 200,
+ 132
+ ],
+ "category_id": 10,
+ "area": 209529,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22429,
+ "image_id": 4117,
+ "bbox": [
+ 120,
+ 198,
+ 126,
+ 137
+ ],
+ "category_id": 9,
+ "area": 61498,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22434,
+ "image_id": 4118,
+ "bbox": [
+ 204,
+ 196,
+ 119,
+ 104
+ ],
+ "category_id": 8,
+ "area": 84483,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22435,
+ "image_id": 4118,
+ "bbox": [
+ 45,
+ 169,
+ 81,
+ 88
+ ],
+ "category_id": 6,
+ "area": 48320,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22436,
+ "image_id": 4118,
+ "bbox": [
+ 28,
+ 222,
+ 124,
+ 176
+ ],
+ "category_id": 6,
+ "area": 147697,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22437,
+ "image_id": 4118,
+ "bbox": [
+ 158,
+ 367,
+ 347,
+ 143
+ ],
+ "category_id": 6,
+ "area": 335405,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22440,
+ "image_id": 4119,
+ "bbox": [
+ 0,
+ 2,
+ 117,
+ 122
+ ],
+ "category_id": 6,
+ "area": 124593,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22441,
+ "image_id": 4119,
+ "bbox": [
+ 0,
+ 115,
+ 42,
+ 75
+ ],
+ "category_id": 6,
+ "area": 27606,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22442,
+ "image_id": 4119,
+ "bbox": [
+ 59,
+ 99,
+ 73,
+ 58
+ ],
+ "category_id": 6,
+ "area": 37632,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22443,
+ "image_id": 4119,
+ "bbox": [
+ 179,
+ 22,
+ 119,
+ 118
+ ],
+ "category_id": 6,
+ "area": 122718,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22444,
+ "image_id": 4119,
+ "bbox": [
+ 193,
+ 172,
+ 72,
+ 96
+ ],
+ "category_id": 6,
+ "area": 60996,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22445,
+ "image_id": 4119,
+ "bbox": [
+ 255,
+ 1,
+ 92,
+ 104
+ ],
+ "category_id": 6,
+ "area": 83738,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22446,
+ "image_id": 4119,
+ "bbox": [
+ 387,
+ 88,
+ 88,
+ 68
+ ],
+ "category_id": 6,
+ "area": 52332,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22447,
+ "image_id": 4119,
+ "bbox": [
+ 192,
+ 342,
+ 100,
+ 99
+ ],
+ "category_id": 6,
+ "area": 86598,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22448,
+ "image_id": 4119,
+ "bbox": [
+ 93,
+ 147,
+ 71,
+ 109
+ ],
+ "category_id": 6,
+ "area": 67704,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22449,
+ "image_id": 4119,
+ "bbox": [
+ 310,
+ 287,
+ 165,
+ 166
+ ],
+ "category_id": 6,
+ "area": 239428,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22458,
+ "image_id": 4121,
+ "bbox": [
+ 117,
+ 145,
+ 263,
+ 205
+ ],
+ "category_id": 9,
+ "area": 189419,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22461,
+ "image_id": 4122,
+ "bbox": [
+ 122,
+ 85,
+ 49,
+ 64
+ ],
+ "category_id": 8,
+ "area": 21712,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22462,
+ "image_id": 4122,
+ "bbox": [
+ 110,
+ 143,
+ 86,
+ 67
+ ],
+ "category_id": 8,
+ "area": 39852,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22463,
+ "image_id": 4122,
+ "bbox": [
+ 206,
+ 81,
+ 24,
+ 36
+ ],
+ "category_id": 8,
+ "area": 6072,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22464,
+ "image_id": 4122,
+ "bbox": [
+ 237,
+ 199,
+ 53,
+ 122
+ ],
+ "category_id": 8,
+ "area": 44622,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22465,
+ "image_id": 4122,
+ "bbox": [
+ 321,
+ 116,
+ 48,
+ 79
+ ],
+ "category_id": 8,
+ "area": 26535,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22466,
+ "image_id": 4123,
+ "bbox": [
+ 0,
+ 275,
+ 262,
+ 165
+ ],
+ "category_id": 8,
+ "area": 151264,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22467,
+ "image_id": 4123,
+ "bbox": [
+ 216,
+ 125,
+ 223,
+ 276
+ ],
+ "category_id": 8,
+ "area": 214230,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22497,
+ "image_id": 4124,
+ "bbox": [
+ 0,
+ 15,
+ 148,
+ 163
+ ],
+ "category_id": 8,
+ "area": 192165,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22498,
+ "image_id": 4124,
+ "bbox": [
+ 137,
+ 26,
+ 144,
+ 178
+ ],
+ "category_id": 8,
+ "area": 203792,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22499,
+ "image_id": 4124,
+ "bbox": [
+ 173,
+ 90,
+ 147,
+ 116
+ ],
+ "category_id": 8,
+ "area": 136284,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22519,
+ "image_id": 4130,
+ "bbox": [
+ 100,
+ 170,
+ 364,
+ 212
+ ],
+ "category_id": 9,
+ "area": 49128,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22527,
+ "image_id": 4132,
+ "bbox": [
+ 142,
+ 49,
+ 368,
+ 460
+ ],
+ "category_id": 10,
+ "area": 1343304,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22528,
+ "image_id": 4133,
+ "bbox": [
+ 79,
+ 152,
+ 123,
+ 67
+ ],
+ "category_id": 8,
+ "area": 29260,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22529,
+ "image_id": 4133,
+ "bbox": [
+ 252,
+ 194,
+ 143,
+ 140
+ ],
+ "category_id": 8,
+ "area": 70526,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22534,
+ "image_id": 4135,
+ "bbox": [
+ 185,
+ 155,
+ 278,
+ 356
+ ],
+ "category_id": 8,
+ "area": 115551,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22536,
+ "image_id": 4136,
+ "bbox": [
+ 16,
+ 213,
+ 148,
+ 175
+ ],
+ "category_id": 6,
+ "area": 231361,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22537,
+ "image_id": 4136,
+ "bbox": [
+ 165,
+ 55,
+ 255,
+ 280
+ ],
+ "category_id": 6,
+ "area": 638330,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22538,
+ "image_id": 4137,
+ "bbox": [
+ 120,
+ 35,
+ 197,
+ 476
+ ],
+ "category_id": 9,
+ "area": 238800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22539,
+ "image_id": 4137,
+ "bbox": [
+ 257,
+ 121,
+ 177,
+ 322
+ ],
+ "category_id": 10,
+ "area": 145036,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22540,
+ "image_id": 4138,
+ "bbox": [
+ 88,
+ 0,
+ 57,
+ 51
+ ],
+ "category_id": 6,
+ "area": 6900,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22541,
+ "image_id": 4138,
+ "bbox": [
+ 232,
+ 146,
+ 208,
+ 213
+ ],
+ "category_id": 9,
+ "area": 102808,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22552,
+ "image_id": 4140,
+ "bbox": [
+ 131,
+ 326,
+ 242,
+ 185
+ ],
+ "category_id": 6,
+ "area": 157905,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22553,
+ "image_id": 4140,
+ "bbox": [
+ 0,
+ 270,
+ 118,
+ 196
+ ],
+ "category_id": 6,
+ "area": 81420,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22556,
+ "image_id": 4140,
+ "bbox": [
+ 352,
+ 265,
+ 69,
+ 98
+ ],
+ "category_id": 6,
+ "area": 24186,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22557,
+ "image_id": 4140,
+ "bbox": [
+ 409,
+ 267,
+ 102,
+ 221
+ ],
+ "category_id": 6,
+ "area": 80184,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22558,
+ "image_id": 4140,
+ "bbox": [
+ 0,
+ 147,
+ 192,
+ 249
+ ],
+ "category_id": 6,
+ "area": 168480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22559,
+ "image_id": 4140,
+ "bbox": [
+ 204,
+ 246,
+ 153,
+ 133
+ ],
+ "category_id": 6,
+ "area": 72192,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22561,
+ "image_id": 4140,
+ "bbox": [
+ 242,
+ 278,
+ 109,
+ 159
+ ],
+ "category_id": 8,
+ "area": 61376,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22563,
+ "image_id": 4142,
+ "bbox": [
+ 51,
+ 231,
+ 389,
+ 204
+ ],
+ "category_id": 8,
+ "area": 280224,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22564,
+ "image_id": 4142,
+ "bbox": [
+ 31,
+ 0,
+ 292,
+ 254
+ ],
+ "category_id": 8,
+ "area": 262056,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22603,
+ "image_id": 4147,
+ "bbox": [
+ 125,
+ 340,
+ 253,
+ 171
+ ],
+ "category_id": 9,
+ "area": 48552,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22605,
+ "image_id": 4149,
+ "bbox": [
+ 1,
+ 135,
+ 358,
+ 374
+ ],
+ "category_id": 9,
+ "area": 471296,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22606,
+ "image_id": 4149,
+ "bbox": [
+ 190,
+ 0,
+ 320,
+ 157
+ ],
+ "category_id": 9,
+ "area": 176800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22607,
+ "image_id": 4149,
+ "bbox": [
+ 190,
+ 50,
+ 318,
+ 344
+ ],
+ "category_id": 9,
+ "area": 386545,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22627,
+ "image_id": 4151,
+ "bbox": [
+ 83,
+ 315,
+ 131,
+ 132
+ ],
+ "category_id": 9,
+ "area": 27750,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22629,
+ "image_id": 4152,
+ "bbox": [
+ 132,
+ 404,
+ 198,
+ 102
+ ],
+ "category_id": 8,
+ "area": 140556,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22630,
+ "image_id": 4152,
+ "bbox": [
+ 29,
+ 180,
+ 480,
+ 321
+ ],
+ "category_id": 8,
+ "area": 1063172,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22631,
+ "image_id": 4152,
+ "bbox": [
+ 260,
+ 9,
+ 170,
+ 329
+ ],
+ "category_id": 8,
+ "area": 387138,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22652,
+ "image_id": 4154,
+ "bbox": [
+ 0,
+ 340,
+ 81,
+ 168
+ ],
+ "category_id": 6,
+ "area": 76104,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22653,
+ "image_id": 4155,
+ "bbox": [
+ 65,
+ 182,
+ 171,
+ 275
+ ],
+ "category_id": 10,
+ "area": 77480,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22654,
+ "image_id": 4155,
+ "bbox": [
+ 201,
+ 448,
+ 48,
+ 63
+ ],
+ "category_id": 10,
+ "area": 5100,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22655,
+ "image_id": 4155,
+ "bbox": [
+ 348,
+ 334,
+ 86,
+ 113
+ ],
+ "category_id": 10,
+ "area": 16050,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22656,
+ "image_id": 4155,
+ "bbox": [
+ 357,
+ 395,
+ 107,
+ 115
+ ],
+ "category_id": 10,
+ "area": 20383,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22657,
+ "image_id": 4155,
+ "bbox": [
+ 298,
+ 0,
+ 152,
+ 182
+ ],
+ "category_id": 10,
+ "area": 45752,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22658,
+ "image_id": 4155,
+ "bbox": [
+ 58,
+ 284,
+ 67,
+ 87
+ ],
+ "category_id": 10,
+ "area": 9711,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22659,
+ "image_id": 4155,
+ "bbox": [
+ 116,
+ 447,
+ 60,
+ 53
+ ],
+ "category_id": 10,
+ "area": 5300,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22669,
+ "image_id": 4158,
+ "bbox": [
+ 144,
+ 405,
+ 249,
+ 105
+ ],
+ "category_id": 6,
+ "area": 122928,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22672,
+ "image_id": 4159,
+ "bbox": [
+ 15,
+ 125,
+ 199,
+ 191
+ ],
+ "category_id": 8,
+ "area": 93015,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22673,
+ "image_id": 4159,
+ "bbox": [
+ 138,
+ 182,
+ 134,
+ 184
+ ],
+ "category_id": 8,
+ "area": 60672,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22674,
+ "image_id": 4159,
+ "bbox": [
+ 213,
+ 222,
+ 227,
+ 174
+ ],
+ "category_id": 8,
+ "area": 97042,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22699,
+ "image_id": 4162,
+ "bbox": [
+ 6,
+ 119,
+ 336,
+ 235
+ ],
+ "category_id": 6,
+ "area": 50337,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22700,
+ "image_id": 4162,
+ "bbox": [
+ 273,
+ 372,
+ 139,
+ 110
+ ],
+ "category_id": 6,
+ "area": 9792,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22720,
+ "image_id": 4167,
+ "bbox": [
+ 104,
+ 180,
+ 407,
+ 269
+ ],
+ "category_id": 9,
+ "area": 185176,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22721,
+ "image_id": 4167,
+ "bbox": [
+ 98,
+ 284,
+ 412,
+ 226
+ ],
+ "category_id": 6,
+ "area": 157440,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22722,
+ "image_id": 4167,
+ "bbox": [
+ 375,
+ 95,
+ 136,
+ 174
+ ],
+ "category_id": 6,
+ "area": 40090,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22723,
+ "image_id": 4168,
+ "bbox": [
+ 92,
+ 0,
+ 346,
+ 258
+ ],
+ "category_id": 9,
+ "area": 85410,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22724,
+ "image_id": 4168,
+ "bbox": [
+ 286,
+ 232,
+ 225,
+ 276
+ ],
+ "category_id": 9,
+ "area": 59436,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22744,
+ "image_id": 4170,
+ "bbox": [
+ 108,
+ 83,
+ 53,
+ 54
+ ],
+ "category_id": 8,
+ "area": 19701,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22745,
+ "image_id": 4170,
+ "bbox": [
+ 116,
+ 144,
+ 86,
+ 72
+ ],
+ "category_id": 8,
+ "area": 42768,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22746,
+ "image_id": 4170,
+ "bbox": [
+ 233,
+ 185,
+ 53,
+ 122
+ ],
+ "category_id": 8,
+ "area": 44823,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22747,
+ "image_id": 4170,
+ "bbox": [
+ 215,
+ 77,
+ 21,
+ 35
+ ],
+ "category_id": 8,
+ "area": 5184,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22748,
+ "image_id": 4170,
+ "bbox": [
+ 320,
+ 116,
+ 52,
+ 73
+ ],
+ "category_id": 8,
+ "area": 26532,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22749,
+ "image_id": 4171,
+ "bbox": [
+ 163,
+ 184,
+ 134,
+ 203
+ ],
+ "category_id": 10,
+ "area": 88944,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22750,
+ "image_id": 4171,
+ "bbox": [
+ 250,
+ 119,
+ 118,
+ 156
+ ],
+ "category_id": 9,
+ "area": 60192,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22790,
+ "image_id": 4176,
+ "bbox": [
+ 56,
+ 73,
+ 136,
+ 259
+ ],
+ "category_id": 8,
+ "area": 124830,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22791,
+ "image_id": 4176,
+ "bbox": [
+ 46,
+ 221,
+ 215,
+ 251
+ ],
+ "category_id": 8,
+ "area": 190267,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22792,
+ "image_id": 4176,
+ "bbox": [
+ 299,
+ 113,
+ 212,
+ 166
+ ],
+ "category_id": 8,
+ "area": 124254,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22793,
+ "image_id": 4176,
+ "bbox": [
+ 411,
+ 0,
+ 100,
+ 177
+ ],
+ "category_id": 6,
+ "area": 62499,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22809,
+ "image_id": 4179,
+ "bbox": [
+ 50,
+ 0,
+ 199,
+ 245
+ ],
+ "category_id": 6,
+ "area": 61152,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22810,
+ "image_id": 4179,
+ "bbox": [
+ 102,
+ 206,
+ 346,
+ 274
+ ],
+ "category_id": 9,
+ "area": 118088,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22811,
+ "image_id": 4180,
+ "bbox": [
+ 156,
+ 243,
+ 90,
+ 63
+ ],
+ "category_id": 8,
+ "area": 16720,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22812,
+ "image_id": 4180,
+ "bbox": [
+ 252,
+ 277,
+ 74,
+ 55
+ ],
+ "category_id": 8,
+ "area": 12168,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22814,
+ "image_id": 4181,
+ "bbox": [
+ 76,
+ 241,
+ 203,
+ 135
+ ],
+ "category_id": 8,
+ "area": 63903,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22815,
+ "image_id": 4181,
+ "bbox": [
+ 179,
+ 103,
+ 188,
+ 113
+ ],
+ "category_id": 8,
+ "area": 49800,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22817,
+ "image_id": 4183,
+ "bbox": [
+ 2,
+ 0,
+ 372,
+ 511
+ ],
+ "category_id": 6,
+ "area": 666596,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22818,
+ "image_id": 4183,
+ "bbox": [
+ 180,
+ 256,
+ 252,
+ 225
+ ],
+ "category_id": 8,
+ "area": 199396,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22819,
+ "image_id": 4183,
+ "bbox": [
+ 204,
+ 9,
+ 289,
+ 290
+ ],
+ "category_id": 8,
+ "area": 294261,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22820,
+ "image_id": 4184,
+ "bbox": [
+ 0,
+ 28,
+ 25,
+ 85
+ ],
+ "category_id": 8,
+ "area": 14570,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22821,
+ "image_id": 4184,
+ "bbox": [
+ 87,
+ 91,
+ 132,
+ 142
+ ],
+ "category_id": 8,
+ "area": 127452,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22822,
+ "image_id": 4184,
+ "bbox": [
+ 139,
+ 173,
+ 110,
+ 207
+ ],
+ "category_id": 8,
+ "area": 154912,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22823,
+ "image_id": 4184,
+ "bbox": [
+ 289,
+ 72,
+ 52,
+ 62
+ ],
+ "category_id": 8,
+ "area": 22458,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22824,
+ "image_id": 4184,
+ "bbox": [
+ 317,
+ 161,
+ 104,
+ 163
+ ],
+ "category_id": 8,
+ "area": 116424,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22825,
+ "image_id": 4184,
+ "bbox": [
+ 417,
+ 177,
+ 45,
+ 85
+ ],
+ "category_id": 8,
+ "area": 26195,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22830,
+ "image_id": 4185,
+ "bbox": [
+ 103,
+ 474,
+ 19,
+ 37
+ ],
+ "category_id": 6,
+ "area": 5120,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22831,
+ "image_id": 4185,
+ "bbox": [
+ 348,
+ 256,
+ 19,
+ 47
+ ],
+ "category_id": 6,
+ "area": 6600,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22832,
+ "image_id": 4185,
+ "bbox": [
+ 331,
+ 101,
+ 30,
+ 44
+ ],
+ "category_id": 6,
+ "area": 9486,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22833,
+ "image_id": 4185,
+ "bbox": [
+ 419,
+ 0,
+ 26,
+ 35
+ ],
+ "category_id": 6,
+ "area": 6586,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22834,
+ "image_id": 4185,
+ "bbox": [
+ 452,
+ 305,
+ 20,
+ 50
+ ],
+ "category_id": 6,
+ "area": 7383,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22837,
+ "image_id": 4187,
+ "bbox": [
+ 268,
+ 85,
+ 216,
+ 268
+ ],
+ "category_id": 9,
+ "area": 63700,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22841,
+ "image_id": 4189,
+ "bbox": [
+ 80,
+ 189,
+ 162,
+ 200
+ ],
+ "category_id": 8,
+ "area": 113805,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22842,
+ "image_id": 4189,
+ "bbox": [
+ 244,
+ 192,
+ 214,
+ 214
+ ],
+ "category_id": 8,
+ "area": 161336,
+ "iscrowd": 0,
+ "ignore": 0
+ },
+ {
+ "id": 22966,
+ "image_id": 4200,
+ "bbox": [
+ 0,
+ 0,
+ 488,
+ 293
+ ],
+ "category_id": 6,
+ "area": 360426,
+ "iscrowd": 0,
+ "ignore": 0
+ }
+ ],
+ "categories": [
+ {
+ "id": 1,
+ "name": "holothurian"
+ },
+ {
+ "id": 2,
+ "name": "echinus"
+ },
+ {
+ "id": 3,
+ "name": "scallop"
+ },
+ {
+ "id": 4,
+ "name": "starfish"
+ },
+ {
+ "id": 5,
+ "name": "fish"
+ },
+ {
+ "id": 6,
+ "name": "corals"
+ },
+ {
+ "id": 7,
+ "name": "diver"
+ },
+ {
+ "id": 8,
+ "name": "cuttlefish"
+ },
+ {
+ "id": 9,
+ "name": "turtle"
+ },
+ {
+ "id": 10,
+ "name": "jellyfish"
+ }
+ ]
+}
\ No newline at end of file
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..3ac884ac8b40c1543ed840dfcafe367fbe4bda62
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/__init__.py
@@ -0,0 +1,27 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import mmcv
+import mmengine
+from mmengine.utils import digit_version
+
+from .version import __version__, version_info
+
+mmcv_minimum_version = '2.0.0rc4'
+mmcv_maximum_version = '2.2.0'
+mmcv_version = digit_version(mmcv.__version__)
+
+mmengine_minimum_version = '0.7.1'
+mmengine_maximum_version = '1.0.0'
+mmengine_version = digit_version(mmengine.__version__)
+
+assert (mmcv_version >= digit_version(mmcv_minimum_version)
+ and mmcv_version < digit_version(mmcv_maximum_version)), \
+ f'MMCV=={mmcv.__version__} is used but incompatible. ' \
+ f'Please install mmcv>={mmcv_minimum_version}, <{mmcv_maximum_version}.'
+
+assert (mmengine_version >= digit_version(mmengine_minimum_version)
+ and mmengine_version < digit_version(mmengine_maximum_version)), \
+ f'MMEngine=={mmengine.__version__} is used but incompatible. ' \
+ f'Please install mmengine>={mmengine_minimum_version}, ' \
+ f'<{mmengine_maximum_version}.'
+
+__all__ = ['__version__', 'version_info', 'digit_version']
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/apis/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/apis/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..c89dc72914b11a73e91dc7e9404f41bf10b93c6c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/apis/__init__.py
@@ -0,0 +1,9 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .det_inferencer import DetInferencer
+from .inference import (async_inference_detector, inference_detector,
+ inference_mot, init_detector, init_track_model)
+
+__all__ = [
+ 'init_detector', 'async_inference_detector', 'inference_detector',
+ 'DetInferencer', 'inference_mot', 'init_track_model'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/apis/det_inferencer.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/apis/det_inferencer.py
new file mode 100644
index 0000000000000000000000000000000000000000..ce8532eb786558ca3807195781d8e380741cea00
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/apis/det_inferencer.py
@@ -0,0 +1,652 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import os.path as osp
+import warnings
+from typing import Dict, Iterable, List, Optional, Sequence, Tuple, Union
+
+import mmcv
+import mmengine
+import numpy as np
+import torch.nn as nn
+from mmcv.transforms import LoadImageFromFile
+from mmengine.dataset import Compose
+from mmengine.fileio import (get_file_backend, isdir, join_path,
+ list_dir_or_file)
+from mmengine.infer.infer import BaseInferencer, ModelType
+from mmengine.model.utils import revert_sync_batchnorm
+from mmengine.registry import init_default_scope
+from mmengine.runner.checkpoint import _load_checkpoint_to_model
+from mmengine.visualization import Visualizer
+from rich.progress import track
+
+from mmdet.evaluation import INSTANCE_OFFSET
+from mmdet.registry import DATASETS
+from mmdet.structures import DetDataSample
+from mmdet.structures.mask import encode_mask_results, mask2bbox
+from mmdet.utils import ConfigType
+from ..evaluation import get_classes
+
+try:
+ from panopticapi.evaluation import VOID
+ from panopticapi.utils import id2rgb
+except ImportError:
+ id2rgb = None
+ VOID = None
+
+InputType = Union[str, np.ndarray]
+InputsType = Union[InputType, Sequence[InputType]]
+PredType = List[DetDataSample]
+ImgType = Union[np.ndarray, Sequence[np.ndarray]]
+
+IMG_EXTENSIONS = ('.jpg', '.jpeg', '.png', '.ppm', '.bmp', '.pgm', '.tif',
+ '.tiff', '.webp')
+
+
+class DetInferencer(BaseInferencer):
+ """Object Detection Inferencer.
+
+ Args:
+ model (str, optional): Path to the config file or the model name
+ defined in metafile. For example, it could be
+ "rtmdet-s" or 'rtmdet_s_8xb32-300e_coco' or
+ "configs/rtmdet/rtmdet_s_8xb32-300e_coco.py".
+ If model is not specified, user must provide the
+ `weights` saved by MMEngine which contains the config string.
+ Defaults to None.
+ weights (str, optional): Path to the checkpoint. If it is not specified
+ and model is a model name of metafile, the weights will be loaded
+ from metafile. Defaults to None.
+ device (str, optional): Device to run inference. If None, the available
+ device will be automatically used. Defaults to None.
+ scope (str, optional): The scope of the model. Defaults to mmdet.
+ palette (str): Color palette used for visualization. The order of
+ priority is palette -> config -> checkpoint. Defaults to 'none'.
+ show_progress (bool): Control whether to display the progress
+ bar during the inference process. Defaults to True.
+ """
+
+ preprocess_kwargs: set = set()
+ forward_kwargs: set = set()
+ visualize_kwargs: set = {
+ 'return_vis',
+ 'show',
+ 'wait_time',
+ 'draw_pred',
+ 'pred_score_thr',
+ 'img_out_dir',
+ 'no_save_vis',
+ }
+ postprocess_kwargs: set = {
+ 'print_result',
+ 'pred_out_dir',
+ 'return_datasamples',
+ 'no_save_pred',
+ }
+
+ def __init__(self,
+ model: Optional[Union[ModelType, str]] = None,
+ weights: Optional[str] = None,
+ device: Optional[str] = None,
+ scope: Optional[str] = 'mmdet',
+ palette: str = 'none',
+ show_progress: bool = True) -> None:
+ # A global counter tracking the number of images processed, for
+ # naming of the output images
+ self.num_visualized_imgs = 0
+ self.num_predicted_imgs = 0
+ self.palette = palette
+ init_default_scope(scope)
+ super().__init__(
+ model=model, weights=weights, device=device, scope=scope)
+ self.model = revert_sync_batchnorm(self.model)
+ self.show_progress = show_progress
+
+ def _load_weights_to_model(self, model: nn.Module,
+ checkpoint: Optional[dict],
+ cfg: Optional[ConfigType]) -> None:
+ """Loading model weights and meta information from cfg and checkpoint.
+
+ Args:
+ model (nn.Module): Model to load weights and meta information.
+ checkpoint (dict, optional): The loaded checkpoint.
+ cfg (Config or ConfigDict, optional): The loaded config.
+ """
+
+ if checkpoint is not None:
+ _load_checkpoint_to_model(model, checkpoint)
+ checkpoint_meta = checkpoint.get('meta', {})
+ # save the dataset_meta in the model for convenience
+ if 'dataset_meta' in checkpoint_meta:
+ # mmdet 3.x, all keys should be lowercase
+ model.dataset_meta = {
+ k.lower(): v
+ for k, v in checkpoint_meta['dataset_meta'].items()
+ }
+ elif 'CLASSES' in checkpoint_meta:
+ # < mmdet 3.x
+ classes = checkpoint_meta['CLASSES']
+ model.dataset_meta = {'classes': classes}
+ else:
+ warnings.warn(
+ 'dataset_meta or class names are not saved in the '
+ 'checkpoint\'s meta data, use COCO classes by default.')
+ model.dataset_meta = {'classes': get_classes('coco')}
+ else:
+ warnings.warn('Checkpoint is not loaded, and the inference '
+ 'result is calculated by the randomly initialized '
+ 'model!')
+ warnings.warn('weights is None, use COCO classes by default.')
+ model.dataset_meta = {'classes': get_classes('coco')}
+
+ # Priority: args.palette -> config -> checkpoint
+ if self.palette != 'none':
+ model.dataset_meta['palette'] = self.palette
+ else:
+ test_dataset_cfg = copy.deepcopy(cfg.test_dataloader.dataset)
+ # lazy init. We only need the metainfo.
+ test_dataset_cfg['lazy_init'] = True
+ metainfo = DATASETS.build(test_dataset_cfg).metainfo
+ cfg_palette = metainfo.get('palette', None)
+ if cfg_palette is not None:
+ model.dataset_meta['palette'] = cfg_palette
+ else:
+ if 'palette' not in model.dataset_meta:
+ warnings.warn(
+ 'palette does not exist, random is used by default. '
+ 'You can also set the palette to customize.')
+ model.dataset_meta['palette'] = 'random'
+
+ def _init_pipeline(self, cfg: ConfigType) -> Compose:
+ """Initialize the test pipeline."""
+ pipeline_cfg = cfg.test_dataloader.dataset.pipeline
+
+ # For inference, the key of ``img_id`` is not used.
+ if 'meta_keys' in pipeline_cfg[-1]:
+ pipeline_cfg[-1]['meta_keys'] = tuple(
+ meta_key for meta_key in pipeline_cfg[-1]['meta_keys']
+ if meta_key != 'img_id')
+
+ load_img_idx = self._get_transform_idx(
+ pipeline_cfg, ('LoadImageFromFile', LoadImageFromFile))
+ if load_img_idx == -1:
+ raise ValueError(
+ 'LoadImageFromFile is not found in the test pipeline')
+ pipeline_cfg[load_img_idx]['type'] = 'mmdet.InferencerLoader'
+ return Compose(pipeline_cfg)
+
+ def _get_transform_idx(self, pipeline_cfg: ConfigType,
+ name: Union[str, Tuple[str, type]]) -> int:
+ """Returns the index of the transform in a pipeline.
+
+ If the transform is not found, returns -1.
+ """
+ for i, transform in enumerate(pipeline_cfg):
+ if transform['type'] in name:
+ return i
+ return -1
+
+ def _init_visualizer(self, cfg: ConfigType) -> Optional[Visualizer]:
+ """Initialize visualizers.
+
+ Args:
+ cfg (ConfigType): Config containing the visualizer information.
+
+ Returns:
+ Visualizer or None: Visualizer initialized with config.
+ """
+ visualizer = super()._init_visualizer(cfg)
+ visualizer.dataset_meta = self.model.dataset_meta
+ return visualizer
+
+ def _inputs_to_list(self, inputs: InputsType) -> list:
+ """Preprocess the inputs to a list.
+
+ Preprocess inputs to a list according to its type:
+
+ - list or tuple: return inputs
+ - str:
+ - Directory path: return all files in the directory
+ - other cases: return a list containing the string. The string
+ could be a path to file, a url or other types of string according
+ to the task.
+
+ Args:
+ inputs (InputsType): Inputs for the inferencer.
+
+ Returns:
+ list: List of input for the :meth:`preprocess`.
+ """
+ if isinstance(inputs, str):
+ backend = get_file_backend(inputs)
+ if hasattr(backend, 'isdir') and isdir(inputs):
+ # Backends like HttpsBackend do not implement `isdir`, so only
+ # those backends that implement `isdir` could accept the inputs
+ # as a directory
+ filename_list = list_dir_or_file(
+ inputs, list_dir=False, suffix=IMG_EXTENSIONS)
+ inputs = [
+ join_path(inputs, filename) for filename in filename_list
+ ]
+
+ if not isinstance(inputs, (list, tuple)):
+ inputs = [inputs]
+
+ return list(inputs)
+
+ def preprocess(self, inputs: InputsType, batch_size: int = 1, **kwargs):
+ """Process the inputs into a model-feedable format.
+
+ Customize your preprocess by overriding this method. Preprocess should
+ return an iterable object, of which each item will be used as the
+ input of ``model.test_step``.
+
+ ``BaseInferencer.preprocess`` will return an iterable chunked data,
+ which will be used in __call__ like this:
+
+ .. code-block:: python
+
+ def __call__(self, inputs, batch_size=1, **kwargs):
+ chunked_data = self.preprocess(inputs, batch_size, **kwargs)
+ for batch in chunked_data:
+ preds = self.forward(batch, **kwargs)
+
+ Args:
+ inputs (InputsType): Inputs given by user.
+ batch_size (int): batch size. Defaults to 1.
+
+ Yields:
+ Any: Data processed by the ``pipeline`` and ``collate_fn``.
+ """
+ chunked_data = self._get_chunk_data(inputs, batch_size)
+ yield from map(self.collate_fn, chunked_data)
+
+ def _get_chunk_data(self, inputs: Iterable, chunk_size: int):
+ """Get batch data from inputs.
+
+ Args:
+ inputs (Iterable): An iterable dataset.
+ chunk_size (int): Equivalent to batch size.
+
+ Yields:
+ list: batch data.
+ """
+ inputs_iter = iter(inputs)
+ while True:
+ try:
+ chunk_data = []
+ for _ in range(chunk_size):
+ inputs_ = next(inputs_iter)
+ if isinstance(inputs_, dict):
+ if 'img' in inputs_:
+ ori_inputs_ = inputs_['img']
+ else:
+ ori_inputs_ = inputs_['img_path']
+ chunk_data.append(
+ (ori_inputs_,
+ self.pipeline(copy.deepcopy(inputs_))))
+ else:
+ chunk_data.append((inputs_, self.pipeline(inputs_)))
+ yield chunk_data
+ except StopIteration:
+ if chunk_data:
+ yield chunk_data
+ break
+
+ # TODO: Video and Webcam are currently not supported and
+ # may consume too much memory if your input folder has a lot of images.
+ # We will be optimized later.
+ def __call__(
+ self,
+ inputs: InputsType,
+ batch_size: int = 1,
+ return_vis: bool = False,
+ show: bool = False,
+ wait_time: int = 0,
+ no_save_vis: bool = False,
+ draw_pred: bool = True,
+ pred_score_thr: float = 0.3,
+ return_datasamples: bool = False,
+ print_result: bool = False,
+ no_save_pred: bool = True,
+ out_dir: str = '',
+ # by open image task
+ texts: Optional[Union[str, list]] = None,
+ # by open panoptic task
+ stuff_texts: Optional[Union[str, list]] = None,
+ # by GLIP and Grounding DINO
+ custom_entities: bool = False,
+ # by Grounding DINO
+ tokens_positive: Optional[Union[int, list]] = None,
+ **kwargs) -> dict:
+ """Call the inferencer.
+
+ Args:
+ inputs (InputsType): Inputs for the inferencer.
+ batch_size (int): Inference batch size. Defaults to 1.
+ show (bool): Whether to display the visualization results in a
+ popup window. Defaults to False.
+ wait_time (float): The interval of show (s). Defaults to 0.
+ no_save_vis (bool): Whether to force not to save prediction
+ vis results. Defaults to False.
+ draw_pred (bool): Whether to draw predicted bounding boxes.
+ Defaults to True.
+ pred_score_thr (float): Minimum score of bboxes to draw.
+ Defaults to 0.3.
+ return_datasamples (bool): Whether to return results as
+ :obj:`DetDataSample`. Defaults to False.
+ print_result (bool): Whether to print the inference result w/o
+ visualization to the console. Defaults to False.
+ no_save_pred (bool): Whether to force not to save prediction
+ results. Defaults to True.
+ out_dir: Dir to save the inference results or
+ visualization. If left as empty, no file will be saved.
+ Defaults to ''.
+ texts (str | list[str]): Text prompts. Defaults to None.
+ stuff_texts (str | list[str]): Stuff text prompts of open
+ panoptic task. Defaults to None.
+ custom_entities (bool): Whether to use custom entities.
+ Defaults to False. Only used in GLIP and Grounding DINO.
+ **kwargs: Other keyword arguments passed to :meth:`preprocess`,
+ :meth:`forward`, :meth:`visualize` and :meth:`postprocess`.
+ Each key in kwargs should be in the corresponding set of
+ ``preprocess_kwargs``, ``forward_kwargs``, ``visualize_kwargs``
+ and ``postprocess_kwargs``.
+
+ Returns:
+ dict: Inference and visualization results.
+ """
+ (
+ preprocess_kwargs,
+ forward_kwargs,
+ visualize_kwargs,
+ postprocess_kwargs,
+ ) = self._dispatch_kwargs(**kwargs)
+
+ ori_inputs = self._inputs_to_list(inputs)
+
+ if texts is not None and isinstance(texts, str):
+ texts = [texts] * len(ori_inputs)
+ if stuff_texts is not None and isinstance(stuff_texts, str):
+ stuff_texts = [stuff_texts] * len(ori_inputs)
+
+ # Currently only supports bs=1
+ tokens_positive = [tokens_positive] * len(ori_inputs)
+
+ if texts is not None:
+ assert len(texts) == len(ori_inputs)
+ for i in range(len(texts)):
+ if isinstance(ori_inputs[i], str):
+ ori_inputs[i] = {
+ 'text': texts[i],
+ 'img_path': ori_inputs[i],
+ 'custom_entities': custom_entities,
+ 'tokens_positive': tokens_positive[i]
+ }
+ else:
+ ori_inputs[i] = {
+ 'text': texts[i],
+ 'img': ori_inputs[i],
+ 'custom_entities': custom_entities,
+ 'tokens_positive': tokens_positive[i]
+ }
+ if stuff_texts is not None:
+ assert len(stuff_texts) == len(ori_inputs)
+ for i in range(len(stuff_texts)):
+ ori_inputs[i]['stuff_text'] = stuff_texts[i]
+
+ inputs = self.preprocess(
+ ori_inputs, batch_size=batch_size, **preprocess_kwargs)
+
+ results_dict = {'predictions': [], 'visualization': []}
+ for ori_imgs, data in (track(inputs, description='Inference')
+ if self.show_progress else inputs):
+ preds = self.forward(data, **forward_kwargs)
+ visualization = self.visualize(
+ ori_imgs,
+ preds,
+ return_vis=return_vis,
+ show=show,
+ wait_time=wait_time,
+ draw_pred=draw_pred,
+ pred_score_thr=pred_score_thr,
+ no_save_vis=no_save_vis,
+ img_out_dir=out_dir,
+ **visualize_kwargs)
+ results = self.postprocess(
+ preds,
+ visualization,
+ return_datasamples=return_datasamples,
+ print_result=print_result,
+ no_save_pred=no_save_pred,
+ pred_out_dir=out_dir,
+ **postprocess_kwargs)
+ results_dict['predictions'].extend(results['predictions'])
+ if results['visualization'] is not None:
+ results_dict['visualization'].extend(results['visualization'])
+ return results_dict
+
+ def visualize(self,
+ inputs: InputsType,
+ preds: PredType,
+ return_vis: bool = False,
+ show: bool = False,
+ wait_time: int = 0,
+ draw_pred: bool = True,
+ pred_score_thr: float = 0.3,
+ no_save_vis: bool = False,
+ img_out_dir: str = '',
+ **kwargs) -> Union[List[np.ndarray], None]:
+ """Visualize predictions.
+
+ Args:
+ inputs (List[Union[str, np.ndarray]]): Inputs for the inferencer.
+ preds (List[:obj:`DetDataSample`]): Predictions of the model.
+ return_vis (bool): Whether to return the visualization result.
+ Defaults to False.
+ show (bool): Whether to display the image in a popup window.
+ Defaults to False.
+ wait_time (float): The interval of show (s). Defaults to 0.
+ draw_pred (bool): Whether to draw predicted bounding boxes.
+ Defaults to True.
+ pred_score_thr (float): Minimum score of bboxes to draw.
+ Defaults to 0.3.
+ no_save_vis (bool): Whether to force not to save prediction
+ vis results. Defaults to False.
+ img_out_dir (str): Output directory of visualization results.
+ If left as empty, no file will be saved. Defaults to ''.
+
+ Returns:
+ List[np.ndarray] or None: Returns visualization results only if
+ applicable.
+ """
+ if no_save_vis is True:
+ img_out_dir = ''
+
+ if not show and img_out_dir == '' and not return_vis:
+ return None
+
+ if self.visualizer is None:
+ raise ValueError('Visualization needs the "visualizer" term'
+ 'defined in the config, but got None.')
+
+ results = []
+
+ for single_input, pred in zip(inputs, preds):
+ if isinstance(single_input, str):
+ img_bytes = mmengine.fileio.get(single_input)
+ img = mmcv.imfrombytes(img_bytes)
+ img = img[:, :, ::-1]
+ img_name = osp.basename(single_input)
+ elif isinstance(single_input, np.ndarray):
+ img = single_input.copy()
+ img_num = str(self.num_visualized_imgs).zfill(8)
+ img_name = f'{img_num}.jpg'
+ else:
+ raise ValueError('Unsupported input type: '
+ f'{type(single_input)}')
+
+ out_file = osp.join(img_out_dir, 'vis',
+ img_name) if img_out_dir != '' else None
+
+ self.visualizer.add_datasample(
+ img_name,
+ img,
+ pred,
+ show=show,
+ wait_time=wait_time,
+ draw_gt=False,
+ draw_pred=draw_pred,
+ pred_score_thr=pred_score_thr,
+ out_file=out_file,
+ )
+ results.append(self.visualizer.get_image())
+ self.num_visualized_imgs += 1
+
+ return results
+
+ def postprocess(
+ self,
+ preds: PredType,
+ visualization: Optional[List[np.ndarray]] = None,
+ return_datasamples: bool = False,
+ print_result: bool = False,
+ no_save_pred: bool = False,
+ pred_out_dir: str = '',
+ **kwargs,
+ ) -> Dict:
+ """Process the predictions and visualization results from ``forward``
+ and ``visualize``.
+
+ This method should be responsible for the following tasks:
+
+ 1. Convert datasamples into a json-serializable dict if needed.
+ 2. Pack the predictions and visualization results and return them.
+ 3. Dump or log the predictions.
+
+ Args:
+ preds (List[:obj:`DetDataSample`]): Predictions of the model.
+ visualization (Optional[np.ndarray]): Visualized predictions.
+ return_datasamples (bool): Whether to use Datasample to store
+ inference results. If False, dict will be used.
+ print_result (bool): Whether to print the inference result w/o
+ visualization to the console. Defaults to False.
+ no_save_pred (bool): Whether to force not to save prediction
+ results. Defaults to False.
+ pred_out_dir: Dir to save the inference results w/o
+ visualization. If left as empty, no file will be saved.
+ Defaults to ''.
+
+ Returns:
+ dict: Inference and visualization results with key ``predictions``
+ and ``visualization``.
+
+ - ``visualization`` (Any): Returned by :meth:`visualize`.
+ - ``predictions`` (dict or DataSample): Returned by
+ :meth:`forward` and processed in :meth:`postprocess`.
+ If ``return_datasamples=False``, it usually should be a
+ json-serializable dict containing only basic data elements such
+ as strings and numbers.
+ """
+ if no_save_pred is True:
+ pred_out_dir = ''
+
+ result_dict = {}
+ results = preds
+ if not return_datasamples:
+ results = []
+ for pred in preds:
+ result = self.pred2dict(pred, pred_out_dir)
+ results.append(result)
+ elif pred_out_dir != '':
+ warnings.warn('Currently does not support saving datasample '
+ 'when return_datasamples is set to True. '
+ 'Prediction results are not saved!')
+ # Add img to the results after printing and dumping
+ result_dict['predictions'] = results
+ if print_result:
+ print(result_dict)
+ result_dict['visualization'] = visualization
+ return result_dict
+
+ # TODO: The data format and fields saved in json need further discussion.
+ # Maybe should include model name, timestamp, filename, image info etc.
+ def pred2dict(self,
+ data_sample: DetDataSample,
+ pred_out_dir: str = '') -> Dict:
+ """Extract elements necessary to represent a prediction into a
+ dictionary.
+
+ It's better to contain only basic data elements such as strings and
+ numbers in order to guarantee it's json-serializable.
+
+ Args:
+ data_sample (:obj:`DetDataSample`): Predictions of the model.
+ pred_out_dir: Dir to save the inference results w/o
+ visualization. If left as empty, no file will be saved.
+ Defaults to ''.
+
+ Returns:
+ dict: Prediction results.
+ """
+ is_save_pred = True
+ if pred_out_dir == '':
+ is_save_pred = False
+
+ if is_save_pred and 'img_path' in data_sample:
+ img_path = osp.basename(data_sample.img_path)
+ img_path = osp.splitext(img_path)[0]
+ out_img_path = osp.join(pred_out_dir, 'preds',
+ img_path + '_panoptic_seg.png')
+ out_json_path = osp.join(pred_out_dir, 'preds', img_path + '.json')
+ elif is_save_pred:
+ out_img_path = osp.join(
+ pred_out_dir, 'preds',
+ f'{self.num_predicted_imgs}_panoptic_seg.png')
+ out_json_path = osp.join(pred_out_dir, 'preds',
+ f'{self.num_predicted_imgs}.json')
+ self.num_predicted_imgs += 1
+
+ result = {}
+ if 'pred_instances' in data_sample:
+ masks = data_sample.pred_instances.get('masks')
+ pred_instances = data_sample.pred_instances.numpy()
+ result = {
+ 'labels': pred_instances.labels.tolist(),
+ 'scores': pred_instances.scores.tolist()
+ }
+ if 'bboxes' in pred_instances:
+ result['bboxes'] = pred_instances.bboxes.tolist()
+ if masks is not None:
+ if 'bboxes' not in pred_instances or pred_instances.bboxes.sum(
+ ) == 0:
+ # Fake bbox, such as the SOLO.
+ bboxes = mask2bbox(masks.cpu()).numpy().tolist()
+ result['bboxes'] = bboxes
+ encode_masks = encode_mask_results(pred_instances.masks)
+ for encode_mask in encode_masks:
+ if isinstance(encode_mask['counts'], bytes):
+ encode_mask['counts'] = encode_mask['counts'].decode()
+ result['masks'] = encode_masks
+
+ if 'pred_panoptic_seg' in data_sample:
+ if VOID is None:
+ raise RuntimeError(
+ 'panopticapi is not installed, please install it by: '
+ 'pip install git+https://github.com/cocodataset/'
+ 'panopticapi.git.')
+
+ pan = data_sample.pred_panoptic_seg.sem_seg.cpu().numpy()[0]
+ pan[pan % INSTANCE_OFFSET == len(
+ self.model.dataset_meta['classes'])] = VOID
+ pan = id2rgb(pan).astype(np.uint8)
+
+ if is_save_pred:
+ mmcv.imwrite(pan[:, :, ::-1], out_img_path)
+ result['panoptic_seg_path'] = out_img_path
+ else:
+ result['panoptic_seg'] = pan
+
+ if is_save_pred:
+ mmengine.dump(result, out_json_path)
+
+ return result
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/apis/inference.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/apis/inference.py
new file mode 100644
index 0000000000000000000000000000000000000000..7e6f914ecabf4b9c110a4fd15310bc97d0197db9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/apis/inference.py
@@ -0,0 +1,372 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import warnings
+from pathlib import Path
+from typing import Optional, Sequence, Union
+
+import numpy as np
+import torch
+import torch.nn as nn
+from mmcv.ops import RoIPool
+from mmcv.transforms import Compose
+from mmengine.config import Config
+from mmengine.dataset import default_collate
+from mmengine.model.utils import revert_sync_batchnorm
+from mmengine.registry import init_default_scope
+from mmengine.runner import load_checkpoint
+
+from mmdet.registry import DATASETS
+from mmdet.utils import ConfigType
+from ..evaluation import get_classes
+from ..registry import MODELS
+from ..structures import DetDataSample, SampleList
+from ..utils import get_test_pipeline_cfg
+
+
+def init_detector(
+ config: Union[str, Path, Config],
+ checkpoint: Optional[str] = None,
+ palette: str = 'none',
+ device: str = 'cuda:0',
+ cfg_options: Optional[dict] = None,
+) -> nn.Module:
+ """Initialize a detector from config file.
+
+ Args:
+ config (str, :obj:`Path`, or :obj:`mmengine.Config`): Config file path,
+ :obj:`Path`, or the config object.
+ checkpoint (str, optional): Checkpoint path. If left as None, the model
+ will not load any weights.
+ palette (str): Color palette used for visualization. If palette
+ is stored in checkpoint, use checkpoint's palette first, otherwise
+ use externally passed palette. Currently, supports 'coco', 'voc',
+ 'citys' and 'random'. Defaults to none.
+ device (str): The device where the anchors will be put on.
+ Defaults to cuda:0.
+ cfg_options (dict, optional): Options to override some settings in
+ the used config.
+
+ Returns:
+ nn.Module: The constructed detector.
+ """
+ if isinstance(config, (str, Path)):
+ config = Config.fromfile(config)
+ elif not isinstance(config, Config):
+ raise TypeError('config must be a filename or Config object, '
+ f'but got {type(config)}')
+ if cfg_options is not None:
+ config.merge_from_dict(cfg_options)
+ elif 'init_cfg' in config.model.backbone:
+ config.model.backbone.init_cfg = None
+
+ scope = config.get('default_scope', 'mmdet')
+ if scope is not None:
+ init_default_scope(config.get('default_scope', 'mmdet'))
+
+ model = MODELS.build(config.model)
+ model = revert_sync_batchnorm(model)
+ if checkpoint is None:
+ warnings.simplefilter('once')
+ warnings.warn('checkpoint is None, use COCO classes by default.')
+ model.dataset_meta = {'classes': get_classes('coco')}
+ else:
+ checkpoint = load_checkpoint(model, checkpoint, map_location='cpu')
+ # Weights converted from elsewhere may not have meta fields.
+ checkpoint_meta = checkpoint.get('meta', {})
+
+ # save the dataset_meta in the model for convenience
+ if 'dataset_meta' in checkpoint_meta:
+ # mmdet 3.x, all keys should be lowercase
+ model.dataset_meta = {
+ k.lower(): v
+ for k, v in checkpoint_meta['dataset_meta'].items()
+ }
+ elif 'CLASSES' in checkpoint_meta:
+ # < mmdet 3.x
+ classes = checkpoint_meta['CLASSES']
+ model.dataset_meta = {'classes': classes}
+ else:
+ warnings.simplefilter('once')
+ warnings.warn(
+ 'dataset_meta or class names are not saved in the '
+ 'checkpoint\'s meta data, use COCO classes by default.')
+ model.dataset_meta = {'classes': get_classes('coco')}
+
+ # Priority: args.palette -> config -> checkpoint
+ if palette != 'none':
+ model.dataset_meta['palette'] = palette
+ else:
+ test_dataset_cfg = copy.deepcopy(config.test_dataloader.dataset)
+ # lazy init. We only need the metainfo.
+ test_dataset_cfg['lazy_init'] = True
+ metainfo = DATASETS.build(test_dataset_cfg).metainfo
+ cfg_palette = metainfo.get('palette', None)
+ if cfg_palette is not None:
+ model.dataset_meta['palette'] = cfg_palette
+ else:
+ if 'palette' not in model.dataset_meta:
+ warnings.warn(
+ 'palette does not exist, random is used by default. '
+ 'You can also set the palette to customize.')
+ model.dataset_meta['palette'] = 'random'
+
+ model.cfg = config # save the config in the model for convenience
+ model.to(device)
+ model.eval()
+ return model
+
+
+ImagesType = Union[str, np.ndarray, Sequence[str], Sequence[np.ndarray]]
+
+
+def inference_detector(
+ model: nn.Module,
+ imgs: ImagesType,
+ test_pipeline: Optional[Compose] = None,
+ text_prompt: Optional[str] = None,
+ custom_entities: bool = False,
+) -> Union[DetDataSample, SampleList]:
+ """Inference image(s) with the detector.
+
+ Args:
+ model (nn.Module): The loaded detector.
+ imgs (str, ndarray, Sequence[str/ndarray]):
+ Either image files or loaded images.
+ test_pipeline (:obj:`Compose`): Test pipeline.
+
+ Returns:
+ :obj:`DetDataSample` or list[:obj:`DetDataSample`]:
+ If imgs is a list or tuple, the same length list type results
+ will be returned, otherwise return the detection results directly.
+ """
+
+ if isinstance(imgs, (list, tuple)):
+ is_batch = True
+ else:
+ imgs = [imgs]
+ is_batch = False
+
+ cfg = model.cfg
+
+ if test_pipeline is None:
+ cfg = cfg.copy()
+ test_pipeline = get_test_pipeline_cfg(cfg)
+ if isinstance(imgs[0], np.ndarray):
+ # Calling this method across libraries will result
+ # in module unregistered error if not prefixed with mmdet.
+ test_pipeline[0].type = 'mmdet.LoadImageFromNDArray'
+
+ test_pipeline = Compose(test_pipeline)
+
+ if model.data_preprocessor.device.type == 'cpu':
+ for m in model.modules():
+ assert not isinstance(
+ m, RoIPool
+ ), 'CPU inference with RoIPool is not supported currently.'
+
+ result_list = []
+ for i, img in enumerate(imgs):
+ # prepare data
+ if isinstance(img, np.ndarray):
+ # TODO: remove img_id.
+ data_ = dict(img=img, img_id=0)
+ else:
+ # TODO: remove img_id.
+ data_ = dict(img_path=img, img_id=0)
+
+ if text_prompt:
+ data_['text'] = text_prompt
+ data_['custom_entities'] = custom_entities
+
+ # build the data pipeline
+ data_ = test_pipeline(data_)
+
+ data_['inputs'] = [data_['inputs']]
+ data_['data_samples'] = [data_['data_samples']]
+
+ # forward the model
+ with torch.no_grad():
+ results = model.test_step(data_)[0]
+
+ result_list.append(results)
+
+ if not is_batch:
+ return result_list[0]
+ else:
+ return result_list
+
+
+# TODO: Awaiting refactoring
+async def async_inference_detector(model, imgs):
+ """Async inference image(s) with the detector.
+
+ Args:
+ model (nn.Module): The loaded detector.
+ img (str | ndarray): Either image files or loaded images.
+
+ Returns:
+ Awaitable detection results.
+ """
+ if not isinstance(imgs, (list, tuple)):
+ imgs = [imgs]
+
+ cfg = model.cfg
+
+ if isinstance(imgs[0], np.ndarray):
+ cfg = cfg.copy()
+ # set loading pipeline type
+ cfg.data.test.pipeline[0].type = 'LoadImageFromNDArray'
+
+ # cfg.data.test.pipeline = replace_ImageToTensor(cfg.data.test.pipeline)
+ test_pipeline = Compose(cfg.data.test.pipeline)
+
+ datas = []
+ for img in imgs:
+ # prepare data
+ if isinstance(img, np.ndarray):
+ # directly add img
+ data = dict(img=img)
+ else:
+ # add information into dict
+ data = dict(img_info=dict(filename=img), img_prefix=None)
+ # build the data pipeline
+ data = test_pipeline(data)
+ datas.append(data)
+
+ for m in model.modules():
+ assert not isinstance(
+ m,
+ RoIPool), 'CPU inference with RoIPool is not supported currently.'
+
+ # We don't restore `torch.is_grad_enabled()` value during concurrent
+ # inference since execution can overlap
+ torch.set_grad_enabled(False)
+ results = await model.aforward_test(data, rescale=True)
+ return results
+
+
+def build_test_pipeline(cfg: ConfigType) -> ConfigType:
+ """Build test_pipeline for mot/vis demo. In mot/vis infer, original
+ test_pipeline should remove the "LoadImageFromFile" and
+ "LoadTrackAnnotations".
+
+ Args:
+ cfg (ConfigDict): The loaded config.
+ Returns:
+ ConfigType: new test_pipeline
+ """
+ # remove the "LoadImageFromFile" and "LoadTrackAnnotations" in pipeline
+ transform_broadcaster = cfg.test_dataloader.dataset.pipeline[0].copy()
+ for transform in transform_broadcaster['transforms']:
+ if transform['type'] == 'Resize':
+ transform_broadcaster['transforms'] = transform
+ pack_track_inputs = cfg.test_dataloader.dataset.pipeline[-1].copy()
+ test_pipeline = Compose([transform_broadcaster, pack_track_inputs])
+
+ return test_pipeline
+
+
+def inference_mot(model: nn.Module, img: np.ndarray, frame_id: int,
+ video_len: int) -> SampleList:
+ """Inference image(s) with the mot model.
+
+ Args:
+ model (nn.Module): The loaded mot model.
+ img (np.ndarray): Loaded image.
+ frame_id (int): frame id.
+ video_len (int): demo video length
+ Returns:
+ SampleList: The tracking data samples.
+ """
+ cfg = model.cfg
+ data = dict(
+ img=[img.astype(np.float32)],
+ frame_id=[frame_id],
+ ori_shape=[img.shape[:2]],
+ img_id=[frame_id + 1],
+ ori_video_length=[video_len])
+
+ test_pipeline = build_test_pipeline(cfg)
+ data = test_pipeline(data)
+
+ if not next(model.parameters()).is_cuda:
+ for m in model.modules():
+ assert not isinstance(
+ m, RoIPool
+ ), 'CPU inference with RoIPool is not supported currently.'
+
+ # forward the model
+ with torch.no_grad():
+ data = default_collate([data])
+ result = model.test_step(data)[0]
+ return result
+
+
+def init_track_model(config: Union[str, Config],
+ checkpoint: Optional[str] = None,
+ detector: Optional[str] = None,
+ reid: Optional[str] = None,
+ device: str = 'cuda:0',
+ cfg_options: Optional[dict] = None) -> nn.Module:
+ """Initialize a model from config file.
+
+ Args:
+ config (str or :obj:`mmengine.Config`): Config file path or the config
+ object.
+ checkpoint (Optional[str], optional): Checkpoint path. Defaults to
+ None.
+ detector (Optional[str], optional): Detector Checkpoint path, use in
+ some tracking algorithms like sort. Defaults to None.
+ reid (Optional[str], optional): Reid checkpoint path. use in
+ some tracking algorithms like sort. Defaults to None.
+ device (str, optional): The device that the model inferences on.
+ Defaults to `cuda:0`.
+ cfg_options (Optional[dict], optional): Options to override some
+ settings in the used config. Defaults to None.
+
+ Returns:
+ nn.Module: The constructed model.
+ """
+ if isinstance(config, str):
+ config = Config.fromfile(config)
+ elif not isinstance(config, Config):
+ raise TypeError('config must be a filename or Config object, '
+ f'but got {type(config)}')
+ if cfg_options is not None:
+ config.merge_from_dict(cfg_options)
+
+ model = MODELS.build(config.model)
+
+ if checkpoint is not None:
+ checkpoint = load_checkpoint(model, checkpoint, map_location='cpu')
+ # Weights converted from elsewhere may not have meta fields.
+ checkpoint_meta = checkpoint.get('meta', {})
+ # save the dataset_meta in the model for convenience
+ if 'dataset_meta' in checkpoint_meta:
+ if 'CLASSES' in checkpoint_meta['dataset_meta']:
+ value = checkpoint_meta['dataset_meta'].pop('CLASSES')
+ checkpoint_meta['dataset_meta']['classes'] = value
+ model.dataset_meta = checkpoint_meta['dataset_meta']
+
+ if detector is not None:
+ assert not (checkpoint and detector), \
+ 'Error: checkpoint and detector checkpoint cannot both exist'
+ load_checkpoint(model.detector, detector, map_location='cpu')
+
+ if reid is not None:
+ assert not (checkpoint and reid), \
+ 'Error: checkpoint and reid checkpoint cannot both exist'
+ load_checkpoint(model.reid, reid, map_location='cpu')
+
+ # Some methods don't load checkpoints or checkpoints don't contain
+ # 'dataset_meta'
+ # VIS need dataset_meta, MOT don't need dataset_meta
+ if not hasattr(model, 'dataset_meta'):
+ warnings.warn('dataset_meta or class names are missed, '
+ 'use None by default.')
+ model.dataset_meta = {'classes': None}
+
+ model.cfg = config # save the config in the model for convenience
+ model.to(device)
+ model.eval()
+ return model
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/datasets/coco_detection.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/datasets/coco_detection.py
new file mode 100644
index 0000000000000000000000000000000000000000..45041f6d236be95eb7592035d31f155c61bfcb25
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/datasets/coco_detection.py
@@ -0,0 +1,104 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.transforms import LoadImageFromFile
+from mmengine.dataset.sampler import DefaultSampler
+
+from mmdet.datasets import AspectRatioBatchSampler, CocoDataset
+from mmdet.datasets.transforms import (LoadAnnotations, PackDetInputs,
+ RandomFlip, Resize)
+from mmdet.evaluation import CocoMetric
+
+# dataset settings
+dataset_type = CocoDataset
+data_root = 'data/coco/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=LoadAnnotations, with_bbox=True),
+ dict(type=Resize, scale=(1333, 800), keep_ratio=True),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=PackDetInputs)
+]
+test_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=Resize, scale=(1333, 800), keep_ratio=True),
+ # If you don't have a gt annotation, delete the pipeline
+ dict(type=LoadAnnotations, with_bbox=True),
+ dict(
+ type=PackDetInputs,
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type=DefaultSampler, shuffle=True),
+ batch_sampler=dict(type=AspectRatioBatchSampler),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type=DefaultSampler, shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type=CocoMetric,
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric='bbox',
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+# inference on test dataset and
+# format the output results for submission.
+# test_dataloader = dict(
+# batch_size=1,
+# num_workers=2,
+# persistent_workers=True,
+# drop_last=False,
+# sampler=dict(type=DefaultSampler, shuffle=False),
+# dataset=dict(
+# type=dataset_type,
+# data_root=data_root,
+# ann_file=data_root + 'annotations/image_info_test-dev2017.json',
+# data_prefix=dict(img='test2017/'),
+# test_mode=True,
+# pipeline=test_pipeline))
+# test_evaluator = dict(
+# type=CocoMetric,
+# metric='bbox',
+# format_only=True,
+# ann_file=data_root + 'annotations/image_info_test-dev2017.json',
+# outfile_prefix='./work_dirs/coco_detection/test')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/datasets/coco_instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/datasets/coco_instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..b9575432e26b7e861c4dfcf535773b7a1990eeab
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/datasets/coco_instance.py
@@ -0,0 +1,106 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.transforms.loading import LoadImageFromFile
+from mmengine.dataset.sampler import DefaultSampler
+
+from mmdet.datasets.coco import CocoDataset
+from mmdet.datasets.samplers.batch_sampler import AspectRatioBatchSampler
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import LoadAnnotations
+from mmdet.datasets.transforms.transforms import RandomFlip, Resize
+from mmdet.evaluation.metrics.coco_metric import CocoMetric
+
+# dataset settings
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=LoadAnnotations, with_bbox=True, with_mask=True),
+ dict(type=Resize, scale=(1333, 800), keep_ratio=True),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=PackDetInputs)
+]
+test_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=Resize, scale=(1333, 800), keep_ratio=True),
+ # If you don't have a gt annotation, delete the pipeline
+ dict(type=LoadAnnotations, with_bbox=True, with_mask=True),
+ dict(
+ type=PackDetInputs,
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type=DefaultSampler, shuffle=True),
+ batch_sampler=dict(type=AspectRatioBatchSampler),
+ dataset=dict(
+ type=CocoDataset,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type=DefaultSampler, shuffle=False),
+ dataset=dict(
+ type=CocoDataset,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type=CocoMetric,
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric=['bbox', 'segm'],
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+# inference on test dataset and
+# format the output results for submission.
+# test_dataloader = dict(
+# batch_size=1,
+# num_workers=2,
+# persistent_workers=True,
+# drop_last=False,
+# sampler=dict(type=DefaultSampler, shuffle=False),
+# dataset=dict(
+# type=CocoDataset,
+# data_root=data_root,
+# ann_file=data_root + 'annotations/image_info_test-dev2017.json',
+# data_prefix=dict(img='test2017/'),
+# test_mode=True,
+# pipeline=test_pipeline))
+# test_evaluator = dict(
+# type=CocoMetric,
+# metric=['bbox', 'segm'],
+# format_only=True,
+# ann_file=data_root + 'annotations/image_info_test-dev2017.json',
+# outfile_prefix='./work_dirs/coco_instance/test')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/datasets/coco_instance_semantic.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/datasets/coco_instance_semantic.py
new file mode 100644
index 0000000000000000000000000000000000000000..7cf5b2cfab8a98a6c97e23a8df663e8f1e90b355
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/datasets/coco_instance_semantic.py
@@ -0,0 +1,87 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.transforms.loading import LoadImageFromFile
+from mmengine.dataset.sampler import DefaultSampler
+
+from mmdet.datasets.coco import CocoDataset
+from mmdet.datasets.samplers.batch_sampler import AspectRatioBatchSampler
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import LoadAnnotations
+from mmdet.datasets.transforms.transforms import RandomFlip, Resize
+from mmdet.evaluation.metrics.coco_metric import CocoMetric
+
+# dataset settings
+dataset_type = 'CocoDataset'
+data_root = 'data/coco/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=LoadAnnotations, with_bbox=True, with_mask=True, with_seg=True),
+ dict(type=Resize, scale=(1333, 800), keep_ratio=True),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=PackDetInputs)
+]
+test_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=Resize, scale=(1333, 800), keep_ratio=True),
+ # If you don't have a gt annotation, delete the pipeline
+ dict(type=LoadAnnotations, with_bbox=True, with_mask=True, with_seg=True),
+ dict(
+ type=PackDetInputs,
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type=DefaultSampler, shuffle=True),
+ batch_sampler=dict(type=AspectRatioBatchSampler),
+ dataset=dict(
+ type=CocoDataset,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/', seg='stuffthingmaps/train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args))
+
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type=DefaultSampler, shuffle=False),
+ dataset=dict(
+ type=CocoDataset,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type=CocoMetric,
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric=['bbox', 'segm'],
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/datasets/coco_panoptic.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/datasets/coco_panoptic.py
new file mode 100644
index 0000000000000000000000000000000000000000..29d655ff619c74c5976d5f06c0c623a0d3459997
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/datasets/coco_panoptic.py
@@ -0,0 +1,105 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.transforms.loading import LoadImageFromFile
+from mmengine.dataset.sampler import DefaultSampler
+
+from mmdet.datasets.coco_panoptic import CocoPanopticDataset
+from mmdet.datasets.samplers.batch_sampler import AspectRatioBatchSampler
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import LoadPanopticAnnotations
+from mmdet.datasets.transforms.transforms import RandomFlip, Resize
+from mmdet.evaluation.metrics.coco_panoptic_metric import CocoPanopticMetric
+
+# dataset settings
+dataset_type = 'CocoPanopticDataset'
+data_root = 'data/coco/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=LoadPanopticAnnotations, backend_args=backend_args),
+ dict(type=Resize, scale=(1333, 800), keep_ratio=True),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=PackDetInputs)
+]
+test_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=Resize, scale=(1333, 800), keep_ratio=True),
+ dict(type=LoadPanopticAnnotations, backend_args=backend_args),
+ dict(
+ type=PackDetInputs,
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type=DefaultSampler, shuffle=True),
+ batch_sampler=dict(type=AspectRatioBatchSampler),
+ dataset=dict(
+ type=CocoPanopticDataset,
+ data_root=data_root,
+ ann_file='annotations/panoptic_train2017.json',
+ data_prefix=dict(
+ img='train2017/', seg='annotations/panoptic_train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type=DefaultSampler, shuffle=False),
+ dataset=dict(
+ type=CocoPanopticDataset,
+ data_root=data_root,
+ ann_file='annotations/panoptic_val2017.json',
+ data_prefix=dict(img='val2017/', seg='annotations/panoptic_val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type=CocoPanopticMetric,
+ ann_file=data_root + 'annotations/panoptic_val2017.json',
+ seg_prefix=data_root + 'annotations/panoptic_val2017/',
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+# inference on test dataset and
+# format the output results for submission.
+# test_dataloader = dict(
+# batch_size=1,
+# num_workers=1,
+# persistent_workers=True,
+# drop_last=False,
+# sampler=dict(type=DefaultSampler, shuffle=False),
+# dataset=dict(
+# type=CocoPanopticDataset,
+# data_root=data_root,
+# ann_file='annotations/panoptic_image_info_test-dev2017.json',
+# data_prefix=dict(img='test2017/'),
+# test_mode=True,
+# pipeline=test_pipeline))
+# test_evaluator = dict(
+# type=CocoPanopticMetric,
+# format_only=True,
+# ann_file=data_root + 'annotations/panoptic_image_info_test-dev2017.json',
+# outfile_prefix='./work_dirs/coco_panoptic/test')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/datasets/mot_challenge.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/datasets/mot_challenge.py
new file mode 100644
index 0000000000000000000000000000000000000000..a71520a84e52a812f83862920040d96746829285
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/datasets/mot_challenge.py
@@ -0,0 +1,101 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.transforms import (LoadImageFromFile, RandomResize,
+ TransformBroadcaster)
+
+from mmdet.datasets import MOTChallengeDataset
+from mmdet.datasets.samplers import TrackImgSampler
+from mmdet.datasets.transforms import (LoadTrackAnnotations, PackTrackInputs,
+ PhotoMetricDistortion, RandomCrop,
+ RandomFlip, Resize,
+ UniformRefFrameSample)
+from mmdet.evaluation import MOTChallengeMetric
+
+# dataset settings
+dataset_type = MOTChallengeDataset
+data_root = 'data/MOT17/'
+img_scale = (1088, 1088)
+
+backend_args = None
+# data pipeline
+train_pipeline = [
+ dict(
+ type=UniformRefFrameSample,
+ num_ref_imgs=1,
+ frame_range=10,
+ filter_key_img=True),
+ dict(
+ type=TransformBroadcaster,
+ share_random_params=True,
+ transforms=[
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=LoadTrackAnnotations),
+ dict(
+ type=RandomResize,
+ scale=img_scale,
+ ratio_range=(0.8, 1.2),
+ keep_ratio=True,
+ clip_object_border=False),
+ dict(type=PhotoMetricDistortion)
+ ]),
+ dict(
+ type=TransformBroadcaster,
+ # different cropped positions for different frames
+ share_random_params=False,
+ transforms=[
+ dict(type=RandomCrop, crop_size=img_scale, bbox_clip_border=False)
+ ]),
+ dict(
+ type=TransformBroadcaster,
+ share_random_params=True,
+ transforms=[
+ dict(type=RandomFlip, prob=0.5),
+ ]),
+ dict(type=PackTrackInputs)
+]
+
+test_pipeline = [
+ dict(
+ type=TransformBroadcaster,
+ transforms=[
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=Resize, scale=img_scale, keep_ratio=True),
+ dict(type=LoadTrackAnnotations)
+ ]),
+ dict(type=PackTrackInputs)
+]
+
+# dataloader
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type=TrackImgSampler), # image-based sampling
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ visibility_thr=-1,
+ ann_file='annotations/half-train_cocoformat.json',
+ data_prefix=dict(img_path='train'),
+ metainfo=dict(classes=('pedestrian', )),
+ pipeline=train_pipeline))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ # Now we support two ways to test, image_based and video_based
+ # if you want to use video_based sampling, you can use as follows
+ # sampler=dict(type='DefaultSampler', shuffle=False, round_up=False),
+ sampler=dict(type=TrackImgSampler), # image-based sampling
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/half-val_cocoformat.json',
+ data_prefix=dict(img_path='train'),
+ test_mode=True,
+ pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+# evaluator
+val_evaluator = dict(
+ type=MOTChallengeMetric, metric=['HOTA', 'CLEAR', 'Identity'])
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/default_runtime.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/default_runtime.py
new file mode 100644
index 0000000000000000000000000000000000000000..ff96dbf29f3c90266a268d3831878b0a437d98b2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/default_runtime.py
@@ -0,0 +1,33 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.hooks import (CheckpointHook, DistSamplerSeedHook, IterTimerHook,
+ LoggerHook, ParamSchedulerHook)
+from mmengine.runner import LogProcessor
+from mmengine.visualization import LocalVisBackend
+
+from mmdet.engine.hooks import DetVisualizationHook
+from mmdet.visualization import DetLocalVisualizer
+
+default_scope = None
+
+default_hooks = dict(
+ timer=dict(type=IterTimerHook),
+ logger=dict(type=LoggerHook, interval=50),
+ param_scheduler=dict(type=ParamSchedulerHook),
+ checkpoint=dict(type=CheckpointHook, interval=1),
+ sampler_seed=dict(type=DistSamplerSeedHook),
+ visualization=dict(type=DetVisualizationHook))
+
+env_cfg = dict(
+ cudnn_benchmark=False,
+ mp_cfg=dict(mp_start_method='fork', opencv_num_threads=0),
+ dist_cfg=dict(backend='nccl'),
+)
+
+vis_backends = [dict(type=LocalVisBackend)]
+visualizer = dict(
+ type=DetLocalVisualizer, vis_backends=vis_backends, name='visualizer')
+log_processor = dict(type=LogProcessor, window_size=50, by_epoch=True)
+
+log_level = 'INFO'
+load_from = None
+resume = False
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/models/cascade_mask_rcnn_r50_fpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/models/cascade_mask_rcnn_r50_fpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..b9132ac40330c67e03ebc608f9527c678c72210e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/models/cascade_mask_rcnn_r50_fpn.py
@@ -0,0 +1,220 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.ops import RoIAlign, nms
+from torch.nn import BatchNorm2d
+
+from mmdet.models.backbones.resnet import ResNet
+from mmdet.models.data_preprocessors.data_preprocessor import \
+ DetDataPreprocessor
+from mmdet.models.dense_heads.rpn_head import RPNHead
+from mmdet.models.detectors.cascade_rcnn import CascadeRCNN
+from mmdet.models.losses.cross_entropy_loss import CrossEntropyLoss
+from mmdet.models.losses.smooth_l1_loss import SmoothL1Loss
+from mmdet.models.necks.fpn import FPN
+from mmdet.models.roi_heads.bbox_heads.convfc_bbox_head import \
+ Shared2FCBBoxHead
+from mmdet.models.roi_heads.cascade_roi_head import CascadeRoIHead
+from mmdet.models.roi_heads.mask_heads.fcn_mask_head import FCNMaskHead
+from mmdet.models.roi_heads.roi_extractors.single_level_roi_extractor import \
+ SingleRoIExtractor
+from mmdet.models.task_modules.assigners.max_iou_assigner import MaxIoUAssigner
+from mmdet.models.task_modules.coders.delta_xywh_bbox_coder import \
+ DeltaXYWHBBoxCoder
+from mmdet.models.task_modules.prior_generators.anchor_generator import \
+ AnchorGenerator
+from mmdet.models.task_modules.samplers.random_sampler import RandomSampler
+
+# model settings
+model = dict(
+ type=CascadeRCNN,
+ data_preprocessor=dict(
+ type=DetDataPreprocessor,
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_mask=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type=ResNet,
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type=BatchNorm2d, requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type=FPN,
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5),
+ rpn_head=dict(
+ type=RPNHead,
+ in_channels=256,
+ feat_channels=256,
+ anchor_generator=dict(
+ type=AnchorGenerator,
+ scales=[8],
+ ratios=[0.5, 1.0, 2.0],
+ strides=[4, 8, 16, 32, 64]),
+ bbox_coder=dict(
+ type=DeltaXYWHBBoxCoder,
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loss_cls=dict(
+ type=CrossEntropyLoss, use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type=SmoothL1Loss, beta=1.0 / 9.0, loss_weight=1.0)),
+ roi_head=dict(
+ type=CascadeRoIHead,
+ num_stages=3,
+ stage_loss_weights=[1, 0.5, 0.25],
+ bbox_roi_extractor=dict(
+ type=SingleRoIExtractor,
+ roi_layer=dict(type=RoIAlign, output_size=7, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ bbox_head=[
+ dict(
+ type=Shared2FCBBoxHead,
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type=DeltaXYWHBBoxCoder,
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type=CrossEntropyLoss, use_sigmoid=False, loss_weight=1.0),
+ loss_bbox=dict(type=SmoothL1Loss, beta=1.0, loss_weight=1.0)),
+ dict(
+ type=Shared2FCBBoxHead,
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type=DeltaXYWHBBoxCoder,
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.05, 0.05, 0.1, 0.1]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type=CrossEntropyLoss, use_sigmoid=False, loss_weight=1.0),
+ loss_bbox=dict(type=SmoothL1Loss, beta=1.0, loss_weight=1.0)),
+ dict(
+ type=Shared2FCBBoxHead,
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type=DeltaXYWHBBoxCoder,
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.033, 0.033, 0.067, 0.067]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type=CrossEntropyLoss, use_sigmoid=False, loss_weight=1.0),
+ loss_bbox=dict(type=SmoothL1Loss, beta=1.0, loss_weight=1.0))
+ ],
+ mask_roi_extractor=dict(
+ type=SingleRoIExtractor,
+ roi_layer=dict(type=RoIAlign, output_size=14, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ mask_head=dict(
+ type=FCNMaskHead,
+ num_convs=4,
+ in_channels=256,
+ conv_out_channels=256,
+ num_classes=80,
+ loss_mask=dict(
+ type=CrossEntropyLoss, use_mask=True, loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ assigner=dict(
+ type=MaxIoUAssigner,
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ match_low_quality=True,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type=RandomSampler,
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=0,
+ pos_weight=-1,
+ debug=False),
+ rpn_proposal=dict(
+ nms_pre=2000,
+ max_per_img=2000,
+ nms=dict(type=nms, iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=[
+ dict(
+ assigner=dict(
+ type=MaxIoUAssigner,
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type=RandomSampler,
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ mask_size=28,
+ pos_weight=-1,
+ debug=False),
+ dict(
+ assigner=dict(
+ type=MaxIoUAssigner,
+ pos_iou_thr=0.6,
+ neg_iou_thr=0.6,
+ min_pos_iou=0.6,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type=RandomSampler,
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ mask_size=28,
+ pos_weight=-1,
+ debug=False),
+ dict(
+ assigner=dict(
+ type=MaxIoUAssigner,
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.7,
+ min_pos_iou=0.7,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type=RandomSampler,
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ mask_size=28,
+ pos_weight=-1,
+ debug=False)
+ ]),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=1000,
+ max_per_img=1000,
+ nms=dict(type=nms, iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ score_thr=0.05,
+ nms=dict(type=nms, iou_threshold=0.5),
+ max_per_img=100,
+ mask_thr_binary=0.5)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/models/cascade_rcnn_r50_fpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/models/cascade_rcnn_r50_fpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..8e6654f381f4993a57b81e6ed1f86c0558b56616
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/models/cascade_rcnn_r50_fpn.py
@@ -0,0 +1,201 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.ops import RoIAlign, nms
+from torch.nn import BatchNorm2d
+
+from mmdet.models.backbones.resnet import ResNet
+from mmdet.models.data_preprocessors.data_preprocessor import \
+ DetDataPreprocessor
+from mmdet.models.dense_heads.rpn_head import RPNHead
+from mmdet.models.detectors.cascade_rcnn import CascadeRCNN
+from mmdet.models.losses.cross_entropy_loss import CrossEntropyLoss
+from mmdet.models.losses.smooth_l1_loss import SmoothL1Loss
+from mmdet.models.necks.fpn import FPN
+from mmdet.models.roi_heads.bbox_heads.convfc_bbox_head import \
+ Shared2FCBBoxHead
+from mmdet.models.roi_heads.cascade_roi_head import CascadeRoIHead
+from mmdet.models.roi_heads.roi_extractors.single_level_roi_extractor import \
+ SingleRoIExtractor
+from mmdet.models.task_modules.assigners.max_iou_assigner import MaxIoUAssigner
+from mmdet.models.task_modules.coders.delta_xywh_bbox_coder import \
+ DeltaXYWHBBoxCoder
+from mmdet.models.task_modules.prior_generators.anchor_generator import \
+ AnchorGenerator
+from mmdet.models.task_modules.samplers.random_sampler import RandomSampler
+
+# model settings
+model = dict(
+ type=CascadeRCNN,
+ data_preprocessor=dict(
+ type=DetDataPreprocessor,
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type=ResNet,
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type=BatchNorm2d, requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type=FPN,
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5),
+ rpn_head=dict(
+ type=RPNHead,
+ in_channels=256,
+ feat_channels=256,
+ anchor_generator=dict(
+ type=AnchorGenerator,
+ scales=[8],
+ ratios=[0.5, 1.0, 2.0],
+ strides=[4, 8, 16, 32, 64]),
+ bbox_coder=dict(
+ type=DeltaXYWHBBoxCoder,
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loss_cls=dict(
+ type=CrossEntropyLoss, use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type=SmoothL1Loss, beta=1.0 / 9.0, loss_weight=1.0)),
+ roi_head=dict(
+ type=CascadeRoIHead,
+ num_stages=3,
+ stage_loss_weights=[1, 0.5, 0.25],
+ bbox_roi_extractor=dict(
+ type=SingleRoIExtractor,
+ roi_layer=dict(type=RoIAlign, output_size=7, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ bbox_head=[
+ dict(
+ type=Shared2FCBBoxHead,
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type=DeltaXYWHBBoxCoder,
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type=CrossEntropyLoss, use_sigmoid=False, loss_weight=1.0),
+ loss_bbox=dict(type=SmoothL1Loss, beta=1.0, loss_weight=1.0)),
+ dict(
+ type=Shared2FCBBoxHead,
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type=DeltaXYWHBBoxCoder,
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.05, 0.05, 0.1, 0.1]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type=CrossEntropyLoss, use_sigmoid=False, loss_weight=1.0),
+ loss_bbox=dict(type=SmoothL1Loss, beta=1.0, loss_weight=1.0)),
+ dict(
+ type=Shared2FCBBoxHead,
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type=DeltaXYWHBBoxCoder,
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.033, 0.033, 0.067, 0.067]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type=CrossEntropyLoss, use_sigmoid=False, loss_weight=1.0),
+ loss_bbox=dict(type=SmoothL1Loss, beta=1.0, loss_weight=1.0))
+ ]),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ assigner=dict(
+ type=MaxIoUAssigner,
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ match_low_quality=True,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type=RandomSampler,
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=0,
+ pos_weight=-1,
+ debug=False),
+ rpn_proposal=dict(
+ nms_pre=2000,
+ max_per_img=2000,
+ nms=dict(type=nms, iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=[
+ dict(
+ assigner=dict(
+ type=MaxIoUAssigner,
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type=RandomSampler,
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ pos_weight=-1,
+ debug=False),
+ dict(
+ assigner=dict(
+ type=MaxIoUAssigner,
+ pos_iou_thr=0.6,
+ neg_iou_thr=0.6,
+ min_pos_iou=0.6,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type=RandomSampler,
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ pos_weight=-1,
+ debug=False),
+ dict(
+ assigner=dict(
+ type=MaxIoUAssigner,
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.7,
+ min_pos_iou=0.7,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type=RandomSampler,
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ pos_weight=-1,
+ debug=False)
+ ]),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=1000,
+ max_per_img=1000,
+ nms=dict(type=nms, iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ score_thr=0.05,
+ nms=dict(type=nms, iou_threshold=0.5),
+ max_per_img=100)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/models/faster_rcnn_r50_fpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/models/faster_rcnn_r50_fpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..7e18de2224d5b4d2cd16a930daf3a9b360455b36
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/models/faster_rcnn_r50_fpn.py
@@ -0,0 +1,138 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.ops import RoIAlign, nms
+from torch.nn import BatchNorm2d
+
+from mmdet.models.backbones.resnet import ResNet
+from mmdet.models.data_preprocessors.data_preprocessor import \
+ DetDataPreprocessor
+from mmdet.models.dense_heads.rpn_head import RPNHead
+from mmdet.models.detectors.faster_rcnn import FasterRCNN
+from mmdet.models.losses.cross_entropy_loss import CrossEntropyLoss
+from mmdet.models.losses.smooth_l1_loss import L1Loss
+from mmdet.models.necks.fpn import FPN
+from mmdet.models.roi_heads.bbox_heads.convfc_bbox_head import \
+ Shared2FCBBoxHead
+from mmdet.models.roi_heads.roi_extractors.single_level_roi_extractor import \
+ SingleRoIExtractor
+from mmdet.models.roi_heads.standard_roi_head import StandardRoIHead
+from mmdet.models.task_modules.assigners.max_iou_assigner import MaxIoUAssigner
+from mmdet.models.task_modules.coders.delta_xywh_bbox_coder import \
+ DeltaXYWHBBoxCoder
+from mmdet.models.task_modules.prior_generators.anchor_generator import \
+ AnchorGenerator
+from mmdet.models.task_modules.samplers.random_sampler import RandomSampler
+
+# model settings
+model = dict(
+ type=FasterRCNN,
+ data_preprocessor=dict(
+ type=DetDataPreprocessor,
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type=ResNet,
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type=BatchNorm2d, requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type=FPN,
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5),
+ rpn_head=dict(
+ type=RPNHead,
+ in_channels=256,
+ feat_channels=256,
+ anchor_generator=dict(
+ type=AnchorGenerator,
+ scales=[8],
+ ratios=[0.5, 1.0, 2.0],
+ strides=[4, 8, 16, 32, 64]),
+ bbox_coder=dict(
+ type=DeltaXYWHBBoxCoder,
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loss_cls=dict(
+ type=CrossEntropyLoss, use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type=L1Loss, loss_weight=1.0)),
+ roi_head=dict(
+ type=StandardRoIHead,
+ bbox_roi_extractor=dict(
+ type=SingleRoIExtractor,
+ roi_layer=dict(type=RoIAlign, output_size=7, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ bbox_head=dict(
+ type=Shared2FCBBoxHead,
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type=DeltaXYWHBBoxCoder,
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=False,
+ loss_cls=dict(
+ type=CrossEntropyLoss, use_sigmoid=False, loss_weight=1.0),
+ loss_bbox=dict(type=L1Loss, loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ assigner=dict(
+ type=MaxIoUAssigner,
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ match_low_quality=True,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type=RandomSampler,
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ rpn_proposal=dict(
+ nms_pre=2000,
+ max_per_img=1000,
+ nms=dict(type=nms, iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ assigner=dict(
+ type=MaxIoUAssigner,
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type=RandomSampler,
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ pos_weight=-1,
+ debug=False)),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=1000,
+ max_per_img=1000,
+ nms=dict(type=nms, iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ score_thr=0.05,
+ nms=dict(type=nms, iou_threshold=0.5),
+ max_per_img=100)
+ # soft-nms is also supported for rcnn testing
+ # e.g., nms=dict(type='soft_nms', iou_threshold=0.5, min_score=0.05)
+ ))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/models/mask_rcnn_r50_caffe_c4.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/models/mask_rcnn_r50_caffe_c4.py
new file mode 100644
index 0000000000000000000000000000000000000000..3054818375f708826ee41901650a11bbbe3afca9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/models/mask_rcnn_r50_caffe_c4.py
@@ -0,0 +1,158 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.ops import RoIAlign, nms
+from mmengine.model.weight_init import PretrainedInit
+from torch.nn import BatchNorm2d
+
+from mmdet.models.backbones.resnet import ResNet
+from mmdet.models.data_preprocessors.data_preprocessor import \
+ DetDataPreprocessor
+from mmdet.models.dense_heads.rpn_head import RPNHead
+from mmdet.models.detectors.mask_rcnn import MaskRCNN
+from mmdet.models.layers import ResLayer
+from mmdet.models.losses.cross_entropy_loss import CrossEntropyLoss
+from mmdet.models.losses.smooth_l1_loss import L1Loss
+from mmdet.models.roi_heads.bbox_heads.bbox_head import BBoxHead
+from mmdet.models.roi_heads.mask_heads.fcn_mask_head import FCNMaskHead
+from mmdet.models.roi_heads.roi_extractors.single_level_roi_extractor import \
+ SingleRoIExtractor
+from mmdet.models.roi_heads.standard_roi_head import StandardRoIHead
+from mmdet.models.task_modules.assigners.max_iou_assigner import MaxIoUAssigner
+from mmdet.models.task_modules.coders.delta_xywh_bbox_coder import \
+ DeltaXYWHBBoxCoder
+from mmdet.models.task_modules.prior_generators.anchor_generator import \
+ AnchorGenerator
+from mmdet.models.task_modules.samplers.random_sampler import RandomSampler
+
+# model settings
+norm_cfg = dict(type=BatchNorm2d, requires_grad=False)
+# model settings
+model = dict(
+ type=MaskRCNN,
+ data_preprocessor=dict(
+ type=DetDataPreprocessor,
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_mask=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type=ResNet,
+ depth=50,
+ num_stages=3,
+ strides=(1, 2, 2),
+ dilations=(1, 1, 1),
+ out_indices=(2, ),
+ frozen_stages=1,
+ norm_cfg=dict(type=BatchNorm2d, requires_grad=True),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type=PretrainedInit,
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')),
+ rpn_head=dict(
+ type=RPNHead,
+ in_channels=1024,
+ feat_channels=1024,
+ anchor_generator=dict(
+ type=AnchorGenerator,
+ scales=[2, 4, 8, 16, 32],
+ ratios=[0.5, 1.0, 2.0],
+ strides=[16]),
+ bbox_coder=dict(
+ type=DeltaXYWHBBoxCoder,
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loss_cls=dict(
+ type=CrossEntropyLoss, use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type=L1Loss, loss_weight=1.0)),
+ roi_head=dict(
+ type=StandardRoIHead,
+ shared_head=dict(
+ type=ResLayer,
+ depth=50,
+ stage=3,
+ stride=2,
+ dilation=1,
+ style='caffe',
+ norm_cfg=norm_cfg,
+ norm_eval=True),
+ bbox_roi_extractor=dict(
+ type=SingleRoIExtractor,
+ roi_layer=dict(type=RoIAlign, output_size=14, sampling_ratio=0),
+ out_channels=1024,
+ featmap_strides=[16]),
+ bbox_head=dict(
+ type=BBoxHead,
+ with_avg_pool=True,
+ roi_feat_size=7,
+ in_channels=2048,
+ num_classes=80,
+ bbox_coder=dict(
+ type=DeltaXYWHBBoxCoder,
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=False,
+ loss_cls=dict(
+ type=CrossEntropyLoss, use_sigmoid=False, loss_weight=1.0),
+ loss_bbox=dict(type=L1Loss, loss_weight=1.0)),
+ mask_roi_extractor=None,
+ mask_head=dict(
+ type=FCNMaskHead,
+ num_convs=0,
+ in_channels=2048,
+ conv_out_channels=256,
+ num_classes=80,
+ loss_mask=dict(
+ type=CrossEntropyLoss, use_mask=True, loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ assigner=dict(
+ type=MaxIoUAssigner,
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ match_low_quality=True,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type=RandomSampler,
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=0,
+ pos_weight=-1,
+ debug=False),
+ rpn_proposal=dict(
+ nms_pre=12000,
+ max_per_img=2000,
+ nms=dict(type=nms, iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ assigner=dict(
+ type=MaxIoUAssigner,
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type=RandomSampler,
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ mask_size=14,
+ pos_weight=-1,
+ debug=False)),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=6000,
+ max_per_img=1000,
+ nms=dict(type=nms, iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ score_thr=0.05,
+ nms=dict(type=nms, iou_threshold=0.5),
+ max_per_img=100,
+ mask_thr_binary=0.5)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/models/mask_rcnn_r50_fpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/models/mask_rcnn_r50_fpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..c8a0b031da51c8147c8ed5c5f29502bd0c4bbe7f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/models/mask_rcnn_r50_fpn.py
@@ -0,0 +1,154 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.ops import RoIAlign, nms
+from mmengine.model.weight_init import PretrainedInit
+from torch.nn import BatchNorm2d
+
+from mmdet.models.backbones.resnet import ResNet
+from mmdet.models.data_preprocessors.data_preprocessor import \
+ DetDataPreprocessor
+from mmdet.models.dense_heads.rpn_head import RPNHead
+from mmdet.models.detectors.mask_rcnn import MaskRCNN
+from mmdet.models.losses.cross_entropy_loss import CrossEntropyLoss
+from mmdet.models.losses.smooth_l1_loss import L1Loss
+from mmdet.models.necks.fpn import FPN
+from mmdet.models.roi_heads.bbox_heads.convfc_bbox_head import \
+ Shared2FCBBoxHead
+from mmdet.models.roi_heads.mask_heads.fcn_mask_head import FCNMaskHead
+from mmdet.models.roi_heads.roi_extractors.single_level_roi_extractor import \
+ SingleRoIExtractor
+from mmdet.models.roi_heads.standard_roi_head import StandardRoIHead
+from mmdet.models.task_modules.assigners.max_iou_assigner import MaxIoUAssigner
+from mmdet.models.task_modules.coders.delta_xywh_bbox_coder import \
+ DeltaXYWHBBoxCoder
+from mmdet.models.task_modules.prior_generators.anchor_generator import \
+ AnchorGenerator
+from mmdet.models.task_modules.samplers.random_sampler import RandomSampler
+
+# model settings
+model = dict(
+ type=MaskRCNN,
+ data_preprocessor=dict(
+ type=DetDataPreprocessor,
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_mask=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type=ResNet,
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type=BatchNorm2d, requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type=PretrainedInit, checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type=FPN,
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ num_outs=5),
+ rpn_head=dict(
+ type=RPNHead,
+ in_channels=256,
+ feat_channels=256,
+ anchor_generator=dict(
+ type=AnchorGenerator,
+ scales=[8],
+ ratios=[0.5, 1.0, 2.0],
+ strides=[4, 8, 16, 32, 64]),
+ bbox_coder=dict(
+ type=DeltaXYWHBBoxCoder,
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loss_cls=dict(
+ type=CrossEntropyLoss, use_sigmoid=True, loss_weight=1.0),
+ loss_bbox=dict(type=L1Loss, loss_weight=1.0)),
+ roi_head=dict(
+ type=StandardRoIHead,
+ bbox_roi_extractor=dict(
+ type=SingleRoIExtractor,
+ roi_layer=dict(type=RoIAlign, output_size=7, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ bbox_head=dict(
+ type=Shared2FCBBoxHead,
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=80,
+ bbox_coder=dict(
+ type=DeltaXYWHBBoxCoder,
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=False,
+ loss_cls=dict(
+ type=CrossEntropyLoss, use_sigmoid=False, loss_weight=1.0),
+ loss_bbox=dict(type=L1Loss, loss_weight=1.0)),
+ mask_roi_extractor=dict(
+ type=SingleRoIExtractor,
+ roi_layer=dict(type=RoIAlign, output_size=14, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ mask_head=dict(
+ type=FCNMaskHead,
+ num_convs=4,
+ in_channels=256,
+ conv_out_channels=256,
+ num_classes=80,
+ loss_mask=dict(
+ type=CrossEntropyLoss, use_mask=True, loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ assigner=dict(
+ type=MaxIoUAssigner,
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ match_low_quality=True,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type=RandomSampler,
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ rpn_proposal=dict(
+ nms_pre=2000,
+ max_per_img=1000,
+ nms=dict(type=nms, iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ assigner=dict(
+ type=MaxIoUAssigner,
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ match_low_quality=True,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type=RandomSampler,
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ mask_size=28,
+ pos_weight=-1,
+ debug=False)),
+ test_cfg=dict(
+ rpn=dict(
+ nms_pre=1000,
+ max_per_img=1000,
+ nms=dict(type=nms, iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ score_thr=0.05,
+ nms=dict(type=nms, iou_threshold=0.5),
+ max_per_img=100,
+ mask_thr_binary=0.5)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/models/retinanet_r50_fpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/models/retinanet_r50_fpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..33e5cc4f1fe69f66801abdfedc578293e96cd23d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/models/retinanet_r50_fpn.py
@@ -0,0 +1,77 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.ops import nms
+from torch.nn import BatchNorm2d
+
+from mmdet.models import (FPN, DetDataPreprocessor, FocalLoss, L1Loss, ResNet,
+ RetinaHead, RetinaNet)
+from mmdet.models.task_modules import (AnchorGenerator, DeltaXYWHBBoxCoder,
+ MaxIoUAssigner, PseudoSampler)
+
+# model settings
+model = dict(
+ type=RetinaNet,
+ data_preprocessor=dict(
+ type=DetDataPreprocessor,
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type=ResNet,
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type=BatchNorm2d, requires_grad=True),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type=FPN,
+ in_channels=[256, 512, 1024, 2048],
+ out_channels=256,
+ start_level=1,
+ add_extra_convs='on_input',
+ num_outs=5),
+ bbox_head=dict(
+ type=RetinaHead,
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ anchor_generator=dict(
+ type=AnchorGenerator,
+ octave_base_scale=4,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[8, 16, 32, 64, 128]),
+ bbox_coder=dict(
+ type=DeltaXYWHBBoxCoder,
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loss_cls=dict(
+ type=FocalLoss,
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox=dict(type=L1Loss, loss_weight=1.0)),
+ # model training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type=MaxIoUAssigner,
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.4,
+ min_pos_iou=0,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type=PseudoSampler), # Focal loss should use PseudoSampler
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type=nms, iou_threshold=0.5),
+ max_per_img=100))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/schedules/schedule_1x.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/schedules/schedule_1x.py
new file mode 100644
index 0000000000000000000000000000000000000000..47d1fa6a4852c40f3f9962a47ec90e365671c61c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/schedules/schedule_1x.py
@@ -0,0 +1,33 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.optim.optimizer.optimizer_wrapper import OptimWrapper
+from mmengine.optim.scheduler.lr_scheduler import LinearLR, MultiStepLR
+from mmengine.runner.loops import EpochBasedTrainLoop, TestLoop, ValLoop
+from torch.optim.sgd import SGD
+
+# training schedule for 1x
+train_cfg = dict(type=EpochBasedTrainLoop, max_epochs=12, val_interval=1)
+val_cfg = dict(type=ValLoop)
+test_cfg = dict(type=TestLoop)
+
+# learning rate
+param_scheduler = [
+ dict(type=LinearLR, start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=12,
+ by_epoch=True,
+ milestones=[8, 11],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type=OptimWrapper,
+ optimizer=dict(type=SGD, lr=0.02, momentum=0.9, weight_decay=0.0001))
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/schedules/schedule_2x.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/schedules/schedule_2x.py
new file mode 100644
index 0000000000000000000000000000000000000000..51ba09a4723bc6ba41b8b4cb6e623ade7db26511
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/_base_/schedules/schedule_2x.py
@@ -0,0 +1,33 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.optim.optimizer.optimizer_wrapper import OptimWrapper
+from mmengine.optim.scheduler.lr_scheduler import LinearLR, MultiStepLR
+from mmengine.runner.loops import EpochBasedTrainLoop, TestLoop, ValLoop
+from torch.optim.sgd import SGD
+
+# training schedule for 1x
+train_cfg = dict(type=EpochBasedTrainLoop, max_epochs=24, val_interval=1)
+val_cfg = dict(type=ValLoop)
+test_cfg = dict(type=TestLoop)
+
+# learning rate
+param_scheduler = [
+ dict(type=LinearLR, start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=24,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type=OptimWrapper,
+ optimizer=dict(type=SGD, lr=0.02, momentum=0.9, weight_decay=0.0001))
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/cascade_rcnn/cascade_mask_rcnn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/cascade_rcnn/cascade_mask_rcnn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..a81c25af8b9506acfa8755ff4ec99d33c661442b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/cascade_rcnn/cascade_mask_rcnn_r50_fpn_1x_coco.py
@@ -0,0 +1,13 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.datasets.coco_instance import *
+ from .._base_.default_runtime import *
+ from .._base_.models.cascade_mask_rcnn_r50_fpn import *
+ from .._base_.schedules.schedule_1x import *
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/cascade_rcnn/cascade_rcnn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/cascade_rcnn/cascade_rcnn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..883f09be67066283e1b59484d3483e73d82af776
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/cascade_rcnn/cascade_rcnn_r50_fpn_1x_coco.py
@@ -0,0 +1,13 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.datasets.coco_detection import *
+ from .._base_.default_runtime import *
+ from .._base_.models.cascade_rcnn_r50_fpn import *
+ from .._base_.schedules.schedule_1x import *
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/lsj_100e_coco_detection.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/lsj_100e_coco_detection.py
new file mode 100644
index 0000000000000000000000000000000000000000..ea2d6bad7f500417ad1eb3e16ca7761c6cadca0e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/lsj_100e_coco_detection.py
@@ -0,0 +1,134 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.default_runtime import *
+
+from mmengine.dataset.sampler import DefaultSampler
+from mmengine.optim import OptimWrapper
+from mmengine.optim.scheduler.lr_scheduler import LinearLR, MultiStepLR
+from mmengine.runner.loops import EpochBasedTrainLoop, TestLoop, ValLoop
+from torch.optim import SGD
+
+from mmdet.datasets import CocoDataset, RepeatDataset
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import (FilterAnnotations,
+ LoadAnnotations,
+ LoadImageFromFile)
+from mmdet.datasets.transforms.transforms import (CachedMixUp, CachedMosaic,
+ Pad, RandomCrop, RandomFlip,
+ RandomResize, Resize)
+from mmdet.evaluation import CocoMetric
+
+# dataset settings
+dataset_type = CocoDataset
+data_root = 'data/coco/'
+image_size = (1024, 1024)
+
+backend_args = None
+
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=LoadAnnotations, with_bbox=True, with_mask=True),
+ dict(
+ type=RandomResize,
+ scale=image_size,
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(
+ type=RandomCrop,
+ crop_type='absolute_range',
+ crop_size=image_size,
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type=FilterAnnotations, min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=PackDetInputs)
+]
+test_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=Resize, scale=(1333, 800), keep_ratio=True),
+ dict(type=LoadAnnotations, with_bbox=True),
+ dict(
+ type=PackDetInputs,
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+# Use RepeatDataset to speed up training
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type=DefaultSampler, shuffle=True),
+ dataset=dict(
+ type=RepeatDataset,
+ times=4, # simply change this from 2 to 16 for 50e - 400e training.
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args)))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type=DefaultSampler, shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type=CocoMetric,
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric=['bbox', 'segm'],
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+max_epochs = 25
+
+train_cfg = dict(
+ type=EpochBasedTrainLoop, max_epochs=max_epochs, val_interval=5)
+val_cfg = dict(type=ValLoop)
+test_cfg = dict(type=TestLoop)
+
+# optimizer assumes bs=64
+optim_wrapper = dict(
+ type=OptimWrapper,
+ optimizer=dict(type=SGD, lr=0.1, momentum=0.9, weight_decay=0.00004))
+
+# learning rate
+param_scheduler = [
+ dict(type=LinearLR, start_factor=0.067, by_epoch=False, begin=0, end=500),
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[22, 24],
+ gamma=0.1)
+]
+
+# only keep latest 2 checkpoints
+default_hooks.update(dict(checkpoint=dict(max_keep_ckpts=2)))
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (32 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/lsj_100e_coco_instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/lsj_100e_coco_instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..90104ee503b22ef395a9b87d74ee80431575d90c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/lsj_100e_coco_instance.py
@@ -0,0 +1,134 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.default_runtime import *
+
+from mmengine.dataset.sampler import DefaultSampler
+from mmengine.optim import OptimWrapper
+from mmengine.optim.scheduler.lr_scheduler import LinearLR, MultiStepLR
+from mmengine.runner.loops import EpochBasedTrainLoop, TestLoop, ValLoop
+from torch.optim import SGD
+
+from mmdet.datasets import CocoDataset, RepeatDataset
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import (FilterAnnotations,
+ LoadAnnotations,
+ LoadImageFromFile)
+from mmdet.datasets.transforms.transforms import (CachedMixUp, CachedMosaic,
+ Pad, RandomCrop, RandomFlip,
+ RandomResize, Resize)
+from mmdet.evaluation import CocoMetric
+
+# dataset settings
+dataset_type = CocoDataset
+data_root = 'data/coco/'
+image_size = (1024, 1024)
+
+backend_args = None
+
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=LoadAnnotations, with_bbox=True, with_mask=True),
+ dict(
+ type=RandomResize,
+ scale=image_size,
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(
+ type=RandomCrop,
+ crop_type='absolute_range',
+ crop_size=image_size,
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type=FilterAnnotations, min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=PackDetInputs)
+]
+test_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=Resize, scale=(1333, 800), keep_ratio=True),
+ dict(type=LoadAnnotations, with_bbox=True, with_mask=True),
+ dict(
+ type=PackDetInputs,
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+# Use RepeatDataset to speed up training
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type=DefaultSampler, shuffle=True),
+ dataset=dict(
+ type=RepeatDataset,
+ times=4, # simply change this from 2 to 16 for 50e - 400e training.
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args)))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type=DefaultSampler, shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type=CocoMetric,
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric=['bbox', 'segm'],
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+max_epochs = 25
+
+train_cfg = dict(
+ type=EpochBasedTrainLoop, max_epochs=max_epochs, val_interval=5)
+val_cfg = dict(type=ValLoop)
+test_cfg = dict(type=TestLoop)
+
+# optimizer assumes bs=64
+optim_wrapper = dict(
+ type=OptimWrapper,
+ optimizer=dict(type=SGD, lr=0.1, momentum=0.9, weight_decay=0.00004))
+
+# learning rate
+param_scheduler = [
+ dict(type=LinearLR, start_factor=0.067, by_epoch=False, begin=0, end=500),
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[22, 24],
+ gamma=0.1)
+]
+
+# only keep latest 2 checkpoints
+default_hooks.update(dict(checkpoint=dict(max_keep_ckpts=2)))
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (32 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/lsj_200e_coco_detection.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/lsj_200e_coco_detection.py
new file mode 100644
index 0000000000000000000000000000000000000000..5759499e95dde6ef99246ab00c21264192ff511c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/lsj_200e_coco_detection.py
@@ -0,0 +1,25 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .lsj_100e_coco_detection import *
+
+# 8x25=200e
+train_dataloader.update(dict(dataset=dict(times=8)))
+
+# learning rate
+param_scheduler = [
+ dict(type=LinearLR, start_factor=0.067, by_epoch=False, begin=0, end=1000),
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=25,
+ by_epoch=True,
+ milestones=[22, 24],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/lsj_200e_coco_instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/lsj_200e_coco_instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..77c5cdd44c488a763d320768e80b314f999ac555
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/lsj_200e_coco_instance.py
@@ -0,0 +1,25 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .lsj_100e_coco_instance import *
+
+# 8x25=200e
+train_dataloader.update(dict(dataset=dict(times=8)))
+
+# learning rate
+param_scheduler = [
+ dict(type=LinearLR, start_factor=0.067, by_epoch=False, begin=0, end=1000),
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=25,
+ by_epoch=True,
+ milestones=[22, 24],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ms_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ms_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c32b24d96aeed59a7340cd7e743dd16b7c728bf1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ms_3x_coco.py
@@ -0,0 +1,130 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.default_runtime import *
+
+from mmcv.transforms import RandomResize
+from mmengine.dataset import RepeatDataset
+from mmengine.dataset.sampler import DefaultSampler
+from mmengine.optim import OptimWrapper
+from mmengine.optim.scheduler.lr_scheduler import LinearLR, MultiStepLR
+from mmengine.runner.loops import EpochBasedTrainLoop, TestLoop, ValLoop
+from torch.optim import SGD
+
+from mmdet.datasets import AspectRatioBatchSampler, CocoDataset
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import (LoadAnnotations,
+ LoadImageFromFile)
+from mmdet.datasets.transforms.transforms import RandomFlip, Resize
+from mmdet.evaluation import CocoMetric
+
+# dataset settings
+dataset_type = CocoDataset
+data_root = 'data/coco/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+# In mstrain 3x config, img_scale=[(1333, 640), (1333, 800)],
+# multiscale_mode='range'
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=LoadAnnotations, with_bbox=True),
+ dict(type=RandomResize, scale=[(1333, 640), (1333, 800)], keep_ratio=True),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=PackDetInputs)
+]
+test_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=Resize, scale=(1333, 800), keep_ratio=True),
+ dict(type=LoadAnnotations, with_bbox=True),
+ dict(
+ type=PackDetInputs,
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader = dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ pin_memory=True,
+ sampler=dict(type=DefaultSampler, shuffle=True),
+ batch_sampler=dict(type=AspectRatioBatchSampler),
+ dataset=dict(
+ type=RepeatDataset,
+ times=3,
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args)))
+val_dataloader = dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type=DefaultSampler, shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args))
+test_dataloader = val_dataloader
+
+val_evaluator = dict(
+ type=CocoMetric,
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric='bbox',
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+# training schedule for 3x with `RepeatDataset`
+train_cfg = dict(type=EpochBasedTrainLoop, max_iters=12, val_interval=1)
+val_cfg = dict(type=ValLoop)
+test_cfg = dict(type=TestLoop)
+
+# learning rate
+param_scheduler = [
+ dict(type=LinearLR, start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=12,
+ by_epoch=False,
+ milestones=[9, 11],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper = dict(
+ type=OptimWrapper,
+ optimizer=dict(type=SGD, lr=0.02, momentum=0.9, weight_decay=0.0001))
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ms_3x_coco_instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ms_3x_coco_instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..3c78909df80173eb37ff83c4ba12614e73848f29
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ms_3x_coco_instance.py
@@ -0,0 +1,136 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.default_runtime import *
+
+from mmcv.transforms import RandomChoiceResize
+from mmengine.dataset import RepeatDataset
+from mmengine.dataset.sampler import DefaultSampler, InfiniteSampler
+from mmengine.optim import OptimWrapper
+from mmengine.optim.scheduler.lr_scheduler import LinearLR, MultiStepLR
+from mmengine.runner.loops import IterBasedTrainLoop, TestLoop, ValLoop
+from torch.optim import SGD
+
+from mmdet.datasets import AspectRatioBatchSampler, CocoDataset
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import (FilterAnnotations,
+ LoadAnnotations,
+ LoadImageFromFile)
+from mmdet.datasets.transforms.transforms import (CachedMixUp, CachedMosaic,
+ Pad, RandomCrop, RandomFlip,
+ RandomResize, Resize)
+from mmdet.evaluation import CocoMetric
+
+# dataset settings
+dataset_type = CocoDataset
+data_root = 'data/coco/'
+
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=LoadAnnotations, with_bbox=True, with_mask=True),
+ dict(
+ type='RandomResize', scale=[(1333, 640), (1333, 800)],
+ keep_ratio=True),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=PackDetInputs)
+]
+test_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=Resize, scale=(1333, 800), keep_ratio=True),
+ dict(type=LoadAnnotations, with_bbox=True, with_mask=True),
+ dict(
+ type=PackDetInputs,
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader.update(
+ dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type=DefaultSampler, shuffle=True),
+ batch_sampler=dict(type=AspectRatioBatchSampler),
+ dataset=dict(
+ type=RepeatDataset,
+ times=3,
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args))))
+val_dataloader.update(
+ dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type=DefaultSampler, shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args)))
+test_dataloader = val_dataloader
+
+val_evaluator.update(
+ dict(
+ type=CocoMetric,
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric='bbox',
+ backend_args=backend_args))
+test_evaluator = val_evaluator
+
+# training schedule for 3x with `RepeatDataset`
+train_cfg.update(dict(type=EpochBasedTrainLoop, max_epochs=12, val_interval=1))
+val_cfg.update(dict(type=ValLoop))
+test_cfg.update(dict(type=TestLoop))
+
+# learning rate
+param_scheduler = [
+ dict(type=LinearLR, start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=12,
+ by_epoch=False,
+ milestones=[9, 11],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper.update(
+ dict(
+ type=OptimWrapper,
+ optimizer=dict(type=SGD, lr=0.02, momentum=0.9, weight_decay=0.0001)))
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr.update(dict(enable=False, base_batch_size=16))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ms_90k_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ms_90k_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3abf1d4a4a8cf53a4abfa43722e306ac04770e18
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ms_90k_coco.py
@@ -0,0 +1,151 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.default_runtime import *
+
+from mmcv.transforms import RandomChoiceResize
+from mmengine.dataset import RepeatDataset
+from mmengine.dataset.sampler import DefaultSampler, InfiniteSampler
+from mmengine.optim import OptimWrapper
+from mmengine.optim.scheduler.lr_scheduler import LinearLR, MultiStepLR
+from mmengine.runner.loops import IterBasedTrainLoop, TestLoop, ValLoop
+from torch.optim import SGD
+
+from mmdet.datasets import AspectRatioBatchSampler, CocoDataset
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import (FilterAnnotations,
+ LoadAnnotations,
+ LoadImageFromFile)
+from mmdet.datasets.transforms.transforms import (CachedMixUp, CachedMosaic,
+ Pad, RandomCrop, RandomFlip,
+ RandomResize, Resize)
+from mmdet.evaluation import CocoMetric
+
+# dataset settings
+dataset_type = CocoDataset
+data_root = 'data/coco/'
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+# Align with Detectron2
+backend = 'pillow'
+train_pipeline = [
+ dict(
+ type=LoadImageFromFile,
+ backend_args=backend_args,
+ imdecode_backend=backend),
+ dict(type=LoadAnnotations, with_bbox=True),
+ dict(
+ type=RandomChoiceResize,
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True,
+ backend=backend),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=PackDetInputs)
+]
+test_pipeline = [
+ dict(
+ type=LoadImageFromFile,
+ backend_args=backend_args,
+ imdecode_backend=backend),
+ dict(type=Resize, scale=(1333, 800), keep_ratio=True, backend=backend),
+ dict(type=LoadAnnotations, with_bbox=True),
+ dict(
+ type=PackDetInputs,
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader.update(
+ dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ pin_memory=True,
+ sampler=dict(type=InfiniteSampler, shuffle=True),
+ batch_sampler=dict(type=AspectRatioBatchSampler),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args)))
+val_dataloader.update(
+ dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ pin_memory=True,
+ sampler=dict(type=DefaultSampler, shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args)))
+test_dataloader = val_dataloader
+
+val_evaluator.update(
+ dict(
+ type=CocoMetric,
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric='bbox',
+ format_only=False,
+ backend_args=backend_args))
+test_evaluator = val_evaluator
+
+# training schedule for 90k
+max_iter = 90000
+train_cfg.update(
+ dict(type=IterBasedTrainLoop, max_iters=max_iter, val_interval=10000))
+val_cfg.update(dict(type=ValLoop))
+test_cfg.update(dict(type=TestLoop))
+
+# learning rate
+param_scheduler = [
+ dict(type=LinearLR, start_factor=0.001, by_epoch=False, begin=0, end=1000),
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=max_iter,
+ by_epoch=False,
+ milestones=[60000, 80000],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper.update(
+ dict(
+ type=OptimWrapper,
+ optimizer=dict(type=SGD, lr=0.02, momentum=0.9, weight_decay=0.0001)))
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr.update(dict(enable=False, base_batch_size=16))
+
+default_hooks.update(dict(checkpoint=dict(by_epoch=False, interval=10000)))
+log_processor.update(dict(by_epoch=False))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ms_poly_3x_coco_instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ms_poly_3x_coco_instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..53913a059a4db9230ebd777934cc8db5595479fe
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ms_poly_3x_coco_instance.py
@@ -0,0 +1,138 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.default_runtime import *
+
+from mmcv.transforms import RandomChoiceResize
+from mmengine.dataset import RepeatDataset
+from mmengine.dataset.sampler import DefaultSampler, InfiniteSampler
+from mmengine.optim import OptimWrapper
+from mmengine.optim.scheduler.lr_scheduler import LinearLR, MultiStepLR
+from mmengine.runner.loops import IterBasedTrainLoop, TestLoop, ValLoop
+from torch.optim import SGD
+
+from mmdet.datasets import AspectRatioBatchSampler, CocoDataset
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import (FilterAnnotations,
+ LoadAnnotations,
+ LoadImageFromFile)
+from mmdet.datasets.transforms.transforms import (CachedMixUp, CachedMosaic,
+ Pad, RandomCrop, RandomFlip,
+ RandomResize, Resize)
+from mmdet.evaluation import CocoMetric
+
+# dataset settings
+dataset_type = CocoDataset
+data_root = 'data/coco/'
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+# In mstrain 3x config, img_scale=[(1333, 640), (1333, 800)],
+# multiscale_mode='range'
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(
+ type=LoadAnnotations, with_bbox=True, with_mask=True, poly2mask=False),
+ dict(
+ type='RandomResize', scale=[(1333, 640), (1333, 800)],
+ keep_ratio=True),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=PackDetInputs)
+]
+test_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=Resize, scale=(1333, 800), keep_ratio=True),
+ dict(
+ type=LoadAnnotations, with_bbox=True, with_mask=True, poly2mask=False),
+ dict(
+ type=PackDetInputs,
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader.update(
+ dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ pin_memory=True,
+ sampler=dict(type=DefaultSampler, shuffle=True),
+ batch_sampler=dict(type=AspectRatioBatchSampler),
+ dataset=dict(
+ type=RepeatDataset,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args)))
+val_dataloader.update(
+ dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ pin_memory=True,
+ sampler=dict(type=DefaultSampler, shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args)))
+test_dataloader = val_dataloader
+
+val_evaluator.update(
+ dict(
+ type=CocoMetric,
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric=['bbox', 'segm'],
+ backend_args=backend_args))
+test_evaluator = val_evaluator
+
+# training schedule for 3x with `RepeatDataset`
+train_cfg.update(dict(type=EpochBasedTrainLoop, max_iters=12, val_interval=1))
+val_cfg.update(dict(type=ValLoop))
+test_cfg.update(dict(type=TestLoop))
+
+# learning rate
+param_scheduler = [
+ dict(type=LinearLR, start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=12,
+ by_epoch=False,
+ milestones=[9, 11],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper.update(
+ dict(
+ type=OptimWrapper,
+ optimizer=dict(type=SGD, lr=0.02, momentum=0.9, weight_decay=0.0001)))
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr.update(dict(enable=False, base_batch_size=16))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ms_poly_90k_coco_instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ms_poly_90k_coco_instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..52367350137035604ea167e5732a791c2e9cae87
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ms_poly_90k_coco_instance.py
@@ -0,0 +1,153 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.default_runtime import *
+
+from mmcv.transforms import RandomChoiceResize
+from mmengine.dataset import RepeatDataset
+from mmengine.dataset.sampler import DefaultSampler, InfiniteSampler
+from mmengine.optim import OptimWrapper
+from mmengine.optim.scheduler.lr_scheduler import LinearLR, MultiStepLR
+from mmengine.runner.loops import IterBasedTrainLoop, TestLoop, ValLoop
+from torch.optim import SGD
+
+from mmdet.datasets import AspectRatioBatchSampler, CocoDataset
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import (FilterAnnotations,
+ LoadAnnotations,
+ LoadImageFromFile)
+from mmdet.datasets.transforms.transforms import (CachedMixUp, CachedMosaic,
+ Pad, RandomCrop, RandomFlip,
+ RandomResize, Resize)
+from mmdet.evaluation import CocoMetric
+
+# dataset settings
+dataset_type = CocoDataset
+data_root = 'data/coco/'
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+# Align with Detectron2
+backend = 'pillow'
+train_pipeline = [
+ dict(
+ type=LoadImageFromFile,
+ backend_args=backend_args,
+ imdecode_backend=backend),
+ dict(
+ type=LoadAnnotations, with_bbox=True, with_mask=True, poly2mask=False),
+ dict(
+ type=RandomChoiceResize,
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True,
+ backend=backend),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=PackDetInputs)
+]
+test_pipeline = [
+ dict(
+ type=LoadImageFromFile,
+ backend_args=backend_args,
+ imdecode_backend=backend),
+ dict(type=Resize, scale=(1333, 800), keep_ratio=True, backend=backend),
+ dict(
+ type=LoadAnnotations, with_bbox=True, with_mask=True, poly2mask=False),
+ dict(
+ type=PackDetInputs,
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader.update(
+ dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ pin_memory=True,
+ sampler=dict(type=InfiniteSampler, shuffle=True),
+ batch_sampler=dict(type=AspectRatioBatchSampler),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args)))
+val_dataloader.update(
+ dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ pin_memory=True,
+ sampler=dict(type=DefaultSampler, shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args)))
+test_dataloader = val_dataloader
+
+val_evaluator.update(
+ dict(
+ type=CocoMetric,
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric=['bbox', 'segm'],
+ format_only=False,
+ backend_args=backend_args))
+test_evaluator = val_evaluator
+
+# training schedule for 90k
+max_iter = 90000
+train_cfg.update(
+ dict(type=IterBasedTrainLoop, max_iters=max_iter, val_interval=10000))
+val_cfg.update(dict(type=ValLoop))
+test_cfg.update(dict(type=TestLoop))
+
+# learning rate
+param_scheduler = [
+ dict(type=LinearLR, start_factor=0.001, by_epoch=False, begin=0, end=1000),
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=max_iter,
+ by_epoch=False,
+ milestones=[60000, 80000],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper.update(
+ dict(
+ type=OptimWrapper,
+ optimizer=dict(type=SGD, lr=0.02, momentum=0.9, weight_decay=0.0001)))
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr.update(dict(enable=False, base_batch_size=16))
+
+default_hooks.update(dict(checkpoint=dict(by_epoch=False, interval=10000)))
+log_processor.update(dict(by_epoch=False))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ssj_270_coco_instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ssj_270_coco_instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..ee86fdad4eca5b87ac0066b635e098d6a927bb49
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ssj_270_coco_instance.py
@@ -0,0 +1,158 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.default_runtime import *
+
+from mmcv.transforms import RandomChoiceResize
+from mmengine.dataset import RepeatDataset
+from mmengine.dataset.sampler import DefaultSampler, InfiniteSampler
+from mmengine.optim import OptimWrapper
+from mmengine.optim.scheduler.lr_scheduler import LinearLR, MultiStepLR
+from mmengine.runner.loops import IterBasedTrainLoop, TestLoop, ValLoop
+from torch.optim import SGD
+
+from mmdet.datasets import AspectRatioBatchSampler, CocoDataset
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import (FilterAnnotations,
+ LoadAnnotations,
+ LoadImageFromFile)
+from mmdet.datasets.transforms.transforms import (CachedMixUp, CachedMosaic,
+ Pad, RandomCrop, RandomFlip,
+ RandomResize, Resize)
+from mmdet.evaluation import CocoMetric
+
+# dataset settings
+dataset_type = CocoDataset
+data_root = 'data/coco/'
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+# Standard Scale Jittering (SSJ) resizes and crops an image
+# with a resize range of 0.8 to 1.25 of the original image size.
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=LoadAnnotations, with_bbox=True, with_mask=True),
+ dict(
+ type=RandomResize,
+ scale=image_size,
+ ratio_range=(0.8, 1.25),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size,
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=PackDetInputs)
+]
+test_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=Resize, scale=(1333, 800), keep_ratio=True),
+ dict(type=LoadAnnotations, with_bbox=True, with_mask=True),
+ dict(
+ type=PackDetInputs,
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+train_dataloader.update(
+ dict(
+ batch_size=2,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type=InfiniteSampler),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args)))
+val_dataloader.update(
+ dict(
+ batch_size=1,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ sampler=dict(type=DefaultSampler, shuffle=False),
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_val2017.json',
+ data_prefix=dict(img='val2017/'),
+ test_mode=True,
+ pipeline=test_pipeline,
+ backend_args=backend_args)))
+test_dataloader = val_dataloader
+
+val_evaluator.update(
+ dict(
+ type=CocoMetric,
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric=['bbox', 'segm'],
+ format_only=False,
+ backend_args=backend_args))
+test_evaluator = val_evaluator
+
+val_evaluator = dict(
+ type=CocoMetric,
+ ann_file=data_root + 'annotations/instances_val2017.json',
+ metric=['bbox', 'segm'],
+ format_only=False,
+ backend_args=backend_args)
+test_evaluator = val_evaluator
+
+# The model is trained by 270k iterations with batch_size 64,
+# which is roughly equivalent to 144 epochs.
+
+max_iter = 270000
+train_cfg.update(
+ dict(type=IterBasedTrainLoop, max_iters=max_iter, val_interval=10000))
+val_cfg.update(dict(type=ValLoop))
+test_cfg.update(dict(type=TestLoop))
+
+# learning rate
+param_scheduler = [
+ dict(type=LinearLR, start_factor=0.001, by_epoch=False, begin=0, end=1000),
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=max_iter,
+ by_epoch=False,
+ milestones=[243000, 256500, 263250],
+ gamma=0.1)
+]
+
+# optimizer
+optim_wrapper.update(
+ dict(
+ type=OptimWrapper,
+ optimizer=dict(type=SGD, lr=0.1, momentum=0.9, weight_decay=0.00004)))
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (8 GPUs) x (2 samples per GPU).
+auto_scale_lr.update(dict(base_batch_size=64))
+
+default_hooks.update(dict(checkpoint=dict(by_epoch=False, interval=10000)))
+log_processor.update(dict(by_epoch=False))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ssj_scp_270k_coco_instance.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ssj_scp_270k_coco_instance.py
new file mode 100644
index 0000000000000000000000000000000000000000..68bb1f0904fcb4de3e2f892355e489f52f53d960
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/common/ssj_scp_270k_coco_instance.py
@@ -0,0 +1,70 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .ssj_270_coco_instance import *
+
+from mmdet.datasets import MultiImageMixDataset
+from mmdet.datasets.transforms import CopyPaste
+
+# dataset settings
+dataset_type = CocoDataset
+data_root = 'data/coco/'
+image_size = (1024, 1024)
+# Example to use different file client
+# Method 1: simply set the data root and let the file I/O module
+# automatically infer from prefix (not support LMDB and Memcache yet)
+
+# data_root = 's3://openmmlab/datasets/detection/coco/'
+
+# Method 2: Use `backend_args`, `file_client_args` in versions before 3.0.0rc6
+# backend_args = dict(
+# backend='petrel',
+# path_mapping=dict({
+# './data/': 's3://openmmlab/datasets/detection/',
+# 'data/': 's3://openmmlab/datasets/detection/'
+# }))
+backend_args = None
+
+# Standard Scale Jittering (SSJ) resizes and crops an image
+# with a resize range of 0.8 to 1.25 of the original image size.
+load_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=LoadAnnotations, with_bbox=True, with_mask=True),
+ dict(
+ type=RandomResize,
+ scale=image_size,
+ ratio_range=(0.8, 1.25),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size,
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=Pad, size=image_size),
+]
+train_pipeline = [
+ dict(type=CopyPaste, max_num_pasted=100),
+ dict(type=PackDetInputs)
+]
+
+train_dataloader.update(
+ dict(
+ type=MultiImageMixDataset,
+ dataset=dict(
+ type=dataset_type,
+ data_root=data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=load_pipeline,
+ backend_args=backend_args),
+ pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/deformable_detr/deformable_detr_r50_16xb2_50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/deformable_detr/deformable_detr_r50_16xb2_50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ee2a41639d84ed8e278af45229b451b742ac8974
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/deformable_detr/deformable_detr_r50_16xb2_50e_coco.py
@@ -0,0 +1,186 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.datasets.coco_detection import *
+ from .._base_.default_runtime import *
+
+from mmcv.transforms import LoadImageFromFile, RandomChoice, RandomChoiceResize
+from mmengine.optim.optimizer import OptimWrapper
+from mmengine.optim.scheduler import MultiStepLR
+from mmengine.runner.loops import EpochBasedTrainLoop, TestLoop, ValLoop
+from torch.optim.adamw import AdamW
+
+from mmdet.datasets.transforms import (LoadAnnotations, PackDetInputs,
+ RandomCrop, RandomFlip, Resize)
+from mmdet.models.backbones import ResNet
+from mmdet.models.data_preprocessors import DetDataPreprocessor
+from mmdet.models.dense_heads import DeformableDETRHead
+from mmdet.models.detectors import DeformableDETR
+from mmdet.models.losses import FocalLoss, GIoULoss, L1Loss
+from mmdet.models.necks import ChannelMapper
+from mmdet.models.task_modules import (BBoxL1Cost, FocalLossCost,
+ HungarianAssigner, IoUCost)
+
+model = dict(
+ type=DeformableDETR,
+ num_queries=300,
+ num_feature_levels=4,
+ with_box_refine=False,
+ as_two_stage=False,
+ data_preprocessor=dict(
+ type=DetDataPreprocessor,
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=1),
+ backbone=dict(
+ type=ResNet,
+ depth=50,
+ num_stages=4,
+ out_indices=(1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type=ChannelMapper,
+ in_channels=[512, 1024, 2048],
+ kernel_size=1,
+ out_channels=256,
+ act_cfg=None,
+ norm_cfg=dict(type='GN', num_groups=32),
+ num_outs=4),
+ encoder=dict( # DeformableDetrTransformerEncoder
+ num_layers=6,
+ layer_cfg=dict( # DeformableDetrTransformerEncoderLayer
+ self_attn_cfg=dict( # MultiScaleDeformableAttention
+ embed_dims=256,
+ batch_first=True),
+ ffn_cfg=dict(
+ embed_dims=256, feedforward_channels=1024, ffn_drop=0.1))),
+ decoder=dict( # DeformableDetrTransformerDecoder
+ num_layers=6,
+ return_intermediate=True,
+ layer_cfg=dict( # DeformableDetrTransformerDecoderLayer
+ self_attn_cfg=dict( # MultiheadAttention
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.1,
+ batch_first=True),
+ cross_attn_cfg=dict( # MultiScaleDeformableAttention
+ embed_dims=256,
+ batch_first=True),
+ ffn_cfg=dict(
+ embed_dims=256, feedforward_channels=1024, ffn_drop=0.1)),
+ post_norm_cfg=None),
+ positional_encoding=dict(num_feats=128, normalize=True, offset=-0.5),
+ bbox_head=dict(
+ type=DeformableDETRHead,
+ num_classes=80,
+ sync_cls_avg_factor=True,
+ loss_cls=dict(
+ type=FocalLoss,
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=2.0),
+ loss_bbox=dict(type=L1Loss, loss_weight=5.0),
+ loss_iou=dict(type=GIoULoss, loss_weight=2.0)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type=HungarianAssigner,
+ match_costs=[
+ dict(type=FocalLossCost, weight=2.0),
+ dict(type=BBoxL1Cost, weight=5.0, box_format='xywh'),
+ dict(type=IoUCost, iou_mode='giou', weight=2.0)
+ ])),
+ test_cfg=dict(max_per_img=100))
+
+# train_pipeline, NOTE the img_scale and the Pad's size_divisor is different
+# from the default setting in mmdet.
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=LoadAnnotations, with_bbox=True),
+ dict(type=RandomFlip, prob=0.5),
+ dict(
+ type=RandomChoice,
+ transforms=[
+ [
+ dict(
+ type=RandomChoiceResize,
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ resize_type=Resize,
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type=RandomChoiceResize,
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ resize_type=Resize,
+ keep_ratio=True),
+ dict(
+ type=RandomCrop,
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type=RandomChoiceResize,
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ resize_type=Resize,
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type=PackDetInputs)
+]
+train_dataloader.update(
+ dict(
+ dataset=dict(
+ filter_cfg=dict(filter_empty_gt=False), pipeline=train_pipeline)))
+
+# optimizer
+optim_wrapper = dict(
+ type=OptimWrapper,
+ optimizer=dict(type=AdamW, lr=0.0002, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'backbone': dict(lr_mult=0.1),
+ 'sampling_offsets': dict(lr_mult=0.1),
+ 'reference_points': dict(lr_mult=0.1)
+ }))
+
+# learning policy
+max_epochs = 50
+train_cfg = dict(
+ type=EpochBasedTrainLoop, max_epochs=max_epochs, val_interval=1)
+val_cfg = dict(type=ValLoop)
+test_cfg = dict(type=TestLoop)
+
+param_scheduler = [
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[40],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (16 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=32)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/deformable_detr/deformable_detr_refine_r50_16xb2_50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/deformable_detr/deformable_detr_refine_r50_16xb2_50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..4f232d6111026488020e586440852c012dd94608
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/deformable_detr/deformable_detr_refine_r50_16xb2_50e_coco.py
@@ -0,0 +1,12 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .deformable_detr_r50_16xb2_50e_coco import *
+
+model.update(dict(with_box_refine=True))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/deformable_detr/deformable_detr_refine_twostage_r50_16xb2_50e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/deformable_detr/deformable_detr_refine_twostage_r50_16xb2_50e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1fac4d8c4f2020b6d87857fbe157419e4c4f0712
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/deformable_detr/deformable_detr_refine_twostage_r50_16xb2_50e_coco.py
@@ -0,0 +1,12 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .deformable_detr_refine_r50_16xb2_50e_coco import *
+
+model.update(dict(as_two_stage=True))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/detr/detr_r101_8xb2_500e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/detr/detr_r101_8xb2_500e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b961468114ce3adb0582378ac422649ef3bd5013
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/detr/detr_r101_8xb2_500e_coco.py
@@ -0,0 +1,13 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.config import read_base
+from mmengine.model.weight_init import PretrainedInit
+
+with read_base():
+ from .detr_r50_8xb2_500e_coco import *
+
+model.update(
+ dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type=PretrainedInit, checkpoint='torchvision://resnet101'))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/detr/detr_r18_8xb2_500e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/detr/detr_r18_8xb2_500e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..11360af18de729bfd9e8d8cb6597067a588852c9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/detr/detr_r18_8xb2_500e_coco.py
@@ -0,0 +1,14 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.config import read_base
+from mmengine.model.weight_init import PretrainedInit
+
+with read_base():
+ from .detr_r50_8xb2_500e_coco import *
+
+model.update(
+ dict(
+ backbone=dict(
+ depth=18,
+ init_cfg=dict(
+ type=PretrainedInit, checkpoint='torchvision://resnet18')),
+ neck=dict(in_channels=[512])))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/detr/detr_r50_8xb2_150e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/detr/detr_r50_8xb2_150e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c50726c7890cb59bee4b921179be1949ff12199e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/detr/detr_r50_8xb2_150e_coco.py
@@ -0,0 +1,182 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.transforms import RandomChoice, RandomChoiceResize
+from mmcv.transforms.loading import LoadImageFromFile
+from mmengine.config import read_base
+from mmengine.model.weight_init import PretrainedInit
+from mmengine.optim.optimizer.optimizer_wrapper import OptimWrapper
+from mmengine.optim.scheduler.lr_scheduler import MultiStepLR
+from mmengine.runner.loops import EpochBasedTrainLoop, TestLoop, ValLoop
+from torch.nn.modules.activation import ReLU
+from torch.nn.modules.batchnorm import BatchNorm2d
+from torch.optim.adamw import AdamW
+
+from mmdet.datasets.transforms import (LoadAnnotations, PackDetInputs,
+ RandomCrop, RandomFlip, Resize)
+from mmdet.models import (DETR, ChannelMapper, DetDataPreprocessor, DETRHead,
+ ResNet)
+from mmdet.models.losses.cross_entropy_loss import CrossEntropyLoss
+from mmdet.models.losses.iou_loss import GIoULoss
+from mmdet.models.losses.smooth_l1_loss import L1Loss
+from mmdet.models.task_modules import (BBoxL1Cost, ClassificationCost,
+ HungarianAssigner, IoUCost)
+
+with read_base():
+ from .._base_.datasets.coco_detection import *
+ from .._base_.default_runtime import *
+
+model = dict(
+ type=DETR,
+ num_queries=100,
+ data_preprocessor=dict(
+ type=DetDataPreprocessor,
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=1),
+ backbone=dict(
+ type=ResNet,
+ depth=50,
+ num_stages=4,
+ out_indices=(3, ),
+ frozen_stages=1,
+ norm_cfg=dict(type=BatchNorm2d, requires_grad=False),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type=PretrainedInit, checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type=ChannelMapper,
+ in_channels=[2048],
+ kernel_size=1,
+ out_channels=256,
+ act_cfg=None,
+ norm_cfg=None,
+ num_outs=1),
+ encoder=dict( # DetrTransformerEncoder
+ num_layers=6,
+ layer_cfg=dict( # DetrTransformerEncoderLayer
+ self_attn_cfg=dict( # MultiheadAttention
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.1,
+ batch_first=True),
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048,
+ num_fcs=2,
+ ffn_drop=0.1,
+ act_cfg=dict(type=ReLU, inplace=True)))),
+ decoder=dict( # DetrTransformerDecoder
+ num_layers=6,
+ layer_cfg=dict( # DetrTransformerDecoderLayer
+ self_attn_cfg=dict( # MultiheadAttention
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.1,
+ batch_first=True),
+ cross_attn_cfg=dict( # MultiheadAttention
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.1,
+ batch_first=True),
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048,
+ num_fcs=2,
+ ffn_drop=0.1,
+ act_cfg=dict(type=ReLU, inplace=True))),
+ return_intermediate=True),
+ positional_encoding=dict(num_feats=128, normalize=True),
+ bbox_head=dict(
+ type=DETRHead,
+ num_classes=80,
+ embed_dims=256,
+ loss_cls=dict(
+ type=CrossEntropyLoss,
+ bg_cls_weight=0.1,
+ use_sigmoid=False,
+ loss_weight=1.0,
+ class_weight=1.0),
+ loss_bbox=dict(type=L1Loss, loss_weight=5.0),
+ loss_iou=dict(type=GIoULoss, loss_weight=2.0)),
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type=HungarianAssigner,
+ match_costs=[
+ dict(type=ClassificationCost, weight=1.),
+ dict(type=BBoxL1Cost, weight=5.0, box_format='xywh'),
+ dict(type=IoUCost, iou_mode='giou', weight=2.0)
+ ])),
+ test_cfg=dict(max_per_img=100))
+
+# train_pipeline, NOTE the img_scale and the Pad's size_divisor is different
+# from the default setting in mmdet.
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=LoadAnnotations, with_bbox=True),
+ dict(type=RandomFlip, prob=0.5),
+ dict(
+ type=RandomChoice,
+ transforms=[[
+ dict(
+ type=RandomChoiceResize,
+ resize_type=Resize,
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type=RandomChoiceResize,
+ resize_type=Resize,
+ scales=[(400, 1333), (500, 1333), (600, 1333)],
+ keep_ratio=True),
+ dict(
+ type=RandomCrop,
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type=RandomChoiceResize,
+ resize_type=Resize,
+ scales=[(480, 1333), (512, 1333), (544, 1333),
+ (576, 1333), (608, 1333), (640, 1333),
+ (672, 1333), (704, 1333), (736, 1333),
+ (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]]),
+ dict(type=PackDetInputs)
+]
+train_dataloader.update(dataset=dict(pipeline=train_pipeline))
+
+# optimizer
+optim_wrapper = dict(
+ type=OptimWrapper,
+ optimizer=dict(type=AdamW, lr=0.0001, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(
+ custom_keys={'backbone': dict(lr_mult=0.1, decay_mult=1.0)}))
+
+# learning policy
+max_epochs = 150
+train_cfg = dict(
+ type=EpochBasedTrainLoop, max_epochs=max_epochs, val_interval=1)
+val_cfg = dict(type=ValLoop)
+test_cfg = dict(type=TestLoop)
+
+param_scheduler = [
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[100],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/detr/detr_r50_8xb2_500e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/detr/detr_r50_8xb2_500e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d7d0817766255a84237f0aea917806e191d161df
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/detr/detr_r50_8xb2_500e_coco.py
@@ -0,0 +1,25 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.config import read_base
+from mmengine.optim.scheduler.lr_scheduler import MultiStepLR
+from mmengine.runner.loops import EpochBasedTrainLoop
+
+with read_base():
+ from .detr_r50_8xb2_150e_coco import *
+
+# learning policy
+max_epochs = 500
+train_cfg.update(
+ type=EpochBasedTrainLoop, max_epochs=max_epochs, val_interval=10)
+
+param_scheduler = [
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[334],
+ gamma=0.1)
+]
+
+# only keep latest 2 checkpoints
+default_hooks.update(checkpoint=dict(max_keep_ckpts=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/dino/dino_4scale_r50_8xb2_12e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/dino/dino_4scale_r50_8xb2_12e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ab8e95a9a76c0cedb78c66993fc7fb7f4623029c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/dino/dino_4scale_r50_8xb2_12e_coco.py
@@ -0,0 +1,190 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.transforms import RandomChoice, RandomChoiceResize
+from mmcv.transforms.loading import LoadImageFromFile
+from mmengine.config import read_base
+from mmengine.model.weight_init import PretrainedInit
+from mmengine.optim.optimizer.optimizer_wrapper import OptimWrapper
+from mmengine.optim.scheduler.lr_scheduler import MultiStepLR
+from mmengine.runner.loops import EpochBasedTrainLoop, TestLoop, ValLoop
+from torch.nn.modules.batchnorm import BatchNorm2d
+from torch.nn.modules.normalization import GroupNorm
+from torch.optim.adamw import AdamW
+
+from mmdet.datasets.transforms import (LoadAnnotations, PackDetInputs,
+ RandomCrop, RandomFlip, Resize)
+from mmdet.models import (DINO, ChannelMapper, DetDataPreprocessor, DINOHead,
+ ResNet)
+from mmdet.models.losses.focal_loss import FocalLoss
+from mmdet.models.losses.iou_loss import GIoULoss
+from mmdet.models.losses.smooth_l1_loss import L1Loss
+from mmdet.models.task_modules import (BBoxL1Cost, FocalLossCost,
+ HungarianAssigner, IoUCost)
+
+with read_base():
+ from .._base_.datasets.coco_detection import *
+ from .._base_.default_runtime import *
+
+model = dict(
+ type=DINO,
+ num_queries=900, # num_matching_queries
+ with_box_refine=True,
+ as_two_stage=True,
+ data_preprocessor=dict(
+ type=DetDataPreprocessor,
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=1),
+ backbone=dict(
+ type=ResNet,
+ depth=50,
+ num_stages=4,
+ out_indices=(1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type=BatchNorm2d, requires_grad=False),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type=PretrainedInit, checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type=ChannelMapper,
+ in_channels=[512, 1024, 2048],
+ kernel_size=1,
+ out_channels=256,
+ act_cfg=None,
+ norm_cfg=dict(type=GroupNorm, num_groups=32),
+ num_outs=4),
+ encoder=dict(
+ num_layers=6,
+ layer_cfg=dict(
+ self_attn_cfg=dict(embed_dims=256, num_levels=4,
+ dropout=0.0), # 0.1 for DeformDETR
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048, # 1024 for DeformDETR
+ ffn_drop=0.0))), # 0.1 for DeformDETR
+ decoder=dict(
+ num_layers=6,
+ return_intermediate=True,
+ layer_cfg=dict(
+ self_attn_cfg=dict(embed_dims=256, num_heads=8,
+ dropout=0.0), # 0.1 for DeformDETR
+ cross_attn_cfg=dict(embed_dims=256, num_levels=4,
+ dropout=0.0), # 0.1 for DeformDETR
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048, # 1024 for DeformDETR
+ ffn_drop=0.0)), # 0.1 for DeformDETR
+ post_norm_cfg=None),
+ positional_encoding=dict(
+ num_feats=128,
+ normalize=True,
+ offset=0.0, # -0.5 for DeformDETR
+ temperature=20), # 10000 for DeformDETR
+ bbox_head=dict(
+ type=DINOHead,
+ num_classes=80,
+ sync_cls_avg_factor=True,
+ loss_cls=dict(
+ type=FocalLoss,
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0), # 2.0 in DeformDETR
+ loss_bbox=dict(type=L1Loss, loss_weight=5.0),
+ loss_iou=dict(type=GIoULoss, loss_weight=2.0)),
+ dn_cfg=dict( # TODO: Move to model.train_cfg ?
+ label_noise_scale=0.5,
+ box_noise_scale=1.0, # 0.4 for DN-DETR
+ group_cfg=dict(dynamic=True, num_groups=None,
+ num_dn_queries=100)), # TODO: half num_dn_queries
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type=HungarianAssigner,
+ match_costs=[
+ dict(type=FocalLossCost, weight=2.0),
+ dict(type=BBoxL1Cost, weight=5.0, box_format='xywh'),
+ dict(type=IoUCost, iou_mode='giou', weight=2.0)
+ ])),
+ test_cfg=dict(max_per_img=300)) # 100 for DeformDETR
+
+# train_pipeline, NOTE the img_scale and the Pad's size_divisor is different
+# from the default setting in mmdet.
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=LoadAnnotations, with_bbox=True),
+ dict(type=RandomFlip, prob=0.5),
+ dict(
+ type=RandomChoice,
+ transforms=[
+ [
+ dict(
+ type=RandomChoiceResize,
+ resize_type=Resize,
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type=RandomChoiceResize,
+ resize_type=Resize,
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type=RandomCrop,
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type=RandomChoiceResize,
+ resize_type=Resize,
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type=PackDetInputs)
+]
+train_dataloader.update(
+ dataset=dict(
+ filter_cfg=dict(filter_empty_gt=False), pipeline=train_pipeline))
+
+# optimizer
+optim_wrapper = dict(
+ type=OptimWrapper,
+ optimizer=dict(
+ type=AdamW,
+ lr=0.0001, # 0.0002 for DeformDETR
+ weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(custom_keys={'backbone': dict(lr_mult=0.1)})
+) # custom_keys contains sampling_offsets and reference_points in DeformDETR # noqa
+
+# learning policy
+max_epochs = 12
+train_cfg = dict(
+ type=EpochBasedTrainLoop, max_epochs=max_epochs, val_interval=1)
+
+val_cfg = dict(type=ValLoop)
+test_cfg = dict(type=TestLoop)
+
+param_scheduler = [
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[11],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/dino/dino_4scale_r50_8xb2_24e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/dino/dino_4scale_r50_8xb2_24e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c10cc2184de8f71571759ecbeac56696afceb5eb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/dino/dino_4scale_r50_8xb2_24e_coco.py
@@ -0,0 +1,12 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.config import read_base
+from mmengine.runner.loops import EpochBasedTrainLoop
+
+with read_base():
+ from .dino_4scale_r50_8xb2_12e_coco import *
+
+max_epochs = 24
+train_cfg.update(
+ dict(type=EpochBasedTrainLoop, max_epochs=max_epochs, val_interval=1))
+
+param_scheduler[0].update(dict(milestones=[20]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/dino/dino_4scale_r50_8xb2_36e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/dino/dino_4scale_r50_8xb2_36e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..3779744322a19d2865f1e6299aba564c4ec1e3d5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/dino/dino_4scale_r50_8xb2_36e_coco.py
@@ -0,0 +1,12 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.config import read_base
+from mmengine.runner.loops import EpochBasedTrainLoop
+
+with read_base():
+ from .dino_4scale_r50_8xb2_12e_coco import *
+
+max_epochs = 36
+train_cfg.update(
+ dict(type=EpochBasedTrainLoop, max_epochs=max_epochs, val_interval=1))
+
+param_scheduler[0].update(dict(milestones=[30]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/dino/dino_4scale_r50_improved_8xb2_12e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/dino/dino_4scale_r50_improved_8xb2_12e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..43c07201079fcdbad3c9ea7a471306080e006cdc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/dino/dino_4scale_r50_improved_8xb2_12e_coco.py
@@ -0,0 +1,24 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.config import read_base
+
+with read_base():
+ from .dino_4scale_r50_8xb2_12e_coco import *
+
+# from deformable detr hyper
+model.update(
+ dict(
+ backbone=dict(frozen_stages=-1),
+ bbox_head=dict(loss_cls=dict(loss_weight=2.0)),
+ positional_encoding=dict(offset=-0.5, temperature=10000),
+ dn_cfg=dict(group_cfg=dict(num_dn_queries=300))))
+
+# optimizer
+optim_wrapper.update(
+ dict(
+ optimizer=dict(lr=0.0002),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'backbone': dict(lr_mult=0.1),
+ 'sampling_offsets': dict(lr_mult=0.1),
+ 'reference_points': dict(lr_mult=0.1)
+ })))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/dino/dino_5scale_swin_l_8xb2_12e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/dino/dino_5scale_swin_l_8xb2_12e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..25aac0187ab2472dd062514fecf988dcd47504a5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/dino/dino_5scale_swin_l_8xb2_12e_coco.py
@@ -0,0 +1,40 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.config import read_base
+from mmengine.model.weight_init import PretrainedInit
+
+from mmdet.models import SwinTransformer
+
+with read_base():
+ from .dino_4scale_r50_8xb2_12e_coco import *
+
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_large_patch4_window12_384_22k.pth' # noqa
+num_levels = 5
+model.merge(
+ dict(
+ num_feature_levels=num_levels,
+ backbone=dict(
+ _delete_=True,
+ type=SwinTransformer,
+ pretrain_img_size=384,
+ embed_dims=192,
+ depths=[2, 2, 18, 2],
+ num_heads=[6, 12, 24, 48],
+ window_size=12,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.2,
+ patch_norm=True,
+ out_indices=(0, 1, 2, 3),
+ # Please only add indices that would be used
+ # in FPN, otherwise some parameter will not be used
+ with_cp=True,
+ convert_weights=True,
+ init_cfg=dict(type=PretrainedInit, checkpoint=pretrained)),
+ neck=dict(in_channels=[192, 384, 768, 1536], num_outs=num_levels),
+ encoder=dict(
+ layer_cfg=dict(self_attn_cfg=dict(num_levels=num_levels))),
+ decoder=dict(
+ layer_cfg=dict(cross_attn_cfg=dict(num_levels=num_levels)))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/dino/dino_5scale_swin_l_8xb2_36e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/dino/dino_5scale_swin_l_8xb2_36e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..494acf59f1c31fe419415920e8b65fbfb9267df1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/dino/dino_5scale_swin_l_8xb2_36e_coco.py
@@ -0,0 +1,12 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.config import read_base
+from mmengine.runner.loops import EpochBasedTrainLoop
+
+with read_base():
+ from .dino_5scale_swin_l_8xb2_12e_coco import *
+
+max_epochs = 36
+train_cfg.update(
+ dict(type=EpochBasedTrainLoop, max_epochs=max_epochs, val_interval=1))
+
+param_scheduler[0].update(dict(milestones=[27, 33]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/faster_rcnn/faster_rcnn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/faster_rcnn/faster_rcnn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f0a6d5a21470752fd26fa162edf5c2241afb1fed
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/faster_rcnn/faster_rcnn_r50_fpn_1x_coco.py
@@ -0,0 +1,13 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.datasets.coco_detection import *
+ from .._base_.default_runtime import *
+ from .._base_.models.faster_rcnn_r50_fpn import *
+ from .._base_.schedules.schedule_1x import *
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r101_caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r101_caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..2780f4afddc05ccd4ae1746206a6a6ad8cece39e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r101_caffe_fpn_1x_coco.py
@@ -0,0 +1,19 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .mask_rcnn_r50_fpn_poly_1x_coco import *
+
+from mmengine.model.weight_init import PretrainedInit
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type=PretrainedInit,
+ checkpoint='open-mmlab://detectron2/resnet101_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r101_caffe_fpn_ms_poly_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r101_caffe_fpn_ms_poly_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8a1badfc4f04f6ad5466d9ec3aa2d07708887927
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r101_caffe_fpn_ms_poly_3x_coco.py
@@ -0,0 +1,28 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from ..common.ms_poly_3x_coco_instance import *
+ from .._base_.models.mask_rcnn_r50_fpn import *
+
+from mmengine.model.weight_init import PretrainedInit
+
+model = dict(
+ # use caffe img_norm
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False),
+ backbone=dict(
+ depth=101,
+ norm_cfg=dict(requires_grad=False),
+ norm_eval=True,
+ style='caffe',
+ init_cfg=dict(
+ type=PretrainedInit,
+ checkpoint='open-mmlab://detectron2/resnet101_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6770cec8eebe8c5130abd15f9bc44d5b5c5db875
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r101_fpn_1x_coco.py
@@ -0,0 +1,18 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.models.mask_rcnn_r50_fpn import *
+
+from mmengine.model.weight_init import PretrainedInit
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type=PretrainedInit, checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r101_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r101_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..fd2aafb912ca84f776637e498d2743213a05d18a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r101_fpn_2x_coco.py
@@ -0,0 +1,18 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .mask_rcnn_r50_fpn_2x_coco import *
+
+from mmengine.model.weight_init import PretrainedInit
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type=PretrainedInit, checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r101_fpn_8xb8_amp_lsj_200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r101_fpn_8xb8_amp_lsj_200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..665808d5dc479ecb7c5a328af3861f59e460ac78
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r101_fpn_8xb8_amp_lsj_200e_coco.py
@@ -0,0 +1,18 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .mask_rcnn_r18_fpn_8xb8_amp_lsj_200e_coco import *
+
+from mmengine.model.weight_init import PretrainedInit
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type=PretrainedInit, checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r101_fpn_ms_poly_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r101_fpn_ms_poly_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..14688795963cb28018f5897429b191b235a86b6b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r101_fpn_ms_poly_3x_coco.py
@@ -0,0 +1,19 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from ..common.ms_poly_3x_coco_instance import *
+ from .._base_.models.mask_rcnn_r50_fpn import *
+
+from mmengine.model.weight_init import PretrainedInit
+
+model = dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type=PretrainedInit, checkpoint='torchvision://resnet101')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r18_fpn_8xb8_amp_lsj_200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r18_fpn_8xb8_amp_lsj_200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..67bd86fa0e8f8b414eec681852511db3b3d4c9c6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r18_fpn_8xb8_amp_lsj_200e_coco.py
@@ -0,0 +1,19 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .mask_rcnn_r50_fpn_8xb8_amp_lsj_200e_coco import *
+
+from mmengine.model.weight_init import PretrainedInit
+
+model = dict(
+ backbone=dict(
+ depth=18,
+ init_cfg=dict(
+ type=PretrainedInit, checkpoint='torchvision://resnet18')),
+ neck=dict(in_channels=[64, 128, 256, 512]))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_c4_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_c4_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..494e6ba593efa663f06e1383ceba8b57b9d097b5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_c4_1x_coco.py
@@ -0,0 +1,13 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.datasets.coco_instance import *
+ from .._base_.default_runtime import *
+ from .._base_.models.mask_rcnn_r50_caffe_c4 import *
+ from .._base_.schedules.schedule_1x import *
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6481fcfd49eeac603eced8e46ee3a8705add8367
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_fpn_1x_coco.py
@@ -0,0 +1,25 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .mask_rcnn_r50_fpn_1x_coco import *
+
+from mmengine.model.weight_init import PretrainedInit
+
+model = dict(
+ # use caffe img_norm
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False),
+ backbone=dict(
+ norm_cfg=dict(requires_grad=False),
+ style='caffe',
+ init_cfg=dict(
+ type=PretrainedInit,
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_fpn_ms_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_fpn_ms_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5952ed587a431740bc3d17ac9d2e6b5a3d326061
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_fpn_ms_1x_coco.py
@@ -0,0 +1,40 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .mask_rcnn_r50_fpn_1x_coco import *
+
+from mmcv.transforms import RandomChoiceResize
+from mmengine.model.weight_init import PretrainedInit
+
+model = dict(
+ # use caffe img_norm
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False),
+ backbone=dict(
+ norm_cfg=dict(requires_grad=False),
+ style='caffe',
+ init_cfg=dict(
+ type=PretrainedInit,
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')))
+
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args={{_base_.backend_args}}),
+ dict(type=LoadAnnotations, with_bbox=True, with_mask=True),
+ dict(
+ type=RandomChoiceResize,
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=PackDetInputs),
+]
+
+train_dataloader.update(dict(dataset=dict(pipeline=train_pipeline)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_fpn_ms_poly_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_fpn_ms_poly_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d62b9ebe958b3a8f790a6e9581942494f42bf7d6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_fpn_ms_poly_1x_coco.py
@@ -0,0 +1,40 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .mask_rcnn_r50_fpn_1x_coco import *
+
+from mmcv.transforms import RandomChoiceResize
+from mmengine.model.weight_init import PretrainedInit
+
+model = dict(
+ # use caffe img_norm
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False),
+ backbone=dict(
+ norm_cfg=dict(requires_grad=False),
+ style='caffe',
+ init_cfg=dict(
+ type=PretrainedInit,
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')))
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args={{_base_.backend_args}}),
+ dict(
+ type=LoadAnnotations, with_bbox=True, with_mask=True, poly2mask=False),
+ dict(
+ type=RandomChoiceResize,
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=PackDetInputs)
+]
+
+train_dataloader.update(dict(dataset=dict(pipeline=train_pipeline)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_fpn_ms_poly_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_fpn_ms_poly_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..fa41b7e00ca153814f28ac29638cc497e7a2d3e9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_fpn_ms_poly_2x_coco.py
@@ -0,0 +1,23 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .mask_rcnn_r50_caffe_fpn_ms_poly_1x_coco import *
+
+train_cfg = dict(max_epochs=24)
+# learning rate
+param_scheduler = [
+ dict(type=LinearLR, start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=24,
+ by_epoch=True,
+ milestones=[16, 22],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_fpn_ms_poly_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_fpn_ms_poly_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c5f9b977b2dfaebfe01d834ac4ad8cf4522fe9c0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_fpn_ms_poly_3x_coco.py
@@ -0,0 +1,23 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .mask_rcnn_r50_caffe_fpn_ms_poly_1x_coco import *
+
+train_cfg = dict(max_epochs=36)
+# learning rate
+param_scheduler = [
+ dict(type=LinearLR, start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=24,
+ by_epoch=True,
+ milestones=[28, 34],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_fpn_poly_1x_coco_v1.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_fpn_poly_1x_coco_v1.py
new file mode 100644
index 0000000000000000000000000000000000000000..28ba7c77ddf10d295d371db2f46d6c1f117ac7c6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_caffe_fpn_poly_1x_coco_v1.py
@@ -0,0 +1,40 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .mask_rcnn_r50_fpn_1x_coco import *
+
+from mmengine.model.weight_init import PretrainedInit
+
+from mmdet.models.losses import SmoothL1Loss
+
+model = dict(
+ # use caffe img_norm
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False),
+ backbone=dict(
+ norm_cfg=dict(requires_grad=False),
+ style='caffe',
+ init_cfg=dict(
+ type=PretrainedInit,
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')),
+ rpn_head=dict(
+ loss_bbox=dict(type=SmoothL1Loss, beta=1.0 / 9.0, loss_weight=1.0)),
+ roi_head=dict(
+ bbox_roi_extractor=dict(
+ roi_layer=dict(
+ type=RoIAlign, output_size=7, sampling_ratio=2,
+ aligned=False)),
+ bbox_head=dict(
+ loss_bbox=dict(type=SmoothL1Loss, beta=1.0, loss_weight=1.0)),
+ mask_roi_extractor=dict(
+ roi_layer=dict(
+ type=RoIAlign, output_size=14, sampling_ratio=2,
+ aligned=False))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8145d08fee85c1758d3794cee952a3b7200b14bd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_1x_coco.py
@@ -0,0 +1,13 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.datasets.coco_instance import *
+ from .._base_.default_runtime import *
+ from .._base_.models.mask_rcnn_r50_fpn import *
+ from .._base_.schedules.schedule_1x import *
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_1x_wandb_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_1x_wandb_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d2c0876541289d10832a1b26ddb6e91f6a66d89a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_1x_wandb_coco.py
@@ -0,0 +1,31 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.datasets.coco_instance import *
+ from .._base_.default_runtime import *
+ from .._base_.models.mask_rcnn_r50_fpn import *
+ from .._base_.schedules.schedule_1x import *
+
+from mmengine.visualization import LocalVisBackend, WandbVisBackend
+
+vis_backends.update(dict(type=WandbVisBackend))
+vis_backends.update(dict(type=LocalVisBackend))
+visualizer.update(dict(vis_backends=vis_backends))
+
+# MMEngine support the following two ways, users can choose
+# according to convenience
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+default_hooks.update(dict(checkpoint=dict(interval=4)))
+
+train_cfg.update(dict(val_interval=2))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..6be010b4508d6ba300a1305a1d405ec9a265ae07
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_2x_coco.py
@@ -0,0 +1,13 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.datasets.coco_instance import *
+ from .._base_.default_runtime import *
+ from .._base_.models.mask_rcnn_r50_fpn import *
+ from .._base_.schedules.schedule_2x import *
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_8xb8_amp_lsj_200e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_8xb8_amp_lsj_200e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ef101fec61e72abc0eb90266d453b5b22331378d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_8xb8_amp_lsj_200e_coco.py
@@ -0,0 +1 @@
+# Copyright (c) OpenMMLab. All rights reserved.
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_amp_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_amp_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..110c3c475429701a92321676d17f829f82cbfb76
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_amp_1x_coco.py
@@ -0,0 +1,14 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .mask_rcnn_r50_fpn_1x_coco import *
+
+from mmengine.optim.optimizer.amp_optimizer_wrapper import AmpOptimWrapper
+
+optim_wrapper.update(dict(type=AmpOptimWrapper))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_ms_poly_-3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_ms_poly_-3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ff4eec6d2be0f4bd61c7bd04057fd58b303120c8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_ms_poly_-3x_coco.py
@@ -0,0 +1,11 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.models.mask_rcnn_r50_fpn import *
+ from ..common.ms_poly_3x_coco_instance import *
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_poly_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_poly_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..012e711cb96f9aa67460b86694838d592fd1ae25
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_r50_fpn_poly_1x_coco.py
@@ -0,0 +1,23 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.datasets.coco_instance import *
+ from .._base_.default_runtime import *
+ from .._base_.models.mask_rcnn_r50_fpn import *
+ from .._base_.schedules.schedule_1x import *
+
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(
+ type=LoadAnnotations, with_bbox=True, with_mask=True, poly2mask=False),
+ dict(type=Resize, scale=(1333, 800), keep_ratio=True),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=PackDetInputs),
+]
+train_dataloader.update(dict(dataset=dict(pipeline=train_pipeline)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_32x4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_32x4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5429b1bd5a62f4786936d19e65d6281807d800bf
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_32x4d_fpn_1x_coco.py
@@ -0,0 +1,28 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .mask_rcnn_r101_fpn_1x_coco import *
+
+from mmengine.model.weight_init import PretrainedInit
+
+from mmdet.models.backbones.resnext import ResNeXt
+
+model = dict(
+ backbone=dict(
+ type=ResNeXt,
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type=BatchNorm2d, requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type=PretrainedInit, checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_32x4d_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_32x4d_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ebae6c1dbc3a234ede68ba5b7a6e199edf966ead
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_32x4d_fpn_2x_coco.py
@@ -0,0 +1,28 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .mask_rcnn_r50_fpn_2x_coco import *
+
+from mmengine.model.weight_init import PretrainedInit
+
+from mmdet.models import ResNeXt
+
+model = dict(
+ backbone=dict(
+ type=ResNeXt,
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type=BatchNorm2d, requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type=PretrainedInit, checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_32x4d_fpn_ms_poly_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_32x4d_fpn_ms_poly_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..aff45d89f351037cec3115271feab678eac3382f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_32x4d_fpn_ms_poly_3x_coco.py
@@ -0,0 +1,29 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from ..common.ms_poly_3x_coco_instance import *
+ from .._base_.models.mask_rcnn_r50_fpn import *
+
+from mmengine.model.weight_init import PretrainedInit
+
+from mmdet.models.backbones import ResNeXt
+
+model = dict(
+ backbone=dict(
+ type=ResNeXt,
+ depth=101,
+ groups=32,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type=BatchNorm2d, requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type=PretrainedInit, checkpoint='open-mmlab://resnext101_32x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_32x8d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_32x8d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d9f2095dc2dff7396896a9b2af2fb05bcd765c69
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_32x8d_fpn_1x_coco.py
@@ -0,0 +1,31 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .mask_rcnn_x101_32x4d_fpn_1x_coco import *
+
+model = dict(
+ # ResNeXt-101-32x8d model trained with Caffe2 at FB,
+ # so the mean and std need to be changed.
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[57.375, 57.120, 58.395],
+ bgr_to_rgb=False),
+ backbone=dict(
+ type=ResNeXt,
+ depth=101,
+ groups=32,
+ base_width=8,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type=BatchNorm2d, requires_grad=False),
+ style='pytorch',
+ init_cfg=dict(
+ type=PretrainedInit,
+ checkpoint='open-mmlab://detectron2/resnext101_32x8d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_32x8d_fpn_ms_poly_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_32x8d_fpn_ms_poly_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8eded941751ce71b9c63baa565275802c7ee9bb2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_32x8d_fpn_ms_poly_1x_coco.py
@@ -0,0 +1,54 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .mask_rcnn_r101_fpn_1x_coco import *
+
+from mmcv.transforms import RandomChoiceResize, RandomFlip
+from mmcv.transforms.loading import LoadImageFromFile
+
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import LoadAnnotations
+from mmdet.models.backbones import ResNeXt
+
+model = dict(
+ # ResNeXt-101-32x8d model trained with Caffe2 at FB,
+ # so the mean and std need to be changed.
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[57.375, 57.120, 58.395],
+ bgr_to_rgb=False),
+ backbone=dict(
+ type=ResNeXt,
+ depth=101,
+ groups=32,
+ base_width=8,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type=BatchNorm2d, requires_grad=False),
+ style='pytorch',
+ init_cfg=dict(
+ type=PretrainedInit,
+ checkpoint='open-mmlab://detectron2/resnext101_32x8d')))
+
+backend_args = None
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(
+ type=LoadAnnotations, with_bbox=True, with_mask=True, poly2mask=False),
+ dict(
+ type=RandomChoiceResize,
+ scales=[(1333, 640), (1333, 672), (1333, 704), (1333, 736),
+ (1333, 768), (1333, 800)],
+ keep_ratio=True),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=PackDetInputs),
+]
+
+train_dataloader = dict(dataset=dict(pipeline=train_pipeline))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_32x8d_fpn_ms_poly_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_32x8d_fpn_ms_poly_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..b3f584675f6da93ec7c188753c2f0478bac25ba8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_32x8d_fpn_ms_poly_3x_coco.py
@@ -0,0 +1,34 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from ..common.ms_poly_3x_coco_instance import *
+ from .._base_.models.mask_rcnn_r50_fpn import *
+
+from mmdet.models.backbones import ResNeXt
+
+model = dict(
+ # ResNeXt-101-32x8d model trained with Caffe2 at FB,
+ # so the mean and std need to be changed.
+ data_preprocessor=dict(
+ mean=[103.530, 116.280, 123.675],
+ std=[57.375, 57.120, 58.395],
+ bgr_to_rgb=False),
+ backbone=dict(
+ type=ResNeXt,
+ depth=101,
+ groups=32,
+ base_width=8,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type=BatchNorm2d, requires_grad=False),
+ style='pytorch',
+ init_cfg=dict(
+ type=PretrainedInit,
+ checkpoint='open-mmlab://detectron2/resnext101_32x8d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_64_4d_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_64_4d_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..8bb6f636e641138b902f69a543da0bd8a656db3d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_64_4d_fpn_1x_coco.py
@@ -0,0 +1,24 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .mask_rcnn_x101_32x4d_fpn_1x_coco import *
+
+model = dict(
+ backbone=dict(
+ type=ResNeXt,
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type=BatchNorm2d, requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type=PretrainedInit, checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_64x4d_fpn_2x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_64x4d_fpn_2x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d661076dcf37df9668d6cdf726ecfc5720c561df
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_64x4d_fpn_2x_coco.py
@@ -0,0 +1,24 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .mask_rcnn_x101_32x4d_fpn_2x_coco import *
+
+model = dict(
+ backbone=dict(
+ type=ResNeXt,
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type=BatchNorm2d, requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type=PretrainedInit, checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_64x4d_fpn_ms_poly_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_64x4d_fpn_ms_poly_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d9ab3643ec27665f7a9411d95c7e01711dfe7623
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/mask_rcnn/mask_rcnn_x101_64x4d_fpn_ms_poly_3x_coco.py
@@ -0,0 +1,27 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from ..common.ms_poly_3x_coco_instance import *
+ from .._base_.models.mask_rcnn_r50_fpn import *
+
+from mmdet.models.backbones import ResNeXt
+
+model = dict(
+ backbone=dict(
+ type=ResNeXt,
+ depth=101,
+ groups=64,
+ base_width=4,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type=BatchNorm2d, requires_grad=True),
+ style='pytorch',
+ init_cfg=dict(
+ type=PretrainedInit, checkpoint='open-mmlab://resnext101_64x4d')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/maskformer/maskformer_r50_ms_16xb1_75e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/maskformer/maskformer_r50_ms_16xb1_75e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..70744013afcad76834d05ccc8aa6303dc6399bc0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/maskformer/maskformer_r50_ms_16xb1_75e_coco.py
@@ -0,0 +1,249 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.transforms import RandomChoice, RandomChoiceResize
+from mmengine.config import read_base
+from mmengine.model.weight_init import PretrainedInit
+from mmengine.optim.optimizer import OptimWrapper
+from mmengine.optim.scheduler import MultiStepLR
+from mmengine.runner import EpochBasedTrainLoop, TestLoop, ValLoop
+from torch.nn.modules.activation import ReLU
+from torch.nn.modules.batchnorm import BatchNorm2d
+from torch.nn.modules.normalization import GroupNorm
+from torch.optim.adamw import AdamW
+
+from mmdet.datasets.transforms.transforms import RandomCrop
+from mmdet.models import MaskFormer
+from mmdet.models.backbones import ResNet
+from mmdet.models.data_preprocessors.data_preprocessor import \
+ DetDataPreprocessor
+from mmdet.models.dense_heads.maskformer_head import MaskFormerHead
+from mmdet.models.layers.pixel_decoder import TransformerEncoderPixelDecoder
+from mmdet.models.losses import CrossEntropyLoss, DiceLoss, FocalLoss
+from mmdet.models.seg_heads.panoptic_fusion_heads import MaskFormerFusionHead
+from mmdet.models.task_modules.assigners.hungarian_assigner import \
+ HungarianAssigner
+from mmdet.models.task_modules.assigners.match_cost import (ClassificationCost,
+ DiceCost,
+ FocalLossCost)
+from mmdet.models.task_modules.samplers import MaskPseudoSampler
+
+with read_base():
+ from .._base_.datasets.coco_panoptic import *
+ from .._base_.default_runtime import *
+
+data_preprocessor = dict(
+ type=DetDataPreprocessor,
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=1,
+ pad_mask=True,
+ mask_pad_value=0,
+ pad_seg=True,
+ seg_pad_value=255)
+
+num_things_classes = 80
+num_stuff_classes = 53
+num_classes = num_things_classes + num_stuff_classes
+model = dict(
+ type=MaskFormer,
+ data_preprocessor=data_preprocessor,
+ backbone=dict(
+ type=ResNet,
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=-1,
+ norm_cfg=dict(type=BatchNorm2d, requires_grad=False),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(
+ type=PretrainedInit, checkpoint='torchvision://resnet50')),
+ panoptic_head=dict(
+ type=MaskFormerHead,
+ in_channels=[256, 512, 1024, 2048], # pass to pixel_decoder inside
+ feat_channels=256,
+ out_channels=256,
+ num_things_classes=num_things_classes,
+ num_stuff_classes=num_stuff_classes,
+ num_queries=100,
+ pixel_decoder=dict(
+ type=TransformerEncoderPixelDecoder,
+ norm_cfg=dict(type=GroupNorm, num_groups=32),
+ act_cfg=dict(type=ReLU),
+ encoder=dict( # DetrTransformerEncoder
+ num_layers=6,
+ layer_cfg=dict( # DetrTransformerEncoderLayer
+ self_attn_cfg=dict( # MultiheadAttention
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.1,
+ batch_first=True),
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048,
+ num_fcs=2,
+ ffn_drop=0.1,
+ act_cfg=dict(type=ReLU, inplace=True)))),
+ positional_encoding=dict(num_feats=128, normalize=True)),
+ enforce_decoder_input_project=False,
+ positional_encoding=dict(num_feats=128, normalize=True),
+ transformer_decoder=dict( # DetrTransformerDecoder
+ num_layers=6,
+ layer_cfg=dict( # DetrTransformerDecoderLayer
+ self_attn_cfg=dict( # MultiheadAttention
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.1,
+ batch_first=True),
+ cross_attn_cfg=dict( # MultiheadAttention
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.1,
+ batch_first=True),
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048,
+ num_fcs=2,
+ ffn_drop=0.1,
+ act_cfg=dict(type=ReLU, inplace=True))),
+ return_intermediate=True),
+ loss_cls=dict(
+ type=CrossEntropyLoss,
+ use_sigmoid=False,
+ loss_weight=1.0,
+ reduction='mean',
+ class_weight=[1.0] * num_classes + [0.1]),
+ loss_mask=dict(
+ type=FocalLoss,
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ reduction='mean',
+ loss_weight=20.0),
+ loss_dice=dict(
+ type=DiceLoss,
+ use_sigmoid=True,
+ activate=True,
+ reduction='mean',
+ naive_dice=True,
+ eps=1.0,
+ loss_weight=1.0)),
+ panoptic_fusion_head=dict(
+ type=MaskFormerFusionHead,
+ num_things_classes=num_things_classes,
+ num_stuff_classes=num_stuff_classes,
+ loss_panoptic=None,
+ init_cfg=None),
+ train_cfg=dict(
+ assigner=dict(
+ type=HungarianAssigner,
+ match_costs=[
+ dict(type=ClassificationCost, weight=1.0),
+ dict(type=FocalLossCost, weight=20.0, binary_input=True),
+ dict(type=DiceCost, weight=1.0, pred_act=True, eps=1.0)
+ ]),
+ sampler=dict(type=MaskPseudoSampler)),
+ test_cfg=dict(
+ panoptic_on=True,
+ # For now, the dataset does not support
+ # evaluating semantic segmentation metric.
+ semantic_on=False,
+ instance_on=False,
+ # max_per_image is for instance segmentation.
+ max_per_image=100,
+ object_mask_thr=0.8,
+ iou_thr=0.8,
+ # In MaskFormer's panoptic postprocessing,
+ # it will not filter masks whose score is smaller than 0.5 .
+ filter_low_score=False),
+ init_cfg=None)
+
+# dataset settings
+train_pipeline = [
+ dict(type=LoadImageFromFile),
+ dict(
+ type=LoadPanopticAnnotations,
+ with_bbox=True,
+ with_mask=True,
+ with_seg=True),
+ dict(type=RandomFlip, prob=0.5),
+ # dict(type=Resize, scale=(1333, 800), keep_ratio=True),
+ dict(
+ type=RandomChoice,
+ transforms=[[
+ dict(
+ type=RandomChoiceResize,
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ resize_type=Resize,
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type=RandomChoiceResize,
+ scales=[(400, 1333), (500, 1333), (600, 1333)],
+ resize_type=Resize,
+ keep_ratio=True),
+ dict(
+ type=RandomCrop,
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type=RandomChoiceResize,
+ scales=[(480, 1333), (512, 1333), (544, 1333),
+ (576, 1333), (608, 1333), (640, 1333),
+ (672, 1333), (704, 1333), (736, 1333),
+ (768, 1333), (800, 1333)],
+ resize_type=Resize,
+ keep_ratio=True)
+ ]]),
+ dict(type=PackDetInputs)
+]
+
+train_dataloader.update(
+ dict(batch_size=1, num_workers=1, dataset=dict(pipeline=train_pipeline)))
+
+val_dataloader.update(dict(batch_size=1, num_workers=1))
+
+test_dataloader = val_dataloader
+
+# optimizer
+optim_wrapper = dict(
+ type=OptimWrapper,
+ optimizer=dict(
+ type=AdamW,
+ lr=0.0001,
+ weight_decay=0.0001,
+ eps=1e-8,
+ betas=(0.9, 0.999)),
+ paramwise_cfg=dict(
+ custom_keys={
+ 'backbone': dict(lr_mult=0.1, decay_mult=1.0),
+ 'query_embed': dict(lr_mult=1.0, decay_mult=0.0)
+ },
+ norm_decay_mult=0.0),
+ clip_grad=dict(max_norm=0.01, norm_type=2))
+
+max_epochs = 75
+
+# learning rate
+param_scheduler = dict(
+ type=MultiStepLR,
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[50],
+ gamma=0.1)
+
+train_cfg = dict(
+ type=EpochBasedTrainLoop, max_epochs=max_epochs, val_interval=1)
+val_cfg = dict(type=ValLoop)
+test_cfg = dict(type=TestLoop)
+
+# Default setting for scaling LR automatically
+# - `enable` means enable scaling LR automatically
+# or not by default.
+# - `base_batch_size` = (16 GPUs) x (1 samples per GPU).
+auto_scale_lr = dict(enable=False, base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/maskformer/maskformer_swin_l_p4_w12_64xb1_ms_300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/maskformer/maskformer_swin_l_p4_w12_64xb1_ms_300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..2affe520918d0f26c0a858f97bb69646a2860f87
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/maskformer/maskformer_swin_l_p4_w12_64xb1_ms_300e_coco.py
@@ -0,0 +1,82 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.config import read_base
+from mmengine.optim.scheduler import LinearLR
+
+from mmdet.models.backbones import SwinTransformer
+from mmdet.models.layers import PixelDecoder
+
+with read_base():
+ from .maskformer_r50_ms_16xb1_75e_coco import *
+
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_large_patch4_window12_384_22k.pth' # noqa
+depths = [2, 2, 18, 2]
+model.update(
+ dict(
+ backbone=dict(
+ _delete_=True,
+ type=SwinTransformer,
+ pretrain_img_size=384,
+ embed_dims=192,
+ patch_size=4,
+ window_size=12,
+ mlp_ratio=4,
+ depths=depths,
+ num_heads=[6, 12, 24, 48],
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(0, 1, 2, 3),
+ with_cp=False,
+ convert_weights=True,
+ init_cfg=dict(type=PretrainedInit, checkpoint=pretrained)),
+ panoptic_head=dict(
+ in_channels=[192, 384, 768, 1536], # pass to pixel_decoder inside
+ pixel_decoder=dict(
+ _delete_=True,
+ type=PixelDecoder,
+ norm_cfg=dict(type=GroupNorm, num_groups=32),
+ act_cfg=dict(type=ReLU)),
+ enforce_decoder_input_project=True)))
+
+# optimizer
+
+# weight_decay = 0.01
+# norm_weight_decay = 0.0
+# embed_weight_decay = 0.0
+embed_multi = dict(lr_mult=1.0, decay_mult=0.0)
+norm_multi = dict(lr_mult=1.0, decay_mult=0.0)
+custom_keys = {
+ 'norm': norm_multi,
+ 'absolute_pos_embed': embed_multi,
+ 'relative_position_bias_table': embed_multi,
+ 'query_embed': embed_multi
+}
+
+optim_wrapper.update(
+ dict(
+ optimizer=dict(lr=6e-5, weight_decay=0.01),
+ paramwise_cfg=dict(custom_keys=custom_keys, norm_decay_mult=0.0)))
+
+max_epochs = 300
+
+# learning rate
+param_scheduler = [
+ dict(type=LinearLR, start_factor=1e-6, by_epoch=False, begin=0, end=1500),
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[250],
+ gamma=0.1)
+]
+
+train_cfg.update(dict(max_epochs=max_epochs))
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (64 GPUs) x (1 samples per GPU)
+auto_scale_lr.update(dict(base_batch_size=64))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/panoptic_fpn/panoptic_fpn_r101_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/panoptic_fpn/panoptic_fpn_r101_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c6059780da15b24d4845cdac9ad33d65a6b24e75
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/panoptic_fpn/panoptic_fpn_r101_fpn_1x_coco.py
@@ -0,0 +1,13 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.config import read_base
+from mmengine.model.weight_init import PretrainedInit
+
+with read_base():
+ from .panoptic_fpn_r50_fpn_1x_coco import *
+
+model.update(
+ dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type=PretrainedInit, checkpoint='torchvision://resnet101'))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/panoptic_fpn/panoptic_fpn_r101_fpn_ms_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/panoptic_fpn/panoptic_fpn_r101_fpn_ms_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c02c3237f81df5823ebf60a6d485365cdb655e32
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/panoptic_fpn/panoptic_fpn_r101_fpn_ms_3x_coco.py
@@ -0,0 +1,13 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.config import read_base
+from mmengine.model.weight_init import PretrainedInit
+
+with read_base():
+ from .panoptic_fpn_r50_fpn_ms_3x_coco import *
+
+model.update(
+ dict(
+ backbone=dict(
+ depth=101,
+ init_cfg=dict(
+ type=PretrainedInit, checkpoint='torchvision://resnet101'))))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/panoptic_fpn/panoptic_fpn_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/panoptic_fpn/panoptic_fpn_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..fc8932803ca0d1fd52bee7d450fc12898e0ec7b3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/panoptic_fpn/panoptic_fpn_r50_fpn_1x_coco.py
@@ -0,0 +1,64 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.models.mask_rcnn_r50_fpn import *
+ from .._base_.datasets.coco_panoptic import *
+ from .._base_.schedules.schedule_1x import *
+ from .._base_.default_runtime import *
+
+from mmcv.ops import nms
+from torch.nn import GroupNorm
+
+from mmdet.models.data_preprocessors.data_preprocessor import \
+ DetDataPreprocessor
+from mmdet.models.detectors.panoptic_fpn import PanopticFPN
+from mmdet.models.losses.cross_entropy_loss import CrossEntropyLoss
+from mmdet.models.seg_heads.panoptic_fpn_head import PanopticFPNHead
+from mmdet.models.seg_heads.panoptic_fusion_heads import HeuristicFusionHead
+
+model.update(
+ dict(
+ type=PanopticFPN,
+ data_preprocessor=dict(
+ type=DetDataPreprocessor,
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32,
+ pad_mask=True,
+ mask_pad_value=0,
+ pad_seg=True,
+ seg_pad_value=255),
+ semantic_head=dict(
+ type=PanopticFPNHead,
+ num_things_classes=80,
+ num_stuff_classes=53,
+ in_channels=256,
+ inner_channels=128,
+ start_level=0,
+ end_level=4,
+ norm_cfg=dict(type=GroupNorm, num_groups=32, requires_grad=True),
+ conv_cfg=None,
+ loss_seg=dict(
+ type=CrossEntropyLoss, ignore_index=255, loss_weight=0.5)),
+ panoptic_fusion_head=dict(
+ type=HeuristicFusionHead,
+ num_things_classes=80,
+ num_stuff_classes=53),
+ test_cfg=dict(
+ rcnn=dict(
+ score_thr=0.6,
+ nms=dict(type=nms, iou_threshold=0.5, class_agnostic=True),
+ max_per_img=100,
+ mask_thr_binary=0.5),
+ # used in HeuristicFusionHead
+ panoptic=dict(mask_overlap=0.5, stuff_area_limit=4096))))
+
+# Forced to remove NumClassCheckHook
+custom_hooks = []
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/panoptic_fpn/panoptic_fpn_r50_fpn_ms_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/panoptic_fpn/panoptic_fpn_r50_fpn_ms_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..25ebe5d67c44831b5a95978ffc0bfacec7c15de6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/panoptic_fpn/panoptic_fpn_r50_fpn_ms_3x_coco.py
@@ -0,0 +1,45 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.config import read_base
+from mmengine.optim.scheduler.lr_scheduler import LinearLR, MultiStepLR
+
+with read_base():
+ from .panoptic_fpn_r50_fpn_1x_coco import *
+
+from mmcv.transforms import RandomResize
+from mmcv.transforms.loading import LoadImageFromFile
+
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import LoadPanopticAnnotations
+from mmdet.datasets.transforms.transforms import RandomFlip
+
+# In mstrain 3x config, img_scale=[(1333, 640), (1333, 800)],
+# multiscale_mode='range'
+train_pipeline = [
+ dict(type=LoadImageFromFile),
+ dict(
+ type=LoadPanopticAnnotations,
+ with_bbox=True,
+ with_mask=True,
+ with_seg=True),
+ dict(type=RandomResize, scale=[(1333, 640), (1333, 800)], keep_ratio=True),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=PackDetInputs)
+]
+
+train_dataloader.update(dict(dataset=dict(pipeline=train_pipeline)))
+
+# TODO: Use RepeatDataset to speed up training
+# training schedule for 3x
+train_cfg.update(dict(max_epochs=36, val_interval=3))
+
+# learning rate
+param_scheduler = [
+ dict(type=LinearLR, start_factor=0.001, by_epoch=False, begin=0, end=500),
+ dict(
+ type=MultiStepLR,
+ begin=0,
+ end=36,
+ by_epoch=True,
+ milestones=[24, 33],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/qdtrack/qdtrack_faster_rcnn_r50_fpn_4e_base.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/qdtrack/qdtrack_faster_rcnn_r50_fpn_4e_base.py
new file mode 100644
index 0000000000000000000000000000000000000000..c672e82c6498092b57c389be01af64a9e26d14bc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/qdtrack/qdtrack_faster_rcnn_r50_fpn_4e_base.py
@@ -0,0 +1,141 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.models.faster_rcnn_r50_fpn import *
+ from .._base_.models.faster_rcnn_r50_fpn import model
+ from .._base_.default_runtime import *
+
+from mmcv.ops import RoIAlign
+from mmengine.hooks import LoggerHook, SyncBuffersHook
+from mmengine.model.weight_init import PretrainedInit
+from mmengine.optim import MultiStepLR, OptimWrapper
+from mmengine.runner.runner import EpochBasedTrainLoop, TestLoop, ValLoop
+from torch.nn.modules.batchnorm import BatchNorm2d
+from torch.nn.modules.normalization import GroupNorm
+from torch.optim import SGD
+
+from mmdet.engine.hooks import TrackVisualizationHook
+from mmdet.models import (QDTrack, QuasiDenseEmbedHead, QuasiDenseTracker,
+ QuasiDenseTrackHead, SingleRoIExtractor,
+ TrackDataPreprocessor)
+from mmdet.models.losses import (L1Loss, MarginL2Loss,
+ MultiPosCrossEntropyLoss, SmoothL1Loss)
+from mmdet.models.task_modules import (CombinedSampler,
+ InstanceBalancedPosSampler,
+ MaxIoUAssigner, RandomSampler)
+from mmdet.visualization import TrackLocalVisualizer
+
+detector = model
+detector.pop('data_preprocessor')
+
+detector['backbone'].update(
+ dict(
+ norm_cfg=dict(type=BatchNorm2d, requires_grad=False),
+ style='caffe',
+ init_cfg=dict(
+ type=PretrainedInit,
+ checkpoint='open-mmlab://detectron2/resnet50_caffe')))
+detector.rpn_head.loss_bbox.update(
+ dict(type=SmoothL1Loss, beta=1.0 / 9.0, loss_weight=1.0))
+detector.rpn_head.bbox_coder.update(dict(clip_border=False))
+detector.roi_head.bbox_head.update(dict(num_classes=1))
+detector.roi_head.bbox_head.bbox_coder.update(dict(clip_border=False))
+detector['init_cfg'] = dict(
+ type=PretrainedInit,
+ checkpoint= # noqa: E251
+ 'https://download.openmmlab.com/mmdetection/v2.0/faster_rcnn/'
+ 'faster_rcnn_r50_fpn_1x_coco-person/'
+ 'faster_rcnn_r50_fpn_1x_coco-person_20201216_175929-d022e227.pth'
+ # noqa: E501
+)
+del model
+
+model = dict(
+ type=QDTrack,
+ data_preprocessor=dict(
+ type=TrackDataPreprocessor,
+ mean=[103.530, 116.280, 123.675],
+ std=[1.0, 1.0, 1.0],
+ bgr_to_rgb=False,
+ pad_size_divisor=32),
+ detector=detector,
+ track_head=dict(
+ type=QuasiDenseTrackHead,
+ roi_extractor=dict(
+ type=SingleRoIExtractor,
+ roi_layer=dict(type=RoIAlign, output_size=7, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ embed_head=dict(
+ type=QuasiDenseEmbedHead,
+ num_convs=4,
+ num_fcs=1,
+ embed_channels=256,
+ norm_cfg=dict(type=GroupNorm, num_groups=32),
+ loss_track=dict(type=MultiPosCrossEntropyLoss, loss_weight=0.25),
+ loss_track_aux=dict(
+ type=MarginL2Loss,
+ neg_pos_ub=3,
+ pos_margin=0,
+ neg_margin=0.1,
+ hard_mining=True,
+ loss_weight=1.0)),
+ loss_bbox=dict(type=L1Loss, loss_weight=1.0),
+ train_cfg=dict(
+ assigner=dict(
+ type=MaxIoUAssigner,
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type=CombinedSampler,
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=3,
+ add_gt_as_proposals=True,
+ pos_sampler=dict(type=InstanceBalancedPosSampler),
+ neg_sampler=dict(type=RandomSampler)))),
+ tracker=dict(
+ type=QuasiDenseTracker,
+ init_score_thr=0.9,
+ obj_score_thr=0.5,
+ match_score_thr=0.5,
+ memo_tracklet_frames=30,
+ memo_backdrop_frames=1,
+ memo_momentum=0.8,
+ nms_conf_thr=0.5,
+ nms_backdrop_iou_thr=0.3,
+ nms_class_iou_thr=0.7,
+ with_cats=True,
+ match_metric='bisoftmax'))
+# optimizer
+optim_wrapper = dict(
+ type=OptimWrapper,
+ optimizer=dict(type=SGD, lr=0.02, momentum=0.9, weight_decay=0.0001),
+ clip_grad=dict(max_norm=35, norm_type=2))
+# learning policy
+param_scheduler = [
+ dict(type=MultiStepLR, begin=0, end=4, by_epoch=True, milestones=[3])
+]
+
+# runtime settings
+train_cfg = dict(type=EpochBasedTrainLoop, max_epochs=4, val_interval=4)
+val_cfg = dict(type=ValLoop)
+test_cfg = dict(type=TestLoop)
+
+default_hooks.update(
+ logger=dict(type=LoggerHook, interval=50),
+ visualization=dict(type=TrackVisualizationHook, draw=False))
+
+visualizer.update(
+ type=TrackLocalVisualizer, vis_backends=vis_backends, name='visualizer')
+
+# custom hooks
+custom_hooks = [
+ # Synchronize model buffers such as running_mean and running_var in BN
+ # at the end of each epoch
+ dict(type=SyncBuffersHook)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/qdtrack/qdtrack_faster_rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/qdtrack/qdtrack_faster_rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py
new file mode 100644
index 0000000000000000000000000000000000000000..2fa715e1b3806f9f9816e3b23a100167c791f0b8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/qdtrack/qdtrack_faster_rcnn_r50_fpn_8xb2-4e_mot17halftrain_test-mot17halfval.py
@@ -0,0 +1,14 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.datasets.mot_challenge import *
+ from .qdtrack_faster_rcnn_r50_fpn_4e_base import *
+
+from mmdet.evaluation import CocoVideoMetric, MOTChallengeMetric
+
+# evaluator
+val_evaluator = [
+ dict(type=CocoVideoMetric, metric=['bbox'], classwise=True),
+ dict(type=MOTChallengeMetric, metric=['HOTA', 'CLEAR', 'Identity'])
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/retinanet/retinanet_r50_fpn_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/retinanet/retinanet_r50_fpn_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..847600e61b3daf556ff24d06af2f08249deb2284
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/retinanet/retinanet_r50_fpn_1x_coco.py
@@ -0,0 +1,20 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.models.retinanet_r50_fpn import *
+ from .._base_.datasets.coco_detection import *
+ from .._base_.schedules.schedule_1x import *
+ from .._base_.default_runtime import *
+ from .retinanet_tta import *
+
+from torch.optim.sgd import SGD
+
+# optimizer
+optim_wrapper.update(
+ dict(optimizer=dict(type=SGD, lr=0.01, momentum=0.9, weight_decay=0.0001)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/retinanet/retinanet_tta.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/retinanet/retinanet_tta.py
new file mode 100644
index 0000000000000000000000000000000000000000..4e340e5854e58a332ee174b0e69e7f3f9ec2c486
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/retinanet/retinanet_tta.py
@@ -0,0 +1,31 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.transforms.loading import LoadImageFromFile
+from mmcv.transforms.processing import TestTimeAug
+
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import LoadAnnotations
+from mmdet.datasets.transforms.transforms import RandomFlip, Resize
+from mmdet.models.test_time_augs.det_tta import DetTTAModel
+
+tta_model = dict(
+ type=DetTTAModel,
+ tta_cfg=dict(nms=dict(type='nms', iou_threshold=0.5), max_per_img=100))
+
+img_scales = [(1333, 800), (666, 400), (2000, 1200)]
+tta_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=None),
+ dict(
+ type=TestTimeAug,
+ transforms=[
+ [dict(type=Resize, scale=s, keep_ratio=True) for s in img_scales],
+ [dict(type=RandomFlip, prob=1.),
+ dict(type=RandomFlip, prob=0.)],
+ [dict(type=LoadAnnotations, with_bbox=True)],
+ [
+ dict(
+ type=PackDetInputs,
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction'))
+ ]
+ ])
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_ins_l_8xb32_300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_ins_l_8xb32_300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..302d7cda110b7a598ba525549e1d96d27ee51990
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_ins_l_8xb32_300e_coco.py
@@ -0,0 +1,134 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .rtmdet_l_8xb32_300e_coco import *
+
+from mmcv.transforms.loading import LoadImageFromFile
+from mmcv.transforms.processing import RandomResize
+from mmengine.hooks.ema_hook import EMAHook
+from torch.nn.modules.activation import SiLU
+
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import (FilterAnnotations,
+ LoadAnnotations)
+from mmdet.datasets.transforms.transforms import (CachedMixUp, CachedMosaic,
+ Pad, RandomCrop, RandomFlip,
+ Resize, YOLOXHSVRandomAug)
+from mmdet.engine.hooks.pipeline_switch_hook import PipelineSwitchHook
+from mmdet.models.dense_heads.rtmdet_ins_head import RTMDetInsSepBNHead
+from mmdet.models.layers.ema import ExpMomentumEMA
+from mmdet.models.losses.dice_loss import DiceLoss
+from mmdet.models.losses.gfocal_loss import QualityFocalLoss
+from mmdet.models.losses.iou_loss import GIoULoss
+from mmdet.models.task_modules.coders.distance_point_bbox_coder import \
+ DistancePointBBoxCoder
+from mmdet.models.task_modules.prior_generators.point_generator import \
+ MlvlPointGenerator
+
+model.merge(
+ dict(
+ bbox_head=dict(
+ _delete_=True,
+ type=RTMDetInsSepBNHead,
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=2,
+ share_conv=True,
+ pred_kernel_size=1,
+ feat_channels=256,
+ act_cfg=dict(type=SiLU, inplace=True),
+ norm_cfg=dict(type='SyncBN', requires_grad=True),
+ anchor_generator=dict(
+ type=MlvlPointGenerator, offset=0, strides=[8, 16, 32]),
+ bbox_coder=dict(type=DistancePointBBoxCoder),
+ loss_cls=dict(
+ type=QualityFocalLoss,
+ use_sigmoid=True,
+ beta=2.0,
+ loss_weight=1.0),
+ loss_bbox=dict(type=GIoULoss, loss_weight=2.0),
+ loss_mask=dict(
+ type=DiceLoss, loss_weight=2.0, eps=5e-6, reduction='mean')),
+ test_cfg=dict(
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.05,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100,
+ mask_thr_binary=0.5),
+ ))
+
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(
+ type=LoadAnnotations, with_bbox=True, with_mask=True, poly2mask=False),
+ dict(type=CachedMosaic, img_scale=(640, 640), pad_val=114.0),
+ dict(
+ type=RandomResize,
+ scale=(1280, 1280),
+ ratio_range=(0.1, 2.0),
+ resize_type=Resize,
+ keep_ratio=True),
+ dict(
+ type=RandomCrop,
+ crop_size=(640, 640),
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type=YOLOXHSVRandomAug),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=Pad, size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(
+ type=CachedMixUp,
+ img_scale=(640, 640),
+ ratio_range=(1.0, 1.0),
+ max_cached_images=20,
+ pad_val=(114, 114, 114)),
+ dict(type=FilterAnnotations, min_gt_bbox_wh=(1, 1)),
+ dict(type=PackDetInputs)
+]
+
+train_dataloader.update(
+ dict(pin_memory=True, dataset=dict(pipeline=train_pipeline)))
+
+train_pipeline_stage2 = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(
+ type=LoadAnnotations, with_bbox=True, with_mask=True, poly2mask=False),
+ dict(
+ type=RandomResize,
+ scale=(640, 640),
+ ratio_range=(0.1, 2.0),
+ resize_type=Resize,
+ keep_ratio=True),
+ dict(
+ type=RandomCrop,
+ crop_size=(640, 640),
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type=FilterAnnotations, min_gt_bbox_wh=(1, 1)),
+ dict(type=YOLOXHSVRandomAug),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=Pad, size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(type=PackDetInputs)
+]
+custom_hooks = [
+ dict(
+ type=EMAHook,
+ ema_type=ExpMomentumEMA,
+ momentum=0.0002,
+ update_buffers=True,
+ priority=49),
+ dict(
+ type=PipelineSwitchHook,
+ switch_epoch=280,
+ switch_pipeline=train_pipeline_stage2)
+]
+
+val_evaluator.update(dict(metric=['bbox', 'segm']))
+test_evaluator = val_evaluator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_ins_m_8xb32_300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_ins_m_8xb32_300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d90be9293a18cfd703ee1d9993b03237fb3c3dab
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_ins_m_8xb32_300e_coco.py
@@ -0,0 +1,17 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .rtmdet_ins_l_8xb32_300e_coco import *
+
+model.update(
+ dict(
+ backbone=dict(deepen_factor=0.67, widen_factor=0.75),
+ neck=dict(
+ in_channels=[192, 384, 768], out_channels=192, num_csp_blocks=2),
+ bbox_head=dict(in_channels=192, feat_channels=192)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_ins_s_8xb32_300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_ins_s_8xb32_300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..58b5b1aff0cff8d770798288b74237bc5183d37b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_ins_s_8xb32_300e_coco.py
@@ -0,0 +1,101 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .rtmdet_ins_l_8xb32_300e_coco import *
+
+from mmcv.transforms.loading import LoadImageFromFile
+from mmcv.transforms.processing import RandomResize
+from mmengine.hooks.ema_hook import EMAHook
+
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import (FilterAnnotations,
+ LoadAnnotations)
+from mmdet.datasets.transforms.transforms import (CachedMixUp, CachedMosaic,
+ Pad, RandomCrop, RandomFlip,
+ Resize, YOLOXHSVRandomAug)
+from mmdet.engine.hooks.pipeline_switch_hook import PipelineSwitchHook
+from mmdet.models.layers.ema import ExpMomentumEMA
+
+checkpoint = 'https://download.openmmlab.com/mmdetection/v3.0/rtmdet/cspnext_rsb_pretrain/cspnext-s_imagenet_600e.pth' # noqa
+model.update(
+ dict(
+ backbone=dict(
+ deepen_factor=0.33,
+ widen_factor=0.5,
+ init_cfg=dict(
+ type='Pretrained', prefix='backbone.', checkpoint=checkpoint)),
+ neck=dict(
+ in_channels=[128, 256, 512], out_channels=128, num_csp_blocks=1),
+ bbox_head=dict(in_channels=128, feat_channels=128)))
+
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(
+ type=LoadAnnotations, with_bbox=True, with_mask=True, poly2mask=False),
+ dict(type=CachedMosaic, img_scale=(640, 640), pad_val=114.0),
+ dict(
+ type=RandomResize,
+ scale=(1280, 1280),
+ ratio_range=(0.5, 2.0),
+ resize_type=Resize,
+ keep_ratio=True),
+ dict(
+ type=RandomCrop,
+ crop_size=(640, 640),
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type=YOLOXHSVRandomAug),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=Pad, size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(
+ type=CachedMixUp,
+ img_scale=(640, 640),
+ ratio_range=(1.0, 1.0),
+ max_cached_images=20,
+ pad_val=(114, 114, 114)),
+ dict(type=FilterAnnotations, min_gt_bbox_wh=(1, 1)),
+ dict(type=PackDetInputs)
+]
+
+train_pipeline_stage2 = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(
+ type=LoadAnnotations, with_bbox=True, with_mask=True, poly2mask=False),
+ dict(
+ type=RandomResize,
+ scale=(640, 640),
+ ratio_range=(0.5, 2.0),
+ resize_type=Resize,
+ keep_ratio=True),
+ dict(
+ type=RandomCrop,
+ crop_size=(640, 640),
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type=FilterAnnotations, min_gt_bbox_wh=(1, 1)),
+ dict(type=YOLOXHSVRandomAug),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=Pad, size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(type=PackDetInputs)
+]
+
+train_dataloader.update(dict(dataset=dict(pipeline=train_pipeline)))
+
+custom_hooks = [
+ dict(
+ type=EMAHook,
+ ema_type=ExpMomentumEMA,
+ momentum=0.0002,
+ update_buffers=True,
+ priority=49),
+ dict(
+ type=PipelineSwitchHook,
+ switch_epoch=280,
+ switch_pipeline=train_pipeline_stage2)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_ins_tiny_8xb32_300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_ins_tiny_8xb32_300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..0356b1951da584034cf65014a39c7440fc3da56d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_ins_tiny_8xb32_300e_coco.py
@@ -0,0 +1,67 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .rtmdet_ins_s_8xb32_300e_coco import *
+
+from mmcv.transforms.loading import LoadImageFromFile
+from mmcv.transforms.processing import RandomResize
+
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import (FilterAnnotations,
+ LoadAnnotations)
+from mmdet.datasets.transforms.transforms import (CachedMixUp, CachedMosaic,
+ Pad, RandomCrop, RandomFlip,
+ Resize, YOLOXHSVRandomAug)
+
+checkpoint = 'https://download.openmmlab.com/mmdetection/v3.0/rtmdet/cspnext_rsb_pretrain/cspnext-tiny_imagenet_600e.pth' # noqa
+
+model.update(
+ dict(
+ backbone=dict(
+ deepen_factor=0.167,
+ widen_factor=0.375,
+ init_cfg=dict(
+ type='Pretrained', prefix='backbone.', checkpoint=checkpoint)),
+ neck=dict(
+ in_channels=[96, 192, 384], out_channels=96, num_csp_blocks=1),
+ bbox_head=dict(in_channels=96, feat_channels=96)))
+
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(
+ type=LoadAnnotations, with_bbox=True, with_mask=True, poly2mask=False),
+ dict(
+ type=CachedMosaic,
+ img_scale=(640, 640),
+ pad_val=114.0,
+ max_cached_images=20,
+ random_pop=False),
+ dict(
+ type=RandomResize,
+ scale=(1280, 1280),
+ ratio_range=(0.5, 2.0),
+ resize_type=Resize,
+ keep_ratio=True),
+ dict(type=RandomCrop, crop_size=(640, 640)),
+ dict(type=YOLOXHSVRandomAug),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=Pad, size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(
+ type=CachedMixUp,
+ img_scale=(640, 640),
+ ratio_range=(1.0, 1.0),
+ max_cached_images=10,
+ random_pop=False,
+ pad_val=(114, 114, 114),
+ prob=0.5),
+ dict(type=FilterAnnotations, min_gt_bbox_wh=(1, 1)),
+ dict(type=PackDetInputs)
+]
+
+train_dataloader.update(dict(dataset=dict(pipeline=train_pipeline)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_ins_x_8xb16_300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_ins_x_8xb16_300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..555b10102f67ee625d65dbfe0894eb4b41198595
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_ins_x_8xb16_300e_coco.py
@@ -0,0 +1,38 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .rtmdet_ins_l_8xb32_300e_coco import *
+from mmengine.optim.scheduler.lr_scheduler import CosineAnnealingLR, LinearLR
+
+model.update(
+ dict(
+ backbone=dict(deepen_factor=1.33, widen_factor=1.25),
+ neck=dict(
+ in_channels=[320, 640, 1280], out_channels=320, num_csp_blocks=4),
+ bbox_head=dict(in_channels=320, feat_channels=320)))
+
+base_lr = 0.002
+
+# optimizer
+optim_wrapper.update(dict(optimizer=dict(lr=base_lr)))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type=LinearLR, start_factor=1.0e-5, by_epoch=False, begin=0, end=1000),
+ dict(
+ # use cosine lr from 150 to 300 epoch
+ type=CosineAnnealingLR,
+ eta_min=base_lr * 0.05,
+ begin=max_epochs // 2,
+ end=max_epochs,
+ T_max=max_epochs // 2,
+ by_epoch=True,
+ convert_to_iter_based=True),
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_l_8xb32_300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_l_8xb32_300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..5dcda7bf994db9f3f5c785d8dea824b3ab8e56a2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_l_8xb32_300e_coco.py
@@ -0,0 +1,220 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .._base_.default_runtime import *
+ from .._base_.schedules.schedule_1x import *
+ from .._base_.datasets.coco_detection import *
+ from .rtmdet_tta import *
+
+from mmcv.ops import nms
+from mmcv.transforms.loading import LoadImageFromFile
+from mmcv.transforms.processing import RandomResize
+from mmengine.hooks.ema_hook import EMAHook
+from mmengine.optim.optimizer.optimizer_wrapper import OptimWrapper
+from mmengine.optim.scheduler.lr_scheduler import CosineAnnealingLR, LinearLR
+from torch.nn import SyncBatchNorm
+from torch.nn.modules.activation import SiLU
+from torch.optim.adamw import AdamW
+
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import LoadAnnotations
+from mmdet.datasets.transforms.transforms import (CachedMixUp, CachedMosaic,
+ Pad, RandomCrop, RandomFlip,
+ Resize, YOLOXHSVRandomAug)
+from mmdet.engine.hooks.pipeline_switch_hook import PipelineSwitchHook
+from mmdet.models.backbones.cspnext import CSPNeXt
+from mmdet.models.data_preprocessors.data_preprocessor import \
+ DetDataPreprocessor
+from mmdet.models.dense_heads.rtmdet_head import RTMDetSepBNHead
+from mmdet.models.detectors.rtmdet import RTMDet
+from mmdet.models.layers.ema import ExpMomentumEMA
+from mmdet.models.losses.gfocal_loss import QualityFocalLoss
+from mmdet.models.losses.iou_loss import GIoULoss
+from mmdet.models.necks.cspnext_pafpn import CSPNeXtPAFPN
+from mmdet.models.task_modules.assigners.dynamic_soft_label_assigner import \
+ DynamicSoftLabelAssigner
+from mmdet.models.task_modules.coders.distance_point_bbox_coder import \
+ DistancePointBBoxCoder
+from mmdet.models.task_modules.prior_generators.point_generator import \
+ MlvlPointGenerator
+
+model = dict(
+ type=RTMDet,
+ data_preprocessor=dict(
+ type=DetDataPreprocessor,
+ mean=[103.53, 116.28, 123.675],
+ std=[57.375, 57.12, 58.395],
+ bgr_to_rgb=False,
+ batch_augments=None),
+ backbone=dict(
+ type=CSPNeXt,
+ arch='P5',
+ expand_ratio=0.5,
+ deepen_factor=1,
+ widen_factor=1,
+ channel_attention=True,
+ norm_cfg=dict(type=SyncBatchNorm),
+ act_cfg=dict(type=SiLU, inplace=True)),
+ neck=dict(
+ type=CSPNeXtPAFPN,
+ in_channels=[256, 512, 1024],
+ out_channels=256,
+ num_csp_blocks=3,
+ expand_ratio=0.5,
+ norm_cfg=dict(type=SyncBatchNorm),
+ act_cfg=dict(type=SiLU, inplace=True)),
+ bbox_head=dict(
+ type=RTMDetSepBNHead,
+ num_classes=80,
+ in_channels=256,
+ stacked_convs=2,
+ feat_channels=256,
+ anchor_generator=dict(
+ type=MlvlPointGenerator, offset=0, strides=[8, 16, 32]),
+ bbox_coder=dict(type=DistancePointBBoxCoder),
+ loss_cls=dict(
+ type=QualityFocalLoss, use_sigmoid=True, beta=2.0,
+ loss_weight=1.0),
+ loss_bbox=dict(type=GIoULoss, loss_weight=2.0),
+ with_objectness=False,
+ exp_on_reg=True,
+ share_conv=True,
+ pred_kernel_size=1,
+ norm_cfg=dict(type=SyncBatchNorm),
+ act_cfg=dict(type=SiLU, inplace=True)),
+ train_cfg=dict(
+ assigner=dict(type=DynamicSoftLabelAssigner, topk=13),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ test_cfg=dict(
+ nms_pre=30000,
+ min_bbox_size=0,
+ score_thr=0.001,
+ nms=dict(type=nms, iou_threshold=0.65),
+ max_per_img=300),
+)
+
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=LoadAnnotations, with_bbox=True),
+ dict(type=CachedMosaic, img_scale=(640, 640), pad_val=114.0),
+ dict(
+ type=RandomResize,
+ scale=(1280, 1280),
+ ratio_range=(0.1, 2.0),
+ resize_type=Resize,
+ keep_ratio=True),
+ dict(type=RandomCrop, crop_size=(640, 640)),
+ dict(type=YOLOXHSVRandomAug),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=Pad, size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(
+ type=CachedMixUp,
+ img_scale=(640, 640),
+ ratio_range=(1.0, 1.0),
+ max_cached_images=20,
+ pad_val=(114, 114, 114)),
+ dict(type=PackDetInputs)
+]
+
+train_pipeline_stage2 = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=LoadAnnotations, with_bbox=True),
+ dict(
+ type=RandomResize,
+ scale=(640, 640),
+ ratio_range=(0.1, 2.0),
+ resize_type=Resize,
+ keep_ratio=True),
+ dict(type=RandomCrop, crop_size=(640, 640)),
+ dict(type=YOLOXHSVRandomAug),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=Pad, size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(type=PackDetInputs)
+]
+
+test_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=Resize, scale=(640, 640), keep_ratio=True),
+ dict(type=Pad, size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(type=LoadAnnotations, with_bbox=True),
+ dict(
+ type=PackDetInputs,
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader.update(
+ dict(
+ batch_size=32,
+ num_workers=10,
+ batch_sampler=None,
+ pin_memory=True,
+ dataset=dict(pipeline=train_pipeline)))
+val_dataloader.update(
+ dict(batch_size=5, num_workers=10, dataset=dict(pipeline=test_pipeline)))
+test_dataloader = val_dataloader
+
+max_epochs = 300
+stage2_num_epochs = 20
+base_lr = 0.004
+interval = 10
+
+train_cfg.update(
+ dict(
+ max_epochs=max_epochs,
+ val_interval=interval,
+ dynamic_intervals=[(max_epochs - stage2_num_epochs, 1)]))
+
+val_evaluator.update(dict(proposal_nums=(100, 1, 10)))
+test_evaluator = val_evaluator
+
+# optimizer
+optim_wrapper = dict(
+ type=OptimWrapper,
+ optimizer=dict(type=AdamW, lr=base_lr, weight_decay=0.05),
+ paramwise_cfg=dict(
+ norm_decay_mult=0, bias_decay_mult=0, bypass_duplicate=True))
+
+# learning rate
+param_scheduler = [
+ dict(
+ type=LinearLR, start_factor=1.0e-5, by_epoch=False, begin=0, end=1000),
+ dict(
+ # use cosine lr from 150 to 300 epoch
+ type=CosineAnnealingLR,
+ eta_min=base_lr * 0.05,
+ begin=max_epochs // 2,
+ end=max_epochs,
+ T_max=max_epochs // 2,
+ by_epoch=True,
+ convert_to_iter_based=True),
+]
+
+# hooks
+default_hooks.update(
+ dict(
+ checkpoint=dict(
+ interval=interval,
+ max_keep_ckpts=3 # only keep latest 3 checkpoints
+ )))
+
+custom_hooks = [
+ dict(
+ type=EMAHook,
+ ema_type=ExpMomentumEMA,
+ momentum=0.0002,
+ update_buffers=True,
+ priority=49),
+ dict(
+ type=PipelineSwitchHook,
+ switch_epoch=max_epochs - stage2_num_epochs,
+ switch_pipeline=train_pipeline_stage2)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_m_8xb32_300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_m_8xb32_300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..e741d8220fe8831894b7b803060031e18dbac62b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_m_8xb32_300e_coco.py
@@ -0,0 +1,17 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .rtmdet_l_8xb32_300e_coco import *
+
+model.update(
+ dict(
+ backbone=dict(deepen_factor=0.67, widen_factor=0.75),
+ neck=dict(
+ in_channels=[192, 384, 768], out_channels=192, num_csp_blocks=2),
+ bbox_head=dict(in_channels=192, feat_channels=192)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_s_8xb32_300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_s_8xb32_300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..db21b747e95a15c69af1c17c16a5e6cfd4a2be78
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_s_8xb32_300e_coco.py
@@ -0,0 +1,88 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .rtmdet_l_8xb32_300e_coco import *
+
+from mmcv.transforms.loading import LoadImageFromFile
+from mmcv.transforms.processing import RandomResize
+from mmengine.hooks.ema_hook import EMAHook
+
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import LoadAnnotations
+from mmdet.datasets.transforms.transforms import (CachedMixUp, CachedMosaic,
+ Pad, RandomCrop, RandomFlip,
+ Resize, YOLOXHSVRandomAug)
+from mmdet.engine.hooks.pipeline_switch_hook import PipelineSwitchHook
+from mmdet.models.layers.ema import ExpMomentumEMA
+
+checkpoint = 'https://download.openmmlab.com/mmdetection/v3.0/rtmdet/cspnext_rsb_pretrain/cspnext-s_imagenet_600e.pth' # noqa
+model.update(
+ dict(
+ backbone=dict(
+ deepen_factor=0.33,
+ widen_factor=0.5,
+ init_cfg=dict(
+ type='Pretrained', prefix='backbone.', checkpoint=checkpoint)),
+ neck=dict(
+ in_channels=[128, 256, 512], out_channels=128, num_csp_blocks=1),
+ bbox_head=dict(in_channels=128, feat_channels=128, exp_on_reg=False)))
+
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=LoadAnnotations, with_bbox=True),
+ dict(type=CachedMosaic, img_scale=(640, 640), pad_val=114.0),
+ dict(
+ type=RandomResize,
+ scale=(1280, 1280),
+ ratio_range=(0.5, 2.0),
+ resize_type=Resize,
+ keep_ratio=True),
+ dict(type=RandomCrop, crop_size=(640, 640)),
+ dict(type=YOLOXHSVRandomAug),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=Pad, size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(
+ type=CachedMixUp,
+ img_scale=(640, 640),
+ ratio_range=(1.0, 1.0),
+ max_cached_images=20,
+ pad_val=(114, 114, 114)),
+ dict(type=PackDetInputs)
+]
+
+train_pipeline_stage2 = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=LoadAnnotations, with_bbox=True),
+ dict(
+ type=RandomResize,
+ scale=(640, 640),
+ ratio_range=(0.5, 2.0),
+ resize_type=Resize,
+ keep_ratio=True),
+ dict(type=RandomCrop, crop_size=(640, 640)),
+ dict(type=YOLOXHSVRandomAug),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=Pad, size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(type=PackDetInputs)
+]
+
+train_dataloader.update(dict(dataset=dict(pipeline=train_pipeline)))
+
+custom_hooks = [
+ dict(
+ type=EMAHook,
+ ema_type=ExpMomentumEMA,
+ momentum=0.0002,
+ update_buffers=True,
+ priority=49),
+ dict(
+ type=PipelineSwitchHook,
+ switch_epoch=280,
+ switch_pipeline=train_pipeline_stage2)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_tiny_8xb32_300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_tiny_8xb32_300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..949d056f16303751d121aeba8f3d859de07b06d2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_tiny_8xb32_300e_coco.py
@@ -0,0 +1,64 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .rtmdet_s_8xb32_300e_coco import *
+
+from mmcv.transforms.loading import LoadImageFromFile
+from mmcv.transforms.processing import RandomResize
+
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import LoadAnnotations
+from mmdet.datasets.transforms.transforms import (CachedMixUp, CachedMosaic,
+ Pad, RandomCrop, RandomFlip,
+ Resize, YOLOXHSVRandomAug)
+
+checkpoint = 'https://download.openmmlab.com/mmdetection/v3.0/rtmdet/cspnext_rsb_pretrain/cspnext-tiny_imagenet_600e.pth' # noqa
+
+model.update(
+ dict(
+ backbone=dict(
+ deepen_factor=0.167,
+ widen_factor=0.375,
+ init_cfg=dict(
+ type='Pretrained', prefix='backbone.', checkpoint=checkpoint)),
+ neck=dict(
+ in_channels=[96, 192, 384], out_channels=96, num_csp_blocks=1),
+ bbox_head=dict(in_channels=96, feat_channels=96, exp_on_reg=False)))
+
+train_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=backend_args),
+ dict(type=LoadAnnotations, with_bbox=True),
+ dict(
+ type=CachedMosaic,
+ img_scale=(640, 640),
+ pad_val=114.0,
+ max_cached_images=20,
+ random_pop=False),
+ dict(
+ type=RandomResize,
+ scale=(1280, 1280),
+ ratio_range=(0.5, 2.0),
+ resize_type=Resize,
+ keep_ratio=True),
+ dict(type=RandomCrop, crop_size=(640, 640)),
+ dict(type=YOLOXHSVRandomAug),
+ dict(type=RandomFlip, prob=0.5),
+ dict(type=Pad, size=(640, 640), pad_val=dict(img=(114, 114, 114))),
+ dict(
+ type=CachedMixUp,
+ img_scale=(640, 640),
+ ratio_range=(1.0, 1.0),
+ max_cached_images=10,
+ random_pop=False,
+ pad_val=(114, 114, 114),
+ prob=0.5),
+ dict(type=PackDetInputs)
+]
+
+train_dataloader.update(dict(dataset=dict(pipeline=train_pipeline)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_tta.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_tta.py
new file mode 100644
index 0000000000000000000000000000000000000000..f27b7aa4a3bf13a28cab3e25be755a9792620ece
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_tta.py
@@ -0,0 +1,43 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.transforms.loading import LoadImageFromFile
+from mmcv.transforms.processing import TestTimeAug
+
+from mmdet.datasets.transforms.formatting import PackDetInputs
+from mmdet.datasets.transforms.loading import LoadAnnotations
+from mmdet.datasets.transforms.transforms import Pad, RandomFlip, Resize
+from mmdet.models.test_time_augs.det_tta import DetTTAModel
+
+tta_model = dict(
+ type=DetTTAModel,
+ tta_cfg=dict(nms=dict(type='nms', iou_threshold=0.6), max_per_img=100))
+
+img_scales = [(640, 640), (320, 320), (960, 960)]
+
+tta_pipeline = [
+ dict(type=LoadImageFromFile, backend_args=None),
+ dict(
+ type=TestTimeAug,
+ transforms=[
+ [dict(type=Resize, scale=s, keep_ratio=True) for s in img_scales],
+ [
+ # ``RandomFlip`` must be placed before ``Pad``, otherwise
+ # bounding box coordinates after flipping cannot be
+ # recovered correctly.
+ dict(type=RandomFlip, prob=1.),
+ dict(type=RandomFlip, prob=0.)
+ ],
+ [
+ dict(
+ type=Pad,
+ size=(960, 960),
+ pad_val=dict(img=(114, 114, 114))),
+ ],
+ [dict(type=LoadAnnotations, with_bbox=True)],
+ [
+ dict(
+ type=PackDetInputs,
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction'))
+ ]
+ ])
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_x_8xb32_300e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_x_8xb32_300e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..04d67d0ca8f08860462eb0eafc645c403e792394
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/configs/rtmdet/rtmdet_x_8xb32_300e_coco.py
@@ -0,0 +1,17 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Please refer to https://mmengine.readthedocs.io/en/latest/advanced_tutorials/config.html#a-pure-python-style-configuration-file-beta for more details. # noqa
+# mmcv >= 2.0.1
+# mmengine >= 0.8.0
+
+from mmengine.config import read_base
+
+with read_base():
+ from .rtmdet_l_8xb32_300e_coco import *
+
+model.update(
+ dict(
+ backbone=dict(deepen_factor=1.33, widen_factor=1.25),
+ neck=dict(
+ in_channels=[320, 640, 1280], out_channels=320, num_csp_blocks=4),
+ bbox_head=dict(in_channels=320, feat_channels=320)))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..670c207cacf9ed0f9fee88bada119ee3aaa85eae
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/__init__.py
@@ -0,0 +1,53 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .ade20k import (ADE20KInstanceDataset, ADE20KPanopticDataset,
+ ADE20KSegDataset)
+from .base_det_dataset import BaseDetDataset
+from .base_semseg_dataset import BaseSegDataset
+from .base_video_dataset import BaseVideoDataset
+from .cityscapes import CityscapesDataset
+from .coco import CocoDataset
+from .coco_caption import CocoCaptionDataset
+from .coco_panoptic import CocoPanopticDataset
+from .coco_semantic import CocoSegDataset
+from .crowdhuman import CrowdHumanDataset
+from .dataset_wrappers import ConcatDataset, MultiImageMixDataset
+from .deepfashion import DeepFashionDataset
+from .dod import DODDataset
+from .dsdl import DSDLDetDataset
+from .flickr30k import Flickr30kDataset
+from .isaid import iSAIDDataset
+from .lvis import LVISDataset, LVISV1Dataset, LVISV05Dataset
+from .mdetr_style_refcoco import MDETRStyleRefCocoDataset
+from .mot_challenge_dataset import MOTChallengeDataset
+from .objects365 import Objects365V1Dataset, Objects365V2Dataset
+from .odvg import ODVGDataset
+from .openimages import OpenImagesChallengeDataset, OpenImagesDataset
+from .refcoco import RefCocoDataset
+from .reid_dataset import ReIDDataset
+from .samplers import (AspectRatioBatchSampler, ClassAwareSampler,
+ CustomSampleSizeSampler, GroupMultiSourceSampler,
+ MultiSourceSampler, TrackAspectRatioBatchSampler,
+ TrackImgSampler)
+from .utils import get_loading_pipeline
+from .v3det import V3DetDataset
+from .voc import VOCDataset
+from .wider_face import WIDERFaceDataset
+from .xml_style import XMLDataset
+from .youtube_vis_dataset import YouTubeVISDataset
+
+__all__ = [
+ 'XMLDataset', 'CocoDataset', 'DeepFashionDataset', 'VOCDataset',
+ 'CityscapesDataset', 'LVISDataset', 'LVISV05Dataset', 'LVISV1Dataset',
+ 'WIDERFaceDataset', 'get_loading_pipeline', 'CocoPanopticDataset',
+ 'MultiImageMixDataset', 'OpenImagesDataset', 'OpenImagesChallengeDataset',
+ 'AspectRatioBatchSampler', 'ClassAwareSampler', 'MultiSourceSampler',
+ 'GroupMultiSourceSampler', 'BaseDetDataset', 'CrowdHumanDataset',
+ 'Objects365V1Dataset', 'Objects365V2Dataset', 'DSDLDetDataset',
+ 'BaseVideoDataset', 'MOTChallengeDataset', 'TrackImgSampler',
+ 'ReIDDataset', 'YouTubeVISDataset', 'TrackAspectRatioBatchSampler',
+ 'ADE20KPanopticDataset', 'CocoCaptionDataset', 'RefCocoDataset',
+ 'BaseSegDataset', 'ADE20KSegDataset', 'CocoSegDataset',
+ 'ADE20KInstanceDataset', 'iSAIDDataset', 'V3DetDataset', 'ConcatDataset',
+ 'ODVGDataset', 'MDETRStyleRefCocoDataset', 'DODDataset',
+ 'CustomSampleSizeSampler', 'Flickr30kDataset'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/ade20k.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/ade20k.py
new file mode 100644
index 0000000000000000000000000000000000000000..573271cb5d0cb83571564272895bddde9a5f6ad7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/ade20k.py
@@ -0,0 +1,260 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import os.path as osp
+from typing import List
+
+from mmengine import fileio
+
+from mmdet.registry import DATASETS
+from .base_semseg_dataset import BaseSegDataset
+from .coco import CocoDataset
+from .coco_panoptic import CocoPanopticDataset
+
+ADE_PALETTE = [(120, 120, 120), (180, 120, 120), (6, 230, 230), (80, 50, 50),
+ (4, 200, 3), (120, 120, 80), (140, 140, 140), (204, 5, 255),
+ (230, 230, 230), (4, 250, 7), (224, 5, 255), (235, 255, 7),
+ (150, 5, 61), (120, 120, 70), (8, 255, 51), (255, 6, 82),
+ (143, 255, 140), (204, 255, 4), (255, 51, 7), (204, 70, 3),
+ (0, 102, 200), (61, 230, 250), (255, 6, 51), (11, 102, 255),
+ (255, 7, 71), (255, 9, 224), (9, 7, 230), (220, 220, 220),
+ (255, 9, 92), (112, 9, 255), (8, 255, 214), (7, 255, 224),
+ (255, 184, 6), (10, 255, 71), (255, 41, 10), (7, 255, 255),
+ (224, 255, 8), (102, 8, 255), (255, 61, 6), (255, 194, 7),
+ (255, 122, 8), (0, 255, 20), (255, 8, 41), (255, 5, 153),
+ (6, 51, 255), (235, 12, 255), (160, 150, 20), (0, 163, 255),
+ (140, 140, 140), (250, 10, 15), (20, 255, 0), (31, 255, 0),
+ (255, 31, 0), (255, 224, 0), (153, 255, 0), (0, 0, 255),
+ (255, 71, 0), (0, 235, 255), (0, 173, 255), (31, 0, 255),
+ (11, 200, 200), (255, 82, 0), (0, 255, 245), (0, 61, 255),
+ (0, 255, 112), (0, 255, 133), (255, 0, 0), (255, 163, 0),
+ (255, 102, 0), (194, 255, 0), (0, 143, 255), (51, 255, 0),
+ (0, 82, 255), (0, 255, 41), (0, 255, 173), (10, 0, 255),
+ (173, 255, 0), (0, 255, 153), (255, 92, 0), (255, 0, 255),
+ (255, 0, 245), (255, 0, 102), (255, 173, 0), (255, 0, 20),
+ (255, 184, 184), (0, 31, 255), (0, 255, 61), (0, 71, 255),
+ (255, 0, 204), (0, 255, 194), (0, 255, 82), (0, 10, 255),
+ (0, 112, 255), (51, 0, 255), (0, 194, 255), (0, 122, 255),
+ (0, 255, 163), (255, 153, 0), (0, 255, 10), (255, 112, 0),
+ (143, 255, 0), (82, 0, 255), (163, 255, 0), (255, 235, 0),
+ (8, 184, 170), (133, 0, 255), (0, 255, 92), (184, 0, 255),
+ (255, 0, 31), (0, 184, 255), (0, 214, 255), (255, 0, 112),
+ (92, 255, 0), (0, 224, 255), (112, 224, 255), (70, 184, 160),
+ (163, 0, 255), (153, 0, 255), (71, 255, 0), (255, 0, 163),
+ (255, 204, 0), (255, 0, 143), (0, 255, 235), (133, 255, 0),
+ (255, 0, 235), (245, 0, 255), (255, 0, 122), (255, 245, 0),
+ (10, 190, 212), (214, 255, 0), (0, 204, 255), (20, 0, 255),
+ (255, 255, 0), (0, 153, 255), (0, 41, 255), (0, 255, 204),
+ (41, 0, 255), (41, 255, 0), (173, 0, 255), (0, 245, 255),
+ (71, 0, 255), (122, 0, 255), (0, 255, 184), (0, 92, 255),
+ (184, 255, 0), (0, 133, 255), (255, 214, 0), (25, 194, 194),
+ (102, 255, 0), (92, 0, 255)]
+
+
+@DATASETS.register_module()
+class ADE20KPanopticDataset(CocoPanopticDataset):
+ METAINFO = {
+ 'classes':
+ ('bed', 'window', 'cabinet', 'person', 'door', 'table', 'curtain',
+ 'chair', 'car', 'painting, picture', 'sofa', 'shelf', 'mirror',
+ 'armchair', 'seat', 'fence', 'desk', 'wardrobe, closet, press',
+ 'lamp', 'tub', 'rail', 'cushion', 'box', 'column, pillar',
+ 'signboard, sign', 'chest of drawers, chest, bureau, dresser',
+ 'counter', 'sink', 'fireplace', 'refrigerator, icebox', 'stairs',
+ 'case, display case, showcase, vitrine',
+ 'pool table, billiard table, snooker table', 'pillow',
+ 'screen door, screen', 'bookcase', 'coffee table',
+ 'toilet, can, commode, crapper, pot, potty, stool, throne', 'flower',
+ 'book', 'bench', 'countertop', 'stove', 'palm, palm tree',
+ 'kitchen island', 'computer', 'swivel chair', 'boat',
+ 'arcade machine', 'bus', 'towel', 'light', 'truck', 'chandelier',
+ 'awning, sunshade, sunblind', 'street lamp', 'booth', 'tv',
+ 'airplane', 'clothes', 'pole',
+ 'bannister, banister, balustrade, balusters, handrail',
+ 'ottoman, pouf, pouffe, puff, hassock', 'bottle', 'van', 'ship',
+ 'fountain', 'washer, automatic washer, washing machine',
+ 'plaything, toy', 'stool', 'barrel, cask', 'basket, handbasket',
+ 'bag', 'minibike, motorbike', 'oven', 'ball', 'food, solid food',
+ 'step, stair', 'trade name', 'microwave', 'pot', 'animal', 'bicycle',
+ 'dishwasher', 'screen', 'sculpture', 'hood, exhaust hood', 'sconce',
+ 'vase', 'traffic light', 'tray', 'trash can', 'fan', 'plate',
+ 'monitor', 'bulletin board', 'radiator', 'glass, drinking glass',
+ 'clock', 'flag', 'wall', 'building', 'sky', 'floor', 'tree',
+ 'ceiling', 'road, route', 'grass', 'sidewalk, pavement',
+ 'earth, ground', 'mountain, mount', 'plant', 'water', 'house', 'sea',
+ 'rug', 'field', 'rock, stone', 'base, pedestal, stand', 'sand',
+ 'skyscraper', 'grandstand, covered stand', 'path', 'runway',
+ 'stairway, staircase', 'river', 'bridge, span', 'blind, screen',
+ 'hill', 'bar', 'hovel, hut, hutch, shack, shanty', 'tower',
+ 'dirt track', 'land, ground, soil',
+ 'escalator, moving staircase, moving stairway',
+ 'buffet, counter, sideboard',
+ 'poster, posting, placard, notice, bill, card', 'stage',
+ 'conveyer belt, conveyor belt, conveyer, conveyor, transporter',
+ 'canopy', 'pool', 'falls', 'tent', 'cradle', 'tank, storage tank',
+ 'lake', 'blanket, cover', 'pier', 'crt screen', 'shower'),
+ 'thing_classes':
+ ('bed', 'window', 'cabinet', 'person', 'door', 'table', 'curtain',
+ 'chair', 'car', 'painting, picture', 'sofa', 'shelf', 'mirror',
+ 'armchair', 'seat', 'fence', 'desk', 'wardrobe, closet, press',
+ 'lamp', 'tub', 'rail', 'cushion', 'box', 'column, pillar',
+ 'signboard, sign', 'chest of drawers, chest, bureau, dresser',
+ 'counter', 'sink', 'fireplace', 'refrigerator, icebox', 'stairs',
+ 'case, display case, showcase, vitrine',
+ 'pool table, billiard table, snooker table', 'pillow',
+ 'screen door, screen', 'bookcase', 'coffee table',
+ 'toilet, can, commode, crapper, pot, potty, stool, throne', 'flower',
+ 'book', 'bench', 'countertop', 'stove', 'palm, palm tree',
+ 'kitchen island', 'computer', 'swivel chair', 'boat',
+ 'arcade machine', 'bus', 'towel', 'light', 'truck', 'chandelier',
+ 'awning, sunshade, sunblind', 'street lamp', 'booth', 'tv',
+ 'airplane', 'clothes', 'pole',
+ 'bannister, banister, balustrade, balusters, handrail',
+ 'ottoman, pouf, pouffe, puff, hassock', 'bottle', 'van', 'ship',
+ 'fountain', 'washer, automatic washer, washing machine',
+ 'plaything, toy', 'stool', 'barrel, cask', 'basket, handbasket',
+ 'bag', 'minibike, motorbike', 'oven', 'ball', 'food, solid food',
+ 'step, stair', 'trade name', 'microwave', 'pot', 'animal', 'bicycle',
+ 'dishwasher', 'screen', 'sculpture', 'hood, exhaust hood', 'sconce',
+ 'vase', 'traffic light', 'tray', 'trash can', 'fan', 'plate',
+ 'monitor', 'bulletin board', 'radiator', 'glass, drinking glass',
+ 'clock', 'flag'),
+ 'stuff_classes':
+ ('wall', 'building', 'sky', 'floor', 'tree', 'ceiling', 'road, route',
+ 'grass', 'sidewalk, pavement', 'earth, ground', 'mountain, mount',
+ 'plant', 'water', 'house', 'sea', 'rug', 'field', 'rock, stone',
+ 'base, pedestal, stand', 'sand', 'skyscraper',
+ 'grandstand, covered stand', 'path', 'runway', 'stairway, staircase',
+ 'river', 'bridge, span', 'blind, screen', 'hill', 'bar',
+ 'hovel, hut, hutch, shack, shanty', 'tower', 'dirt track',
+ 'land, ground, soil', 'escalator, moving staircase, moving stairway',
+ 'buffet, counter, sideboard',
+ 'poster, posting, placard, notice, bill, card', 'stage',
+ 'conveyer belt, conveyor belt, conveyer, conveyor, transporter',
+ 'canopy', 'pool', 'falls', 'tent', 'cradle', 'tank, storage tank',
+ 'lake', 'blanket, cover', 'pier', 'crt screen', 'shower'),
+ 'palette':
+ ADE_PALETTE
+ }
+
+
+@DATASETS.register_module()
+class ADE20KInstanceDataset(CocoDataset):
+ METAINFO = {
+ 'classes':
+ ('bed', 'windowpane', 'cabinet', 'person', 'door', 'table', 'curtain',
+ 'chair', 'car', 'painting', 'sofa', 'shelf', 'mirror', 'armchair',
+ 'seat', 'fence', 'desk', 'wardrobe', 'lamp', 'bathtub', 'railing',
+ 'cushion', 'box', 'column', 'signboard', 'chest of drawers',
+ 'counter', 'sink', 'fireplace', 'refrigerator', 'stairs', 'case',
+ 'pool table', 'pillow', 'screen door', 'bookcase', 'coffee table',
+ 'toilet', 'flower', 'book', 'bench', 'countertop', 'stove', 'palm',
+ 'kitchen island', 'computer', 'swivel chair', 'boat',
+ 'arcade machine', 'bus', 'towel', 'light', 'truck', 'chandelier',
+ 'awning', 'streetlight', 'booth', 'television receiver', 'airplane',
+ 'apparel', 'pole', 'bannister', 'ottoman', 'bottle', 'van', 'ship',
+ 'fountain', 'washer', 'plaything', 'stool', 'barrel', 'basket', 'bag',
+ 'minibike', 'oven', 'ball', 'food', 'step', 'trade name', 'microwave',
+ 'pot', 'animal', 'bicycle', 'dishwasher', 'screen', 'sculpture',
+ 'hood', 'sconce', 'vase', 'traffic light', 'tray', 'ashcan', 'fan',
+ 'plate', 'monitor', 'bulletin board', 'radiator', 'glass', 'clock',
+ 'flag'),
+ 'palette': [(204, 5, 255), (230, 230, 230), (224, 5, 255),
+ (150, 5, 61), (8, 255, 51), (255, 6, 82), (255, 51, 7),
+ (204, 70, 3), (0, 102, 200), (255, 6, 51), (11, 102, 255),
+ (255, 7, 71), (220, 220, 220), (8, 255, 214),
+ (7, 255, 224), (255, 184, 6), (10, 255, 71), (7, 255, 255),
+ (224, 255, 8), (102, 8, 255), (255, 61, 6), (255, 194, 7),
+ (0, 255, 20), (255, 8, 41), (255, 5, 153), (6, 51, 255),
+ (235, 12, 255), (0, 163, 255), (250, 10, 15), (20, 255, 0),
+ (255, 224, 0), (0, 0, 255), (255, 71, 0), (0, 235, 255),
+ (0, 173, 255), (0, 255, 245), (0, 255, 112), (0, 255, 133),
+ (255, 0, 0), (255, 163, 0), (194, 255, 0), (0, 143, 255),
+ (51, 255, 0), (0, 82, 255), (0, 255, 41), (0, 255, 173),
+ (10, 0, 255), (173, 255, 0), (255, 92, 0), (255, 0, 245),
+ (255, 0, 102), (255, 173, 0), (255, 0, 20), (0, 31, 255),
+ (0, 255, 61), (0, 71, 255), (255, 0, 204), (0, 255, 194),
+ (0, 255, 82), (0, 112, 255), (51, 0, 255), (0, 122, 255),
+ (255, 153, 0), (0, 255, 10), (163, 255, 0), (255, 235, 0),
+ (8, 184, 170), (184, 0, 255), (255, 0, 31), (0, 214, 255),
+ (255, 0, 112), (92, 255, 0), (70, 184, 160), (163, 0, 255),
+ (71, 255, 0), (255, 0, 163), (255, 204, 0), (255, 0, 143),
+ (133, 255, 0), (255, 0, 235), (245, 0, 255), (255, 0, 122),
+ (255, 245, 0), (214, 255, 0), (0, 204, 255), (255, 255, 0),
+ (0, 153, 255), (0, 41, 255), (0, 255, 204), (41, 0, 255),
+ (41, 255, 0), (173, 0, 255), (0, 245, 255), (0, 255, 184),
+ (0, 92, 255), (184, 255, 0), (255, 214, 0), (25, 194, 194),
+ (102, 255, 0), (92, 0, 255)],
+ }
+
+
+@DATASETS.register_module()
+class ADE20KSegDataset(BaseSegDataset):
+ """ADE20K dataset.
+
+ In segmentation map annotation for ADE20K, 0 stands for background, which
+ is not included in 150 categories. The ``img_suffix`` is fixed to '.jpg',
+ and ``seg_map_suffix`` is fixed to '.png'.
+ """
+ METAINFO = dict(
+ classes=('wall', 'building', 'sky', 'floor', 'tree', 'ceiling', 'road',
+ 'bed ', 'windowpane', 'grass', 'cabinet', 'sidewalk',
+ 'person', 'earth', 'door', 'table', 'mountain', 'plant',
+ 'curtain', 'chair', 'car', 'water', 'painting', 'sofa',
+ 'shelf', 'house', 'sea', 'mirror', 'rug', 'field', 'armchair',
+ 'seat', 'fence', 'desk', 'rock', 'wardrobe', 'lamp',
+ 'bathtub', 'railing', 'cushion', 'base', 'box', 'column',
+ 'signboard', 'chest of drawers', 'counter', 'sand', 'sink',
+ 'skyscraper', 'fireplace', 'refrigerator', 'grandstand',
+ 'path', 'stairs', 'runway', 'case', 'pool table', 'pillow',
+ 'screen door', 'stairway', 'river', 'bridge', 'bookcase',
+ 'blind', 'coffee table', 'toilet', 'flower', 'book', 'hill',
+ 'bench', 'countertop', 'stove', 'palm', 'kitchen island',
+ 'computer', 'swivel chair', 'boat', 'bar', 'arcade machine',
+ 'hovel', 'bus', 'towel', 'light', 'truck', 'tower',
+ 'chandelier', 'awning', 'streetlight', 'booth',
+ 'television receiver', 'airplane', 'dirt track', 'apparel',
+ 'pole', 'land', 'bannister', 'escalator', 'ottoman', 'bottle',
+ 'buffet', 'poster', 'stage', 'van', 'ship', 'fountain',
+ 'conveyer belt', 'canopy', 'washer', 'plaything',
+ 'swimming pool', 'stool', 'barrel', 'basket', 'waterfall',
+ 'tent', 'bag', 'minibike', 'cradle', 'oven', 'ball', 'food',
+ 'step', 'tank', 'trade name', 'microwave', 'pot', 'animal',
+ 'bicycle', 'lake', 'dishwasher', 'screen', 'blanket',
+ 'sculpture', 'hood', 'sconce', 'vase', 'traffic light',
+ 'tray', 'ashcan', 'fan', 'pier', 'crt screen', 'plate',
+ 'monitor', 'bulletin board', 'shower', 'radiator', 'glass',
+ 'clock', 'flag'),
+ palette=ADE_PALETTE)
+
+ def __init__(self,
+ img_suffix='.jpg',
+ seg_map_suffix='.png',
+ return_classes=False,
+ **kwargs) -> None:
+ self.return_classes = return_classes
+ super().__init__(
+ img_suffix=img_suffix, seg_map_suffix=seg_map_suffix, **kwargs)
+
+ def load_data_list(self) -> List[dict]:
+ """Load annotation from directory or annotation file.
+
+ Returns:
+ List[dict]: All data info of dataset.
+ """
+ data_list = []
+ img_dir = self.data_prefix.get('img_path', None)
+ ann_dir = self.data_prefix.get('seg_map_path', None)
+ for img in fileio.list_dir_or_file(
+ dir_path=img_dir,
+ list_dir=False,
+ suffix=self.img_suffix,
+ recursive=True,
+ backend_args=self.backend_args):
+ data_info = dict(img_path=osp.join(img_dir, img))
+ if ann_dir is not None:
+ seg_map = img.replace(self.img_suffix, self.seg_map_suffix)
+ data_info['seg_map_path'] = osp.join(ann_dir, seg_map)
+ data_info['label_map'] = self.label_map
+ if self.return_classes:
+ data_info['text'] = list(self._metainfo['classes'])
+ data_list.append(data_info)
+ return data_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/api_wrappers/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/api_wrappers/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..8e3c41a2f87b14d10339955208e0502aeeeb7082
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/api_wrappers/__init__.py
@@ -0,0 +1,5 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .coco_api import COCO, COCOeval, COCOPanoptic
+from .cocoeval_mp import COCOevalMP
+
+__all__ = ['COCO', 'COCOeval', 'COCOPanoptic', 'COCOevalMP']
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/api_wrappers/coco_api.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/api_wrappers/coco_api.py
new file mode 100644
index 0000000000000000000000000000000000000000..b2d11a122e1860d1b097710ff98adfddc1508c5a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/api_wrappers/coco_api.py
@@ -0,0 +1,137 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+# This file add snake case alias for coco api
+
+import warnings
+from collections import defaultdict
+from typing import List, Optional, Union
+
+import pycocotools
+from pycocotools.coco import COCO as _COCO
+from pycocotools.cocoeval import COCOeval as _COCOeval
+
+
+class COCO(_COCO):
+ """This class is almost the same as official pycocotools package.
+
+ It implements some snake case function aliases. So that the COCO class has
+ the same interface as LVIS class.
+ """
+
+ def __init__(self, annotation_file=None):
+ if getattr(pycocotools, '__version__', '0') >= '12.0.2':
+ warnings.warn(
+ 'mmpycocotools is deprecated. Please install official pycocotools by "pip install pycocotools"', # noqa: E501
+ UserWarning)
+ super().__init__(annotation_file=annotation_file)
+ self.img_ann_map = self.imgToAnns
+ self.cat_img_map = self.catToImgs
+
+ def get_ann_ids(self, img_ids=[], cat_ids=[], area_rng=[], iscrowd=None):
+ return self.getAnnIds(img_ids, cat_ids, area_rng, iscrowd)
+
+ def get_cat_ids(self, cat_names=[], sup_names=[], cat_ids=[]):
+ return self.getCatIds(cat_names, sup_names, cat_ids)
+
+ def get_img_ids(self, img_ids=[], cat_ids=[]):
+ return self.getImgIds(img_ids, cat_ids)
+
+ def load_anns(self, ids):
+ return self.loadAnns(ids)
+
+ def load_cats(self, ids):
+ return self.loadCats(ids)
+
+ def load_imgs(self, ids):
+ return self.loadImgs(ids)
+
+
+# just for the ease of import
+COCOeval = _COCOeval
+
+
+class COCOPanoptic(COCO):
+ """This wrapper is for loading the panoptic style annotation file.
+
+ The format is shown in the CocoPanopticDataset class.
+
+ Args:
+ annotation_file (str, optional): Path of annotation file.
+ Defaults to None.
+ """
+
+ def __init__(self, annotation_file: Optional[str] = None) -> None:
+ super(COCOPanoptic, self).__init__(annotation_file)
+
+ def createIndex(self) -> None:
+ """Create index."""
+ # create index
+ print('creating index...')
+ # anns stores 'segment_id -> annotation'
+ anns, cats, imgs = {}, {}, {}
+ img_to_anns, cat_to_imgs = defaultdict(list), defaultdict(list)
+ if 'annotations' in self.dataset:
+ for ann in self.dataset['annotations']:
+ for seg_ann in ann['segments_info']:
+ # to match with instance.json
+ seg_ann['image_id'] = ann['image_id']
+ img_to_anns[ann['image_id']].append(seg_ann)
+ # segment_id is not unique in coco dataset orz...
+ # annotations from different images but
+ # may have same segment_id
+ if seg_ann['id'] in anns.keys():
+ anns[seg_ann['id']].append(seg_ann)
+ else:
+ anns[seg_ann['id']] = [seg_ann]
+
+ # filter out annotations from other images
+ img_to_anns_ = defaultdict(list)
+ for k, v in img_to_anns.items():
+ img_to_anns_[k] = [x for x in v if x['image_id'] == k]
+ img_to_anns = img_to_anns_
+
+ if 'images' in self.dataset:
+ for img_info in self.dataset['images']:
+ img_info['segm_file'] = img_info['file_name'].replace(
+ '.jpg', '.png')
+ imgs[img_info['id']] = img_info
+
+ if 'categories' in self.dataset:
+ for cat in self.dataset['categories']:
+ cats[cat['id']] = cat
+
+ if 'annotations' in self.dataset and 'categories' in self.dataset:
+ for ann in self.dataset['annotations']:
+ for seg_ann in ann['segments_info']:
+ cat_to_imgs[seg_ann['category_id']].append(ann['image_id'])
+
+ print('index created!')
+
+ self.anns = anns
+ self.imgToAnns = img_to_anns
+ self.catToImgs = cat_to_imgs
+ self.imgs = imgs
+ self.cats = cats
+
+ def load_anns(self,
+ ids: Union[List[int], int] = []) -> Optional[List[dict]]:
+ """Load anns with the specified ids.
+
+ ``self.anns`` is a list of annotation lists instead of a
+ list of annotations.
+
+ Args:
+ ids (Union[List[int], int]): Integer ids specifying anns.
+
+ Returns:
+ anns (List[dict], optional): Loaded ann objects.
+ """
+ anns = []
+
+ if hasattr(ids, '__iter__') and hasattr(ids, '__len__'):
+ # self.anns is a list of annotation lists instead of
+ # a list of annotations
+ for id in ids:
+ anns += self.anns[id]
+ return anns
+ elif type(ids) == int:
+ return self.anns[ids]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/api_wrappers/cocoeval_mp.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/api_wrappers/cocoeval_mp.py
new file mode 100644
index 0000000000000000000000000000000000000000..b3673ea7a7edc593cb49fb336f352a20c1b1015b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/api_wrappers/cocoeval_mp.py
@@ -0,0 +1,296 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import itertools
+import time
+from collections import defaultdict
+
+import numpy as np
+import torch.multiprocessing as mp
+from mmengine.logging import MMLogger
+from pycocotools.cocoeval import COCOeval
+from tqdm import tqdm
+
+
+class COCOevalMP(COCOeval):
+
+ def _prepare(self):
+ '''
+ Prepare ._gts and ._dts for evaluation based on params
+ :return: None
+ '''
+
+ def _toMask(anns, coco):
+ # modify ann['segmentation'] by reference
+ for ann in anns:
+ rle = coco.annToRLE(ann)
+ ann['segmentation'] = rle
+
+ p = self.params
+ if p.useCats:
+ gts = []
+ dts = []
+ img_ids = set(p.imgIds)
+ cat_ids = set(p.catIds)
+ for gt in self.cocoGt.dataset['annotations']:
+ if (gt['category_id'] in cat_ids) and (gt['image_id']
+ in img_ids):
+ gts.append(gt)
+ for dt in self.cocoDt.dataset['annotations']:
+ if (dt['category_id'] in cat_ids) and (dt['image_id']
+ in img_ids):
+ dts.append(dt)
+ # gts=self.cocoGt.loadAnns(self.cocoGt.getAnnIds(imgIds=p.imgIds, catIds=p.catIds)) # noqa
+ # dts=self.cocoDt.loadAnns(self.cocoDt.getAnnIds(imgIds=p.imgIds, catIds=p.catIds)) # noqa
+ # gts=self.cocoGt.dataset['annotations']
+ # dts=self.cocoDt.dataset['annotations']
+ else:
+ gts = self.cocoGt.loadAnns(self.cocoGt.getAnnIds(imgIds=p.imgIds))
+ dts = self.cocoDt.loadAnns(self.cocoDt.getAnnIds(imgIds=p.imgIds))
+
+ # convert ground truth to mask if iouType == 'segm'
+ if p.iouType == 'segm':
+ _toMask(gts, self.cocoGt)
+ _toMask(dts, self.cocoDt)
+ # set ignore flag
+ for gt in gts:
+ gt['ignore'] = gt['ignore'] if 'ignore' in gt else 0
+ gt['ignore'] = 'iscrowd' in gt and gt['iscrowd']
+ if p.iouType == 'keypoints':
+ gt['ignore'] = (gt['num_keypoints'] == 0) or gt['ignore']
+ self._gts = defaultdict(list) # gt for evaluation
+ self._dts = defaultdict(list) # dt for evaluation
+ for gt in gts:
+ self._gts[gt['image_id'], gt['category_id']].append(gt)
+ for dt in dts:
+ self._dts[dt['image_id'], dt['category_id']].append(dt)
+ self.evalImgs = defaultdict(
+ list) # per-image per-category evaluation results
+ self.eval = {} # accumulated evaluation results
+
+ def evaluate(self):
+ """Run per image evaluation on given images and store results (a list
+ of dict) in self.evalImgs.
+
+ :return: None
+ """
+ tic = time.time()
+ print('Running per image evaluation...')
+ p = self.params
+ # add backward compatibility if useSegm is specified in params
+ if p.useSegm is not None:
+ p.iouType = 'segm' if p.useSegm == 1 else 'bbox'
+ print('useSegm (deprecated) is not None. Running {} evaluation'.
+ format(p.iouType))
+ print('Evaluate annotation type *{}*'.format(p.iouType))
+ p.imgIds = list(np.unique(p.imgIds))
+ if p.useCats:
+ p.catIds = list(np.unique(p.catIds))
+ p.maxDets = sorted(p.maxDets)
+ self.params = p
+
+ # loop through images, area range, max detection number
+ catIds = p.catIds if p.useCats else [-1]
+
+ nproc = 8
+ split_size = len(catIds) // nproc
+ mp_params = []
+ for i in range(nproc):
+ begin = i * split_size
+ end = (i + 1) * split_size
+ if i == nproc - 1:
+ end = len(catIds)
+ mp_params.append((catIds[begin:end], ))
+
+ MMLogger.get_current_instance().info(
+ 'start multi processing evaluation ...')
+ with mp.Pool(nproc) as pool:
+ self.evalImgs = pool.starmap(self._evaluateImg, mp_params)
+
+ self.evalImgs = list(itertools.chain(*self.evalImgs))
+
+ self._paramsEval = copy.deepcopy(self.params)
+ toc = time.time()
+ print('DONE (t={:0.2f}s).'.format(toc - tic))
+
+ def _evaluateImg(self, catids_chunk):
+ self._prepare()
+ p = self.params
+ maxDet = max(p.maxDets)
+ all_params = []
+ for catId in catids_chunk:
+ for areaRng in p.areaRng:
+ for imgId in p.imgIds:
+ all_params.append((catId, areaRng, imgId))
+ evalImgs = [
+ self.evaluateImg(imgId, catId, areaRng, maxDet)
+ for catId, areaRng, imgId in tqdm(all_params)
+ ]
+ return evalImgs
+
+ def evaluateImg(self, imgId, catId, aRng, maxDet):
+ p = self.params
+ if p.useCats:
+ gt = self._gts[imgId, catId]
+ dt = self._dts[imgId, catId]
+ else:
+ gt = [_ for cId in p.catIds for _ in self._gts[imgId, cId]]
+ dt = [_ for cId in p.catIds for _ in self._dts[imgId, cId]]
+ if len(gt) == 0 and len(dt) == 0:
+ return None
+
+ for g in gt:
+ if g['ignore'] or (g['area'] < aRng[0] or g['area'] > aRng[1]):
+ g['_ignore'] = 1
+ else:
+ g['_ignore'] = 0
+
+ # sort dt highest score first, sort gt ignore last
+ gtind = np.argsort([g['_ignore'] for g in gt], kind='mergesort')
+ gt = [gt[i] for i in gtind]
+ dtind = np.argsort([-d['score'] for d in dt], kind='mergesort')
+ dt = [dt[i] for i in dtind[0:maxDet]]
+ iscrowd = [int(o['iscrowd']) for o in gt]
+ # load computed ious
+ # ious = self.ious[imgId, catId][:, gtind] if len(self.ious[imgId, catId]) > 0 else self.ious[imgId, catId] # noqa
+ ious = self.computeIoU(imgId, catId)
+ ious = ious[:, gtind] if len(ious) > 0 else ious
+
+ T = len(p.iouThrs)
+ G = len(gt)
+ D = len(dt)
+ gtm = np.zeros((T, G))
+ dtm = np.zeros((T, D))
+ gtIg = np.array([g['_ignore'] for g in gt])
+ dtIg = np.zeros((T, D))
+ if not len(ious) == 0:
+ for tind, t in enumerate(p.iouThrs):
+ for dind, d in enumerate(dt):
+ # information about best match so far (m=-1 -> unmatched)
+ iou = min([t, 1 - 1e-10])
+ m = -1
+ for gind, g in enumerate(gt):
+ # if this gt already matched, and not a crowd, continue
+ if gtm[tind, gind] > 0 and not iscrowd[gind]:
+ continue
+ # if dt matched to reg gt, and on ignore gt, stop
+ if m > -1 and gtIg[m] == 0 and gtIg[gind] == 1:
+ break
+ # continue to next gt unless better match made
+ if ious[dind, gind] < iou:
+ continue
+ # if match successful and best so far,
+ # store appropriately
+ iou = ious[dind, gind]
+ m = gind
+ # if match made store id of match for both dt and gt
+ if m == -1:
+ continue
+ dtIg[tind, dind] = gtIg[m]
+ dtm[tind, dind] = gt[m]['id']
+ gtm[tind, m] = d['id']
+ # set unmatched detections outside of area range to ignore
+ a = np.array([d['area'] < aRng[0] or d['area'] > aRng[1]
+ for d in dt]).reshape((1, len(dt)))
+ dtIg = np.logical_or(dtIg, np.logical_and(dtm == 0, np.repeat(a, T,
+ 0)))
+ # store results for given image and category
+
+ return {
+ 'image_id': imgId,
+ 'category_id': catId,
+ 'aRng': aRng,
+ 'maxDet': maxDet,
+ 'dtIds': [d['id'] for d in dt],
+ 'gtIds': [g['id'] for g in gt],
+ 'dtMatches': dtm,
+ 'gtMatches': gtm,
+ 'dtScores': [d['score'] for d in dt],
+ 'gtIgnore': gtIg,
+ 'dtIgnore': dtIg,
+ }
+
+ def summarize(self):
+ """Compute and display summary metrics for evaluation results.
+
+ Note this function can *only* be applied on the default parameter
+ setting
+ """
+
+ def _summarize(ap=1, iouThr=None, areaRng='all', maxDets=100):
+ p = self.params
+ iStr = ' {:<18} {} @[ IoU={:<9} | area={:>6s} | maxDets={:>3d} ] = {:0.3f}' # noqa
+ titleStr = 'Average Precision' if ap == 1 else 'Average Recall'
+ typeStr = '(AP)' if ap == 1 else '(AR)'
+ iouStr = '{:0.2f}:{:0.2f}'.format(p.iouThrs[0], p.iouThrs[-1]) \
+ if iouThr is None else '{:0.2f}'.format(iouThr)
+
+ aind = [
+ i for i, aRng in enumerate(p.areaRngLbl) if aRng == areaRng
+ ]
+ mind = [i for i, mDet in enumerate(p.maxDets) if mDet == maxDets]
+ if ap == 1:
+ # dimension of precision: [TxRxKxAxM]
+ s = self.eval['precision']
+ # IoU
+ if iouThr is not None:
+ t = np.where(iouThr == p.iouThrs)[0]
+ s = s[t]
+ s = s[:, :, :, aind, mind]
+ else:
+ # dimension of recall: [TxKxAxM]
+ s = self.eval['recall']
+ if iouThr is not None:
+ t = np.where(iouThr == p.iouThrs)[0]
+ s = s[t]
+ s = s[:, :, aind, mind]
+ if len(s[s > -1]) == 0:
+ mean_s = -1
+ else:
+ mean_s = np.mean(s[s > -1])
+ print(
+ iStr.format(titleStr, typeStr, iouStr, areaRng, maxDets,
+ mean_s))
+ return mean_s
+
+ def _summarizeDets():
+ stats = []
+ stats.append(_summarize(1, maxDets=self.params.maxDets[-1]))
+ stats.append(
+ _summarize(1, iouThr=.5, maxDets=self.params.maxDets[-1]))
+ stats.append(
+ _summarize(1, iouThr=.75, maxDets=self.params.maxDets[-1]))
+ for area_rng in ('small', 'medium', 'large'):
+ stats.append(
+ _summarize(
+ 1, areaRng=area_rng, maxDets=self.params.maxDets[-1]))
+ for max_det in self.params.maxDets:
+ stats.append(_summarize(0, maxDets=max_det))
+ for area_rng in ('small', 'medium', 'large'):
+ stats.append(
+ _summarize(
+ 0, areaRng=area_rng, maxDets=self.params.maxDets[-1]))
+ stats = np.array(stats)
+ return stats
+
+ def _summarizeKps():
+ stats = np.zeros((10, ))
+ stats[0] = _summarize(1, maxDets=20)
+ stats[1] = _summarize(1, maxDets=20, iouThr=.5)
+ stats[2] = _summarize(1, maxDets=20, iouThr=.75)
+ stats[3] = _summarize(1, maxDets=20, areaRng='medium')
+ stats[4] = _summarize(1, maxDets=20, areaRng='large')
+ stats[5] = _summarize(0, maxDets=20)
+ stats[6] = _summarize(0, maxDets=20, iouThr=.5)
+ stats[7] = _summarize(0, maxDets=20, iouThr=.75)
+ stats[8] = _summarize(0, maxDets=20, areaRng='medium')
+ stats[9] = _summarize(0, maxDets=20, areaRng='large')
+ return stats
+
+ if not self.eval:
+ raise Exception('Please run accumulate() first')
+ iouType = self.params.iouType
+ if iouType == 'segm' or iouType == 'bbox':
+ summarize = _summarizeDets
+ elif iouType == 'keypoints':
+ summarize = _summarizeKps
+ self.stats = summarize()
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/base_det_dataset.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/base_det_dataset.py
new file mode 100644
index 0000000000000000000000000000000000000000..8b3876d5c06eb7d3741a29fe8b0963a7e425ec1b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/base_det_dataset.py
@@ -0,0 +1,131 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import os.path as osp
+from typing import List, Optional
+
+from mmengine.dataset import BaseDataset
+from mmengine.fileio import load
+from mmengine.utils import is_abs
+
+from ..registry import DATASETS
+
+
+@DATASETS.register_module()
+class BaseDetDataset(BaseDataset):
+ """Base dataset for detection.
+
+ Args:
+ proposal_file (str, optional): Proposals file path. Defaults to None.
+ file_client_args (dict): Arguments to instantiate the
+ corresponding backend in mmdet <= 3.0.0rc6. Defaults to None.
+ backend_args (dict, optional): Arguments to instantiate the
+ corresponding backend. Defaults to None.
+ return_classes (bool): Whether to return class information
+ for open vocabulary-based algorithms. Defaults to False.
+ caption_prompt (dict, optional): Prompt for captioning.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ *args,
+ seg_map_suffix: str = '.png',
+ proposal_file: Optional[str] = None,
+ file_client_args: dict = None,
+ backend_args: dict = None,
+ return_classes: bool = False,
+ caption_prompt: Optional[dict] = None,
+ **kwargs) -> None:
+ self.seg_map_suffix = seg_map_suffix
+ self.proposal_file = proposal_file
+ self.backend_args = backend_args
+ self.return_classes = return_classes
+ self.caption_prompt = caption_prompt
+ if self.caption_prompt is not None:
+ assert self.return_classes, \
+ 'return_classes must be True when using caption_prompt'
+ if file_client_args is not None:
+ raise RuntimeError(
+ 'The `file_client_args` is deprecated, '
+ 'please use `backend_args` instead, please refer to'
+ 'https://github.com/open-mmlab/mmdetection/blob/main/configs/_base_/datasets/coco_detection.py' # noqa: E501
+ )
+ super().__init__(*args, **kwargs)
+
+ def full_init(self) -> None:
+ """Load annotation file and set ``BaseDataset._fully_initialized`` to
+ True.
+
+ If ``lazy_init=False``, ``full_init`` will be called during the
+ instantiation and ``self._fully_initialized`` will be set to True. If
+ ``obj._fully_initialized=False``, the class method decorated by
+ ``force_full_init`` will call ``full_init`` automatically.
+
+ Several steps to initialize annotation:
+
+ - load_data_list: Load annotations from annotation file.
+ - load_proposals: Load proposals from proposal file, if
+ `self.proposal_file` is not None.
+ - filter data information: Filter annotations according to
+ filter_cfg.
+ - slice_data: Slice dataset according to ``self._indices``
+ - serialize_data: Serialize ``self.data_list`` if
+ ``self.serialize_data`` is True.
+ """
+ if self._fully_initialized:
+ return
+ # load data information
+ self.data_list = self.load_data_list()
+ # get proposals from file
+ if self.proposal_file is not None:
+ self.load_proposals()
+ # filter illegal data, such as data that has no annotations.
+ self.data_list = self.filter_data()
+
+ # Get subset data according to indices.
+ if self._indices is not None:
+ self.data_list = self._get_unserialized_subset(self._indices)
+
+ # serialize data_list
+ if self.serialize_data:
+ self.data_bytes, self.data_address = self._serialize_data()
+
+ self._fully_initialized = True
+
+ def load_proposals(self) -> None:
+ """Load proposals from proposals file.
+
+ The `proposals_list` should be a dict[img_path: proposals]
+ with the same length as `data_list`. And the `proposals` should be
+ a `dict` or :obj:`InstanceData` usually contains following keys.
+
+ - bboxes (np.ndarry): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - scores (np.ndarry): Classification scores, has a shape
+ (num_instance, ).
+ """
+ # TODO: Add Unit Test after fully support Dump-Proposal Metric
+ if not is_abs(self.proposal_file):
+ self.proposal_file = osp.join(self.data_root, self.proposal_file)
+ proposals_list = load(
+ self.proposal_file, backend_args=self.backend_args)
+ assert len(self.data_list) == len(proposals_list)
+ for data_info in self.data_list:
+ img_path = data_info['img_path']
+ # `file_name` is the key to obtain the proposals from the
+ # `proposals_list`.
+ file_name = osp.join(
+ osp.split(osp.split(img_path)[0])[-1],
+ osp.split(img_path)[-1])
+ proposals = proposals_list[file_name]
+ data_info['proposals'] = proposals
+
+ def get_cat_ids(self, idx: int) -> List[int]:
+ """Get COCO category ids by index.
+
+ Args:
+ idx (int): Index of data.
+
+ Returns:
+ List[int]: All categories in the image of specified index.
+ """
+ instances = self.get_data_info(idx)['instances']
+ return [instance['bbox_label'] for instance in instances]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/base_semseg_dataset.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/base_semseg_dataset.py
new file mode 100644
index 0000000000000000000000000000000000000000..d10f762a21a897ab8274fbe9eefab054691a7c60
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/base_semseg_dataset.py
@@ -0,0 +1,265 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import os.path as osp
+from typing import Callable, Dict, List, Optional, Sequence, Union
+
+import mmengine
+import mmengine.fileio as fileio
+import numpy as np
+from mmengine.dataset import BaseDataset, Compose
+
+from mmdet.registry import DATASETS
+
+
+@DATASETS.register_module()
+class BaseSegDataset(BaseDataset):
+ """Custom dataset for semantic segmentation. An example of file structure
+ is as followed.
+
+ .. code-block:: none
+
+ ├── data
+ │ ├── my_dataset
+ │ │ ├── img_dir
+ │ │ │ ├── train
+ │ │ │ │ ├── xxx{img_suffix}
+ │ │ │ │ ├── yyy{img_suffix}
+ │ │ │ │ ├── zzz{img_suffix}
+ │ │ │ ├── val
+ │ │ ├── ann_dir
+ │ │ │ ├── train
+ │ │ │ │ ├── xxx{seg_map_suffix}
+ │ │ │ │ ├── yyy{seg_map_suffix}
+ │ │ │ │ ├── zzz{seg_map_suffix}
+ │ │ │ ├── val
+
+ The img/gt_semantic_seg pair of BaseSegDataset should be of the same
+ except suffix. A valid img/gt_semantic_seg filename pair should be like
+ ``xxx{img_suffix}`` and ``xxx{seg_map_suffix}`` (extension is also included
+ in the suffix). If split is given, then ``xxx`` is specified in txt file.
+ Otherwise, all files in ``img_dir/``and ``ann_dir`` will be loaded.
+ Please refer to ``docs/en/tutorials/new_dataset.md`` for more details.
+
+
+ Args:
+ ann_file (str): Annotation file path. Defaults to ''.
+ metainfo (dict, optional): Meta information for dataset, such as
+ specify classes to load. Defaults to None.
+ data_root (str, optional): The root directory for ``data_prefix`` and
+ ``ann_file``. Defaults to None.
+ data_prefix (dict, optional): Prefix for training data. Defaults to
+ dict(img_path=None, seg_map_path=None).
+ img_suffix (str): Suffix of images. Default: '.jpg'
+ seg_map_suffix (str): Suffix of segmentation maps. Default: '.png'
+ filter_cfg (dict, optional): Config for filter data. Defaults to None.
+ indices (int or Sequence[int], optional): Support using first few
+ data in annotation file to facilitate training/testing on a smaller
+ dataset. Defaults to None which means using all ``data_infos``.
+ serialize_data (bool, optional): Whether to hold memory using
+ serialized objects, when enabled, data loader workers can use
+ shared RAM from master process instead of making a copy. Defaults
+ to True.
+ pipeline (list, optional): Processing pipeline. Defaults to [].
+ test_mode (bool, optional): ``test_mode=True`` means in test phase.
+ Defaults to False.
+ lazy_init (bool, optional): Whether to load annotation during
+ instantiation. In some cases, such as visualization, only the meta
+ information of the dataset is needed, which is not necessary to
+ load annotation file. ``Basedataset`` can skip load annotations to
+ save time by set ``lazy_init=True``. Defaults to False.
+ use_label_map (bool, optional): Whether to use label map.
+ Defaults to False.
+ max_refetch (int, optional): If ``Basedataset.prepare_data`` get a
+ None img. The maximum extra number of cycles to get a valid
+ image. Defaults to 1000.
+ backend_args (dict, Optional): Arguments to instantiate a file backend.
+ See https://mmengine.readthedocs.io/en/latest/api/fileio.htm
+ for details. Defaults to None.
+ Notes: mmcv>=2.0.0rc4 required.
+ """
+ METAINFO: dict = dict()
+
+ def __init__(self,
+ ann_file: str = '',
+ img_suffix='.jpg',
+ seg_map_suffix='.png',
+ metainfo: Optional[dict] = None,
+ data_root: Optional[str] = None,
+ data_prefix: dict = dict(img_path='', seg_map_path=''),
+ filter_cfg: Optional[dict] = None,
+ indices: Optional[Union[int, Sequence[int]]] = None,
+ serialize_data: bool = True,
+ pipeline: List[Union[dict, Callable]] = [],
+ test_mode: bool = False,
+ lazy_init: bool = False,
+ use_label_map: bool = False,
+ max_refetch: int = 1000,
+ backend_args: Optional[dict] = None) -> None:
+
+ self.img_suffix = img_suffix
+ self.seg_map_suffix = seg_map_suffix
+ self.backend_args = backend_args.copy() if backend_args else None
+
+ self.data_root = data_root
+ self.data_prefix = copy.copy(data_prefix)
+ self.ann_file = ann_file
+ self.filter_cfg = copy.deepcopy(filter_cfg)
+ self._indices = indices
+ self.serialize_data = serialize_data
+ self.test_mode = test_mode
+ self.max_refetch = max_refetch
+ self.data_list: List[dict] = []
+ self.data_bytes: np.ndarray
+
+ # Set meta information.
+ self._metainfo = self._load_metainfo(copy.deepcopy(metainfo))
+
+ # Get label map for custom classes
+ new_classes = self._metainfo.get('classes', None)
+ self.label_map = self.get_label_map(
+ new_classes) if use_label_map else None
+ self._metainfo.update(dict(label_map=self.label_map))
+
+ # Update palette based on label map or generate palette
+ # if it is not defined
+ updated_palette = self._update_palette()
+ self._metainfo.update(dict(palette=updated_palette))
+
+ # Join paths.
+ if self.data_root is not None:
+ self._join_prefix()
+
+ # Build pipeline.
+ self.pipeline = Compose(pipeline)
+ # Full initialize the dataset.
+ if not lazy_init:
+ self.full_init()
+
+ if test_mode:
+ assert self._metainfo.get('classes') is not None, \
+ 'dataset metainfo `classes` should be specified when testing'
+
+ @classmethod
+ def get_label_map(cls,
+ new_classes: Optional[Sequence] = None
+ ) -> Union[Dict, None]:
+ """Require label mapping.
+
+ The ``label_map`` is a dictionary, its keys are the old label ids and
+ its values are the new label ids, and is used for changing pixel
+ labels in load_annotations. If and only if old classes in cls.METAINFO
+ is not equal to new classes in self._metainfo and nether of them is not
+ None, `label_map` is not None.
+
+ Args:
+ new_classes (list, tuple, optional): The new classes name from
+ metainfo. Default to None.
+
+
+ Returns:
+ dict, optional: The mapping from old classes in cls.METAINFO to
+ new classes in self._metainfo
+ """
+ old_classes = cls.METAINFO.get('classes', None)
+ if (new_classes is not None and old_classes is not None
+ and list(new_classes) != list(old_classes)):
+
+ label_map = {}
+ if not set(new_classes).issubset(cls.METAINFO['classes']):
+ raise ValueError(
+ f'new classes {new_classes} is not a '
+ f'subset of classes {old_classes} in METAINFO.')
+ for i, c in enumerate(old_classes):
+ if c not in new_classes:
+ # 0 is background
+ label_map[i] = 0
+ else:
+ label_map[i] = new_classes.index(c)
+ return label_map
+ else:
+ return None
+
+ def _update_palette(self) -> list:
+ """Update palette after loading metainfo.
+
+ If length of palette is equal to classes, just return the palette.
+ If palette is not defined, it will randomly generate a palette.
+ If classes is updated by customer, it will return the subset of
+ palette.
+
+ Returns:
+ Sequence: Palette for current dataset.
+ """
+ palette = self._metainfo.get('palette', [])
+ classes = self._metainfo.get('classes', [])
+ # palette does match classes
+ if len(palette) == len(classes):
+ return palette
+
+ if len(palette) == 0:
+ # Get random state before set seed, and restore
+ # random state later.
+ # It will prevent loss of randomness, as the palette
+ # may be different in each iteration if not specified.
+ # See: https://github.com/open-mmlab/mmdetection/issues/5844
+ state = np.random.get_state()
+ np.random.seed(42)
+ # random palette
+ new_palette = np.random.randint(
+ 0, 255, size=(len(classes), 3)).tolist()
+ np.random.set_state(state)
+ elif len(palette) >= len(classes) and self.label_map is not None:
+ new_palette = []
+ # return subset of palette
+ for old_id, new_id in sorted(
+ self.label_map.items(), key=lambda x: x[1]):
+ # 0 is background
+ if new_id != 0:
+ new_palette.append(palette[old_id])
+ new_palette = type(palette)(new_palette)
+ elif len(palette) >= len(classes):
+ # Allow palette length is greater than classes.
+ return palette
+ else:
+ raise ValueError('palette does not match classes '
+ f'as metainfo is {self._metainfo}.')
+ return new_palette
+
+ def load_data_list(self) -> List[dict]:
+ """Load annotation from directory or annotation file.
+
+ Returns:
+ list[dict]: All data info of dataset.
+ """
+ data_list = []
+ img_dir = self.data_prefix.get('img_path', None)
+ ann_dir = self.data_prefix.get('seg_map_path', None)
+ if not osp.isdir(self.ann_file) and self.ann_file:
+ assert osp.isfile(self.ann_file), \
+ f'Failed to load `ann_file` {self.ann_file}'
+ lines = mmengine.list_from_file(
+ self.ann_file, backend_args=self.backend_args)
+ for line in lines:
+ img_name = line.strip()
+ data_info = dict(
+ img_path=osp.join(img_dir, img_name + self.img_suffix))
+ if ann_dir is not None:
+ seg_map = img_name + self.seg_map_suffix
+ data_info['seg_map_path'] = osp.join(ann_dir, seg_map)
+ data_info['label_map'] = self.label_map
+ data_list.append(data_info)
+ else:
+ for img in fileio.list_dir_or_file(
+ dir_path=img_dir,
+ list_dir=False,
+ suffix=self.img_suffix,
+ recursive=True,
+ backend_args=self.backend_args):
+ data_info = dict(img_path=osp.join(img_dir, img))
+ if ann_dir is not None:
+ seg_map = img.replace(self.img_suffix, self.seg_map_suffix)
+ data_info['seg_map_path'] = osp.join(ann_dir, seg_map)
+ data_info['label_map'] = self.label_map
+ data_list.append(data_info)
+ data_list = sorted(data_list, key=lambda x: x['img_path'])
+ return data_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/base_video_dataset.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/base_video_dataset.py
new file mode 100644
index 0000000000000000000000000000000000000000..0a4a7a25f16206f06c7b64a7ce4c3588efd5455e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/base_video_dataset.py
@@ -0,0 +1,304 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import os.path as osp
+from collections import defaultdict
+from typing import Any, List, Tuple
+
+import mmengine.fileio as fileio
+from mmengine.dataset import BaseDataset
+from mmengine.logging import print_log
+
+from mmdet.datasets.api_wrappers import COCO
+from mmdet.registry import DATASETS
+
+
+@DATASETS.register_module()
+class BaseVideoDataset(BaseDataset):
+ """Base video dataset for VID, MOT and VIS tasks."""
+
+ META = dict(classes=None)
+ # ann_id is unique in coco dataset.
+ ANN_ID_UNIQUE = True
+
+ def __init__(self, *args, backend_args: dict = None, **kwargs):
+ self.backend_args = backend_args
+ super().__init__(*args, **kwargs)
+
+ def load_data_list(self) -> Tuple[List[dict], List]:
+ """Load annotations from an annotation file named as ``self.ann_file``.
+
+ Returns:
+ tuple(list[dict], list): A list of annotation and a list of
+ valid data indices.
+ """
+ with fileio.get_local_path(self.ann_file) as local_path:
+ self.coco = COCO(local_path)
+ # The order of returned `cat_ids` will not
+ # change with the order of the classes
+ self.cat_ids = self.coco.get_cat_ids(
+ cat_names=self.metainfo['classes'])
+ self.cat2label = {cat_id: i for i, cat_id in enumerate(self.cat_ids)}
+ self.cat_img_map = copy.deepcopy(self.coco.cat_img_map)
+ # used in `filter_data`
+ self.img_ids_with_ann = set()
+
+ img_ids = self.coco.get_img_ids()
+ total_ann_ids = []
+ # if ``video_id`` is not in the annotation file, we will assign a big
+ # unique video_id for this video.
+ single_video_id = 100000
+ videos = {}
+ for img_id in img_ids:
+ raw_img_info = self.coco.load_imgs([img_id])[0]
+ raw_img_info['img_id'] = img_id
+ if 'video_id' not in raw_img_info:
+ single_video_id = single_video_id + 1
+ video_id = single_video_id
+ else:
+ video_id = raw_img_info['video_id']
+
+ if video_id not in videos:
+ videos[video_id] = {
+ 'video_id': video_id,
+ 'images': [],
+ 'video_length': 0
+ }
+
+ videos[video_id]['video_length'] += 1
+ ann_ids = self.coco.get_ann_ids(
+ img_ids=[img_id], cat_ids=self.cat_ids)
+ raw_ann_info = self.coco.load_anns(ann_ids)
+ total_ann_ids.extend(ann_ids)
+
+ parsed_data_info = self.parse_data_info(
+ dict(raw_img_info=raw_img_info, raw_ann_info=raw_ann_info))
+
+ if len(parsed_data_info['instances']) > 0:
+ self.img_ids_with_ann.add(parsed_data_info['img_id'])
+
+ videos[video_id]['images'].append(parsed_data_info)
+
+ data_list = [v for v in videos.values()]
+
+ if self.ANN_ID_UNIQUE:
+ assert len(set(total_ann_ids)) == len(
+ total_ann_ids
+ ), f"Annotation ids in '{self.ann_file}' are not unique!"
+
+ del self.coco
+
+ return data_list
+
+ def parse_data_info(self, raw_data_info: dict) -> dict:
+ """Parse raw annotation to target format.
+
+ Args:
+ raw_data_info (dict): Raw data information loaded from
+ ``ann_file``.
+
+ Returns:
+ dict: Parsed annotation.
+ """
+ img_info = raw_data_info['raw_img_info']
+ ann_info = raw_data_info['raw_ann_info']
+ data_info = {}
+
+ data_info.update(img_info)
+ if self.data_prefix.get('img_path', None) is not None:
+ img_path = osp.join(self.data_prefix['img_path'],
+ img_info['file_name'])
+ else:
+ img_path = img_info['file_name']
+ data_info['img_path'] = img_path
+
+ instances = []
+ for i, ann in enumerate(ann_info):
+ instance = {}
+
+ if ann.get('ignore', False):
+ continue
+ x1, y1, w, h = ann['bbox']
+ inter_w = max(0, min(x1 + w, img_info['width']) - max(x1, 0))
+ inter_h = max(0, min(y1 + h, img_info['height']) - max(y1, 0))
+ if inter_w * inter_h == 0:
+ continue
+ if ann['area'] <= 0 or w < 1 or h < 1:
+ continue
+ if ann['category_id'] not in self.cat_ids:
+ continue
+ bbox = [x1, y1, x1 + w, y1 + h]
+
+ if ann.get('iscrowd', False):
+ instance['ignore_flag'] = 1
+ else:
+ instance['ignore_flag'] = 0
+ instance['bbox'] = bbox
+ instance['bbox_label'] = self.cat2label[ann['category_id']]
+ if ann.get('segmentation', None):
+ instance['mask'] = ann['segmentation']
+ if ann.get('instance_id', None):
+ instance['instance_id'] = ann['instance_id']
+ else:
+ # image dataset usually has no `instance_id`.
+ # Therefore, we set it to `i`.
+ instance['instance_id'] = i
+ instances.append(instance)
+ data_info['instances'] = instances
+ return data_info
+
+ def filter_data(self) -> List[int]:
+ """Filter image annotations according to filter_cfg.
+
+ Returns:
+ list[int]: Filtered results.
+ """
+ if self.test_mode:
+ return self.data_list
+
+ num_imgs_before_filter = sum(
+ [len(info['images']) for info in self.data_list])
+ num_imgs_after_filter = 0
+
+ # obtain images that contain annotations of the required categories
+ ids_in_cat = set()
+ for i, class_id in enumerate(self.cat_ids):
+ ids_in_cat |= set(self.cat_img_map[class_id])
+ # merge the image id sets of the two conditions and use the merged set
+ # to filter out images if self.filter_empty_gt=True
+ ids_in_cat &= self.img_ids_with_ann
+
+ new_data_list = []
+ for video_data_info in self.data_list:
+ imgs_data_info = video_data_info['images']
+ valid_imgs_data_info = []
+
+ for data_info in imgs_data_info:
+ img_id = data_info['img_id']
+ width = data_info['width']
+ height = data_info['height']
+ # TODO: simplify these conditions
+ if self.filter_cfg is None:
+ if img_id not in ids_in_cat:
+ video_data_info['video_length'] -= 1
+ continue
+ if min(width, height) >= 32:
+ valid_imgs_data_info.append(data_info)
+ num_imgs_after_filter += 1
+ else:
+ video_data_info['video_length'] -= 1
+ else:
+ if self.filter_cfg.get('filter_empty_gt',
+ True) and img_id not in ids_in_cat:
+ video_data_info['video_length'] -= 1
+ continue
+ if min(width, height) >= self.filter_cfg.get(
+ 'min_size', 32):
+ valid_imgs_data_info.append(data_info)
+ num_imgs_after_filter += 1
+ else:
+ video_data_info['video_length'] -= 1
+ video_data_info['images'] = valid_imgs_data_info
+ new_data_list.append(video_data_info)
+
+ print_log(
+ 'The number of samples before and after filtering: '
+ f'{num_imgs_before_filter} / {num_imgs_after_filter}', 'current')
+ return new_data_list
+
+ def prepare_data(self, idx) -> Any:
+ """Get date processed by ``self.pipeline``. Note that ``idx`` is a
+ video index in default since the base element of video dataset is a
+ video. However, in some cases, we need to specific both the video index
+ and frame index. For example, in traing mode, we may want to sample the
+ specific frames and all the frames must be sampled once in a epoch; in
+ test mode, we may want to output data of a single image rather than the
+ whole video for saving memory.
+
+ Args:
+ idx (int): The index of ``data_info``.
+
+ Returns:
+ Any: Depends on ``self.pipeline``.
+ """
+ if isinstance(idx, tuple):
+ assert len(idx) == 2, 'The length of idx must be 2: '
+ '(video_index, frame_index)'
+ video_idx, frame_idx = idx[0], idx[1]
+ else:
+ video_idx, frame_idx = idx, None
+
+ data_info = self.get_data_info(video_idx)
+ if self.test_mode:
+ # Support two test_mode: frame-level and video-level
+ final_data_info = defaultdict(list)
+ if frame_idx is None:
+ frames_idx_list = list(range(data_info['video_length']))
+ else:
+ frames_idx_list = [frame_idx]
+ for index in frames_idx_list:
+ frame_ann = data_info['images'][index]
+ frame_ann['video_id'] = data_info['video_id']
+ # Collate data_list (list of dict to dict of list)
+ for key, value in frame_ann.items():
+ final_data_info[key].append(value)
+ # copy the info in video-level into img-level
+ # TODO: the value of this key is the same as that of
+ # `video_length` in test mode
+ final_data_info['ori_video_length'].append(
+ data_info['video_length'])
+
+ final_data_info['video_length'] = [len(frames_idx_list)
+ ] * len(frames_idx_list)
+ return self.pipeline(final_data_info)
+ else:
+ # Specify `key_frame_id` for the frame sampling in the pipeline
+ if frame_idx is not None:
+ data_info['key_frame_id'] = frame_idx
+ return self.pipeline(data_info)
+
+ def get_cat_ids(self, index) -> List[int]:
+ """Following image detection, we provide this interface function. Get
+ category ids by video index and frame index.
+
+ Args:
+ index: The index of the dataset. It support two kinds of inputs:
+ Tuple:
+ video_idx (int): Index of video.
+ frame_idx (int): Index of frame.
+ Int: Index of video.
+
+ Returns:
+ List[int]: All categories in the image of specified video index
+ and frame index.
+ """
+ if isinstance(index, tuple):
+ assert len(
+ index
+ ) == 2, f'Expect the length of index is 2, but got {len(index)}'
+ video_idx, frame_idx = index
+ instances = self.get_data_info(
+ video_idx)['images'][frame_idx]['instances']
+ return [instance['bbox_label'] for instance in instances]
+ else:
+ cat_ids = []
+ for img in self.get_data_info(index)['images']:
+ for instance in img['instances']:
+ cat_ids.append(instance['bbox_label'])
+ return cat_ids
+
+ @property
+ def num_all_imgs(self):
+ """Get the number of all the images in this video dataset."""
+ return sum(
+ [len(self.get_data_info(i)['images']) for i in range(len(self))])
+
+ def get_len_per_video(self, idx):
+ """Get length of one video.
+
+ Args:
+ idx (int): Index of video.
+
+ Returns:
+ int (int): The length of the video.
+ """
+ return len(self.get_data_info(idx)['images'])
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/cityscapes.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/cityscapes.py
new file mode 100644
index 0000000000000000000000000000000000000000..09755eb1e8b0f0c278085bd2fafbb7247a3fc946
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/cityscapes.py
@@ -0,0 +1,61 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+# Modified from https://github.com/facebookresearch/detectron2/blob/master/detectron2/data/datasets/cityscapes.py # noqa
+# and https://github.com/mcordts/cityscapesScripts/blob/master/cityscapesscripts/evaluation/evalInstanceLevelSemanticLabeling.py # noqa
+
+from typing import List
+
+from mmdet.registry import DATASETS
+from .coco import CocoDataset
+
+
+@DATASETS.register_module()
+class CityscapesDataset(CocoDataset):
+ """Dataset for Cityscapes."""
+
+ METAINFO = {
+ 'classes': ('person', 'rider', 'car', 'truck', 'bus', 'train',
+ 'motorcycle', 'bicycle'),
+ 'palette': [(220, 20, 60), (255, 0, 0), (0, 0, 142), (0, 0, 70),
+ (0, 60, 100), (0, 80, 100), (0, 0, 230), (119, 11, 32)]
+ }
+
+ def filter_data(self) -> List[dict]:
+ """Filter annotations according to filter_cfg.
+
+ Returns:
+ List[dict]: Filtered results.
+ """
+ if self.test_mode:
+ return self.data_list
+
+ if self.filter_cfg is None:
+ return self.data_list
+
+ filter_empty_gt = self.filter_cfg.get('filter_empty_gt', False)
+ min_size = self.filter_cfg.get('min_size', 0)
+
+ # obtain images that contain annotation
+ ids_with_ann = set(data_info['img_id'] for data_info in self.data_list)
+ # obtain images that contain annotations of the required categories
+ ids_in_cat = set()
+ for i, class_id in enumerate(self.cat_ids):
+ ids_in_cat |= set(self.cat_img_map[class_id])
+ # merge the image id sets of the two conditions and use the merged set
+ # to filter out images if self.filter_empty_gt=True
+ ids_in_cat &= ids_with_ann
+
+ valid_data_infos = []
+ for i, data_info in enumerate(self.data_list):
+ img_id = data_info['img_id']
+ width = data_info['width']
+ height = data_info['height']
+ all_is_crowd = all([
+ instance['ignore_flag'] == 1
+ for instance in data_info['instances']
+ ])
+ if filter_empty_gt and (img_id not in ids_in_cat or all_is_crowd):
+ continue
+ if min(width, height) >= min_size:
+ valid_data_infos.append(data_info)
+
+ return valid_data_infos
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..ff40a3288591f011baaa134c2741cbe2b4b3d51e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/coco.py
@@ -0,0 +1,203 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import os.path as osp
+from typing import List, Union
+
+from mmengine.fileio import get_local_path
+
+from mmdet.registry import DATASETS
+from .api_wrappers import COCO
+from .base_det_dataset import BaseDetDataset
+
+
+@DATASETS.register_module()
+class CocoDataset(BaseDetDataset):
+ """Dataset for COCO."""
+
+ METAINFO = {
+ 'classes':
+ ('person', 'bicycle', 'car', 'motorcycle', 'airplane', 'bus', 'train',
+ 'truck', 'boat', 'traffic light', 'fire hydrant', 'stop sign',
+ 'parking meter', 'bench', 'bird', 'cat', 'dog', 'horse', 'sheep',
+ 'cow', 'elephant', 'bear', 'zebra', 'giraffe', 'backpack', 'umbrella',
+ 'handbag', 'tie', 'suitcase', 'frisbee', 'skis', 'snowboard',
+ 'sports ball', 'kite', 'baseball bat', 'baseball glove', 'skateboard',
+ 'surfboard', 'tennis racket', 'bottle', 'wine glass', 'cup', 'fork',
+ 'knife', 'spoon', 'bowl', 'banana', 'apple', 'sandwich', 'orange',
+ 'broccoli', 'carrot', 'hot dog', 'pizza', 'donut', 'cake', 'chair',
+ 'couch', 'potted plant', 'bed', 'dining table', 'toilet', 'tv',
+ 'laptop', 'mouse', 'remote', 'keyboard', 'cell phone', 'microwave',
+ 'oven', 'toaster', 'sink', 'refrigerator', 'book', 'clock', 'vase',
+ 'scissors', 'teddy bear', 'hair drier', 'toothbrush'),
+ # palette is a list of color tuples, which is used for visualization.
+ 'palette':
+ [(220, 20, 60), (119, 11, 32), (0, 0, 142), (0, 0, 230), (106, 0, 228),
+ (0, 60, 100), (0, 80, 100), (0, 0, 70), (0, 0, 192), (250, 170, 30),
+ (100, 170, 30), (220, 220, 0), (175, 116, 175), (250, 0, 30),
+ (165, 42, 42), (255, 77, 255), (0, 226, 252), (182, 182, 255),
+ (0, 82, 0), (120, 166, 157), (110, 76, 0), (174, 57, 255),
+ (199, 100, 0), (72, 0, 118), (255, 179, 240), (0, 125, 92),
+ (209, 0, 151), (188, 208, 182), (0, 220, 176), (255, 99, 164),
+ (92, 0, 73), (133, 129, 255), (78, 180, 255), (0, 228, 0),
+ (174, 255, 243), (45, 89, 255), (134, 134, 103), (145, 148, 174),
+ (255, 208, 186), (197, 226, 255), (171, 134, 1), (109, 63, 54),
+ (207, 138, 255), (151, 0, 95), (9, 80, 61), (84, 105, 51),
+ (74, 65, 105), (166, 196, 102), (208, 195, 210), (255, 109, 65),
+ (0, 143, 149), (179, 0, 194), (209, 99, 106), (5, 121, 0),
+ (227, 255, 205), (147, 186, 208), (153, 69, 1), (3, 95, 161),
+ (163, 255, 0), (119, 0, 170), (0, 182, 199), (0, 165, 120),
+ (183, 130, 88), (95, 32, 0), (130, 114, 135), (110, 129, 133),
+ (166, 74, 118), (219, 142, 185), (79, 210, 114), (178, 90, 62),
+ (65, 70, 15), (127, 167, 115), (59, 105, 106), (142, 108, 45),
+ (196, 172, 0), (95, 54, 80), (128, 76, 255), (201, 57, 1),
+ (246, 0, 122), (191, 162, 208)]
+ }
+ COCOAPI = COCO
+ # ann_id is unique in coco dataset.
+ ANN_ID_UNIQUE = True
+
+ def load_data_list(self) -> List[dict]:
+ """Load annotations from an annotation file named as ``self.ann_file``
+
+ Returns:
+ List[dict]: A list of annotation.
+ """ # noqa: E501
+ with get_local_path(
+ self.ann_file, backend_args=self.backend_args) as local_path:
+ self.coco = self.COCOAPI(local_path)
+ # The order of returned `cat_ids` will not
+ # change with the order of the `classes`
+ self.cat_ids = self.coco.get_cat_ids(
+ cat_names=self.metainfo['classes'])
+ self.cat2label = {cat_id: i for i, cat_id in enumerate(self.cat_ids)}
+ self.cat_img_map = copy.deepcopy(self.coco.cat_img_map)
+
+ img_ids = self.coco.get_img_ids()
+ data_list = []
+ total_ann_ids = []
+ for img_id in img_ids:
+ raw_img_info = self.coco.load_imgs([img_id])[0]
+ raw_img_info['img_id'] = img_id
+
+ ann_ids = self.coco.get_ann_ids(img_ids=[img_id])
+ raw_ann_info = self.coco.load_anns(ann_ids)
+ total_ann_ids.extend(ann_ids)
+
+ parsed_data_info = self.parse_data_info({
+ 'raw_ann_info':
+ raw_ann_info,
+ 'raw_img_info':
+ raw_img_info
+ })
+ data_list.append(parsed_data_info)
+ if self.ANN_ID_UNIQUE:
+ assert len(set(total_ann_ids)) == len(
+ total_ann_ids
+ ), f"Annotation ids in '{self.ann_file}' are not unique!"
+
+ del self.coco
+
+ return data_list
+
+ def parse_data_info(self, raw_data_info: dict) -> Union[dict, List[dict]]:
+ """Parse raw annotation to target format.
+
+ Args:
+ raw_data_info (dict): Raw data information load from ``ann_file``
+
+ Returns:
+ Union[dict, List[dict]]: Parsed annotation.
+ """
+ img_info = raw_data_info['raw_img_info']
+ ann_info = raw_data_info['raw_ann_info']
+
+ data_info = {}
+
+ # TODO: need to change data_prefix['img'] to data_prefix['img_path']
+ img_path = osp.join(self.data_prefix['img'], img_info['file_name'])
+ if self.data_prefix.get('seg', None):
+ seg_map_path = osp.join(
+ self.data_prefix['seg'],
+ img_info['file_name'].rsplit('.', 1)[0] + self.seg_map_suffix)
+ else:
+ seg_map_path = None
+ data_info['img_path'] = img_path
+ data_info['img_id'] = img_info['img_id']
+ data_info['seg_map_path'] = seg_map_path
+ data_info['height'] = img_info['height']
+ data_info['width'] = img_info['width']
+
+ if self.return_classes:
+ data_info['text'] = self.metainfo['classes']
+ data_info['caption_prompt'] = self.caption_prompt
+ data_info['custom_entities'] = True
+
+ instances = []
+ for i, ann in enumerate(ann_info):
+ instance = {}
+
+ if ann.get('ignore', False):
+ continue
+ x1, y1, w, h = ann['bbox']
+ inter_w = max(0, min(x1 + w, img_info['width']) - max(x1, 0))
+ inter_h = max(0, min(y1 + h, img_info['height']) - max(y1, 0))
+ if inter_w * inter_h == 0:
+ continue
+ if ann['area'] <= 0 or w < 1 or h < 1:
+ continue
+ if ann['category_id'] not in self.cat_ids:
+ continue
+ bbox = [x1, y1, x1 + w, y1 + h]
+
+ if ann.get('iscrowd', False):
+ instance['ignore_flag'] = 1
+ else:
+ instance['ignore_flag'] = 0
+ instance['bbox'] = bbox
+ instance['bbox_label'] = self.cat2label[ann['category_id']]
+
+ if ann.get('segmentation', None):
+ instance['mask'] = ann['segmentation']
+
+ instances.append(instance)
+ data_info['instances'] = instances
+ return data_info
+
+ def filter_data(self) -> List[dict]:
+ """Filter annotations according to filter_cfg.
+
+ Returns:
+ List[dict]: Filtered results.
+ """
+ if self.test_mode:
+ self.data_list = list(filter(lambda x: osp.exists(x['img_path']),
+ self.data_list))
+ return self.data_list
+
+ if self.filter_cfg is None:
+ return self.data_list
+
+ filter_empty_gt = self.filter_cfg.get('filter_empty_gt', False)
+ min_size = self.filter_cfg.get('min_size', 0)
+
+ # obtain images that contain annotation
+ ids_with_ann = set(data_info['img_id'] for data_info in self.data_list)
+ # obtain images that contain annotations of the required categories
+ ids_in_cat = set()
+ for i, class_id in enumerate(self.cat_ids):
+ ids_in_cat |= set(self.cat_img_map[class_id])
+ # merge the image id sets of the two conditions and use the merged set
+ # to filter out images if self.filter_empty_gt=True
+ ids_in_cat &= ids_with_ann
+
+ valid_data_infos = []
+ for i, data_info in enumerate(self.data_list):
+ img_id = data_info['img_id']
+ width = data_info['width']
+ height = data_info['height']
+ if filter_empty_gt and img_id not in ids_in_cat:
+ continue
+ if min(width, height) >= min_size:
+ valid_data_infos.append(data_info)
+
+ return valid_data_infos
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/coco.py.bak b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/coco.py.bak
new file mode 100644
index 0000000000000000000000000000000000000000..1cf21c4e667e3b565ea01d1eb95bcdbf171b90d0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/coco.py.bak
@@ -0,0 +1,201 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import os.path as osp
+from typing import List, Union
+
+from mmengine.fileio import get_local_path
+
+from mmdet.registry import DATASETS
+from .api_wrappers import COCO
+from .base_det_dataset import BaseDetDataset
+
+
+@DATASETS.register_module()
+class CocoDataset(BaseDetDataset):
+ """Dataset for COCO."""
+
+ METAINFO = {
+ 'classes':
+ ('person', 'bicycle', 'car', 'motorcycle', 'airplane', 'bus', 'train',
+ 'truck', 'boat', 'traffic light', 'fire hydrant', 'stop sign',
+ 'parking meter', 'bench', 'bird', 'cat', 'dog', 'horse', 'sheep',
+ 'cow', 'elephant', 'bear', 'zebra', 'giraffe', 'backpack', 'umbrella',
+ 'handbag', 'tie', 'suitcase', 'frisbee', 'skis', 'snowboard',
+ 'sports ball', 'kite', 'baseball bat', 'baseball glove', 'skateboard',
+ 'surfboard', 'tennis racket', 'bottle', 'wine glass', 'cup', 'fork',
+ 'knife', 'spoon', 'bowl', 'banana', 'apple', 'sandwich', 'orange',
+ 'broccoli', 'carrot', 'hot dog', 'pizza', 'donut', 'cake', 'chair',
+ 'couch', 'potted plant', 'bed', 'dining table', 'toilet', 'tv',
+ 'laptop', 'mouse', 'remote', 'keyboard', 'cell phone', 'microwave',
+ 'oven', 'toaster', 'sink', 'refrigerator', 'book', 'clock', 'vase',
+ 'scissors', 'teddy bear', 'hair drier', 'toothbrush'),
+ # palette is a list of color tuples, which is used for visualization.
+ 'palette':
+ [(220, 20, 60), (119, 11, 32), (0, 0, 142), (0, 0, 230), (106, 0, 228),
+ (0, 60, 100), (0, 80, 100), (0, 0, 70), (0, 0, 192), (250, 170, 30),
+ (100, 170, 30), (220, 220, 0), (175, 116, 175), (250, 0, 30),
+ (165, 42, 42), (255, 77, 255), (0, 226, 252), (182, 182, 255),
+ (0, 82, 0), (120, 166, 157), (110, 76, 0), (174, 57, 255),
+ (199, 100, 0), (72, 0, 118), (255, 179, 240), (0, 125, 92),
+ (209, 0, 151), (188, 208, 182), (0, 220, 176), (255, 99, 164),
+ (92, 0, 73), (133, 129, 255), (78, 180, 255), (0, 228, 0),
+ (174, 255, 243), (45, 89, 255), (134, 134, 103), (145, 148, 174),
+ (255, 208, 186), (197, 226, 255), (171, 134, 1), (109, 63, 54),
+ (207, 138, 255), (151, 0, 95), (9, 80, 61), (84, 105, 51),
+ (74, 65, 105), (166, 196, 102), (208, 195, 210), (255, 109, 65),
+ (0, 143, 149), (179, 0, 194), (209, 99, 106), (5, 121, 0),
+ (227, 255, 205), (147, 186, 208), (153, 69, 1), (3, 95, 161),
+ (163, 255, 0), (119, 0, 170), (0, 182, 199), (0, 165, 120),
+ (183, 130, 88), (95, 32, 0), (130, 114, 135), (110, 129, 133),
+ (166, 74, 118), (219, 142, 185), (79, 210, 114), (178, 90, 62),
+ (65, 70, 15), (127, 167, 115), (59, 105, 106), (142, 108, 45),
+ (196, 172, 0), (95, 54, 80), (128, 76, 255), (201, 57, 1),
+ (246, 0, 122), (191, 162, 208)]
+ }
+ COCOAPI = COCO
+ # ann_id is unique in coco dataset.
+ ANN_ID_UNIQUE = True
+
+ def load_data_list(self) -> List[dict]:
+ """Load annotations from an annotation file named as ``self.ann_file``
+
+ Returns:
+ List[dict]: A list of annotation.
+ """ # noqa: E501
+ with get_local_path(
+ self.ann_file, backend_args=self.backend_args) as local_path:
+ self.coco = self.COCOAPI(local_path)
+ # The order of returned `cat_ids` will not
+ # change with the order of the `classes`
+ self.cat_ids = self.coco.get_cat_ids(
+ cat_names=self.metainfo['classes'])
+ self.cat2label = {cat_id: i for i, cat_id in enumerate(self.cat_ids)}
+ self.cat_img_map = copy.deepcopy(self.coco.cat_img_map)
+
+ img_ids = self.coco.get_img_ids()
+ data_list = []
+ total_ann_ids = []
+ for img_id in img_ids:
+ raw_img_info = self.coco.load_imgs([img_id])[0]
+ raw_img_info['img_id'] = img_id
+
+ ann_ids = self.coco.get_ann_ids(img_ids=[img_id])
+ raw_ann_info = self.coco.load_anns(ann_ids)
+ total_ann_ids.extend(ann_ids)
+
+ parsed_data_info = self.parse_data_info({
+ 'raw_ann_info':
+ raw_ann_info,
+ 'raw_img_info':
+ raw_img_info
+ })
+ data_list.append(parsed_data_info)
+ if self.ANN_ID_UNIQUE:
+ assert len(set(total_ann_ids)) == len(
+ total_ann_ids
+ ), f"Annotation ids in '{self.ann_file}' are not unique!"
+
+ del self.coco
+
+ return data_list
+
+ def parse_data_info(self, raw_data_info: dict) -> Union[dict, List[dict]]:
+ """Parse raw annotation to target format.
+
+ Args:
+ raw_data_info (dict): Raw data information load from ``ann_file``
+
+ Returns:
+ Union[dict, List[dict]]: Parsed annotation.
+ """
+ img_info = raw_data_info['raw_img_info']
+ ann_info = raw_data_info['raw_ann_info']
+
+ data_info = {}
+
+ # TODO: need to change data_prefix['img'] to data_prefix['img_path']
+ img_path = osp.join(self.data_prefix['img'], img_info['file_name'])
+ if self.data_prefix.get('seg', None):
+ seg_map_path = osp.join(
+ self.data_prefix['seg'],
+ img_info['file_name'].rsplit('.', 1)[0] + self.seg_map_suffix)
+ else:
+ seg_map_path = None
+ data_info['img_path'] = img_path
+ data_info['img_id'] = img_info['img_id']
+ data_info['seg_map_path'] = seg_map_path
+ data_info['height'] = img_info['height']
+ data_info['width'] = img_info['width']
+
+ if self.return_classes:
+ data_info['text'] = self.metainfo['classes']
+ data_info['caption_prompt'] = self.caption_prompt
+ data_info['custom_entities'] = True
+
+ instances = []
+ for i, ann in enumerate(ann_info):
+ instance = {}
+
+ if ann.get('ignore', False):
+ continue
+ x1, y1, w, h = ann['bbox']
+ inter_w = max(0, min(x1 + w, img_info['width']) - max(x1, 0))
+ inter_h = max(0, min(y1 + h, img_info['height']) - max(y1, 0))
+ if inter_w * inter_h == 0:
+ continue
+ if ann['area'] <= 0 or w < 1 or h < 1:
+ continue
+ if ann['category_id'] not in self.cat_ids:
+ continue
+ bbox = [x1, y1, x1 + w, y1 + h]
+
+ if ann.get('iscrowd', False):
+ instance['ignore_flag'] = 1
+ else:
+ instance['ignore_flag'] = 0
+ instance['bbox'] = bbox
+ instance['bbox_label'] = self.cat2label[ann['category_id']]
+
+ if ann.get('segmentation', None):
+ instance['mask'] = ann['segmentation']
+
+ instances.append(instance)
+ data_info['instances'] = instances
+ return data_info
+
+ def filter_data(self) -> List[dict]:
+ """Filter annotations according to filter_cfg.
+
+ Returns:
+ List[dict]: Filtered results.
+ """
+ if self.test_mode:
+ return self.data_list
+
+ if self.filter_cfg is None:
+ return self.data_list
+
+ filter_empty_gt = self.filter_cfg.get('filter_empty_gt', False)
+ min_size = self.filter_cfg.get('min_size', 0)
+
+ # obtain images that contain annotation
+ ids_with_ann = set(data_info['img_id'] for data_info in self.data_list)
+ # obtain images that contain annotations of the required categories
+ ids_in_cat = set()
+ for i, class_id in enumerate(self.cat_ids):
+ ids_in_cat |= set(self.cat_img_map[class_id])
+ # merge the image id sets of the two conditions and use the merged set
+ # to filter out images if self.filter_empty_gt=True
+ ids_in_cat &= ids_with_ann
+
+ valid_data_infos = []
+ for i, data_info in enumerate(self.data_list):
+ img_id = data_info['img_id']
+ width = data_info['width']
+ height = data_info['height']
+ if filter_empty_gt and img_id not in ids_in_cat:
+ continue
+ if min(width, height) >= min_size:
+ valid_data_infos.append(data_info)
+
+ return valid_data_infos
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/coco_caption.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/coco_caption.py
new file mode 100644
index 0000000000000000000000000000000000000000..ee695fe9a768f2be5345c6ad6bafc74177f252c0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/coco_caption.py
@@ -0,0 +1,32 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from pathlib import Path
+from typing import List
+
+import mmengine
+from mmengine.dataset import BaseDataset
+from mmengine.fileio import get_file_backend
+
+from mmdet.registry import DATASETS
+
+
+@DATASETS.register_module()
+class CocoCaptionDataset(BaseDataset):
+ """COCO2014 Caption dataset."""
+
+ def load_data_list(self) -> List[dict]:
+ """Load data list."""
+ img_prefix = self.data_prefix['img_path']
+ annotations = mmengine.load(self.ann_file)
+ file_backend = get_file_backend(img_prefix)
+
+ data_list = []
+ for ann in annotations:
+ data_info = {
+ 'img_id': Path(ann['image']).stem.split('_')[-1],
+ 'img_path': file_backend.join_path(img_prefix, ann['image']),
+ 'gt_caption': ann['caption'],
+ }
+
+ data_list.append(data_info)
+
+ return data_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/coco_panoptic.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/coco_panoptic.py
new file mode 100644
index 0000000000000000000000000000000000000000..b7a200e01d323e998afa782797e1cc92f75c70cf
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/coco_panoptic.py
@@ -0,0 +1,292 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import os.path as osp
+from typing import Callable, List, Optional, Sequence, Union
+
+from mmdet.registry import DATASETS
+from .api_wrappers import COCOPanoptic
+from .coco import CocoDataset
+
+
+@DATASETS.register_module()
+class CocoPanopticDataset(CocoDataset):
+ """Coco dataset for Panoptic segmentation.
+
+ The annotation format is shown as follows. The `ann` field is optional
+ for testing.
+
+ .. code-block:: none
+
+ [
+ {
+ 'filename': f'{image_id:012}.png',
+ 'image_id':9
+ 'segments_info':
+ [
+ {
+ 'id': 8345037, (segment_id in panoptic png,
+ convert from rgb)
+ 'category_id': 51,
+ 'iscrowd': 0,
+ 'bbox': (x1, y1, w, h),
+ 'area': 24315
+ },
+ ...
+ ]
+ },
+ ...
+ ]
+
+ Args:
+ ann_file (str): Annotation file path. Defaults to ''.
+ metainfo (dict, optional): Meta information for dataset, such as class
+ information. Defaults to None.
+ data_root (str, optional): The root directory for ``data_prefix`` and
+ ``ann_file``. Defaults to None.
+ data_prefix (dict, optional): Prefix for training data. Defaults to
+ ``dict(img=None, ann=None, seg=None)``. The prefix ``seg`` which is
+ for panoptic segmentation map must be not None.
+ filter_cfg (dict, optional): Config for filter data. Defaults to None.
+ indices (int or Sequence[int], optional): Support using first few
+ data in annotation file to facilitate training/testing on a smaller
+ dataset. Defaults to None which means using all ``data_infos``.
+ serialize_data (bool, optional): Whether to hold memory using
+ serialized objects, when enabled, data loader workers can use
+ shared RAM from master process instead of making a copy. Defaults
+ to True.
+ pipeline (list, optional): Processing pipeline. Defaults to [].
+ test_mode (bool, optional): ``test_mode=True`` means in test phase.
+ Defaults to False.
+ lazy_init (bool, optional): Whether to load annotation during
+ instantiation. In some cases, such as visualization, only the meta
+ information of the dataset is needed, which is not necessary to
+ load annotation file. ``Basedataset`` can skip load annotations to
+ save time by set ``lazy_init=False``. Defaults to False.
+ max_refetch (int, optional): If ``Basedataset.prepare_data`` get a
+ None img. The maximum extra number of cycles to get a valid
+ image. Defaults to 1000.
+ """
+
+ METAINFO = {
+ 'classes':
+ ('person', 'bicycle', 'car', 'motorcycle', 'airplane', 'bus', 'train',
+ 'truck', 'boat', 'traffic light', 'fire hydrant', 'stop sign',
+ 'parking meter', 'bench', 'bird', 'cat', 'dog', 'horse', 'sheep',
+ 'cow', 'elephant', 'bear', 'zebra', 'giraffe', 'backpack', 'umbrella',
+ 'handbag', 'tie', 'suitcase', 'frisbee', 'skis', 'snowboard',
+ 'sports ball', 'kite', 'baseball bat', 'baseball glove', 'skateboard',
+ 'surfboard', 'tennis racket', 'bottle', 'wine glass', 'cup', 'fork',
+ 'knife', 'spoon', 'bowl', 'banana', 'apple', 'sandwich', 'orange',
+ 'broccoli', 'carrot', 'hot dog', 'pizza', 'donut', 'cake', 'chair',
+ 'couch', 'potted plant', 'bed', 'dining table', 'toilet', 'tv',
+ 'laptop', 'mouse', 'remote', 'keyboard', 'cell phone', 'microwave',
+ 'oven', 'toaster', 'sink', 'refrigerator', 'book', 'clock', 'vase',
+ 'scissors', 'teddy bear', 'hair drier', 'toothbrush', 'banner',
+ 'blanket', 'bridge', 'cardboard', 'counter', 'curtain', 'door-stuff',
+ 'floor-wood', 'flower', 'fruit', 'gravel', 'house', 'light',
+ 'mirror-stuff', 'net', 'pillow', 'platform', 'playingfield',
+ 'railroad', 'river', 'road', 'roof', 'sand', 'sea', 'shelf', 'snow',
+ 'stairs', 'tent', 'towel', 'wall-brick', 'wall-stone', 'wall-tile',
+ 'wall-wood', 'water-other', 'window-blind', 'window-other',
+ 'tree-merged', 'fence-merged', 'ceiling-merged', 'sky-other-merged',
+ 'cabinet-merged', 'table-merged', 'floor-other-merged',
+ 'pavement-merged', 'mountain-merged', 'grass-merged', 'dirt-merged',
+ 'paper-merged', 'food-other-merged', 'building-other-merged',
+ 'rock-merged', 'wall-other-merged', 'rug-merged'),
+ 'thing_classes':
+ ('person', 'bicycle', 'car', 'motorcycle', 'airplane', 'bus', 'train',
+ 'truck', 'boat', 'traffic light', 'fire hydrant', 'stop sign',
+ 'parking meter', 'bench', 'bird', 'cat', 'dog', 'horse', 'sheep',
+ 'cow', 'elephant', 'bear', 'zebra', 'giraffe', 'backpack', 'umbrella',
+ 'handbag', 'tie', 'suitcase', 'frisbee', 'skis', 'snowboard',
+ 'sports ball', 'kite', 'baseball bat', 'baseball glove', 'skateboard',
+ 'surfboard', 'tennis racket', 'bottle', 'wine glass', 'cup', 'fork',
+ 'knife', 'spoon', 'bowl', 'banana', 'apple', 'sandwich', 'orange',
+ 'broccoli', 'carrot', 'hot dog', 'pizza', 'donut', 'cake', 'chair',
+ 'couch', 'potted plant', 'bed', 'dining table', 'toilet', 'tv',
+ 'laptop', 'mouse', 'remote', 'keyboard', 'cell phone', 'microwave',
+ 'oven', 'toaster', 'sink', 'refrigerator', 'book', 'clock', 'vase',
+ 'scissors', 'teddy bear', 'hair drier', 'toothbrush'),
+ 'stuff_classes':
+ ('banner', 'blanket', 'bridge', 'cardboard', 'counter', 'curtain',
+ 'door-stuff', 'floor-wood', 'flower', 'fruit', 'gravel', 'house',
+ 'light', 'mirror-stuff', 'net', 'pillow', 'platform', 'playingfield',
+ 'railroad', 'river', 'road', 'roof', 'sand', 'sea', 'shelf', 'snow',
+ 'stairs', 'tent', 'towel', 'wall-brick', 'wall-stone', 'wall-tile',
+ 'wall-wood', 'water-other', 'window-blind', 'window-other',
+ 'tree-merged', 'fence-merged', 'ceiling-merged', 'sky-other-merged',
+ 'cabinet-merged', 'table-merged', 'floor-other-merged',
+ 'pavement-merged', 'mountain-merged', 'grass-merged', 'dirt-merged',
+ 'paper-merged', 'food-other-merged', 'building-other-merged',
+ 'rock-merged', 'wall-other-merged', 'rug-merged'),
+ 'palette':
+ [(220, 20, 60), (119, 11, 32), (0, 0, 142), (0, 0, 230), (106, 0, 228),
+ (0, 60, 100), (0, 80, 100), (0, 0, 70), (0, 0, 192), (250, 170, 30),
+ (100, 170, 30), (220, 220, 0), (175, 116, 175), (250, 0, 30),
+ (165, 42, 42), (255, 77, 255), (0, 226, 252), (182, 182, 255),
+ (0, 82, 0), (120, 166, 157), (110, 76, 0), (174, 57, 255),
+ (199, 100, 0), (72, 0, 118), (255, 179, 240), (0, 125, 92),
+ (209, 0, 151), (188, 208, 182), (0, 220, 176), (255, 99, 164),
+ (92, 0, 73), (133, 129, 255), (78, 180, 255), (0, 228, 0),
+ (174, 255, 243), (45, 89, 255), (134, 134, 103), (145, 148, 174),
+ (255, 208, 186), (197, 226, 255), (171, 134, 1), (109, 63, 54),
+ (207, 138, 255), (151, 0, 95), (9, 80, 61), (84, 105, 51),
+ (74, 65, 105), (166, 196, 102), (208, 195, 210), (255, 109, 65),
+ (0, 143, 149), (179, 0, 194), (209, 99, 106), (5, 121, 0),
+ (227, 255, 205), (147, 186, 208), (153, 69, 1), (3, 95, 161),
+ (163, 255, 0), (119, 0, 170), (0, 182, 199), (0, 165, 120),
+ (183, 130, 88), (95, 32, 0), (130, 114, 135), (110, 129, 133),
+ (166, 74, 118), (219, 142, 185), (79, 210, 114), (178, 90, 62),
+ (65, 70, 15), (127, 167, 115), (59, 105, 106), (142, 108, 45),
+ (196, 172, 0), (95, 54, 80), (128, 76, 255), (201, 57, 1),
+ (246, 0, 122), (191, 162, 208), (255, 255, 128), (147, 211, 203),
+ (150, 100, 100), (168, 171, 172), (146, 112, 198), (210, 170, 100),
+ (92, 136, 89), (218, 88, 184), (241, 129, 0), (217, 17, 255),
+ (124, 74, 181), (70, 70, 70), (255, 228, 255), (154, 208, 0),
+ (193, 0, 92), (76, 91, 113), (255, 180, 195), (106, 154, 176),
+ (230, 150, 140), (60, 143, 255), (128, 64, 128), (92, 82, 55),
+ (254, 212, 124), (73, 77, 174), (255, 160, 98), (255, 255, 255),
+ (104, 84, 109), (169, 164, 131), (225, 199, 255), (137, 54, 74),
+ (135, 158, 223), (7, 246, 231), (107, 255, 200), (58, 41, 149),
+ (183, 121, 142), (255, 73, 97), (107, 142, 35), (190, 153, 153),
+ (146, 139, 141), (70, 130, 180), (134, 199, 156), (209, 226, 140),
+ (96, 36, 108), (96, 96, 96), (64, 170, 64), (152, 251, 152),
+ (208, 229, 228), (206, 186, 171), (152, 161, 64), (116, 112, 0),
+ (0, 114, 143), (102, 102, 156), (250, 141, 255)]
+ }
+ COCOAPI = COCOPanoptic
+ # ann_id is not unique in coco panoptic dataset.
+ ANN_ID_UNIQUE = False
+
+ def __init__(self,
+ ann_file: str = '',
+ metainfo: Optional[dict] = None,
+ data_root: Optional[str] = None,
+ data_prefix: dict = dict(img=None, ann=None, seg=None),
+ filter_cfg: Optional[dict] = None,
+ indices: Optional[Union[int, Sequence[int]]] = None,
+ serialize_data: bool = True,
+ pipeline: List[Union[dict, Callable]] = [],
+ test_mode: bool = False,
+ lazy_init: bool = False,
+ max_refetch: int = 1000,
+ backend_args: dict = None,
+ **kwargs) -> None:
+ super().__init__(
+ ann_file=ann_file,
+ metainfo=metainfo,
+ data_root=data_root,
+ data_prefix=data_prefix,
+ filter_cfg=filter_cfg,
+ indices=indices,
+ serialize_data=serialize_data,
+ pipeline=pipeline,
+ test_mode=test_mode,
+ lazy_init=lazy_init,
+ max_refetch=max_refetch,
+ backend_args=backend_args,
+ **kwargs)
+
+ def parse_data_info(self, raw_data_info: dict) -> dict:
+ """Parse raw annotation to target format.
+
+ Args:
+ raw_data_info (dict): Raw data information load from ``ann_file``.
+
+ Returns:
+ dict: Parsed annotation.
+ """
+ img_info = raw_data_info['raw_img_info']
+ ann_info = raw_data_info['raw_ann_info']
+ # filter out unmatched annotations which have
+ # same segment_id but belong to other image
+ ann_info = [
+ ann for ann in ann_info if ann['image_id'] == img_info['img_id']
+ ]
+ data_info = {}
+
+ img_path = osp.join(self.data_prefix['img'], img_info['file_name'])
+ if self.data_prefix.get('seg', None):
+ seg_map_path = osp.join(
+ self.data_prefix['seg'],
+ img_info['file_name'].replace('.jpg', '.png'))
+ else:
+ seg_map_path = None
+ data_info['img_path'] = img_path
+ data_info['img_id'] = img_info['img_id']
+ data_info['seg_map_path'] = seg_map_path
+ data_info['height'] = img_info['height']
+ data_info['width'] = img_info['width']
+
+ if self.return_classes:
+ data_info['text'] = self.metainfo['thing_classes']
+ data_info['stuff_text'] = self.metainfo['stuff_classes']
+ data_info['custom_entities'] = True # no important
+
+ instances = []
+ segments_info = []
+ for ann in ann_info:
+ instance = {}
+ x1, y1, w, h = ann['bbox']
+ if ann['area'] <= 0 or w < 1 or h < 1:
+ continue
+ bbox = [x1, y1, x1 + w, y1 + h]
+ category_id = ann['category_id']
+ contiguous_cat_id = self.cat2label[category_id]
+
+ is_thing = self.coco.load_cats(ids=category_id)[0]['isthing']
+ if is_thing:
+ is_crowd = ann.get('iscrowd', False)
+ instance['bbox'] = bbox
+ instance['bbox_label'] = contiguous_cat_id
+ if not is_crowd:
+ instance['ignore_flag'] = 0
+ else:
+ instance['ignore_flag'] = 1
+ is_thing = False
+
+ segment_info = {
+ 'id': ann['id'],
+ 'category': contiguous_cat_id,
+ 'is_thing': is_thing
+ }
+ segments_info.append(segment_info)
+ if len(instance) > 0 and is_thing:
+ instances.append(instance)
+ data_info['instances'] = instances
+ data_info['segments_info'] = segments_info
+ return data_info
+
+ def filter_data(self) -> List[dict]:
+ """Filter images too small or without ground truth.
+
+ Returns:
+ List[dict]: ``self.data_list`` after filtering.
+ """
+ if self.test_mode:
+ return self.data_list
+
+ if self.filter_cfg is None:
+ return self.data_list
+
+ filter_empty_gt = self.filter_cfg.get('filter_empty_gt', False)
+ min_size = self.filter_cfg.get('min_size', 0)
+
+ ids_with_ann = set()
+ # check whether images have legal thing annotations.
+ for data_info in self.data_list:
+ for segment_info in data_info['segments_info']:
+ if not segment_info['is_thing']:
+ continue
+ ids_with_ann.add(data_info['img_id'])
+
+ valid_data_list = []
+ for data_info in self.data_list:
+ img_id = data_info['img_id']
+ width = data_info['width']
+ height = data_info['height']
+ if filter_empty_gt and img_id not in ids_with_ann:
+ continue
+ if min(width, height) >= min_size:
+ valid_data_list.append(data_info)
+
+ return valid_data_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/coco_semantic.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/coco_semantic.py
new file mode 100644
index 0000000000000000000000000000000000000000..752568454456c1e5edcb2a24c6c2b46f042cb334
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/coco_semantic.py
@@ -0,0 +1,90 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import DATASETS
+from .ade20k import ADE20KSegDataset
+
+
+@DATASETS.register_module()
+class CocoSegDataset(ADE20KSegDataset):
+ """COCO dataset.
+
+ In segmentation map annotation for COCO. The ``img_suffix`` is fixed to
+ '.jpg', and ``seg_map_suffix`` is fixed to '.png'.
+ """
+
+ METAINFO = dict(
+ classes=(
+ 'person', 'bicycle', 'car', 'motorcycle', 'airplane', 'bus',
+ 'train', 'truck', 'boat', 'traffic light', 'fire hydrant',
+ 'stop sign', 'parking meter', 'bench', 'bird', 'cat', 'dog',
+ 'horse', 'sheep', 'cow', 'elephant', 'bear', 'zebra', 'giraffe',
+ 'backpack', 'umbrella', 'handbag', 'tie', 'suitcase', 'frisbee',
+ 'skis', 'snowboard', 'sports ball', 'kite', 'baseball bat',
+ 'baseball glove', 'skateboard', 'surfboard', 'tennis racket',
+ 'bottle', 'wine glass', 'cup', 'fork', 'knife', 'spoon', 'bowl',
+ 'banana', 'apple', 'sandwich', 'orange', 'broccoli', 'carrot',
+ 'hot dog', 'pizza', 'donut', 'cake', 'chair', 'couch',
+ 'potted plant', 'bed', 'dining table', 'toilet', 'tv', 'laptop',
+ 'mouse', 'remote', 'keyboard', 'cell phone', 'microwave', 'oven',
+ 'toaster', 'sink', 'refrigerator', 'book', 'clock', 'vase',
+ 'scissors', 'teddy bear', 'hair drier', 'toothbrush', 'banner',
+ 'blanket', 'branch', 'bridge', 'building-other', 'bush', 'cabinet',
+ 'cage', 'cardboard', 'carpet', 'ceiling-other', 'ceiling-tile',
+ 'cloth', 'clothes', 'clouds', 'counter', 'cupboard', 'curtain',
+ 'desk-stuff', 'dirt', 'door-stuff', 'fence', 'floor-marble',
+ 'floor-other', 'floor-stone', 'floor-tile', 'floor-wood', 'flower',
+ 'fog', 'food-other', 'fruit', 'furniture-other', 'grass', 'gravel',
+ 'ground-other', 'hill', 'house', 'leaves', 'light', 'mat', 'metal',
+ 'mirror-stuff', 'moss', 'mountain', 'mud', 'napkin', 'net',
+ 'paper', 'pavement', 'pillow', 'plant-other', 'plastic',
+ 'platform', 'playingfield', 'railing', 'railroad', 'river', 'road',
+ 'rock', 'roof', 'rug', 'salad', 'sand', 'sea', 'shelf',
+ 'sky-other', 'skyscraper', 'snow', 'solid-other', 'stairs',
+ 'stone', 'straw', 'structural-other', 'table', 'tent',
+ 'textile-other', 'towel', 'tree', 'vegetable', 'wall-brick',
+ 'wall-concrete', 'wall-other', 'wall-panel', 'wall-stone',
+ 'wall-tile', 'wall-wood', 'water-other', 'waterdrops',
+ 'window-blind', 'window-other', 'wood'),
+ palette=[(120, 120, 120), (180, 120, 120), (6, 230, 230), (80, 50, 50),
+ (4, 200, 3), (120, 120, 80), (140, 140, 140), (204, 5, 255),
+ (230, 230, 230), (4, 250, 7), (224, 5, 255), (235, 255, 7),
+ (150, 5, 61), (120, 120, 70), (8, 255, 51), (255, 6, 82),
+ (143, 255, 140), (204, 255, 4), (255, 51, 7), (204, 70, 3),
+ (0, 102, 200), (61, 230, 250), (255, 6, 51), (11, 102, 255),
+ (255, 7, 71), (255, 9, 224), (9, 7, 230), (220, 220, 220),
+ (255, 9, 92), (112, 9, 255), (8, 255, 214), (7, 255, 224),
+ (255, 184, 6), (10, 255, 71), (255, 41, 10), (7, 255, 255),
+ (224, 255, 8), (102, 8, 255), (255, 61, 6), (255, 194, 7),
+ (255, 122, 8), (0, 255, 20), (255, 8, 41), (255, 5, 153),
+ (6, 51, 255), (235, 12, 255), (160, 150, 20), (0, 163, 255),
+ (140, 140, 140), (250, 10, 15), (20, 255, 0), (31, 255, 0),
+ (255, 31, 0), (255, 224, 0), (153, 255, 0), (0, 0, 255),
+ (255, 71, 0), (0, 235, 255), (0, 173, 255), (31, 0, 255),
+ (11, 200, 200), (255, 82, 0), (0, 255, 245), (0, 61, 255),
+ (0, 255, 112), (0, 255, 133), (255, 0, 0), (255, 163, 0),
+ (255, 102, 0), (194, 255, 0), (0, 143, 255), (51, 255, 0),
+ (0, 82, 255), (0, 255, 41), (0, 255, 173), (10, 0, 255),
+ (173, 255, 0), (0, 255, 153), (255, 92, 0), (255, 0, 255),
+ (255, 0, 245), (255, 0, 102), (255, 173, 0), (255, 0, 20),
+ (255, 184, 184), (0, 31, 255), (0, 255, 61), (0, 71, 255),
+ (255, 0, 204), (0, 255, 194), (0, 255, 82), (0, 10, 255),
+ (0, 112, 255), (51, 0, 255), (0, 194, 255), (0, 122, 255),
+ (0, 255, 163), (255, 153, 0), (0, 255, 10), (255, 112, 0),
+ (143, 255, 0), (82, 0, 255), (163, 255, 0), (255, 235, 0),
+ (8, 184, 170), (133, 0, 255), (0, 255, 92), (184, 0, 255),
+ (255, 0, 31), (0, 184, 255), (0, 214, 255), (255, 0, 112),
+ (92, 255, 0), (0, 224, 255), (112, 224, 255), (70, 184, 160),
+ (163, 0, 255), (153, 0, 255), (71, 255, 0), (255, 0, 163),
+ (255, 204, 0), (255, 0, 143), (0, 255, 235), (133, 255, 0),
+ (255, 0, 235), (245, 0, 255), (255, 0, 122), (255, 245, 0),
+ (10, 190, 212), (214, 255, 0), (0, 204, 255), (20, 0, 255),
+ (255, 255, 0), (0, 153, 255), (0, 41, 255), (0, 255, 204),
+ (41, 0, 255), (41, 255, 0), (173, 0, 255), (0, 245, 255),
+ (71, 0, 255), (122, 0, 255), (0, 255, 184), (0, 92, 255),
+ (184, 255, 0), (0, 133, 255), (255, 214, 0), (25, 194, 194),
+ (102, 255, 0), (92, 0, 255), (107, 255, 200), (58, 41, 149),
+ (183, 121, 142), (255, 73, 97), (107, 142, 35),
+ (190, 153, 153), (146, 139, 141), (70, 130, 180),
+ (134, 199, 156), (209, 226, 140), (96, 36, 108), (96, 96, 96),
+ (64, 170, 64), (152, 251, 152), (208, 229, 228),
+ (206, 186, 171), (152, 161, 64), (116, 112, 0), (0, 114, 143),
+ (102, 102, 156), (250, 141, 255)])
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/crowdhuman.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/crowdhuman.py
new file mode 100644
index 0000000000000000000000000000000000000000..650176ee545ba6a10a816517553b3b77718d945b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/crowdhuman.py
@@ -0,0 +1,159 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import json
+import logging
+import os.path as osp
+import warnings
+from typing import List, Union
+
+import mmcv
+from mmengine.dist import get_rank
+from mmengine.fileio import dump, get, get_text, load
+from mmengine.logging import print_log
+from mmengine.utils import ProgressBar
+
+from mmdet.registry import DATASETS
+from .base_det_dataset import BaseDetDataset
+
+
+@DATASETS.register_module()
+class CrowdHumanDataset(BaseDetDataset):
+ r"""Dataset for CrowdHuman.
+
+ Args:
+ data_root (str): The root directory for
+ ``data_prefix`` and ``ann_file``.
+ ann_file (str): Annotation file path.
+ extra_ann_file (str | optional):The path of extra image metas
+ for CrowdHuman. It can be created by CrowdHumanDataset
+ automatically or by tools/misc/get_crowdhuman_id_hw.py
+ manually. Defaults to None.
+ """
+
+ METAINFO = {
+ 'classes': ('person', ),
+ # palette is a list of color tuples, which is used for visualization.
+ 'palette': [(220, 20, 60)]
+ }
+
+ def __init__(self, data_root, ann_file, extra_ann_file=None, **kwargs):
+ # extra_ann_file record the size of each image. This file is
+ # automatically created when you first load the CrowdHuman
+ # dataset by mmdet.
+ if extra_ann_file is not None:
+ self.extra_ann_exist = True
+ self.extra_anns = load(extra_ann_file)
+ else:
+ ann_file_name = osp.basename(ann_file)
+ if 'train' in ann_file_name:
+ self.extra_ann_file = osp.join(data_root, 'id_hw_train.json')
+ elif 'val' in ann_file_name:
+ self.extra_ann_file = osp.join(data_root, 'id_hw_val.json')
+ self.extra_ann_exist = False
+ if not osp.isfile(self.extra_ann_file):
+ print_log(
+ 'extra_ann_file does not exist, prepare to collect '
+ 'image height and width...',
+ level=logging.INFO)
+ self.extra_anns = {}
+ else:
+ self.extra_ann_exist = True
+ self.extra_anns = load(self.extra_ann_file)
+ super().__init__(data_root=data_root, ann_file=ann_file, **kwargs)
+
+ def load_data_list(self) -> List[dict]:
+ """Load annotations from an annotation file named as ``self.ann_file``
+
+ Returns:
+ List[dict]: A list of annotation.
+ """ # noqa: E501
+ anno_strs = get_text(
+ self.ann_file, backend_args=self.backend_args).strip().split('\n')
+ print_log('loading CrowdHuman annotation...', level=logging.INFO)
+ data_list = []
+ prog_bar = ProgressBar(len(anno_strs))
+ for i, anno_str in enumerate(anno_strs):
+ anno_dict = json.loads(anno_str)
+ parsed_data_info = self.parse_data_info(anno_dict)
+ data_list.append(parsed_data_info)
+ prog_bar.update()
+ if not self.extra_ann_exist and get_rank() == 0:
+ # TODO: support file client
+ try:
+ dump(self.extra_anns, self.extra_ann_file, file_format='json')
+ except: # noqa
+ warnings.warn(
+ 'Cache files can not be saved automatically! To speed up'
+ 'loading the dataset, please manually generate the cache'
+ ' file by file tools/misc/get_crowdhuman_id_hw.py')
+
+ print_log(
+ f'\nsave extra_ann_file in {self.data_root}',
+ level=logging.INFO)
+
+ del self.extra_anns
+ print_log('\nDone', level=logging.INFO)
+ return data_list
+
+ def parse_data_info(self, raw_data_info: dict) -> Union[dict, List[dict]]:
+ """Parse raw annotation to target format.
+
+ Args:
+ raw_data_info (dict): Raw data information load from ``ann_file``
+
+ Returns:
+ Union[dict, List[dict]]: Parsed annotation.
+ """
+ data_info = {}
+ img_path = osp.join(self.data_prefix['img'],
+ f"{raw_data_info['ID']}.jpg")
+ data_info['img_path'] = img_path
+ data_info['img_id'] = raw_data_info['ID']
+
+ if not self.extra_ann_exist:
+ img_bytes = get(img_path, backend_args=self.backend_args)
+ img = mmcv.imfrombytes(img_bytes, backend='cv2')
+ data_info['height'], data_info['width'] = img.shape[:2]
+ self.extra_anns[raw_data_info['ID']] = img.shape[:2]
+ del img, img_bytes
+ else:
+ data_info['height'], data_info['width'] = self.extra_anns[
+ raw_data_info['ID']]
+
+ instances = []
+ for i, ann in enumerate(raw_data_info['gtboxes']):
+ instance = {}
+ if ann['tag'] not in self.metainfo['classes']:
+ instance['bbox_label'] = -1
+ instance['ignore_flag'] = 1
+ else:
+ instance['bbox_label'] = self.metainfo['classes'].index(
+ ann['tag'])
+ instance['ignore_flag'] = 0
+ if 'extra' in ann:
+ if 'ignore' in ann['extra']:
+ if ann['extra']['ignore'] != 0:
+ instance['bbox_label'] = -1
+ instance['ignore_flag'] = 1
+
+ x1, y1, w, h = ann['fbox']
+ bbox = [x1, y1, x1 + w, y1 + h]
+ instance['bbox'] = bbox
+
+ # Record the full bbox(fbox), head bbox(hbox) and visible
+ # bbox(vbox) as additional information. If you need to use
+ # this information, you just need to design the pipeline
+ # instead of overriding the CrowdHumanDataset.
+ instance['fbox'] = bbox
+ hbox = ann['hbox']
+ instance['hbox'] = [
+ hbox[0], hbox[1], hbox[0] + hbox[2], hbox[1] + hbox[3]
+ ]
+ vbox = ann['vbox']
+ instance['vbox'] = [
+ vbox[0], vbox[1], vbox[0] + vbox[2], vbox[1] + vbox[3]
+ ]
+
+ instances.append(instance)
+
+ data_info['instances'] = instances
+ return data_info
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/dataset_wrappers.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/dataset_wrappers.py
new file mode 100644
index 0000000000000000000000000000000000000000..d4e26e07c0f8a9e9f106bcd351f71e7b24d6ccf9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/dataset_wrappers.py
@@ -0,0 +1,260 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import collections
+import copy
+from typing import List, Sequence, Union
+
+from mmengine.dataset import BaseDataset
+from mmengine.dataset import ConcatDataset as MMENGINE_ConcatDataset
+from mmengine.dataset import force_full_init
+
+from mmdet.registry import DATASETS, TRANSFORMS
+
+
+@DATASETS.register_module()
+class MultiImageMixDataset:
+ """A wrapper of multiple images mixed dataset.
+
+ Suitable for training on multiple images mixed data augmentation like
+ mosaic and mixup. For the augmentation pipeline of mixed image data,
+ the `get_indexes` method needs to be provided to obtain the image
+ indexes, and you can set `skip_flags` to change the pipeline running
+ process. At the same time, we provide the `dynamic_scale` parameter
+ to dynamically change the output image size.
+
+ Args:
+ dataset (:obj:`CustomDataset`): The dataset to be mixed.
+ pipeline (Sequence[dict]): Sequence of transform object or
+ config dict to be composed.
+ dynamic_scale (tuple[int], optional): The image scale can be changed
+ dynamically. Default to None. It is deprecated.
+ skip_type_keys (list[str], optional): Sequence of type string to
+ be skip pipeline. Default to None.
+ max_refetch (int): The maximum number of retry iterations for getting
+ valid results from the pipeline. If the number of iterations is
+ greater than `max_refetch`, but results is still None, then the
+ iteration is terminated and raise the error. Default: 15.
+ """
+
+ def __init__(self,
+ dataset: Union[BaseDataset, dict],
+ pipeline: Sequence[str],
+ skip_type_keys: Union[Sequence[str], None] = None,
+ max_refetch: int = 15,
+ lazy_init: bool = False) -> None:
+ assert isinstance(pipeline, collections.abc.Sequence)
+ if skip_type_keys is not None:
+ assert all([
+ isinstance(skip_type_key, str)
+ for skip_type_key in skip_type_keys
+ ])
+ self._skip_type_keys = skip_type_keys
+
+ self.pipeline = []
+ self.pipeline_types = []
+ for transform in pipeline:
+ if isinstance(transform, dict):
+ self.pipeline_types.append(transform['type'])
+ transform = TRANSFORMS.build(transform)
+ self.pipeline.append(transform)
+ else:
+ raise TypeError('pipeline must be a dict')
+
+ self.dataset: BaseDataset
+ if isinstance(dataset, dict):
+ self.dataset = DATASETS.build(dataset)
+ elif isinstance(dataset, BaseDataset):
+ self.dataset = dataset
+ else:
+ raise TypeError(
+ 'elements in datasets sequence should be config or '
+ f'`BaseDataset` instance, but got {type(dataset)}')
+
+ self._metainfo = self.dataset.metainfo
+ if hasattr(self.dataset, 'flag'):
+ self.flag = self.dataset.flag
+ self.num_samples = len(self.dataset)
+ self.max_refetch = max_refetch
+
+ self._fully_initialized = False
+ if not lazy_init:
+ self.full_init()
+
+ @property
+ def metainfo(self) -> dict:
+ """Get the meta information of the multi-image-mixed dataset.
+
+ Returns:
+ dict: The meta information of multi-image-mixed dataset.
+ """
+ return copy.deepcopy(self._metainfo)
+
+ def full_init(self):
+ """Loop to ``full_init`` each dataset."""
+ if self._fully_initialized:
+ return
+
+ self.dataset.full_init()
+ self._ori_len = len(self.dataset)
+ self._fully_initialized = True
+
+ @force_full_init
+ def get_data_info(self, idx: int) -> dict:
+ """Get annotation by index.
+
+ Args:
+ idx (int): Global index of ``ConcatDataset``.
+
+ Returns:
+ dict: The idx-th annotation of the datasets.
+ """
+ return self.dataset.get_data_info(idx)
+
+ @force_full_init
+ def __len__(self):
+ return self.num_samples
+
+ def __getitem__(self, idx):
+ results = copy.deepcopy(self.dataset[idx])
+ for (transform, transform_type) in zip(self.pipeline,
+ self.pipeline_types):
+ if self._skip_type_keys is not None and \
+ transform_type in self._skip_type_keys:
+ continue
+
+ if hasattr(transform, 'get_indexes'):
+ for i in range(self.max_refetch):
+ # Make sure the results passed the loading pipeline
+ # of the original dataset is not None.
+ indexes = transform.get_indexes(self.dataset)
+ if not isinstance(indexes, collections.abc.Sequence):
+ indexes = [indexes]
+ mix_results = [
+ copy.deepcopy(self.dataset[index]) for index in indexes
+ ]
+ if None not in mix_results:
+ results['mix_results'] = mix_results
+ break
+ else:
+ raise RuntimeError(
+ 'The loading pipeline of the original dataset'
+ ' always return None. Please check the correctness '
+ 'of the dataset and its pipeline.')
+
+ for i in range(self.max_refetch):
+ # To confirm the results passed the training pipeline
+ # of the wrapper is not None.
+ updated_results = transform(copy.deepcopy(results))
+ if updated_results is not None:
+ results = updated_results
+ break
+ else:
+ raise RuntimeError(
+ 'The training pipeline of the dataset wrapper'
+ ' always return None.Please check the correctness '
+ 'of the dataset and its pipeline.')
+
+ if 'mix_results' in results:
+ results.pop('mix_results')
+
+ return results
+
+ def update_skip_type_keys(self, skip_type_keys):
+ """Update skip_type_keys. It is called by an external hook.
+
+ Args:
+ skip_type_keys (list[str], optional): Sequence of type
+ string to be skip pipeline.
+ """
+ assert all([
+ isinstance(skip_type_key, str) for skip_type_key in skip_type_keys
+ ])
+ self._skip_type_keys = skip_type_keys
+
+
+@DATASETS.register_module()
+class ConcatDataset(MMENGINE_ConcatDataset):
+ """A wrapper of concatenated dataset.
+
+ Same as ``torch.utils.data.dataset.ConcatDataset``, support
+ lazy_init and get_dataset_source.
+
+ Note:
+ ``ConcatDataset`` should not inherit from ``BaseDataset`` since
+ ``get_subset`` and ``get_subset_`` could produce ambiguous meaning
+ sub-dataset which conflicts with original dataset. If you want to use
+ a sub-dataset of ``ConcatDataset``, you should set ``indices``
+ arguments for wrapped dataset which inherit from ``BaseDataset``.
+
+ Args:
+ datasets (Sequence[BaseDataset] or Sequence[dict]): A list of datasets
+ which will be concatenated.
+ lazy_init (bool, optional): Whether to load annotation during
+ instantiation. Defaults to False.
+ ignore_keys (List[str] or str): Ignore the keys that can be
+ unequal in `dataset.metainfo`. Defaults to None.
+ `New in version 0.3.0.`
+ """
+
+ def __init__(self,
+ datasets: Sequence[Union[BaseDataset, dict]],
+ lazy_init: bool = False,
+ ignore_keys: Union[str, List[str], None] = None):
+ self.datasets: List[BaseDataset] = []
+ for i, dataset in enumerate(datasets):
+ if isinstance(dataset, dict):
+ self.datasets.append(DATASETS.build(dataset))
+ elif isinstance(dataset, BaseDataset):
+ self.datasets.append(dataset)
+ else:
+ raise TypeError(
+ 'elements in datasets sequence should be config or '
+ f'`BaseDataset` instance, but got {type(dataset)}')
+ if ignore_keys is None:
+ self.ignore_keys = []
+ elif isinstance(ignore_keys, str):
+ self.ignore_keys = [ignore_keys]
+ elif isinstance(ignore_keys, list):
+ self.ignore_keys = ignore_keys
+ else:
+ raise TypeError('ignore_keys should be a list or str, '
+ f'but got {type(ignore_keys)}')
+
+ meta_keys: set = set()
+ for dataset in self.datasets:
+ meta_keys |= dataset.metainfo.keys()
+ # if the metainfo of multiple datasets are the same, use metainfo
+ # of the first dataset, else the metainfo is a list with metainfo
+ # of all the datasets
+ is_all_same = True
+ self._metainfo_first = self.datasets[0].metainfo
+ for i, dataset in enumerate(self.datasets, 1):
+ for key in meta_keys:
+ if key in self.ignore_keys:
+ continue
+ if key not in dataset.metainfo:
+ is_all_same = False
+ break
+ if self._metainfo_first[key] != dataset.metainfo[key]:
+ is_all_same = False
+ break
+
+ if is_all_same:
+ self._metainfo = self.datasets[0].metainfo
+ else:
+ self._metainfo = [dataset.metainfo for dataset in self.datasets]
+
+ self._fully_initialized = False
+ if not lazy_init:
+ self.full_init()
+
+ if is_all_same:
+ self._metainfo.update(
+ dict(cumulative_sizes=self.cumulative_sizes))
+ else:
+ for i, dataset in enumerate(self.datasets):
+ self._metainfo[i].update(
+ dict(cumulative_sizes=self.cumulative_sizes))
+
+ def get_dataset_source(self, idx: int) -> int:
+ dataset_idx, _ = self._get_ori_dataset_idx(idx)
+ return dataset_idx
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/deepfashion.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/deepfashion.py
new file mode 100644
index 0000000000000000000000000000000000000000..f853fc63398d598b90a88323e660ba6f4d81e2df
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/deepfashion.py
@@ -0,0 +1,19 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import DATASETS
+from .coco import CocoDataset
+
+
+@DATASETS.register_module()
+class DeepFashionDataset(CocoDataset):
+ """Dataset for DeepFashion."""
+
+ METAINFO = {
+ 'classes': ('top', 'skirt', 'leggings', 'dress', 'outer', 'pants',
+ 'bag', 'neckwear', 'headwear', 'eyeglass', 'belt',
+ 'footwear', 'hair', 'skin', 'face'),
+ # palette is a list of color tuples, which is used for visualization.
+ 'palette': [(0, 192, 64), (0, 64, 96), (128, 192, 192), (0, 64, 64),
+ (0, 192, 224), (0, 192, 192), (128, 192, 64), (0, 192, 96),
+ (128, 32, 192), (0, 0, 224), (0, 0, 64), (0, 160, 192),
+ (128, 0, 96), (128, 0, 192), (0, 32, 192)]
+ }
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/dod.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/dod.py
new file mode 100644
index 0000000000000000000000000000000000000000..152d32aaf70c7fb5e3730d46d26e150fc1204f22
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/dod.py
@@ -0,0 +1,78 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import os.path as osp
+from typing import List, Optional
+
+import numpy as np
+
+from mmdet.registry import DATASETS
+from .base_det_dataset import BaseDetDataset
+
+try:
+ from d_cube import D3
+except ImportError:
+ D3 = None
+from .api_wrappers import COCO
+
+
+@DATASETS.register_module()
+class DODDataset(BaseDetDataset):
+
+ def __init__(self,
+ *args,
+ data_root: Optional[str] = '',
+ data_prefix: dict = dict(img_path=''),
+ **kwargs) -> None:
+ if D3 is None:
+ raise ImportError(
+ 'Please install d3 by `pip install ddd-dataset`.')
+ pkl_anno_path = osp.join(data_root, data_prefix['anno'])
+ self.img_root = osp.join(data_root, data_prefix['img'])
+ self.d3 = D3(self.img_root, pkl_anno_path)
+
+ sent_infos = self.d3.load_sents()
+ classes = tuple([sent_info['raw_sent'] for sent_info in sent_infos])
+ super().__init__(
+ *args,
+ data_root=data_root,
+ data_prefix=data_prefix,
+ metainfo={'classes': classes},
+ **kwargs)
+
+ def load_data_list(self) -> List[dict]:
+ coco = COCO(self.ann_file)
+ data_list = []
+ img_ids = self.d3.get_img_ids()
+ for img_id in img_ids:
+ data_info = {}
+
+ img_info = self.d3.load_imgs(img_id)[0]
+ file_name = img_info['file_name']
+ img_path = osp.join(self.img_root, file_name)
+ data_info['img_path'] = img_path
+ data_info['img_id'] = img_id
+ data_info['height'] = img_info['height']
+ data_info['width'] = img_info['width']
+
+ group_ids = self.d3.get_group_ids(img_ids=[img_id])
+ sent_ids = self.d3.get_sent_ids(group_ids=group_ids)
+ sent_list = self.d3.load_sents(sent_ids=sent_ids)
+ text_list = [sent['raw_sent'] for sent in sent_list]
+ ann_ids = coco.get_ann_ids(img_ids=[img_id])
+ anno = coco.load_anns(ann_ids)
+
+ data_info['text'] = text_list
+ data_info['sent_ids'] = np.array([s for s in sent_ids])
+ data_info['custom_entities'] = True
+
+ instances = []
+ for i, ann in enumerate(anno):
+ instance = {}
+ x1, y1, w, h = ann['bbox']
+ bbox = [x1, y1, x1 + w, y1 + h]
+ instance['ignore_flag'] = 0
+ instance['bbox'] = bbox
+ instance['bbox_label'] = ann['category_id'] - 1
+ instances.append(instance)
+ data_info['instances'] = instances
+ data_list.append(data_info)
+ return data_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/dsdl.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/dsdl.py
new file mode 100644
index 0000000000000000000000000000000000000000..75570a2a6396e0e7a4ce5cac5dbf2a23cd164629
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/dsdl.py
@@ -0,0 +1,192 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import os
+from typing import List
+
+from mmdet.registry import DATASETS
+from .base_det_dataset import BaseDetDataset
+
+try:
+ from dsdl.dataset import DSDLDataset
+except ImportError:
+ DSDLDataset = None
+
+
+@DATASETS.register_module()
+class DSDLDetDataset(BaseDetDataset):
+ """Dataset for dsdl detection.
+
+ Args:
+ with_bbox(bool): Load bbox or not, defaults to be True.
+ with_polygon(bool): Load polygon or not, defaults to be False.
+ with_mask(bool): Load seg map mask or not, defaults to be False.
+ with_imagelevel_label(bool): Load image level label or not,
+ defaults to be False.
+ with_hierarchy(bool): Load hierarchy information or not,
+ defaults to be False.
+ specific_key_path(dict): Path of specific key which can not
+ be loaded by it's field name.
+ pre_transform(dict): pre-transform functions before loading.
+ """
+
+ METAINFO = {}
+
+ def __init__(self,
+ with_bbox: bool = True,
+ with_polygon: bool = False,
+ with_mask: bool = False,
+ with_imagelevel_label: bool = False,
+ with_hierarchy: bool = False,
+ specific_key_path: dict = {},
+ pre_transform: dict = {},
+ **kwargs) -> None:
+
+ if DSDLDataset is None:
+ raise RuntimeError(
+ 'Package dsdl is not installed. Please run "pip install dsdl".'
+ )
+
+ self.with_hierarchy = with_hierarchy
+ self.specific_key_path = specific_key_path
+
+ loc_config = dict(type='LocalFileReader', working_dir='')
+ if kwargs.get('data_root'):
+ kwargs['ann_file'] = os.path.join(kwargs['data_root'],
+ kwargs['ann_file'])
+ self.required_fields = ['Image', 'ImageShape', 'Label', 'ignore_flag']
+ if with_bbox:
+ self.required_fields.append('Bbox')
+ if with_polygon:
+ self.required_fields.append('Polygon')
+ if with_mask:
+ self.required_fields.append('LabelMap')
+ if with_imagelevel_label:
+ self.required_fields.append('image_level_labels')
+ assert 'image_level_labels' in specific_key_path.keys(
+ ), '`image_level_labels` not specified in `specific_key_path` !'
+
+ self.extra_keys = [
+ key for key in self.specific_key_path.keys()
+ if key not in self.required_fields
+ ]
+
+ self.dsdldataset = DSDLDataset(
+ dsdl_yaml=kwargs['ann_file'],
+ location_config=loc_config,
+ required_fields=self.required_fields,
+ specific_key_path=specific_key_path,
+ transform=pre_transform,
+ )
+
+ BaseDetDataset.__init__(self, **kwargs)
+
+ def load_data_list(self) -> List[dict]:
+ """Load data info from an dsdl yaml file named as ``self.ann_file``
+
+ Returns:
+ List[dict]: A list of data info.
+ """
+ if self.with_hierarchy:
+ # get classes_names and relation_matrix
+ classes_names, relation_matrix = \
+ self.dsdldataset.class_dom.get_hierarchy_info()
+ self._metainfo['classes'] = tuple(classes_names)
+ self._metainfo['RELATION_MATRIX'] = relation_matrix
+
+ else:
+ self._metainfo['classes'] = tuple(self.dsdldataset.class_names)
+
+ data_list = []
+
+ for i, data in enumerate(self.dsdldataset):
+ # basic image info, including image id, path and size.
+ datainfo = dict(
+ img_id=i,
+ img_path=os.path.join(self.data_prefix['img_path'],
+ data['Image'][0].location),
+ width=data['ImageShape'][0].width,
+ height=data['ImageShape'][0].height,
+ )
+
+ # get image label info
+ if 'image_level_labels' in data.keys():
+ if self.with_hierarchy:
+ # get leaf node name when using hierarchy classes
+ datainfo['image_level_labels'] = [
+ self._metainfo['classes'].index(i.leaf_node_name)
+ for i in data['image_level_labels']
+ ]
+ else:
+ datainfo['image_level_labels'] = [
+ self._metainfo['classes'].index(i.name)
+ for i in data['image_level_labels']
+ ]
+
+ # get semantic segmentation info
+ if 'LabelMap' in data.keys():
+ datainfo['seg_map_path'] = data['LabelMap']
+
+ # load instance info
+ instances = []
+ if 'Bbox' in data.keys():
+ for idx in range(len(data['Bbox'])):
+ bbox = data['Bbox'][idx]
+ if self.with_hierarchy:
+ # get leaf node name when using hierarchy classes
+ label = data['Label'][idx].leaf_node_name
+ label_index = self._metainfo['classes'].index(label)
+ else:
+ label = data['Label'][idx].name
+ label_index = self._metainfo['classes'].index(label)
+
+ instance = {}
+ instance['bbox'] = bbox.xyxy
+ instance['bbox_label'] = label_index
+
+ if 'ignore_flag' in data.keys():
+ # get ignore flag
+ instance['ignore_flag'] = data['ignore_flag'][idx]
+ else:
+ instance['ignore_flag'] = 0
+
+ if 'Polygon' in data.keys():
+ # get polygon info
+ polygon = data['Polygon'][idx]
+ instance['mask'] = polygon.openmmlabformat
+
+ for key in self.extra_keys:
+ # load extra instance info
+ instance[key] = data[key][idx]
+
+ instances.append(instance)
+
+ datainfo['instances'] = instances
+ # append a standard sample in data list
+ if len(datainfo['instances']) > 0:
+ data_list.append(datainfo)
+
+ return data_list
+
+ def filter_data(self) -> List[dict]:
+ """Filter annotations according to filter_cfg.
+
+ Returns:
+ List[dict]: Filtered results.
+ """
+ if self.test_mode:
+ return self.data_list
+
+ filter_empty_gt = self.filter_cfg.get('filter_empty_gt', False) \
+ if self.filter_cfg is not None else False
+ min_size = self.filter_cfg.get('min_size', 0) \
+ if self.filter_cfg is not None else 0
+
+ valid_data_list = []
+ for i, data_info in enumerate(self.data_list):
+ width = data_info['width']
+ height = data_info['height']
+ if filter_empty_gt and len(data_info['instances']) == 0:
+ continue
+ if min(width, height) >= min_size:
+ valid_data_list.append(data_info)
+
+ return valid_data_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/flickr30k.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/flickr30k.py
new file mode 100644
index 0000000000000000000000000000000000000000..0c76a41bc965bb0e8348c3d13e77d5c6e8ca08ce
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/flickr30k.py
@@ -0,0 +1,81 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import os.path as osp
+from typing import List
+
+from pycocotools.coco import COCO
+
+from mmdet.registry import DATASETS
+from .base_det_dataset import BaseDetDataset
+
+
+def convert_phrase_ids(phrase_ids: list) -> list:
+ unique_elements = sorted(set(phrase_ids))
+ element_to_new_label = {
+ element: label
+ for label, element in enumerate(unique_elements)
+ }
+ phrase_ids = [element_to_new_label[element] for element in phrase_ids]
+ return phrase_ids
+
+
+@DATASETS.register_module()
+class Flickr30kDataset(BaseDetDataset):
+ """Flickr30K Dataset."""
+
+ def load_data_list(self) -> List[dict]:
+
+ self.coco = COCO(self.ann_file)
+
+ self.ids = sorted(list(self.coco.imgs.keys()))
+
+ data_list = []
+ for img_id in self.ids:
+ if isinstance(img_id, str):
+ ann_ids = self.coco.getAnnIds(imgIds=[img_id], iscrowd=None)
+ else:
+ ann_ids = self.coco.getAnnIds(imgIds=img_id, iscrowd=None)
+
+ coco_img = self.coco.loadImgs(img_id)[0]
+
+ caption = coco_img['caption']
+ file_name = coco_img['file_name']
+ img_path = osp.join(self.data_prefix['img'], file_name)
+ width = coco_img['width']
+ height = coco_img['height']
+ tokens_positive = coco_img['tokens_positive_eval']
+ phrases = [caption[i[0][0]:i[0][1]] for i in tokens_positive]
+ phrase_ids = []
+
+ instances = []
+ annos = self.coco.loadAnns(ann_ids)
+ for anno in annos:
+ instance = {
+ 'bbox': [
+ anno['bbox'][0], anno['bbox'][1],
+ anno['bbox'][0] + anno['bbox'][2],
+ anno['bbox'][1] + anno['bbox'][3]
+ ],
+ 'bbox_label':
+ anno['category_id'],
+ 'ignore_flag':
+ anno['iscrowd']
+ }
+ phrase_ids.append(anno['phrase_ids'])
+ instances.append(instance)
+
+ phrase_ids = convert_phrase_ids(phrase_ids)
+
+ data_list.append(
+ dict(
+ img_path=img_path,
+ img_id=img_id,
+ height=height,
+ width=width,
+ instances=instances,
+ text=caption,
+ phrase_ids=phrase_ids,
+ tokens_positive=tokens_positive,
+ phrases=phrases,
+ ))
+
+ return data_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/isaid.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/isaid.py
new file mode 100644
index 0000000000000000000000000000000000000000..87067d8459c4dd6e80e5f808f613e0bd600b5f2f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/isaid.py
@@ -0,0 +1,25 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import DATASETS
+from .coco import CocoDataset
+
+
+@DATASETS.register_module()
+class iSAIDDataset(CocoDataset):
+ """Dataset for iSAID instance segmentation.
+
+ iSAID: A Large-scale Dataset for Instance Segmentation
+ in Aerial Images.
+
+ For more detail, please refer to "projects/iSAID/README.md"
+ """
+
+ METAINFO = dict(
+ classes=('background', 'ship', 'store_tank', 'baseball_diamond',
+ 'tennis_court', 'basketball_court', 'Ground_Track_Field',
+ 'Bridge', 'Large_Vehicle', 'Small_Vehicle', 'Helicopter',
+ 'Swimming_pool', 'Roundabout', 'Soccer_ball_field', 'plane',
+ 'Harbor'),
+ palette=[(0, 0, 0), (0, 0, 63), (0, 63, 63), (0, 63, 0), (0, 63, 127),
+ (0, 63, 191), (0, 63, 255), (0, 127, 63), (0, 127, 127),
+ (0, 0, 127), (0, 0, 191), (0, 0, 255), (0, 191, 127),
+ (0, 127, 191), (0, 127, 255), (0, 100, 155)])
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/lvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/lvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..b9629f5d463da183f0b4ab4c5d0f7ff7b07e4348
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/lvis.py
@@ -0,0 +1,638 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import warnings
+from typing import List
+
+from mmengine.fileio import get_local_path
+
+from mmdet.registry import DATASETS
+from .coco import CocoDataset
+
+
+@DATASETS.register_module()
+class LVISV05Dataset(CocoDataset):
+ """LVIS v0.5 dataset for detection."""
+
+ METAINFO = {
+ 'classes':
+ ('acorn', 'aerosol_can', 'air_conditioner', 'airplane', 'alarm_clock',
+ 'alcohol', 'alligator', 'almond', 'ambulance', 'amplifier', 'anklet',
+ 'antenna', 'apple', 'apple_juice', 'applesauce', 'apricot', 'apron',
+ 'aquarium', 'armband', 'armchair', 'armoire', 'armor', 'artichoke',
+ 'trash_can', 'ashtray', 'asparagus', 'atomizer', 'avocado', 'award',
+ 'awning', 'ax', 'baby_buggy', 'basketball_backboard', 'backpack',
+ 'handbag', 'suitcase', 'bagel', 'bagpipe', 'baguet', 'bait', 'ball',
+ 'ballet_skirt', 'balloon', 'bamboo', 'banana', 'Band_Aid', 'bandage',
+ 'bandanna', 'banjo', 'banner', 'barbell', 'barge', 'barrel',
+ 'barrette', 'barrow', 'baseball_base', 'baseball', 'baseball_bat',
+ 'baseball_cap', 'baseball_glove', 'basket', 'basketball_hoop',
+ 'basketball', 'bass_horn', 'bat_(animal)', 'bath_mat', 'bath_towel',
+ 'bathrobe', 'bathtub', 'batter_(food)', 'battery', 'beachball',
+ 'bead', 'beaker', 'bean_curd', 'beanbag', 'beanie', 'bear', 'bed',
+ 'bedspread', 'cow', 'beef_(food)', 'beeper', 'beer_bottle',
+ 'beer_can', 'beetle', 'bell', 'bell_pepper', 'belt', 'belt_buckle',
+ 'bench', 'beret', 'bib', 'Bible', 'bicycle', 'visor', 'binder',
+ 'binoculars', 'bird', 'birdfeeder', 'birdbath', 'birdcage',
+ 'birdhouse', 'birthday_cake', 'birthday_card', 'biscuit_(bread)',
+ 'pirate_flag', 'black_sheep', 'blackboard', 'blanket', 'blazer',
+ 'blender', 'blimp', 'blinker', 'blueberry', 'boar', 'gameboard',
+ 'boat', 'bobbin', 'bobby_pin', 'boiled_egg', 'bolo_tie', 'deadbolt',
+ 'bolt', 'bonnet', 'book', 'book_bag', 'bookcase', 'booklet',
+ 'bookmark', 'boom_microphone', 'boot', 'bottle', 'bottle_opener',
+ 'bouquet', 'bow_(weapon)', 'bow_(decorative_ribbons)', 'bow-tie',
+ 'bowl', 'pipe_bowl', 'bowler_hat', 'bowling_ball', 'bowling_pin',
+ 'boxing_glove', 'suspenders', 'bracelet', 'brass_plaque', 'brassiere',
+ 'bread-bin', 'breechcloth', 'bridal_gown', 'briefcase',
+ 'bristle_brush', 'broccoli', 'broach', 'broom', 'brownie',
+ 'brussels_sprouts', 'bubble_gum', 'bucket', 'horse_buggy', 'bull',
+ 'bulldog', 'bulldozer', 'bullet_train', 'bulletin_board',
+ 'bulletproof_vest', 'bullhorn', 'corned_beef', 'bun', 'bunk_bed',
+ 'buoy', 'burrito', 'bus_(vehicle)', 'business_card', 'butcher_knife',
+ 'butter', 'butterfly', 'button', 'cab_(taxi)', 'cabana', 'cabin_car',
+ 'cabinet', 'locker', 'cake', 'calculator', 'calendar', 'calf',
+ 'camcorder', 'camel', 'camera', 'camera_lens', 'camper_(vehicle)',
+ 'can', 'can_opener', 'candelabrum', 'candle', 'candle_holder',
+ 'candy_bar', 'candy_cane', 'walking_cane', 'canister', 'cannon',
+ 'canoe', 'cantaloup', 'canteen', 'cap_(headwear)', 'bottle_cap',
+ 'cape', 'cappuccino', 'car_(automobile)', 'railcar_(part_of_a_train)',
+ 'elevator_car', 'car_battery', 'identity_card', 'card', 'cardigan',
+ 'cargo_ship', 'carnation', 'horse_carriage', 'carrot', 'tote_bag',
+ 'cart', 'carton', 'cash_register', 'casserole', 'cassette', 'cast',
+ 'cat', 'cauliflower', 'caviar', 'cayenne_(spice)', 'CD_player',
+ 'celery', 'cellular_telephone', 'chain_mail', 'chair',
+ 'chaise_longue', 'champagne', 'chandelier', 'chap', 'checkbook',
+ 'checkerboard', 'cherry', 'chessboard',
+ 'chest_of_drawers_(furniture)', 'chicken_(animal)', 'chicken_wire',
+ 'chickpea', 'Chihuahua', 'chili_(vegetable)', 'chime', 'chinaware',
+ 'crisp_(potato_chip)', 'poker_chip', 'chocolate_bar',
+ 'chocolate_cake', 'chocolate_milk', 'chocolate_mousse', 'choker',
+ 'chopping_board', 'chopstick', 'Christmas_tree', 'slide', 'cider',
+ 'cigar_box', 'cigarette', 'cigarette_case', 'cistern', 'clarinet',
+ 'clasp', 'cleansing_agent', 'clementine', 'clip', 'clipboard',
+ 'clock', 'clock_tower', 'clothes_hamper', 'clothespin', 'clutch_bag',
+ 'coaster', 'coat', 'coat_hanger', 'coatrack', 'cock', 'coconut',
+ 'coffee_filter', 'coffee_maker', 'coffee_table', 'coffeepot', 'coil',
+ 'coin', 'colander', 'coleslaw', 'coloring_material',
+ 'combination_lock', 'pacifier', 'comic_book', 'computer_keyboard',
+ 'concrete_mixer', 'cone', 'control', 'convertible_(automobile)',
+ 'sofa_bed', 'cookie', 'cookie_jar', 'cooking_utensil',
+ 'cooler_(for_food)', 'cork_(bottle_plug)', 'corkboard', 'corkscrew',
+ 'edible_corn', 'cornbread', 'cornet', 'cornice', 'cornmeal', 'corset',
+ 'romaine_lettuce', 'costume', 'cougar', 'coverall', 'cowbell',
+ 'cowboy_hat', 'crab_(animal)', 'cracker', 'crape', 'crate', 'crayon',
+ 'cream_pitcher', 'credit_card', 'crescent_roll', 'crib', 'crock_pot',
+ 'crossbar', 'crouton', 'crow', 'crown', 'crucifix', 'cruise_ship',
+ 'police_cruiser', 'crumb', 'crutch', 'cub_(animal)', 'cube',
+ 'cucumber', 'cufflink', 'cup', 'trophy_cup', 'cupcake', 'hair_curler',
+ 'curling_iron', 'curtain', 'cushion', 'custard', 'cutting_tool',
+ 'cylinder', 'cymbal', 'dachshund', 'dagger', 'dartboard',
+ 'date_(fruit)', 'deck_chair', 'deer', 'dental_floss', 'desk',
+ 'detergent', 'diaper', 'diary', 'die', 'dinghy', 'dining_table',
+ 'tux', 'dish', 'dish_antenna', 'dishrag', 'dishtowel', 'dishwasher',
+ 'dishwasher_detergent', 'diskette', 'dispenser', 'Dixie_cup', 'dog',
+ 'dog_collar', 'doll', 'dollar', 'dolphin', 'domestic_ass', 'eye_mask',
+ 'doorbell', 'doorknob', 'doormat', 'doughnut', 'dove', 'dragonfly',
+ 'drawer', 'underdrawers', 'dress', 'dress_hat', 'dress_suit',
+ 'dresser', 'drill', 'drinking_fountain', 'drone', 'dropper',
+ 'drum_(musical_instrument)', 'drumstick', 'duck', 'duckling',
+ 'duct_tape', 'duffel_bag', 'dumbbell', 'dumpster', 'dustpan',
+ 'Dutch_oven', 'eagle', 'earphone', 'earplug', 'earring', 'easel',
+ 'eclair', 'eel', 'egg', 'egg_roll', 'egg_yolk', 'eggbeater',
+ 'eggplant', 'electric_chair', 'refrigerator', 'elephant', 'elk',
+ 'envelope', 'eraser', 'escargot', 'eyepatch', 'falcon', 'fan',
+ 'faucet', 'fedora', 'ferret', 'Ferris_wheel', 'ferry', 'fig_(fruit)',
+ 'fighter_jet', 'figurine', 'file_cabinet', 'file_(tool)',
+ 'fire_alarm', 'fire_engine', 'fire_extinguisher', 'fire_hose',
+ 'fireplace', 'fireplug', 'fish', 'fish_(food)', 'fishbowl',
+ 'fishing_boat', 'fishing_rod', 'flag', 'flagpole', 'flamingo',
+ 'flannel', 'flash', 'flashlight', 'fleece', 'flip-flop_(sandal)',
+ 'flipper_(footwear)', 'flower_arrangement', 'flute_glass', 'foal',
+ 'folding_chair', 'food_processor', 'football_(American)',
+ 'football_helmet', 'footstool', 'fork', 'forklift', 'freight_car',
+ 'French_toast', 'freshener', 'frisbee', 'frog', 'fruit_juice',
+ 'fruit_salad', 'frying_pan', 'fudge', 'funnel', 'futon', 'gag',
+ 'garbage', 'garbage_truck', 'garden_hose', 'gargle', 'gargoyle',
+ 'garlic', 'gasmask', 'gazelle', 'gelatin', 'gemstone', 'giant_panda',
+ 'gift_wrap', 'ginger', 'giraffe', 'cincture',
+ 'glass_(drink_container)', 'globe', 'glove', 'goat', 'goggles',
+ 'goldfish', 'golf_club', 'golfcart', 'gondola_(boat)', 'goose',
+ 'gorilla', 'gourd', 'surgical_gown', 'grape', 'grasshopper', 'grater',
+ 'gravestone', 'gravy_boat', 'green_bean', 'green_onion', 'griddle',
+ 'grillroom', 'grinder_(tool)', 'grits', 'grizzly', 'grocery_bag',
+ 'guacamole', 'guitar', 'gull', 'gun', 'hair_spray', 'hairbrush',
+ 'hairnet', 'hairpin', 'ham', 'hamburger', 'hammer', 'hammock',
+ 'hamper', 'hamster', 'hair_dryer', 'hand_glass', 'hand_towel',
+ 'handcart', 'handcuff', 'handkerchief', 'handle', 'handsaw',
+ 'hardback_book', 'harmonium', 'hat', 'hatbox', 'hatch', 'veil',
+ 'headband', 'headboard', 'headlight', 'headscarf', 'headset',
+ 'headstall_(for_horses)', 'hearing_aid', 'heart', 'heater',
+ 'helicopter', 'helmet', 'heron', 'highchair', 'hinge', 'hippopotamus',
+ 'hockey_stick', 'hog', 'home_plate_(baseball)', 'honey', 'fume_hood',
+ 'hook', 'horse', 'hose', 'hot-air_balloon', 'hotplate', 'hot_sauce',
+ 'hourglass', 'houseboat', 'hummingbird', 'hummus', 'polar_bear',
+ 'icecream', 'popsicle', 'ice_maker', 'ice_pack', 'ice_skate',
+ 'ice_tea', 'igniter', 'incense', 'inhaler', 'iPod',
+ 'iron_(for_clothing)', 'ironing_board', 'jacket', 'jam', 'jean',
+ 'jeep', 'jelly_bean', 'jersey', 'jet_plane', 'jewelry', 'joystick',
+ 'jumpsuit', 'kayak', 'keg', 'kennel', 'kettle', 'key', 'keycard',
+ 'kilt', 'kimono', 'kitchen_sink', 'kitchen_table', 'kite', 'kitten',
+ 'kiwi_fruit', 'knee_pad', 'knife', 'knight_(chess_piece)',
+ 'knitting_needle', 'knob', 'knocker_(on_a_door)', 'koala', 'lab_coat',
+ 'ladder', 'ladle', 'ladybug', 'lamb_(animal)', 'lamb-chop', 'lamp',
+ 'lamppost', 'lampshade', 'lantern', 'lanyard', 'laptop_computer',
+ 'lasagna', 'latch', 'lawn_mower', 'leather', 'legging_(clothing)',
+ 'Lego', 'lemon', 'lemonade', 'lettuce', 'license_plate', 'life_buoy',
+ 'life_jacket', 'lightbulb', 'lightning_rod', 'lime', 'limousine',
+ 'linen_paper', 'lion', 'lip_balm', 'lipstick', 'liquor', 'lizard',
+ 'Loafer_(type_of_shoe)', 'log', 'lollipop', 'lotion',
+ 'speaker_(stereo_equipment)', 'loveseat', 'machine_gun', 'magazine',
+ 'magnet', 'mail_slot', 'mailbox_(at_home)', 'mallet', 'mammoth',
+ 'mandarin_orange', 'manger', 'manhole', 'map', 'marker', 'martini',
+ 'mascot', 'mashed_potato', 'masher', 'mask', 'mast',
+ 'mat_(gym_equipment)', 'matchbox', 'mattress', 'measuring_cup',
+ 'measuring_stick', 'meatball', 'medicine', 'melon', 'microphone',
+ 'microscope', 'microwave_oven', 'milestone', 'milk', 'minivan',
+ 'mint_candy', 'mirror', 'mitten', 'mixer_(kitchen_tool)', 'money',
+ 'monitor_(computer_equipment) computer_monitor', 'monkey', 'motor',
+ 'motor_scooter', 'motor_vehicle', 'motorboat', 'motorcycle',
+ 'mound_(baseball)', 'mouse_(animal_rodent)',
+ 'mouse_(computer_equipment)', 'mousepad', 'muffin', 'mug', 'mushroom',
+ 'music_stool', 'musical_instrument', 'nailfile', 'nameplate',
+ 'napkin', 'neckerchief', 'necklace', 'necktie', 'needle', 'nest',
+ 'newsstand', 'nightshirt', 'nosebag_(for_animals)',
+ 'noseband_(for_animals)', 'notebook', 'notepad', 'nut', 'nutcracker',
+ 'oar', 'octopus_(food)', 'octopus_(animal)', 'oil_lamp', 'olive_oil',
+ 'omelet', 'onion', 'orange_(fruit)', 'orange_juice', 'oregano',
+ 'ostrich', 'ottoman', 'overalls_(clothing)', 'owl', 'packet',
+ 'inkpad', 'pad', 'paddle', 'padlock', 'paintbox', 'paintbrush',
+ 'painting', 'pajamas', 'palette', 'pan_(for_cooking)',
+ 'pan_(metal_container)', 'pancake', 'pantyhose', 'papaya',
+ 'paperclip', 'paper_plate', 'paper_towel', 'paperback_book',
+ 'paperweight', 'parachute', 'parakeet', 'parasail_(sports)',
+ 'parchment', 'parka', 'parking_meter', 'parrot',
+ 'passenger_car_(part_of_a_train)', 'passenger_ship', 'passport',
+ 'pastry', 'patty_(food)', 'pea_(food)', 'peach', 'peanut_butter',
+ 'pear', 'peeler_(tool_for_fruit_and_vegetables)', 'pegboard',
+ 'pelican', 'pen', 'pencil', 'pencil_box', 'pencil_sharpener',
+ 'pendulum', 'penguin', 'pennant', 'penny_(coin)', 'pepper',
+ 'pepper_mill', 'perfume', 'persimmon', 'baby', 'pet', 'petfood',
+ 'pew_(church_bench)', 'phonebook', 'phonograph_record', 'piano',
+ 'pickle', 'pickup_truck', 'pie', 'pigeon', 'piggy_bank', 'pillow',
+ 'pin_(non_jewelry)', 'pineapple', 'pinecone', 'ping-pong_ball',
+ 'pinwheel', 'tobacco_pipe', 'pipe', 'pistol', 'pita_(bread)',
+ 'pitcher_(vessel_for_liquid)', 'pitchfork', 'pizza', 'place_mat',
+ 'plate', 'platter', 'playing_card', 'playpen', 'pliers',
+ 'plow_(farm_equipment)', 'pocket_watch', 'pocketknife',
+ 'poker_(fire_stirring_tool)', 'pole', 'police_van', 'polo_shirt',
+ 'poncho', 'pony', 'pool_table', 'pop_(soda)', 'portrait',
+ 'postbox_(public)', 'postcard', 'poster', 'pot', 'flowerpot',
+ 'potato', 'potholder', 'pottery', 'pouch', 'power_shovel', 'prawn',
+ 'printer', 'projectile_(weapon)', 'projector', 'propeller', 'prune',
+ 'pudding', 'puffer_(fish)', 'puffin', 'pug-dog', 'pumpkin', 'puncher',
+ 'puppet', 'puppy', 'quesadilla', 'quiche', 'quilt', 'rabbit',
+ 'race_car', 'racket', 'radar', 'radiator', 'radio_receiver', 'radish',
+ 'raft', 'rag_doll', 'raincoat', 'ram_(animal)', 'raspberry', 'rat',
+ 'razorblade', 'reamer_(juicer)', 'rearview_mirror', 'receipt',
+ 'recliner', 'record_player', 'red_cabbage', 'reflector',
+ 'remote_control', 'rhinoceros', 'rib_(food)', 'rifle', 'ring',
+ 'river_boat', 'road_map', 'robe', 'rocking_chair', 'roller_skate',
+ 'Rollerblade', 'rolling_pin', 'root_beer',
+ 'router_(computer_equipment)', 'rubber_band', 'runner_(carpet)',
+ 'plastic_bag', 'saddle_(on_an_animal)', 'saddle_blanket', 'saddlebag',
+ 'safety_pin', 'sail', 'salad', 'salad_plate', 'salami',
+ 'salmon_(fish)', 'salmon_(food)', 'salsa', 'saltshaker',
+ 'sandal_(type_of_shoe)', 'sandwich', 'satchel', 'saucepan', 'saucer',
+ 'sausage', 'sawhorse', 'saxophone', 'scale_(measuring_instrument)',
+ 'scarecrow', 'scarf', 'school_bus', 'scissors', 'scoreboard',
+ 'scrambled_eggs', 'scraper', 'scratcher', 'screwdriver',
+ 'scrubbing_brush', 'sculpture', 'seabird', 'seahorse', 'seaplane',
+ 'seashell', 'seedling', 'serving_dish', 'sewing_machine', 'shaker',
+ 'shampoo', 'shark', 'sharpener', 'Sharpie', 'shaver_(electric)',
+ 'shaving_cream', 'shawl', 'shears', 'sheep', 'shepherd_dog',
+ 'sherbert', 'shield', 'shirt', 'shoe', 'shopping_bag',
+ 'shopping_cart', 'short_pants', 'shot_glass', 'shoulder_bag',
+ 'shovel', 'shower_head', 'shower_curtain', 'shredder_(for_paper)',
+ 'sieve', 'signboard', 'silo', 'sink', 'skateboard', 'skewer', 'ski',
+ 'ski_boot', 'ski_parka', 'ski_pole', 'skirt', 'sled', 'sleeping_bag',
+ 'sling_(bandage)', 'slipper_(footwear)', 'smoothie', 'snake',
+ 'snowboard', 'snowman', 'snowmobile', 'soap', 'soccer_ball', 'sock',
+ 'soda_fountain', 'carbonated_water', 'sofa', 'softball',
+ 'solar_array', 'sombrero', 'soup', 'soup_bowl', 'soupspoon',
+ 'sour_cream', 'soya_milk', 'space_shuttle', 'sparkler_(fireworks)',
+ 'spatula', 'spear', 'spectacles', 'spice_rack', 'spider', 'sponge',
+ 'spoon', 'sportswear', 'spotlight', 'squirrel',
+ 'stapler_(stapling_machine)', 'starfish', 'statue_(sculpture)',
+ 'steak_(food)', 'steak_knife', 'steamer_(kitchen_appliance)',
+ 'steering_wheel', 'stencil', 'stepladder', 'step_stool',
+ 'stereo_(sound_system)', 'stew', 'stirrer', 'stirrup',
+ 'stockings_(leg_wear)', 'stool', 'stop_sign', 'brake_light', 'stove',
+ 'strainer', 'strap', 'straw_(for_drinking)', 'strawberry',
+ 'street_sign', 'streetlight', 'string_cheese', 'stylus', 'subwoofer',
+ 'sugar_bowl', 'sugarcane_(plant)', 'suit_(clothing)', 'sunflower',
+ 'sunglasses', 'sunhat', 'sunscreen', 'surfboard', 'sushi', 'mop',
+ 'sweat_pants', 'sweatband', 'sweater', 'sweatshirt', 'sweet_potato',
+ 'swimsuit', 'sword', 'syringe', 'Tabasco_sauce', 'table-tennis_table',
+ 'table', 'table_lamp', 'tablecloth', 'tachometer', 'taco', 'tag',
+ 'taillight', 'tambourine', 'army_tank', 'tank_(storage_vessel)',
+ 'tank_top_(clothing)', 'tape_(sticky_cloth_or_paper)', 'tape_measure',
+ 'tapestry', 'tarp', 'tartan', 'tassel', 'tea_bag', 'teacup',
+ 'teakettle', 'teapot', 'teddy_bear', 'telephone', 'telephone_booth',
+ 'telephone_pole', 'telephoto_lens', 'television_camera',
+ 'television_set', 'tennis_ball', 'tennis_racket', 'tequila',
+ 'thermometer', 'thermos_bottle', 'thermostat', 'thimble', 'thread',
+ 'thumbtack', 'tiara', 'tiger', 'tights_(clothing)', 'timer',
+ 'tinfoil', 'tinsel', 'tissue_paper', 'toast_(food)', 'toaster',
+ 'toaster_oven', 'toilet', 'toilet_tissue', 'tomato', 'tongs',
+ 'toolbox', 'toothbrush', 'toothpaste', 'toothpick', 'cover',
+ 'tortilla', 'tow_truck', 'towel', 'towel_rack', 'toy',
+ 'tractor_(farm_equipment)', 'traffic_light', 'dirt_bike',
+ 'trailer_truck', 'train_(railroad_vehicle)', 'trampoline', 'tray',
+ 'tree_house', 'trench_coat', 'triangle_(musical_instrument)',
+ 'tricycle', 'tripod', 'trousers', 'truck', 'truffle_(chocolate)',
+ 'trunk', 'vat', 'turban', 'turkey_(bird)', 'turkey_(food)', 'turnip',
+ 'turtle', 'turtleneck_(clothing)', 'typewriter', 'umbrella',
+ 'underwear', 'unicycle', 'urinal', 'urn', 'vacuum_cleaner', 'valve',
+ 'vase', 'vending_machine', 'vent', 'videotape', 'vinegar', 'violin',
+ 'vodka', 'volleyball', 'vulture', 'waffle', 'waffle_iron', 'wagon',
+ 'wagon_wheel', 'walking_stick', 'wall_clock', 'wall_socket', 'wallet',
+ 'walrus', 'wardrobe', 'wasabi', 'automatic_washer', 'watch',
+ 'water_bottle', 'water_cooler', 'water_faucet', 'water_filter',
+ 'water_heater', 'water_jug', 'water_gun', 'water_scooter',
+ 'water_ski', 'water_tower', 'watering_can', 'watermelon',
+ 'weathervane', 'webcam', 'wedding_cake', 'wedding_ring', 'wet_suit',
+ 'wheel', 'wheelchair', 'whipped_cream', 'whiskey', 'whistle', 'wick',
+ 'wig', 'wind_chime', 'windmill', 'window_box_(for_plants)',
+ 'windshield_wiper', 'windsock', 'wine_bottle', 'wine_bucket',
+ 'wineglass', 'wing_chair', 'blinder_(for_horses)', 'wok', 'wolf',
+ 'wooden_spoon', 'wreath', 'wrench', 'wristband', 'wristlet', 'yacht',
+ 'yak', 'yogurt', 'yoke_(animal_equipment)', 'zebra', 'zucchini'),
+ 'palette':
+ None
+ }
+
+ def load_data_list(self) -> List[dict]:
+ """Load annotations from an annotation file named as ``self.ann_file``
+
+ Returns:
+ List[dict]: A list of annotation.
+ """ # noqa: E501
+ try:
+ import lvis
+ if getattr(lvis, '__version__', '0') >= '10.5.3':
+ warnings.warn(
+ 'mmlvis is deprecated, please install official lvis-api by "pip install git+https://github.com/lvis-dataset/lvis-api.git"', # noqa: E501
+ UserWarning)
+ from lvis import LVIS
+ except ImportError:
+ raise ImportError(
+ 'Package lvis is not installed. Please run "pip install git+https://github.com/lvis-dataset/lvis-api.git".' # noqa: E501
+ )
+ with get_local_path(
+ self.ann_file, backend_args=self.backend_args) as local_path:
+ self.lvis = LVIS(local_path)
+ self.cat_ids = self.lvis.get_cat_ids()
+ self.cat2label = {cat_id: i for i, cat_id in enumerate(self.cat_ids)}
+ self.cat_img_map = copy.deepcopy(self.lvis.cat_img_map)
+
+ img_ids = self.lvis.get_img_ids()
+ data_list = []
+ total_ann_ids = []
+ for img_id in img_ids:
+ raw_img_info = self.lvis.load_imgs([img_id])[0]
+ raw_img_info['img_id'] = img_id
+ if raw_img_info['file_name'].startswith('COCO'):
+ # Convert form the COCO 2014 file naming convention of
+ # COCO_[train/val/test]2014_000000000000.jpg to the 2017
+ # naming convention of 000000000000.jpg
+ # (LVIS v1 will fix this naming issue)
+ raw_img_info['file_name'] = raw_img_info['file_name'][-16:]
+ ann_ids = self.lvis.get_ann_ids(img_ids=[img_id])
+ raw_ann_info = self.lvis.load_anns(ann_ids)
+ total_ann_ids.extend(ann_ids)
+
+ parsed_data_info = self.parse_data_info({
+ 'raw_ann_info':
+ raw_ann_info,
+ 'raw_img_info':
+ raw_img_info
+ })
+ data_list.append(parsed_data_info)
+ if self.ANN_ID_UNIQUE:
+ assert len(set(total_ann_ids)) == len(
+ total_ann_ids
+ ), f"Annotation ids in '{self.ann_file}' are not unique!"
+
+ del self.lvis
+
+ return data_list
+
+
+LVISDataset = LVISV05Dataset
+DATASETS.register_module(name='LVISDataset', module=LVISDataset)
+
+
+@DATASETS.register_module()
+class LVISV1Dataset(LVISDataset):
+ """LVIS v1 dataset for detection."""
+
+ METAINFO = {
+ 'classes':
+ ('aerosol_can', 'air_conditioner', 'airplane', 'alarm_clock',
+ 'alcohol', 'alligator', 'almond', 'ambulance', 'amplifier', 'anklet',
+ 'antenna', 'apple', 'applesauce', 'apricot', 'apron', 'aquarium',
+ 'arctic_(type_of_shoe)', 'armband', 'armchair', 'armoire', 'armor',
+ 'artichoke', 'trash_can', 'ashtray', 'asparagus', 'atomizer',
+ 'avocado', 'award', 'awning', 'ax', 'baboon', 'baby_buggy',
+ 'basketball_backboard', 'backpack', 'handbag', 'suitcase', 'bagel',
+ 'bagpipe', 'baguet', 'bait', 'ball', 'ballet_skirt', 'balloon',
+ 'bamboo', 'banana', 'Band_Aid', 'bandage', 'bandanna', 'banjo',
+ 'banner', 'barbell', 'barge', 'barrel', 'barrette', 'barrow',
+ 'baseball_base', 'baseball', 'baseball_bat', 'baseball_cap',
+ 'baseball_glove', 'basket', 'basketball', 'bass_horn', 'bat_(animal)',
+ 'bath_mat', 'bath_towel', 'bathrobe', 'bathtub', 'batter_(food)',
+ 'battery', 'beachball', 'bead', 'bean_curd', 'beanbag', 'beanie',
+ 'bear', 'bed', 'bedpan', 'bedspread', 'cow', 'beef_(food)', 'beeper',
+ 'beer_bottle', 'beer_can', 'beetle', 'bell', 'bell_pepper', 'belt',
+ 'belt_buckle', 'bench', 'beret', 'bib', 'Bible', 'bicycle', 'visor',
+ 'billboard', 'binder', 'binoculars', 'bird', 'birdfeeder', 'birdbath',
+ 'birdcage', 'birdhouse', 'birthday_cake', 'birthday_card',
+ 'pirate_flag', 'black_sheep', 'blackberry', 'blackboard', 'blanket',
+ 'blazer', 'blender', 'blimp', 'blinker', 'blouse', 'blueberry',
+ 'gameboard', 'boat', 'bob', 'bobbin', 'bobby_pin', 'boiled_egg',
+ 'bolo_tie', 'deadbolt', 'bolt', 'bonnet', 'book', 'bookcase',
+ 'booklet', 'bookmark', 'boom_microphone', 'boot', 'bottle',
+ 'bottle_opener', 'bouquet', 'bow_(weapon)',
+ 'bow_(decorative_ribbons)', 'bow-tie', 'bowl', 'pipe_bowl',
+ 'bowler_hat', 'bowling_ball', 'box', 'boxing_glove', 'suspenders',
+ 'bracelet', 'brass_plaque', 'brassiere', 'bread-bin', 'bread',
+ 'breechcloth', 'bridal_gown', 'briefcase', 'broccoli', 'broach',
+ 'broom', 'brownie', 'brussels_sprouts', 'bubble_gum', 'bucket',
+ 'horse_buggy', 'bull', 'bulldog', 'bulldozer', 'bullet_train',
+ 'bulletin_board', 'bulletproof_vest', 'bullhorn', 'bun', 'bunk_bed',
+ 'buoy', 'burrito', 'bus_(vehicle)', 'business_card', 'butter',
+ 'butterfly', 'button', 'cab_(taxi)', 'cabana', 'cabin_car', 'cabinet',
+ 'locker', 'cake', 'calculator', 'calendar', 'calf', 'camcorder',
+ 'camel', 'camera', 'camera_lens', 'camper_(vehicle)', 'can',
+ 'can_opener', 'candle', 'candle_holder', 'candy_bar', 'candy_cane',
+ 'walking_cane', 'canister', 'canoe', 'cantaloup', 'canteen',
+ 'cap_(headwear)', 'bottle_cap', 'cape', 'cappuccino',
+ 'car_(automobile)', 'railcar_(part_of_a_train)', 'elevator_car',
+ 'car_battery', 'identity_card', 'card', 'cardigan', 'cargo_ship',
+ 'carnation', 'horse_carriage', 'carrot', 'tote_bag', 'cart', 'carton',
+ 'cash_register', 'casserole', 'cassette', 'cast', 'cat',
+ 'cauliflower', 'cayenne_(spice)', 'CD_player', 'celery',
+ 'cellular_telephone', 'chain_mail', 'chair', 'chaise_longue',
+ 'chalice', 'chandelier', 'chap', 'checkbook', 'checkerboard',
+ 'cherry', 'chessboard', 'chicken_(animal)', 'chickpea',
+ 'chili_(vegetable)', 'chime', 'chinaware', 'crisp_(potato_chip)',
+ 'poker_chip', 'chocolate_bar', 'chocolate_cake', 'chocolate_milk',
+ 'chocolate_mousse', 'choker', 'chopping_board', 'chopstick',
+ 'Christmas_tree', 'slide', 'cider', 'cigar_box', 'cigarette',
+ 'cigarette_case', 'cistern', 'clarinet', 'clasp', 'cleansing_agent',
+ 'cleat_(for_securing_rope)', 'clementine', 'clip', 'clipboard',
+ 'clippers_(for_plants)', 'cloak', 'clock', 'clock_tower',
+ 'clothes_hamper', 'clothespin', 'clutch_bag', 'coaster', 'coat',
+ 'coat_hanger', 'coatrack', 'cock', 'cockroach', 'cocoa_(beverage)',
+ 'coconut', 'coffee_maker', 'coffee_table', 'coffeepot', 'coil',
+ 'coin', 'colander', 'coleslaw', 'coloring_material',
+ 'combination_lock', 'pacifier', 'comic_book', 'compass',
+ 'computer_keyboard', 'condiment', 'cone', 'control',
+ 'convertible_(automobile)', 'sofa_bed', 'cooker', 'cookie',
+ 'cooking_utensil', 'cooler_(for_food)', 'cork_(bottle_plug)',
+ 'corkboard', 'corkscrew', 'edible_corn', 'cornbread', 'cornet',
+ 'cornice', 'cornmeal', 'corset', 'costume', 'cougar', 'coverall',
+ 'cowbell', 'cowboy_hat', 'crab_(animal)', 'crabmeat', 'cracker',
+ 'crape', 'crate', 'crayon', 'cream_pitcher', 'crescent_roll', 'crib',
+ 'crock_pot', 'crossbar', 'crouton', 'crow', 'crowbar', 'crown',
+ 'crucifix', 'cruise_ship', 'police_cruiser', 'crumb', 'crutch',
+ 'cub_(animal)', 'cube', 'cucumber', 'cufflink', 'cup', 'trophy_cup',
+ 'cupboard', 'cupcake', 'hair_curler', 'curling_iron', 'curtain',
+ 'cushion', 'cylinder', 'cymbal', 'dagger', 'dalmatian', 'dartboard',
+ 'date_(fruit)', 'deck_chair', 'deer', 'dental_floss', 'desk',
+ 'detergent', 'diaper', 'diary', 'die', 'dinghy', 'dining_table',
+ 'tux', 'dish', 'dish_antenna', 'dishrag', 'dishtowel', 'dishwasher',
+ 'dishwasher_detergent', 'dispenser', 'diving_board', 'Dixie_cup',
+ 'dog', 'dog_collar', 'doll', 'dollar', 'dollhouse', 'dolphin',
+ 'domestic_ass', 'doorknob', 'doormat', 'doughnut', 'dove',
+ 'dragonfly', 'drawer', 'underdrawers', 'dress', 'dress_hat',
+ 'dress_suit', 'dresser', 'drill', 'drone', 'dropper',
+ 'drum_(musical_instrument)', 'drumstick', 'duck', 'duckling',
+ 'duct_tape', 'duffel_bag', 'dumbbell', 'dumpster', 'dustpan', 'eagle',
+ 'earphone', 'earplug', 'earring', 'easel', 'eclair', 'eel', 'egg',
+ 'egg_roll', 'egg_yolk', 'eggbeater', 'eggplant', 'electric_chair',
+ 'refrigerator', 'elephant', 'elk', 'envelope', 'eraser', 'escargot',
+ 'eyepatch', 'falcon', 'fan', 'faucet', 'fedora', 'ferret',
+ 'Ferris_wheel', 'ferry', 'fig_(fruit)', 'fighter_jet', 'figurine',
+ 'file_cabinet', 'file_(tool)', 'fire_alarm', 'fire_engine',
+ 'fire_extinguisher', 'fire_hose', 'fireplace', 'fireplug',
+ 'first-aid_kit', 'fish', 'fish_(food)', 'fishbowl', 'fishing_rod',
+ 'flag', 'flagpole', 'flamingo', 'flannel', 'flap', 'flash',
+ 'flashlight', 'fleece', 'flip-flop_(sandal)', 'flipper_(footwear)',
+ 'flower_arrangement', 'flute_glass', 'foal', 'folding_chair',
+ 'food_processor', 'football_(American)', 'football_helmet',
+ 'footstool', 'fork', 'forklift', 'freight_car', 'French_toast',
+ 'freshener', 'frisbee', 'frog', 'fruit_juice', 'frying_pan', 'fudge',
+ 'funnel', 'futon', 'gag', 'garbage', 'garbage_truck', 'garden_hose',
+ 'gargle', 'gargoyle', 'garlic', 'gasmask', 'gazelle', 'gelatin',
+ 'gemstone', 'generator', 'giant_panda', 'gift_wrap', 'ginger',
+ 'giraffe', 'cincture', 'glass_(drink_container)', 'globe', 'glove',
+ 'goat', 'goggles', 'goldfish', 'golf_club', 'golfcart',
+ 'gondola_(boat)', 'goose', 'gorilla', 'gourd', 'grape', 'grater',
+ 'gravestone', 'gravy_boat', 'green_bean', 'green_onion', 'griddle',
+ 'grill', 'grits', 'grizzly', 'grocery_bag', 'guitar', 'gull', 'gun',
+ 'hairbrush', 'hairnet', 'hairpin', 'halter_top', 'ham', 'hamburger',
+ 'hammer', 'hammock', 'hamper', 'hamster', 'hair_dryer', 'hand_glass',
+ 'hand_towel', 'handcart', 'handcuff', 'handkerchief', 'handle',
+ 'handsaw', 'hardback_book', 'harmonium', 'hat', 'hatbox', 'veil',
+ 'headband', 'headboard', 'headlight', 'headscarf', 'headset',
+ 'headstall_(for_horses)', 'heart', 'heater', 'helicopter', 'helmet',
+ 'heron', 'highchair', 'hinge', 'hippopotamus', 'hockey_stick', 'hog',
+ 'home_plate_(baseball)', 'honey', 'fume_hood', 'hook', 'hookah',
+ 'hornet', 'horse', 'hose', 'hot-air_balloon', 'hotplate', 'hot_sauce',
+ 'hourglass', 'houseboat', 'hummingbird', 'hummus', 'polar_bear',
+ 'icecream', 'popsicle', 'ice_maker', 'ice_pack', 'ice_skate',
+ 'igniter', 'inhaler', 'iPod', 'iron_(for_clothing)', 'ironing_board',
+ 'jacket', 'jam', 'jar', 'jean', 'jeep', 'jelly_bean', 'jersey',
+ 'jet_plane', 'jewel', 'jewelry', 'joystick', 'jumpsuit', 'kayak',
+ 'keg', 'kennel', 'kettle', 'key', 'keycard', 'kilt', 'kimono',
+ 'kitchen_sink', 'kitchen_table', 'kite', 'kitten', 'kiwi_fruit',
+ 'knee_pad', 'knife', 'knitting_needle', 'knob', 'knocker_(on_a_door)',
+ 'koala', 'lab_coat', 'ladder', 'ladle', 'ladybug', 'lamb_(animal)',
+ 'lamb-chop', 'lamp', 'lamppost', 'lampshade', 'lantern', 'lanyard',
+ 'laptop_computer', 'lasagna', 'latch', 'lawn_mower', 'leather',
+ 'legging_(clothing)', 'Lego', 'legume', 'lemon', 'lemonade',
+ 'lettuce', 'license_plate', 'life_buoy', 'life_jacket', 'lightbulb',
+ 'lightning_rod', 'lime', 'limousine', 'lion', 'lip_balm', 'liquor',
+ 'lizard', 'log', 'lollipop', 'speaker_(stereo_equipment)', 'loveseat',
+ 'machine_gun', 'magazine', 'magnet', 'mail_slot', 'mailbox_(at_home)',
+ 'mallard', 'mallet', 'mammoth', 'manatee', 'mandarin_orange',
+ 'manger', 'manhole', 'map', 'marker', 'martini', 'mascot',
+ 'mashed_potato', 'masher', 'mask', 'mast', 'mat_(gym_equipment)',
+ 'matchbox', 'mattress', 'measuring_cup', 'measuring_stick',
+ 'meatball', 'medicine', 'melon', 'microphone', 'microscope',
+ 'microwave_oven', 'milestone', 'milk', 'milk_can', 'milkshake',
+ 'minivan', 'mint_candy', 'mirror', 'mitten', 'mixer_(kitchen_tool)',
+ 'money', 'monitor_(computer_equipment) computer_monitor', 'monkey',
+ 'motor', 'motor_scooter', 'motor_vehicle', 'motorcycle',
+ 'mound_(baseball)', 'mouse_(computer_equipment)', 'mousepad',
+ 'muffin', 'mug', 'mushroom', 'music_stool', 'musical_instrument',
+ 'nailfile', 'napkin', 'neckerchief', 'necklace', 'necktie', 'needle',
+ 'nest', 'newspaper', 'newsstand', 'nightshirt',
+ 'nosebag_(for_animals)', 'noseband_(for_animals)', 'notebook',
+ 'notepad', 'nut', 'nutcracker', 'oar', 'octopus_(food)',
+ 'octopus_(animal)', 'oil_lamp', 'olive_oil', 'omelet', 'onion',
+ 'orange_(fruit)', 'orange_juice', 'ostrich', 'ottoman', 'oven',
+ 'overalls_(clothing)', 'owl', 'packet', 'inkpad', 'pad', 'paddle',
+ 'padlock', 'paintbrush', 'painting', 'pajamas', 'palette',
+ 'pan_(for_cooking)', 'pan_(metal_container)', 'pancake', 'pantyhose',
+ 'papaya', 'paper_plate', 'paper_towel', 'paperback_book',
+ 'paperweight', 'parachute', 'parakeet', 'parasail_(sports)',
+ 'parasol', 'parchment', 'parka', 'parking_meter', 'parrot',
+ 'passenger_car_(part_of_a_train)', 'passenger_ship', 'passport',
+ 'pastry', 'patty_(food)', 'pea_(food)', 'peach', 'peanut_butter',
+ 'pear', 'peeler_(tool_for_fruit_and_vegetables)', 'wooden_leg',
+ 'pegboard', 'pelican', 'pen', 'pencil', 'pencil_box',
+ 'pencil_sharpener', 'pendulum', 'penguin', 'pennant', 'penny_(coin)',
+ 'pepper', 'pepper_mill', 'perfume', 'persimmon', 'person', 'pet',
+ 'pew_(church_bench)', 'phonebook', 'phonograph_record', 'piano',
+ 'pickle', 'pickup_truck', 'pie', 'pigeon', 'piggy_bank', 'pillow',
+ 'pin_(non_jewelry)', 'pineapple', 'pinecone', 'ping-pong_ball',
+ 'pinwheel', 'tobacco_pipe', 'pipe', 'pistol', 'pita_(bread)',
+ 'pitcher_(vessel_for_liquid)', 'pitchfork', 'pizza', 'place_mat',
+ 'plate', 'platter', 'playpen', 'pliers', 'plow_(farm_equipment)',
+ 'plume', 'pocket_watch', 'pocketknife', 'poker_(fire_stirring_tool)',
+ 'pole', 'polo_shirt', 'poncho', 'pony', 'pool_table', 'pop_(soda)',
+ 'postbox_(public)', 'postcard', 'poster', 'pot', 'flowerpot',
+ 'potato', 'potholder', 'pottery', 'pouch', 'power_shovel', 'prawn',
+ 'pretzel', 'printer', 'projectile_(weapon)', 'projector', 'propeller',
+ 'prune', 'pudding', 'puffer_(fish)', 'puffin', 'pug-dog', 'pumpkin',
+ 'puncher', 'puppet', 'puppy', 'quesadilla', 'quiche', 'quilt',
+ 'rabbit', 'race_car', 'racket', 'radar', 'radiator', 'radio_receiver',
+ 'radish', 'raft', 'rag_doll', 'raincoat', 'ram_(animal)', 'raspberry',
+ 'rat', 'razorblade', 'reamer_(juicer)', 'rearview_mirror', 'receipt',
+ 'recliner', 'record_player', 'reflector', 'remote_control',
+ 'rhinoceros', 'rib_(food)', 'rifle', 'ring', 'river_boat', 'road_map',
+ 'robe', 'rocking_chair', 'rodent', 'roller_skate', 'Rollerblade',
+ 'rolling_pin', 'root_beer', 'router_(computer_equipment)',
+ 'rubber_band', 'runner_(carpet)', 'plastic_bag',
+ 'saddle_(on_an_animal)', 'saddle_blanket', 'saddlebag', 'safety_pin',
+ 'sail', 'salad', 'salad_plate', 'salami', 'salmon_(fish)',
+ 'salmon_(food)', 'salsa', 'saltshaker', 'sandal_(type_of_shoe)',
+ 'sandwich', 'satchel', 'saucepan', 'saucer', 'sausage', 'sawhorse',
+ 'saxophone', 'scale_(measuring_instrument)', 'scarecrow', 'scarf',
+ 'school_bus', 'scissors', 'scoreboard', 'scraper', 'screwdriver',
+ 'scrubbing_brush', 'sculpture', 'seabird', 'seahorse', 'seaplane',
+ 'seashell', 'sewing_machine', 'shaker', 'shampoo', 'shark',
+ 'sharpener', 'Sharpie', 'shaver_(electric)', 'shaving_cream', 'shawl',
+ 'shears', 'sheep', 'shepherd_dog', 'sherbert', 'shield', 'shirt',
+ 'shoe', 'shopping_bag', 'shopping_cart', 'short_pants', 'shot_glass',
+ 'shoulder_bag', 'shovel', 'shower_head', 'shower_cap',
+ 'shower_curtain', 'shredder_(for_paper)', 'signboard', 'silo', 'sink',
+ 'skateboard', 'skewer', 'ski', 'ski_boot', 'ski_parka', 'ski_pole',
+ 'skirt', 'skullcap', 'sled', 'sleeping_bag', 'sling_(bandage)',
+ 'slipper_(footwear)', 'smoothie', 'snake', 'snowboard', 'snowman',
+ 'snowmobile', 'soap', 'soccer_ball', 'sock', 'sofa', 'softball',
+ 'solar_array', 'sombrero', 'soup', 'soup_bowl', 'soupspoon',
+ 'sour_cream', 'soya_milk', 'space_shuttle', 'sparkler_(fireworks)',
+ 'spatula', 'spear', 'spectacles', 'spice_rack', 'spider', 'crawfish',
+ 'sponge', 'spoon', 'sportswear', 'spotlight', 'squid_(food)',
+ 'squirrel', 'stagecoach', 'stapler_(stapling_machine)', 'starfish',
+ 'statue_(sculpture)', 'steak_(food)', 'steak_knife', 'steering_wheel',
+ 'stepladder', 'step_stool', 'stereo_(sound_system)', 'stew',
+ 'stirrer', 'stirrup', 'stool', 'stop_sign', 'brake_light', 'stove',
+ 'strainer', 'strap', 'straw_(for_drinking)', 'strawberry',
+ 'street_sign', 'streetlight', 'string_cheese', 'stylus', 'subwoofer',
+ 'sugar_bowl', 'sugarcane_(plant)', 'suit_(clothing)', 'sunflower',
+ 'sunglasses', 'sunhat', 'surfboard', 'sushi', 'mop', 'sweat_pants',
+ 'sweatband', 'sweater', 'sweatshirt', 'sweet_potato', 'swimsuit',
+ 'sword', 'syringe', 'Tabasco_sauce', 'table-tennis_table', 'table',
+ 'table_lamp', 'tablecloth', 'tachometer', 'taco', 'tag', 'taillight',
+ 'tambourine', 'army_tank', 'tank_(storage_vessel)',
+ 'tank_top_(clothing)', 'tape_(sticky_cloth_or_paper)', 'tape_measure',
+ 'tapestry', 'tarp', 'tartan', 'tassel', 'tea_bag', 'teacup',
+ 'teakettle', 'teapot', 'teddy_bear', 'telephone', 'telephone_booth',
+ 'telephone_pole', 'telephoto_lens', 'television_camera',
+ 'television_set', 'tennis_ball', 'tennis_racket', 'tequila',
+ 'thermometer', 'thermos_bottle', 'thermostat', 'thimble', 'thread',
+ 'thumbtack', 'tiara', 'tiger', 'tights_(clothing)', 'timer',
+ 'tinfoil', 'tinsel', 'tissue_paper', 'toast_(food)', 'toaster',
+ 'toaster_oven', 'toilet', 'toilet_tissue', 'tomato', 'tongs',
+ 'toolbox', 'toothbrush', 'toothpaste', 'toothpick', 'cover',
+ 'tortilla', 'tow_truck', 'towel', 'towel_rack', 'toy',
+ 'tractor_(farm_equipment)', 'traffic_light', 'dirt_bike',
+ 'trailer_truck', 'train_(railroad_vehicle)', 'trampoline', 'tray',
+ 'trench_coat', 'triangle_(musical_instrument)', 'tricycle', 'tripod',
+ 'trousers', 'truck', 'truffle_(chocolate)', 'trunk', 'vat', 'turban',
+ 'turkey_(food)', 'turnip', 'turtle', 'turtleneck_(clothing)',
+ 'typewriter', 'umbrella', 'underwear', 'unicycle', 'urinal', 'urn',
+ 'vacuum_cleaner', 'vase', 'vending_machine', 'vent', 'vest',
+ 'videotape', 'vinegar', 'violin', 'vodka', 'volleyball', 'vulture',
+ 'waffle', 'waffle_iron', 'wagon', 'wagon_wheel', 'walking_stick',
+ 'wall_clock', 'wall_socket', 'wallet', 'walrus', 'wardrobe',
+ 'washbasin', 'automatic_washer', 'watch', 'water_bottle',
+ 'water_cooler', 'water_faucet', 'water_heater', 'water_jug',
+ 'water_gun', 'water_scooter', 'water_ski', 'water_tower',
+ 'watering_can', 'watermelon', 'weathervane', 'webcam', 'wedding_cake',
+ 'wedding_ring', 'wet_suit', 'wheel', 'wheelchair', 'whipped_cream',
+ 'whistle', 'wig', 'wind_chime', 'windmill', 'window_box_(for_plants)',
+ 'windshield_wiper', 'windsock', 'wine_bottle', 'wine_bucket',
+ 'wineglass', 'blinder_(for_horses)', 'wok', 'wolf', 'wooden_spoon',
+ 'wreath', 'wrench', 'wristband', 'wristlet', 'yacht', 'yogurt',
+ 'yoke_(animal_equipment)', 'zebra', 'zucchini'),
+ 'palette':
+ None
+ }
+
+ def load_data_list(self) -> List[dict]:
+ """Load annotations from an annotation file named as ``self.ann_file``
+
+ Returns:
+ List[dict]: A list of annotation.
+ """ # noqa: E501
+ try:
+ import lvis
+ if getattr(lvis, '__version__', '0') >= '10.5.3':
+ warnings.warn(
+ 'mmlvis is deprecated, please install official lvis-api by "pip install git+https://github.com/lvis-dataset/lvis-api.git"', # noqa: E501
+ UserWarning)
+ from lvis import LVIS
+ except ImportError:
+ raise ImportError(
+ 'Package lvis is not installed. Please run "pip install git+https://github.com/lvis-dataset/lvis-api.git".' # noqa: E501
+ )
+ with get_local_path(
+ self.ann_file, backend_args=self.backend_args) as local_path:
+ self.lvis = LVIS(local_path)
+ self.cat_ids = self.lvis.get_cat_ids()
+ self.cat2label = {cat_id: i for i, cat_id in enumerate(self.cat_ids)}
+ self.cat_img_map = copy.deepcopy(self.lvis.cat_img_map)
+
+ img_ids = self.lvis.get_img_ids()
+ data_list = []
+ total_ann_ids = []
+ for img_id in img_ids:
+ raw_img_info = self.lvis.load_imgs([img_id])[0]
+ raw_img_info['img_id'] = img_id
+ # coco_url is used in LVISv1 instead of file_name
+ # e.g. http://images.cocodataset.org/train2017/000000391895.jpg
+ # train/val split in specified in url
+ raw_img_info['file_name'] = raw_img_info['coco_url'].replace(
+ 'http://images.cocodataset.org/', '')
+ ann_ids = self.lvis.get_ann_ids(img_ids=[img_id])
+ raw_ann_info = self.lvis.load_anns(ann_ids)
+ total_ann_ids.extend(ann_ids)
+ parsed_data_info = self.parse_data_info({
+ 'raw_ann_info':
+ raw_ann_info,
+ 'raw_img_info':
+ raw_img_info
+ })
+ data_list.append(parsed_data_info)
+ if self.ANN_ID_UNIQUE:
+ assert len(set(total_ann_ids)) == len(
+ total_ann_ids
+ ), f"Annotation ids in '{self.ann_file}' are not unique!"
+
+ del self.lvis
+
+ return data_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/mdetr_style_refcoco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/mdetr_style_refcoco.py
new file mode 100644
index 0000000000000000000000000000000000000000..cc56dec49db72daddf929bcc65471ffc2ca6fb4d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/mdetr_style_refcoco.py
@@ -0,0 +1,57 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import os.path as osp
+from typing import List
+
+from mmengine.fileio import get_local_path
+
+from mmdet.datasets import BaseDetDataset
+from mmdet.registry import DATASETS
+from .api_wrappers import COCO
+
+
+@DATASETS.register_module()
+class MDETRStyleRefCocoDataset(BaseDetDataset):
+ """RefCOCO dataset.
+
+ Only support evaluation now.
+ """
+
+ def load_data_list(self) -> List[dict]:
+ with get_local_path(
+ self.ann_file, backend_args=self.backend_args) as local_path:
+ coco = COCO(local_path)
+
+ img_ids = coco.get_img_ids()
+
+ data_infos = []
+ for img_id in img_ids:
+ raw_img_info = coco.load_imgs([img_id])[0]
+ ann_ids = coco.get_ann_ids(img_ids=[img_id])
+ raw_ann_info = coco.load_anns(ann_ids)
+
+ data_info = {}
+ img_path = osp.join(self.data_prefix['img'],
+ raw_img_info['file_name'])
+ data_info['img_path'] = img_path
+ data_info['img_id'] = img_id
+ data_info['height'] = raw_img_info['height']
+ data_info['width'] = raw_img_info['width']
+ data_info['dataset_mode'] = raw_img_info['dataset_name']
+
+ data_info['text'] = raw_img_info['caption']
+ data_info['custom_entities'] = False
+ data_info['tokens_positive'] = -1
+
+ instances = []
+ for i, ann in enumerate(raw_ann_info):
+ instance = {}
+ x1, y1, w, h = ann['bbox']
+ bbox = [x1, y1, x1 + w, y1 + h]
+ instance['bbox'] = bbox
+ instance['bbox_label'] = ann['category_id']
+ instance['ignore_flag'] = 0
+ instances.append(instance)
+
+ data_info['instances'] = instances
+ data_infos.append(data_info)
+ return data_infos
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/mot_challenge_dataset.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/mot_challenge_dataset.py
new file mode 100644
index 0000000000000000000000000000000000000000..ffbdc48ebf8d4a4ba11a605c8bc2a479cf2a0c96
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/mot_challenge_dataset.py
@@ -0,0 +1,88 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import os.path as osp
+from typing import List, Union
+
+from mmdet.registry import DATASETS
+from .base_video_dataset import BaseVideoDataset
+
+
+@DATASETS.register_module()
+class MOTChallengeDataset(BaseVideoDataset):
+ """Dataset for MOTChallenge.
+
+ Args:
+ visibility_thr (float, optional): The minimum visibility
+ for the objects during training. Default to -1.
+ """
+
+ METAINFO = {
+ 'classes':
+ ('pedestrian', 'person_on_vehicle', 'car', 'bicycle', 'motorbike',
+ 'non_mot_vehicle', 'static_person', 'distractor', 'occluder',
+ 'occluder_on_ground', 'occluder_full', 'reflection', 'crowd')
+ }
+
+ def __init__(self, visibility_thr: float = -1, *args, **kwargs):
+ self.visibility_thr = visibility_thr
+ super().__init__(*args, **kwargs)
+
+ def parse_data_info(self, raw_data_info: dict) -> Union[dict, List[dict]]:
+ """Parse raw annotation to target format. The difference between this
+ function and the one in ``BaseVideoDataset`` is that the parsing here
+ adds ``visibility`` and ``mot_conf``.
+
+ Args:
+ raw_data_info (dict): Raw data information load from ``ann_file``
+
+ Returns:
+ Union[dict, List[dict]]: Parsed annotation.
+ """
+ img_info = raw_data_info['raw_img_info']
+ ann_info = raw_data_info['raw_ann_info']
+ data_info = {}
+
+ data_info.update(img_info)
+ if self.data_prefix.get('img_path', None) is not None:
+ img_path = osp.join(self.data_prefix['img_path'],
+ img_info['file_name'])
+ else:
+ img_path = img_info['file_name']
+ data_info['img_path'] = img_path
+
+ instances = []
+ for i, ann in enumerate(ann_info):
+ instance = {}
+
+ if (not self.test_mode) and (ann['visibility'] <
+ self.visibility_thr):
+ continue
+ if ann.get('ignore', False):
+ continue
+ x1, y1, w, h = ann['bbox']
+ inter_w = max(0, min(x1 + w, img_info['width']) - max(x1, 0))
+ inter_h = max(0, min(y1 + h, img_info['height']) - max(y1, 0))
+ if inter_w * inter_h == 0:
+ continue
+ if ann['area'] <= 0 or w < 1 or h < 1:
+ continue
+ if ann['category_id'] not in self.cat_ids:
+ continue
+ bbox = [x1, y1, x1 + w, y1 + h]
+
+ if ann.get('iscrowd', False):
+ instance['ignore_flag'] = 1
+ else:
+ instance['ignore_flag'] = 0
+ instance['bbox'] = bbox
+ instance['bbox_label'] = self.cat2label[ann['category_id']]
+ instance['instance_id'] = ann['instance_id']
+ instance['category_id'] = ann['category_id']
+ instance['mot_conf'] = ann['mot_conf']
+ instance['visibility'] = ann['visibility']
+ if len(instance) > 0:
+ instances.append(instance)
+ if not self.test_mode:
+ assert len(instances) > 0, f'No valid instances found in ' \
+ f'image {data_info["img_path"]}!'
+ data_info['instances'] = instances
+ return data_info
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/objects365.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/objects365.py
new file mode 100644
index 0000000000000000000000000000000000000000..e99869bfa309635af3c03cbfa77f732db3f50637
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/objects365.py
@@ -0,0 +1,284 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import os.path as osp
+from typing import List
+
+from mmengine.fileio import get_local_path
+
+from mmdet.registry import DATASETS
+from .api_wrappers import COCO
+from .coco import CocoDataset
+
+# images exist in annotations but not in image folder.
+objv2_ignore_list = [
+ osp.join('patch16', 'objects365_v2_00908726.jpg'),
+ osp.join('patch6', 'objects365_v1_00320532.jpg'),
+ osp.join('patch6', 'objects365_v1_00320534.jpg'),
+]
+
+
+@DATASETS.register_module()
+class Objects365V1Dataset(CocoDataset):
+ """Objects365 v1 dataset for detection."""
+
+ METAINFO = {
+ 'classes':
+ ('person', 'sneakers', 'chair', 'hat', 'lamp', 'bottle',
+ 'cabinet/shelf', 'cup', 'car', 'glasses', 'picture/frame', 'desk',
+ 'handbag', 'street lights', 'book', 'plate', 'helmet',
+ 'leather shoes', 'pillow', 'glove', 'potted plant', 'bracelet',
+ 'flower', 'tv', 'storage box', 'vase', 'bench', 'wine glass', 'boots',
+ 'bowl', 'dining table', 'umbrella', 'boat', 'flag', 'speaker',
+ 'trash bin/can', 'stool', 'backpack', 'couch', 'belt', 'carpet',
+ 'basket', 'towel/napkin', 'slippers', 'barrel/bucket', 'coffee table',
+ 'suv', 'toy', 'tie', 'bed', 'traffic light', 'pen/pencil',
+ 'microphone', 'sandals', 'canned', 'necklace', 'mirror', 'faucet',
+ 'bicycle', 'bread', 'high heels', 'ring', 'van', 'watch', 'sink',
+ 'horse', 'fish', 'apple', 'camera', 'candle', 'teddy bear', 'cake',
+ 'motorcycle', 'wild bird', 'laptop', 'knife', 'traffic sign',
+ 'cell phone', 'paddle', 'truck', 'cow', 'power outlet', 'clock',
+ 'drum', 'fork', 'bus', 'hanger', 'nightstand', 'pot/pan', 'sheep',
+ 'guitar', 'traffic cone', 'tea pot', 'keyboard', 'tripod', 'hockey',
+ 'fan', 'dog', 'spoon', 'blackboard/whiteboard', 'balloon',
+ 'air conditioner', 'cymbal', 'mouse', 'telephone', 'pickup truck',
+ 'orange', 'banana', 'airplane', 'luggage', 'skis', 'soccer',
+ 'trolley', 'oven', 'remote', 'baseball glove', 'paper towel',
+ 'refrigerator', 'train', 'tomato', 'machinery vehicle', 'tent',
+ 'shampoo/shower gel', 'head phone', 'lantern', 'donut',
+ 'cleaning products', 'sailboat', 'tangerine', 'pizza', 'kite',
+ 'computer box', 'elephant', 'toiletries', 'gas stove', 'broccoli',
+ 'toilet', 'stroller', 'shovel', 'baseball bat', 'microwave',
+ 'skateboard', 'surfboard', 'surveillance camera', 'gun', 'life saver',
+ 'cat', 'lemon', 'liquid soap', 'zebra', 'duck', 'sports car',
+ 'giraffe', 'pumpkin', 'piano', 'stop sign', 'radiator', 'converter',
+ 'tissue ', 'carrot', 'washing machine', 'vent', 'cookies',
+ 'cutting/chopping board', 'tennis racket', 'candy',
+ 'skating and skiing shoes', 'scissors', 'folder', 'baseball',
+ 'strawberry', 'bow tie', 'pigeon', 'pepper', 'coffee machine',
+ 'bathtub', 'snowboard', 'suitcase', 'grapes', 'ladder', 'pear',
+ 'american football', 'basketball', 'potato', 'paint brush', 'printer',
+ 'billiards', 'fire hydrant', 'goose', 'projector', 'sausage',
+ 'fire extinguisher', 'extension cord', 'facial mask', 'tennis ball',
+ 'chopsticks', 'electronic stove and gas stove', 'pie', 'frisbee',
+ 'kettle', 'hamburger', 'golf club', 'cucumber', 'clutch', 'blender',
+ 'tong', 'slide', 'hot dog', 'toothbrush', 'facial cleanser', 'mango',
+ 'deer', 'egg', 'violin', 'marker', 'ship', 'chicken', 'onion',
+ 'ice cream', 'tape', 'wheelchair', 'plum', 'bar soap', 'scale',
+ 'watermelon', 'cabbage', 'router/modem', 'golf ball', 'pine apple',
+ 'crane', 'fire truck', 'peach', 'cello', 'notepaper', 'tricycle',
+ 'toaster', 'helicopter', 'green beans', 'brush', 'carriage', 'cigar',
+ 'earphone', 'penguin', 'hurdle', 'swing', 'radio', 'CD',
+ 'parking meter', 'swan', 'garlic', 'french fries', 'horn', 'avocado',
+ 'saxophone', 'trumpet', 'sandwich', 'cue', 'kiwi fruit', 'bear',
+ 'fishing rod', 'cherry', 'tablet', 'green vegetables', 'nuts', 'corn',
+ 'key', 'screwdriver', 'globe', 'broom', 'pliers', 'volleyball',
+ 'hammer', 'eggplant', 'trophy', 'dates', 'board eraser', 'rice',
+ 'tape measure/ruler', 'dumbbell', 'hamimelon', 'stapler', 'camel',
+ 'lettuce', 'goldfish', 'meat balls', 'medal', 'toothpaste',
+ 'antelope', 'shrimp', 'rickshaw', 'trombone', 'pomegranate',
+ 'coconut', 'jellyfish', 'mushroom', 'calculator', 'treadmill',
+ 'butterfly', 'egg tart', 'cheese', 'pig', 'pomelo', 'race car',
+ 'rice cooker', 'tuba', 'crosswalk sign', 'papaya', 'hair drier',
+ 'green onion', 'chips', 'dolphin', 'sushi', 'urinal', 'donkey',
+ 'electric drill', 'spring rolls', 'tortoise/turtle', 'parrot',
+ 'flute', 'measuring cup', 'shark', 'steak', 'poker card',
+ 'binoculars', 'llama', 'radish', 'noodles', 'yak', 'mop', 'crab',
+ 'microscope', 'barbell', 'bread/bun', 'baozi', 'lion', 'red cabbage',
+ 'polar bear', 'lighter', 'seal', 'mangosteen', 'comb', 'eraser',
+ 'pitaya', 'scallop', 'pencil case', 'saw', 'table tennis paddle',
+ 'okra', 'starfish', 'eagle', 'monkey', 'durian', 'game board',
+ 'rabbit', 'french horn', 'ambulance', 'asparagus', 'hoverboard',
+ 'pasta', 'target', 'hotair balloon', 'chainsaw', 'lobster', 'iron',
+ 'flashlight'),
+ 'palette':
+ None
+ }
+
+ COCOAPI = COCO
+ # ann_id is unique in coco dataset.
+ ANN_ID_UNIQUE = True
+
+ def load_data_list(self) -> List[dict]:
+ """Load annotations from an annotation file named as ``self.ann_file``
+
+ Returns:
+ List[dict]: A list of annotation.
+ """ # noqa: E501
+ with get_local_path(
+ self.ann_file, backend_args=self.backend_args) as local_path:
+ self.coco = self.COCOAPI(local_path)
+
+ # 'categories' list in objects365_train.json and objects365_val.json
+ # is inconsistent, need sort list(or dict) before get cat_ids.
+ cats = self.coco.cats
+ sorted_cats = {i: cats[i] for i in sorted(cats)}
+ self.coco.cats = sorted_cats
+ categories = self.coco.dataset['categories']
+ sorted_categories = sorted(categories, key=lambda i: i['id'])
+ self.coco.dataset['categories'] = sorted_categories
+ # The order of returned `cat_ids` will not
+ # change with the order of the `classes`
+ self.cat_ids = self.coco.get_cat_ids(
+ cat_names=self.metainfo['classes'])
+ self.cat2label = {cat_id: i for i, cat_id in enumerate(self.cat_ids)}
+ self.cat_img_map = copy.deepcopy(self.coco.cat_img_map)
+
+ img_ids = self.coco.get_img_ids()
+ data_list = []
+ total_ann_ids = []
+ for img_id in img_ids:
+ raw_img_info = self.coco.load_imgs([img_id])[0]
+ raw_img_info['img_id'] = img_id
+
+ ann_ids = self.coco.get_ann_ids(img_ids=[img_id])
+ raw_ann_info = self.coco.load_anns(ann_ids)
+ total_ann_ids.extend(ann_ids)
+
+ parsed_data_info = self.parse_data_info({
+ 'raw_ann_info':
+ raw_ann_info,
+ 'raw_img_info':
+ raw_img_info
+ })
+ data_list.append(parsed_data_info)
+ if self.ANN_ID_UNIQUE:
+ assert len(set(total_ann_ids)) == len(
+ total_ann_ids
+ ), f"Annotation ids in '{self.ann_file}' are not unique!"
+
+ del self.coco
+
+ return data_list
+
+
+@DATASETS.register_module()
+class Objects365V2Dataset(CocoDataset):
+ """Objects365 v2 dataset for detection."""
+ METAINFO = {
+ 'classes':
+ ('Person', 'Sneakers', 'Chair', 'Other Shoes', 'Hat', 'Car', 'Lamp',
+ 'Glasses', 'Bottle', 'Desk', 'Cup', 'Street Lights', 'Cabinet/shelf',
+ 'Handbag/Satchel', 'Bracelet', 'Plate', 'Picture/Frame', 'Helmet',
+ 'Book', 'Gloves', 'Storage box', 'Boat', 'Leather Shoes', 'Flower',
+ 'Bench', 'Potted Plant', 'Bowl/Basin', 'Flag', 'Pillow', 'Boots',
+ 'Vase', 'Microphone', 'Necklace', 'Ring', 'SUV', 'Wine Glass', 'Belt',
+ 'Moniter/TV', 'Backpack', 'Umbrella', 'Traffic Light', 'Speaker',
+ 'Watch', 'Tie', 'Trash bin Can', 'Slippers', 'Bicycle', 'Stool',
+ 'Barrel/bucket', 'Van', 'Couch', 'Sandals', 'Bakset', 'Drum',
+ 'Pen/Pencil', 'Bus', 'Wild Bird', 'High Heels', 'Motorcycle',
+ 'Guitar', 'Carpet', 'Cell Phone', 'Bread', 'Camera', 'Canned',
+ 'Truck', 'Traffic cone', 'Cymbal', 'Lifesaver', 'Towel',
+ 'Stuffed Toy', 'Candle', 'Sailboat', 'Laptop', 'Awning', 'Bed',
+ 'Faucet', 'Tent', 'Horse', 'Mirror', 'Power outlet', 'Sink', 'Apple',
+ 'Air Conditioner', 'Knife', 'Hockey Stick', 'Paddle', 'Pickup Truck',
+ 'Fork', 'Traffic Sign', 'Ballon', 'Tripod', 'Dog', 'Spoon', 'Clock',
+ 'Pot', 'Cow', 'Cake', 'Dinning Table', 'Sheep', 'Hanger',
+ 'Blackboard/Whiteboard', 'Napkin', 'Other Fish', 'Orange/Tangerine',
+ 'Toiletry', 'Keyboard', 'Tomato', 'Lantern', 'Machinery Vehicle',
+ 'Fan', 'Green Vegetables', 'Banana', 'Baseball Glove', 'Airplane',
+ 'Mouse', 'Train', 'Pumpkin', 'Soccer', 'Skiboard', 'Luggage',
+ 'Nightstand', 'Tea pot', 'Telephone', 'Trolley', 'Head Phone',
+ 'Sports Car', 'Stop Sign', 'Dessert', 'Scooter', 'Stroller', 'Crane',
+ 'Remote', 'Refrigerator', 'Oven', 'Lemon', 'Duck', 'Baseball Bat',
+ 'Surveillance Camera', 'Cat', 'Jug', 'Broccoli', 'Piano', 'Pizza',
+ 'Elephant', 'Skateboard', 'Surfboard', 'Gun',
+ 'Skating and Skiing shoes', 'Gas stove', 'Donut', 'Bow Tie', 'Carrot',
+ 'Toilet', 'Kite', 'Strawberry', 'Other Balls', 'Shovel', 'Pepper',
+ 'Computer Box', 'Toilet Paper', 'Cleaning Products', 'Chopsticks',
+ 'Microwave', 'Pigeon', 'Baseball', 'Cutting/chopping Board',
+ 'Coffee Table', 'Side Table', 'Scissors', 'Marker', 'Pie', 'Ladder',
+ 'Snowboard', 'Cookies', 'Radiator', 'Fire Hydrant', 'Basketball',
+ 'Zebra', 'Grape', 'Giraffe', 'Potato', 'Sausage', 'Tricycle',
+ 'Violin', 'Egg', 'Fire Extinguisher', 'Candy', 'Fire Truck',
+ 'Billards', 'Converter', 'Bathtub', 'Wheelchair', 'Golf Club',
+ 'Briefcase', 'Cucumber', 'Cigar/Cigarette ', 'Paint Brush', 'Pear',
+ 'Heavy Truck', 'Hamburger', 'Extractor', 'Extention Cord', 'Tong',
+ 'Tennis Racket', 'Folder', 'American Football', 'earphone', 'Mask',
+ 'Kettle', 'Tennis', 'Ship', 'Swing', 'Coffee Machine', 'Slide',
+ 'Carriage', 'Onion', 'Green beans', 'Projector', 'Frisbee',
+ 'Washing Machine/Drying Machine', 'Chicken', 'Printer', 'Watermelon',
+ 'Saxophone', 'Tissue', 'Toothbrush', 'Ice cream', 'Hotair ballon',
+ 'Cello', 'French Fries', 'Scale', 'Trophy', 'Cabbage', 'Hot dog',
+ 'Blender', 'Peach', 'Rice', 'Wallet/Purse', 'Volleyball', 'Deer',
+ 'Goose', 'Tape', 'Tablet', 'Cosmetics', 'Trumpet', 'Pineapple',
+ 'Golf Ball', 'Ambulance', 'Parking meter', 'Mango', 'Key', 'Hurdle',
+ 'Fishing Rod', 'Medal', 'Flute', 'Brush', 'Penguin', 'Megaphone',
+ 'Corn', 'Lettuce', 'Garlic', 'Swan', 'Helicopter', 'Green Onion',
+ 'Sandwich', 'Nuts', 'Speed Limit Sign', 'Induction Cooker', 'Broom',
+ 'Trombone', 'Plum', 'Rickshaw', 'Goldfish', 'Kiwi fruit',
+ 'Router/modem', 'Poker Card', 'Toaster', 'Shrimp', 'Sushi', 'Cheese',
+ 'Notepaper', 'Cherry', 'Pliers', 'CD', 'Pasta', 'Hammer', 'Cue',
+ 'Avocado', 'Hamimelon', 'Flask', 'Mushroon', 'Screwdriver', 'Soap',
+ 'Recorder', 'Bear', 'Eggplant', 'Board Eraser', 'Coconut',
+ 'Tape Measur/ Ruler', 'Pig', 'Showerhead', 'Globe', 'Chips', 'Steak',
+ 'Crosswalk Sign', 'Stapler', 'Campel', 'Formula 1 ', 'Pomegranate',
+ 'Dishwasher', 'Crab', 'Hoverboard', 'Meat ball', 'Rice Cooker',
+ 'Tuba', 'Calculator', 'Papaya', 'Antelope', 'Parrot', 'Seal',
+ 'Buttefly', 'Dumbbell', 'Donkey', 'Lion', 'Urinal', 'Dolphin',
+ 'Electric Drill', 'Hair Dryer', 'Egg tart', 'Jellyfish', 'Treadmill',
+ 'Lighter', 'Grapefruit', 'Game board', 'Mop', 'Radish', 'Baozi',
+ 'Target', 'French', 'Spring Rolls', 'Monkey', 'Rabbit', 'Pencil Case',
+ 'Yak', 'Red Cabbage', 'Binoculars', 'Asparagus', 'Barbell', 'Scallop',
+ 'Noddles', 'Comb', 'Dumpling', 'Oyster', 'Table Teniis paddle',
+ 'Cosmetics Brush/Eyeliner Pencil', 'Chainsaw', 'Eraser', 'Lobster',
+ 'Durian', 'Okra', 'Lipstick', 'Cosmetics Mirror', 'Curling',
+ 'Table Tennis '),
+ 'palette':
+ None
+ }
+
+ COCOAPI = COCO
+ # ann_id is unique in coco dataset.
+ ANN_ID_UNIQUE = True
+
+ def load_data_list(self) -> List[dict]:
+ """Load annotations from an annotation file named as ``self.ann_file``
+
+ Returns:
+ List[dict]: A list of annotation.
+ """ # noqa: E501
+ with get_local_path(
+ self.ann_file, backend_args=self.backend_args) as local_path:
+ self.coco = self.COCOAPI(local_path)
+ # The order of returned `cat_ids` will not
+ # change with the order of the `classes`
+ self.cat_ids = self.coco.get_cat_ids(
+ cat_names=self.metainfo['classes'])
+ self.cat2label = {cat_id: i for i, cat_id in enumerate(self.cat_ids)}
+ self.cat_img_map = copy.deepcopy(self.coco.cat_img_map)
+
+ img_ids = self.coco.get_img_ids()
+ data_list = []
+ total_ann_ids = []
+ for img_id in img_ids:
+ raw_img_info = self.coco.load_imgs([img_id])[0]
+ raw_img_info['img_id'] = img_id
+
+ ann_ids = self.coco.get_ann_ids(img_ids=[img_id])
+ raw_ann_info = self.coco.load_anns(ann_ids)
+ total_ann_ids.extend(ann_ids)
+
+ # file_name should be `patchX/xxx.jpg`
+ file_name = osp.join(
+ osp.split(osp.split(raw_img_info['file_name'])[0])[-1],
+ osp.split(raw_img_info['file_name'])[-1])
+
+ if file_name in objv2_ignore_list:
+ continue
+
+ raw_img_info['file_name'] = file_name
+ parsed_data_info = self.parse_data_info({
+ 'raw_ann_info':
+ raw_ann_info,
+ 'raw_img_info':
+ raw_img_info
+ })
+ data_list.append(parsed_data_info)
+ if self.ANN_ID_UNIQUE:
+ assert len(set(total_ann_ids)) == len(
+ total_ann_ids
+ ), f"Annotation ids in '{self.ann_file}' are not unique!"
+
+ del self.coco
+
+ return data_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/odvg.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/odvg.py
new file mode 100644
index 0000000000000000000000000000000000000000..c73865f2ea724205640bea2c701c355bbd9135e3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/odvg.py
@@ -0,0 +1,106 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import json
+import os.path as osp
+from typing import List, Optional
+
+from mmengine.fileio import get_local_path
+
+from mmdet.registry import DATASETS
+from .base_det_dataset import BaseDetDataset
+
+
+@DATASETS.register_module()
+class ODVGDataset(BaseDetDataset):
+ """object detection and visual grounding dataset."""
+
+ def __init__(self,
+ *args,
+ data_root: str = '',
+ label_map_file: Optional[str] = None,
+ need_text: bool = True,
+ **kwargs) -> None:
+ self.dataset_mode = 'VG'
+ self.need_text = need_text
+ if label_map_file:
+ label_map_file = osp.join(data_root, label_map_file)
+ with open(label_map_file, 'r') as file:
+ self.label_map = json.load(file)
+ self.dataset_mode = 'OD'
+ super().__init__(*args, data_root=data_root, **kwargs)
+ assert self.return_classes is True
+
+ def load_data_list(self) -> List[dict]:
+ with get_local_path(
+ self.ann_file, backend_args=self.backend_args) as local_path:
+ with open(local_path, 'r') as f:
+ data_list = [json.loads(line) for line in f]
+
+ out_data_list = []
+ for data in data_list:
+ data_info = {}
+ img_path = osp.join(self.data_prefix['img'], data['filename'])
+ data_info['img_path'] = img_path
+ data_info['height'] = data['height']
+ data_info['width'] = data['width']
+ if self.dataset_mode == 'OD':
+ if self.need_text:
+ data_info['text'] = self.label_map
+ anno = data.get('detection', {})
+ instances = [obj for obj in anno.get('instances', [])]
+ bboxes = [obj['bbox'] for obj in instances]
+ bbox_labels = [str(obj['label']) for obj in instances]
+
+ instances = []
+ for bbox, label in zip(bboxes, bbox_labels):
+ instance = {}
+ x1, y1, x2, y2 = bbox
+ inter_w = max(0, min(x2, data['width']) - max(x1, 0))
+ inter_h = max(0, min(y2, data['height']) - max(y1, 0))
+ if inter_w * inter_h == 0:
+ continue
+ if (x2 - x1) < 1 or (y2 - y1) < 1:
+ continue
+ instance['ignore_flag'] = 0
+ instance['bbox'] = bbox
+ instance['bbox_label'] = int(label)
+ instances.append(instance)
+ data_info['instances'] = instances
+ data_info['dataset_mode'] = self.dataset_mode
+ out_data_list.append(data_info)
+ else:
+ anno = data['grounding']
+ data_info['text'] = anno['caption']
+ regions = anno['regions']
+
+ instances = []
+ phrases = {}
+ for i, region in enumerate(regions):
+ bbox = region['bbox']
+ phrase = region['phrase']
+ tokens_positive = region['tokens_positive']
+ if not isinstance(bbox[0], list):
+ bbox = [bbox]
+ for box in bbox:
+ instance = {}
+ x1, y1, x2, y2 = box
+ inter_w = max(0, min(x2, data['width']) - max(x1, 0))
+ inter_h = max(0, min(y2, data['height']) - max(y1, 0))
+ if inter_w * inter_h == 0:
+ continue
+ if (x2 - x1) < 1 or (y2 - y1) < 1:
+ continue
+ instance['ignore_flag'] = 0
+ instance['bbox'] = box
+ instance['bbox_label'] = i
+ phrases[i] = {
+ 'phrase': phrase,
+ 'tokens_positive': tokens_positive
+ }
+ instances.append(instance)
+ data_info['instances'] = instances
+ data_info['phrases'] = phrases
+ data_info['dataset_mode'] = self.dataset_mode
+ out_data_list.append(data_info)
+
+ del data_list
+ return out_data_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/openimages.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/openimages.py
new file mode 100644
index 0000000000000000000000000000000000000000..a3c6c8ec44fdfe86a653fc6a716009836f7d471c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/openimages.py
@@ -0,0 +1,484 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import csv
+import os.path as osp
+from collections import defaultdict
+from typing import Dict, List, Optional
+
+import numpy as np
+from mmengine.fileio import get_local_path, load
+from mmengine.utils import is_abs
+
+from mmdet.registry import DATASETS
+from .base_det_dataset import BaseDetDataset
+
+
+@DATASETS.register_module()
+class OpenImagesDataset(BaseDetDataset):
+ """Open Images dataset for detection.
+
+ Args:
+ ann_file (str): Annotation file path.
+ label_file (str): File path of the label description file that
+ maps the classes names in MID format to their short
+ descriptions.
+ meta_file (str): File path to get image metas.
+ hierarchy_file (str): The file path of the class hierarchy.
+ image_level_ann_file (str): Human-verified image level annotation,
+ which is used in evaluation.
+ backend_args (dict, optional): Arguments to instantiate the
+ corresponding backend. Defaults to None.
+ """
+
+ METAINFO: dict = dict(dataset_type='oid_v6')
+
+ def __init__(self,
+ label_file: str,
+ meta_file: str,
+ hierarchy_file: str,
+ image_level_ann_file: Optional[str] = None,
+ **kwargs) -> None:
+ self.label_file = label_file
+ self.meta_file = meta_file
+ self.hierarchy_file = hierarchy_file
+ self.image_level_ann_file = image_level_ann_file
+ super().__init__(**kwargs)
+
+ def load_data_list(self) -> List[dict]:
+ """Load annotations from an annotation file named as ``self.ann_file``
+
+ Returns:
+ List[dict]: A list of annotation.
+ """
+ classes_names, label_id_mapping = self._parse_label_file(
+ self.label_file)
+ self._metainfo['classes'] = classes_names
+ self.label_id_mapping = label_id_mapping
+
+ if self.image_level_ann_file is not None:
+ img_level_anns = self._parse_img_level_ann(
+ self.image_level_ann_file)
+ else:
+ img_level_anns = None
+
+ # OpenImagesMetric can get the relation matrix from the dataset meta
+ relation_matrix = self._get_relation_matrix(self.hierarchy_file)
+ self._metainfo['RELATION_MATRIX'] = relation_matrix
+
+ data_list = []
+ with get_local_path(
+ self.ann_file, backend_args=self.backend_args) as local_path:
+ with open(local_path, 'r') as f:
+ reader = csv.reader(f)
+ last_img_id = None
+ instances = []
+ for i, line in enumerate(reader):
+ if i == 0:
+ continue
+ img_id = line[0]
+ if last_img_id is None:
+ last_img_id = img_id
+ label_id = line[2]
+ assert label_id in self.label_id_mapping
+ label = int(self.label_id_mapping[label_id])
+ bbox = [
+ float(line[4]), # xmin
+ float(line[6]), # ymin
+ float(line[5]), # xmax
+ float(line[7]) # ymax
+ ]
+ is_occluded = True if int(line[8]) == 1 else False
+ is_truncated = True if int(line[9]) == 1 else False
+ is_group_of = True if int(line[10]) == 1 else False
+ is_depiction = True if int(line[11]) == 1 else False
+ is_inside = True if int(line[12]) == 1 else False
+
+ instance = dict(
+ bbox=bbox,
+ bbox_label=label,
+ ignore_flag=0,
+ is_occluded=is_occluded,
+ is_truncated=is_truncated,
+ is_group_of=is_group_of,
+ is_depiction=is_depiction,
+ is_inside=is_inside)
+ last_img_path = osp.join(self.data_prefix['img'],
+ f'{last_img_id}.jpg')
+ if img_id != last_img_id:
+ # switch to a new image, record previous image's data.
+ data_info = dict(
+ img_path=last_img_path,
+ img_id=last_img_id,
+ instances=instances,
+ )
+ data_list.append(data_info)
+ instances = []
+ instances.append(instance)
+ last_img_id = img_id
+ data_list.append(
+ dict(
+ img_path=last_img_path,
+ img_id=last_img_id,
+ instances=instances,
+ ))
+
+ # add image metas to data list
+ img_metas = load(
+ self.meta_file, file_format='pkl', backend_args=self.backend_args)
+ assert len(img_metas) == len(data_list)
+ for i, meta in enumerate(img_metas):
+ img_id = data_list[i]['img_id']
+ assert f'{img_id}.jpg' == osp.split(meta['filename'])[-1]
+ h, w = meta['ori_shape'][:2]
+ data_list[i]['height'] = h
+ data_list[i]['width'] = w
+ # denormalize bboxes
+ for j in range(len(data_list[i]['instances'])):
+ data_list[i]['instances'][j]['bbox'][0] *= w
+ data_list[i]['instances'][j]['bbox'][2] *= w
+ data_list[i]['instances'][j]['bbox'][1] *= h
+ data_list[i]['instances'][j]['bbox'][3] *= h
+ # add image-level annotation
+ if img_level_anns is not None:
+ img_labels = []
+ confidences = []
+ img_ann_list = img_level_anns.get(img_id, [])
+ for ann in img_ann_list:
+ img_labels.append(int(ann['image_level_label']))
+ confidences.append(float(ann['confidence']))
+ data_list[i]['image_level_labels'] = np.array(
+ img_labels, dtype=np.int64)
+ data_list[i]['confidences'] = np.array(
+ confidences, dtype=np.float32)
+ return data_list
+
+ def _parse_label_file(self, label_file: str) -> tuple:
+ """Get classes name and index mapping from cls-label-description file.
+
+ Args:
+ label_file (str): File path of the label description file that
+ maps the classes names in MID format to their short
+ descriptions.
+
+ Returns:
+ tuple: Class name of OpenImages.
+ """
+
+ index_list = []
+ classes_names = []
+ with get_local_path(
+ label_file, backend_args=self.backend_args) as local_path:
+ with open(local_path, 'r') as f:
+ reader = csv.reader(f)
+ for line in reader:
+ # self.cat2label[line[0]] = line[1]
+ classes_names.append(line[1])
+ index_list.append(line[0])
+ index_mapping = {index: i for i, index in enumerate(index_list)}
+ return classes_names, index_mapping
+
+ def _parse_img_level_ann(self,
+ img_level_ann_file: str) -> Dict[str, List[dict]]:
+ """Parse image level annotations from csv style ann_file.
+
+ Args:
+ img_level_ann_file (str): CSV style image level annotation
+ file path.
+
+ Returns:
+ Dict[str, List[dict]]: Annotations where item of the defaultdict
+ indicates an image, each of which has (n) dicts.
+ Keys of dicts are:
+
+ - `image_level_label` (int): Label id.
+ - `confidence` (float): Labels that are human-verified to be
+ present in an image have confidence = 1 (positive labels).
+ Labels that are human-verified to be absent from an image
+ have confidence = 0 (negative labels). Machine-generated
+ labels have fractional confidences, generally >= 0.5.
+ The higher the confidence, the smaller the chance for
+ the label to be a false positive.
+ """
+
+ item_lists = defaultdict(list)
+ with get_local_path(
+ img_level_ann_file,
+ backend_args=self.backend_args) as local_path:
+ with open(local_path, 'r') as f:
+ reader = csv.reader(f)
+ for i, line in enumerate(reader):
+ if i == 0:
+ continue
+ img_id = line[0]
+ item_lists[img_id].append(
+ dict(
+ image_level_label=int(
+ self.label_id_mapping[line[2]]),
+ confidence=float(line[3])))
+ return item_lists
+
+ def _get_relation_matrix(self, hierarchy_file: str) -> np.ndarray:
+ """Get the matrix of class hierarchy from the hierarchy file. Hierarchy
+ for 600 classes can be found at https://storage.googleapis.com/openimag
+ es/2018_04/bbox_labels_600_hierarchy_visualizer/circle.html.
+
+ Args:
+ hierarchy_file (str): File path to the hierarchy for classes.
+
+ Returns:
+ np.ndarray: The matrix of the corresponding relationship between
+ the parent class and the child class, of shape
+ (class_num, class_num).
+ """ # noqa
+
+ hierarchy = load(
+ hierarchy_file, file_format='json', backend_args=self.backend_args)
+ class_num = len(self._metainfo['classes'])
+ relation_matrix = np.eye(class_num, class_num)
+ relation_matrix = self._convert_hierarchy_tree(hierarchy,
+ relation_matrix)
+ return relation_matrix
+
+ def _convert_hierarchy_tree(self,
+ hierarchy_map: dict,
+ relation_matrix: np.ndarray,
+ parents: list = [],
+ get_all_parents: bool = True) -> np.ndarray:
+ """Get matrix of the corresponding relationship between the parent
+ class and the child class.
+
+ Args:
+ hierarchy_map (dict): Including label name and corresponding
+ subcategory. Keys of dicts are:
+
+ - `LabeName` (str): Name of the label.
+ - `Subcategory` (dict | list): Corresponding subcategory(ies).
+ relation_matrix (ndarray): The matrix of the corresponding
+ relationship between the parent class and the child class,
+ of shape (class_num, class_num).
+ parents (list): Corresponding parent class.
+ get_all_parents (bool): Whether get all parent names.
+ Default: True
+
+ Returns:
+ ndarray: The matrix of the corresponding relationship between
+ the parent class and the child class, of shape
+ (class_num, class_num).
+ """
+
+ if 'Subcategory' in hierarchy_map:
+ for node in hierarchy_map['Subcategory']:
+ if 'LabelName' in node:
+ children_name = node['LabelName']
+ children_index = self.label_id_mapping[children_name]
+ children = [children_index]
+ else:
+ continue
+ if len(parents) > 0:
+ for parent_index in parents:
+ if get_all_parents:
+ children.append(parent_index)
+ relation_matrix[children_index, parent_index] = 1
+ relation_matrix = self._convert_hierarchy_tree(
+ node, relation_matrix, parents=children)
+ return relation_matrix
+
+ def _join_prefix(self):
+ """Join ``self.data_root`` with annotation path."""
+ super()._join_prefix()
+ if not is_abs(self.label_file) and self.label_file:
+ self.label_file = osp.join(self.data_root, self.label_file)
+ if not is_abs(self.meta_file) and self.meta_file:
+ self.meta_file = osp.join(self.data_root, self.meta_file)
+ if not is_abs(self.hierarchy_file) and self.hierarchy_file:
+ self.hierarchy_file = osp.join(self.data_root, self.hierarchy_file)
+ if self.image_level_ann_file and not is_abs(self.image_level_ann_file):
+ self.image_level_ann_file = osp.join(self.data_root,
+ self.image_level_ann_file)
+
+
+@DATASETS.register_module()
+class OpenImagesChallengeDataset(OpenImagesDataset):
+ """Open Images Challenge dataset for detection.
+
+ Args:
+ ann_file (str): Open Images Challenge box annotation in txt format.
+ """
+
+ METAINFO: dict = dict(dataset_type='oid_challenge')
+
+ def __init__(self, ann_file: str, **kwargs) -> None:
+ if not ann_file.endswith('txt'):
+ raise TypeError('The annotation file of Open Images Challenge '
+ 'should be a txt file.')
+
+ super().__init__(ann_file=ann_file, **kwargs)
+
+ def load_data_list(self) -> List[dict]:
+ """Load annotations from an annotation file named as ``self.ann_file``
+
+ Returns:
+ List[dict]: A list of annotation.
+ """
+ classes_names, label_id_mapping = self._parse_label_file(
+ self.label_file)
+ self._metainfo['classes'] = classes_names
+ self.label_id_mapping = label_id_mapping
+
+ if self.image_level_ann_file is not None:
+ img_level_anns = self._parse_img_level_ann(
+ self.image_level_ann_file)
+ else:
+ img_level_anns = None
+
+ # OpenImagesMetric can get the relation matrix from the dataset meta
+ relation_matrix = self._get_relation_matrix(self.hierarchy_file)
+ self._metainfo['RELATION_MATRIX'] = relation_matrix
+
+ data_list = []
+ with get_local_path(
+ self.ann_file, backend_args=self.backend_args) as local_path:
+ with open(local_path, 'r') as f:
+ lines = f.readlines()
+ i = 0
+ while i < len(lines):
+ instances = []
+ filename = lines[i].rstrip()
+ i += 2
+ img_gt_size = int(lines[i])
+ i += 1
+ for j in range(img_gt_size):
+ sp = lines[i + j].split()
+ instances.append(
+ dict(
+ bbox=[
+ float(sp[1]),
+ float(sp[2]),
+ float(sp[3]),
+ float(sp[4])
+ ],
+ bbox_label=int(sp[0]) - 1, # labels begin from 1
+ ignore_flag=0,
+ is_group_ofs=True if int(sp[5]) == 1 else False))
+ i += img_gt_size
+ data_list.append(
+ dict(
+ img_path=osp.join(self.data_prefix['img'], filename),
+ instances=instances,
+ ))
+
+ # add image metas to data list
+ img_metas = load(
+ self.meta_file, file_format='pkl', backend_args=self.backend_args)
+ assert len(img_metas) == len(data_list)
+ for i, meta in enumerate(img_metas):
+ img_id = osp.split(data_list[i]['img_path'])[-1][:-4]
+ assert img_id == osp.split(meta['filename'])[-1][:-4]
+ h, w = meta['ori_shape'][:2]
+ data_list[i]['height'] = h
+ data_list[i]['width'] = w
+ data_list[i]['img_id'] = img_id
+ # denormalize bboxes
+ for j in range(len(data_list[i]['instances'])):
+ data_list[i]['instances'][j]['bbox'][0] *= w
+ data_list[i]['instances'][j]['bbox'][2] *= w
+ data_list[i]['instances'][j]['bbox'][1] *= h
+ data_list[i]['instances'][j]['bbox'][3] *= h
+ # add image-level annotation
+ if img_level_anns is not None:
+ img_labels = []
+ confidences = []
+ img_ann_list = img_level_anns.get(img_id, [])
+ for ann in img_ann_list:
+ img_labels.append(int(ann['image_level_label']))
+ confidences.append(float(ann['confidence']))
+ data_list[i]['image_level_labels'] = np.array(
+ img_labels, dtype=np.int64)
+ data_list[i]['confidences'] = np.array(
+ confidences, dtype=np.float32)
+ return data_list
+
+ def _parse_label_file(self, label_file: str) -> tuple:
+ """Get classes name and index mapping from cls-label-description file.
+
+ Args:
+ label_file (str): File path of the label description file that
+ maps the classes names in MID format to their short
+ descriptions.
+
+ Returns:
+ tuple: Class name of OpenImages.
+ """
+ label_list = []
+ id_list = []
+ index_mapping = {}
+ with get_local_path(
+ label_file, backend_args=self.backend_args) as local_path:
+ with open(local_path, 'r') as f:
+ reader = csv.reader(f)
+ for line in reader:
+ label_name = line[0]
+ label_id = int(line[2])
+ label_list.append(line[1])
+ id_list.append(label_id)
+ index_mapping[label_name] = label_id - 1
+ indexes = np.argsort(id_list)
+ classes_names = []
+ for index in indexes:
+ classes_names.append(label_list[index])
+ return classes_names, index_mapping
+
+ def _parse_img_level_ann(self, image_level_ann_file):
+ """Parse image level annotations from csv style ann_file.
+
+ Args:
+ image_level_ann_file (str): CSV style image level annotation
+ file path.
+
+ Returns:
+ defaultdict[list[dict]]: Annotations where item of the defaultdict
+ indicates an image, each of which has (n) dicts.
+ Keys of dicts are:
+
+ - `image_level_label` (int): of shape 1.
+ - `confidence` (float): of shape 1.
+ """
+
+ item_lists = defaultdict(list)
+ with get_local_path(
+ image_level_ann_file,
+ backend_args=self.backend_args) as local_path:
+ with open(local_path, 'r') as f:
+ reader = csv.reader(f)
+ i = -1
+ for line in reader:
+ i += 1
+ if i == 0:
+ continue
+ else:
+ img_id = line[0]
+ label_id = line[1]
+ assert label_id in self.label_id_mapping
+ image_level_label = int(
+ self.label_id_mapping[label_id])
+ confidence = float(line[2])
+ item_lists[img_id].append(
+ dict(
+ image_level_label=image_level_label,
+ confidence=confidence))
+ return item_lists
+
+ def _get_relation_matrix(self, hierarchy_file: str) -> np.ndarray:
+ """Get the matrix of class hierarchy from the hierarchy file.
+
+ Args:
+ hierarchy_file (str): File path to the hierarchy for classes.
+
+ Returns:
+ np.ndarray: The matrix of the corresponding
+ relationship between the parent class and the child class,
+ of shape (class_num, class_num).
+ """
+ with get_local_path(
+ hierarchy_file, backend_args=self.backend_args) as local_path:
+ class_label_tree = np.load(local_path, allow_pickle=True)
+ return class_label_tree[1:, 1:]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/refcoco.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/refcoco.py
new file mode 100644
index 0000000000000000000000000000000000000000..0dae75fd547216a5b69033cc821b93a1d9ac6abc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/refcoco.py
@@ -0,0 +1,163 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import collections
+import os.path as osp
+import random
+from typing import Dict, List
+
+import mmengine
+from mmengine.dataset import BaseDataset
+
+from mmdet.registry import DATASETS
+
+
+@DATASETS.register_module()
+class RefCocoDataset(BaseDataset):
+ """RefCOCO dataset.
+
+ The `Refcoco` and `Refcoco+` dataset is based on
+ `ReferItGame: Referring to Objects in Photographs of Natural Scenes
+ `_.
+
+ The `Refcocog` dataset is based on
+ `Generation and Comprehension of Unambiguous Object Descriptions
+ `_.
+
+ Args:
+ ann_file (str): Annotation file path.
+ data_root (str): The root directory for ``data_prefix`` and
+ ``ann_file``. Defaults to ''.
+ data_prefix (str): Prefix for training data.
+ split_file (str): Split file path.
+ split (str): Split name. Defaults to 'train'.
+ text_mode (str): Text mode. Defaults to 'random'.
+ **kwargs: Other keyword arguments in :class:`BaseDataset`.
+ """
+
+ def __init__(self,
+ data_root: str,
+ ann_file: str,
+ split_file: str,
+ data_prefix: Dict,
+ split: str = 'train',
+ text_mode: str = 'random',
+ **kwargs):
+ self.split_file = split_file
+ self.split = split
+
+ assert text_mode in ['original', 'random', 'concat', 'select_first']
+ self.text_mode = text_mode
+ super().__init__(
+ data_root=data_root,
+ data_prefix=data_prefix,
+ ann_file=ann_file,
+ **kwargs,
+ )
+
+ def _join_prefix(self):
+ if not mmengine.is_abs(self.split_file) and self.split_file:
+ self.split_file = osp.join(self.data_root, self.split_file)
+
+ return super()._join_prefix()
+
+ def _init_refs(self):
+ """Initialize the refs for RefCOCO."""
+ anns, imgs = {}, {}
+ for ann in self.instances['annotations']:
+ anns[ann['id']] = ann
+ for img in self.instances['images']:
+ imgs[img['id']] = img
+
+ refs, ref_to_ann = {}, {}
+ for ref in self.splits:
+ # ids
+ ref_id = ref['ref_id']
+ ann_id = ref['ann_id']
+ # add mapping related to ref
+ refs[ref_id] = ref
+ ref_to_ann[ref_id] = anns[ann_id]
+
+ self.refs = refs
+ self.ref_to_ann = ref_to_ann
+
+ def load_data_list(self) -> List[dict]:
+ """Load data list."""
+ self.splits = mmengine.load(self.split_file, file_format='pkl')
+ self.instances = mmengine.load(self.ann_file, file_format='json')
+ self._init_refs()
+ img_prefix = self.data_prefix['img_path']
+
+ ref_ids = [
+ ref['ref_id'] for ref in self.splits if ref['split'] == self.split
+ ]
+ full_anno = []
+ for ref_id in ref_ids:
+ ref = self.refs[ref_id]
+ ann = self.ref_to_ann[ref_id]
+ ann.update(ref)
+ full_anno.append(ann)
+
+ image_id_list = []
+ final_anno = {}
+ for anno in full_anno:
+ image_id_list.append(anno['image_id'])
+ final_anno[anno['ann_id']] = anno
+ annotations = [value for key, value in final_anno.items()]
+
+ coco_train_id = []
+ image_annot = {}
+ for i in range(len(self.instances['images'])):
+ coco_train_id.append(self.instances['images'][i]['id'])
+ image_annot[self.instances['images'][i]
+ ['id']] = self.instances['images'][i]
+
+ images = []
+ for image_id in list(set(image_id_list)):
+ images += [image_annot[image_id]]
+
+ data_list = []
+
+ grounding_dict = collections.defaultdict(list)
+ for anno in annotations:
+ image_id = int(anno['image_id'])
+ grounding_dict[image_id].append(anno)
+
+ join_path = mmengine.fileio.get_file_backend(img_prefix).join_path
+ for image in images:
+ img_id = image['id']
+ instances = []
+ sentences = []
+ for grounding_anno in grounding_dict[img_id]:
+ texts = [x['raw'].lower() for x in grounding_anno['sentences']]
+ # random select one text
+ if self.text_mode == 'random':
+ idx = random.randint(0, len(texts) - 1)
+ text = [texts[idx]]
+ # concat all texts
+ elif self.text_mode == 'concat':
+ text = [''.join(texts)]
+ # select the first text
+ elif self.text_mode == 'select_first':
+ text = [texts[0]]
+ # use all texts
+ elif self.text_mode == 'original':
+ text = texts
+ else:
+ raise ValueError(f'Invalid text mode "{self.text_mode}".')
+ ins = [{
+ 'mask': grounding_anno['segmentation'],
+ 'ignore_flag': 0
+ }] * len(text)
+ instances.extend(ins)
+ sentences.extend(text)
+ data_info = {
+ 'img_path': join_path(img_prefix, image['file_name']),
+ 'img_id': img_id,
+ 'instances': instances,
+ 'text': sentences
+ }
+ data_list.append(data_info)
+
+ if len(data_list) == 0:
+ raise ValueError(f'No sample in split "{self.split}".')
+
+ return data_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/reid_dataset.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/reid_dataset.py
new file mode 100644
index 0000000000000000000000000000000000000000..1eed3ee4f0358edf59d19695c2b28394336dffd3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/reid_dataset.py
@@ -0,0 +1,127 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import os.path as osp
+from collections import defaultdict
+from typing import Any, Dict, List
+
+import numpy as np
+from mmengine.dataset import BaseDataset
+from mmengine.utils import check_file_exist
+
+from mmdet.registry import DATASETS
+
+
+@DATASETS.register_module()
+class ReIDDataset(BaseDataset):
+ """Dataset for ReID.
+
+ Args:
+ triplet_sampler (dict, optional): The sampler for hard mining
+ triplet loss. Defaults to None.
+ keys: num_ids (int): The number of person ids.
+ ins_per_id (int): The number of image for each person.
+ """
+
+ def __init__(self, triplet_sampler: dict = None, *args, **kwargs):
+ self.triplet_sampler = triplet_sampler
+ super().__init__(*args, **kwargs)
+
+ def load_data_list(self) -> List[dict]:
+ """Load annotations from an annotation file named as ''self.ann_file''.
+
+ Returns:
+ list[dict]: A list of annotation.
+ """
+ assert isinstance(self.ann_file, str)
+ check_file_exist(self.ann_file)
+ data_list = []
+ with open(self.ann_file) as f:
+ samples = [x.strip().split(' ') for x in f.readlines()]
+ for filename, gt_label in samples:
+ info = dict(img_prefix=self.data_prefix)
+ if self.data_prefix['img_path'] is not None:
+ info['img_path'] = osp.join(self.data_prefix['img_path'],
+ filename)
+ else:
+ info['img_path'] = filename
+ info['gt_label'] = np.array(gt_label, dtype=np.int64)
+ data_list.append(info)
+ self._parse_ann_info(data_list)
+ return data_list
+
+ def _parse_ann_info(self, data_list: List[dict]):
+ """Parse person id annotations."""
+ index_tmp_dic = defaultdict(list) # pid->[idx1,...,idxN]
+ self.index_dic = dict() # pid->array([idx1,...,idxN])
+ for idx, info in enumerate(data_list):
+ pid = info['gt_label']
+ index_tmp_dic[int(pid)].append(idx)
+ for pid, idxs in index_tmp_dic.items():
+ self.index_dic[pid] = np.asarray(idxs, dtype=np.int64)
+ self.pids = np.asarray(list(self.index_dic.keys()), dtype=np.int64)
+
+ def prepare_data(self, idx: int) -> Any:
+ """Get data processed by ''self.pipeline''.
+
+ Args:
+ idx (int): The index of ''data_info''
+
+ Returns:
+ Any: Depends on ''self.pipeline''
+ """
+ data_info = self.get_data_info(idx)
+ if self.triplet_sampler is not None:
+ img_info = self.triplet_sampling(data_info['gt_label'],
+ **self.triplet_sampler)
+ data_info = copy.deepcopy(img_info) # triplet -> list
+ else:
+ data_info = copy.deepcopy(data_info) # no triplet -> dict
+ return self.pipeline(data_info)
+
+ def triplet_sampling(self,
+ pos_pid,
+ num_ids: int = 8,
+ ins_per_id: int = 4) -> Dict:
+ """Triplet sampler for hard mining triplet loss. First, for one
+ pos_pid, random sample ins_per_id images with same person id.
+
+ Then, random sample num_ids - 1 images for each negative id.
+ Finally, random sample ins_per_id images for each negative id.
+
+ Args:
+ pos_pid (ndarray): The person id of the anchor.
+ num_ids (int): The number of person ids.
+ ins_per_id (int): The number of images for each person.
+
+ Returns:
+ Dict: Annotation information of num_ids X ins_per_id images.
+ """
+ assert len(self.pids) >= num_ids, \
+ 'The number of person ids in the training set must ' \
+ 'be greater than the number of person ids in the sample.'
+
+ pos_idxs = self.index_dic[int(
+ pos_pid)] # all positive idxs for pos_pid
+ idxs_list = []
+ # select positive samplers
+ idxs_list.extend(pos_idxs[np.random.choice(
+ pos_idxs.shape[0], ins_per_id, replace=True)])
+ # select negative ids
+ neg_pids = np.random.choice(
+ [i for i, _ in enumerate(self.pids) if i != pos_pid],
+ num_ids - 1,
+ replace=False)
+ # select negative samplers for each negative id
+ for neg_pid in neg_pids:
+ neg_idxs = self.index_dic[neg_pid]
+ idxs_list.extend(neg_idxs[np.random.choice(
+ neg_idxs.shape[0], ins_per_id, replace=True)])
+ # return the final triplet batch
+ triplet_img_infos = []
+ for idx in idxs_list:
+ triplet_img_infos.append(copy.deepcopy(self.get_data_info(idx)))
+ # Collect data_list scatters (list of dict -> dict of list)
+ out = dict()
+ for key in triplet_img_infos[0].keys():
+ out[key] = [_info[key] for _info in triplet_img_infos]
+ return out
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..9ea0e4cb0628fc23bc034c51e503d8ceca5ee90c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/__init__.py
@@ -0,0 +1,16 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .batch_sampler import (AspectRatioBatchSampler,
+ MultiDataAspectRatioBatchSampler,
+ TrackAspectRatioBatchSampler)
+from .class_aware_sampler import ClassAwareSampler
+from .custom_sample_size_sampler import CustomSampleSizeSampler
+from .multi_data_sampler import MultiDataSampler
+from .multi_source_sampler import GroupMultiSourceSampler, MultiSourceSampler
+from .track_img_sampler import TrackImgSampler
+
+__all__ = [
+ 'ClassAwareSampler', 'AspectRatioBatchSampler', 'MultiSourceSampler',
+ 'GroupMultiSourceSampler', 'TrackImgSampler',
+ 'TrackAspectRatioBatchSampler', 'MultiDataSampler',
+ 'MultiDataAspectRatioBatchSampler', 'CustomSampleSizeSampler'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/batch_sampler.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/batch_sampler.py
new file mode 100644
index 0000000000000000000000000000000000000000..c17789c4e3ea51f1fa140d039a679f797a7660f6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/batch_sampler.py
@@ -0,0 +1,193 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Sequence
+
+from torch.utils.data import BatchSampler, Sampler
+
+from mmdet.datasets.samplers.track_img_sampler import TrackImgSampler
+from mmdet.registry import DATA_SAMPLERS
+
+
+# TODO: maybe replace with a data_loader wrapper
+@DATA_SAMPLERS.register_module()
+class AspectRatioBatchSampler(BatchSampler):
+ """A sampler wrapper for grouping images with similar aspect ratio (< 1 or.
+
+ >= 1) into a same batch.
+
+ Args:
+ sampler (Sampler): Base sampler.
+ batch_size (int): Size of mini-batch.
+ drop_last (bool): If ``True``, the sampler will drop the last batch if
+ its size would be less than ``batch_size``.
+ """
+
+ def __init__(self,
+ sampler: Sampler,
+ batch_size: int,
+ drop_last: bool = False) -> None:
+ if not isinstance(sampler, Sampler):
+ raise TypeError('sampler should be an instance of ``Sampler``, '
+ f'but got {sampler}')
+ if not isinstance(batch_size, int) or batch_size <= 0:
+ raise ValueError('batch_size should be a positive integer value, '
+ f'but got batch_size={batch_size}')
+ self.sampler = sampler
+ self.batch_size = batch_size
+ self.drop_last = drop_last
+ # two groups for w < h and w >= h
+ self._aspect_ratio_buckets = [[] for _ in range(2)]
+
+ def __iter__(self) -> Sequence[int]:
+ for idx in self.sampler:
+ data_info = self.sampler.dataset.get_data_info(idx)
+ width, height = data_info['width'], data_info['height']
+ bucket_id = 0 if width < height else 1
+ bucket = self._aspect_ratio_buckets[bucket_id]
+ bucket.append(idx)
+ # yield a batch of indices in the same aspect ratio group
+ if len(bucket) == self.batch_size:
+ yield bucket[:]
+ del bucket[:]
+
+ # yield the rest data and reset the bucket
+ left_data = self._aspect_ratio_buckets[0] + self._aspect_ratio_buckets[
+ 1]
+ self._aspect_ratio_buckets = [[] for _ in range(2)]
+ while len(left_data) > 0:
+ if len(left_data) <= self.batch_size:
+ if not self.drop_last:
+ yield left_data[:]
+ left_data = []
+ else:
+ yield left_data[:self.batch_size]
+ left_data = left_data[self.batch_size:]
+
+ def __len__(self) -> int:
+ if self.drop_last:
+ return len(self.sampler) // self.batch_size
+ else:
+ return (len(self.sampler) + self.batch_size - 1) // self.batch_size
+
+
+@DATA_SAMPLERS.register_module()
+class TrackAspectRatioBatchSampler(AspectRatioBatchSampler):
+ """A sampler wrapper for grouping images with similar aspect ratio (< 1 or.
+
+ >= 1) into a same batch.
+
+ Args:
+ sampler (Sampler): Base sampler.
+ batch_size (int): Size of mini-batch.
+ drop_last (bool): If ``True``, the sampler will drop the last batch if
+ its size would be less than ``batch_size``.
+ """
+
+ def __iter__(self) -> Sequence[int]:
+ for idx in self.sampler:
+ # hard code to solve TrackImgSampler
+ if isinstance(self.sampler, TrackImgSampler):
+ video_idx, _ = idx
+ else:
+ video_idx = idx
+ # video_idx
+ data_info = self.sampler.dataset.get_data_info(video_idx)
+ # data_info {video_id, images, video_length}
+ img_data_info = data_info['images'][0]
+ width, height = img_data_info['width'], img_data_info['height']
+ bucket_id = 0 if width < height else 1
+ bucket = self._aspect_ratio_buckets[bucket_id]
+ bucket.append(idx)
+ # yield a batch of indices in the same aspect ratio group
+ if len(bucket) == self.batch_size:
+ yield bucket[:]
+ del bucket[:]
+
+ # yield the rest data and reset the bucket
+ left_data = self._aspect_ratio_buckets[0] + self._aspect_ratio_buckets[
+ 1]
+ self._aspect_ratio_buckets = [[] for _ in range(2)]
+ while len(left_data) > 0:
+ if len(left_data) <= self.batch_size:
+ if not self.drop_last:
+ yield left_data[:]
+ left_data = []
+ else:
+ yield left_data[:self.batch_size]
+ left_data = left_data[self.batch_size:]
+
+
+@DATA_SAMPLERS.register_module()
+class MultiDataAspectRatioBatchSampler(BatchSampler):
+ """A sampler wrapper for grouping images with similar aspect ratio (< 1 or.
+
+ >= 1) into a same batch for multi-source datasets.
+
+ Args:
+ sampler (Sampler): Base sampler.
+ batch_size (Sequence(int)): Size of mini-batch for multi-source
+ datasets.
+ num_datasets(int): Number of multi-source datasets.
+ drop_last (bool): If ``True``, the sampler will drop the last batch if
+ its size would be less than ``batch_size``.
+ """
+
+ def __init__(self,
+ sampler: Sampler,
+ batch_size: Sequence[int],
+ num_datasets: int,
+ drop_last: bool = True) -> None:
+ if not isinstance(sampler, Sampler):
+ raise TypeError('sampler should be an instance of ``Sampler``, '
+ f'but got {sampler}')
+ self.sampler = sampler
+ self.batch_size = batch_size
+ self.num_datasets = num_datasets
+ self.drop_last = drop_last
+ # two groups for w < h and w >= h for each dataset --> 2 * num_datasets
+ self._buckets = [[] for _ in range(2 * self.num_datasets)]
+
+ def __iter__(self) -> Sequence[int]:
+ for idx in self.sampler:
+ data_info = self.sampler.dataset.get_data_info(idx)
+ width, height = data_info['width'], data_info['height']
+ dataset_source_idx = self.sampler.dataset.get_dataset_source(idx)
+ aspect_ratio_bucket_id = 0 if width < height else 1
+ bucket_id = dataset_source_idx * 2 + aspect_ratio_bucket_id
+ bucket = self._buckets[bucket_id]
+ bucket.append(idx)
+ # yield a batch of indices in the same aspect ratio group
+ if len(bucket) == self.batch_size[dataset_source_idx]:
+ yield bucket[:]
+ del bucket[:]
+
+ # yield the rest data and reset the bucket
+ for i in range(self.num_datasets):
+ left_data = self._buckets[i * 2 + 0] + self._buckets[i * 2 + 1]
+ while len(left_data) > 0:
+ if len(left_data) <= self.batch_size[i]:
+ if not self.drop_last:
+ yield left_data[:]
+ left_data = []
+ else:
+ yield left_data[:self.batch_size[i]]
+ left_data = left_data[self.batch_size[i]:]
+
+ self._buckets = [[] for _ in range(2 * self.num_datasets)]
+
+ def __len__(self) -> int:
+ sizes = [0 for _ in range(self.num_datasets)]
+ for idx in self.sampler:
+ dataset_source_idx = self.sampler.dataset.get_dataset_source(idx)
+ sizes[dataset_source_idx] += 1
+
+ if self.drop_last:
+ lens = 0
+ for i in range(self.num_datasets):
+ lens += sizes[i] // self.batch_size[i]
+ return lens
+ else:
+ lens = 0
+ for i in range(self.num_datasets):
+ lens += (sizes[i] + self.batch_size[i] -
+ 1) // self.batch_size[i]
+ return lens
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/class_aware_sampler.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/class_aware_sampler.py
new file mode 100644
index 0000000000000000000000000000000000000000..6ca2f9b3ffb7c780ab25cc3704b67589763259e0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/class_aware_sampler.py
@@ -0,0 +1,192 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+from typing import Dict, Iterator, Optional, Union
+
+import numpy as np
+import torch
+from mmengine.dataset import BaseDataset
+from mmengine.dist import get_dist_info, sync_random_seed
+from torch.utils.data import Sampler
+
+from mmdet.registry import DATA_SAMPLERS
+
+
+@DATA_SAMPLERS.register_module()
+class ClassAwareSampler(Sampler):
+ r"""Sampler that restricts data loading to the label of the dataset.
+
+ A class-aware sampling strategy to effectively tackle the
+ non-uniform class distribution. The length of the training data is
+ consistent with source data. Simple improvements based on `Relay
+ Backpropagation for Effective Learning of Deep Convolutional
+ Neural Networks `_
+
+ The implementation logic is referred to
+ https://github.com/Sense-X/TSD/blob/master/mmdet/datasets/samplers/distributed_classaware_sampler.py
+
+ Args:
+ dataset: Dataset used for sampling.
+ seed (int, optional): random seed used to shuffle the sampler.
+ This number should be identical across all
+ processes in the distributed group. Defaults to None.
+ num_sample_class (int): The number of samples taken from each
+ per-label list. Defaults to 1.
+ """
+
+ def __init__(self,
+ dataset: BaseDataset,
+ seed: Optional[int] = None,
+ num_sample_class: int = 1) -> None:
+ rank, world_size = get_dist_info()
+ self.rank = rank
+ self.world_size = world_size
+
+ self.dataset = dataset
+ self.epoch = 0
+ # Must be the same across all workers. If None, will use a
+ # random seed shared among workers
+ # (require synchronization among all workers)
+ if seed is None:
+ seed = sync_random_seed()
+ self.seed = seed
+
+ # The number of samples taken from each per-label list
+ assert num_sample_class > 0 and isinstance(num_sample_class, int)
+ self.num_sample_class = num_sample_class
+ # Get per-label image list from dataset
+ self.cat_dict = self.get_cat2imgs()
+
+ self.num_samples = int(math.ceil(len(self.dataset) * 1.0 / world_size))
+ self.total_size = self.num_samples * self.world_size
+
+ # get number of images containing each category
+ self.num_cat_imgs = [len(x) for x in self.cat_dict.values()]
+ # filter labels without images
+ self.valid_cat_inds = [
+ i for i, length in enumerate(self.num_cat_imgs) if length != 0
+ ]
+ self.num_classes = len(self.valid_cat_inds)
+
+ def get_cat2imgs(self) -> Dict[int, list]:
+ """Get a dict with class as key and img_ids as values.
+
+ Returns:
+ dict[int, list]: A dict of per-label image list,
+ the item of the dict indicates a label index,
+ corresponds to the image index that contains the label.
+ """
+ classes = self.dataset.metainfo.get('classes', None)
+ if classes is None:
+ raise ValueError('dataset metainfo must contain `classes`')
+ # sort the label index
+ cat2imgs = {i: [] for i in range(len(classes))}
+ for i in range(len(self.dataset)):
+ cat_ids = set(self.dataset.get_cat_ids(i))
+ for cat in cat_ids:
+ cat2imgs[cat].append(i)
+ return cat2imgs
+
+ def __iter__(self) -> Iterator[int]:
+ # deterministically shuffle based on epoch
+ g = torch.Generator()
+ g.manual_seed(self.epoch + self.seed)
+
+ # initialize label list
+ label_iter_list = RandomCycleIter(self.valid_cat_inds, generator=g)
+ # initialize each per-label image list
+ data_iter_dict = dict()
+ for i in self.valid_cat_inds:
+ data_iter_dict[i] = RandomCycleIter(self.cat_dict[i], generator=g)
+
+ def gen_cat_img_inds(cls_list, data_dict, num_sample_cls):
+ """Traverse the categories and extract `num_sample_cls` image
+ indexes of the corresponding categories one by one."""
+ id_indices = []
+ for _ in range(len(cls_list)):
+ cls_idx = next(cls_list)
+ for _ in range(num_sample_cls):
+ id = next(data_dict[cls_idx])
+ id_indices.append(id)
+ return id_indices
+
+ # deterministically shuffle based on epoch
+ num_bins = int(
+ math.ceil(self.total_size * 1.0 / self.num_classes /
+ self.num_sample_class))
+ indices = []
+ for i in range(num_bins):
+ indices += gen_cat_img_inds(label_iter_list, data_iter_dict,
+ self.num_sample_class)
+
+ # fix extra samples to make it evenly divisible
+ if len(indices) >= self.total_size:
+ indices = indices[:self.total_size]
+ else:
+ indices += indices[:(self.total_size - len(indices))]
+ assert len(indices) == self.total_size
+
+ # subsample
+ offset = self.num_samples * self.rank
+ indices = indices[offset:offset + self.num_samples]
+ assert len(indices) == self.num_samples
+
+ return iter(indices)
+
+ def __len__(self) -> int:
+ """The number of samples in this rank."""
+ return self.num_samples
+
+ def set_epoch(self, epoch: int) -> None:
+ """Sets the epoch for this sampler.
+
+ When :attr:`shuffle=True`, this ensures all replicas use a different
+ random ordering for each epoch. Otherwise, the next iteration of this
+ sampler will yield the same ordering.
+
+ Args:
+ epoch (int): Epoch number.
+ """
+ self.epoch = epoch
+
+
+class RandomCycleIter:
+ """Shuffle the list and do it again after the list have traversed.
+
+ The implementation logic is referred to
+ https://github.com/wutong16/DistributionBalancedLoss/blob/master/mllt/datasets/loader/sampler.py
+
+ Example:
+ >>> label_list = [0, 1, 2, 4, 5]
+ >>> g = torch.Generator()
+ >>> g.manual_seed(0)
+ >>> label_iter_list = RandomCycleIter(label_list, generator=g)
+ >>> index = next(label_iter_list)
+ Args:
+ data (list or ndarray): The data that needs to be shuffled.
+ generator: An torch.Generator object, which is used in setting the seed
+ for generating random numbers.
+ """ # noqa: W605
+
+ def __init__(self,
+ data: Union[list, np.ndarray],
+ generator: torch.Generator = None) -> None:
+ self.data = data
+ self.length = len(data)
+ self.index = torch.randperm(self.length, generator=generator).numpy()
+ self.i = 0
+ self.generator = generator
+
+ def __iter__(self) -> Iterator:
+ return self
+
+ def __len__(self) -> int:
+ return len(self.data)
+
+ def __next__(self):
+ if self.i == self.length:
+ self.index = torch.randperm(
+ self.length, generator=self.generator).numpy()
+ self.i = 0
+ idx = self.data[self.index[self.i]]
+ self.i += 1
+ return idx
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/custom_sample_size_sampler.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/custom_sample_size_sampler.py
new file mode 100644
index 0000000000000000000000000000000000000000..6bedf6c66be81b091a6424bae6788953ba7763a3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/custom_sample_size_sampler.py
@@ -0,0 +1,111 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+from typing import Iterator, Optional, Sequence, Sized
+
+import torch
+from mmengine.dist import get_dist_info, sync_random_seed
+from torch.utils.data import Sampler
+
+from mmdet.registry import DATA_SAMPLERS
+from .class_aware_sampler import RandomCycleIter
+
+
+@DATA_SAMPLERS.register_module()
+class CustomSampleSizeSampler(Sampler):
+
+ def __init__(self,
+ dataset: Sized,
+ dataset_size: Sequence[int],
+ ratio_mode: bool = False,
+ seed: Optional[int] = None,
+ round_up: bool = True) -> None:
+ assert len(dataset.datasets) == len(dataset_size)
+ rank, world_size = get_dist_info()
+ self.rank = rank
+ self.world_size = world_size
+
+ self.dataset = dataset
+ if seed is None:
+ seed = sync_random_seed()
+ self.seed = seed
+ self.epoch = 0
+ self.round_up = round_up
+
+ total_size = 0
+ total_size_fake = 0
+ self.dataset_index = []
+ self.dataset_cycle_iter = []
+ new_dataset_size = []
+ for dataset, size in zip(dataset.datasets, dataset_size):
+ self.dataset_index.append(
+ list(range(total_size_fake,
+ len(dataset) + total_size_fake)))
+ total_size_fake += len(dataset)
+ if size == -1:
+ total_size += len(dataset)
+ self.dataset_cycle_iter.append(None)
+ new_dataset_size.append(-1)
+ else:
+ if ratio_mode:
+ size = int(size * len(dataset))
+ assert size <= len(
+ dataset
+ ), f'dataset size {size} is larger than ' \
+ f'dataset length {len(dataset)}'
+ total_size += size
+ new_dataset_size.append(size)
+
+ g = torch.Generator()
+ g.manual_seed(self.seed)
+ self.dataset_cycle_iter.append(
+ RandomCycleIter(self.dataset_index[-1], generator=g))
+ self.dataset_size = new_dataset_size
+
+ if self.round_up:
+ self.num_samples = math.ceil(total_size / world_size)
+ self.total_size = self.num_samples * self.world_size
+ else:
+ self.num_samples = math.ceil((total_size - rank) / world_size)
+ self.total_size = total_size
+
+ def __iter__(self) -> Iterator[int]:
+ """Iterate the indices."""
+ # deterministically shuffle based on epoch and seed
+ g = torch.Generator()
+ g.manual_seed(self.seed + self.epoch)
+
+ out_index = []
+ for data_size, data_index, cycle_iter in zip(self.dataset_size,
+ self.dataset_index,
+ self.dataset_cycle_iter):
+ if data_size == -1:
+ out_index += data_index
+ else:
+ index = [next(cycle_iter) for _ in range(data_size)]
+ out_index += index
+
+ index = torch.randperm(len(out_index), generator=g).numpy().tolist()
+ indices = [out_index[i] for i in index]
+
+ if self.round_up:
+ indices = (
+ indices *
+ int(self.total_size / len(indices) + 1))[:self.total_size]
+ indices = indices[self.rank:self.total_size:self.world_size]
+ return iter(indices)
+
+ def __len__(self) -> int:
+ """The number of samples in this rank."""
+ return self.num_samples
+
+ def set_epoch(self, epoch: int) -> None:
+ """Sets the epoch for this sampler.
+
+ When :attr:`shuffle=True`, this ensures all replicas use a different
+ random ordering for each epoch. Otherwise, the next iteration of this
+ sampler will yield the same ordering.
+
+ Args:
+ epoch (int): Epoch number.
+ """
+ self.epoch = epoch
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/multi_data_sampler.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/multi_data_sampler.py
new file mode 100644
index 0000000000000000000000000000000000000000..c3a4b60d84122ce9eb2090095e9744c2bd73cc3d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/multi_data_sampler.py
@@ -0,0 +1,110 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+from typing import Iterator, Optional, Sequence, Sized
+
+import torch
+from mmengine.dist import get_dist_info, sync_random_seed
+from mmengine.registry import DATA_SAMPLERS
+from torch.utils.data import Sampler
+
+
+@DATA_SAMPLERS.register_module()
+class MultiDataSampler(Sampler):
+ """The default data sampler for both distributed and non-distributed
+ environment.
+
+ It has several differences from the PyTorch ``DistributedSampler`` as
+ below:
+
+ 1. This sampler supports non-distributed environment.
+
+ 2. The round up behaviors are a little different.
+
+ - If ``round_up=True``, this sampler will add extra samples to make the
+ number of samples is evenly divisible by the world size. And
+ this behavior is the same as the ``DistributedSampler`` with
+ ``drop_last=False``.
+ - If ``round_up=False``, this sampler won't remove or add any samples
+ while the ``DistributedSampler`` with ``drop_last=True`` will remove
+ tail samples.
+
+ Args:
+ dataset (Sized): The dataset.
+ dataset_ratio (Sequence(int)) The ratios of different datasets.
+ seed (int, optional): Random seed used to shuffle the sampler if
+ :attr:`shuffle=True`. This number should be identical across all
+ processes in the distributed group. Defaults to None.
+ round_up (bool): Whether to add extra samples to make the number of
+ samples evenly divisible by the world size. Defaults to True.
+ """
+
+ def __init__(self,
+ dataset: Sized,
+ dataset_ratio: Sequence[int],
+ seed: Optional[int] = None,
+ round_up: bool = True) -> None:
+ rank, world_size = get_dist_info()
+ self.rank = rank
+ self.world_size = world_size
+
+ self.dataset = dataset
+ self.dataset_ratio = dataset_ratio
+
+ if seed is None:
+ seed = sync_random_seed()
+ self.seed = seed
+ self.epoch = 0
+ self.round_up = round_up
+
+ if self.round_up:
+ self.num_samples = math.ceil(len(self.dataset) / world_size)
+ self.total_size = self.num_samples * self.world_size
+ else:
+ self.num_samples = math.ceil(
+ (len(self.dataset) - rank) / world_size)
+ self.total_size = len(self.dataset)
+
+ self.sizes = [len(dataset) for dataset in self.dataset.datasets]
+
+ dataset_weight = [
+ torch.ones(s) * max(self.sizes) / s * r / sum(self.dataset_ratio)
+ for i, (r, s) in enumerate(zip(self.dataset_ratio, self.sizes))
+ ]
+ self.weights = torch.cat(dataset_weight)
+
+ def __iter__(self) -> Iterator[int]:
+ """Iterate the indices."""
+ # deterministically shuffle based on epoch and seed
+ g = torch.Generator()
+ g.manual_seed(self.seed + self.epoch)
+
+ indices = torch.multinomial(
+ self.weights, len(self.weights), generator=g,
+ replacement=True).tolist()
+
+ # add extra samples to make it evenly divisible
+ if self.round_up:
+ indices = (
+ indices *
+ int(self.total_size / len(indices) + 1))[:self.total_size]
+
+ # subsample
+ indices = indices[self.rank:self.total_size:self.world_size]
+
+ return iter(indices)
+
+ def __len__(self) -> int:
+ """The number of samples in this rank."""
+ return self.num_samples
+
+ def set_epoch(self, epoch: int) -> None:
+ """Sets the epoch for this sampler.
+
+ When :attr:`shuffle=True`, this ensures all replicas use a different
+ random ordering for each epoch. Otherwise, the next iteration of this
+ sampler will yield the same ordering.
+
+ Args:
+ epoch (int): Epoch number.
+ """
+ self.epoch = epoch
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/multi_source_sampler.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/multi_source_sampler.py
new file mode 100644
index 0000000000000000000000000000000000000000..6efcde35e1375547239825a8f78a9e74f7825290
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/multi_source_sampler.py
@@ -0,0 +1,214 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import itertools
+from typing import Iterator, List, Optional, Sized, Union
+
+import numpy as np
+import torch
+from mmengine.dataset import BaseDataset
+from mmengine.dist import get_dist_info, sync_random_seed
+from torch.utils.data import Sampler
+
+from mmdet.registry import DATA_SAMPLERS
+
+
+@DATA_SAMPLERS.register_module()
+class MultiSourceSampler(Sampler):
+ r"""Multi-Source Infinite Sampler.
+
+ According to the sampling ratio, sample data from different
+ datasets to form batches.
+
+ Args:
+ dataset (Sized): The dataset.
+ batch_size (int): Size of mini-batch.
+ source_ratio (list[int | float]): The sampling ratio of different
+ source datasets in a mini-batch.
+ shuffle (bool): Whether shuffle the dataset or not. Defaults to True.
+ seed (int, optional): Random seed. If None, set a random seed.
+ Defaults to None.
+
+ Examples:
+ >>> dataset_type = 'ConcatDataset'
+ >>> sub_dataset_type = 'CocoDataset'
+ >>> data_root = 'data/coco/'
+ >>> sup_ann = '../coco_semi_annos/instances_train2017.1@10.json'
+ >>> unsup_ann = '../coco_semi_annos/' \
+ >>> 'instances_train2017.1@10-unlabeled.json'
+ >>> dataset = dict(type=dataset_type,
+ >>> datasets=[
+ >>> dict(
+ >>> type=sub_dataset_type,
+ >>> data_root=data_root,
+ >>> ann_file=sup_ann,
+ >>> data_prefix=dict(img='train2017/'),
+ >>> filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ >>> pipeline=sup_pipeline),
+ >>> dict(
+ >>> type=sub_dataset_type,
+ >>> data_root=data_root,
+ >>> ann_file=unsup_ann,
+ >>> data_prefix=dict(img='train2017/'),
+ >>> filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ >>> pipeline=unsup_pipeline),
+ >>> ])
+ >>> train_dataloader = dict(
+ >>> batch_size=5,
+ >>> num_workers=5,
+ >>> persistent_workers=True,
+ >>> sampler=dict(type='MultiSourceSampler',
+ >>> batch_size=5, source_ratio=[1, 4]),
+ >>> batch_sampler=None,
+ >>> dataset=dataset)
+ """
+
+ def __init__(self,
+ dataset: Sized,
+ batch_size: int,
+ source_ratio: List[Union[int, float]],
+ shuffle: bool = True,
+ seed: Optional[int] = None) -> None:
+
+ assert hasattr(dataset, 'cumulative_sizes'),\
+ f'The dataset must be ConcatDataset, but get {dataset}'
+ assert isinstance(batch_size, int) and batch_size > 0, \
+ 'batch_size must be a positive integer value, ' \
+ f'but got batch_size={batch_size}'
+ assert isinstance(source_ratio, list), \
+ f'source_ratio must be a list, but got source_ratio={source_ratio}'
+ assert len(source_ratio) == len(dataset.cumulative_sizes), \
+ 'The length of source_ratio must be equal to ' \
+ f'the number of datasets, but got source_ratio={source_ratio}'
+
+ rank, world_size = get_dist_info()
+ self.rank = rank
+ self.world_size = world_size
+
+ self.dataset = dataset
+ self.cumulative_sizes = [0] + dataset.cumulative_sizes
+ self.batch_size = batch_size
+ self.source_ratio = source_ratio
+
+ self.num_per_source = [
+ int(batch_size * sr / sum(source_ratio)) for sr in source_ratio
+ ]
+ self.num_per_source[0] = batch_size - sum(self.num_per_source[1:])
+
+ assert sum(self.num_per_source) == batch_size, \
+ 'The sum of num_per_source must be equal to ' \
+ f'batch_size, but get {self.num_per_source}'
+
+ self.seed = sync_random_seed() if seed is None else seed
+ self.shuffle = shuffle
+ self.source2inds = {
+ source: self._indices_of_rank(len(ds))
+ for source, ds in enumerate(dataset.datasets)
+ }
+
+ def _infinite_indices(self, sample_size: int) -> Iterator[int]:
+ """Infinitely yield a sequence of indices."""
+ g = torch.Generator()
+ g.manual_seed(self.seed)
+ while True:
+ if self.shuffle:
+ yield from torch.randperm(sample_size, generator=g).tolist()
+ else:
+ yield from torch.arange(sample_size).tolist()
+
+ def _indices_of_rank(self, sample_size: int) -> Iterator[int]:
+ """Slice the infinite indices by rank."""
+ yield from itertools.islice(
+ self._infinite_indices(sample_size), self.rank, None,
+ self.world_size)
+
+ def __iter__(self) -> Iterator[int]:
+ batch_buffer = []
+ while True:
+ for source, num in enumerate(self.num_per_source):
+ batch_buffer_per_source = []
+ for idx in self.source2inds[source]:
+ idx += self.cumulative_sizes[source]
+ batch_buffer_per_source.append(idx)
+ if len(batch_buffer_per_source) == num:
+ batch_buffer += batch_buffer_per_source
+ break
+ yield from batch_buffer
+ batch_buffer = []
+
+ def __len__(self) -> int:
+ return len(self.dataset)
+
+ def set_epoch(self, epoch: int) -> None:
+ """Not supported in `epoch-based runner."""
+ pass
+
+
+@DATA_SAMPLERS.register_module()
+class GroupMultiSourceSampler(MultiSourceSampler):
+ r"""Group Multi-Source Infinite Sampler.
+
+ According to the sampling ratio, sample data from different
+ datasets but the same group to form batches.
+
+ Args:
+ dataset (Sized): The dataset.
+ batch_size (int): Size of mini-batch.
+ source_ratio (list[int | float]): The sampling ratio of different
+ source datasets in a mini-batch.
+ shuffle (bool): Whether shuffle the dataset or not. Defaults to True.
+ seed (int, optional): Random seed. If None, set a random seed.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ dataset: BaseDataset,
+ batch_size: int,
+ source_ratio: List[Union[int, float]],
+ shuffle: bool = True,
+ seed: Optional[int] = None) -> None:
+ super().__init__(
+ dataset=dataset,
+ batch_size=batch_size,
+ source_ratio=source_ratio,
+ shuffle=shuffle,
+ seed=seed)
+
+ self._get_source_group_info()
+ self.group_source2inds = [{
+ source:
+ self._indices_of_rank(self.group2size_per_source[source][group])
+ for source in range(len(dataset.datasets))
+ } for group in range(len(self.group_ratio))]
+
+ def _get_source_group_info(self) -> None:
+ self.group2size_per_source = [{0: 0, 1: 0}, {0: 0, 1: 0}]
+ self.group2inds_per_source = [{0: [], 1: []}, {0: [], 1: []}]
+ for source, dataset in enumerate(self.dataset.datasets):
+ for idx in range(len(dataset)):
+ data_info = dataset.get_data_info(idx)
+ width, height = data_info['width'], data_info['height']
+ group = 0 if width < height else 1
+ self.group2size_per_source[source][group] += 1
+ self.group2inds_per_source[source][group].append(idx)
+
+ self.group_sizes = np.zeros(2, dtype=np.int64)
+ for group2size in self.group2size_per_source:
+ for group, size in group2size.items():
+ self.group_sizes[group] += size
+ self.group_ratio = self.group_sizes / sum(self.group_sizes)
+
+ def __iter__(self) -> Iterator[int]:
+ batch_buffer = []
+ while True:
+ group = np.random.choice(
+ list(range(len(self.group_ratio))), p=self.group_ratio)
+ for source, num in enumerate(self.num_per_source):
+ batch_buffer_per_source = []
+ for idx in self.group_source2inds[group][source]:
+ idx = self.group2inds_per_source[source][group][
+ idx] + self.cumulative_sizes[source]
+ batch_buffer_per_source.append(idx)
+ if len(batch_buffer_per_source) == num:
+ batch_buffer += batch_buffer_per_source
+ break
+ yield from batch_buffer
+ batch_buffer = []
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/track_img_sampler.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/track_img_sampler.py
new file mode 100644
index 0000000000000000000000000000000000000000..d7db629f40f3f24bdf14cd852ccc4472d1d50f1b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/samplers/track_img_sampler.py
@@ -0,0 +1,146 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+import random
+from typing import Iterator, Optional, Sized
+
+import numpy as np
+from mmengine.dataset import ClassBalancedDataset, ConcatDataset
+from mmengine.dist import get_dist_info, sync_random_seed
+from torch.utils.data import Sampler
+
+from mmdet.registry import DATA_SAMPLERS
+from ..base_video_dataset import BaseVideoDataset
+
+
+@DATA_SAMPLERS.register_module()
+class TrackImgSampler(Sampler):
+ """Sampler that providing image-level sampling outputs for video datasets
+ in tracking tasks. It could be both used in both distributed and
+ non-distributed environment.
+ If using the default sampler in pytorch, the subsequent data receiver will
+ get one video, which is not desired in some cases:
+ (Take a non-distributed environment as an example)
+ 1. In test mode, we want only one image is fed into the data pipeline. This
+ is in consideration of memory usage since feeding the whole video commonly
+ requires a large amount of memory (>=20G on MOTChallenge17 dataset), which
+ is not available in some machines.
+ 2. In training mode, we may want to make sure all the images in one video
+ are randomly sampled once in one epoch and this can not be guaranteed in
+ the default sampler in pytorch.
+
+ Args:
+ dataset (Sized): Dataset used for sampling.
+ seed (int, optional): random seed used to shuffle the sampler. This
+ number should be identical across all processes in the distributed
+ group. Defaults to None.
+ """
+
+ def __init__(
+ self,
+ dataset: Sized,
+ seed: Optional[int] = None,
+ ) -> None:
+ rank, world_size = get_dist_info()
+ self.rank = rank
+ self.world_size = world_size
+ self.epoch = 0
+ if seed is None:
+ self.seed = sync_random_seed()
+ else:
+ self.seed = seed
+
+ self.dataset = dataset
+ self.indices = []
+ # Hard code here to handle different dataset wrapper
+ if isinstance(self.dataset, ConcatDataset):
+ cat_datasets = self.dataset.datasets
+ assert isinstance(
+ cat_datasets[0], BaseVideoDataset
+ ), f'expected BaseVideoDataset, but got {type(cat_datasets[0])}'
+ self.test_mode = cat_datasets[0].test_mode
+ assert not self.test_mode, "'ConcatDataset' should not exist in "
+ 'test mode'
+ for dataset in cat_datasets:
+ num_videos = len(dataset)
+ for video_ind in range(num_videos):
+ self.indices.extend([
+ (video_ind, frame_ind) for frame_ind in range(
+ dataset.get_len_per_video(video_ind))
+ ])
+ elif isinstance(self.dataset, ClassBalancedDataset):
+ ori_dataset = self.dataset.dataset
+ assert isinstance(
+ ori_dataset, BaseVideoDataset
+ ), f'expected BaseVideoDataset, but got {type(ori_dataset)}'
+ self.test_mode = ori_dataset.test_mode
+ assert not self.test_mode, "'ClassBalancedDataset' should not "
+ 'exist in test mode'
+ video_indices = self.dataset.repeat_indices
+ for index in video_indices:
+ self.indices.extend([(index, frame_ind) for frame_ind in range(
+ ori_dataset.get_len_per_video(index))])
+ else:
+ assert isinstance(
+ self.dataset, BaseVideoDataset
+ ), 'TrackImgSampler is only supported in BaseVideoDataset or '
+ 'dataset wrapper: ClassBalancedDataset and ConcatDataset, but '
+ f'got {type(self.dataset)} '
+ self.test_mode = self.dataset.test_mode
+ num_videos = len(self.dataset)
+
+ if self.test_mode:
+ # in test mode, the images belong to the same video must be put
+ # on the same device.
+ if num_videos < self.world_size:
+ raise ValueError(f'only {num_videos} videos loaded,'
+ f'but {self.world_size} gpus were given.')
+ chunks = np.array_split(
+ list(range(num_videos)), self.world_size)
+ for videos_inds in chunks:
+ indices_chunk = []
+ for video_ind in videos_inds:
+ indices_chunk.extend([
+ (video_ind, frame_ind) for frame_ind in range(
+ self.dataset.get_len_per_video(video_ind))
+ ])
+ self.indices.append(indices_chunk)
+ else:
+ for video_ind in range(num_videos):
+ self.indices.extend([
+ (video_ind, frame_ind) for frame_ind in range(
+ self.dataset.get_len_per_video(video_ind))
+ ])
+
+ if self.test_mode:
+ self.num_samples = len(self.indices[self.rank])
+ self.total_size = sum(
+ [len(index_list) for index_list in self.indices])
+ else:
+ self.num_samples = int(
+ math.ceil(len(self.indices) * 1.0 / self.world_size))
+ self.total_size = self.num_samples * self.world_size
+
+ def __iter__(self) -> Iterator:
+ if self.test_mode:
+ # in test mode, the order of frames can not be shuffled.
+ indices = self.indices[self.rank]
+ else:
+ # deterministically shuffle based on epoch
+ rng = random.Random(self.epoch + self.seed)
+ indices = rng.sample(self.indices, len(self.indices))
+
+ # add extra samples to make it evenly divisible
+ indices += indices[:(self.total_size - len(indices))]
+ assert len(indices) == self.total_size
+
+ # subsample
+ indices = indices[self.rank:self.total_size:self.world_size]
+ assert len(indices) == self.num_samples
+
+ return iter(indices)
+
+ def __len__(self):
+ return self.num_samples
+
+ def set_epoch(self, epoch):
+ self.epoch = epoch
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..ab3478feb008443cb0e56bf5084261370e38327d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/__init__.py
@@ -0,0 +1,45 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .augment_wrappers import AutoAugment, RandAugment
+from .colorspace import (AutoContrast, Brightness, Color, ColorTransform,
+ Contrast, Equalize, Invert, Posterize, Sharpness,
+ Solarize, SolarizeAdd)
+from .formatting import (ImageToTensor, PackDetInputs, PackReIDInputs,
+ PackTrackInputs, ToTensor, Transpose)
+from .frame_sampling import BaseFrameSample, UniformRefFrameSample
+from .geometric import (GeomTransform, Rotate, ShearX, ShearY, TranslateX,
+ TranslateY)
+from .instaboost import InstaBoost
+from .loading import (FilterAnnotations, InferencerLoader, LoadAnnotations,
+ LoadEmptyAnnotations, LoadImageFromNDArray,
+ LoadMultiChannelImageFromFiles, LoadPanopticAnnotations,
+ LoadProposals, LoadTrackAnnotations)
+from .text_transformers import LoadTextAnnotations, RandomSamplingNegPos
+from .transformers_glip import GTBoxSubOne_GLIP, RandomFlip_GLIP
+from .transforms import (Albu, CachedMixUp, CachedMosaic, CopyPaste, CutOut,
+ Expand, FixScaleResize, FixShapeResize,
+ MinIoURandomCrop, MixUp, Mosaic, Pad,
+ PhotoMetricDistortion, RandomAffine,
+ RandomCenterCropPad, RandomCrop, RandomErasing,
+ RandomFlip, RandomShift, Resize, ResizeShortestEdge,
+ SegRescale, YOLOXHSVRandomAug)
+from .wrappers import MultiBranch, ProposalBroadcaster, RandomOrder
+
+__all__ = [
+ 'PackDetInputs', 'ToTensor', 'ImageToTensor', 'Transpose',
+ 'LoadImageFromNDArray', 'LoadAnnotations', 'LoadPanopticAnnotations',
+ 'LoadMultiChannelImageFromFiles', 'LoadProposals', 'Resize', 'RandomFlip',
+ 'RandomCrop', 'SegRescale', 'MinIoURandomCrop', 'Expand',
+ 'PhotoMetricDistortion', 'Albu', 'InstaBoost', 'RandomCenterCropPad',
+ 'AutoAugment', 'CutOut', 'ShearX', 'ShearY', 'Rotate', 'Color', 'Equalize',
+ 'Brightness', 'Contrast', 'TranslateX', 'TranslateY', 'RandomShift',
+ 'Mosaic', 'MixUp', 'RandomAffine', 'YOLOXHSVRandomAug', 'CopyPaste',
+ 'FilterAnnotations', 'Pad', 'GeomTransform', 'ColorTransform',
+ 'RandAugment', 'Sharpness', 'Solarize', 'SolarizeAdd', 'Posterize',
+ 'AutoContrast', 'Invert', 'MultiBranch', 'RandomErasing',
+ 'LoadEmptyAnnotations', 'RandomOrder', 'CachedMosaic', 'CachedMixUp',
+ 'FixShapeResize', 'ProposalBroadcaster', 'InferencerLoader',
+ 'LoadTrackAnnotations', 'BaseFrameSample', 'UniformRefFrameSample',
+ 'PackTrackInputs', 'PackReIDInputs', 'FixScaleResize',
+ 'ResizeShortestEdge', 'GTBoxSubOne_GLIP', 'RandomFlip_GLIP',
+ 'RandomSamplingNegPos', 'LoadTextAnnotations'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/augment_wrappers.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/augment_wrappers.py
new file mode 100644
index 0000000000000000000000000000000000000000..19fae6efdf66aa4c26bb85a2f2c96a1e079320b8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/augment_wrappers.py
@@ -0,0 +1,264 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Union
+
+import numpy as np
+from mmcv.transforms import RandomChoice
+from mmcv.transforms.utils import cache_randomness
+from mmengine.config import ConfigDict
+
+from mmdet.registry import TRANSFORMS
+
+# AutoAugment uses reinforcement learning to search for
+# some widely useful data augmentation strategies,
+# here we provide AUTOAUG_POLICIES_V0.
+# For AUTOAUG_POLICIES_V0, each tuple is an augmentation
+# operation of the form (operation, probability, magnitude).
+# Each element in policies is a policy that will be applied
+# sequentially on the image.
+
+# RandAugment defines a data augmentation search space, RANDAUG_SPACE,
+# sampling 1~3 data augmentations each time, and
+# setting the magnitude of each data augmentation randomly,
+# which will be applied sequentially on the image.
+
+_MAX_LEVEL = 10
+
+AUTOAUG_POLICIES_V0 = [
+ [('Equalize', 0.8, 1), ('ShearY', 0.8, 4)],
+ [('Color', 0.4, 9), ('Equalize', 0.6, 3)],
+ [('Color', 0.4, 1), ('Rotate', 0.6, 8)],
+ [('Solarize', 0.8, 3), ('Equalize', 0.4, 7)],
+ [('Solarize', 0.4, 2), ('Solarize', 0.6, 2)],
+ [('Color', 0.2, 0), ('Equalize', 0.8, 8)],
+ [('Equalize', 0.4, 8), ('SolarizeAdd', 0.8, 3)],
+ [('ShearX', 0.2, 9), ('Rotate', 0.6, 8)],
+ [('Color', 0.6, 1), ('Equalize', 1.0, 2)],
+ [('Invert', 0.4, 9), ('Rotate', 0.6, 0)],
+ [('Equalize', 1.0, 9), ('ShearY', 0.6, 3)],
+ [('Color', 0.4, 7), ('Equalize', 0.6, 0)],
+ [('Posterize', 0.4, 6), ('AutoContrast', 0.4, 7)],
+ [('Solarize', 0.6, 8), ('Color', 0.6, 9)],
+ [('Solarize', 0.2, 4), ('Rotate', 0.8, 9)],
+ [('Rotate', 1.0, 7), ('TranslateY', 0.8, 9)],
+ [('ShearX', 0.0, 0), ('Solarize', 0.8, 4)],
+ [('ShearY', 0.8, 0), ('Color', 0.6, 4)],
+ [('Color', 1.0, 0), ('Rotate', 0.6, 2)],
+ [('Equalize', 0.8, 4), ('Equalize', 0.0, 8)],
+ [('Equalize', 1.0, 4), ('AutoContrast', 0.6, 2)],
+ [('ShearY', 0.4, 7), ('SolarizeAdd', 0.6, 7)],
+ [('Posterize', 0.8, 2), ('Solarize', 0.6, 10)],
+ [('Solarize', 0.6, 8), ('Equalize', 0.6, 1)],
+ [('Color', 0.8, 6), ('Rotate', 0.4, 5)],
+]
+
+
+def policies_v0():
+ """Autoaugment policies that was used in AutoAugment Paper."""
+ policies = list()
+ for policy_args in AUTOAUG_POLICIES_V0:
+ policy = list()
+ for args in policy_args:
+ policy.append(dict(type=args[0], prob=args[1], level=args[2]))
+ policies.append(policy)
+ return policies
+
+
+RANDAUG_SPACE = [[dict(type='AutoContrast')], [dict(type='Equalize')],
+ [dict(type='Invert')], [dict(type='Rotate')],
+ [dict(type='Posterize')], [dict(type='Solarize')],
+ [dict(type='SolarizeAdd')], [dict(type='Color')],
+ [dict(type='Contrast')], [dict(type='Brightness')],
+ [dict(type='Sharpness')], [dict(type='ShearX')],
+ [dict(type='ShearY')], [dict(type='TranslateX')],
+ [dict(type='TranslateY')]]
+
+
+def level_to_mag(level: Optional[int], min_mag: float,
+ max_mag: float) -> float:
+ """Map from level to magnitude."""
+ if level is None:
+ return round(np.random.rand() * (max_mag - min_mag) + min_mag, 1)
+ else:
+ return round(level / _MAX_LEVEL * (max_mag - min_mag) + min_mag, 1)
+
+
+@TRANSFORMS.register_module()
+class AutoAugment(RandomChoice):
+ """Auto augmentation.
+
+ This data augmentation is proposed in `AutoAugment: Learning
+ Augmentation Policies from Data `_
+ and in `Learning Data Augmentation Strategies for Object Detection
+ `_.
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_bboxes_labels (np.int64) (optional)
+ - gt_masks (BitmapMasks | PolygonMasks) (optional)
+ - gt_ignore_flags (bool) (optional)
+ - gt_seg_map (np.uint8) (optional)
+
+ Modified Keys:
+
+ - img
+ - img_shape
+ - gt_bboxes
+ - gt_bboxes_labels
+ - gt_masks
+ - gt_ignore_flags
+ - gt_seg_map
+
+ Added Keys:
+
+ - homography_matrix
+
+ Args:
+ policies (List[List[Union[dict, ConfigDict]]]):
+ The policies of auto augmentation.Each policy in ``policies``
+ is a specific augmentation policy, and is composed by several
+ augmentations. When AutoAugment is called, a random policy in
+ ``policies`` will be selected to augment images.
+ Defaults to policy_v0().
+ prob (list[float], optional): The probabilities associated
+ with each policy. The length should be equal to the policy
+ number and the sum should be 1. If not given, a uniform
+ distribution will be assumed. Defaults to None.
+
+ Examples:
+ >>> policies = [
+ >>> [
+ >>> dict(type='Sharpness', prob=0.0, level=8),
+ >>> dict(type='ShearX', prob=0.4, level=0,)
+ >>> ],
+ >>> [
+ >>> dict(type='Rotate', prob=0.6, level=10),
+ >>> dict(type='Color', prob=1.0, level=6)
+ >>> ]
+ >>> ]
+ >>> augmentation = AutoAugment(policies)
+ >>> img = np.ones(100, 100, 3)
+ >>> gt_bboxes = np.ones(10, 4)
+ >>> results = dict(img=img, gt_bboxes=gt_bboxes)
+ >>> results = augmentation(results)
+ """
+
+ def __init__(self,
+ policies: List[List[Union[dict, ConfigDict]]] = policies_v0(),
+ prob: Optional[List[float]] = None) -> None:
+ assert isinstance(policies, list) and len(policies) > 0, \
+ 'Policies must be a non-empty list.'
+ for policy in policies:
+ assert isinstance(policy, list) and len(policy) > 0, \
+ 'Each policy in policies must be a non-empty list.'
+ for augment in policy:
+ assert isinstance(augment, dict) and 'type' in augment, \
+ 'Each specific augmentation must be a dict with key' \
+ ' "type".'
+ super().__init__(transforms=policies, prob=prob)
+ self.policies = policies
+
+ def __repr__(self) -> str:
+ return f'{self.__class__.__name__}(policies={self.policies}, ' \
+ f'prob={self.prob})'
+
+
+@TRANSFORMS.register_module()
+class RandAugment(RandomChoice):
+ """Rand augmentation.
+
+ This data augmentation is proposed in `RandAugment:
+ Practical automated data augmentation with a reduced
+ search space `_.
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_bboxes_labels (np.int64) (optional)
+ - gt_masks (BitmapMasks | PolygonMasks) (optional)
+ - gt_ignore_flags (bool) (optional)
+ - gt_seg_map (np.uint8) (optional)
+
+ Modified Keys:
+
+ - img
+ - img_shape
+ - gt_bboxes
+ - gt_bboxes_labels
+ - gt_masks
+ - gt_ignore_flags
+ - gt_seg_map
+
+ Added Keys:
+
+ - homography_matrix
+
+ Args:
+ aug_space (List[List[Union[dict, ConfigDict]]]): The augmentation space
+ of rand augmentation. Each augmentation transform in ``aug_space``
+ is a specific transform, and is composed by several augmentations.
+ When RandAugment is called, a random transform in ``aug_space``
+ will be selected to augment images. Defaults to aug_space.
+ aug_num (int): Number of augmentation to apply equentially.
+ Defaults to 2.
+ prob (list[float], optional): The probabilities associated with
+ each augmentation. The length should be equal to the
+ augmentation space and the sum should be 1. If not given,
+ a uniform distribution will be assumed. Defaults to None.
+
+ Examples:
+ >>> aug_space = [
+ >>> dict(type='Sharpness'),
+ >>> dict(type='ShearX'),
+ >>> dict(type='Color'),
+ >>> ],
+ >>> augmentation = RandAugment(aug_space)
+ >>> img = np.ones(100, 100, 3)
+ >>> gt_bboxes = np.ones(10, 4)
+ >>> results = dict(img=img, gt_bboxes=gt_bboxes)
+ >>> results = augmentation(results)
+ """
+
+ def __init__(self,
+ aug_space: List[Union[dict, ConfigDict]] = RANDAUG_SPACE,
+ aug_num: int = 2,
+ prob: Optional[List[float]] = None) -> None:
+ assert isinstance(aug_space, list) and len(aug_space) > 0, \
+ 'Augmentation space must be a non-empty list.'
+ for aug in aug_space:
+ assert isinstance(aug, list) and len(aug) == 1, \
+ 'Each augmentation in aug_space must be a list.'
+ for transform in aug:
+ assert isinstance(transform, dict) and 'type' in transform, \
+ 'Each specific transform must be a dict with key' \
+ ' "type".'
+ super().__init__(transforms=aug_space, prob=prob)
+ self.aug_space = aug_space
+ self.aug_num = aug_num
+
+ @cache_randomness
+ def random_pipeline_index(self):
+ indices = np.arange(len(self.transforms))
+ return np.random.choice(
+ indices, self.aug_num, p=self.prob, replace=False)
+
+ def transform(self, results: dict) -> dict:
+ """Transform function to use RandAugment.
+
+ Args:
+ results (dict): Result dict from loading pipeline.
+
+ Returns:
+ dict: Result dict with RandAugment.
+ """
+ for idx in self.random_pipeline_index():
+ results = self.transforms[idx](results)
+ return results
+
+ def __repr__(self) -> str:
+ return f'{self.__class__.__name__}(' \
+ f'aug_space={self.aug_space}, '\
+ f'aug_num={self.aug_num}, ' \
+ f'prob={self.prob})'
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/colorspace.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/colorspace.py
new file mode 100644
index 0000000000000000000000000000000000000000..e0ba2e97c7eedf65df5ab8942ee461f48a785f39
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/colorspace.py
@@ -0,0 +1,493 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+from typing import Optional
+
+import mmcv
+import numpy as np
+from mmcv.transforms import BaseTransform
+from mmcv.transforms.utils import cache_randomness
+
+from mmdet.registry import TRANSFORMS
+from .augment_wrappers import _MAX_LEVEL, level_to_mag
+
+
+@TRANSFORMS.register_module()
+class ColorTransform(BaseTransform):
+ """Base class for color transformations. All color transformations need to
+ inherit from this base class. ``ColorTransform`` unifies the class
+ attributes and class functions of color transformations (Color, Brightness,
+ Contrast, Sharpness, Solarize, SolarizeAdd, Equalize, AutoContrast, Invert,
+ and Posterize), and only distort color channels, without impacting the
+ locations of the instances.
+
+ Required Keys:
+
+ - img
+
+ Modified Keys:
+
+ - img
+
+ Args:
+ prob (float): The probability for performing the geometric
+ transformation and should be in range [0, 1]. Defaults to 1.0.
+ level (int, optional): The level should be in range [0, _MAX_LEVEL].
+ If level is None, it will generate from [0, _MAX_LEVEL] randomly.
+ Defaults to None.
+ min_mag (float): The minimum magnitude for color transformation.
+ Defaults to 0.1.
+ max_mag (float): The maximum magnitude for color transformation.
+ Defaults to 1.9.
+ """
+
+ def __init__(self,
+ prob: float = 1.0,
+ level: Optional[int] = None,
+ min_mag: float = 0.1,
+ max_mag: float = 1.9) -> None:
+ assert 0 <= prob <= 1.0, f'The probability of the transformation ' \
+ f'should be in range [0,1], got {prob}.'
+ assert level is None or isinstance(level, int), \
+ f'The level should be None or type int, got {type(level)}.'
+ assert level is None or 0 <= level <= _MAX_LEVEL, \
+ f'The level should be in range [0,{_MAX_LEVEL}], got {level}.'
+ assert isinstance(min_mag, float), \
+ f'min_mag should be type float, got {type(min_mag)}.'
+ assert isinstance(max_mag, float), \
+ f'max_mag should be type float, got {type(max_mag)}.'
+ assert min_mag <= max_mag, \
+ f'min_mag should smaller than max_mag, ' \
+ f'got min_mag={min_mag} and max_mag={max_mag}'
+ self.prob = prob
+ self.level = level
+ self.min_mag = min_mag
+ self.max_mag = max_mag
+
+ def _transform_img(self, results: dict, mag: float) -> None:
+ """Transform the image."""
+ pass
+
+ @cache_randomness
+ def _random_disable(self):
+ """Randomly disable the transform."""
+ return np.random.rand() > self.prob
+
+ @cache_randomness
+ def _get_mag(self):
+ """Get the magnitude of the transform."""
+ return level_to_mag(self.level, self.min_mag, self.max_mag)
+
+ def transform(self, results: dict) -> dict:
+ """Transform function for images.
+
+ Args:
+ results (dict): Result dict from loading pipeline.
+
+ Returns:
+ dict: Transformed results.
+ """
+
+ if self._random_disable():
+ return results
+ mag = self._get_mag()
+ self._transform_img(results, mag)
+ return results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'(prob={self.prob}, '
+ repr_str += f'level={self.level}, '
+ repr_str += f'min_mag={self.min_mag}, '
+ repr_str += f'max_mag={self.max_mag})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class Color(ColorTransform):
+ """Adjust the color balance of the image, in a manner similar to the
+ controls on a colour TV set. A magnitude=0 gives a black & white image,
+ whereas magnitude=1 gives the original image. The bboxes, masks and
+ segmentations are not modified.
+
+ Required Keys:
+
+ - img
+
+ Modified Keys:
+
+ - img
+
+ Args:
+ prob (float): The probability for performing Color transformation.
+ Defaults to 1.0.
+ level (int, optional): Should be in range [0,_MAX_LEVEL].
+ If level is None, it will generate from [0, _MAX_LEVEL] randomly.
+ Defaults to None.
+ min_mag (float): The minimum magnitude for Color transformation.
+ Defaults to 0.1.
+ max_mag (float): The maximum magnitude for Color transformation.
+ Defaults to 1.9.
+ """
+
+ def __init__(self,
+ prob: float = 1.0,
+ level: Optional[int] = None,
+ min_mag: float = 0.1,
+ max_mag: float = 1.9) -> None:
+ assert 0. <= min_mag <= 2.0, \
+ f'min_mag for Color should be in range [0,2], got {min_mag}.'
+ assert 0. <= max_mag <= 2.0, \
+ f'max_mag for Color should be in range [0,2], got {max_mag}.'
+ super().__init__(
+ prob=prob, level=level, min_mag=min_mag, max_mag=max_mag)
+
+ def _transform_img(self, results: dict, mag: float) -> None:
+ """Apply Color transformation to image."""
+ # NOTE defaultly the image should be BGR format
+ img = results['img']
+ results['img'] = mmcv.adjust_color(img, mag).astype(img.dtype)
+
+
+@TRANSFORMS.register_module()
+class Brightness(ColorTransform):
+ """Adjust the brightness of the image. A magnitude=0 gives a black image,
+ whereas magnitude=1 gives the original image. The bboxes, masks and
+ segmentations are not modified.
+
+ Required Keys:
+
+ - img
+
+ Modified Keys:
+
+ - img
+
+ Args:
+ prob (float): The probability for performing Brightness transformation.
+ Defaults to 1.0.
+ level (int, optional): Should be in range [0,_MAX_LEVEL].
+ If level is None, it will generate from [0, _MAX_LEVEL] randomly.
+ Defaults to None.
+ min_mag (float): The minimum magnitude for Brightness transformation.
+ Defaults to 0.1.
+ max_mag (float): The maximum magnitude for Brightness transformation.
+ Defaults to 1.9.
+ """
+
+ def __init__(self,
+ prob: float = 1.0,
+ level: Optional[int] = None,
+ min_mag: float = 0.1,
+ max_mag: float = 1.9) -> None:
+ assert 0. <= min_mag <= 2.0, \
+ f'min_mag for Brightness should be in range [0,2], got {min_mag}.'
+ assert 0. <= max_mag <= 2.0, \
+ f'max_mag for Brightness should be in range [0,2], got {max_mag}.'
+ super().__init__(
+ prob=prob, level=level, min_mag=min_mag, max_mag=max_mag)
+
+ def _transform_img(self, results: dict, mag: float) -> None:
+ """Adjust the brightness of image."""
+ img = results['img']
+ results['img'] = mmcv.adjust_brightness(img, mag).astype(img.dtype)
+
+
+@TRANSFORMS.register_module()
+class Contrast(ColorTransform):
+ """Control the contrast of the image. A magnitude=0 gives a gray image,
+ whereas magnitude=1 gives the original imageThe bboxes, masks and
+ segmentations are not modified.
+
+ Required Keys:
+
+ - img
+
+ Modified Keys:
+
+ - img
+
+ Args:
+ prob (float): The probability for performing Contrast transformation.
+ Defaults to 1.0.
+ level (int, optional): Should be in range [0,_MAX_LEVEL].
+ If level is None, it will generate from [0, _MAX_LEVEL] randomly.
+ Defaults to None.
+ min_mag (float): The minimum magnitude for Contrast transformation.
+ Defaults to 0.1.
+ max_mag (float): The maximum magnitude for Contrast transformation.
+ Defaults to 1.9.
+ """
+
+ def __init__(self,
+ prob: float = 1.0,
+ level: Optional[int] = None,
+ min_mag: float = 0.1,
+ max_mag: float = 1.9) -> None:
+ assert 0. <= min_mag <= 2.0, \
+ f'min_mag for Contrast should be in range [0,2], got {min_mag}.'
+ assert 0. <= max_mag <= 2.0, \
+ f'max_mag for Contrast should be in range [0,2], got {max_mag}.'
+ super().__init__(
+ prob=prob, level=level, min_mag=min_mag, max_mag=max_mag)
+
+ def _transform_img(self, results: dict, mag: float) -> None:
+ """Adjust the image contrast."""
+ img = results['img']
+ results['img'] = mmcv.adjust_contrast(img, mag).astype(img.dtype)
+
+
+@TRANSFORMS.register_module()
+class Sharpness(ColorTransform):
+ """Adjust images sharpness. A positive magnitude would enhance the
+ sharpness and a negative magnitude would make the image blurry. A
+ magnitude=0 gives the origin img.
+
+ Required Keys:
+
+ - img
+
+ Modified Keys:
+
+ - img
+
+ Args:
+ prob (float): The probability for performing Sharpness transformation.
+ Defaults to 1.0.
+ level (int, optional): Should be in range [0,_MAX_LEVEL].
+ If level is None, it will generate from [0, _MAX_LEVEL] randomly.
+ Defaults to None.
+ min_mag (float): The minimum magnitude for Sharpness transformation.
+ Defaults to 0.1.
+ max_mag (float): The maximum magnitude for Sharpness transformation.
+ Defaults to 1.9.
+ """
+
+ def __init__(self,
+ prob: float = 1.0,
+ level: Optional[int] = None,
+ min_mag: float = 0.1,
+ max_mag: float = 1.9) -> None:
+ assert 0. <= min_mag <= 2.0, \
+ f'min_mag for Sharpness should be in range [0,2], got {min_mag}.'
+ assert 0. <= max_mag <= 2.0, \
+ f'max_mag for Sharpness should be in range [0,2], got {max_mag}.'
+ super().__init__(
+ prob=prob, level=level, min_mag=min_mag, max_mag=max_mag)
+
+ def _transform_img(self, results: dict, mag: float) -> None:
+ """Adjust the image sharpness."""
+ img = results['img']
+ results['img'] = mmcv.adjust_sharpness(img, mag).astype(img.dtype)
+
+
+@TRANSFORMS.register_module()
+class Solarize(ColorTransform):
+ """Solarize images (Invert all pixels above a threshold value of
+ magnitude.).
+
+ Required Keys:
+
+ - img
+
+ Modified Keys:
+
+ - img
+
+ Args:
+ prob (float): The probability for performing Solarize transformation.
+ Defaults to 1.0.
+ level (int, optional): Should be in range [0,_MAX_LEVEL].
+ If level is None, it will generate from [0, _MAX_LEVEL] randomly.
+ Defaults to None.
+ min_mag (float): The minimum magnitude for Solarize transformation.
+ Defaults to 0.0.
+ max_mag (float): The maximum magnitude for Solarize transformation.
+ Defaults to 256.0.
+ """
+
+ def __init__(self,
+ prob: float = 1.0,
+ level: Optional[int] = None,
+ min_mag: float = 0.0,
+ max_mag: float = 256.0) -> None:
+ assert 0. <= min_mag <= 256.0, f'min_mag for Solarize should be ' \
+ f'in range [0, 256], got {min_mag}.'
+ assert 0. <= max_mag <= 256.0, f'max_mag for Solarize should be ' \
+ f'in range [0, 256], got {max_mag}.'
+ super().__init__(
+ prob=prob, level=level, min_mag=min_mag, max_mag=max_mag)
+
+ def _transform_img(self, results: dict, mag: float) -> None:
+ """Invert all pixel values above magnitude."""
+ img = results['img']
+ results['img'] = mmcv.solarize(img, mag).astype(img.dtype)
+
+
+@TRANSFORMS.register_module()
+class SolarizeAdd(ColorTransform):
+ """SolarizeAdd images. For each pixel in the image that is less than 128,
+ add an additional amount to it decided by the magnitude.
+
+ Required Keys:
+
+ - img
+
+ Modified Keys:
+
+ - img
+
+ Args:
+ prob (float): The probability for performing SolarizeAdd
+ transformation. Defaults to 1.0.
+ level (int, optional): Should be in range [0,_MAX_LEVEL].
+ If level is None, it will generate from [0, _MAX_LEVEL] randomly.
+ Defaults to None.
+ min_mag (float): The minimum magnitude for SolarizeAdd transformation.
+ Defaults to 0.0.
+ max_mag (float): The maximum magnitude for SolarizeAdd transformation.
+ Defaults to 110.0.
+ """
+
+ def __init__(self,
+ prob: float = 1.0,
+ level: Optional[int] = None,
+ min_mag: float = 0.0,
+ max_mag: float = 110.0) -> None:
+ assert 0. <= min_mag <= 110.0, f'min_mag for SolarizeAdd should be ' \
+ f'in range [0, 110], got {min_mag}.'
+ assert 0. <= max_mag <= 110.0, f'max_mag for SolarizeAdd should be ' \
+ f'in range [0, 110], got {max_mag}.'
+ super().__init__(
+ prob=prob, level=level, min_mag=min_mag, max_mag=max_mag)
+
+ def _transform_img(self, results: dict, mag: float) -> None:
+ """SolarizeAdd the image."""
+ img = results['img']
+ img_solarized = np.where(img < 128, np.minimum(img + mag, 255), img)
+ results['img'] = img_solarized.astype(img.dtype)
+
+
+@TRANSFORMS.register_module()
+class Posterize(ColorTransform):
+ """Posterize images (reduce the number of bits for each color channel).
+
+ Required Keys:
+
+ - img
+
+ Modified Keys:
+
+ - img
+
+ Args:
+ prob (float): The probability for performing Posterize
+ transformation. Defaults to 1.0.
+ level (int, optional): Should be in range [0,_MAX_LEVEL].
+ If level is None, it will generate from [0, _MAX_LEVEL] randomly.
+ Defaults to None.
+ min_mag (float): The minimum magnitude for Posterize transformation.
+ Defaults to 0.0.
+ max_mag (float): The maximum magnitude for Posterize transformation.
+ Defaults to 4.0.
+ """
+
+ def __init__(self,
+ prob: float = 1.0,
+ level: Optional[int] = None,
+ min_mag: float = 0.0,
+ max_mag: float = 4.0) -> None:
+ assert 0. <= min_mag <= 8.0, f'min_mag for Posterize should be ' \
+ f'in range [0, 8], got {min_mag}.'
+ assert 0. <= max_mag <= 8.0, f'max_mag for Posterize should be ' \
+ f'in range [0, 8], got {max_mag}.'
+ super().__init__(
+ prob=prob, level=level, min_mag=min_mag, max_mag=max_mag)
+
+ def _transform_img(self, results: dict, mag: float) -> None:
+ """Posterize the image."""
+ img = results['img']
+ results['img'] = mmcv.posterize(img, math.ceil(mag)).astype(img.dtype)
+
+
+@TRANSFORMS.register_module()
+class Equalize(ColorTransform):
+ """Equalize the image histogram. The bboxes, masks and segmentations are
+ not modified.
+
+ Required Keys:
+
+ - img
+
+ Modified Keys:
+
+ - img
+
+ Args:
+ prob (float): The probability for performing Equalize transformation.
+ Defaults to 1.0.
+ level (int, optional): No use for Equalize transformation.
+ Defaults to None.
+ min_mag (float): No use for Equalize transformation. Defaults to 0.1.
+ max_mag (float): No use for Equalize transformation. Defaults to 1.9.
+ """
+
+ def _transform_img(self, results: dict, mag: float) -> None:
+ """Equalizes the histogram of one image."""
+ img = results['img']
+ results['img'] = mmcv.imequalize(img).astype(img.dtype)
+
+
+@TRANSFORMS.register_module()
+class AutoContrast(ColorTransform):
+ """Auto adjust image contrast.
+
+ Required Keys:
+
+ - img
+
+ Modified Keys:
+
+ - img
+
+ Args:
+ prob (float): The probability for performing AutoContrast should
+ be in range [0, 1]. Defaults to 1.0.
+ level (int, optional): No use for AutoContrast transformation.
+ Defaults to None.
+ min_mag (float): No use for AutoContrast transformation.
+ Defaults to 0.1.
+ max_mag (float): No use for AutoContrast transformation.
+ Defaults to 1.9.
+ """
+
+ def _transform_img(self, results: dict, mag: float) -> None:
+ """Auto adjust image contrast."""
+ img = results['img']
+ results['img'] = mmcv.auto_contrast(img).astype(img.dtype)
+
+
+@TRANSFORMS.register_module()
+class Invert(ColorTransform):
+ """Invert images.
+
+ Required Keys:
+
+ - img
+
+ Modified Keys:
+
+ - img
+
+ Args:
+ prob (float): The probability for performing invert therefore should
+ be in range [0, 1]. Defaults to 1.0.
+ level (int, optional): No use for Invert transformation.
+ Defaults to None.
+ min_mag (float): No use for Invert transformation. Defaults to 0.1.
+ max_mag (float): No use for Invert transformation. Defaults to 1.9.
+ """
+
+ def _transform_img(self, results: dict, mag: float) -> None:
+ """Invert the image."""
+ img = results['img']
+ results['img'] = mmcv.iminvert(img).astype(img.dtype)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/formatting.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/formatting.py
new file mode 100644
index 0000000000000000000000000000000000000000..05263807c0eab470b0c73f435d327ad8cadb60b3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/formatting.py
@@ -0,0 +1,512 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Sequence
+
+import numpy as np
+from mmcv.transforms import to_tensor
+from mmcv.transforms.base import BaseTransform
+from mmengine.structures import InstanceData, PixelData
+
+from mmdet.registry import TRANSFORMS
+from mmdet.structures import DetDataSample, ReIDDataSample, TrackDataSample
+from mmdet.structures.bbox import BaseBoxes
+
+
+@TRANSFORMS.register_module()
+class PackDetInputs(BaseTransform):
+ """Pack the inputs data for the detection / semantic segmentation /
+ panoptic segmentation.
+
+ The ``img_meta`` item is always populated. The contents of the
+ ``img_meta`` dictionary depends on ``meta_keys``. By default this includes:
+
+ - ``img_id``: id of the image
+
+ - ``img_path``: path to the image file
+
+ - ``ori_shape``: original shape of the image as a tuple (h, w)
+
+ - ``img_shape``: shape of the image input to the network as a tuple \
+ (h, w). Note that images may be zero padded on the \
+ bottom/right if the batch tensor is larger than this shape.
+
+ - ``scale_factor``: a float indicating the preprocessing scale
+
+ - ``flip``: a boolean indicating if image flip transform was used
+
+ - ``flip_direction``: the flipping direction
+
+ Args:
+ meta_keys (Sequence[str], optional): Meta keys to be converted to
+ ``mmcv.DataContainer`` and collected in ``data[img_metas]``.
+ Default: ``('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction')``
+ """
+ mapping_table = {
+ 'gt_bboxes': 'bboxes',
+ 'gt_bboxes_labels': 'labels',
+ 'gt_masks': 'masks'
+ }
+
+ def __init__(self,
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'flip', 'flip_direction')):
+ self.meta_keys = meta_keys
+
+ def transform(self, results: dict) -> dict:
+ """Method to pack the input data.
+
+ Args:
+ results (dict): Result dict from the data pipeline.
+
+ Returns:
+ dict:
+
+ - 'inputs' (obj:`torch.Tensor`): The forward data of models.
+ - 'data_sample' (obj:`DetDataSample`): The annotation info of the
+ sample.
+ """
+ packed_results = dict()
+ if 'img' in results:
+ img = results['img']
+ if len(img.shape) < 3:
+ img = np.expand_dims(img, -1)
+ # To improve the computational speed by by 3-5 times, apply:
+ # If image is not contiguous, use
+ # `numpy.transpose()` followed by `numpy.ascontiguousarray()`
+ # If image is already contiguous, use
+ # `torch.permute()` followed by `torch.contiguous()`
+ # Refer to https://github.com/open-mmlab/mmdetection/pull/9533
+ # for more details
+ if not img.flags.c_contiguous:
+ img = np.ascontiguousarray(img.transpose(2, 0, 1))
+ img = to_tensor(img)
+ else:
+ img = to_tensor(img).permute(2, 0, 1).contiguous()
+
+ packed_results['inputs'] = img
+
+ if 'gt_ignore_flags' in results:
+ valid_idx = np.where(results['gt_ignore_flags'] == 0)[0]
+ ignore_idx = np.where(results['gt_ignore_flags'] == 1)[0]
+
+ data_sample = DetDataSample()
+ instance_data = InstanceData()
+ ignore_instance_data = InstanceData()
+
+ for key in self.mapping_table.keys():
+ if key not in results:
+ continue
+ if key == 'gt_masks' or isinstance(results[key], BaseBoxes):
+ if 'gt_ignore_flags' in results:
+ instance_data[
+ self.mapping_table[key]] = results[key][valid_idx]
+ ignore_instance_data[
+ self.mapping_table[key]] = results[key][ignore_idx]
+ else:
+ instance_data[self.mapping_table[key]] = results[key]
+ else:
+ if 'gt_ignore_flags' in results:
+ instance_data[self.mapping_table[key]] = to_tensor(
+ results[key][valid_idx])
+ ignore_instance_data[self.mapping_table[key]] = to_tensor(
+ results[key][ignore_idx])
+ else:
+ instance_data[self.mapping_table[key]] = to_tensor(
+ results[key])
+ data_sample.gt_instances = instance_data
+ data_sample.ignored_instances = ignore_instance_data
+
+ if 'proposals' in results:
+ proposals = InstanceData(
+ bboxes=to_tensor(results['proposals']),
+ scores=to_tensor(results['proposals_scores']))
+ data_sample.proposals = proposals
+
+ if 'gt_seg_map' in results:
+ gt_sem_seg_data = dict(
+ sem_seg=to_tensor(results['gt_seg_map'][None, ...].copy()))
+ gt_sem_seg_data = PixelData(**gt_sem_seg_data)
+ if 'ignore_index' in results:
+ metainfo = dict(ignore_index=results['ignore_index'])
+ gt_sem_seg_data.set_metainfo(metainfo)
+ data_sample.gt_sem_seg = gt_sem_seg_data
+
+ img_meta = {}
+ for key in self.meta_keys:
+ if key in results:
+ img_meta[key] = results[key]
+ data_sample.set_metainfo(img_meta)
+ packed_results['data_samples'] = data_sample
+
+ return packed_results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'(meta_keys={self.meta_keys})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class ToTensor:
+ """Convert some results to :obj:`torch.Tensor` by given keys.
+
+ Args:
+ keys (Sequence[str]): Keys that need to be converted to Tensor.
+ """
+
+ def __init__(self, keys):
+ self.keys = keys
+
+ def __call__(self, results):
+ """Call function to convert data in results to :obj:`torch.Tensor`.
+
+ Args:
+ results (dict): Result dict contains the data to convert.
+
+ Returns:
+ dict: The result dict contains the data converted
+ to :obj:`torch.Tensor`.
+ """
+ for key in self.keys:
+ results[key] = to_tensor(results[key])
+ return results
+
+ def __repr__(self):
+ return self.__class__.__name__ + f'(keys={self.keys})'
+
+
+@TRANSFORMS.register_module()
+class ImageToTensor:
+ """Convert image to :obj:`torch.Tensor` by given keys.
+
+ The dimension order of input image is (H, W, C). The pipeline will convert
+ it to (C, H, W). If only 2 dimension (H, W) is given, the output would be
+ (1, H, W).
+
+ Args:
+ keys (Sequence[str]): Key of images to be converted to Tensor.
+ """
+
+ def __init__(self, keys):
+ self.keys = keys
+
+ def __call__(self, results):
+ """Call function to convert image in results to :obj:`torch.Tensor` and
+ transpose the channel order.
+
+ Args:
+ results (dict): Result dict contains the image data to convert.
+
+ Returns:
+ dict: The result dict contains the image converted
+ to :obj:`torch.Tensor` and permuted to (C, H, W) order.
+ """
+ for key in self.keys:
+ img = results[key]
+ if len(img.shape) < 3:
+ img = np.expand_dims(img, -1)
+ results[key] = to_tensor(img).permute(2, 0, 1).contiguous()
+
+ return results
+
+ def __repr__(self):
+ return self.__class__.__name__ + f'(keys={self.keys})'
+
+
+@TRANSFORMS.register_module()
+class Transpose:
+ """Transpose some results by given keys.
+
+ Args:
+ keys (Sequence[str]): Keys of results to be transposed.
+ order (Sequence[int]): Order of transpose.
+ """
+
+ def __init__(self, keys, order):
+ self.keys = keys
+ self.order = order
+
+ def __call__(self, results):
+ """Call function to transpose the channel order of data in results.
+
+ Args:
+ results (dict): Result dict contains the data to transpose.
+
+ Returns:
+ dict: The result dict contains the data transposed to \
+ ``self.order``.
+ """
+ for key in self.keys:
+ results[key] = results[key].transpose(self.order)
+ return results
+
+ def __repr__(self):
+ return self.__class__.__name__ + \
+ f'(keys={self.keys}, order={self.order})'
+
+
+@TRANSFORMS.register_module()
+class WrapFieldsToLists:
+ """Wrap fields of the data dictionary into lists for evaluation.
+
+ This class can be used as a last step of a test or validation
+ pipeline for single image evaluation or inference.
+
+ Example:
+ >>> test_pipeline = [
+ >>> dict(type='LoadImageFromFile'),
+ >>> dict(type='Normalize',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ to_rgb=True),
+ >>> dict(type='Pad', size_divisor=32),
+ >>> dict(type='ImageToTensor', keys=['img']),
+ >>> dict(type='Collect', keys=['img']),
+ >>> dict(type='WrapFieldsToLists')
+ >>> ]
+ """
+
+ def __call__(self, results):
+ """Call function to wrap fields into lists.
+
+ Args:
+ results (dict): Result dict contains the data to wrap.
+
+ Returns:
+ dict: The result dict where value of ``self.keys`` are wrapped \
+ into list.
+ """
+
+ # Wrap dict fields into lists
+ for key, val in results.items():
+ results[key] = [val]
+ return results
+
+ def __repr__(self):
+ return f'{self.__class__.__name__}()'
+
+
+@TRANSFORMS.register_module()
+class PackTrackInputs(BaseTransform):
+ """Pack the inputs data for the multi object tracking and video instance
+ segmentation. All the information of images are packed to ``inputs``. All
+ the information except images are packed to ``data_samples``. In order to
+ get the original annotaiton and meta info, we add `instances` key into meta
+ keys.
+
+ Args:
+ meta_keys (Sequence[str]): Meta keys to be collected in
+ ``data_sample.metainfo``. Defaults to None.
+ default_meta_keys (tuple): Default meta keys. Defaults to ('img_id',
+ 'img_path', 'ori_shape', 'img_shape', 'scale_factor',
+ 'flip', 'flip_direction', 'frame_id', 'is_video_data',
+ 'video_id', 'video_length', 'instances').
+ """
+ mapping_table = {
+ 'gt_bboxes': 'bboxes',
+ 'gt_bboxes_labels': 'labels',
+ 'gt_masks': 'masks',
+ 'gt_instances_ids': 'instances_ids'
+ }
+
+ def __init__(self,
+ meta_keys: Optional[dict] = None,
+ default_meta_keys: tuple = ('img_id', 'img_path', 'ori_shape',
+ 'img_shape', 'scale_factor',
+ 'flip', 'flip_direction',
+ 'frame_id', 'video_id',
+ 'video_length',
+ 'ori_video_length', 'instances')):
+ self.meta_keys = default_meta_keys
+ if meta_keys is not None:
+ if isinstance(meta_keys, str):
+ meta_keys = (meta_keys, )
+ else:
+ assert isinstance(meta_keys, tuple), \
+ 'meta_keys must be str or tuple'
+ self.meta_keys += meta_keys
+
+ def transform(self, results: dict) -> dict:
+ """Method to pack the input data.
+ Args:
+ results (dict): Result dict from the data pipeline.
+ Returns:
+ dict:
+ - 'inputs' (dict[Tensor]): The forward data of models.
+ - 'data_samples' (obj:`TrackDataSample`): The annotation info of
+ the samples.
+ """
+ packed_results = dict()
+ packed_results['inputs'] = dict()
+
+ # 1. Pack images
+ if 'img' in results:
+ imgs = results['img']
+ imgs = np.stack(imgs, axis=0)
+ imgs = imgs.transpose(0, 3, 1, 2)
+ packed_results['inputs'] = to_tensor(imgs)
+
+ # 2. Pack InstanceData
+ if 'gt_ignore_flags' in results:
+ gt_ignore_flags_list = results['gt_ignore_flags']
+ valid_idx_list, ignore_idx_list = [], []
+ for gt_ignore_flags in gt_ignore_flags_list:
+ valid_idx = np.where(gt_ignore_flags == 0)[0]
+ ignore_idx = np.where(gt_ignore_flags == 1)[0]
+ valid_idx_list.append(valid_idx)
+ ignore_idx_list.append(ignore_idx)
+
+ assert 'img_id' in results, "'img_id' must contained in the results "
+ 'for counting the number of images'
+
+ num_imgs = len(results['img_id'])
+ instance_data_list = [InstanceData() for _ in range(num_imgs)]
+ ignore_instance_data_list = [InstanceData() for _ in range(num_imgs)]
+
+ for key in self.mapping_table.keys():
+ if key not in results:
+ continue
+ if key == 'gt_masks':
+ mapped_key = self.mapping_table[key]
+ gt_masks_list = results[key]
+ if 'gt_ignore_flags' in results:
+ for i, gt_mask in enumerate(gt_masks_list):
+ valid_idx, ignore_idx = valid_idx_list[
+ i], ignore_idx_list[i]
+ instance_data_list[i][mapped_key] = gt_mask[valid_idx]
+ ignore_instance_data_list[i][mapped_key] = gt_mask[
+ ignore_idx]
+
+ else:
+ for i, gt_mask in enumerate(gt_masks_list):
+ instance_data_list[i][mapped_key] = gt_mask
+
+ else:
+ anns_list = results[key]
+ if 'gt_ignore_flags' in results:
+ for i, ann in enumerate(anns_list):
+ valid_idx, ignore_idx = valid_idx_list[
+ i], ignore_idx_list[i]
+ instance_data_list[i][
+ self.mapping_table[key]] = to_tensor(
+ ann[valid_idx])
+ ignore_instance_data_list[i][
+ self.mapping_table[key]] = to_tensor(
+ ann[ignore_idx])
+ else:
+ for i, ann in enumerate(anns_list):
+ instance_data_list[i][
+ self.mapping_table[key]] = to_tensor(ann)
+
+ det_data_samples_list = []
+ for i in range(num_imgs):
+ det_data_sample = DetDataSample()
+ det_data_sample.gt_instances = instance_data_list[i]
+ det_data_sample.ignored_instances = ignore_instance_data_list[i]
+ det_data_samples_list.append(det_data_sample)
+
+ # 3. Pack metainfo
+ for key in self.meta_keys:
+ if key not in results:
+ continue
+ img_metas_list = results[key]
+ for i, img_meta in enumerate(img_metas_list):
+ det_data_samples_list[i].set_metainfo({f'{key}': img_meta})
+
+ track_data_sample = TrackDataSample()
+ track_data_sample.video_data_samples = det_data_samples_list
+ if 'key_frame_flags' in results:
+ key_frame_flags = np.asarray(results['key_frame_flags'])
+ key_frames_inds = np.where(key_frame_flags)[0].tolist()
+ ref_frames_inds = np.where(~key_frame_flags)[0].tolist()
+ track_data_sample.set_metainfo(
+ dict(key_frames_inds=key_frames_inds))
+ track_data_sample.set_metainfo(
+ dict(ref_frames_inds=ref_frames_inds))
+
+ packed_results['data_samples'] = track_data_sample
+ return packed_results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'meta_keys={self.meta_keys}, '
+ repr_str += f'default_meta_keys={self.default_meta_keys})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class PackReIDInputs(BaseTransform):
+ """Pack the inputs data for the ReID. The ``meta_info`` item is always
+ populated. The contents of the ``meta_info`` dictionary depends on
+ ``meta_keys``. By default this includes:
+
+ - ``img_path``: path to the image file.
+ - ``ori_shape``: original shape of the image as a tuple (H, W).
+ - ``img_shape``: shape of the image input to the network as a tuple
+ (H, W). Note that images may be zero padded on the bottom/right
+ if the batch tensor is larger than this shape.
+ - ``scale``: scale of the image as a tuple (W, H).
+ - ``scale_factor``: a float indicating the pre-processing scale.
+ - ``flip``: a boolean indicating if image flip transform was used.
+ - ``flip_direction``: the flipping direction.
+ Args:
+ meta_keys (Sequence[str], optional): The meta keys to saved in the
+ ``metainfo`` of the packed ``data_sample``.
+ """
+ default_meta_keys = ('img_path', 'ori_shape', 'img_shape', 'scale',
+ 'scale_factor')
+
+ def __init__(self, meta_keys: Sequence[str] = ()) -> None:
+ self.meta_keys = self.default_meta_keys
+ if meta_keys is not None:
+ if isinstance(meta_keys, str):
+ meta_keys = (meta_keys, )
+ else:
+ assert isinstance(meta_keys, tuple), \
+ 'meta_keys must be str or tuple.'
+ self.meta_keys += meta_keys
+
+ def transform(self, results: dict) -> dict:
+ """Method to pack the input data.
+ Args:
+ results (dict): Result dict from the data pipeline.
+ Returns:
+ dict:
+ - 'inputs' (dict[Tensor]): The forward data of models.
+ - 'data_samples' (obj:`ReIDDataSample`): The meta info of the
+ sample.
+ """
+ packed_results = dict(inputs=dict(), data_samples=None)
+ assert 'img' in results, 'Missing the key ``img``.'
+ _type = type(results['img'])
+ label = results['gt_label']
+
+ if _type == list:
+ img = results['img']
+ label = np.stack(label, axis=0) # (N,)
+ assert all([type(v) == _type for v in results.values()]), \
+ 'All items in the results must have the same type.'
+ else:
+ img = [results['img']]
+
+ img = np.stack(img, axis=3) # (H, W, C, N)
+ img = img.transpose(3, 2, 0, 1) # (N, C, H, W)
+ img = np.ascontiguousarray(img)
+
+ packed_results['inputs'] = to_tensor(img)
+
+ data_sample = ReIDDataSample()
+ data_sample.set_gt_label(label)
+
+ meta_info = dict()
+ for key in self.meta_keys:
+ meta_info[key] = results[key]
+ data_sample.set_metainfo(meta_info)
+ packed_results['data_samples'] = data_sample
+
+ return packed_results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'(meta_keys={self.meta_keys})'
+ return repr_str
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/frame_sampling.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/frame_sampling.py
new file mode 100644
index 0000000000000000000000000000000000000000..a91f1e7880f8f061f183dc30a01758d97b7d03da
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/frame_sampling.py
@@ -0,0 +1,177 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import random
+from collections import defaultdict
+from typing import Dict, List, Optional, Union
+
+from mmcv.transforms import BaseTransform
+
+from mmdet.registry import TRANSFORMS
+
+
+@TRANSFORMS.register_module()
+class BaseFrameSample(BaseTransform):
+ """Directly get the key frame, no reference frames.
+
+ Args:
+ collect_video_keys (list[str]): The keys of video info to be
+ collected.
+ """
+
+ def __init__(self,
+ collect_video_keys: List[str] = ['video_id', 'video_length']):
+ self.collect_video_keys = collect_video_keys
+
+ def prepare_data(self, video_infos: dict,
+ sampled_inds: List[int]) -> Dict[str, List]:
+ """Prepare data for the subsequent pipeline.
+
+ Args:
+ video_infos (dict): The whole video information.
+ sampled_inds (list[int]): The sampled frame indices.
+
+ Returns:
+ dict: The processed data information.
+ """
+ frames_anns = video_infos['images']
+ final_data_info = defaultdict(list)
+ # for data in frames_anns:
+ for index in sampled_inds:
+ data = frames_anns[index]
+ # copy the info in video-level into img-level
+ for key in self.collect_video_keys:
+ if key == 'video_length':
+ data['ori_video_length'] = video_infos[key]
+ data['video_length'] = len(sampled_inds)
+ else:
+ data[key] = video_infos[key]
+ # Collate data_list (list of dict to dict of list)
+ for key, value in data.items():
+ final_data_info[key].append(value)
+
+ return final_data_info
+
+ def transform(self, video_infos: dict) -> Optional[Dict[str, List]]:
+ """Transform the video information.
+
+ Args:
+ video_infos (dict): The whole video information.
+
+ Returns:
+ dict: The data information of the key frames.
+ """
+ if 'key_frame_id' in video_infos:
+ key_frame_id = video_infos['key_frame_id']
+ assert isinstance(video_infos['key_frame_id'], int)
+ else:
+ key_frame_id = random.sample(
+ list(range(video_infos['video_length'])), 1)[0]
+ results = self.prepare_data(video_infos, [key_frame_id])
+
+ return results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'(collect_video_keys={self.collect_video_keys})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class UniformRefFrameSample(BaseFrameSample):
+ """Uniformly sample reference frames.
+
+ Args:
+ num_ref_imgs (int): Number of reference frames to be sampled.
+ frame_range (int | list[int]): Range of frames to be sampled around
+ key frame. If int, the range is [-frame_range, frame_range].
+ Defaults to 10.
+ filter_key_img (bool): Whether to filter the key frame when
+ sampling reference frames. Defaults to True.
+ collect_video_keys (list[str]): The keys of video info to be
+ collected.
+ """
+
+ def __init__(self,
+ num_ref_imgs: int = 1,
+ frame_range: Union[int, List[int]] = 10,
+ filter_key_img: bool = True,
+ collect_video_keys: List[str] = ['video_id', 'video_length']):
+ self.num_ref_imgs = num_ref_imgs
+ self.filter_key_img = filter_key_img
+ if isinstance(frame_range, int):
+ assert frame_range >= 0, 'frame_range can not be a negative value.'
+ frame_range = [-frame_range, frame_range]
+ elif isinstance(frame_range, list):
+ assert len(frame_range) == 2, 'The length must be 2.'
+ assert frame_range[0] <= 0 and frame_range[1] >= 0
+ for i in frame_range:
+ assert isinstance(i, int), 'Each element must be int.'
+ else:
+ raise TypeError('The type of frame_range must be int or list.')
+ self.frame_range = frame_range
+ super().__init__(collect_video_keys=collect_video_keys)
+
+ def sampling_frames(self, video_length: int, key_frame_id: int):
+ """Sampling frames.
+
+ Args:
+ video_length (int): The length of the video.
+ key_frame_id (int): The key frame id.
+
+ Returns:
+ list[int]: The sampled frame indices.
+ """
+ if video_length > 1:
+ left = max(0, key_frame_id + self.frame_range[0])
+ right = min(key_frame_id + self.frame_range[1], video_length - 1)
+ frame_ids = list(range(0, video_length))
+
+ valid_ids = frame_ids[left:right + 1]
+ if self.filter_key_img and key_frame_id in valid_ids:
+ valid_ids.remove(key_frame_id)
+ assert len(
+ valid_ids
+ ) > 0, 'After filtering key frame, there are no valid frames'
+ if len(valid_ids) < self.num_ref_imgs:
+ valid_ids = valid_ids * self.num_ref_imgs
+ ref_frame_ids = random.sample(valid_ids, self.num_ref_imgs)
+ else:
+ ref_frame_ids = [key_frame_id] * self.num_ref_imgs
+
+ sampled_frames_ids = [key_frame_id] + ref_frame_ids
+ sampled_frames_ids = sorted(sampled_frames_ids)
+
+ key_frames_ind = sampled_frames_ids.index(key_frame_id)
+ key_frame_flags = [False] * len(sampled_frames_ids)
+ key_frame_flags[key_frames_ind] = True
+ return sampled_frames_ids, key_frame_flags
+
+ def transform(self, video_infos: dict) -> Optional[Dict[str, List]]:
+ """Transform the video information.
+
+ Args:
+ video_infos (dict): The whole video information.
+
+ Returns:
+ dict: The data information of the sampled frames.
+ """
+ if 'key_frame_id' in video_infos:
+ key_frame_id = video_infos['key_frame_id']
+ assert isinstance(video_infos['key_frame_id'], int)
+ else:
+ key_frame_id = random.sample(
+ list(range(video_infos['video_length'])), 1)[0]
+
+ (sampled_frames_ids, key_frame_flags) = self.sampling_frames(
+ video_infos['video_length'], key_frame_id=key_frame_id)
+ results = self.prepare_data(video_infos, sampled_frames_ids)
+ results['key_frame_flags'] = key_frame_flags
+
+ return results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'(num_ref_imgs={self.num_ref_imgs}, '
+ repr_str += f'frame_range={self.frame_range}, '
+ repr_str += f'filter_key_img={self.filter_key_img}, '
+ repr_str += f'collect_video_keys={self.collect_video_keys})'
+ return repr_str
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/geometric.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/geometric.py
new file mode 100644
index 0000000000000000000000000000000000000000..d2cd6be258f73a69aa2c2b36fef64c6c4e46a2a4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/geometric.py
@@ -0,0 +1,754 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+from typing import Optional, Union
+
+import cv2
+import mmcv
+import numpy as np
+from mmcv.transforms import BaseTransform
+from mmcv.transforms.utils import cache_randomness
+
+from mmdet.registry import TRANSFORMS
+from mmdet.structures.bbox import autocast_box_type
+from .augment_wrappers import _MAX_LEVEL, level_to_mag
+
+
+@TRANSFORMS.register_module()
+class GeomTransform(BaseTransform):
+ """Base class for geometric transformations. All geometric transformations
+ need to inherit from this base class. ``GeomTransform`` unifies the class
+ attributes and class functions of geometric transformations (ShearX,
+ ShearY, Rotate, TranslateX, and TranslateY), and records the homography
+ matrix.
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_masks (BitmapMasks | PolygonMasks) (optional)
+ - gt_seg_map (np.uint8) (optional)
+
+ Modified Keys:
+
+ - img
+ - gt_bboxes
+ - gt_masks
+ - gt_seg_map
+
+ Added Keys:
+
+ - homography_matrix
+
+ Args:
+ prob (float): The probability for performing the geometric
+ transformation and should be in range [0, 1]. Defaults to 1.0.
+ level (int, optional): The level should be in range [0, _MAX_LEVEL].
+ If level is None, it will generate from [0, _MAX_LEVEL] randomly.
+ Defaults to None.
+ min_mag (float): The minimum magnitude for geometric transformation.
+ Defaults to 0.0.
+ max_mag (float): The maximum magnitude for geometric transformation.
+ Defaults to 1.0.
+ reversal_prob (float): The probability that reverses the geometric
+ transformation magnitude. Should be in range [0,1].
+ Defaults to 0.5.
+ img_border_value (int | float | tuple): The filled values for
+ image border. If float, the same fill value will be used for
+ all the three channels of image. If tuple, it should be 3 elements.
+ Defaults to 128.
+ mask_border_value (int): The fill value used for masks. Defaults to 0.
+ seg_ignore_label (int): The fill value used for segmentation map.
+ Note this value must equals ``ignore_label`` in ``semantic_head``
+ of the corresponding config. Defaults to 255.
+ interpolation (str): Interpolation method, accepted values are
+ "nearest", "bilinear", "bicubic", "area", "lanczos" for 'cv2'
+ backend, "nearest", "bilinear" for 'pillow' backend. Defaults
+ to 'bilinear'.
+ """
+
+ def __init__(self,
+ prob: float = 1.0,
+ level: Optional[int] = None,
+ min_mag: float = 0.0,
+ max_mag: float = 1.0,
+ reversal_prob: float = 0.5,
+ img_border_value: Union[int, float, tuple] = 128,
+ mask_border_value: int = 0,
+ seg_ignore_label: int = 255,
+ interpolation: str = 'bilinear') -> None:
+ assert 0 <= prob <= 1.0, f'The probability of the transformation ' \
+ f'should be in range [0,1], got {prob}.'
+ assert level is None or isinstance(level, int), \
+ f'The level should be None or type int, got {type(level)}.'
+ assert level is None or 0 <= level <= _MAX_LEVEL, \
+ f'The level should be in range [0,{_MAX_LEVEL}], got {level}.'
+ assert isinstance(min_mag, float), \
+ f'min_mag should be type float, got {type(min_mag)}.'
+ assert isinstance(max_mag, float), \
+ f'max_mag should be type float, got {type(max_mag)}.'
+ assert min_mag <= max_mag, \
+ f'min_mag should smaller than max_mag, ' \
+ f'got min_mag={min_mag} and max_mag={max_mag}'
+ assert isinstance(reversal_prob, float), \
+ f'reversal_prob should be type float, got {type(max_mag)}.'
+ assert 0 <= reversal_prob <= 1.0, \
+ f'The reversal probability of the transformation magnitude ' \
+ f'should be type float, got {type(reversal_prob)}.'
+ if isinstance(img_border_value, (float, int)):
+ img_border_value = tuple([float(img_border_value)] * 3)
+ elif isinstance(img_border_value, tuple):
+ assert len(img_border_value) == 3, \
+ f'img_border_value as tuple must have 3 elements, ' \
+ f'got {len(img_border_value)}.'
+ img_border_value = tuple([float(val) for val in img_border_value])
+ else:
+ raise ValueError(
+ 'img_border_value must be float or tuple with 3 elements.')
+ assert np.all([0 <= val <= 255 for val in img_border_value]), 'all ' \
+ 'elements of img_border_value should between range [0,255].' \
+ f'got {img_border_value}.'
+ self.prob = prob
+ self.level = level
+ self.min_mag = min_mag
+ self.max_mag = max_mag
+ self.reversal_prob = reversal_prob
+ self.img_border_value = img_border_value
+ self.mask_border_value = mask_border_value
+ self.seg_ignore_label = seg_ignore_label
+ self.interpolation = interpolation
+
+ def _transform_img(self, results: dict, mag: float) -> None:
+ """Transform the image."""
+ pass
+
+ def _transform_masks(self, results: dict, mag: float) -> None:
+ """Transform the masks."""
+ pass
+
+ def _transform_seg(self, results: dict, mag: float) -> None:
+ """Transform the segmentation map."""
+ pass
+
+ def _get_homography_matrix(self, results: dict, mag: float) -> np.ndarray:
+ """Get the homography matrix for the geometric transformation."""
+ return np.eye(3, dtype=np.float32)
+
+ def _transform_bboxes(self, results: dict, mag: float) -> None:
+ """Transform the bboxes."""
+ results['gt_bboxes'].project_(self.homography_matrix)
+ results['gt_bboxes'].clip_(results['img_shape'])
+
+ def _record_homography_matrix(self, results: dict) -> None:
+ """Record the homography matrix for the geometric transformation."""
+ if results.get('homography_matrix', None) is None:
+ results['homography_matrix'] = self.homography_matrix
+ else:
+ results['homography_matrix'] = self.homography_matrix @ results[
+ 'homography_matrix']
+
+ @cache_randomness
+ def _random_disable(self):
+ """Randomly disable the transform."""
+ return np.random.rand() > self.prob
+
+ @cache_randomness
+ def _get_mag(self):
+ """Get the magnitude of the transform."""
+ mag = level_to_mag(self.level, self.min_mag, self.max_mag)
+ return -mag if np.random.rand() > self.reversal_prob else mag
+
+ @autocast_box_type()
+ def transform(self, results: dict) -> dict:
+ """Transform function for images, bounding boxes, masks and semantic
+ segmentation map.
+
+ Args:
+ results (dict): Result dict from loading pipeline.
+
+ Returns:
+ dict: Transformed results.
+ """
+
+ if self._random_disable():
+ return results
+ mag = self._get_mag()
+ self.homography_matrix = self._get_homography_matrix(results, mag)
+ self._record_homography_matrix(results)
+ self._transform_img(results, mag)
+ if results.get('gt_bboxes', None) is not None:
+ self._transform_bboxes(results, mag)
+ if results.get('gt_masks', None) is not None:
+ self._transform_masks(results, mag)
+ if results.get('gt_seg_map', None) is not None:
+ self._transform_seg(results, mag)
+ return results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'(prob={self.prob}, '
+ repr_str += f'level={self.level}, '
+ repr_str += f'min_mag={self.min_mag}, '
+ repr_str += f'max_mag={self.max_mag}, '
+ repr_str += f'reversal_prob={self.reversal_prob}, '
+ repr_str += f'img_border_value={self.img_border_value}, '
+ repr_str += f'mask_border_value={self.mask_border_value}, '
+ repr_str += f'seg_ignore_label={self.seg_ignore_label}, '
+ repr_str += f'interpolation={self.interpolation})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class ShearX(GeomTransform):
+ """Shear the images, bboxes, masks and segmentation map horizontally.
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_masks (BitmapMasks | PolygonMasks) (optional)
+ - gt_seg_map (np.uint8) (optional)
+
+ Modified Keys:
+
+ - img
+ - gt_bboxes
+ - gt_masks
+ - gt_seg_map
+
+ Added Keys:
+
+ - homography_matrix
+
+ Args:
+ prob (float): The probability for performing Shear and should be in
+ range [0, 1]. Defaults to 1.0.
+ level (int, optional): The level should be in range [0, _MAX_LEVEL].
+ If level is None, it will generate from [0, _MAX_LEVEL] randomly.
+ Defaults to None.
+ min_mag (float): The minimum angle for the horizontal shear.
+ Defaults to 0.0.
+ max_mag (float): The maximum angle for the horizontal shear.
+ Defaults to 30.0.
+ reversal_prob (float): The probability that reverses the horizontal
+ shear magnitude. Should be in range [0,1]. Defaults to 0.5.
+ img_border_value (int | float | tuple): The filled values for
+ image border. If float, the same fill value will be used for
+ all the three channels of image. If tuple, it should be 3 elements.
+ Defaults to 128.
+ mask_border_value (int): The fill value used for masks. Defaults to 0.
+ seg_ignore_label (int): The fill value used for segmentation map.
+ Note this value must equals ``ignore_label`` in ``semantic_head``
+ of the corresponding config. Defaults to 255.
+ interpolation (str): Interpolation method, accepted values are
+ "nearest", "bilinear", "bicubic", "area", "lanczos" for 'cv2'
+ backend, "nearest", "bilinear" for 'pillow' backend. Defaults
+ to 'bilinear'.
+ """
+
+ def __init__(self,
+ prob: float = 1.0,
+ level: Optional[int] = None,
+ min_mag: float = 0.0,
+ max_mag: float = 30.0,
+ reversal_prob: float = 0.5,
+ img_border_value: Union[int, float, tuple] = 128,
+ mask_border_value: int = 0,
+ seg_ignore_label: int = 255,
+ interpolation: str = 'bilinear') -> None:
+ assert 0. <= min_mag <= 90., \
+ f'min_mag angle for ShearX should be ' \
+ f'in range [0, 90], got {min_mag}.'
+ assert 0. <= max_mag <= 90., \
+ f'max_mag angle for ShearX should be ' \
+ f'in range [0, 90], got {max_mag}.'
+ super().__init__(
+ prob=prob,
+ level=level,
+ min_mag=min_mag,
+ max_mag=max_mag,
+ reversal_prob=reversal_prob,
+ img_border_value=img_border_value,
+ mask_border_value=mask_border_value,
+ seg_ignore_label=seg_ignore_label,
+ interpolation=interpolation)
+
+ @cache_randomness
+ def _get_mag(self):
+ """Get the magnitude of the transform."""
+ mag = level_to_mag(self.level, self.min_mag, self.max_mag)
+ mag = np.tan(mag * np.pi / 180)
+ return -mag if np.random.rand() > self.reversal_prob else mag
+
+ def _get_homography_matrix(self, results: dict, mag: float) -> np.ndarray:
+ """Get the homography matrix for ShearX."""
+ return np.array([[1, mag, 0], [0, 1, 0], [0, 0, 1]], dtype=np.float32)
+
+ def _transform_img(self, results: dict, mag: float) -> None:
+ """Shear the image horizontally."""
+ results['img'] = mmcv.imshear(
+ results['img'],
+ mag,
+ direction='horizontal',
+ border_value=self.img_border_value,
+ interpolation=self.interpolation)
+
+ def _transform_masks(self, results: dict, mag: float) -> None:
+ """Shear the masks horizontally."""
+ results['gt_masks'] = results['gt_masks'].shear(
+ results['img_shape'],
+ mag,
+ direction='horizontal',
+ border_value=self.mask_border_value,
+ interpolation=self.interpolation)
+
+ def _transform_seg(self, results: dict, mag: float) -> None:
+ """Shear the segmentation map horizontally."""
+ results['gt_seg_map'] = mmcv.imshear(
+ results['gt_seg_map'],
+ mag,
+ direction='horizontal',
+ border_value=self.seg_ignore_label,
+ interpolation='nearest')
+
+
+@TRANSFORMS.register_module()
+class ShearY(GeomTransform):
+ """Shear the images, bboxes, masks and segmentation map vertically.
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_masks (BitmapMasks | PolygonMasks) (optional)
+ - gt_seg_map (np.uint8) (optional)
+
+ Modified Keys:
+
+ - img
+ - gt_bboxes
+ - gt_masks
+ - gt_seg_map
+
+ Added Keys:
+
+ - homography_matrix
+
+ Args:
+ prob (float): The probability for performing ShearY and should be in
+ range [0, 1]. Defaults to 1.0.
+ level (int, optional): The level should be in range [0,_MAX_LEVEL].
+ If level is None, it will generate from [0, _MAX_LEVEL] randomly.
+ Defaults to None.
+ min_mag (float): The minimum angle for the vertical shear.
+ Defaults to 0.0.
+ max_mag (float): The maximum angle for the vertical shear.
+ Defaults to 30.0.
+ reversal_prob (float): The probability that reverses the vertical
+ shear magnitude. Should be in range [0,1]. Defaults to 0.5.
+ img_border_value (int | float | tuple): The filled values for
+ image border. If float, the same fill value will be used for
+ all the three channels of image. If tuple, it should be 3 elements.
+ Defaults to 128.
+ mask_border_value (int): The fill value used for masks. Defaults to 0.
+ seg_ignore_label (int): The fill value used for segmentation map.
+ Note this value must equals ``ignore_label`` in ``semantic_head``
+ of the corresponding config. Defaults to 255.
+ interpolation (str): Interpolation method, accepted values are
+ "nearest", "bilinear", "bicubic", "area", "lanczos" for 'cv2'
+ backend, "nearest", "bilinear" for 'pillow' backend. Defaults
+ to 'bilinear'.
+ """
+
+ def __init__(self,
+ prob: float = 1.0,
+ level: Optional[int] = None,
+ min_mag: float = 0.0,
+ max_mag: float = 30.,
+ reversal_prob: float = 0.5,
+ img_border_value: Union[int, float, tuple] = 128,
+ mask_border_value: int = 0,
+ seg_ignore_label: int = 255,
+ interpolation: str = 'bilinear') -> None:
+ assert 0. <= min_mag <= 90., \
+ f'min_mag angle for ShearY should be ' \
+ f'in range [0, 90], got {min_mag}.'
+ assert 0. <= max_mag <= 90., \
+ f'max_mag angle for ShearY should be ' \
+ f'in range [0, 90], got {max_mag}.'
+ super().__init__(
+ prob=prob,
+ level=level,
+ min_mag=min_mag,
+ max_mag=max_mag,
+ reversal_prob=reversal_prob,
+ img_border_value=img_border_value,
+ mask_border_value=mask_border_value,
+ seg_ignore_label=seg_ignore_label,
+ interpolation=interpolation)
+
+ @cache_randomness
+ def _get_mag(self):
+ """Get the magnitude of the transform."""
+ mag = level_to_mag(self.level, self.min_mag, self.max_mag)
+ mag = np.tan(mag * np.pi / 180)
+ return -mag if np.random.rand() > self.reversal_prob else mag
+
+ def _get_homography_matrix(self, results: dict, mag: float) -> np.ndarray:
+ """Get the homography matrix for ShearY."""
+ return np.array([[1, 0, 0], [mag, 1, 0], [0, 0, 1]], dtype=np.float32)
+
+ def _transform_img(self, results: dict, mag: float) -> None:
+ """Shear the image vertically."""
+ results['img'] = mmcv.imshear(
+ results['img'],
+ mag,
+ direction='vertical',
+ border_value=self.img_border_value,
+ interpolation=self.interpolation)
+
+ def _transform_masks(self, results: dict, mag: float) -> None:
+ """Shear the masks vertically."""
+ results['gt_masks'] = results['gt_masks'].shear(
+ results['img_shape'],
+ mag,
+ direction='vertical',
+ border_value=self.mask_border_value,
+ interpolation=self.interpolation)
+
+ def _transform_seg(self, results: dict, mag: float) -> None:
+ """Shear the segmentation map vertically."""
+ results['gt_seg_map'] = mmcv.imshear(
+ results['gt_seg_map'],
+ mag,
+ direction='vertical',
+ border_value=self.seg_ignore_label,
+ interpolation='nearest')
+
+
+@TRANSFORMS.register_module()
+class Rotate(GeomTransform):
+ """Rotate the images, bboxes, masks and segmentation map.
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_masks (BitmapMasks | PolygonMasks) (optional)
+ - gt_seg_map (np.uint8) (optional)
+
+ Modified Keys:
+
+ - img
+ - gt_bboxes
+ - gt_masks
+ - gt_seg_map
+
+ Added Keys:
+
+ - homography_matrix
+
+ Args:
+ prob (float): The probability for perform transformation and
+ should be in range 0 to 1. Defaults to 1.0.
+ level (int, optional): The level should be in range [0, _MAX_LEVEL].
+ If level is None, it will generate from [0, _MAX_LEVEL] randomly.
+ Defaults to None.
+ min_mag (float): The maximum angle for rotation.
+ Defaults to 0.0.
+ max_mag (float): The maximum angle for rotation.
+ Defaults to 30.0.
+ reversal_prob (float): The probability that reverses the rotation
+ magnitude. Should be in range [0,1]. Defaults to 0.5.
+ img_border_value (int | float | tuple): The filled values for
+ image border. If float, the same fill value will be used for
+ all the three channels of image. If tuple, it should be 3 elements.
+ Defaults to 128.
+ mask_border_value (int): The fill value used for masks. Defaults to 0.
+ seg_ignore_label (int): The fill value used for segmentation map.
+ Note this value must equals ``ignore_label`` in ``semantic_head``
+ of the corresponding config. Defaults to 255.
+ interpolation (str): Interpolation method, accepted values are
+ "nearest", "bilinear", "bicubic", "area", "lanczos" for 'cv2'
+ backend, "nearest", "bilinear" for 'pillow' backend. Defaults
+ to 'bilinear'.
+ """
+
+ def __init__(self,
+ prob: float = 1.0,
+ level: Optional[int] = None,
+ min_mag: float = 0.0,
+ max_mag: float = 30.0,
+ reversal_prob: float = 0.5,
+ img_border_value: Union[int, float, tuple] = 128,
+ mask_border_value: int = 0,
+ seg_ignore_label: int = 255,
+ interpolation: str = 'bilinear') -> None:
+ assert 0. <= min_mag <= 180., \
+ f'min_mag for Rotate should be in range [0,180], got {min_mag}.'
+ assert 0. <= max_mag <= 180., \
+ f'max_mag for Rotate should be in range [0,180], got {max_mag}.'
+ super().__init__(
+ prob=prob,
+ level=level,
+ min_mag=min_mag,
+ max_mag=max_mag,
+ reversal_prob=reversal_prob,
+ img_border_value=img_border_value,
+ mask_border_value=mask_border_value,
+ seg_ignore_label=seg_ignore_label,
+ interpolation=interpolation)
+
+ def _get_homography_matrix(self, results: dict, mag: float) -> np.ndarray:
+ """Get the homography matrix for Rotate."""
+ img_shape = results['img_shape']
+ center = ((img_shape[1] - 1) * 0.5, (img_shape[0] - 1) * 0.5)
+ cv2_rotation_matrix = cv2.getRotationMatrix2D(center, -mag, 1.0)
+ return np.concatenate(
+ [cv2_rotation_matrix,
+ np.array([0, 0, 1]).reshape((1, 3))]).astype(np.float32)
+
+ def _transform_img(self, results: dict, mag: float) -> None:
+ """Rotate the image."""
+ results['img'] = mmcv.imrotate(
+ results['img'],
+ mag,
+ border_value=self.img_border_value,
+ interpolation=self.interpolation)
+
+ def _transform_masks(self, results: dict, mag: float) -> None:
+ """Rotate the masks."""
+ results['gt_masks'] = results['gt_masks'].rotate(
+ results['img_shape'],
+ mag,
+ border_value=self.mask_border_value,
+ interpolation=self.interpolation)
+
+ def _transform_seg(self, results: dict, mag: float) -> None:
+ """Rotate the segmentation map."""
+ results['gt_seg_map'] = mmcv.imrotate(
+ results['gt_seg_map'],
+ mag,
+ border_value=self.seg_ignore_label,
+ interpolation='nearest')
+
+
+@TRANSFORMS.register_module()
+class TranslateX(GeomTransform):
+ """Translate the images, bboxes, masks and segmentation map horizontally.
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_masks (BitmapMasks | PolygonMasks) (optional)
+ - gt_seg_map (np.uint8) (optional)
+
+ Modified Keys:
+
+ - img
+ - gt_bboxes
+ - gt_masks
+ - gt_seg_map
+
+ Added Keys:
+
+ - homography_matrix
+
+ Args:
+ prob (float): The probability for perform transformation and
+ should be in range 0 to 1. Defaults to 1.0.
+ level (int, optional): The level should be in range [0, _MAX_LEVEL].
+ If level is None, it will generate from [0, _MAX_LEVEL] randomly.
+ Defaults to None.
+ min_mag (float): The minimum pixel's offset ratio for horizontal
+ translation. Defaults to 0.0.
+ max_mag (float): The maximum pixel's offset ratio for horizontal
+ translation. Defaults to 0.1.
+ reversal_prob (float): The probability that reverses the horizontal
+ translation magnitude. Should be in range [0,1]. Defaults to 0.5.
+ img_border_value (int | float | tuple): The filled values for
+ image border. If float, the same fill value will be used for
+ all the three channels of image. If tuple, it should be 3 elements.
+ Defaults to 128.
+ mask_border_value (int): The fill value used for masks. Defaults to 0.
+ seg_ignore_label (int): The fill value used for segmentation map.
+ Note this value must equals ``ignore_label`` in ``semantic_head``
+ of the corresponding config. Defaults to 255.
+ interpolation (str): Interpolation method, accepted values are
+ "nearest", "bilinear", "bicubic", "area", "lanczos" for 'cv2'
+ backend, "nearest", "bilinear" for 'pillow' backend. Defaults
+ to 'bilinear'.
+ """
+
+ def __init__(self,
+ prob: float = 1.0,
+ level: Optional[int] = None,
+ min_mag: float = 0.0,
+ max_mag: float = 0.1,
+ reversal_prob: float = 0.5,
+ img_border_value: Union[int, float, tuple] = 128,
+ mask_border_value: int = 0,
+ seg_ignore_label: int = 255,
+ interpolation: str = 'bilinear') -> None:
+ assert 0. <= min_mag <= 1., \
+ f'min_mag ratio for TranslateX should be ' \
+ f'in range [0, 1], got {min_mag}.'
+ assert 0. <= max_mag <= 1., \
+ f'max_mag ratio for TranslateX should be ' \
+ f'in range [0, 1], got {max_mag}.'
+ super().__init__(
+ prob=prob,
+ level=level,
+ min_mag=min_mag,
+ max_mag=max_mag,
+ reversal_prob=reversal_prob,
+ img_border_value=img_border_value,
+ mask_border_value=mask_border_value,
+ seg_ignore_label=seg_ignore_label,
+ interpolation=interpolation)
+
+ def _get_homography_matrix(self, results: dict, mag: float) -> np.ndarray:
+ """Get the homography matrix for TranslateX."""
+ mag = int(results['img_shape'][1] * mag)
+ return np.array([[1, 0, mag], [0, 1, 0], [0, 0, 1]], dtype=np.float32)
+
+ def _transform_img(self, results: dict, mag: float) -> None:
+ """Translate the image horizontally."""
+ mag = int(results['img_shape'][1] * mag)
+ results['img'] = mmcv.imtranslate(
+ results['img'],
+ mag,
+ direction='horizontal',
+ border_value=self.img_border_value,
+ interpolation=self.interpolation)
+
+ def _transform_masks(self, results: dict, mag: float) -> None:
+ """Translate the masks horizontally."""
+ mag = int(results['img_shape'][1] * mag)
+ results['gt_masks'] = results['gt_masks'].translate(
+ results['img_shape'],
+ mag,
+ direction='horizontal',
+ border_value=self.mask_border_value,
+ interpolation=self.interpolation)
+
+ def _transform_seg(self, results: dict, mag: float) -> None:
+ """Translate the segmentation map horizontally."""
+ mag = int(results['img_shape'][1] * mag)
+ results['gt_seg_map'] = mmcv.imtranslate(
+ results['gt_seg_map'],
+ mag,
+ direction='horizontal',
+ border_value=self.seg_ignore_label,
+ interpolation='nearest')
+
+
+@TRANSFORMS.register_module()
+class TranslateY(GeomTransform):
+ """Translate the images, bboxes, masks and segmentation map vertically.
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_masks (BitmapMasks | PolygonMasks) (optional)
+ - gt_seg_map (np.uint8) (optional)
+
+ Modified Keys:
+
+ - img
+ - gt_bboxes
+ - gt_masks
+ - gt_seg_map
+
+ Added Keys:
+
+ - homography_matrix
+
+ Args:
+ prob (float): The probability for perform transformation and
+ should be in range 0 to 1. Defaults to 1.0.
+ level (int, optional): The level should be in range [0, _MAX_LEVEL].
+ If level is None, it will generate from [0, _MAX_LEVEL] randomly.
+ Defaults to None.
+ min_mag (float): The minimum pixel's offset ratio for vertical
+ translation. Defaults to 0.0.
+ max_mag (float): The maximum pixel's offset ratio for vertical
+ translation. Defaults to 0.1.
+ reversal_prob (float): The probability that reverses the vertical
+ translation magnitude. Should be in range [0,1]. Defaults to 0.5.
+ img_border_value (int | float | tuple): The filled values for
+ image border. If float, the same fill value will be used for
+ all the three channels of image. If tuple, it should be 3 elements.
+ Defaults to 128.
+ mask_border_value (int): The fill value used for masks. Defaults to 0.
+ seg_ignore_label (int): The fill value used for segmentation map.
+ Note this value must equals ``ignore_label`` in ``semantic_head``
+ of the corresponding config. Defaults to 255.
+ interpolation (str): Interpolation method, accepted values are
+ "nearest", "bilinear", "bicubic", "area", "lanczos" for 'cv2'
+ backend, "nearest", "bilinear" for 'pillow' backend. Defaults
+ to 'bilinear'.
+ """
+
+ def __init__(self,
+ prob: float = 1.0,
+ level: Optional[int] = None,
+ min_mag: float = 0.0,
+ max_mag: float = 0.1,
+ reversal_prob: float = 0.5,
+ img_border_value: Union[int, float, tuple] = 128,
+ mask_border_value: int = 0,
+ seg_ignore_label: int = 255,
+ interpolation: str = 'bilinear') -> None:
+ assert 0. <= min_mag <= 1., \
+ f'min_mag ratio for TranslateY should be ' \
+ f'in range [0,1], got {min_mag}.'
+ assert 0. <= max_mag <= 1., \
+ f'max_mag ratio for TranslateY should be ' \
+ f'in range [0,1], got {max_mag}.'
+ super().__init__(
+ prob=prob,
+ level=level,
+ min_mag=min_mag,
+ max_mag=max_mag,
+ reversal_prob=reversal_prob,
+ img_border_value=img_border_value,
+ mask_border_value=mask_border_value,
+ seg_ignore_label=seg_ignore_label,
+ interpolation=interpolation)
+
+ def _get_homography_matrix(self, results: dict, mag: float) -> np.ndarray:
+ """Get the homography matrix for TranslateY."""
+ mag = int(results['img_shape'][0] * mag)
+ return np.array([[1, 0, 0], [0, 1, mag], [0, 0, 1]], dtype=np.float32)
+
+ def _transform_img(self, results: dict, mag: float) -> None:
+ """Translate the image vertically."""
+ mag = int(results['img_shape'][0] * mag)
+ results['img'] = mmcv.imtranslate(
+ results['img'],
+ mag,
+ direction='vertical',
+ border_value=self.img_border_value,
+ interpolation=self.interpolation)
+
+ def _transform_masks(self, results: dict, mag: float) -> None:
+ """Translate masks vertically."""
+ mag = int(results['img_shape'][0] * mag)
+ results['gt_masks'] = results['gt_masks'].translate(
+ results['img_shape'],
+ mag,
+ direction='vertical',
+ border_value=self.mask_border_value,
+ interpolation=self.interpolation)
+
+ def _transform_seg(self, results: dict, mag: float) -> None:
+ """Translate segmentation map vertically."""
+ mag = int(results['img_shape'][0] * mag)
+ results['gt_seg_map'] = mmcv.imtranslate(
+ results['gt_seg_map'],
+ mag,
+ direction='vertical',
+ border_value=self.seg_ignore_label,
+ interpolation='nearest')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/instaboost.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/instaboost.py
new file mode 100644
index 0000000000000000000000000000000000000000..30dc1603643ec8d398bfade95f5ec1c9b8f89c8d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/instaboost.py
@@ -0,0 +1,150 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Tuple
+
+import numpy as np
+from mmcv.transforms import BaseTransform
+
+from mmdet.registry import TRANSFORMS
+
+
+@TRANSFORMS.register_module()
+class InstaBoost(BaseTransform):
+ r"""Data augmentation method in `InstaBoost: Boosting Instance
+ Segmentation Via Probability Map Guided Copy-Pasting
+ `_.
+
+ Refer to https://github.com/GothicAi/Instaboost for implementation details.
+
+
+ Required Keys:
+
+ - img (np.uint8)
+ - instances
+
+ Modified Keys:
+
+ - img (np.uint8)
+ - instances
+
+ Args:
+ action_candidate (tuple): Action candidates. "normal", "horizontal", \
+ "vertical", "skip" are supported. Defaults to ('normal', \
+ 'horizontal', 'skip').
+ action_prob (tuple): Corresponding action probabilities. Should be \
+ the same length as action_candidate. Defaults to (1, 0, 0).
+ scale (tuple): (min scale, max scale). Defaults to (0.8, 1.2).
+ dx (int): The maximum x-axis shift will be (instance width) / dx.
+ Defaults to 15.
+ dy (int): The maximum y-axis shift will be (instance height) / dy.
+ Defaults to 15.
+ theta (tuple): (min rotation degree, max rotation degree). \
+ Defaults to (-1, 1).
+ color_prob (float): Probability of images for color augmentation.
+ Defaults to 0.5.
+ hflag (bool): Whether to use heatmap guided. Defaults to False.
+ aug_ratio (float): Probability of applying this transformation. \
+ Defaults to 0.5.
+ """
+
+ def __init__(self,
+ action_candidate: tuple = ('normal', 'horizontal', 'skip'),
+ action_prob: tuple = (1, 0, 0),
+ scale: tuple = (0.8, 1.2),
+ dx: int = 15,
+ dy: int = 15,
+ theta: tuple = (-1, 1),
+ color_prob: float = 0.5,
+ hflag: bool = False,
+ aug_ratio: float = 0.5) -> None:
+
+ import matplotlib
+ import matplotlib.pyplot as plt
+ default_backend = plt.get_backend()
+
+ try:
+ import instaboostfast as instaboost
+ except ImportError:
+ raise ImportError(
+ 'Please run "pip install instaboostfast" '
+ 'to install instaboostfast first for instaboost augmentation.')
+
+ # instaboost will modify the default backend
+ # and cause visualization to fail.
+ matplotlib.use(default_backend)
+
+ self.cfg = instaboost.InstaBoostConfig(action_candidate, action_prob,
+ scale, dx, dy, theta,
+ color_prob, hflag)
+ self.aug_ratio = aug_ratio
+
+ def _load_anns(self, results: dict) -> Tuple[list, list]:
+ """Convert raw anns to instaboost expected input format."""
+ anns = []
+ ignore_anns = []
+ for instance in results['instances']:
+ label = instance['bbox_label']
+ bbox = instance['bbox']
+ mask = instance['mask']
+ x1, y1, x2, y2 = bbox
+ # assert (x2 - x1) >= 1 and (y2 - y1) >= 1
+ bbox = [x1, y1, x2 - x1, y2 - y1]
+
+ if instance['ignore_flag'] == 0:
+ anns.append({
+ 'category_id': label,
+ 'segmentation': mask,
+ 'bbox': bbox
+ })
+ else:
+ # Ignore instances without data augmentation
+ ignore_anns.append(instance)
+ return anns, ignore_anns
+
+ def _parse_anns(self, results: dict, anns: list, ignore_anns: list,
+ img: np.ndarray) -> dict:
+ """Restore the result of instaboost processing to the original anns
+ format."""
+ instances = []
+ for ann in anns:
+ x1, y1, w, h = ann['bbox']
+ # TODO: more essential bug need to be fixed in instaboost
+ if w <= 0 or h <= 0:
+ continue
+ bbox = [x1, y1, x1 + w, y1 + h]
+ instances.append(
+ dict(
+ bbox=bbox,
+ bbox_label=ann['category_id'],
+ mask=ann['segmentation'],
+ ignore_flag=0))
+
+ instances.extend(ignore_anns)
+ results['img'] = img
+ results['instances'] = instances
+ return results
+
+ def transform(self, results) -> dict:
+ """The transform function."""
+ img = results['img']
+ ori_type = img.dtype
+ if 'instances' not in results or len(results['instances']) == 0:
+ return results
+
+ anns, ignore_anns = self._load_anns(results)
+ if np.random.choice([0, 1], p=[1 - self.aug_ratio, self.aug_ratio]):
+ try:
+ import instaboostfast as instaboost
+ except ImportError:
+ raise ImportError('Please run "pip install instaboostfast" '
+ 'to install instaboostfast first.')
+ anns, img = instaboost.get_new_data(
+ anns, img.astype(np.uint8), self.cfg, background=None)
+
+ results = self._parse_anns(results, anns, ignore_anns,
+ img.astype(ori_type))
+ return results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'(aug_ratio={self.aug_ratio})'
+ return repr_str
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/loading.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/loading.py
new file mode 100644
index 0000000000000000000000000000000000000000..722d4b0e7c830dfde2412746db1258b880167a2f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/loading.py
@@ -0,0 +1,1074 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Tuple, Union
+
+import mmcv
+import numpy as np
+import pycocotools.mask as maskUtils
+import torch
+from mmcv.transforms import BaseTransform
+from mmcv.transforms import LoadAnnotations as MMCV_LoadAnnotations
+from mmcv.transforms import LoadImageFromFile
+from mmengine.fileio import get
+from mmengine.structures import BaseDataElement
+
+from mmdet.registry import TRANSFORMS
+from mmdet.structures.bbox import get_box_type
+from mmdet.structures.bbox.box_type import autocast_box_type
+from mmdet.structures.mask import BitmapMasks, PolygonMasks
+
+
+@TRANSFORMS.register_module()
+class LoadImageFromNDArray(LoadImageFromFile):
+ """Load an image from ``results['img']``.
+
+ Similar with :obj:`LoadImageFromFile`, but the image has been loaded as
+ :obj:`np.ndarray` in ``results['img']``. Can be used when loading image
+ from webcam.
+
+ Required Keys:
+
+ - img
+
+ Modified Keys:
+
+ - img
+ - img_path
+ - img_shape
+ - ori_shape
+
+ Args:
+ to_float32 (bool): Whether to convert the loaded image to a float32
+ numpy array. If set to False, the loaded image is an uint8 array.
+ Defaults to False.
+ """
+
+ def transform(self, results: dict) -> dict:
+ """Transform function to add image meta information.
+
+ Args:
+ results (dict): Result dict with Webcam read image in
+ ``results['img']``.
+
+ Returns:
+ dict: The dict contains loaded image and meta information.
+ """
+
+ img = results['img']
+ if self.to_float32:
+ img = img.astype(np.float32)
+
+ results['img_path'] = None
+ results['img'] = img
+ results['img_shape'] = img.shape[:2]
+ results['ori_shape'] = img.shape[:2]
+ return results
+
+
+@TRANSFORMS.register_module()
+class LoadMultiChannelImageFromFiles(BaseTransform):
+ """Load multi-channel images from a list of separate channel files.
+
+ Required Keys:
+
+ - img_path
+
+ Modified Keys:
+
+ - img
+ - img_shape
+ - ori_shape
+
+ Args:
+ to_float32 (bool): Whether to convert the loaded image to a float32
+ numpy array. If set to False, the loaded image is an uint8 array.
+ Defaults to False.
+ color_type (str): The flag argument for :func:``mmcv.imfrombytes``.
+ Defaults to 'unchanged'.
+ imdecode_backend (str): The image decoding backend type. The backend
+ argument for :func:``mmcv.imfrombytes``.
+ See :func:``mmcv.imfrombytes`` for details.
+ Defaults to 'cv2'.
+ file_client_args (dict): Arguments to instantiate the
+ corresponding backend in mmdet <= 3.0.0rc6. Defaults to None.
+ backend_args (dict, optional): Arguments to instantiate the
+ corresponding backend in mmdet >= 3.0.0rc7. Defaults to None.
+ """
+
+ def __init__(
+ self,
+ to_float32: bool = False,
+ color_type: str = 'unchanged',
+ imdecode_backend: str = 'cv2',
+ file_client_args: dict = None,
+ backend_args: dict = None,
+ ) -> None:
+ self.to_float32 = to_float32
+ self.color_type = color_type
+ self.imdecode_backend = imdecode_backend
+ self.backend_args = backend_args
+ if file_client_args is not None:
+ raise RuntimeError(
+ 'The `file_client_args` is deprecated, '
+ 'please use `backend_args` instead, please refer to'
+ 'https://github.com/open-mmlab/mmdetection/blob/main/configs/_base_/datasets/coco_detection.py' # noqa: E501
+ )
+
+ def transform(self, results: dict) -> dict:
+ """Transform functions to load multiple images and get images meta
+ information.
+
+ Args:
+ results (dict): Result dict from :obj:`mmdet.CustomDataset`.
+
+ Returns:
+ dict: The dict contains loaded images and meta information.
+ """
+
+ assert isinstance(results['img_path'], list)
+ img = []
+ for name in results['img_path']:
+ img_bytes = get(name, backend_args=self.backend_args)
+ img.append(
+ mmcv.imfrombytes(
+ img_bytes,
+ flag=self.color_type,
+ backend=self.imdecode_backend))
+ img = np.stack(img, axis=-1)
+ if self.to_float32:
+ img = img.astype(np.float32)
+
+ results['img'] = img
+ results['img_shape'] = img.shape[:2]
+ results['ori_shape'] = img.shape[:2]
+ return results
+
+ def __repr__(self):
+ repr_str = (f'{self.__class__.__name__}('
+ f'to_float32={self.to_float32}, '
+ f"color_type='{self.color_type}', "
+ f"imdecode_backend='{self.imdecode_backend}', "
+ f'backend_args={self.backend_args})')
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class LoadAnnotations(MMCV_LoadAnnotations):
+ """Load and process the ``instances`` and ``seg_map`` annotation provided
+ by dataset.
+
+ The annotation format is as the following:
+
+ .. code-block:: python
+
+ {
+ 'instances':
+ [
+ {
+ # List of 4 numbers representing the bounding box of the
+ # instance, in (x1, y1, x2, y2) order.
+ 'bbox': [x1, y1, x2, y2],
+
+ # Label of image classification.
+ 'bbox_label': 1,
+
+ # Used in instance/panoptic segmentation. The segmentation mask
+ # of the instance or the information of segments.
+ # 1. If list[list[float]], it represents a list of polygons,
+ # one for each connected component of the object. Each
+ # list[float] is one simple polygon in the format of
+ # [x1, y1, ..., xn, yn] (n >= 3). The Xs and Ys are absolute
+ # coordinates in unit of pixels.
+ # 2. If dict, it represents the per-pixel segmentation mask in
+ # COCO's compressed RLE format. The dict should have keys
+ # “size” and “counts”. Can be loaded by pycocotools
+ 'mask': list[list[float]] or dict,
+
+ }
+ ]
+ # Filename of semantic or panoptic segmentation ground truth file.
+ 'seg_map_path': 'a/b/c'
+ }
+
+ After this module, the annotation has been changed to the format below:
+
+ .. code-block:: python
+
+ {
+ # In (x1, y1, x2, y2) order, float type. N is the number of bboxes
+ # in an image
+ 'gt_bboxes': BaseBoxes(N, 4)
+ # In int type.
+ 'gt_bboxes_labels': np.ndarray(N, )
+ # In built-in class
+ 'gt_masks': PolygonMasks (H, W) or BitmapMasks (H, W)
+ # In uint8 type.
+ 'gt_seg_map': np.ndarray (H, W)
+ # in (x, y, v) order, float type.
+ }
+
+ Required Keys:
+
+ - height
+ - width
+ - instances
+
+ - bbox (optional)
+ - bbox_label
+ - mask (optional)
+ - ignore_flag
+
+ - seg_map_path (optional)
+
+ Added Keys:
+
+ - gt_bboxes (BaseBoxes[torch.float32])
+ - gt_bboxes_labels (np.int64)
+ - gt_masks (BitmapMasks | PolygonMasks)
+ - gt_seg_map (np.uint8)
+ - gt_ignore_flags (bool)
+
+ Args:
+ with_bbox (bool): Whether to parse and load the bbox annotation.
+ Defaults to True.
+ with_label (bool): Whether to parse and load the label annotation.
+ Defaults to True.
+ with_mask (bool): Whether to parse and load the mask annotation.
+ Default: False.
+ with_seg (bool): Whether to parse and load the semantic segmentation
+ annotation. Defaults to False.
+ poly2mask (bool): Whether to convert mask to bitmap. Default: True.
+ box_type (str): The box type used to wrap the bboxes. If ``box_type``
+ is None, gt_bboxes will keep being np.ndarray. Defaults to 'hbox'.
+ reduce_zero_label (bool): Whether reduce all label value
+ by 1. Usually used for datasets where 0 is background label.
+ Defaults to False.
+ ignore_index (int): The label index to be ignored.
+ Valid only if reduce_zero_label is true. Defaults is 255.
+ imdecode_backend (str): The image decoding backend type. The backend
+ argument for :func:``mmcv.imfrombytes``.
+ See :fun:``mmcv.imfrombytes`` for details.
+ Defaults to 'cv2'.
+ backend_args (dict, optional): Arguments to instantiate the
+ corresponding backend. Defaults to None.
+ """
+
+ def __init__(
+ self,
+ with_mask: bool = False,
+ poly2mask: bool = True,
+ box_type: str = 'hbox',
+ # use for semseg
+ reduce_zero_label: bool = False,
+ ignore_index: int = 255,
+ **kwargs) -> None:
+ super(LoadAnnotations, self).__init__(**kwargs)
+ self.with_mask = with_mask
+ self.poly2mask = poly2mask
+ self.box_type = box_type
+ self.reduce_zero_label = reduce_zero_label
+ self.ignore_index = ignore_index
+
+ def _load_bboxes(self, results: dict) -> None:
+ """Private function to load bounding box annotations.
+
+ Args:
+ results (dict): Result dict from :obj:``mmengine.BaseDataset``.
+ Returns:
+ dict: The dict contains loaded bounding box annotations.
+ """
+ gt_bboxes = []
+ gt_ignore_flags = []
+ for instance in results.get('instances', []):
+ gt_bboxes.append(instance['bbox'])
+ gt_ignore_flags.append(instance['ignore_flag'])
+ if self.box_type is None:
+ results['gt_bboxes'] = np.array(
+ gt_bboxes, dtype=np.float32).reshape((-1, 4))
+ else:
+ _, box_type_cls = get_box_type(self.box_type)
+ results['gt_bboxes'] = box_type_cls(gt_bboxes, dtype=torch.float32)
+ results['gt_ignore_flags'] = np.array(gt_ignore_flags, dtype=bool)
+
+ def _load_labels(self, results: dict) -> None:
+ """Private function to load label annotations.
+
+ Args:
+ results (dict): Result dict from :obj:``mmengine.BaseDataset``.
+
+ Returns:
+ dict: The dict contains loaded label annotations.
+ """
+ gt_bboxes_labels = []
+ for instance in results.get('instances', []):
+ gt_bboxes_labels.append(instance['bbox_label'])
+ # TODO: Inconsistent with mmcv, consider how to deal with it later.
+ results['gt_bboxes_labels'] = np.array(
+ gt_bboxes_labels, dtype=np.int64)
+
+ def _poly2mask(self, mask_ann: Union[list, dict], img_h: int,
+ img_w: int) -> np.ndarray:
+ """Private function to convert masks represented with polygon to
+ bitmaps.
+
+ Args:
+ mask_ann (list | dict): Polygon mask annotation input.
+ img_h (int): The height of output mask.
+ img_w (int): The width of output mask.
+
+ Returns:
+ np.ndarray: The decode bitmap mask of shape (img_h, img_w).
+ """
+
+ if isinstance(mask_ann, list):
+ # polygon -- a single object might consist of multiple parts
+ # we merge all parts into one mask rle code
+ rles = maskUtils.frPyObjects(mask_ann, img_h, img_w)
+ rle = maskUtils.merge(rles)
+ elif isinstance(mask_ann['counts'], list):
+ # uncompressed RLE
+ rle = maskUtils.frPyObjects(mask_ann, img_h, img_w)
+ else:
+ # rle
+ rle = mask_ann
+ mask = maskUtils.decode(rle)
+ return mask
+
+ def _process_masks(self, results: dict) -> list:
+ """Process gt_masks and filter invalid polygons.
+
+ Args:
+ results (dict): Result dict from :obj:``mmengine.BaseDataset``.
+
+ Returns:
+ list: Processed gt_masks.
+ """
+ gt_masks = []
+ gt_ignore_flags = []
+ for instance in results.get('instances', []):
+ gt_mask = instance['mask']
+ # If the annotation of segmentation mask is invalid,
+ # ignore the whole instance.
+ if isinstance(gt_mask, list):
+ gt_mask = [
+ np.array(polygon) for polygon in gt_mask
+ if len(polygon) % 2 == 0 and len(polygon) >= 6
+ ]
+ if len(gt_mask) == 0:
+ # ignore this instance and set gt_mask to a fake mask
+ instance['ignore_flag'] = 1
+ gt_mask = [np.zeros(6)]
+ elif not self.poly2mask:
+ # `PolygonMasks` requires a ploygon of format List[np.array],
+ # other formats are invalid.
+ instance['ignore_flag'] = 1
+ gt_mask = [np.zeros(6)]
+ elif isinstance(gt_mask, dict) and \
+ not (gt_mask.get('counts') is not None and
+ gt_mask.get('size') is not None and
+ isinstance(gt_mask['counts'], (list, str))):
+ # if gt_mask is a dict, it should include `counts` and `size`,
+ # so that `BitmapMasks` can uncompressed RLE
+ instance['ignore_flag'] = 1
+ gt_mask = [np.zeros(6)]
+ gt_masks.append(gt_mask)
+ # re-process gt_ignore_flags
+ gt_ignore_flags.append(instance['ignore_flag'])
+ results['gt_ignore_flags'] = np.array(gt_ignore_flags, dtype=bool)
+ return gt_masks
+
+ def _load_masks(self, results: dict) -> None:
+ """Private function to load mask annotations.
+
+ Args:
+ results (dict): Result dict from :obj:``mmengine.BaseDataset``.
+ """
+ h, w = results['ori_shape']
+ gt_masks = self._process_masks(results)
+ if self.poly2mask:
+ gt_masks = BitmapMasks(
+ [self._poly2mask(mask, h, w) for mask in gt_masks], h, w)
+ else:
+ # fake polygon masks will be ignored in `PackDetInputs`
+ gt_masks = PolygonMasks([mask for mask in gt_masks], h, w)
+ results['gt_masks'] = gt_masks
+
+ def _load_seg_map(self, results: dict) -> None:
+ """Private function to load semantic segmentation annotations.
+
+ Args:
+ results (dict): Result dict from :obj:``mmcv.BaseDataset``.
+
+ Returns:
+ dict: The dict contains loaded semantic segmentation annotations.
+ """
+ if results.get('seg_map_path', None) is None:
+ return
+
+ img_bytes = get(
+ results['seg_map_path'], backend_args=self.backend_args)
+ gt_semantic_seg = mmcv.imfrombytes(
+ img_bytes, flag='unchanged',
+ backend=self.imdecode_backend).squeeze()
+
+ if self.reduce_zero_label:
+ # avoid using underflow conversion
+ gt_semantic_seg[gt_semantic_seg == 0] = self.ignore_index
+ gt_semantic_seg = gt_semantic_seg - 1
+ gt_semantic_seg[gt_semantic_seg == self.ignore_index -
+ 1] = self.ignore_index
+
+ # modify if custom classes
+ if results.get('label_map', None) is not None:
+ # Add deep copy to solve bug of repeatedly
+ # replace `gt_semantic_seg`, which is reported in
+ # https://github.com/open-mmlab/mmsegmentation/pull/1445/
+ gt_semantic_seg_copy = gt_semantic_seg.copy()
+ for old_id, new_id in results['label_map'].items():
+ gt_semantic_seg[gt_semantic_seg_copy == old_id] = new_id
+ results['gt_seg_map'] = gt_semantic_seg
+ results['ignore_index'] = self.ignore_index
+
+ def transform(self, results: dict) -> dict:
+ """Function to load multiple types annotations.
+
+ Args:
+ results (dict): Result dict from :obj:``mmengine.BaseDataset``.
+
+ Returns:
+ dict: The dict contains loaded bounding box, label and
+ semantic segmentation.
+ """
+
+ if self.with_bbox:
+ self._load_bboxes(results)
+ if self.with_label:
+ self._load_labels(results)
+ if self.with_mask:
+ self._load_masks(results)
+ if self.with_seg:
+ self._load_seg_map(results)
+ return results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'(with_bbox={self.with_bbox}, '
+ repr_str += f'with_label={self.with_label}, '
+ repr_str += f'with_mask={self.with_mask}, '
+ repr_str += f'with_seg={self.with_seg}, '
+ repr_str += f'poly2mask={self.poly2mask}, '
+ repr_str += f"imdecode_backend='{self.imdecode_backend}', "
+ repr_str += f'backend_args={self.backend_args})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class LoadPanopticAnnotations(LoadAnnotations):
+ """Load multiple types of panoptic annotations.
+
+ The annotation format is as the following:
+
+ .. code-block:: python
+
+ {
+ 'instances':
+ [
+ {
+ # List of 4 numbers representing the bounding box of the
+ # instance, in (x1, y1, x2, y2) order.
+ 'bbox': [x1, y1, x2, y2],
+
+ # Label of image classification.
+ 'bbox_label': 1,
+ },
+ ...
+ ]
+ 'segments_info':
+ [
+ {
+ # id = cls_id + instance_id * INSTANCE_OFFSET
+ 'id': int,
+
+ # Contiguous category id defined in dataset.
+ 'category': int
+
+ # Thing flag.
+ 'is_thing': bool
+ },
+ ...
+ ]
+
+ # Filename of semantic or panoptic segmentation ground truth file.
+ 'seg_map_path': 'a/b/c'
+ }
+
+ After this module, the annotation has been changed to the format below:
+
+ .. code-block:: python
+
+ {
+ # In (x1, y1, x2, y2) order, float type. N is the number of bboxes
+ # in an image
+ 'gt_bboxes': BaseBoxes(N, 4)
+ # In int type.
+ 'gt_bboxes_labels': np.ndarray(N, )
+ # In built-in class
+ 'gt_masks': PolygonMasks (H, W) or BitmapMasks (H, W)
+ # In uint8 type.
+ 'gt_seg_map': np.ndarray (H, W)
+ # in (x, y, v) order, float type.
+ }
+
+ Required Keys:
+
+ - height
+ - width
+ - instances
+ - bbox
+ - bbox_label
+ - ignore_flag
+ - segments_info
+ - id
+ - category
+ - is_thing
+ - seg_map_path
+
+ Added Keys:
+
+ - gt_bboxes (BaseBoxes[torch.float32])
+ - gt_bboxes_labels (np.int64)
+ - gt_masks (BitmapMasks | PolygonMasks)
+ - gt_seg_map (np.uint8)
+ - gt_ignore_flags (bool)
+
+ Args:
+ with_bbox (bool): Whether to parse and load the bbox annotation.
+ Defaults to True.
+ with_label (bool): Whether to parse and load the label annotation.
+ Defaults to True.
+ with_mask (bool): Whether to parse and load the mask annotation.
+ Defaults to True.
+ with_seg (bool): Whether to parse and load the semantic segmentation
+ annotation. Defaults to False.
+ box_type (str): The box mode used to wrap the bboxes.
+ imdecode_backend (str): The image decoding backend type. The backend
+ argument for :func:``mmcv.imfrombytes``.
+ See :fun:``mmcv.imfrombytes`` for details.
+ Defaults to 'cv2'.
+ backend_args (dict, optional): Arguments to instantiate the
+ corresponding backend in mmdet >= 3.0.0rc7. Defaults to None.
+ """
+
+ def __init__(self,
+ with_bbox: bool = True,
+ with_label: bool = True,
+ with_mask: bool = True,
+ with_seg: bool = True,
+ box_type: str = 'hbox',
+ imdecode_backend: str = 'cv2',
+ backend_args: dict = None) -> None:
+ try:
+ from panopticapi import utils
+ except ImportError:
+ raise ImportError(
+ 'panopticapi is not installed, please install it by: '
+ 'pip install git+https://github.com/cocodataset/'
+ 'panopticapi.git.')
+ self.rgb2id = utils.rgb2id
+
+ super(LoadPanopticAnnotations, self).__init__(
+ with_bbox=with_bbox,
+ with_label=with_label,
+ with_mask=with_mask,
+ with_seg=with_seg,
+ with_keypoints=False,
+ box_type=box_type,
+ imdecode_backend=imdecode_backend,
+ backend_args=backend_args)
+
+ def _load_masks_and_semantic_segs(self, results: dict) -> None:
+ """Private function to load mask and semantic segmentation annotations.
+
+ In gt_semantic_seg, the foreground label is from ``0`` to
+ ``num_things - 1``, the background label is from ``num_things`` to
+ ``num_things + num_stuff - 1``, 255 means the ignored label (``VOID``).
+
+ Args:
+ results (dict): Result dict from :obj:``mmdet.CustomDataset``.
+ """
+ # seg_map_path is None, when inference on the dataset without gts.
+ if results.get('seg_map_path', None) is None:
+ return
+
+ img_bytes = get(
+ results['seg_map_path'], backend_args=self.backend_args)
+ pan_png = mmcv.imfrombytes(
+ img_bytes, flag='color', channel_order='rgb').squeeze()
+ pan_png = self.rgb2id(pan_png)
+
+ gt_masks = []
+ gt_seg = np.zeros_like(pan_png) + 255 # 255 as ignore
+
+ for segment_info in results['segments_info']:
+ mask = (pan_png == segment_info['id'])
+ gt_seg = np.where(mask, segment_info['category'], gt_seg)
+
+ # The legal thing masks
+ if segment_info.get('is_thing'):
+ gt_masks.append(mask.astype(np.uint8))
+
+ if self.with_mask:
+ h, w = results['ori_shape']
+ gt_masks = BitmapMasks(gt_masks, h, w)
+ results['gt_masks'] = gt_masks
+
+ if self.with_seg:
+ results['gt_seg_map'] = gt_seg
+
+ def transform(self, results: dict) -> dict:
+ """Function to load multiple types panoptic annotations.
+
+ Args:
+ results (dict): Result dict from :obj:``mmdet.CustomDataset``.
+
+ Returns:
+ dict: The dict contains loaded bounding box, label, mask and
+ semantic segmentation annotations.
+ """
+
+ if self.with_bbox:
+ self._load_bboxes(results)
+ if self.with_label:
+ self._load_labels(results)
+ if self.with_mask or self.with_seg:
+ # The tasks completed by '_load_masks' and '_load_semantic_segs'
+ # in LoadAnnotations are merged to one function.
+ self._load_masks_and_semantic_segs(results)
+
+ return results
+
+
+@TRANSFORMS.register_module()
+class LoadProposals(BaseTransform):
+ """Load proposal pipeline.
+
+ Required Keys:
+
+ - proposals
+
+ Modified Keys:
+
+ - proposals
+
+ Args:
+ num_max_proposals (int, optional): Maximum number of proposals to load.
+ If not specified, all proposals will be loaded.
+ """
+
+ def __init__(self, num_max_proposals: Optional[int] = None) -> None:
+ self.num_max_proposals = num_max_proposals
+
+ def transform(self, results: dict) -> dict:
+ """Transform function to load proposals from file.
+
+ Args:
+ results (dict): Result dict from :obj:`mmdet.CustomDataset`.
+
+ Returns:
+ dict: The dict contains loaded proposal annotations.
+ """
+
+ proposals = results['proposals']
+ # the type of proposals should be `dict` or `InstanceData`
+ assert isinstance(proposals, dict) \
+ or isinstance(proposals, BaseDataElement)
+ bboxes = proposals['bboxes'].astype(np.float32)
+ assert bboxes.shape[1] == 4, \
+ f'Proposals should have shapes (n, 4), but found {bboxes.shape}'
+
+ if 'scores' in proposals:
+ scores = proposals['scores'].astype(np.float32)
+ assert bboxes.shape[0] == scores.shape[0]
+ else:
+ scores = np.zeros(bboxes.shape[0], dtype=np.float32)
+
+ if self.num_max_proposals is not None:
+ # proposals should sort by scores during dumping the proposals
+ bboxes = bboxes[:self.num_max_proposals]
+ scores = scores[:self.num_max_proposals]
+
+ if len(bboxes) == 0:
+ bboxes = np.zeros((0, 4), dtype=np.float32)
+ scores = np.zeros(0, dtype=np.float32)
+
+ results['proposals'] = bboxes
+ results['proposals_scores'] = scores
+ return results
+
+ def __repr__(self):
+ return self.__class__.__name__ + \
+ f'(num_max_proposals={self.num_max_proposals})'
+
+
+@TRANSFORMS.register_module()
+class FilterAnnotations(BaseTransform):
+ """Filter invalid annotations.
+
+ Required Keys:
+
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_bboxes_labels (np.int64) (optional)
+ - gt_masks (BitmapMasks | PolygonMasks) (optional)
+ - gt_ignore_flags (bool) (optional)
+
+ Modified Keys:
+
+ - gt_bboxes (optional)
+ - gt_bboxes_labels (optional)
+ - gt_masks (optional)
+ - gt_ignore_flags (optional)
+
+ Args:
+ min_gt_bbox_wh (tuple[float]): Minimum width and height of ground truth
+ boxes. Default: (1., 1.)
+ min_gt_mask_area (int): Minimum foreground area of ground truth masks.
+ Default: 1
+ by_box (bool): Filter instances with bounding boxes not meeting the
+ min_gt_bbox_wh threshold. Default: True
+ by_mask (bool): Filter instances with masks not meeting
+ min_gt_mask_area threshold. Default: False
+ keep_empty (bool): Whether to return None when it
+ becomes an empty bbox after filtering. Defaults to True.
+ """
+
+ def __init__(self,
+ min_gt_bbox_wh: Tuple[int, int] = (1, 1),
+ min_gt_mask_area: int = 1,
+ by_box: bool = True,
+ by_mask: bool = False,
+ keep_empty: bool = True) -> None:
+ # TODO: add more filter options
+ assert by_box or by_mask
+ self.min_gt_bbox_wh = min_gt_bbox_wh
+ self.min_gt_mask_area = min_gt_mask_area
+ self.by_box = by_box
+ self.by_mask = by_mask
+ self.keep_empty = keep_empty
+
+ @autocast_box_type()
+ def transform(self, results: dict) -> Union[dict, None]:
+ """Transform function to filter annotations.
+
+ Args:
+ results (dict): Result dict.
+
+ Returns:
+ dict: Updated result dict.
+ """
+ assert 'gt_bboxes' in results
+ gt_bboxes = results['gt_bboxes']
+ if gt_bboxes.shape[0] == 0:
+ return results
+
+ tests = []
+ if self.by_box:
+ tests.append(
+ ((gt_bboxes.widths > self.min_gt_bbox_wh[0]) &
+ (gt_bboxes.heights > self.min_gt_bbox_wh[1])).numpy())
+ if self.by_mask:
+ assert 'gt_masks' in results
+ gt_masks = results['gt_masks']
+ tests.append(gt_masks.areas >= self.min_gt_mask_area)
+
+ keep = tests[0]
+ for t in tests[1:]:
+ keep = keep & t
+
+ if not keep.any():
+ if self.keep_empty:
+ return None
+
+ keys = ('gt_bboxes', 'gt_bboxes_labels', 'gt_masks', 'gt_ignore_flags')
+ for key in keys:
+ if key in results:
+ results[key] = results[key][keep]
+
+ return results
+
+ def __repr__(self):
+ return self.__class__.__name__ + \
+ f'(min_gt_bbox_wh={self.min_gt_bbox_wh}, ' \
+ f'keep_empty={self.keep_empty})'
+
+
+@TRANSFORMS.register_module()
+class LoadEmptyAnnotations(BaseTransform):
+ """Load Empty Annotations for unlabeled images.
+
+ Added Keys:
+ - gt_bboxes (np.float32)
+ - gt_bboxes_labels (np.int64)
+ - gt_masks (BitmapMasks | PolygonMasks)
+ - gt_seg_map (np.uint8)
+ - gt_ignore_flags (bool)
+
+ Args:
+ with_bbox (bool): Whether to load the pseudo bbox annotation.
+ Defaults to True.
+ with_label (bool): Whether to load the pseudo label annotation.
+ Defaults to True.
+ with_mask (bool): Whether to load the pseudo mask annotation.
+ Default: False.
+ with_seg (bool): Whether to load the pseudo semantic segmentation
+ annotation. Defaults to False.
+ seg_ignore_label (int): The fill value used for segmentation map.
+ Note this value must equals ``ignore_label`` in ``semantic_head``
+ of the corresponding config. Defaults to 255.
+ """
+
+ def __init__(self,
+ with_bbox: bool = True,
+ with_label: bool = True,
+ with_mask: bool = False,
+ with_seg: bool = False,
+ seg_ignore_label: int = 255) -> None:
+ self.with_bbox = with_bbox
+ self.with_label = with_label
+ self.with_mask = with_mask
+ self.with_seg = with_seg
+ self.seg_ignore_label = seg_ignore_label
+
+ def transform(self, results: dict) -> dict:
+ """Transform function to load empty annotations.
+
+ Args:
+ results (dict): Result dict.
+ Returns:
+ dict: Updated result dict.
+ """
+ if self.with_bbox:
+ results['gt_bboxes'] = np.zeros((0, 4), dtype=np.float32)
+ results['gt_ignore_flags'] = np.zeros((0, ), dtype=bool)
+ if self.with_label:
+ results['gt_bboxes_labels'] = np.zeros((0, ), dtype=np.int64)
+ if self.with_mask:
+ # TODO: support PolygonMasks
+ h, w = results['img_shape']
+ gt_masks = np.zeros((0, h, w), dtype=np.uint8)
+ results['gt_masks'] = BitmapMasks(gt_masks, h, w)
+ if self.with_seg:
+ h, w = results['img_shape']
+ results['gt_seg_map'] = self.seg_ignore_label * np.ones(
+ (h, w), dtype=np.uint8)
+ return results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'(with_bbox={self.with_bbox}, '
+ repr_str += f'with_label={self.with_label}, '
+ repr_str += f'with_mask={self.with_mask}, '
+ repr_str += f'with_seg={self.with_seg}, '
+ repr_str += f'seg_ignore_label={self.seg_ignore_label})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class InferencerLoader(BaseTransform):
+ """Load an image from ``results['img']``.
+
+ Similar with :obj:`LoadImageFromFile`, but the image has been loaded as
+ :obj:`np.ndarray` in ``results['img']``. Can be used when loading image
+ from webcam.
+
+ Required Keys:
+
+ - img
+
+ Modified Keys:
+
+ - img
+ - img_path
+ - img_shape
+ - ori_shape
+
+ Args:
+ to_float32 (bool): Whether to convert the loaded image to a float32
+ numpy array. If set to False, the loaded image is an uint8 array.
+ Defaults to False.
+ """
+
+ def __init__(self, **kwargs) -> None:
+ super().__init__()
+ self.from_file = TRANSFORMS.build(
+ dict(type='LoadImageFromFile', **kwargs))
+ self.from_ndarray = TRANSFORMS.build(
+ dict(type='mmdet.LoadImageFromNDArray', **kwargs))
+
+ def transform(self, results: Union[str, np.ndarray, dict]) -> dict:
+ """Transform function to add image meta information.
+
+ Args:
+ results (str, np.ndarray or dict): The result.
+
+ Returns:
+ dict: The dict contains loaded image and meta information.
+ """
+ if isinstance(results, str):
+ inputs = dict(img_path=results)
+ elif isinstance(results, np.ndarray):
+ inputs = dict(img=results)
+ elif isinstance(results, dict):
+ inputs = results
+ else:
+ raise NotImplementedError
+
+ if 'img' in inputs:
+ return self.from_ndarray(inputs)
+ return self.from_file(inputs)
+
+
+@TRANSFORMS.register_module()
+class LoadTrackAnnotations(LoadAnnotations):
+ """Load and process the ``instances`` and ``seg_map`` annotation provided
+ by dataset. It must load ``instances_ids`` which is only used in the
+ tracking tasks. The annotation format is as the following:
+
+ .. code-block:: python
+ {
+ 'instances':
+ [
+ {
+ # List of 4 numbers representing the bounding box of the
+ # instance, in (x1, y1, x2, y2) order.
+ 'bbox': [x1, y1, x2, y2],
+ # Label of image classification.
+ 'bbox_label': 1,
+ # Used in tracking.
+ # Id of instances.
+ 'instance_id': 100,
+ # Used in instance/panoptic segmentation. The segmentation mask
+ # of the instance or the information of segments.
+ # 1. If list[list[float]], it represents a list of polygons,
+ # one for each connected component of the object. Each
+ # list[float] is one simple polygon in the format of
+ # [x1, y1, ..., xn, yn] (n >= 3). The Xs and Ys are absolute
+ # coordinates in unit of pixels.
+ # 2. If dict, it represents the per-pixel segmentation mask in
+ # COCO's compressed RLE format. The dict should have keys
+ # “size” and “counts”. Can be loaded by pycocotools
+ 'mask': list[list[float]] or dict,
+ }
+ ]
+ # Filename of semantic or panoptic segmentation ground truth file.
+ 'seg_map_path': 'a/b/c'
+ }
+
+ After this module, the annotation has been changed to the format below:
+ .. code-block:: python
+ {
+ # In (x1, y1, x2, y2) order, float type. N is the number of bboxes
+ # in an image
+ 'gt_bboxes': np.ndarray(N, 4)
+ # In int type.
+ 'gt_bboxes_labels': np.ndarray(N, )
+ # In built-in class
+ 'gt_masks': PolygonMasks (H, W) or BitmapMasks (H, W)
+ # In uint8 type.
+ 'gt_seg_map': np.ndarray (H, W)
+ # in (x, y, v) order, float type.
+ }
+
+ Required Keys:
+
+ - height (optional)
+ - width (optional)
+ - instances
+ - bbox (optional)
+ - bbox_label
+ - instance_id (optional)
+ - mask (optional)
+ - ignore_flag (optional)
+ - seg_map_path (optional)
+
+ Added Keys:
+
+ - gt_bboxes (np.float32)
+ - gt_bboxes_labels (np.int32)
+ - gt_instances_ids (np.int32)
+ - gt_masks (BitmapMasks | PolygonMasks)
+ - gt_seg_map (np.uint8)
+ - gt_ignore_flags (np.bool)
+ """
+
+ def __init__(self, **kwargs) -> None:
+ super().__init__(**kwargs)
+
+ def _load_bboxes(self, results: dict) -> None:
+ """Private function to load bounding box annotations.
+
+ Args:
+ results (dict): Result dict from :obj:``mmcv.BaseDataset``.
+
+ Returns:
+ dict: The dict contains loaded bounding box annotations.
+ """
+ gt_bboxes = []
+ gt_ignore_flags = []
+ # TODO: use bbox_type
+ for instance in results['instances']:
+ # The datasets which are only format in evaluation don't have
+ # groundtruth boxes.
+ if 'bbox' in instance:
+ gt_bboxes.append(instance['bbox'])
+ if 'ignore_flag' in instance:
+ gt_ignore_flags.append(instance['ignore_flag'])
+
+ # TODO: check this case
+ if len(gt_bboxes) != len(gt_ignore_flags):
+ # There may be no ``gt_ignore_flags`` in some cases, we treat them
+ # as all False in order to keep the length of ``gt_bboxes`` and
+ # ``gt_ignore_flags`` the same
+ gt_ignore_flags = [False] * len(gt_bboxes)
+
+ results['gt_bboxes'] = np.array(
+ gt_bboxes, dtype=np.float32).reshape(-1, 4)
+ results['gt_ignore_flags'] = np.array(gt_ignore_flags, dtype=bool)
+
+ def _load_instances_ids(self, results: dict) -> None:
+ """Private function to load instances id annotations.
+
+ Args:
+ results (dict): Result dict from :obj :obj:``mmcv.BaseDataset``.
+
+ Returns:
+ dict: The dict containing instances id annotations.
+ """
+ gt_instances_ids = []
+ for instance in results['instances']:
+ gt_instances_ids.append(instance['instance_id'])
+ results['gt_instances_ids'] = np.array(
+ gt_instances_ids, dtype=np.int32)
+
+ def transform(self, results: dict) -> dict:
+ """Function to load multiple types annotations.
+
+ Args:
+ results (dict): Result dict from :obj:``mmcv.BaseDataset``.
+
+ Returns:
+ dict: The dict contains loaded bounding box, label, instances id
+ and semantic segmentation and keypoints annotations.
+ """
+ results = super().transform(results)
+ self._load_instances_ids(results)
+ return results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'(with_bbox={self.with_bbox}, '
+ repr_str += f'with_label={self.with_label}, '
+ repr_str += f'with_mask={self.with_mask}, '
+ repr_str += f'with_seg={self.with_seg}, '
+ repr_str += f'poly2mask={self.poly2mask}, '
+ repr_str += f"imdecode_backend='{self.imdecode_backend}', "
+ repr_str += f'file_client_args={self.file_client_args})'
+ return repr_str
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/text_transformers.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/text_transformers.py
new file mode 100644
index 0000000000000000000000000000000000000000..12a0e57db3d41baa6f5b7d1834ba74538ad9ca19
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/text_transformers.py
@@ -0,0 +1,255 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import json
+
+from mmcv.transforms import BaseTransform
+
+from mmdet.registry import TRANSFORMS
+from mmdet.structures.bbox import BaseBoxes
+
+try:
+ from transformers import AutoTokenizer
+ from transformers import BertModel as HFBertModel
+except ImportError:
+ AutoTokenizer = None
+ HFBertModel = None
+
+import random
+import re
+
+import numpy as np
+
+
+def clean_name(name):
+ name = re.sub(r'\(.*\)', '', name)
+ name = re.sub(r'_', ' ', name)
+ name = re.sub(r' ', ' ', name)
+ name = name.lower()
+ return name
+
+
+def check_for_positive_overflow(gt_bboxes, gt_labels, text, tokenizer,
+ max_tokens):
+ # Check if we have too many positive labels
+ # generate a caption by appending the positive labels
+ positive_label_list = np.unique(gt_labels).tolist()
+ # random shuffule so we can sample different annotations
+ # at different epochs
+ random.shuffle(positive_label_list)
+
+ kept_lables = []
+ length = 0
+
+ for index, label in enumerate(positive_label_list):
+
+ label_text = clean_name(text[str(label)]) + '. '
+
+ tokenized = tokenizer.tokenize(label_text)
+
+ length += len(tokenized)
+
+ if length > max_tokens:
+ break
+ else:
+ kept_lables.append(label)
+
+ keep_box_index = []
+ keep_gt_labels = []
+ for i in range(len(gt_labels)):
+ if gt_labels[i] in kept_lables:
+ keep_box_index.append(i)
+ keep_gt_labels.append(gt_labels[i])
+
+ return gt_bboxes[keep_box_index], np.array(
+ keep_gt_labels, dtype=np.long), length
+
+
+def generate_senetence_given_labels(positive_label_list, negative_label_list,
+ text):
+ label_to_positions = {}
+
+ label_list = negative_label_list + positive_label_list
+
+ random.shuffle(label_list)
+
+ pheso_caption = ''
+
+ label_remap_dict = {}
+ for index, label in enumerate(label_list):
+
+ start_index = len(pheso_caption)
+
+ pheso_caption += clean_name(text[str(label)])
+
+ end_index = len(pheso_caption)
+
+ if label in positive_label_list:
+ label_to_positions[index] = [[start_index, end_index]]
+ label_remap_dict[int(label)] = index
+
+ # if index != len(label_list) - 1:
+ # pheso_caption += '. '
+ pheso_caption += '. '
+
+ return label_to_positions, pheso_caption, label_remap_dict
+
+
+@TRANSFORMS.register_module()
+class RandomSamplingNegPos(BaseTransform):
+
+ def __init__(self,
+ tokenizer_name,
+ num_sample_negative=85,
+ max_tokens=256,
+ full_sampling_prob=0.5,
+ label_map_file=None):
+ if AutoTokenizer is None:
+ raise RuntimeError(
+ 'transformers is not installed, please install it by: '
+ 'pip install transformers.')
+
+ self.tokenizer = AutoTokenizer.from_pretrained(tokenizer_name)
+ self.num_sample_negative = num_sample_negative
+ self.full_sampling_prob = full_sampling_prob
+ self.max_tokens = max_tokens
+ self.label_map = None
+ if label_map_file:
+ with open(label_map_file, 'r') as file:
+ self.label_map = json.load(file)
+
+ def transform(self, results: dict) -> dict:
+ if 'phrases' in results:
+ return self.vg_aug(results)
+ else:
+ return self.od_aug(results)
+
+ def vg_aug(self, results):
+ gt_bboxes = results['gt_bboxes']
+ if isinstance(gt_bboxes, BaseBoxes):
+ gt_bboxes = gt_bboxes.tensor
+ gt_labels = results['gt_bboxes_labels']
+ text = results['text'].lower().strip()
+ if not text.endswith('.'):
+ text = text + '. '
+
+ phrases = results['phrases']
+ # TODO: add neg
+ positive_label_list = np.unique(gt_labels).tolist()
+ label_to_positions = {}
+ for label in positive_label_list:
+ label_to_positions[label] = phrases[label]['tokens_positive']
+
+ results['gt_bboxes'] = gt_bboxes
+ results['gt_bboxes_labels'] = gt_labels
+
+ results['text'] = text
+ results['tokens_positive'] = label_to_positions
+ return results
+
+ def od_aug(self, results):
+ gt_bboxes = results['gt_bboxes']
+ if isinstance(gt_bboxes, BaseBoxes):
+ gt_bboxes = gt_bboxes.tensor
+ gt_labels = results['gt_bboxes_labels']
+
+ if 'text' not in results:
+ assert self.label_map is not None
+ text = self.label_map
+ else:
+ text = results['text']
+
+ original_box_num = len(gt_labels)
+ # If the category name is in the format of 'a/b' (in object365),
+ # we randomly select one of them.
+ for key, value in text.items():
+ if '/' in value:
+ text[key] = random.choice(value.split('/')).strip()
+
+ gt_bboxes, gt_labels, positive_caption_length = \
+ check_for_positive_overflow(gt_bboxes, gt_labels,
+ text, self.tokenizer, self.max_tokens)
+
+ if len(gt_bboxes) < original_box_num:
+ print('WARNING: removed {} boxes due to positive caption overflow'.
+ format(original_box_num - len(gt_bboxes)))
+
+ valid_negative_indexes = list(text.keys())
+
+ positive_label_list = np.unique(gt_labels).tolist()
+ full_negative = self.num_sample_negative
+
+ if full_negative > len(valid_negative_indexes):
+ full_negative = len(valid_negative_indexes)
+
+ outer_prob = random.random()
+
+ if outer_prob < self.full_sampling_prob:
+ # c. probability_full: add both all positive and all negatives
+ num_negatives = full_negative
+ else:
+ if random.random() < 1.0:
+ num_negatives = np.random.choice(max(1, full_negative)) + 1
+ else:
+ num_negatives = full_negative
+
+ # Keep some negatives
+ negative_label_list = set()
+ if num_negatives != -1:
+ if num_negatives > len(valid_negative_indexes):
+ num_negatives = len(valid_negative_indexes)
+
+ for i in np.random.choice(
+ valid_negative_indexes, size=num_negatives, replace=False):
+ if int(i) not in positive_label_list:
+ negative_label_list.add(i)
+
+ random.shuffle(positive_label_list)
+
+ negative_label_list = list(negative_label_list)
+ random.shuffle(negative_label_list)
+
+ negative_max_length = self.max_tokens - positive_caption_length
+ screened_negative_label_list = []
+
+ for negative_label in negative_label_list:
+ label_text = clean_name(text[str(negative_label)]) + '. '
+
+ tokenized = self.tokenizer.tokenize(label_text)
+
+ negative_max_length -= len(tokenized)
+
+ if negative_max_length > 0:
+ screened_negative_label_list.append(negative_label)
+ else:
+ break
+ negative_label_list = screened_negative_label_list
+ label_to_positions, pheso_caption, label_remap_dict = \
+ generate_senetence_given_labels(positive_label_list,
+ negative_label_list, text)
+
+ # label remap
+ if len(gt_labels) > 0:
+ gt_labels = np.vectorize(lambda x: label_remap_dict[x])(gt_labels)
+
+ results['gt_bboxes'] = gt_bboxes
+ results['gt_bboxes_labels'] = gt_labels
+
+ results['text'] = pheso_caption
+ results['tokens_positive'] = label_to_positions
+
+ return results
+
+
+@TRANSFORMS.register_module()
+class LoadTextAnnotations(BaseTransform):
+
+ def transform(self, results: dict) -> dict:
+ if 'phrases' in results:
+ tokens_positive = [
+ phrase['tokens_positive']
+ for phrase in results['phrases'].values()
+ ]
+ results['tokens_positive'] = tokens_positive
+ else:
+ text = results['text']
+ results['text'] = list(text.values())
+ return results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/transformers_glip.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/transformers_glip.py
new file mode 100644
index 0000000000000000000000000000000000000000..60c4f87d1b86c13f886da27584114b6420b8b8cb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/transformers_glip.py
@@ -0,0 +1,66 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import mmcv
+import numpy as np
+from mmcv.transforms import BaseTransform
+
+from mmdet.registry import TRANSFORMS
+from mmdet.structures.bbox import HorizontalBoxes, autocast_box_type
+from .transforms import RandomFlip
+
+
+@TRANSFORMS.register_module()
+class GTBoxSubOne_GLIP(BaseTransform):
+ """Subtract 1 from the x2 and y2 coordinates of the gt_bboxes."""
+
+ def transform(self, results: dict) -> dict:
+ if 'gt_bboxes' in results:
+ gt_bboxes = results['gt_bboxes']
+ if isinstance(gt_bboxes, np.ndarray):
+ gt_bboxes[:, 2:] -= 1
+ results['gt_bboxes'] = gt_bboxes
+ elif isinstance(gt_bboxes, HorizontalBoxes):
+ gt_bboxes = results['gt_bboxes'].tensor
+ gt_bboxes[:, 2:] -= 1
+ results['gt_bboxes'] = HorizontalBoxes(gt_bboxes)
+ else:
+ raise NotImplementedError
+ return results
+
+
+@TRANSFORMS.register_module()
+class RandomFlip_GLIP(RandomFlip):
+ """Flip the image & bboxes & masks & segs horizontally or vertically.
+
+ When using horizontal flipping, the corresponding bbox x-coordinate needs
+ to be additionally subtracted by one.
+ """
+
+ @autocast_box_type()
+ def _flip(self, results: dict) -> None:
+ """Flip images, bounding boxes, and semantic segmentation map."""
+ # flip image
+ results['img'] = mmcv.imflip(
+ results['img'], direction=results['flip_direction'])
+
+ img_shape = results['img'].shape[:2]
+
+ # flip bboxes
+ if results.get('gt_bboxes', None) is not None:
+ results['gt_bboxes'].flip_(img_shape, results['flip_direction'])
+ # Only change this line
+ if results['flip_direction'] == 'horizontal':
+ results['gt_bboxes'].translate_([-1, 0])
+
+ # TODO: check it
+ # flip masks
+ if results.get('gt_masks', None) is not None:
+ results['gt_masks'] = results['gt_masks'].flip(
+ results['flip_direction'])
+
+ # flip segs
+ if results.get('gt_seg_map', None) is not None:
+ results['gt_seg_map'] = mmcv.imflip(
+ results['gt_seg_map'], direction=results['flip_direction'])
+
+ # record homography matrix for flip
+ self._record_homography_matrix(results)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/transforms.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/transforms.py
new file mode 100644
index 0000000000000000000000000000000000000000..c50b987db33c91f759f6c89580f605631ce4f558
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/transforms.py
@@ -0,0 +1,3856 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import inspect
+import math
+import warnings
+from typing import List, Optional, Sequence, Tuple, Union
+
+import cv2
+import mmcv
+import numpy as np
+from mmcv.image import imresize
+from mmcv.image.geometric import _scale_size
+from mmcv.transforms import BaseTransform
+from mmcv.transforms import Pad as MMCV_Pad
+from mmcv.transforms import RandomFlip as MMCV_RandomFlip
+from mmcv.transforms import Resize as MMCV_Resize
+from mmcv.transforms.utils import avoid_cache_randomness, cache_randomness
+from mmengine.dataset import BaseDataset
+from mmengine.utils import is_str
+from numpy import random
+
+from mmdet.registry import TRANSFORMS
+from mmdet.structures.bbox import HorizontalBoxes, autocast_box_type
+from mmdet.structures.mask import BitmapMasks, PolygonMasks
+from mmdet.utils import log_img_scale
+
+try:
+ from imagecorruptions import corrupt
+except ImportError:
+ corrupt = None
+
+try:
+ import albumentations
+ from albumentations import Compose
+except ImportError:
+ albumentations = None
+ Compose = None
+
+Number = Union[int, float]
+
+
+def _fixed_scale_size(
+ size: Tuple[int, int],
+ scale: Union[float, int, tuple],
+) -> Tuple[int, int]:
+ """Rescale a size by a ratio.
+
+ Args:
+ size (tuple[int]): (w, h).
+ scale (float | tuple(float)): Scaling factor.
+
+ Returns:
+ tuple[int]: scaled size.
+ """
+ if isinstance(scale, (float, int)):
+ scale = (scale, scale)
+ w, h = size
+ # don't need o.5 offset
+ return int(w * float(scale[0])), int(h * float(scale[1]))
+
+
+def rescale_size(old_size: tuple,
+ scale: Union[float, int, tuple],
+ return_scale: bool = False) -> tuple:
+ """Calculate the new size to be rescaled to.
+
+ Args:
+ old_size (tuple[int]): The old size (w, h) of image.
+ scale (float | tuple[int]): The scaling factor or maximum size.
+ If it is a float number, then the image will be rescaled by this
+ factor, else if it is a tuple of 2 integers, then the image will
+ be rescaled as large as possible within the scale.
+ return_scale (bool): Whether to return the scaling factor besides the
+ rescaled image size.
+
+ Returns:
+ tuple[int]: The new rescaled image size.
+ """
+ w, h = old_size
+ if isinstance(scale, (float, int)):
+ if scale <= 0:
+ raise ValueError(f'Invalid scale {scale}, must be positive.')
+ scale_factor = scale
+ elif isinstance(scale, tuple):
+ max_long_edge = max(scale)
+ max_short_edge = min(scale)
+ scale_factor = min(max_long_edge / max(h, w),
+ max_short_edge / min(h, w))
+ else:
+ raise TypeError(
+ f'Scale must be a number or tuple of int, but got {type(scale)}')
+ # only change this
+ new_size = _fixed_scale_size((w, h), scale_factor)
+
+ if return_scale:
+ return new_size, scale_factor
+ else:
+ return new_size
+
+
+def imrescale(
+ img: np.ndarray,
+ scale: Union[float, Tuple[int, int]],
+ return_scale: bool = False,
+ interpolation: str = 'bilinear',
+ backend: Optional[str] = None
+) -> Union[np.ndarray, Tuple[np.ndarray, float]]:
+ """Resize image while keeping the aspect ratio.
+
+ Args:
+ img (ndarray): The input image.
+ scale (float | tuple[int]): The scaling factor or maximum size.
+ If it is a float number, then the image will be rescaled by this
+ factor, else if it is a tuple of 2 integers, then the image will
+ be rescaled as large as possible within the scale.
+ return_scale (bool): Whether to return the scaling factor besides the
+ rescaled image.
+ interpolation (str): Same as :func:`resize`.
+ backend (str | None): Same as :func:`resize`.
+
+ Returns:
+ ndarray: The rescaled image.
+ """
+ h, w = img.shape[:2]
+ new_size, scale_factor = rescale_size((w, h), scale, return_scale=True)
+ rescaled_img = imresize(
+ img, new_size, interpolation=interpolation, backend=backend)
+ if return_scale:
+ return rescaled_img, scale_factor
+ else:
+ return rescaled_img
+
+
+@TRANSFORMS.register_module()
+class Resize(MMCV_Resize):
+ """Resize images & bbox & seg.
+
+ This transform resizes the input image according to ``scale`` or
+ ``scale_factor``. Bboxes, masks, and seg map are then resized
+ with the same scale factor.
+ if ``scale`` and ``scale_factor`` are both set, it will use ``scale`` to
+ resize.
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_masks (BitmapMasks | PolygonMasks) (optional)
+ - gt_seg_map (np.uint8) (optional)
+
+ Modified Keys:
+
+ - img
+ - img_shape
+ - gt_bboxes
+ - gt_masks
+ - gt_seg_map
+
+
+ Added Keys:
+
+ - scale
+ - scale_factor
+ - keep_ratio
+ - homography_matrix
+
+ Args:
+ scale (int or tuple): Images scales for resizing. Defaults to None
+ scale_factor (float or tuple[float]): Scale factors for resizing.
+ Defaults to None.
+ keep_ratio (bool): Whether to keep the aspect ratio when resizing the
+ image. Defaults to False.
+ clip_object_border (bool): Whether to clip the objects
+ outside the border of the image. In some dataset like MOT17, the gt
+ bboxes are allowed to cross the border of images. Therefore, we
+ don't need to clip the gt bboxes in these cases. Defaults to True.
+ backend (str): Image resize backend, choices are 'cv2' and 'pillow'.
+ These two backends generates slightly different results. Defaults
+ to 'cv2'.
+ interpolation (str): Interpolation method, accepted values are
+ "nearest", "bilinear", "bicubic", "area", "lanczos" for 'cv2'
+ backend, "nearest", "bilinear" for 'pillow' backend. Defaults
+ to 'bilinear'.
+ """
+
+ def _resize_masks(self, results: dict) -> None:
+ """Resize masks with ``results['scale']``"""
+ if results.get('gt_masks', None) is not None:
+ if self.keep_ratio:
+ results['gt_masks'] = results['gt_masks'].rescale(
+ results['scale'])
+ else:
+ results['gt_masks'] = results['gt_masks'].resize(
+ results['img_shape'])
+
+ def _resize_bboxes(self, results: dict) -> None:
+ """Resize bounding boxes with ``results['scale_factor']``."""
+ if results.get('gt_bboxes', None) is not None:
+ results['gt_bboxes'].rescale_(results['scale_factor'])
+ if self.clip_object_border:
+ results['gt_bboxes'].clip_(results['img_shape'])
+
+ def _record_homography_matrix(self, results: dict) -> None:
+ """Record the homography matrix for the Resize."""
+ w_scale, h_scale = results['scale_factor']
+ homography_matrix = np.array(
+ [[w_scale, 0, 0], [0, h_scale, 0], [0, 0, 1]], dtype=np.float32)
+ if results.get('homography_matrix', None) is None:
+ results['homography_matrix'] = homography_matrix
+ else:
+ results['homography_matrix'] = homography_matrix @ results[
+ 'homography_matrix']
+
+ @autocast_box_type()
+ def transform(self, results: dict) -> dict:
+ """Transform function to resize images, bounding boxes and semantic
+ segmentation map.
+
+ Args:
+ results (dict): Result dict from loading pipeline.
+ Returns:
+ dict: Resized results, 'img', 'gt_bboxes', 'gt_seg_map',
+ 'scale', 'scale_factor', 'height', 'width', and 'keep_ratio' keys
+ are updated in result dict.
+ """
+ if self.scale:
+ results['scale'] = self.scale
+ else:
+ img_shape = results['img'].shape[:2]
+ results['scale'] = _scale_size(img_shape[::-1], self.scale_factor)
+ self._resize_img(results)
+ self._resize_bboxes(results)
+ self._resize_masks(results)
+ self._resize_seg(results)
+ self._record_homography_matrix(results)
+ return results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'(scale={self.scale}, '
+ repr_str += f'scale_factor={self.scale_factor}, '
+ repr_str += f'keep_ratio={self.keep_ratio}, '
+ repr_str += f'clip_object_border={self.clip_object_border}), '
+ repr_str += f'backend={self.backend}), '
+ repr_str += f'interpolation={self.interpolation})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class FixScaleResize(Resize):
+ """Compared to Resize, FixScaleResize fixes the scaling issue when
+ `keep_ratio=true`."""
+
+ def _resize_img(self, results):
+ """Resize images with ``results['scale']``."""
+ if results.get('img', None) is not None:
+ if self.keep_ratio:
+ img, scale_factor = imrescale(
+ results['img'],
+ results['scale'],
+ interpolation=self.interpolation,
+ return_scale=True,
+ backend=self.backend)
+ new_h, new_w = img.shape[:2]
+ h, w = results['img'].shape[:2]
+ w_scale = new_w / w
+ h_scale = new_h / h
+ else:
+ img, w_scale, h_scale = mmcv.imresize(
+ results['img'],
+ results['scale'],
+ interpolation=self.interpolation,
+ return_scale=True,
+ backend=self.backend)
+ results['img'] = img
+ results['img_shape'] = img.shape[:2]
+ results['scale_factor'] = (w_scale, h_scale)
+ results['keep_ratio'] = self.keep_ratio
+
+
+@TRANSFORMS.register_module()
+class ResizeShortestEdge(BaseTransform):
+ """Resize the image and mask while keeping the aspect ratio unchanged.
+
+ Modified from https://github.com/facebookresearch/detectron2/blob/main/detectron2/data/transforms/augmentation_impl.py#L130 # noqa:E501
+
+ This transform attempts to scale the shorter edge to the given
+ `scale`, as long as the longer edge does not exceed `max_size`.
+ If `max_size` is reached, then downscale so that the longer
+ edge does not exceed `max_size`.
+
+ Required Keys:
+ - img
+ - gt_seg_map (optional)
+ Modified Keys:
+ - img
+ - img_shape
+ - gt_seg_map (optional))
+ Added Keys:
+ - scale
+ - scale_factor
+ - keep_ratio
+
+ Args:
+ scale (Union[int, Tuple[int, int]]): The target short edge length.
+ If it's tuple, will select the min value as the short edge length.
+ max_size (int): The maximum allowed longest edge length.
+ """
+
+ def __init__(self,
+ scale: Union[int, Tuple[int, int]],
+ max_size: Optional[int] = None,
+ resize_type: str = 'Resize',
+ **resize_kwargs) -> None:
+ super().__init__()
+ self.scale = scale
+ self.max_size = max_size
+
+ self.resize_cfg = dict(type=resize_type, **resize_kwargs)
+ self.resize = TRANSFORMS.build({'scale': 0, **self.resize_cfg})
+
+ def _get_output_shape(
+ self, img: np.ndarray,
+ short_edge_length: Union[int, Tuple[int, int]]) -> Tuple[int, int]:
+ """Compute the target image shape with the given `short_edge_length`.
+
+ Args:
+ img (np.ndarray): The input image.
+ short_edge_length (Union[int, Tuple[int, int]]): The target short
+ edge length. If it's tuple, will select the min value as the
+ short edge length.
+ """
+ h, w = img.shape[:2]
+ if isinstance(short_edge_length, int):
+ size = short_edge_length * 1.0
+ elif isinstance(short_edge_length, tuple):
+ size = min(short_edge_length) * 1.0
+ scale = size / min(h, w)
+ if h < w:
+ new_h, new_w = size, scale * w
+ else:
+ new_h, new_w = scale * h, size
+
+ if self.max_size and max(new_h, new_w) > self.max_size:
+ scale = self.max_size * 1.0 / max(new_h, new_w)
+ new_h *= scale
+ new_w *= scale
+
+ new_h = int(new_h + 0.5)
+ new_w = int(new_w + 0.5)
+ return new_w, new_h
+
+ def transform(self, results: dict) -> dict:
+ self.resize.scale = self._get_output_shape(results['img'], self.scale)
+ return self.resize(results)
+
+
+@TRANSFORMS.register_module()
+class FixShapeResize(Resize):
+ """Resize images & bbox & seg to the specified size.
+
+ This transform resizes the input image according to ``width`` and
+ ``height``. Bboxes, masks, and seg map are then resized
+ with the same parameters.
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_masks (BitmapMasks | PolygonMasks) (optional)
+ - gt_seg_map (np.uint8) (optional)
+
+ Modified Keys:
+
+ - img
+ - img_shape
+ - gt_bboxes
+ - gt_masks
+ - gt_seg_map
+
+
+ Added Keys:
+
+ - scale
+ - scale_factor
+ - keep_ratio
+ - homography_matrix
+
+ Args:
+ width (int): width for resizing.
+ height (int): height for resizing.
+ Defaults to None.
+ pad_val (Number | dict[str, Number], optional): Padding value for if
+ the pad_mode is "constant". If it is a single number, the value
+ to pad the image is the number and to pad the semantic
+ segmentation map is 255. If it is a dict, it should have the
+ following keys:
+
+ - img: The value to pad the image.
+ - seg: The value to pad the semantic segmentation map.
+ Defaults to dict(img=0, seg=255).
+ keep_ratio (bool): Whether to keep the aspect ratio when resizing the
+ image. Defaults to False.
+ clip_object_border (bool): Whether to clip the objects
+ outside the border of the image. In some dataset like MOT17, the gt
+ bboxes are allowed to cross the border of images. Therefore, we
+ don't need to clip the gt bboxes in these cases. Defaults to True.
+ backend (str): Image resize backend, choices are 'cv2' and 'pillow'.
+ These two backends generates slightly different results. Defaults
+ to 'cv2'.
+ interpolation (str): Interpolation method, accepted values are
+ "nearest", "bilinear", "bicubic", "area", "lanczos" for 'cv2'
+ backend, "nearest", "bilinear" for 'pillow' backend. Defaults
+ to 'bilinear'.
+ """
+
+ def __init__(self,
+ width: int,
+ height: int,
+ pad_val: Union[Number, dict] = dict(img=0, seg=255),
+ keep_ratio: bool = False,
+ clip_object_border: bool = True,
+ backend: str = 'cv2',
+ interpolation: str = 'bilinear') -> None:
+ assert width is not None and height is not None, (
+ '`width` and'
+ '`height` can not be `None`')
+
+ self.width = width
+ self.height = height
+ self.scale = (width, height)
+
+ self.backend = backend
+ self.interpolation = interpolation
+ self.keep_ratio = keep_ratio
+ self.clip_object_border = clip_object_border
+
+ if keep_ratio is True:
+ # padding to the fixed size when keep_ratio=True
+ self.pad_transform = Pad(size=self.scale, pad_val=pad_val)
+
+ @autocast_box_type()
+ def transform(self, results: dict) -> dict:
+ """Transform function to resize images, bounding boxes and semantic
+ segmentation map.
+
+ Args:
+ results (dict): Result dict from loading pipeline.
+ Returns:
+ dict: Resized results, 'img', 'gt_bboxes', 'gt_seg_map',
+ 'scale', 'scale_factor', 'height', 'width', and 'keep_ratio' keys
+ are updated in result dict.
+ """
+ img = results['img']
+ h, w = img.shape[:2]
+ if self.keep_ratio:
+ scale_factor = min(self.width / w, self.height / h)
+ results['scale_factor'] = (scale_factor, scale_factor)
+ real_w, real_h = int(w * float(scale_factor) +
+ 0.5), int(h * float(scale_factor) + 0.5)
+ img, scale_factor = mmcv.imrescale(
+ results['img'], (real_w, real_h),
+ interpolation=self.interpolation,
+ return_scale=True,
+ backend=self.backend)
+ # the w_scale and h_scale has minor difference
+ # a real fix should be done in the mmcv.imrescale in the future
+ results['img'] = img
+ results['img_shape'] = img.shape[:2]
+ results['keep_ratio'] = self.keep_ratio
+ results['scale'] = (real_w, real_h)
+ else:
+ results['scale'] = (self.width, self.height)
+ results['scale_factor'] = (self.width / w, self.height / h)
+ super()._resize_img(results)
+
+ self._resize_bboxes(results)
+ self._resize_masks(results)
+ self._resize_seg(results)
+ self._record_homography_matrix(results)
+ if self.keep_ratio:
+ self.pad_transform(results)
+ return results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'(width={self.width}, height={self.height}, '
+ repr_str += f'keep_ratio={self.keep_ratio}, '
+ repr_str += f'clip_object_border={self.clip_object_border}), '
+ repr_str += f'backend={self.backend}), '
+ repr_str += f'interpolation={self.interpolation})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class RandomFlip(MMCV_RandomFlip):
+ """Flip the image & bbox & mask & segmentation map. Added or Updated keys:
+ flip, flip_direction, img, gt_bboxes, and gt_seg_map. There are 3 flip
+ modes:
+
+ - ``prob`` is float, ``direction`` is string: the image will be
+ ``direction``ly flipped with probability of ``prob`` .
+ E.g., ``prob=0.5``, ``direction='horizontal'``,
+ then image will be horizontally flipped with probability of 0.5.
+ - ``prob`` is float, ``direction`` is list of string: the image will
+ be ``direction[i]``ly flipped with probability of
+ ``prob/len(direction)``.
+ E.g., ``prob=0.5``, ``direction=['horizontal', 'vertical']``,
+ then image will be horizontally flipped with probability of 0.25,
+ vertically with probability of 0.25.
+ - ``prob`` is list of float, ``direction`` is list of string:
+ given ``len(prob) == len(direction)``, the image will
+ be ``direction[i]``ly flipped with probability of ``prob[i]``.
+ E.g., ``prob=[0.3, 0.5]``, ``direction=['horizontal',
+ 'vertical']``, then image will be horizontally flipped with
+ probability of 0.3, vertically with probability of 0.5.
+
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_masks (BitmapMasks | PolygonMasks) (optional)
+ - gt_seg_map (np.uint8) (optional)
+
+ Modified Keys:
+
+ - img
+ - gt_bboxes
+ - gt_masks
+ - gt_seg_map
+
+ Added Keys:
+
+ - flip
+ - flip_direction
+ - homography_matrix
+
+
+ Args:
+ prob (float | list[float], optional): The flipping probability.
+ Defaults to None.
+ direction(str | list[str]): The flipping direction. Options
+ If input is a list, the length must equal ``prob``. Each
+ element in ``prob`` indicates the flip probability of
+ corresponding direction. Defaults to 'horizontal'.
+ """
+
+ def _record_homography_matrix(self, results: dict) -> None:
+ """Record the homography matrix for the RandomFlip."""
+ cur_dir = results['flip_direction']
+ h, w = results['img'].shape[:2]
+
+ if cur_dir == 'horizontal':
+ homography_matrix = np.array([[-1, 0, w], [0, 1, 0], [0, 0, 1]],
+ dtype=np.float32)
+ elif cur_dir == 'vertical':
+ homography_matrix = np.array([[1, 0, 0], [0, -1, h], [0, 0, 1]],
+ dtype=np.float32)
+ elif cur_dir == 'diagonal':
+ homography_matrix = np.array([[-1, 0, w], [0, -1, h], [0, 0, 1]],
+ dtype=np.float32)
+ else:
+ homography_matrix = np.eye(3, dtype=np.float32)
+
+ if results.get('homography_matrix', None) is None:
+ results['homography_matrix'] = homography_matrix
+ else:
+ results['homography_matrix'] = homography_matrix @ results[
+ 'homography_matrix']
+
+ @autocast_box_type()
+ def _flip(self, results: dict) -> None:
+ """Flip images, bounding boxes, and semantic segmentation map."""
+ # flip image
+ results['img'] = mmcv.imflip(
+ results['img'], direction=results['flip_direction'])
+
+ img_shape = results['img'].shape[:2]
+
+ # flip bboxes
+ if results.get('gt_bboxes', None) is not None:
+ results['gt_bboxes'].flip_(img_shape, results['flip_direction'])
+
+ # flip masks
+ if results.get('gt_masks', None) is not None:
+ results['gt_masks'] = results['gt_masks'].flip(
+ results['flip_direction'])
+
+ # flip segs
+ if results.get('gt_seg_map', None) is not None:
+ results['gt_seg_map'] = mmcv.imflip(
+ results['gt_seg_map'], direction=results['flip_direction'])
+
+ # record homography matrix for flip
+ self._record_homography_matrix(results)
+
+
+@TRANSFORMS.register_module()
+class RandomShift(BaseTransform):
+ """Shift the image and box given shift pixels and probability.
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (BaseBoxes[torch.float32])
+ - gt_bboxes_labels (np.int64)
+ - gt_ignore_flags (bool) (optional)
+
+ Modified Keys:
+
+ - img
+ - gt_bboxes
+ - gt_bboxes_labels
+ - gt_ignore_flags (bool) (optional)
+
+ Args:
+ prob (float): Probability of shifts. Defaults to 0.5.
+ max_shift_px (int): The max pixels for shifting. Defaults to 32.
+ filter_thr_px (int): The width and height threshold for filtering.
+ The bbox and the rest of the targets below the width and
+ height threshold will be filtered. Defaults to 1.
+ """
+
+ def __init__(self,
+ prob: float = 0.5,
+ max_shift_px: int = 32,
+ filter_thr_px: int = 1) -> None:
+ assert 0 <= prob <= 1
+ assert max_shift_px >= 0
+ self.prob = prob
+ self.max_shift_px = max_shift_px
+ self.filter_thr_px = int(filter_thr_px)
+
+ @cache_randomness
+ def _random_prob(self) -> float:
+ return random.uniform(0, 1)
+
+ @autocast_box_type()
+ def transform(self, results: dict) -> dict:
+ """Transform function to random shift images, bounding boxes.
+
+ Args:
+ results (dict): Result dict from loading pipeline.
+
+ Returns:
+ dict: Shift results.
+ """
+ if self._random_prob() < self.prob:
+ img_shape = results['img'].shape[:2]
+
+ random_shift_x = random.randint(-self.max_shift_px,
+ self.max_shift_px)
+ random_shift_y = random.randint(-self.max_shift_px,
+ self.max_shift_px)
+ new_x = max(0, random_shift_x)
+ ori_x = max(0, -random_shift_x)
+ new_y = max(0, random_shift_y)
+ ori_y = max(0, -random_shift_y)
+
+ # TODO: support mask and semantic segmentation maps.
+ bboxes = results['gt_bboxes'].clone()
+ bboxes.translate_([random_shift_x, random_shift_y])
+
+ # clip border
+ bboxes.clip_(img_shape)
+
+ # remove invalid bboxes
+ valid_inds = (bboxes.widths > self.filter_thr_px).numpy() & (
+ bboxes.heights > self.filter_thr_px).numpy()
+ # If the shift does not contain any gt-bbox area, skip this
+ # image.
+ if not valid_inds.any():
+ return results
+ bboxes = bboxes[valid_inds]
+ results['gt_bboxes'] = bboxes
+ results['gt_bboxes_labels'] = results['gt_bboxes_labels'][
+ valid_inds]
+
+ if results.get('gt_ignore_flags', None) is not None:
+ results['gt_ignore_flags'] = \
+ results['gt_ignore_flags'][valid_inds]
+
+ # shift img
+ img = results['img']
+ new_img = np.zeros_like(img)
+ img_h, img_w = img.shape[:2]
+ new_h = img_h - np.abs(random_shift_y)
+ new_w = img_w - np.abs(random_shift_x)
+ new_img[new_y:new_y + new_h, new_x:new_x + new_w] \
+ = img[ori_y:ori_y + new_h, ori_x:ori_x + new_w]
+ results['img'] = new_img
+
+ return results
+
+ def __repr__(self):
+ repr_str = self.__class__.__name__
+ repr_str += f'(prob={self.prob}, '
+ repr_str += f'max_shift_px={self.max_shift_px}, '
+ repr_str += f'filter_thr_px={self.filter_thr_px})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class Pad(MMCV_Pad):
+ """Pad the image & segmentation map.
+
+ There are three padding modes: (1) pad to a fixed size and (2) pad to the
+ minimum size that is divisible by some number. and (3)pad to square. Also,
+ pad to square and pad to the minimum size can be used as the same time.
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_masks (BitmapMasks | PolygonMasks) (optional)
+ - gt_seg_map (np.uint8) (optional)
+
+ Modified Keys:
+
+ - img
+ - img_shape
+ - gt_masks
+ - gt_seg_map
+
+ Added Keys:
+
+ - pad_shape
+ - pad_fixed_size
+ - pad_size_divisor
+
+ Args:
+ size (tuple, optional): Fixed padding size.
+ Expected padding shape (width, height). Defaults to None.
+ size_divisor (int, optional): The divisor of padded size. Defaults to
+ None.
+ pad_to_square (bool): Whether to pad the image into a square.
+ Currently only used for YOLOX. Defaults to False.
+ pad_val (Number | dict[str, Number], optional) - Padding value for if
+ the pad_mode is "constant". If it is a single number, the value
+ to pad the image is the number and to pad the semantic
+ segmentation map is 255. If it is a dict, it should have the
+ following keys:
+
+ - img: The value to pad the image.
+ - seg: The value to pad the semantic segmentation map.
+ Defaults to dict(img=0, seg=255).
+ padding_mode (str): Type of padding. Should be: constant, edge,
+ reflect or symmetric. Defaults to 'constant'.
+
+ - constant: pads with a constant value, this value is specified
+ with pad_val.
+ - edge: pads with the last value at the edge of the image.
+ - reflect: pads with reflection of image without repeating the last
+ value on the edge. For example, padding [1, 2, 3, 4] with 2
+ elements on both sides in reflect mode will result in
+ [3, 2, 1, 2, 3, 4, 3, 2].
+ - symmetric: pads with reflection of image repeating the last value
+ on the edge. For example, padding [1, 2, 3, 4] with 2 elements on
+ both sides in symmetric mode will result in
+ [2, 1, 1, 2, 3, 4, 4, 3]
+ """
+
+ def _pad_masks(self, results: dict) -> None:
+ """Pad masks according to ``results['pad_shape']``."""
+ if results.get('gt_masks', None) is not None:
+ pad_val = self.pad_val.get('masks', 0)
+ pad_shape = results['pad_shape'][:2]
+ results['gt_masks'] = results['gt_masks'].pad(
+ pad_shape, pad_val=pad_val)
+
+ def transform(self, results: dict) -> dict:
+ """Call function to pad images, masks, semantic segmentation maps.
+
+ Args:
+ results (dict): Result dict from loading pipeline.
+
+ Returns:
+ dict: Updated result dict.
+ """
+ self._pad_img(results)
+ self._pad_seg(results)
+ self._pad_masks(results)
+ return results
+
+
+@TRANSFORMS.register_module()
+class RandomCrop(BaseTransform):
+ """Random crop the image & bboxes & masks.
+
+ The absolute ``crop_size`` is sampled based on ``crop_type`` and
+ ``image_size``, then the cropped results are generated.
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_bboxes_labels (np.int64) (optional)
+ - gt_masks (BitmapMasks | PolygonMasks) (optional)
+ - gt_ignore_flags (bool) (optional)
+ - gt_seg_map (np.uint8) (optional)
+
+ Modified Keys:
+
+ - img
+ - img_shape
+ - gt_bboxes (optional)
+ - gt_bboxes_labels (optional)
+ - gt_masks (optional)
+ - gt_ignore_flags (optional)
+ - gt_seg_map (optional)
+ - gt_instances_ids (options, only used in MOT/VIS)
+
+ Added Keys:
+
+ - homography_matrix
+
+ Args:
+ crop_size (tuple): The relative ratio or absolute pixels of
+ (width, height).
+ crop_type (str, optional): One of "relative_range", "relative",
+ "absolute", "absolute_range". "relative" randomly crops
+ (h * crop_size[0], w * crop_size[1]) part from an input of size
+ (h, w). "relative_range" uniformly samples relative crop size from
+ range [crop_size[0], 1] and [crop_size[1], 1] for height and width
+ respectively. "absolute" crops from an input with absolute size
+ (crop_size[0], crop_size[1]). "absolute_range" uniformly samples
+ crop_h in range [crop_size[0], min(h, crop_size[1])] and crop_w
+ in range [crop_size[0], min(w, crop_size[1])].
+ Defaults to "absolute".
+ allow_negative_crop (bool, optional): Whether to allow a crop that does
+ not contain any bbox area. Defaults to False.
+ recompute_bbox (bool, optional): Whether to re-compute the boxes based
+ on cropped instance masks. Defaults to False.
+ bbox_clip_border (bool, optional): Whether clip the objects outside
+ the border of the image. Defaults to True.
+
+ Note:
+ - If the image is smaller than the absolute crop size, return the
+ original image.
+ - The keys for bboxes, labels and masks must be aligned. That is,
+ ``gt_bboxes`` corresponds to ``gt_labels`` and ``gt_masks``, and
+ ``gt_bboxes_ignore`` corresponds to ``gt_labels_ignore`` and
+ ``gt_masks_ignore``.
+ - If the crop does not contain any gt-bbox region and
+ ``allow_negative_crop`` is set to False, skip this image.
+ """
+
+ def __init__(self,
+ crop_size: tuple,
+ crop_type: str = 'absolute',
+ allow_negative_crop: bool = False,
+ recompute_bbox: bool = False,
+ bbox_clip_border: bool = True) -> None:
+ if crop_type not in [
+ 'relative_range', 'relative', 'absolute', 'absolute_range'
+ ]:
+ raise ValueError(f'Invalid crop_type {crop_type}.')
+ if crop_type in ['absolute', 'absolute_range']:
+ assert crop_size[0] > 0 and crop_size[1] > 0
+ assert isinstance(crop_size[0], int) and isinstance(
+ crop_size[1], int)
+ if crop_type == 'absolute_range':
+ assert crop_size[0] <= crop_size[1]
+ else:
+ assert 0 < crop_size[0] <= 1 and 0 < crop_size[1] <= 1
+ self.crop_size = crop_size
+ self.crop_type = crop_type
+ self.allow_negative_crop = allow_negative_crop
+ self.bbox_clip_border = bbox_clip_border
+ self.recompute_bbox = recompute_bbox
+
+ def _crop_data(self, results: dict, crop_size: Tuple[int, int],
+ allow_negative_crop: bool) -> Union[dict, None]:
+ """Function to randomly crop images, bounding boxes, masks, semantic
+ segmentation maps.
+
+ Args:
+ results (dict): Result dict from loading pipeline.
+ crop_size (Tuple[int, int]): Expected absolute size after
+ cropping, (h, w).
+ allow_negative_crop (bool): Whether to allow a crop that does not
+ contain any bbox area.
+
+ Returns:
+ results (Union[dict, None]): Randomly cropped results, 'img_shape'
+ key in result dict is updated according to crop size. None will
+ be returned when there is no valid bbox after cropping.
+ """
+ assert crop_size[0] > 0 and crop_size[1] > 0
+ img = results['img']
+ margin_h = max(img.shape[0] - crop_size[0], 0)
+ margin_w = max(img.shape[1] - crop_size[1], 0)
+ offset_h, offset_w = self._rand_offset((margin_h, margin_w))
+ crop_y1, crop_y2 = offset_h, offset_h + crop_size[0]
+ crop_x1, crop_x2 = offset_w, offset_w + crop_size[1]
+
+ # Record the homography matrix for the RandomCrop
+ homography_matrix = np.array(
+ [[1, 0, -offset_w], [0, 1, -offset_h], [0, 0, 1]],
+ dtype=np.float32)
+ if results.get('homography_matrix', None) is None:
+ results['homography_matrix'] = homography_matrix
+ else:
+ results['homography_matrix'] = homography_matrix @ results[
+ 'homography_matrix']
+
+ # crop the image
+ img = img[crop_y1:crop_y2, crop_x1:crop_x2, ...]
+ img_shape = img.shape
+ results['img'] = img
+ results['img_shape'] = img_shape[:2]
+
+ # crop bboxes accordingly and clip to the image boundary
+ if results.get('gt_bboxes', None) is not None:
+ bboxes = results['gt_bboxes']
+ bboxes.translate_([-offset_w, -offset_h])
+ if self.bbox_clip_border:
+ bboxes.clip_(img_shape[:2])
+ valid_inds = bboxes.is_inside(img_shape[:2]).numpy()
+ # If the crop does not contain any gt-bbox area and
+ # allow_negative_crop is False, skip this image.
+ if (not valid_inds.any() and not allow_negative_crop):
+ return None
+
+ results['gt_bboxes'] = bboxes[valid_inds]
+
+ if results.get('gt_ignore_flags', None) is not None:
+ results['gt_ignore_flags'] = \
+ results['gt_ignore_flags'][valid_inds]
+
+ if results.get('gt_bboxes_labels', None) is not None:
+ results['gt_bboxes_labels'] = \
+ results['gt_bboxes_labels'][valid_inds]
+
+ if results.get('gt_masks', None) is not None:
+ results['gt_masks'] = results['gt_masks'][
+ valid_inds.nonzero()[0]].crop(
+ np.asarray([crop_x1, crop_y1, crop_x2, crop_y2]))
+ if self.recompute_bbox:
+ results['gt_bboxes'] = results['gt_masks'].get_bboxes(
+ type(results['gt_bboxes']))
+
+ # We should remove the instance ids corresponding to invalid boxes.
+ if results.get('gt_instances_ids', None) is not None:
+ results['gt_instances_ids'] = \
+ results['gt_instances_ids'][valid_inds]
+
+ # crop semantic seg
+ if results.get('gt_seg_map', None) is not None:
+ results['gt_seg_map'] = results['gt_seg_map'][crop_y1:crop_y2,
+ crop_x1:crop_x2]
+
+ return results
+
+ @cache_randomness
+ def _rand_offset(self, margin: Tuple[int, int]) -> Tuple[int, int]:
+ """Randomly generate crop offset.
+
+ Args:
+ margin (Tuple[int, int]): The upper bound for the offset generated
+ randomly.
+
+ Returns:
+ Tuple[int, int]: The random offset for the crop.
+ """
+ margin_h, margin_w = margin
+ offset_h = np.random.randint(0, margin_h + 1)
+ offset_w = np.random.randint(0, margin_w + 1)
+
+ return offset_h, offset_w
+
+ @cache_randomness
+ def _get_crop_size(self, image_size: Tuple[int, int]) -> Tuple[int, int]:
+ """Randomly generates the absolute crop size based on `crop_type` and
+ `image_size`.
+
+ Args:
+ image_size (Tuple[int, int]): (h, w).
+
+ Returns:
+ crop_size (Tuple[int, int]): (crop_h, crop_w) in absolute pixels.
+ """
+ h, w = image_size
+ if self.crop_type == 'absolute':
+ return min(self.crop_size[1], h), min(self.crop_size[0], w)
+ elif self.crop_type == 'absolute_range':
+ crop_h = np.random.randint(
+ min(h, self.crop_size[0]),
+ min(h, self.crop_size[1]) + 1)
+ crop_w = np.random.randint(
+ min(w, self.crop_size[0]),
+ min(w, self.crop_size[1]) + 1)
+ return crop_h, crop_w
+ elif self.crop_type == 'relative':
+ crop_w, crop_h = self.crop_size
+ return int(h * crop_h + 0.5), int(w * crop_w + 0.5)
+ else:
+ # 'relative_range'
+ crop_size = np.asarray(self.crop_size, dtype=np.float32)
+ crop_h, crop_w = crop_size + np.random.rand(2) * (1 - crop_size)
+ return int(h * crop_h + 0.5), int(w * crop_w + 0.5)
+
+ @autocast_box_type()
+ def transform(self, results: dict) -> Union[dict, None]:
+ """Transform function to randomly crop images, bounding boxes, masks,
+ semantic segmentation maps.
+
+ Args:
+ results (dict): Result dict from loading pipeline.
+
+ Returns:
+ results (Union[dict, None]): Randomly cropped results, 'img_shape'
+ key in result dict is updated according to crop size. None will
+ be returned when there is no valid bbox after cropping.
+ """
+ image_size = results['img'].shape[:2]
+ crop_size = self._get_crop_size(image_size)
+ results = self._crop_data(results, crop_size, self.allow_negative_crop)
+ return results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'(crop_size={self.crop_size}, '
+ repr_str += f'crop_type={self.crop_type}, '
+ repr_str += f'allow_negative_crop={self.allow_negative_crop}, '
+ repr_str += f'recompute_bbox={self.recompute_bbox}, '
+ repr_str += f'bbox_clip_border={self.bbox_clip_border})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class SegRescale(BaseTransform):
+ """Rescale semantic segmentation maps.
+
+ This transform rescale the ``gt_seg_map`` according to ``scale_factor``.
+
+ Required Keys:
+
+ - gt_seg_map
+
+ Modified Keys:
+
+ - gt_seg_map
+
+ Args:
+ scale_factor (float): The scale factor of the final output. Defaults
+ to 1.
+ backend (str): Image rescale backend, choices are 'cv2' and 'pillow'.
+ These two backends generates slightly different results. Defaults
+ to 'cv2'.
+ """
+
+ def __init__(self, scale_factor: float = 1, backend: str = 'cv2') -> None:
+ self.scale_factor = scale_factor
+ self.backend = backend
+
+ def transform(self, results: dict) -> dict:
+ """Transform function to scale the semantic segmentation map.
+
+ Args:
+ results (dict): Result dict from loading pipeline.
+
+ Returns:
+ dict: Result dict with semantic segmentation map scaled.
+ """
+ if self.scale_factor != 1:
+ results['gt_seg_map'] = mmcv.imrescale(
+ results['gt_seg_map'],
+ self.scale_factor,
+ interpolation='nearest',
+ backend=self.backend)
+
+ return results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'(scale_factor={self.scale_factor}, '
+ repr_str += f'backend={self.backend})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class PhotoMetricDistortion(BaseTransform):
+ """Apply photometric distortion to image sequentially, every transformation
+ is applied with a probability of 0.5. The position of random contrast is in
+ second or second to last.
+
+ 1. random brightness
+ 2. random contrast (mode 0)
+ 3. convert color from BGR to HSV
+ 4. random saturation
+ 5. random hue
+ 6. convert color from HSV to BGR
+ 7. random contrast (mode 1)
+ 8. randomly swap channels
+
+ Required Keys:
+
+ - img (np.uint8)
+
+ Modified Keys:
+
+ - img (np.float32)
+
+ Args:
+ brightness_delta (int): delta of brightness.
+ contrast_range (sequence): range of contrast.
+ saturation_range (sequence): range of saturation.
+ hue_delta (int): delta of hue.
+ """
+
+ def __init__(self,
+ brightness_delta: int = 32,
+ contrast_range: Sequence[Number] = (0.5, 1.5),
+ saturation_range: Sequence[Number] = (0.5, 1.5),
+ hue_delta: int = 18) -> None:
+ self.brightness_delta = brightness_delta
+ self.contrast_lower, self.contrast_upper = contrast_range
+ self.saturation_lower, self.saturation_upper = saturation_range
+ self.hue_delta = hue_delta
+
+ @cache_randomness
+ def _random_flags(self) -> Sequence[Number]:
+ mode = random.randint(2)
+ brightness_flag = random.randint(2)
+ contrast_flag = random.randint(2)
+ saturation_flag = random.randint(2)
+ hue_flag = random.randint(2)
+ swap_flag = random.randint(2)
+ delta_value = random.uniform(-self.brightness_delta,
+ self.brightness_delta)
+ alpha_value = random.uniform(self.contrast_lower, self.contrast_upper)
+ saturation_value = random.uniform(self.saturation_lower,
+ self.saturation_upper)
+ hue_value = random.uniform(-self.hue_delta, self.hue_delta)
+ swap_value = random.permutation(3)
+
+ return (mode, brightness_flag, contrast_flag, saturation_flag,
+ hue_flag, swap_flag, delta_value, alpha_value,
+ saturation_value, hue_value, swap_value)
+
+ def transform(self, results: dict) -> dict:
+ """Transform function to perform photometric distortion on images.
+
+ Args:
+ results (dict): Result dict from loading pipeline.
+
+ Returns:
+ dict: Result dict with images distorted.
+ """
+ assert 'img' in results, '`img` is not found in results'
+ img = results['img']
+ img = img.astype(np.float32)
+
+ (mode, brightness_flag, contrast_flag, saturation_flag, hue_flag,
+ swap_flag, delta_value, alpha_value, saturation_value, hue_value,
+ swap_value) = self._random_flags()
+
+ # random brightness
+ if brightness_flag:
+ img += delta_value
+
+ # mode == 0 --> do random contrast first
+ # mode == 1 --> do random contrast last
+ if mode == 1:
+ if contrast_flag:
+ img *= alpha_value
+
+ # convert color from BGR to HSV
+ img = mmcv.bgr2hsv(img)
+
+ # random saturation
+ if saturation_flag:
+ img[..., 1] *= saturation_value
+ # For image(type=float32), after convert bgr to hsv by opencv,
+ # valid saturation value range is [0, 1]
+ if saturation_value > 1:
+ img[..., 1] = img[..., 1].clip(0, 1)
+
+ # random hue
+ if hue_flag:
+ img[..., 0] += hue_value
+ img[..., 0][img[..., 0] > 360] -= 360
+ img[..., 0][img[..., 0] < 0] += 360
+
+ # convert color from HSV to BGR
+ img = mmcv.hsv2bgr(img)
+
+ # random contrast
+ if mode == 0:
+ if contrast_flag:
+ img *= alpha_value
+
+ # randomly swap channels
+ if swap_flag:
+ img = img[..., swap_value]
+
+ results['img'] = img
+ return results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'(brightness_delta={self.brightness_delta}, '
+ repr_str += 'contrast_range='
+ repr_str += f'{(self.contrast_lower, self.contrast_upper)}, '
+ repr_str += 'saturation_range='
+ repr_str += f'{(self.saturation_lower, self.saturation_upper)}, '
+ repr_str += f'hue_delta={self.hue_delta})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class Expand(BaseTransform):
+ """Random expand the image & bboxes & masks & segmentation map.
+
+ Randomly place the original image on a canvas of ``ratio`` x original image
+ size filled with mean values. The ratio is in the range of ratio_range.
+
+ Required Keys:
+
+ - img
+ - img_shape
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_masks (BitmapMasks | PolygonMasks) (optional)
+ - gt_seg_map (np.uint8) (optional)
+
+ Modified Keys:
+
+ - img
+ - img_shape
+ - gt_bboxes
+ - gt_masks
+ - gt_seg_map
+
+
+ Args:
+ mean (sequence): mean value of dataset.
+ to_rgb (bool): if need to convert the order of mean to align with RGB.
+ ratio_range (sequence)): range of expand ratio.
+ seg_ignore_label (int): label of ignore segmentation map.
+ prob (float): probability of applying this transformation
+ """
+
+ def __init__(self,
+ mean: Sequence[Number] = (0, 0, 0),
+ to_rgb: bool = True,
+ ratio_range: Sequence[Number] = (1, 4),
+ seg_ignore_label: int = None,
+ prob: float = 0.5) -> None:
+ self.to_rgb = to_rgb
+ self.ratio_range = ratio_range
+ if to_rgb:
+ self.mean = mean[::-1]
+ else:
+ self.mean = mean
+ self.min_ratio, self.max_ratio = ratio_range
+ self.seg_ignore_label = seg_ignore_label
+ self.prob = prob
+
+ @cache_randomness
+ def _random_prob(self) -> float:
+ return random.uniform(0, 1)
+
+ @cache_randomness
+ def _random_ratio(self) -> float:
+ return random.uniform(self.min_ratio, self.max_ratio)
+
+ @cache_randomness
+ def _random_left_top(self, ratio: float, h: int,
+ w: int) -> Tuple[int, int]:
+ left = int(random.uniform(0, w * ratio - w))
+ top = int(random.uniform(0, h * ratio - h))
+ return left, top
+
+ @autocast_box_type()
+ def transform(self, results: dict) -> dict:
+ """Transform function to expand images, bounding boxes, masks,
+ segmentation map.
+
+ Args:
+ results (dict): Result dict from loading pipeline.
+
+ Returns:
+ dict: Result dict with images, bounding boxes, masks, segmentation
+ map expanded.
+ """
+ if self._random_prob() > self.prob:
+ return results
+ assert 'img' in results, '`img` is not found in results'
+ img = results['img']
+ h, w, c = img.shape
+ ratio = self._random_ratio()
+ # speedup expand when meets large image
+ if np.all(self.mean == self.mean[0]):
+ expand_img = np.empty((int(h * ratio), int(w * ratio), c),
+ img.dtype)
+ expand_img.fill(self.mean[0])
+ else:
+ expand_img = np.full((int(h * ratio), int(w * ratio), c),
+ self.mean,
+ dtype=img.dtype)
+ left, top = self._random_left_top(ratio, h, w)
+ expand_img[top:top + h, left:left + w] = img
+ results['img'] = expand_img
+ results['img_shape'] = expand_img.shape[:2]
+
+ # expand bboxes
+ if results.get('gt_bboxes', None) is not None:
+ results['gt_bboxes'].translate_([left, top])
+
+ # expand masks
+ if results.get('gt_masks', None) is not None:
+ results['gt_masks'] = results['gt_masks'].expand(
+ int(h * ratio), int(w * ratio), top, left)
+
+ # expand segmentation map
+ if results.get('gt_seg_map', None) is not None:
+ gt_seg = results['gt_seg_map']
+ expand_gt_seg = np.full((int(h * ratio), int(w * ratio)),
+ self.seg_ignore_label,
+ dtype=gt_seg.dtype)
+ expand_gt_seg[top:top + h, left:left + w] = gt_seg
+ results['gt_seg_map'] = expand_gt_seg
+
+ return results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'(mean={self.mean}, to_rgb={self.to_rgb}, '
+ repr_str += f'ratio_range={self.ratio_range}, '
+ repr_str += f'seg_ignore_label={self.seg_ignore_label}, '
+ repr_str += f'prob={self.prob})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class MinIoURandomCrop(BaseTransform):
+ """Random crop the image & bboxes & masks & segmentation map, the cropped
+ patches have minimum IoU requirement with original image & bboxes & masks.
+
+ & segmentation map, the IoU threshold is randomly selected from min_ious.
+
+
+ Required Keys:
+
+ - img
+ - img_shape
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_bboxes_labels (np.int64) (optional)
+ - gt_masks (BitmapMasks | PolygonMasks) (optional)
+ - gt_ignore_flags (bool) (optional)
+ - gt_seg_map (np.uint8) (optional)
+
+ Modified Keys:
+
+ - img
+ - img_shape
+ - gt_bboxes
+ - gt_bboxes_labels
+ - gt_masks
+ - gt_ignore_flags
+ - gt_seg_map
+
+
+ Args:
+ min_ious (Sequence[float]): minimum IoU threshold for all intersections
+ with bounding boxes.
+ min_crop_size (float): minimum crop's size (i.e. h,w := a*h, a*w,
+ where a >= min_crop_size).
+ bbox_clip_border (bool, optional): Whether clip the objects outside
+ the border of the image. Defaults to True.
+ """
+
+ def __init__(self,
+ min_ious: Sequence[float] = (0.1, 0.3, 0.5, 0.7, 0.9),
+ min_crop_size: float = 0.3,
+ bbox_clip_border: bool = True) -> None:
+
+ self.min_ious = min_ious
+ self.sample_mode = (1, *min_ious, 0)
+ self.min_crop_size = min_crop_size
+ self.bbox_clip_border = bbox_clip_border
+
+ @cache_randomness
+ def _random_mode(self) -> Number:
+ return random.choice(self.sample_mode)
+
+ @autocast_box_type()
+ def transform(self, results: dict) -> dict:
+ """Transform function to crop images and bounding boxes with minimum
+ IoU constraint.
+
+ Args:
+ results (dict): Result dict from loading pipeline.
+
+ Returns:
+ dict: Result dict with images and bounding boxes cropped, \
+ 'img_shape' key is updated.
+ """
+ assert 'img' in results, '`img` is not found in results'
+ assert 'gt_bboxes' in results, '`gt_bboxes` is not found in results'
+ img = results['img']
+ boxes = results['gt_bboxes']
+ h, w, c = img.shape
+ while True:
+ mode = self._random_mode()
+ self.mode = mode
+ if mode == 1:
+ return results
+
+ min_iou = self.mode
+ for i in range(50):
+ new_w = random.uniform(self.min_crop_size * w, w)
+ new_h = random.uniform(self.min_crop_size * h, h)
+
+ # h / w in [0.5, 2]
+ if new_h / new_w < 0.5 or new_h / new_w > 2:
+ continue
+
+ left = random.uniform(w - new_w)
+ top = random.uniform(h - new_h)
+
+ patch = np.array(
+ (int(left), int(top), int(left + new_w), int(top + new_h)))
+ # Line or point crop is not allowed
+ if patch[2] == patch[0] or patch[3] == patch[1]:
+ continue
+ overlaps = boxes.overlaps(
+ HorizontalBoxes(patch.reshape(-1, 4).astype(np.float32)),
+ boxes).numpy().reshape(-1)
+ if len(overlaps) > 0 and overlaps.min() < min_iou:
+ continue
+
+ # center of boxes should inside the crop img
+ # only adjust boxes and instance masks when the gt is not empty
+ if len(overlaps) > 0:
+ # adjust boxes
+ def is_center_of_bboxes_in_patch(boxes, patch):
+ centers = boxes.centers.numpy()
+ mask = ((centers[:, 0] > patch[0]) *
+ (centers[:, 1] > patch[1]) *
+ (centers[:, 0] < patch[2]) *
+ (centers[:, 1] < patch[3]))
+ return mask
+
+ mask = is_center_of_bboxes_in_patch(boxes, patch)
+ if not mask.any():
+ continue
+ if results.get('gt_bboxes', None) is not None:
+ boxes = results['gt_bboxes']
+ mask = is_center_of_bboxes_in_patch(boxes, patch)
+ boxes = boxes[mask]
+ boxes.translate_([-patch[0], -patch[1]])
+ if self.bbox_clip_border:
+ boxes.clip_(
+ [patch[3] - patch[1], patch[2] - patch[0]])
+ results['gt_bboxes'] = boxes
+
+ # ignore_flags
+ if results.get('gt_ignore_flags', None) is not None:
+ results['gt_ignore_flags'] = \
+ results['gt_ignore_flags'][mask]
+
+ # labels
+ if results.get('gt_bboxes_labels', None) is not None:
+ results['gt_bboxes_labels'] = results[
+ 'gt_bboxes_labels'][mask]
+
+ # mask fields
+ if results.get('gt_masks', None) is not None:
+ results['gt_masks'] = results['gt_masks'][
+ mask.nonzero()[0]].crop(patch)
+ # adjust the img no matter whether the gt is empty before crop
+ img = img[patch[1]:patch[3], patch[0]:patch[2]]
+ results['img'] = img
+ results['img_shape'] = img.shape[:2]
+
+ # seg fields
+ if results.get('gt_seg_map', None) is not None:
+ results['gt_seg_map'] = results['gt_seg_map'][
+ patch[1]:patch[3], patch[0]:patch[2]]
+ return results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'(min_ious={self.min_ious}, '
+ repr_str += f'min_crop_size={self.min_crop_size}, '
+ repr_str += f'bbox_clip_border={self.bbox_clip_border})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class Corrupt(BaseTransform):
+ """Corruption augmentation.
+
+ Corruption transforms implemented based on
+ `imagecorruptions `_.
+
+ Required Keys:
+
+ - img (np.uint8)
+
+
+ Modified Keys:
+
+ - img (np.uint8)
+
+
+ Args:
+ corruption (str): Corruption name.
+ severity (int): The severity of corruption. Defaults to 1.
+ """
+
+ def __init__(self, corruption: str, severity: int = 1) -> None:
+ self.corruption = corruption
+ self.severity = severity
+
+ def transform(self, results: dict) -> dict:
+ """Call function to corrupt image.
+
+ Args:
+ results (dict): Result dict from loading pipeline.
+
+ Returns:
+ dict: Result dict with images corrupted.
+ """
+
+ if corrupt is None:
+ raise RuntimeError('imagecorruptions is not installed')
+ results['img'] = corrupt(
+ results['img'].astype(np.uint8),
+ corruption_name=self.corruption,
+ severity=self.severity)
+ return results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'(corruption={self.corruption}, '
+ repr_str += f'severity={self.severity})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+@avoid_cache_randomness
+class Albu(BaseTransform):
+ """Albumentation augmentation.
+
+ Adds custom transformations from Albumentations library.
+ Please, visit `https://albumentations.readthedocs.io`
+ to get more information.
+
+ Required Keys:
+
+ - img (np.uint8)
+ - gt_bboxes (HorizontalBoxes[torch.float32]) (optional)
+ - gt_masks (BitmapMasks | PolygonMasks) (optional)
+
+ Modified Keys:
+
+ - img (np.uint8)
+ - gt_bboxes (HorizontalBoxes[torch.float32]) (optional)
+ - gt_masks (BitmapMasks | PolygonMasks) (optional)
+ - img_shape (tuple)
+
+ An example of ``transforms`` is as followed:
+
+ .. code-block::
+
+ [
+ dict(
+ type='ShiftScaleRotate',
+ shift_limit=0.0625,
+ scale_limit=0.0,
+ rotate_limit=0,
+ interpolation=1,
+ p=0.5),
+ dict(
+ type='RandomBrightnessContrast',
+ brightness_limit=[0.1, 0.3],
+ contrast_limit=[0.1, 0.3],
+ p=0.2),
+ dict(type='ChannelShuffle', p=0.1),
+ dict(
+ type='OneOf',
+ transforms=[
+ dict(type='Blur', blur_limit=3, p=1.0),
+ dict(type='MedianBlur', blur_limit=3, p=1.0)
+ ],
+ p=0.1),
+ ]
+
+ Args:
+ transforms (list[dict]): A list of albu transformations
+ bbox_params (dict, optional): Bbox_params for albumentation `Compose`
+ keymap (dict, optional): Contains
+ {'input key':'albumentation-style key'}
+ skip_img_without_anno (bool): Whether to skip the image if no ann left
+ after aug. Defaults to False.
+ """
+
+ def __init__(self,
+ transforms: List[dict],
+ bbox_params: Optional[dict] = None,
+ keymap: Optional[dict] = None,
+ skip_img_without_anno: bool = False) -> None:
+ if Compose is None:
+ raise RuntimeError('albumentations is not installed')
+
+ # Args will be modified later, copying it will be safer
+ transforms = copy.deepcopy(transforms)
+ if bbox_params is not None:
+ bbox_params = copy.deepcopy(bbox_params)
+ if keymap is not None:
+ keymap = copy.deepcopy(keymap)
+ self.transforms = transforms
+ self.filter_lost_elements = False
+ self.skip_img_without_anno = skip_img_without_anno
+
+ # A simple workaround to remove masks without boxes
+ if (isinstance(bbox_params, dict) and 'label_fields' in bbox_params
+ and 'filter_lost_elements' in bbox_params):
+ self.filter_lost_elements = True
+ self.origin_label_fields = bbox_params['label_fields']
+ bbox_params['label_fields'] = ['idx_mapper']
+ del bbox_params['filter_lost_elements']
+
+ self.bbox_params = (
+ self.albu_builder(bbox_params) if bbox_params else None)
+ self.aug = Compose([self.albu_builder(t) for t in self.transforms],
+ bbox_params=self.bbox_params)
+
+ if not keymap:
+ self.keymap_to_albu = {
+ 'img': 'image',
+ 'gt_masks': 'masks',
+ 'gt_bboxes': 'bboxes'
+ }
+ else:
+ self.keymap_to_albu = keymap
+ self.keymap_back = {v: k for k, v in self.keymap_to_albu.items()}
+
+ def albu_builder(self, cfg: dict) -> albumentations:
+ """Import a module from albumentations.
+
+ It inherits some of :func:`build_from_cfg` logic.
+
+ Args:
+ cfg (dict): Config dict. It should at least contain the key "type".
+
+ Returns:
+ obj: The constructed object.
+ """
+
+ assert isinstance(cfg, dict) and 'type' in cfg
+ args = cfg.copy()
+ obj_type = args.pop('type')
+ if is_str(obj_type):
+ if albumentations is None:
+ raise RuntimeError('albumentations is not installed')
+ obj_cls = getattr(albumentations, obj_type)
+ elif inspect.isclass(obj_type):
+ obj_cls = obj_type
+ else:
+ raise TypeError(
+ f'type must be a str or valid type, but got {type(obj_type)}')
+
+ if 'transforms' in args:
+ args['transforms'] = [
+ self.albu_builder(transform)
+ for transform in args['transforms']
+ ]
+
+ return obj_cls(**args)
+
+ @staticmethod
+ def mapper(d: dict, keymap: dict) -> dict:
+ """Dictionary mapper. Renames keys according to keymap provided.
+
+ Args:
+ d (dict): old dict
+ keymap (dict): {'old_key':'new_key'}
+ Returns:
+ dict: new dict.
+ """
+ updated_dict = {}
+ for k, v in zip(d.keys(), d.values()):
+ new_k = keymap.get(k, k)
+ updated_dict[new_k] = d[k]
+ return updated_dict
+
+ @autocast_box_type()
+ def transform(self, results: dict) -> Union[dict, None]:
+ """Transform function of Albu."""
+ # TODO: gt_seg_map is not currently supported
+ # dict to albumentations format
+ results = self.mapper(results, self.keymap_to_albu)
+ results, ori_masks = self._preprocess_results(results)
+ results = self.aug(**results)
+ results = self._postprocess_results(results, ori_masks)
+ if results is None:
+ return None
+ # back to the original format
+ results = self.mapper(results, self.keymap_back)
+ results['img_shape'] = results['img'].shape[:2]
+ return results
+
+ def _preprocess_results(self, results: dict) -> tuple:
+ """Pre-processing results to facilitate the use of Albu."""
+ if 'bboxes' in results:
+ # to list of boxes
+ if not isinstance(results['bboxes'], HorizontalBoxes):
+ raise NotImplementedError(
+ 'Albu only supports horizontal boxes now')
+ bboxes = results['bboxes'].numpy()
+ results['bboxes'] = [x for x in bboxes]
+ # add pseudo-field for filtration
+ if self.filter_lost_elements:
+ results['idx_mapper'] = np.arange(len(results['bboxes']))
+
+ # TODO: Support mask structure in albu
+ ori_masks = None
+ if 'masks' in results:
+ if isinstance(results['masks'], PolygonMasks):
+ raise NotImplementedError(
+ 'Albu only supports BitMap masks now')
+ ori_masks = results['masks']
+ if albumentations.__version__ < '0.5':
+ results['masks'] = results['masks'].masks
+ else:
+ results['masks'] = [mask for mask in results['masks'].masks]
+
+ return results, ori_masks
+
+ def _postprocess_results(
+ self,
+ results: dict,
+ ori_masks: Optional[Union[BitmapMasks,
+ PolygonMasks]] = None) -> dict:
+ """Post-processing Albu output."""
+ # albumentations may return np.array or list on different versions
+ if 'gt_bboxes_labels' in results and isinstance(
+ results['gt_bboxes_labels'], list):
+ results['gt_bboxes_labels'] = np.array(
+ results['gt_bboxes_labels'], dtype=np.int64)
+ if 'gt_ignore_flags' in results and isinstance(
+ results['gt_ignore_flags'], list):
+ results['gt_ignore_flags'] = np.array(
+ results['gt_ignore_flags'], dtype=bool)
+
+ if 'bboxes' in results:
+ if isinstance(results['bboxes'], list):
+ results['bboxes'] = np.array(
+ results['bboxes'], dtype=np.float32)
+ results['bboxes'] = results['bboxes'].reshape(-1, 4)
+ results['bboxes'] = HorizontalBoxes(results['bboxes'])
+
+ # filter label_fields
+ if self.filter_lost_elements:
+
+ for label in self.origin_label_fields:
+ results[label] = np.array(
+ [results[label][i] for i in results['idx_mapper']])
+ if 'masks' in results:
+ assert ori_masks is not None
+ results['masks'] = np.array(
+ [results['masks'][i] for i in results['idx_mapper']])
+ results['masks'] = ori_masks.__class__(
+ results['masks'],
+ results['masks'][0].shape[0],
+ results['masks'][0].shape[1],
+ )
+ if (not len(results['idx_mapper'])
+ and self.skip_img_without_anno):
+ return None
+ elif 'masks' in results:
+ results['masks'] = ori_masks.__class__(results['masks'],
+ ori_masks.height,
+ ori_masks.width)
+
+ return results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__ + f'(transforms={self.transforms})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+@avoid_cache_randomness
+class RandomCenterCropPad(BaseTransform):
+ """Random center crop and random around padding for CornerNet.
+
+ This operation generates randomly cropped image from the original image and
+ pads it simultaneously. Different from :class:`RandomCrop`, the output
+ shape may not equal to ``crop_size`` strictly. We choose a random value
+ from ``ratios`` and the output shape could be larger or smaller than
+ ``crop_size``. The padding operation is also different from :class:`Pad`,
+ here we use around padding instead of right-bottom padding.
+
+ The relation between output image (padding image) and original image:
+
+ .. code:: text
+
+ output image
+
+ +----------------------------+
+ | padded area |
+ +------|----------------------------|----------+
+ | | cropped area | |
+ | | +---------------+ | |
+ | | | . center | | | original image
+ | | | range | | |
+ | | +---------------+ | |
+ +------|----------------------------|----------+
+ | padded area |
+ +----------------------------+
+
+ There are 5 main areas in the figure:
+
+ - output image: output image of this operation, also called padding
+ image in following instruction.
+ - original image: input image of this operation.
+ - padded area: non-intersect area of output image and original image.
+ - cropped area: the overlap of output image and original image.
+ - center range: a smaller area where random center chosen from.
+ center range is computed by ``border`` and original image's shape
+ to avoid our random center is too close to original image's border.
+
+ Also this operation act differently in train and test mode, the summary
+ pipeline is listed below.
+
+ Train pipeline:
+
+ 1. Choose a ``random_ratio`` from ``ratios``, the shape of padding image
+ will be ``random_ratio * crop_size``.
+ 2. Choose a ``random_center`` in center range.
+ 3. Generate padding image with center matches the ``random_center``.
+ 4. Initialize the padding image with pixel value equals to ``mean``.
+ 5. Copy the cropped area to padding image.
+ 6. Refine annotations.
+
+ Test pipeline:
+
+ 1. Compute output shape according to ``test_pad_mode``.
+ 2. Generate padding image with center matches the original image
+ center.
+ 3. Initialize the padding image with pixel value equals to ``mean``.
+ 4. Copy the ``cropped area`` to padding image.
+
+ Required Keys:
+
+ - img (np.float32)
+ - img_shape (tuple)
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_bboxes_labels (np.int64) (optional)
+ - gt_ignore_flags (bool) (optional)
+
+ Modified Keys:
+
+ - img (np.float32)
+ - img_shape (tuple)
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_bboxes_labels (np.int64) (optional)
+ - gt_ignore_flags (bool) (optional)
+
+ Args:
+ crop_size (tuple, optional): expected size after crop, final size will
+ computed according to ratio. Requires (width, height)
+ in train mode, and None in test mode.
+ ratios (tuple, optional): random select a ratio from tuple and crop
+ image to (crop_size[0] * ratio) * (crop_size[1] * ratio).
+ Only available in train mode. Defaults to (0.9, 1.0, 1.1).
+ border (int, optional): max distance from center select area to image
+ border. Only available in train mode. Defaults to 128.
+ mean (sequence, optional): Mean values of 3 channels.
+ std (sequence, optional): Std values of 3 channels.
+ to_rgb (bool, optional): Whether to convert the image from BGR to RGB.
+ test_mode (bool): whether involve random variables in transform.
+ In train mode, crop_size is fixed, center coords and ratio is
+ random selected from predefined lists. In test mode, crop_size
+ is image's original shape, center coords and ratio is fixed.
+ Defaults to False.
+ test_pad_mode (tuple, optional): padding method and padding shape
+ value, only available in test mode. Default is using
+ 'logical_or' with 127 as padding shape value.
+
+ - 'logical_or': final_shape = input_shape | padding_shape_value
+ - 'size_divisor': final_shape = int(
+ ceil(input_shape / padding_shape_value) * padding_shape_value)
+
+ Defaults to ('logical_or', 127).
+ test_pad_add_pix (int): Extra padding pixel in test mode.
+ Defaults to 0.
+ bbox_clip_border (bool): Whether clip the objects outside
+ the border of the image. Defaults to True.
+ """
+
+ def __init__(self,
+ crop_size: Optional[tuple] = None,
+ ratios: Optional[tuple] = (0.9, 1.0, 1.1),
+ border: Optional[int] = 128,
+ mean: Optional[Sequence] = None,
+ std: Optional[Sequence] = None,
+ to_rgb: Optional[bool] = None,
+ test_mode: bool = False,
+ test_pad_mode: Optional[tuple] = ('logical_or', 127),
+ test_pad_add_pix: int = 0,
+ bbox_clip_border: bool = True) -> None:
+ if test_mode:
+ assert crop_size is None, 'crop_size must be None in test mode'
+ assert ratios is None, 'ratios must be None in test mode'
+ assert border is None, 'border must be None in test mode'
+ assert isinstance(test_pad_mode, (list, tuple))
+ assert test_pad_mode[0] in ['logical_or', 'size_divisor']
+ else:
+ assert isinstance(crop_size, (list, tuple))
+ assert crop_size[0] > 0 and crop_size[1] > 0, (
+ 'crop_size must > 0 in train mode')
+ assert isinstance(ratios, (list, tuple))
+ assert test_pad_mode is None, (
+ 'test_pad_mode must be None in train mode')
+
+ self.crop_size = crop_size
+ self.ratios = ratios
+ self.border = border
+ # We do not set default value to mean, std and to_rgb because these
+ # hyper-parameters are easy to forget but could affect the performance.
+ # Please use the same setting as Normalize for performance assurance.
+ assert mean is not None and std is not None and to_rgb is not None
+ self.to_rgb = to_rgb
+ self.input_mean = mean
+ self.input_std = std
+ if to_rgb:
+ self.mean = mean[::-1]
+ self.std = std[::-1]
+ else:
+ self.mean = mean
+ self.std = std
+ self.test_mode = test_mode
+ self.test_pad_mode = test_pad_mode
+ self.test_pad_add_pix = test_pad_add_pix
+ self.bbox_clip_border = bbox_clip_border
+
+ def _get_border(self, border, size):
+ """Get final border for the target size.
+
+ This function generates a ``final_border`` according to image's shape.
+ The area between ``final_border`` and ``size - final_border`` is the
+ ``center range``. We randomly choose center from the ``center range``
+ to avoid our random center is too close to original image's border.
+ Also ``center range`` should be larger than 0.
+
+ Args:
+ border (int): The initial border, default is 128.
+ size (int): The width or height of original image.
+ Returns:
+ int: The final border.
+ """
+ k = 2 * border / size
+ i = pow(2, np.ceil(np.log2(np.ceil(k))) + (k == int(k)))
+ return border // i
+
+ def _filter_boxes(self, patch, boxes):
+ """Check whether the center of each box is in the patch.
+
+ Args:
+ patch (list[int]): The cropped area, [left, top, right, bottom].
+ boxes (numpy array, (N x 4)): Ground truth boxes.
+
+ Returns:
+ mask (numpy array, (N,)): Each box is inside or outside the patch.
+ """
+ center = boxes.centers.numpy()
+ mask = (center[:, 0] > patch[0]) * (center[:, 1] > patch[1]) * (
+ center[:, 0] < patch[2]) * (
+ center[:, 1] < patch[3])
+ return mask
+
+ def _crop_image_and_paste(self, image, center, size):
+ """Crop image with a given center and size, then paste the cropped
+ image to a blank image with two centers align.
+
+ This function is equivalent to generating a blank image with ``size``
+ as its shape. Then cover it on the original image with two centers (
+ the center of blank image and the random center of original image)
+ aligned. The overlap area is paste from the original image and the
+ outside area is filled with ``mean pixel``.
+
+ Args:
+ image (np array, H x W x C): Original image.
+ center (list[int]): Target crop center coord.
+ size (list[int]): Target crop size. [target_h, target_w]
+
+ Returns:
+ cropped_img (np array, target_h x target_w x C): Cropped image.
+ border (np array, 4): The distance of four border of
+ ``cropped_img`` to the original image area, [top, bottom,
+ left, right]
+ patch (list[int]): The cropped area, [left, top, right, bottom].
+ """
+ center_y, center_x = center
+ target_h, target_w = size
+ img_h, img_w, img_c = image.shape
+
+ x0 = max(0, center_x - target_w // 2)
+ x1 = min(center_x + target_w // 2, img_w)
+ y0 = max(0, center_y - target_h // 2)
+ y1 = min(center_y + target_h // 2, img_h)
+ patch = np.array((int(x0), int(y0), int(x1), int(y1)))
+
+ left, right = center_x - x0, x1 - center_x
+ top, bottom = center_y - y0, y1 - center_y
+
+ cropped_center_y, cropped_center_x = target_h // 2, target_w // 2
+ cropped_img = np.zeros((target_h, target_w, img_c), dtype=image.dtype)
+ for i in range(img_c):
+ cropped_img[:, :, i] += self.mean[i]
+ y_slice = slice(cropped_center_y - top, cropped_center_y + bottom)
+ x_slice = slice(cropped_center_x - left, cropped_center_x + right)
+ cropped_img[y_slice, x_slice, :] = image[y0:y1, x0:x1, :]
+
+ border = np.array([
+ cropped_center_y - top, cropped_center_y + bottom,
+ cropped_center_x - left, cropped_center_x + right
+ ],
+ dtype=np.float32)
+
+ return cropped_img, border, patch
+
+ def _train_aug(self, results):
+ """Random crop and around padding the original image.
+
+ Args:
+ results (dict): Image infomations in the augment pipeline.
+
+ Returns:
+ results (dict): The updated dict.
+ """
+ img = results['img']
+ h, w, c = img.shape
+ gt_bboxes = results['gt_bboxes']
+ while True:
+ scale = random.choice(self.ratios)
+ new_h = int(self.crop_size[1] * scale)
+ new_w = int(self.crop_size[0] * scale)
+ h_border = self._get_border(self.border, h)
+ w_border = self._get_border(self.border, w)
+
+ for i in range(50):
+ center_x = random.randint(low=w_border, high=w - w_border)
+ center_y = random.randint(low=h_border, high=h - h_border)
+
+ cropped_img, border, patch = self._crop_image_and_paste(
+ img, [center_y, center_x], [new_h, new_w])
+
+ if len(gt_bboxes) == 0:
+ results['img'] = cropped_img
+ results['img_shape'] = cropped_img.shape[:2]
+ return results
+
+ # if image do not have valid bbox, any crop patch is valid.
+ mask = self._filter_boxes(patch, gt_bboxes)
+ if not mask.any():
+ continue
+
+ results['img'] = cropped_img
+ results['img_shape'] = cropped_img.shape[:2]
+
+ x0, y0, x1, y1 = patch
+
+ left_w, top_h = center_x - x0, center_y - y0
+ cropped_center_x, cropped_center_y = new_w // 2, new_h // 2
+
+ # crop bboxes accordingly and clip to the image boundary
+ gt_bboxes = gt_bboxes[mask]
+ gt_bboxes.translate_([
+ cropped_center_x - left_w - x0,
+ cropped_center_y - top_h - y0
+ ])
+ if self.bbox_clip_border:
+ gt_bboxes.clip_([new_h, new_w])
+ keep = gt_bboxes.is_inside([new_h, new_w]).numpy()
+ gt_bboxes = gt_bboxes[keep]
+
+ results['gt_bboxes'] = gt_bboxes
+
+ # ignore_flags
+ if results.get('gt_ignore_flags', None) is not None:
+ gt_ignore_flags = results['gt_ignore_flags'][mask]
+ results['gt_ignore_flags'] = \
+ gt_ignore_flags[keep]
+
+ # labels
+ if results.get('gt_bboxes_labels', None) is not None:
+ gt_labels = results['gt_bboxes_labels'][mask]
+ results['gt_bboxes_labels'] = gt_labels[keep]
+
+ if 'gt_masks' in results or 'gt_seg_map' in results:
+ raise NotImplementedError(
+ 'RandomCenterCropPad only supports bbox.')
+
+ return results
+
+ def _test_aug(self, results):
+ """Around padding the original image without cropping.
+
+ The padding mode and value are from ``test_pad_mode``.
+
+ Args:
+ results (dict): Image infomations in the augment pipeline.
+
+ Returns:
+ results (dict): The updated dict.
+ """
+ img = results['img']
+ h, w, c = img.shape
+ if self.test_pad_mode[0] in ['logical_or']:
+ # self.test_pad_add_pix is only used for centernet
+ target_h = (h | self.test_pad_mode[1]) + self.test_pad_add_pix
+ target_w = (w | self.test_pad_mode[1]) + self.test_pad_add_pix
+ elif self.test_pad_mode[0] in ['size_divisor']:
+ divisor = self.test_pad_mode[1]
+ target_h = int(np.ceil(h / divisor)) * divisor
+ target_w = int(np.ceil(w / divisor)) * divisor
+ else:
+ raise NotImplementedError(
+ 'RandomCenterCropPad only support two testing pad mode:'
+ 'logical-or and size_divisor.')
+
+ cropped_img, border, _ = self._crop_image_and_paste(
+ img, [h // 2, w // 2], [target_h, target_w])
+ results['img'] = cropped_img
+ results['img_shape'] = cropped_img.shape[:2]
+ results['border'] = border
+ return results
+
+ @autocast_box_type()
+ def transform(self, results: dict) -> dict:
+ img = results['img']
+ assert img.dtype == np.float32, (
+ 'RandomCenterCropPad needs the input image of dtype np.float32,'
+ ' please set "to_float32=True" in "LoadImageFromFile" pipeline')
+ h, w, c = img.shape
+ assert c == len(self.mean)
+ if self.test_mode:
+ return self._test_aug(results)
+ else:
+ return self._train_aug(results)
+
+ def __repr__(self):
+ repr_str = self.__class__.__name__
+ repr_str += f'(crop_size={self.crop_size}, '
+ repr_str += f'ratios={self.ratios}, '
+ repr_str += f'border={self.border}, '
+ repr_str += f'mean={self.input_mean}, '
+ repr_str += f'std={self.input_std}, '
+ repr_str += f'to_rgb={self.to_rgb}, '
+ repr_str += f'test_mode={self.test_mode}, '
+ repr_str += f'test_pad_mode={self.test_pad_mode}, '
+ repr_str += f'bbox_clip_border={self.bbox_clip_border})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class CutOut(BaseTransform):
+ """CutOut operation.
+
+ Randomly drop some regions of image used in
+ `Cutout `_.
+
+ Required Keys:
+
+ - img
+
+ Modified Keys:
+
+ - img
+
+ Args:
+ n_holes (int or tuple[int, int]): Number of regions to be dropped.
+ If it is given as a list, number of holes will be randomly
+ selected from the closed interval [``n_holes[0]``, ``n_holes[1]``].
+ cutout_shape (tuple[int, int] or list[tuple[int, int]], optional):
+ The candidate shape of dropped regions. It can be
+ ``tuple[int, int]`` to use a fixed cutout shape, or
+ ``list[tuple[int, int]]`` to randomly choose shape
+ from the list. Defaults to None.
+ cutout_ratio (tuple[float, float] or list[tuple[float, float]],
+ optional): The candidate ratio of dropped regions. It can be
+ ``tuple[float, float]`` to use a fixed ratio or
+ ``list[tuple[float, float]]`` to randomly choose ratio
+ from the list. Please note that ``cutout_shape`` and
+ ``cutout_ratio`` cannot be both given at the same time.
+ Defaults to None.
+ fill_in (tuple[float, float, float] or tuple[int, int, int]): The value
+ of pixel to fill in the dropped regions. Defaults to (0, 0, 0).
+ """
+
+ def __init__(
+ self,
+ n_holes: Union[int, Tuple[int, int]],
+ cutout_shape: Optional[Union[Tuple[int, int],
+ List[Tuple[int, int]]]] = None,
+ cutout_ratio: Optional[Union[Tuple[float, float],
+ List[Tuple[float, float]]]] = None,
+ fill_in: Union[Tuple[float, float, float], Tuple[int, int,
+ int]] = (0, 0, 0)
+ ) -> None:
+
+ assert (cutout_shape is None) ^ (cutout_ratio is None), \
+ 'Either cutout_shape or cutout_ratio should be specified.'
+ assert (isinstance(cutout_shape, (list, tuple))
+ or isinstance(cutout_ratio, (list, tuple)))
+ if isinstance(n_holes, tuple):
+ assert len(n_holes) == 2 and 0 <= n_holes[0] < n_holes[1]
+ else:
+ n_holes = (n_holes, n_holes)
+ self.n_holes = n_holes
+ self.fill_in = fill_in
+ self.with_ratio = cutout_ratio is not None
+ self.candidates = cutout_ratio if self.with_ratio else cutout_shape
+ if not isinstance(self.candidates, list):
+ self.candidates = [self.candidates]
+
+ @autocast_box_type()
+ def transform(self, results: dict) -> dict:
+ """Call function to drop some regions of image."""
+ h, w, c = results['img'].shape
+ n_holes = np.random.randint(self.n_holes[0], self.n_holes[1] + 1)
+ for _ in range(n_holes):
+ x1 = np.random.randint(0, w)
+ y1 = np.random.randint(0, h)
+ index = np.random.randint(0, len(self.candidates))
+ if not self.with_ratio:
+ cutout_w, cutout_h = self.candidates[index]
+ else:
+ cutout_w = int(self.candidates[index][0] * w)
+ cutout_h = int(self.candidates[index][1] * h)
+
+ x2 = np.clip(x1 + cutout_w, 0, w)
+ y2 = np.clip(y1 + cutout_h, 0, h)
+ results['img'][y1:y2, x1:x2, :] = self.fill_in
+
+ return results
+
+ def __repr__(self):
+ repr_str = self.__class__.__name__
+ repr_str += f'(n_holes={self.n_holes}, '
+ repr_str += (f'cutout_ratio={self.candidates}, ' if self.with_ratio
+ else f'cutout_shape={self.candidates}, ')
+ repr_str += f'fill_in={self.fill_in})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class Mosaic(BaseTransform):
+ """Mosaic augmentation.
+
+ Given 4 images, mosaic transform combines them into
+ one output image. The output image is composed of the parts from each sub-
+ image.
+
+ .. code:: text
+
+ mosaic transform
+ center_x
+ +------------------------------+
+ | pad | pad |
+ | +-----------+ |
+ | | | |
+ | | image1 |--------+ |
+ | | | | |
+ | | | image2 | |
+ center_y |----+-------------+-----------|
+ | | cropped | |
+ |pad | image3 | image4 |
+ | | | |
+ +----|-------------+-----------+
+ | |
+ +-------------+
+
+ The mosaic transform steps are as follows:
+
+ 1. Choose the mosaic center as the intersections of 4 images
+ 2. Get the left top image according to the index, and randomly
+ sample another 3 images from the custom dataset.
+ 3. Sub image will be cropped if image is larger than mosaic patch
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_bboxes_labels (np.int64) (optional)
+ - gt_ignore_flags (bool) (optional)
+ - mix_results (List[dict])
+
+ Modified Keys:
+
+ - img
+ - img_shape
+ - gt_bboxes (optional)
+ - gt_bboxes_labels (optional)
+ - gt_ignore_flags (optional)
+
+ Args:
+ img_scale (Sequence[int]): Image size before mosaic pipeline of single
+ image. The shape order should be (width, height).
+ Defaults to (640, 640).
+ center_ratio_range (Sequence[float]): Center ratio range of mosaic
+ output. Defaults to (0.5, 1.5).
+ bbox_clip_border (bool, optional): Whether to clip the objects outside
+ the border of the image. In some dataset like MOT17, the gt bboxes
+ are allowed to cross the border of images. Therefore, we don't
+ need to clip the gt bboxes in these cases. Defaults to True.
+ pad_val (int): Pad value. Defaults to 114.
+ prob (float): Probability of applying this transformation.
+ Defaults to 1.0.
+ """
+
+ def __init__(self,
+ img_scale: Tuple[int, int] = (640, 640),
+ center_ratio_range: Tuple[float, float] = (0.5, 1.5),
+ bbox_clip_border: bool = True,
+ pad_val: float = 114.0,
+ prob: float = 1.0) -> None:
+ assert isinstance(img_scale, tuple)
+ assert 0 <= prob <= 1.0, 'The probability should be in range [0,1]. ' \
+ f'got {prob}.'
+
+ log_img_scale(img_scale, skip_square=True, shape_order='wh')
+ self.img_scale = img_scale
+ self.center_ratio_range = center_ratio_range
+ self.bbox_clip_border = bbox_clip_border
+ self.pad_val = pad_val
+ self.prob = prob
+
+ @cache_randomness
+ def get_indexes(self, dataset: BaseDataset) -> int:
+ """Call function to collect indexes.
+
+ Args:
+ dataset (:obj:`MultiImageMixDataset`): The dataset.
+
+ Returns:
+ list: indexes.
+ """
+
+ indexes = [random.randint(0, len(dataset)) for _ in range(3)]
+ return indexes
+
+ @autocast_box_type()
+ def transform(self, results: dict) -> dict:
+ """Mosaic transform function.
+
+ Args:
+ results (dict): Result dict.
+
+ Returns:
+ dict: Updated result dict.
+ """
+ if random.uniform(0, 1) > self.prob:
+ return results
+
+ assert 'mix_results' in results
+ mosaic_bboxes = []
+ mosaic_bboxes_labels = []
+ mosaic_ignore_flags = []
+ if len(results['img'].shape) == 3:
+ mosaic_img = np.full(
+ (int(self.img_scale[1] * 2), int(self.img_scale[0] * 2), 3),
+ self.pad_val,
+ dtype=results['img'].dtype)
+ else:
+ mosaic_img = np.full(
+ (int(self.img_scale[1] * 2), int(self.img_scale[0] * 2)),
+ self.pad_val,
+ dtype=results['img'].dtype)
+
+ # mosaic center x, y
+ center_x = int(
+ random.uniform(*self.center_ratio_range) * self.img_scale[0])
+ center_y = int(
+ random.uniform(*self.center_ratio_range) * self.img_scale[1])
+ center_position = (center_x, center_y)
+
+ loc_strs = ('top_left', 'top_right', 'bottom_left', 'bottom_right')
+ for i, loc in enumerate(loc_strs):
+ if loc == 'top_left':
+ results_patch = copy.deepcopy(results)
+ else:
+ results_patch = copy.deepcopy(results['mix_results'][i - 1])
+
+ img_i = results_patch['img']
+ h_i, w_i = img_i.shape[:2]
+ # keep_ratio resize
+ scale_ratio_i = min(self.img_scale[1] / h_i,
+ self.img_scale[0] / w_i)
+ img_i = mmcv.imresize(
+ img_i, (int(w_i * scale_ratio_i), int(h_i * scale_ratio_i)))
+
+ # compute the combine parameters
+ paste_coord, crop_coord = self._mosaic_combine(
+ loc, center_position, img_i.shape[:2][::-1])
+ x1_p, y1_p, x2_p, y2_p = paste_coord
+ x1_c, y1_c, x2_c, y2_c = crop_coord
+
+ # crop and paste image
+ mosaic_img[y1_p:y2_p, x1_p:x2_p] = img_i[y1_c:y2_c, x1_c:x2_c]
+
+ # adjust coordinate
+ gt_bboxes_i = results_patch['gt_bboxes']
+ gt_bboxes_labels_i = results_patch['gt_bboxes_labels']
+ gt_ignore_flags_i = results_patch['gt_ignore_flags']
+
+ padw = x1_p - x1_c
+ padh = y1_p - y1_c
+ gt_bboxes_i.rescale_([scale_ratio_i, scale_ratio_i])
+ gt_bboxes_i.translate_([padw, padh])
+ mosaic_bboxes.append(gt_bboxes_i)
+ mosaic_bboxes_labels.append(gt_bboxes_labels_i)
+ mosaic_ignore_flags.append(gt_ignore_flags_i)
+
+ mosaic_bboxes = mosaic_bboxes[0].cat(mosaic_bboxes, 0)
+ mosaic_bboxes_labels = np.concatenate(mosaic_bboxes_labels, 0)
+ mosaic_ignore_flags = np.concatenate(mosaic_ignore_flags, 0)
+
+ if self.bbox_clip_border:
+ mosaic_bboxes.clip_([2 * self.img_scale[1], 2 * self.img_scale[0]])
+ # remove outside bboxes
+ inside_inds = mosaic_bboxes.is_inside(
+ [2 * self.img_scale[1], 2 * self.img_scale[0]]).numpy()
+ mosaic_bboxes = mosaic_bboxes[inside_inds]
+ mosaic_bboxes_labels = mosaic_bboxes_labels[inside_inds]
+ mosaic_ignore_flags = mosaic_ignore_flags[inside_inds]
+
+ results['img'] = mosaic_img
+ results['img_shape'] = mosaic_img.shape[:2]
+ results['gt_bboxes'] = mosaic_bboxes
+ results['gt_bboxes_labels'] = mosaic_bboxes_labels
+ results['gt_ignore_flags'] = mosaic_ignore_flags
+ return results
+
+ def _mosaic_combine(
+ self, loc: str, center_position_xy: Sequence[float],
+ img_shape_wh: Sequence[int]) -> Tuple[Tuple[int], Tuple[int]]:
+ """Calculate global coordinate of mosaic image and local coordinate of
+ cropped sub-image.
+
+ Args:
+ loc (str): Index for the sub-image, loc in ('top_left',
+ 'top_right', 'bottom_left', 'bottom_right').
+ center_position_xy (Sequence[float]): Mixing center for 4 images,
+ (x, y).
+ img_shape_wh (Sequence[int]): Width and height of sub-image
+
+ Returns:
+ tuple[tuple[float]]: Corresponding coordinate of pasting and
+ cropping
+ - paste_coord (tuple): paste corner coordinate in mosaic image.
+ - crop_coord (tuple): crop corner coordinate in mosaic image.
+ """
+ assert loc in ('top_left', 'top_right', 'bottom_left', 'bottom_right')
+ if loc == 'top_left':
+ # index0 to top left part of image
+ x1, y1, x2, y2 = max(center_position_xy[0] - img_shape_wh[0], 0), \
+ max(center_position_xy[1] - img_shape_wh[1], 0), \
+ center_position_xy[0], \
+ center_position_xy[1]
+ crop_coord = img_shape_wh[0] - (x2 - x1), img_shape_wh[1] - (
+ y2 - y1), img_shape_wh[0], img_shape_wh[1]
+
+ elif loc == 'top_right':
+ # index1 to top right part of image
+ x1, y1, x2, y2 = center_position_xy[0], \
+ max(center_position_xy[1] - img_shape_wh[1], 0), \
+ min(center_position_xy[0] + img_shape_wh[0],
+ self.img_scale[0] * 2), \
+ center_position_xy[1]
+ crop_coord = 0, img_shape_wh[1] - (y2 - y1), min(
+ img_shape_wh[0], x2 - x1), img_shape_wh[1]
+
+ elif loc == 'bottom_left':
+ # index2 to bottom left part of image
+ x1, y1, x2, y2 = max(center_position_xy[0] - img_shape_wh[0], 0), \
+ center_position_xy[1], \
+ center_position_xy[0], \
+ min(self.img_scale[1] * 2, center_position_xy[1] +
+ img_shape_wh[1])
+ crop_coord = img_shape_wh[0] - (x2 - x1), 0, img_shape_wh[0], min(
+ y2 - y1, img_shape_wh[1])
+
+ else:
+ # index3 to bottom right part of image
+ x1, y1, x2, y2 = center_position_xy[0], \
+ center_position_xy[1], \
+ min(center_position_xy[0] + img_shape_wh[0],
+ self.img_scale[0] * 2), \
+ min(self.img_scale[1] * 2, center_position_xy[1] +
+ img_shape_wh[1])
+ crop_coord = 0, 0, min(img_shape_wh[0],
+ x2 - x1), min(y2 - y1, img_shape_wh[1])
+
+ paste_coord = x1, y1, x2, y2
+ return paste_coord, crop_coord
+
+ def __repr__(self):
+ repr_str = self.__class__.__name__
+ repr_str += f'(img_scale={self.img_scale}, '
+ repr_str += f'center_ratio_range={self.center_ratio_range}, '
+ repr_str += f'pad_val={self.pad_val}, '
+ repr_str += f'prob={self.prob})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class MixUp(BaseTransform):
+ """MixUp data augmentation.
+
+ .. code:: text
+
+ mixup transform
+ +------------------------------+
+ | mixup image | |
+ | +--------|--------+ |
+ | | | | |
+ |---------------+ | |
+ | | | |
+ | | image | |
+ | | | |
+ | | | |
+ | |-----------------+ |
+ | pad |
+ +------------------------------+
+
+ The mixup transform steps are as follows:
+
+ 1. Another random image is picked by dataset and embedded in
+ the top left patch(after padding and resizing)
+ 2. The target of mixup transform is the weighted average of mixup
+ image and origin image.
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_bboxes_labels (np.int64) (optional)
+ - gt_ignore_flags (bool) (optional)
+ - mix_results (List[dict])
+
+
+ Modified Keys:
+
+ - img
+ - img_shape
+ - gt_bboxes (optional)
+ - gt_bboxes_labels (optional)
+ - gt_ignore_flags (optional)
+
+
+ Args:
+ img_scale (Sequence[int]): Image output size after mixup pipeline.
+ The shape order should be (width, height). Defaults to (640, 640).
+ ratio_range (Sequence[float]): Scale ratio of mixup image.
+ Defaults to (0.5, 1.5).
+ flip_ratio (float): Horizontal flip ratio of mixup image.
+ Defaults to 0.5.
+ pad_val (int): Pad value. Defaults to 114.
+ max_iters (int): The maximum number of iterations. If the number of
+ iterations is greater than `max_iters`, but gt_bbox is still
+ empty, then the iteration is terminated. Defaults to 15.
+ bbox_clip_border (bool, optional): Whether to clip the objects outside
+ the border of the image. In some dataset like MOT17, the gt bboxes
+ are allowed to cross the border of images. Therefore, we don't
+ need to clip the gt bboxes in these cases. Defaults to True.
+ """
+
+ def __init__(self,
+ img_scale: Tuple[int, int] = (640, 640),
+ ratio_range: Tuple[float, float] = (0.5, 1.5),
+ flip_ratio: float = 0.5,
+ pad_val: float = 114.0,
+ max_iters: int = 15,
+ bbox_clip_border: bool = True) -> None:
+ assert isinstance(img_scale, tuple)
+ log_img_scale(img_scale, skip_square=True, shape_order='wh')
+ self.dynamic_scale = img_scale
+ self.ratio_range = ratio_range
+ self.flip_ratio = flip_ratio
+ self.pad_val = pad_val
+ self.max_iters = max_iters
+ self.bbox_clip_border = bbox_clip_border
+
+ @cache_randomness
+ def get_indexes(self, dataset: BaseDataset) -> int:
+ """Call function to collect indexes.
+
+ Args:
+ dataset (:obj:`MultiImageMixDataset`): The dataset.
+
+ Returns:
+ list: indexes.
+ """
+
+ for i in range(self.max_iters):
+ index = random.randint(0, len(dataset))
+ gt_bboxes_i = dataset[index]['gt_bboxes']
+ if len(gt_bboxes_i) != 0:
+ break
+
+ return index
+
+ @autocast_box_type()
+ def transform(self, results: dict) -> dict:
+ """MixUp transform function.
+
+ Args:
+ results (dict): Result dict.
+
+ Returns:
+ dict: Updated result dict.
+ """
+
+ assert 'mix_results' in results
+ assert len(
+ results['mix_results']) == 1, 'MixUp only support 2 images now !'
+
+ if results['mix_results'][0]['gt_bboxes'].shape[0] == 0:
+ # empty bbox
+ return results
+
+ retrieve_results = results['mix_results'][0]
+ retrieve_img = retrieve_results['img']
+
+ jit_factor = random.uniform(*self.ratio_range)
+ is_flip = random.uniform(0, 1) > self.flip_ratio
+
+ if len(retrieve_img.shape) == 3:
+ out_img = np.ones(
+ (self.dynamic_scale[1], self.dynamic_scale[0], 3),
+ dtype=retrieve_img.dtype) * self.pad_val
+ else:
+ out_img = np.ones(
+ self.dynamic_scale[::-1],
+ dtype=retrieve_img.dtype) * self.pad_val
+
+ # 1. keep_ratio resize
+ scale_ratio = min(self.dynamic_scale[1] / retrieve_img.shape[0],
+ self.dynamic_scale[0] / retrieve_img.shape[1])
+ retrieve_img = mmcv.imresize(
+ retrieve_img, (int(retrieve_img.shape[1] * scale_ratio),
+ int(retrieve_img.shape[0] * scale_ratio)))
+
+ # 2. paste
+ out_img[:retrieve_img.shape[0], :retrieve_img.shape[1]] = retrieve_img
+
+ # 3. scale jit
+ scale_ratio *= jit_factor
+ out_img = mmcv.imresize(out_img, (int(out_img.shape[1] * jit_factor),
+ int(out_img.shape[0] * jit_factor)))
+
+ # 4. flip
+ if is_flip:
+ out_img = out_img[:, ::-1, :]
+
+ # 5. random crop
+ ori_img = results['img']
+ origin_h, origin_w = out_img.shape[:2]
+ target_h, target_w = ori_img.shape[:2]
+ padded_img = np.ones((max(origin_h, target_h), max(
+ origin_w, target_w), 3)) * self.pad_val
+ padded_img = padded_img.astype(np.uint8)
+ padded_img[:origin_h, :origin_w] = out_img
+
+ x_offset, y_offset = 0, 0
+ if padded_img.shape[0] > target_h:
+ y_offset = random.randint(0, padded_img.shape[0] - target_h)
+ if padded_img.shape[1] > target_w:
+ x_offset = random.randint(0, padded_img.shape[1] - target_w)
+ padded_cropped_img = padded_img[y_offset:y_offset + target_h,
+ x_offset:x_offset + target_w]
+
+ # 6. adjust bbox
+ retrieve_gt_bboxes = retrieve_results['gt_bboxes']
+ retrieve_gt_bboxes.rescale_([scale_ratio, scale_ratio])
+ if self.bbox_clip_border:
+ retrieve_gt_bboxes.clip_([origin_h, origin_w])
+
+ if is_flip:
+ retrieve_gt_bboxes.flip_([origin_h, origin_w],
+ direction='horizontal')
+
+ # 7. filter
+ cp_retrieve_gt_bboxes = retrieve_gt_bboxes.clone()
+ cp_retrieve_gt_bboxes.translate_([-x_offset, -y_offset])
+ if self.bbox_clip_border:
+ cp_retrieve_gt_bboxes.clip_([target_h, target_w])
+
+ # 8. mix up
+ ori_img = ori_img.astype(np.float32)
+ mixup_img = 0.5 * ori_img + 0.5 * padded_cropped_img.astype(np.float32)
+
+ retrieve_gt_bboxes_labels = retrieve_results['gt_bboxes_labels']
+ retrieve_gt_ignore_flags = retrieve_results['gt_ignore_flags']
+
+ mixup_gt_bboxes = cp_retrieve_gt_bboxes.cat(
+ (results['gt_bboxes'], cp_retrieve_gt_bboxes), dim=0)
+ mixup_gt_bboxes_labels = np.concatenate(
+ (results['gt_bboxes_labels'], retrieve_gt_bboxes_labels), axis=0)
+ mixup_gt_ignore_flags = np.concatenate(
+ (results['gt_ignore_flags'], retrieve_gt_ignore_flags), axis=0)
+
+ # remove outside bbox
+ inside_inds = mixup_gt_bboxes.is_inside([target_h, target_w]).numpy()
+ mixup_gt_bboxes = mixup_gt_bboxes[inside_inds]
+ mixup_gt_bboxes_labels = mixup_gt_bboxes_labels[inside_inds]
+ mixup_gt_ignore_flags = mixup_gt_ignore_flags[inside_inds]
+
+ results['img'] = mixup_img.astype(np.uint8)
+ results['img_shape'] = mixup_img.shape[:2]
+ results['gt_bboxes'] = mixup_gt_bboxes
+ results['gt_bboxes_labels'] = mixup_gt_bboxes_labels
+ results['gt_ignore_flags'] = mixup_gt_ignore_flags
+
+ return results
+
+ def __repr__(self):
+ repr_str = self.__class__.__name__
+ repr_str += f'(dynamic_scale={self.dynamic_scale}, '
+ repr_str += f'ratio_range={self.ratio_range}, '
+ repr_str += f'flip_ratio={self.flip_ratio}, '
+ repr_str += f'pad_val={self.pad_val}, '
+ repr_str += f'max_iters={self.max_iters}, '
+ repr_str += f'bbox_clip_border={self.bbox_clip_border})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class RandomAffine(BaseTransform):
+ """Random affine transform data augmentation.
+
+ This operation randomly generates affine transform matrix which including
+ rotation, translation, shear and scaling transforms.
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_bboxes_labels (np.int64) (optional)
+ - gt_ignore_flags (bool) (optional)
+
+ Modified Keys:
+
+ - img
+ - img_shape
+ - gt_bboxes (optional)
+ - gt_bboxes_labels (optional)
+ - gt_ignore_flags (optional)
+
+ Args:
+ max_rotate_degree (float): Maximum degrees of rotation transform.
+ Defaults to 10.
+ max_translate_ratio (float): Maximum ratio of translation.
+ Defaults to 0.1.
+ scaling_ratio_range (tuple[float]): Min and max ratio of
+ scaling transform. Defaults to (0.5, 1.5).
+ max_shear_degree (float): Maximum degrees of shear
+ transform. Defaults to 2.
+ border (tuple[int]): Distance from width and height sides of input
+ image to adjust output shape. Only used in mosaic dataset.
+ Defaults to (0, 0).
+ border_val (tuple[int]): Border padding values of 3 channels.
+ Defaults to (114, 114, 114).
+ bbox_clip_border (bool, optional): Whether to clip the objects outside
+ the border of the image. In some dataset like MOT17, the gt bboxes
+ are allowed to cross the border of images. Therefore, we don't
+ need to clip the gt bboxes in these cases. Defaults to True.
+ """
+
+ def __init__(self,
+ max_rotate_degree: float = 10.0,
+ max_translate_ratio: float = 0.1,
+ scaling_ratio_range: Tuple[float, float] = (0.5, 1.5),
+ max_shear_degree: float = 2.0,
+ border: Tuple[int, int] = (0, 0),
+ border_val: Tuple[int, int, int] = (114, 114, 114),
+ bbox_clip_border: bool = True) -> None:
+ assert 0 <= max_translate_ratio <= 1
+ assert scaling_ratio_range[0] <= scaling_ratio_range[1]
+ assert scaling_ratio_range[0] > 0
+ self.max_rotate_degree = max_rotate_degree
+ self.max_translate_ratio = max_translate_ratio
+ self.scaling_ratio_range = scaling_ratio_range
+ self.max_shear_degree = max_shear_degree
+ self.border = border
+ self.border_val = border_val
+ self.bbox_clip_border = bbox_clip_border
+
+ @cache_randomness
+ def _get_random_homography_matrix(self, height, width):
+ # Rotation
+ rotation_degree = random.uniform(-self.max_rotate_degree,
+ self.max_rotate_degree)
+ rotation_matrix = self._get_rotation_matrix(rotation_degree)
+
+ # Scaling
+ scaling_ratio = random.uniform(self.scaling_ratio_range[0],
+ self.scaling_ratio_range[1])
+ scaling_matrix = self._get_scaling_matrix(scaling_ratio)
+
+ # Shear
+ x_degree = random.uniform(-self.max_shear_degree,
+ self.max_shear_degree)
+ y_degree = random.uniform(-self.max_shear_degree,
+ self.max_shear_degree)
+ shear_matrix = self._get_shear_matrix(x_degree, y_degree)
+
+ # Translation
+ trans_x = random.uniform(-self.max_translate_ratio,
+ self.max_translate_ratio) * width
+ trans_y = random.uniform(-self.max_translate_ratio,
+ self.max_translate_ratio) * height
+ translate_matrix = self._get_translation_matrix(trans_x, trans_y)
+
+ warp_matrix = (
+ translate_matrix @ shear_matrix @ rotation_matrix @ scaling_matrix)
+ return warp_matrix
+
+ @autocast_box_type()
+ def transform(self, results: dict) -> dict:
+ img = results['img']
+ height = img.shape[0] + self.border[1] * 2
+ width = img.shape[1] + self.border[0] * 2
+
+ warp_matrix = self._get_random_homography_matrix(height, width)
+
+ img = cv2.warpPerspective(
+ img,
+ warp_matrix,
+ dsize=(width, height),
+ borderValue=self.border_val)
+ results['img'] = img
+ results['img_shape'] = img.shape[:2]
+
+ bboxes = results['gt_bboxes']
+ num_bboxes = len(bboxes)
+ if num_bboxes:
+ bboxes.project_(warp_matrix)
+ if self.bbox_clip_border:
+ bboxes.clip_([height, width])
+ # remove outside bbox
+ valid_index = bboxes.is_inside([height, width]).numpy()
+ results['gt_bboxes'] = bboxes[valid_index]
+ results['gt_bboxes_labels'] = results['gt_bboxes_labels'][
+ valid_index]
+ results['gt_ignore_flags'] = results['gt_ignore_flags'][
+ valid_index]
+
+ if 'gt_masks' in results:
+ raise NotImplementedError('RandomAffine only supports bbox.')
+ return results
+
+ def __repr__(self):
+ repr_str = self.__class__.__name__
+ repr_str += f'(max_rotate_degree={self.max_rotate_degree}, '
+ repr_str += f'max_translate_ratio={self.max_translate_ratio}, '
+ repr_str += f'scaling_ratio_range={self.scaling_ratio_range}, '
+ repr_str += f'max_shear_degree={self.max_shear_degree}, '
+ repr_str += f'border={self.border}, '
+ repr_str += f'border_val={self.border_val}, '
+ repr_str += f'bbox_clip_border={self.bbox_clip_border})'
+ return repr_str
+
+ @staticmethod
+ def _get_rotation_matrix(rotate_degrees: float) -> np.ndarray:
+ radian = math.radians(rotate_degrees)
+ rotation_matrix = np.array(
+ [[np.cos(radian), -np.sin(radian), 0.],
+ [np.sin(radian), np.cos(radian), 0.], [0., 0., 1.]],
+ dtype=np.float32)
+ return rotation_matrix
+
+ @staticmethod
+ def _get_scaling_matrix(scale_ratio: float) -> np.ndarray:
+ scaling_matrix = np.array(
+ [[scale_ratio, 0., 0.], [0., scale_ratio, 0.], [0., 0., 1.]],
+ dtype=np.float32)
+ return scaling_matrix
+
+ @staticmethod
+ def _get_shear_matrix(x_shear_degrees: float,
+ y_shear_degrees: float) -> np.ndarray:
+ x_radian = math.radians(x_shear_degrees)
+ y_radian = math.radians(y_shear_degrees)
+ shear_matrix = np.array([[1, np.tan(x_radian), 0.],
+ [np.tan(y_radian), 1, 0.], [0., 0., 1.]],
+ dtype=np.float32)
+ return shear_matrix
+
+ @staticmethod
+ def _get_translation_matrix(x: float, y: float) -> np.ndarray:
+ translation_matrix = np.array([[1, 0., x], [0., 1, y], [0., 0., 1.]],
+ dtype=np.float32)
+ return translation_matrix
+
+
+@TRANSFORMS.register_module()
+class YOLOXHSVRandomAug(BaseTransform):
+ """Apply HSV augmentation to image sequentially. It is referenced from
+ https://github.com/Megvii-
+ BaseDetection/YOLOX/blob/main/yolox/data/data_augment.py#L21.
+
+ Required Keys:
+
+ - img
+
+ Modified Keys:
+
+ - img
+
+ Args:
+ hue_delta (int): delta of hue. Defaults to 5.
+ saturation_delta (int): delta of saturation. Defaults to 30.
+ value_delta (int): delat of value. Defaults to 30.
+ """
+
+ def __init__(self,
+ hue_delta: int = 5,
+ saturation_delta: int = 30,
+ value_delta: int = 30) -> None:
+ self.hue_delta = hue_delta
+ self.saturation_delta = saturation_delta
+ self.value_delta = value_delta
+
+ @cache_randomness
+ def _get_hsv_gains(self):
+ hsv_gains = np.random.uniform(-1, 1, 3) * [
+ self.hue_delta, self.saturation_delta, self.value_delta
+ ]
+ # random selection of h, s, v
+ hsv_gains *= np.random.randint(0, 2, 3)
+ # prevent overflow
+ hsv_gains = hsv_gains.astype(np.int16)
+ return hsv_gains
+
+ def transform(self, results: dict) -> dict:
+ img = results['img']
+ hsv_gains = self._get_hsv_gains()
+ img_hsv = cv2.cvtColor(img, cv2.COLOR_BGR2HSV).astype(np.int16)
+
+ img_hsv[..., 0] = (img_hsv[..., 0] + hsv_gains[0]) % 180
+ img_hsv[..., 1] = np.clip(img_hsv[..., 1] + hsv_gains[1], 0, 255)
+ img_hsv[..., 2] = np.clip(img_hsv[..., 2] + hsv_gains[2], 0, 255)
+ cv2.cvtColor(img_hsv.astype(img.dtype), cv2.COLOR_HSV2BGR, dst=img)
+
+ results['img'] = img
+ return results
+
+ def __repr__(self):
+ repr_str = self.__class__.__name__
+ repr_str += f'(hue_delta={self.hue_delta}, '
+ repr_str += f'saturation_delta={self.saturation_delta}, '
+ repr_str += f'value_delta={self.value_delta})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class CopyPaste(BaseTransform):
+ """Simple Copy-Paste is a Strong Data Augmentation Method for Instance
+ Segmentation The simple copy-paste transform steps are as follows:
+
+ 1. The destination image is already resized with aspect ratio kept,
+ cropped and padded.
+ 2. Randomly select a source image, which is also already resized
+ with aspect ratio kept, cropped and padded in a similar way
+ as the destination image.
+ 3. Randomly select some objects from the source image.
+ 4. Paste these source objects to the destination image directly,
+ due to the source and destination image have the same size.
+ 5. Update object masks of the destination image, for some origin objects
+ may be occluded.
+ 6. Generate bboxes from the updated destination masks and
+ filter some objects which are totally occluded, and adjust bboxes
+ which are partly occluded.
+ 7. Append selected source bboxes, masks, and labels.
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (BaseBoxes[torch.float32]) (optional)
+ - gt_bboxes_labels (np.int64) (optional)
+ - gt_ignore_flags (bool) (optional)
+ - gt_masks (BitmapMasks) (optional)
+
+ Modified Keys:
+
+ - img
+ - gt_bboxes (optional)
+ - gt_bboxes_labels (optional)
+ - gt_ignore_flags (optional)
+ - gt_masks (optional)
+
+ Args:
+ max_num_pasted (int): The maximum number of pasted objects.
+ Defaults to 100.
+ bbox_occluded_thr (int): The threshold of occluded bbox.
+ Defaults to 10.
+ mask_occluded_thr (int): The threshold of occluded mask.
+ Defaults to 300.
+ selected (bool): Whether select objects or not. If select is False,
+ all objects of the source image will be pasted to the
+ destination image.
+ Defaults to True.
+ paste_by_box (bool): Whether use boxes as masks when masks are not
+ available.
+ Defaults to False.
+ """
+
+ def __init__(
+ self,
+ max_num_pasted: int = 100,
+ bbox_occluded_thr: int = 10,
+ mask_occluded_thr: int = 300,
+ selected: bool = True,
+ paste_by_box: bool = False,
+ ) -> None:
+ self.max_num_pasted = max_num_pasted
+ self.bbox_occluded_thr = bbox_occluded_thr
+ self.mask_occluded_thr = mask_occluded_thr
+ self.selected = selected
+ self.paste_by_box = paste_by_box
+
+ @cache_randomness
+ def get_indexes(self, dataset: BaseDataset) -> int:
+ """Call function to collect indexes.s.
+
+ Args:
+ dataset (:obj:`MultiImageMixDataset`): The dataset.
+ Returns:
+ list: Indexes.
+ """
+ return random.randint(0, len(dataset))
+
+ @autocast_box_type()
+ def transform(self, results: dict) -> dict:
+ """Transform function to make a copy-paste of image.
+
+ Args:
+ results (dict): Result dict.
+ Returns:
+ dict: Result dict with copy-paste transformed.
+ """
+
+ assert 'mix_results' in results
+ num_images = len(results['mix_results'])
+ assert num_images == 1, \
+ f'CopyPaste only supports processing 2 images, got {num_images}'
+ if self.selected:
+ selected_results = self._select_object(results['mix_results'][0])
+ else:
+ selected_results = results['mix_results'][0]
+ return self._copy_paste(results, selected_results)
+
+ @cache_randomness
+ def _get_selected_inds(self, num_bboxes: int) -> np.ndarray:
+ max_num_pasted = min(num_bboxes + 1, self.max_num_pasted)
+ num_pasted = np.random.randint(0, max_num_pasted)
+ return np.random.choice(num_bboxes, size=num_pasted, replace=False)
+
+ def get_gt_masks(self, results: dict) -> BitmapMasks:
+ """Get gt_masks originally or generated based on bboxes.
+
+ If gt_masks is not contained in results,
+ it will be generated based on gt_bboxes.
+ Args:
+ results (dict): Result dict.
+ Returns:
+ BitmapMasks: gt_masks, originally or generated based on bboxes.
+ """
+ if results.get('gt_masks', None) is not None:
+ if self.paste_by_box:
+ warnings.warn('gt_masks is already contained in results, '
+ 'so paste_by_box is disabled.')
+ return results['gt_masks']
+ else:
+ if not self.paste_by_box:
+ raise RuntimeError('results does not contain masks.')
+ return results['gt_bboxes'].create_masks(results['img'].shape[:2])
+
+ def _select_object(self, results: dict) -> dict:
+ """Select some objects from the source results."""
+ bboxes = results['gt_bboxes']
+ labels = results['gt_bboxes_labels']
+ masks = self.get_gt_masks(results)
+ ignore_flags = results['gt_ignore_flags']
+
+ selected_inds = self._get_selected_inds(bboxes.shape[0])
+
+ selected_bboxes = bboxes[selected_inds]
+ selected_labels = labels[selected_inds]
+ selected_masks = masks[selected_inds]
+ selected_ignore_flags = ignore_flags[selected_inds]
+
+ results['gt_bboxes'] = selected_bboxes
+ results['gt_bboxes_labels'] = selected_labels
+ results['gt_masks'] = selected_masks
+ results['gt_ignore_flags'] = selected_ignore_flags
+ return results
+
+ def _copy_paste(self, dst_results: dict, src_results: dict) -> dict:
+ """CopyPaste transform function.
+
+ Args:
+ dst_results (dict): Result dict of the destination image.
+ src_results (dict): Result dict of the source image.
+ Returns:
+ dict: Updated result dict.
+ """
+ dst_img = dst_results['img']
+ dst_bboxes = dst_results['gt_bboxes']
+ dst_labels = dst_results['gt_bboxes_labels']
+ dst_masks = self.get_gt_masks(dst_results)
+ dst_ignore_flags = dst_results['gt_ignore_flags']
+
+ src_img = src_results['img']
+ src_bboxes = src_results['gt_bboxes']
+ src_labels = src_results['gt_bboxes_labels']
+ src_masks = src_results['gt_masks']
+ src_ignore_flags = src_results['gt_ignore_flags']
+
+ if len(src_bboxes) == 0:
+ return dst_results
+
+ # update masks and generate bboxes from updated masks
+ composed_mask = np.where(np.any(src_masks.masks, axis=0), 1, 0)
+ updated_dst_masks = self._get_updated_masks(dst_masks, composed_mask)
+ updated_dst_bboxes = updated_dst_masks.get_bboxes(type(dst_bboxes))
+ assert len(updated_dst_bboxes) == len(updated_dst_masks)
+
+ # filter totally occluded objects
+ l1_distance = (updated_dst_bboxes.tensor - dst_bboxes.tensor).abs()
+ bboxes_inds = (l1_distance <= self.bbox_occluded_thr).all(
+ dim=-1).numpy()
+ masks_inds = updated_dst_masks.masks.sum(
+ axis=(1, 2)) > self.mask_occluded_thr
+ valid_inds = bboxes_inds | masks_inds
+
+ # Paste source objects to destination image directly
+ img = dst_img * (1 - composed_mask[..., np.newaxis]
+ ) + src_img * composed_mask[..., np.newaxis]
+ bboxes = src_bboxes.cat([updated_dst_bboxes[valid_inds], src_bboxes])
+ labels = np.concatenate([dst_labels[valid_inds], src_labels])
+ masks = np.concatenate(
+ [updated_dst_masks.masks[valid_inds], src_masks.masks])
+ ignore_flags = np.concatenate(
+ [dst_ignore_flags[valid_inds], src_ignore_flags])
+
+ dst_results['img'] = img
+ dst_results['gt_bboxes'] = bboxes
+ dst_results['gt_bboxes_labels'] = labels
+ dst_results['gt_masks'] = BitmapMasks(masks, masks.shape[1],
+ masks.shape[2])
+ dst_results['gt_ignore_flags'] = ignore_flags
+
+ return dst_results
+
+ def _get_updated_masks(self, masks: BitmapMasks,
+ composed_mask: np.ndarray) -> BitmapMasks:
+ """Update masks with composed mask."""
+ assert masks.masks.shape[-2:] == composed_mask.shape[-2:], \
+ 'Cannot compare two arrays of different size'
+ masks.masks = np.where(composed_mask, 0, masks.masks)
+ return masks
+
+ def __repr__(self):
+ repr_str = self.__class__.__name__
+ repr_str += f'(max_num_pasted={self.max_num_pasted}, '
+ repr_str += f'bbox_occluded_thr={self.bbox_occluded_thr}, '
+ repr_str += f'mask_occluded_thr={self.mask_occluded_thr}, '
+ repr_str += f'selected={self.selected}), '
+ repr_str += f'paste_by_box={self.paste_by_box})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class RandomErasing(BaseTransform):
+ """RandomErasing operation.
+
+ Random Erasing randomly selects a rectangle region
+ in an image and erases its pixels with random values.
+ `RandomErasing `_.
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (HorizontalBoxes[torch.float32]) (optional)
+ - gt_bboxes_labels (np.int64) (optional)
+ - gt_ignore_flags (bool) (optional)
+ - gt_masks (BitmapMasks) (optional)
+
+ Modified Keys:
+ - img
+ - gt_bboxes (optional)
+ - gt_bboxes_labels (optional)
+ - gt_ignore_flags (optional)
+ - gt_masks (optional)
+
+ Args:
+ n_patches (int or tuple[int, int]): Number of regions to be dropped.
+ If it is given as a tuple, number of patches will be randomly
+ selected from the closed interval [``n_patches[0]``,
+ ``n_patches[1]``].
+ ratio (float or tuple[float, float]): The ratio of erased regions.
+ It can be ``float`` to use a fixed ratio or ``tuple[float, float]``
+ to randomly choose ratio from the interval.
+ squared (bool): Whether to erase square region. Defaults to True.
+ bbox_erased_thr (float): The threshold for the maximum area proportion
+ of the bbox to be erased. When the proportion of the area where the
+ bbox is erased is greater than the threshold, the bbox will be
+ removed. Defaults to 0.9.
+ img_border_value (int or float or tuple): The filled values for
+ image border. If float, the same fill value will be used for
+ all the three channels of image. If tuple, it should be 3 elements.
+ Defaults to 128.
+ mask_border_value (int): The fill value used for masks. Defaults to 0.
+ seg_ignore_label (int): The fill value used for segmentation map.
+ Note this value must equals ``ignore_label`` in ``semantic_head``
+ of the corresponding config. Defaults to 255.
+ """
+
+ def __init__(
+ self,
+ n_patches: Union[int, Tuple[int, int]],
+ ratio: Union[float, Tuple[float, float]],
+ squared: bool = True,
+ bbox_erased_thr: float = 0.9,
+ img_border_value: Union[int, float, tuple] = 128,
+ mask_border_value: int = 0,
+ seg_ignore_label: int = 255,
+ ) -> None:
+ if isinstance(n_patches, tuple):
+ assert len(n_patches) == 2 and 0 <= n_patches[0] < n_patches[1]
+ else:
+ n_patches = (n_patches, n_patches)
+ if isinstance(ratio, tuple):
+ assert len(ratio) == 2 and 0 <= ratio[0] < ratio[1] <= 1
+ else:
+ ratio = (ratio, ratio)
+
+ self.n_patches = n_patches
+ self.ratio = ratio
+ self.squared = squared
+ self.bbox_erased_thr = bbox_erased_thr
+ self.img_border_value = img_border_value
+ self.mask_border_value = mask_border_value
+ self.seg_ignore_label = seg_ignore_label
+
+ @cache_randomness
+ def _get_patches(self, img_shape: Tuple[int, int]) -> List[list]:
+ """Get patches for random erasing."""
+ patches = []
+ n_patches = np.random.randint(self.n_patches[0], self.n_patches[1] + 1)
+ for _ in range(n_patches):
+ if self.squared:
+ ratio = np.random.random() * (self.ratio[1] -
+ self.ratio[0]) + self.ratio[0]
+ ratio = (ratio, ratio)
+ else:
+ ratio = (np.random.random() * (self.ratio[1] - self.ratio[0]) +
+ self.ratio[0], np.random.random() *
+ (self.ratio[1] - self.ratio[0]) + self.ratio[0])
+ ph, pw = int(img_shape[0] * ratio[0]), int(img_shape[1] * ratio[1])
+ px1, py1 = np.random.randint(0,
+ img_shape[1] - pw), np.random.randint(
+ 0, img_shape[0] - ph)
+ px2, py2 = px1 + pw, py1 + ph
+ patches.append([px1, py1, px2, py2])
+ return np.array(patches)
+
+ def _transform_img(self, results: dict, patches: List[list]) -> None:
+ """Random erasing the image."""
+ for patch in patches:
+ px1, py1, px2, py2 = patch
+ results['img'][py1:py2, px1:px2, :] = self.img_border_value
+
+ def _transform_bboxes(self, results: dict, patches: List[list]) -> None:
+ """Random erasing the bboxes."""
+ bboxes = results['gt_bboxes']
+ # TODO: unify the logic by using operators in BaseBoxes.
+ assert isinstance(bboxes, HorizontalBoxes)
+ bboxes = bboxes.numpy()
+ left_top = np.maximum(bboxes[:, None, :2], patches[:, :2])
+ right_bottom = np.minimum(bboxes[:, None, 2:], patches[:, 2:])
+ wh = np.maximum(right_bottom - left_top, 0)
+ inter_areas = wh[:, :, 0] * wh[:, :, 1]
+ bbox_areas = (bboxes[:, 2] - bboxes[:, 0]) * (
+ bboxes[:, 3] - bboxes[:, 1])
+ bboxes_erased_ratio = inter_areas.sum(-1) / (bbox_areas + 1e-7)
+ valid_inds = bboxes_erased_ratio < self.bbox_erased_thr
+ results['gt_bboxes'] = HorizontalBoxes(bboxes[valid_inds])
+ results['gt_bboxes_labels'] = results['gt_bboxes_labels'][valid_inds]
+ results['gt_ignore_flags'] = results['gt_ignore_flags'][valid_inds]
+ if results.get('gt_masks', None) is not None:
+ results['gt_masks'] = results['gt_masks'][valid_inds]
+
+ def _transform_masks(self, results: dict, patches: List[list]) -> None:
+ """Random erasing the masks."""
+ for patch in patches:
+ px1, py1, px2, py2 = patch
+ results['gt_masks'].masks[:, py1:py2,
+ px1:px2] = self.mask_border_value
+
+ def _transform_seg(self, results: dict, patches: List[list]) -> None:
+ """Random erasing the segmentation map."""
+ for patch in patches:
+ px1, py1, px2, py2 = patch
+ results['gt_seg_map'][py1:py2, px1:px2] = self.seg_ignore_label
+
+ @autocast_box_type()
+ def transform(self, results: dict) -> dict:
+ """Transform function to erase some regions of image."""
+ patches = self._get_patches(results['img_shape'])
+ self._transform_img(results, patches)
+ if results.get('gt_bboxes', None) is not None:
+ self._transform_bboxes(results, patches)
+ if results.get('gt_masks', None) is not None:
+ self._transform_masks(results, patches)
+ if results.get('gt_seg_map', None) is not None:
+ self._transform_seg(results, patches)
+ return results
+
+ def __repr__(self):
+ repr_str = self.__class__.__name__
+ repr_str += f'(n_patches={self.n_patches}, '
+ repr_str += f'ratio={self.ratio}, '
+ repr_str += f'squared={self.squared}, '
+ repr_str += f'bbox_erased_thr={self.bbox_erased_thr}, '
+ repr_str += f'img_border_value={self.img_border_value}, '
+ repr_str += f'mask_border_value={self.mask_border_value}, '
+ repr_str += f'seg_ignore_label={self.seg_ignore_label})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class CachedMosaic(Mosaic):
+ """Cached mosaic augmentation.
+
+ Cached mosaic transform will random select images from the cache
+ and combine them into one output image.
+
+ .. code:: text
+
+ mosaic transform
+ center_x
+ +------------------------------+
+ | pad | pad |
+ | +-----------+ |
+ | | | |
+ | | image1 |--------+ |
+ | | | | |
+ | | | image2 | |
+ center_y |----+-------------+-----------|
+ | | cropped | |
+ |pad | image3 | image4 |
+ | | | |
+ +----|-------------+-----------+
+ | |
+ +-------------+
+
+ The cached mosaic transform steps are as follows:
+
+ 1. Append the results from the last transform into the cache.
+ 2. Choose the mosaic center as the intersections of 4 images
+ 3. Get the left top image according to the index, and randomly
+ sample another 3 images from the result cache.
+ 4. Sub image will be cropped if image is larger than mosaic patch
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (np.float32) (optional)
+ - gt_bboxes_labels (np.int64) (optional)
+ - gt_ignore_flags (bool) (optional)
+
+ Modified Keys:
+
+ - img
+ - img_shape
+ - gt_bboxes (optional)
+ - gt_bboxes_labels (optional)
+ - gt_ignore_flags (optional)
+
+ Args:
+ img_scale (Sequence[int]): Image size before mosaic pipeline of single
+ image. The shape order should be (width, height).
+ Defaults to (640, 640).
+ center_ratio_range (Sequence[float]): Center ratio range of mosaic
+ output. Defaults to (0.5, 1.5).
+ bbox_clip_border (bool, optional): Whether to clip the objects outside
+ the border of the image. In some dataset like MOT17, the gt bboxes
+ are allowed to cross the border of images. Therefore, we don't
+ need to clip the gt bboxes in these cases. Defaults to True.
+ pad_val (int): Pad value. Defaults to 114.
+ prob (float): Probability of applying this transformation.
+ Defaults to 1.0.
+ max_cached_images (int): The maximum length of the cache. The larger
+ the cache, the stronger the randomness of this transform. As a
+ rule of thumb, providing 10 caches for each image suffices for
+ randomness. Defaults to 40.
+ random_pop (bool): Whether to randomly pop a result from the cache
+ when the cache is full. If set to False, use FIFO popping method.
+ Defaults to True.
+ """
+
+ def __init__(self,
+ *args,
+ max_cached_images: int = 40,
+ random_pop: bool = True,
+ **kwargs) -> None:
+ super().__init__(*args, **kwargs)
+ self.results_cache = []
+ self.random_pop = random_pop
+ assert max_cached_images >= 4, 'The length of cache must >= 4, ' \
+ f'but got {max_cached_images}.'
+ self.max_cached_images = max_cached_images
+
+ @cache_randomness
+ def get_indexes(self, cache: list) -> list:
+ """Call function to collect indexes.
+
+ Args:
+ cache (list): The results cache.
+
+ Returns:
+ list: indexes.
+ """
+
+ indexes = [random.randint(0, len(cache) - 1) for _ in range(3)]
+ return indexes
+
+ @autocast_box_type()
+ def transform(self, results: dict) -> dict:
+ """Mosaic transform function.
+
+ Args:
+ results (dict): Result dict.
+
+ Returns:
+ dict: Updated result dict.
+ """
+ # cache and pop images
+ self.results_cache.append(copy.deepcopy(results))
+ if len(self.results_cache) > self.max_cached_images:
+ if self.random_pop:
+ index = random.randint(0, len(self.results_cache) - 1)
+ else:
+ index = 0
+ self.results_cache.pop(index)
+
+ if len(self.results_cache) <= 4:
+ return results
+
+ if random.uniform(0, 1) > self.prob:
+ return results
+ indices = self.get_indexes(self.results_cache)
+ mix_results = [copy.deepcopy(self.results_cache[i]) for i in indices]
+
+ # TODO: refactor mosaic to reuse these code.
+ mosaic_bboxes = []
+ mosaic_bboxes_labels = []
+ mosaic_ignore_flags = []
+ mosaic_masks = []
+ with_mask = True if 'gt_masks' in results else False
+
+ if len(results['img'].shape) == 3:
+ mosaic_img = np.full(
+ (int(self.img_scale[1] * 2), int(self.img_scale[0] * 2), 3),
+ self.pad_val,
+ dtype=results['img'].dtype)
+ else:
+ mosaic_img = np.full(
+ (int(self.img_scale[1] * 2), int(self.img_scale[0] * 2)),
+ self.pad_val,
+ dtype=results['img'].dtype)
+
+ # mosaic center x, y
+ center_x = int(
+ random.uniform(*self.center_ratio_range) * self.img_scale[0])
+ center_y = int(
+ random.uniform(*self.center_ratio_range) * self.img_scale[1])
+ center_position = (center_x, center_y)
+
+ loc_strs = ('top_left', 'top_right', 'bottom_left', 'bottom_right')
+ for i, loc in enumerate(loc_strs):
+ if loc == 'top_left':
+ results_patch = copy.deepcopy(results)
+ else:
+ results_patch = copy.deepcopy(mix_results[i - 1])
+
+ img_i = results_patch['img']
+ h_i, w_i = img_i.shape[:2]
+ # keep_ratio resize
+ scale_ratio_i = min(self.img_scale[1] / h_i,
+ self.img_scale[0] / w_i)
+ img_i = mmcv.imresize(
+ img_i, (int(w_i * scale_ratio_i), int(h_i * scale_ratio_i)))
+
+ # compute the combine parameters
+ paste_coord, crop_coord = self._mosaic_combine(
+ loc, center_position, img_i.shape[:2][::-1])
+ x1_p, y1_p, x2_p, y2_p = paste_coord
+ x1_c, y1_c, x2_c, y2_c = crop_coord
+
+ # crop and paste image
+ mosaic_img[y1_p:y2_p, x1_p:x2_p] = img_i[y1_c:y2_c, x1_c:x2_c]
+
+ # adjust coordinate
+ gt_bboxes_i = results_patch['gt_bboxes']
+ gt_bboxes_labels_i = results_patch['gt_bboxes_labels']
+ gt_ignore_flags_i = results_patch['gt_ignore_flags']
+
+ padw = x1_p - x1_c
+ padh = y1_p - y1_c
+ gt_bboxes_i.rescale_([scale_ratio_i, scale_ratio_i])
+ gt_bboxes_i.translate_([padw, padh])
+ mosaic_bboxes.append(gt_bboxes_i)
+ mosaic_bboxes_labels.append(gt_bboxes_labels_i)
+ mosaic_ignore_flags.append(gt_ignore_flags_i)
+ if with_mask and results_patch.get('gt_masks', None) is not None:
+ gt_masks_i = results_patch['gt_masks']
+ gt_masks_i = gt_masks_i.rescale(float(scale_ratio_i))
+ gt_masks_i = gt_masks_i.translate(
+ out_shape=(int(self.img_scale[0] * 2),
+ int(self.img_scale[1] * 2)),
+ offset=padw,
+ direction='horizontal')
+ gt_masks_i = gt_masks_i.translate(
+ out_shape=(int(self.img_scale[0] * 2),
+ int(self.img_scale[1] * 2)),
+ offset=padh,
+ direction='vertical')
+ mosaic_masks.append(gt_masks_i)
+
+ mosaic_bboxes = mosaic_bboxes[0].cat(mosaic_bboxes, 0)
+ mosaic_bboxes_labels = np.concatenate(mosaic_bboxes_labels, 0)
+ mosaic_ignore_flags = np.concatenate(mosaic_ignore_flags, 0)
+
+ if self.bbox_clip_border:
+ mosaic_bboxes.clip_([2 * self.img_scale[1], 2 * self.img_scale[0]])
+ # remove outside bboxes
+ inside_inds = mosaic_bboxes.is_inside(
+ [2 * self.img_scale[1], 2 * self.img_scale[0]]).numpy()
+ mosaic_bboxes = mosaic_bboxes[inside_inds]
+ mosaic_bboxes_labels = mosaic_bboxes_labels[inside_inds]
+ mosaic_ignore_flags = mosaic_ignore_flags[inside_inds]
+
+ results['img'] = mosaic_img
+ results['img_shape'] = mosaic_img.shape[:2]
+ results['gt_bboxes'] = mosaic_bboxes
+ results['gt_bboxes_labels'] = mosaic_bboxes_labels
+ results['gt_ignore_flags'] = mosaic_ignore_flags
+
+ if with_mask:
+ mosaic_masks = mosaic_masks[0].cat(mosaic_masks)
+ results['gt_masks'] = mosaic_masks[inside_inds]
+ return results
+
+ def __repr__(self):
+ repr_str = self.__class__.__name__
+ repr_str += f'(img_scale={self.img_scale}, '
+ repr_str += f'center_ratio_range={self.center_ratio_range}, '
+ repr_str += f'pad_val={self.pad_val}, '
+ repr_str += f'prob={self.prob}, '
+ repr_str += f'max_cached_images={self.max_cached_images}, '
+ repr_str += f'random_pop={self.random_pop})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class CachedMixUp(BaseTransform):
+ """Cached mixup data augmentation.
+
+ .. code:: text
+
+ mixup transform
+ +------------------------------+
+ | mixup image | |
+ | +--------|--------+ |
+ | | | | |
+ |---------------+ | |
+ | | | |
+ | | image | |
+ | | | |
+ | | | |
+ | |-----------------+ |
+ | pad |
+ +------------------------------+
+
+ The cached mixup transform steps are as follows:
+
+ 1. Append the results from the last transform into the cache.
+ 2. Another random image is picked from the cache and embedded in
+ the top left patch(after padding and resizing)
+ 3. The target of mixup transform is the weighted average of mixup
+ image and origin image.
+
+ Required Keys:
+
+ - img
+ - gt_bboxes (np.float32) (optional)
+ - gt_bboxes_labels (np.int64) (optional)
+ - gt_ignore_flags (bool) (optional)
+ - mix_results (List[dict])
+
+
+ Modified Keys:
+
+ - img
+ - img_shape
+ - gt_bboxes (optional)
+ - gt_bboxes_labels (optional)
+ - gt_ignore_flags (optional)
+
+
+ Args:
+ img_scale (Sequence[int]): Image output size after mixup pipeline.
+ The shape order should be (width, height). Defaults to (640, 640).
+ ratio_range (Sequence[float]): Scale ratio of mixup image.
+ Defaults to (0.5, 1.5).
+ flip_ratio (float): Horizontal flip ratio of mixup image.
+ Defaults to 0.5.
+ pad_val (int): Pad value. Defaults to 114.
+ max_iters (int): The maximum number of iterations. If the number of
+ iterations is greater than `max_iters`, but gt_bbox is still
+ empty, then the iteration is terminated. Defaults to 15.
+ bbox_clip_border (bool, optional): Whether to clip the objects outside
+ the border of the image. In some dataset like MOT17, the gt bboxes
+ are allowed to cross the border of images. Therefore, we don't
+ need to clip the gt bboxes in these cases. Defaults to True.
+ max_cached_images (int): The maximum length of the cache. The larger
+ the cache, the stronger the randomness of this transform. As a
+ rule of thumb, providing 10 caches for each image suffices for
+ randomness. Defaults to 20.
+ random_pop (bool): Whether to randomly pop a result from the cache
+ when the cache is full. If set to False, use FIFO popping method.
+ Defaults to True.
+ prob (float): Probability of applying this transformation.
+ Defaults to 1.0.
+ """
+
+ def __init__(self,
+ img_scale: Tuple[int, int] = (640, 640),
+ ratio_range: Tuple[float, float] = (0.5, 1.5),
+ flip_ratio: float = 0.5,
+ pad_val: float = 114.0,
+ max_iters: int = 15,
+ bbox_clip_border: bool = True,
+ max_cached_images: int = 20,
+ random_pop: bool = True,
+ prob: float = 1.0) -> None:
+ assert isinstance(img_scale, tuple)
+ assert max_cached_images >= 2, 'The length of cache must >= 2, ' \
+ f'but got {max_cached_images}.'
+ assert 0 <= prob <= 1.0, 'The probability should be in range [0,1]. ' \
+ f'got {prob}.'
+ self.dynamic_scale = img_scale
+ self.ratio_range = ratio_range
+ self.flip_ratio = flip_ratio
+ self.pad_val = pad_val
+ self.max_iters = max_iters
+ self.bbox_clip_border = bbox_clip_border
+ self.results_cache = []
+
+ self.max_cached_images = max_cached_images
+ self.random_pop = random_pop
+ self.prob = prob
+
+ @cache_randomness
+ def get_indexes(self, cache: list) -> int:
+ """Call function to collect indexes.
+
+ Args:
+ cache (list): The result cache.
+
+ Returns:
+ int: index.
+ """
+
+ for i in range(self.max_iters):
+ index = random.randint(0, len(cache) - 1)
+ gt_bboxes_i = cache[index]['gt_bboxes']
+ if len(gt_bboxes_i) != 0:
+ break
+ return index
+
+ @autocast_box_type()
+ def transform(self, results: dict) -> dict:
+ """MixUp transform function.
+
+ Args:
+ results (dict): Result dict.
+
+ Returns:
+ dict: Updated result dict.
+ """
+ # cache and pop images
+ self.results_cache.append(copy.deepcopy(results))
+ if len(self.results_cache) > self.max_cached_images:
+ if self.random_pop:
+ index = random.randint(0, len(self.results_cache) - 1)
+ else:
+ index = 0
+ self.results_cache.pop(index)
+
+ if len(self.results_cache) <= 1:
+ return results
+
+ if random.uniform(0, 1) > self.prob:
+ return results
+
+ index = self.get_indexes(self.results_cache)
+ retrieve_results = copy.deepcopy(self.results_cache[index])
+
+ # TODO: refactor mixup to reuse these code.
+ if retrieve_results['gt_bboxes'].shape[0] == 0:
+ # empty bbox
+ return results
+
+ retrieve_img = retrieve_results['img']
+ with_mask = True if 'gt_masks' in results else False
+
+ jit_factor = random.uniform(*self.ratio_range)
+ is_flip = random.uniform(0, 1) > self.flip_ratio
+
+ if len(retrieve_img.shape) == 3:
+ out_img = np.ones(
+ (self.dynamic_scale[1], self.dynamic_scale[0], 3),
+ dtype=retrieve_img.dtype) * self.pad_val
+ else:
+ out_img = np.ones(
+ self.dynamic_scale[::-1],
+ dtype=retrieve_img.dtype) * self.pad_val
+
+ # 1. keep_ratio resize
+ scale_ratio = min(self.dynamic_scale[1] / retrieve_img.shape[0],
+ self.dynamic_scale[0] / retrieve_img.shape[1])
+ retrieve_img = mmcv.imresize(
+ retrieve_img, (int(retrieve_img.shape[1] * scale_ratio),
+ int(retrieve_img.shape[0] * scale_ratio)))
+
+ # 2. paste
+ out_img[:retrieve_img.shape[0], :retrieve_img.shape[1]] = retrieve_img
+
+ # 3. scale jit
+ scale_ratio *= jit_factor
+ out_img = mmcv.imresize(out_img, (int(out_img.shape[1] * jit_factor),
+ int(out_img.shape[0] * jit_factor)))
+
+ # 4. flip
+ if is_flip:
+ out_img = out_img[:, ::-1, :]
+
+ # 5. random crop
+ ori_img = results['img']
+ origin_h, origin_w = out_img.shape[:2]
+ target_h, target_w = ori_img.shape[:2]
+ padded_img = np.ones((max(origin_h, target_h), max(
+ origin_w, target_w), 3)) * self.pad_val
+ padded_img = padded_img.astype(np.uint8)
+ padded_img[:origin_h, :origin_w] = out_img
+
+ x_offset, y_offset = 0, 0
+ if padded_img.shape[0] > target_h:
+ y_offset = random.randint(0, padded_img.shape[0] - target_h)
+ if padded_img.shape[1] > target_w:
+ x_offset = random.randint(0, padded_img.shape[1] - target_w)
+ padded_cropped_img = padded_img[y_offset:y_offset + target_h,
+ x_offset:x_offset + target_w]
+
+ # 6. adjust bbox
+ retrieve_gt_bboxes = retrieve_results['gt_bboxes']
+ retrieve_gt_bboxes.rescale_([scale_ratio, scale_ratio])
+ if with_mask:
+ retrieve_gt_masks = retrieve_results['gt_masks'].rescale(
+ scale_ratio)
+
+ if self.bbox_clip_border:
+ retrieve_gt_bboxes.clip_([origin_h, origin_w])
+
+ if is_flip:
+ retrieve_gt_bboxes.flip_([origin_h, origin_w],
+ direction='horizontal')
+ if with_mask:
+ retrieve_gt_masks = retrieve_gt_masks.flip()
+
+ # 7. filter
+ cp_retrieve_gt_bboxes = retrieve_gt_bboxes.clone()
+ cp_retrieve_gt_bboxes.translate_([-x_offset, -y_offset])
+ if with_mask:
+ retrieve_gt_masks = retrieve_gt_masks.translate(
+ out_shape=(target_h, target_w),
+ offset=-x_offset,
+ direction='horizontal')
+ retrieve_gt_masks = retrieve_gt_masks.translate(
+ out_shape=(target_h, target_w),
+ offset=-y_offset,
+ direction='vertical')
+
+ if self.bbox_clip_border:
+ cp_retrieve_gt_bboxes.clip_([target_h, target_w])
+
+ # 8. mix up
+ ori_img = ori_img.astype(np.float32)
+ mixup_img = 0.5 * ori_img + 0.5 * padded_cropped_img.astype(np.float32)
+
+ retrieve_gt_bboxes_labels = retrieve_results['gt_bboxes_labels']
+ retrieve_gt_ignore_flags = retrieve_results['gt_ignore_flags']
+
+ mixup_gt_bboxes = cp_retrieve_gt_bboxes.cat(
+ (results['gt_bboxes'], cp_retrieve_gt_bboxes), dim=0)
+ mixup_gt_bboxes_labels = np.concatenate(
+ (results['gt_bboxes_labels'], retrieve_gt_bboxes_labels), axis=0)
+ mixup_gt_ignore_flags = np.concatenate(
+ (results['gt_ignore_flags'], retrieve_gt_ignore_flags), axis=0)
+ if with_mask:
+ mixup_gt_masks = retrieve_gt_masks.cat(
+ [results['gt_masks'], retrieve_gt_masks])
+
+ # remove outside bbox
+ inside_inds = mixup_gt_bboxes.is_inside([target_h, target_w]).numpy()
+ mixup_gt_bboxes = mixup_gt_bboxes[inside_inds]
+ mixup_gt_bboxes_labels = mixup_gt_bboxes_labels[inside_inds]
+ mixup_gt_ignore_flags = mixup_gt_ignore_flags[inside_inds]
+ if with_mask:
+ mixup_gt_masks = mixup_gt_masks[inside_inds]
+
+ results['img'] = mixup_img.astype(np.uint8)
+ results['img_shape'] = mixup_img.shape[:2]
+ results['gt_bboxes'] = mixup_gt_bboxes
+ results['gt_bboxes_labels'] = mixup_gt_bboxes_labels
+ results['gt_ignore_flags'] = mixup_gt_ignore_flags
+ if with_mask:
+ results['gt_masks'] = mixup_gt_masks
+ return results
+
+ def __repr__(self):
+ repr_str = self.__class__.__name__
+ repr_str += f'(dynamic_scale={self.dynamic_scale}, '
+ repr_str += f'ratio_range={self.ratio_range}, '
+ repr_str += f'flip_ratio={self.flip_ratio}, '
+ repr_str += f'pad_val={self.pad_val}, '
+ repr_str += f'max_iters={self.max_iters}, '
+ repr_str += f'bbox_clip_border={self.bbox_clip_border}, '
+ repr_str += f'max_cached_images={self.max_cached_images}, '
+ repr_str += f'random_pop={self.random_pop}, '
+ repr_str += f'prob={self.prob})'
+ return repr_str
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/wrappers.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/wrappers.py
new file mode 100644
index 0000000000000000000000000000000000000000..3a17711c06bfbd4dc0038dce9ea7796d1476c37e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/transforms/wrappers.py
@@ -0,0 +1,277 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from typing import Callable, Dict, List, Optional, Union
+
+import numpy as np
+from mmcv.transforms import BaseTransform, Compose
+from mmcv.transforms.utils import cache_random_params, cache_randomness
+
+from mmdet.registry import TRANSFORMS
+
+
+@TRANSFORMS.register_module()
+class MultiBranch(BaseTransform):
+ r"""Multiple branch pipeline wrapper.
+
+ Generate multiple data-augmented versions of the same image.
+ `MultiBranch` needs to specify the branch names of all
+ pipelines of the dataset, perform corresponding data augmentation
+ for the current branch, and return None for other branches,
+ which ensures the consistency of return format across
+ different samples.
+
+ Args:
+ branch_field (list): List of branch names.
+ branch_pipelines (dict): Dict of different pipeline configs
+ to be composed.
+
+ Examples:
+ >>> branch_field = ['sup', 'unsup_teacher', 'unsup_student']
+ >>> sup_pipeline = [
+ >>> dict(type='LoadImageFromFile'),
+ >>> dict(type='LoadAnnotations', with_bbox=True),
+ >>> dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ >>> dict(type='RandomFlip', prob=0.5),
+ >>> dict(
+ >>> type='MultiBranch',
+ >>> branch_field=branch_field,
+ >>> sup=dict(type='PackDetInputs'))
+ >>> ]
+ >>> weak_pipeline = [
+ >>> dict(type='LoadImageFromFile'),
+ >>> dict(type='LoadAnnotations', with_bbox=True),
+ >>> dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ >>> dict(type='RandomFlip', prob=0.0),
+ >>> dict(
+ >>> type='MultiBranch',
+ >>> branch_field=branch_field,
+ >>> sup=dict(type='PackDetInputs'))
+ >>> ]
+ >>> strong_pipeline = [
+ >>> dict(type='LoadImageFromFile'),
+ >>> dict(type='LoadAnnotations', with_bbox=True),
+ >>> dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ >>> dict(type='RandomFlip', prob=1.0),
+ >>> dict(
+ >>> type='MultiBranch',
+ >>> branch_field=branch_field,
+ >>> sup=dict(type='PackDetInputs'))
+ >>> ]
+ >>> unsup_pipeline = [
+ >>> dict(type='LoadImageFromFile'),
+ >>> dict(type='LoadEmptyAnnotations'),
+ >>> dict(
+ >>> type='MultiBranch',
+ >>> branch_field=branch_field,
+ >>> unsup_teacher=weak_pipeline,
+ >>> unsup_student=strong_pipeline)
+ >>> ]
+ >>> from mmcv.transforms import Compose
+ >>> sup_branch = Compose(sup_pipeline)
+ >>> unsup_branch = Compose(unsup_pipeline)
+ >>> print(sup_branch)
+ >>> Compose(
+ >>> LoadImageFromFile(ignore_empty=False, to_float32=False, color_type='color', imdecode_backend='cv2') # noqa
+ >>> LoadAnnotations(with_bbox=True, with_label=True, with_mask=False, with_seg=False, poly2mask=True, imdecode_backend='cv2') # noqa
+ >>> Resize(scale=(1333, 800), scale_factor=None, keep_ratio=True, clip_object_border=True), backend=cv2), interpolation=bilinear) # noqa
+ >>> RandomFlip(prob=0.5, direction=horizontal)
+ >>> MultiBranch(branch_pipelines=['sup'])
+ >>> )
+ >>> print(unsup_branch)
+ >>> Compose(
+ >>> LoadImageFromFile(ignore_empty=False, to_float32=False, color_type='color', imdecode_backend='cv2') # noqa
+ >>> LoadEmptyAnnotations(with_bbox=True, with_label=True, with_mask=False, with_seg=False, seg_ignore_label=255) # noqa
+ >>> MultiBranch(branch_pipelines=['unsup_teacher', 'unsup_student'])
+ >>> )
+ """
+
+ def __init__(self, branch_field: List[str],
+ **branch_pipelines: dict) -> None:
+ self.branch_field = branch_field
+ self.branch_pipelines = {
+ branch: Compose(pipeline)
+ for branch, pipeline in branch_pipelines.items()
+ }
+
+ def transform(self, results: dict) -> dict:
+ """Transform function to apply transforms sequentially.
+
+ Args:
+ results (dict): Result dict from loading pipeline.
+
+ Returns:
+ dict:
+
+ - 'inputs' (Dict[str, obj:`torch.Tensor`]): The forward data of
+ models from different branches.
+ - 'data_sample' (Dict[str,obj:`DetDataSample`]): The annotation
+ info of the sample from different branches.
+ """
+
+ multi_results = {}
+ for branch in self.branch_field:
+ multi_results[branch] = {'inputs': None, 'data_samples': None}
+ for branch, pipeline in self.branch_pipelines.items():
+ branch_results = pipeline(copy.deepcopy(results))
+ # If one branch pipeline returns None,
+ # it will sample another data from dataset.
+ if branch_results is None:
+ return None
+ multi_results[branch] = branch_results
+
+ format_results = {}
+ for branch, results in multi_results.items():
+ for key in results.keys():
+ if format_results.get(key, None) is None:
+ format_results[key] = {branch: results[key]}
+ else:
+ format_results[key][branch] = results[key]
+ return format_results
+
+ def __repr__(self) -> str:
+ repr_str = self.__class__.__name__
+ repr_str += f'(branch_pipelines={list(self.branch_pipelines.keys())})'
+ return repr_str
+
+
+@TRANSFORMS.register_module()
+class RandomOrder(Compose):
+ """Shuffle the transform Sequence."""
+
+ @cache_randomness
+ def _random_permutation(self):
+ return np.random.permutation(len(self.transforms))
+
+ def transform(self, results: Dict) -> Optional[Dict]:
+ """Transform function to apply transforms in random order.
+
+ Args:
+ results (dict): A result dict contains the results to transform.
+
+ Returns:
+ dict or None: Transformed results.
+ """
+ inds = self._random_permutation()
+ for idx in inds:
+ t = self.transforms[idx]
+ results = t(results)
+ if results is None:
+ return None
+ return results
+
+ def __repr__(self):
+ """Compute the string representation."""
+ format_string = self.__class__.__name__ + '('
+ for t in self.transforms:
+ format_string += f'{t.__class__.__name__}, '
+ format_string += ')'
+ return format_string
+
+
+@TRANSFORMS.register_module()
+class ProposalBroadcaster(BaseTransform):
+ """A transform wrapper to apply the wrapped transforms to process both
+ `gt_bboxes` and `proposals` without adding any codes. It will do the
+ following steps:
+
+ 1. Scatter the broadcasting targets to a list of inputs of the wrapped
+ transforms. The type of the list should be list[dict, dict], which
+ the first is the original inputs, the second is the processing
+ results that `gt_bboxes` being rewritten by the `proposals`.
+ 2. Apply ``self.transforms``, with same random parameters, which is
+ sharing with a context manager. The type of the outputs is a
+ list[dict, dict].
+ 3. Gather the outputs, update the `proposals` in the first item of
+ the outputs with the `gt_bboxes` in the second .
+
+ Args:
+ transforms (list, optional): Sequence of transform
+ object or config dict to be wrapped. Defaults to [].
+
+ Note: The `TransformBroadcaster` in MMCV can achieve the same operation as
+ `ProposalBroadcaster`, but need to set more complex parameters.
+
+ Examples:
+ >>> pipeline = [
+ >>> dict(type='LoadImageFromFile'),
+ >>> dict(type='LoadProposals', num_max_proposals=2000),
+ >>> dict(type='LoadAnnotations', with_bbox=True),
+ >>> dict(
+ >>> type='ProposalBroadcaster',
+ >>> transforms=[
+ >>> dict(type='Resize', scale=(1333, 800),
+ >>> keep_ratio=True),
+ >>> dict(type='RandomFlip', prob=0.5),
+ >>> ]),
+ >>> dict(type='PackDetInputs')]
+ """
+
+ def __init__(self, transforms: List[Union[dict, Callable]] = []) -> None:
+ self.transforms = Compose(transforms)
+
+ def transform(self, results: dict) -> dict:
+ """Apply wrapped transform functions to process both `gt_bboxes` and
+ `proposals`.
+
+ Args:
+ results (dict): Result dict from loading pipeline.
+
+ Returns:
+ dict: Updated result dict.
+ """
+ assert results.get('proposals', None) is not None, \
+ '`proposals` should be in the results, please delete ' \
+ '`ProposalBroadcaster` in your configs, or check whether ' \
+ 'you have load proposals successfully.'
+
+ inputs = self._process_input(results)
+ outputs = self._apply_transforms(inputs)
+ outputs = self._process_output(outputs)
+ return outputs
+
+ def _process_input(self, data: dict) -> list:
+ """Scatter the broadcasting targets to a list of inputs of the wrapped
+ transforms.
+
+ Args:
+ data (dict): The original input data.
+
+ Returns:
+ list[dict]: A list of input data.
+ """
+ cp_data = copy.deepcopy(data)
+ cp_data['gt_bboxes'] = cp_data['proposals']
+ scatters = [data, cp_data]
+ return scatters
+
+ def _apply_transforms(self, inputs: list) -> list:
+ """Apply ``self.transforms``.
+
+ Args:
+ inputs (list[dict, dict]): list of input data.
+
+ Returns:
+ list[dict]: The output of the wrapped pipeline.
+ """
+ assert len(inputs) == 2
+ ctx = cache_random_params
+ with ctx(self.transforms):
+ output_scatters = [self.transforms(_input) for _input in inputs]
+ return output_scatters
+
+ def _process_output(self, output_scatters: list) -> dict:
+ """Gathering and renaming data items.
+
+ Args:
+ output_scatters (list[dict, dict]): The output of the wrapped
+ pipeline.
+
+ Returns:
+ dict: Updated result dict.
+ """
+ assert isinstance(output_scatters, list) and \
+ isinstance(output_scatters[0], dict) and \
+ len(output_scatters) == 2
+ outputs = output_scatters[0]
+ outputs['proposals'] = output_scatters[1]['gt_bboxes']
+ return outputs
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/utils.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/utils.py
new file mode 100644
index 0000000000000000000000000000000000000000..d794eb4b06ec9db56ff3a5fc7b817d1d9332a989
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/utils.py
@@ -0,0 +1,48 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+from mmcv.transforms import LoadImageFromFile
+
+from mmdet.datasets.transforms import LoadAnnotations, LoadPanopticAnnotations
+from mmdet.registry import TRANSFORMS
+
+
+def get_loading_pipeline(pipeline):
+ """Only keep loading image and annotations related configuration.
+
+ Args:
+ pipeline (list[dict]): Data pipeline configs.
+
+ Returns:
+ list[dict]: The new pipeline list with only keep
+ loading image and annotations related configuration.
+
+ Examples:
+ >>> pipelines = [
+ ... dict(type='LoadImageFromFile'),
+ ... dict(type='LoadAnnotations', with_bbox=True),
+ ... dict(type='Resize', img_scale=(1333, 800), keep_ratio=True),
+ ... dict(type='RandomFlip', flip_ratio=0.5),
+ ... dict(type='Normalize', **img_norm_cfg),
+ ... dict(type='Pad', size_divisor=32),
+ ... dict(type='DefaultFormatBundle'),
+ ... dict(type='Collect', keys=['img', 'gt_bboxes', 'gt_labels'])
+ ... ]
+ >>> expected_pipelines = [
+ ... dict(type='LoadImageFromFile'),
+ ... dict(type='LoadAnnotations', with_bbox=True)
+ ... ]
+ >>> assert expected_pipelines ==\
+ ... get_loading_pipeline(pipelines)
+ """
+ loading_pipeline_cfg = []
+ for cfg in pipeline:
+ obj_cls = TRANSFORMS.get(cfg['type'])
+ # TODO:use more elegant way to distinguish loading modules
+ if obj_cls is not None and obj_cls in (LoadImageFromFile,
+ LoadAnnotations,
+ LoadPanopticAnnotations):
+ loading_pipeline_cfg.append(cfg)
+ assert len(loading_pipeline_cfg) == 2, \
+ 'The data pipeline in your config file must include ' \
+ 'loading image and annotations related pipeline.'
+ return loading_pipeline_cfg
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/v3det.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/v3det.py
new file mode 100644
index 0000000000000000000000000000000000000000..25bfe3bc718841143653c54954240186c3376955
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/v3det.py
@@ -0,0 +1,32 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import os.path
+from typing import Optional
+
+import mmengine
+
+from mmdet.registry import DATASETS
+from .coco import CocoDataset
+
+
+@DATASETS.register_module()
+class V3DetDataset(CocoDataset):
+ """Dataset for V3Det."""
+
+ METAINFO = {
+ 'classes': None,
+ 'palette': None,
+ }
+
+ def __init__(
+ self,
+ *args,
+ metainfo: Optional[dict] = None,
+ data_root: str = '',
+ label_file='annotations/category_name_13204_v3det_2023_v1.txt', # noqa
+ **kwargs) -> None:
+ class_names = tuple(
+ mmengine.list_from_file(os.path.join(data_root, label_file)))
+ if metainfo is None:
+ metainfo = {'classes': class_names}
+ super().__init__(
+ *args, data_root=data_root, metainfo=metainfo, **kwargs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/voc.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/voc.py
new file mode 100644
index 0000000000000000000000000000000000000000..65e73f2f0bd4f2b16d5237cd3b5f342e44cf0438
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/voc.py
@@ -0,0 +1,31 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import DATASETS
+from .xml_style import XMLDataset
+
+
+@DATASETS.register_module()
+class VOCDataset(XMLDataset):
+ """Dataset for PASCAL VOC."""
+
+ METAINFO = {
+ 'classes':
+ ('aeroplane', 'bicycle', 'bird', 'boat', 'bottle', 'bus', 'car', 'cat',
+ 'chair', 'cow', 'diningtable', 'dog', 'horse', 'motorbike', 'person',
+ 'pottedplant', 'sheep', 'sofa', 'train', 'tvmonitor'),
+ # palette is a list of color tuples, which is used for visualization.
+ 'palette': [(106, 0, 228), (119, 11, 32), (165, 42, 42), (0, 0, 192),
+ (197, 226, 255), (0, 60, 100), (0, 0, 142), (255, 77, 255),
+ (153, 69, 1), (120, 166, 157), (0, 182, 199),
+ (0, 226, 252), (182, 182, 255), (0, 0, 230), (220, 20, 60),
+ (163, 255, 0), (0, 82, 0), (3, 95, 161), (0, 80, 100),
+ (183, 130, 88)]
+ }
+
+ def __init__(self, **kwargs):
+ super().__init__(**kwargs)
+ if 'VOC2007' in self.sub_data_root:
+ self._metainfo['dataset_type'] = 'VOC2007'
+ elif 'VOC2012' in self.sub_data_root:
+ self._metainfo['dataset_type'] = 'VOC2012'
+ else:
+ self._metainfo['dataset_type'] = None
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/wider_face.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/wider_face.py
new file mode 100644
index 0000000000000000000000000000000000000000..62c7fff869ab970b6f96908a998ba6feb25ea205
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/wider_face.py
@@ -0,0 +1,90 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import os.path as osp
+import xml.etree.ElementTree as ET
+
+from mmengine.dist import is_main_process
+from mmengine.fileio import get_local_path, list_from_file
+from mmengine.utils import ProgressBar
+
+from mmdet.registry import DATASETS
+from mmdet.utils.typing_utils import List, Union
+from .xml_style import XMLDataset
+
+
+@DATASETS.register_module()
+class WIDERFaceDataset(XMLDataset):
+ """Reader for the WIDER Face dataset in PASCAL VOC format.
+
+ Conversion scripts can be found in
+ https://github.com/sovrasov/wider-face-pascal-voc-annotations
+ """
+ METAINFO = {'classes': ('face', ), 'palette': [(0, 255, 0)]}
+
+ def load_data_list(self) -> List[dict]:
+ """Load annotation from XML style ann_file.
+
+ Returns:
+ list[dict]: Annotation info from XML file.
+ """
+ assert self._metainfo.get('classes', None) is not None, \
+ 'classes in `XMLDataset` can not be None.'
+ self.cat2label = {
+ cat: i
+ for i, cat in enumerate(self._metainfo['classes'])
+ }
+
+ data_list = []
+ img_ids = list_from_file(self.ann_file, backend_args=self.backend_args)
+
+ # loading process takes around 10 mins
+ if is_main_process():
+ prog_bar = ProgressBar(len(img_ids))
+
+ for img_id in img_ids:
+ raw_img_info = {}
+ raw_img_info['img_id'] = img_id
+ raw_img_info['file_name'] = f'{img_id}.jpg'
+ parsed_data_info = self.parse_data_info(raw_img_info)
+ data_list.append(parsed_data_info)
+
+ if is_main_process():
+ prog_bar.update()
+ return data_list
+
+ def parse_data_info(self, img_info: dict) -> Union[dict, List[dict]]:
+ """Parse raw annotation to target format.
+
+ Args:
+ img_info (dict): Raw image information, usually it includes
+ `img_id`, `file_name`, and `xml_path`.
+
+ Returns:
+ Union[dict, List[dict]]: Parsed annotation.
+ """
+ data_info = {}
+ img_id = img_info['img_id']
+ xml_path = osp.join(self.data_prefix['img'], 'Annotations',
+ f'{img_id}.xml')
+ data_info['img_id'] = img_id
+ data_info['xml_path'] = xml_path
+
+ # deal with xml file
+ with get_local_path(
+ xml_path, backend_args=self.backend_args) as local_path:
+ raw_ann_info = ET.parse(local_path)
+ root = raw_ann_info.getroot()
+ size = root.find('size')
+ width = int(size.find('width').text)
+ height = int(size.find('height').text)
+ folder = root.find('folder').text
+ img_path = osp.join(self.data_prefix['img'], folder,
+ img_info['file_name'])
+ data_info['img_path'] = img_path
+
+ data_info['height'] = height
+ data_info['width'] = width
+
+ # Coordinates are in range [0, width - 1 or height - 1]
+ data_info['instances'] = self._parse_instance_info(
+ raw_ann_info, minus_one=False)
+ return data_info
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/xml_style.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/xml_style.py
new file mode 100644
index 0000000000000000000000000000000000000000..06045ea0092238abdac9622511b336586858f8f5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/xml_style.py
@@ -0,0 +1,186 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import os.path as osp
+import xml.etree.ElementTree as ET
+from typing import List, Optional, Union
+
+import mmcv
+from mmengine.fileio import get, get_local_path, list_from_file
+
+from mmdet.registry import DATASETS
+from .base_det_dataset import BaseDetDataset
+
+
+@DATASETS.register_module()
+class XMLDataset(BaseDetDataset):
+ """XML dataset for detection.
+
+ Args:
+ img_subdir (str): Subdir where images are stored. Default: JPEGImages.
+ ann_subdir (str): Subdir where annotations are. Default: Annotations.
+ backend_args (dict, optional): Arguments to instantiate the
+ corresponding backend. Defaults to None.
+ """
+
+ def __init__(self,
+ img_subdir: str = 'JPEGImages',
+ ann_subdir: str = 'Annotations',
+ **kwargs) -> None:
+ self.img_subdir = img_subdir
+ self.ann_subdir = ann_subdir
+ super().__init__(**kwargs)
+
+ @property
+ def sub_data_root(self) -> str:
+ """Return the sub data root."""
+ return self.data_prefix.get('sub_data_root', '')
+
+ def load_data_list(self) -> List[dict]:
+ """Load annotation from XML style ann_file.
+
+ Returns:
+ list[dict]: Annotation info from XML file.
+ """
+ assert self._metainfo.get('classes', None) is not None, \
+ '`classes` in `XMLDataset` can not be None.'
+ self.cat2label = {
+ cat: i
+ for i, cat in enumerate(self._metainfo['classes'])
+ }
+
+ data_list = []
+ img_ids = list_from_file(self.ann_file, backend_args=self.backend_args)
+ for img_id in img_ids:
+ file_name = osp.join(self.img_subdir, f'{img_id}.jpg')
+ xml_path = osp.join(self.sub_data_root, self.ann_subdir,
+ f'{img_id}.xml')
+
+ raw_img_info = {}
+ raw_img_info['img_id'] = img_id
+ raw_img_info['file_name'] = file_name
+ raw_img_info['xml_path'] = xml_path
+
+ parsed_data_info = self.parse_data_info(raw_img_info)
+ data_list.append(parsed_data_info)
+ return data_list
+
+ @property
+ def bbox_min_size(self) -> Optional[int]:
+ """Return the minimum size of bounding boxes in the images."""
+ if self.filter_cfg is not None:
+ return self.filter_cfg.get('bbox_min_size', None)
+ else:
+ return None
+
+ def parse_data_info(self, img_info: dict) -> Union[dict, List[dict]]:
+ """Parse raw annotation to target format.
+
+ Args:
+ img_info (dict): Raw image information, usually it includes
+ `img_id`, `file_name`, and `xml_path`.
+
+ Returns:
+ Union[dict, List[dict]]: Parsed annotation.
+ """
+ data_info = {}
+ img_path = osp.join(self.sub_data_root, img_info['file_name'])
+ data_info['img_path'] = img_path
+ data_info['img_id'] = img_info['img_id']
+ data_info['xml_path'] = img_info['xml_path']
+
+ # deal with xml file
+ with get_local_path(
+ img_info['xml_path'],
+ backend_args=self.backend_args) as local_path:
+ raw_ann_info = ET.parse(local_path)
+ root = raw_ann_info.getroot()
+ size = root.find('size')
+ if size is not None:
+ width = int(size.find('width').text)
+ height = int(size.find('height').text)
+ else:
+ img_bytes = get(img_path, backend_args=self.backend_args)
+ img = mmcv.imfrombytes(img_bytes, backend='cv2')
+ height, width = img.shape[:2]
+ del img, img_bytes
+
+ data_info['height'] = height
+ data_info['width'] = width
+
+ data_info['instances'] = self._parse_instance_info(
+ raw_ann_info, minus_one=True)
+
+ return data_info
+
+ def _parse_instance_info(self,
+ raw_ann_info: ET,
+ minus_one: bool = True) -> List[dict]:
+ """parse instance information.
+
+ Args:
+ raw_ann_info (ElementTree): ElementTree object.
+ minus_one (bool): Whether to subtract 1 from the coordinates.
+ Defaults to True.
+
+ Returns:
+ List[dict]: List of instances.
+ """
+ instances = []
+ for obj in raw_ann_info.findall('object'):
+ instance = {}
+ name = obj.find('name').text
+ if name not in self._metainfo['classes']:
+ continue
+ difficult = obj.find('difficult')
+ difficult = 0 if difficult is None else int(difficult.text)
+ bnd_box = obj.find('bndbox')
+ bbox = [
+ int(float(bnd_box.find('xmin').text)),
+ int(float(bnd_box.find('ymin').text)),
+ int(float(bnd_box.find('xmax').text)),
+ int(float(bnd_box.find('ymax').text))
+ ]
+
+ # VOC needs to subtract 1 from the coordinates
+ if minus_one:
+ bbox = [x - 1 for x in bbox]
+
+ ignore = False
+ if self.bbox_min_size is not None:
+ assert not self.test_mode
+ w = bbox[2] - bbox[0]
+ h = bbox[3] - bbox[1]
+ if w < self.bbox_min_size or h < self.bbox_min_size:
+ ignore = True
+ if difficult or ignore:
+ instance['ignore_flag'] = 1
+ else:
+ instance['ignore_flag'] = 0
+ instance['bbox'] = bbox
+ instance['bbox_label'] = self.cat2label[name]
+ instances.append(instance)
+ return instances
+
+ def filter_data(self) -> List[dict]:
+ """Filter annotations according to filter_cfg.
+
+ Returns:
+ List[dict]: Filtered results.
+ """
+ if self.test_mode:
+ return self.data_list
+
+ filter_empty_gt = self.filter_cfg.get('filter_empty_gt', False) \
+ if self.filter_cfg is not None else False
+ min_size = self.filter_cfg.get('min_size', 0) \
+ if self.filter_cfg is not None else 0
+
+ valid_data_infos = []
+ for i, data_info in enumerate(self.data_list):
+ width = data_info['width']
+ height = data_info['height']
+ if filter_empty_gt and len(data_info['instances']) == 0:
+ continue
+ if min(width, height) >= min_size:
+ valid_data_infos.append(data_info)
+
+ return valid_data_infos
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/youtube_vis_dataset.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/youtube_vis_dataset.py
new file mode 100644
index 0000000000000000000000000000000000000000..38c3d3909f1b8fd795c181546094056c54c9c4b2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/datasets/youtube_vis_dataset.py
@@ -0,0 +1,52 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import DATASETS
+from .base_video_dataset import BaseVideoDataset
+
+
+@DATASETS.register_module()
+class YouTubeVISDataset(BaseVideoDataset):
+ """YouTube VIS dataset for video instance segmentation.
+
+ Args:
+ dataset_version (str): Select dataset year version.
+ """
+
+ def __init__(self, dataset_version: str, *args, **kwargs):
+ self.set_dataset_classes(dataset_version)
+ super().__init__(*args, **kwargs)
+
+ @classmethod
+ def set_dataset_classes(cls, dataset_version: str) -> None:
+ """Pass the category of the corresponding year to metainfo.
+
+ Args:
+ dataset_version (str): Select dataset year version.
+ """
+ classes_2019_version = ('person', 'giant_panda', 'lizard', 'parrot',
+ 'skateboard', 'sedan', 'ape', 'dog', 'snake',
+ 'monkey', 'hand', 'rabbit', 'duck', 'cat',
+ 'cow', 'fish', 'train', 'horse', 'turtle',
+ 'bear', 'motorbike', 'giraffe', 'leopard',
+ 'fox', 'deer', 'owl', 'surfboard', 'airplane',
+ 'truck', 'zebra', 'tiger', 'elephant',
+ 'snowboard', 'boat', 'shark', 'mouse', 'frog',
+ 'eagle', 'earless_seal', 'tennis_racket')
+
+ classes_2021_version = ('airplane', 'bear', 'bird', 'boat', 'car',
+ 'cat', 'cow', 'deer', 'dog', 'duck',
+ 'earless_seal', 'elephant', 'fish',
+ 'flying_disc', 'fox', 'frog', 'giant_panda',
+ 'giraffe', 'horse', 'leopard', 'lizard',
+ 'monkey', 'motorbike', 'mouse', 'parrot',
+ 'person', 'rabbit', 'shark', 'skateboard',
+ 'snake', 'snowboard', 'squirrel', 'surfboard',
+ 'tennis_racket', 'tiger', 'train', 'truck',
+ 'turtle', 'whale', 'zebra')
+
+ if dataset_version == '2019':
+ cls.METAINFO = dict(classes=classes_2019_version)
+ elif dataset_version == '2021':
+ cls.METAINFO = dict(classes=classes_2021_version)
+ else:
+ raise NotImplementedError('Not supported YouTubeVIS dataset'
+ f'version: {dataset_version}')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..c91ace6ffa20948af572d3a0fd594e8a0b091775
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/__init__.py
@@ -0,0 +1,5 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .hooks import * # noqa: F401, F403
+from .optimizers import * # noqa: F401, F403
+from .runner import * # noqa: F401, F403
+from .schedulers import * # noqa: F401, F403
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..889fa557adef87e2251c625a7353503226beb079
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/__init__.py
@@ -0,0 +1,21 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .checkloss_hook import CheckInvalidLossHook
+from .mean_teacher_hook import MeanTeacherHook
+from .memory_profiler_hook import MemoryProfilerHook
+from .num_class_check_hook import NumClassCheckHook
+from .pipeline_switch_hook import PipelineSwitchHook
+from .set_epoch_info_hook import SetEpochInfoHook
+from .sync_norm_hook import SyncNormHook
+from .utils import trigger_visualization_hook
+from .visualization_hook import (DetVisualizationHook,
+ GroundingVisualizationHook,
+ TrackVisualizationHook)
+from .yolox_mode_switch_hook import YOLOXModeSwitchHook
+
+__all__ = [
+ 'YOLOXModeSwitchHook', 'SyncNormHook', 'CheckInvalidLossHook',
+ 'SetEpochInfoHook', 'MemoryProfilerHook', 'DetVisualizationHook',
+ 'NumClassCheckHook', 'MeanTeacherHook', 'trigger_visualization_hook',
+ 'PipelineSwitchHook', 'TrackVisualizationHook',
+ 'GroundingVisualizationHook'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/checkloss_hook.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/checkloss_hook.py
new file mode 100644
index 0000000000000000000000000000000000000000..3ebfcd5dfcd7ae329399723d3a9c0fc0a0d722ef
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/checkloss_hook.py
@@ -0,0 +1,42 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional
+
+import torch
+from mmengine.hooks import Hook
+from mmengine.runner import Runner
+
+from mmdet.registry import HOOKS
+
+
+@HOOKS.register_module()
+class CheckInvalidLossHook(Hook):
+ """Check invalid loss hook.
+
+ This hook will regularly check whether the loss is valid
+ during training.
+
+ Args:
+ interval (int): Checking interval (every k iterations).
+ Default: 50.
+ """
+
+ def __init__(self, interval: int = 50) -> None:
+ self.interval = interval
+
+ def after_train_iter(self,
+ runner: Runner,
+ batch_idx: int,
+ data_batch: Optional[dict] = None,
+ outputs: Optional[dict] = None) -> None:
+ """Regularly check whether the loss is valid every n iterations.
+
+ Args:
+ runner (:obj:`Runner`): The runner of the training process.
+ batch_idx (int): The index of the current batch in the train loop.
+ data_batch (dict, Optional): Data from dataloader.
+ Defaults to None.
+ outputs (dict, Optional): Outputs from model. Defaults to None.
+ """
+ if self.every_n_train_iters(runner, self.interval):
+ assert torch.isfinite(outputs['loss']), \
+ runner.logger.info('loss become infinite or NaN!')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/mean_teacher_hook.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/mean_teacher_hook.py
new file mode 100644
index 0000000000000000000000000000000000000000..b924c0a5934248d05e7ce1add50e7574b739b9c7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/mean_teacher_hook.py
@@ -0,0 +1,87 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional
+
+import torch.nn as nn
+from mmengine.hooks import Hook
+from mmengine.model import is_model_wrapper
+from mmengine.runner import Runner
+
+from mmdet.registry import HOOKS
+
+
+@HOOKS.register_module()
+class MeanTeacherHook(Hook):
+ """Mean Teacher Hook.
+
+ Mean Teacher is an efficient semi-supervised learning method in
+ `Mean Teacher `_.
+ This method requires two models with exactly the same structure,
+ as the student model and the teacher model, respectively.
+ The student model updates the parameters through gradient descent,
+ and the teacher model updates the parameters through
+ exponential moving average of the student model.
+ Compared with the student model, the teacher model
+ is smoother and accumulates more knowledge.
+
+ Args:
+ momentum (float): The momentum used for updating teacher's parameter.
+ Teacher's parameter are updated with the formula:
+ `teacher = (1-momentum) * teacher + momentum * student`.
+ Defaults to 0.001.
+ interval (int): Update teacher's parameter every interval iteration.
+ Defaults to 1.
+ skip_buffers (bool): Whether to skip the model buffers, such as
+ batchnorm running stats (running_mean, running_var), it does not
+ perform the ema operation. Default to True.
+ """
+
+ def __init__(self,
+ momentum: float = 0.001,
+ interval: int = 1,
+ skip_buffer=True) -> None:
+ assert 0 < momentum < 1
+ self.momentum = momentum
+ self.interval = interval
+ self.skip_buffers = skip_buffer
+
+ def before_train(self, runner: Runner) -> None:
+ """To check that teacher model and student model exist."""
+ model = runner.model
+ if is_model_wrapper(model):
+ model = model.module
+ assert hasattr(model, 'teacher')
+ assert hasattr(model, 'student')
+ # only do it at initial stage
+ if runner.iter == 0:
+ self.momentum_update(model, 1)
+
+ def after_train_iter(self,
+ runner: Runner,
+ batch_idx: int,
+ data_batch: Optional[dict] = None,
+ outputs: Optional[dict] = None) -> None:
+ """Update teacher's parameter every self.interval iterations."""
+ if (runner.iter + 1) % self.interval != 0:
+ return
+ model = runner.model
+ if is_model_wrapper(model):
+ model = model.module
+ self.momentum_update(model, self.momentum)
+
+ def momentum_update(self, model: nn.Module, momentum: float) -> None:
+ """Compute the moving average of the parameters using exponential
+ moving average."""
+ if self.skip_buffers:
+ for (src_name, src_parm), (dst_name, dst_parm) in zip(
+ model.student.named_parameters(),
+ model.teacher.named_parameters()):
+ dst_parm.data.mul_(1 - momentum).add_(
+ src_parm.data, alpha=momentum)
+ else:
+ for (src_parm,
+ dst_parm) in zip(model.student.state_dict().values(),
+ model.teacher.state_dict().values()):
+ # exclude num_tracking
+ if dst_parm.dtype.is_floating_point:
+ dst_parm.data.mul_(1 - momentum).add_(
+ src_parm.data, alpha=momentum)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/memory_profiler_hook.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/memory_profiler_hook.py
new file mode 100644
index 0000000000000000000000000000000000000000..3dcdcae0b669ade46026d28c46b35f35d90b504b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/memory_profiler_hook.py
@@ -0,0 +1,121 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Sequence
+
+from mmengine.hooks import Hook
+from mmengine.runner import Runner
+
+from mmdet.registry import HOOKS
+from mmdet.structures import DetDataSample
+
+
+@HOOKS.register_module()
+class MemoryProfilerHook(Hook):
+ """Memory profiler hook recording memory information including virtual
+ memory, swap memory, and the memory of the current process.
+
+ Args:
+ interval (int): Checking interval (every k iterations).
+ Default: 50.
+ """
+
+ def __init__(self, interval: int = 50) -> None:
+ try:
+ from psutil import swap_memory, virtual_memory
+ self._swap_memory = swap_memory
+ self._virtual_memory = virtual_memory
+ except ImportError:
+ raise ImportError('psutil is not installed, please install it by: '
+ 'pip install psutil')
+
+ try:
+ from memory_profiler import memory_usage
+ self._memory_usage = memory_usage
+ except ImportError:
+ raise ImportError(
+ 'memory_profiler is not installed, please install it by: '
+ 'pip install memory_profiler')
+
+ self.interval = interval
+
+ def _record_memory_information(self, runner: Runner) -> None:
+ """Regularly record memory information.
+
+ Args:
+ runner (:obj:`Runner`): The runner of the training or evaluation
+ process.
+ """
+ # in Byte
+ virtual_memory = self._virtual_memory()
+ swap_memory = self._swap_memory()
+ # in MB
+ process_memory = self._memory_usage()[0]
+ factor = 1024 * 1024
+ runner.logger.info(
+ 'Memory information '
+ 'available_memory: '
+ f'{round(virtual_memory.available / factor)} MB, '
+ 'used_memory: '
+ f'{round(virtual_memory.used / factor)} MB, '
+ f'memory_utilization: {virtual_memory.percent} %, '
+ 'available_swap_memory: '
+ f'{round((swap_memory.total - swap_memory.used) / factor)}'
+ ' MB, '
+ f'used_swap_memory: {round(swap_memory.used / factor)} MB, '
+ f'swap_memory_utilization: {swap_memory.percent} %, '
+ 'current_process_memory: '
+ f'{round(process_memory)} MB')
+
+ def after_train_iter(self,
+ runner: Runner,
+ batch_idx: int,
+ data_batch: Optional[dict] = None,
+ outputs: Optional[dict] = None) -> None:
+ """Regularly record memory information.
+
+ Args:
+ runner (:obj:`Runner`): The runner of the training process.
+ batch_idx (int): The index of the current batch in the train loop.
+ data_batch (dict, optional): Data from dataloader.
+ Defaults to None.
+ outputs (dict, optional): Outputs from model. Defaults to None.
+ """
+ if self.every_n_inner_iters(batch_idx, self.interval):
+ self._record_memory_information(runner)
+
+ def after_val_iter(
+ self,
+ runner: Runner,
+ batch_idx: int,
+ data_batch: Optional[dict] = None,
+ outputs: Optional[Sequence[DetDataSample]] = None) -> None:
+ """Regularly record memory information.
+
+ Args:
+ runner (:obj:`Runner`): The runner of the validation process.
+ batch_idx (int): The index of the current batch in the val loop.
+ data_batch (dict, optional): Data from dataloader.
+ Defaults to None.
+ outputs (Sequence[:obj:`DetDataSample`], optional):
+ Outputs from model. Defaults to None.
+ """
+ if self.every_n_inner_iters(batch_idx, self.interval):
+ self._record_memory_information(runner)
+
+ def after_test_iter(
+ self,
+ runner: Runner,
+ batch_idx: int,
+ data_batch: Optional[dict] = None,
+ outputs: Optional[Sequence[DetDataSample]] = None) -> None:
+ """Regularly record memory information.
+
+ Args:
+ runner (:obj:`Runner`): The runner of the testing process.
+ batch_idx (int): The index of the current batch in the test loop.
+ data_batch (dict, optional): Data from dataloader.
+ Defaults to None.
+ outputs (Sequence[:obj:`DetDataSample`], optional):
+ Outputs from model. Defaults to None.
+ """
+ if self.every_n_inner_iters(batch_idx, self.interval):
+ self._record_memory_information(runner)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/num_class_check_hook.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/num_class_check_hook.py
new file mode 100644
index 0000000000000000000000000000000000000000..6588473acfbd3ffe8e80eb163aa7ee449332e6b8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/num_class_check_hook.py
@@ -0,0 +1,68 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.cnn import VGG
+from mmengine.hooks import Hook
+from mmengine.runner import Runner
+
+from mmdet.registry import HOOKS
+
+
+@HOOKS.register_module()
+class NumClassCheckHook(Hook):
+ """Check whether the `num_classes` in head matches the length of `classes`
+ in `dataset.metainfo`."""
+
+ def _check_head(self, runner: Runner, mode: str) -> None:
+ """Check whether the `num_classes` in head matches the length of
+ `classes` in `dataset.metainfo`.
+
+ Args:
+ runner (:obj:`Runner`): The runner of the training or evaluation
+ process.
+ """
+ assert mode in ['train', 'val']
+ model = runner.model
+ dataset = runner.train_dataloader.dataset if mode == 'train' else \
+ runner.val_dataloader.dataset
+ if dataset.metainfo.get('classes', None) is None:
+ runner.logger.warning(
+ f'Please set `classes` '
+ f'in the {dataset.__class__.__name__} `metainfo` and'
+ f'check if it is consistent with the `num_classes` '
+ f'of head')
+ else:
+ classes = dataset.metainfo['classes']
+ assert type(classes) is not str, \
+ (f'`classes` in {dataset.__class__.__name__}'
+ f'should be a tuple of str.'
+ f'Add comma if number of classes is 1 as '
+ f'classes = ({classes},)')
+ from mmdet.models.roi_heads.mask_heads import FusedSemanticHead
+ for name, module in model.named_modules():
+ if hasattr(module, 'num_classes') and not name.endswith(
+ 'rpn_head') and not isinstance(
+ module, (VGG, FusedSemanticHead)):
+ assert module.num_classes == len(classes), \
+ (f'The `num_classes` ({module.num_classes}) in '
+ f'{module.__class__.__name__} of '
+ f'{model.__class__.__name__} does not matches '
+ f'the length of `classes` '
+ f'{len(classes)}) in '
+ f'{dataset.__class__.__name__}')
+
+ def before_train_epoch(self, runner: Runner) -> None:
+ """Check whether the training dataset is compatible with head.
+
+ Args:
+ runner (:obj:`Runner`): The runner of the training or evaluation
+ process.
+ """
+ self._check_head(runner, 'train')
+
+ def before_val_epoch(self, runner: Runner) -> None:
+ """Check whether the dataset in val epoch is compatible with head.
+
+ Args:
+ runner (:obj:`Runner`): The runner of the training or evaluation
+ process.
+ """
+ self._check_head(runner, 'val')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/pipeline_switch_hook.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/pipeline_switch_hook.py
new file mode 100644
index 0000000000000000000000000000000000000000..a5abd897803b11793ebace86e45aac8f59938545
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/pipeline_switch_hook.py
@@ -0,0 +1,43 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.transforms import Compose
+from mmengine.hooks import Hook
+
+from mmdet.registry import HOOKS
+
+
+@HOOKS.register_module()
+class PipelineSwitchHook(Hook):
+ """Switch data pipeline at switch_epoch.
+
+ Args:
+ switch_epoch (int): switch pipeline at this epoch.
+ switch_pipeline (list[dict]): the pipeline to switch to.
+ """
+
+ def __init__(self, switch_epoch, switch_pipeline):
+ self.switch_epoch = switch_epoch
+ self.switch_pipeline = switch_pipeline
+ self._restart_dataloader = False
+ self._has_switched = False
+
+ def before_train_epoch(self, runner):
+ """switch pipeline."""
+ epoch = runner.epoch
+ train_loader = runner.train_dataloader
+ if epoch >= self.switch_epoch and not self._has_switched:
+ runner.logger.info('Switch pipeline now!')
+ # The dataset pipeline cannot be updated when persistent_workers
+ # is True, so we need to force the dataloader's multi-process
+ # restart. This is a very hacky approach.
+ train_loader.dataset.pipeline = Compose(self.switch_pipeline)
+ if hasattr(train_loader, 'persistent_workers'
+ ) and train_loader.persistent_workers is True:
+ train_loader._DataLoader__initialized = False
+ train_loader._iterator = None
+ self._restart_dataloader = True
+ self._has_switched = True
+ else:
+ # Once the restart is complete, we need to restore
+ # the initialization flag.
+ if self._restart_dataloader:
+ train_loader._DataLoader__initialized = True
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/set_epoch_info_hook.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/set_epoch_info_hook.py
new file mode 100644
index 0000000000000000000000000000000000000000..183f3167445dc0818e4fa37bdd2049d3876ed031
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/set_epoch_info_hook.py
@@ -0,0 +1,17 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.hooks import Hook
+from mmengine.model.wrappers import is_model_wrapper
+
+from mmdet.registry import HOOKS
+
+
+@HOOKS.register_module()
+class SetEpochInfoHook(Hook):
+ """Set runner's epoch information to the model."""
+
+ def before_train_epoch(self, runner):
+ epoch = runner.epoch
+ model = runner.model
+ if is_model_wrapper(model):
+ model = model.module
+ model.set_epoch(epoch)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/sync_norm_hook.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/sync_norm_hook.py
new file mode 100644
index 0000000000000000000000000000000000000000..a1734380c83157c911568098abfce761fb3c9a1f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/sync_norm_hook.py
@@ -0,0 +1,37 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from collections import OrderedDict
+
+from mmengine.dist import get_dist_info
+from mmengine.hooks import Hook
+from torch import nn
+
+from mmdet.registry import HOOKS
+from mmdet.utils import all_reduce_dict
+
+
+def get_norm_states(module: nn.Module) -> OrderedDict:
+ """Get the state_dict of batch norms in the module."""
+ async_norm_states = OrderedDict()
+ for name, child in module.named_modules():
+ if isinstance(child, nn.modules.batchnorm._NormBase):
+ for k, v in child.state_dict().items():
+ async_norm_states['.'.join([name, k])] = v
+ return async_norm_states
+
+
+@HOOKS.register_module()
+class SyncNormHook(Hook):
+ """Synchronize Norm states before validation, currently used in YOLOX."""
+
+ def before_val_epoch(self, runner):
+ """Synchronizing norm."""
+ module = runner.model
+ _, world_size = get_dist_info()
+ if world_size == 1:
+ return
+ norm_states = get_norm_states(module)
+ if len(norm_states) == 0:
+ return
+ # TODO: use `all_reduce_dict` in mmengine
+ norm_states = all_reduce_dict(norm_states, op='mean')
+ module.load_state_dict(norm_states, strict=False)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/utils.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/utils.py
new file mode 100644
index 0000000000000000000000000000000000000000..d267cfe77be163c0520568b7b7936f4453914aab
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/utils.py
@@ -0,0 +1,19 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+def trigger_visualization_hook(cfg, args):
+ default_hooks = cfg.default_hooks
+ if 'visualization' in default_hooks:
+ visualization_hook = default_hooks['visualization']
+ # Turn on visualization
+ visualization_hook['draw'] = True
+ if args.show:
+ visualization_hook['show'] = True
+ visualization_hook['wait_time'] = args.wait_time
+ if args.show_dir:
+ visualization_hook['test_out_dir'] = args.show_dir
+ else:
+ raise RuntimeError(
+ 'VisualizationHook must be included in default_hooks.'
+ 'refer to usage '
+ '"visualization=dict(type=\'VisualizationHook\')"')
+
+ return cfg
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/visualization_hook.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/visualization_hook.py
new file mode 100644
index 0000000000000000000000000000000000000000..3408186b6ef9c4195745b0c740519541572d27d2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/visualization_hook.py
@@ -0,0 +1,515 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import os.path as osp
+import warnings
+from typing import Optional, Sequence
+
+import mmcv
+import numpy as np
+from mmengine.fileio import get
+from mmengine.hooks import Hook
+from mmengine.runner import Runner
+from mmengine.utils import mkdir_or_exist
+from mmengine.visualization import Visualizer
+
+from mmdet.datasets.samplers import TrackImgSampler
+from mmdet.registry import HOOKS
+from mmdet.structures import DetDataSample, TrackDataSample
+from mmdet.structures.bbox import BaseBoxes
+from mmdet.visualization.palette import _get_adaptive_scales
+
+
+@HOOKS.register_module()
+class DetVisualizationHook(Hook):
+ """Detection Visualization Hook. Used to visualize validation and testing
+ process prediction results.
+
+ In the testing phase:
+
+ 1. If ``show`` is True, it means that only the prediction results are
+ visualized without storing data, so ``vis_backends`` needs to
+ be excluded.
+ 2. If ``test_out_dir`` is specified, it means that the prediction results
+ need to be saved to ``test_out_dir``. In order to avoid vis_backends
+ also storing data, so ``vis_backends`` needs to be excluded.
+ 3. ``vis_backends`` takes effect if the user does not specify ``show``
+ and `test_out_dir``. You can set ``vis_backends`` to WandbVisBackend or
+ TensorboardVisBackend to store the prediction result in Wandb or
+ Tensorboard.
+
+ Args:
+ draw (bool): whether to draw prediction results. If it is False,
+ it means that no drawing will be done. Defaults to False.
+ interval (int): The interval of visualization. Defaults to 50.
+ score_thr (float): The threshold to visualize the bboxes
+ and masks. Defaults to 0.3.
+ show (bool): Whether to display the drawn image. Default to False.
+ wait_time (float): The interval of show (s). Defaults to 0.
+ test_out_dir (str, optional): directory where painted images
+ will be saved in testing process.
+ backend_args (dict, optional): Arguments to instantiate the
+ corresponding backend. Defaults to None.
+ """
+
+ def __init__(self,
+ draw: bool = False,
+ interval: int = 50,
+ score_thr: float = 0.3,
+ show: bool = False,
+ wait_time: float = 0.,
+ test_out_dir: Optional[str] = None,
+ backend_args: dict = None):
+ self._visualizer: Visualizer = Visualizer.get_current_instance()
+ self.interval = interval
+ self.score_thr = score_thr
+ self.show = show
+ if self.show:
+ # No need to think about vis backends.
+ self._visualizer._vis_backends = {}
+ warnings.warn('The show is True, it means that only '
+ 'the prediction results are visualized '
+ 'without storing data, so vis_backends '
+ 'needs to be excluded.')
+
+ self.wait_time = wait_time
+ self.backend_args = backend_args
+ self.draw = draw
+ self.test_out_dir = test_out_dir
+ self._test_index = 0
+
+ def after_val_iter(self, runner: Runner, batch_idx: int, data_batch: dict,
+ outputs: Sequence[DetDataSample]) -> None:
+ """Run after every ``self.interval`` validation iterations.
+
+ Args:
+ runner (:obj:`Runner`): The runner of the validation process.
+ batch_idx (int): The index of the current batch in the val loop.
+ data_batch (dict): Data from dataloader.
+ outputs (Sequence[:obj:`DetDataSample`]]): A batch of data samples
+ that contain annotations and predictions.
+ """
+ if self.draw is False:
+ return
+
+ # There is no guarantee that the same batch of images
+ # is visualized for each evaluation.
+ total_curr_iter = runner.iter + batch_idx
+
+ # Visualize only the first data
+ img_path = outputs[0].img_path
+ img_bytes = get(img_path, backend_args=self.backend_args)
+ img = mmcv.imfrombytes(img_bytes, channel_order='rgb')
+
+ if total_curr_iter % self.interval == 0:
+ self._visualizer.add_datasample(
+ osp.basename(img_path) if self.show else 'val_img',
+ img,
+ data_sample=outputs[0],
+ show=self.show,
+ wait_time=self.wait_time,
+ pred_score_thr=self.score_thr,
+ step=total_curr_iter)
+
+ def after_test_iter(self, runner: Runner, batch_idx: int, data_batch: dict,
+ outputs: Sequence[DetDataSample]) -> None:
+ """Run after every testing iterations.
+
+ Args:
+ runner (:obj:`Runner`): The runner of the testing process.
+ batch_idx (int): The index of the current batch in the val loop.
+ data_batch (dict): Data from dataloader.
+ outputs (Sequence[:obj:`DetDataSample`]): A batch of data samples
+ that contain annotations and predictions.
+ """
+ if self.draw is False:
+ return
+
+ if self.test_out_dir is not None:
+ self.test_out_dir = osp.join(runner.work_dir, runner.timestamp,
+ self.test_out_dir)
+ mkdir_or_exist(self.test_out_dir)
+
+ for data_sample in outputs:
+ self._test_index += 1
+
+ img_path = data_sample.img_path
+ img_bytes = get(img_path, backend_args=self.backend_args)
+ img = mmcv.imfrombytes(img_bytes, channel_order='rgb')
+
+ out_file = None
+ if self.test_out_dir is not None:
+ out_file = osp.basename(img_path)
+ out_file = osp.join(self.test_out_dir, out_file)
+
+ self._visualizer.add_datasample(
+ osp.basename(img_path) if self.show else 'test_img',
+ img,
+ data_sample=data_sample,
+ show=self.show,
+ wait_time=self.wait_time,
+ pred_score_thr=self.score_thr,
+ out_file=out_file,
+ step=self._test_index)
+
+
+@HOOKS.register_module()
+class TrackVisualizationHook(Hook):
+ """Tracking Visualization Hook. Used to visualize validation and testing
+ process prediction results.
+
+ In the testing phase:
+
+ 1. If ``show`` is True, it means that only the prediction results are
+ visualized without storing data, so ``vis_backends`` needs to
+ be excluded.
+ 2. If ``test_out_dir`` is specified, it means that the prediction results
+ need to be saved to ``test_out_dir``. In order to avoid vis_backends
+ also storing data, so ``vis_backends`` needs to be excluded.
+ 3. ``vis_backends`` takes effect if the user does not specify ``show``
+ and `test_out_dir``. You can set ``vis_backends`` to WandbVisBackend or
+ TensorboardVisBackend to store the prediction result in Wandb or
+ Tensorboard.
+
+ Args:
+ draw (bool): whether to draw prediction results. If it is False,
+ it means that no drawing will be done. Defaults to False.
+ frame_interval (int): The interval of visualization. Defaults to 30.
+ score_thr (float): The threshold to visualize the bboxes
+ and masks. Defaults to 0.3.
+ show (bool): Whether to display the drawn image. Default to False.
+ wait_time (float): The interval of show (s). Defaults to 0.
+ test_out_dir (str, optional): directory where painted images
+ will be saved in testing process.
+ backend_args (dict): Arguments to instantiate a file client.
+ Defaults to ``None``.
+ """
+
+ def __init__(self,
+ draw: bool = False,
+ frame_interval: int = 30,
+ score_thr: float = 0.3,
+ show: bool = False,
+ wait_time: float = 0.,
+ test_out_dir: Optional[str] = None,
+ backend_args: dict = None) -> None:
+ self._visualizer: Visualizer = Visualizer.get_current_instance()
+ self.frame_interval = frame_interval
+ self.score_thr = score_thr
+ self.show = show
+ if self.show:
+ # No need to think about vis backends.
+ self._visualizer._vis_backends = {}
+ warnings.warn('The show is True, it means that only '
+ 'the prediction results are visualized '
+ 'without storing data, so vis_backends '
+ 'needs to be excluded.')
+
+ self.wait_time = wait_time
+ self.backend_args = backend_args
+ self.draw = draw
+ self.test_out_dir = test_out_dir
+ self.image_idx = 0
+
+ def after_val_iter(self, runner: Runner, batch_idx: int, data_batch: dict,
+ outputs: Sequence[TrackDataSample]) -> None:
+ """Run after every ``self.interval`` validation iteration.
+
+ Args:
+ runner (:obj:`Runner`): The runner of the validation process.
+ batch_idx (int): The index of the current batch in the val loop.
+ data_batch (dict): Data from dataloader.
+ outputs (Sequence[:obj:`TrackDataSample`]): Outputs from model.
+ """
+ if self.draw is False:
+ return
+
+ assert len(outputs) == 1, \
+ 'only batch_size=1 is supported while validating.'
+
+ sampler = runner.val_dataloader.sampler
+ if isinstance(sampler, TrackImgSampler):
+ if self.every_n_inner_iters(batch_idx, self.frame_interval):
+ total_curr_iter = runner.iter + batch_idx
+ track_data_sample = outputs[0]
+ self.visualize_single_image(track_data_sample[0],
+ total_curr_iter)
+ else:
+ # video visualization DefaultSampler
+ if self.every_n_inner_iters(batch_idx, 1):
+ track_data_sample = outputs[0]
+ video_length = len(track_data_sample)
+
+ for frame_id in range(video_length):
+ if frame_id % self.frame_interval == 0:
+ total_curr_iter = runner.iter + self.image_idx + \
+ frame_id
+ img_data_sample = track_data_sample[frame_id]
+ self.visualize_single_image(img_data_sample,
+ total_curr_iter)
+ self.image_idx = self.image_idx + video_length
+
+ def after_test_iter(self, runner: Runner, batch_idx: int, data_batch: dict,
+ outputs: Sequence[TrackDataSample]) -> None:
+ """Run after every testing iteration.
+
+ Args:
+ runner (:obj:`Runner`): The runner of the testing process.
+ batch_idx (int): The index of the current batch in the test loop.
+ data_batch (dict): Data from dataloader.
+ outputs (Sequence[:obj:`TrackDataSample`]): Outputs from model.
+ """
+ if self.draw is False:
+ return
+
+ assert len(outputs) == 1, \
+ 'only batch_size=1 is supported while testing.'
+
+ if self.test_out_dir is not None:
+ self.test_out_dir = osp.join(runner.work_dir, runner.timestamp,
+ self.test_out_dir)
+ mkdir_or_exist(self.test_out_dir)
+
+ sampler = runner.test_dataloader.sampler
+ if isinstance(sampler, TrackImgSampler):
+ if self.every_n_inner_iters(batch_idx, self.frame_interval):
+ track_data_sample = outputs[0]
+ self.visualize_single_image(track_data_sample[0], batch_idx)
+ else:
+ # video visualization DefaultSampler
+ if self.every_n_inner_iters(batch_idx, 1):
+ track_data_sample = outputs[0]
+ video_length = len(track_data_sample)
+
+ for frame_id in range(video_length):
+ if frame_id % self.frame_interval == 0:
+ img_data_sample = track_data_sample[frame_id]
+ self.visualize_single_image(img_data_sample,
+ self.image_idx + frame_id)
+ self.image_idx = self.image_idx + video_length
+
+ def visualize_single_image(self, img_data_sample: DetDataSample,
+ step: int) -> None:
+ """
+ Args:
+ img_data_sample (DetDataSample): single image output.
+ step (int): The index of the current image.
+ """
+ img_path = img_data_sample.img_path
+ img_bytes = get(img_path, backend_args=self.backend_args)
+ img = mmcv.imfrombytes(img_bytes, channel_order='rgb')
+
+ out_file = None
+ if self.test_out_dir is not None:
+ video_name = img_path.split('/')[-3]
+ mkdir_or_exist(osp.join(self.test_out_dir, video_name))
+ out_file = osp.join(self.test_out_dir, video_name,
+ osp.basename(img_path))
+
+ self._visualizer.add_datasample(
+ osp.basename(img_path) if self.show else 'test_img',
+ img,
+ data_sample=img_data_sample,
+ show=self.show,
+ wait_time=self.wait_time,
+ pred_score_thr=self.score_thr,
+ out_file=out_file,
+ step=step)
+
+
+def draw_all_character(visualizer, characters, w):
+ start_index = 2
+ y_index = 5
+ for char in characters:
+ if isinstance(char, str):
+ visualizer.draw_texts(
+ str(char),
+ positions=np.array([start_index, y_index]),
+ colors=(0, 0, 0),
+ font_families='monospace')
+ start_index += len(char) * 8
+ else:
+ visualizer.draw_texts(
+ str(char[0]),
+ positions=np.array([start_index, y_index]),
+ colors=char[1],
+ font_families='monospace')
+ start_index += len(char[0]) * 8
+
+ if start_index > w - 10:
+ start_index = 2
+ y_index += 15
+
+ drawn_text = visualizer.get_image()
+ return drawn_text
+
+
+@HOOKS.register_module()
+class GroundingVisualizationHook(DetVisualizationHook):
+
+ def after_test_iter(self, runner: Runner, batch_idx: int, data_batch: dict,
+ outputs: Sequence[DetDataSample]) -> None:
+ """Run after every testing iterations.
+
+ Args:
+ runner (:obj:`Runner`): The runner of the testing process.
+ batch_idx (int): The index of the current batch in the val loop.
+ data_batch (dict): Data from dataloader.
+ outputs (Sequence[:obj:`DetDataSample`]): A batch of data samples
+ that contain annotations and predictions.
+ """
+ if self.draw is False:
+ return
+
+ if self.test_out_dir is not None:
+ self.test_out_dir = osp.join(runner.work_dir, runner.timestamp,
+ self.test_out_dir)
+ mkdir_or_exist(self.test_out_dir)
+
+ for data_sample in outputs:
+ data_sample = data_sample.cpu()
+
+ self._test_index += 1
+
+ img_path = data_sample.img_path
+ img_bytes = get(img_path, backend_args=self.backend_args)
+ img = mmcv.imfrombytes(img_bytes, channel_order='rgb')
+
+ out_file = None
+ if self.test_out_dir is not None:
+ out_file = osp.basename(img_path)
+ out_file = osp.join(self.test_out_dir, out_file)
+
+ text = data_sample.text
+ if isinstance(text, str): # VG
+ gt_instances = data_sample.gt_instances
+ tokens_positive = data_sample.tokens_positive
+ if 'phrase_ids' in data_sample:
+ # flickr30k
+ gt_labels = data_sample.phrase_ids
+ else:
+ gt_labels = gt_instances.labels
+ gt_bboxes = gt_instances.get('bboxes', None)
+ if gt_bboxes is not None and isinstance(gt_bboxes, BaseBoxes):
+ gt_instances.bboxes = gt_bboxes.tensor
+ print(gt_labels, tokens_positive, gt_bboxes, img_path)
+ pred_instances = data_sample.pred_instances
+ pred_instances = pred_instances[
+ pred_instances.scores > self.score_thr]
+ pred_labels = pred_instances.labels
+ pred_bboxes = pred_instances.bboxes
+ pred_scores = pred_instances.scores
+
+ max_label = 0
+ if len(gt_labels) > 0:
+ max_label = max(gt_labels)
+ if len(pred_labels) > 0:
+ max_label = max(max(pred_labels), max_label)
+
+ max_label = int(max(max_label, 0))
+ palette = np.random.randint(0, 256, size=(max_label + 1, 3))
+ bbox_palette = [tuple(c) for c in palette]
+ # bbox_palette = get_palette('random', max_label + 1)
+ if len(gt_labels) >= len(pred_labels):
+ colors = [bbox_palette[label] for label in gt_labels]
+ else:
+ colors = [bbox_palette[label] for label in pred_labels]
+
+ self._visualizer.set_image(img)
+
+ for label, bbox, color in zip(gt_labels, gt_bboxes, colors):
+ self._visualizer.draw_bboxes(
+ bbox, edge_colors=color, face_colors=color, alpha=0.3)
+ self._visualizer.draw_bboxes(
+ bbox, edge_colors=color, alpha=1)
+
+ drawn_img = self._visualizer.get_image()
+
+ new_image = np.ones(
+ (100, img.shape[1], 3), dtype=np.uint8) * 255
+ self._visualizer.set_image(new_image)
+
+ if tokens_positive == -1: # REC
+ gt_tokens_positive = [[]]
+ else: # Phrase Grounding
+ gt_tokens_positive = [
+ tokens_positive[label] for label in gt_labels
+ ]
+ split_by_character = [char for char in text]
+ characters = []
+ start_index = 0
+ end_index = 0
+ for w in split_by_character:
+ end_index += len(w)
+ is_find = False
+ for i, positive in enumerate(gt_tokens_positive):
+ for p in positive:
+ if start_index >= p[0] and end_index <= p[1]:
+ characters.append([w, colors[i]])
+ is_find = True
+ break
+ if is_find:
+ break
+ if not is_find:
+ characters.append([w, (0, 0, 0)])
+ start_index = end_index
+
+ drawn_text = draw_all_character(self._visualizer, characters,
+ img.shape[1])
+ drawn_gt_img = np.concatenate((drawn_img, drawn_text), axis=0)
+
+ self._visualizer.set_image(img)
+
+ for label, bbox, color in zip(pred_labels, pred_bboxes,
+ colors):
+ self._visualizer.draw_bboxes(
+ bbox, edge_colors=color, face_colors=color, alpha=0.3)
+ self._visualizer.draw_bboxes(
+ bbox, edge_colors=color, alpha=1)
+ print(pred_labels, pred_bboxes, pred_scores, colors)
+ areas = (pred_bboxes[:, 3] - pred_bboxes[:, 1]) * (
+ pred_bboxes[:, 2] - pred_bboxes[:, 0])
+ scales = _get_adaptive_scales(areas)
+ score = [str(round(s.item(), 2)) for s in pred_scores]
+ font_sizes = [int(13 * scales[i]) for i in range(len(scales))]
+ self._visualizer.draw_texts(
+ score,
+ pred_bboxes[:, :2].int(),
+ colors=(255, 255, 255),
+ font_sizes=font_sizes,
+ bboxes=[{
+ 'facecolor': 'black',
+ 'alpha': 0.8,
+ 'pad': 0.7,
+ 'edgecolor': 'none'
+ }] * len(pred_bboxes))
+
+ drawn_img = self._visualizer.get_image()
+
+ new_image = np.ones(
+ (100, img.shape[1], 3), dtype=np.uint8) * 255
+ self._visualizer.set_image(new_image)
+ drawn_text = draw_all_character(self._visualizer, characters,
+ img.shape[1])
+ drawn_pred_img = np.concatenate((drawn_img, drawn_text),
+ axis=0)
+ drawn_img = np.concatenate((drawn_gt_img, drawn_pred_img),
+ axis=1)
+
+ if self.show:
+ self._visualizer.show(
+ drawn_img,
+ win_name=osp.basename(img_path),
+ wait_time=self.wait_time)
+ if out_file is not None:
+ mmcv.imwrite(drawn_img[..., ::-1], out_file)
+ else:
+ self.add_image('test_img', drawn_img, self._test_index)
+ else: # OD
+ self._visualizer.add_datasample(
+ osp.basename(img_path) if self.show else 'test_img',
+ img,
+ data_sample=data_sample,
+ show=self.show,
+ wait_time=self.wait_time,
+ pred_score_thr=self.score_thr,
+ out_file=out_file,
+ step=self._test_index)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/yolox_mode_switch_hook.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/yolox_mode_switch_hook.py
new file mode 100644
index 0000000000000000000000000000000000000000..05a2c69068bedd1c6fb3836e1fc34568e9f6bc83
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/hooks/yolox_mode_switch_hook.py
@@ -0,0 +1,66 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Sequence
+
+from mmengine.hooks import Hook
+from mmengine.model import is_model_wrapper
+
+from mmdet.registry import HOOKS
+
+
+@HOOKS.register_module()
+class YOLOXModeSwitchHook(Hook):
+ """Switch the mode of YOLOX during training.
+
+ This hook turns off the mosaic and mixup data augmentation and switches
+ to use L1 loss in bbox_head.
+
+ Args:
+ num_last_epochs (int): The number of latter epochs in the end of the
+ training to close the data augmentation and switch to L1 loss.
+ Defaults to 15.
+ skip_type_keys (Sequence[str], optional): Sequence of type string to be
+ skip pipeline. Defaults to ('Mosaic', 'RandomAffine', 'MixUp').
+ """
+
+ def __init__(
+ self,
+ num_last_epochs: int = 15,
+ skip_type_keys: Sequence[str] = ('Mosaic', 'RandomAffine', 'MixUp')
+ ) -> None:
+ self.num_last_epochs = num_last_epochs
+ self.skip_type_keys = skip_type_keys
+ self._restart_dataloader = False
+ self._has_switched = False
+
+ def before_train_epoch(self, runner) -> None:
+ """Close mosaic and mixup augmentation and switches to use L1 loss."""
+ epoch = runner.epoch
+ train_loader = runner.train_dataloader
+ model = runner.model
+ # TODO: refactor after mmengine using model wrapper
+ if is_model_wrapper(model):
+ model = model.module
+ epoch_to_be_switched = ((epoch + 1) >=
+ runner.max_epochs - self.num_last_epochs)
+ if epoch_to_be_switched and not self._has_switched:
+ runner.logger.info('No mosaic and mixup aug now!')
+ # The dataset pipeline cannot be updated when persistent_workers
+ # is True, so we need to force the dataloader's multi-process
+ # restart. This is a very hacky approach.
+ train_loader.dataset.update_skip_type_keys(self.skip_type_keys)
+ if hasattr(train_loader, 'persistent_workers'
+ ) and train_loader.persistent_workers is True:
+ train_loader._DataLoader__initialized = False
+ train_loader._iterator = None
+ self._restart_dataloader = True
+ runner.logger.info('Add additional L1 loss now!')
+ if hasattr(model, 'detector'):
+ model.detector.bbox_head.use_l1 = True
+ else:
+ model.bbox_head.use_l1 = True
+ self._has_switched = True
+ else:
+ # Once the restart is complete, we need to restore
+ # the initialization flag.
+ if self._restart_dataloader:
+ train_loader._DataLoader__initialized = True
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/optimizers/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/optimizers/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..83db069ee34cad0888bbf388d3cc7030ba49bbbb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/optimizers/__init__.py
@@ -0,0 +1,5 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .layer_decay_optimizer_constructor import \
+ LearningRateDecayOptimizerConstructor
+
+__all__ = ['LearningRateDecayOptimizerConstructor']
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/optimizers/layer_decay_optimizer_constructor.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/optimizers/layer_decay_optimizer_constructor.py
new file mode 100644
index 0000000000000000000000000000000000000000..73028a0aef698d63dcba8c4935d6ef6c577d0f46
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/optimizers/layer_decay_optimizer_constructor.py
@@ -0,0 +1,158 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import json
+from typing import List
+
+import torch.nn as nn
+from mmengine.dist import get_dist_info
+from mmengine.logging import MMLogger
+from mmengine.optim import DefaultOptimWrapperConstructor
+
+from mmdet.registry import OPTIM_WRAPPER_CONSTRUCTORS
+
+
+def get_layer_id_for_convnext(var_name, max_layer_id):
+ """Get the layer id to set the different learning rates in ``layer_wise``
+ decay_type.
+
+ Args:
+ var_name (str): The key of the model.
+ max_layer_id (int): Maximum layer id.
+
+ Returns:
+ int: The id number corresponding to different learning rate in
+ ``LearningRateDecayOptimizerConstructor``.
+ """
+
+ if var_name in ('backbone.cls_token', 'backbone.mask_token',
+ 'backbone.pos_embed'):
+ return 0
+ elif var_name.startswith('backbone.downsample_layers'):
+ stage_id = int(var_name.split('.')[2])
+ if stage_id == 0:
+ layer_id = 0
+ elif stage_id == 1:
+ layer_id = 2
+ elif stage_id == 2:
+ layer_id = 3
+ elif stage_id == 3:
+ layer_id = max_layer_id
+ return layer_id
+ elif var_name.startswith('backbone.stages'):
+ stage_id = int(var_name.split('.')[2])
+ block_id = int(var_name.split('.')[3])
+ if stage_id == 0:
+ layer_id = 1
+ elif stage_id == 1:
+ layer_id = 2
+ elif stage_id == 2:
+ layer_id = 3 + block_id // 3
+ elif stage_id == 3:
+ layer_id = max_layer_id
+ return layer_id
+ else:
+ return max_layer_id + 1
+
+
+def get_stage_id_for_convnext(var_name, max_stage_id):
+ """Get the stage id to set the different learning rates in ``stage_wise``
+ decay_type.
+
+ Args:
+ var_name (str): The key of the model.
+ max_stage_id (int): Maximum stage id.
+
+ Returns:
+ int: The id number corresponding to different learning rate in
+ ``LearningRateDecayOptimizerConstructor``.
+ """
+
+ if var_name in ('backbone.cls_token', 'backbone.mask_token',
+ 'backbone.pos_embed'):
+ return 0
+ elif var_name.startswith('backbone.downsample_layers'):
+ return 0
+ elif var_name.startswith('backbone.stages'):
+ stage_id = int(var_name.split('.')[2])
+ return stage_id + 1
+ else:
+ return max_stage_id - 1
+
+
+@OPTIM_WRAPPER_CONSTRUCTORS.register_module()
+class LearningRateDecayOptimizerConstructor(DefaultOptimWrapperConstructor):
+ # Different learning rates are set for different layers of backbone.
+ # Note: Currently, this optimizer constructor is built for ConvNeXt.
+
+ def add_params(self, params: List[dict], module: nn.Module,
+ **kwargs) -> None:
+ """Add all parameters of module to the params list.
+
+ The parameters of the given module will be added to the list of param
+ groups, with specific rules defined by paramwise_cfg.
+
+ Args:
+ params (list[dict]): A list of param groups, it will be modified
+ in place.
+ module (nn.Module): The module to be added.
+ """
+ logger = MMLogger.get_current_instance()
+
+ parameter_groups = {}
+ logger.info(f'self.paramwise_cfg is {self.paramwise_cfg}')
+ num_layers = self.paramwise_cfg.get('num_layers') + 2
+ decay_rate = self.paramwise_cfg.get('decay_rate')
+ decay_type = self.paramwise_cfg.get('decay_type', 'layer_wise')
+ logger.info('Build LearningRateDecayOptimizerConstructor '
+ f'{decay_type} {decay_rate} - {num_layers}')
+ weight_decay = self.base_wd
+ for name, param in module.named_parameters():
+ if not param.requires_grad:
+ continue # frozen weights
+ if len(param.shape) == 1 or name.endswith('.bias') or name in (
+ 'pos_embed', 'cls_token'):
+ group_name = 'no_decay'
+ this_weight_decay = 0.
+ else:
+ group_name = 'decay'
+ this_weight_decay = weight_decay
+ if 'layer_wise' in decay_type:
+ if 'ConvNeXt' in module.backbone.__class__.__name__:
+ layer_id = get_layer_id_for_convnext(
+ name, self.paramwise_cfg.get('num_layers'))
+ logger.info(f'set param {name} as id {layer_id}')
+ else:
+ raise NotImplementedError()
+ elif decay_type == 'stage_wise':
+ if 'ConvNeXt' in module.backbone.__class__.__name__:
+ layer_id = get_stage_id_for_convnext(name, num_layers)
+ logger.info(f'set param {name} as id {layer_id}')
+ else:
+ raise NotImplementedError()
+ group_name = f'layer_{layer_id}_{group_name}'
+
+ if group_name not in parameter_groups:
+ scale = decay_rate**(num_layers - layer_id - 1)
+
+ parameter_groups[group_name] = {
+ 'weight_decay': this_weight_decay,
+ 'params': [],
+ 'param_names': [],
+ 'lr_scale': scale,
+ 'group_name': group_name,
+ 'lr': scale * self.base_lr,
+ }
+
+ parameter_groups[group_name]['params'].append(param)
+ parameter_groups[group_name]['param_names'].append(name)
+ rank, _ = get_dist_info()
+ if rank == 0:
+ to_display = {}
+ for key in parameter_groups:
+ to_display[key] = {
+ 'param_names': parameter_groups[key]['param_names'],
+ 'lr_scale': parameter_groups[key]['lr_scale'],
+ 'lr': parameter_groups[key]['lr'],
+ 'weight_decay': parameter_groups[key]['weight_decay'],
+ }
+ logger.info(f'Param groups = {json.dumps(to_display, indent=2)}')
+ params.extend(parameter_groups.values())
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/runner/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/runner/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..e8bcce4448e48e2d64354ba6770f9f426fb3d869
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/runner/__init__.py
@@ -0,0 +1,4 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .loops import TeacherStudentValLoop
+
+__all__ = ['TeacherStudentValLoop']
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/runner/loops.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/runner/loops.py
new file mode 100644
index 0000000000000000000000000000000000000000..afe53afa5c80facf3ba6c224bd358e0859dade32
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/runner/loops.py
@@ -0,0 +1,38 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.model import is_model_wrapper
+from mmengine.runner import ValLoop
+
+from mmdet.registry import LOOPS
+
+
+@LOOPS.register_module()
+class TeacherStudentValLoop(ValLoop):
+ """Loop for validation of model teacher and student."""
+
+ def run(self):
+ """Launch validation for model teacher and student."""
+ self.runner.call_hook('before_val')
+ self.runner.call_hook('before_val_epoch')
+ self.runner.model.eval()
+
+ model = self.runner.model
+ if is_model_wrapper(model):
+ model = model.module
+ assert hasattr(model, 'teacher')
+ assert hasattr(model, 'student')
+
+ predict_on = model.semi_test_cfg.get('predict_on', None)
+ multi_metrics = dict()
+ for _predict_on in ['teacher', 'student']:
+ model.semi_test_cfg['predict_on'] = _predict_on
+ for idx, data_batch in enumerate(self.dataloader):
+ self.run_iter(idx, data_batch)
+ # compute metrics
+ metrics = self.evaluator.evaluate(len(self.dataloader.dataset))
+ multi_metrics.update(
+ {'/'.join((_predict_on, k)): v
+ for k, v in metrics.items()})
+ model.semi_test_cfg['predict_on'] = predict_on
+
+ self.runner.call_hook('after_val_epoch', metrics=multi_metrics)
+ self.runner.call_hook('after_val')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/schedulers/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/schedulers/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..01261646fa8255c643e86ba0517019760a50d387
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/schedulers/__init__.py
@@ -0,0 +1,8 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .quadratic_warmup import (QuadraticWarmupLR, QuadraticWarmupMomentum,
+ QuadraticWarmupParamScheduler)
+
+__all__ = [
+ 'QuadraticWarmupParamScheduler', 'QuadraticWarmupMomentum',
+ 'QuadraticWarmupLR'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/schedulers/quadratic_warmup.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/schedulers/quadratic_warmup.py
new file mode 100644
index 0000000000000000000000000000000000000000..639b47854887786bf3f81d6d0a375033d190d91e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/engine/schedulers/quadratic_warmup.py
@@ -0,0 +1,131 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.optim.scheduler.lr_scheduler import LRSchedulerMixin
+from mmengine.optim.scheduler.momentum_scheduler import MomentumSchedulerMixin
+from mmengine.optim.scheduler.param_scheduler import INF, _ParamScheduler
+from torch.optim import Optimizer
+
+from mmdet.registry import PARAM_SCHEDULERS
+
+
+@PARAM_SCHEDULERS.register_module()
+class QuadraticWarmupParamScheduler(_ParamScheduler):
+ r"""Warm up the parameter value of each parameter group by quadratic
+ formula:
+
+ .. math::
+
+ X_{t} = X_{t-1} + \frac{2t+1}{{(end-begin)}^{2}} \times X_{base}
+
+ Args:
+ optimizer (Optimizer): Wrapped optimizer.
+ param_name (str): Name of the parameter to be adjusted, such as
+ ``lr``, ``momentum``.
+ begin (int): Step at which to start updating the parameters.
+ Defaults to 0.
+ end (int): Step at which to stop updating the parameters.
+ Defaults to INF.
+ last_step (int): The index of last step. Used for resume without
+ state dict. Defaults to -1.
+ by_epoch (bool): Whether the scheduled parameters are updated by
+ epochs. Defaults to True.
+ verbose (bool): Whether to print the value for each update.
+ Defaults to False.
+ """
+
+ def __init__(self,
+ optimizer: Optimizer,
+ param_name: str,
+ begin: int = 0,
+ end: int = INF,
+ last_step: int = -1,
+ by_epoch: bool = True,
+ verbose: bool = False):
+ if end >= INF:
+ raise ValueError('``end`` must be less than infinity,'
+ 'Please set ``end`` parameter of '
+ '``QuadraticWarmupScheduler`` as the '
+ 'number of warmup end.')
+ self.total_iters = end - begin
+ super().__init__(
+ optimizer=optimizer,
+ param_name=param_name,
+ begin=begin,
+ end=end,
+ last_step=last_step,
+ by_epoch=by_epoch,
+ verbose=verbose)
+
+ @classmethod
+ def build_iter_from_epoch(cls,
+ *args,
+ begin=0,
+ end=INF,
+ by_epoch=True,
+ epoch_length=None,
+ **kwargs):
+ """Build an iter-based instance of this scheduler from an epoch-based
+ config."""
+ assert by_epoch, 'Only epoch-based kwargs whose `by_epoch=True` can ' \
+ 'be converted to iter-based.'
+ assert epoch_length is not None and epoch_length > 0, \
+ f'`epoch_length` must be a positive integer, ' \
+ f'but got {epoch_length}.'
+ by_epoch = False
+ begin = begin * epoch_length
+ if end != INF:
+ end = end * epoch_length
+ return cls(*args, begin=begin, end=end, by_epoch=by_epoch, **kwargs)
+
+ def _get_value(self):
+ """Compute value using chainable form of the scheduler."""
+ if self.last_step == 0:
+ return [
+ base_value * (2 * self.last_step + 1) / self.total_iters**2
+ for base_value in self.base_values
+ ]
+
+ return [
+ group[self.param_name] + base_value *
+ (2 * self.last_step + 1) / self.total_iters**2
+ for base_value, group in zip(self.base_values,
+ self.optimizer.param_groups)
+ ]
+
+
+@PARAM_SCHEDULERS.register_module()
+class QuadraticWarmupLR(LRSchedulerMixin, QuadraticWarmupParamScheduler):
+ """Warm up the learning rate of each parameter group by quadratic formula.
+
+ Args:
+ optimizer (Optimizer): Wrapped optimizer.
+ begin (int): Step at which to start updating the parameters.
+ Defaults to 0.
+ end (int): Step at which to stop updating the parameters.
+ Defaults to INF.
+ last_step (int): The index of last step. Used for resume without
+ state dict. Defaults to -1.
+ by_epoch (bool): Whether the scheduled parameters are updated by
+ epochs. Defaults to True.
+ verbose (bool): Whether to print the value for each update.
+ Defaults to False.
+ """
+
+
+@PARAM_SCHEDULERS.register_module()
+class QuadraticWarmupMomentum(MomentumSchedulerMixin,
+ QuadraticWarmupParamScheduler):
+ """Warm up the momentum value of each parameter group by quadratic formula.
+
+ Args:
+ optimizer (Optimizer): Wrapped optimizer.
+ begin (int): Step at which to start updating the parameters.
+ Defaults to 0.
+ end (int): Step at which to stop updating the parameters.
+ Defaults to INF.
+ last_step (int): The index of last step. Used for resume without
+ state dict. Defaults to -1.
+ by_epoch (bool): Whether the scheduled parameters are updated by
+ epochs. Defaults to True.
+ verbose (bool): Whether to print the value for each update.
+ Defaults to False.
+ """
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..126dea092eb1a4affab9fbe3fb043f5b373607ee
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/__init__.py
@@ -0,0 +1,4 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .evaluator import * # noqa: F401,F403
+from .functional import * # noqa: F401,F403
+from .metrics import * # noqa: F401,F403
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/evaluator/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/evaluator/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..6b13fe99548e7e2e4c6e196a2da22b9c8cbec8a3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/evaluator/__init__.py
@@ -0,0 +1,4 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .multi_datasets_evaluator import MultiDatasetsEvaluator
+
+__all__ = ['MultiDatasetsEvaluator']
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/evaluator/multi_datasets_evaluator.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/evaluator/multi_datasets_evaluator.py
new file mode 100644
index 0000000000000000000000000000000000000000..5cff1cf210e644e11b348f3aa757119ac579170d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/evaluator/multi_datasets_evaluator.py
@@ -0,0 +1,111 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+from collections import OrderedDict
+from typing import Sequence, Union
+
+from mmengine.dist import (broadcast_object_list, collect_results,
+ is_main_process)
+from mmengine.evaluator import BaseMetric, Evaluator
+from mmengine.evaluator.metric import _to_cpu
+from mmengine.registry import EVALUATOR
+
+from mmdet.utils import ConfigType
+
+
+@EVALUATOR.register_module()
+class MultiDatasetsEvaluator(Evaluator):
+ """Wrapper class to compose class: `ConcatDataset` and multiple
+ :class:`BaseMetric` instances.
+ The metrics will be evaluated on each dataset slice separately. The name of
+ the each metric is the concatenation of the dataset prefix, the metric
+ prefix and the key of metric - e.g.
+ `dataset_prefix/metric_prefix/accuracy`.
+
+ Args:
+ metrics (dict or BaseMetric or Sequence): The config of metrics.
+ dataset_prefixes (Sequence[str]): The prefix of each dataset. The
+ length of this sequence should be the same as the length of the
+ datasets.
+ """
+
+ def __init__(self, metrics: Union[ConfigType, BaseMetric, Sequence],
+ dataset_prefixes: Sequence[str]) -> None:
+ super().__init__(metrics)
+ self.dataset_prefixes = dataset_prefixes
+ self._setups = False
+
+ def _get_cumulative_sizes(self):
+ # ConcatDataset have a property `cumulative_sizes`
+ if isinstance(self.dataset_meta, Sequence):
+ dataset_slices = self.dataset_meta[0]['cumulative_sizes']
+ if not self._setups:
+ self._setups = True
+ for dataset_meta, metric in zip(self.dataset_meta,
+ self.metrics):
+ metric.dataset_meta = dataset_meta
+ else:
+ dataset_slices = self.dataset_meta['cumulative_sizes']
+ return dataset_slices
+
+ def evaluate(self, size: int) -> dict:
+ """Invoke ``evaluate`` method of each metric and collect the metrics
+ dictionary.
+
+ Args:
+ size (int): Length of the entire validation dataset. When batch
+ size > 1, the dataloader may pad some data samples to make
+ sure all ranks have the same length of dataset slice. The
+ ``collect_results`` function will drop the padded data based on
+ this size.
+
+ Returns:
+ dict: Evaluation results of all metrics. The keys are the names
+ of the metrics, and the values are corresponding results.
+ """
+ metrics_results = OrderedDict()
+ dataset_slices = self._get_cumulative_sizes()
+ assert len(dataset_slices) == len(self.dataset_prefixes)
+
+ for dataset_prefix, start, end, metric in zip(
+ self.dataset_prefixes, [0] + dataset_slices[:-1],
+ dataset_slices, self.metrics):
+ if len(metric.results) == 0:
+ warnings.warn(
+ f'{metric.__class__.__name__} got empty `self.results`.'
+ 'Please ensure that the processed results are properly '
+ 'added into `self.results` in `process` method.')
+
+ results = collect_results(metric.results, size,
+ metric.collect_device)
+
+ if is_main_process():
+ # cast all tensors in results list to cpu
+ results = _to_cpu(results)
+ _metrics = metric.compute_metrics(
+ results[start:end]) # type: ignore
+
+ if metric.prefix:
+ final_prefix = '/'.join((dataset_prefix, metric.prefix))
+ else:
+ final_prefix = dataset_prefix
+ print(f'================{final_prefix}================')
+ metric_results = {
+ '/'.join((final_prefix, k)): v
+ for k, v in _metrics.items()
+ }
+
+ # Check metric name conflicts
+ for name in metric_results.keys():
+ if name in metrics_results:
+ raise ValueError(
+ 'There are multiple evaluation results with '
+ f'the same metric name {name}. Please make '
+ 'sure all metrics have different prefixes.')
+ metrics_results.update(metric_results)
+ metric.results.clear()
+ if is_main_process():
+ metrics_results = [metrics_results]
+ else:
+ metrics_results = [None] # type: ignore
+ broadcast_object_list(metrics_results)
+ return metrics_results[0]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..96d58ebd3ab0dd714a6f361622a7faf2a09486cb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/__init__.py
@@ -0,0 +1,26 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .bbox_overlaps import bbox_overlaps
+from .cityscapes_utils import evaluateImgLists
+from .class_names import (cityscapes_classes, coco_classes,
+ coco_panoptic_classes, dataset_aliases, get_classes,
+ imagenet_det_classes, imagenet_vid_classes,
+ objects365v1_classes, objects365v2_classes,
+ oid_challenge_classes, oid_v6_classes, voc_classes)
+from .mean_ap import average_precision, eval_map, print_map_summary
+from .panoptic_utils import (INSTANCE_OFFSET, pq_compute_multi_core,
+ pq_compute_single_core)
+from .recall import (eval_recalls, plot_iou_recall, plot_num_recall,
+ print_recall_summary)
+from .ytvis import YTVIS
+from .ytviseval import YTVISeval
+
+__all__ = [
+ 'voc_classes', 'imagenet_det_classes', 'imagenet_vid_classes',
+ 'coco_classes', 'cityscapes_classes', 'dataset_aliases', 'get_classes',
+ 'average_precision', 'eval_map', 'print_map_summary', 'eval_recalls',
+ 'print_recall_summary', 'plot_num_recall', 'plot_iou_recall',
+ 'oid_v6_classes', 'oid_challenge_classes', 'INSTANCE_OFFSET',
+ 'pq_compute_single_core', 'pq_compute_multi_core', 'bbox_overlaps',
+ 'objects365v1_classes', 'objects365v2_classes', 'coco_panoptic_classes',
+ 'evaluateImgLists', 'YTVIS', 'YTVISeval'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/bbox_overlaps.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/bbox_overlaps.py
new file mode 100644
index 0000000000000000000000000000000000000000..5d6eb82fcfc8d5444dd2a13b7d95b978f8206a55
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/bbox_overlaps.py
@@ -0,0 +1,65 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import numpy as np
+
+
+def bbox_overlaps(bboxes1,
+ bboxes2,
+ mode='iou',
+ eps=1e-6,
+ use_legacy_coordinate=False):
+ """Calculate the ious between each bbox of bboxes1 and bboxes2.
+
+ Args:
+ bboxes1 (ndarray): Shape (n, 4)
+ bboxes2 (ndarray): Shape (k, 4)
+ mode (str): IOU (intersection over union) or IOF (intersection
+ over foreground)
+ use_legacy_coordinate (bool): Whether to use coordinate system in
+ mmdet v1.x. which means width, height should be
+ calculated as 'x2 - x1 + 1` and 'y2 - y1 + 1' respectively.
+ Note when function is used in `VOCDataset`, it should be
+ True to align with the official implementation
+ `http://host.robots.ox.ac.uk/pascal/VOC/voc2012/VOCdevkit_18-May-2011.tar`
+ Default: False.
+
+ Returns:
+ ious (ndarray): Shape (n, k)
+ """
+
+ assert mode in ['iou', 'iof']
+ if not use_legacy_coordinate:
+ extra_length = 0.
+ else:
+ extra_length = 1.
+ bboxes1 = bboxes1.astype(np.float32)
+ bboxes2 = bboxes2.astype(np.float32)
+ rows = bboxes1.shape[0]
+ cols = bboxes2.shape[0]
+ ious = np.zeros((rows, cols), dtype=np.float32)
+ if rows * cols == 0:
+ return ious
+ exchange = False
+ if bboxes1.shape[0] > bboxes2.shape[0]:
+ bboxes1, bboxes2 = bboxes2, bboxes1
+ ious = np.zeros((cols, rows), dtype=np.float32)
+ exchange = True
+ area1 = (bboxes1[:, 2] - bboxes1[:, 0] + extra_length) * (
+ bboxes1[:, 3] - bboxes1[:, 1] + extra_length)
+ area2 = (bboxes2[:, 2] - bboxes2[:, 0] + extra_length) * (
+ bboxes2[:, 3] - bboxes2[:, 1] + extra_length)
+ for i in range(bboxes1.shape[0]):
+ x_start = np.maximum(bboxes1[i, 0], bboxes2[:, 0])
+ y_start = np.maximum(bboxes1[i, 1], bboxes2[:, 1])
+ x_end = np.minimum(bboxes1[i, 2], bboxes2[:, 2])
+ y_end = np.minimum(bboxes1[i, 3], bboxes2[:, 3])
+ overlap = np.maximum(x_end - x_start + extra_length, 0) * np.maximum(
+ y_end - y_start + extra_length, 0)
+ if mode == 'iou':
+ union = area1[i] + area2 - overlap
+ else:
+ union = area1[i] if not exchange else area2
+ union = np.maximum(union, eps)
+ ious[i, :] = overlap / union
+ if exchange:
+ ious = ious.T
+ return ious
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/cityscapes_utils.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/cityscapes_utils.py
new file mode 100644
index 0000000000000000000000000000000000000000..5ced3680deefe333af7cca3675a6359c02dd96f8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/cityscapes_utils.py
@@ -0,0 +1,302 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+# Copyright (c) https://github.com/mcordts/cityscapesScripts
+# A wrapper of `cityscapesscripts` which supports loading groundtruth
+# image from `backend_args`.
+import json
+import os
+import sys
+from pathlib import Path
+from typing import Optional, Union
+
+import mmcv
+import numpy as np
+from mmengine.fileio import get
+
+try:
+ import cityscapesscripts.evaluation.evalInstanceLevelSemanticLabeling as CSEval # noqa: E501
+ from cityscapesscripts.evaluation.evalInstanceLevelSemanticLabeling import \
+ CArgs # noqa: E501
+ from cityscapesscripts.evaluation.instance import Instance
+ from cityscapesscripts.helpers.csHelpers import (id2label, labels,
+ writeDict2JSON)
+ HAS_CITYSCAPESAPI = True
+except ImportError:
+ CArgs = object
+ HAS_CITYSCAPESAPI = False
+
+
+def evaluateImgLists(prediction_list: list,
+ groundtruth_list: list,
+ args: CArgs,
+ backend_args: Optional[dict] = None,
+ dump_matches: bool = False) -> dict:
+ """A wrapper of obj:``cityscapesscripts.evaluation.
+
+ evalInstanceLevelSemanticLabeling.evaluateImgLists``. Support loading
+ groundtruth image from file backend.
+ Args:
+ prediction_list (list): A list of prediction txt file.
+ groundtruth_list (list): A list of groundtruth image file.
+ args (CArgs): A global object setting in
+ obj:``cityscapesscripts.evaluation.
+ evalInstanceLevelSemanticLabeling``
+ backend_args (dict, optional): Arguments to instantiate the
+ preifx of uri corresponding backend. Defaults to None.
+ dump_matches (bool): whether dump matches.json. Defaults to False.
+ Returns:
+ dict: The computed metric.
+ """
+ if not HAS_CITYSCAPESAPI:
+ raise RuntimeError('Failed to import `cityscapesscripts`.'
+ 'Please try to install official '
+ 'cityscapesscripts by '
+ '"pip install cityscapesscripts"')
+ # determine labels of interest
+ CSEval.setInstanceLabels(args)
+ # get dictionary of all ground truth instances
+ gt_instances = getGtInstances(
+ groundtruth_list, args, backend_args=backend_args)
+ # match predictions and ground truth
+ matches = matchGtWithPreds(prediction_list, groundtruth_list, gt_instances,
+ args, backend_args)
+ if dump_matches:
+ CSEval.writeDict2JSON(matches, 'matches.json')
+ # evaluate matches
+ apScores = CSEval.evaluateMatches(matches, args)
+ # averages
+ avgDict = CSEval.computeAverages(apScores, args)
+ # result dict
+ resDict = CSEval.prepareJSONDataForResults(avgDict, apScores, args)
+ if args.JSONOutput:
+ # create output folder if necessary
+ path = os.path.dirname(args.exportFile)
+ CSEval.ensurePath(path)
+ # Write APs to JSON
+ CSEval.writeDict2JSON(resDict, args.exportFile)
+
+ CSEval.printResults(avgDict, args)
+
+ return resDict
+
+
+def matchGtWithPreds(prediction_list: list,
+ groundtruth_list: list,
+ gt_instances: dict,
+ args: CArgs,
+ backend_args=None):
+ """A wrapper of obj:``cityscapesscripts.evaluation.
+
+ evalInstanceLevelSemanticLabeling.matchGtWithPreds``. Support loading
+ groundtruth image from file backend.
+ Args:
+ prediction_list (list): A list of prediction txt file.
+ groundtruth_list (list): A list of groundtruth image file.
+ gt_instances (dict): Groundtruth dict.
+ args (CArgs): A global object setting in
+ obj:``cityscapesscripts.evaluation.
+ evalInstanceLevelSemanticLabeling``
+ backend_args (dict, optional): Arguments to instantiate the
+ preifx of uri corresponding backend. Defaults to None.
+ Returns:
+ dict: The processed prediction and groundtruth result.
+ """
+ if not HAS_CITYSCAPESAPI:
+ raise RuntimeError('Failed to import `cityscapesscripts`.'
+ 'Please try to install official '
+ 'cityscapesscripts by '
+ '"pip install cityscapesscripts"')
+ matches: dict = dict()
+ if not args.quiet:
+ print(f'Matching {len(prediction_list)} pairs of images...')
+
+ count = 0
+ for (pred, gt) in zip(prediction_list, groundtruth_list):
+ # Read input files
+ gt_image = readGTImage(gt, backend_args)
+ pred_info = readPredInfo(pred)
+ # Get and filter ground truth instances
+ unfiltered_instances = gt_instances[gt]
+ cur_gt_instances_orig = CSEval.filterGtInstances(
+ unfiltered_instances, args)
+
+ # Try to assign all predictions
+ (cur_gt_instances,
+ cur_pred_instances) = CSEval.assignGt2Preds(cur_gt_instances_orig,
+ gt_image, pred_info, args)
+
+ # append to global dict
+ matches[gt] = {}
+ matches[gt]['groundTruth'] = cur_gt_instances
+ matches[gt]['prediction'] = cur_pred_instances
+
+ count += 1
+ if not args.quiet:
+ print(f'\rImages Processed: {count}', end=' ')
+ sys.stdout.flush()
+
+ if not args.quiet:
+ print('')
+
+ return matches
+
+
+def readGTImage(image_file: Union[str, Path],
+ backend_args: Optional[dict] = None) -> np.ndarray:
+ """Read an image from path.
+
+ Same as obj:``cityscapesscripts.evaluation.
+ evalInstanceLevelSemanticLabeling.readGTImage``, but support loading
+ groundtruth image from file backend.
+ Args:
+ image_file (str or Path): Either a str or pathlib.Path.
+ backend_args (dict, optional): Instantiates the corresponding file
+ backend. It may contain `backend` key to specify the file
+ backend. If it contains, the file backend corresponding to this
+ value will be used and initialized with the remaining values,
+ otherwise the corresponding file backend will be selected
+ based on the prefix of the file path. Defaults to None.
+ Returns:
+ np.ndarray: The groundtruth image.
+ """
+ img_bytes = get(image_file, backend_args=backend_args)
+ img = mmcv.imfrombytes(img_bytes, flag='unchanged', backend='pillow')
+ return img
+
+
+def readPredInfo(prediction_file: str) -> dict:
+ """A wrapper of obj:``cityscapesscripts.evaluation.
+
+ evalInstanceLevelSemanticLabeling.readPredInfo``.
+ Args:
+ prediction_file (str): The prediction txt file.
+ Returns:
+ dict: The processed prediction results.
+ """
+ if not HAS_CITYSCAPESAPI:
+ raise RuntimeError('Failed to import `cityscapesscripts`.'
+ 'Please try to install official '
+ 'cityscapesscripts by '
+ '"pip install cityscapesscripts"')
+ printError = CSEval.printError
+
+ predInfo = {}
+ if (not os.path.isfile(prediction_file)):
+ printError(f"Infofile '{prediction_file}' "
+ 'for the predictions not found.')
+ with open(prediction_file) as f:
+ for line in f:
+ splittedLine = line.split(' ')
+ if len(splittedLine) != 3:
+ printError('Invalid prediction file. Expected content: '
+ 'relPathPrediction1 labelIDPrediction1 '
+ 'confidencePrediction1')
+ if os.path.isabs(splittedLine[0]):
+ printError('Invalid prediction file. First entry in each '
+ 'line must be a relative path.')
+
+ filename = os.path.join(
+ os.path.dirname(prediction_file), splittedLine[0])
+
+ imageInfo = {}
+ imageInfo['labelID'] = int(float(splittedLine[1]))
+ imageInfo['conf'] = float(splittedLine[2]) # type: ignore
+ predInfo[filename] = imageInfo
+
+ return predInfo
+
+
+def getGtInstances(groundtruth_list: list,
+ args: CArgs,
+ backend_args: Optional[dict] = None) -> dict:
+ """A wrapper of obj:``cityscapesscripts.evaluation.
+
+ evalInstanceLevelSemanticLabeling.getGtInstances``. Support loading
+ groundtruth image from file backend.
+ Args:
+ groundtruth_list (list): A list of groundtruth image file.
+ args (CArgs): A global object setting in
+ obj:``cityscapesscripts.evaluation.
+ evalInstanceLevelSemanticLabeling``
+ backend_args (dict, optional): Arguments to instantiate the
+ preifx of uri corresponding backend. Defaults to None.
+ Returns:
+ dict: The computed metric.
+ """
+ if not HAS_CITYSCAPESAPI:
+ raise RuntimeError('Failed to import `cityscapesscripts`.'
+ 'Please try to install official '
+ 'cityscapesscripts by '
+ '"pip install cityscapesscripts"')
+ # if there is a global statistics json, then load it
+ if (os.path.isfile(args.gtInstancesFile)):
+ if not args.quiet:
+ print('Loading ground truth instances from JSON.')
+ with open(args.gtInstancesFile) as json_file:
+ gt_instances = json.load(json_file)
+ # otherwise create it
+ else:
+ if (not args.quiet):
+ print('Creating ground truth instances from png files.')
+ gt_instances = instances2dict(
+ groundtruth_list, args, backend_args=backend_args)
+ writeDict2JSON(gt_instances, args.gtInstancesFile)
+
+ return gt_instances
+
+
+def instances2dict(image_list: list,
+ args: CArgs,
+ backend_args: Optional[dict] = None) -> dict:
+ """A wrapper of obj:``cityscapesscripts.evaluation.
+
+ evalInstanceLevelSemanticLabeling.instances2dict``. Support loading
+ groundtruth image from file backend.
+ Args:
+ image_list (list): A list of image file.
+ args (CArgs): A global object setting in
+ obj:``cityscapesscripts.evaluation.
+ evalInstanceLevelSemanticLabeling``
+ backend_args (dict, optional): Arguments to instantiate the
+ preifx of uri corresponding backend. Defaults to None.
+ Returns:
+ dict: The processed groundtruth results.
+ """
+ if not HAS_CITYSCAPESAPI:
+ raise RuntimeError('Failed to import `cityscapesscripts`.'
+ 'Please try to install official '
+ 'cityscapesscripts by '
+ '"pip install cityscapesscripts"')
+ imgCount = 0
+ instanceDict = {}
+
+ if not isinstance(image_list, list):
+ image_list = [image_list]
+
+ if not args.quiet:
+ print(f'Processing {len(image_list)} images...')
+
+ for image_name in image_list:
+ # Load image
+ img_bytes = get(image_name, backend_args=backend_args)
+ imgNp = mmcv.imfrombytes(img_bytes, flag='unchanged', backend='pillow')
+
+ # Initialize label categories
+ instances: dict = {}
+ for label in labels:
+ instances[label.name] = []
+
+ # Loop through all instance ids in instance image
+ for instanceId in np.unique(imgNp):
+ instanceObj = Instance(imgNp, instanceId)
+
+ instances[id2label[instanceObj.labelID].name].append(
+ instanceObj.toDict())
+
+ instanceDict[image_name] = instances
+ imgCount += 1
+
+ if not args.quiet:
+ print(f'\rImages Processed: {imgCount}', end=' ')
+ sys.stdout.flush()
+
+ return instanceDict
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/class_names.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/class_names.py
new file mode 100644
index 0000000000000000000000000000000000000000..623a89cfdc06ab04831afd3423d5f725acc881f0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/class_names.py
@@ -0,0 +1,762 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.utils import is_str
+
+
+def wider_face_classes() -> list:
+ """Class names of WIDERFace."""
+ return ['face']
+
+
+def voc_classes() -> list:
+ """Class names of PASCAL VOC."""
+ return [
+ 'aeroplane', 'bicycle', 'bird', 'boat', 'bottle', 'bus', 'car', 'cat',
+ 'chair', 'cow', 'diningtable', 'dog', 'horse', 'motorbike', 'person',
+ 'pottedplant', 'sheep', 'sofa', 'train', 'tvmonitor'
+ ]
+
+
+def imagenet_det_classes() -> list:
+ """Class names of ImageNet Det."""
+ return [
+ 'accordion', 'airplane', 'ant', 'antelope', 'apple', 'armadillo',
+ 'artichoke', 'axe', 'baby_bed', 'backpack', 'bagel', 'balance_beam',
+ 'banana', 'band_aid', 'banjo', 'baseball', 'basketball', 'bathing_cap',
+ 'beaker', 'bear', 'bee', 'bell_pepper', 'bench', 'bicycle', 'binder',
+ 'bird', 'bookshelf', 'bow_tie', 'bow', 'bowl', 'brassiere', 'burrito',
+ 'bus', 'butterfly', 'camel', 'can_opener', 'car', 'cart', 'cattle',
+ 'cello', 'centipede', 'chain_saw', 'chair', 'chime', 'cocktail_shaker',
+ 'coffee_maker', 'computer_keyboard', 'computer_mouse', 'corkscrew',
+ 'cream', 'croquet_ball', 'crutch', 'cucumber', 'cup_or_mug', 'diaper',
+ 'digital_clock', 'dishwasher', 'dog', 'domestic_cat', 'dragonfly',
+ 'drum', 'dumbbell', 'electric_fan', 'elephant', 'face_powder', 'fig',
+ 'filing_cabinet', 'flower_pot', 'flute', 'fox', 'french_horn', 'frog',
+ 'frying_pan', 'giant_panda', 'goldfish', 'golf_ball', 'golfcart',
+ 'guacamole', 'guitar', 'hair_dryer', 'hair_spray', 'hamburger',
+ 'hammer', 'hamster', 'harmonica', 'harp', 'hat_with_a_wide_brim',
+ 'head_cabbage', 'helmet', 'hippopotamus', 'horizontal_bar', 'horse',
+ 'hotdog', 'iPod', 'isopod', 'jellyfish', 'koala_bear', 'ladle',
+ 'ladybug', 'lamp', 'laptop', 'lemon', 'lion', 'lipstick', 'lizard',
+ 'lobster', 'maillot', 'maraca', 'microphone', 'microwave', 'milk_can',
+ 'miniskirt', 'monkey', 'motorcycle', 'mushroom', 'nail', 'neck_brace',
+ 'oboe', 'orange', 'otter', 'pencil_box', 'pencil_sharpener', 'perfume',
+ 'person', 'piano', 'pineapple', 'ping-pong_ball', 'pitcher', 'pizza',
+ 'plastic_bag', 'plate_rack', 'pomegranate', 'popsicle', 'porcupine',
+ 'power_drill', 'pretzel', 'printer', 'puck', 'punching_bag', 'purse',
+ 'rabbit', 'racket', 'ray', 'red_panda', 'refrigerator',
+ 'remote_control', 'rubber_eraser', 'rugby_ball', 'ruler',
+ 'salt_or_pepper_shaker', 'saxophone', 'scorpion', 'screwdriver',
+ 'seal', 'sheep', 'ski', 'skunk', 'snail', 'snake', 'snowmobile',
+ 'snowplow', 'soap_dispenser', 'soccer_ball', 'sofa', 'spatula',
+ 'squirrel', 'starfish', 'stethoscope', 'stove', 'strainer',
+ 'strawberry', 'stretcher', 'sunglasses', 'swimming_trunks', 'swine',
+ 'syringe', 'table', 'tape_player', 'tennis_ball', 'tick', 'tie',
+ 'tiger', 'toaster', 'traffic_light', 'train', 'trombone', 'trumpet',
+ 'turtle', 'tv_or_monitor', 'unicycle', 'vacuum', 'violin',
+ 'volleyball', 'waffle_iron', 'washer', 'water_bottle', 'watercraft',
+ 'whale', 'wine_bottle', 'zebra'
+ ]
+
+
+def imagenet_vid_classes() -> list:
+ """Class names of ImageNet VID."""
+ return [
+ 'airplane', 'antelope', 'bear', 'bicycle', 'bird', 'bus', 'car',
+ 'cattle', 'dog', 'domestic_cat', 'elephant', 'fox', 'giant_panda',
+ 'hamster', 'horse', 'lion', 'lizard', 'monkey', 'motorcycle', 'rabbit',
+ 'red_panda', 'sheep', 'snake', 'squirrel', 'tiger', 'train', 'turtle',
+ 'watercraft', 'whale', 'zebra'
+ ]
+
+
+def coco_classes() -> list:
+ """Class names of COCO."""
+ return [
+ 'person', 'bicycle', 'car', 'motorcycle', 'airplane', 'bus', 'train',
+ 'truck', 'boat', 'traffic_light', 'fire_hydrant', 'stop_sign',
+ 'parking_meter', 'bench', 'bird', 'cat', 'dog', 'horse', 'sheep',
+ 'cow', 'elephant', 'bear', 'zebra', 'giraffe', 'backpack', 'umbrella',
+ 'handbag', 'tie', 'suitcase', 'frisbee', 'skis', 'snowboard',
+ 'sports_ball', 'kite', 'baseball_bat', 'baseball_glove', 'skateboard',
+ 'surfboard', 'tennis_racket', 'bottle', 'wine_glass', 'cup', 'fork',
+ 'knife', 'spoon', 'bowl', 'banana', 'apple', 'sandwich', 'orange',
+ 'broccoli', 'carrot', 'hot_dog', 'pizza', 'donut', 'cake', 'chair',
+ 'couch', 'potted_plant', 'bed', 'dining_table', 'toilet', 'tv',
+ 'laptop', 'mouse', 'remote', 'keyboard', 'cell_phone', 'microwave',
+ 'oven', 'toaster', 'sink', 'refrigerator', 'book', 'clock', 'vase',
+ 'scissors', 'teddy_bear', 'hair_drier', 'toothbrush'
+ ]
+
+
+def coco_panoptic_classes() -> list:
+ """Class names of COCO panoptic."""
+ return [
+ 'person', 'bicycle', 'car', 'motorcycle', 'airplane', 'bus', 'train',
+ 'truck', 'boat', 'traffic light', 'fire hydrant', 'stop sign',
+ 'parking meter', 'bench', 'bird', 'cat', 'dog', 'horse', 'sheep',
+ 'cow', 'elephant', 'bear', 'zebra', 'giraffe', 'backpack', 'umbrella',
+ 'handbag', 'tie', 'suitcase', 'frisbee', 'skis', 'snowboard',
+ 'sports ball', 'kite', 'baseball bat', 'baseball glove', 'skateboard',
+ 'surfboard', 'tennis racket', 'bottle', 'wine glass', 'cup', 'fork',
+ 'knife', 'spoon', 'bowl', 'banana', 'apple', 'sandwich', 'orange',
+ 'broccoli', 'carrot', 'hot dog', 'pizza', 'donut', 'cake', 'chair',
+ 'couch', 'potted plant', 'bed', 'dining table', 'toilet', 'tv',
+ 'laptop', 'mouse', 'remote', 'keyboard', 'cell phone', 'microwave',
+ 'oven', 'toaster', 'sink', 'refrigerator', 'book', 'clock', 'vase',
+ 'scissors', 'teddy bear', 'hair drier', 'toothbrush', 'banner',
+ 'blanket', 'bridge', 'cardboard', 'counter', 'curtain', 'door-stuff',
+ 'floor-wood', 'flower', 'fruit', 'gravel', 'house', 'light',
+ 'mirror-stuff', 'net', 'pillow', 'platform', 'playingfield',
+ 'railroad', 'river', 'road', 'roof', 'sand', 'sea', 'shelf', 'snow',
+ 'stairs', 'tent', 'towel', 'wall-brick', 'wall-stone', 'wall-tile',
+ 'wall-wood', 'water-other', 'window-blind', 'window-other',
+ 'tree-merged', 'fence-merged', 'ceiling-merged', 'sky-other-merged',
+ 'cabinet-merged', 'table-merged', 'floor-other-merged',
+ 'pavement-merged', 'mountain-merged', 'grass-merged', 'dirt-merged',
+ 'paper-merged', 'food-other-merged', 'building-other-merged',
+ 'rock-merged', 'wall-other-merged', 'rug-merged'
+ ]
+
+
+def cityscapes_classes() -> list:
+ """Class names of Cityscapes."""
+ return [
+ 'person', 'rider', 'car', 'truck', 'bus', 'train', 'motorcycle',
+ 'bicycle'
+ ]
+
+
+def oid_challenge_classes() -> list:
+ """Class names of Open Images Challenge."""
+ return [
+ 'Footwear', 'Jeans', 'House', 'Tree', 'Woman', 'Man', 'Land vehicle',
+ 'Person', 'Wheel', 'Bus', 'Human face', 'Bird', 'Dress', 'Girl',
+ 'Vehicle', 'Building', 'Cat', 'Car', 'Belt', 'Elephant', 'Dessert',
+ 'Butterfly', 'Train', 'Guitar', 'Poster', 'Book', 'Boy', 'Bee',
+ 'Flower', 'Window', 'Hat', 'Human head', 'Dog', 'Human arm', 'Drink',
+ 'Human mouth', 'Human hair', 'Human nose', 'Human hand', 'Table',
+ 'Marine invertebrates', 'Fish', 'Sculpture', 'Rose', 'Street light',
+ 'Glasses', 'Fountain', 'Skyscraper', 'Swimwear', 'Brassiere', 'Drum',
+ 'Duck', 'Countertop', 'Furniture', 'Ball', 'Human leg', 'Boat',
+ 'Balloon', 'Bicycle helmet', 'Goggles', 'Door', 'Human eye', 'Shirt',
+ 'Toy', 'Teddy bear', 'Pasta', 'Tomato', 'Human ear',
+ 'Vehicle registration plate', 'Microphone', 'Musical keyboard',
+ 'Tower', 'Houseplant', 'Flowerpot', 'Fruit', 'Vegetable',
+ 'Musical instrument', 'Suit', 'Motorcycle', 'Bagel', 'French fries',
+ 'Hamburger', 'Chair', 'Salt and pepper shakers', 'Snail', 'Airplane',
+ 'Horse', 'Laptop', 'Computer keyboard', 'Football helmet', 'Cocktail',
+ 'Juice', 'Tie', 'Computer monitor', 'Human beard', 'Bottle',
+ 'Saxophone', 'Lemon', 'Mouse', 'Sock', 'Cowboy hat', 'Sun hat',
+ 'Football', 'Porch', 'Sunglasses', 'Lobster', 'Crab', 'Picture frame',
+ 'Van', 'Crocodile', 'Surfboard', 'Shorts', 'Helicopter', 'Helmet',
+ 'Sports uniform', 'Taxi', 'Swan', 'Goose', 'Coat', 'Jacket', 'Handbag',
+ 'Flag', 'Skateboard', 'Television', 'Tire', 'Spoon', 'Palm tree',
+ 'Stairs', 'Salad', 'Castle', 'Oven', 'Microwave oven', 'Wine',
+ 'Ceiling fan', 'Mechanical fan', 'Cattle', 'Truck', 'Box', 'Ambulance',
+ 'Desk', 'Wine glass', 'Reptile', 'Tank', 'Traffic light', 'Billboard',
+ 'Tent', 'Insect', 'Spider', 'Treadmill', 'Cupboard', 'Shelf',
+ 'Seat belt', 'Human foot', 'Bicycle', 'Bicycle wheel', 'Couch',
+ 'Bookcase', 'Fedora', 'Backpack', 'Bench', 'Oyster',
+ 'Moths and butterflies', 'Lavender', 'Waffle', 'Fork', 'Animal',
+ 'Accordion', 'Mobile phone', 'Plate', 'Coffee cup', 'Saucer',
+ 'Platter', 'Dagger', 'Knife', 'Bull', 'Tortoise', 'Sea turtle', 'Deer',
+ 'Weapon', 'Apple', 'Ski', 'Taco', 'Traffic sign', 'Beer', 'Necklace',
+ 'Sunflower', 'Piano', 'Organ', 'Harpsichord', 'Bed', 'Cabinetry',
+ 'Nightstand', 'Curtain', 'Chest of drawers', 'Drawer', 'Parrot',
+ 'Sandal', 'High heels', 'Tableware', 'Cart', 'Mushroom', 'Kite',
+ 'Missile', 'Seafood', 'Camera', 'Paper towel', 'Toilet paper',
+ 'Sombrero', 'Radish', 'Lighthouse', 'Segway', 'Pig', 'Watercraft',
+ 'Golf cart', 'studio couch', 'Dolphin', 'Whale', 'Earrings', 'Otter',
+ 'Sea lion', 'Whiteboard', 'Monkey', 'Gondola', 'Zebra',
+ 'Baseball glove', 'Scarf', 'Adhesive tape', 'Trousers', 'Scoreboard',
+ 'Lily', 'Carnivore', 'Power plugs and sockets', 'Office building',
+ 'Sandwich', 'Swimming pool', 'Headphones', 'Tin can', 'Crown', 'Doll',
+ 'Cake', 'Frog', 'Beetle', 'Ant', 'Gas stove', 'Canoe', 'Falcon',
+ 'Blue jay', 'Egg', 'Fire hydrant', 'Raccoon', 'Muffin', 'Wall clock',
+ 'Coffee', 'Mug', 'Tea', 'Bear', 'Waste container', 'Home appliance',
+ 'Candle', 'Lion', 'Mirror', 'Starfish', 'Marine mammal', 'Wheelchair',
+ 'Umbrella', 'Alpaca', 'Violin', 'Cello', 'Brown bear', 'Canary', 'Bat',
+ 'Ruler', 'Plastic bag', 'Penguin', 'Watermelon', 'Harbor seal', 'Pen',
+ 'Pumpkin', 'Harp', 'Kitchen appliance', 'Roller skates', 'Bust',
+ 'Coffee table', 'Tennis ball', 'Tennis racket', 'Ladder', 'Boot',
+ 'Bowl', 'Stop sign', 'Volleyball', 'Eagle', 'Paddle', 'Chicken',
+ 'Skull', 'Lamp', 'Beehive', 'Maple', 'Sink', 'Goldfish', 'Tripod',
+ 'Coconut', 'Bidet', 'Tap', 'Bathroom cabinet', 'Toilet',
+ 'Filing cabinet', 'Pretzel', 'Table tennis racket', 'Bronze sculpture',
+ 'Rocket', 'Mouse', 'Hamster', 'Lizard', 'Lifejacket', 'Goat',
+ 'Washing machine', 'Trumpet', 'Horn', 'Trombone', 'Sheep',
+ 'Tablet computer', 'Pillow', 'Kitchen & dining room table',
+ 'Parachute', 'Raven', 'Glove', 'Loveseat', 'Christmas tree',
+ 'Shellfish', 'Rifle', 'Shotgun', 'Sushi', 'Sparrow', 'Bread',
+ 'Toaster', 'Watch', 'Asparagus', 'Artichoke', 'Suitcase', 'Antelope',
+ 'Broccoli', 'Ice cream', 'Racket', 'Banana', 'Cookie', 'Cucumber',
+ 'Dragonfly', 'Lynx', 'Caterpillar', 'Light bulb', 'Office supplies',
+ 'Miniskirt', 'Skirt', 'Fireplace', 'Potato', 'Light switch',
+ 'Croissant', 'Cabbage', 'Ladybug', 'Handgun', 'Luggage and bags',
+ 'Window blind', 'Snowboard', 'Baseball bat', 'Digital clock',
+ 'Serving tray', 'Infant bed', 'Sofa bed', 'Guacamole', 'Fox', 'Pizza',
+ 'Snowplow', 'Jet ski', 'Refrigerator', 'Lantern', 'Convenience store',
+ 'Sword', 'Rugby ball', 'Owl', 'Ostrich', 'Pancake', 'Strawberry',
+ 'Carrot', 'Tart', 'Dice', 'Turkey', 'Rabbit', 'Invertebrate', 'Vase',
+ 'Stool', 'Swim cap', 'Shower', 'Clock', 'Jellyfish', 'Aircraft',
+ 'Chopsticks', 'Orange', 'Snake', 'Sewing machine', 'Kangaroo', 'Mixer',
+ 'Food processor', 'Shrimp', 'Towel', 'Porcupine', 'Jaguar', 'Cannon',
+ 'Limousine', 'Mule', 'Squirrel', 'Kitchen knife', 'Tiara', 'Tiger',
+ 'Bow and arrow', 'Candy', 'Rhinoceros', 'Shark', 'Cricket ball',
+ 'Doughnut', 'Plumbing fixture', 'Camel', 'Polar bear', 'Coin',
+ 'Printer', 'Blender', 'Giraffe', 'Billiard table', 'Kettle',
+ 'Dinosaur', 'Pineapple', 'Zucchini', 'Jug', 'Barge', 'Teapot',
+ 'Golf ball', 'Binoculars', 'Scissors', 'Hot dog', 'Door handle',
+ 'Seahorse', 'Bathtub', 'Leopard', 'Centipede', 'Grapefruit', 'Snowman',
+ 'Cheetah', 'Alarm clock', 'Grape', 'Wrench', 'Wok', 'Bell pepper',
+ 'Cake stand', 'Barrel', 'Woodpecker', 'Flute', 'Corded phone',
+ 'Willow', 'Punching bag', 'Pomegranate', 'Telephone', 'Pear',
+ 'Common fig', 'Bench', 'Wood-burning stove', 'Burrito', 'Nail',
+ 'Turtle', 'Submarine sandwich', 'Drinking straw', 'Peach', 'Popcorn',
+ 'Frying pan', 'Picnic basket', 'Honeycomb', 'Envelope', 'Mango',
+ 'Cutting board', 'Pitcher', 'Stationary bicycle', 'Dumbbell',
+ 'Personal care', 'Dog bed', 'Snowmobile', 'Oboe', 'Briefcase',
+ 'Squash', 'Tick', 'Slow cooker', 'Coffeemaker', 'Measuring cup',
+ 'Crutch', 'Stretcher', 'Screwdriver', 'Flashlight', 'Spatula',
+ 'Pressure cooker', 'Ring binder', 'Beaker', 'Torch', 'Winter melon'
+ ]
+
+
+def oid_v6_classes() -> list:
+ """Class names of Open Images V6."""
+ return [
+ 'Tortoise', 'Container', 'Magpie', 'Sea turtle', 'Football',
+ 'Ambulance', 'Ladder', 'Toothbrush', 'Syringe', 'Sink', 'Toy',
+ 'Organ (Musical Instrument)', 'Cassette deck', 'Apple', 'Human eye',
+ 'Cosmetics', 'Paddle', 'Snowman', 'Beer', 'Chopsticks', 'Human beard',
+ 'Bird', 'Parking meter', 'Traffic light', 'Croissant', 'Cucumber',
+ 'Radish', 'Towel', 'Doll', 'Skull', 'Washing machine', 'Glove', 'Tick',
+ 'Belt', 'Sunglasses', 'Banjo', 'Cart', 'Ball', 'Backpack', 'Bicycle',
+ 'Home appliance', 'Centipede', 'Boat', 'Surfboard', 'Boot',
+ 'Headphones', 'Hot dog', 'Shorts', 'Fast food', 'Bus', 'Boy',
+ 'Screwdriver', 'Bicycle wheel', 'Barge', 'Laptop', 'Miniskirt',
+ 'Drill (Tool)', 'Dress', 'Bear', 'Waffle', 'Pancake', 'Brown bear',
+ 'Woodpecker', 'Blue jay', 'Pretzel', 'Bagel', 'Tower', 'Teapot',
+ 'Person', 'Bow and arrow', 'Swimwear', 'Beehive', 'Brassiere', 'Bee',
+ 'Bat (Animal)', 'Starfish', 'Popcorn', 'Burrito', 'Chainsaw',
+ 'Balloon', 'Wrench', 'Tent', 'Vehicle registration plate', 'Lantern',
+ 'Toaster', 'Flashlight', 'Billboard', 'Tiara', 'Limousine', 'Necklace',
+ 'Carnivore', 'Scissors', 'Stairs', 'Computer keyboard', 'Printer',
+ 'Traffic sign', 'Chair', 'Shirt', 'Poster', 'Cheese', 'Sock',
+ 'Fire hydrant', 'Land vehicle', 'Earrings', 'Tie', 'Watercraft',
+ 'Cabinetry', 'Suitcase', 'Muffin', 'Bidet', 'Snack', 'Snowmobile',
+ 'Clock', 'Medical equipment', 'Cattle', 'Cello', 'Jet ski', 'Camel',
+ 'Coat', 'Suit', 'Desk', 'Cat', 'Bronze sculpture', 'Juice', 'Gondola',
+ 'Beetle', 'Cannon', 'Computer mouse', 'Cookie', 'Office building',
+ 'Fountain', 'Coin', 'Calculator', 'Cocktail', 'Computer monitor',
+ 'Box', 'Stapler', 'Christmas tree', 'Cowboy hat', 'Hiking equipment',
+ 'Studio couch', 'Drum', 'Dessert', 'Wine rack', 'Drink', 'Zucchini',
+ 'Ladle', 'Human mouth', 'Dairy Product', 'Dice', 'Oven', 'Dinosaur',
+ 'Ratchet (Device)', 'Couch', 'Cricket ball', 'Winter melon', 'Spatula',
+ 'Whiteboard', 'Pencil sharpener', 'Door', 'Hat', 'Shower', 'Eraser',
+ 'Fedora', 'Guacamole', 'Dagger', 'Scarf', 'Dolphin', 'Sombrero',
+ 'Tin can', 'Mug', 'Tap', 'Harbor seal', 'Stretcher', 'Can opener',
+ 'Goggles', 'Human body', 'Roller skates', 'Coffee cup',
+ 'Cutting board', 'Blender', 'Plumbing fixture', 'Stop sign',
+ 'Office supplies', 'Volleyball (Ball)', 'Vase', 'Slow cooker',
+ 'Wardrobe', 'Coffee', 'Whisk', 'Paper towel', 'Personal care', 'Food',
+ 'Sun hat', 'Tree house', 'Flying disc', 'Skirt', 'Gas stove',
+ 'Salt and pepper shakers', 'Mechanical fan', 'Face powder', 'Fax',
+ 'Fruit', 'French fries', 'Nightstand', 'Barrel', 'Kite', 'Tart',
+ 'Treadmill', 'Fox', 'Flag', 'French horn', 'Window blind',
+ 'Human foot', 'Golf cart', 'Jacket', 'Egg (Food)', 'Street light',
+ 'Guitar', 'Pillow', 'Human leg', 'Isopod', 'Grape', 'Human ear',
+ 'Power plugs and sockets', 'Panda', 'Giraffe', 'Woman', 'Door handle',
+ 'Rhinoceros', 'Bathtub', 'Goldfish', 'Houseplant', 'Goat',
+ 'Baseball bat', 'Baseball glove', 'Mixing bowl',
+ 'Marine invertebrates', 'Kitchen utensil', 'Light switch', 'House',
+ 'Horse', 'Stationary bicycle', 'Hammer', 'Ceiling fan', 'Sofa bed',
+ 'Adhesive tape', 'Harp', 'Sandal', 'Bicycle helmet', 'Saucer',
+ 'Harpsichord', 'Human hair', 'Heater', 'Harmonica', 'Hamster',
+ 'Curtain', 'Bed', 'Kettle', 'Fireplace', 'Scale', 'Drinking straw',
+ 'Insect', 'Hair dryer', 'Kitchenware', 'Indoor rower', 'Invertebrate',
+ 'Food processor', 'Bookcase', 'Refrigerator', 'Wood-burning stove',
+ 'Punching bag', 'Common fig', 'Cocktail shaker', 'Jaguar (Animal)',
+ 'Golf ball', 'Fashion accessory', 'Alarm clock', 'Filing cabinet',
+ 'Artichoke', 'Table', 'Tableware', 'Kangaroo', 'Koala', 'Knife',
+ 'Bottle', 'Bottle opener', 'Lynx', 'Lavender (Plant)', 'Lighthouse',
+ 'Dumbbell', 'Human head', 'Bowl', 'Humidifier', 'Porch', 'Lizard',
+ 'Billiard table', 'Mammal', 'Mouse', 'Motorcycle',
+ 'Musical instrument', 'Swim cap', 'Frying pan', 'Snowplow',
+ 'Bathroom cabinet', 'Missile', 'Bust', 'Man', 'Waffle iron', 'Milk',
+ 'Ring binder', 'Plate', 'Mobile phone', 'Baked goods', 'Mushroom',
+ 'Crutch', 'Pitcher (Container)', 'Mirror', 'Personal flotation device',
+ 'Table tennis racket', 'Pencil case', 'Musical keyboard', 'Scoreboard',
+ 'Briefcase', 'Kitchen knife', 'Nail (Construction)', 'Tennis ball',
+ 'Plastic bag', 'Oboe', 'Chest of drawers', 'Ostrich', 'Piano', 'Girl',
+ 'Plant', 'Potato', 'Hair spray', 'Sports equipment', 'Pasta',
+ 'Penguin', 'Pumpkin', 'Pear', 'Infant bed', 'Polar bear', 'Mixer',
+ 'Cupboard', 'Jacuzzi', 'Pizza', 'Digital clock', 'Pig', 'Reptile',
+ 'Rifle', 'Lipstick', 'Skateboard', 'Raven', 'High heels', 'Red panda',
+ 'Rose', 'Rabbit', 'Sculpture', 'Saxophone', 'Shotgun', 'Seafood',
+ 'Submarine sandwich', 'Snowboard', 'Sword', 'Picture frame', 'Sushi',
+ 'Loveseat', 'Ski', 'Squirrel', 'Tripod', 'Stethoscope', 'Submarine',
+ 'Scorpion', 'Segway', 'Training bench', 'Snake', 'Coffee table',
+ 'Skyscraper', 'Sheep', 'Television', 'Trombone', 'Tea', 'Tank', 'Taco',
+ 'Telephone', 'Torch', 'Tiger', 'Strawberry', 'Trumpet', 'Tree',
+ 'Tomato', 'Train', 'Tool', 'Picnic basket', 'Cooking spray',
+ 'Trousers', 'Bowling equipment', 'Football helmet', 'Truck',
+ 'Measuring cup', 'Coffeemaker', 'Violin', 'Vehicle', 'Handbag',
+ 'Paper cutter', 'Wine', 'Weapon', 'Wheel', 'Worm', 'Wok', 'Whale',
+ 'Zebra', 'Auto part', 'Jug', 'Pizza cutter', 'Cream', 'Monkey', 'Lion',
+ 'Bread', 'Platter', 'Chicken', 'Eagle', 'Helicopter', 'Owl', 'Duck',
+ 'Turtle', 'Hippopotamus', 'Crocodile', 'Toilet', 'Toilet paper',
+ 'Squid', 'Clothing', 'Footwear', 'Lemon', 'Spider', 'Deer', 'Frog',
+ 'Banana', 'Rocket', 'Wine glass', 'Countertop', 'Tablet computer',
+ 'Waste container', 'Swimming pool', 'Dog', 'Book', 'Elephant', 'Shark',
+ 'Candle', 'Leopard', 'Axe', 'Hand dryer', 'Soap dispenser',
+ 'Porcupine', 'Flower', 'Canary', 'Cheetah', 'Palm tree', 'Hamburger',
+ 'Maple', 'Building', 'Fish', 'Lobster', 'Garden Asparagus',
+ 'Furniture', 'Hedgehog', 'Airplane', 'Spoon', 'Otter', 'Bull',
+ 'Oyster', 'Horizontal bar', 'Convenience store', 'Bomb', 'Bench',
+ 'Ice cream', 'Caterpillar', 'Butterfly', 'Parachute', 'Orange',
+ 'Antelope', 'Beaker', 'Moths and butterflies', 'Window', 'Closet',
+ 'Castle', 'Jellyfish', 'Goose', 'Mule', 'Swan', 'Peach', 'Coconut',
+ 'Seat belt', 'Raccoon', 'Chisel', 'Fork', 'Lamp', 'Camera',
+ 'Squash (Plant)', 'Racket', 'Human face', 'Human arm', 'Vegetable',
+ 'Diaper', 'Unicycle', 'Falcon', 'Chime', 'Snail', 'Shellfish',
+ 'Cabbage', 'Carrot', 'Mango', 'Jeans', 'Flowerpot', 'Pineapple',
+ 'Drawer', 'Stool', 'Envelope', 'Cake', 'Dragonfly', 'Common sunflower',
+ 'Microwave oven', 'Honeycomb', 'Marine mammal', 'Sea lion', 'Ladybug',
+ 'Shelf', 'Watch', 'Candy', 'Salad', 'Parrot', 'Handgun', 'Sparrow',
+ 'Van', 'Grinder', 'Spice rack', 'Light bulb', 'Corded phone',
+ 'Sports uniform', 'Tennis racket', 'Wall clock', 'Serving tray',
+ 'Kitchen & dining room table', 'Dog bed', 'Cake stand',
+ 'Cat furniture', 'Bathroom accessory', 'Facial tissue holder',
+ 'Pressure cooker', 'Kitchen appliance', 'Tire', 'Ruler',
+ 'Luggage and bags', 'Microphone', 'Broccoli', 'Umbrella', 'Pastry',
+ 'Grapefruit', 'Band-aid', 'Animal', 'Bell pepper', 'Turkey', 'Lily',
+ 'Pomegranate', 'Doughnut', 'Glasses', 'Human nose', 'Pen', 'Ant',
+ 'Car', 'Aircraft', 'Human hand', 'Skunk', 'Teddy bear', 'Watermelon',
+ 'Cantaloupe', 'Dishwasher', 'Flute', 'Balance beam', 'Sandwich',
+ 'Shrimp', 'Sewing machine', 'Binoculars', 'Rays and skates', 'Ipod',
+ 'Accordion', 'Willow', 'Crab', 'Crown', 'Seahorse', 'Perfume',
+ 'Alpaca', 'Taxi', 'Canoe', 'Remote control', 'Wheelchair',
+ 'Rugby ball', 'Armadillo', 'Maracas', 'Helmet'
+ ]
+
+
+def objects365v1_classes() -> list:
+ """Class names of Objects365 V1."""
+ return [
+ 'person', 'sneakers', 'chair', 'hat', 'lamp', 'bottle',
+ 'cabinet/shelf', 'cup', 'car', 'glasses', 'picture/frame', 'desk',
+ 'handbag', 'street lights', 'book', 'plate', 'helmet', 'leather shoes',
+ 'pillow', 'glove', 'potted plant', 'bracelet', 'flower', 'tv',
+ 'storage box', 'vase', 'bench', 'wine glass', 'boots', 'bowl',
+ 'dining table', 'umbrella', 'boat', 'flag', 'speaker', 'trash bin/can',
+ 'stool', 'backpack', 'couch', 'belt', 'carpet', 'basket',
+ 'towel/napkin', 'slippers', 'barrel/bucket', 'coffee table', 'suv',
+ 'toy', 'tie', 'bed', 'traffic light', 'pen/pencil', 'microphone',
+ 'sandals', 'canned', 'necklace', 'mirror', 'faucet', 'bicycle',
+ 'bread', 'high heels', 'ring', 'van', 'watch', 'sink', 'horse', 'fish',
+ 'apple', 'camera', 'candle', 'teddy bear', 'cake', 'motorcycle',
+ 'wild bird', 'laptop', 'knife', 'traffic sign', 'cell phone', 'paddle',
+ 'truck', 'cow', 'power outlet', 'clock', 'drum', 'fork', 'bus',
+ 'hanger', 'nightstand', 'pot/pan', 'sheep', 'guitar', 'traffic cone',
+ 'tea pot', 'keyboard', 'tripod', 'hockey', 'fan', 'dog', 'spoon',
+ 'blackboard/whiteboard', 'balloon', 'air conditioner', 'cymbal',
+ 'mouse', 'telephone', 'pickup truck', 'orange', 'banana', 'airplane',
+ 'luggage', 'skis', 'soccer', 'trolley', 'oven', 'remote',
+ 'baseball glove', 'paper towel', 'refrigerator', 'train', 'tomato',
+ 'machinery vehicle', 'tent', 'shampoo/shower gel', 'head phone',
+ 'lantern', 'donut', 'cleaning products', 'sailboat', 'tangerine',
+ 'pizza', 'kite', 'computer box', 'elephant', 'toiletries', 'gas stove',
+ 'broccoli', 'toilet', 'stroller', 'shovel', 'baseball bat',
+ 'microwave', 'skateboard', 'surfboard', 'surveillance camera', 'gun',
+ 'life saver', 'cat', 'lemon', 'liquid soap', 'zebra', 'duck',
+ 'sports car', 'giraffe', 'pumpkin', 'piano', 'stop sign', 'radiator',
+ 'converter', 'tissue ', 'carrot', 'washing machine', 'vent', 'cookies',
+ 'cutting/chopping board', 'tennis racket', 'candy',
+ 'skating and skiing shoes', 'scissors', 'folder', 'baseball',
+ 'strawberry', 'bow tie', 'pigeon', 'pepper', 'coffee machine',
+ 'bathtub', 'snowboard', 'suitcase', 'grapes', 'ladder', 'pear',
+ 'american football', 'basketball', 'potato', 'paint brush', 'printer',
+ 'billiards', 'fire hydrant', 'goose', 'projector', 'sausage',
+ 'fire extinguisher', 'extension cord', 'facial mask', 'tennis ball',
+ 'chopsticks', 'electronic stove and gas stove', 'pie', 'frisbee',
+ 'kettle', 'hamburger', 'golf club', 'cucumber', 'clutch', 'blender',
+ 'tong', 'slide', 'hot dog', 'toothbrush', 'facial cleanser', 'mango',
+ 'deer', 'egg', 'violin', 'marker', 'ship', 'chicken', 'onion',
+ 'ice cream', 'tape', 'wheelchair', 'plum', 'bar soap', 'scale',
+ 'watermelon', 'cabbage', 'router/modem', 'golf ball', 'pine apple',
+ 'crane', 'fire truck', 'peach', 'cello', 'notepaper', 'tricycle',
+ 'toaster', 'helicopter', 'green beans', 'brush', 'carriage', 'cigar',
+ 'earphone', 'penguin', 'hurdle', 'swing', 'radio', 'CD',
+ 'parking meter', 'swan', 'garlic', 'french fries', 'horn', 'avocado',
+ 'saxophone', 'trumpet', 'sandwich', 'cue', 'kiwi fruit', 'bear',
+ 'fishing rod', 'cherry', 'tablet', 'green vegetables', 'nuts', 'corn',
+ 'key', 'screwdriver', 'globe', 'broom', 'pliers', 'volleyball',
+ 'hammer', 'eggplant', 'trophy', 'dates', 'board eraser', 'rice',
+ 'tape measure/ruler', 'dumbbell', 'hamimelon', 'stapler', 'camel',
+ 'lettuce', 'goldfish', 'meat balls', 'medal', 'toothpaste', 'antelope',
+ 'shrimp', 'rickshaw', 'trombone', 'pomegranate', 'coconut',
+ 'jellyfish', 'mushroom', 'calculator', 'treadmill', 'butterfly',
+ 'egg tart', 'cheese', 'pig', 'pomelo', 'race car', 'rice cooker',
+ 'tuba', 'crosswalk sign', 'papaya', 'hair drier', 'green onion',
+ 'chips', 'dolphin', 'sushi', 'urinal', 'donkey', 'electric drill',
+ 'spring rolls', 'tortoise/turtle', 'parrot', 'flute', 'measuring cup',
+ 'shark', 'steak', 'poker card', 'binoculars', 'llama', 'radish',
+ 'noodles', 'yak', 'mop', 'crab', 'microscope', 'barbell', 'bread/bun',
+ 'baozi', 'lion', 'red cabbage', 'polar bear', 'lighter', 'seal',
+ 'mangosteen', 'comb', 'eraser', 'pitaya', 'scallop', 'pencil case',
+ 'saw', 'table tennis paddle', 'okra', 'starfish', 'eagle', 'monkey',
+ 'durian', 'game board', 'rabbit', 'french horn', 'ambulance',
+ 'asparagus', 'hoverboard', 'pasta', 'target', 'hotair balloon',
+ 'chainsaw', 'lobster', 'iron', 'flashlight'
+ ]
+
+
+def objects365v2_classes() -> list:
+ """Class names of Objects365 V2."""
+ return [
+ 'Person', 'Sneakers', 'Chair', 'Other Shoes', 'Hat', 'Car', 'Lamp',
+ 'Glasses', 'Bottle', 'Desk', 'Cup', 'Street Lights', 'Cabinet/shelf',
+ 'Handbag/Satchel', 'Bracelet', 'Plate', 'Picture/Frame', 'Helmet',
+ 'Book', 'Gloves', 'Storage box', 'Boat', 'Leather Shoes', 'Flower',
+ 'Bench', 'Potted Plant', 'Bowl/Basin', 'Flag', 'Pillow', 'Boots',
+ 'Vase', 'Microphone', 'Necklace', 'Ring', 'SUV', 'Wine Glass', 'Belt',
+ 'Moniter/TV', 'Backpack', 'Umbrella', 'Traffic Light', 'Speaker',
+ 'Watch', 'Tie', 'Trash bin Can', 'Slippers', 'Bicycle', 'Stool',
+ 'Barrel/bucket', 'Van', 'Couch', 'Sandals', 'Bakset', 'Drum',
+ 'Pen/Pencil', 'Bus', 'Wild Bird', 'High Heels', 'Motorcycle', 'Guitar',
+ 'Carpet', 'Cell Phone', 'Bread', 'Camera', 'Canned', 'Truck',
+ 'Traffic cone', 'Cymbal', 'Lifesaver', 'Towel', 'Stuffed Toy',
+ 'Candle', 'Sailboat', 'Laptop', 'Awning', 'Bed', 'Faucet', 'Tent',
+ 'Horse', 'Mirror', 'Power outlet', 'Sink', 'Apple', 'Air Conditioner',
+ 'Knife', 'Hockey Stick', 'Paddle', 'Pickup Truck', 'Fork',
+ 'Traffic Sign', 'Ballon', 'Tripod', 'Dog', 'Spoon', 'Clock', 'Pot',
+ 'Cow', 'Cake', 'Dinning Table', 'Sheep', 'Hanger',
+ 'Blackboard/Whiteboard', 'Napkin', 'Other Fish', 'Orange/Tangerine',
+ 'Toiletry', 'Keyboard', 'Tomato', 'Lantern', 'Machinery Vehicle',
+ 'Fan', 'Green Vegetables', 'Banana', 'Baseball Glove', 'Airplane',
+ 'Mouse', 'Train', 'Pumpkin', 'Soccer', 'Skiboard', 'Luggage',
+ 'Nightstand', 'Tea pot', 'Telephone', 'Trolley', 'Head Phone',
+ 'Sports Car', 'Stop Sign', 'Dessert', 'Scooter', 'Stroller', 'Crane',
+ 'Remote', 'Refrigerator', 'Oven', 'Lemon', 'Duck', 'Baseball Bat',
+ 'Surveillance Camera', 'Cat', 'Jug', 'Broccoli', 'Piano', 'Pizza',
+ 'Elephant', 'Skateboard', 'Surfboard', 'Gun',
+ 'Skating and Skiing shoes', 'Gas stove', 'Donut', 'Bow Tie', 'Carrot',
+ 'Toilet', 'Kite', 'Strawberry', 'Other Balls', 'Shovel', 'Pepper',
+ 'Computer Box', 'Toilet Paper', 'Cleaning Products', 'Chopsticks',
+ 'Microwave', 'Pigeon', 'Baseball', 'Cutting/chopping Board',
+ 'Coffee Table', 'Side Table', 'Scissors', 'Marker', 'Pie', 'Ladder',
+ 'Snowboard', 'Cookies', 'Radiator', 'Fire Hydrant', 'Basketball',
+ 'Zebra', 'Grape', 'Giraffe', 'Potato', 'Sausage', 'Tricycle', 'Violin',
+ 'Egg', 'Fire Extinguisher', 'Candy', 'Fire Truck', 'Billards',
+ 'Converter', 'Bathtub', 'Wheelchair', 'Golf Club', 'Briefcase',
+ 'Cucumber', 'Cigar/Cigarette ', 'Paint Brush', 'Pear', 'Heavy Truck',
+ 'Hamburger', 'Extractor', 'Extention Cord', 'Tong', 'Tennis Racket',
+ 'Folder', 'American Football', 'earphone', 'Mask', 'Kettle', 'Tennis',
+ 'Ship', 'Swing', 'Coffee Machine', 'Slide', 'Carriage', 'Onion',
+ 'Green beans', 'Projector', 'Frisbee',
+ 'Washing Machine/Drying Machine', 'Chicken', 'Printer', 'Watermelon',
+ 'Saxophone', 'Tissue', 'Toothbrush', 'Ice cream', 'Hotair ballon',
+ 'Cello', 'French Fries', 'Scale', 'Trophy', 'Cabbage', 'Hot dog',
+ 'Blender', 'Peach', 'Rice', 'Wallet/Purse', 'Volleyball', 'Deer',
+ 'Goose', 'Tape', 'Tablet', 'Cosmetics', 'Trumpet', 'Pineapple',
+ 'Golf Ball', 'Ambulance', 'Parking meter', 'Mango', 'Key', 'Hurdle',
+ 'Fishing Rod', 'Medal', 'Flute', 'Brush', 'Penguin', 'Megaphone',
+ 'Corn', 'Lettuce', 'Garlic', 'Swan', 'Helicopter', 'Green Onion',
+ 'Sandwich', 'Nuts', 'Speed Limit Sign', 'Induction Cooker', 'Broom',
+ 'Trombone', 'Plum', 'Rickshaw', 'Goldfish', 'Kiwi fruit',
+ 'Router/modem', 'Poker Card', 'Toaster', 'Shrimp', 'Sushi', 'Cheese',
+ 'Notepaper', 'Cherry', 'Pliers', 'CD', 'Pasta', 'Hammer', 'Cue',
+ 'Avocado', 'Hamimelon', 'Flask', 'Mushroon', 'Screwdriver', 'Soap',
+ 'Recorder', 'Bear', 'Eggplant', 'Board Eraser', 'Coconut',
+ 'Tape Measur/ Ruler', 'Pig', 'Showerhead', 'Globe', 'Chips', 'Steak',
+ 'Crosswalk Sign', 'Stapler', 'Campel', 'Formula 1 ', 'Pomegranate',
+ 'Dishwasher', 'Crab', 'Hoverboard', 'Meat ball', 'Rice Cooker', 'Tuba',
+ 'Calculator', 'Papaya', 'Antelope', 'Parrot', 'Seal', 'Buttefly',
+ 'Dumbbell', 'Donkey', 'Lion', 'Urinal', 'Dolphin', 'Electric Drill',
+ 'Hair Dryer', 'Egg tart', 'Jellyfish', 'Treadmill', 'Lighter',
+ 'Grapefruit', 'Game board', 'Mop', 'Radish', 'Baozi', 'Target',
+ 'French', 'Spring Rolls', 'Monkey', 'Rabbit', 'Pencil Case', 'Yak',
+ 'Red Cabbage', 'Binoculars', 'Asparagus', 'Barbell', 'Scallop',
+ 'Noddles', 'Comb', 'Dumpling', 'Oyster', 'Table Teniis paddle',
+ 'Cosmetics Brush/Eyeliner Pencil', 'Chainsaw', 'Eraser', 'Lobster',
+ 'Durian', 'Okra', 'Lipstick', 'Cosmetics Mirror', 'Curling',
+ 'Table Tennis '
+ ]
+
+
+def lvis_classes() -> list:
+ """Class names of LVIS."""
+ return [
+ 'aerosol_can', 'air_conditioner', 'airplane', 'alarm_clock', 'alcohol',
+ 'alligator', 'almond', 'ambulance', 'amplifier', 'anklet', 'antenna',
+ 'apple', 'applesauce', 'apricot', 'apron', 'aquarium',
+ 'arctic_(type_of_shoe)', 'armband', 'armchair', 'armoire', 'armor',
+ 'artichoke', 'trash_can', 'ashtray', 'asparagus', 'atomizer',
+ 'avocado', 'award', 'awning', 'ax', 'baboon', 'baby_buggy',
+ 'basketball_backboard', 'backpack', 'handbag', 'suitcase', 'bagel',
+ 'bagpipe', 'baguet', 'bait', 'ball', 'ballet_skirt', 'balloon',
+ 'bamboo', 'banana', 'Band_Aid', 'bandage', 'bandanna', 'banjo',
+ 'banner', 'barbell', 'barge', 'barrel', 'barrette', 'barrow',
+ 'baseball_base', 'baseball', 'baseball_bat', 'baseball_cap',
+ 'baseball_glove', 'basket', 'basketball', 'bass_horn', 'bat_(animal)',
+ 'bath_mat', 'bath_towel', 'bathrobe', 'bathtub', 'batter_(food)',
+ 'battery', 'beachball', 'bead', 'bean_curd', 'beanbag', 'beanie',
+ 'bear', 'bed', 'bedpan', 'bedspread', 'cow', 'beef_(food)', 'beeper',
+ 'beer_bottle', 'beer_can', 'beetle', 'bell', 'bell_pepper', 'belt',
+ 'belt_buckle', 'bench', 'beret', 'bib', 'Bible', 'bicycle', 'visor',
+ 'billboard', 'binder', 'binoculars', 'bird', 'birdfeeder', 'birdbath',
+ 'birdcage', 'birdhouse', 'birthday_cake', 'birthday_card',
+ 'pirate_flag', 'black_sheep', 'blackberry', 'blackboard', 'blanket',
+ 'blazer', 'blender', 'blimp', 'blinker', 'blouse', 'blueberry',
+ 'gameboard', 'boat', 'bob', 'bobbin', 'bobby_pin', 'boiled_egg',
+ 'bolo_tie', 'deadbolt', 'bolt', 'bonnet', 'book', 'bookcase',
+ 'booklet', 'bookmark', 'boom_microphone', 'boot', 'bottle',
+ 'bottle_opener', 'bouquet', 'bow_(weapon)', 'bow_(decorative_ribbons)',
+ 'bow-tie', 'bowl', 'pipe_bowl', 'bowler_hat', 'bowling_ball', 'box',
+ 'boxing_glove', 'suspenders', 'bracelet', 'brass_plaque', 'brassiere',
+ 'bread-bin', 'bread', 'breechcloth', 'bridal_gown', 'briefcase',
+ 'broccoli', 'broach', 'broom', 'brownie', 'brussels_sprouts',
+ 'bubble_gum', 'bucket', 'horse_buggy', 'bull', 'bulldog', 'bulldozer',
+ 'bullet_train', 'bulletin_board', 'bulletproof_vest', 'bullhorn',
+ 'bun', 'bunk_bed', 'buoy', 'burrito', 'bus_(vehicle)', 'business_card',
+ 'butter', 'butterfly', 'button', 'cab_(taxi)', 'cabana', 'cabin_car',
+ 'cabinet', 'locker', 'cake', 'calculator', 'calendar', 'calf',
+ 'camcorder', 'camel', 'camera', 'camera_lens', 'camper_(vehicle)',
+ 'can', 'can_opener', 'candle', 'candle_holder', 'candy_bar',
+ 'candy_cane', 'walking_cane', 'canister', 'canoe', 'cantaloup',
+ 'canteen', 'cap_(headwear)', 'bottle_cap', 'cape', 'cappuccino',
+ 'car_(automobile)', 'railcar_(part_of_a_train)', 'elevator_car',
+ 'car_battery', 'identity_card', 'card', 'cardigan', 'cargo_ship',
+ 'carnation', 'horse_carriage', 'carrot', 'tote_bag', 'cart', 'carton',
+ 'cash_register', 'casserole', 'cassette', 'cast', 'cat', 'cauliflower',
+ 'cayenne_(spice)', 'CD_player', 'celery', 'cellular_telephone',
+ 'chain_mail', 'chair', 'chaise_longue', 'chalice', 'chandelier',
+ 'chap', 'checkbook', 'checkerboard', 'cherry', 'chessboard',
+ 'chicken_(animal)', 'chickpea', 'chili_(vegetable)', 'chime',
+ 'chinaware', 'crisp_(potato_chip)', 'poker_chip', 'chocolate_bar',
+ 'chocolate_cake', 'chocolate_milk', 'chocolate_mousse', 'choker',
+ 'chopping_board', 'chopstick', 'Christmas_tree', 'slide', 'cider',
+ 'cigar_box', 'cigarette', 'cigarette_case', 'cistern', 'clarinet',
+ 'clasp', 'cleansing_agent', 'cleat_(for_securing_rope)', 'clementine',
+ 'clip', 'clipboard', 'clippers_(for_plants)', 'cloak', 'clock',
+ 'clock_tower', 'clothes_hamper', 'clothespin', 'clutch_bag', 'coaster',
+ 'coat', 'coat_hanger', 'coatrack', 'cock', 'cockroach',
+ 'cocoa_(beverage)', 'coconut', 'coffee_maker', 'coffee_table',
+ 'coffeepot', 'coil', 'coin', 'colander', 'coleslaw',
+ 'coloring_material', 'combination_lock', 'pacifier', 'comic_book',
+ 'compass', 'computer_keyboard', 'condiment', 'cone', 'control',
+ 'convertible_(automobile)', 'sofa_bed', 'cooker', 'cookie',
+ 'cooking_utensil', 'cooler_(for_food)', 'cork_(bottle_plug)',
+ 'corkboard', 'corkscrew', 'edible_corn', 'cornbread', 'cornet',
+ 'cornice', 'cornmeal', 'corset', 'costume', 'cougar', 'coverall',
+ 'cowbell', 'cowboy_hat', 'crab_(animal)', 'crabmeat', 'cracker',
+ 'crape', 'crate', 'crayon', 'cream_pitcher', 'crescent_roll', 'crib',
+ 'crock_pot', 'crossbar', 'crouton', 'crow', 'crowbar', 'crown',
+ 'crucifix', 'cruise_ship', 'police_cruiser', 'crumb', 'crutch',
+ 'cub_(animal)', 'cube', 'cucumber', 'cufflink', 'cup', 'trophy_cup',
+ 'cupboard', 'cupcake', 'hair_curler', 'curling_iron', 'curtain',
+ 'cushion', 'cylinder', 'cymbal', 'dagger', 'dalmatian', 'dartboard',
+ 'date_(fruit)', 'deck_chair', 'deer', 'dental_floss', 'desk',
+ 'detergent', 'diaper', 'diary', 'die', 'dinghy', 'dining_table', 'tux',
+ 'dish', 'dish_antenna', 'dishrag', 'dishtowel', 'dishwasher',
+ 'dishwasher_detergent', 'dispenser', 'diving_board', 'Dixie_cup',
+ 'dog', 'dog_collar', 'doll', 'dollar', 'dollhouse', 'dolphin',
+ 'domestic_ass', 'doorknob', 'doormat', 'doughnut', 'dove', 'dragonfly',
+ 'drawer', 'underdrawers', 'dress', 'dress_hat', 'dress_suit',
+ 'dresser', 'drill', 'drone', 'dropper', 'drum_(musical_instrument)',
+ 'drumstick', 'duck', 'duckling', 'duct_tape', 'duffel_bag', 'dumbbell',
+ 'dumpster', 'dustpan', 'eagle', 'earphone', 'earplug', 'earring',
+ 'easel', 'eclair', 'eel', 'egg', 'egg_roll', 'egg_yolk', 'eggbeater',
+ 'eggplant', 'electric_chair', 'refrigerator', 'elephant', 'elk',
+ 'envelope', 'eraser', 'escargot', 'eyepatch', 'falcon', 'fan',
+ 'faucet', 'fedora', 'ferret', 'Ferris_wheel', 'ferry', 'fig_(fruit)',
+ 'fighter_jet', 'figurine', 'file_cabinet', 'file_(tool)', 'fire_alarm',
+ 'fire_engine', 'fire_extinguisher', 'fire_hose', 'fireplace',
+ 'fireplug', 'first-aid_kit', 'fish', 'fish_(food)', 'fishbowl',
+ 'fishing_rod', 'flag', 'flagpole', 'flamingo', 'flannel', 'flap',
+ 'flash', 'flashlight', 'fleece', 'flip-flop_(sandal)',
+ 'flipper_(footwear)', 'flower_arrangement', 'flute_glass', 'foal',
+ 'folding_chair', 'food_processor', 'football_(American)',
+ 'football_helmet', 'footstool', 'fork', 'forklift', 'freight_car',
+ 'French_toast', 'freshener', 'frisbee', 'frog', 'fruit_juice',
+ 'frying_pan', 'fudge', 'funnel', 'futon', 'gag', 'garbage',
+ 'garbage_truck', 'garden_hose', 'gargle', 'gargoyle', 'garlic',
+ 'gasmask', 'gazelle', 'gelatin', 'gemstone', 'generator',
+ 'giant_panda', 'gift_wrap', 'ginger', 'giraffe', 'cincture',
+ 'glass_(drink_container)', 'globe', 'glove', 'goat', 'goggles',
+ 'goldfish', 'golf_club', 'golfcart', 'gondola_(boat)', 'goose',
+ 'gorilla', 'gourd', 'grape', 'grater', 'gravestone', 'gravy_boat',
+ 'green_bean', 'green_onion', 'griddle', 'grill', 'grits', 'grizzly',
+ 'grocery_bag', 'guitar', 'gull', 'gun', 'hairbrush', 'hairnet',
+ 'hairpin', 'halter_top', 'ham', 'hamburger', 'hammer', 'hammock',
+ 'hamper', 'hamster', 'hair_dryer', 'hand_glass', 'hand_towel',
+ 'handcart', 'handcuff', 'handkerchief', 'handle', 'handsaw',
+ 'hardback_book', 'harmonium', 'hat', 'hatbox', 'veil', 'headband',
+ 'headboard', 'headlight', 'headscarf', 'headset',
+ 'headstall_(for_horses)', 'heart', 'heater', 'helicopter', 'helmet',
+ 'heron', 'highchair', 'hinge', 'hippopotamus', 'hockey_stick', 'hog',
+ 'home_plate_(baseball)', 'honey', 'fume_hood', 'hook', 'hookah',
+ 'hornet', 'horse', 'hose', 'hot-air_balloon', 'hotplate', 'hot_sauce',
+ 'hourglass', 'houseboat', 'hummingbird', 'hummus', 'polar_bear',
+ 'icecream', 'popsicle', 'ice_maker', 'ice_pack', 'ice_skate',
+ 'igniter', 'inhaler', 'iPod', 'iron_(for_clothing)', 'ironing_board',
+ 'jacket', 'jam', 'jar', 'jean', 'jeep', 'jelly_bean', 'jersey',
+ 'jet_plane', 'jewel', 'jewelry', 'joystick', 'jumpsuit', 'kayak',
+ 'keg', 'kennel', 'kettle', 'key', 'keycard', 'kilt', 'kimono',
+ 'kitchen_sink', 'kitchen_table', 'kite', 'kitten', 'kiwi_fruit',
+ 'knee_pad', 'knife', 'knitting_needle', 'knob', 'knocker_(on_a_door)',
+ 'koala', 'lab_coat', 'ladder', 'ladle', 'ladybug', 'lamb_(animal)',
+ 'lamb-chop', 'lamp', 'lamppost', 'lampshade', 'lantern', 'lanyard',
+ 'laptop_computer', 'lasagna', 'latch', 'lawn_mower', 'leather',
+ 'legging_(clothing)', 'Lego', 'legume', 'lemon', 'lemonade', 'lettuce',
+ 'license_plate', 'life_buoy', 'life_jacket', 'lightbulb',
+ 'lightning_rod', 'lime', 'limousine', 'lion', 'lip_balm', 'liquor',
+ 'lizard', 'log', 'lollipop', 'speaker_(stereo_equipment)', 'loveseat',
+ 'machine_gun', 'magazine', 'magnet', 'mail_slot', 'mailbox_(at_home)',
+ 'mallard', 'mallet', 'mammoth', 'manatee', 'mandarin_orange', 'manger',
+ 'manhole', 'map', 'marker', 'martini', 'mascot', 'mashed_potato',
+ 'masher', 'mask', 'mast', 'mat_(gym_equipment)', 'matchbox',
+ 'mattress', 'measuring_cup', 'measuring_stick', 'meatball', 'medicine',
+ 'melon', 'microphone', 'microscope', 'microwave_oven', 'milestone',
+ 'milk', 'milk_can', 'milkshake', 'minivan', 'mint_candy', 'mirror',
+ 'mitten', 'mixer_(kitchen_tool)', 'money',
+ 'monitor_(computer_equipment) computer_monitor', 'monkey', 'motor',
+ 'motor_scooter', 'motor_vehicle', 'motorcycle', 'mound_(baseball)',
+ 'mouse_(computer_equipment)', 'mousepad', 'muffin', 'mug', 'mushroom',
+ 'music_stool', 'musical_instrument', 'nailfile', 'napkin',
+ 'neckerchief', 'necklace', 'necktie', 'needle', 'nest', 'newspaper',
+ 'newsstand', 'nightshirt', 'nosebag_(for_animals)',
+ 'noseband_(for_animals)', 'notebook', 'notepad', 'nut', 'nutcracker',
+ 'oar', 'octopus_(food)', 'octopus_(animal)', 'oil_lamp', 'olive_oil',
+ 'omelet', 'onion', 'orange_(fruit)', 'orange_juice', 'ostrich',
+ 'ottoman', 'oven', 'overalls_(clothing)', 'owl', 'packet', 'inkpad',
+ 'pad', 'paddle', 'padlock', 'paintbrush', 'painting', 'pajamas',
+ 'palette', 'pan_(for_cooking)', 'pan_(metal_container)', 'pancake',
+ 'pantyhose', 'papaya', 'paper_plate', 'paper_towel', 'paperback_book',
+ 'paperweight', 'parachute', 'parakeet', 'parasail_(sports)', 'parasol',
+ 'parchment', 'parka', 'parking_meter', 'parrot',
+ 'passenger_car_(part_of_a_train)', 'passenger_ship', 'passport',
+ 'pastry', 'patty_(food)', 'pea_(food)', 'peach', 'peanut_butter',
+ 'pear', 'peeler_(tool_for_fruit_and_vegetables)', 'wooden_leg',
+ 'pegboard', 'pelican', 'pen', 'pencil', 'pencil_box',
+ 'pencil_sharpener', 'pendulum', 'penguin', 'pennant', 'penny_(coin)',
+ 'pepper', 'pepper_mill', 'perfume', 'persimmon', 'person', 'pet',
+ 'pew_(church_bench)', 'phonebook', 'phonograph_record', 'piano',
+ 'pickle', 'pickup_truck', 'pie', 'pigeon', 'piggy_bank', 'pillow',
+ 'pin_(non_jewelry)', 'pineapple', 'pinecone', 'ping-pong_ball',
+ 'pinwheel', 'tobacco_pipe', 'pipe', 'pistol', 'pita_(bread)',
+ 'pitcher_(vessel_for_liquid)', 'pitchfork', 'pizza', 'place_mat',
+ 'plate', 'platter', 'playpen', 'pliers', 'plow_(farm_equipment)',
+ 'plume', 'pocket_watch', 'pocketknife', 'poker_(fire_stirring_tool)',
+ 'pole', 'polo_shirt', 'poncho', 'pony', 'pool_table', 'pop_(soda)',
+ 'postbox_(public)', 'postcard', 'poster', 'pot', 'flowerpot', 'potato',
+ 'potholder', 'pottery', 'pouch', 'power_shovel', 'prawn', 'pretzel',
+ 'printer', 'projectile_(weapon)', 'projector', 'propeller', 'prune',
+ 'pudding', 'puffer_(fish)', 'puffin', 'pug-dog', 'pumpkin', 'puncher',
+ 'puppet', 'puppy', 'quesadilla', 'quiche', 'quilt', 'rabbit',
+ 'race_car', 'racket', 'radar', 'radiator', 'radio_receiver', 'radish',
+ 'raft', 'rag_doll', 'raincoat', 'ram_(animal)', 'raspberry', 'rat',
+ 'razorblade', 'reamer_(juicer)', 'rearview_mirror', 'receipt',
+ 'recliner', 'record_player', 'reflector', 'remote_control',
+ 'rhinoceros', 'rib_(food)', 'rifle', 'ring', 'river_boat', 'road_map',
+ 'robe', 'rocking_chair', 'rodent', 'roller_skate', 'Rollerblade',
+ 'rolling_pin', 'root_beer', 'router_(computer_equipment)',
+ 'rubber_band', 'runner_(carpet)', 'plastic_bag',
+ 'saddle_(on_an_animal)', 'saddle_blanket', 'saddlebag', 'safety_pin',
+ 'sail', 'salad', 'salad_plate', 'salami', 'salmon_(fish)',
+ 'salmon_(food)', 'salsa', 'saltshaker', 'sandal_(type_of_shoe)',
+ 'sandwich', 'satchel', 'saucepan', 'saucer', 'sausage', 'sawhorse',
+ 'saxophone', 'scale_(measuring_instrument)', 'scarecrow', 'scarf',
+ 'school_bus', 'scissors', 'scoreboard', 'scraper', 'screwdriver',
+ 'scrubbing_brush', 'sculpture', 'seabird', 'seahorse', 'seaplane',
+ 'seashell', 'sewing_machine', 'shaker', 'shampoo', 'shark',
+ 'sharpener', 'Sharpie', 'shaver_(electric)', 'shaving_cream', 'shawl',
+ 'shears', 'sheep', 'shepherd_dog', 'sherbert', 'shield', 'shirt',
+ 'shoe', 'shopping_bag', 'shopping_cart', 'short_pants', 'shot_glass',
+ 'shoulder_bag', 'shovel', 'shower_head', 'shower_cap',
+ 'shower_curtain', 'shredder_(for_paper)', 'signboard', 'silo', 'sink',
+ 'skateboard', 'skewer', 'ski', 'ski_boot', 'ski_parka', 'ski_pole',
+ 'skirt', 'skullcap', 'sled', 'sleeping_bag', 'sling_(bandage)',
+ 'slipper_(footwear)', 'smoothie', 'snake', 'snowboard', 'snowman',
+ 'snowmobile', 'soap', 'soccer_ball', 'sock', 'sofa', 'softball',
+ 'solar_array', 'sombrero', 'soup', 'soup_bowl', 'soupspoon',
+ 'sour_cream', 'soya_milk', 'space_shuttle', 'sparkler_(fireworks)',
+ 'spatula', 'spear', 'spectacles', 'spice_rack', 'spider', 'crawfish',
+ 'sponge', 'spoon', 'sportswear', 'spotlight', 'squid_(food)',
+ 'squirrel', 'stagecoach', 'stapler_(stapling_machine)', 'starfish',
+ 'statue_(sculpture)', 'steak_(food)', 'steak_knife', 'steering_wheel',
+ 'stepladder', 'step_stool', 'stereo_(sound_system)', 'stew', 'stirrer',
+ 'stirrup', 'stool', 'stop_sign', 'brake_light', 'stove', 'strainer',
+ 'strap', 'straw_(for_drinking)', 'strawberry', 'street_sign',
+ 'streetlight', 'string_cheese', 'stylus', 'subwoofer', 'sugar_bowl',
+ 'sugarcane_(plant)', 'suit_(clothing)', 'sunflower', 'sunglasses',
+ 'sunhat', 'surfboard', 'sushi', 'mop', 'sweat_pants', 'sweatband',
+ 'sweater', 'sweatshirt', 'sweet_potato', 'swimsuit', 'sword',
+ 'syringe', 'Tabasco_sauce', 'table-tennis_table', 'table',
+ 'table_lamp', 'tablecloth', 'tachometer', 'taco', 'tag', 'taillight',
+ 'tambourine', 'army_tank', 'tank_(storage_vessel)',
+ 'tank_top_(clothing)', 'tape_(sticky_cloth_or_paper)', 'tape_measure',
+ 'tapestry', 'tarp', 'tartan', 'tassel', 'tea_bag', 'teacup',
+ 'teakettle', 'teapot', 'teddy_bear', 'telephone', 'telephone_booth',
+ 'telephone_pole', 'telephoto_lens', 'television_camera',
+ 'television_set', 'tennis_ball', 'tennis_racket', 'tequila',
+ 'thermometer', 'thermos_bottle', 'thermostat', 'thimble', 'thread',
+ 'thumbtack', 'tiara', 'tiger', 'tights_(clothing)', 'timer', 'tinfoil',
+ 'tinsel', 'tissue_paper', 'toast_(food)', 'toaster', 'toaster_oven',
+ 'toilet', 'toilet_tissue', 'tomato', 'tongs', 'toolbox', 'toothbrush',
+ 'toothpaste', 'toothpick', 'cover', 'tortilla', 'tow_truck', 'towel',
+ 'towel_rack', 'toy', 'tractor_(farm_equipment)', 'traffic_light',
+ 'dirt_bike', 'trailer_truck', 'train_(railroad_vehicle)', 'trampoline',
+ 'tray', 'trench_coat', 'triangle_(musical_instrument)', 'tricycle',
+ 'tripod', 'trousers', 'truck', 'truffle_(chocolate)', 'trunk', 'vat',
+ 'turban', 'turkey_(food)', 'turnip', 'turtle', 'turtleneck_(clothing)',
+ 'typewriter', 'umbrella', 'underwear', 'unicycle', 'urinal', 'urn',
+ 'vacuum_cleaner', 'vase', 'vending_machine', 'vent', 'vest',
+ 'videotape', 'vinegar', 'violin', 'vodka', 'volleyball', 'vulture',
+ 'waffle', 'waffle_iron', 'wagon', 'wagon_wheel', 'walking_stick',
+ 'wall_clock', 'wall_socket', 'wallet', 'walrus', 'wardrobe',
+ 'washbasin', 'automatic_washer', 'watch', 'water_bottle',
+ 'water_cooler', 'water_faucet', 'water_heater', 'water_jug',
+ 'water_gun', 'water_scooter', 'water_ski', 'water_tower',
+ 'watering_can', 'watermelon', 'weathervane', 'webcam', 'wedding_cake',
+ 'wedding_ring', 'wet_suit', 'wheel', 'wheelchair', 'whipped_cream',
+ 'whistle', 'wig', 'wind_chime', 'windmill', 'window_box_(for_plants)',
+ 'windshield_wiper', 'windsock', 'wine_bottle', 'wine_bucket',
+ 'wineglass', 'blinder_(for_horses)', 'wok', 'wolf', 'wooden_spoon',
+ 'wreath', 'wrench', 'wristband', 'wristlet', 'yacht', 'yogurt',
+ 'yoke_(animal_equipment)', 'zebra', 'zucchini'
+ ]
+
+
+dataset_aliases = {
+ 'voc': ['voc', 'pascal_voc', 'voc07', 'voc12'],
+ 'imagenet_det': ['det', 'imagenet_det', 'ilsvrc_det'],
+ 'imagenet_vid': ['vid', 'imagenet_vid', 'ilsvrc_vid'],
+ 'coco': ['coco', 'mscoco', 'ms_coco'],
+ 'coco_panoptic': ['coco_panoptic', 'panoptic'],
+ 'wider_face': ['WIDERFaceDataset', 'wider_face', 'WIDERFace'],
+ 'cityscapes': ['cityscapes'],
+ 'oid_challenge': ['oid_challenge', 'openimages_challenge'],
+ 'oid_v6': ['oid_v6', 'openimages_v6'],
+ 'objects365v1': ['objects365v1', 'obj365v1'],
+ 'objects365v2': ['objects365v2', 'obj365v2'],
+ 'lvis': ['lvis', 'lvis_v1'],
+}
+
+
+def get_classes(dataset) -> list:
+ """Get class names of a dataset."""
+ alias2name = {}
+ for name, aliases in dataset_aliases.items():
+ for alias in aliases:
+ alias2name[alias] = name
+
+ if is_str(dataset):
+ if dataset in alias2name:
+ labels = eval(alias2name[dataset] + '_classes()')
+ else:
+ raise ValueError(f'Unrecognized dataset: {dataset}')
+ else:
+ raise TypeError(f'dataset must a str, but got {type(dataset)}')
+ return labels
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/mean_ap.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/mean_ap.py
new file mode 100644
index 0000000000000000000000000000000000000000..989972a48467f74fa915fa6f3807d0db3becdba2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/mean_ap.py
@@ -0,0 +1,792 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from multiprocessing import Pool
+
+import numpy as np
+from mmengine.logging import print_log
+from mmengine.utils import is_str
+from terminaltables import AsciiTable
+
+from .bbox_overlaps import bbox_overlaps
+from .class_names import get_classes
+
+
+def average_precision(recalls, precisions, mode='area'):
+ """Calculate average precision (for single or multiple scales).
+
+ Args:
+ recalls (ndarray): shape (num_scales, num_dets) or (num_dets, )
+ precisions (ndarray): shape (num_scales, num_dets) or (num_dets, )
+ mode (str): 'area' or '11points', 'area' means calculating the area
+ under precision-recall curve, '11points' means calculating
+ the average precision of recalls at [0, 0.1, ..., 1]
+
+ Returns:
+ float or ndarray: calculated average precision
+ """
+ no_scale = False
+ if recalls.ndim == 1:
+ no_scale = True
+ recalls = recalls[np.newaxis, :]
+ precisions = precisions[np.newaxis, :]
+ assert recalls.shape == precisions.shape and recalls.ndim == 2
+ num_scales = recalls.shape[0]
+ ap = np.zeros(num_scales, dtype=np.float32)
+ if mode == 'area':
+ zeros = np.zeros((num_scales, 1), dtype=recalls.dtype)
+ ones = np.ones((num_scales, 1), dtype=recalls.dtype)
+ mrec = np.hstack((zeros, recalls, ones))
+ mpre = np.hstack((zeros, precisions, zeros))
+ for i in range(mpre.shape[1] - 1, 0, -1):
+ mpre[:, i - 1] = np.maximum(mpre[:, i - 1], mpre[:, i])
+ for i in range(num_scales):
+ ind = np.where(mrec[i, 1:] != mrec[i, :-1])[0]
+ ap[i] = np.sum(
+ (mrec[i, ind + 1] - mrec[i, ind]) * mpre[i, ind + 1])
+ elif mode == '11points':
+ for i in range(num_scales):
+ for thr in np.arange(0, 1 + 1e-3, 0.1):
+ precs = precisions[i, recalls[i, :] >= thr]
+ prec = precs.max() if precs.size > 0 else 0
+ ap[i] += prec
+ ap /= 11
+ else:
+ raise ValueError(
+ 'Unrecognized mode, only "area" and "11points" are supported')
+ if no_scale:
+ ap = ap[0]
+ return ap
+
+
+def tpfp_imagenet(det_bboxes,
+ gt_bboxes,
+ gt_bboxes_ignore=None,
+ default_iou_thr=0.5,
+ area_ranges=None,
+ use_legacy_coordinate=False,
+ **kwargs):
+ """Check if detected bboxes are true positive or false positive.
+
+ Args:
+ det_bbox (ndarray): Detected bboxes of this image, of shape (m, 5).
+ gt_bboxes (ndarray): GT bboxes of this image, of shape (n, 4).
+ gt_bboxes_ignore (ndarray): Ignored gt bboxes of this image,
+ of shape (k, 4). Defaults to None
+ default_iou_thr (float): IoU threshold to be considered as matched for
+ medium and large bboxes (small ones have special rules).
+ Defaults to 0.5.
+ area_ranges (list[tuple] | None): Range of bbox areas to be evaluated,
+ in the format [(min1, max1), (min2, max2), ...]. Defaults to None.
+ use_legacy_coordinate (bool): Whether to use coordinate system in
+ mmdet v1.x. which means width, height should be
+ calculated as 'x2 - x1 + 1` and 'y2 - y1 + 1' respectively.
+ Defaults to False.
+
+ Returns:
+ tuple[np.ndarray]: (tp, fp) whose elements are 0 and 1. The shape of
+ each array is (num_scales, m).
+ """
+
+ if not use_legacy_coordinate:
+ extra_length = 0.
+ else:
+ extra_length = 1.
+
+ # an indicator of ignored gts
+ gt_ignore_inds = np.concatenate(
+ (np.zeros(gt_bboxes.shape[0],
+ dtype=bool), np.ones(gt_bboxes_ignore.shape[0], dtype=bool)))
+ # stack gt_bboxes and gt_bboxes_ignore for convenience
+ gt_bboxes = np.vstack((gt_bboxes, gt_bboxes_ignore))
+
+ num_dets = det_bboxes.shape[0]
+ num_gts = gt_bboxes.shape[0]
+ if area_ranges is None:
+ area_ranges = [(None, None)]
+ num_scales = len(area_ranges)
+ # tp and fp are of shape (num_scales, num_gts), each row is tp or fp
+ # of a certain scale.
+ tp = np.zeros((num_scales, num_dets), dtype=np.float32)
+ fp = np.zeros((num_scales, num_dets), dtype=np.float32)
+ if gt_bboxes.shape[0] == 0:
+ if area_ranges == [(None, None)]:
+ fp[...] = 1
+ else:
+ det_areas = (
+ det_bboxes[:, 2] - det_bboxes[:, 0] + extra_length) * (
+ det_bboxes[:, 3] - det_bboxes[:, 1] + extra_length)
+ for i, (min_area, max_area) in enumerate(area_ranges):
+ fp[i, (det_areas >= min_area) & (det_areas < max_area)] = 1
+ return tp, fp
+ ious = bbox_overlaps(
+ det_bboxes, gt_bboxes - 1, use_legacy_coordinate=use_legacy_coordinate)
+ gt_w = gt_bboxes[:, 2] - gt_bboxes[:, 0] + extra_length
+ gt_h = gt_bboxes[:, 3] - gt_bboxes[:, 1] + extra_length
+ iou_thrs = np.minimum((gt_w * gt_h) / ((gt_w + 10.0) * (gt_h + 10.0)),
+ default_iou_thr)
+ # sort all detections by scores in descending order
+ sort_inds = np.argsort(-det_bboxes[:, -1])
+ for k, (min_area, max_area) in enumerate(area_ranges):
+ gt_covered = np.zeros(num_gts, dtype=bool)
+ # if no area range is specified, gt_area_ignore is all False
+ if min_area is None:
+ gt_area_ignore = np.zeros_like(gt_ignore_inds, dtype=bool)
+ else:
+ gt_areas = gt_w * gt_h
+ gt_area_ignore = (gt_areas < min_area) | (gt_areas >= max_area)
+ for i in sort_inds:
+ max_iou = -1
+ matched_gt = -1
+ # find best overlapped available gt
+ for j in range(num_gts):
+ # different from PASCAL VOC: allow finding other gts if the
+ # best overlapped ones are already matched by other det bboxes
+ if gt_covered[j]:
+ continue
+ elif ious[i, j] >= iou_thrs[j] and ious[i, j] > max_iou:
+ max_iou = ious[i, j]
+ matched_gt = j
+ # there are 4 cases for a det bbox:
+ # 1. it matches a gt, tp = 1, fp = 0
+ # 2. it matches an ignored gt, tp = 0, fp = 0
+ # 3. it matches no gt and within area range, tp = 0, fp = 1
+ # 4. it matches no gt but is beyond area range, tp = 0, fp = 0
+ if matched_gt >= 0:
+ gt_covered[matched_gt] = 1
+ if not (gt_ignore_inds[matched_gt]
+ or gt_area_ignore[matched_gt]):
+ tp[k, i] = 1
+ elif min_area is None:
+ fp[k, i] = 1
+ else:
+ bbox = det_bboxes[i, :4]
+ area = (bbox[2] - bbox[0] + extra_length) * (
+ bbox[3] - bbox[1] + extra_length)
+ if area >= min_area and area < max_area:
+ fp[k, i] = 1
+ return tp, fp
+
+
+def tpfp_default(det_bboxes,
+ gt_bboxes,
+ gt_bboxes_ignore=None,
+ iou_thr=0.5,
+ area_ranges=None,
+ use_legacy_coordinate=False,
+ **kwargs):
+ """Check if detected bboxes are true positive or false positive.
+
+ Args:
+ det_bbox (ndarray): Detected bboxes of this image, of shape (m, 5).
+ gt_bboxes (ndarray): GT bboxes of this image, of shape (n, 4).
+ gt_bboxes_ignore (ndarray): Ignored gt bboxes of this image,
+ of shape (k, 4). Defaults to None
+ iou_thr (float): IoU threshold to be considered as matched.
+ Defaults to 0.5.
+ area_ranges (list[tuple] | None): Range of bbox areas to be
+ evaluated, in the format [(min1, max1), (min2, max2), ...].
+ Defaults to None.
+ use_legacy_coordinate (bool): Whether to use coordinate system in
+ mmdet v1.x. which means width, height should be
+ calculated as 'x2 - x1 + 1` and 'y2 - y1 + 1' respectively.
+ Defaults to False.
+
+ Returns:
+ tuple[np.ndarray]: (tp, fp) whose elements are 0 and 1. The shape of
+ each array is (num_scales, m).
+ """
+
+ if not use_legacy_coordinate:
+ extra_length = 0.
+ else:
+ extra_length = 1.
+
+ # an indicator of ignored gts
+ gt_ignore_inds = np.concatenate(
+ (np.zeros(gt_bboxes.shape[0],
+ dtype=bool), np.ones(gt_bboxes_ignore.shape[0], dtype=bool)))
+ # stack gt_bboxes and gt_bboxes_ignore for convenience
+ gt_bboxes = np.vstack((gt_bboxes, gt_bboxes_ignore))
+
+ num_dets = det_bboxes.shape[0]
+ num_gts = gt_bboxes.shape[0]
+ if area_ranges is None:
+ area_ranges = [(None, None)]
+ num_scales = len(area_ranges)
+ # tp and fp are of shape (num_scales, num_gts), each row is tp or fp of
+ # a certain scale
+ tp = np.zeros((num_scales, num_dets), dtype=np.float32)
+ fp = np.zeros((num_scales, num_dets), dtype=np.float32)
+
+ # if there is no gt bboxes in this image, then all det bboxes
+ # within area range are false positives
+ if gt_bboxes.shape[0] == 0:
+ if area_ranges == [(None, None)]:
+ fp[...] = 1
+ else:
+ det_areas = (
+ det_bboxes[:, 2] - det_bboxes[:, 0] + extra_length) * (
+ det_bboxes[:, 3] - det_bboxes[:, 1] + extra_length)
+ for i, (min_area, max_area) in enumerate(area_ranges):
+ fp[i, (det_areas >= min_area) & (det_areas < max_area)] = 1
+ return tp, fp
+
+ ious = bbox_overlaps(
+ det_bboxes, gt_bboxes, use_legacy_coordinate=use_legacy_coordinate)
+ # for each det, the max iou with all gts
+ ious_max = ious.max(axis=1)
+ # for each det, which gt overlaps most with it
+ ious_argmax = ious.argmax(axis=1)
+ # sort all dets in descending order by scores
+ sort_inds = np.argsort(-det_bboxes[:, -1])
+ for k, (min_area, max_area) in enumerate(area_ranges):
+ gt_covered = np.zeros(num_gts, dtype=bool)
+ # if no area range is specified, gt_area_ignore is all False
+ if min_area is None:
+ gt_area_ignore = np.zeros_like(gt_ignore_inds, dtype=bool)
+ else:
+ gt_areas = (gt_bboxes[:, 2] - gt_bboxes[:, 0] + extra_length) * (
+ gt_bboxes[:, 3] - gt_bboxes[:, 1] + extra_length)
+ gt_area_ignore = (gt_areas < min_area) | (gt_areas >= max_area)
+ for i in sort_inds:
+ if ious_max[i] >= iou_thr:
+ matched_gt = ious_argmax[i]
+ if not (gt_ignore_inds[matched_gt]
+ or gt_area_ignore[matched_gt]):
+ if not gt_covered[matched_gt]:
+ gt_covered[matched_gt] = True
+ tp[k, i] = 1
+ else:
+ fp[k, i] = 1
+ # otherwise ignore this detected bbox, tp = 0, fp = 0
+ elif min_area is None:
+ fp[k, i] = 1
+ else:
+ bbox = det_bboxes[i, :4]
+ area = (bbox[2] - bbox[0] + extra_length) * (
+ bbox[3] - bbox[1] + extra_length)
+ if area >= min_area and area < max_area:
+ fp[k, i] = 1
+ return tp, fp
+
+
+def tpfp_openimages(det_bboxes,
+ gt_bboxes,
+ gt_bboxes_ignore=None,
+ iou_thr=0.5,
+ area_ranges=None,
+ use_legacy_coordinate=False,
+ gt_bboxes_group_of=None,
+ use_group_of=True,
+ ioa_thr=0.5,
+ **kwargs):
+ """Check if detected bboxes are true positive or false positive.
+
+ Args:
+ det_bbox (ndarray): Detected bboxes of this image, of shape (m, 5).
+ gt_bboxes (ndarray): GT bboxes of this image, of shape (n, 4).
+ gt_bboxes_ignore (ndarray): Ignored gt bboxes of this image,
+ of shape (k, 4). Defaults to None
+ iou_thr (float): IoU threshold to be considered as matched.
+ Defaults to 0.5.
+ area_ranges (list[tuple] | None): Range of bbox areas to be
+ evaluated, in the format [(min1, max1), (min2, max2), ...].
+ Defaults to None.
+ use_legacy_coordinate (bool): Whether to use coordinate system in
+ mmdet v1.x. which means width, height should be
+ calculated as 'x2 - x1 + 1` and 'y2 - y1 + 1' respectively.
+ Defaults to False.
+ gt_bboxes_group_of (ndarray): GT group_of of this image, of shape
+ (k, 1). Defaults to None
+ use_group_of (bool): Whether to use group of when calculate TP and FP,
+ which only used in OpenImages evaluation. Defaults to True.
+ ioa_thr (float | None): IoA threshold to be considered as matched,
+ which only used in OpenImages evaluation. Defaults to 0.5.
+
+ Returns:
+ tuple[np.ndarray]: Returns a tuple (tp, fp, det_bboxes), where
+ (tp, fp) whose elements are 0 and 1. The shape of each array is
+ (num_scales, m). (det_bboxes) whose will filter those are not
+ matched by group of gts when processing Open Images evaluation.
+ The shape is (num_scales, m).
+ """
+
+ if not use_legacy_coordinate:
+ extra_length = 0.
+ else:
+ extra_length = 1.
+
+ # an indicator of ignored gts
+ gt_ignore_inds = np.concatenate(
+ (np.zeros(gt_bboxes.shape[0],
+ dtype=bool), np.ones(gt_bboxes_ignore.shape[0], dtype=bool)))
+ # stack gt_bboxes and gt_bboxes_ignore for convenience
+ gt_bboxes = np.vstack((gt_bboxes, gt_bboxes_ignore))
+
+ num_dets = det_bboxes.shape[0]
+ num_gts = gt_bboxes.shape[0]
+ if area_ranges is None:
+ area_ranges = [(None, None)]
+ num_scales = len(area_ranges)
+ # tp and fp are of shape (num_scales, num_gts), each row is tp or fp of
+ # a certain scale
+ tp = np.zeros((num_scales, num_dets), dtype=np.float32)
+ fp = np.zeros((num_scales, num_dets), dtype=np.float32)
+
+ # if there is no gt bboxes in this image, then all det bboxes
+ # within area range are false positives
+ if gt_bboxes.shape[0] == 0:
+ if area_ranges == [(None, None)]:
+ fp[...] = 1
+ else:
+ det_areas = (
+ det_bboxes[:, 2] - det_bboxes[:, 0] + extra_length) * (
+ det_bboxes[:, 3] - det_bboxes[:, 1] + extra_length)
+ for i, (min_area, max_area) in enumerate(area_ranges):
+ fp[i, (det_areas >= min_area) & (det_areas < max_area)] = 1
+ return tp, fp, det_bboxes
+
+ if gt_bboxes_group_of is not None and use_group_of:
+ # if handle group-of boxes, divided gt boxes into two parts:
+ # non-group-of and group-of.Then calculate ious and ioas through
+ # non-group-of group-of gts respectively. This only used in
+ # OpenImages evaluation.
+ assert gt_bboxes_group_of.shape[0] == gt_bboxes.shape[0]
+ non_group_gt_bboxes = gt_bboxes[~gt_bboxes_group_of]
+ group_gt_bboxes = gt_bboxes[gt_bboxes_group_of]
+ num_gts_group = group_gt_bboxes.shape[0]
+ ious = bbox_overlaps(det_bboxes, non_group_gt_bboxes)
+ ioas = bbox_overlaps(det_bboxes, group_gt_bboxes, mode='iof')
+ else:
+ # if not consider group-of boxes, only calculate ious through gt boxes
+ ious = bbox_overlaps(
+ det_bboxes, gt_bboxes, use_legacy_coordinate=use_legacy_coordinate)
+ ioas = None
+
+ if ious.shape[1] > 0:
+ # for each det, the max iou with all gts
+ ious_max = ious.max(axis=1)
+ # for each det, which gt overlaps most with it
+ ious_argmax = ious.argmax(axis=1)
+ # sort all dets in descending order by scores
+ sort_inds = np.argsort(-det_bboxes[:, -1])
+ for k, (min_area, max_area) in enumerate(area_ranges):
+ gt_covered = np.zeros(num_gts, dtype=bool)
+ # if no area range is specified, gt_area_ignore is all False
+ if min_area is None:
+ gt_area_ignore = np.zeros_like(gt_ignore_inds, dtype=bool)
+ else:
+ gt_areas = (
+ gt_bboxes[:, 2] - gt_bboxes[:, 0] + extra_length) * (
+ gt_bboxes[:, 3] - gt_bboxes[:, 1] + extra_length)
+ gt_area_ignore = (gt_areas < min_area) | (gt_areas >= max_area)
+ for i in sort_inds:
+ if ious_max[i] >= iou_thr:
+ matched_gt = ious_argmax[i]
+ if not (gt_ignore_inds[matched_gt]
+ or gt_area_ignore[matched_gt]):
+ if not gt_covered[matched_gt]:
+ gt_covered[matched_gt] = True
+ tp[k, i] = 1
+ else:
+ fp[k, i] = 1
+ # otherwise ignore this detected bbox, tp = 0, fp = 0
+ elif min_area is None:
+ fp[k, i] = 1
+ else:
+ bbox = det_bboxes[i, :4]
+ area = (bbox[2] - bbox[0] + extra_length) * (
+ bbox[3] - bbox[1] + extra_length)
+ if area >= min_area and area < max_area:
+ fp[k, i] = 1
+ else:
+ # if there is no no-group-of gt bboxes in this image,
+ # then all det bboxes within area range are false positives.
+ # Only used in OpenImages evaluation.
+ if area_ranges == [(None, None)]:
+ fp[...] = 1
+ else:
+ det_areas = (
+ det_bboxes[:, 2] - det_bboxes[:, 0] + extra_length) * (
+ det_bboxes[:, 3] - det_bboxes[:, 1] + extra_length)
+ for i, (min_area, max_area) in enumerate(area_ranges):
+ fp[i, (det_areas >= min_area) & (det_areas < max_area)] = 1
+
+ if ioas is None or ioas.shape[1] <= 0:
+ return tp, fp, det_bboxes
+ else:
+ # The evaluation of group-of TP and FP are done in two stages:
+ # 1. All detections are first matched to non group-of boxes; true
+ # positives are determined.
+ # 2. Detections that are determined as false positives are matched
+ # against group-of boxes and calculated group-of TP and FP.
+ # Only used in OpenImages evaluation.
+ det_bboxes_group = np.zeros(
+ (num_scales, ioas.shape[1], det_bboxes.shape[1]), dtype=float)
+ match_group_of = np.zeros((num_scales, num_dets), dtype=bool)
+ tp_group = np.zeros((num_scales, num_gts_group), dtype=np.float32)
+ ioas_max = ioas.max(axis=1)
+ # for each det, which gt overlaps most with it
+ ioas_argmax = ioas.argmax(axis=1)
+ # sort all dets in descending order by scores
+ sort_inds = np.argsort(-det_bboxes[:, -1])
+ for k, (min_area, max_area) in enumerate(area_ranges):
+ box_is_covered = tp[k]
+ # if no area range is specified, gt_area_ignore is all False
+ if min_area is None:
+ gt_area_ignore = np.zeros_like(gt_ignore_inds, dtype=bool)
+ else:
+ gt_areas = (gt_bboxes[:, 2] - gt_bboxes[:, 0]) * (
+ gt_bboxes[:, 3] - gt_bboxes[:, 1])
+ gt_area_ignore = (gt_areas < min_area) | (gt_areas >= max_area)
+ for i in sort_inds:
+ matched_gt = ioas_argmax[i]
+ if not box_is_covered[i]:
+ if ioas_max[i] >= ioa_thr:
+ if not (gt_ignore_inds[matched_gt]
+ or gt_area_ignore[matched_gt]):
+ if not tp_group[k, matched_gt]:
+ tp_group[k, matched_gt] = 1
+ match_group_of[k, i] = True
+ else:
+ match_group_of[k, i] = True
+
+ if det_bboxes_group[k, matched_gt, -1] < \
+ det_bboxes[i, -1]:
+ det_bboxes_group[k, matched_gt] = \
+ det_bboxes[i]
+
+ fp_group = (tp_group <= 0).astype(float)
+ tps = []
+ fps = []
+ # concatenate tp, fp, and det-boxes which not matched group of
+ # gt boxes and tp_group, fp_group, and det_bboxes_group which
+ # matched group of boxes respectively.
+ for i in range(num_scales):
+ tps.append(
+ np.concatenate((tp[i][~match_group_of[i]], tp_group[i])))
+ fps.append(
+ np.concatenate((fp[i][~match_group_of[i]], fp_group[i])))
+ det_bboxes = np.concatenate(
+ (det_bboxes[~match_group_of[i]], det_bboxes_group[i]))
+
+ tp = np.vstack(tps)
+ fp = np.vstack(fps)
+ return tp, fp, det_bboxes
+
+
+def get_cls_results(det_results, annotations, class_id):
+ """Get det results and gt information of a certain class.
+
+ Args:
+ det_results (list[list]): Same as `eval_map()`.
+ annotations (list[dict]): Same as `eval_map()`.
+ class_id (int): ID of a specific class.
+
+ Returns:
+ tuple[list[np.ndarray]]: detected bboxes, gt bboxes, ignored gt bboxes
+ """
+ cls_dets = [img_res[class_id] for img_res in det_results]
+ cls_gts = []
+ cls_gts_ignore = []
+ for ann in annotations:
+ gt_inds = ann['labels'] == class_id
+ cls_gts.append(ann['bboxes'][gt_inds, :])
+
+ if ann.get('labels_ignore', None) is not None:
+ ignore_inds = ann['labels_ignore'] == class_id
+ cls_gts_ignore.append(ann['bboxes_ignore'][ignore_inds, :])
+ else:
+ cls_gts_ignore.append(np.empty((0, 4), dtype=np.float32))
+
+ return cls_dets, cls_gts, cls_gts_ignore
+
+
+def get_cls_group_ofs(annotations, class_id):
+ """Get `gt_group_of` of a certain class, which is used in Open Images.
+
+ Args:
+ annotations (list[dict]): Same as `eval_map()`.
+ class_id (int): ID of a specific class.
+
+ Returns:
+ list[np.ndarray]: `gt_group_of` of a certain class.
+ """
+ gt_group_ofs = []
+ for ann in annotations:
+ gt_inds = ann['labels'] == class_id
+ if ann.get('gt_is_group_ofs', None) is not None:
+ gt_group_ofs.append(ann['gt_is_group_ofs'][gt_inds])
+ else:
+ gt_group_ofs.append(np.empty((0, 1), dtype=bool))
+
+ return gt_group_ofs
+
+
+def eval_map(det_results,
+ annotations,
+ scale_ranges=None,
+ iou_thr=0.5,
+ ioa_thr=None,
+ dataset=None,
+ logger=None,
+ tpfp_fn=None,
+ nproc=4,
+ use_legacy_coordinate=False,
+ use_group_of=False,
+ eval_mode='area'):
+ """Evaluate mAP of a dataset.
+
+ Args:
+ det_results (list[list]): [[cls1_det, cls2_det, ...], ...].
+ The outer list indicates images, and the inner list indicates
+ per-class detected bboxes.
+ annotations (list[dict]): Ground truth annotations where each item of
+ the list indicates an image. Keys of annotations are:
+
+ - `bboxes`: numpy array of shape (n, 4)
+ - `labels`: numpy array of shape (n, )
+ - `bboxes_ignore` (optional): numpy array of shape (k, 4)
+ - `labels_ignore` (optional): numpy array of shape (k, )
+ scale_ranges (list[tuple] | None): Range of scales to be evaluated,
+ in the format [(min1, max1), (min2, max2), ...]. A range of
+ (32, 64) means the area range between (32**2, 64**2).
+ Defaults to None.
+ iou_thr (float): IoU threshold to be considered as matched.
+ Defaults to 0.5.
+ ioa_thr (float | None): IoA threshold to be considered as matched,
+ which only used in OpenImages evaluation. Defaults to None.
+ dataset (list[str] | str | None): Dataset name or dataset classes,
+ there are minor differences in metrics for different datasets, e.g.
+ "voc", "imagenet_det", etc. Defaults to None.
+ logger (logging.Logger | str | None): The way to print the mAP
+ summary. See `mmengine.logging.print_log()` for details.
+ Defaults to None.
+ tpfp_fn (callable | None): The function used to determine true/
+ false positives. If None, :func:`tpfp_default` is used as default
+ unless dataset is 'det' or 'vid' (:func:`tpfp_imagenet` in this
+ case). If it is given as a function, then this function is used
+ to evaluate tp & fp. Default None.
+ nproc (int): Processes used for computing TP and FP.
+ Defaults to 4.
+ use_legacy_coordinate (bool): Whether to use coordinate system in
+ mmdet v1.x. which means width, height should be
+ calculated as 'x2 - x1 + 1` and 'y2 - y1 + 1' respectively.
+ Defaults to False.
+ use_group_of (bool): Whether to use group of when calculate TP and FP,
+ which only used in OpenImages evaluation. Defaults to False.
+ eval_mode (str): 'area' or '11points', 'area' means calculating the
+ area under precision-recall curve, '11points' means calculating
+ the average precision of recalls at [0, 0.1, ..., 1],
+ PASCAL VOC2007 uses `11points` as default evaluate mode, while
+ others are 'area'. Defaults to 'area'.
+
+ Returns:
+ tuple: (mAP, [dict, dict, ...])
+ """
+ assert len(det_results) == len(annotations)
+ assert eval_mode in ['area', '11points'], \
+ f'Unrecognized {eval_mode} mode, only "area" and "11points" ' \
+ 'are supported'
+ if not use_legacy_coordinate:
+ extra_length = 0.
+ else:
+ extra_length = 1.
+
+ num_imgs = len(det_results)
+ num_scales = len(scale_ranges) if scale_ranges is not None else 1
+ num_classes = len(det_results[0]) # positive class num
+ area_ranges = ([(rg[0]**2, rg[1]**2) for rg in scale_ranges]
+ if scale_ranges is not None else None)
+
+ # There is no need to use multi processes to process
+ # when num_imgs = 1 .
+ if num_imgs > 1:
+ assert nproc > 0, 'nproc must be at least one.'
+ nproc = min(nproc, num_imgs)
+ pool = Pool(nproc)
+
+ eval_results = []
+ for i in range(num_classes):
+ # get gt and det bboxes of this class
+ cls_dets, cls_gts, cls_gts_ignore = get_cls_results(
+ det_results, annotations, i)
+ # choose proper function according to datasets to compute tp and fp
+ if tpfp_fn is None:
+ if dataset in ['det', 'vid']:
+ tpfp_fn = tpfp_imagenet
+ elif dataset in ['oid_challenge', 'oid_v6'] \
+ or use_group_of is True:
+ tpfp_fn = tpfp_openimages
+ else:
+ tpfp_fn = tpfp_default
+ if not callable(tpfp_fn):
+ raise ValueError(
+ f'tpfp_fn has to be a function or None, but got {tpfp_fn}')
+
+ if num_imgs > 1:
+ # compute tp and fp for each image with multiple processes
+ args = []
+ if use_group_of:
+ # used in Open Images Dataset evaluation
+ gt_group_ofs = get_cls_group_ofs(annotations, i)
+ args.append(gt_group_ofs)
+ args.append([use_group_of for _ in range(num_imgs)])
+ if ioa_thr is not None:
+ args.append([ioa_thr for _ in range(num_imgs)])
+
+ tpfp = pool.starmap(
+ tpfp_fn,
+ zip(cls_dets, cls_gts, cls_gts_ignore,
+ [iou_thr for _ in range(num_imgs)],
+ [area_ranges for _ in range(num_imgs)],
+ [use_legacy_coordinate for _ in range(num_imgs)], *args))
+ else:
+ tpfp = tpfp_fn(
+ cls_dets[0],
+ cls_gts[0],
+ cls_gts_ignore[0],
+ iou_thr,
+ area_ranges,
+ use_legacy_coordinate,
+ gt_bboxes_group_of=(get_cls_group_ofs(annotations, i)[0]
+ if use_group_of else None),
+ use_group_of=use_group_of,
+ ioa_thr=ioa_thr)
+ tpfp = [tpfp]
+
+ if use_group_of:
+ tp, fp, cls_dets = tuple(zip(*tpfp))
+ else:
+ tp, fp = tuple(zip(*tpfp))
+ # calculate gt number of each scale
+ # ignored gts or gts beyond the specific scale are not counted
+ num_gts = np.zeros(num_scales, dtype=int)
+ for j, bbox in enumerate(cls_gts):
+ if area_ranges is None:
+ num_gts[0] += bbox.shape[0]
+ else:
+ gt_areas = (bbox[:, 2] - bbox[:, 0] + extra_length) * (
+ bbox[:, 3] - bbox[:, 1] + extra_length)
+ for k, (min_area, max_area) in enumerate(area_ranges):
+ num_gts[k] += np.sum((gt_areas >= min_area)
+ & (gt_areas < max_area))
+ # sort all det bboxes by score, also sort tp and fp
+ cls_dets = np.vstack(cls_dets)
+ num_dets = cls_dets.shape[0]
+ sort_inds = np.argsort(-cls_dets[:, -1])
+ tp = np.hstack(tp)[:, sort_inds]
+ fp = np.hstack(fp)[:, sort_inds]
+ # calculate recall and precision with tp and fp
+ tp = np.cumsum(tp, axis=1)
+ fp = np.cumsum(fp, axis=1)
+ eps = np.finfo(np.float32).eps
+ recalls = tp / np.maximum(num_gts[:, np.newaxis], eps)
+ precisions = tp / np.maximum((tp + fp), eps)
+ # calculate AP
+ if scale_ranges is None:
+ recalls = recalls[0, :]
+ precisions = precisions[0, :]
+ num_gts = num_gts.item()
+ ap = average_precision(recalls, precisions, eval_mode)
+ eval_results.append({
+ 'num_gts': num_gts,
+ 'num_dets': num_dets,
+ 'recall': recalls,
+ 'precision': precisions,
+ 'ap': ap
+ })
+
+ if num_imgs > 1:
+ pool.close()
+
+ if scale_ranges is not None:
+ # shape (num_classes, num_scales)
+ all_ap = np.vstack([cls_result['ap'] for cls_result in eval_results])
+ all_num_gts = np.vstack(
+ [cls_result['num_gts'] for cls_result in eval_results])
+ mean_ap = []
+ for i in range(num_scales):
+ if np.any(all_num_gts[:, i] > 0):
+ mean_ap.append(all_ap[all_num_gts[:, i] > 0, i].mean())
+ else:
+ mean_ap.append(0.0)
+ else:
+ aps = []
+ for cls_result in eval_results:
+ if cls_result['num_gts'] > 0:
+ aps.append(cls_result['ap'])
+ mean_ap = np.array(aps).mean().item() if aps else 0.0
+
+ print_map_summary(
+ mean_ap, eval_results, dataset, area_ranges, logger=logger)
+
+ return mean_ap, eval_results
+
+
+def print_map_summary(mean_ap,
+ results,
+ dataset=None,
+ scale_ranges=None,
+ logger=None):
+ """Print mAP and results of each class.
+
+ A table will be printed to show the gts/dets/recall/AP of each class and
+ the mAP.
+
+ Args:
+ mean_ap (float): Calculated from `eval_map()`.
+ results (list[dict]): Calculated from `eval_map()`.
+ dataset (list[str] | str | None): Dataset name or dataset classes.
+ scale_ranges (list[tuple] | None): Range of scales to be evaluated.
+ logger (logging.Logger | str | None): The way to print the mAP
+ summary. See `mmengine.logging.print_log()` for details.
+ Defaults to None.
+ """
+
+ if logger == 'silent':
+ return
+
+ if isinstance(results[0]['ap'], np.ndarray):
+ num_scales = len(results[0]['ap'])
+ else:
+ num_scales = 1
+
+ if scale_ranges is not None:
+ assert len(scale_ranges) == num_scales
+
+ num_classes = len(results)
+
+ recalls = np.zeros((num_scales, num_classes), dtype=np.float32)
+ aps = np.zeros((num_scales, num_classes), dtype=np.float32)
+ num_gts = np.zeros((num_scales, num_classes), dtype=int)
+ for i, cls_result in enumerate(results):
+ if cls_result['recall'].size > 0:
+ recalls[:, i] = np.array(cls_result['recall'], ndmin=2)[:, -1]
+ aps[:, i] = cls_result['ap']
+ num_gts[:, i] = cls_result['num_gts']
+
+ if dataset is None:
+ label_names = [str(i) for i in range(num_classes)]
+ elif is_str(dataset):
+ label_names = get_classes(dataset)
+ else:
+ label_names = dataset
+
+ if not isinstance(mean_ap, list):
+ mean_ap = [mean_ap]
+
+ header = ['class', 'gts', 'dets', 'recall', 'ap']
+ for i in range(num_scales):
+ if scale_ranges is not None:
+ print_log(f'Scale range {scale_ranges[i]}', logger=logger)
+ table_data = [header]
+ for j in range(num_classes):
+ row_data = [
+ label_names[j], num_gts[i, j], results[j]['num_dets'],
+ f'{recalls[i, j]:.3f}', f'{aps[i, j]:.3f}'
+ ]
+ table_data.append(row_data)
+ table_data.append(['mAP', '', '', '', f'{mean_ap[i]:.3f}'])
+ table = AsciiTable(table_data)
+ table.inner_footing_row_border = True
+ print_log('\n' + table.table, logger=logger)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/panoptic_utils.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/panoptic_utils.py
new file mode 100644
index 0000000000000000000000000000000000000000..6faa8ed52bc46c2cb74b1974b8daa521e616e996
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/panoptic_utils.py
@@ -0,0 +1,228 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+# Copyright (c) 2018, Alexander Kirillov
+# This file supports `backend_args` for `panopticapi`,
+# the source code is copied from `panopticapi`,
+# only the way to load the gt images is modified.
+import multiprocessing
+import os
+
+import mmcv
+import numpy as np
+from mmengine.fileio import get
+
+# A custom value to distinguish instance ID and category ID; need to
+# be greater than the number of categories.
+# For a pixel in the panoptic result map:
+# pan_id = ins_id * INSTANCE_OFFSET + cat_id
+INSTANCE_OFFSET = 1000
+
+try:
+ from panopticapi.evaluation import OFFSET, VOID, PQStat
+ from panopticapi.utils import rgb2id
+except ImportError:
+ PQStat = None
+ rgb2id = None
+ VOID = 0
+ OFFSET = 256 * 256 * 256
+
+
+def pq_compute_single_core(proc_id,
+ annotation_set,
+ gt_folder,
+ pred_folder,
+ categories,
+ backend_args=None,
+ print_log=False):
+ """The single core function to evaluate the metric of Panoptic
+ Segmentation.
+
+ Same as the function with the same name in `panopticapi`. Only the function
+ to load the images is changed to use the file client.
+
+ Args:
+ proc_id (int): The id of the mini process.
+ gt_folder (str): The path of the ground truth images.
+ pred_folder (str): The path of the prediction images.
+ categories (str): The categories of the dataset.
+ backend_args (object): The Backend of the dataset. If None,
+ the backend will be set to `local`.
+ print_log (bool): Whether to print the log. Defaults to False.
+ """
+ if PQStat is None:
+ raise RuntimeError(
+ 'panopticapi is not installed, please install it by: '
+ 'pip install git+https://github.com/cocodataset/'
+ 'panopticapi.git.')
+
+ pq_stat = PQStat()
+
+ idx = 0
+ for gt_ann, pred_ann in annotation_set:
+ if print_log and idx % 100 == 0:
+ print('Core: {}, {} from {} images processed'.format(
+ proc_id, idx, len(annotation_set)))
+ idx += 1
+ # The gt images can be on the local disk or `ceph`, so we use
+ # backend here.
+ img_bytes = get(
+ os.path.join(gt_folder, gt_ann['file_name']),
+ backend_args=backend_args)
+ pan_gt = mmcv.imfrombytes(img_bytes, flag='color', channel_order='rgb')
+ pan_gt = rgb2id(pan_gt)
+
+ # The predictions can only be on the local dist now.
+ pan_pred = mmcv.imread(
+ os.path.join(pred_folder, pred_ann['file_name']),
+ flag='color',
+ channel_order='rgb')
+ pan_pred = rgb2id(pan_pred)
+
+ gt_segms = {el['id']: el for el in gt_ann['segments_info']}
+ pred_segms = {el['id']: el for el in pred_ann['segments_info']}
+
+ # predicted segments area calculation + prediction sanity checks
+ pred_labels_set = set(el['id'] for el in pred_ann['segments_info'])
+ labels, labels_cnt = np.unique(pan_pred, return_counts=True)
+ for label, label_cnt in zip(labels, labels_cnt):
+ if label not in pred_segms:
+ if label == VOID:
+ continue
+ raise KeyError(
+ 'In the image with ID {} segment with ID {} is '
+ 'presented in PNG and not presented in JSON.'.format(
+ gt_ann['image_id'], label))
+ pred_segms[label]['area'] = label_cnt
+ pred_labels_set.remove(label)
+ if pred_segms[label]['category_id'] not in categories:
+ raise KeyError(
+ 'In the image with ID {} segment with ID {} has '
+ 'unknown category_id {}.'.format(
+ gt_ann['image_id'], label,
+ pred_segms[label]['category_id']))
+ if len(pred_labels_set) != 0:
+ raise KeyError(
+ 'In the image with ID {} the following segment IDs {} '
+ 'are presented in JSON and not presented in PNG.'.format(
+ gt_ann['image_id'], list(pred_labels_set)))
+
+ # confusion matrix calculation
+ pan_gt_pred = pan_gt.astype(np.uint64) * OFFSET + pan_pred.astype(
+ np.uint64)
+ gt_pred_map = {}
+ labels, labels_cnt = np.unique(pan_gt_pred, return_counts=True)
+ for label, intersection in zip(labels, labels_cnt):
+ gt_id = label // OFFSET
+ pred_id = label % OFFSET
+ gt_pred_map[(gt_id, pred_id)] = intersection
+
+ # count all matched pairs
+ gt_matched = set()
+ pred_matched = set()
+ for label_tuple, intersection in gt_pred_map.items():
+ gt_label, pred_label = label_tuple
+ if gt_label not in gt_segms:
+ continue
+ if pred_label not in pred_segms:
+ continue
+ if gt_segms[gt_label]['iscrowd'] == 1:
+ continue
+ if gt_segms[gt_label]['category_id'] != pred_segms[pred_label][
+ 'category_id']:
+ continue
+
+ union = pred_segms[pred_label]['area'] + gt_segms[gt_label][
+ 'area'] - intersection - gt_pred_map.get((VOID, pred_label), 0)
+ iou = intersection / union
+ if iou > 0.5:
+ pq_stat[gt_segms[gt_label]['category_id']].tp += 1
+ pq_stat[gt_segms[gt_label]['category_id']].iou += iou
+ gt_matched.add(gt_label)
+ pred_matched.add(pred_label)
+
+ # count false positives
+ crowd_labels_dict = {}
+ for gt_label, gt_info in gt_segms.items():
+ if gt_label in gt_matched:
+ continue
+ # crowd segments are ignored
+ if gt_info['iscrowd'] == 1:
+ crowd_labels_dict[gt_info['category_id']] = gt_label
+ continue
+ pq_stat[gt_info['category_id']].fn += 1
+
+ # count false positives
+ for pred_label, pred_info in pred_segms.items():
+ if pred_label in pred_matched:
+ continue
+ # intersection of the segment with VOID
+ intersection = gt_pred_map.get((VOID, pred_label), 0)
+ # plus intersection with corresponding CROWD region if it exists
+ if pred_info['category_id'] in crowd_labels_dict:
+ intersection += gt_pred_map.get(
+ (crowd_labels_dict[pred_info['category_id']], pred_label),
+ 0)
+ # predicted segment is ignored if more than half of
+ # the segment correspond to VOID and CROWD regions
+ if intersection / pred_info['area'] > 0.5:
+ continue
+ pq_stat[pred_info['category_id']].fp += 1
+
+ if print_log:
+ print('Core: {}, all {} images processed'.format(
+ proc_id, len(annotation_set)))
+ return pq_stat
+
+
+def pq_compute_multi_core(matched_annotations_list,
+ gt_folder,
+ pred_folder,
+ categories,
+ backend_args=None,
+ nproc=32):
+ """Evaluate the metrics of Panoptic Segmentation with multithreading.
+
+ Same as the function with the same name in `panopticapi`.
+
+ Args:
+ matched_annotations_list (list): The matched annotation list. Each
+ element is a tuple of annotations of the same image with the
+ format (gt_anns, pred_anns).
+ gt_folder (str): The path of the ground truth images.
+ pred_folder (str): The path of the prediction images.
+ categories (str): The categories of the dataset.
+ backend_args (object): The file client of the dataset. If None,
+ the backend will be set to `local`.
+ nproc (int): Number of processes for panoptic quality computing.
+ Defaults to 32. When `nproc` exceeds the number of cpu cores,
+ the number of cpu cores is used.
+ """
+ if PQStat is None:
+ raise RuntimeError(
+ 'panopticapi is not installed, please install it by: '
+ 'pip install git+https://github.com/cocodataset/'
+ 'panopticapi.git.')
+
+ cpu_num = min(nproc, multiprocessing.cpu_count())
+
+ annotations_split = np.array_split(matched_annotations_list, cpu_num)
+ print('Number of cores: {}, images per core: {}'.format(
+ cpu_num, len(annotations_split[0])))
+ workers = multiprocessing.Pool(processes=cpu_num)
+ processes = []
+ for proc_id, annotation_set in enumerate(annotations_split):
+ p = workers.apply_async(pq_compute_single_core,
+ (proc_id, annotation_set, gt_folder,
+ pred_folder, categories, backend_args))
+ processes.append(p)
+
+ # Close the process pool, otherwise it will lead to memory
+ # leaking problems.
+ workers.close()
+ workers.join()
+
+ pq_stat = PQStat()
+ for p in processes:
+ pq_stat += p.get()
+
+ return pq_stat
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/recall.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/recall.py
new file mode 100644
index 0000000000000000000000000000000000000000..4bce2bf3614ab454dbbdf48efc4650018cc71b13
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/recall.py
@@ -0,0 +1,199 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from collections.abc import Sequence
+
+import numpy as np
+from mmengine.logging import print_log
+from terminaltables import AsciiTable
+
+from .bbox_overlaps import bbox_overlaps
+
+
+def _recalls(all_ious, proposal_nums, thrs):
+
+ img_num = all_ious.shape[0]
+ total_gt_num = sum([ious.shape[0] for ious in all_ious])
+
+ _ious = np.zeros((proposal_nums.size, total_gt_num), dtype=np.float32)
+ for k, proposal_num in enumerate(proposal_nums):
+ tmp_ious = np.zeros(0)
+ for i in range(img_num):
+ ious = all_ious[i][:, :proposal_num].copy()
+ gt_ious = np.zeros((ious.shape[0]))
+ if ious.size == 0:
+ tmp_ious = np.hstack((tmp_ious, gt_ious))
+ continue
+ for j in range(ious.shape[0]):
+ gt_max_overlaps = ious.argmax(axis=1)
+ max_ious = ious[np.arange(0, ious.shape[0]), gt_max_overlaps]
+ gt_idx = max_ious.argmax()
+ gt_ious[j] = max_ious[gt_idx]
+ box_idx = gt_max_overlaps[gt_idx]
+ ious[gt_idx, :] = -1
+ ious[:, box_idx] = -1
+ tmp_ious = np.hstack((tmp_ious, gt_ious))
+ _ious[k, :] = tmp_ious
+
+ _ious = np.fliplr(np.sort(_ious, axis=1))
+ recalls = np.zeros((proposal_nums.size, thrs.size))
+ for i, thr in enumerate(thrs):
+ recalls[:, i] = (_ious >= thr).sum(axis=1) / float(total_gt_num)
+
+ return recalls
+
+
+def set_recall_param(proposal_nums, iou_thrs):
+ """Check proposal_nums and iou_thrs and set correct format."""
+ if isinstance(proposal_nums, Sequence):
+ _proposal_nums = np.array(proposal_nums)
+ elif isinstance(proposal_nums, int):
+ _proposal_nums = np.array([proposal_nums])
+ else:
+ _proposal_nums = proposal_nums
+
+ if iou_thrs is None:
+ _iou_thrs = np.array([0.5])
+ elif isinstance(iou_thrs, Sequence):
+ _iou_thrs = np.array(iou_thrs)
+ elif isinstance(iou_thrs, float):
+ _iou_thrs = np.array([iou_thrs])
+ else:
+ _iou_thrs = iou_thrs
+
+ return _proposal_nums, _iou_thrs
+
+
+def eval_recalls(gts,
+ proposals,
+ proposal_nums=None,
+ iou_thrs=0.5,
+ logger=None,
+ use_legacy_coordinate=False):
+ """Calculate recalls.
+
+ Args:
+ gts (list[ndarray]): a list of arrays of shape (n, 4)
+ proposals (list[ndarray]): a list of arrays of shape (k, 4) or (k, 5)
+ proposal_nums (int | Sequence[int]): Top N proposals to be evaluated.
+ iou_thrs (float | Sequence[float]): IoU thresholds. Default: 0.5.
+ logger (logging.Logger | str | None): The way to print the recall
+ summary. See `mmengine.logging.print_log()` for details.
+ Default: None.
+ use_legacy_coordinate (bool): Whether use coordinate system
+ in mmdet v1.x. "1" was added to both height and width
+ which means w, h should be
+ computed as 'x2 - x1 + 1` and 'y2 - y1 + 1'. Default: False.
+
+
+ Returns:
+ ndarray: recalls of different ious and proposal nums
+ """
+
+ img_num = len(gts)
+ assert img_num == len(proposals)
+ proposal_nums, iou_thrs = set_recall_param(proposal_nums, iou_thrs)
+ all_ious = []
+ for i in range(img_num):
+ if proposals[i].ndim == 2 and proposals[i].shape[1] == 5:
+ scores = proposals[i][:, 4]
+ sort_idx = np.argsort(scores)[::-1]
+ img_proposal = proposals[i][sort_idx, :]
+ else:
+ img_proposal = proposals[i]
+ prop_num = min(img_proposal.shape[0], proposal_nums[-1])
+ if gts[i] is None or gts[i].shape[0] == 0:
+ ious = np.zeros((0, img_proposal.shape[0]), dtype=np.float32)
+ else:
+ ious = bbox_overlaps(
+ gts[i],
+ img_proposal[:prop_num, :4],
+ use_legacy_coordinate=use_legacy_coordinate)
+ all_ious.append(ious)
+ all_ious = np.array(all_ious)
+ recalls = _recalls(all_ious, proposal_nums, iou_thrs)
+
+ print_recall_summary(recalls, proposal_nums, iou_thrs, logger=logger)
+ return recalls
+
+
+def print_recall_summary(recalls,
+ proposal_nums,
+ iou_thrs,
+ row_idxs=None,
+ col_idxs=None,
+ logger=None):
+ """Print recalls in a table.
+
+ Args:
+ recalls (ndarray): calculated from `bbox_recalls`
+ proposal_nums (ndarray or list): top N proposals
+ iou_thrs (ndarray or list): iou thresholds
+ row_idxs (ndarray): which rows(proposal nums) to print
+ col_idxs (ndarray): which cols(iou thresholds) to print
+ logger (logging.Logger | str | None): The way to print the recall
+ summary. See `mmengine.logging.print_log()` for details.
+ Default: None.
+ """
+ proposal_nums = np.array(proposal_nums, dtype=np.int32)
+ iou_thrs = np.array(iou_thrs)
+ if row_idxs is None:
+ row_idxs = np.arange(proposal_nums.size)
+ if col_idxs is None:
+ col_idxs = np.arange(iou_thrs.size)
+ row_header = [''] + iou_thrs[col_idxs].tolist()
+ table_data = [row_header]
+ for i, num in enumerate(proposal_nums[row_idxs]):
+ row = [f'{val:.3f}' for val in recalls[row_idxs[i], col_idxs].tolist()]
+ row.insert(0, num)
+ table_data.append(row)
+ table = AsciiTable(table_data)
+ print_log('\n' + table.table, logger=logger)
+
+
+def plot_num_recall(recalls, proposal_nums):
+ """Plot Proposal_num-Recalls curve.
+
+ Args:
+ recalls(ndarray or list): shape (k,)
+ proposal_nums(ndarray or list): same shape as `recalls`
+ """
+ if isinstance(proposal_nums, np.ndarray):
+ _proposal_nums = proposal_nums.tolist()
+ else:
+ _proposal_nums = proposal_nums
+ if isinstance(recalls, np.ndarray):
+ _recalls = recalls.tolist()
+ else:
+ _recalls = recalls
+
+ import matplotlib.pyplot as plt
+ f = plt.figure()
+ plt.plot([0] + _proposal_nums, [0] + _recalls)
+ plt.xlabel('Proposal num')
+ plt.ylabel('Recall')
+ plt.axis([0, proposal_nums.max(), 0, 1])
+ f.show()
+
+
+def plot_iou_recall(recalls, iou_thrs):
+ """Plot IoU-Recalls curve.
+
+ Args:
+ recalls(ndarray or list): shape (k,)
+ iou_thrs(ndarray or list): same shape as `recalls`
+ """
+ if isinstance(iou_thrs, np.ndarray):
+ _iou_thrs = iou_thrs.tolist()
+ else:
+ _iou_thrs = iou_thrs
+ if isinstance(recalls, np.ndarray):
+ _recalls = recalls.tolist()
+ else:
+ _recalls = recalls
+
+ import matplotlib.pyplot as plt
+ f = plt.figure()
+ plt.plot(_iou_thrs + [1.0], _recalls + [0.])
+ plt.xlabel('IoU')
+ plt.ylabel('Recall')
+ plt.axis([iou_thrs.min(), 1, 0, 1])
+ f.show()
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/ytvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/ytvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..c65a7e9bc956c7de42e0d6e511dabb3d7325782d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/ytvis.py
@@ -0,0 +1,305 @@
+# Copyright (c) Github URL
+# Copied from
+# https://github.com/youtubevos/cocoapi/blob/master/PythonAPI/pycocotools/ytvos.py
+__author__ = 'ychfan'
+# Interface for accessing the YouTubeVIS dataset.
+
+# The following API functions are defined:
+# YTVIS - YTVIS api class that loads YouTubeVIS annotation file
+# and prepare data structures.
+# decodeMask - Decode binary mask M encoded via run-length encoding.
+# encodeMask - Encode binary mask M using run-length encoding.
+# getAnnIds - Get ann ids that satisfy given filter conditions.
+# getCatIds - Get cat ids that satisfy given filter conditions.
+# getImgIds - Get img ids that satisfy given filter conditions.
+# loadAnns - Load anns with the specified ids.
+# loadCats - Load cats with the specified ids.
+# loadImgs - Load imgs with the specified ids.
+# annToMask - Convert segmentation in an annotation to binary mask.
+# loadRes - Load algorithm results and create API for accessing them.
+
+# Microsoft COCO Toolbox. version 2.0
+# Data, paper, and tutorials available at: http://mscoco.org/
+# Code written by Piotr Dollar and Tsung-Yi Lin, 2014.
+# Licensed under the Simplified BSD License [see bsd.txt]
+
+import copy
+import itertools
+import json
+import sys
+import time
+from collections import defaultdict
+
+import numpy as np
+from pycocotools import mask as maskUtils
+
+PYTHON_VERSION = sys.version_info[0]
+
+
+def _isArrayLike(obj):
+ return hasattr(obj, '__iter__') and hasattr(obj, '__len__')
+
+
+class YTVIS:
+
+ def __init__(self, annotation_file=None):
+ """Constructor of Microsoft COCO helper class for reading and
+ visualizing annotations.
+
+ :param annotation_file (str | dict): location of annotation file or
+ dict results.
+ :param image_folder (str): location to the folder that hosts images.
+ :return:
+ """
+ # load dataset
+ self.dataset, self.anns, self.cats, self.vids = dict(), dict(), dict(
+ ), dict()
+ self.vidToAnns, self.catToVids = defaultdict(list), defaultdict(list)
+ if annotation_file is not None:
+ print('loading annotations into memory...')
+ tic = time.time()
+ if type(annotation_file) == str:
+ dataset = json.load(open(annotation_file, 'r'))
+ else:
+ dataset = annotation_file
+ assert type(
+ dataset
+ ) == dict, 'annotation file format {} not supported'.format(
+ type(dataset))
+ print('Done (t={:0.2f}s)'.format(time.time() - tic))
+ self.dataset = dataset
+ self.createIndex()
+
+ def createIndex(self):
+ # create index
+ print('creating index...')
+ anns, cats, vids = {}, {}, {}
+ vidToAnns, catToVids = defaultdict(list), defaultdict(list)
+ if 'annotations' in self.dataset:
+ for ann in self.dataset['annotations']:
+ vidToAnns[ann['video_id']].append(ann)
+ anns[ann['id']] = ann
+
+ if 'videos' in self.dataset:
+ for vid in self.dataset['videos']:
+ vids[vid['id']] = vid
+
+ if 'categories' in self.dataset:
+ for cat in self.dataset['categories']:
+ cats[cat['id']] = cat
+
+ if 'annotations' in self.dataset and 'categories' in self.dataset:
+ for ann in self.dataset['annotations']:
+ catToVids[ann['category_id']].append(ann['video_id'])
+
+ print('index created!')
+
+ # create class members
+ self.anns = anns
+ self.vidToAnns = vidToAnns
+ self.catToVids = catToVids
+ self.vids = vids
+ self.cats = cats
+
+ def getAnnIds(self, vidIds=[], catIds=[], areaRng=[], iscrowd=None):
+ """Get ann ids that satisfy given filter conditions. default skips that
+ filter.
+
+ :param vidIds (int array) : get anns for given vids
+ catIds (int array) : get anns for given cats
+ areaRng (float array) : get anns for given area range
+ iscrowd (boolean) : get anns for given crowd label
+ :return: ids (int array) : integer array of ann ids
+ """
+ vidIds = vidIds if _isArrayLike(vidIds) else [vidIds]
+ catIds = catIds if _isArrayLike(catIds) else [catIds]
+
+ if len(vidIds) == len(catIds) == len(areaRng) == 0:
+ anns = self.dataset['annotations']
+ else:
+ if not len(vidIds) == 0:
+ lists = [
+ self.vidToAnns[vidId] for vidId in vidIds
+ if vidId in self.vidToAnns
+ ]
+ anns = list(itertools.chain.from_iterable(lists))
+ else:
+ anns = self.dataset['annotations']
+ anns = anns if len(catIds) == 0 else [
+ ann for ann in anns if ann['category_id'] in catIds
+ ]
+ anns = anns if len(areaRng) == 0 else [
+ ann for ann in anns if ann['avg_area'] > areaRng[0]
+ and ann['avg_area'] < areaRng[1]
+ ]
+ if iscrowd is not None:
+ ids = [ann['id'] for ann in anns if ann['iscrowd'] == iscrowd]
+ else:
+ ids = [ann['id'] for ann in anns]
+ return ids
+
+ def getCatIds(self, catNms=[], supNms=[], catIds=[]):
+ """filtering parameters. default skips that filter.
+
+ :param catNms (str array) : get cats for given cat names
+ :param supNms (str array) : get cats for given supercategory names
+ :param catIds (int array) : get cats for given cat ids
+ :return: ids (int array) : integer array of cat ids
+ """
+ catNms = catNms if _isArrayLike(catNms) else [catNms]
+ supNms = supNms if _isArrayLike(supNms) else [supNms]
+ catIds = catIds if _isArrayLike(catIds) else [catIds]
+
+ if len(catNms) == len(supNms) == len(catIds) == 0:
+ cats = self.dataset['categories']
+ else:
+ cats = self.dataset['categories']
+ cats = cats if len(catNms) == 0 else [
+ cat for cat in cats if cat['name'] in catNms
+ ]
+ cats = cats if len(supNms) == 0 else [
+ cat for cat in cats if cat['supercategory'] in supNms
+ ]
+ cats = cats if len(catIds) == 0 else [
+ cat for cat in cats if cat['id'] in catIds
+ ]
+ ids = [cat['id'] for cat in cats]
+ return ids
+
+ def getVidIds(self, vidIds=[], catIds=[]):
+ """Get vid ids that satisfy given filter conditions.
+
+ :param vidIds (int array) : get vids for given ids
+ :param catIds (int array) : get vids with all given cats
+ :return: ids (int array) : integer array of vid ids
+ """
+ vidIds = vidIds if _isArrayLike(vidIds) else [vidIds]
+ catIds = catIds if _isArrayLike(catIds) else [catIds]
+
+ if len(vidIds) == len(catIds) == 0:
+ ids = self.vids.keys()
+ else:
+ ids = set(vidIds)
+ for i, catId in enumerate(catIds):
+ if i == 0 and len(ids) == 0:
+ ids = set(self.catToVids[catId])
+ else:
+ ids &= set(self.catToVids[catId])
+ return list(ids)
+
+ def loadAnns(self, ids=[]):
+ """Load anns with the specified ids.
+
+ :param ids (int array) : integer ids specifying anns
+ :return: anns (object array) : loaded ann objects
+ """
+ if _isArrayLike(ids):
+ return [self.anns[id] for id in ids]
+ elif type(ids) == int:
+ return [self.anns[ids]]
+
+ def loadCats(self, ids=[]):
+ """Load cats with the specified ids.
+
+ :param ids (int array) : integer ids specifying cats
+ :return: cats (object array) : loaded cat objects
+ """
+ if _isArrayLike(ids):
+ return [self.cats[id] for id in ids]
+ elif type(ids) == int:
+ return [self.cats[ids]]
+
+ def loadVids(self, ids=[]):
+ """Load anns with the specified ids.
+
+ :param ids (int array) : integer ids specifying vid
+ :return: vids (object array) : loaded vid objects
+ """
+ if _isArrayLike(ids):
+ return [self.vids[id] for id in ids]
+ elif type(ids) == int:
+ return [self.vids[ids]]
+
+ def loadRes(self, resFile):
+ """Load result file and return a result api object.
+
+ :param resFile (str) : file name of result file
+ :return: res (obj) : result api object
+ """
+ res = YTVIS()
+ res.dataset['videos'] = [img for img in self.dataset['videos']]
+
+ print('Loading and preparing results...')
+ tic = time.time()
+ if type(resFile) == str or (PYTHON_VERSION == 2
+ and type(resFile) == str):
+ anns = json.load(open(resFile))
+ elif type(resFile) == np.ndarray:
+ anns = self.loadNumpyAnnotations(resFile)
+ else:
+ anns = resFile
+ assert type(anns) == list, 'results in not an array of objects'
+ annsVidIds = [ann['video_id'] for ann in anns]
+ assert set(annsVidIds) == (set(annsVidIds) & set(self.getVidIds())), \
+ 'Results do not correspond to current coco set'
+ if 'segmentations' in anns[0]:
+ res.dataset['categories'] = copy.deepcopy(
+ self.dataset['categories'])
+ for id, ann in enumerate(anns):
+ ann['areas'] = []
+ if 'bboxes' not in ann:
+ ann['bboxes'] = []
+ for seg in ann['segmentations']:
+ # now only support compressed RLE format
+ # as segmentation results
+ if seg:
+ ann['areas'].append(maskUtils.area(seg))
+ if len(ann['bboxes']) < len(ann['areas']):
+ ann['bboxes'].append(maskUtils.toBbox(seg))
+ else:
+ ann['areas'].append(None)
+ if len(ann['bboxes']) < len(ann['areas']):
+ ann['bboxes'].append(None)
+ ann['id'] = id + 1
+ l_ori = [a for a in ann['areas'] if a]
+ if len(l_ori) == 0:
+ ann['avg_area'] = 0
+ else:
+ ann['avg_area'] = np.array(l_ori).mean()
+ ann['iscrowd'] = 0
+ print('DONE (t={:0.2f}s)'.format(time.time() - tic))
+
+ res.dataset['annotations'] = anns
+ res.createIndex()
+ return res
+
+ def annToRLE(self, ann, frameId):
+ """Convert annotation which can be polygons, uncompressed RLE to RLE.
+
+ :return: binary mask (numpy 2D array)
+ """
+ t = self.vids[ann['video_id']]
+ h, w = t['height'], t['width']
+ segm = ann['segmentations'][frameId]
+ if type(segm) == list:
+ # polygon -- a single object might consist of multiple parts
+ # we merge all parts into one mask rle code
+ rles = maskUtils.frPyObjects(segm, h, w)
+ rle = maskUtils.merge(rles)
+ elif type(segm['counts']) == list:
+ # uncompressed RLE
+ rle = maskUtils.frPyObjects(segm, h, w)
+ else:
+ # rle
+ rle = segm
+ return rle
+
+ def annToMask(self, ann, frameId):
+ """Convert annotation which can be polygons, uncompressed RLE, or RLE
+ to binary mask.
+
+ :return: binary mask (numpy 2D array)
+ """
+ rle = self.annToRLE(ann, frameId)
+ m = maskUtils.decode(rle)
+ return m
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/ytviseval.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/ytviseval.py
new file mode 100644
index 0000000000000000000000000000000000000000..fdaf110d37c61b4e02873a4dc83e1722a70a29f1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/evaluation/functional/ytviseval.py
@@ -0,0 +1,623 @@
+# Copyright (c) Github URL
+# Copied from
+# https://github.com/youtubevos/cocoapi/blob/master/PythonAPI/pycocotools/ytvoseval.py
+__author__ = 'ychfan'
+
+import copy
+import datetime
+import time
+from collections import defaultdict
+
+import numpy as np
+from pycocotools import mask as maskUtils
+
+
+class YTVISeval:
+ # Interface for evaluating video instance segmentation on
+ # the YouTubeVIS dataset.
+ #
+ # The usage for YTVISeval is as follows:
+ # cocoGt=..., cocoDt=... # load dataset and results
+ # E = YTVISeval(cocoGt,cocoDt); # initialize YTVISeval object
+ # E.params.recThrs = ...; # set parameters as desired
+ # E.evaluate(); # run per image evaluation
+ # E.accumulate(); # accumulate per image results
+ # E.summarize(); # display summary metrics of results
+ # For example usage see evalDemo.m and http://mscoco.org/.
+ #
+ # The evaluation parameters are as follows (defaults in brackets):
+ # imgIds - [all] N img ids to use for evaluation
+ # catIds - [all] K cat ids to use for evaluation
+ # iouThrs - [.5:.05:.95] T=10 IoU thresholds for evaluation
+ # recThrs - [0:.01:1] R=101 recall thresholds for evaluation
+ # areaRng - [...] A=4 object area ranges for evaluation
+ # maxDets - [1 10 100] M=3 thresholds on max detections per image
+ # iouType - ['segm'] set iouType to 'segm', 'bbox' or 'keypoints'
+ # iouType replaced the now DEPRECATED useSegm parameter.
+ # useCats - [1] if true use category labels for evaluation
+ # Note: if useCats=0 category labels are ignored as in proposal scoring.
+ # Note: multiple areaRngs [Ax2] and maxDets [Mx1] can be specified.
+ #
+ # evaluate(): evaluates detections on every image and every category and
+ # concats the results into the "evalImgs" with fields:
+ # dtIds - [1xD] id for each of the D detections (dt)
+ # gtIds - [1xG] id for each of the G ground truths (gt)
+ # dtMatches - [TxD] matching gt id at each IoU or 0
+ # gtMatches - [TxG] matching dt id at each IoU or 0
+ # dtScores - [1xD] confidence of each dt
+ # gtIgnore - [1xG] ignore flag for each gt
+ # dtIgnore - [TxD] ignore flag for each dt at each IoU
+ #
+ # accumulate(): accumulates the per-image, per-category evaluation
+ # results in "evalImgs" into the dictionary "eval" with fields:
+ # params - parameters used for evaluation
+ # date - date evaluation was performed
+ # counts - [T,R,K,A,M] parameter dimensions (see above)
+ # precision - [TxRxKxAxM] precision for every evaluation setting
+ # recall - [TxKxAxM] max recall for every evaluation setting
+ # Note: precision and recall==-1 for settings with no gt objects.
+ #
+ # See also coco, mask, pycocoDemo, pycocoEvalDemo
+ #
+ # Microsoft COCO Toolbox. version 2.0
+ # Data, paper, and tutorials available at: http://mscoco.org/
+ # Code written by Piotr Dollar and Tsung-Yi Lin, 2015.
+ # Licensed under the Simplified BSD License [see coco/license.txt]
+ def __init__(self, cocoGt=None, cocoDt=None, iouType='segm'):
+ """Initialize CocoEval using coco APIs for gt and dt.
+
+ :param cocoGt: coco object with ground truth annotations
+ :param cocoDt: coco object with detection results
+ :return: None
+ """
+ if not iouType:
+ print('iouType not specified. use default iouType segm')
+ self.cocoGt = cocoGt # ground truth COCO API
+ self.cocoDt = cocoDt # detections COCO API
+ self.params = {} # evaluation parameters
+ self.evalVids = defaultdict(
+ list) # per-image per-category evaluation results [KxAxI] elements
+ self.eval = {} # accumulated evaluation results
+ self._gts = defaultdict(list) # gt for evaluation
+ self._dts = defaultdict(list) # dt for evaluation
+ self.params = Params(iouType=iouType) # parameters
+ self._paramsEval = {} # parameters for evaluation
+ self.stats = [] # result summarization
+ self.ious = {} # ious between all gts and dts
+ if cocoGt is not None:
+ self.params.vidIds = sorted(cocoGt.getVidIds())
+ self.params.catIds = sorted(cocoGt.getCatIds())
+
+ def _prepare(self):
+ '''
+ Prepare ._gts and ._dts for evaluation based on params
+ :return: None
+ '''
+
+ def _toMask(anns, coco):
+ # modify ann['segmentation'] by reference
+ for ann in anns:
+ for i, a in enumerate(ann['segmentations']):
+ if a:
+ rle = coco.annToRLE(ann, i)
+ ann['segmentations'][i] = rle
+ l_ori = [a for a in ann['areas'] if a]
+ if len(l_ori) == 0:
+ ann['avg_area'] = 0
+ else:
+ ann['avg_area'] = np.array(l_ori).mean()
+
+ p = self.params
+ if p.useCats:
+ gts = self.cocoGt.loadAnns(
+ self.cocoGt.getAnnIds(vidIds=p.vidIds, catIds=p.catIds))
+ dts = self.cocoDt.loadAnns(
+ self.cocoDt.getAnnIds(vidIds=p.vidIds, catIds=p.catIds))
+ else:
+ gts = self.cocoGt.loadAnns(self.cocoGt.getAnnIds(vidIds=p.vidIds))
+ dts = self.cocoDt.loadAnns(self.cocoDt.getAnnIds(vidIds=p.vidIds))
+
+ # convert ground truth to mask if iouType == 'segm'
+ if p.iouType == 'segm':
+ _toMask(gts, self.cocoGt)
+ _toMask(dts, self.cocoDt)
+ # set ignore flag
+ for gt in gts:
+ gt['ignore'] = gt['ignore'] if 'ignore' in gt else 0
+ gt['ignore'] = 'iscrowd' in gt and gt['iscrowd']
+ if p.iouType == 'keypoints':
+ gt['ignore'] = (gt['num_keypoints'] == 0) or gt['ignore']
+ self._gts = defaultdict(list) # gt for evaluation
+ self._dts = defaultdict(list) # dt for evaluation
+ for gt in gts:
+ self._gts[gt['video_id'], gt['category_id']].append(gt)
+ for dt in dts:
+ self._dts[dt['video_id'], dt['category_id']].append(dt)
+ self.evalVids = defaultdict(
+ list) # per-image per-category evaluation results
+ self.eval = {} # accumulated evaluation results
+
+ def evaluate(self):
+ '''
+ Run per image evaluation on given images and store
+ results (a list of dict) in self.evalVids
+ :return: None
+ '''
+ tic = time.time()
+ print('Running per image evaluation...')
+ p = self.params
+ # add backward compatibility if useSegm is specified in params
+ if p.useSegm is not None:
+ p.iouType = 'segm' if p.useSegm == 1 else 'bbox'
+ print('useSegm (deprecated) is not None. Running {} evaluation'.
+ format(p.iouType))
+ print('Evaluate annotation type *{}*'.format(p.iouType))
+ p.vidIds = list(np.unique(p.vidIds))
+ if p.useCats:
+ p.catIds = list(np.unique(p.catIds))
+ p.maxDets = sorted(p.maxDets)
+ self.params = p
+
+ self._prepare()
+ # loop through images, area range, max detection number
+ catIds = p.catIds if p.useCats else [-1]
+
+ if p.iouType == 'segm' or p.iouType == 'bbox':
+ computeIoU = self.computeIoU
+ elif p.iouType == 'keypoints':
+ computeIoU = self.computeOks
+ self.ious = {(vidId, catId): computeIoU(vidId, catId)
+ for vidId in p.vidIds for catId in catIds}
+
+ evaluateVid = self.evaluateVid
+ maxDet = p.maxDets[-1]
+
+ self.evalImgs = [
+ evaluateVid(vidId, catId, areaRng, maxDet) for catId in catIds
+ for areaRng in p.areaRng for vidId in p.vidIds
+ ]
+ self._paramsEval = copy.deepcopy(self.params)
+ toc = time.time()
+ print('DONE (t={:0.2f}s).'.format(toc - tic))
+
+ def computeIoU(self, vidId, catId):
+ p = self.params
+ if p.useCats:
+ gt = self._gts[vidId, catId]
+ dt = self._dts[vidId, catId]
+ else:
+ gt = [_ for cId in p.catIds for _ in self._gts[vidId, cId]]
+ dt = [_ for cId in p.catIds for _ in self._dts[vidId, cId]]
+ if len(gt) == 0 and len(dt) == 0:
+ return []
+ inds = np.argsort([-d['score'] for d in dt], kind='mergesort')
+ dt = [dt[i] for i in inds]
+ if len(dt) > p.maxDets[-1]:
+ dt = dt[0:p.maxDets[-1]]
+
+ if p.iouType == 'segm':
+ g = [g['segmentations'] for g in gt]
+ d = [d['segmentations'] for d in dt]
+ elif p.iouType == 'bbox':
+ g = [g['bboxes'] for g in gt]
+ d = [d['bboxes'] for d in dt]
+ else:
+ raise Exception('unknown iouType for iou computation')
+
+ # compute iou between each dt and gt region
+
+ def iou_seq(d_seq, g_seq):
+ i = .0
+ u = .0
+ for d, g in zip(d_seq, g_seq):
+ if d and g:
+ i += maskUtils.area(maskUtils.merge([d, g], True))
+ u += maskUtils.area(maskUtils.merge([d, g], False))
+ elif not d and g:
+ u += maskUtils.area(g)
+ elif d and not g:
+ u += maskUtils.area(d)
+ if not u > .0:
+ print('Mask sizes in video {} and category {} may not match!'.
+ format(vidId, catId))
+ iou = i / u if u > .0 else .0
+ return iou
+
+ ious = np.zeros([len(d), len(g)])
+ for i, j in np.ndindex(ious.shape):
+ ious[i, j] = iou_seq(d[i], g[j])
+
+ return ious
+
+ def computeOks(self, imgId, catId):
+ p = self.params
+
+ gts = self._gts[imgId, catId]
+ dts = self._dts[imgId, catId]
+ inds = np.argsort([-d['score'] for d in dts], kind='mergesort')
+ dts = [dts[i] for i in inds]
+ if len(dts) > p.maxDets[-1]:
+ dts = dts[0:p.maxDets[-1]]
+ # if len(gts) == 0 and len(dts) == 0:
+ if len(gts) == 0 or len(dts) == 0:
+ return []
+ ious = np.zeros((len(dts), len(gts)))
+ sigmas = np.array([
+ .26, .25, .25, .35, .35, .79, .79, .72, .72, .62, .62, 1.07, 1.07,
+ .87, .87, .89, .89
+ ]) / 10.0
+ vars = (sigmas * 2)**2
+ k = len(sigmas)
+ # compute oks between each detection and ground truth object
+ for j, gt in enumerate(gts):
+ # create bounds for ignore regions(double the gt bbox)
+ g = np.array(gt['keypoints'])
+ xg = g[0::3]
+ yg = g[1::3]
+ vg = g[2::3]
+ k1 = np.count_nonzero(vg > 0)
+ bb = gt['bbox']
+ x0 = bb[0] - bb[2]
+ x1 = bb[0] + bb[2] * 2
+ y0 = bb[1] - bb[3]
+ y1 = bb[1] + bb[3] * 2
+ for i, dt in enumerate(dts):
+ d = np.array(dt['keypoints'])
+ xd = d[0::3]
+ yd = d[1::3]
+ if k1 > 0:
+ # measure the per-keypoint distance if keypoints visible
+ dx = xd - xg
+ dy = yd - yg
+ else:
+ # measure minimum distance to keypoints
+ z = np.zeros((k))
+ dx = np.max((z, x0 - xd), axis=0) + np.max(
+ (z, xd - x1), axis=0)
+ dy = np.max((z, y0 - yd), axis=0) + np.max(
+ (z, yd - y1), axis=0)
+ e = (dx**2 + dy**2) / vars / (gt['avg_area'] +
+ np.spacing(1)) / 2
+ if k1 > 0:
+ e = e[vg > 0]
+ ious[i, j] = np.sum(np.exp(-e)) / e.shape[0]
+ return ious
+
+ def evaluateVid(self, vidId, catId, aRng, maxDet):
+ '''
+ perform evaluation for single category and image
+ :return: dict (single image results)
+ '''
+ p = self.params
+ if p.useCats:
+ gt = self._gts[vidId, catId]
+ dt = self._dts[vidId, catId]
+ else:
+ gt = [_ for cId in p.catIds for _ in self._gts[vidId, cId]]
+ dt = [_ for cId in p.catIds for _ in self._dts[vidId, cId]]
+ if len(gt) == 0 and len(dt) == 0:
+ return None
+
+ for g in gt:
+ if g['ignore'] or (g['avg_area'] < aRng[0]
+ or g['avg_area'] > aRng[1]):
+ g['_ignore'] = 1
+ else:
+ g['_ignore'] = 0
+
+ # sort dt highest score first, sort gt ignore last
+ gtind = np.argsort([g['_ignore'] for g in gt], kind='mergesort')
+ gt = [gt[i] for i in gtind]
+ dtind = np.argsort([-d['score'] for d in dt], kind='mergesort')
+ dt = [dt[i] for i in dtind[0:maxDet]]
+ iscrowd = [int(o['iscrowd']) for o in gt]
+ # load computed ious
+ ious = self.ious[vidId, catId][:, gtind] if len(
+ self.ious[vidId, catId]) > 0 else self.ious[vidId, catId]
+
+ T = len(p.iouThrs)
+ G = len(gt)
+ D = len(dt)
+ gtm = np.zeros((T, G))
+ dtm = np.zeros((T, D))
+ gtIg = np.array([g['_ignore'] for g in gt])
+ dtIg = np.zeros((T, D))
+ if not len(ious) == 0:
+ for tind, t in enumerate(p.iouThrs):
+ for dind, d in enumerate(dt):
+ # information about best match so far (m=-1 -> unmatched)
+ iou = min([t, 1 - 1e-10])
+ m = -1
+ for gind, g in enumerate(gt):
+ # if this gt already matched, and not a crowd, continue
+ if gtm[tind, gind] > 0 and not iscrowd[gind]:
+ continue
+ # if dt matched to reg gt, and on ignore gt, stop
+ if m > -1 and gtIg[m] == 0 and gtIg[gind] == 1:
+ break
+ # continue to next gt unless better match made
+ if ious[dind, gind] < iou:
+ continue
+ # if match successful and best so far,
+ # store appropriately
+ iou = ious[dind, gind]
+ m = gind
+ # if match made store id of match for both dt and gt
+ if m == -1:
+ continue
+ dtIg[tind, dind] = gtIg[m]
+ dtm[tind, dind] = gt[m]['id']
+ gtm[tind, m] = d['id']
+ # set unmatched detections outside of area range to ignore
+ a = np.array([
+ d['avg_area'] < aRng[0] or d['avg_area'] > aRng[1] for d in dt
+ ]).reshape((1, len(dt)))
+ dtIg = np.logical_or(dtIg, np.logical_and(dtm == 0, np.repeat(a, T,
+ 0)))
+ # store results for given image and category
+ return {
+ 'video_id': vidId,
+ 'category_id': catId,
+ 'aRng': aRng,
+ 'maxDet': maxDet,
+ 'dtIds': [d['id'] for d in dt],
+ 'gtIds': [g['id'] for g in gt],
+ 'dtMatches': dtm,
+ 'gtMatches': gtm,
+ 'dtScores': [d['score'] for d in dt],
+ 'gtIgnore': gtIg,
+ 'dtIgnore': dtIg,
+ }
+
+ def accumulate(self, p=None):
+ """Accumulate per image evaluation results and store the result in
+ self.eval.
+
+ :param p: input params for evaluation
+ :return: None
+ """
+ print('Accumulating evaluation results...')
+ tic = time.time()
+ if not self.evalImgs:
+ print('Please run evaluate() first')
+ # allows input customized parameters
+ if p is None:
+ p = self.params
+ p.catIds = p.catIds if p.useCats == 1 else [-1]
+ T = len(p.iouThrs)
+ R = len(p.recThrs)
+ K = len(p.catIds) if p.useCats else 1
+ A = len(p.areaRng)
+ M = len(p.maxDets)
+ precision = -np.ones(
+ (T, R, K, A, M)) # -1 for the precision of absent categories
+ recall = -np.ones((T, K, A, M))
+ scores = -np.ones((T, R, K, A, M))
+
+ # create dictionary for future indexing
+ _pe = self._paramsEval
+ catIds = _pe.catIds if _pe.useCats else [-1]
+ setK = set(catIds)
+ setA = set(map(tuple, _pe.areaRng))
+ setM = set(_pe.maxDets)
+ setI = set(_pe.vidIds)
+ # get inds to evaluate
+ k_list = [n for n, k in enumerate(p.catIds) if k in setK]
+ m_list = [m for n, m in enumerate(p.maxDets) if m in setM]
+ a_list = [
+ n for n, a in enumerate(map(lambda x: tuple(x), p.areaRng))
+ if a in setA
+ ]
+ i_list = [n for n, i in enumerate(p.vidIds) if i in setI]
+ I0 = len(_pe.vidIds)
+ A0 = len(_pe.areaRng)
+ # retrieve E at each category, area range, and max number of detections
+ for k, k0 in enumerate(k_list):
+ Nk = k0 * A0 * I0
+ for a, a0 in enumerate(a_list):
+ Na = a0 * I0
+ for m, maxDet in enumerate(m_list):
+ E = [self.evalImgs[Nk + Na + i] for i in i_list]
+ E = [e for e in E if e is not None]
+ if len(E) == 0:
+ continue
+ dtScores = np.concatenate(
+ [e['dtScores'][0:maxDet] for e in E])
+
+ inds = np.argsort(-dtScores, kind='mergesort')
+ dtScoresSorted = dtScores[inds]
+
+ dtm = np.concatenate(
+ [e['dtMatches'][:, 0:maxDet] for e in E], axis=1)[:,
+ inds]
+ dtIg = np.concatenate(
+ [e['dtIgnore'][:, 0:maxDet] for e in E], axis=1)[:,
+ inds]
+ gtIg = np.concatenate([e['gtIgnore'] for e in E])
+ npig = np.count_nonzero(gtIg == 0)
+ if npig == 0:
+ continue
+ tps = np.logical_and(dtm, np.logical_not(dtIg))
+ fps = np.logical_and(
+ np.logical_not(dtm), np.logical_not(dtIg))
+
+ tp_sum = np.cumsum(tps, axis=1).astype(dtype=np.float)
+ fp_sum = np.cumsum(fps, axis=1).astype(dtype=np.float)
+ for t, (tp, fp) in enumerate(zip(tp_sum, fp_sum)):
+ tp = np.array(tp)
+ fp = np.array(fp)
+ nd_ori = len(tp)
+ rc = tp / npig
+ pr = tp / (fp + tp + np.spacing(1))
+ q = np.zeros((R, ))
+ ss = np.zeros((R, ))
+
+ if nd_ori:
+ recall[t, k, a, m] = rc[-1]
+ else:
+ recall[t, k, a, m] = 0
+
+ # use python array gets significant speed improvement
+ pr = pr.tolist()
+ q = q.tolist()
+
+ for i in range(nd_ori - 1, 0, -1):
+ if pr[i] > pr[i - 1]:
+ pr[i - 1] = pr[i]
+
+ inds = np.searchsorted(rc, p.recThrs, side='left')
+ try:
+ for ri, pi in enumerate(inds):
+ q[ri] = pr[pi]
+ ss[ri] = dtScoresSorted[pi]
+ except Exception:
+ pass
+ precision[t, :, k, a, m] = np.array(q)
+ scores[t, :, k, a, m] = np.array(ss)
+ self.eval = {
+ 'params': p,
+ 'counts': [T, R, K, A, M],
+ 'date': datetime.datetime.now().strftime('%Y-%m-%d %H:%M:%S'),
+ 'precision': precision,
+ 'recall': recall,
+ 'scores': scores,
+ }
+ toc = time.time()
+ print('DONE (t={:0.2f}s).'.format(toc - tic))
+
+ def summarize(self):
+ """Compute and display summary metrics for evaluation results.
+
+ Note this function can *only* be applied on the default parameter
+ setting
+ """
+
+ def _summarize(ap=1, iouThr=None, areaRng='all', maxDets=100):
+ p = self.params
+ iStr = ' {:<18} {} @[ IoU={:<9} | area={:>6s} | ' \
+ 'maxDets={:>3d} ] = {:0.3f}'
+ titleStr = 'Average Precision' if ap == 1 else 'Average Recall'
+ typeStr = '(AP)' if ap == 1 else '(AR)'
+ iouStr = '{:0.2f}:{:0.2f}'.format(p.iouThrs[0], p.iouThrs[-1]) \
+ if iouThr is None else '{:0.2f}'.format(iouThr)
+
+ aind = [
+ i for i, aRng in enumerate(p.areaRngLbl) if aRng == areaRng
+ ]
+ mind = [i for i, mDet in enumerate(p.maxDets) if mDet == maxDets]
+ if ap == 1:
+ # dimension of precision: [TxRxKxAxM]
+ s = self.eval['precision']
+ # IoU
+ if iouThr is not None:
+ t = np.where(iouThr == p.iouThrs)[0]
+ s = s[t]
+ s = s[:, :, :, aind, mind]
+ else:
+ # dimension of recall: [TxKxAxM]
+ s = self.eval['recall']
+ if iouThr is not None:
+ t = np.where(iouThr == p.iouThrs)[0]
+ s = s[t]
+ s = s[:, :, aind, mind]
+ if len(s[s > -1]) == 0:
+ mean_s = -1
+ else:
+ mean_s = np.mean(s[s > -1])
+ print(
+ iStr.format(titleStr, typeStr, iouStr, areaRng, maxDets,
+ mean_s))
+ return mean_s
+
+ def _summarizeDets():
+ stats = np.zeros((12, ))
+ stats[0] = _summarize(1)
+ stats[1] = _summarize(1, iouThr=.5, maxDets=self.params.maxDets[2])
+ stats[2] = _summarize(
+ 1, iouThr=.75, maxDets=self.params.maxDets[2])
+ stats[3] = _summarize(
+ 1, areaRng='small', maxDets=self.params.maxDets[2])
+ stats[4] = _summarize(
+ 1, areaRng='medium', maxDets=self.params.maxDets[2])
+ stats[5] = _summarize(
+ 1, areaRng='large', maxDets=self.params.maxDets[2])
+ stats[6] = _summarize(0, maxDets=self.params.maxDets[0])
+ stats[7] = _summarize(0, maxDets=self.params.maxDets[1])
+ stats[8] = _summarize(0, maxDets=self.params.maxDets[2])
+ stats[9] = _summarize(
+ 0, areaRng='small', maxDets=self.params.maxDets[2])
+ stats[10] = _summarize(
+ 0, areaRng='medium', maxDets=self.params.maxDets[2])
+ stats[11] = _summarize(
+ 0, areaRng='large', maxDets=self.params.maxDets[2])
+ return stats
+
+ def _summarizeKps():
+ stats = np.zeros((10, ))
+ stats[0] = _summarize(1, maxDets=20)
+ stats[1] = _summarize(1, maxDets=20, iouThr=.5)
+ stats[2] = _summarize(1, maxDets=20, iouThr=.75)
+ stats[3] = _summarize(1, maxDets=20, areaRng='medium')
+ stats[4] = _summarize(1, maxDets=20, areaRng='large')
+ stats[5] = _summarize(0, maxDets=20)
+ stats[6] = _summarize(0, maxDets=20, iouThr=.5)
+ stats[7] = _summarize(0, maxDets=20, iouThr=.75)
+ stats[8] = _summarize(0, maxDets=20, areaRng='medium')
+ stats[9] = _summarize(0, maxDets=20, areaRng='large')
+ return stats
+
+ if not self.eval:
+ raise Exception('Please run accumulate() first')
+ iouType = self.params.iouType
+ if iouType == 'segm' or iouType == 'bbox':
+ summarize = _summarizeDets
+ elif iouType == 'keypoints':
+ summarize = _summarizeKps
+ self.stats = summarize()
+
+ def __str__(self):
+ self.summarize()
+
+
+class Params:
+ """Params for coco evaluation api."""
+
+ def setDetParams(self):
+ self.vidIds = []
+ self.catIds = []
+ # np.arange causes trouble. the data point on arange
+ # is slightly larger than the true value
+ self.iouThrs = np.linspace(
+ .5, 0.95, int(np.round((0.95 - .5) / .05)) + 1, endpoint=True)
+ self.recThrs = np.linspace(
+ .0, 1.00, int(np.round((1.00 - .0) / .01)) + 1, endpoint=True)
+ self.maxDets = [1, 10, 100]
+ self.areaRng = [[0**2, 1e5**2], [0**2, 128**2], [128**2, 256**2],
+ [256**2, 1e5**2]]
+ self.areaRngLbl = ['all', 'small', 'medium', 'large']
+ self.useCats = 1
+
+ def setKpParams(self):
+ self.vidIds = []
+ self.catIds = []
+ # np.arange causes trouble. the data point on arange
+ # is slightly larger than the true value
+ self.iouThrs = np.linspace(
+ .5, 0.95, int(np.round((0.95 - .5) / .05)) + 1, endpoint=True)
+ self.recThrs = np.linspace(
+ .0, 1.00, int(np.round((1.00 - .0) / .01)) + 1, endpoint=True)
+ self.maxDets = [20]
+ self.areaRng = [[0**2, 1e5**2], [32**2, 96**2], [96**2, 1e5**2]]
+ self.areaRngLbl = ['all', 'medium', 'large']
+ self.useCats = 1
+
+ def __init__(self, iouType='segm'):
+ if iouType == 'segm' or iouType == 'bbox':
+ self.setDetParams()
+ elif iouType == 'keypoints':
+ self.setKpParams()
+ else:
+ raise Exception('iouType not supported')
+ self.iouType = iouType
+ # useSegm is deprecated
+ self.useSegm = None
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..c0a0d5e8d350d81e72787ff73fd85c2176783b43
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/__init__.py
@@ -0,0 +1,18 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .backbones import * # noqa: F401,F403
+from .data_preprocessors import * # noqa: F401,F403
+from .dense_heads import * # noqa: F401,F403
+from .detectors import * # noqa: F401,F403
+from .language_models import * # noqa: F401,F403
+from .layers import * # noqa: F401,F403
+from .losses import * # noqa: F401,F403
+from .mot import * # noqa: F401,F403
+from .necks import * # noqa: F401,F403
+from .reid import * # noqa: F401,F403
+from .roi_heads import * # noqa: F401,F403
+from .seg_heads import * # noqa: F401,F403
+from .task_modules import * # noqa: F401,F403
+from .test_time_augs import * # noqa: F401,F403
+from .trackers import * # noqa: F401,F403
+from .tracking_heads import * # noqa: F401,F403
+from .vis import * # noqa: F401,F403
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..e16ff85f7037b36fb2046fcbcd3af523050a6516
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/__init__.py
@@ -0,0 +1,27 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .csp_darknet import CSPDarknet
+from .cspnext import CSPNeXt
+from .darknet import Darknet
+from .detectors_resnet import DetectoRS_ResNet
+from .detectors_resnext import DetectoRS_ResNeXt
+from .efficientnet import EfficientNet
+from .hourglass import HourglassNet
+from .hrnet import HRNet
+from .mobilenet_v2 import MobileNetV2
+from .pvt import PyramidVisionTransformer, PyramidVisionTransformerV2
+from .regnet import RegNet
+from .res2net import Res2Net
+from .resnest import ResNeSt
+from .resnet import ResNet, ResNetV1d
+from .resnext import ResNeXt
+from .ssd_vgg import SSDVGG
+from .swin import SwinTransformer
+from .trident_resnet import TridentResNet
+
+__all__ = [
+ 'RegNet', 'ResNet', 'ResNetV1d', 'ResNeXt', 'SSDVGG', 'HRNet',
+ 'MobileNetV2', 'Res2Net', 'HourglassNet', 'DetectoRS_ResNet',
+ 'DetectoRS_ResNeXt', 'Darknet', 'ResNeSt', 'TridentResNet', 'CSPDarknet',
+ 'SwinTransformer', 'PyramidVisionTransformer',
+ 'PyramidVisionTransformerV2', 'EfficientNet', 'CSPNeXt'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/csp_darknet.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/csp_darknet.py
new file mode 100644
index 0000000000000000000000000000000000000000..a890b486f255befa23fe5a3e9746f8f9298ac33f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/csp_darknet.py
@@ -0,0 +1,286 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule, DepthwiseSeparableConvModule
+from mmengine.model import BaseModule
+from torch.nn.modules.batchnorm import _BatchNorm
+
+from mmdet.registry import MODELS
+from ..layers import CSPLayer
+
+
+class Focus(nn.Module):
+ """Focus width and height information into channel space.
+
+ Args:
+ in_channels (int): The input channels of this Module.
+ out_channels (int): The output channels of this Module.
+ kernel_size (int): The kernel size of the convolution. Default: 1
+ stride (int): The stride of the convolution. Default: 1
+ conv_cfg (dict): Config dict for convolution layer. Default: None,
+ which means using conv2d.
+ norm_cfg (dict): Config dict for normalization layer.
+ Default: dict(type='BN', momentum=0.03, eps=0.001).
+ act_cfg (dict): Config dict for activation layer.
+ Default: dict(type='Swish').
+ """
+
+ def __init__(self,
+ in_channels,
+ out_channels,
+ kernel_size=1,
+ stride=1,
+ conv_cfg=None,
+ norm_cfg=dict(type='BN', momentum=0.03, eps=0.001),
+ act_cfg=dict(type='Swish')):
+ super().__init__()
+ self.conv = ConvModule(
+ in_channels * 4,
+ out_channels,
+ kernel_size,
+ stride,
+ padding=(kernel_size - 1) // 2,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+
+ def forward(self, x):
+ # shape of x (b,c,w,h) -> y(b,4c,w/2,h/2)
+ patch_top_left = x[..., ::2, ::2]
+ patch_top_right = x[..., ::2, 1::2]
+ patch_bot_left = x[..., 1::2, ::2]
+ patch_bot_right = x[..., 1::2, 1::2]
+ x = torch.cat(
+ (
+ patch_top_left,
+ patch_bot_left,
+ patch_top_right,
+ patch_bot_right,
+ ),
+ dim=1,
+ )
+ return self.conv(x)
+
+
+class SPPBottleneck(BaseModule):
+ """Spatial pyramid pooling layer used in YOLOv3-SPP.
+
+ Args:
+ in_channels (int): The input channels of this Module.
+ out_channels (int): The output channels of this Module.
+ kernel_sizes (tuple[int]): Sequential of kernel sizes of pooling
+ layers. Default: (5, 9, 13).
+ conv_cfg (dict): Config dict for convolution layer. Default: None,
+ which means using conv2d.
+ norm_cfg (dict): Config dict for normalization layer.
+ Default: dict(type='BN').
+ act_cfg (dict): Config dict for activation layer.
+ Default: dict(type='Swish').
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None.
+ """
+
+ def __init__(self,
+ in_channels,
+ out_channels,
+ kernel_sizes=(5, 9, 13),
+ conv_cfg=None,
+ norm_cfg=dict(type='BN', momentum=0.03, eps=0.001),
+ act_cfg=dict(type='Swish'),
+ init_cfg=None):
+ super().__init__(init_cfg)
+ mid_channels = in_channels // 2
+ self.conv1 = ConvModule(
+ in_channels,
+ mid_channels,
+ 1,
+ stride=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+ self.poolings = nn.ModuleList([
+ nn.MaxPool2d(kernel_size=ks, stride=1, padding=ks // 2)
+ for ks in kernel_sizes
+ ])
+ conv2_channels = mid_channels * (len(kernel_sizes) + 1)
+ self.conv2 = ConvModule(
+ conv2_channels,
+ out_channels,
+ 1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+
+ def forward(self, x):
+ x = self.conv1(x)
+ with torch.cuda.amp.autocast(enabled=False):
+ x = torch.cat(
+ [x] + [pooling(x) for pooling in self.poolings], dim=1)
+ x = self.conv2(x)
+ return x
+
+
+@MODELS.register_module()
+class CSPDarknet(BaseModule):
+ """CSP-Darknet backbone used in YOLOv5 and YOLOX.
+
+ Args:
+ arch (str): Architecture of CSP-Darknet, from {P5, P6}.
+ Default: P5.
+ deepen_factor (float): Depth multiplier, multiply number of
+ blocks in CSP layer by this amount. Default: 1.0.
+ widen_factor (float): Width multiplier, multiply number of
+ channels in each layer by this amount. Default: 1.0.
+ out_indices (Sequence[int]): Output from which stages.
+ Default: (2, 3, 4).
+ frozen_stages (int): Stages to be frozen (stop grad and set eval
+ mode). -1 means not freezing any parameters. Default: -1.
+ use_depthwise (bool): Whether to use depthwise separable convolution.
+ Default: False.
+ arch_ovewrite(list): Overwrite default arch settings. Default: None.
+ spp_kernal_sizes: (tuple[int]): Sequential of kernel sizes of SPP
+ layers. Default: (5, 9, 13).
+ conv_cfg (dict): Config dict for convolution layer. Default: None.
+ norm_cfg (dict): Dictionary to construct and config norm layer.
+ Default: dict(type='BN', requires_grad=True).
+ act_cfg (dict): Config dict for activation layer.
+ Default: dict(type='LeakyReLU', negative_slope=0.1).
+ norm_eval (bool): Whether to set norm layers to eval mode, namely,
+ freeze running stats (mean and var). Note: Effect on Batch Norm
+ and its variants only.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None.
+ Example:
+ >>> from mmdet.models import CSPDarknet
+ >>> import torch
+ >>> self = CSPDarknet(depth=53)
+ >>> self.eval()
+ >>> inputs = torch.rand(1, 3, 416, 416)
+ >>> level_outputs = self.forward(inputs)
+ >>> for level_out in level_outputs:
+ ... print(tuple(level_out.shape))
+ ...
+ (1, 256, 52, 52)
+ (1, 512, 26, 26)
+ (1, 1024, 13, 13)
+ """
+ # From left to right:
+ # in_channels, out_channels, num_blocks, add_identity, use_spp
+ arch_settings = {
+ 'P5': [[64, 128, 3, True, False], [128, 256, 9, True, False],
+ [256, 512, 9, True, False], [512, 1024, 3, False, True]],
+ 'P6': [[64, 128, 3, True, False], [128, 256, 9, True, False],
+ [256, 512, 9, True, False], [512, 768, 3, True, False],
+ [768, 1024, 3, False, True]]
+ }
+
+ def __init__(self,
+ arch='P5',
+ deepen_factor=1.0,
+ widen_factor=1.0,
+ out_indices=(2, 3, 4),
+ frozen_stages=-1,
+ use_depthwise=False,
+ arch_ovewrite=None,
+ spp_kernal_sizes=(5, 9, 13),
+ conv_cfg=None,
+ norm_cfg=dict(type='BN', momentum=0.03, eps=0.001),
+ act_cfg=dict(type='Swish'),
+ norm_eval=False,
+ init_cfg=dict(
+ type='Kaiming',
+ layer='Conv2d',
+ a=math.sqrt(5),
+ distribution='uniform',
+ mode='fan_in',
+ nonlinearity='leaky_relu')):
+ super().__init__(init_cfg)
+ arch_setting = self.arch_settings[arch]
+ if arch_ovewrite:
+ arch_setting = arch_ovewrite
+ assert set(out_indices).issubset(
+ i for i in range(len(arch_setting) + 1))
+ if frozen_stages not in range(-1, len(arch_setting) + 1):
+ raise ValueError('frozen_stages must be in range(-1, '
+ 'len(arch_setting) + 1). But received '
+ f'{frozen_stages}')
+
+ self.out_indices = out_indices
+ self.frozen_stages = frozen_stages
+ self.use_depthwise = use_depthwise
+ self.norm_eval = norm_eval
+ conv = DepthwiseSeparableConvModule if use_depthwise else ConvModule
+
+ self.stem = Focus(
+ 3,
+ int(arch_setting[0][0] * widen_factor),
+ kernel_size=3,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+ self.layers = ['stem']
+
+ for i, (in_channels, out_channels, num_blocks, add_identity,
+ use_spp) in enumerate(arch_setting):
+ in_channels = int(in_channels * widen_factor)
+ out_channels = int(out_channels * widen_factor)
+ num_blocks = max(round(num_blocks * deepen_factor), 1)
+ stage = []
+ conv_layer = conv(
+ in_channels,
+ out_channels,
+ 3,
+ stride=2,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+ stage.append(conv_layer)
+ if use_spp:
+ spp = SPPBottleneck(
+ out_channels,
+ out_channels,
+ kernel_sizes=spp_kernal_sizes,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+ stage.append(spp)
+ csp_layer = CSPLayer(
+ out_channels,
+ out_channels,
+ num_blocks=num_blocks,
+ add_identity=add_identity,
+ use_depthwise=use_depthwise,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+ stage.append(csp_layer)
+ self.add_module(f'stage{i + 1}', nn.Sequential(*stage))
+ self.layers.append(f'stage{i + 1}')
+
+ def _freeze_stages(self):
+ if self.frozen_stages >= 0:
+ for i in range(self.frozen_stages + 1):
+ m = getattr(self, self.layers[i])
+ m.eval()
+ for param in m.parameters():
+ param.requires_grad = False
+
+ def train(self, mode=True):
+ super(CSPDarknet, self).train(mode)
+ self._freeze_stages()
+ if mode and self.norm_eval:
+ for m in self.modules():
+ if isinstance(m, _BatchNorm):
+ m.eval()
+
+ def forward(self, x):
+ outs = []
+ for i, layer_name in enumerate(self.layers):
+ layer = getattr(self, layer_name)
+ x = layer(x)
+ if i in self.out_indices:
+ outs.append(x)
+ return tuple(outs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/cspnext.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/cspnext.py
new file mode 100644
index 0000000000000000000000000000000000000000..269725a70224047a1f7f7564ba8199e38df25cc8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/cspnext.py
@@ -0,0 +1,195 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+from typing import Sequence, Tuple
+
+import torch.nn as nn
+from mmcv.cnn import ConvModule, DepthwiseSeparableConvModule
+from mmengine.model import BaseModule
+from torch import Tensor
+from torch.nn.modules.batchnorm import _BatchNorm
+
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from ..layers import CSPLayer
+from .csp_darknet import SPPBottleneck
+
+
+@MODELS.register_module()
+class CSPNeXt(BaseModule):
+ """CSPNeXt backbone used in RTMDet.
+
+ Args:
+ arch (str): Architecture of CSPNeXt, from {P5, P6}.
+ Defaults to P5.
+ expand_ratio (float): Ratio to adjust the number of channels of the
+ hidden layer. Defaults to 0.5.
+ deepen_factor (float): Depth multiplier, multiply number of
+ blocks in CSP layer by this amount. Defaults to 1.0.
+ widen_factor (float): Width multiplier, multiply number of
+ channels in each layer by this amount. Defaults to 1.0.
+ out_indices (Sequence[int]): Output from which stages.
+ Defaults to (2, 3, 4).
+ frozen_stages (int): Stages to be frozen (stop grad and set eval
+ mode). -1 means not freezing any parameters. Defaults to -1.
+ use_depthwise (bool): Whether to use depthwise separable convolution.
+ Defaults to False.
+ arch_ovewrite (list): Overwrite default arch settings.
+ Defaults to None.
+ spp_kernel_sizes: (tuple[int]): Sequential of kernel sizes of SPP
+ layers. Defaults to (5, 9, 13).
+ channel_attention (bool): Whether to add channel attention in each
+ stage. Defaults to True.
+ conv_cfg (:obj:`ConfigDict` or dict, optional): Config dict for
+ convolution layer. Defaults to None.
+ norm_cfg (:obj:`ConfigDict` or dict): Dictionary to construct and
+ config norm layer. Defaults to dict(type='BN', requires_grad=True).
+ act_cfg (:obj:`ConfigDict` or dict): Config dict for activation layer.
+ Defaults to dict(type='SiLU').
+ norm_eval (bool): Whether to set norm layers to eval mode, namely,
+ freeze running stats (mean and var). Note: Effect on Batch Norm
+ and its variants only.
+ init_cfg (:obj:`ConfigDict` or dict or list[dict] or
+ list[:obj:`ConfigDict`]): Initialization config dict.
+ """
+ # From left to right:
+ # in_channels, out_channels, num_blocks, add_identity, use_spp
+ arch_settings = {
+ 'P5': [[64, 128, 3, True, False], [128, 256, 6, True, False],
+ [256, 512, 6, True, False], [512, 1024, 3, False, True]],
+ 'P6': [[64, 128, 3, True, False], [128, 256, 6, True, False],
+ [256, 512, 6, True, False], [512, 768, 3, True, False],
+ [768, 1024, 3, False, True]]
+ }
+
+ def __init__(
+ self,
+ arch: str = 'P5',
+ deepen_factor: float = 1.0,
+ widen_factor: float = 1.0,
+ out_indices: Sequence[int] = (2, 3, 4),
+ frozen_stages: int = -1,
+ use_depthwise: bool = False,
+ expand_ratio: float = 0.5,
+ arch_ovewrite: dict = None,
+ spp_kernel_sizes: Sequence[int] = (5, 9, 13),
+ channel_attention: bool = True,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: ConfigType = dict(type='BN', momentum=0.03, eps=0.001),
+ act_cfg: ConfigType = dict(type='SiLU'),
+ norm_eval: bool = False,
+ init_cfg: OptMultiConfig = dict(
+ type='Kaiming',
+ layer='Conv2d',
+ a=math.sqrt(5),
+ distribution='uniform',
+ mode='fan_in',
+ nonlinearity='leaky_relu')
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ arch_setting = self.arch_settings[arch]
+ if arch_ovewrite:
+ arch_setting = arch_ovewrite
+ assert set(out_indices).issubset(
+ i for i in range(len(arch_setting) + 1))
+ if frozen_stages not in range(-1, len(arch_setting) + 1):
+ raise ValueError('frozen_stages must be in range(-1, '
+ 'len(arch_setting) + 1). But received '
+ f'{frozen_stages}')
+
+ self.out_indices = out_indices
+ self.frozen_stages = frozen_stages
+ self.use_depthwise = use_depthwise
+ self.norm_eval = norm_eval
+ conv = DepthwiseSeparableConvModule if use_depthwise else ConvModule
+ self.stem = nn.Sequential(
+ ConvModule(
+ 3,
+ int(arch_setting[0][0] * widen_factor // 2),
+ 3,
+ padding=1,
+ stride=2,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg),
+ ConvModule(
+ int(arch_setting[0][0] * widen_factor // 2),
+ int(arch_setting[0][0] * widen_factor // 2),
+ 3,
+ padding=1,
+ stride=1,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg),
+ ConvModule(
+ int(arch_setting[0][0] * widen_factor // 2),
+ int(arch_setting[0][0] * widen_factor),
+ 3,
+ padding=1,
+ stride=1,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg))
+ self.layers = ['stem']
+
+ for i, (in_channels, out_channels, num_blocks, add_identity,
+ use_spp) in enumerate(arch_setting):
+ in_channels = int(in_channels * widen_factor)
+ out_channels = int(out_channels * widen_factor)
+ num_blocks = max(round(num_blocks * deepen_factor), 1)
+ stage = []
+ conv_layer = conv(
+ in_channels,
+ out_channels,
+ 3,
+ stride=2,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+ stage.append(conv_layer)
+ if use_spp:
+ spp = SPPBottleneck(
+ out_channels,
+ out_channels,
+ kernel_sizes=spp_kernel_sizes,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+ stage.append(spp)
+ csp_layer = CSPLayer(
+ out_channels,
+ out_channels,
+ num_blocks=num_blocks,
+ add_identity=add_identity,
+ use_depthwise=use_depthwise,
+ use_cspnext_block=True,
+ expand_ratio=expand_ratio,
+ channel_attention=channel_attention,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+ stage.append(csp_layer)
+ self.add_module(f'stage{i + 1}', nn.Sequential(*stage))
+ self.layers.append(f'stage{i + 1}')
+
+ def _freeze_stages(self) -> None:
+ if self.frozen_stages >= 0:
+ for i in range(self.frozen_stages + 1):
+ m = getattr(self, self.layers[i])
+ m.eval()
+ for param in m.parameters():
+ param.requires_grad = False
+
+ def train(self, mode=True) -> None:
+ super().train(mode)
+ self._freeze_stages()
+ if mode and self.norm_eval:
+ for m in self.modules():
+ if isinstance(m, _BatchNorm):
+ m.eval()
+
+ def forward(self, x: Tuple[Tensor, ...]) -> Tuple[Tensor, ...]:
+ outs = []
+ for i, layer_name in enumerate(self.layers):
+ layer = getattr(self, layer_name)
+ x = layer(x)
+ if i in self.out_indices:
+ outs.append(x)
+ return tuple(outs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/darknet.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/darknet.py
new file mode 100644
index 0000000000000000000000000000000000000000..1d44da1e03f04a7e0801c10e5338277cf6244ab1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/darknet.py
@@ -0,0 +1,213 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+# Copyright (c) 2019 Western Digital Corporation or its affiliates.
+
+import warnings
+
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from mmengine.model import BaseModule
+from torch.nn.modules.batchnorm import _BatchNorm
+
+from mmdet.registry import MODELS
+
+
+class ResBlock(BaseModule):
+ """The basic residual block used in Darknet. Each ResBlock consists of two
+ ConvModules and the input is added to the final output. Each ConvModule is
+ composed of Conv, BN, and LeakyReLU. In YoloV3 paper, the first convLayer
+ has half of the number of the filters as much as the second convLayer. The
+ first convLayer has filter size of 1x1 and the second one has the filter
+ size of 3x3.
+
+ Args:
+ in_channels (int): The input channels. Must be even.
+ conv_cfg (dict): Config dict for convolution layer. Default: None.
+ norm_cfg (dict): Dictionary to construct and config norm layer.
+ Default: dict(type='BN', requires_grad=True)
+ act_cfg (dict): Config dict for activation layer.
+ Default: dict(type='LeakyReLU', negative_slope=0.1).
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None
+ """
+
+ def __init__(self,
+ in_channels,
+ conv_cfg=None,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ act_cfg=dict(type='LeakyReLU', negative_slope=0.1),
+ init_cfg=None):
+ super(ResBlock, self).__init__(init_cfg)
+ assert in_channels % 2 == 0 # ensure the in_channels is even
+ half_in_channels = in_channels // 2
+
+ # shortcut
+ cfg = dict(conv_cfg=conv_cfg, norm_cfg=norm_cfg, act_cfg=act_cfg)
+
+ self.conv1 = ConvModule(in_channels, half_in_channels, 1, **cfg)
+ self.conv2 = ConvModule(
+ half_in_channels, in_channels, 3, padding=1, **cfg)
+
+ def forward(self, x):
+ residual = x
+ out = self.conv1(x)
+ out = self.conv2(out)
+ out = out + residual
+
+ return out
+
+
+@MODELS.register_module()
+class Darknet(BaseModule):
+ """Darknet backbone.
+
+ Args:
+ depth (int): Depth of Darknet. Currently only support 53.
+ out_indices (Sequence[int]): Output from which stages.
+ frozen_stages (int): Stages to be frozen (stop grad and set eval mode).
+ -1 means not freezing any parameters. Default: -1.
+ conv_cfg (dict): Config dict for convolution layer. Default: None.
+ norm_cfg (dict): Dictionary to construct and config norm layer.
+ Default: dict(type='BN', requires_grad=True)
+ act_cfg (dict): Config dict for activation layer.
+ Default: dict(type='LeakyReLU', negative_slope=0.1).
+ norm_eval (bool): Whether to set norm layers to eval mode, namely,
+ freeze running stats (mean and var). Note: Effect on Batch Norm
+ and its variants only.
+ pretrained (str, optional): model pretrained path. Default: None
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None
+
+ Example:
+ >>> from mmdet.models import Darknet
+ >>> import torch
+ >>> self = Darknet(depth=53)
+ >>> self.eval()
+ >>> inputs = torch.rand(1, 3, 416, 416)
+ >>> level_outputs = self.forward(inputs)
+ >>> for level_out in level_outputs:
+ ... print(tuple(level_out.shape))
+ ...
+ (1, 256, 52, 52)
+ (1, 512, 26, 26)
+ (1, 1024, 13, 13)
+ """
+
+ # Dict(depth: (layers, channels))
+ arch_settings = {
+ 53: ((1, 2, 8, 8, 4), ((32, 64), (64, 128), (128, 256), (256, 512),
+ (512, 1024)))
+ }
+
+ def __init__(self,
+ depth=53,
+ out_indices=(3, 4, 5),
+ frozen_stages=-1,
+ conv_cfg=None,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ act_cfg=dict(type='LeakyReLU', negative_slope=0.1),
+ norm_eval=True,
+ pretrained=None,
+ init_cfg=None):
+ super(Darknet, self).__init__(init_cfg)
+ if depth not in self.arch_settings:
+ raise KeyError(f'invalid depth {depth} for darknet')
+
+ self.depth = depth
+ self.out_indices = out_indices
+ self.frozen_stages = frozen_stages
+ self.layers, self.channels = self.arch_settings[depth]
+
+ cfg = dict(conv_cfg=conv_cfg, norm_cfg=norm_cfg, act_cfg=act_cfg)
+
+ self.conv1 = ConvModule(3, 32, 3, padding=1, **cfg)
+
+ self.cr_blocks = ['conv1']
+ for i, n_layers in enumerate(self.layers):
+ layer_name = f'conv_res_block{i + 1}'
+ in_c, out_c = self.channels[i]
+ self.add_module(
+ layer_name,
+ self.make_conv_res_block(in_c, out_c, n_layers, **cfg))
+ self.cr_blocks.append(layer_name)
+
+ self.norm_eval = norm_eval
+
+ assert not (init_cfg and pretrained), \
+ 'init_cfg and pretrained cannot be specified at the same time'
+ if isinstance(pretrained, str):
+ warnings.warn('DeprecationWarning: pretrained is deprecated, '
+ 'please use "init_cfg" instead')
+ self.init_cfg = dict(type='Pretrained', checkpoint=pretrained)
+ elif pretrained is None:
+ if init_cfg is None:
+ self.init_cfg = [
+ dict(type='Kaiming', layer='Conv2d'),
+ dict(
+ type='Constant',
+ val=1,
+ layer=['_BatchNorm', 'GroupNorm'])
+ ]
+ else:
+ raise TypeError('pretrained must be a str or None')
+
+ def forward(self, x):
+ outs = []
+ for i, layer_name in enumerate(self.cr_blocks):
+ cr_block = getattr(self, layer_name)
+ x = cr_block(x)
+ if i in self.out_indices:
+ outs.append(x)
+
+ return tuple(outs)
+
+ def _freeze_stages(self):
+ if self.frozen_stages >= 0:
+ for i in range(self.frozen_stages):
+ m = getattr(self, self.cr_blocks[i])
+ m.eval()
+ for param in m.parameters():
+ param.requires_grad = False
+
+ def train(self, mode=True):
+ super(Darknet, self).train(mode)
+ self._freeze_stages()
+ if mode and self.norm_eval:
+ for m in self.modules():
+ if isinstance(m, _BatchNorm):
+ m.eval()
+
+ @staticmethod
+ def make_conv_res_block(in_channels,
+ out_channels,
+ res_repeat,
+ conv_cfg=None,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ act_cfg=dict(type='LeakyReLU',
+ negative_slope=0.1)):
+ """In Darknet backbone, ConvLayer is usually followed by ResBlock. This
+ function will make that. The Conv layers always have 3x3 filters with
+ stride=2. The number of the filters in Conv layer is the same as the
+ out channels of the ResBlock.
+
+ Args:
+ in_channels (int): The number of input channels.
+ out_channels (int): The number of output channels.
+ res_repeat (int): The number of ResBlocks.
+ conv_cfg (dict): Config dict for convolution layer. Default: None.
+ norm_cfg (dict): Dictionary to construct and config norm layer.
+ Default: dict(type='BN', requires_grad=True)
+ act_cfg (dict): Config dict for activation layer.
+ Default: dict(type='LeakyReLU', negative_slope=0.1).
+ """
+
+ cfg = dict(conv_cfg=conv_cfg, norm_cfg=norm_cfg, act_cfg=act_cfg)
+
+ model = nn.Sequential()
+ model.add_module(
+ 'conv',
+ ConvModule(
+ in_channels, out_channels, 3, stride=2, padding=1, **cfg))
+ for idx in range(res_repeat):
+ model.add_module('res{}'.format(idx),
+ ResBlock(out_channels, **cfg))
+ return model
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/detectors_resnet.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/detectors_resnet.py
new file mode 100644
index 0000000000000000000000000000000000000000..f33424fce4a933d675f1f1d3d4ad89e0173c5f9e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/detectors_resnet.py
@@ -0,0 +1,353 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch.nn as nn
+import torch.utils.checkpoint as cp
+from mmcv.cnn import build_conv_layer, build_norm_layer
+from mmengine.logging import MMLogger
+from mmengine.model import Sequential, constant_init, kaiming_init
+from mmengine.runner.checkpoint import load_checkpoint
+from torch.nn.modules.batchnorm import _BatchNorm
+
+from mmdet.registry import MODELS
+from .resnet import BasicBlock
+from .resnet import Bottleneck as _Bottleneck
+from .resnet import ResNet
+
+
+class Bottleneck(_Bottleneck):
+ r"""Bottleneck for the ResNet backbone in `DetectoRS
+ `_.
+
+ This bottleneck allows the users to specify whether to use
+ SAC (Switchable Atrous Convolution) and RFP (Recursive Feature Pyramid).
+
+ Args:
+ inplanes (int): The number of input channels.
+ planes (int): The number of output channels before expansion.
+ rfp_inplanes (int, optional): The number of channels from RFP.
+ Default: None. If specified, an additional conv layer will be
+ added for ``rfp_feat``. Otherwise, the structure is the same as
+ base class.
+ sac (dict, optional): Dictionary to construct SAC. Default: None.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None
+ """
+ expansion = 4
+
+ def __init__(self,
+ inplanes,
+ planes,
+ rfp_inplanes=None,
+ sac=None,
+ init_cfg=None,
+ **kwargs):
+ super(Bottleneck, self).__init__(
+ inplanes, planes, init_cfg=init_cfg, **kwargs)
+
+ assert sac is None or isinstance(sac, dict)
+ self.sac = sac
+ self.with_sac = sac is not None
+ if self.with_sac:
+ self.conv2 = build_conv_layer(
+ self.sac,
+ planes,
+ planes,
+ kernel_size=3,
+ stride=self.conv2_stride,
+ padding=self.dilation,
+ dilation=self.dilation,
+ bias=False)
+
+ self.rfp_inplanes = rfp_inplanes
+ if self.rfp_inplanes:
+ self.rfp_conv = build_conv_layer(
+ None,
+ self.rfp_inplanes,
+ planes * self.expansion,
+ 1,
+ stride=1,
+ bias=True)
+ if init_cfg is None:
+ self.init_cfg = dict(
+ type='Constant', val=0, override=dict(name='rfp_conv'))
+
+ def rfp_forward(self, x, rfp_feat):
+ """The forward function that also takes the RFP features as input."""
+
+ def _inner_forward(x):
+ identity = x
+
+ out = self.conv1(x)
+ out = self.norm1(out)
+ out = self.relu(out)
+
+ if self.with_plugins:
+ out = self.forward_plugin(out, self.after_conv1_plugin_names)
+
+ out = self.conv2(out)
+ out = self.norm2(out)
+ out = self.relu(out)
+
+ if self.with_plugins:
+ out = self.forward_plugin(out, self.after_conv2_plugin_names)
+
+ out = self.conv3(out)
+ out = self.norm3(out)
+
+ if self.with_plugins:
+ out = self.forward_plugin(out, self.after_conv3_plugin_names)
+
+ if self.downsample is not None:
+ identity = self.downsample(x)
+
+ out += identity
+
+ return out
+
+ if self.with_cp and x.requires_grad:
+ out = cp.checkpoint(_inner_forward, x)
+ else:
+ out = _inner_forward(x)
+
+ if self.rfp_inplanes:
+ rfp_feat = self.rfp_conv(rfp_feat)
+ out = out + rfp_feat
+
+ out = self.relu(out)
+
+ return out
+
+
+class ResLayer(Sequential):
+ """ResLayer to build ResNet style backbone for RPF in detectoRS.
+
+ The difference between this module and base class is that we pass
+ ``rfp_inplanes`` to the first block.
+
+ Args:
+ block (nn.Module): block used to build ResLayer.
+ inplanes (int): inplanes of block.
+ planes (int): planes of block.
+ num_blocks (int): number of blocks.
+ stride (int): stride of the first block. Default: 1
+ avg_down (bool): Use AvgPool instead of stride conv when
+ downsampling in the bottleneck. Default: False
+ conv_cfg (dict): dictionary to construct and config conv layer.
+ Default: None
+ norm_cfg (dict): dictionary to construct and config norm layer.
+ Default: dict(type='BN')
+ downsample_first (bool): Downsample at the first block or last block.
+ False for Hourglass, True for ResNet. Default: True
+ rfp_inplanes (int, optional): The number of channels from RFP.
+ Default: None. If specified, an additional conv layer will be
+ added for ``rfp_feat``. Otherwise, the structure is the same as
+ base class.
+ """
+
+ def __init__(self,
+ block,
+ inplanes,
+ planes,
+ num_blocks,
+ stride=1,
+ avg_down=False,
+ conv_cfg=None,
+ norm_cfg=dict(type='BN'),
+ downsample_first=True,
+ rfp_inplanes=None,
+ **kwargs):
+ self.block = block
+ assert downsample_first, f'downsample_first={downsample_first} is ' \
+ 'not supported in DetectoRS'
+
+ downsample = None
+ if stride != 1 or inplanes != planes * block.expansion:
+ downsample = []
+ conv_stride = stride
+ if avg_down and stride != 1:
+ conv_stride = 1
+ downsample.append(
+ nn.AvgPool2d(
+ kernel_size=stride,
+ stride=stride,
+ ceil_mode=True,
+ count_include_pad=False))
+ downsample.extend([
+ build_conv_layer(
+ conv_cfg,
+ inplanes,
+ planes * block.expansion,
+ kernel_size=1,
+ stride=conv_stride,
+ bias=False),
+ build_norm_layer(norm_cfg, planes * block.expansion)[1]
+ ])
+ downsample = nn.Sequential(*downsample)
+
+ layers = []
+ layers.append(
+ block(
+ inplanes=inplanes,
+ planes=planes,
+ stride=stride,
+ downsample=downsample,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ rfp_inplanes=rfp_inplanes,
+ **kwargs))
+ inplanes = planes * block.expansion
+ for _ in range(1, num_blocks):
+ layers.append(
+ block(
+ inplanes=inplanes,
+ planes=planes,
+ stride=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ **kwargs))
+
+ super(ResLayer, self).__init__(*layers)
+
+
+@MODELS.register_module()
+class DetectoRS_ResNet(ResNet):
+ """ResNet backbone for DetectoRS.
+
+ Args:
+ sac (dict, optional): Dictionary to construct SAC (Switchable Atrous
+ Convolution). Default: None.
+ stage_with_sac (list): Which stage to use sac. Default: (False, False,
+ False, False).
+ rfp_inplanes (int, optional): The number of channels from RFP.
+ Default: None. If specified, an additional conv layer will be
+ added for ``rfp_feat``. Otherwise, the structure is the same as
+ base class.
+ output_img (bool): If ``True``, the input image will be inserted into
+ the starting position of output. Default: False.
+ """
+
+ arch_settings = {
+ 50: (Bottleneck, (3, 4, 6, 3)),
+ 101: (Bottleneck, (3, 4, 23, 3)),
+ 152: (Bottleneck, (3, 8, 36, 3))
+ }
+
+ def __init__(self,
+ sac=None,
+ stage_with_sac=(False, False, False, False),
+ rfp_inplanes=None,
+ output_img=False,
+ pretrained=None,
+ init_cfg=None,
+ **kwargs):
+ assert not (init_cfg and pretrained), \
+ 'init_cfg and pretrained cannot be specified at the same time'
+ self.pretrained = pretrained
+ if init_cfg is not None:
+ assert isinstance(init_cfg, dict), \
+ f'init_cfg must be a dict, but got {type(init_cfg)}'
+ if 'type' in init_cfg:
+ assert init_cfg.get('type') == 'Pretrained', \
+ 'Only can initialize module by loading a pretrained model'
+ else:
+ raise KeyError('`init_cfg` must contain the key "type"')
+ self.pretrained = init_cfg.get('checkpoint')
+ self.sac = sac
+ self.stage_with_sac = stage_with_sac
+ self.rfp_inplanes = rfp_inplanes
+ self.output_img = output_img
+ super(DetectoRS_ResNet, self).__init__(**kwargs)
+
+ self.inplanes = self.stem_channels
+ self.res_layers = []
+ for i, num_blocks in enumerate(self.stage_blocks):
+ stride = self.strides[i]
+ dilation = self.dilations[i]
+ dcn = self.dcn if self.stage_with_dcn[i] else None
+ sac = self.sac if self.stage_with_sac[i] else None
+ if self.plugins is not None:
+ stage_plugins = self.make_stage_plugins(self.plugins, i)
+ else:
+ stage_plugins = None
+ planes = self.base_channels * 2**i
+ res_layer = self.make_res_layer(
+ block=self.block,
+ inplanes=self.inplanes,
+ planes=planes,
+ num_blocks=num_blocks,
+ stride=stride,
+ dilation=dilation,
+ style=self.style,
+ avg_down=self.avg_down,
+ with_cp=self.with_cp,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ dcn=dcn,
+ sac=sac,
+ rfp_inplanes=rfp_inplanes if i > 0 else None,
+ plugins=stage_plugins)
+ self.inplanes = planes * self.block.expansion
+ layer_name = f'layer{i + 1}'
+ self.add_module(layer_name, res_layer)
+ self.res_layers.append(layer_name)
+
+ self._freeze_stages()
+
+ # In order to be properly initialized by RFP
+ def init_weights(self):
+ # Calling this method will cause parameter initialization exception
+ # super(DetectoRS_ResNet, self).init_weights()
+
+ if isinstance(self.pretrained, str):
+ logger = MMLogger.get_current_instance()
+ load_checkpoint(self, self.pretrained, strict=False, logger=logger)
+ elif self.pretrained is None:
+ for m in self.modules():
+ if isinstance(m, nn.Conv2d):
+ kaiming_init(m)
+ elif isinstance(m, (_BatchNorm, nn.GroupNorm)):
+ constant_init(m, 1)
+
+ if self.dcn is not None:
+ for m in self.modules():
+ if isinstance(m, Bottleneck) and hasattr(
+ m.conv2, 'conv_offset'):
+ constant_init(m.conv2.conv_offset, 0)
+
+ if self.zero_init_residual:
+ for m in self.modules():
+ if isinstance(m, Bottleneck):
+ constant_init(m.norm3, 0)
+ elif isinstance(m, BasicBlock):
+ constant_init(m.norm2, 0)
+ else:
+ raise TypeError('pretrained must be a str or None')
+
+ def make_res_layer(self, **kwargs):
+ """Pack all blocks in a stage into a ``ResLayer`` for DetectoRS."""
+ return ResLayer(**kwargs)
+
+ def forward(self, x):
+ """Forward function."""
+ outs = list(super(DetectoRS_ResNet, self).forward(x))
+ if self.output_img:
+ outs.insert(0, x)
+ return tuple(outs)
+
+ def rfp_forward(self, x, rfp_feats):
+ """Forward function for RFP."""
+ if self.deep_stem:
+ x = self.stem(x)
+ else:
+ x = self.conv1(x)
+ x = self.norm1(x)
+ x = self.relu(x)
+ x = self.maxpool(x)
+ outs = []
+ for i, layer_name in enumerate(self.res_layers):
+ res_layer = getattr(self, layer_name)
+ rfp_feat = rfp_feats[i] if i > 0 else None
+ for layer in res_layer:
+ x = layer.rfp_forward(x, rfp_feat)
+ if i in self.out_indices:
+ outs.append(x)
+ return tuple(outs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/detectors_resnext.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/detectors_resnext.py
new file mode 100644
index 0000000000000000000000000000000000000000..4bbd63154bb47910e27cf6a75e4b359e050063e1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/detectors_resnext.py
@@ -0,0 +1,123 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+
+from mmcv.cnn import build_conv_layer, build_norm_layer
+
+from mmdet.registry import MODELS
+from .detectors_resnet import Bottleneck as _Bottleneck
+from .detectors_resnet import DetectoRS_ResNet
+
+
+class Bottleneck(_Bottleneck):
+ expansion = 4
+
+ def __init__(self,
+ inplanes,
+ planes,
+ groups=1,
+ base_width=4,
+ base_channels=64,
+ **kwargs):
+ """Bottleneck block for ResNeXt.
+
+ If style is "pytorch", the stride-two layer is the 3x3 conv layer, if
+ it is "caffe", the stride-two layer is the first 1x1 conv layer.
+ """
+ super(Bottleneck, self).__init__(inplanes, planes, **kwargs)
+
+ if groups == 1:
+ width = self.planes
+ else:
+ width = math.floor(self.planes *
+ (base_width / base_channels)) * groups
+
+ self.norm1_name, norm1 = build_norm_layer(
+ self.norm_cfg, width, postfix=1)
+ self.norm2_name, norm2 = build_norm_layer(
+ self.norm_cfg, width, postfix=2)
+ self.norm3_name, norm3 = build_norm_layer(
+ self.norm_cfg, self.planes * self.expansion, postfix=3)
+
+ self.conv1 = build_conv_layer(
+ self.conv_cfg,
+ self.inplanes,
+ width,
+ kernel_size=1,
+ stride=self.conv1_stride,
+ bias=False)
+ self.add_module(self.norm1_name, norm1)
+ fallback_on_stride = False
+ self.with_modulated_dcn = False
+ if self.with_dcn:
+ fallback_on_stride = self.dcn.pop('fallback_on_stride', False)
+ if self.with_sac:
+ self.conv2 = build_conv_layer(
+ self.sac,
+ width,
+ width,
+ kernel_size=3,
+ stride=self.conv2_stride,
+ padding=self.dilation,
+ dilation=self.dilation,
+ groups=groups,
+ bias=False)
+ elif not self.with_dcn or fallback_on_stride:
+ self.conv2 = build_conv_layer(
+ self.conv_cfg,
+ width,
+ width,
+ kernel_size=3,
+ stride=self.conv2_stride,
+ padding=self.dilation,
+ dilation=self.dilation,
+ groups=groups,
+ bias=False)
+ else:
+ assert self.conv_cfg is None, 'conv_cfg must be None for DCN'
+ self.conv2 = build_conv_layer(
+ self.dcn,
+ width,
+ width,
+ kernel_size=3,
+ stride=self.conv2_stride,
+ padding=self.dilation,
+ dilation=self.dilation,
+ groups=groups,
+ bias=False)
+
+ self.add_module(self.norm2_name, norm2)
+ self.conv3 = build_conv_layer(
+ self.conv_cfg,
+ width,
+ self.planes * self.expansion,
+ kernel_size=1,
+ bias=False)
+ self.add_module(self.norm3_name, norm3)
+
+
+@MODELS.register_module()
+class DetectoRS_ResNeXt(DetectoRS_ResNet):
+ """ResNeXt backbone for DetectoRS.
+
+ Args:
+ groups (int): The number of groups in ResNeXt.
+ base_width (int): The base width of ResNeXt.
+ """
+
+ arch_settings = {
+ 50: (Bottleneck, (3, 4, 6, 3)),
+ 101: (Bottleneck, (3, 4, 23, 3)),
+ 152: (Bottleneck, (3, 8, 36, 3))
+ }
+
+ def __init__(self, groups=1, base_width=4, **kwargs):
+ self.groups = groups
+ self.base_width = base_width
+ super(DetectoRS_ResNeXt, self).__init__(**kwargs)
+
+ def make_res_layer(self, **kwargs):
+ return super().make_res_layer(
+ groups=self.groups,
+ base_width=self.base_width,
+ base_channels=self.base_channels,
+ **kwargs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/efficientnet.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/efficientnet.py
new file mode 100644
index 0000000000000000000000000000000000000000..8484afe2e34e2bf8327e8aefedb968bd9a1e7792
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/efficientnet.py
@@ -0,0 +1,418 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import math
+from functools import partial
+
+import torch
+import torch.nn as nn
+import torch.utils.checkpoint as cp
+from mmcv.cnn.bricks import ConvModule, DropPath
+from mmengine.model import BaseModule, Sequential
+
+from mmdet.registry import MODELS
+from ..layers import InvertedResidual, SELayer
+from ..utils import make_divisible
+
+
+class EdgeResidual(BaseModule):
+ """Edge Residual Block.
+
+ Args:
+ in_channels (int): The input channels of this module.
+ out_channels (int): The output channels of this module.
+ mid_channels (int): The input channels of the second convolution.
+ kernel_size (int): The kernel size of the first convolution.
+ Defaults to 3.
+ stride (int): The stride of the first convolution. Defaults to 1.
+ se_cfg (dict, optional): Config dict for se layer. Defaults to None,
+ which means no se layer.
+ with_residual (bool): Use residual connection. Defaults to True.
+ conv_cfg (dict, optional): Config dict for convolution layer.
+ Defaults to None, which means using conv2d.
+ norm_cfg (dict): Config dict for normalization layer.
+ Defaults to ``dict(type='BN')``.
+ act_cfg (dict): Config dict for activation layer.
+ Defaults to ``dict(type='ReLU')``.
+ drop_path_rate (float): stochastic depth rate. Defaults to 0.
+ with_cp (bool): Use checkpoint or not. Using checkpoint will save some
+ memory while slowing down the training speed. Defaults to False.
+ init_cfg (dict | list[dict], optional): Initialization config dict.
+ """
+
+ def __init__(self,
+ in_channels,
+ out_channels,
+ mid_channels,
+ kernel_size=3,
+ stride=1,
+ se_cfg=None,
+ with_residual=True,
+ conv_cfg=None,
+ norm_cfg=dict(type='BN'),
+ act_cfg=dict(type='ReLU'),
+ drop_path_rate=0.,
+ with_cp=False,
+ init_cfg=None,
+ **kwargs):
+ super(EdgeResidual, self).__init__(init_cfg=init_cfg)
+ assert stride in [1, 2]
+ self.with_cp = with_cp
+ self.drop_path = DropPath(
+ drop_path_rate) if drop_path_rate > 0 else nn.Identity()
+ self.with_se = se_cfg is not None
+ self.with_residual = (
+ stride == 1 and in_channels == out_channels and with_residual)
+
+ if self.with_se:
+ assert isinstance(se_cfg, dict)
+
+ self.conv1 = ConvModule(
+ in_channels=in_channels,
+ out_channels=mid_channels,
+ kernel_size=kernel_size,
+ stride=1,
+ padding=kernel_size // 2,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+
+ if self.with_se:
+ self.se = SELayer(**se_cfg)
+
+ self.conv2 = ConvModule(
+ in_channels=mid_channels,
+ out_channels=out_channels,
+ kernel_size=1,
+ stride=stride,
+ padding=0,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=None)
+
+ def forward(self, x):
+
+ def _inner_forward(x):
+ out = x
+ out = self.conv1(out)
+
+ if self.with_se:
+ out = self.se(out)
+
+ out = self.conv2(out)
+
+ if self.with_residual:
+ return x + self.drop_path(out)
+ else:
+ return out
+
+ if self.with_cp and x.requires_grad:
+ out = cp.checkpoint(_inner_forward, x)
+ else:
+ out = _inner_forward(x)
+
+ return out
+
+
+def model_scaling(layer_setting, arch_setting):
+ """Scaling operation to the layer's parameters according to the
+ arch_setting."""
+ # scale width
+ new_layer_setting = copy.deepcopy(layer_setting)
+ for layer_cfg in new_layer_setting:
+ for block_cfg in layer_cfg:
+ block_cfg[1] = make_divisible(block_cfg[1] * arch_setting[0], 8)
+
+ # scale depth
+ split_layer_setting = [new_layer_setting[0]]
+ for layer_cfg in new_layer_setting[1:-1]:
+ tmp_index = [0]
+ for i in range(len(layer_cfg) - 1):
+ if layer_cfg[i + 1][1] != layer_cfg[i][1]:
+ tmp_index.append(i + 1)
+ tmp_index.append(len(layer_cfg))
+ for i in range(len(tmp_index) - 1):
+ split_layer_setting.append(layer_cfg[tmp_index[i]:tmp_index[i +
+ 1]])
+ split_layer_setting.append(new_layer_setting[-1])
+
+ num_of_layers = [len(layer_cfg) for layer_cfg in split_layer_setting[1:-1]]
+ new_layers = [
+ int(math.ceil(arch_setting[1] * num)) for num in num_of_layers
+ ]
+
+ merge_layer_setting = [split_layer_setting[0]]
+ for i, layer_cfg in enumerate(split_layer_setting[1:-1]):
+ if new_layers[i] <= num_of_layers[i]:
+ tmp_layer_cfg = layer_cfg[:new_layers[i]]
+ else:
+ tmp_layer_cfg = copy.deepcopy(layer_cfg) + [layer_cfg[-1]] * (
+ new_layers[i] - num_of_layers[i])
+ if tmp_layer_cfg[0][3] == 1 and i != 0:
+ merge_layer_setting[-1] += tmp_layer_cfg.copy()
+ else:
+ merge_layer_setting.append(tmp_layer_cfg.copy())
+ merge_layer_setting.append(split_layer_setting[-1])
+
+ return merge_layer_setting
+
+
+@MODELS.register_module()
+class EfficientNet(BaseModule):
+ """EfficientNet backbone.
+
+ Args:
+ arch (str): Architecture of efficientnet. Defaults to b0.
+ out_indices (Sequence[int]): Output from which stages.
+ Defaults to (6, ).
+ frozen_stages (int): Stages to be frozen (all param fixed).
+ Defaults to 0, which means not freezing any parameters.
+ conv_cfg (dict): Config dict for convolution layer.
+ Defaults to None, which means using conv2d.
+ norm_cfg (dict): Config dict for normalization layer.
+ Defaults to dict(type='BN').
+ act_cfg (dict): Config dict for activation layer.
+ Defaults to dict(type='Swish').
+ norm_eval (bool): Whether to set norm layers to eval mode, namely,
+ freeze running stats (mean and var). Note: Effect on Batch Norm
+ and its variants only. Defaults to False.
+ with_cp (bool): Use checkpoint or not. Using checkpoint will save some
+ memory while slowing down the training speed. Defaults to False.
+ """
+
+ # Parameters to build layers.
+ # 'b' represents the architecture of normal EfficientNet family includes
+ # 'b0', 'b1', 'b2', 'b3', 'b4', 'b5', 'b6', 'b7', 'b8'.
+ # 'e' represents the architecture of EfficientNet-EdgeTPU including 'es',
+ # 'em', 'el'.
+ # 6 parameters are needed to construct a layer, From left to right:
+ # - kernel_size: The kernel size of the block
+ # - out_channel: The number of out_channels of the block
+ # - se_ratio: The sequeeze ratio of SELayer.
+ # - stride: The stride of the block
+ # - expand_ratio: The expand_ratio of the mid_channels
+ # - block_type: -1: Not a block, 0: InvertedResidual, 1: EdgeResidual
+ layer_settings = {
+ 'b': [[[3, 32, 0, 2, 0, -1]],
+ [[3, 16, 4, 1, 1, 0]],
+ [[3, 24, 4, 2, 6, 0],
+ [3, 24, 4, 1, 6, 0]],
+ [[5, 40, 4, 2, 6, 0],
+ [5, 40, 4, 1, 6, 0]],
+ [[3, 80, 4, 2, 6, 0],
+ [3, 80, 4, 1, 6, 0],
+ [3, 80, 4, 1, 6, 0],
+ [5, 112, 4, 1, 6, 0],
+ [5, 112, 4, 1, 6, 0],
+ [5, 112, 4, 1, 6, 0]],
+ [[5, 192, 4, 2, 6, 0],
+ [5, 192, 4, 1, 6, 0],
+ [5, 192, 4, 1, 6, 0],
+ [5, 192, 4, 1, 6, 0],
+ [3, 320, 4, 1, 6, 0]],
+ [[1, 1280, 0, 1, 0, -1]]
+ ],
+ 'e': [[[3, 32, 0, 2, 0, -1]],
+ [[3, 24, 0, 1, 3, 1]],
+ [[3, 32, 0, 2, 8, 1],
+ [3, 32, 0, 1, 8, 1]],
+ [[3, 48, 0, 2, 8, 1],
+ [3, 48, 0, 1, 8, 1],
+ [3, 48, 0, 1, 8, 1],
+ [3, 48, 0, 1, 8, 1]],
+ [[5, 96, 0, 2, 8, 0],
+ [5, 96, 0, 1, 8, 0],
+ [5, 96, 0, 1, 8, 0],
+ [5, 96, 0, 1, 8, 0],
+ [5, 96, 0, 1, 8, 0],
+ [5, 144, 0, 1, 8, 0],
+ [5, 144, 0, 1, 8, 0],
+ [5, 144, 0, 1, 8, 0],
+ [5, 144, 0, 1, 8, 0]],
+ [[5, 192, 0, 2, 8, 0],
+ [5, 192, 0, 1, 8, 0]],
+ [[1, 1280, 0, 1, 0, -1]]
+ ]
+ } # yapf: disable
+
+ # Parameters to build different kinds of architecture.
+ # From left to right: scaling factor for width, scaling factor for depth,
+ # resolution.
+ arch_settings = {
+ 'b0': (1.0, 1.0, 224),
+ 'b1': (1.0, 1.1, 240),
+ 'b2': (1.1, 1.2, 260),
+ 'b3': (1.2, 1.4, 300),
+ 'b4': (1.4, 1.8, 380),
+ 'b5': (1.6, 2.2, 456),
+ 'b6': (1.8, 2.6, 528),
+ 'b7': (2.0, 3.1, 600),
+ 'b8': (2.2, 3.6, 672),
+ 'es': (1.0, 1.0, 224),
+ 'em': (1.0, 1.1, 240),
+ 'el': (1.2, 1.4, 300)
+ }
+
+ def __init__(self,
+ arch='b0',
+ drop_path_rate=0.,
+ out_indices=(6, ),
+ frozen_stages=0,
+ conv_cfg=dict(type='Conv2dAdaptivePadding'),
+ norm_cfg=dict(type='BN', eps=1e-3),
+ act_cfg=dict(type='Swish'),
+ norm_eval=False,
+ with_cp=False,
+ init_cfg=[
+ dict(type='Kaiming', layer='Conv2d'),
+ dict(
+ type='Constant',
+ layer=['_BatchNorm', 'GroupNorm'],
+ val=1)
+ ]):
+ super(EfficientNet, self).__init__(init_cfg)
+ assert arch in self.arch_settings, \
+ f'"{arch}" is not one of the arch_settings ' \
+ f'({", ".join(self.arch_settings.keys())})'
+ self.arch_setting = self.arch_settings[arch]
+ self.layer_setting = self.layer_settings[arch[:1]]
+ for index in out_indices:
+ if index not in range(0, len(self.layer_setting)):
+ raise ValueError('the item in out_indices must in '
+ f'range(0, {len(self.layer_setting)}). '
+ f'But received {index}')
+
+ if frozen_stages not in range(len(self.layer_setting) + 1):
+ raise ValueError('frozen_stages must be in range(0, '
+ f'{len(self.layer_setting) + 1}). '
+ f'But received {frozen_stages}')
+ self.drop_path_rate = drop_path_rate
+ self.out_indices = out_indices
+ self.frozen_stages = frozen_stages
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self.act_cfg = act_cfg
+ self.norm_eval = norm_eval
+ self.with_cp = with_cp
+
+ self.layer_setting = model_scaling(self.layer_setting,
+ self.arch_setting)
+ block_cfg_0 = self.layer_setting[0][0]
+ block_cfg_last = self.layer_setting[-1][0]
+ self.in_channels = make_divisible(block_cfg_0[1], 8)
+ self.out_channels = block_cfg_last[1]
+ self.layers = nn.ModuleList()
+ self.layers.append(
+ ConvModule(
+ in_channels=3,
+ out_channels=self.in_channels,
+ kernel_size=block_cfg_0[0],
+ stride=block_cfg_0[3],
+ padding=block_cfg_0[0] // 2,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg))
+ self.make_layer()
+ # Avoid building unused layers in mmdetection.
+ if len(self.layers) < max(self.out_indices) + 1:
+ self.layers.append(
+ ConvModule(
+ in_channels=self.in_channels,
+ out_channels=self.out_channels,
+ kernel_size=block_cfg_last[0],
+ stride=block_cfg_last[3],
+ padding=block_cfg_last[0] // 2,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg))
+
+ def make_layer(self):
+ # Without the first and the final conv block.
+ layer_setting = self.layer_setting[1:-1]
+
+ total_num_blocks = sum([len(x) for x in layer_setting])
+ block_idx = 0
+ dpr = [
+ x.item()
+ for x in torch.linspace(0, self.drop_path_rate, total_num_blocks)
+ ] # stochastic depth decay rule
+
+ for i, layer_cfg in enumerate(layer_setting):
+ # Avoid building unused layers in mmdetection.
+ if i > max(self.out_indices) - 1:
+ break
+ layer = []
+ for i, block_cfg in enumerate(layer_cfg):
+ (kernel_size, out_channels, se_ratio, stride, expand_ratio,
+ block_type) = block_cfg
+
+ mid_channels = int(self.in_channels * expand_ratio)
+ out_channels = make_divisible(out_channels, 8)
+ if se_ratio <= 0:
+ se_cfg = None
+ else:
+ # In mmdetection, the `divisor` is deleted to align
+ # the logic of SELayer with mmpretrain.
+ se_cfg = dict(
+ channels=mid_channels,
+ ratio=expand_ratio * se_ratio,
+ act_cfg=(self.act_cfg, dict(type='Sigmoid')))
+ if block_type == 1: # edge tpu
+ if i > 0 and expand_ratio == 3:
+ with_residual = False
+ expand_ratio = 4
+ else:
+ with_residual = True
+ mid_channels = int(self.in_channels * expand_ratio)
+ if se_cfg is not None:
+ # In mmdetection, the `divisor` is deleted to align
+ # the logic of SELayer with mmpretrain.
+ se_cfg = dict(
+ channels=mid_channels,
+ ratio=se_ratio * expand_ratio,
+ act_cfg=(self.act_cfg, dict(type='Sigmoid')))
+ block = partial(EdgeResidual, with_residual=with_residual)
+ else:
+ block = InvertedResidual
+ layer.append(
+ block(
+ in_channels=self.in_channels,
+ out_channels=out_channels,
+ mid_channels=mid_channels,
+ kernel_size=kernel_size,
+ stride=stride,
+ se_cfg=se_cfg,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg,
+ drop_path_rate=dpr[block_idx],
+ with_cp=self.with_cp,
+ # In mmdetection, `with_expand_conv` is set to align
+ # the logic of InvertedResidual with mmpretrain.
+ with_expand_conv=(mid_channels != self.in_channels)))
+ self.in_channels = out_channels
+ block_idx += 1
+ self.layers.append(Sequential(*layer))
+
+ def forward(self, x):
+ outs = []
+ for i, layer in enumerate(self.layers):
+ x = layer(x)
+ if i in self.out_indices:
+ outs.append(x)
+
+ return tuple(outs)
+
+ def _freeze_stages(self):
+ for i in range(self.frozen_stages):
+ m = self.layers[i]
+ m.eval()
+ for param in m.parameters():
+ param.requires_grad = False
+
+ def train(self, mode=True):
+ super(EfficientNet, self).train(mode)
+ self._freeze_stages()
+ if mode and self.norm_eval:
+ for m in self.modules():
+ if isinstance(m, nn.BatchNorm2d):
+ m.eval()
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/hourglass.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/hourglass.py
new file mode 100644
index 0000000000000000000000000000000000000000..bb58799f7b32138b3f58383419ddce9aa6d5ca18
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/hourglass.py
@@ -0,0 +1,225 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Sequence
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule
+from mmengine.model import BaseModule
+
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptMultiConfig
+from ..layers import ResLayer
+from .resnet import BasicBlock
+
+
+class HourglassModule(BaseModule):
+ """Hourglass Module for HourglassNet backbone.
+
+ Generate module recursively and use BasicBlock as the base unit.
+
+ Args:
+ depth (int): Depth of current HourglassModule.
+ stage_channels (list[int]): Feature channels of sub-modules in current
+ and follow-up HourglassModule.
+ stage_blocks (list[int]): Number of sub-modules stacked in current and
+ follow-up HourglassModule.
+ norm_cfg (ConfigType): Dictionary to construct and config norm layer.
+ Defaults to `dict(type='BN', requires_grad=True)`
+ upsample_cfg (ConfigType): Config dict for interpolate layer.
+ Defaults to `dict(mode='nearest')`
+ init_cfg (dict or ConfigDict, optional): the config to control the
+ initialization.
+ """
+
+ def __init__(self,
+ depth: int,
+ stage_channels: List[int],
+ stage_blocks: List[int],
+ norm_cfg: ConfigType = dict(type='BN', requires_grad=True),
+ upsample_cfg: ConfigType = dict(mode='nearest'),
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg)
+
+ self.depth = depth
+
+ cur_block = stage_blocks[0]
+ next_block = stage_blocks[1]
+
+ cur_channel = stage_channels[0]
+ next_channel = stage_channels[1]
+
+ self.up1 = ResLayer(
+ BasicBlock, cur_channel, cur_channel, cur_block, norm_cfg=norm_cfg)
+
+ self.low1 = ResLayer(
+ BasicBlock,
+ cur_channel,
+ next_channel,
+ cur_block,
+ stride=2,
+ norm_cfg=norm_cfg)
+
+ if self.depth > 1:
+ self.low2 = HourglassModule(depth - 1, stage_channels[1:],
+ stage_blocks[1:])
+ else:
+ self.low2 = ResLayer(
+ BasicBlock,
+ next_channel,
+ next_channel,
+ next_block,
+ norm_cfg=norm_cfg)
+
+ self.low3 = ResLayer(
+ BasicBlock,
+ next_channel,
+ cur_channel,
+ cur_block,
+ norm_cfg=norm_cfg,
+ downsample_first=False)
+
+ self.up2 = F.interpolate
+ self.upsample_cfg = upsample_cfg
+
+ def forward(self, x: torch.Tensor) -> nn.Module:
+ """Forward function."""
+ up1 = self.up1(x)
+ low1 = self.low1(x)
+ low2 = self.low2(low1)
+ low3 = self.low3(low2)
+ # Fixing `scale factor` (e.g. 2) is common for upsampling, but
+ # in some cases the spatial size is mismatched and error will arise.
+ if 'scale_factor' in self.upsample_cfg:
+ up2 = self.up2(low3, **self.upsample_cfg)
+ else:
+ shape = up1.shape[2:]
+ up2 = self.up2(low3, size=shape, **self.upsample_cfg)
+ return up1 + up2
+
+
+@MODELS.register_module()
+class HourglassNet(BaseModule):
+ """HourglassNet backbone.
+
+ Stacked Hourglass Networks for Human Pose Estimation.
+ More details can be found in the `paper
+ `_ .
+
+ Args:
+ downsample_times (int): Downsample times in a HourglassModule.
+ num_stacks (int): Number of HourglassModule modules stacked,
+ 1 for Hourglass-52, 2 for Hourglass-104.
+ stage_channels (Sequence[int]): Feature channel of each sub-module in a
+ HourglassModule.
+ stage_blocks (Sequence[int]): Number of sub-modules stacked in a
+ HourglassModule.
+ feat_channel (int): Feature channel of conv after a HourglassModule.
+ norm_cfg (norm_cfg): Dictionary to construct and config norm layer.
+ init_cfg (dict or ConfigDict, optional): the config to control the
+ initialization.
+
+ Example:
+ >>> from mmdet.models import HourglassNet
+ >>> import torch
+ >>> self = HourglassNet()
+ >>> self.eval()
+ >>> inputs = torch.rand(1, 3, 511, 511)
+ >>> level_outputs = self.forward(inputs)
+ >>> for level_output in level_outputs:
+ ... print(tuple(level_output.shape))
+ (1, 256, 128, 128)
+ (1, 256, 128, 128)
+ """
+
+ def __init__(self,
+ downsample_times: int = 5,
+ num_stacks: int = 2,
+ stage_channels: Sequence = (256, 256, 384, 384, 384, 512),
+ stage_blocks: Sequence = (2, 2, 2, 2, 2, 4),
+ feat_channel: int = 256,
+ norm_cfg: ConfigType = dict(type='BN', requires_grad=True),
+ init_cfg: OptMultiConfig = None) -> None:
+ assert init_cfg is None, 'To prevent abnormal initialization ' \
+ 'behavior, init_cfg is not allowed to be set'
+ super().__init__(init_cfg)
+
+ self.num_stacks = num_stacks
+ assert self.num_stacks >= 1
+ assert len(stage_channels) == len(stage_blocks)
+ assert len(stage_channels) > downsample_times
+
+ cur_channel = stage_channels[0]
+
+ self.stem = nn.Sequential(
+ ConvModule(
+ 3, cur_channel // 2, 7, padding=3, stride=2,
+ norm_cfg=norm_cfg),
+ ResLayer(
+ BasicBlock,
+ cur_channel // 2,
+ cur_channel,
+ 1,
+ stride=2,
+ norm_cfg=norm_cfg))
+
+ self.hourglass_modules = nn.ModuleList([
+ HourglassModule(downsample_times, stage_channels, stage_blocks)
+ for _ in range(num_stacks)
+ ])
+
+ self.inters = ResLayer(
+ BasicBlock,
+ cur_channel,
+ cur_channel,
+ num_stacks - 1,
+ norm_cfg=norm_cfg)
+
+ self.conv1x1s = nn.ModuleList([
+ ConvModule(
+ cur_channel, cur_channel, 1, norm_cfg=norm_cfg, act_cfg=None)
+ for _ in range(num_stacks - 1)
+ ])
+
+ self.out_convs = nn.ModuleList([
+ ConvModule(
+ cur_channel, feat_channel, 3, padding=1, norm_cfg=norm_cfg)
+ for _ in range(num_stacks)
+ ])
+
+ self.remap_convs = nn.ModuleList([
+ ConvModule(
+ feat_channel, cur_channel, 1, norm_cfg=norm_cfg, act_cfg=None)
+ for _ in range(num_stacks - 1)
+ ])
+
+ self.relu = nn.ReLU(inplace=True)
+
+ def init_weights(self) -> None:
+ """Init module weights."""
+ # Training Centripetal Model needs to reset parameters for Conv2d
+ super().init_weights()
+ for m in self.modules():
+ if isinstance(m, nn.Conv2d):
+ m.reset_parameters()
+
+ def forward(self, x: torch.Tensor) -> List[torch.Tensor]:
+ """Forward function."""
+ inter_feat = self.stem(x)
+ out_feats = []
+
+ for ind in range(self.num_stacks):
+ single_hourglass = self.hourglass_modules[ind]
+ out_conv = self.out_convs[ind]
+
+ hourglass_feat = single_hourglass(inter_feat)
+ out_feat = out_conv(hourglass_feat)
+ out_feats.append(out_feat)
+
+ if ind < self.num_stacks - 1:
+ inter_feat = self.conv1x1s[ind](
+ inter_feat) + self.remap_convs[ind](
+ out_feat)
+ inter_feat = self.inters[ind](self.relu(inter_feat))
+
+ return out_feats
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/hrnet.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/hrnet.py
new file mode 100644
index 0000000000000000000000000000000000000000..77bd3cc7125bb7ba03cd201ab3a55174b01dde50
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/hrnet.py
@@ -0,0 +1,589 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+
+import torch.nn as nn
+from mmcv.cnn import build_conv_layer, build_norm_layer
+from mmengine.model import BaseModule, ModuleList, Sequential
+from torch.nn.modules.batchnorm import _BatchNorm
+
+from mmdet.registry import MODELS
+from .resnet import BasicBlock, Bottleneck
+
+
+class HRModule(BaseModule):
+ """High-Resolution Module for HRNet.
+
+ In this module, every branch has 4 BasicBlocks/Bottlenecks. Fusion/Exchange
+ is in this module.
+ """
+
+ def __init__(self,
+ num_branches,
+ blocks,
+ num_blocks,
+ in_channels,
+ num_channels,
+ multiscale_output=True,
+ with_cp=False,
+ conv_cfg=None,
+ norm_cfg=dict(type='BN'),
+ block_init_cfg=None,
+ init_cfg=None):
+ super(HRModule, self).__init__(init_cfg)
+ self.block_init_cfg = block_init_cfg
+ self._check_branches(num_branches, num_blocks, in_channels,
+ num_channels)
+
+ self.in_channels = in_channels
+ self.num_branches = num_branches
+
+ self.multiscale_output = multiscale_output
+ self.norm_cfg = norm_cfg
+ self.conv_cfg = conv_cfg
+ self.with_cp = with_cp
+ self.branches = self._make_branches(num_branches, blocks, num_blocks,
+ num_channels)
+ self.fuse_layers = self._make_fuse_layers()
+ self.relu = nn.ReLU(inplace=False)
+
+ def _check_branches(self, num_branches, num_blocks, in_channels,
+ num_channels):
+ if num_branches != len(num_blocks):
+ error_msg = f'NUM_BRANCHES({num_branches}) ' \
+ f'!= NUM_BLOCKS({len(num_blocks)})'
+ raise ValueError(error_msg)
+
+ if num_branches != len(num_channels):
+ error_msg = f'NUM_BRANCHES({num_branches}) ' \
+ f'!= NUM_CHANNELS({len(num_channels)})'
+ raise ValueError(error_msg)
+
+ if num_branches != len(in_channels):
+ error_msg = f'NUM_BRANCHES({num_branches}) ' \
+ f'!= NUM_INCHANNELS({len(in_channels)})'
+ raise ValueError(error_msg)
+
+ def _make_one_branch(self,
+ branch_index,
+ block,
+ num_blocks,
+ num_channels,
+ stride=1):
+ downsample = None
+ if stride != 1 or \
+ self.in_channels[branch_index] != \
+ num_channels[branch_index] * block.expansion:
+ downsample = nn.Sequential(
+ build_conv_layer(
+ self.conv_cfg,
+ self.in_channels[branch_index],
+ num_channels[branch_index] * block.expansion,
+ kernel_size=1,
+ stride=stride,
+ bias=False),
+ build_norm_layer(self.norm_cfg, num_channels[branch_index] *
+ block.expansion)[1])
+
+ layers = []
+ layers.append(
+ block(
+ self.in_channels[branch_index],
+ num_channels[branch_index],
+ stride,
+ downsample=downsample,
+ with_cp=self.with_cp,
+ norm_cfg=self.norm_cfg,
+ conv_cfg=self.conv_cfg,
+ init_cfg=self.block_init_cfg))
+ self.in_channels[branch_index] = \
+ num_channels[branch_index] * block.expansion
+ for i in range(1, num_blocks[branch_index]):
+ layers.append(
+ block(
+ self.in_channels[branch_index],
+ num_channels[branch_index],
+ with_cp=self.with_cp,
+ norm_cfg=self.norm_cfg,
+ conv_cfg=self.conv_cfg,
+ init_cfg=self.block_init_cfg))
+
+ return Sequential(*layers)
+
+ def _make_branches(self, num_branches, block, num_blocks, num_channels):
+ branches = []
+
+ for i in range(num_branches):
+ branches.append(
+ self._make_one_branch(i, block, num_blocks, num_channels))
+
+ return ModuleList(branches)
+
+ def _make_fuse_layers(self):
+ if self.num_branches == 1:
+ return None
+
+ num_branches = self.num_branches
+ in_channels = self.in_channels
+ fuse_layers = []
+ num_out_branches = num_branches if self.multiscale_output else 1
+ for i in range(num_out_branches):
+ fuse_layer = []
+ for j in range(num_branches):
+ if j > i:
+ fuse_layer.append(
+ nn.Sequential(
+ build_conv_layer(
+ self.conv_cfg,
+ in_channels[j],
+ in_channels[i],
+ kernel_size=1,
+ stride=1,
+ padding=0,
+ bias=False),
+ build_norm_layer(self.norm_cfg, in_channels[i])[1],
+ nn.Upsample(
+ scale_factor=2**(j - i), mode='nearest')))
+ elif j == i:
+ fuse_layer.append(None)
+ else:
+ conv_downsamples = []
+ for k in range(i - j):
+ if k == i - j - 1:
+ conv_downsamples.append(
+ nn.Sequential(
+ build_conv_layer(
+ self.conv_cfg,
+ in_channels[j],
+ in_channels[i],
+ kernel_size=3,
+ stride=2,
+ padding=1,
+ bias=False),
+ build_norm_layer(self.norm_cfg,
+ in_channels[i])[1]))
+ else:
+ conv_downsamples.append(
+ nn.Sequential(
+ build_conv_layer(
+ self.conv_cfg,
+ in_channels[j],
+ in_channels[j],
+ kernel_size=3,
+ stride=2,
+ padding=1,
+ bias=False),
+ build_norm_layer(self.norm_cfg,
+ in_channels[j])[1],
+ nn.ReLU(inplace=False)))
+ fuse_layer.append(nn.Sequential(*conv_downsamples))
+ fuse_layers.append(nn.ModuleList(fuse_layer))
+
+ return nn.ModuleList(fuse_layers)
+
+ def forward(self, x):
+ """Forward function."""
+ if self.num_branches == 1:
+ return [self.branches[0](x[0])]
+
+ for i in range(self.num_branches):
+ x[i] = self.branches[i](x[i])
+
+ x_fuse = []
+ for i in range(len(self.fuse_layers)):
+ y = 0
+ for j in range(self.num_branches):
+ if i == j:
+ y += x[j]
+ else:
+ y += self.fuse_layers[i][j](x[j])
+ x_fuse.append(self.relu(y))
+ return x_fuse
+
+
+@MODELS.register_module()
+class HRNet(BaseModule):
+ """HRNet backbone.
+
+ `High-Resolution Representations for Labeling Pixels and Regions
+ arXiv: `_.
+
+ Args:
+ extra (dict): Detailed configuration for each stage of HRNet.
+ There must be 4 stages, the configuration for each stage must have
+ 5 keys:
+
+ - num_modules(int): The number of HRModule in this stage.
+ - num_branches(int): The number of branches in the HRModule.
+ - block(str): The type of convolution block.
+ - num_blocks(tuple): The number of blocks in each branch.
+ The length must be equal to num_branches.
+ - num_channels(tuple): The number of channels in each branch.
+ The length must be equal to num_branches.
+ in_channels (int): Number of input image channels. Default: 3.
+ conv_cfg (dict): Dictionary to construct and config conv layer.
+ norm_cfg (dict): Dictionary to construct and config norm layer.
+ norm_eval (bool): Whether to set norm layers to eval mode, namely,
+ freeze running stats (mean and var). Note: Effect on Batch Norm
+ and its variants only. Default: True.
+ with_cp (bool): Use checkpoint or not. Using checkpoint will save some
+ memory while slowing down the training speed. Default: False.
+ zero_init_residual (bool): Whether to use zero init for last norm layer
+ in resblocks to let them behave as identity. Default: False.
+ multiscale_output (bool): Whether to output multi-level features
+ produced by multiple branches. If False, only the first level
+ feature will be output. Default: True.
+ pretrained (str, optional): Model pretrained path. Default: None.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None.
+
+ Example:
+ >>> from mmdet.models import HRNet
+ >>> import torch
+ >>> extra = dict(
+ >>> stage1=dict(
+ >>> num_modules=1,
+ >>> num_branches=1,
+ >>> block='BOTTLENECK',
+ >>> num_blocks=(4, ),
+ >>> num_channels=(64, )),
+ >>> stage2=dict(
+ >>> num_modules=1,
+ >>> num_branches=2,
+ >>> block='BASIC',
+ >>> num_blocks=(4, 4),
+ >>> num_channels=(32, 64)),
+ >>> stage3=dict(
+ >>> num_modules=4,
+ >>> num_branches=3,
+ >>> block='BASIC',
+ >>> num_blocks=(4, 4, 4),
+ >>> num_channels=(32, 64, 128)),
+ >>> stage4=dict(
+ >>> num_modules=3,
+ >>> num_branches=4,
+ >>> block='BASIC',
+ >>> num_blocks=(4, 4, 4, 4),
+ >>> num_channels=(32, 64, 128, 256)))
+ >>> self = HRNet(extra, in_channels=1)
+ >>> self.eval()
+ >>> inputs = torch.rand(1, 1, 32, 32)
+ >>> level_outputs = self.forward(inputs)
+ >>> for level_out in level_outputs:
+ ... print(tuple(level_out.shape))
+ (1, 32, 8, 8)
+ (1, 64, 4, 4)
+ (1, 128, 2, 2)
+ (1, 256, 1, 1)
+ """
+
+ blocks_dict = {'BASIC': BasicBlock, 'BOTTLENECK': Bottleneck}
+
+ def __init__(self,
+ extra,
+ in_channels=3,
+ conv_cfg=None,
+ norm_cfg=dict(type='BN'),
+ norm_eval=True,
+ with_cp=False,
+ zero_init_residual=False,
+ multiscale_output=True,
+ pretrained=None,
+ init_cfg=None):
+ super(HRNet, self).__init__(init_cfg)
+
+ self.pretrained = pretrained
+ assert not (init_cfg and pretrained), \
+ 'init_cfg and pretrained cannot be specified at the same time'
+ if isinstance(pretrained, str):
+ warnings.warn('DeprecationWarning: pretrained is deprecated, '
+ 'please use "init_cfg" instead')
+ self.init_cfg = dict(type='Pretrained', checkpoint=pretrained)
+ elif pretrained is None:
+ if init_cfg is None:
+ self.init_cfg = [
+ dict(type='Kaiming', layer='Conv2d'),
+ dict(
+ type='Constant',
+ val=1,
+ layer=['_BatchNorm', 'GroupNorm'])
+ ]
+ else:
+ raise TypeError('pretrained must be a str or None')
+
+ # Assert configurations of 4 stages are in extra
+ assert 'stage1' in extra and 'stage2' in extra \
+ and 'stage3' in extra and 'stage4' in extra
+ # Assert whether the length of `num_blocks` and `num_channels` are
+ # equal to `num_branches`
+ for i in range(4):
+ cfg = extra[f'stage{i + 1}']
+ assert len(cfg['num_blocks']) == cfg['num_branches'] and \
+ len(cfg['num_channels']) == cfg['num_branches']
+
+ self.extra = extra
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self.norm_eval = norm_eval
+ self.with_cp = with_cp
+ self.zero_init_residual = zero_init_residual
+
+ # stem net
+ self.norm1_name, norm1 = build_norm_layer(self.norm_cfg, 64, postfix=1)
+ self.norm2_name, norm2 = build_norm_layer(self.norm_cfg, 64, postfix=2)
+
+ self.conv1 = build_conv_layer(
+ self.conv_cfg,
+ in_channels,
+ 64,
+ kernel_size=3,
+ stride=2,
+ padding=1,
+ bias=False)
+
+ self.add_module(self.norm1_name, norm1)
+ self.conv2 = build_conv_layer(
+ self.conv_cfg,
+ 64,
+ 64,
+ kernel_size=3,
+ stride=2,
+ padding=1,
+ bias=False)
+
+ self.add_module(self.norm2_name, norm2)
+ self.relu = nn.ReLU(inplace=True)
+
+ # stage 1
+ self.stage1_cfg = self.extra['stage1']
+ num_channels = self.stage1_cfg['num_channels'][0]
+ block_type = self.stage1_cfg['block']
+ num_blocks = self.stage1_cfg['num_blocks'][0]
+
+ block = self.blocks_dict[block_type]
+ stage1_out_channels = num_channels * block.expansion
+ self.layer1 = self._make_layer(block, 64, num_channels, num_blocks)
+
+ # stage 2
+ self.stage2_cfg = self.extra['stage2']
+ num_channels = self.stage2_cfg['num_channels']
+ block_type = self.stage2_cfg['block']
+
+ block = self.blocks_dict[block_type]
+ num_channels = [channel * block.expansion for channel in num_channels]
+ self.transition1 = self._make_transition_layer([stage1_out_channels],
+ num_channels)
+ self.stage2, pre_stage_channels = self._make_stage(
+ self.stage2_cfg, num_channels)
+
+ # stage 3
+ self.stage3_cfg = self.extra['stage3']
+ num_channels = self.stage3_cfg['num_channels']
+ block_type = self.stage3_cfg['block']
+
+ block = self.blocks_dict[block_type]
+ num_channels = [channel * block.expansion for channel in num_channels]
+ self.transition2 = self._make_transition_layer(pre_stage_channels,
+ num_channels)
+ self.stage3, pre_stage_channels = self._make_stage(
+ self.stage3_cfg, num_channels)
+
+ # stage 4
+ self.stage4_cfg = self.extra['stage4']
+ num_channels = self.stage4_cfg['num_channels']
+ block_type = self.stage4_cfg['block']
+
+ block = self.blocks_dict[block_type]
+ num_channels = [channel * block.expansion for channel in num_channels]
+ self.transition3 = self._make_transition_layer(pre_stage_channels,
+ num_channels)
+ self.stage4, pre_stage_channels = self._make_stage(
+ self.stage4_cfg, num_channels, multiscale_output=multiscale_output)
+
+ @property
+ def norm1(self):
+ """nn.Module: the normalization layer named "norm1" """
+ return getattr(self, self.norm1_name)
+
+ @property
+ def norm2(self):
+ """nn.Module: the normalization layer named "norm2" """
+ return getattr(self, self.norm2_name)
+
+ def _make_transition_layer(self, num_channels_pre_layer,
+ num_channels_cur_layer):
+ num_branches_cur = len(num_channels_cur_layer)
+ num_branches_pre = len(num_channels_pre_layer)
+
+ transition_layers = []
+ for i in range(num_branches_cur):
+ if i < num_branches_pre:
+ if num_channels_cur_layer[i] != num_channels_pre_layer[i]:
+ transition_layers.append(
+ nn.Sequential(
+ build_conv_layer(
+ self.conv_cfg,
+ num_channels_pre_layer[i],
+ num_channels_cur_layer[i],
+ kernel_size=3,
+ stride=1,
+ padding=1,
+ bias=False),
+ build_norm_layer(self.norm_cfg,
+ num_channels_cur_layer[i])[1],
+ nn.ReLU(inplace=True)))
+ else:
+ transition_layers.append(None)
+ else:
+ conv_downsamples = []
+ for j in range(i + 1 - num_branches_pre):
+ in_channels = num_channels_pre_layer[-1]
+ out_channels = num_channels_cur_layer[i] \
+ if j == i - num_branches_pre else in_channels
+ conv_downsamples.append(
+ nn.Sequential(
+ build_conv_layer(
+ self.conv_cfg,
+ in_channels,
+ out_channels,
+ kernel_size=3,
+ stride=2,
+ padding=1,
+ bias=False),
+ build_norm_layer(self.norm_cfg, out_channels)[1],
+ nn.ReLU(inplace=True)))
+ transition_layers.append(nn.Sequential(*conv_downsamples))
+
+ return nn.ModuleList(transition_layers)
+
+ def _make_layer(self, block, inplanes, planes, blocks, stride=1):
+ downsample = None
+ if stride != 1 or inplanes != planes * block.expansion:
+ downsample = nn.Sequential(
+ build_conv_layer(
+ self.conv_cfg,
+ inplanes,
+ planes * block.expansion,
+ kernel_size=1,
+ stride=stride,
+ bias=False),
+ build_norm_layer(self.norm_cfg, planes * block.expansion)[1])
+
+ layers = []
+ block_init_cfg = None
+ if self.pretrained is None and not hasattr(
+ self, 'init_cfg') and self.zero_init_residual:
+ if block is BasicBlock:
+ block_init_cfg = dict(
+ type='Constant', val=0, override=dict(name='norm2'))
+ elif block is Bottleneck:
+ block_init_cfg = dict(
+ type='Constant', val=0, override=dict(name='norm3'))
+ layers.append(
+ block(
+ inplanes,
+ planes,
+ stride,
+ downsample=downsample,
+ with_cp=self.with_cp,
+ norm_cfg=self.norm_cfg,
+ conv_cfg=self.conv_cfg,
+ init_cfg=block_init_cfg,
+ ))
+ inplanes = planes * block.expansion
+ for i in range(1, blocks):
+ layers.append(
+ block(
+ inplanes,
+ planes,
+ with_cp=self.with_cp,
+ norm_cfg=self.norm_cfg,
+ conv_cfg=self.conv_cfg,
+ init_cfg=block_init_cfg))
+
+ return Sequential(*layers)
+
+ def _make_stage(self, layer_config, in_channels, multiscale_output=True):
+ num_modules = layer_config['num_modules']
+ num_branches = layer_config['num_branches']
+ num_blocks = layer_config['num_blocks']
+ num_channels = layer_config['num_channels']
+ block = self.blocks_dict[layer_config['block']]
+
+ hr_modules = []
+ block_init_cfg = None
+ if self.pretrained is None and not hasattr(
+ self, 'init_cfg') and self.zero_init_residual:
+ if block is BasicBlock:
+ block_init_cfg = dict(
+ type='Constant', val=0, override=dict(name='norm2'))
+ elif block is Bottleneck:
+ block_init_cfg = dict(
+ type='Constant', val=0, override=dict(name='norm3'))
+
+ for i in range(num_modules):
+ # multi_scale_output is only used for the last module
+ if not multiscale_output and i == num_modules - 1:
+ reset_multiscale_output = False
+ else:
+ reset_multiscale_output = True
+
+ hr_modules.append(
+ HRModule(
+ num_branches,
+ block,
+ num_blocks,
+ in_channels,
+ num_channels,
+ reset_multiscale_output,
+ with_cp=self.with_cp,
+ norm_cfg=self.norm_cfg,
+ conv_cfg=self.conv_cfg,
+ block_init_cfg=block_init_cfg))
+
+ return Sequential(*hr_modules), in_channels
+
+ def forward(self, x):
+ """Forward function."""
+ x = self.conv1(x)
+ x = self.norm1(x)
+ x = self.relu(x)
+ x = self.conv2(x)
+ x = self.norm2(x)
+ x = self.relu(x)
+ x = self.layer1(x)
+
+ x_list = []
+ for i in range(self.stage2_cfg['num_branches']):
+ if self.transition1[i] is not None:
+ x_list.append(self.transition1[i](x))
+ else:
+ x_list.append(x)
+ y_list = self.stage2(x_list)
+
+ x_list = []
+ for i in range(self.stage3_cfg['num_branches']):
+ if self.transition2[i] is not None:
+ x_list.append(self.transition2[i](y_list[-1]))
+ else:
+ x_list.append(y_list[i])
+ y_list = self.stage3(x_list)
+
+ x_list = []
+ for i in range(self.stage4_cfg['num_branches']):
+ if self.transition3[i] is not None:
+ x_list.append(self.transition3[i](y_list[-1]))
+ else:
+ x_list.append(y_list[i])
+ y_list = self.stage4(x_list)
+
+ return y_list
+
+ def train(self, mode=True):
+ """Convert the model into training mode will keeping the normalization
+ layer freezed."""
+ super(HRNet, self).train(mode)
+ if mode and self.norm_eval:
+ for m in self.modules():
+ # trick: eval have effect on BatchNorm only
+ if isinstance(m, _BatchNorm):
+ m.eval()
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/mobilenet_v2.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/mobilenet_v2.py
new file mode 100644
index 0000000000000000000000000000000000000000..a4fd0519ad4d5106e1acb82624d6393052596ce8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/mobilenet_v2.py
@@ -0,0 +1,198 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from mmengine.model import BaseModule
+from torch.nn.modules.batchnorm import _BatchNorm
+
+from mmdet.registry import MODELS
+from ..layers import InvertedResidual
+from ..utils import make_divisible
+
+
+@MODELS.register_module()
+class MobileNetV2(BaseModule):
+ """MobileNetV2 backbone.
+
+ Args:
+ widen_factor (float): Width multiplier, multiply number of
+ channels in each layer by this amount. Default: 1.0.
+ out_indices (Sequence[int], optional): Output from which stages.
+ Default: (1, 2, 4, 7).
+ frozen_stages (int): Stages to be frozen (all param fixed).
+ Default: -1, which means not freezing any parameters.
+ conv_cfg (dict, optional): Config dict for convolution layer.
+ Default: None, which means using conv2d.
+ norm_cfg (dict): Config dict for normalization layer.
+ Default: dict(type='BN').
+ act_cfg (dict): Config dict for activation layer.
+ Default: dict(type='ReLU6').
+ norm_eval (bool): Whether to set norm layers to eval mode, namely,
+ freeze running stats (mean and var). Note: Effect on Batch Norm
+ and its variants only. Default: False.
+ with_cp (bool): Use checkpoint or not. Using checkpoint will save some
+ memory while slowing down the training speed. Default: False.
+ pretrained (str, optional): model pretrained path. Default: None
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None
+ """
+
+ # Parameters to build layers. 4 parameters are needed to construct a
+ # layer, from left to right: expand_ratio, channel, num_blocks, stride.
+ arch_settings = [[1, 16, 1, 1], [6, 24, 2, 2], [6, 32, 3, 2],
+ [6, 64, 4, 2], [6, 96, 3, 1], [6, 160, 3, 2],
+ [6, 320, 1, 1]]
+
+ def __init__(self,
+ widen_factor=1.,
+ out_indices=(1, 2, 4, 7),
+ frozen_stages=-1,
+ conv_cfg=None,
+ norm_cfg=dict(type='BN'),
+ act_cfg=dict(type='ReLU6'),
+ norm_eval=False,
+ with_cp=False,
+ pretrained=None,
+ init_cfg=None):
+ super(MobileNetV2, self).__init__(init_cfg)
+
+ self.pretrained = pretrained
+ assert not (init_cfg and pretrained), \
+ 'init_cfg and pretrained cannot be specified at the same time'
+ if isinstance(pretrained, str):
+ warnings.warn('DeprecationWarning: pretrained is deprecated, '
+ 'please use "init_cfg" instead')
+ self.init_cfg = dict(type='Pretrained', checkpoint=pretrained)
+ elif pretrained is None:
+ if init_cfg is None:
+ self.init_cfg = [
+ dict(type='Kaiming', layer='Conv2d'),
+ dict(
+ type='Constant',
+ val=1,
+ layer=['_BatchNorm', 'GroupNorm'])
+ ]
+ else:
+ raise TypeError('pretrained must be a str or None')
+
+ self.widen_factor = widen_factor
+ self.out_indices = out_indices
+ if not set(out_indices).issubset(set(range(0, 8))):
+ raise ValueError('out_indices must be a subset of range'
+ f'(0, 8). But received {out_indices}')
+
+ if frozen_stages not in range(-1, 8):
+ raise ValueError('frozen_stages must be in range(-1, 8). '
+ f'But received {frozen_stages}')
+ self.out_indices = out_indices
+ self.frozen_stages = frozen_stages
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self.act_cfg = act_cfg
+ self.norm_eval = norm_eval
+ self.with_cp = with_cp
+
+ self.in_channels = make_divisible(32 * widen_factor, 8)
+
+ self.conv1 = ConvModule(
+ in_channels=3,
+ out_channels=self.in_channels,
+ kernel_size=3,
+ stride=2,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg)
+
+ self.layers = []
+
+ for i, layer_cfg in enumerate(self.arch_settings):
+ expand_ratio, channel, num_blocks, stride = layer_cfg
+ out_channels = make_divisible(channel * widen_factor, 8)
+ inverted_res_layer = self.make_layer(
+ out_channels=out_channels,
+ num_blocks=num_blocks,
+ stride=stride,
+ expand_ratio=expand_ratio)
+ layer_name = f'layer{i + 1}'
+ self.add_module(layer_name, inverted_res_layer)
+ self.layers.append(layer_name)
+
+ if widen_factor > 1.0:
+ self.out_channel = int(1280 * widen_factor)
+ else:
+ self.out_channel = 1280
+
+ layer = ConvModule(
+ in_channels=self.in_channels,
+ out_channels=self.out_channel,
+ kernel_size=1,
+ stride=1,
+ padding=0,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg)
+ self.add_module('conv2', layer)
+ self.layers.append('conv2')
+
+ def make_layer(self, out_channels, num_blocks, stride, expand_ratio):
+ """Stack InvertedResidual blocks to build a layer for MobileNetV2.
+
+ Args:
+ out_channels (int): out_channels of block.
+ num_blocks (int): number of blocks.
+ stride (int): stride of the first block. Default: 1
+ expand_ratio (int): Expand the number of channels of the
+ hidden layer in InvertedResidual by this ratio. Default: 6.
+ """
+ layers = []
+ for i in range(num_blocks):
+ if i >= 1:
+ stride = 1
+ layers.append(
+ InvertedResidual(
+ self.in_channels,
+ out_channels,
+ mid_channels=int(round(self.in_channels * expand_ratio)),
+ stride=stride,
+ with_expand_conv=expand_ratio != 1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg,
+ with_cp=self.with_cp))
+ self.in_channels = out_channels
+
+ return nn.Sequential(*layers)
+
+ def _freeze_stages(self):
+ if self.frozen_stages >= 0:
+ for param in self.conv1.parameters():
+ param.requires_grad = False
+ for i in range(1, self.frozen_stages + 1):
+ layer = getattr(self, f'layer{i}')
+ layer.eval()
+ for param in layer.parameters():
+ param.requires_grad = False
+
+ def forward(self, x):
+ """Forward function."""
+ x = self.conv1(x)
+ outs = []
+ for i, layer_name in enumerate(self.layers):
+ layer = getattr(self, layer_name)
+ x = layer(x)
+ if i in self.out_indices:
+ outs.append(x)
+ return tuple(outs)
+
+ def train(self, mode=True):
+ """Convert the model into training mode while keep normalization layer
+ frozen."""
+ super(MobileNetV2, self).train(mode)
+ self._freeze_stages()
+ if mode and self.norm_eval:
+ for m in self.modules():
+ # trick: eval have effect on BatchNorm only
+ if isinstance(m, _BatchNorm):
+ m.eval()
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/pvt.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/pvt.py
new file mode 100644
index 0000000000000000000000000000000000000000..8b250f63c1b22f21a892faf4c41ccc2d20e83e13
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/pvt.py
@@ -0,0 +1,665 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+import warnings
+from collections import OrderedDict
+
+import numpy as np
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import Conv2d, build_activation_layer, build_norm_layer
+from mmcv.cnn.bricks.drop import build_dropout
+from mmcv.cnn.bricks.transformer import MultiheadAttention
+from mmengine.logging import MMLogger
+from mmengine.model import (BaseModule, ModuleList, Sequential, constant_init,
+ normal_init, trunc_normal_init)
+from mmengine.model.weight_init import trunc_normal_
+from mmengine.runner.checkpoint import CheckpointLoader, load_state_dict
+from torch.nn.modules.utils import _pair as to_2tuple
+
+from mmdet.registry import MODELS
+from ..layers import PatchEmbed, nchw_to_nlc, nlc_to_nchw
+
+
+class MixFFN(BaseModule):
+ """An implementation of MixFFN of PVT.
+
+ The differences between MixFFN & FFN:
+ 1. Use 1X1 Conv to replace Linear layer.
+ 2. Introduce 3X3 Depth-wise Conv to encode positional information.
+
+ Args:
+ embed_dims (int): The feature dimension. Same as
+ `MultiheadAttention`.
+ feedforward_channels (int): The hidden dimension of FFNs.
+ act_cfg (dict, optional): The activation config for FFNs.
+ Default: dict(type='GELU').
+ ffn_drop (float, optional): Probability of an element to be
+ zeroed in FFN. Default 0.0.
+ dropout_layer (obj:`ConfigDict`): The dropout_layer used
+ when adding the shortcut.
+ Default: None.
+ use_conv (bool): If True, add 3x3 DWConv between two Linear layers.
+ Defaults: False.
+ init_cfg (obj:`mmengine.ConfigDict`): The Config for initialization.
+ Default: None.
+ """
+
+ def __init__(self,
+ embed_dims,
+ feedforward_channels,
+ act_cfg=dict(type='GELU'),
+ ffn_drop=0.,
+ dropout_layer=None,
+ use_conv=False,
+ init_cfg=None):
+ super(MixFFN, self).__init__(init_cfg=init_cfg)
+
+ self.embed_dims = embed_dims
+ self.feedforward_channels = feedforward_channels
+ self.act_cfg = act_cfg
+ activate = build_activation_layer(act_cfg)
+
+ in_channels = embed_dims
+ fc1 = Conv2d(
+ in_channels=in_channels,
+ out_channels=feedforward_channels,
+ kernel_size=1,
+ stride=1,
+ bias=True)
+ if use_conv:
+ # 3x3 depth wise conv to provide positional encode information
+ dw_conv = Conv2d(
+ in_channels=feedforward_channels,
+ out_channels=feedforward_channels,
+ kernel_size=3,
+ stride=1,
+ padding=(3 - 1) // 2,
+ bias=True,
+ groups=feedforward_channels)
+ fc2 = Conv2d(
+ in_channels=feedforward_channels,
+ out_channels=in_channels,
+ kernel_size=1,
+ stride=1,
+ bias=True)
+ drop = nn.Dropout(ffn_drop)
+ layers = [fc1, activate, drop, fc2, drop]
+ if use_conv:
+ layers.insert(1, dw_conv)
+ self.layers = Sequential(*layers)
+ self.dropout_layer = build_dropout(
+ dropout_layer) if dropout_layer else torch.nn.Identity()
+
+ def forward(self, x, hw_shape, identity=None):
+ out = nlc_to_nchw(x, hw_shape)
+ out = self.layers(out)
+ out = nchw_to_nlc(out)
+ if identity is None:
+ identity = x
+ return identity + self.dropout_layer(out)
+
+
+class SpatialReductionAttention(MultiheadAttention):
+ """An implementation of Spatial Reduction Attention of PVT.
+
+ This module is modified from MultiheadAttention which is a module from
+ mmcv.cnn.bricks.transformer.
+
+ Args:
+ embed_dims (int): The embedding dimension.
+ num_heads (int): Parallel attention heads.
+ attn_drop (float): A Dropout layer on attn_output_weights.
+ Default: 0.0.
+ proj_drop (float): A Dropout layer after `nn.MultiheadAttention`.
+ Default: 0.0.
+ dropout_layer (obj:`ConfigDict`): The dropout_layer used
+ when adding the shortcut. Default: None.
+ batch_first (bool): Key, Query and Value are shape of
+ (batch, n, embed_dim)
+ or (n, batch, embed_dim). Default: False.
+ qkv_bias (bool): enable bias for qkv if True. Default: True.
+ norm_cfg (dict): Config dict for normalization layer.
+ Default: dict(type='LN').
+ sr_ratio (int): The ratio of spatial reduction of Spatial Reduction
+ Attention of PVT. Default: 1.
+ init_cfg (obj:`mmengine.ConfigDict`): The Config for initialization.
+ Default: None.
+ """
+
+ def __init__(self,
+ embed_dims,
+ num_heads,
+ attn_drop=0.,
+ proj_drop=0.,
+ dropout_layer=None,
+ batch_first=True,
+ qkv_bias=True,
+ norm_cfg=dict(type='LN'),
+ sr_ratio=1,
+ init_cfg=None):
+ super().__init__(
+ embed_dims,
+ num_heads,
+ attn_drop,
+ proj_drop,
+ batch_first=batch_first,
+ dropout_layer=dropout_layer,
+ bias=qkv_bias,
+ init_cfg=init_cfg)
+
+ self.sr_ratio = sr_ratio
+ if sr_ratio > 1:
+ self.sr = Conv2d(
+ in_channels=embed_dims,
+ out_channels=embed_dims,
+ kernel_size=sr_ratio,
+ stride=sr_ratio)
+ # The ret[0] of build_norm_layer is norm name.
+ self.norm = build_norm_layer(norm_cfg, embed_dims)[1]
+
+ # handle the BC-breaking from https://github.com/open-mmlab/mmcv/pull/1418 # noqa
+ from mmdet import digit_version, mmcv_version
+ if mmcv_version < digit_version('1.3.17'):
+ warnings.warn('The legacy version of forward function in'
+ 'SpatialReductionAttention is deprecated in'
+ 'mmcv>=1.3.17 and will no longer support in the'
+ 'future. Please upgrade your mmcv.')
+ self.forward = self.legacy_forward
+
+ def forward(self, x, hw_shape, identity=None):
+
+ x_q = x
+ if self.sr_ratio > 1:
+ x_kv = nlc_to_nchw(x, hw_shape)
+ x_kv = self.sr(x_kv)
+ x_kv = nchw_to_nlc(x_kv)
+ x_kv = self.norm(x_kv)
+ else:
+ x_kv = x
+
+ if identity is None:
+ identity = x_q
+
+ # Because the dataflow('key', 'query', 'value') of
+ # ``torch.nn.MultiheadAttention`` is (num_queries, batch,
+ # embed_dims), We should adjust the shape of dataflow from
+ # batch_first (batch, num_queries, embed_dims) to num_queries_first
+ # (num_queries ,batch, embed_dims), and recover ``attn_output``
+ # from num_queries_first to batch_first.
+ if self.batch_first:
+ x_q = x_q.transpose(0, 1)
+ x_kv = x_kv.transpose(0, 1)
+
+ out = self.attn(query=x_q, key=x_kv, value=x_kv)[0]
+
+ if self.batch_first:
+ out = out.transpose(0, 1)
+
+ return identity + self.dropout_layer(self.proj_drop(out))
+
+ def legacy_forward(self, x, hw_shape, identity=None):
+ """multi head attention forward in mmcv version < 1.3.17."""
+ x_q = x
+ if self.sr_ratio > 1:
+ x_kv = nlc_to_nchw(x, hw_shape)
+ x_kv = self.sr(x_kv)
+ x_kv = nchw_to_nlc(x_kv)
+ x_kv = self.norm(x_kv)
+ else:
+ x_kv = x
+
+ if identity is None:
+ identity = x_q
+
+ out = self.attn(query=x_q, key=x_kv, value=x_kv)[0]
+
+ return identity + self.dropout_layer(self.proj_drop(out))
+
+
+class PVTEncoderLayer(BaseModule):
+ """Implements one encoder layer in PVT.
+
+ Args:
+ embed_dims (int): The feature dimension.
+ num_heads (int): Parallel attention heads.
+ feedforward_channels (int): The hidden dimension for FFNs.
+ drop_rate (float): Probability of an element to be zeroed.
+ after the feed forward layer. Default: 0.0.
+ attn_drop_rate (float): The drop out rate for attention layer.
+ Default: 0.0.
+ drop_path_rate (float): stochastic depth rate. Default: 0.0.
+ qkv_bias (bool): enable bias for qkv if True.
+ Default: True.
+ act_cfg (dict): The activation config for FFNs.
+ Default: dict(type='GELU').
+ norm_cfg (dict): Config dict for normalization layer.
+ Default: dict(type='LN').
+ sr_ratio (int): The ratio of spatial reduction of Spatial Reduction
+ Attention of PVT. Default: 1.
+ use_conv_ffn (bool): If True, use Convolutional FFN to replace FFN.
+ Default: False.
+ init_cfg (dict, optional): Initialization config dict.
+ Default: None.
+ """
+
+ def __init__(self,
+ embed_dims,
+ num_heads,
+ feedforward_channels,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.,
+ qkv_bias=True,
+ act_cfg=dict(type='GELU'),
+ norm_cfg=dict(type='LN'),
+ sr_ratio=1,
+ use_conv_ffn=False,
+ init_cfg=None):
+ super(PVTEncoderLayer, self).__init__(init_cfg=init_cfg)
+
+ # The ret[0] of build_norm_layer is norm name.
+ self.norm1 = build_norm_layer(norm_cfg, embed_dims)[1]
+
+ self.attn = SpatialReductionAttention(
+ embed_dims=embed_dims,
+ num_heads=num_heads,
+ attn_drop=attn_drop_rate,
+ proj_drop=drop_rate,
+ dropout_layer=dict(type='DropPath', drop_prob=drop_path_rate),
+ qkv_bias=qkv_bias,
+ norm_cfg=norm_cfg,
+ sr_ratio=sr_ratio)
+
+ # The ret[0] of build_norm_layer is norm name.
+ self.norm2 = build_norm_layer(norm_cfg, embed_dims)[1]
+
+ self.ffn = MixFFN(
+ embed_dims=embed_dims,
+ feedforward_channels=feedforward_channels,
+ ffn_drop=drop_rate,
+ dropout_layer=dict(type='DropPath', drop_prob=drop_path_rate),
+ use_conv=use_conv_ffn,
+ act_cfg=act_cfg)
+
+ def forward(self, x, hw_shape):
+ x = self.attn(self.norm1(x), hw_shape, identity=x)
+ x = self.ffn(self.norm2(x), hw_shape, identity=x)
+
+ return x
+
+
+class AbsolutePositionEmbedding(BaseModule):
+ """An implementation of the absolute position embedding in PVT.
+
+ Args:
+ pos_shape (int): The shape of the absolute position embedding.
+ pos_dim (int): The dimension of the absolute position embedding.
+ drop_rate (float): Probability of an element to be zeroed.
+ Default: 0.0.
+ """
+
+ def __init__(self, pos_shape, pos_dim, drop_rate=0., init_cfg=None):
+ super().__init__(init_cfg=init_cfg)
+
+ if isinstance(pos_shape, int):
+ pos_shape = to_2tuple(pos_shape)
+ elif isinstance(pos_shape, tuple):
+ if len(pos_shape) == 1:
+ pos_shape = to_2tuple(pos_shape[0])
+ assert len(pos_shape) == 2, \
+ f'The size of image should have length 1 or 2, ' \
+ f'but got {len(pos_shape)}'
+ self.pos_shape = pos_shape
+ self.pos_dim = pos_dim
+
+ self.pos_embed = nn.Parameter(
+ torch.zeros(1, pos_shape[0] * pos_shape[1], pos_dim))
+ self.drop = nn.Dropout(p=drop_rate)
+
+ def init_weights(self):
+ trunc_normal_(self.pos_embed, std=0.02)
+
+ def resize_pos_embed(self, pos_embed, input_shape, mode='bilinear'):
+ """Resize pos_embed weights.
+
+ Resize pos_embed using bilinear interpolate method.
+
+ Args:
+ pos_embed (torch.Tensor): Position embedding weights.
+ input_shape (tuple): Tuple for (downsampled input image height,
+ downsampled input image width).
+ mode (str): Algorithm used for upsampling:
+ ``'nearest'`` | ``'linear'`` | ``'bilinear'`` | ``'bicubic'`` |
+ ``'trilinear'``. Default: ``'bilinear'``.
+
+ Return:
+ torch.Tensor: The resized pos_embed of shape [B, L_new, C].
+ """
+ assert pos_embed.ndim == 3, 'shape of pos_embed must be [B, L, C]'
+ pos_h, pos_w = self.pos_shape
+ pos_embed_weight = pos_embed[:, (-1 * pos_h * pos_w):]
+ pos_embed_weight = pos_embed_weight.reshape(
+ 1, pos_h, pos_w, self.pos_dim).permute(0, 3, 1, 2).contiguous()
+ pos_embed_weight = F.interpolate(
+ pos_embed_weight, size=input_shape, mode=mode)
+ pos_embed_weight = torch.flatten(pos_embed_weight,
+ 2).transpose(1, 2).contiguous()
+ pos_embed = pos_embed_weight
+
+ return pos_embed
+
+ def forward(self, x, hw_shape, mode='bilinear'):
+ pos_embed = self.resize_pos_embed(self.pos_embed, hw_shape, mode)
+ return self.drop(x + pos_embed)
+
+
+@MODELS.register_module()
+class PyramidVisionTransformer(BaseModule):
+ """Pyramid Vision Transformer (PVT)
+
+ Implementation of `Pyramid Vision Transformer: A Versatile Backbone for
+ Dense Prediction without Convolutions
+ `_.
+
+ Args:
+ pretrain_img_size (int | tuple[int]): The size of input image when
+ pretrain. Defaults: 224.
+ in_channels (int): Number of input channels. Default: 3.
+ embed_dims (int): Embedding dimension. Default: 64.
+ num_stags (int): The num of stages. Default: 4.
+ num_layers (Sequence[int]): The layer number of each transformer encode
+ layer. Default: [3, 4, 6, 3].
+ num_heads (Sequence[int]): The attention heads of each transformer
+ encode layer. Default: [1, 2, 5, 8].
+ patch_sizes (Sequence[int]): The patch_size of each patch embedding.
+ Default: [4, 2, 2, 2].
+ strides (Sequence[int]): The stride of each patch embedding.
+ Default: [4, 2, 2, 2].
+ paddings (Sequence[int]): The padding of each patch embedding.
+ Default: [0, 0, 0, 0].
+ sr_ratios (Sequence[int]): The spatial reduction rate of each
+ transformer encode layer. Default: [8, 4, 2, 1].
+ out_indices (Sequence[int] | int): Output from which stages.
+ Default: (0, 1, 2, 3).
+ mlp_ratios (Sequence[int]): The ratio of the mlp hidden dim to the
+ embedding dim of each transformer encode layer.
+ Default: [8, 8, 4, 4].
+ qkv_bias (bool): Enable bias for qkv if True. Default: True.
+ drop_rate (float): Probability of an element to be zeroed.
+ Default 0.0.
+ attn_drop_rate (float): The drop out rate for attention layer.
+ Default 0.0.
+ drop_path_rate (float): stochastic depth rate. Default 0.1.
+ use_abs_pos_embed (bool): If True, add absolute position embedding to
+ the patch embedding. Defaults: True.
+ use_conv_ffn (bool): If True, use Convolutional FFN to replace FFN.
+ Default: False.
+ act_cfg (dict): The activation config for FFNs.
+ Default: dict(type='GELU').
+ norm_cfg (dict): Config dict for normalization layer.
+ Default: dict(type='LN').
+ pretrained (str, optional): model pretrained path. Default: None.
+ convert_weights (bool): The flag indicates whether the
+ pre-trained model is from the original repo. We may need
+ to convert some keys to make it compatible.
+ Default: True.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None.
+ """
+
+ def __init__(self,
+ pretrain_img_size=224,
+ in_channels=3,
+ embed_dims=64,
+ num_stages=4,
+ num_layers=[3, 4, 6, 3],
+ num_heads=[1, 2, 5, 8],
+ patch_sizes=[4, 2, 2, 2],
+ strides=[4, 2, 2, 2],
+ paddings=[0, 0, 0, 0],
+ sr_ratios=[8, 4, 2, 1],
+ out_indices=(0, 1, 2, 3),
+ mlp_ratios=[8, 8, 4, 4],
+ qkv_bias=True,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.1,
+ use_abs_pos_embed=True,
+ norm_after_stage=False,
+ use_conv_ffn=False,
+ act_cfg=dict(type='GELU'),
+ norm_cfg=dict(type='LN', eps=1e-6),
+ pretrained=None,
+ convert_weights=True,
+ init_cfg=None):
+ super().__init__(init_cfg=init_cfg)
+
+ self.convert_weights = convert_weights
+ if isinstance(pretrain_img_size, int):
+ pretrain_img_size = to_2tuple(pretrain_img_size)
+ elif isinstance(pretrain_img_size, tuple):
+ if len(pretrain_img_size) == 1:
+ pretrain_img_size = to_2tuple(pretrain_img_size[0])
+ assert len(pretrain_img_size) == 2, \
+ f'The size of image should have length 1 or 2, ' \
+ f'but got {len(pretrain_img_size)}'
+
+ assert not (init_cfg and pretrained), \
+ 'init_cfg and pretrained cannot be setting at the same time'
+ if isinstance(pretrained, str):
+ warnings.warn('DeprecationWarning: pretrained is deprecated, '
+ 'please use "init_cfg" instead')
+ self.init_cfg = dict(type='Pretrained', checkpoint=pretrained)
+ elif pretrained is None:
+ self.init_cfg = init_cfg
+ else:
+ raise TypeError('pretrained must be a str or None')
+
+ self.embed_dims = embed_dims
+
+ self.num_stages = num_stages
+ self.num_layers = num_layers
+ self.num_heads = num_heads
+ self.patch_sizes = patch_sizes
+ self.strides = strides
+ self.sr_ratios = sr_ratios
+ assert num_stages == len(num_layers) == len(num_heads) \
+ == len(patch_sizes) == len(strides) == len(sr_ratios)
+
+ self.out_indices = out_indices
+ assert max(out_indices) < self.num_stages
+ self.pretrained = pretrained
+
+ # transformer encoder
+ dpr = [
+ x.item()
+ for x in torch.linspace(0, drop_path_rate, sum(num_layers))
+ ] # stochastic num_layer decay rule
+
+ cur = 0
+ self.layers = ModuleList()
+ for i, num_layer in enumerate(num_layers):
+ embed_dims_i = embed_dims * num_heads[i]
+ patch_embed = PatchEmbed(
+ in_channels=in_channels,
+ embed_dims=embed_dims_i,
+ kernel_size=patch_sizes[i],
+ stride=strides[i],
+ padding=paddings[i],
+ bias=True,
+ norm_cfg=norm_cfg)
+
+ layers = ModuleList()
+ if use_abs_pos_embed:
+ pos_shape = pretrain_img_size // np.prod(patch_sizes[:i + 1])
+ pos_embed = AbsolutePositionEmbedding(
+ pos_shape=pos_shape,
+ pos_dim=embed_dims_i,
+ drop_rate=drop_rate)
+ layers.append(pos_embed)
+ layers.extend([
+ PVTEncoderLayer(
+ embed_dims=embed_dims_i,
+ num_heads=num_heads[i],
+ feedforward_channels=mlp_ratios[i] * embed_dims_i,
+ drop_rate=drop_rate,
+ attn_drop_rate=attn_drop_rate,
+ drop_path_rate=dpr[cur + idx],
+ qkv_bias=qkv_bias,
+ act_cfg=act_cfg,
+ norm_cfg=norm_cfg,
+ sr_ratio=sr_ratios[i],
+ use_conv_ffn=use_conv_ffn) for idx in range(num_layer)
+ ])
+ in_channels = embed_dims_i
+ # The ret[0] of build_norm_layer is norm name.
+ if norm_after_stage:
+ norm = build_norm_layer(norm_cfg, embed_dims_i)[1]
+ else:
+ norm = nn.Identity()
+ self.layers.append(ModuleList([patch_embed, layers, norm]))
+ cur += num_layer
+
+ def init_weights(self):
+ logger = MMLogger.get_current_instance()
+ if self.init_cfg is None:
+ logger.warn(f'No pre-trained weights for '
+ f'{self.__class__.__name__}, '
+ f'training start from scratch')
+ for m in self.modules():
+ if isinstance(m, nn.Linear):
+ trunc_normal_init(m, std=.02, bias=0.)
+ elif isinstance(m, nn.LayerNorm):
+ constant_init(m, 1.0)
+ elif isinstance(m, nn.Conv2d):
+ fan_out = m.kernel_size[0] * m.kernel_size[
+ 1] * m.out_channels
+ fan_out //= m.groups
+ normal_init(m, 0, math.sqrt(2.0 / fan_out))
+ elif isinstance(m, AbsolutePositionEmbedding):
+ m.init_weights()
+ else:
+ assert 'checkpoint' in self.init_cfg, f'Only support ' \
+ f'specify `Pretrained` in ' \
+ f'`init_cfg` in ' \
+ f'{self.__class__.__name__} '
+ checkpoint = CheckpointLoader.load_checkpoint(
+ self.init_cfg.checkpoint, logger=logger, map_location='cpu')
+ logger.warn(f'Load pre-trained model for '
+ f'{self.__class__.__name__} from original repo')
+ if 'state_dict' in checkpoint:
+ state_dict = checkpoint['state_dict']
+ elif 'model' in checkpoint:
+ state_dict = checkpoint['model']
+ else:
+ state_dict = checkpoint
+ if self.convert_weights:
+ # Because pvt backbones are not supported by mmpretrain,
+ # so we need to convert pre-trained weights to match this
+ # implementation.
+ state_dict = pvt_convert(state_dict)
+ load_state_dict(self, state_dict, strict=False, logger=logger)
+
+ def forward(self, x):
+ outs = []
+
+ for i, layer in enumerate(self.layers):
+ x, hw_shape = layer[0](x)
+
+ for block in layer[1]:
+ x = block(x, hw_shape)
+ x = layer[2](x)
+ x = nlc_to_nchw(x, hw_shape)
+ if i in self.out_indices:
+ outs.append(x)
+
+ return outs
+
+
+@MODELS.register_module()
+class PyramidVisionTransformerV2(PyramidVisionTransformer):
+ """Implementation of `PVTv2: Improved Baselines with Pyramid Vision
+ Transformer `_."""
+
+ def __init__(self, **kwargs):
+ super(PyramidVisionTransformerV2, self).__init__(
+ patch_sizes=[7, 3, 3, 3],
+ paddings=[3, 1, 1, 1],
+ use_abs_pos_embed=False,
+ norm_after_stage=True,
+ use_conv_ffn=True,
+ **kwargs)
+
+
+def pvt_convert(ckpt):
+ new_ckpt = OrderedDict()
+ # Process the concat between q linear weights and kv linear weights
+ use_abs_pos_embed = False
+ use_conv_ffn = False
+ for k in ckpt.keys():
+ if k.startswith('pos_embed'):
+ use_abs_pos_embed = True
+ if k.find('dwconv') >= 0:
+ use_conv_ffn = True
+ for k, v in ckpt.items():
+ if k.startswith('head'):
+ continue
+ if k.startswith('norm.'):
+ continue
+ if k.startswith('cls_token'):
+ continue
+ if k.startswith('pos_embed'):
+ stage_i = int(k.replace('pos_embed', ''))
+ new_k = k.replace(f'pos_embed{stage_i}',
+ f'layers.{stage_i - 1}.1.0.pos_embed')
+ if stage_i == 4 and v.size(1) == 50: # 1 (cls token) + 7 * 7
+ new_v = v[:, 1:, :] # remove cls token
+ else:
+ new_v = v
+ elif k.startswith('patch_embed'):
+ stage_i = int(k.split('.')[0].replace('patch_embed', ''))
+ new_k = k.replace(f'patch_embed{stage_i}',
+ f'layers.{stage_i - 1}.0')
+ new_v = v
+ if 'proj.' in new_k:
+ new_k = new_k.replace('proj.', 'projection.')
+ elif k.startswith('block'):
+ stage_i = int(k.split('.')[0].replace('block', ''))
+ layer_i = int(k.split('.')[1])
+ new_layer_i = layer_i + use_abs_pos_embed
+ new_k = k.replace(f'block{stage_i}.{layer_i}',
+ f'layers.{stage_i - 1}.1.{new_layer_i}')
+ new_v = v
+ if 'attn.q.' in new_k:
+ sub_item_k = k.replace('q.', 'kv.')
+ new_k = new_k.replace('q.', 'attn.in_proj_')
+ new_v = torch.cat([v, ckpt[sub_item_k]], dim=0)
+ elif 'attn.kv.' in new_k:
+ continue
+ elif 'attn.proj.' in new_k:
+ new_k = new_k.replace('proj.', 'attn.out_proj.')
+ elif 'attn.sr.' in new_k:
+ new_k = new_k.replace('sr.', 'sr.')
+ elif 'mlp.' in new_k:
+ string = f'{new_k}-'
+ new_k = new_k.replace('mlp.', 'ffn.layers.')
+ if 'fc1.weight' in new_k or 'fc2.weight' in new_k:
+ new_v = v.reshape((*v.shape, 1, 1))
+ new_k = new_k.replace('fc1.', '0.')
+ new_k = new_k.replace('dwconv.dwconv.', '1.')
+ if use_conv_ffn:
+ new_k = new_k.replace('fc2.', '4.')
+ else:
+ new_k = new_k.replace('fc2.', '3.')
+ string += f'{new_k} {v.shape}-{new_v.shape}'
+ elif k.startswith('norm'):
+ stage_i = int(k[4])
+ new_k = k.replace(f'norm{stage_i}', f'layers.{stage_i - 1}.2')
+ new_v = v
+ else:
+ new_k = k
+ new_v = v
+ new_ckpt[new_k] = new_v
+
+ return new_ckpt
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/regnet.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/regnet.py
new file mode 100644
index 0000000000000000000000000000000000000000..55d3ce075f0cec68de4537a71ed569151d684562
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/regnet.py
@@ -0,0 +1,356 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+
+import numpy as np
+import torch.nn as nn
+from mmcv.cnn import build_conv_layer, build_norm_layer
+
+from mmdet.registry import MODELS
+from .resnet import ResNet
+from .resnext import Bottleneck
+
+
+@MODELS.register_module()
+class RegNet(ResNet):
+ """RegNet backbone.
+
+ More details can be found in `paper `_ .
+
+ Args:
+ arch (dict): The parameter of RegNets.
+
+ - w0 (int): initial width
+ - wa (float): slope of width
+ - wm (float): quantization parameter to quantize the width
+ - depth (int): depth of the backbone
+ - group_w (int): width of group
+ - bot_mul (float): bottleneck ratio, i.e. expansion of bottleneck.
+ strides (Sequence[int]): Strides of the first block of each stage.
+ base_channels (int): Base channels after stem layer.
+ in_channels (int): Number of input image channels. Default: 3.
+ dilations (Sequence[int]): Dilation of each stage.
+ out_indices (Sequence[int]): Output from which stages.
+ style (str): `pytorch` or `caffe`. If set to "pytorch", the stride-two
+ layer is the 3x3 conv layer, otherwise the stride-two layer is
+ the first 1x1 conv layer.
+ frozen_stages (int): Stages to be frozen (all param fixed). -1 means
+ not freezing any parameters.
+ norm_cfg (dict): dictionary to construct and config norm layer.
+ norm_eval (bool): Whether to set norm layers to eval mode, namely,
+ freeze running stats (mean and var). Note: Effect on Batch Norm
+ and its variants only.
+ with_cp (bool): Use checkpoint or not. Using checkpoint will save some
+ memory while slowing down the training speed.
+ zero_init_residual (bool): whether to use zero init for last norm layer
+ in resblocks to let them behave as identity.
+ pretrained (str, optional): model pretrained path. Default: None
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None
+
+ Example:
+ >>> from mmdet.models import RegNet
+ >>> import torch
+ >>> self = RegNet(
+ arch=dict(
+ w0=88,
+ wa=26.31,
+ wm=2.25,
+ group_w=48,
+ depth=25,
+ bot_mul=1.0))
+ >>> self.eval()
+ >>> inputs = torch.rand(1, 3, 32, 32)
+ >>> level_outputs = self.forward(inputs)
+ >>> for level_out in level_outputs:
+ ... print(tuple(level_out.shape))
+ (1, 96, 8, 8)
+ (1, 192, 4, 4)
+ (1, 432, 2, 2)
+ (1, 1008, 1, 1)
+ """
+ arch_settings = {
+ 'regnetx_400mf':
+ dict(w0=24, wa=24.48, wm=2.54, group_w=16, depth=22, bot_mul=1.0),
+ 'regnetx_800mf':
+ dict(w0=56, wa=35.73, wm=2.28, group_w=16, depth=16, bot_mul=1.0),
+ 'regnetx_1.6gf':
+ dict(w0=80, wa=34.01, wm=2.25, group_w=24, depth=18, bot_mul=1.0),
+ 'regnetx_3.2gf':
+ dict(w0=88, wa=26.31, wm=2.25, group_w=48, depth=25, bot_mul=1.0),
+ 'regnetx_4.0gf':
+ dict(w0=96, wa=38.65, wm=2.43, group_w=40, depth=23, bot_mul=1.0),
+ 'regnetx_6.4gf':
+ dict(w0=184, wa=60.83, wm=2.07, group_w=56, depth=17, bot_mul=1.0),
+ 'regnetx_8.0gf':
+ dict(w0=80, wa=49.56, wm=2.88, group_w=120, depth=23, bot_mul=1.0),
+ 'regnetx_12gf':
+ dict(w0=168, wa=73.36, wm=2.37, group_w=112, depth=19, bot_mul=1.0),
+ }
+
+ def __init__(self,
+ arch,
+ in_channels=3,
+ stem_channels=32,
+ base_channels=32,
+ strides=(2, 2, 2, 2),
+ dilations=(1, 1, 1, 1),
+ out_indices=(0, 1, 2, 3),
+ style='pytorch',
+ deep_stem=False,
+ avg_down=False,
+ frozen_stages=-1,
+ conv_cfg=None,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ dcn=None,
+ stage_with_dcn=(False, False, False, False),
+ plugins=None,
+ with_cp=False,
+ zero_init_residual=True,
+ pretrained=None,
+ init_cfg=None):
+ super(ResNet, self).__init__(init_cfg)
+
+ # Generate RegNet parameters first
+ if isinstance(arch, str):
+ assert arch in self.arch_settings, \
+ f'"arch": "{arch}" is not one of the' \
+ ' arch_settings'
+ arch = self.arch_settings[arch]
+ elif not isinstance(arch, dict):
+ raise ValueError('Expect "arch" to be either a string '
+ f'or a dict, got {type(arch)}')
+
+ widths, num_stages = self.generate_regnet(
+ arch['w0'],
+ arch['wa'],
+ arch['wm'],
+ arch['depth'],
+ )
+ # Convert to per stage format
+ stage_widths, stage_blocks = self.get_stages_from_blocks(widths)
+ # Generate group widths and bot muls
+ group_widths = [arch['group_w'] for _ in range(num_stages)]
+ self.bottleneck_ratio = [arch['bot_mul'] for _ in range(num_stages)]
+ # Adjust the compatibility of stage_widths and group_widths
+ stage_widths, group_widths = self.adjust_width_group(
+ stage_widths, self.bottleneck_ratio, group_widths)
+
+ # Group params by stage
+ self.stage_widths = stage_widths
+ self.group_widths = group_widths
+ self.depth = sum(stage_blocks)
+ self.stem_channels = stem_channels
+ self.base_channels = base_channels
+ self.num_stages = num_stages
+ assert num_stages >= 1 and num_stages <= 4
+ self.strides = strides
+ self.dilations = dilations
+ assert len(strides) == len(dilations) == num_stages
+ self.out_indices = out_indices
+ assert max(out_indices) < num_stages
+ self.style = style
+ self.deep_stem = deep_stem
+ self.avg_down = avg_down
+ self.frozen_stages = frozen_stages
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self.with_cp = with_cp
+ self.norm_eval = norm_eval
+ self.dcn = dcn
+ self.stage_with_dcn = stage_with_dcn
+ if dcn is not None:
+ assert len(stage_with_dcn) == num_stages
+ self.plugins = plugins
+ self.zero_init_residual = zero_init_residual
+ self.block = Bottleneck
+ expansion_bak = self.block.expansion
+ self.block.expansion = 1
+ self.stage_blocks = stage_blocks[:num_stages]
+
+ self._make_stem_layer(in_channels, stem_channels)
+
+ block_init_cfg = None
+ assert not (init_cfg and pretrained), \
+ 'init_cfg and pretrained cannot be specified at the same time'
+ if isinstance(pretrained, str):
+ warnings.warn('DeprecationWarning: pretrained is deprecated, '
+ 'please use "init_cfg" instead')
+ self.init_cfg = dict(type='Pretrained', checkpoint=pretrained)
+ elif pretrained is None:
+ if init_cfg is None:
+ self.init_cfg = [
+ dict(type='Kaiming', layer='Conv2d'),
+ dict(
+ type='Constant',
+ val=1,
+ layer=['_BatchNorm', 'GroupNorm'])
+ ]
+ if self.zero_init_residual:
+ block_init_cfg = dict(
+ type='Constant', val=0, override=dict(name='norm3'))
+ else:
+ raise TypeError('pretrained must be a str or None')
+
+ self.inplanes = stem_channels
+ self.res_layers = []
+ for i, num_blocks in enumerate(self.stage_blocks):
+ stride = self.strides[i]
+ dilation = self.dilations[i]
+ group_width = self.group_widths[i]
+ width = int(round(self.stage_widths[i] * self.bottleneck_ratio[i]))
+ stage_groups = width // group_width
+
+ dcn = self.dcn if self.stage_with_dcn[i] else None
+ if self.plugins is not None:
+ stage_plugins = self.make_stage_plugins(self.plugins, i)
+ else:
+ stage_plugins = None
+
+ res_layer = self.make_res_layer(
+ block=self.block,
+ inplanes=self.inplanes,
+ planes=self.stage_widths[i],
+ num_blocks=num_blocks,
+ stride=stride,
+ dilation=dilation,
+ style=self.style,
+ avg_down=self.avg_down,
+ with_cp=self.with_cp,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ dcn=dcn,
+ plugins=stage_plugins,
+ groups=stage_groups,
+ base_width=group_width,
+ base_channels=self.stage_widths[i],
+ init_cfg=block_init_cfg)
+ self.inplanes = self.stage_widths[i]
+ layer_name = f'layer{i + 1}'
+ self.add_module(layer_name, res_layer)
+ self.res_layers.append(layer_name)
+
+ self._freeze_stages()
+
+ self.feat_dim = stage_widths[-1]
+ self.block.expansion = expansion_bak
+
+ def _make_stem_layer(self, in_channels, base_channels):
+ self.conv1 = build_conv_layer(
+ self.conv_cfg,
+ in_channels,
+ base_channels,
+ kernel_size=3,
+ stride=2,
+ padding=1,
+ bias=False)
+ self.norm1_name, norm1 = build_norm_layer(
+ self.norm_cfg, base_channels, postfix=1)
+ self.add_module(self.norm1_name, norm1)
+ self.relu = nn.ReLU(inplace=True)
+
+ def generate_regnet(self,
+ initial_width,
+ width_slope,
+ width_parameter,
+ depth,
+ divisor=8):
+ """Generates per block width from RegNet parameters.
+
+ Args:
+ initial_width ([int]): Initial width of the backbone
+ width_slope ([float]): Slope of the quantized linear function
+ width_parameter ([int]): Parameter used to quantize the width.
+ depth ([int]): Depth of the backbone.
+ divisor (int, optional): The divisor of channels. Defaults to 8.
+
+ Returns:
+ list, int: return a list of widths of each stage and the number \
+ of stages
+ """
+ assert width_slope >= 0
+ assert initial_width > 0
+ assert width_parameter > 1
+ assert initial_width % divisor == 0
+ widths_cont = np.arange(depth) * width_slope + initial_width
+ ks = np.round(
+ np.log(widths_cont / initial_width) / np.log(width_parameter))
+ widths = initial_width * np.power(width_parameter, ks)
+ widths = np.round(np.divide(widths, divisor)) * divisor
+ num_stages = len(np.unique(widths))
+ widths, widths_cont = widths.astype(int).tolist(), widths_cont.tolist()
+ return widths, num_stages
+
+ @staticmethod
+ def quantize_float(number, divisor):
+ """Converts a float to closest non-zero int divisible by divisor.
+
+ Args:
+ number (int): Original number to be quantized.
+ divisor (int): Divisor used to quantize the number.
+
+ Returns:
+ int: quantized number that is divisible by devisor.
+ """
+ return int(round(number / divisor) * divisor)
+
+ def adjust_width_group(self, widths, bottleneck_ratio, groups):
+ """Adjusts the compatibility of widths and groups.
+
+ Args:
+ widths (list[int]): Width of each stage.
+ bottleneck_ratio (float): Bottleneck ratio.
+ groups (int): number of groups in each stage
+
+ Returns:
+ tuple(list): The adjusted widths and groups of each stage.
+ """
+ bottleneck_width = [
+ int(w * b) for w, b in zip(widths, bottleneck_ratio)
+ ]
+ groups = [min(g, w_bot) for g, w_bot in zip(groups, bottleneck_width)]
+ bottleneck_width = [
+ self.quantize_float(w_bot, g)
+ for w_bot, g in zip(bottleneck_width, groups)
+ ]
+ widths = [
+ int(w_bot / b)
+ for w_bot, b in zip(bottleneck_width, bottleneck_ratio)
+ ]
+ return widths, groups
+
+ def get_stages_from_blocks(self, widths):
+ """Gets widths/stage_blocks of network at each stage.
+
+ Args:
+ widths (list[int]): Width in each stage.
+
+ Returns:
+ tuple(list): width and depth of each stage
+ """
+ width_diff = [
+ width != width_prev
+ for width, width_prev in zip(widths + [0], [0] + widths)
+ ]
+ stage_widths = [
+ width for width, diff in zip(widths, width_diff[:-1]) if diff
+ ]
+ stage_blocks = np.diff([
+ depth for depth, diff in zip(range(len(width_diff)), width_diff)
+ if diff
+ ]).tolist()
+ return stage_widths, stage_blocks
+
+ def forward(self, x):
+ """Forward function."""
+ x = self.conv1(x)
+ x = self.norm1(x)
+ x = self.relu(x)
+
+ outs = []
+ for i, layer_name in enumerate(self.res_layers):
+ res_layer = getattr(self, layer_name)
+ x = res_layer(x)
+ if i in self.out_indices:
+ outs.append(x)
+ return tuple(outs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/res2net.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/res2net.py
new file mode 100644
index 0000000000000000000000000000000000000000..958fc88465c6769cb4c50907c92335331e8b7834
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/res2net.py
@@ -0,0 +1,327 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+
+import torch
+import torch.nn as nn
+import torch.utils.checkpoint as cp
+from mmcv.cnn import build_conv_layer, build_norm_layer
+from mmengine.model import Sequential
+
+from mmdet.registry import MODELS
+from .resnet import Bottleneck as _Bottleneck
+from .resnet import ResNet
+
+
+class Bottle2neck(_Bottleneck):
+ expansion = 4
+
+ def __init__(self,
+ inplanes,
+ planes,
+ scales=4,
+ base_width=26,
+ base_channels=64,
+ stage_type='normal',
+ **kwargs):
+ """Bottle2neck block for Res2Net.
+
+ If style is "pytorch", the stride-two layer is the 3x3 conv layer, if
+ it is "caffe", the stride-two layer is the first 1x1 conv layer.
+ """
+ super(Bottle2neck, self).__init__(inplanes, planes, **kwargs)
+ assert scales > 1, 'Res2Net degenerates to ResNet when scales = 1.'
+ width = int(math.floor(self.planes * (base_width / base_channels)))
+
+ self.norm1_name, norm1 = build_norm_layer(
+ self.norm_cfg, width * scales, postfix=1)
+ self.norm3_name, norm3 = build_norm_layer(
+ self.norm_cfg, self.planes * self.expansion, postfix=3)
+
+ self.conv1 = build_conv_layer(
+ self.conv_cfg,
+ self.inplanes,
+ width * scales,
+ kernel_size=1,
+ stride=self.conv1_stride,
+ bias=False)
+ self.add_module(self.norm1_name, norm1)
+
+ if stage_type == 'stage' and self.conv2_stride != 1:
+ self.pool = nn.AvgPool2d(
+ kernel_size=3, stride=self.conv2_stride, padding=1)
+ convs = []
+ bns = []
+
+ fallback_on_stride = False
+ if self.with_dcn:
+ fallback_on_stride = self.dcn.pop('fallback_on_stride', False)
+ if not self.with_dcn or fallback_on_stride:
+ for i in range(scales - 1):
+ convs.append(
+ build_conv_layer(
+ self.conv_cfg,
+ width,
+ width,
+ kernel_size=3,
+ stride=self.conv2_stride,
+ padding=self.dilation,
+ dilation=self.dilation,
+ bias=False))
+ bns.append(
+ build_norm_layer(self.norm_cfg, width, postfix=i + 1)[1])
+ self.convs = nn.ModuleList(convs)
+ self.bns = nn.ModuleList(bns)
+ else:
+ assert self.conv_cfg is None, 'conv_cfg must be None for DCN'
+ for i in range(scales - 1):
+ convs.append(
+ build_conv_layer(
+ self.dcn,
+ width,
+ width,
+ kernel_size=3,
+ stride=self.conv2_stride,
+ padding=self.dilation,
+ dilation=self.dilation,
+ bias=False))
+ bns.append(
+ build_norm_layer(self.norm_cfg, width, postfix=i + 1)[1])
+ self.convs = nn.ModuleList(convs)
+ self.bns = nn.ModuleList(bns)
+
+ self.conv3 = build_conv_layer(
+ self.conv_cfg,
+ width * scales,
+ self.planes * self.expansion,
+ kernel_size=1,
+ bias=False)
+ self.add_module(self.norm3_name, norm3)
+
+ self.stage_type = stage_type
+ self.scales = scales
+ self.width = width
+ delattr(self, 'conv2')
+ delattr(self, self.norm2_name)
+
+ def forward(self, x):
+ """Forward function."""
+
+ def _inner_forward(x):
+ identity = x
+
+ out = self.conv1(x)
+ out = self.norm1(out)
+ out = self.relu(out)
+
+ if self.with_plugins:
+ out = self.forward_plugin(out, self.after_conv1_plugin_names)
+
+ spx = torch.split(out, self.width, 1)
+ sp = self.convs[0](spx[0].contiguous())
+ sp = self.relu(self.bns[0](sp))
+ out = sp
+ for i in range(1, self.scales - 1):
+ if self.stage_type == 'stage':
+ sp = spx[i]
+ else:
+ sp = sp + spx[i]
+ sp = self.convs[i](sp.contiguous())
+ sp = self.relu(self.bns[i](sp))
+ out = torch.cat((out, sp), 1)
+
+ if self.stage_type == 'normal' or self.conv2_stride == 1:
+ out = torch.cat((out, spx[self.scales - 1]), 1)
+ elif self.stage_type == 'stage':
+ out = torch.cat((out, self.pool(spx[self.scales - 1])), 1)
+
+ if self.with_plugins:
+ out = self.forward_plugin(out, self.after_conv2_plugin_names)
+
+ out = self.conv3(out)
+ out = self.norm3(out)
+
+ if self.with_plugins:
+ out = self.forward_plugin(out, self.after_conv3_plugin_names)
+
+ if self.downsample is not None:
+ identity = self.downsample(x)
+
+ out += identity
+
+ return out
+
+ if self.with_cp and x.requires_grad:
+ out = cp.checkpoint(_inner_forward, x)
+ else:
+ out = _inner_forward(x)
+
+ out = self.relu(out)
+
+ return out
+
+
+class Res2Layer(Sequential):
+ """Res2Layer to build Res2Net style backbone.
+
+ Args:
+ block (nn.Module): block used to build ResLayer.
+ inplanes (int): inplanes of block.
+ planes (int): planes of block.
+ num_blocks (int): number of blocks.
+ stride (int): stride of the first block. Default: 1
+ avg_down (bool): Use AvgPool instead of stride conv when
+ downsampling in the bottle2neck. Default: False
+ conv_cfg (dict): dictionary to construct and config conv layer.
+ Default: None
+ norm_cfg (dict): dictionary to construct and config norm layer.
+ Default: dict(type='BN')
+ scales (int): Scales used in Res2Net. Default: 4
+ base_width (int): Basic width of each scale. Default: 26
+ """
+
+ def __init__(self,
+ block,
+ inplanes,
+ planes,
+ num_blocks,
+ stride=1,
+ avg_down=True,
+ conv_cfg=None,
+ norm_cfg=dict(type='BN'),
+ scales=4,
+ base_width=26,
+ **kwargs):
+ self.block = block
+
+ downsample = None
+ if stride != 1 or inplanes != planes * block.expansion:
+ downsample = nn.Sequential(
+ nn.AvgPool2d(
+ kernel_size=stride,
+ stride=stride,
+ ceil_mode=True,
+ count_include_pad=False),
+ build_conv_layer(
+ conv_cfg,
+ inplanes,
+ planes * block.expansion,
+ kernel_size=1,
+ stride=1,
+ bias=False),
+ build_norm_layer(norm_cfg, planes * block.expansion)[1],
+ )
+
+ layers = []
+ layers.append(
+ block(
+ inplanes=inplanes,
+ planes=planes,
+ stride=stride,
+ downsample=downsample,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ scales=scales,
+ base_width=base_width,
+ stage_type='stage',
+ **kwargs))
+ inplanes = planes * block.expansion
+ for i in range(1, num_blocks):
+ layers.append(
+ block(
+ inplanes=inplanes,
+ planes=planes,
+ stride=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ scales=scales,
+ base_width=base_width,
+ **kwargs))
+ super(Res2Layer, self).__init__(*layers)
+
+
+@MODELS.register_module()
+class Res2Net(ResNet):
+ """Res2Net backbone.
+
+ Args:
+ scales (int): Scales used in Res2Net. Default: 4
+ base_width (int): Basic width of each scale. Default: 26
+ depth (int): Depth of res2net, from {50, 101, 152}.
+ in_channels (int): Number of input image channels. Default: 3.
+ num_stages (int): Res2net stages. Default: 4.
+ strides (Sequence[int]): Strides of the first block of each stage.
+ dilations (Sequence[int]): Dilation of each stage.
+ out_indices (Sequence[int]): Output from which stages.
+ style (str): `pytorch` or `caffe`. If set to "pytorch", the stride-two
+ layer is the 3x3 conv layer, otherwise the stride-two layer is
+ the first 1x1 conv layer.
+ deep_stem (bool): Replace 7x7 conv in input stem with 3 3x3 conv
+ avg_down (bool): Use AvgPool instead of stride conv when
+ downsampling in the bottle2neck.
+ frozen_stages (int): Stages to be frozen (stop grad and set eval mode).
+ -1 means not freezing any parameters.
+ norm_cfg (dict): Dictionary to construct and config norm layer.
+ norm_eval (bool): Whether to set norm layers to eval mode, namely,
+ freeze running stats (mean and var). Note: Effect on Batch Norm
+ and its variants only.
+ plugins (list[dict]): List of plugins for stages, each dict contains:
+
+ - cfg (dict, required): Cfg dict to build plugin.
+ - position (str, required): Position inside block to insert
+ plugin, options are 'after_conv1', 'after_conv2', 'after_conv3'.
+ - stages (tuple[bool], optional): Stages to apply plugin, length
+ should be same as 'num_stages'.
+ with_cp (bool): Use checkpoint or not. Using checkpoint will save some
+ memory while slowing down the training speed.
+ zero_init_residual (bool): Whether to use zero init for last norm layer
+ in resblocks to let them behave as identity.
+ pretrained (str, optional): model pretrained path. Default: None
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None
+
+ Example:
+ >>> from mmdet.models import Res2Net
+ >>> import torch
+ >>> self = Res2Net(depth=50, scales=4, base_width=26)
+ >>> self.eval()
+ >>> inputs = torch.rand(1, 3, 32, 32)
+ >>> level_outputs = self.forward(inputs)
+ >>> for level_out in level_outputs:
+ ... print(tuple(level_out.shape))
+ (1, 256, 8, 8)
+ (1, 512, 4, 4)
+ (1, 1024, 2, 2)
+ (1, 2048, 1, 1)
+ """
+
+ arch_settings = {
+ 50: (Bottle2neck, (3, 4, 6, 3)),
+ 101: (Bottle2neck, (3, 4, 23, 3)),
+ 152: (Bottle2neck, (3, 8, 36, 3))
+ }
+
+ def __init__(self,
+ scales=4,
+ base_width=26,
+ style='pytorch',
+ deep_stem=True,
+ avg_down=True,
+ pretrained=None,
+ init_cfg=None,
+ **kwargs):
+ self.scales = scales
+ self.base_width = base_width
+ super(Res2Net, self).__init__(
+ style='pytorch',
+ deep_stem=True,
+ avg_down=True,
+ pretrained=pretrained,
+ init_cfg=init_cfg,
+ **kwargs)
+
+ def make_res_layer(self, **kwargs):
+ return Res2Layer(
+ scales=self.scales,
+ base_width=self.base_width,
+ base_channels=self.base_channels,
+ **kwargs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/resnest.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/resnest.py
new file mode 100644
index 0000000000000000000000000000000000000000..d4466c4cc416237bee1f870b52e3c20a849c5a60
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/resnest.py
@@ -0,0 +1,322 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+import torch.utils.checkpoint as cp
+from mmcv.cnn import build_conv_layer, build_norm_layer
+from mmengine.model import BaseModule
+
+from mmdet.registry import MODELS
+from ..layers import ResLayer
+from .resnet import Bottleneck as _Bottleneck
+from .resnet import ResNetV1d
+
+
+class RSoftmax(nn.Module):
+ """Radix Softmax module in ``SplitAttentionConv2d``.
+
+ Args:
+ radix (int): Radix of input.
+ groups (int): Groups of input.
+ """
+
+ def __init__(self, radix, groups):
+ super().__init__()
+ self.radix = radix
+ self.groups = groups
+
+ def forward(self, x):
+ batch = x.size(0)
+ if self.radix > 1:
+ x = x.view(batch, self.groups, self.radix, -1).transpose(1, 2)
+ x = F.softmax(x, dim=1)
+ x = x.reshape(batch, -1)
+ else:
+ x = torch.sigmoid(x)
+ return x
+
+
+class SplitAttentionConv2d(BaseModule):
+ """Split-Attention Conv2d in ResNeSt.
+
+ Args:
+ in_channels (int): Number of channels in the input feature map.
+ channels (int): Number of intermediate channels.
+ kernel_size (int | tuple[int]): Size of the convolution kernel.
+ stride (int | tuple[int]): Stride of the convolution.
+ padding (int | tuple[int]): Zero-padding added to both sides of
+ dilation (int | tuple[int]): Spacing between kernel elements.
+ groups (int): Number of blocked connections from input channels to
+ output channels.
+ groups (int): Same as nn.Conv2d.
+ radix (int): Radix of SpltAtConv2d. Default: 2
+ reduction_factor (int): Reduction factor of inter_channels. Default: 4.
+ conv_cfg (dict): Config dict for convolution layer. Default: None,
+ which means using conv2d.
+ norm_cfg (dict): Config dict for normalization layer. Default: None.
+ dcn (dict): Config dict for DCN. Default: None.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None
+ """
+
+ def __init__(self,
+ in_channels,
+ channels,
+ kernel_size,
+ stride=1,
+ padding=0,
+ dilation=1,
+ groups=1,
+ radix=2,
+ reduction_factor=4,
+ conv_cfg=None,
+ norm_cfg=dict(type='BN'),
+ dcn=None,
+ init_cfg=None):
+ super(SplitAttentionConv2d, self).__init__(init_cfg)
+ inter_channels = max(in_channels * radix // reduction_factor, 32)
+ self.radix = radix
+ self.groups = groups
+ self.channels = channels
+ self.with_dcn = dcn is not None
+ self.dcn = dcn
+ fallback_on_stride = False
+ if self.with_dcn:
+ fallback_on_stride = self.dcn.pop('fallback_on_stride', False)
+ if self.with_dcn and not fallback_on_stride:
+ assert conv_cfg is None, 'conv_cfg must be None for DCN'
+ conv_cfg = dcn
+ self.conv = build_conv_layer(
+ conv_cfg,
+ in_channels,
+ channels * radix,
+ kernel_size,
+ stride=stride,
+ padding=padding,
+ dilation=dilation,
+ groups=groups * radix,
+ bias=False)
+ # To be consistent with original implementation, starting from 0
+ self.norm0_name, norm0 = build_norm_layer(
+ norm_cfg, channels * radix, postfix=0)
+ self.add_module(self.norm0_name, norm0)
+ self.relu = nn.ReLU(inplace=True)
+ self.fc1 = build_conv_layer(
+ None, channels, inter_channels, 1, groups=self.groups)
+ self.norm1_name, norm1 = build_norm_layer(
+ norm_cfg, inter_channels, postfix=1)
+ self.add_module(self.norm1_name, norm1)
+ self.fc2 = build_conv_layer(
+ None, inter_channels, channels * radix, 1, groups=self.groups)
+ self.rsoftmax = RSoftmax(radix, groups)
+
+ @property
+ def norm0(self):
+ """nn.Module: the normalization layer named "norm0" """
+ return getattr(self, self.norm0_name)
+
+ @property
+ def norm1(self):
+ """nn.Module: the normalization layer named "norm1" """
+ return getattr(self, self.norm1_name)
+
+ def forward(self, x):
+ x = self.conv(x)
+ x = self.norm0(x)
+ x = self.relu(x)
+
+ batch, rchannel = x.shape[:2]
+ batch = x.size(0)
+ if self.radix > 1:
+ splits = x.view(batch, self.radix, -1, *x.shape[2:])
+ gap = splits.sum(dim=1)
+ else:
+ gap = x
+ gap = F.adaptive_avg_pool2d(gap, 1)
+ gap = self.fc1(gap)
+
+ gap = self.norm1(gap)
+ gap = self.relu(gap)
+
+ atten = self.fc2(gap)
+ atten = self.rsoftmax(atten).view(batch, -1, 1, 1)
+
+ if self.radix > 1:
+ attens = atten.view(batch, self.radix, -1, *atten.shape[2:])
+ out = torch.sum(attens * splits, dim=1)
+ else:
+ out = atten * x
+ return out.contiguous()
+
+
+class Bottleneck(_Bottleneck):
+ """Bottleneck block for ResNeSt.
+
+ Args:
+ inplane (int): Input planes of this block.
+ planes (int): Middle planes of this block.
+ groups (int): Groups of conv2.
+ base_width (int): Base of width in terms of base channels. Default: 4.
+ base_channels (int): Base of channels for calculating width.
+ Default: 64.
+ radix (int): Radix of SpltAtConv2d. Default: 2
+ reduction_factor (int): Reduction factor of inter_channels in
+ SplitAttentionConv2d. Default: 4.
+ avg_down_stride (bool): Whether to use average pool for stride in
+ Bottleneck. Default: True.
+ kwargs (dict): Key word arguments for base class.
+ """
+ expansion = 4
+
+ def __init__(self,
+ inplanes,
+ planes,
+ groups=1,
+ base_width=4,
+ base_channels=64,
+ radix=2,
+ reduction_factor=4,
+ avg_down_stride=True,
+ **kwargs):
+ """Bottleneck block for ResNeSt."""
+ super(Bottleneck, self).__init__(inplanes, planes, **kwargs)
+
+ if groups == 1:
+ width = self.planes
+ else:
+ width = math.floor(self.planes *
+ (base_width / base_channels)) * groups
+
+ self.avg_down_stride = avg_down_stride and self.conv2_stride > 1
+
+ self.norm1_name, norm1 = build_norm_layer(
+ self.norm_cfg, width, postfix=1)
+ self.norm3_name, norm3 = build_norm_layer(
+ self.norm_cfg, self.planes * self.expansion, postfix=3)
+
+ self.conv1 = build_conv_layer(
+ self.conv_cfg,
+ self.inplanes,
+ width,
+ kernel_size=1,
+ stride=self.conv1_stride,
+ bias=False)
+ self.add_module(self.norm1_name, norm1)
+ self.with_modulated_dcn = False
+ self.conv2 = SplitAttentionConv2d(
+ width,
+ width,
+ kernel_size=3,
+ stride=1 if self.avg_down_stride else self.conv2_stride,
+ padding=self.dilation,
+ dilation=self.dilation,
+ groups=groups,
+ radix=radix,
+ reduction_factor=reduction_factor,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ dcn=self.dcn)
+ delattr(self, self.norm2_name)
+
+ if self.avg_down_stride:
+ self.avd_layer = nn.AvgPool2d(3, self.conv2_stride, padding=1)
+
+ self.conv3 = build_conv_layer(
+ self.conv_cfg,
+ width,
+ self.planes * self.expansion,
+ kernel_size=1,
+ bias=False)
+ self.add_module(self.norm3_name, norm3)
+
+ def forward(self, x):
+
+ def _inner_forward(x):
+ identity = x
+
+ out = self.conv1(x)
+ out = self.norm1(out)
+ out = self.relu(out)
+
+ if self.with_plugins:
+ out = self.forward_plugin(out, self.after_conv1_plugin_names)
+
+ out = self.conv2(out)
+
+ if self.avg_down_stride:
+ out = self.avd_layer(out)
+
+ if self.with_plugins:
+ out = self.forward_plugin(out, self.after_conv2_plugin_names)
+
+ out = self.conv3(out)
+ out = self.norm3(out)
+
+ if self.with_plugins:
+ out = self.forward_plugin(out, self.after_conv3_plugin_names)
+
+ if self.downsample is not None:
+ identity = self.downsample(x)
+
+ out += identity
+
+ return out
+
+ if self.with_cp and x.requires_grad:
+ out = cp.checkpoint(_inner_forward, x)
+ else:
+ out = _inner_forward(x)
+
+ out = self.relu(out)
+
+ return out
+
+
+@MODELS.register_module()
+class ResNeSt(ResNetV1d):
+ """ResNeSt backbone.
+
+ Args:
+ groups (int): Number of groups of Bottleneck. Default: 1
+ base_width (int): Base width of Bottleneck. Default: 4
+ radix (int): Radix of SplitAttentionConv2d. Default: 2
+ reduction_factor (int): Reduction factor of inter_channels in
+ SplitAttentionConv2d. Default: 4.
+ avg_down_stride (bool): Whether to use average pool for stride in
+ Bottleneck. Default: True.
+ kwargs (dict): Keyword arguments for ResNet.
+ """
+
+ arch_settings = {
+ 50: (Bottleneck, (3, 4, 6, 3)),
+ 101: (Bottleneck, (3, 4, 23, 3)),
+ 152: (Bottleneck, (3, 8, 36, 3)),
+ 200: (Bottleneck, (3, 24, 36, 3))
+ }
+
+ def __init__(self,
+ groups=1,
+ base_width=4,
+ radix=2,
+ reduction_factor=4,
+ avg_down_stride=True,
+ **kwargs):
+ self.groups = groups
+ self.base_width = base_width
+ self.radix = radix
+ self.reduction_factor = reduction_factor
+ self.avg_down_stride = avg_down_stride
+ super(ResNeSt, self).__init__(**kwargs)
+
+ def make_res_layer(self, **kwargs):
+ """Pack all blocks in a stage into a ``ResLayer``."""
+ return ResLayer(
+ groups=self.groups,
+ base_width=self.base_width,
+ base_channels=self.base_channels,
+ radix=self.radix,
+ reduction_factor=self.reduction_factor,
+ avg_down_stride=self.avg_down_stride,
+ **kwargs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/resnet.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/resnet.py
new file mode 100644
index 0000000000000000000000000000000000000000..1d6f48f94f286e3c5e3179f752a7b36ea77c0d45
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/resnet.py
@@ -0,0 +1,672 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+
+import torch.nn as nn
+import torch.utils.checkpoint as cp
+from mmcv.cnn import build_conv_layer, build_norm_layer, build_plugin_layer
+from mmengine.model import BaseModule
+from torch.nn.modules.batchnorm import _BatchNorm
+
+from mmdet.registry import MODELS
+from ..layers import ResLayer
+
+
+class BasicBlock(BaseModule):
+ expansion = 1
+
+ def __init__(self,
+ inplanes,
+ planes,
+ stride=1,
+ dilation=1,
+ downsample=None,
+ style='pytorch',
+ with_cp=False,
+ conv_cfg=None,
+ norm_cfg=dict(type='BN'),
+ dcn=None,
+ plugins=None,
+ init_cfg=None):
+ super(BasicBlock, self).__init__(init_cfg)
+ assert dcn is None, 'Not implemented yet.'
+ assert plugins is None, 'Not implemented yet.'
+
+ self.norm1_name, norm1 = build_norm_layer(norm_cfg, planes, postfix=1)
+ self.norm2_name, norm2 = build_norm_layer(norm_cfg, planes, postfix=2)
+
+ self.conv1 = build_conv_layer(
+ conv_cfg,
+ inplanes,
+ planes,
+ 3,
+ stride=stride,
+ padding=dilation,
+ dilation=dilation,
+ bias=False)
+ self.add_module(self.norm1_name, norm1)
+ self.conv2 = build_conv_layer(
+ conv_cfg, planes, planes, 3, padding=1, bias=False)
+ self.add_module(self.norm2_name, norm2)
+
+ self.relu = nn.ReLU(inplace=True)
+ self.downsample = downsample
+ self.stride = stride
+ self.dilation = dilation
+ self.with_cp = with_cp
+
+ @property
+ def norm1(self):
+ """nn.Module: normalization layer after the first convolution layer"""
+ return getattr(self, self.norm1_name)
+
+ @property
+ def norm2(self):
+ """nn.Module: normalization layer after the second convolution layer"""
+ return getattr(self, self.norm2_name)
+
+ def forward(self, x):
+ """Forward function."""
+
+ def _inner_forward(x):
+ identity = x
+
+ out = self.conv1(x)
+ out = self.norm1(out)
+ out = self.relu(out)
+
+ out = self.conv2(out)
+ out = self.norm2(out)
+
+ if self.downsample is not None:
+ identity = self.downsample(x)
+
+ out += identity
+
+ return out
+
+ if self.with_cp and x.requires_grad:
+ out = cp.checkpoint(_inner_forward, x)
+ else:
+ out = _inner_forward(x)
+
+ out = self.relu(out)
+
+ return out
+
+
+class Bottleneck(BaseModule):
+ expansion = 4
+
+ def __init__(self,
+ inplanes,
+ planes,
+ stride=1,
+ dilation=1,
+ downsample=None,
+ style='pytorch',
+ with_cp=False,
+ conv_cfg=None,
+ norm_cfg=dict(type='BN'),
+ dcn=None,
+ plugins=None,
+ init_cfg=None):
+ """Bottleneck block for ResNet.
+
+ If style is "pytorch", the stride-two layer is the 3x3 conv layer, if
+ it is "caffe", the stride-two layer is the first 1x1 conv layer.
+ """
+ super(Bottleneck, self).__init__(init_cfg)
+ assert style in ['pytorch', 'caffe']
+ assert dcn is None or isinstance(dcn, dict)
+ assert plugins is None or isinstance(plugins, list)
+ if plugins is not None:
+ allowed_position = ['after_conv1', 'after_conv2', 'after_conv3']
+ assert all(p['position'] in allowed_position for p in plugins)
+
+ self.inplanes = inplanes
+ self.planes = planes
+ self.stride = stride
+ self.dilation = dilation
+ self.style = style
+ self.with_cp = with_cp
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self.dcn = dcn
+ self.with_dcn = dcn is not None
+ self.plugins = plugins
+ self.with_plugins = plugins is not None
+
+ if self.with_plugins:
+ # collect plugins for conv1/conv2/conv3
+ self.after_conv1_plugins = [
+ plugin['cfg'] for plugin in plugins
+ if plugin['position'] == 'after_conv1'
+ ]
+ self.after_conv2_plugins = [
+ plugin['cfg'] for plugin in plugins
+ if plugin['position'] == 'after_conv2'
+ ]
+ self.after_conv3_plugins = [
+ plugin['cfg'] for plugin in plugins
+ if plugin['position'] == 'after_conv3'
+ ]
+
+ if self.style == 'pytorch':
+ self.conv1_stride = 1
+ self.conv2_stride = stride
+ else:
+ self.conv1_stride = stride
+ self.conv2_stride = 1
+
+ self.norm1_name, norm1 = build_norm_layer(norm_cfg, planes, postfix=1)
+ self.norm2_name, norm2 = build_norm_layer(norm_cfg, planes, postfix=2)
+ self.norm3_name, norm3 = build_norm_layer(
+ norm_cfg, planes * self.expansion, postfix=3)
+
+ self.conv1 = build_conv_layer(
+ conv_cfg,
+ inplanes,
+ planes,
+ kernel_size=1,
+ stride=self.conv1_stride,
+ bias=False)
+ self.add_module(self.norm1_name, norm1)
+ fallback_on_stride = False
+ if self.with_dcn:
+ fallback_on_stride = dcn.pop('fallback_on_stride', False)
+ if not self.with_dcn or fallback_on_stride:
+ self.conv2 = build_conv_layer(
+ conv_cfg,
+ planes,
+ planes,
+ kernel_size=3,
+ stride=self.conv2_stride,
+ padding=dilation,
+ dilation=dilation,
+ bias=False)
+ else:
+ assert self.conv_cfg is None, 'conv_cfg must be None for DCN'
+ self.conv2 = build_conv_layer(
+ dcn,
+ planes,
+ planes,
+ kernel_size=3,
+ stride=self.conv2_stride,
+ padding=dilation,
+ dilation=dilation,
+ bias=False)
+
+ self.add_module(self.norm2_name, norm2)
+ self.conv3 = build_conv_layer(
+ conv_cfg,
+ planes,
+ planes * self.expansion,
+ kernel_size=1,
+ bias=False)
+ self.add_module(self.norm3_name, norm3)
+
+ self.relu = nn.ReLU(inplace=True)
+ self.downsample = downsample
+
+ if self.with_plugins:
+ self.after_conv1_plugin_names = self.make_block_plugins(
+ planes, self.after_conv1_plugins)
+ self.after_conv2_plugin_names = self.make_block_plugins(
+ planes, self.after_conv2_plugins)
+ self.after_conv3_plugin_names = self.make_block_plugins(
+ planes * self.expansion, self.after_conv3_plugins)
+
+ def make_block_plugins(self, in_channels, plugins):
+ """make plugins for block.
+
+ Args:
+ in_channels (int): Input channels of plugin.
+ plugins (list[dict]): List of plugins cfg to build.
+
+ Returns:
+ list[str]: List of the names of plugin.
+ """
+ assert isinstance(plugins, list)
+ plugin_names = []
+ for plugin in plugins:
+ plugin = plugin.copy()
+ name, layer = build_plugin_layer(
+ plugin,
+ in_channels=in_channels,
+ postfix=plugin.pop('postfix', ''))
+ assert not hasattr(self, name), f'duplicate plugin {name}'
+ self.add_module(name, layer)
+ plugin_names.append(name)
+ return plugin_names
+
+ def forward_plugin(self, x, plugin_names):
+ out = x
+ for name in plugin_names:
+ out = getattr(self, name)(out)
+ return out
+
+ @property
+ def norm1(self):
+ """nn.Module: normalization layer after the first convolution layer"""
+ return getattr(self, self.norm1_name)
+
+ @property
+ def norm2(self):
+ """nn.Module: normalization layer after the second convolution layer"""
+ return getattr(self, self.norm2_name)
+
+ @property
+ def norm3(self):
+ """nn.Module: normalization layer after the third convolution layer"""
+ return getattr(self, self.norm3_name)
+
+ def forward(self, x):
+ """Forward function."""
+
+ def _inner_forward(x):
+ identity = x
+ out = self.conv1(x)
+ out = self.norm1(out)
+ out = self.relu(out)
+
+ if self.with_plugins:
+ out = self.forward_plugin(out, self.after_conv1_plugin_names)
+
+ out = self.conv2(out)
+ out = self.norm2(out)
+ out = self.relu(out)
+
+ if self.with_plugins:
+ out = self.forward_plugin(out, self.after_conv2_plugin_names)
+
+ out = self.conv3(out)
+ out = self.norm3(out)
+
+ if self.with_plugins:
+ out = self.forward_plugin(out, self.after_conv3_plugin_names)
+
+ if self.downsample is not None:
+ identity = self.downsample(x)
+
+ out += identity
+
+ return out
+
+ if self.with_cp and x.requires_grad:
+ out = cp.checkpoint(_inner_forward, x)
+ else:
+ out = _inner_forward(x)
+
+ out = self.relu(out)
+
+ return out
+
+
+@MODELS.register_module()
+class ResNet(BaseModule):
+ """ResNet backbone.
+
+ Args:
+ depth (int): Depth of resnet, from {18, 34, 50, 101, 152}.
+ stem_channels (int | None): Number of stem channels. If not specified,
+ it will be the same as `base_channels`. Default: None.
+ base_channels (int): Number of base channels of res layer. Default: 64.
+ in_channels (int): Number of input image channels. Default: 3.
+ num_stages (int): Resnet stages. Default: 4.
+ strides (Sequence[int]): Strides of the first block of each stage.
+ dilations (Sequence[int]): Dilation of each stage.
+ out_indices (Sequence[int]): Output from which stages.
+ style (str): `pytorch` or `caffe`. If set to "pytorch", the stride-two
+ layer is the 3x3 conv layer, otherwise the stride-two layer is
+ the first 1x1 conv layer.
+ deep_stem (bool): Replace 7x7 conv in input stem with 3 3x3 conv
+ avg_down (bool): Use AvgPool instead of stride conv when
+ downsampling in the bottleneck.
+ frozen_stages (int): Stages to be frozen (stop grad and set eval mode).
+ -1 means not freezing any parameters.
+ norm_cfg (dict): Dictionary to construct and config norm layer.
+ norm_eval (bool): Whether to set norm layers to eval mode, namely,
+ freeze running stats (mean and var). Note: Effect on Batch Norm
+ and its variants only.
+ plugins (list[dict]): List of plugins for stages, each dict contains:
+
+ - cfg (dict, required): Cfg dict to build plugin.
+ - position (str, required): Position inside block to insert
+ plugin, options are 'after_conv1', 'after_conv2', 'after_conv3'.
+ - stages (tuple[bool], optional): Stages to apply plugin, length
+ should be same as 'num_stages'.
+ with_cp (bool): Use checkpoint or not. Using checkpoint will save some
+ memory while slowing down the training speed.
+ zero_init_residual (bool): Whether to use zero init for last norm layer
+ in resblocks to let them behave as identity.
+ pretrained (str, optional): model pretrained path. Default: None
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None
+
+ Example:
+ >>> from mmdet.models import ResNet
+ >>> import torch
+ >>> self = ResNet(depth=18)
+ >>> self.eval()
+ >>> inputs = torch.rand(1, 3, 32, 32)
+ >>> level_outputs = self.forward(inputs)
+ >>> for level_out in level_outputs:
+ ... print(tuple(level_out.shape))
+ (1, 64, 8, 8)
+ (1, 128, 4, 4)
+ (1, 256, 2, 2)
+ (1, 512, 1, 1)
+ """
+
+ arch_settings = {
+ 18: (BasicBlock, (2, 2, 2, 2)),
+ 34: (BasicBlock, (3, 4, 6, 3)),
+ 50: (Bottleneck, (3, 4, 6, 3)),
+ 101: (Bottleneck, (3, 4, 23, 3)),
+ 152: (Bottleneck, (3, 8, 36, 3))
+ }
+
+ def __init__(self,
+ depth,
+ in_channels=3,
+ stem_channels=None,
+ base_channels=64,
+ num_stages=4,
+ strides=(1, 2, 2, 2),
+ dilations=(1, 1, 1, 1),
+ out_indices=(0, 1, 2, 3),
+ style='pytorch',
+ deep_stem=False,
+ avg_down=False,
+ frozen_stages=-1,
+ conv_cfg=None,
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ dcn=None,
+ stage_with_dcn=(False, False, False, False),
+ plugins=None,
+ with_cp=False,
+ zero_init_residual=True,
+ pretrained=None,
+ init_cfg=None):
+ super(ResNet, self).__init__(init_cfg)
+ self.zero_init_residual = zero_init_residual
+ if depth not in self.arch_settings:
+ raise KeyError(f'invalid depth {depth} for resnet')
+
+ block_init_cfg = None
+ assert not (init_cfg and pretrained), \
+ 'init_cfg and pretrained cannot be specified at the same time'
+ if isinstance(pretrained, str):
+ warnings.warn('DeprecationWarning: pretrained is deprecated, '
+ 'please use "init_cfg" instead')
+ self.init_cfg = dict(type='Pretrained', checkpoint=pretrained)
+ elif pretrained is None:
+ if init_cfg is None:
+ self.init_cfg = [
+ dict(type='Kaiming', layer='Conv2d'),
+ dict(
+ type='Constant',
+ val=1,
+ layer=['_BatchNorm', 'GroupNorm'])
+ ]
+ block = self.arch_settings[depth][0]
+ if self.zero_init_residual:
+ if block is BasicBlock:
+ block_init_cfg = dict(
+ type='Constant',
+ val=0,
+ override=dict(name='norm2'))
+ elif block is Bottleneck:
+ block_init_cfg = dict(
+ type='Constant',
+ val=0,
+ override=dict(name='norm3'))
+ else:
+ raise TypeError('pretrained must be a str or None')
+
+ self.depth = depth
+ if stem_channels is None:
+ stem_channels = base_channels
+ self.stem_channels = stem_channels
+ self.base_channels = base_channels
+ self.num_stages = num_stages
+ assert num_stages >= 1 and num_stages <= 4
+ self.strides = strides
+ self.dilations = dilations
+ assert len(strides) == len(dilations) == num_stages
+ self.out_indices = out_indices
+ assert max(out_indices) < num_stages
+ self.style = style
+ self.deep_stem = deep_stem
+ self.avg_down = avg_down
+ self.frozen_stages = frozen_stages
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self.with_cp = with_cp
+ self.norm_eval = norm_eval
+ self.dcn = dcn
+ self.stage_with_dcn = stage_with_dcn
+ if dcn is not None:
+ assert len(stage_with_dcn) == num_stages
+ self.plugins = plugins
+ self.block, stage_blocks = self.arch_settings[depth]
+ self.stage_blocks = stage_blocks[:num_stages]
+ self.inplanes = stem_channels
+
+ self._make_stem_layer(in_channels, stem_channels)
+
+ self.res_layers = []
+ for i, num_blocks in enumerate(self.stage_blocks):
+ stride = strides[i]
+ dilation = dilations[i]
+ dcn = self.dcn if self.stage_with_dcn[i] else None
+ if plugins is not None:
+ stage_plugins = self.make_stage_plugins(plugins, i)
+ else:
+ stage_plugins = None
+ planes = base_channels * 2**i
+ res_layer = self.make_res_layer(
+ block=self.block,
+ inplanes=self.inplanes,
+ planes=planes,
+ num_blocks=num_blocks,
+ stride=stride,
+ dilation=dilation,
+ style=self.style,
+ avg_down=self.avg_down,
+ with_cp=with_cp,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ dcn=dcn,
+ plugins=stage_plugins,
+ init_cfg=block_init_cfg)
+ self.inplanes = planes * self.block.expansion
+ layer_name = f'layer{i + 1}'
+ self.add_module(layer_name, res_layer)
+ self.res_layers.append(layer_name)
+
+ self._freeze_stages()
+
+ self.feat_dim = self.block.expansion * base_channels * 2**(
+ len(self.stage_blocks) - 1)
+
+ def make_stage_plugins(self, plugins, stage_idx):
+ """Make plugins for ResNet ``stage_idx`` th stage.
+
+ Currently we support to insert ``context_block``,
+ ``empirical_attention_block``, ``nonlocal_block`` into the backbone
+ like ResNet/ResNeXt. They could be inserted after conv1/conv2/conv3 of
+ Bottleneck.
+
+ An example of plugins format could be:
+
+ Examples:
+ >>> plugins=[
+ ... dict(cfg=dict(type='xxx', arg1='xxx'),
+ ... stages=(False, True, True, True),
+ ... position='after_conv2'),
+ ... dict(cfg=dict(type='yyy'),
+ ... stages=(True, True, True, True),
+ ... position='after_conv3'),
+ ... dict(cfg=dict(type='zzz', postfix='1'),
+ ... stages=(True, True, True, True),
+ ... position='after_conv3'),
+ ... dict(cfg=dict(type='zzz', postfix='2'),
+ ... stages=(True, True, True, True),
+ ... position='after_conv3')
+ ... ]
+ >>> self = ResNet(depth=18)
+ >>> stage_plugins = self.make_stage_plugins(plugins, 0)
+ >>> assert len(stage_plugins) == 3
+
+ Suppose ``stage_idx=0``, the structure of blocks in the stage would be:
+
+ .. code-block:: none
+
+ conv1-> conv2->conv3->yyy->zzz1->zzz2
+
+ Suppose 'stage_idx=1', the structure of blocks in the stage would be:
+
+ .. code-block:: none
+
+ conv1-> conv2->xxx->conv3->yyy->zzz1->zzz2
+
+ If stages is missing, the plugin would be applied to all stages.
+
+ Args:
+ plugins (list[dict]): List of plugins cfg to build. The postfix is
+ required if multiple same type plugins are inserted.
+ stage_idx (int): Index of stage to build
+
+ Returns:
+ list[dict]: Plugins for current stage
+ """
+ stage_plugins = []
+ for plugin in plugins:
+ plugin = plugin.copy()
+ stages = plugin.pop('stages', None)
+ assert stages is None or len(stages) == self.num_stages
+ # whether to insert plugin into current stage
+ if stages is None or stages[stage_idx]:
+ stage_plugins.append(plugin)
+
+ return stage_plugins
+
+ def make_res_layer(self, **kwargs):
+ """Pack all blocks in a stage into a ``ResLayer``."""
+ return ResLayer(**kwargs)
+
+ @property
+ def norm1(self):
+ """nn.Module: the normalization layer named "norm1" """
+ return getattr(self, self.norm1_name)
+
+ def _make_stem_layer(self, in_channels, stem_channels):
+ if self.deep_stem:
+ self.stem = nn.Sequential(
+ build_conv_layer(
+ self.conv_cfg,
+ in_channels,
+ stem_channels // 2,
+ kernel_size=3,
+ stride=2,
+ padding=1,
+ bias=False),
+ build_norm_layer(self.norm_cfg, stem_channels // 2)[1],
+ nn.ReLU(inplace=True),
+ build_conv_layer(
+ self.conv_cfg,
+ stem_channels // 2,
+ stem_channels // 2,
+ kernel_size=3,
+ stride=1,
+ padding=1,
+ bias=False),
+ build_norm_layer(self.norm_cfg, stem_channels // 2)[1],
+ nn.ReLU(inplace=True),
+ build_conv_layer(
+ self.conv_cfg,
+ stem_channels // 2,
+ stem_channels,
+ kernel_size=3,
+ stride=1,
+ padding=1,
+ bias=False),
+ build_norm_layer(self.norm_cfg, stem_channels)[1],
+ nn.ReLU(inplace=True))
+ else:
+ self.conv1 = build_conv_layer(
+ self.conv_cfg,
+ in_channels,
+ stem_channels,
+ kernel_size=7,
+ stride=2,
+ padding=3,
+ bias=False)
+ self.norm1_name, norm1 = build_norm_layer(
+ self.norm_cfg, stem_channels, postfix=1)
+ self.add_module(self.norm1_name, norm1)
+ self.relu = nn.ReLU(inplace=True)
+ self.maxpool = nn.MaxPool2d(kernel_size=3, stride=2, padding=1)
+
+ def _freeze_stages(self):
+ if self.frozen_stages >= 0:
+ if self.deep_stem:
+ self.stem.eval()
+ for param in self.stem.parameters():
+ param.requires_grad = False
+ else:
+ self.norm1.eval()
+ for m in [self.conv1, self.norm1]:
+ for param in m.parameters():
+ param.requires_grad = False
+
+ for i in range(1, self.frozen_stages + 1):
+ m = getattr(self, f'layer{i}')
+ m.eval()
+ for param in m.parameters():
+ param.requires_grad = False
+
+ def forward(self, x):
+ """Forward function."""
+ if self.deep_stem:
+ x = self.stem(x)
+ else:
+ x = self.conv1(x)
+ x = self.norm1(x)
+ x = self.relu(x)
+ x = self.maxpool(x)
+ outs = []
+ for i, layer_name in enumerate(self.res_layers):
+ res_layer = getattr(self, layer_name)
+ x = res_layer(x)
+ if i in self.out_indices:
+ outs.append(x)
+ return tuple(outs)
+
+ def train(self, mode=True):
+ """Convert the model into training mode while keep normalization layer
+ freezed."""
+ super(ResNet, self).train(mode)
+ self._freeze_stages()
+ if mode and self.norm_eval:
+ for m in self.modules():
+ # trick: eval have effect on BatchNorm only
+ if isinstance(m, _BatchNorm):
+ m.eval()
+
+
+@MODELS.register_module()
+class ResNetV1d(ResNet):
+ r"""ResNetV1d variant described in `Bag of Tricks
+ `_.
+
+ Compared with default ResNet(ResNetV1b), ResNetV1d replaces the 7x7 conv in
+ the input stem with three 3x3 convs. And in the downsampling block, a 2x2
+ avg_pool with stride 2 is added before conv, whose stride is changed to 1.
+ """
+
+ def __init__(self, **kwargs):
+ super(ResNetV1d, self).__init__(
+ deep_stem=True, avg_down=True, **kwargs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/resnext.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/resnext.py
new file mode 100644
index 0000000000000000000000000000000000000000..df3d79e046c3ab9b289bcfeb6f937c87f6c09bfa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/resnext.py
@@ -0,0 +1,154 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+
+from mmcv.cnn import build_conv_layer, build_norm_layer
+
+from mmdet.registry import MODELS
+from ..layers import ResLayer
+from .resnet import Bottleneck as _Bottleneck
+from .resnet import ResNet
+
+
+class Bottleneck(_Bottleneck):
+ expansion = 4
+
+ def __init__(self,
+ inplanes,
+ planes,
+ groups=1,
+ base_width=4,
+ base_channels=64,
+ **kwargs):
+ """Bottleneck block for ResNeXt.
+
+ If style is "pytorch", the stride-two layer is the 3x3 conv layer, if
+ it is "caffe", the stride-two layer is the first 1x1 conv layer.
+ """
+ super(Bottleneck, self).__init__(inplanes, planes, **kwargs)
+
+ if groups == 1:
+ width = self.planes
+ else:
+ width = math.floor(self.planes *
+ (base_width / base_channels)) * groups
+
+ self.norm1_name, norm1 = build_norm_layer(
+ self.norm_cfg, width, postfix=1)
+ self.norm2_name, norm2 = build_norm_layer(
+ self.norm_cfg, width, postfix=2)
+ self.norm3_name, norm3 = build_norm_layer(
+ self.norm_cfg, self.planes * self.expansion, postfix=3)
+
+ self.conv1 = build_conv_layer(
+ self.conv_cfg,
+ self.inplanes,
+ width,
+ kernel_size=1,
+ stride=self.conv1_stride,
+ bias=False)
+ self.add_module(self.norm1_name, norm1)
+ fallback_on_stride = False
+ self.with_modulated_dcn = False
+ if self.with_dcn:
+ fallback_on_stride = self.dcn.pop('fallback_on_stride', False)
+ if not self.with_dcn or fallback_on_stride:
+ self.conv2 = build_conv_layer(
+ self.conv_cfg,
+ width,
+ width,
+ kernel_size=3,
+ stride=self.conv2_stride,
+ padding=self.dilation,
+ dilation=self.dilation,
+ groups=groups,
+ bias=False)
+ else:
+ assert self.conv_cfg is None, 'conv_cfg must be None for DCN'
+ self.conv2 = build_conv_layer(
+ self.dcn,
+ width,
+ width,
+ kernel_size=3,
+ stride=self.conv2_stride,
+ padding=self.dilation,
+ dilation=self.dilation,
+ groups=groups,
+ bias=False)
+
+ self.add_module(self.norm2_name, norm2)
+ self.conv3 = build_conv_layer(
+ self.conv_cfg,
+ width,
+ self.planes * self.expansion,
+ kernel_size=1,
+ bias=False)
+ self.add_module(self.norm3_name, norm3)
+
+ if self.with_plugins:
+ self._del_block_plugins(self.after_conv1_plugin_names +
+ self.after_conv2_plugin_names +
+ self.after_conv3_plugin_names)
+ self.after_conv1_plugin_names = self.make_block_plugins(
+ width, self.after_conv1_plugins)
+ self.after_conv2_plugin_names = self.make_block_plugins(
+ width, self.after_conv2_plugins)
+ self.after_conv3_plugin_names = self.make_block_plugins(
+ self.planes * self.expansion, self.after_conv3_plugins)
+
+ def _del_block_plugins(self, plugin_names):
+ """delete plugins for block if exist.
+
+ Args:
+ plugin_names (list[str]): List of plugins name to delete.
+ """
+ assert isinstance(plugin_names, list)
+ for plugin_name in plugin_names:
+ del self._modules[plugin_name]
+
+
+@MODELS.register_module()
+class ResNeXt(ResNet):
+ """ResNeXt backbone.
+
+ Args:
+ depth (int): Depth of resnet, from {18, 34, 50, 101, 152}.
+ in_channels (int): Number of input image channels. Default: 3.
+ num_stages (int): Resnet stages. Default: 4.
+ groups (int): Group of resnext.
+ base_width (int): Base width of resnext.
+ strides (Sequence[int]): Strides of the first block of each stage.
+ dilations (Sequence[int]): Dilation of each stage.
+ out_indices (Sequence[int]): Output from which stages.
+ style (str): `pytorch` or `caffe`. If set to "pytorch", the stride-two
+ layer is the 3x3 conv layer, otherwise the stride-two layer is
+ the first 1x1 conv layer.
+ frozen_stages (int): Stages to be frozen (all param fixed). -1 means
+ not freezing any parameters.
+ norm_cfg (dict): dictionary to construct and config norm layer.
+ norm_eval (bool): Whether to set norm layers to eval mode, namely,
+ freeze running stats (mean and var). Note: Effect on Batch Norm
+ and its variants only.
+ with_cp (bool): Use checkpoint or not. Using checkpoint will save some
+ memory while slowing down the training speed.
+ zero_init_residual (bool): whether to use zero init for last norm layer
+ in resblocks to let them behave as identity.
+ """
+
+ arch_settings = {
+ 50: (Bottleneck, (3, 4, 6, 3)),
+ 101: (Bottleneck, (3, 4, 23, 3)),
+ 152: (Bottleneck, (3, 8, 36, 3))
+ }
+
+ def __init__(self, groups=1, base_width=4, **kwargs):
+ self.groups = groups
+ self.base_width = base_width
+ super(ResNeXt, self).__init__(**kwargs)
+
+ def make_res_layer(self, **kwargs):
+ """Pack all blocks in a stage into a ``ResLayer``"""
+ return ResLayer(
+ groups=self.groups,
+ base_width=self.base_width,
+ base_channels=self.base_channels,
+ **kwargs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/ssd_vgg.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/ssd_vgg.py
new file mode 100644
index 0000000000000000000000000000000000000000..843e82e2722f93b9b2abb5180c827c8f2a430b48
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/ssd_vgg.py
@@ -0,0 +1,128 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+
+import torch.nn as nn
+from mmcv.cnn import VGG
+from mmengine.model import BaseModule
+
+from mmdet.registry import MODELS
+from ..necks import ssd_neck
+
+
+@MODELS.register_module()
+class SSDVGG(VGG, BaseModule):
+ """VGG Backbone network for single-shot-detection.
+
+ Args:
+ depth (int): Depth of vgg, from {11, 13, 16, 19}.
+ with_last_pool (bool): Whether to add a pooling layer at the last
+ of the model
+ ceil_mode (bool): When True, will use `ceil` instead of `floor`
+ to compute the output shape.
+ out_indices (Sequence[int]): Output from which stages.
+ out_feature_indices (Sequence[int]): Output from which feature map.
+ pretrained (str, optional): model pretrained path. Default: None
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None
+ input_size (int, optional): Deprecated argumment.
+ Width and height of input, from {300, 512}.
+ l2_norm_scale (float, optional) : Deprecated argumment.
+ L2 normalization layer init scale.
+
+ Example:
+ >>> self = SSDVGG(input_size=300, depth=11)
+ >>> self.eval()
+ >>> inputs = torch.rand(1, 3, 300, 300)
+ >>> level_outputs = self.forward(inputs)
+ >>> for level_out in level_outputs:
+ ... print(tuple(level_out.shape))
+ (1, 1024, 19, 19)
+ (1, 512, 10, 10)
+ (1, 256, 5, 5)
+ (1, 256, 3, 3)
+ (1, 256, 1, 1)
+ """
+ extra_setting = {
+ 300: (256, 'S', 512, 128, 'S', 256, 128, 256, 128, 256),
+ 512: (256, 'S', 512, 128, 'S', 256, 128, 'S', 256, 128, 'S', 256, 128),
+ }
+
+ def __init__(self,
+ depth,
+ with_last_pool=False,
+ ceil_mode=True,
+ out_indices=(3, 4),
+ out_feature_indices=(22, 34),
+ pretrained=None,
+ init_cfg=None,
+ input_size=None,
+ l2_norm_scale=None):
+ # TODO: in_channels for mmcv.VGG
+ super(SSDVGG, self).__init__(
+ depth,
+ with_last_pool=with_last_pool,
+ ceil_mode=ceil_mode,
+ out_indices=out_indices)
+
+ self.features.add_module(
+ str(len(self.features)),
+ nn.MaxPool2d(kernel_size=3, stride=1, padding=1))
+ self.features.add_module(
+ str(len(self.features)),
+ nn.Conv2d(512, 1024, kernel_size=3, padding=6, dilation=6))
+ self.features.add_module(
+ str(len(self.features)), nn.ReLU(inplace=True))
+ self.features.add_module(
+ str(len(self.features)), nn.Conv2d(1024, 1024, kernel_size=1))
+ self.features.add_module(
+ str(len(self.features)), nn.ReLU(inplace=True))
+ self.out_feature_indices = out_feature_indices
+
+ assert not (init_cfg and pretrained), \
+ 'init_cfg and pretrained cannot be specified at the same time'
+
+ if init_cfg is not None:
+ self.init_cfg = init_cfg
+ elif isinstance(pretrained, str):
+ warnings.warn('DeprecationWarning: pretrained is deprecated, '
+ 'please use "init_cfg" instead')
+ self.init_cfg = dict(type='Pretrained', checkpoint=pretrained)
+ elif pretrained is None:
+ self.init_cfg = [
+ dict(type='Kaiming', layer='Conv2d'),
+ dict(type='Constant', val=1, layer='BatchNorm2d'),
+ dict(type='Normal', std=0.01, layer='Linear'),
+ ]
+ else:
+ raise TypeError('pretrained must be a str or None')
+
+ if input_size is not None:
+ warnings.warn('DeprecationWarning: input_size is deprecated')
+ if l2_norm_scale is not None:
+ warnings.warn('DeprecationWarning: l2_norm_scale in VGG is '
+ 'deprecated, it has been moved to SSDNeck.')
+
+ def init_weights(self, pretrained=None):
+ super(VGG, self).init_weights()
+
+ def forward(self, x):
+ """Forward function."""
+ outs = []
+ for i, layer in enumerate(self.features):
+ x = layer(x)
+ if i in self.out_feature_indices:
+ outs.append(x)
+
+ if len(outs) == 1:
+ return outs[0]
+ else:
+ return tuple(outs)
+
+
+class L2Norm(ssd_neck.L2Norm):
+
+ def __init__(self, **kwargs):
+ super(L2Norm, self).__init__(**kwargs)
+ warnings.warn('DeprecationWarning: L2Norm in ssd_vgg.py '
+ 'is deprecated, please use L2Norm in '
+ 'mmdet/models/necks/ssd_neck.py instead')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/swin.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/swin.py
new file mode 100644
index 0000000000000000000000000000000000000000..062190fa077d7b01e0c1db76bea0cfb5dc7b6620
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/swin.py
@@ -0,0 +1,819 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+from collections import OrderedDict
+from copy import deepcopy
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+import torch.utils.checkpoint as cp
+from mmcv.cnn import build_norm_layer
+from mmcv.cnn.bricks.transformer import FFN, build_dropout
+from mmengine.logging import MMLogger
+from mmengine.model import BaseModule, ModuleList
+from mmengine.model.weight_init import (constant_init, trunc_normal_,
+ trunc_normal_init)
+from mmengine.runner.checkpoint import CheckpointLoader
+from mmengine.utils import to_2tuple
+
+from mmdet.registry import MODELS
+from ..layers import PatchEmbed, PatchMerging
+
+
+class WindowMSA(BaseModule):
+ """Window based multi-head self-attention (W-MSA) module with relative
+ position bias.
+
+ Args:
+ embed_dims (int): Number of input channels.
+ num_heads (int): Number of attention heads.
+ window_size (tuple[int]): The height and width of the window.
+ qkv_bias (bool, optional): If True, add a learnable bias to q, k, v.
+ Default: True.
+ qk_scale (float | None, optional): Override default qk scale of
+ head_dim ** -0.5 if set. Default: None.
+ attn_drop_rate (float, optional): Dropout ratio of attention weight.
+ Default: 0.0
+ proj_drop_rate (float, optional): Dropout ratio of output. Default: 0.
+ init_cfg (dict | None, optional): The Config for initialization.
+ Default: None.
+ """
+
+ def __init__(self,
+ embed_dims,
+ num_heads,
+ window_size,
+ qkv_bias=True,
+ qk_scale=None,
+ attn_drop_rate=0.,
+ proj_drop_rate=0.,
+ init_cfg=None):
+
+ super().__init__()
+ self.embed_dims = embed_dims
+ self.window_size = window_size # Wh, Ww
+ self.num_heads = num_heads
+ head_embed_dims = embed_dims // num_heads
+ self.scale = qk_scale or head_embed_dims**-0.5
+ self.init_cfg = init_cfg
+
+ # define a parameter table of relative position bias
+ self.relative_position_bias_table = nn.Parameter(
+ torch.zeros((2 * window_size[0] - 1) * (2 * window_size[1] - 1),
+ num_heads)) # 2*Wh-1 * 2*Ww-1, nH
+
+ # About 2x faster than original impl
+ Wh, Ww = self.window_size
+ rel_index_coords = self.double_step_seq(2 * Ww - 1, Wh, 1, Ww)
+ rel_position_index = rel_index_coords + rel_index_coords.T
+ rel_position_index = rel_position_index.flip(1).contiguous()
+ self.register_buffer('relative_position_index', rel_position_index)
+
+ self.qkv = nn.Linear(embed_dims, embed_dims * 3, bias=qkv_bias)
+ self.attn_drop = nn.Dropout(attn_drop_rate)
+ self.proj = nn.Linear(embed_dims, embed_dims)
+ self.proj_drop = nn.Dropout(proj_drop_rate)
+
+ self.softmax = nn.Softmax(dim=-1)
+
+ def init_weights(self):
+ trunc_normal_(self.relative_position_bias_table, std=0.02)
+
+ def forward(self, x, mask=None):
+ """
+ Args:
+
+ x (tensor): input features with shape of (num_windows*B, N, C)
+ mask (tensor | None, Optional): mask with shape of (num_windows,
+ Wh*Ww, Wh*Ww), value should be between (-inf, 0].
+ """
+ B, N, C = x.shape
+ qkv = self.qkv(x).reshape(B, N, 3, self.num_heads,
+ C // self.num_heads).permute(2, 0, 3, 1, 4)
+ # make torchscript happy (cannot use tensor as tuple)
+ q, k, v = qkv[0], qkv[1], qkv[2]
+
+ q = q * self.scale
+ attn = (q @ k.transpose(-2, -1))
+
+ relative_position_bias = self.relative_position_bias_table[
+ self.relative_position_index.view(-1)].view(
+ self.window_size[0] * self.window_size[1],
+ self.window_size[0] * self.window_size[1],
+ -1) # Wh*Ww,Wh*Ww,nH
+ relative_position_bias = relative_position_bias.permute(
+ 2, 0, 1).contiguous() # nH, Wh*Ww, Wh*Ww
+ attn = attn + relative_position_bias.unsqueeze(0)
+
+ if mask is not None:
+ nW = mask.shape[0]
+ attn = attn.view(B // nW, nW, self.num_heads, N,
+ N) + mask.unsqueeze(1).unsqueeze(0)
+ attn = attn.view(-1, self.num_heads, N, N)
+ attn = self.softmax(attn)
+
+ attn = self.attn_drop(attn)
+
+ x = (attn @ v).transpose(1, 2).reshape(B, N, C)
+ x = self.proj(x)
+ x = self.proj_drop(x)
+ return x
+
+ @staticmethod
+ def double_step_seq(step1, len1, step2, len2):
+ seq1 = torch.arange(0, step1 * len1, step1)
+ seq2 = torch.arange(0, step2 * len2, step2)
+ return (seq1[:, None] + seq2[None, :]).reshape(1, -1)
+
+
+class ShiftWindowMSA(BaseModule):
+ """Shifted Window Multihead Self-Attention Module.
+
+ Args:
+ embed_dims (int): Number of input channels.
+ num_heads (int): Number of attention heads.
+ window_size (int): The height and width of the window.
+ shift_size (int, optional): The shift step of each window towards
+ right-bottom. If zero, act as regular window-msa. Defaults to 0.
+ qkv_bias (bool, optional): If True, add a learnable bias to q, k, v.
+ Default: True
+ qk_scale (float | None, optional): Override default qk scale of
+ head_dim ** -0.5 if set. Defaults: None.
+ attn_drop_rate (float, optional): Dropout ratio of attention weight.
+ Defaults: 0.
+ proj_drop_rate (float, optional): Dropout ratio of output.
+ Defaults: 0.
+ dropout_layer (dict, optional): The dropout_layer used before output.
+ Defaults: dict(type='DropPath', drop_prob=0.).
+ init_cfg (dict, optional): The extra config for initialization.
+ Default: None.
+ """
+
+ def __init__(self,
+ embed_dims,
+ num_heads,
+ window_size,
+ shift_size=0,
+ qkv_bias=True,
+ qk_scale=None,
+ attn_drop_rate=0,
+ proj_drop_rate=0,
+ dropout_layer=dict(type='DropPath', drop_prob=0.),
+ init_cfg=None):
+ super().__init__(init_cfg)
+
+ self.window_size = window_size
+ self.shift_size = shift_size
+ assert 0 <= self.shift_size < self.window_size
+
+ self.w_msa = WindowMSA(
+ embed_dims=embed_dims,
+ num_heads=num_heads,
+ window_size=to_2tuple(window_size),
+ qkv_bias=qkv_bias,
+ qk_scale=qk_scale,
+ attn_drop_rate=attn_drop_rate,
+ proj_drop_rate=proj_drop_rate,
+ init_cfg=None)
+
+ self.drop = build_dropout(dropout_layer)
+
+ def forward(self, query, hw_shape):
+ B, L, C = query.shape
+ H, W = hw_shape
+ assert L == H * W, 'input feature has wrong size'
+ query = query.view(B, H, W, C)
+
+ # pad feature maps to multiples of window size
+ pad_r = (self.window_size - W % self.window_size) % self.window_size
+ pad_b = (self.window_size - H % self.window_size) % self.window_size
+ query = F.pad(query, (0, 0, 0, pad_r, 0, pad_b))
+ H_pad, W_pad = query.shape[1], query.shape[2]
+
+ # cyclic shift
+ if self.shift_size > 0:
+ shifted_query = torch.roll(
+ query,
+ shifts=(-self.shift_size, -self.shift_size),
+ dims=(1, 2))
+
+ # calculate attention mask for SW-MSA
+ img_mask = torch.zeros((1, H_pad, W_pad, 1), device=query.device)
+ h_slices = (slice(0, -self.window_size),
+ slice(-self.window_size,
+ -self.shift_size), slice(-self.shift_size, None))
+ w_slices = (slice(0, -self.window_size),
+ slice(-self.window_size,
+ -self.shift_size), slice(-self.shift_size, None))
+ cnt = 0
+ for h in h_slices:
+ for w in w_slices:
+ img_mask[:, h, w, :] = cnt
+ cnt += 1
+
+ # nW, window_size, window_size, 1
+ mask_windows = self.window_partition(img_mask)
+ mask_windows = mask_windows.view(
+ -1, self.window_size * self.window_size)
+ attn_mask = mask_windows.unsqueeze(1) - mask_windows.unsqueeze(2)
+ attn_mask = attn_mask.masked_fill(attn_mask != 0,
+ float(-100.0)).masked_fill(
+ attn_mask == 0, float(0.0))
+ else:
+ shifted_query = query
+ attn_mask = None
+
+ # nW*B, window_size, window_size, C
+ query_windows = self.window_partition(shifted_query)
+ # nW*B, window_size*window_size, C
+ query_windows = query_windows.view(-1, self.window_size**2, C)
+
+ # W-MSA/SW-MSA (nW*B, window_size*window_size, C)
+ attn_windows = self.w_msa(query_windows, mask=attn_mask)
+
+ # merge windows
+ attn_windows = attn_windows.view(-1, self.window_size,
+ self.window_size, C)
+
+ # B H' W' C
+ shifted_x = self.window_reverse(attn_windows, H_pad, W_pad)
+ # reverse cyclic shift
+ if self.shift_size > 0:
+ x = torch.roll(
+ shifted_x,
+ shifts=(self.shift_size, self.shift_size),
+ dims=(1, 2))
+ else:
+ x = shifted_x
+
+ if pad_r > 0 or pad_b:
+ x = x[:, :H, :W, :].contiguous()
+
+ x = x.view(B, H * W, C)
+
+ x = self.drop(x)
+ return x
+
+ def window_reverse(self, windows, H, W):
+ """
+ Args:
+ windows: (num_windows*B, window_size, window_size, C)
+ H (int): Height of image
+ W (int): Width of image
+ Returns:
+ x: (B, H, W, C)
+ """
+ window_size = self.window_size
+ B = int(windows.shape[0] / (H * W / window_size / window_size))
+ x = windows.view(B, H // window_size, W // window_size, window_size,
+ window_size, -1)
+ x = x.permute(0, 1, 3, 2, 4, 5).contiguous().view(B, H, W, -1)
+ return x
+
+ def window_partition(self, x):
+ """
+ Args:
+ x: (B, H, W, C)
+ Returns:
+ windows: (num_windows*B, window_size, window_size, C)
+ """
+ B, H, W, C = x.shape
+ window_size = self.window_size
+ x = x.view(B, H // window_size, window_size, W // window_size,
+ window_size, C)
+ windows = x.permute(0, 1, 3, 2, 4, 5).contiguous()
+ windows = windows.view(-1, window_size, window_size, C)
+ return windows
+
+
+class SwinBlock(BaseModule):
+ """"
+ Args:
+ embed_dims (int): The feature dimension.
+ num_heads (int): Parallel attention heads.
+ feedforward_channels (int): The hidden dimension for FFNs.
+ window_size (int, optional): The local window scale. Default: 7.
+ shift (bool, optional): whether to shift window or not. Default False.
+ qkv_bias (bool, optional): enable bias for qkv if True. Default: True.
+ qk_scale (float | None, optional): Override default qk scale of
+ head_dim ** -0.5 if set. Default: None.
+ drop_rate (float, optional): Dropout rate. Default: 0.
+ attn_drop_rate (float, optional): Attention dropout rate. Default: 0.
+ drop_path_rate (float, optional): Stochastic depth rate. Default: 0.
+ act_cfg (dict, optional): The config dict of activation function.
+ Default: dict(type='GELU').
+ norm_cfg (dict, optional): The config dict of normalization.
+ Default: dict(type='LN').
+ with_cp (bool, optional): Use checkpoint or not. Using checkpoint
+ will save some memory while slowing down the training speed.
+ Default: False.
+ init_cfg (dict | list | None, optional): The init config.
+ Default: None.
+ """
+
+ def __init__(self,
+ embed_dims,
+ num_heads,
+ feedforward_channels,
+ window_size=7,
+ shift=False,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.,
+ act_cfg=dict(type='GELU'),
+ norm_cfg=dict(type='LN'),
+ with_cp=False,
+ init_cfg=None):
+
+ super(SwinBlock, self).__init__()
+
+ self.init_cfg = init_cfg
+ self.with_cp = with_cp
+
+ self.norm1 = build_norm_layer(norm_cfg, embed_dims)[1]
+ self.attn = ShiftWindowMSA(
+ embed_dims=embed_dims,
+ num_heads=num_heads,
+ window_size=window_size,
+ shift_size=window_size // 2 if shift else 0,
+ qkv_bias=qkv_bias,
+ qk_scale=qk_scale,
+ attn_drop_rate=attn_drop_rate,
+ proj_drop_rate=drop_rate,
+ dropout_layer=dict(type='DropPath', drop_prob=drop_path_rate),
+ init_cfg=None)
+
+ self.norm2 = build_norm_layer(norm_cfg, embed_dims)[1]
+ self.ffn = FFN(
+ embed_dims=embed_dims,
+ feedforward_channels=feedforward_channels,
+ num_fcs=2,
+ ffn_drop=drop_rate,
+ dropout_layer=dict(type='DropPath', drop_prob=drop_path_rate),
+ act_cfg=act_cfg,
+ add_identity=True,
+ init_cfg=None)
+
+ def forward(self, x, hw_shape):
+
+ def _inner_forward(x):
+ identity = x
+ x = self.norm1(x)
+ x = self.attn(x, hw_shape)
+
+ x = x + identity
+
+ identity = x
+ x = self.norm2(x)
+ x = self.ffn(x, identity=identity)
+
+ return x
+
+ if self.with_cp and x.requires_grad:
+ x = cp.checkpoint(_inner_forward, x)
+ else:
+ x = _inner_forward(x)
+
+ return x
+
+
+class SwinBlockSequence(BaseModule):
+ """Implements one stage in Swin Transformer.
+
+ Args:
+ embed_dims (int): The feature dimension.
+ num_heads (int): Parallel attention heads.
+ feedforward_channels (int): The hidden dimension for FFNs.
+ depth (int): The number of blocks in this stage.
+ window_size (int, optional): The local window scale. Default: 7.
+ qkv_bias (bool, optional): enable bias for qkv if True. Default: True.
+ qk_scale (float | None, optional): Override default qk scale of
+ head_dim ** -0.5 if set. Default: None.
+ drop_rate (float, optional): Dropout rate. Default: 0.
+ attn_drop_rate (float, optional): Attention dropout rate. Default: 0.
+ drop_path_rate (float | list[float], optional): Stochastic depth
+ rate. Default: 0.
+ downsample (BaseModule | None, optional): The downsample operation
+ module. Default: None.
+ act_cfg (dict, optional): The config dict of activation function.
+ Default: dict(type='GELU').
+ norm_cfg (dict, optional): The config dict of normalization.
+ Default: dict(type='LN').
+ with_cp (bool, optional): Use checkpoint or not. Using checkpoint
+ will save some memory while slowing down the training speed.
+ Default: False.
+ init_cfg (dict | list | None, optional): The init config.
+ Default: None.
+ """
+
+ def __init__(self,
+ embed_dims,
+ num_heads,
+ feedforward_channels,
+ depth,
+ window_size=7,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.,
+ downsample=None,
+ act_cfg=dict(type='GELU'),
+ norm_cfg=dict(type='LN'),
+ with_cp=False,
+ init_cfg=None):
+ super().__init__(init_cfg=init_cfg)
+
+ if isinstance(drop_path_rate, list):
+ drop_path_rates = drop_path_rate
+ assert len(drop_path_rates) == depth
+ else:
+ drop_path_rates = [deepcopy(drop_path_rate) for _ in range(depth)]
+
+ self.blocks = ModuleList()
+ for i in range(depth):
+ block = SwinBlock(
+ embed_dims=embed_dims,
+ num_heads=num_heads,
+ feedforward_channels=feedforward_channels,
+ window_size=window_size,
+ shift=False if i % 2 == 0 else True,
+ qkv_bias=qkv_bias,
+ qk_scale=qk_scale,
+ drop_rate=drop_rate,
+ attn_drop_rate=attn_drop_rate,
+ drop_path_rate=drop_path_rates[i],
+ act_cfg=act_cfg,
+ norm_cfg=norm_cfg,
+ with_cp=with_cp,
+ init_cfg=None)
+ self.blocks.append(block)
+
+ self.downsample = downsample
+
+ def forward(self, x, hw_shape):
+ for block in self.blocks:
+ x = block(x, hw_shape)
+
+ if self.downsample:
+ x_down, down_hw_shape = self.downsample(x, hw_shape)
+ return x_down, down_hw_shape, x, hw_shape
+ else:
+ return x, hw_shape, x, hw_shape
+
+
+@MODELS.register_module()
+class SwinTransformer(BaseModule):
+ """ Swin Transformer
+ A PyTorch implement of : `Swin Transformer:
+ Hierarchical Vision Transformer using Shifted Windows` -
+ https://arxiv.org/abs/2103.14030
+
+ Inspiration from
+ https://github.com/microsoft/Swin-Transformer
+
+ Args:
+ pretrain_img_size (int | tuple[int]): The size of input image when
+ pretrain. Defaults: 224.
+ in_channels (int): The num of input channels.
+ Defaults: 3.
+ embed_dims (int): The feature dimension. Default: 96.
+ patch_size (int | tuple[int]): Patch size. Default: 4.
+ window_size (int): Window size. Default: 7.
+ mlp_ratio (int): Ratio of mlp hidden dim to embedding dim.
+ Default: 4.
+ depths (tuple[int]): Depths of each Swin Transformer stage.
+ Default: (2, 2, 6, 2).
+ num_heads (tuple[int]): Parallel attention heads of each Swin
+ Transformer stage. Default: (3, 6, 12, 24).
+ strides (tuple[int]): The patch merging or patch embedding stride of
+ each Swin Transformer stage. (In swin, we set kernel size equal to
+ stride.) Default: (4, 2, 2, 2).
+ out_indices (tuple[int]): Output from which stages.
+ Default: (0, 1, 2, 3).
+ qkv_bias (bool, optional): If True, add a learnable bias to query, key,
+ value. Default: True
+ qk_scale (float | None, optional): Override default qk scale of
+ head_dim ** -0.5 if set. Default: None.
+ patch_norm (bool): If add a norm layer for patch embed and patch
+ merging. Default: True.
+ drop_rate (float): Dropout rate. Defaults: 0.
+ attn_drop_rate (float): Attention dropout rate. Default: 0.
+ drop_path_rate (float): Stochastic depth rate. Defaults: 0.1.
+ use_abs_pos_embed (bool): If True, add absolute position embedding to
+ the patch embedding. Defaults: False.
+ act_cfg (dict): Config dict for activation layer.
+ Default: dict(type='GELU').
+ norm_cfg (dict): Config dict for normalization layer at
+ output of backone. Defaults: dict(type='LN').
+ with_cp (bool, optional): Use checkpoint or not. Using checkpoint
+ will save some memory while slowing down the training speed.
+ Default: False.
+ pretrained (str, optional): model pretrained path. Default: None.
+ convert_weights (bool): The flag indicates whether the
+ pre-trained model is from the original repo. We may need
+ to convert some keys to make it compatible.
+ Default: False.
+ frozen_stages (int): Stages to be frozen (stop grad and set eval mode).
+ Default: -1 (-1 means not freezing any parameters).
+ init_cfg (dict, optional): The Config for initialization.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ pretrain_img_size=224,
+ in_channels=3,
+ embed_dims=96,
+ patch_size=4,
+ window_size=7,
+ mlp_ratio=4,
+ depths=(2, 2, 6, 2),
+ num_heads=(3, 6, 12, 24),
+ strides=(4, 2, 2, 2),
+ out_indices=(0, 1, 2, 3),
+ qkv_bias=True,
+ qk_scale=None,
+ patch_norm=True,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.1,
+ use_abs_pos_embed=False,
+ act_cfg=dict(type='GELU'),
+ norm_cfg=dict(type='LN'),
+ with_cp=False,
+ pretrained=None,
+ convert_weights=False,
+ frozen_stages=-1,
+ init_cfg=None):
+ self.convert_weights = convert_weights
+ self.frozen_stages = frozen_stages
+ if isinstance(pretrain_img_size, int):
+ pretrain_img_size = to_2tuple(pretrain_img_size)
+ elif isinstance(pretrain_img_size, tuple):
+ if len(pretrain_img_size) == 1:
+ pretrain_img_size = to_2tuple(pretrain_img_size[0])
+ assert len(pretrain_img_size) == 2, \
+ f'The size of image should have length 1 or 2, ' \
+ f'but got {len(pretrain_img_size)}'
+
+ assert not (init_cfg and pretrained), \
+ 'init_cfg and pretrained cannot be specified at the same time'
+ if isinstance(pretrained, str):
+ warnings.warn('DeprecationWarning: pretrained is deprecated, '
+ 'please use "init_cfg" instead')
+ self.init_cfg = dict(type='Pretrained', checkpoint=pretrained)
+ elif pretrained is None:
+ self.init_cfg = init_cfg
+ else:
+ raise TypeError('pretrained must be a str or None')
+
+ super(SwinTransformer, self).__init__(init_cfg=init_cfg)
+
+ num_layers = len(depths)
+ self.out_indices = out_indices
+ self.use_abs_pos_embed = use_abs_pos_embed
+
+ assert strides[0] == patch_size, 'Use non-overlapping patch embed.'
+
+ self.patch_embed = PatchEmbed(
+ in_channels=in_channels,
+ embed_dims=embed_dims,
+ conv_type='Conv2d',
+ kernel_size=patch_size,
+ stride=strides[0],
+ norm_cfg=norm_cfg if patch_norm else None,
+ init_cfg=None)
+
+ if self.use_abs_pos_embed:
+ patch_row = pretrain_img_size[0] // patch_size
+ patch_col = pretrain_img_size[1] // patch_size
+ num_patches = patch_row * patch_col
+ self.absolute_pos_embed = nn.Parameter(
+ torch.zeros((1, num_patches, embed_dims)))
+
+ self.drop_after_pos = nn.Dropout(p=drop_rate)
+
+ # set stochastic depth decay rule
+ total_depth = sum(depths)
+ dpr = [
+ x.item() for x in torch.linspace(0, drop_path_rate, total_depth)
+ ]
+
+ self.stages = ModuleList()
+ in_channels = embed_dims
+ for i in range(num_layers):
+ if i < num_layers - 1:
+ downsample = PatchMerging(
+ in_channels=in_channels,
+ out_channels=2 * in_channels,
+ stride=strides[i + 1],
+ norm_cfg=norm_cfg if patch_norm else None,
+ init_cfg=None)
+ else:
+ downsample = None
+
+ stage = SwinBlockSequence(
+ embed_dims=in_channels,
+ num_heads=num_heads[i],
+ feedforward_channels=mlp_ratio * in_channels,
+ depth=depths[i],
+ window_size=window_size,
+ qkv_bias=qkv_bias,
+ qk_scale=qk_scale,
+ drop_rate=drop_rate,
+ attn_drop_rate=attn_drop_rate,
+ drop_path_rate=dpr[sum(depths[:i]):sum(depths[:i + 1])],
+ downsample=downsample,
+ act_cfg=act_cfg,
+ norm_cfg=norm_cfg,
+ with_cp=with_cp,
+ init_cfg=None)
+ self.stages.append(stage)
+ if downsample:
+ in_channels = downsample.out_channels
+
+ self.num_features = [int(embed_dims * 2**i) for i in range(num_layers)]
+ # Add a norm layer for each output
+ for i in out_indices:
+ layer = build_norm_layer(norm_cfg, self.num_features[i])[1]
+ layer_name = f'norm{i}'
+ self.add_module(layer_name, layer)
+
+ def train(self, mode=True):
+ """Convert the model into training mode while keep layers freezed."""
+ super(SwinTransformer, self).train(mode)
+ self._freeze_stages()
+
+ def _freeze_stages(self):
+ if self.frozen_stages >= 0:
+ self.patch_embed.eval()
+ for param in self.patch_embed.parameters():
+ param.requires_grad = False
+ if self.use_abs_pos_embed:
+ self.absolute_pos_embed.requires_grad = False
+ self.drop_after_pos.eval()
+
+ for i in range(1, self.frozen_stages + 1):
+
+ if (i - 1) in self.out_indices:
+ norm_layer = getattr(self, f'norm{i-1}')
+ norm_layer.eval()
+ for param in norm_layer.parameters():
+ param.requires_grad = False
+
+ m = self.stages[i - 1]
+ m.eval()
+ for param in m.parameters():
+ param.requires_grad = False
+
+ def init_weights(self):
+ logger = MMLogger.get_current_instance()
+ if self.init_cfg is None:
+ logger.warn(f'No pre-trained weights for '
+ f'{self.__class__.__name__}, '
+ f'training start from scratch')
+ if self.use_abs_pos_embed:
+ trunc_normal_(self.absolute_pos_embed, std=0.02)
+ for m in self.modules():
+ if isinstance(m, nn.Linear):
+ trunc_normal_init(m, std=.02, bias=0.)
+ elif isinstance(m, nn.LayerNorm):
+ constant_init(m, 1.0)
+ else:
+ assert 'checkpoint' in self.init_cfg, f'Only support ' \
+ f'specify `Pretrained` in ' \
+ f'`init_cfg` in ' \
+ f'{self.__class__.__name__} '
+ ckpt = CheckpointLoader.load_checkpoint(
+ self.init_cfg.checkpoint, logger=logger, map_location='cpu')
+ if 'state_dict' in ckpt:
+ _state_dict = ckpt['state_dict']
+ elif 'model' in ckpt:
+ _state_dict = ckpt['model']
+ else:
+ _state_dict = ckpt
+ if self.convert_weights:
+ # supported loading weight from original repo,
+ _state_dict = swin_converter(_state_dict)
+
+ state_dict = OrderedDict()
+ for k, v in _state_dict.items():
+ if k.startswith('backbone.'):
+ state_dict[k[9:]] = v
+
+ # strip prefix of state_dict
+ if list(state_dict.keys())[0].startswith('module.'):
+ state_dict = {k[7:]: v for k, v in state_dict.items()}
+
+ # reshape absolute position embedding
+ if state_dict.get('absolute_pos_embed') is not None:
+ absolute_pos_embed = state_dict['absolute_pos_embed']
+ N1, L, C1 = absolute_pos_embed.size()
+ N2, C2, H, W = self.absolute_pos_embed.size()
+ if N1 != N2 or C1 != C2 or L != H * W:
+ logger.warning('Error in loading absolute_pos_embed, pass')
+ else:
+ state_dict['absolute_pos_embed'] = absolute_pos_embed.view(
+ N2, H, W, C2).permute(0, 3, 1, 2).contiguous()
+
+ # interpolate position bias table if needed
+ relative_position_bias_table_keys = [
+ k for k in state_dict.keys()
+ if 'relative_position_bias_table' in k
+ ]
+ for table_key in relative_position_bias_table_keys:
+ table_pretrained = state_dict[table_key]
+ table_current = self.state_dict()[table_key]
+ L1, nH1 = table_pretrained.size()
+ L2, nH2 = table_current.size()
+ if nH1 != nH2:
+ logger.warning(f'Error in loading {table_key}, pass')
+ elif L1 != L2:
+ S1 = int(L1**0.5)
+ S2 = int(L2**0.5)
+ table_pretrained_resized = F.interpolate(
+ table_pretrained.permute(1, 0).reshape(1, nH1, S1, S1),
+ size=(S2, S2),
+ mode='bicubic')
+ state_dict[table_key] = table_pretrained_resized.view(
+ nH2, L2).permute(1, 0).contiguous()
+
+ # load state_dict
+ self.load_state_dict(state_dict, False)
+
+ def forward(self, x):
+ x, hw_shape = self.patch_embed(x)
+
+ if self.use_abs_pos_embed:
+ x = x + self.absolute_pos_embed
+ x = self.drop_after_pos(x)
+
+ outs = []
+ for i, stage in enumerate(self.stages):
+ x, hw_shape, out, out_hw_shape = stage(x, hw_shape)
+ if i in self.out_indices:
+ norm_layer = getattr(self, f'norm{i}')
+ out = norm_layer(out)
+ out = out.view(-1, *out_hw_shape,
+ self.num_features[i]).permute(0, 3, 1,
+ 2).contiguous()
+ outs.append(out)
+
+ return outs
+
+
+def swin_converter(ckpt):
+
+ new_ckpt = OrderedDict()
+
+ def correct_unfold_reduction_order(x):
+ out_channel, in_channel = x.shape
+ x = x.reshape(out_channel, 4, in_channel // 4)
+ x = x[:, [0, 2, 1, 3], :].transpose(1,
+ 2).reshape(out_channel, in_channel)
+ return x
+
+ def correct_unfold_norm_order(x):
+ in_channel = x.shape[0]
+ x = x.reshape(4, in_channel // 4)
+ x = x[[0, 2, 1, 3], :].transpose(0, 1).reshape(in_channel)
+ return x
+
+ for k, v in ckpt.items():
+ if k.startswith('head'):
+ continue
+ elif k.startswith('layers'):
+ new_v = v
+ if 'attn.' in k:
+ new_k = k.replace('attn.', 'attn.w_msa.')
+ elif 'mlp.' in k:
+ if 'mlp.fc1.' in k:
+ new_k = k.replace('mlp.fc1.', 'ffn.layers.0.0.')
+ elif 'mlp.fc2.' in k:
+ new_k = k.replace('mlp.fc2.', 'ffn.layers.1.')
+ else:
+ new_k = k.replace('mlp.', 'ffn.')
+ elif 'downsample' in k:
+ new_k = k
+ if 'reduction.' in k:
+ new_v = correct_unfold_reduction_order(v)
+ elif 'norm.' in k:
+ new_v = correct_unfold_norm_order(v)
+ else:
+ new_k = k
+ new_k = new_k.replace('layers', 'stages', 1)
+ elif k.startswith('patch_embed'):
+ new_v = v
+ if 'proj' in k:
+ new_k = k.replace('proj', 'projection')
+ else:
+ new_k = k
+ else:
+ new_v = v
+ new_k = k
+
+ new_ckpt['backbone.' + new_k] = new_v
+
+ return new_ckpt
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/trident_resnet.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/trident_resnet.py
new file mode 100644
index 0000000000000000000000000000000000000000..22c76354522ff8533b094df6858ec361ba400c1e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/backbones/trident_resnet.py
@@ -0,0 +1,298 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+import torch.utils.checkpoint as cp
+from mmcv.cnn import build_conv_layer, build_norm_layer
+from mmengine.model import BaseModule
+from torch.nn.modules.utils import _pair
+
+from mmdet.models.backbones.resnet import Bottleneck, ResNet
+from mmdet.registry import MODELS
+
+
+class TridentConv(BaseModule):
+ """Trident Convolution Module.
+
+ Args:
+ in_channels (int): Number of channels in input.
+ out_channels (int): Number of channels in output.
+ kernel_size (int): Size of convolution kernel.
+ stride (int, optional): Convolution stride. Default: 1.
+ trident_dilations (tuple[int, int, int], optional): Dilations of
+ different trident branch. Default: (1, 2, 3).
+ test_branch_idx (int, optional): In inference, all 3 branches will
+ be used if `test_branch_idx==-1`, otherwise only branch with
+ index `test_branch_idx` will be used. Default: 1.
+ bias (bool, optional): Whether to use bias in convolution or not.
+ Default: False.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None
+ """
+
+ def __init__(self,
+ in_channels,
+ out_channels,
+ kernel_size,
+ stride=1,
+ trident_dilations=(1, 2, 3),
+ test_branch_idx=1,
+ bias=False,
+ init_cfg=None):
+ super(TridentConv, self).__init__(init_cfg)
+ self.num_branch = len(trident_dilations)
+ self.with_bias = bias
+ self.test_branch_idx = test_branch_idx
+ self.stride = _pair(stride)
+ self.kernel_size = _pair(kernel_size)
+ self.paddings = _pair(trident_dilations)
+ self.dilations = trident_dilations
+ self.in_channels = in_channels
+ self.out_channels = out_channels
+ self.bias = bias
+
+ self.weight = nn.Parameter(
+ torch.Tensor(out_channels, in_channels, *self.kernel_size))
+ if bias:
+ self.bias = nn.Parameter(torch.Tensor(out_channels))
+ else:
+ self.bias = None
+
+ def extra_repr(self):
+ tmpstr = f'in_channels={self.in_channels}'
+ tmpstr += f', out_channels={self.out_channels}'
+ tmpstr += f', kernel_size={self.kernel_size}'
+ tmpstr += f', num_branch={self.num_branch}'
+ tmpstr += f', test_branch_idx={self.test_branch_idx}'
+ tmpstr += f', stride={self.stride}'
+ tmpstr += f', paddings={self.paddings}'
+ tmpstr += f', dilations={self.dilations}'
+ tmpstr += f', bias={self.bias}'
+ return tmpstr
+
+ def forward(self, inputs):
+ if self.training or self.test_branch_idx == -1:
+ outputs = [
+ F.conv2d(input, self.weight, self.bias, self.stride, padding,
+ dilation) for input, dilation, padding in zip(
+ inputs, self.dilations, self.paddings)
+ ]
+ else:
+ assert len(inputs) == 1
+ outputs = [
+ F.conv2d(inputs[0], self.weight, self.bias, self.stride,
+ self.paddings[self.test_branch_idx],
+ self.dilations[self.test_branch_idx])
+ ]
+
+ return outputs
+
+
+# Since TridentNet is defined over ResNet50 and ResNet101, here we
+# only support TridentBottleneckBlock.
+class TridentBottleneck(Bottleneck):
+ """BottleBlock for TridentResNet.
+
+ Args:
+ trident_dilations (tuple[int, int, int]): Dilations of different
+ trident branch.
+ test_branch_idx (int): In inference, all 3 branches will be used
+ if `test_branch_idx==-1`, otherwise only branch with index
+ `test_branch_idx` will be used.
+ concat_output (bool): Whether to concat the output list to a Tensor.
+ `True` only in the last Block.
+ """
+
+ def __init__(self, trident_dilations, test_branch_idx, concat_output,
+ **kwargs):
+
+ super(TridentBottleneck, self).__init__(**kwargs)
+ self.trident_dilations = trident_dilations
+ self.num_branch = len(trident_dilations)
+ self.concat_output = concat_output
+ self.test_branch_idx = test_branch_idx
+ self.conv2 = TridentConv(
+ self.planes,
+ self.planes,
+ kernel_size=3,
+ stride=self.conv2_stride,
+ bias=False,
+ trident_dilations=self.trident_dilations,
+ test_branch_idx=test_branch_idx,
+ init_cfg=dict(
+ type='Kaiming',
+ distribution='uniform',
+ mode='fan_in',
+ override=dict(name='conv2')))
+
+ def forward(self, x):
+
+ def _inner_forward(x):
+ num_branch = (
+ self.num_branch
+ if self.training or self.test_branch_idx == -1 else 1)
+ identity = x
+ if not isinstance(x, list):
+ x = (x, ) * num_branch
+ identity = x
+ if self.downsample is not None:
+ identity = [self.downsample(b) for b in x]
+
+ out = [self.conv1(b) for b in x]
+ out = [self.norm1(b) for b in out]
+ out = [self.relu(b) for b in out]
+
+ if self.with_plugins:
+ for k in range(len(out)):
+ out[k] = self.forward_plugin(out[k],
+ self.after_conv1_plugin_names)
+
+ out = self.conv2(out)
+ out = [self.norm2(b) for b in out]
+ out = [self.relu(b) for b in out]
+ if self.with_plugins:
+ for k in range(len(out)):
+ out[k] = self.forward_plugin(out[k],
+ self.after_conv2_plugin_names)
+
+ out = [self.conv3(b) for b in out]
+ out = [self.norm3(b) for b in out]
+
+ if self.with_plugins:
+ for k in range(len(out)):
+ out[k] = self.forward_plugin(out[k],
+ self.after_conv3_plugin_names)
+
+ out = [
+ out_b + identity_b for out_b, identity_b in zip(out, identity)
+ ]
+ return out
+
+ if self.with_cp and x.requires_grad:
+ out = cp.checkpoint(_inner_forward, x)
+ else:
+ out = _inner_forward(x)
+
+ out = [self.relu(b) for b in out]
+ if self.concat_output:
+ out = torch.cat(out, dim=0)
+ return out
+
+
+def make_trident_res_layer(block,
+ inplanes,
+ planes,
+ num_blocks,
+ stride=1,
+ trident_dilations=(1, 2, 3),
+ style='pytorch',
+ with_cp=False,
+ conv_cfg=None,
+ norm_cfg=dict(type='BN'),
+ dcn=None,
+ plugins=None,
+ test_branch_idx=-1):
+ """Build Trident Res Layers."""
+
+ downsample = None
+ if stride != 1 or inplanes != planes * block.expansion:
+ downsample = []
+ conv_stride = stride
+ downsample.extend([
+ build_conv_layer(
+ conv_cfg,
+ inplanes,
+ planes * block.expansion,
+ kernel_size=1,
+ stride=conv_stride,
+ bias=False),
+ build_norm_layer(norm_cfg, planes * block.expansion)[1]
+ ])
+ downsample = nn.Sequential(*downsample)
+
+ layers = []
+ for i in range(num_blocks):
+ layers.append(
+ block(
+ inplanes=inplanes,
+ planes=planes,
+ stride=stride if i == 0 else 1,
+ trident_dilations=trident_dilations,
+ downsample=downsample if i == 0 else None,
+ style=style,
+ with_cp=with_cp,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ dcn=dcn,
+ plugins=plugins,
+ test_branch_idx=test_branch_idx,
+ concat_output=True if i == num_blocks - 1 else False))
+ inplanes = planes * block.expansion
+ return nn.Sequential(*layers)
+
+
+@MODELS.register_module()
+class TridentResNet(ResNet):
+ """The stem layer, stage 1 and stage 2 in Trident ResNet are identical to
+ ResNet, while in stage 3, Trident BottleBlock is utilized to replace the
+ normal BottleBlock to yield trident output. Different branch shares the
+ convolution weight but uses different dilations to achieve multi-scale
+ output.
+
+ / stage3(b0) \
+ x - stem - stage1 - stage2 - stage3(b1) - output
+ \ stage3(b2) /
+
+ Args:
+ depth (int): Depth of resnet, from {50, 101, 152}.
+ num_branch (int): Number of branches in TridentNet.
+ test_branch_idx (int): In inference, all 3 branches will be used
+ if `test_branch_idx==-1`, otherwise only branch with index
+ `test_branch_idx` will be used.
+ trident_dilations (tuple[int]): Dilations of different trident branch.
+ len(trident_dilations) should be equal to num_branch.
+ """ # noqa
+
+ def __init__(self, depth, num_branch, test_branch_idx, trident_dilations,
+ **kwargs):
+
+ assert num_branch == len(trident_dilations)
+ assert depth in (50, 101, 152)
+ super(TridentResNet, self).__init__(depth, **kwargs)
+ assert self.num_stages == 3
+ self.test_branch_idx = test_branch_idx
+ self.num_branch = num_branch
+
+ last_stage_idx = self.num_stages - 1
+ stride = self.strides[last_stage_idx]
+ dilation = trident_dilations
+ dcn = self.dcn if self.stage_with_dcn[last_stage_idx] else None
+ if self.plugins is not None:
+ stage_plugins = self.make_stage_plugins(self.plugins,
+ last_stage_idx)
+ else:
+ stage_plugins = None
+ planes = self.base_channels * 2**last_stage_idx
+ res_layer = make_trident_res_layer(
+ TridentBottleneck,
+ inplanes=(self.block.expansion * self.base_channels *
+ 2**(last_stage_idx - 1)),
+ planes=planes,
+ num_blocks=self.stage_blocks[last_stage_idx],
+ stride=stride,
+ trident_dilations=dilation,
+ style=self.style,
+ with_cp=self.with_cp,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ dcn=dcn,
+ plugins=stage_plugins,
+ test_branch_idx=self.test_branch_idx)
+
+ layer_name = f'layer{last_stage_idx + 1}'
+
+ self.__setattr__(layer_name, res_layer)
+ self.res_layers.pop(last_stage_idx)
+ self.res_layers.insert(last_stage_idx, layer_name)
+
+ self._freeze_stages()
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/data_preprocessors/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/data_preprocessors/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..201a1da6a4f320a17cea9c65d5c102bfdd7700d8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/data_preprocessors/__init__.py
@@ -0,0 +1,13 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .data_preprocessor import (BatchFixedSizePad, BatchResize,
+ BatchSyncRandomResize, BoxInstDataPreprocessor,
+ DetDataPreprocessor,
+ MultiBranchDataPreprocessor)
+from .reid_data_preprocessor import ReIDDataPreprocessor
+from .track_data_preprocessor import TrackDataPreprocessor
+
+__all__ = [
+ 'DetDataPreprocessor', 'BatchSyncRandomResize', 'BatchFixedSizePad',
+ 'MultiBranchDataPreprocessor', 'BatchResize', 'BoxInstDataPreprocessor',
+ 'TrackDataPreprocessor', 'ReIDDataPreprocessor'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/data_preprocessors/data_preprocessor.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/data_preprocessors/data_preprocessor.py
new file mode 100644
index 0000000000000000000000000000000000000000..55b5c35b3a4888c95c6646df3fa080347afe4704
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/data_preprocessors/data_preprocessor.py
@@ -0,0 +1,793 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import random
+from numbers import Number
+from typing import List, Optional, Sequence, Tuple, Union
+
+import numpy as np
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmengine.dist import barrier, broadcast, get_dist_info
+from mmengine.logging import MessageHub
+from mmengine.model import BaseDataPreprocessor, ImgDataPreprocessor
+from mmengine.structures import PixelData
+from mmengine.utils import is_seq_of
+from torch import Tensor
+
+from mmdet.models.utils import unfold_wo_center
+from mmdet.models.utils.misc import samplelist_boxtype2tensor
+from mmdet.registry import MODELS
+from mmdet.structures import DetDataSample
+from mmdet.structures.mask import BitmapMasks
+from mmdet.utils import ConfigType
+
+try:
+ import skimage
+except ImportError:
+ skimage = None
+
+
+@MODELS.register_module()
+class DetDataPreprocessor(ImgDataPreprocessor):
+ """Image pre-processor for detection tasks.
+
+ Comparing with the :class:`mmengine.ImgDataPreprocessor`,
+
+ 1. It supports batch augmentations.
+ 2. It will additionally append batch_input_shape and pad_shape
+ to data_samples considering the object detection task.
+
+ It provides the data pre-processing as follows
+
+ - Collate and move data to the target device.
+ - Pad inputs to the maximum size of current batch with defined
+ ``pad_value``. The padding size can be divisible by a defined
+ ``pad_size_divisor``
+ - Stack inputs to batch_inputs.
+ - Convert inputs from bgr to rgb if the shape of input is (3, H, W).
+ - Normalize image with defined std and mean.
+ - Do batch augmentations during training.
+
+ Args:
+ mean (Sequence[Number], optional): The pixel mean of R, G, B channels.
+ Defaults to None.
+ std (Sequence[Number], optional): The pixel standard deviation of
+ R, G, B channels. Defaults to None.
+ pad_size_divisor (int): The size of padded image should be
+ divisible by ``pad_size_divisor``. Defaults to 1.
+ pad_value (Number): The padded pixel value. Defaults to 0.
+ pad_mask (bool): Whether to pad instance masks. Defaults to False.
+ mask_pad_value (int): The padded pixel value for instance masks.
+ Defaults to 0.
+ pad_seg (bool): Whether to pad semantic segmentation maps.
+ Defaults to False.
+ seg_pad_value (int): The padded pixel value for semantic
+ segmentation maps. Defaults to 255.
+ bgr_to_rgb (bool): whether to convert image from BGR to RGB.
+ Defaults to False.
+ rgb_to_bgr (bool): whether to convert image from RGB to RGB.
+ Defaults to False.
+ boxtype2tensor (bool): Whether to convert the ``BaseBoxes`` type of
+ bboxes data to ``Tensor`` type. Defaults to True.
+ non_blocking (bool): Whether block current process
+ when transferring data to device. Defaults to False.
+ batch_augments (list[dict], optional): Batch-level augmentations
+ """
+
+ def __init__(self,
+ mean: Sequence[Number] = None,
+ std: Sequence[Number] = None,
+ pad_size_divisor: int = 1,
+ pad_value: Union[float, int] = 0,
+ pad_mask: bool = False,
+ mask_pad_value: int = 0,
+ pad_seg: bool = False,
+ seg_pad_value: int = 255,
+ bgr_to_rgb: bool = False,
+ rgb_to_bgr: bool = False,
+ boxtype2tensor: bool = True,
+ non_blocking: Optional[bool] = False,
+ batch_augments: Optional[List[dict]] = None):
+ super().__init__(
+ mean=mean,
+ std=std,
+ pad_size_divisor=pad_size_divisor,
+ pad_value=pad_value,
+ bgr_to_rgb=bgr_to_rgb,
+ rgb_to_bgr=rgb_to_bgr,
+ non_blocking=non_blocking)
+ if batch_augments is not None:
+ self.batch_augments = nn.ModuleList(
+ [MODELS.build(aug) for aug in batch_augments])
+ else:
+ self.batch_augments = None
+ self.pad_mask = pad_mask
+ self.mask_pad_value = mask_pad_value
+ self.pad_seg = pad_seg
+ self.seg_pad_value = seg_pad_value
+ self.boxtype2tensor = boxtype2tensor
+
+ def forward(self, data: dict, training: bool = False) -> dict:
+ """Perform normalization,padding and bgr2rgb conversion based on
+ ``BaseDataPreprocessor``.
+
+ Args:
+ data (dict): Data sampled from dataloader.
+ training (bool): Whether to enable training time augmentation.
+
+ Returns:
+ dict: Data in the same format as the model input.
+ """
+ batch_pad_shape = self._get_pad_shape(data)
+ data = super().forward(data=data, training=training)
+ inputs, data_samples = data['inputs'], data['data_samples']
+
+ if data_samples is not None:
+ # NOTE the batched image size information may be useful, e.g.
+ # in DETR, this is needed for the construction of masks, which is
+ # then used for the transformer_head.
+ batch_input_shape = tuple(inputs[0].size()[-2:])
+ for data_sample, pad_shape in zip(data_samples, batch_pad_shape):
+ data_sample.set_metainfo({
+ 'batch_input_shape': batch_input_shape,
+ 'pad_shape': pad_shape
+ })
+
+ if self.boxtype2tensor:
+ samplelist_boxtype2tensor(data_samples)
+
+ if self.pad_mask and training:
+ self.pad_gt_masks(data_samples)
+
+ if self.pad_seg and training:
+ self.pad_gt_sem_seg(data_samples)
+
+ if training and self.batch_augments is not None:
+ for batch_aug in self.batch_augments:
+ inputs, data_samples = batch_aug(inputs, data_samples)
+
+ return {'inputs': inputs, 'data_samples': data_samples}
+
+ def _get_pad_shape(self, data: dict) -> List[tuple]:
+ """Get the pad_shape of each image based on data and
+ pad_size_divisor."""
+ _batch_inputs = data['inputs']
+ # Process data with `pseudo_collate`.
+ if is_seq_of(_batch_inputs, torch.Tensor):
+ batch_pad_shape = []
+ for ori_input in _batch_inputs:
+ pad_h = int(
+ np.ceil(ori_input.shape[1] /
+ self.pad_size_divisor)) * self.pad_size_divisor
+ pad_w = int(
+ np.ceil(ori_input.shape[2] /
+ self.pad_size_divisor)) * self.pad_size_divisor
+ batch_pad_shape.append((pad_h, pad_w))
+ # Process data with `default_collate`.
+ elif isinstance(_batch_inputs, torch.Tensor):
+ assert _batch_inputs.dim() == 4, (
+ 'The input of `ImgDataPreprocessor` should be a NCHW tensor '
+ 'or a list of tensor, but got a tensor with shape: '
+ f'{_batch_inputs.shape}')
+ pad_h = int(
+ np.ceil(_batch_inputs.shape[2] /
+ self.pad_size_divisor)) * self.pad_size_divisor
+ pad_w = int(
+ np.ceil(_batch_inputs.shape[3] /
+ self.pad_size_divisor)) * self.pad_size_divisor
+ batch_pad_shape = [(pad_h, pad_w)] * _batch_inputs.shape[0]
+ else:
+ raise TypeError('Output of `cast_data` should be a dict '
+ 'or a tuple with inputs and data_samples, but got'
+ f'{type(data)}: {data}')
+ return batch_pad_shape
+
+ def pad_gt_masks(self,
+ batch_data_samples: Sequence[DetDataSample]) -> None:
+ """Pad gt_masks to shape of batch_input_shape."""
+ if 'masks' in batch_data_samples[0].gt_instances:
+ for data_samples in batch_data_samples:
+ masks = data_samples.gt_instances.masks
+ data_samples.gt_instances.masks = masks.pad(
+ data_samples.batch_input_shape,
+ pad_val=self.mask_pad_value)
+
+ def pad_gt_sem_seg(self,
+ batch_data_samples: Sequence[DetDataSample]) -> None:
+ """Pad gt_sem_seg to shape of batch_input_shape."""
+ if 'gt_sem_seg' in batch_data_samples[0]:
+ for data_samples in batch_data_samples:
+ gt_sem_seg = data_samples.gt_sem_seg.sem_seg
+ h, w = gt_sem_seg.shape[-2:]
+ pad_h, pad_w = data_samples.batch_input_shape
+ gt_sem_seg = F.pad(
+ gt_sem_seg,
+ pad=(0, max(pad_w - w, 0), 0, max(pad_h - h, 0)),
+ mode='constant',
+ value=self.seg_pad_value)
+ data_samples.gt_sem_seg = PixelData(sem_seg=gt_sem_seg)
+
+
+@MODELS.register_module()
+class BatchSyncRandomResize(nn.Module):
+ """Batch random resize which synchronizes the random size across ranks.
+
+ Args:
+ random_size_range (tuple): The multi-scale random range during
+ multi-scale training.
+ interval (int): The iter interval of change
+ image size. Defaults to 10.
+ size_divisor (int): Image size divisible factor.
+ Defaults to 32.
+ """
+
+ def __init__(self,
+ random_size_range: Tuple[int, int],
+ interval: int = 10,
+ size_divisor: int = 32) -> None:
+ super().__init__()
+ self.rank, self.world_size = get_dist_info()
+ self._input_size = None
+ self._random_size_range = (round(random_size_range[0] / size_divisor),
+ round(random_size_range[1] / size_divisor))
+ self._interval = interval
+ self._size_divisor = size_divisor
+
+ def forward(
+ self, inputs: Tensor, data_samples: List[DetDataSample]
+ ) -> Tuple[Tensor, List[DetDataSample]]:
+ """resize a batch of images and bboxes to shape ``self._input_size``"""
+ h, w = inputs.shape[-2:]
+ if self._input_size is None:
+ self._input_size = (h, w)
+ scale_y = self._input_size[0] / h
+ scale_x = self._input_size[1] / w
+ if scale_x != 1 or scale_y != 1:
+ inputs = F.interpolate(
+ inputs,
+ size=self._input_size,
+ mode='bilinear',
+ align_corners=False)
+ for data_sample in data_samples:
+ img_shape = (int(data_sample.img_shape[0] * scale_y),
+ int(data_sample.img_shape[1] * scale_x))
+ pad_shape = (int(data_sample.pad_shape[0] * scale_y),
+ int(data_sample.pad_shape[1] * scale_x))
+ data_sample.set_metainfo({
+ 'img_shape': img_shape,
+ 'pad_shape': pad_shape,
+ 'batch_input_shape': self._input_size
+ })
+ data_sample.gt_instances.bboxes[
+ ...,
+ 0::2] = data_sample.gt_instances.bboxes[...,
+ 0::2] * scale_x
+ data_sample.gt_instances.bboxes[
+ ...,
+ 1::2] = data_sample.gt_instances.bboxes[...,
+ 1::2] * scale_y
+ if 'ignored_instances' in data_sample:
+ data_sample.ignored_instances.bboxes[
+ ..., 0::2] = data_sample.ignored_instances.bboxes[
+ ..., 0::2] * scale_x
+ data_sample.ignored_instances.bboxes[
+ ..., 1::2] = data_sample.ignored_instances.bboxes[
+ ..., 1::2] * scale_y
+ message_hub = MessageHub.get_current_instance()
+ if (message_hub.get_info('iter') + 1) % self._interval == 0:
+ self._input_size = self._get_random_size(
+ aspect_ratio=float(w / h), device=inputs.device)
+ return inputs, data_samples
+
+ def _get_random_size(self, aspect_ratio: float,
+ device: torch.device) -> Tuple[int, int]:
+ """Randomly generate a shape in ``_random_size_range`` and broadcast to
+ all ranks."""
+ tensor = torch.LongTensor(2).to(device)
+ if self.rank == 0:
+ size = random.randint(*self._random_size_range)
+ size = (self._size_divisor * size,
+ self._size_divisor * int(aspect_ratio * size))
+ tensor[0] = size[0]
+ tensor[1] = size[1]
+ barrier()
+ broadcast(tensor, 0)
+ input_size = (tensor[0].item(), tensor[1].item())
+ return input_size
+
+
+@MODELS.register_module()
+class BatchFixedSizePad(nn.Module):
+ """Fixed size padding for batch images.
+
+ Args:
+ size (Tuple[int, int]): Fixed padding size. Expected padding
+ shape (h, w). Defaults to None.
+ img_pad_value (int): The padded pixel value for images.
+ Defaults to 0.
+ pad_mask (bool): Whether to pad instance masks. Defaults to False.
+ mask_pad_value (int): The padded pixel value for instance masks.
+ Defaults to 0.
+ pad_seg (bool): Whether to pad semantic segmentation maps.
+ Defaults to False.
+ seg_pad_value (int): The padded pixel value for semantic
+ segmentation maps. Defaults to 255.
+ """
+
+ def __init__(self,
+ size: Tuple[int, int],
+ img_pad_value: int = 0,
+ pad_mask: bool = False,
+ mask_pad_value: int = 0,
+ pad_seg: bool = False,
+ seg_pad_value: int = 255) -> None:
+ super().__init__()
+ self.size = size
+ self.pad_mask = pad_mask
+ self.pad_seg = pad_seg
+ self.img_pad_value = img_pad_value
+ self.mask_pad_value = mask_pad_value
+ self.seg_pad_value = seg_pad_value
+
+ def forward(
+ self,
+ inputs: Tensor,
+ data_samples: Optional[List[dict]] = None
+ ) -> Tuple[Tensor, Optional[List[dict]]]:
+ """Pad image, instance masks, segmantic segmentation maps."""
+ src_h, src_w = inputs.shape[-2:]
+ dst_h, dst_w = self.size
+
+ if src_h >= dst_h and src_w >= dst_w:
+ return inputs, data_samples
+
+ inputs = F.pad(
+ inputs,
+ pad=(0, max(0, dst_w - src_w), 0, max(0, dst_h - src_h)),
+ mode='constant',
+ value=self.img_pad_value)
+
+ if data_samples is not None:
+ # update batch_input_shape
+ for data_sample in data_samples:
+ data_sample.set_metainfo({
+ 'batch_input_shape': (dst_h, dst_w),
+ 'pad_shape': (dst_h, dst_w)
+ })
+
+ if self.pad_mask:
+ for data_sample in data_samples:
+ masks = data_sample.gt_instances.masks
+ data_sample.gt_instances.masks = masks.pad(
+ (dst_h, dst_w), pad_val=self.mask_pad_value)
+
+ if self.pad_seg:
+ for data_sample in data_samples:
+ gt_sem_seg = data_sample.gt_sem_seg.sem_seg
+ h, w = gt_sem_seg.shape[-2:]
+ gt_sem_seg = F.pad(
+ gt_sem_seg,
+ pad=(0, max(0, dst_w - w), 0, max(0, dst_h - h)),
+ mode='constant',
+ value=self.seg_pad_value)
+ data_sample.gt_sem_seg = PixelData(sem_seg=gt_sem_seg)
+
+ return inputs, data_samples
+
+
+@MODELS.register_module()
+class MultiBranchDataPreprocessor(BaseDataPreprocessor):
+ """DataPreprocessor wrapper for multi-branch data.
+
+ Take semi-supervised object detection as an example, assume that
+ the ratio of labeled data and unlabeled data in a batch is 1:2,
+ `sup` indicates the branch where the labeled data is augmented,
+ `unsup_teacher` and `unsup_student` indicate the branches where
+ the unlabeled data is augmented by different pipeline.
+
+ The input format of multi-branch data is shown as below :
+
+ .. code-block:: none
+ {
+ 'inputs':
+ {
+ 'sup': [Tensor, None, None],
+ 'unsup_teacher': [None, Tensor, Tensor],
+ 'unsup_student': [None, Tensor, Tensor],
+ },
+ 'data_sample':
+ {
+ 'sup': [DetDataSample, None, None],
+ 'unsup_teacher': [None, DetDataSample, DetDataSample],
+ 'unsup_student': [NOne, DetDataSample, DetDataSample],
+ }
+ }
+
+ The format of multi-branch data
+ after filtering None is shown as below :
+
+ .. code-block:: none
+ {
+ 'inputs':
+ {
+ 'sup': [Tensor],
+ 'unsup_teacher': [Tensor, Tensor],
+ 'unsup_student': [Tensor, Tensor],
+ },
+ 'data_sample':
+ {
+ 'sup': [DetDataSample],
+ 'unsup_teacher': [DetDataSample, DetDataSample],
+ 'unsup_student': [DetDataSample, DetDataSample],
+ }
+ }
+
+ In order to reuse `DetDataPreprocessor` for the data
+ from different branches, the format of multi-branch data
+ grouped by branch is as below :
+
+ .. code-block:: none
+ {
+ 'sup':
+ {
+ 'inputs': [Tensor]
+ 'data_sample': [DetDataSample, DetDataSample]
+ },
+ 'unsup_teacher':
+ {
+ 'inputs': [Tensor, Tensor]
+ 'data_sample': [DetDataSample, DetDataSample]
+ },
+ 'unsup_student':
+ {
+ 'inputs': [Tensor, Tensor]
+ 'data_sample': [DetDataSample, DetDataSample]
+ },
+ }
+
+ After preprocessing data from different branches,
+ the multi-branch data needs to be reformatted as:
+
+ .. code-block:: none
+ {
+ 'inputs':
+ {
+ 'sup': [Tensor],
+ 'unsup_teacher': [Tensor, Tensor],
+ 'unsup_student': [Tensor, Tensor],
+ },
+ 'data_sample':
+ {
+ 'sup': [DetDataSample],
+ 'unsup_teacher': [DetDataSample, DetDataSample],
+ 'unsup_student': [DetDataSample, DetDataSample],
+ }
+ }
+
+ Args:
+ data_preprocessor (:obj:`ConfigDict` or dict): Config of
+ :class:`DetDataPreprocessor` to process the input data.
+ """
+
+ def __init__(self, data_preprocessor: ConfigType) -> None:
+ super().__init__()
+ self.data_preprocessor = MODELS.build(data_preprocessor)
+
+ def forward(self, data: dict, training: bool = False) -> dict:
+ """Perform normalization,padding and bgr2rgb conversion based on
+ ``BaseDataPreprocessor`` for multi-branch data.
+
+ Args:
+ data (dict): Data sampled from dataloader.
+ training (bool): Whether to enable training time augmentation.
+
+ Returns:
+ dict:
+
+ - 'inputs' (Dict[str, obj:`torch.Tensor`]): The forward data of
+ models from different branches.
+ - 'data_sample' (Dict[str, obj:`DetDataSample`]): The annotation
+ info of the sample from different branches.
+ """
+
+ if training is False:
+ return self.data_preprocessor(data, training)
+
+ # Filter out branches with a value of None
+ for key in data.keys():
+ for branch in data[key].keys():
+ data[key][branch] = list(
+ filter(lambda x: x is not None, data[key][branch]))
+
+ # Group data by branch
+ multi_branch_data = {}
+ for key in data.keys():
+ for branch in data[key].keys():
+ if multi_branch_data.get(branch, None) is None:
+ multi_branch_data[branch] = {key: data[key][branch]}
+ elif multi_branch_data[branch].get(key, None) is None:
+ multi_branch_data[branch][key] = data[key][branch]
+ else:
+ multi_branch_data[branch][key].append(data[key][branch])
+
+ # Preprocess data from different branches
+ for branch, _data in multi_branch_data.items():
+ multi_branch_data[branch] = self.data_preprocessor(_data, training)
+
+ # Format data by inputs and data_samples
+ format_data = {}
+ for branch in multi_branch_data.keys():
+ for key in multi_branch_data[branch].keys():
+ if format_data.get(key, None) is None:
+ format_data[key] = {branch: multi_branch_data[branch][key]}
+ elif format_data[key].get(branch, None) is None:
+ format_data[key][branch] = multi_branch_data[branch][key]
+ else:
+ format_data[key][branch].append(
+ multi_branch_data[branch][key])
+
+ return format_data
+
+ @property
+ def device(self):
+ return self.data_preprocessor.device
+
+ def to(self, device: Optional[Union[int, torch.device]], *args,
+ **kwargs) -> nn.Module:
+ """Overrides this method to set the :attr:`device`
+
+ Args:
+ device (int or torch.device, optional): The desired device of the
+ parameters and buffers in this module.
+
+ Returns:
+ nn.Module: The model itself.
+ """
+
+ return self.data_preprocessor.to(device, *args, **kwargs)
+
+ def cuda(self, *args, **kwargs) -> nn.Module:
+ """Overrides this method to set the :attr:`device`
+
+ Returns:
+ nn.Module: The model itself.
+ """
+
+ return self.data_preprocessor.cuda(*args, **kwargs)
+
+ def cpu(self, *args, **kwargs) -> nn.Module:
+ """Overrides this method to set the :attr:`device`
+
+ Returns:
+ nn.Module: The model itself.
+ """
+
+ return self.data_preprocessor.cpu(*args, **kwargs)
+
+
+@MODELS.register_module()
+class BatchResize(nn.Module):
+ """Batch resize during training. This implementation is modified from
+ https://github.com/Purkialo/CrowdDet/blob/master/lib/data/CrowdHuman.py.
+
+ It provides the data pre-processing as follows:
+ - A batch of all images will pad to a uniform size and stack them into
+ a torch.Tensor by `DetDataPreprocessor`.
+ - `BatchFixShapeResize` resize all images to the target size.
+ - Padding images to make sure the size of image can be divisible by
+ ``pad_size_divisor``.
+
+ Args:
+ scale (tuple): Images scales for resizing.
+ pad_size_divisor (int): Image size divisible factor.
+ Defaults to 1.
+ pad_value (Number): The padded pixel value. Defaults to 0.
+ """
+
+ def __init__(
+ self,
+ scale: tuple,
+ pad_size_divisor: int = 1,
+ pad_value: Union[float, int] = 0,
+ ) -> None:
+ super().__init__()
+ self.min_size = min(scale)
+ self.max_size = max(scale)
+ self.pad_size_divisor = pad_size_divisor
+ self.pad_value = pad_value
+
+ def forward(
+ self, inputs: Tensor, data_samples: List[DetDataSample]
+ ) -> Tuple[Tensor, List[DetDataSample]]:
+ """resize a batch of images and bboxes."""
+
+ batch_height, batch_width = inputs.shape[-2:]
+ target_height, target_width, scale = self.get_target_size(
+ batch_height, batch_width)
+
+ inputs = F.interpolate(
+ inputs,
+ size=(target_height, target_width),
+ mode='bilinear',
+ align_corners=False)
+
+ inputs = self.get_padded_tensor(inputs, self.pad_value)
+
+ if data_samples is not None:
+ batch_input_shape = tuple(inputs.size()[-2:])
+ for data_sample in data_samples:
+ img_shape = [
+ int(scale * _) for _ in list(data_sample.img_shape)
+ ]
+ data_sample.set_metainfo({
+ 'img_shape': tuple(img_shape),
+ 'batch_input_shape': batch_input_shape,
+ 'pad_shape': batch_input_shape,
+ 'scale_factor': (scale, scale)
+ })
+
+ data_sample.gt_instances.bboxes *= scale
+ data_sample.ignored_instances.bboxes *= scale
+
+ return inputs, data_samples
+
+ def get_target_size(self, height: int,
+ width: int) -> Tuple[int, int, float]:
+ """Get the target size of a batch of images based on data and scale."""
+ im_size_min = np.min([height, width])
+ im_size_max = np.max([height, width])
+ scale = self.min_size / im_size_min
+ if scale * im_size_max > self.max_size:
+ scale = self.max_size / im_size_max
+ target_height, target_width = int(round(height * scale)), int(
+ round(width * scale))
+ return target_height, target_width, scale
+
+ def get_padded_tensor(self, tensor: Tensor, pad_value: int) -> Tensor:
+ """Pad images according to pad_size_divisor."""
+ assert tensor.ndim == 4
+ target_height, target_width = tensor.shape[-2], tensor.shape[-1]
+ divisor = self.pad_size_divisor
+ padded_height = (target_height + divisor - 1) // divisor * divisor
+ padded_width = (target_width + divisor - 1) // divisor * divisor
+ padded_tensor = torch.ones([
+ tensor.shape[0], tensor.shape[1], padded_height, padded_width
+ ]) * pad_value
+ padded_tensor = padded_tensor.type_as(tensor)
+ padded_tensor[:, :, :target_height, :target_width] = tensor
+ return padded_tensor
+
+
+@MODELS.register_module()
+class BoxInstDataPreprocessor(DetDataPreprocessor):
+ """Pseudo mask pre-processor for BoxInst.
+
+ Comparing with the :class:`mmdet.DetDataPreprocessor`,
+
+ 1. It generates masks using box annotations.
+ 2. It computes the images color similarity in LAB color space.
+
+ Args:
+ mask_stride (int): The mask output stride in boxinst. Defaults to 4.
+ pairwise_size (int): The size of neighborhood for each pixel.
+ Defaults to 3.
+ pairwise_dilation (int): The dilation of neighborhood for each pixel.
+ Defaults to 2.
+ pairwise_color_thresh (float): The thresh of image color similarity.
+ Defaults to 0.3.
+ bottom_pixels_removed (int): The length of removed pixels in bottom.
+ It is caused by the annotation error in coco dataset.
+ Defaults to 10.
+ """
+
+ def __init__(self,
+ *arg,
+ mask_stride: int = 4,
+ pairwise_size: int = 3,
+ pairwise_dilation: int = 2,
+ pairwise_color_thresh: float = 0.3,
+ bottom_pixels_removed: int = 10,
+ **kwargs) -> None:
+ super().__init__(*arg, **kwargs)
+ self.mask_stride = mask_stride
+ self.pairwise_size = pairwise_size
+ self.pairwise_dilation = pairwise_dilation
+ self.pairwise_color_thresh = pairwise_color_thresh
+ self.bottom_pixels_removed = bottom_pixels_removed
+
+ if skimage is None:
+ raise RuntimeError('skimage is not installed,\
+ please install it by: pip install scikit-image')
+
+ def get_images_color_similarity(self, inputs: Tensor,
+ image_masks: Tensor) -> Tensor:
+ """Compute the image color similarity in LAB color space."""
+ assert inputs.dim() == 4
+ assert inputs.size(0) == 1
+
+ unfolded_images = unfold_wo_center(
+ inputs,
+ kernel_size=self.pairwise_size,
+ dilation=self.pairwise_dilation)
+ diff = inputs[:, :, None] - unfolded_images
+ similarity = torch.exp(-torch.norm(diff, dim=1) * 0.5)
+
+ unfolded_weights = unfold_wo_center(
+ image_masks[None, None],
+ kernel_size=self.pairwise_size,
+ dilation=self.pairwise_dilation)
+ unfolded_weights = torch.max(unfolded_weights, dim=1)[0]
+
+ return similarity * unfolded_weights
+
+ def forward(self, data: dict, training: bool = False) -> dict:
+ """Get pseudo mask labels using color similarity."""
+ det_data = super().forward(data, training)
+ inputs, data_samples = det_data['inputs'], det_data['data_samples']
+
+ if training:
+ # get image masks and remove bottom pixels
+ b_img_h, b_img_w = data_samples[0].batch_input_shape
+ img_masks = []
+ for i in range(inputs.shape[0]):
+ img_h, img_w = data_samples[i].img_shape
+ img_mask = inputs.new_ones((img_h, img_w))
+ pixels_removed = int(self.bottom_pixels_removed *
+ float(img_h) / float(b_img_h))
+ if pixels_removed > 0:
+ img_mask[-pixels_removed:, :] = 0
+ pad_w = b_img_w - img_w
+ pad_h = b_img_h - img_h
+ img_mask = F.pad(img_mask, (0, pad_w, 0, pad_h), 'constant',
+ 0.)
+ img_masks.append(img_mask)
+ img_masks = torch.stack(img_masks, dim=0)
+ start = int(self.mask_stride // 2)
+ img_masks = img_masks[:, start::self.mask_stride,
+ start::self.mask_stride]
+
+ # Get origin rgb image for color similarity
+ ori_imgs = inputs * self.std + self.mean
+ downsampled_imgs = F.avg_pool2d(
+ ori_imgs.float(),
+ kernel_size=self.mask_stride,
+ stride=self.mask_stride,
+ padding=0)
+
+ # Compute color similarity for pseudo mask generation
+ for im_i, data_sample in enumerate(data_samples):
+ # TODO: Support rgb2lab in mmengine?
+ images_lab = skimage.color.rgb2lab(
+ downsampled_imgs[im_i].byte().permute(1, 2,
+ 0).cpu().numpy())
+ images_lab = torch.as_tensor(
+ images_lab, device=ori_imgs.device, dtype=torch.float32)
+ images_lab = images_lab.permute(2, 0, 1)[None]
+ images_color_similarity = self.get_images_color_similarity(
+ images_lab, img_masks[im_i])
+ pairwise_mask = (images_color_similarity >=
+ self.pairwise_color_thresh).float()
+
+ per_im_bboxes = data_sample.gt_instances.bboxes
+ if per_im_bboxes.shape[0] > 0:
+ per_im_masks = []
+ for per_box in per_im_bboxes:
+ mask_full = torch.zeros((b_img_h, b_img_w),
+ device=self.device).float()
+ mask_full[int(per_box[1]):int(per_box[3] + 1),
+ int(per_box[0]):int(per_box[2] + 1)] = 1.0
+ per_im_masks.append(mask_full)
+ per_im_masks = torch.stack(per_im_masks, dim=0)
+ pairwise_masks = torch.cat(
+ [pairwise_mask for _ in range(per_im_bboxes.shape[0])],
+ dim=0)
+ else:
+ per_im_masks = torch.zeros((0, b_img_h, b_img_w))
+ pairwise_masks = torch.zeros(
+ (0, self.pairwise_size**2 - 1, b_img_h, b_img_w))
+
+ # TODO: Support BitmapMasks with tensor?
+ data_sample.gt_instances.masks = BitmapMasks(
+ per_im_masks.cpu().numpy(), b_img_h, b_img_w)
+ data_sample.gt_instances.pairwise_masks = pairwise_masks
+ return {'inputs': inputs, 'data_samples': data_samples}
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/data_preprocessors/reid_data_preprocessor.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/data_preprocessors/reid_data_preprocessor.py
new file mode 100644
index 0000000000000000000000000000000000000000..3d0a1d45d97ba350e8845c6620f3b73f05545e61
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/data_preprocessors/reid_data_preprocessor.py
@@ -0,0 +1,216 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+from numbers import Number
+from typing import Optional, Sequence
+
+import torch
+import torch.nn.functional as F
+from mmengine.model import BaseDataPreprocessor, stack_batch
+
+from mmdet.registry import MODELS
+
+try:
+ import mmpretrain
+ from mmpretrain.models.utils.batch_augments import RandomBatchAugment
+ from mmpretrain.structures import (batch_label_to_onehot, cat_batch_labels,
+ tensor_split)
+except ImportError:
+ mmpretrain = None
+
+
+def stack_batch_scores(elements, device=None):
+ """Stack the ``score`` of a batch of :obj:`LabelData` to a tensor.
+
+ Args:
+ elements (List[LabelData]): A batch of :obj`LabelData`.
+ device (torch.device, optional): The output device of the batch label.
+ Defaults to None.
+ Returns:
+ torch.Tensor: The stacked score tensor.
+ """
+ item = elements[0]
+ if 'score' not in item._data_fields:
+ return None
+
+ batch_score = torch.stack([element.score for element in elements])
+ if device is not None:
+ batch_score = batch_score.to(device)
+ return batch_score
+
+
+@MODELS.register_module()
+class ReIDDataPreprocessor(BaseDataPreprocessor):
+ """Image pre-processor for classification tasks.
+
+ Comparing with the :class:`mmengine.model.ImgDataPreprocessor`,
+
+ 1. It won't do normalization if ``mean`` is not specified.
+ 2. It does normalization and color space conversion after stacking batch.
+ 3. It supports batch augmentations like mixup and cutmix.
+
+ It provides the data pre-processing as follows
+
+ - Collate and move data to the target device.
+ - Pad inputs to the maximum size of current batch with defined
+ ``pad_value``. The padding size can be divisible by a defined
+ ``pad_size_divisor``
+ - Stack inputs to batch_inputs.
+ - Convert inputs from bgr to rgb if the shape of input is (3, H, W).
+ - Normalize image with defined std and mean.
+ - Do batch augmentations like Mixup and Cutmix during training.
+
+ Args:
+ mean (Sequence[Number], optional): The pixel mean of R, G, B channels.
+ Defaults to None.
+ std (Sequence[Number], optional): The pixel standard deviation of
+ R, G, B channels. Defaults to None.
+ pad_size_divisor (int): The size of padded image should be
+ divisible by ``pad_size_divisor``. Defaults to 1.
+ pad_value (Number): The padded pixel value. Defaults to 0.
+ to_rgb (bool): whether to convert image from BGR to RGB.
+ Defaults to False.
+ to_onehot (bool): Whether to generate one-hot format gt-labels and set
+ to data samples. Defaults to False.
+ num_classes (int, optional): The number of classes. Defaults to None.
+ batch_augments (dict, optional): The batch augmentations settings,
+ including "augments" and "probs". For more details, see
+ :class:`mmpretrain.models.RandomBatchAugment`.
+ """
+
+ def __init__(self,
+ mean: Sequence[Number] = None,
+ std: Sequence[Number] = None,
+ pad_size_divisor: int = 1,
+ pad_value: Number = 0,
+ to_rgb: bool = False,
+ to_onehot: bool = False,
+ num_classes: Optional[int] = None,
+ batch_augments: Optional[dict] = None):
+ if mmpretrain is None:
+ raise RuntimeError('Please run "pip install openmim" and '
+ 'run "mim install mmpretrain" to '
+ 'install mmpretrain first.')
+ super().__init__()
+ self.pad_size_divisor = pad_size_divisor
+ self.pad_value = pad_value
+ self.to_rgb = to_rgb
+ self.to_onehot = to_onehot
+ self.num_classes = num_classes
+
+ if mean is not None:
+ assert std is not None, 'To enable the normalization in ' \
+ 'preprocessing, please specify both `mean` and `std`.'
+ # Enable the normalization in preprocessing.
+ self._enable_normalize = True
+ self.register_buffer('mean',
+ torch.tensor(mean).view(-1, 1, 1), False)
+ self.register_buffer('std',
+ torch.tensor(std).view(-1, 1, 1), False)
+ else:
+ self._enable_normalize = False
+
+ if batch_augments is not None:
+ self.batch_augments = RandomBatchAugment(**batch_augments)
+ if not self.to_onehot:
+ from mmengine.logging import MMLogger
+ MMLogger.get_current_instance().info(
+ 'Because batch augmentations are enabled, the data '
+ 'preprocessor automatically enables the `to_onehot` '
+ 'option to generate one-hot format labels.')
+ self.to_onehot = True
+ else:
+ self.batch_augments = None
+
+ def forward(self, data: dict, training: bool = False) -> dict:
+ """Perform normalization, padding, bgr2rgb conversion and batch
+ augmentation based on ``BaseDataPreprocessor``.
+
+ Args:
+ data (dict): data sampled from dataloader.
+ training (bool): Whether to enable training time augmentation.
+
+ Returns:
+ dict: Data in the same format as the model input.
+ """
+ inputs = self.cast_data(data['inputs'])
+
+ if isinstance(inputs, torch.Tensor):
+ # The branch if use `default_collate` as the collate_fn in the
+ # dataloader.
+
+ # ------ To RGB ------
+ if self.to_rgb and inputs.size(1) == 3:
+ inputs = inputs.flip(1)
+
+ # -- Normalization ---
+ inputs = inputs.float()
+ if self._enable_normalize:
+ inputs = (inputs - self.mean) / self.std
+
+ # ------ Padding -----
+ if self.pad_size_divisor > 1:
+ h, w = inputs.shape[-2:]
+
+ target_h = math.ceil(
+ h / self.pad_size_divisor) * self.pad_size_divisor
+ target_w = math.ceil(
+ w / self.pad_size_divisor) * self.pad_size_divisor
+ pad_h = target_h - h
+ pad_w = target_w - w
+ inputs = F.pad(inputs, (0, pad_w, 0, pad_h), 'constant',
+ self.pad_value)
+ else:
+ # The branch if use `pseudo_collate` as the collate_fn in the
+ # dataloader.
+
+ processed_inputs = []
+ for input_ in inputs:
+ # ------ To RGB ------
+ if self.to_rgb and input_.size(0) == 3:
+ input_ = input_.flip(0)
+
+ # -- Normalization ---
+ input_ = input_.float()
+ if self._enable_normalize:
+ input_ = (input_ - self.mean) / self.std
+
+ processed_inputs.append(input_)
+ # Combine padding and stack
+ inputs = stack_batch(processed_inputs, self.pad_size_divisor,
+ self.pad_value)
+
+ data_samples = data.get('data_samples', None)
+ sample_item = data_samples[0] if data_samples is not None else None
+ if 'gt_label' in sample_item:
+ gt_labels = [sample.gt_label for sample in data_samples]
+ gt_labels_tensor = [gt_label.label for gt_label in gt_labels]
+ batch_label, label_indices = cat_batch_labels(gt_labels_tensor)
+ batch_label = batch_label.to(self.device)
+
+ batch_score = stack_batch_scores(gt_labels, device=self.device)
+ if batch_score is None and self.to_onehot:
+ assert batch_label is not None, \
+ 'Cannot generate onehot format labels because no labels.'
+ num_classes = self.num_classes or data_samples[0].get(
+ 'num_classes')
+ assert num_classes is not None, \
+ 'Cannot generate one-hot format labels because not set ' \
+ '`num_classes` in `data_preprocessor`.'
+ batch_score = batch_label_to_onehot(batch_label, label_indices,
+ num_classes)
+
+ # ----- Batch Augmentations ----
+ if training and self.batch_augments is not None:
+ inputs, batch_score = self.batch_augments(inputs, batch_score)
+
+ # ----- scatter labels and scores to data samples ---
+ if batch_label is not None:
+ for sample, label in zip(
+ data_samples, tensor_split(batch_label,
+ label_indices)):
+ sample.set_gt_label(label)
+ if batch_score is not None:
+ for sample, score in zip(data_samples, batch_score):
+ sample.set_gt_score(score)
+
+ return {'inputs': inputs, 'data_samples': data_samples}
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/data_preprocessors/track_data_preprocessor.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/data_preprocessors/track_data_preprocessor.py
new file mode 100644
index 0000000000000000000000000000000000000000..40a65b8eaebacdaddd574768fbb00e8c5a072d85
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/data_preprocessors/track_data_preprocessor.py
@@ -0,0 +1,266 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, List, Optional, Sequence, Union
+
+import numpy as np
+import torch
+import torch.nn.functional as F
+from mmengine.model.utils import stack_batch
+
+from mmdet.models.utils.misc import samplelist_boxtype2tensor
+from mmdet.registry import MODELS
+from mmdet.structures import TrackDataSample
+from mmdet.structures.mask import BitmapMasks
+from .data_preprocessor import DetDataPreprocessor
+
+
+@MODELS.register_module()
+class TrackDataPreprocessor(DetDataPreprocessor):
+ """Image pre-processor for tracking tasks.
+
+ Accepts the data sampled by the dataloader, and preprocesses
+ it into the format of the model input. ``TrackDataPreprocessor``
+ provides the tracking data pre-processing as follows:
+
+ - Collate and move data to the target device.
+ - Pad inputs to the maximum size of current batch with defined
+ ``pad_value``. The padding size can be divisible by a defined
+ ``pad_size_divisor``
+ - Stack inputs to inputs.
+ - Convert inputs from bgr to rgb if the shape of input is (1, 3, H, W).
+ - Normalize image with defined std and mean.
+ - Do batch augmentations during training.
+ - Record the information of ``batch_input_shape`` and ``pad_shape``.
+
+ Args:
+ mean (Sequence[Number], optional): The pixel mean of R, G, B
+ channels. Defaults to None.
+ std (Sequence[Number], optional): The pixel standard deviation of
+ R, G, B channels. Defaults to None.
+ pad_size_divisor (int): The size of padded image should be
+ divisible by ``pad_size_divisor``. Defaults to 1.
+ pad_value (Number): The padded pixel value. Defaults to 0.
+ pad_mask (bool): Whether to pad instance masks. Defaults to False.
+ mask_pad_value (int): The padded pixel value for instance masks.
+ Defaults to 0.
+ bgr_to_rgb (bool): whether to convert image from BGR to RGB.
+ Defaults to False.
+ rgb_to_bgr (bool): whether to convert image from RGB to RGB.
+ Defaults to False.
+ use_det_processor: (bool): whether to use DetDataPreprocessor
+ in training phrase. This is mainly for some tracking models
+ fed into one image rather than a group of image in training.
+ Defaults to False.
+ . boxtype2tensor (bool): Whether to convert the ``BaseBoxes`` type of
+ bboxes data to ``Tensor`` type. Defaults to True.
+ batch_augments (list[dict], optional): Batch-level augmentations
+ """
+
+ def __init__(self,
+ mean: Optional[Sequence[Union[float, int]]] = None,
+ std: Optional[Sequence[Union[float, int]]] = None,
+ use_det_processor: bool = False,
+ **kwargs):
+ super().__init__(mean=mean, std=std, **kwargs)
+ self.use_det_processor = use_det_processor
+ if mean is not None and not self.use_det_processor:
+ # overwrite the ``register_bufffer`` in ``ImgDataPreprocessor``
+ # since the shape of ``mean`` and ``std`` in tracking tasks must be
+ # (T, C, H, W), which T is the temporal length of the video.
+ self.register_buffer('mean',
+ torch.tensor(mean).view(1, -1, 1, 1), False)
+ self.register_buffer('std',
+ torch.tensor(std).view(1, -1, 1, 1), False)
+
+ def forward(self, data: dict, training: bool = False) -> Dict:
+ """Perform normalization,padding and bgr2rgb conversion based on
+ ``TrackDataPreprocessor``.
+
+ Args:
+ data (dict): data sampled from dataloader.
+ training (bool): Whether to enable training time augmentation.
+
+ Returns:
+ Tuple[Dict[str, List[torch.Tensor]], OptSampleList]: Data in the
+ same format as the model input.
+ """
+ if self.use_det_processor and training:
+ batch_pad_shape = self._get_pad_shape(data)
+ else:
+ batch_pad_shape = self._get_track_pad_shape(data)
+
+ data = self.cast_data(data)
+ imgs, data_samples = data['inputs'], data['data_samples']
+
+ if self.use_det_processor and training:
+ assert imgs[0].dim() == 3, \
+ 'Only support the 3 dims when use detpreprocessor in training'
+ if self._channel_conversion:
+ imgs = [_img[[2, 1, 0], ...] for _img in imgs]
+ # Convert to `float`
+ imgs = [_img.float() for _img in imgs]
+ if self._enable_normalize:
+ imgs = [(_img - self.mean) / self.std for _img in imgs]
+ inputs = stack_batch(imgs, self.pad_size_divisor, self.pad_value)
+ else:
+ assert imgs[0].dim() == 4, \
+ 'Only support the 4 dims when use trackprocessor in training'
+ # The shape of imgs[0] is (T, C, H, W).
+ channel = imgs[0].size(1)
+ if self._channel_conversion and channel == 3:
+ imgs = [_img[:, [2, 1, 0], ...] for _img in imgs]
+ # change to `float`
+ imgs = [_img.float() for _img in imgs]
+ if self._enable_normalize:
+ imgs = [(_img - self.mean) / self.std for _img in imgs]
+ inputs = stack_track_batch(imgs, self.pad_size_divisor,
+ self.pad_value)
+
+ if data_samples is not None:
+ # NOTE the batched image size information may be useful, e.g.
+ # in DETR, this is needed for the construction of masks, which is
+ # then used for the transformer_head.
+ batch_input_shape = tuple(inputs.size()[-2:])
+ if self.use_det_processor and training:
+ for data_sample, pad_shape in zip(data_samples,
+ batch_pad_shape):
+ data_sample.set_metainfo({
+ 'batch_input_shape': batch_input_shape,
+ 'pad_shape': pad_shape
+ })
+ if self.boxtype2tensor:
+ samplelist_boxtype2tensor(data_samples)
+ if self.pad_mask:
+ self.pad_gt_masks(data_samples)
+ else:
+ for track_data_sample, pad_shapes in zip(
+ data_samples, batch_pad_shape):
+ for i in range(len(track_data_sample)):
+ det_data_sample = track_data_sample[i]
+ det_data_sample.set_metainfo({
+ 'batch_input_shape': batch_input_shape,
+ 'pad_shape': pad_shapes[i]
+ })
+ if self.pad_mask and training:
+ self.pad_track_gt_masks(data_samples)
+
+ if training and self.batch_augments is not None:
+ for batch_aug in self.batch_augments:
+ if self.use_det_processor and training:
+ inputs, data_samples = batch_aug(inputs, data_samples)
+ else:
+ # we only support T==1 when using batch augments.
+ # Only yolox need batch_aug, and yolox can only process
+ # (N, C, H, W) shape.
+ # The shape of `inputs` is (N, T, C, H, W), hence, we use
+ # inputs[:, 0] to change the shape to (N, C, H, W).
+ assert inputs.size(1) == 1 and len(
+ data_samples[0]
+ ) == 1, 'Only support the number of sequence images equals to 1 when using batch augment.' # noqa: E501
+ det_data_samples = [
+ track_data_sample[0]
+ for track_data_sample in data_samples
+ ]
+ aug_inputs, aug_det_samples = batch_aug(
+ inputs[:, 0], det_data_samples)
+ inputs = aug_inputs.unsqueeze(1)
+ for track_data_sample, det_sample in zip(
+ data_samples, aug_det_samples):
+ track_data_sample.video_data_samples = [det_sample]
+
+ # Note: inputs may contain large number of frames, so we must make
+ # sure that the mmeory is contiguous for stable forward
+ inputs = inputs.contiguous()
+
+ return dict(inputs=inputs, data_samples=data_samples)
+
+ def _get_track_pad_shape(self, data: dict) -> Dict[str, List]:
+ """Get the pad_shape of each image based on data and pad_size_divisor.
+
+ Args:
+ data (dict): Data sampled from dataloader.
+
+ Returns:
+ Dict[str, List]: The shape of padding.
+ """
+ batch_pad_shape = dict()
+ batch_pad_shape = []
+ for imgs in data['inputs']:
+ # The sequence images in one sample among a batch have the same
+ # original shape
+ pad_h = int(np.ceil(imgs.shape[-2] /
+ self.pad_size_divisor)) * self.pad_size_divisor
+ pad_w = int(np.ceil(imgs.shape[-1] /
+ self.pad_size_divisor)) * self.pad_size_divisor
+ pad_shapes = [(pad_h, pad_w)] * imgs.size(0)
+ batch_pad_shape.append(pad_shapes)
+ return batch_pad_shape
+
+ def pad_track_gt_masks(self,
+ data_samples: Sequence[TrackDataSample]) -> None:
+ """Pad gt_masks to shape of batch_input_shape."""
+ if 'masks' in data_samples[0][0].get('gt_instances', None):
+ for track_data_sample in data_samples:
+ for i in range(len(track_data_sample)):
+ det_data_sample = track_data_sample[i]
+ masks = det_data_sample.gt_instances.masks
+ # TODO: whether to use BitmapMasks
+ assert isinstance(masks, BitmapMasks)
+ batch_input_shape = det_data_sample.batch_input_shape
+ det_data_sample.gt_instances.masks = masks.pad(
+ batch_input_shape, pad_val=self.mask_pad_value)
+
+
+def stack_track_batch(tensors: List[torch.Tensor],
+ pad_size_divisor: int = 0,
+ pad_value: Union[int, float] = 0) -> torch.Tensor:
+ """Stack multiple tensors to form a batch and pad the images to the max
+ shape use the right bottom padding mode in these images. If
+ ``pad_size_divisor > 0``, add padding to ensure the common height and width
+ is divisible by ``pad_size_divisor``. The difference between this function
+ and ``stack_batch`` in MMEngine is that this function can process batch
+ sequence images with shape (N, T, C, H, W).
+
+ Args:
+ tensors (List[Tensor]): The input multiple tensors. each is a
+ TCHW 4D-tensor. T denotes the number of key/reference frames.
+ pad_size_divisor (int): If ``pad_size_divisor > 0``, add padding
+ to ensure the common height and width is divisible by
+ ``pad_size_divisor``. This depends on the model, and many
+ models need a divisibility of 32. Defaults to 0
+ pad_value (int, float): The padding value. Defaults to 0
+
+ Returns:
+ Tensor: The NTCHW 5D-tensor. N denotes the batch size.
+ """
+ assert isinstance(tensors, list), \
+ f'Expected input type to be list, but got {type(tensors)}'
+ assert len(set([tensor.ndim for tensor in tensors])) == 1, \
+ f'Expected the dimensions of all tensors must be the same, ' \
+ f'but got {[tensor.ndim for tensor in tensors]}'
+ assert tensors[0].ndim == 4, f'Expected tensor dimension to be 4, ' \
+ f'but got {tensors[0].ndim}'
+ assert len(set([tensor.shape[0] for tensor in tensors])) == 1, \
+ f'Expected the channels of all tensors must be the same, ' \
+ f'but got {[tensor.shape[0] for tensor in tensors]}'
+
+ tensor_sizes = [(tensor.shape[-2], tensor.shape[-1]) for tensor in tensors]
+ max_size = np.stack(tensor_sizes).max(0)
+
+ if pad_size_divisor > 1:
+ # the last two dims are H,W, both subject to divisibility requirement
+ max_size = (
+ max_size +
+ (pad_size_divisor - 1)) // pad_size_divisor * pad_size_divisor
+
+ padded_samples = []
+ for tensor in tensors:
+ padding_size = [
+ 0, max_size[-1] - tensor.shape[-1], 0,
+ max_size[-2] - tensor.shape[-2]
+ ]
+ if sum(padding_size) == 0:
+ padded_samples.append(tensor)
+ else:
+ padded_samples.append(F.pad(tensor, padding_size, value=pad_value))
+
+ return torch.stack(padded_samples, dim=0)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..c9b55ec2a4230a741e9a2c696ec434bf9cc8bafa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/__init__.py
@@ -0,0 +1,72 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .anchor_free_head import AnchorFreeHead
+from .anchor_head import AnchorHead
+from .atss_head import ATSSHead
+from .atss_vlfusion_head import ATSSVLFusionHead
+from .autoassign_head import AutoAssignHead
+from .boxinst_head import BoxInstBboxHead, BoxInstMaskHead
+from .cascade_rpn_head import CascadeRPNHead, StageCascadeRPNHead
+from .centernet_head import CenterNetHead
+from .centernet_update_head import CenterNetUpdateHead
+from .centripetal_head import CentripetalHead
+from .condinst_head import CondInstBboxHead, CondInstMaskHead
+from .conditional_detr_head import ConditionalDETRHead
+from .corner_head import CornerHead
+from .dab_detr_head import DABDETRHead
+from .ddod_head import DDODHead
+from .ddq_detr_head import DDQDETRHead
+from .deformable_detr_head import DeformableDETRHead
+from .detr_head import DETRHead
+from .dino_head import DINOHead
+from .embedding_rpn_head import EmbeddingRPNHead
+from .fcos_head import FCOSHead
+from .fovea_head import FoveaHead
+from .free_anchor_retina_head import FreeAnchorRetinaHead
+from .fsaf_head import FSAFHead
+from .ga_retina_head import GARetinaHead
+from .ga_rpn_head import GARPNHead
+from .gfl_head import GFLHead
+from .grounding_dino_head import GroundingDINOHead
+from .guided_anchor_head import FeatureAdaption, GuidedAnchorHead
+from .lad_head import LADHead
+from .ld_head import LDHead
+from .mask2former_head import Mask2FormerHead
+from .maskformer_head import MaskFormerHead
+from .nasfcos_head import NASFCOSHead
+from .paa_head import PAAHead
+from .pisa_retinanet_head import PISARetinaHead
+from .pisa_ssd_head import PISASSDHead
+from .reppoints_head import RepPointsHead
+from .retina_head import RetinaHead
+from .retina_sepbn_head import RetinaSepBNHead
+from .rpn_head import RPNHead
+from .rtmdet_head import RTMDetHead, RTMDetSepBNHead
+from .rtmdet_ins_head import RTMDetInsHead, RTMDetInsSepBNHead
+from .sabl_retina_head import SABLRetinaHead
+from .solo_head import DecoupledSOLOHead, DecoupledSOLOLightHead, SOLOHead
+from .solov2_head import SOLOV2Head
+from .ssd_head import SSDHead
+from .tood_head import TOODHead
+from .vfnet_head import VFNetHead
+from .yolact_head import YOLACTHead, YOLACTProtonet
+from .yolo_head import YOLOV3Head
+from .yolof_head import YOLOFHead
+from .yolox_head import YOLOXHead
+
+__all__ = [
+ 'AnchorFreeHead', 'AnchorHead', 'GuidedAnchorHead', 'FeatureAdaption',
+ 'RPNHead', 'GARPNHead', 'RetinaHead', 'RetinaSepBNHead', 'GARetinaHead',
+ 'SSDHead', 'FCOSHead', 'RepPointsHead', 'FoveaHead',
+ 'FreeAnchorRetinaHead', 'ATSSHead', 'FSAFHead', 'NASFCOSHead',
+ 'PISARetinaHead', 'PISASSDHead', 'GFLHead', 'CornerHead', 'YOLACTHead',
+ 'YOLACTProtonet', 'YOLOV3Head', 'PAAHead', 'SABLRetinaHead',
+ 'CentripetalHead', 'VFNetHead', 'StageCascadeRPNHead', 'CascadeRPNHead',
+ 'EmbeddingRPNHead', 'LDHead', 'AutoAssignHead', 'DETRHead', 'YOLOFHead',
+ 'DeformableDETRHead', 'CenterNetHead', 'YOLOXHead', 'SOLOHead',
+ 'DecoupledSOLOHead', 'DecoupledSOLOLightHead', 'SOLOV2Head', 'LADHead',
+ 'TOODHead', 'MaskFormerHead', 'Mask2FormerHead', 'DDODHead',
+ 'CenterNetUpdateHead', 'RTMDetHead', 'RTMDetSepBNHead', 'CondInstBboxHead',
+ 'CondInstMaskHead', 'RTMDetInsHead', 'RTMDetInsSepBNHead',
+ 'BoxInstBboxHead', 'BoxInstMaskHead', 'ConditionalDETRHead', 'DINOHead',
+ 'ATSSVLFusionHead', 'DABDETRHead', 'DDQDETRHead', 'GroundingDINOHead'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/anchor_free_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/anchor_free_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..90a9b3625b8fef12a2ee3a964c89597b597cb2ec
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/anchor_free_head.py
@@ -0,0 +1,317 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from abc import abstractmethod
+from typing import Any, List, Sequence, Tuple, Union
+
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from numpy import ndarray
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.utils import (ConfigType, InstanceList, MultiConfig, OptConfigType,
+ OptInstanceList)
+from ..task_modules.prior_generators import MlvlPointGenerator
+from ..utils import multi_apply
+from .base_dense_head import BaseDenseHead
+
+StrideType = Union[Sequence[int], Sequence[Tuple[int, int]]]
+
+
+@MODELS.register_module()
+class AnchorFreeHead(BaseDenseHead):
+ """Anchor-free head (FCOS, Fovea, RepPoints, etc.).
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ feat_channels (int): Number of hidden channels. Used in child classes.
+ stacked_convs (int): Number of stacking convs of the head.
+ strides (Sequence[int] or Sequence[Tuple[int, int]]): Downsample
+ factor of each feature map.
+ dcn_on_last_conv (bool): If true, use dcn in the last layer of
+ towers. Defaults to False.
+ conv_bias (bool or str): If specified as `auto`, it will be decided by
+ the norm_cfg. Bias of conv will be set as True if `norm_cfg` is
+ None, otherwise False. Default: "auto".
+ loss_cls (:obj:`ConfigDict` or dict): Config of classification loss.
+ loss_bbox (:obj:`ConfigDict` or dict): Config of localization loss.
+ bbox_coder (:obj:`ConfigDict` or dict): Config of bbox coder. Defaults
+ 'DistancePointBBoxCoder'.
+ conv_cfg (:obj:`ConfigDict` or dict, Optional): Config dict for
+ convolution layer. Defaults to None.
+ norm_cfg (:obj:`ConfigDict` or dict, Optional): Config dict for
+ normalization layer. Defaults to None.
+ train_cfg (:obj:`ConfigDict` or dict, Optional): Training config of
+ anchor-free head.
+ test_cfg (:obj:`ConfigDict` or dict, Optional): Testing config of
+ anchor-free head.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict]): Initialization config dict.
+ """ # noqa: W605
+
+ _version = 1
+
+ def __init__(
+ self,
+ num_classes: int,
+ in_channels: int,
+ feat_channels: int = 256,
+ stacked_convs: int = 4,
+ strides: StrideType = (4, 8, 16, 32, 64),
+ dcn_on_last_conv: bool = False,
+ conv_bias: Union[bool, str] = 'auto',
+ loss_cls: ConfigType = dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox: ConfigType = dict(type='IoULoss', loss_weight=1.0),
+ bbox_coder: ConfigType = dict(type='DistancePointBBoxCoder'),
+ conv_cfg: OptConfigType = None,
+ norm_cfg: OptConfigType = None,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ init_cfg: MultiConfig = dict(
+ type='Normal',
+ layer='Conv2d',
+ std=0.01,
+ override=dict(
+ type='Normal', name='conv_cls', std=0.01, bias_prob=0.01))
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.num_classes = num_classes
+ self.use_sigmoid_cls = loss_cls.get('use_sigmoid', False)
+ if self.use_sigmoid_cls:
+ self.cls_out_channels = num_classes
+ else:
+ self.cls_out_channels = num_classes + 1
+ self.in_channels = in_channels
+ self.feat_channels = feat_channels
+ self.stacked_convs = stacked_convs
+ self.strides = strides
+ self.dcn_on_last_conv = dcn_on_last_conv
+ assert conv_bias == 'auto' or isinstance(conv_bias, bool)
+ self.conv_bias = conv_bias
+ self.loss_cls = MODELS.build(loss_cls)
+ self.loss_bbox = MODELS.build(loss_bbox)
+ self.bbox_coder = TASK_UTILS.build(bbox_coder)
+
+ self.prior_generator = MlvlPointGenerator(strides)
+
+ # In order to keep a more general interface and be consistent with
+ # anchor_head. We can think of point like one anchor
+ self.num_base_priors = self.prior_generator.num_base_priors[0]
+
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self.fp16_enabled = False
+
+ self._init_layers()
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self._init_cls_convs()
+ self._init_reg_convs()
+ self._init_predictor()
+
+ def _init_cls_convs(self) -> None:
+ """Initialize classification conv layers of the head."""
+ self.cls_convs = nn.ModuleList()
+ for i in range(self.stacked_convs):
+ chn = self.in_channels if i == 0 else self.feat_channels
+ if self.dcn_on_last_conv and i == self.stacked_convs - 1:
+ conv_cfg = dict(type='DCNv2')
+ else:
+ conv_cfg = self.conv_cfg
+ self.cls_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=self.norm_cfg,
+ bias=self.conv_bias))
+
+ def _init_reg_convs(self) -> None:
+ """Initialize bbox regression conv layers of the head."""
+ self.reg_convs = nn.ModuleList()
+ for i in range(self.stacked_convs):
+ chn = self.in_channels if i == 0 else self.feat_channels
+ if self.dcn_on_last_conv and i == self.stacked_convs - 1:
+ conv_cfg = dict(type='DCNv2')
+ else:
+ conv_cfg = self.conv_cfg
+ self.reg_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=self.norm_cfg,
+ bias=self.conv_bias))
+
+ def _init_predictor(self) -> None:
+ """Initialize predictor layers of the head."""
+ self.conv_cls = nn.Conv2d(
+ self.feat_channels, self.cls_out_channels, 3, padding=1)
+ self.conv_reg = nn.Conv2d(self.feat_channels, 4, 3, padding=1)
+
+ def _load_from_state_dict(self, state_dict: dict, prefix: str,
+ local_metadata: dict, strict: bool,
+ missing_keys: Union[List[str], str],
+ unexpected_keys: Union[List[str], str],
+ error_msgs: Union[List[str], str]) -> None:
+ """Hack some keys of the model state dict so that can load checkpoints
+ of previous version."""
+ version = local_metadata.get('version', None)
+ if version is None:
+ # the key is different in early versions
+ # for example, 'fcos_cls' become 'conv_cls' now
+ bbox_head_keys = [
+ k for k in state_dict.keys() if k.startswith(prefix)
+ ]
+ ori_predictor_keys = []
+ new_predictor_keys = []
+ # e.g. 'fcos_cls' or 'fcos_reg'
+ for key in bbox_head_keys:
+ ori_predictor_keys.append(key)
+ key = key.split('.')
+ if len(key) < 2:
+ conv_name = None
+ elif key[1].endswith('cls'):
+ conv_name = 'conv_cls'
+ elif key[1].endswith('reg'):
+ conv_name = 'conv_reg'
+ elif key[1].endswith('centerness'):
+ conv_name = 'conv_centerness'
+ else:
+ conv_name = None
+ if conv_name is not None:
+ key[1] = conv_name
+ new_predictor_keys.append('.'.join(key))
+ else:
+ ori_predictor_keys.pop(-1)
+ for i in range(len(new_predictor_keys)):
+ state_dict[new_predictor_keys[i]] = state_dict.pop(
+ ori_predictor_keys[i])
+ super()._load_from_state_dict(state_dict, prefix, local_metadata,
+ strict, missing_keys, unexpected_keys,
+ error_msgs)
+
+ def forward(self, x: Tuple[Tensor]) -> Tuple[List[Tensor], List[Tensor]]:
+ """Forward features from the upstream network.
+
+ Args:
+ feats (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: Usually contain classification scores and bbox predictions.
+
+ - cls_scores (list[Tensor]): Box scores for each scale level, \
+ each is a 4D-tensor, the channel number is \
+ num_points * num_classes.
+ - bbox_preds (list[Tensor]): Box energies / deltas for each scale \
+ level, each is a 4D-tensor, the channel number is num_points * 4.
+ """
+ return multi_apply(self.forward_single, x)[:2]
+
+ def forward_single(self, x: Tensor) -> Tuple[Tensor, ...]:
+ """Forward features of a single scale level.
+
+ Args:
+ x (Tensor): FPN feature maps of the specified stride.
+
+ Returns:
+ tuple: Scores for each class, bbox predictions, features
+ after classification and regression conv layers, some
+ models needs these features like FCOS.
+ """
+ cls_feat = x
+ reg_feat = x
+
+ for cls_layer in self.cls_convs:
+ cls_feat = cls_layer(cls_feat)
+ cls_score = self.conv_cls(cls_feat)
+
+ for reg_layer in self.reg_convs:
+ reg_feat = reg_layer(reg_feat)
+ bbox_pred = self.conv_reg(reg_feat)
+ return cls_score, bbox_pred, cls_feat, reg_feat
+
+ @abstractmethod
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level,
+ each is a 4D-tensor, the channel number is
+ num_points * num_classes.
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level, each is a 4D-tensor, the channel number is
+ num_points * 4.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ """
+
+ raise NotImplementedError
+
+ @abstractmethod
+ def get_targets(self, points: List[Tensor],
+ batch_gt_instances: InstanceList) -> Any:
+ """Compute regression, classification and centerness targets for points
+ in multiple images.
+
+ Args:
+ points (list[Tensor]): Points of each fpn level, each has shape
+ (num_points, 2).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ """
+ raise NotImplementedError
+
+ # TODO refactor aug_test
+ def aug_test(self,
+ aug_batch_feats: List[Tensor],
+ aug_batch_img_metas: List[List[Tensor]],
+ rescale: bool = False) -> List[ndarray]:
+ """Test function with test time augmentation.
+
+ Args:
+ aug_batch_feats (list[Tensor]): the outer list indicates test-time
+ augmentations and inner Tensor should have a shape NxCxHxW,
+ which contains features for all images in the batch.
+ aug_batch_img_metas (list[list[dict]]): the outer list indicates
+ test-time augs (multiscale, flip, etc.) and the inner list
+ indicates images in a batch. each dict has image information.
+ rescale (bool, optional): Whether to rescale the results.
+ Defaults to False.
+
+ Returns:
+ list[ndarray]: bbox results of each class
+ """
+ return self.aug_test_bboxes(
+ aug_batch_feats, aug_batch_img_metas, rescale=rescale)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/anchor_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/anchor_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..4578caca818550397875a0df34c128f461e6ec75
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/anchor_head.py
@@ -0,0 +1,530 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+from typing import List, Optional, Tuple, Union
+
+import torch
+import torch.nn as nn
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures.bbox import BaseBoxes, cat_boxes, get_box_tensor
+from mmdet.utils import (ConfigType, InstanceList, OptConfigType,
+ OptInstanceList, OptMultiConfig)
+from ..task_modules.prior_generators import (AnchorGenerator,
+ anchor_inside_flags)
+from ..task_modules.samplers import PseudoSampler
+from ..utils import images_to_levels, multi_apply, unmap
+from .base_dense_head import BaseDenseHead
+
+
+@MODELS.register_module()
+class AnchorHead(BaseDenseHead):
+ """Anchor-based head (RPN, RetinaNet, SSD, etc.).
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ feat_channels (int): Number of hidden channels. Used in child classes.
+ anchor_generator (dict): Config dict for anchor generator
+ bbox_coder (dict): Config of bounding box coder.
+ reg_decoded_bbox (bool): If true, the regression loss would be
+ applied directly on decoded bounding boxes, converting both
+ the predicted boxes and regression targets to absolute
+ coordinates format. Default False. It should be `True` when
+ using `IoULoss`, `GIoULoss`, or `DIoULoss` in the bbox head.
+ loss_cls (dict): Config of classification loss.
+ loss_bbox (dict): Config of localization loss.
+ train_cfg (dict): Training config of anchor head.
+ test_cfg (dict): Testing config of anchor head.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ """ # noqa: W605
+
+ def __init__(
+ self,
+ num_classes: int,
+ in_channels: int,
+ feat_channels: int = 256,
+ anchor_generator: ConfigType = dict(
+ type='AnchorGenerator',
+ scales=[8, 16, 32],
+ ratios=[0.5, 1.0, 2.0],
+ strides=[4, 8, 16, 32, 64]),
+ bbox_coder: ConfigType = dict(
+ type='DeltaXYWHBBoxCoder',
+ clip_border=True,
+ target_means=(.0, .0, .0, .0),
+ target_stds=(1.0, 1.0, 1.0, 1.0)),
+ reg_decoded_bbox: bool = False,
+ loss_cls: ConfigType = dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox: ConfigType = dict(
+ type='SmoothL1Loss', beta=1.0 / 9.0, loss_weight=1.0),
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ init_cfg: OptMultiConfig = dict(
+ type='Normal', layer='Conv2d', std=0.01)
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.in_channels = in_channels
+ self.num_classes = num_classes
+ self.feat_channels = feat_channels
+ self.use_sigmoid_cls = loss_cls.get('use_sigmoid', False)
+ if self.use_sigmoid_cls:
+ self.cls_out_channels = num_classes
+ else:
+ self.cls_out_channels = num_classes + 1
+
+ if self.cls_out_channels <= 0:
+ raise ValueError(f'num_classes={num_classes} is too small')
+ self.reg_decoded_bbox = reg_decoded_bbox
+
+ self.bbox_coder = TASK_UTILS.build(bbox_coder)
+ self.loss_cls = MODELS.build(loss_cls)
+ self.loss_bbox = MODELS.build(loss_bbox)
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+ if self.train_cfg:
+ self.assigner = TASK_UTILS.build(self.train_cfg['assigner'])
+ if train_cfg.get('sampler', None) is not None:
+ self.sampler = TASK_UTILS.build(
+ self.train_cfg['sampler'], default_args=dict(context=self))
+ else:
+ self.sampler = PseudoSampler(context=self)
+
+ self.fp16_enabled = False
+
+ self.prior_generator = TASK_UTILS.build(anchor_generator)
+
+ # Usually the numbers of anchors for each level are the same
+ # except SSD detectors. So it is an int in the most dense
+ # heads but a list of int in SSDHead
+ self.num_base_priors = self.prior_generator.num_base_priors[0]
+ self._init_layers()
+
+ @property
+ def num_anchors(self) -> int:
+ warnings.warn('DeprecationWarning: `num_anchors` is deprecated, '
+ 'for consistency or also use '
+ '`num_base_priors` instead')
+ return self.prior_generator.num_base_priors[0]
+
+ @property
+ def anchor_generator(self) -> AnchorGenerator:
+ warnings.warn('DeprecationWarning: anchor_generator is deprecated, '
+ 'please use "prior_generator" instead')
+ return self.prior_generator
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self.conv_cls = nn.Conv2d(self.in_channels,
+ self.num_base_priors * self.cls_out_channels,
+ 1)
+ reg_dim = self.bbox_coder.encode_size
+ self.conv_reg = nn.Conv2d(self.in_channels,
+ self.num_base_priors * reg_dim, 1)
+
+ def forward_single(self, x: Tensor) -> Tuple[Tensor, Tensor]:
+ """Forward feature of a single scale level.
+
+ Args:
+ x (Tensor): Features of a single scale level.
+
+ Returns:
+ tuple:
+ cls_score (Tensor): Cls scores for a single scale level \
+ the channels number is num_base_priors * num_classes.
+ bbox_pred (Tensor): Box energies / deltas for a single scale \
+ level, the channels number is num_base_priors * 4.
+ """
+ cls_score = self.conv_cls(x)
+ bbox_pred = self.conv_reg(x)
+ return cls_score, bbox_pred
+
+ def forward(self, x: Tuple[Tensor]) -> Tuple[List[Tensor]]:
+ """Forward features from the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: A tuple of classification scores and bbox prediction.
+
+ - cls_scores (list[Tensor]): Classification scores for all \
+ scale levels, each is a 4D-tensor, the channels number \
+ is num_base_priors * num_classes.
+ - bbox_preds (list[Tensor]): Box energies / deltas for all \
+ scale levels, each is a 4D-tensor, the channels number \
+ is num_base_priors * 4.
+ """
+ return multi_apply(self.forward_single, x)
+
+ def get_anchors(self,
+ featmap_sizes: List[tuple],
+ batch_img_metas: List[dict],
+ device: Union[torch.device, str] = 'cuda') \
+ -> Tuple[List[List[Tensor]], List[List[Tensor]]]:
+ """Get anchors according to feature map sizes.
+
+ Args:
+ featmap_sizes (list[tuple]): Multi-level feature map sizes.
+ batch_img_metas (list[dict]): Image meta info.
+ device (torch.device | str): Device for returned tensors.
+ Defaults to cuda.
+
+ Returns:
+ tuple:
+
+ - anchor_list (list[list[Tensor]]): Anchors of each image.
+ - valid_flag_list (list[list[Tensor]]): Valid flags of each
+ image.
+ """
+ num_imgs = len(batch_img_metas)
+
+ # since feature map sizes of all images are the same, we only compute
+ # anchors for one time
+ multi_level_anchors = self.prior_generator.grid_priors(
+ featmap_sizes, device=device)
+ anchor_list = [multi_level_anchors for _ in range(num_imgs)]
+
+ # for each image, we compute valid flags of multi level anchors
+ valid_flag_list = []
+ for img_id, img_meta in enumerate(batch_img_metas):
+ multi_level_flags = self.prior_generator.valid_flags(
+ featmap_sizes, img_meta['pad_shape'], device)
+ valid_flag_list.append(multi_level_flags)
+
+ return anchor_list, valid_flag_list
+
+ def _get_targets_single(self,
+ flat_anchors: Union[Tensor, BaseBoxes],
+ valid_flags: Tensor,
+ gt_instances: InstanceData,
+ img_meta: dict,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ unmap_outputs: bool = True) -> tuple:
+ """Compute regression and classification targets for anchors in a
+ single image.
+
+ Args:
+ flat_anchors (Tensor or :obj:`BaseBoxes`): Multi-level anchors
+ of the image, which are concatenated into a single tensor
+ or box type of shape (num_anchors, 4)
+ valid_flags (Tensor): Multi level valid flags of the image,
+ which are concatenated into a single tensor of
+ shape (num_anchors, ).
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes`` and ``labels``
+ attributes.
+ img_meta (dict): Meta information for current image.
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ unmap_outputs (bool): Whether to map outputs back to the original
+ set of anchors. Defaults to True.
+
+ Returns:
+ tuple:
+
+ - labels (Tensor): Labels of each level.
+ - label_weights (Tensor): Label weights of each level.
+ - bbox_targets (Tensor): BBox targets of each level.
+ - bbox_weights (Tensor): BBox weights of each level.
+ - pos_inds (Tensor): positive samples indexes.
+ - neg_inds (Tensor): negative samples indexes.
+ - sampling_result (:obj:`SamplingResult`): Sampling results.
+ """
+ inside_flags = anchor_inside_flags(flat_anchors, valid_flags,
+ img_meta['img_shape'][:2],
+ self.train_cfg['allowed_border'])
+ if not inside_flags.any():
+ raise ValueError(
+ 'There is no valid anchor inside the image boundary. Please '
+ 'check the image size and anchor sizes, or set '
+ '``allowed_border`` to -1 to skip the condition.')
+ # assign gt and sample anchors
+ anchors = flat_anchors[inside_flags]
+
+ pred_instances = InstanceData(priors=anchors)
+ assign_result = self.assigner.assign(pred_instances, gt_instances,
+ gt_instances_ignore)
+ # No sampling is required except for RPN and
+ # Guided Anchoring algorithms
+ sampling_result = self.sampler.sample(assign_result, pred_instances,
+ gt_instances)
+
+ num_valid_anchors = anchors.shape[0]
+ target_dim = gt_instances.bboxes.size(-1) if self.reg_decoded_bbox \
+ else self.bbox_coder.encode_size
+ bbox_targets = anchors.new_zeros(num_valid_anchors, target_dim)
+ bbox_weights = anchors.new_zeros(num_valid_anchors, target_dim)
+
+ # TODO: Considering saving memory, is it necessary to be long?
+ labels = anchors.new_full((num_valid_anchors, ),
+ self.num_classes,
+ dtype=torch.long)
+ label_weights = anchors.new_zeros(num_valid_anchors, dtype=torch.float)
+
+ pos_inds = sampling_result.pos_inds
+ neg_inds = sampling_result.neg_inds
+ # `bbox_coder.encode` accepts tensor or box type inputs and generates
+ # tensor targets. If regressing decoded boxes, the code will convert
+ # box type `pos_bbox_targets` to tensor.
+ if len(pos_inds) > 0:
+ if not self.reg_decoded_bbox:
+ pos_bbox_targets = self.bbox_coder.encode(
+ sampling_result.pos_priors, sampling_result.pos_gt_bboxes)
+ else:
+ pos_bbox_targets = sampling_result.pos_gt_bboxes
+ pos_bbox_targets = get_box_tensor(pos_bbox_targets)
+ bbox_targets[pos_inds, :] = pos_bbox_targets
+ bbox_weights[pos_inds, :] = 1.0
+
+ labels[pos_inds] = sampling_result.pos_gt_labels
+ if self.train_cfg['pos_weight'] <= 0:
+ label_weights[pos_inds] = 1.0
+ else:
+ label_weights[pos_inds] = self.train_cfg['pos_weight']
+ if len(neg_inds) > 0:
+ label_weights[neg_inds] = 1.0
+
+ # map up to original set of anchors
+ if unmap_outputs:
+ num_total_anchors = flat_anchors.size(0)
+ labels = unmap(
+ labels, num_total_anchors, inside_flags,
+ fill=self.num_classes) # fill bg label
+ label_weights = unmap(label_weights, num_total_anchors,
+ inside_flags)
+ bbox_targets = unmap(bbox_targets, num_total_anchors, inside_flags)
+ bbox_weights = unmap(bbox_weights, num_total_anchors, inside_flags)
+
+ return (labels, label_weights, bbox_targets, bbox_weights, pos_inds,
+ neg_inds, sampling_result)
+
+ def get_targets(self,
+ anchor_list: List[List[Tensor]],
+ valid_flag_list: List[List[Tensor]],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None,
+ unmap_outputs: bool = True,
+ return_sampling_results: bool = False) -> tuple:
+ """Compute regression and classification targets for anchors in
+ multiple images.
+
+ Args:
+ anchor_list (list[list[Tensor]]): Multi level anchors of each
+ image. The outer list indicates images, and the inner list
+ corresponds to feature levels of the image. Each element of
+ the inner list is a tensor of shape (num_anchors, 4).
+ valid_flag_list (list[list[Tensor]]): Multi level valid flags of
+ each image. The outer list indicates images, and the inner list
+ corresponds to feature levels of the image. Each element of
+ the inner list is a tensor of shape (num_anchors, )
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ unmap_outputs (bool): Whether to map outputs back to the original
+ set of anchors. Defaults to True.
+ return_sampling_results (bool): Whether to return the sampling
+ results. Defaults to False.
+
+ Returns:
+ tuple: Usually returns a tuple containing learning targets.
+
+ - labels_list (list[Tensor]): Labels of each level.
+ - label_weights_list (list[Tensor]): Label weights of each
+ level.
+ - bbox_targets_list (list[Tensor]): BBox targets of each level.
+ - bbox_weights_list (list[Tensor]): BBox weights of each level.
+ - avg_factor (int): Average factor that is used to average
+ the loss. When using sampling method, avg_factor is usually
+ the sum of positive and negative priors. When using
+ `PseudoSampler`, `avg_factor` is usually equal to the number
+ of positive priors.
+
+ additional_returns: This function enables user-defined returns from
+ `self._get_targets_single`. These returns are currently refined
+ to properties at each feature map (i.e. having HxW dimension).
+ The results will be concatenated after the end
+ """
+ num_imgs = len(batch_img_metas)
+ assert len(anchor_list) == len(valid_flag_list) == num_imgs
+
+ if batch_gt_instances_ignore is None:
+ batch_gt_instances_ignore = [None] * num_imgs
+
+ # anchor number of multi levels
+ num_level_anchors = [anchors.size(0) for anchors in anchor_list[0]]
+ # concat all level anchors to a single tensor
+ concat_anchor_list = []
+ concat_valid_flag_list = []
+ for i in range(num_imgs):
+ assert len(anchor_list[i]) == len(valid_flag_list[i])
+ concat_anchor_list.append(cat_boxes(anchor_list[i]))
+ concat_valid_flag_list.append(torch.cat(valid_flag_list[i]))
+
+ # compute targets for each image
+ results = multi_apply(
+ self._get_targets_single,
+ concat_anchor_list,
+ concat_valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore,
+ unmap_outputs=unmap_outputs)
+ (all_labels, all_label_weights, all_bbox_targets, all_bbox_weights,
+ pos_inds_list, neg_inds_list, sampling_results_list) = results[:7]
+ rest_results = list(results[7:]) # user-added return values
+ # Get `avg_factor` of all images, which calculate in `SamplingResult`.
+ # When using sampling method, avg_factor is usually the sum of
+ # positive and negative priors. When using `PseudoSampler`,
+ # `avg_factor` is usually equal to the number of positive priors.
+ avg_factor = sum(
+ [results.avg_factor for results in sampling_results_list])
+ # update `_raw_positive_infos`, which will be used when calling
+ # `get_positive_infos`.
+ self._raw_positive_infos.update(sampling_results=sampling_results_list)
+ # split targets to a list w.r.t. multiple levels
+ labels_list = images_to_levels(all_labels, num_level_anchors)
+ label_weights_list = images_to_levels(all_label_weights,
+ num_level_anchors)
+ bbox_targets_list = images_to_levels(all_bbox_targets,
+ num_level_anchors)
+ bbox_weights_list = images_to_levels(all_bbox_weights,
+ num_level_anchors)
+ res = (labels_list, label_weights_list, bbox_targets_list,
+ bbox_weights_list, avg_factor)
+ if return_sampling_results:
+ res = res + (sampling_results_list, )
+ for i, r in enumerate(rest_results): # user-added return values
+ rest_results[i] = images_to_levels(r, num_level_anchors)
+
+ return res + tuple(rest_results)
+
+ def loss_by_feat_single(self, cls_score: Tensor, bbox_pred: Tensor,
+ anchors: Tensor, labels: Tensor,
+ label_weights: Tensor, bbox_targets: Tensor,
+ bbox_weights: Tensor, avg_factor: int) -> tuple:
+ """Calculate the loss of a single scale level based on the features
+ extracted by the detection head.
+
+ Args:
+ cls_score (Tensor): Box scores for each scale level
+ Has shape (N, num_anchors * num_classes, H, W).
+ bbox_pred (Tensor): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W).
+ anchors (Tensor): Box reference for each scale level with shape
+ (N, num_total_anchors, 4).
+ labels (Tensor): Labels of each anchors with shape
+ (N, num_total_anchors).
+ label_weights (Tensor): Label weights of each anchor with shape
+ (N, num_total_anchors)
+ bbox_targets (Tensor): BBox regression targets of each anchor
+ weight shape (N, num_total_anchors, 4).
+ bbox_weights (Tensor): BBox regression loss weights of each anchor
+ with shape (N, num_total_anchors, 4).
+ avg_factor (int): Average factor that is used to average the loss.
+
+ Returns:
+ tuple: loss components.
+ """
+ # classification loss
+ labels = labels.reshape(-1)
+ label_weights = label_weights.reshape(-1)
+ cls_score = cls_score.permute(0, 2, 3,
+ 1).reshape(-1, self.cls_out_channels)
+ loss_cls = self.loss_cls(
+ cls_score, labels, label_weights, avg_factor=avg_factor)
+ # regression loss
+ target_dim = bbox_targets.size(-1)
+ bbox_targets = bbox_targets.reshape(-1, target_dim)
+ bbox_weights = bbox_weights.reshape(-1, target_dim)
+ bbox_pred = bbox_pred.permute(0, 2, 3,
+ 1).reshape(-1,
+ self.bbox_coder.encode_size)
+ if self.reg_decoded_bbox:
+ # When the regression loss (e.g. `IouLoss`, `GIouLoss`)
+ # is applied directly on the decoded bounding boxes, it
+ # decodes the already encoded coordinates to absolute format.
+ anchors = anchors.reshape(-1, anchors.size(-1))
+ bbox_pred = self.bbox_coder.decode(anchors, bbox_pred)
+ bbox_pred = get_box_tensor(bbox_pred)
+ loss_bbox = self.loss_bbox(
+ bbox_pred, bbox_targets, bbox_weights, avg_factor=avg_factor)
+ return loss_cls, loss_bbox
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ has shape (N, num_anchors * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ assert len(featmap_sizes) == self.prior_generator.num_levels
+
+ device = cls_scores[0].device
+
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+ cls_reg_targets = self.get_targets(
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore)
+ (labels_list, label_weights_list, bbox_targets_list, bbox_weights_list,
+ avg_factor) = cls_reg_targets
+
+ # anchor number of multi levels
+ num_level_anchors = [anchors.size(0) for anchors in anchor_list[0]]
+ # concat all level anchors and flags to a single tensor
+ concat_anchor_list = []
+ for i in range(len(anchor_list)):
+ concat_anchor_list.append(cat_boxes(anchor_list[i]))
+ all_anchor_list = images_to_levels(concat_anchor_list,
+ num_level_anchors)
+
+ losses_cls, losses_bbox = multi_apply(
+ self.loss_by_feat_single,
+ cls_scores,
+ bbox_preds,
+ all_anchor_list,
+ labels_list,
+ label_weights_list,
+ bbox_targets_list,
+ bbox_weights_list,
+ avg_factor=avg_factor)
+ return dict(loss_cls=losses_cls, loss_bbox=losses_bbox)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/atss_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/atss_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..2ce71b3eff5e0ed624ec7ae16e8db80c90e8ffa1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/atss_head.py
@@ -0,0 +1,524 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Sequence, Tuple
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule, Scale
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import (ConfigType, InstanceList, MultiConfig, OptConfigType,
+ OptInstanceList, reduce_mean)
+from ..task_modules.prior_generators import anchor_inside_flags
+from ..utils import images_to_levels, multi_apply, unmap
+from .anchor_head import AnchorHead
+
+
+@MODELS.register_module()
+class ATSSHead(AnchorHead):
+ """Detection Head of `ATSS `_.
+
+ ATSS head structure is similar with FCOS, however ATSS use anchor boxes
+ and assign label by Adaptive Training Sample Selection instead max-iou.
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ pred_kernel_size (int): Kernel size of ``nn.Conv2d``
+ stacked_convs (int): Number of stacking convs of the head.
+ conv_cfg (:obj:`ConfigDict` or dict, optional): Config dict for
+ convolution layer. Defaults to None.
+ norm_cfg (:obj:`ConfigDict` or dict): Config dict for normalization
+ layer. Defaults to ``dict(type='GN', num_groups=32,
+ requires_grad=True)``.
+ reg_decoded_bbox (bool): If true, the regression loss would be
+ applied directly on decoded bounding boxes, converting both
+ the predicted boxes and regression targets to absolute
+ coordinates format. Defaults to False. It should be `True` when
+ using `IoULoss`, `GIoULoss`, or `DIoULoss` in the bbox head.
+ loss_centerness (:obj:`ConfigDict` or dict): Config of centerness loss.
+ Defaults to ``dict(type='CrossEntropyLoss', use_sigmoid=True,
+ loss_weight=1.0)``.
+ init_cfg (:obj:`ConfigDict` or dict or list[dict] or
+ list[:obj:`ConfigDict`]): Initialization config dict.
+ """
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: int,
+ pred_kernel_size: int = 3,
+ stacked_convs: int = 4,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: ConfigType = dict(
+ type='GN', num_groups=32, requires_grad=True),
+ reg_decoded_bbox: bool = True,
+ loss_centerness: ConfigType = dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ loss_weight=1.0),
+ init_cfg: MultiConfig = dict(
+ type='Normal',
+ layer='Conv2d',
+ std=0.01,
+ override=dict(
+ type='Normal',
+ name='atss_cls',
+ std=0.01,
+ bias_prob=0.01)),
+ **kwargs) -> None:
+ self.pred_kernel_size = pred_kernel_size
+ self.stacked_convs = stacked_convs
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ super().__init__(
+ num_classes=num_classes,
+ in_channels=in_channels,
+ reg_decoded_bbox=reg_decoded_bbox,
+ init_cfg=init_cfg,
+ **kwargs)
+
+ self.sampling = False
+ self.loss_centerness = MODELS.build(loss_centerness)
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self.relu = nn.ReLU(inplace=True)
+ self.cls_convs = nn.ModuleList()
+ self.reg_convs = nn.ModuleList()
+ for i in range(self.stacked_convs):
+ chn = self.in_channels if i == 0 else self.feat_channels
+ self.cls_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ self.reg_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ pred_pad_size = self.pred_kernel_size // 2
+ self.atss_cls = nn.Conv2d(
+ self.feat_channels,
+ self.num_anchors * self.cls_out_channels,
+ self.pred_kernel_size,
+ padding=pred_pad_size)
+ self.atss_reg = nn.Conv2d(
+ self.feat_channels,
+ self.num_base_priors * 4,
+ self.pred_kernel_size,
+ padding=pred_pad_size)
+ self.atss_centerness = nn.Conv2d(
+ self.feat_channels,
+ self.num_base_priors * 1,
+ self.pred_kernel_size,
+ padding=pred_pad_size)
+ self.scales = nn.ModuleList(
+ [Scale(1.0) for _ in self.prior_generator.strides])
+
+ def forward(self, x: Tuple[Tensor]) -> Tuple[List[Tensor]]:
+ """Forward features from the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: Usually a tuple of classification scores and bbox prediction
+ cls_scores (list[Tensor]): Classification scores for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_anchors * num_classes.
+ bbox_preds (list[Tensor]): Box energies / deltas for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_anchors * 4.
+ """
+ return multi_apply(self.forward_single, x, self.scales)
+
+ def forward_single(self, x: Tensor, scale: Scale) -> Sequence[Tensor]:
+ """Forward feature of a single scale level.
+
+ Args:
+ x (Tensor): Features of a single scale level.
+ scale (:obj: `mmcv.cnn.Scale`): Learnable scale module to resize
+ the bbox prediction.
+
+ Returns:
+ tuple:
+ cls_score (Tensor): Cls scores for a single scale level
+ the channels number is num_anchors * num_classes.
+ bbox_pred (Tensor): Box energies / deltas for a single scale
+ level, the channels number is num_anchors * 4.
+ centerness (Tensor): Centerness for a single scale level, the
+ channel number is (N, num_anchors * 1, H, W).
+ """
+ cls_feat = x
+ reg_feat = x
+ for cls_conv in self.cls_convs:
+ cls_feat = cls_conv(cls_feat)
+ for reg_conv in self.reg_convs:
+ reg_feat = reg_conv(reg_feat)
+ cls_score = self.atss_cls(cls_feat)
+ # we just follow atss, not apply exp in bbox_pred
+ bbox_pred = scale(self.atss_reg(reg_feat)).float()
+ centerness = self.atss_centerness(reg_feat)
+ return cls_score, bbox_pred, centerness
+
+ def loss_by_feat_single(self, anchors: Tensor, cls_score: Tensor,
+ bbox_pred: Tensor, centerness: Tensor,
+ labels: Tensor, label_weights: Tensor,
+ bbox_targets: Tensor, avg_factor: float) -> dict:
+ """Calculate the loss of a single scale level based on the features
+ extracted by the detection head.
+
+ Args:
+ cls_score (Tensor): Box scores for each scale level
+ Has shape (N, num_anchors * num_classes, H, W).
+ bbox_pred (Tensor): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W).
+ anchors (Tensor): Box reference for each scale level with shape
+ (N, num_total_anchors, 4).
+ labels (Tensor): Labels of each anchors with shape
+ (N, num_total_anchors).
+ label_weights (Tensor): Label weights of each anchor with shape
+ (N, num_total_anchors)
+ bbox_targets (Tensor): BBox regression targets of each anchor with
+ shape (N, num_total_anchors, 4).
+ avg_factor (float): Average factor that is used to average
+ the loss. When using sampling method, avg_factor is usually
+ the sum of positive and negative priors. When using
+ `PseudoSampler`, `avg_factor` is usually equal to the number
+ of positive priors.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+
+ anchors = anchors.reshape(-1, 4)
+ cls_score = cls_score.permute(0, 2, 3, 1).reshape(
+ -1, self.cls_out_channels).contiguous()
+ bbox_pred = bbox_pred.permute(0, 2, 3, 1).reshape(-1, 4)
+ centerness = centerness.permute(0, 2, 3, 1).reshape(-1)
+ bbox_targets = bbox_targets.reshape(-1, 4)
+ labels = labels.reshape(-1)
+ label_weights = label_weights.reshape(-1)
+
+ # classification loss
+ loss_cls = self.loss_cls(
+ cls_score, labels, label_weights, avg_factor=avg_factor)
+
+ # FG cat_id: [0, num_classes -1], BG cat_id: num_classes
+ bg_class_ind = self.num_classes
+ pos_inds = ((labels >= 0)
+ & (labels < bg_class_ind)).nonzero().squeeze(1)
+
+ if len(pos_inds) > 0:
+ pos_bbox_targets = bbox_targets[pos_inds]
+ pos_bbox_pred = bbox_pred[pos_inds]
+ pos_anchors = anchors[pos_inds]
+ pos_centerness = centerness[pos_inds]
+
+ centerness_targets = self.centerness_target(
+ pos_anchors, pos_bbox_targets)
+ pos_decode_bbox_pred = self.bbox_coder.decode(
+ pos_anchors, pos_bbox_pred)
+
+ # regression loss
+ loss_bbox = self.loss_bbox(
+ pos_decode_bbox_pred,
+ pos_bbox_targets,
+ weight=centerness_targets,
+ avg_factor=1.0)
+
+ # centerness loss
+ loss_centerness = self.loss_centerness(
+ pos_centerness, centerness_targets, avg_factor=avg_factor)
+
+ else:
+ loss_bbox = bbox_pred.sum() * 0
+ loss_centerness = centerness.sum() * 0
+ centerness_targets = bbox_targets.new_tensor(0.)
+
+ return loss_cls, loss_bbox, loss_centerness, centerness_targets.sum()
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ centernesses: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ Has shape (N, num_anchors * num_classes, H, W)
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W)
+ centernesses (list[Tensor]): Centerness for each scale
+ level with shape (N, num_anchors * 1, H, W)
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ featmap_sizes = [featmap.size()[-2:] for featmap in bbox_preds]
+ assert len(featmap_sizes) == self.prior_generator.num_levels
+
+ device = cls_scores[0].device
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+
+ cls_reg_targets = self.get_targets(
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore)
+
+ (anchor_list, labels_list, label_weights_list, bbox_targets_list,
+ bbox_weights_list, avg_factor) = cls_reg_targets
+ avg_factor = reduce_mean(
+ torch.tensor(avg_factor, dtype=torch.float, device=device)).item()
+
+ losses_cls, losses_bbox, loss_centerness, \
+ bbox_avg_factor = multi_apply(
+ self.loss_by_feat_single,
+ anchor_list,
+ cls_scores,
+ bbox_preds,
+ centernesses,
+ labels_list,
+ label_weights_list,
+ bbox_targets_list,
+ avg_factor=avg_factor)
+
+ bbox_avg_factor = sum(bbox_avg_factor)
+ bbox_avg_factor = reduce_mean(bbox_avg_factor).clamp_(min=1).item()
+ losses_bbox = list(map(lambda x: x / bbox_avg_factor, losses_bbox))
+ return dict(
+ loss_cls=losses_cls,
+ loss_bbox=losses_bbox,
+ loss_centerness=loss_centerness)
+
+ def centerness_target(self, anchors: Tensor, gts: Tensor) -> Tensor:
+ """Calculate the centerness between anchors and gts.
+
+ Only calculate pos centerness targets, otherwise there may be nan.
+
+ Args:
+ anchors (Tensor): Anchors with shape (N, 4), "xyxy" format.
+ gts (Tensor): Ground truth bboxes with shape (N, 4), "xyxy" format.
+
+ Returns:
+ Tensor: Centerness between anchors and gts.
+ """
+ anchors_cx = (anchors[:, 2] + anchors[:, 0]) / 2
+ anchors_cy = (anchors[:, 3] + anchors[:, 1]) / 2
+ l_ = anchors_cx - gts[:, 0]
+ t_ = anchors_cy - gts[:, 1]
+ r_ = gts[:, 2] - anchors_cx
+ b_ = gts[:, 3] - anchors_cy
+
+ left_right = torch.stack([l_, r_], dim=1)
+ top_bottom = torch.stack([t_, b_], dim=1)
+ centerness = torch.sqrt(
+ (left_right.min(dim=-1)[0] / left_right.max(dim=-1)[0]) *
+ (top_bottom.min(dim=-1)[0] / top_bottom.max(dim=-1)[0]))
+ assert not torch.isnan(centerness).any()
+ return centerness
+
+ def get_targets(self,
+ anchor_list: List[List[Tensor]],
+ valid_flag_list: List[List[Tensor]],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None,
+ unmap_outputs: bool = True) -> tuple:
+ """Get targets for ATSS head.
+
+ This method is almost the same as `AnchorHead.get_targets()`. Besides
+ returning the targets as the parent method does, it also returns the
+ anchors as the first element of the returned tuple.
+ """
+ num_imgs = len(batch_img_metas)
+ assert len(anchor_list) == len(valid_flag_list) == num_imgs
+
+ # anchor number of multi levels
+ num_level_anchors = [anchors.size(0) for anchors in anchor_list[0]]
+ num_level_anchors_list = [num_level_anchors] * num_imgs
+
+ # concat all level anchors and flags to a single tensor
+ for i in range(num_imgs):
+ assert len(anchor_list[i]) == len(valid_flag_list[i])
+ anchor_list[i] = torch.cat(anchor_list[i])
+ valid_flag_list[i] = torch.cat(valid_flag_list[i])
+
+ # compute targets for each image
+ if batch_gt_instances_ignore is None:
+ batch_gt_instances_ignore = [None] * num_imgs
+ (all_anchors, all_labels, all_label_weights, all_bbox_targets,
+ all_bbox_weights, pos_inds_list, neg_inds_list,
+ sampling_results_list) = multi_apply(
+ self._get_targets_single,
+ anchor_list,
+ valid_flag_list,
+ num_level_anchors_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore,
+ unmap_outputs=unmap_outputs)
+ # Get `avg_factor` of all images, which calculate in `SamplingResult`.
+ # When using sampling method, avg_factor is usually the sum of
+ # positive and negative priors. When using `PseudoSampler`,
+ # `avg_factor` is usually equal to the number of positive priors.
+ avg_factor = sum(
+ [results.avg_factor for results in sampling_results_list])
+ # split targets to a list w.r.t. multiple levels
+ anchors_list = images_to_levels(all_anchors, num_level_anchors)
+ labels_list = images_to_levels(all_labels, num_level_anchors)
+ label_weights_list = images_to_levels(all_label_weights,
+ num_level_anchors)
+ bbox_targets_list = images_to_levels(all_bbox_targets,
+ num_level_anchors)
+ bbox_weights_list = images_to_levels(all_bbox_weights,
+ num_level_anchors)
+ return (anchors_list, labels_list, label_weights_list,
+ bbox_targets_list, bbox_weights_list, avg_factor)
+
+ def _get_targets_single(self,
+ flat_anchors: Tensor,
+ valid_flags: Tensor,
+ num_level_anchors: List[int],
+ gt_instances: InstanceData,
+ img_meta: dict,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ unmap_outputs: bool = True) -> tuple:
+ """Compute regression, classification targets for anchors in a single
+ image.
+
+ Args:
+ flat_anchors (Tensor): Multi-level anchors of the image, which are
+ concatenated into a single tensor of shape (num_anchors ,4)
+ valid_flags (Tensor): Multi level valid flags of the image,
+ which are concatenated into a single tensor of
+ shape (num_anchors,).
+ num_level_anchors (List[int]): Number of anchors of each scale
+ level.
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ img_meta (dict): Meta information for current image.
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ unmap_outputs (bool): Whether to map outputs back to the original
+ set of anchors.
+
+ Returns:
+ tuple: N is the number of total anchors in the image.
+ labels (Tensor): Labels of all anchors in the image with shape
+ (N,).
+ label_weights (Tensor): Label weights of all anchor in the
+ image with shape (N,).
+ bbox_targets (Tensor): BBox targets of all anchors in the
+ image with shape (N, 4).
+ bbox_weights (Tensor): BBox weights of all anchors in the
+ image with shape (N, 4)
+ pos_inds (Tensor): Indices of positive anchor with shape
+ (num_pos,).
+ neg_inds (Tensor): Indices of negative anchor with shape
+ (num_neg,).
+ sampling_result (:obj:`SamplingResult`): Sampling results.
+ """
+ inside_flags = anchor_inside_flags(flat_anchors, valid_flags,
+ img_meta['img_shape'][:2],
+ self.train_cfg['allowed_border'])
+ if not inside_flags.any():
+ raise ValueError(
+ 'There is no valid anchor inside the image boundary. Please '
+ 'check the image size and anchor sizes, or set '
+ '``allowed_border`` to -1 to skip the condition.')
+ # assign gt and sample anchors
+ anchors = flat_anchors[inside_flags, :]
+
+ num_level_anchors_inside = self.get_num_level_anchors_inside(
+ num_level_anchors, inside_flags)
+ pred_instances = InstanceData(priors=anchors)
+ assign_result = self.assigner.assign(pred_instances,
+ num_level_anchors_inside,
+ gt_instances, gt_instances_ignore)
+
+ sampling_result = self.sampler.sample(assign_result, pred_instances,
+ gt_instances)
+
+ num_valid_anchors = anchors.shape[0]
+ bbox_targets = torch.zeros_like(anchors)
+ bbox_weights = torch.zeros_like(anchors)
+ labels = anchors.new_full((num_valid_anchors, ),
+ self.num_classes,
+ dtype=torch.long)
+ label_weights = anchors.new_zeros(num_valid_anchors, dtype=torch.float)
+
+ pos_inds = sampling_result.pos_inds
+ neg_inds = sampling_result.neg_inds
+ if len(pos_inds) > 0:
+ if self.reg_decoded_bbox:
+ pos_bbox_targets = sampling_result.pos_gt_bboxes
+ else:
+ pos_bbox_targets = self.bbox_coder.encode(
+ sampling_result.pos_priors, sampling_result.pos_gt_bboxes)
+
+ bbox_targets[pos_inds, :] = pos_bbox_targets
+ bbox_weights[pos_inds, :] = 1.0
+
+ labels[pos_inds] = sampling_result.pos_gt_labels
+ if self.train_cfg['pos_weight'] <= 0:
+ label_weights[pos_inds] = 1.0
+ else:
+ label_weights[pos_inds] = self.train_cfg['pos_weight']
+ if len(neg_inds) > 0:
+ label_weights[neg_inds] = 1.0
+
+ # map up to original set of anchors
+ if unmap_outputs:
+ num_total_anchors = flat_anchors.size(0)
+ anchors = unmap(anchors, num_total_anchors, inside_flags)
+ labels = unmap(
+ labels, num_total_anchors, inside_flags, fill=self.num_classes)
+ label_weights = unmap(label_weights, num_total_anchors,
+ inside_flags)
+ bbox_targets = unmap(bbox_targets, num_total_anchors, inside_flags)
+ bbox_weights = unmap(bbox_weights, num_total_anchors, inside_flags)
+
+ return (anchors, labels, label_weights, bbox_targets, bbox_weights,
+ pos_inds, neg_inds, sampling_result)
+
+ def get_num_level_anchors_inside(self, num_level_anchors, inside_flags):
+ """Get the number of valid anchors in every level."""
+
+ split_inside_flags = torch.split(inside_flags, num_level_anchors)
+ num_level_anchors_inside = [
+ int(flags.sum()) for flags in split_inside_flags
+ ]
+ return num_level_anchors_inside
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/atss_vlfusion_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/atss_vlfusion_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..c5cd28b4a040ba447130aed07629f6312f95dcf3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/atss_vlfusion_head.py
@@ -0,0 +1,949 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import math
+from typing import Callable, List, Optional, Sequence, Tuple, Union
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import Scale
+from mmcv.ops.modulated_deform_conv import ModulatedDeformConv2d
+from mmengine.config import ConfigDict
+from mmengine.model import BaseModel
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+try:
+ from transformers import BertConfig
+except ImportError:
+ BertConfig = None
+
+from mmdet.registry import MODELS
+from mmdet.structures.bbox import cat_boxes
+from mmdet.utils import InstanceList, OptInstanceList, reduce_mean
+from ..utils import (BertEncoderLayer, VLFuse, filter_scores_and_topk,
+ permute_and_flatten, select_single_mlvl,
+ unpack_gt_instances)
+from ..utils.vlfuse_helper import MAX_CLAMP_VALUE
+from .atss_head import ATSSHead
+
+
+def convert_grounding_to_cls_scores(logits: Tensor,
+ positive_maps: List[dict]) -> Tensor:
+ """Convert logits to class scores."""
+ assert len(positive_maps) == logits.shape[0] # batch size
+
+ scores = torch.zeros(logits.shape[0], logits.shape[1],
+ len(positive_maps[0])).to(logits.device)
+ if positive_maps is not None:
+ if all(x == positive_maps[0] for x in positive_maps):
+ # only need to compute once
+ positive_map = positive_maps[0]
+ for label_j in positive_map:
+ scores[:, :, label_j -
+ 1] = logits[:, :,
+ torch.LongTensor(positive_map[label_j]
+ )].mean(-1)
+ else:
+ for i, positive_map in enumerate(positive_maps):
+ for label_j in positive_map:
+ scores[i, :, label_j - 1] = logits[
+ i, :, torch.LongTensor(positive_map[label_j])].mean(-1)
+ return scores
+
+
+class Conv3x3Norm(nn.Module):
+ """Conv3x3 and norm."""
+
+ def __init__(self,
+ in_channels: int,
+ out_channels: int,
+ stride: int,
+ groups: int = 1,
+ use_dcn: bool = False,
+ norm_type: Optional[Union[Sequence, str]] = None):
+ super().__init__()
+
+ if use_dcn:
+ self.conv = ModulatedDeformConv2d(
+ in_channels,
+ out_channels,
+ kernel_size=3,
+ stride=stride,
+ padding=1,
+ groups=groups)
+ else:
+ self.conv = nn.Conv2d(
+ in_channels,
+ out_channels,
+ kernel_size=3,
+ stride=stride,
+ padding=1,
+ groups=groups)
+
+ if isinstance(norm_type, Sequence):
+ assert len(norm_type) == 2
+ assert norm_type[0] == 'gn'
+ gn_group = norm_type[1]
+ norm_type = norm_type[0]
+
+ if norm_type == 'bn':
+ bn_op = nn.BatchNorm2d(out_channels)
+ elif norm_type == 'gn':
+ bn_op = nn.GroupNorm(
+ num_groups=gn_group, num_channels=out_channels)
+ if norm_type is not None:
+ self.bn = bn_op
+ else:
+ self.bn = None
+
+ def forward(self, x, **kwargs):
+ x = self.conv(x, **kwargs)
+ if self.bn:
+ x = self.bn(x)
+ return x
+
+
+class DyReLU(nn.Module):
+ """Dynamic ReLU."""
+
+ def __init__(self,
+ in_channels: int,
+ out_channels: int,
+ expand_ratio: int = 4):
+ super().__init__()
+ self.avg_pool = nn.AdaptiveAvgPool2d(1)
+ self.expand_ratio = expand_ratio
+ self.out_channels = out_channels
+
+ self.fc = nn.Sequential(
+ nn.Linear(in_channels, in_channels // expand_ratio),
+ nn.ReLU(inplace=True),
+ nn.Linear(in_channels // expand_ratio,
+ out_channels * self.expand_ratio),
+ nn.Hardsigmoid(inplace=True))
+
+ def forward(self, x) -> Tensor:
+ x_out = x
+ b, c, h, w = x.size()
+ x = self.avg_pool(x).view(b, c)
+ x = self.fc(x).view(b, -1, 1, 1)
+
+ a1, b1, a2, b2 = torch.split(x, self.out_channels, dim=1)
+ a1 = (a1 - 0.5) * 2 + 1.0
+ a2 = (a2 - 0.5) * 2
+ b1 = b1 - 0.5
+ b2 = b2 - 0.5
+ out = torch.max(x_out * a1 + b1, x_out * a2 + b2)
+ return out
+
+
+class DyConv(nn.Module):
+ """Dynamic Convolution."""
+
+ def __init__(self,
+ conv_func: Callable,
+ in_channels: int,
+ out_channels: int,
+ use_dyfuse: bool = True,
+ use_dyrelu: bool = False,
+ use_dcn: bool = False):
+ super().__init__()
+
+ self.dyconvs = nn.ModuleList()
+ self.dyconvs.append(conv_func(in_channels, out_channels, 1))
+ self.dyconvs.append(conv_func(in_channels, out_channels, 1))
+ self.dyconvs.append(conv_func(in_channels, out_channels, 2))
+
+ if use_dyfuse:
+ self.attnconv = nn.Sequential(
+ nn.AdaptiveAvgPool2d(1),
+ nn.Conv2d(in_channels, 1, kernel_size=1),
+ nn.ReLU(inplace=True))
+ self.h_sigmoid = nn.Hardsigmoid(inplace=True)
+ else:
+ self.attnconv = None
+
+ if use_dyrelu:
+ self.relu = DyReLU(in_channels, out_channels)
+ else:
+ self.relu = nn.ReLU()
+
+ if use_dcn:
+ self.offset = nn.Conv2d(
+ in_channels, 27, kernel_size=3, stride=1, padding=1)
+ else:
+ self.offset = None
+
+ self.init_weights()
+
+ def init_weights(self):
+ for m in self.dyconvs.modules():
+ if isinstance(m, nn.Conv2d):
+ nn.init.normal_(m.weight.data, 0, 0.01)
+ if m.bias is not None:
+ m.bias.data.zero_()
+ if self.attnconv is not None:
+ for m in self.attnconv.modules():
+ if isinstance(m, nn.Conv2d):
+ nn.init.normal_(m.weight.data, 0, 0.01)
+ if m.bias is not None:
+ m.bias.data.zero_()
+
+ def forward(self, inputs: dict) -> dict:
+ visual_feats = inputs['visual']
+
+ out_vis_feats = []
+ for level, feature in enumerate(visual_feats):
+
+ offset_conv_args = {}
+ if self.offset is not None:
+ offset_mask = self.offset(feature)
+ offset = offset_mask[:, :18, :, :]
+ mask = offset_mask[:, 18:, :, :].sigmoid()
+ offset_conv_args = dict(offset=offset, mask=mask)
+
+ temp_feats = [self.dyconvs[1](feature, **offset_conv_args)]
+
+ if level > 0:
+ temp_feats.append(self.dyconvs[2](visual_feats[level - 1],
+ **offset_conv_args))
+ if level < len(visual_feats) - 1:
+ temp_feats.append(
+ F.upsample_bilinear(
+ self.dyconvs[0](visual_feats[level + 1],
+ **offset_conv_args),
+ size=[feature.size(2),
+ feature.size(3)]))
+ mean_feats = torch.mean(
+ torch.stack(temp_feats), dim=0, keepdim=False)
+
+ if self.attnconv is not None:
+ attn_feat = []
+ res_feat = []
+ for feat in temp_feats:
+ res_feat.append(feat)
+ attn_feat.append(self.attnconv(feat))
+
+ res_feat = torch.stack(res_feat)
+ spa_pyr_attn = self.h_sigmoid(torch.stack(attn_feat))
+
+ mean_feats = torch.mean(
+ res_feat * spa_pyr_attn, dim=0, keepdim=False)
+
+ out_vis_feats.append(mean_feats)
+
+ out_vis_feats = [self.relu(item) for item in out_vis_feats]
+
+ features_dict = {'visual': out_vis_feats, 'lang': inputs['lang']}
+
+ return features_dict
+
+
+class VLFusionModule(BaseModel):
+ """Visual-lang Fusion Module."""
+
+ def __init__(self,
+ in_channels: int,
+ feat_channels: int,
+ num_base_priors: int,
+ early_fuse: bool = False,
+ num_dyhead_blocks: int = 6,
+ lang_model_name: str = 'bert-base-uncased',
+ use_dyrelu: bool = True,
+ use_dyfuse: bool = True,
+ use_dcn: bool = True,
+ use_checkpoint: bool = False,
+ **kwargs) -> None:
+ super().__init__(**kwargs)
+ if BertConfig is None:
+ raise RuntimeError(
+ 'transformers is not installed, please install it by: '
+ 'pip install transformers.')
+ self.in_channels = in_channels
+ self.feat_channels = feat_channels
+ self.num_base_priors = num_base_priors
+ self.early_fuse = early_fuse
+ self.num_dyhead_blocks = num_dyhead_blocks
+ self.use_dyrelu = use_dyrelu
+ self.use_dyfuse = use_dyfuse
+ self.use_dcn = use_dcn
+ self.use_checkpoint = use_checkpoint
+
+ self.lang_cfg = BertConfig.from_pretrained(lang_model_name)
+ self.lang_dim = self.lang_cfg.hidden_size
+ self._init_layers()
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the model."""
+ bias_value = -math.log((1 - 0.01) / 0.01)
+
+ dyhead_tower = []
+ for i in range(self.num_dyhead_blocks):
+ if self.early_fuse:
+ # cross-modality fusion
+ dyhead_tower.append(VLFuse(use_checkpoint=self.use_checkpoint))
+ # lang branch
+ dyhead_tower.append(
+ BertEncoderLayer(
+ self.lang_cfg,
+ clamp_min_for_underflow=True,
+ clamp_max_for_overflow=True))
+
+ # vision branch
+ dyhead_tower.append(
+ DyConv(
+ lambda i, o, s: Conv3x3Norm(
+ i, o, s, use_dcn=self.use_dcn, norm_type=['gn', 16]),
+ self.in_channels if i == 0 else self.feat_channels,
+ self.feat_channels,
+ use_dyrelu=(self.use_dyrelu
+ and self.in_channels == self.feat_channels)
+ if i == 0 else self.use_dyrelu,
+ use_dyfuse=(self.use_dyfuse
+ and self.in_channels == self.feat_channels)
+ if i == 0 else self.use_dyfuse,
+ use_dcn=(self.use_dcn
+ and self.in_channels == self.feat_channels)
+ if i == 0 else self.use_dcn,
+ ))
+
+ self.add_module('dyhead_tower', nn.Sequential(*dyhead_tower))
+
+ self.bbox_pred = nn.Conv2d(
+ self.feat_channels, self.num_base_priors * 4, kernel_size=1)
+ self.centerness = nn.Conv2d(
+ self.feat_channels, self.num_base_priors * 1, kernel_size=1)
+ self.dot_product_projection_text = nn.Linear(
+ self.lang_dim,
+ self.num_base_priors * self.feat_channels,
+ bias=True)
+ self.log_scale = nn.Parameter(torch.Tensor([0.0]), requires_grad=True)
+ self.bias_lang = nn.Parameter(
+ torch.zeros(self.lang_dim), requires_grad=True)
+ self.bias0 = nn.Parameter(
+ torch.Tensor([bias_value]), requires_grad=True)
+ self.scales = nn.ModuleList([Scale(1.0) for _ in range(5)])
+
+ def forward(self, visual_feats: Tuple[Tensor],
+ language_feats: dict) -> Tuple:
+ feat_inputs = {'visual': visual_feats, 'lang': language_feats}
+ dyhead_tower = self.dyhead_tower(feat_inputs)
+
+ if self.early_fuse:
+ embedding = dyhead_tower['lang']['hidden']
+ else:
+ embedding = language_feats['embedded']
+
+ embedding = F.normalize(embedding, p=2, dim=-1)
+ dot_product_proj_tokens = self.dot_product_projection_text(embedding /
+ 2.0)
+ dot_product_proj_tokens_bias = torch.matmul(
+ embedding, self.bias_lang) + self.bias0
+
+ bbox_preds = []
+ centerness = []
+ cls_logits = []
+
+ for i, feature in enumerate(visual_feats):
+ visual = dyhead_tower['visual'][i]
+ B, C, H, W = visual.shape
+
+ bbox_pred = self.scales[i](self.bbox_pred(visual))
+ bbox_preds.append(bbox_pred)
+ centerness.append(self.centerness(visual))
+
+ dot_product_proj_queries = permute_and_flatten(
+ visual, B, self.num_base_priors, C, H, W)
+
+ bias = dot_product_proj_tokens_bias.unsqueeze(1).repeat(
+ 1, self.num_base_priors, 1)
+ dot_product_logit = (
+ torch.matmul(dot_product_proj_queries,
+ dot_product_proj_tokens.transpose(-1, -2)) /
+ self.log_scale.exp()) + bias
+ dot_product_logit = torch.clamp(
+ dot_product_logit, max=MAX_CLAMP_VALUE)
+ dot_product_logit = torch.clamp(
+ dot_product_logit, min=-MAX_CLAMP_VALUE)
+ cls_logits.append(dot_product_logit)
+
+ return bbox_preds, centerness, cls_logits
+
+
+@MODELS.register_module()
+class ATSSVLFusionHead(ATSSHead):
+ """ATSS head with visual-language fusion module.
+
+ Args:
+ early_fuse (bool): Whether to fuse visual and language features
+ Defaults to False.
+ use_checkpoint (bool): Whether to use checkpoint. Defaults to False.
+ num_dyhead_blocks (int): Number of dynamic head blocks. Defaults to 6.
+ lang_model_name (str): Name of the language model.
+ Defaults to 'bert-base-uncased'.
+ """
+
+ def __init__(self,
+ *args,
+ early_fuse: bool = False,
+ use_checkpoint: bool = False,
+ num_dyhead_blocks: int = 6,
+ lang_model_name: str = 'bert-base-uncased',
+ init_cfg=None,
+ **kwargs):
+ super().__init__(*args, **kwargs, init_cfg=init_cfg)
+ self.head = VLFusionModule(
+ in_channels=self.in_channels,
+ feat_channels=self.feat_channels,
+ num_base_priors=self.num_base_priors,
+ early_fuse=early_fuse,
+ use_checkpoint=use_checkpoint,
+ num_dyhead_blocks=num_dyhead_blocks,
+ lang_model_name=lang_model_name)
+ self.text_masks = None
+
+ def _init_layers(self) -> None:
+ """No need to initialize the ATSS head layer."""
+ pass
+
+ def forward(self, visual_feats: Tuple[Tensor],
+ language_feats: dict) -> Tuple[Tensor]:
+ """Forward function."""
+ bbox_preds, centerness, cls_logits = self.head(visual_feats,
+ language_feats)
+ return cls_logits, bbox_preds, centerness
+
+ def loss(self, visual_feats: Tuple[Tensor], language_feats: dict,
+ batch_data_samples):
+ outputs = unpack_gt_instances(batch_data_samples)
+ (batch_gt_instances, batch_gt_instances_ignore,
+ batch_img_metas) = outputs
+
+ outs = self(visual_feats, language_feats)
+ self.text_masks = language_feats['masks']
+ loss_inputs = outs + (batch_gt_instances, batch_img_metas,
+ batch_gt_instances_ignore)
+ losses = self.loss_by_feat(*loss_inputs)
+ return losses
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ centernesses: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ Has shape (N, num_anchors * num_classes, H, W)
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W)
+ centernesses (list[Tensor]): Centerness for each scale
+ level with shape (N, num_anchors * 1, H, W)
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ featmap_sizes = [featmap.size()[-2:] for featmap in bbox_preds]
+ assert len(featmap_sizes) == self.prior_generator.num_levels
+
+ device = cls_scores[0].device
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+
+ cls_reg_targets = self.get_targets(
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore)
+
+ (anchor_list, labels_list, label_weights_list, bbox_targets_list,
+ bbox_weights_list, avg_factor) = cls_reg_targets
+ avg_factor = reduce_mean(
+ torch.tensor(avg_factor, dtype=torch.float, device=device)).item()
+
+ anchors = torch.cat(anchor_list, dim=1)
+ labels = torch.cat(labels_list, dim=1)
+ label_weights = torch.cat(label_weights_list, dim=1)
+ bbox_targets = torch.cat(bbox_targets_list, dim=1)
+ cls_scores = torch.cat(cls_scores, dim=1)
+
+ centernesses_ = []
+ bbox_preds_ = []
+ for bbox_pred, centerness in zip(bbox_preds, centernesses):
+ centernesses_.append(
+ centerness.permute(0, 2, 3,
+ 1).reshape(cls_scores.size(0), -1, 1))
+ bbox_preds_.append(
+ bbox_pred.permute(0, 2, 3,
+ 1).reshape(cls_scores.size(0), -1, 4))
+ bbox_preds = torch.cat(bbox_preds_, dim=1)
+ centernesses = torch.cat(centernesses_, dim=1)
+
+ losses_cls, losses_bbox, loss_centerness, bbox_avg_factor = \
+ self._loss_by_feat(
+ anchors,
+ cls_scores,
+ bbox_preds,
+ centernesses,
+ labels,
+ label_weights,
+ bbox_targets,
+ avg_factor=avg_factor)
+
+ bbox_avg_factor = reduce_mean(bbox_avg_factor).clamp_(min=1).item()
+ losses_bbox = losses_bbox / bbox_avg_factor
+ return dict(
+ loss_cls=losses_cls,
+ loss_bbox=losses_bbox,
+ loss_centerness=loss_centerness)
+
+ def _loss_by_feat(self, anchors: Tensor, cls_score: Tensor,
+ bbox_pred: Tensor, centerness: Tensor, labels: Tensor,
+ label_weights: Tensor, bbox_targets: Tensor,
+ avg_factor: float) -> dict:
+ """Calculate the loss of all scale level based on the features
+ extracted by the detection head.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+
+ anchors = anchors.reshape(-1, 4)
+
+ # ===== this change =====
+ pos_inds = (labels.sum(-1) > 0).reshape(-1)
+
+ # Loss is not computed for the padded regions of the text.
+ assert (self.text_masks.dim() == 2)
+ text_mask = (self.text_masks > 0).unsqueeze(1)
+ text_mask = text_mask.repeat(1, cls_score.size(1), 1)
+ cls_score = torch.masked_select(cls_score, text_mask).contiguous()
+ labels = torch.masked_select(labels, text_mask)
+ label_weights = label_weights[...,
+ None].repeat(1, 1, text_mask.size(-1))
+ label_weights = torch.masked_select(label_weights, text_mask)
+
+ bbox_pred = bbox_pred.reshape(-1, 4)
+ centerness = centerness.reshape(-1)
+ bbox_targets = bbox_targets.reshape(-1, 4)
+ labels = labels.reshape(-1)
+ label_weights = label_weights.reshape(-1)
+
+ # classification loss
+ loss_cls = self.loss_cls(
+ cls_score, labels, label_weights, avg_factor=avg_factor)
+
+ if pos_inds.sum() > 0:
+ pos_bbox_targets = bbox_targets[pos_inds]
+ pos_bbox_pred = bbox_pred[pos_inds]
+ pos_anchors = anchors[pos_inds]
+ pos_centerness = centerness[pos_inds]
+
+ centerness_targets = self.centerness_target(
+ pos_anchors, pos_bbox_targets)
+
+ if torch.isnan(centerness_targets).any():
+ print('=====Centerness includes NaN=====')
+ mask = ~torch.isnan(centerness_targets)
+ centerness_targets = centerness_targets[mask]
+ pos_centerness = pos_centerness[mask]
+ pos_anchors = pos_anchors[mask]
+ pos_bbox_targets = pos_bbox_targets[mask]
+ pos_bbox_pred = pos_bbox_pred[mask]
+
+ if pos_bbox_targets.shape[0] == 0:
+ loss_bbox = bbox_pred.sum() * 0
+ loss_centerness = centerness.sum() * 0
+ centerness_targets = bbox_targets.new_tensor(0.)
+ return loss_cls, loss_bbox, loss_centerness, \
+ centerness_targets.sum()
+
+ # The decoding process takes the offset into consideration.
+ pos_anchors[:, 2:] += 1
+ pos_decode_bbox_pred = self.bbox_coder.decode(
+ pos_anchors, pos_bbox_pred)
+
+ # regression loss
+ loss_bbox = self.loss_bbox(
+ pos_decode_bbox_pred,
+ pos_bbox_targets,
+ weight=centerness_targets,
+ avg_factor=1.0)
+
+ # centerness loss
+ loss_centerness = self.loss_centerness(
+ pos_centerness, centerness_targets, avg_factor=avg_factor)
+ else:
+ loss_bbox = bbox_pred.sum() * 0
+ loss_centerness = centerness.sum() * 0
+ centerness_targets = bbox_targets.new_tensor(0.)
+
+ return loss_cls, loss_bbox, loss_centerness, centerness_targets.sum()
+
+ def _get_targets_single(self,
+ flat_anchors: Tensor,
+ valid_flags: Tensor,
+ num_level_anchors: List[int],
+ gt_instances: InstanceData,
+ img_meta: dict,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ unmap_outputs: bool = True) -> tuple:
+ """Compute regression, classification targets for anchors in a single
+ image.
+
+ Args:
+ flat_anchors (Tensor): Multi-level anchors of the image, which are
+ concatenated into a single tensor of shape (num_anchors ,4)
+ valid_flags (Tensor): Multi level valid flags of the image,
+ which are concatenated into a single tensor of
+ shape (num_anchors,).
+ num_level_anchors (List[int]): Number of anchors of each scale
+ level.
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ img_meta (dict): Meta information for current image.
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ unmap_outputs (bool): Whether to map outputs back to the original
+ set of anchors.
+
+ Returns:
+ tuple: N is the number of total anchors in the image.
+ labels (Tensor): Labels of all anchors in the image with shape
+ (N,).
+ label_weights (Tensor): Label weights of all anchor in the
+ image with shape (N,).
+ bbox_targets (Tensor): BBox targets of all anchors in the
+ image with shape (N, 4).
+ bbox_weights (Tensor): BBox weights of all anchors in the
+ image with shape (N, 4)
+ pos_inds (Tensor): Indices of positive anchor with shape
+ (num_pos,).
+ neg_inds (Tensor): Indices of negative anchor with shape
+ (num_neg,).
+ sampling_result (:obj:`SamplingResult`): Sampling results.
+ """
+ anchors = flat_anchors
+ # Align the official implementation
+ anchors[:, 2:] -= 1
+
+ num_level_anchors_inside = num_level_anchors
+ pred_instances = InstanceData(priors=anchors)
+ assign_result = self.assigner.assign(pred_instances,
+ num_level_anchors_inside,
+ gt_instances, gt_instances_ignore)
+
+ sampling_result = self.sampler.sample(assign_result, pred_instances,
+ gt_instances)
+
+ num_valid_anchors = anchors.shape[0]
+ bbox_targets = torch.zeros_like(anchors)
+ bbox_weights = torch.zeros_like(anchors)
+
+ # ===== this change =====
+ labels = anchors.new_full((num_valid_anchors, self.feat_channels),
+ 0,
+ dtype=torch.float32)
+ label_weights = anchors.new_zeros(num_valid_anchors, dtype=torch.float)
+ pos_inds = sampling_result.pos_inds
+ neg_inds = sampling_result.neg_inds
+ if len(pos_inds) > 0:
+ if self.reg_decoded_bbox:
+ pos_bbox_targets = sampling_result.pos_gt_bboxes
+ else:
+ pos_bbox_targets = self.bbox_coder.encode(
+ sampling_result.pos_priors, sampling_result.pos_gt_bboxes)
+
+ bbox_targets[pos_inds, :] = pos_bbox_targets
+ bbox_weights[pos_inds, :] = 1.0
+
+ # ===== this change =====
+ labels[pos_inds] = gt_instances.positive_maps[
+ sampling_result.pos_assigned_gt_inds]
+ if self.train_cfg['pos_weight'] <= 0:
+ label_weights[pos_inds] = 1.0
+ else:
+ label_weights[pos_inds] = self.train_cfg['pos_weight']
+ if len(neg_inds) > 0:
+ label_weights[neg_inds] = 1.0
+
+ return (anchors, labels, label_weights, bbox_targets, bbox_weights,
+ pos_inds, neg_inds, sampling_result)
+
+ def centerness_target(self, anchors: Tensor, gts: Tensor) -> Tensor:
+ """Calculate the centerness between anchors and gts.
+
+ Only calculate pos centerness targets, otherwise there may be nan.
+
+ Args:
+ anchors (Tensor): Anchors with shape (N, 4), "xyxy" format.
+ gts (Tensor): Ground truth bboxes with shape (N, 4), "xyxy" format.
+
+ Returns:
+ Tensor: Centerness between anchors and gts.
+ """
+ anchors_cx = (anchors[:, 2] + anchors[:, 0]) / 2
+ anchors_cy = (anchors[:, 3] + anchors[:, 1]) / 2
+ l_ = anchors_cx - gts[:, 0]
+ t_ = anchors_cy - gts[:, 1]
+ r_ = gts[:, 2] - anchors_cx
+ b_ = gts[:, 3] - anchors_cy
+
+ left_right = torch.stack([l_, r_], dim=1)
+ top_bottom = torch.stack([t_, b_], dim=1)
+ centerness = torch.sqrt(
+ (left_right.min(dim=-1)[0] / left_right.max(dim=-1)[0]) *
+ (top_bottom.min(dim=-1)[0] / top_bottom.max(dim=-1)[0]))
+ # assert not torch.isnan(centerness).any()
+ return centerness
+
+ def predict(self,
+ visual_feats: Tuple[Tensor],
+ language_feats: dict,
+ batch_data_samples,
+ rescale: bool = True):
+ """Perform forward propagation of the detection head and predict
+ detection results on the features of the upstream network.
+
+ Args:
+ visual_feats (tuple[Tensor]): Multi-level visual features from the
+ upstream network, each is a 4D-tensor.
+ language_feats (dict): Language features from the upstream network.
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool, optional): Whether to rescale the results.
+ Defaults to False.
+
+ Returns:
+ list[obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ """
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+ batch_token_positive_maps = [
+ data_samples.token_positive_map
+ for data_samples in batch_data_samples
+ ]
+ outs = self(visual_feats, language_feats)
+
+ predictions = self.predict_by_feat(
+ *outs,
+ batch_img_metas=batch_img_metas,
+ batch_token_positive_maps=batch_token_positive_maps,
+ rescale=rescale)
+ return predictions
+
+ def predict_by_feat(self,
+ cls_logits: List[Tensor],
+ bbox_preds: List[Tensor],
+ score_factors: List[Tensor],
+ batch_img_metas: Optional[List[dict]] = None,
+ batch_token_positive_maps: Optional[List[dict]] = None,
+ cfg: Optional[ConfigDict] = None,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ bbox results.
+
+ Note: When score_factors is not None, the cls_scores are
+ usually multiplied by it then obtain the real score used in NMS,
+ such as CenterNess in FCOS, IoU branch in ATSS.
+
+ Args:
+ cls_logits (list[Tensor]): Classification scores for all
+ scale levels, each is a 4D-tensor, has shape
+ (batch_size, num_priors * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas for all
+ scale levels, each is a 4D-tensor, has shape
+ (batch_size, num_priors * 4, H, W).
+ score_factors (list[Tensor], optional): Score factor for
+ all scale level, each is a 4D-tensor, has shape
+ (batch_size, num_priors * 1, H, W). Defaults to None.
+ batch_img_metas (list[dict], Optional): Batch image meta info.
+ Defaults to None.
+ batch_token_positive_maps (list[dict], Optional): Batch token
+ positive map. Defaults to None.
+ cfg (ConfigDict, optional): Test / postprocessing
+ configuration, if None, test_cfg would be used.
+ Defaults to None.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ list[:obj:`InstanceData`]: Object detection results of each image
+ after the post process. Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ assert len(bbox_preds) == len(score_factors)
+ num_levels = len(bbox_preds)
+
+ featmap_sizes = [bbox_preds[i].shape[-2:] for i in range(num_levels)]
+ mlvl_priors = self.prior_generator.grid_priors(
+ featmap_sizes,
+ dtype=bbox_preds[0].dtype,
+ device=bbox_preds[0].device)
+
+ result_list = []
+
+ for img_id in range(len(batch_img_metas)):
+ img_meta = batch_img_metas[img_id]
+ token_positive_maps = batch_token_positive_maps[img_id]
+ bbox_pred_list = select_single_mlvl(
+ bbox_preds, img_id, detach=True)
+ score_factor_list = select_single_mlvl(
+ score_factors, img_id, detach=True)
+ cls_logit_list = select_single_mlvl(
+ cls_logits, img_id, detach=True)
+
+ results = self._predict_by_feat_single(
+ bbox_pred_list=bbox_pred_list,
+ score_factor_list=score_factor_list,
+ cls_logit_list=cls_logit_list,
+ mlvl_priors=mlvl_priors,
+ token_positive_maps=token_positive_maps,
+ img_meta=img_meta,
+ cfg=cfg,
+ rescale=rescale,
+ with_nms=with_nms)
+ result_list.append(results)
+ return result_list
+
+ def _predict_by_feat_single(self,
+ bbox_pred_list: List[Tensor],
+ score_factor_list: List[Tensor],
+ cls_logit_list: List[Tensor],
+ mlvl_priors: List[Tensor],
+ token_positive_maps: dict,
+ img_meta: dict,
+ cfg: ConfigDict,
+ rescale: bool = True,
+ with_nms: bool = True) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results.
+
+ Args:
+ bbox_pred_list (list[Tensor]): Box energies / deltas from
+ all scale levels of a single image, each item has shape
+ (num_priors * 4, H, W).
+ score_factor_list (list[Tensor]): Score factor from all scale
+ levels of a single image, each item has shape
+ (num_priors * 1, H, W).
+ cls_logit_list (list[Tensor]): Box scores from all scale
+ levels of a single image, each item has shape
+ (num_priors * num_classes, H, W).
+ mlvl_priors (list[Tensor]): Each element in the list is
+ the priors of a single level in feature pyramid. In all
+ anchor-based methods, it has shape (num_priors, 4). In
+ all anchor-free methods, it has shape (num_priors, 2)
+ when `with_stride=True`, otherwise it still has shape
+ (num_priors, 4).
+ token_positive_maps (dict): Token positive map.
+ img_meta (dict): Image meta info.
+ cfg (mmengine.Config): Test / postprocessing configuration,
+ if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ cfg = self.test_cfg if cfg is None else cfg
+ cfg = copy.deepcopy(cfg)
+ img_shape = img_meta['img_shape']
+ nms_pre = cfg.get('nms_pre', -1)
+ score_thr = cfg.get('score_thr', 0)
+
+ mlvl_bbox_preds = []
+ mlvl_valid_priors = []
+ mlvl_scores = []
+ mlvl_labels = []
+
+ for level_idx, (bbox_pred, score_factor, cls_logit, priors) in \
+ enumerate(zip(bbox_pred_list,
+ score_factor_list, cls_logit_list, mlvl_priors)):
+ bbox_pred = bbox_pred.permute(1, 2, 0).reshape(
+ -1, self.bbox_coder.encode_size)
+ score_factor = score_factor.permute(1, 2, 0).reshape(-1).sigmoid()
+
+ scores = convert_grounding_to_cls_scores(
+ logits=cls_logit.sigmoid()[None],
+ positive_maps=[token_positive_maps])[0]
+
+ results = filter_scores_and_topk(
+ scores, score_thr, nms_pre,
+ dict(bbox_pred=bbox_pred, priors=priors))
+
+ scores, labels, keep_idxs, filtered_results = results
+
+ bbox_pred = filtered_results['bbox_pred']
+ priors = filtered_results['priors']
+ score_factor = score_factor[keep_idxs]
+ scores = torch.sqrt(scores * score_factor)
+
+ mlvl_bbox_preds.append(bbox_pred)
+ mlvl_valid_priors.append(priors)
+ mlvl_scores.append(scores)
+ mlvl_labels.append(labels)
+
+ bbox_pred = torch.cat(mlvl_bbox_preds)
+ priors = cat_boxes(mlvl_valid_priors)
+ bboxes = self.bbox_coder.decode(priors, bbox_pred, max_shape=img_shape)
+
+ results = InstanceData()
+ results.bboxes = bboxes
+ results.scores = torch.cat(mlvl_scores)
+ results.labels = torch.cat(mlvl_labels)
+
+ predictions = self._bbox_post_process(
+ results=results,
+ cfg=cfg,
+ rescale=rescale,
+ with_nms=with_nms,
+ img_meta=img_meta)
+
+ if len(predictions) > 0:
+ # Note: GLIP adopts a very strange bbox decoder logic,
+ # and if 1 is not added here, it will not align with
+ # the official mAP.
+ predictions.bboxes[:, 2:] = predictions.bboxes[:, 2:] + 1
+ return predictions
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/autoassign_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/autoassign_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..a2b30ff0d7d41205f0a92ede7b8eb10a234c5942
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/autoassign_head.py
@@ -0,0 +1,524 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, List, Sequence, Tuple
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import Scale
+from mmengine.model import bias_init_with_prob, normal_init
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures.bbox import bbox_overlaps
+from mmdet.utils import InstanceList, OptInstanceList, reduce_mean
+from ..task_modules.prior_generators import MlvlPointGenerator
+from ..utils import levels_to_images, multi_apply
+from .fcos_head import FCOSHead
+
+EPS = 1e-12
+
+
+class CenterPrior(nn.Module):
+ """Center Weighting module to adjust the category-specific prior
+ distributions.
+
+ Args:
+ force_topk (bool): When no point falls into gt_bbox, forcibly
+ select the k points closest to the center to calculate
+ the center prior. Defaults to False.
+ topk (int): The number of points used to calculate the
+ center prior when no point falls in gt_bbox. Only work when
+ force_topk if True. Defaults to 9.
+ num_classes (int): The class number of dataset. Defaults to 80.
+ strides (Sequence[int]): The stride of each input feature map.
+ Defaults to (8, 16, 32, 64, 128).
+ """
+
+ def __init__(
+ self,
+ force_topk: bool = False,
+ topk: int = 9,
+ num_classes: int = 80,
+ strides: Sequence[int] = (8, 16, 32, 64, 128)
+ ) -> None:
+ super().__init__()
+ self.mean = nn.Parameter(torch.zeros(num_classes, 2))
+ self.sigma = nn.Parameter(torch.ones(num_classes, 2))
+ self.strides = strides
+ self.force_topk = force_topk
+ self.topk = topk
+
+ def forward(self, anchor_points_list: List[Tensor],
+ gt_instances: InstanceData,
+ inside_gt_bbox_mask: Tensor) -> Tuple[Tensor, Tensor]:
+ """Get the center prior of each point on the feature map for each
+ instance.
+
+ Args:
+ anchor_points_list (list[Tensor]): list of coordinate
+ of points on feature map. Each with shape
+ (num_points, 2).
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes`` and ``labels``
+ attributes.
+ inside_gt_bbox_mask (Tensor): Tensor of bool type,
+ with shape of (num_points, num_gt), each
+ value is used to mark whether this point falls
+ within a certain gt.
+
+ Returns:
+ tuple[Tensor, Tensor]:
+
+ - center_prior_weights(Tensor): Float tensor with shape of \
+ (num_points, num_gt). Each value represents the center \
+ weighting coefficient.
+ - inside_gt_bbox_mask (Tensor): Tensor of bool type, with shape \
+ of (num_points, num_gt), each value is used to mark whether this \
+ point falls within a certain gt or is the topk nearest points for \
+ a specific gt_bbox.
+ """
+ gt_bboxes = gt_instances.bboxes
+ labels = gt_instances.labels
+
+ inside_gt_bbox_mask = inside_gt_bbox_mask.clone()
+ num_gts = len(labels)
+ num_points = sum([len(item) for item in anchor_points_list])
+ if num_gts == 0:
+ return gt_bboxes.new_zeros(num_points,
+ num_gts), inside_gt_bbox_mask
+ center_prior_list = []
+ for slvl_points, stride in zip(anchor_points_list, self.strides):
+ # slvl_points: points from single level in FPN, has shape (h*w, 2)
+ # single_level_points has shape (h*w, num_gt, 2)
+ single_level_points = slvl_points[:, None, :].expand(
+ (slvl_points.size(0), len(gt_bboxes), 2))
+ gt_center_x = ((gt_bboxes[:, 0] + gt_bboxes[:, 2]) / 2)
+ gt_center_y = ((gt_bboxes[:, 1] + gt_bboxes[:, 3]) / 2)
+ gt_center = torch.stack((gt_center_x, gt_center_y), dim=1)
+ gt_center = gt_center[None]
+ # instance_center has shape (1, num_gt, 2)
+ instance_center = self.mean[labels][None]
+ # instance_sigma has shape (1, num_gt, 2)
+ instance_sigma = self.sigma[labels][None]
+ # distance has shape (num_points, num_gt, 2)
+ distance = (((single_level_points - gt_center) / float(stride) -
+ instance_center)**2)
+ center_prior = torch.exp(-distance /
+ (2 * instance_sigma**2)).prod(dim=-1)
+ center_prior_list.append(center_prior)
+ center_prior_weights = torch.cat(center_prior_list, dim=0)
+
+ if self.force_topk:
+ gt_inds_no_points_inside = torch.nonzero(
+ inside_gt_bbox_mask.sum(0) == 0).reshape(-1)
+ if gt_inds_no_points_inside.numel():
+ topk_center_index = \
+ center_prior_weights[:, gt_inds_no_points_inside].topk(
+ self.topk,
+ dim=0)[1]
+ temp_mask = inside_gt_bbox_mask[:, gt_inds_no_points_inside]
+ inside_gt_bbox_mask[:, gt_inds_no_points_inside] = \
+ torch.scatter(temp_mask,
+ dim=0,
+ index=topk_center_index,
+ src=torch.ones_like(
+ topk_center_index,
+ dtype=torch.bool))
+
+ center_prior_weights[~inside_gt_bbox_mask] = 0
+ return center_prior_weights, inside_gt_bbox_mask
+
+
+@MODELS.register_module()
+class AutoAssignHead(FCOSHead):
+ """AutoAssignHead head used in AutoAssign.
+
+ More details can be found in the `paper
+ `_ .
+
+ Args:
+ force_topk (bool): Used in center prior initialization to
+ handle extremely small gt. Default is False.
+ topk (int): The number of points used to calculate the
+ center prior when no point falls in gt_bbox. Only work when
+ force_topk if True. Defaults to 9.
+ pos_loss_weight (float): The loss weight of positive loss
+ and with default value 0.25.
+ neg_loss_weight (float): The loss weight of negative loss
+ and with default value 0.75.
+ center_loss_weight (float): The loss weight of center prior
+ loss and with default value 0.75.
+ """
+
+ def __init__(self,
+ *args,
+ force_topk: bool = False,
+ topk: int = 9,
+ pos_loss_weight: float = 0.25,
+ neg_loss_weight: float = 0.75,
+ center_loss_weight: float = 0.75,
+ **kwargs) -> None:
+ super().__init__(*args, conv_bias=True, **kwargs)
+ self.center_prior = CenterPrior(
+ force_topk=force_topk,
+ topk=topk,
+ num_classes=self.num_classes,
+ strides=self.strides)
+ self.pos_loss_weight = pos_loss_weight
+ self.neg_loss_weight = neg_loss_weight
+ self.center_loss_weight = center_loss_weight
+ self.prior_generator = MlvlPointGenerator(self.strides, offset=0)
+
+ def init_weights(self) -> None:
+ """Initialize weights of the head.
+
+ In particular, we have special initialization for classified conv's and
+ regression conv's bias
+ """
+
+ super(AutoAssignHead, self).init_weights()
+ bias_cls = bias_init_with_prob(0.02)
+ normal_init(self.conv_cls, std=0.01, bias=bias_cls)
+ normal_init(self.conv_reg, std=0.01, bias=4.0)
+
+ def forward_single(self, x: Tensor, scale: Scale,
+ stride: int) -> Tuple[Tensor, Tensor, Tensor]:
+ """Forward features of a single scale level.
+
+ Args:
+ x (Tensor): FPN feature maps of the specified stride.
+ scale (:obj:`mmcv.cnn.Scale`): Learnable scale module to resize
+ the bbox prediction.
+ stride (int): The corresponding stride for feature maps, only
+ used to normalize the bbox prediction when self.norm_on_bbox
+ is True.
+
+ Returns:
+ tuple[Tensor, Tensor, Tensor]: scores for each class, bbox
+ predictions and centerness predictions of input feature maps.
+ """
+ cls_score, bbox_pred, cls_feat, reg_feat = super(
+ FCOSHead, self).forward_single(x)
+ centerness = self.conv_centerness(reg_feat)
+ # scale the bbox_pred of different level
+ # float to avoid overflow when enabling FP16
+ bbox_pred = scale(bbox_pred).float()
+ # bbox_pred needed for gradient computation has been modified
+ # by F.relu(bbox_pred) when run with PyTorch 1.10. So replace
+ # F.relu(bbox_pred) with bbox_pred.clamp(min=0)
+ bbox_pred = bbox_pred.clamp(min=0)
+ bbox_pred *= stride
+ return cls_score, bbox_pred, centerness
+
+ def get_pos_loss_single(self, cls_score: Tensor, objectness: Tensor,
+ reg_loss: Tensor, gt_instances: InstanceData,
+ center_prior_weights: Tensor) -> Tuple[Tensor]:
+ """Calculate the positive loss of all points in gt_bboxes.
+
+ Args:
+ cls_score (Tensor): All category scores for each point on
+ the feature map. The shape is (num_points, num_class).
+ objectness (Tensor): Foreground probability of all points,
+ has shape (num_points, 1).
+ reg_loss (Tensor): The regression loss of each gt_bbox and each
+ prediction box, has shape of (num_points, num_gt).
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes`` and ``labels``
+ attributes.
+ center_prior_weights (Tensor): Float tensor with shape
+ of (num_points, num_gt). Each value represents
+ the center weighting coefficient.
+
+ Returns:
+ tuple[Tensor]:
+
+ - pos_loss (Tensor): The positive loss of all points in the \
+ gt_bboxes.
+ """
+ gt_labels = gt_instances.labels
+ # p_loc: localization confidence
+ p_loc = torch.exp(-reg_loss)
+ # p_cls: classification confidence
+ p_cls = (cls_score * objectness)[:, gt_labels]
+ # p_pos: joint confidence indicator
+ p_pos = p_cls * p_loc
+
+ # 3 is a hyper-parameter to control the contributions of high and
+ # low confidence locations towards positive losses.
+ confidence_weight = torch.exp(p_pos * 3)
+ p_pos_weight = (confidence_weight * center_prior_weights) / (
+ (confidence_weight * center_prior_weights).sum(
+ 0, keepdim=True)).clamp(min=EPS)
+ reweighted_p_pos = (p_pos * p_pos_weight).sum(0)
+ pos_loss = F.binary_cross_entropy(
+ reweighted_p_pos,
+ torch.ones_like(reweighted_p_pos),
+ reduction='none')
+ pos_loss = pos_loss.sum() * self.pos_loss_weight
+ return pos_loss,
+
+ def get_neg_loss_single(self, cls_score: Tensor, objectness: Tensor,
+ gt_instances: InstanceData, ious: Tensor,
+ inside_gt_bbox_mask: Tensor) -> Tuple[Tensor]:
+ """Calculate the negative loss of all points in feature map.
+
+ Args:
+ cls_score (Tensor): All category scores for each point on
+ the feature map. The shape is (num_points, num_class).
+ objectness (Tensor): Foreground probability of all points
+ and is shape of (num_points, 1).
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes`` and ``labels``
+ attributes.
+ ious (Tensor): Float tensor with shape of (num_points, num_gt).
+ Each value represent the iou of pred_bbox and gt_bboxes.
+ inside_gt_bbox_mask (Tensor): Tensor of bool type,
+ with shape of (num_points, num_gt), each
+ value is used to mark whether this point falls
+ within a certain gt.
+
+ Returns:
+ tuple[Tensor]:
+
+ - neg_loss (Tensor): The negative loss of all points in the \
+ feature map.
+ """
+ gt_labels = gt_instances.labels
+ num_gts = len(gt_labels)
+ joint_conf = (cls_score * objectness)
+ p_neg_weight = torch.ones_like(joint_conf)
+ if num_gts > 0:
+ # the order of dinmension would affect the value of
+ # p_neg_weight, we strictly follow the original
+ # implementation.
+ inside_gt_bbox_mask = inside_gt_bbox_mask.permute(1, 0)
+ ious = ious.permute(1, 0)
+
+ foreground_idxs = torch.nonzero(inside_gt_bbox_mask, as_tuple=True)
+ temp_weight = (1 / (1 - ious[foreground_idxs]).clamp_(EPS))
+
+ def normalize(x):
+ return (x - x.min() + EPS) / (x.max() - x.min() + EPS)
+
+ for instance_idx in range(num_gts):
+ idxs = foreground_idxs[0] == instance_idx
+ if idxs.any():
+ temp_weight[idxs] = normalize(temp_weight[idxs])
+
+ p_neg_weight[foreground_idxs[1],
+ gt_labels[foreground_idxs[0]]] = 1 - temp_weight
+
+ logits = (joint_conf * p_neg_weight)
+ neg_loss = (
+ logits**2 * F.binary_cross_entropy(
+ logits, torch.zeros_like(logits), reduction='none'))
+ neg_loss = neg_loss.sum() * self.neg_loss_weight
+ return neg_loss,
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ objectnesses: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None
+ ) -> Dict[str, Tensor]:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level,
+ each is a 4D-tensor, the channel number is
+ num_points * num_classes.
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level, each is a 4D-tensor, the channel number is
+ num_points * 4.
+ objectnesses (list[Tensor]): objectness for each scale level, each
+ is a 4D-tensor, the channel number is num_points * 1.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+
+ assert len(cls_scores) == len(bbox_preds) == len(objectnesses)
+ all_num_gt = sum([len(item) for item in batch_gt_instances])
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ all_level_points = self.prior_generator.grid_priors(
+ featmap_sizes,
+ dtype=bbox_preds[0].dtype,
+ device=bbox_preds[0].device)
+ inside_gt_bbox_mask_list, bbox_targets_list = self.get_targets(
+ all_level_points, batch_gt_instances)
+
+ center_prior_weight_list = []
+ temp_inside_gt_bbox_mask_list = []
+ for gt_instances, inside_gt_bbox_mask in zip(batch_gt_instances,
+ inside_gt_bbox_mask_list):
+ center_prior_weight, inside_gt_bbox_mask = \
+ self.center_prior(all_level_points, gt_instances,
+ inside_gt_bbox_mask)
+ center_prior_weight_list.append(center_prior_weight)
+ temp_inside_gt_bbox_mask_list.append(inside_gt_bbox_mask)
+ inside_gt_bbox_mask_list = temp_inside_gt_bbox_mask_list
+ mlvl_points = torch.cat(all_level_points, dim=0)
+ bbox_preds = levels_to_images(bbox_preds)
+ cls_scores = levels_to_images(cls_scores)
+ objectnesses = levels_to_images(objectnesses)
+
+ reg_loss_list = []
+ ious_list = []
+ num_points = len(mlvl_points)
+
+ for bbox_pred, encoded_targets, inside_gt_bbox_mask in zip(
+ bbox_preds, bbox_targets_list, inside_gt_bbox_mask_list):
+ temp_num_gt = encoded_targets.size(1)
+ expand_mlvl_points = mlvl_points[:, None, :].expand(
+ num_points, temp_num_gt, 2).reshape(-1, 2)
+ encoded_targets = encoded_targets.reshape(-1, 4)
+ expand_bbox_pred = bbox_pred[:, None, :].expand(
+ num_points, temp_num_gt, 4).reshape(-1, 4)
+ decoded_bbox_preds = self.bbox_coder.decode(
+ expand_mlvl_points, expand_bbox_pred)
+ decoded_target_preds = self.bbox_coder.decode(
+ expand_mlvl_points, encoded_targets)
+ with torch.no_grad():
+ ious = bbox_overlaps(
+ decoded_bbox_preds, decoded_target_preds, is_aligned=True)
+ ious = ious.reshape(num_points, temp_num_gt)
+ if temp_num_gt:
+ ious = ious.max(
+ dim=-1, keepdim=True).values.repeat(1, temp_num_gt)
+ else:
+ ious = ious.new_zeros(num_points, temp_num_gt)
+ ious[~inside_gt_bbox_mask] = 0
+ ious_list.append(ious)
+ loss_bbox = self.loss_bbox(
+ decoded_bbox_preds,
+ decoded_target_preds,
+ weight=None,
+ reduction_override='none')
+ reg_loss_list.append(loss_bbox.reshape(num_points, temp_num_gt))
+
+ cls_scores = [item.sigmoid() for item in cls_scores]
+ objectnesses = [item.sigmoid() for item in objectnesses]
+ pos_loss_list, = multi_apply(self.get_pos_loss_single, cls_scores,
+ objectnesses, reg_loss_list,
+ batch_gt_instances,
+ center_prior_weight_list)
+ pos_avg_factor = reduce_mean(
+ bbox_pred.new_tensor(all_num_gt)).clamp_(min=1)
+ pos_loss = sum(pos_loss_list) / pos_avg_factor
+
+ neg_loss_list, = multi_apply(self.get_neg_loss_single, cls_scores,
+ objectnesses, batch_gt_instances,
+ ious_list, inside_gt_bbox_mask_list)
+ neg_avg_factor = sum(item.data.sum()
+ for item in center_prior_weight_list)
+ neg_avg_factor = reduce_mean(neg_avg_factor).clamp_(min=1)
+ neg_loss = sum(neg_loss_list) / neg_avg_factor
+
+ center_loss = []
+ for i in range(len(batch_img_metas)):
+
+ if inside_gt_bbox_mask_list[i].any():
+ center_loss.append(
+ len(batch_gt_instances[i]) /
+ center_prior_weight_list[i].sum().clamp_(min=EPS))
+ # when width or height of gt_bbox is smaller than stride of p3
+ else:
+ center_loss.append(center_prior_weight_list[i].sum() * 0)
+
+ center_loss = torch.stack(center_loss).mean() * self.center_loss_weight
+
+ # avoid dead lock in DDP
+ if all_num_gt == 0:
+ pos_loss = bbox_preds[0].sum() * 0
+ dummy_center_prior_loss = self.center_prior.mean.sum(
+ ) * 0 + self.center_prior.sigma.sum() * 0
+ center_loss = objectnesses[0].sum() * 0 + dummy_center_prior_loss
+
+ loss = dict(
+ loss_pos=pos_loss, loss_neg=neg_loss, loss_center=center_loss)
+
+ return loss
+
+ def get_targets(
+ self, points: List[Tensor], batch_gt_instances: InstanceList
+ ) -> Tuple[List[Tensor], List[Tensor]]:
+ """Compute regression targets and each point inside or outside gt_bbox
+ in multiple images.
+
+ Args:
+ points (list[Tensor]): Points of all fpn level, each has shape
+ (num_points, 2).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+
+ Returns:
+ tuple(list[Tensor], list[Tensor]):
+
+ - inside_gt_bbox_mask_list (list[Tensor]): Each Tensor is with \
+ bool type and shape of (num_points, num_gt), each value is used \
+ to mark whether this point falls within a certain gt.
+ - concat_lvl_bbox_targets (list[Tensor]): BBox targets of each \
+ level. Each tensor has shape (num_points, num_gt, 4).
+ """
+
+ concat_points = torch.cat(points, dim=0)
+ # the number of points per img, per lvl
+ inside_gt_bbox_mask_list, bbox_targets_list = multi_apply(
+ self._get_targets_single, batch_gt_instances, points=concat_points)
+ return inside_gt_bbox_mask_list, bbox_targets_list
+
+ def _get_targets_single(self, gt_instances: InstanceData,
+ points: Tensor) -> Tuple[Tensor, Tensor]:
+ """Compute regression targets and each point inside or outside gt_bbox
+ for a single image.
+
+ Args:
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes`` and ``labels``
+ attributes.
+ points (Tensor): Points of all fpn level, has shape
+ (num_points, 2).
+
+ Returns:
+ tuple[Tensor, Tensor]: Containing the following Tensors:
+
+ - inside_gt_bbox_mask (Tensor): Bool tensor with shape \
+ (num_points, num_gt), each value is used to mark whether this \
+ point falls within a certain gt.
+ - bbox_targets (Tensor): BBox targets of each points with each \
+ gt_bboxes, has shape (num_points, num_gt, 4).
+ """
+ gt_bboxes = gt_instances.bboxes
+ num_points = points.size(0)
+ num_gts = gt_bboxes.size(0)
+ gt_bboxes = gt_bboxes[None].expand(num_points, num_gts, 4)
+ xs, ys = points[:, 0], points[:, 1]
+ xs = xs[:, None]
+ ys = ys[:, None]
+ left = xs - gt_bboxes[..., 0]
+ right = gt_bboxes[..., 2] - xs
+ top = ys - gt_bboxes[..., 1]
+ bottom = gt_bboxes[..., 3] - ys
+ bbox_targets = torch.stack((left, top, right, bottom), -1)
+ if num_gts:
+ inside_gt_bbox_mask = bbox_targets.min(-1)[0] > 0
+ else:
+ inside_gt_bbox_mask = bbox_targets.new_zeros((num_points, num_gts),
+ dtype=torch.bool)
+
+ return inside_gt_bbox_mask, bbox_targets
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/base_dense_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/base_dense_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..d0a4469e02c469d029cc2791289dbf41554d6a53
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/base_dense_head.py
@@ -0,0 +1,583 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from abc import ABCMeta, abstractmethod
+from inspect import signature
+from typing import List, Optional, Tuple
+
+import torch
+from mmcv.ops import batched_nms
+from mmengine.config import ConfigDict
+from mmengine.model import BaseModule, constant_init
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import (cat_boxes, get_box_tensor, get_box_wh,
+ scale_boxes)
+from mmdet.utils import InstanceList, OptMultiConfig
+from ..test_time_augs import merge_aug_results
+from ..utils import (filter_scores_and_topk, select_single_mlvl,
+ unpack_gt_instances)
+
+
+class BaseDenseHead(BaseModule, metaclass=ABCMeta):
+ """Base class for DenseHeads.
+
+ 1. The ``init_weights`` method is used to initialize densehead's
+ model parameters. After detector initialization, ``init_weights``
+ is triggered when ``detector.init_weights()`` is called externally.
+
+ 2. The ``loss`` method is used to calculate the loss of densehead,
+ which includes two steps: (1) the densehead model performs forward
+ propagation to obtain the feature maps (2) The ``loss_by_feat`` method
+ is called based on the feature maps to calculate the loss.
+
+ .. code:: text
+
+ loss(): forward() -> loss_by_feat()
+
+ 3. The ``predict`` method is used to predict detection results,
+ which includes two steps: (1) the densehead model performs forward
+ propagation to obtain the feature maps (2) The ``predict_by_feat`` method
+ is called based on the feature maps to predict detection results including
+ post-processing.
+
+ .. code:: text
+
+ predict(): forward() -> predict_by_feat()
+
+ 4. The ``loss_and_predict`` method is used to return loss and detection
+ results at the same time. It will call densehead's ``forward``,
+ ``loss_by_feat`` and ``predict_by_feat`` methods in order. If one-stage is
+ used as RPN, the densehead needs to return both losses and predictions.
+ This predictions is used as the proposal of roihead.
+
+ .. code:: text
+
+ loss_and_predict(): forward() -> loss_by_feat() -> predict_by_feat()
+ """
+
+ def __init__(self, init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ # `_raw_positive_infos` will be used in `get_positive_infos`, which
+ # can get positive information.
+ self._raw_positive_infos = dict()
+
+ def init_weights(self) -> None:
+ """Initialize the weights."""
+ super().init_weights()
+ # avoid init_cfg overwrite the initialization of `conv_offset`
+ for m in self.modules():
+ # DeformConv2dPack, ModulatedDeformConv2dPack
+ if hasattr(m, 'conv_offset'):
+ constant_init(m.conv_offset, 0)
+
+ def get_positive_infos(self) -> InstanceList:
+ """Get positive information from sampling results.
+
+ Returns:
+ list[:obj:`InstanceData`]: Positive information of each image,
+ usually including positive bboxes, positive labels, positive
+ priors, etc.
+ """
+ if len(self._raw_positive_infos) == 0:
+ return None
+
+ sampling_results = self._raw_positive_infos.get(
+ 'sampling_results', None)
+ assert sampling_results is not None
+ positive_infos = []
+ for sampling_result in enumerate(sampling_results):
+ pos_info = InstanceData()
+ pos_info.bboxes = sampling_result.pos_gt_bboxes
+ pos_info.labels = sampling_result.pos_gt_labels
+ pos_info.priors = sampling_result.pos_priors
+ pos_info.pos_assigned_gt_inds = \
+ sampling_result.pos_assigned_gt_inds
+ pos_info.pos_inds = sampling_result.pos_inds
+ positive_infos.append(pos_info)
+ return positive_infos
+
+ def loss(self, x: Tuple[Tensor], batch_data_samples: SampleList) -> dict:
+ """Perform forward propagation and loss calculation of the detection
+ head on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ outs = self(x)
+
+ outputs = unpack_gt_instances(batch_data_samples)
+ (batch_gt_instances, batch_gt_instances_ignore,
+ batch_img_metas) = outputs
+
+ loss_inputs = outs + (batch_gt_instances, batch_img_metas,
+ batch_gt_instances_ignore)
+ losses = self.loss_by_feat(*loss_inputs)
+ return losses
+
+ @abstractmethod
+ def loss_by_feat(self, **kwargs) -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head."""
+ pass
+
+ def loss_and_predict(
+ self,
+ x: Tuple[Tensor],
+ batch_data_samples: SampleList,
+ proposal_cfg: Optional[ConfigDict] = None
+ ) -> Tuple[dict, InstanceList]:
+ """Perform forward propagation of the head, then calculate loss and
+ predictions from the features and data samples.
+
+ Args:
+ x (tuple[Tensor]): Features from FPN.
+ batch_data_samples (list[:obj:`DetDataSample`]): Each item contains
+ the meta information of each image and corresponding
+ annotations.
+ proposal_cfg (ConfigDict, optional): Test / postprocessing
+ configuration, if None, test_cfg would be used.
+ Defaults to None.
+
+ Returns:
+ tuple: the return value is a tuple contains:
+
+ - losses: (dict[str, Tensor]): A dictionary of loss components.
+ - predictions (list[:obj:`InstanceData`]): Detection
+ results of each image after the post process.
+ """
+ outputs = unpack_gt_instances(batch_data_samples)
+ (batch_gt_instances, batch_gt_instances_ignore,
+ batch_img_metas) = outputs
+
+ outs = self(x)
+
+ loss_inputs = outs + (batch_gt_instances, batch_img_metas,
+ batch_gt_instances_ignore)
+ losses = self.loss_by_feat(*loss_inputs)
+
+ predictions = self.predict_by_feat(
+ *outs, batch_img_metas=batch_img_metas, cfg=proposal_cfg)
+ return losses, predictions
+
+ def predict(self,
+ x: Tuple[Tensor],
+ batch_data_samples: SampleList,
+ rescale: bool = False) -> InstanceList:
+ """Perform forward propagation of the detection head and predict
+ detection results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Multi-level features from the
+ upstream network, each is a 4D-tensor.
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool, optional): Whether to rescale the results.
+ Defaults to False.
+
+ Returns:
+ list[obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ """
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+
+ outs = self(x)
+
+ predictions = self.predict_by_feat(
+ *outs, batch_img_metas=batch_img_metas, rescale=rescale)
+ return predictions
+
+ def predict_by_feat(self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ score_factors: Optional[List[Tensor]] = None,
+ batch_img_metas: Optional[List[dict]] = None,
+ cfg: Optional[ConfigDict] = None,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ bbox results.
+
+ Note: When score_factors is not None, the cls_scores are
+ usually multiplied by it then obtain the real score used in NMS,
+ such as CenterNess in FCOS, IoU branch in ATSS.
+
+ Args:
+ cls_scores (list[Tensor]): Classification scores for all
+ scale levels, each is a 4D-tensor, has shape
+ (batch_size, num_priors * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas for all
+ scale levels, each is a 4D-tensor, has shape
+ (batch_size, num_priors * 4, H, W).
+ score_factors (list[Tensor], optional): Score factor for
+ all scale level, each is a 4D-tensor, has shape
+ (batch_size, num_priors * 1, H, W). Defaults to None.
+ batch_img_metas (list[dict], Optional): Batch image meta info.
+ Defaults to None.
+ cfg (ConfigDict, optional): Test / postprocessing
+ configuration, if None, test_cfg would be used.
+ Defaults to None.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ list[:obj:`InstanceData`]: Object detection results of each image
+ after the post process. Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ assert len(cls_scores) == len(bbox_preds)
+
+ if score_factors is None:
+ # e.g. Retina, FreeAnchor, Foveabox, etc.
+ with_score_factors = False
+ else:
+ # e.g. FCOS, PAA, ATSS, AutoAssign, etc.
+ with_score_factors = True
+ assert len(cls_scores) == len(score_factors)
+
+ num_levels = len(cls_scores)
+
+ featmap_sizes = [cls_scores[i].shape[-2:] for i in range(num_levels)]
+ mlvl_priors = self.prior_generator.grid_priors(
+ featmap_sizes,
+ dtype=cls_scores[0].dtype,
+ device=cls_scores[0].device)
+
+ result_list = []
+
+ for img_id in range(len(batch_img_metas)):
+ img_meta = batch_img_metas[img_id]
+ cls_score_list = select_single_mlvl(
+ cls_scores, img_id, detach=True)
+ bbox_pred_list = select_single_mlvl(
+ bbox_preds, img_id, detach=True)
+ if with_score_factors:
+ score_factor_list = select_single_mlvl(
+ score_factors, img_id, detach=True)
+ else:
+ score_factor_list = [None for _ in range(num_levels)]
+
+ results = self._predict_by_feat_single(
+ cls_score_list=cls_score_list,
+ bbox_pred_list=bbox_pred_list,
+ score_factor_list=score_factor_list,
+ mlvl_priors=mlvl_priors,
+ img_meta=img_meta,
+ cfg=cfg,
+ rescale=rescale,
+ with_nms=with_nms)
+ result_list.append(results)
+ return result_list
+
+ def _predict_by_feat_single(self,
+ cls_score_list: List[Tensor],
+ bbox_pred_list: List[Tensor],
+ score_factor_list: List[Tensor],
+ mlvl_priors: List[Tensor],
+ img_meta: dict,
+ cfg: ConfigDict,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results.
+
+ Args:
+ cls_score_list (list[Tensor]): Box scores from all scale
+ levels of a single image, each item has shape
+ (num_priors * num_classes, H, W).
+ bbox_pred_list (list[Tensor]): Box energies / deltas from
+ all scale levels of a single image, each item has shape
+ (num_priors * 4, H, W).
+ score_factor_list (list[Tensor]): Score factor from all scale
+ levels of a single image, each item has shape
+ (num_priors * 1, H, W).
+ mlvl_priors (list[Tensor]): Each element in the list is
+ the priors of a single level in feature pyramid. In all
+ anchor-based methods, it has shape (num_priors, 4). In
+ all anchor-free methods, it has shape (num_priors, 2)
+ when `with_stride=True`, otherwise it still has shape
+ (num_priors, 4).
+ img_meta (dict): Image meta info.
+ cfg (mmengine.Config): Test / postprocessing configuration,
+ if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ if score_factor_list[0] is None:
+ # e.g. Retina, FreeAnchor, etc.
+ with_score_factors = False
+ else:
+ # e.g. FCOS, PAA, ATSS, etc.
+ with_score_factors = True
+
+ cfg = self.test_cfg if cfg is None else cfg
+ cfg = copy.deepcopy(cfg)
+ img_shape = img_meta['img_shape']
+ nms_pre = cfg.get('nms_pre', -1)
+
+ mlvl_bbox_preds = []
+ mlvl_valid_priors = []
+ mlvl_scores = []
+ mlvl_labels = []
+ if with_score_factors:
+ mlvl_score_factors = []
+ else:
+ mlvl_score_factors = None
+ for level_idx, (cls_score, bbox_pred, score_factor, priors) in \
+ enumerate(zip(cls_score_list, bbox_pred_list,
+ score_factor_list, mlvl_priors)):
+
+ assert cls_score.size()[-2:] == bbox_pred.size()[-2:]
+
+ dim = self.bbox_coder.encode_size
+ bbox_pred = bbox_pred.permute(1, 2, 0).reshape(-1, dim)
+ if with_score_factors:
+ score_factor = score_factor.permute(1, 2,
+ 0).reshape(-1).sigmoid()
+ cls_score = cls_score.permute(1, 2,
+ 0).reshape(-1, self.cls_out_channels)
+
+ # the `custom_cls_channels` parameter is derived from
+ # CrossEntropyCustomLoss and FocalCustomLoss, and is currently used
+ # in v3det.
+ if getattr(self.loss_cls, 'custom_cls_channels', False):
+ scores = self.loss_cls.get_activation(cls_score)
+ elif self.use_sigmoid_cls:
+ scores = cls_score.sigmoid()
+ else:
+ # remind that we set FG labels to [0, num_class-1]
+ # since mmdet v2.0
+ # BG cat_id: num_class
+ scores = cls_score.softmax(-1)[:, :-1]
+
+ # After https://github.com/open-mmlab/mmdetection/pull/6268/,
+ # this operation keeps fewer bboxes under the same `nms_pre`.
+ # There is no difference in performance for most models. If you
+ # find a slight drop in performance, you can set a larger
+ # `nms_pre` than before.
+ score_thr = cfg.get('score_thr', 0)
+
+ results = filter_scores_and_topk(
+ scores, score_thr, nms_pre,
+ dict(bbox_pred=bbox_pred, priors=priors))
+ scores, labels, keep_idxs, filtered_results = results
+
+ bbox_pred = filtered_results['bbox_pred']
+ priors = filtered_results['priors']
+
+ if with_score_factors:
+ score_factor = score_factor[keep_idxs]
+
+ mlvl_bbox_preds.append(bbox_pred)
+ mlvl_valid_priors.append(priors)
+ mlvl_scores.append(scores)
+ mlvl_labels.append(labels)
+
+ if with_score_factors:
+ mlvl_score_factors.append(score_factor)
+
+ bbox_pred = torch.cat(mlvl_bbox_preds)
+ priors = cat_boxes(mlvl_valid_priors)
+ bboxes = self.bbox_coder.decode(priors, bbox_pred, max_shape=img_shape)
+
+ results = InstanceData()
+ results.bboxes = bboxes
+ results.scores = torch.cat(mlvl_scores)
+ results.labels = torch.cat(mlvl_labels)
+ if with_score_factors:
+ results.score_factors = torch.cat(mlvl_score_factors)
+
+ return self._bbox_post_process(
+ results=results,
+ cfg=cfg,
+ rescale=rescale,
+ with_nms=with_nms,
+ img_meta=img_meta)
+
+ def _bbox_post_process(self,
+ results: InstanceData,
+ cfg: ConfigDict,
+ rescale: bool = False,
+ with_nms: bool = True,
+ img_meta: Optional[dict] = None) -> InstanceData:
+ """bbox post-processing method.
+
+ The boxes would be rescaled to the original image scale and do
+ the nms operation. Usually `with_nms` is False is used for aug test.
+
+ Args:
+ results (:obj:`InstaceData`): Detection instance results,
+ each item has shape (num_bboxes, ).
+ cfg (ConfigDict): Test / postprocessing configuration,
+ if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Default to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Default to True.
+ img_meta (dict, optional): Image meta info. Defaults to None.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ if rescale:
+ assert img_meta.get('scale_factor') is not None
+ scale_factor = [1 / s for s in img_meta['scale_factor']]
+ results.bboxes = scale_boxes(results.bboxes, scale_factor)
+
+ if hasattr(results, 'score_factors'):
+ # TODO: Add sqrt operation in order to be consistent with
+ # the paper.
+ score_factors = results.pop('score_factors')
+ results.scores = results.scores * score_factors
+
+ # filter small size bboxes
+ if cfg.get('min_bbox_size', -1) >= 0:
+ w, h = get_box_wh(results.bboxes)
+ valid_mask = (w > cfg.min_bbox_size) & (h > cfg.min_bbox_size)
+ if not valid_mask.all():
+ results = results[valid_mask]
+
+ # TODO: deal with `with_nms` and `nms_cfg=None` in test_cfg
+ if with_nms and results.bboxes.numel() > 0:
+ bboxes = get_box_tensor(results.bboxes)
+ det_bboxes, keep_idxs = batched_nms(bboxes, results.scores,
+ results.labels, cfg.nms)
+ results = results[keep_idxs]
+ # some nms would reweight the score, such as softnms
+ results.scores = det_bboxes[:, -1]
+ results = results[:cfg.max_per_img]
+
+ return results
+
+ def aug_test(self,
+ aug_batch_feats,
+ aug_batch_img_metas,
+ rescale=False,
+ with_ori_nms=False,
+ **kwargs):
+ """Test function with test time augmentation.
+
+ Args:
+ aug_batch_feats (list[tuple[Tensor]]): The outer list
+ indicates test-time augmentations and inner tuple
+ indicate the multi-level feats from
+ FPN, each Tensor should have a shape (B, C, H, W),
+ aug_batch_img_metas (list[list[dict]]): Meta information
+ of images under the different test-time augs
+ (multiscale, flip, etc.). The outer list indicate
+ the
+ rescale (bool, optional): Whether to rescale the results.
+ Defaults to False.
+ with_ori_nms (bool): Whether execute the nms in original head.
+ Defaults to False. It will be `True` when the head is
+ adopted as `rpn_head`.
+
+ Returns:
+ list(obj:`InstanceData`): Detection results of the
+ input images. Each item usually contains\
+ following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance,)
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances,).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ # TODO: remove this for detr and deformdetr
+ sig_of_get_results = signature(self.get_results)
+ get_results_args = [
+ p.name for p in sig_of_get_results.parameters.values()
+ ]
+ get_results_single_sig = signature(self._get_results_single)
+ get_results_single_sig_args = [
+ p.name for p in get_results_single_sig.parameters.values()
+ ]
+ assert ('with_nms' in get_results_args) and \
+ ('with_nms' in get_results_single_sig_args), \
+ f'{self.__class__.__name__}' \
+ 'does not support test-time augmentation '
+
+ num_imgs = len(aug_batch_img_metas[0])
+ aug_batch_results = []
+ for x, img_metas in zip(aug_batch_feats, aug_batch_img_metas):
+ outs = self.forward(x)
+ batch_instance_results = self.get_results(
+ *outs,
+ img_metas=img_metas,
+ cfg=self.test_cfg,
+ rescale=False,
+ with_nms=with_ori_nms,
+ **kwargs)
+ aug_batch_results.append(batch_instance_results)
+
+ # after merging, bboxes will be rescaled to the original image
+ batch_results = merge_aug_results(aug_batch_results,
+ aug_batch_img_metas)
+
+ final_results = []
+ for img_id in range(num_imgs):
+ results = batch_results[img_id]
+ det_bboxes, keep_idxs = batched_nms(results.bboxes, results.scores,
+ results.labels,
+ self.test_cfg.nms)
+ results = results[keep_idxs]
+ # some nms operation may reweight the score such as softnms
+ results.scores = det_bboxes[:, -1]
+ results = results[:self.test_cfg.max_per_img]
+ if rescale:
+ # all results have been mapped to the original scale
+ # in `merge_aug_results`, so just pass
+ pass
+ else:
+ # map to the first aug image scale
+ scale_factor = results.bboxes.new_tensor(
+ aug_batch_img_metas[0][img_id]['scale_factor'])
+ results.bboxes = \
+ results.bboxes * scale_factor
+
+ final_results.append(results)
+
+ return final_results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/base_mask_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/base_mask_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..7183d782829aa15bf12b9e2f7ade999c84d0593f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/base_mask_head.py
@@ -0,0 +1,128 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from abc import ABCMeta, abstractmethod
+from typing import List, Tuple, Union
+
+from mmengine.model import BaseModule
+from torch import Tensor
+
+from mmdet.structures import SampleList
+from mmdet.utils import InstanceList, OptInstanceList, OptMultiConfig
+from ..utils import unpack_gt_instances
+
+
+class BaseMaskHead(BaseModule, metaclass=ABCMeta):
+ """Base class for mask heads used in One-Stage Instance Segmentation."""
+
+ def __init__(self, init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+
+ @abstractmethod
+ def loss_by_feat(self, *args, **kwargs):
+ """Calculate the loss based on the features extracted by the mask
+ head."""
+ pass
+
+ @abstractmethod
+ def predict_by_feat(self, *args, **kwargs):
+ """Transform a batch of output features extracted from the head into
+ mask results."""
+ pass
+
+ def loss(self,
+ x: Union[List[Tensor], Tuple[Tensor]],
+ batch_data_samples: SampleList,
+ positive_infos: OptInstanceList = None,
+ **kwargs) -> dict:
+ """Perform forward propagation and loss calculation of the mask head on
+ the features of the upstream network.
+
+ Args:
+ x (list[Tensor] | tuple[Tensor]): Features from FPN.
+ Each has a shape (B, C, H, W).
+ batch_data_samples (list[:obj:`DetDataSample`]): Each item contains
+ the meta information of each image and corresponding
+ annotations.
+ positive_infos (list[:obj:`InstanceData`], optional): Information
+ of positive samples. Used when the label assignment is
+ done outside the MaskHead, e.g., BboxHead in
+ YOLACT or CondInst, etc. When the label assignment is done in
+ MaskHead, it would be None, like SOLO or SOLOv2. All values
+ in it should have shape (num_positive_samples, *).
+
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ if positive_infos is None:
+ outs = self(x)
+ else:
+ outs = self(x, positive_infos)
+
+ assert isinstance(outs, tuple), 'Forward results should be a tuple, ' \
+ 'even if only one item is returned'
+
+ outputs = unpack_gt_instances(batch_data_samples)
+ batch_gt_instances, batch_gt_instances_ignore, batch_img_metas \
+ = outputs
+ for gt_instances, img_metas in zip(batch_gt_instances,
+ batch_img_metas):
+ img_shape = img_metas['batch_input_shape']
+ gt_masks = gt_instances.masks.pad(img_shape)
+ gt_instances.masks = gt_masks
+
+ losses = self.loss_by_feat(
+ *outs,
+ batch_gt_instances=batch_gt_instances,
+ batch_img_metas=batch_img_metas,
+ positive_infos=positive_infos,
+ batch_gt_instances_ignore=batch_gt_instances_ignore,
+ **kwargs)
+ return losses
+
+ def predict(self,
+ x: Tuple[Tensor],
+ batch_data_samples: SampleList,
+ rescale: bool = False,
+ results_list: OptInstanceList = None,
+ **kwargs) -> InstanceList:
+ """Test function without test-time augmentation.
+
+ Args:
+ x (tuple[Tensor]): Multi-level features from the
+ upstream network, each is a 4D-tensor.
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool, optional): Whether to rescale the results.
+ Defaults to False.
+ results_list (list[obj:`InstanceData`], optional): Detection
+ results of each image after the post process. Only exist
+ if there is a `bbox_head`, like `YOLACT`, `CondInst`, etc.
+
+ Returns:
+ list[obj:`InstanceData`]: Instance segmentation
+ results of each image after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance,)
+ - labels (Tensor): Has a shape (num_instances,).
+ - masks (Tensor): Processed mask results, has a
+ shape (num_instances, h, w).
+ """
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+ if results_list is None:
+ outs = self(x)
+ else:
+ outs = self(x, results_list)
+
+ results_list = self.predict_by_feat(
+ *outs,
+ batch_img_metas=batch_img_metas,
+ rescale=rescale,
+ results_list=results_list,
+ **kwargs)
+
+ return results_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/boxinst_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/boxinst_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..7d6e8f7777a852cad89b709e59af2d8e12b343a6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/boxinst_head.py
@@ -0,0 +1,252 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List
+
+import torch
+import torch.nn.functional as F
+from mmengine import MessageHub
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import InstanceList
+from ..utils.misc import unfold_wo_center
+from .condinst_head import CondInstBboxHead, CondInstMaskHead
+
+
+@MODELS.register_module()
+class BoxInstBboxHead(CondInstBboxHead):
+ """BoxInst box head used in https://arxiv.org/abs/2012.02310."""
+
+ def __init__(self, *args, **kwargs) -> None:
+ super().__init__(*args, **kwargs)
+
+
+@MODELS.register_module()
+class BoxInstMaskHead(CondInstMaskHead):
+ """BoxInst mask head used in https://arxiv.org/abs/2012.02310.
+
+ This head outputs the mask for BoxInst.
+
+ Args:
+ pairwise_size (dict): The size of neighborhood for each pixel.
+ Defaults to 3.
+ pairwise_dilation (int): The dilation of neighborhood for each pixel.
+ Defaults to 2.
+ warmup_iters (int): Warmup iterations for pair-wise loss.
+ Defaults to 10000.
+ """
+
+ def __init__(self,
+ *arg,
+ pairwise_size: int = 3,
+ pairwise_dilation: int = 2,
+ warmup_iters: int = 10000,
+ **kwargs) -> None:
+ self.pairwise_size = pairwise_size
+ self.pairwise_dilation = pairwise_dilation
+ self.warmup_iters = warmup_iters
+ super().__init__(*arg, **kwargs)
+
+ def get_pairwise_affinity(self, mask_logits: Tensor) -> Tensor:
+ """Compute the pairwise affinity for each pixel."""
+ log_fg_prob = F.logsigmoid(mask_logits).unsqueeze(1)
+ log_bg_prob = F.logsigmoid(-mask_logits).unsqueeze(1)
+
+ log_fg_prob_unfold = unfold_wo_center(
+ log_fg_prob,
+ kernel_size=self.pairwise_size,
+ dilation=self.pairwise_dilation)
+ log_bg_prob_unfold = unfold_wo_center(
+ log_bg_prob,
+ kernel_size=self.pairwise_size,
+ dilation=self.pairwise_dilation)
+
+ # the probability of making the same prediction:
+ # p_i * p_j + (1 - p_i) * (1 - p_j)
+ # we compute the the probability in log space
+ # to avoid numerical instability
+ log_same_fg_prob = log_fg_prob[:, :, None] + log_fg_prob_unfold
+ log_same_bg_prob = log_bg_prob[:, :, None] + log_bg_prob_unfold
+
+ # TODO: Figure out the difference between it and directly sum
+ max_ = torch.max(log_same_fg_prob, log_same_bg_prob)
+ log_same_prob = torch.log(
+ torch.exp(log_same_fg_prob - max_) +
+ torch.exp(log_same_bg_prob - max_)) + max_
+
+ return -log_same_prob[:, 0]
+
+ def loss_by_feat(self, mask_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict], positive_infos: InstanceList,
+ **kwargs) -> dict:
+ """Calculate the loss based on the features extracted by the mask head.
+
+ Args:
+ mask_preds (list[Tensor]): List of predicted masks, each has
+ shape (num_classes, H, W).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``masks``,
+ and ``labels`` attributes.
+ batch_img_metas (list[dict]): Meta information of multiple images.
+ positive_infos (List[:obj:``InstanceData``]): Information of
+ positive samples of each image that are assigned in detection
+ head.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ assert positive_infos is not None, \
+ 'positive_infos should not be None in `BoxInstMaskHead`'
+ losses = dict()
+
+ loss_mask_project = 0.
+ loss_mask_pairwise = 0.
+ num_imgs = len(mask_preds)
+ total_pos = 0.
+ avg_fatcor = 0.
+
+ for idx in range(num_imgs):
+ (mask_pred, pos_mask_targets, pos_pairwise_masks, num_pos) = \
+ self._get_targets_single(
+ mask_preds[idx], batch_gt_instances[idx],
+ positive_infos[idx])
+ # mask loss
+ total_pos += num_pos
+ if num_pos == 0 or pos_mask_targets is None:
+ loss_project = mask_pred.new_zeros(1).mean()
+ loss_pairwise = mask_pred.new_zeros(1).mean()
+ avg_fatcor += 0.
+ else:
+ # compute the project term
+ loss_project_x = self.loss_mask(
+ mask_pred.max(dim=1, keepdim=True)[0],
+ pos_mask_targets.max(dim=1, keepdim=True)[0],
+ reduction_override='none').sum()
+ loss_project_y = self.loss_mask(
+ mask_pred.max(dim=2, keepdim=True)[0],
+ pos_mask_targets.max(dim=2, keepdim=True)[0],
+ reduction_override='none').sum()
+ loss_project = loss_project_x + loss_project_y
+ # compute the pairwise term
+ pairwise_affinity = self.get_pairwise_affinity(mask_pred)
+ avg_fatcor += pos_pairwise_masks.sum().clamp(min=1.0)
+ loss_pairwise = (pairwise_affinity * pos_pairwise_masks).sum()
+
+ loss_mask_project += loss_project
+ loss_mask_pairwise += loss_pairwise
+
+ if total_pos == 0:
+ total_pos += 1 # avoid nan
+ if avg_fatcor == 0:
+ avg_fatcor += 1 # avoid nan
+ loss_mask_project = loss_mask_project / total_pos
+ loss_mask_pairwise = loss_mask_pairwise / avg_fatcor
+ message_hub = MessageHub.get_current_instance()
+ iter = message_hub.get_info('iter')
+ warmup_factor = min(iter / float(self.warmup_iters), 1.0)
+ loss_mask_pairwise *= warmup_factor
+
+ losses.update(
+ loss_mask_project=loss_mask_project,
+ loss_mask_pairwise=loss_mask_pairwise)
+ return losses
+
+ def _get_targets_single(self, mask_preds: Tensor,
+ gt_instances: InstanceData,
+ positive_info: InstanceData):
+ """Compute targets for predictions of single image.
+
+ Args:
+ mask_preds (Tensor): Predicted prototypes with shape
+ (num_classes, H, W).
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes``, ``labels``,
+ and ``masks`` attributes.
+ positive_info (:obj:`InstanceData`): Information of positive
+ samples that are assigned in detection head. It usually
+ contains following keys.
+
+ - pos_assigned_gt_inds (Tensor): Assigner GT indexes of
+ positive proposals, has shape (num_pos, )
+ - pos_inds (Tensor): Positive index of image, has
+ shape (num_pos, ).
+ - param_pred (Tensor): Positive param preditions
+ with shape (num_pos, num_params).
+
+ Returns:
+ tuple: Usually returns a tuple containing learning targets.
+
+ - mask_preds (Tensor): Positive predicted mask with shape
+ (num_pos, mask_h, mask_w).
+ - pos_mask_targets (Tensor): Positive mask targets with shape
+ (num_pos, mask_h, mask_w).
+ - pos_pairwise_masks (Tensor): Positive pairwise masks with
+ shape: (num_pos, num_neighborhood, mask_h, mask_w).
+ - num_pos (int): Positive numbers.
+ """
+ gt_bboxes = gt_instances.bboxes
+ device = gt_bboxes.device
+ # Note that gt_masks are generated by full box
+ # from BoxInstDataPreprocessor
+ gt_masks = gt_instances.masks.to_tensor(
+ dtype=torch.bool, device=device).float()
+ # Note that pairwise_masks are generated by image color similarity
+ # from BoxInstDataPreprocessor
+ pairwise_masks = gt_instances.pairwise_masks
+ pairwise_masks = pairwise_masks.to(device=device)
+
+ # process with mask targets
+ pos_assigned_gt_inds = positive_info.get('pos_assigned_gt_inds')
+ scores = positive_info.get('scores')
+ centernesses = positive_info.get('centernesses')
+ num_pos = pos_assigned_gt_inds.size(0)
+
+ if gt_masks.size(0) == 0 or num_pos == 0:
+ return mask_preds, None, None, 0
+ # Since we're producing (near) full image masks,
+ # it'd take too much vram to backprop on every single mask.
+ # Thus we select only a subset.
+ if (self.max_masks_to_train != -1) and \
+ (num_pos > self.max_masks_to_train):
+ perm = torch.randperm(num_pos)
+ select = perm[:self.max_masks_to_train]
+ mask_preds = mask_preds[select]
+ pos_assigned_gt_inds = pos_assigned_gt_inds[select]
+ num_pos = self.max_masks_to_train
+ elif self.topk_masks_per_img != -1:
+ unique_gt_inds = pos_assigned_gt_inds.unique()
+ num_inst_per_gt = max(
+ int(self.topk_masks_per_img / len(unique_gt_inds)), 1)
+
+ keep_mask_preds = []
+ keep_pos_assigned_gt_inds = []
+ for gt_ind in unique_gt_inds:
+ per_inst_pos_inds = (pos_assigned_gt_inds == gt_ind)
+ mask_preds_per_inst = mask_preds[per_inst_pos_inds]
+ gt_inds_per_inst = pos_assigned_gt_inds[per_inst_pos_inds]
+ if sum(per_inst_pos_inds) > num_inst_per_gt:
+ per_inst_scores = scores[per_inst_pos_inds].sigmoid().max(
+ dim=1)[0]
+ per_inst_centerness = centernesses[
+ per_inst_pos_inds].sigmoid().reshape(-1, )
+ select = (per_inst_scores * per_inst_centerness).topk(
+ k=num_inst_per_gt, dim=0)[1]
+ mask_preds_per_inst = mask_preds_per_inst[select]
+ gt_inds_per_inst = gt_inds_per_inst[select]
+ keep_mask_preds.append(mask_preds_per_inst)
+ keep_pos_assigned_gt_inds.append(gt_inds_per_inst)
+ mask_preds = torch.cat(keep_mask_preds)
+ pos_assigned_gt_inds = torch.cat(keep_pos_assigned_gt_inds)
+ num_pos = pos_assigned_gt_inds.size(0)
+
+ # Follow the origin implement
+ start = int(self.mask_out_stride // 2)
+ gt_masks = gt_masks[:, start::self.mask_out_stride,
+ start::self.mask_out_stride]
+ gt_masks = gt_masks.gt(0.5).float()
+ pos_mask_targets = gt_masks[pos_assigned_gt_inds]
+ pos_pairwise_masks = pairwise_masks[pos_assigned_gt_inds]
+ pos_pairwise_masks = pos_pairwise_masks * pos_mask_targets.unsqueeze(1)
+
+ return (mask_preds, pos_mask_targets, pos_pairwise_masks, num_pos)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/cascade_rpn_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/cascade_rpn_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..a8686cc2c9118094df34a04fdeabd87daa636707
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/cascade_rpn_head.py
@@ -0,0 +1,1110 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from __future__ import division
+import copy
+from typing import Dict, List, Optional, Tuple, Union
+
+import torch
+import torch.nn as nn
+from mmcv.ops import DeformConv2d
+from mmengine.config import ConfigDict
+from mmengine.model import BaseModule, ModuleList
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures import SampleList
+from mmdet.utils import (ConfigType, InstanceList, MultiConfig,
+ OptInstanceList, OptMultiConfig)
+from ..task_modules.assigners import RegionAssigner
+from ..task_modules.samplers import PseudoSampler
+from ..utils import (images_to_levels, multi_apply, select_single_mlvl,
+ unpack_gt_instances)
+from .base_dense_head import BaseDenseHead
+from .rpn_head import RPNHead
+
+
+class AdaptiveConv(BaseModule):
+ """AdaptiveConv used to adapt the sampling location with the anchors.
+
+ Args:
+ in_channels (int): Number of channels in the input image.
+ out_channels (int): Number of channels produced by the convolution.
+ kernel_size (int or tuple[int]): Size of the conv kernel.
+ Defaults to 3.
+ stride (int or tuple[int]): Stride of the convolution. Defaults to 1.
+ padding (int or tuple[int]): Zero-padding added to both sides of
+ the input. Defaults to 1.
+ dilation (int or tuple[int]): Spacing between kernel elements.
+ Defaults to 3.
+ groups (int): Number of blocked connections from input channels to
+ output channels. Defaults to 1.
+ bias (bool): If set True, adds a learnable bias to the output.
+ Defaults to False.
+ adapt_type (str): Type of adaptive conv, can be either ``offset``
+ (arbitrary anchors) or 'dilation' (uniform anchor).
+ Defaults to 'dilation'.
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or \
+ list[dict]): Initialization config dict.
+ """
+
+ def __init__(
+ self,
+ in_channels: int,
+ out_channels: int,
+ kernel_size: Union[int, Tuple[int]] = 3,
+ stride: Union[int, Tuple[int]] = 1,
+ padding: Union[int, Tuple[int]] = 1,
+ dilation: Union[int, Tuple[int]] = 3,
+ groups: int = 1,
+ bias: bool = False,
+ adapt_type: str = 'dilation',
+ init_cfg: MultiConfig = dict(
+ type='Normal', std=0.01, override=dict(name='conv'))
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ assert adapt_type in ['offset', 'dilation']
+ self.adapt_type = adapt_type
+
+ assert kernel_size == 3, 'Adaptive conv only supports kernels 3'
+ if self.adapt_type == 'offset':
+ assert stride == 1 and padding == 1 and groups == 1, \
+ 'Adaptive conv offset mode only supports padding: {1}, ' \
+ f'stride: {1}, groups: {1}'
+ self.conv = DeformConv2d(
+ in_channels,
+ out_channels,
+ kernel_size,
+ padding=padding,
+ stride=stride,
+ groups=groups,
+ bias=bias)
+ else:
+ self.conv = nn.Conv2d(
+ in_channels,
+ out_channels,
+ kernel_size,
+ padding=dilation,
+ dilation=dilation)
+
+ def forward(self, x: Tensor, offset: Tensor) -> Tensor:
+ """Forward function."""
+ if self.adapt_type == 'offset':
+ N, _, H, W = x.shape
+ assert offset is not None
+ assert H * W == offset.shape[1]
+ # reshape [N, NA, 18] to (N, 18, H, W)
+ offset = offset.permute(0, 2, 1).reshape(N, -1, H, W)
+ offset = offset.contiguous()
+ x = self.conv(x, offset)
+ else:
+ assert offset is None
+ x = self.conv(x)
+ return x
+
+
+@MODELS.register_module()
+class StageCascadeRPNHead(RPNHead):
+ """Stage of CascadeRPNHead.
+
+ Args:
+ in_channels (int): Number of channels in the input feature map.
+ anchor_generator (:obj:`ConfigDict` or dict): anchor generator config.
+ adapt_cfg (:obj:`ConfigDict` or dict): adaptation config.
+ bridged_feature (bool): whether update rpn feature. Defaults to False.
+ with_cls (bool): whether use classification branch. Defaults to True.
+ init_cfg :obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or
+ list[dict], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ in_channels: int,
+ anchor_generator: ConfigType = dict(
+ type='AnchorGenerator',
+ scales=[8],
+ ratios=[1.0],
+ strides=[4, 8, 16, 32, 64]),
+ adapt_cfg: ConfigType = dict(type='dilation', dilation=3),
+ bridged_feature: bool = False,
+ with_cls: bool = True,
+ init_cfg: OptMultiConfig = None,
+ **kwargs) -> None:
+ self.with_cls = with_cls
+ self.anchor_strides = anchor_generator['strides']
+ self.anchor_scales = anchor_generator['scales']
+ self.bridged_feature = bridged_feature
+ self.adapt_cfg = adapt_cfg
+ super().__init__(
+ in_channels=in_channels,
+ anchor_generator=anchor_generator,
+ init_cfg=init_cfg,
+ **kwargs)
+
+ # override sampling and sampler
+ if self.train_cfg:
+ self.assigner = TASK_UTILS.build(self.train_cfg['assigner'])
+ # use PseudoSampler when sampling is False
+ if self.train_cfg.get('sampler', None) is not None:
+ self.sampler = TASK_UTILS.build(
+ self.train_cfg['sampler'], default_args=dict(context=self))
+ else:
+ self.sampler = PseudoSampler(context=self)
+
+ if init_cfg is None:
+ self.init_cfg = dict(
+ type='Normal', std=0.01, override=[dict(name='rpn_reg')])
+ if self.with_cls:
+ self.init_cfg['override'].append(dict(name='rpn_cls'))
+
+ def _init_layers(self) -> None:
+ """Init layers of a CascadeRPN stage."""
+ adapt_cfg = copy.deepcopy(self.adapt_cfg)
+ adapt_cfg['adapt_type'] = adapt_cfg.pop('type')
+ self.rpn_conv = AdaptiveConv(self.in_channels, self.feat_channels,
+ **adapt_cfg)
+ if self.with_cls:
+ self.rpn_cls = nn.Conv2d(self.feat_channels,
+ self.num_anchors * self.cls_out_channels,
+ 1)
+ self.rpn_reg = nn.Conv2d(self.feat_channels, self.num_anchors * 4, 1)
+ self.relu = nn.ReLU(inplace=True)
+
+ def forward_single(self, x: Tensor, offset: Tensor) -> Tuple[Tensor]:
+ """Forward function of single scale."""
+ bridged_x = x
+ x = self.relu(self.rpn_conv(x, offset))
+ if self.bridged_feature:
+ bridged_x = x # update feature
+ cls_score = self.rpn_cls(x) if self.with_cls else None
+ bbox_pred = self.rpn_reg(x)
+ return bridged_x, cls_score, bbox_pred
+
+ def forward(
+ self,
+ feats: List[Tensor],
+ offset_list: Optional[List[Tensor]] = None) -> Tuple[List[Tensor]]:
+ """Forward function."""
+ if offset_list is None:
+ offset_list = [None for _ in range(len(feats))]
+ return multi_apply(self.forward_single, feats, offset_list)
+
+ def _region_targets_single(self, flat_anchors: Tensor, valid_flags: Tensor,
+ gt_instances: InstanceData, img_meta: dict,
+ gt_instances_ignore: InstanceData,
+ featmap_sizes: List[Tuple[int, int]],
+ num_level_anchors: List[int]) -> tuple:
+ """Get anchor targets based on region for single level.
+
+ Args:
+ flat_anchors (Tensor): Multi-level anchors of the image, which are
+ concatenated into a single tensor of shape (num_anchors, 4)
+ valid_flags (Tensor): Multi level valid flags of the image,
+ which are concatenated into a single tensor of
+ shape (num_anchors, ).
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes`` and ``labels``
+ attributes.
+ img_meta (dict): Meta information for current image.
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ featmap_sizes (list[Tuple[int, int]]): Feature map size each level.
+ num_level_anchors (list[int]): The number of anchors in each level.
+
+ Returns:
+ tuple:
+
+ - labels (Tensor): Labels of each level.
+ - label_weights (Tensor): Label weights of each level.
+ - bbox_targets (Tensor): BBox targets of each level.
+ - bbox_weights (Tensor): BBox weights of each level.
+ - pos_inds (Tensor): positive samples indexes.
+ - neg_inds (Tensor): negative samples indexes.
+ - sampling_result (:obj:`SamplingResult`): Sampling results.
+ """
+ pred_instances = InstanceData()
+ pred_instances.priors = flat_anchors
+ pred_instances.valid_flags = valid_flags
+
+ assign_result = self.assigner.assign(
+ pred_instances,
+ gt_instances,
+ img_meta,
+ featmap_sizes,
+ num_level_anchors,
+ self.anchor_scales[0],
+ self.anchor_strides,
+ gt_instances_ignore=gt_instances_ignore,
+ allowed_border=self.train_cfg['allowed_border'])
+ sampling_result = self.sampler.sample(assign_result, pred_instances,
+ gt_instances)
+
+ num_anchors = flat_anchors.shape[0]
+ bbox_targets = torch.zeros_like(flat_anchors)
+ bbox_weights = torch.zeros_like(flat_anchors)
+ labels = flat_anchors.new_zeros(num_anchors, dtype=torch.long)
+ label_weights = flat_anchors.new_zeros(num_anchors, dtype=torch.float)
+
+ pos_inds = sampling_result.pos_inds
+ neg_inds = sampling_result.neg_inds
+ if len(pos_inds) > 0:
+ if not self.reg_decoded_bbox:
+ pos_bbox_targets = self.bbox_coder.encode(
+ sampling_result.pos_bboxes, sampling_result.pos_gt_bboxes)
+ else:
+ pos_bbox_targets = sampling_result.pos_gt_bboxes
+ bbox_targets[pos_inds, :] = pos_bbox_targets
+ bbox_weights[pos_inds, :] = 1.0
+ labels[pos_inds] = sampling_result.pos_gt_labels
+ if self.train_cfg['pos_weight'] <= 0:
+ label_weights[pos_inds] = 1.0
+ else:
+ label_weights[pos_inds] = self.train_cfg['pos_weight']
+ if len(neg_inds) > 0:
+ label_weights[neg_inds] = 1.0
+
+ return (labels, label_weights, bbox_targets, bbox_weights, pos_inds,
+ neg_inds, sampling_result)
+
+ def region_targets(
+ self,
+ anchor_list: List[List[Tensor]],
+ valid_flag_list: List[List[Tensor]],
+ featmap_sizes: List[Tuple[int, int]],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None,
+ return_sampling_results: bool = False,
+ ) -> tuple:
+ """Compute regression and classification targets for anchors when using
+ RegionAssigner.
+
+ Args:
+ anchor_list (list[list[Tensor]]): Multi level anchors of each
+ image.
+ valid_flag_list (list[list[Tensor]]): Multi level valid flags of
+ each image.
+ featmap_sizes (list[Tuple[int, int]]): Feature map size each level.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ tuple:
+
+ - labels_list (list[Tensor]): Labels of each level.
+ - label_weights_list (list[Tensor]): Label weights of each
+ level.
+ - bbox_targets_list (list[Tensor]): BBox targets of each level.
+ - bbox_weights_list (list[Tensor]): BBox weights of each level.
+ - avg_factor (int): Average factor that is used to average
+ the loss. When using sampling method, avg_factor is usually
+ the sum of positive and negative priors. When using
+ ``PseudoSampler``, ``avg_factor`` is usually equal to the
+ number of positive priors.
+ """
+ num_imgs = len(batch_img_metas)
+ assert len(anchor_list) == len(valid_flag_list) == num_imgs
+
+ if batch_gt_instances_ignore is None:
+ batch_gt_instances_ignore = [None] * num_imgs
+
+ # anchor number of multi levels
+ num_level_anchors = [anchors.size(0) for anchors in anchor_list[0]]
+ # concat all level anchors to a single tensor
+ concat_anchor_list = []
+ concat_valid_flag_list = []
+ for i in range(num_imgs):
+ assert len(anchor_list[i]) == len(valid_flag_list[i])
+ concat_anchor_list.append(torch.cat(anchor_list[i]))
+ concat_valid_flag_list.append(torch.cat(valid_flag_list[i]))
+
+ # compute targets for each image
+ (all_labels, all_label_weights, all_bbox_targets, all_bbox_weights,
+ pos_inds_list, neg_inds_list, sampling_results_list) = multi_apply(
+ self._region_targets_single,
+ concat_anchor_list,
+ concat_valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore,
+ featmap_sizes=featmap_sizes,
+ num_level_anchors=num_level_anchors)
+ # no valid anchors
+ if any([labels is None for labels in all_labels]):
+ return None
+ # sampled anchors of all images
+ avg_factor = sum(
+ [results.avg_factor for results in sampling_results_list])
+ # split targets to a list w.r.t. multiple levels
+ labels_list = images_to_levels(all_labels, num_level_anchors)
+ label_weights_list = images_to_levels(all_label_weights,
+ num_level_anchors)
+ bbox_targets_list = images_to_levels(all_bbox_targets,
+ num_level_anchors)
+ bbox_weights_list = images_to_levels(all_bbox_weights,
+ num_level_anchors)
+ res = (labels_list, label_weights_list, bbox_targets_list,
+ bbox_weights_list, avg_factor)
+ if return_sampling_results:
+ res = res + (sampling_results_list, )
+ return res
+
+ def get_targets(
+ self,
+ anchor_list: List[List[Tensor]],
+ valid_flag_list: List[List[Tensor]],
+ featmap_sizes: List[Tuple[int, int]],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None,
+ return_sampling_results: bool = False,
+ ) -> tuple:
+ """Compute regression and classification targets for anchors.
+
+ Args:
+ anchor_list (list[list[Tensor]]): Multi level anchors of each
+ image.
+ valid_flag_list (list[list[Tensor]]): Multi level valid flags of
+ each image.
+ featmap_sizes (list[Tuple[int, int]]): Feature map size each level.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ return_sampling_results (bool): Whether to return the sampling
+ results. Defaults to False.
+
+ Returns:
+ tuple:
+
+ - labels_list (list[Tensor]): Labels of each level.
+ - label_weights_list (list[Tensor]): Label weights of each
+ level.
+ - bbox_targets_list (list[Tensor]): BBox targets of each level.
+ - bbox_weights_list (list[Tensor]): BBox weights of each level.
+ - avg_factor (int): Average factor that is used to average
+ the loss. When using sampling method, avg_factor is usually
+ the sum of positive and negative priors. When using
+ ``PseudoSampler``, ``avg_factor`` is usually equal to the
+ number of positive priors.
+ """
+ if isinstance(self.assigner, RegionAssigner):
+ cls_reg_targets = self.region_targets(
+ anchor_list,
+ valid_flag_list,
+ featmap_sizes,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore,
+ return_sampling_results=return_sampling_results)
+ else:
+ cls_reg_targets = super().get_targets(
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore,
+ return_sampling_results=return_sampling_results)
+ return cls_reg_targets
+
+ def anchor_offset(self, anchor_list: List[List[Tensor]],
+ anchor_strides: List[int],
+ featmap_sizes: List[Tuple[int, int]]) -> List[Tensor]:
+ """ Get offset for deformable conv based on anchor shape
+ NOTE: currently support deformable kernel_size=3 and dilation=1
+
+ Args:
+ anchor_list (list[list[tensor])): [NI, NLVL, NA, 4] list of
+ multi-level anchors
+ anchor_strides (list[int]): anchor stride of each level
+
+ Returns:
+ list[tensor]: offset of DeformConv kernel with shapes of
+ [NLVL, NA, 2, 18].
+ """
+
+ def _shape_offset(anchors, stride, ks=3, dilation=1):
+ # currently support kernel_size=3 and dilation=1
+ assert ks == 3 and dilation == 1
+ pad = (ks - 1) // 2
+ idx = torch.arange(-pad, pad + 1, dtype=dtype, device=device)
+ yy, xx = torch.meshgrid(idx, idx) # return order matters
+ xx = xx.reshape(-1)
+ yy = yy.reshape(-1)
+ w = (anchors[:, 2] - anchors[:, 0]) / stride
+ h = (anchors[:, 3] - anchors[:, 1]) / stride
+ w = w / (ks - 1) - dilation
+ h = h / (ks - 1) - dilation
+ offset_x = w[:, None] * xx # (NA, ks**2)
+ offset_y = h[:, None] * yy # (NA, ks**2)
+ return offset_x, offset_y
+
+ def _ctr_offset(anchors, stride, featmap_size):
+ feat_h, feat_w = featmap_size
+ assert len(anchors) == feat_h * feat_w
+
+ x = (anchors[:, 0] + anchors[:, 2]) * 0.5
+ y = (anchors[:, 1] + anchors[:, 3]) * 0.5
+ # compute centers on feature map
+ x = x / stride
+ y = y / stride
+ # compute predefine centers
+ xx = torch.arange(0, feat_w, device=anchors.device)
+ yy = torch.arange(0, feat_h, device=anchors.device)
+ yy, xx = torch.meshgrid(yy, xx)
+ xx = xx.reshape(-1).type_as(x)
+ yy = yy.reshape(-1).type_as(y)
+
+ offset_x = x - xx # (NA, )
+ offset_y = y - yy # (NA, )
+ return offset_x, offset_y
+
+ num_imgs = len(anchor_list)
+ num_lvls = len(anchor_list[0])
+ dtype = anchor_list[0][0].dtype
+ device = anchor_list[0][0].device
+ num_level_anchors = [anchors.size(0) for anchors in anchor_list[0]]
+
+ offset_list = []
+ for i in range(num_imgs):
+ mlvl_offset = []
+ for lvl in range(num_lvls):
+ c_offset_x, c_offset_y = _ctr_offset(anchor_list[i][lvl],
+ anchor_strides[lvl],
+ featmap_sizes[lvl])
+ s_offset_x, s_offset_y = _shape_offset(anchor_list[i][lvl],
+ anchor_strides[lvl])
+
+ # offset = ctr_offset + shape_offset
+ offset_x = s_offset_x + c_offset_x[:, None]
+ offset_y = s_offset_y + c_offset_y[:, None]
+
+ # offset order (y0, x0, y1, x2, .., y8, x8, y9, x9)
+ offset = torch.stack([offset_y, offset_x], dim=-1)
+ offset = offset.reshape(offset.size(0), -1) # [NA, 2*ks**2]
+ mlvl_offset.append(offset)
+ offset_list.append(torch.cat(mlvl_offset)) # [totalNA, 2*ks**2]
+ offset_list = images_to_levels(offset_list, num_level_anchors)
+ return offset_list
+
+ def loss_by_feat_single(self, cls_score: Tensor, bbox_pred: Tensor,
+ anchors: Tensor, labels: Tensor,
+ label_weights: Tensor, bbox_targets: Tensor,
+ bbox_weights: Tensor, avg_factor: int) -> tuple:
+ """Loss function on single scale."""
+ # classification loss
+ if self.with_cls:
+ labels = labels.reshape(-1)
+ label_weights = label_weights.reshape(-1)
+ cls_score = cls_score.permute(0, 2, 3,
+ 1).reshape(-1, self.cls_out_channels)
+ loss_cls = self.loss_cls(
+ cls_score, labels, label_weights, avg_factor=avg_factor)
+ # regression loss
+ bbox_targets = bbox_targets.reshape(-1, 4)
+ bbox_weights = bbox_weights.reshape(-1, 4)
+ bbox_pred = bbox_pred.permute(0, 2, 3, 1).reshape(-1, 4)
+ if self.reg_decoded_bbox:
+ # When the regression loss (e.g. `IouLoss`, `GIouLoss`)
+ # is applied directly on the decoded bounding boxes, it
+ # decodes the already encoded coordinates to absolute format.
+ anchors = anchors.reshape(-1, 4)
+ bbox_pred = self.bbox_coder.decode(anchors, bbox_pred)
+ loss_reg = self.loss_bbox(
+ bbox_pred, bbox_targets, bbox_weights, avg_factor=avg_factor)
+ if self.with_cls:
+ return loss_cls, loss_reg
+ return None, loss_reg
+
+ def loss_by_feat(
+ self,
+ anchor_list: List[List[Tensor]],
+ valid_flag_list: List[List[Tensor]],
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None
+ ) -> Dict[str, Tensor]:
+ """Compute losses of the head.
+
+ Args:
+ anchor_list (list[list[Tensor]]): Multi level anchors of each
+ image.
+ valid_flag_list (list[list[Tensor]]): Multi level valid flags of
+ each image. The outer list indicates images, and the inner list
+ corresponds to feature levels of the image. Each element of
+ the inner list is a tensor of shape (num_anchors, )
+ cls_scores (list[Tensor]): Box scores for each scale level
+ Has shape (N, num_anchors * num_classes, H, W)
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W)
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ featmap_sizes = [featmap.size()[-2:] for featmap in bbox_preds]
+ cls_reg_targets = self.get_targets(
+ anchor_list,
+ valid_flag_list,
+ featmap_sizes,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore,
+ return_sampling_results=True)
+ (labels_list, label_weights_list, bbox_targets_list, bbox_weights_list,
+ avg_factor, sampling_results_list) = cls_reg_targets
+ if not sampling_results_list[0].avg_factor_with_neg:
+ # 200 is hard-coded average factor,
+ # which follows guided anchoring.
+ avg_factor = sum([label.numel() for label in labels_list]) / 200.0
+
+ # change per image, per level anchor_list to per_level, per_image
+ mlvl_anchor_list = list(zip(*anchor_list))
+ # concat mlvl_anchor_list
+ mlvl_anchor_list = [
+ torch.cat(anchors, dim=0) for anchors in mlvl_anchor_list
+ ]
+
+ losses = multi_apply(
+ self.loss_by_feat_single,
+ cls_scores,
+ bbox_preds,
+ mlvl_anchor_list,
+ labels_list,
+ label_weights_list,
+ bbox_targets_list,
+ bbox_weights_list,
+ avg_factor=avg_factor)
+ if self.with_cls:
+ return dict(loss_rpn_cls=losses[0], loss_rpn_reg=losses[1])
+ return dict(loss_rpn_reg=losses[1])
+
+ def predict_by_feat(self,
+ anchor_list: List[List[Tensor]],
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_img_metas: List[dict],
+ cfg: Optional[ConfigDict] = None,
+ rescale: bool = False) -> InstanceList:
+ """Get proposal predict. Overriding to enable input ``anchor_list``
+ from outside.
+
+ Args:
+ anchor_list (list[list[Tensor]]): Multi level anchors of each
+ image.
+ cls_scores (list[Tensor]): Classification scores for all
+ scale levels, each is a 4D-tensor, has shape
+ (batch_size, num_priors * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas for all
+ scale levels, each is a 4D-tensor, has shape
+ (batch_size, num_priors * 4, H, W).
+ batch_img_metas (list[dict], Optional): Image meta info.
+ cfg (:obj:`ConfigDict`, optional): Test / postprocessing
+ configuration, if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Object detection results of each image
+ after the post process. Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ assert len(cls_scores) == len(bbox_preds)
+
+ result_list = []
+ for img_id in range(len(batch_img_metas)):
+ cls_score_list = select_single_mlvl(cls_scores, img_id)
+ bbox_pred_list = select_single_mlvl(bbox_preds, img_id)
+ proposals = self._predict_by_feat_single(
+ cls_scores=cls_score_list,
+ bbox_preds=bbox_pred_list,
+ mlvl_anchors=anchor_list[img_id],
+ img_meta=batch_img_metas[img_id],
+ cfg=cfg,
+ rescale=rescale)
+ result_list.append(proposals)
+ return result_list
+
+ def _predict_by_feat_single(self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ mlvl_anchors: List[Tensor],
+ img_meta: dict,
+ cfg: ConfigDict,
+ rescale: bool = False) -> InstanceData:
+ """Transform outputs of a single image into bbox predictions.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores from all scale
+ levels of a single image, each item has shape
+ (num_anchors * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas from
+ all scale levels of a single image, each item has
+ shape (num_anchors * 4, H, W).
+ mlvl_anchors (list[Tensor]): Box reference from all scale
+ levels of a single image, each item has shape
+ (num_total_anchors, 4).
+ img_shape (tuple[int]): Shape of the input image,
+ (height, width, 3).
+ scale_factor (ndarray): Scale factor of the image arange as
+ (w_scale, h_scale, w_scale, h_scale).
+ cfg (:obj:`ConfigDict`): Test / postprocessing configuration,
+ if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ cfg = self.test_cfg if cfg is None else cfg
+ cfg = copy.deepcopy(cfg)
+ # bboxes from different level should be independent during NMS,
+ # level_ids are used as labels for batched NMS to separate them
+ level_ids = []
+ mlvl_scores = []
+ mlvl_bbox_preds = []
+ mlvl_valid_anchors = []
+ nms_pre = cfg.get('nms_pre', -1)
+ for idx in range(len(cls_scores)):
+ rpn_cls_score = cls_scores[idx]
+ rpn_bbox_pred = bbox_preds[idx]
+ assert rpn_cls_score.size()[-2:] == rpn_bbox_pred.size()[-2:]
+ rpn_cls_score = rpn_cls_score.permute(1, 2, 0)
+ if self.use_sigmoid_cls:
+ rpn_cls_score = rpn_cls_score.reshape(-1)
+ scores = rpn_cls_score.sigmoid()
+ else:
+ rpn_cls_score = rpn_cls_score.reshape(-1, 2)
+ # We set FG labels to [0, num_class-1] and BG label to
+ # num_class in RPN head since mmdet v2.5, which is unified to
+ # be consistent with other head since mmdet v2.0. In mmdet v2.0
+ # to v2.4 we keep BG label as 0 and FG label as 1 in rpn head.
+ scores = rpn_cls_score.softmax(dim=1)[:, 0]
+ rpn_bbox_pred = rpn_bbox_pred.permute(1, 2, 0).reshape(-1, 4)
+ anchors = mlvl_anchors[idx]
+
+ if 0 < nms_pre < scores.shape[0]:
+ # sort is faster than topk
+ # _, topk_inds = scores.topk(cfg.nms_pre)
+ ranked_scores, rank_inds = scores.sort(descending=True)
+ topk_inds = rank_inds[:nms_pre]
+ scores = ranked_scores[:nms_pre]
+ rpn_bbox_pred = rpn_bbox_pred[topk_inds, :]
+ anchors = anchors[topk_inds, :]
+ mlvl_scores.append(scores)
+ mlvl_bbox_preds.append(rpn_bbox_pred)
+ mlvl_valid_anchors.append(anchors)
+ level_ids.append(
+ scores.new_full((scores.size(0), ), idx, dtype=torch.long))
+
+ anchors = torch.cat(mlvl_valid_anchors)
+ rpn_bbox_pred = torch.cat(mlvl_bbox_preds)
+ bboxes = self.bbox_coder.decode(
+ anchors, rpn_bbox_pred, max_shape=img_meta['img_shape'])
+
+ proposals = InstanceData()
+ proposals.bboxes = bboxes
+ proposals.scores = torch.cat(mlvl_scores)
+ proposals.level_ids = torch.cat(level_ids)
+
+ return self._bbox_post_process(
+ results=proposals, cfg=cfg, rescale=rescale, img_meta=img_meta)
+
+ def refine_bboxes(self, anchor_list: List[List[Tensor]],
+ bbox_preds: List[Tensor],
+ img_metas: List[dict]) -> List[List[Tensor]]:
+ """Refine bboxes through stages."""
+ num_levels = len(bbox_preds)
+ new_anchor_list = []
+ for img_id in range(len(img_metas)):
+ mlvl_anchors = []
+ for i in range(num_levels):
+ bbox_pred = bbox_preds[i][img_id].detach()
+ bbox_pred = bbox_pred.permute(1, 2, 0).reshape(-1, 4)
+ img_shape = img_metas[img_id]['img_shape']
+ bboxes = self.bbox_coder.decode(anchor_list[img_id][i],
+ bbox_pred, img_shape)
+ mlvl_anchors.append(bboxes)
+ new_anchor_list.append(mlvl_anchors)
+ return new_anchor_list
+
+ def loss(self, x: Tuple[Tensor], batch_data_samples: SampleList) -> dict:
+ """Perform forward propagation and loss calculation of the detection
+ head on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ outputs = unpack_gt_instances(batch_data_samples)
+ batch_gt_instances, _, batch_img_metas = outputs
+
+ featmap_sizes = [featmap.size()[-2:] for featmap in x]
+ device = x[0].device
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+
+ if self.adapt_cfg['type'] == 'offset':
+ offset_list = self.anchor_offset(anchor_list, self.anchor_strides,
+ featmap_sizes)
+ else:
+ offset_list = None
+
+ x, cls_score, bbox_pred = self(x, offset_list)
+ rpn_loss_inputs = (anchor_list, valid_flag_list, cls_score, bbox_pred,
+ batch_gt_instances, batch_img_metas)
+ losses = self.loss_by_feat(*rpn_loss_inputs)
+
+ return losses
+
+ def loss_and_predict(
+ self,
+ x: Tuple[Tensor],
+ batch_data_samples: SampleList,
+ proposal_cfg: Optional[ConfigDict] = None,
+ ) -> Tuple[dict, InstanceList]:
+ """Perform forward propagation of the head, then calculate loss and
+ predictions from the features and data samples.
+
+ Args:
+ x (tuple[Tensor]): Features from FPN.
+ batch_data_samples (list[:obj:`DetDataSample`]): Each item contains
+ the meta information of each image and corresponding
+ annotations.
+ proposal_cfg (:obj`ConfigDict`, optional): Test / postprocessing
+ configuration, if None, test_cfg would be used.
+ Defaults to None.
+
+ Returns:
+ tuple: the return value is a tuple contains:
+
+ - losses: (dict[str, Tensor]): A dictionary of loss components.
+ - predictions (list[:obj:`InstanceData`]): Detection
+ results of each image after the post process.
+ """
+ outputs = unpack_gt_instances(batch_data_samples)
+ batch_gt_instances, _, batch_img_metas = outputs
+
+ featmap_sizes = [featmap.size()[-2:] for featmap in x]
+ device = x[0].device
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+
+ if self.adapt_cfg['type'] == 'offset':
+ offset_list = self.anchor_offset(anchor_list, self.anchor_strides,
+ featmap_sizes)
+ else:
+ offset_list = None
+
+ x, cls_score, bbox_pred = self(x, offset_list)
+ rpn_loss_inputs = (anchor_list, valid_flag_list, cls_score, bbox_pred,
+ batch_gt_instances, batch_img_metas)
+ losses = self.loss_by_feat(*rpn_loss_inputs)
+
+ predictions = self.predict_by_feat(
+ anchor_list,
+ cls_score,
+ bbox_pred,
+ batch_img_metas=batch_img_metas,
+ cfg=proposal_cfg)
+ return losses, predictions
+
+ def predict(self,
+ x: Tuple[Tensor],
+ batch_data_samples: SampleList,
+ rescale: bool = False) -> InstanceList:
+ """Perform forward propagation of the detection head and predict
+ detection results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Multi-level features from the
+ upstream network, each is a 4D-tensor.
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool, optional): Whether to rescale the results.
+ Defaults to False.
+
+ Returns:
+ list[obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ """
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+
+ featmap_sizes = [featmap.size()[-2:] for featmap in x]
+ device = x[0].device
+ anchor_list, _ = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+
+ if self.adapt_cfg['type'] == 'offset':
+ offset_list = self.anchor_offset(anchor_list, self.anchor_strides,
+ featmap_sizes)
+ else:
+ offset_list = None
+
+ x, cls_score, bbox_pred = self(x, offset_list)
+ predictions = self.stages[-1].predict_by_feat(
+ anchor_list,
+ cls_score,
+ bbox_pred,
+ batch_img_metas=batch_img_metas,
+ rescale=rescale)
+ return predictions
+
+
+@MODELS.register_module()
+class CascadeRPNHead(BaseDenseHead):
+ """The CascadeRPNHead will predict more accurate region proposals, which is
+ required for two-stage detectors (such as Fast/Faster R-CNN). CascadeRPN
+ consists of a sequence of RPNStage to progressively improve the accuracy of
+ the detected proposals.
+
+ More details can be found in ``https://arxiv.org/abs/1909.06720``.
+
+ Args:
+ num_stages (int): number of CascadeRPN stages.
+ stages (list[:obj:`ConfigDict` or dict]): list of configs to build
+ the stages.
+ train_cfg (list[:obj:`ConfigDict` or dict]): list of configs at
+ training time each stage.
+ test_cfg (:obj:`ConfigDict` or dict): config at testing time.
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or \
+ list[dict]): Initialization config dict.
+ """
+
+ def __init__(self,
+ num_classes: int,
+ num_stages: int,
+ stages: List[ConfigType],
+ train_cfg: List[ConfigType],
+ test_cfg: ConfigType,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ assert num_classes == 1, 'Only support num_classes == 1'
+ assert num_stages == len(stages)
+ self.num_stages = num_stages
+ # Be careful! Pretrained weights cannot be loaded when use
+ # nn.ModuleList
+ self.stages = ModuleList()
+ for i in range(len(stages)):
+ train_cfg_i = train_cfg[i] if train_cfg is not None else None
+ stages[i].update(train_cfg=train_cfg_i)
+ stages[i].update(test_cfg=test_cfg)
+ self.stages.append(MODELS.build(stages[i]))
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+
+ def loss_by_feat(self):
+ """loss_by_feat() is implemented in StageCascadeRPNHead."""
+ pass
+
+ def predict_by_feat(self):
+ """predict_by_feat() is implemented in StageCascadeRPNHead."""
+ pass
+
+ def loss(self, x: Tuple[Tensor], batch_data_samples: SampleList) -> dict:
+ """Perform forward propagation and loss calculation of the detection
+ head on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ outputs = unpack_gt_instances(batch_data_samples)
+ batch_gt_instances, _, batch_img_metas = outputs
+
+ featmap_sizes = [featmap.size()[-2:] for featmap in x]
+ device = x[0].device
+ anchor_list, valid_flag_list = self.stages[0].get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+
+ losses = dict()
+
+ for i in range(self.num_stages):
+ stage = self.stages[i]
+
+ if stage.adapt_cfg['type'] == 'offset':
+ offset_list = stage.anchor_offset(anchor_list,
+ stage.anchor_strides,
+ featmap_sizes)
+ else:
+ offset_list = None
+ x, cls_score, bbox_pred = stage(x, offset_list)
+ rpn_loss_inputs = (anchor_list, valid_flag_list, cls_score,
+ bbox_pred, batch_gt_instances, batch_img_metas)
+ stage_loss = stage.loss_by_feat(*rpn_loss_inputs)
+ for name, value in stage_loss.items():
+ losses['s{}.{}'.format(i, name)] = value
+
+ # refine boxes
+ if i < self.num_stages - 1:
+ anchor_list = stage.refine_bboxes(anchor_list, bbox_pred,
+ batch_img_metas)
+
+ return losses
+
+ def loss_and_predict(
+ self,
+ x: Tuple[Tensor],
+ batch_data_samples: SampleList,
+ proposal_cfg: Optional[ConfigDict] = None,
+ ) -> Tuple[dict, InstanceList]:
+ """Perform forward propagation of the head, then calculate loss and
+ predictions from the features and data samples.
+
+ Args:
+ x (tuple[Tensor]): Features from FPN.
+ batch_data_samples (list[:obj:`DetDataSample`]): Each item contains
+ the meta information of each image and corresponding
+ annotations.
+ proposal_cfg (ConfigDict, optional): Test / postprocessing
+ configuration, if None, test_cfg would be used.
+ Defaults to None.
+
+ Returns:
+ tuple: the return value is a tuple contains:
+
+ - losses: (dict[str, Tensor]): A dictionary of loss components.
+ - predictions (list[:obj:`InstanceData`]): Detection
+ results of each image after the post process.
+ """
+ outputs = unpack_gt_instances(batch_data_samples)
+ batch_gt_instances, _, batch_img_metas = outputs
+
+ featmap_sizes = [featmap.size()[-2:] for featmap in x]
+ device = x[0].device
+ anchor_list, valid_flag_list = self.stages[0].get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+
+ losses = dict()
+
+ for i in range(self.num_stages):
+ stage = self.stages[i]
+
+ if stage.adapt_cfg['type'] == 'offset':
+ offset_list = stage.anchor_offset(anchor_list,
+ stage.anchor_strides,
+ featmap_sizes)
+ else:
+ offset_list = None
+ x, cls_score, bbox_pred = stage(x, offset_list)
+ rpn_loss_inputs = (anchor_list, valid_flag_list, cls_score,
+ bbox_pred, batch_gt_instances, batch_img_metas)
+ stage_loss = stage.loss_by_feat(*rpn_loss_inputs)
+ for name, value in stage_loss.items():
+ losses['s{}.{}'.format(i, name)] = value
+
+ # refine boxes
+ if i < self.num_stages - 1:
+ anchor_list = stage.refine_bboxes(anchor_list, bbox_pred,
+ batch_img_metas)
+
+ predictions = self.stages[-1].predict_by_feat(
+ anchor_list,
+ cls_score,
+ bbox_pred,
+ batch_img_metas=batch_img_metas,
+ cfg=proposal_cfg)
+ return losses, predictions
+
+ def predict(self,
+ x: Tuple[Tensor],
+ batch_data_samples: SampleList,
+ rescale: bool = False) -> InstanceList:
+ """Perform forward propagation of the detection head and predict
+ detection results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Multi-level features from the
+ upstream network, each is a 4D-tensor.
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool, optional): Whether to rescale the results.
+ Defaults to False.
+
+ Returns:
+ list[obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ """
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+
+ featmap_sizes = [featmap.size()[-2:] for featmap in x]
+ device = x[0].device
+ anchor_list, _ = self.stages[0].get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+
+ for i in range(self.num_stages):
+ stage = self.stages[i]
+ if stage.adapt_cfg['type'] == 'offset':
+ offset_list = stage.anchor_offset(anchor_list,
+ stage.anchor_strides,
+ featmap_sizes)
+ else:
+ offset_list = None
+ x, cls_score, bbox_pred = stage(x, offset_list)
+ if i < self.num_stages - 1:
+ anchor_list = stage.refine_bboxes(anchor_list, bbox_pred,
+ batch_img_metas)
+
+ predictions = self.stages[-1].predict_by_feat(
+ anchor_list,
+ cls_score,
+ bbox_pred,
+ batch_img_metas=batch_img_metas,
+ rescale=rescale)
+ return predictions
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/centernet_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/centernet_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..09f3e599eb176965e53f270014cbd326858b7c17
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/centernet_head.py
@@ -0,0 +1,447 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple
+
+import torch
+import torch.nn as nn
+from mmcv.ops import batched_nms
+from mmengine.config import ConfigDict
+from mmengine.model import bias_init_with_prob, normal_init
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import (ConfigType, InstanceList, OptConfigType,
+ OptInstanceList, OptMultiConfig)
+from ..utils import (gaussian_radius, gen_gaussian_target, get_local_maximum,
+ get_topk_from_heatmap, multi_apply,
+ transpose_and_gather_feat)
+from .base_dense_head import BaseDenseHead
+
+
+@MODELS.register_module()
+class CenterNetHead(BaseDenseHead):
+ """Objects as Points Head. CenterHead use center_point to indicate object's
+ position. Paper link
+
+ Args:
+ in_channels (int): Number of channel in the input feature map.
+ feat_channels (int): Number of channel in the intermediate feature map.
+ num_classes (int): Number of categories excluding the background
+ category.
+ loss_center_heatmap (:obj:`ConfigDict` or dict): Config of center
+ heatmap loss. Defaults to
+ dict(type='GaussianFocalLoss', loss_weight=1.0)
+ loss_wh (:obj:`ConfigDict` or dict): Config of wh loss. Defaults to
+ dict(type='L1Loss', loss_weight=0.1).
+ loss_offset (:obj:`ConfigDict` or dict): Config of offset loss.
+ Defaults to dict(type='L1Loss', loss_weight=1.0).
+ train_cfg (:obj:`ConfigDict` or dict, optional): Training config.
+ Useless in CenterNet, but we keep this variable for
+ SingleStageDetector.
+ test_cfg (:obj:`ConfigDict` or dict, optional): Testing config
+ of CenterNet.
+ init_cfg (:obj:`ConfigDict` or dict or list[dict] or
+ list[:obj:`ConfigDict`], optional): Initialization
+ config dict.
+ """
+
+ def __init__(self,
+ in_channels: int,
+ feat_channels: int,
+ num_classes: int,
+ loss_center_heatmap: ConfigType = dict(
+ type='GaussianFocalLoss', loss_weight=1.0),
+ loss_wh: ConfigType = dict(type='L1Loss', loss_weight=0.1),
+ loss_offset: ConfigType = dict(
+ type='L1Loss', loss_weight=1.0),
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.num_classes = num_classes
+ self.heatmap_head = self._build_head(in_channels, feat_channels,
+ num_classes)
+ self.wh_head = self._build_head(in_channels, feat_channels, 2)
+ self.offset_head = self._build_head(in_channels, feat_channels, 2)
+
+ self.loss_center_heatmap = MODELS.build(loss_center_heatmap)
+ self.loss_wh = MODELS.build(loss_wh)
+ self.loss_offset = MODELS.build(loss_offset)
+
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+ self.fp16_enabled = False
+
+ def _build_head(self, in_channels: int, feat_channels: int,
+ out_channels: int) -> nn.Sequential:
+ """Build head for each branch."""
+ layer = nn.Sequential(
+ nn.Conv2d(in_channels, feat_channels, kernel_size=3, padding=1),
+ nn.ReLU(inplace=True),
+ nn.Conv2d(feat_channels, out_channels, kernel_size=1))
+ return layer
+
+ def init_weights(self) -> None:
+ """Initialize weights of the head."""
+ bias_init = bias_init_with_prob(0.1)
+ self.heatmap_head[-1].bias.data.fill_(bias_init)
+ for head in [self.wh_head, self.offset_head]:
+ for m in head.modules():
+ if isinstance(m, nn.Conv2d):
+ normal_init(m, std=0.001)
+
+ def forward(self, x: Tuple[Tensor, ...]) -> Tuple[List[Tensor]]:
+ """Forward features. Notice CenterNet head does not use FPN.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ center_heatmap_preds (list[Tensor]): center predict heatmaps for
+ all levels, the channels number is num_classes.
+ wh_preds (list[Tensor]): wh predicts for all levels, the channels
+ number is 2.
+ offset_preds (list[Tensor]): offset predicts for all levels, the
+ channels number is 2.
+ """
+ return multi_apply(self.forward_single, x)
+
+ def forward_single(self, x: Tensor) -> Tuple[Tensor, ...]:
+ """Forward feature of a single level.
+
+ Args:
+ x (Tensor): Feature of a single level.
+
+ Returns:
+ center_heatmap_pred (Tensor): center predict heatmaps, the
+ channels number is num_classes.
+ wh_pred (Tensor): wh predicts, the channels number is 2.
+ offset_pred (Tensor): offset predicts, the channels number is 2.
+ """
+ center_heatmap_pred = self.heatmap_head(x).sigmoid()
+ wh_pred = self.wh_head(x)
+ offset_pred = self.offset_head(x)
+ return center_heatmap_pred, wh_pred, offset_pred
+
+ def loss_by_feat(
+ self,
+ center_heatmap_preds: List[Tensor],
+ wh_preds: List[Tensor],
+ offset_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Compute losses of the head.
+
+ Args:
+ center_heatmap_preds (list[Tensor]): center predict heatmaps for
+ all levels with shape (B, num_classes, H, W).
+ wh_preds (list[Tensor]): wh predicts for all levels with
+ shape (B, 2, H, W).
+ offset_preds (list[Tensor]): offset predicts for all levels
+ with shape (B, 2, H, W).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: which has components below:
+ - loss_center_heatmap (Tensor): loss of center heatmap.
+ - loss_wh (Tensor): loss of hw heatmap
+ - loss_offset (Tensor): loss of offset heatmap.
+ """
+ assert len(center_heatmap_preds) == len(wh_preds) == len(
+ offset_preds) == 1
+ center_heatmap_pred = center_heatmap_preds[0]
+ wh_pred = wh_preds[0]
+ offset_pred = offset_preds[0]
+
+ gt_bboxes = [
+ gt_instances.bboxes for gt_instances in batch_gt_instances
+ ]
+ gt_labels = [
+ gt_instances.labels for gt_instances in batch_gt_instances
+ ]
+ img_shape = batch_img_metas[0]['batch_input_shape']
+ target_result, avg_factor = self.get_targets(gt_bboxes, gt_labels,
+ center_heatmap_pred.shape,
+ img_shape)
+
+ center_heatmap_target = target_result['center_heatmap_target']
+ wh_target = target_result['wh_target']
+ offset_target = target_result['offset_target']
+ wh_offset_target_weight = target_result['wh_offset_target_weight']
+
+ # Since the channel of wh_target and offset_target is 2, the avg_factor
+ # of loss_center_heatmap is always 1/2 of loss_wh and loss_offset.
+ loss_center_heatmap = self.loss_center_heatmap(
+ center_heatmap_pred, center_heatmap_target, avg_factor=avg_factor)
+ loss_wh = self.loss_wh(
+ wh_pred,
+ wh_target,
+ wh_offset_target_weight,
+ avg_factor=avg_factor * 2)
+ loss_offset = self.loss_offset(
+ offset_pred,
+ offset_target,
+ wh_offset_target_weight,
+ avg_factor=avg_factor * 2)
+ return dict(
+ loss_center_heatmap=loss_center_heatmap,
+ loss_wh=loss_wh,
+ loss_offset=loss_offset)
+
+ def get_targets(self, gt_bboxes: List[Tensor], gt_labels: List[Tensor],
+ feat_shape: tuple, img_shape: tuple) -> Tuple[dict, int]:
+ """Compute regression and classification targets in multiple images.
+
+ Args:
+ gt_bboxes (list[Tensor]): Ground truth bboxes for each image with
+ shape (num_gts, 4) in [tl_x, tl_y, br_x, br_y] format.
+ gt_labels (list[Tensor]): class indices corresponding to each box.
+ feat_shape (tuple): feature map shape with value [B, _, H, W]
+ img_shape (tuple): image shape.
+
+ Returns:
+ tuple[dict, float]: The float value is mean avg_factor, the dict
+ has components below:
+ - center_heatmap_target (Tensor): targets of center heatmap, \
+ shape (B, num_classes, H, W).
+ - wh_target (Tensor): targets of wh predict, shape \
+ (B, 2, H, W).
+ - offset_target (Tensor): targets of offset predict, shape \
+ (B, 2, H, W).
+ - wh_offset_target_weight (Tensor): weights of wh and offset \
+ predict, shape (B, 2, H, W).
+ """
+ img_h, img_w = img_shape[:2]
+ bs, _, feat_h, feat_w = feat_shape
+
+ width_ratio = float(feat_w / img_w)
+ height_ratio = float(feat_h / img_h)
+
+ center_heatmap_target = gt_bboxes[-1].new_zeros(
+ [bs, self.num_classes, feat_h, feat_w])
+ wh_target = gt_bboxes[-1].new_zeros([bs, 2, feat_h, feat_w])
+ offset_target = gt_bboxes[-1].new_zeros([bs, 2, feat_h, feat_w])
+ wh_offset_target_weight = gt_bboxes[-1].new_zeros(
+ [bs, 2, feat_h, feat_w])
+
+ for batch_id in range(bs):
+ gt_bbox = gt_bboxes[batch_id]
+ gt_label = gt_labels[batch_id]
+ center_x = (gt_bbox[:, [0]] + gt_bbox[:, [2]]) * width_ratio / 2
+ center_y = (gt_bbox[:, [1]] + gt_bbox[:, [3]]) * height_ratio / 2
+ gt_centers = torch.cat((center_x, center_y), dim=1)
+
+ for j, ct in enumerate(gt_centers):
+ ctx_int, cty_int = ct.int()
+ ctx, cty = ct
+ scale_box_h = (gt_bbox[j][3] - gt_bbox[j][1]) * height_ratio
+ scale_box_w = (gt_bbox[j][2] - gt_bbox[j][0]) * width_ratio
+ radius = gaussian_radius([scale_box_h, scale_box_w],
+ min_overlap=0.3)
+ radius = max(0, int(radius))
+ ind = gt_label[j]
+ gen_gaussian_target(center_heatmap_target[batch_id, ind],
+ [ctx_int, cty_int], radius)
+
+ wh_target[batch_id, 0, cty_int, ctx_int] = scale_box_w
+ wh_target[batch_id, 1, cty_int, ctx_int] = scale_box_h
+
+ offset_target[batch_id, 0, cty_int, ctx_int] = ctx - ctx_int
+ offset_target[batch_id, 1, cty_int, ctx_int] = cty - cty_int
+
+ wh_offset_target_weight[batch_id, :, cty_int, ctx_int] = 1
+
+ avg_factor = max(1, center_heatmap_target.eq(1).sum())
+ target_result = dict(
+ center_heatmap_target=center_heatmap_target,
+ wh_target=wh_target,
+ offset_target=offset_target,
+ wh_offset_target_weight=wh_offset_target_weight)
+ return target_result, avg_factor
+
+ def predict_by_feat(self,
+ center_heatmap_preds: List[Tensor],
+ wh_preds: List[Tensor],
+ offset_preds: List[Tensor],
+ batch_img_metas: Optional[List[dict]] = None,
+ rescale: bool = True,
+ with_nms: bool = False) -> InstanceList:
+ """Transform network output for a batch into bbox predictions.
+
+ Args:
+ center_heatmap_preds (list[Tensor]): Center predict heatmaps for
+ all levels with shape (B, num_classes, H, W).
+ wh_preds (list[Tensor]): WH predicts for all levels with
+ shape (B, 2, H, W).
+ offset_preds (list[Tensor]): Offset predicts for all levels
+ with shape (B, 2, H, W).
+ batch_img_metas (list[dict], optional): Batch image meta info.
+ Defaults to None.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to True.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Instance segmentation
+ results of each image after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ assert len(center_heatmap_preds) == len(wh_preds) == len(
+ offset_preds) == 1
+ result_list = []
+ for img_id in range(len(batch_img_metas)):
+ result_list.append(
+ self._predict_by_feat_single(
+ center_heatmap_preds[0][img_id:img_id + 1, ...],
+ wh_preds[0][img_id:img_id + 1, ...],
+ offset_preds[0][img_id:img_id + 1, ...],
+ batch_img_metas[img_id],
+ rescale=rescale,
+ with_nms=with_nms))
+ return result_list
+
+ def _predict_by_feat_single(self,
+ center_heatmap_pred: Tensor,
+ wh_pred: Tensor,
+ offset_pred: Tensor,
+ img_meta: dict,
+ rescale: bool = True,
+ with_nms: bool = False) -> InstanceData:
+ """Transform outputs of a single image into bbox results.
+
+ Args:
+ center_heatmap_pred (Tensor): Center heatmap for current level with
+ shape (1, num_classes, H, W).
+ wh_pred (Tensor): WH heatmap for current level with shape
+ (1, num_classes, H, W).
+ offset_pred (Tensor): Offset for current level with shape
+ (1, corner_offset_channels, H, W).
+ img_meta (dict): Meta information of current image, e.g.,
+ image size, scaling factor, etc.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to True.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to False.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ batch_det_bboxes, batch_labels = self._decode_heatmap(
+ center_heatmap_pred,
+ wh_pred,
+ offset_pred,
+ img_meta['batch_input_shape'],
+ k=self.test_cfg.topk,
+ kernel=self.test_cfg.local_maximum_kernel)
+
+ det_bboxes = batch_det_bboxes.view([-1, 5])
+ det_labels = batch_labels.view(-1)
+
+ batch_border = det_bboxes.new_tensor(img_meta['border'])[...,
+ [2, 0, 2, 0]]
+ det_bboxes[..., :4] -= batch_border
+
+ if rescale and 'scale_factor' in img_meta:
+ det_bboxes[..., :4] /= det_bboxes.new_tensor(
+ img_meta['scale_factor']).repeat((1, 2))
+
+ if with_nms:
+ det_bboxes, det_labels = self._bboxes_nms(det_bboxes, det_labels,
+ self.test_cfg)
+ results = InstanceData()
+ results.bboxes = det_bboxes[..., :4]
+ results.scores = det_bboxes[..., 4]
+ results.labels = det_labels
+ return results
+
+ def _decode_heatmap(self,
+ center_heatmap_pred: Tensor,
+ wh_pred: Tensor,
+ offset_pred: Tensor,
+ img_shape: tuple,
+ k: int = 100,
+ kernel: int = 3) -> Tuple[Tensor, Tensor]:
+ """Transform outputs into detections raw bbox prediction.
+
+ Args:
+ center_heatmap_pred (Tensor): center predict heatmap,
+ shape (B, num_classes, H, W).
+ wh_pred (Tensor): wh predict, shape (B, 2, H, W).
+ offset_pred (Tensor): offset predict, shape (B, 2, H, W).
+ img_shape (tuple): image shape in hw format.
+ k (int): Get top k center keypoints from heatmap. Defaults to 100.
+ kernel (int): Max pooling kernel for extract local maximum pixels.
+ Defaults to 3.
+
+ Returns:
+ tuple[Tensor]: Decoded output of CenterNetHead, containing
+ the following Tensors:
+
+ - batch_bboxes (Tensor): Coords of each box with shape (B, k, 5)
+ - batch_topk_labels (Tensor): Categories of each box with \
+ shape (B, k)
+ """
+ height, width = center_heatmap_pred.shape[2:]
+ inp_h, inp_w = img_shape
+
+ center_heatmap_pred = get_local_maximum(
+ center_heatmap_pred, kernel=kernel)
+
+ *batch_dets, topk_ys, topk_xs = get_topk_from_heatmap(
+ center_heatmap_pred, k=k)
+ batch_scores, batch_index, batch_topk_labels = batch_dets
+
+ wh = transpose_and_gather_feat(wh_pred, batch_index)
+ offset = transpose_and_gather_feat(offset_pred, batch_index)
+ topk_xs = topk_xs + offset[..., 0]
+ topk_ys = topk_ys + offset[..., 1]
+ tl_x = (topk_xs - wh[..., 0] / 2) * (inp_w / width)
+ tl_y = (topk_ys - wh[..., 1] / 2) * (inp_h / height)
+ br_x = (topk_xs + wh[..., 0] / 2) * (inp_w / width)
+ br_y = (topk_ys + wh[..., 1] / 2) * (inp_h / height)
+
+ batch_bboxes = torch.stack([tl_x, tl_y, br_x, br_y], dim=2)
+ batch_bboxes = torch.cat((batch_bboxes, batch_scores[..., None]),
+ dim=-1)
+ return batch_bboxes, batch_topk_labels
+
+ def _bboxes_nms(self, bboxes: Tensor, labels: Tensor,
+ cfg: ConfigDict) -> Tuple[Tensor, Tensor]:
+ """bboxes nms."""
+ if labels.numel() > 0:
+ max_num = cfg.max_per_img
+ bboxes, keep = batched_nms(bboxes[:, :4], bboxes[:,
+ -1].contiguous(),
+ labels, cfg.nms)
+ if max_num > 0:
+ bboxes = bboxes[:max_num]
+ labels = labels[keep][:max_num]
+
+ return bboxes, labels
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/centernet_update_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/centernet_update_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..00cfcb89806209c9416b1bd7e9a14d82a4911175
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/centernet_update_head.py
@@ -0,0 +1,624 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, List, Optional, Sequence, Tuple
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import Scale
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures.bbox import bbox2distance
+from mmdet.utils import (ConfigType, InstanceList, OptConfigType,
+ OptInstanceList, reduce_mean)
+from ..utils import multi_apply
+from .anchor_free_head import AnchorFreeHead
+
+INF = 1000000000
+RangeType = Sequence[Tuple[int, int]]
+
+
+def _transpose(tensor_list: List[Tensor],
+ num_point_list: list) -> List[Tensor]:
+ """This function is used to transpose image first tensors to level first
+ ones."""
+ for img_idx in range(len(tensor_list)):
+ tensor_list[img_idx] = torch.split(
+ tensor_list[img_idx], num_point_list, dim=0)
+
+ tensors_level_first = []
+ for targets_per_level in zip(*tensor_list):
+ tensors_level_first.append(torch.cat(targets_per_level, dim=0))
+ return tensors_level_first
+
+
+@MODELS.register_module()
+class CenterNetUpdateHead(AnchorFreeHead):
+ """CenterNetUpdateHead is an improved version of CenterNet in CenterNet2.
+ Paper link ``_.
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channel in the input feature map.
+ regress_ranges (Sequence[Tuple[int, int]]): Regress range of multiple
+ level points.
+ hm_min_radius (int): Heatmap target minimum radius of cls branch.
+ Defaults to 4.
+ hm_min_overlap (float): Heatmap target minimum overlap of cls branch.
+ Defaults to 0.8.
+ more_pos_thresh (float): The filtering threshold when the cls branch
+ adds more positive samples. Defaults to 0.2.
+ more_pos_topk (int): The maximum number of additional positive samples
+ added to each gt. Defaults to 9.
+ soft_weight_on_reg (bool): Whether to use the soft target of the
+ cls branch as the soft weight of the bbox branch.
+ Defaults to False.
+ loss_cls (:obj:`ConfigDict` or dict): Config of cls loss. Defaults to
+ dict(type='GaussianFocalLoss', loss_weight=1.0)
+ loss_bbox (:obj:`ConfigDict` or dict): Config of bbox loss. Defaults to
+ dict(type='GIoULoss', loss_weight=2.0).
+ norm_cfg (:obj:`ConfigDict` or dict, optional): dictionary to construct
+ and config norm layer. Defaults to
+ ``norm_cfg=dict(type='GN', num_groups=32, requires_grad=True)``.
+ train_cfg (:obj:`ConfigDict` or dict, optional): Training config.
+ Unused in CenterNet. Reserved for compatibility with
+ SingleStageDetector.
+ test_cfg (:obj:`ConfigDict` or dict, optional): Testing config
+ of CenterNet.
+ """
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: int,
+ regress_ranges: RangeType = ((0, 80), (64, 160), (128, 320),
+ (256, 640), (512, INF)),
+ hm_min_radius: int = 4,
+ hm_min_overlap: float = 0.8,
+ more_pos_thresh: float = 0.2,
+ more_pos_topk: int = 9,
+ soft_weight_on_reg: bool = False,
+ loss_cls: ConfigType = dict(
+ type='GaussianFocalLoss',
+ pos_weight=0.25,
+ neg_weight=0.75,
+ loss_weight=1.0),
+ loss_bbox: ConfigType = dict(
+ type='GIoULoss', loss_weight=2.0),
+ norm_cfg: OptConfigType = dict(
+ type='GN', num_groups=32, requires_grad=True),
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ **kwargs) -> None:
+ super().__init__(
+ num_classes=num_classes,
+ in_channels=in_channels,
+ loss_cls=loss_cls,
+ loss_bbox=loss_bbox,
+ norm_cfg=norm_cfg,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ **kwargs)
+ self.soft_weight_on_reg = soft_weight_on_reg
+ self.hm_min_radius = hm_min_radius
+ self.more_pos_thresh = more_pos_thresh
+ self.more_pos_topk = more_pos_topk
+ self.delta = (1 - hm_min_overlap) / (1 + hm_min_overlap)
+ self.sigmoid_clamp = 0.0001
+
+ # GaussianFocalLoss must be sigmoid mode
+ self.use_sigmoid_cls = True
+ self.cls_out_channels = num_classes
+
+ self.regress_ranges = regress_ranges
+ self.scales = nn.ModuleList([Scale(1.0) for _ in self.strides])
+
+ def _init_predictor(self) -> None:
+ """Initialize predictor layers of the head."""
+ self.conv_cls = nn.Conv2d(
+ self.feat_channels, self.num_classes, 3, padding=1)
+ self.conv_reg = nn.Conv2d(self.feat_channels, 4, 3, padding=1)
+
+ def forward(self, x: Tuple[Tensor]) -> Tuple[List[Tensor], List[Tensor]]:
+ """Forward features from the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: A tuple of each level outputs.
+
+ - cls_scores (list[Tensor]): Box scores for each scale level, \
+ each is a 4D-tensor, the channel number is num_classes.
+ - bbox_preds (list[Tensor]): Box energies / deltas for each \
+ scale level, each is a 4D-tensor, the channel number is 4.
+ """
+ return multi_apply(self.forward_single, x, self.scales, self.strides)
+
+ def forward_single(self, x: Tensor, scale: Scale,
+ stride: int) -> Tuple[Tensor, Tensor]:
+ """Forward features of a single scale level.
+
+ Args:
+ x (Tensor): FPN feature maps of the specified stride.
+ scale (:obj:`mmcv.cnn.Scale`): Learnable scale module to resize
+ the bbox prediction.
+ stride (int): The corresponding stride for feature maps.
+
+ Returns:
+ tuple: scores for each class, bbox predictions of
+ input feature maps.
+ """
+ cls_score, bbox_pred, _, _ = super().forward_single(x)
+ # scale the bbox_pred of different level
+ # float to avoid overflow when enabling FP16
+ bbox_pred = scale(bbox_pred).float()
+ # bbox_pred needed for gradient computation has been modified
+ # by F.relu(bbox_pred) when run with PyTorch 1.10. So replace
+ # F.relu(bbox_pred) with bbox_pred.clamp(min=0)
+ bbox_pred = bbox_pred.clamp(min=0)
+ if not self.training:
+ bbox_pred *= stride
+ return cls_score, bbox_pred
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None
+ ) -> Dict[str, Tensor]:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level,
+ each is a 4D-tensor, the channel number is num_classes.
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level, each is a 4D-tensor, the channel number is 4.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ num_imgs = cls_scores[0].size(0)
+ assert len(cls_scores) == len(bbox_preds)
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ all_level_points = self.prior_generator.grid_priors(
+ featmap_sizes,
+ dtype=bbox_preds[0].dtype,
+ device=bbox_preds[0].device)
+
+ # 1 flatten outputs
+ flatten_cls_scores = [
+ cls_score.permute(0, 2, 3, 1).reshape(-1, self.cls_out_channels)
+ for cls_score in cls_scores
+ ]
+ flatten_bbox_preds = [
+ bbox_pred.permute(0, 2, 3, 1).reshape(-1, 4)
+ for bbox_pred in bbox_preds
+ ]
+ flatten_cls_scores = torch.cat(flatten_cls_scores)
+ flatten_bbox_preds = torch.cat(flatten_bbox_preds)
+
+ # repeat points to align with bbox_preds
+ flatten_points = torch.cat(
+ [points.repeat(num_imgs, 1) for points in all_level_points])
+
+ assert (torch.isfinite(flatten_bbox_preds).all().item())
+
+ # 2 calc reg and cls branch targets
+ cls_targets, bbox_targets = self.get_targets(all_level_points,
+ batch_gt_instances)
+
+ # 3 add more pos index for cls branch
+ featmap_sizes = flatten_points.new_tensor(featmap_sizes)
+ pos_inds, cls_labels = self.add_cls_pos_inds(flatten_points,
+ flatten_bbox_preds,
+ featmap_sizes,
+ batch_gt_instances)
+
+ # 4 calc cls loss
+ if pos_inds is None:
+ # num_gts=0
+ num_pos_cls = bbox_preds[0].new_tensor(0, dtype=torch.float)
+ else:
+ num_pos_cls = bbox_preds[0].new_tensor(
+ len(pos_inds), dtype=torch.float)
+ num_pos_cls = max(reduce_mean(num_pos_cls), 1.0)
+ flatten_cls_scores = flatten_cls_scores.sigmoid().clamp(
+ min=self.sigmoid_clamp, max=1 - self.sigmoid_clamp)
+ cls_loss = self.loss_cls(
+ flatten_cls_scores,
+ cls_targets,
+ pos_inds=pos_inds,
+ pos_labels=cls_labels,
+ avg_factor=num_pos_cls)
+
+ # 5 calc reg loss
+ pos_bbox_inds = torch.nonzero(
+ bbox_targets.max(dim=1)[0] >= 0).squeeze(1)
+ pos_bbox_preds = flatten_bbox_preds[pos_bbox_inds]
+ pos_bbox_targets = bbox_targets[pos_bbox_inds]
+
+ bbox_weight_map = cls_targets.max(dim=1)[0]
+ bbox_weight_map = bbox_weight_map[pos_bbox_inds]
+ bbox_weight_map = bbox_weight_map if self.soft_weight_on_reg \
+ else torch.ones_like(bbox_weight_map)
+ num_pos_bbox = max(reduce_mean(bbox_weight_map.sum()), 1.0)
+
+ if len(pos_bbox_inds) > 0:
+ pos_points = flatten_points[pos_bbox_inds]
+ pos_decoded_bbox_preds = self.bbox_coder.decode(
+ pos_points, pos_bbox_preds)
+ pos_decoded_target_preds = self.bbox_coder.decode(
+ pos_points, pos_bbox_targets)
+ bbox_loss = self.loss_bbox(
+ pos_decoded_bbox_preds,
+ pos_decoded_target_preds,
+ weight=bbox_weight_map,
+ avg_factor=num_pos_bbox)
+ else:
+ bbox_loss = flatten_bbox_preds.sum() * 0
+
+ return dict(loss_cls=cls_loss, loss_bbox=bbox_loss)
+
+ def get_targets(
+ self,
+ points: List[Tensor],
+ batch_gt_instances: InstanceList,
+ ) -> Tuple[Tensor, Tensor]:
+ """Compute classification and bbox targets for points in multiple
+ images.
+
+ Args:
+ points (list[Tensor]): Points of each fpn level, each has shape
+ (num_points, 2).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+
+ Returns:
+ tuple: Targets of each level.
+
+ - concat_lvl_labels (Tensor): Labels of all level and batch.
+ - concat_lvl_bbox_targets (Tensor): BBox targets of all \
+ level and batch.
+ """
+ assert len(points) == len(self.regress_ranges)
+
+ num_levels = len(points)
+ # the number of points per img, per lvl
+ num_points = [center.size(0) for center in points]
+
+ # expand regress ranges to align with points
+ expanded_regress_ranges = [
+ points[i].new_tensor(self.regress_ranges[i])[None].expand_as(
+ points[i]) for i in range(num_levels)
+ ]
+ # concat all levels points and regress ranges
+ concat_regress_ranges = torch.cat(expanded_regress_ranges, dim=0)
+ concat_points = torch.cat(points, dim=0)
+ concat_strides = torch.cat([
+ concat_points.new_ones(num_points[i]) * self.strides[i]
+ for i in range(num_levels)
+ ])
+
+ # get labels and bbox_targets of each image
+ cls_targets_list, bbox_targets_list = multi_apply(
+ self._get_targets_single,
+ batch_gt_instances,
+ points=concat_points,
+ regress_ranges=concat_regress_ranges,
+ strides=concat_strides)
+
+ bbox_targets_list = _transpose(bbox_targets_list, num_points)
+ cls_targets_list = _transpose(cls_targets_list, num_points)
+ concat_lvl_bbox_targets = torch.cat(bbox_targets_list, 0)
+ concat_lvl_cls_targets = torch.cat(cls_targets_list, dim=0)
+ return concat_lvl_cls_targets, concat_lvl_bbox_targets
+
+ def _get_targets_single(self, gt_instances: InstanceData, points: Tensor,
+ regress_ranges: Tensor,
+ strides: Tensor) -> Tuple[Tensor, Tensor]:
+ """Compute classification and bbox targets for a single image."""
+ num_points = points.size(0)
+ num_gts = len(gt_instances)
+ gt_bboxes = gt_instances.bboxes
+ gt_labels = gt_instances.labels
+
+ if num_gts == 0:
+ return gt_labels.new_full((num_points,
+ self.num_classes),
+ self.num_classes), \
+ gt_bboxes.new_full((num_points, 4), -1)
+
+ # Calculate the regression tblr target corresponding to all points
+ points = points[:, None].expand(num_points, num_gts, 2)
+ gt_bboxes = gt_bboxes[None].expand(num_points, num_gts, 4)
+ strides = strides[:, None, None].expand(num_points, num_gts, 2)
+
+ bbox_target = bbox2distance(points, gt_bboxes) # M x N x 4
+
+ # condition1: inside a gt bbox
+ inside_gt_bbox_mask = bbox_target.min(dim=2)[0] > 0 # M x N
+
+ # condition2: Calculate the nearest points from
+ # the upper, lower, left and right ranges from
+ # the center of the gt bbox
+ centers = ((gt_bboxes[..., [0, 1]] + gt_bboxes[..., [2, 3]]) / 2)
+ centers_discret = ((centers / strides).int() * strides).float() + \
+ strides / 2
+
+ centers_discret_dist = points - centers_discret
+ dist_x = centers_discret_dist[..., 0].abs()
+ dist_y = centers_discret_dist[..., 1].abs()
+ inside_gt_center3x3_mask = (dist_x <= strides[..., 0]) & \
+ (dist_y <= strides[..., 0])
+
+ # condition3: limit the regression range for each location
+ bbox_target_wh = bbox_target[..., :2] + bbox_target[..., 2:]
+ crit = (bbox_target_wh**2).sum(dim=2)**0.5 / 2
+ inside_fpn_level_mask = (crit >= regress_ranges[:, [0]]) & \
+ (crit <= regress_ranges[:, [1]])
+ bbox_target_mask = inside_gt_bbox_mask & \
+ inside_gt_center3x3_mask & \
+ inside_fpn_level_mask
+
+ # Calculate the distance weight map
+ gt_center_peak_mask = ((centers_discret_dist**2).sum(dim=2) == 0)
+ weighted_dist = ((points - centers)**2).sum(dim=2) # M x N
+ weighted_dist[gt_center_peak_mask] = 0
+
+ areas = (gt_bboxes[..., 2] - gt_bboxes[..., 0]) * (
+ gt_bboxes[..., 3] - gt_bboxes[..., 1])
+ radius = self.delta**2 * 2 * areas
+ radius = torch.clamp(radius, min=self.hm_min_radius**2)
+ weighted_dist = weighted_dist / radius
+
+ # Calculate bbox_target
+ bbox_weighted_dist = weighted_dist.clone()
+ bbox_weighted_dist[bbox_target_mask == 0] = INF * 1.0
+ min_dist, min_inds = bbox_weighted_dist.min(dim=1)
+ bbox_target = bbox_target[range(len(bbox_target)),
+ min_inds] # M x N x 4 --> M x 4
+ bbox_target[min_dist == INF] = -INF
+
+ # Convert to feature map scale
+ bbox_target /= strides[:, 0, :].repeat(1, 2)
+
+ # Calculate cls_target
+ cls_target = self._create_heatmaps_from_dist(weighted_dist, gt_labels)
+
+ return cls_target, bbox_target
+
+ @torch.no_grad()
+ def add_cls_pos_inds(
+ self, flatten_points: Tensor, flatten_bbox_preds: Tensor,
+ featmap_sizes: Tensor, batch_gt_instances: InstanceList
+ ) -> Tuple[Optional[Tensor], Optional[Tensor]]:
+ """Provide additional adaptive positive samples to the classification
+ branch.
+
+ Args:
+ flatten_points (Tensor): The point after flatten, including
+ batch image and all levels. The shape is (N, 2).
+ flatten_bbox_preds (Tensor): The bbox predicts after flatten,
+ including batch image and all levels. The shape is (N, 4).
+ featmap_sizes (Tensor): Feature map size of all layers.
+ The shape is (5, 2).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+
+ Returns:
+ tuple:
+
+ - pos_inds (Tensor): Adaptively selected positive sample index.
+ - cls_labels (Tensor): Corresponding positive class label.
+ """
+ outputs = self._get_center3x3_region_index_targets(
+ batch_gt_instances, featmap_sizes)
+ cls_labels, fpn_level_masks, center3x3_inds, \
+ center3x3_bbox_targets, center3x3_masks = outputs
+
+ num_gts, total_level, K = cls_labels.shape[0], len(
+ self.strides), center3x3_masks.shape[-1]
+
+ if num_gts == 0:
+ return None, None
+
+ # The out-of-bounds index is forcibly set to 0
+ # to prevent loss calculation errors
+ center3x3_inds[center3x3_masks == 0] = 0
+ reg_pred_center3x3 = flatten_bbox_preds[center3x3_inds]
+ center3x3_points = flatten_points[center3x3_inds].view(-1, 2)
+
+ center3x3_bbox_targets_expand = center3x3_bbox_targets.view(
+ -1, 4).clamp(min=0)
+
+ pos_decoded_bbox_preds = self.bbox_coder.decode(
+ center3x3_points, reg_pred_center3x3.view(-1, 4))
+ pos_decoded_target_preds = self.bbox_coder.decode(
+ center3x3_points, center3x3_bbox_targets_expand)
+ center3x3_bbox_loss = self.loss_bbox(
+ pos_decoded_bbox_preds,
+ pos_decoded_target_preds,
+ None,
+ reduction_override='none').view(num_gts, total_level,
+ K) / self.loss_bbox.loss_weight
+
+ # Invalid index Loss set to infinity
+ center3x3_bbox_loss[center3x3_masks == 0] = INF
+
+ # 4 is the center point of the sampled 9 points, the center point
+ # of gt bbox after discretization.
+ # The center point of gt bbox after discretization
+ # must be a positive sample, so we force its loss to be set to 0.
+ center3x3_bbox_loss.view(-1, K)[fpn_level_masks.view(-1), 4] = 0
+ center3x3_bbox_loss = center3x3_bbox_loss.view(num_gts, -1)
+
+ loss_thr = torch.kthvalue(
+ center3x3_bbox_loss, self.more_pos_topk, dim=1)[0]
+
+ loss_thr[loss_thr > self.more_pos_thresh] = self.more_pos_thresh
+ new_pos = center3x3_bbox_loss < loss_thr.view(num_gts, 1)
+ pos_inds = center3x3_inds.view(num_gts, -1)[new_pos]
+ cls_labels = cls_labels.view(num_gts,
+ 1).expand(num_gts,
+ total_level * K)[new_pos]
+ return pos_inds, cls_labels
+
+ def _create_heatmaps_from_dist(self, weighted_dist: Tensor,
+ cls_labels: Tensor) -> Tensor:
+ """Generate heatmaps of classification branch based on weighted
+ distance map."""
+ heatmaps = weighted_dist.new_zeros(
+ (weighted_dist.shape[0], self.num_classes))
+ for c in range(self.num_classes):
+ inds = (cls_labels == c) # N
+ if inds.int().sum() == 0:
+ continue
+ heatmaps[:, c] = torch.exp(-weighted_dist[:, inds].min(dim=1)[0])
+ zeros = heatmaps[:, c] < 1e-4
+ heatmaps[zeros, c] = 0
+ return heatmaps
+
+ def _get_center3x3_region_index_targets(self,
+ bacth_gt_instances: InstanceList,
+ shapes_per_level: Tensor) -> tuple:
+ """Get the center (and the 3x3 region near center) locations and target
+ of each objects."""
+ cls_labels = []
+ inside_fpn_level_masks = []
+ center3x3_inds = []
+ center3x3_masks = []
+ center3x3_bbox_targets = []
+
+ total_levels = len(self.strides)
+ batch = len(bacth_gt_instances)
+
+ shapes_per_level = shapes_per_level.long()
+ area_per_level = (shapes_per_level[:, 0] * shapes_per_level[:, 1])
+
+ # Select a total of 9 positions of 3x3 in the center of the gt bbox
+ # as candidate positive samples
+ K = 9
+ dx = shapes_per_level.new_tensor([-1, 0, 1, -1, 0, 1, -1, 0,
+ 1]).view(1, 1, K)
+ dy = shapes_per_level.new_tensor([-1, -1, -1, 0, 0, 0, 1, 1,
+ 1]).view(1, 1, K)
+
+ regress_ranges = shapes_per_level.new_tensor(self.regress_ranges).view(
+ len(self.regress_ranges), 2) # L x 2
+ strides = shapes_per_level.new_tensor(self.strides)
+
+ start_coord_pre_level = []
+ _start = 0
+ for level in range(total_levels):
+ start_coord_pre_level.append(_start)
+ _start = _start + batch * area_per_level[level]
+ start_coord_pre_level = shapes_per_level.new_tensor(
+ start_coord_pre_level).view(1, total_levels, 1)
+ area_per_level = area_per_level.view(1, total_levels, 1)
+
+ for im_i in range(batch):
+ gt_instance = bacth_gt_instances[im_i]
+ gt_bboxes = gt_instance.bboxes
+ gt_labels = gt_instance.labels
+ num_gts = gt_bboxes.shape[0]
+ if num_gts == 0:
+ continue
+
+ cls_labels.append(gt_labels)
+
+ gt_bboxes = gt_bboxes[:, None].expand(num_gts, total_levels, 4)
+ expanded_strides = strides[None, :,
+ None].expand(num_gts, total_levels, 2)
+ expanded_regress_ranges = regress_ranges[None].expand(
+ num_gts, total_levels, 2)
+ expanded_shapes_per_level = shapes_per_level[None].expand(
+ num_gts, total_levels, 2)
+
+ # calc reg_target
+ centers = ((gt_bboxes[..., [0, 1]] + gt_bboxes[..., [2, 3]]) / 2)
+ centers_inds = (centers / expanded_strides).long()
+ centers_discret = centers_inds * expanded_strides \
+ + expanded_strides // 2
+
+ bbox_target = bbox2distance(centers_discret,
+ gt_bboxes) # M x N x 4
+
+ # calc inside_fpn_level_mask
+ bbox_target_wh = bbox_target[..., :2] + bbox_target[..., 2:]
+ crit = (bbox_target_wh**2).sum(dim=2)**0.5 / 2
+ inside_fpn_level_mask = \
+ (crit >= expanded_regress_ranges[..., 0]) & \
+ (crit <= expanded_regress_ranges[..., 1])
+
+ inside_gt_bbox_mask = bbox_target.min(dim=2)[0] >= 0
+ inside_fpn_level_mask = inside_gt_bbox_mask & inside_fpn_level_mask
+ inside_fpn_level_masks.append(inside_fpn_level_mask)
+
+ # calc center3x3_ind and mask
+ expand_ws = expanded_shapes_per_level[..., 1:2].expand(
+ num_gts, total_levels, K)
+ expand_hs = expanded_shapes_per_level[..., 0:1].expand(
+ num_gts, total_levels, K)
+ centers_inds_x = centers_inds[..., 0:1]
+ centers_inds_y = centers_inds[..., 1:2]
+
+ center3x3_idx = start_coord_pre_level + \
+ im_i * area_per_level + \
+ (centers_inds_y + dy) * expand_ws + \
+ (centers_inds_x + dx)
+ center3x3_mask = \
+ ((centers_inds_y + dy) < expand_hs) & \
+ ((centers_inds_y + dy) >= 0) & \
+ ((centers_inds_x + dx) < expand_ws) & \
+ ((centers_inds_x + dx) >= 0)
+
+ # recalc center3x3 region reg target
+ bbox_target = bbox_target / expanded_strides.repeat(1, 1, 2)
+ center3x3_bbox_target = bbox_target[..., None, :].expand(
+ num_gts, total_levels, K, 4).clone()
+ center3x3_bbox_target[..., 0] += dx
+ center3x3_bbox_target[..., 1] += dy
+ center3x3_bbox_target[..., 2] -= dx
+ center3x3_bbox_target[..., 3] -= dy
+ # update center3x3_mask
+ center3x3_mask = center3x3_mask & (
+ center3x3_bbox_target.min(dim=3)[0] >= 0) # n x L x K
+
+ center3x3_inds.append(center3x3_idx)
+ center3x3_masks.append(center3x3_mask)
+ center3x3_bbox_targets.append(center3x3_bbox_target)
+
+ if len(inside_fpn_level_masks) > 0:
+ cls_labels = torch.cat(cls_labels, dim=0)
+ inside_fpn_level_masks = torch.cat(inside_fpn_level_masks, dim=0)
+ center3x3_inds = torch.cat(center3x3_inds, dim=0).long()
+ center3x3_bbox_targets = torch.cat(center3x3_bbox_targets, dim=0)
+ center3x3_masks = torch.cat(center3x3_masks, dim=0)
+ else:
+ cls_labels = shapes_per_level.new_zeros(0).long()
+ inside_fpn_level_masks = shapes_per_level.new_zeros(
+ (0, total_levels)).bool()
+ center3x3_inds = shapes_per_level.new_zeros(
+ (0, total_levels, K)).long()
+ center3x3_bbox_targets = shapes_per_level.new_zeros(
+ (0, total_levels, K, 4)).float()
+ center3x3_masks = shapes_per_level.new_zeros(
+ (0, total_levels, K)).bool()
+ return cls_labels, inside_fpn_level_masks, center3x3_inds, \
+ center3x3_bbox_targets, center3x3_masks
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/centripetal_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/centripetal_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..18f6601ff82394864d53351b10b40f51eb2aec6b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/centripetal_head.py
@@ -0,0 +1,459 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple
+
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from mmcv.ops import DeformConv2d
+from mmengine.model import normal_init
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import (ConfigType, InstanceList, OptInstanceList,
+ OptMultiConfig)
+from ..utils import multi_apply
+from .corner_head import CornerHead
+
+
+@MODELS.register_module()
+class CentripetalHead(CornerHead):
+ """Head of CentripetalNet: Pursuing High-quality Keypoint Pairs for Object
+ Detection.
+
+ CentripetalHead inherits from :class:`CornerHead`. It removes the
+ embedding branch and adds guiding shift and centripetal shift branches.
+ More details can be found in the `paper
+ `_ .
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ num_feat_levels (int): Levels of feature from the previous module.
+ 2 for HourglassNet-104 and 1 for HourglassNet-52. HourglassNet-104
+ outputs the final feature and intermediate supervision feature and
+ HourglassNet-52 only outputs the final feature. Defaults to 2.
+ corner_emb_channels (int): Channel of embedding vector. Defaults to 1.
+ train_cfg (:obj:`ConfigDict` or dict, optional): Training config.
+ Useless in CornerHead, but we keep this variable for
+ SingleStageDetector.
+ test_cfg (:obj:`ConfigDict` or dict, optional): Testing config of
+ CornerHead.
+ loss_heatmap (:obj:`ConfigDict` or dict): Config of corner heatmap
+ loss. Defaults to GaussianFocalLoss.
+ loss_embedding (:obj:`ConfigDict` or dict): Config of corner embedding
+ loss. Defaults to AssociativeEmbeddingLoss.
+ loss_offset (:obj:`ConfigDict` or dict): Config of corner offset loss.
+ Defaults to SmoothL1Loss.
+ loss_guiding_shift (:obj:`ConfigDict` or dict): Config of
+ guiding shift loss. Defaults to SmoothL1Loss.
+ loss_centripetal_shift (:obj:`ConfigDict` or dict): Config of
+ centripetal shift loss. Defaults to SmoothL1Loss.
+ init_cfg (:obj:`ConfigDict` or dict, optional): the config to control
+ the initialization.
+ """
+
+ def __init__(self,
+ *args,
+ centripetal_shift_channels: int = 2,
+ guiding_shift_channels: int = 2,
+ feat_adaption_conv_kernel: int = 3,
+ loss_guiding_shift: ConfigType = dict(
+ type='SmoothL1Loss', beta=1.0, loss_weight=0.05),
+ loss_centripetal_shift: ConfigType = dict(
+ type='SmoothL1Loss', beta=1.0, loss_weight=1),
+ init_cfg: OptMultiConfig = None,
+ **kwargs) -> None:
+ assert init_cfg is None, 'To prevent abnormal initialization ' \
+ 'behavior, init_cfg is not allowed to be set'
+ assert centripetal_shift_channels == 2, (
+ 'CentripetalHead only support centripetal_shift_channels == 2')
+ self.centripetal_shift_channels = centripetal_shift_channels
+ assert guiding_shift_channels == 2, (
+ 'CentripetalHead only support guiding_shift_channels == 2')
+ self.guiding_shift_channels = guiding_shift_channels
+ self.feat_adaption_conv_kernel = feat_adaption_conv_kernel
+ super().__init__(*args, init_cfg=init_cfg, **kwargs)
+ self.loss_guiding_shift = MODELS.build(loss_guiding_shift)
+ self.loss_centripetal_shift = MODELS.build(loss_centripetal_shift)
+
+ def _init_centripetal_layers(self) -> None:
+ """Initialize centripetal layers.
+
+ Including feature adaption deform convs (feat_adaption), deform offset
+ prediction convs (dcn_off), guiding shift (guiding_shift) and
+ centripetal shift ( centripetal_shift). Each branch has two parts:
+ prefix `tl_` for top-left and `br_` for bottom-right.
+ """
+ self.tl_feat_adaption = nn.ModuleList()
+ self.br_feat_adaption = nn.ModuleList()
+ self.tl_dcn_offset = nn.ModuleList()
+ self.br_dcn_offset = nn.ModuleList()
+ self.tl_guiding_shift = nn.ModuleList()
+ self.br_guiding_shift = nn.ModuleList()
+ self.tl_centripetal_shift = nn.ModuleList()
+ self.br_centripetal_shift = nn.ModuleList()
+
+ for _ in range(self.num_feat_levels):
+ self.tl_feat_adaption.append(
+ DeformConv2d(self.in_channels, self.in_channels,
+ self.feat_adaption_conv_kernel, 1, 1))
+ self.br_feat_adaption.append(
+ DeformConv2d(self.in_channels, self.in_channels,
+ self.feat_adaption_conv_kernel, 1, 1))
+
+ self.tl_guiding_shift.append(
+ self._make_layers(
+ out_channels=self.guiding_shift_channels,
+ in_channels=self.in_channels))
+ self.br_guiding_shift.append(
+ self._make_layers(
+ out_channels=self.guiding_shift_channels,
+ in_channels=self.in_channels))
+
+ self.tl_dcn_offset.append(
+ ConvModule(
+ self.guiding_shift_channels,
+ self.feat_adaption_conv_kernel**2 *
+ self.guiding_shift_channels,
+ 1,
+ bias=False,
+ act_cfg=None))
+ self.br_dcn_offset.append(
+ ConvModule(
+ self.guiding_shift_channels,
+ self.feat_adaption_conv_kernel**2 *
+ self.guiding_shift_channels,
+ 1,
+ bias=False,
+ act_cfg=None))
+
+ self.tl_centripetal_shift.append(
+ self._make_layers(
+ out_channels=self.centripetal_shift_channels,
+ in_channels=self.in_channels))
+ self.br_centripetal_shift.append(
+ self._make_layers(
+ out_channels=self.centripetal_shift_channels,
+ in_channels=self.in_channels))
+
+ def _init_layers(self) -> None:
+ """Initialize layers for CentripetalHead.
+
+ Including two parts: CornerHead layers and CentripetalHead layers
+ """
+ super()._init_layers() # using _init_layers in CornerHead
+ self._init_centripetal_layers()
+
+ def init_weights(self) -> None:
+ super().init_weights()
+ for i in range(self.num_feat_levels):
+ normal_init(self.tl_feat_adaption[i], std=0.01)
+ normal_init(self.br_feat_adaption[i], std=0.01)
+ normal_init(self.tl_dcn_offset[i].conv, std=0.1)
+ normal_init(self.br_dcn_offset[i].conv, std=0.1)
+ _ = [x.conv.reset_parameters() for x in self.tl_guiding_shift[i]]
+ _ = [x.conv.reset_parameters() for x in self.br_guiding_shift[i]]
+ _ = [
+ x.conv.reset_parameters() for x in self.tl_centripetal_shift[i]
+ ]
+ _ = [
+ x.conv.reset_parameters() for x in self.br_centripetal_shift[i]
+ ]
+
+ def forward_single(self, x: Tensor, lvl_ind: int) -> List[Tensor]:
+ """Forward feature of a single level.
+
+ Args:
+ x (Tensor): Feature of a single level.
+ lvl_ind (int): Level index of current feature.
+
+ Returns:
+ tuple[Tensor]: A tuple of CentripetalHead's output for current
+ feature level. Containing the following Tensors:
+
+ - tl_heat (Tensor): Predicted top-left corner heatmap.
+ - br_heat (Tensor): Predicted bottom-right corner heatmap.
+ - tl_off (Tensor): Predicted top-left offset heatmap.
+ - br_off (Tensor): Predicted bottom-right offset heatmap.
+ - tl_guiding_shift (Tensor): Predicted top-left guiding shift
+ heatmap.
+ - br_guiding_shift (Tensor): Predicted bottom-right guiding
+ shift heatmap.
+ - tl_centripetal_shift (Tensor): Predicted top-left centripetal
+ shift heatmap.
+ - br_centripetal_shift (Tensor): Predicted bottom-right
+ centripetal shift heatmap.
+ """
+ tl_heat, br_heat, _, _, tl_off, br_off, tl_pool, br_pool = super(
+ ).forward_single(
+ x, lvl_ind, return_pool=True)
+
+ tl_guiding_shift = self.tl_guiding_shift[lvl_ind](tl_pool)
+ br_guiding_shift = self.br_guiding_shift[lvl_ind](br_pool)
+
+ tl_dcn_offset = self.tl_dcn_offset[lvl_ind](tl_guiding_shift.detach())
+ br_dcn_offset = self.br_dcn_offset[lvl_ind](br_guiding_shift.detach())
+
+ tl_feat_adaption = self.tl_feat_adaption[lvl_ind](tl_pool,
+ tl_dcn_offset)
+ br_feat_adaption = self.br_feat_adaption[lvl_ind](br_pool,
+ br_dcn_offset)
+
+ tl_centripetal_shift = self.tl_centripetal_shift[lvl_ind](
+ tl_feat_adaption)
+ br_centripetal_shift = self.br_centripetal_shift[lvl_ind](
+ br_feat_adaption)
+
+ result_list = [
+ tl_heat, br_heat, tl_off, br_off, tl_guiding_shift,
+ br_guiding_shift, tl_centripetal_shift, br_centripetal_shift
+ ]
+ return result_list
+
+ def loss_by_feat(
+ self,
+ tl_heats: List[Tensor],
+ br_heats: List[Tensor],
+ tl_offs: List[Tensor],
+ br_offs: List[Tensor],
+ tl_guiding_shifts: List[Tensor],
+ br_guiding_shifts: List[Tensor],
+ tl_centripetal_shifts: List[Tensor],
+ br_centripetal_shifts: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ tl_heats (list[Tensor]): Top-left corner heatmaps for each level
+ with shape (N, num_classes, H, W).
+ br_heats (list[Tensor]): Bottom-right corner heatmaps for each
+ level with shape (N, num_classes, H, W).
+ tl_offs (list[Tensor]): Top-left corner offsets for each level
+ with shape (N, corner_offset_channels, H, W).
+ br_offs (list[Tensor]): Bottom-right corner offsets for each level
+ with shape (N, corner_offset_channels, H, W).
+ tl_guiding_shifts (list[Tensor]): Top-left guiding shifts for each
+ level with shape (N, guiding_shift_channels, H, W).
+ br_guiding_shifts (list[Tensor]): Bottom-right guiding shifts for
+ each level with shape (N, guiding_shift_channels, H, W).
+ tl_centripetal_shifts (list[Tensor]): Top-left centripetal shifts
+ for each level with shape (N, centripetal_shift_channels, H,
+ W).
+ br_centripetal_shifts (list[Tensor]): Bottom-right centripetal
+ shifts for each level with shape (N,
+ centripetal_shift_channels, H, W).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Specify which bounding boxes can be ignored when computing
+ the loss.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components. Containing the
+ following losses:
+
+ - det_loss (list[Tensor]): Corner keypoint losses of all
+ feature levels.
+ - off_loss (list[Tensor]): Corner offset losses of all feature
+ levels.
+ - guiding_loss (list[Tensor]): Guiding shift losses of all
+ feature levels.
+ - centripetal_loss (list[Tensor]): Centripetal shift losses of
+ all feature levels.
+ """
+ gt_bboxes = [
+ gt_instances.bboxes for gt_instances in batch_gt_instances
+ ]
+ gt_labels = [
+ gt_instances.labels for gt_instances in batch_gt_instances
+ ]
+
+ targets = self.get_targets(
+ gt_bboxes,
+ gt_labels,
+ tl_heats[-1].shape,
+ batch_img_metas[0]['batch_input_shape'],
+ with_corner_emb=self.with_corner_emb,
+ with_guiding_shift=True,
+ with_centripetal_shift=True)
+ mlvl_targets = [targets for _ in range(self.num_feat_levels)]
+ [det_losses, off_losses, guiding_losses, centripetal_losses
+ ] = multi_apply(self.loss_by_feat_single, tl_heats, br_heats, tl_offs,
+ br_offs, tl_guiding_shifts, br_guiding_shifts,
+ tl_centripetal_shifts, br_centripetal_shifts,
+ mlvl_targets)
+ loss_dict = dict(
+ det_loss=det_losses,
+ off_loss=off_losses,
+ guiding_loss=guiding_losses,
+ centripetal_loss=centripetal_losses)
+ return loss_dict
+
+ def loss_by_feat_single(self, tl_hmp: Tensor, br_hmp: Tensor,
+ tl_off: Tensor, br_off: Tensor,
+ tl_guiding_shift: Tensor, br_guiding_shift: Tensor,
+ tl_centripetal_shift: Tensor,
+ br_centripetal_shift: Tensor,
+ targets: dict) -> Tuple[Tensor, ...]:
+ """Calculate the loss of a single scale level based on the features
+ extracted by the detection head.
+
+ Args:
+ tl_hmp (Tensor): Top-left corner heatmap for current level with
+ shape (N, num_classes, H, W).
+ br_hmp (Tensor): Bottom-right corner heatmap for current level with
+ shape (N, num_classes, H, W).
+ tl_off (Tensor): Top-left corner offset for current level with
+ shape (N, corner_offset_channels, H, W).
+ br_off (Tensor): Bottom-right corner offset for current level with
+ shape (N, corner_offset_channels, H, W).
+ tl_guiding_shift (Tensor): Top-left guiding shift for current level
+ with shape (N, guiding_shift_channels, H, W).
+ br_guiding_shift (Tensor): Bottom-right guiding shift for current
+ level with shape (N, guiding_shift_channels, H, W).
+ tl_centripetal_shift (Tensor): Top-left centripetal shift for
+ current level with shape (N, centripetal_shift_channels, H, W).
+ br_centripetal_shift (Tensor): Bottom-right centripetal shift for
+ current level with shape (N, centripetal_shift_channels, H, W).
+ targets (dict): Corner target generated by `get_targets`.
+
+ Returns:
+ tuple[torch.Tensor]: Losses of the head's different branches
+ containing the following losses:
+
+ - det_loss (Tensor): Corner keypoint loss.
+ - off_loss (Tensor): Corner offset loss.
+ - guiding_loss (Tensor): Guiding shift loss.
+ - centripetal_loss (Tensor): Centripetal shift loss.
+ """
+ targets['corner_embedding'] = None
+
+ det_loss, _, _, off_loss = super().loss_by_feat_single(
+ tl_hmp, br_hmp, None, None, tl_off, br_off, targets)
+
+ gt_tl_guiding_shift = targets['topleft_guiding_shift']
+ gt_br_guiding_shift = targets['bottomright_guiding_shift']
+ gt_tl_centripetal_shift = targets['topleft_centripetal_shift']
+ gt_br_centripetal_shift = targets['bottomright_centripetal_shift']
+
+ gt_tl_heatmap = targets['topleft_heatmap']
+ gt_br_heatmap = targets['bottomright_heatmap']
+ # We only compute the offset loss at the real corner position.
+ # The value of real corner would be 1 in heatmap ground truth.
+ # The mask is computed in class agnostic mode and its shape is
+ # batch * 1 * width * height.
+ tl_mask = gt_tl_heatmap.eq(1).sum(1).gt(0).unsqueeze(1).type_as(
+ gt_tl_heatmap)
+ br_mask = gt_br_heatmap.eq(1).sum(1).gt(0).unsqueeze(1).type_as(
+ gt_br_heatmap)
+
+ # Guiding shift loss
+ tl_guiding_loss = self.loss_guiding_shift(
+ tl_guiding_shift,
+ gt_tl_guiding_shift,
+ tl_mask,
+ avg_factor=tl_mask.sum())
+ br_guiding_loss = self.loss_guiding_shift(
+ br_guiding_shift,
+ gt_br_guiding_shift,
+ br_mask,
+ avg_factor=br_mask.sum())
+ guiding_loss = (tl_guiding_loss + br_guiding_loss) / 2.0
+ # Centripetal shift loss
+ tl_centripetal_loss = self.loss_centripetal_shift(
+ tl_centripetal_shift,
+ gt_tl_centripetal_shift,
+ tl_mask,
+ avg_factor=tl_mask.sum())
+ br_centripetal_loss = self.loss_centripetal_shift(
+ br_centripetal_shift,
+ gt_br_centripetal_shift,
+ br_mask,
+ avg_factor=br_mask.sum())
+ centripetal_loss = (tl_centripetal_loss + br_centripetal_loss) / 2.0
+
+ return det_loss, off_loss, guiding_loss, centripetal_loss
+
+ def predict_by_feat(self,
+ tl_heats: List[Tensor],
+ br_heats: List[Tensor],
+ tl_offs: List[Tensor],
+ br_offs: List[Tensor],
+ tl_guiding_shifts: List[Tensor],
+ br_guiding_shifts: List[Tensor],
+ tl_centripetal_shifts: List[Tensor],
+ br_centripetal_shifts: List[Tensor],
+ batch_img_metas: Optional[List[dict]] = None,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ bbox results.
+
+ Args:
+ tl_heats (list[Tensor]): Top-left corner heatmaps for each level
+ with shape (N, num_classes, H, W).
+ br_heats (list[Tensor]): Bottom-right corner heatmaps for each
+ level with shape (N, num_classes, H, W).
+ tl_offs (list[Tensor]): Top-left corner offsets for each level
+ with shape (N, corner_offset_channels, H, W).
+ br_offs (list[Tensor]): Bottom-right corner offsets for each level
+ with shape (N, corner_offset_channels, H, W).
+ tl_guiding_shifts (list[Tensor]): Top-left guiding shifts for each
+ level with shape (N, guiding_shift_channels, H, W). Useless in
+ this function, we keep this arg because it's the raw output
+ from CentripetalHead.
+ br_guiding_shifts (list[Tensor]): Bottom-right guiding shifts for
+ each level with shape (N, guiding_shift_channels, H, W).
+ Useless in this function, we keep this arg because it's the
+ raw output from CentripetalHead.
+ tl_centripetal_shifts (list[Tensor]): Top-left centripetal shifts
+ for each level with shape (N, centripetal_shift_channels, H,
+ W).
+ br_centripetal_shifts (list[Tensor]): Bottom-right centripetal
+ shifts for each level with shape (N,
+ centripetal_shift_channels, H, W).
+ batch_img_metas (list[dict], optional): Batch image meta info.
+ Defaults to None.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ list[:obj:`InstanceData`]: Object detection results of each image
+ after the post process. Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ assert tl_heats[-1].shape[0] == br_heats[-1].shape[0] == len(
+ batch_img_metas)
+ result_list = []
+ for img_id in range(len(batch_img_metas)):
+ result_list.append(
+ self._predict_by_feat_single(
+ tl_heats[-1][img_id:img_id + 1, :],
+ br_heats[-1][img_id:img_id + 1, :],
+ tl_offs[-1][img_id:img_id + 1, :],
+ br_offs[-1][img_id:img_id + 1, :],
+ batch_img_metas[img_id],
+ tl_emb=None,
+ br_emb=None,
+ tl_centripetal_shift=tl_centripetal_shifts[-1][
+ img_id:img_id + 1, :],
+ br_centripetal_shift=br_centripetal_shifts[-1][
+ img_id:img_id + 1, :],
+ rescale=rescale,
+ with_nms=with_nms))
+
+ return result_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/condinst_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/condinst_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..35a25e6339a8161314cb0523e7181f9d400023ac
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/condinst_head.py
@@ -0,0 +1,1226 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from typing import Dict, List, Optional, Tuple
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule, Scale
+from mmengine.config import ConfigDict
+from mmengine.model import BaseModule, kaiming_init
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures.bbox import cat_boxes
+from mmdet.utils import (ConfigType, InstanceList, MultiConfig, OptConfigType,
+ OptInstanceList, reduce_mean)
+from ..task_modules.prior_generators import MlvlPointGenerator
+from ..utils import (aligned_bilinear, filter_scores_and_topk, multi_apply,
+ relative_coordinate_maps, select_single_mlvl)
+from ..utils.misc import empty_instances
+from .base_mask_head import BaseMaskHead
+from .fcos_head import FCOSHead
+
+INF = 1e8
+
+
+@MODELS.register_module()
+class CondInstBboxHead(FCOSHead):
+ """CondInst box head used in https://arxiv.org/abs/1904.02689.
+
+ Note that CondInst Bbox Head is a extension of FCOS head.
+ Two differences are described as follows:
+
+ 1. CondInst box head predicts a set of params for each instance.
+ 2. CondInst box head return the pos_gt_inds and pos_inds.
+
+ Args:
+ num_params (int): Number of params for instance segmentation.
+ """
+
+ def __init__(self, *args, num_params: int = 169, **kwargs) -> None:
+ self.num_params = num_params
+ super().__init__(*args, **kwargs)
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ super()._init_layers()
+ self.controller = nn.Conv2d(
+ self.feat_channels, self.num_params, 3, padding=1)
+
+ def forward_single(self, x: Tensor, scale: Scale,
+ stride: int) -> Tuple[Tensor, Tensor, Tensor, Tensor]:
+ """Forward features of a single scale level.
+
+ Args:
+ x (Tensor): FPN feature maps of the specified stride.
+ scale (:obj:`mmcv.cnn.Scale`): Learnable scale module to resize
+ the bbox prediction.
+ stride (int): The corresponding stride for feature maps, only
+ used to normalize the bbox prediction when self.norm_on_bbox
+ is True.
+
+ Returns:
+ tuple: scores for each class, bbox predictions, centerness
+ predictions and param predictions of input feature maps.
+ """
+ cls_score, bbox_pred, cls_feat, reg_feat = \
+ super(FCOSHead, self).forward_single(x)
+ if self.centerness_on_reg:
+ centerness = self.conv_centerness(reg_feat)
+ else:
+ centerness = self.conv_centerness(cls_feat)
+ # scale the bbox_pred of different level
+ # float to avoid overflow when enabling FP16
+ bbox_pred = scale(bbox_pred).float()
+ if self.norm_on_bbox:
+ # bbox_pred needed for gradient computation has been modified
+ # by F.relu(bbox_pred) when run with PyTorch 1.10. So replace
+ # F.relu(bbox_pred) with bbox_pred.clamp(min=0)
+ bbox_pred = bbox_pred.clamp(min=0)
+ if not self.training:
+ bbox_pred *= stride
+ else:
+ bbox_pred = bbox_pred.exp()
+ param_pred = self.controller(reg_feat)
+ return cls_score, bbox_pred, centerness, param_pred
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ centernesses: List[Tensor],
+ param_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None
+ ) -> Dict[str, Tensor]:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level,
+ each is a 4D-tensor, the channel number is
+ num_points * num_classes.
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level, each is a 4D-tensor, the channel number is
+ num_points * 4.
+ centernesses (list[Tensor]): centerness for each scale level, each
+ is a 4D-tensor, the channel number is num_points * 1.
+ param_preds (List[Tensor]): param_pred for each scale level, each
+ is a 4D-tensor, the channel number is num_params.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ assert len(cls_scores) == len(bbox_preds) == len(centernesses)
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ # Need stride for rel coord compute
+ all_level_points_strides = self.prior_generator.grid_priors(
+ featmap_sizes,
+ dtype=bbox_preds[0].dtype,
+ device=bbox_preds[0].device,
+ with_stride=True)
+ all_level_points = [i[:, :2] for i in all_level_points_strides]
+ all_level_strides = [i[:, 2] for i in all_level_points_strides]
+ labels, bbox_targets, pos_inds_list, pos_gt_inds_list = \
+ self.get_targets(all_level_points, batch_gt_instances)
+
+ num_imgs = cls_scores[0].size(0)
+ # flatten cls_scores, bbox_preds and centerness
+ flatten_cls_scores = [
+ cls_score.permute(0, 2, 3, 1).reshape(-1, self.cls_out_channels)
+ for cls_score in cls_scores
+ ]
+ flatten_bbox_preds = [
+ bbox_pred.permute(0, 2, 3, 1).reshape(-1, 4)
+ for bbox_pred in bbox_preds
+ ]
+ flatten_centerness = [
+ centerness.permute(0, 2, 3, 1).reshape(-1)
+ for centerness in centernesses
+ ]
+ flatten_cls_scores = torch.cat(flatten_cls_scores)
+ flatten_bbox_preds = torch.cat(flatten_bbox_preds)
+ flatten_centerness = torch.cat(flatten_centerness)
+ flatten_labels = torch.cat(labels)
+ flatten_bbox_targets = torch.cat(bbox_targets)
+ # repeat points to align with bbox_preds
+ flatten_points = torch.cat(
+ [points.repeat(num_imgs, 1) for points in all_level_points])
+
+ # FG cat_id: [0, num_classes -1], BG cat_id: num_classes
+ bg_class_ind = self.num_classes
+ pos_inds = ((flatten_labels >= 0)
+ & (flatten_labels < bg_class_ind)).nonzero().reshape(-1)
+ num_pos = torch.tensor(
+ len(pos_inds), dtype=torch.float, device=bbox_preds[0].device)
+ num_pos = max(reduce_mean(num_pos), 1.0)
+ loss_cls = self.loss_cls(
+ flatten_cls_scores, flatten_labels, avg_factor=num_pos)
+
+ pos_bbox_preds = flatten_bbox_preds[pos_inds]
+ pos_centerness = flatten_centerness[pos_inds]
+ pos_bbox_targets = flatten_bbox_targets[pos_inds]
+ pos_centerness_targets = self.centerness_target(pos_bbox_targets)
+ # centerness weighted iou loss
+ centerness_denorm = max(
+ reduce_mean(pos_centerness_targets.sum().detach()), 1e-6)
+
+ if len(pos_inds) > 0:
+ pos_points = flatten_points[pos_inds]
+ pos_decoded_bbox_preds = self.bbox_coder.decode(
+ pos_points, pos_bbox_preds)
+ pos_decoded_target_preds = self.bbox_coder.decode(
+ pos_points, pos_bbox_targets)
+ loss_bbox = self.loss_bbox(
+ pos_decoded_bbox_preds,
+ pos_decoded_target_preds,
+ weight=pos_centerness_targets,
+ avg_factor=centerness_denorm)
+ loss_centerness = self.loss_centerness(
+ pos_centerness, pos_centerness_targets, avg_factor=num_pos)
+ else:
+ loss_bbox = pos_bbox_preds.sum()
+ loss_centerness = pos_centerness.sum()
+
+ self._raw_positive_infos.update(cls_scores=cls_scores)
+ self._raw_positive_infos.update(centernesses=centernesses)
+ self._raw_positive_infos.update(param_preds=param_preds)
+ self._raw_positive_infos.update(all_level_points=all_level_points)
+ self._raw_positive_infos.update(all_level_strides=all_level_strides)
+ self._raw_positive_infos.update(pos_gt_inds_list=pos_gt_inds_list)
+ self._raw_positive_infos.update(pos_inds_list=pos_inds_list)
+
+ return dict(
+ loss_cls=loss_cls,
+ loss_bbox=loss_bbox,
+ loss_centerness=loss_centerness)
+
+ def get_targets(
+ self, points: List[Tensor], batch_gt_instances: InstanceList
+ ) -> Tuple[List[Tensor], List[Tensor], List[Tensor], List[Tensor]]:
+ """Compute regression, classification and centerness targets for points
+ in multiple images.
+
+ Args:
+ points (list[Tensor]): Points of each fpn level, each has shape
+ (num_points, 2).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+
+ Returns:
+ tuple: Targets of each level.
+
+ - concat_lvl_labels (list[Tensor]): Labels of each level.
+ - concat_lvl_bbox_targets (list[Tensor]): BBox targets of each \
+ level.
+ - pos_inds_list (list[Tensor]): pos_inds of each image.
+ - pos_gt_inds_list (List[Tensor]): pos_gt_inds of each image.
+ """
+ assert len(points) == len(self.regress_ranges)
+ num_levels = len(points)
+ # expand regress ranges to align with points
+ expanded_regress_ranges = [
+ points[i].new_tensor(self.regress_ranges[i])[None].expand_as(
+ points[i]) for i in range(num_levels)
+ ]
+ # concat all levels points and regress ranges
+ concat_regress_ranges = torch.cat(expanded_regress_ranges, dim=0)
+ concat_points = torch.cat(points, dim=0)
+
+ # the number of points per img, per lvl
+ num_points = [center.size(0) for center in points]
+
+ # get labels and bbox_targets of each image
+ labels_list, bbox_targets_list, pos_inds_list, pos_gt_inds_list = \
+ multi_apply(
+ self._get_targets_single,
+ batch_gt_instances,
+ points=concat_points,
+ regress_ranges=concat_regress_ranges,
+ num_points_per_lvl=num_points)
+
+ # split to per img, per level
+ labels_list = [labels.split(num_points, 0) for labels in labels_list]
+ bbox_targets_list = [
+ bbox_targets.split(num_points, 0)
+ for bbox_targets in bbox_targets_list
+ ]
+
+ # concat per level image
+ concat_lvl_labels = []
+ concat_lvl_bbox_targets = []
+ for i in range(num_levels):
+ concat_lvl_labels.append(
+ torch.cat([labels[i] for labels in labels_list]))
+ bbox_targets = torch.cat(
+ [bbox_targets[i] for bbox_targets in bbox_targets_list])
+ if self.norm_on_bbox:
+ bbox_targets = bbox_targets / self.strides[i]
+ concat_lvl_bbox_targets.append(bbox_targets)
+ return (concat_lvl_labels, concat_lvl_bbox_targets, pos_inds_list,
+ pos_gt_inds_list)
+
+ def _get_targets_single(
+ self, gt_instances: InstanceData, points: Tensor,
+ regress_ranges: Tensor, num_points_per_lvl: List[int]
+ ) -> Tuple[Tensor, Tensor, Tensor, Tensor]:
+ """Compute regression and classification targets for a single image."""
+ num_points = points.size(0)
+ num_gts = len(gt_instances)
+ gt_bboxes = gt_instances.bboxes
+ gt_labels = gt_instances.labels
+ gt_masks = gt_instances.get('masks', None)
+
+ if num_gts == 0:
+ return gt_labels.new_full((num_points,), self.num_classes), \
+ gt_bboxes.new_zeros((num_points, 4)), \
+ gt_bboxes.new_zeros((0,), dtype=torch.int64), \
+ gt_bboxes.new_zeros((0,), dtype=torch.int64)
+
+ areas = (gt_bboxes[:, 2] - gt_bboxes[:, 0]) * (
+ gt_bboxes[:, 3] - gt_bboxes[:, 1])
+ # TODO: figure out why these two are different
+ # areas = areas[None].expand(num_points, num_gts)
+ areas = areas[None].repeat(num_points, 1)
+ regress_ranges = regress_ranges[:, None, :].expand(
+ num_points, num_gts, 2)
+ gt_bboxes = gt_bboxes[None].expand(num_points, num_gts, 4)
+ xs, ys = points[:, 0], points[:, 1]
+ xs = xs[:, None].expand(num_points, num_gts)
+ ys = ys[:, None].expand(num_points, num_gts)
+
+ left = xs - gt_bboxes[..., 0]
+ right = gt_bboxes[..., 2] - xs
+ top = ys - gt_bboxes[..., 1]
+ bottom = gt_bboxes[..., 3] - ys
+ bbox_targets = torch.stack((left, top, right, bottom), -1)
+
+ if self.center_sampling:
+ # condition1: inside a `center bbox`
+ radius = self.center_sample_radius
+ # if gt_mask not None, use gt mask's centroid to determine
+ # the center region rather than gt_bbox center
+ if gt_masks is None:
+ center_xs = (gt_bboxes[..., 0] + gt_bboxes[..., 2]) / 2
+ center_ys = (gt_bboxes[..., 1] + gt_bboxes[..., 3]) / 2
+ else:
+ h, w = gt_masks.height, gt_masks.width
+ masks = gt_masks.to_tensor(
+ dtype=torch.bool, device=gt_bboxes.device)
+ yys = torch.arange(
+ 0, h, dtype=torch.float32, device=masks.device)
+ xxs = torch.arange(
+ 0, w, dtype=torch.float32, device=masks.device)
+ # m00/m10/m01 represent the moments of a contour
+ # centroid is computed by m00/m10 and m00/m01
+ m00 = masks.sum(dim=-1).sum(dim=-1).clamp(min=1e-6)
+ m10 = (masks * xxs).sum(dim=-1).sum(dim=-1)
+ m01 = (masks * yys[:, None]).sum(dim=-1).sum(dim=-1)
+ center_xs = m10 / m00
+ center_ys = m01 / m00
+
+ center_xs = center_xs[None].expand(num_points, num_gts)
+ center_ys = center_ys[None].expand(num_points, num_gts)
+ center_gts = torch.zeros_like(gt_bboxes)
+ stride = center_xs.new_zeros(center_xs.shape)
+
+ # project the points on current lvl back to the `original` sizes
+ lvl_begin = 0
+ for lvl_idx, num_points_lvl in enumerate(num_points_per_lvl):
+ lvl_end = lvl_begin + num_points_lvl
+ stride[lvl_begin:lvl_end] = self.strides[lvl_idx] * radius
+ lvl_begin = lvl_end
+
+ x_mins = center_xs - stride
+ y_mins = center_ys - stride
+ x_maxs = center_xs + stride
+ y_maxs = center_ys + stride
+ center_gts[..., 0] = torch.where(x_mins > gt_bboxes[..., 0],
+ x_mins, gt_bboxes[..., 0])
+ center_gts[..., 1] = torch.where(y_mins > gt_bboxes[..., 1],
+ y_mins, gt_bboxes[..., 1])
+ center_gts[..., 2] = torch.where(x_maxs > gt_bboxes[..., 2],
+ gt_bboxes[..., 2], x_maxs)
+ center_gts[..., 3] = torch.where(y_maxs > gt_bboxes[..., 3],
+ gt_bboxes[..., 3], y_maxs)
+
+ cb_dist_left = xs - center_gts[..., 0]
+ cb_dist_right = center_gts[..., 2] - xs
+ cb_dist_top = ys - center_gts[..., 1]
+ cb_dist_bottom = center_gts[..., 3] - ys
+ center_bbox = torch.stack(
+ (cb_dist_left, cb_dist_top, cb_dist_right, cb_dist_bottom), -1)
+ inside_gt_bbox_mask = center_bbox.min(-1)[0] > 0
+ else:
+ # condition1: inside a gt bbox
+ inside_gt_bbox_mask = bbox_targets.min(-1)[0] > 0
+
+ # condition2: limit the regression range for each location
+ max_regress_distance = bbox_targets.max(-1)[0]
+ inside_regress_range = (
+ (max_regress_distance >= regress_ranges[..., 0])
+ & (max_regress_distance <= regress_ranges[..., 1]))
+
+ # if there are still more than one objects for a location,
+ # we choose the one with minimal area
+ areas[inside_gt_bbox_mask == 0] = INF
+ areas[inside_regress_range == 0] = INF
+ min_area, min_area_inds = areas.min(dim=1)
+
+ labels = gt_labels[min_area_inds]
+ labels[min_area == INF] = self.num_classes # set as BG
+ bbox_targets = bbox_targets[range(num_points), min_area_inds]
+
+ # return pos_inds & pos_gt_inds
+ bg_class_ind = self.num_classes
+ pos_inds = ((labels >= 0)
+ & (labels < bg_class_ind)).nonzero().reshape(-1)
+ pos_gt_inds = min_area_inds[labels < self.num_classes]
+ return labels, bbox_targets, pos_inds, pos_gt_inds
+
+ def get_positive_infos(self) -> InstanceList:
+ """Get positive information from sampling results.
+
+ Returns:
+ list[:obj:`InstanceData`]: Positive information of each image,
+ usually including positive bboxes, positive labels, positive
+ priors, etc.
+ """
+ assert len(self._raw_positive_infos) > 0
+
+ pos_gt_inds_list = self._raw_positive_infos['pos_gt_inds_list']
+ pos_inds_list = self._raw_positive_infos['pos_inds_list']
+ num_imgs = len(pos_gt_inds_list)
+
+ cls_score_list = []
+ centerness_list = []
+ param_pred_list = []
+ point_list = []
+ stride_list = []
+ for cls_score_per_lvl, centerness_per_lvl, param_pred_per_lvl,\
+ point_per_lvl, stride_per_lvl in \
+ zip(self._raw_positive_infos['cls_scores'],
+ self._raw_positive_infos['centernesses'],
+ self._raw_positive_infos['param_preds'],
+ self._raw_positive_infos['all_level_points'],
+ self._raw_positive_infos['all_level_strides']):
+ cls_score_per_lvl = \
+ cls_score_per_lvl.permute(
+ 0, 2, 3, 1).reshape(num_imgs, -1, self.num_classes)
+ centerness_per_lvl = \
+ centerness_per_lvl.permute(
+ 0, 2, 3, 1).reshape(num_imgs, -1, 1)
+ param_pred_per_lvl = \
+ param_pred_per_lvl.permute(
+ 0, 2, 3, 1).reshape(num_imgs, -1, self.num_params)
+ point_per_lvl = point_per_lvl.unsqueeze(0).repeat(num_imgs, 1, 1)
+ stride_per_lvl = stride_per_lvl.unsqueeze(0).repeat(num_imgs, 1)
+
+ cls_score_list.append(cls_score_per_lvl)
+ centerness_list.append(centerness_per_lvl)
+ param_pred_list.append(param_pred_per_lvl)
+ point_list.append(point_per_lvl)
+ stride_list.append(stride_per_lvl)
+ cls_scores = torch.cat(cls_score_list, dim=1)
+ centernesses = torch.cat(centerness_list, dim=1)
+ param_preds = torch.cat(param_pred_list, dim=1)
+ all_points = torch.cat(point_list, dim=1)
+ all_strides = torch.cat(stride_list, dim=1)
+
+ positive_infos = []
+ for i, (pos_gt_inds,
+ pos_inds) in enumerate(zip(pos_gt_inds_list, pos_inds_list)):
+ pos_info = InstanceData()
+ pos_info.points = all_points[i][pos_inds]
+ pos_info.strides = all_strides[i][pos_inds]
+ pos_info.scores = cls_scores[i][pos_inds]
+ pos_info.centernesses = centernesses[i][pos_inds]
+ pos_info.param_preds = param_preds[i][pos_inds]
+ pos_info.pos_assigned_gt_inds = pos_gt_inds
+ pos_info.pos_inds = pos_inds
+ positive_infos.append(pos_info)
+ return positive_infos
+
+ def predict_by_feat(self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ score_factors: Optional[List[Tensor]] = None,
+ param_preds: Optional[List[Tensor]] = None,
+ batch_img_metas: Optional[List[dict]] = None,
+ cfg: Optional[ConfigDict] = None,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ bbox results.
+
+ Note: When score_factors is not None, the cls_scores are
+ usually multiplied by it then obtain the real score used in NMS,
+ such as CenterNess in FCOS, IoU branch in ATSS.
+
+ Args:
+ cls_scores (list[Tensor]): Classification scores for all
+ scale levels, each is a 4D-tensor, has shape
+ (batch_size, num_priors * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas for all
+ scale levels, each is a 4D-tensor, has shape
+ (batch_size, num_priors * 4, H, W).
+ score_factors (list[Tensor], optional): Score factor for
+ all scale level, each is a 4D-tensor, has shape
+ (batch_size, num_priors * 1, H, W). Defaults to None.
+ param_preds (list[Tensor], optional): Params for all scale
+ level, each is a 4D-tensor, has shape
+ (batch_size, num_priors * num_params, H, W)
+ batch_img_metas (list[dict], Optional): Batch image meta info.
+ Defaults to None.
+ cfg (ConfigDict, optional): Test / postprocessing
+ configuration, if None, test_cfg would be used.
+ Defaults to None.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ list[:obj:`InstanceData`]: Object detection results of each image
+ after the post process. Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ assert len(cls_scores) == len(bbox_preds)
+
+ if score_factors is None:
+ # e.g. Retina, FreeAnchor, Foveabox, etc.
+ with_score_factors = False
+ else:
+ # e.g. FCOS, PAA, ATSS, AutoAssign, etc.
+ with_score_factors = True
+ assert len(cls_scores) == len(score_factors)
+
+ num_levels = len(cls_scores)
+
+ featmap_sizes = [cls_scores[i].shape[-2:] for i in range(num_levels)]
+ all_level_points_strides = self.prior_generator.grid_priors(
+ featmap_sizes,
+ dtype=bbox_preds[0].dtype,
+ device=bbox_preds[0].device,
+ with_stride=True)
+ all_level_points = [i[:, :2] for i in all_level_points_strides]
+ all_level_strides = [i[:, 2] for i in all_level_points_strides]
+
+ result_list = []
+
+ for img_id in range(len(batch_img_metas)):
+ img_meta = batch_img_metas[img_id]
+ cls_score_list = select_single_mlvl(
+ cls_scores, img_id, detach=True)
+ bbox_pred_list = select_single_mlvl(
+ bbox_preds, img_id, detach=True)
+ if with_score_factors:
+ score_factor_list = select_single_mlvl(
+ score_factors, img_id, detach=True)
+ else:
+ score_factor_list = [None for _ in range(num_levels)]
+ param_pred_list = select_single_mlvl(
+ param_preds, img_id, detach=True)
+
+ results = self._predict_by_feat_single(
+ cls_score_list=cls_score_list,
+ bbox_pred_list=bbox_pred_list,
+ score_factor_list=score_factor_list,
+ param_pred_list=param_pred_list,
+ mlvl_points=all_level_points,
+ mlvl_strides=all_level_strides,
+ img_meta=img_meta,
+ cfg=cfg,
+ rescale=rescale,
+ with_nms=with_nms)
+ result_list.append(results)
+ return result_list
+
+ def _predict_by_feat_single(self,
+ cls_score_list: List[Tensor],
+ bbox_pred_list: List[Tensor],
+ score_factor_list: List[Tensor],
+ param_pred_list: List[Tensor],
+ mlvl_points: List[Tensor],
+ mlvl_strides: List[Tensor],
+ img_meta: dict,
+ cfg: ConfigDict,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results.
+
+ Args:
+ cls_score_list (list[Tensor]): Box scores from all scale
+ levels of a single image, each item has shape
+ (num_priors * num_classes, H, W).
+ bbox_pred_list (list[Tensor]): Box energies / deltas from
+ all scale levels of a single image, each item has shape
+ (num_priors * 4, H, W).
+ score_factor_list (list[Tensor]): Score factor from all scale
+ levels of a single image, each item has shape
+ (num_priors * 1, H, W).
+ param_pred_list (List[Tensor]): Param predition from all scale
+ levels of a single image, each item has shape
+ (num_priors * num_params, H, W).
+ mlvl_points (list[Tensor]): Each element in the list is
+ the priors of a single level in feature pyramid.
+ It has shape (num_priors, 2)
+ mlvl_strides (List[Tensor]): Each element in the list is
+ the stride of a single level in feature pyramid.
+ It has shape (num_priors, 1)
+ img_meta (dict): Image meta info.
+ cfg (mmengine.Config): Test / postprocessing configuration,
+ if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ if score_factor_list[0] is None:
+ # e.g. Retina, FreeAnchor, etc.
+ with_score_factors = False
+ else:
+ # e.g. FCOS, PAA, ATSS, etc.
+ with_score_factors = True
+
+ cfg = self.test_cfg if cfg is None else cfg
+ cfg = copy.deepcopy(cfg)
+ img_shape = img_meta['img_shape']
+ nms_pre = cfg.get('nms_pre', -1)
+
+ mlvl_bbox_preds = []
+ mlvl_param_preds = []
+ mlvl_valid_points = []
+ mlvl_valid_strides = []
+ mlvl_scores = []
+ mlvl_labels = []
+ if with_score_factors:
+ mlvl_score_factors = []
+ else:
+ mlvl_score_factors = None
+ for level_idx, (cls_score, bbox_pred, score_factor,
+ param_pred, points, strides) in \
+ enumerate(zip(cls_score_list, bbox_pred_list,
+ score_factor_list, param_pred_list,
+ mlvl_points, mlvl_strides)):
+
+ assert cls_score.size()[-2:] == bbox_pred.size()[-2:]
+
+ dim = self.bbox_coder.encode_size
+ bbox_pred = bbox_pred.permute(1, 2, 0).reshape(-1, dim)
+ if with_score_factors:
+ score_factor = score_factor.permute(1, 2,
+ 0).reshape(-1).sigmoid()
+ cls_score = cls_score.permute(1, 2,
+ 0).reshape(-1, self.cls_out_channels)
+ if self.use_sigmoid_cls:
+ scores = cls_score.sigmoid()
+ else:
+ # remind that we set FG labels to [0, num_class-1]
+ # since mmdet v2.0
+ # BG cat_id: num_class
+ scores = cls_score.softmax(-1)[:, :-1]
+
+ param_pred = param_pred.permute(1, 2,
+ 0).reshape(-1, self.num_params)
+
+ # After https://github.com/open-mmlab/mmdetection/pull/6268/,
+ # this operation keeps fewer bboxes under the same `nms_pre`.
+ # There is no difference in performance for most models. If you
+ # find a slight drop in performance, you can set a larger
+ # `nms_pre` than before.
+ score_thr = cfg.get('score_thr', 0)
+
+ results = filter_scores_and_topk(
+ scores, score_thr, nms_pre,
+ dict(
+ bbox_pred=bbox_pred,
+ param_pred=param_pred,
+ points=points,
+ strides=strides))
+ scores, labels, keep_idxs, filtered_results = results
+
+ bbox_pred = filtered_results['bbox_pred']
+ param_pred = filtered_results['param_pred']
+ points = filtered_results['points']
+ strides = filtered_results['strides']
+
+ if with_score_factors:
+ score_factor = score_factor[keep_idxs]
+
+ mlvl_bbox_preds.append(bbox_pred)
+ mlvl_param_preds.append(param_pred)
+ mlvl_valid_points.append(points)
+ mlvl_valid_strides.append(strides)
+ mlvl_scores.append(scores)
+ mlvl_labels.append(labels)
+
+ if with_score_factors:
+ mlvl_score_factors.append(score_factor)
+
+ bbox_pred = torch.cat(mlvl_bbox_preds)
+ priors = cat_boxes(mlvl_valid_points)
+ bboxes = self.bbox_coder.decode(priors, bbox_pred, max_shape=img_shape)
+
+ results = InstanceData()
+ results.bboxes = bboxes
+ results.scores = torch.cat(mlvl_scores)
+ results.labels = torch.cat(mlvl_labels)
+ results.param_preds = torch.cat(mlvl_param_preds)
+ results.points = torch.cat(mlvl_valid_points)
+ results.strides = torch.cat(mlvl_valid_strides)
+ if with_score_factors:
+ results.score_factors = torch.cat(mlvl_score_factors)
+
+ return self._bbox_post_process(
+ results=results,
+ cfg=cfg,
+ rescale=rescale,
+ with_nms=with_nms,
+ img_meta=img_meta)
+
+
+class MaskFeatModule(BaseModule):
+ """CondInst mask feature map branch used in \
+ https://arxiv.org/abs/1904.02689.
+
+ Args:
+ in_channels (int): Number of channels in the input feature map.
+ feat_channels (int): Number of hidden channels of the mask feature
+ map branch.
+ start_level (int): The starting feature map level from RPN that
+ will be used to predict the mask feature map.
+ end_level (int): The ending feature map level from rpn that
+ will be used to predict the mask feature map.
+ out_channels (int): Number of output channels of the mask feature
+ map branch. This is the channel count of the mask
+ feature map that to be dynamically convolved with the predicted
+ kernel.
+ mask_stride (int): Downsample factor of the mask feature map output.
+ Defaults to 4.
+ num_stacked_convs (int): Number of convs in mask feature branch.
+ conv_cfg (dict): Config dict for convolution layer. Default: None.
+ norm_cfg (dict): Config dict for normalization layer. Default: None.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ """
+
+ def __init__(self,
+ in_channels: int,
+ feat_channels: int,
+ start_level: int,
+ end_level: int,
+ out_channels: int,
+ mask_stride: int = 4,
+ num_stacked_convs: int = 4,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: OptConfigType = None,
+ init_cfg: MultiConfig = [
+ dict(type='Normal', layer='Conv2d', std=0.01)
+ ],
+ **kwargs) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.in_channels = in_channels
+ self.feat_channels = feat_channels
+ self.start_level = start_level
+ self.end_level = end_level
+ self.mask_stride = mask_stride
+ self.num_stacked_convs = num_stacked_convs
+ assert start_level >= 0 and end_level >= start_level
+ self.out_channels = out_channels
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self._init_layers()
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self.convs_all_levels = nn.ModuleList()
+ for i in range(self.start_level, self.end_level + 1):
+ convs_per_level = nn.Sequential()
+ convs_per_level.add_module(
+ f'conv{i}',
+ ConvModule(
+ self.in_channels,
+ self.feat_channels,
+ 3,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ inplace=False,
+ bias=False))
+ self.convs_all_levels.append(convs_per_level)
+
+ conv_branch = []
+ for _ in range(self.num_stacked_convs):
+ conv_branch.append(
+ ConvModule(
+ self.feat_channels,
+ self.feat_channels,
+ 3,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ bias=False))
+ self.conv_branch = nn.Sequential(*conv_branch)
+
+ self.conv_pred = nn.Conv2d(
+ self.feat_channels, self.out_channels, 1, stride=1)
+
+ def init_weights(self) -> None:
+ """Initialize weights of the head."""
+ super().init_weights()
+ kaiming_init(self.convs_all_levels, a=1, distribution='uniform')
+ kaiming_init(self.conv_branch, a=1, distribution='uniform')
+ kaiming_init(self.conv_pred, a=1, distribution='uniform')
+
+ def forward(self, x: Tuple[Tensor]) -> Tensor:
+ """Forward features from the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ Tensor: The predicted mask feature map.
+ """
+ inputs = x[self.start_level:self.end_level + 1]
+ assert len(inputs) == (self.end_level - self.start_level + 1)
+ feature_add_all_level = self.convs_all_levels[0](inputs[0])
+ target_h, target_w = feature_add_all_level.size()[2:]
+ for i in range(1, len(inputs)):
+ input_p = inputs[i]
+ x_p = self.convs_all_levels[i](input_p)
+ h, w = x_p.size()[2:]
+ factor_h = target_h // h
+ factor_w = target_w // w
+ assert factor_h == factor_w
+ feature_per_level = aligned_bilinear(x_p, factor_h)
+ feature_add_all_level = feature_add_all_level + \
+ feature_per_level
+
+ feature_add_all_level = self.conv_branch(feature_add_all_level)
+ feature_pred = self.conv_pred(feature_add_all_level)
+ return feature_pred
+
+
+@MODELS.register_module()
+class CondInstMaskHead(BaseMaskHead):
+ """CondInst mask head used in https://arxiv.org/abs/1904.02689.
+
+ This head outputs the mask for CondInst.
+
+ Args:
+ mask_feature_head (dict): Config of CondInstMaskFeatHead.
+ num_layers (int): Number of dynamic conv layers.
+ feat_channels (int): Number of channels in the dynamic conv.
+ mask_out_stride (int): The stride of the mask feat.
+ size_of_interest (int): The size of the region used in rel coord.
+ max_masks_to_train (int): Maximum number of masks to train for
+ each image.
+ loss_segm (:obj:`ConfigDict` or dict, optional): Config of
+ segmentation loss.
+ train_cfg (:obj:`ConfigDict` or dict, optional): Training config
+ of head.
+ test_cfg (:obj:`ConfigDict` or dict, optional): Testing config of
+ head.
+ """
+
+ def __init__(self,
+ mask_feature_head: ConfigType,
+ num_layers: int = 3,
+ feat_channels: int = 8,
+ mask_out_stride: int = 4,
+ size_of_interest: int = 8,
+ max_masks_to_train: int = -1,
+ topk_masks_per_img: int = -1,
+ loss_mask: ConfigType = None,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None) -> None:
+ super().__init__()
+ self.mask_feature_head = MaskFeatModule(**mask_feature_head)
+ self.mask_feat_stride = self.mask_feature_head.mask_stride
+ self.in_channels = self.mask_feature_head.out_channels
+ self.num_layers = num_layers
+ self.feat_channels = feat_channels
+ self.size_of_interest = size_of_interest
+ self.mask_out_stride = mask_out_stride
+ self.max_masks_to_train = max_masks_to_train
+ self.topk_masks_per_img = topk_masks_per_img
+ self.prior_generator = MlvlPointGenerator([self.mask_feat_stride])
+
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+ self.loss_mask = MODELS.build(loss_mask)
+ self._init_layers()
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ weight_nums, bias_nums = [], []
+ for i in range(self.num_layers):
+ if i == 0:
+ weight_nums.append((self.in_channels + 2) * self.feat_channels)
+ bias_nums.append(self.feat_channels)
+ elif i == self.num_layers - 1:
+ weight_nums.append(self.feat_channels * 1)
+ bias_nums.append(1)
+ else:
+ weight_nums.append(self.feat_channels * self.feat_channels)
+ bias_nums.append(self.feat_channels)
+
+ self.weight_nums = weight_nums
+ self.bias_nums = bias_nums
+ self.num_params = sum(weight_nums) + sum(bias_nums)
+
+ def parse_dynamic_params(
+ self, params: Tensor) -> Tuple[List[Tensor], List[Tensor]]:
+ """parse the dynamic params for dynamic conv."""
+ num_insts = params.size(0)
+ params_splits = list(
+ torch.split_with_sizes(
+ params, self.weight_nums + self.bias_nums, dim=1))
+ weight_splits = params_splits[:self.num_layers]
+ bias_splits = params_splits[self.num_layers:]
+ for i in range(self.num_layers):
+ if i < self.num_layers - 1:
+ weight_splits[i] = weight_splits[i].reshape(
+ num_insts * self.in_channels, -1, 1, 1)
+ bias_splits[i] = bias_splits[i].reshape(num_insts *
+ self.in_channels)
+ else:
+ # out_channels x in_channels x 1 x 1
+ weight_splits[i] = weight_splits[i].reshape(
+ num_insts * 1, -1, 1, 1)
+ bias_splits[i] = bias_splits[i].reshape(num_insts)
+
+ return weight_splits, bias_splits
+
+ def dynamic_conv_forward(self, features: Tensor, weights: List[Tensor],
+ biases: List[Tensor], num_insts: int) -> Tensor:
+ """dynamic forward, each layer follow a relu."""
+ n_layers = len(weights)
+ x = features
+ for i, (w, b) in enumerate(zip(weights, biases)):
+ x = F.conv2d(x, w, bias=b, stride=1, padding=0, groups=num_insts)
+ if i < n_layers - 1:
+ x = F.relu(x)
+ return x
+
+ def forward(self, x: tuple, positive_infos: InstanceList) -> tuple:
+ """Forward feature from the upstream network to get prototypes and
+ linearly combine the prototypes, using masks coefficients, into
+ instance masks. Finally, crop the instance masks with given bboxes.
+
+ Args:
+ x (Tuple[Tensor]): Feature from the upstream network, which is
+ a 4D-tensor.
+ positive_infos (List[:obj:``InstanceData``]): Positive information
+ that calculate from detect head.
+
+ Returns:
+ tuple: Predicted instance segmentation masks
+ """
+ mask_feats = self.mask_feature_head(x)
+ return multi_apply(self.forward_single, mask_feats, positive_infos)
+
+ def forward_single(self, mask_feat: Tensor,
+ positive_info: InstanceData) -> Tensor:
+ """Forward features of a each image."""
+ pos_param_preds = positive_info.get('param_preds')
+ pos_points = positive_info.get('points')
+ pos_strides = positive_info.get('strides')
+
+ num_inst = pos_param_preds.shape[0]
+ mask_feat = mask_feat[None].repeat(num_inst, 1, 1, 1)
+ _, _, H, W = mask_feat.size()
+ if num_inst == 0:
+ return (pos_param_preds.new_zeros((0, 1, H, W)), )
+
+ locations = self.prior_generator.single_level_grid_priors(
+ mask_feat.size()[2:], 0, device=mask_feat.device)
+
+ rel_coords = relative_coordinate_maps(locations, pos_points,
+ pos_strides,
+ self.size_of_interest,
+ mask_feat.size()[2:])
+ mask_head_inputs = torch.cat([rel_coords, mask_feat], dim=1)
+ mask_head_inputs = mask_head_inputs.reshape(1, -1, H, W)
+
+ weights, biases = self.parse_dynamic_params(pos_param_preds)
+ mask_preds = self.dynamic_conv_forward(mask_head_inputs, weights,
+ biases, num_inst)
+ mask_preds = mask_preds.reshape(-1, H, W)
+ mask_preds = aligned_bilinear(
+ mask_preds.unsqueeze(0),
+ int(self.mask_feat_stride / self.mask_out_stride)).squeeze(0)
+
+ return (mask_preds, )
+
+ def loss_by_feat(self, mask_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict], positive_infos: InstanceList,
+ **kwargs) -> dict:
+ """Calculate the loss based on the features extracted by the mask head.
+
+ Args:
+ mask_preds (list[Tensor]): List of predicted masks, each has
+ shape (num_classes, H, W).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``masks``,
+ and ``labels`` attributes.
+ batch_img_metas (list[dict]): Meta information of multiple images.
+ positive_infos (List[:obj:``InstanceData``]): Information of
+ positive samples of each image that are assigned in detection
+ head.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ assert positive_infos is not None, \
+ 'positive_infos should not be None in `CondInstMaskHead`'
+ losses = dict()
+
+ loss_mask = 0.
+ num_imgs = len(mask_preds)
+ total_pos = 0
+
+ for idx in range(num_imgs):
+ (mask_pred, pos_mask_targets, num_pos) = \
+ self._get_targets_single(
+ mask_preds[idx], batch_gt_instances[idx],
+ positive_infos[idx])
+ # mask loss
+ total_pos += num_pos
+ if num_pos == 0 or pos_mask_targets is None:
+ loss = mask_pred.new_zeros(1).mean()
+ else:
+ loss = self.loss_mask(
+ mask_pred, pos_mask_targets,
+ reduction_override='none').sum()
+ loss_mask += loss
+
+ if total_pos == 0:
+ total_pos += 1 # avoid nan
+ loss_mask = loss_mask / total_pos
+ losses.update(loss_mask=loss_mask)
+ return losses
+
+ def _get_targets_single(self, mask_preds: Tensor,
+ gt_instances: InstanceData,
+ positive_info: InstanceData):
+ """Compute targets for predictions of single image.
+
+ Args:
+ mask_preds (Tensor): Predicted prototypes with shape
+ (num_classes, H, W).
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes``, ``labels``,
+ and ``masks`` attributes.
+ positive_info (:obj:`InstanceData`): Information of positive
+ samples that are assigned in detection head. It usually
+ contains following keys.
+
+ - pos_assigned_gt_inds (Tensor): Assigner GT indexes of
+ positive proposals, has shape (num_pos, )
+ - pos_inds (Tensor): Positive index of image, has
+ shape (num_pos, ).
+ - param_pred (Tensor): Positive param preditions
+ with shape (num_pos, num_params).
+
+ Returns:
+ tuple: Usually returns a tuple containing learning targets.
+
+ - mask_preds (Tensor): Positive predicted mask with shape
+ (num_pos, mask_h, mask_w).
+ - pos_mask_targets (Tensor): Positive mask targets with shape
+ (num_pos, mask_h, mask_w).
+ - num_pos (int): Positive numbers.
+ """
+ gt_bboxes = gt_instances.bboxes
+ device = gt_bboxes.device
+ gt_masks = gt_instances.masks.to_tensor(
+ dtype=torch.bool, device=device).float()
+
+ # process with mask targets
+ pos_assigned_gt_inds = positive_info.get('pos_assigned_gt_inds')
+ scores = positive_info.get('scores')
+ centernesses = positive_info.get('centernesses')
+ num_pos = pos_assigned_gt_inds.size(0)
+
+ if gt_masks.size(0) == 0 or num_pos == 0:
+ return mask_preds, None, 0
+ # Since we're producing (near) full image masks,
+ # it'd take too much vram to backprop on every single mask.
+ # Thus we select only a subset.
+ if (self.max_masks_to_train != -1) and \
+ (num_pos > self.max_masks_to_train):
+ perm = torch.randperm(num_pos)
+ select = perm[:self.max_masks_to_train]
+ mask_preds = mask_preds[select]
+ pos_assigned_gt_inds = pos_assigned_gt_inds[select]
+ num_pos = self.max_masks_to_train
+ elif self.topk_masks_per_img != -1:
+ unique_gt_inds = pos_assigned_gt_inds.unique()
+ num_inst_per_gt = max(
+ int(self.topk_masks_per_img / len(unique_gt_inds)), 1)
+
+ keep_mask_preds = []
+ keep_pos_assigned_gt_inds = []
+ for gt_ind in unique_gt_inds:
+ per_inst_pos_inds = (pos_assigned_gt_inds == gt_ind)
+ mask_preds_per_inst = mask_preds[per_inst_pos_inds]
+ gt_inds_per_inst = pos_assigned_gt_inds[per_inst_pos_inds]
+ if sum(per_inst_pos_inds) > num_inst_per_gt:
+ per_inst_scores = scores[per_inst_pos_inds].sigmoid().max(
+ dim=1)[0]
+ per_inst_centerness = centernesses[
+ per_inst_pos_inds].sigmoid().reshape(-1, )
+ select = (per_inst_scores * per_inst_centerness).topk(
+ k=num_inst_per_gt, dim=0)[1]
+ mask_preds_per_inst = mask_preds_per_inst[select]
+ gt_inds_per_inst = gt_inds_per_inst[select]
+ keep_mask_preds.append(mask_preds_per_inst)
+ keep_pos_assigned_gt_inds.append(gt_inds_per_inst)
+ mask_preds = torch.cat(keep_mask_preds)
+ pos_assigned_gt_inds = torch.cat(keep_pos_assigned_gt_inds)
+ num_pos = pos_assigned_gt_inds.size(0)
+
+ # Follow the origin implement
+ start = int(self.mask_out_stride // 2)
+ gt_masks = gt_masks[:, start::self.mask_out_stride,
+ start::self.mask_out_stride]
+ gt_masks = gt_masks.gt(0.5).float()
+ pos_mask_targets = gt_masks[pos_assigned_gt_inds]
+
+ return (mask_preds, pos_mask_targets, num_pos)
+
+ def predict_by_feat(self,
+ mask_preds: List[Tensor],
+ results_list: InstanceList,
+ batch_img_metas: List[dict],
+ rescale: bool = True,
+ **kwargs) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ mask results.
+
+ Args:
+ mask_preds (list[Tensor]): Predicted prototypes with shape
+ (num_classes, H, W).
+ results_list (List[:obj:``InstanceData``]): BBoxHead results.
+ batch_img_metas (list[dict]): Meta information of all images.
+ rescale (bool, optional): Whether to rescale the results.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Processed results of multiple
+ images.Each :obj:`InstanceData` usually contains
+ following keys.
+
+ - scores (Tensor): Classification scores, has shape
+ (num_instance,).
+ - labels (Tensor): Has shape (num_instances,).
+ - masks (Tensor): Processed mask results, has
+ shape (num_instances, h, w).
+ """
+ assert len(mask_preds) == len(results_list) == len(batch_img_metas)
+
+ for img_id in range(len(batch_img_metas)):
+ img_meta = batch_img_metas[img_id]
+ results = results_list[img_id]
+ bboxes = results.bboxes
+ mask_pred = mask_preds[img_id]
+ if bboxes.shape[0] == 0 or mask_pred.shape[0] == 0:
+ results_list[img_id] = empty_instances(
+ [img_meta],
+ bboxes.device,
+ task_type='mask',
+ instance_results=[results])[0]
+ else:
+ im_mask = self._predict_by_feat_single(
+ mask_preds=mask_pred,
+ bboxes=bboxes,
+ img_meta=img_meta,
+ rescale=rescale)
+ results.masks = im_mask
+ return results_list
+
+ def _predict_by_feat_single(self,
+ mask_preds: Tensor,
+ bboxes: Tensor,
+ img_meta: dict,
+ rescale: bool,
+ cfg: OptConfigType = None):
+ """Transform a single image's features extracted from the head into
+ mask results.
+
+ Args:
+ mask_preds (Tensor): Predicted prototypes, has shape [H, W, N].
+ img_meta (dict): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ rescale (bool): If rescale is False, then returned masks will
+ fit the scale of imgs[0].
+ cfg (dict, optional): Config used in test phase.
+ Defaults to None.
+
+ Returns:
+ :obj:`InstanceData`: Processed results of single image.
+ it usually contains following keys.
+
+ - scores (Tensor): Classification scores, has shape
+ (num_instance,).
+ - labels (Tensor): Has shape (num_instances,).
+ - masks (Tensor): Processed mask results, has
+ shape (num_instances, h, w).
+ """
+ cfg = self.test_cfg if cfg is None else cfg
+ scale_factor = bboxes.new_tensor(img_meta['scale_factor']).repeat(
+ (1, 2))
+ img_h, img_w = img_meta['img_shape'][:2]
+ ori_h, ori_w = img_meta['ori_shape'][:2]
+
+ mask_preds = mask_preds.sigmoid().unsqueeze(0)
+ mask_preds = aligned_bilinear(mask_preds, self.mask_out_stride)
+ mask_preds = mask_preds[:, :, :img_h, :img_w]
+ if rescale: # in-placed rescale the bboxes
+ scale_factor = bboxes.new_tensor(img_meta['scale_factor']).repeat(
+ (1, 2))
+ bboxes /= scale_factor
+
+ masks = F.interpolate(
+ mask_preds, (ori_h, ori_w),
+ mode='bilinear',
+ align_corners=False).squeeze(0) > cfg.mask_thr
+ else:
+ masks = mask_preds.squeeze(0) > cfg.mask_thr
+
+ return masks
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/conditional_detr_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/conditional_detr_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..cc2df2c215667121c5fe329f369510ecd4666faf
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/conditional_detr_head.py
@@ -0,0 +1,168 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Tuple
+
+import torch
+import torch.nn as nn
+from mmengine.model import bias_init_with_prob
+from torch import Tensor
+
+from mmdet.models.layers.transformer import inverse_sigmoid
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.utils import InstanceList
+from .detr_head import DETRHead
+
+
+@MODELS.register_module()
+class ConditionalDETRHead(DETRHead):
+ """Head of Conditional DETR. Conditional DETR: Conditional DETR for Fast
+ Training Convergence. More details can be found in the `paper.
+
+ `_ .
+ """
+
+ def init_weights(self):
+ """Initialize weights of the transformer head."""
+ super().init_weights()
+ # The initialization below for transformer head is very
+ # important as we use Focal_loss for loss_cls
+ if self.loss_cls.use_sigmoid:
+ bias_init = bias_init_with_prob(0.01)
+ nn.init.constant_(self.fc_cls.bias, bias_init)
+
+ def forward(self, hidden_states: Tensor,
+ references: Tensor) -> Tuple[Tensor, Tensor]:
+ """"Forward function.
+
+ Args:
+ hidden_states (Tensor): Features from transformer decoder. If
+ `return_intermediate_dec` is True output has shape
+ (num_decoder_layers, bs, num_queries, dim), else has shape (1,
+ bs, num_queries, dim) which only contains the last layer
+ outputs.
+ references (Tensor): References from transformer decoder, has
+ shape (bs, num_queries, 2).
+ Returns:
+ tuple[Tensor]: results of head containing the following tensor.
+
+ - layers_cls_scores (Tensor): Outputs from the classification head,
+ shape (num_decoder_layers, bs, num_queries, cls_out_channels).
+ Note cls_out_channels should include background.
+ - layers_bbox_preds (Tensor): Sigmoid outputs from the regression
+ head with normalized coordinate format (cx, cy, w, h), has shape
+ (num_decoder_layers, bs, num_queries, 4).
+ """
+
+ references_unsigmoid = inverse_sigmoid(references)
+ layers_bbox_preds = []
+ for layer_id in range(hidden_states.shape[0]):
+ tmp_reg_preds = self.fc_reg(
+ self.activate(self.reg_ffn(hidden_states[layer_id])))
+ tmp_reg_preds[..., :2] += references_unsigmoid
+ outputs_coord = tmp_reg_preds.sigmoid()
+ layers_bbox_preds.append(outputs_coord)
+ layers_bbox_preds = torch.stack(layers_bbox_preds)
+
+ layers_cls_scores = self.fc_cls(hidden_states)
+ return layers_cls_scores, layers_bbox_preds
+
+ def loss(self, hidden_states: Tensor, references: Tensor,
+ batch_data_samples: SampleList) -> dict:
+ """Perform forward propagation and loss calculation of the detection
+ head on the features of the upstream network.
+
+ Args:
+ hidden_states (Tensor): Features from the transformer decoder, has
+ shape (num_decoder_layers, bs, num_queries, dim).
+ references (Tensor): References from the transformer decoder, has
+ shape (num_decoder_layers, bs, num_queries, 2).
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ batch_gt_instances = []
+ batch_img_metas = []
+ for data_sample in batch_data_samples:
+ batch_img_metas.append(data_sample.metainfo)
+ batch_gt_instances.append(data_sample.gt_instances)
+
+ outs = self(hidden_states, references)
+ loss_inputs = outs + (batch_gt_instances, batch_img_metas)
+ losses = self.loss_by_feat(*loss_inputs)
+ return losses
+
+ def loss_and_predict(
+ self, hidden_states: Tensor, references: Tensor,
+ batch_data_samples: SampleList) -> Tuple[dict, InstanceList]:
+ """Perform forward propagation of the head, then calculate loss and
+ predictions from the features and data samples. Over-write because
+ img_metas are needed as inputs for bbox_head.
+
+ Args:
+ hidden_states (Tensor): Features from the transformer decoder, has
+ shape (num_decoder_layers, bs, num_queries, dim).
+ references (Tensor): References from the transformer decoder, has
+ shape (num_decoder_layers, bs, num_queries, 2).
+ batch_data_samples (list[:obj:`DetDataSample`]): Each item contains
+ the meta information of each image and corresponding
+ annotations.
+
+ Returns:
+ tuple: The return value is a tuple contains:
+
+ - losses: (dict[str, Tensor]): A dictionary of loss components.
+ - predictions (list[:obj:`InstanceData`]): Detection
+ results of each image after the post process.
+ """
+ batch_gt_instances = []
+ batch_img_metas = []
+ for data_sample in batch_data_samples:
+ batch_img_metas.append(data_sample.metainfo)
+ batch_gt_instances.append(data_sample.gt_instances)
+
+ outs = self(hidden_states, references)
+ loss_inputs = outs + (batch_gt_instances, batch_img_metas)
+ losses = self.loss_by_feat(*loss_inputs)
+
+ predictions = self.predict_by_feat(
+ *outs, batch_img_metas=batch_img_metas)
+ return losses, predictions
+
+ def predict(self,
+ hidden_states: Tensor,
+ references: Tensor,
+ batch_data_samples: SampleList,
+ rescale: bool = True) -> InstanceList:
+ """Perform forward propagation of the detection head and predict
+ detection results on the features of the upstream network. Over-write
+ because img_metas are needed as inputs for bbox_head.
+
+ Args:
+ hidden_states (Tensor): Features from the transformer decoder, has
+ shape (num_decoder_layers, bs, num_queries, dim).
+ references (Tensor): References from the transformer decoder, has
+ shape (num_decoder_layers, bs, num_queries, 2).
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool, optional): Whether to rescale the results.
+ Defaults to True.
+
+ Returns:
+ list[obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ """
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+
+ last_layer_hidden_state = hidden_states[-1].unsqueeze(0)
+ outs = self(last_layer_hidden_state, references)
+
+ predictions = self.predict_by_feat(
+ *outs, batch_img_metas=batch_img_metas, rescale=rescale)
+
+ return predictions
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/corner_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/corner_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..0cec71d50947ff58224ae698ec9c2f9406b58efb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/corner_head.py
@@ -0,0 +1,1084 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from logging import warning
+from math import ceil, log
+from typing import List, Optional, Sequence, Tuple
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from mmcv.ops import CornerPool, batched_nms
+from mmengine.config import ConfigDict
+from mmengine.model import BaseModule, bias_init_with_prob
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import (ConfigType, InstanceList, OptConfigType,
+ OptInstanceList, OptMultiConfig)
+from ..utils import (gather_feat, gaussian_radius, gen_gaussian_target,
+ get_local_maximum, get_topk_from_heatmap, multi_apply,
+ transpose_and_gather_feat)
+from .base_dense_head import BaseDenseHead
+
+
+class BiCornerPool(BaseModule):
+ """Bidirectional Corner Pooling Module (TopLeft, BottomRight, etc.)
+
+ Args:
+ in_channels (int): Input channels of module.
+ directions (list[str]): Directions of two CornerPools.
+ out_channels (int): Output channels of module.
+ feat_channels (int): Feature channels of module.
+ norm_cfg (:obj:`ConfigDict` or dict): Dictionary to construct
+ and config norm layer.
+ init_cfg (:obj:`ConfigDict` or dict, optional): the config to
+ control the initialization.
+ """
+
+ def __init__(self,
+ in_channels: int,
+ directions: List[int],
+ feat_channels: int = 128,
+ out_channels: int = 128,
+ norm_cfg: ConfigType = dict(type='BN', requires_grad=True),
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg)
+ self.direction1_conv = ConvModule(
+ in_channels, feat_channels, 3, padding=1, norm_cfg=norm_cfg)
+ self.direction2_conv = ConvModule(
+ in_channels, feat_channels, 3, padding=1, norm_cfg=norm_cfg)
+
+ self.aftpool_conv = ConvModule(
+ feat_channels,
+ out_channels,
+ 3,
+ padding=1,
+ norm_cfg=norm_cfg,
+ act_cfg=None)
+
+ self.conv1 = ConvModule(
+ in_channels, out_channels, 1, norm_cfg=norm_cfg, act_cfg=None)
+ self.conv2 = ConvModule(
+ in_channels, out_channels, 3, padding=1, norm_cfg=norm_cfg)
+
+ self.direction1_pool = CornerPool(directions[0])
+ self.direction2_pool = CornerPool(directions[1])
+ self.relu = nn.ReLU(inplace=True)
+
+ def forward(self, x: Tensor) -> Tensor:
+ """Forward features from the upstream network.
+
+ Args:
+ x (tensor): Input feature of BiCornerPool.
+
+ Returns:
+ conv2 (tensor): Output feature of BiCornerPool.
+ """
+ direction1_conv = self.direction1_conv(x)
+ direction2_conv = self.direction2_conv(x)
+ direction1_feat = self.direction1_pool(direction1_conv)
+ direction2_feat = self.direction2_pool(direction2_conv)
+ aftpool_conv = self.aftpool_conv(direction1_feat + direction2_feat)
+ conv1 = self.conv1(x)
+ relu = self.relu(aftpool_conv + conv1)
+ conv2 = self.conv2(relu)
+ return conv2
+
+
+@MODELS.register_module()
+class CornerHead(BaseDenseHead):
+ """Head of CornerNet: Detecting Objects as Paired Keypoints.
+
+ Code is modified from the `official github repo
+ `_ .
+
+ More details can be found in the `paper
+ `_ .
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ num_feat_levels (int): Levels of feature from the previous module.
+ 2 for HourglassNet-104 and 1 for HourglassNet-52. Because
+ HourglassNet-104 outputs the final feature and intermediate
+ supervision feature and HourglassNet-52 only outputs the final
+ feature. Defaults to 2.
+ corner_emb_channels (int): Channel of embedding vector. Defaults to 1.
+ train_cfg (:obj:`ConfigDict` or dict, optional): Training config.
+ Useless in CornerHead, but we keep this variable for
+ SingleStageDetector.
+ test_cfg (:obj:`ConfigDict` or dict, optional): Testing config of
+ CornerHead.
+ loss_heatmap (:obj:`ConfigDict` or dict): Config of corner heatmap
+ loss. Defaults to GaussianFocalLoss.
+ loss_embedding (:obj:`ConfigDict` or dict): Config of corner embedding
+ loss. Defaults to AssociativeEmbeddingLoss.
+ loss_offset (:obj:`ConfigDict` or dict): Config of corner offset loss.
+ Defaults to SmoothL1Loss.
+ init_cfg (:obj:`ConfigDict` or dict, optional): the config to control
+ the initialization.
+ """
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: int,
+ num_feat_levels: int = 2,
+ corner_emb_channels: int = 1,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ loss_heatmap: ConfigType = dict(
+ type='GaussianFocalLoss',
+ alpha=2.0,
+ gamma=4.0,
+ loss_weight=1),
+ loss_embedding: ConfigType = dict(
+ type='AssociativeEmbeddingLoss',
+ pull_weight=0.25,
+ push_weight=0.25),
+ loss_offset: ConfigType = dict(
+ type='SmoothL1Loss', beta=1.0, loss_weight=1),
+ init_cfg: OptMultiConfig = None) -> None:
+ assert init_cfg is None, 'To prevent abnormal initialization ' \
+ 'behavior, init_cfg is not allowed to be set'
+ super().__init__(init_cfg=init_cfg)
+ self.num_classes = num_classes
+ self.in_channels = in_channels
+ self.corner_emb_channels = corner_emb_channels
+ self.with_corner_emb = self.corner_emb_channels > 0
+ self.corner_offset_channels = 2
+ self.num_feat_levels = num_feat_levels
+ self.loss_heatmap = MODELS.build(
+ loss_heatmap) if loss_heatmap is not None else None
+ self.loss_embedding = MODELS.build(
+ loss_embedding) if loss_embedding is not None else None
+ self.loss_offset = MODELS.build(
+ loss_offset) if loss_offset is not None else None
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+
+ self._init_layers()
+
+ def _make_layers(self,
+ out_channels: int,
+ in_channels: int = 256,
+ feat_channels: int = 256) -> nn.Sequential:
+ """Initialize conv sequential for CornerHead."""
+ return nn.Sequential(
+ ConvModule(in_channels, feat_channels, 3, padding=1),
+ ConvModule(
+ feat_channels, out_channels, 1, norm_cfg=None, act_cfg=None))
+
+ def _init_corner_kpt_layers(self) -> None:
+ """Initialize corner keypoint layers.
+
+ Including corner heatmap branch and corner offset branch. Each branch
+ has two parts: prefix `tl_` for top-left and `br_` for bottom-right.
+ """
+ self.tl_pool, self.br_pool = nn.ModuleList(), nn.ModuleList()
+ self.tl_heat, self.br_heat = nn.ModuleList(), nn.ModuleList()
+ self.tl_off, self.br_off = nn.ModuleList(), nn.ModuleList()
+
+ for _ in range(self.num_feat_levels):
+ self.tl_pool.append(
+ BiCornerPool(
+ self.in_channels, ['top', 'left'],
+ out_channels=self.in_channels))
+ self.br_pool.append(
+ BiCornerPool(
+ self.in_channels, ['bottom', 'right'],
+ out_channels=self.in_channels))
+
+ self.tl_heat.append(
+ self._make_layers(
+ out_channels=self.num_classes,
+ in_channels=self.in_channels))
+ self.br_heat.append(
+ self._make_layers(
+ out_channels=self.num_classes,
+ in_channels=self.in_channels))
+
+ self.tl_off.append(
+ self._make_layers(
+ out_channels=self.corner_offset_channels,
+ in_channels=self.in_channels))
+ self.br_off.append(
+ self._make_layers(
+ out_channels=self.corner_offset_channels,
+ in_channels=self.in_channels))
+
+ def _init_corner_emb_layers(self) -> None:
+ """Initialize corner embedding layers.
+
+ Only include corner embedding branch with two parts: prefix `tl_` for
+ top-left and `br_` for bottom-right.
+ """
+ self.tl_emb, self.br_emb = nn.ModuleList(), nn.ModuleList()
+
+ for _ in range(self.num_feat_levels):
+ self.tl_emb.append(
+ self._make_layers(
+ out_channels=self.corner_emb_channels,
+ in_channels=self.in_channels))
+ self.br_emb.append(
+ self._make_layers(
+ out_channels=self.corner_emb_channels,
+ in_channels=self.in_channels))
+
+ def _init_layers(self) -> None:
+ """Initialize layers for CornerHead.
+
+ Including two parts: corner keypoint layers and corner embedding layers
+ """
+ self._init_corner_kpt_layers()
+ if self.with_corner_emb:
+ self._init_corner_emb_layers()
+
+ def init_weights(self) -> None:
+ super().init_weights()
+ bias_init = bias_init_with_prob(0.1)
+ for i in range(self.num_feat_levels):
+ # The initialization of parameters are different between
+ # nn.Conv2d and ConvModule. Our experiments show that
+ # using the original initialization of nn.Conv2d increases
+ # the final mAP by about 0.2%
+ self.tl_heat[i][-1].conv.reset_parameters()
+ self.tl_heat[i][-1].conv.bias.data.fill_(bias_init)
+ self.br_heat[i][-1].conv.reset_parameters()
+ self.br_heat[i][-1].conv.bias.data.fill_(bias_init)
+ self.tl_off[i][-1].conv.reset_parameters()
+ self.br_off[i][-1].conv.reset_parameters()
+ if self.with_corner_emb:
+ self.tl_emb[i][-1].conv.reset_parameters()
+ self.br_emb[i][-1].conv.reset_parameters()
+
+ def forward(self, feats: Tuple[Tensor]) -> tuple:
+ """Forward features from the upstream network.
+
+ Args:
+ feats (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: Usually a tuple of corner heatmaps, offset heatmaps and
+ embedding heatmaps.
+ - tl_heats (list[Tensor]): Top-left corner heatmaps for all
+ levels, each is a 4D-tensor, the channels number is
+ num_classes.
+ - br_heats (list[Tensor]): Bottom-right corner heatmaps for all
+ levels, each is a 4D-tensor, the channels number is
+ num_classes.
+ - tl_embs (list[Tensor] | list[None]): Top-left embedding
+ heatmaps for all levels, each is a 4D-tensor or None.
+ If not None, the channels number is corner_emb_channels.
+ - br_embs (list[Tensor] | list[None]): Bottom-right embedding
+ heatmaps for all levels, each is a 4D-tensor or None.
+ If not None, the channels number is corner_emb_channels.
+ - tl_offs (list[Tensor]): Top-left offset heatmaps for all
+ levels, each is a 4D-tensor. The channels number is
+ corner_offset_channels.
+ - br_offs (list[Tensor]): Bottom-right offset heatmaps for all
+ levels, each is a 4D-tensor. The channels number is
+ corner_offset_channels.
+ """
+ lvl_ind = list(range(self.num_feat_levels))
+ return multi_apply(self.forward_single, feats, lvl_ind)
+
+ def forward_single(self,
+ x: Tensor,
+ lvl_ind: int,
+ return_pool: bool = False) -> List[Tensor]:
+ """Forward feature of a single level.
+
+ Args:
+ x (Tensor): Feature of a single level.
+ lvl_ind (int): Level index of current feature.
+ return_pool (bool): Return corner pool feature or not.
+ Defaults to False.
+
+ Returns:
+ tuple[Tensor]: A tuple of CornerHead's output for current feature
+ level. Containing the following Tensors:
+
+ - tl_heat (Tensor): Predicted top-left corner heatmap.
+ - br_heat (Tensor): Predicted bottom-right corner heatmap.
+ - tl_emb (Tensor | None): Predicted top-left embedding heatmap.
+ None for `self.with_corner_emb == False`.
+ - br_emb (Tensor | None): Predicted bottom-right embedding
+ heatmap. None for `self.with_corner_emb == False`.
+ - tl_off (Tensor): Predicted top-left offset heatmap.
+ - br_off (Tensor): Predicted bottom-right offset heatmap.
+ - tl_pool (Tensor): Top-left corner pool feature. Not must
+ have.
+ - br_pool (Tensor): Bottom-right corner pool feature. Not must
+ have.
+ """
+ tl_pool = self.tl_pool[lvl_ind](x)
+ tl_heat = self.tl_heat[lvl_ind](tl_pool)
+ br_pool = self.br_pool[lvl_ind](x)
+ br_heat = self.br_heat[lvl_ind](br_pool)
+
+ tl_emb, br_emb = None, None
+ if self.with_corner_emb:
+ tl_emb = self.tl_emb[lvl_ind](tl_pool)
+ br_emb = self.br_emb[lvl_ind](br_pool)
+
+ tl_off = self.tl_off[lvl_ind](tl_pool)
+ br_off = self.br_off[lvl_ind](br_pool)
+
+ result_list = [tl_heat, br_heat, tl_emb, br_emb, tl_off, br_off]
+ if return_pool:
+ result_list.append(tl_pool)
+ result_list.append(br_pool)
+
+ return result_list
+
+ def get_targets(self,
+ gt_bboxes: List[Tensor],
+ gt_labels: List[Tensor],
+ feat_shape: Sequence[int],
+ img_shape: Sequence[int],
+ with_corner_emb: bool = False,
+ with_guiding_shift: bool = False,
+ with_centripetal_shift: bool = False) -> dict:
+ """Generate corner targets.
+
+ Including corner heatmap, corner offset.
+
+ Optional: corner embedding, corner guiding shift, centripetal shift.
+
+ For CornerNet, we generate corner heatmap, corner offset and corner
+ embedding from this function.
+
+ For CentripetalNet, we generate corner heatmap, corner offset, guiding
+ shift and centripetal shift from this function.
+
+ Args:
+ gt_bboxes (list[Tensor]): Ground truth bboxes of each image, each
+ has shape (num_gt, 4).
+ gt_labels (list[Tensor]): Ground truth labels of each box, each has
+ shape (num_gt, ).
+ feat_shape (Sequence[int]): Shape of output feature,
+ [batch, channel, height, width].
+ img_shape (Sequence[int]): Shape of input image,
+ [height, width, channel].
+ with_corner_emb (bool): Generate corner embedding target or not.
+ Defaults to False.
+ with_guiding_shift (bool): Generate guiding shift target or not.
+ Defaults to False.
+ with_centripetal_shift (bool): Generate centripetal shift target or
+ not. Defaults to False.
+
+ Returns:
+ dict: Ground truth of corner heatmap, corner offset, corner
+ embedding, guiding shift and centripetal shift. Containing the
+ following keys:
+
+ - topleft_heatmap (Tensor): Ground truth top-left corner
+ heatmap.
+ - bottomright_heatmap (Tensor): Ground truth bottom-right
+ corner heatmap.
+ - topleft_offset (Tensor): Ground truth top-left corner offset.
+ - bottomright_offset (Tensor): Ground truth bottom-right corner
+ offset.
+ - corner_embedding (list[list[list[int]]]): Ground truth corner
+ embedding. Not must have.
+ - topleft_guiding_shift (Tensor): Ground truth top-left corner
+ guiding shift. Not must have.
+ - bottomright_guiding_shift (Tensor): Ground truth bottom-right
+ corner guiding shift. Not must have.
+ - topleft_centripetal_shift (Tensor): Ground truth top-left
+ corner centripetal shift. Not must have.
+ - bottomright_centripetal_shift (Tensor): Ground truth
+ bottom-right corner centripetal shift. Not must have.
+ """
+ batch_size, _, height, width = feat_shape
+ img_h, img_w = img_shape[:2]
+
+ width_ratio = float(width / img_w)
+ height_ratio = float(height / img_h)
+
+ gt_tl_heatmap = gt_bboxes[-1].new_zeros(
+ [batch_size, self.num_classes, height, width])
+ gt_br_heatmap = gt_bboxes[-1].new_zeros(
+ [batch_size, self.num_classes, height, width])
+ gt_tl_offset = gt_bboxes[-1].new_zeros([batch_size, 2, height, width])
+ gt_br_offset = gt_bboxes[-1].new_zeros([batch_size, 2, height, width])
+
+ if with_corner_emb:
+ match = []
+
+ # Guiding shift is a kind of offset, from center to corner
+ if with_guiding_shift:
+ gt_tl_guiding_shift = gt_bboxes[-1].new_zeros(
+ [batch_size, 2, height, width])
+ gt_br_guiding_shift = gt_bboxes[-1].new_zeros(
+ [batch_size, 2, height, width])
+ # Centripetal shift is also a kind of offset, from center to corner
+ # and normalized by log.
+ if with_centripetal_shift:
+ gt_tl_centripetal_shift = gt_bboxes[-1].new_zeros(
+ [batch_size, 2, height, width])
+ gt_br_centripetal_shift = gt_bboxes[-1].new_zeros(
+ [batch_size, 2, height, width])
+
+ for batch_id in range(batch_size):
+ # Ground truth of corner embedding per image is a list of coord set
+ corner_match = []
+ for box_id in range(len(gt_labels[batch_id])):
+ left, top, right, bottom = gt_bboxes[batch_id][box_id]
+ center_x = (left + right) / 2.0
+ center_y = (top + bottom) / 2.0
+ label = gt_labels[batch_id][box_id]
+
+ # Use coords in the feature level to generate ground truth
+ scale_left = left * width_ratio
+ scale_right = right * width_ratio
+ scale_top = top * height_ratio
+ scale_bottom = bottom * height_ratio
+ scale_center_x = center_x * width_ratio
+ scale_center_y = center_y * height_ratio
+
+ # Int coords on feature map/ground truth tensor
+ left_idx = int(min(scale_left, width - 1))
+ right_idx = int(min(scale_right, width - 1))
+ top_idx = int(min(scale_top, height - 1))
+ bottom_idx = int(min(scale_bottom, height - 1))
+
+ # Generate gaussian heatmap
+ scale_box_width = ceil(scale_right - scale_left)
+ scale_box_height = ceil(scale_bottom - scale_top)
+ radius = gaussian_radius((scale_box_height, scale_box_width),
+ min_overlap=0.3)
+ radius = max(0, int(radius))
+ gt_tl_heatmap[batch_id, label] = gen_gaussian_target(
+ gt_tl_heatmap[batch_id, label], [left_idx, top_idx],
+ radius)
+ gt_br_heatmap[batch_id, label] = gen_gaussian_target(
+ gt_br_heatmap[batch_id, label], [right_idx, bottom_idx],
+ radius)
+
+ # Generate corner offset
+ left_offset = scale_left - left_idx
+ top_offset = scale_top - top_idx
+ right_offset = scale_right - right_idx
+ bottom_offset = scale_bottom - bottom_idx
+ gt_tl_offset[batch_id, 0, top_idx, left_idx] = left_offset
+ gt_tl_offset[batch_id, 1, top_idx, left_idx] = top_offset
+ gt_br_offset[batch_id, 0, bottom_idx, right_idx] = right_offset
+ gt_br_offset[batch_id, 1, bottom_idx,
+ right_idx] = bottom_offset
+
+ # Generate corner embedding
+ if with_corner_emb:
+ corner_match.append([[top_idx, left_idx],
+ [bottom_idx, right_idx]])
+ # Generate guiding shift
+ if with_guiding_shift:
+ gt_tl_guiding_shift[batch_id, 0, top_idx,
+ left_idx] = scale_center_x - left_idx
+ gt_tl_guiding_shift[batch_id, 1, top_idx,
+ left_idx] = scale_center_y - top_idx
+ gt_br_guiding_shift[batch_id, 0, bottom_idx,
+ right_idx] = right_idx - scale_center_x
+ gt_br_guiding_shift[
+ batch_id, 1, bottom_idx,
+ right_idx] = bottom_idx - scale_center_y
+ # Generate centripetal shift
+ if with_centripetal_shift:
+ gt_tl_centripetal_shift[batch_id, 0, top_idx,
+ left_idx] = log(scale_center_x -
+ scale_left)
+ gt_tl_centripetal_shift[batch_id, 1, top_idx,
+ left_idx] = log(scale_center_y -
+ scale_top)
+ gt_br_centripetal_shift[batch_id, 0, bottom_idx,
+ right_idx] = log(scale_right -
+ scale_center_x)
+ gt_br_centripetal_shift[batch_id, 1, bottom_idx,
+ right_idx] = log(scale_bottom -
+ scale_center_y)
+
+ if with_corner_emb:
+ match.append(corner_match)
+
+ target_result = dict(
+ topleft_heatmap=gt_tl_heatmap,
+ topleft_offset=gt_tl_offset,
+ bottomright_heatmap=gt_br_heatmap,
+ bottomright_offset=gt_br_offset)
+
+ if with_corner_emb:
+ target_result.update(corner_embedding=match)
+ if with_guiding_shift:
+ target_result.update(
+ topleft_guiding_shift=gt_tl_guiding_shift,
+ bottomright_guiding_shift=gt_br_guiding_shift)
+ if with_centripetal_shift:
+ target_result.update(
+ topleft_centripetal_shift=gt_tl_centripetal_shift,
+ bottomright_centripetal_shift=gt_br_centripetal_shift)
+
+ return target_result
+
+ def loss_by_feat(
+ self,
+ tl_heats: List[Tensor],
+ br_heats: List[Tensor],
+ tl_embs: List[Tensor],
+ br_embs: List[Tensor],
+ tl_offs: List[Tensor],
+ br_offs: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ tl_heats (list[Tensor]): Top-left corner heatmaps for each level
+ with shape (N, num_classes, H, W).
+ br_heats (list[Tensor]): Bottom-right corner heatmaps for each
+ level with shape (N, num_classes, H, W).
+ tl_embs (list[Tensor]): Top-left corner embeddings for each level
+ with shape (N, corner_emb_channels, H, W).
+ br_embs (list[Tensor]): Bottom-right corner embeddings for each
+ level with shape (N, corner_emb_channels, H, W).
+ tl_offs (list[Tensor]): Top-left corner offsets for each level
+ with shape (N, corner_offset_channels, H, W).
+ br_offs (list[Tensor]): Bottom-right corner offsets for each level
+ with shape (N, corner_offset_channels, H, W).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Specify which bounding boxes can be ignored when computing
+ the loss.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components. Containing the
+ following losses:
+
+ - det_loss (list[Tensor]): Corner keypoint losses of all
+ feature levels.
+ - pull_loss (list[Tensor]): Part one of AssociativeEmbedding
+ losses of all feature levels.
+ - push_loss (list[Tensor]): Part two of AssociativeEmbedding
+ losses of all feature levels.
+ - off_loss (list[Tensor]): Corner offset losses of all feature
+ levels.
+ """
+ gt_bboxes = [
+ gt_instances.bboxes for gt_instances in batch_gt_instances
+ ]
+ gt_labels = [
+ gt_instances.labels for gt_instances in batch_gt_instances
+ ]
+
+ targets = self.get_targets(
+ gt_bboxes,
+ gt_labels,
+ tl_heats[-1].shape,
+ batch_img_metas[0]['batch_input_shape'],
+ with_corner_emb=self.with_corner_emb)
+ mlvl_targets = [targets for _ in range(self.num_feat_levels)]
+ det_losses, pull_losses, push_losses, off_losses = multi_apply(
+ self.loss_by_feat_single, tl_heats, br_heats, tl_embs, br_embs,
+ tl_offs, br_offs, mlvl_targets)
+ loss_dict = dict(det_loss=det_losses, off_loss=off_losses)
+ if self.with_corner_emb:
+ loss_dict.update(pull_loss=pull_losses, push_loss=push_losses)
+ return loss_dict
+
+ def loss_by_feat_single(self, tl_hmp: Tensor, br_hmp: Tensor,
+ tl_emb: Optional[Tensor], br_emb: Optional[Tensor],
+ tl_off: Tensor, br_off: Tensor,
+ targets: dict) -> Tuple[Tensor, ...]:
+ """Calculate the loss of a single scale level based on the features
+ extracted by the detection head.
+
+ Args:
+ tl_hmp (Tensor): Top-left corner heatmap for current level with
+ shape (N, num_classes, H, W).
+ br_hmp (Tensor): Bottom-right corner heatmap for current level with
+ shape (N, num_classes, H, W).
+ tl_emb (Tensor, optional): Top-left corner embedding for current
+ level with shape (N, corner_emb_channels, H, W).
+ br_emb (Tensor, optional): Bottom-right corner embedding for
+ current level with shape (N, corner_emb_channels, H, W).
+ tl_off (Tensor): Top-left corner offset for current level with
+ shape (N, corner_offset_channels, H, W).
+ br_off (Tensor): Bottom-right corner offset for current level with
+ shape (N, corner_offset_channels, H, W).
+ targets (dict): Corner target generated by `get_targets`.
+
+ Returns:
+ tuple[torch.Tensor]: Losses of the head's different branches
+ containing the following losses:
+
+ - det_loss (Tensor): Corner keypoint loss.
+ - pull_loss (Tensor): Part one of AssociativeEmbedding loss.
+ - push_loss (Tensor): Part two of AssociativeEmbedding loss.
+ - off_loss (Tensor): Corner offset loss.
+ """
+ gt_tl_hmp = targets['topleft_heatmap']
+ gt_br_hmp = targets['bottomright_heatmap']
+ gt_tl_off = targets['topleft_offset']
+ gt_br_off = targets['bottomright_offset']
+ gt_embedding = targets['corner_embedding']
+
+ # Detection loss
+ tl_det_loss = self.loss_heatmap(
+ tl_hmp.sigmoid(),
+ gt_tl_hmp,
+ avg_factor=max(1,
+ gt_tl_hmp.eq(1).sum()))
+ br_det_loss = self.loss_heatmap(
+ br_hmp.sigmoid(),
+ gt_br_hmp,
+ avg_factor=max(1,
+ gt_br_hmp.eq(1).sum()))
+ det_loss = (tl_det_loss + br_det_loss) / 2.0
+
+ # AssociativeEmbedding loss
+ if self.with_corner_emb and self.loss_embedding is not None:
+ pull_loss, push_loss = self.loss_embedding(tl_emb, br_emb,
+ gt_embedding)
+ else:
+ pull_loss, push_loss = None, None
+
+ # Offset loss
+ # We only compute the offset loss at the real corner position.
+ # The value of real corner would be 1 in heatmap ground truth.
+ # The mask is computed in class agnostic mode and its shape is
+ # batch * 1 * width * height.
+ tl_off_mask = gt_tl_hmp.eq(1).sum(1).gt(0).unsqueeze(1).type_as(
+ gt_tl_hmp)
+ br_off_mask = gt_br_hmp.eq(1).sum(1).gt(0).unsqueeze(1).type_as(
+ gt_br_hmp)
+ tl_off_loss = self.loss_offset(
+ tl_off,
+ gt_tl_off,
+ tl_off_mask,
+ avg_factor=max(1, tl_off_mask.sum()))
+ br_off_loss = self.loss_offset(
+ br_off,
+ gt_br_off,
+ br_off_mask,
+ avg_factor=max(1, br_off_mask.sum()))
+
+ off_loss = (tl_off_loss + br_off_loss) / 2.0
+
+ return det_loss, pull_loss, push_loss, off_loss
+
+ def predict_by_feat(self,
+ tl_heats: List[Tensor],
+ br_heats: List[Tensor],
+ tl_embs: List[Tensor],
+ br_embs: List[Tensor],
+ tl_offs: List[Tensor],
+ br_offs: List[Tensor],
+ batch_img_metas: Optional[List[dict]] = None,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ bbox results.
+
+ Args:
+ tl_heats (list[Tensor]): Top-left corner heatmaps for each level
+ with shape (N, num_classes, H, W).
+ br_heats (list[Tensor]): Bottom-right corner heatmaps for each
+ level with shape (N, num_classes, H, W).
+ tl_embs (list[Tensor]): Top-left corner embeddings for each level
+ with shape (N, corner_emb_channels, H, W).
+ br_embs (list[Tensor]): Bottom-right corner embeddings for each
+ level with shape (N, corner_emb_channels, H, W).
+ tl_offs (list[Tensor]): Top-left corner offsets for each level
+ with shape (N, corner_offset_channels, H, W).
+ br_offs (list[Tensor]): Bottom-right corner offsets for each level
+ with shape (N, corner_offset_channels, H, W).
+ batch_img_metas (list[dict], optional): Batch image meta info.
+ Defaults to None.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ list[:obj:`InstanceData`]: Object detection results of each image
+ after the post process. Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ assert tl_heats[-1].shape[0] == br_heats[-1].shape[0] == len(
+ batch_img_metas)
+ result_list = []
+ for img_id in range(len(batch_img_metas)):
+ result_list.append(
+ self._predict_by_feat_single(
+ tl_heats[-1][img_id:img_id + 1, :],
+ br_heats[-1][img_id:img_id + 1, :],
+ tl_offs[-1][img_id:img_id + 1, :],
+ br_offs[-1][img_id:img_id + 1, :],
+ batch_img_metas[img_id],
+ tl_emb=tl_embs[-1][img_id:img_id + 1, :],
+ br_emb=br_embs[-1][img_id:img_id + 1, :],
+ rescale=rescale,
+ with_nms=with_nms))
+
+ return result_list
+
+ def _predict_by_feat_single(self,
+ tl_heat: Tensor,
+ br_heat: Tensor,
+ tl_off: Tensor,
+ br_off: Tensor,
+ img_meta: dict,
+ tl_emb: Optional[Tensor] = None,
+ br_emb: Optional[Tensor] = None,
+ tl_centripetal_shift: Optional[Tensor] = None,
+ br_centripetal_shift: Optional[Tensor] = None,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results.
+
+ Args:
+ tl_heat (Tensor): Top-left corner heatmap for current level with
+ shape (N, num_classes, H, W).
+ br_heat (Tensor): Bottom-right corner heatmap for current level
+ with shape (N, num_classes, H, W).
+ tl_off (Tensor): Top-left corner offset for current level with
+ shape (N, corner_offset_channels, H, W).
+ br_off (Tensor): Bottom-right corner offset for current level with
+ shape (N, corner_offset_channels, H, W).
+ img_meta (dict): Meta information of current image, e.g.,
+ image size, scaling factor, etc.
+ tl_emb (Tensor): Top-left corner embedding for current level with
+ shape (N, corner_emb_channels, H, W).
+ br_emb (Tensor): Bottom-right corner embedding for current level
+ with shape (N, corner_emb_channels, H, W).
+ tl_centripetal_shift: Top-left corner's centripetal shift for
+ current level with shape (N, 2, H, W).
+ br_centripetal_shift: Bottom-right corner's centripetal shift for
+ current level with shape (N, 2, H, W).
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ if isinstance(img_meta, (list, tuple)):
+ img_meta = img_meta[0]
+
+ batch_bboxes, batch_scores, batch_clses = self._decode_heatmap(
+ tl_heat=tl_heat.sigmoid(),
+ br_heat=br_heat.sigmoid(),
+ tl_off=tl_off,
+ br_off=br_off,
+ tl_emb=tl_emb,
+ br_emb=br_emb,
+ tl_centripetal_shift=tl_centripetal_shift,
+ br_centripetal_shift=br_centripetal_shift,
+ img_meta=img_meta,
+ k=self.test_cfg.corner_topk,
+ kernel=self.test_cfg.local_maximum_kernel,
+ distance_threshold=self.test_cfg.distance_threshold)
+
+ if rescale and 'scale_factor' in img_meta:
+ batch_bboxes /= batch_bboxes.new_tensor(
+ img_meta['scale_factor']).repeat((1, 2))
+
+ bboxes = batch_bboxes.view([-1, 4])
+ scores = batch_scores.view(-1)
+ clses = batch_clses.view(-1)
+
+ det_bboxes = torch.cat([bboxes, scores.unsqueeze(-1)], -1)
+ keepinds = (det_bboxes[:, -1] > -0.1)
+ det_bboxes = det_bboxes[keepinds]
+ det_labels = clses[keepinds]
+
+ if with_nms:
+ det_bboxes, det_labels = self._bboxes_nms(det_bboxes, det_labels,
+ self.test_cfg)
+
+ results = InstanceData()
+ results.bboxes = det_bboxes[..., :4]
+ results.scores = det_bboxes[..., 4]
+ results.labels = det_labels
+ return results
+
+ def _bboxes_nms(self, bboxes: Tensor, labels: Tensor,
+ cfg: ConfigDict) -> Tuple[Tensor, Tensor]:
+ """bboxes nms."""
+ if 'nms_cfg' in cfg:
+ warning.warn('nms_cfg in test_cfg will be deprecated. '
+ 'Please rename it as nms')
+ if 'nms' not in cfg:
+ cfg.nms = cfg.nms_cfg
+
+ if labels.numel() > 0:
+ max_num = cfg.max_per_img
+ bboxes, keep = batched_nms(bboxes[:, :4], bboxes[:,
+ -1].contiguous(),
+ labels, cfg.nms)
+ if max_num > 0:
+ bboxes = bboxes[:max_num]
+ labels = labels[keep][:max_num]
+
+ return bboxes, labels
+
+ def _decode_heatmap(self,
+ tl_heat: Tensor,
+ br_heat: Tensor,
+ tl_off: Tensor,
+ br_off: Tensor,
+ tl_emb: Optional[Tensor] = None,
+ br_emb: Optional[Tensor] = None,
+ tl_centripetal_shift: Optional[Tensor] = None,
+ br_centripetal_shift: Optional[Tensor] = None,
+ img_meta: Optional[dict] = None,
+ k: int = 100,
+ kernel: int = 3,
+ distance_threshold: float = 0.5,
+ num_dets: int = 1000) -> Tuple[Tensor, Tensor, Tensor]:
+ """Transform outputs into detections raw bbox prediction.
+
+ Args:
+ tl_heat (Tensor): Top-left corner heatmap for current level with
+ shape (N, num_classes, H, W).
+ br_heat (Tensor): Bottom-right corner heatmap for current level
+ with shape (N, num_classes, H, W).
+ tl_off (Tensor): Top-left corner offset for current level with
+ shape (N, corner_offset_channels, H, W).
+ br_off (Tensor): Bottom-right corner offset for current level with
+ shape (N, corner_offset_channels, H, W).
+ tl_emb (Tensor, Optional): Top-left corner embedding for current
+ level with shape (N, corner_emb_channels, H, W).
+ br_emb (Tensor, Optional): Bottom-right corner embedding for
+ current level with shape (N, corner_emb_channels, H, W).
+ tl_centripetal_shift (Tensor, Optional): Top-left centripetal shift
+ for current level with shape (N, 2, H, W).
+ br_centripetal_shift (Tensor, Optional): Bottom-right centripetal
+ shift for current level with shape (N, 2, H, W).
+ img_meta (dict): Meta information of current image, e.g.,
+ image size, scaling factor, etc.
+ k (int): Get top k corner keypoints from heatmap.
+ kernel (int): Max pooling kernel for extract local maximum pixels.
+ distance_threshold (float): Distance threshold. Top-left and
+ bottom-right corner keypoints with feature distance less than
+ the threshold will be regarded as keypoints from same object.
+ num_dets (int): Num of raw boxes before doing nms.
+
+ Returns:
+ tuple[torch.Tensor]: Decoded output of CornerHead, containing the
+ following Tensors:
+
+ - bboxes (Tensor): Coords of each box.
+ - scores (Tensor): Scores of each box.
+ - clses (Tensor): Categories of each box.
+ """
+ with_embedding = tl_emb is not None and br_emb is not None
+ with_centripetal_shift = (
+ tl_centripetal_shift is not None
+ and br_centripetal_shift is not None)
+ assert with_embedding + with_centripetal_shift == 1
+ batch, _, height, width = tl_heat.size()
+ if torch.onnx.is_in_onnx_export():
+ inp_h, inp_w = img_meta['pad_shape_for_onnx'][:2]
+ else:
+ inp_h, inp_w = img_meta['batch_input_shape'][:2]
+
+ # perform nms on heatmaps
+ tl_heat = get_local_maximum(tl_heat, kernel=kernel)
+ br_heat = get_local_maximum(br_heat, kernel=kernel)
+
+ tl_scores, tl_inds, tl_clses, tl_ys, tl_xs = get_topk_from_heatmap(
+ tl_heat, k=k)
+ br_scores, br_inds, br_clses, br_ys, br_xs = get_topk_from_heatmap(
+ br_heat, k=k)
+
+ # We use repeat instead of expand here because expand is a
+ # shallow-copy function. Thus it could cause unexpected testing result
+ # sometimes. Using expand will decrease about 10% mAP during testing
+ # compared to repeat.
+ tl_ys = tl_ys.view(batch, k, 1).repeat(1, 1, k)
+ tl_xs = tl_xs.view(batch, k, 1).repeat(1, 1, k)
+ br_ys = br_ys.view(batch, 1, k).repeat(1, k, 1)
+ br_xs = br_xs.view(batch, 1, k).repeat(1, k, 1)
+
+ tl_off = transpose_and_gather_feat(tl_off, tl_inds)
+ tl_off = tl_off.view(batch, k, 1, 2)
+ br_off = transpose_and_gather_feat(br_off, br_inds)
+ br_off = br_off.view(batch, 1, k, 2)
+
+ tl_xs = tl_xs + tl_off[..., 0]
+ tl_ys = tl_ys + tl_off[..., 1]
+ br_xs = br_xs + br_off[..., 0]
+ br_ys = br_ys + br_off[..., 1]
+
+ if with_centripetal_shift:
+ tl_centripetal_shift = transpose_and_gather_feat(
+ tl_centripetal_shift, tl_inds).view(batch, k, 1, 2).exp()
+ br_centripetal_shift = transpose_and_gather_feat(
+ br_centripetal_shift, br_inds).view(batch, 1, k, 2).exp()
+
+ tl_ctxs = tl_xs + tl_centripetal_shift[..., 0]
+ tl_ctys = tl_ys + tl_centripetal_shift[..., 1]
+ br_ctxs = br_xs - br_centripetal_shift[..., 0]
+ br_ctys = br_ys - br_centripetal_shift[..., 1]
+
+ # all possible boxes based on top k corners (ignoring class)
+ tl_xs *= (inp_w / width)
+ tl_ys *= (inp_h / height)
+ br_xs *= (inp_w / width)
+ br_ys *= (inp_h / height)
+
+ if with_centripetal_shift:
+ tl_ctxs *= (inp_w / width)
+ tl_ctys *= (inp_h / height)
+ br_ctxs *= (inp_w / width)
+ br_ctys *= (inp_h / height)
+
+ x_off, y_off = 0, 0 # no crop
+ if not torch.onnx.is_in_onnx_export():
+ # since `RandomCenterCropPad` is done on CPU with numpy and it's
+ # not dynamic traceable when exporting to ONNX, thus 'border'
+ # does not appears as key in 'img_meta'. As a tmp solution,
+ # we move this 'border' handle part to the postprocess after
+ # finished exporting to ONNX, which is handle in
+ # `mmdet/core/export/model_wrappers.py`. Though difference between
+ # pytorch and exported onnx model, it might be ignored since
+ # comparable performance is achieved between them (e.g. 40.4 vs
+ # 40.6 on COCO val2017, for CornerNet without test-time flip)
+ if 'border' in img_meta:
+ x_off = img_meta['border'][2]
+ y_off = img_meta['border'][0]
+
+ tl_xs -= x_off
+ tl_ys -= y_off
+ br_xs -= x_off
+ br_ys -= y_off
+
+ zeros = tl_xs.new_zeros(*tl_xs.size())
+ tl_xs = torch.where(tl_xs > 0.0, tl_xs, zeros)
+ tl_ys = torch.where(tl_ys > 0.0, tl_ys, zeros)
+ br_xs = torch.where(br_xs > 0.0, br_xs, zeros)
+ br_ys = torch.where(br_ys > 0.0, br_ys, zeros)
+
+ bboxes = torch.stack((tl_xs, tl_ys, br_xs, br_ys), dim=3)
+ area_bboxes = ((br_xs - tl_xs) * (br_ys - tl_ys)).abs()
+
+ if with_centripetal_shift:
+ tl_ctxs -= x_off
+ tl_ctys -= y_off
+ br_ctxs -= x_off
+ br_ctys -= y_off
+
+ tl_ctxs *= tl_ctxs.gt(0.0).type_as(tl_ctxs)
+ tl_ctys *= tl_ctys.gt(0.0).type_as(tl_ctys)
+ br_ctxs *= br_ctxs.gt(0.0).type_as(br_ctxs)
+ br_ctys *= br_ctys.gt(0.0).type_as(br_ctys)
+
+ ct_bboxes = torch.stack((tl_ctxs, tl_ctys, br_ctxs, br_ctys),
+ dim=3)
+ area_ct_bboxes = ((br_ctxs - tl_ctxs) * (br_ctys - tl_ctys)).abs()
+
+ rcentral = torch.zeros_like(ct_bboxes)
+ # magic nums from paper section 4.1
+ mu = torch.ones_like(area_bboxes) / 2.4
+ mu[area_bboxes > 3500] = 1 / 2.1 # large bbox have smaller mu
+
+ bboxes_center_x = (bboxes[..., 0] + bboxes[..., 2]) / 2
+ bboxes_center_y = (bboxes[..., 1] + bboxes[..., 3]) / 2
+ rcentral[..., 0] = bboxes_center_x - mu * (bboxes[..., 2] -
+ bboxes[..., 0]) / 2
+ rcentral[..., 1] = bboxes_center_y - mu * (bboxes[..., 3] -
+ bboxes[..., 1]) / 2
+ rcentral[..., 2] = bboxes_center_x + mu * (bboxes[..., 2] -
+ bboxes[..., 0]) / 2
+ rcentral[..., 3] = bboxes_center_y + mu * (bboxes[..., 3] -
+ bboxes[..., 1]) / 2
+ area_rcentral = ((rcentral[..., 2] - rcentral[..., 0]) *
+ (rcentral[..., 3] - rcentral[..., 1])).abs()
+ dists = area_ct_bboxes / area_rcentral
+
+ tl_ctx_inds = (ct_bboxes[..., 0] <= rcentral[..., 0]) | (
+ ct_bboxes[..., 0] >= rcentral[..., 2])
+ tl_cty_inds = (ct_bboxes[..., 1] <= rcentral[..., 1]) | (
+ ct_bboxes[..., 1] >= rcentral[..., 3])
+ br_ctx_inds = (ct_bboxes[..., 2] <= rcentral[..., 0]) | (
+ ct_bboxes[..., 2] >= rcentral[..., 2])
+ br_cty_inds = (ct_bboxes[..., 3] <= rcentral[..., 1]) | (
+ ct_bboxes[..., 3] >= rcentral[..., 3])
+
+ if with_embedding:
+ tl_emb = transpose_and_gather_feat(tl_emb, tl_inds)
+ tl_emb = tl_emb.view(batch, k, 1)
+ br_emb = transpose_and_gather_feat(br_emb, br_inds)
+ br_emb = br_emb.view(batch, 1, k)
+ dists = torch.abs(tl_emb - br_emb)
+
+ tl_scores = tl_scores.view(batch, k, 1).repeat(1, 1, k)
+ br_scores = br_scores.view(batch, 1, k).repeat(1, k, 1)
+
+ scores = (tl_scores + br_scores) / 2 # scores for all possible boxes
+
+ # tl and br should have same class
+ tl_clses = tl_clses.view(batch, k, 1).repeat(1, 1, k)
+ br_clses = br_clses.view(batch, 1, k).repeat(1, k, 1)
+ cls_inds = (tl_clses != br_clses)
+
+ # reject boxes based on distances
+ dist_inds = dists > distance_threshold
+
+ # reject boxes based on widths and heights
+ width_inds = (br_xs <= tl_xs)
+ height_inds = (br_ys <= tl_ys)
+
+ # No use `scores[cls_inds]`, instead we use `torch.where` here.
+ # Since only 1-D indices with type 'tensor(bool)' are supported
+ # when exporting to ONNX, any other bool indices with more dimensions
+ # (e.g. 2-D bool tensor) as input parameter in node is invalid
+ negative_scores = -1 * torch.ones_like(scores)
+ scores = torch.where(cls_inds, negative_scores, scores)
+ scores = torch.where(width_inds, negative_scores, scores)
+ scores = torch.where(height_inds, negative_scores, scores)
+ scores = torch.where(dist_inds, negative_scores, scores)
+
+ if with_centripetal_shift:
+ scores[tl_ctx_inds] = -1
+ scores[tl_cty_inds] = -1
+ scores[br_ctx_inds] = -1
+ scores[br_cty_inds] = -1
+
+ scores = scores.view(batch, -1)
+ scores, inds = torch.topk(scores, num_dets)
+ scores = scores.unsqueeze(2)
+
+ bboxes = bboxes.view(batch, -1, 4)
+ bboxes = gather_feat(bboxes, inds)
+
+ clses = tl_clses.contiguous().view(batch, -1, 1)
+ clses = gather_feat(clses, inds)
+
+ return bboxes, scores, clses
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/dab_detr_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/dab_detr_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..892833ffce5f17f6f9e82e67b7d32c6b9c1bafc0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/dab_detr_head.py
@@ -0,0 +1,106 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Tuple
+
+import torch.nn as nn
+from mmcv.cnn import Linear
+from mmengine.model import bias_init_with_prob, constant_init
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.utils import InstanceList
+from ..layers import MLP, inverse_sigmoid
+from .conditional_detr_head import ConditionalDETRHead
+
+
+@MODELS.register_module()
+class DABDETRHead(ConditionalDETRHead):
+ """Head of DAB-DETR. DAB-DETR: Dynamic Anchor Boxes are Better Queries for
+ DETR.
+
+ More details can be found in the `paper
+ `_ .
+ """
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the transformer head."""
+ # cls branch
+ self.fc_cls = Linear(self.embed_dims, self.cls_out_channels)
+ # reg branch
+ self.fc_reg = MLP(self.embed_dims, self.embed_dims, 4, 3)
+
+ def init_weights(self) -> None:
+ """initialize weights."""
+ if self.loss_cls.use_sigmoid:
+ bias_init = bias_init_with_prob(0.01)
+ nn.init.constant_(self.fc_cls.bias, bias_init)
+ constant_init(self.fc_reg.layers[-1], 0., bias=0.)
+
+ def forward(self, hidden_states: Tensor,
+ references: Tensor) -> Tuple[Tensor, Tensor]:
+ """"Forward function.
+
+ Args:
+ hidden_states (Tensor): Features from transformer decoder. If
+ `return_intermediate_dec` is True output has shape
+ (num_decoder_layers, bs, num_queries, dim), else has shape (1,
+ bs, num_queries, dim) which only contains the last layer
+ outputs.
+ references (Tensor): References from transformer decoder. If
+ `return_intermediate_dec` is True output has shape
+ (num_decoder_layers, bs, num_queries, 2/4), else has shape (1,
+ bs, num_queries, 2/4)
+ which only contains the last layer reference.
+ Returns:
+ tuple[Tensor]: results of head containing the following tensor.
+
+ - layers_cls_scores (Tensor): Outputs from the classification head,
+ shape (num_decoder_layers, bs, num_queries, cls_out_channels).
+ Note cls_out_channels should include background.
+ - layers_bbox_preds (Tensor): Sigmoid outputs from the regression
+ head with normalized coordinate format (cx, cy, w, h), has shape
+ (num_decoder_layers, bs, num_queries, 4).
+ """
+ layers_cls_scores = self.fc_cls(hidden_states)
+ references_before_sigmoid = inverse_sigmoid(references, eps=1e-3)
+ tmp_reg_preds = self.fc_reg(hidden_states)
+ tmp_reg_preds[..., :references_before_sigmoid.
+ size(-1)] += references_before_sigmoid
+ layers_bbox_preds = tmp_reg_preds.sigmoid()
+ return layers_cls_scores, layers_bbox_preds
+
+ def predict(self,
+ hidden_states: Tensor,
+ references: Tensor,
+ batch_data_samples: SampleList,
+ rescale: bool = True) -> InstanceList:
+ """Perform forward propagation of the detection head and predict
+ detection results on the features of the upstream network. Over-write
+ because img_metas are needed as inputs for bbox_head.
+
+ Args:
+ hidden_states (Tensor): Feature from the transformer decoder, has
+ shape (num_decoder_layers, bs, num_queries, dim).
+ references (Tensor): references from the transformer decoder, has
+ shape (num_decoder_layers, bs, num_queries, 2/4).
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool, optional): Whether to rescale the results.
+ Defaults to True.
+
+ Returns:
+ list[obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ """
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+
+ last_layer_hidden_state = hidden_states[-1].unsqueeze(0)
+ last_layer_reference = references[-1].unsqueeze(0)
+ outs = self(last_layer_hidden_state, last_layer_reference)
+
+ predictions = self.predict_by_feat(
+ *outs, batch_img_metas=batch_img_metas, rescale=rescale)
+ return predictions
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/ddod_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/ddod_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..64e91ff0135230a8d634c5964eb520e1461c872a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/ddod_head.py
@@ -0,0 +1,794 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Sequence, Tuple
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule, Scale
+from mmengine.model import bias_init_with_prob, normal_init
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures.bbox import bbox_overlaps
+from mmdet.utils import (ConfigType, InstanceList, OptConfigType,
+ OptInstanceList, reduce_mean)
+from ..task_modules.prior_generators import anchor_inside_flags
+from ..utils import images_to_levels, multi_apply, unmap
+from .anchor_head import AnchorHead
+
+EPS = 1e-12
+
+
+@MODELS.register_module()
+class DDODHead(AnchorHead):
+ """Detection Head of `DDOD `_.
+
+ DDOD head decomposes conjunctions lying in most current one-stage
+ detectors via label assignment disentanglement, spatial feature
+ disentanglement, and pyramid supervision disentanglement.
+
+ Args:
+ num_classes (int): Number of categories excluding the
+ background category.
+ in_channels (int): Number of channels in the input feature map.
+ stacked_convs (int): The number of stacked Conv. Defaults to 4.
+ conv_cfg (:obj:`ConfigDict` or dict, optional): Config dict for
+ convolution layer. Defaults to None.
+ use_dcn (bool): Use dcn, Same as ATSS when False. Defaults to True.
+ norm_cfg (:obj:`ConfigDict` or dict): Normal config of ddod head.
+ Defaults to dict(type='GN', num_groups=32, requires_grad=True).
+ loss_iou (:obj:`ConfigDict` or dict): Config of IoU loss. Defaults to
+ dict(type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0).
+ """
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: int,
+ stacked_convs: int = 4,
+ conv_cfg: OptConfigType = None,
+ use_dcn: bool = True,
+ norm_cfg: ConfigType = dict(
+ type='GN', num_groups=32, requires_grad=True),
+ loss_iou: ConfigType = dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ loss_weight=1.0),
+ **kwargs) -> None:
+ self.stacked_convs = stacked_convs
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self.use_dcn = use_dcn
+ super().__init__(num_classes, in_channels, **kwargs)
+
+ if self.train_cfg:
+ self.cls_assigner = TASK_UTILS.build(self.train_cfg['assigner'])
+ self.reg_assigner = TASK_UTILS.build(
+ self.train_cfg['reg_assigner'])
+ self.loss_iou = MODELS.build(loss_iou)
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self.relu = nn.ReLU(inplace=True)
+ self.cls_convs = nn.ModuleList()
+ self.reg_convs = nn.ModuleList()
+ for i in range(self.stacked_convs):
+ chn = self.in_channels if i == 0 else self.feat_channels
+ self.cls_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=dict(type='DCN', deform_groups=1)
+ if i == 0 and self.use_dcn else self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ self.reg_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=dict(type='DCN', deform_groups=1)
+ if i == 0 and self.use_dcn else self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ self.atss_cls = nn.Conv2d(
+ self.feat_channels,
+ self.num_base_priors * self.cls_out_channels,
+ 3,
+ padding=1)
+ self.atss_reg = nn.Conv2d(
+ self.feat_channels, self.num_base_priors * 4, 3, padding=1)
+ self.atss_iou = nn.Conv2d(
+ self.feat_channels, self.num_base_priors * 1, 3, padding=1)
+ self.scales = nn.ModuleList(
+ [Scale(1.0) for _ in self.prior_generator.strides])
+
+ # we use the global list in loss
+ self.cls_num_pos_samples_per_level = [
+ 0. for _ in range(len(self.prior_generator.strides))
+ ]
+ self.reg_num_pos_samples_per_level = [
+ 0. for _ in range(len(self.prior_generator.strides))
+ ]
+
+ def init_weights(self) -> None:
+ """Initialize weights of the head."""
+ for m in self.cls_convs:
+ normal_init(m.conv, std=0.01)
+ for m in self.reg_convs:
+ normal_init(m.conv, std=0.01)
+ normal_init(self.atss_reg, std=0.01)
+ normal_init(self.atss_iou, std=0.01)
+ bias_cls = bias_init_with_prob(0.01)
+ normal_init(self.atss_cls, std=0.01, bias=bias_cls)
+
+ def forward(self, x: Tuple[Tensor]) -> Tuple[List[Tensor]]:
+ """Forward features from the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: A tuple of classification scores, bbox predictions,
+ and iou predictions.
+
+ - cls_scores (list[Tensor]): Classification scores for all \
+ scale levels, each is a 4D-tensor, the channels number is \
+ num_base_priors * num_classes.
+ - bbox_preds (list[Tensor]): Box energies / deltas for all \
+ scale levels, each is a 4D-tensor, the channels number is \
+ num_base_priors * 4.
+ - iou_preds (list[Tensor]): IoU scores for all scale levels, \
+ each is a 4D-tensor, the channels number is num_base_priors * 1.
+ """
+ return multi_apply(self.forward_single, x, self.scales)
+
+ def forward_single(self, x: Tensor, scale: Scale) -> Sequence[Tensor]:
+ """Forward feature of a single scale level.
+
+ Args:
+ x (Tensor): Features of a single scale level.
+ scale (:obj: `mmcv.cnn.Scale`): Learnable scale module to resize
+ the bbox prediction.
+
+ Returns:
+ tuple:
+
+ - cls_score (Tensor): Cls scores for a single scale level \
+ the channels number is num_base_priors * num_classes.
+ - bbox_pred (Tensor): Box energies / deltas for a single \
+ scale level, the channels number is num_base_priors * 4.
+ - iou_pred (Tensor): Iou for a single scale level, the \
+ channel number is (N, num_base_priors * 1, H, W).
+ """
+ cls_feat = x
+ reg_feat = x
+ for cls_conv in self.cls_convs:
+ cls_feat = cls_conv(cls_feat)
+ for reg_conv in self.reg_convs:
+ reg_feat = reg_conv(reg_feat)
+ cls_score = self.atss_cls(cls_feat)
+ # we just follow atss, not apply exp in bbox_pred
+ bbox_pred = scale(self.atss_reg(reg_feat)).float()
+ iou_pred = self.atss_iou(reg_feat)
+ return cls_score, bbox_pred, iou_pred
+
+ def loss_cls_by_feat_single(self, cls_score: Tensor, labels: Tensor,
+ label_weights: Tensor,
+ reweight_factor: List[float],
+ avg_factor: float) -> Tuple[Tensor]:
+ """Compute cls loss of a single scale level.
+
+ Args:
+ cls_score (Tensor): Box scores for each scale level
+ Has shape (N, num_base_priors * num_classes, H, W).
+ labels (Tensor): Labels of each anchors with shape
+ (N, num_total_anchors).
+ label_weights (Tensor): Label weights of each anchor with shape
+ (N, num_total_anchors)
+ reweight_factor (List[float]): Reweight factor for cls and reg
+ loss.
+ avg_factor (float): Average factor that is used to average
+ the loss. When using sampling method, avg_factor is usually
+ the sum of positive and negative priors. When using
+ `PseudoSampler`, `avg_factor` is usually equal to the number
+ of positive priors.
+
+ Returns:
+ Tuple[Tensor]: A tuple of loss components.
+ """
+ cls_score = cls_score.permute(0, 2, 3, 1).reshape(
+ -1, self.cls_out_channels).contiguous()
+ labels = labels.reshape(-1)
+ label_weights = label_weights.reshape(-1)
+ loss_cls = self.loss_cls(
+ cls_score, labels, label_weights, avg_factor=avg_factor)
+ return reweight_factor * loss_cls,
+
+ def loss_reg_by_feat_single(self, anchors: Tensor, bbox_pred: Tensor,
+ iou_pred: Tensor, labels,
+ label_weights: Tensor, bbox_targets: Tensor,
+ bbox_weights: Tensor,
+ reweight_factor: List[float],
+ avg_factor: float) -> Tuple[Tensor, Tensor]:
+ """Compute reg loss of a single scale level based on the features
+ extracted by the detection head.
+
+ Args:
+ anchors (Tensor): Box reference for each scale level with shape
+ (N, num_total_anchors, 4).
+ bbox_pred (Tensor): Box energies / deltas for each scale
+ level with shape (N, num_base_priors * 4, H, W).
+ iou_pred (Tensor): Iou for a single scale level, the
+ channel number is (N, num_base_priors * 1, H, W).
+ labels (Tensor): Labels of each anchors with shape
+ (N, num_total_anchors).
+ label_weights (Tensor): Label weights of each anchor with shape
+ (N, num_total_anchors)
+ bbox_targets (Tensor): BBox regression targets of each anchor with
+ shape (N, num_total_anchors, 4).
+ bbox_weights (Tensor): BBox weights of all anchors in the
+ image with shape (N, 4)
+ reweight_factor (List[float]): Reweight factor for cls and reg
+ loss.
+ avg_factor (float): Average factor that is used to average
+ the loss. When using sampling method, avg_factor is usually
+ the sum of positive and negative priors. When using
+ `PseudoSampler`, `avg_factor` is usually equal to the number
+ of positive priors.
+ Returns:
+ Tuple[Tensor, Tensor]: A tuple of loss components.
+ """
+ anchors = anchors.reshape(-1, 4)
+ bbox_pred = bbox_pred.permute(0, 2, 3, 1).reshape(-1, 4)
+ iou_pred = iou_pred.permute(0, 2, 3, 1).reshape(-1, )
+ bbox_targets = bbox_targets.reshape(-1, 4)
+ bbox_weights = bbox_weights.reshape(-1, 4)
+ labels = labels.reshape(-1)
+ label_weights = label_weights.reshape(-1)
+
+ iou_targets = label_weights.new_zeros(labels.shape)
+ iou_weights = label_weights.new_zeros(labels.shape)
+ iou_weights[(bbox_weights.sum(axis=1) > 0).nonzero(
+ as_tuple=False)] = 1.
+
+ # FG cat_id: [0, num_classes -1], BG cat_id: num_classes
+ bg_class_ind = self.num_classes
+ pos_inds = ((labels >= 0)
+ &
+ (labels < bg_class_ind)).nonzero(as_tuple=False).squeeze(1)
+
+ if len(pos_inds) > 0:
+ pos_bbox_targets = bbox_targets[pos_inds]
+ pos_bbox_pred = bbox_pred[pos_inds]
+ pos_anchors = anchors[pos_inds]
+
+ pos_decode_bbox_pred = self.bbox_coder.decode(
+ pos_anchors, pos_bbox_pred)
+ pos_decode_bbox_targets = self.bbox_coder.decode(
+ pos_anchors, pos_bbox_targets)
+
+ # regression loss
+ loss_bbox = self.loss_bbox(
+ pos_decode_bbox_pred,
+ pos_decode_bbox_targets,
+ avg_factor=avg_factor)
+
+ iou_targets[pos_inds] = bbox_overlaps(
+ pos_decode_bbox_pred.detach(),
+ pos_decode_bbox_targets,
+ is_aligned=True)
+ loss_iou = self.loss_iou(
+ iou_pred, iou_targets, iou_weights, avg_factor=avg_factor)
+ else:
+ loss_bbox = bbox_pred.sum() * 0
+ loss_iou = iou_pred.sum() * 0
+
+ return reweight_factor * loss_bbox, reweight_factor * loss_iou
+
+ def calc_reweight_factor(self, labels_list: List[Tensor]) -> List[float]:
+ """Compute reweight_factor for regression and classification loss."""
+ # get pos samples for each level
+ bg_class_ind = self.num_classes
+ for ii, each_level_label in enumerate(labels_list):
+ pos_inds = ((each_level_label >= 0) &
+ (each_level_label < bg_class_ind)).nonzero(
+ as_tuple=False).squeeze(1)
+ self.cls_num_pos_samples_per_level[ii] += len(pos_inds)
+ # get reweight factor from 1 ~ 2 with bilinear interpolation
+ min_pos_samples = min(self.cls_num_pos_samples_per_level)
+ max_pos_samples = max(self.cls_num_pos_samples_per_level)
+ interval = 1. / (max_pos_samples - min_pos_samples + 1e-10)
+ reweight_factor_per_level = []
+ for pos_samples in self.cls_num_pos_samples_per_level:
+ factor = 2. - (pos_samples - min_pos_samples) * interval
+ reweight_factor_per_level.append(factor)
+ return reweight_factor_per_level
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ iou_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ Has shape (N, num_base_priors * num_classes, H, W)
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_base_priors * 4, H, W)
+ iou_preds (list[Tensor]): Score factor for all scale level,
+ each is a 4D-tensor, has shape (batch_size, 1, H, W).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ assert len(featmap_sizes) == self.prior_generator.num_levels
+
+ device = cls_scores[0].device
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+
+ # calculate common vars for cls and reg assigners at once
+ targets_com = self.process_predictions_and_anchors(
+ anchor_list, valid_flag_list, cls_scores, bbox_preds,
+ batch_img_metas, batch_gt_instances_ignore)
+ (anchor_list, valid_flag_list, num_level_anchors_list, cls_score_list,
+ bbox_pred_list, batch_gt_instances_ignore) = targets_com
+
+ # classification branch assigner
+ cls_targets = self.get_cls_targets(
+ anchor_list,
+ valid_flag_list,
+ num_level_anchors_list,
+ cls_score_list,
+ bbox_pred_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore)
+
+ (cls_anchor_list, labels_list, label_weights_list, bbox_targets_list,
+ bbox_weights_list, avg_factor) = cls_targets
+
+ avg_factor = reduce_mean(
+ torch.tensor(avg_factor, dtype=torch.float, device=device)).item()
+ avg_factor = max(avg_factor, 1.0)
+
+ reweight_factor_per_level = self.calc_reweight_factor(labels_list)
+
+ cls_losses_cls, = multi_apply(
+ self.loss_cls_by_feat_single,
+ cls_scores,
+ labels_list,
+ label_weights_list,
+ reweight_factor_per_level,
+ avg_factor=avg_factor)
+
+ # regression branch assigner
+ reg_targets = self.get_reg_targets(
+ anchor_list,
+ valid_flag_list,
+ num_level_anchors_list,
+ cls_score_list,
+ bbox_pred_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore)
+
+ (reg_anchor_list, labels_list, label_weights_list, bbox_targets_list,
+ bbox_weights_list, avg_factor) = reg_targets
+
+ avg_factor = reduce_mean(
+ torch.tensor(avg_factor, dtype=torch.float, device=device)).item()
+ avg_factor = max(avg_factor, 1.0)
+
+ reweight_factor_per_level = self.calc_reweight_factor(labels_list)
+
+ reg_losses_bbox, reg_losses_iou = multi_apply(
+ self.loss_reg_by_feat_single,
+ reg_anchor_list,
+ bbox_preds,
+ iou_preds,
+ labels_list,
+ label_weights_list,
+ bbox_targets_list,
+ bbox_weights_list,
+ reweight_factor_per_level,
+ avg_factor=avg_factor)
+
+ return dict(
+ loss_cls=cls_losses_cls,
+ loss_bbox=reg_losses_bbox,
+ loss_iou=reg_losses_iou)
+
+ def process_predictions_and_anchors(
+ self,
+ anchor_list: List[List[Tensor]],
+ valid_flag_list: List[List[Tensor]],
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> tuple:
+ """Compute common vars for regression and classification targets.
+
+ Args:
+ anchor_list (List[List[Tensor]]): anchors of each image.
+ valid_flag_list (List[List[Tensor]]): Valid flags of each image.
+ cls_scores (List[Tensor]): Classification scores for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_base_priors * num_classes.
+ bbox_preds (list[Tensor]): Box energies / deltas for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_base_priors * 4.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Return:
+ tuple[Tensor]: A tuple of common loss vars.
+ """
+ num_imgs = len(batch_img_metas)
+ assert len(anchor_list) == len(valid_flag_list) == num_imgs
+
+ # anchor number of multi levels
+ num_level_anchors = [anchors.size(0) for anchors in anchor_list[0]]
+ num_level_anchors_list = [num_level_anchors] * num_imgs
+
+ anchor_list_ = []
+ valid_flag_list_ = []
+ # concat all level anchors and flags to a single tensor
+ for i in range(num_imgs):
+ assert len(anchor_list[i]) == len(valid_flag_list[i])
+ anchor_list_.append(torch.cat(anchor_list[i]))
+ valid_flag_list_.append(torch.cat(valid_flag_list[i]))
+
+ # compute targets for each image
+ if batch_gt_instances_ignore is None:
+ batch_gt_instances_ignore = [None for _ in range(num_imgs)]
+
+ num_levels = len(cls_scores)
+ cls_score_list = []
+ bbox_pred_list = []
+
+ mlvl_cls_score_list = [
+ cls_score.permute(0, 2, 3, 1).reshape(
+ num_imgs, -1, self.num_base_priors * self.cls_out_channels)
+ for cls_score in cls_scores
+ ]
+ mlvl_bbox_pred_list = [
+ bbox_pred.permute(0, 2, 3, 1).reshape(num_imgs, -1,
+ self.num_base_priors * 4)
+ for bbox_pred in bbox_preds
+ ]
+
+ for i in range(num_imgs):
+ mlvl_cls_tensor_list = [
+ mlvl_cls_score_list[j][i] for j in range(num_levels)
+ ]
+ mlvl_bbox_tensor_list = [
+ mlvl_bbox_pred_list[j][i] for j in range(num_levels)
+ ]
+ cat_mlvl_cls_score = torch.cat(mlvl_cls_tensor_list, dim=0)
+ cat_mlvl_bbox_pred = torch.cat(mlvl_bbox_tensor_list, dim=0)
+ cls_score_list.append(cat_mlvl_cls_score)
+ bbox_pred_list.append(cat_mlvl_bbox_pred)
+ return (anchor_list_, valid_flag_list_, num_level_anchors_list,
+ cls_score_list, bbox_pred_list, batch_gt_instances_ignore)
+
+ def get_cls_targets(self,
+ anchor_list: List[Tensor],
+ valid_flag_list: List[Tensor],
+ num_level_anchors_list: List[int],
+ cls_score_list: List[Tensor],
+ bbox_pred_list: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None,
+ unmap_outputs: bool = True) -> tuple:
+ """Get cls targets for DDOD head.
+
+ This method is almost the same as `AnchorHead.get_targets()`.
+ Besides returning the targets as the parent method does,
+ it also returns the anchors as the first element of the
+ returned tuple.
+
+ Args:
+ anchor_list (list[Tensor]): anchors of each image.
+ valid_flag_list (list[Tensor]): Valid flags of each image.
+ num_level_anchors_list (list[Tensor]): Number of anchors of each
+ scale level of all image.
+ cls_score_list (list[Tensor]): Classification scores for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_base_priors * num_classes.
+ bbox_pred_list (list[Tensor]): Box energies / deltas for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_base_priors * 4.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ unmap_outputs (bool): Whether to map outputs back to the original
+ set of anchors.
+
+ Return:
+ tuple[Tensor]: A tuple of cls targets components.
+ """
+ (all_anchors, all_labels, all_label_weights, all_bbox_targets,
+ all_bbox_weights, pos_inds_list, neg_inds_list,
+ sampling_results_list) = multi_apply(
+ self._get_targets_single,
+ anchor_list,
+ valid_flag_list,
+ cls_score_list,
+ bbox_pred_list,
+ num_level_anchors_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore,
+ unmap_outputs=unmap_outputs,
+ is_cls_assigner=True)
+ # Get `avg_factor` of all images, which calculate in `SamplingResult`.
+ # When using sampling method, avg_factor is usually the sum of
+ # positive and negative priors. When using `PseudoSampler`,
+ # `avg_factor` is usually equal to the number of positive priors.
+ avg_factor = sum(
+ [results.avg_factor for results in sampling_results_list])
+ # split targets to a list w.r.t. multiple levels
+ anchors_list = images_to_levels(all_anchors, num_level_anchors_list[0])
+ labels_list = images_to_levels(all_labels, num_level_anchors_list[0])
+ label_weights_list = images_to_levels(all_label_weights,
+ num_level_anchors_list[0])
+ bbox_targets_list = images_to_levels(all_bbox_targets,
+ num_level_anchors_list[0])
+ bbox_weights_list = images_to_levels(all_bbox_weights,
+ num_level_anchors_list[0])
+ return (anchors_list, labels_list, label_weights_list,
+ bbox_targets_list, bbox_weights_list, avg_factor)
+
+ def get_reg_targets(self,
+ anchor_list: List[Tensor],
+ valid_flag_list: List[Tensor],
+ num_level_anchors_list: List[int],
+ cls_score_list: List[Tensor],
+ bbox_pred_list: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None,
+ unmap_outputs: bool = True) -> tuple:
+ """Get reg targets for DDOD head.
+
+ This method is almost the same as `AnchorHead.get_targets()` when
+ is_cls_assigner is False. Besides returning the targets as the parent
+ method does, it also returns the anchors as the first element of the
+ returned tuple.
+
+ Args:
+ anchor_list (list[Tensor]): anchors of each image.
+ valid_flag_list (list[Tensor]): Valid flags of each image.
+ num_level_anchors_list (list[Tensor]): Number of anchors of each
+ scale level of all image.
+ cls_score_list (list[Tensor]): Classification scores for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_base_priors * num_classes.
+ bbox_pred_list (list[Tensor]): Box energies / deltas for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_base_priors * 4.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ unmap_outputs (bool): Whether to map outputs back to the original
+ set of anchors.
+
+ Return:
+ tuple[Tensor]: A tuple of reg targets components.
+ """
+ (all_anchors, all_labels, all_label_weights, all_bbox_targets,
+ all_bbox_weights, pos_inds_list, neg_inds_list,
+ sampling_results_list) = multi_apply(
+ self._get_targets_single,
+ anchor_list,
+ valid_flag_list,
+ cls_score_list,
+ bbox_pred_list,
+ num_level_anchors_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore,
+ unmap_outputs=unmap_outputs,
+ is_cls_assigner=False)
+ # Get `avg_factor` of all images, which calculate in `SamplingResult`.
+ # When using sampling method, avg_factor is usually the sum of
+ # positive and negative priors. When using `PseudoSampler`,
+ # `avg_factor` is usually equal to the number of positive priors.
+ avg_factor = sum(
+ [results.avg_factor for results in sampling_results_list])
+ # split targets to a list w.r.t. multiple levels
+ anchors_list = images_to_levels(all_anchors, num_level_anchors_list[0])
+ labels_list = images_to_levels(all_labels, num_level_anchors_list[0])
+ label_weights_list = images_to_levels(all_label_weights,
+ num_level_anchors_list[0])
+ bbox_targets_list = images_to_levels(all_bbox_targets,
+ num_level_anchors_list[0])
+ bbox_weights_list = images_to_levels(all_bbox_weights,
+ num_level_anchors_list[0])
+ return (anchors_list, labels_list, label_weights_list,
+ bbox_targets_list, bbox_weights_list, avg_factor)
+
+ def _get_targets_single(self,
+ flat_anchors: Tensor,
+ valid_flags: Tensor,
+ cls_scores: Tensor,
+ bbox_preds: Tensor,
+ num_level_anchors: List[int],
+ gt_instances: InstanceData,
+ img_meta: dict,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ unmap_outputs: bool = True,
+ is_cls_assigner: bool = True) -> tuple:
+ """Compute regression, classification targets for anchors in a single
+ image.
+
+ Args:
+ flat_anchors (Tensor): Multi-level anchors of the image,
+ which are concatenated into a single tensor of shape
+ (num_base_priors, 4).
+ valid_flags (Tensor): Multi level valid flags of the image,
+ which are concatenated into a single tensor of
+ shape (num_base_priors,).
+ cls_scores (Tensor): Classification scores for all scale
+ levels of the image.
+ bbox_preds (Tensor): Box energies / deltas for all scale
+ levels of the image.
+ num_level_anchors (List[int]): Number of anchors of each
+ scale level.
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ img_meta (dict): Meta information for current image.
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ unmap_outputs (bool): Whether to map outputs back to the original
+ set of anchors. Defaults to True.
+ is_cls_assigner (bool): Classification or regression.
+ Defaults to True.
+
+ Returns:
+ tuple: N is the number of total anchors in the image.
+ - anchors (Tensor): all anchors in the image with shape (N, 4).
+ - labels (Tensor): Labels of all anchors in the image with \
+ shape (N, ).
+ - label_weights (Tensor): Label weights of all anchor in the \
+ image with shape (N, ).
+ - bbox_targets (Tensor): BBox targets of all anchors in the \
+ image with shape (N, 4).
+ - bbox_weights (Tensor): BBox weights of all anchors in the \
+ image with shape (N, 4)
+ - pos_inds (Tensor): Indices of positive anchor with shape \
+ (num_pos, ).
+ - neg_inds (Tensor): Indices of negative anchor with shape \
+ (num_neg, ).
+ - sampling_result (:obj:`SamplingResult`): Sampling results.
+ """
+ inside_flags = anchor_inside_flags(flat_anchors, valid_flags,
+ img_meta['img_shape'][:2],
+ self.train_cfg['allowed_border'])
+ if not inside_flags.any():
+ raise ValueError(
+ 'There is no valid anchor inside the image boundary. Please '
+ 'check the image size and anchor sizes, or set '
+ '``allowed_border`` to -1 to skip the condition.')
+ # assign gt and sample anchors
+ anchors = flat_anchors[inside_flags, :]
+
+ num_level_anchors_inside = self.get_num_level_anchors_inside(
+ num_level_anchors, inside_flags)
+ bbox_preds_valid = bbox_preds[inside_flags, :]
+ cls_scores_valid = cls_scores[inside_flags, :]
+
+ assigner = self.cls_assigner if is_cls_assigner else self.reg_assigner
+
+ # decode prediction out of assigner
+ bbox_preds_valid = self.bbox_coder.decode(anchors, bbox_preds_valid)
+ pred_instances = InstanceData(
+ priors=anchors, bboxes=bbox_preds_valid, scores=cls_scores_valid)
+
+ assign_result = assigner.assign(
+ pred_instances=pred_instances,
+ num_level_priors=num_level_anchors_inside,
+ gt_instances=gt_instances,
+ gt_instances_ignore=gt_instances_ignore)
+ sampling_result = self.sampler.sample(
+ assign_result=assign_result,
+ pred_instances=pred_instances,
+ gt_instances=gt_instances)
+
+ num_valid_anchors = anchors.shape[0]
+ bbox_targets = torch.zeros_like(anchors)
+ bbox_weights = torch.zeros_like(anchors)
+ labels = anchors.new_full((num_valid_anchors, ),
+ self.num_classes,
+ dtype=torch.long)
+ label_weights = anchors.new_zeros(num_valid_anchors, dtype=torch.float)
+
+ pos_inds = sampling_result.pos_inds
+ neg_inds = sampling_result.neg_inds
+ if len(pos_inds) > 0:
+ pos_bbox_targets = self.bbox_coder.encode(
+ sampling_result.pos_bboxes, sampling_result.pos_gt_bboxes)
+ bbox_targets[pos_inds, :] = pos_bbox_targets
+ bbox_weights[pos_inds, :] = 1.0
+
+ labels[pos_inds] = sampling_result.pos_gt_labels
+ if self.train_cfg['pos_weight'] <= 0:
+ label_weights[pos_inds] = 1.0
+ else:
+ label_weights[pos_inds] = self.train_cfg['pos_weight']
+ if len(neg_inds) > 0:
+ label_weights[neg_inds] = 1.0
+
+ # map up to original set of anchors
+ if unmap_outputs:
+ num_total_anchors = flat_anchors.size(0)
+ anchors = unmap(anchors, num_total_anchors, inside_flags)
+ labels = unmap(
+ labels, num_total_anchors, inside_flags, fill=self.num_classes)
+ label_weights = unmap(label_weights, num_total_anchors,
+ inside_flags)
+ bbox_targets = unmap(bbox_targets, num_total_anchors, inside_flags)
+ bbox_weights = unmap(bbox_weights, num_total_anchors, inside_flags)
+
+ return (anchors, labels, label_weights, bbox_targets, bbox_weights,
+ pos_inds, neg_inds, sampling_result)
+
+ def get_num_level_anchors_inside(self, num_level_anchors: List[int],
+ inside_flags: Tensor) -> List[int]:
+ """Get the anchors of each scale level inside.
+
+ Args:
+ num_level_anchors (list[int]): Number of anchors of each
+ scale level.
+ inside_flags (Tensor): Multi level inside flags of the image,
+ which are concatenated into a single tensor of
+ shape (num_base_priors,).
+
+ Returns:
+ list[int]: Number of anchors of each scale level inside.
+ """
+ split_inside_flags = torch.split(inside_flags, num_level_anchors)
+ num_level_anchors_inside = [
+ int(flags.sum()) for flags in split_inside_flags
+ ]
+ return num_level_anchors_inside
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/ddq_detr_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/ddq_detr_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..0580653ac264ea0a597eec76624ab7eb3c7f6a10
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/ddq_detr_head.py
@@ -0,0 +1,550 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from typing import Dict, List, Tuple
+
+import torch
+from mmengine.model import bias_init_with_prob, constant_init
+from torch import Tensor, nn
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import bbox_cxcywh_to_xyxy
+from mmdet.utils import InstanceList, OptInstanceList, reduce_mean
+from ..layers import inverse_sigmoid
+from ..losses import DDQAuxLoss
+from ..utils import multi_apply
+from .dino_head import DINOHead
+
+
+@MODELS.register_module()
+class DDQDETRHead(DINOHead):
+ r"""Head of DDQDETR: Dense Distinct Query for
+ End-to-End Object Detection.
+
+ Code is modified from the `official github repo
+ `_.
+
+ More details can be found in the `paper
+ `_ .
+
+ Args:
+ aux_num_pos (int): Number of positive targets assigned to a
+ perdicted object. Defaults to 4.
+ """
+
+ def __init__(self, *args, aux_num_pos=4, **kwargs):
+ super(DDQDETRHead, self).__init__(*args, **kwargs)
+ self.aux_loss_for_dense = DDQAuxLoss(
+ train_cfg=dict(
+ assigner=dict(type='TopkHungarianAssigner', topk=aux_num_pos),
+ alpha=1,
+ beta=6))
+
+ def _init_layers(self) -> None:
+ """Initialize classification branch and regression branch of aux head
+ for dense queries."""
+ super(DDQDETRHead, self)._init_layers()
+ # If decoder `num_layers` = 6 and `as_two_stage` = True, then:
+ # 1) 6 main heads are required for
+ # each decoder output of distinct queries.
+ # 2) 1 main head is required for `output_memory` of distinct queries.
+ # 3) 1 aux head is required for `output_memory` of dense queries,
+ # which is done by code below this comment.
+ # So 8 heads are required in sum.
+ # aux head for dense queries on encoder feature map
+ self.cls_branches.append(copy.deepcopy(self.cls_branches[-1]))
+ self.reg_branches.append(copy.deepcopy(self.reg_branches[-1]))
+
+ # If decoder `num_layers` = 6 and `as_two_stage` = True, then:
+ # 6 aux heads are required for each decoder output of dense queries.
+ # So 8 + 6 = 14 heads and heads are requires in sum.
+ # self.num_pred_layer is 7
+ # aux head for dense queries in decoder
+ self.aux_cls_branches = nn.ModuleList([
+ copy.deepcopy(self.cls_branches[-1])
+ for _ in range(self.num_pred_layer - 1)
+ ])
+ self.aux_reg_branches = nn.ModuleList([
+ copy.deepcopy(self.reg_branches[-1])
+ for _ in range(self.num_pred_layer - 1)
+ ])
+
+ def init_weights(self) -> None:
+ """Initialize weights of the Deformable DETR head."""
+ bias_init = bias_init_with_prob(0.01)
+ for m in self.cls_branches:
+ nn.init.constant_(m.bias, bias_init)
+ for m in self.aux_cls_branches:
+ nn.init.constant_(m.bias, bias_init)
+ for m in self.reg_branches:
+ constant_init(m[-1], 0, bias=0)
+ for m in self.reg_branches:
+ nn.init.constant_(m[-1].bias.data[2:], 0.0)
+
+ for m in self.aux_reg_branches:
+ constant_init(m[-1], 0, bias=0)
+
+ for m in self.aux_reg_branches:
+ nn.init.constant_(m[-1].bias.data[2:], 0.0)
+
+ def forward(self, hidden_states: Tensor,
+ references: List[Tensor]) -> Tuple[Tensor]:
+ """Forward function.
+
+ Args:
+ hidden_states (Tensor): Hidden states output from each decoder
+ layer, has shape (num_decoder_layers, bs, num_queries_total,
+ dim), where `num_queries_total` is the sum of
+ `num_denoising_queries`, `num_queries` and `num_dense_queries`
+ when `self.training` is `True`, else `num_queries`.
+ references (list[Tensor]): List of the reference from the decoder.
+ The first reference is the `init_reference` (initial) and the
+ other num_decoder_layers(6) references are `inter_references`
+ (intermediate). Each reference has shape (bs,
+ num_queries_total, 4) with the last dimension arranged as
+ (cx, cy, w, h).
+
+ Returns:
+ tuple[Tensor]: results of head containing the following tensors.
+
+ - all_layers_outputs_classes (Tensor): Outputs from the
+ classification head, has shape (num_decoder_layers, bs,
+ num_queries_total, cls_out_channels).
+ - all_layers_outputs_coords (Tensor): Sigmoid outputs from the
+ regression head with normalized coordinate format (cx, cy, w,
+ h), has shape (num_decoder_layers, bs, num_queries_total, 4)
+ with the last dimension arranged as (cx, cy, w, h).
+ """
+ all_layers_outputs_classes = []
+ all_layers_outputs_coords = []
+ if self.training:
+ num_dense = self.cache_dict['num_dense_queries']
+ for layer_id in range(hidden_states.shape[0]):
+ reference = inverse_sigmoid(references[layer_id])
+ hidden_state = hidden_states[layer_id]
+ if self.training:
+ dense_hidden_state = hidden_state[:, -num_dense:]
+ hidden_state = hidden_state[:, :-num_dense]
+
+ outputs_class = self.cls_branches[layer_id](hidden_state)
+ tmp_reg_preds = self.reg_branches[layer_id](hidden_state)
+ if self.training:
+ dense_outputs_class = self.aux_cls_branches[layer_id](
+ dense_hidden_state)
+ dense_tmp_reg_preds = self.aux_reg_branches[layer_id](
+ dense_hidden_state)
+ outputs_class = torch.cat([outputs_class, dense_outputs_class],
+ dim=1)
+ tmp_reg_preds = torch.cat([tmp_reg_preds, dense_tmp_reg_preds],
+ dim=1)
+
+ if reference.shape[-1] == 4:
+ tmp_reg_preds += reference
+ else:
+ assert reference.shape[-1] == 2
+ tmp_reg_preds[..., :2] += reference
+ outputs_coord = tmp_reg_preds.sigmoid()
+ all_layers_outputs_classes.append(outputs_class)
+ all_layers_outputs_coords.append(outputs_coord)
+
+ all_layers_outputs_classes = torch.stack(all_layers_outputs_classes)
+ all_layers_outputs_coords = torch.stack(all_layers_outputs_coords)
+
+ return all_layers_outputs_classes, all_layers_outputs_coords
+
+ def loss(self,
+ hidden_states: Tensor,
+ references: List[Tensor],
+ enc_outputs_class: Tensor,
+ enc_outputs_coord: Tensor,
+ batch_data_samples: SampleList,
+ dn_meta: Dict[str, int],
+ aux_enc_outputs_class=None,
+ aux_enc_outputs_coord=None) -> dict:
+ """Perform forward propagation and loss calculation of the detection
+ head on the queries of the upstream network.
+
+ Args:
+ hidden_states (Tensor): Hidden states output from each decoder
+ layer, has shape (num_decoder_layers, bs, num_queries_total,
+ dim), where `num_queries_total` is the sum of
+ `num_denoising_queries`, `num_queries` and `num_dense_queries`
+ when `self.training` is `True`, else `num_queries`.
+ references (list[Tensor]): List of the reference from the decoder.
+ The first reference is the `init_reference` (initial) and the
+ other num_decoder_layers(6) references are `inter_references`
+ (intermediate). Each reference has shape (bs,
+ num_queries_total, 4) with the last dimension arranged as
+ (cx, cy, w, h).
+ enc_outputs_class (Tensor): The top k classification score of
+ each point on encoder feature map, has shape (bs, num_queries,
+ cls_out_channels).
+ enc_outputs_coord (Tensor): The proposal generated from points
+ with top k score, has shape (bs, num_queries, 4) with the
+ last dimension arranged as (cx, cy, w, h).
+ batch_data_samples (list[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ dn_meta (Dict[str, int]): The dictionary saves information about
+ group collation, including 'num_denoising_queries' and
+ 'num_denoising_groups'. It will be used for split outputs of
+ denoising and matching parts and loss calculation.
+ aux_enc_outputs_class (Tensor): The `dense_topk` classification
+ score of each point on encoder feature map, has shape (bs,
+ num_dense_queries, cls_out_channels).
+ It is `None` when `self.training` is `False`.
+ aux_enc_outputs_coord (Tensor): The proposal generated from points
+ with `dense_topk` score, has shape (bs, num_dense_queries, 4)
+ with the last dimension arranged as (cx, cy, w, h).
+ It is `None` when `self.training` is `False`.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ batch_gt_instances = []
+ batch_img_metas = []
+ for data_sample in batch_data_samples:
+ batch_img_metas.append(data_sample.metainfo)
+ batch_gt_instances.append(data_sample.gt_instances)
+
+ outs = self(hidden_states, references)
+ loss_inputs = outs + (enc_outputs_class, enc_outputs_coord,
+ batch_gt_instances, batch_img_metas, dn_meta)
+ losses = self.loss_by_feat(*loss_inputs)
+
+ aux_enc_outputs_coord = bbox_cxcywh_to_xyxy(aux_enc_outputs_coord)
+ aux_enc_outputs_coord_list = []
+ for img_id in range(len(aux_enc_outputs_coord)):
+ det_bboxes = aux_enc_outputs_coord[img_id]
+ img_shape = batch_img_metas[img_id]['img_shape']
+ det_bboxes[:, 0::2] = det_bboxes[:, 0::2] * img_shape[1]
+ det_bboxes[:, 1::2] = det_bboxes[:, 1::2] * img_shape[0]
+ aux_enc_outputs_coord_list.append(det_bboxes)
+ aux_enc_outputs_coord = torch.stack(aux_enc_outputs_coord_list)
+ aux_loss = self.aux_loss_for_dense.loss(
+ aux_enc_outputs_class.sigmoid(), aux_enc_outputs_coord,
+ [item.bboxes for item in batch_gt_instances],
+ [item.labels for item in batch_gt_instances], batch_img_metas)
+ for k, v in aux_loss.items():
+ losses[f'aux_enc_{k}'] = v
+
+ return losses
+
+ def loss_by_feat(
+ self,
+ all_layers_cls_scores: Tensor,
+ all_layers_bbox_preds: Tensor,
+ enc_cls_scores: Tensor,
+ enc_bbox_preds: Tensor,
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ dn_meta: Dict[str, int],
+ batch_gt_instances_ignore: OptInstanceList = None
+ ) -> Dict[str, Tensor]:
+ """Loss function.
+
+ Args:
+ all_layers_cls_scores (Tensor): Classification scores of all
+ decoder layers, has shape (num_decoder_layers, bs,
+ num_queries_total, cls_out_channels).
+ all_layers_bbox_preds (Tensor): Bbox coordinates of all decoder
+ layers. Each has shape (num_decoder_layers, bs,
+ num_queries_total, 4) with normalized coordinate format
+ (cx, cy, w, h).
+ enc_cls_scores (Tensor): The top k score of each point on
+ encoder feature map, has shape (bs, num_queries,
+ cls_out_channels).
+ enc_bbox_preds (Tensor): The proposal generated from points
+ with top k score, has shape (bs, num_queries, 4) with the
+ last dimension arranged as (cx, cy, w, h).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image,
+ e.g., image size, scaling factor, etc.
+ dn_meta (Dict[str, int]): The dictionary saves information about
+ group collation, including 'num_denoising_queries' and
+ 'num_denoising_groups'. It will be used for split outputs of
+ denoising and matching parts and loss calculation.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ (all_layers_matching_cls_scores, all_layers_matching_bbox_preds,
+ all_layers_denoising_cls_scores, all_layers_denoising_bbox_preds) = \
+ self.split_outputs(
+ all_layers_cls_scores, all_layers_bbox_preds, dn_meta)
+
+ num_dense_queries = dn_meta['num_dense_queries']
+ num_layer = all_layers_matching_bbox_preds.size(0)
+ dense_all_layers_matching_cls_scores = all_layers_matching_cls_scores[:, :, # noqa: E501
+ -num_dense_queries:] # noqa: E501
+ dense_all_layers_matching_bbox_preds = all_layers_matching_bbox_preds[:, :, # noqa: E501
+ -num_dense_queries:] # noqa: E501
+
+ all_layers_matching_cls_scores = all_layers_matching_cls_scores[:, :, : # noqa: E501
+ -num_dense_queries] # noqa: E501
+ all_layers_matching_bbox_preds = all_layers_matching_bbox_preds[:, :, : # noqa: E501
+ -num_dense_queries] # noqa: E501
+
+ loss_dict = self.loss_for_distinct_queries(
+ all_layers_matching_cls_scores, all_layers_matching_bbox_preds,
+ batch_gt_instances, batch_img_metas, batch_gt_instances_ignore)
+
+ if enc_cls_scores is not None:
+
+ enc_loss_cls, enc_losses_bbox, enc_losses_iou = \
+ self.loss_by_feat_single(
+ enc_cls_scores, enc_bbox_preds,
+ batch_gt_instances=batch_gt_instances,
+ batch_img_metas=batch_img_metas)
+ loss_dict['enc_loss_cls'] = enc_loss_cls
+ loss_dict['enc_loss_bbox'] = enc_losses_bbox
+ loss_dict['enc_loss_iou'] = enc_losses_iou
+
+ if all_layers_denoising_cls_scores is not None:
+ dn_losses_cls, dn_losses_bbox, dn_losses_iou = self.loss_dn(
+ all_layers_denoising_cls_scores,
+ all_layers_denoising_bbox_preds,
+ batch_gt_instances=batch_gt_instances,
+ batch_img_metas=batch_img_metas,
+ dn_meta=dn_meta)
+ loss_dict['dn_loss_cls'] = dn_losses_cls[-1]
+ loss_dict['dn_loss_bbox'] = dn_losses_bbox[-1]
+ loss_dict['dn_loss_iou'] = dn_losses_iou[-1]
+ for num_dec_layer, (loss_cls_i, loss_bbox_i, loss_iou_i) in \
+ enumerate(zip(dn_losses_cls[:-1], dn_losses_bbox[:-1],
+ dn_losses_iou[:-1])):
+ loss_dict[f'd{num_dec_layer}.dn_loss_cls'] = loss_cls_i
+ loss_dict[f'd{num_dec_layer}.dn_loss_bbox'] = loss_bbox_i
+ loss_dict[f'd{num_dec_layer}.dn_loss_iou'] = loss_iou_i
+
+ for l_id in range(num_layer):
+ cls_scores = dense_all_layers_matching_cls_scores[l_id].sigmoid()
+ bbox_preds = dense_all_layers_matching_bbox_preds[l_id]
+
+ bbox_preds = bbox_cxcywh_to_xyxy(bbox_preds)
+ bbox_preds_list = []
+ for img_id in range(len(bbox_preds)):
+ det_bboxes = bbox_preds[img_id]
+ img_shape = batch_img_metas[img_id]['img_shape']
+ det_bboxes[:, 0::2] = det_bboxes[:, 0::2] * img_shape[1]
+ det_bboxes[:, 1::2] = det_bboxes[:, 1::2] * img_shape[0]
+ bbox_preds_list.append(det_bboxes)
+ bbox_preds = torch.stack(bbox_preds_list)
+ aux_loss = self.aux_loss_for_dense.loss(
+ cls_scores, bbox_preds,
+ [item.bboxes for item in batch_gt_instances],
+ [item.labels for item in batch_gt_instances], batch_img_metas)
+ for k, v in aux_loss.items():
+ loss_dict[f'{l_id}_aux_{k}'] = v
+
+ return loss_dict
+
+ def loss_for_distinct_queries(
+ self,
+ all_layers_cls_scores: Tensor,
+ all_layers_bbox_preds: Tensor,
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None
+ ) -> Dict[str, Tensor]:
+ """Calculate the loss of distinct queries, that is, excluding denoising
+ and dense queries. Only select the distinct queries in decoder for
+ loss.
+
+ Args:
+ all_layers_cls_scores (Tensor): Classification scores of all
+ decoder layers, has shape (num_decoder_layers, bs,
+ num_queries, cls_out_channels).
+ all_layers_bbox_preds (Tensor): Bbox coordinates of all decoder
+ layers. It has shape (num_decoder_layers, bs,
+ num_queries, 4) with the last dimension arranged as
+ (cx, cy, w, h).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image,
+ e.g., image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ assert batch_gt_instances_ignore is None, \
+ f'{self.__class__.__name__} only supports ' \
+ 'for batch_gt_instances_ignore setting to None.'
+
+ losses_cls, losses_bbox, losses_iou = multi_apply(
+ self._loss_for_distinct_queries_single,
+ all_layers_cls_scores,
+ all_layers_bbox_preds,
+ [i for i in range(len(all_layers_bbox_preds))],
+ batch_gt_instances=batch_gt_instances,
+ batch_img_metas=batch_img_metas)
+
+ loss_dict = dict()
+ # loss from the last decoder layer
+ loss_dict['loss_cls'] = losses_cls[-1]
+ loss_dict['loss_bbox'] = losses_bbox[-1]
+ loss_dict['loss_iou'] = losses_iou[-1]
+ # loss from other decoder layers
+ num_dec_layer = 0
+ for loss_cls_i, loss_bbox_i, loss_iou_i in \
+ zip(losses_cls[:-1], losses_bbox[:-1], losses_iou[:-1]):
+ loss_dict[f'd{num_dec_layer}.loss_cls'] = loss_cls_i
+ loss_dict[f'd{num_dec_layer}.loss_bbox'] = loss_bbox_i
+ loss_dict[f'd{num_dec_layer}.loss_iou'] = loss_iou_i
+ num_dec_layer += 1
+ return loss_dict
+
+ def _loss_for_distinct_queries_single(self, cls_scores, bbox_preds, l_id,
+ batch_gt_instances, batch_img_metas):
+ """Calculate the loss for outputs from a single decoder layer of
+ distinct queries, that is, excluding denoising and dense queries. Only
+ select the distinct queries in decoder for loss.
+
+ Args:
+ cls_scores (Tensor): Classification scores of a single
+ decoder layer, has shape (bs, num_queries, cls_out_channels).
+ bbox_preds (Tensor): Bbox coordinates of a single decoder
+ layer. It has shape (bs, num_queries, 4) with the last
+ dimension arranged as (cx, cy, w, h).
+ l_id (int): Decoder layer index for these outputs.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image,
+ e.g., image size, scaling factor, etc.
+
+ Returns:
+ Tuple[Tensor]: A tuple including `loss_cls`, `loss_box` and
+ `loss_iou`.
+ """
+ num_imgs = cls_scores.size(0)
+ if 0 < l_id:
+ batch_mask = [
+ self.cache_dict['distinct_query_mask'][l_id - 1][
+ img_id * self.cache_dict['num_heads']][0]
+ for img_id in range(num_imgs)
+ ]
+ else:
+ batch_mask = [
+ torch.ones(len(cls_scores[i]),
+ device=cls_scores.device).bool()
+ for i in range(num_imgs)
+ ]
+ # only select the distinct queries in decoder for loss
+ cls_scores_list = [
+ cls_scores[i][batch_mask[i]] for i in range(num_imgs)
+ ]
+ bbox_preds_list = [
+ bbox_preds[i][batch_mask[i]] for i in range(num_imgs)
+ ]
+ cls_scores = torch.cat(cls_scores_list)
+
+ cls_reg_targets = self.get_targets(cls_scores_list, bbox_preds_list,
+ batch_gt_instances, batch_img_metas)
+ (labels_list, label_weights_list, bbox_targets_list, bbox_weights_list,
+ num_total_pos, num_total_neg) = cls_reg_targets
+ labels = torch.cat(labels_list, 0)
+ label_weights = torch.cat(label_weights_list, 0)
+ bbox_targets = torch.cat(bbox_targets_list, 0)
+ bbox_weights = torch.cat(bbox_weights_list, 0)
+
+ # classification loss
+ cls_scores = cls_scores.reshape(-1, self.cls_out_channels)
+ # construct weighted avg_factor to match with the official DETR repo
+ cls_avg_factor = num_total_pos * 1.0 + \
+ num_total_neg * self.bg_cls_weight
+ if self.sync_cls_avg_factor:
+ cls_avg_factor = reduce_mean(
+ cls_scores.new_tensor([cls_avg_factor]))
+ cls_avg_factor = max(cls_avg_factor, 1)
+
+ loss_cls = self.loss_cls(
+ cls_scores, labels, label_weights, avg_factor=cls_avg_factor)
+
+ # Compute the average number of gt boxes across all gpus, for
+ # normalization purposes
+ num_total_pos = loss_cls.new_tensor([num_total_pos])
+ num_total_pos = torch.clamp(reduce_mean(num_total_pos), min=1).item()
+
+ # construct factors used for rescale bboxes
+ factors = []
+ for img_meta, bbox_pred in zip(batch_img_metas, bbox_preds_list):
+ img_h, img_w, = img_meta['img_shape']
+ factor = bbox_pred.new_tensor([img_w, img_h, img_w,
+ img_h]).unsqueeze(0).repeat(
+ bbox_pred.size(0), 1)
+ factors.append(factor)
+ factors = torch.cat(factors, 0)
+
+ # DETR regress the relative position of boxes (cxcywh) in the image,
+ # thus the learning target is normalized by the image size. So here
+ # we need to re-scale them for calculating IoU loss
+ bbox_preds = torch.cat(bbox_preds_list)
+ bbox_preds = bbox_preds.reshape(-1, 4)
+ bboxes = bbox_cxcywh_to_xyxy(bbox_preds) * factors
+ bboxes_gt = bbox_cxcywh_to_xyxy(bbox_targets) * factors
+
+ # regression IoU loss, defaultly GIoU loss
+ loss_iou = self.loss_iou(
+ bboxes, bboxes_gt, bbox_weights, avg_factor=num_total_pos)
+
+ # regression L1 loss
+ loss_bbox = self.loss_bbox(
+ bbox_preds, bbox_targets, bbox_weights, avg_factor=num_total_pos)
+ return loss_cls, loss_bbox, loss_iou
+
+ def predict_by_feat(self,
+ layer_cls_scores: Tensor,
+ layer_bbox_preds: Tensor,
+ batch_img_metas: List[dict],
+ rescale: bool = True) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ bbox results.
+
+ Args:
+ layer_cls_scores (Tensor): Classification scores of all
+ decoder layers, has shape (num_decoder_layers, bs,
+ num_queries, cls_out_channels).
+ layer_bbox_preds (Tensor): Bbox coordinates of all decoder layers.
+ Each has shape (num_decoder_layers, bs, num_queries, 4)
+ with normalized coordinate format (cx, cy, w, h).
+ batch_img_metas (list[dict]): Meta information of each image.
+ rescale (bool, optional): If `True`, return boxes in original
+ image space. Default `False`.
+
+ Returns:
+ list[obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ """
+ cls_scores = layer_cls_scores[-1]
+ bbox_preds = layer_bbox_preds[-1]
+
+ num_imgs = cls_scores.size(0)
+ # -1 is last layer input query mask
+
+ batch_mask = [
+ self.cache_dict['distinct_query_mask'][-1][
+ img_id * self.cache_dict['num_heads']][0]
+ for img_id in range(num_imgs)
+ ]
+
+ result_list = []
+ for img_id in range(len(batch_img_metas)):
+ cls_score = cls_scores[img_id][batch_mask[img_id]]
+ bbox_pred = bbox_preds[img_id][batch_mask[img_id]]
+ img_meta = batch_img_metas[img_id]
+ results = self._predict_by_feat_single(cls_score, bbox_pred,
+ img_meta, rescale)
+ result_list.append(results)
+ return result_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/deformable_detr_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/deformable_detr_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..adedd4aa6b533bcfece618eed4045c95bf0fdebb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/deformable_detr_head.py
@@ -0,0 +1,329 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from typing import Dict, List, Tuple
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import Linear
+from mmengine.model import bias_init_with_prob, constant_init
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.utils import InstanceList, OptInstanceList
+from ..layers import inverse_sigmoid
+from .detr_head import DETRHead
+
+
+@MODELS.register_module()
+class DeformableDETRHead(DETRHead):
+ r"""Head of DeformDETR: Deformable DETR: Deformable Transformers for
+ End-to-End Object Detection.
+
+ Code is modified from the `official github repo
+ `_.
+
+ More details can be found in the `paper
+ `_ .
+
+ Args:
+ share_pred_layer (bool): Whether to share parameters for all the
+ prediction layers. Defaults to `False`.
+ num_pred_layer (int): The number of the prediction layers.
+ Defaults to 6.
+ as_two_stage (bool, optional): Whether to generate the proposal
+ from the outputs of encoder. Defaults to `False`.
+ """
+
+ def __init__(self,
+ *args,
+ share_pred_layer: bool = False,
+ num_pred_layer: int = 6,
+ as_two_stage: bool = False,
+ **kwargs) -> None:
+ self.share_pred_layer = share_pred_layer
+ self.num_pred_layer = num_pred_layer
+ self.as_two_stage = as_two_stage
+
+ super().__init__(*args, **kwargs)
+
+ def _init_layers(self) -> None:
+ """Initialize classification branch and regression branch of head."""
+ fc_cls = Linear(self.embed_dims, self.cls_out_channels)
+ reg_branch = []
+ for _ in range(self.num_reg_fcs):
+ reg_branch.append(Linear(self.embed_dims, self.embed_dims))
+ reg_branch.append(nn.ReLU())
+ reg_branch.append(Linear(self.embed_dims, 4))
+ reg_branch = nn.Sequential(*reg_branch)
+
+ if self.share_pred_layer:
+ self.cls_branches = nn.ModuleList(
+ [fc_cls for _ in range(self.num_pred_layer)])
+ self.reg_branches = nn.ModuleList(
+ [reg_branch for _ in range(self.num_pred_layer)])
+ else:
+ self.cls_branches = nn.ModuleList(
+ [copy.deepcopy(fc_cls) for _ in range(self.num_pred_layer)])
+ self.reg_branches = nn.ModuleList([
+ copy.deepcopy(reg_branch) for _ in range(self.num_pred_layer)
+ ])
+
+ def init_weights(self) -> None:
+ """Initialize weights of the Deformable DETR head."""
+ if self.loss_cls.use_sigmoid:
+ bias_init = bias_init_with_prob(0.01)
+ for m in self.cls_branches:
+ if hasattr(m, 'bias') and m.bias is not None:
+ nn.init.constant_(m.bias, bias_init)
+ for m in self.reg_branches:
+ constant_init(m[-1], 0, bias=0)
+ nn.init.constant_(self.reg_branches[0][-1].bias.data[2:], -2.0)
+ if self.as_two_stage:
+ for m in self.reg_branches:
+ nn.init.constant_(m[-1].bias.data[2:], 0.0)
+
+ def forward(self, hidden_states: Tensor,
+ references: List[Tensor]) -> Tuple[Tensor, Tensor]:
+ """Forward function.
+
+ Args:
+ hidden_states (Tensor): Hidden states output from each decoder
+ layer, has shape (num_decoder_layers, bs, num_queries, dim).
+ references (list[Tensor]): List of the reference from the decoder.
+ The first reference is the `init_reference` (initial) and the
+ other num_decoder_layers(6) references are `inter_references`
+ (intermediate). The `init_reference` has shape (bs,
+ num_queries, 4) when `as_two_stage` of the detector is `True`,
+ otherwise (bs, num_queries, 2). Each `inter_reference` has
+ shape (bs, num_queries, 4) when `with_box_refine` of the
+ detector is `True`, otherwise (bs, num_queries, 2). The
+ coordinates are arranged as (cx, cy) when the last dimension is
+ 2, and (cx, cy, w, h) when it is 4.
+
+ Returns:
+ tuple[Tensor]: results of head containing the following tensor.
+
+ - all_layers_outputs_classes (Tensor): Outputs from the
+ classification head, has shape (num_decoder_layers, bs,
+ num_queries, cls_out_channels).
+ - all_layers_outputs_coords (Tensor): Sigmoid outputs from the
+ regression head with normalized coordinate format (cx, cy, w,
+ h), has shape (num_decoder_layers, bs, num_queries, 4) with the
+ last dimension arranged as (cx, cy, w, h).
+ """
+ all_layers_outputs_classes = []
+ all_layers_outputs_coords = []
+
+ for layer_id in range(hidden_states.shape[0]):
+ reference = inverse_sigmoid(references[layer_id])
+ # NOTE The last reference will not be used.
+ hidden_state = hidden_states[layer_id]
+ outputs_class = self.cls_branches[layer_id](hidden_state)
+ tmp_reg_preds = self.reg_branches[layer_id](hidden_state)
+ if reference.shape[-1] == 4:
+ # When `layer` is 0 and `as_two_stage` of the detector
+ # is `True`, or when `layer` is greater than 0 and
+ # `with_box_refine` of the detector is `True`.
+ tmp_reg_preds += reference
+ else:
+ # When `layer` is 0 and `as_two_stage` of the detector
+ # is `False`, or when `layer` is greater than 0 and
+ # `with_box_refine` of the detector is `False`.
+ assert reference.shape[-1] == 2
+ tmp_reg_preds[..., :2] += reference
+ outputs_coord = tmp_reg_preds.sigmoid()
+ all_layers_outputs_classes.append(outputs_class)
+ all_layers_outputs_coords.append(outputs_coord)
+
+ all_layers_outputs_classes = torch.stack(all_layers_outputs_classes)
+ all_layers_outputs_coords = torch.stack(all_layers_outputs_coords)
+
+ return all_layers_outputs_classes, all_layers_outputs_coords
+
+ def loss(self, hidden_states: Tensor, references: List[Tensor],
+ enc_outputs_class: Tensor, enc_outputs_coord: Tensor,
+ batch_data_samples: SampleList) -> dict:
+ """Perform forward propagation and loss calculation of the detection
+ head on the queries of the upstream network.
+
+ Args:
+ hidden_states (Tensor): Hidden states output from each decoder
+ layer, has shape (num_decoder_layers, num_queries, bs, dim).
+ references (list[Tensor]): List of the reference from the decoder.
+ The first reference is the `init_reference` (initial) and the
+ other num_decoder_layers(6) references are `inter_references`
+ (intermediate). The `init_reference` has shape (bs,
+ num_queries, 4) when `as_two_stage` of the detector is `True`,
+ otherwise (bs, num_queries, 2). Each `inter_reference` has
+ shape (bs, num_queries, 4) when `with_box_refine` of the
+ detector is `True`, otherwise (bs, num_queries, 2). The
+ coordinates are arranged as (cx, cy) when the last dimension is
+ 2, and (cx, cy, w, h) when it is 4.
+ enc_outputs_class (Tensor): The score of each point on encode
+ feature map, has shape (bs, num_feat_points, cls_out_channels).
+ Only when `as_two_stage` is `True` it would be passed in,
+ otherwise it would be `None`.
+ enc_outputs_coord (Tensor): The proposal generate from the encode
+ feature map, has shape (bs, num_feat_points, 4) with the last
+ dimension arranged as (cx, cy, w, h). Only when `as_two_stage`
+ is `True` it would be passed in, otherwise it would be `None`.
+ batch_data_samples (list[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ batch_gt_instances = []
+ batch_img_metas = []
+ for data_sample in batch_data_samples:
+ batch_img_metas.append(data_sample.metainfo)
+ batch_gt_instances.append(data_sample.gt_instances)
+
+ outs = self(hidden_states, references)
+ loss_inputs = outs + (enc_outputs_class, enc_outputs_coord,
+ batch_gt_instances, batch_img_metas)
+ losses = self.loss_by_feat(*loss_inputs)
+ return losses
+
+ def loss_by_feat(
+ self,
+ all_layers_cls_scores: Tensor,
+ all_layers_bbox_preds: Tensor,
+ enc_cls_scores: Tensor,
+ enc_bbox_preds: Tensor,
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None
+ ) -> Dict[str, Tensor]:
+ """Loss function.
+
+ Args:
+ all_layers_cls_scores (Tensor): Classification scores of all
+ decoder layers, has shape (num_decoder_layers, bs, num_queries,
+ cls_out_channels).
+ all_layers_bbox_preds (Tensor): Regression outputs of all decoder
+ layers. Each is a 4D-tensor with normalized coordinate format
+ (cx, cy, w, h) and has shape (num_decoder_layers, bs,
+ num_queries, 4) with the last dimension arranged as
+ (cx, cy, w, h).
+ enc_cls_scores (Tensor): The score of each point on encode
+ feature map, has shape (bs, num_feat_points, cls_out_channels).
+ Only when `as_two_stage` is `True` it would be passes in,
+ otherwise, it would be `None`.
+ enc_bbox_preds (Tensor): The proposal generate from the encode
+ feature map, has shape (bs, num_feat_points, 4) with the last
+ dimension arranged as (cx, cy, w, h). Only when `as_two_stage`
+ is `True` it would be passed in, otherwise it would be `None`.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ loss_dict = super().loss_by_feat(all_layers_cls_scores,
+ all_layers_bbox_preds,
+ batch_gt_instances, batch_img_metas,
+ batch_gt_instances_ignore)
+
+ # loss of proposal generated from encode feature map.
+ if enc_cls_scores is not None:
+ proposal_gt_instances = copy.deepcopy(batch_gt_instances)
+ for i in range(len(proposal_gt_instances)):
+ proposal_gt_instances[i].labels = torch.zeros_like(
+ proposal_gt_instances[i].labels)
+ enc_loss_cls, enc_losses_bbox, enc_losses_iou = \
+ self.loss_by_feat_single(
+ enc_cls_scores, enc_bbox_preds,
+ batch_gt_instances=proposal_gt_instances,
+ batch_img_metas=batch_img_metas)
+ loss_dict['enc_loss_cls'] = enc_loss_cls
+ loss_dict['enc_loss_bbox'] = enc_losses_bbox
+ loss_dict['enc_loss_iou'] = enc_losses_iou
+ return loss_dict
+
+ def predict(self,
+ hidden_states: Tensor,
+ references: List[Tensor],
+ batch_data_samples: SampleList,
+ rescale: bool = True) -> InstanceList:
+ """Perform forward propagation and loss calculation of the detection
+ head on the queries of the upstream network.
+
+ Args:
+ hidden_states (Tensor): Hidden states output from each decoder
+ layer, has shape (num_decoder_layers, num_queries, bs, dim).
+ references (list[Tensor]): List of the reference from the decoder.
+ The first reference is the `init_reference` (initial) and the
+ other num_decoder_layers(6) references are `inter_references`
+ (intermediate). The `init_reference` has shape (bs,
+ num_queries, 4) when `as_two_stage` of the detector is `True`,
+ otherwise (bs, num_queries, 2). Each `inter_reference` has
+ shape (bs, num_queries, 4) when `with_box_refine` of the
+ detector is `True`, otherwise (bs, num_queries, 2). The
+ coordinates are arranged as (cx, cy) when the last dimension is
+ 2, and (cx, cy, w, h) when it is 4.
+ batch_data_samples (list[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool, optional): If `True`, return boxes in original
+ image space. Defaults to `True`.
+
+ Returns:
+ list[obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ """
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+
+ outs = self(hidden_states, references)
+
+ predictions = self.predict_by_feat(
+ *outs, batch_img_metas=batch_img_metas, rescale=rescale)
+ return predictions
+
+ def predict_by_feat(self,
+ all_layers_cls_scores: Tensor,
+ all_layers_bbox_preds: Tensor,
+ batch_img_metas: List[Dict],
+ rescale: bool = False) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ bbox results.
+
+ Args:
+ all_layers_cls_scores (Tensor): Classification scores of all
+ decoder layers, has shape (num_decoder_layers, bs, num_queries,
+ cls_out_channels).
+ all_layers_bbox_preds (Tensor): Regression outputs of all decoder
+ layers. Each is a 4D-tensor with normalized coordinate format
+ (cx, cy, w, h) and shape (num_decoder_layers, bs, num_queries,
+ 4) with the last dimension arranged as (cx, cy, w, h).
+ batch_img_metas (list[dict]): Meta information of each image.
+ rescale (bool, optional): If `True`, return boxes in original
+ image space. Default `False`.
+
+ Returns:
+ list[obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ """
+ cls_scores = all_layers_cls_scores[-1]
+ bbox_preds = all_layers_bbox_preds[-1]
+
+ result_list = []
+ for img_id in range(len(batch_img_metas)):
+ cls_score = cls_scores[img_id]
+ bbox_pred = bbox_preds[img_id]
+ img_meta = batch_img_metas[img_id]
+ results = self._predict_by_feat_single(cls_score, bbox_pred,
+ img_meta, rescale)
+ result_list.append(results)
+ return result_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/dense_test_mixins.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/dense_test_mixins.py
new file mode 100644
index 0000000000000000000000000000000000000000..a7526d48430d6bc6b82777980d0bef418e80b91c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/dense_test_mixins.py
@@ -0,0 +1,215 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import sys
+import warnings
+from inspect import signature
+
+import torch
+from mmcv.ops import batched_nms
+from mmengine.structures import InstanceData
+
+from mmdet.structures.bbox import bbox_mapping_back
+from ..test_time_augs import merge_aug_proposals
+
+if sys.version_info >= (3, 7):
+ from mmdet.utils.contextmanagers import completed
+
+
+class BBoxTestMixin(object):
+ """Mixin class for testing det bboxes via DenseHead."""
+
+ def simple_test_bboxes(self, feats, img_metas, rescale=False):
+ """Test det bboxes without test-time augmentation, can be applied in
+ DenseHead except for ``RPNHead`` and its variants, e.g., ``GARPNHead``,
+ etc.
+
+ Args:
+ feats (tuple[torch.Tensor]): Multi-level features from the
+ upstream network, each is a 4D-tensor.
+ img_metas (list[dict]): List of image information.
+ rescale (bool, optional): Whether to rescale the results.
+ Defaults to False.
+
+ Returns:
+ list[obj:`InstanceData`]: Detection results of each
+ image after the post process. \
+ Each item usually contains following keys. \
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance,)
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances,).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ warnings.warn('You are calling `simple_test_bboxes` in '
+ '`dense_test_mixins`, but the `dense_test_mixins`'
+ 'will be deprecated soon. Please use '
+ '`simple_test` instead.')
+ outs = self.forward(feats)
+ results_list = self.get_results(
+ *outs, img_metas=img_metas, rescale=rescale)
+ return results_list
+
+ def aug_test_bboxes(self, feats, img_metas, rescale=False):
+ """Test det bboxes with test time augmentation, can be applied in
+ DenseHead except for ``RPNHead`` and its variants, e.g., ``GARPNHead``,
+ etc.
+
+ Args:
+ feats (list[Tensor]): the outer list indicates test-time
+ augmentations and inner Tensor should have a shape NxCxHxW,
+ which contains features for all images in the batch.
+ img_metas (list[list[dict]]): the outer list indicates test-time
+ augs (multiscale, flip, etc.) and the inner list indicates
+ images in a batch. each dict has image information.
+ rescale (bool, optional): Whether to rescale the results.
+ Defaults to False.
+
+ Returns:
+ list[tuple[Tensor, Tensor]]: Each item in result_list is 2-tuple.
+ The first item is ``bboxes`` with shape (n, 5),
+ where 5 represent (tl_x, tl_y, br_x, br_y, score).
+ The shape of the second tensor in the tuple is ``labels``
+ with shape (n,). The length of list should always be 1.
+ """
+
+ warnings.warn('You are calling `aug_test_bboxes` in '
+ '`dense_test_mixins`, but the `dense_test_mixins`'
+ 'will be deprecated soon. Please use '
+ '`aug_test` instead.')
+ # check with_nms argument
+ gb_sig = signature(self.get_results)
+ gb_args = [p.name for p in gb_sig.parameters.values()]
+ gbs_sig = signature(self._get_results_single)
+ gbs_args = [p.name for p in gbs_sig.parameters.values()]
+ assert ('with_nms' in gb_args) and ('with_nms' in gbs_args), \
+ f'{self.__class__.__name__}' \
+ ' does not support test-time augmentation'
+
+ aug_bboxes = []
+ aug_scores = []
+ aug_labels = []
+ for x, img_meta in zip(feats, img_metas):
+ # only one image in the batch
+ outs = self.forward(x)
+ bbox_outputs = self.get_results(
+ *outs,
+ img_metas=img_meta,
+ cfg=self.test_cfg,
+ rescale=False,
+ with_nms=False)[0]
+ aug_bboxes.append(bbox_outputs.bboxes)
+ aug_scores.append(bbox_outputs.scores)
+ if len(bbox_outputs) >= 3:
+ aug_labels.append(bbox_outputs.labels)
+
+ # after merging, bboxes will be rescaled to the original image size
+ merged_bboxes, merged_scores = self.merge_aug_bboxes(
+ aug_bboxes, aug_scores, img_metas)
+ merged_labels = torch.cat(aug_labels, dim=0) if aug_labels else None
+
+ if merged_bboxes.numel() == 0:
+ det_bboxes = torch.cat([merged_bboxes, merged_scores[:, None]], -1)
+ return [
+ (det_bboxes, merged_labels),
+ ]
+
+ det_bboxes, keep_idxs = batched_nms(merged_bboxes, merged_scores,
+ merged_labels, self.test_cfg.nms)
+ det_bboxes = det_bboxes[:self.test_cfg.max_per_img]
+ det_labels = merged_labels[keep_idxs][:self.test_cfg.max_per_img]
+
+ if rescale:
+ _det_bboxes = det_bboxes
+ else:
+ _det_bboxes = det_bboxes.clone()
+ _det_bboxes[:, :4] *= det_bboxes.new_tensor(
+ img_metas[0][0]['scale_factor'])
+
+ results = InstanceData()
+ results.bboxes = _det_bboxes[:, :4]
+ results.scores = _det_bboxes[:, 4]
+ results.labels = det_labels
+ return [results]
+
+ def aug_test_rpn(self, feats, img_metas):
+ """Test with augmentation for only for ``RPNHead`` and its variants,
+ e.g., ``GARPNHead``, etc.
+
+ Args:
+ feats (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+ img_metas (list[dict]): Meta info of each image.
+
+ Returns:
+ list[Tensor]: Proposals of each image, each item has shape (n, 5),
+ where 5 represent (tl_x, tl_y, br_x, br_y, score).
+ """
+ samples_per_gpu = len(img_metas[0])
+ aug_proposals = [[] for _ in range(samples_per_gpu)]
+ for x, img_meta in zip(feats, img_metas):
+ results_list = self.simple_test_rpn(x, img_meta)
+ for i, results in enumerate(results_list):
+ proposals = torch.cat(
+ [results.bboxes, results.scores[:, None]], dim=-1)
+ aug_proposals[i].append(proposals)
+ # reorganize the order of 'img_metas' to match the dimensions
+ # of 'aug_proposals'
+ aug_img_metas = []
+ for i in range(samples_per_gpu):
+ aug_img_meta = []
+ for j in range(len(img_metas)):
+ aug_img_meta.append(img_metas[j][i])
+ aug_img_metas.append(aug_img_meta)
+ # after merging, proposals will be rescaled to the original image size
+
+ merged_proposals = []
+ for proposals, aug_img_meta in zip(aug_proposals, aug_img_metas):
+ merged_proposal = merge_aug_proposals(proposals, aug_img_meta,
+ self.test_cfg)
+ results = InstanceData()
+ results.bboxes = merged_proposal[:, :4]
+ results.scores = merged_proposal[:, 4]
+ merged_proposals.append(results)
+ return merged_proposals
+
+ if sys.version_info >= (3, 7):
+
+ async def async_simple_test_rpn(self, x, img_metas):
+ sleep_interval = self.test_cfg.pop('async_sleep_interval', 0.025)
+ async with completed(
+ __name__, 'rpn_head_forward',
+ sleep_interval=sleep_interval):
+ rpn_outs = self(x)
+
+ proposal_list = self.get_results(*rpn_outs, img_metas=img_metas)
+ return proposal_list
+
+ def merge_aug_bboxes(self, aug_bboxes, aug_scores, img_metas):
+ """Merge augmented detection bboxes and scores.
+
+ Args:
+ aug_bboxes (list[Tensor]): shape (n, 4*#class)
+ aug_scores (list[Tensor] or None): shape (n, #class)
+ img_shapes (list[Tensor]): shape (3, ).
+
+ Returns:
+ tuple[Tensor]: ``bboxes`` with shape (n,4), where
+ 4 represent (tl_x, tl_y, br_x, br_y)
+ and ``scores`` with shape (n,).
+ """
+ recovered_bboxes = []
+ for bboxes, img_info in zip(aug_bboxes, img_metas):
+ img_shape = img_info[0]['img_shape']
+ scale_factor = img_info[0]['scale_factor']
+ flip = img_info[0]['flip']
+ flip_direction = img_info[0]['flip_direction']
+ bboxes = bbox_mapping_back(bboxes, img_shape, scale_factor, flip,
+ flip_direction)
+ recovered_bboxes.append(bboxes)
+ bboxes = torch.cat(recovered_bboxes, dim=0)
+ if aug_scores is None:
+ return bboxes
+ else:
+ scores = torch.cat(aug_scores, dim=0)
+ return bboxes, scores
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/detr_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/detr_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..9daeb4740057c1f07095ffbf97b73ea40fc93106
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/detr_head.py
@@ -0,0 +1,634 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, List, Tuple
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import Linear
+from mmcv.cnn.bricks.transformer import FFN
+from mmengine.model import BaseModule
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import (bbox_cxcywh_to_xyxy, bbox_overlaps,
+ bbox_xyxy_to_cxcywh)
+from mmdet.utils import (ConfigType, InstanceList, OptInstanceList,
+ OptMultiConfig, reduce_mean)
+from ..losses import QualityFocalLoss
+from ..utils import multi_apply
+
+
+@MODELS.register_module()
+class DETRHead(BaseModule):
+ r"""Head of DETR. DETR:End-to-End Object Detection with Transformers.
+
+ More details can be found in the `paper
+ `_ .
+
+ Args:
+ num_classes (int): Number of categories excluding the background.
+ embed_dims (int): The dims of Transformer embedding.
+ num_reg_fcs (int): Number of fully-connected layers used in `FFN`,
+ which is then used for the regression head. Defaults to 2.
+ sync_cls_avg_factor (bool): Whether to sync the `avg_factor` of
+ all ranks. Default to `False`.
+ loss_cls (:obj:`ConfigDict` or dict): Config of the classification
+ loss. Defaults to `CrossEntropyLoss`.
+ loss_bbox (:obj:`ConfigDict` or dict): Config of the regression bbox
+ loss. Defaults to `L1Loss`.
+ loss_iou (:obj:`ConfigDict` or dict): Config of the regression iou
+ loss. Defaults to `GIoULoss`.
+ train_cfg (:obj:`ConfigDict` or dict): Training config of transformer
+ head.
+ test_cfg (:obj:`ConfigDict` or dict): Testing config of transformer
+ head.
+ init_cfg (:obj:`ConfigDict` or dict, optional): the config to control
+ the initialization. Defaults to None.
+ """
+
+ _version = 2
+
+ def __init__(
+ self,
+ num_classes: int,
+ embed_dims: int = 256,
+ num_reg_fcs: int = 2,
+ sync_cls_avg_factor: bool = False,
+ loss_cls: ConfigType = dict(
+ type='CrossEntropyLoss',
+ bg_cls_weight=0.1,
+ use_sigmoid=False,
+ loss_weight=1.0,
+ class_weight=1.0),
+ loss_bbox: ConfigType = dict(type='L1Loss', loss_weight=5.0),
+ loss_iou: ConfigType = dict(type='GIoULoss', loss_weight=2.0),
+ train_cfg: ConfigType = dict(
+ assigner=dict(
+ type='HungarianAssigner',
+ match_costs=[
+ dict(type='ClassificationCost', weight=1.),
+ dict(type='BBoxL1Cost', weight=5.0, box_format='xywh'),
+ dict(type='IoUCost', iou_mode='giou', weight=2.0)
+ ])),
+ test_cfg: ConfigType = dict(max_per_img=100),
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.bg_cls_weight = 0
+ self.sync_cls_avg_factor = sync_cls_avg_factor
+ class_weight = loss_cls.get('class_weight', None)
+ if class_weight is not None and (self.__class__ is DETRHead):
+ assert isinstance(class_weight, float), 'Expected ' \
+ 'class_weight to have type float. Found ' \
+ f'{type(class_weight)}.'
+ # NOTE following the official DETR repo, bg_cls_weight means
+ # relative classification weight of the no-object class.
+ bg_cls_weight = loss_cls.get('bg_cls_weight', class_weight)
+ assert isinstance(bg_cls_weight, float), 'Expected ' \
+ 'bg_cls_weight to have type float. Found ' \
+ f'{type(bg_cls_weight)}.'
+ class_weight = torch.ones(num_classes + 1) * class_weight
+ # set background class as the last indice
+ class_weight[num_classes] = bg_cls_weight
+ loss_cls.update({'class_weight': class_weight})
+ if 'bg_cls_weight' in loss_cls:
+ loss_cls.pop('bg_cls_weight')
+ self.bg_cls_weight = bg_cls_weight
+
+ if train_cfg:
+ assert 'assigner' in train_cfg, 'assigner should be provided ' \
+ 'when train_cfg is set.'
+ assigner = train_cfg['assigner']
+ self.assigner = TASK_UTILS.build(assigner)
+ if train_cfg.get('sampler', None) is not None:
+ raise RuntimeError('DETR do not build sampler.')
+ self.num_classes = num_classes
+ self.embed_dims = embed_dims
+ self.num_reg_fcs = num_reg_fcs
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+ self.loss_cls = MODELS.build(loss_cls)
+ self.loss_bbox = MODELS.build(loss_bbox)
+ self.loss_iou = MODELS.build(loss_iou)
+
+ if self.loss_cls.use_sigmoid:
+ self.cls_out_channels = num_classes
+ else:
+ self.cls_out_channels = num_classes + 1
+
+ self._init_layers()
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the transformer head."""
+ # cls branch
+ self.fc_cls = Linear(self.embed_dims, self.cls_out_channels)
+ # reg branch
+ self.activate = nn.ReLU()
+ self.reg_ffn = FFN(
+ self.embed_dims,
+ self.embed_dims,
+ self.num_reg_fcs,
+ dict(type='ReLU', inplace=True),
+ dropout=0.0,
+ add_residual=False)
+ # NOTE the activations of reg_branch here is the same as
+ # those in transformer, but they are actually different
+ # in DAB-DETR (prelu in transformer and relu in reg_branch)
+ self.fc_reg = Linear(self.embed_dims, 4)
+
+ def forward(self, hidden_states: Tensor) -> Tuple[Tensor]:
+ """"Forward function.
+
+ Args:
+ hidden_states (Tensor): Features from transformer decoder. If
+ `return_intermediate_dec` in detr.py is True output has shape
+ (num_decoder_layers, bs, num_queries, dim), else has shape
+ (1, bs, num_queries, dim) which only contains the last layer
+ outputs.
+ Returns:
+ tuple[Tensor]: results of head containing the following tensor.
+
+ - layers_cls_scores (Tensor): Outputs from the classification head,
+ shape (num_decoder_layers, bs, num_queries, cls_out_channels).
+ Note cls_out_channels should include background.
+ - layers_bbox_preds (Tensor): Sigmoid outputs from the regression
+ head with normalized coordinate format (cx, cy, w, h), has shape
+ (num_decoder_layers, bs, num_queries, 4).
+ """
+ layers_cls_scores = self.fc_cls(hidden_states)
+ layers_bbox_preds = self.fc_reg(
+ self.activate(self.reg_ffn(hidden_states))).sigmoid()
+ return layers_cls_scores, layers_bbox_preds
+
+ def loss(self, hidden_states: Tensor,
+ batch_data_samples: SampleList) -> dict:
+ """Perform forward propagation and loss calculation of the detection
+ head on the features of the upstream network.
+
+ Args:
+ hidden_states (Tensor): Feature from the transformer decoder, has
+ shape (num_decoder_layers, bs, num_queries, cls_out_channels)
+ or (num_decoder_layers, num_queries, bs, cls_out_channels).
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ batch_gt_instances = []
+ batch_img_metas = []
+ for data_sample in batch_data_samples:
+ batch_img_metas.append(data_sample.metainfo)
+ batch_gt_instances.append(data_sample.gt_instances)
+
+ outs = self(hidden_states)
+ loss_inputs = outs + (batch_gt_instances, batch_img_metas)
+ losses = self.loss_by_feat(*loss_inputs)
+ return losses
+
+ def loss_by_feat(
+ self,
+ all_layers_cls_scores: Tensor,
+ all_layers_bbox_preds: Tensor,
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None
+ ) -> Dict[str, Tensor]:
+ """"Loss function.
+
+ Only outputs from the last feature level are used for computing
+ losses by default.
+
+ Args:
+ all_layers_cls_scores (Tensor): Classification outputs
+ of each decoder layers. Each is a 4D-tensor, has shape
+ (num_decoder_layers, bs, num_queries, cls_out_channels).
+ all_layers_bbox_preds (Tensor): Sigmoid regression
+ outputs of each decoder layers. Each is a 4D-tensor with
+ normalized coordinate format (cx, cy, w, h) and shape
+ (num_decoder_layers, bs, num_queries, 4).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ assert batch_gt_instances_ignore is None, \
+ f'{self.__class__.__name__} only supports ' \
+ 'for batch_gt_instances_ignore setting to None.'
+
+ losses_cls, losses_bbox, losses_iou = multi_apply(
+ self.loss_by_feat_single,
+ all_layers_cls_scores,
+ all_layers_bbox_preds,
+ batch_gt_instances=batch_gt_instances,
+ batch_img_metas=batch_img_metas)
+
+ loss_dict = dict()
+ # loss from the last decoder layer
+ loss_dict['loss_cls'] = losses_cls[-1]
+ loss_dict['loss_bbox'] = losses_bbox[-1]
+ loss_dict['loss_iou'] = losses_iou[-1]
+ # loss from other decoder layers
+ num_dec_layer = 0
+ for loss_cls_i, loss_bbox_i, loss_iou_i in \
+ zip(losses_cls[:-1], losses_bbox[:-1], losses_iou[:-1]):
+ loss_dict[f'd{num_dec_layer}.loss_cls'] = loss_cls_i
+ loss_dict[f'd{num_dec_layer}.loss_bbox'] = loss_bbox_i
+ loss_dict[f'd{num_dec_layer}.loss_iou'] = loss_iou_i
+ num_dec_layer += 1
+ return loss_dict
+
+ def loss_by_feat_single(self, cls_scores: Tensor, bbox_preds: Tensor,
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict]) -> Tuple[Tensor]:
+ """Loss function for outputs from a single decoder layer of a single
+ feature level.
+
+ Args:
+ cls_scores (Tensor): Box score logits from a single decoder layer
+ for all images, has shape (bs, num_queries, cls_out_channels).
+ bbox_preds (Tensor): Sigmoid outputs from a single decoder layer
+ for all images, with normalized coordinate (cx, cy, w, h) and
+ shape (bs, num_queries, 4).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+
+ Returns:
+ Tuple[Tensor]: A tuple including `loss_cls`, `loss_box` and
+ `loss_iou`.
+ """
+ num_imgs = cls_scores.size(0)
+ cls_scores_list = [cls_scores[i] for i in range(num_imgs)]
+ bbox_preds_list = [bbox_preds[i] for i in range(num_imgs)]
+ cls_reg_targets = self.get_targets(cls_scores_list, bbox_preds_list,
+ batch_gt_instances, batch_img_metas)
+ (labels_list, label_weights_list, bbox_targets_list, bbox_weights_list,
+ num_total_pos, num_total_neg) = cls_reg_targets
+ labels = torch.cat(labels_list, 0)
+ label_weights = torch.cat(label_weights_list, 0)
+ bbox_targets = torch.cat(bbox_targets_list, 0)
+ bbox_weights = torch.cat(bbox_weights_list, 0)
+
+ # classification loss
+ cls_scores = cls_scores.reshape(-1, self.cls_out_channels)
+ # construct weighted avg_factor to match with the official DETR repo
+ cls_avg_factor = num_total_pos * 1.0 + \
+ num_total_neg * self.bg_cls_weight
+ if self.sync_cls_avg_factor:
+ cls_avg_factor = reduce_mean(
+ cls_scores.new_tensor([cls_avg_factor]))
+ cls_avg_factor = max(cls_avg_factor, 1)
+
+ if isinstance(self.loss_cls, QualityFocalLoss):
+ bg_class_ind = self.num_classes
+ pos_inds = ((labels >= 0)
+ & (labels < bg_class_ind)).nonzero().squeeze(1)
+ scores = label_weights.new_zeros(labels.shape)
+ pos_bbox_targets = bbox_targets[pos_inds]
+ pos_decode_bbox_targets = bbox_cxcywh_to_xyxy(pos_bbox_targets)
+ pos_bbox_pred = bbox_preds.reshape(-1, 4)[pos_inds]
+ pos_decode_bbox_pred = bbox_cxcywh_to_xyxy(pos_bbox_pred)
+ scores[pos_inds] = bbox_overlaps(
+ pos_decode_bbox_pred.detach(),
+ pos_decode_bbox_targets,
+ is_aligned=True)
+ loss_cls = self.loss_cls(
+ cls_scores, (labels, scores),
+ label_weights,
+ avg_factor=cls_avg_factor)
+ else:
+ loss_cls = self.loss_cls(
+ cls_scores, labels, label_weights, avg_factor=cls_avg_factor)
+
+ # Compute the average number of gt boxes across all gpus, for
+ # normalization purposes
+ num_total_pos = loss_cls.new_tensor([num_total_pos])
+ num_total_pos = torch.clamp(reduce_mean(num_total_pos), min=1).item()
+
+ # construct factors used for rescale bboxes
+ factors = []
+ for img_meta, bbox_pred in zip(batch_img_metas, bbox_preds):
+ img_h, img_w, = img_meta['img_shape']
+ factor = bbox_pred.new_tensor([img_w, img_h, img_w,
+ img_h]).unsqueeze(0).repeat(
+ bbox_pred.size(0), 1)
+ factors.append(factor)
+ factors = torch.cat(factors, 0)
+
+ # DETR regress the relative position of boxes (cxcywh) in the image,
+ # thus the learning target is normalized by the image size. So here
+ # we need to re-scale them for calculating IoU loss
+ bbox_preds = bbox_preds.reshape(-1, 4)
+ bboxes = bbox_cxcywh_to_xyxy(bbox_preds) * factors
+ bboxes_gt = bbox_cxcywh_to_xyxy(bbox_targets) * factors
+
+ # regression IoU loss, defaultly GIoU loss
+ loss_iou = self.loss_iou(
+ bboxes, bboxes_gt, bbox_weights, avg_factor=num_total_pos)
+
+ # regression L1 loss
+ loss_bbox = self.loss_bbox(
+ bbox_preds, bbox_targets, bbox_weights, avg_factor=num_total_pos)
+ return loss_cls, loss_bbox, loss_iou
+
+ def get_targets(self, cls_scores_list: List[Tensor],
+ bbox_preds_list: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict]) -> tuple:
+ """Compute regression and classification targets for a batch image.
+
+ Outputs from a single decoder layer of a single feature level are used.
+
+ Args:
+ cls_scores_list (list[Tensor]): Box score logits from a single
+ decoder layer for each image, has shape [num_queries,
+ cls_out_channels].
+ bbox_preds_list (list[Tensor]): Sigmoid outputs from a single
+ decoder layer for each image, with normalized coordinate
+ (cx, cy, w, h) and shape [num_queries, 4].
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+
+ Returns:
+ tuple: a tuple containing the following targets.
+
+ - labels_list (list[Tensor]): Labels for all images.
+ - label_weights_list (list[Tensor]): Label weights for all images.
+ - bbox_targets_list (list[Tensor]): BBox targets for all images.
+ - bbox_weights_list (list[Tensor]): BBox weights for all images.
+ - num_total_pos (int): Number of positive samples in all images.
+ - num_total_neg (int): Number of negative samples in all images.
+ """
+ (labels_list, label_weights_list, bbox_targets_list, bbox_weights_list,
+ pos_inds_list,
+ neg_inds_list) = multi_apply(self._get_targets_single,
+ cls_scores_list, bbox_preds_list,
+ batch_gt_instances, batch_img_metas)
+ num_total_pos = sum((inds.numel() for inds in pos_inds_list))
+ num_total_neg = sum((inds.numel() for inds in neg_inds_list))
+ return (labels_list, label_weights_list, bbox_targets_list,
+ bbox_weights_list, num_total_pos, num_total_neg)
+
+ def _get_targets_single(self, cls_score: Tensor, bbox_pred: Tensor,
+ gt_instances: InstanceData,
+ img_meta: dict) -> tuple:
+ """Compute regression and classification targets for one image.
+
+ Outputs from a single decoder layer of a single feature level are used.
+
+ Args:
+ cls_score (Tensor): Box score logits from a single decoder layer
+ for one image. Shape [num_queries, cls_out_channels].
+ bbox_pred (Tensor): Sigmoid outputs from a single decoder layer
+ for one image, with normalized coordinate (cx, cy, w, h) and
+ shape [num_queries, 4].
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes`` and ``labels``
+ attributes.
+ img_meta (dict): Meta information for one image.
+
+ Returns:
+ tuple[Tensor]: a tuple containing the following for one image.
+
+ - labels (Tensor): Labels of each image.
+ - label_weights (Tensor]): Label weights of each image.
+ - bbox_targets (Tensor): BBox targets of each image.
+ - bbox_weights (Tensor): BBox weights of each image.
+ - pos_inds (Tensor): Sampled positive indices for each image.
+ - neg_inds (Tensor): Sampled negative indices for each image.
+ """
+ img_h, img_w = img_meta['img_shape']
+ factor = bbox_pred.new_tensor([img_w, img_h, img_w,
+ img_h]).unsqueeze(0)
+ num_bboxes = bbox_pred.size(0)
+ # convert bbox_pred from xywh, normalized to xyxy, unnormalized
+ bbox_pred = bbox_cxcywh_to_xyxy(bbox_pred)
+ bbox_pred = bbox_pred * factor
+
+ pred_instances = InstanceData(scores=cls_score, bboxes=bbox_pred)
+ # assigner and sampler
+ assign_result = self.assigner.assign(
+ pred_instances=pred_instances,
+ gt_instances=gt_instances,
+ img_meta=img_meta)
+
+ gt_bboxes = gt_instances.bboxes
+ gt_labels = gt_instances.labels
+ pos_inds = torch.nonzero(
+ assign_result.gt_inds > 0, as_tuple=False).squeeze(-1).unique()
+ neg_inds = torch.nonzero(
+ assign_result.gt_inds == 0, as_tuple=False).squeeze(-1).unique()
+ pos_assigned_gt_inds = assign_result.gt_inds[pos_inds] - 1
+ pos_gt_bboxes = gt_bboxes[pos_assigned_gt_inds.long(), :]
+
+ # label targets
+ labels = gt_bboxes.new_full((num_bboxes, ),
+ self.num_classes,
+ dtype=torch.long)
+ labels[pos_inds] = gt_labels[pos_assigned_gt_inds]
+ label_weights = gt_bboxes.new_ones(num_bboxes)
+
+ # bbox targets
+ bbox_targets = torch.zeros_like(bbox_pred, dtype=gt_bboxes.dtype)
+ bbox_weights = torch.zeros_like(bbox_pred, dtype=gt_bboxes.dtype)
+ bbox_weights[pos_inds] = 1.0
+
+ # DETR regress the relative position of boxes (cxcywh) in the image.
+ # Thus the learning target should be normalized by the image size, also
+ # the box format should be converted from defaultly x1y1x2y2 to cxcywh.
+ pos_gt_bboxes_normalized = pos_gt_bboxes / factor
+ pos_gt_bboxes_targets = bbox_xyxy_to_cxcywh(pos_gt_bboxes_normalized)
+ bbox_targets[pos_inds] = pos_gt_bboxes_targets
+ return (labels, label_weights, bbox_targets, bbox_weights, pos_inds,
+ neg_inds)
+
+ def loss_and_predict(
+ self, hidden_states: Tuple[Tensor],
+ batch_data_samples: SampleList) -> Tuple[dict, InstanceList]:
+ """Perform forward propagation of the head, then calculate loss and
+ predictions from the features and data samples. Over-write because
+ img_metas are needed as inputs for bbox_head.
+
+ Args:
+ hidden_states (tuple[Tensor]): Feature from the transformer
+ decoder, has shape (num_decoder_layers, bs, num_queries, dim).
+ batch_data_samples (list[:obj:`DetDataSample`]): Each item contains
+ the meta information of each image and corresponding
+ annotations.
+
+ Returns:
+ tuple: the return value is a tuple contains:
+
+ - losses: (dict[str, Tensor]): A dictionary of loss components.
+ - predictions (list[:obj:`InstanceData`]): Detection
+ results of each image after the post process.
+ """
+ batch_gt_instances = []
+ batch_img_metas = []
+ for data_sample in batch_data_samples:
+ batch_img_metas.append(data_sample.metainfo)
+ batch_gt_instances.append(data_sample.gt_instances)
+
+ outs = self(hidden_states)
+ loss_inputs = outs + (batch_gt_instances, batch_img_metas)
+ losses = self.loss_by_feat(*loss_inputs)
+
+ predictions = self.predict_by_feat(
+ *outs, batch_img_metas=batch_img_metas)
+ return losses, predictions
+
+ def predict(self,
+ hidden_states: Tuple[Tensor],
+ batch_data_samples: SampleList,
+ rescale: bool = True) -> InstanceList:
+ """Perform forward propagation of the detection head and predict
+ detection results on the features of the upstream network. Over-write
+ because img_metas are needed as inputs for bbox_head.
+
+ Args:
+ hidden_states (tuple[Tensor]): Multi-level features from the
+ upstream network, each is a 4D-tensor.
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool, optional): Whether to rescale the results.
+ Defaults to True.
+
+ Returns:
+ list[obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ """
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+
+ last_layer_hidden_state = hidden_states[-1].unsqueeze(0)
+ outs = self(last_layer_hidden_state)
+
+ predictions = self.predict_by_feat(
+ *outs, batch_img_metas=batch_img_metas, rescale=rescale)
+
+ return predictions
+
+ def predict_by_feat(self,
+ layer_cls_scores: Tensor,
+ layer_bbox_preds: Tensor,
+ batch_img_metas: List[dict],
+ rescale: bool = True) -> InstanceList:
+ """Transform network outputs for a batch into bbox predictions.
+
+ Args:
+ layer_cls_scores (Tensor): Classification outputs of the last or
+ all decoder layer. Each is a 4D-tensor, has shape
+ (num_decoder_layers, bs, num_queries, cls_out_channels).
+ layer_bbox_preds (Tensor): Sigmoid regression outputs of the last
+ or all decoder layer. Each is a 4D-tensor with normalized
+ coordinate format (cx, cy, w, h) and shape
+ (num_decoder_layers, bs, num_queries, 4).
+ batch_img_metas (list[dict]): Meta information of each image.
+ rescale (bool, optional): If `True`, return boxes in original
+ image space. Defaults to `True`.
+
+ Returns:
+ list[:obj:`InstanceData`]: Object detection results of each image
+ after the post process. Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ # NOTE only using outputs from the last feature level,
+ # and only the outputs from the last decoder layer is used.
+ cls_scores = layer_cls_scores[-1]
+ bbox_preds = layer_bbox_preds[-1]
+
+ result_list = []
+ for img_id in range(len(batch_img_metas)):
+ cls_score = cls_scores[img_id]
+ bbox_pred = bbox_preds[img_id]
+ img_meta = batch_img_metas[img_id]
+ results = self._predict_by_feat_single(cls_score, bbox_pred,
+ img_meta, rescale)
+ result_list.append(results)
+ return result_list
+
+ def _predict_by_feat_single(self,
+ cls_score: Tensor,
+ bbox_pred: Tensor,
+ img_meta: dict,
+ rescale: bool = True) -> InstanceData:
+ """Transform outputs from the last decoder layer into bbox predictions
+ for each image.
+
+ Args:
+ cls_score (Tensor): Box score logits from the last decoder layer
+ for each image. Shape [num_queries, cls_out_channels].
+ bbox_pred (Tensor): Sigmoid outputs from the last decoder layer
+ for each image, with coordinate format (cx, cy, w, h) and
+ shape [num_queries, 4].
+ img_meta (dict): Image meta info.
+ rescale (bool): If True, return boxes in original image
+ space. Default True.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ assert len(cls_score) == len(bbox_pred) # num_queries
+ max_per_img = self.test_cfg.get('max_per_img', len(cls_score))
+ img_shape = img_meta['img_shape']
+ # exclude background
+ if self.loss_cls.use_sigmoid:
+ cls_score = cls_score.sigmoid()
+ scores, indexes = cls_score.view(-1).topk(max_per_img)
+ det_labels = indexes % self.num_classes
+ bbox_index = indexes // self.num_classes
+ bbox_pred = bbox_pred[bbox_index]
+ else:
+ scores, det_labels = F.softmax(cls_score, dim=-1)[..., :-1].max(-1)
+ scores, bbox_index = scores.topk(max_per_img)
+ bbox_pred = bbox_pred[bbox_index]
+ det_labels = det_labels[bbox_index]
+
+ det_bboxes = bbox_cxcywh_to_xyxy(bbox_pred)
+ det_bboxes[:, 0::2] = det_bboxes[:, 0::2] * img_shape[1]
+ det_bboxes[:, 1::2] = det_bboxes[:, 1::2] * img_shape[0]
+ det_bboxes[:, 0::2].clamp_(min=0, max=img_shape[1])
+ det_bboxes[:, 1::2].clamp_(min=0, max=img_shape[0])
+ if rescale:
+ assert img_meta.get('scale_factor') is not None
+ det_bboxes /= det_bboxes.new_tensor(
+ img_meta['scale_factor']).repeat((1, 2))
+
+ results = InstanceData()
+ results.bboxes = det_bboxes
+ results.scores = scores
+ results.labels = det_labels
+ return results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/dino_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/dino_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..54f46d1474f97f2d183926a6dc68a0be79f7cef1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/dino_head.py
@@ -0,0 +1,479 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, List, Tuple
+
+import torch
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import (bbox_cxcywh_to_xyxy, bbox_overlaps,
+ bbox_xyxy_to_cxcywh)
+from mmdet.utils import InstanceList, OptInstanceList, reduce_mean
+from ..losses import QualityFocalLoss
+from ..utils import multi_apply
+from .deformable_detr_head import DeformableDETRHead
+
+
+@MODELS.register_module()
+class DINOHead(DeformableDETRHead):
+ r"""Head of the DINO: DETR with Improved DeNoising Anchor Boxes
+ for End-to-End Object Detection
+
+ Code is modified from the `official github repo
+ `_.
+
+ More details can be found in the `paper
+ `_ .
+ """
+
+ def loss(self, hidden_states: Tensor, references: List[Tensor],
+ enc_outputs_class: Tensor, enc_outputs_coord: Tensor,
+ batch_data_samples: SampleList, dn_meta: Dict[str, int]) -> dict:
+ """Perform forward propagation and loss calculation of the detection
+ head on the queries of the upstream network.
+
+ Args:
+ hidden_states (Tensor): Hidden states output from each decoder
+ layer, has shape (num_decoder_layers, bs, num_queries_total,
+ dim), where `num_queries_total` is the sum of
+ `num_denoising_queries` and `num_matching_queries` when
+ `self.training` is `True`, else `num_matching_queries`.
+ references (list[Tensor]): List of the reference from the decoder.
+ The first reference is the `init_reference` (initial) and the
+ other num_decoder_layers(6) references are `inter_references`
+ (intermediate). The `init_reference` has shape (bs,
+ num_queries_total, 4) and each `inter_reference` has shape
+ (bs, num_queries, 4) with the last dimension arranged as
+ (cx, cy, w, h).
+ enc_outputs_class (Tensor): The score of each point on encode
+ feature map, has shape (bs, num_feat_points, cls_out_channels).
+ enc_outputs_coord (Tensor): The proposal generate from the
+ encode feature map, has shape (bs, num_feat_points, 4) with the
+ last dimension arranged as (cx, cy, w, h).
+ batch_data_samples (list[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ dn_meta (Dict[str, int]): The dictionary saves information about
+ group collation, including 'num_denoising_queries' and
+ 'num_denoising_groups'. It will be used for split outputs of
+ denoising and matching parts and loss calculation.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ batch_gt_instances = []
+ batch_img_metas = []
+ for data_sample in batch_data_samples:
+ batch_img_metas.append(data_sample.metainfo)
+ batch_gt_instances.append(data_sample.gt_instances)
+
+ outs = self(hidden_states, references)
+ loss_inputs = outs + (enc_outputs_class, enc_outputs_coord,
+ batch_gt_instances, batch_img_metas, dn_meta)
+ losses = self.loss_by_feat(*loss_inputs)
+ return losses
+
+ def loss_by_feat(
+ self,
+ all_layers_cls_scores: Tensor,
+ all_layers_bbox_preds: Tensor,
+ enc_cls_scores: Tensor,
+ enc_bbox_preds: Tensor,
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ dn_meta: Dict[str, int],
+ batch_gt_instances_ignore: OptInstanceList = None
+ ) -> Dict[str, Tensor]:
+ """Loss function.
+
+ Args:
+ all_layers_cls_scores (Tensor): Classification scores of all
+ decoder layers, has shape (num_decoder_layers, bs,
+ num_queries_total, cls_out_channels), where
+ `num_queries_total` is the sum of `num_denoising_queries`
+ and `num_matching_queries`.
+ all_layers_bbox_preds (Tensor): Regression outputs of all decoder
+ layers. Each is a 4D-tensor with normalized coordinate format
+ (cx, cy, w, h) and has shape (num_decoder_layers, bs,
+ num_queries_total, 4).
+ enc_cls_scores (Tensor): The score of each point on encode
+ feature map, has shape (bs, num_feat_points, cls_out_channels).
+ enc_bbox_preds (Tensor): The proposal generate from the encode
+ feature map, has shape (bs, num_feat_points, 4) with the last
+ dimension arranged as (cx, cy, w, h).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ dn_meta (Dict[str, int]): The dictionary saves information about
+ group collation, including 'num_denoising_queries' and
+ 'num_denoising_groups'. It will be used for split outputs of
+ denoising and matching parts and loss calculation.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ # extract denoising and matching part of outputs
+ (all_layers_matching_cls_scores, all_layers_matching_bbox_preds,
+ all_layers_denoising_cls_scores, all_layers_denoising_bbox_preds) = \
+ self.split_outputs(
+ all_layers_cls_scores, all_layers_bbox_preds, dn_meta)
+
+ loss_dict = super(DeformableDETRHead, self).loss_by_feat(
+ all_layers_matching_cls_scores, all_layers_matching_bbox_preds,
+ batch_gt_instances, batch_img_metas, batch_gt_instances_ignore)
+ # NOTE DETRHead.loss_by_feat but not DeformableDETRHead.loss_by_feat
+ # is called, because the encoder loss calculations are different
+ # between DINO and DeformableDETR.
+
+ # loss of proposal generated from encode feature map.
+ if enc_cls_scores is not None:
+ # NOTE The enc_loss calculation of the DINO is
+ # different from that of Deformable DETR.
+ enc_loss_cls, enc_losses_bbox, enc_losses_iou = \
+ self.loss_by_feat_single(
+ enc_cls_scores, enc_bbox_preds,
+ batch_gt_instances=batch_gt_instances,
+ batch_img_metas=batch_img_metas)
+ loss_dict['enc_loss_cls'] = enc_loss_cls
+ loss_dict['enc_loss_bbox'] = enc_losses_bbox
+ loss_dict['enc_loss_iou'] = enc_losses_iou
+
+ if all_layers_denoising_cls_scores is not None:
+ # calculate denoising loss from all decoder layers
+ dn_losses_cls, dn_losses_bbox, dn_losses_iou = self.loss_dn(
+ all_layers_denoising_cls_scores,
+ all_layers_denoising_bbox_preds,
+ batch_gt_instances=batch_gt_instances,
+ batch_img_metas=batch_img_metas,
+ dn_meta=dn_meta)
+ # collate denoising loss
+ loss_dict['dn_loss_cls'] = dn_losses_cls[-1]
+ loss_dict['dn_loss_bbox'] = dn_losses_bbox[-1]
+ loss_dict['dn_loss_iou'] = dn_losses_iou[-1]
+ for num_dec_layer, (loss_cls_i, loss_bbox_i, loss_iou_i) in \
+ enumerate(zip(dn_losses_cls[:-1], dn_losses_bbox[:-1],
+ dn_losses_iou[:-1])):
+ loss_dict[f'd{num_dec_layer}.dn_loss_cls'] = loss_cls_i
+ loss_dict[f'd{num_dec_layer}.dn_loss_bbox'] = loss_bbox_i
+ loss_dict[f'd{num_dec_layer}.dn_loss_iou'] = loss_iou_i
+ return loss_dict
+
+ def loss_dn(self, all_layers_denoising_cls_scores: Tensor,
+ all_layers_denoising_bbox_preds: Tensor,
+ batch_gt_instances: InstanceList, batch_img_metas: List[dict],
+ dn_meta: Dict[str, int]) -> Tuple[List[Tensor]]:
+ """Calculate denoising loss.
+
+ Args:
+ all_layers_denoising_cls_scores (Tensor): Classification scores of
+ all decoder layers in denoising part, has shape (
+ num_decoder_layers, bs, num_denoising_queries,
+ cls_out_channels).
+ all_layers_denoising_bbox_preds (Tensor): Regression outputs of all
+ decoder layers in denoising part. Each is a 4D-tensor with
+ normalized coordinate format (cx, cy, w, h) and has shape
+ (num_decoder_layers, bs, num_denoising_queries, 4).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ dn_meta (Dict[str, int]): The dictionary saves information about
+ group collation, including 'num_denoising_queries' and
+ 'num_denoising_groups'. It will be used for split outputs of
+ denoising and matching parts and loss calculation.
+
+ Returns:
+ Tuple[List[Tensor]]: The loss_dn_cls, loss_dn_bbox, and loss_dn_iou
+ of each decoder layers.
+ """
+ return multi_apply(
+ self._loss_dn_single,
+ all_layers_denoising_cls_scores,
+ all_layers_denoising_bbox_preds,
+ batch_gt_instances=batch_gt_instances,
+ batch_img_metas=batch_img_metas,
+ dn_meta=dn_meta)
+
+ def _loss_dn_single(self, dn_cls_scores: Tensor, dn_bbox_preds: Tensor,
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ dn_meta: Dict[str, int]) -> Tuple[Tensor]:
+ """Denoising loss for outputs from a single decoder layer.
+
+ Args:
+ dn_cls_scores (Tensor): Classification scores of a single decoder
+ layer in denoising part, has shape (bs, num_denoising_queries,
+ cls_out_channels).
+ dn_bbox_preds (Tensor): Regression outputs of a single decoder
+ layer in denoising part. Each is a 4D-tensor with normalized
+ coordinate format (cx, cy, w, h) and has shape
+ (bs, num_denoising_queries, 4).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ dn_meta (Dict[str, int]): The dictionary saves information about
+ group collation, including 'num_denoising_queries' and
+ 'num_denoising_groups'. It will be used for split outputs of
+ denoising and matching parts and loss calculation.
+
+ Returns:
+ Tuple[Tensor]: A tuple including `loss_cls`, `loss_box` and
+ `loss_iou`.
+ """
+ cls_reg_targets = self.get_dn_targets(batch_gt_instances,
+ batch_img_metas, dn_meta)
+ (labels_list, label_weights_list, bbox_targets_list, bbox_weights_list,
+ num_total_pos, num_total_neg) = cls_reg_targets
+ labels = torch.cat(labels_list, 0)
+ label_weights = torch.cat(label_weights_list, 0)
+ bbox_targets = torch.cat(bbox_targets_list, 0)
+ bbox_weights = torch.cat(bbox_weights_list, 0)
+
+ # classification loss
+ cls_scores = dn_cls_scores.reshape(-1, self.cls_out_channels)
+ # construct weighted avg_factor to match with the official DETR repo
+ cls_avg_factor = \
+ num_total_pos * 1.0 + num_total_neg * self.bg_cls_weight
+ if self.sync_cls_avg_factor:
+ cls_avg_factor = reduce_mean(
+ cls_scores.new_tensor([cls_avg_factor]))
+ cls_avg_factor = max(cls_avg_factor, 1)
+
+ if len(cls_scores) > 0:
+ if isinstance(self.loss_cls, QualityFocalLoss):
+ bg_class_ind = self.num_classes
+ pos_inds = ((labels >= 0)
+ & (labels < bg_class_ind)).nonzero().squeeze(1)
+ scores = label_weights.new_zeros(labels.shape)
+ pos_bbox_targets = bbox_targets[pos_inds]
+ pos_decode_bbox_targets = bbox_cxcywh_to_xyxy(pos_bbox_targets)
+ pos_bbox_pred = dn_bbox_preds.reshape(-1, 4)[pos_inds]
+ pos_decode_bbox_pred = bbox_cxcywh_to_xyxy(pos_bbox_pred)
+ scores[pos_inds] = bbox_overlaps(
+ pos_decode_bbox_pred.detach(),
+ pos_decode_bbox_targets,
+ is_aligned=True)
+ loss_cls = self.loss_cls(
+ cls_scores, (labels, scores),
+ weight=label_weights,
+ avg_factor=cls_avg_factor)
+ else:
+ loss_cls = self.loss_cls(
+ cls_scores,
+ labels,
+ label_weights,
+ avg_factor=cls_avg_factor)
+ else:
+ loss_cls = torch.zeros(
+ 1, dtype=cls_scores.dtype, device=cls_scores.device)
+
+ # Compute the average number of gt boxes across all gpus, for
+ # normalization purposes
+ num_total_pos = loss_cls.new_tensor([num_total_pos])
+ num_total_pos = torch.clamp(reduce_mean(num_total_pos), min=1).item()
+
+ # construct factors used for rescale bboxes
+ factors = []
+ for img_meta, bbox_pred in zip(batch_img_metas, dn_bbox_preds):
+ img_h, img_w = img_meta['img_shape']
+ factor = bbox_pred.new_tensor([img_w, img_h, img_w,
+ img_h]).unsqueeze(0).repeat(
+ bbox_pred.size(0), 1)
+ factors.append(factor)
+ factors = torch.cat(factors)
+
+ # DETR regress the relative position of boxes (cxcywh) in the image,
+ # thus the learning target is normalized by the image size. So here
+ # we need to re-scale them for calculating IoU loss
+ bbox_preds = dn_bbox_preds.reshape(-1, 4)
+ bboxes = bbox_cxcywh_to_xyxy(bbox_preds) * factors
+ bboxes_gt = bbox_cxcywh_to_xyxy(bbox_targets) * factors
+
+ # regression IoU loss, defaultly GIoU loss
+ loss_iou = self.loss_iou(
+ bboxes, bboxes_gt, bbox_weights, avg_factor=num_total_pos)
+
+ # regression L1 loss
+ loss_bbox = self.loss_bbox(
+ bbox_preds, bbox_targets, bbox_weights, avg_factor=num_total_pos)
+ return loss_cls, loss_bbox, loss_iou
+
+ def get_dn_targets(self, batch_gt_instances: InstanceList,
+ batch_img_metas: dict, dn_meta: Dict[str,
+ int]) -> tuple:
+ """Get targets in denoising part for a batch of images.
+
+ Args:
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ dn_meta (Dict[str, int]): The dictionary saves information about
+ group collation, including 'num_denoising_queries' and
+ 'num_denoising_groups'. It will be used for split outputs of
+ denoising and matching parts and loss calculation.
+
+ Returns:
+ tuple: a tuple containing the following targets.
+
+ - labels_list (list[Tensor]): Labels for all images.
+ - label_weights_list (list[Tensor]): Label weights for all images.
+ - bbox_targets_list (list[Tensor]): BBox targets for all images.
+ - bbox_weights_list (list[Tensor]): BBox weights for all images.
+ - num_total_pos (int): Number of positive samples in all images.
+ - num_total_neg (int): Number of negative samples in all images.
+ """
+ (labels_list, label_weights_list, bbox_targets_list, bbox_weights_list,
+ pos_inds_list, neg_inds_list) = multi_apply(
+ self._get_dn_targets_single,
+ batch_gt_instances,
+ batch_img_metas,
+ dn_meta=dn_meta)
+ num_total_pos = sum((inds.numel() for inds in pos_inds_list))
+ num_total_neg = sum((inds.numel() for inds in neg_inds_list))
+ return (labels_list, label_weights_list, bbox_targets_list,
+ bbox_weights_list, num_total_pos, num_total_neg)
+
+ def _get_dn_targets_single(self, gt_instances: InstanceData,
+ img_meta: dict, dn_meta: Dict[str,
+ int]) -> tuple:
+ """Get targets in denoising part for one image.
+
+ Args:
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes`` and ``labels``
+ attributes.
+ img_meta (dict): Meta information for one image.
+ dn_meta (Dict[str, int]): The dictionary saves information about
+ group collation, including 'num_denoising_queries' and
+ 'num_denoising_groups'. It will be used for split outputs of
+ denoising and matching parts and loss calculation.
+
+ Returns:
+ tuple[Tensor]: a tuple containing the following for one image.
+
+ - labels (Tensor): Labels of each image.
+ - label_weights (Tensor]): Label weights of each image.
+ - bbox_targets (Tensor): BBox targets of each image.
+ - bbox_weights (Tensor): BBox weights of each image.
+ - pos_inds (Tensor): Sampled positive indices for each image.
+ - neg_inds (Tensor): Sampled negative indices for each image.
+ """
+ gt_bboxes = gt_instances.bboxes
+ gt_labels = gt_instances.labels
+ num_groups = dn_meta['num_denoising_groups']
+ num_denoising_queries = dn_meta['num_denoising_queries']
+ num_queries_each_group = int(num_denoising_queries / num_groups)
+ device = gt_bboxes.device
+
+ if len(gt_labels) > 0:
+ t = torch.arange(len(gt_labels), dtype=torch.long, device=device)
+ t = t.unsqueeze(0).repeat(num_groups, 1)
+ pos_assigned_gt_inds = t.flatten()
+ pos_inds = torch.arange(
+ num_groups, dtype=torch.long, device=device)
+ pos_inds = pos_inds.unsqueeze(1) * num_queries_each_group + t
+ pos_inds = pos_inds.flatten()
+ else:
+ pos_inds = pos_assigned_gt_inds = \
+ gt_bboxes.new_tensor([], dtype=torch.long)
+
+ neg_inds = pos_inds + num_queries_each_group // 2
+
+ # label targets
+ labels = gt_bboxes.new_full((num_denoising_queries, ),
+ self.num_classes,
+ dtype=torch.long)
+ labels[pos_inds] = gt_labels[pos_assigned_gt_inds]
+ label_weights = gt_bboxes.new_ones(num_denoising_queries)
+
+ # bbox targets
+ bbox_targets = torch.zeros(num_denoising_queries, 4, device=device)
+ bbox_weights = torch.zeros(num_denoising_queries, 4, device=device)
+ bbox_weights[pos_inds] = 1.0
+ img_h, img_w = img_meta['img_shape']
+
+ # DETR regress the relative position of boxes (cxcywh) in the image.
+ # Thus the learning target should be normalized by the image size, also
+ # the box format should be converted from defaultly x1y1x2y2 to cxcywh.
+ factor = gt_bboxes.new_tensor([img_w, img_h, img_w,
+ img_h]).unsqueeze(0)
+ gt_bboxes_normalized = gt_bboxes / factor
+ gt_bboxes_targets = bbox_xyxy_to_cxcywh(gt_bboxes_normalized)
+ bbox_targets[pos_inds] = gt_bboxes_targets.repeat([num_groups, 1])
+
+ return (labels, label_weights, bbox_targets, bbox_weights, pos_inds,
+ neg_inds)
+
+ @staticmethod
+ def split_outputs(all_layers_cls_scores: Tensor,
+ all_layers_bbox_preds: Tensor,
+ dn_meta: Dict[str, int]) -> Tuple[Tensor]:
+ """Split outputs of the denoising part and the matching part.
+
+ For the total outputs of `num_queries_total` length, the former
+ `num_denoising_queries` outputs are from denoising queries, and
+ the rest `num_matching_queries` ones are from matching queries,
+ where `num_queries_total` is the sum of `num_denoising_queries` and
+ `num_matching_queries`.
+
+ Args:
+ all_layers_cls_scores (Tensor): Classification scores of all
+ decoder layers, has shape (num_decoder_layers, bs,
+ num_queries_total, cls_out_channels).
+ all_layers_bbox_preds (Tensor): Regression outputs of all decoder
+ layers. Each is a 4D-tensor with normalized coordinate format
+ (cx, cy, w, h) and has shape (num_decoder_layers, bs,
+ num_queries_total, 4).
+ dn_meta (Dict[str, int]): The dictionary saves information about
+ group collation, including 'num_denoising_queries' and
+ 'num_denoising_groups'.
+
+ Returns:
+ Tuple[Tensor]: a tuple containing the following outputs.
+
+ - all_layers_matching_cls_scores (Tensor): Classification scores
+ of all decoder layers in matching part, has shape
+ (num_decoder_layers, bs, num_matching_queries, cls_out_channels).
+ - all_layers_matching_bbox_preds (Tensor): Regression outputs of
+ all decoder layers in matching part. Each is a 4D-tensor with
+ normalized coordinate format (cx, cy, w, h) and has shape
+ (num_decoder_layers, bs, num_matching_queries, 4).
+ - all_layers_denoising_cls_scores (Tensor): Classification scores
+ of all decoder layers in denoising part, has shape
+ (num_decoder_layers, bs, num_denoising_queries,
+ cls_out_channels).
+ - all_layers_denoising_bbox_preds (Tensor): Regression outputs of
+ all decoder layers in denoising part. Each is a 4D-tensor with
+ normalized coordinate format (cx, cy, w, h) and has shape
+ (num_decoder_layers, bs, num_denoising_queries, 4).
+ """
+ num_denoising_queries = dn_meta['num_denoising_queries']
+ if dn_meta is not None:
+ all_layers_denoising_cls_scores = \
+ all_layers_cls_scores[:, :, : num_denoising_queries, :]
+ all_layers_denoising_bbox_preds = \
+ all_layers_bbox_preds[:, :, : num_denoising_queries, :]
+ all_layers_matching_cls_scores = \
+ all_layers_cls_scores[:, :, num_denoising_queries:, :]
+ all_layers_matching_bbox_preds = \
+ all_layers_bbox_preds[:, :, num_denoising_queries:, :]
+ else:
+ all_layers_denoising_cls_scores = None
+ all_layers_denoising_bbox_preds = None
+ all_layers_matching_cls_scores = all_layers_cls_scores
+ all_layers_matching_bbox_preds = all_layers_bbox_preds
+ return (all_layers_matching_cls_scores, all_layers_matching_bbox_preds,
+ all_layers_denoising_cls_scores,
+ all_layers_denoising_bbox_preds)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/embedding_rpn_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/embedding_rpn_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..97e84fa83b892c0274615d582fe43a6693541617
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/embedding_rpn_head.py
@@ -0,0 +1,132 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List
+
+import torch
+import torch.nn as nn
+from mmengine.model import BaseModule
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures.bbox import bbox_cxcywh_to_xyxy
+from mmdet.structures.det_data_sample import SampleList
+from mmdet.utils import InstanceList, OptConfigType
+
+
+@MODELS.register_module()
+class EmbeddingRPNHead(BaseModule):
+ """RPNHead in the `Sparse R-CNN `_ .
+
+ Unlike traditional RPNHead, this module does not need FPN input, but just
+ decode `init_proposal_bboxes` and expand the first dimension of
+ `init_proposal_bboxes` and `init_proposal_features` to the batch_size.
+
+ Args:
+ num_proposals (int): Number of init_proposals. Defaults to 100.
+ proposal_feature_channel (int): Channel number of
+ init_proposal_feature. Defaults to 256.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict]): Initialization config dict. Defaults to None.
+ """
+
+ def __init__(self,
+ num_proposals: int = 100,
+ proposal_feature_channel: int = 256,
+ init_cfg: OptConfigType = None,
+ **kwargs) -> None:
+ # `**kwargs` is necessary to avoid some potential error.
+ assert init_cfg is None, 'To prevent abnormal initialization ' \
+ 'behavior, init_cfg is not allowed to be set'
+ super().__init__(init_cfg=init_cfg)
+ self.num_proposals = num_proposals
+ self.proposal_feature_channel = proposal_feature_channel
+ self._init_layers()
+
+ def _init_layers(self) -> None:
+ """Initialize a sparse set of proposal boxes and proposal features."""
+ self.init_proposal_bboxes = nn.Embedding(self.num_proposals, 4)
+ self.init_proposal_features = nn.Embedding(
+ self.num_proposals, self.proposal_feature_channel)
+
+ def init_weights(self) -> None:
+ """Initialize the init_proposal_bboxes as normalized.
+
+ [c_x, c_y, w, h], and we initialize it to the size of the entire
+ image.
+ """
+ super().init_weights()
+ nn.init.constant_(self.init_proposal_bboxes.weight[:, :2], 0.5)
+ nn.init.constant_(self.init_proposal_bboxes.weight[:, 2:], 1)
+
+ def _decode_init_proposals(self, x: List[Tensor],
+ batch_data_samples: SampleList) -> InstanceList:
+ """Decode init_proposal_bboxes according to the size of images and
+ expand dimension of init_proposal_features to batch_size.
+
+ Args:
+ x (list[Tensor]): List of FPN features.
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+
+ Returns:
+ List[:obj:`InstanceData`:] Detection results of each image.
+ Each item usually contains following keys.
+
+ - proposals: Decoded proposal bboxes,
+ has shape (num_proposals, 4).
+ - features: init_proposal_features, expanded proposal
+ features, has shape
+ (num_proposals, proposal_feature_channel).
+ - imgs_whwh: Tensor with shape
+ (num_proposals, 4), the dimension means
+ [img_width, img_height, img_width, img_height].
+ """
+ batch_img_metas = []
+ for data_sample in batch_data_samples:
+ batch_img_metas.append(data_sample.metainfo)
+
+ proposals = self.init_proposal_bboxes.weight.clone()
+ proposals = bbox_cxcywh_to_xyxy(proposals)
+ imgs_whwh = []
+ for meta in batch_img_metas:
+ h, w = meta['img_shape'][:2]
+ imgs_whwh.append(x[0].new_tensor([[w, h, w, h]]))
+ imgs_whwh = torch.cat(imgs_whwh, dim=0)
+ imgs_whwh = imgs_whwh[:, None, :]
+ proposals = proposals * imgs_whwh
+
+ rpn_results_list = []
+ for idx in range(len(batch_img_metas)):
+ rpn_results = InstanceData()
+ rpn_results.bboxes = proposals[idx]
+ rpn_results.imgs_whwh = imgs_whwh[idx].repeat(
+ self.num_proposals, 1)
+ rpn_results.features = self.init_proposal_features.weight.clone()
+ rpn_results_list.append(rpn_results)
+ return rpn_results_list
+
+ def loss(self, *args, **kwargs):
+ """Perform forward propagation and loss calculation of the detection
+ head on the features of the upstream network."""
+ raise NotImplementedError(
+ 'EmbeddingRPNHead does not have `loss`, please use '
+ '`predict` or `loss_and_predict` instead.')
+
+ def predict(self, x: List[Tensor], batch_data_samples: SampleList,
+ **kwargs) -> InstanceList:
+ """Perform forward propagation of the detection head and predict
+ detection results on the features of the upstream network."""
+ # `**kwargs` is necessary to avoid some potential error.
+ return self._decode_init_proposals(
+ x=x, batch_data_samples=batch_data_samples)
+
+ def loss_and_predict(self, x: List[Tensor], batch_data_samples: SampleList,
+ **kwargs) -> tuple:
+ """Perform forward propagation of the head, then calculate loss and
+ predictions from the features and data samples."""
+ # `**kwargs` is necessary to avoid some potential error.
+ predictions = self._decode_init_proposals(
+ x=x, batch_data_samples=batch_data_samples)
+
+ return dict(), predictions
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/fcos_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/fcos_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..ba4d4640010c7e8e7c6a4db3e0fce887b4105217
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/fcos_head.py
@@ -0,0 +1,476 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, List, Tuple
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import Scale
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.models.layers import NormedConv2d
+from mmdet.registry import MODELS
+from mmdet.utils import (ConfigType, InstanceList, MultiConfig,
+ OptInstanceList, RangeType, reduce_mean)
+from ..utils import multi_apply
+from .anchor_free_head import AnchorFreeHead
+
+INF = 1e8
+
+
+@MODELS.register_module()
+class FCOSHead(AnchorFreeHead):
+ """Anchor-free head used in `FCOS `_.
+
+ The FCOS head does not use anchor boxes. Instead bounding boxes are
+ predicted at each pixel and a centerness measure is used to suppress
+ low-quality predictions.
+ Here norm_on_bbox, centerness_on_reg, dcn_on_last_conv are training
+ tricks used in official repo, which will bring remarkable mAP gains
+ of up to 4.9. Please see https://github.com/tianzhi0549/FCOS for
+ more detail.
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ strides (Sequence[int] or Sequence[Tuple[int, int]]): Strides of points
+ in multiple feature levels. Defaults to (4, 8, 16, 32, 64).
+ regress_ranges (Sequence[Tuple[int, int]]): Regress range of multiple
+ level points.
+ center_sampling (bool): If true, use center sampling.
+ Defaults to False.
+ center_sample_radius (float): Radius of center sampling.
+ Defaults to 1.5.
+ norm_on_bbox (bool): If true, normalize the regression targets with
+ FPN strides. Defaults to False.
+ centerness_on_reg (bool): If true, position centerness on the
+ regress branch. Please refer to https://github.com/tianzhi0549/FCOS/issues/89#issuecomment-516877042.
+ Defaults to False.
+ conv_bias (bool or str): If specified as `auto`, it will be decided by
+ the norm_cfg. Bias of conv will be set as True if `norm_cfg` is
+ None, otherwise False. Defaults to "auto".
+ loss_cls (:obj:`ConfigDict` or dict): Config of classification loss.
+ loss_bbox (:obj:`ConfigDict` or dict): Config of localization loss.
+ loss_centerness (:obj:`ConfigDict`, or dict): Config of centerness
+ loss.
+ norm_cfg (:obj:`ConfigDict` or dict): dictionary to construct and
+ config norm layer. Defaults to
+ ``norm_cfg=dict(type='GN', num_groups=32, requires_grad=True)``.
+ cls_predictor_cfg (:obj:`ConfigDict` or dict): dictionary to construct and
+ config conv_cls. Defaults to None.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict]): Initialization config dict.
+
+ Example:
+ >>> self = FCOSHead(11, 7)
+ >>> feats = [torch.rand(1, 7, s, s) for s in [4, 8, 16, 32, 64]]
+ >>> cls_score, bbox_pred, centerness = self.forward(feats)
+ >>> assert len(cls_score) == len(self.scales)
+ """ # noqa: E501
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: int,
+ regress_ranges: RangeType = ((-1, 64), (64, 128), (128, 256),
+ (256, 512), (512, INF)),
+ center_sampling: bool = False,
+ center_sample_radius: float = 1.5,
+ norm_on_bbox: bool = False,
+ centerness_on_reg: bool = False,
+ loss_cls: ConfigType = dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox: ConfigType = dict(type='IoULoss', loss_weight=1.0),
+ loss_centerness: ConfigType = dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ loss_weight=1.0),
+ norm_cfg: ConfigType = dict(
+ type='GN', num_groups=32, requires_grad=True),
+ cls_predictor_cfg=None,
+ init_cfg: MultiConfig = dict(
+ type='Normal',
+ layer='Conv2d',
+ std=0.01,
+ override=dict(
+ type='Normal',
+ name='conv_cls',
+ std=0.01,
+ bias_prob=0.01)),
+ **kwargs) -> None:
+ self.regress_ranges = regress_ranges
+ self.center_sampling = center_sampling
+ self.center_sample_radius = center_sample_radius
+ self.norm_on_bbox = norm_on_bbox
+ self.centerness_on_reg = centerness_on_reg
+ self.cls_predictor_cfg = cls_predictor_cfg
+ super().__init__(
+ num_classes=num_classes,
+ in_channels=in_channels,
+ loss_cls=loss_cls,
+ loss_bbox=loss_bbox,
+ norm_cfg=norm_cfg,
+ init_cfg=init_cfg,
+ **kwargs)
+ self.loss_centerness = MODELS.build(loss_centerness)
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ super()._init_layers()
+ self.conv_centerness = nn.Conv2d(self.feat_channels, 1, 3, padding=1)
+ self.scales = nn.ModuleList([Scale(1.0) for _ in self.strides])
+ if self.cls_predictor_cfg is not None:
+ self.cls_predictor_cfg.pop('type')
+ self.conv_cls = NormedConv2d(
+ self.feat_channels,
+ self.cls_out_channels,
+ 1,
+ padding=0,
+ **self.cls_predictor_cfg)
+
+ def forward(
+ self, x: Tuple[Tensor]
+ ) -> Tuple[List[Tensor], List[Tensor], List[Tensor]]:
+ """Forward features from the upstream network.
+
+ Args:
+ feats (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: A tuple of each level outputs.
+
+ - cls_scores (list[Tensor]): Box scores for each scale level, \
+ each is a 4D-tensor, the channel number is \
+ num_points * num_classes.
+ - bbox_preds (list[Tensor]): Box energies / deltas for each \
+ scale level, each is a 4D-tensor, the channel number is \
+ num_points * 4.
+ - centernesses (list[Tensor]): centerness for each scale level, \
+ each is a 4D-tensor, the channel number is num_points * 1.
+ """
+ return multi_apply(self.forward_single, x, self.scales, self.strides)
+
+ def forward_single(self, x: Tensor, scale: Scale,
+ stride: int) -> Tuple[Tensor, Tensor, Tensor]:
+ """Forward features of a single scale level.
+
+ Args:
+ x (Tensor): FPN feature maps of the specified stride.
+ scale (:obj:`mmcv.cnn.Scale`): Learnable scale module to resize
+ the bbox prediction.
+ stride (int): The corresponding stride for feature maps, only
+ used to normalize the bbox prediction when self.norm_on_bbox
+ is True.
+
+ Returns:
+ tuple: scores for each class, bbox predictions and centerness
+ predictions of input feature maps.
+ """
+ cls_score, bbox_pred, cls_feat, reg_feat = super().forward_single(x)
+ if self.centerness_on_reg:
+ centerness = self.conv_centerness(reg_feat)
+ else:
+ centerness = self.conv_centerness(cls_feat)
+ # scale the bbox_pred of different level
+ # float to avoid overflow when enabling FP16
+ bbox_pred = scale(bbox_pred).float()
+ if self.norm_on_bbox:
+ # bbox_pred needed for gradient computation has been modified
+ # by F.relu(bbox_pred) when run with PyTorch 1.10. So replace
+ # F.relu(bbox_pred) with bbox_pred.clamp(min=0)
+ bbox_pred = bbox_pred.clamp(min=0)
+ if not self.training:
+ bbox_pred *= stride
+ else:
+ bbox_pred = bbox_pred.exp()
+ return cls_score, bbox_pred, centerness
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ centernesses: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None
+ ) -> Dict[str, Tensor]:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level,
+ each is a 4D-tensor, the channel number is
+ num_points * num_classes.
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level, each is a 4D-tensor, the channel number is
+ num_points * 4.
+ centernesses (list[Tensor]): centerness for each scale level, each
+ is a 4D-tensor, the channel number is num_points * 1.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ assert len(cls_scores) == len(bbox_preds) == len(centernesses)
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ all_level_points = self.prior_generator.grid_priors(
+ featmap_sizes,
+ dtype=bbox_preds[0].dtype,
+ device=bbox_preds[0].device)
+ labels, bbox_targets = self.get_targets(all_level_points,
+ batch_gt_instances)
+
+ num_imgs = cls_scores[0].size(0)
+ # flatten cls_scores, bbox_preds and centerness
+ flatten_cls_scores = [
+ cls_score.permute(0, 2, 3, 1).reshape(-1, self.cls_out_channels)
+ for cls_score in cls_scores
+ ]
+ flatten_bbox_preds = [
+ bbox_pred.permute(0, 2, 3, 1).reshape(-1, 4)
+ for bbox_pred in bbox_preds
+ ]
+ flatten_centerness = [
+ centerness.permute(0, 2, 3, 1).reshape(-1)
+ for centerness in centernesses
+ ]
+ flatten_cls_scores = torch.cat(flatten_cls_scores)
+ flatten_bbox_preds = torch.cat(flatten_bbox_preds)
+ flatten_centerness = torch.cat(flatten_centerness)
+ flatten_labels = torch.cat(labels)
+ flatten_bbox_targets = torch.cat(bbox_targets)
+ # repeat points to align with bbox_preds
+ flatten_points = torch.cat(
+ [points.repeat(num_imgs, 1) for points in all_level_points])
+
+ losses = dict()
+
+ # FG cat_id: [0, num_classes -1], BG cat_id: num_classes
+ bg_class_ind = self.num_classes
+ pos_inds = ((flatten_labels >= 0)
+ & (flatten_labels < bg_class_ind)).nonzero().reshape(-1)
+ num_pos = torch.tensor(
+ len(pos_inds), dtype=torch.float, device=bbox_preds[0].device)
+ num_pos = max(reduce_mean(num_pos), 1.0)
+ loss_cls = self.loss_cls(
+ flatten_cls_scores, flatten_labels, avg_factor=num_pos)
+
+ if getattr(self.loss_cls, 'custom_accuracy', False):
+ acc = self.loss_cls.get_accuracy(flatten_cls_scores,
+ flatten_labels)
+ losses.update(acc)
+
+ pos_bbox_preds = flatten_bbox_preds[pos_inds]
+ pos_centerness = flatten_centerness[pos_inds]
+ pos_bbox_targets = flatten_bbox_targets[pos_inds]
+ pos_centerness_targets = self.centerness_target(pos_bbox_targets)
+ # centerness weighted iou loss
+ centerness_denorm = max(
+ reduce_mean(pos_centerness_targets.sum().detach()), 1e-6)
+
+ if len(pos_inds) > 0:
+ pos_points = flatten_points[pos_inds]
+ pos_decoded_bbox_preds = self.bbox_coder.decode(
+ pos_points, pos_bbox_preds)
+ pos_decoded_target_preds = self.bbox_coder.decode(
+ pos_points, pos_bbox_targets)
+ loss_bbox = self.loss_bbox(
+ pos_decoded_bbox_preds,
+ pos_decoded_target_preds,
+ weight=pos_centerness_targets,
+ avg_factor=centerness_denorm)
+ loss_centerness = self.loss_centerness(
+ pos_centerness, pos_centerness_targets, avg_factor=num_pos)
+ else:
+ loss_bbox = pos_bbox_preds.sum()
+ loss_centerness = pos_centerness.sum()
+
+ losses['loss_cls'] = loss_cls
+ losses['loss_bbox'] = loss_bbox
+ losses['loss_centerness'] = loss_centerness
+
+ return losses
+
+ def get_targets(
+ self, points: List[Tensor], batch_gt_instances: InstanceList
+ ) -> Tuple[List[Tensor], List[Tensor]]:
+ """Compute regression, classification and centerness targets for points
+ in multiple images.
+
+ Args:
+ points (list[Tensor]): Points of each fpn level, each has shape
+ (num_points, 2).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+
+ Returns:
+ tuple: Targets of each level.
+
+ - concat_lvl_labels (list[Tensor]): Labels of each level.
+ - concat_lvl_bbox_targets (list[Tensor]): BBox targets of each \
+ level.
+ """
+ assert len(points) == len(self.regress_ranges)
+ num_levels = len(points)
+ # expand regress ranges to align with points
+ expanded_regress_ranges = [
+ points[i].new_tensor(self.regress_ranges[i])[None].expand_as(
+ points[i]) for i in range(num_levels)
+ ]
+ # concat all levels points and regress ranges
+ concat_regress_ranges = torch.cat(expanded_regress_ranges, dim=0)
+ concat_points = torch.cat(points, dim=0)
+
+ # the number of points per img, per lvl
+ num_points = [center.size(0) for center in points]
+
+ # get labels and bbox_targets of each image
+ labels_list, bbox_targets_list = multi_apply(
+ self._get_targets_single,
+ batch_gt_instances,
+ points=concat_points,
+ regress_ranges=concat_regress_ranges,
+ num_points_per_lvl=num_points)
+
+ # split to per img, per level
+ labels_list = [labels.split(num_points, 0) for labels in labels_list]
+ bbox_targets_list = [
+ bbox_targets.split(num_points, 0)
+ for bbox_targets in bbox_targets_list
+ ]
+
+ # concat per level image
+ concat_lvl_labels = []
+ concat_lvl_bbox_targets = []
+ for i in range(num_levels):
+ concat_lvl_labels.append(
+ torch.cat([labels[i] for labels in labels_list]))
+ bbox_targets = torch.cat(
+ [bbox_targets[i] for bbox_targets in bbox_targets_list])
+ if self.norm_on_bbox:
+ bbox_targets = bbox_targets / self.strides[i]
+ concat_lvl_bbox_targets.append(bbox_targets)
+ return concat_lvl_labels, concat_lvl_bbox_targets
+
+ def _get_targets_single(
+ self, gt_instances: InstanceData, points: Tensor,
+ regress_ranges: Tensor,
+ num_points_per_lvl: List[int]) -> Tuple[Tensor, Tensor]:
+ """Compute regression and classification targets for a single image."""
+ num_points = points.size(0)
+ num_gts = len(gt_instances)
+ gt_bboxes = gt_instances.bboxes
+ gt_labels = gt_instances.labels
+
+ if num_gts == 0:
+ return gt_labels.new_full((num_points,), self.num_classes), \
+ gt_bboxes.new_zeros((num_points, 4))
+
+ areas = (gt_bboxes[:, 2] - gt_bboxes[:, 0]) * (
+ gt_bboxes[:, 3] - gt_bboxes[:, 1])
+ # TODO: figure out why these two are different
+ # areas = areas[None].expand(num_points, num_gts)
+ areas = areas[None].repeat(num_points, 1)
+ regress_ranges = regress_ranges[:, None, :].expand(
+ num_points, num_gts, 2)
+ gt_bboxes = gt_bboxes[None].expand(num_points, num_gts, 4)
+ xs, ys = points[:, 0], points[:, 1]
+ xs = xs[:, None].expand(num_points, num_gts)
+ ys = ys[:, None].expand(num_points, num_gts)
+
+ left = xs - gt_bboxes[..., 0]
+ right = gt_bboxes[..., 2] - xs
+ top = ys - gt_bboxes[..., 1]
+ bottom = gt_bboxes[..., 3] - ys
+ bbox_targets = torch.stack((left, top, right, bottom), -1)
+
+ if self.center_sampling:
+ # condition1: inside a `center bbox`
+ radius = self.center_sample_radius
+ center_xs = (gt_bboxes[..., 0] + gt_bboxes[..., 2]) / 2
+ center_ys = (gt_bboxes[..., 1] + gt_bboxes[..., 3]) / 2
+ center_gts = torch.zeros_like(gt_bboxes)
+ stride = center_xs.new_zeros(center_xs.shape)
+
+ # project the points on current lvl back to the `original` sizes
+ lvl_begin = 0
+ for lvl_idx, num_points_lvl in enumerate(num_points_per_lvl):
+ lvl_end = lvl_begin + num_points_lvl
+ stride[lvl_begin:lvl_end] = self.strides[lvl_idx] * radius
+ lvl_begin = lvl_end
+
+ x_mins = center_xs - stride
+ y_mins = center_ys - stride
+ x_maxs = center_xs + stride
+ y_maxs = center_ys + stride
+ center_gts[..., 0] = torch.where(x_mins > gt_bboxes[..., 0],
+ x_mins, gt_bboxes[..., 0])
+ center_gts[..., 1] = torch.where(y_mins > gt_bboxes[..., 1],
+ y_mins, gt_bboxes[..., 1])
+ center_gts[..., 2] = torch.where(x_maxs > gt_bboxes[..., 2],
+ gt_bboxes[..., 2], x_maxs)
+ center_gts[..., 3] = torch.where(y_maxs > gt_bboxes[..., 3],
+ gt_bboxes[..., 3], y_maxs)
+
+ cb_dist_left = xs - center_gts[..., 0]
+ cb_dist_right = center_gts[..., 2] - xs
+ cb_dist_top = ys - center_gts[..., 1]
+ cb_dist_bottom = center_gts[..., 3] - ys
+ center_bbox = torch.stack(
+ (cb_dist_left, cb_dist_top, cb_dist_right, cb_dist_bottom), -1)
+ inside_gt_bbox_mask = center_bbox.min(-1)[0] > 0
+ else:
+ # condition1: inside a gt bbox
+ inside_gt_bbox_mask = bbox_targets.min(-1)[0] > 0
+
+ # condition2: limit the regression range for each location
+ max_regress_distance = bbox_targets.max(-1)[0]
+ inside_regress_range = (
+ (max_regress_distance >= regress_ranges[..., 0])
+ & (max_regress_distance <= regress_ranges[..., 1]))
+
+ # if there are still more than one objects for a location,
+ # we choose the one with minimal area
+ areas[inside_gt_bbox_mask == 0] = INF
+ areas[inside_regress_range == 0] = INF
+ min_area, min_area_inds = areas.min(dim=1)
+
+ labels = gt_labels[min_area_inds]
+ labels[min_area == INF] = self.num_classes # set as BG
+ bbox_targets = bbox_targets[range(num_points), min_area_inds]
+
+ return labels, bbox_targets
+
+ def centerness_target(self, pos_bbox_targets: Tensor) -> Tensor:
+ """Compute centerness targets.
+
+ Args:
+ pos_bbox_targets (Tensor): BBox targets of positive bboxes in shape
+ (num_pos, 4)
+
+ Returns:
+ Tensor: Centerness target.
+ """
+ # only calculate pos centerness targets, otherwise there may be nan
+ left_right = pos_bbox_targets[:, [0, 2]]
+ top_bottom = pos_bbox_targets[:, [1, 3]]
+ if len(left_right) == 0:
+ centerness_targets = left_right[..., 0]
+ else:
+ centerness_targets = (
+ left_right.min(dim=-1)[0] / left_right.max(dim=-1)[0]) * (
+ top_bottom.min(dim=-1)[0] / top_bottom.max(dim=-1)[0])
+ return torch.sqrt(centerness_targets)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/fovea_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/fovea_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..89353deac7f0189c1e464288521ee8e4238f0107
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/fovea_head.py
@@ -0,0 +1,509 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, List, Optional, Tuple
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from mmcv.ops import DeformConv2d
+from mmengine.config import ConfigDict
+from mmengine.model import BaseModule
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import InstanceList, OptInstanceList, OptMultiConfig
+from ..utils import filter_scores_and_topk, multi_apply
+from .anchor_free_head import AnchorFreeHead
+
+INF = 1e8
+
+
+class FeatureAlign(BaseModule):
+ """Feature Align Module.
+
+ Feature Align Module is implemented based on DCN v1.
+ It uses anchor shape prediction rather than feature map to
+ predict offsets of deform conv layer.
+
+ Args:
+ in_channels (int): Number of channels in the input feature map.
+ out_channels (int): Number of channels in the output feature map.
+ kernel_size (int): Size of the convolution kernel.
+ ``norm_cfg=dict(type='GN', num_groups=32, requires_grad=True)``.
+ deform_groups: (int): Group number of DCN in
+ FeatureAdaption module.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict], optional): Initialization config dict.
+ """
+
+ def __init__(
+ self,
+ in_channels: int,
+ out_channels: int,
+ kernel_size: int = 3,
+ deform_groups: int = 4,
+ init_cfg: OptMultiConfig = dict(
+ type='Normal',
+ layer='Conv2d',
+ std=0.1,
+ override=dict(type='Normal', name='conv_adaption', std=0.01))
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ offset_channels = kernel_size * kernel_size * 2
+ self.conv_offset = nn.Conv2d(
+ 4, deform_groups * offset_channels, 1, bias=False)
+ self.conv_adaption = DeformConv2d(
+ in_channels,
+ out_channels,
+ kernel_size=kernel_size,
+ padding=(kernel_size - 1) // 2,
+ deform_groups=deform_groups)
+ self.relu = nn.ReLU(inplace=True)
+
+ def forward(self, x: Tensor, shape: Tensor) -> Tensor:
+ """Forward function of feature align module.
+
+ Args:
+ x (Tensor): Features from the upstream network.
+ shape (Tensor): Exponential of bbox predictions.
+
+ Returns:
+ x (Tensor): The aligned features.
+ """
+ offset = self.conv_offset(shape)
+ x = self.relu(self.conv_adaption(x, offset))
+ return x
+
+
+@MODELS.register_module()
+class FoveaHead(AnchorFreeHead):
+ """Detection Head of `FoveaBox: Beyond Anchor-based Object Detector.
+
+ `_.
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ base_edge_list (list[int]): List of edges.
+ scale_ranges (list[tuple]): Range of scales.
+ sigma (float): Super parameter of ``FoveaHead``.
+ with_deform (bool): Whether use deform conv.
+ deform_groups (int): Deformable conv group size.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict], optional): Initialization config dict.
+ """
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: int,
+ base_edge_list: List[int] = (16, 32, 64, 128, 256),
+ scale_ranges: List[tuple] = ((8, 32), (16, 64), (32, 128),
+ (64, 256), (128, 512)),
+ sigma: float = 0.4,
+ with_deform: bool = False,
+ deform_groups: int = 4,
+ init_cfg: OptMultiConfig = dict(
+ type='Normal',
+ layer='Conv2d',
+ std=0.01,
+ override=dict(
+ type='Normal',
+ name='conv_cls',
+ std=0.01,
+ bias_prob=0.01)),
+ **kwargs) -> None:
+ self.base_edge_list = base_edge_list
+ self.scale_ranges = scale_ranges
+ self.sigma = sigma
+ self.with_deform = with_deform
+ self.deform_groups = deform_groups
+ super().__init__(
+ num_classes=num_classes,
+ in_channels=in_channels,
+ init_cfg=init_cfg,
+ **kwargs)
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ # box branch
+ super()._init_reg_convs()
+ self.conv_reg = nn.Conv2d(self.feat_channels, 4, 3, padding=1)
+
+ # cls branch
+ if not self.with_deform:
+ super()._init_cls_convs()
+ self.conv_cls = nn.Conv2d(
+ self.feat_channels, self.cls_out_channels, 3, padding=1)
+ else:
+ self.cls_convs = nn.ModuleList()
+ self.cls_convs.append(
+ ConvModule(
+ self.feat_channels, (self.feat_channels * 4),
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ bias=self.norm_cfg is None))
+ self.cls_convs.append(
+ ConvModule((self.feat_channels * 4), (self.feat_channels * 4),
+ 1,
+ stride=1,
+ padding=0,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ bias=self.norm_cfg is None))
+ self.feature_adaption = FeatureAlign(
+ self.feat_channels,
+ self.feat_channels,
+ kernel_size=3,
+ deform_groups=self.deform_groups)
+ self.conv_cls = nn.Conv2d(
+ int(self.feat_channels * 4),
+ self.cls_out_channels,
+ 3,
+ padding=1)
+
+ def forward_single(self, x: Tensor) -> Tuple[Tensor, Tensor]:
+ """Forward features of a single scale level.
+
+ Args:
+ x (Tensor): FPN feature maps of the specified stride.
+
+ Returns:
+ tuple: scores for each class and bbox predictions of input
+ feature maps.
+ """
+ cls_feat = x
+ reg_feat = x
+ for reg_layer in self.reg_convs:
+ reg_feat = reg_layer(reg_feat)
+ bbox_pred = self.conv_reg(reg_feat)
+ if self.with_deform:
+ cls_feat = self.feature_adaption(cls_feat, bbox_pred.exp())
+ for cls_layer in self.cls_convs:
+ cls_feat = cls_layer(cls_feat)
+ cls_score = self.conv_cls(cls_feat)
+ return cls_score, bbox_pred
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None
+ ) -> Dict[str, Tensor]:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level,
+ each is a 4D-tensor, the channel number is
+ num_priors * num_classes.
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level, each is a 4D-tensor, the channel number is
+ num_priors * 4.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ assert len(cls_scores) == len(bbox_preds)
+
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ priors = self.prior_generator.grid_priors(
+ featmap_sizes,
+ dtype=bbox_preds[0].dtype,
+ device=bbox_preds[0].device)
+ num_imgs = cls_scores[0].size(0)
+ flatten_cls_scores = [
+ cls_score.permute(0, 2, 3, 1).reshape(-1, self.cls_out_channels)
+ for cls_score in cls_scores
+ ]
+ flatten_bbox_preds = [
+ bbox_pred.permute(0, 2, 3, 1).reshape(-1, 4)
+ for bbox_pred in bbox_preds
+ ]
+ flatten_cls_scores = torch.cat(flatten_cls_scores)
+ flatten_bbox_preds = torch.cat(flatten_bbox_preds)
+ flatten_labels, flatten_bbox_targets = self.get_targets(
+ batch_gt_instances, featmap_sizes, priors)
+
+ # FG cat_id: [0, num_classes -1], BG cat_id: num_classes
+ pos_inds = ((flatten_labels >= 0)
+ & (flatten_labels < self.num_classes)).nonzero().view(-1)
+ num_pos = len(pos_inds)
+
+ loss_cls = self.loss_cls(
+ flatten_cls_scores, flatten_labels, avg_factor=num_pos + num_imgs)
+ if num_pos > 0:
+ pos_bbox_preds = flatten_bbox_preds[pos_inds]
+ pos_bbox_targets = flatten_bbox_targets[pos_inds]
+ pos_weights = pos_bbox_targets.new_ones(pos_bbox_targets.size())
+ loss_bbox = self.loss_bbox(
+ pos_bbox_preds,
+ pos_bbox_targets,
+ pos_weights,
+ avg_factor=num_pos)
+ else:
+ loss_bbox = torch.tensor(
+ 0,
+ dtype=flatten_bbox_preds.dtype,
+ device=flatten_bbox_preds.device)
+ return dict(loss_cls=loss_cls, loss_bbox=loss_bbox)
+
+ def get_targets(
+ self, batch_gt_instances: InstanceList, featmap_sizes: List[tuple],
+ priors_list: List[Tensor]) -> Tuple[List[Tensor], List[Tensor]]:
+ """Compute regression and classification for priors in multiple images.
+
+ Args:
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ featmap_sizes (list[tuple]): Size tuple of feature maps.
+ priors_list (list[Tensor]): Priors list of each fpn level, each has
+ shape (num_priors, 2).
+
+ Returns:
+ tuple: Targets of each level.
+
+ - flatten_labels (list[Tensor]): Labels of each level.
+ - flatten_bbox_targets (list[Tensor]): BBox targets of each
+ level.
+ """
+ label_list, bbox_target_list = multi_apply(
+ self._get_targets_single,
+ batch_gt_instances,
+ featmap_size_list=featmap_sizes,
+ priors_list=priors_list)
+ flatten_labels = [
+ torch.cat([
+ labels_level_img.flatten() for labels_level_img in labels_level
+ ]) for labels_level in zip(*label_list)
+ ]
+ flatten_bbox_targets = [
+ torch.cat([
+ bbox_targets_level_img.reshape(-1, 4)
+ for bbox_targets_level_img in bbox_targets_level
+ ]) for bbox_targets_level in zip(*bbox_target_list)
+ ]
+ flatten_labels = torch.cat(flatten_labels)
+ flatten_bbox_targets = torch.cat(flatten_bbox_targets)
+ return flatten_labels, flatten_bbox_targets
+
+ def _get_targets_single(self,
+ gt_instances: InstanceData,
+ featmap_size_list: List[tuple] = None,
+ priors_list: List[Tensor] = None) -> tuple:
+ """Compute regression and classification targets for a single image.
+
+ Args:
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ featmap_size_list (list[tuple]): Size tuple of feature maps.
+ priors_list (list[Tensor]): Priors of each fpn level, each has
+ shape (num_priors, 2).
+
+ Returns:
+ tuple:
+
+ - label_list (list[Tensor]): Labels of all anchors in the image.
+ - box_target_list (list[Tensor]): BBox targets of all anchors in
+ the image.
+ """
+ gt_bboxes_raw = gt_instances.bboxes
+ gt_labels_raw = gt_instances.labels
+ gt_areas = torch.sqrt((gt_bboxes_raw[:, 2] - gt_bboxes_raw[:, 0]) *
+ (gt_bboxes_raw[:, 3] - gt_bboxes_raw[:, 1]))
+ label_list = []
+ bbox_target_list = []
+ # for each pyramid, find the cls and box target
+ for base_len, (lower_bound, upper_bound), stride, featmap_size, \
+ priors in zip(self.base_edge_list, self.scale_ranges,
+ self.strides, featmap_size_list, priors_list):
+ # FG cat_id: [0, num_classes -1], BG cat_id: num_classes
+ priors = priors.view(*featmap_size, 2)
+ x, y = priors[..., 0], priors[..., 1]
+ labels = gt_labels_raw.new_full(featmap_size, self.num_classes)
+ bbox_targets = gt_bboxes_raw.new_ones(featmap_size[0],
+ featmap_size[1], 4)
+ # scale assignment
+ hit_indices = ((gt_areas >= lower_bound) &
+ (gt_areas <= upper_bound)).nonzero().flatten()
+ if len(hit_indices) == 0:
+ label_list.append(labels)
+ bbox_target_list.append(torch.log(bbox_targets))
+ continue
+ _, hit_index_order = torch.sort(-gt_areas[hit_indices])
+ hit_indices = hit_indices[hit_index_order]
+ gt_bboxes = gt_bboxes_raw[hit_indices, :] / stride
+ gt_labels = gt_labels_raw[hit_indices]
+ half_w = 0.5 * (gt_bboxes[:, 2] - gt_bboxes[:, 0])
+ half_h = 0.5 * (gt_bboxes[:, 3] - gt_bboxes[:, 1])
+ # valid fovea area: left, right, top, down
+ pos_left = torch.ceil(
+ gt_bboxes[:, 0] + (1 - self.sigma) * half_w - 0.5).long(). \
+ clamp(0, featmap_size[1] - 1)
+ pos_right = torch.floor(
+ gt_bboxes[:, 0] + (1 + self.sigma) * half_w - 0.5).long(). \
+ clamp(0, featmap_size[1] - 1)
+ pos_top = torch.ceil(
+ gt_bboxes[:, 1] + (1 - self.sigma) * half_h - 0.5).long(). \
+ clamp(0, featmap_size[0] - 1)
+ pos_down = torch.floor(
+ gt_bboxes[:, 1] + (1 + self.sigma) * half_h - 0.5).long(). \
+ clamp(0, featmap_size[0] - 1)
+ for px1, py1, px2, py2, label, (gt_x1, gt_y1, gt_x2, gt_y2) in \
+ zip(pos_left, pos_top, pos_right, pos_down, gt_labels,
+ gt_bboxes_raw[hit_indices, :]):
+ labels[py1:py2 + 1, px1:px2 + 1] = label
+ bbox_targets[py1:py2 + 1, px1:px2 + 1, 0] = \
+ (x[py1:py2 + 1, px1:px2 + 1] - gt_x1) / base_len
+ bbox_targets[py1:py2 + 1, px1:px2 + 1, 1] = \
+ (y[py1:py2 + 1, px1:px2 + 1] - gt_y1) / base_len
+ bbox_targets[py1:py2 + 1, px1:px2 + 1, 2] = \
+ (gt_x2 - x[py1:py2 + 1, px1:px2 + 1]) / base_len
+ bbox_targets[py1:py2 + 1, px1:px2 + 1, 3] = \
+ (gt_y2 - y[py1:py2 + 1, px1:px2 + 1]) / base_len
+ bbox_targets = bbox_targets.clamp(min=1. / 16, max=16.)
+ label_list.append(labels)
+ bbox_target_list.append(torch.log(bbox_targets))
+ return label_list, bbox_target_list
+
+ # Same as base_dense_head/_predict_by_feat_single except self._bbox_decode
+ def _predict_by_feat_single(self,
+ cls_score_list: List[Tensor],
+ bbox_pred_list: List[Tensor],
+ score_factor_list: List[Tensor],
+ mlvl_priors: List[Tensor],
+ img_meta: dict,
+ cfg: Optional[ConfigDict] = None,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results.
+
+ Args:
+ cls_score_list (list[Tensor]): Box scores from all scale
+ levels of a single image, each item has shape
+ (num_priors * num_classes, H, W).
+ bbox_pred_list (list[Tensor]): Box energies / deltas from
+ all scale levels of a single image, each item has shape
+ (num_priors * 4, H, W).
+ score_factor_list (list[Tensor]): Score factor from all scale
+ levels of a single image, each item has shape
+ (num_priors * 1, H, W).
+ mlvl_priors (list[Tensor]): Each element in the list is
+ the priors of a single level in feature pyramid, has shape
+ (num_priors, 2).
+ img_meta (dict): Image meta info.
+ cfg (ConfigDict, optional): Test / postprocessing
+ configuration, if None, test_cfg would be used.
+ Defaults to None.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ cfg = self.test_cfg if cfg is None else cfg
+ assert len(cls_score_list) == len(bbox_pred_list)
+ img_shape = img_meta['img_shape']
+ nms_pre = cfg.get('nms_pre', -1)
+
+ mlvl_bboxes = []
+ mlvl_scores = []
+ mlvl_labels = []
+ for level_idx, (cls_score, bbox_pred, stride, base_len, priors) in \
+ enumerate(zip(cls_score_list, bbox_pred_list, self.strides,
+ self.base_edge_list, mlvl_priors)):
+ assert cls_score.size()[-2:] == bbox_pred.size()[-2:]
+ bbox_pred = bbox_pred.permute(1, 2, 0).reshape(-1, 4)
+
+ scores = cls_score.permute(1, 2, 0).reshape(
+ -1, self.cls_out_channels).sigmoid()
+
+ # After https://github.com/open-mmlab/mmdetection/pull/6268/,
+ # this operation keeps fewer bboxes under the same `nms_pre`.
+ # There is no difference in performance for most models. If you
+ # find a slight drop in performance, you can set a larger
+ # `nms_pre` than before.
+ results = filter_scores_and_topk(
+ scores, cfg.score_thr, nms_pre,
+ dict(bbox_pred=bbox_pred, priors=priors))
+ scores, labels, _, filtered_results = results
+
+ bbox_pred = filtered_results['bbox_pred']
+ priors = filtered_results['priors']
+
+ bboxes = self._bbox_decode(priors, bbox_pred, base_len, img_shape)
+
+ mlvl_bboxes.append(bboxes)
+ mlvl_scores.append(scores)
+ mlvl_labels.append(labels)
+
+ results = InstanceData()
+ results.bboxes = torch.cat(mlvl_bboxes)
+ results.scores = torch.cat(mlvl_scores)
+ results.labels = torch.cat(mlvl_labels)
+
+ return self._bbox_post_process(
+ results=results,
+ cfg=cfg,
+ rescale=rescale,
+ with_nms=with_nms,
+ img_meta=img_meta)
+
+ def _bbox_decode(self, priors: Tensor, bbox_pred: Tensor, base_len: int,
+ max_shape: int) -> Tensor:
+ """Function to decode bbox.
+
+ Args:
+ priors (Tensor): Center proiors of an image, has shape
+ (num_instances, 2).
+ bbox_preds (Tensor): Box energies / deltas for all instances,
+ has shape (batch_size, num_instances, 4).
+ base_len (int): The base length.
+ max_shape (int): The max shape of bbox.
+
+ Returns:
+ Tensor: Decoded bboxes in (tl_x, tl_y, br_x, br_y) format. Has
+ shape (batch_size, num_instances, 4).
+ """
+ bbox_pred = bbox_pred.exp()
+
+ y = priors[:, 1]
+ x = priors[:, 0]
+ x1 = (x - base_len * bbox_pred[:, 0]). \
+ clamp(min=0, max=max_shape[1] - 1)
+ y1 = (y - base_len * bbox_pred[:, 1]). \
+ clamp(min=0, max=max_shape[0] - 1)
+ x2 = (x + base_len * bbox_pred[:, 2]). \
+ clamp(min=0, max=max_shape[1] - 1)
+ y2 = (y + base_len * bbox_pred[:, 3]). \
+ clamp(min=0, max=max_shape[0] - 1)
+ decoded_bboxes = torch.stack([x1, y1, x2, y2], -1)
+ return decoded_bboxes
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/free_anchor_retina_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/free_anchor_retina_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..df6fb9202c32735121bf7738e332fbfc5ac7e6bd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/free_anchor_retina_head.py
@@ -0,0 +1,312 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List
+
+import torch
+import torch.nn.functional as F
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures.bbox import bbox_overlaps
+from mmdet.utils import InstanceList, OptConfigType, OptInstanceList
+from ..utils import multi_apply
+from .retina_head import RetinaHead
+
+EPS = 1e-12
+
+
+@MODELS.register_module()
+class FreeAnchorRetinaHead(RetinaHead):
+ """FreeAnchor RetinaHead used in https://arxiv.org/abs/1909.02466.
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ stacked_convs (int): Number of conv layers in cls and reg tower.
+ Defaults to 4.
+ conv_cfg (:obj:`ConfigDict` or dict, optional): dictionary to
+ construct and config conv layer. Defaults to None.
+ norm_cfg (:obj:`ConfigDict` or dict, optional): dictionary to
+ construct and config norm layer. Defaults to
+ norm_cfg=dict(type='GN', num_groups=32, requires_grad=True).
+ pre_anchor_topk (int): Number of boxes that be token in each bag.
+ Defaults to 50
+ bbox_thr (float): The threshold of the saturated linear function.
+ It is usually the same with the IoU threshold used in NMS.
+ Defaults to 0.6.
+ gamma (float): Gamma parameter in focal loss. Defaults to 2.0.
+ alpha (float): Alpha parameter in focal loss. Defaults to 0.5.
+ """
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: int,
+ stacked_convs: int = 4,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: OptConfigType = None,
+ pre_anchor_topk: int = 50,
+ bbox_thr: float = 0.6,
+ gamma: float = 2.0,
+ alpha: float = 0.5,
+ **kwargs) -> None:
+ super().__init__(
+ num_classes=num_classes,
+ in_channels=in_channels,
+ stacked_convs=stacked_convs,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ **kwargs)
+
+ self.pre_anchor_topk = pre_anchor_topk
+ self.bbox_thr = bbox_thr
+ self.gamma = gamma
+ self.alpha = alpha
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ has shape (N, num_anchors * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ assert len(featmap_sizes) == self.prior_generator.num_levels
+
+ device = cls_scores[0].device
+
+ anchor_list, _ = self.get_anchors(
+ featmap_sizes=featmap_sizes,
+ batch_img_metas=batch_img_metas,
+ device=device)
+ concat_anchor_list = [torch.cat(anchor) for anchor in anchor_list]
+
+ # concatenate each level
+ cls_scores = [
+ cls.permute(0, 2, 3,
+ 1).reshape(cls.size(0), -1, self.cls_out_channels)
+ for cls in cls_scores
+ ]
+ bbox_preds = [
+ bbox_pred.permute(0, 2, 3, 1).reshape(bbox_pred.size(0), -1, 4)
+ for bbox_pred in bbox_preds
+ ]
+ cls_scores = torch.cat(cls_scores, dim=1)
+ cls_probs = torch.sigmoid(cls_scores)
+ bbox_preds = torch.cat(bbox_preds, dim=1)
+
+ box_probs, positive_losses, num_pos_list = multi_apply(
+ self.positive_loss_single, cls_probs, bbox_preds,
+ concat_anchor_list, batch_gt_instances)
+
+ num_pos = sum(num_pos_list)
+ positive_loss = torch.cat(positive_losses).sum() / max(1, num_pos)
+
+ # box_prob: P{a_{j} \in A_{+}}
+ box_probs = torch.stack(box_probs, dim=0)
+
+ # negative_loss:
+ # \sum_{j}{ FL((1 - P{a_{j} \in A_{+}}) * (1 - P_{j}^{bg})) } / n||B||
+ negative_loss = self.negative_bag_loss(cls_probs, box_probs).sum() / \
+ max(1, num_pos * self.pre_anchor_topk)
+
+ # avoid the absence of gradients in regression subnet
+ # when no ground-truth in a batch
+ if num_pos == 0:
+ positive_loss = bbox_preds.sum() * 0
+
+ losses = {
+ 'positive_bag_loss': positive_loss,
+ 'negative_bag_loss': negative_loss
+ }
+ return losses
+
+ def positive_loss_single(self, cls_prob: Tensor, bbox_pred: Tensor,
+ flat_anchors: Tensor,
+ gt_instances: InstanceData) -> tuple:
+ """Compute positive loss.
+
+ Args:
+ cls_prob (Tensor): Classification probability of shape
+ (num_anchors, num_classes).
+ bbox_pred (Tensor): Box probability of shape (num_anchors, 4).
+ flat_anchors (Tensor): Multi-level anchors of the image, which are
+ concatenated into a single tensor of shape (num_anchors, 4)
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes`` and ``labels``
+ attributes.
+
+ Returns:
+ tuple:
+
+ - box_prob (Tensor): Box probability of shape (num_anchors, 4).
+ - positive_loss (Tensor): Positive loss of shape (num_pos, ).
+ - num_pos (int): positive samples indexes.
+ """
+
+ gt_bboxes = gt_instances.bboxes
+ gt_labels = gt_instances.labels
+ with torch.no_grad():
+ if len(gt_bboxes) == 0:
+ image_box_prob = torch.zeros(
+ flat_anchors.size(0),
+ self.cls_out_channels).type_as(bbox_pred)
+ else:
+ # box_localization: a_{j}^{loc}, shape: [j, 4]
+ pred_boxes = self.bbox_coder.decode(flat_anchors, bbox_pred)
+
+ # object_box_iou: IoU_{ij}^{loc}, shape: [i, j]
+ object_box_iou = bbox_overlaps(gt_bboxes, pred_boxes)
+
+ # object_box_prob: P{a_{j} -> b_{i}}, shape: [i, j]
+ t1 = self.bbox_thr
+ t2 = object_box_iou.max(
+ dim=1, keepdim=True).values.clamp(min=t1 + 1e-12)
+ object_box_prob = ((object_box_iou - t1) / (t2 - t1)).clamp(
+ min=0, max=1)
+
+ # object_cls_box_prob: P{a_{j} -> b_{i}}, shape: [i, c, j]
+ num_obj = gt_labels.size(0)
+ indices = torch.stack(
+ [torch.arange(num_obj).type_as(gt_labels), gt_labels],
+ dim=0)
+ object_cls_box_prob = torch.sparse_coo_tensor(
+ indices, object_box_prob)
+
+ # image_box_iou: P{a_{j} \in A_{+}}, shape: [c, j]
+ """
+ from "start" to "end" implement:
+ image_box_iou = torch.sparse.max(object_cls_box_prob,
+ dim=0).t()
+
+ """
+ # start
+ box_cls_prob = torch.sparse.sum(
+ object_cls_box_prob, dim=0).to_dense()
+
+ indices = torch.nonzero(box_cls_prob, as_tuple=False).t_()
+ if indices.numel() == 0:
+ image_box_prob = torch.zeros(
+ flat_anchors.size(0),
+ self.cls_out_channels).type_as(object_box_prob)
+ else:
+ nonzero_box_prob = torch.where(
+ (gt_labels.unsqueeze(dim=-1) == indices[0]),
+ object_box_prob[:, indices[1]],
+ torch.tensor(
+ [0]).type_as(object_box_prob)).max(dim=0).values
+
+ # upmap to shape [j, c]
+ image_box_prob = torch.sparse_coo_tensor(
+ indices.flip([0]),
+ nonzero_box_prob,
+ size=(flat_anchors.size(0),
+ self.cls_out_channels)).to_dense()
+ # end
+ box_prob = image_box_prob
+
+ # construct bags for objects
+ match_quality_matrix = bbox_overlaps(gt_bboxes, flat_anchors)
+ _, matched = torch.topk(
+ match_quality_matrix, self.pre_anchor_topk, dim=1, sorted=False)
+ del match_quality_matrix
+
+ # matched_cls_prob: P_{ij}^{cls}
+ matched_cls_prob = torch.gather(
+ cls_prob[matched], 2,
+ gt_labels.view(-1, 1, 1).repeat(1, self.pre_anchor_topk,
+ 1)).squeeze(2)
+
+ # matched_box_prob: P_{ij}^{loc}
+ matched_anchors = flat_anchors[matched]
+ matched_object_targets = self.bbox_coder.encode(
+ matched_anchors,
+ gt_bboxes.unsqueeze(dim=1).expand_as(matched_anchors))
+ loss_bbox = self.loss_bbox(
+ bbox_pred[matched],
+ matched_object_targets,
+ reduction_override='none').sum(-1)
+ matched_box_prob = torch.exp(-loss_bbox)
+
+ # positive_losses: {-log( Mean-max(P_{ij}^{cls} * P_{ij}^{loc}) )}
+ num_pos = len(gt_bboxes)
+ positive_loss = self.positive_bag_loss(matched_cls_prob,
+ matched_box_prob)
+
+ return box_prob, positive_loss, num_pos
+
+ def positive_bag_loss(self, matched_cls_prob: Tensor,
+ matched_box_prob: Tensor) -> Tensor:
+ """Compute positive bag loss.
+
+ :math:`-log( Mean-max(P_{ij}^{cls} * P_{ij}^{loc}) )`.
+
+ :math:`P_{ij}^{cls}`: matched_cls_prob, classification probability of matched samples.
+
+ :math:`P_{ij}^{loc}`: matched_box_prob, box probability of matched samples.
+
+ Args:
+ matched_cls_prob (Tensor): Classification probability of matched
+ samples in shape (num_gt, pre_anchor_topk).
+ matched_box_prob (Tensor): BBox probability of matched samples,
+ in shape (num_gt, pre_anchor_topk).
+
+ Returns:
+ Tensor: Positive bag loss in shape (num_gt,).
+ """ # noqa: E501, W605
+ # bag_prob = Mean-max(matched_prob)
+ matched_prob = matched_cls_prob * matched_box_prob
+ weight = 1 / torch.clamp(1 - matched_prob, 1e-12, None)
+ weight /= weight.sum(dim=1).unsqueeze(dim=-1)
+ bag_prob = (weight * matched_prob).sum(dim=1)
+ # positive_bag_loss = -self.alpha * log(bag_prob)
+ return self.alpha * F.binary_cross_entropy(
+ bag_prob, torch.ones_like(bag_prob), reduction='none')
+
+ def negative_bag_loss(self, cls_prob: Tensor, box_prob: Tensor) -> Tensor:
+ """Compute negative bag loss.
+
+ :math:`FL((1 - P_{a_{j} \in A_{+}}) * (1 - P_{j}^{bg}))`.
+
+ :math:`P_{a_{j} \in A_{+}}`: Box_probability of matched samples.
+
+ :math:`P_{j}^{bg}`: Classification probability of negative samples.
+
+ Args:
+ cls_prob (Tensor): Classification probability, in shape
+ (num_img, num_anchors, num_classes).
+ box_prob (Tensor): Box probability, in shape
+ (num_img, num_anchors, num_classes).
+
+ Returns:
+ Tensor: Negative bag loss in shape (num_img, num_anchors,
+ num_classes).
+ """ # noqa: E501, W605
+ prob = cls_prob * (1 - box_prob)
+ # There are some cases when neg_prob = 0.
+ # This will cause the neg_prob.log() to be inf without clamp.
+ prob = prob.clamp(min=EPS, max=1 - EPS)
+ negative_bag_loss = prob**self.gamma * F.binary_cross_entropy(
+ prob, torch.zeros_like(prob), reduction='none')
+ return (1 - self.alpha) * negative_bag_loss
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/fsaf_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/fsaf_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..0a01c487406693253eb17b883cac9ed06cf95802
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/fsaf_head.py
@@ -0,0 +1,458 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, List, Optional, Tuple
+
+import numpy as np
+import torch
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import InstanceList, OptInstanceList, OptMultiConfig
+from ..losses.accuracy import accuracy
+from ..losses.utils import weight_reduce_loss
+from ..task_modules.prior_generators import anchor_inside_flags
+from ..utils import images_to_levels, multi_apply, unmap
+from .retina_head import RetinaHead
+
+
+@MODELS.register_module()
+class FSAFHead(RetinaHead):
+ """Anchor-free head used in `FSAF `_.
+
+ The head contains two subnetworks. The first classifies anchor boxes and
+ the second regresses deltas for the anchors (num_anchors is 1 for anchor-
+ free methods)
+
+ Args:
+ *args: Same as its base class in :class:`RetinaHead`
+ score_threshold (float, optional): The score_threshold to calculate
+ positive recall. If given, prediction scores lower than this value
+ is counted as incorrect prediction. Defaults to None.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict]): Initialization config dict.
+ **kwargs: Same as its base class in :class:`RetinaHead`
+
+ Example:
+ >>> import torch
+ >>> self = FSAFHead(11, 7)
+ >>> x = torch.rand(1, 7, 32, 32)
+ >>> cls_score, bbox_pred = self.forward_single(x)
+ >>> # Each anchor predicts a score for each class except background
+ >>> cls_per_anchor = cls_score.shape[1] / self.num_anchors
+ >>> box_per_anchor = bbox_pred.shape[1] / self.num_anchors
+ >>> assert cls_per_anchor == self.num_classes
+ >>> assert box_per_anchor == 4
+ """
+
+ def __init__(self,
+ *args,
+ score_threshold: Optional[float] = None,
+ init_cfg: OptMultiConfig = None,
+ **kwargs) -> None:
+ # The positive bias in self.retina_reg conv is to prevent predicted \
+ # bbox with 0 area
+ if init_cfg is None:
+ init_cfg = dict(
+ type='Normal',
+ layer='Conv2d',
+ std=0.01,
+ override=[
+ dict(
+ type='Normal',
+ name='retina_cls',
+ std=0.01,
+ bias_prob=0.01),
+ dict(
+ type='Normal', name='retina_reg', std=0.01, bias=0.25)
+ ])
+ super().__init__(*args, init_cfg=init_cfg, **kwargs)
+ self.score_threshold = score_threshold
+
+ def forward_single(self, x: Tensor) -> Tuple[Tensor, Tensor]:
+ """Forward feature map of a single scale level.
+
+ Args:
+ x (Tensor): Feature map of a single scale level.
+
+ Returns:
+ tuple[Tensor, Tensor]:
+
+ - cls_score (Tensor): Box scores for each scale level Has \
+ shape (N, num_points * num_classes, H, W).
+ - bbox_pred (Tensor): Box energies / deltas for each scale \
+ level with shape (N, num_points * 4, H, W).
+ """
+ cls_score, bbox_pred = super().forward_single(x)
+ # relu: TBLR encoder only accepts positive bbox_pred
+ return cls_score, self.relu(bbox_pred)
+
+ def _get_targets_single(self,
+ flat_anchors: Tensor,
+ valid_flags: Tensor,
+ gt_instances: InstanceData,
+ img_meta: dict,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ unmap_outputs: bool = True) -> tuple:
+ """Compute regression and classification targets for anchors in a
+ single image.
+
+ Most of the codes are the same with the base class :obj: `AnchorHead`,
+ except that it also collects and returns the matched gt index in the
+ image (from 0 to num_gt-1). If the anchor bbox is not matched to any
+ gt, the corresponding value in pos_gt_inds is -1.
+
+ Args:
+ flat_anchors (Tensor): Multi-level anchors of the image, which are
+ concatenated into a single tensor of shape (num_anchors, 4)
+ valid_flags (Tensor): Multi level valid flags of the image,
+ which are concatenated into a single tensor of
+ shape (num_anchors, ).
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes`` and ``labels``
+ attributes.
+ img_meta (dict): Meta information for current image.
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ unmap_outputs (bool): Whether to map outputs back to the original
+ set of anchors. Defaults to True.
+ """
+ inside_flags = anchor_inside_flags(flat_anchors, valid_flags,
+ img_meta['img_shape'][:2],
+ self.train_cfg['allowed_border'])
+ if not inside_flags.any():
+ raise ValueError(
+ 'There is no valid anchor inside the image boundary. Please '
+ 'check the image size and anchor sizes, or set '
+ '``allowed_border`` to -1 to skip the condition.')
+ # Assign gt and sample anchors
+ anchors = flat_anchors[inside_flags.type(torch.bool), :]
+
+ pred_instances = InstanceData(priors=anchors)
+ assign_result = self.assigner.assign(pred_instances, gt_instances,
+ gt_instances_ignore)
+ sampling_result = self.sampler.sample(assign_result, pred_instances,
+ gt_instances)
+
+ num_valid_anchors = anchors.shape[0]
+ bbox_targets = torch.zeros_like(anchors)
+ bbox_weights = torch.zeros_like(anchors)
+ labels = anchors.new_full((num_valid_anchors, ),
+ self.num_classes,
+ dtype=torch.long)
+ label_weights = anchors.new_zeros(
+ (num_valid_anchors, self.cls_out_channels), dtype=torch.float)
+ pos_gt_inds = anchors.new_full((num_valid_anchors, ),
+ -1,
+ dtype=torch.long)
+
+ pos_inds = sampling_result.pos_inds
+ neg_inds = sampling_result.neg_inds
+
+ if len(pos_inds) > 0:
+ if not self.reg_decoded_bbox:
+ pos_bbox_targets = self.bbox_coder.encode(
+ sampling_result.pos_bboxes, sampling_result.pos_gt_bboxes)
+ else:
+ # When the regression loss (e.g. `IouLoss`, `GIouLoss`)
+ # is applied directly on the decoded bounding boxes, both
+ # the predicted boxes and regression targets should be with
+ # absolute coordinate format.
+ pos_bbox_targets = sampling_result.pos_gt_bboxes
+ bbox_targets[pos_inds, :] = pos_bbox_targets
+ bbox_weights[pos_inds, :] = 1.0
+ # The assigned gt_index for each anchor. (0-based)
+ pos_gt_inds[pos_inds] = sampling_result.pos_assigned_gt_inds
+ labels[pos_inds] = sampling_result.pos_gt_labels
+ if self.train_cfg['pos_weight'] <= 0:
+ label_weights[pos_inds] = 1.0
+ else:
+ label_weights[pos_inds] = self.train_cfg['pos_weight']
+
+ if len(neg_inds) > 0:
+ label_weights[neg_inds] = 1.0
+
+ # shadowed_labels is a tensor composed of tuples
+ # (anchor_inds, class_label) that indicate those anchors lying in the
+ # outer region of a gt or overlapped by another gt with a smaller
+ # area.
+ #
+ # Therefore, only the shadowed labels are ignored for loss calculation.
+ # the key `shadowed_labels` is defined in :obj:`CenterRegionAssigner`
+ shadowed_labels = assign_result.get_extra_property('shadowed_labels')
+ if shadowed_labels is not None and shadowed_labels.numel():
+ if len(shadowed_labels.shape) == 2:
+ idx_, label_ = shadowed_labels[:, 0], shadowed_labels[:, 1]
+ assert (labels[idx_] != label_).all(), \
+ 'One label cannot be both positive and ignored'
+ label_weights[idx_, label_] = 0
+ else:
+ label_weights[shadowed_labels] = 0
+
+ # map up to original set of anchors
+ if unmap_outputs:
+ num_total_anchors = flat_anchors.size(0)
+ labels = unmap(
+ labels, num_total_anchors, inside_flags,
+ fill=self.num_classes) # fill bg label
+ label_weights = unmap(label_weights, num_total_anchors,
+ inside_flags)
+ bbox_targets = unmap(bbox_targets, num_total_anchors, inside_flags)
+ bbox_weights = unmap(bbox_weights, num_total_anchors, inside_flags)
+ pos_gt_inds = unmap(
+ pos_gt_inds, num_total_anchors, inside_flags, fill=-1)
+
+ return (labels, label_weights, bbox_targets, bbox_weights, pos_inds,
+ neg_inds, sampling_result, pos_gt_inds)
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None
+ ) -> Dict[str, Tensor]:
+ """Compute loss of the head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ Has shape (N, num_points * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_points * 4, H, W).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ for i in range(len(bbox_preds)): # loop over fpn level
+ # avoid 0 area of the predicted bbox
+ bbox_preds[i] = bbox_preds[i].clamp(min=1e-4)
+ # TODO: It may directly use the base-class loss function.
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ assert len(featmap_sizes) == self.prior_generator.num_levels
+ batch_size = len(batch_img_metas)
+ device = cls_scores[0].device
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+ cls_reg_targets = self.get_targets(
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore,
+ return_sampling_results=True)
+ (labels_list, label_weights_list, bbox_targets_list, bbox_weights_list,
+ avg_factor, sampling_results_list,
+ pos_assigned_gt_inds_list) = cls_reg_targets
+
+ num_gts = np.array(list(map(len, batch_gt_instances)))
+ # anchor number of multi levels
+ num_level_anchors = [anchors.size(0) for anchors in anchor_list[0]]
+ # concat all level anchors and flags to a single tensor
+ concat_anchor_list = []
+ for i in range(len(anchor_list)):
+ concat_anchor_list.append(torch.cat(anchor_list[i]))
+ all_anchor_list = images_to_levels(concat_anchor_list,
+ num_level_anchors)
+ losses_cls, losses_bbox = multi_apply(
+ self.loss_by_feat_single,
+ cls_scores,
+ bbox_preds,
+ all_anchor_list,
+ labels_list,
+ label_weights_list,
+ bbox_targets_list,
+ bbox_weights_list,
+ avg_factor=avg_factor)
+
+ # `pos_assigned_gt_inds_list` (length: fpn_levels) stores the assigned
+ # gt index of each anchor bbox in each fpn level.
+ cum_num_gts = list(np.cumsum(num_gts)) # length of batch_size
+ for i, assign in enumerate(pos_assigned_gt_inds_list):
+ # loop over fpn levels
+ for j in range(1, batch_size):
+ # loop over batch size
+ # Convert gt indices in each img to those in the batch
+ assign[j][assign[j] >= 0] += int(cum_num_gts[j - 1])
+ pos_assigned_gt_inds_list[i] = assign.flatten()
+ labels_list[i] = labels_list[i].flatten()
+ num_gts = num_gts.sum() # total number of gt in the batch
+ # The unique label index of each gt in the batch
+ label_sequence = torch.arange(num_gts, device=device)
+ # Collect the average loss of each gt in each level
+ with torch.no_grad():
+ loss_levels, = multi_apply(
+ self.collect_loss_level_single,
+ losses_cls,
+ losses_bbox,
+ pos_assigned_gt_inds_list,
+ labels_seq=label_sequence)
+ # Shape: (fpn_levels, num_gts). Loss of each gt at each fpn level
+ loss_levels = torch.stack(loss_levels, dim=0)
+ # Locate the best fpn level for loss back-propagation
+ if loss_levels.numel() == 0: # zero gt
+ argmin = loss_levels.new_empty((num_gts, ), dtype=torch.long)
+ else:
+ _, argmin = loss_levels.min(dim=0)
+
+ # Reweight the loss of each (anchor, label) pair, so that only those
+ # at the best gt level are back-propagated.
+ losses_cls, losses_bbox, pos_inds = multi_apply(
+ self.reweight_loss_single,
+ losses_cls,
+ losses_bbox,
+ pos_assigned_gt_inds_list,
+ labels_list,
+ list(range(len(losses_cls))),
+ min_levels=argmin)
+ num_pos = torch.cat(pos_inds, 0).sum().float()
+ pos_recall = self.calculate_pos_recall(cls_scores, labels_list,
+ pos_inds)
+
+ if num_pos == 0: # No gt
+ num_total_neg = sum(
+ [results.num_neg for results in sampling_results_list])
+ avg_factor = num_pos + num_total_neg
+ else:
+ avg_factor = num_pos
+ for i in range(len(losses_cls)):
+ losses_cls[i] /= avg_factor
+ losses_bbox[i] /= avg_factor
+ return dict(
+ loss_cls=losses_cls,
+ loss_bbox=losses_bbox,
+ num_pos=num_pos / batch_size,
+ pos_recall=pos_recall)
+
+ def calculate_pos_recall(self, cls_scores: List[Tensor],
+ labels_list: List[Tensor],
+ pos_inds: List[Tensor]) -> Tensor:
+ """Calculate positive recall with score threshold.
+
+ Args:
+ cls_scores (list[Tensor]): Classification scores at all fpn levels.
+ Each tensor is in shape (N, num_classes * num_anchors, H, W)
+ labels_list (list[Tensor]): The label that each anchor is assigned
+ to. Shape (N * H * W * num_anchors, )
+ pos_inds (list[Tensor]): List of bool tensors indicating whether
+ the anchor is assigned to a positive label.
+ Shape (N * H * W * num_anchors, )
+
+ Returns:
+ Tensor: A single float number indicating the positive recall.
+ """
+ with torch.no_grad():
+ num_class = self.num_classes
+ scores = [
+ cls.permute(0, 2, 3, 1).reshape(-1, num_class)[pos]
+ for cls, pos in zip(cls_scores, pos_inds)
+ ]
+ labels = [
+ label.reshape(-1)[pos]
+ for label, pos in zip(labels_list, pos_inds)
+ ]
+ scores = torch.cat(scores, dim=0)
+ labels = torch.cat(labels, dim=0)
+ if self.use_sigmoid_cls:
+ scores = scores.sigmoid()
+ else:
+ scores = scores.softmax(dim=1)
+
+ return accuracy(scores, labels, thresh=self.score_threshold)
+
+ def collect_loss_level_single(self, cls_loss: Tensor, reg_loss: Tensor,
+ assigned_gt_inds: Tensor,
+ labels_seq: Tensor) -> Tensor:
+ """Get the average loss in each FPN level w.r.t. each gt label.
+
+ Args:
+ cls_loss (Tensor): Classification loss of each feature map pixel,
+ shape (num_anchor, num_class)
+ reg_loss (Tensor): Regression loss of each feature map pixel,
+ shape (num_anchor, 4)
+ assigned_gt_inds (Tensor): It indicates which gt the prior is
+ assigned to (0-based, -1: no assignment). shape (num_anchor),
+ labels_seq: The rank of labels. shape (num_gt)
+
+ Returns:
+ Tensor: shape (num_gt), average loss of each gt in this level
+ """
+ if len(reg_loss.shape) == 2: # iou loss has shape (num_prior, 4)
+ reg_loss = reg_loss.sum(dim=-1) # sum loss in tblr dims
+ if len(cls_loss.shape) == 2:
+ cls_loss = cls_loss.sum(dim=-1) # sum loss in class dims
+ loss = cls_loss + reg_loss
+ assert loss.size(0) == assigned_gt_inds.size(0)
+ # Default loss value is 1e6 for a layer where no anchor is positive
+ # to ensure it will not be chosen to back-propagate gradient
+ losses_ = loss.new_full(labels_seq.shape, 1e6)
+ for i, l in enumerate(labels_seq):
+ match = assigned_gt_inds == l
+ if match.any():
+ losses_[i] = loss[match].mean()
+ return losses_,
+
+ def reweight_loss_single(self, cls_loss: Tensor, reg_loss: Tensor,
+ assigned_gt_inds: Tensor, labels: Tensor,
+ level: int, min_levels: Tensor) -> tuple:
+ """Reweight loss values at each level.
+
+ Reassign loss values at each level by masking those where the
+ pre-calculated loss is too large. Then return the reduced losses.
+
+ Args:
+ cls_loss (Tensor): Element-wise classification loss.
+ Shape: (num_anchors, num_classes)
+ reg_loss (Tensor): Element-wise regression loss.
+ Shape: (num_anchors, 4)
+ assigned_gt_inds (Tensor): The gt indices that each anchor bbox
+ is assigned to. -1 denotes a negative anchor, otherwise it is the
+ gt index (0-based). Shape: (num_anchors, ),
+ labels (Tensor): Label assigned to anchors. Shape: (num_anchors, ).
+ level (int): The current level index in the pyramid
+ (0-4 for RetinaNet)
+ min_levels (Tensor): The best-matching level for each gt.
+ Shape: (num_gts, ),
+
+ Returns:
+ tuple:
+
+ - cls_loss: Reduced corrected classification loss. Scalar.
+ - reg_loss: Reduced corrected regression loss. Scalar.
+ - pos_flags (Tensor): Corrected bool tensor indicating the \
+ final positive anchors. Shape: (num_anchors, ).
+ """
+ loc_weight = torch.ones_like(reg_loss)
+ cls_weight = torch.ones_like(cls_loss)
+ pos_flags = assigned_gt_inds >= 0 # positive pixel flag
+ pos_indices = torch.nonzero(pos_flags, as_tuple=False).flatten()
+
+ if pos_flags.any(): # pos pixels exist
+ pos_assigned_gt_inds = assigned_gt_inds[pos_flags]
+ zeroing_indices = (min_levels[pos_assigned_gt_inds] != level)
+ neg_indices = pos_indices[zeroing_indices]
+
+ if neg_indices.numel():
+ pos_flags[neg_indices] = 0
+ loc_weight[neg_indices] = 0
+ # Only the weight corresponding to the label is
+ # zeroed out if not selected
+ zeroing_labels = labels[neg_indices]
+ assert (zeroing_labels >= 0).all()
+ cls_weight[neg_indices, zeroing_labels] = 0
+
+ # Weighted loss for both cls and reg loss
+ cls_loss = weight_reduce_loss(cls_loss, cls_weight, reduction='sum')
+ reg_loss = weight_reduce_loss(reg_loss, loc_weight, reduction='sum')
+
+ return cls_loss, reg_loss, pos_flags
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/ga_retina_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/ga_retina_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..569910b365126e90638256f0d10addfa230fd141
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/ga_retina_head.py
@@ -0,0 +1,120 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Tuple
+
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from mmcv.ops import MaskedConv2d
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import OptConfigType, OptMultiConfig
+from .guided_anchor_head import FeatureAdaption, GuidedAnchorHead
+
+
+@MODELS.register_module()
+class GARetinaHead(GuidedAnchorHead):
+ """Guided-Anchor-based RetinaNet head."""
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: int,
+ stacked_convs: int = 4,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: OptConfigType = None,
+ init_cfg: OptMultiConfig = None,
+ **kwargs) -> None:
+ if init_cfg is None:
+ init_cfg = dict(
+ type='Normal',
+ layer='Conv2d',
+ std=0.01,
+ override=[
+ dict(
+ type='Normal',
+ name='conv_loc',
+ std=0.01,
+ bias_prob=0.01),
+ dict(
+ type='Normal',
+ name='retina_cls',
+ std=0.01,
+ bias_prob=0.01)
+ ])
+ self.stacked_convs = stacked_convs
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ super().__init__(
+ num_classes=num_classes,
+ in_channels=in_channels,
+ init_cfg=init_cfg,
+ **kwargs)
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self.relu = nn.ReLU(inplace=True)
+ self.cls_convs = nn.ModuleList()
+ self.reg_convs = nn.ModuleList()
+ for i in range(self.stacked_convs):
+ chn = self.in_channels if i == 0 else self.feat_channels
+ self.cls_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ self.reg_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+
+ self.conv_loc = nn.Conv2d(self.feat_channels, 1, 1)
+ num_anchors = self.square_anchor_generator.num_base_priors[0]
+ self.conv_shape = nn.Conv2d(self.feat_channels, num_anchors * 2, 1)
+ self.feature_adaption_cls = FeatureAdaption(
+ self.feat_channels,
+ self.feat_channels,
+ kernel_size=3,
+ deform_groups=self.deform_groups)
+ self.feature_adaption_reg = FeatureAdaption(
+ self.feat_channels,
+ self.feat_channels,
+ kernel_size=3,
+ deform_groups=self.deform_groups)
+ self.retina_cls = MaskedConv2d(
+ self.feat_channels,
+ self.num_base_priors * self.cls_out_channels,
+ 3,
+ padding=1)
+ self.retina_reg = MaskedConv2d(
+ self.feat_channels, self.num_base_priors * 4, 3, padding=1)
+
+ def forward_single(self, x: Tensor) -> Tuple[Tensor]:
+ """Forward feature map of a single scale level."""
+ cls_feat = x
+ reg_feat = x
+ for cls_conv in self.cls_convs:
+ cls_feat = cls_conv(cls_feat)
+ for reg_conv in self.reg_convs:
+ reg_feat = reg_conv(reg_feat)
+
+ loc_pred = self.conv_loc(cls_feat)
+ shape_pred = self.conv_shape(reg_feat)
+
+ cls_feat = self.feature_adaption_cls(cls_feat, shape_pred)
+ reg_feat = self.feature_adaption_reg(reg_feat, shape_pred)
+
+ if not self.training:
+ mask = loc_pred.sigmoid()[0] >= self.loc_filter_thr
+ else:
+ mask = None
+ cls_score = self.retina_cls(cls_feat, mask)
+ bbox_pred = self.retina_reg(reg_feat, mask)
+ return cls_score, bbox_pred, shape_pred, loc_pred
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/ga_rpn_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/ga_rpn_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..9614463165533358b8465420a87dfa47e7de1177
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/ga_rpn_head.py
@@ -0,0 +1,222 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from typing import List, Tuple
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.ops import nms
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, InstanceList, MultiConfig, OptInstanceList
+from .guided_anchor_head import GuidedAnchorHead
+
+
+@MODELS.register_module()
+class GARPNHead(GuidedAnchorHead):
+ """Guided-Anchor-based RPN head."""
+
+ def __init__(self,
+ in_channels: int,
+ num_classes: int = 1,
+ init_cfg: MultiConfig = dict(
+ type='Normal',
+ layer='Conv2d',
+ std=0.01,
+ override=dict(
+ type='Normal',
+ name='conv_loc',
+ std=0.01,
+ bias_prob=0.01)),
+ **kwargs) -> None:
+ super().__init__(
+ num_classes=num_classes,
+ in_channels=in_channels,
+ init_cfg=init_cfg,
+ **kwargs)
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self.rpn_conv = nn.Conv2d(
+ self.in_channels, self.feat_channels, 3, padding=1)
+ super(GARPNHead, self)._init_layers()
+
+ def forward_single(self, x: Tensor) -> Tuple[Tensor]:
+ """Forward feature of a single scale level."""
+
+ x = self.rpn_conv(x)
+ x = F.relu(x, inplace=True)
+ (cls_score, bbox_pred, shape_pred,
+ loc_pred) = super().forward_single(x)
+ return cls_score, bbox_pred, shape_pred, loc_pred
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ shape_preds: List[Tensor],
+ loc_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ has shape (N, num_anchors * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W).
+ shape_preds (list[Tensor]): shape predictions for each scale
+ level with shape (N, 1, H, W).
+ loc_preds (list[Tensor]): location predictions for each scale
+ level with shape (N, num_anchors * 2, H, W).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ losses = super().loss_by_feat(
+ cls_scores,
+ bbox_preds,
+ shape_preds,
+ loc_preds,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore)
+ return dict(
+ loss_rpn_cls=losses['loss_cls'],
+ loss_rpn_bbox=losses['loss_bbox'],
+ loss_anchor_shape=losses['loss_shape'],
+ loss_anchor_loc=losses['loss_loc'])
+
+ def _predict_by_feat_single(self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ mlvl_anchors: List[Tensor],
+ mlvl_masks: List[Tensor],
+ img_meta: dict,
+ cfg: ConfigType,
+ rescale: bool = False) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores from all scale
+ levels of a single image, each item has shape
+ (num_priors * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas from
+ all scale levels of a single image, each item has shape
+ (num_priors * 4, H, W).
+ mlvl_anchors (list[Tensor]): Each element in the list is
+ the anchors of a single level in feature pyramid. it has
+ shape (num_priors, 4).
+ mlvl_masks (list[Tensor]): Each element in the list is location
+ masks of a single level.
+ img_meta (dict): Image meta info.
+ cfg (:obj:`ConfigDict` or dict): Test / postprocessing
+ configuration, if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4), the last
+ dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ cfg = self.test_cfg if cfg is None else cfg
+ cfg = copy.deepcopy(cfg)
+ assert cfg.nms.get('type', 'nms') == 'nms', 'GARPNHead only support ' \
+ 'naive nms.'
+
+ mlvl_proposals = []
+ for idx in range(len(cls_scores)):
+ rpn_cls_score = cls_scores[idx]
+ rpn_bbox_pred = bbox_preds[idx]
+ anchors = mlvl_anchors[idx]
+ mask = mlvl_masks[idx]
+ assert rpn_cls_score.size()[-2:] == rpn_bbox_pred.size()[-2:]
+ # if no location is kept, end.
+ if mask.sum() == 0:
+ continue
+ rpn_cls_score = rpn_cls_score.permute(1, 2, 0)
+ if self.use_sigmoid_cls:
+ rpn_cls_score = rpn_cls_score.reshape(-1)
+ scores = rpn_cls_score.sigmoid()
+ else:
+ rpn_cls_score = rpn_cls_score.reshape(-1, 2)
+ # remind that we set FG labels to [0, num_class-1]
+ # since mmdet v2.0
+ # BG cat_id: num_class
+ scores = rpn_cls_score.softmax(dim=1)[:, :-1]
+ # filter scores, bbox_pred w.r.t. mask.
+ # anchors are filtered in get_anchors() beforehand.
+ scores = scores[mask]
+ rpn_bbox_pred = rpn_bbox_pred.permute(1, 2, 0).reshape(-1,
+ 4)[mask, :]
+ if scores.dim() == 0:
+ rpn_bbox_pred = rpn_bbox_pred.unsqueeze(0)
+ anchors = anchors.unsqueeze(0)
+ scores = scores.unsqueeze(0)
+ # filter anchors, bbox_pred, scores w.r.t. scores
+ if cfg.nms_pre > 0 and scores.shape[0] > cfg.nms_pre:
+ _, topk_inds = scores.topk(cfg.nms_pre)
+ rpn_bbox_pred = rpn_bbox_pred[topk_inds, :]
+ anchors = anchors[topk_inds, :]
+ scores = scores[topk_inds]
+ # get proposals w.r.t. anchors and rpn_bbox_pred
+ proposals = self.bbox_coder.decode(
+ anchors, rpn_bbox_pred, max_shape=img_meta['img_shape'])
+ # filter out too small bboxes
+ if cfg.min_bbox_size >= 0:
+ w = proposals[:, 2] - proposals[:, 0]
+ h = proposals[:, 3] - proposals[:, 1]
+ valid_mask = (w > cfg.min_bbox_size) & (h > cfg.min_bbox_size)
+ if not valid_mask.all():
+ proposals = proposals[valid_mask]
+ scores = scores[valid_mask]
+
+ # NMS in current level
+ proposals, _ = nms(proposals, scores, cfg.nms.iou_threshold)
+ proposals = proposals[:cfg.nms_post, :]
+ mlvl_proposals.append(proposals)
+ proposals = torch.cat(mlvl_proposals, 0)
+ if cfg.get('nms_across_levels', False):
+ # NMS across multi levels
+ proposals, _ = nms(proposals[:, :4], proposals[:, -1],
+ cfg.nms.iou_threshold)
+ proposals = proposals[:cfg.max_per_img, :]
+ else:
+ scores = proposals[:, 4]
+ num = min(cfg.max_per_img, proposals.shape[0])
+ _, topk_inds = scores.topk(num)
+ proposals = proposals[topk_inds, :]
+
+ bboxes = proposals[:, :-1]
+ scores = proposals[:, -1]
+ if rescale:
+ assert img_meta.get('scale_factor') is not None
+ bboxes /= bboxes.new_tensor(img_meta['scale_factor']).repeat(
+ (1, 2))
+
+ results = InstanceData()
+ results.bboxes = bboxes
+ results.scores = scores
+ results.labels = scores.new_zeros(scores.size(0), dtype=torch.long)
+ return results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/gfl_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/gfl_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..be43d9b4da39da602b3b87bd3c9739c67367615b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/gfl_head.py
@@ -0,0 +1,667 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Sequence, Tuple
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule, Scale
+from mmengine.config import ConfigDict
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures.bbox import bbox_overlaps
+from mmdet.utils import (ConfigType, InstanceList, MultiConfig, OptConfigType,
+ OptInstanceList, reduce_mean)
+from ..task_modules.prior_generators import anchor_inside_flags
+from ..task_modules.samplers import PseudoSampler
+from ..utils import (filter_scores_and_topk, images_to_levels, multi_apply,
+ unmap)
+from .anchor_head import AnchorHead
+
+
+class Integral(nn.Module):
+ """A fixed layer for calculating integral result from distribution.
+
+ This layer calculates the target location by :math: ``sum{P(y_i) * y_i}``,
+ P(y_i) denotes the softmax vector that represents the discrete distribution
+ y_i denotes the discrete set, usually {0, 1, 2, ..., reg_max}
+
+ Args:
+ reg_max (int): The maximal value of the discrete set. Defaults to 16.
+ You may want to reset it according to your new dataset or related
+ settings.
+ """
+
+ def __init__(self, reg_max: int = 16) -> None:
+ super().__init__()
+ self.reg_max = reg_max
+ self.register_buffer('project',
+ torch.linspace(0, self.reg_max, self.reg_max + 1))
+
+ def forward(self, x: Tensor) -> Tensor:
+ """Forward feature from the regression head to get integral result of
+ bounding box location.
+
+ Args:
+ x (Tensor): Features of the regression head, shape (N, 4*(n+1)),
+ n is self.reg_max.
+
+ Returns:
+ x (Tensor): Integral result of box locations, i.e., distance
+ offsets from the box center in four directions, shape (N, 4).
+ """
+ x = F.softmax(x.reshape(-1, self.reg_max + 1), dim=1)
+ x = F.linear(x, self.project.type_as(x)).reshape(-1, 4)
+ return x
+
+
+@MODELS.register_module()
+class GFLHead(AnchorHead):
+ """Generalized Focal Loss: Learning Qualified and Distributed Bounding
+ Boxes for Dense Object Detection.
+
+ GFL head structure is similar with ATSS, however GFL uses
+ 1) joint representation for classification and localization quality, and
+ 2) flexible General distribution for bounding box locations,
+ which are supervised by
+ Quality Focal Loss (QFL) and Distribution Focal Loss (DFL), respectively
+
+ https://arxiv.org/abs/2006.04388
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ stacked_convs (int): Number of conv layers in cls and reg tower.
+ Defaults to 4.
+ conv_cfg (:obj:`ConfigDict` or dict, optional): dictionary to construct
+ and config conv layer. Defaults to None.
+ norm_cfg (:obj:`ConfigDict` or dict): dictionary to construct and
+ config norm layer. Default: dict(type='GN', num_groups=32,
+ requires_grad=True).
+ loss_qfl (:obj:`ConfigDict` or dict): Config of Quality Focal Loss
+ (QFL).
+ bbox_coder (:obj:`ConfigDict` or dict): Config of bbox coder. Defaults
+ to 'DistancePointBBoxCoder'.
+ reg_max (int): Max value of integral set :math: ``{0, ..., reg_max}``
+ in QFL setting. Defaults to 16.
+ init_cfg (:obj:`ConfigDict` or dict or list[dict] or
+ list[:obj:`ConfigDict`]): Initialization config dict.
+ Example:
+ >>> self = GFLHead(11, 7)
+ >>> feats = [torch.rand(1, 7, s, s) for s in [4, 8, 16, 32, 64]]
+ >>> cls_quality_score, bbox_pred = self.forward(feats)
+ >>> assert len(cls_quality_score) == len(self.scales)
+ """
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: int,
+ stacked_convs: int = 4,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: ConfigType = dict(
+ type='GN', num_groups=32, requires_grad=True),
+ loss_dfl: ConfigType = dict(
+ type='DistributionFocalLoss', loss_weight=0.25),
+ bbox_coder: ConfigType = dict(type='DistancePointBBoxCoder'),
+ reg_max: int = 16,
+ init_cfg: MultiConfig = dict(
+ type='Normal',
+ layer='Conv2d',
+ std=0.01,
+ override=dict(
+ type='Normal',
+ name='gfl_cls',
+ std=0.01,
+ bias_prob=0.01)),
+ **kwargs) -> None:
+ self.stacked_convs = stacked_convs
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self.reg_max = reg_max
+ super().__init__(
+ num_classes=num_classes,
+ in_channels=in_channels,
+ bbox_coder=bbox_coder,
+ init_cfg=init_cfg,
+ **kwargs)
+
+ if self.train_cfg:
+ self.assigner = TASK_UTILS.build(self.train_cfg['assigner'])
+ if self.train_cfg.get('sampler', None) is not None:
+ self.sampler = TASK_UTILS.build(
+ self.train_cfg['sampler'], default_args=dict(context=self))
+ else:
+ self.sampler = PseudoSampler(context=self)
+
+ self.integral = Integral(self.reg_max)
+ self.loss_dfl = MODELS.build(loss_dfl)
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self.relu = nn.ReLU()
+ self.cls_convs = nn.ModuleList()
+ self.reg_convs = nn.ModuleList()
+ for i in range(self.stacked_convs):
+ chn = self.in_channels if i == 0 else self.feat_channels
+ self.cls_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ self.reg_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ assert self.num_anchors == 1, 'anchor free version'
+ self.gfl_cls = nn.Conv2d(
+ self.feat_channels, self.cls_out_channels, 3, padding=1)
+ self.gfl_reg = nn.Conv2d(
+ self.feat_channels, 4 * (self.reg_max + 1), 3, padding=1)
+ self.scales = nn.ModuleList(
+ [Scale(1.0) for _ in self.prior_generator.strides])
+
+ def forward(self, x: Tuple[Tensor]) -> Tuple[List[Tensor]]:
+ """Forward features from the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: Usually a tuple of classification scores and bbox prediction
+
+ - cls_scores (list[Tensor]): Classification and quality (IoU)
+ joint scores for all scale levels, each is a 4D-tensor,
+ the channel number is num_classes.
+ - bbox_preds (list[Tensor]): Box distribution logits for all
+ scale levels, each is a 4D-tensor, the channel number is
+ 4*(n+1), n is max value of integral set.
+ """
+ return multi_apply(self.forward_single, x, self.scales)
+
+ def forward_single(self, x: Tensor, scale: Scale) -> Sequence[Tensor]:
+ """Forward feature of a single scale level.
+
+ Args:
+ x (Tensor): Features of a single scale level.
+ scale (:obj: `mmcv.cnn.Scale`): Learnable scale module to resize
+ the bbox prediction.
+
+ Returns:
+ tuple:
+
+ - cls_score (Tensor): Cls and quality joint scores for a single
+ scale level the channel number is num_classes.
+ - bbox_pred (Tensor): Box distribution logits for a single scale
+ level, the channel number is 4*(n+1), n is max value of
+ integral set.
+ """
+ cls_feat = x
+ reg_feat = x
+ for cls_conv in self.cls_convs:
+ cls_feat = cls_conv(cls_feat)
+ for reg_conv in self.reg_convs:
+ reg_feat = reg_conv(reg_feat)
+ cls_score = self.gfl_cls(cls_feat)
+ bbox_pred = scale(self.gfl_reg(reg_feat)).float()
+ return cls_score, bbox_pred
+
+ def anchor_center(self, anchors: Tensor) -> Tensor:
+ """Get anchor centers from anchors.
+
+ Args:
+ anchors (Tensor): Anchor list with shape (N, 4), ``xyxy`` format.
+
+ Returns:
+ Tensor: Anchor centers with shape (N, 2), ``xy`` format.
+ """
+ anchors_cx = (anchors[..., 2] + anchors[..., 0]) / 2
+ anchors_cy = (anchors[..., 3] + anchors[..., 1]) / 2
+ return torch.stack([anchors_cx, anchors_cy], dim=-1)
+
+ def loss_by_feat_single(self, anchors: Tensor, cls_score: Tensor,
+ bbox_pred: Tensor, labels: Tensor,
+ label_weights: Tensor, bbox_targets: Tensor,
+ stride: Tuple[int], avg_factor: int) -> dict:
+ """Calculate the loss of a single scale level based on the features
+ extracted by the detection head.
+
+ Args:
+ anchors (Tensor): Box reference for each scale level with shape
+ (N, num_total_anchors, 4).
+ cls_score (Tensor): Cls and quality joint scores for each scale
+ level has shape (N, num_classes, H, W).
+ bbox_pred (Tensor): Box distribution logits for each scale
+ level with shape (N, 4*(n+1), H, W), n is max value of integral
+ set.
+ labels (Tensor): Labels of each anchors with shape
+ (N, num_total_anchors).
+ label_weights (Tensor): Label weights of each anchor with shape
+ (N, num_total_anchors)
+ bbox_targets (Tensor): BBox regression targets of each anchor with
+ shape (N, num_total_anchors, 4).
+ stride (Tuple[int]): Stride in this scale level.
+ avg_factor (int): Average factor that is used to average
+ the loss. When using sampling method, avg_factor is usually
+ the sum of positive and negative priors. When using
+ `PseudoSampler`, `avg_factor` is usually equal to the number
+ of positive priors.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ assert stride[0] == stride[1], 'h stride is not equal to w stride!'
+ anchors = anchors.reshape(-1, 4)
+ cls_score = cls_score.permute(0, 2, 3,
+ 1).reshape(-1, self.cls_out_channels)
+ bbox_pred = bbox_pred.permute(0, 2, 3,
+ 1).reshape(-1, 4 * (self.reg_max + 1))
+ bbox_targets = bbox_targets.reshape(-1, 4)
+ labels = labels.reshape(-1)
+ label_weights = label_weights.reshape(-1)
+
+ # FG cat_id: [0, num_classes -1], BG cat_id: num_classes
+ bg_class_ind = self.num_classes
+ pos_inds = ((labels >= 0)
+ & (labels < bg_class_ind)).nonzero().squeeze(1)
+ score = label_weights.new_zeros(labels.shape)
+
+ if len(pos_inds) > 0:
+ pos_bbox_targets = bbox_targets[pos_inds]
+ pos_bbox_pred = bbox_pred[pos_inds]
+ pos_anchors = anchors[pos_inds]
+ pos_anchor_centers = self.anchor_center(pos_anchors) / stride[0]
+
+ weight_targets = cls_score.detach().sigmoid()
+ weight_targets = weight_targets.max(dim=1)[0][pos_inds]
+ pos_bbox_pred_corners = self.integral(pos_bbox_pred)
+ pos_decode_bbox_pred = self.bbox_coder.decode(
+ pos_anchor_centers, pos_bbox_pred_corners)
+ pos_decode_bbox_targets = pos_bbox_targets / stride[0]
+ score[pos_inds] = bbox_overlaps(
+ pos_decode_bbox_pred.detach(),
+ pos_decode_bbox_targets,
+ is_aligned=True)
+ pred_corners = pos_bbox_pred.reshape(-1, self.reg_max + 1)
+ target_corners = self.bbox_coder.encode(pos_anchor_centers,
+ pos_decode_bbox_targets,
+ self.reg_max).reshape(-1)
+
+ # regression loss
+ loss_bbox = self.loss_bbox(
+ pos_decode_bbox_pred,
+ pos_decode_bbox_targets,
+ weight=weight_targets,
+ avg_factor=1.0)
+
+ # dfl loss
+ loss_dfl = self.loss_dfl(
+ pred_corners,
+ target_corners,
+ weight=weight_targets[:, None].expand(-1, 4).reshape(-1),
+ avg_factor=4.0)
+ else:
+ loss_bbox = bbox_pred.sum() * 0
+ loss_dfl = bbox_pred.sum() * 0
+ weight_targets = bbox_pred.new_tensor(0)
+
+ # cls (qfl) loss
+ loss_cls = self.loss_cls(
+ cls_score, (labels, score),
+ weight=label_weights,
+ avg_factor=avg_factor)
+
+ return loss_cls, loss_bbox, loss_dfl, weight_targets.sum()
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Cls and quality scores for each scale
+ level has shape (N, num_classes, H, W).
+ bbox_preds (list[Tensor]): Box distribution logits for each scale
+ level with shape (N, 4*(n+1), H, W), n is max value of integral
+ set.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ assert len(featmap_sizes) == self.prior_generator.num_levels
+
+ device = cls_scores[0].device
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+
+ cls_reg_targets = self.get_targets(
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore)
+
+ (anchor_list, labels_list, label_weights_list, bbox_targets_list,
+ bbox_weights_list, avg_factor) = cls_reg_targets
+
+ avg_factor = reduce_mean(
+ torch.tensor(avg_factor, dtype=torch.float, device=device)).item()
+
+ losses_cls, losses_bbox, losses_dfl,\
+ avg_factor = multi_apply(
+ self.loss_by_feat_single,
+ anchor_list,
+ cls_scores,
+ bbox_preds,
+ labels_list,
+ label_weights_list,
+ bbox_targets_list,
+ self.prior_generator.strides,
+ avg_factor=avg_factor)
+
+ avg_factor = sum(avg_factor)
+ avg_factor = reduce_mean(avg_factor).clamp_(min=1).item()
+ losses_bbox = list(map(lambda x: x / avg_factor, losses_bbox))
+ losses_dfl = list(map(lambda x: x / avg_factor, losses_dfl))
+ return dict(
+ loss_cls=losses_cls, loss_bbox=losses_bbox, loss_dfl=losses_dfl)
+
+ def _predict_by_feat_single(self,
+ cls_score_list: List[Tensor],
+ bbox_pred_list: List[Tensor],
+ score_factor_list: List[Tensor],
+ mlvl_priors: List[Tensor],
+ img_meta: dict,
+ cfg: ConfigDict,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results.
+
+ Args:
+ cls_score_list (list[Tensor]): Box scores from all scale
+ levels of a single image, each item has shape
+ (num_priors * num_classes, H, W).
+ bbox_pred_list (list[Tensor]): Box energies / deltas from
+ all scale levels of a single image, each item has shape
+ (num_priors * 4, H, W).
+ score_factor_list (list[Tensor]): Score factor from all scale
+ levels of a single image. GFL head does not need this value.
+ mlvl_priors (list[Tensor]): Each element in the list is
+ the priors of a single level in feature pyramid, has shape
+ (num_priors, 4).
+ img_meta (dict): Image meta info.
+ cfg (:obj: `ConfigDict`): Test / postprocessing configuration,
+ if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ tuple[Tensor]: Results of detected bboxes and labels. If with_nms
+ is False and mlvl_score_factor is None, return mlvl_bboxes and
+ mlvl_scores, else return mlvl_bboxes, mlvl_scores and
+ mlvl_score_factor. Usually with_nms is False is used for aug
+ test. If with_nms is True, then return the following format
+
+ - det_bboxes (Tensor): Predicted bboxes with shape
+ [num_bboxes, 5], where the first 4 columns are bounding
+ box positions (tl_x, tl_y, br_x, br_y) and the 5-th
+ column are scores between 0 and 1.
+ - det_labels (Tensor): Predicted labels of the corresponding
+ box with shape [num_bboxes].
+ """
+ cfg = self.test_cfg if cfg is None else cfg
+ img_shape = img_meta['img_shape']
+ nms_pre = cfg.get('nms_pre', -1)
+
+ mlvl_bboxes = []
+ mlvl_scores = []
+ mlvl_labels = []
+ for level_idx, (cls_score, bbox_pred, stride, priors) in enumerate(
+ zip(cls_score_list, bbox_pred_list,
+ self.prior_generator.strides, mlvl_priors)):
+ assert cls_score.size()[-2:] == bbox_pred.size()[-2:]
+ assert stride[0] == stride[1]
+
+ bbox_pred = bbox_pred.permute(1, 2, 0)
+ bbox_pred = self.integral(bbox_pred) * stride[0]
+
+ scores = cls_score.permute(1, 2, 0).reshape(
+ -1, self.cls_out_channels).sigmoid()
+
+ # After https://github.com/open-mmlab/mmdetection/pull/6268/,
+ # this operation keeps fewer bboxes under the same `nms_pre`.
+ # There is no difference in performance for most models. If you
+ # find a slight drop in performance, you can set a larger
+ # `nms_pre` than before.
+ results = filter_scores_and_topk(
+ scores, cfg.score_thr, nms_pre,
+ dict(bbox_pred=bbox_pred, priors=priors))
+ scores, labels, _, filtered_results = results
+
+ bbox_pred = filtered_results['bbox_pred']
+ priors = filtered_results['priors']
+
+ bboxes = self.bbox_coder.decode(
+ self.anchor_center(priors), bbox_pred, max_shape=img_shape)
+ mlvl_bboxes.append(bboxes)
+ mlvl_scores.append(scores)
+ mlvl_labels.append(labels)
+
+ results = InstanceData()
+ results.bboxes = torch.cat(mlvl_bboxes)
+ results.scores = torch.cat(mlvl_scores)
+ results.labels = torch.cat(mlvl_labels)
+
+ return self._bbox_post_process(
+ results=results,
+ cfg=cfg,
+ rescale=rescale,
+ with_nms=with_nms,
+ img_meta=img_meta)
+
+ def get_targets(self,
+ anchor_list: List[Tensor],
+ valid_flag_list: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None,
+ unmap_outputs=True) -> tuple:
+ """Get targets for GFL head.
+
+ This method is almost the same as `AnchorHead.get_targets()`. Besides
+ returning the targets as the parent method does, it also returns the
+ anchors as the first element of the returned tuple.
+ """
+ num_imgs = len(batch_img_metas)
+ assert len(anchor_list) == len(valid_flag_list) == num_imgs
+
+ # anchor number of multi levels
+ num_level_anchors = [anchors.size(0) for anchors in anchor_list[0]]
+ num_level_anchors_list = [num_level_anchors] * num_imgs
+
+ # concat all level anchors and flags to a single tensor
+ for i in range(num_imgs):
+ assert len(anchor_list[i]) == len(valid_flag_list[i])
+ anchor_list[i] = torch.cat(anchor_list[i])
+ valid_flag_list[i] = torch.cat(valid_flag_list[i])
+
+ # compute targets for each image
+ if batch_gt_instances_ignore is None:
+ batch_gt_instances_ignore = [None] * num_imgs
+ (all_anchors, all_labels, all_label_weights, all_bbox_targets,
+ all_bbox_weights, pos_inds_list, neg_inds_list,
+ sampling_results_list) = multi_apply(
+ self._get_targets_single,
+ anchor_list,
+ valid_flag_list,
+ num_level_anchors_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore,
+ unmap_outputs=unmap_outputs)
+ # Get `avg_factor` of all images, which calculate in `SamplingResult`.
+ # When using sampling method, avg_factor is usually the sum of
+ # positive and negative priors. When using `PseudoSampler`,
+ # `avg_factor` is usually equal to the number of positive priors.
+ avg_factor = sum(
+ [results.avg_factor for results in sampling_results_list])
+ # split targets to a list w.r.t. multiple levels
+ anchors_list = images_to_levels(all_anchors, num_level_anchors)
+ labels_list = images_to_levels(all_labels, num_level_anchors)
+ label_weights_list = images_to_levels(all_label_weights,
+ num_level_anchors)
+ bbox_targets_list = images_to_levels(all_bbox_targets,
+ num_level_anchors)
+ bbox_weights_list = images_to_levels(all_bbox_weights,
+ num_level_anchors)
+ return (anchors_list, labels_list, label_weights_list,
+ bbox_targets_list, bbox_weights_list, avg_factor)
+
+ def _get_targets_single(self,
+ flat_anchors: Tensor,
+ valid_flags: Tensor,
+ num_level_anchors: List[int],
+ gt_instances: InstanceData,
+ img_meta: dict,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ unmap_outputs: bool = True) -> tuple:
+ """Compute regression, classification targets for anchors in a single
+ image.
+
+ Args:
+ flat_anchors (Tensor): Multi-level anchors of the image, which are
+ concatenated into a single tensor of shape (num_anchors, 4)
+ valid_flags (Tensor): Multi level valid flags of the image,
+ which are concatenated into a single tensor of
+ shape (num_anchors,).
+ num_level_anchors (list[int]): Number of anchors of each scale
+ level.
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ img_meta (dict): Meta information for current image.
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ unmap_outputs (bool): Whether to map outputs back to the original
+ set of anchors. Defaults to True.
+
+ Returns:
+ tuple: N is the number of total anchors in the image.
+
+ - anchors (Tensor): All anchors in the image with shape (N, 4).
+ - labels (Tensor): Labels of all anchors in the image with
+ shape (N,).
+ - label_weights (Tensor): Label weights of all anchor in the
+ image with shape (N,).
+ - bbox_targets (Tensor): BBox targets of all anchors in the
+ image with shape (N, 4).
+ - bbox_weights (Tensor): BBox weights of all anchors in the
+ image with shape (N, 4).
+ - pos_inds (Tensor): Indices of positive anchor with shape
+ (num_pos,).
+ - neg_inds (Tensor): Indices of negative anchor with shape
+ (num_neg,).
+ - sampling_result (:obj:`SamplingResult`): Sampling results.
+ """
+ inside_flags = anchor_inside_flags(flat_anchors, valid_flags,
+ img_meta['img_shape'][:2],
+ self.train_cfg['allowed_border'])
+ if not inside_flags.any():
+ raise ValueError(
+ 'There is no valid anchor inside the image boundary. Please '
+ 'check the image size and anchor sizes, or set '
+ '``allowed_border`` to -1 to skip the condition.')
+ # assign gt and sample anchors
+ anchors = flat_anchors[inside_flags, :]
+ num_level_anchors_inside = self.get_num_level_anchors_inside(
+ num_level_anchors, inside_flags)
+ pred_instances = InstanceData(priors=anchors)
+ assign_result = self.assigner.assign(
+ pred_instances=pred_instances,
+ num_level_priors=num_level_anchors_inside,
+ gt_instances=gt_instances,
+ gt_instances_ignore=gt_instances_ignore)
+
+ sampling_result = self.sampler.sample(
+ assign_result=assign_result,
+ pred_instances=pred_instances,
+ gt_instances=gt_instances)
+
+ num_valid_anchors = anchors.shape[0]
+ bbox_targets = torch.zeros_like(anchors)
+ bbox_weights = torch.zeros_like(anchors)
+ labels = anchors.new_full((num_valid_anchors, ),
+ self.num_classes,
+ dtype=torch.long)
+ label_weights = anchors.new_zeros(num_valid_anchors, dtype=torch.float)
+
+ pos_inds = sampling_result.pos_inds
+ neg_inds = sampling_result.neg_inds
+ if len(pos_inds) > 0:
+ pos_bbox_targets = sampling_result.pos_gt_bboxes
+ bbox_targets[pos_inds, :] = pos_bbox_targets
+ bbox_weights[pos_inds, :] = 1.0
+
+ labels[pos_inds] = sampling_result.pos_gt_labels
+ if self.train_cfg['pos_weight'] <= 0:
+ label_weights[pos_inds] = 1.0
+ else:
+ label_weights[pos_inds] = self.train_cfg['pos_weight']
+ if len(neg_inds) > 0:
+ label_weights[neg_inds] = 1.0
+
+ # map up to original set of anchors
+ if unmap_outputs:
+ num_total_anchors = flat_anchors.size(0)
+ anchors = unmap(anchors, num_total_anchors, inside_flags)
+ labels = unmap(
+ labels, num_total_anchors, inside_flags, fill=self.num_classes)
+ label_weights = unmap(label_weights, num_total_anchors,
+ inside_flags)
+ bbox_targets = unmap(bbox_targets, num_total_anchors, inside_flags)
+ bbox_weights = unmap(bbox_weights, num_total_anchors, inside_flags)
+
+ return (anchors, labels, label_weights, bbox_targets, bbox_weights,
+ pos_inds, neg_inds, sampling_result)
+
+ def get_num_level_anchors_inside(self, num_level_anchors: List[int],
+ inside_flags: Tensor) -> List[int]:
+ """Get the number of valid anchors in every level."""
+
+ split_inside_flags = torch.split(inside_flags, num_level_anchors)
+ num_level_anchors_inside = [
+ int(flags.sum()) for flags in split_inside_flags
+ ]
+ return num_level_anchors_inside
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/grounding_dino_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/grounding_dino_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..8088322546f24ae6f3e60aff1378d5c2feefdcf0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/grounding_dino_head.py
@@ -0,0 +1,774 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import math
+from typing import Dict, List, Optional, Tuple, Union
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import Linear
+from mmengine.model import constant_init
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.models.losses import QualityFocalLoss
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import bbox_cxcywh_to_xyxy, bbox_xyxy_to_cxcywh
+from mmdet.utils import InstanceList, reduce_mean
+from ..layers import inverse_sigmoid
+from .atss_vlfusion_head import convert_grounding_to_cls_scores
+from .dino_head import DINOHead
+
+
+class ContrastiveEmbed(nn.Module):
+ """text visual ContrastiveEmbed layer.
+
+ Args:
+ max_text_len (int, optional): Maximum length of text.
+ log_scale (Optional[Union[str, float]]): The initial value of a
+ learnable parameter to multiply with the similarity
+ matrix to normalize the output. Defaults to 0.0.
+ - If set to 'auto', the similarity matrix will be normalized by
+ a fixed value ``sqrt(d_c)`` where ``d_c`` is the channel number.
+ - If set to 'none' or ``None``, there is no normalization applied.
+ - If set to a float number, the similarity matrix will be multiplied
+ by ``exp(log_scale)``, where ``log_scale`` is learnable.
+ bias (bool, optional): Whether to add bias to the output.
+ If set to ``True``, a learnable bias that is initialized as -4.6
+ will be added to the output. Useful when training from scratch.
+ Defaults to False.
+ """
+
+ def __init__(self,
+ max_text_len: int = 256,
+ log_scale: Optional[Union[str, float]] = None,
+ bias: bool = False):
+ super().__init__()
+ self.max_text_len = max_text_len
+ self.log_scale = log_scale
+ if isinstance(log_scale, float):
+ self.log_scale = nn.Parameter(
+ torch.Tensor([float(log_scale)]), requires_grad=True)
+ elif log_scale not in ['auto', 'none', None]:
+ raise ValueError(f'log_scale should be one of '
+ f'"auto", "none", None, but got {log_scale}')
+
+ self.bias = None
+ if bias:
+ bias_value = -math.log((1 - 0.01) / 0.01)
+ self.bias = nn.Parameter(
+ torch.Tensor([bias_value]), requires_grad=True)
+
+ def forward(self, visual_feat: Tensor, text_feat: Tensor,
+ text_token_mask: Tensor) -> Tensor:
+ """Forward function.
+
+ Args:
+ visual_feat (Tensor): Visual features.
+ text_feat (Tensor): Text features.
+ text_token_mask (Tensor): A mask used for text feats.
+
+ Returns:
+ Tensor: Classification score.
+ """
+ res = visual_feat @ text_feat.transpose(-1, -2)
+ if isinstance(self.log_scale, nn.Parameter):
+ res = res * self.log_scale.exp()
+ elif self.log_scale == 'auto':
+ # NOTE: similar to the normalizer in self-attention
+ res = res / math.sqrt(visual_feat.shape[-1])
+ if self.bias is not None:
+ res = res + self.bias
+ res.masked_fill_(~text_token_mask[:, None, :], float('-inf'))
+
+ new_res = torch.full((*res.shape[:-1], self.max_text_len),
+ float('-inf'),
+ device=res.device)
+ new_res[..., :res.shape[-1]] = res
+
+ return new_res
+
+
+@MODELS.register_module()
+class GroundingDINOHead(DINOHead):
+ """Head of the Grounding DINO: Marrying DINO with Grounded Pre-Training for
+ Open-Set Object Detection.
+
+ Args:
+ contrastive_cfg (dict, optional): Contrastive config that contains
+ keys like ``max_text_len``. Defaults to dict(max_text_len=256).
+ """
+
+ def __init__(self, contrastive_cfg=dict(max_text_len=256), **kwargs):
+ self.contrastive_cfg = contrastive_cfg
+ self.max_text_len = contrastive_cfg.get('max_text_len', 256)
+ super().__init__(**kwargs)
+
+ def _init_layers(self) -> None:
+ """Initialize classification branch and regression branch of head."""
+ fc_cls = ContrastiveEmbed(**self.contrastive_cfg)
+ reg_branch = []
+ for _ in range(self.num_reg_fcs):
+ reg_branch.append(Linear(self.embed_dims, self.embed_dims))
+ reg_branch.append(nn.ReLU())
+ reg_branch.append(Linear(self.embed_dims, 4))
+ reg_branch = nn.Sequential(*reg_branch)
+
+ # NOTE: due to the fc_cls is a contrastive embedding and don't
+ # have any trainable parameters,we do not need to copy it.
+ if self.share_pred_layer:
+ self.cls_branches = nn.ModuleList(
+ [fc_cls for _ in range(self.num_pred_layer)])
+ self.reg_branches = nn.ModuleList(
+ [reg_branch for _ in range(self.num_pred_layer)])
+ else:
+ self.cls_branches = nn.ModuleList(
+ [copy.deepcopy(fc_cls) for _ in range(self.num_pred_layer)])
+ self.reg_branches = nn.ModuleList([
+ copy.deepcopy(reg_branch) for _ in range(self.num_pred_layer)
+ ])
+
+ def init_weights(self) -> None:
+ """Initialize weights of the Deformable DETR head."""
+ for m in self.reg_branches:
+ constant_init(m[-1], 0, bias=0)
+ nn.init.constant_(self.reg_branches[0][-1].bias.data[2:], -2.0)
+ if self.as_two_stage:
+ for m in self.reg_branches:
+ nn.init.constant_(m[-1].bias.data[2:], 0.0)
+
+ def _get_targets_single(self, cls_score: Tensor, bbox_pred: Tensor,
+ gt_instances: InstanceData,
+ img_meta: dict) -> tuple:
+ """Compute regression and classification targets for one image.
+
+ Outputs from a single decoder layer of a single feature level are used.
+
+ Args:
+ cls_score (Tensor): Box score logits from a single decoder layer
+ for one image. Shape [num_queries, cls_out_channels].
+ bbox_pred (Tensor): Sigmoid outputs from a single decoder layer
+ for one image, with normalized coordinate (cx, cy, w, h) and
+ shape [num_queries, 4].
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes`` and ``labels``
+ attributes.
+ img_meta (dict): Meta information for one image.
+
+ Returns:
+ tuple[Tensor]: a tuple containing the following for one image.
+
+ - labels (Tensor): Labels of each image.
+ - label_weights (Tensor]): Label weights of each image.
+ - bbox_targets (Tensor): BBox targets of each image.
+ - bbox_weights (Tensor): BBox weights of each image.
+ - pos_inds (Tensor): Sampled positive indices for each image.
+ - neg_inds (Tensor): Sampled negative indices for each image.
+ """
+ img_h, img_w = img_meta['img_shape']
+ factor = bbox_pred.new_tensor([img_w, img_h, img_w,
+ img_h]).unsqueeze(0)
+ num_bboxes = bbox_pred.size(0)
+ # convert bbox_pred from xywh, normalized to xyxy, unnormalized
+ bbox_pred = bbox_cxcywh_to_xyxy(bbox_pred)
+ bbox_pred = bbox_pred * factor
+
+ pred_instances = InstanceData(scores=cls_score, bboxes=bbox_pred)
+ # assigner and sampler
+ assign_result = self.assigner.assign(
+ pred_instances=pred_instances,
+ gt_instances=gt_instances,
+ img_meta=img_meta)
+ gt_bboxes = gt_instances.bboxes
+
+ pos_inds = torch.nonzero(
+ assign_result.gt_inds > 0, as_tuple=False).squeeze(-1).unique()
+ neg_inds = torch.nonzero(
+ assign_result.gt_inds == 0, as_tuple=False).squeeze(-1).unique()
+ pos_assigned_gt_inds = assign_result.gt_inds[pos_inds] - 1
+ pos_gt_bboxes = gt_bboxes[pos_assigned_gt_inds.long(), :]
+
+ # Major changes. The labels are 0-1 binary labels for each bbox
+ # and text tokens.
+ labels = gt_bboxes.new_full((num_bboxes, self.max_text_len),
+ 0,
+ dtype=torch.float32)
+ labels[pos_inds] = gt_instances.positive_maps[pos_assigned_gt_inds]
+ label_weights = gt_bboxes.new_ones(num_bboxes)
+
+ # bbox targets
+ bbox_targets = torch.zeros_like(bbox_pred, dtype=gt_bboxes.dtype)
+ bbox_weights = torch.zeros_like(bbox_pred, dtype=gt_bboxes.dtype)
+ bbox_weights[pos_inds] = 1.0
+
+ # DETR regress the relative position of boxes (cxcywh) in the image.
+ # Thus the learning target should be normalized by the image size, also
+ # the box format should be converted from defaultly x1y1x2y2 to cxcywh.
+ pos_gt_bboxes_normalized = pos_gt_bboxes / factor
+ pos_gt_bboxes_targets = bbox_xyxy_to_cxcywh(pos_gt_bboxes_normalized)
+ bbox_targets[pos_inds] = pos_gt_bboxes_targets
+ return (labels, label_weights, bbox_targets, bbox_weights, pos_inds,
+ neg_inds)
+
+ def forward(
+ self,
+ hidden_states: Tensor,
+ references: List[Tensor],
+ memory_text: Tensor,
+ text_token_mask: Tensor,
+ ) -> Tuple[Tensor]:
+ """Forward function.
+
+ Args:
+ hidden_states (Tensor): Hidden states output from each decoder
+ layer, has shape (num_decoder_layers, bs, num_queries, dim).
+ references (List[Tensor]): List of the reference from the decoder.
+ The first reference is the `init_reference` (initial) and the
+ other num_decoder_layers(6) references are `inter_references`
+ (intermediate). The `init_reference` has shape (bs,
+ num_queries, 4) when `as_two_stage` of the detector is `True`,
+ otherwise (bs, num_queries, 2). Each `inter_reference` has
+ shape (bs, num_queries, 4) when `with_box_refine` of the
+ detector is `True`, otherwise (bs, num_queries, 2). The
+ coordinates are arranged as (cx, cy) when the last dimension is
+ 2, and (cx, cy, w, h) when it is 4.
+ memory_text (Tensor): Memory text. It has shape (bs, len_text,
+ text_embed_dims).
+ text_token_mask (Tensor): Text token mask. It has shape (bs,
+ len_text).
+
+ Returns:
+ tuple[Tensor]: results of head containing the following tensor.
+
+ - all_layers_outputs_classes (Tensor): Outputs from the
+ classification head, has shape (num_decoder_layers, bs,
+ num_queries, cls_out_channels).
+ - all_layers_outputs_coords (Tensor): Sigmoid outputs from the
+ regression head with normalized coordinate format (cx, cy, w,
+ h), has shape (num_decoder_layers, bs, num_queries, 4) with the
+ last dimension arranged as (cx, cy, w, h).
+ """
+ all_layers_outputs_classes = []
+ all_layers_outputs_coords = []
+
+ for layer_id in range(hidden_states.shape[0]):
+ reference = inverse_sigmoid(references[layer_id])
+ # NOTE The last reference will not be used.
+ hidden_state = hidden_states[layer_id]
+ outputs_class = self.cls_branches[layer_id](hidden_state,
+ memory_text,
+ text_token_mask)
+ tmp_reg_preds = self.reg_branches[layer_id](hidden_state)
+ if reference.shape[-1] == 4:
+ # When `layer` is 0 and `as_two_stage` of the detector
+ # is `True`, or when `layer` is greater than 0 and
+ # `with_box_refine` of the detector is `True`.
+ tmp_reg_preds += reference
+ else:
+ # When `layer` is 0 and `as_two_stage` of the detector
+ # is `False`, or when `layer` is greater than 0 and
+ # `with_box_refine` of the detector is `False`.
+ assert reference.shape[-1] == 2
+ tmp_reg_preds[..., :2] += reference
+ outputs_coord = tmp_reg_preds.sigmoid()
+ all_layers_outputs_classes.append(outputs_class)
+ all_layers_outputs_coords.append(outputs_coord)
+
+ all_layers_outputs_classes = torch.stack(all_layers_outputs_classes)
+ all_layers_outputs_coords = torch.stack(all_layers_outputs_coords)
+
+ return all_layers_outputs_classes, all_layers_outputs_coords
+
+ def predict(self,
+ hidden_states: Tensor,
+ references: List[Tensor],
+ memory_text: Tensor,
+ text_token_mask: Tensor,
+ batch_data_samples: SampleList,
+ rescale: bool = True) -> InstanceList:
+ """Perform forward propagation and loss calculation of the detection
+ head on the queries of the upstream network.
+
+ Args:
+ hidden_states (Tensor): Hidden states output from each decoder
+ layer, has shape (num_decoder_layers, num_queries, bs, dim).
+ references (List[Tensor]): List of the reference from the decoder.
+ The first reference is the `init_reference` (initial) and the
+ other num_decoder_layers(6) references are `inter_references`
+ (intermediate). The `init_reference` has shape (bs,
+ num_queries, 4) when `as_two_stage` of the detector is `True`,
+ otherwise (bs, num_queries, 2). Each `inter_reference` has
+ shape (bs, num_queries, 4) when `with_box_refine` of the
+ detector is `True`, otherwise (bs, num_queries, 2). The
+ coordinates are arranged as (cx, cy) when the last dimension is
+ 2, and (cx, cy, w, h) when it is 4.
+ memory_text (Tensor): Memory text. It has shape (bs, len_text,
+ text_embed_dims).
+ text_token_mask (Tensor): Text token mask. It has shape (bs,
+ len_text).
+ batch_data_samples (SampleList): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool, optional): If `True`, return boxes in original
+ image space. Defaults to `True`.
+
+ Returns:
+ InstanceList: Detection results of each image
+ after the post process.
+ """
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+ batch_token_positive_maps = [
+ data_samples.token_positive_map
+ for data_samples in batch_data_samples
+ ]
+
+ outs = self(hidden_states, references, memory_text, text_token_mask)
+
+ predictions = self.predict_by_feat(
+ *outs,
+ batch_img_metas=batch_img_metas,
+ batch_token_positive_maps=batch_token_positive_maps,
+ rescale=rescale)
+ return predictions
+
+ def predict_by_feat(self,
+ all_layers_cls_scores: Tensor,
+ all_layers_bbox_preds: Tensor,
+ batch_img_metas: List[Dict],
+ batch_token_positive_maps: Optional[List[dict]] = None,
+ rescale: bool = False) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ bbox results.
+
+ Args:
+ all_layers_cls_scores (Tensor): Classification scores of all
+ decoder layers, has shape (num_decoder_layers, bs, num_queries,
+ cls_out_channels).
+ all_layers_bbox_preds (Tensor): Regression outputs of all decoder
+ layers. Each is a 4D-tensor with normalized coordinate format
+ (cx, cy, w, h) and shape (num_decoder_layers, bs, num_queries,
+ 4) with the last dimension arranged as (cx, cy, w, h).
+ batch_img_metas (List[Dict]): _description_
+ batch_token_positive_maps (list[dict], Optional): Batch token
+ positive map. Defaults to None.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Object detection results of each image
+ after the post process. Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ cls_scores = all_layers_cls_scores[-1]
+ bbox_preds = all_layers_bbox_preds[-1]
+ result_list = []
+ for img_id in range(len(batch_img_metas)):
+ cls_score = cls_scores[img_id]
+ bbox_pred = bbox_preds[img_id]
+ img_meta = batch_img_metas[img_id]
+ token_positive_maps = batch_token_positive_maps[img_id]
+ results = self._predict_by_feat_single(cls_score, bbox_pred,
+ token_positive_maps,
+ img_meta, rescale)
+ result_list.append(results)
+ return result_list
+
+ def _predict_by_feat_single(self,
+ cls_score: Tensor,
+ bbox_pred: Tensor,
+ token_positive_maps: dict,
+ img_meta: dict,
+ rescale: bool = True) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results.
+
+ Args:
+ cls_score (Tensor): Box score logits from the last decoder layer
+ for each image. Shape [num_queries, cls_out_channels].
+ bbox_pred (Tensor): Sigmoid outputs from the last decoder layer
+ for each image, with coordinate format (cx, cy, w, h) and
+ shape [num_queries, 4].
+ token_positive_maps (dict): Token positive map.
+ img_meta (dict): Image meta info.
+ rescale (bool, optional): If True, return boxes in original image
+ space. Default True.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ assert len(cls_score) == len(bbox_pred) # num_queries
+ max_per_img = self.test_cfg.get('max_per_img', len(cls_score))
+ img_shape = img_meta['img_shape']
+
+ if token_positive_maps is not None:
+ cls_score = convert_grounding_to_cls_scores(
+ logits=cls_score.sigmoid()[None],
+ positive_maps=[token_positive_maps])[0]
+ scores, indexes = cls_score.view(-1).topk(max_per_img)
+ num_classes = cls_score.shape[-1]
+ det_labels = indexes % num_classes
+ bbox_index = indexes // num_classes
+ bbox_pred = bbox_pred[bbox_index]
+ else:
+ cls_score = cls_score.sigmoid()
+ scores, _ = cls_score.max(-1)
+ scores, indexes = scores.topk(max_per_img)
+ bbox_pred = bbox_pred[indexes]
+ det_labels = scores.new_zeros(scores.shape, dtype=torch.long)
+
+ det_bboxes = bbox_cxcywh_to_xyxy(bbox_pred)
+ det_bboxes[:, 0::2] = det_bboxes[:, 0::2] * img_shape[1]
+ det_bboxes[:, 1::2] = det_bboxes[:, 1::2] * img_shape[0]
+ det_bboxes[:, 0::2].clamp_(min=0, max=img_shape[1])
+ det_bboxes[:, 1::2].clamp_(min=0, max=img_shape[0])
+ if rescale:
+ assert img_meta.get('scale_factor') is not None
+ det_bboxes /= det_bboxes.new_tensor(
+ img_meta['scale_factor']).repeat((1, 2))
+ results = InstanceData()
+ results.bboxes = det_bboxes
+ results.scores = scores
+ results.labels = det_labels
+ return results
+
+ def loss(self, hidden_states: Tensor, references: List[Tensor],
+ memory_text: Tensor, text_token_mask: Tensor,
+ enc_outputs_class: Tensor, enc_outputs_coord: Tensor,
+ batch_data_samples: SampleList, dn_meta: Dict[str, int]) -> dict:
+ """Perform forward propagation and loss calculation of the detection
+ head on the queries of the upstream network.
+
+ Args:
+ hidden_states (Tensor): Hidden states output from each decoder
+ layer, has shape (num_decoder_layers, bs, num_queries_total,
+ dim), where `num_queries_total` is the sum of
+ `num_denoising_queries` and `num_matching_queries` when
+ `self.training` is `True`, else `num_matching_queries`.
+ references (list[Tensor]): List of the reference from the decoder.
+ The first reference is the `init_reference` (initial) and the
+ other num_decoder_layers(6) references are `inter_references`
+ (intermediate). The `init_reference` has shape (bs,
+ num_queries_total, 4) and each `inter_reference` has shape
+ (bs, num_queries, 4) with the last dimension arranged as
+ (cx, cy, w, h).
+ memory_text (Tensor): Memory text. It has shape (bs, len_text,
+ text_embed_dims).
+ enc_outputs_class (Tensor): The score of each point on encode
+ feature map, has shape (bs, num_feat_points, cls_out_channels).
+ enc_outputs_coord (Tensor): The proposal generate from the
+ encode feature map, has shape (bs, num_feat_points, 4) with the
+ last dimension arranged as (cx, cy, w, h).
+ batch_data_samples (list[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ dn_meta (Dict[str, int]): The dictionary saves information about
+ group collation, including 'num_denoising_queries' and
+ 'num_denoising_groups'. It will be used for split outputs of
+ denoising and matching parts and loss calculation.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ batch_gt_instances = []
+ batch_img_metas = []
+ for data_sample in batch_data_samples:
+ batch_img_metas.append(data_sample.metainfo)
+ batch_gt_instances.append(data_sample.gt_instances)
+
+ outs = self(hidden_states, references, memory_text, text_token_mask)
+ self.text_masks = text_token_mask
+ loss_inputs = outs + (enc_outputs_class, enc_outputs_coord,
+ batch_gt_instances, batch_img_metas, dn_meta)
+ losses = self.loss_by_feat(*loss_inputs)
+ return losses
+
+ def loss_by_feat_single(self, cls_scores: Tensor, bbox_preds: Tensor,
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict]) -> Tuple[Tensor]:
+ """Loss function for outputs from a single decoder layer of a single
+ feature level.
+
+ Args:
+ cls_scores (Tensor): Box score logits from a single decoder layer
+ for all images, has shape (bs, num_queries, cls_out_channels).
+ bbox_preds (Tensor): Sigmoid outputs from a single decoder layer
+ for all images, with normalized coordinate (cx, cy, w, h) and
+ shape (bs, num_queries, 4).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+
+ Returns:
+ Tuple[Tensor]: A tuple including `loss_cls`, `loss_box` and
+ `loss_iou`.
+ """
+ num_imgs = cls_scores.size(0)
+ cls_scores_list = [cls_scores[i] for i in range(num_imgs)]
+ bbox_preds_list = [bbox_preds[i] for i in range(num_imgs)]
+ with torch.no_grad():
+ cls_reg_targets = self.get_targets(cls_scores_list,
+ bbox_preds_list,
+ batch_gt_instances,
+ batch_img_metas)
+ (labels_list, label_weights_list, bbox_targets_list, bbox_weights_list,
+ num_total_pos, num_total_neg) = cls_reg_targets
+ labels = torch.stack(labels_list, 0)
+ label_weights = torch.stack(label_weights_list, 0)
+ bbox_targets = torch.cat(bbox_targets_list, 0)
+ bbox_weights = torch.cat(bbox_weights_list, 0)
+
+ # ===== this change =====
+ # Loss is not computed for the padded regions of the text.
+ assert (self.text_masks.dim() == 2)
+ text_masks = self.text_masks.new_zeros(
+ (self.text_masks.size(0), self.max_text_len))
+ text_masks[:, :self.text_masks.size(1)] = self.text_masks
+ text_mask = (text_masks > 0).unsqueeze(1)
+ text_mask = text_mask.repeat(1, cls_scores.size(1), 1)
+ cls_scores = torch.masked_select(cls_scores, text_mask).contiguous()
+
+ labels = torch.masked_select(labels, text_mask)
+ label_weights = label_weights[...,
+ None].repeat(1, 1, text_mask.size(-1))
+ label_weights = torch.masked_select(label_weights, text_mask)
+
+ # classification loss
+ # construct weighted avg_factor to match with the official DETR repo
+ cls_avg_factor = num_total_pos * 1.0 + \
+ num_total_neg * self.bg_cls_weight
+ if self.sync_cls_avg_factor:
+ cls_avg_factor = reduce_mean(
+ cls_scores.new_tensor([cls_avg_factor]))
+ cls_avg_factor = max(cls_avg_factor, 1)
+
+ if isinstance(self.loss_cls, QualityFocalLoss):
+ raise NotImplementedError(
+ 'QualityFocalLoss for GroundingDINOHead is not supported yet.')
+ else:
+ loss_cls = self.loss_cls(
+ cls_scores, labels, label_weights, avg_factor=cls_avg_factor)
+
+ # Compute the average number of gt boxes across all gpus, for
+ # normalization purposes
+ num_total_pos = loss_cls.new_tensor([num_total_pos])
+ num_total_pos = torch.clamp(reduce_mean(num_total_pos), min=1).item()
+
+ # construct factors used for rescale bboxes
+ factors = []
+ for img_meta, bbox_pred in zip(batch_img_metas, bbox_preds):
+ img_h, img_w, = img_meta['img_shape']
+ factor = bbox_pred.new_tensor([img_w, img_h, img_w,
+ img_h]).unsqueeze(0).repeat(
+ bbox_pred.size(0), 1)
+ factors.append(factor)
+ factors = torch.cat(factors, 0)
+
+ # DETR regress the relative position of boxes (cxcywh) in the image,
+ # thus the learning target is normalized by the image size. So here
+ # we need to re-scale them for calculating IoU loss
+ bbox_preds = bbox_preds.reshape(-1, 4)
+ bboxes = bbox_cxcywh_to_xyxy(bbox_preds) * factors
+ bboxes_gt = bbox_cxcywh_to_xyxy(bbox_targets) * factors
+
+ # regression IoU loss, defaultly GIoU loss
+ loss_iou = self.loss_iou(
+ bboxes, bboxes_gt, bbox_weights, avg_factor=num_total_pos)
+
+ # regression L1 loss
+ loss_bbox = self.loss_bbox(
+ bbox_preds, bbox_targets, bbox_weights, avg_factor=num_total_pos)
+ return loss_cls, loss_bbox, loss_iou
+
+ def _loss_dn_single(self, dn_cls_scores: Tensor, dn_bbox_preds: Tensor,
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ dn_meta: Dict[str, int]) -> Tuple[Tensor]:
+ """Denoising loss for outputs from a single decoder layer.
+
+ Args:
+ dn_cls_scores (Tensor): Classification scores of a single decoder
+ layer in denoising part, has shape (bs, num_denoising_queries,
+ cls_out_channels).
+ dn_bbox_preds (Tensor): Regression outputs of a single decoder
+ layer in denoising part. Each is a 4D-tensor with normalized
+ coordinate format (cx, cy, w, h) and has shape
+ (bs, num_denoising_queries, 4).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ dn_meta (Dict[str, int]): The dictionary saves information about
+ group collation, including 'num_denoising_queries' and
+ 'num_denoising_groups'. It will be used for split outputs of
+ denoising and matching parts and loss calculation.
+
+ Returns:
+ Tuple[Tensor]: A tuple including `loss_cls`, `loss_box` and
+ `loss_iou`.
+ """
+ cls_reg_targets = self.get_dn_targets(batch_gt_instances,
+ batch_img_metas, dn_meta)
+ (labels_list, label_weights_list, bbox_targets_list, bbox_weights_list,
+ num_total_pos, num_total_neg) = cls_reg_targets
+ labels = torch.stack(labels_list, 0)
+ label_weights = torch.stack(label_weights_list, 0)
+ bbox_targets = torch.cat(bbox_targets_list, 0)
+ bbox_weights = torch.cat(bbox_weights_list, 0)
+ # ===== this change =====
+ # Loss is not computed for the padded regions of the text.
+ assert (self.text_masks.dim() == 2)
+ text_masks = self.text_masks.new_zeros(
+ (self.text_masks.size(0), self.max_text_len))
+ text_masks[:, :self.text_masks.size(1)] = self.text_masks
+ text_mask = (text_masks > 0).unsqueeze(1)
+ text_mask = text_mask.repeat(1, dn_cls_scores.size(1), 1)
+ cls_scores = torch.masked_select(dn_cls_scores, text_mask).contiguous()
+ labels = torch.masked_select(labels, text_mask)
+ label_weights = label_weights[...,
+ None].repeat(1, 1, text_mask.size(-1))
+ label_weights = torch.masked_select(label_weights, text_mask)
+ # =======================
+
+ # classification loss
+ # construct weighted avg_factor to match with the official DETR repo
+ cls_avg_factor = \
+ num_total_pos * 1.0 + num_total_neg * self.bg_cls_weight
+ if self.sync_cls_avg_factor:
+ cls_avg_factor = reduce_mean(
+ cls_scores.new_tensor([cls_avg_factor]))
+ cls_avg_factor = max(cls_avg_factor, 1)
+
+ if len(cls_scores) > 0:
+ if isinstance(self.loss_cls, QualityFocalLoss):
+ raise NotImplementedError('QualityFocalLoss is not supported')
+ else:
+ loss_cls = self.loss_cls(
+ cls_scores,
+ labels,
+ label_weights,
+ avg_factor=cls_avg_factor)
+ else:
+ loss_cls = torch.zeros(
+ 1, dtype=cls_scores.dtype, device=cls_scores.device)
+
+ # Compute the average number of gt boxes across all gpus, for
+ # normalization purposes
+ num_total_pos = loss_cls.new_tensor([num_total_pos])
+ num_total_pos = torch.clamp(reduce_mean(num_total_pos), min=1).item()
+
+ # construct factors used for rescale bboxes
+ factors = []
+ for img_meta, bbox_pred in zip(batch_img_metas, dn_bbox_preds):
+ img_h, img_w = img_meta['img_shape']
+ factor = bbox_pred.new_tensor([img_w, img_h, img_w,
+ img_h]).unsqueeze(0).repeat(
+ bbox_pred.size(0), 1)
+ factors.append(factor)
+ factors = torch.cat(factors)
+
+ # DETR regress the relative position of boxes (cxcywh) in the image,
+ # thus the learning target is normalized by the image size. So here
+ # we need to re-scale them for calculating IoU loss
+ bbox_preds = dn_bbox_preds.reshape(-1, 4)
+ bboxes = bbox_cxcywh_to_xyxy(bbox_preds) * factors
+ bboxes_gt = bbox_cxcywh_to_xyxy(bbox_targets) * factors
+
+ # regression IoU loss, defaultly GIoU loss
+ loss_iou = self.loss_iou(
+ bboxes, bboxes_gt, bbox_weights, avg_factor=num_total_pos)
+
+ # regression L1 loss
+ loss_bbox = self.loss_bbox(
+ bbox_preds, bbox_targets, bbox_weights, avg_factor=num_total_pos)
+ return loss_cls, loss_bbox, loss_iou
+
+ def _get_dn_targets_single(self, gt_instances: InstanceData,
+ img_meta: dict, dn_meta: Dict[str,
+ int]) -> tuple:
+ """Get targets in denoising part for one image.
+
+ Args:
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes`` and ``labels``
+ attributes.
+ img_meta (dict): Meta information for one image.
+ dn_meta (Dict[str, int]): The dictionary saves information about
+ group collation, including 'num_denoising_queries' and
+ 'num_denoising_groups'. It will be used for split outputs of
+ denoising and matching parts and loss calculation.
+
+ Returns:
+ tuple[Tensor]: a tuple containing the following for one image.
+
+ - labels (Tensor): Labels of each image.
+ - label_weights (Tensor]): Label weights of each image.
+ - bbox_targets (Tensor): BBox targets of each image.
+ - bbox_weights (Tensor): BBox weights of each image.
+ - pos_inds (Tensor): Sampled positive indices for each image.
+ - neg_inds (Tensor): Sampled negative indices for each image.
+ """
+ gt_bboxes = gt_instances.bboxes
+ gt_labels = gt_instances.labels
+ num_groups = dn_meta['num_denoising_groups']
+ num_denoising_queries = dn_meta['num_denoising_queries']
+ num_queries_each_group = int(num_denoising_queries / num_groups)
+ device = gt_bboxes.device
+
+ if len(gt_labels) > 0:
+ t = torch.arange(len(gt_labels), dtype=torch.long, device=device)
+ t = t.unsqueeze(0).repeat(num_groups, 1)
+ pos_assigned_gt_inds = t.flatten()
+ pos_inds = torch.arange(
+ num_groups, dtype=torch.long, device=device)
+ pos_inds = pos_inds.unsqueeze(1) * num_queries_each_group + t
+ pos_inds = pos_inds.flatten()
+ else:
+ pos_inds = pos_assigned_gt_inds = \
+ gt_bboxes.new_tensor([], dtype=torch.long)
+
+ neg_inds = pos_inds + num_queries_each_group // 2
+ # label targets
+ # this change
+ labels = gt_bboxes.new_full((num_denoising_queries, self.max_text_len),
+ 0,
+ dtype=torch.float32)
+ labels[pos_inds] = gt_instances.positive_maps[pos_assigned_gt_inds]
+ label_weights = gt_bboxes.new_ones(num_denoising_queries)
+
+ # bbox targets
+ bbox_targets = torch.zeros(num_denoising_queries, 4, device=device)
+ bbox_weights = torch.zeros(num_denoising_queries, 4, device=device)
+ bbox_weights[pos_inds] = 1.0
+ img_h, img_w = img_meta['img_shape']
+
+ # DETR regress the relative position of boxes (cxcywh) in the image.
+ # Thus the learning target should be normalized by the image size, also
+ # the box format should be converted from defaultly x1y1x2y2 to cxcywh.
+ factor = gt_bboxes.new_tensor([img_w, img_h, img_w,
+ img_h]).unsqueeze(0)
+ gt_bboxes_normalized = gt_bboxes / factor
+ gt_bboxes_targets = bbox_xyxy_to_cxcywh(gt_bboxes_normalized)
+ bbox_targets[pos_inds] = gt_bboxes_targets.repeat([num_groups, 1])
+
+ return (labels, label_weights, bbox_targets, bbox_weights, pos_inds,
+ neg_inds)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/guided_anchor_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/guided_anchor_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..59f6dd3336e66065dc88b702e925965d4089c72f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/guided_anchor_head.py
@@ -0,0 +1,994 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple
+
+import torch
+import torch.nn as nn
+from mmcv.ops import DeformConv2d, MaskedConv2d
+from mmengine.model import BaseModule
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.utils import (ConfigType, InstanceList, MultiConfig, OptConfigType,
+ OptInstanceList)
+from ..layers import multiclass_nms
+from ..task_modules.prior_generators import anchor_inside_flags, calc_region
+from ..task_modules.samplers import PseudoSampler
+from ..utils import images_to_levels, multi_apply, unmap
+from .anchor_head import AnchorHead
+
+
+class FeatureAdaption(BaseModule):
+ """Feature Adaption Module.
+
+ Feature Adaption Module is implemented based on DCN v1.
+ It uses anchor shape prediction rather than feature map to
+ predict offsets of deform conv layer.
+
+ Args:
+ in_channels (int): Number of channels in the input feature map.
+ out_channels (int): Number of channels in the output feature map.
+ kernel_size (int): Deformable conv kernel size. Defaults to 3.
+ deform_groups (int): Deformable conv group size. Defaults to 4.
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or \
+ list[dict], optional): Initialization config dict.
+ """
+
+ def __init__(
+ self,
+ in_channels: int,
+ out_channels: int,
+ kernel_size: int = 3,
+ deform_groups: int = 4,
+ init_cfg: MultiConfig = dict(
+ type='Normal',
+ layer='Conv2d',
+ std=0.1,
+ override=dict(type='Normal', name='conv_adaption', std=0.01))
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ offset_channels = kernel_size * kernel_size * 2
+ self.conv_offset = nn.Conv2d(
+ 2, deform_groups * offset_channels, 1, bias=False)
+ self.conv_adaption = DeformConv2d(
+ in_channels,
+ out_channels,
+ kernel_size=kernel_size,
+ padding=(kernel_size - 1) // 2,
+ deform_groups=deform_groups)
+ self.relu = nn.ReLU(inplace=True)
+
+ def forward(self, x: Tensor, shape: Tensor) -> Tensor:
+ offset = self.conv_offset(shape.detach())
+ x = self.relu(self.conv_adaption(x, offset))
+ return x
+
+
+@MODELS.register_module()
+class GuidedAnchorHead(AnchorHead):
+ """Guided-Anchor-based head (GA-RPN, GA-RetinaNet, etc.).
+
+ This GuidedAnchorHead will predict high-quality feature guided
+ anchors and locations where anchors will be kept in inference.
+ There are mainly 3 categories of bounding-boxes.
+
+ - Sampled 9 pairs for target assignment. (approxes)
+ - The square boxes where the predicted anchors are based on. (squares)
+ - Guided anchors.
+
+ Please refer to https://arxiv.org/abs/1901.03278 for more details.
+
+ Args:
+ num_classes (int): Number of classes.
+ in_channels (int): Number of channels in the input feature map.
+ feat_channels (int): Number of hidden channels. Defaults to 256.
+ approx_anchor_generator (:obj:`ConfigDict` or dict): Config dict
+ for approx generator
+ square_anchor_generator (:obj:`ConfigDict` or dict): Config dict
+ for square generator
+ anchor_coder (:obj:`ConfigDict` or dict): Config dict for anchor coder
+ bbox_coder (:obj:`ConfigDict` or dict): Config dict for bbox coder
+ reg_decoded_bbox (bool): If true, the regression loss would be
+ applied directly on decoded bounding boxes, converting both
+ the predicted boxes and regression targets to absolute
+ coordinates format. Defaults to False. It should be `True` when
+ using `IoULoss`, `GIoULoss`, or `DIoULoss` in the bbox head.
+ deform_groups: (int): Group number of DCN in FeatureAdaption module.
+ Defaults to 4.
+ loc_filter_thr (float): Threshold to filter out unconcerned regions.
+ Defaults to 0.01.
+ loss_loc (:obj:`ConfigDict` or dict): Config of location loss.
+ loss_shape (:obj:`ConfigDict` or dict): Config of anchor shape loss.
+ loss_cls (:obj:`ConfigDict` or dict): Config of classification loss.
+ loss_bbox (:obj:`ConfigDict` or dict): Config of bbox regression loss.
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or \
+ list[dict], optional): Initialization config dict.
+ """
+
+ def __init__(
+ self,
+ num_classes: int,
+ in_channels: int,
+ feat_channels: int = 256,
+ approx_anchor_generator: ConfigType = dict(
+ type='AnchorGenerator',
+ octave_base_scale=8,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[4, 8, 16, 32, 64]),
+ square_anchor_generator: ConfigType = dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ scales=[8],
+ strides=[4, 8, 16, 32, 64]),
+ anchor_coder: ConfigType = dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ bbox_coder: ConfigType = dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ reg_decoded_bbox: bool = False,
+ deform_groups: int = 4,
+ loc_filter_thr: float = 0.01,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ loss_loc: ConfigType = dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_shape: ConfigType = dict(
+ type='BoundedIoULoss', beta=0.2, loss_weight=1.0),
+ loss_cls: ConfigType = dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ loss_bbox: ConfigType = dict(
+ type='SmoothL1Loss', beta=1.0, loss_weight=1.0),
+ init_cfg: MultiConfig = dict(
+ type='Normal',
+ layer='Conv2d',
+ std=0.01,
+ override=dict(
+ type='Normal', name='conv_loc', std=0.01, lbias_prob=0.01))
+ ) -> None:
+ super(AnchorHead, self).__init__(init_cfg=init_cfg)
+ self.in_channels = in_channels
+ self.num_classes = num_classes
+ self.feat_channels = feat_channels
+ self.deform_groups = deform_groups
+ self.loc_filter_thr = loc_filter_thr
+
+ # build approx_anchor_generator and square_anchor_generator
+ assert (approx_anchor_generator['octave_base_scale'] ==
+ square_anchor_generator['scales'][0])
+ assert (approx_anchor_generator['strides'] ==
+ square_anchor_generator['strides'])
+ self.approx_anchor_generator = TASK_UTILS.build(
+ approx_anchor_generator)
+ self.square_anchor_generator = TASK_UTILS.build(
+ square_anchor_generator)
+ self.approxs_per_octave = self.approx_anchor_generator \
+ .num_base_priors[0]
+
+ self.reg_decoded_bbox = reg_decoded_bbox
+
+ # one anchor per location
+ self.num_base_priors = self.square_anchor_generator.num_base_priors[0]
+
+ self.use_sigmoid_cls = loss_cls.get('use_sigmoid', False)
+ self.loc_focal_loss = loss_loc['type'] in ['FocalLoss']
+ if self.use_sigmoid_cls:
+ self.cls_out_channels = self.num_classes
+ else:
+ self.cls_out_channels = self.num_classes + 1
+
+ # build bbox_coder
+ self.anchor_coder = TASK_UTILS.build(anchor_coder)
+ self.bbox_coder = TASK_UTILS.build(bbox_coder)
+
+ # build losses
+ self.loss_loc = MODELS.build(loss_loc)
+ self.loss_shape = MODELS.build(loss_shape)
+ self.loss_cls = MODELS.build(loss_cls)
+ self.loss_bbox = MODELS.build(loss_bbox)
+
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+
+ if self.train_cfg:
+ self.assigner = TASK_UTILS.build(self.train_cfg['assigner'])
+ # use PseudoSampler when no sampler in train_cfg
+ if train_cfg.get('sampler', None) is not None:
+ self.sampler = TASK_UTILS.build(
+ self.train_cfg['sampler'], default_args=dict(context=self))
+ else:
+ self.sampler = PseudoSampler()
+
+ self.ga_assigner = TASK_UTILS.build(self.train_cfg['ga_assigner'])
+ if train_cfg.get('ga_sampler', None) is not None:
+ self.ga_sampler = TASK_UTILS.build(
+ self.train_cfg['ga_sampler'],
+ default_args=dict(context=self))
+ else:
+ self.ga_sampler = PseudoSampler()
+
+ self._init_layers()
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self.relu = nn.ReLU(inplace=True)
+ self.conv_loc = nn.Conv2d(self.in_channels, 1, 1)
+ self.conv_shape = nn.Conv2d(self.in_channels, self.num_base_priors * 2,
+ 1)
+ self.feature_adaption = FeatureAdaption(
+ self.in_channels,
+ self.feat_channels,
+ kernel_size=3,
+ deform_groups=self.deform_groups)
+ self.conv_cls = MaskedConv2d(
+ self.feat_channels, self.num_base_priors * self.cls_out_channels,
+ 1)
+ self.conv_reg = MaskedConv2d(self.feat_channels,
+ self.num_base_priors * 4, 1)
+
+ def forward_single(self, x: Tensor) -> Tuple[Tensor]:
+ """Forward feature of a single scale level."""
+ loc_pred = self.conv_loc(x)
+ shape_pred = self.conv_shape(x)
+ x = self.feature_adaption(x, shape_pred)
+ # masked conv is only used during inference for speed-up
+ if not self.training:
+ mask = loc_pred.sigmoid()[0] >= self.loc_filter_thr
+ else:
+ mask = None
+ cls_score = self.conv_cls(x, mask)
+ bbox_pred = self.conv_reg(x, mask)
+ return cls_score, bbox_pred, shape_pred, loc_pred
+
+ def forward(self, x: List[Tensor]) -> Tuple[List[Tensor]]:
+ """Forward features from the upstream network."""
+ return multi_apply(self.forward_single, x)
+
+ def get_sampled_approxs(self,
+ featmap_sizes: List[Tuple[int, int]],
+ batch_img_metas: List[dict],
+ device: str = 'cuda') -> tuple:
+ """Get sampled approxs and inside flags according to feature map sizes.
+
+ Args:
+ featmap_sizes (list[tuple]): Multi-level feature map sizes.
+ batch_img_metas (list[dict]): Image meta info.
+ device (str): device for returned tensors
+
+ Returns:
+ tuple: approxes of each image, inside flags of each image
+ """
+ num_imgs = len(batch_img_metas)
+
+ # since feature map sizes of all images are the same, we only compute
+ # approxes for one time
+ multi_level_approxs = self.approx_anchor_generator.grid_priors(
+ featmap_sizes, device=device)
+ approxs_list = [multi_level_approxs for _ in range(num_imgs)]
+
+ # for each image, we compute inside flags of multi level approxes
+ inside_flag_list = []
+ for img_id, img_meta in enumerate(batch_img_metas):
+ multi_level_flags = []
+ multi_level_approxs = approxs_list[img_id]
+
+ # obtain valid flags for each approx first
+ multi_level_approx_flags = self.approx_anchor_generator \
+ .valid_flags(featmap_sizes,
+ img_meta['pad_shape'],
+ device=device)
+
+ for i, flags in enumerate(multi_level_approx_flags):
+ approxs = multi_level_approxs[i]
+ inside_flags_list = []
+ for j in range(self.approxs_per_octave):
+ split_valid_flags = flags[j::self.approxs_per_octave]
+ split_approxs = approxs[j::self.approxs_per_octave, :]
+ inside_flags = anchor_inside_flags(
+ split_approxs, split_valid_flags,
+ img_meta['img_shape'][:2],
+ self.train_cfg['allowed_border'])
+ inside_flags_list.append(inside_flags)
+ # inside_flag for a position is true if any anchor in this
+ # position is true
+ inside_flags = (
+ torch.stack(inside_flags_list, 0).sum(dim=0) > 0)
+ multi_level_flags.append(inside_flags)
+ inside_flag_list.append(multi_level_flags)
+ return approxs_list, inside_flag_list
+
+ def get_anchors(self,
+ featmap_sizes: List[Tuple[int, int]],
+ shape_preds: List[Tensor],
+ loc_preds: List[Tensor],
+ batch_img_metas: List[dict],
+ use_loc_filter: bool = False,
+ device: str = 'cuda') -> tuple:
+ """Get squares according to feature map sizes and guided anchors.
+
+ Args:
+ featmap_sizes (list[tuple]): Multi-level feature map sizes.
+ shape_preds (list[tensor]): Multi-level shape predictions.
+ loc_preds (list[tensor]): Multi-level location predictions.
+ batch_img_metas (list[dict]): Image meta info.
+ use_loc_filter (bool): Use loc filter or not. Defaults to False
+ device (str): device for returned tensors.
+ Defaults to `cuda`.
+
+ Returns:
+ tuple: square approxs of each image, guided anchors of each image,
+ loc masks of each image.
+ """
+ num_imgs = len(batch_img_metas)
+ num_levels = len(featmap_sizes)
+
+ # since feature map sizes of all images are the same, we only compute
+ # squares for one time
+ multi_level_squares = self.square_anchor_generator.grid_priors(
+ featmap_sizes, device=device)
+ squares_list = [multi_level_squares for _ in range(num_imgs)]
+
+ # for each image, we compute multi level guided anchors
+ guided_anchors_list = []
+ loc_mask_list = []
+ for img_id, img_meta in enumerate(batch_img_metas):
+ multi_level_guided_anchors = []
+ multi_level_loc_mask = []
+ for i in range(num_levels):
+ squares = squares_list[img_id][i]
+ shape_pred = shape_preds[i][img_id]
+ loc_pred = loc_preds[i][img_id]
+ guided_anchors, loc_mask = self._get_guided_anchors_single(
+ squares,
+ shape_pred,
+ loc_pred,
+ use_loc_filter=use_loc_filter)
+ multi_level_guided_anchors.append(guided_anchors)
+ multi_level_loc_mask.append(loc_mask)
+ guided_anchors_list.append(multi_level_guided_anchors)
+ loc_mask_list.append(multi_level_loc_mask)
+ return squares_list, guided_anchors_list, loc_mask_list
+
+ def _get_guided_anchors_single(
+ self,
+ squares: Tensor,
+ shape_pred: Tensor,
+ loc_pred: Tensor,
+ use_loc_filter: bool = False) -> Tuple[Tensor]:
+ """Get guided anchors and loc masks for a single level.
+
+ Args:
+ squares (tensor): Squares of a single level.
+ shape_pred (tensor): Shape predictions of a single level.
+ loc_pred (tensor): Loc predictions of a single level.
+ use_loc_filter (list[tensor]): Use loc filter or not.
+ Defaults to False.
+
+ Returns:
+ tuple: guided anchors, location masks
+ """
+ # calculate location filtering mask
+ loc_pred = loc_pred.sigmoid().detach()
+ if use_loc_filter:
+ loc_mask = loc_pred >= self.loc_filter_thr
+ else:
+ loc_mask = loc_pred >= 0.0
+ mask = loc_mask.permute(1, 2, 0).expand(-1, -1, self.num_base_priors)
+ mask = mask.contiguous().view(-1)
+ # calculate guided anchors
+ squares = squares[mask]
+ anchor_deltas = shape_pred.permute(1, 2, 0).contiguous().view(
+ -1, 2).detach()[mask]
+ bbox_deltas = anchor_deltas.new_full(squares.size(), 0)
+ bbox_deltas[:, 2:] = anchor_deltas
+ guided_anchors = self.anchor_coder.decode(
+ squares, bbox_deltas, wh_ratio_clip=1e-6)
+ return guided_anchors, mask
+
+ def ga_loc_targets(self, batch_gt_instances: InstanceList,
+ featmap_sizes: List[Tuple[int, int]]) -> tuple:
+ """Compute location targets for guided anchoring.
+
+ Each feature map is divided into positive, negative and ignore regions.
+ - positive regions: target 1, weight 1
+ - ignore regions: target 0, weight 0
+ - negative regions: target 0, weight 0.1
+
+ Args:
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ featmap_sizes (list[tuple]): Multi level sizes of each feature
+ maps.
+
+ Returns:
+ tuple: Returns a tuple containing location targets.
+ """
+ anchor_scale = self.approx_anchor_generator.octave_base_scale
+ anchor_strides = self.approx_anchor_generator.strides
+ # Currently only supports same stride in x and y direction.
+ for stride in anchor_strides:
+ assert (stride[0] == stride[1])
+ anchor_strides = [stride[0] for stride in anchor_strides]
+
+ center_ratio = self.train_cfg['center_ratio']
+ ignore_ratio = self.train_cfg['ignore_ratio']
+ img_per_gpu = len(batch_gt_instances)
+ num_lvls = len(featmap_sizes)
+ r1 = (1 - center_ratio) / 2
+ r2 = (1 - ignore_ratio) / 2
+ all_loc_targets = []
+ all_loc_weights = []
+ all_ignore_map = []
+ for lvl_id in range(num_lvls):
+ h, w = featmap_sizes[lvl_id]
+ loc_targets = torch.zeros(
+ img_per_gpu,
+ 1,
+ h,
+ w,
+ device=batch_gt_instances[0].bboxes.device,
+ dtype=torch.float32)
+ loc_weights = torch.full_like(loc_targets, -1)
+ ignore_map = torch.zeros_like(loc_targets)
+ all_loc_targets.append(loc_targets)
+ all_loc_weights.append(loc_weights)
+ all_ignore_map.append(ignore_map)
+ for img_id in range(img_per_gpu):
+ gt_bboxes = batch_gt_instances[img_id].bboxes
+ scale = torch.sqrt((gt_bboxes[:, 2] - gt_bboxes[:, 0]) *
+ (gt_bboxes[:, 3] - gt_bboxes[:, 1]))
+ min_anchor_size = scale.new_full(
+ (1, ), float(anchor_scale * anchor_strides[0]))
+ # assign gt bboxes to different feature levels w.r.t. their scales
+ target_lvls = torch.floor(
+ torch.log2(scale) - torch.log2(min_anchor_size) + 0.5)
+ target_lvls = target_lvls.clamp(min=0, max=num_lvls - 1).long()
+ for gt_id in range(gt_bboxes.size(0)):
+ lvl = target_lvls[gt_id].item()
+ # rescaled to corresponding feature map
+ gt_ = gt_bboxes[gt_id, :4] / anchor_strides[lvl]
+ # calculate ignore regions
+ ignore_x1, ignore_y1, ignore_x2, ignore_y2 = calc_region(
+ gt_, r2, featmap_sizes[lvl])
+ # calculate positive (center) regions
+ ctr_x1, ctr_y1, ctr_x2, ctr_y2 = calc_region(
+ gt_, r1, featmap_sizes[lvl])
+ all_loc_targets[lvl][img_id, 0, ctr_y1:ctr_y2 + 1,
+ ctr_x1:ctr_x2 + 1] = 1
+ all_loc_weights[lvl][img_id, 0, ignore_y1:ignore_y2 + 1,
+ ignore_x1:ignore_x2 + 1] = 0
+ all_loc_weights[lvl][img_id, 0, ctr_y1:ctr_y2 + 1,
+ ctr_x1:ctr_x2 + 1] = 1
+ # calculate ignore map on nearby low level feature
+ if lvl > 0:
+ d_lvl = lvl - 1
+ # rescaled to corresponding feature map
+ gt_ = gt_bboxes[gt_id, :4] / anchor_strides[d_lvl]
+ ignore_x1, ignore_y1, ignore_x2, ignore_y2 = calc_region(
+ gt_, r2, featmap_sizes[d_lvl])
+ all_ignore_map[d_lvl][img_id, 0, ignore_y1:ignore_y2 + 1,
+ ignore_x1:ignore_x2 + 1] = 1
+ # calculate ignore map on nearby high level feature
+ if lvl < num_lvls - 1:
+ u_lvl = lvl + 1
+ # rescaled to corresponding feature map
+ gt_ = gt_bboxes[gt_id, :4] / anchor_strides[u_lvl]
+ ignore_x1, ignore_y1, ignore_x2, ignore_y2 = calc_region(
+ gt_, r2, featmap_sizes[u_lvl])
+ all_ignore_map[u_lvl][img_id, 0, ignore_y1:ignore_y2 + 1,
+ ignore_x1:ignore_x2 + 1] = 1
+ for lvl_id in range(num_lvls):
+ # ignore negative regions w.r.t. ignore map
+ all_loc_weights[lvl_id][(all_loc_weights[lvl_id] < 0)
+ & (all_ignore_map[lvl_id] > 0)] = 0
+ # set negative regions with weight 0.1
+ all_loc_weights[lvl_id][all_loc_weights[lvl_id] < 0] = 0.1
+ # loc average factor to balance loss
+ loc_avg_factor = sum(
+ [t.size(0) * t.size(-1) * t.size(-2)
+ for t in all_loc_targets]) / 200
+ return all_loc_targets, all_loc_weights, loc_avg_factor
+
+ def _ga_shape_target_single(self,
+ flat_approxs: Tensor,
+ inside_flags: Tensor,
+ flat_squares: Tensor,
+ gt_instances: InstanceData,
+ gt_instances_ignore: Optional[InstanceData],
+ img_meta: dict,
+ unmap_outputs: bool = True) -> tuple:
+ """Compute guided anchoring targets.
+
+ This function returns sampled anchors and gt bboxes directly
+ rather than calculates regression targets.
+
+ Args:
+ flat_approxs (Tensor): flat approxs of a single image,
+ shape (n, 4)
+ inside_flags (Tensor): inside flags of a single image,
+ shape (n, ).
+ flat_squares (Tensor): flat squares of a single image,
+ shape (approxs_per_octave * n, 4)
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ img_meta (dict): Meta info of a single image.
+ unmap_outputs (bool): unmap outputs or not.
+
+ Returns:
+ tuple: Returns a tuple containing shape targets of each image.
+ """
+ if not inside_flags.any():
+ raise ValueError(
+ 'There is no valid anchor inside the image boundary. Please '
+ 'check the image size and anchor sizes, or set '
+ '``allowed_border`` to -1 to skip the condition.')
+ # assign gt and sample anchors
+ num_square = flat_squares.size(0)
+ approxs = flat_approxs.view(num_square, self.approxs_per_octave, 4)
+ approxs = approxs[inside_flags, ...]
+ squares = flat_squares[inside_flags, :]
+
+ pred_instances = InstanceData()
+ pred_instances.priors = squares
+ pred_instances.approxs = approxs
+
+ assign_result = self.ga_assigner.assign(
+ pred_instances=pred_instances,
+ gt_instances=gt_instances,
+ gt_instances_ignore=gt_instances_ignore)
+ sampling_result = self.ga_sampler.sample(
+ assign_result=assign_result,
+ pred_instances=pred_instances,
+ gt_instances=gt_instances)
+
+ bbox_anchors = torch.zeros_like(squares)
+ bbox_gts = torch.zeros_like(squares)
+ bbox_weights = torch.zeros_like(squares)
+
+ pos_inds = sampling_result.pos_inds
+ neg_inds = sampling_result.neg_inds
+ if len(pos_inds) > 0:
+ bbox_anchors[pos_inds, :] = sampling_result.pos_bboxes
+ bbox_gts[pos_inds, :] = sampling_result.pos_gt_bboxes
+ bbox_weights[pos_inds, :] = 1.0
+
+ # map up to original set of anchors
+ if unmap_outputs:
+ num_total_anchors = flat_squares.size(0)
+ bbox_anchors = unmap(bbox_anchors, num_total_anchors, inside_flags)
+ bbox_gts = unmap(bbox_gts, num_total_anchors, inside_flags)
+ bbox_weights = unmap(bbox_weights, num_total_anchors, inside_flags)
+
+ return (bbox_anchors, bbox_gts, bbox_weights, pos_inds, neg_inds,
+ sampling_result)
+
+ def ga_shape_targets(self,
+ approx_list: List[List[Tensor]],
+ inside_flag_list: List[List[Tensor]],
+ square_list: List[List[Tensor]],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None,
+ unmap_outputs: bool = True) -> tuple:
+ """Compute guided anchoring targets.
+
+ Args:
+ approx_list (list[list[Tensor]]): Multi level approxs of each
+ image.
+ inside_flag_list (list[list[Tensor]]): Multi level inside flags
+ of each image.
+ square_list (list[list[Tensor]]): Multi level squares of each
+ image.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ unmap_outputs (bool): unmap outputs or not. Defaults to None.
+
+ Returns:
+ tuple: Returns a tuple containing shape targets.
+ """
+ num_imgs = len(batch_img_metas)
+ assert len(approx_list) == len(inside_flag_list) == len(
+ square_list) == num_imgs
+ # anchor number of multi levels
+ num_level_squares = [squares.size(0) for squares in square_list[0]]
+ # concat all level anchors and flags to a single tensor
+ inside_flag_flat_list = []
+ approx_flat_list = []
+ square_flat_list = []
+ for i in range(num_imgs):
+ assert len(square_list[i]) == len(inside_flag_list[i])
+ inside_flag_flat_list.append(torch.cat(inside_flag_list[i]))
+ approx_flat_list.append(torch.cat(approx_list[i]))
+ square_flat_list.append(torch.cat(square_list[i]))
+
+ # compute targets for each image
+ if batch_gt_instances_ignore is None:
+ batch_gt_instances_ignore = [None for _ in range(num_imgs)]
+ (all_bbox_anchors, all_bbox_gts, all_bbox_weights, pos_inds_list,
+ neg_inds_list, sampling_results_list) = multi_apply(
+ self._ga_shape_target_single,
+ approx_flat_list,
+ inside_flag_flat_list,
+ square_flat_list,
+ batch_gt_instances,
+ batch_gt_instances_ignore,
+ batch_img_metas,
+ unmap_outputs=unmap_outputs)
+ # sampled anchors of all images
+ avg_factor = sum(
+ [results.avg_factor for results in sampling_results_list])
+ # split targets to a list w.r.t. multiple levels
+ bbox_anchors_list = images_to_levels(all_bbox_anchors,
+ num_level_squares)
+ bbox_gts_list = images_to_levels(all_bbox_gts, num_level_squares)
+ bbox_weights_list = images_to_levels(all_bbox_weights,
+ num_level_squares)
+ return (bbox_anchors_list, bbox_gts_list, bbox_weights_list,
+ avg_factor)
+
+ def loss_shape_single(self, shape_pred: Tensor, bbox_anchors: Tensor,
+ bbox_gts: Tensor, anchor_weights: Tensor,
+ avg_factor: int) -> Tensor:
+ """Compute shape loss in single level."""
+ shape_pred = shape_pred.permute(0, 2, 3, 1).contiguous().view(-1, 2)
+ bbox_anchors = bbox_anchors.contiguous().view(-1, 4)
+ bbox_gts = bbox_gts.contiguous().view(-1, 4)
+ anchor_weights = anchor_weights.contiguous().view(-1, 4)
+ bbox_deltas = bbox_anchors.new_full(bbox_anchors.size(), 0)
+ bbox_deltas[:, 2:] += shape_pred
+ # filter out negative samples to speed-up weighted_bounded_iou_loss
+ inds = torch.nonzero(
+ anchor_weights[:, 0] > 0, as_tuple=False).squeeze(1)
+ bbox_deltas_ = bbox_deltas[inds]
+ bbox_anchors_ = bbox_anchors[inds]
+ bbox_gts_ = bbox_gts[inds]
+ anchor_weights_ = anchor_weights[inds]
+ pred_anchors_ = self.anchor_coder.decode(
+ bbox_anchors_, bbox_deltas_, wh_ratio_clip=1e-6)
+ loss_shape = self.loss_shape(
+ pred_anchors_, bbox_gts_, anchor_weights_, avg_factor=avg_factor)
+ return loss_shape
+
+ def loss_loc_single(self, loc_pred: Tensor, loc_target: Tensor,
+ loc_weight: Tensor, avg_factor: float) -> Tensor:
+ """Compute location loss in single level."""
+ loss_loc = self.loss_loc(
+ loc_pred.reshape(-1, 1),
+ loc_target.reshape(-1).long(),
+ loc_weight.reshape(-1),
+ avg_factor=avg_factor)
+ return loss_loc
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ shape_preds: List[Tensor],
+ loc_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ has shape (N, num_anchors * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W).
+ shape_preds (list[Tensor]): shape predictions for each scale
+ level with shape (N, 1, H, W).
+ loc_preds (list[Tensor]): location predictions for each scale
+ level with shape (N, num_anchors * 2, H, W).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ assert len(featmap_sizes) == self.approx_anchor_generator.num_levels
+
+ device = cls_scores[0].device
+
+ # get loc targets
+ loc_targets, loc_weights, loc_avg_factor = self.ga_loc_targets(
+ batch_gt_instances, featmap_sizes)
+
+ # get sampled approxes
+ approxs_list, inside_flag_list = self.get_sampled_approxs(
+ featmap_sizes, batch_img_metas, device=device)
+ # get squares and guided anchors
+ squares_list, guided_anchors_list, _ = self.get_anchors(
+ featmap_sizes,
+ shape_preds,
+ loc_preds,
+ batch_img_metas,
+ device=device)
+
+ # get shape targets
+ shape_targets = self.ga_shape_targets(approxs_list, inside_flag_list,
+ squares_list, batch_gt_instances,
+ batch_img_metas)
+ (bbox_anchors_list, bbox_gts_list, anchor_weights_list,
+ ga_avg_factor) = shape_targets
+
+ # get anchor targets
+ cls_reg_targets = self.get_targets(
+ guided_anchors_list,
+ inside_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore)
+ (labels_list, label_weights_list, bbox_targets_list, bbox_weights_list,
+ avg_factor) = cls_reg_targets
+
+ # anchor number of multi levels
+ num_level_anchors = [
+ anchors.size(0) for anchors in guided_anchors_list[0]
+ ]
+ # concat all level anchors to a single tensor
+ concat_anchor_list = []
+ for i in range(len(guided_anchors_list)):
+ concat_anchor_list.append(torch.cat(guided_anchors_list[i]))
+ all_anchor_list = images_to_levels(concat_anchor_list,
+ num_level_anchors)
+
+ # get classification and bbox regression losses
+ losses_cls, losses_bbox = multi_apply(
+ self.loss_by_feat_single,
+ cls_scores,
+ bbox_preds,
+ all_anchor_list,
+ labels_list,
+ label_weights_list,
+ bbox_targets_list,
+ bbox_weights_list,
+ avg_factor=avg_factor)
+
+ # get anchor location loss
+ losses_loc = []
+ for i in range(len(loc_preds)):
+ loss_loc = self.loss_loc_single(
+ loc_preds[i],
+ loc_targets[i],
+ loc_weights[i],
+ avg_factor=loc_avg_factor)
+ losses_loc.append(loss_loc)
+
+ # get anchor shape loss
+ losses_shape = []
+ for i in range(len(shape_preds)):
+ loss_shape = self.loss_shape_single(
+ shape_preds[i],
+ bbox_anchors_list[i],
+ bbox_gts_list[i],
+ anchor_weights_list[i],
+ avg_factor=ga_avg_factor)
+ losses_shape.append(loss_shape)
+
+ return dict(
+ loss_cls=losses_cls,
+ loss_bbox=losses_bbox,
+ loss_shape=losses_shape,
+ loss_loc=losses_loc)
+
+ def predict_by_feat(self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ shape_preds: List[Tensor],
+ loc_preds: List[Tensor],
+ batch_img_metas: List[dict],
+ cfg: OptConfigType = None,
+ rescale: bool = False) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ bbox results.
+
+ Args:
+ cls_scores (list[Tensor]): Classification scores for all
+ scale levels, each is a 4D-tensor, has shape
+ (batch_size, num_priors * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas for all
+ scale levels, each is a 4D-tensor, has shape
+ (batch_size, num_priors * 4, H, W).
+ shape_preds (list[Tensor]): shape predictions for each scale
+ level with shape (N, 1, H, W).
+ loc_preds (list[Tensor]): location predictions for each scale
+ level with shape (N, num_anchors * 2, H, W).
+ batch_img_metas (list[dict], Optional): Batch image meta info.
+ Defaults to None.
+ cfg (ConfigDict, optional): Test / postprocessing
+ configuration, if None, test_cfg would be used.
+ Defaults to None.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Object detection results of each image
+ after the post process. Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4), the last
+ dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ assert len(cls_scores) == len(bbox_preds) == len(shape_preds) == len(
+ loc_preds)
+ num_levels = len(cls_scores)
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ device = cls_scores[0].device
+ # get guided anchors
+ _, guided_anchors, loc_masks = self.get_anchors(
+ featmap_sizes,
+ shape_preds,
+ loc_preds,
+ batch_img_metas,
+ use_loc_filter=not self.training,
+ device=device)
+ result_list = []
+ for img_id in range(len(batch_img_metas)):
+ cls_score_list = [
+ cls_scores[i][img_id].detach() for i in range(num_levels)
+ ]
+ bbox_pred_list = [
+ bbox_preds[i][img_id].detach() for i in range(num_levels)
+ ]
+ guided_anchor_list = [
+ guided_anchors[img_id][i].detach() for i in range(num_levels)
+ ]
+ loc_mask_list = [
+ loc_masks[img_id][i].detach() for i in range(num_levels)
+ ]
+ proposals = self._predict_by_feat_single(
+ cls_scores=cls_score_list,
+ bbox_preds=bbox_pred_list,
+ mlvl_anchors=guided_anchor_list,
+ mlvl_masks=loc_mask_list,
+ img_meta=batch_img_metas[img_id],
+ cfg=cfg,
+ rescale=rescale)
+ result_list.append(proposals)
+ return result_list
+
+ def _predict_by_feat_single(self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ mlvl_anchors: List[Tensor],
+ mlvl_masks: List[Tensor],
+ img_meta: dict,
+ cfg: ConfigType,
+ rescale: bool = False) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores from all scale
+ levels of a single image, each item has shape
+ (num_priors * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas from
+ all scale levels of a single image, each item has shape
+ (num_priors * 4, H, W).
+ mlvl_anchors (list[Tensor]): Each element in the list is
+ the anchors of a single level in feature pyramid. it has
+ shape (num_priors, 4).
+ mlvl_masks (list[Tensor]): Each element in the list is location
+ masks of a single level.
+ img_meta (dict): Image meta info.
+ cfg (:obj:`ConfigDict` or dict): Test / postprocessing
+ configuration, if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4), the last
+ dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ cfg = self.test_cfg if cfg is None else cfg
+ assert len(cls_scores) == len(bbox_preds) == len(mlvl_anchors)
+ mlvl_bbox_preds = []
+ mlvl_valid_anchors = []
+ mlvl_scores = []
+ for cls_score, bbox_pred, anchors, mask in zip(cls_scores, bbox_preds,
+ mlvl_anchors,
+ mlvl_masks):
+ assert cls_score.size()[-2:] == bbox_pred.size()[-2:]
+ # if no location is kept, end.
+ if mask.sum() == 0:
+ continue
+ # reshape scores and bbox_pred
+ cls_score = cls_score.permute(1, 2,
+ 0).reshape(-1, self.cls_out_channels)
+ if self.use_sigmoid_cls:
+ scores = cls_score.sigmoid()
+ else:
+ scores = cls_score.softmax(-1)
+ bbox_pred = bbox_pred.permute(1, 2, 0).reshape(-1, 4)
+ # filter scores, bbox_pred w.r.t. mask.
+ # anchors are filtered in get_anchors() beforehand.
+ scores = scores[mask, :]
+ bbox_pred = bbox_pred[mask, :]
+ if scores.dim() == 0:
+ anchors = anchors.unsqueeze(0)
+ scores = scores.unsqueeze(0)
+ bbox_pred = bbox_pred.unsqueeze(0)
+ # filter anchors, bbox_pred, scores w.r.t. scores
+ nms_pre = cfg.get('nms_pre', -1)
+ if nms_pre > 0 and scores.shape[0] > nms_pre:
+ if self.use_sigmoid_cls:
+ max_scores, _ = scores.max(dim=1)
+ else:
+ # remind that we set FG labels to [0, num_class-1]
+ # since mmdet v2.0
+ # BG cat_id: num_class
+ max_scores, _ = scores[:, :-1].max(dim=1)
+ _, topk_inds = max_scores.topk(nms_pre)
+ anchors = anchors[topk_inds, :]
+ bbox_pred = bbox_pred[topk_inds, :]
+ scores = scores[topk_inds, :]
+
+ mlvl_bbox_preds.append(bbox_pred)
+ mlvl_valid_anchors.append(anchors)
+ mlvl_scores.append(scores)
+
+ mlvl_bbox_preds = torch.cat(mlvl_bbox_preds)
+ mlvl_anchors = torch.cat(mlvl_valid_anchors)
+ mlvl_scores = torch.cat(mlvl_scores)
+ mlvl_bboxes = self.bbox_coder.decode(
+ mlvl_anchors, mlvl_bbox_preds, max_shape=img_meta['img_shape'])
+
+ if rescale:
+ assert img_meta.get('scale_factor') is not None
+ mlvl_bboxes /= mlvl_bboxes.new_tensor(
+ img_meta['scale_factor']).repeat((1, 2))
+
+ if self.use_sigmoid_cls:
+ # Add a dummy background class to the backend when using sigmoid
+ # remind that we set FG labels to [0, num_class-1] since mmdet v2.0
+ # BG cat_id: num_class
+ padding = mlvl_scores.new_zeros(mlvl_scores.shape[0], 1)
+ mlvl_scores = torch.cat([mlvl_scores, padding], dim=1)
+ # multi class NMS
+ det_bboxes, det_labels = multiclass_nms(mlvl_bboxes, mlvl_scores,
+ cfg.score_thr, cfg.nms,
+ cfg.max_per_img)
+
+ results = InstanceData()
+ results.bboxes = det_bboxes[:, :-1]
+ results.scores = det_bboxes[:, -1]
+ results.labels = det_labels
+ return results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/lad_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/lad_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..d1218e1f88206704d4f414d151ccd34a189ac5d0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/lad_head.py
@@ -0,0 +1,226 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional
+
+import torch
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import bbox_overlaps
+from mmdet.utils import InstanceList, OptInstanceList
+from ..utils import levels_to_images, multi_apply, unpack_gt_instances
+from .paa_head import PAAHead
+
+
+@MODELS.register_module()
+class LADHead(PAAHead):
+ """Label Assignment Head from the paper: `Improving Object Detection by
+ Label Assignment Distillation `_"""
+
+ def get_label_assignment(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ iou_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> tuple:
+ """Get label assignment (from teacher).
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ Has shape (N, num_anchors * num_classes, H, W)
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W)
+ iou_preds (list[Tensor]): iou_preds for each scale
+ level with shape (N, num_anchors * 1, H, W)
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ tuple: Returns a tuple containing label assignment variables.
+
+ - labels (Tensor): Labels of all anchors, each with
+ shape (num_anchors,).
+ - labels_weight (Tensor): Label weights of all anchor.
+ each with shape (num_anchors,).
+ - bboxes_target (Tensor): BBox targets of all anchors.
+ each with shape (num_anchors, 4).
+ - bboxes_weight (Tensor): BBox weights of all anchors.
+ each with shape (num_anchors, 4).
+ - pos_inds_flatten (Tensor): Contains all index of positive
+ sample in all anchor.
+ - pos_anchors (Tensor): Positive anchors.
+ - num_pos (int): Number of positive anchors.
+ """
+
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ assert len(featmap_sizes) == self.prior_generator.num_levels
+
+ device = cls_scores[0].device
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+ cls_reg_targets = self.get_targets(
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore,
+ )
+ (labels, labels_weight, bboxes_target, bboxes_weight, pos_inds,
+ pos_gt_index) = cls_reg_targets
+ cls_scores = levels_to_images(cls_scores)
+ cls_scores = [
+ item.reshape(-1, self.cls_out_channels) for item in cls_scores
+ ]
+ bbox_preds = levels_to_images(bbox_preds)
+ bbox_preds = [item.reshape(-1, 4) for item in bbox_preds]
+ pos_losses_list, = multi_apply(self.get_pos_loss, anchor_list,
+ cls_scores, bbox_preds, labels,
+ labels_weight, bboxes_target,
+ bboxes_weight, pos_inds)
+
+ with torch.no_grad():
+ reassign_labels, reassign_label_weight, \
+ reassign_bbox_weights, num_pos = multi_apply(
+ self.paa_reassign,
+ pos_losses_list,
+ labels,
+ labels_weight,
+ bboxes_weight,
+ pos_inds,
+ pos_gt_index,
+ anchor_list)
+ num_pos = sum(num_pos)
+ # convert all tensor list to a flatten tensor
+ labels = torch.cat(reassign_labels, 0).view(-1)
+ flatten_anchors = torch.cat(
+ [torch.cat(item, 0) for item in anchor_list])
+ labels_weight = torch.cat(reassign_label_weight, 0).view(-1)
+ bboxes_target = torch.cat(bboxes_target,
+ 0).view(-1, bboxes_target[0].size(-1))
+
+ pos_inds_flatten = ((labels >= 0)
+ &
+ (labels < self.num_classes)).nonzero().reshape(-1)
+
+ if num_pos:
+ pos_anchors = flatten_anchors[pos_inds_flatten]
+ else:
+ pos_anchors = None
+
+ label_assignment_results = (labels, labels_weight, bboxes_target,
+ bboxes_weight, pos_inds_flatten,
+ pos_anchors, num_pos)
+ return label_assignment_results
+
+ def loss(self, x: List[Tensor], label_assignment_results: tuple,
+ batch_data_samples: SampleList) -> dict:
+ """Forward train with the available label assignment (student receives
+ from teacher).
+
+ Args:
+ x (list[Tensor]): Features from FPN.
+ label_assignment_results (tuple): As the outputs defined in the
+ function `self.get_label_assignment`.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ losses: (dict[str, Tensor]): A dictionary of loss components.
+ """
+ outputs = unpack_gt_instances(batch_data_samples)
+ batch_gt_instances, batch_gt_instances_ignore, batch_img_metas \
+ = outputs
+
+ outs = self(x)
+ loss_inputs = outs + (batch_gt_instances, batch_img_metas)
+ losses = self.loss_by_feat(
+ *loss_inputs,
+ batch_gt_instances_ignore=batch_gt_instances_ignore,
+ label_assignment_results=label_assignment_results)
+ return losses
+
+ def loss_by_feat(self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ iou_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None,
+ label_assignment_results: Optional[tuple] = None) -> dict:
+ """Compute losses of the head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ Has shape (N, num_anchors * num_classes, H, W)
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W)
+ iou_preds (list[Tensor]): iou_preds for each scale
+ level with shape (N, num_anchors * 1, H, W)
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ label_assignment_results (tuple, optional): As the outputs defined
+ in the function `self.get_
+ label_assignment`.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss gmm_assignment.
+ """
+
+ (labels, labels_weight, bboxes_target, bboxes_weight, pos_inds_flatten,
+ pos_anchors, num_pos) = label_assignment_results
+
+ cls_scores = levels_to_images(cls_scores)
+ cls_scores = [
+ item.reshape(-1, self.cls_out_channels) for item in cls_scores
+ ]
+ bbox_preds = levels_to_images(bbox_preds)
+ bbox_preds = [item.reshape(-1, 4) for item in bbox_preds]
+ iou_preds = levels_to_images(iou_preds)
+ iou_preds = [item.reshape(-1, 1) for item in iou_preds]
+
+ # convert all tensor list to a flatten tensor
+ cls_scores = torch.cat(cls_scores, 0).view(-1, cls_scores[0].size(-1))
+ bbox_preds = torch.cat(bbox_preds, 0).view(-1, bbox_preds[0].size(-1))
+ iou_preds = torch.cat(iou_preds, 0).view(-1, iou_preds[0].size(-1))
+
+ losses_cls = self.loss_cls(
+ cls_scores,
+ labels,
+ labels_weight,
+ avg_factor=max(num_pos, len(batch_img_metas))) # avoid num_pos=0
+ if num_pos:
+ pos_bbox_pred = self.bbox_coder.decode(
+ pos_anchors, bbox_preds[pos_inds_flatten])
+ pos_bbox_target = bboxes_target[pos_inds_flatten]
+ iou_target = bbox_overlaps(
+ pos_bbox_pred.detach(), pos_bbox_target, is_aligned=True)
+ losses_iou = self.loss_centerness(
+ iou_preds[pos_inds_flatten],
+ iou_target.unsqueeze(-1),
+ avg_factor=num_pos)
+ losses_bbox = self.loss_bbox(
+ pos_bbox_pred, pos_bbox_target, avg_factor=num_pos)
+
+ else:
+ losses_iou = iou_preds.sum() * 0
+ losses_bbox = bbox_preds.sum() * 0
+
+ return dict(
+ loss_cls=losses_cls, loss_bbox=losses_bbox, loss_iou=losses_iou)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/ld_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/ld_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..2558fac97ee26ff89c5fa1b386f5ce68c3ad384d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/ld_head.py
@@ -0,0 +1,257 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple
+
+import torch
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import bbox_overlaps
+from mmdet.utils import ConfigType, InstanceList, OptInstanceList, reduce_mean
+from ..utils import multi_apply, unpack_gt_instances
+from .gfl_head import GFLHead
+
+
+@MODELS.register_module()
+class LDHead(GFLHead):
+ """Localization distillation Head. (Short description)
+
+ It utilizes the learned bbox distributions to transfer the localization
+ dark knowledge from teacher to student. Original paper: `Localization
+ Distillation for Object Detection. `_
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ loss_ld (:obj:`ConfigDict` or dict): Config of Localization
+ Distillation Loss (LD), T is the temperature for distillation.
+ """
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: int,
+ loss_ld: ConfigType = dict(
+ type='LocalizationDistillationLoss',
+ loss_weight=0.25,
+ T=10),
+ **kwargs) -> dict:
+
+ super().__init__(
+ num_classes=num_classes, in_channels=in_channels, **kwargs)
+ self.loss_ld = MODELS.build(loss_ld)
+
+ def loss_by_feat_single(self, anchors: Tensor, cls_score: Tensor,
+ bbox_pred: Tensor, labels: Tensor,
+ label_weights: Tensor, bbox_targets: Tensor,
+ stride: Tuple[int], soft_targets: Tensor,
+ avg_factor: int):
+ """Calculate the loss of a single scale level based on the features
+ extracted by the detection head.
+
+ Args:
+ anchors (Tensor): Box reference for each scale level with shape
+ (N, num_total_anchors, 4).
+ cls_score (Tensor): Cls and quality joint scores for each scale
+ level has shape (N, num_classes, H, W).
+ bbox_pred (Tensor): Box distribution logits for each scale
+ level with shape (N, 4*(n+1), H, W), n is max value of integral
+ set.
+ labels (Tensor): Labels of each anchors with shape
+ (N, num_total_anchors).
+ label_weights (Tensor): Label weights of each anchor with shape
+ (N, num_total_anchors)
+ bbox_targets (Tensor): BBox regression targets of each anchor with
+ shape (N, num_total_anchors, 4).
+ stride (tuple): Stride in this scale level.
+ soft_targets (Tensor): Soft BBox regression targets.
+ avg_factor (int): Average factor that is used to average
+ the loss. When using sampling method, avg_factor is usually
+ the sum of positive and negative priors. When using
+ `PseudoSampler`, `avg_factor` is usually equal to the number
+ of positive priors.
+
+ Returns:
+ dict[tuple, Tensor]: Loss components and weight targets.
+ """
+ assert stride[0] == stride[1], 'h stride is not equal to w stride!'
+ anchors = anchors.reshape(-1, 4)
+ cls_score = cls_score.permute(0, 2, 3,
+ 1).reshape(-1, self.cls_out_channels)
+ bbox_pred = bbox_pred.permute(0, 2, 3,
+ 1).reshape(-1, 4 * (self.reg_max + 1))
+ soft_targets = soft_targets.permute(0, 2, 3,
+ 1).reshape(-1,
+ 4 * (self.reg_max + 1))
+
+ bbox_targets = bbox_targets.reshape(-1, 4)
+ labels = labels.reshape(-1)
+ label_weights = label_weights.reshape(-1)
+
+ # FG cat_id: [0, num_classes -1], BG cat_id: num_classes
+ bg_class_ind = self.num_classes
+ pos_inds = ((labels >= 0)
+ & (labels < bg_class_ind)).nonzero().squeeze(1)
+ score = label_weights.new_zeros(labels.shape)
+
+ if len(pos_inds) > 0:
+ pos_bbox_targets = bbox_targets[pos_inds]
+ pos_bbox_pred = bbox_pred[pos_inds]
+ pos_anchors = anchors[pos_inds]
+ pos_anchor_centers = self.anchor_center(pos_anchors) / stride[0]
+
+ weight_targets = cls_score.detach().sigmoid()
+ weight_targets = weight_targets.max(dim=1)[0][pos_inds]
+ pos_bbox_pred_corners = self.integral(pos_bbox_pred)
+ pos_decode_bbox_pred = self.bbox_coder.decode(
+ pos_anchor_centers, pos_bbox_pred_corners)
+ pos_decode_bbox_targets = pos_bbox_targets / stride[0]
+ score[pos_inds] = bbox_overlaps(
+ pos_decode_bbox_pred.detach(),
+ pos_decode_bbox_targets,
+ is_aligned=True)
+ pred_corners = pos_bbox_pred.reshape(-1, self.reg_max + 1)
+ pos_soft_targets = soft_targets[pos_inds]
+ soft_corners = pos_soft_targets.reshape(-1, self.reg_max + 1)
+
+ target_corners = self.bbox_coder.encode(pos_anchor_centers,
+ pos_decode_bbox_targets,
+ self.reg_max).reshape(-1)
+
+ # regression loss
+ loss_bbox = self.loss_bbox(
+ pos_decode_bbox_pred,
+ pos_decode_bbox_targets,
+ weight=weight_targets,
+ avg_factor=1.0)
+
+ # dfl loss
+ loss_dfl = self.loss_dfl(
+ pred_corners,
+ target_corners,
+ weight=weight_targets[:, None].expand(-1, 4).reshape(-1),
+ avg_factor=4.0)
+
+ # ld loss
+ loss_ld = self.loss_ld(
+ pred_corners,
+ soft_corners,
+ weight=weight_targets[:, None].expand(-1, 4).reshape(-1),
+ avg_factor=4.0)
+
+ else:
+ loss_ld = bbox_pred.sum() * 0
+ loss_bbox = bbox_pred.sum() * 0
+ loss_dfl = bbox_pred.sum() * 0
+ weight_targets = bbox_pred.new_tensor(0)
+
+ # cls (qfl) loss
+ loss_cls = self.loss_cls(
+ cls_score, (labels, score),
+ weight=label_weights,
+ avg_factor=avg_factor)
+
+ return loss_cls, loss_bbox, loss_dfl, loss_ld, weight_targets.sum()
+
+ def loss(self, x: List[Tensor], out_teacher: Tuple[Tensor],
+ batch_data_samples: SampleList) -> dict:
+ """
+ Args:
+ x (list[Tensor]): Features from FPN.
+ out_teacher (tuple[Tensor]): The output of teacher.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ tuple[dict, list]: The loss components and proposals of each image.
+
+ - losses (dict[str, Tensor]): A dictionary of loss components.
+ - proposal_list (list[Tensor]): Proposals of each image.
+ """
+ outputs = unpack_gt_instances(batch_data_samples)
+ batch_gt_instances, batch_gt_instances_ignore, batch_img_metas \
+ = outputs
+
+ outs = self(x)
+ soft_targets = out_teacher[1]
+ loss_inputs = outs + (batch_gt_instances, batch_img_metas,
+ soft_targets)
+ losses = self.loss_by_feat(
+ *loss_inputs, batch_gt_instances_ignore=batch_gt_instances_ignore)
+
+ return losses
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ soft_targets: List[Tensor],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Compute losses of the head.
+
+ Args:
+ cls_scores (list[Tensor]): Cls and quality scores for each scale
+ level has shape (N, num_classes, H, W).
+ bbox_preds (list[Tensor]): Box distribution logits for each scale
+ level with shape (N, 4*(n+1), H, W), n is max value of integral
+ set.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ soft_targets (list[Tensor]): Soft BBox regression targets.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ assert len(featmap_sizes) == self.prior_generator.num_levels
+
+ device = cls_scores[0].device
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+
+ cls_reg_targets = self.get_targets(
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore)
+
+ (anchor_list, labels_list, label_weights_list, bbox_targets_list,
+ bbox_weights_list, avg_factor) = cls_reg_targets
+
+ avg_factor = reduce_mean(
+ torch.tensor(avg_factor, dtype=torch.float, device=device)).item()
+
+ losses_cls, losses_bbox, losses_dfl, losses_ld, \
+ avg_factor = multi_apply(
+ self.loss_by_feat_single,
+ anchor_list,
+ cls_scores,
+ bbox_preds,
+ labels_list,
+ label_weights_list,
+ bbox_targets_list,
+ self.prior_generator.strides,
+ soft_targets,
+ avg_factor=avg_factor)
+
+ avg_factor = sum(avg_factor) + 1e-6
+ avg_factor = reduce_mean(avg_factor).item()
+ losses_bbox = [x / avg_factor for x in losses_bbox]
+ losses_dfl = [x / avg_factor for x in losses_dfl]
+ return dict(
+ loss_cls=losses_cls,
+ loss_bbox=losses_bbox,
+ loss_dfl=losses_dfl,
+ loss_ld=losses_ld)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/mask2former_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/mask2former_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..12d47c655255f92819646b8ea304b9736ec30660
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/mask2former_head.py
@@ -0,0 +1,459 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from typing import List, Tuple
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import Conv2d
+from mmcv.ops import point_sample
+from mmengine.model import ModuleList, caffe2_xavier_init
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures import SampleList
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig, reduce_mean
+from ..layers import Mask2FormerTransformerDecoder, SinePositionalEncoding
+from ..utils import get_uncertain_point_coords_with_randomness
+from .anchor_free_head import AnchorFreeHead
+from .maskformer_head import MaskFormerHead
+
+
+@MODELS.register_module()
+class Mask2FormerHead(MaskFormerHead):
+ """Implements the Mask2Former head.
+
+ See `Masked-attention Mask Transformer for Universal Image
+ Segmentation `_ for details.
+
+ Args:
+ in_channels (list[int]): Number of channels in the input feature map.
+ feat_channels (int): Number of channels for features.
+ out_channels (int): Number of channels for output.
+ num_things_classes (int): Number of things.
+ num_stuff_classes (int): Number of stuff.
+ num_queries (int): Number of query in Transformer decoder.
+ pixel_decoder (:obj:`ConfigDict` or dict): Config for pixel
+ decoder. Defaults to None.
+ enforce_decoder_input_project (bool, optional): Whether to add
+ a layer to change the embed_dim of tranformer encoder in
+ pixel decoder to the embed_dim of transformer decoder.
+ Defaults to False.
+ transformer_decoder (:obj:`ConfigDict` or dict): Config for
+ transformer decoder. Defaults to None.
+ positional_encoding (:obj:`ConfigDict` or dict): Config for
+ transformer decoder position encoding. Defaults to
+ dict(num_feats=128, normalize=True).
+ loss_cls (:obj:`ConfigDict` or dict): Config of the classification
+ loss. Defaults to None.
+ loss_mask (:obj:`ConfigDict` or dict): Config of the mask loss.
+ Defaults to None.
+ loss_dice (:obj:`ConfigDict` or dict): Config of the dice loss.
+ Defaults to None.
+ train_cfg (:obj:`ConfigDict` or dict, optional): Training config of
+ Mask2Former head.
+ test_cfg (:obj:`ConfigDict` or dict, optional): Testing config of
+ Mask2Former head.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict], optional): Initialization config dict. Defaults to None.
+ """
+
+ def __init__(self,
+ in_channels: List[int],
+ feat_channels: int,
+ out_channels: int,
+ num_things_classes: int = 80,
+ num_stuff_classes: int = 53,
+ num_queries: int = 100,
+ num_transformer_feat_level: int = 3,
+ pixel_decoder: ConfigType = ...,
+ enforce_decoder_input_project: bool = False,
+ transformer_decoder: ConfigType = ...,
+ positional_encoding: ConfigType = dict(
+ num_feats=128, normalize=True),
+ loss_cls: ConfigType = dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=2.0,
+ reduction='mean',
+ class_weight=[1.0] * 133 + [0.1]),
+ loss_mask: ConfigType = dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ reduction='mean',
+ loss_weight=5.0),
+ loss_dice: ConfigType = dict(
+ type='DiceLoss',
+ use_sigmoid=True,
+ activate=True,
+ reduction='mean',
+ naive_dice=True,
+ eps=1.0,
+ loss_weight=5.0),
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ init_cfg: OptMultiConfig = None,
+ **kwargs) -> None:
+ super(AnchorFreeHead, self).__init__(init_cfg=init_cfg)
+ self.num_things_classes = num_things_classes
+ self.num_stuff_classes = num_stuff_classes
+ self.num_classes = self.num_things_classes + self.num_stuff_classes
+ self.num_queries = num_queries
+ self.num_transformer_feat_level = num_transformer_feat_level
+ self.num_heads = transformer_decoder.layer_cfg.cross_attn_cfg.num_heads
+ self.num_transformer_decoder_layers = transformer_decoder.num_layers
+ assert pixel_decoder.encoder.layer_cfg. \
+ self_attn_cfg.num_levels == num_transformer_feat_level
+ pixel_decoder_ = copy.deepcopy(pixel_decoder)
+ pixel_decoder_.update(
+ in_channels=in_channels,
+ feat_channels=feat_channels,
+ out_channels=out_channels)
+ self.pixel_decoder = MODELS.build(pixel_decoder_)
+ self.transformer_decoder = Mask2FormerTransformerDecoder(
+ **transformer_decoder)
+ self.decoder_embed_dims = self.transformer_decoder.embed_dims
+
+ self.decoder_input_projs = ModuleList()
+ # from low resolution to high resolution
+ for _ in range(num_transformer_feat_level):
+ if (self.decoder_embed_dims != feat_channels
+ or enforce_decoder_input_project):
+ self.decoder_input_projs.append(
+ Conv2d(
+ feat_channels, self.decoder_embed_dims, kernel_size=1))
+ else:
+ self.decoder_input_projs.append(nn.Identity())
+ self.decoder_positional_encoding = SinePositionalEncoding(
+ **positional_encoding)
+ self.query_embed = nn.Embedding(self.num_queries, feat_channels)
+ self.query_feat = nn.Embedding(self.num_queries, feat_channels)
+ # from low resolution to high resolution
+ self.level_embed = nn.Embedding(self.num_transformer_feat_level,
+ feat_channels)
+
+ self.cls_embed = nn.Linear(feat_channels, self.num_classes + 1)
+ self.mask_embed = nn.Sequential(
+ nn.Linear(feat_channels, feat_channels), nn.ReLU(inplace=True),
+ nn.Linear(feat_channels, feat_channels), nn.ReLU(inplace=True),
+ nn.Linear(feat_channels, out_channels))
+
+ self.test_cfg = test_cfg
+ self.train_cfg = train_cfg
+ if train_cfg:
+ self.assigner = TASK_UTILS.build(self.train_cfg['assigner'])
+ self.sampler = TASK_UTILS.build(
+ self.train_cfg['sampler'], default_args=dict(context=self))
+ self.num_points = self.train_cfg.get('num_points', 12544)
+ self.oversample_ratio = self.train_cfg.get('oversample_ratio', 3.0)
+ self.importance_sample_ratio = self.train_cfg.get(
+ 'importance_sample_ratio', 0.75)
+
+ self.class_weight = loss_cls.class_weight
+ self.loss_cls = MODELS.build(loss_cls)
+ self.loss_mask = MODELS.build(loss_mask)
+ self.loss_dice = MODELS.build(loss_dice)
+
+ def init_weights(self) -> None:
+ for m in self.decoder_input_projs:
+ if isinstance(m, Conv2d):
+ caffe2_xavier_init(m, bias=0)
+
+ self.pixel_decoder.init_weights()
+
+ for p in self.transformer_decoder.parameters():
+ if p.dim() > 1:
+ nn.init.xavier_normal_(p)
+
+ def _get_targets_single(self, cls_score: Tensor, mask_pred: Tensor,
+ gt_instances: InstanceData,
+ img_meta: dict) -> Tuple[Tensor]:
+ """Compute classification and mask targets for one image.
+
+ Args:
+ cls_score (Tensor): Mask score logits from a single decoder layer
+ for one image. Shape (num_queries, cls_out_channels).
+ mask_pred (Tensor): Mask logits for a single decoder layer for one
+ image. Shape (num_queries, h, w).
+ gt_instances (:obj:`InstanceData`): It contains ``labels`` and
+ ``masks``.
+ img_meta (dict): Image informtation.
+
+ Returns:
+ tuple[Tensor]: A tuple containing the following for one image.
+
+ - labels (Tensor): Labels of each image. \
+ shape (num_queries, ).
+ - label_weights (Tensor): Label weights of each image. \
+ shape (num_queries, ).
+ - mask_targets (Tensor): Mask targets of each image. \
+ shape (num_queries, h, w).
+ - mask_weights (Tensor): Mask weights of each image. \
+ shape (num_queries, ).
+ - pos_inds (Tensor): Sampled positive indices for each \
+ image.
+ - neg_inds (Tensor): Sampled negative indices for each \
+ image.
+ - sampling_result (:obj:`SamplingResult`): Sampling results.
+ """
+ gt_labels = gt_instances.labels
+ gt_masks = gt_instances.masks
+ # sample points
+ num_queries = cls_score.shape[0]
+ num_gts = gt_labels.shape[0]
+
+ point_coords = torch.rand((1, self.num_points, 2),
+ device=cls_score.device)
+ # shape (num_queries, num_points)
+ mask_points_pred = point_sample(
+ mask_pred.unsqueeze(1), point_coords.repeat(num_queries, 1,
+ 1)).squeeze(1)
+ # shape (num_gts, num_points)
+ gt_points_masks = point_sample(
+ gt_masks.unsqueeze(1).float(), point_coords.repeat(num_gts, 1,
+ 1)).squeeze(1)
+
+ sampled_gt_instances = InstanceData(
+ labels=gt_labels, masks=gt_points_masks)
+ sampled_pred_instances = InstanceData(
+ scores=cls_score, masks=mask_points_pred)
+ # assign and sample
+ assign_result = self.assigner.assign(
+ pred_instances=sampled_pred_instances,
+ gt_instances=sampled_gt_instances,
+ img_meta=img_meta)
+ pred_instances = InstanceData(scores=cls_score, masks=mask_pred)
+ sampling_result = self.sampler.sample(
+ assign_result=assign_result,
+ pred_instances=pred_instances,
+ gt_instances=gt_instances)
+ pos_inds = sampling_result.pos_inds
+ neg_inds = sampling_result.neg_inds
+
+ # label target
+ labels = gt_labels.new_full((self.num_queries, ),
+ self.num_classes,
+ dtype=torch.long)
+ labels[pos_inds] = gt_labels[sampling_result.pos_assigned_gt_inds]
+ label_weights = gt_labels.new_ones((self.num_queries, ))
+
+ # mask target
+ mask_targets = gt_masks[sampling_result.pos_assigned_gt_inds]
+ mask_weights = mask_pred.new_zeros((self.num_queries, ))
+ mask_weights[pos_inds] = 1.0
+
+ return (labels, label_weights, mask_targets, mask_weights, pos_inds,
+ neg_inds, sampling_result)
+
+ def _loss_by_feat_single(self, cls_scores: Tensor, mask_preds: Tensor,
+ batch_gt_instances: List[InstanceData],
+ batch_img_metas: List[dict]) -> Tuple[Tensor]:
+ """Loss function for outputs from a single decoder layer.
+
+ Args:
+ cls_scores (Tensor): Mask score logits from a single decoder layer
+ for all images. Shape (batch_size, num_queries,
+ cls_out_channels). Note `cls_out_channels` should includes
+ background.
+ mask_preds (Tensor): Mask logits for a pixel decoder for all
+ images. Shape (batch_size, num_queries, h, w).
+ batch_gt_instances (list[obj:`InstanceData`]): each contains
+ ``labels`` and ``masks``.
+ batch_img_metas (list[dict]): List of image meta information.
+
+ Returns:
+ tuple[Tensor]: Loss components for outputs from a single \
+ decoder layer.
+ """
+ num_imgs = cls_scores.size(0)
+ cls_scores_list = [cls_scores[i] for i in range(num_imgs)]
+ mask_preds_list = [mask_preds[i] for i in range(num_imgs)]
+ (labels_list, label_weights_list, mask_targets_list, mask_weights_list,
+ avg_factor) = self.get_targets(cls_scores_list, mask_preds_list,
+ batch_gt_instances, batch_img_metas)
+ # shape (batch_size, num_queries)
+ labels = torch.stack(labels_list, dim=0)
+ # shape (batch_size, num_queries)
+ label_weights = torch.stack(label_weights_list, dim=0)
+ # shape (num_total_gts, h, w)
+ mask_targets = torch.cat(mask_targets_list, dim=0)
+ # shape (batch_size, num_queries)
+ mask_weights = torch.stack(mask_weights_list, dim=0)
+
+ # classfication loss
+ # shape (batch_size * num_queries, )
+ cls_scores = cls_scores.flatten(0, 1)
+ labels = labels.flatten(0, 1)
+ label_weights = label_weights.flatten(0, 1)
+
+ class_weight = cls_scores.new_tensor(self.class_weight)
+ loss_cls = self.loss_cls(
+ cls_scores,
+ labels,
+ label_weights,
+ avg_factor=class_weight[labels].sum())
+
+ num_total_masks = reduce_mean(cls_scores.new_tensor([avg_factor]))
+ num_total_masks = max(num_total_masks, 1)
+
+ # extract positive ones
+ # shape (batch_size, num_queries, h, w) -> (num_total_gts, h, w)
+ mask_preds = mask_preds[mask_weights > 0]
+
+ if mask_targets.shape[0] == 0:
+ # zero match
+ loss_dice = mask_preds.sum()
+ loss_mask = mask_preds.sum()
+ return loss_cls, loss_mask, loss_dice
+
+ with torch.no_grad():
+ points_coords = get_uncertain_point_coords_with_randomness(
+ mask_preds.unsqueeze(1), None, self.num_points,
+ self.oversample_ratio, self.importance_sample_ratio)
+ # shape (num_total_gts, h, w) -> (num_total_gts, num_points)
+ mask_point_targets = point_sample(
+ mask_targets.unsqueeze(1).float(), points_coords).squeeze(1)
+ # shape (num_queries, h, w) -> (num_queries, num_points)
+ mask_point_preds = point_sample(
+ mask_preds.unsqueeze(1), points_coords).squeeze(1)
+
+ # dice loss
+ loss_dice = self.loss_dice(
+ mask_point_preds, mask_point_targets, avg_factor=num_total_masks)
+
+ # mask loss
+ # shape (num_queries, num_points) -> (num_queries * num_points, )
+ mask_point_preds = mask_point_preds.reshape(-1)
+ # shape (num_total_gts, num_points) -> (num_total_gts * num_points, )
+ mask_point_targets = mask_point_targets.reshape(-1)
+ loss_mask = self.loss_mask(
+ mask_point_preds,
+ mask_point_targets,
+ avg_factor=num_total_masks * self.num_points)
+
+ return loss_cls, loss_mask, loss_dice
+
+ def _forward_head(self, decoder_out: Tensor, mask_feature: Tensor,
+ attn_mask_target_size: Tuple[int, int]) -> Tuple[Tensor]:
+ """Forward for head part which is called after every decoder layer.
+
+ Args:
+ decoder_out (Tensor): in shape (batch_size, num_queries, c).
+ mask_feature (Tensor): in shape (batch_size, c, h, w).
+ attn_mask_target_size (tuple[int, int]): target attention
+ mask size.
+
+ Returns:
+ tuple: A tuple contain three elements.
+
+ - cls_pred (Tensor): Classification scores in shape \
+ (batch_size, num_queries, cls_out_channels). \
+ Note `cls_out_channels` should includes background.
+ - mask_pred (Tensor): Mask scores in shape \
+ (batch_size, num_queries,h, w).
+ - attn_mask (Tensor): Attention mask in shape \
+ (batch_size * num_heads, num_queries, h, w).
+ """
+ decoder_out = self.transformer_decoder.post_norm(decoder_out)
+ # shape (num_queries, batch_size, c)
+ cls_pred = self.cls_embed(decoder_out)
+ # shape (num_queries, batch_size, c)
+ mask_embed = self.mask_embed(decoder_out)
+ # shape (num_queries, batch_size, h, w)
+ mask_pred = torch.einsum('bqc,bchw->bqhw', mask_embed, mask_feature)
+ attn_mask = F.interpolate(
+ mask_pred,
+ attn_mask_target_size,
+ mode='bilinear',
+ align_corners=False)
+ # shape (num_queries, batch_size, h, w) ->
+ # (batch_size * num_head, num_queries, h, w)
+ attn_mask = attn_mask.flatten(2).unsqueeze(1).repeat(
+ (1, self.num_heads, 1, 1)).flatten(0, 1)
+ attn_mask = attn_mask.sigmoid() < 0.5
+ attn_mask = attn_mask.detach()
+
+ return cls_pred, mask_pred, attn_mask
+
+ def forward(self, x: List[Tensor],
+ batch_data_samples: SampleList) -> Tuple[List[Tensor]]:
+ """Forward function.
+
+ Args:
+ x (list[Tensor]): Multi scale Features from the
+ upstream network, each is a 4D-tensor.
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+
+ Returns:
+ tuple[list[Tensor]]: A tuple contains two elements.
+
+ - cls_pred_list (list[Tensor)]: Classification logits \
+ for each decoder layer. Each is a 3D-tensor with shape \
+ (batch_size, num_queries, cls_out_channels). \
+ Note `cls_out_channels` should includes background.
+ - mask_pred_list (list[Tensor]): Mask logits for each \
+ decoder layer. Each with shape (batch_size, num_queries, \
+ h, w).
+ """
+ batch_size = x[0].shape[0]
+ mask_features, multi_scale_memorys = self.pixel_decoder(x)
+ # multi_scale_memorys (from low resolution to high resolution)
+ decoder_inputs = []
+ decoder_positional_encodings = []
+ for i in range(self.num_transformer_feat_level):
+ decoder_input = self.decoder_input_projs[i](multi_scale_memorys[i])
+ # shape (batch_size, c, h, w) -> (batch_size, h*w, c)
+ decoder_input = decoder_input.flatten(2).permute(0, 2, 1)
+ level_embed = self.level_embed.weight[i].view(1, 1, -1)
+ decoder_input = decoder_input + level_embed
+ # shape (batch_size, c, h, w) -> (batch_size, h*w, c)
+ mask = decoder_input.new_zeros(
+ (batch_size, ) + multi_scale_memorys[i].shape[-2:],
+ dtype=torch.bool)
+ decoder_positional_encoding = self.decoder_positional_encoding(
+ mask)
+ decoder_positional_encoding = decoder_positional_encoding.flatten(
+ 2).permute(0, 2, 1)
+ decoder_inputs.append(decoder_input)
+ decoder_positional_encodings.append(decoder_positional_encoding)
+ # shape (num_queries, c) -> (batch_size, num_queries, c)
+ query_feat = self.query_feat.weight.unsqueeze(0).repeat(
+ (batch_size, 1, 1))
+ query_embed = self.query_embed.weight.unsqueeze(0).repeat(
+ (batch_size, 1, 1))
+
+ cls_pred_list = []
+ mask_pred_list = []
+ cls_pred, mask_pred, attn_mask = self._forward_head(
+ query_feat, mask_features, multi_scale_memorys[0].shape[-2:])
+ cls_pred_list.append(cls_pred)
+ mask_pred_list.append(mask_pred)
+
+ for i in range(self.num_transformer_decoder_layers):
+ level_idx = i % self.num_transformer_feat_level
+ # if a mask is all True(all background), then set it all False.
+ mask_sum = (attn_mask.sum(-1) != attn_mask.shape[-1]).unsqueeze(-1)
+ attn_mask = attn_mask & mask_sum
+ # cross_attn + self_attn
+ layer = self.transformer_decoder.layers[i]
+ query_feat = layer(
+ query=query_feat,
+ key=decoder_inputs[level_idx],
+ value=decoder_inputs[level_idx],
+ query_pos=query_embed,
+ key_pos=decoder_positional_encodings[level_idx],
+ cross_attn_mask=attn_mask,
+ query_key_padding_mask=None,
+ # here we do not apply masking on padded region
+ key_padding_mask=None)
+ cls_pred, mask_pred, attn_mask = self._forward_head(
+ query_feat, mask_features, multi_scale_memorys[
+ (i + 1) % self.num_transformer_feat_level].shape[-2:])
+
+ cls_pred_list.append(cls_pred)
+ mask_pred_list.append(mask_pred)
+
+ return cls_pred_list, mask_pred_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/maskformer_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/maskformer_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..24c0655ee1c36e0110cf6578d1c095c50a297d81
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/maskformer_head.py
@@ -0,0 +1,601 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, List, Optional, Tuple, Union
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import Conv2d
+from mmengine.model import caffe2_xavier_init
+from mmengine.structures import InstanceData, PixelData
+from torch import Tensor
+
+from mmdet.models.layers.pixel_decoder import PixelDecoder
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures import SampleList
+from mmdet.utils import (ConfigType, InstanceList, OptConfigType,
+ OptMultiConfig, reduce_mean)
+from ..layers import DetrTransformerDecoder, SinePositionalEncoding
+from ..utils import multi_apply, preprocess_panoptic_gt
+from .anchor_free_head import AnchorFreeHead
+
+
+@MODELS.register_module()
+class MaskFormerHead(AnchorFreeHead):
+ """Implements the MaskFormer head.
+
+ See `Per-Pixel Classification is Not All You Need for Semantic
+ Segmentation `_ for details.
+
+ Args:
+ in_channels (list[int]): Number of channels in the input feature map.
+ feat_channels (int): Number of channels for feature.
+ out_channels (int): Number of channels for output.
+ num_things_classes (int): Number of things.
+ num_stuff_classes (int): Number of stuff.
+ num_queries (int): Number of query in Transformer.
+ pixel_decoder (:obj:`ConfigDict` or dict): Config for pixel
+ decoder.
+ enforce_decoder_input_project (bool): Whether to add a layer
+ to change the embed_dim of transformer encoder in pixel decoder to
+ the embed_dim of transformer decoder. Defaults to False.
+ transformer_decoder (:obj:`ConfigDict` or dict): Config for
+ transformer decoder.
+ positional_encoding (:obj:`ConfigDict` or dict): Config for
+ transformer decoder position encoding.
+ loss_cls (:obj:`ConfigDict` or dict): Config of the classification
+ loss. Defaults to `CrossEntropyLoss`.
+ loss_mask (:obj:`ConfigDict` or dict): Config of the mask loss.
+ Defaults to `FocalLoss`.
+ loss_dice (:obj:`ConfigDict` or dict): Config of the dice loss.
+ Defaults to `DiceLoss`.
+ train_cfg (:obj:`ConfigDict` or dict, optional): Training config of
+ MaskFormer head.
+ test_cfg (:obj:`ConfigDict` or dict, optional): Testing config of
+ MaskFormer head.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict], optional): Initialization config dict. Defaults to None.
+ """
+
+ def __init__(self,
+ in_channels: List[int],
+ feat_channels: int,
+ out_channels: int,
+ num_things_classes: int = 80,
+ num_stuff_classes: int = 53,
+ num_queries: int = 100,
+ pixel_decoder: ConfigType = ...,
+ enforce_decoder_input_project: bool = False,
+ transformer_decoder: ConfigType = ...,
+ positional_encoding: ConfigType = dict(
+ num_feats=128, normalize=True),
+ loss_cls: ConfigType = dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0,
+ class_weight=[1.0] * 133 + [0.1]),
+ loss_mask: ConfigType = dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=20.0),
+ loss_dice: ConfigType = dict(
+ type='DiceLoss',
+ use_sigmoid=True,
+ activate=True,
+ naive_dice=True,
+ loss_weight=1.0),
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ init_cfg: OptMultiConfig = None,
+ **kwargs) -> None:
+ super(AnchorFreeHead, self).__init__(init_cfg=init_cfg)
+ self.num_things_classes = num_things_classes
+ self.num_stuff_classes = num_stuff_classes
+ self.num_classes = self.num_things_classes + self.num_stuff_classes
+ self.num_queries = num_queries
+
+ pixel_decoder.update(
+ in_channels=in_channels,
+ feat_channels=feat_channels,
+ out_channels=out_channels)
+ self.pixel_decoder = MODELS.build(pixel_decoder)
+ self.transformer_decoder = DetrTransformerDecoder(
+ **transformer_decoder)
+ self.decoder_embed_dims = self.transformer_decoder.embed_dims
+ if type(self.pixel_decoder) == PixelDecoder and (
+ self.decoder_embed_dims != in_channels[-1]
+ or enforce_decoder_input_project):
+ self.decoder_input_proj = Conv2d(
+ in_channels[-1], self.decoder_embed_dims, kernel_size=1)
+ else:
+ self.decoder_input_proj = nn.Identity()
+ self.decoder_pe = SinePositionalEncoding(**positional_encoding)
+ self.query_embed = nn.Embedding(self.num_queries, out_channels)
+
+ self.cls_embed = nn.Linear(feat_channels, self.num_classes + 1)
+ self.mask_embed = nn.Sequential(
+ nn.Linear(feat_channels, feat_channels), nn.ReLU(inplace=True),
+ nn.Linear(feat_channels, feat_channels), nn.ReLU(inplace=True),
+ nn.Linear(feat_channels, out_channels))
+
+ self.test_cfg = test_cfg
+ self.train_cfg = train_cfg
+ if train_cfg:
+ self.assigner = TASK_UTILS.build(train_cfg['assigner'])
+ self.sampler = TASK_UTILS.build(
+ train_cfg['sampler'], default_args=dict(context=self))
+
+ self.class_weight = loss_cls.class_weight
+ self.loss_cls = MODELS.build(loss_cls)
+ self.loss_mask = MODELS.build(loss_mask)
+ self.loss_dice = MODELS.build(loss_dice)
+
+ def init_weights(self) -> None:
+ if isinstance(self.decoder_input_proj, Conv2d):
+ caffe2_xavier_init(self.decoder_input_proj, bias=0)
+
+ self.pixel_decoder.init_weights()
+
+ for p in self.transformer_decoder.parameters():
+ if p.dim() > 1:
+ nn.init.xavier_uniform_(p)
+
+ def preprocess_gt(
+ self, batch_gt_instances: InstanceList,
+ batch_gt_semantic_segs: List[Optional[PixelData]]) -> InstanceList:
+ """Preprocess the ground truth for all images.
+
+ Args:
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``labels``, each is
+ ground truth labels of each bbox, with shape (num_gts, )
+ and ``masks``, each is ground truth masks of each instances
+ of a image, shape (num_gts, h, w).
+ gt_semantic_seg (list[Optional[PixelData]]): Ground truth of
+ semantic segmentation, each with the shape (1, h, w).
+ [0, num_thing_class - 1] means things,
+ [num_thing_class, num_class-1] means stuff,
+ 255 means VOID. It's None when training instance segmentation.
+
+ Returns:
+ list[obj:`InstanceData`]: each contains the following keys
+
+ - labels (Tensor): Ground truth class indices\
+ for a image, with shape (n, ), n is the sum of\
+ number of stuff type and number of instance in a image.
+ - masks (Tensor): Ground truth mask for a\
+ image, with shape (n, h, w).
+ """
+ num_things_list = [self.num_things_classes] * len(batch_gt_instances)
+ num_stuff_list = [self.num_stuff_classes] * len(batch_gt_instances)
+ gt_labels_list = [
+ gt_instances['labels'] for gt_instances in batch_gt_instances
+ ]
+ gt_masks_list = [
+ gt_instances['masks'] for gt_instances in batch_gt_instances
+ ]
+ gt_semantic_segs = [
+ None if gt_semantic_seg is None else gt_semantic_seg.sem_seg
+ for gt_semantic_seg in batch_gt_semantic_segs
+ ]
+ targets = multi_apply(preprocess_panoptic_gt, gt_labels_list,
+ gt_masks_list, gt_semantic_segs, num_things_list,
+ num_stuff_list)
+ labels, masks = targets
+ batch_gt_instances = [
+ InstanceData(labels=label, masks=mask)
+ for label, mask in zip(labels, masks)
+ ]
+ return batch_gt_instances
+
+ def get_targets(
+ self,
+ cls_scores_list: List[Tensor],
+ mask_preds_list: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ return_sampling_results: bool = False
+ ) -> Tuple[List[Union[Tensor, int]]]:
+ """Compute classification and mask targets for all images for a decoder
+ layer.
+
+ Args:
+ cls_scores_list (list[Tensor]): Mask score logits from a single
+ decoder layer for all images. Each with shape (num_queries,
+ cls_out_channels).
+ mask_preds_list (list[Tensor]): Mask logits from a single decoder
+ layer for all images. Each with shape (num_queries, h, w).
+ batch_gt_instances (list[obj:`InstanceData`]): each contains
+ ``labels`` and ``masks``.
+ batch_img_metas (list[dict]): List of image meta information.
+ return_sampling_results (bool): Whether to return the sampling
+ results. Defaults to False.
+
+ Returns:
+ tuple: a tuple containing the following targets.
+
+ - labels_list (list[Tensor]): Labels of all images.\
+ Each with shape (num_queries, ).
+ - label_weights_list (list[Tensor]): Label weights\
+ of all images. Each with shape (num_queries, ).
+ - mask_targets_list (list[Tensor]): Mask targets of\
+ all images. Each with shape (num_queries, h, w).
+ - mask_weights_list (list[Tensor]): Mask weights of\
+ all images. Each with shape (num_queries, ).
+ - avg_factor (int): Average factor that is used to average\
+ the loss. When using sampling method, avg_factor is
+ usually the sum of positive and negative priors. When
+ using `MaskPseudoSampler`, `avg_factor` is usually equal
+ to the number of positive priors.
+
+ additional_returns: This function enables user-defined returns from
+ `self._get_targets_single`. These returns are currently refined
+ to properties at each feature map (i.e. having HxW dimension).
+ The results will be concatenated after the end.
+ """
+ results = multi_apply(self._get_targets_single, cls_scores_list,
+ mask_preds_list, batch_gt_instances,
+ batch_img_metas)
+ (labels_list, label_weights_list, mask_targets_list, mask_weights_list,
+ pos_inds_list, neg_inds_list, sampling_results_list) = results[:7]
+ rest_results = list(results[7:])
+
+ avg_factor = sum(
+ [results.avg_factor for results in sampling_results_list])
+
+ res = (labels_list, label_weights_list, mask_targets_list,
+ mask_weights_list, avg_factor)
+ if return_sampling_results:
+ res = res + (sampling_results_list)
+
+ return res + tuple(rest_results)
+
+ def _get_targets_single(self, cls_score: Tensor, mask_pred: Tensor,
+ gt_instances: InstanceData,
+ img_meta: dict) -> Tuple[Tensor]:
+ """Compute classification and mask targets for one image.
+
+ Args:
+ cls_score (Tensor): Mask score logits from a single decoder layer
+ for one image. Shape (num_queries, cls_out_channels).
+ mask_pred (Tensor): Mask logits for a single decoder layer for one
+ image. Shape (num_queries, h, w).
+ gt_instances (:obj:`InstanceData`): It contains ``labels`` and
+ ``masks``.
+ img_meta (dict): Image informtation.
+
+ Returns:
+ tuple: a tuple containing the following for one image.
+
+ - labels (Tensor): Labels of each image.
+ shape (num_queries, ).
+ - label_weights (Tensor): Label weights of each image.
+ shape (num_queries, ).
+ - mask_targets (Tensor): Mask targets of each image.
+ shape (num_queries, h, w).
+ - mask_weights (Tensor): Mask weights of each image.
+ shape (num_queries, ).
+ - pos_inds (Tensor): Sampled positive indices for each image.
+ - neg_inds (Tensor): Sampled negative indices for each image.
+ - sampling_result (:obj:`SamplingResult`): Sampling results.
+ """
+ gt_masks = gt_instances.masks
+ gt_labels = gt_instances.labels
+
+ target_shape = mask_pred.shape[-2:]
+ if gt_masks.shape[0] > 0:
+ gt_masks_downsampled = F.interpolate(
+ gt_masks.unsqueeze(1).float(), target_shape,
+ mode='nearest').squeeze(1).long()
+ else:
+ gt_masks_downsampled = gt_masks
+
+ pred_instances = InstanceData(scores=cls_score, masks=mask_pred)
+ downsampled_gt_instances = InstanceData(
+ labels=gt_labels, masks=gt_masks_downsampled)
+ # assign and sample
+ assign_result = self.assigner.assign(
+ pred_instances=pred_instances,
+ gt_instances=downsampled_gt_instances,
+ img_meta=img_meta)
+ sampling_result = self.sampler.sample(
+ assign_result=assign_result,
+ pred_instances=pred_instances,
+ gt_instances=gt_instances)
+ pos_inds = sampling_result.pos_inds
+ neg_inds = sampling_result.neg_inds
+
+ # label target
+ labels = gt_labels.new_full((self.num_queries, ),
+ self.num_classes,
+ dtype=torch.long)
+ labels[pos_inds] = gt_labels[sampling_result.pos_assigned_gt_inds]
+ label_weights = gt_labels.new_ones(self.num_queries)
+
+ # mask target
+ mask_targets = gt_masks[sampling_result.pos_assigned_gt_inds]
+ mask_weights = mask_pred.new_zeros((self.num_queries, ))
+ mask_weights[pos_inds] = 1.0
+
+ return (labels, label_weights, mask_targets, mask_weights, pos_inds,
+ neg_inds, sampling_result)
+
+ def loss_by_feat(self, all_cls_scores: Tensor, all_mask_preds: Tensor,
+ batch_gt_instances: List[InstanceData],
+ batch_img_metas: List[dict]) -> Dict[str, Tensor]:
+ """Loss function.
+
+ Args:
+ all_cls_scores (Tensor): Classification scores for all decoder
+ layers with shape (num_decoder, batch_size, num_queries,
+ cls_out_channels). Note `cls_out_channels` should includes
+ background.
+ all_mask_preds (Tensor): Mask scores for all decoder layers with
+ shape (num_decoder, batch_size, num_queries, h, w).
+ batch_gt_instances (list[obj:`InstanceData`]): each contains
+ ``labels`` and ``masks``.
+ batch_img_metas (list[dict]): List of image meta information.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ num_dec_layers = len(all_cls_scores)
+ batch_gt_instances_list = [
+ batch_gt_instances for _ in range(num_dec_layers)
+ ]
+ img_metas_list = [batch_img_metas for _ in range(num_dec_layers)]
+ losses_cls, losses_mask, losses_dice = multi_apply(
+ self._loss_by_feat_single, all_cls_scores, all_mask_preds,
+ batch_gt_instances_list, img_metas_list)
+
+ loss_dict = dict()
+ # loss from the last decoder layer
+ loss_dict['loss_cls'] = losses_cls[-1]
+ loss_dict['loss_mask'] = losses_mask[-1]
+ loss_dict['loss_dice'] = losses_dice[-1]
+ # loss from other decoder layers
+ num_dec_layer = 0
+ for loss_cls_i, loss_mask_i, loss_dice_i in zip(
+ losses_cls[:-1], losses_mask[:-1], losses_dice[:-1]):
+ loss_dict[f'd{num_dec_layer}.loss_cls'] = loss_cls_i
+ loss_dict[f'd{num_dec_layer}.loss_mask'] = loss_mask_i
+ loss_dict[f'd{num_dec_layer}.loss_dice'] = loss_dice_i
+ num_dec_layer += 1
+ return loss_dict
+
+ def _loss_by_feat_single(self, cls_scores: Tensor, mask_preds: Tensor,
+ batch_gt_instances: List[InstanceData],
+ batch_img_metas: List[dict]) -> Tuple[Tensor]:
+ """Loss function for outputs from a single decoder layer.
+
+ Args:
+ cls_scores (Tensor): Mask score logits from a single decoder layer
+ for all images. Shape (batch_size, num_queries,
+ cls_out_channels). Note `cls_out_channels` should includes
+ background.
+ mask_preds (Tensor): Mask logits for a pixel decoder for all
+ images. Shape (batch_size, num_queries, h, w).
+ batch_gt_instances (list[obj:`InstanceData`]): each contains
+ ``labels`` and ``masks``.
+ batch_img_metas (list[dict]): List of image meta information.
+
+ Returns:
+ tuple[Tensor]: Loss components for outputs from a single decoder\
+ layer.
+ """
+ num_imgs = cls_scores.size(0)
+ cls_scores_list = [cls_scores[i] for i in range(num_imgs)]
+ mask_preds_list = [mask_preds[i] for i in range(num_imgs)]
+
+ (labels_list, label_weights_list, mask_targets_list, mask_weights_list,
+ avg_factor) = self.get_targets(cls_scores_list, mask_preds_list,
+ batch_gt_instances, batch_img_metas)
+ # shape (batch_size, num_queries)
+ labels = torch.stack(labels_list, dim=0)
+ # shape (batch_size, num_queries)
+ label_weights = torch.stack(label_weights_list, dim=0)
+ # shape (num_total_gts, h, w)
+ mask_targets = torch.cat(mask_targets_list, dim=0)
+ # shape (batch_size, num_queries)
+ mask_weights = torch.stack(mask_weights_list, dim=0)
+
+ # classfication loss
+ # shape (batch_size * num_queries, )
+ cls_scores = cls_scores.flatten(0, 1)
+ labels = labels.flatten(0, 1)
+ label_weights = label_weights.flatten(0, 1)
+
+ class_weight = cls_scores.new_tensor(self.class_weight)
+ loss_cls = self.loss_cls(
+ cls_scores,
+ labels,
+ label_weights,
+ avg_factor=class_weight[labels].sum())
+
+ num_total_masks = reduce_mean(cls_scores.new_tensor([avg_factor]))
+ num_total_masks = max(num_total_masks, 1)
+
+ # extract positive ones
+ # shape (batch_size, num_queries, h, w) -> (num_total_gts, h, w)
+ mask_preds = mask_preds[mask_weights > 0]
+ target_shape = mask_targets.shape[-2:]
+
+ if mask_targets.shape[0] == 0:
+ # zero match
+ loss_dice = mask_preds.sum()
+ loss_mask = mask_preds.sum()
+ return loss_cls, loss_mask, loss_dice
+
+ # upsample to shape of target
+ # shape (num_total_gts, h, w)
+ mask_preds = F.interpolate(
+ mask_preds.unsqueeze(1),
+ target_shape,
+ mode='bilinear',
+ align_corners=False).squeeze(1)
+
+ # dice loss
+ loss_dice = self.loss_dice(
+ mask_preds, mask_targets, avg_factor=num_total_masks)
+
+ # mask loss
+ # FocalLoss support input of shape (n, num_class)
+ h, w = mask_preds.shape[-2:]
+ # shape (num_total_gts, h, w) -> (num_total_gts * h * w, 1)
+ mask_preds = mask_preds.reshape(-1, 1)
+ # shape (num_total_gts, h, w) -> (num_total_gts * h * w)
+ mask_targets = mask_targets.reshape(-1)
+ # target is (1 - mask_targets) !!!
+ loss_mask = self.loss_mask(
+ mask_preds, 1 - mask_targets, avg_factor=num_total_masks * h * w)
+
+ return loss_cls, loss_mask, loss_dice
+
+ def forward(self, x: Tuple[Tensor],
+ batch_data_samples: SampleList) -> Tuple[Tensor]:
+ """Forward function.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each
+ is a 4D-tensor.
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+
+ Returns:
+ tuple[Tensor]: a tuple contains two elements.
+
+ - all_cls_scores (Tensor): Classification scores for each\
+ scale level. Each is a 4D-tensor with shape\
+ (num_decoder, batch_size, num_queries, cls_out_channels).\
+ Note `cls_out_channels` should includes background.
+ - all_mask_preds (Tensor): Mask scores for each decoder\
+ layer. Each with shape (num_decoder, batch_size,\
+ num_queries, h, w).
+ """
+ batch_img_metas = [
+ data_sample.metainfo for data_sample in batch_data_samples
+ ]
+ batch_size = x[0].shape[0]
+ input_img_h, input_img_w = batch_img_metas[0]['batch_input_shape']
+ padding_mask = x[-1].new_ones((batch_size, input_img_h, input_img_w),
+ dtype=torch.float32)
+ for i in range(batch_size):
+ img_h, img_w = batch_img_metas[i]['img_shape']
+ padding_mask[i, :img_h, :img_w] = 0
+ padding_mask = F.interpolate(
+ padding_mask.unsqueeze(1), size=x[-1].shape[-2:],
+ mode='nearest').to(torch.bool).squeeze(1)
+ # when backbone is swin, memory is output of last stage of swin.
+ # when backbone is r50, memory is output of tranformer encoder.
+ mask_features, memory = self.pixel_decoder(x, batch_img_metas)
+ pos_embed = self.decoder_pe(padding_mask)
+ memory = self.decoder_input_proj(memory)
+ # shape (batch_size, c, h, w) -> (batch_size, h*w, c)
+ memory = memory.flatten(2).permute(0, 2, 1)
+ pos_embed = pos_embed.flatten(2).permute(0, 2, 1)
+ # shape (batch_size, h * w)
+ padding_mask = padding_mask.flatten(1)
+ # shape = (num_queries, embed_dims)
+ query_embed = self.query_embed.weight
+ # shape = (batch_size, num_queries, embed_dims)
+ query_embed = query_embed.unsqueeze(0).repeat(batch_size, 1, 1)
+ target = torch.zeros_like(query_embed)
+ # shape (num_decoder, num_queries, batch_size, embed_dims)
+ out_dec = self.transformer_decoder(
+ query=target,
+ key=memory,
+ value=memory,
+ query_pos=query_embed,
+ key_pos=pos_embed,
+ key_padding_mask=padding_mask)
+
+ # cls_scores
+ all_cls_scores = self.cls_embed(out_dec)
+
+ # mask_preds
+ mask_embed = self.mask_embed(out_dec)
+ all_mask_preds = torch.einsum('lbqc,bchw->lbqhw', mask_embed,
+ mask_features)
+
+ return all_cls_scores, all_mask_preds
+
+ def loss(
+ self,
+ x: Tuple[Tensor],
+ batch_data_samples: SampleList,
+ ) -> Dict[str, Tensor]:
+ """Perform forward propagation and loss calculation of the panoptic
+ head on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Multi-level features from the upstream
+ network, each is a 4D-tensor.
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+
+ Returns:
+ dict[str, Tensor]: a dictionary of loss components
+ """
+ batch_img_metas = []
+ batch_gt_instances = []
+ batch_gt_semantic_segs = []
+ for data_sample in batch_data_samples:
+ batch_img_metas.append(data_sample.metainfo)
+ batch_gt_instances.append(data_sample.gt_instances)
+ if 'gt_sem_seg' in data_sample:
+ batch_gt_semantic_segs.append(data_sample.gt_sem_seg)
+ else:
+ batch_gt_semantic_segs.append(None)
+
+ # forward
+ all_cls_scores, all_mask_preds = self(x, batch_data_samples)
+
+ # preprocess ground truth
+ batch_gt_instances = self.preprocess_gt(batch_gt_instances,
+ batch_gt_semantic_segs)
+
+ # loss
+ losses = self.loss_by_feat(all_cls_scores, all_mask_preds,
+ batch_gt_instances, batch_img_metas)
+
+ return losses
+
+ def predict(self, x: Tuple[Tensor],
+ batch_data_samples: SampleList) -> Tuple[Tensor]:
+ """Test without augmentaton.
+
+ Args:
+ x (tuple[Tensor]): Multi-level features from the
+ upstream network, each is a 4D-tensor.
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+
+ Returns:
+ tuple[Tensor]: A tuple contains two tensors.
+
+ - mask_cls_results (Tensor): Mask classification logits,\
+ shape (batch_size, num_queries, cls_out_channels).
+ Note `cls_out_channels` should includes background.
+ - mask_pred_results (Tensor): Mask logits, shape \
+ (batch_size, num_queries, h, w).
+ """
+ batch_img_metas = [
+ data_sample.metainfo for data_sample in batch_data_samples
+ ]
+ all_cls_scores, all_mask_preds = self(x, batch_data_samples)
+ mask_cls_results = all_cls_scores[-1]
+ mask_pred_results = all_mask_preds[-1]
+
+ # upsample masks
+ img_shape = batch_img_metas[0]['batch_input_shape']
+ mask_pred_results = F.interpolate(
+ mask_pred_results,
+ size=(img_shape[0], img_shape[1]),
+ mode='bilinear',
+ align_corners=False)
+
+ return mask_cls_results, mask_pred_results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/nasfcos_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/nasfcos_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..14ee62a7910d90a108fefb2acef00c91ab83ecc8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/nasfcos_head.py
@@ -0,0 +1,114 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+
+import torch.nn as nn
+from mmcv.cnn import ConvModule, Scale
+
+from mmdet.models.dense_heads.fcos_head import FCOSHead
+from mmdet.registry import MODELS
+from mmdet.utils import OptMultiConfig
+
+
+@MODELS.register_module()
+class NASFCOSHead(FCOSHead):
+ """Anchor-free head used in `NASFCOS `_.
+
+ It is quite similar with FCOS head, except for the searched structure of
+ classification branch and bbox regression branch, where a structure of
+ "dconv3x3, conv3x3, dconv3x3, conv1x1" is utilized instead.
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ strides (Sequence[int] or Sequence[Tuple[int, int]]): Strides of points
+ in multiple feature levels. Defaults to (4, 8, 16, 32, 64).
+ regress_ranges (Sequence[Tuple[int, int]]): Regress range of multiple
+ level points.
+ center_sampling (bool): If true, use center sampling.
+ Defaults to False.
+ center_sample_radius (float): Radius of center sampling.
+ Defaults to 1.5.
+ norm_on_bbox (bool): If true, normalize the regression targets with
+ FPN strides. Defaults to False.
+ centerness_on_reg (bool): If true, position centerness on the
+ regress branch. Please refer to https://github.com/tianzhi0549/FCOS/issues/89#issuecomment-516877042.
+ Defaults to False.
+ conv_bias (bool or str): If specified as `auto`, it will be decided by
+ the norm_cfg. Bias of conv will be set as True if `norm_cfg` is
+ None, otherwise False. Defaults to "auto".
+ loss_cls (:obj:`ConfigDict` or dict): Config of classification loss.
+ loss_bbox (:obj:`ConfigDict` or dict): Config of localization loss.
+ loss_centerness (:obj:`ConfigDict`, or dict): Config of centerness
+ loss.
+ norm_cfg (:obj:`ConfigDict` or dict): dictionary to construct and
+ config norm layer. Defaults to
+ ``norm_cfg=dict(type='GN', num_groups=32, requires_grad=True)``.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict], opitonal): Initialization config dict.
+ """ # noqa: E501
+
+ def __init__(self,
+ *args,
+ init_cfg: OptMultiConfig = None,
+ **kwargs) -> None:
+ if init_cfg is None:
+ init_cfg = [
+ dict(type='Caffe2Xavier', layer=['ConvModule', 'Conv2d']),
+ dict(
+ type='Normal',
+ std=0.01,
+ override=[
+ dict(name='conv_reg'),
+ dict(name='conv_centerness'),
+ dict(
+ name='conv_cls',
+ type='Normal',
+ std=0.01,
+ bias_prob=0.01)
+ ]),
+ ]
+ super().__init__(*args, init_cfg=init_cfg, **kwargs)
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ dconv3x3_config = dict(
+ type='DCNv2',
+ kernel_size=3,
+ use_bias=True,
+ deform_groups=2,
+ padding=1)
+ conv3x3_config = dict(type='Conv', kernel_size=3, padding=1)
+ conv1x1_config = dict(type='Conv', kernel_size=1)
+
+ self.arch_config = [
+ dconv3x3_config, conv3x3_config, dconv3x3_config, conv1x1_config
+ ]
+ self.cls_convs = nn.ModuleList()
+ self.reg_convs = nn.ModuleList()
+ for i, op_ in enumerate(self.arch_config):
+ op = copy.deepcopy(op_)
+ chn = self.in_channels if i == 0 else self.feat_channels
+ assert isinstance(op, dict)
+ use_bias = op.pop('use_bias', False)
+ padding = op.pop('padding', 0)
+ kernel_size = op.pop('kernel_size')
+ module = ConvModule(
+ chn,
+ self.feat_channels,
+ kernel_size,
+ stride=1,
+ padding=padding,
+ norm_cfg=self.norm_cfg,
+ bias=use_bias,
+ conv_cfg=op)
+
+ self.cls_convs.append(copy.deepcopy(module))
+ self.reg_convs.append(copy.deepcopy(module))
+
+ self.conv_cls = nn.Conv2d(
+ self.feat_channels, self.cls_out_channels, 3, padding=1)
+ self.conv_reg = nn.Conv2d(self.feat_channels, 4, 3, padding=1)
+ self.conv_centerness = nn.Conv2d(self.feat_channels, 1, 3, padding=1)
+
+ self.scales = nn.ModuleList([Scale(1.0) for _ in self.strides])
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/paa_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/paa_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..3c1f453d2788b354970254e8875068e824c370d4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/paa_head.py
@@ -0,0 +1,730 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple
+
+import numpy as np
+import torch
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures.bbox import bbox_overlaps
+from mmdet.utils import (ConfigType, InstanceList, OptConfigType,
+ OptInstanceList)
+from ..layers import multiclass_nms
+from ..utils import levels_to_images, multi_apply
+from . import ATSSHead
+
+EPS = 1e-12
+try:
+ import sklearn.mixture as skm
+except ImportError:
+ skm = None
+
+
+@MODELS.register_module()
+class PAAHead(ATSSHead):
+ """Head of PAAAssignment: Probabilistic Anchor Assignment with IoU
+ Prediction for Object Detection.
+
+ Code is modified from the `official github repo
+ `_.
+
+ More details can be found in the `paper
+ `_ .
+
+ Args:
+ topk (int): Select topk samples with smallest loss in
+ each level.
+ score_voting (bool): Whether to use score voting in post-process.
+ covariance_type : String describing the type of covariance parameters
+ to be used in :class:`sklearn.mixture.GaussianMixture`.
+ It must be one of:
+
+ - 'full': each component has its own general covariance matrix
+ - 'tied': all components share the same general covariance matrix
+ - 'diag': each component has its own diagonal covariance matrix
+ - 'spherical': each component has its own single variance
+ Default: 'diag'. From 'full' to 'spherical', the gmm fitting
+ process is faster yet the performance could be influenced. For most
+ cases, 'diag' should be a good choice.
+ """
+
+ def __init__(self,
+ *args,
+ topk: int = 9,
+ score_voting: bool = True,
+ covariance_type: str = 'diag',
+ **kwargs):
+ # topk used in paa reassign process
+ self.topk = topk
+ self.with_score_voting = score_voting
+ self.covariance_type = covariance_type
+ super().__init__(*args, **kwargs)
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ iou_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ Has shape (N, num_anchors * num_classes, H, W)
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W)
+ iou_preds (list[Tensor]): iou_preds for each scale
+ level with shape (N, num_anchors * 1, H, W)
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss gmm_assignment.
+ """
+
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ assert len(featmap_sizes) == self.prior_generator.num_levels
+
+ device = cls_scores[0].device
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+ cls_reg_targets = self.get_targets(
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore,
+ )
+ (labels, labels_weight, bboxes_target, bboxes_weight, pos_inds,
+ pos_gt_index) = cls_reg_targets
+ cls_scores = levels_to_images(cls_scores)
+ cls_scores = [
+ item.reshape(-1, self.cls_out_channels) for item in cls_scores
+ ]
+ bbox_preds = levels_to_images(bbox_preds)
+ bbox_preds = [item.reshape(-1, 4) for item in bbox_preds]
+ iou_preds = levels_to_images(iou_preds)
+ iou_preds = [item.reshape(-1, 1) for item in iou_preds]
+ pos_losses_list, = multi_apply(self.get_pos_loss, anchor_list,
+ cls_scores, bbox_preds, labels,
+ labels_weight, bboxes_target,
+ bboxes_weight, pos_inds)
+
+ with torch.no_grad():
+ reassign_labels, reassign_label_weight, \
+ reassign_bbox_weights, num_pos = multi_apply(
+ self.paa_reassign,
+ pos_losses_list,
+ labels,
+ labels_weight,
+ bboxes_weight,
+ pos_inds,
+ pos_gt_index,
+ anchor_list)
+ num_pos = sum(num_pos)
+ # convert all tensor list to a flatten tensor
+ cls_scores = torch.cat(cls_scores, 0).view(-1, cls_scores[0].size(-1))
+ bbox_preds = torch.cat(bbox_preds, 0).view(-1, bbox_preds[0].size(-1))
+ iou_preds = torch.cat(iou_preds, 0).view(-1, iou_preds[0].size(-1))
+ labels = torch.cat(reassign_labels, 0).view(-1)
+ flatten_anchors = torch.cat(
+ [torch.cat(item, 0) for item in anchor_list])
+ labels_weight = torch.cat(reassign_label_weight, 0).view(-1)
+ bboxes_target = torch.cat(bboxes_target,
+ 0).view(-1, bboxes_target[0].size(-1))
+
+ pos_inds_flatten = ((labels >= 0)
+ &
+ (labels < self.num_classes)).nonzero().reshape(-1)
+
+ losses_cls = self.loss_cls(
+ cls_scores,
+ labels,
+ labels_weight,
+ avg_factor=max(num_pos, len(batch_img_metas))) # avoid num_pos=0
+ if num_pos:
+ pos_bbox_pred = self.bbox_coder.decode(
+ flatten_anchors[pos_inds_flatten],
+ bbox_preds[pos_inds_flatten])
+ pos_bbox_target = bboxes_target[pos_inds_flatten]
+ iou_target = bbox_overlaps(
+ pos_bbox_pred.detach(), pos_bbox_target, is_aligned=True)
+ losses_iou = self.loss_centerness(
+ iou_preds[pos_inds_flatten],
+ iou_target.unsqueeze(-1),
+ avg_factor=num_pos)
+ losses_bbox = self.loss_bbox(
+ pos_bbox_pred,
+ pos_bbox_target,
+ iou_target.clamp(min=EPS),
+ avg_factor=iou_target.sum())
+ else:
+ losses_iou = iou_preds.sum() * 0
+ losses_bbox = bbox_preds.sum() * 0
+
+ return dict(
+ loss_cls=losses_cls, loss_bbox=losses_bbox, loss_iou=losses_iou)
+
+ def get_pos_loss(self, anchors: List[Tensor], cls_score: Tensor,
+ bbox_pred: Tensor, label: Tensor, label_weight: Tensor,
+ bbox_target: dict, bbox_weight: Tensor,
+ pos_inds: Tensor) -> Tensor:
+ """Calculate loss of all potential positive samples obtained from first
+ match process.
+
+ Args:
+ anchors (list[Tensor]): Anchors of each scale.
+ cls_score (Tensor): Box scores of single image with shape
+ (num_anchors, num_classes)
+ bbox_pred (Tensor): Box energies / deltas of single image
+ with shape (num_anchors, 4)
+ label (Tensor): classification target of each anchor with
+ shape (num_anchors,)
+ label_weight (Tensor): Classification loss weight of each
+ anchor with shape (num_anchors).
+ bbox_target (dict): Regression target of each anchor with
+ shape (num_anchors, 4).
+ bbox_weight (Tensor): Bbox weight of each anchor with shape
+ (num_anchors, 4).
+ pos_inds (Tensor): Index of all positive samples got from
+ first assign process.
+
+ Returns:
+ Tensor: Losses of all positive samples in single image.
+ """
+ if not len(pos_inds):
+ return cls_score.new([]),
+ anchors_all_level = torch.cat(anchors, 0)
+ pos_scores = cls_score[pos_inds]
+ pos_bbox_pred = bbox_pred[pos_inds]
+ pos_label = label[pos_inds]
+ pos_label_weight = label_weight[pos_inds]
+ pos_bbox_target = bbox_target[pos_inds]
+ pos_bbox_weight = bbox_weight[pos_inds]
+ pos_anchors = anchors_all_level[pos_inds]
+ pos_bbox_pred = self.bbox_coder.decode(pos_anchors, pos_bbox_pred)
+
+ # to keep loss dimension
+ loss_cls = self.loss_cls(
+ pos_scores,
+ pos_label,
+ pos_label_weight,
+ avg_factor=1.0,
+ reduction_override='none')
+
+ loss_bbox = self.loss_bbox(
+ pos_bbox_pred,
+ pos_bbox_target,
+ pos_bbox_weight,
+ avg_factor=1.0, # keep same loss weight before reassign
+ reduction_override='none')
+
+ loss_cls = loss_cls.sum(-1)
+ pos_loss = loss_bbox + loss_cls
+ return pos_loss,
+
+ def paa_reassign(self, pos_losses: Tensor, label: Tensor,
+ label_weight: Tensor, bbox_weight: Tensor,
+ pos_inds: Tensor, pos_gt_inds: Tensor,
+ anchors: List[Tensor]) -> tuple:
+ """Fit loss to GMM distribution and separate positive, ignore, negative
+ samples again with GMM model.
+
+ Args:
+ pos_losses (Tensor): Losses of all positive samples in
+ single image.
+ label (Tensor): classification target of each anchor with
+ shape (num_anchors,)
+ label_weight (Tensor): Classification loss weight of each
+ anchor with shape (num_anchors).
+ bbox_weight (Tensor): Bbox weight of each anchor with shape
+ (num_anchors, 4).
+ pos_inds (Tensor): Index of all positive samples got from
+ first assign process.
+ pos_gt_inds (Tensor): Gt_index of all positive samples got
+ from first assign process.
+ anchors (list[Tensor]): Anchors of each scale.
+
+ Returns:
+ tuple: Usually returns a tuple containing learning targets.
+
+ - label (Tensor): classification target of each anchor after
+ paa assign, with shape (num_anchors,)
+ - label_weight (Tensor): Classification loss weight of each
+ anchor after paa assign, with shape (num_anchors).
+ - bbox_weight (Tensor): Bbox weight of each anchor with shape
+ (num_anchors, 4).
+ - num_pos (int): The number of positive samples after paa
+ assign.
+ """
+ if not len(pos_inds):
+ return label, label_weight, bbox_weight, 0
+ label = label.clone()
+ label_weight = label_weight.clone()
+ bbox_weight = bbox_weight.clone()
+ num_gt = pos_gt_inds.max() + 1
+ num_level = len(anchors)
+ num_anchors_each_level = [item.size(0) for item in anchors]
+ num_anchors_each_level.insert(0, 0)
+ inds_level_interval = np.cumsum(num_anchors_each_level)
+ pos_level_mask = []
+ for i in range(num_level):
+ mask = (pos_inds >= inds_level_interval[i]) & (
+ pos_inds < inds_level_interval[i + 1])
+ pos_level_mask.append(mask)
+ pos_inds_after_paa = [label.new_tensor([])]
+ ignore_inds_after_paa = [label.new_tensor([])]
+ for gt_ind in range(num_gt):
+ pos_inds_gmm = []
+ pos_loss_gmm = []
+ gt_mask = pos_gt_inds == gt_ind
+ for level in range(num_level):
+ level_mask = pos_level_mask[level]
+ level_gt_mask = level_mask & gt_mask
+ value, topk_inds = pos_losses[level_gt_mask].topk(
+ min(level_gt_mask.sum(), self.topk), largest=False)
+ pos_inds_gmm.append(pos_inds[level_gt_mask][topk_inds])
+ pos_loss_gmm.append(value)
+ pos_inds_gmm = torch.cat(pos_inds_gmm)
+ pos_loss_gmm = torch.cat(pos_loss_gmm)
+ # fix gmm need at least two sample
+ if len(pos_inds_gmm) < 2:
+ continue
+ device = pos_inds_gmm.device
+ pos_loss_gmm, sort_inds = pos_loss_gmm.sort()
+ pos_inds_gmm = pos_inds_gmm[sort_inds]
+ pos_loss_gmm = pos_loss_gmm.view(-1, 1).cpu().numpy()
+ min_loss, max_loss = pos_loss_gmm.min(), pos_loss_gmm.max()
+ means_init = np.array([min_loss, max_loss]).reshape(2, 1)
+ weights_init = np.array([0.5, 0.5])
+ precisions_init = np.array([1.0, 1.0]).reshape(2, 1, 1) # full
+ if self.covariance_type == 'spherical':
+ precisions_init = precisions_init.reshape(2)
+ elif self.covariance_type == 'diag':
+ precisions_init = precisions_init.reshape(2, 1)
+ elif self.covariance_type == 'tied':
+ precisions_init = np.array([[1.0]])
+ if skm is None:
+ raise ImportError('Please run "pip install sklearn" '
+ 'to install sklearn first.')
+ gmm = skm.GaussianMixture(
+ 2,
+ weights_init=weights_init,
+ means_init=means_init,
+ precisions_init=precisions_init,
+ covariance_type=self.covariance_type)
+ gmm.fit(pos_loss_gmm)
+ gmm_assignment = gmm.predict(pos_loss_gmm)
+ scores = gmm.score_samples(pos_loss_gmm)
+ gmm_assignment = torch.from_numpy(gmm_assignment).to(device)
+ scores = torch.from_numpy(scores).to(device)
+
+ pos_inds_temp, ignore_inds_temp = self.gmm_separation_scheme(
+ gmm_assignment, scores, pos_inds_gmm)
+ pos_inds_after_paa.append(pos_inds_temp)
+ ignore_inds_after_paa.append(ignore_inds_temp)
+
+ pos_inds_after_paa = torch.cat(pos_inds_after_paa)
+ ignore_inds_after_paa = torch.cat(ignore_inds_after_paa)
+ reassign_mask = (pos_inds.unsqueeze(1) != pos_inds_after_paa).all(1)
+ reassign_ids = pos_inds[reassign_mask]
+ label[reassign_ids] = self.num_classes
+ label_weight[ignore_inds_after_paa] = 0
+ bbox_weight[reassign_ids] = 0
+ num_pos = len(pos_inds_after_paa)
+ return label, label_weight, bbox_weight, num_pos
+
+ def gmm_separation_scheme(self, gmm_assignment: Tensor, scores: Tensor,
+ pos_inds_gmm: Tensor) -> Tuple[Tensor, Tensor]:
+ """A general separation scheme for gmm model.
+
+ It separates a GMM distribution of candidate samples into three
+ parts, 0 1 and uncertain areas, and you can implement other
+ separation schemes by rewriting this function.
+
+ Args:
+ gmm_assignment (Tensor): The prediction of GMM which is of shape
+ (num_samples,). The 0/1 value indicates the distribution
+ that each sample comes from.
+ scores (Tensor): The probability of sample coming from the
+ fit GMM distribution. The tensor is of shape (num_samples,).
+ pos_inds_gmm (Tensor): All the indexes of samples which are used
+ to fit GMM model. The tensor is of shape (num_samples,)
+
+ Returns:
+ tuple[Tensor, Tensor]: The indices of positive and ignored samples.
+
+ - pos_inds_temp (Tensor): Indices of positive samples.
+ - ignore_inds_temp (Tensor): Indices of ignore samples.
+ """
+ # The implementation is (c) in Fig.3 in origin paper instead of (b).
+ # You can refer to issues such as
+ # https://github.com/kkhoot/PAA/issues/8 and
+ # https://github.com/kkhoot/PAA/issues/9.
+ fgs = gmm_assignment == 0
+ pos_inds_temp = fgs.new_tensor([], dtype=torch.long)
+ ignore_inds_temp = fgs.new_tensor([], dtype=torch.long)
+ if fgs.nonzero().numel():
+ _, pos_thr_ind = scores[fgs].topk(1)
+ pos_inds_temp = pos_inds_gmm[fgs][:pos_thr_ind + 1]
+ ignore_inds_temp = pos_inds_gmm.new_tensor([])
+ return pos_inds_temp, ignore_inds_temp
+
+ def get_targets(self,
+ anchor_list: List[List[Tensor]],
+ valid_flag_list: List[List[Tensor]],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None,
+ unmap_outputs: bool = True) -> tuple:
+ """Get targets for PAA head.
+
+ This method is almost the same as `AnchorHead.get_targets()`. We direct
+ return the results from _get_targets_single instead map it to levels
+ by images_to_levels function.
+
+ Args:
+ anchor_list (list[list[Tensor]]): Multi level anchors of each
+ image. The outer list indicates images, and the inner list
+ corresponds to feature levels of the image. Each element of
+ the inner list is a tensor of shape (num_anchors, 4).
+ valid_flag_list (list[list[Tensor]]): Multi level valid flags of
+ each image. The outer list indicates images, and the inner list
+ corresponds to feature levels of the image. Each element of
+ the inner list is a tensor of shape (num_anchors, )
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ unmap_outputs (bool): Whether to map outputs back to the original
+ set of anchors. Defaults to True.
+
+ Returns:
+ tuple: Usually returns a tuple containing learning targets.
+
+ - labels (list[Tensor]): Labels of all anchors, each with
+ shape (num_anchors,).
+ - label_weights (list[Tensor]): Label weights of all anchor.
+ each with shape (num_anchors,).
+ - bbox_targets (list[Tensor]): BBox targets of all anchors.
+ each with shape (num_anchors, 4).
+ - bbox_weights (list[Tensor]): BBox weights of all anchors.
+ each with shape (num_anchors, 4).
+ - pos_inds (list[Tensor]): Contains all index of positive
+ sample in all anchor.
+ - gt_inds (list[Tensor]): Contains all gt_index of positive
+ sample in all anchor.
+ """
+
+ num_imgs = len(batch_img_metas)
+ assert len(anchor_list) == len(valid_flag_list) == num_imgs
+ concat_anchor_list = []
+ concat_valid_flag_list = []
+ for i in range(num_imgs):
+ assert len(anchor_list[i]) == len(valid_flag_list[i])
+ concat_anchor_list.append(torch.cat(anchor_list[i]))
+ concat_valid_flag_list.append(torch.cat(valid_flag_list[i]))
+
+ # compute targets for each image
+ if batch_gt_instances_ignore is None:
+ batch_gt_instances_ignore = [None] * num_imgs
+ results = multi_apply(
+ self._get_targets_single,
+ concat_anchor_list,
+ concat_valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore,
+ unmap_outputs=unmap_outputs)
+
+ (labels, label_weights, bbox_targets, bbox_weights, valid_pos_inds,
+ valid_neg_inds, sampling_result) = results
+
+ # Due to valid flag of anchors, we have to calculate the real pos_inds
+ # in origin anchor set.
+ pos_inds = []
+ for i, single_labels in enumerate(labels):
+ pos_mask = (0 <= single_labels) & (
+ single_labels < self.num_classes)
+ pos_inds.append(pos_mask.nonzero().view(-1))
+
+ gt_inds = [item.pos_assigned_gt_inds for item in sampling_result]
+ return (labels, label_weights, bbox_targets, bbox_weights, pos_inds,
+ gt_inds)
+
+ def _get_targets_single(self,
+ flat_anchors: Tensor,
+ valid_flags: Tensor,
+ gt_instances: InstanceData,
+ img_meta: dict,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ unmap_outputs: bool = True) -> tuple:
+ """Compute regression and classification targets for anchors in a
+ single image.
+
+ This method is same as `AnchorHead._get_targets_single()`.
+ """
+ assert unmap_outputs, 'We must map outputs back to the original' \
+ 'set of anchors in PAAhead'
+ return super(ATSSHead, self)._get_targets_single(
+ flat_anchors,
+ valid_flags,
+ gt_instances,
+ img_meta,
+ gt_instances_ignore,
+ unmap_outputs=True)
+
+ def predict_by_feat(self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ score_factors: Optional[List[Tensor]] = None,
+ batch_img_metas: Optional[List[dict]] = None,
+ cfg: OptConfigType = None,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ bbox results.
+
+ This method is same as `BaseDenseHead.get_results()`.
+ """
+ assert with_nms, 'PAA only supports "with_nms=True" now and it ' \
+ 'means PAAHead does not support ' \
+ 'test-time augmentation'
+ return super().predict_by_feat(
+ cls_scores=cls_scores,
+ bbox_preds=bbox_preds,
+ score_factors=score_factors,
+ batch_img_metas=batch_img_metas,
+ cfg=cfg,
+ rescale=rescale,
+ with_nms=with_nms)
+
+ def _predict_by_feat_single(self,
+ cls_score_list: List[Tensor],
+ bbox_pred_list: List[Tensor],
+ score_factor_list: List[Tensor],
+ mlvl_priors: List[Tensor],
+ img_meta: dict,
+ cfg: OptConfigType = None,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results.
+
+ Args:
+ cls_score_list (list[Tensor]): Box scores from all scale
+ levels of a single image, each item has shape
+ (num_priors * num_classes, H, W).
+ bbox_pred_list (list[Tensor]): Box energies / deltas from
+ all scale levels of a single image, each item has shape
+ (num_priors * 4, H, W).
+ score_factor_list (list[Tensor]): Score factors from all scale
+ levels of a single image, each item has shape
+ (num_priors * 1, H, W).
+ mlvl_priors (list[Tensor]): Each element in the list is
+ the priors of a single level in feature pyramid, has shape
+ (num_priors, 4).
+ img_meta (dict): Image meta info.
+ cfg (:obj:`ConfigDict` or dict, optional): Test / postprocessing
+ configuration, if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Default: False.
+ with_nms (bool): If True, do nms before return boxes.
+ Default: True.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ cfg = self.test_cfg if cfg is None else cfg
+ img_shape = img_meta['img_shape']
+ nms_pre = cfg.get('nms_pre', -1)
+
+ mlvl_bboxes = []
+ mlvl_scores = []
+ mlvl_score_factors = []
+ for level_idx, (cls_score, bbox_pred, score_factor, priors) in \
+ enumerate(zip(cls_score_list, bbox_pred_list,
+ score_factor_list, mlvl_priors)):
+ assert cls_score.size()[-2:] == bbox_pred.size()[-2:]
+
+ scores = cls_score.permute(1, 2, 0).reshape(
+ -1, self.cls_out_channels).sigmoid()
+ bbox_pred = bbox_pred.permute(1, 2, 0).reshape(-1, 4)
+ score_factor = score_factor.permute(1, 2, 0).reshape(-1).sigmoid()
+
+ if 0 < nms_pre < scores.shape[0]:
+ max_scores, _ = (scores *
+ score_factor[:, None]).sqrt().max(dim=1)
+ _, topk_inds = max_scores.topk(nms_pre)
+ priors = priors[topk_inds, :]
+ bbox_pred = bbox_pred[topk_inds, :]
+ scores = scores[topk_inds, :]
+ score_factor = score_factor[topk_inds]
+
+ bboxes = self.bbox_coder.decode(
+ priors, bbox_pred, max_shape=img_shape)
+ mlvl_bboxes.append(bboxes)
+ mlvl_scores.append(scores)
+ mlvl_score_factors.append(score_factor)
+
+ results = InstanceData()
+ results.bboxes = torch.cat(mlvl_bboxes)
+ results.scores = torch.cat(mlvl_scores)
+ results.score_factors = torch.cat(mlvl_score_factors)
+
+ return self._bbox_post_process(results, cfg, rescale, with_nms,
+ img_meta)
+
+ def _bbox_post_process(self,
+ results: InstanceData,
+ cfg: ConfigType,
+ rescale: bool = False,
+ with_nms: bool = True,
+ img_meta: Optional[dict] = None):
+ """bbox post-processing method.
+
+ The boxes would be rescaled to the original image scale and do
+ the nms operation. Usually with_nms is False is used for aug test.
+
+ Args:
+ results (:obj:`InstaceData`): Detection instance results,
+ each item has shape (num_bboxes, ).
+ cfg (:obj:`ConfigDict` or dict): Test / postprocessing
+ configuration, if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Default: False.
+ with_nms (bool): If True, do nms before return boxes.
+ Default: True.
+ img_meta (dict, optional): Image meta info. Defaults to None.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ if rescale:
+ results.bboxes /= results.bboxes.new_tensor(
+ img_meta['scale_factor']).repeat((1, 2))
+ # Add a dummy background class to the backend when using sigmoid
+ # remind that we set FG labels to [0, num_class-1] since mmdet v2.0
+ # BG cat_id: num_class
+ padding = results.scores.new_zeros(results.scores.shape[0], 1)
+ mlvl_scores = torch.cat([results.scores, padding], dim=1)
+
+ mlvl_nms_scores = (mlvl_scores * results.score_factors[:, None]).sqrt()
+ det_bboxes, det_labels = multiclass_nms(
+ results.bboxes,
+ mlvl_nms_scores,
+ cfg.score_thr,
+ cfg.nms,
+ cfg.max_per_img,
+ score_factors=None)
+ if self.with_score_voting and len(det_bboxes) > 0:
+ det_bboxes, det_labels = self.score_voting(det_bboxes, det_labels,
+ results.bboxes,
+ mlvl_nms_scores,
+ cfg.score_thr)
+ nms_results = InstanceData()
+ nms_results.bboxes = det_bboxes[:, :-1]
+ nms_results.scores = det_bboxes[:, -1]
+ nms_results.labels = det_labels
+ return nms_results
+
+ def score_voting(self, det_bboxes: Tensor, det_labels: Tensor,
+ mlvl_bboxes: Tensor, mlvl_nms_scores: Tensor,
+ score_thr: float) -> Tuple[Tensor, Tensor]:
+ """Implementation of score voting method works on each remaining boxes
+ after NMS procedure.
+
+ Args:
+ det_bboxes (Tensor): Remaining boxes after NMS procedure,
+ with shape (k, 5), each dimension means
+ (x1, y1, x2, y2, score).
+ det_labels (Tensor): The label of remaining boxes, with shape
+ (k, 1),Labels are 0-based.
+ mlvl_bboxes (Tensor): All boxes before the NMS procedure,
+ with shape (num_anchors,4).
+ mlvl_nms_scores (Tensor): The scores of all boxes which is used
+ in the NMS procedure, with shape (num_anchors, num_class)
+ score_thr (float): The score threshold of bboxes.
+
+ Returns:
+ tuple: Usually returns a tuple containing voting results.
+
+ - det_bboxes_voted (Tensor): Remaining boxes after
+ score voting procedure, with shape (k, 5), each
+ dimension means (x1, y1, x2, y2, score).
+ - det_labels_voted (Tensor): Label of remaining bboxes
+ after voting, with shape (num_anchors,).
+ """
+ candidate_mask = mlvl_nms_scores > score_thr
+ candidate_mask_nonzeros = candidate_mask.nonzero(as_tuple=False)
+ candidate_inds = candidate_mask_nonzeros[:, 0]
+ candidate_labels = candidate_mask_nonzeros[:, 1]
+ candidate_bboxes = mlvl_bboxes[candidate_inds]
+ candidate_scores = mlvl_nms_scores[candidate_mask]
+ det_bboxes_voted = []
+ det_labels_voted = []
+ for cls in range(self.cls_out_channels):
+ candidate_cls_mask = candidate_labels == cls
+ if not candidate_cls_mask.any():
+ continue
+ candidate_cls_scores = candidate_scores[candidate_cls_mask]
+ candidate_cls_bboxes = candidate_bboxes[candidate_cls_mask]
+ det_cls_mask = det_labels == cls
+ det_cls_bboxes = det_bboxes[det_cls_mask].view(
+ -1, det_bboxes.size(-1))
+ det_candidate_ious = bbox_overlaps(det_cls_bboxes[:, :4],
+ candidate_cls_bboxes)
+ for det_ind in range(len(det_cls_bboxes)):
+ single_det_ious = det_candidate_ious[det_ind]
+ pos_ious_mask = single_det_ious > 0.01
+ pos_ious = single_det_ious[pos_ious_mask]
+ pos_bboxes = candidate_cls_bboxes[pos_ious_mask]
+ pos_scores = candidate_cls_scores[pos_ious_mask]
+ pis = (torch.exp(-(1 - pos_ious)**2 / 0.025) *
+ pos_scores)[:, None]
+ voted_box = torch.sum(
+ pis * pos_bboxes, dim=0) / torch.sum(
+ pis, dim=0)
+ voted_score = det_cls_bboxes[det_ind][-1:][None, :]
+ det_bboxes_voted.append(
+ torch.cat((voted_box[None, :], voted_score), dim=1))
+ det_labels_voted.append(cls)
+
+ det_bboxes_voted = torch.cat(det_bboxes_voted, dim=0)
+ det_labels_voted = det_labels.new_tensor(det_labels_voted)
+ return det_bboxes_voted, det_labels_voted
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/pisa_retinanet_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/pisa_retinanet_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..85fd54f5be3605d0994c2a2d4d9d7deac4c0f284
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/pisa_retinanet_head.py
@@ -0,0 +1,154 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List
+
+import torch
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import InstanceList, OptInstanceList
+from ..losses import carl_loss, isr_p
+from ..utils import images_to_levels
+from .retina_head import RetinaHead
+
+
+@MODELS.register_module()
+class PISARetinaHead(RetinaHead):
+ """PISA Retinanet Head.
+
+ The head owns the same structure with Retinanet Head, but differs in two
+ aspects:
+ 1. Importance-based Sample Reweighting Positive (ISR-P) is applied to
+ change the positive loss weights.
+ 2. Classification-aware regression loss is adopted as a third loss.
+ """
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Compute losses of the head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ Has shape (N, num_anchors * num_classes, H, W)
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W)
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict: Loss dict, comprise classification loss, regression loss and
+ carl loss.
+ """
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ assert len(featmap_sizes) == self.prior_generator.num_levels
+
+ device = cls_scores[0].device
+
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+ label_channels = self.cls_out_channels if self.use_sigmoid_cls else 1
+ cls_reg_targets = self.get_targets(
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore,
+ return_sampling_results=True)
+ if cls_reg_targets is None:
+ return None
+ (labels_list, label_weights_list, bbox_targets_list, bbox_weights_list,
+ avg_factor, sampling_results_list) = cls_reg_targets
+
+ # anchor number of multi levels
+ num_level_anchors = [anchors.size(0) for anchors in anchor_list[0]]
+ # concat all level anchors and flags to a single tensor
+ concat_anchor_list = []
+ for i in range(len(anchor_list)):
+ concat_anchor_list.append(torch.cat(anchor_list[i]))
+ all_anchor_list = images_to_levels(concat_anchor_list,
+ num_level_anchors)
+
+ num_imgs = len(batch_img_metas)
+ flatten_cls_scores = [
+ cls_score.permute(0, 2, 3, 1).reshape(num_imgs, -1, label_channels)
+ for cls_score in cls_scores
+ ]
+ flatten_cls_scores = torch.cat(
+ flatten_cls_scores, dim=1).reshape(-1,
+ flatten_cls_scores[0].size(-1))
+ flatten_bbox_preds = [
+ bbox_pred.permute(0, 2, 3, 1).reshape(num_imgs, -1, 4)
+ for bbox_pred in bbox_preds
+ ]
+ flatten_bbox_preds = torch.cat(
+ flatten_bbox_preds, dim=1).view(-1, flatten_bbox_preds[0].size(-1))
+ flatten_labels = torch.cat(labels_list, dim=1).reshape(-1)
+ flatten_label_weights = torch.cat(
+ label_weights_list, dim=1).reshape(-1)
+ flatten_anchors = torch.cat(all_anchor_list, dim=1).reshape(-1, 4)
+ flatten_bbox_targets = torch.cat(
+ bbox_targets_list, dim=1).reshape(-1, 4)
+ flatten_bbox_weights = torch.cat(
+ bbox_weights_list, dim=1).reshape(-1, 4)
+
+ # Apply ISR-P
+ isr_cfg = self.train_cfg.get('isr', None)
+ if isr_cfg is not None:
+ all_targets = (flatten_labels, flatten_label_weights,
+ flatten_bbox_targets, flatten_bbox_weights)
+ with torch.no_grad():
+ all_targets = isr_p(
+ flatten_cls_scores,
+ flatten_bbox_preds,
+ all_targets,
+ flatten_anchors,
+ sampling_results_list,
+ bbox_coder=self.bbox_coder,
+ loss_cls=self.loss_cls,
+ num_class=self.num_classes,
+ **self.train_cfg['isr'])
+ (flatten_labels, flatten_label_weights, flatten_bbox_targets,
+ flatten_bbox_weights) = all_targets
+
+ # For convenience we compute loss once instead separating by fpn level,
+ # so that we don't need to separate the weights by level again.
+ # The result should be the same
+ losses_cls = self.loss_cls(
+ flatten_cls_scores,
+ flatten_labels,
+ flatten_label_weights,
+ avg_factor=avg_factor)
+ losses_bbox = self.loss_bbox(
+ flatten_bbox_preds,
+ flatten_bbox_targets,
+ flatten_bbox_weights,
+ avg_factor=avg_factor)
+ loss_dict = dict(loss_cls=losses_cls, loss_bbox=losses_bbox)
+
+ # CARL Loss
+ carl_cfg = self.train_cfg.get('carl', None)
+ if carl_cfg is not None:
+ loss_carl = carl_loss(
+ flatten_cls_scores,
+ flatten_labels,
+ flatten_bbox_preds,
+ flatten_bbox_targets,
+ self.loss_bbox,
+ **self.train_cfg['carl'],
+ avg_factor=avg_factor,
+ sigmoid=True,
+ num_class=self.num_classes)
+ loss_dict.update(loss_carl)
+
+ return loss_dict
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/pisa_ssd_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/pisa_ssd_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..ec09cb40a9c95d3f9889d736b80dfccef07f6fd1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/pisa_ssd_head.py
@@ -0,0 +1,182 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, List, Union
+
+import torch
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import InstanceList, OptInstanceList
+from ..losses import CrossEntropyLoss, SmoothL1Loss, carl_loss, isr_p
+from ..utils import multi_apply
+from .ssd_head import SSDHead
+
+
+# TODO: add loss evaluator for SSD
+@MODELS.register_module()
+class PISASSDHead(SSDHead):
+ """Implementation of `PISA SSD head `_
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (Sequence[int]): Number of channels in the input feature
+ map.
+ stacked_convs (int): Number of conv layers in cls and reg tower.
+ Defaults to 0.
+ feat_channels (int): Number of hidden channels when stacked_convs
+ > 0. Defaults to 256.
+ use_depthwise (bool): Whether to use DepthwiseSeparableConv.
+ Defaults to False.
+ conv_cfg (:obj:`ConfigDict` or dict, Optional): Dictionary to construct
+ and config conv layer. Defaults to None.
+ norm_cfg (:obj:`ConfigDict` or dict, Optional): Dictionary to construct
+ and config norm layer. Defaults to None.
+ act_cfg (:obj:`ConfigDict` or dict, Optional): Dictionary to construct
+ and config activation layer. Defaults to None.
+ anchor_generator (:obj:`ConfigDict` or dict): Config dict for anchor
+ generator.
+ bbox_coder (:obj:`ConfigDict` or dict): Config of bounding box coder.
+ reg_decoded_bbox (bool): If true, the regression loss would be
+ applied directly on decoded bounding boxes, converting both
+ the predicted boxes and regression targets to absolute
+ coordinates format. Defaults to False. It should be `True` when
+ using `IoULoss`, `GIoULoss`, or `DIoULoss` in the bbox head.
+ train_cfg (:obj:`ConfigDict` or dict, Optional): Training config of
+ anchor head.
+ test_cfg (:obj:`ConfigDict` or dict, Optional): Testing config of
+ anchor head.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict], Optional): Initialization config dict.
+ """ # noqa: W605
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None
+ ) -> Dict[str, Union[List[Tensor], Tensor]]:
+ """Compute losses of the head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ Has shape (N, num_anchors * num_classes, H, W)
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W)
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Union[List[Tensor], Tensor]]: A dictionary of loss
+ components. the dict has components below:
+
+ - loss_cls (list[Tensor]): A list containing each feature map \
+ classification loss.
+ - loss_bbox (list[Tensor]): A list containing each feature map \
+ regression loss.
+ - loss_carl (Tensor): The loss of CARL.
+ """
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ assert len(featmap_sizes) == self.prior_generator.num_levels
+
+ device = cls_scores[0].device
+
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+ cls_reg_targets = self.get_targets(
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore,
+ unmap_outputs=False,
+ return_sampling_results=True)
+ (labels_list, label_weights_list, bbox_targets_list, bbox_weights_list,
+ avg_factor, sampling_results_list) = cls_reg_targets
+
+ num_images = len(batch_img_metas)
+ all_cls_scores = torch.cat([
+ s.permute(0, 2, 3, 1).reshape(
+ num_images, -1, self.cls_out_channels) for s in cls_scores
+ ], 1)
+ all_labels = torch.cat(labels_list, -1).view(num_images, -1)
+ all_label_weights = torch.cat(label_weights_list,
+ -1).view(num_images, -1)
+ all_bbox_preds = torch.cat([
+ b.permute(0, 2, 3, 1).reshape(num_images, -1, 4)
+ for b in bbox_preds
+ ], -2)
+ all_bbox_targets = torch.cat(bbox_targets_list,
+ -2).view(num_images, -1, 4)
+ all_bbox_weights = torch.cat(bbox_weights_list,
+ -2).view(num_images, -1, 4)
+
+ # concat all level anchors to a single tensor
+ all_anchors = []
+ for i in range(num_images):
+ all_anchors.append(torch.cat(anchor_list[i]))
+
+ isr_cfg = self.train_cfg.get('isr', None)
+ all_targets = (all_labels.view(-1), all_label_weights.view(-1),
+ all_bbox_targets.view(-1,
+ 4), all_bbox_weights.view(-1, 4))
+ # apply ISR-P
+ if isr_cfg is not None:
+ all_targets = isr_p(
+ all_cls_scores.view(-1, all_cls_scores.size(-1)),
+ all_bbox_preds.view(-1, 4),
+ all_targets,
+ torch.cat(all_anchors),
+ sampling_results_list,
+ loss_cls=CrossEntropyLoss(),
+ bbox_coder=self.bbox_coder,
+ **self.train_cfg['isr'],
+ num_class=self.num_classes)
+ (new_labels, new_label_weights, new_bbox_targets,
+ new_bbox_weights) = all_targets
+ all_labels = new_labels.view(all_labels.shape)
+ all_label_weights = new_label_weights.view(all_label_weights.shape)
+ all_bbox_targets = new_bbox_targets.view(all_bbox_targets.shape)
+ all_bbox_weights = new_bbox_weights.view(all_bbox_weights.shape)
+
+ # add CARL loss
+ carl_loss_cfg = self.train_cfg.get('carl', None)
+ if carl_loss_cfg is not None:
+ loss_carl = carl_loss(
+ all_cls_scores.view(-1, all_cls_scores.size(-1)),
+ all_targets[0],
+ all_bbox_preds.view(-1, 4),
+ all_targets[2],
+ SmoothL1Loss(beta=1.),
+ **self.train_cfg['carl'],
+ avg_factor=avg_factor,
+ num_class=self.num_classes)
+
+ # check NaN and Inf
+ assert torch.isfinite(all_cls_scores).all().item(), \
+ 'classification scores become infinite or NaN!'
+ assert torch.isfinite(all_bbox_preds).all().item(), \
+ 'bbox predications become infinite or NaN!'
+
+ losses_cls, losses_bbox = multi_apply(
+ self.loss_by_feat_single,
+ all_cls_scores,
+ all_bbox_preds,
+ all_anchors,
+ all_labels,
+ all_label_weights,
+ all_bbox_targets,
+ all_bbox_weights,
+ avg_factor=avg_factor)
+ loss_dict = dict(loss_cls=losses_cls, loss_bbox=losses_bbox)
+ if carl_loss_cfg is not None:
+ loss_dict.update(loss_carl)
+ return loss_dict
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/reppoints_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/reppoints_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..22f3e3401a4abd9cc35b41d24efe23e5655a905e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/reppoints_head.py
@@ -0,0 +1,885 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, List, Sequence, Tuple
+
+import numpy as np
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from mmcv.ops import DeformConv2d
+from mmengine.config import ConfigDict
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.utils import ConfigType, InstanceList, MultiConfig, OptInstanceList
+from ..task_modules.prior_generators import MlvlPointGenerator
+from ..task_modules.samplers import PseudoSampler
+from ..utils import (filter_scores_and_topk, images_to_levels, multi_apply,
+ unmap)
+from .anchor_free_head import AnchorFreeHead
+
+
+@MODELS.register_module()
+class RepPointsHead(AnchorFreeHead):
+ """RepPoint head.
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ point_feat_channels (int): Number of channels of points features.
+ num_points (int): Number of points.
+ gradient_mul (float): The multiplier to gradients from
+ points refinement and recognition.
+ point_strides (Sequence[int]): points strides.
+ point_base_scale (int): bbox scale for assigning labels.
+ loss_cls (:obj:`ConfigDict` or dict): Config of classification loss.
+ loss_bbox_init (:obj:`ConfigDict` or dict): Config of initial points
+ loss.
+ loss_bbox_refine (:obj:`ConfigDict` or dict): Config of points loss in
+ refinement.
+ use_grid_points (bool): If we use bounding box representation, the
+ reppoints is represented as grid points on the bounding box.
+ center_init (bool): Whether to use center point assignment.
+ transform_method (str): The methods to transform RepPoints to bbox.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict]): Initialization config dict.
+ """ # noqa: W605
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: int,
+ point_feat_channels: int = 256,
+ num_points: int = 9,
+ gradient_mul: float = 0.1,
+ point_strides: Sequence[int] = [8, 16, 32, 64, 128],
+ point_base_scale: int = 4,
+ loss_cls: ConfigType = dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox_init: ConfigType = dict(
+ type='SmoothL1Loss', beta=1.0 / 9.0, loss_weight=0.5),
+ loss_bbox_refine: ConfigType = dict(
+ type='SmoothL1Loss', beta=1.0 / 9.0, loss_weight=1.0),
+ use_grid_points: bool = False,
+ center_init: bool = True,
+ transform_method: str = 'moment',
+ moment_mul: float = 0.01,
+ init_cfg: MultiConfig = dict(
+ type='Normal',
+ layer='Conv2d',
+ std=0.01,
+ override=dict(
+ type='Normal',
+ name='reppoints_cls_out',
+ std=0.01,
+ bias_prob=0.01)),
+ **kwargs) -> None:
+ self.num_points = num_points
+ self.point_feat_channels = point_feat_channels
+ self.use_grid_points = use_grid_points
+ self.center_init = center_init
+
+ # we use deform conv to extract points features
+ self.dcn_kernel = int(np.sqrt(num_points))
+ self.dcn_pad = int((self.dcn_kernel - 1) / 2)
+ assert self.dcn_kernel * self.dcn_kernel == num_points, \
+ 'The points number should be a square number.'
+ assert self.dcn_kernel % 2 == 1, \
+ 'The points number should be an odd square number.'
+ dcn_base = np.arange(-self.dcn_pad,
+ self.dcn_pad + 1).astype(np.float64)
+ dcn_base_y = np.repeat(dcn_base, self.dcn_kernel)
+ dcn_base_x = np.tile(dcn_base, self.dcn_kernel)
+ dcn_base_offset = np.stack([dcn_base_y, dcn_base_x], axis=1).reshape(
+ (-1))
+ self.dcn_base_offset = torch.tensor(dcn_base_offset).view(1, -1, 1, 1)
+
+ super().__init__(
+ num_classes=num_classes,
+ in_channels=in_channels,
+ loss_cls=loss_cls,
+ init_cfg=init_cfg,
+ **kwargs)
+
+ self.gradient_mul = gradient_mul
+ self.point_base_scale = point_base_scale
+ self.point_strides = point_strides
+ self.prior_generator = MlvlPointGenerator(
+ self.point_strides, offset=0.)
+
+ if self.train_cfg:
+ self.init_assigner = TASK_UTILS.build(
+ self.train_cfg['init']['assigner'])
+ self.refine_assigner = TASK_UTILS.build(
+ self.train_cfg['refine']['assigner'])
+
+ if self.train_cfg.get('sampler', None) is not None:
+ self.sampler = TASK_UTILS.build(
+ self.train_cfg['sampler'], default_args=dict(context=self))
+ else:
+ self.sampler = PseudoSampler(context=self)
+
+ self.transform_method = transform_method
+ if self.transform_method == 'moment':
+ self.moment_transfer = nn.Parameter(
+ data=torch.zeros(2), requires_grad=True)
+ self.moment_mul = moment_mul
+
+ self.use_sigmoid_cls = loss_cls.get('use_sigmoid', False)
+ if self.use_sigmoid_cls:
+ self.cls_out_channels = self.num_classes
+ else:
+ self.cls_out_channels = self.num_classes + 1
+ self.loss_bbox_init = MODELS.build(loss_bbox_init)
+ self.loss_bbox_refine = MODELS.build(loss_bbox_refine)
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self.relu = nn.ReLU(inplace=True)
+ self.cls_convs = nn.ModuleList()
+ self.reg_convs = nn.ModuleList()
+ for i in range(self.stacked_convs):
+ chn = self.in_channels if i == 0 else self.feat_channels
+ self.cls_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ self.reg_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ pts_out_dim = 4 if self.use_grid_points else 2 * self.num_points
+ self.reppoints_cls_conv = DeformConv2d(self.feat_channels,
+ self.point_feat_channels,
+ self.dcn_kernel, 1,
+ self.dcn_pad)
+ self.reppoints_cls_out = nn.Conv2d(self.point_feat_channels,
+ self.cls_out_channels, 1, 1, 0)
+ self.reppoints_pts_init_conv = nn.Conv2d(self.feat_channels,
+ self.point_feat_channels, 3,
+ 1, 1)
+ self.reppoints_pts_init_out = nn.Conv2d(self.point_feat_channels,
+ pts_out_dim, 1, 1, 0)
+ self.reppoints_pts_refine_conv = DeformConv2d(self.feat_channels,
+ self.point_feat_channels,
+ self.dcn_kernel, 1,
+ self.dcn_pad)
+ self.reppoints_pts_refine_out = nn.Conv2d(self.point_feat_channels,
+ pts_out_dim, 1, 1, 0)
+
+ def points2bbox(self, pts: Tensor, y_first: bool = True) -> Tensor:
+ """Converting the points set into bounding box.
+
+ Args:
+ pts (Tensor): the input points sets (fields), each points
+ set (fields) is represented as 2n scalar.
+ y_first (bool): if y_first=True, the point set is
+ represented as [y1, x1, y2, x2 ... yn, xn], otherwise
+ the point set is represented as
+ [x1, y1, x2, y2 ... xn, yn]. Defaults to True.
+
+ Returns:
+ Tensor: each points set is converting to a bbox [x1, y1, x2, y2].
+ """
+ pts_reshape = pts.view(pts.shape[0], -1, 2, *pts.shape[2:])
+ pts_y = pts_reshape[:, :, 0, ...] if y_first else pts_reshape[:, :, 1,
+ ...]
+ pts_x = pts_reshape[:, :, 1, ...] if y_first else pts_reshape[:, :, 0,
+ ...]
+ if self.transform_method == 'minmax':
+ bbox_left = pts_x.min(dim=1, keepdim=True)[0]
+ bbox_right = pts_x.max(dim=1, keepdim=True)[0]
+ bbox_up = pts_y.min(dim=1, keepdim=True)[0]
+ bbox_bottom = pts_y.max(dim=1, keepdim=True)[0]
+ bbox = torch.cat([bbox_left, bbox_up, bbox_right, bbox_bottom],
+ dim=1)
+ elif self.transform_method == 'partial_minmax':
+ pts_y = pts_y[:, :4, ...]
+ pts_x = pts_x[:, :4, ...]
+ bbox_left = pts_x.min(dim=1, keepdim=True)[0]
+ bbox_right = pts_x.max(dim=1, keepdim=True)[0]
+ bbox_up = pts_y.min(dim=1, keepdim=True)[0]
+ bbox_bottom = pts_y.max(dim=1, keepdim=True)[0]
+ bbox = torch.cat([bbox_left, bbox_up, bbox_right, bbox_bottom],
+ dim=1)
+ elif self.transform_method == 'moment':
+ pts_y_mean = pts_y.mean(dim=1, keepdim=True)
+ pts_x_mean = pts_x.mean(dim=1, keepdim=True)
+ pts_y_std = torch.std(pts_y - pts_y_mean, dim=1, keepdim=True)
+ pts_x_std = torch.std(pts_x - pts_x_mean, dim=1, keepdim=True)
+ moment_transfer = (self.moment_transfer * self.moment_mul) + (
+ self.moment_transfer.detach() * (1 - self.moment_mul))
+ moment_width_transfer = moment_transfer[0]
+ moment_height_transfer = moment_transfer[1]
+ half_width = pts_x_std * torch.exp(moment_width_transfer)
+ half_height = pts_y_std * torch.exp(moment_height_transfer)
+ bbox = torch.cat([
+ pts_x_mean - half_width, pts_y_mean - half_height,
+ pts_x_mean + half_width, pts_y_mean + half_height
+ ],
+ dim=1)
+ else:
+ raise NotImplementedError
+ return bbox
+
+ def gen_grid_from_reg(self, reg: Tensor,
+ previous_boxes: Tensor) -> Tuple[Tensor]:
+ """Base on the previous bboxes and regression values, we compute the
+ regressed bboxes and generate the grids on the bboxes.
+
+ Args:
+ reg (Tensor): the regression value to previous bboxes.
+ previous_boxes (Tensor): previous bboxes.
+
+ Returns:
+ Tuple[Tensor]: generate grids on the regressed bboxes.
+ """
+ b, _, h, w = reg.shape
+ bxy = (previous_boxes[:, :2, ...] + previous_boxes[:, 2:, ...]) / 2.
+ bwh = (previous_boxes[:, 2:, ...] -
+ previous_boxes[:, :2, ...]).clamp(min=1e-6)
+ grid_topleft = bxy + bwh * reg[:, :2, ...] - 0.5 * bwh * torch.exp(
+ reg[:, 2:, ...])
+ grid_wh = bwh * torch.exp(reg[:, 2:, ...])
+ grid_left = grid_topleft[:, [0], ...]
+ grid_top = grid_topleft[:, [1], ...]
+ grid_width = grid_wh[:, [0], ...]
+ grid_height = grid_wh[:, [1], ...]
+ intervel = torch.linspace(0., 1., self.dcn_kernel).view(
+ 1, self.dcn_kernel, 1, 1).type_as(reg)
+ grid_x = grid_left + grid_width * intervel
+ grid_x = grid_x.unsqueeze(1).repeat(1, self.dcn_kernel, 1, 1, 1)
+ grid_x = grid_x.view(b, -1, h, w)
+ grid_y = grid_top + grid_height * intervel
+ grid_y = grid_y.unsqueeze(2).repeat(1, 1, self.dcn_kernel, 1, 1)
+ grid_y = grid_y.view(b, -1, h, w)
+ grid_yx = torch.stack([grid_y, grid_x], dim=2)
+ grid_yx = grid_yx.view(b, -1, h, w)
+ regressed_bbox = torch.cat([
+ grid_left, grid_top, grid_left + grid_width, grid_top + grid_height
+ ], 1)
+ return grid_yx, regressed_bbox
+
+ def forward(self, feats: Tuple[Tensor]) -> Tuple[Tensor]:
+ return multi_apply(self.forward_single, feats)
+
+ def forward_single(self, x: Tensor) -> Tuple[Tensor]:
+ """Forward feature map of a single FPN level."""
+ dcn_base_offset = self.dcn_base_offset.type_as(x)
+ # If we use center_init, the initial reppoints is from center points.
+ # If we use bounding bbox representation, the initial reppoints is
+ # from regular grid placed on a pre-defined bbox.
+ if self.use_grid_points or not self.center_init:
+ scale = self.point_base_scale / 2
+ points_init = dcn_base_offset / dcn_base_offset.max() * scale
+ bbox_init = x.new_tensor([-scale, -scale, scale,
+ scale]).view(1, 4, 1, 1)
+ else:
+ points_init = 0
+ cls_feat = x
+ pts_feat = x
+ for cls_conv in self.cls_convs:
+ cls_feat = cls_conv(cls_feat)
+ for reg_conv in self.reg_convs:
+ pts_feat = reg_conv(pts_feat)
+ # initialize reppoints
+ pts_out_init = self.reppoints_pts_init_out(
+ self.relu(self.reppoints_pts_init_conv(pts_feat)))
+ if self.use_grid_points:
+ pts_out_init, bbox_out_init = self.gen_grid_from_reg(
+ pts_out_init, bbox_init.detach())
+ else:
+ pts_out_init = pts_out_init + points_init
+ # refine and classify reppoints
+ pts_out_init_grad_mul = (1 - self.gradient_mul) * pts_out_init.detach(
+ ) + self.gradient_mul * pts_out_init
+ dcn_offset = pts_out_init_grad_mul - dcn_base_offset
+ cls_out = self.reppoints_cls_out(
+ self.relu(self.reppoints_cls_conv(cls_feat, dcn_offset)))
+ pts_out_refine = self.reppoints_pts_refine_out(
+ self.relu(self.reppoints_pts_refine_conv(pts_feat, dcn_offset)))
+ if self.use_grid_points:
+ pts_out_refine, bbox_out_refine = self.gen_grid_from_reg(
+ pts_out_refine, bbox_out_init.detach())
+ else:
+ pts_out_refine = pts_out_refine + pts_out_init.detach()
+
+ if self.training:
+ return cls_out, pts_out_init, pts_out_refine
+ else:
+ return cls_out, self.points2bbox(pts_out_refine)
+
+ def get_points(self, featmap_sizes: List[Tuple[int]],
+ batch_img_metas: List[dict], device: str) -> tuple:
+ """Get points according to feature map sizes.
+
+ Args:
+ featmap_sizes (list[tuple]): Multi-level feature map sizes.
+ batch_img_metas (list[dict]): Image meta info.
+
+ Returns:
+ tuple: points of each image, valid flags of each image
+ """
+ num_imgs = len(batch_img_metas)
+
+ # since feature map sizes of all images are the same, we only compute
+ # points center for one time
+ multi_level_points = self.prior_generator.grid_priors(
+ featmap_sizes, device=device, with_stride=True)
+ points_list = [[point.clone() for point in multi_level_points]
+ for _ in range(num_imgs)]
+
+ # for each image, we compute valid flags of multi level grids
+ valid_flag_list = []
+ for img_id, img_meta in enumerate(batch_img_metas):
+ multi_level_flags = self.prior_generator.valid_flags(
+ featmap_sizes, img_meta['pad_shape'], device=device)
+ valid_flag_list.append(multi_level_flags)
+
+ return points_list, valid_flag_list
+
+ def centers_to_bboxes(self, point_list: List[Tensor]) -> List[Tensor]:
+ """Get bboxes according to center points.
+
+ Only used in :class:`MaxIoUAssigner`.
+ """
+ bbox_list = []
+ for i_img, point in enumerate(point_list):
+ bbox = []
+ for i_lvl in range(len(self.point_strides)):
+ scale = self.point_base_scale * self.point_strides[i_lvl] * 0.5
+ bbox_shift = torch.Tensor([-scale, -scale, scale,
+ scale]).view(1, 4).type_as(point[0])
+ bbox_center = torch.cat(
+ [point[i_lvl][:, :2], point[i_lvl][:, :2]], dim=1)
+ bbox.append(bbox_center + bbox_shift)
+ bbox_list.append(bbox)
+ return bbox_list
+
+ def offset_to_pts(self, center_list: List[Tensor],
+ pred_list: List[Tensor]) -> List[Tensor]:
+ """Change from point offset to point coordinate."""
+ pts_list = []
+ for i_lvl in range(len(self.point_strides)):
+ pts_lvl = []
+ for i_img in range(len(center_list)):
+ pts_center = center_list[i_img][i_lvl][:, :2].repeat(
+ 1, self.num_points)
+ pts_shift = pred_list[i_lvl][i_img]
+ yx_pts_shift = pts_shift.permute(1, 2, 0).view(
+ -1, 2 * self.num_points)
+ y_pts_shift = yx_pts_shift[..., 0::2]
+ x_pts_shift = yx_pts_shift[..., 1::2]
+ xy_pts_shift = torch.stack([x_pts_shift, y_pts_shift], -1)
+ xy_pts_shift = xy_pts_shift.view(*yx_pts_shift.shape[:-1], -1)
+ pts = xy_pts_shift * self.point_strides[i_lvl] + pts_center
+ pts_lvl.append(pts)
+ pts_lvl = torch.stack(pts_lvl, 0)
+ pts_list.append(pts_lvl)
+ return pts_list
+
+ def _get_targets_single(self,
+ flat_proposals: Tensor,
+ valid_flags: Tensor,
+ gt_instances: InstanceData,
+ gt_instances_ignore: InstanceData,
+ stage: str = 'init',
+ unmap_outputs: bool = True) -> tuple:
+ """Compute corresponding GT box and classification targets for
+ proposals.
+
+ Args:
+ flat_proposals (Tensor): Multi level points of a image.
+ valid_flags (Tensor): Multi level valid flags of a image.
+ gt_instances (InstanceData): It usually includes ``bboxes`` and
+ ``labels`` attributes.
+ gt_instances_ignore (InstanceData): It includes ``bboxes``
+ attribute data that is ignored during training and testing.
+ stage (str): 'init' or 'refine'. Generate target for
+ init stage or refine stage. Defaults to 'init'.
+ unmap_outputs (bool): Whether to map outputs back to
+ the original set of anchors. Defaults to True.
+
+ Returns:
+ tuple:
+
+ - labels (Tensor): Labels of each level.
+ - label_weights (Tensor): Label weights of each level.
+ - bbox_targets (Tensor): BBox targets of each level.
+ - bbox_weights (Tensor): BBox weights of each level.
+ - pos_inds (Tensor): positive samples indexes.
+ - neg_inds (Tensor): negative samples indexes.
+ - sampling_result (:obj:`SamplingResult`): Sampling results.
+ """
+ inside_flags = valid_flags
+ if not inside_flags.any():
+ raise ValueError(
+ 'There is no valid proposal inside the image boundary. Please '
+ 'check the image size.')
+ # assign gt and sample proposals
+ proposals = flat_proposals[inside_flags, :]
+ pred_instances = InstanceData(priors=proposals)
+
+ if stage == 'init':
+ assigner = self.init_assigner
+ pos_weight = self.train_cfg['init']['pos_weight']
+ else:
+ assigner = self.refine_assigner
+ pos_weight = self.train_cfg['refine']['pos_weight']
+
+ assign_result = assigner.assign(pred_instances, gt_instances,
+ gt_instances_ignore)
+ sampling_result = self.sampler.sample(assign_result, pred_instances,
+ gt_instances)
+
+ num_valid_proposals = proposals.shape[0]
+ bbox_gt = proposals.new_zeros([num_valid_proposals, 4])
+ pos_proposals = torch.zeros_like(proposals)
+ proposals_weights = proposals.new_zeros([num_valid_proposals, 4])
+ labels = proposals.new_full((num_valid_proposals, ),
+ self.num_classes,
+ dtype=torch.long)
+ label_weights = proposals.new_zeros(
+ num_valid_proposals, dtype=torch.float)
+
+ pos_inds = sampling_result.pos_inds
+ neg_inds = sampling_result.neg_inds
+ if len(pos_inds) > 0:
+ bbox_gt[pos_inds, :] = sampling_result.pos_gt_bboxes
+ pos_proposals[pos_inds, :] = proposals[pos_inds, :]
+ proposals_weights[pos_inds, :] = 1.0
+
+ labels[pos_inds] = sampling_result.pos_gt_labels
+ if pos_weight <= 0:
+ label_weights[pos_inds] = 1.0
+ else:
+ label_weights[pos_inds] = pos_weight
+ if len(neg_inds) > 0:
+ label_weights[neg_inds] = 1.0
+
+ # map up to original set of proposals
+ if unmap_outputs:
+ num_total_proposals = flat_proposals.size(0)
+ labels = unmap(
+ labels,
+ num_total_proposals,
+ inside_flags,
+ fill=self.num_classes) # fill bg label
+ label_weights = unmap(label_weights, num_total_proposals,
+ inside_flags)
+ bbox_gt = unmap(bbox_gt, num_total_proposals, inside_flags)
+ pos_proposals = unmap(pos_proposals, num_total_proposals,
+ inside_flags)
+ proposals_weights = unmap(proposals_weights, num_total_proposals,
+ inside_flags)
+
+ return (labels, label_weights, bbox_gt, pos_proposals,
+ proposals_weights, pos_inds, neg_inds, sampling_result)
+
+ def get_targets(self,
+ proposals_list: List[Tensor],
+ valid_flag_list: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None,
+ stage: str = 'init',
+ unmap_outputs: bool = True,
+ return_sampling_results: bool = False) -> tuple:
+ """Compute corresponding GT box and classification targets for
+ proposals.
+
+ Args:
+ proposals_list (list[Tensor]): Multi level points/bboxes of each
+ image.
+ valid_flag_list (list[Tensor]): Multi level valid flags of each
+ image.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ stage (str): 'init' or 'refine'. Generate target for init stage or
+ refine stage.
+ unmap_outputs (bool): Whether to map outputs back to the original
+ set of anchors.
+ return_sampling_results (bool): Whether to return the sampling
+ results. Defaults to False.
+
+ Returns:
+ tuple:
+
+ - labels_list (list[Tensor]): Labels of each level.
+ - label_weights_list (list[Tensor]): Label weights of each
+ level.
+ - bbox_gt_list (list[Tensor]): Ground truth bbox of each level.
+ - proposals_list (list[Tensor]): Proposals(points/bboxes) of
+ each level.
+ - proposal_weights_list (list[Tensor]): Proposal weights of
+ each level.
+ - avg_factor (int): Average factor that is used to average
+ the loss. When using sampling method, avg_factor is usually
+ the sum of positive and negative priors. When using
+ `PseudoSampler`, `avg_factor` is usually equal to the number
+ of positive priors.
+ """
+ assert stage in ['init', 'refine']
+ num_imgs = len(batch_img_metas)
+ assert len(proposals_list) == len(valid_flag_list) == num_imgs
+
+ # points number of multi levels
+ num_level_proposals = [points.size(0) for points in proposals_list[0]]
+
+ # concat all level points and flags to a single tensor
+ for i in range(num_imgs):
+ assert len(proposals_list[i]) == len(valid_flag_list[i])
+ proposals_list[i] = torch.cat(proposals_list[i])
+ valid_flag_list[i] = torch.cat(valid_flag_list[i])
+
+ if batch_gt_instances_ignore is None:
+ batch_gt_instances_ignore = [None] * num_imgs
+
+ (all_labels, all_label_weights, all_bbox_gt, all_proposals,
+ all_proposal_weights, pos_inds_list, neg_inds_list,
+ sampling_results_list) = multi_apply(
+ self._get_targets_single,
+ proposals_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_gt_instances_ignore,
+ stage=stage,
+ unmap_outputs=unmap_outputs)
+
+ # sampled points of all images
+ avg_refactor = sum(
+ [results.avg_factor for results in sampling_results_list])
+ labels_list = images_to_levels(all_labels, num_level_proposals)
+ label_weights_list = images_to_levels(all_label_weights,
+ num_level_proposals)
+ bbox_gt_list = images_to_levels(all_bbox_gt, num_level_proposals)
+ proposals_list = images_to_levels(all_proposals, num_level_proposals)
+ proposal_weights_list = images_to_levels(all_proposal_weights,
+ num_level_proposals)
+ res = (labels_list, label_weights_list, bbox_gt_list, proposals_list,
+ proposal_weights_list, avg_refactor)
+ if return_sampling_results:
+ res = res + (sampling_results_list, )
+
+ return res
+
+ def loss_by_feat_single(self, cls_score: Tensor, pts_pred_init: Tensor,
+ pts_pred_refine: Tensor, labels: Tensor,
+ label_weights, bbox_gt_init: Tensor,
+ bbox_weights_init: Tensor, bbox_gt_refine: Tensor,
+ bbox_weights_refine: Tensor, stride: int,
+ avg_factor_init: int,
+ avg_factor_refine: int) -> Tuple[Tensor]:
+ """Calculate the loss of a single scale level based on the features
+ extracted by the detection head.
+
+ Args:
+ cls_score (Tensor): Box scores for each scale level
+ Has shape (N, num_classes, h_i, w_i).
+ pts_pred_init (Tensor): Points of shape
+ (batch_size, h_i * w_i, num_points * 2).
+ pts_pred_refine (Tensor): Points refined of shape
+ (batch_size, h_i * w_i, num_points * 2).
+ labels (Tensor): Ground truth class indices with shape
+ (batch_size, h_i * w_i).
+ label_weights (Tensor): Label weights of shape
+ (batch_size, h_i * w_i).
+ bbox_gt_init (Tensor): BBox regression targets in the init stage
+ of shape (batch_size, h_i * w_i, 4).
+ bbox_weights_init (Tensor): BBox regression loss weights in the
+ init stage of shape (batch_size, h_i * w_i, 4).
+ bbox_gt_refine (Tensor): BBox regression targets in the refine
+ stage of shape (batch_size, h_i * w_i, 4).
+ bbox_weights_refine (Tensor): BBox regression loss weights in the
+ refine stage of shape (batch_size, h_i * w_i, 4).
+ stride (int): Point stride.
+ avg_factor_init (int): Average factor that is used to average
+ the loss in the init stage.
+ avg_factor_refine (int): Average factor that is used to average
+ the loss in the refine stage.
+
+ Returns:
+ Tuple[Tensor]: loss components.
+ """
+ # classification loss
+ labels = labels.reshape(-1)
+ label_weights = label_weights.reshape(-1)
+ cls_score = cls_score.permute(0, 2, 3,
+ 1).reshape(-1, self.cls_out_channels)
+ cls_score = cls_score.contiguous()
+ loss_cls = self.loss_cls(
+ cls_score, labels, label_weights, avg_factor=avg_factor_refine)
+
+ # points loss
+ bbox_gt_init = bbox_gt_init.reshape(-1, 4)
+ bbox_weights_init = bbox_weights_init.reshape(-1, 4)
+ bbox_pred_init = self.points2bbox(
+ pts_pred_init.reshape(-1, 2 * self.num_points), y_first=False)
+ bbox_gt_refine = bbox_gt_refine.reshape(-1, 4)
+ bbox_weights_refine = bbox_weights_refine.reshape(-1, 4)
+ bbox_pred_refine = self.points2bbox(
+ pts_pred_refine.reshape(-1, 2 * self.num_points), y_first=False)
+ normalize_term = self.point_base_scale * stride
+ loss_pts_init = self.loss_bbox_init(
+ bbox_pred_init / normalize_term,
+ bbox_gt_init / normalize_term,
+ bbox_weights_init,
+ avg_factor=avg_factor_init)
+ loss_pts_refine = self.loss_bbox_refine(
+ bbox_pred_refine / normalize_term,
+ bbox_gt_refine / normalize_term,
+ bbox_weights_refine,
+ avg_factor=avg_factor_refine)
+ return loss_cls, loss_pts_init, loss_pts_refine
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ pts_preds_init: List[Tensor],
+ pts_preds_refine: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None
+ ) -> Dict[str, Tensor]:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level,
+ each is a 4D-tensor, of shape (batch_size, num_classes, h, w).
+ pts_preds_init (list[Tensor]): Points for each scale level, each is
+ a 3D-tensor, of shape (batch_size, h_i * w_i, num_points * 2).
+ pts_preds_refine (list[Tensor]): Points refined for each scale
+ level, each is a 3D-tensor, of shape
+ (batch_size, h_i * w_i, num_points * 2).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ device = cls_scores[0].device
+
+ # target for initial stage
+ center_list, valid_flag_list = self.get_points(featmap_sizes,
+ batch_img_metas, device)
+ pts_coordinate_preds_init = self.offset_to_pts(center_list,
+ pts_preds_init)
+ if self.train_cfg['init']['assigner']['type'] == 'PointAssigner':
+ # Assign target for center list
+ candidate_list = center_list
+ else:
+ # transform center list to bbox list and
+ # assign target for bbox list
+ bbox_list = self.centers_to_bboxes(center_list)
+ candidate_list = bbox_list
+ cls_reg_targets_init = self.get_targets(
+ proposals_list=candidate_list,
+ valid_flag_list=valid_flag_list,
+ batch_gt_instances=batch_gt_instances,
+ batch_img_metas=batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore,
+ stage='init',
+ return_sampling_results=False)
+ (*_, bbox_gt_list_init, candidate_list_init, bbox_weights_list_init,
+ avg_factor_init) = cls_reg_targets_init
+
+ # target for refinement stage
+ center_list, valid_flag_list = self.get_points(featmap_sizes,
+ batch_img_metas, device)
+ pts_coordinate_preds_refine = self.offset_to_pts(
+ center_list, pts_preds_refine)
+ bbox_list = []
+ for i_img, center in enumerate(center_list):
+ bbox = []
+ for i_lvl in range(len(pts_preds_refine)):
+ bbox_preds_init = self.points2bbox(
+ pts_preds_init[i_lvl].detach())
+ bbox_shift = bbox_preds_init * self.point_strides[i_lvl]
+ bbox_center = torch.cat(
+ [center[i_lvl][:, :2], center[i_lvl][:, :2]], dim=1)
+ bbox.append(bbox_center +
+ bbox_shift[i_img].permute(1, 2, 0).reshape(-1, 4))
+ bbox_list.append(bbox)
+ cls_reg_targets_refine = self.get_targets(
+ proposals_list=bbox_list,
+ valid_flag_list=valid_flag_list,
+ batch_gt_instances=batch_gt_instances,
+ batch_img_metas=batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore,
+ stage='refine',
+ return_sampling_results=False)
+ (labels_list, label_weights_list, bbox_gt_list_refine,
+ candidate_list_refine, bbox_weights_list_refine,
+ avg_factor_refine) = cls_reg_targets_refine
+
+ # compute loss
+ losses_cls, losses_pts_init, losses_pts_refine = multi_apply(
+ self.loss_by_feat_single,
+ cls_scores,
+ pts_coordinate_preds_init,
+ pts_coordinate_preds_refine,
+ labels_list,
+ label_weights_list,
+ bbox_gt_list_init,
+ bbox_weights_list_init,
+ bbox_gt_list_refine,
+ bbox_weights_list_refine,
+ self.point_strides,
+ avg_factor_init=avg_factor_init,
+ avg_factor_refine=avg_factor_refine)
+ loss_dict_all = {
+ 'loss_cls': losses_cls,
+ 'loss_pts_init': losses_pts_init,
+ 'loss_pts_refine': losses_pts_refine
+ }
+ return loss_dict_all
+
+ # Same as base_dense_head/_get_bboxes_single except self._bbox_decode
+ def _predict_by_feat_single(self,
+ cls_score_list: List[Tensor],
+ bbox_pred_list: List[Tensor],
+ score_factor_list: List[Tensor],
+ mlvl_priors: List[Tensor],
+ img_meta: dict,
+ cfg: ConfigDict,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceData:
+ """Transform outputs of a single image into bbox predictions.
+
+ Args:
+ cls_score_list (list[Tensor]): Box scores from all scale
+ levels of a single image, each item has shape
+ (num_priors * num_classes, H, W).
+ bbox_pred_list (list[Tensor]): Box energies / deltas from
+ all scale levels of a single image, each item has shape
+ (num_priors * 4, H, W).
+ score_factor_list (list[Tensor]): Score factor from all scale
+ levels of a single image. RepPoints head does not need
+ this value.
+ mlvl_priors (list[Tensor]): Each element in the list is
+ the priors of a single level in feature pyramid, has shape
+ (num_priors, 2).
+ img_meta (dict): Image meta info.
+ cfg (:obj:`ConfigDict`): Test / postprocessing configuration,
+ if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ cfg = self.test_cfg if cfg is None else cfg
+ assert len(cls_score_list) == len(bbox_pred_list)
+ img_shape = img_meta['img_shape']
+ nms_pre = cfg.get('nms_pre', -1)
+
+ mlvl_bboxes = []
+ mlvl_scores = []
+ mlvl_labels = []
+ for level_idx, (cls_score, bbox_pred, priors) in enumerate(
+ zip(cls_score_list, bbox_pred_list, mlvl_priors)):
+ assert cls_score.size()[-2:] == bbox_pred.size()[-2:]
+ bbox_pred = bbox_pred.permute(1, 2, 0).reshape(-1, 4)
+
+ cls_score = cls_score.permute(1, 2,
+ 0).reshape(-1, self.cls_out_channels)
+ if self.use_sigmoid_cls:
+ scores = cls_score.sigmoid()
+ else:
+ scores = cls_score.softmax(-1)[:, :-1]
+
+ # After https://github.com/open-mmlab/mmdetection/pull/6268/,
+ # this operation keeps fewer bboxes under the same `nms_pre`.
+ # There is no difference in performance for most models. If you
+ # find a slight drop in performance, you can set a larger
+ # `nms_pre` than before.
+ results = filter_scores_and_topk(
+ scores, cfg.score_thr, nms_pre,
+ dict(bbox_pred=bbox_pred, priors=priors))
+ scores, labels, _, filtered_results = results
+
+ bbox_pred = filtered_results['bbox_pred']
+ priors = filtered_results['priors']
+
+ bboxes = self._bbox_decode(priors, bbox_pred,
+ self.point_strides[level_idx],
+ img_shape)
+
+ mlvl_bboxes.append(bboxes)
+ mlvl_scores.append(scores)
+ mlvl_labels.append(labels)
+
+ results = InstanceData()
+ results.bboxes = torch.cat(mlvl_bboxes)
+ results.scores = torch.cat(mlvl_scores)
+ results.labels = torch.cat(mlvl_labels)
+
+ return self._bbox_post_process(
+ results=results,
+ cfg=cfg,
+ rescale=rescale,
+ with_nms=with_nms,
+ img_meta=img_meta)
+
+ def _bbox_decode(self, points: Tensor, bbox_pred: Tensor, stride: int,
+ max_shape: Tuple[int, int]) -> Tensor:
+ """Decode the prediction to bounding box.
+
+ Args:
+ points (Tensor): shape (h_i * w_i, 2).
+ bbox_pred (Tensor): shape (h_i * w_i, 4).
+ stride (int): Stride for bbox_pred in different level.
+ max_shape (Tuple[int, int]): image shape.
+
+ Returns:
+ Tensor: Bounding boxes decoded.
+ """
+ bbox_pos_center = torch.cat([points[:, :2], points[:, :2]], dim=1)
+ bboxes = bbox_pred * stride + bbox_pos_center
+ x1 = bboxes[:, 0].clamp(min=0, max=max_shape[1])
+ y1 = bboxes[:, 1].clamp(min=0, max=max_shape[0])
+ x2 = bboxes[:, 2].clamp(min=0, max=max_shape[1])
+ y2 = bboxes[:, 3].clamp(min=0, max=max_shape[0])
+ decoded_bboxes = torch.stack([x1, y1, x2, y2], dim=-1)
+ return decoded_bboxes
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/retina_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/retina_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..be3ae74d81ba38609646f0d0406098ecbdcef688
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/retina_head.py
@@ -0,0 +1,120 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+
+from mmdet.registry import MODELS
+from .anchor_head import AnchorHead
+
+
+@MODELS.register_module()
+class RetinaHead(AnchorHead):
+ r"""An anchor-based head used in `RetinaNet
+ `_.
+
+ The head contains two subnetworks. The first classifies anchor boxes and
+ the second regresses deltas for the anchors.
+
+ Example:
+ >>> import torch
+ >>> self = RetinaHead(11, 7)
+ >>> x = torch.rand(1, 7, 32, 32)
+ >>> cls_score, bbox_pred = self.forward_single(x)
+ >>> # Each anchor predicts a score for each class except background
+ >>> cls_per_anchor = cls_score.shape[1] / self.num_anchors
+ >>> box_per_anchor = bbox_pred.shape[1] / self.num_anchors
+ >>> assert cls_per_anchor == (self.num_classes)
+ >>> assert box_per_anchor == 4
+ """
+
+ def __init__(self,
+ num_classes,
+ in_channels,
+ stacked_convs=4,
+ conv_cfg=None,
+ norm_cfg=None,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ octave_base_scale=4,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[8, 16, 32, 64, 128]),
+ init_cfg=dict(
+ type='Normal',
+ layer='Conv2d',
+ std=0.01,
+ override=dict(
+ type='Normal',
+ name='retina_cls',
+ std=0.01,
+ bias_prob=0.01)),
+ **kwargs):
+ assert stacked_convs >= 0, \
+ '`stacked_convs` must be non-negative integers, ' \
+ f'but got {stacked_convs} instead.'
+ self.stacked_convs = stacked_convs
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ super(RetinaHead, self).__init__(
+ num_classes,
+ in_channels,
+ anchor_generator=anchor_generator,
+ init_cfg=init_cfg,
+ **kwargs)
+
+ def _init_layers(self):
+ """Initialize layers of the head."""
+ self.relu = nn.ReLU(inplace=True)
+ self.cls_convs = nn.ModuleList()
+ self.reg_convs = nn.ModuleList()
+ in_channels = self.in_channels
+ for i in range(self.stacked_convs):
+ self.cls_convs.append(
+ ConvModule(
+ in_channels,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ self.reg_convs.append(
+ ConvModule(
+ in_channels,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ in_channels = self.feat_channels
+ self.retina_cls = nn.Conv2d(
+ in_channels,
+ self.num_base_priors * self.cls_out_channels,
+ 3,
+ padding=1)
+ reg_dim = self.bbox_coder.encode_size
+ self.retina_reg = nn.Conv2d(
+ in_channels, self.num_base_priors * reg_dim, 3, padding=1)
+
+ def forward_single(self, x):
+ """Forward feature of a single scale level.
+
+ Args:
+ x (Tensor): Features of a single scale level.
+
+ Returns:
+ tuple:
+ cls_score (Tensor): Cls scores for a single scale level
+ the channels number is num_anchors * num_classes.
+ bbox_pred (Tensor): Box energies / deltas for a single scale
+ level, the channels number is num_anchors * 4.
+ """
+ cls_feat = x
+ reg_feat = x
+ for cls_conv in self.cls_convs:
+ cls_feat = cls_conv(cls_feat)
+ for reg_conv in self.reg_convs:
+ reg_feat = reg_conv(reg_feat)
+ cls_score = self.retina_cls(cls_feat)
+ bbox_pred = self.retina_reg(reg_feat)
+ return cls_score, bbox_pred
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/retina_sepbn_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/retina_sepbn_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..681a39983a08670adaa3e24a4099c4f26bc967ce
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/retina_sepbn_head.py
@@ -0,0 +1,127 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Tuple
+
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from mmengine.model import bias_init_with_prob, normal_init
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import OptConfigType, OptMultiConfig
+from .anchor_head import AnchorHead
+
+
+@MODELS.register_module()
+class RetinaSepBNHead(AnchorHead):
+ """"RetinaHead with separate BN.
+
+ In RetinaHead, conv/norm layers are shared across different FPN levels,
+ while in RetinaSepBNHead, conv layers are shared across different FPN
+ levels, but BN layers are separated.
+ """
+
+ def __init__(self,
+ num_classes: int,
+ num_ins: int,
+ in_channels: int,
+ stacked_convs: int = 4,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: OptConfigType = None,
+ init_cfg: OptMultiConfig = None,
+ **kwargs) -> None:
+ assert init_cfg is None, 'To prevent abnormal initialization ' \
+ 'behavior, init_cfg is not allowed to be set'
+ self.stacked_convs = stacked_convs
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self.num_ins = num_ins
+ super().__init__(
+ num_classes=num_classes,
+ in_channels=in_channels,
+ init_cfg=init_cfg,
+ **kwargs)
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self.relu = nn.ReLU(inplace=True)
+ self.cls_convs = nn.ModuleList()
+ self.reg_convs = nn.ModuleList()
+ for i in range(self.num_ins):
+ cls_convs = nn.ModuleList()
+ reg_convs = nn.ModuleList()
+ for j in range(self.stacked_convs):
+ chn = self.in_channels if j == 0 else self.feat_channels
+ cls_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ reg_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ self.cls_convs.append(cls_convs)
+ self.reg_convs.append(reg_convs)
+ for i in range(self.stacked_convs):
+ for j in range(1, self.num_ins):
+ self.cls_convs[j][i].conv = self.cls_convs[0][i].conv
+ self.reg_convs[j][i].conv = self.reg_convs[0][i].conv
+ self.retina_cls = nn.Conv2d(
+ self.feat_channels,
+ self.num_base_priors * self.cls_out_channels,
+ 3,
+ padding=1)
+ self.retina_reg = nn.Conv2d(
+ self.feat_channels, self.num_base_priors * 4, 3, padding=1)
+
+ def init_weights(self) -> None:
+ """Initialize weights of the head."""
+ super().init_weights()
+ for m in self.cls_convs[0]:
+ normal_init(m.conv, std=0.01)
+ for m in self.reg_convs[0]:
+ normal_init(m.conv, std=0.01)
+ bias_cls = bias_init_with_prob(0.01)
+ normal_init(self.retina_cls, std=0.01, bias=bias_cls)
+ normal_init(self.retina_reg, std=0.01)
+
+ def forward(self, feats: Tuple[Tensor]) -> tuple:
+ """Forward features from the upstream network.
+
+ Args:
+ feats (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: Usually a tuple of classification scores and bbox prediction
+
+ - cls_scores (list[Tensor]): Classification scores for all
+ scale levels, each is a 4D-tensor, the channels number is
+ num_anchors * num_classes.
+ - bbox_preds (list[Tensor]): Box energies / deltas for all
+ scale levels, each is a 4D-tensor, the channels number is
+ num_anchors * 4.
+ """
+ cls_scores = []
+ bbox_preds = []
+ for i, x in enumerate(feats):
+ cls_feat = feats[i]
+ reg_feat = feats[i]
+ for cls_conv in self.cls_convs[i]:
+ cls_feat = cls_conv(cls_feat)
+ for reg_conv in self.reg_convs[i]:
+ reg_feat = reg_conv(reg_feat)
+ cls_score = self.retina_cls(cls_feat)
+ bbox_pred = self.retina_reg(reg_feat)
+ cls_scores.append(cls_score)
+ bbox_preds.append(bbox_pred)
+ return cls_scores, bbox_preds
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/rpn_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/rpn_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..6b544009d2ffc4c3c9065707a0a8a72c577eb432
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/rpn_head.py
@@ -0,0 +1,302 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from typing import List, Optional, Tuple
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule
+from mmcv.ops import batched_nms
+from mmengine.config import ConfigDict
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures.bbox import (cat_boxes, empty_box_as, get_box_tensor,
+ get_box_wh, scale_boxes)
+from mmdet.utils import InstanceList, MultiConfig, OptInstanceList
+from .anchor_head import AnchorHead
+
+
+@MODELS.register_module()
+class RPNHead(AnchorHead):
+ """Implementation of RPN head.
+
+ Args:
+ in_channels (int): Number of channels in the input feature map.
+ num_classes (int): Number of categories excluding the background
+ category. Defaults to 1.
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or \
+ list[dict]): Initialization config dict.
+ num_convs (int): Number of convolution layers in the head.
+ Defaults to 1.
+ """ # noqa: W605
+
+ def __init__(self,
+ in_channels: int,
+ num_classes: int = 1,
+ init_cfg: MultiConfig = dict(
+ type='Normal', layer='Conv2d', std=0.01),
+ num_convs: int = 1,
+ **kwargs) -> None:
+ self.num_convs = num_convs
+ assert num_classes == 1
+ super().__init__(
+ num_classes=num_classes,
+ in_channels=in_channels,
+ init_cfg=init_cfg,
+ **kwargs)
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ if self.num_convs > 1:
+ rpn_convs = []
+ for i in range(self.num_convs):
+ if i == 0:
+ in_channels = self.in_channels
+ else:
+ in_channels = self.feat_channels
+ # use ``inplace=False`` to avoid error: one of the variables
+ # needed for gradient computation has been modified by an
+ # inplace operation.
+ rpn_convs.append(
+ ConvModule(
+ in_channels,
+ self.feat_channels,
+ 3,
+ padding=1,
+ inplace=False))
+ self.rpn_conv = nn.Sequential(*rpn_convs)
+ else:
+ self.rpn_conv = nn.Conv2d(
+ self.in_channels, self.feat_channels, 3, padding=1)
+ self.rpn_cls = nn.Conv2d(self.feat_channels,
+ self.num_base_priors * self.cls_out_channels,
+ 1)
+ reg_dim = self.bbox_coder.encode_size
+ self.rpn_reg = nn.Conv2d(self.feat_channels,
+ self.num_base_priors * reg_dim, 1)
+
+ def forward_single(self, x: Tensor) -> Tuple[Tensor, Tensor]:
+ """Forward feature of a single scale level.
+
+ Args:
+ x (Tensor): Features of a single scale level.
+
+ Returns:
+ tuple:
+ cls_score (Tensor): Cls scores for a single scale level \
+ the channels number is num_base_priors * num_classes.
+ bbox_pred (Tensor): Box energies / deltas for a single scale \
+ level, the channels number is num_base_priors * 4.
+ """
+ x = self.rpn_conv(x)
+ x = F.relu(x)
+ rpn_cls_score = self.rpn_cls(x)
+ rpn_bbox_pred = self.rpn_reg(x)
+ return rpn_cls_score, rpn_bbox_pred
+
+ def loss_by_feat(self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) \
+ -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level,
+ has shape (N, num_anchors * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W).
+ batch_gt_instances (list[obj:InstanceData]): Batch of gt_instance.
+ It usually includes ``bboxes`` and ``labels`` attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[obj:InstanceData], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ losses = super().loss_by_feat(
+ cls_scores,
+ bbox_preds,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore)
+ return dict(
+ loss_rpn_cls=losses['loss_cls'], loss_rpn_bbox=losses['loss_bbox'])
+
+ def _predict_by_feat_single(self,
+ cls_score_list: List[Tensor],
+ bbox_pred_list: List[Tensor],
+ score_factor_list: List[Tensor],
+ mlvl_priors: List[Tensor],
+ img_meta: dict,
+ cfg: ConfigDict,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results.
+
+ Args:
+ cls_score_list (list[Tensor]): Box scores from all scale
+ levels of a single image, each item has shape
+ (num_priors * num_classes, H, W).
+ bbox_pred_list (list[Tensor]): Box energies / deltas from
+ all scale levels of a single image, each item has shape
+ (num_priors * 4, H, W).
+ score_factor_list (list[Tensor]): Be compatible with
+ BaseDenseHead. Not used in RPNHead.
+ mlvl_priors (list[Tensor]): Each element in the list is
+ the priors of a single level in feature pyramid. In all
+ anchor-based methods, it has shape (num_priors, 4). In
+ all anchor-free methods, it has shape (num_priors, 2)
+ when `with_stride=True`, otherwise it still has shape
+ (num_priors, 4).
+ img_meta (dict): Image meta info.
+ cfg (ConfigDict, optional): Test / postprocessing configuration,
+ if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ cfg = self.test_cfg if cfg is None else cfg
+ cfg = copy.deepcopy(cfg)
+ img_shape = img_meta['img_shape']
+ nms_pre = cfg.get('nms_pre', -1)
+
+ mlvl_bbox_preds = []
+ mlvl_valid_priors = []
+ mlvl_scores = []
+ level_ids = []
+ for level_idx, (cls_score, bbox_pred, priors) in \
+ enumerate(zip(cls_score_list, bbox_pred_list,
+ mlvl_priors)):
+ assert cls_score.size()[-2:] == bbox_pred.size()[-2:]
+
+ reg_dim = self.bbox_coder.encode_size
+ bbox_pred = bbox_pred.permute(1, 2, 0).reshape(-1, reg_dim)
+ cls_score = cls_score.permute(1, 2,
+ 0).reshape(-1, self.cls_out_channels)
+ if self.use_sigmoid_cls:
+ scores = cls_score.sigmoid()
+ else:
+ # remind that we set FG labels to [0] since mmdet v2.0
+ # BG cat_id: 1
+ scores = cls_score.softmax(-1)[:, :-1]
+
+ scores = torch.squeeze(scores)
+ if 0 < nms_pre < scores.shape[0]:
+ # sort is faster than topk
+ # _, topk_inds = scores.topk(cfg.nms_pre)
+ ranked_scores, rank_inds = scores.sort(descending=True)
+ topk_inds = rank_inds[:nms_pre]
+ scores = ranked_scores[:nms_pre]
+ bbox_pred = bbox_pred[topk_inds, :]
+ priors = priors[topk_inds]
+
+ mlvl_bbox_preds.append(bbox_pred)
+ mlvl_valid_priors.append(priors)
+ mlvl_scores.append(scores)
+
+ # use level id to implement the separate level nms
+ level_ids.append(
+ scores.new_full((scores.size(0), ),
+ level_idx,
+ dtype=torch.long))
+
+ bbox_pred = torch.cat(mlvl_bbox_preds)
+ priors = cat_boxes(mlvl_valid_priors)
+ bboxes = self.bbox_coder.decode(priors, bbox_pred, max_shape=img_shape)
+
+ results = InstanceData()
+ results.bboxes = bboxes
+ results.scores = torch.cat(mlvl_scores)
+ results.level_ids = torch.cat(level_ids)
+
+ return self._bbox_post_process(
+ results=results, cfg=cfg, rescale=rescale, img_meta=img_meta)
+
+ def _bbox_post_process(self,
+ results: InstanceData,
+ cfg: ConfigDict,
+ rescale: bool = False,
+ with_nms: bool = True,
+ img_meta: Optional[dict] = None) -> InstanceData:
+ """bbox post-processing method.
+
+ The boxes would be rescaled to the original image scale and do
+ the nms operation.
+
+ Args:
+ results (:obj:`InstaceData`): Detection instance results,
+ each item has shape (num_bboxes, ).
+ cfg (ConfigDict): Test / postprocessing configuration.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Default to True.
+ img_meta (dict, optional): Image meta info. Defaults to None.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ assert with_nms, '`with_nms` must be True in RPNHead'
+ if rescale:
+ assert img_meta.get('scale_factor') is not None
+ scale_factor = [1 / s for s in img_meta['scale_factor']]
+ results.bboxes = scale_boxes(results.bboxes, scale_factor)
+
+ # filter small size bboxes
+ if cfg.get('min_bbox_size', -1) >= 0:
+ w, h = get_box_wh(results.bboxes)
+ valid_mask = (w > cfg.min_bbox_size) & (h > cfg.min_bbox_size)
+ if not valid_mask.all():
+ results = results[valid_mask]
+
+ if results.bboxes.numel() > 0:
+ bboxes = get_box_tensor(results.bboxes)
+ det_bboxes, keep_idxs = batched_nms(bboxes, results.scores,
+ results.level_ids, cfg.nms)
+ results = results[keep_idxs]
+ # some nms would reweight the score, such as softnms
+ results.scores = det_bboxes[:, -1]
+ results = results[:cfg.max_per_img]
+ # TODO: This would unreasonably show the 0th class label
+ # in visualization
+ results.labels = results.scores.new_zeros(
+ len(results), dtype=torch.long)
+ del results.level_ids
+ else:
+ # To avoid some potential error
+ results_ = InstanceData()
+ results_.bboxes = empty_box_as(results.bboxes)
+ results_.scores = results.scores.new_zeros(0)
+ results_.labels = results.scores.new_zeros(0)
+ results = results_
+ return results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/rtmdet_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/rtmdet_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..ae0ee6d2f35a0fa46ba0b8de21054433d0420b65
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/rtmdet_head.py
@@ -0,0 +1,692 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple, Union
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule, DepthwiseSeparableConvModule, Scale, is_norm
+from mmengine.model import bias_init_with_prob, constant_init, normal_init
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures.bbox import distance2bbox
+from mmdet.utils import ConfigType, InstanceList, OptInstanceList, reduce_mean
+from ..layers.transformer import inverse_sigmoid
+from ..task_modules import anchor_inside_flags
+from ..utils import (images_to_levels, multi_apply, sigmoid_geometric_mean,
+ unmap)
+from .atss_head import ATSSHead
+
+
+@MODELS.register_module()
+class RTMDetHead(ATSSHead):
+ """Detection Head of RTMDet.
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ with_objectness (bool): Whether to add an objectness branch.
+ Defaults to True.
+ act_cfg (:obj:`ConfigDict` or dict): Config dict for activation layer.
+ Default: dict(type='ReLU')
+ """
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: int,
+ with_objectness: bool = True,
+ act_cfg: ConfigType = dict(type='ReLU'),
+ **kwargs) -> None:
+ self.act_cfg = act_cfg
+ self.with_objectness = with_objectness
+ super().__init__(num_classes, in_channels, **kwargs)
+ if self.train_cfg:
+ self.assigner = TASK_UTILS.build(self.train_cfg['assigner'])
+
+ def _init_layers(self):
+ """Initialize layers of the head."""
+ self.cls_convs = nn.ModuleList()
+ self.reg_convs = nn.ModuleList()
+ for i in range(self.stacked_convs):
+ chn = self.in_channels if i == 0 else self.feat_channels
+ self.cls_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg))
+ self.reg_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg))
+ pred_pad_size = self.pred_kernel_size // 2
+ self.rtm_cls = nn.Conv2d(
+ self.feat_channels,
+ self.num_base_priors * self.cls_out_channels,
+ self.pred_kernel_size,
+ padding=pred_pad_size)
+ self.rtm_reg = nn.Conv2d(
+ self.feat_channels,
+ self.num_base_priors * 4,
+ self.pred_kernel_size,
+ padding=pred_pad_size)
+ if self.with_objectness:
+ self.rtm_obj = nn.Conv2d(
+ self.feat_channels,
+ 1,
+ self.pred_kernel_size,
+ padding=pred_pad_size)
+
+ self.scales = nn.ModuleList(
+ [Scale(1.0) for _ in self.prior_generator.strides])
+
+ def init_weights(self) -> None:
+ """Initialize weights of the head."""
+ for m in self.modules():
+ if isinstance(m, nn.Conv2d):
+ normal_init(m, mean=0, std=0.01)
+ if is_norm(m):
+ constant_init(m, 1)
+ bias_cls = bias_init_with_prob(0.01)
+ normal_init(self.rtm_cls, std=0.01, bias=bias_cls)
+ normal_init(self.rtm_reg, std=0.01)
+ if self.with_objectness:
+ normal_init(self.rtm_obj, std=0.01, bias=bias_cls)
+
+ def forward(self, feats: Tuple[Tensor, ...]) -> tuple:
+ """Forward features from the upstream network.
+
+ Args:
+ feats (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: Usually a tuple of classification scores and bbox prediction
+ - cls_scores (list[Tensor]): Classification scores for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_base_priors * num_classes.
+ - bbox_preds (list[Tensor]): Box energies / deltas for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_base_priors * 4.
+ """
+
+ cls_scores = []
+ bbox_preds = []
+ for idx, (x, scale, stride) in enumerate(
+ zip(feats, self.scales, self.prior_generator.strides)):
+ cls_feat = x
+ reg_feat = x
+
+ for cls_layer in self.cls_convs:
+ cls_feat = cls_layer(cls_feat)
+ cls_score = self.rtm_cls(cls_feat)
+
+ for reg_layer in self.reg_convs:
+ reg_feat = reg_layer(reg_feat)
+
+ if self.with_objectness:
+ objectness = self.rtm_obj(reg_feat)
+ cls_score = inverse_sigmoid(
+ sigmoid_geometric_mean(cls_score, objectness))
+
+ reg_dist = scale(self.rtm_reg(reg_feat).exp()).float() * stride[0]
+
+ cls_scores.append(cls_score)
+ bbox_preds.append(reg_dist)
+ return tuple(cls_scores), tuple(bbox_preds)
+
+ def loss_by_feat_single(self, cls_score: Tensor, bbox_pred: Tensor,
+ labels: Tensor, label_weights: Tensor,
+ bbox_targets: Tensor, assign_metrics: Tensor,
+ stride: List[int]):
+ """Compute loss of a single scale level.
+
+ Args:
+ cls_score (Tensor): Box scores for each scale level
+ Has shape (N, num_anchors * num_classes, H, W).
+ bbox_pred (Tensor): Decoded bboxes for each scale
+ level with shape (N, num_anchors * 4, H, W).
+ labels (Tensor): Labels of each anchors with shape
+ (N, num_total_anchors).
+ label_weights (Tensor): Label weights of each anchor with shape
+ (N, num_total_anchors).
+ bbox_targets (Tensor): BBox regression targets of each anchor with
+ shape (N, num_total_anchors, 4).
+ assign_metrics (Tensor): Assign metrics with shape
+ (N, num_total_anchors).
+ stride (List[int]): Downsample stride of the feature map.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ assert stride[0] == stride[1], 'h stride is not equal to w stride!'
+ cls_score = cls_score.permute(0, 2, 3, 1).reshape(
+ -1, self.cls_out_channels).contiguous()
+ bbox_pred = bbox_pred.reshape(-1, 4)
+ bbox_targets = bbox_targets.reshape(-1, 4)
+ labels = labels.reshape(-1)
+ assign_metrics = assign_metrics.reshape(-1)
+ label_weights = label_weights.reshape(-1)
+ targets = (labels, assign_metrics)
+
+ loss_cls = self.loss_cls(
+ cls_score, targets, label_weights, avg_factor=1.0)
+
+ # FG cat_id: [0, num_classes -1], BG cat_id: num_classes
+ bg_class_ind = self.num_classes
+ pos_inds = ((labels >= 0)
+ & (labels < bg_class_ind)).nonzero().squeeze(1)
+
+ if len(pos_inds) > 0:
+ pos_bbox_targets = bbox_targets[pos_inds]
+ pos_bbox_pred = bbox_pred[pos_inds]
+
+ pos_decode_bbox_pred = pos_bbox_pred
+ pos_decode_bbox_targets = pos_bbox_targets
+
+ # regression loss
+ pos_bbox_weight = assign_metrics[pos_inds]
+
+ loss_bbox = self.loss_bbox(
+ pos_decode_bbox_pred,
+ pos_decode_bbox_targets,
+ weight=pos_bbox_weight,
+ avg_factor=1.0)
+ else:
+ loss_bbox = bbox_pred.sum() * 0
+ pos_bbox_weight = bbox_targets.new_tensor(0.)
+
+ return loss_cls, loss_bbox, assign_metrics.sum(), pos_bbox_weight.sum()
+
+ def loss_by_feat(self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None):
+ """Compute losses of the head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ Has shape (N, num_anchors * num_classes, H, W)
+ bbox_preds (list[Tensor]): Decoded box for each scale
+ level with shape (N, num_anchors * 4, H, W) in
+ [tl_x, tl_y, br_x, br_y] format.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ num_imgs = len(batch_img_metas)
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ assert len(featmap_sizes) == self.prior_generator.num_levels
+
+ device = cls_scores[0].device
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+ flatten_cls_scores = torch.cat([
+ cls_score.permute(0, 2, 3, 1).reshape(num_imgs, -1,
+ self.cls_out_channels)
+ for cls_score in cls_scores
+ ], 1)
+ decoded_bboxes = []
+ for anchor, bbox_pred in zip(anchor_list[0], bbox_preds):
+ anchor = anchor.reshape(-1, 4)
+ bbox_pred = bbox_pred.permute(0, 2, 3, 1).reshape(num_imgs, -1, 4)
+ bbox_pred = distance2bbox(anchor, bbox_pred)
+ decoded_bboxes.append(bbox_pred)
+
+ flatten_bboxes = torch.cat(decoded_bboxes, 1)
+
+ cls_reg_targets = self.get_targets(
+ flatten_cls_scores,
+ flatten_bboxes,
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore)
+ (anchor_list, labels_list, label_weights_list, bbox_targets_list,
+ assign_metrics_list, sampling_results_list) = cls_reg_targets
+
+ losses_cls, losses_bbox,\
+ cls_avg_factors, bbox_avg_factors = multi_apply(
+ self.loss_by_feat_single,
+ cls_scores,
+ decoded_bboxes,
+ labels_list,
+ label_weights_list,
+ bbox_targets_list,
+ assign_metrics_list,
+ self.prior_generator.strides)
+
+ cls_avg_factor = reduce_mean(sum(cls_avg_factors)).clamp_(min=1).item()
+ losses_cls = list(map(lambda x: x / cls_avg_factor, losses_cls))
+
+ bbox_avg_factor = reduce_mean(
+ sum(bbox_avg_factors)).clamp_(min=1).item()
+ losses_bbox = list(map(lambda x: x / bbox_avg_factor, losses_bbox))
+ return dict(loss_cls=losses_cls, loss_bbox=losses_bbox)
+
+ def get_targets(self,
+ cls_scores: Tensor,
+ bbox_preds: Tensor,
+ anchor_list: List[List[Tensor]],
+ valid_flag_list: List[List[Tensor]],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None,
+ unmap_outputs=True):
+ """Compute regression and classification targets for anchors in
+ multiple images.
+
+ Args:
+ cls_scores (Tensor): Classification predictions of images,
+ a 3D-Tensor with shape [num_imgs, num_priors, num_classes].
+ bbox_preds (Tensor): Decoded bboxes predictions of one image,
+ a 3D-Tensor with shape [num_imgs, num_priors, 4] in [tl_x,
+ tl_y, br_x, br_y] format.
+ anchor_list (list[list[Tensor]]): Multi level anchors of each
+ image. The outer list indicates images, and the inner list
+ corresponds to feature levels of the image. Each element of
+ the inner list is a tensor of shape (num_anchors, 4).
+ valid_flag_list (list[list[Tensor]]): Multi level valid flags of
+ each image. The outer list indicates images, and the inner list
+ corresponds to feature levels of the image. Each element of
+ the inner list is a tensor of shape (num_anchors, )
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ unmap_outputs (bool): Whether to map outputs back to the original
+ set of anchors. Defaults to True.
+
+ Returns:
+ tuple: a tuple containing learning targets.
+
+ - anchors_list (list[list[Tensor]]): Anchors of each level.
+ - labels_list (list[Tensor]): Labels of each level.
+ - label_weights_list (list[Tensor]): Label weights of each
+ level.
+ - bbox_targets_list (list[Tensor]): BBox targets of each level.
+ - assign_metrics_list (list[Tensor]): alignment metrics of each
+ level.
+ """
+ num_imgs = len(batch_img_metas)
+ assert len(anchor_list) == len(valid_flag_list) == num_imgs
+
+ # anchor number of multi levels
+ num_level_anchors = [anchors.size(0) for anchors in anchor_list[0]]
+
+ # concat all level anchors and flags to a single tensor
+ for i in range(num_imgs):
+ assert len(anchor_list[i]) == len(valid_flag_list[i])
+ anchor_list[i] = torch.cat(anchor_list[i])
+ valid_flag_list[i] = torch.cat(valid_flag_list[i])
+
+ # compute targets for each image
+ if batch_gt_instances_ignore is None:
+ batch_gt_instances_ignore = [None] * num_imgs
+ # anchor_list: list(b * [-1, 4])
+ (all_anchors, all_labels, all_label_weights, all_bbox_targets,
+ all_assign_metrics, sampling_results_list) = multi_apply(
+ self._get_targets_single,
+ cls_scores.detach(),
+ bbox_preds.detach(),
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore,
+ unmap_outputs=unmap_outputs)
+ # no valid anchors
+ if any([labels is None for labels in all_labels]):
+ return None
+
+ # split targets to a list w.r.t. multiple levels
+ anchors_list = images_to_levels(all_anchors, num_level_anchors)
+ labels_list = images_to_levels(all_labels, num_level_anchors)
+ label_weights_list = images_to_levels(all_label_weights,
+ num_level_anchors)
+ bbox_targets_list = images_to_levels(all_bbox_targets,
+ num_level_anchors)
+ assign_metrics_list = images_to_levels(all_assign_metrics,
+ num_level_anchors)
+
+ return (anchors_list, labels_list, label_weights_list,
+ bbox_targets_list, assign_metrics_list, sampling_results_list)
+
+ def _get_targets_single(self,
+ cls_scores: Tensor,
+ bbox_preds: Tensor,
+ flat_anchors: Tensor,
+ valid_flags: Tensor,
+ gt_instances: InstanceData,
+ img_meta: dict,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ unmap_outputs=True):
+ """Compute regression, classification targets for anchors in a single
+ image.
+
+ Args:
+ cls_scores (list(Tensor)): Box scores for each image.
+ bbox_preds (list(Tensor)): Box energies / deltas for each image.
+ flat_anchors (Tensor): Multi-level anchors of the image, which are
+ concatenated into a single tensor of shape (num_anchors ,4)
+ valid_flags (Tensor): Multi level valid flags of the image,
+ which are concatenated into a single tensor of
+ shape (num_anchors,).
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ img_meta (dict): Meta information for current image.
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ unmap_outputs (bool): Whether to map outputs back to the original
+ set of anchors. Defaults to True.
+
+ Returns:
+ tuple: N is the number of total anchors in the image.
+
+ - anchors (Tensor): All anchors in the image with shape (N, 4).
+ - labels (Tensor): Labels of all anchors in the image with shape
+ (N,).
+ - label_weights (Tensor): Label weights of all anchor in the
+ image with shape (N,).
+ - bbox_targets (Tensor): BBox targets of all anchors in the
+ image with shape (N, 4).
+ - norm_alignment_metrics (Tensor): Normalized alignment metrics
+ of all priors in the image with shape (N,).
+ """
+ inside_flags = anchor_inside_flags(flat_anchors, valid_flags,
+ img_meta['img_shape'][:2],
+ self.train_cfg['allowed_border'])
+ if not inside_flags.any():
+ return (None, ) * 7
+ # assign gt and sample anchors
+ anchors = flat_anchors[inside_flags, :]
+
+ pred_instances = InstanceData(
+ scores=cls_scores[inside_flags, :],
+ bboxes=bbox_preds[inside_flags, :],
+ priors=anchors)
+
+ assign_result = self.assigner.assign(pred_instances, gt_instances,
+ gt_instances_ignore)
+
+ sampling_result = self.sampler.sample(assign_result, pred_instances,
+ gt_instances)
+
+ num_valid_anchors = anchors.shape[0]
+ bbox_targets = torch.zeros_like(anchors)
+ labels = anchors.new_full((num_valid_anchors, ),
+ self.num_classes,
+ dtype=torch.long)
+ label_weights = anchors.new_zeros(num_valid_anchors, dtype=torch.float)
+ assign_metrics = anchors.new_zeros(
+ num_valid_anchors, dtype=torch.float)
+
+ pos_inds = sampling_result.pos_inds
+ neg_inds = sampling_result.neg_inds
+ if len(pos_inds) > 0:
+ # point-based
+ pos_bbox_targets = sampling_result.pos_gt_bboxes
+ bbox_targets[pos_inds, :] = pos_bbox_targets
+
+ labels[pos_inds] = sampling_result.pos_gt_labels
+ if self.train_cfg['pos_weight'] <= 0:
+ label_weights[pos_inds] = 1.0
+ else:
+ label_weights[pos_inds] = self.train_cfg['pos_weight']
+ if len(neg_inds) > 0:
+ label_weights[neg_inds] = 1.0
+
+ class_assigned_gt_inds = torch.unique(
+ sampling_result.pos_assigned_gt_inds)
+ for gt_inds in class_assigned_gt_inds:
+ gt_class_inds = pos_inds[sampling_result.pos_assigned_gt_inds ==
+ gt_inds]
+ assign_metrics[gt_class_inds] = assign_result.max_overlaps[
+ gt_class_inds]
+
+ # map up to original set of anchors
+ if unmap_outputs:
+ num_total_anchors = flat_anchors.size(0)
+ anchors = unmap(anchors, num_total_anchors, inside_flags)
+ labels = unmap(
+ labels, num_total_anchors, inside_flags, fill=self.num_classes)
+ label_weights = unmap(label_weights, num_total_anchors,
+ inside_flags)
+ bbox_targets = unmap(bbox_targets, num_total_anchors, inside_flags)
+ assign_metrics = unmap(assign_metrics, num_total_anchors,
+ inside_flags)
+ return (anchors, labels, label_weights, bbox_targets, assign_metrics,
+ sampling_result)
+
+ def get_anchors(self,
+ featmap_sizes: List[tuple],
+ batch_img_metas: List[dict],
+ device: Union[torch.device, str] = 'cuda') \
+ -> Tuple[List[List[Tensor]], List[List[Tensor]]]:
+ """Get anchors according to feature map sizes.
+
+ Args:
+ featmap_sizes (list[tuple]): Multi-level feature map sizes.
+ batch_img_metas (list[dict]): Image meta info.
+ device (torch.device or str): Device for returned tensors.
+ Defaults to cuda.
+
+ Returns:
+ tuple:
+
+ - anchor_list (list[list[Tensor]]): Anchors of each image.
+ - valid_flag_list (list[list[Tensor]]): Valid flags of each
+ image.
+ """
+ num_imgs = len(batch_img_metas)
+
+ # since feature map sizes of all images are the same, we only compute
+ # anchors for one time
+ multi_level_anchors = self.prior_generator.grid_priors(
+ featmap_sizes, device=device, with_stride=True)
+ anchor_list = [multi_level_anchors for _ in range(num_imgs)]
+
+ # for each image, we compute valid flags of multi level anchors
+ valid_flag_list = []
+ for img_id, img_meta in enumerate(batch_img_metas):
+ multi_level_flags = self.prior_generator.valid_flags(
+ featmap_sizes, img_meta['pad_shape'], device)
+ valid_flag_list.append(multi_level_flags)
+ return anchor_list, valid_flag_list
+
+
+@MODELS.register_module()
+class RTMDetSepBNHead(RTMDetHead):
+ """RTMDetHead with separated BN layers and shared conv layers.
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ share_conv (bool): Whether to share conv layers between stages.
+ Defaults to True.
+ use_depthwise (bool): Whether to use depthwise separable convolution in
+ head. Defaults to False.
+ norm_cfg (:obj:`ConfigDict` or dict)): Config dict for normalization
+ layer. Defaults to dict(type='BN', momentum=0.03, eps=0.001).
+ act_cfg (:obj:`ConfigDict` or dict)): Config dict for activation layer.
+ Defaults to dict(type='SiLU').
+ pred_kernel_size (int): Kernel size of prediction layer. Defaults to 1.
+ """
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: int,
+ share_conv: bool = True,
+ use_depthwise: bool = False,
+ norm_cfg: ConfigType = dict(
+ type='BN', momentum=0.03, eps=0.001),
+ act_cfg: ConfigType = dict(type='SiLU'),
+ pred_kernel_size: int = 1,
+ exp_on_reg=False,
+ **kwargs) -> None:
+ self.share_conv = share_conv
+ self.exp_on_reg = exp_on_reg
+ self.use_depthwise = use_depthwise
+ super().__init__(
+ num_classes,
+ in_channels,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg,
+ pred_kernel_size=pred_kernel_size,
+ **kwargs)
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ conv = DepthwiseSeparableConvModule \
+ if self.use_depthwise else ConvModule
+ self.cls_convs = nn.ModuleList()
+ self.reg_convs = nn.ModuleList()
+
+ self.rtm_cls = nn.ModuleList()
+ self.rtm_reg = nn.ModuleList()
+ if self.with_objectness:
+ self.rtm_obj = nn.ModuleList()
+ for n in range(len(self.prior_generator.strides)):
+ cls_convs = nn.ModuleList()
+ reg_convs = nn.ModuleList()
+ for i in range(self.stacked_convs):
+ chn = self.in_channels if i == 0 else self.feat_channels
+ cls_convs.append(
+ conv(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg))
+ reg_convs.append(
+ conv(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg))
+ self.cls_convs.append(cls_convs)
+ self.reg_convs.append(reg_convs)
+
+ self.rtm_cls.append(
+ nn.Conv2d(
+ self.feat_channels,
+ self.num_base_priors * self.cls_out_channels,
+ self.pred_kernel_size,
+ padding=self.pred_kernel_size // 2))
+ self.rtm_reg.append(
+ nn.Conv2d(
+ self.feat_channels,
+ self.num_base_priors * 4,
+ self.pred_kernel_size,
+ padding=self.pred_kernel_size // 2))
+ if self.with_objectness:
+ self.rtm_obj.append(
+ nn.Conv2d(
+ self.feat_channels,
+ 1,
+ self.pred_kernel_size,
+ padding=self.pred_kernel_size // 2))
+
+ if self.share_conv:
+ for n in range(len(self.prior_generator.strides)):
+ for i in range(self.stacked_convs):
+ self.cls_convs[n][i].conv = self.cls_convs[0][i].conv
+ self.reg_convs[n][i].conv = self.reg_convs[0][i].conv
+
+ def init_weights(self) -> None:
+ """Initialize weights of the head."""
+ for m in self.modules():
+ if isinstance(m, nn.Conv2d):
+ normal_init(m, mean=0, std=0.01)
+ if is_norm(m):
+ constant_init(m, 1)
+ bias_cls = bias_init_with_prob(0.01)
+ for rtm_cls, rtm_reg in zip(self.rtm_cls, self.rtm_reg):
+ normal_init(rtm_cls, std=0.01, bias=bias_cls)
+ normal_init(rtm_reg, std=0.01)
+ if self.with_objectness:
+ for rtm_obj in self.rtm_obj:
+ normal_init(rtm_obj, std=0.01, bias=bias_cls)
+
+ def forward(self, feats: Tuple[Tensor, ...]) -> tuple:
+ """Forward features from the upstream network.
+
+ Args:
+ feats (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: Usually a tuple of classification scores and bbox prediction
+
+ - cls_scores (tuple[Tensor]): Classification scores for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_anchors * num_classes.
+ - bbox_preds (tuple[Tensor]): Box energies / deltas for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_anchors * 4.
+ """
+
+ cls_scores = []
+ bbox_preds = []
+ for idx, (x, stride) in enumerate(
+ zip(feats, self.prior_generator.strides)):
+ cls_feat = x
+ reg_feat = x
+
+ for cls_layer in self.cls_convs[idx]:
+ cls_feat = cls_layer(cls_feat)
+ cls_score = self.rtm_cls[idx](cls_feat)
+
+ for reg_layer in self.reg_convs[idx]:
+ reg_feat = reg_layer(reg_feat)
+
+ if self.with_objectness:
+ objectness = self.rtm_obj[idx](reg_feat)
+ cls_score = inverse_sigmoid(
+ sigmoid_geometric_mean(cls_score, objectness))
+ if self.exp_on_reg:
+ reg_dist = self.rtm_reg[idx](reg_feat).exp() * stride[0]
+ else:
+ reg_dist = self.rtm_reg[idx](reg_feat) * stride[0]
+ cls_scores.append(cls_score)
+ bbox_preds.append(reg_dist)
+ return tuple(cls_scores), tuple(bbox_preds)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/rtmdet_ins_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/rtmdet_ins_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..261a57fe485245dcbe41696c9237258f829ca25a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/rtmdet_ins_head.py
@@ -0,0 +1,1034 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import math
+from typing import List, Optional, Tuple
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule, is_norm
+from mmcv.ops import batched_nms
+from mmengine.model import (BaseModule, bias_init_with_prob, constant_init,
+ normal_init)
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.models.layers.transformer import inverse_sigmoid
+from mmdet.models.utils import (filter_scores_and_topk, multi_apply,
+ select_single_mlvl, sigmoid_geometric_mean)
+from mmdet.registry import MODELS
+from mmdet.structures.bbox import (cat_boxes, distance2bbox, get_box_tensor,
+ get_box_wh, scale_boxes)
+from mmdet.utils import ConfigType, InstanceList, OptInstanceList, reduce_mean
+from .rtmdet_head import RTMDetHead
+
+
+@MODELS.register_module()
+class RTMDetInsHead(RTMDetHead):
+ """Detection Head of RTMDet-Ins.
+
+ Args:
+ num_prototypes (int): Number of mask prototype features extracted
+ from the mask head. Defaults to 8.
+ dyconv_channels (int): Channel of the dynamic conv layers.
+ Defaults to 8.
+ num_dyconvs (int): Number of the dynamic convolution layers.
+ Defaults to 3.
+ mask_loss_stride (int): Down sample stride of the masks for loss
+ computation. Defaults to 4.
+ loss_mask (:obj:`ConfigDict` or dict): Config dict for mask loss.
+ """
+
+ def __init__(self,
+ *args,
+ num_prototypes: int = 8,
+ dyconv_channels: int = 8,
+ num_dyconvs: int = 3,
+ mask_loss_stride: int = 4,
+ loss_mask=dict(
+ type='DiceLoss',
+ loss_weight=2.0,
+ eps=5e-6,
+ reduction='mean'),
+ **kwargs) -> None:
+ self.num_prototypes = num_prototypes
+ self.num_dyconvs = num_dyconvs
+ self.dyconv_channels = dyconv_channels
+ self.mask_loss_stride = mask_loss_stride
+ super().__init__(*args, **kwargs)
+ self.loss_mask = MODELS.build(loss_mask)
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ super()._init_layers()
+ # a branch to predict kernels of dynamic convs
+ self.kernel_convs = nn.ModuleList()
+ # calculate num dynamic parameters
+ weight_nums, bias_nums = [], []
+ for i in range(self.num_dyconvs):
+ if i == 0:
+ weight_nums.append(
+ # mask prototype and coordinate features
+ (self.num_prototypes + 2) * self.dyconv_channels)
+ bias_nums.append(self.dyconv_channels * 1)
+ elif i == self.num_dyconvs - 1:
+ weight_nums.append(self.dyconv_channels * 1)
+ bias_nums.append(1)
+ else:
+ weight_nums.append(self.dyconv_channels * self.dyconv_channels)
+ bias_nums.append(self.dyconv_channels * 1)
+ self.weight_nums = weight_nums
+ self.bias_nums = bias_nums
+ self.num_gen_params = sum(weight_nums) + sum(bias_nums)
+
+ for i in range(self.stacked_convs):
+ chn = self.in_channels if i == 0 else self.feat_channels
+ self.kernel_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg))
+ pred_pad_size = self.pred_kernel_size // 2
+ self.rtm_kernel = nn.Conv2d(
+ self.feat_channels,
+ self.num_gen_params,
+ self.pred_kernel_size,
+ padding=pred_pad_size)
+ self.mask_head = MaskFeatModule(
+ in_channels=self.in_channels,
+ feat_channels=self.feat_channels,
+ stacked_convs=4,
+ num_levels=len(self.prior_generator.strides),
+ num_prototypes=self.num_prototypes,
+ act_cfg=self.act_cfg,
+ norm_cfg=self.norm_cfg)
+
+ def forward(self, feats: Tuple[Tensor, ...]) -> tuple:
+ """Forward features from the upstream network.
+
+ Args:
+ feats (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: Usually a tuple of classification scores and bbox prediction
+ - cls_scores (list[Tensor]): Classification scores for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_base_priors * num_classes.
+ - bbox_preds (list[Tensor]): Box energies / deltas for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_base_priors * 4.
+ - kernel_preds (list[Tensor]): Dynamic conv kernels for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_gen_params.
+ - mask_feat (Tensor): Output feature of the mask head. Each is a
+ 4D-tensor, the channels number is num_prototypes.
+ """
+ mask_feat = self.mask_head(feats)
+
+ cls_scores = []
+ bbox_preds = []
+ kernel_preds = []
+ for idx, (x, scale, stride) in enumerate(
+ zip(feats, self.scales, self.prior_generator.strides)):
+ cls_feat = x
+ reg_feat = x
+ kernel_feat = x
+
+ for cls_layer in self.cls_convs:
+ cls_feat = cls_layer(cls_feat)
+ cls_score = self.rtm_cls(cls_feat)
+
+ for kernel_layer in self.kernel_convs:
+ kernel_feat = kernel_layer(kernel_feat)
+ kernel_pred = self.rtm_kernel(kernel_feat)
+
+ for reg_layer in self.reg_convs:
+ reg_feat = reg_layer(reg_feat)
+
+ if self.with_objectness:
+ objectness = self.rtm_obj(reg_feat)
+ cls_score = inverse_sigmoid(
+ sigmoid_geometric_mean(cls_score, objectness))
+
+ reg_dist = scale(self.rtm_reg(reg_feat)) * stride[0]
+
+ cls_scores.append(cls_score)
+ bbox_preds.append(reg_dist)
+ kernel_preds.append(kernel_pred)
+ return tuple(cls_scores), tuple(bbox_preds), tuple(
+ kernel_preds), mask_feat
+
+ def predict_by_feat(self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ kernel_preds: List[Tensor],
+ mask_feat: Tensor,
+ score_factors: Optional[List[Tensor]] = None,
+ batch_img_metas: Optional[List[dict]] = None,
+ cfg: Optional[ConfigType] = None,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ bbox results.
+
+ Note: When score_factors is not None, the cls_scores are
+ usually multiplied by it then obtain the real score used in NMS,
+ such as CenterNess in FCOS, IoU branch in ATSS.
+
+ Args:
+ cls_scores (list[Tensor]): Classification scores for all
+ scale levels, each is a 4D-tensor, has shape
+ (batch_size, num_priors * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas for all
+ scale levels, each is a 4D-tensor, has shape
+ (batch_size, num_priors * 4, H, W).
+ kernel_preds (list[Tensor]): Kernel predictions of dynamic
+ convs for all scale levels, each is a 4D-tensor, has shape
+ (batch_size, num_params, H, W).
+ mask_feat (Tensor): Mask prototype features extracted from the
+ mask head, has shape (batch_size, num_prototypes, H, W).
+ score_factors (list[Tensor], optional): Score factor for
+ all scale level, each is a 4D-tensor, has shape
+ (batch_size, num_priors * 1, H, W). Defaults to None.
+ batch_img_metas (list[dict], Optional): Batch image meta info.
+ Defaults to None.
+ cfg (ConfigDict, optional): Test / postprocessing
+ configuration, if None, test_cfg would be used.
+ Defaults to None.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ list[:obj:`InstanceData`]: Object detection results of each image
+ after the post process. Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, h, w).
+ """
+ assert len(cls_scores) == len(bbox_preds)
+
+ if score_factors is None:
+ # e.g. Retina, FreeAnchor, Foveabox, etc.
+ with_score_factors = False
+ else:
+ # e.g. FCOS, PAA, ATSS, AutoAssign, etc.
+ with_score_factors = True
+ assert len(cls_scores) == len(score_factors)
+
+ num_levels = len(cls_scores)
+
+ featmap_sizes = [cls_scores[i].shape[-2:] for i in range(num_levels)]
+ mlvl_priors = self.prior_generator.grid_priors(
+ featmap_sizes,
+ dtype=cls_scores[0].dtype,
+ device=cls_scores[0].device,
+ with_stride=True)
+
+ result_list = []
+
+ for img_id in range(len(batch_img_metas)):
+ img_meta = batch_img_metas[img_id]
+ cls_score_list = select_single_mlvl(
+ cls_scores, img_id, detach=True)
+ bbox_pred_list = select_single_mlvl(
+ bbox_preds, img_id, detach=True)
+ kernel_pred_list = select_single_mlvl(
+ kernel_preds, img_id, detach=True)
+ if with_score_factors:
+ score_factor_list = select_single_mlvl(
+ score_factors, img_id, detach=True)
+ else:
+ score_factor_list = [None for _ in range(num_levels)]
+
+ results = self._predict_by_feat_single(
+ cls_score_list=cls_score_list,
+ bbox_pred_list=bbox_pred_list,
+ kernel_pred_list=kernel_pred_list,
+ mask_feat=mask_feat[img_id],
+ score_factor_list=score_factor_list,
+ mlvl_priors=mlvl_priors,
+ img_meta=img_meta,
+ cfg=cfg,
+ rescale=rescale,
+ with_nms=with_nms)
+ result_list.append(results)
+ return result_list
+
+ def _predict_by_feat_single(self,
+ cls_score_list: List[Tensor],
+ bbox_pred_list: List[Tensor],
+ kernel_pred_list: List[Tensor],
+ mask_feat: Tensor,
+ score_factor_list: List[Tensor],
+ mlvl_priors: List[Tensor],
+ img_meta: dict,
+ cfg: ConfigType,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox and mask results.
+
+ Args:
+ cls_score_list (list[Tensor]): Box scores from all scale
+ levels of a single image, each item has shape
+ (num_priors * num_classes, H, W).
+ bbox_pred_list (list[Tensor]): Box energies / deltas from
+ all scale levels of a single image, each item has shape
+ (num_priors * 4, H, W).
+ kernel_preds (list[Tensor]): Kernel predictions of dynamic
+ convs for all scale levels of a single image, each is a
+ 4D-tensor, has shape (num_params, H, W).
+ mask_feat (Tensor): Mask prototype features of a single image
+ extracted from the mask head, has shape (num_prototypes, H, W).
+ score_factor_list (list[Tensor]): Score factor from all scale
+ levels of a single image, each item has shape
+ (num_priors * 1, H, W).
+ mlvl_priors (list[Tensor]): Each element in the list is
+ the priors of a single level in feature pyramid. In all
+ anchor-based methods, it has shape (num_priors, 4). In
+ all anchor-free methods, it has shape (num_priors, 2)
+ when `with_stride=True`, otherwise it still has shape
+ (num_priors, 4).
+ img_meta (dict): Image meta info.
+ cfg (mmengine.Config): Test / postprocessing configuration,
+ if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, h, w).
+ """
+ if score_factor_list[0] is None:
+ # e.g. Retina, FreeAnchor, etc.
+ with_score_factors = False
+ else:
+ # e.g. FCOS, PAA, ATSS, etc.
+ with_score_factors = True
+
+ cfg = self.test_cfg if cfg is None else cfg
+ cfg = copy.deepcopy(cfg)
+ img_shape = img_meta['img_shape']
+ nms_pre = cfg.get('nms_pre', -1)
+
+ mlvl_bbox_preds = []
+ mlvl_kernels = []
+ mlvl_valid_priors = []
+ mlvl_scores = []
+ mlvl_labels = []
+ if with_score_factors:
+ mlvl_score_factors = []
+ else:
+ mlvl_score_factors = None
+
+ for level_idx, (cls_score, bbox_pred, kernel_pred,
+ score_factor, priors) in \
+ enumerate(zip(cls_score_list, bbox_pred_list, kernel_pred_list,
+ score_factor_list, mlvl_priors)):
+
+ assert cls_score.size()[-2:] == bbox_pred.size()[-2:]
+
+ dim = self.bbox_coder.encode_size
+ bbox_pred = bbox_pred.permute(1, 2, 0).reshape(-1, dim)
+ if with_score_factors:
+ score_factor = score_factor.permute(1, 2,
+ 0).reshape(-1).sigmoid()
+ cls_score = cls_score.permute(1, 2,
+ 0).reshape(-1, self.cls_out_channels)
+ kernel_pred = kernel_pred.permute(1, 2, 0).reshape(
+ -1, self.num_gen_params)
+ if self.use_sigmoid_cls:
+ scores = cls_score.sigmoid()
+ else:
+ # remind that we set FG labels to [0, num_class-1]
+ # since mmdet v2.0
+ # BG cat_id: num_class
+ scores = cls_score.softmax(-1)[:, :-1]
+
+ # After https://github.com/open-mmlab/mmdetection/pull/6268/,
+ # this operation keeps fewer bboxes under the same `nms_pre`.
+ # There is no difference in performance for most models. If you
+ # find a slight drop in performance, you can set a larger
+ # `nms_pre` than before.
+ score_thr = cfg.get('score_thr', 0)
+
+ results = filter_scores_and_topk(
+ scores, score_thr, nms_pre,
+ dict(
+ bbox_pred=bbox_pred,
+ priors=priors,
+ kernel_pred=kernel_pred))
+ scores, labels, keep_idxs, filtered_results = results
+
+ bbox_pred = filtered_results['bbox_pred']
+ priors = filtered_results['priors']
+ kernel_pred = filtered_results['kernel_pred']
+
+ if with_score_factors:
+ score_factor = score_factor[keep_idxs]
+
+ mlvl_bbox_preds.append(bbox_pred)
+ mlvl_valid_priors.append(priors)
+ mlvl_scores.append(scores)
+ mlvl_labels.append(labels)
+ mlvl_kernels.append(kernel_pred)
+
+ if with_score_factors:
+ mlvl_score_factors.append(score_factor)
+
+ bbox_pred = torch.cat(mlvl_bbox_preds)
+ priors = cat_boxes(mlvl_valid_priors)
+ bboxes = self.bbox_coder.decode(
+ priors[..., :2], bbox_pred, max_shape=img_shape)
+
+ results = InstanceData()
+ results.bboxes = bboxes
+ results.priors = priors
+ results.scores = torch.cat(mlvl_scores)
+ results.labels = torch.cat(mlvl_labels)
+ results.kernels = torch.cat(mlvl_kernels)
+ if with_score_factors:
+ results.score_factors = torch.cat(mlvl_score_factors)
+
+ return self._bbox_mask_post_process(
+ results=results,
+ mask_feat=mask_feat,
+ cfg=cfg,
+ rescale=rescale,
+ with_nms=with_nms,
+ img_meta=img_meta)
+
+ def _bbox_mask_post_process(
+ self,
+ results: InstanceData,
+ mask_feat,
+ cfg: ConfigType,
+ rescale: bool = False,
+ with_nms: bool = True,
+ img_meta: Optional[dict] = None) -> InstanceData:
+ """bbox and mask post-processing method.
+
+ The boxes would be rescaled to the original image scale and do
+ the nms operation. Usually `with_nms` is False is used for aug test.
+
+ Args:
+ results (:obj:`InstaceData`): Detection instance results,
+ each item has shape (num_bboxes, ).
+ cfg (ConfigDict): Test / postprocessing configuration,
+ if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Default to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Default to True.
+ img_meta (dict, optional): Image meta info. Defaults to None.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, h, w).
+ """
+ stride = self.prior_generator.strides[0][0]
+ if rescale:
+ assert img_meta.get('scale_factor') is not None
+ scale_factor = [1 / s for s in img_meta['scale_factor']]
+ results.bboxes = scale_boxes(results.bboxes, scale_factor)
+
+ if hasattr(results, 'score_factors'):
+ # TODO: Add sqrt operation in order to be consistent with
+ # the paper.
+ score_factors = results.pop('score_factors')
+ results.scores = results.scores * score_factors
+
+ # filter small size bboxes
+ if cfg.get('min_bbox_size', -1) >= 0:
+ w, h = get_box_wh(results.bboxes)
+ valid_mask = (w > cfg.min_bbox_size) & (h > cfg.min_bbox_size)
+ if not valid_mask.all():
+ results = results[valid_mask]
+
+ # TODO: deal with `with_nms` and `nms_cfg=None` in test_cfg
+ assert with_nms, 'with_nms must be True for RTMDet-Ins'
+ if results.bboxes.numel() > 0:
+ bboxes = get_box_tensor(results.bboxes)
+ det_bboxes, keep_idxs = batched_nms(bboxes, results.scores,
+ results.labels, cfg.nms)
+ results = results[keep_idxs]
+ # some nms would reweight the score, such as softnms
+ results.scores = det_bboxes[:, -1]
+ results = results[:cfg.max_per_img]
+
+ # process masks
+ mask_logits = self._mask_predict_by_feat_single(
+ mask_feat, results.kernels, results.priors)
+
+ mask_logits = F.interpolate(
+ mask_logits.unsqueeze(0), scale_factor=stride, mode='bilinear')
+ if rescale:
+ ori_h, ori_w = img_meta['ori_shape'][:2]
+ mask_logits = F.interpolate(
+ mask_logits,
+ size=[
+ math.ceil(mask_logits.shape[-2] * scale_factor[0]),
+ math.ceil(mask_logits.shape[-1] * scale_factor[1])
+ ],
+ mode='bilinear',
+ align_corners=False)[..., :ori_h, :ori_w]
+ masks = mask_logits.sigmoid().squeeze(0)
+ masks = masks > cfg.mask_thr_binary
+ results.masks = masks
+ else:
+ h, w = img_meta['ori_shape'][:2] if rescale else img_meta[
+ 'img_shape'][:2]
+ results.masks = torch.zeros(
+ size=(results.bboxes.shape[0], h, w),
+ dtype=torch.bool,
+ device=results.bboxes.device)
+
+ return results
+
+ def parse_dynamic_params(self, flatten_kernels: Tensor) -> tuple:
+ """split kernel head prediction to conv weight and bias."""
+ n_inst = flatten_kernels.size(0)
+ n_layers = len(self.weight_nums)
+ params_splits = list(
+ torch.split_with_sizes(
+ flatten_kernels, self.weight_nums + self.bias_nums, dim=1))
+ weight_splits = params_splits[:n_layers]
+ bias_splits = params_splits[n_layers:]
+ for i in range(n_layers):
+ if i < n_layers - 1:
+ weight_splits[i] = weight_splits[i].reshape(
+ n_inst * self.dyconv_channels, -1, 1, 1)
+ bias_splits[i] = bias_splits[i].reshape(n_inst *
+ self.dyconv_channels)
+ else:
+ weight_splits[i] = weight_splits[i].reshape(n_inst, -1, 1, 1)
+ bias_splits[i] = bias_splits[i].reshape(n_inst)
+
+ return weight_splits, bias_splits
+
+ def _mask_predict_by_feat_single(self, mask_feat: Tensor, kernels: Tensor,
+ priors: Tensor) -> Tensor:
+ """Generate mask logits from mask features with dynamic convs.
+
+ Args:
+ mask_feat (Tensor): Mask prototype features.
+ Has shape (num_prototypes, H, W).
+ kernels (Tensor): Kernel parameters for each instance.
+ Has shape (num_instance, num_params)
+ priors (Tensor): Center priors for each instance.
+ Has shape (num_instance, 4).
+ Returns:
+ Tensor: Instance segmentation masks for each instance.
+ Has shape (num_instance, H, W).
+ """
+ num_inst = priors.shape[0]
+ h, w = mask_feat.size()[-2:]
+ if num_inst < 1:
+ return torch.empty(
+ size=(num_inst, h, w),
+ dtype=mask_feat.dtype,
+ device=mask_feat.device)
+ if len(mask_feat.shape) < 4:
+ mask_feat.unsqueeze(0)
+
+ coord = self.prior_generator.single_level_grid_priors(
+ (h, w), level_idx=0, device=mask_feat.device).reshape(1, -1, 2)
+ num_inst = priors.shape[0]
+ points = priors[:, :2].reshape(-1, 1, 2)
+ strides = priors[:, 2:].reshape(-1, 1, 2)
+ relative_coord = (points - coord).permute(0, 2, 1) / (
+ strides[..., 0].reshape(-1, 1, 1) * 8)
+ relative_coord = relative_coord.reshape(num_inst, 2, h, w)
+
+ mask_feat = torch.cat(
+ [relative_coord,
+ mask_feat.repeat(num_inst, 1, 1, 1)], dim=1)
+ weights, biases = self.parse_dynamic_params(kernels)
+
+ n_layers = len(weights)
+ x = mask_feat.reshape(1, -1, h, w)
+ for i, (weight, bias) in enumerate(zip(weights, biases)):
+ x = F.conv2d(
+ x, weight, bias=bias, stride=1, padding=0, groups=num_inst)
+ if i < n_layers - 1:
+ x = F.relu(x)
+ x = x.reshape(num_inst, h, w)
+ return x
+
+ def loss_mask_by_feat(self, mask_feats: Tensor, flatten_kernels: Tensor,
+ sampling_results_list: list,
+ batch_gt_instances: InstanceList) -> Tensor:
+ """Compute instance segmentation loss.
+
+ Args:
+ mask_feats (list[Tensor]): Mask prototype features extracted from
+ the mask head. Has shape (N, num_prototypes, H, W)
+ flatten_kernels (list[Tensor]): Kernels of the dynamic conv layers.
+ Has shape (N, num_instances, num_params)
+ sampling_results_list (list[:obj:`SamplingResults`]) Batch of
+ assignment results.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+
+ Returns:
+ Tensor: The mask loss tensor.
+ """
+ batch_pos_mask_logits = []
+ pos_gt_masks = []
+ for idx, (mask_feat, kernels, sampling_results,
+ gt_instances) in enumerate(
+ zip(mask_feats, flatten_kernels, sampling_results_list,
+ batch_gt_instances)):
+ pos_priors = sampling_results.pos_priors
+ pos_inds = sampling_results.pos_inds
+ pos_kernels = kernels[pos_inds] # n_pos, num_gen_params
+ pos_mask_logits = self._mask_predict_by_feat_single(
+ mask_feat, pos_kernels, pos_priors)
+ if gt_instances.masks.numel() == 0:
+ gt_masks = torch.empty_like(gt_instances.masks)
+ else:
+ gt_masks = gt_instances.masks[
+ sampling_results.pos_assigned_gt_inds, :]
+ batch_pos_mask_logits.append(pos_mask_logits)
+ pos_gt_masks.append(gt_masks)
+
+ pos_gt_masks = torch.cat(pos_gt_masks, 0)
+ batch_pos_mask_logits = torch.cat(batch_pos_mask_logits, 0)
+
+ # avg_factor
+ num_pos = batch_pos_mask_logits.shape[0]
+ num_pos = reduce_mean(mask_feats.new_tensor([num_pos
+ ])).clamp_(min=1).item()
+
+ if batch_pos_mask_logits.shape[0] == 0:
+ return mask_feats.sum() * 0
+
+ scale = self.prior_generator.strides[0][0] // self.mask_loss_stride
+ # upsample pred masks
+ batch_pos_mask_logits = F.interpolate(
+ batch_pos_mask_logits.unsqueeze(0),
+ scale_factor=scale,
+ mode='bilinear',
+ align_corners=False).squeeze(0)
+ # downsample gt masks
+ pos_gt_masks = pos_gt_masks[:, self.mask_loss_stride //
+ 2::self.mask_loss_stride,
+ self.mask_loss_stride //
+ 2::self.mask_loss_stride]
+
+ loss_mask = self.loss_mask(
+ batch_pos_mask_logits,
+ pos_gt_masks,
+ weight=None,
+ avg_factor=num_pos)
+
+ return loss_mask
+
+ def loss_by_feat(self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ kernel_preds: List[Tensor],
+ mask_feat: Tensor,
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None):
+ """Compute losses of the head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ Has shape (N, num_anchors * num_classes, H, W)
+ bbox_preds (list[Tensor]): Decoded box for each scale
+ level with shape (N, num_anchors * 4, H, W) in
+ [tl_x, tl_y, br_x, br_y] format.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ num_imgs = len(batch_img_metas)
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ assert len(featmap_sizes) == self.prior_generator.num_levels
+
+ device = cls_scores[0].device
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+ flatten_cls_scores = torch.cat([
+ cls_score.permute(0, 2, 3, 1).reshape(num_imgs, -1,
+ self.cls_out_channels)
+ for cls_score in cls_scores
+ ], 1)
+ flatten_kernels = torch.cat([
+ kernel_pred.permute(0, 2, 3, 1).reshape(num_imgs, -1,
+ self.num_gen_params)
+ for kernel_pred in kernel_preds
+ ], 1)
+ decoded_bboxes = []
+ for anchor, bbox_pred in zip(anchor_list[0], bbox_preds):
+ anchor = anchor.reshape(-1, 4)
+ bbox_pred = bbox_pred.permute(0, 2, 3, 1).reshape(num_imgs, -1, 4)
+ bbox_pred = distance2bbox(anchor, bbox_pred)
+ decoded_bboxes.append(bbox_pred)
+
+ flatten_bboxes = torch.cat(decoded_bboxes, 1)
+ for gt_instances in batch_gt_instances:
+ gt_instances.masks = gt_instances.masks.to_tensor(
+ dtype=torch.bool, device=device)
+
+ cls_reg_targets = self.get_targets(
+ flatten_cls_scores,
+ flatten_bboxes,
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore)
+ (anchor_list, labels_list, label_weights_list, bbox_targets_list,
+ assign_metrics_list, sampling_results_list) = cls_reg_targets
+
+ losses_cls, losses_bbox,\
+ cls_avg_factors, bbox_avg_factors = multi_apply(
+ self.loss_by_feat_single,
+ cls_scores,
+ decoded_bboxes,
+ labels_list,
+ label_weights_list,
+ bbox_targets_list,
+ assign_metrics_list,
+ self.prior_generator.strides)
+
+ cls_avg_factor = reduce_mean(sum(cls_avg_factors)).clamp_(min=1).item()
+ losses_cls = list(map(lambda x: x / cls_avg_factor, losses_cls))
+
+ bbox_avg_factor = reduce_mean(
+ sum(bbox_avg_factors)).clamp_(min=1).item()
+ losses_bbox = list(map(lambda x: x / bbox_avg_factor, losses_bbox))
+
+ loss_mask = self.loss_mask_by_feat(mask_feat, flatten_kernels,
+ sampling_results_list,
+ batch_gt_instances)
+ loss = dict(
+ loss_cls=losses_cls, loss_bbox=losses_bbox, loss_mask=loss_mask)
+ return loss
+
+
+class MaskFeatModule(BaseModule):
+ """Mask feature head used in RTMDet-Ins.
+
+ Args:
+ in_channels (int): Number of channels in the input feature map.
+ feat_channels (int): Number of hidden channels of the mask feature
+ map branch.
+ num_levels (int): The starting feature map level from RPN that
+ will be used to predict the mask feature map.
+ num_prototypes (int): Number of output channel of the mask feature
+ map branch. This is the channel count of the mask
+ feature map that to be dynamically convolved with the predicted
+ kernel.
+ stacked_convs (int): Number of convs in mask feature branch.
+ act_cfg (:obj:`ConfigDict` or dict): Config dict for activation layer.
+ Default: dict(type='ReLU', inplace=True)
+ norm_cfg (dict): Config dict for normalization layer. Default: None.
+ """
+
+ def __init__(
+ self,
+ in_channels: int,
+ feat_channels: int = 256,
+ stacked_convs: int = 4,
+ num_levels: int = 3,
+ num_prototypes: int = 8,
+ act_cfg: ConfigType = dict(type='ReLU', inplace=True),
+ norm_cfg: ConfigType = dict(type='BN')
+ ) -> None:
+ super().__init__(init_cfg=None)
+ self.num_levels = num_levels
+ self.fusion_conv = nn.Conv2d(num_levels * in_channels, in_channels, 1)
+ convs = []
+ for i in range(stacked_convs):
+ in_c = in_channels if i == 0 else feat_channels
+ convs.append(
+ ConvModule(
+ in_c,
+ feat_channels,
+ 3,
+ padding=1,
+ act_cfg=act_cfg,
+ norm_cfg=norm_cfg))
+ self.stacked_convs = nn.Sequential(*convs)
+ self.projection = nn.Conv2d(
+ feat_channels, num_prototypes, kernel_size=1)
+
+ def forward(self, features: Tuple[Tensor, ...]) -> Tensor:
+ # multi-level feature fusion
+ fusion_feats = [features[0]]
+ size = features[0].shape[-2:]
+ for i in range(1, self.num_levels):
+ f = F.interpolate(features[i], size=size, mode='bilinear')
+ fusion_feats.append(f)
+ fusion_feats = torch.cat(fusion_feats, dim=1)
+ fusion_feats = self.fusion_conv(fusion_feats)
+ # pred mask feats
+ mask_features = self.stacked_convs(fusion_feats)
+ mask_features = self.projection(mask_features)
+ return mask_features
+
+
+@MODELS.register_module()
+class RTMDetInsSepBNHead(RTMDetInsHead):
+ """Detection Head of RTMDet-Ins with sep-bn layers.
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ share_conv (bool): Whether to share conv layers between stages.
+ Defaults to True.
+ norm_cfg (:obj:`ConfigDict` or dict)): Config dict for normalization
+ layer. Defaults to dict(type='BN').
+ act_cfg (:obj:`ConfigDict` or dict)): Config dict for activation layer.
+ Defaults to dict(type='SiLU', inplace=True).
+ pred_kernel_size (int): Kernel size of prediction layer. Defaults to 1.
+ """
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: int,
+ share_conv: bool = True,
+ with_objectness: bool = False,
+ norm_cfg: ConfigType = dict(type='BN', requires_grad=True),
+ act_cfg: ConfigType = dict(type='SiLU', inplace=True),
+ pred_kernel_size: int = 1,
+ **kwargs) -> None:
+ self.share_conv = share_conv
+ super().__init__(
+ num_classes,
+ in_channels,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg,
+ pred_kernel_size=pred_kernel_size,
+ with_objectness=with_objectness,
+ **kwargs)
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self.cls_convs = nn.ModuleList()
+ self.reg_convs = nn.ModuleList()
+ self.kernel_convs = nn.ModuleList()
+
+ self.rtm_cls = nn.ModuleList()
+ self.rtm_reg = nn.ModuleList()
+ self.rtm_kernel = nn.ModuleList()
+ self.rtm_obj = nn.ModuleList()
+
+ # calculate num dynamic parameters
+ weight_nums, bias_nums = [], []
+ for i in range(self.num_dyconvs):
+ if i == 0:
+ weight_nums.append(
+ (self.num_prototypes + 2) * self.dyconv_channels)
+ bias_nums.append(self.dyconv_channels)
+ elif i == self.num_dyconvs - 1:
+ weight_nums.append(self.dyconv_channels)
+ bias_nums.append(1)
+ else:
+ weight_nums.append(self.dyconv_channels * self.dyconv_channels)
+ bias_nums.append(self.dyconv_channels)
+ self.weight_nums = weight_nums
+ self.bias_nums = bias_nums
+ self.num_gen_params = sum(weight_nums) + sum(bias_nums)
+ pred_pad_size = self.pred_kernel_size // 2
+
+ for n in range(len(self.prior_generator.strides)):
+ cls_convs = nn.ModuleList()
+ reg_convs = nn.ModuleList()
+ kernel_convs = nn.ModuleList()
+ for i in range(self.stacked_convs):
+ chn = self.in_channels if i == 0 else self.feat_channels
+ cls_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg))
+ reg_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg))
+ kernel_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg))
+ self.cls_convs.append(cls_convs)
+ self.reg_convs.append(cls_convs)
+ self.kernel_convs.append(kernel_convs)
+
+ self.rtm_cls.append(
+ nn.Conv2d(
+ self.feat_channels,
+ self.num_base_priors * self.cls_out_channels,
+ self.pred_kernel_size,
+ padding=pred_pad_size))
+ self.rtm_reg.append(
+ nn.Conv2d(
+ self.feat_channels,
+ self.num_base_priors * 4,
+ self.pred_kernel_size,
+ padding=pred_pad_size))
+ self.rtm_kernel.append(
+ nn.Conv2d(
+ self.feat_channels,
+ self.num_gen_params,
+ self.pred_kernel_size,
+ padding=pred_pad_size))
+ if self.with_objectness:
+ self.rtm_obj.append(
+ nn.Conv2d(
+ self.feat_channels,
+ 1,
+ self.pred_kernel_size,
+ padding=pred_pad_size))
+
+ if self.share_conv:
+ for n in range(len(self.prior_generator.strides)):
+ for i in range(self.stacked_convs):
+ self.cls_convs[n][i].conv = self.cls_convs[0][i].conv
+ self.reg_convs[n][i].conv = self.reg_convs[0][i].conv
+
+ self.mask_head = MaskFeatModule(
+ in_channels=self.in_channels,
+ feat_channels=self.feat_channels,
+ stacked_convs=4,
+ num_levels=len(self.prior_generator.strides),
+ num_prototypes=self.num_prototypes,
+ act_cfg=self.act_cfg,
+ norm_cfg=self.norm_cfg)
+
+ def init_weights(self) -> None:
+ """Initialize weights of the head."""
+ for m in self.modules():
+ if isinstance(m, nn.Conv2d):
+ normal_init(m, mean=0, std=0.01)
+ if is_norm(m):
+ constant_init(m, 1)
+ bias_cls = bias_init_with_prob(0.01)
+ for rtm_cls, rtm_reg, rtm_kernel in zip(self.rtm_cls, self.rtm_reg,
+ self.rtm_kernel):
+ normal_init(rtm_cls, std=0.01, bias=bias_cls)
+ normal_init(rtm_reg, std=0.01, bias=1)
+ if self.with_objectness:
+ for rtm_obj in self.rtm_obj:
+ normal_init(rtm_obj, std=0.01, bias=bias_cls)
+
+ def forward(self, feats: Tuple[Tensor, ...]) -> tuple:
+ """Forward features from the upstream network.
+
+ Args:
+ feats (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: Usually a tuple of classification scores and bbox prediction
+ - cls_scores (list[Tensor]): Classification scores for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_base_priors * num_classes.
+ - bbox_preds (list[Tensor]): Box energies / deltas for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_base_priors * 4.
+ - kernel_preds (list[Tensor]): Dynamic conv kernels for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_gen_params.
+ - mask_feat (Tensor): Output feature of the mask head. Each is a
+ 4D-tensor, the channels number is num_prototypes.
+ """
+ mask_feat = self.mask_head(feats)
+
+ cls_scores = []
+ bbox_preds = []
+ kernel_preds = []
+ for idx, (x, stride) in enumerate(
+ zip(feats, self.prior_generator.strides)):
+ cls_feat = x
+ reg_feat = x
+ kernel_feat = x
+
+ for cls_layer in self.cls_convs[idx]:
+ cls_feat = cls_layer(cls_feat)
+ cls_score = self.rtm_cls[idx](cls_feat)
+
+ for kernel_layer in self.kernel_convs[idx]:
+ kernel_feat = kernel_layer(kernel_feat)
+ kernel_pred = self.rtm_kernel[idx](kernel_feat)
+
+ for reg_layer in self.reg_convs[idx]:
+ reg_feat = reg_layer(reg_feat)
+
+ if self.with_objectness:
+ objectness = self.rtm_obj[idx](reg_feat)
+ cls_score = inverse_sigmoid(
+ sigmoid_geometric_mean(cls_score, objectness))
+
+ reg_dist = F.relu(self.rtm_reg[idx](reg_feat)) * stride[0]
+
+ cls_scores.append(cls_score)
+ bbox_preds.append(reg_dist)
+ kernel_preds.append(kernel_pred)
+ return tuple(cls_scores), tuple(bbox_preds), tuple(
+ kernel_preds), mask_feat
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/sabl_retina_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/sabl_retina_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..8cd1b71cc2c80035a0378180da70caddf853375d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/sabl_retina_head.py
@@ -0,0 +1,706 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple, Union
+
+import numpy as np
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from mmengine.config import ConfigDict
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.utils import (ConfigType, InstanceList, MultiConfig, OptConfigType,
+ OptInstanceList)
+from ..task_modules.samplers import PseudoSampler
+from ..utils import (filter_scores_and_topk, images_to_levels, multi_apply,
+ unmap)
+from .base_dense_head import BaseDenseHead
+from .guided_anchor_head import GuidedAnchorHead
+
+
+@MODELS.register_module()
+class SABLRetinaHead(BaseDenseHead):
+ """Side-Aware Boundary Localization (SABL) for RetinaNet.
+
+ The anchor generation, assigning and sampling in SABLRetinaHead
+ are the same as GuidedAnchorHead for guided anchoring.
+
+ Please refer to https://arxiv.org/abs/1912.04260 for more details.
+
+ Args:
+ num_classes (int): Number of classes.
+ in_channels (int): Number of channels in the input feature map.
+ stacked_convs (int): Number of Convs for classification and
+ regression branches. Defaults to 4.
+ feat_channels (int): Number of hidden channels. Defaults to 256.
+ approx_anchor_generator (:obj:`ConfigType` or dict): Config dict for
+ approx generator.
+ square_anchor_generator (:obj:`ConfigDict` or dict): Config dict for
+ square generator.
+ conv_cfg (:obj:`ConfigDict` or dict, optional): Config dict for
+ ConvModule. Defaults to None.
+ norm_cfg (:obj:`ConfigDict` or dict, optional): Config dict for
+ Norm Layer. Defaults to None.
+ bbox_coder (:obj:`ConfigDict` or dict): Config dict for bbox coder.
+ reg_decoded_bbox (bool): If true, the regression loss would be
+ applied directly on decoded bounding boxes, converting both
+ the predicted boxes and regression targets to absolute
+ coordinates format. Default False. It should be ``True`` when
+ using ``IoULoss``, ``GIoULoss``, or ``DIoULoss`` in the bbox head.
+ train_cfg (:obj:`ConfigDict` or dict, optional): Training config of
+ SABLRetinaHead.
+ test_cfg (:obj:`ConfigDict` or dict, optional): Testing config of
+ SABLRetinaHead.
+ loss_cls (:obj:`ConfigDict` or dict): Config of classification loss.
+ loss_bbox_cls (:obj:`ConfigDict` or dict): Config of classification
+ loss for bbox branch.
+ loss_bbox_reg (:obj:`ConfigDict` or dict): Config of regression loss
+ for bbox branch.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict], optional): Initialization config dict.
+ """
+
+ def __init__(
+ self,
+ num_classes: int,
+ in_channels: int,
+ stacked_convs: int = 4,
+ feat_channels: int = 256,
+ approx_anchor_generator: ConfigType = dict(
+ type='AnchorGenerator',
+ octave_base_scale=4,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[8, 16, 32, 64, 128]),
+ square_anchor_generator: ConfigType = dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ scales=[4],
+ strides=[8, 16, 32, 64, 128]),
+ conv_cfg: OptConfigType = None,
+ norm_cfg: OptConfigType = None,
+ bbox_coder: ConfigType = dict(
+ type='BucketingBBoxCoder', num_buckets=14, scale_factor=3.0),
+ reg_decoded_bbox: bool = False,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ loss_cls: ConfigType = dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ loss_bbox_cls: ConfigType = dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.5),
+ loss_bbox_reg: ConfigType = dict(
+ type='SmoothL1Loss', beta=1.0 / 9.0, loss_weight=1.5),
+ init_cfg: MultiConfig = dict(
+ type='Normal',
+ layer='Conv2d',
+ std=0.01,
+ override=dict(
+ type='Normal', name='retina_cls', std=0.01, bias_prob=0.01))
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.in_channels = in_channels
+ self.num_classes = num_classes
+ self.feat_channels = feat_channels
+ self.num_buckets = bbox_coder['num_buckets']
+ self.side_num = int(np.ceil(self.num_buckets / 2))
+
+ assert (approx_anchor_generator['octave_base_scale'] ==
+ square_anchor_generator['scales'][0])
+ assert (approx_anchor_generator['strides'] ==
+ square_anchor_generator['strides'])
+
+ self.approx_anchor_generator = TASK_UTILS.build(
+ approx_anchor_generator)
+ self.square_anchor_generator = TASK_UTILS.build(
+ square_anchor_generator)
+ self.approxs_per_octave = (
+ self.approx_anchor_generator.num_base_priors[0])
+
+ # one anchor per location
+ self.num_base_priors = self.square_anchor_generator.num_base_priors[0]
+
+ self.stacked_convs = stacked_convs
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+
+ self.reg_decoded_bbox = reg_decoded_bbox
+
+ self.use_sigmoid_cls = loss_cls.get('use_sigmoid', False)
+ if self.use_sigmoid_cls:
+ self.cls_out_channels = num_classes
+ else:
+ self.cls_out_channels = num_classes + 1
+
+ self.bbox_coder = TASK_UTILS.build(bbox_coder)
+ self.loss_cls = MODELS.build(loss_cls)
+ self.loss_bbox_cls = MODELS.build(loss_bbox_cls)
+ self.loss_bbox_reg = MODELS.build(loss_bbox_reg)
+
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+
+ if self.train_cfg:
+ self.assigner = TASK_UTILS.build(self.train_cfg['assigner'])
+ # use PseudoSampler when sampling is False
+ if 'sampler' in self.train_cfg:
+ self.sampler = TASK_UTILS.build(
+ self.train_cfg['sampler'], default_args=dict(context=self))
+ else:
+ self.sampler = PseudoSampler(context=self)
+
+ self._init_layers()
+
+ def _init_layers(self) -> None:
+ self.relu = nn.ReLU(inplace=True)
+ self.cls_convs = nn.ModuleList()
+ self.reg_convs = nn.ModuleList()
+ for i in range(self.stacked_convs):
+ chn = self.in_channels if i == 0 else self.feat_channels
+ self.cls_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ self.reg_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ self.retina_cls = nn.Conv2d(
+ self.feat_channels, self.cls_out_channels, 3, padding=1)
+ self.retina_bbox_reg = nn.Conv2d(
+ self.feat_channels, self.side_num * 4, 3, padding=1)
+ self.retina_bbox_cls = nn.Conv2d(
+ self.feat_channels, self.side_num * 4, 3, padding=1)
+
+ def forward_single(self, x: Tensor) -> Tuple[Tensor, Tensor]:
+ cls_feat = x
+ reg_feat = x
+ for cls_conv in self.cls_convs:
+ cls_feat = cls_conv(cls_feat)
+ for reg_conv in self.reg_convs:
+ reg_feat = reg_conv(reg_feat)
+ cls_score = self.retina_cls(cls_feat)
+ bbox_cls_pred = self.retina_bbox_cls(reg_feat)
+ bbox_reg_pred = self.retina_bbox_reg(reg_feat)
+ bbox_pred = (bbox_cls_pred, bbox_reg_pred)
+ return cls_score, bbox_pred
+
+ def forward(self, feats: List[Tensor]) -> Tuple[List[Tensor]]:
+ return multi_apply(self.forward_single, feats)
+
+ def get_anchors(
+ self,
+ featmap_sizes: List[tuple],
+ img_metas: List[dict],
+ device: Union[torch.device, str] = 'cuda'
+ ) -> Tuple[List[List[Tensor]], List[List[Tensor]]]:
+ """Get squares according to feature map sizes and guided anchors.
+
+ Args:
+ featmap_sizes (list[tuple]): Multi-level feature map sizes.
+ img_metas (list[dict]): Image meta info.
+ device (torch.device | str): device for returned tensors
+
+ Returns:
+ tuple: square approxs of each image
+ """
+ num_imgs = len(img_metas)
+
+ # since feature map sizes of all images are the same, we only compute
+ # squares for one time
+ multi_level_squares = self.square_anchor_generator.grid_priors(
+ featmap_sizes, device=device)
+ squares_list = [multi_level_squares for _ in range(num_imgs)]
+
+ return squares_list
+
+ def get_targets(self,
+ approx_list: List[List[Tensor]],
+ inside_flag_list: List[List[Tensor]],
+ square_list: List[List[Tensor]],
+ batch_gt_instances: InstanceList,
+ batch_img_metas,
+ batch_gt_instances_ignore: OptInstanceList = None,
+ unmap_outputs=True) -> tuple:
+ """Compute bucketing targets.
+
+ Args:
+ approx_list (list[list[Tensor]]): Multi level approxs of each
+ image.
+ inside_flag_list (list[list[Tensor]]): Multi level inside flags of
+ each image.
+ square_list (list[list[Tensor]]): Multi level squares of each
+ image.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ unmap_outputs (bool): Whether to map outputs back to the original
+ set of anchors. Defaults to True.
+
+ Returns:
+ tuple: Returns a tuple containing learning targets.
+
+ - labels_list (list[Tensor]): Labels of each level.
+ - label_weights_list (list[Tensor]): Label weights of each level.
+ - bbox_cls_targets_list (list[Tensor]): BBox cls targets of \
+ each level.
+ - bbox_cls_weights_list (list[Tensor]): BBox cls weights of \
+ each level.
+ - bbox_reg_targets_list (list[Tensor]): BBox reg targets of \
+ each level.
+ - bbox_reg_weights_list (list[Tensor]): BBox reg weights of \
+ each level.
+ - num_total_pos (int): Number of positive samples in all images.
+ - num_total_neg (int): Number of negative samples in all images.
+ """
+ num_imgs = len(batch_img_metas)
+ assert len(approx_list) == len(inside_flag_list) == len(
+ square_list) == num_imgs
+ # anchor number of multi levels
+ num_level_squares = [squares.size(0) for squares in square_list[0]]
+ # concat all level anchors and flags to a single tensor
+ inside_flag_flat_list = []
+ approx_flat_list = []
+ square_flat_list = []
+ for i in range(num_imgs):
+ assert len(square_list[i]) == len(inside_flag_list[i])
+ inside_flag_flat_list.append(torch.cat(inside_flag_list[i]))
+ approx_flat_list.append(torch.cat(approx_list[i]))
+ square_flat_list.append(torch.cat(square_list[i]))
+
+ # compute targets for each image
+ if batch_gt_instances_ignore is None:
+ batch_gt_instances_ignore = [None for _ in range(num_imgs)]
+ (all_labels, all_label_weights, all_bbox_cls_targets,
+ all_bbox_cls_weights, all_bbox_reg_targets, all_bbox_reg_weights,
+ pos_inds_list, neg_inds_list, sampling_results_list) = multi_apply(
+ self._get_targets_single,
+ approx_flat_list,
+ inside_flag_flat_list,
+ square_flat_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore,
+ unmap_outputs=unmap_outputs)
+
+ # sampled anchors of all images
+ avg_factor = sum(
+ [results.avg_factor for results in sampling_results_list])
+ # split targets to a list w.r.t. multiple levels
+ labels_list = images_to_levels(all_labels, num_level_squares)
+ label_weights_list = images_to_levels(all_label_weights,
+ num_level_squares)
+ bbox_cls_targets_list = images_to_levels(all_bbox_cls_targets,
+ num_level_squares)
+ bbox_cls_weights_list = images_to_levels(all_bbox_cls_weights,
+ num_level_squares)
+ bbox_reg_targets_list = images_to_levels(all_bbox_reg_targets,
+ num_level_squares)
+ bbox_reg_weights_list = images_to_levels(all_bbox_reg_weights,
+ num_level_squares)
+ return (labels_list, label_weights_list, bbox_cls_targets_list,
+ bbox_cls_weights_list, bbox_reg_targets_list,
+ bbox_reg_weights_list, avg_factor)
+
+ def _get_targets_single(self,
+ flat_approxs: Tensor,
+ inside_flags: Tensor,
+ flat_squares: Tensor,
+ gt_instances: InstanceData,
+ img_meta: dict,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ unmap_outputs: bool = True) -> tuple:
+ """Compute regression and classification targets for anchors in a
+ single image.
+
+ Args:
+ flat_approxs (Tensor): flat approxs of a single image,
+ shape (n, 4)
+ inside_flags (Tensor): inside flags of a single image,
+ shape (n, ).
+ flat_squares (Tensor): flat squares of a single image,
+ shape (approxs_per_octave * n, 4)
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes`` and ``labels``
+ attributes.
+ img_meta (dict): Meta information for current image.
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ unmap_outputs (bool): Whether to map outputs back to the original
+ set of anchors. Defaults to True.
+
+ Returns:
+ tuple:
+
+ - labels_list (Tensor): Labels in a single image.
+ - label_weights (Tensor): Label weights in a single image.
+ - bbox_cls_targets (Tensor): BBox cls targets in a single image.
+ - bbox_cls_weights (Tensor): BBox cls weights in a single image.
+ - bbox_reg_targets (Tensor): BBox reg targets in a single image.
+ - bbox_reg_weights (Tensor): BBox reg weights in a single image.
+ - num_total_pos (int): Number of positive samples in a single \
+ image.
+ - num_total_neg (int): Number of negative samples in a single \
+ image.
+ - sampling_result (:obj:`SamplingResult`): Sampling result object.
+ """
+ if not inside_flags.any():
+ raise ValueError(
+ 'There is no valid anchor inside the image boundary. Please '
+ 'check the image size and anchor sizes, or set '
+ '``allowed_border`` to -1 to skip the condition.')
+ # assign gt and sample anchors
+ num_square = flat_squares.size(0)
+ approxs = flat_approxs.view(num_square, self.approxs_per_octave, 4)
+ approxs = approxs[inside_flags, ...]
+ squares = flat_squares[inside_flags, :]
+
+ pred_instances = InstanceData()
+ pred_instances.priors = squares
+ pred_instances.approxs = approxs
+ assign_result = self.assigner.assign(pred_instances, gt_instances,
+ gt_instances_ignore)
+ sampling_result = self.sampler.sample(assign_result, pred_instances,
+ gt_instances)
+
+ num_valid_squares = squares.shape[0]
+ bbox_cls_targets = squares.new_zeros(
+ (num_valid_squares, self.side_num * 4))
+ bbox_cls_weights = squares.new_zeros(
+ (num_valid_squares, self.side_num * 4))
+ bbox_reg_targets = squares.new_zeros(
+ (num_valid_squares, self.side_num * 4))
+ bbox_reg_weights = squares.new_zeros(
+ (num_valid_squares, self.side_num * 4))
+ labels = squares.new_full((num_valid_squares, ),
+ self.num_classes,
+ dtype=torch.long)
+ label_weights = squares.new_zeros(num_valid_squares, dtype=torch.float)
+
+ pos_inds = sampling_result.pos_inds
+ neg_inds = sampling_result.neg_inds
+ if len(pos_inds) > 0:
+ (pos_bbox_reg_targets, pos_bbox_reg_weights, pos_bbox_cls_targets,
+ pos_bbox_cls_weights) = self.bbox_coder.encode(
+ sampling_result.pos_bboxes, sampling_result.pos_gt_bboxes)
+
+ bbox_cls_targets[pos_inds, :] = pos_bbox_cls_targets
+ bbox_reg_targets[pos_inds, :] = pos_bbox_reg_targets
+ bbox_cls_weights[pos_inds, :] = pos_bbox_cls_weights
+ bbox_reg_weights[pos_inds, :] = pos_bbox_reg_weights
+ labels[pos_inds] = sampling_result.pos_gt_labels
+ if self.train_cfg['pos_weight'] <= 0:
+ label_weights[pos_inds] = 1.0
+ else:
+ label_weights[pos_inds] = self.train_cfg['pos_weight']
+ if len(neg_inds) > 0:
+ label_weights[neg_inds] = 1.0
+
+ # map up to original set of anchors
+ if unmap_outputs:
+ num_total_anchors = flat_squares.size(0)
+ labels = unmap(
+ labels, num_total_anchors, inside_flags, fill=self.num_classes)
+ label_weights = unmap(label_weights, num_total_anchors,
+ inside_flags)
+ bbox_cls_targets = unmap(bbox_cls_targets, num_total_anchors,
+ inside_flags)
+ bbox_cls_weights = unmap(bbox_cls_weights, num_total_anchors,
+ inside_flags)
+ bbox_reg_targets = unmap(bbox_reg_targets, num_total_anchors,
+ inside_flags)
+ bbox_reg_weights = unmap(bbox_reg_weights, num_total_anchors,
+ inside_flags)
+ return (labels, label_weights, bbox_cls_targets, bbox_cls_weights,
+ bbox_reg_targets, bbox_reg_weights, pos_inds, neg_inds,
+ sampling_result)
+
+ def loss_by_feat_single(self, cls_score: Tensor, bbox_pred: Tensor,
+ labels: Tensor, label_weights: Tensor,
+ bbox_cls_targets: Tensor, bbox_cls_weights: Tensor,
+ bbox_reg_targets: Tensor, bbox_reg_weights: Tensor,
+ avg_factor: float) -> Tuple[Tensor]:
+ """Calculate the loss of a single scale level based on the features
+ extracted by the detection head.
+
+ Args:
+ cls_score (Tensor): Box scores for each scale level
+ Has shape (N, num_anchors * num_classes, H, W).
+ bbox_pred (Tensor): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W).
+ labels (Tensor): Labels in a single image.
+ label_weights (Tensor): Label weights in a single level.
+ bbox_cls_targets (Tensor): BBox cls targets in a single level.
+ bbox_cls_weights (Tensor): BBox cls weights in a single level.
+ bbox_reg_targets (Tensor): BBox reg targets in a single level.
+ bbox_reg_weights (Tensor): BBox reg weights in a single level.
+ avg_factor (int): Average factor that is used to average the loss.
+
+ Returns:
+ tuple: loss components.
+ """
+ # classification loss
+ labels = labels.reshape(-1)
+ label_weights = label_weights.reshape(-1)
+ cls_score = cls_score.permute(0, 2, 3,
+ 1).reshape(-1, self.cls_out_channels)
+ loss_cls = self.loss_cls(
+ cls_score, labels, label_weights, avg_factor=avg_factor)
+ # regression loss
+ bbox_cls_targets = bbox_cls_targets.reshape(-1, self.side_num * 4)
+ bbox_cls_weights = bbox_cls_weights.reshape(-1, self.side_num * 4)
+ bbox_reg_targets = bbox_reg_targets.reshape(-1, self.side_num * 4)
+ bbox_reg_weights = bbox_reg_weights.reshape(-1, self.side_num * 4)
+ (bbox_cls_pred, bbox_reg_pred) = bbox_pred
+ bbox_cls_pred = bbox_cls_pred.permute(0, 2, 3, 1).reshape(
+ -1, self.side_num * 4)
+ bbox_reg_pred = bbox_reg_pred.permute(0, 2, 3, 1).reshape(
+ -1, self.side_num * 4)
+ loss_bbox_cls = self.loss_bbox_cls(
+ bbox_cls_pred,
+ bbox_cls_targets.long(),
+ bbox_cls_weights,
+ avg_factor=avg_factor * 4 * self.side_num)
+ loss_bbox_reg = self.loss_bbox_reg(
+ bbox_reg_pred,
+ bbox_reg_targets,
+ bbox_reg_weights,
+ avg_factor=avg_factor * 4 * self.bbox_coder.offset_topk)
+ return loss_cls, loss_bbox_cls, loss_bbox_reg
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ has shape (N, num_anchors * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ assert len(featmap_sizes) == self.approx_anchor_generator.num_levels
+
+ device = cls_scores[0].device
+
+ # get sampled approxes
+ approxs_list, inside_flag_list = GuidedAnchorHead.get_sampled_approxs(
+ self, featmap_sizes, batch_img_metas, device=device)
+
+ square_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+
+ cls_reg_targets = self.get_targets(
+ approxs_list,
+ inside_flag_list,
+ square_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore)
+ (labels_list, label_weights_list, bbox_cls_targets_list,
+ bbox_cls_weights_list, bbox_reg_targets_list, bbox_reg_weights_list,
+ avg_factor) = cls_reg_targets
+
+ losses_cls, losses_bbox_cls, losses_bbox_reg = multi_apply(
+ self.loss_by_feat_single,
+ cls_scores,
+ bbox_preds,
+ labels_list,
+ label_weights_list,
+ bbox_cls_targets_list,
+ bbox_cls_weights_list,
+ bbox_reg_targets_list,
+ bbox_reg_weights_list,
+ avg_factor=avg_factor)
+ return dict(
+ loss_cls=losses_cls,
+ loss_bbox_cls=losses_bbox_cls,
+ loss_bbox_reg=losses_bbox_reg)
+
+ def predict_by_feat(self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_img_metas: List[dict],
+ cfg: Optional[ConfigDict] = None,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ bbox results.
+
+ Note: When score_factors is not None, the cls_scores are
+ usually multiplied by it then obtain the real score used in NMS,
+ such as CenterNess in FCOS, IoU branch in ATSS.
+
+ Args:
+ cls_scores (list[Tensor]): Classification scores for all
+ scale levels, each is a 4D-tensor, has shape
+ (batch_size, num_priors * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas for all
+ scale levels, each is a 4D-tensor, has shape
+ (batch_size, num_priors * 4, H, W).
+ batch_img_metas (list[dict], Optional): Batch image meta info.
+ cfg (:obj:`ConfigDict`, optional): Test / postprocessing
+ configuration, if None, test_cfg would be used.
+ Defaults to None.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ list[:obj:`InstanceData`]: Object detection results of each image
+ after the post process. Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ assert len(cls_scores) == len(bbox_preds)
+ num_levels = len(cls_scores)
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+
+ device = cls_scores[0].device
+ mlvl_anchors = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+ result_list = []
+ for img_id in range(len(batch_img_metas)):
+ cls_score_list = [
+ cls_scores[i][img_id].detach() for i in range(num_levels)
+ ]
+ bbox_cls_pred_list = [
+ bbox_preds[i][0][img_id].detach() for i in range(num_levels)
+ ]
+ bbox_reg_pred_list = [
+ bbox_preds[i][1][img_id].detach() for i in range(num_levels)
+ ]
+ proposals = self._predict_by_feat_single(
+ cls_scores=cls_score_list,
+ bbox_cls_preds=bbox_cls_pred_list,
+ bbox_reg_preds=bbox_reg_pred_list,
+ mlvl_anchors=mlvl_anchors[img_id],
+ img_meta=batch_img_metas[img_id],
+ cfg=cfg,
+ rescale=rescale,
+ with_nms=with_nms)
+ result_list.append(proposals)
+ return result_list
+
+ def _predict_by_feat_single(self,
+ cls_scores: List[Tensor],
+ bbox_cls_preds: List[Tensor],
+ bbox_reg_preds: List[Tensor],
+ mlvl_anchors: List[Tensor],
+ img_meta: dict,
+ cfg: ConfigDict,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceData:
+ cfg = self.test_cfg if cfg is None else cfg
+ nms_pre = cfg.get('nms_pre', -1)
+
+ mlvl_bboxes = []
+ mlvl_scores = []
+ mlvl_confids = []
+ mlvl_labels = []
+ assert len(cls_scores) == len(bbox_cls_preds) == len(
+ bbox_reg_preds) == len(mlvl_anchors)
+ for cls_score, bbox_cls_pred, bbox_reg_pred, anchors in zip(
+ cls_scores, bbox_cls_preds, bbox_reg_preds, mlvl_anchors):
+ assert cls_score.size()[-2:] == bbox_cls_pred.size(
+ )[-2:] == bbox_reg_pred.size()[-2::]
+ cls_score = cls_score.permute(1, 2,
+ 0).reshape(-1, self.cls_out_channels)
+ if self.use_sigmoid_cls:
+ scores = cls_score.sigmoid()
+ else:
+ scores = cls_score.softmax(-1)[:, :-1]
+ bbox_cls_pred = bbox_cls_pred.permute(1, 2, 0).reshape(
+ -1, self.side_num * 4)
+ bbox_reg_pred = bbox_reg_pred.permute(1, 2, 0).reshape(
+ -1, self.side_num * 4)
+
+ # After https://github.com/open-mmlab/mmdetection/pull/6268/,
+ # this operation keeps fewer bboxes under the same `nms_pre`.
+ # There is no difference in performance for most models. If you
+ # find a slight drop in performance, you can set a larger
+ # `nms_pre` than before.
+ results = filter_scores_and_topk(
+ scores, cfg.score_thr, nms_pre,
+ dict(
+ anchors=anchors,
+ bbox_cls_pred=bbox_cls_pred,
+ bbox_reg_pred=bbox_reg_pred))
+ scores, labels, _, filtered_results = results
+
+ anchors = filtered_results['anchors']
+ bbox_cls_pred = filtered_results['bbox_cls_pred']
+ bbox_reg_pred = filtered_results['bbox_reg_pred']
+
+ bbox_preds = [
+ bbox_cls_pred.contiguous(),
+ bbox_reg_pred.contiguous()
+ ]
+ bboxes, confids = self.bbox_coder.decode(
+ anchors.contiguous(),
+ bbox_preds,
+ max_shape=img_meta['img_shape'])
+
+ mlvl_bboxes.append(bboxes)
+ mlvl_scores.append(scores)
+ mlvl_confids.append(confids)
+ mlvl_labels.append(labels)
+
+ results = InstanceData()
+ results.bboxes = torch.cat(mlvl_bboxes)
+ results.scores = torch.cat(mlvl_scores)
+ results.score_factors = torch.cat(mlvl_confids)
+ results.labels = torch.cat(mlvl_labels)
+
+ return self._bbox_post_process(
+ results=results,
+ cfg=cfg,
+ rescale=rescale,
+ with_nms=with_nms,
+ img_meta=img_meta)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/solo_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/solo_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..8cf338451358b01899faa4b299d33fafd7262d21
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/solo_head.py
@@ -0,0 +1,1263 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple
+
+import mmcv
+import numpy as np
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.models.utils.misc import floordiv
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, InstanceList, MultiConfig, OptConfigType
+from ..layers import mask_matrix_nms
+from ..utils import center_of_mass, generate_coordinate, multi_apply
+from .base_mask_head import BaseMaskHead
+
+
+@MODELS.register_module()
+class SOLOHead(BaseMaskHead):
+ """SOLO mask head used in `SOLO: Segmenting Objects by Locations.
+
+ `_
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ feat_channels (int): Number of hidden channels. Used in child classes.
+ Defaults to 256.
+ stacked_convs (int): Number of stacking convs of the head.
+ Defaults to 4.
+ strides (tuple): Downsample factor of each feature map.
+ scale_ranges (tuple[tuple[int, int]]): Area range of multiple
+ level masks, in the format [(min1, max1), (min2, max2), ...].
+ A range of (16, 64) means the area range between (16, 64).
+ pos_scale (float): Constant scale factor to control the center region.
+ num_grids (list[int]): Divided image into a uniform grids, each
+ feature map has a different grid value. The number of output
+ channels is grid ** 2. Defaults to [40, 36, 24, 16, 12].
+ cls_down_index (int): The index of downsample operation in
+ classification branch. Defaults to 0.
+ loss_mask (dict): Config of mask loss.
+ loss_cls (dict): Config of classification loss.
+ norm_cfg (dict): Dictionary to construct and config norm layer.
+ Defaults to norm_cfg=dict(type='GN', num_groups=32,
+ requires_grad=True).
+ train_cfg (dict): Training config of head.
+ test_cfg (dict): Testing config of head.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ """
+
+ def __init__(
+ self,
+ num_classes: int,
+ in_channels: int,
+ feat_channels: int = 256,
+ stacked_convs: int = 4,
+ strides: tuple = (4, 8, 16, 32, 64),
+ scale_ranges: tuple = ((8, 32), (16, 64), (32, 128), (64, 256), (128,
+ 512)),
+ pos_scale: float = 0.2,
+ num_grids: list = [40, 36, 24, 16, 12],
+ cls_down_index: int = 0,
+ loss_mask: ConfigType = dict(
+ type='DiceLoss', use_sigmoid=True, loss_weight=3.0),
+ loss_cls: ConfigType = dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ norm_cfg: ConfigType = dict(
+ type='GN', num_groups=32, requires_grad=True),
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ init_cfg: MultiConfig = [
+ dict(type='Normal', layer='Conv2d', std=0.01),
+ dict(
+ type='Normal',
+ std=0.01,
+ bias_prob=0.01,
+ override=dict(name='conv_mask_list')),
+ dict(
+ type='Normal',
+ std=0.01,
+ bias_prob=0.01,
+ override=dict(name='conv_cls'))
+ ]
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.num_classes = num_classes
+ self.cls_out_channels = self.num_classes
+ self.in_channels = in_channels
+ self.feat_channels = feat_channels
+ self.stacked_convs = stacked_convs
+ self.strides = strides
+ self.num_grids = num_grids
+ # number of FPN feats
+ self.num_levels = len(strides)
+ assert self.num_levels == len(scale_ranges) == len(num_grids)
+ self.scale_ranges = scale_ranges
+ self.pos_scale = pos_scale
+
+ self.cls_down_index = cls_down_index
+ self.loss_cls = MODELS.build(loss_cls)
+ self.loss_mask = MODELS.build(loss_mask)
+ self.norm_cfg = norm_cfg
+ self.init_cfg = init_cfg
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+ self._init_layers()
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self.mask_convs = nn.ModuleList()
+ self.cls_convs = nn.ModuleList()
+ for i in range(self.stacked_convs):
+ chn = self.in_channels + 2 if i == 0 else self.feat_channels
+ self.mask_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ norm_cfg=self.norm_cfg))
+ chn = self.in_channels if i == 0 else self.feat_channels
+ self.cls_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ norm_cfg=self.norm_cfg))
+ self.conv_mask_list = nn.ModuleList()
+ for num_grid in self.num_grids:
+ self.conv_mask_list.append(
+ nn.Conv2d(self.feat_channels, num_grid**2, 1))
+
+ self.conv_cls = nn.Conv2d(
+ self.feat_channels, self.cls_out_channels, 3, padding=1)
+
+ def resize_feats(self, x: Tuple[Tensor]) -> List[Tensor]:
+ """Downsample the first feat and upsample last feat in feats.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ list[Tensor]: Features after resizing, each is a 4D-tensor.
+ """
+ out = []
+ for i in range(len(x)):
+ if i == 0:
+ out.append(
+ F.interpolate(x[0], scale_factor=0.5, mode='bilinear'))
+ elif i == len(x) - 1:
+ out.append(
+ F.interpolate(
+ x[i], size=x[i - 1].shape[-2:], mode='bilinear'))
+ else:
+ out.append(x[i])
+ return out
+
+ def forward(self, x: Tuple[Tensor]) -> tuple:
+ """Forward features from the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: A tuple of classification scores and mask prediction.
+
+ - mlvl_mask_preds (list[Tensor]): Multi-level mask prediction.
+ Each element in the list has shape
+ (batch_size, num_grids**2 ,h ,w).
+ - mlvl_cls_preds (list[Tensor]): Multi-level scores.
+ Each element in the list has shape
+ (batch_size, num_classes, num_grids ,num_grids).
+ """
+ assert len(x) == self.num_levels
+ feats = self.resize_feats(x)
+ mlvl_mask_preds = []
+ mlvl_cls_preds = []
+ for i in range(self.num_levels):
+ x = feats[i]
+ mask_feat = x
+ cls_feat = x
+ # generate and concat the coordinate
+ coord_feat = generate_coordinate(mask_feat.size(),
+ mask_feat.device)
+ mask_feat = torch.cat([mask_feat, coord_feat], 1)
+
+ for mask_layer in (self.mask_convs):
+ mask_feat = mask_layer(mask_feat)
+
+ mask_feat = F.interpolate(
+ mask_feat, scale_factor=2, mode='bilinear')
+ mask_preds = self.conv_mask_list[i](mask_feat)
+
+ # cls branch
+ for j, cls_layer in enumerate(self.cls_convs):
+ if j == self.cls_down_index:
+ num_grid = self.num_grids[i]
+ cls_feat = F.interpolate(
+ cls_feat, size=num_grid, mode='bilinear')
+ cls_feat = cls_layer(cls_feat)
+
+ cls_pred = self.conv_cls(cls_feat)
+
+ if not self.training:
+ feat_wh = feats[0].size()[-2:]
+ upsampled_size = (feat_wh[0] * 2, feat_wh[1] * 2)
+ mask_preds = F.interpolate(
+ mask_preds.sigmoid(), size=upsampled_size, mode='bilinear')
+ cls_pred = cls_pred.sigmoid()
+ # get local maximum
+ local_max = F.max_pool2d(cls_pred, 2, stride=1, padding=1)
+ keep_mask = local_max[:, :, :-1, :-1] == cls_pred
+ cls_pred = cls_pred * keep_mask
+
+ mlvl_mask_preds.append(mask_preds)
+ mlvl_cls_preds.append(cls_pred)
+ return mlvl_mask_preds, mlvl_cls_preds
+
+ def loss_by_feat(self, mlvl_mask_preds: List[Tensor],
+ mlvl_cls_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict], **kwargs) -> dict:
+ """Calculate the loss based on the features extracted by the mask head.
+
+ Args:
+ mlvl_mask_preds (list[Tensor]): Multi-level mask prediction.
+ Each element in the list has shape
+ (batch_size, num_grids**2 ,h ,w).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``masks``,
+ and ``labels`` attributes.
+ batch_img_metas (list[dict]): Meta information of multiple images.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ num_levels = self.num_levels
+ num_imgs = len(batch_img_metas)
+
+ featmap_sizes = [featmap.size()[-2:] for featmap in mlvl_mask_preds]
+
+ # `BoolTensor` in `pos_masks` represent
+ # whether the corresponding point is
+ # positive
+ pos_mask_targets, labels, pos_masks = multi_apply(
+ self._get_targets_single,
+ batch_gt_instances,
+ featmap_sizes=featmap_sizes)
+
+ # change from the outside list meaning multi images
+ # to the outside list meaning multi levels
+ mlvl_pos_mask_targets = [[] for _ in range(num_levels)]
+ mlvl_pos_mask_preds = [[] for _ in range(num_levels)]
+ mlvl_pos_masks = [[] for _ in range(num_levels)]
+ mlvl_labels = [[] for _ in range(num_levels)]
+ for img_id in range(num_imgs):
+ assert num_levels == len(pos_mask_targets[img_id])
+ for lvl in range(num_levels):
+ mlvl_pos_mask_targets[lvl].append(
+ pos_mask_targets[img_id][lvl])
+ mlvl_pos_mask_preds[lvl].append(
+ mlvl_mask_preds[lvl][img_id, pos_masks[img_id][lvl], ...])
+ mlvl_pos_masks[lvl].append(pos_masks[img_id][lvl].flatten())
+ mlvl_labels[lvl].append(labels[img_id][lvl].flatten())
+
+ # cat multiple image
+ temp_mlvl_cls_preds = []
+ for lvl in range(num_levels):
+ mlvl_pos_mask_targets[lvl] = torch.cat(
+ mlvl_pos_mask_targets[lvl], dim=0)
+ mlvl_pos_mask_preds[lvl] = torch.cat(
+ mlvl_pos_mask_preds[lvl], dim=0)
+ mlvl_pos_masks[lvl] = torch.cat(mlvl_pos_masks[lvl], dim=0)
+ mlvl_labels[lvl] = torch.cat(mlvl_labels[lvl], dim=0)
+ temp_mlvl_cls_preds.append(mlvl_cls_preds[lvl].permute(
+ 0, 2, 3, 1).reshape(-1, self.cls_out_channels))
+
+ num_pos = sum(item.sum() for item in mlvl_pos_masks)
+ # dice loss
+ loss_mask = []
+ for pred, target in zip(mlvl_pos_mask_preds, mlvl_pos_mask_targets):
+ if pred.size()[0] == 0:
+ loss_mask.append(pred.sum().unsqueeze(0))
+ continue
+ loss_mask.append(
+ self.loss_mask(pred, target, reduction_override='none'))
+ if num_pos > 0:
+ loss_mask = torch.cat(loss_mask).sum() / num_pos
+ else:
+ loss_mask = torch.cat(loss_mask).mean()
+
+ flatten_labels = torch.cat(mlvl_labels)
+ flatten_cls_preds = torch.cat(temp_mlvl_cls_preds)
+ loss_cls = self.loss_cls(
+ flatten_cls_preds, flatten_labels, avg_factor=num_pos + 1)
+ return dict(loss_mask=loss_mask, loss_cls=loss_cls)
+
+ def _get_targets_single(self,
+ gt_instances: InstanceData,
+ featmap_sizes: Optional[list] = None) -> tuple:
+ """Compute targets for predictions of single image.
+
+ Args:
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes``, ``labels``,
+ and ``masks`` attributes.
+ featmap_sizes (list[:obj:`torch.size`]): Size of each
+ feature map from feature pyramid, each element
+ means (feat_h, feat_w). Defaults to None.
+
+ Returns:
+ Tuple: Usually returns a tuple containing targets for predictions.
+
+ - mlvl_pos_mask_targets (list[Tensor]): Each element represent
+ the binary mask targets for positive points in this
+ level, has shape (num_pos, out_h, out_w).
+ - mlvl_labels (list[Tensor]): Each element is
+ classification labels for all
+ points in this level, has shape
+ (num_grid, num_grid).
+ - mlvl_pos_masks (list[Tensor]): Each element is
+ a `BoolTensor` to represent whether the
+ corresponding point in single level
+ is positive, has shape (num_grid **2).
+ """
+ gt_labels = gt_instances.labels
+ device = gt_labels.device
+
+ gt_bboxes = gt_instances.bboxes
+ gt_areas = torch.sqrt((gt_bboxes[:, 2] - gt_bboxes[:, 0]) *
+ (gt_bboxes[:, 3] - gt_bboxes[:, 1]))
+
+ gt_masks = gt_instances.masks.to_tensor(
+ dtype=torch.bool, device=device)
+
+ mlvl_pos_mask_targets = []
+ mlvl_labels = []
+ mlvl_pos_masks = []
+ for (lower_bound, upper_bound), stride, featmap_size, num_grid \
+ in zip(self.scale_ranges, self.strides,
+ featmap_sizes, self.num_grids):
+
+ mask_target = torch.zeros(
+ [num_grid**2, featmap_size[0], featmap_size[1]],
+ dtype=torch.uint8,
+ device=device)
+ # FG cat_id: [0, num_classes -1], BG cat_id: num_classes
+ labels = torch.zeros([num_grid, num_grid],
+ dtype=torch.int64,
+ device=device) + self.num_classes
+ pos_mask = torch.zeros([num_grid**2],
+ dtype=torch.bool,
+ device=device)
+
+ gt_inds = ((gt_areas >= lower_bound) &
+ (gt_areas <= upper_bound)).nonzero().flatten()
+ if len(gt_inds) == 0:
+ mlvl_pos_mask_targets.append(
+ mask_target.new_zeros(0, featmap_size[0], featmap_size[1]))
+ mlvl_labels.append(labels)
+ mlvl_pos_masks.append(pos_mask)
+ continue
+ hit_gt_bboxes = gt_bboxes[gt_inds]
+ hit_gt_labels = gt_labels[gt_inds]
+ hit_gt_masks = gt_masks[gt_inds, ...]
+
+ pos_w_ranges = 0.5 * (hit_gt_bboxes[:, 2] -
+ hit_gt_bboxes[:, 0]) * self.pos_scale
+ pos_h_ranges = 0.5 * (hit_gt_bboxes[:, 3] -
+ hit_gt_bboxes[:, 1]) * self.pos_scale
+
+ # Make sure hit_gt_masks has a value
+ valid_mask_flags = hit_gt_masks.sum(dim=-1).sum(dim=-1) > 0
+ output_stride = stride / 2
+
+ for gt_mask, gt_label, pos_h_range, pos_w_range, \
+ valid_mask_flag in \
+ zip(hit_gt_masks, hit_gt_labels, pos_h_ranges,
+ pos_w_ranges, valid_mask_flags):
+ if not valid_mask_flag:
+ continue
+ upsampled_size = (featmap_sizes[0][0] * 4,
+ featmap_sizes[0][1] * 4)
+ center_h, center_w = center_of_mass(gt_mask)
+
+ coord_w = int(
+ floordiv((center_w / upsampled_size[1]), (1. / num_grid),
+ rounding_mode='trunc'))
+ coord_h = int(
+ floordiv((center_h / upsampled_size[0]), (1. / num_grid),
+ rounding_mode='trunc'))
+
+ # left, top, right, down
+ top_box = max(
+ 0,
+ int(
+ floordiv(
+ (center_h - pos_h_range) / upsampled_size[0],
+ (1. / num_grid),
+ rounding_mode='trunc')))
+ down_box = min(
+ num_grid - 1,
+ int(
+ floordiv(
+ (center_h + pos_h_range) / upsampled_size[0],
+ (1. / num_grid),
+ rounding_mode='trunc')))
+ left_box = max(
+ 0,
+ int(
+ floordiv(
+ (center_w - pos_w_range) / upsampled_size[1],
+ (1. / num_grid),
+ rounding_mode='trunc')))
+ right_box = min(
+ num_grid - 1,
+ int(
+ floordiv(
+ (center_w + pos_w_range) / upsampled_size[1],
+ (1. / num_grid),
+ rounding_mode='trunc')))
+
+ top = max(top_box, coord_h - 1)
+ down = min(down_box, coord_h + 1)
+ left = max(coord_w - 1, left_box)
+ right = min(right_box, coord_w + 1)
+
+ labels[top:(down + 1), left:(right + 1)] = gt_label
+ # ins
+ gt_mask = np.uint8(gt_mask.cpu().numpy())
+ # Follow the original implementation, F.interpolate is
+ # different from cv2 and opencv
+ gt_mask = mmcv.imrescale(gt_mask, scale=1. / output_stride)
+ gt_mask = torch.from_numpy(gt_mask).to(device=device)
+
+ for i in range(top, down + 1):
+ for j in range(left, right + 1):
+ index = int(i * num_grid + j)
+ mask_target[index, :gt_mask.shape[0], :gt_mask.
+ shape[1]] = gt_mask
+ pos_mask[index] = True
+ mlvl_pos_mask_targets.append(mask_target[pos_mask])
+ mlvl_labels.append(labels)
+ mlvl_pos_masks.append(pos_mask)
+ return mlvl_pos_mask_targets, mlvl_labels, mlvl_pos_masks
+
+ def predict_by_feat(self, mlvl_mask_preds: List[Tensor],
+ mlvl_cls_scores: List[Tensor],
+ batch_img_metas: List[dict], **kwargs) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ mask results.
+
+ Args:
+ mlvl_mask_preds (list[Tensor]): Multi-level mask prediction.
+ Each element in the list has shape
+ (batch_size, num_grids**2 ,h ,w).
+ mlvl_cls_scores (list[Tensor]): Multi-level scores. Each element
+ in the list has shape
+ (batch_size, num_classes, num_grids ,num_grids).
+ batch_img_metas (list[dict]): Meta information of all images.
+
+ Returns:
+ list[:obj:`InstanceData`]: Processed results of multiple
+ images.Each :obj:`InstanceData` usually contains
+ following keys.
+
+ - scores (Tensor): Classification scores, has shape
+ (num_instance,).
+ - labels (Tensor): Has shape (num_instances,).
+ - masks (Tensor): Processed mask results, has
+ shape (num_instances, h, w).
+ """
+ mlvl_cls_scores = [
+ item.permute(0, 2, 3, 1) for item in mlvl_cls_scores
+ ]
+ assert len(mlvl_mask_preds) == len(mlvl_cls_scores)
+ num_levels = len(mlvl_cls_scores)
+
+ results_list = []
+ for img_id in range(len(batch_img_metas)):
+ cls_pred_list = [
+ mlvl_cls_scores[lvl][img_id].view(-1, self.cls_out_channels)
+ for lvl in range(num_levels)
+ ]
+ mask_pred_list = [
+ mlvl_mask_preds[lvl][img_id] for lvl in range(num_levels)
+ ]
+
+ cls_pred_list = torch.cat(cls_pred_list, dim=0)
+ mask_pred_list = torch.cat(mask_pred_list, dim=0)
+ img_meta = batch_img_metas[img_id]
+
+ results = self._predict_by_feat_single(
+ cls_pred_list, mask_pred_list, img_meta=img_meta)
+ results_list.append(results)
+
+ return results_list
+
+ def _predict_by_feat_single(self,
+ cls_scores: Tensor,
+ mask_preds: Tensor,
+ img_meta: dict,
+ cfg: OptConfigType = None) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ mask results.
+
+ Args:
+ cls_scores (Tensor): Classification score of all points
+ in single image, has shape (num_points, num_classes).
+ mask_preds (Tensor): Mask prediction of all points in
+ single image, has shape (num_points, feat_h, feat_w).
+ img_meta (dict): Meta information of corresponding image.
+ cfg (dict, optional): Config used in test phase.
+ Defaults to None.
+
+ Returns:
+ :obj:`InstanceData`: Processed results of single image.
+ it usually contains following keys.
+
+ - scores (Tensor): Classification scores, has shape
+ (num_instance,).
+ - labels (Tensor): Has shape (num_instances,).
+ - masks (Tensor): Processed mask results, has
+ shape (num_instances, h, w).
+ """
+
+ def empty_results(cls_scores, ori_shape):
+ """Generate a empty results."""
+ results = InstanceData()
+ results.scores = cls_scores.new_ones(0)
+ results.masks = cls_scores.new_zeros(0, *ori_shape)
+ results.labels = cls_scores.new_ones(0)
+ results.bboxes = cls_scores.new_zeros(0, 4)
+ return results
+
+ cfg = self.test_cfg if cfg is None else cfg
+ assert len(cls_scores) == len(mask_preds)
+
+ featmap_size = mask_preds.size()[-2:]
+
+ h, w = img_meta['img_shape'][:2]
+ upsampled_size = (featmap_size[0] * 4, featmap_size[1] * 4)
+
+ score_mask = (cls_scores > cfg.score_thr)
+ cls_scores = cls_scores[score_mask]
+ if len(cls_scores) == 0:
+ return empty_results(cls_scores, img_meta['ori_shape'][:2])
+
+ inds = score_mask.nonzero()
+ cls_labels = inds[:, 1]
+
+ # Filter the mask mask with an area is smaller than
+ # stride of corresponding feature level
+ lvl_interval = cls_labels.new_tensor(self.num_grids).pow(2).cumsum(0)
+ strides = cls_scores.new_ones(lvl_interval[-1])
+ strides[:lvl_interval[0]] *= self.strides[0]
+ for lvl in range(1, self.num_levels):
+ strides[lvl_interval[lvl -
+ 1]:lvl_interval[lvl]] *= self.strides[lvl]
+ strides = strides[inds[:, 0]]
+ mask_preds = mask_preds[inds[:, 0]]
+
+ masks = mask_preds > cfg.mask_thr
+ sum_masks = masks.sum((1, 2)).float()
+ keep = sum_masks > strides
+ if keep.sum() == 0:
+ return empty_results(cls_scores, img_meta['ori_shape'][:2])
+ masks = masks[keep]
+ mask_preds = mask_preds[keep]
+ sum_masks = sum_masks[keep]
+ cls_scores = cls_scores[keep]
+ cls_labels = cls_labels[keep]
+
+ # maskness.
+ mask_scores = (mask_preds * masks).sum((1, 2)) / sum_masks
+ cls_scores *= mask_scores
+
+ scores, labels, _, keep_inds = mask_matrix_nms(
+ masks,
+ cls_labels,
+ cls_scores,
+ mask_area=sum_masks,
+ nms_pre=cfg.nms_pre,
+ max_num=cfg.max_per_img,
+ kernel=cfg.kernel,
+ sigma=cfg.sigma,
+ filter_thr=cfg.filter_thr)
+ # mask_matrix_nms may return an empty Tensor
+ if len(keep_inds) == 0:
+ return empty_results(cls_scores, img_meta['ori_shape'][:2])
+ mask_preds = mask_preds[keep_inds]
+ mask_preds = F.interpolate(
+ mask_preds.unsqueeze(0), size=upsampled_size,
+ mode='bilinear')[:, :, :h, :w]
+ mask_preds = F.interpolate(
+ mask_preds, size=img_meta['ori_shape'][:2],
+ mode='bilinear').squeeze(0)
+ masks = mask_preds > cfg.mask_thr
+
+ results = InstanceData()
+ results.masks = masks
+ results.labels = labels
+ results.scores = scores
+ # create an empty bbox in InstanceData to avoid bugs when
+ # calculating metrics.
+ results.bboxes = results.scores.new_zeros(len(scores), 4)
+ return results
+
+
+@MODELS.register_module()
+class DecoupledSOLOHead(SOLOHead):
+ """Decoupled SOLO mask head used in `SOLO: Segmenting Objects by Locations.
+
+ `_
+
+ Args:
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ """
+
+ def __init__(self,
+ *args,
+ init_cfg: MultiConfig = [
+ dict(type='Normal', layer='Conv2d', std=0.01),
+ dict(
+ type='Normal',
+ std=0.01,
+ bias_prob=0.01,
+ override=dict(name='conv_mask_list_x')),
+ dict(
+ type='Normal',
+ std=0.01,
+ bias_prob=0.01,
+ override=dict(name='conv_mask_list_y')),
+ dict(
+ type='Normal',
+ std=0.01,
+ bias_prob=0.01,
+ override=dict(name='conv_cls'))
+ ],
+ **kwargs) -> None:
+ super().__init__(*args, init_cfg=init_cfg, **kwargs)
+
+ def _init_layers(self) -> None:
+ self.mask_convs_x = nn.ModuleList()
+ self.mask_convs_y = nn.ModuleList()
+ self.cls_convs = nn.ModuleList()
+
+ for i in range(self.stacked_convs):
+ chn = self.in_channels + 1 if i == 0 else self.feat_channels
+ self.mask_convs_x.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ norm_cfg=self.norm_cfg))
+ self.mask_convs_y.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ norm_cfg=self.norm_cfg))
+
+ chn = self.in_channels if i == 0 else self.feat_channels
+ self.cls_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ norm_cfg=self.norm_cfg))
+
+ self.conv_mask_list_x = nn.ModuleList()
+ self.conv_mask_list_y = nn.ModuleList()
+ for num_grid in self.num_grids:
+ self.conv_mask_list_x.append(
+ nn.Conv2d(self.feat_channels, num_grid, 3, padding=1))
+ self.conv_mask_list_y.append(
+ nn.Conv2d(self.feat_channels, num_grid, 3, padding=1))
+ self.conv_cls = nn.Conv2d(
+ self.feat_channels, self.cls_out_channels, 3, padding=1)
+
+ def forward(self, x: Tuple[Tensor]) -> Tuple:
+ """Forward features from the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: A tuple of classification scores and mask prediction.
+
+ - mlvl_mask_preds_x (list[Tensor]): Multi-level mask prediction
+ from x branch. Each element in the list has shape
+ (batch_size, num_grids ,h ,w).
+ - mlvl_mask_preds_y (list[Tensor]): Multi-level mask prediction
+ from y branch. Each element in the list has shape
+ (batch_size, num_grids ,h ,w).
+ - mlvl_cls_preds (list[Tensor]): Multi-level scores.
+ Each element in the list has shape
+ (batch_size, num_classes, num_grids ,num_grids).
+ """
+ assert len(x) == self.num_levels
+ feats = self.resize_feats(x)
+ mask_preds_x = []
+ mask_preds_y = []
+ cls_preds = []
+ for i in range(self.num_levels):
+ x = feats[i]
+ mask_feat = x
+ cls_feat = x
+ # generate and concat the coordinate
+ coord_feat = generate_coordinate(mask_feat.size(),
+ mask_feat.device)
+ mask_feat_x = torch.cat([mask_feat, coord_feat[:, 0:1, ...]], 1)
+ mask_feat_y = torch.cat([mask_feat, coord_feat[:, 1:2, ...]], 1)
+
+ for mask_layer_x, mask_layer_y in \
+ zip(self.mask_convs_x, self.mask_convs_y):
+ mask_feat_x = mask_layer_x(mask_feat_x)
+ mask_feat_y = mask_layer_y(mask_feat_y)
+
+ mask_feat_x = F.interpolate(
+ mask_feat_x, scale_factor=2, mode='bilinear')
+ mask_feat_y = F.interpolate(
+ mask_feat_y, scale_factor=2, mode='bilinear')
+
+ mask_pred_x = self.conv_mask_list_x[i](mask_feat_x)
+ mask_pred_y = self.conv_mask_list_y[i](mask_feat_y)
+
+ # cls branch
+ for j, cls_layer in enumerate(self.cls_convs):
+ if j == self.cls_down_index:
+ num_grid = self.num_grids[i]
+ cls_feat = F.interpolate(
+ cls_feat, size=num_grid, mode='bilinear')
+ cls_feat = cls_layer(cls_feat)
+
+ cls_pred = self.conv_cls(cls_feat)
+
+ if not self.training:
+ feat_wh = feats[0].size()[-2:]
+ upsampled_size = (feat_wh[0] * 2, feat_wh[1] * 2)
+ mask_pred_x = F.interpolate(
+ mask_pred_x.sigmoid(),
+ size=upsampled_size,
+ mode='bilinear')
+ mask_pred_y = F.interpolate(
+ mask_pred_y.sigmoid(),
+ size=upsampled_size,
+ mode='bilinear')
+ cls_pred = cls_pred.sigmoid()
+ # get local maximum
+ local_max = F.max_pool2d(cls_pred, 2, stride=1, padding=1)
+ keep_mask = local_max[:, :, :-1, :-1] == cls_pred
+ cls_pred = cls_pred * keep_mask
+
+ mask_preds_x.append(mask_pred_x)
+ mask_preds_y.append(mask_pred_y)
+ cls_preds.append(cls_pred)
+ return mask_preds_x, mask_preds_y, cls_preds
+
+ def loss_by_feat(self, mlvl_mask_preds_x: List[Tensor],
+ mlvl_mask_preds_y: List[Tensor],
+ mlvl_cls_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict], **kwargs) -> dict:
+ """Calculate the loss based on the features extracted by the mask head.
+
+ Args:
+ mlvl_mask_preds_x (list[Tensor]): Multi-level mask prediction
+ from x branch. Each element in the list has shape
+ (batch_size, num_grids ,h ,w).
+ mlvl_mask_preds_y (list[Tensor]): Multi-level mask prediction
+ from y branch. Each element in the list has shape
+ (batch_size, num_grids ,h ,w).
+ mlvl_cls_preds (list[Tensor]): Multi-level scores. Each element
+ in the list has shape
+ (batch_size, num_classes, num_grids ,num_grids).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``masks``,
+ and ``labels`` attributes.
+ batch_img_metas (list[dict]): Meta information of multiple images.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ num_levels = self.num_levels
+ num_imgs = len(batch_img_metas)
+ featmap_sizes = [featmap.size()[-2:] for featmap in mlvl_mask_preds_x]
+
+ pos_mask_targets, labels, xy_pos_indexes = multi_apply(
+ self._get_targets_single,
+ batch_gt_instances,
+ featmap_sizes=featmap_sizes)
+
+ # change from the outside list meaning multi images
+ # to the outside list meaning multi levels
+ mlvl_pos_mask_targets = [[] for _ in range(num_levels)]
+ mlvl_pos_mask_preds_x = [[] for _ in range(num_levels)]
+ mlvl_pos_mask_preds_y = [[] for _ in range(num_levels)]
+ mlvl_labels = [[] for _ in range(num_levels)]
+ for img_id in range(num_imgs):
+
+ for lvl in range(num_levels):
+ mlvl_pos_mask_targets[lvl].append(
+ pos_mask_targets[img_id][lvl])
+ mlvl_pos_mask_preds_x[lvl].append(
+ mlvl_mask_preds_x[lvl][img_id,
+ xy_pos_indexes[img_id][lvl][:, 1]])
+ mlvl_pos_mask_preds_y[lvl].append(
+ mlvl_mask_preds_y[lvl][img_id,
+ xy_pos_indexes[img_id][lvl][:, 0]])
+ mlvl_labels[lvl].append(labels[img_id][lvl].flatten())
+
+ # cat multiple image
+ temp_mlvl_cls_preds = []
+ for lvl in range(num_levels):
+ mlvl_pos_mask_targets[lvl] = torch.cat(
+ mlvl_pos_mask_targets[lvl], dim=0)
+ mlvl_pos_mask_preds_x[lvl] = torch.cat(
+ mlvl_pos_mask_preds_x[lvl], dim=0)
+ mlvl_pos_mask_preds_y[lvl] = torch.cat(
+ mlvl_pos_mask_preds_y[lvl], dim=0)
+ mlvl_labels[lvl] = torch.cat(mlvl_labels[lvl], dim=0)
+ temp_mlvl_cls_preds.append(mlvl_cls_preds[lvl].permute(
+ 0, 2, 3, 1).reshape(-1, self.cls_out_channels))
+
+ num_pos = 0.
+ # dice loss
+ loss_mask = []
+ for pred_x, pred_y, target in \
+ zip(mlvl_pos_mask_preds_x,
+ mlvl_pos_mask_preds_y, mlvl_pos_mask_targets):
+ num_masks = pred_x.size(0)
+ if num_masks == 0:
+ # make sure can get grad
+ loss_mask.append((pred_x.sum() + pred_y.sum()).unsqueeze(0))
+ continue
+ num_pos += num_masks
+ pred_mask = pred_y.sigmoid() * pred_x.sigmoid()
+ loss_mask.append(
+ self.loss_mask(pred_mask, target, reduction_override='none'))
+ if num_pos > 0:
+ loss_mask = torch.cat(loss_mask).sum() / num_pos
+ else:
+ loss_mask = torch.cat(loss_mask).mean()
+
+ # cate
+ flatten_labels = torch.cat(mlvl_labels)
+ flatten_cls_preds = torch.cat(temp_mlvl_cls_preds)
+
+ loss_cls = self.loss_cls(
+ flatten_cls_preds, flatten_labels, avg_factor=num_pos + 1)
+ return dict(loss_mask=loss_mask, loss_cls=loss_cls)
+
+ def _get_targets_single(self,
+ gt_instances: InstanceData,
+ featmap_sizes: Optional[list] = None) -> tuple:
+ """Compute targets for predictions of single image.
+
+ Args:
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes``, ``labels``,
+ and ``masks`` attributes.
+ featmap_sizes (list[:obj:`torch.size`]): Size of each
+ feature map from feature pyramid, each element
+ means (feat_h, feat_w). Defaults to None.
+
+ Returns:
+ Tuple: Usually returns a tuple containing targets for predictions.
+
+ - mlvl_pos_mask_targets (list[Tensor]): Each element represent
+ the binary mask targets for positive points in this
+ level, has shape (num_pos, out_h, out_w).
+ - mlvl_labels (list[Tensor]): Each element is
+ classification labels for all
+ points in this level, has shape
+ (num_grid, num_grid).
+ - mlvl_xy_pos_indexes (list[Tensor]): Each element
+ in the list contains the index of positive samples in
+ corresponding level, has shape (num_pos, 2), last
+ dimension 2 present (index_x, index_y).
+ """
+ mlvl_pos_mask_targets, mlvl_labels, mlvl_pos_masks = \
+ super()._get_targets_single(gt_instances,
+ featmap_sizes=featmap_sizes)
+
+ mlvl_xy_pos_indexes = [(item - self.num_classes).nonzero()
+ for item in mlvl_labels]
+
+ return mlvl_pos_mask_targets, mlvl_labels, mlvl_xy_pos_indexes
+
+ def predict_by_feat(self, mlvl_mask_preds_x: List[Tensor],
+ mlvl_mask_preds_y: List[Tensor],
+ mlvl_cls_scores: List[Tensor],
+ batch_img_metas: List[dict], **kwargs) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ mask results.
+
+ Args:
+ mlvl_mask_preds_x (list[Tensor]): Multi-level mask prediction
+ from x branch. Each element in the list has shape
+ (batch_size, num_grids ,h ,w).
+ mlvl_mask_preds_y (list[Tensor]): Multi-level mask prediction
+ from y branch. Each element in the list has shape
+ (batch_size, num_grids ,h ,w).
+ mlvl_cls_scores (list[Tensor]): Multi-level scores. Each element
+ in the list has shape
+ (batch_size, num_classes ,num_grids ,num_grids).
+ batch_img_metas (list[dict]): Meta information of all images.
+
+ Returns:
+ list[:obj:`InstanceData`]: Processed results of multiple
+ images.Each :obj:`InstanceData` usually contains
+ following keys.
+
+ - scores (Tensor): Classification scores, has shape
+ (num_instance,).
+ - labels (Tensor): Has shape (num_instances,).
+ - masks (Tensor): Processed mask results, has
+ shape (num_instances, h, w).
+ """
+ mlvl_cls_scores = [
+ item.permute(0, 2, 3, 1) for item in mlvl_cls_scores
+ ]
+ assert len(mlvl_mask_preds_x) == len(mlvl_cls_scores)
+ num_levels = len(mlvl_cls_scores)
+
+ results_list = []
+ for img_id in range(len(batch_img_metas)):
+ cls_pred_list = [
+ mlvl_cls_scores[i][img_id].view(
+ -1, self.cls_out_channels).detach()
+ for i in range(num_levels)
+ ]
+ mask_pred_list_x = [
+ mlvl_mask_preds_x[i][img_id] for i in range(num_levels)
+ ]
+ mask_pred_list_y = [
+ mlvl_mask_preds_y[i][img_id] for i in range(num_levels)
+ ]
+
+ cls_pred_list = torch.cat(cls_pred_list, dim=0)
+ mask_pred_list_x = torch.cat(mask_pred_list_x, dim=0)
+ mask_pred_list_y = torch.cat(mask_pred_list_y, dim=0)
+ img_meta = batch_img_metas[img_id]
+
+ results = self._predict_by_feat_single(
+ cls_pred_list,
+ mask_pred_list_x,
+ mask_pred_list_y,
+ img_meta=img_meta)
+ results_list.append(results)
+ return results_list
+
+ def _predict_by_feat_single(self,
+ cls_scores: Tensor,
+ mask_preds_x: Tensor,
+ mask_preds_y: Tensor,
+ img_meta: dict,
+ cfg: OptConfigType = None) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ mask results.
+
+ Args:
+ cls_scores (Tensor): Classification score of all points
+ in single image, has shape (num_points, num_classes).
+ mask_preds_x (Tensor): Mask prediction of x branch of
+ all points in single image, has shape
+ (sum_num_grids, feat_h, feat_w).
+ mask_preds_y (Tensor): Mask prediction of y branch of
+ all points in single image, has shape
+ (sum_num_grids, feat_h, feat_w).
+ img_meta (dict): Meta information of corresponding image.
+ cfg (dict): Config used in test phase.
+
+ Returns:
+ :obj:`InstanceData`: Processed results of single image.
+ it usually contains following keys.
+
+ - scores (Tensor): Classification scores, has shape
+ (num_instance,).
+ - labels (Tensor): Has shape (num_instances,).
+ - masks (Tensor): Processed mask results, has
+ shape (num_instances, h, w).
+ """
+
+ def empty_results(cls_scores, ori_shape):
+ """Generate a empty results."""
+ results = InstanceData()
+ results.scores = cls_scores.new_ones(0)
+ results.masks = cls_scores.new_zeros(0, *ori_shape)
+ results.labels = cls_scores.new_ones(0)
+ results.bboxes = cls_scores.new_zeros(0, 4)
+ return results
+
+ cfg = self.test_cfg if cfg is None else cfg
+
+ featmap_size = mask_preds_x.size()[-2:]
+
+ h, w = img_meta['img_shape'][:2]
+ upsampled_size = (featmap_size[0] * 4, featmap_size[1] * 4)
+
+ score_mask = (cls_scores > cfg.score_thr)
+ cls_scores = cls_scores[score_mask]
+ inds = score_mask.nonzero()
+ lvl_interval = inds.new_tensor(self.num_grids).pow(2).cumsum(0)
+ num_all_points = lvl_interval[-1]
+ lvl_start_index = inds.new_ones(num_all_points)
+ num_grids = inds.new_ones(num_all_points)
+ seg_size = inds.new_tensor(self.num_grids).cumsum(0)
+ mask_lvl_start_index = inds.new_ones(num_all_points)
+ strides = inds.new_ones(num_all_points)
+
+ lvl_start_index[:lvl_interval[0]] *= 0
+ mask_lvl_start_index[:lvl_interval[0]] *= 0
+ num_grids[:lvl_interval[0]] *= self.num_grids[0]
+ strides[:lvl_interval[0]] *= self.strides[0]
+
+ for lvl in range(1, self.num_levels):
+ lvl_start_index[lvl_interval[lvl - 1]:lvl_interval[lvl]] *= \
+ lvl_interval[lvl - 1]
+ mask_lvl_start_index[lvl_interval[lvl - 1]:lvl_interval[lvl]] *= \
+ seg_size[lvl - 1]
+ num_grids[lvl_interval[lvl - 1]:lvl_interval[lvl]] *= \
+ self.num_grids[lvl]
+ strides[lvl_interval[lvl - 1]:lvl_interval[lvl]] *= \
+ self.strides[lvl]
+
+ lvl_start_index = lvl_start_index[inds[:, 0]]
+ mask_lvl_start_index = mask_lvl_start_index[inds[:, 0]]
+ num_grids = num_grids[inds[:, 0]]
+ strides = strides[inds[:, 0]]
+
+ y_lvl_offset = (inds[:, 0] - lvl_start_index) // num_grids
+ x_lvl_offset = (inds[:, 0] - lvl_start_index) % num_grids
+ y_inds = mask_lvl_start_index + y_lvl_offset
+ x_inds = mask_lvl_start_index + x_lvl_offset
+
+ cls_labels = inds[:, 1]
+ mask_preds = mask_preds_x[x_inds, ...] * mask_preds_y[y_inds, ...]
+
+ masks = mask_preds > cfg.mask_thr
+ sum_masks = masks.sum((1, 2)).float()
+ keep = sum_masks > strides
+ if keep.sum() == 0:
+ return empty_results(cls_scores, img_meta['ori_shape'][:2])
+
+ masks = masks[keep]
+ mask_preds = mask_preds[keep]
+ sum_masks = sum_masks[keep]
+ cls_scores = cls_scores[keep]
+ cls_labels = cls_labels[keep]
+
+ # maskness.
+ mask_scores = (mask_preds * masks).sum((1, 2)) / sum_masks
+ cls_scores *= mask_scores
+
+ scores, labels, _, keep_inds = mask_matrix_nms(
+ masks,
+ cls_labels,
+ cls_scores,
+ mask_area=sum_masks,
+ nms_pre=cfg.nms_pre,
+ max_num=cfg.max_per_img,
+ kernel=cfg.kernel,
+ sigma=cfg.sigma,
+ filter_thr=cfg.filter_thr)
+ # mask_matrix_nms may return an empty Tensor
+ if len(keep_inds) == 0:
+ return empty_results(cls_scores, img_meta['ori_shape'][:2])
+ mask_preds = mask_preds[keep_inds]
+ mask_preds = F.interpolate(
+ mask_preds.unsqueeze(0), size=upsampled_size,
+ mode='bilinear')[:, :, :h, :w]
+ mask_preds = F.interpolate(
+ mask_preds, size=img_meta['ori_shape'][:2],
+ mode='bilinear').squeeze(0)
+ masks = mask_preds > cfg.mask_thr
+
+ results = InstanceData()
+ results.masks = masks
+ results.labels = labels
+ results.scores = scores
+ # create an empty bbox in InstanceData to avoid bugs when
+ # calculating metrics.
+ results.bboxes = results.scores.new_zeros(len(scores), 4)
+
+ return results
+
+
+@MODELS.register_module()
+class DecoupledSOLOLightHead(DecoupledSOLOHead):
+ """Decoupled Light SOLO mask head used in `SOLO: Segmenting Objects by
+ Locations `_
+
+ Args:
+ with_dcn (bool): Whether use dcn in mask_convs and cls_convs,
+ Defaults to False.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ """
+
+ def __init__(self,
+ *args,
+ dcn_cfg: OptConfigType = None,
+ init_cfg: MultiConfig = [
+ dict(type='Normal', layer='Conv2d', std=0.01),
+ dict(
+ type='Normal',
+ std=0.01,
+ bias_prob=0.01,
+ override=dict(name='conv_mask_list_x')),
+ dict(
+ type='Normal',
+ std=0.01,
+ bias_prob=0.01,
+ override=dict(name='conv_mask_list_y')),
+ dict(
+ type='Normal',
+ std=0.01,
+ bias_prob=0.01,
+ override=dict(name='conv_cls'))
+ ],
+ **kwargs) -> None:
+ assert dcn_cfg is None or isinstance(dcn_cfg, dict)
+ self.dcn_cfg = dcn_cfg
+ super().__init__(*args, init_cfg=init_cfg, **kwargs)
+
+ def _init_layers(self) -> None:
+ self.mask_convs = nn.ModuleList()
+ self.cls_convs = nn.ModuleList()
+
+ for i in range(self.stacked_convs):
+ if self.dcn_cfg is not None \
+ and i == self.stacked_convs - 1:
+ conv_cfg = self.dcn_cfg
+ else:
+ conv_cfg = None
+
+ chn = self.in_channels + 2 if i == 0 else self.feat_channels
+ self.mask_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=self.norm_cfg))
+
+ chn = self.in_channels if i == 0 else self.feat_channels
+ self.cls_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=self.norm_cfg))
+
+ self.conv_mask_list_x = nn.ModuleList()
+ self.conv_mask_list_y = nn.ModuleList()
+ for num_grid in self.num_grids:
+ self.conv_mask_list_x.append(
+ nn.Conv2d(self.feat_channels, num_grid, 3, padding=1))
+ self.conv_mask_list_y.append(
+ nn.Conv2d(self.feat_channels, num_grid, 3, padding=1))
+ self.conv_cls = nn.Conv2d(
+ self.feat_channels, self.cls_out_channels, 3, padding=1)
+
+ def forward(self, x: Tuple[Tensor]) -> Tuple:
+ """Forward features from the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: A tuple of classification scores and mask prediction.
+
+ - mlvl_mask_preds_x (list[Tensor]): Multi-level mask prediction
+ from x branch. Each element in the list has shape
+ (batch_size, num_grids ,h ,w).
+ - mlvl_mask_preds_y (list[Tensor]): Multi-level mask prediction
+ from y branch. Each element in the list has shape
+ (batch_size, num_grids ,h ,w).
+ - mlvl_cls_preds (list[Tensor]): Multi-level scores.
+ Each element in the list has shape
+ (batch_size, num_classes, num_grids ,num_grids).
+ """
+ assert len(x) == self.num_levels
+ feats = self.resize_feats(x)
+ mask_preds_x = []
+ mask_preds_y = []
+ cls_preds = []
+ for i in range(self.num_levels):
+ x = feats[i]
+ mask_feat = x
+ cls_feat = x
+ # generate and concat the coordinate
+ coord_feat = generate_coordinate(mask_feat.size(),
+ mask_feat.device)
+ mask_feat = torch.cat([mask_feat, coord_feat], 1)
+
+ for mask_layer in self.mask_convs:
+ mask_feat = mask_layer(mask_feat)
+
+ mask_feat = F.interpolate(
+ mask_feat, scale_factor=2, mode='bilinear')
+
+ mask_pred_x = self.conv_mask_list_x[i](mask_feat)
+ mask_pred_y = self.conv_mask_list_y[i](mask_feat)
+
+ # cls branch
+ for j, cls_layer in enumerate(self.cls_convs):
+ if j == self.cls_down_index:
+ num_grid = self.num_grids[i]
+ cls_feat = F.interpolate(
+ cls_feat, size=num_grid, mode='bilinear')
+ cls_feat = cls_layer(cls_feat)
+
+ cls_pred = self.conv_cls(cls_feat)
+
+ if not self.training:
+ feat_wh = feats[0].size()[-2:]
+ upsampled_size = (feat_wh[0] * 2, feat_wh[1] * 2)
+ mask_pred_x = F.interpolate(
+ mask_pred_x.sigmoid(),
+ size=upsampled_size,
+ mode='bilinear')
+ mask_pred_y = F.interpolate(
+ mask_pred_y.sigmoid(),
+ size=upsampled_size,
+ mode='bilinear')
+ cls_pred = cls_pred.sigmoid()
+ # get local maximum
+ local_max = F.max_pool2d(cls_pred, 2, stride=1, padding=1)
+ keep_mask = local_max[:, :, :-1, :-1] == cls_pred
+ cls_pred = cls_pred * keep_mask
+
+ mask_preds_x.append(mask_pred_x)
+ mask_preds_y.append(mask_pred_y)
+ cls_preds.append(cls_pred)
+ return mask_preds_x, mask_preds_y, cls_preds
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/solov2_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/solov2_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..35b9df0c45148cb18e8afb659b10dd0b9e866b99
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/solov2_head.py
@@ -0,0 +1,799 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+from typing import List, Optional, Tuple
+
+import mmcv
+import numpy as np
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule
+from mmengine.model import BaseModule
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.models.utils.misc import floordiv
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, InstanceList, MultiConfig, OptConfigType
+from ..layers import mask_matrix_nms
+from ..utils import center_of_mass, generate_coordinate, multi_apply
+from .solo_head import SOLOHead
+
+
+class MaskFeatModule(BaseModule):
+ """SOLOv2 mask feature map branch used in `SOLOv2: Dynamic and Fast
+ Instance Segmentation. `_
+
+ Args:
+ in_channels (int): Number of channels in the input feature map.
+ feat_channels (int): Number of hidden channels of the mask feature
+ map branch.
+ start_level (int): The starting feature map level from RPN that
+ will be used to predict the mask feature map.
+ end_level (int): The ending feature map level from rpn that
+ will be used to predict the mask feature map.
+ out_channels (int): Number of output channels of the mask feature
+ map branch. This is the channel count of the mask
+ feature map that to be dynamically convolved with the predicted
+ kernel.
+ mask_stride (int): Downsample factor of the mask feature map output.
+ Defaults to 4.
+ conv_cfg (dict): Config dict for convolution layer. Default: None.
+ norm_cfg (dict): Config dict for normalization layer. Default: None.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ """
+
+ def __init__(
+ self,
+ in_channels: int,
+ feat_channels: int,
+ start_level: int,
+ end_level: int,
+ out_channels: int,
+ mask_stride: int = 4,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: OptConfigType = None,
+ init_cfg: MultiConfig = [
+ dict(type='Normal', layer='Conv2d', std=0.01)
+ ]
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.in_channels = in_channels
+ self.feat_channels = feat_channels
+ self.start_level = start_level
+ self.end_level = end_level
+ self.mask_stride = mask_stride
+ assert start_level >= 0 and end_level >= start_level
+ self.out_channels = out_channels
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self._init_layers()
+ self.fp16_enabled = False
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self.convs_all_levels = nn.ModuleList()
+ for i in range(self.start_level, self.end_level + 1):
+ convs_per_level = nn.Sequential()
+ if i == 0:
+ convs_per_level.add_module(
+ f'conv{i}',
+ ConvModule(
+ self.in_channels,
+ self.feat_channels,
+ 3,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ inplace=False))
+ self.convs_all_levels.append(convs_per_level)
+ continue
+
+ for j in range(i):
+ if j == 0:
+ if i == self.end_level:
+ chn = self.in_channels + 2
+ else:
+ chn = self.in_channels
+ convs_per_level.add_module(
+ f'conv{j}',
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ inplace=False))
+ convs_per_level.add_module(
+ f'upsample{j}',
+ nn.Upsample(
+ scale_factor=2,
+ mode='bilinear',
+ align_corners=False))
+ continue
+
+ convs_per_level.add_module(
+ f'conv{j}',
+ ConvModule(
+ self.feat_channels,
+ self.feat_channels,
+ 3,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ inplace=False))
+ convs_per_level.add_module(
+ f'upsample{j}',
+ nn.Upsample(
+ scale_factor=2, mode='bilinear', align_corners=False))
+
+ self.convs_all_levels.append(convs_per_level)
+
+ self.conv_pred = ConvModule(
+ self.feat_channels,
+ self.out_channels,
+ 1,
+ padding=0,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg)
+
+ def forward(self, x: Tuple[Tensor]) -> Tensor:
+ """Forward features from the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ Tensor: The predicted mask feature map.
+ """
+ inputs = x[self.start_level:self.end_level + 1]
+ assert len(inputs) == (self.end_level - self.start_level + 1)
+ feature_add_all_level = self.convs_all_levels[0](inputs[0])
+ for i in range(1, len(inputs)):
+ input_p = inputs[i]
+ if i == len(inputs) - 1:
+ coord_feat = generate_coordinate(input_p.size(),
+ input_p.device)
+ input_p = torch.cat([input_p, coord_feat], 1)
+
+ feature_add_all_level = feature_add_all_level + \
+ self.convs_all_levels[i](input_p)
+
+ feature_pred = self.conv_pred(feature_add_all_level)
+ return feature_pred
+
+
+@MODELS.register_module()
+class SOLOV2Head(SOLOHead):
+ """SOLOv2 mask head used in `SOLOv2: Dynamic and Fast Instance
+ Segmentation. `_
+
+ Args:
+ mask_feature_head (dict): Config of SOLOv2MaskFeatHead.
+ dynamic_conv_size (int): Dynamic Conv kernel size. Defaults to 1.
+ dcn_cfg (dict): Dcn conv configurations in kernel_convs and cls_conv.
+ Defaults to None.
+ dcn_apply_to_all_conv (bool): Whether to use dcn in every layer of
+ kernel_convs and cls_convs, or only the last layer. It shall be set
+ `True` for the normal version of SOLOv2 and `False` for the
+ light-weight version. Defaults to True.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ """
+
+ def __init__(self,
+ *args,
+ mask_feature_head: ConfigType,
+ dynamic_conv_size: int = 1,
+ dcn_cfg: OptConfigType = None,
+ dcn_apply_to_all_conv: bool = True,
+ init_cfg: MultiConfig = [
+ dict(type='Normal', layer='Conv2d', std=0.01),
+ dict(
+ type='Normal',
+ std=0.01,
+ bias_prob=0.01,
+ override=dict(name='conv_cls'))
+ ],
+ **kwargs) -> None:
+ assert dcn_cfg is None or isinstance(dcn_cfg, dict)
+ self.dcn_cfg = dcn_cfg
+ self.with_dcn = dcn_cfg is not None
+ self.dcn_apply_to_all_conv = dcn_apply_to_all_conv
+ self.dynamic_conv_size = dynamic_conv_size
+ mask_out_channels = mask_feature_head.get('out_channels')
+ self.kernel_out_channels = \
+ mask_out_channels * self.dynamic_conv_size * self.dynamic_conv_size
+
+ super().__init__(*args, init_cfg=init_cfg, **kwargs)
+
+ # update the in_channels of mask_feature_head
+ if mask_feature_head.get('in_channels', None) is not None:
+ if mask_feature_head.in_channels != self.in_channels:
+ warnings.warn('The `in_channels` of SOLOv2MaskFeatHead and '
+ 'SOLOv2Head should be same, changing '
+ 'mask_feature_head.in_channels to '
+ f'{self.in_channels}')
+ mask_feature_head.update(in_channels=self.in_channels)
+ else:
+ mask_feature_head.update(in_channels=self.in_channels)
+
+ self.mask_feature_head = MaskFeatModule(**mask_feature_head)
+ self.mask_stride = self.mask_feature_head.mask_stride
+ self.fp16_enabled = False
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self.cls_convs = nn.ModuleList()
+ self.kernel_convs = nn.ModuleList()
+ conv_cfg = None
+ for i in range(self.stacked_convs):
+ if self.with_dcn:
+ if self.dcn_apply_to_all_conv:
+ conv_cfg = self.dcn_cfg
+ elif i == self.stacked_convs - 1:
+ # light head
+ conv_cfg = self.dcn_cfg
+
+ chn = self.in_channels + 2 if i == 0 else self.feat_channels
+ self.kernel_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=self.norm_cfg,
+ bias=self.norm_cfg is None))
+
+ chn = self.in_channels if i == 0 else self.feat_channels
+ self.cls_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=self.norm_cfg,
+ bias=self.norm_cfg is None))
+
+ self.conv_cls = nn.Conv2d(
+ self.feat_channels, self.cls_out_channels, 3, padding=1)
+
+ self.conv_kernel = nn.Conv2d(
+ self.feat_channels, self.kernel_out_channels, 3, padding=1)
+
+ def forward(self, x):
+ """Forward features from the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: A tuple of classification scores, mask prediction,
+ and mask features.
+
+ - mlvl_kernel_preds (list[Tensor]): Multi-level dynamic kernel
+ prediction. The kernel is used to generate instance
+ segmentation masks by dynamic convolution. Each element in
+ the list has shape
+ (batch_size, kernel_out_channels, num_grids, num_grids).
+ - mlvl_cls_preds (list[Tensor]): Multi-level scores. Each
+ element in the list has shape
+ (batch_size, num_classes, num_grids, num_grids).
+ - mask_feats (Tensor): Unified mask feature map used to
+ generate instance segmentation masks by dynamic convolution.
+ Has shape (batch_size, mask_out_channels, h, w).
+ """
+ assert len(x) == self.num_levels
+ mask_feats = self.mask_feature_head(x)
+ ins_kernel_feats = self.resize_feats(x)
+ mlvl_kernel_preds = []
+ mlvl_cls_preds = []
+ for i in range(self.num_levels):
+ ins_kernel_feat = ins_kernel_feats[i]
+ # ins branch
+ # concat coord
+ coord_feat = generate_coordinate(ins_kernel_feat.size(),
+ ins_kernel_feat.device)
+ ins_kernel_feat = torch.cat([ins_kernel_feat, coord_feat], 1)
+
+ # kernel branch
+ kernel_feat = ins_kernel_feat
+ kernel_feat = F.interpolate(
+ kernel_feat,
+ size=self.num_grids[i],
+ mode='bilinear',
+ align_corners=False)
+
+ cate_feat = kernel_feat[:, :-2, :, :]
+
+ kernel_feat = kernel_feat.contiguous()
+ for i, kernel_conv in enumerate(self.kernel_convs):
+ kernel_feat = kernel_conv(kernel_feat)
+ kernel_pred = self.conv_kernel(kernel_feat)
+
+ # cate branch
+ cate_feat = cate_feat.contiguous()
+ for i, cls_conv in enumerate(self.cls_convs):
+ cate_feat = cls_conv(cate_feat)
+ cate_pred = self.conv_cls(cate_feat)
+
+ mlvl_kernel_preds.append(kernel_pred)
+ mlvl_cls_preds.append(cate_pred)
+
+ return mlvl_kernel_preds, mlvl_cls_preds, mask_feats
+
+ def _get_targets_single(self,
+ gt_instances: InstanceData,
+ featmap_sizes: Optional[list] = None) -> tuple:
+ """Compute targets for predictions of single image.
+
+ Args:
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes``, ``labels``,
+ and ``masks`` attributes.
+ featmap_sizes (list[:obj:`torch.size`]): Size of each
+ feature map from feature pyramid, each element
+ means (feat_h, feat_w). Defaults to None.
+
+ Returns:
+ Tuple: Usually returns a tuple containing targets for predictions.
+
+ - mlvl_pos_mask_targets (list[Tensor]): Each element represent
+ the binary mask targets for positive points in this
+ level, has shape (num_pos, out_h, out_w).
+ - mlvl_labels (list[Tensor]): Each element is
+ classification labels for all
+ points in this level, has shape
+ (num_grid, num_grid).
+ - mlvl_pos_masks (list[Tensor]): Each element is
+ a `BoolTensor` to represent whether the
+ corresponding point in single level
+ is positive, has shape (num_grid **2).
+ - mlvl_pos_indexes (list[list]): Each element
+ in the list contains the positive index in
+ corresponding level, has shape (num_pos).
+ """
+ gt_labels = gt_instances.labels
+ device = gt_labels.device
+
+ gt_bboxes = gt_instances.bboxes
+ gt_areas = torch.sqrt((gt_bboxes[:, 2] - gt_bboxes[:, 0]) *
+ (gt_bboxes[:, 3] - gt_bboxes[:, 1]))
+ gt_masks = gt_instances.masks.to_tensor(
+ dtype=torch.bool, device=device)
+
+ mlvl_pos_mask_targets = []
+ mlvl_pos_indexes = []
+ mlvl_labels = []
+ mlvl_pos_masks = []
+ for (lower_bound, upper_bound), num_grid \
+ in zip(self.scale_ranges, self.num_grids):
+ mask_target = []
+ # FG cat_id: [0, num_classes -1], BG cat_id: num_classes
+ pos_index = []
+ labels = torch.zeros([num_grid, num_grid],
+ dtype=torch.int64,
+ device=device) + self.num_classes
+ pos_mask = torch.zeros([num_grid**2],
+ dtype=torch.bool,
+ device=device)
+
+ gt_inds = ((gt_areas >= lower_bound) &
+ (gt_areas <= upper_bound)).nonzero().flatten()
+ if len(gt_inds) == 0:
+ mlvl_pos_mask_targets.append(
+ torch.zeros([0, featmap_sizes[0], featmap_sizes[1]],
+ dtype=torch.uint8,
+ device=device))
+ mlvl_labels.append(labels)
+ mlvl_pos_masks.append(pos_mask)
+ mlvl_pos_indexes.append([])
+ continue
+ hit_gt_bboxes = gt_bboxes[gt_inds]
+ hit_gt_labels = gt_labels[gt_inds]
+ hit_gt_masks = gt_masks[gt_inds, ...]
+
+ pos_w_ranges = 0.5 * (hit_gt_bboxes[:, 2] -
+ hit_gt_bboxes[:, 0]) * self.pos_scale
+ pos_h_ranges = 0.5 * (hit_gt_bboxes[:, 3] -
+ hit_gt_bboxes[:, 1]) * self.pos_scale
+
+ # Make sure hit_gt_masks has a value
+ valid_mask_flags = hit_gt_masks.sum(dim=-1).sum(dim=-1) > 0
+
+ for gt_mask, gt_label, pos_h_range, pos_w_range, \
+ valid_mask_flag in \
+ zip(hit_gt_masks, hit_gt_labels, pos_h_ranges,
+ pos_w_ranges, valid_mask_flags):
+ if not valid_mask_flag:
+ continue
+ upsampled_size = (featmap_sizes[0] * self.mask_stride,
+ featmap_sizes[1] * self.mask_stride)
+ center_h, center_w = center_of_mass(gt_mask)
+
+ coord_w = int(
+ floordiv((center_w / upsampled_size[1]), (1. / num_grid),
+ rounding_mode='trunc'))
+ coord_h = int(
+ floordiv((center_h / upsampled_size[0]), (1. / num_grid),
+ rounding_mode='trunc'))
+
+ # left, top, right, down
+ top_box = max(
+ 0,
+ int(
+ floordiv(
+ (center_h - pos_h_range) / upsampled_size[0],
+ (1. / num_grid),
+ rounding_mode='trunc')))
+ down_box = min(
+ num_grid - 1,
+ int(
+ floordiv(
+ (center_h + pos_h_range) / upsampled_size[0],
+ (1. / num_grid),
+ rounding_mode='trunc')))
+ left_box = max(
+ 0,
+ int(
+ floordiv(
+ (center_w - pos_w_range) / upsampled_size[1],
+ (1. / num_grid),
+ rounding_mode='trunc')))
+ right_box = min(
+ num_grid - 1,
+ int(
+ floordiv(
+ (center_w + pos_w_range) / upsampled_size[1],
+ (1. / num_grid),
+ rounding_mode='trunc')))
+
+ top = max(top_box, coord_h - 1)
+ down = min(down_box, coord_h + 1)
+ left = max(coord_w - 1, left_box)
+ right = min(right_box, coord_w + 1)
+
+ labels[top:(down + 1), left:(right + 1)] = gt_label
+ # ins
+ gt_mask = np.uint8(gt_mask.cpu().numpy())
+ # Follow the original implementation, F.interpolate is
+ # different from cv2 and opencv
+ gt_mask = mmcv.imrescale(gt_mask, scale=1. / self.mask_stride)
+ gt_mask = torch.from_numpy(gt_mask).to(device=device)
+
+ for i in range(top, down + 1):
+ for j in range(left, right + 1):
+ index = int(i * num_grid + j)
+ this_mask_target = torch.zeros(
+ [featmap_sizes[0], featmap_sizes[1]],
+ dtype=torch.uint8,
+ device=device)
+ this_mask_target[:gt_mask.shape[0], :gt_mask.
+ shape[1]] = gt_mask
+ mask_target.append(this_mask_target)
+ pos_mask[index] = True
+ pos_index.append(index)
+ if len(mask_target) == 0:
+ mask_target = torch.zeros(
+ [0, featmap_sizes[0], featmap_sizes[1]],
+ dtype=torch.uint8,
+ device=device)
+ else:
+ mask_target = torch.stack(mask_target, 0)
+ mlvl_pos_mask_targets.append(mask_target)
+ mlvl_labels.append(labels)
+ mlvl_pos_masks.append(pos_mask)
+ mlvl_pos_indexes.append(pos_index)
+ return (mlvl_pos_mask_targets, mlvl_labels, mlvl_pos_masks,
+ mlvl_pos_indexes)
+
+ def loss_by_feat(self, mlvl_kernel_preds: List[Tensor],
+ mlvl_cls_preds: List[Tensor], mask_feats: Tensor,
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict], **kwargs) -> dict:
+ """Calculate the loss based on the features extracted by the mask head.
+
+ Args:
+ mlvl_kernel_preds (list[Tensor]): Multi-level dynamic kernel
+ prediction. The kernel is used to generate instance
+ segmentation masks by dynamic convolution. Each element in the
+ list has shape
+ (batch_size, kernel_out_channels, num_grids, num_grids).
+ mlvl_cls_preds (list[Tensor]): Multi-level scores. Each element
+ in the list has shape
+ (batch_size, num_classes, num_grids, num_grids).
+ mask_feats (Tensor): Unified mask feature map used to generate
+ instance segmentation masks by dynamic convolution. Has shape
+ (batch_size, mask_out_channels, h, w).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``masks``,
+ and ``labels`` attributes.
+ batch_img_metas (list[dict]): Meta information of multiple images.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ featmap_sizes = mask_feats.size()[-2:]
+
+ pos_mask_targets, labels, pos_masks, pos_indexes = multi_apply(
+ self._get_targets_single,
+ batch_gt_instances,
+ featmap_sizes=featmap_sizes)
+
+ mlvl_mask_targets = [
+ torch.cat(lvl_mask_targets, 0)
+ for lvl_mask_targets in zip(*pos_mask_targets)
+ ]
+
+ mlvl_pos_kernel_preds = []
+ for lvl_kernel_preds, lvl_pos_indexes in zip(mlvl_kernel_preds,
+ zip(*pos_indexes)):
+ lvl_pos_kernel_preds = []
+ for img_lvl_kernel_preds, img_lvl_pos_indexes in zip(
+ lvl_kernel_preds, lvl_pos_indexes):
+ img_lvl_pos_kernel_preds = img_lvl_kernel_preds.view(
+ img_lvl_kernel_preds.shape[0], -1)[:, img_lvl_pos_indexes]
+ lvl_pos_kernel_preds.append(img_lvl_pos_kernel_preds)
+ mlvl_pos_kernel_preds.append(lvl_pos_kernel_preds)
+
+ # make multilevel mlvl_mask_pred
+ mlvl_mask_preds = []
+ for lvl_pos_kernel_preds in mlvl_pos_kernel_preds:
+ lvl_mask_preds = []
+ for img_id, img_lvl_pos_kernel_pred in enumerate(
+ lvl_pos_kernel_preds):
+ if img_lvl_pos_kernel_pred.size()[-1] == 0:
+ continue
+ img_mask_feats = mask_feats[[img_id]]
+ h, w = img_mask_feats.shape[-2:]
+ num_kernel = img_lvl_pos_kernel_pred.shape[1]
+ img_lvl_mask_pred = F.conv2d(
+ img_mask_feats,
+ img_lvl_pos_kernel_pred.permute(1, 0).view(
+ num_kernel, -1, self.dynamic_conv_size,
+ self.dynamic_conv_size),
+ stride=1).view(-1, h, w)
+ lvl_mask_preds.append(img_lvl_mask_pred)
+ if len(lvl_mask_preds) == 0:
+ lvl_mask_preds = None
+ else:
+ lvl_mask_preds = torch.cat(lvl_mask_preds, 0)
+ mlvl_mask_preds.append(lvl_mask_preds)
+ # dice loss
+ num_pos = 0
+ for img_pos_masks in pos_masks:
+ for lvl_img_pos_masks in img_pos_masks:
+ # Fix `Tensor` object has no attribute `count_nonzero()`
+ # in PyTorch 1.6, the type of `lvl_img_pos_masks`
+ # should be `torch.bool`.
+ num_pos += lvl_img_pos_masks.nonzero().numel()
+ loss_mask = []
+ for lvl_mask_preds, lvl_mask_targets in zip(mlvl_mask_preds,
+ mlvl_mask_targets):
+ if lvl_mask_preds is None:
+ continue
+ loss_mask.append(
+ self.loss_mask(
+ lvl_mask_preds,
+ lvl_mask_targets,
+ reduction_override='none'))
+ if num_pos > 0:
+ loss_mask = torch.cat(loss_mask).sum() / num_pos
+ else:
+ loss_mask = mask_feats.sum() * 0
+
+ # cate
+ flatten_labels = [
+ torch.cat(
+ [img_lvl_labels.flatten() for img_lvl_labels in lvl_labels])
+ for lvl_labels in zip(*labels)
+ ]
+ flatten_labels = torch.cat(flatten_labels)
+
+ flatten_cls_preds = [
+ lvl_cls_preds.permute(0, 2, 3, 1).reshape(-1, self.num_classes)
+ for lvl_cls_preds in mlvl_cls_preds
+ ]
+ flatten_cls_preds = torch.cat(flatten_cls_preds)
+
+ loss_cls = self.loss_cls(
+ flatten_cls_preds, flatten_labels, avg_factor=num_pos + 1)
+ return dict(loss_mask=loss_mask, loss_cls=loss_cls)
+
+ def predict_by_feat(self, mlvl_kernel_preds: List[Tensor],
+ mlvl_cls_scores: List[Tensor], mask_feats: Tensor,
+ batch_img_metas: List[dict], **kwargs) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ mask results.
+
+ Args:
+ mlvl_kernel_preds (list[Tensor]): Multi-level dynamic kernel
+ prediction. The kernel is used to generate instance
+ segmentation masks by dynamic convolution. Each element in the
+ list has shape
+ (batch_size, kernel_out_channels, num_grids, num_grids).
+ mlvl_cls_scores (list[Tensor]): Multi-level scores. Each element
+ in the list has shape
+ (batch_size, num_classes, num_grids, num_grids).
+ mask_feats (Tensor): Unified mask feature map used to generate
+ instance segmentation masks by dynamic convolution. Has shape
+ (batch_size, mask_out_channels, h, w).
+ batch_img_metas (list[dict]): Meta information of all images.
+
+ Returns:
+ list[:obj:`InstanceData`]: Processed results of multiple
+ images.Each :obj:`InstanceData` usually contains
+ following keys.
+
+ - scores (Tensor): Classification scores, has shape
+ (num_instance,).
+ - labels (Tensor): Has shape (num_instances,).
+ - masks (Tensor): Processed mask results, has
+ shape (num_instances, h, w).
+ """
+ num_levels = len(mlvl_cls_scores)
+ assert len(mlvl_kernel_preds) == len(mlvl_cls_scores)
+
+ for lvl in range(num_levels):
+ cls_scores = mlvl_cls_scores[lvl]
+ cls_scores = cls_scores.sigmoid()
+ local_max = F.max_pool2d(cls_scores, 2, stride=1, padding=1)
+ keep_mask = local_max[:, :, :-1, :-1] == cls_scores
+ cls_scores = cls_scores * keep_mask
+ mlvl_cls_scores[lvl] = cls_scores.permute(0, 2, 3, 1)
+
+ result_list = []
+ for img_id in range(len(batch_img_metas)):
+ img_cls_pred = [
+ mlvl_cls_scores[lvl][img_id].view(-1, self.cls_out_channels)
+ for lvl in range(num_levels)
+ ]
+ img_mask_feats = mask_feats[[img_id]]
+ img_kernel_pred = [
+ mlvl_kernel_preds[lvl][img_id].permute(1, 2, 0).view(
+ -1, self.kernel_out_channels) for lvl in range(num_levels)
+ ]
+ img_cls_pred = torch.cat(img_cls_pred, dim=0)
+ img_kernel_pred = torch.cat(img_kernel_pred, dim=0)
+ result = self._predict_by_feat_single(
+ img_kernel_pred,
+ img_cls_pred,
+ img_mask_feats,
+ img_meta=batch_img_metas[img_id])
+ result_list.append(result)
+ return result_list
+
+ def _predict_by_feat_single(self,
+ kernel_preds: Tensor,
+ cls_scores: Tensor,
+ mask_feats: Tensor,
+ img_meta: dict,
+ cfg: OptConfigType = None) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ mask results.
+
+ Args:
+ kernel_preds (Tensor): Dynamic kernel prediction of all points
+ in single image, has shape
+ (num_points, kernel_out_channels).
+ cls_scores (Tensor): Classification score of all points
+ in single image, has shape (num_points, num_classes).
+ mask_feats (Tensor): Mask prediction of all points in
+ single image, has shape (num_points, feat_h, feat_w).
+ img_meta (dict): Meta information of corresponding image.
+ cfg (dict, optional): Config used in test phase.
+ Defaults to None.
+
+ Returns:
+ :obj:`InstanceData`: Processed results of single image.
+ it usually contains following keys.
+
+ - scores (Tensor): Classification scores, has shape
+ (num_instance,).
+ - labels (Tensor): Has shape (num_instances,).
+ - masks (Tensor): Processed mask results, has
+ shape (num_instances, h, w).
+ """
+
+ def empty_results(cls_scores, ori_shape):
+ """Generate a empty results."""
+ results = InstanceData()
+ results.scores = cls_scores.new_ones(0)
+ results.masks = cls_scores.new_zeros(0, *ori_shape)
+ results.labels = cls_scores.new_ones(0)
+ results.bboxes = cls_scores.new_zeros(0, 4)
+ return results
+
+ cfg = self.test_cfg if cfg is None else cfg
+ assert len(kernel_preds) == len(cls_scores)
+
+ featmap_size = mask_feats.size()[-2:]
+
+ # overall info
+ h, w = img_meta['img_shape'][:2]
+ upsampled_size = (featmap_size[0] * self.mask_stride,
+ featmap_size[1] * self.mask_stride)
+
+ # process.
+ score_mask = (cls_scores > cfg.score_thr)
+ cls_scores = cls_scores[score_mask]
+ if len(cls_scores) == 0:
+ return empty_results(cls_scores, img_meta['ori_shape'][:2])
+
+ # cate_labels & kernel_preds
+ inds = score_mask.nonzero()
+ cls_labels = inds[:, 1]
+ kernel_preds = kernel_preds[inds[:, 0]]
+
+ # trans vector.
+ lvl_interval = cls_labels.new_tensor(self.num_grids).pow(2).cumsum(0)
+ strides = kernel_preds.new_ones(lvl_interval[-1])
+
+ strides[:lvl_interval[0]] *= self.strides[0]
+ for lvl in range(1, self.num_levels):
+ strides[lvl_interval[lvl -
+ 1]:lvl_interval[lvl]] *= self.strides[lvl]
+ strides = strides[inds[:, 0]]
+
+ # mask encoding.
+ kernel_preds = kernel_preds.view(
+ kernel_preds.size(0), -1, self.dynamic_conv_size,
+ self.dynamic_conv_size)
+ mask_preds = F.conv2d(
+ mask_feats, kernel_preds, stride=1).squeeze(0).sigmoid()
+ # mask.
+ masks = mask_preds > cfg.mask_thr
+ sum_masks = masks.sum((1, 2)).float()
+ keep = sum_masks > strides
+ if keep.sum() == 0:
+ return empty_results(cls_scores, img_meta['ori_shape'][:2])
+ masks = masks[keep]
+ mask_preds = mask_preds[keep]
+ sum_masks = sum_masks[keep]
+ cls_scores = cls_scores[keep]
+ cls_labels = cls_labels[keep]
+
+ # maskness.
+ mask_scores = (mask_preds * masks).sum((1, 2)) / sum_masks
+ cls_scores *= mask_scores
+
+ scores, labels, _, keep_inds = mask_matrix_nms(
+ masks,
+ cls_labels,
+ cls_scores,
+ mask_area=sum_masks,
+ nms_pre=cfg.nms_pre,
+ max_num=cfg.max_per_img,
+ kernel=cfg.kernel,
+ sigma=cfg.sigma,
+ filter_thr=cfg.filter_thr)
+ if len(keep_inds) == 0:
+ return empty_results(cls_scores, img_meta['ori_shape'][:2])
+ mask_preds = mask_preds[keep_inds]
+ mask_preds = F.interpolate(
+ mask_preds.unsqueeze(0),
+ size=upsampled_size,
+ mode='bilinear',
+ align_corners=False)[:, :, :h, :w]
+ mask_preds = F.interpolate(
+ mask_preds,
+ size=img_meta['ori_shape'][:2],
+ mode='bilinear',
+ align_corners=False).squeeze(0)
+ masks = mask_preds > cfg.mask_thr
+
+ results = InstanceData()
+ results.masks = masks
+ results.labels = labels
+ results.scores = scores
+ # create an empty bbox in InstanceData to avoid bugs when
+ # calculating metrics.
+ results.bboxes = results.scores.new_zeros(len(scores), 4)
+
+ return results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/ssd_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/ssd_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..950df29110d914cc888bc16c6cbf1856f604a1de
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/ssd_head.py
@@ -0,0 +1,362 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, List, Optional, Sequence, Tuple
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule, DepthwiseSeparableConvModule
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.utils import ConfigType, InstanceList, MultiConfig, OptInstanceList
+from ..losses import smooth_l1_loss
+from ..task_modules.samplers import PseudoSampler
+from ..utils import multi_apply
+from .anchor_head import AnchorHead
+
+
+# TODO: add loss evaluator for SSD
+@MODELS.register_module()
+class SSDHead(AnchorHead):
+ """Implementation of `SSD head `_
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (Sequence[int]): Number of channels in the input feature
+ map.
+ stacked_convs (int): Number of conv layers in cls and reg tower.
+ Defaults to 0.
+ feat_channels (int): Number of hidden channels when stacked_convs
+ > 0. Defaults to 256.
+ use_depthwise (bool): Whether to use DepthwiseSeparableConv.
+ Defaults to False.
+ conv_cfg (:obj:`ConfigDict` or dict, Optional): Dictionary to construct
+ and config conv layer. Defaults to None.
+ norm_cfg (:obj:`ConfigDict` or dict, Optional): Dictionary to construct
+ and config norm layer. Defaults to None.
+ act_cfg (:obj:`ConfigDict` or dict, Optional): Dictionary to construct
+ and config activation layer. Defaults to None.
+ anchor_generator (:obj:`ConfigDict` or dict): Config dict for anchor
+ generator.
+ bbox_coder (:obj:`ConfigDict` or dict): Config of bounding box coder.
+ reg_decoded_bbox (bool): If true, the regression loss would be
+ applied directly on decoded bounding boxes, converting both
+ the predicted boxes and regression targets to absolute
+ coordinates format. Defaults to False. It should be `True` when
+ using `IoULoss`, `GIoULoss`, or `DIoULoss` in the bbox head.
+ train_cfg (:obj:`ConfigDict` or dict, Optional): Training config of
+ anchor head.
+ test_cfg (:obj:`ConfigDict` or dict, Optional): Testing config of
+ anchor head.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict], Optional): Initialization config dict.
+ """ # noqa: W605
+
+ def __init__(
+ self,
+ num_classes: int = 80,
+ in_channels: Sequence[int] = (512, 1024, 512, 256, 256, 256),
+ stacked_convs: int = 0,
+ feat_channels: int = 256,
+ use_depthwise: bool = False,
+ conv_cfg: Optional[ConfigType] = None,
+ norm_cfg: Optional[ConfigType] = None,
+ act_cfg: Optional[ConfigType] = None,
+ anchor_generator: ConfigType = dict(
+ type='SSDAnchorGenerator',
+ scale_major=False,
+ input_size=300,
+ strides=[8, 16, 32, 64, 100, 300],
+ ratios=([2], [2, 3], [2, 3], [2, 3], [2], [2]),
+ basesize_ratio_range=(0.1, 0.9)),
+ bbox_coder: ConfigType = dict(
+ type='DeltaXYWHBBoxCoder',
+ clip_border=True,
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0],
+ ),
+ reg_decoded_bbox: bool = False,
+ train_cfg: Optional[ConfigType] = None,
+ test_cfg: Optional[ConfigType] = None,
+ init_cfg: MultiConfig = dict(
+ type='Xavier', layer='Conv2d', distribution='uniform', bias=0)
+ ) -> None:
+ super(AnchorHead, self).__init__(init_cfg=init_cfg)
+ self.num_classes = num_classes
+ self.in_channels = in_channels
+ self.stacked_convs = stacked_convs
+ self.feat_channels = feat_channels
+ self.use_depthwise = use_depthwise
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self.act_cfg = act_cfg
+
+ self.cls_out_channels = num_classes + 1 # add background class
+ self.prior_generator = TASK_UTILS.build(anchor_generator)
+
+ # Usually the numbers of anchors for each level are the same
+ # except SSD detectors. So it is an int in the most dense
+ # heads but a list of int in SSDHead
+ self.num_base_priors = self.prior_generator.num_base_priors
+
+ self._init_layers()
+
+ self.bbox_coder = TASK_UTILS.build(bbox_coder)
+ self.reg_decoded_bbox = reg_decoded_bbox
+ self.use_sigmoid_cls = False
+ self.cls_focal_loss = False
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+ if self.train_cfg:
+ self.assigner = TASK_UTILS.build(self.train_cfg['assigner'])
+ if self.train_cfg.get('sampler', None) is not None:
+ self.sampler = TASK_UTILS.build(
+ self.train_cfg['sampler'], default_args=dict(context=self))
+ else:
+ self.sampler = PseudoSampler(context=self)
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self.cls_convs = nn.ModuleList()
+ self.reg_convs = nn.ModuleList()
+ # TODO: Use registry to choose ConvModule type
+ conv = DepthwiseSeparableConvModule \
+ if self.use_depthwise else ConvModule
+
+ for channel, num_base_priors in zip(self.in_channels,
+ self.num_base_priors):
+ cls_layers = []
+ reg_layers = []
+ in_channel = channel
+ # build stacked conv tower, not used in default ssd
+ for i in range(self.stacked_convs):
+ cls_layers.append(
+ conv(
+ in_channel,
+ self.feat_channels,
+ 3,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg))
+ reg_layers.append(
+ conv(
+ in_channel,
+ self.feat_channels,
+ 3,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg))
+ in_channel = self.feat_channels
+ # SSD-Lite head
+ if self.use_depthwise:
+ cls_layers.append(
+ ConvModule(
+ in_channel,
+ in_channel,
+ 3,
+ padding=1,
+ groups=in_channel,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg))
+ reg_layers.append(
+ ConvModule(
+ in_channel,
+ in_channel,
+ 3,
+ padding=1,
+ groups=in_channel,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg))
+ cls_layers.append(
+ nn.Conv2d(
+ in_channel,
+ num_base_priors * self.cls_out_channels,
+ kernel_size=1 if self.use_depthwise else 3,
+ padding=0 if self.use_depthwise else 1))
+ reg_layers.append(
+ nn.Conv2d(
+ in_channel,
+ num_base_priors * 4,
+ kernel_size=1 if self.use_depthwise else 3,
+ padding=0 if self.use_depthwise else 1))
+ self.cls_convs.append(nn.Sequential(*cls_layers))
+ self.reg_convs.append(nn.Sequential(*reg_layers))
+
+ def forward(self, x: Tuple[Tensor]) -> Tuple[List[Tensor], List[Tensor]]:
+ """Forward features from the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple[list[Tensor], list[Tensor]]: A tuple of cls_scores list and
+ bbox_preds list.
+
+ - cls_scores (list[Tensor]): Classification scores for all scale \
+ levels, each is a 4D-tensor, the channels number is \
+ num_anchors * num_classes.
+ - bbox_preds (list[Tensor]): Box energies / deltas for all scale \
+ levels, each is a 4D-tensor, the channels number is \
+ num_anchors * 4.
+ """
+ cls_scores = []
+ bbox_preds = []
+ for feat, reg_conv, cls_conv in zip(x, self.reg_convs, self.cls_convs):
+ cls_scores.append(cls_conv(feat))
+ bbox_preds.append(reg_conv(feat))
+ return cls_scores, bbox_preds
+
+ def loss_by_feat_single(self, cls_score: Tensor, bbox_pred: Tensor,
+ anchor: Tensor, labels: Tensor,
+ label_weights: Tensor, bbox_targets: Tensor,
+ bbox_weights: Tensor,
+ avg_factor: int) -> Tuple[Tensor, Tensor]:
+ """Compute loss of a single image.
+
+ Args:
+ cls_score (Tensor): Box scores for eachimage
+ Has shape (num_total_anchors, num_classes).
+ bbox_pred (Tensor): Box energies / deltas for each image
+ level with shape (num_total_anchors, 4).
+ anchors (Tensor): Box reference for each scale level with shape
+ (num_total_anchors, 4).
+ labels (Tensor): Labels of each anchors with shape
+ (num_total_anchors,).
+ label_weights (Tensor): Label weights of each anchor with shape
+ (num_total_anchors,)
+ bbox_targets (Tensor): BBox regression targets of each anchor with
+ shape (num_total_anchors, 4).
+ bbox_weights (Tensor): BBox regression loss weights of each anchor
+ with shape (num_total_anchors, 4).
+ avg_factor (int): Average factor that is used to average
+ the loss. When using sampling method, avg_factor is usually
+ the sum of positive and negative priors. When using
+ `PseudoSampler`, `avg_factor` is usually equal to the number
+ of positive priors.
+
+ Returns:
+ Tuple[Tensor, Tensor]: A tuple of cls loss and bbox loss of one
+ feature map.
+ """
+
+ loss_cls_all = F.cross_entropy(
+ cls_score, labels, reduction='none') * label_weights
+ # FG cat_id: [0, num_classes -1], BG cat_id: num_classes
+ pos_inds = ((labels >= 0) & (labels < self.num_classes)).nonzero(
+ as_tuple=False).reshape(-1)
+ neg_inds = (labels == self.num_classes).nonzero(
+ as_tuple=False).view(-1)
+
+ num_pos_samples = pos_inds.size(0)
+ num_neg_samples = self.train_cfg['neg_pos_ratio'] * num_pos_samples
+ if num_neg_samples > neg_inds.size(0):
+ num_neg_samples = neg_inds.size(0)
+ topk_loss_cls_neg, _ = loss_cls_all[neg_inds].topk(num_neg_samples)
+ loss_cls_pos = loss_cls_all[pos_inds].sum()
+ loss_cls_neg = topk_loss_cls_neg.sum()
+ loss_cls = (loss_cls_pos + loss_cls_neg) / avg_factor
+
+ if self.reg_decoded_bbox:
+ # When the regression loss (e.g. `IouLoss`, `GIouLoss`)
+ # is applied directly on the decoded bounding boxes, it
+ # decodes the already encoded coordinates to absolute format.
+ bbox_pred = self.bbox_coder.decode(anchor, bbox_pred)
+
+ loss_bbox = smooth_l1_loss(
+ bbox_pred,
+ bbox_targets,
+ bbox_weights,
+ beta=self.train_cfg['smoothl1_beta'],
+ avg_factor=avg_factor)
+ return loss_cls[None], loss_bbox
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None
+ ) -> Dict[str, List[Tensor]]:
+ """Compute losses of the head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ Has shape (N, num_anchors * num_classes, H, W)
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W)
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, list[Tensor]]: A dictionary of loss components. the dict
+ has components below:
+
+ - loss_cls (list[Tensor]): A list containing each feature map \
+ classification loss.
+ - loss_bbox (list[Tensor]): A list containing each feature map \
+ regression loss.
+ """
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ assert len(featmap_sizes) == self.prior_generator.num_levels
+
+ device = cls_scores[0].device
+
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+ cls_reg_targets = self.get_targets(
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore,
+ unmap_outputs=True)
+ (labels_list, label_weights_list, bbox_targets_list, bbox_weights_list,
+ avg_factor) = cls_reg_targets
+
+ num_images = len(batch_img_metas)
+ all_cls_scores = torch.cat([
+ s.permute(0, 2, 3, 1).reshape(
+ num_images, -1, self.cls_out_channels) for s in cls_scores
+ ], 1)
+ all_labels = torch.cat(labels_list, -1).view(num_images, -1)
+ all_label_weights = torch.cat(label_weights_list,
+ -1).view(num_images, -1)
+ all_bbox_preds = torch.cat([
+ b.permute(0, 2, 3, 1).reshape(num_images, -1, 4)
+ for b in bbox_preds
+ ], -2)
+ all_bbox_targets = torch.cat(bbox_targets_list,
+ -2).view(num_images, -1, 4)
+ all_bbox_weights = torch.cat(bbox_weights_list,
+ -2).view(num_images, -1, 4)
+
+ # concat all level anchors to a single tensor
+ all_anchors = []
+ for i in range(num_images):
+ all_anchors.append(torch.cat(anchor_list[i]))
+
+ losses_cls, losses_bbox = multi_apply(
+ self.loss_by_feat_single,
+ all_cls_scores,
+ all_bbox_preds,
+ all_anchors,
+ all_labels,
+ all_label_weights,
+ all_bbox_targets,
+ all_bbox_weights,
+ avg_factor=avg_factor)
+ return dict(loss_cls=losses_cls, loss_bbox=losses_bbox)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/tood_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/tood_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..8c59598d89289df6d1a87c7b6fde112429ac8f45
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/tood_head.py
@@ -0,0 +1,805 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule, Scale
+from mmcv.ops import deform_conv2d
+from mmengine import MessageHub
+from mmengine.config import ConfigDict
+from mmengine.model import bias_init_with_prob, normal_init
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures.bbox import distance2bbox
+from mmdet.utils import (ConfigType, InstanceList, OptConfigType,
+ OptInstanceList, reduce_mean)
+from ..task_modules.prior_generators import anchor_inside_flags
+from ..utils import (filter_scores_and_topk, images_to_levels, multi_apply,
+ sigmoid_geometric_mean, unmap)
+from .atss_head import ATSSHead
+
+
+class TaskDecomposition(nn.Module):
+ """Task decomposition module in task-aligned predictor of TOOD.
+
+ Args:
+ feat_channels (int): Number of feature channels in TOOD head.
+ stacked_convs (int): Number of conv layers in TOOD head.
+ la_down_rate (int): Downsample rate of layer attention.
+ Defaults to 8.
+ conv_cfg (:obj:`ConfigDict` or dict, optional): Config dict for
+ convolution layer. Defaults to None.
+ norm_cfg (:obj:`ConfigDict` or dict, optional): Config dict for
+ normalization layer. Defaults to None.
+ """
+
+ def __init__(self,
+ feat_channels: int,
+ stacked_convs: int,
+ la_down_rate: int = 8,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: OptConfigType = None) -> None:
+ super().__init__()
+ self.feat_channels = feat_channels
+ self.stacked_convs = stacked_convs
+ self.in_channels = self.feat_channels * self.stacked_convs
+ self.norm_cfg = norm_cfg
+ self.layer_attention = nn.Sequential(
+ nn.Conv2d(self.in_channels, self.in_channels // la_down_rate, 1),
+ nn.ReLU(inplace=True),
+ nn.Conv2d(
+ self.in_channels // la_down_rate,
+ self.stacked_convs,
+ 1,
+ padding=0), nn.Sigmoid())
+
+ self.reduction_conv = ConvModule(
+ self.in_channels,
+ self.feat_channels,
+ 1,
+ stride=1,
+ padding=0,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ bias=norm_cfg is None)
+
+ def init_weights(self) -> None:
+ """Initialize the parameters."""
+ for m in self.layer_attention.modules():
+ if isinstance(m, nn.Conv2d):
+ normal_init(m, std=0.001)
+ normal_init(self.reduction_conv.conv, std=0.01)
+
+ def forward(self,
+ feat: Tensor,
+ avg_feat: Optional[Tensor] = None) -> Tensor:
+ """Forward function of task decomposition module."""
+ b, c, h, w = feat.shape
+ if avg_feat is None:
+ avg_feat = F.adaptive_avg_pool2d(feat, (1, 1))
+ weight = self.layer_attention(avg_feat)
+
+ # here we first compute the product between layer attention weight and
+ # conv weight, and then compute the convolution between new conv weight
+ # and feature map, in order to save memory and FLOPs.
+ conv_weight = weight.reshape(
+ b, 1, self.stacked_convs,
+ 1) * self.reduction_conv.conv.weight.reshape(
+ 1, self.feat_channels, self.stacked_convs, self.feat_channels)
+ conv_weight = conv_weight.reshape(b, self.feat_channels,
+ self.in_channels)
+ feat = feat.reshape(b, self.in_channels, h * w)
+ feat = torch.bmm(conv_weight, feat).reshape(b, self.feat_channels, h,
+ w)
+ if self.norm_cfg is not None:
+ feat = self.reduction_conv.norm(feat)
+ feat = self.reduction_conv.activate(feat)
+
+ return feat
+
+
+@MODELS.register_module()
+class TOODHead(ATSSHead):
+ """TOODHead used in `TOOD: Task-aligned One-stage Object Detection.
+
+ `_.
+
+ TOOD uses Task-aligned head (T-head) and is optimized by Task Alignment
+ Learning (TAL).
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ num_dcn (int): Number of deformable convolution in the head.
+ Defaults to 0.
+ anchor_type (str): If set to ``anchor_free``, the head will use centers
+ to regress bboxes. If set to ``anchor_based``, the head will
+ regress bboxes based on anchors. Defaults to ``anchor_free``.
+ initial_loss_cls (:obj:`ConfigDict` or dict): Config of initial loss.
+
+ Example:
+ >>> self = TOODHead(11, 7)
+ >>> feats = [torch.rand(1, 7, s, s) for s in [4, 8, 16, 32, 64]]
+ >>> cls_score, bbox_pred = self.forward(feats)
+ >>> assert len(cls_score) == len(self.scales)
+ """
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: int,
+ num_dcn: int = 0,
+ anchor_type: str = 'anchor_free',
+ initial_loss_cls: ConfigType = dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ activated=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ **kwargs) -> None:
+ assert anchor_type in ['anchor_free', 'anchor_based']
+ self.num_dcn = num_dcn
+ self.anchor_type = anchor_type
+ super().__init__(
+ num_classes=num_classes, in_channels=in_channels, **kwargs)
+
+ if self.train_cfg:
+ self.initial_epoch = self.train_cfg['initial_epoch']
+ self.initial_assigner = TASK_UTILS.build(
+ self.train_cfg['initial_assigner'])
+ self.initial_loss_cls = MODELS.build(initial_loss_cls)
+ self.assigner = self.initial_assigner
+ self.alignment_assigner = TASK_UTILS.build(
+ self.train_cfg['assigner'])
+ self.alpha = self.train_cfg['alpha']
+ self.beta = self.train_cfg['beta']
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self.relu = nn.ReLU(inplace=True)
+ self.inter_convs = nn.ModuleList()
+ for i in range(self.stacked_convs):
+ if i < self.num_dcn:
+ conv_cfg = dict(type='DCNv2', deform_groups=4)
+ else:
+ conv_cfg = self.conv_cfg
+ chn = self.in_channels if i == 0 else self.feat_channels
+ self.inter_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=self.norm_cfg))
+
+ self.cls_decomp = TaskDecomposition(self.feat_channels,
+ self.stacked_convs,
+ self.stacked_convs * 8,
+ self.conv_cfg, self.norm_cfg)
+ self.reg_decomp = TaskDecomposition(self.feat_channels,
+ self.stacked_convs,
+ self.stacked_convs * 8,
+ self.conv_cfg, self.norm_cfg)
+
+ self.tood_cls = nn.Conv2d(
+ self.feat_channels,
+ self.num_base_priors * self.cls_out_channels,
+ 3,
+ padding=1)
+ self.tood_reg = nn.Conv2d(
+ self.feat_channels, self.num_base_priors * 4, 3, padding=1)
+
+ self.cls_prob_module = nn.Sequential(
+ nn.Conv2d(self.feat_channels * self.stacked_convs,
+ self.feat_channels // 4, 1), nn.ReLU(inplace=True),
+ nn.Conv2d(self.feat_channels // 4, 1, 3, padding=1))
+ self.reg_offset_module = nn.Sequential(
+ nn.Conv2d(self.feat_channels * self.stacked_convs,
+ self.feat_channels // 4, 1), nn.ReLU(inplace=True),
+ nn.Conv2d(self.feat_channels // 4, 4 * 2, 3, padding=1))
+
+ self.scales = nn.ModuleList(
+ [Scale(1.0) for _ in self.prior_generator.strides])
+
+ def init_weights(self) -> None:
+ """Initialize weights of the head."""
+ bias_cls = bias_init_with_prob(0.01)
+ for m in self.inter_convs:
+ normal_init(m.conv, std=0.01)
+ for m in self.cls_prob_module:
+ if isinstance(m, nn.Conv2d):
+ normal_init(m, std=0.01)
+ for m in self.reg_offset_module:
+ if isinstance(m, nn.Conv2d):
+ normal_init(m, std=0.001)
+ normal_init(self.cls_prob_module[-1], std=0.01, bias=bias_cls)
+
+ self.cls_decomp.init_weights()
+ self.reg_decomp.init_weights()
+
+ normal_init(self.tood_cls, std=0.01, bias=bias_cls)
+ normal_init(self.tood_reg, std=0.01)
+
+ def forward(self, feats: Tuple[Tensor]) -> Tuple[List[Tensor]]:
+ """Forward features from the upstream network.
+
+ Args:
+ feats (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: Usually a tuple of classification scores and bbox prediction
+ cls_scores (list[Tensor]): Classification scores for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_anchors * num_classes.
+ bbox_preds (list[Tensor]): Decoded box for all scale levels,
+ each is a 4D-tensor, the channels number is
+ num_anchors * 4. In [tl_x, tl_y, br_x, br_y] format.
+ """
+ cls_scores = []
+ bbox_preds = []
+ for idx, (x, scale, stride) in enumerate(
+ zip(feats, self.scales, self.prior_generator.strides)):
+ b, c, h, w = x.shape
+ anchor = self.prior_generator.single_level_grid_priors(
+ (h, w), idx, device=x.device)
+ anchor = torch.cat([anchor for _ in range(b)])
+ # extract task interactive features
+ inter_feats = []
+ for inter_conv in self.inter_convs:
+ x = inter_conv(x)
+ inter_feats.append(x)
+ feat = torch.cat(inter_feats, 1)
+
+ # task decomposition
+ avg_feat = F.adaptive_avg_pool2d(feat, (1, 1))
+ cls_feat = self.cls_decomp(feat, avg_feat)
+ reg_feat = self.reg_decomp(feat, avg_feat)
+
+ # cls prediction and alignment
+ cls_logits = self.tood_cls(cls_feat)
+ cls_prob = self.cls_prob_module(feat)
+ cls_score = sigmoid_geometric_mean(cls_logits, cls_prob)
+
+ # reg prediction and alignment
+ if self.anchor_type == 'anchor_free':
+ reg_dist = scale(self.tood_reg(reg_feat).exp()).float()
+ reg_dist = reg_dist.permute(0, 2, 3, 1).reshape(-1, 4)
+ reg_bbox = distance2bbox(
+ self.anchor_center(anchor) / stride[0],
+ reg_dist).reshape(b, h, w, 4).permute(0, 3, 1,
+ 2) # (b, c, h, w)
+ elif self.anchor_type == 'anchor_based':
+ reg_dist = scale(self.tood_reg(reg_feat)).float()
+ reg_dist = reg_dist.permute(0, 2, 3, 1).reshape(-1, 4)
+ reg_bbox = self.bbox_coder.decode(anchor, reg_dist).reshape(
+ b, h, w, 4).permute(0, 3, 1, 2) / stride[0]
+ else:
+ raise NotImplementedError(
+ f'Unknown anchor type: {self.anchor_type}.'
+ f'Please use `anchor_free` or `anchor_based`.')
+ reg_offset = self.reg_offset_module(feat)
+ bbox_pred = self.deform_sampling(reg_bbox.contiguous(),
+ reg_offset.contiguous())
+
+ # After deform_sampling, some boxes will become invalid (The
+ # left-top point is at the right or bottom of the right-bottom
+ # point), which will make the GIoULoss negative.
+ invalid_bbox_idx = (bbox_pred[:, [0]] > bbox_pred[:, [2]]) | \
+ (bbox_pred[:, [1]] > bbox_pred[:, [3]])
+ invalid_bbox_idx = invalid_bbox_idx.expand_as(bbox_pred)
+ bbox_pred = torch.where(invalid_bbox_idx, reg_bbox, bbox_pred)
+
+ cls_scores.append(cls_score)
+ bbox_preds.append(bbox_pred)
+ return tuple(cls_scores), tuple(bbox_preds)
+
+ def deform_sampling(self, feat: Tensor, offset: Tensor) -> Tensor:
+ """Sampling the feature x according to offset.
+
+ Args:
+ feat (Tensor): Feature
+ offset (Tensor): Spatial offset for feature sampling
+ """
+ # it is an equivalent implementation of bilinear interpolation
+ b, c, h, w = feat.shape
+ weight = feat.new_ones(c, 1, 1, 1)
+ y = deform_conv2d(feat, offset, weight, 1, 0, 1, c, c)
+ return y
+
+ def anchor_center(self, anchors: Tensor) -> Tensor:
+ """Get anchor centers from anchors.
+
+ Args:
+ anchors (Tensor): Anchor list with shape (N, 4), "xyxy" format.
+
+ Returns:
+ Tensor: Anchor centers with shape (N, 2), "xy" format.
+ """
+ anchors_cx = (anchors[:, 2] + anchors[:, 0]) / 2
+ anchors_cy = (anchors[:, 3] + anchors[:, 1]) / 2
+ return torch.stack([anchors_cx, anchors_cy], dim=-1)
+
+ def loss_by_feat_single(self, anchors: Tensor, cls_score: Tensor,
+ bbox_pred: Tensor, labels: Tensor,
+ label_weights: Tensor, bbox_targets: Tensor,
+ alignment_metrics: Tensor,
+ stride: Tuple[int, int]) -> dict:
+ """Calculate the loss of a single scale level based on the features
+ extracted by the detection head.
+
+ Args:
+ anchors (Tensor): Box reference for each scale level with shape
+ (N, num_total_anchors, 4).
+ cls_score (Tensor): Box scores for each scale level
+ Has shape (N, num_anchors * num_classes, H, W).
+ bbox_pred (Tensor): Decoded bboxes for each scale
+ level with shape (N, num_anchors * 4, H, W).
+ labels (Tensor): Labels of each anchors with shape
+ (N, num_total_anchors).
+ label_weights (Tensor): Label weights of each anchor with shape
+ (N, num_total_anchors).
+ bbox_targets (Tensor): BBox regression targets of each anchor with
+ shape (N, num_total_anchors, 4).
+ alignment_metrics (Tensor): Alignment metrics with shape
+ (N, num_total_anchors).
+ stride (Tuple[int, int]): Downsample stride of the feature map.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ assert stride[0] == stride[1], 'h stride is not equal to w stride!'
+ anchors = anchors.reshape(-1, 4)
+ cls_score = cls_score.permute(0, 2, 3, 1).reshape(
+ -1, self.cls_out_channels).contiguous()
+ bbox_pred = bbox_pred.permute(0, 2, 3, 1).reshape(-1, 4)
+ bbox_targets = bbox_targets.reshape(-1, 4)
+ labels = labels.reshape(-1)
+ alignment_metrics = alignment_metrics.reshape(-1)
+ label_weights = label_weights.reshape(-1)
+ targets = labels if self.epoch < self.initial_epoch else (
+ labels, alignment_metrics)
+ cls_loss_func = self.initial_loss_cls \
+ if self.epoch < self.initial_epoch else self.loss_cls
+
+ loss_cls = cls_loss_func(
+ cls_score, targets, label_weights, avg_factor=1.0)
+
+ # FG cat_id: [0, num_classes -1], BG cat_id: num_classes
+ bg_class_ind = self.num_classes
+ pos_inds = ((labels >= 0)
+ & (labels < bg_class_ind)).nonzero().squeeze(1)
+
+ if len(pos_inds) > 0:
+ pos_bbox_targets = bbox_targets[pos_inds]
+ pos_bbox_pred = bbox_pred[pos_inds]
+ pos_anchors = anchors[pos_inds]
+
+ pos_decode_bbox_pred = pos_bbox_pred
+ pos_decode_bbox_targets = pos_bbox_targets / stride[0]
+
+ # regression loss
+ pos_bbox_weight = self.centerness_target(
+ pos_anchors, pos_bbox_targets
+ ) if self.epoch < self.initial_epoch else alignment_metrics[
+ pos_inds]
+
+ loss_bbox = self.loss_bbox(
+ pos_decode_bbox_pred,
+ pos_decode_bbox_targets,
+ weight=pos_bbox_weight,
+ avg_factor=1.0)
+ else:
+ loss_bbox = bbox_pred.sum() * 0
+ pos_bbox_weight = bbox_targets.new_tensor(0.)
+
+ return loss_cls, loss_bbox, alignment_metrics.sum(
+ ), pos_bbox_weight.sum()
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ Has shape (N, num_anchors * num_classes, H, W)
+ bbox_preds (list[Tensor]): Decoded box for each scale
+ level with shape (N, num_anchors * 4, H, W) in
+ [tl_x, tl_y, br_x, br_y] format.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ num_imgs = len(batch_img_metas)
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ assert len(featmap_sizes) == self.prior_generator.num_levels
+
+ device = cls_scores[0].device
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+
+ flatten_cls_scores = torch.cat([
+ cls_score.permute(0, 2, 3, 1).reshape(num_imgs, -1,
+ self.cls_out_channels)
+ for cls_score in cls_scores
+ ], 1)
+ flatten_bbox_preds = torch.cat([
+ bbox_pred.permute(0, 2, 3, 1).reshape(num_imgs, -1, 4) * stride[0]
+ for bbox_pred, stride in zip(bbox_preds,
+ self.prior_generator.strides)
+ ], 1)
+
+ cls_reg_targets = self.get_targets(
+ flatten_cls_scores,
+ flatten_bbox_preds,
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore)
+ (anchor_list, labels_list, label_weights_list, bbox_targets_list,
+ alignment_metrics_list) = cls_reg_targets
+
+ losses_cls, losses_bbox, \
+ cls_avg_factors, bbox_avg_factors = multi_apply(
+ self.loss_by_feat_single,
+ anchor_list,
+ cls_scores,
+ bbox_preds,
+ labels_list,
+ label_weights_list,
+ bbox_targets_list,
+ alignment_metrics_list,
+ self.prior_generator.strides)
+
+ cls_avg_factor = reduce_mean(sum(cls_avg_factors)).clamp_(min=1).item()
+ losses_cls = list(map(lambda x: x / cls_avg_factor, losses_cls))
+
+ bbox_avg_factor = reduce_mean(
+ sum(bbox_avg_factors)).clamp_(min=1).item()
+ losses_bbox = list(map(lambda x: x / bbox_avg_factor, losses_bbox))
+ return dict(loss_cls=losses_cls, loss_bbox=losses_bbox)
+
+ def _predict_by_feat_single(self,
+ cls_score_list: List[Tensor],
+ bbox_pred_list: List[Tensor],
+ score_factor_list: List[Tensor],
+ mlvl_priors: List[Tensor],
+ img_meta: dict,
+ cfg: Optional[ConfigDict] = None,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results.
+
+ Args:
+ cls_score_list (list[Tensor]): Box scores from all scale
+ levels of a single image, each item has shape
+ (num_priors * num_classes, H, W).
+ bbox_pred_list (list[Tensor]): Box energies / deltas from
+ all scale levels of a single image, each item has shape
+ (num_priors * 4, H, W).
+ score_factor_list (list[Tensor]): Score factor from all scale
+ levels of a single image, each item has shape
+ (num_priors * 1, H, W).
+ mlvl_priors (list[Tensor]): Each element in the list is
+ the priors of a single level in feature pyramid. In all
+ anchor-based methods, it has shape (num_priors, 4). In
+ all anchor-free methods, it has shape (num_priors, 2)
+ when `with_stride=True`, otherwise it still has shape
+ (num_priors, 4).
+ img_meta (dict): Image meta info.
+ cfg (:obj:`ConfigDict`, optional): Test / postprocessing
+ configuration, if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ tuple[Tensor]: Results of detected bboxes and labels. If with_nms
+ is False and mlvl_score_factor is None, return mlvl_bboxes and
+ mlvl_scores, else return mlvl_bboxes, mlvl_scores and
+ mlvl_score_factor. Usually with_nms is False is used for aug
+ test. If with_nms is True, then return the following format
+
+ - det_bboxes (Tensor): Predicted bboxes with shape \
+ [num_bboxes, 5], where the first 4 columns are bounding \
+ box positions (tl_x, tl_y, br_x, br_y) and the 5-th \
+ column are scores between 0 and 1.
+ - det_labels (Tensor): Predicted labels of the corresponding \
+ box with shape [num_bboxes].
+ """
+
+ cfg = self.test_cfg if cfg is None else cfg
+ nms_pre = cfg.get('nms_pre', -1)
+
+ mlvl_bboxes = []
+ mlvl_scores = []
+ mlvl_labels = []
+ for cls_score, bbox_pred, priors, stride in zip(
+ cls_score_list, bbox_pred_list, mlvl_priors,
+ self.prior_generator.strides):
+ assert cls_score.size()[-2:] == bbox_pred.size()[-2:]
+
+ bbox_pred = bbox_pred.permute(1, 2, 0).reshape(-1, 4) * stride[0]
+ scores = cls_score.permute(1, 2,
+ 0).reshape(-1, self.cls_out_channels)
+
+ # After https://github.com/open-mmlab/mmdetection/pull/6268/,
+ # this operation keeps fewer bboxes under the same `nms_pre`.
+ # There is no difference in performance for most models. If you
+ # find a slight drop in performance, you can set a larger
+ # `nms_pre` than before.
+ results = filter_scores_and_topk(
+ scores, cfg.score_thr, nms_pre,
+ dict(bbox_pred=bbox_pred, priors=priors))
+ scores, labels, keep_idxs, filtered_results = results
+
+ bboxes = filtered_results['bbox_pred']
+
+ mlvl_bboxes.append(bboxes)
+ mlvl_scores.append(scores)
+ mlvl_labels.append(labels)
+
+ results = InstanceData()
+ results.bboxes = torch.cat(mlvl_bboxes)
+ results.scores = torch.cat(mlvl_scores)
+ results.labels = torch.cat(mlvl_labels)
+
+ return self._bbox_post_process(
+ results=results,
+ cfg=cfg,
+ rescale=rescale,
+ with_nms=with_nms,
+ img_meta=img_meta)
+
+ def get_targets(self,
+ cls_scores: List[List[Tensor]],
+ bbox_preds: List[List[Tensor]],
+ anchor_list: List[List[Tensor]],
+ valid_flag_list: List[List[Tensor]],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None,
+ unmap_outputs: bool = True) -> tuple:
+ """Compute regression and classification targets for anchors in
+ multiple images.
+
+ Args:
+ cls_scores (list[list[Tensor]]): Classification predictions of
+ images, a 3D-Tensor with shape [num_imgs, num_priors,
+ num_classes].
+ bbox_preds (list[list[Tensor]]): Decoded bboxes predictions of one
+ image, a 3D-Tensor with shape [num_imgs, num_priors, 4] in
+ [tl_x, tl_y, br_x, br_y] format.
+ anchor_list (list[list[Tensor]]): Multi level anchors of each
+ image. The outer list indicates images, and the inner list
+ corresponds to feature levels of the image. Each element of
+ the inner list is a tensor of shape (num_anchors, 4).
+ valid_flag_list (list[list[Tensor]]): Multi level valid flags of
+ each image. The outer list indicates images, and the inner list
+ corresponds to feature levels of the image. Each element of
+ the inner list is a tensor of shape (num_anchors, )
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ unmap_outputs (bool): Whether to map outputs back to the original
+ set of anchors.
+
+ Returns:
+ tuple: a tuple containing learning targets.
+
+ - anchors_list (list[list[Tensor]]): Anchors of each level.
+ - labels_list (list[Tensor]): Labels of each level.
+ - label_weights_list (list[Tensor]): Label weights of each
+ level.
+ - bbox_targets_list (list[Tensor]): BBox targets of each level.
+ - norm_alignment_metrics_list (list[Tensor]): Normalized
+ alignment metrics of each level.
+ """
+ num_imgs = len(batch_img_metas)
+ assert len(anchor_list) == len(valid_flag_list) == num_imgs
+
+ # anchor number of multi levels
+ num_level_anchors = [anchors.size(0) for anchors in anchor_list[0]]
+ num_level_anchors_list = [num_level_anchors] * num_imgs
+
+ # concat all level anchors and flags to a single tensor
+ for i in range(num_imgs):
+ assert len(anchor_list[i]) == len(valid_flag_list[i])
+ anchor_list[i] = torch.cat(anchor_list[i])
+ valid_flag_list[i] = torch.cat(valid_flag_list[i])
+
+ # compute targets for each image
+ if batch_gt_instances_ignore is None:
+ batch_gt_instances_ignore = [None] * num_imgs
+ # anchor_list: list(b * [-1, 4])
+
+ # get epoch information from message hub
+ message_hub = MessageHub.get_current_instance()
+ self.epoch = message_hub.get_info('epoch')
+
+ if self.epoch < self.initial_epoch:
+ (all_anchors, all_labels, all_label_weights, all_bbox_targets,
+ all_bbox_weights, pos_inds_list, neg_inds_list,
+ sampling_result) = multi_apply(
+ super()._get_targets_single,
+ anchor_list,
+ valid_flag_list,
+ num_level_anchors_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore,
+ unmap_outputs=unmap_outputs)
+ all_assign_metrics = [
+ weight[..., 0] for weight in all_bbox_weights
+ ]
+ else:
+ (all_anchors, all_labels, all_label_weights, all_bbox_targets,
+ all_assign_metrics) = multi_apply(
+ self._get_targets_single,
+ cls_scores,
+ bbox_preds,
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore,
+ unmap_outputs=unmap_outputs)
+
+ # split targets to a list w.r.t. multiple levels
+ anchors_list = images_to_levels(all_anchors, num_level_anchors)
+ labels_list = images_to_levels(all_labels, num_level_anchors)
+ label_weights_list = images_to_levels(all_label_weights,
+ num_level_anchors)
+ bbox_targets_list = images_to_levels(all_bbox_targets,
+ num_level_anchors)
+ norm_alignment_metrics_list = images_to_levels(all_assign_metrics,
+ num_level_anchors)
+
+ return (anchors_list, labels_list, label_weights_list,
+ bbox_targets_list, norm_alignment_metrics_list)
+
+ def _get_targets_single(self,
+ cls_scores: Tensor,
+ bbox_preds: Tensor,
+ flat_anchors: Tensor,
+ valid_flags: Tensor,
+ gt_instances: InstanceData,
+ img_meta: dict,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ unmap_outputs: bool = True) -> tuple:
+ """Compute regression, classification targets for anchors in a single
+ image.
+
+ Args:
+ cls_scores (Tensor): Box scores for each image.
+ bbox_preds (Tensor): Box energies / deltas for each image.
+ flat_anchors (Tensor): Multi-level anchors of the image, which are
+ concatenated into a single tensor of shape (num_anchors ,4)
+ valid_flags (Tensor): Multi level valid flags of the image,
+ which are concatenated into a single tensor of
+ shape (num_anchors,).
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ img_meta (dict): Meta information for current image.
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ unmap_outputs (bool): Whether to map outputs back to the original
+ set of anchors.
+
+ Returns:
+ tuple: N is the number of total anchors in the image.
+ anchors (Tensor): All anchors in the image with shape (N, 4).
+ labels (Tensor): Labels of all anchors in the image with shape
+ (N,).
+ label_weights (Tensor): Label weights of all anchor in the
+ image with shape (N,).
+ bbox_targets (Tensor): BBox targets of all anchors in the
+ image with shape (N, 4).
+ norm_alignment_metrics (Tensor): Normalized alignment metrics
+ of all priors in the image with shape (N,).
+ """
+ inside_flags = anchor_inside_flags(flat_anchors, valid_flags,
+ img_meta['img_shape'][:2],
+ self.train_cfg['allowed_border'])
+ if not inside_flags.any():
+ raise ValueError(
+ 'There is no valid anchor inside the image boundary. Please '
+ 'check the image size and anchor sizes, or set '
+ '``allowed_border`` to -1 to skip the condition.')
+ # assign gt and sample anchors
+ anchors = flat_anchors[inside_flags, :]
+ pred_instances = InstanceData(
+ priors=anchors,
+ scores=cls_scores[inside_flags, :],
+ bboxes=bbox_preds[inside_flags, :])
+ assign_result = self.alignment_assigner.assign(pred_instances,
+ gt_instances,
+ gt_instances_ignore,
+ self.alpha, self.beta)
+ assign_ious = assign_result.max_overlaps
+ assign_metrics = assign_result.assign_metrics
+
+ sampling_result = self.sampler.sample(assign_result, pred_instances,
+ gt_instances)
+
+ num_valid_anchors = anchors.shape[0]
+ bbox_targets = torch.zeros_like(anchors)
+ labels = anchors.new_full((num_valid_anchors, ),
+ self.num_classes,
+ dtype=torch.long)
+ label_weights = anchors.new_zeros(num_valid_anchors, dtype=torch.float)
+ norm_alignment_metrics = anchors.new_zeros(
+ num_valid_anchors, dtype=torch.float)
+
+ pos_inds = sampling_result.pos_inds
+ neg_inds = sampling_result.neg_inds
+ if len(pos_inds) > 0:
+ # point-based
+ pos_bbox_targets = sampling_result.pos_gt_bboxes
+ bbox_targets[pos_inds, :] = pos_bbox_targets
+
+ labels[pos_inds] = sampling_result.pos_gt_labels
+ if self.train_cfg['pos_weight'] <= 0:
+ label_weights[pos_inds] = 1.0
+ else:
+ label_weights[pos_inds] = self.train_cfg['pos_weight']
+ if len(neg_inds) > 0:
+ label_weights[neg_inds] = 1.0
+
+ class_assigned_gt_inds = torch.unique(
+ sampling_result.pos_assigned_gt_inds)
+ for gt_inds in class_assigned_gt_inds:
+ gt_class_inds = pos_inds[sampling_result.pos_assigned_gt_inds ==
+ gt_inds]
+ pos_alignment_metrics = assign_metrics[gt_class_inds]
+ pos_ious = assign_ious[gt_class_inds]
+ pos_norm_alignment_metrics = pos_alignment_metrics / (
+ pos_alignment_metrics.max() + 10e-8) * pos_ious.max()
+ norm_alignment_metrics[gt_class_inds] = pos_norm_alignment_metrics
+
+ # map up to original set of anchors
+ if unmap_outputs:
+ num_total_anchors = flat_anchors.size(0)
+ anchors = unmap(anchors, num_total_anchors, inside_flags)
+ labels = unmap(
+ labels, num_total_anchors, inside_flags, fill=self.num_classes)
+ label_weights = unmap(label_weights, num_total_anchors,
+ inside_flags)
+ bbox_targets = unmap(bbox_targets, num_total_anchors, inside_flags)
+ norm_alignment_metrics = unmap(norm_alignment_metrics,
+ num_total_anchors, inside_flags)
+ return (anchors, labels, label_weights, bbox_targets,
+ norm_alignment_metrics)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/vfnet_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/vfnet_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..430b06d085d94760d56a7ea083eaf23bd32b1f53
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/vfnet_head.py
@@ -0,0 +1,722 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple, Union
+
+import numpy as np
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule, Scale
+from mmcv.ops import DeformConv2d
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures.bbox import bbox_overlaps
+from mmdet.utils import (ConfigType, InstanceList, MultiConfig,
+ OptInstanceList, RangeType, reduce_mean)
+from ..task_modules.prior_generators import MlvlPointGenerator
+from ..task_modules.samplers import PseudoSampler
+from ..utils import multi_apply
+from .atss_head import ATSSHead
+from .fcos_head import FCOSHead
+
+INF = 1e8
+
+
+@MODELS.register_module()
+class VFNetHead(ATSSHead, FCOSHead):
+ """Head of `VarifocalNet (VFNet): An IoU-aware Dense Object
+ Detector.`_.
+
+ The VFNet predicts IoU-aware classification scores which mix the
+ object presence confidence and object localization accuracy as the
+ detection score. It is built on the FCOS architecture and uses ATSS
+ for defining positive/negative training examples. The VFNet is trained
+ with Varifocal Loss and empolys star-shaped deformable convolution to
+ extract features for a bbox.
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ regress_ranges (Sequence[Tuple[int, int]]): Regress range of multiple
+ level points.
+ center_sampling (bool): If true, use center sampling. Defaults to False.
+ center_sample_radius (float): Radius of center sampling. Defaults to 1.5.
+ sync_num_pos (bool): If true, synchronize the number of positive
+ examples across GPUs. Defaults to True
+ gradient_mul (float): The multiplier to gradients from bbox refinement
+ and recognition. Defaults to 0.1.
+ bbox_norm_type (str): The bbox normalization type, 'reg_denom' or
+ 'stride'. Defaults to reg_denom
+ loss_cls_fl (:obj:`ConfigDict` or dict): Config of focal loss.
+ use_vfl (bool): If true, use varifocal loss for training.
+ Defaults to True.
+ loss_cls (:obj:`ConfigDict` or dict): Config of varifocal loss.
+ loss_bbox (:obj:`ConfigDict` or dict): Config of localization loss,
+ GIoU Loss.
+ loss_bbox (:obj:`ConfigDict` or dict): Config of localization
+ refinement loss, GIoU Loss.
+ norm_cfg (:obj:`ConfigDict` or dict): dictionary to construct and
+ config norm layer. Defaults to norm_cfg=dict(type='GN',
+ num_groups=32, requires_grad=True).
+ use_atss (bool): If true, use ATSS to define positive/negative
+ examples. Defaults to True.
+ anchor_generator (:obj:`ConfigDict` or dict): Config of anchor
+ generator for ATSS.
+ init_cfg (:obj:`ConfigDict` or dict or list[dict] or
+ list[:obj:`ConfigDict`]): Initialization config dict.
+
+ Example:
+ >>> self = VFNetHead(11, 7)
+ >>> feats = [torch.rand(1, 7, s, s) for s in [4, 8, 16, 32, 64]]
+ >>> cls_score, bbox_pred, bbox_pred_refine= self.forward(feats)
+ >>> assert len(cls_score) == len(self.scales)
+ """ # noqa: E501
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: int,
+ regress_ranges: RangeType = ((-1, 64), (64, 128), (128, 256),
+ (256, 512), (512, INF)),
+ center_sampling: bool = False,
+ center_sample_radius: float = 1.5,
+ sync_num_pos: bool = True,
+ gradient_mul: float = 0.1,
+ bbox_norm_type: str = 'reg_denom',
+ loss_cls_fl: ConfigType = dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0),
+ use_vfl: bool = True,
+ loss_cls: ConfigType = dict(
+ type='VarifocalLoss',
+ use_sigmoid=True,
+ alpha=0.75,
+ gamma=2.0,
+ iou_weighted=True,
+ loss_weight=1.0),
+ loss_bbox: ConfigType = dict(
+ type='GIoULoss', loss_weight=1.5),
+ loss_bbox_refine: ConfigType = dict(
+ type='GIoULoss', loss_weight=2.0),
+ norm_cfg: ConfigType = dict(
+ type='GN', num_groups=32, requires_grad=True),
+ use_atss: bool = True,
+ reg_decoded_bbox: bool = True,
+ anchor_generator: ConfigType = dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ octave_base_scale=8,
+ scales_per_octave=1,
+ center_offset=0.0,
+ strides=[8, 16, 32, 64, 128]),
+ init_cfg: MultiConfig = dict(
+ type='Normal',
+ layer='Conv2d',
+ std=0.01,
+ override=dict(
+ type='Normal',
+ name='vfnet_cls',
+ std=0.01,
+ bias_prob=0.01)),
+ **kwargs) -> None:
+ # dcn base offsets, adapted from reppoints_head.py
+ self.num_dconv_points = 9
+ self.dcn_kernel = int(np.sqrt(self.num_dconv_points))
+ self.dcn_pad = int((self.dcn_kernel - 1) / 2)
+ dcn_base = np.arange(-self.dcn_pad,
+ self.dcn_pad + 1).astype(np.float64)
+ dcn_base_y = np.repeat(dcn_base, self.dcn_kernel)
+ dcn_base_x = np.tile(dcn_base, self.dcn_kernel)
+ dcn_base_offset = np.stack([dcn_base_y, dcn_base_x], axis=1).reshape(
+ (-1))
+ self.dcn_base_offset = torch.tensor(dcn_base_offset).view(1, -1, 1, 1)
+
+ super(FCOSHead, self).__init__(
+ num_classes=num_classes,
+ in_channels=in_channels,
+ norm_cfg=norm_cfg,
+ init_cfg=init_cfg,
+ **kwargs)
+ self.regress_ranges = regress_ranges
+ self.reg_denoms = [
+ regress_range[-1] for regress_range in regress_ranges
+ ]
+ self.reg_denoms[-1] = self.reg_denoms[-2] * 2
+ self.center_sampling = center_sampling
+ self.center_sample_radius = center_sample_radius
+ self.sync_num_pos = sync_num_pos
+ self.bbox_norm_type = bbox_norm_type
+ self.gradient_mul = gradient_mul
+ self.use_vfl = use_vfl
+ if self.use_vfl:
+ self.loss_cls = MODELS.build(loss_cls)
+ else:
+ self.loss_cls = MODELS.build(loss_cls_fl)
+ self.loss_bbox = MODELS.build(loss_bbox)
+ self.loss_bbox_refine = MODELS.build(loss_bbox_refine)
+
+ # for getting ATSS targets
+ self.use_atss = use_atss
+ self.reg_decoded_bbox = reg_decoded_bbox
+ self.use_sigmoid_cls = loss_cls.get('use_sigmoid', False)
+
+ self.anchor_center_offset = anchor_generator['center_offset']
+
+ self.num_base_priors = self.prior_generator.num_base_priors[0]
+
+ if self.train_cfg:
+ self.assigner = TASK_UTILS.build(self.train_cfg['assigner'])
+ if self.train_cfg.get('sampler', None) is not None:
+ self.sampler = TASK_UTILS.build(
+ self.train_cfg['sampler'], default_args=dict(context=self))
+ else:
+ self.sampler = PseudoSampler()
+ # only be used in `get_atss_targets` when `use_atss` is True
+ self.atss_prior_generator = TASK_UTILS.build(anchor_generator)
+
+ self.fcos_prior_generator = MlvlPointGenerator(
+ anchor_generator['strides'],
+ self.anchor_center_offset if self.use_atss else 0.5)
+
+ # In order to reuse the `get_bboxes` in `BaseDenseHead.
+ # Only be used in testing phase.
+ self.prior_generator = self.fcos_prior_generator
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ super(FCOSHead, self)._init_cls_convs()
+ super(FCOSHead, self)._init_reg_convs()
+ self.relu = nn.ReLU()
+ self.vfnet_reg_conv = ConvModule(
+ self.feat_channels,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ bias=self.conv_bias)
+ self.vfnet_reg = nn.Conv2d(self.feat_channels, 4, 3, padding=1)
+ self.scales = nn.ModuleList([Scale(1.0) for _ in self.strides])
+
+ self.vfnet_reg_refine_dconv = DeformConv2d(
+ self.feat_channels,
+ self.feat_channels,
+ self.dcn_kernel,
+ 1,
+ padding=self.dcn_pad)
+ self.vfnet_reg_refine = nn.Conv2d(self.feat_channels, 4, 3, padding=1)
+ self.scales_refine = nn.ModuleList([Scale(1.0) for _ in self.strides])
+
+ self.vfnet_cls_dconv = DeformConv2d(
+ self.feat_channels,
+ self.feat_channels,
+ self.dcn_kernel,
+ 1,
+ padding=self.dcn_pad)
+ self.vfnet_cls = nn.Conv2d(
+ self.feat_channels, self.cls_out_channels, 3, padding=1)
+
+ def forward(self, x: Tuple[Tensor]) -> Tuple[List[Tensor]]:
+ """Forward features from the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple:
+
+ - cls_scores (list[Tensor]): Box iou-aware scores for each scale
+ level, each is a 4D-tensor, the channel number is
+ num_points * num_classes.
+ - bbox_preds (list[Tensor]): Box offsets for each
+ scale level, each is a 4D-tensor, the channel number is
+ num_points * 4.
+ - bbox_preds_refine (list[Tensor]): Refined Box offsets for
+ each scale level, each is a 4D-tensor, the channel
+ number is num_points * 4.
+ """
+ return multi_apply(self.forward_single, x, self.scales,
+ self.scales_refine, self.strides, self.reg_denoms)
+
+ def forward_single(self, x: Tensor, scale: Scale, scale_refine: Scale,
+ stride: int, reg_denom: int) -> tuple:
+ """Forward features of a single scale level.
+
+ Args:
+ x (Tensor): FPN feature maps of the specified stride.
+ scale (:obj: `mmcv.cnn.Scale`): Learnable scale module to resize
+ the bbox prediction.
+ scale_refine (:obj: `mmcv.cnn.Scale`): Learnable scale module to
+ resize the refined bbox prediction.
+ stride (int): The corresponding stride for feature maps,
+ used to normalize the bbox prediction when
+ bbox_norm_type = 'stride'.
+ reg_denom (int): The corresponding regression range for feature
+ maps, only used to normalize the bbox prediction when
+ bbox_norm_type = 'reg_denom'.
+
+ Returns:
+ tuple: iou-aware cls scores for each box, bbox predictions and
+ refined bbox predictions of input feature maps.
+ """
+ cls_feat = x
+ reg_feat = x
+
+ for cls_layer in self.cls_convs:
+ cls_feat = cls_layer(cls_feat)
+
+ for reg_layer in self.reg_convs:
+ reg_feat = reg_layer(reg_feat)
+
+ # predict the bbox_pred of different level
+ reg_feat_init = self.vfnet_reg_conv(reg_feat)
+ if self.bbox_norm_type == 'reg_denom':
+ bbox_pred = scale(
+ self.vfnet_reg(reg_feat_init)).float().exp() * reg_denom
+ elif self.bbox_norm_type == 'stride':
+ bbox_pred = scale(
+ self.vfnet_reg(reg_feat_init)).float().exp() * stride
+ else:
+ raise NotImplementedError
+
+ # compute star deformable convolution offsets
+ # converting dcn_offset to reg_feat.dtype thus VFNet can be
+ # trained with FP16
+ dcn_offset = self.star_dcn_offset(bbox_pred, self.gradient_mul,
+ stride).to(reg_feat.dtype)
+
+ # refine the bbox_pred
+ reg_feat = self.relu(self.vfnet_reg_refine_dconv(reg_feat, dcn_offset))
+ bbox_pred_refine = scale_refine(
+ self.vfnet_reg_refine(reg_feat)).float().exp()
+ bbox_pred_refine = bbox_pred_refine * bbox_pred.detach()
+
+ # predict the iou-aware cls score
+ cls_feat = self.relu(self.vfnet_cls_dconv(cls_feat, dcn_offset))
+ cls_score = self.vfnet_cls(cls_feat)
+
+ if self.training:
+ return cls_score, bbox_pred, bbox_pred_refine
+ else:
+ return cls_score, bbox_pred_refine
+
+ def star_dcn_offset(self, bbox_pred: Tensor, gradient_mul: float,
+ stride: int) -> Tensor:
+ """Compute the star deformable conv offsets.
+
+ Args:
+ bbox_pred (Tensor): Predicted bbox distance offsets (l, r, t, b).
+ gradient_mul (float): Gradient multiplier.
+ stride (int): The corresponding stride for feature maps,
+ used to project the bbox onto the feature map.
+
+ Returns:
+ Tensor: The offsets for deformable convolution.
+ """
+ dcn_base_offset = self.dcn_base_offset.type_as(bbox_pred)
+ bbox_pred_grad_mul = (1 - gradient_mul) * bbox_pred.detach() + \
+ gradient_mul * bbox_pred
+ # map to the feature map scale
+ bbox_pred_grad_mul = bbox_pred_grad_mul / stride
+ N, C, H, W = bbox_pred.size()
+
+ x1 = bbox_pred_grad_mul[:, 0, :, :]
+ y1 = bbox_pred_grad_mul[:, 1, :, :]
+ x2 = bbox_pred_grad_mul[:, 2, :, :]
+ y2 = bbox_pred_grad_mul[:, 3, :, :]
+ bbox_pred_grad_mul_offset = bbox_pred.new_zeros(
+ N, 2 * self.num_dconv_points, H, W)
+ bbox_pred_grad_mul_offset[:, 0, :, :] = -1.0 * y1 # -y1
+ bbox_pred_grad_mul_offset[:, 1, :, :] = -1.0 * x1 # -x1
+ bbox_pred_grad_mul_offset[:, 2, :, :] = -1.0 * y1 # -y1
+ bbox_pred_grad_mul_offset[:, 4, :, :] = -1.0 * y1 # -y1
+ bbox_pred_grad_mul_offset[:, 5, :, :] = x2 # x2
+ bbox_pred_grad_mul_offset[:, 7, :, :] = -1.0 * x1 # -x1
+ bbox_pred_grad_mul_offset[:, 11, :, :] = x2 # x2
+ bbox_pred_grad_mul_offset[:, 12, :, :] = y2 # y2
+ bbox_pred_grad_mul_offset[:, 13, :, :] = -1.0 * x1 # -x1
+ bbox_pred_grad_mul_offset[:, 14, :, :] = y2 # y2
+ bbox_pred_grad_mul_offset[:, 16, :, :] = y2 # y2
+ bbox_pred_grad_mul_offset[:, 17, :, :] = x2 # x2
+ dcn_offset = bbox_pred_grad_mul_offset - dcn_base_offset
+
+ return dcn_offset
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ bbox_preds_refine: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Compute loss of the head.
+
+ Args:
+ cls_scores (list[Tensor]): Box iou-aware scores for each scale
+ level, each is a 4D-tensor, the channel number is
+ num_points * num_classes.
+ bbox_preds (list[Tensor]): Box offsets for each
+ scale level, each is a 4D-tensor, the channel number is
+ num_points * 4.
+ bbox_preds_refine (list[Tensor]): Refined Box offsets for
+ each scale level, each is a 4D-tensor, the channel
+ number is num_points * 4.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ assert len(cls_scores) == len(bbox_preds) == len(bbox_preds_refine)
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ all_level_points = self.fcos_prior_generator.grid_priors(
+ featmap_sizes, bbox_preds[0].dtype, bbox_preds[0].device)
+ labels, label_weights, bbox_targets, bbox_weights = self.get_targets(
+ cls_scores,
+ all_level_points,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore)
+
+ num_imgs = cls_scores[0].size(0)
+ # flatten cls_scores, bbox_preds and bbox_preds_refine
+ flatten_cls_scores = [
+ cls_score.permute(0, 2, 3,
+ 1).reshape(-1,
+ self.cls_out_channels).contiguous()
+ for cls_score in cls_scores
+ ]
+ flatten_bbox_preds = [
+ bbox_pred.permute(0, 2, 3, 1).reshape(-1, 4).contiguous()
+ for bbox_pred in bbox_preds
+ ]
+ flatten_bbox_preds_refine = [
+ bbox_pred_refine.permute(0, 2, 3, 1).reshape(-1, 4).contiguous()
+ for bbox_pred_refine in bbox_preds_refine
+ ]
+ flatten_cls_scores = torch.cat(flatten_cls_scores)
+ flatten_bbox_preds = torch.cat(flatten_bbox_preds)
+ flatten_bbox_preds_refine = torch.cat(flatten_bbox_preds_refine)
+ flatten_labels = torch.cat(labels)
+ flatten_bbox_targets = torch.cat(bbox_targets)
+ # repeat points to align with bbox_preds
+ flatten_points = torch.cat(
+ [points.repeat(num_imgs, 1) for points in all_level_points])
+
+ # FG cat_id: [0, num_classes - 1], BG cat_id: num_classes
+ bg_class_ind = self.num_classes
+ pos_inds = torch.where(
+ ((flatten_labels >= 0) & (flatten_labels < bg_class_ind)) > 0)[0]
+ num_pos = len(pos_inds)
+
+ pos_bbox_preds = flatten_bbox_preds[pos_inds]
+ pos_bbox_preds_refine = flatten_bbox_preds_refine[pos_inds]
+ pos_labels = flatten_labels[pos_inds]
+
+ # sync num_pos across all gpus
+ if self.sync_num_pos:
+ num_pos_avg_per_gpu = reduce_mean(
+ pos_inds.new_tensor(num_pos).float()).item()
+ num_pos_avg_per_gpu = max(num_pos_avg_per_gpu, 1.0)
+ else:
+ num_pos_avg_per_gpu = num_pos
+
+ pos_bbox_targets = flatten_bbox_targets[pos_inds]
+ pos_points = flatten_points[pos_inds]
+
+ pos_decoded_bbox_preds = self.bbox_coder.decode(
+ pos_points, pos_bbox_preds)
+ pos_decoded_target_preds = self.bbox_coder.decode(
+ pos_points, pos_bbox_targets)
+ iou_targets_ini = bbox_overlaps(
+ pos_decoded_bbox_preds,
+ pos_decoded_target_preds.detach(),
+ is_aligned=True).clamp(min=1e-6)
+ bbox_weights_ini = iou_targets_ini.clone().detach()
+ bbox_avg_factor_ini = reduce_mean(
+ bbox_weights_ini.sum()).clamp_(min=1).item()
+
+ pos_decoded_bbox_preds_refine = \
+ self.bbox_coder.decode(pos_points, pos_bbox_preds_refine)
+ iou_targets_rf = bbox_overlaps(
+ pos_decoded_bbox_preds_refine,
+ pos_decoded_target_preds.detach(),
+ is_aligned=True).clamp(min=1e-6)
+ bbox_weights_rf = iou_targets_rf.clone().detach()
+ bbox_avg_factor_rf = reduce_mean(
+ bbox_weights_rf.sum()).clamp_(min=1).item()
+
+ if num_pos > 0:
+ loss_bbox = self.loss_bbox(
+ pos_decoded_bbox_preds,
+ pos_decoded_target_preds.detach(),
+ weight=bbox_weights_ini,
+ avg_factor=bbox_avg_factor_ini)
+
+ loss_bbox_refine = self.loss_bbox_refine(
+ pos_decoded_bbox_preds_refine,
+ pos_decoded_target_preds.detach(),
+ weight=bbox_weights_rf,
+ avg_factor=bbox_avg_factor_rf)
+
+ # build IoU-aware cls_score targets
+ if self.use_vfl:
+ pos_ious = iou_targets_rf.clone().detach()
+ cls_iou_targets = torch.zeros_like(flatten_cls_scores)
+ cls_iou_targets[pos_inds, pos_labels] = pos_ious
+ else:
+ loss_bbox = pos_bbox_preds.sum() * 0
+ loss_bbox_refine = pos_bbox_preds_refine.sum() * 0
+ if self.use_vfl:
+ cls_iou_targets = torch.zeros_like(flatten_cls_scores)
+
+ if self.use_vfl:
+ loss_cls = self.loss_cls(
+ flatten_cls_scores,
+ cls_iou_targets,
+ avg_factor=num_pos_avg_per_gpu)
+ else:
+ loss_cls = self.loss_cls(
+ flatten_cls_scores,
+ flatten_labels,
+ weight=label_weights,
+ avg_factor=num_pos_avg_per_gpu)
+
+ return dict(
+ loss_cls=loss_cls,
+ loss_bbox=loss_bbox,
+ loss_bbox_rf=loss_bbox_refine)
+
+ def get_targets(
+ self,
+ cls_scores: List[Tensor],
+ mlvl_points: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> tuple:
+ """A wrapper for computing ATSS and FCOS targets for points in multiple
+ images.
+
+ Args:
+ cls_scores (list[Tensor]): Box iou-aware scores for each scale
+ level with shape (N, num_points * num_classes, H, W).
+ mlvl_points (list[Tensor]): Points of each fpn level, each has
+ shape (num_points, 2).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ tuple:
+
+ - labels_list (list[Tensor]): Labels of each level.
+ - label_weights (Tensor/None): Label weights of all levels.
+ - bbox_targets_list (list[Tensor]): Regression targets of each
+ level, (l, t, r, b).
+ - bbox_weights (Tensor/None): Bbox weights of all levels.
+ """
+ if self.use_atss:
+ return self.get_atss_targets(cls_scores, mlvl_points,
+ batch_gt_instances, batch_img_metas,
+ batch_gt_instances_ignore)
+ else:
+ self.norm_on_bbox = False
+ return self.get_fcos_targets(mlvl_points, batch_gt_instances)
+
+ def _get_targets_single(self, *args, **kwargs):
+ """Avoid ambiguity in multiple inheritance."""
+ if self.use_atss:
+ return ATSSHead._get_targets_single(self, *args, **kwargs)
+ else:
+ return FCOSHead._get_targets_single(self, *args, **kwargs)
+
+ def get_fcos_targets(self, points: List[Tensor],
+ batch_gt_instances: InstanceList) -> tuple:
+ """Compute FCOS regression and classification targets for points in
+ multiple images.
+
+ Args:
+ points (list[Tensor]): Points of each fpn level, each has shape
+ (num_points, 2).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+
+ Returns:
+ tuple:
+
+ - labels (list[Tensor]): Labels of each level.
+ - label_weights: None, to be compatible with ATSS targets.
+ - bbox_targets (list[Tensor]): BBox targets of each level.
+ - bbox_weights: None, to be compatible with ATSS targets.
+ """
+ labels, bbox_targets = FCOSHead.get_targets(self, points,
+ batch_gt_instances)
+ label_weights = None
+ bbox_weights = None
+ return labels, label_weights, bbox_targets, bbox_weights
+
+ def get_anchors(self,
+ featmap_sizes: List[Tuple],
+ batch_img_metas: List[dict],
+ device: str = 'cuda') -> tuple:
+ """Get anchors according to feature map sizes.
+
+ Args:
+ featmap_sizes (list[tuple]): Multi-level feature map sizes.
+ batch_img_metas (list[dict]): Image meta info.
+ device (str): Device for returned tensors
+
+ Returns:
+ tuple:
+
+ - anchor_list (list[Tensor]): Anchors of each image.
+ - valid_flag_list (list[Tensor]): Valid flags of each image.
+ """
+ num_imgs = len(batch_img_metas)
+
+ # since feature map sizes of all images are the same, we only compute
+ # anchors for one time
+ multi_level_anchors = self.atss_prior_generator.grid_priors(
+ featmap_sizes, device=device)
+ anchor_list = [multi_level_anchors for _ in range(num_imgs)]
+
+ # for each image, we compute valid flags of multi level anchors
+ valid_flag_list = []
+ for img_id, img_meta in enumerate(batch_img_metas):
+ multi_level_flags = self.atss_prior_generator.valid_flags(
+ featmap_sizes, img_meta['pad_shape'], device=device)
+ valid_flag_list.append(multi_level_flags)
+
+ return anchor_list, valid_flag_list
+
+ def get_atss_targets(
+ self,
+ cls_scores: List[Tensor],
+ mlvl_points: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> tuple:
+ """A wrapper for computing ATSS targets for points in multiple images.
+
+ Args:
+ cls_scores (list[Tensor]): Box iou-aware scores for each scale
+ level with shape (N, num_points * num_classes, H, W).
+ mlvl_points (list[Tensor]): Points of each fpn level, each has
+ shape (num_points, 2).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ tuple:
+
+ - labels_list (list[Tensor]): Labels of each level.
+ - label_weights (Tensor): Label weights of all levels.
+ - bbox_targets_list (list[Tensor]): Regression targets of each
+ level, (l, t, r, b).
+ - bbox_weights (Tensor): Bbox weights of all levels.
+ """
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ assert len(
+ featmap_sizes
+ ) == self.atss_prior_generator.num_levels == \
+ self.fcos_prior_generator.num_levels
+
+ device = cls_scores[0].device
+
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+
+ cls_reg_targets = ATSSHead.get_targets(
+ self,
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore,
+ unmap_outputs=True)
+
+ (anchor_list, labels_list, label_weights_list, bbox_targets_list,
+ bbox_weights_list, avg_factor) = cls_reg_targets
+
+ bbox_targets_list = [
+ bbox_targets.reshape(-1, 4) for bbox_targets in bbox_targets_list
+ ]
+
+ num_imgs = len(batch_img_metas)
+ # transform bbox_targets (x1, y1, x2, y2) into (l, t, r, b) format
+ bbox_targets_list = self.transform_bbox_targets(
+ bbox_targets_list, mlvl_points, num_imgs)
+
+ labels_list = [labels.reshape(-1) for labels in labels_list]
+ label_weights_list = [
+ label_weights.reshape(-1) for label_weights in label_weights_list
+ ]
+ bbox_weights_list = [
+ bbox_weights.reshape(-1) for bbox_weights in bbox_weights_list
+ ]
+ label_weights = torch.cat(label_weights_list)
+ bbox_weights = torch.cat(bbox_weights_list)
+ return labels_list, label_weights, bbox_targets_list, bbox_weights
+
+ def transform_bbox_targets(self, decoded_bboxes: List[Tensor],
+ mlvl_points: List[Tensor],
+ num_imgs: int) -> List[Tensor]:
+ """Transform bbox_targets (x1, y1, x2, y2) into (l, t, r, b) format.
+
+ Args:
+ decoded_bboxes (list[Tensor]): Regression targets of each level,
+ in the form of (x1, y1, x2, y2).
+ mlvl_points (list[Tensor]): Points of each fpn level, each has
+ shape (num_points, 2).
+ num_imgs (int): the number of images in a batch.
+
+ Returns:
+ bbox_targets (list[Tensor]): Regression targets of each level in
+ the form of (l, t, r, b).
+ """
+ # TODO: Re-implemented in Class PointCoder
+ assert len(decoded_bboxes) == len(mlvl_points)
+ num_levels = len(decoded_bboxes)
+ mlvl_points = [points.repeat(num_imgs, 1) for points in mlvl_points]
+ bbox_targets = []
+ for i in range(num_levels):
+ bbox_target = self.bbox_coder.encode(mlvl_points[i],
+ decoded_bboxes[i])
+ bbox_targets.append(bbox_target)
+
+ return bbox_targets
+
+ def _load_from_state_dict(self, state_dict: dict, prefix: str,
+ local_metadata: dict, strict: bool,
+ missing_keys: Union[List[str], str],
+ unexpected_keys: Union[List[str], str],
+ error_msgs: Union[List[str], str]) -> None:
+ """Override the method in the parent class to avoid changing para's
+ name."""
+ pass
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/yolact_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/yolact_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..3390c136a31bee81134667eb28ad8829ddb84cc3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/yolact_head.py
@@ -0,0 +1,1193 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from typing import List, Optional
+
+import numpy as np
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule
+from mmengine.model import BaseModule, ModuleList
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import (ConfigType, InstanceList, OptConfigType,
+ OptInstanceList, OptMultiConfig)
+from ..layers import fast_nms
+from ..utils import images_to_levels, multi_apply, select_single_mlvl
+from ..utils.misc import empty_instances
+from .anchor_head import AnchorHead
+from .base_mask_head import BaseMaskHead
+
+
+@MODELS.register_module()
+class YOLACTHead(AnchorHead):
+ """YOLACT box head used in https://arxiv.org/abs/1904.02689.
+
+ Note that YOLACT head is a light version of RetinaNet head.
+ Four differences are described as follows:
+
+ 1. YOLACT box head has three-times fewer anchors.
+ 2. YOLACT box head shares the convs for box and cls branches.
+ 3. YOLACT box head uses OHEM instead of Focal loss.
+ 4. YOLACT box head predicts a set of mask coefficients for each box.
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ anchor_generator (:obj:`ConfigDict` or dict): Config dict for
+ anchor generator
+ loss_cls (:obj:`ConfigDict` or dict): Config of classification loss.
+ loss_bbox (:obj:`ConfigDict` or dict): Config of localization loss.
+ num_head_convs (int): Number of the conv layers shared by
+ box and cls branches.
+ num_protos (int): Number of the mask coefficients.
+ use_ohem (bool): If true, ``loss_single_OHEM`` will be used for
+ cls loss calculation. If false, ``loss_single`` will be used.
+ conv_cfg (:obj:`ConfigDict` or dict, optional): Dictionary to
+ construct and config conv layer.
+ norm_cfg (:obj:`ConfigDict` or dict, optional): Dictionary to
+ construct and config norm layer.
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or
+ list[dict], optional): Initialization config dict.
+ """
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: int,
+ anchor_generator: ConfigType = dict(
+ type='AnchorGenerator',
+ octave_base_scale=3,
+ scales_per_octave=1,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[8, 16, 32, 64, 128]),
+ loss_cls: ConfigType = dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ reduction='none',
+ loss_weight=1.0),
+ loss_bbox: ConfigType = dict(
+ type='SmoothL1Loss', beta=1.0, loss_weight=1.5),
+ num_head_convs: int = 1,
+ num_protos: int = 32,
+ use_ohem: bool = True,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: OptConfigType = None,
+ init_cfg: OptMultiConfig = dict(
+ type='Xavier',
+ distribution='uniform',
+ bias=0,
+ layer='Conv2d'),
+ **kwargs) -> None:
+ self.num_head_convs = num_head_convs
+ self.num_protos = num_protos
+ self.use_ohem = use_ohem
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ super().__init__(
+ num_classes=num_classes,
+ in_channels=in_channels,
+ loss_cls=loss_cls,
+ loss_bbox=loss_bbox,
+ anchor_generator=anchor_generator,
+ init_cfg=init_cfg,
+ **kwargs)
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self.relu = nn.ReLU(inplace=True)
+ self.head_convs = ModuleList()
+ for i in range(self.num_head_convs):
+ chn = self.in_channels if i == 0 else self.feat_channels
+ self.head_convs.append(
+ ConvModule(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ self.conv_cls = nn.Conv2d(
+ self.feat_channels,
+ self.num_base_priors * self.cls_out_channels,
+ 3,
+ padding=1)
+ self.conv_reg = nn.Conv2d(
+ self.feat_channels, self.num_base_priors * 4, 3, padding=1)
+ self.conv_coeff = nn.Conv2d(
+ self.feat_channels,
+ self.num_base_priors * self.num_protos,
+ 3,
+ padding=1)
+
+ def forward_single(self, x: Tensor) -> tuple:
+ """Forward feature of a single scale level.
+
+ Args:
+ x (Tensor): Features of a single scale level.
+
+ Returns:
+ tuple:
+
+ - cls_score (Tensor): Cls scores for a single scale level
+ the channels number is num_anchors * num_classes.
+ - bbox_pred (Tensor): Box energies / deltas for a single scale
+ level, the channels number is num_anchors * 4.
+ - coeff_pred (Tensor): Mask coefficients for a single scale
+ level, the channels number is num_anchors * num_protos.
+ """
+ for head_conv in self.head_convs:
+ x = head_conv(x)
+ cls_score = self.conv_cls(x)
+ bbox_pred = self.conv_reg(x)
+ coeff_pred = self.conv_coeff(x).tanh()
+ return cls_score, bbox_pred, coeff_pred
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ coeff_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Calculate the loss based on the features extracted by the bbox head.
+
+ When ``self.use_ohem == True``, it functions like ``SSDHead.loss``,
+ otherwise, it follows ``AnchorHead.loss``.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ has shape (N, num_anchors * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W).
+ coeff_preds (list[Tensor]): Mask coefficients for each scale
+ level with shape (N, num_anchors * num_protos, H, W)
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ assert len(featmap_sizes) == self.prior_generator.num_levels
+
+ device = cls_scores[0].device
+
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+ cls_reg_targets = self.get_targets(
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore,
+ unmap_outputs=not self.use_ohem,
+ return_sampling_results=True)
+ (labels_list, label_weights_list, bbox_targets_list, bbox_weights_list,
+ avg_factor, sampling_results) = cls_reg_targets
+
+ if self.use_ohem:
+ num_images = len(batch_img_metas)
+ all_cls_scores = torch.cat([
+ s.permute(0, 2, 3, 1).reshape(
+ num_images, -1, self.cls_out_channels) for s in cls_scores
+ ], 1)
+ all_labels = torch.cat(labels_list, -1).view(num_images, -1)
+ all_label_weights = torch.cat(label_weights_list,
+ -1).view(num_images, -1)
+ all_bbox_preds = torch.cat([
+ b.permute(0, 2, 3, 1).reshape(num_images, -1, 4)
+ for b in bbox_preds
+ ], -2)
+ all_bbox_targets = torch.cat(bbox_targets_list,
+ -2).view(num_images, -1, 4)
+ all_bbox_weights = torch.cat(bbox_weights_list,
+ -2).view(num_images, -1, 4)
+
+ # concat all level anchors to a single tensor
+ all_anchors = []
+ for i in range(num_images):
+ all_anchors.append(torch.cat(anchor_list[i]))
+
+ # check NaN and Inf
+ assert torch.isfinite(all_cls_scores).all().item(), \
+ 'classification scores become infinite or NaN!'
+ assert torch.isfinite(all_bbox_preds).all().item(), \
+ 'bbox predications become infinite or NaN!'
+
+ losses_cls, losses_bbox = multi_apply(
+ self.OHEMloss_by_feat_single,
+ all_cls_scores,
+ all_bbox_preds,
+ all_anchors,
+ all_labels,
+ all_label_weights,
+ all_bbox_targets,
+ all_bbox_weights,
+ avg_factor=avg_factor)
+ else:
+ # anchor number of multi levels
+ num_level_anchors = [anchors.size(0) for anchors in anchor_list[0]]
+ # concat all level anchors and flags to a single tensor
+ concat_anchor_list = []
+ for i in range(len(anchor_list)):
+ concat_anchor_list.append(torch.cat(anchor_list[i]))
+ all_anchor_list = images_to_levels(concat_anchor_list,
+ num_level_anchors)
+ losses_cls, losses_bbox = multi_apply(
+ self.loss_by_feat_single,
+ cls_scores,
+ bbox_preds,
+ all_anchor_list,
+ labels_list,
+ label_weights_list,
+ bbox_targets_list,
+ bbox_weights_list,
+ avg_factor=avg_factor)
+ losses = dict(loss_cls=losses_cls, loss_bbox=losses_bbox)
+ # update `_raw_positive_infos`, which will be used when calling
+ # `get_positive_infos`.
+ self._raw_positive_infos.update(coeff_preds=coeff_preds)
+ return losses
+
+ def OHEMloss_by_feat_single(self, cls_score: Tensor, bbox_pred: Tensor,
+ anchors: Tensor, labels: Tensor,
+ label_weights: Tensor, bbox_targets: Tensor,
+ bbox_weights: Tensor,
+ avg_factor: int) -> tuple:
+ """Compute loss of a single image. Similar to
+ func:``SSDHead.loss_by_feat_single``
+
+ Args:
+ cls_score (Tensor): Box scores for eachimage
+ Has shape (num_total_anchors, num_classes).
+ bbox_pred (Tensor): Box energies / deltas for each image
+ level with shape (num_total_anchors, 4).
+ anchors (Tensor): Box reference for each scale level with shape
+ (num_total_anchors, 4).
+ labels (Tensor): Labels of each anchors with shape
+ (num_total_anchors,).
+ label_weights (Tensor): Label weights of each anchor with shape
+ (num_total_anchors,)
+ bbox_targets (Tensor): BBox regression targets of each anchor with
+ shape (num_total_anchors, 4).
+ bbox_weights (Tensor): BBox regression loss weights of each anchor
+ with shape (num_total_anchors, 4).
+ avg_factor (int): Average factor that is used to average
+ the loss. When using sampling method, avg_factor is usually
+ the sum of positive and negative priors. When using
+ `PseudoSampler`, `avg_factor` is usually equal to the number
+ of positive priors.
+
+ Returns:
+ Tuple[Tensor, Tensor]: A tuple of cls loss and bbox loss of one
+ feature map.
+ """
+
+ loss_cls_all = self.loss_cls(cls_score, labels, label_weights)
+
+ # FG cat_id: [0, num_classes -1], BG cat_id: num_classes
+ pos_inds = ((labels >= 0) & (labels < self.num_classes)).nonzero(
+ as_tuple=False).reshape(-1)
+ neg_inds = (labels == self.num_classes).nonzero(
+ as_tuple=False).view(-1)
+
+ num_pos_samples = pos_inds.size(0)
+ if num_pos_samples == 0:
+ num_neg_samples = neg_inds.size(0)
+ else:
+ num_neg_samples = self.train_cfg['neg_pos_ratio'] * \
+ num_pos_samples
+ if num_neg_samples > neg_inds.size(0):
+ num_neg_samples = neg_inds.size(0)
+ topk_loss_cls_neg, _ = loss_cls_all[neg_inds].topk(num_neg_samples)
+ loss_cls_pos = loss_cls_all[pos_inds].sum()
+ loss_cls_neg = topk_loss_cls_neg.sum()
+ loss_cls = (loss_cls_pos + loss_cls_neg) / avg_factor
+ if self.reg_decoded_bbox:
+ # When the regression loss (e.g. `IouLoss`, `GIouLoss`)
+ # is applied directly on the decoded bounding boxes, it
+ # decodes the already encoded coordinates to absolute format.
+ bbox_pred = self.bbox_coder.decode(anchors, bbox_pred)
+ loss_bbox = self.loss_bbox(
+ bbox_pred, bbox_targets, bbox_weights, avg_factor=avg_factor)
+ return loss_cls[None], loss_bbox
+
+ def get_positive_infos(self) -> InstanceList:
+ """Get positive information from sampling results.
+
+ Returns:
+ list[:obj:`InstanceData`]: Positive Information of each image,
+ usually including positive bboxes, positive labels, positive
+ priors, positive coeffs, etc.
+ """
+ assert len(self._raw_positive_infos) > 0
+ sampling_results = self._raw_positive_infos['sampling_results']
+ num_imgs = len(sampling_results)
+
+ coeff_pred_list = []
+ for coeff_pred_per_level in self._raw_positive_infos['coeff_preds']:
+ coeff_pred_per_level = \
+ coeff_pred_per_level.permute(
+ 0, 2, 3, 1).reshape(num_imgs, -1, self.num_protos)
+ coeff_pred_list.append(coeff_pred_per_level)
+ coeff_preds = torch.cat(coeff_pred_list, dim=1)
+
+ pos_info_list = []
+ for idx, sampling_result in enumerate(sampling_results):
+ pos_info = InstanceData()
+ coeff_preds_single = coeff_preds[idx]
+ pos_info.pos_assigned_gt_inds = \
+ sampling_result.pos_assigned_gt_inds
+ pos_info.pos_inds = sampling_result.pos_inds
+ pos_info.coeffs = coeff_preds_single[sampling_result.pos_inds]
+ pos_info.bboxes = sampling_result.pos_gt_bboxes
+ pos_info_list.append(pos_info)
+ return pos_info_list
+
+ def predict_by_feat(self,
+ cls_scores,
+ bbox_preds,
+ coeff_preds,
+ batch_img_metas,
+ cfg=None,
+ rescale=True,
+ **kwargs):
+ """Similar to func:``AnchorHead.get_bboxes``, but additionally
+ processes coeff_preds.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ with shape (N, num_anchors * num_classes, H, W)
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W)
+ coeff_preds (list[Tensor]): Mask coefficients for each scale
+ level with shape (N, num_anchors * num_protos, H, W)
+ batch_img_metas (list[dict]): Batch image meta info.
+ cfg (:obj:`Config` | None): Test / postprocessing configuration,
+ if None, test_cfg would be used
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to True.
+
+ Returns:
+ list[:obj:`InstanceData`]: Object detection results of each image
+ after the post process. Each item usually contains following keys.
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - coeffs (Tensor): the predicted mask coefficients of
+ instance inside the corresponding box has a shape
+ (n, num_protos).
+ """
+ assert len(cls_scores) == len(bbox_preds)
+ num_levels = len(cls_scores)
+
+ device = cls_scores[0].device
+ featmap_sizes = [cls_scores[i].shape[-2:] for i in range(num_levels)]
+ mlvl_priors = self.prior_generator.grid_priors(
+ featmap_sizes, device=device)
+
+ result_list = []
+ for img_id in range(len(batch_img_metas)):
+ img_meta = batch_img_metas[img_id]
+ cls_score_list = select_single_mlvl(cls_scores, img_id)
+ bbox_pred_list = select_single_mlvl(bbox_preds, img_id)
+ coeff_pred_list = select_single_mlvl(coeff_preds, img_id)
+ results = self._predict_by_feat_single(
+ cls_score_list=cls_score_list,
+ bbox_pred_list=bbox_pred_list,
+ coeff_preds_list=coeff_pred_list,
+ mlvl_priors=mlvl_priors,
+ img_meta=img_meta,
+ cfg=cfg,
+ rescale=rescale)
+ result_list.append(results)
+ return result_list
+
+ def _predict_by_feat_single(self,
+ cls_score_list: List[Tensor],
+ bbox_pred_list: List[Tensor],
+ coeff_preds_list: List[Tensor],
+ mlvl_priors: List[Tensor],
+ img_meta: dict,
+ cfg: ConfigType,
+ rescale: bool = True) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results. Similar to func:``AnchorHead._predict_by_feat_single``,
+ but additionally processes coeff_preds_list and uses fast NMS instead
+ of traditional NMS.
+
+ Args:
+ cls_score_list (list[Tensor]): Box scores for a single scale level
+ Has shape (num_priors * num_classes, H, W).
+ bbox_pred_list (list[Tensor]): Box energies / deltas for a single
+ scale level with shape (num_priors * 4, H, W).
+ coeff_preds_list (list[Tensor]): Mask coefficients for a single
+ scale level with shape (num_priors * num_protos, H, W).
+ mlvl_priors (list[Tensor]): Each element in the list is
+ the priors of a single level in feature pyramid,
+ has shape (num_priors, 4).
+ img_meta (dict): Image meta info.
+ cfg (mmengine.Config): Test / postprocessing configuration,
+ if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - coeffs (Tensor): the predicted mask coefficients of
+ instance inside the corresponding box has a shape
+ (n, num_protos).
+ """
+ assert len(cls_score_list) == len(bbox_pred_list) == len(mlvl_priors)
+
+ cfg = self.test_cfg if cfg is None else cfg
+ cfg = copy.deepcopy(cfg)
+ img_shape = img_meta['img_shape']
+ nms_pre = cfg.get('nms_pre', -1)
+
+ mlvl_bbox_preds = []
+ mlvl_valid_priors = []
+ mlvl_scores = []
+ mlvl_coeffs = []
+ for cls_score, bbox_pred, coeff_pred, priors in \
+ zip(cls_score_list, bbox_pred_list,
+ coeff_preds_list, mlvl_priors):
+ assert cls_score.size()[-2:] == bbox_pred.size()[-2:]
+ cls_score = cls_score.permute(1, 2,
+ 0).reshape(-1, self.cls_out_channels)
+ if self.use_sigmoid_cls:
+ scores = cls_score.sigmoid()
+ else:
+ scores = cls_score.softmax(-1)
+ bbox_pred = bbox_pred.permute(1, 2, 0).reshape(-1, 4)
+ coeff_pred = coeff_pred.permute(1, 2,
+ 0).reshape(-1, self.num_protos)
+
+ if 0 < nms_pre < scores.shape[0]:
+ # Get maximum scores for foreground classes.
+ if self.use_sigmoid_cls:
+ max_scores, _ = scores.max(dim=1)
+ else:
+ # remind that we set FG labels to [0, num_class-1]
+ # since mmdet v2.0
+ # BG cat_id: num_class
+ max_scores, _ = scores[:, :-1].max(dim=1)
+ _, topk_inds = max_scores.topk(nms_pre)
+ priors = priors[topk_inds, :]
+ bbox_pred = bbox_pred[topk_inds, :]
+ scores = scores[topk_inds, :]
+ coeff_pred = coeff_pred[topk_inds, :]
+
+ mlvl_bbox_preds.append(bbox_pred)
+ mlvl_valid_priors.append(priors)
+ mlvl_scores.append(scores)
+ mlvl_coeffs.append(coeff_pred)
+
+ bbox_pred = torch.cat(mlvl_bbox_preds)
+ priors = torch.cat(mlvl_valid_priors)
+ multi_bboxes = self.bbox_coder.decode(
+ priors, bbox_pred, max_shape=img_shape)
+
+ multi_scores = torch.cat(mlvl_scores)
+ multi_coeffs = torch.cat(mlvl_coeffs)
+
+ return self._bbox_post_process(
+ multi_bboxes=multi_bboxes,
+ multi_scores=multi_scores,
+ multi_coeffs=multi_coeffs,
+ cfg=cfg,
+ rescale=rescale,
+ img_meta=img_meta)
+
+ def _bbox_post_process(self,
+ multi_bboxes: Tensor,
+ multi_scores: Tensor,
+ multi_coeffs: Tensor,
+ cfg: ConfigType,
+ rescale: bool = False,
+ img_meta: Optional[dict] = None,
+ **kwargs) -> InstanceData:
+ """bbox post-processing method.
+
+ The boxes would be rescaled to the original image scale and do
+ the nms operation. Usually `with_nms` is False is used for aug test.
+
+ Args:
+ multi_bboxes (Tensor): Predicted bbox that concat all levels.
+ multi_scores (Tensor): Bbox scores that concat all levels.
+ multi_coeffs (Tensor): Mask coefficients that concat all levels.
+ cfg (ConfigDict): Test / postprocessing configuration,
+ if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Default to False.
+ img_meta (dict, optional): Image meta info. Defaults to None.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - coeffs (Tensor): the predicted mask coefficients of
+ instance inside the corresponding box has a shape
+ (n, num_protos).
+ """
+ if rescale:
+ assert img_meta.get('scale_factor') is not None
+ multi_bboxes /= multi_bboxes.new_tensor(
+ img_meta['scale_factor']).repeat((1, 2))
+ # mlvl_bboxes /= mlvl_bboxes.new_tensor(scale_factor)
+
+ if self.use_sigmoid_cls:
+ # Add a dummy background class to the backend when using sigmoid
+ # remind that we set FG labels to [0, num_class-1] since mmdet v2.0
+ # BG cat_id: num_class
+
+ padding = multi_scores.new_zeros(multi_scores.shape[0], 1)
+ multi_scores = torch.cat([multi_scores, padding], dim=1)
+ det_bboxes, det_labels, det_coeffs = fast_nms(
+ multi_bboxes, multi_scores, multi_coeffs, cfg.score_thr,
+ cfg.iou_thr, cfg.top_k, cfg.max_per_img)
+ results = InstanceData()
+ results.bboxes = det_bboxes[:, :4]
+ results.scores = det_bboxes[:, -1]
+ results.labels = det_labels
+ results.coeffs = det_coeffs
+ return results
+
+
+@MODELS.register_module()
+class YOLACTProtonet(BaseMaskHead):
+ """YOLACT mask head used in https://arxiv.org/abs/1904.02689.
+
+ This head outputs the mask prototypes for YOLACT.
+
+ Args:
+ in_channels (int): Number of channels in the input feature map.
+ proto_channels (tuple[int]): Output channels of protonet convs.
+ proto_kernel_sizes (tuple[int]): Kernel sizes of protonet convs.
+ include_last_relu (bool): If keep the last relu of protonet.
+ num_protos (int): Number of prototypes.
+ num_classes (int): Number of categories excluding the background
+ category.
+ loss_mask_weight (float): Reweight the mask loss by this factor.
+ max_masks_to_train (int): Maximum number of masks to train for
+ each image.
+ with_seg_branch (bool): Whether to apply a semantic segmentation
+ branch and calculate loss during training to increase
+ performance with no speed penalty. Defaults to True.
+ loss_segm (:obj:`ConfigDict` or dict, optional): Config of
+ semantic segmentation loss.
+ train_cfg (:obj:`ConfigDict` or dict, optional): Training config
+ of head.
+ test_cfg (:obj:`ConfigDict` or dict, optional): Testing config of
+ head.
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or
+ list[dict], optional): Initialization config dict.
+ """
+
+ def __init__(
+ self,
+ num_classes: int,
+ in_channels: int = 256,
+ proto_channels: tuple = (256, 256, 256, None, 256, 32),
+ proto_kernel_sizes: tuple = (3, 3, 3, -2, 3, 1),
+ include_last_relu: bool = True,
+ num_protos: int = 32,
+ loss_mask_weight: float = 1.0,
+ max_masks_to_train: int = 100,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ with_seg_branch: bool = True,
+ loss_segm: ConfigType = dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=1.0),
+ init_cfg=dict(
+ type='Xavier',
+ distribution='uniform',
+ override=dict(name='protonet'))
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.in_channels = in_channels
+ self.proto_channels = proto_channels
+ self.proto_kernel_sizes = proto_kernel_sizes
+ self.include_last_relu = include_last_relu
+
+ # Segmentation branch
+ self.with_seg_branch = with_seg_branch
+ self.segm_branch = SegmentationModule(
+ num_classes=num_classes, in_channels=in_channels) \
+ if with_seg_branch else None
+ self.loss_segm = MODELS.build(loss_segm) if with_seg_branch else None
+
+ self.loss_mask_weight = loss_mask_weight
+ self.num_protos = num_protos
+ self.num_classes = num_classes
+ self.max_masks_to_train = max_masks_to_train
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+ self._init_layers()
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ # Possible patterns:
+ # ( 256, 3) -> conv
+ # ( 256,-2) -> deconv
+ # (None,-2) -> bilinear interpolate
+ in_channels = self.in_channels
+ protonets = ModuleList()
+ for num_channels, kernel_size in zip(self.proto_channels,
+ self.proto_kernel_sizes):
+ if kernel_size > 0:
+ layer = nn.Conv2d(
+ in_channels,
+ num_channels,
+ kernel_size,
+ padding=kernel_size // 2)
+ else:
+ if num_channels is None:
+ layer = InterpolateModule(
+ scale_factor=-kernel_size,
+ mode='bilinear',
+ align_corners=False)
+ else:
+ layer = nn.ConvTranspose2d(
+ in_channels,
+ num_channels,
+ -kernel_size,
+ padding=kernel_size // 2)
+ protonets.append(layer)
+ protonets.append(nn.ReLU(inplace=True))
+ in_channels = num_channels if num_channels is not None \
+ else in_channels
+ if not self.include_last_relu:
+ protonets = protonets[:-1]
+ self.protonet = nn.Sequential(*protonets)
+
+ def forward(self, x: tuple, positive_infos: InstanceList) -> tuple:
+ """Forward feature from the upstream network to get prototypes and
+ linearly combine the prototypes, using masks coefficients, into
+ instance masks. Finally, crop the instance masks with given bboxes.
+
+ Args:
+ x (Tuple[Tensor]): Feature from the upstream network, which is
+ a 4D-tensor.
+ positive_infos (List[:obj:``InstanceData``]): Positive information
+ that calculate from detect head.
+
+ Returns:
+ tuple: Predicted instance segmentation masks and
+ semantic segmentation map.
+ """
+ # YOLACT used single feature map to get segmentation masks
+ single_x = x[0]
+
+ # YOLACT segmentation branch, if not training or segmentation branch
+ # is None, will not process the forward function.
+ if self.segm_branch is not None and self.training:
+ segm_preds = self.segm_branch(single_x)
+ else:
+ segm_preds = None
+ # YOLACT mask head
+ prototypes = self.protonet(single_x)
+ prototypes = prototypes.permute(0, 2, 3, 1).contiguous()
+
+ num_imgs = single_x.size(0)
+
+ mask_pred_list = []
+ for idx in range(num_imgs):
+ cur_prototypes = prototypes[idx]
+ pos_coeffs = positive_infos[idx].coeffs
+
+ # Linearly combine the prototypes with the mask coefficients
+ mask_preds = cur_prototypes @ pos_coeffs.t()
+ mask_preds = torch.sigmoid(mask_preds)
+ mask_pred_list.append(mask_preds)
+ return mask_pred_list, segm_preds
+
+ def loss_by_feat(self, mask_preds: List[Tensor], segm_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict], positive_infos: InstanceList,
+ **kwargs) -> dict:
+ """Calculate the loss based on the features extracted by the mask head.
+
+ Args:
+ mask_preds (list[Tensor]): List of predicted prototypes, each has
+ shape (num_classes, H, W).
+ segm_preds (Tensor): Predicted semantic segmentation map with
+ shape (N, num_classes, H, W)
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``masks``,
+ and ``labels`` attributes.
+ batch_img_metas (list[dict]): Meta information of multiple images.
+ positive_infos (List[:obj:``InstanceData``]): Information of
+ positive samples of each image that are assigned in detection
+ head.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ assert positive_infos is not None, \
+ 'positive_infos should not be None in `YOLACTProtonet`'
+ losses = dict()
+
+ # crop
+ croped_mask_pred = self.crop_mask_preds(mask_preds, batch_img_metas,
+ positive_infos)
+
+ loss_mask = []
+ loss_segm = []
+ num_imgs, _, mask_h, mask_w = segm_preds.size()
+ assert num_imgs == len(croped_mask_pred)
+ segm_avg_factor = num_imgs * mask_h * mask_w
+ total_pos = 0
+
+ if self.segm_branch is not None:
+ assert segm_preds is not None
+
+ for idx in range(num_imgs):
+ img_meta = batch_img_metas[idx]
+
+ (mask_preds, pos_mask_targets, segm_targets, num_pos,
+ gt_bboxes_for_reweight) = self._get_targets_single(
+ croped_mask_pred[idx], segm_preds[idx],
+ batch_gt_instances[idx], positive_infos[idx])
+
+ # segmentation loss
+ if self.with_seg_branch:
+ if segm_targets is None:
+ loss = segm_preds[idx].sum() * 0.
+ else:
+ loss = self.loss_segm(
+ segm_preds[idx],
+ segm_targets,
+ avg_factor=segm_avg_factor)
+ loss_segm.append(loss)
+ # mask loss
+ total_pos += num_pos
+ if num_pos == 0 or pos_mask_targets is None:
+ loss = mask_preds.sum() * 0.
+ else:
+ mask_preds = torch.clamp(mask_preds, 0, 1)
+ loss = F.binary_cross_entropy(
+ mask_preds, pos_mask_targets,
+ reduction='none') * self.loss_mask_weight
+
+ h, w = img_meta['img_shape'][:2]
+ gt_bboxes_width = (gt_bboxes_for_reweight[:, 2] -
+ gt_bboxes_for_reweight[:, 0]) / w
+ gt_bboxes_height = (gt_bboxes_for_reweight[:, 3] -
+ gt_bboxes_for_reweight[:, 1]) / h
+ loss = loss.mean(dim=(1,
+ 2)) / gt_bboxes_width / gt_bboxes_height
+ loss = torch.sum(loss)
+ loss_mask.append(loss)
+
+ if total_pos == 0:
+ total_pos += 1 # avoid nan
+ loss_mask = [x / total_pos for x in loss_mask]
+
+ losses.update(loss_mask=loss_mask)
+ if self.with_seg_branch:
+ losses.update(loss_segm=loss_segm)
+
+ return losses
+
+ def _get_targets_single(self, mask_preds: Tensor, segm_pred: Tensor,
+ gt_instances: InstanceData,
+ positive_info: InstanceData):
+ """Compute targets for predictions of single image.
+
+ Args:
+ mask_preds (Tensor): Predicted prototypes with shape
+ (num_classes, H, W).
+ segm_pred (Tensor): Predicted semantic segmentation map
+ with shape (num_classes, H, W).
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes``, ``labels``,
+ and ``masks`` attributes.
+ positive_info (:obj:`InstanceData`): Information of positive
+ samples that are assigned in detection head. It usually
+ contains following keys.
+
+ - pos_assigned_gt_inds (Tensor): Assigner GT indexes of
+ positive proposals, has shape (num_pos, )
+ - pos_inds (Tensor): Positive index of image, has
+ shape (num_pos, ).
+ - coeffs (Tensor): Positive mask coefficients
+ with shape (num_pos, num_protos).
+ - bboxes (Tensor): Positive bboxes with shape
+ (num_pos, 4)
+
+ Returns:
+ tuple: Usually returns a tuple containing learning targets.
+
+ - mask_preds (Tensor): Positive predicted mask with shape
+ (num_pos, mask_h, mask_w).
+ - pos_mask_targets (Tensor): Positive mask targets with shape
+ (num_pos, mask_h, mask_w).
+ - segm_targets (Tensor): Semantic segmentation targets with shape
+ (num_classes, segm_h, segm_w).
+ - num_pos (int): Positive numbers.
+ - gt_bboxes_for_reweight (Tensor): GT bboxes that match to the
+ positive priors has shape (num_pos, 4).
+ """
+ gt_bboxes = gt_instances.bboxes
+ gt_labels = gt_instances.labels
+ device = gt_bboxes.device
+ gt_masks = gt_instances.masks.to_tensor(
+ dtype=torch.bool, device=device).float()
+ if gt_masks.size(0) == 0:
+ return mask_preds, None, None, 0, None
+
+ # process with semantic segmentation targets
+ if segm_pred is not None:
+ num_classes, segm_h, segm_w = segm_pred.size()
+ with torch.no_grad():
+ downsampled_masks = F.interpolate(
+ gt_masks.unsqueeze(0), (segm_h, segm_w),
+ mode='bilinear',
+ align_corners=False).squeeze(0)
+ downsampled_masks = downsampled_masks.gt(0.5).float()
+ segm_targets = torch.zeros_like(segm_pred, requires_grad=False)
+ for obj_idx in range(downsampled_masks.size(0)):
+ segm_targets[gt_labels[obj_idx] - 1] = torch.max(
+ segm_targets[gt_labels[obj_idx] - 1],
+ downsampled_masks[obj_idx])
+ else:
+ segm_targets = None
+ # process with mask targets
+ pos_assigned_gt_inds = positive_info.pos_assigned_gt_inds
+ num_pos = pos_assigned_gt_inds.size(0)
+ # Since we're producing (near) full image masks,
+ # it'd take too much vram to backprop on every single mask.
+ # Thus we select only a subset.
+ if num_pos > self.max_masks_to_train:
+ perm = torch.randperm(num_pos)
+ select = perm[:self.max_masks_to_train]
+ mask_preds = mask_preds[select]
+ pos_assigned_gt_inds = pos_assigned_gt_inds[select]
+ num_pos = self.max_masks_to_train
+
+ gt_bboxes_for_reweight = gt_bboxes[pos_assigned_gt_inds]
+
+ mask_h, mask_w = mask_preds.shape[-2:]
+ gt_masks = F.interpolate(
+ gt_masks.unsqueeze(0), (mask_h, mask_w),
+ mode='bilinear',
+ align_corners=False).squeeze(0)
+ gt_masks = gt_masks.gt(0.5).float()
+ pos_mask_targets = gt_masks[pos_assigned_gt_inds]
+
+ return (mask_preds, pos_mask_targets, segm_targets, num_pos,
+ gt_bboxes_for_reweight)
+
+ def crop_mask_preds(self, mask_preds: List[Tensor],
+ batch_img_metas: List[dict],
+ positive_infos: InstanceList) -> list:
+ """Crop predicted masks by zeroing out everything not in the predicted
+ bbox.
+
+ Args:
+ mask_preds (list[Tensor]): Predicted prototypes with shape
+ (num_classes, H, W).
+ batch_img_metas (list[dict]): Meta information of multiple images.
+ positive_infos (List[:obj:``InstanceData``]): Positive
+ information that calculate from detect head.
+
+ Returns:
+ list: The cropped masks.
+ """
+ croped_mask_preds = []
+ for img_meta, mask_preds, cur_info in zip(batch_img_metas, mask_preds,
+ positive_infos):
+ bboxes_for_cropping = copy.deepcopy(cur_info.bboxes)
+ h, w = img_meta['img_shape'][:2]
+ bboxes_for_cropping[:, 0::2] /= w
+ bboxes_for_cropping[:, 1::2] /= h
+ mask_preds = self.crop_single(mask_preds, bboxes_for_cropping)
+ mask_preds = mask_preds.permute(2, 0, 1).contiguous()
+ croped_mask_preds.append(mask_preds)
+ return croped_mask_preds
+
+ def crop_single(self,
+ masks: Tensor,
+ boxes: Tensor,
+ padding: int = 1) -> Tensor:
+ """Crop single predicted masks by zeroing out everything not in the
+ predicted bbox.
+
+ Args:
+ masks (Tensor): Predicted prototypes, has shape [H, W, N].
+ boxes (Tensor): Bbox coords in relative point form with
+ shape [N, 4].
+ padding (int): Image padding size.
+
+ Return:
+ Tensor: The cropped masks.
+ """
+ h, w, n = masks.size()
+ x1, x2 = self.sanitize_coordinates(
+ boxes[:, 0], boxes[:, 2], w, padding, cast=False)
+ y1, y2 = self.sanitize_coordinates(
+ boxes[:, 1], boxes[:, 3], h, padding, cast=False)
+
+ rows = torch.arange(
+ w, device=masks.device, dtype=x1.dtype).view(1, -1,
+ 1).expand(h, w, n)
+ cols = torch.arange(
+ h, device=masks.device, dtype=x1.dtype).view(-1, 1,
+ 1).expand(h, w, n)
+
+ masks_left = rows >= x1.view(1, 1, -1)
+ masks_right = rows < x2.view(1, 1, -1)
+ masks_up = cols >= y1.view(1, 1, -1)
+ masks_down = cols < y2.view(1, 1, -1)
+
+ crop_mask = masks_left * masks_right * masks_up * masks_down
+
+ return masks * crop_mask.float()
+
+ def sanitize_coordinates(self,
+ x1: Tensor,
+ x2: Tensor,
+ img_size: int,
+ padding: int = 0,
+ cast: bool = True) -> tuple:
+ """Sanitizes the input coordinates so that x1 < x2, x1 != x2, x1 >= 0,
+ and x2 <= image_size. Also converts from relative to absolute
+ coordinates and casts the results to long tensors.
+
+ Warning: this does things in-place behind the scenes so
+ copy if necessary.
+
+ Args:
+ x1 (Tensor): shape (N, ).
+ x2 (Tensor): shape (N, ).
+ img_size (int): Size of the input image.
+ padding (int): x1 >= padding, x2 <= image_size-padding.
+ cast (bool): If cast is false, the result won't be cast to longs.
+
+ Returns:
+ tuple:
+
+ - x1 (Tensor): Sanitized _x1.
+ - x2 (Tensor): Sanitized _x2.
+ """
+ x1 = x1 * img_size
+ x2 = x2 * img_size
+ if cast:
+ x1 = x1.long()
+ x2 = x2.long()
+ x1 = torch.min(x1, x2)
+ x2 = torch.max(x1, x2)
+ x1 = torch.clamp(x1 - padding, min=0)
+ x2 = torch.clamp(x2 + padding, max=img_size)
+ return x1, x2
+
+ def predict_by_feat(self,
+ mask_preds: List[Tensor],
+ segm_preds: Tensor,
+ results_list: InstanceList,
+ batch_img_metas: List[dict],
+ rescale: bool = True,
+ **kwargs) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ mask results.
+
+ Args:
+ mask_preds (list[Tensor]): Predicted prototypes with shape
+ (num_classes, H, W).
+ results_list (List[:obj:``InstanceData``]): BBoxHead results.
+ batch_img_metas (list[dict]): Meta information of all images.
+ rescale (bool, optional): Whether to rescale the results.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Processed results of multiple
+ images.Each :obj:`InstanceData` usually contains
+ following keys.
+
+ - scores (Tensor): Classification scores, has shape
+ (num_instance,).
+ - labels (Tensor): Has shape (num_instances,).
+ - masks (Tensor): Processed mask results, has
+ shape (num_instances, h, w).
+ """
+ assert len(mask_preds) == len(results_list) == len(batch_img_metas)
+
+ croped_mask_pred = self.crop_mask_preds(mask_preds, batch_img_metas,
+ results_list)
+
+ for img_id in range(len(batch_img_metas)):
+ img_meta = batch_img_metas[img_id]
+ results = results_list[img_id]
+ bboxes = results.bboxes
+ mask_preds = croped_mask_pred[img_id]
+ if bboxes.shape[0] == 0 or mask_preds.shape[0] == 0:
+ results_list[img_id] = empty_instances(
+ [img_meta],
+ bboxes.device,
+ task_type='mask',
+ instance_results=[results])[0]
+ else:
+ im_mask = self._predict_by_feat_single(
+ mask_preds=croped_mask_pred[img_id],
+ bboxes=bboxes,
+ img_meta=img_meta,
+ rescale=rescale)
+ results.masks = im_mask
+ return results_list
+
+ def _predict_by_feat_single(self,
+ mask_preds: Tensor,
+ bboxes: Tensor,
+ img_meta: dict,
+ rescale: bool,
+ cfg: OptConfigType = None):
+ """Transform a single image's features extracted from the head into
+ mask results.
+
+ Args:
+ mask_preds (Tensor): Predicted prototypes, has shape [H, W, N].
+ bboxes (Tensor): Bbox coords in relative point form with
+ shape [N, 4].
+ img_meta (dict): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ rescale (bool): If rescale is False, then returned masks will
+ fit the scale of imgs[0].
+ cfg (dict, optional): Config used in test phase.
+ Defaults to None.
+
+ Returns:
+ :obj:`InstanceData`: Processed results of single image.
+ it usually contains following keys.
+
+ - scores (Tensor): Classification scores, has shape
+ (num_instance,).
+ - labels (Tensor): Has shape (num_instances,).
+ - masks (Tensor): Processed mask results, has
+ shape (num_instances, h, w).
+ """
+ cfg = self.test_cfg if cfg is None else cfg
+ scale_factor = bboxes.new_tensor(img_meta['scale_factor']).repeat(
+ (1, 2))
+ img_h, img_w = img_meta['ori_shape'][:2]
+ if rescale: # in-placed rescale the bboxes
+ scale_factor = bboxes.new_tensor(img_meta['scale_factor']).repeat(
+ (1, 2))
+ bboxes /= scale_factor
+ else:
+ w_scale, h_scale = scale_factor[0, 0], scale_factor[0, 1]
+ img_h = np.round(img_h * h_scale.item()).astype(np.int32)
+ img_w = np.round(img_w * w_scale.item()).astype(np.int32)
+
+ masks = F.interpolate(
+ mask_preds.unsqueeze(0), (img_h, img_w),
+ mode='bilinear',
+ align_corners=False).squeeze(0) > cfg.mask_thr
+
+ if cfg.mask_thr_binary < 0:
+ # for visualization and debugging
+ masks = (masks * 255).to(dtype=torch.uint8)
+
+ return masks
+
+
+class SegmentationModule(BaseModule):
+ """YOLACT segmentation branch used in `_
+
+ In mmdet v2.x `segm_loss` is calculated in YOLACTSegmHead, while in
+ mmdet v3.x `SegmentationModule` is used to obtain the predicted semantic
+ segmentation map and `segm_loss` is calculated in YOLACTProtonet.
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ """
+
+ def __init__(
+ self,
+ num_classes: int,
+ in_channels: int = 256,
+ init_cfg: ConfigType = dict(
+ type='Xavier',
+ distribution='uniform',
+ override=dict(name='segm_conv'))
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.in_channels = in_channels
+ self.num_classes = num_classes
+ self._init_layers()
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self.segm_conv = nn.Conv2d(
+ self.in_channels, self.num_classes, kernel_size=1)
+
+ def forward(self, x: Tensor) -> Tensor:
+ """Forward feature from the upstream network.
+
+ Args:
+ x (Tensor): Feature from the upstream network, which is
+ a 4D-tensor.
+
+ Returns:
+ Tensor: Predicted semantic segmentation map with shape
+ (N, num_classes, H, W).
+ """
+ return self.segm_conv(x)
+
+
+class InterpolateModule(BaseModule):
+ """This is a module version of F.interpolate.
+
+ Any arguments you give it just get passed along for the ride.
+ """
+
+ def __init__(self, *args, init_cfg=None, **kwargs) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.args = args
+ self.kwargs = kwargs
+
+ def forward(self, x: Tensor) -> Tensor:
+ """Forward features from the upstream network.
+
+ Args:
+ x (Tensor): Feature from the upstream network, which is
+ a 4D-tensor.
+
+ Returns:
+ Tensor: A 4D-tensor feature map.
+ """
+ return F.interpolate(x, *self.args, **self.kwargs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/yolo_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/yolo_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..0f63afbbc94353e16e4c67ec5bc0b6cd1200de07
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/yolo_head.py
@@ -0,0 +1,527 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+# Copyright (c) 2019 Western Digital Corporation or its affiliates.
+
+import copy
+import warnings
+from typing import List, Optional, Sequence, Tuple
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule, is_norm
+from mmengine.model import bias_init_with_prob, constant_init, normal_init
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.utils import (ConfigType, InstanceList, OptConfigType,
+ OptInstanceList)
+from ..task_modules.samplers import PseudoSampler
+from ..utils import filter_scores_and_topk, images_to_levels, multi_apply
+from .base_dense_head import BaseDenseHead
+
+
+@MODELS.register_module()
+class YOLOV3Head(BaseDenseHead):
+ """YOLOV3Head Paper link: https://arxiv.org/abs/1804.02767.
+
+ Args:
+ num_classes (int): The number of object classes (w/o background)
+ in_channels (Sequence[int]): Number of input channels per scale.
+ out_channels (Sequence[int]): The number of output channels per scale
+ before the final 1x1 layer. Default: (1024, 512, 256).
+ anchor_generator (:obj:`ConfigDict` or dict): Config dict for anchor
+ generator.
+ bbox_coder (:obj:`ConfigDict` or dict): Config of bounding box coder.
+ featmap_strides (Sequence[int]): The stride of each scale.
+ Should be in descending order. Defaults to (32, 16, 8).
+ one_hot_smoother (float): Set a non-zero value to enable label-smooth
+ Defaults to 0.
+ conv_cfg (:obj:`ConfigDict` or dict, optional): Config dict for
+ convolution layer. Defaults to None.
+ norm_cfg (:obj:`ConfigDict` or dict): Dictionary to construct and
+ config norm layer. Defaults to dict(type='BN', requires_grad=True).
+ act_cfg (:obj:`ConfigDict` or dict): Config dict for activation layer.
+ Defaults to dict(type='LeakyReLU', negative_slope=0.1).
+ loss_cls (:obj:`ConfigDict` or dict): Config of classification loss.
+ loss_conf (:obj:`ConfigDict` or dict): Config of confidence loss.
+ loss_xy (:obj:`ConfigDict` or dict): Config of xy coordinate loss.
+ loss_wh (:obj:`ConfigDict` or dict): Config of wh coordinate loss.
+ train_cfg (:obj:`ConfigDict` or dict, optional): Training config of
+ YOLOV3 head. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): Testing config of
+ YOLOV3 head. Defaults to None.
+ """
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: Sequence[int],
+ out_channels: Sequence[int] = (1024, 512, 256),
+ anchor_generator: ConfigType = dict(
+ type='YOLOAnchorGenerator',
+ base_sizes=[[(116, 90), (156, 198), (373, 326)],
+ [(30, 61), (62, 45), (59, 119)],
+ [(10, 13), (16, 30), (33, 23)]],
+ strides=[32, 16, 8]),
+ bbox_coder: ConfigType = dict(type='YOLOBBoxCoder'),
+ featmap_strides: Sequence[int] = (32, 16, 8),
+ one_hot_smoother: float = 0.,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: ConfigType = dict(type='BN', requires_grad=True),
+ act_cfg: ConfigType = dict(
+ type='LeakyReLU', negative_slope=0.1),
+ loss_cls: ConfigType = dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ loss_weight=1.0),
+ loss_conf: ConfigType = dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ loss_weight=1.0),
+ loss_xy: ConfigType = dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ loss_weight=1.0),
+ loss_wh: ConfigType = dict(type='MSELoss', loss_weight=1.0),
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None) -> None:
+ super().__init__(init_cfg=None)
+ # Check params
+ assert (len(in_channels) == len(out_channels) == len(featmap_strides))
+
+ self.num_classes = num_classes
+ self.in_channels = in_channels
+ self.out_channels = out_channels
+ self.featmap_strides = featmap_strides
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+ if self.train_cfg:
+ self.assigner = TASK_UTILS.build(self.train_cfg['assigner'])
+ if train_cfg.get('sampler', None) is not None:
+ self.sampler = TASK_UTILS.build(
+ self.train_cfg['sampler'], context=self)
+ else:
+ self.sampler = PseudoSampler()
+
+ self.one_hot_smoother = one_hot_smoother
+
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self.act_cfg = act_cfg
+
+ self.bbox_coder = TASK_UTILS.build(bbox_coder)
+
+ self.prior_generator = TASK_UTILS.build(anchor_generator)
+
+ self.loss_cls = MODELS.build(loss_cls)
+ self.loss_conf = MODELS.build(loss_conf)
+ self.loss_xy = MODELS.build(loss_xy)
+ self.loss_wh = MODELS.build(loss_wh)
+
+ self.num_base_priors = self.prior_generator.num_base_priors[0]
+ assert len(
+ self.prior_generator.num_base_priors) == len(featmap_strides)
+ self._init_layers()
+
+ @property
+ def num_levels(self) -> int:
+ """int: number of feature map levels"""
+ return len(self.featmap_strides)
+
+ @property
+ def num_attrib(self) -> int:
+ """int: number of attributes in pred_map, bboxes (4) +
+ objectness (1) + num_classes"""
+
+ return 5 + self.num_classes
+
+ def _init_layers(self) -> None:
+ """initialize conv layers in YOLOv3 head."""
+ self.convs_bridge = nn.ModuleList()
+ self.convs_pred = nn.ModuleList()
+ for i in range(self.num_levels):
+ conv_bridge = ConvModule(
+ self.in_channels[i],
+ self.out_channels[i],
+ 3,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg)
+ conv_pred = nn.Conv2d(self.out_channels[i],
+ self.num_base_priors * self.num_attrib, 1)
+
+ self.convs_bridge.append(conv_bridge)
+ self.convs_pred.append(conv_pred)
+
+ def init_weights(self) -> None:
+ """initialize weights."""
+ for m in self.modules():
+ if isinstance(m, nn.Conv2d):
+ normal_init(m, mean=0, std=0.01)
+ if is_norm(m):
+ constant_init(m, 1)
+
+ # Use prior in model initialization to improve stability
+ for conv_pred, stride in zip(self.convs_pred, self.featmap_strides):
+ bias = conv_pred.bias.reshape(self.num_base_priors, -1)
+ # init objectness with prior of 8 objects per feature map
+ # refer to https://github.com/ultralytics/yolov3
+ nn.init.constant_(bias.data[:, 4],
+ bias_init_with_prob(8 / (608 / stride)**2))
+ nn.init.constant_(bias.data[:, 5:], bias_init_with_prob(0.01))
+
+ def forward(self, x: Tuple[Tensor, ...]) -> tuple:
+ """Forward features from the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple[Tensor]: A tuple of multi-level predication map, each is a
+ 4D-tensor of shape (batch_size, 5+num_classes, height, width).
+ """
+
+ assert len(x) == self.num_levels
+ pred_maps = []
+ for i in range(self.num_levels):
+ feat = x[i]
+ feat = self.convs_bridge[i](feat)
+ pred_map = self.convs_pred[i](feat)
+ pred_maps.append(pred_map)
+
+ return tuple(pred_maps),
+
+ def predict_by_feat(self,
+ pred_maps: Sequence[Tensor],
+ batch_img_metas: Optional[List[dict]],
+ cfg: OptConfigType = None,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ bbox results. It has been accelerated since PR #5991.
+
+ Args:
+ pred_maps (Sequence[Tensor]): Raw predictions for a batch of
+ images.
+ batch_img_metas (list[dict], Optional): Batch image meta info.
+ Defaults to None.
+ cfg (:obj:`ConfigDict` or dict, optional): Test / postprocessing
+ configuration, if None, test_cfg would be used.
+ Defaults to None.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ list[:obj:`InstanceData`]: Object detection results of each image
+ after the post process. Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ assert len(pred_maps) == self.num_levels
+ cfg = self.test_cfg if cfg is None else cfg
+ cfg = copy.deepcopy(cfg)
+
+ num_imgs = len(batch_img_metas)
+ featmap_sizes = [pred_map.shape[-2:] for pred_map in pred_maps]
+
+ mlvl_anchors = self.prior_generator.grid_priors(
+ featmap_sizes, device=pred_maps[0].device)
+ flatten_preds = []
+ flatten_strides = []
+ for pred, stride in zip(pred_maps, self.featmap_strides):
+ pred = pred.permute(0, 2, 3, 1).reshape(num_imgs, -1,
+ self.num_attrib)
+ pred[..., :2].sigmoid_()
+ flatten_preds.append(pred)
+ flatten_strides.append(
+ pred.new_tensor(stride).expand(pred.size(1)))
+
+ flatten_preds = torch.cat(flatten_preds, dim=1)
+ flatten_bbox_preds = flatten_preds[..., :4]
+ flatten_objectness = flatten_preds[..., 4].sigmoid()
+ flatten_cls_scores = flatten_preds[..., 5:].sigmoid()
+ flatten_anchors = torch.cat(mlvl_anchors)
+ flatten_strides = torch.cat(flatten_strides)
+ flatten_bboxes = self.bbox_coder.decode(flatten_anchors,
+ flatten_bbox_preds,
+ flatten_strides.unsqueeze(-1))
+ results_list = []
+ for (bboxes, scores, objectness,
+ img_meta) in zip(flatten_bboxes, flatten_cls_scores,
+ flatten_objectness, batch_img_metas):
+ # Filtering out all predictions with conf < conf_thr
+ conf_thr = cfg.get('conf_thr', -1)
+ if conf_thr > 0:
+ conf_inds = objectness >= conf_thr
+ bboxes = bboxes[conf_inds, :]
+ scores = scores[conf_inds, :]
+ objectness = objectness[conf_inds]
+
+ score_thr = cfg.get('score_thr', 0)
+ nms_pre = cfg.get('nms_pre', -1)
+ scores, labels, keep_idxs, _ = filter_scores_and_topk(
+ scores, score_thr, nms_pre)
+
+ results = InstanceData(
+ scores=scores,
+ labels=labels,
+ bboxes=bboxes[keep_idxs],
+ score_factors=objectness[keep_idxs],
+ )
+ results = self._bbox_post_process(
+ results=results,
+ cfg=cfg,
+ rescale=rescale,
+ with_nms=with_nms,
+ img_meta=img_meta)
+ results_list.append(results)
+ return results_list
+
+ def loss_by_feat(
+ self,
+ pred_maps: Sequence[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ pred_maps (list[Tensor]): Prediction map for each scale level,
+ shape (N, num_anchors * num_attrib, H, W)
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ num_imgs = len(batch_img_metas)
+ device = pred_maps[0][0].device
+
+ featmap_sizes = [
+ pred_maps[i].shape[-2:] for i in range(self.num_levels)
+ ]
+ mlvl_anchors = self.prior_generator.grid_priors(
+ featmap_sizes, device=device)
+ anchor_list = [mlvl_anchors for _ in range(num_imgs)]
+
+ responsible_flag_list = []
+ for img_id in range(num_imgs):
+ responsible_flag_list.append(
+ self.responsible_flags(featmap_sizes,
+ batch_gt_instances[img_id].bboxes,
+ device))
+
+ target_maps_list, neg_maps_list = self.get_targets(
+ anchor_list, responsible_flag_list, batch_gt_instances)
+
+ losses_cls, losses_conf, losses_xy, losses_wh = multi_apply(
+ self.loss_by_feat_single, pred_maps, target_maps_list,
+ neg_maps_list)
+
+ return dict(
+ loss_cls=losses_cls,
+ loss_conf=losses_conf,
+ loss_xy=losses_xy,
+ loss_wh=losses_wh)
+
+ def loss_by_feat_single(self, pred_map: Tensor, target_map: Tensor,
+ neg_map: Tensor) -> tuple:
+ """Calculate the loss of a single scale level based on the features
+ extracted by the detection head.
+
+ Args:
+ pred_map (Tensor): Raw predictions for a single level.
+ target_map (Tensor): The Ground-Truth target for a single level.
+ neg_map (Tensor): The negative masks for a single level.
+
+ Returns:
+ tuple:
+ loss_cls (Tensor): Classification loss.
+ loss_conf (Tensor): Confidence loss.
+ loss_xy (Tensor): Regression loss of x, y coordinate.
+ loss_wh (Tensor): Regression loss of w, h coordinate.
+ """
+
+ num_imgs = len(pred_map)
+ pred_map = pred_map.permute(0, 2, 3,
+ 1).reshape(num_imgs, -1, self.num_attrib)
+ neg_mask = neg_map.float()
+ pos_mask = target_map[..., 4]
+ pos_and_neg_mask = neg_mask + pos_mask
+ pos_mask = pos_mask.unsqueeze(dim=-1)
+ if torch.max(pos_and_neg_mask) > 1.:
+ warnings.warn('There is overlap between pos and neg sample.')
+ pos_and_neg_mask = pos_and_neg_mask.clamp(min=0., max=1.)
+
+ pred_xy = pred_map[..., :2]
+ pred_wh = pred_map[..., 2:4]
+ pred_conf = pred_map[..., 4]
+ pred_label = pred_map[..., 5:]
+
+ target_xy = target_map[..., :2]
+ target_wh = target_map[..., 2:4]
+ target_conf = target_map[..., 4]
+ target_label = target_map[..., 5:]
+
+ loss_cls = self.loss_cls(pred_label, target_label, weight=pos_mask)
+ loss_conf = self.loss_conf(
+ pred_conf, target_conf, weight=pos_and_neg_mask)
+ loss_xy = self.loss_xy(pred_xy, target_xy, weight=pos_mask)
+ loss_wh = self.loss_wh(pred_wh, target_wh, weight=pos_mask)
+
+ return loss_cls, loss_conf, loss_xy, loss_wh
+
+ def get_targets(self, anchor_list: List[List[Tensor]],
+ responsible_flag_list: List[List[Tensor]],
+ batch_gt_instances: List[InstanceData]) -> tuple:
+ """Compute target maps for anchors in multiple images.
+
+ Args:
+ anchor_list (list[list[Tensor]]): Multi level anchors of each
+ image. The outer list indicates images, and the inner list
+ corresponds to feature levels of the image. Each element of
+ the inner list is a tensor of shape (num_total_anchors, 4).
+ responsible_flag_list (list[list[Tensor]]): Multi level responsible
+ flags of each image. Each element is a tensor of shape
+ (num_total_anchors, )
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+
+ Returns:
+ tuple: Usually returns a tuple containing learning targets.
+ - target_map_list (list[Tensor]): Target map of each level.
+ - neg_map_list (list[Tensor]): Negative map of each level.
+ """
+ num_imgs = len(anchor_list)
+
+ # anchor number of multi levels
+ num_level_anchors = [anchors.size(0) for anchors in anchor_list[0]]
+
+ results = multi_apply(self._get_targets_single, anchor_list,
+ responsible_flag_list, batch_gt_instances)
+
+ all_target_maps, all_neg_maps = results
+ assert num_imgs == len(all_target_maps) == len(all_neg_maps)
+ target_maps_list = images_to_levels(all_target_maps, num_level_anchors)
+ neg_maps_list = images_to_levels(all_neg_maps, num_level_anchors)
+
+ return target_maps_list, neg_maps_list
+
+ def _get_targets_single(self, anchors: List[Tensor],
+ responsible_flags: List[Tensor],
+ gt_instances: InstanceData) -> tuple:
+ """Generate matching bounding box prior and converted GT.
+
+ Args:
+ anchors (List[Tensor]): Multi-level anchors of the image.
+ responsible_flags (List[Tensor]): Multi-level responsible flags of
+ anchors
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes`` and ``labels``
+ attributes.
+
+ Returns:
+ tuple:
+ target_map (Tensor): Predication target map of each
+ scale level, shape (num_total_anchors,
+ 5+num_classes)
+ neg_map (Tensor): Negative map of each scale level,
+ shape (num_total_anchors,)
+ """
+ gt_bboxes = gt_instances.bboxes
+ gt_labels = gt_instances.labels
+ anchor_strides = []
+ for i in range(len(anchors)):
+ anchor_strides.append(
+ torch.tensor(self.featmap_strides[i],
+ device=gt_bboxes.device).repeat(len(anchors[i])))
+ concat_anchors = torch.cat(anchors)
+ concat_responsible_flags = torch.cat(responsible_flags)
+
+ anchor_strides = torch.cat(anchor_strides)
+ assert len(anchor_strides) == len(concat_anchors) == \
+ len(concat_responsible_flags)
+ pred_instances = InstanceData(
+ priors=concat_anchors, responsible_flags=concat_responsible_flags)
+
+ assign_result = self.assigner.assign(pred_instances, gt_instances)
+ sampling_result = self.sampler.sample(assign_result, pred_instances,
+ gt_instances)
+
+ target_map = concat_anchors.new_zeros(
+ concat_anchors.size(0), self.num_attrib)
+
+ target_map[sampling_result.pos_inds, :4] = self.bbox_coder.encode(
+ sampling_result.pos_priors, sampling_result.pos_gt_bboxes,
+ anchor_strides[sampling_result.pos_inds])
+
+ target_map[sampling_result.pos_inds, 4] = 1
+
+ gt_labels_one_hot = F.one_hot(
+ gt_labels, num_classes=self.num_classes).float()
+ if self.one_hot_smoother != 0: # label smooth
+ gt_labels_one_hot = gt_labels_one_hot * (
+ 1 - self.one_hot_smoother
+ ) + self.one_hot_smoother / self.num_classes
+ target_map[sampling_result.pos_inds, 5:] = gt_labels_one_hot[
+ sampling_result.pos_assigned_gt_inds]
+
+ neg_map = concat_anchors.new_zeros(
+ concat_anchors.size(0), dtype=torch.uint8)
+ neg_map[sampling_result.neg_inds] = 1
+
+ return target_map, neg_map
+
+ def responsible_flags(self, featmap_sizes: List[tuple], gt_bboxes: Tensor,
+ device: str) -> List[Tensor]:
+ """Generate responsible anchor flags of grid cells in multiple scales.
+
+ Args:
+ featmap_sizes (List[tuple]): List of feature map sizes in multiple
+ feature levels.
+ gt_bboxes (Tensor): Ground truth boxes, shape (n, 4).
+ device (str): Device where the anchors will be put on.
+
+ Return:
+ List[Tensor]: responsible flags of anchors in multiple level
+ """
+ assert self.num_levels == len(featmap_sizes)
+ multi_level_responsible_flags = []
+ for i in range(self.num_levels):
+ anchor_stride = self.prior_generator.strides[i]
+ feat_h, feat_w = featmap_sizes[i]
+ gt_cx = ((gt_bboxes[:, 0] + gt_bboxes[:, 2]) * 0.5).to(device)
+ gt_cy = ((gt_bboxes[:, 1] + gt_bboxes[:, 3]) * 0.5).to(device)
+ gt_grid_x = torch.floor(gt_cx / anchor_stride[0]).long()
+ gt_grid_y = torch.floor(gt_cy / anchor_stride[1]).long()
+ # row major indexing
+ gt_bboxes_grid_idx = gt_grid_y * feat_w + gt_grid_x
+
+ responsible_grid = torch.zeros(
+ feat_h * feat_w, dtype=torch.uint8, device=device)
+ responsible_grid[gt_bboxes_grid_idx] = 1
+
+ responsible_grid = responsible_grid[:, None].expand(
+ responsible_grid.size(0),
+ self.prior_generator.num_base_priors[i]).contiguous().view(-1)
+
+ multi_level_responsible_flags.append(responsible_grid)
+ return multi_level_responsible_flags
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/yolof_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/yolof_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..b5e5e6b7a92861bcd2ba3824df1f94270ba51160
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/yolof_head.py
@@ -0,0 +1,399 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule, is_norm
+from mmengine.model import bias_init_with_prob, constant_init, normal_init
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, InstanceList, OptInstanceList, reduce_mean
+from ..task_modules.prior_generators import anchor_inside_flags
+from ..utils import levels_to_images, multi_apply, unmap
+from .anchor_head import AnchorHead
+
+INF = 1e8
+
+
+@MODELS.register_module()
+class YOLOFHead(AnchorHead):
+ """Detection Head of `YOLOF `_
+
+ Args:
+ num_classes (int): The number of object classes (w/o background)
+ in_channels (list[int]): The number of input channels per scale.
+ cls_num_convs (int): The number of convolutions of cls branch.
+ Defaults to 2.
+ reg_num_convs (int): The number of convolutions of reg branch.
+ Defaults to 4.
+ norm_cfg (:obj:`ConfigDict` or dict): Config dict for normalization
+ layer. Defaults to ``dict(type='BN', requires_grad=True)``.
+ """
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: List[int],
+ num_cls_convs: int = 2,
+ num_reg_convs: int = 4,
+ norm_cfg: ConfigType = dict(type='BN', requires_grad=True),
+ **kwargs) -> None:
+ self.num_cls_convs = num_cls_convs
+ self.num_reg_convs = num_reg_convs
+ self.norm_cfg = norm_cfg
+ super().__init__(
+ num_classes=num_classes, in_channels=in_channels, **kwargs)
+
+ def _init_layers(self) -> None:
+ cls_subnet = []
+ bbox_subnet = []
+ for i in range(self.num_cls_convs):
+ cls_subnet.append(
+ ConvModule(
+ self.in_channels,
+ self.in_channels,
+ kernel_size=3,
+ padding=1,
+ norm_cfg=self.norm_cfg))
+ for i in range(self.num_reg_convs):
+ bbox_subnet.append(
+ ConvModule(
+ self.in_channels,
+ self.in_channels,
+ kernel_size=3,
+ padding=1,
+ norm_cfg=self.norm_cfg))
+ self.cls_subnet = nn.Sequential(*cls_subnet)
+ self.bbox_subnet = nn.Sequential(*bbox_subnet)
+ self.cls_score = nn.Conv2d(
+ self.in_channels,
+ self.num_base_priors * self.num_classes,
+ kernel_size=3,
+ stride=1,
+ padding=1)
+ self.bbox_pred = nn.Conv2d(
+ self.in_channels,
+ self.num_base_priors * 4,
+ kernel_size=3,
+ stride=1,
+ padding=1)
+ self.object_pred = nn.Conv2d(
+ self.in_channels,
+ self.num_base_priors,
+ kernel_size=3,
+ stride=1,
+ padding=1)
+
+ def init_weights(self) -> None:
+ for m in self.modules():
+ if isinstance(m, nn.Conv2d):
+ normal_init(m, mean=0, std=0.01)
+ if is_norm(m):
+ constant_init(m, 1)
+
+ # Use prior in model initialization to improve stability
+ bias_cls = bias_init_with_prob(0.01)
+ torch.nn.init.constant_(self.cls_score.bias, bias_cls)
+
+ def forward_single(self, x: Tensor) -> Tuple[Tensor, Tensor]:
+ """Forward feature of a single scale level.
+
+ Args:
+ x (Tensor): Features of a single scale level.
+
+ Returns:
+ tuple:
+ normalized_cls_score (Tensor): Normalized Cls scores for a \
+ single scale level, the channels number is \
+ num_base_priors * num_classes.
+ bbox_reg (Tensor): Box energies / deltas for a single scale \
+ level, the channels number is num_base_priors * 4.
+ """
+ cls_score = self.cls_score(self.cls_subnet(x))
+ N, _, H, W = cls_score.shape
+ cls_score = cls_score.view(N, -1, self.num_classes, H, W)
+
+ reg_feat = self.bbox_subnet(x)
+ bbox_reg = self.bbox_pred(reg_feat)
+ objectness = self.object_pred(reg_feat)
+
+ # implicit objectness
+ objectness = objectness.view(N, -1, 1, H, W)
+ normalized_cls_score = cls_score + objectness - torch.log(
+ 1. + torch.clamp(cls_score.exp(), max=INF) +
+ torch.clamp(objectness.exp(), max=INF))
+ normalized_cls_score = normalized_cls_score.view(N, -1, H, W)
+ return normalized_cls_score, bbox_reg
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ has shape (N, num_anchors * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ assert len(cls_scores) == 1
+ assert self.prior_generator.num_levels == 1
+
+ device = cls_scores[0].device
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+
+ # The output level is always 1
+ anchor_list = [anchors[0] for anchors in anchor_list]
+ valid_flag_list = [valid_flags[0] for valid_flags in valid_flag_list]
+
+ cls_scores_list = levels_to_images(cls_scores)
+ bbox_preds_list = levels_to_images(bbox_preds)
+
+ cls_reg_targets = self.get_targets(
+ cls_scores_list,
+ bbox_preds_list,
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore)
+ if cls_reg_targets is None:
+ return None
+ (batch_labels, batch_label_weights, avg_factor, batch_bbox_weights,
+ batch_pos_predicted_boxes, batch_target_boxes) = cls_reg_targets
+
+ flatten_labels = batch_labels.reshape(-1)
+ batch_label_weights = batch_label_weights.reshape(-1)
+ cls_score = cls_scores[0].permute(0, 2, 3,
+ 1).reshape(-1, self.cls_out_channels)
+
+ avg_factor = reduce_mean(
+ torch.tensor(avg_factor, dtype=torch.float, device=device)).item()
+
+ # classification loss
+ loss_cls = self.loss_cls(
+ cls_score,
+ flatten_labels,
+ batch_label_weights,
+ avg_factor=avg_factor)
+
+ # regression loss
+ if batch_pos_predicted_boxes.shape[0] == 0:
+ # no pos sample
+ loss_bbox = batch_pos_predicted_boxes.sum() * 0
+ else:
+ loss_bbox = self.loss_bbox(
+ batch_pos_predicted_boxes,
+ batch_target_boxes,
+ batch_bbox_weights.float(),
+ avg_factor=avg_factor)
+
+ return dict(loss_cls=loss_cls, loss_bbox=loss_bbox)
+
+ def get_targets(self,
+ cls_scores_list: List[Tensor],
+ bbox_preds_list: List[Tensor],
+ anchor_list: List[Tensor],
+ valid_flag_list: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None,
+ unmap_outputs: bool = True):
+ """Compute regression and classification targets for anchors in
+ multiple images.
+
+ Args:
+ cls_scores_list (list[Tensor]): Classification scores of
+ each image. each is a 4D-tensor, the shape is
+ (h * w, num_anchors * num_classes).
+ bbox_preds_list (list[Tensor]): Bbox preds of each image.
+ each is a 4D-tensor, the shape is (h * w, num_anchors * 4).
+ anchor_list (list[Tensor]): Anchors of each image. Each element of
+ is a tensor of shape (h * w * num_anchors, 4).
+ valid_flag_list (list[Tensor]): Valid flags of each image. Each
+ element of is a tensor of shape (h * w * num_anchors, )
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ unmap_outputs (bool): Whether to map outputs back to the original
+ set of anchors.
+
+ Returns:
+ tuple: Usually returns a tuple containing learning targets.
+
+ - batch_labels (Tensor): Label of all images. Each element \
+ of is a tensor of shape (batch, h * w * num_anchors)
+ - batch_label_weights (Tensor): Label weights of all images \
+ of is a tensor of shape (batch, h * w * num_anchors)
+ - num_total_pos (int): Number of positive samples in all \
+ images.
+ - num_total_neg (int): Number of negative samples in all \
+ images.
+ additional_returns: This function enables user-defined returns from
+ `self._get_targets_single`. These returns are currently refined
+ to properties at each feature map (i.e. having HxW dimension).
+ The results will be concatenated after the end
+ """
+ num_imgs = len(batch_img_metas)
+ assert len(anchor_list) == len(valid_flag_list) == num_imgs
+
+ # compute targets for each image
+ if batch_gt_instances_ignore is None:
+ batch_gt_instances_ignore = [None] * num_imgs
+ results = multi_apply(
+ self._get_targets_single,
+ bbox_preds_list,
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore,
+ unmap_outputs=unmap_outputs)
+ (all_labels, all_label_weights, pos_inds, neg_inds,
+ sampling_results_list) = results[:5]
+ # Get `avg_factor` of all images, which calculate in `SamplingResult`.
+ # When using sampling method, avg_factor is usually the sum of
+ # positive and negative priors. When using `PseudoSampler`,
+ # `avg_factor` is usually equal to the number of positive priors.
+ avg_factor = sum(
+ [results.avg_factor for results in sampling_results_list])
+ rest_results = list(results[5:]) # user-added return values
+
+ batch_labels = torch.stack(all_labels, 0)
+ batch_label_weights = torch.stack(all_label_weights, 0)
+
+ res = (batch_labels, batch_label_weights, avg_factor)
+ for i, rests in enumerate(rest_results): # user-added return values
+ rest_results[i] = torch.cat(rests, 0)
+
+ return res + tuple(rest_results)
+
+ def _get_targets_single(self,
+ bbox_preds: Tensor,
+ flat_anchors: Tensor,
+ valid_flags: Tensor,
+ gt_instances: InstanceData,
+ img_meta: dict,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ unmap_outputs: bool = True) -> tuple:
+ """Compute regression and classification targets for anchors in a
+ single image.
+
+ Args:
+ bbox_preds (Tensor): Bbox prediction of the image, which
+ shape is (h * w ,4)
+ flat_anchors (Tensor): Anchors of the image, which shape is
+ (h * w * num_anchors ,4)
+ valid_flags (Tensor): Valid flags of the image, which shape is
+ (h * w * num_anchors,).
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes`` and ``labels``
+ attributes.
+ img_meta (dict): Meta information for current image.
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ unmap_outputs (bool): Whether to map outputs back to the original
+ set of anchors.
+
+ Returns:
+ tuple:
+ labels (Tensor): Labels of image, which shape is
+ (h * w * num_anchors, ).
+ label_weights (Tensor): Label weights of image, which shape is
+ (h * w * num_anchors, ).
+ pos_inds (Tensor): Pos index of image.
+ neg_inds (Tensor): Neg index of image.
+ sampling_result (obj:`SamplingResult`): Sampling result.
+ pos_bbox_weights (Tensor): The Weight of using to calculate
+ the bbox branch loss, which shape is (num, ).
+ pos_predicted_boxes (Tensor): boxes predicted value of
+ using to calculate the bbox branch loss, which shape is
+ (num, 4).
+ pos_target_boxes (Tensor): boxes target value of
+ using to calculate the bbox branch loss, which shape is
+ (num, 4).
+ """
+ inside_flags = anchor_inside_flags(flat_anchors, valid_flags,
+ img_meta['img_shape'][:2],
+ self.train_cfg['allowed_border'])
+ if not inside_flags.any():
+ raise ValueError(
+ 'There is no valid anchor inside the image boundary. Please '
+ 'check the image size and anchor sizes, or set '
+ '``allowed_border`` to -1 to skip the condition.')
+
+ # assign gt and sample anchors
+ anchors = flat_anchors[inside_flags, :]
+ bbox_preds = bbox_preds.reshape(-1, 4)
+ bbox_preds = bbox_preds[inside_flags, :]
+
+ # decoded bbox
+ decoder_bbox_preds = self.bbox_coder.decode(anchors, bbox_preds)
+ pred_instances = InstanceData(
+ priors=anchors, decoder_priors=decoder_bbox_preds)
+ assign_result = self.assigner.assign(pred_instances, gt_instances,
+ gt_instances_ignore)
+
+ pos_bbox_weights = assign_result.get_extra_property('pos_idx')
+ pos_predicted_boxes = assign_result.get_extra_property(
+ 'pos_predicted_boxes')
+ pos_target_boxes = assign_result.get_extra_property('target_boxes')
+
+ sampling_result = self.sampler.sample(assign_result, pred_instances,
+ gt_instances)
+ num_valid_anchors = anchors.shape[0]
+ labels = anchors.new_full((num_valid_anchors, ),
+ self.num_classes,
+ dtype=torch.long)
+ label_weights = anchors.new_zeros(num_valid_anchors, dtype=torch.float)
+
+ pos_inds = sampling_result.pos_inds
+ neg_inds = sampling_result.neg_inds
+ if len(pos_inds) > 0:
+ labels[pos_inds] = sampling_result.pos_gt_labels
+ if self.train_cfg['pos_weight'] <= 0:
+ label_weights[pos_inds] = 1.0
+ else:
+ label_weights[pos_inds] = self.train_cfg['pos_weight']
+ if len(neg_inds) > 0:
+ label_weights[neg_inds] = 1.0
+
+ # map up to original set of anchors
+ if unmap_outputs:
+ num_total_anchors = flat_anchors.size(0)
+ labels = unmap(
+ labels, num_total_anchors, inside_flags,
+ fill=self.num_classes) # fill bg label
+ label_weights = unmap(label_weights, num_total_anchors,
+ inside_flags)
+
+ return (labels, label_weights, pos_inds, neg_inds, sampling_result,
+ pos_bbox_weights, pos_predicted_boxes, pos_target_boxes)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/yolox_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/yolox_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..00fe1e42766e4ca0052cf31d2e940dfab73fb200
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/dense_heads/yolox_head.py
@@ -0,0 +1,618 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+from typing import List, Optional, Sequence, Tuple, Union
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule, DepthwiseSeparableConvModule
+from mmcv.ops.nms import batched_nms
+from mmengine.config import ConfigDict
+from mmengine.model import bias_init_with_prob
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures.bbox import bbox_xyxy_to_cxcywh
+from mmdet.utils import (ConfigType, OptConfigType, OptInstanceList,
+ OptMultiConfig, reduce_mean)
+from ..task_modules.prior_generators import MlvlPointGenerator
+from ..task_modules.samplers import PseudoSampler
+from ..utils import multi_apply
+from .base_dense_head import BaseDenseHead
+
+
+@MODELS.register_module()
+class YOLOXHead(BaseDenseHead):
+ """YOLOXHead head used in `YOLOX `_.
+
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channels in the input feature map.
+ feat_channels (int): Number of hidden channels in stacking convs.
+ Defaults to 256
+ stacked_convs (int): Number of stacking convs of the head.
+ Defaults to (8, 16, 32).
+ strides (Sequence[int]): Downsample factor of each feature map.
+ Defaults to None.
+ use_depthwise (bool): Whether to depthwise separable convolution in
+ blocks. Defaults to False.
+ dcn_on_last_conv (bool): If true, use dcn in the last layer of
+ towers. Defaults to False.
+ conv_bias (bool or str): If specified as `auto`, it will be decided by
+ the norm_cfg. Bias of conv will be set as True if `norm_cfg` is
+ None, otherwise False. Defaults to "auto".
+ conv_cfg (:obj:`ConfigDict` or dict, optional): Config dict for
+ convolution layer. Defaults to None.
+ norm_cfg (:obj:`ConfigDict` or dict): Config dict for normalization
+ layer. Defaults to dict(type='BN', momentum=0.03, eps=0.001).
+ act_cfg (:obj:`ConfigDict` or dict): Config dict for activation layer.
+ Defaults to None.
+ loss_cls (:obj:`ConfigDict` or dict): Config of classification loss.
+ loss_bbox (:obj:`ConfigDict` or dict): Config of localization loss.
+ loss_obj (:obj:`ConfigDict` or dict): Config of objectness loss.
+ loss_l1 (:obj:`ConfigDict` or dict): Config of L1 loss.
+ train_cfg (:obj:`ConfigDict` or dict, optional): Training config of
+ anchor head. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): Testing config of
+ anchor head. Defaults to None.
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or
+ list[dict], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def __init__(
+ self,
+ num_classes: int,
+ in_channels: int,
+ feat_channels: int = 256,
+ stacked_convs: int = 2,
+ strides: Sequence[int] = (8, 16, 32),
+ use_depthwise: bool = False,
+ dcn_on_last_conv: bool = False,
+ conv_bias: Union[bool, str] = 'auto',
+ conv_cfg: OptConfigType = None,
+ norm_cfg: ConfigType = dict(type='BN', momentum=0.03, eps=0.001),
+ act_cfg: ConfigType = dict(type='Swish'),
+ loss_cls: ConfigType = dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ reduction='sum',
+ loss_weight=1.0),
+ loss_bbox: ConfigType = dict(
+ type='IoULoss',
+ mode='square',
+ eps=1e-16,
+ reduction='sum',
+ loss_weight=5.0),
+ loss_obj: ConfigType = dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ reduction='sum',
+ loss_weight=1.0),
+ loss_l1: ConfigType = dict(
+ type='L1Loss', reduction='sum', loss_weight=1.0),
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ init_cfg: OptMultiConfig = dict(
+ type='Kaiming',
+ layer='Conv2d',
+ a=math.sqrt(5),
+ distribution='uniform',
+ mode='fan_in',
+ nonlinearity='leaky_relu')
+ ) -> None:
+
+ super().__init__(init_cfg=init_cfg)
+ self.num_classes = num_classes
+ self.cls_out_channels = num_classes
+ self.in_channels = in_channels
+ self.feat_channels = feat_channels
+ self.stacked_convs = stacked_convs
+ self.strides = strides
+ self.use_depthwise = use_depthwise
+ self.dcn_on_last_conv = dcn_on_last_conv
+ assert conv_bias == 'auto' or isinstance(conv_bias, bool)
+ self.conv_bias = conv_bias
+ self.use_sigmoid_cls = True
+
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self.act_cfg = act_cfg
+
+ self.loss_cls: nn.Module = MODELS.build(loss_cls)
+ self.loss_bbox: nn.Module = MODELS.build(loss_bbox)
+ self.loss_obj: nn.Module = MODELS.build(loss_obj)
+
+ self.use_l1 = False # This flag will be modified by hooks.
+ self.loss_l1: nn.Module = MODELS.build(loss_l1)
+
+ self.prior_generator = MlvlPointGenerator(strides, offset=0)
+
+ self.test_cfg = test_cfg
+ self.train_cfg = train_cfg
+
+ if self.train_cfg:
+ self.assigner = TASK_UTILS.build(self.train_cfg['assigner'])
+ # YOLOX does not support sampling
+ self.sampler = PseudoSampler()
+
+ self._init_layers()
+
+ def _init_layers(self) -> None:
+ """Initialize heads for all level feature maps."""
+ self.multi_level_cls_convs = nn.ModuleList()
+ self.multi_level_reg_convs = nn.ModuleList()
+ self.multi_level_conv_cls = nn.ModuleList()
+ self.multi_level_conv_reg = nn.ModuleList()
+ self.multi_level_conv_obj = nn.ModuleList()
+ for _ in self.strides:
+ self.multi_level_cls_convs.append(self._build_stacked_convs())
+ self.multi_level_reg_convs.append(self._build_stacked_convs())
+ conv_cls, conv_reg, conv_obj = self._build_predictor()
+ self.multi_level_conv_cls.append(conv_cls)
+ self.multi_level_conv_reg.append(conv_reg)
+ self.multi_level_conv_obj.append(conv_obj)
+
+ def _build_stacked_convs(self) -> nn.Sequential:
+ """Initialize conv layers of a single level head."""
+ conv = DepthwiseSeparableConvModule \
+ if self.use_depthwise else ConvModule
+ stacked_convs = []
+ for i in range(self.stacked_convs):
+ chn = self.in_channels if i == 0 else self.feat_channels
+ if self.dcn_on_last_conv and i == self.stacked_convs - 1:
+ conv_cfg = dict(type='DCNv2')
+ else:
+ conv_cfg = self.conv_cfg
+ stacked_convs.append(
+ conv(
+ chn,
+ self.feat_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=self.norm_cfg,
+ act_cfg=self.act_cfg,
+ bias=self.conv_bias))
+ return nn.Sequential(*stacked_convs)
+
+ def _build_predictor(self) -> Tuple[nn.Module, nn.Module, nn.Module]:
+ """Initialize predictor layers of a single level head."""
+ conv_cls = nn.Conv2d(self.feat_channels, self.cls_out_channels, 1)
+ conv_reg = nn.Conv2d(self.feat_channels, 4, 1)
+ conv_obj = nn.Conv2d(self.feat_channels, 1, 1)
+ return conv_cls, conv_reg, conv_obj
+
+ def init_weights(self) -> None:
+ """Initialize weights of the head."""
+ super(YOLOXHead, self).init_weights()
+ # Use prior in model initialization to improve stability
+ bias_init = bias_init_with_prob(0.01)
+ for conv_cls, conv_obj in zip(self.multi_level_conv_cls,
+ self.multi_level_conv_obj):
+ conv_cls.bias.data.fill_(bias_init)
+ conv_obj.bias.data.fill_(bias_init)
+
+ def forward_single(self, x: Tensor, cls_convs: nn.Module,
+ reg_convs: nn.Module, conv_cls: nn.Module,
+ conv_reg: nn.Module,
+ conv_obj: nn.Module) -> Tuple[Tensor, Tensor, Tensor]:
+ """Forward feature of a single scale level."""
+
+ cls_feat = cls_convs(x)
+ reg_feat = reg_convs(x)
+
+ cls_score = conv_cls(cls_feat)
+ bbox_pred = conv_reg(reg_feat)
+ objectness = conv_obj(reg_feat)
+
+ return cls_score, bbox_pred, objectness
+
+ def forward(self, x: Tuple[Tensor]) -> Tuple[List]:
+ """Forward features from the upstream network.
+
+ Args:
+ x (Tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+ Returns:
+ Tuple[List]: A tuple of multi-level classification scores, bbox
+ predictions, and objectnesses.
+ """
+
+ return multi_apply(self.forward_single, x, self.multi_level_cls_convs,
+ self.multi_level_reg_convs,
+ self.multi_level_conv_cls,
+ self.multi_level_conv_reg,
+ self.multi_level_conv_obj)
+
+ def predict_by_feat(self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ objectnesses: Optional[List[Tensor]],
+ batch_img_metas: Optional[List[dict]] = None,
+ cfg: Optional[ConfigDict] = None,
+ rescale: bool = False,
+ with_nms: bool = True) -> List[InstanceData]:
+ """Transform a batch of output features extracted by the head into
+ bbox results.
+ Args:
+ cls_scores (list[Tensor]): Classification scores for all
+ scale levels, each is a 4D-tensor, has shape
+ (batch_size, num_priors * num_classes, H, W).
+ bbox_preds (list[Tensor]): Box energies / deltas for all
+ scale levels, each is a 4D-tensor, has shape
+ (batch_size, num_priors * 4, H, W).
+ objectnesses (list[Tensor], Optional): Score factor for
+ all scale level, each is a 4D-tensor, has shape
+ (batch_size, 1, H, W).
+ batch_img_metas (list[dict], Optional): Batch image meta info.
+ Defaults to None.
+ cfg (ConfigDict, optional): Test / postprocessing
+ configuration, if None, test_cfg would be used.
+ Defaults to None.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ list[:obj:`InstanceData`]: Object detection results of each image
+ after the post process. Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ assert len(cls_scores) == len(bbox_preds) == len(objectnesses)
+ cfg = self.test_cfg if cfg is None else cfg
+
+ num_imgs = len(batch_img_metas)
+ featmap_sizes = [cls_score.shape[2:] for cls_score in cls_scores]
+ mlvl_priors = self.prior_generator.grid_priors(
+ featmap_sizes,
+ dtype=cls_scores[0].dtype,
+ device=cls_scores[0].device,
+ with_stride=True)
+
+ # flatten cls_scores, bbox_preds and objectness
+ flatten_cls_scores = [
+ cls_score.permute(0, 2, 3, 1).reshape(num_imgs, -1,
+ self.cls_out_channels)
+ for cls_score in cls_scores
+ ]
+ flatten_bbox_preds = [
+ bbox_pred.permute(0, 2, 3, 1).reshape(num_imgs, -1, 4)
+ for bbox_pred in bbox_preds
+ ]
+ flatten_objectness = [
+ objectness.permute(0, 2, 3, 1).reshape(num_imgs, -1)
+ for objectness in objectnesses
+ ]
+
+ flatten_cls_scores = torch.cat(flatten_cls_scores, dim=1).sigmoid()
+ flatten_bbox_preds = torch.cat(flatten_bbox_preds, dim=1)
+ flatten_objectness = torch.cat(flatten_objectness, dim=1).sigmoid()
+ flatten_priors = torch.cat(mlvl_priors)
+
+ flatten_bboxes = self._bbox_decode(flatten_priors, flatten_bbox_preds)
+
+ result_list = []
+ for img_id, img_meta in enumerate(batch_img_metas):
+ max_scores, labels = torch.max(flatten_cls_scores[img_id], 1)
+ valid_mask = flatten_objectness[
+ img_id] * max_scores >= cfg.score_thr
+ results = InstanceData(
+ bboxes=flatten_bboxes[img_id][valid_mask],
+ scores=max_scores[valid_mask] *
+ flatten_objectness[img_id][valid_mask],
+ labels=labels[valid_mask])
+
+ result_list.append(
+ self._bbox_post_process(
+ results=results,
+ cfg=cfg,
+ rescale=rescale,
+ with_nms=with_nms,
+ img_meta=img_meta))
+
+ return result_list
+
+ def _bbox_decode(self, priors: Tensor, bbox_preds: Tensor) -> Tensor:
+ """Decode regression results (delta_x, delta_x, w, h) to bboxes (tl_x,
+ tl_y, br_x, br_y).
+
+ Args:
+ priors (Tensor): Center proiors of an image, has shape
+ (num_instances, 2).
+ bbox_preds (Tensor): Box energies / deltas for all instances,
+ has shape (batch_size, num_instances, 4).
+
+ Returns:
+ Tensor: Decoded bboxes in (tl_x, tl_y, br_x, br_y) format. Has
+ shape (batch_size, num_instances, 4).
+ """
+ xys = (bbox_preds[..., :2] * priors[:, 2:]) + priors[:, :2]
+ whs = bbox_preds[..., 2:].exp() * priors[:, 2:]
+
+ tl_x = (xys[..., 0] - whs[..., 0] / 2)
+ tl_y = (xys[..., 1] - whs[..., 1] / 2)
+ br_x = (xys[..., 0] + whs[..., 0] / 2)
+ br_y = (xys[..., 1] + whs[..., 1] / 2)
+
+ decoded_bboxes = torch.stack([tl_x, tl_y, br_x, br_y], -1)
+ return decoded_bboxes
+
+ def _bbox_post_process(self,
+ results: InstanceData,
+ cfg: ConfigDict,
+ rescale: bool = False,
+ with_nms: bool = True,
+ img_meta: Optional[dict] = None) -> InstanceData:
+ """bbox post-processing method.
+
+ The boxes would be rescaled to the original image scale and do
+ the nms operation. Usually `with_nms` is False is used for aug test.
+
+ Args:
+ results (:obj:`InstaceData`): Detection instance results,
+ each item has shape (num_bboxes, ).
+ cfg (mmengine.Config): Test / postprocessing configuration,
+ if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Default to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Default to True.
+ img_meta (dict, optional): Image meta info. Defaults to None.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+
+ if rescale:
+ assert img_meta.get('scale_factor') is not None
+ results.bboxes /= results.bboxes.new_tensor(
+ img_meta['scale_factor']).repeat((1, 2))
+
+ if with_nms and results.bboxes.numel() > 0:
+ det_bboxes, keep_idxs = batched_nms(results.bboxes, results.scores,
+ results.labels, cfg.nms)
+ results = results[keep_idxs]
+ # some nms would reweight the score, such as softnms
+ results.scores = det_bboxes[:, -1]
+ return results
+
+ def loss_by_feat(
+ self,
+ cls_scores: Sequence[Tensor],
+ bbox_preds: Sequence[Tensor],
+ objectnesses: Sequence[Tensor],
+ batch_gt_instances: Sequence[InstanceData],
+ batch_img_metas: Sequence[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (Sequence[Tensor]): Box scores for each scale level,
+ each is a 4D-tensor, the channel number is
+ num_priors * num_classes.
+ bbox_preds (Sequence[Tensor]): Box energies / deltas for each scale
+ level, each is a 4D-tensor, the channel number is
+ num_priors * 4.
+ objectnesses (Sequence[Tensor]): Score factor for
+ all scale level, each is a 4D-tensor, has shape
+ (batch_size, 1, H, W).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ Returns:
+ dict[str, Tensor]: A dictionary of losses.
+ """
+ num_imgs = len(batch_img_metas)
+ if batch_gt_instances_ignore is None:
+ batch_gt_instances_ignore = [None] * num_imgs
+
+ featmap_sizes = [cls_score.shape[2:] for cls_score in cls_scores]
+ mlvl_priors = self.prior_generator.grid_priors(
+ featmap_sizes,
+ dtype=cls_scores[0].dtype,
+ device=cls_scores[0].device,
+ with_stride=True)
+
+ flatten_cls_preds = [
+ cls_pred.permute(0, 2, 3, 1).reshape(num_imgs, -1,
+ self.cls_out_channels)
+ for cls_pred in cls_scores
+ ]
+ flatten_bbox_preds = [
+ bbox_pred.permute(0, 2, 3, 1).reshape(num_imgs, -1, 4)
+ for bbox_pred in bbox_preds
+ ]
+ flatten_objectness = [
+ objectness.permute(0, 2, 3, 1).reshape(num_imgs, -1)
+ for objectness in objectnesses
+ ]
+
+ flatten_cls_preds = torch.cat(flatten_cls_preds, dim=1)
+ flatten_bbox_preds = torch.cat(flatten_bbox_preds, dim=1)
+ flatten_objectness = torch.cat(flatten_objectness, dim=1)
+ flatten_priors = torch.cat(mlvl_priors)
+ flatten_bboxes = self._bbox_decode(flatten_priors, flatten_bbox_preds)
+
+ (pos_masks, cls_targets, obj_targets, bbox_targets, l1_targets,
+ num_fg_imgs) = multi_apply(
+ self._get_targets_single,
+ flatten_priors.unsqueeze(0).repeat(num_imgs, 1, 1),
+ flatten_cls_preds.detach(), flatten_bboxes.detach(),
+ flatten_objectness.detach(), batch_gt_instances, batch_img_metas,
+ batch_gt_instances_ignore)
+
+ # The experimental results show that 'reduce_mean' can improve
+ # performance on the COCO dataset.
+ num_pos = torch.tensor(
+ sum(num_fg_imgs),
+ dtype=torch.float,
+ device=flatten_cls_preds.device)
+ num_total_samples = max(reduce_mean(num_pos), 1.0)
+
+ pos_masks = torch.cat(pos_masks, 0)
+ cls_targets = torch.cat(cls_targets, 0)
+ obj_targets = torch.cat(obj_targets, 0)
+ bbox_targets = torch.cat(bbox_targets, 0)
+ if self.use_l1:
+ l1_targets = torch.cat(l1_targets, 0)
+
+ loss_obj = self.loss_obj(flatten_objectness.view(-1, 1),
+ obj_targets) / num_total_samples
+ if num_pos > 0:
+ loss_cls = self.loss_cls(
+ flatten_cls_preds.view(-1, self.num_classes)[pos_masks],
+ cls_targets) / num_total_samples
+ loss_bbox = self.loss_bbox(
+ flatten_bboxes.view(-1, 4)[pos_masks],
+ bbox_targets) / num_total_samples
+ else:
+ # Avoid cls and reg branch not participating in the gradient
+ # propagation when there is no ground-truth in the images.
+ # For more details, please refer to
+ # https://github.com/open-mmlab/mmdetection/issues/7298
+ loss_cls = flatten_cls_preds.sum() * 0
+ loss_bbox = flatten_bboxes.sum() * 0
+
+ loss_dict = dict(
+ loss_cls=loss_cls, loss_bbox=loss_bbox, loss_obj=loss_obj)
+
+ if self.use_l1:
+ if num_pos > 0:
+ loss_l1 = self.loss_l1(
+ flatten_bbox_preds.view(-1, 4)[pos_masks],
+ l1_targets) / num_total_samples
+ else:
+ # Avoid cls and reg branch not participating in the gradient
+ # propagation when there is no ground-truth in the images.
+ # For more details, please refer to
+ # https://github.com/open-mmlab/mmdetection/issues/7298
+ loss_l1 = flatten_bbox_preds.sum() * 0
+ loss_dict.update(loss_l1=loss_l1)
+
+ return loss_dict
+
+ @torch.no_grad()
+ def _get_targets_single(
+ self,
+ priors: Tensor,
+ cls_preds: Tensor,
+ decoded_bboxes: Tensor,
+ objectness: Tensor,
+ gt_instances: InstanceData,
+ img_meta: dict,
+ gt_instances_ignore: Optional[InstanceData] = None) -> tuple:
+ """Compute classification, regression, and objectness targets for
+ priors in a single image.
+
+ Args:
+ priors (Tensor): All priors of one image, a 2D-Tensor with shape
+ [num_priors, 4] in [cx, xy, stride_w, stride_y] format.
+ cls_preds (Tensor): Classification predictions of one image,
+ a 2D-Tensor with shape [num_priors, num_classes]
+ decoded_bboxes (Tensor): Decoded bboxes predictions of one image,
+ a 2D-Tensor with shape [num_priors, 4] in [tl_x, tl_y,
+ br_x, br_y] format.
+ objectness (Tensor): Objectness predictions of one image,
+ a 1D-Tensor with shape [num_priors]
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes`` and ``labels``
+ attributes.
+ img_meta (dict): Meta information for current image.
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ Returns:
+ tuple:
+ foreground_mask (list[Tensor]): Binary mask of foreground
+ targets.
+ cls_target (list[Tensor]): Classification targets of an image.
+ obj_target (list[Tensor]): Objectness targets of an image.
+ bbox_target (list[Tensor]): BBox targets of an image.
+ l1_target (int): BBox L1 targets of an image.
+ num_pos_per_img (int): Number of positive samples in an image.
+ """
+
+ num_priors = priors.size(0)
+ num_gts = len(gt_instances)
+ # No target
+ if num_gts == 0:
+ cls_target = cls_preds.new_zeros((0, self.num_classes))
+ bbox_target = cls_preds.new_zeros((0, 4))
+ l1_target = cls_preds.new_zeros((0, 4))
+ obj_target = cls_preds.new_zeros((num_priors, 1))
+ foreground_mask = cls_preds.new_zeros(num_priors).bool()
+ return (foreground_mask, cls_target, obj_target, bbox_target,
+ l1_target, 0)
+
+ # YOLOX uses center priors with 0.5 offset to assign targets,
+ # but use center priors without offset to regress bboxes.
+ offset_priors = torch.cat(
+ [priors[:, :2] + priors[:, 2:] * 0.5, priors[:, 2:]], dim=-1)
+
+ scores = cls_preds.sigmoid() * objectness.unsqueeze(1).sigmoid()
+ pred_instances = InstanceData(
+ bboxes=decoded_bboxes, scores=scores.sqrt_(), priors=offset_priors)
+ assign_result = self.assigner.assign(
+ pred_instances=pred_instances,
+ gt_instances=gt_instances,
+ gt_instances_ignore=gt_instances_ignore)
+
+ sampling_result = self.sampler.sample(assign_result, pred_instances,
+ gt_instances)
+ pos_inds = sampling_result.pos_inds
+ num_pos_per_img = pos_inds.size(0)
+
+ pos_ious = assign_result.max_overlaps[pos_inds]
+ # IOU aware classification score
+ cls_target = F.one_hot(sampling_result.pos_gt_labels,
+ self.num_classes) * pos_ious.unsqueeze(-1)
+ obj_target = torch.zeros_like(objectness).unsqueeze(-1)
+ obj_target[pos_inds] = 1
+ bbox_target = sampling_result.pos_gt_bboxes
+ l1_target = cls_preds.new_zeros((num_pos_per_img, 4))
+ if self.use_l1:
+ l1_target = self._get_l1_target(l1_target, bbox_target,
+ priors[pos_inds])
+ foreground_mask = torch.zeros_like(objectness).to(torch.bool)
+ foreground_mask[pos_inds] = 1
+ return (foreground_mask, cls_target, obj_target, bbox_target,
+ l1_target, num_pos_per_img)
+
+ def _get_l1_target(self,
+ l1_target: Tensor,
+ gt_bboxes: Tensor,
+ priors: Tensor,
+ eps: float = 1e-8) -> Tensor:
+ """Convert gt bboxes to center offset and log width height."""
+ gt_cxcywh = bbox_xyxy_to_cxcywh(gt_bboxes)
+ l1_target[:, :2] = (gt_cxcywh[:, :2] - priors[:, :2]) / priors[:, 2:]
+ l1_target[:, 2:] = torch.log(gt_cxcywh[:, 2:] / priors[:, 2:] + eps)
+ return l1_target
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..e5a06d2813c810504e12592506be9347111d6696
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/__init__.py
@@ -0,0 +1,75 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .atss import ATSS
+from .autoassign import AutoAssign
+from .base import BaseDetector
+from .base_detr import DetectionTransformer
+from .boxinst import BoxInst
+from .cascade_rcnn import CascadeRCNN
+from .centernet import CenterNet
+from .condinst import CondInst
+from .conditional_detr import ConditionalDETR
+from .cornernet import CornerNet
+from .crowddet import CrowdDet
+from .d2_wrapper import Detectron2Wrapper
+from .dab_detr import DABDETR
+from .ddod import DDOD
+from .ddq_detr import DDQDETR
+from .deformable_detr import DeformableDETR
+from .detr import DETR
+from .dino import DINO
+from .fast_rcnn import FastRCNN
+from .faster_rcnn import FasterRCNN
+from .fcos import FCOS
+from .fovea import FOVEA
+from .fsaf import FSAF
+from .gfl import GFL
+from .glip import GLIP
+from .grid_rcnn import GridRCNN
+from .grounding_dino import GroundingDINO
+from .htc import HybridTaskCascade
+from .kd_one_stage import KnowledgeDistillationSingleStageDetector
+from .lad import LAD
+from .mask2former import Mask2Former
+from .mask_rcnn import MaskRCNN
+from .mask_scoring_rcnn import MaskScoringRCNN
+from .maskformer import MaskFormer
+from .nasfcos import NASFCOS
+from .paa import PAA
+from .panoptic_fpn import PanopticFPN
+from .panoptic_two_stage_segmentor import TwoStagePanopticSegmentor
+from .point_rend import PointRend
+from .queryinst import QueryInst
+from .reppoints_detector import RepPointsDetector
+from .retinanet import RetinaNet
+from .rpn import RPN
+from .rtmdet import RTMDet
+from .scnet import SCNet
+from .semi_base import SemiBaseDetector
+from .single_stage import SingleStageDetector
+from .soft_teacher import SoftTeacher
+from .solo import SOLO
+from .solov2 import SOLOv2
+from .sparse_rcnn import SparseRCNN
+from .tood import TOOD
+from .trident_faster_rcnn import TridentFasterRCNN
+from .two_stage import TwoStageDetector
+from .vfnet import VFNet
+from .yolact import YOLACT
+from .yolo import YOLOV3
+from .yolof import YOLOF
+from .yolox import YOLOX
+
+__all__ = [
+ 'ATSS', 'BaseDetector', 'SingleStageDetector', 'TwoStageDetector', 'RPN',
+ 'KnowledgeDistillationSingleStageDetector', 'FastRCNN', 'FasterRCNN',
+ 'MaskRCNN', 'CascadeRCNN', 'HybridTaskCascade', 'RetinaNet', 'FCOS',
+ 'GridRCNN', 'MaskScoringRCNN', 'RepPointsDetector', 'FOVEA', 'FSAF',
+ 'NASFCOS', 'PointRend', 'GFL', 'CornerNet', 'PAA', 'YOLOV3', 'YOLACT',
+ 'VFNet', 'DETR', 'TridentFasterRCNN', 'SparseRCNN', 'SCNet', 'SOLO',
+ 'SOLOv2', 'DeformableDETR', 'AutoAssign', 'YOLOF', 'CenterNet', 'YOLOX',
+ 'TwoStagePanopticSegmentor', 'PanopticFPN', 'QueryInst', 'LAD', 'TOOD',
+ 'MaskFormer', 'DDOD', 'Mask2Former', 'SemiBaseDetector', 'SoftTeacher',
+ 'RTMDet', 'Detectron2Wrapper', 'CrowdDet', 'CondInst', 'BoxInst',
+ 'DetectionTransformer', 'ConditionalDETR', 'DINO', 'DABDETR', 'GLIP',
+ 'DDQDETR', 'GroundingDINO'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/atss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/atss.py
new file mode 100644
index 0000000000000000000000000000000000000000..0bfcc728dc4cc33c0b705a2ab22a4e3f4ad7386d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/atss.py
@@ -0,0 +1,41 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class ATSS(SingleStageDetector):
+ """Implementation of `ATSS `_
+
+ Args:
+ backbone (:obj:`ConfigDict` or dict): The backbone module.
+ neck (:obj:`ConfigDict` or dict): The neck module.
+ bbox_head (:obj:`ConfigDict` or dict): The bbox head module.
+ train_cfg (:obj:`ConfigDict` or dict, optional): The training config
+ of ATSS. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): The testing config
+ of ATSS. Defaults to None.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional): Config of
+ :class:`DetDataPreprocessor` to process the input data.
+ Defaults to None.
+ init_cfg (:obj:`ConfigDict` or dict, optional): the config to control
+ the initialization. Defaults to None.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/autoassign.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/autoassign.py
new file mode 100644
index 0000000000000000000000000000000000000000..a0b3570fe6e0c3812a72bc677038bb4e76b05576
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/autoassign.py
@@ -0,0 +1,43 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class AutoAssign(SingleStageDetector):
+ """Implementation of `AutoAssign: Differentiable Label Assignment for Dense
+ Object Detection `_
+
+ Args:
+ backbone (:obj:`ConfigDict` or dict): The backbone config.
+ neck (:obj:`ConfigDict` or dict): The neck config.
+ bbox_head (:obj:`ConfigDict` or dict): The bbox head config.
+ train_cfg (:obj:`ConfigDict` or dict, optional): The training config
+ of AutoAssign. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): The testing config
+ of AutoAssign. Defaults to None.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional): Config of
+ :class:`DetDataPreprocessor` to process the input data.
+ Defaults to None.
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or
+ list[dict], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None):
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/base.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/base.py
new file mode 100644
index 0000000000000000000000000000000000000000..1a193b0ca9ca3d2b42fda452004d5c97421f426c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/base.py
@@ -0,0 +1,156 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from abc import ABCMeta, abstractmethod
+from typing import Dict, List, Tuple, Union
+
+import torch
+from mmengine.model import BaseModel
+from torch import Tensor
+
+from mmdet.structures import DetDataSample, OptSampleList, SampleList
+from mmdet.utils import InstanceList, OptConfigType, OptMultiConfig
+from ..utils import samplelist_boxtype2tensor
+
+ForwardResults = Union[Dict[str, torch.Tensor], List[DetDataSample],
+ Tuple[torch.Tensor], torch.Tensor]
+
+
+class BaseDetector(BaseModel, metaclass=ABCMeta):
+ """Base class for detectors.
+
+ Args:
+ data_preprocessor (dict or ConfigDict, optional): The pre-process
+ config of :class:`BaseDataPreprocessor`. it usually includes,
+ ``pad_size_divisor``, ``pad_value``, ``mean`` and ``std``.
+ init_cfg (dict or ConfigDict, optional): the config to control the
+ initialization. Defaults to None.
+ """
+
+ def __init__(self,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None):
+ super().__init__(
+ data_preprocessor=data_preprocessor, init_cfg=init_cfg)
+
+ @property
+ def with_neck(self) -> bool:
+ """bool: whether the detector has a neck"""
+ return hasattr(self, 'neck') and self.neck is not None
+
+ # TODO: these properties need to be carefully handled
+ # for both single stage & two stage detectors
+ @property
+ def with_shared_head(self) -> bool:
+ """bool: whether the detector has a shared head in the RoI Head"""
+ return hasattr(self, 'roi_head') and self.roi_head.with_shared_head
+
+ @property
+ def with_bbox(self) -> bool:
+ """bool: whether the detector has a bbox head"""
+ return ((hasattr(self, 'roi_head') and self.roi_head.with_bbox)
+ or (hasattr(self, 'bbox_head') and self.bbox_head is not None))
+
+ @property
+ def with_mask(self) -> bool:
+ """bool: whether the detector has a mask head"""
+ return ((hasattr(self, 'roi_head') and self.roi_head.with_mask)
+ or (hasattr(self, 'mask_head') and self.mask_head is not None))
+
+ def forward(self,
+ inputs: torch.Tensor,
+ data_samples: OptSampleList = None,
+ mode: str = 'tensor') -> ForwardResults:
+ """The unified entry for a forward process in both training and test.
+
+ The method should accept three modes: "tensor", "predict" and "loss":
+
+ - "tensor": Forward the whole network and return tensor or tuple of
+ tensor without any post-processing, same as a common nn.Module.
+ - "predict": Forward and return the predictions, which are fully
+ processed to a list of :obj:`DetDataSample`.
+ - "loss": Forward and return a dict of losses according to the given
+ inputs and data samples.
+
+ Note that this method doesn't handle either back propagation or
+ parameter update, which are supposed to be done in :meth:`train_step`.
+
+ Args:
+ inputs (torch.Tensor): The input tensor with shape
+ (N, C, ...) in general.
+ data_samples (list[:obj:`DetDataSample`], optional): A batch of
+ data samples that contain annotations and predictions.
+ Defaults to None.
+ mode (str): Return what kind of value. Defaults to 'tensor'.
+
+ Returns:
+ The return type depends on ``mode``.
+
+ - If ``mode="tensor"``, return a tensor or a tuple of tensor.
+ - If ``mode="predict"``, return a list of :obj:`DetDataSample`.
+ - If ``mode="loss"``, return a dict of tensor.
+ """
+ if mode == 'loss':
+ return self.loss(inputs, data_samples)
+ elif mode == 'predict':
+ return self.predict(inputs, data_samples)
+ elif mode == 'tensor':
+ return self._forward(inputs, data_samples)
+ else:
+ raise RuntimeError(f'Invalid mode "{mode}". '
+ 'Only supports loss, predict and tensor mode')
+
+ @abstractmethod
+ def loss(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> Union[dict, tuple]:
+ """Calculate losses from a batch of inputs and data samples."""
+ pass
+
+ @abstractmethod
+ def predict(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> SampleList:
+ """Predict results from a batch of inputs and data samples with post-
+ processing."""
+ pass
+
+ @abstractmethod
+ def _forward(self,
+ batch_inputs: Tensor,
+ batch_data_samples: OptSampleList = None):
+ """Network forward process.
+
+ Usually includes backbone, neck and head forward without any post-
+ processing.
+ """
+ pass
+
+ @abstractmethod
+ def extract_feat(self, batch_inputs: Tensor):
+ """Extract features from images."""
+ pass
+
+ def add_pred_to_datasample(self, data_samples: SampleList,
+ results_list: InstanceList) -> SampleList:
+ """Add predictions to `DetDataSample`.
+
+ Args:
+ data_samples (list[:obj:`DetDataSample`], optional): A batch of
+ data samples that contain annotations and predictions.
+ results_list (list[:obj:`InstanceData`]): Detection results of
+ each image.
+
+ Returns:
+ list[:obj:`DetDataSample`]: Detection results of the
+ input images. Each DetDataSample usually contain
+ 'pred_instances'. And the ``pred_instances`` usually
+ contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ for data_sample, pred_instances in zip(data_samples, results_list):
+ data_sample.pred_instances = pred_instances
+ samplelist_boxtype2tensor(data_samples)
+ return data_samples
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/base_detr.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/base_detr.py
new file mode 100644
index 0000000000000000000000000000000000000000..88f00ec7408c389a1eb06beac6b383007f80b893
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/base_detr.py
@@ -0,0 +1,332 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from abc import ABCMeta, abstractmethod
+from typing import Dict, List, Tuple, Union
+
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import OptSampleList, SampleList
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .base import BaseDetector
+
+
+@MODELS.register_module()
+class DetectionTransformer(BaseDetector, metaclass=ABCMeta):
+ r"""Base class for Detection Transformer.
+
+ In Detection Transformer, an encoder is used to process output features of
+ neck, then several queries interact with the encoder features using a
+ decoder and do the regression and classification with the bounding box
+ head.
+
+ Args:
+ backbone (:obj:`ConfigDict` or dict): Config of the backbone.
+ neck (:obj:`ConfigDict` or dict, optional): Config of the neck.
+ Defaults to None.
+ encoder (:obj:`ConfigDict` or dict, optional): Config of the
+ Transformer encoder. Defaults to None.
+ decoder (:obj:`ConfigDict` or dict, optional): Config of the
+ Transformer decoder. Defaults to None.
+ bbox_head (:obj:`ConfigDict` or dict, optional): Config for the
+ bounding box head module. Defaults to None.
+ positional_encoding (:obj:`ConfigDict` or dict, optional): Config
+ of the positional encoding module. Defaults to None.
+ num_queries (int, optional): Number of decoder query in Transformer.
+ Defaults to 100.
+ train_cfg (:obj:`ConfigDict` or dict, optional): Training config of
+ the bounding box head module. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): Testing config of
+ the bounding box head module. Defaults to None.
+ data_preprocessor (dict or ConfigDict, optional): The pre-process
+ config of :class:`BaseDataPreprocessor`. it usually includes,
+ ``pad_size_divisor``, ``pad_value``, ``mean`` and ``std``.
+ Defaults to None.
+ init_cfg (:obj:`ConfigDict` or dict, optional): the config to control
+ the initialization. Defaults to None.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: OptConfigType = None,
+ encoder: OptConfigType = None,
+ decoder: OptConfigType = None,
+ bbox_head: OptConfigType = None,
+ positional_encoding: OptConfigType = None,
+ num_queries: int = 100,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ data_preprocessor=data_preprocessor, init_cfg=init_cfg)
+ # process args
+ bbox_head.update(train_cfg=train_cfg)
+ bbox_head.update(test_cfg=test_cfg)
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+ self.encoder = encoder
+ self.decoder = decoder
+ self.positional_encoding = positional_encoding
+ self.num_queries = num_queries
+
+ # init model layers
+ self.backbone = MODELS.build(backbone)
+ if neck is not None:
+ self.neck = MODELS.build(neck)
+ self.bbox_head = MODELS.build(bbox_head)
+ self._init_layers()
+
+ @abstractmethod
+ def _init_layers(self) -> None:
+ """Initialize layers except for backbone, neck and bbox_head."""
+ pass
+
+ def loss(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> Union[dict, list]:
+ """Calculate losses from a batch of inputs and data samples.
+
+ Args:
+ batch_inputs (Tensor): Input images of shape (bs, dim, H, W).
+ These should usually be mean centered and std scaled.
+ batch_data_samples (List[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict: A dictionary of loss components
+ """
+ img_feats = self.extract_feat(batch_inputs)
+ head_inputs_dict = self.forward_transformer(img_feats,
+ batch_data_samples)
+ losses = self.bbox_head.loss(
+ **head_inputs_dict, batch_data_samples=batch_data_samples)
+
+ return losses
+
+ def predict(self,
+ batch_inputs: Tensor,
+ batch_data_samples: SampleList,
+ rescale: bool = True) -> SampleList:
+ """Predict results from a batch of inputs and data samples with post-
+ processing.
+
+ Args:
+ batch_inputs (Tensor): Inputs, has shape (bs, dim, H, W).
+ batch_data_samples (List[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+ rescale (bool): Whether to rescale the results.
+ Defaults to True.
+
+ Returns:
+ list[:obj:`DetDataSample`]: Detection results of the input images.
+ Each DetDataSample usually contain 'pred_instances'. And the
+ `pred_instances` usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ img_feats = self.extract_feat(batch_inputs)
+ head_inputs_dict = self.forward_transformer(img_feats,
+ batch_data_samples)
+ results_list = self.bbox_head.predict(
+ **head_inputs_dict,
+ rescale=rescale,
+ batch_data_samples=batch_data_samples)
+ batch_data_samples = self.add_pred_to_datasample(
+ batch_data_samples, results_list)
+ return batch_data_samples
+
+ def _forward(
+ self,
+ batch_inputs: Tensor,
+ batch_data_samples: OptSampleList = None) -> Tuple[List[Tensor]]:
+ """Network forward process. Usually includes backbone, neck and head
+ forward without any post-processing.
+
+ Args:
+ batch_inputs (Tensor): Inputs, has shape (bs, dim, H, W).
+ batch_data_samples (List[:obj:`DetDataSample`], optional): The
+ batch data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+ Defaults to None.
+
+ Returns:
+ tuple[Tensor]: A tuple of features from ``bbox_head`` forward.
+ """
+ img_feats = self.extract_feat(batch_inputs)
+ head_inputs_dict = self.forward_transformer(img_feats,
+ batch_data_samples)
+ results = self.bbox_head.forward(**head_inputs_dict)
+ return results
+
+ def forward_transformer(self,
+ img_feats: Tuple[Tensor],
+ batch_data_samples: OptSampleList = None) -> Dict:
+ """Forward process of Transformer, which includes four steps:
+ 'pre_transformer' -> 'encoder' -> 'pre_decoder' -> 'decoder'. We
+ summarized the parameters flow of the existing DETR-like detector,
+ which can be illustrated as follow:
+
+ .. code:: text
+
+ img_feats & batch_data_samples
+ |
+ V
+ +-----------------+
+ | pre_transformer |
+ +-----------------+
+ | |
+ | V
+ | +-----------------+
+ | | forward_encoder |
+ | +-----------------+
+ | |
+ | V
+ | +---------------+
+ | | pre_decoder |
+ | +---------------+
+ | | |
+ V V |
+ +-----------------+ |
+ | forward_decoder | |
+ +-----------------+ |
+ | |
+ V V
+ head_inputs_dict
+
+ Args:
+ img_feats (tuple[Tensor]): Tuple of feature maps from neck. Each
+ feature map has shape (bs, dim, H, W).
+ batch_data_samples (list[:obj:`DetDataSample`], optional): The
+ batch data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+ Defaults to None.
+
+ Returns:
+ dict: The dictionary of bbox_head function inputs, which always
+ includes the `hidden_states` of the decoder output and may contain
+ `references` including the initial and intermediate references.
+ """
+ encoder_inputs_dict, decoder_inputs_dict = self.pre_transformer(
+ img_feats, batch_data_samples)
+
+ encoder_outputs_dict = self.forward_encoder(**encoder_inputs_dict)
+
+ tmp_dec_in, head_inputs_dict = self.pre_decoder(**encoder_outputs_dict)
+ decoder_inputs_dict.update(tmp_dec_in)
+
+ decoder_outputs_dict = self.forward_decoder(**decoder_inputs_dict)
+ head_inputs_dict.update(decoder_outputs_dict)
+ return head_inputs_dict
+
+ def extract_feat(self, batch_inputs: Tensor) -> Tuple[Tensor]:
+ """Extract features.
+
+ Args:
+ batch_inputs (Tensor): Image tensor, has shape (bs, dim, H, W).
+
+ Returns:
+ tuple[Tensor]: Tuple of feature maps from neck. Each feature map
+ has shape (bs, dim, H, W).
+ """
+ x = self.backbone(batch_inputs)
+ if self.with_neck:
+ x = self.neck(x)
+ return x
+
+ @abstractmethod
+ def pre_transformer(
+ self,
+ img_feats: Tuple[Tensor],
+ batch_data_samples: OptSampleList = None) -> Tuple[Dict, Dict]:
+ """Process image features before feeding them to the transformer.
+
+ Args:
+ img_feats (tuple[Tensor]): Tuple of feature maps from neck. Each
+ feature map has shape (bs, dim, H, W).
+ batch_data_samples (list[:obj:`DetDataSample`], optional): The
+ batch data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+ Defaults to None.
+
+ Returns:
+ tuple[dict, dict]: The first dict contains the inputs of encoder
+ and the second dict contains the inputs of decoder.
+
+ - encoder_inputs_dict (dict): The keyword args dictionary of
+ `self.forward_encoder()`, which includes 'feat', 'feat_mask',
+ 'feat_pos', and other algorithm-specific arguments.
+ - decoder_inputs_dict (dict): The keyword args dictionary of
+ `self.forward_decoder()`, which includes 'memory_mask', and
+ other algorithm-specific arguments.
+ """
+ pass
+
+ @abstractmethod
+ def forward_encoder(self, feat: Tensor, feat_mask: Tensor,
+ feat_pos: Tensor, **kwargs) -> Dict:
+ """Forward with Transformer encoder.
+
+ Args:
+ feat (Tensor): Sequential features, has shape (bs, num_feat_points,
+ dim).
+ feat_mask (Tensor): ByteTensor, the padding mask of the features,
+ has shape (bs, num_feat_points).
+ feat_pos (Tensor): The positional embeddings of the features, has
+ shape (bs, num_feat_points, dim).
+
+ Returns:
+ dict: The dictionary of encoder outputs, which includes the
+ `memory` of the encoder output and other algorithm-specific
+ arguments.
+ """
+ pass
+
+ @abstractmethod
+ def pre_decoder(self, memory: Tensor, **kwargs) -> Tuple[Dict, Dict]:
+ """Prepare intermediate variables before entering Transformer decoder,
+ such as `query`, `query_pos`, and `reference_points`.
+
+ Args:
+ memory (Tensor): The output embeddings of the Transformer encoder,
+ has shape (bs, num_feat_points, dim).
+
+ Returns:
+ tuple[dict, dict]: The first dict contains the inputs of decoder
+ and the second dict contains the inputs of the bbox_head function.
+
+ - decoder_inputs_dict (dict): The keyword dictionary args of
+ `self.forward_decoder()`, which includes 'query', 'query_pos',
+ 'memory', and other algorithm-specific arguments.
+ - head_inputs_dict (dict): The keyword dictionary args of the
+ bbox_head functions, which is usually empty, or includes
+ `enc_outputs_class` and `enc_outputs_class` when the detector
+ support 'two stage' or 'query selection' strategies.
+ """
+ pass
+
+ @abstractmethod
+ def forward_decoder(self, query: Tensor, query_pos: Tensor, memory: Tensor,
+ **kwargs) -> Dict:
+ """Forward with Transformer decoder.
+
+ Args:
+ query (Tensor): The queries of decoder inputs, has shape
+ (bs, num_queries, dim).
+ query_pos (Tensor): The positional queries of decoder inputs,
+ has shape (bs, num_queries, dim).
+ memory (Tensor): The output embeddings of the Transformer encoder,
+ has shape (bs, num_feat_points, dim).
+
+ Returns:
+ dict: The dictionary of decoder outputs, which includes the
+ `hidden_states` of the decoder output, `references` including
+ the initial and intermediate reference_points, and other
+ algorithm-specific arguments.
+ """
+ pass
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/boxinst.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/boxinst.py
new file mode 100644
index 0000000000000000000000000000000000000000..ca6b0bdd90a2a7e78f429a6822dbde6f809426da
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/boxinst.py
@@ -0,0 +1,28 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage_instance_seg import SingleStageInstanceSegmentor
+
+
+@MODELS.register_module()
+class BoxInst(SingleStageInstanceSegmentor):
+ """Implementation of `BoxInst `_"""
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ mask_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ mask_head=mask_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/cascade_rcnn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/cascade_rcnn.py
new file mode 100644
index 0000000000000000000000000000000000000000..ecf733ff104b99436fcc74130b0ccea12a0fa6d0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/cascade_rcnn.py
@@ -0,0 +1,29 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .two_stage import TwoStageDetector
+
+
+@MODELS.register_module()
+class CascadeRCNN(TwoStageDetector):
+ r"""Implementation of `Cascade R-CNN: Delving into High Quality Object
+ Detection `_"""
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: OptConfigType = None,
+ rpn_head: OptConfigType = None,
+ roi_head: OptConfigType = None,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ rpn_head=rpn_head,
+ roi_head=roi_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/centernet.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/centernet.py
new file mode 100644
index 0000000000000000000000000000000000000000..9c6622d6280227ecba9ede4aabf72c22a764e11d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/centernet.py
@@ -0,0 +1,29 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class CenterNet(SingleStageDetector):
+ """Implementation of CenterNet(Objects as Points)
+
+ .
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/condinst.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/condinst.py
new file mode 100644
index 0000000000000000000000000000000000000000..ed2dc99eea3faf7b03a3970d46a372d28eb89fe1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/condinst.py
@@ -0,0 +1,28 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage_instance_seg import SingleStageInstanceSegmentor
+
+
+@MODELS.register_module()
+class CondInst(SingleStageInstanceSegmentor):
+ """Implementation of `CondInst `_"""
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ mask_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ mask_head=mask_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/conditional_detr.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/conditional_detr.py
new file mode 100644
index 0000000000000000000000000000000000000000..d57868e63a2ece085a7e5b67ee93c921ba334830
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/conditional_detr.py
@@ -0,0 +1,74 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict
+
+import torch.nn as nn
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from ..layers import (ConditionalDetrTransformerDecoder,
+ DetrTransformerEncoder, SinePositionalEncoding)
+from .detr import DETR
+
+
+@MODELS.register_module()
+class ConditionalDETR(DETR):
+ r"""Implementation of `Conditional DETR for Fast Training Convergence.
+
+ `_.
+
+ Code is modified from the `official github repo
+ `_.
+ """
+
+ def _init_layers(self) -> None:
+ """Initialize layers except for backbone, neck and bbox_head."""
+ self.positional_encoding = SinePositionalEncoding(
+ **self.positional_encoding)
+ self.encoder = DetrTransformerEncoder(**self.encoder)
+ self.decoder = ConditionalDetrTransformerDecoder(**self.decoder)
+ self.embed_dims = self.encoder.embed_dims
+ # NOTE The embed_dims is typically passed from the inside out.
+ # For example in DETR, The embed_dims is passed as
+ # self_attn -> the first encoder layer -> encoder -> detector.
+ self.query_embedding = nn.Embedding(self.num_queries, self.embed_dims)
+
+ num_feats = self.positional_encoding.num_feats
+ assert num_feats * 2 == self.embed_dims, \
+ f'embed_dims should be exactly 2 times of num_feats. ' \
+ f'Found {self.embed_dims} and {num_feats}.'
+
+ def forward_decoder(self, query: Tensor, query_pos: Tensor, memory: Tensor,
+ memory_mask: Tensor, memory_pos: Tensor) -> Dict:
+ """Forward with Transformer decoder.
+
+ Args:
+ query (Tensor): The queries of decoder inputs, has shape
+ (bs, num_queries, dim).
+ query_pos (Tensor): The positional queries of decoder inputs,
+ has shape (bs, num_queries, dim).
+ memory (Tensor): The output embeddings of the Transformer encoder,
+ has shape (bs, num_feat_points, dim).
+ memory_mask (Tensor): ByteTensor, the padding mask of the memory,
+ has shape (bs, num_feat_points).
+ memory_pos (Tensor): The positional embeddings of memory, has
+ shape (bs, num_feat_points, dim).
+
+ Returns:
+ dict: The dictionary of decoder outputs, which includes the
+ `hidden_states` and `references` of the decoder output.
+
+ - hidden_states (Tensor): Has shape
+ (num_decoder_layers, bs, num_queries, dim)
+ - references (Tensor): Has shape
+ (bs, num_queries, 2)
+ """
+
+ hidden_states, references = self.decoder(
+ query=query,
+ key=memory,
+ query_pos=query_pos,
+ key_pos=memory_pos,
+ key_padding_mask=memory_mask)
+ head_inputs_dict = dict(
+ hidden_states=hidden_states, references=references)
+ return head_inputs_dict
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/cornernet.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/cornernet.py
new file mode 100644
index 0000000000000000000000000000000000000000..946af4dbe6ae339d44f8db265ff7f11b9e02d239
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/cornernet.py
@@ -0,0 +1,30 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class CornerNet(SingleStageDetector):
+ """CornerNet.
+
+ This detector is the implementation of the paper `CornerNet: Detecting
+ Objects as Paired Keypoints `_ .
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/crowddet.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/crowddet.py
new file mode 100644
index 0000000000000000000000000000000000000000..4f43bc08aa95756324381ee4182f001a008613c8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/crowddet.py
@@ -0,0 +1,45 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .two_stage import TwoStageDetector
+
+
+@MODELS.register_module()
+class CrowdDet(TwoStageDetector):
+ """Implementation of `CrowdDet `_
+
+ Args:
+ backbone (:obj:`ConfigDict` or dict): The backbone config.
+ rpn_head (:obj:`ConfigDict` or dict): The rpn config.
+ roi_head (:obj:`ConfigDict` or dict): The roi config.
+ train_cfg (:obj:`ConfigDict` or dict, optional): The training config
+ of FCOS. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): The testing config
+ of FCOS. Defaults to None.
+ neck (:obj:`ConfigDict` or dict): The neck config.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional): Config of
+ :class:`DetDataPreprocessor` to process the input data.
+ Defaults to None.
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or
+ list[dict], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ rpn_head: ConfigType,
+ roi_head: ConfigType,
+ train_cfg: ConfigType,
+ test_cfg: ConfigType,
+ neck: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ rpn_head=rpn_head,
+ roi_head=roi_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ init_cfg=init_cfg,
+ data_preprocessor=data_preprocessor)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/d2_wrapper.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/d2_wrapper.py
new file mode 100644
index 0000000000000000000000000000000000000000..3a2daa413e8fe0397ec37008d781ce449e7a26fd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/d2_wrapper.py
@@ -0,0 +1,291 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Union
+
+from mmengine.config import ConfigDict
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import BaseBoxes
+from mmdet.structures.mask import BitmapMasks, PolygonMasks
+from mmdet.utils import ConfigType
+from .base import BaseDetector
+
+try:
+ import detectron2
+ from detectron2.config import get_cfg
+ from detectron2.modeling import build_model
+ from detectron2.structures.masks import BitMasks as D2_BitMasks
+ from detectron2.structures.masks import PolygonMasks as D2_PolygonMasks
+ from detectron2.utils.events import EventStorage
+except ImportError:
+ detectron2 = None
+
+
+def _to_cfgnode_list(cfg: ConfigType,
+ config_list: list = [],
+ father_name: str = 'MODEL') -> tuple:
+ """Convert the key and value of mmengine.ConfigDict into a list.
+
+ Args:
+ cfg (ConfigDict): The detectron2 model config.
+ config_list (list): A list contains the key and value of ConfigDict.
+ Defaults to [].
+ father_name (str): The father name add before the key.
+ Defaults to "MODEL".
+
+ Returns:
+ tuple:
+
+ - config_list: A list contains the key and value of ConfigDict.
+ - father_name (str): The father name add before the key.
+ Defaults to "MODEL".
+ """
+ for key, value in cfg.items():
+ name = f'{father_name}.{key.upper()}'
+ if isinstance(value, ConfigDict) or isinstance(value, dict):
+ config_list, fater_name = \
+ _to_cfgnode_list(value, config_list, name)
+ else:
+ config_list.append(name)
+ config_list.append(value)
+
+ return config_list, father_name
+
+
+def convert_d2_pred_to_datasample(data_samples: SampleList,
+ d2_results_list: list) -> SampleList:
+ """Convert the Detectron2's result to DetDataSample.
+
+ Args:
+ data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+ d2_results_list (list): The list of the results of Detectron2's model.
+
+ Returns:
+ list[:obj:`DetDataSample`]: Detection results of the
+ input images. Each DetDataSample usually contain
+ 'pred_instances'. And the ``pred_instances`` usually
+ contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ assert len(data_samples) == len(d2_results_list)
+ for data_sample, d2_results in zip(data_samples, d2_results_list):
+ d2_instance = d2_results['instances']
+
+ results = InstanceData()
+ results.bboxes = d2_instance.pred_boxes.tensor
+ results.scores = d2_instance.scores
+ results.labels = d2_instance.pred_classes
+
+ if d2_instance.has('pred_masks'):
+ results.masks = d2_instance.pred_masks
+ data_sample.pred_instances = results
+
+ return data_samples
+
+
+@MODELS.register_module()
+class Detectron2Wrapper(BaseDetector):
+ """Wrapper of a Detectron2 model. Input/output formats of this class follow
+ MMDetection's convention, so a Detectron2 model can be trained and
+ evaluated in MMDetection.
+
+ Args:
+ detector (:obj:`ConfigDict` or dict): The module config of
+ Detectron2.
+ bgr_to_rgb (bool): whether to convert image from BGR to RGB.
+ Defaults to False.
+ rgb_to_bgr (bool): whether to convert image from RGB to BGR.
+ Defaults to False.
+ """
+
+ def __init__(self,
+ detector: ConfigType,
+ bgr_to_rgb: bool = False,
+ rgb_to_bgr: bool = False) -> None:
+ if detectron2 is None:
+ raise ImportError('Please install Detectron2 first')
+ assert not (bgr_to_rgb and rgb_to_bgr), (
+ '`bgr2rgb` and `rgb2bgr` cannot be set to True at the same time')
+ super().__init__()
+ self._channel_conversion = rgb_to_bgr or bgr_to_rgb
+ cfgnode_list, _ = _to_cfgnode_list(detector)
+ self.cfg = get_cfg()
+ self.cfg.merge_from_list(cfgnode_list)
+ self.d2_model = build_model(self.cfg)
+ self.storage = EventStorage()
+
+ def init_weights(self) -> None:
+ """Initialization Backbone.
+
+ NOTE: The initialization of other layers are in Detectron2,
+ if users want to change the initialization way, please
+ change the code in Detectron2.
+ """
+ from detectron2.checkpoint import DetectionCheckpointer
+ checkpointer = DetectionCheckpointer(model=self.d2_model)
+ checkpointer.load(self.cfg.MODEL.WEIGHTS, checkpointables=[])
+
+ def loss(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> Union[dict, tuple]:
+ """Calculate losses from a batch of inputs and data samples.
+
+ The inputs will first convert to the Detectron2 type and feed into
+ D2 models.
+
+ Args:
+ batch_inputs (Tensor): Input images of shape (N, C, H, W).
+ These should usually be mean centered and std scaled.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ d2_batched_inputs = self._convert_to_d2_inputs(
+ batch_inputs=batch_inputs,
+ batch_data_samples=batch_data_samples,
+ training=True)
+
+ with self.storage as storage: # noqa
+ losses = self.d2_model(d2_batched_inputs)
+ # storage contains some training information, such as cls_accuracy.
+ # you can use storage.latest() to get the detail information
+ return losses
+
+ def predict(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> SampleList:
+ """Predict results from a batch of inputs and data samples with post-
+ processing.
+
+ The inputs will first convert to the Detectron2 type and feed into
+ D2 models. And the results will convert back to the MMDet type.
+
+ Args:
+ batch_inputs (Tensor): Input images of shape (N, C, H, W).
+ These should usually be mean centered and std scaled.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+
+ Returns:
+ list[:obj:`DetDataSample`]: Detection results of the
+ input images. Each DetDataSample usually contain
+ 'pred_instances'. And the ``pred_instances`` usually
+ contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ d2_batched_inputs = self._convert_to_d2_inputs(
+ batch_inputs=batch_inputs,
+ batch_data_samples=batch_data_samples,
+ training=False)
+ # results in detectron2 has already rescale
+ d2_results_list = self.d2_model(d2_batched_inputs)
+ batch_data_samples = convert_d2_pred_to_datasample(
+ data_samples=batch_data_samples, d2_results_list=d2_results_list)
+
+ return batch_data_samples
+
+ def _forward(self, *args, **kwargs):
+ """Network forward process.
+
+ Usually includes backbone, neck and head forward without any post-
+ processing.
+ """
+ raise NotImplementedError(
+ f'`_forward` is not implemented in {self.__class__.__name__}')
+
+ def extract_feat(self, *args, **kwargs):
+ """Extract features from images.
+
+ `extract_feat` will not be used in obj:``Detectron2Wrapper``.
+ """
+ pass
+
+ def _convert_to_d2_inputs(self,
+ batch_inputs: Tensor,
+ batch_data_samples: SampleList,
+ training=True) -> list:
+ """Convert inputs type to support Detectron2's model.
+
+ Args:
+ batch_inputs (Tensor): Input images of shape (N, C, H, W).
+ These should usually be mean centered and std scaled.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+ training (bool): Whether to enable training time processing.
+
+ Returns:
+ list[dict]: A list of dict, which will be fed into Detectron2's
+ model. And the dict usually contains following keys.
+
+ - image (Tensor): Image in (C, H, W) format.
+ - instances (Instances): GT Instance.
+ - height (int): the output height resolution of the model
+ - width (int): the output width resolution of the model
+ """
+ from detectron2.data.detection_utils import filter_empty_instances
+ from detectron2.structures import Boxes, Instances
+
+ batched_d2_inputs = []
+ for image, data_samples in zip(batch_inputs, batch_data_samples):
+ d2_inputs = dict()
+ # deal with metainfo
+ meta_info = data_samples.metainfo
+ d2_inputs['file_name'] = meta_info['img_path']
+ d2_inputs['height'], d2_inputs['width'] = meta_info['ori_shape']
+ d2_inputs['image_id'] = meta_info['img_id']
+ # deal with image
+ if self._channel_conversion:
+ image = image[[2, 1, 0], ...]
+ d2_inputs['image'] = image
+ # deal with gt_instances
+ gt_instances = data_samples.gt_instances
+ d2_instances = Instances(meta_info['img_shape'])
+
+ gt_boxes = gt_instances.bboxes
+ # TODO: use mmdet.structures.box.get_box_tensor after PR 8658
+ # has merged
+ if isinstance(gt_boxes, BaseBoxes):
+ gt_boxes = gt_boxes.tensor
+ d2_instances.gt_boxes = Boxes(gt_boxes)
+
+ d2_instances.gt_classes = gt_instances.labels
+ if gt_instances.get('masks', None) is not None:
+ gt_masks = gt_instances.masks
+ if isinstance(gt_masks, PolygonMasks):
+ d2_instances.gt_masks = D2_PolygonMasks(gt_masks.masks)
+ elif isinstance(gt_masks, BitmapMasks):
+ d2_instances.gt_masks = D2_BitMasks(gt_masks.masks)
+ else:
+ raise TypeError('The type of `gt_mask` can be '
+ '`PolygonMasks` or `BitMasks`, but get '
+ f'{type(gt_masks)}.')
+ # convert to cpu and convert back to cuda to avoid
+ # some potential error
+ if training:
+ device = gt_boxes.device
+ d2_instances = filter_empty_instances(
+ d2_instances.to('cpu')).to(device)
+ d2_inputs['instances'] = d2_instances
+ batched_d2_inputs.append(d2_inputs)
+
+ return batched_d2_inputs
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/dab_detr.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/dab_detr.py
new file mode 100644
index 0000000000000000000000000000000000000000..b61301cf6660924f0832f4068841a4664797c585
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/dab_detr.py
@@ -0,0 +1,139 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, Tuple
+
+from mmengine.model import uniform_init
+from torch import Tensor, nn
+
+from mmdet.registry import MODELS
+from ..layers import SinePositionalEncoding
+from ..layers.transformer import (DABDetrTransformerDecoder,
+ DABDetrTransformerEncoder, inverse_sigmoid)
+from .detr import DETR
+
+
+@MODELS.register_module()
+class DABDETR(DETR):
+ r"""Implementation of `DAB-DETR:
+ Dynamic Anchor Boxes are Better Queries for DETR.
+
+ `_.
+
+ Code is modified from the `official github repo
+ `_.
+
+ Args:
+ with_random_refpoints (bool): Whether to randomly initialize query
+ embeddings and not update them during training.
+ Defaults to False.
+ num_patterns (int): Inspired by Anchor-DETR. Defaults to 0.
+ """
+
+ def __init__(self,
+ *args,
+ with_random_refpoints: bool = False,
+ num_patterns: int = 0,
+ **kwargs) -> None:
+ self.with_random_refpoints = with_random_refpoints
+ assert isinstance(num_patterns, int), \
+ f'num_patterns should be int but {num_patterns}.'
+ self.num_patterns = num_patterns
+
+ super().__init__(*args, **kwargs)
+
+ def _init_layers(self) -> None:
+ """Initialize layers except for backbone, neck and bbox_head."""
+ self.positional_encoding = SinePositionalEncoding(
+ **self.positional_encoding)
+ self.encoder = DABDetrTransformerEncoder(**self.encoder)
+ self.decoder = DABDetrTransformerDecoder(**self.decoder)
+ self.embed_dims = self.encoder.embed_dims
+ self.query_dim = self.decoder.query_dim
+ self.query_embedding = nn.Embedding(self.num_queries, self.query_dim)
+ if self.num_patterns > 0:
+ self.patterns = nn.Embedding(self.num_patterns, self.embed_dims)
+
+ num_feats = self.positional_encoding.num_feats
+ assert num_feats * 2 == self.embed_dims, \
+ f'embed_dims should be exactly 2 times of num_feats. ' \
+ f'Found {self.embed_dims} and {num_feats}.'
+
+ def init_weights(self) -> None:
+ """Initialize weights for Transformer and other components."""
+ super(DABDETR, self).init_weights()
+ if self.with_random_refpoints:
+ uniform_init(self.query_embedding)
+ self.query_embedding.weight.data[:, :2] = \
+ inverse_sigmoid(self.query_embedding.weight.data[:, :2])
+ self.query_embedding.weight.data[:, :2].requires_grad = False
+
+ def pre_decoder(self, memory: Tensor) -> Tuple[Dict, Dict]:
+ """Prepare intermediate variables before entering Transformer decoder,
+ such as `query`, `query_pos`.
+
+ Args:
+ memory (Tensor): The output embeddings of the Transformer encoder,
+ has shape (bs, num_feat_points, dim).
+
+ Returns:
+ tuple[dict, dict]: The first dict contains the inputs of decoder
+ and the second dict contains the inputs of the bbox_head function.
+
+ - decoder_inputs_dict (dict): The keyword args dictionary of
+ `self.forward_decoder()`, which includes 'query', 'query_pos',
+ 'memory' and 'reg_branches'.
+ - head_inputs_dict (dict): The keyword args dictionary of the
+ bbox_head functions, which is usually empty, or includes
+ `enc_outputs_class` and `enc_outputs_class` when the detector
+ support 'two stage' or 'query selection' strategies.
+ """
+ batch_size = memory.size(0)
+ query_pos = self.query_embedding.weight
+ query_pos = query_pos.unsqueeze(0).repeat(batch_size, 1, 1)
+ if self.num_patterns == 0:
+ query = query_pos.new_zeros(batch_size, self.num_queries,
+ self.embed_dims)
+ else:
+ query = self.patterns.weight[:, None, None, :]\
+ .repeat(1, self.num_queries, batch_size, 1)\
+ .view(-1, batch_size, self.embed_dims)\
+ .permute(1, 0, 2)
+ query_pos = query_pos.repeat(1, self.num_patterns, 1)
+
+ decoder_inputs_dict = dict(
+ query_pos=query_pos, query=query, memory=memory)
+ head_inputs_dict = dict()
+ return decoder_inputs_dict, head_inputs_dict
+
+ def forward_decoder(self, query: Tensor, query_pos: Tensor, memory: Tensor,
+ memory_mask: Tensor, memory_pos: Tensor) -> Dict:
+ """Forward with Transformer decoder.
+
+ Args:
+ query (Tensor): The queries of decoder inputs, has shape
+ (bs, num_queries, dim).
+ query_pos (Tensor): The positional queries of decoder inputs,
+ has shape (bs, num_queries, dim).
+ memory (Tensor): The output embeddings of the Transformer encoder,
+ has shape (bs, num_feat_points, dim).
+ memory_mask (Tensor): ByteTensor, the padding mask of the memory,
+ has shape (bs, num_feat_points).
+ memory_pos (Tensor): The positional embeddings of memory, has
+ shape (bs, num_feat_points, dim).
+
+ Returns:
+ dict: The dictionary of decoder outputs, which includes the
+ `hidden_states` and `references` of the decoder output.
+ """
+
+ hidden_states, references = self.decoder(
+ query=query,
+ key=memory,
+ query_pos=query_pos,
+ key_pos=memory_pos,
+ key_padding_mask=memory_mask,
+ reg_branches=self.bbox_head.
+ fc_reg # iterative refinement for anchor boxes
+ )
+ head_inputs_dict = dict(
+ hidden_states=hidden_states, references=references)
+ return head_inputs_dict
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/ddod.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/ddod.py
new file mode 100644
index 0000000000000000000000000000000000000000..3503a40c8eb6d6c0496ea0f31740acecf774113a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/ddod.py
@@ -0,0 +1,41 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class DDOD(SingleStageDetector):
+ """Implementation of `DDOD `_.
+
+ Args:
+ backbone (:obj:`ConfigDict` or dict): The backbone module.
+ neck (:obj:`ConfigDict` or dict): The neck module.
+ bbox_head (:obj:`ConfigDict` or dict): The bbox head module.
+ train_cfg (:obj:`ConfigDict` or dict, optional): The training config
+ of ATSS. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): The testing config
+ of ATSS. Defaults to None.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional): Config of
+ :class:`DetDataPreprocessor` to process the input data.
+ Defaults to None.
+ init_cfg (:obj:`ConfigDict` or dict, optional): the config to control
+ the initialization. Defaults to None.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/ddq_detr.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/ddq_detr.py
new file mode 100644
index 0000000000000000000000000000000000000000..57d4959d50ddd7a761d5e5c7a29d1f7f233f838a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/ddq_detr.py
@@ -0,0 +1,274 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, Tuple
+
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+from mmcv.ops import MultiScaleDeformableAttention, batched_nms
+from torch import Tensor, nn
+from torch.nn.init import normal_
+
+from mmdet.registry import MODELS
+from mmdet.structures import OptSampleList
+from mmdet.structures.bbox import bbox_cxcywh_to_xyxy
+from mmdet.utils import OptConfigType
+from ..layers import DDQTransformerDecoder
+from ..utils import align_tensor
+from .deformable_detr import DeformableDETR
+from .dino import DINO
+
+
+@MODELS.register_module()
+class DDQDETR(DINO):
+ r"""Implementation of `Dense Distinct Query for
+ End-to-End Object Detection `_
+
+ Code is modified from the `official github repo
+ `_.
+
+ Args:
+ dense_topk_ratio (float): Ratio of num_dense queries to num_queries.
+ Defaults to 1.5.
+ dqs_cfg (:obj:`ConfigDict` or dict, optional): Config of
+ Distinct Queries Selection. Defaults to nms with
+ `iou_threshold` = 0.8.
+ """
+
+ def __init__(self,
+ *args,
+ dense_topk_ratio: float = 1.5,
+ dqs_cfg: OptConfigType = dict(type='nms', iou_threshold=0.8),
+ **kwargs):
+ self.dense_topk_ratio = dense_topk_ratio
+ self.decoder_cfg = kwargs['decoder']
+ self.dqs_cfg = dqs_cfg
+ super().__init__(*args, **kwargs)
+
+ # a share dict in all moduls
+ # pass some intermediate results and config parameters
+ cache_dict = dict()
+ for m in self.modules():
+ m.cache_dict = cache_dict
+ # first element is the start index of matching queries
+ # second element is the number of matching queries
+ self.cache_dict['dis_query_info'] = [0, 0]
+
+ # mask for distinct queries in each decoder layer
+ self.cache_dict['distinct_query_mask'] = []
+ # pass to decoder do the dqs
+ self.cache_dict['cls_branches'] = self.bbox_head.cls_branches
+ # Used to construct the attention mask after dqs
+ self.cache_dict['num_heads'] = self.encoder.layers[
+ 0].self_attn.num_heads
+ # pass to decoder to do the dqs
+ self.cache_dict['dqs_cfg'] = self.dqs_cfg
+
+ def _init_layers(self) -> None:
+ """Initialize layers except for backbone, neck and bbox_head."""
+ super(DDQDETR, self)._init_layers()
+ self.decoder = DDQTransformerDecoder(**self.decoder_cfg)
+ self.query_embedding = None
+ self.query_map = nn.Linear(self.embed_dims, self.embed_dims)
+
+ def init_weights(self) -> None:
+ """Initialize weights for Transformer and other components."""
+ super(DeformableDETR, self).init_weights()
+ for coder in self.encoder, self.decoder:
+ for p in coder.parameters():
+ if p.dim() > 1:
+ nn.init.xavier_uniform_(p)
+ for m in self.modules():
+ if isinstance(m, MultiScaleDeformableAttention):
+ m.init_weights()
+ nn.init.xavier_uniform_(self.memory_trans_fc.weight)
+ normal_(self.level_embed)
+
+ def pre_decoder(
+ self,
+ memory: Tensor,
+ memory_mask: Tensor,
+ spatial_shapes: Tensor,
+ batch_data_samples: OptSampleList = None,
+ ) -> Tuple[Dict]:
+ """Prepare intermediate variables before entering Transformer decoder,
+ such as `query`, `memory`, and `reference_points`.
+
+ Args:
+ memory (Tensor): The output embeddings of the Transformer encoder,
+ has shape (bs, num_feat_points, dim).
+ memory_mask (Tensor): ByteTensor, the padding mask of the memory,
+ has shape (bs, num_feat_points). Will only be used when
+ `as_two_stage` is `True`.
+ spatial_shapes (Tensor): Spatial shapes of features in all levels.
+ With shape (num_levels, 2), last dimension represents (h, w).
+ Will only be used when `as_two_stage` is `True`.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+ Defaults to None.
+
+ Returns:
+ tuple[dict]: The decoder_inputs_dict and head_inputs_dict.
+
+ - decoder_inputs_dict (dict): The keyword dictionary args of
+ `self.forward_decoder()`, which includes 'query', 'memory',
+ `reference_points`, and `dn_mask`. The reference points of
+ decoder input here are 4D boxes, although it has `points`
+ in its name.
+ - head_inputs_dict (dict): The keyword dictionary args of the
+ bbox_head functions, which includes `topk_score`, `topk_coords`,
+ `dense_topk_score`, `dense_topk_coords`,
+ and `dn_meta`, when `self.training` is `True`, else is empty.
+ """
+ bs, _, c = memory.shape
+ output_memory, output_proposals = self.gen_encoder_output_proposals(
+ memory, memory_mask, spatial_shapes)
+ enc_outputs_class = self.bbox_head.cls_branches[
+ self.decoder.num_layers](
+ output_memory)
+ enc_outputs_coord_unact = self.bbox_head.reg_branches[
+ self.decoder.num_layers](output_memory) + output_proposals
+
+ if self.training:
+ # aux dense branch particularly in DDQ DETR, which doesn't exist
+ # in DINO.
+ # -1 is the aux head for the encoder
+ dense_enc_outputs_class = self.bbox_head.cls_branches[-1](
+ output_memory)
+ dense_enc_outputs_coord_unact = self.bbox_head.reg_branches[-1](
+ output_memory) + output_proposals
+
+ topk = self.num_queries
+ dense_topk = int(topk * self.dense_topk_ratio)
+
+ proposals = enc_outputs_coord_unact.sigmoid()
+ proposals = bbox_cxcywh_to_xyxy(proposals)
+ scores = enc_outputs_class.max(-1)[0].sigmoid()
+
+ if self.training:
+ # aux dense branch particularly in DDQ DETR, which doesn't exist
+ # in DINO.
+ dense_proposals = dense_enc_outputs_coord_unact.sigmoid()
+ dense_proposals = bbox_cxcywh_to_xyxy(dense_proposals)
+ dense_scores = dense_enc_outputs_class.max(-1)[0].sigmoid()
+
+ num_imgs = len(scores)
+ topk_score = []
+ topk_coords_unact = []
+ # Distinct query.
+ query = []
+
+ dense_topk_score = []
+ dense_topk_coords_unact = []
+ dense_query = []
+
+ for img_id in range(num_imgs):
+ single_proposals = proposals[img_id]
+ single_scores = scores[img_id]
+
+ # `batched_nms` of class scores and bbox coordinations is used
+ # particularly by DDQ DETR for region proposal generation,
+ # instead of `topk` of class scores by DINO.
+ _, keep_idxs = batched_nms(
+ single_proposals, single_scores,
+ torch.ones(len(single_scores), device=single_scores.device),
+ self.cache_dict['dqs_cfg'])
+
+ if self.training:
+ # aux dense branch particularly in DDQ DETR, which doesn't
+ # exist in DINO.
+ dense_single_proposals = dense_proposals[img_id]
+ dense_single_scores = dense_scores[img_id]
+ # sort according the score
+ # Only sort by classification score, neither nms nor topk is
+ # required. So input parameter `nms_cfg` = None.
+ _, dense_keep_idxs = batched_nms(
+ dense_single_proposals, dense_single_scores,
+ torch.ones(
+ len(dense_single_scores),
+ device=dense_single_scores.device), None)
+
+ dense_topk_score.append(dense_enc_outputs_class[img_id]
+ [dense_keep_idxs][:dense_topk])
+ dense_topk_coords_unact.append(
+ dense_enc_outputs_coord_unact[img_id][dense_keep_idxs]
+ [:dense_topk])
+
+ topk_score.append(enc_outputs_class[img_id][keep_idxs][:topk])
+
+ # Instead of initializing the content part with transformed
+ # coordinates in Deformable DETR, we fuse the feature map
+ # embedding of distinct positions as the content part, which
+ # makes the initial queries more distinct.
+ topk_coords_unact.append(
+ enc_outputs_coord_unact[img_id][keep_idxs][:topk])
+
+ map_memory = self.query_map(memory[img_id].detach())
+ query.append(map_memory[keep_idxs][:topk])
+ if self.training:
+ # aux dense branch particularly in DDQ DETR, which doesn't
+ # exist in DINO.
+ dense_query.append(map_memory[dense_keep_idxs][:dense_topk])
+
+ topk_score = align_tensor(topk_score, topk)
+ topk_coords_unact = align_tensor(topk_coords_unact, topk)
+ query = align_tensor(query, topk)
+ if self.training:
+ dense_topk_score = align_tensor(dense_topk_score)
+ dense_topk_coords_unact = align_tensor(dense_topk_coords_unact)
+
+ dense_query = align_tensor(dense_query)
+ num_dense_queries = dense_query.size(1)
+ if self.training:
+ query = torch.cat([query, dense_query], dim=1)
+ topk_coords_unact = torch.cat(
+ [topk_coords_unact, dense_topk_coords_unact], dim=1)
+
+ topk_coords = topk_coords_unact.sigmoid()
+ if self.training:
+ dense_topk_coords = topk_coords[:, -num_dense_queries:]
+ topk_coords = topk_coords[:, :-num_dense_queries]
+
+ topk_coords_unact = topk_coords_unact.detach()
+
+ if self.training:
+ dn_label_query, dn_bbox_query, dn_mask, dn_meta = \
+ self.dn_query_generator(batch_data_samples)
+ query = torch.cat([dn_label_query, query], dim=1)
+ reference_points = torch.cat([dn_bbox_query, topk_coords_unact],
+ dim=1)
+
+ # Update `dn_mask` to add mask for dense queries.
+ ori_size = dn_mask.size(-1)
+ new_size = dn_mask.size(-1) + num_dense_queries
+ new_dn_mask = dn_mask.new_ones((new_size, new_size)).bool()
+ dense_mask = torch.zeros(num_dense_queries,
+ num_dense_queries).bool()
+ self.cache_dict['dis_query_info'] = [dn_label_query.size(1), topk]
+
+ new_dn_mask[ori_size:, ori_size:] = dense_mask
+ new_dn_mask[:ori_size, :ori_size] = dn_mask
+ dn_meta['num_dense_queries'] = num_dense_queries
+ dn_mask = new_dn_mask
+ self.cache_dict['num_dense_queries'] = num_dense_queries
+ self.decoder.aux_reg_branches = self.bbox_head.aux_reg_branches
+
+ else:
+ self.cache_dict['dis_query_info'] = [0, topk]
+ reference_points = topk_coords_unact
+ dn_mask, dn_meta = None, None
+
+ reference_points = reference_points.sigmoid()
+
+ decoder_inputs_dict = dict(
+ query=query,
+ memory=memory,
+ reference_points=reference_points,
+ dn_mask=dn_mask)
+ head_inputs_dict = dict(
+ enc_outputs_class=topk_score,
+ enc_outputs_coord=topk_coords,
+ aux_enc_outputs_class=dense_topk_score,
+ aux_enc_outputs_coord=dense_topk_coords,
+ dn_meta=dn_meta) if self.training else dict()
+
+ return decoder_inputs_dict, head_inputs_dict
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/deformable_detr.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/deformable_detr.py
new file mode 100644
index 0000000000000000000000000000000000000000..0eb5cd2f95204542d5a9ace1a6d92e0b858c139f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/deformable_detr.py
@@ -0,0 +1,572 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+from typing import Dict, Tuple
+
+import torch
+import torch.nn.functional as F
+from mmcv.cnn.bricks.transformer import MultiScaleDeformableAttention
+from mmengine.model import xavier_init
+from torch import Tensor, nn
+from torch.nn.init import normal_
+
+from mmdet.registry import MODELS
+from mmdet.structures import OptSampleList
+from mmdet.utils import OptConfigType
+from ..layers import (DeformableDetrTransformerDecoder,
+ DeformableDetrTransformerEncoder, SinePositionalEncoding)
+from .base_detr import DetectionTransformer
+
+
+@MODELS.register_module()
+class DeformableDETR(DetectionTransformer):
+ r"""Implementation of `Deformable DETR: Deformable Transformers for
+ End-to-End Object Detection `_
+
+ Code is modified from the `official github repo
+ `_.
+
+ Args:
+ decoder (:obj:`ConfigDict` or dict, optional): Config of the
+ Transformer decoder. Defaults to None.
+ bbox_head (:obj:`ConfigDict` or dict, optional): Config for the
+ bounding box head module. Defaults to None.
+ with_box_refine (bool, optional): Whether to refine the references
+ in the decoder. Defaults to `False`.
+ as_two_stage (bool, optional): Whether to generate the proposal
+ from the outputs of encoder. Defaults to `False`.
+ num_feature_levels (int, optional): Number of feature levels.
+ Defaults to 4.
+ """
+
+ def __init__(self,
+ *args,
+ decoder: OptConfigType = None,
+ bbox_head: OptConfigType = None,
+ with_box_refine: bool = False,
+ as_two_stage: bool = False,
+ num_feature_levels: int = 4,
+ **kwargs) -> None:
+ self.with_box_refine = with_box_refine
+ self.as_two_stage = as_two_stage
+ self.num_feature_levels = num_feature_levels
+
+ if bbox_head is not None:
+ assert 'share_pred_layer' not in bbox_head and \
+ 'num_pred_layer' not in bbox_head and \
+ 'as_two_stage' not in bbox_head, \
+ 'The two keyword args `share_pred_layer`, `num_pred_layer`, ' \
+ 'and `as_two_stage are set in `detector.__init__()`, users ' \
+ 'should not set them in `bbox_head` config.'
+ # The last prediction layer is used to generate proposal
+ # from encode feature map when `as_two_stage` is `True`.
+ # And all the prediction layers should share parameters
+ # when `with_box_refine` is `True`.
+ bbox_head['share_pred_layer'] = not with_box_refine
+ bbox_head['num_pred_layer'] = (decoder['num_layers'] + 1) \
+ if self.as_two_stage else decoder['num_layers']
+ bbox_head['as_two_stage'] = as_two_stage
+
+ super().__init__(*args, decoder=decoder, bbox_head=bbox_head, **kwargs)
+
+ def _init_layers(self) -> None:
+ """Initialize layers except for backbone, neck and bbox_head."""
+ self.positional_encoding = SinePositionalEncoding(
+ **self.positional_encoding)
+ self.encoder = DeformableDetrTransformerEncoder(**self.encoder)
+ self.decoder = DeformableDetrTransformerDecoder(**self.decoder)
+ self.embed_dims = self.encoder.embed_dims
+ if not self.as_two_stage:
+ self.query_embedding = nn.Embedding(self.num_queries,
+ self.embed_dims * 2)
+ # NOTE The query_embedding will be split into query and query_pos
+ # in self.pre_decoder, hence, the embed_dims are doubled.
+
+ num_feats = self.positional_encoding.num_feats
+ assert num_feats * 2 == self.embed_dims, \
+ 'embed_dims should be exactly 2 times of num_feats. ' \
+ f'Found {self.embed_dims} and {num_feats}.'
+
+ self.level_embed = nn.Parameter(
+ torch.Tensor(self.num_feature_levels, self.embed_dims))
+
+ if self.as_two_stage:
+ self.memory_trans_fc = nn.Linear(self.embed_dims, self.embed_dims)
+ self.memory_trans_norm = nn.LayerNorm(self.embed_dims)
+ self.pos_trans_fc = nn.Linear(self.embed_dims * 2,
+ self.embed_dims * 2)
+ self.pos_trans_norm = nn.LayerNorm(self.embed_dims * 2)
+ else:
+ self.reference_points_fc = nn.Linear(self.embed_dims, 2)
+
+ def init_weights(self) -> None:
+ """Initialize weights for Transformer and other components."""
+ super().init_weights()
+ for coder in self.encoder, self.decoder:
+ for p in coder.parameters():
+ if p.dim() > 1:
+ nn.init.xavier_uniform_(p)
+ for m in self.modules():
+ if isinstance(m, MultiScaleDeformableAttention):
+ m.init_weights()
+ if self.as_two_stage:
+ nn.init.xavier_uniform_(self.memory_trans_fc.weight)
+ nn.init.xavier_uniform_(self.pos_trans_fc.weight)
+ else:
+ xavier_init(
+ self.reference_points_fc, distribution='uniform', bias=0.)
+ normal_(self.level_embed)
+
+ def pre_transformer(
+ self,
+ mlvl_feats: Tuple[Tensor],
+ batch_data_samples: OptSampleList = None) -> Tuple[Dict]:
+ """Process image features before feeding them to the transformer.
+
+ The forward procedure of the transformer is defined as:
+ 'pre_transformer' -> 'encoder' -> 'pre_decoder' -> 'decoder'
+ More details can be found at `TransformerDetector.forward_transformer`
+ in `mmdet/detector/base_detr.py`.
+
+ Args:
+ mlvl_feats (tuple[Tensor]): Multi-level features that may have
+ different resolutions, output from neck. Each feature has
+ shape (bs, dim, h_lvl, w_lvl), where 'lvl' means 'layer'.
+ batch_data_samples (list[:obj:`DetDataSample`], optional): The
+ batch data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+ Defaults to None.
+
+ Returns:
+ tuple[dict]: The first dict contains the inputs of encoder and the
+ second dict contains the inputs of decoder.
+
+ - encoder_inputs_dict (dict): The keyword args dictionary of
+ `self.forward_encoder()`, which includes 'feat', 'feat_mask',
+ and 'feat_pos'.
+ - decoder_inputs_dict (dict): The keyword args dictionary of
+ `self.forward_decoder()`, which includes 'memory_mask'.
+ """
+ batch_size = mlvl_feats[0].size(0)
+
+ # construct binary masks for the transformer.
+ assert batch_data_samples is not None
+ batch_input_shape = batch_data_samples[0].batch_input_shape
+ input_img_h, input_img_w = batch_input_shape
+ img_shape_list = [sample.img_shape for sample in batch_data_samples]
+ same_shape_flag = all([
+ s[0] == input_img_h and s[1] == input_img_w for s in img_shape_list
+ ])
+ # support torch2onnx without feeding masks
+ if torch.onnx.is_in_onnx_export() or same_shape_flag:
+ mlvl_masks = []
+ mlvl_pos_embeds = []
+ for feat in mlvl_feats:
+ mlvl_masks.append(None)
+ mlvl_pos_embeds.append(
+ self.positional_encoding(None, input=feat))
+ else:
+ masks = mlvl_feats[0].new_ones(
+ (batch_size, input_img_h, input_img_w))
+ for img_id in range(batch_size):
+ img_h, img_w = img_shape_list[img_id]
+ masks[img_id, :img_h, :img_w] = 0
+ # NOTE following the official DETR repo, non-zero
+ # values representing ignored positions, while
+ # zero values means valid positions.
+
+ mlvl_masks = []
+ mlvl_pos_embeds = []
+ for feat in mlvl_feats:
+ mlvl_masks.append(
+ F.interpolate(masks[None], size=feat.shape[-2:]).to(
+ torch.bool).squeeze(0))
+ mlvl_pos_embeds.append(
+ self.positional_encoding(mlvl_masks[-1]))
+
+ feat_flatten = []
+ lvl_pos_embed_flatten = []
+ mask_flatten = []
+ spatial_shapes = []
+ for lvl, (feat, mask, pos_embed) in enumerate(
+ zip(mlvl_feats, mlvl_masks, mlvl_pos_embeds)):
+ batch_size, c, h, w = feat.shape
+ spatial_shape = torch._shape_as_tensor(feat)[2:].to(feat.device)
+ # [bs, c, h_lvl, w_lvl] -> [bs, h_lvl*w_lvl, c]
+ feat = feat.view(batch_size, c, -1).permute(0, 2, 1)
+ pos_embed = pos_embed.view(batch_size, c, -1).permute(0, 2, 1)
+ lvl_pos_embed = pos_embed + self.level_embed[lvl].view(1, 1, -1)
+ # [bs, h_lvl, w_lvl] -> [bs, h_lvl*w_lvl]
+ if mask is not None:
+ mask = mask.flatten(1)
+
+ feat_flatten.append(feat)
+ lvl_pos_embed_flatten.append(lvl_pos_embed)
+ mask_flatten.append(mask)
+ spatial_shapes.append(spatial_shape)
+
+ # (bs, num_feat_points, dim)
+ feat_flatten = torch.cat(feat_flatten, 1)
+ lvl_pos_embed_flatten = torch.cat(lvl_pos_embed_flatten, 1)
+ # (bs, num_feat_points), where num_feat_points = sum_lvl(h_lvl*w_lvl)
+ if mask_flatten[0] is not None:
+ mask_flatten = torch.cat(mask_flatten, 1)
+ else:
+ mask_flatten = None
+
+ # (num_level, 2)
+ spatial_shapes = torch.cat(spatial_shapes).view(-1, 2)
+ level_start_index = torch.cat((
+ spatial_shapes.new_zeros((1, )), # (num_level)
+ spatial_shapes.prod(1).cumsum(0)[:-1]))
+ if mlvl_masks[0] is not None:
+ valid_ratios = torch.stack( # (bs, num_level, 2)
+ [self.get_valid_ratio(m) for m in mlvl_masks], 1)
+ else:
+ valid_ratios = mlvl_feats[0].new_ones(batch_size, len(mlvl_feats),
+ 2)
+
+ encoder_inputs_dict = dict(
+ feat=feat_flatten,
+ feat_mask=mask_flatten,
+ feat_pos=lvl_pos_embed_flatten,
+ spatial_shapes=spatial_shapes,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios)
+ decoder_inputs_dict = dict(
+ memory_mask=mask_flatten,
+ spatial_shapes=spatial_shapes,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios)
+ return encoder_inputs_dict, decoder_inputs_dict
+
+ def forward_encoder(self, feat: Tensor, feat_mask: Tensor,
+ feat_pos: Tensor, spatial_shapes: Tensor,
+ level_start_index: Tensor,
+ valid_ratios: Tensor) -> Dict:
+ """Forward with Transformer encoder.
+
+ The forward procedure of the transformer is defined as:
+ 'pre_transformer' -> 'encoder' -> 'pre_decoder' -> 'decoder'
+ More details can be found at `TransformerDetector.forward_transformer`
+ in `mmdet/detector/base_detr.py`.
+
+ Args:
+ feat (Tensor): Sequential features, has shape (bs, num_feat_points,
+ dim).
+ feat_mask (Tensor): ByteTensor, the padding mask of the features,
+ has shape (bs, num_feat_points).
+ feat_pos (Tensor): The positional embeddings of the features, has
+ shape (bs, num_feat_points, dim).
+ spatial_shapes (Tensor): Spatial shapes of features in all levels,
+ has shape (num_levels, 2), last dimension represents (h, w).
+ level_start_index (Tensor): The start index of each level.
+ A tensor has shape (num_levels, ) and can be represented
+ as [0, h_0*w_0, h_0*w_0+h_1*w_1, ...].
+ valid_ratios (Tensor): The ratios of the valid width and the valid
+ height relative to the width and the height of features in all
+ levels, has shape (bs, num_levels, 2).
+
+ Returns:
+ dict: The dictionary of encoder outputs, which includes the
+ `memory` of the encoder output.
+ """
+ memory = self.encoder(
+ query=feat,
+ query_pos=feat_pos,
+ key_padding_mask=feat_mask, # for self_attn
+ spatial_shapes=spatial_shapes,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios)
+ encoder_outputs_dict = dict(
+ memory=memory,
+ memory_mask=feat_mask,
+ spatial_shapes=spatial_shapes)
+ return encoder_outputs_dict
+
+ def pre_decoder(self, memory: Tensor, memory_mask: Tensor,
+ spatial_shapes: Tensor) -> Tuple[Dict, Dict]:
+ """Prepare intermediate variables before entering Transformer decoder,
+ such as `query`, `query_pos`, and `reference_points`.
+
+ The forward procedure of the transformer is defined as:
+ 'pre_transformer' -> 'encoder' -> 'pre_decoder' -> 'decoder'
+ More details can be found at `TransformerDetector.forward_transformer`
+ in `mmdet/detector/base_detr.py`.
+
+ Args:
+ memory (Tensor): The output embeddings of the Transformer encoder,
+ has shape (bs, num_feat_points, dim).
+ memory_mask (Tensor): ByteTensor, the padding mask of the memory,
+ has shape (bs, num_feat_points). It will only be used when
+ `as_two_stage` is `True`.
+ spatial_shapes (Tensor): Spatial shapes of features in all levels,
+ has shape (num_levels, 2), last dimension represents (h, w).
+ It will only be used when `as_two_stage` is `True`.
+
+ Returns:
+ tuple[dict, dict]: The decoder_inputs_dict and head_inputs_dict.
+
+ - decoder_inputs_dict (dict): The keyword dictionary args of
+ `self.forward_decoder()`, which includes 'query', 'query_pos',
+ 'memory', and `reference_points`. The reference_points of
+ decoder input here are 4D boxes when `as_two_stage` is `True`,
+ otherwise 2D points, although it has `points` in its name.
+ The reference_points in encoder is always 2D points.
+ - head_inputs_dict (dict): The keyword dictionary args of the
+ bbox_head functions, which includes `enc_outputs_class` and
+ `enc_outputs_coord`. They are both `None` when 'as_two_stage'
+ is `False`. The dict is empty when `self.training` is `False`.
+ """
+ batch_size, _, c = memory.shape
+ if self.as_two_stage:
+ output_memory, output_proposals = \
+ self.gen_encoder_output_proposals(
+ memory, memory_mask, spatial_shapes)
+ enc_outputs_class = self.bbox_head.cls_branches[
+ self.decoder.num_layers](
+ output_memory)
+ enc_outputs_coord_unact = self.bbox_head.reg_branches[
+ self.decoder.num_layers](output_memory) + output_proposals
+ enc_outputs_coord = enc_outputs_coord_unact.sigmoid()
+ # We only use the first channel in enc_outputs_class as foreground,
+ # the other (num_classes - 1) channels are actually not used.
+ # Its targets are set to be 0s, which indicates the first
+ # class (foreground) because we use [0, num_classes - 1] to
+ # indicate class labels, background class is indicated by
+ # num_classes (similar convention in RPN).
+ # See https://github.com/open-mmlab/mmdetection/blob/master/mmdet/models/dense_heads/deformable_detr_head.py#L241 # noqa
+ # This follows the official implementation of Deformable DETR.
+ topk_proposals = torch.topk(
+ enc_outputs_class[..., 0], self.num_queries, dim=1)[1]
+ topk_coords_unact = torch.gather(
+ enc_outputs_coord_unact, 1,
+ topk_proposals.unsqueeze(-1).repeat(1, 1, 4))
+ topk_coords_unact = topk_coords_unact.detach()
+ reference_points = topk_coords_unact.sigmoid()
+ pos_trans_out = self.pos_trans_fc(
+ self.get_proposal_pos_embed(topk_coords_unact))
+ pos_trans_out = self.pos_trans_norm(pos_trans_out)
+ query_pos, query = torch.split(pos_trans_out, c, dim=2)
+ else:
+ enc_outputs_class, enc_outputs_coord = None, None
+ query_embed = self.query_embedding.weight
+ query_pos, query = torch.split(query_embed, c, dim=1)
+ query_pos = query_pos.unsqueeze(0).expand(batch_size, -1, -1)
+ query = query.unsqueeze(0).expand(batch_size, -1, -1)
+ reference_points = self.reference_points_fc(query_pos).sigmoid()
+
+ decoder_inputs_dict = dict(
+ query=query,
+ query_pos=query_pos,
+ memory=memory,
+ reference_points=reference_points)
+ head_inputs_dict = dict(
+ enc_outputs_class=enc_outputs_class,
+ enc_outputs_coord=enc_outputs_coord) if self.training else dict()
+ return decoder_inputs_dict, head_inputs_dict
+
+ def forward_decoder(self, query: Tensor, query_pos: Tensor, memory: Tensor,
+ memory_mask: Tensor, reference_points: Tensor,
+ spatial_shapes: Tensor, level_start_index: Tensor,
+ valid_ratios: Tensor) -> Dict:
+ """Forward with Transformer decoder.
+
+ The forward procedure of the transformer is defined as:
+ 'pre_transformer' -> 'encoder' -> 'pre_decoder' -> 'decoder'
+ More details can be found at `TransformerDetector.forward_transformer`
+ in `mmdet/detector/base_detr.py`.
+
+ Args:
+ query (Tensor): The queries of decoder inputs, has shape
+ (bs, num_queries, dim).
+ query_pos (Tensor): The positional queries of decoder inputs,
+ has shape (bs, num_queries, dim).
+ memory (Tensor): The output embeddings of the Transformer encoder,
+ has shape (bs, num_feat_points, dim).
+ memory_mask (Tensor): ByteTensor, the padding mask of the memory,
+ has shape (bs, num_feat_points).
+ reference_points (Tensor): The initial reference, has shape
+ (bs, num_queries, 4) with the last dimension arranged as
+ (cx, cy, w, h) when `as_two_stage` is `True`, otherwise has
+ shape (bs, num_queries, 2) with the last dimension arranged as
+ (cx, cy).
+ spatial_shapes (Tensor): Spatial shapes of features in all levels,
+ has shape (num_levels, 2), last dimension represents (h, w).
+ level_start_index (Tensor): The start index of each level.
+ A tensor has shape (num_levels, ) and can be represented
+ as [0, h_0*w_0, h_0*w_0+h_1*w_1, ...].
+ valid_ratios (Tensor): The ratios of the valid width and the valid
+ height relative to the width and the height of features in all
+ levels, has shape (bs, num_levels, 2).
+
+ Returns:
+ dict: The dictionary of decoder outputs, which includes the
+ `hidden_states` of the decoder output and `references` including
+ the initial and intermediate reference_points.
+ """
+ inter_states, inter_references = self.decoder(
+ query=query,
+ value=memory,
+ query_pos=query_pos,
+ key_padding_mask=memory_mask, # for cross_attn
+ reference_points=reference_points,
+ spatial_shapes=spatial_shapes,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios,
+ reg_branches=self.bbox_head.reg_branches
+ if self.with_box_refine else None)
+ references = [reference_points, *inter_references]
+ decoder_outputs_dict = dict(
+ hidden_states=inter_states, references=references)
+ return decoder_outputs_dict
+
+ @staticmethod
+ def get_valid_ratio(mask: Tensor) -> Tensor:
+ """Get the valid radios of feature map in a level.
+
+ .. code:: text
+
+ |---> valid_W <---|
+ ---+-----------------+-----+---
+ A | | | A
+ | | | | |
+ | | | | |
+ valid_H | | | |
+ | | | | H
+ | | | | |
+ V | | | |
+ ---+-----------------+ | |
+ | | V
+ +-----------------------+---
+ |---------> W <---------|
+
+ The valid_ratios are defined as:
+ r_h = valid_H / H, r_w = valid_W / W
+ They are the factors to re-normalize the relative coordinates of the
+ image to the relative coordinates of the current level feature map.
+
+ Args:
+ mask (Tensor): Binary mask of a feature map, has shape (bs, H, W).
+
+ Returns:
+ Tensor: valid ratios [r_w, r_h] of a feature map, has shape (1, 2).
+ """
+ _, H, W = mask.shape
+ valid_H = torch.sum(~mask[:, :, 0], 1)
+ valid_W = torch.sum(~mask[:, 0, :], 1)
+ valid_ratio_h = valid_H.float() / H
+ valid_ratio_w = valid_W.float() / W
+ valid_ratio = torch.stack([valid_ratio_w, valid_ratio_h], -1)
+ return valid_ratio
+
+ def gen_encoder_output_proposals(
+ self, memory: Tensor, memory_mask: Tensor,
+ spatial_shapes: Tensor) -> Tuple[Tensor, Tensor]:
+ """Generate proposals from encoded memory. The function will only be
+ used when `as_two_stage` is `True`.
+
+ Args:
+ memory (Tensor): The output embeddings of the Transformer encoder,
+ has shape (bs, num_feat_points, dim).
+ memory_mask (Tensor): ByteTensor, the padding mask of the memory,
+ has shape (bs, num_feat_points).
+ spatial_shapes (Tensor): Spatial shapes of features in all levels,
+ has shape (num_levels, 2), last dimension represents (h, w).
+
+ Returns:
+ tuple: A tuple of transformed memory and proposals.
+
+ - output_memory (Tensor): The transformed memory for obtaining
+ top-k proposals, has shape (bs, num_feat_points, dim).
+ - output_proposals (Tensor): The inverse-normalized proposal, has
+ shape (batch_size, num_keys, 4) with the last dimension arranged
+ as (cx, cy, w, h).
+ """
+
+ bs = memory.size(0)
+ proposals = []
+ _cur = 0 # start index in the sequence of the current level
+ for lvl, HW in enumerate(spatial_shapes):
+ H, W = HW
+
+ if memory_mask is not None:
+ mask_flatten_ = memory_mask[:, _cur:(_cur + H * W)].view(
+ bs, H, W, 1)
+ valid_H = torch.sum(~mask_flatten_[:, :, 0, 0],
+ 1).unsqueeze(-1)
+ valid_W = torch.sum(~mask_flatten_[:, 0, :, 0],
+ 1).unsqueeze(-1)
+ scale = torch.cat([valid_W, valid_H], 1).view(bs, 1, 1, 2)
+ else:
+ if not isinstance(HW, torch.Tensor):
+ HW = memory.new_tensor(HW)
+ scale = HW.unsqueeze(0).flip(dims=[0, 1]).view(1, 1, 1, 2)
+ grid_y, grid_x = torch.meshgrid(
+ torch.linspace(
+ 0, H - 1, H, dtype=torch.float32, device=memory.device),
+ torch.linspace(
+ 0, W - 1, W, dtype=torch.float32, device=memory.device))
+ grid = torch.cat([grid_x.unsqueeze(-1), grid_y.unsqueeze(-1)], -1)
+ grid = (grid.unsqueeze(0).expand(bs, -1, -1, -1) + 0.5) / scale
+ wh = torch.ones_like(grid) * 0.05 * (2.0**lvl)
+ proposal = torch.cat((grid, wh), -1).view(bs, -1, 4)
+ proposals.append(proposal)
+ _cur += (H * W)
+ output_proposals = torch.cat(proposals, 1)
+ # do not use `all` to make it exportable to onnx
+ output_proposals_valid = (
+ (output_proposals > 0.01) & (output_proposals < 0.99)).sum(
+ -1, keepdim=True) == output_proposals.shape[-1]
+ # inverse_sigmoid
+ output_proposals = torch.log(output_proposals / (1 - output_proposals))
+ if memory_mask is not None:
+ output_proposals = output_proposals.masked_fill(
+ memory_mask.unsqueeze(-1), float('inf'))
+ output_proposals = output_proposals.masked_fill(
+ ~output_proposals_valid, float('inf'))
+
+ output_memory = memory
+ if memory_mask is not None:
+ output_memory = output_memory.masked_fill(
+ memory_mask.unsqueeze(-1), float(0))
+ output_memory = output_memory.masked_fill(~output_proposals_valid,
+ float(0))
+ output_memory = self.memory_trans_fc(output_memory)
+ output_memory = self.memory_trans_norm(output_memory)
+ # [bs, sum(hw), 2]
+ return output_memory, output_proposals
+
+ @staticmethod
+ def get_proposal_pos_embed(proposals: Tensor,
+ num_pos_feats: int = 128,
+ temperature: int = 10000) -> Tensor:
+ """Get the position embedding of the proposal.
+
+ Args:
+ proposals (Tensor): Not normalized proposals, has shape
+ (bs, num_queries, 4) with the last dimension arranged as
+ (cx, cy, w, h).
+ num_pos_feats (int, optional): The feature dimension for each
+ position along x, y, w, and h-axis. Note the final returned
+ dimension for each position is 4 times of num_pos_feats.
+ Default to 128.
+ temperature (int, optional): The temperature used for scaling the
+ position embedding. Defaults to 10000.
+
+ Returns:
+ Tensor: The position embedding of proposal, has shape
+ (bs, num_queries, num_pos_feats * 4), with the last dimension
+ arranged as (cx, cy, w, h)
+ """
+ scale = 2 * math.pi
+ dim_t = torch.arange(
+ num_pos_feats, dtype=torch.float32, device=proposals.device)
+ dim_t = temperature**(2 * (dim_t // 2) / num_pos_feats)
+ # N, L, 4
+ proposals = proposals.sigmoid() * scale
+ # N, L, 4, 128
+ pos = proposals[:, :, :, None] / dim_t
+ # N, L, 4, 64, 2
+ pos = torch.stack((pos[:, :, :, 0::2].sin(), pos[:, :, :, 1::2].cos()),
+ dim=4).flatten(2)
+ return pos
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/detr.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/detr.py
new file mode 100644
index 0000000000000000000000000000000000000000..7895e9ecb4eb66cb75d173c191c2128c3f55c197
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/detr.py
@@ -0,0 +1,225 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, Tuple
+
+import torch
+import torch.nn.functional as F
+from torch import Tensor, nn
+
+from mmdet.registry import MODELS
+from mmdet.structures import OptSampleList
+from ..layers import (DetrTransformerDecoder, DetrTransformerEncoder,
+ SinePositionalEncoding)
+from .base_detr import DetectionTransformer
+
+
+@MODELS.register_module()
+class DETR(DetectionTransformer):
+ r"""Implementation of `DETR: End-to-End Object Detection with Transformers.
+
+ `_.
+
+ Code is modified from the `official github repo
+ `_.
+ """
+
+ def _init_layers(self) -> None:
+ """Initialize layers except for backbone, neck and bbox_head."""
+ self.positional_encoding = SinePositionalEncoding(
+ **self.positional_encoding)
+ self.encoder = DetrTransformerEncoder(**self.encoder)
+ self.decoder = DetrTransformerDecoder(**self.decoder)
+ self.embed_dims = self.encoder.embed_dims
+ # NOTE The embed_dims is typically passed from the inside out.
+ # For example in DETR, The embed_dims is passed as
+ # self_attn -> the first encoder layer -> encoder -> detector.
+ self.query_embedding = nn.Embedding(self.num_queries, self.embed_dims)
+
+ num_feats = self.positional_encoding.num_feats
+ assert num_feats * 2 == self.embed_dims, \
+ 'embed_dims should be exactly 2 times of num_feats. ' \
+ f'Found {self.embed_dims} and {num_feats}.'
+
+ def init_weights(self) -> None:
+ """Initialize weights for Transformer and other components."""
+ super().init_weights()
+ for coder in self.encoder, self.decoder:
+ for p in coder.parameters():
+ if p.dim() > 1:
+ nn.init.xavier_uniform_(p)
+
+ def pre_transformer(
+ self,
+ img_feats: Tuple[Tensor],
+ batch_data_samples: OptSampleList = None) -> Tuple[Dict, Dict]:
+ """Prepare the inputs of the Transformer.
+
+ The forward procedure of the transformer is defined as:
+ 'pre_transformer' -> 'encoder' -> 'pre_decoder' -> 'decoder'
+ More details can be found at `TransformerDetector.forward_transformer`
+ in `mmdet/detector/base_detr.py`.
+
+ Args:
+ img_feats (Tuple[Tensor]): Tuple of features output from the neck,
+ has shape (bs, c, h, w).
+ batch_data_samples (List[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such as
+ `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+ Defaults to None.
+
+ Returns:
+ tuple[dict, dict]: The first dict contains the inputs of encoder
+ and the second dict contains the inputs of decoder.
+
+ - encoder_inputs_dict (dict): The keyword args dictionary of
+ `self.forward_encoder()`, which includes 'feat', 'feat_mask',
+ and 'feat_pos'.
+ - decoder_inputs_dict (dict): The keyword args dictionary of
+ `self.forward_decoder()`, which includes 'memory_mask',
+ and 'memory_pos'.
+ """
+
+ feat = img_feats[-1] # NOTE img_feats contains only one feature.
+ batch_size, feat_dim, _, _ = feat.shape
+ # construct binary masks which for the transformer.
+ assert batch_data_samples is not None
+ batch_input_shape = batch_data_samples[0].batch_input_shape
+ input_img_h, input_img_w = batch_input_shape
+ img_shape_list = [sample.img_shape for sample in batch_data_samples]
+ same_shape_flag = all([
+ s[0] == input_img_h and s[1] == input_img_w for s in img_shape_list
+ ])
+ if torch.onnx.is_in_onnx_export() or same_shape_flag:
+ masks = None
+ # [batch_size, embed_dim, h, w]
+ pos_embed = self.positional_encoding(masks, input=feat)
+ else:
+ masks = feat.new_ones((batch_size, input_img_h, input_img_w))
+ for img_id in range(batch_size):
+ img_h, img_w = img_shape_list[img_id]
+ masks[img_id, :img_h, :img_w] = 0
+ # NOTE following the official DETR repo, non-zero values represent
+ # ignored positions, while zero values mean valid positions.
+
+ masks = F.interpolate(
+ masks.unsqueeze(1),
+ size=feat.shape[-2:]).to(torch.bool).squeeze(1)
+ # [batch_size, embed_dim, h, w]
+ pos_embed = self.positional_encoding(masks)
+
+ # use `view` instead of `flatten` for dynamically exporting to ONNX
+ # [bs, c, h, w] -> [bs, h*w, c]
+ feat = feat.view(batch_size, feat_dim, -1).permute(0, 2, 1)
+ pos_embed = pos_embed.view(batch_size, feat_dim, -1).permute(0, 2, 1)
+ # [bs, h, w] -> [bs, h*w]
+ if masks is not None:
+ masks = masks.view(batch_size, -1)
+
+ # prepare transformer_inputs_dict
+ encoder_inputs_dict = dict(
+ feat=feat, feat_mask=masks, feat_pos=pos_embed)
+ decoder_inputs_dict = dict(memory_mask=masks, memory_pos=pos_embed)
+ return encoder_inputs_dict, decoder_inputs_dict
+
+ def forward_encoder(self, feat: Tensor, feat_mask: Tensor,
+ feat_pos: Tensor) -> Dict:
+ """Forward with Transformer encoder.
+
+ The forward procedure of the transformer is defined as:
+ 'pre_transformer' -> 'encoder' -> 'pre_decoder' -> 'decoder'
+ More details can be found at `TransformerDetector.forward_transformer`
+ in `mmdet/detector/base_detr.py`.
+
+ Args:
+ feat (Tensor): Sequential features, has shape (bs, num_feat_points,
+ dim).
+ feat_mask (Tensor): ByteTensor, the padding mask of the features,
+ has shape (bs, num_feat_points).
+ feat_pos (Tensor): The positional embeddings of the features, has
+ shape (bs, num_feat_points, dim).
+
+ Returns:
+ dict: The dictionary of encoder outputs, which includes the
+ `memory` of the encoder output.
+ """
+ memory = self.encoder(
+ query=feat, query_pos=feat_pos,
+ key_padding_mask=feat_mask) # for self_attn
+ encoder_outputs_dict = dict(memory=memory)
+ return encoder_outputs_dict
+
+ def pre_decoder(self, memory: Tensor) -> Tuple[Dict, Dict]:
+ """Prepare intermediate variables before entering Transformer decoder,
+ such as `query`, `query_pos`.
+
+ The forward procedure of the transformer is defined as:
+ 'pre_transformer' -> 'encoder' -> 'pre_decoder' -> 'decoder'
+ More details can be found at `TransformerDetector.forward_transformer`
+ in `mmdet/detector/base_detr.py`.
+
+ Args:
+ memory (Tensor): The output embeddings of the Transformer encoder,
+ has shape (bs, num_feat_points, dim).
+
+ Returns:
+ tuple[dict, dict]: The first dict contains the inputs of decoder
+ and the second dict contains the inputs of the bbox_head function.
+
+ - decoder_inputs_dict (dict): The keyword args dictionary of
+ `self.forward_decoder()`, which includes 'query', 'query_pos',
+ 'memory'.
+ - head_inputs_dict (dict): The keyword args dictionary of the
+ bbox_head functions, which is usually empty, or includes
+ `enc_outputs_class` and `enc_outputs_class` when the detector
+ support 'two stage' or 'query selection' strategies.
+ """
+
+ batch_size = memory.size(0) # (bs, num_feat_points, dim)
+ query_pos = self.query_embedding.weight
+ # (num_queries, dim) -> (bs, num_queries, dim)
+ query_pos = query_pos.unsqueeze(0).repeat(batch_size, 1, 1)
+ query = torch.zeros_like(query_pos)
+
+ decoder_inputs_dict = dict(
+ query_pos=query_pos, query=query, memory=memory)
+ head_inputs_dict = dict()
+ return decoder_inputs_dict, head_inputs_dict
+
+ def forward_decoder(self, query: Tensor, query_pos: Tensor, memory: Tensor,
+ memory_mask: Tensor, memory_pos: Tensor) -> Dict:
+ """Forward with Transformer decoder.
+
+ The forward procedure of the transformer is defined as:
+ 'pre_transformer' -> 'encoder' -> 'pre_decoder' -> 'decoder'
+ More details can be found at `TransformerDetector.forward_transformer`
+ in `mmdet/detector/base_detr.py`.
+
+ Args:
+ query (Tensor): The queries of decoder inputs, has shape
+ (bs, num_queries, dim).
+ query_pos (Tensor): The positional queries of decoder inputs,
+ has shape (bs, num_queries, dim).
+ memory (Tensor): The output embeddings of the Transformer encoder,
+ has shape (bs, num_feat_points, dim).
+ memory_mask (Tensor): ByteTensor, the padding mask of the memory,
+ has shape (bs, num_feat_points).
+ memory_pos (Tensor): The positional embeddings of memory, has
+ shape (bs, num_feat_points, dim).
+
+ Returns:
+ dict: The dictionary of decoder outputs, which includes the
+ `hidden_states` of the decoder output.
+
+ - hidden_states (Tensor): Has shape
+ (num_decoder_layers, bs, num_queries, dim)
+ """
+
+ hidden_states = self.decoder(
+ query=query,
+ key=memory,
+ value=memory,
+ query_pos=query_pos,
+ key_pos=memory_pos,
+ key_padding_mask=memory_mask) # for cross_attn
+
+ head_inputs_dict = dict(hidden_states=hidden_states)
+ return head_inputs_dict
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/dino.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/dino.py
new file mode 100644
index 0000000000000000000000000000000000000000..ade47f531d27246511cafc2997a07d58677538a7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/dino.py
@@ -0,0 +1,287 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, Optional, Tuple
+
+import torch
+from torch import Tensor, nn
+from torch.nn.init import normal_
+
+from mmdet.registry import MODELS
+from mmdet.structures import OptSampleList
+from mmdet.utils import OptConfigType
+from ..layers import (CdnQueryGenerator, DeformableDetrTransformerEncoder,
+ DinoTransformerDecoder, SinePositionalEncoding)
+from .deformable_detr import DeformableDETR, MultiScaleDeformableAttention
+
+
+@MODELS.register_module()
+class DINO(DeformableDETR):
+ r"""Implementation of `DINO: DETR with Improved DeNoising Anchor Boxes
+ for End-to-End Object Detection `_
+
+ Code is modified from the `official github repo
+ `_.
+
+ Args:
+ dn_cfg (:obj:`ConfigDict` or dict, optional): Config of denoising
+ query generator. Defaults to `None`.
+ """
+
+ def __init__(self, *args, dn_cfg: OptConfigType = None, **kwargs) -> None:
+ super().__init__(*args, **kwargs)
+ assert self.as_two_stage, 'as_two_stage must be True for DINO'
+ assert self.with_box_refine, 'with_box_refine must be True for DINO'
+
+ if dn_cfg is not None:
+ assert 'num_classes' not in dn_cfg and \
+ 'num_queries' not in dn_cfg and \
+ 'hidden_dim' not in dn_cfg, \
+ 'The three keyword args `num_classes`, `embed_dims`, and ' \
+ '`num_matching_queries` are set in `detector.__init__()`, ' \
+ 'users should not set them in `dn_cfg` config.'
+ dn_cfg['num_classes'] = self.bbox_head.num_classes
+ dn_cfg['embed_dims'] = self.embed_dims
+ dn_cfg['num_matching_queries'] = self.num_queries
+ self.dn_query_generator = CdnQueryGenerator(**dn_cfg)
+
+ def _init_layers(self) -> None:
+ """Initialize layers except for backbone, neck and bbox_head."""
+ self.positional_encoding = SinePositionalEncoding(
+ **self.positional_encoding)
+ self.encoder = DeformableDetrTransformerEncoder(**self.encoder)
+ self.decoder = DinoTransformerDecoder(**self.decoder)
+ self.embed_dims = self.encoder.embed_dims
+ self.query_embedding = nn.Embedding(self.num_queries, self.embed_dims)
+ # NOTE In DINO, the query_embedding only contains content
+ # queries, while in Deformable DETR, the query_embedding
+ # contains both content and spatial queries, and in DETR,
+ # it only contains spatial queries.
+
+ num_feats = self.positional_encoding.num_feats
+ assert num_feats * 2 == self.embed_dims, \
+ f'embed_dims should be exactly 2 times of num_feats. ' \
+ f'Found {self.embed_dims} and {num_feats}.'
+
+ self.level_embed = nn.Parameter(
+ torch.Tensor(self.num_feature_levels, self.embed_dims))
+ self.memory_trans_fc = nn.Linear(self.embed_dims, self.embed_dims)
+ self.memory_trans_norm = nn.LayerNorm(self.embed_dims)
+
+ def init_weights(self) -> None:
+ """Initialize weights for Transformer and other components."""
+ super(DeformableDETR, self).init_weights()
+ for coder in self.encoder, self.decoder:
+ for p in coder.parameters():
+ if p.dim() > 1:
+ nn.init.xavier_uniform_(p)
+ for m in self.modules():
+ if isinstance(m, MultiScaleDeformableAttention):
+ m.init_weights()
+ nn.init.xavier_uniform_(self.memory_trans_fc.weight)
+ nn.init.xavier_uniform_(self.query_embedding.weight)
+ normal_(self.level_embed)
+
+ def forward_transformer(
+ self,
+ img_feats: Tuple[Tensor],
+ batch_data_samples: OptSampleList = None,
+ ) -> Dict:
+ """Forward process of Transformer.
+
+ The forward procedure of the transformer is defined as:
+ 'pre_transformer' -> 'encoder' -> 'pre_decoder' -> 'decoder'
+ More details can be found at `TransformerDetector.forward_transformer`
+ in `mmdet/detector/base_detr.py`.
+ The difference is that the ground truth in `batch_data_samples` is
+ required for the `pre_decoder` to prepare the query of DINO.
+ Additionally, DINO inherits the `pre_transformer` method and the
+ `forward_encoder` method of DeformableDETR. More details about the
+ two methods can be found in `mmdet/detector/deformable_detr.py`.
+
+ Args:
+ img_feats (tuple[Tensor]): Tuple of feature maps from neck. Each
+ feature map has shape (bs, dim, H, W).
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+ Defaults to None.
+
+ Returns:
+ dict: The dictionary of bbox_head function inputs, which always
+ includes the `hidden_states` of the decoder output and may contain
+ `references` including the initial and intermediate references.
+ """
+ encoder_inputs_dict, decoder_inputs_dict = self.pre_transformer(
+ img_feats, batch_data_samples)
+
+ encoder_outputs_dict = self.forward_encoder(**encoder_inputs_dict)
+
+ tmp_dec_in, head_inputs_dict = self.pre_decoder(
+ **encoder_outputs_dict, batch_data_samples=batch_data_samples)
+ decoder_inputs_dict.update(tmp_dec_in)
+
+ decoder_outputs_dict = self.forward_decoder(**decoder_inputs_dict)
+ head_inputs_dict.update(decoder_outputs_dict)
+ return head_inputs_dict
+
+ def pre_decoder(
+ self,
+ memory: Tensor,
+ memory_mask: Tensor,
+ spatial_shapes: Tensor,
+ batch_data_samples: OptSampleList = None,
+ ) -> Tuple[Dict]:
+ """Prepare intermediate variables before entering Transformer decoder,
+ such as `query`, `query_pos`, and `reference_points`.
+
+ Args:
+ memory (Tensor): The output embeddings of the Transformer encoder,
+ has shape (bs, num_feat_points, dim).
+ memory_mask (Tensor): ByteTensor, the padding mask of the memory,
+ has shape (bs, num_feat_points). Will only be used when
+ `as_two_stage` is `True`.
+ spatial_shapes (Tensor): Spatial shapes of features in all levels.
+ With shape (num_levels, 2), last dimension represents (h, w).
+ Will only be used when `as_two_stage` is `True`.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+ Defaults to None.
+
+ Returns:
+ tuple[dict]: The decoder_inputs_dict and head_inputs_dict.
+
+ - decoder_inputs_dict (dict): The keyword dictionary args of
+ `self.forward_decoder()`, which includes 'query', 'memory',
+ `reference_points`, and `dn_mask`. The reference points of
+ decoder input here are 4D boxes, although it has `points`
+ in its name.
+ - head_inputs_dict (dict): The keyword dictionary args of the
+ bbox_head functions, which includes `topk_score`, `topk_coords`,
+ and `dn_meta` when `self.training` is `True`, else is empty.
+ """
+ bs, _, c = memory.shape
+ cls_out_features = self.bbox_head.cls_branches[
+ self.decoder.num_layers].out_features
+
+ output_memory, output_proposals = self.gen_encoder_output_proposals(
+ memory, memory_mask, spatial_shapes)
+ enc_outputs_class = self.bbox_head.cls_branches[
+ self.decoder.num_layers](
+ output_memory)
+ enc_outputs_coord_unact = self.bbox_head.reg_branches[
+ self.decoder.num_layers](output_memory) + output_proposals
+
+ # NOTE The DINO selects top-k proposals according to scores of
+ # multi-class classification, while DeformDETR, where the input
+ # is `enc_outputs_class[..., 0]` selects according to scores of
+ # binary classification.
+ topk_indices = torch.topk(
+ enc_outputs_class.max(-1)[0], k=self.num_queries, dim=1)[1]
+ topk_score = torch.gather(
+ enc_outputs_class, 1,
+ topk_indices.unsqueeze(-1).repeat(1, 1, cls_out_features))
+ topk_coords_unact = torch.gather(
+ enc_outputs_coord_unact, 1,
+ topk_indices.unsqueeze(-1).repeat(1, 1, 4))
+ topk_coords = topk_coords_unact.sigmoid()
+ topk_coords_unact = topk_coords_unact.detach()
+
+ query = self.query_embedding.weight[:, None, :]
+ query = query.repeat(1, bs, 1).transpose(0, 1)
+ if self.training:
+ dn_label_query, dn_bbox_query, dn_mask, dn_meta = \
+ self.dn_query_generator(batch_data_samples)
+ query = torch.cat([dn_label_query, query], dim=1)
+ reference_points = torch.cat([dn_bbox_query, topk_coords_unact],
+ dim=1)
+ else:
+ reference_points = topk_coords_unact
+ dn_mask, dn_meta = None, None
+ reference_points = reference_points.sigmoid()
+
+ decoder_inputs_dict = dict(
+ query=query,
+ memory=memory,
+ reference_points=reference_points,
+ dn_mask=dn_mask)
+ # NOTE DINO calculates encoder losses on scores and coordinates
+ # of selected top-k encoder queries, while DeformDETR is of all
+ # encoder queries.
+ head_inputs_dict = dict(
+ enc_outputs_class=topk_score,
+ enc_outputs_coord=topk_coords,
+ dn_meta=dn_meta) if self.training else dict()
+ return decoder_inputs_dict, head_inputs_dict
+
+ def forward_decoder(self,
+ query: Tensor,
+ memory: Tensor,
+ memory_mask: Tensor,
+ reference_points: Tensor,
+ spatial_shapes: Tensor,
+ level_start_index: Tensor,
+ valid_ratios: Tensor,
+ dn_mask: Optional[Tensor] = None,
+ **kwargs) -> Dict:
+ """Forward with Transformer decoder.
+
+ The forward procedure of the transformer is defined as:
+ 'pre_transformer' -> 'encoder' -> 'pre_decoder' -> 'decoder'
+ More details can be found at `TransformerDetector.forward_transformer`
+ in `mmdet/detector/base_detr.py`.
+
+ Args:
+ query (Tensor): The queries of decoder inputs, has shape
+ (bs, num_queries_total, dim), where `num_queries_total` is the
+ sum of `num_denoising_queries` and `num_matching_queries` when
+ `self.training` is `True`, else `num_matching_queries`.
+ memory (Tensor): The output embeddings of the Transformer encoder,
+ has shape (bs, num_feat_points, dim).
+ memory_mask (Tensor): ByteTensor, the padding mask of the memory,
+ has shape (bs, num_feat_points).
+ reference_points (Tensor): The initial reference, has shape
+ (bs, num_queries_total, 4) with the last dimension arranged as
+ (cx, cy, w, h).
+ spatial_shapes (Tensor): Spatial shapes of features in all levels,
+ has shape (num_levels, 2), last dimension represents (h, w).
+ level_start_index (Tensor): The start index of each level.
+ A tensor has shape (num_levels, ) and can be represented
+ as [0, h_0*w_0, h_0*w_0+h_1*w_1, ...].
+ valid_ratios (Tensor): The ratios of the valid width and the valid
+ height relative to the width and the height of features in all
+ levels, has shape (bs, num_levels, 2).
+ dn_mask (Tensor, optional): The attention mask to prevent
+ information leakage from different denoising groups and
+ matching parts, will be used as `self_attn_mask` of the
+ `self.decoder`, has shape (num_queries_total,
+ num_queries_total).
+ It is `None` when `self.training` is `False`.
+
+ Returns:
+ dict: The dictionary of decoder outputs, which includes the
+ `hidden_states` of the decoder output and `references` including
+ the initial and intermediate reference_points.
+ """
+ inter_states, references = self.decoder(
+ query=query,
+ value=memory,
+ key_padding_mask=memory_mask,
+ self_attn_mask=dn_mask,
+ reference_points=reference_points,
+ spatial_shapes=spatial_shapes,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios,
+ reg_branches=self.bbox_head.reg_branches,
+ **kwargs)
+
+ if len(query) == self.num_queries:
+ # NOTE: This is to make sure label_embeding can be involved to
+ # produce loss even if there is no denoising query (no ground truth
+ # target in this GPU), otherwise, this will raise runtime error in
+ # distributed training.
+ inter_states[0] += \
+ self.dn_query_generator.label_embedding.weight[0, 0] * 0.0
+
+ decoder_outputs_dict = dict(
+ hidden_states=inter_states, references=list(references))
+ return decoder_outputs_dict
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/fast_rcnn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/fast_rcnn.py
new file mode 100644
index 0000000000000000000000000000000000000000..5b39050fdc2989eb5c870704e1c1417987d53d46
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/fast_rcnn.py
@@ -0,0 +1,26 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .two_stage import TwoStageDetector
+
+
+@MODELS.register_module()
+class FastRCNN(TwoStageDetector):
+ """Implementation of `Fast R-CNN `_"""
+
+ def __init__(self,
+ backbone: ConfigType,
+ roi_head: ConfigType,
+ train_cfg: ConfigType,
+ test_cfg: ConfigType,
+ neck: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ roi_head=roi_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ init_cfg=init_cfg,
+ data_preprocessor=data_preprocessor)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/faster_rcnn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/faster_rcnn.py
new file mode 100644
index 0000000000000000000000000000000000000000..36109e3200a2d8e7d8a1032f7028e47a7699fb6a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/faster_rcnn.py
@@ -0,0 +1,28 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .two_stage import TwoStageDetector
+
+
+@MODELS.register_module()
+class FasterRCNN(TwoStageDetector):
+ """Implementation of `Faster R-CNN `_"""
+
+ def __init__(self,
+ backbone: ConfigType,
+ rpn_head: ConfigType,
+ roi_head: ConfigType,
+ train_cfg: ConfigType,
+ test_cfg: ConfigType,
+ neck: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ rpn_head=rpn_head,
+ roi_head=roi_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ init_cfg=init_cfg,
+ data_preprocessor=data_preprocessor)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/fcos.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/fcos.py
new file mode 100644
index 0000000000000000000000000000000000000000..c628059313ac80644ec2ba2c806e7baf2e418a41
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/fcos.py
@@ -0,0 +1,42 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class FCOS(SingleStageDetector):
+ """Implementation of `FCOS `_
+
+ Args:
+ backbone (:obj:`ConfigDict` or dict): The backbone config.
+ neck (:obj:`ConfigDict` or dict): The neck config.
+ bbox_head (:obj:`ConfigDict` or dict): The bbox head config.
+ train_cfg (:obj:`ConfigDict` or dict, optional): The training config
+ of FCOS. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): The testing config
+ of FCOS. Defaults to None.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional): Config of
+ :class:`DetDataPreprocessor` to process the input data.
+ Defaults to None.
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or
+ list[dict], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/fovea.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/fovea.py
new file mode 100644
index 0000000000000000000000000000000000000000..5e4f21caa239147e3b81e66280aa1da043715b42
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/fovea.py
@@ -0,0 +1,41 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class FOVEA(SingleStageDetector):
+ """Implementation of `FoveaBox `_
+ Args:
+ backbone (:obj:`ConfigDict` or dict): The backbone config.
+ neck (:obj:`ConfigDict` or dict): The neck config.
+ bbox_head (:obj:`ConfigDict` or dict): The bbox head config.
+ train_cfg (:obj:`ConfigDict` or dict, optional): The training config
+ of FOVEA. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): The testing config
+ of FOVEA. Defaults to None.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional): Config of
+ :class:`DetDataPreprocessor` to process the input data.
+ Defaults to None.
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or
+ list[dict], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/fsaf.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/fsaf.py
new file mode 100644
index 0000000000000000000000000000000000000000..01b40273341f2a85cfa427f8adfc945a1b7da58a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/fsaf.py
@@ -0,0 +1,26 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class FSAF(SingleStageDetector):
+ """Implementation of `FSAF `_"""
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None):
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/gfl.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/gfl.py
new file mode 100644
index 0000000000000000000000000000000000000000..c26821af68c224d4b55a1ca3d2be4c6e1d1b155d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/gfl.py
@@ -0,0 +1,41 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class GFL(SingleStageDetector):
+ """Implementation of `GFL `_
+
+ Args:
+ backbone (:obj:`ConfigDict` or dict): The backbone module.
+ neck (:obj:`ConfigDict` or dict): The neck module.
+ bbox_head (:obj:`ConfigDict` or dict): The bbox head module.
+ train_cfg (:obj:`ConfigDict` or dict, optional): The training config
+ of GFL. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): The testing config
+ of GFL. Defaults to None.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional): Config of
+ :class:`DetDataPreprocessor` to process the input data.
+ Defaults to None.
+ init_cfg (:obj:`ConfigDict` or dict, optional): the config to control
+ the initialization. Defaults to None.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/glip.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/glip.py
new file mode 100644
index 0000000000000000000000000000000000000000..45cfe7d39fd7b8d9e9bc37c49fe369ff87bc68d9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/glip.py
@@ -0,0 +1,590 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import re
+import warnings
+from typing import Optional, Tuple, Union
+
+import torch
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+def find_noun_phrases(caption: str) -> list:
+ """Find noun phrases in a caption using nltk.
+ Args:
+ caption (str): The caption to analyze.
+
+ Returns:
+ list: List of noun phrases found in the caption.
+
+ Examples:
+ >>> caption = 'There is two cat and a remote in the picture'
+ >>> find_noun_phrases(caption) # ['cat', 'a remote', 'the picture']
+ """
+ try:
+ import nltk
+ nltk.download('punkt', download_dir='~/nltk_data')
+ nltk.download('averaged_perceptron_tagger', download_dir='~/nltk_data')
+ except ImportError:
+ raise RuntimeError('nltk is not installed, please install it by: '
+ 'pip install nltk.')
+
+ caption = caption.lower()
+ tokens = nltk.word_tokenize(caption)
+ pos_tags = nltk.pos_tag(tokens)
+
+ grammar = 'NP: {?*+}'
+ cp = nltk.RegexpParser(grammar)
+ result = cp.parse(pos_tags)
+
+ noun_phrases = []
+ for subtree in result.subtrees():
+ if subtree.label() == 'NP':
+ noun_phrases.append(' '.join(t[0] for t in subtree.leaves()))
+
+ return noun_phrases
+
+
+def remove_punctuation(text: str) -> str:
+ """Remove punctuation from a text.
+ Args:
+ text (str): The input text.
+
+ Returns:
+ str: The text with punctuation removed.
+ """
+ punctuation = [
+ '|', ':', ';', '@', '(', ')', '[', ']', '{', '}', '^', '\'', '\"', '’',
+ '`', '?', '$', '%', '#', '!', '&', '*', '+', ',', '.'
+ ]
+ for p in punctuation:
+ text = text.replace(p, '')
+ return text.strip()
+
+
+def run_ner(caption: str) -> Tuple[list, list]:
+ """Run NER on a caption and return the tokens and noun phrases.
+ Args:
+ caption (str): The input caption.
+
+ Returns:
+ Tuple[List, List]: A tuple containing the tokens and noun phrases.
+ - tokens_positive (List): A list of token positions.
+ - noun_phrases (List): A list of noun phrases.
+ """
+ noun_phrases = find_noun_phrases(caption)
+ noun_phrases = [remove_punctuation(phrase) for phrase in noun_phrases]
+ noun_phrases = [phrase for phrase in noun_phrases if phrase != '']
+ print('noun_phrases:', noun_phrases)
+ relevant_phrases = noun_phrases
+ labels = noun_phrases
+
+ tokens_positive = []
+ for entity, label in zip(relevant_phrases, labels):
+ try:
+ # search all occurrences and mark them as different entities
+ # TODO: Not Robust
+ for m in re.finditer(entity, caption.lower()):
+ tokens_positive.append([[m.start(), m.end()]])
+ except Exception:
+ print('noun entities:', noun_phrases)
+ print('entity:', entity)
+ print('caption:', caption.lower())
+ return tokens_positive, noun_phrases
+
+
+def create_positive_map(tokenized,
+ tokens_positive: list,
+ max_num_entities: int = 256) -> Tensor:
+ """construct a map such that positive_map[i,j] = True
+ if box i is associated to token j
+
+ Args:
+ tokenized: The tokenized input.
+ tokens_positive (list): A list of token ranges
+ associated with positive boxes.
+ max_num_entities (int, optional): The maximum number of entities.
+ Defaults to 256.
+
+ Returns:
+ torch.Tensor: The positive map.
+
+ Raises:
+ Exception: If an error occurs during token-to-char mapping.
+ """
+ positive_map = torch.zeros((len(tokens_positive), max_num_entities),
+ dtype=torch.float)
+
+ for j, tok_list in enumerate(tokens_positive):
+ for (beg, end) in tok_list:
+ try:
+ beg_pos = tokenized.char_to_token(beg)
+ end_pos = tokenized.char_to_token(end - 1)
+ except Exception as e:
+ print('beg:', beg, 'end:', end)
+ print('token_positive:', tokens_positive)
+ raise e
+ if beg_pos is None:
+ try:
+ beg_pos = tokenized.char_to_token(beg + 1)
+ if beg_pos is None:
+ beg_pos = tokenized.char_to_token(beg + 2)
+ except Exception:
+ beg_pos = None
+ if end_pos is None:
+ try:
+ end_pos = tokenized.char_to_token(end - 2)
+ if end_pos is None:
+ end_pos = tokenized.char_to_token(end - 3)
+ except Exception:
+ end_pos = None
+ if beg_pos is None or end_pos is None:
+ continue
+
+ assert beg_pos is not None and end_pos is not None
+ positive_map[j, beg_pos:end_pos + 1].fill_(1)
+ return positive_map / (positive_map.sum(-1)[:, None] + 1e-6)
+
+
+def create_positive_map_label_to_token(positive_map: Tensor,
+ plus: int = 0) -> dict:
+ """Create a dictionary mapping the label to the token.
+ Args:
+ positive_map (Tensor): The positive map tensor.
+ plus (int, optional): Value added to the label for indexing.
+ Defaults to 0.
+
+ Returns:
+ dict: The dictionary mapping the label to the token.
+ """
+ positive_map_label_to_token = {}
+ for i in range(len(positive_map)):
+ positive_map_label_to_token[i + plus] = torch.nonzero(
+ positive_map[i], as_tuple=True)[0].tolist()
+ return positive_map_label_to_token
+
+
+def clean_label_name(name: str) -> str:
+ name = re.sub(r'\(.*\)', '', name)
+ name = re.sub(r'_', ' ', name)
+ name = re.sub(r' ', ' ', name)
+ return name
+
+
+def chunks(lst: list, n: int) -> list:
+ """Yield successive n-sized chunks from lst."""
+ all_ = []
+ for i in range(0, len(lst), n):
+ data_index = lst[i:i + n]
+ all_.append(data_index)
+ counter = 0
+ for i in all_:
+ counter += len(i)
+ assert (counter == len(lst))
+
+ return all_
+
+
+@MODELS.register_module()
+class GLIP(SingleStageDetector):
+ """Implementation of `GLIP `_
+ Args:
+ backbone (:obj:`ConfigDict` or dict): The backbone config.
+ neck (:obj:`ConfigDict` or dict): The neck config.
+ bbox_head (:obj:`ConfigDict` or dict): The bbox head config.
+ language_model (:obj:`ConfigDict` or dict): The language model config.
+ train_cfg (:obj:`ConfigDict` or dict, optional): The training config
+ of GLIP. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): The testing config
+ of GLIP. Defaults to None.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional): Config of
+ :class:`DetDataPreprocessor` to process the input data.
+ Defaults to None.
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or
+ list[dict], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ language_model: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
+ self.language_model = MODELS.build(language_model)
+
+ self._special_tokens = '. '
+
+ def to_enhance_text_prompts(self, original_caption, enhanced_text_prompts):
+ caption_string = ''
+ tokens_positive = []
+ for idx, word in enumerate(original_caption):
+ if word in enhanced_text_prompts:
+ enhanced_text_dict = enhanced_text_prompts[word]
+ if 'prefix' in enhanced_text_dict:
+ caption_string += enhanced_text_dict['prefix']
+ start_i = len(caption_string)
+ if 'name' in enhanced_text_dict:
+ caption_string += enhanced_text_dict['name']
+ else:
+ caption_string += word
+ end_i = len(caption_string)
+ tokens_positive.append([[start_i, end_i]])
+
+ if 'suffix' in enhanced_text_dict:
+ caption_string += enhanced_text_dict['suffix']
+ else:
+ tokens_positive.append(
+ [[len(caption_string),
+ len(caption_string) + len(word)]])
+ caption_string += word
+
+ if idx != len(original_caption) - 1:
+ caption_string += self._special_tokens
+ return caption_string, tokens_positive
+
+ def to_plain_text_prompts(self, original_caption):
+ caption_string = ''
+ tokens_positive = []
+ for idx, word in enumerate(original_caption):
+ tokens_positive.append(
+ [[len(caption_string),
+ len(caption_string) + len(word)]])
+ caption_string += word
+ if idx != len(original_caption) - 1:
+ caption_string += self._special_tokens
+ return caption_string, tokens_positive
+
+ def get_tokens_and_prompts(
+ self,
+ original_caption: Union[str, list, tuple],
+ custom_entities: bool = False,
+ enhanced_text_prompts: Optional[ConfigType] = None
+ ) -> Tuple[dict, str, list, list]:
+ """Get the tokens positive and prompts for the caption."""
+ if isinstance(original_caption, (list, tuple)) or custom_entities:
+ if custom_entities and isinstance(original_caption, str):
+ original_caption = original_caption.strip(self._special_tokens)
+ original_caption = original_caption.split(self._special_tokens)
+ original_caption = list(
+ filter(lambda x: len(x) > 0, original_caption))
+
+ original_caption = [clean_label_name(i) for i in original_caption]
+
+ if custom_entities and enhanced_text_prompts is not None:
+ caption_string, tokens_positive = self.to_enhance_text_prompts(
+ original_caption, enhanced_text_prompts)
+ else:
+ caption_string, tokens_positive = self.to_plain_text_prompts(
+ original_caption)
+
+ tokenized = self.language_model.tokenizer([caption_string],
+ return_tensors='pt')
+ entities = original_caption
+ else:
+ original_caption = original_caption.strip(self._special_tokens)
+ tokenized = self.language_model.tokenizer([original_caption],
+ return_tensors='pt')
+ tokens_positive, noun_phrases = run_ner(original_caption)
+ entities = noun_phrases
+ caption_string = original_caption
+
+ return tokenized, caption_string, tokens_positive, entities
+
+ def get_positive_map(self, tokenized, tokens_positive):
+ positive_map = create_positive_map(tokenized, tokens_positive)
+ positive_map_label_to_token = create_positive_map_label_to_token(
+ positive_map, plus=1)
+ return positive_map_label_to_token, positive_map
+
+ def get_tokens_positive_and_prompts(
+ self,
+ original_caption: Union[str, list, tuple],
+ custom_entities: bool = False,
+ enhanced_text_prompt: Optional[ConfigType] = None,
+ tokens_positive: Optional[list] = None,
+ ) -> Tuple[dict, str, Tensor, list]:
+ if tokens_positive is not None:
+ if tokens_positive == -1:
+ if not original_caption.endswith('.'):
+ original_caption = original_caption + self._special_tokens
+ return None, original_caption, None, original_caption
+ else:
+ if not original_caption.endswith('.'):
+ original_caption = original_caption + self._special_tokens
+ tokenized = self.language_model.tokenizer([original_caption],
+ return_tensors='pt')
+ positive_map_label_to_token, positive_map = \
+ self.get_positive_map(tokenized, tokens_positive)
+
+ entities = []
+ for token_positive in tokens_positive:
+ instance_entities = []
+ for t in token_positive:
+ instance_entities.append(original_caption[t[0]:t[1]])
+ entities.append(' / '.join(instance_entities))
+ return positive_map_label_to_token, original_caption, \
+ positive_map, entities
+
+ chunked_size = self.test_cfg.get('chunked_size', -1)
+ if not self.training and chunked_size > 0:
+ assert isinstance(original_caption,
+ (list, tuple)) or custom_entities is True
+ all_output = self.get_tokens_positive_and_prompts_chunked(
+ original_caption, enhanced_text_prompt)
+ positive_map_label_to_token, \
+ caption_string, \
+ positive_map, \
+ entities = all_output
+ else:
+ tokenized, caption_string, tokens_positive, entities = \
+ self.get_tokens_and_prompts(
+ original_caption, custom_entities, enhanced_text_prompt)
+ positive_map_label_to_token, positive_map = self.get_positive_map(
+ tokenized, tokens_positive)
+ if tokenized.input_ids.shape[1] > self.language_model.max_tokens:
+ warnings.warn('Inputting a text that is too long will result '
+ 'in poor prediction performance. '
+ 'Please reduce the text length.')
+ return positive_map_label_to_token, caption_string, \
+ positive_map, entities
+
+ def get_tokens_positive_and_prompts_chunked(
+ self,
+ original_caption: Union[list, tuple],
+ enhanced_text_prompts: Optional[ConfigType] = None):
+ chunked_size = self.test_cfg.get('chunked_size', -1)
+ original_caption = [clean_label_name(i) for i in original_caption]
+
+ original_caption_chunked = chunks(original_caption, chunked_size)
+ ids_chunked = chunks(
+ list(range(1,
+ len(original_caption) + 1)), chunked_size)
+
+ positive_map_label_to_token_chunked = []
+ caption_string_chunked = []
+ positive_map_chunked = []
+ entities_chunked = []
+
+ for i in range(len(ids_chunked)):
+ if enhanced_text_prompts is not None:
+ caption_string, tokens_positive = self.to_enhance_text_prompts(
+ original_caption_chunked[i], enhanced_text_prompts)
+ else:
+ caption_string, tokens_positive = self.to_plain_text_prompts(
+ original_caption_chunked[i])
+ tokenized = self.language_model.tokenizer([caption_string],
+ return_tensors='pt')
+ if tokenized.input_ids.shape[1] > self.language_model.max_tokens:
+ warnings.warn('Inputting a text that is too long will result '
+ 'in poor prediction performance. '
+ 'Please reduce the --chunked-size.')
+ positive_map_label_to_token, positive_map = self.get_positive_map(
+ tokenized, tokens_positive)
+
+ caption_string_chunked.append(caption_string)
+ positive_map_label_to_token_chunked.append(
+ positive_map_label_to_token)
+ positive_map_chunked.append(positive_map)
+ entities_chunked.append(original_caption_chunked[i])
+
+ return positive_map_label_to_token_chunked, \
+ caption_string_chunked, \
+ positive_map_chunked, \
+ entities_chunked
+
+ def loss(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> Union[dict, list]:
+ # TODO: Only open vocabulary tasks are supported for training now.
+ text_prompts = [
+ data_samples.text for data_samples in batch_data_samples
+ ]
+
+ gt_labels = [
+ data_samples.gt_instances.labels
+ for data_samples in batch_data_samples
+ ]
+
+ new_text_prompts = []
+ positive_maps = []
+ if len(set(text_prompts)) == 1:
+ # All the text prompts are the same,
+ # so there is no need to calculate them multiple times.
+ tokenized, caption_string, tokens_positive, _ = \
+ self.get_tokens_and_prompts(
+ text_prompts[0], True)
+ new_text_prompts = [caption_string] * len(batch_inputs)
+ for gt_label in gt_labels:
+ new_tokens_positive = [
+ tokens_positive[label] for label in gt_label
+ ]
+ _, positive_map = self.get_positive_map(
+ tokenized, new_tokens_positive)
+ positive_maps.append(positive_map)
+ else:
+ for text_prompt, gt_label in zip(text_prompts, gt_labels):
+ tokenized, caption_string, tokens_positive, _ = \
+ self.get_tokens_and_prompts(
+ text_prompt, True)
+ new_tokens_positive = [
+ tokens_positive[label] for label in gt_label
+ ]
+ _, positive_map = self.get_positive_map(
+ tokenized, new_tokens_positive)
+ positive_maps.append(positive_map)
+ new_text_prompts.append(caption_string)
+
+ language_dict_features = self.language_model(new_text_prompts)
+ for i, data_samples in enumerate(batch_data_samples):
+ # .bool().float() is very important
+ positive_map = positive_maps[i].to(
+ batch_inputs.device).bool().float()
+ data_samples.gt_instances.positive_maps = positive_map
+
+ visual_features = self.extract_feat(batch_inputs)
+
+ losses = self.bbox_head.loss(visual_features, language_dict_features,
+ batch_data_samples)
+ return losses
+
+ def predict(self,
+ batch_inputs: Tensor,
+ batch_data_samples: SampleList,
+ rescale: bool = True) -> SampleList:
+ """Predict results from a batch of inputs and data samples with post-
+ processing.
+
+ Args:
+ batch_inputs (Tensor): Inputs with shape (N, C, H, W).
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool): Whether to rescale the results.
+ Defaults to True.
+
+ Returns:
+ list[:obj:`DetDataSample`]: Detection results of the
+ input images. Each DetDataSample usually contain
+ 'pred_instances'. And the ``pred_instances`` usually
+ contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - label_names (List[str]): Label names of bboxes.
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ text_prompts = []
+ enhanced_text_prompts = []
+ tokens_positives = []
+ for data_samples in batch_data_samples:
+ text_prompts.append(data_samples.text)
+ if 'caption_prompt' in data_samples:
+ enhanced_text_prompts.append(data_samples.caption_prompt)
+ else:
+ enhanced_text_prompts.append(None)
+ tokens_positives.append(data_samples.get('tokens_positive', None))
+
+ if 'custom_entities' in batch_data_samples[0]:
+ # Assuming that the `custom_entities` flag
+ # inside a batch is always the same. For single image inference
+ custom_entities = batch_data_samples[0].custom_entities
+ else:
+ custom_entities = False
+
+ if len(set(text_prompts)) == 1:
+ # All the text prompts are the same,
+ # so there is no need to calculate them multiple times.
+ _positive_maps_and_prompts = [
+ self.get_tokens_positive_and_prompts(
+ text_prompts[0], custom_entities, enhanced_text_prompts[0],
+ tokens_positives[0])
+ ] * len(batch_inputs)
+ else:
+ _positive_maps_and_prompts = [
+ self.get_tokens_positive_and_prompts(text_prompt,
+ custom_entities,
+ enhanced_text_prompt,
+ tokens_positive)
+ for text_prompt, enhanced_text_prompt, tokens_positive in zip(
+ text_prompts, enhanced_text_prompts, tokens_positives)
+ ]
+
+ token_positive_maps, text_prompts, _, entities = zip(
+ *_positive_maps_and_prompts)
+
+ visual_features = self.extract_feat(batch_inputs)
+
+ if isinstance(text_prompts[0], list):
+ # chunked text prompts, only bs=1 is supported
+ assert len(batch_inputs) == 1
+ count = 0
+ results_list = []
+
+ entities = [[item for lst in entities[0] for item in lst]]
+
+ for b in range(len(text_prompts[0])):
+ text_prompts_once = [text_prompts[0][b]]
+ token_positive_maps_once = token_positive_maps[0][b]
+ language_dict_features = self.language_model(text_prompts_once)
+ batch_data_samples[
+ 0].token_positive_map = token_positive_maps_once
+
+ pred_instances = self.bbox_head.predict(
+ copy.deepcopy(visual_features),
+ language_dict_features,
+ batch_data_samples,
+ rescale=rescale)[0]
+
+ if len(pred_instances) > 0:
+ pred_instances.labels += count
+ count += len(token_positive_maps_once)
+ results_list.append(pred_instances)
+ results_list = [results_list[0].cat(results_list)]
+ else:
+ language_dict_features = self.language_model(list(text_prompts))
+
+ for i, data_samples in enumerate(batch_data_samples):
+ data_samples.token_positive_map = token_positive_maps[i]
+
+ results_list = self.bbox_head.predict(
+ visual_features,
+ language_dict_features,
+ batch_data_samples,
+ rescale=rescale)
+
+ for data_sample, pred_instances, entity in zip(batch_data_samples,
+ results_list, entities):
+ if len(pred_instances) > 0:
+ label_names = []
+ for labels in pred_instances.labels:
+ if labels >= len(entity):
+ warnings.warn(
+ 'The unexpected output indicates an issue with '
+ 'named entity recognition. You can try '
+ 'setting custom_entities=True and running '
+ 'again to see if it helps.')
+ label_names.append('unobject')
+ else:
+ label_names.append(entity[labels])
+ # for visualization
+ pred_instances.label_names = label_names
+ data_sample.pred_instances = pred_instances
+ return batch_data_samples
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/grid_rcnn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/grid_rcnn.py
new file mode 100644
index 0000000000000000000000000000000000000000..7bcb5b033edc620f1cf61b986c345961b719e6f1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/grid_rcnn.py
@@ -0,0 +1,33 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .two_stage import TwoStageDetector
+
+
+@MODELS.register_module()
+class GridRCNN(TwoStageDetector):
+ """Grid R-CNN.
+
+ This detector is the implementation of:
+ - Grid R-CNN (https://arxiv.org/abs/1811.12030)
+ - Grid R-CNN Plus: Faster and Better (https://arxiv.org/abs/1906.05688)
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ rpn_head: ConfigType,
+ roi_head: ConfigType,
+ train_cfg: ConfigType,
+ test_cfg: ConfigType,
+ neck: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ rpn_head=rpn_head,
+ roi_head=roi_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/grounding_dino.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/grounding_dino.py
new file mode 100644
index 0000000000000000000000000000000000000000..b1ab7c2da16453e4aa43020681811a8b24767ad0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/grounding_dino.py
@@ -0,0 +1,621 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import re
+import warnings
+from typing import Dict, Optional, Tuple, Union
+
+import torch
+import torch.nn as nn
+from mmengine.runner.amp import autocast
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import OptSampleList, SampleList
+from mmdet.utils import ConfigType
+from ..layers import SinePositionalEncoding
+from ..layers.transformer.grounding_dino_layers import (
+ GroundingDinoTransformerDecoder, GroundingDinoTransformerEncoder)
+from .dino import DINO
+from .glip import (create_positive_map, create_positive_map_label_to_token,
+ run_ner)
+
+
+def clean_label_name(name: str) -> str:
+ name = re.sub(r'\(.*\)', '', name)
+ name = re.sub(r'_', ' ', name)
+ name = re.sub(r' ', ' ', name)
+ return name
+
+
+def chunks(lst: list, n: int) -> list:
+ """Yield successive n-sized chunks from lst."""
+ all_ = []
+ for i in range(0, len(lst), n):
+ data_index = lst[i:i + n]
+ all_.append(data_index)
+ counter = 0
+ for i in all_:
+ counter += len(i)
+ assert (counter == len(lst))
+
+ return all_
+
+
+@MODELS.register_module()
+class GroundingDINO(DINO):
+ """Implementation of `Grounding DINO: Marrying DINO with Grounded Pre-
+ Training for Open-Set Object Detection.
+
+ `_
+
+ Code is modified from the `official github repo
+ `_.
+ """
+
+ def __init__(self,
+ language_model,
+ *args,
+ use_autocast=False,
+ **kwargs) -> None:
+
+ self.language_model_cfg = language_model
+ self._special_tokens = '. '
+ self.use_autocast = use_autocast
+ super().__init__(*args, **kwargs)
+
+ def _init_layers(self) -> None:
+ """Initialize layers except for backbone, neck and bbox_head."""
+ self.positional_encoding = SinePositionalEncoding(
+ **self.positional_encoding)
+ self.encoder = GroundingDinoTransformerEncoder(**self.encoder)
+ self.decoder = GroundingDinoTransformerDecoder(**self.decoder)
+ self.embed_dims = self.encoder.embed_dims
+ self.query_embedding = nn.Embedding(self.num_queries, self.embed_dims)
+ num_feats = self.positional_encoding.num_feats
+ assert num_feats * 2 == self.embed_dims, \
+ f'embed_dims should be exactly 2 times of num_feats. ' \
+ f'Found {self.embed_dims} and {num_feats}.'
+
+ self.level_embed = nn.Parameter(
+ torch.Tensor(self.num_feature_levels, self.embed_dims))
+ self.memory_trans_fc = nn.Linear(self.embed_dims, self.embed_dims)
+ self.memory_trans_norm = nn.LayerNorm(self.embed_dims)
+
+ # text modules
+ self.language_model = MODELS.build(self.language_model_cfg)
+ self.text_feat_map = nn.Linear(
+ self.language_model.language_backbone.body.language_dim,
+ self.embed_dims,
+ bias=True)
+
+ def init_weights(self) -> None:
+ """Initialize weights for Transformer and other components."""
+ super().init_weights()
+ nn.init.constant_(self.text_feat_map.bias.data, 0)
+ nn.init.xavier_uniform_(self.text_feat_map.weight.data)
+
+ def to_enhance_text_prompts(self, original_caption, enhanced_text_prompts):
+ caption_string = ''
+ tokens_positive = []
+ for idx, word in enumerate(original_caption):
+ if word in enhanced_text_prompts:
+ enhanced_text_dict = enhanced_text_prompts[word]
+ if 'prefix' in enhanced_text_dict:
+ caption_string += enhanced_text_dict['prefix']
+ start_i = len(caption_string)
+ if 'name' in enhanced_text_dict:
+ caption_string += enhanced_text_dict['name']
+ else:
+ caption_string += word
+ end_i = len(caption_string)
+ tokens_positive.append([[start_i, end_i]])
+
+ if 'suffix' in enhanced_text_dict:
+ caption_string += enhanced_text_dict['suffix']
+ else:
+ tokens_positive.append(
+ [[len(caption_string),
+ len(caption_string) + len(word)]])
+ caption_string += word
+ caption_string += self._special_tokens
+ return caption_string, tokens_positive
+
+ def to_plain_text_prompts(self, original_caption):
+ caption_string = ''
+ tokens_positive = []
+ for idx, word in enumerate(original_caption):
+ tokens_positive.append(
+ [[len(caption_string),
+ len(caption_string) + len(word)]])
+ caption_string += word
+ caption_string += self._special_tokens
+ return caption_string, tokens_positive
+
+ def get_tokens_and_prompts(
+ self,
+ original_caption: Union[str, list, tuple],
+ custom_entities: bool = False,
+ enhanced_text_prompts: Optional[ConfigType] = None
+ ) -> Tuple[dict, str, list]:
+ """Get the tokens positive and prompts for the caption."""
+ if isinstance(original_caption, (list, tuple)) or custom_entities:
+ if custom_entities and isinstance(original_caption, str):
+ original_caption = original_caption.strip(self._special_tokens)
+ original_caption = original_caption.split(self._special_tokens)
+ original_caption = list(
+ filter(lambda x: len(x) > 0, original_caption))
+
+ original_caption = [clean_label_name(i) for i in original_caption]
+
+ if custom_entities and enhanced_text_prompts is not None:
+ caption_string, tokens_positive = self.to_enhance_text_prompts(
+ original_caption, enhanced_text_prompts)
+ else:
+ caption_string, tokens_positive = self.to_plain_text_prompts(
+ original_caption)
+
+ # NOTE: Tokenizer in Grounding DINO is different from
+ # that in GLIP. The tokenizer in GLIP will pad the
+ # caption_string to max_length, while the tokenizer
+ # in Grounding DINO will not.
+ tokenized = self.language_model.tokenizer(
+ [caption_string],
+ padding='max_length'
+ if self.language_model.pad_to_max else 'longest',
+ return_tensors='pt')
+ entities = original_caption
+ else:
+ if not original_caption.endswith('.'):
+ original_caption = original_caption + self._special_tokens
+ # NOTE: Tokenizer in Grounding DINO is different from
+ # that in GLIP. The tokenizer in GLIP will pad the
+ # caption_string to max_length, while the tokenizer
+ # in Grounding DINO will not.
+ tokenized = self.language_model.tokenizer(
+ [original_caption],
+ padding='max_length'
+ if self.language_model.pad_to_max else 'longest',
+ return_tensors='pt')
+ tokens_positive, noun_phrases = run_ner(original_caption)
+ entities = noun_phrases
+ caption_string = original_caption
+
+ return tokenized, caption_string, tokens_positive, entities
+
+ def get_positive_map(self, tokenized, tokens_positive):
+ positive_map = create_positive_map(
+ tokenized,
+ tokens_positive,
+ max_num_entities=self.bbox_head.cls_branches[
+ self.decoder.num_layers].max_text_len)
+ positive_map_label_to_token = create_positive_map_label_to_token(
+ positive_map, plus=1)
+ return positive_map_label_to_token, positive_map
+
+ def get_tokens_positive_and_prompts(
+ self,
+ original_caption: Union[str, list, tuple],
+ custom_entities: bool = False,
+ enhanced_text_prompt: Optional[ConfigType] = None,
+ tokens_positive: Optional[list] = None,
+ ) -> Tuple[dict, str, Tensor, list]:
+ """Get the tokens positive and prompts for the caption.
+
+ Args:
+ original_caption (str): The original caption, e.g. 'bench . car .'
+ custom_entities (bool, optional): Whether to use custom entities.
+ If ``True``, the ``original_caption`` should be a list of
+ strings, each of which is a word. Defaults to False.
+
+ Returns:
+ Tuple[dict, str, dict, str]: The dict is a mapping from each entity
+ id, which is numbered from 1, to its positive token id.
+ The str represents the prompts.
+ """
+ if tokens_positive is not None:
+ if tokens_positive == -1:
+ if not original_caption.endswith('.'):
+ original_caption = original_caption + self._special_tokens
+ return None, original_caption, None, original_caption
+ else:
+ if not original_caption.endswith('.'):
+ original_caption = original_caption + self._special_tokens
+ tokenized = self.language_model.tokenizer(
+ [original_caption],
+ padding='max_length'
+ if self.language_model.pad_to_max else 'longest',
+ return_tensors='pt')
+ positive_map_label_to_token, positive_map = \
+ self.get_positive_map(tokenized, tokens_positive)
+
+ entities = []
+ for token_positive in tokens_positive:
+ instance_entities = []
+ for t in token_positive:
+ instance_entities.append(original_caption[t[0]:t[1]])
+ entities.append(' / '.join(instance_entities))
+ return positive_map_label_to_token, original_caption, \
+ positive_map, entities
+
+ chunked_size = self.test_cfg.get('chunked_size', -1)
+ if not self.training and chunked_size > 0:
+ assert isinstance(original_caption,
+ (list, tuple)) or custom_entities is True
+ all_output = self.get_tokens_positive_and_prompts_chunked(
+ original_caption, enhanced_text_prompt)
+ positive_map_label_to_token, \
+ caption_string, \
+ positive_map, \
+ entities = all_output
+ else:
+ tokenized, caption_string, tokens_positive, entities = \
+ self.get_tokens_and_prompts(
+ original_caption, custom_entities, enhanced_text_prompt)
+ positive_map_label_to_token, positive_map = self.get_positive_map(
+ tokenized, tokens_positive)
+ return positive_map_label_to_token, caption_string, \
+ positive_map, entities
+
+ def get_tokens_positive_and_prompts_chunked(
+ self,
+ original_caption: Union[list, tuple],
+ enhanced_text_prompts: Optional[ConfigType] = None):
+ chunked_size = self.test_cfg.get('chunked_size', -1)
+ original_caption = [clean_label_name(i) for i in original_caption]
+
+ original_caption_chunked = chunks(original_caption, chunked_size)
+ ids_chunked = chunks(
+ list(range(1,
+ len(original_caption) + 1)), chunked_size)
+
+ positive_map_label_to_token_chunked = []
+ caption_string_chunked = []
+ positive_map_chunked = []
+ entities_chunked = []
+
+ for i in range(len(ids_chunked)):
+ if enhanced_text_prompts is not None:
+ caption_string, tokens_positive = self.to_enhance_text_prompts(
+ original_caption_chunked[i], enhanced_text_prompts)
+ else:
+ caption_string, tokens_positive = self.to_plain_text_prompts(
+ original_caption_chunked[i])
+ tokenized = self.language_model.tokenizer([caption_string],
+ return_tensors='pt')
+ if tokenized.input_ids.shape[1] > self.language_model.max_tokens:
+ warnings.warn('Inputting a text that is too long will result '
+ 'in poor prediction performance. '
+ 'Please reduce the --chunked-size.')
+ positive_map_label_to_token, positive_map = self.get_positive_map(
+ tokenized, tokens_positive)
+
+ caption_string_chunked.append(caption_string)
+ positive_map_label_to_token_chunked.append(
+ positive_map_label_to_token)
+ positive_map_chunked.append(positive_map)
+ entities_chunked.append(original_caption_chunked[i])
+
+ return positive_map_label_to_token_chunked, \
+ caption_string_chunked, \
+ positive_map_chunked, \
+ entities_chunked
+
+ def forward_transformer(
+ self,
+ img_feats: Tuple[Tensor],
+ text_dict: Dict,
+ batch_data_samples: OptSampleList = None,
+ ) -> Dict:
+ encoder_inputs_dict, decoder_inputs_dict = self.pre_transformer(
+ img_feats, batch_data_samples)
+
+ encoder_outputs_dict = self.forward_encoder(
+ **encoder_inputs_dict, text_dict=text_dict)
+
+ tmp_dec_in, head_inputs_dict = self.pre_decoder(
+ **encoder_outputs_dict, batch_data_samples=batch_data_samples)
+ decoder_inputs_dict.update(tmp_dec_in)
+
+ decoder_outputs_dict = self.forward_decoder(**decoder_inputs_dict)
+ head_inputs_dict.update(decoder_outputs_dict)
+ return head_inputs_dict
+
+ def forward_encoder(self, feat: Tensor, feat_mask: Tensor,
+ feat_pos: Tensor, spatial_shapes: Tensor,
+ level_start_index: Tensor, valid_ratios: Tensor,
+ text_dict: Dict) -> Dict:
+ text_token_mask = text_dict['text_token_mask']
+ memory, memory_text = self.encoder(
+ query=feat,
+ query_pos=feat_pos,
+ key_padding_mask=feat_mask, # for self_attn
+ spatial_shapes=spatial_shapes,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios,
+ # for text encoder
+ memory_text=text_dict['embedded'],
+ text_attention_mask=~text_token_mask,
+ position_ids=text_dict['position_ids'],
+ text_self_attention_masks=text_dict['masks'])
+ encoder_outputs_dict = dict(
+ memory=memory,
+ memory_mask=feat_mask,
+ spatial_shapes=spatial_shapes,
+ memory_text=memory_text,
+ text_token_mask=text_token_mask)
+ return encoder_outputs_dict
+
+ def pre_decoder(
+ self,
+ memory: Tensor,
+ memory_mask: Tensor,
+ spatial_shapes: Tensor,
+ memory_text: Tensor,
+ text_token_mask: Tensor,
+ batch_data_samples: OptSampleList = None,
+ ) -> Tuple[Dict]:
+ bs, _, c = memory.shape
+
+ output_memory, output_proposals = self.gen_encoder_output_proposals(
+ memory, memory_mask, spatial_shapes)
+
+ enc_outputs_class = self.bbox_head.cls_branches[
+ self.decoder.num_layers](output_memory, memory_text,
+ text_token_mask)
+ cls_out_features = self.bbox_head.cls_branches[
+ self.decoder.num_layers].max_text_len
+ enc_outputs_coord_unact = self.bbox_head.reg_branches[
+ self.decoder.num_layers](output_memory) + output_proposals
+
+ # NOTE The DINO selects top-k proposals according to scores of
+ # multi-class classification, while DeformDETR, where the input
+ # is `enc_outputs_class[..., 0]` selects according to scores of
+ # binary classification.
+ topk_indices = torch.topk(
+ enc_outputs_class.max(-1)[0], k=self.num_queries, dim=1)[1]
+
+ topk_score = torch.gather(
+ enc_outputs_class, 1,
+ topk_indices.unsqueeze(-1).repeat(1, 1, cls_out_features))
+ topk_coords_unact = torch.gather(
+ enc_outputs_coord_unact, 1,
+ topk_indices.unsqueeze(-1).repeat(1, 1, 4))
+ topk_coords = topk_coords_unact.sigmoid()
+ topk_coords_unact = topk_coords_unact.detach()
+
+ query = self.query_embedding.weight[:, None, :]
+ query = query.repeat(1, bs, 1).transpose(0, 1)
+ if self.training:
+ dn_label_query, dn_bbox_query, dn_mask, dn_meta = \
+ self.dn_query_generator(batch_data_samples)
+ query = torch.cat([dn_label_query, query], dim=1)
+ reference_points = torch.cat([dn_bbox_query, topk_coords_unact],
+ dim=1)
+ else:
+ reference_points = topk_coords_unact
+ dn_mask, dn_meta = None, None
+ reference_points = reference_points.sigmoid()
+
+ decoder_inputs_dict = dict(
+ query=query,
+ memory=memory,
+ reference_points=reference_points,
+ dn_mask=dn_mask,
+ memory_text=memory_text,
+ text_attention_mask=~text_token_mask,
+ )
+ # NOTE DINO calculates encoder losses on scores and coordinates
+ # of selected top-k encoder queries, while DeformDETR is of all
+ # encoder queries.
+ head_inputs_dict = dict(
+ enc_outputs_class=topk_score,
+ enc_outputs_coord=topk_coords,
+ dn_meta=dn_meta) if self.training else dict()
+ # append text_feats to head_inputs_dict
+ head_inputs_dict['memory_text'] = memory_text
+ head_inputs_dict['text_token_mask'] = text_token_mask
+ return decoder_inputs_dict, head_inputs_dict
+
+ def loss(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> Union[dict, list]:
+ text_prompts = [
+ data_samples.text for data_samples in batch_data_samples
+ ]
+
+ gt_labels = [
+ data_samples.gt_instances.labels
+ for data_samples in batch_data_samples
+ ]
+
+ if 'tokens_positive' in batch_data_samples[0]:
+ tokens_positive = [
+ data_samples.tokens_positive
+ for data_samples in batch_data_samples
+ ]
+ positive_maps = []
+ for token_positive, text_prompt, gt_label in zip(
+ tokens_positive, text_prompts, gt_labels):
+ tokenized = self.language_model.tokenizer(
+ [text_prompt],
+ padding='max_length'
+ if self.language_model.pad_to_max else 'longest',
+ return_tensors='pt')
+ new_tokens_positive = [
+ token_positive[label.item()] for label in gt_label
+ ]
+ _, positive_map = self.get_positive_map(
+ tokenized, new_tokens_positive)
+ positive_maps.append(positive_map)
+ new_text_prompts = text_prompts
+ else:
+ new_text_prompts = []
+ positive_maps = []
+ if len(set(text_prompts)) == 1:
+ # All the text prompts are the same,
+ # so there is no need to calculate them multiple times.
+ tokenized, caption_string, tokens_positive, _ = \
+ self.get_tokens_and_prompts(
+ text_prompts[0], True)
+ new_text_prompts = [caption_string] * len(batch_inputs)
+ for gt_label in gt_labels:
+ new_tokens_positive = [
+ tokens_positive[label] for label in gt_label
+ ]
+ _, positive_map = self.get_positive_map(
+ tokenized, new_tokens_positive)
+ positive_maps.append(positive_map)
+ else:
+ for text_prompt, gt_label in zip(text_prompts, gt_labels):
+ tokenized, caption_string, tokens_positive, _ = \
+ self.get_tokens_and_prompts(
+ text_prompt, True)
+ new_tokens_positive = [
+ tokens_positive[label] for label in gt_label
+ ]
+ _, positive_map = self.get_positive_map(
+ tokenized, new_tokens_positive)
+ positive_maps.append(positive_map)
+ new_text_prompts.append(caption_string)
+
+ text_dict = self.language_model(new_text_prompts)
+ if self.text_feat_map is not None:
+ text_dict['embedded'] = self.text_feat_map(text_dict['embedded'])
+
+ for i, data_samples in enumerate(batch_data_samples):
+ positive_map = positive_maps[i].to(
+ batch_inputs.device).bool().float()
+ text_token_mask = text_dict['text_token_mask'][i]
+ data_samples.gt_instances.positive_maps = positive_map
+ data_samples.gt_instances.text_token_mask = \
+ text_token_mask.unsqueeze(0).repeat(
+ len(positive_map), 1)
+ if self.use_autocast:
+ with autocast(enabled=True):
+ visual_features = self.extract_feat(batch_inputs)
+ else:
+ visual_features = self.extract_feat(batch_inputs)
+ head_inputs_dict = self.forward_transformer(visual_features, text_dict,
+ batch_data_samples)
+
+ losses = self.bbox_head.loss(
+ **head_inputs_dict, batch_data_samples=batch_data_samples)
+ return losses
+
+ def predict(self, batch_inputs, batch_data_samples, rescale: bool = True):
+ text_prompts = []
+ enhanced_text_prompts = []
+ tokens_positives = []
+ for data_samples in batch_data_samples:
+ text_prompts.append(data_samples.text)
+ if 'caption_prompt' in data_samples:
+ enhanced_text_prompts.append(data_samples.caption_prompt)
+ else:
+ enhanced_text_prompts.append(None)
+ tokens_positives.append(data_samples.get('tokens_positive', None))
+
+ if 'custom_entities' in batch_data_samples[0]:
+ # Assuming that the `custom_entities` flag
+ # inside a batch is always the same. For single image inference
+ custom_entities = batch_data_samples[0].custom_entities
+ else:
+ custom_entities = False
+ if len(text_prompts) == 1:
+ # All the text prompts are the same,
+ # so there is no need to calculate them multiple times.
+ _positive_maps_and_prompts = [
+ self.get_tokens_positive_and_prompts(
+ text_prompts[0], custom_entities, enhanced_text_prompts[0],
+ tokens_positives[0])
+ ] * len(batch_inputs)
+ else:
+ _positive_maps_and_prompts = [
+ self.get_tokens_positive_and_prompts(text_prompt,
+ custom_entities,
+ enhanced_text_prompt,
+ tokens_positive)
+ for text_prompt, enhanced_text_prompt, tokens_positive in zip(
+ text_prompts, enhanced_text_prompts, tokens_positives)
+ ]
+ token_positive_maps, text_prompts, _, entities = zip(
+ *_positive_maps_and_prompts)
+
+ # image feature extraction
+ visual_feats = self.extract_feat(batch_inputs)
+
+ if isinstance(text_prompts[0], list):
+ # chunked text prompts, only bs=1 is supported
+ assert len(batch_inputs) == 1
+ count = 0
+ results_list = []
+
+ entities = [[item for lst in entities[0] for item in lst]]
+
+ for b in range(len(text_prompts[0])):
+ text_prompts_once = [text_prompts[0][b]]
+ token_positive_maps_once = token_positive_maps[0][b]
+ text_dict = self.language_model(text_prompts_once)
+ # text feature map layer
+ if self.text_feat_map is not None:
+ text_dict['embedded'] = self.text_feat_map(
+ text_dict['embedded'])
+
+ batch_data_samples[
+ 0].token_positive_map = token_positive_maps_once
+
+ head_inputs_dict = self.forward_transformer(
+ copy.deepcopy(visual_feats), text_dict, batch_data_samples)
+ pred_instances = self.bbox_head.predict(
+ **head_inputs_dict,
+ rescale=rescale,
+ batch_data_samples=batch_data_samples)[0]
+
+ if len(pred_instances) > 0:
+ pred_instances.labels += count
+ count += len(token_positive_maps_once)
+ results_list.append(pred_instances)
+ results_list = [results_list[0].cat(results_list)]
+ is_rec_tasks = [False] * len(results_list)
+ else:
+ # extract text feats
+ text_dict = self.language_model(list(text_prompts))
+ # text feature map layer
+ if self.text_feat_map is not None:
+ text_dict['embedded'] = self.text_feat_map(
+ text_dict['embedded'])
+
+ is_rec_tasks = []
+ for i, data_samples in enumerate(batch_data_samples):
+ if token_positive_maps[i] is not None:
+ is_rec_tasks.append(False)
+ else:
+ is_rec_tasks.append(True)
+ data_samples.token_positive_map = token_positive_maps[i]
+
+ head_inputs_dict = self.forward_transformer(
+ visual_feats, text_dict, batch_data_samples)
+ results_list = self.bbox_head.predict(
+ **head_inputs_dict,
+ rescale=rescale,
+ batch_data_samples=batch_data_samples)
+
+ for data_sample, pred_instances, entity, is_rec_task in zip(
+ batch_data_samples, results_list, entities, is_rec_tasks):
+ if len(pred_instances) > 0:
+ label_names = []
+ for labels in pred_instances.labels:
+ if is_rec_task:
+ label_names.append(entity)
+ continue
+ if labels >= len(entity):
+ warnings.warn(
+ 'The unexpected output indicates an issue with '
+ 'named entity recognition. You can try '
+ 'setting custom_entities=True and running '
+ 'again to see if it helps.')
+ label_names.append('unobject')
+ else:
+ label_names.append(entity[labels])
+ # for visualization
+ pred_instances.label_names = label_names
+ data_sample.pred_instances = pred_instances
+ return batch_data_samples
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/htc.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/htc.py
new file mode 100644
index 0000000000000000000000000000000000000000..22a2aa889a59fd0e0afeb95a7369028def6e4fa9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/htc.py
@@ -0,0 +1,16 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from .cascade_rcnn import CascadeRCNN
+
+
+@MODELS.register_module()
+class HybridTaskCascade(CascadeRCNN):
+ """Implementation of `HTC `_"""
+
+ def __init__(self, **kwargs) -> None:
+ super().__init__(**kwargs)
+
+ @property
+ def with_semantic(self) -> bool:
+ """bool: whether the detector has a semantic head"""
+ return self.roi_head.with_semantic
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/kd_one_stage.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/kd_one_stage.py
new file mode 100644
index 0000000000000000000000000000000000000000..8a4a1bb564c0f6e4cabe32a5c01cfea252ecfb7d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/kd_one_stage.py
@@ -0,0 +1,122 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from pathlib import Path
+from typing import Any, Optional, Union
+
+import torch
+import torch.nn as nn
+from mmengine.config import Config
+from mmengine.runner import load_checkpoint
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.utils import ConfigType, OptConfigType
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class KnowledgeDistillationSingleStageDetector(SingleStageDetector):
+ r"""Implementation of `Distilling the Knowledge in a Neural Network.
+ `_.
+
+ Args:
+ backbone (:obj:`ConfigDict` or dict): The backbone module.
+ neck (:obj:`ConfigDict` or dict): The neck module.
+ bbox_head (:obj:`ConfigDict` or dict): The bbox head module.
+ teacher_config (:obj:`ConfigDict` | dict | str | Path): Config file
+ path or the config object of teacher model.
+ teacher_ckpt (str, optional): Checkpoint path of teacher model.
+ If left as None, the model will not load any weights.
+ Defaults to True.
+ eval_teacher (bool): Set the train mode for teacher.
+ Defaults to True.
+ train_cfg (:obj:`ConfigDict` or dict, optional): The training config
+ of ATSS. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): The testing config
+ of ATSS. Defaults to None.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional): Config of
+ :class:`DetDataPreprocessor` to process the input data.
+ Defaults to None.
+ """
+
+ def __init__(
+ self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ teacher_config: Union[ConfigType, str, Path],
+ teacher_ckpt: Optional[str] = None,
+ eval_teacher: bool = True,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ ) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor)
+ self.eval_teacher = eval_teacher
+ # Build teacher model
+ if isinstance(teacher_config, (str, Path)):
+ teacher_config = Config.fromfile(teacher_config)
+ self.teacher_model = MODELS.build(teacher_config['model'])
+ if teacher_ckpt is not None:
+ load_checkpoint(
+ self.teacher_model, teacher_ckpt, map_location='cpu')
+
+ def loss(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> dict:
+ """
+ Args:
+ batch_inputs (Tensor): Input images of shape (N, C, H, W).
+ These should usually be mean centered and std scaled.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ x = self.extract_feat(batch_inputs)
+ with torch.no_grad():
+ teacher_x = self.teacher_model.extract_feat(batch_inputs)
+ out_teacher = self.teacher_model.bbox_head(teacher_x)
+ losses = self.bbox_head.loss(x, out_teacher, batch_data_samples)
+ return losses
+
+ def cuda(self, device: Optional[str] = None) -> nn.Module:
+ """Since teacher_model is registered as a plain object, it is necessary
+ to put the teacher model to cuda when calling ``cuda`` function."""
+ self.teacher_model.cuda(device=device)
+ return super().cuda(device=device)
+
+ def to(self, device: Optional[str] = None) -> nn.Module:
+ """Since teacher_model is registered as a plain object, it is necessary
+ to put the teacher model to other device when calling ``to``
+ function."""
+ self.teacher_model.to(device=device)
+ return super().to(device=device)
+
+ def train(self, mode: bool = True) -> None:
+ """Set the same train mode for teacher and student model."""
+ if self.eval_teacher:
+ self.teacher_model.train(False)
+ else:
+ self.teacher_model.train(mode)
+ super().train(mode)
+
+ def __setattr__(self, name: str, value: Any) -> None:
+ """Set attribute, i.e. self.name = value
+
+ This reloading prevent the teacher model from being registered as a
+ nn.Module. The teacher module is registered as a plain object, so that
+ the teacher parameters will not show up when calling
+ ``self.parameters``, ``self.modules``, ``self.children`` methods.
+ """
+ if name == 'teacher_model':
+ object.__setattr__(self, name, value)
+ else:
+ super().__setattr__(name, value)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/lad.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/lad.py
new file mode 100644
index 0000000000000000000000000000000000000000..008f898772988715c67783d9218ff39c4dd95d80
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/lad.py
@@ -0,0 +1,93 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional
+
+import torch
+import torch.nn as nn
+from mmengine.runner import load_checkpoint
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.utils import ConfigType, OptConfigType
+from ..utils.misc import unpack_gt_instances
+from .kd_one_stage import KnowledgeDistillationSingleStageDetector
+
+
+@MODELS.register_module()
+class LAD(KnowledgeDistillationSingleStageDetector):
+ """Implementation of `LAD `_."""
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ teacher_backbone: ConfigType,
+ teacher_neck: ConfigType,
+ teacher_bbox_head: ConfigType,
+ teacher_ckpt: Optional[str] = None,
+ eval_teacher: bool = True,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None) -> None:
+ super(KnowledgeDistillationSingleStageDetector, self).__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor)
+ self.eval_teacher = eval_teacher
+ self.teacher_model = nn.Module()
+ self.teacher_model.backbone = MODELS.build(teacher_backbone)
+ if teacher_neck is not None:
+ self.teacher_model.neck = MODELS.build(teacher_neck)
+ teacher_bbox_head.update(train_cfg=train_cfg)
+ teacher_bbox_head.update(test_cfg=test_cfg)
+ self.teacher_model.bbox_head = MODELS.build(teacher_bbox_head)
+ if teacher_ckpt is not None:
+ load_checkpoint(
+ self.teacher_model, teacher_ckpt, map_location='cpu')
+
+ @property
+ def with_teacher_neck(self) -> bool:
+ """bool: whether the detector has a teacher_neck"""
+ return hasattr(self.teacher_model, 'neck') and \
+ self.teacher_model.neck is not None
+
+ def extract_teacher_feat(self, batch_inputs: Tensor) -> Tensor:
+ """Directly extract teacher features from the backbone+neck."""
+ x = self.teacher_model.backbone(batch_inputs)
+ if self.with_teacher_neck:
+ x = self.teacher_model.neck(x)
+ return x
+
+ def loss(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> dict:
+ """
+ Args:
+ batch_inputs (Tensor): Input images of shape (N, C, H, W).
+ These should usually be mean centered and std scaled.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ outputs = unpack_gt_instances(batch_data_samples)
+ batch_gt_instances, batch_gt_instances_ignore, batch_img_metas \
+ = outputs
+ # get label assignment from the teacher
+ with torch.no_grad():
+ x_teacher = self.extract_teacher_feat(batch_inputs)
+ outs_teacher = self.teacher_model.bbox_head(x_teacher)
+ label_assignment_results = \
+ self.teacher_model.bbox_head.get_label_assignment(
+ *outs_teacher, batch_gt_instances, batch_img_metas,
+ batch_gt_instances_ignore)
+
+ # the student use the label assignment from the teacher to learn
+ x = self.extract_feat(batch_inputs)
+ losses = self.bbox_head.loss(x, label_assignment_results,
+ batch_data_samples)
+ return losses
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/mask2former.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/mask2former.py
new file mode 100644
index 0000000000000000000000000000000000000000..4f38ef44e482039fdf7476d048eee5df2a96fd9b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/mask2former.py
@@ -0,0 +1,30 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .maskformer import MaskFormer
+
+
+@MODELS.register_module()
+class Mask2Former(MaskFormer):
+ r"""Implementation of `Masked-attention Mask
+ Transformer for Universal Image Segmentation
+ `_."""
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: OptConfigType = None,
+ panoptic_head: OptConfigType = None,
+ panoptic_fusion_head: OptConfigType = None,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None):
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ panoptic_head=panoptic_head,
+ panoptic_fusion_head=panoptic_fusion_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/mask_rcnn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/mask_rcnn.py
new file mode 100644
index 0000000000000000000000000000000000000000..880ee1e8ac3926d618ef47985549d3214175ee73
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/mask_rcnn.py
@@ -0,0 +1,30 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.config import ConfigDict
+
+from mmdet.registry import MODELS
+from mmdet.utils import OptConfigType, OptMultiConfig
+from .two_stage import TwoStageDetector
+
+
+@MODELS.register_module()
+class MaskRCNN(TwoStageDetector):
+ """Implementation of `Mask R-CNN `_"""
+
+ def __init__(self,
+ backbone: ConfigDict,
+ rpn_head: ConfigDict,
+ roi_head: ConfigDict,
+ train_cfg: ConfigDict,
+ test_cfg: ConfigDict,
+ neck: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ rpn_head=rpn_head,
+ roi_head=roi_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ init_cfg=init_cfg,
+ data_preprocessor=data_preprocessor)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/mask_scoring_rcnn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/mask_scoring_rcnn.py
new file mode 100644
index 0000000000000000000000000000000000000000..e09d3a1041f929113962e42bdf8b169e52dabe25
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/mask_scoring_rcnn.py
@@ -0,0 +1,31 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .two_stage import TwoStageDetector
+
+
+@MODELS.register_module()
+class MaskScoringRCNN(TwoStageDetector):
+ """Mask Scoring RCNN.
+
+ https://arxiv.org/abs/1903.00241
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ rpn_head: ConfigType,
+ roi_head: ConfigType,
+ train_cfg: ConfigType,
+ test_cfg: ConfigType,
+ neck: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ rpn_head=rpn_head,
+ roi_head=roi_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/maskformer.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/maskformer.py
new file mode 100644
index 0000000000000000000000000000000000000000..7493c00e1b87cf9b2fbd2c80f1e642f6eb2bea55
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/maskformer.py
@@ -0,0 +1,170 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, List, Tuple
+
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class MaskFormer(SingleStageDetector):
+ r"""Implementation of `Per-Pixel Classification is
+ NOT All You Need for Semantic Segmentation
+ `_."""
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: OptConfigType = None,
+ panoptic_head: OptConfigType = None,
+ panoptic_fusion_head: OptConfigType = None,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None):
+ super(SingleStageDetector, self).__init__(
+ data_preprocessor=data_preprocessor, init_cfg=init_cfg)
+ self.backbone = MODELS.build(backbone)
+ if neck is not None:
+ self.neck = MODELS.build(neck)
+
+ panoptic_head_ = panoptic_head.deepcopy()
+ panoptic_head_.update(train_cfg=train_cfg)
+ panoptic_head_.update(test_cfg=test_cfg)
+ self.panoptic_head = MODELS.build(panoptic_head_)
+
+ panoptic_fusion_head_ = panoptic_fusion_head.deepcopy()
+ panoptic_fusion_head_.update(test_cfg=test_cfg)
+ self.panoptic_fusion_head = MODELS.build(panoptic_fusion_head_)
+
+ self.num_things_classes = self.panoptic_head.num_things_classes
+ self.num_stuff_classes = self.panoptic_head.num_stuff_classes
+ self.num_classes = self.panoptic_head.num_classes
+
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+
+ def loss(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> Dict[str, Tensor]:
+ """
+ Args:
+ batch_inputs (Tensor): Input images of shape (N, C, H, W).
+ These should usually be mean centered and std scaled.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict[str, Tensor]: a dictionary of loss components
+ """
+ x = self.extract_feat(batch_inputs)
+ losses = self.panoptic_head.loss(x, batch_data_samples)
+ return losses
+
+ def predict(self,
+ batch_inputs: Tensor,
+ batch_data_samples: SampleList,
+ rescale: bool = True) -> SampleList:
+ """Predict results from a batch of inputs and data samples with post-
+ processing.
+
+ Args:
+ batch_inputs (Tensor): Inputs with shape (N, C, H, W).
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool): Whether to rescale the results.
+ Defaults to True.
+
+ Returns:
+ list[:obj:`DetDataSample`]: Detection results of the
+ input images. Each DetDataSample usually contain
+ 'pred_instances' and `pred_panoptic_seg`. And the
+ ``pred_instances`` usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+
+ And the ``pred_panoptic_seg`` contains the following key
+
+ - sem_seg (Tensor): panoptic segmentation mask, has a
+ shape (1, h, w).
+ """
+ feats = self.extract_feat(batch_inputs)
+ mask_cls_results, mask_pred_results = self.panoptic_head.predict(
+ feats, batch_data_samples)
+ results_list = self.panoptic_fusion_head.predict(
+ mask_cls_results,
+ mask_pred_results,
+ batch_data_samples,
+ rescale=rescale)
+ results = self.add_pred_to_datasample(batch_data_samples, results_list)
+
+ return results
+
+ def add_pred_to_datasample(self, data_samples: SampleList,
+ results_list: List[dict]) -> SampleList:
+ """Add predictions to `DetDataSample`.
+
+ Args:
+ data_samples (list[:obj:`DetDataSample`], optional): A batch of
+ data samples that contain annotations and predictions.
+ results_list (List[dict]): Instance segmentation, segmantic
+ segmentation and panoptic segmentation results.
+
+ Returns:
+ list[:obj:`DetDataSample`]: Detection results of the
+ input images. Each DetDataSample usually contain
+ 'pred_instances' and `pred_panoptic_seg`. And the
+ ``pred_instances`` usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+
+ And the ``pred_panoptic_seg`` contains the following key
+
+ - sem_seg (Tensor): panoptic segmentation mask, has a
+ shape (1, h, w).
+ """
+ for data_sample, pred_results in zip(data_samples, results_list):
+ if 'pan_results' in pred_results:
+ data_sample.pred_panoptic_seg = pred_results['pan_results']
+
+ if 'ins_results' in pred_results:
+ data_sample.pred_instances = pred_results['ins_results']
+
+ assert 'sem_results' not in pred_results, 'segmantic ' \
+ 'segmentation results are not supported yet.'
+
+ return data_samples
+
+ def _forward(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> Tuple[List[Tensor]]:
+ """Network forward process. Usually includes backbone, neck and head
+ forward without any post-processing.
+
+ Args:
+ batch_inputs (Tensor): Inputs with shape (N, C, H, W).
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ tuple[List[Tensor]]: A tuple of features from ``panoptic_head``
+ forward.
+ """
+ feats = self.extract_feat(batch_inputs)
+ results = self.panoptic_head.forward(feats, batch_data_samples)
+ return results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/nasfcos.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/nasfcos.py
new file mode 100644
index 0000000000000000000000000000000000000000..da2b911bcfc6b0ba51b00d9b3948a3df7af2e74f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/nasfcos.py
@@ -0,0 +1,43 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class NASFCOS(SingleStageDetector):
+ """Implementation of `NAS-FCOS: Fast Neural Architecture Search for Object
+ Detection. `_
+
+ Args:
+ backbone (:obj:`ConfigDict` or dict): The backbone config.
+ neck (:obj:`ConfigDict` or dict): The neck config.
+ bbox_head (:obj:`ConfigDict` or dict): The bbox head config.
+ train_cfg (:obj:`ConfigDict` or dict, optional): The training config
+ of NASFCOS. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): The testing config
+ of NASFCOS. Defaults to None.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional): Config of
+ :class:`DetDataPreprocessor` to process the input data.
+ Defaults to None.
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or
+ list[dict], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/paa.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/paa.py
new file mode 100644
index 0000000000000000000000000000000000000000..094306b2fbd18ba45536470ec80443e4ff793e67
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/paa.py
@@ -0,0 +1,41 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class PAA(SingleStageDetector):
+ """Implementation of `PAA `_
+
+ Args:
+ backbone (:obj:`ConfigDict` or dict): The backbone module.
+ neck (:obj:`ConfigDict` or dict): The neck module.
+ bbox_head (:obj:`ConfigDict` or dict): The bbox head module.
+ train_cfg (:obj:`ConfigDict` or dict, optional): The training config
+ of PAA. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): The testing config
+ of PAA. Defaults to None.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional): Config of
+ :class:`DetDataPreprocessor` to process the input data.
+ Defaults to None.
+ init_cfg (:obj:`ConfigDict` or dict, optional): the config to control
+ the initialization. Defaults to None.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/panoptic_fpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/panoptic_fpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..ae63ccc38931daa60b4e62f94dcf9f44574d3669
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/panoptic_fpn.py
@@ -0,0 +1,35 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .panoptic_two_stage_segmentor import TwoStagePanopticSegmentor
+
+
+@MODELS.register_module()
+class PanopticFPN(TwoStagePanopticSegmentor):
+ r"""Implementation of `Panoptic feature pyramid
+ networks `_"""
+
+ def __init__(
+ self,
+ backbone: ConfigType,
+ neck: OptConfigType = None,
+ rpn_head: OptConfigType = None,
+ roi_head: OptConfigType = None,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None,
+ # for panoptic segmentation
+ semantic_head: OptConfigType = None,
+ panoptic_fusion_head: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ rpn_head=rpn_head,
+ roi_head=roi_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg,
+ semantic_head=semantic_head,
+ panoptic_fusion_head=panoptic_fusion_head)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/panoptic_two_stage_segmentor.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/panoptic_two_stage_segmentor.py
new file mode 100644
index 0000000000000000000000000000000000000000..879edbe1ac6a0f482fdd740f4058e508e728414d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/panoptic_two_stage_segmentor.py
@@ -0,0 +1,234 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from typing import List
+
+import torch
+from mmengine.structures import PixelData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .two_stage import TwoStageDetector
+
+
+@MODELS.register_module()
+class TwoStagePanopticSegmentor(TwoStageDetector):
+ """Base class of Two-stage Panoptic Segmentor.
+
+ As well as the components in TwoStageDetector, Panoptic Segmentor has extra
+ semantic_head and panoptic_fusion_head.
+ """
+
+ def __init__(
+ self,
+ backbone: ConfigType,
+ neck: OptConfigType = None,
+ rpn_head: OptConfigType = None,
+ roi_head: OptConfigType = None,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None,
+ # for panoptic segmentation
+ semantic_head: OptConfigType = None,
+ panoptic_fusion_head: OptConfigType = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ rpn_head=rpn_head,
+ roi_head=roi_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
+
+ if semantic_head is not None:
+ self.semantic_head = MODELS.build(semantic_head)
+
+ if panoptic_fusion_head is not None:
+ panoptic_cfg = test_cfg.panoptic if test_cfg is not None else None
+ panoptic_fusion_head_ = panoptic_fusion_head.deepcopy()
+ panoptic_fusion_head_.update(test_cfg=panoptic_cfg)
+ self.panoptic_fusion_head = MODELS.build(panoptic_fusion_head_)
+
+ self.num_things_classes = self.panoptic_fusion_head.\
+ num_things_classes
+ self.num_stuff_classes = self.panoptic_fusion_head.\
+ num_stuff_classes
+ self.num_classes = self.panoptic_fusion_head.num_classes
+
+ @property
+ def with_semantic_head(self) -> bool:
+ """bool: whether the detector has semantic head"""
+ return hasattr(self,
+ 'semantic_head') and self.semantic_head is not None
+
+ @property
+ def with_panoptic_fusion_head(self) -> bool:
+ """bool: whether the detector has panoptic fusion head"""
+ return hasattr(self, 'panoptic_fusion_head') and \
+ self.panoptic_fusion_head is not None
+
+ def loss(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> dict:
+ """
+ Args:
+ batch_inputs (Tensor): Input images of shape (N, C, H, W).
+ These should usually be mean centered and std scaled.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ x = self.extract_feat(batch_inputs)
+
+ losses = dict()
+
+ # RPN forward and loss
+ if self.with_rpn:
+ proposal_cfg = self.train_cfg.get('rpn_proposal',
+ self.test_cfg.rpn)
+ rpn_data_samples = copy.deepcopy(batch_data_samples)
+ # set cat_id of gt_labels to 0 in RPN
+ for data_sample in rpn_data_samples:
+ data_sample.gt_instances.labels = \
+ torch.zeros_like(data_sample.gt_instances.labels)
+
+ rpn_losses, rpn_results_list = self.rpn_head.loss_and_predict(
+ x, rpn_data_samples, proposal_cfg=proposal_cfg)
+ # avoid get same name with roi_head loss
+ keys = rpn_losses.keys()
+ for key in list(keys):
+ if 'loss' in key and 'rpn' not in key:
+ rpn_losses[f'rpn_{key}'] = rpn_losses.pop(key)
+ losses.update(rpn_losses)
+ else:
+ # TODO: Not support currently, should have a check at Fast R-CNN
+ assert batch_data_samples[0].get('proposals', None) is not None
+ # use pre-defined proposals in InstanceData for the second stage
+ # to extract ROI features.
+ rpn_results_list = [
+ data_sample.proposals for data_sample in batch_data_samples
+ ]
+
+ roi_losses = self.roi_head.loss(x, rpn_results_list,
+ batch_data_samples)
+ losses.update(roi_losses)
+
+ semantic_loss = self.semantic_head.loss(x, batch_data_samples)
+ losses.update(semantic_loss)
+
+ return losses
+
+ def predict(self,
+ batch_inputs: Tensor,
+ batch_data_samples: SampleList,
+ rescale: bool = True) -> SampleList:
+ """Predict results from a batch of inputs and data samples with post-
+ processing.
+
+ Args:
+ batch_inputs (Tensor): Inputs with shape (N, C, H, W).
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool): Whether to rescale the results.
+ Defaults to True.
+
+ Returns:
+ List[:obj:`DetDataSample`]: Return the packed panoptic segmentation
+ results of input images. Each DetDataSample usually contains
+ 'pred_panoptic_seg'. And the 'pred_panoptic_seg' has a key
+ ``sem_seg``, which is a tensor of shape (1, h, w).
+ """
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+
+ x = self.extract_feat(batch_inputs)
+
+ # If there are no pre-defined proposals, use RPN to get proposals
+ if batch_data_samples[0].get('proposals', None) is None:
+ rpn_results_list = self.rpn_head.predict(
+ x, batch_data_samples, rescale=False)
+ else:
+ rpn_results_list = [
+ data_sample.proposals for data_sample in batch_data_samples
+ ]
+
+ results_list = self.roi_head.predict(
+ x, rpn_results_list, batch_data_samples, rescale=rescale)
+
+ seg_preds = self.semantic_head.predict(x, batch_img_metas, rescale)
+
+ results_list = self.panoptic_fusion_head.predict(
+ results_list, seg_preds)
+
+ batch_data_samples = self.add_pred_to_datasample(
+ batch_data_samples, results_list)
+ return batch_data_samples
+
+ # TODO the code has not been verified and needs to be refactored later.
+ def _forward(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> tuple:
+ """Network forward process. Usually includes backbone, neck and head
+ forward without any post-processing.
+
+ Args:
+ batch_inputs (Tensor): Inputs with shape (N, C, H, W).
+
+ Returns:
+ tuple: A tuple of features from ``rpn_head``, ``roi_head`` and
+ ``semantic_head`` forward.
+ """
+ results = ()
+ x = self.extract_feat(batch_inputs)
+ rpn_outs = self.rpn_head.forward(x)
+ results = results + (rpn_outs)
+
+ # If there are no pre-defined proposals, use RPN to get proposals
+ if batch_data_samples[0].get('proposals', None) is None:
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+ rpn_results_list = self.rpn_head.predict_by_feat(
+ *rpn_outs, batch_img_metas=batch_img_metas, rescale=False)
+ else:
+ # TODO: Not checked currently.
+ rpn_results_list = [
+ data_sample.proposals for data_sample in batch_data_samples
+ ]
+
+ # roi_head
+ roi_outs = self.roi_head(x, rpn_results_list)
+ results = results + (roi_outs)
+
+ # semantic_head
+ sem_outs = self.semantic_head.forward(x)
+ results = results + (sem_outs['seg_preds'], )
+
+ return results
+
+ def add_pred_to_datasample(self, data_samples: SampleList,
+ results_list: List[PixelData]) -> SampleList:
+ """Add predictions to `DetDataSample`.
+
+ Args:
+ data_samples (list[:obj:`DetDataSample`]): The
+ annotation data of every samples.
+ results_list (List[PixelData]): Panoptic segmentation results of
+ each image.
+
+ Returns:
+ List[:obj:`DetDataSample`]: Return the packed panoptic segmentation
+ results of input images. Each DetDataSample usually contains
+ 'pred_panoptic_seg'. And the 'pred_panoptic_seg' has a key
+ ``sem_seg``, which is a tensor of shape (1, h, w).
+ """
+
+ for data_sample, pred_panoptic_seg in zip(data_samples, results_list):
+ data_sample.pred_panoptic_seg = pred_panoptic_seg
+ return data_samples
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/point_rend.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/point_rend.py
new file mode 100644
index 0000000000000000000000000000000000000000..5062ac0c945e79bd53e66e1642aec51113475cad
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/point_rend.py
@@ -0,0 +1,35 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.config import ConfigDict
+
+from mmdet.registry import MODELS
+from mmdet.utils import OptConfigType, OptMultiConfig
+from .two_stage import TwoStageDetector
+
+
+@MODELS.register_module()
+class PointRend(TwoStageDetector):
+ """PointRend: Image Segmentation as Rendering
+
+ This detector is the implementation of
+ `PointRend `_.
+
+ """
+
+ def __init__(self,
+ backbone: ConfigDict,
+ rpn_head: ConfigDict,
+ roi_head: ConfigDict,
+ train_cfg: ConfigDict,
+ test_cfg: ConfigDict,
+ neck: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ rpn_head=rpn_head,
+ roi_head=roi_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ init_cfg=init_cfg,
+ data_preprocessor=data_preprocessor)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/queryinst.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/queryinst.py
new file mode 100644
index 0000000000000000000000000000000000000000..400ce20c01f5c3825e343f2d32accf740c5dd55c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/queryinst.py
@@ -0,0 +1,29 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .sparse_rcnn import SparseRCNN
+
+
+@MODELS.register_module()
+class QueryInst(SparseRCNN):
+ r"""Implementation of
+ `Instances as Queries `_"""
+
+ def __init__(self,
+ backbone: ConfigType,
+ rpn_head: ConfigType,
+ roi_head: ConfigType,
+ train_cfg: ConfigType,
+ test_cfg: ConfigType,
+ neck: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ rpn_head=rpn_head,
+ roi_head=roi_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/reppoints_detector.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/reppoints_detector.py
new file mode 100644
index 0000000000000000000000000000000000000000..d86cec2ecda0671939e227c50f00379e81d3ac9c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/reppoints_detector.py
@@ -0,0 +1,30 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class RepPointsDetector(SingleStageDetector):
+ """RepPoints: Point Set Representation for Object Detection.
+
+ This detector is the implementation of:
+ - RepPoints detector (https://arxiv.org/pdf/1904.11490)
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None):
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/retinanet.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/retinanet.py
new file mode 100644
index 0000000000000000000000000000000000000000..03e3cb20e5bda603e9384d83688a56fa590e6de8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/retinanet.py
@@ -0,0 +1,26 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class RetinaNet(SingleStageDetector):
+ """Implementation of `RetinaNet `_"""
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/rpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/rpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..72fe8521fcc9bc796801b2dd68269bb57aaab984
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/rpn.py
@@ -0,0 +1,81 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import warnings
+
+import torch
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class RPN(SingleStageDetector):
+ """Implementation of Region Proposal Network.
+
+ Args:
+ backbone (:obj:`ConfigDict` or dict): The backbone config.
+ neck (:obj:`ConfigDict` or dict): The neck config.
+ bbox_head (:obj:`ConfigDict` or dict): The bbox head config.
+ train_cfg (:obj:`ConfigDict` or dict, optional): The training config.
+ test_cfg (:obj:`ConfigDict` or dict, optional): The testing config.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional): Config of
+ :class:`DetDataPreprocessor` to process the input data.
+ Defaults to None.
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or
+ list[dict], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ rpn_head: ConfigType,
+ train_cfg: ConfigType,
+ test_cfg: ConfigType,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None,
+ **kwargs) -> None:
+ super(SingleStageDetector, self).__init__(
+ data_preprocessor=data_preprocessor, init_cfg=init_cfg)
+ self.backbone = MODELS.build(backbone)
+ self.neck = MODELS.build(neck) if neck is not None else None
+ rpn_train_cfg = train_cfg['rpn'] if train_cfg is not None else None
+ rpn_head_num_classes = rpn_head.get('num_classes', 1)
+ if rpn_head_num_classes != 1:
+ warnings.warn('The `num_classes` should be 1 in RPN, but get '
+ f'{rpn_head_num_classes}, please set '
+ 'rpn_head.num_classes = 1 in your config file.')
+ rpn_head.update(num_classes=1)
+ rpn_head.update(train_cfg=rpn_train_cfg)
+ rpn_head.update(test_cfg=test_cfg['rpn'])
+ self.bbox_head = MODELS.build(rpn_head)
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+
+ def loss(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> dict:
+ """Calculate losses from a batch of inputs and data samples.
+
+ Args:
+ batch_inputs (Tensor): Input images of shape (N, C, H, W).
+ These should usually be mean centered and std scaled.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ x = self.extract_feat(batch_inputs)
+
+ # set cat_id of gt_labels to 0 in RPN
+ rpn_data_samples = copy.deepcopy(batch_data_samples)
+ for data_sample in rpn_data_samples:
+ data_sample.gt_instances.labels = \
+ torch.zeros_like(data_sample.gt_instances.labels)
+
+ losses = self.bbox_head.loss(x, rpn_data_samples)
+ return losses
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/rtmdet.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/rtmdet.py
new file mode 100644
index 0000000000000000000000000000000000000000..b43e053fc41a4b8400bbc0946fffedfa735b9451
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/rtmdet.py
@@ -0,0 +1,52 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+from mmengine.dist import get_world_size
+from mmengine.logging import print_log
+
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class RTMDet(SingleStageDetector):
+ """Implementation of RTMDet.
+
+ Args:
+ backbone (:obj:`ConfigDict` or dict): The backbone module.
+ neck (:obj:`ConfigDict` or dict): The neck module.
+ bbox_head (:obj:`ConfigDict` or dict): The bbox head module.
+ train_cfg (:obj:`ConfigDict` or dict, optional): The training config
+ of ATSS. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): The testing config
+ of ATSS. Defaults to None.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional): Config of
+ :class:`DetDataPreprocessor` to process the input data.
+ Defaults to None.
+ init_cfg (:obj:`ConfigDict` or dict, optional): the config to control
+ the initialization. Defaults to None.
+ use_syncbn (bool): Whether to use SyncBatchNorm. Defaults to True.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None,
+ use_syncbn: bool = True) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
+
+ # TODO: Waiting for mmengine support
+ if use_syncbn and get_world_size() > 1:
+ torch.nn.SyncBatchNorm.convert_sync_batchnorm(self)
+ print_log('Using SyncBatchNorm()', 'current')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/scnet.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/scnet.py
new file mode 100644
index 0000000000000000000000000000000000000000..606a0203869f1731a21d811f06c4781f5cd90d8d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/scnet.py
@@ -0,0 +1,11 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from .cascade_rcnn import CascadeRCNN
+
+
+@MODELS.register_module()
+class SCNet(CascadeRCNN):
+ """Implementation of `SCNet `_"""
+
+ def __init__(self, **kwargs) -> None:
+ super().__init__(**kwargs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/semi_base.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/semi_base.py
new file mode 100644
index 0000000000000000000000000000000000000000..f3f0c8c030830e188bf3ad245d5b3cb471ecb04f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/semi_base.py
@@ -0,0 +1,266 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from typing import Dict, List, Optional, Tuple, Union
+
+import torch
+import torch.nn as nn
+from torch import Tensor
+
+from mmdet.models.utils import (filter_gt_instances, rename_loss_dict,
+ reweight_loss_dict)
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import bbox_project
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .base import BaseDetector
+
+
+@MODELS.register_module()
+class SemiBaseDetector(BaseDetector):
+ """Base class for semi-supervised detectors.
+
+ Semi-supervised detectors typically consisting of a teacher model
+ updated by exponential moving average and a student model updated
+ by gradient descent.
+
+ Args:
+ detector (:obj:`ConfigDict` or dict): The detector config.
+ semi_train_cfg (:obj:`ConfigDict` or dict, optional):
+ The semi-supervised training config.
+ semi_test_cfg (:obj:`ConfigDict` or dict, optional):
+ The semi-supervised testing config.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional): Config of
+ :class:`DetDataPreprocessor` to process the input data.
+ Defaults to None.
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or
+ list[dict], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ detector: ConfigType,
+ semi_train_cfg: OptConfigType = None,
+ semi_test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ data_preprocessor=data_preprocessor, init_cfg=init_cfg)
+ self.student = MODELS.build(detector)
+ self.teacher = MODELS.build(detector)
+ self.semi_train_cfg = semi_train_cfg
+ self.semi_test_cfg = semi_test_cfg
+ if self.semi_train_cfg.get('freeze_teacher', True) is True:
+ self.freeze(self.teacher)
+
+ @staticmethod
+ def freeze(model: nn.Module):
+ """Freeze the model."""
+ model.eval()
+ for param in model.parameters():
+ param.requires_grad = False
+
+ def loss(self, multi_batch_inputs: Dict[str, Tensor],
+ multi_batch_data_samples: Dict[str, SampleList]) -> dict:
+ """Calculate losses from multi-branch inputs and data samples.
+
+ Args:
+ multi_batch_inputs (Dict[str, Tensor]): The dict of multi-branch
+ input images, each value with shape (N, C, H, W).
+ Each value should usually be mean centered and std scaled.
+ multi_batch_data_samples (Dict[str, List[:obj:`DetDataSample`]]):
+ The dict of multi-branch data samples.
+
+ Returns:
+ dict: A dictionary of loss components
+ """
+ losses = dict()
+ losses.update(**self.loss_by_gt_instances(
+ multi_batch_inputs['sup'], multi_batch_data_samples['sup']))
+
+ origin_pseudo_data_samples, batch_info = self.get_pseudo_instances(
+ multi_batch_inputs['unsup_teacher'],
+ multi_batch_data_samples['unsup_teacher'])
+ multi_batch_data_samples[
+ 'unsup_student'] = self.project_pseudo_instances(
+ origin_pseudo_data_samples,
+ multi_batch_data_samples['unsup_student'])
+ losses.update(**self.loss_by_pseudo_instances(
+ multi_batch_inputs['unsup_student'],
+ multi_batch_data_samples['unsup_student'], batch_info))
+ return losses
+
+ def loss_by_gt_instances(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> dict:
+ """Calculate losses from a batch of inputs and ground-truth data
+ samples.
+
+ Args:
+ batch_inputs (Tensor): Input images of shape (N, C, H, W).
+ These should usually be mean centered and std scaled.
+ batch_data_samples (List[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict: A dictionary of loss components
+ """
+
+ losses = self.student.loss(batch_inputs, batch_data_samples)
+ sup_weight = self.semi_train_cfg.get('sup_weight', 1.)
+ return rename_loss_dict('sup_', reweight_loss_dict(losses, sup_weight))
+
+ def loss_by_pseudo_instances(self,
+ batch_inputs: Tensor,
+ batch_data_samples: SampleList,
+ batch_info: Optional[dict] = None) -> dict:
+ """Calculate losses from a batch of inputs and pseudo data samples.
+
+ Args:
+ batch_inputs (Tensor): Input images of shape (N, C, H, W).
+ These should usually be mean centered and std scaled.
+ batch_data_samples (List[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`,
+ which are `pseudo_instance` or `pseudo_panoptic_seg`
+ or `pseudo_sem_seg` in fact.
+ batch_info (dict): Batch information of teacher model
+ forward propagation process. Defaults to None.
+
+ Returns:
+ dict: A dictionary of loss components
+ """
+ batch_data_samples = filter_gt_instances(
+ batch_data_samples, score_thr=self.semi_train_cfg.cls_pseudo_thr)
+ losses = self.student.loss(batch_inputs, batch_data_samples)
+ pseudo_instances_num = sum([
+ len(data_samples.gt_instances)
+ for data_samples in batch_data_samples
+ ])
+ unsup_weight = self.semi_train_cfg.get(
+ 'unsup_weight', 1.) if pseudo_instances_num > 0 else 0.
+ return rename_loss_dict('unsup_',
+ reweight_loss_dict(losses, unsup_weight))
+
+ @torch.no_grad()
+ def get_pseudo_instances(
+ self, batch_inputs: Tensor, batch_data_samples: SampleList
+ ) -> Tuple[SampleList, Optional[dict]]:
+ """Get pseudo instances from teacher model."""
+ self.teacher.eval()
+ results_list = self.teacher.predict(
+ batch_inputs, batch_data_samples, rescale=False)
+ batch_info = {}
+ for data_samples, results in zip(batch_data_samples, results_list):
+ data_samples.gt_instances = results.pred_instances
+ data_samples.gt_instances.bboxes = bbox_project(
+ data_samples.gt_instances.bboxes,
+ torch.from_numpy(data_samples.homography_matrix).inverse().to(
+ self.data_preprocessor.device), data_samples.ori_shape)
+ return batch_data_samples, batch_info
+
+ def project_pseudo_instances(self, batch_pseudo_instances: SampleList,
+ batch_data_samples: SampleList) -> SampleList:
+ """Project pseudo instances."""
+ for pseudo_instances, data_samples in zip(batch_pseudo_instances,
+ batch_data_samples):
+ data_samples.gt_instances = copy.deepcopy(
+ pseudo_instances.gt_instances)
+ data_samples.gt_instances.bboxes = bbox_project(
+ data_samples.gt_instances.bboxes,
+ torch.tensor(data_samples.homography_matrix).to(
+ self.data_preprocessor.device), data_samples.img_shape)
+ wh_thr = self.semi_train_cfg.get('min_pseudo_bbox_wh', (1e-2, 1e-2))
+ return filter_gt_instances(batch_data_samples, wh_thr=wh_thr)
+
+ def predict(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> SampleList:
+ """Predict results from a batch of inputs and data samples with post-
+ processing.
+
+ Args:
+ batch_inputs (Tensor): Inputs with shape (N, C, H, W).
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool): Whether to rescale the results.
+ Defaults to True.
+
+ Returns:
+ list[:obj:`DetDataSample`]: Return the detection results of the
+ input images. The returns value is DetDataSample,
+ which usually contain 'pred_instances'. And the
+ ``pred_instances`` usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+ """
+ if self.semi_test_cfg.get('predict_on', 'teacher') == 'teacher':
+ return self.teacher(
+ batch_inputs, batch_data_samples, mode='predict')
+ else:
+ return self.student(
+ batch_inputs, batch_data_samples, mode='predict')
+
+ def _forward(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> SampleList:
+ """Network forward process. Usually includes backbone, neck and head
+ forward without any post-processing.
+
+ Args:
+ batch_inputs (Tensor): Inputs with shape (N, C, H, W).
+
+ Returns:
+ tuple: A tuple of features from ``rpn_head`` and ``roi_head``
+ forward.
+ """
+ if self.semi_test_cfg.get('forward_on', 'teacher') == 'teacher':
+ return self.teacher(
+ batch_inputs, batch_data_samples, mode='tensor')
+ else:
+ return self.student(
+ batch_inputs, batch_data_samples, mode='tensor')
+
+ def extract_feat(self, batch_inputs: Tensor) -> Tuple[Tensor]:
+ """Extract features.
+
+ Args:
+ batch_inputs (Tensor): Image tensor with shape (N, C, H ,W).
+
+ Returns:
+ tuple[Tensor]: Multi-level features that may have
+ different resolutions.
+ """
+ if self.semi_test_cfg.get('extract_feat_on', 'teacher') == 'teacher':
+ return self.teacher.extract_feat(batch_inputs)
+ else:
+ return self.student.extract_feat(batch_inputs)
+
+ def _load_from_state_dict(self, state_dict: dict, prefix: str,
+ local_metadata: dict, strict: bool,
+ missing_keys: Union[List[str], str],
+ unexpected_keys: Union[List[str], str],
+ error_msgs: Union[List[str], str]) -> None:
+ """Add teacher and student prefixes to model parameter names."""
+ if not any([
+ 'student' in key or 'teacher' in key
+ for key in state_dict.keys()
+ ]):
+ keys = list(state_dict.keys())
+ state_dict.update({'teacher.' + k: state_dict[k] for k in keys})
+ state_dict.update({'student.' + k: state_dict[k] for k in keys})
+ for k in keys:
+ state_dict.pop(k)
+ return super()._load_from_state_dict(
+ state_dict,
+ prefix,
+ local_metadata,
+ strict,
+ missing_keys,
+ unexpected_keys,
+ error_msgs,
+ )
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/single_stage.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/single_stage.py
new file mode 100644
index 0000000000000000000000000000000000000000..06c074085967bbc9040d93e5eb446b67a006087e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/single_stage.py
@@ -0,0 +1,149 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple, Union
+
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import OptSampleList, SampleList
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .base import BaseDetector
+
+
+@MODELS.register_module()
+class SingleStageDetector(BaseDetector):
+ """Base class for single-stage detectors.
+
+ Single-stage detectors directly and densely predict bounding boxes on the
+ output features of the backbone+neck.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: OptConfigType = None,
+ bbox_head: OptConfigType = None,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ data_preprocessor=data_preprocessor, init_cfg=init_cfg)
+ self.backbone = MODELS.build(backbone)
+ if neck is not None:
+ self.neck = MODELS.build(neck)
+ bbox_head.update(train_cfg=train_cfg)
+ bbox_head.update(test_cfg=test_cfg)
+ self.bbox_head = MODELS.build(bbox_head)
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+
+ def _load_from_state_dict(self, state_dict: dict, prefix: str,
+ local_metadata: dict, strict: bool,
+ missing_keys: Union[List[str], str],
+ unexpected_keys: Union[List[str], str],
+ error_msgs: Union[List[str], str]) -> None:
+ """Exchange bbox_head key to rpn_head key when loading two-stage
+ weights into single-stage model."""
+ bbox_head_prefix = prefix + '.bbox_head' if prefix else 'bbox_head'
+ bbox_head_keys = [
+ k for k in state_dict.keys() if k.startswith(bbox_head_prefix)
+ ]
+ rpn_head_prefix = prefix + '.rpn_head' if prefix else 'rpn_head'
+ rpn_head_keys = [
+ k for k in state_dict.keys() if k.startswith(rpn_head_prefix)
+ ]
+ if len(bbox_head_keys) == 0 and len(rpn_head_keys) != 0:
+ for rpn_head_key in rpn_head_keys:
+ bbox_head_key = bbox_head_prefix + \
+ rpn_head_key[len(rpn_head_prefix):]
+ state_dict[bbox_head_key] = state_dict.pop(rpn_head_key)
+ super()._load_from_state_dict(state_dict, prefix, local_metadata,
+ strict, missing_keys, unexpected_keys,
+ error_msgs)
+
+ def loss(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> Union[dict, list]:
+ """Calculate losses from a batch of inputs and data samples.
+
+ Args:
+ batch_inputs (Tensor): Input images of shape (N, C, H, W).
+ These should usually be mean centered and std scaled.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ x = self.extract_feat(batch_inputs)
+ losses = self.bbox_head.loss(x, batch_data_samples)
+ return losses
+
+ def predict(self,
+ batch_inputs: Tensor,
+ batch_data_samples: SampleList,
+ rescale: bool = True) -> SampleList:
+ """Predict results from a batch of inputs and data samples with post-
+ processing.
+
+ Args:
+ batch_inputs (Tensor): Inputs with shape (N, C, H, W).
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool): Whether to rescale the results.
+ Defaults to True.
+
+ Returns:
+ list[:obj:`DetDataSample`]: Detection results of the
+ input images. Each DetDataSample usually contain
+ 'pred_instances'. And the ``pred_instances`` usually
+ contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ x = self.extract_feat(batch_inputs)
+ results_list = self.bbox_head.predict(
+ x, batch_data_samples, rescale=rescale)
+ batch_data_samples = self.add_pred_to_datasample(
+ batch_data_samples, results_list)
+ return batch_data_samples
+
+ def _forward(
+ self,
+ batch_inputs: Tensor,
+ batch_data_samples: OptSampleList = None) -> Tuple[List[Tensor]]:
+ """Network forward process. Usually includes backbone, neck and head
+ forward without any post-processing.
+
+ Args:
+ batch_inputs (Tensor): Inputs with shape (N, C, H, W).
+ batch_data_samples (list[:obj:`DetDataSample`]): Each item contains
+ the meta information of each image and corresponding
+ annotations.
+
+ Returns:
+ tuple[list]: A tuple of features from ``bbox_head`` forward.
+ """
+ x = self.extract_feat(batch_inputs)
+ results = self.bbox_head.forward(x)
+ return results
+
+ def extract_feat(self, batch_inputs: Tensor) -> Tuple[Tensor]:
+ """Extract features.
+
+ Args:
+ batch_inputs (Tensor): Image tensor with shape (N, C, H ,W).
+
+ Returns:
+ tuple[Tensor]: Multi-level features that may have
+ different resolutions.
+ """
+ x = self.backbone(batch_inputs)
+ if self.with_neck:
+ x = self.neck(x)
+ return x
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/single_stage_instance_seg.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/single_stage_instance_seg.py
new file mode 100644
index 0000000000000000000000000000000000000000..acb5f0d2f8e4636b86b4b66cbf5c4916d0dae16f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/single_stage_instance_seg.py
@@ -0,0 +1,180 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from typing import Tuple
+
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import OptSampleList, SampleList
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .base import BaseDetector
+
+INF = 1e8
+
+
+@MODELS.register_module()
+class SingleStageInstanceSegmentor(BaseDetector):
+ """Base class for single-stage instance segmentors."""
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: OptConfigType = None,
+ bbox_head: OptConfigType = None,
+ mask_head: OptConfigType = None,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ data_preprocessor=data_preprocessor, init_cfg=init_cfg)
+ self.backbone = MODELS.build(backbone)
+ if neck is not None:
+ self.neck = MODELS.build(neck)
+ else:
+ self.neck = None
+ if bbox_head is not None:
+ bbox_head.update(train_cfg=copy.deepcopy(train_cfg))
+ bbox_head.update(test_cfg=copy.deepcopy(test_cfg))
+ self.bbox_head = MODELS.build(bbox_head)
+ else:
+ self.bbox_head = None
+
+ assert mask_head, f'`mask_head` must ' \
+ f'be implemented in {self.__class__.__name__}'
+ mask_head.update(train_cfg=copy.deepcopy(train_cfg))
+ mask_head.update(test_cfg=copy.deepcopy(test_cfg))
+ self.mask_head = MODELS.build(mask_head)
+
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+
+ def extract_feat(self, batch_inputs: Tensor) -> Tuple[Tensor]:
+ """Extract features.
+
+ Args:
+ batch_inputs (Tensor): Image tensor with shape (N, C, H ,W).
+
+ Returns:
+ tuple[Tensor]: Multi-level features that may have different
+ resolutions.
+ """
+ x = self.backbone(batch_inputs)
+ if self.with_neck:
+ x = self.neck(x)
+ return x
+
+ def _forward(self,
+ batch_inputs: Tensor,
+ batch_data_samples: OptSampleList = None,
+ **kwargs) -> tuple:
+ """Network forward process. Usually includes backbone, neck and head
+ forward without any post-processing.
+
+ Args:
+ batch_inputs (Tensor): Inputs with shape (N, C, H, W).
+
+ Returns:
+ tuple: A tuple of features from ``bbox_head`` forward.
+ """
+ outs = ()
+ # backbone
+ x = self.extract_feat(batch_inputs)
+ # bbox_head
+ positive_infos = None
+ if self.with_bbox:
+ assert batch_data_samples is not None
+ bbox_outs = self.bbox_head.forward(x)
+ outs = outs + (bbox_outs, )
+ # It is necessary to use `bbox_head.loss` to update
+ # `_raw_positive_infos` which will be used in `get_positive_infos`
+ # positive_infos will be used in the following mask head.
+ _ = self.bbox_head.loss(x, batch_data_samples, **kwargs)
+ positive_infos = self.bbox_head.get_positive_infos()
+ # mask_head
+ if positive_infos is None:
+ mask_outs = self.mask_head.forward(x)
+ else:
+ mask_outs = self.mask_head.forward(x, positive_infos)
+ outs = outs + (mask_outs, )
+ return outs
+
+ def loss(self, batch_inputs: Tensor, batch_data_samples: SampleList,
+ **kwargs) -> dict:
+ """
+ Args:
+ batch_inputs (Tensor): Input images of shape (N, C, H, W).
+ These should usually be mean centered and std scaled.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ x = self.extract_feat(batch_inputs)
+ losses = dict()
+
+ positive_infos = None
+ # CondInst and YOLACT have bbox_head
+ if self.with_bbox:
+ bbox_losses = self.bbox_head.loss(x, batch_data_samples, **kwargs)
+ losses.update(bbox_losses)
+ # get positive information from bbox head, which will be used
+ # in the following mask head.
+ positive_infos = self.bbox_head.get_positive_infos()
+
+ mask_loss = self.mask_head.loss(
+ x, batch_data_samples, positive_infos=positive_infos, **kwargs)
+ # avoid loss override
+ assert not set(mask_loss.keys()) & set(losses.keys())
+
+ losses.update(mask_loss)
+ return losses
+
+ def predict(self,
+ batch_inputs: Tensor,
+ batch_data_samples: SampleList,
+ rescale: bool = True,
+ **kwargs) -> SampleList:
+ """Perform forward propagation of the mask head and predict mask
+ results on the features of the upstream network.
+
+ Args:
+ batch_inputs (Tensor): Inputs with shape (N, C, H, W).
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool): Whether to rescale the results.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`DetDataSample`]: Detection results of the
+ input images. Each DetDataSample usually contain
+ 'pred_instances'. And the ``pred_instances`` usually
+ contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+ """
+ x = self.extract_feat(batch_inputs)
+ if self.with_bbox:
+ # the bbox branch does not need to be scaled to the original
+ # image scale, because the mask branch will scale both bbox
+ # and mask at the same time.
+ bbox_rescale = rescale if not self.with_mask else False
+ results_list = self.bbox_head.predict(
+ x, batch_data_samples, rescale=bbox_rescale)
+ else:
+ results_list = None
+
+ results_list = self.mask_head.predict(
+ x, batch_data_samples, rescale=rescale, results_list=results_list)
+
+ batch_data_samples = self.add_pred_to_datasample(
+ batch_data_samples, results_list)
+ return batch_data_samples
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/soft_teacher.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/soft_teacher.py
new file mode 100644
index 0000000000000000000000000000000000000000..80853f1d8399c70008923067777a2581671ede0b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/soft_teacher.py
@@ -0,0 +1,378 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from typing import List, Optional, Tuple
+
+import torch
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.models.utils import (filter_gt_instances, rename_loss_dict,
+ reweight_loss_dict)
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import bbox2roi, bbox_project
+from mmdet.utils import ConfigType, InstanceList, OptConfigType, OptMultiConfig
+from ..utils.misc import unpack_gt_instances
+from .semi_base import SemiBaseDetector
+
+
+@MODELS.register_module()
+class SoftTeacher(SemiBaseDetector):
+ r"""Implementation of `End-to-End Semi-Supervised Object Detection
+ with Soft Teacher `_
+
+ Args:
+ detector (:obj:`ConfigDict` or dict): The detector config.
+ semi_train_cfg (:obj:`ConfigDict` or dict, optional):
+ The semi-supervised training config.
+ semi_test_cfg (:obj:`ConfigDict` or dict, optional):
+ The semi-supervised testing config.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional): Config of
+ :class:`DetDataPreprocessor` to process the input data.
+ Defaults to None.
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or
+ list[dict], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ detector: ConfigType,
+ semi_train_cfg: OptConfigType = None,
+ semi_test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ detector=detector,
+ semi_train_cfg=semi_train_cfg,
+ semi_test_cfg=semi_test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
+
+ def loss_by_pseudo_instances(self,
+ batch_inputs: Tensor,
+ batch_data_samples: SampleList,
+ batch_info: Optional[dict] = None) -> dict:
+ """Calculate losses from a batch of inputs and pseudo data samples.
+
+ Args:
+ batch_inputs (Tensor): Input images of shape (N, C, H, W).
+ These should usually be mean centered and std scaled.
+ batch_data_samples (List[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`,
+ which are `pseudo_instance` or `pseudo_panoptic_seg`
+ or `pseudo_sem_seg` in fact.
+ batch_info (dict): Batch information of teacher model
+ forward propagation process. Defaults to None.
+
+ Returns:
+ dict: A dictionary of loss components
+ """
+
+ x = self.student.extract_feat(batch_inputs)
+
+ losses = {}
+ rpn_losses, rpn_results_list = self.rpn_loss_by_pseudo_instances(
+ x, batch_data_samples)
+ losses.update(**rpn_losses)
+ losses.update(**self.rcnn_cls_loss_by_pseudo_instances(
+ x, rpn_results_list, batch_data_samples, batch_info))
+ losses.update(**self.rcnn_reg_loss_by_pseudo_instances(
+ x, rpn_results_list, batch_data_samples))
+ unsup_weight = self.semi_train_cfg.get('unsup_weight', 1.)
+ return rename_loss_dict('unsup_',
+ reweight_loss_dict(losses, unsup_weight))
+
+ @torch.no_grad()
+ def get_pseudo_instances(
+ self, batch_inputs: Tensor, batch_data_samples: SampleList
+ ) -> Tuple[SampleList, Optional[dict]]:
+ """Get pseudo instances from teacher model."""
+ assert self.teacher.with_bbox, 'Bbox head must be implemented.'
+ x = self.teacher.extract_feat(batch_inputs)
+
+ # If there are no pre-defined proposals, use RPN to get proposals
+ if batch_data_samples[0].get('proposals', None) is None:
+ rpn_results_list = self.teacher.rpn_head.predict(
+ x, batch_data_samples, rescale=False)
+ else:
+ rpn_results_list = [
+ data_sample.proposals for data_sample in batch_data_samples
+ ]
+
+ results_list = self.teacher.roi_head.predict(
+ x, rpn_results_list, batch_data_samples, rescale=False)
+
+ for data_samples, results in zip(batch_data_samples, results_list):
+ data_samples.gt_instances = results
+
+ batch_data_samples = filter_gt_instances(
+ batch_data_samples,
+ score_thr=self.semi_train_cfg.pseudo_label_initial_score_thr)
+
+ reg_uncs_list = self.compute_uncertainty_with_aug(
+ x, batch_data_samples)
+
+ for data_samples, reg_uncs in zip(batch_data_samples, reg_uncs_list):
+ data_samples.gt_instances['reg_uncs'] = reg_uncs
+ data_samples.gt_instances.bboxes = bbox_project(
+ data_samples.gt_instances.bboxes,
+ torch.from_numpy(data_samples.homography_matrix).inverse().to(
+ self.data_preprocessor.device), data_samples.ori_shape)
+
+ batch_info = {
+ 'feat': x,
+ 'img_shape': [],
+ 'homography_matrix': [],
+ 'metainfo': []
+ }
+ for data_samples in batch_data_samples:
+ batch_info['img_shape'].append(data_samples.img_shape)
+ batch_info['homography_matrix'].append(
+ torch.from_numpy(data_samples.homography_matrix).to(
+ self.data_preprocessor.device))
+ batch_info['metainfo'].append(data_samples.metainfo)
+ return batch_data_samples, batch_info
+
+ def rpn_loss_by_pseudo_instances(self, x: Tuple[Tensor],
+ batch_data_samples: SampleList) -> dict:
+ """Calculate rpn loss from a batch of inputs and pseudo data samples.
+
+ Args:
+ x (tuple[Tensor]): Features from FPN.
+ batch_data_samples (List[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`,
+ which are `pseudo_instance` or `pseudo_panoptic_seg`
+ or `pseudo_sem_seg` in fact.
+ Returns:
+ dict: A dictionary of rpn loss components
+ """
+
+ rpn_data_samples = copy.deepcopy(batch_data_samples)
+ rpn_data_samples = filter_gt_instances(
+ rpn_data_samples, score_thr=self.semi_train_cfg.rpn_pseudo_thr)
+ proposal_cfg = self.student.train_cfg.get('rpn_proposal',
+ self.student.test_cfg.rpn)
+ # set cat_id of gt_labels to 0 in RPN
+ for data_sample in rpn_data_samples:
+ data_sample.gt_instances.labels = \
+ torch.zeros_like(data_sample.gt_instances.labels)
+
+ rpn_losses, rpn_results_list = self.student.rpn_head.loss_and_predict(
+ x, rpn_data_samples, proposal_cfg=proposal_cfg)
+ for key in rpn_losses.keys():
+ if 'loss' in key and 'rpn' not in key:
+ rpn_losses[f'rpn_{key}'] = rpn_losses.pop(key)
+ return rpn_losses, rpn_results_list
+
+ def rcnn_cls_loss_by_pseudo_instances(self, x: Tuple[Tensor],
+ unsup_rpn_results_list: InstanceList,
+ batch_data_samples: SampleList,
+ batch_info: dict) -> dict:
+ """Calculate classification loss from a batch of inputs and pseudo data
+ samples.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ unsup_rpn_results_list (list[:obj:`InstanceData`]):
+ List of region proposals.
+ batch_data_samples (List[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`,
+ which are `pseudo_instance` or `pseudo_panoptic_seg`
+ or `pseudo_sem_seg` in fact.
+ batch_info (dict): Batch information of teacher model
+ forward propagation process.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of rcnn
+ classification loss components
+ """
+ rpn_results_list = copy.deepcopy(unsup_rpn_results_list)
+ cls_data_samples = copy.deepcopy(batch_data_samples)
+ cls_data_samples = filter_gt_instances(
+ cls_data_samples, score_thr=self.semi_train_cfg.cls_pseudo_thr)
+
+ outputs = unpack_gt_instances(cls_data_samples)
+ batch_gt_instances, batch_gt_instances_ignore, _ = outputs
+
+ # assign gts and sample proposals
+ num_imgs = len(cls_data_samples)
+ sampling_results = []
+ for i in range(num_imgs):
+ # rename rpn_results.bboxes to rpn_results.priors
+ rpn_results = rpn_results_list[i]
+ rpn_results.priors = rpn_results.pop('bboxes')
+ assign_result = self.student.roi_head.bbox_assigner.assign(
+ rpn_results, batch_gt_instances[i],
+ batch_gt_instances_ignore[i])
+ sampling_result = self.student.roi_head.bbox_sampler.sample(
+ assign_result,
+ rpn_results,
+ batch_gt_instances[i],
+ feats=[lvl_feat[i][None] for lvl_feat in x])
+ sampling_results.append(sampling_result)
+
+ selected_bboxes = [res.priors for res in sampling_results]
+ rois = bbox2roi(selected_bboxes)
+ bbox_results = self.student.roi_head._bbox_forward(x, rois)
+ # cls_reg_targets is a tuple of labels, label_weights,
+ # and bbox_targets, bbox_weights
+ cls_reg_targets = self.student.roi_head.bbox_head.get_targets(
+ sampling_results, self.student.train_cfg.rcnn)
+
+ selected_results_list = []
+ for bboxes, data_samples, teacher_matrix, teacher_img_shape in zip(
+ selected_bboxes, batch_data_samples,
+ batch_info['homography_matrix'], batch_info['img_shape']):
+ student_matrix = torch.tensor(
+ data_samples.homography_matrix, device=teacher_matrix.device)
+ homography_matrix = teacher_matrix @ student_matrix.inverse()
+ projected_bboxes = bbox_project(bboxes, homography_matrix,
+ teacher_img_shape)
+ selected_results_list.append(InstanceData(bboxes=projected_bboxes))
+
+ with torch.no_grad():
+ results_list = self.teacher.roi_head.predict_bbox(
+ batch_info['feat'],
+ batch_info['metainfo'],
+ selected_results_list,
+ rcnn_test_cfg=None,
+ rescale=False)
+ bg_score = torch.cat(
+ [results.scores[:, -1] for results in results_list])
+ # cls_reg_targets[0] is labels
+ neg_inds = cls_reg_targets[
+ 0] == self.student.roi_head.bbox_head.num_classes
+ # cls_reg_targets[1] is label_weights
+ cls_reg_targets[1][neg_inds] = bg_score[neg_inds].detach()
+
+ losses = self.student.roi_head.bbox_head.loss(
+ bbox_results['cls_score'], bbox_results['bbox_pred'], rois,
+ *cls_reg_targets)
+ # cls_reg_targets[1] is label_weights
+ losses['loss_cls'] = losses['loss_cls'] * len(
+ cls_reg_targets[1]) / max(sum(cls_reg_targets[1]), 1.0)
+ return losses
+
+ def rcnn_reg_loss_by_pseudo_instances(
+ self, x: Tuple[Tensor], unsup_rpn_results_list: InstanceList,
+ batch_data_samples: SampleList) -> dict:
+ """Calculate rcnn regression loss from a batch of inputs and pseudo
+ data samples.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ unsup_rpn_results_list (list[:obj:`InstanceData`]):
+ List of region proposals.
+ batch_data_samples (List[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`,
+ which are `pseudo_instance` or `pseudo_panoptic_seg`
+ or `pseudo_sem_seg` in fact.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of rcnn
+ regression loss components
+ """
+ rpn_results_list = copy.deepcopy(unsup_rpn_results_list)
+ reg_data_samples = copy.deepcopy(batch_data_samples)
+ for data_samples in reg_data_samples:
+ if data_samples.gt_instances.bboxes.shape[0] > 0:
+ data_samples.gt_instances = data_samples.gt_instances[
+ data_samples.gt_instances.reg_uncs <
+ self.semi_train_cfg.reg_pseudo_thr]
+ roi_losses = self.student.roi_head.loss(x, rpn_results_list,
+ reg_data_samples)
+ return {'loss_bbox': roi_losses['loss_bbox']}
+
+ def compute_uncertainty_with_aug(
+ self, x: Tuple[Tensor],
+ batch_data_samples: SampleList) -> List[Tensor]:
+ """Compute uncertainty with augmented bboxes.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ batch_data_samples (List[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`,
+ which are `pseudo_instance` or `pseudo_panoptic_seg`
+ or `pseudo_sem_seg` in fact.
+
+ Returns:
+ list[Tensor]: A list of uncertainty for pseudo bboxes.
+ """
+ auged_results_list = self.aug_box(batch_data_samples,
+ self.semi_train_cfg.jitter_times,
+ self.semi_train_cfg.jitter_scale)
+ # flatten
+ auged_results_list = [
+ InstanceData(bboxes=auged.reshape(-1, auged.shape[-1]))
+ for auged in auged_results_list
+ ]
+
+ self.teacher.roi_head.test_cfg = None
+ results_list = self.teacher.roi_head.predict(
+ x, auged_results_list, batch_data_samples, rescale=False)
+ self.teacher.roi_head.test_cfg = self.teacher.test_cfg.rcnn
+
+ reg_channel = max(
+ [results.bboxes.shape[-1] for results in results_list]) // 4
+ bboxes = [
+ results.bboxes.reshape(self.semi_train_cfg.jitter_times, -1,
+ results.bboxes.shape[-1])
+ if results.bboxes.numel() > 0 else results.bboxes.new_zeros(
+ self.semi_train_cfg.jitter_times, 0, 4 * reg_channel).float()
+ for results in results_list
+ ]
+
+ box_unc = [bbox.std(dim=0) for bbox in bboxes]
+ bboxes = [bbox.mean(dim=0) for bbox in bboxes]
+ labels = [
+ data_samples.gt_instances.labels
+ for data_samples in batch_data_samples
+ ]
+ if reg_channel != 1:
+ bboxes = [
+ bbox.reshape(bbox.shape[0], reg_channel,
+ 4)[torch.arange(bbox.shape[0]), label]
+ for bbox, label in zip(bboxes, labels)
+ ]
+ box_unc = [
+ unc.reshape(unc.shape[0], reg_channel,
+ 4)[torch.arange(unc.shape[0]), label]
+ for unc, label in zip(box_unc, labels)
+ ]
+
+ box_shape = [(bbox[:, 2:4] - bbox[:, :2]).clamp(min=1.0)
+ for bbox in bboxes]
+ box_unc = [
+ torch.mean(
+ unc / wh[:, None, :].expand(-1, 2, 2).reshape(-1, 4), dim=-1)
+ if wh.numel() > 0 else unc for unc, wh in zip(box_unc, box_shape)
+ ]
+ return box_unc
+
+ @staticmethod
+ def aug_box(batch_data_samples, times, frac):
+ """Augment bboxes with jitter."""
+
+ def _aug_single(box):
+ box_scale = box[:, 2:4] - box[:, :2]
+ box_scale = (
+ box_scale.clamp(min=1)[:, None, :].expand(-1, 2,
+ 2).reshape(-1, 4))
+ aug_scale = box_scale * frac # [n,4]
+
+ offset = (
+ torch.randn(times, box.shape[0], 4, device=box.device) *
+ aug_scale[None, ...])
+ new_box = box.clone()[None, ...].expand(times, box.shape[0],
+ -1) + offset
+ return new_box
+
+ return [
+ _aug_single(data_samples.gt_instances.bboxes)
+ for data_samples in batch_data_samples
+ ]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/solo.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/solo.py
new file mode 100644
index 0000000000000000000000000000000000000000..6bf47ba24941e09fd795b241a3f6aa0b67ae3380
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/solo.py
@@ -0,0 +1,31 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage_instance_seg import SingleStageInstanceSegmentor
+
+
+@MODELS.register_module()
+class SOLO(SingleStageInstanceSegmentor):
+ """`SOLO: Segmenting Objects by Locations
+ `_
+
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: OptConfigType = None,
+ bbox_head: OptConfigType = None,
+ mask_head: OptConfigType = None,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None):
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ mask_head=mask_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/solov2.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/solov2.py
new file mode 100644
index 0000000000000000000000000000000000000000..1eefe4c532267be1480d13b8d73fc54bf694e81c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/solov2.py
@@ -0,0 +1,31 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage_instance_seg import SingleStageInstanceSegmentor
+
+
+@MODELS.register_module()
+class SOLOv2(SingleStageInstanceSegmentor):
+ """`SOLOv2: Dynamic and Fast Instance Segmentation
+ `_
+
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: OptConfigType = None,
+ bbox_head: OptConfigType = None,
+ mask_head: OptConfigType = None,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None):
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ mask_head=mask_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/sparse_rcnn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/sparse_rcnn.py
new file mode 100644
index 0000000000000000000000000000000000000000..75442a69e472953854ded9fc8c30ac4ab30535d3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/sparse_rcnn.py
@@ -0,0 +1,31 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .two_stage import TwoStageDetector
+
+
+@MODELS.register_module()
+class SparseRCNN(TwoStageDetector):
+ r"""Implementation of `Sparse R-CNN: End-to-End Object Detection with
+ Learnable Proposals `_"""
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: OptConfigType = None,
+ rpn_head: OptConfigType = None,
+ roi_head: OptConfigType = None,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ rpn_head=rpn_head,
+ roi_head=roi_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
+ assert self.with_rpn, 'Sparse R-CNN and QueryInst ' \
+ 'do not support external proposals'
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/tood.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/tood.py
new file mode 100644
index 0000000000000000000000000000000000000000..38720482c5451471f5a66a6cf689dbed6100c9fa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/tood.py
@@ -0,0 +1,42 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class TOOD(SingleStageDetector):
+ r"""Implementation of `TOOD: Task-aligned One-stage Object Detection.
+ `_
+
+ Args:
+ backbone (:obj:`ConfigDict` or dict): The backbone module.
+ neck (:obj:`ConfigDict` or dict): The neck module.
+ bbox_head (:obj:`ConfigDict` or dict): The bbox head module.
+ train_cfg (:obj:`ConfigDict` or dict, optional): The training config
+ of TOOD. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): The testing config
+ of TOOD. Defaults to None.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional): Config of
+ :class:`DetDataPreprocessor` to process the input data.
+ Defaults to None.
+ init_cfg (:obj:`ConfigDict` or dict, optional): the config to control
+ the initialization. Defaults to None.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/trident_faster_rcnn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/trident_faster_rcnn.py
new file mode 100644
index 0000000000000000000000000000000000000000..4244925beaebea820f836b41ab5463f5f499f4d0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/trident_faster_rcnn.py
@@ -0,0 +1,81 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .faster_rcnn import FasterRCNN
+
+
+@MODELS.register_module()
+class TridentFasterRCNN(FasterRCNN):
+ """Implementation of `TridentNet `_"""
+
+ def __init__(self,
+ backbone: ConfigType,
+ rpn_head: ConfigType,
+ roi_head: ConfigType,
+ train_cfg: ConfigType,
+ test_cfg: ConfigType,
+ neck: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ rpn_head=rpn_head,
+ roi_head=roi_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
+ assert self.backbone.num_branch == self.roi_head.num_branch
+ assert self.backbone.test_branch_idx == self.roi_head.test_branch_idx
+ self.num_branch = self.backbone.num_branch
+ self.test_branch_idx = self.backbone.test_branch_idx
+
+ def _forward(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> tuple:
+ """copy the ``batch_data_samples`` to fit multi-branch."""
+ num_branch = self.num_branch \
+ if self.training or self.test_branch_idx == -1 else 1
+ trident_data_samples = batch_data_samples * num_branch
+ return super()._forward(
+ batch_inputs=batch_inputs, batch_data_samples=trident_data_samples)
+
+ def loss(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> dict:
+ """copy the ``batch_data_samples`` to fit multi-branch."""
+ num_branch = self.num_branch \
+ if self.training or self.test_branch_idx == -1 else 1
+ trident_data_samples = batch_data_samples * num_branch
+ return super().loss(
+ batch_inputs=batch_inputs, batch_data_samples=trident_data_samples)
+
+ def predict(self,
+ batch_inputs: Tensor,
+ batch_data_samples: SampleList,
+ rescale: bool = True) -> SampleList:
+ """copy the ``batch_data_samples`` to fit multi-branch."""
+ num_branch = self.num_branch \
+ if self.training or self.test_branch_idx == -1 else 1
+ trident_data_samples = batch_data_samples * num_branch
+ return super().predict(
+ batch_inputs=batch_inputs,
+ batch_data_samples=trident_data_samples,
+ rescale=rescale)
+
+ # TODO need to refactor
+ def aug_test(self, imgs, img_metas, rescale=False):
+ """Test with augmentations.
+
+ If rescale is False, then returned bboxes and masks will fit the scale
+ of imgs[0].
+ """
+ x = self.extract_feats(imgs)
+ num_branch = (self.num_branch if self.test_branch_idx == -1 else 1)
+ trident_img_metas = [img_metas * num_branch for img_metas in img_metas]
+ proposal_list = self.rpn_head.aug_test_rpn(x, trident_img_metas)
+ return self.roi_head.aug_test(
+ x, proposal_list, img_metas, rescale=rescale)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/two_stage.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/two_stage.py
new file mode 100644
index 0000000000000000000000000000000000000000..4e83df9eb5ce837636e10c4592fe26a7edce1657
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/two_stage.py
@@ -0,0 +1,243 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import warnings
+from typing import List, Tuple, Union
+
+import torch
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .base import BaseDetector
+
+
+@MODELS.register_module()
+class TwoStageDetector(BaseDetector):
+ """Base class for two-stage detectors.
+
+ Two-stage detectors typically consisting of a region proposal network and a
+ task-specific regression head.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: OptConfigType = None,
+ rpn_head: OptConfigType = None,
+ roi_head: OptConfigType = None,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ data_preprocessor=data_preprocessor, init_cfg=init_cfg)
+ self.backbone = MODELS.build(backbone)
+
+ if neck is not None:
+ self.neck = MODELS.build(neck)
+
+ if rpn_head is not None:
+ rpn_train_cfg = train_cfg.rpn if train_cfg is not None else None
+ rpn_head_ = rpn_head.copy()
+ rpn_head_.update(train_cfg=rpn_train_cfg, test_cfg=test_cfg.rpn)
+ rpn_head_num_classes = rpn_head_.get('num_classes', None)
+ if rpn_head_num_classes is None:
+ rpn_head_.update(num_classes=1)
+ else:
+ if rpn_head_num_classes != 1:
+ warnings.warn(
+ 'The `num_classes` should be 1 in RPN, but get '
+ f'{rpn_head_num_classes}, please set '
+ 'rpn_head.num_classes = 1 in your config file.')
+ rpn_head_.update(num_classes=1)
+ self.rpn_head = MODELS.build(rpn_head_)
+
+ if roi_head is not None:
+ # update train and test cfg here for now
+ # TODO: refactor assigner & sampler
+ rcnn_train_cfg = train_cfg.rcnn if train_cfg is not None else None
+ roi_head.update(train_cfg=rcnn_train_cfg)
+ roi_head.update(test_cfg=test_cfg.rcnn)
+ self.roi_head = MODELS.build(roi_head)
+
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+
+ def _load_from_state_dict(self, state_dict: dict, prefix: str,
+ local_metadata: dict, strict: bool,
+ missing_keys: Union[List[str], str],
+ unexpected_keys: Union[List[str], str],
+ error_msgs: Union[List[str], str]) -> None:
+ """Exchange bbox_head key to rpn_head key when loading single-stage
+ weights into two-stage model."""
+ bbox_head_prefix = prefix + '.bbox_head' if prefix else 'bbox_head'
+ bbox_head_keys = [
+ k for k in state_dict.keys() if k.startswith(bbox_head_prefix)
+ ]
+ rpn_head_prefix = prefix + '.rpn_head' if prefix else 'rpn_head'
+ rpn_head_keys = [
+ k for k in state_dict.keys() if k.startswith(rpn_head_prefix)
+ ]
+ if len(bbox_head_keys) != 0 and len(rpn_head_keys) == 0:
+ for bbox_head_key in bbox_head_keys:
+ rpn_head_key = rpn_head_prefix + \
+ bbox_head_key[len(bbox_head_prefix):]
+ state_dict[rpn_head_key] = state_dict.pop(bbox_head_key)
+ super()._load_from_state_dict(state_dict, prefix, local_metadata,
+ strict, missing_keys, unexpected_keys,
+ error_msgs)
+
+ @property
+ def with_rpn(self) -> bool:
+ """bool: whether the detector has RPN"""
+ return hasattr(self, 'rpn_head') and self.rpn_head is not None
+
+ @property
+ def with_roi_head(self) -> bool:
+ """bool: whether the detector has a RoI head"""
+ return hasattr(self, 'roi_head') and self.roi_head is not None
+
+ def extract_feat(self, batch_inputs: Tensor) -> Tuple[Tensor]:
+ """Extract features.
+
+ Args:
+ batch_inputs (Tensor): Image tensor with shape (N, C, H ,W).
+
+ Returns:
+ tuple[Tensor]: Multi-level features that may have
+ different resolutions.
+ """
+ x = self.backbone(batch_inputs)
+ if self.with_neck:
+ x = self.neck(x)
+ return x
+
+ def _forward(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> tuple:
+ """Network forward process. Usually includes backbone, neck and head
+ forward without any post-processing.
+
+ Args:
+ batch_inputs (Tensor): Inputs with shape (N, C, H, W).
+ batch_data_samples (list[:obj:`DetDataSample`]): Each item contains
+ the meta information of each image and corresponding
+ annotations.
+
+ Returns:
+ tuple: A tuple of features from ``rpn_head`` and ``roi_head``
+ forward.
+ """
+ results = ()
+ x = self.extract_feat(batch_inputs)
+
+ if self.with_rpn:
+ rpn_results_list = self.rpn_head.predict(
+ x, batch_data_samples, rescale=False)
+ else:
+ assert batch_data_samples[0].get('proposals', None) is not None
+ rpn_results_list = [
+ data_sample.proposals for data_sample in batch_data_samples
+ ]
+ roi_outs = self.roi_head.forward(x, rpn_results_list,
+ batch_data_samples)
+ results = results + (roi_outs, )
+ return results
+
+ def loss(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> dict:
+ """Calculate losses from a batch of inputs and data samples.
+
+ Args:
+ batch_inputs (Tensor): Input images of shape (N, C, H, W).
+ These should usually be mean centered and std scaled.
+ batch_data_samples (List[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict: A dictionary of loss components
+ """
+ x = self.extract_feat(batch_inputs)
+
+ losses = dict()
+
+ # RPN forward and loss
+ if self.with_rpn:
+ proposal_cfg = self.train_cfg.get('rpn_proposal',
+ self.test_cfg.rpn)
+ rpn_data_samples = copy.deepcopy(batch_data_samples)
+ # set cat_id of gt_labels to 0 in RPN
+ for data_sample in rpn_data_samples:
+ data_sample.gt_instances.labels = \
+ torch.zeros_like(data_sample.gt_instances.labels)
+
+ rpn_losses, rpn_results_list = self.rpn_head.loss_and_predict(
+ x, rpn_data_samples, proposal_cfg=proposal_cfg)
+ # avoid get same name with roi_head loss
+ keys = rpn_losses.keys()
+ for key in list(keys):
+ if 'loss' in key and 'rpn' not in key:
+ rpn_losses[f'rpn_{key}'] = rpn_losses.pop(key)
+ losses.update(rpn_losses)
+ else:
+ assert batch_data_samples[0].get('proposals', None) is not None
+ # use pre-defined proposals in InstanceData for the second stage
+ # to extract ROI features.
+ rpn_results_list = [
+ data_sample.proposals for data_sample in batch_data_samples
+ ]
+
+ roi_losses = self.roi_head.loss(x, rpn_results_list,
+ batch_data_samples)
+ losses.update(roi_losses)
+
+ return losses
+
+ def predict(self,
+ batch_inputs: Tensor,
+ batch_data_samples: SampleList,
+ rescale: bool = True) -> SampleList:
+ """Predict results from a batch of inputs and data samples with post-
+ processing.
+
+ Args:
+ batch_inputs (Tensor): Inputs with shape (N, C, H, W).
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool): Whether to rescale the results.
+ Defaults to True.
+
+ Returns:
+ list[:obj:`DetDataSample`]: Return the detection results of the
+ input images. The returns value is DetDataSample,
+ which usually contain 'pred_instances'. And the
+ ``pred_instances`` usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+ """
+
+ assert self.with_bbox, 'Bbox head must be implemented.'
+ x = self.extract_feat(batch_inputs)
+
+ # If there are no pre-defined proposals, use RPN to get proposals
+ if batch_data_samples[0].get('proposals', None) is None:
+ rpn_results_list = self.rpn_head.predict(
+ x, batch_data_samples, rescale=False)
+ else:
+ rpn_results_list = [
+ data_sample.proposals for data_sample in batch_data_samples
+ ]
+
+ results_list = self.roi_head.predict(
+ x, rpn_results_list, batch_data_samples, rescale=rescale)
+
+ batch_data_samples = self.add_pred_to_datasample(
+ batch_data_samples, results_list)
+ return batch_data_samples
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/vfnet.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/vfnet.py
new file mode 100644
index 0000000000000000000000000000000000000000..a695513faa7d37756d7716cbca0e457060400518
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/vfnet.py
@@ -0,0 +1,42 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class VFNet(SingleStageDetector):
+ """Implementation of `VarifocalNet
+ (VFNet).`_
+
+ Args:
+ backbone (:obj:`ConfigDict` or dict): The backbone module.
+ neck (:obj:`ConfigDict` or dict): The neck module.
+ bbox_head (:obj:`ConfigDict` or dict): The bbox head module.
+ train_cfg (:obj:`ConfigDict` or dict, optional): The training config
+ of VFNet. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): The testing config
+ of VFNet. Defaults to None.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional): Config of
+ :class:`DetDataPreprocessor` to process the input data.
+ Defaults to None.
+ init_cfg (:obj:`ConfigDict` or dict, optional): the config to control
+ the initialization. Defaults to None.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/yolact.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/yolact.py
new file mode 100644
index 0000000000000000000000000000000000000000..f15fb7b70263b0c4018751067771b1365af96f67
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/yolact.py
@@ -0,0 +1,28 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage_instance_seg import SingleStageInstanceSegmentor
+
+
+@MODELS.register_module()
+class YOLACT(SingleStageInstanceSegmentor):
+ """Implementation of `YOLACT `_"""
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ mask_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ mask_head=mask_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/yolo.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/yolo.py
new file mode 100644
index 0000000000000000000000000000000000000000..5cb9a9cd250a2c26af22032b1ed4bb5a7a8af605
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/yolo.py
@@ -0,0 +1,45 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+# Copyright (c) 2019 Western Digital Corporation or its affiliates.
+
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class YOLOV3(SingleStageDetector):
+ r"""Implementation of `Yolov3: An incremental improvement
+ `_
+
+ Args:
+ backbone (:obj:`ConfigDict` or dict): The backbone module.
+ neck (:obj:`ConfigDict` or dict): The neck module.
+ bbox_head (:obj:`ConfigDict` or dict): The bbox head module.
+ train_cfg (:obj:`ConfigDict` or dict, optional): The training config
+ of YOLOX. Default: None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): The testing config
+ of YOLOX. Default: None.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional):
+ Model preprocessing config for processing the input data.
+ it usually includes ``to_rgb``, ``pad_size_divisor``,
+ ``pad_value``, ``mean`` and ``std``. Defaults to None.
+ init_cfg (:obj:`ConfigDict` or dict, optional): the config to control
+ the initialization. Defaults to None.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/yolof.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/yolof.py
new file mode 100644
index 0000000000000000000000000000000000000000..c6d98b9134a7f422fa7ea1f1a1e0d548d36603e8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/yolof.py
@@ -0,0 +1,43 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class YOLOF(SingleStageDetector):
+ r"""Implementation of `You Only Look One-level Feature
+ `_
+
+ Args:
+ backbone (:obj:`ConfigDict` or dict): The backbone module.
+ neck (:obj:`ConfigDict` or dict): The neck module.
+ bbox_head (:obj:`ConfigDict` or dict): The bbox head module.
+ train_cfg (:obj:`ConfigDict` or dict, optional): The training config
+ of YOLOF. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): The testing config
+ of YOLOF. Defaults to None.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional):
+ Model preprocessing config for processing the input data.
+ it usually includes ``to_rgb``, ``pad_size_divisor``,
+ ``pad_value``, ``mean`` and ``std``. Defaults to None.
+ init_cfg (:obj:`ConfigDict` or dict, optional): the config to control
+ the initialization. Defaults to None.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/yolox.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/yolox.py
new file mode 100644
index 0000000000000000000000000000000000000000..df9190c93f7b043910fbce3bd5ee8dc0ef7b5f68
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/detectors/yolox.py
@@ -0,0 +1,43 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .single_stage import SingleStageDetector
+
+
+@MODELS.register_module()
+class YOLOX(SingleStageDetector):
+ r"""Implementation of `YOLOX: Exceeding YOLO Series in 2021
+ `_
+
+ Args:
+ backbone (:obj:`ConfigDict` or dict): The backbone config.
+ neck (:obj:`ConfigDict` or dict): The neck config.
+ bbox_head (:obj:`ConfigDict` or dict): The bbox head config.
+ train_cfg (:obj:`ConfigDict` or dict, optional): The training config
+ of YOLOX. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): The testing config
+ of YOLOX. Defaults to None.
+ data_preprocessor (:obj:`ConfigDict` or dict, optional): Config of
+ :class:`DetDataPreprocessor` to process the input data.
+ Defaults to None.
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or
+ list[dict], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ backbone: ConfigType,
+ neck: ConfigType,
+ bbox_head: ConfigType,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ backbone=backbone,
+ neck=neck,
+ bbox_head=bbox_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ data_preprocessor=data_preprocessor,
+ init_cfg=init_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/language_models/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/language_models/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..70f1a22c7c01624ba3235f1737f8aea1e26a19fe
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/language_models/__init__.py
@@ -0,0 +1,4 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .bert import BertModel
+
+__all__ = ['BertModel']
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/language_models/bert.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/language_models/bert.py
new file mode 100644
index 0000000000000000000000000000000000000000..efb0f46bad6eb0734a324c32a7b05f2795604265
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/language_models/bert.py
@@ -0,0 +1,231 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from collections import OrderedDict
+from typing import Sequence
+
+import torch
+from mmengine.model import BaseModel
+from torch import nn
+
+try:
+ from transformers import AutoTokenizer, BertConfig
+ from transformers import BertModel as HFBertModel
+except ImportError:
+ AutoTokenizer = None
+ HFBertModel = None
+
+from mmdet.registry import MODELS
+
+
+def generate_masks_with_special_tokens_and_transfer_map(
+ tokenized, special_tokens_list):
+ """Generate attention mask between each pair of special tokens.
+
+ Only token pairs in between two special tokens are attended to
+ and thus the attention mask for these pairs is positive.
+
+ Args:
+ input_ids (torch.Tensor): input ids. Shape: [bs, num_token]
+ special_tokens_mask (list): special tokens mask.
+
+ Returns:
+ Tuple(Tensor, Tensor):
+ - attention_mask is the attention mask between each tokens.
+ Only token pairs in between two special tokens are positive.
+ Shape: [bs, num_token, num_token].
+ - position_ids is the position id of tokens within each valid sentence.
+ The id starts from 0 whenenver a special token is encountered.
+ Shape: [bs, num_token]
+ """
+ input_ids = tokenized['input_ids']
+ bs, num_token = input_ids.shape
+ # special_tokens_mask:
+ # bs, num_token. 1 for special tokens. 0 for normal tokens
+ special_tokens_mask = torch.zeros((bs, num_token),
+ device=input_ids.device).bool()
+
+ for special_token in special_tokens_list:
+ special_tokens_mask |= input_ids == special_token
+
+ # idxs: each row is a list of indices of special tokens
+ idxs = torch.nonzero(special_tokens_mask)
+
+ # generate attention mask and positional ids
+ attention_mask = (
+ torch.eye(num_token,
+ device=input_ids.device).bool().unsqueeze(0).repeat(
+ bs, 1, 1))
+ position_ids = torch.zeros((bs, num_token), device=input_ids.device)
+ previous_col = 0
+ for i in range(idxs.shape[0]):
+ row, col = idxs[i]
+ if (col == 0) or (col == num_token - 1):
+ attention_mask[row, col, col] = True
+ position_ids[row, col] = 0
+ else:
+ attention_mask[row, previous_col + 1:col + 1,
+ previous_col + 1:col + 1] = True
+ position_ids[row, previous_col + 1:col + 1] = torch.arange(
+ 0, col - previous_col, device=input_ids.device)
+ previous_col = col
+
+ return attention_mask, position_ids.to(torch.long)
+
+
+@MODELS.register_module()
+class BertModel(BaseModel):
+ """BERT model for language embedding only encoder.
+
+ Args:
+ name (str, optional): name of the pretrained BERT model from
+ HuggingFace. Defaults to bert-base-uncased.
+ max_tokens (int, optional): maximum number of tokens to be
+ used for BERT. Defaults to 256.
+ pad_to_max (bool, optional): whether to pad the tokens to max_tokens.
+ Defaults to True.
+ use_sub_sentence_represent (bool, optional): whether to use sub
+ sentence represent introduced in `Grounding DINO
+ `. Defaults to False.
+ special_tokens_list (list, optional): special tokens used to split
+ subsentence. It cannot be None when `use_sub_sentence_represent`
+ is True. Defaults to None.
+ add_pooling_layer (bool, optional): whether to adding pooling
+ layer in bert encoder. Defaults to False.
+ num_layers_of_embedded (int, optional): number of layers of
+ the embedded model. Defaults to 1.
+ use_checkpoint (bool, optional): whether to use gradient checkpointing.
+ Defaults to False.
+ """
+
+ def __init__(self,
+ name: str = 'bert-base-uncased',
+ max_tokens: int = 256,
+ pad_to_max: bool = True,
+ use_sub_sentence_represent: bool = False,
+ special_tokens_list: list = None,
+ add_pooling_layer: bool = False,
+ num_layers_of_embedded: int = 1,
+ use_checkpoint: bool = False,
+ **kwargs) -> None:
+
+ super().__init__(**kwargs)
+ self.max_tokens = max_tokens
+ self.pad_to_max = pad_to_max
+
+ if AutoTokenizer is None:
+ raise RuntimeError(
+ 'transformers is not installed, please install it by: '
+ 'pip install transformers.')
+
+ self.tokenizer = AutoTokenizer.from_pretrained(name)
+ self.language_backbone = nn.Sequential(
+ OrderedDict([('body',
+ BertEncoder(
+ name,
+ add_pooling_layer=add_pooling_layer,
+ num_layers_of_embedded=num_layers_of_embedded,
+ use_checkpoint=use_checkpoint))]))
+
+ self.use_sub_sentence_represent = use_sub_sentence_represent
+ if self.use_sub_sentence_represent:
+ assert special_tokens_list is not None, \
+ 'special_tokens should not be None \
+ if use_sub_sentence_represent is True'
+
+ self.special_tokens = self.tokenizer.convert_tokens_to_ids(
+ special_tokens_list)
+
+ def forward(self, captions: Sequence[str], **kwargs) -> dict:
+ """Forward function."""
+ device = next(self.language_backbone.parameters()).device
+ tokenized = self.tokenizer.batch_encode_plus(
+ captions,
+ max_length=self.max_tokens,
+ padding='max_length' if self.pad_to_max else 'longest',
+ return_special_tokens_mask=True,
+ return_tensors='pt',
+ truncation=True).to(device)
+ input_ids = tokenized.input_ids
+ if self.use_sub_sentence_represent:
+ attention_mask, position_ids = \
+ generate_masks_with_special_tokens_and_transfer_map(
+ tokenized, self.special_tokens)
+ token_type_ids = tokenized['token_type_ids']
+
+ else:
+ attention_mask = tokenized.attention_mask
+ position_ids = None
+ token_type_ids = None
+
+ tokenizer_input = {
+ 'input_ids': input_ids,
+ 'attention_mask': attention_mask,
+ 'position_ids': position_ids,
+ 'token_type_ids': token_type_ids
+ }
+ language_dict_features = self.language_backbone(tokenizer_input)
+ if self.use_sub_sentence_represent:
+ language_dict_features['position_ids'] = position_ids
+ language_dict_features[
+ 'text_token_mask'] = tokenized.attention_mask.bool()
+ return language_dict_features
+
+
+class BertEncoder(nn.Module):
+ """BERT encoder for language embedding.
+
+ Args:
+ name (str): name of the pretrained BERT model from HuggingFace.
+ Defaults to bert-base-uncased.
+ add_pooling_layer (bool): whether to add a pooling layer.
+ num_layers_of_embedded (int): number of layers of the embedded model.
+ Defaults to 1.
+ use_checkpoint (bool): whether to use gradient checkpointing.
+ Defaults to False.
+ """
+
+ def __init__(self,
+ name: str,
+ add_pooling_layer: bool = False,
+ num_layers_of_embedded: int = 1,
+ use_checkpoint: bool = False):
+ super().__init__()
+ if BertConfig is None:
+ raise RuntimeError(
+ 'transformers is not installed, please install it by: '
+ 'pip install transformers.')
+ config = BertConfig.from_pretrained(name)
+ config.gradient_checkpointing = use_checkpoint
+ # only encoder
+ self.model = HFBertModel.from_pretrained(
+ name, add_pooling_layer=add_pooling_layer, config=config)
+ self.language_dim = config.hidden_size
+ self.num_layers_of_embedded = num_layers_of_embedded
+
+ def forward(self, x) -> dict:
+ mask = x['attention_mask']
+
+ outputs = self.model(
+ input_ids=x['input_ids'],
+ attention_mask=mask,
+ position_ids=x['position_ids'],
+ token_type_ids=x['token_type_ids'],
+ output_hidden_states=True,
+ )
+
+ # outputs has 13 layers, 1 input layer and 12 hidden layers
+ encoded_layers = outputs.hidden_states[1:]
+ features = torch.stack(encoded_layers[-self.num_layers_of_embedded:],
+ 1).mean(1)
+ # language embedding has shape [len(phrase), seq_len, language_dim]
+ features = features / self.num_layers_of_embedded
+ if mask.dim() == 2:
+ embedded = features * mask.unsqueeze(-1).float()
+ else:
+ embedded = features
+
+ results = {
+ 'embedded': embedded,
+ 'masks': mask,
+ 'hidden': encoded_layers[-1]
+ }
+ return results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..e3c41f64d11bbdb7f2c8e128a2e28b2845159589
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/__init__.py
@@ -0,0 +1,65 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .activations import SiLU
+from .bbox_nms import fast_nms, multiclass_nms
+from .brick_wrappers import (AdaptiveAvgPool2d, FrozenBatchNorm2d,
+ adaptive_avg_pool2d)
+from .conv_upsample import ConvUpsample
+from .csp_layer import CSPLayer
+from .dropblock import DropBlock
+from .ema import ExpMomentumEMA
+from .inverted_residual import InvertedResidual
+from .matrix_nms import mask_matrix_nms
+from .msdeformattn_pixel_decoder import MSDeformAttnPixelDecoder
+from .normed_predictor import NormedConv2d, NormedLinear
+from .pixel_decoder import PixelDecoder, TransformerEncoderPixelDecoder
+from .positional_encoding import (LearnedPositionalEncoding,
+ SinePositionalEncoding,
+ SinePositionalEncoding3D)
+from .res_layer import ResLayer, SimplifiedBasicBlock
+from .se_layer import ChannelAttention, DyReLU, SELayer
+# yapf: disable
+from .transformer import (MLP, AdaptivePadding, CdnQueryGenerator,
+ ConditionalAttention,
+ ConditionalDetrTransformerDecoder,
+ ConditionalDetrTransformerDecoderLayer,
+ DABDetrTransformerDecoder,
+ DABDetrTransformerDecoderLayer,
+ DABDetrTransformerEncoder, DDQTransformerDecoder,
+ DeformableDetrTransformerDecoder,
+ DeformableDetrTransformerDecoderLayer,
+ DeformableDetrTransformerEncoder,
+ DeformableDetrTransformerEncoderLayer,
+ DetrTransformerDecoder, DetrTransformerDecoderLayer,
+ DetrTransformerEncoder, DetrTransformerEncoderLayer,
+ DinoTransformerDecoder, DynamicConv,
+ Mask2FormerTransformerDecoder,
+ Mask2FormerTransformerDecoderLayer,
+ Mask2FormerTransformerEncoder, PatchEmbed,
+ PatchMerging, coordinate_to_encoding,
+ inverse_sigmoid, nchw_to_nlc, nlc_to_nchw)
+
+# yapf: enable
+
+__all__ = [
+ 'fast_nms', 'multiclass_nms', 'mask_matrix_nms', 'DropBlock',
+ 'PixelDecoder', 'TransformerEncoderPixelDecoder',
+ 'MSDeformAttnPixelDecoder', 'ResLayer', 'PatchMerging',
+ 'SinePositionalEncoding', 'LearnedPositionalEncoding', 'DynamicConv',
+ 'SimplifiedBasicBlock', 'NormedLinear', 'NormedConv2d', 'InvertedResidual',
+ 'SELayer', 'ConvUpsample', 'CSPLayer', 'adaptive_avg_pool2d',
+ 'AdaptiveAvgPool2d', 'PatchEmbed', 'nchw_to_nlc', 'nlc_to_nchw', 'DyReLU',
+ 'ExpMomentumEMA', 'inverse_sigmoid', 'ChannelAttention', 'SiLU', 'MLP',
+ 'DetrTransformerEncoderLayer', 'DetrTransformerDecoderLayer',
+ 'DetrTransformerEncoder', 'DetrTransformerDecoder',
+ 'DeformableDetrTransformerEncoder', 'DeformableDetrTransformerDecoder',
+ 'DeformableDetrTransformerEncoderLayer',
+ 'DeformableDetrTransformerDecoderLayer', 'AdaptivePadding',
+ 'coordinate_to_encoding', 'ConditionalAttention',
+ 'DABDetrTransformerDecoderLayer', 'DABDetrTransformerDecoder',
+ 'DABDetrTransformerEncoder', 'DDQTransformerDecoder',
+ 'ConditionalDetrTransformerDecoder',
+ 'ConditionalDetrTransformerDecoderLayer', 'DinoTransformerDecoder',
+ 'CdnQueryGenerator', 'Mask2FormerTransformerEncoder',
+ 'Mask2FormerTransformerDecoderLayer', 'Mask2FormerTransformerDecoder',
+ 'SinePositionalEncoding3D', 'FrozenBatchNorm2d'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/activations.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/activations.py
new file mode 100644
index 0000000000000000000000000000000000000000..9e73ef42180ccd3dddb4bcca224c0b4eb5da807c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/activations.py
@@ -0,0 +1,22 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+import torch.nn as nn
+from mmengine.utils import digit_version
+
+from mmdet.registry import MODELS
+
+if digit_version(torch.__version__) >= digit_version('1.7.0'):
+ from torch.nn import SiLU
+else:
+
+ class SiLU(nn.Module):
+ """Sigmoid Weighted Liner Unit."""
+
+ def __init__(self, inplace=True):
+ super().__init__()
+
+ def forward(self, inputs) -> torch.Tensor:
+ return inputs * torch.sigmoid(inputs)
+
+
+MODELS.register_module(module=SiLU, name='SiLU')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/bbox_nms.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/bbox_nms.py
new file mode 100644
index 0000000000000000000000000000000000000000..fd67a45f60ca98c354e095127ab7dbb9653deca5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/bbox_nms.py
@@ -0,0 +1,184 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Tuple, Union
+
+import torch
+from mmcv.ops.nms import batched_nms
+from torch import Tensor
+
+from mmdet.structures.bbox import bbox_overlaps
+from mmdet.utils import ConfigType
+
+
+def multiclass_nms(
+ multi_bboxes: Tensor,
+ multi_scores: Tensor,
+ score_thr: float,
+ nms_cfg: ConfigType,
+ max_num: int = -1,
+ score_factors: Optional[Tensor] = None,
+ return_inds: bool = False,
+ box_dim: int = 4
+) -> Union[Tuple[Tensor, Tensor, Tensor], Tuple[Tensor, Tensor]]:
+ """NMS for multi-class bboxes.
+
+ Args:
+ multi_bboxes (Tensor): shape (n, #class*4) or (n, 4)
+ multi_scores (Tensor): shape (n, #class), where the last column
+ contains scores of the background class, but this will be ignored.
+ score_thr (float): bbox threshold, bboxes with scores lower than it
+ will not be considered.
+ nms_cfg (Union[:obj:`ConfigDict`, dict]): a dict that contains
+ the arguments of nms operations.
+ max_num (int, optional): if there are more than max_num bboxes after
+ NMS, only top max_num will be kept. Default to -1.
+ score_factors (Tensor, optional): The factors multiplied to scores
+ before applying NMS. Default to None.
+ return_inds (bool, optional): Whether return the indices of kept
+ bboxes. Default to False.
+ box_dim (int): The dimension of boxes. Defaults to 4.
+
+ Returns:
+ Union[Tuple[Tensor, Tensor, Tensor], Tuple[Tensor, Tensor]]:
+ (dets, labels, indices (optional)), tensors of shape (k, 5),
+ (k), and (k). Dets are boxes with scores. Labels are 0-based.
+ """
+ num_classes = multi_scores.size(1) - 1
+ # exclude background category
+ if multi_bboxes.shape[1] > box_dim:
+ bboxes = multi_bboxes.view(multi_scores.size(0), -1, box_dim)
+ else:
+ bboxes = multi_bboxes[:, None].expand(
+ multi_scores.size(0), num_classes, box_dim)
+
+ scores = multi_scores[:, :-1]
+
+ labels = torch.arange(num_classes, dtype=torch.long, device=scores.device)
+ labels = labels.view(1, -1).expand_as(scores)
+
+ bboxes = bboxes.reshape(-1, box_dim)
+ scores = scores.reshape(-1)
+ labels = labels.reshape(-1)
+
+ if not torch.onnx.is_in_onnx_export():
+ # NonZero not supported in TensorRT
+ # remove low scoring boxes
+ valid_mask = scores > score_thr
+ # multiply score_factor after threshold to preserve more bboxes, improve
+ # mAP by 1% for YOLOv3
+ if score_factors is not None:
+ # expand the shape to match original shape of score
+ score_factors = score_factors.view(-1, 1).expand(
+ multi_scores.size(0), num_classes)
+ score_factors = score_factors.reshape(-1)
+ scores = scores * score_factors
+
+ if not torch.onnx.is_in_onnx_export():
+ # NonZero not supported in TensorRT
+ inds = valid_mask.nonzero(as_tuple=False).squeeze(1)
+ bboxes, scores, labels = bboxes[inds], scores[inds], labels[inds]
+ else:
+ # TensorRT NMS plugin has invalid output filled with -1
+ # add dummy data to make detection output correct.
+ bboxes = torch.cat([bboxes, bboxes.new_zeros(1, box_dim)], dim=0)
+ scores = torch.cat([scores, scores.new_zeros(1)], dim=0)
+ labels = torch.cat([labels, labels.new_zeros(1)], dim=0)
+
+ if bboxes.numel() == 0:
+ if torch.onnx.is_in_onnx_export():
+ raise RuntimeError('[ONNX Error] Can not record NMS '
+ 'as it has not been executed this time')
+ dets = torch.cat([bboxes, scores[:, None]], -1)
+ if return_inds:
+ return dets, labels, inds
+ else:
+ return dets, labels
+
+ dets, keep = batched_nms(bboxes, scores, labels, nms_cfg)
+
+ if max_num > 0:
+ dets = dets[:max_num]
+ keep = keep[:max_num]
+
+ if return_inds:
+ return dets, labels[keep], inds[keep]
+ else:
+ return dets, labels[keep]
+
+
+def fast_nms(
+ multi_bboxes: Tensor,
+ multi_scores: Tensor,
+ multi_coeffs: Tensor,
+ score_thr: float,
+ iou_thr: float,
+ top_k: int,
+ max_num: int = -1
+) -> Union[Tuple[Tensor, Tensor, Tensor], Tuple[Tensor, Tensor]]:
+ """Fast NMS in `YOLACT `_.
+
+ Fast NMS allows already-removed detections to suppress other detections so
+ that every instance can be decided to be kept or discarded in parallel,
+ which is not possible in traditional NMS. This relaxation allows us to
+ implement Fast NMS entirely in standard GPU-accelerated matrix operations.
+
+ Args:
+ multi_bboxes (Tensor): shape (n, #class*4) or (n, 4)
+ multi_scores (Tensor): shape (n, #class+1), where the last column
+ contains scores of the background class, but this will be ignored.
+ multi_coeffs (Tensor): shape (n, #class*coeffs_dim).
+ score_thr (float): bbox threshold, bboxes with scores lower than it
+ will not be considered.
+ iou_thr (float): IoU threshold to be considered as conflicted.
+ top_k (int): if there are more than top_k bboxes before NMS,
+ only top top_k will be kept.
+ max_num (int): if there are more than max_num bboxes after NMS,
+ only top max_num will be kept. If -1, keep all the bboxes.
+ Default: -1.
+
+ Returns:
+ Union[Tuple[Tensor, Tensor, Tensor], Tuple[Tensor, Tensor]]:
+ (dets, labels, coefficients), tensors of shape (k, 5), (k, 1),
+ and (k, coeffs_dim). Dets are boxes with scores.
+ Labels are 0-based.
+ """
+
+ scores = multi_scores[:, :-1].t() # [#class, n]
+ scores, idx = scores.sort(1, descending=True)
+
+ idx = idx[:, :top_k].contiguous()
+ scores = scores[:, :top_k] # [#class, topk]
+ num_classes, num_dets = idx.size()
+ boxes = multi_bboxes[idx.view(-1), :].view(num_classes, num_dets, 4)
+ coeffs = multi_coeffs[idx.view(-1), :].view(num_classes, num_dets, -1)
+
+ iou = bbox_overlaps(boxes, boxes) # [#class, topk, topk]
+ iou.triu_(diagonal=1)
+ iou_max, _ = iou.max(dim=1)
+
+ # Now just filter out the ones higher than the threshold
+ keep = iou_max <= iou_thr
+
+ # Second thresholding introduces 0.2 mAP gain at negligible time cost
+ keep *= scores > score_thr
+
+ # Assign each kept detection to its corresponding class
+ classes = torch.arange(
+ num_classes, device=boxes.device)[:, None].expand_as(keep)
+ classes = classes[keep]
+
+ boxes = boxes[keep]
+ coeffs = coeffs[keep]
+ scores = scores[keep]
+
+ # Only keep the top max_num highest scores across all classes
+ scores, idx = scores.sort(0, descending=True)
+ if max_num > 0:
+ idx = idx[:max_num]
+ scores = scores[:max_num]
+
+ classes = classes[idx]
+ boxes = boxes[idx]
+ coeffs = coeffs[idx]
+
+ cls_dets = torch.cat([boxes, scores[:, None]], dim=1)
+ return cls_dets, classes, coeffs
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/brick_wrappers.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/brick_wrappers.py
new file mode 100644
index 0000000000000000000000000000000000000000..5ecb8499de329132561dfedb8f55c36080787b31
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/brick_wrappers.py
@@ -0,0 +1,138 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn.bricks.wrappers import NewEmptyTensorOp, obsolete_torch_version
+
+from mmdet.registry import MODELS
+
+if torch.__version__ == 'parrots':
+ TORCH_VERSION = torch.__version__
+else:
+ # torch.__version__ could be 1.3.1+cu92, we only need the first two
+ # for comparison
+ TORCH_VERSION = tuple(int(x) for x in torch.__version__.split('.')[:2])
+
+
+def adaptive_avg_pool2d(input, output_size):
+ """Handle empty batch dimension to adaptive_avg_pool2d.
+
+ Args:
+ input (tensor): 4D tensor.
+ output_size (int, tuple[int,int]): the target output size.
+ """
+ if input.numel() == 0 and obsolete_torch_version(TORCH_VERSION, (1, 9)):
+ if isinstance(output_size, int):
+ output_size = [output_size, output_size]
+ output_size = [*input.shape[:2], *output_size]
+ empty = NewEmptyTensorOp.apply(input, output_size)
+ return empty
+ else:
+ return F.adaptive_avg_pool2d(input, output_size)
+
+
+class AdaptiveAvgPool2d(nn.AdaptiveAvgPool2d):
+ """Handle empty batch dimension to AdaptiveAvgPool2d."""
+
+ def forward(self, x):
+ # PyTorch 1.9 does not support empty tensor inference yet
+ if x.numel() == 0 and obsolete_torch_version(TORCH_VERSION, (1, 9)):
+ output_size = self.output_size
+ if isinstance(output_size, int):
+ output_size = [output_size, output_size]
+ else:
+ output_size = [
+ v if v is not None else d
+ for v, d in zip(output_size,
+ x.size()[-2:])
+ ]
+ output_size = [*x.shape[:2], *output_size]
+ empty = NewEmptyTensorOp.apply(x, output_size)
+ return empty
+
+ return super().forward(x)
+
+
+# Modified from
+# https://github.com/facebookresearch/detectron2/blob/main/detectron2/layers/batch_norm.py#L13 # noqa
+@MODELS.register_module('FrozenBN')
+class FrozenBatchNorm2d(nn.Module):
+ """BatchNorm2d where the batch statistics and the affine parameters are
+ fixed.
+
+ It contains non-trainable buffers called
+ "weight" and "bias", "running_mean", "running_var",
+ initialized to perform identity transformation.
+ Args:
+ num_features (int): :math:`C` from an expected input of size
+ :math:`(N, C, H, W)`.
+ eps (float): a value added to the denominator for numerical stability.
+ Default: 1e-5
+ """
+
+ def __init__(self, num_features, eps=1e-5, **kwargs):
+ super().__init__()
+ self.num_features = num_features
+ self.eps = eps
+ self.register_buffer('weight', torch.ones(num_features))
+ self.register_buffer('bias', torch.zeros(num_features))
+ self.register_buffer('running_mean', torch.zeros(num_features))
+ self.register_buffer('running_var', torch.ones(num_features) - eps)
+
+ def forward(self, x):
+ if x.requires_grad:
+ # When gradients are needed, F.batch_norm will use extra memory
+ # because its backward op computes gradients for weight/bias
+ # as well.
+ scale = self.weight * (self.running_var + self.eps).rsqrt()
+ bias = self.bias - self.running_mean * scale
+ scale = scale.reshape(1, -1, 1, 1)
+ bias = bias.reshape(1, -1, 1, 1)
+ out_dtype = x.dtype # may be half
+ return x * scale.to(out_dtype) + bias.to(out_dtype)
+ else:
+ # When gradients are not needed, F.batch_norm is a single fused op
+ # and provide more optimization opportunities.
+ return F.batch_norm(
+ x,
+ self.running_mean,
+ self.running_var,
+ self.weight,
+ self.bias,
+ training=False,
+ eps=self.eps,
+ )
+
+ def __repr__(self):
+ return 'FrozenBatchNorm2d(num_features={}, eps={})'.format(
+ self.num_features, self.eps)
+
+ @classmethod
+ def convert_frozen_batchnorm(cls, module):
+ """Convert all BatchNorm/SyncBatchNorm in module into FrozenBatchNorm.
+
+ Args:
+ module (torch.nn.Module):
+ Returns:
+ If module is BatchNorm/SyncBatchNorm, returns a new module.
+ Otherwise, in-place convert module and return it.
+ Similar to convert_sync_batchnorm in
+ https://github.com/pytorch/pytorch/blob/master/torch/nn/modules/batchnorm.py
+ """
+ bn_module = nn.modules.batchnorm
+ bn_module = (bn_module.BatchNorm2d, bn_module.SyncBatchNorm)
+ res = module
+ if isinstance(module, bn_module):
+ res = cls(module.num_features)
+ if module.affine:
+ res.weight.data = module.weight.data.clone().detach()
+ res.bias.data = module.bias.data.clone().detach()
+ res.running_mean.data = module.running_mean.data
+ res.running_var.data = module.running_var.data
+ res.eps = module.eps
+ else:
+ for name, child in module.named_children():
+ new_child = cls.convert_frozen_batchnorm(child)
+ if new_child is not child:
+ res.add_module(name, new_child)
+ return res
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/conv_upsample.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/conv_upsample.py
new file mode 100644
index 0000000000000000000000000000000000000000..32505875a2162330ed7d00455f088d08d94f679e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/conv_upsample.py
@@ -0,0 +1,67 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule
+from mmengine.model import BaseModule, ModuleList
+
+
+class ConvUpsample(BaseModule):
+ """ConvUpsample performs 2x upsampling after Conv.
+
+ There are several `ConvModule` layers. In the first few layers, upsampling
+ will be applied after each layer of convolution. The number of upsampling
+ must be no more than the number of ConvModule layers.
+
+ Args:
+ in_channels (int): Number of channels in the input feature map.
+ inner_channels (int): Number of channels produced by the convolution.
+ num_layers (int): Number of convolution layers.
+ num_upsample (int | optional): Number of upsampling layer. Must be no
+ more than num_layers. Upsampling will be applied after the first
+ ``num_upsample`` layers of convolution. Default: ``num_layers``.
+ conv_cfg (dict): Config dict for convolution layer. Default: None,
+ which means using conv2d.
+ norm_cfg (dict): Config dict for normalization layer. Default: None.
+ init_cfg (dict): Config dict for initialization. Default: None.
+ kwargs (key word augments): Other augments used in ConvModule.
+ """
+
+ def __init__(self,
+ in_channels,
+ inner_channels,
+ num_layers=1,
+ num_upsample=None,
+ conv_cfg=None,
+ norm_cfg=None,
+ init_cfg=None,
+ **kwargs):
+ super(ConvUpsample, self).__init__(init_cfg)
+ if num_upsample is None:
+ num_upsample = num_layers
+ assert num_upsample <= num_layers, \
+ f'num_upsample({num_upsample})must be no more than ' \
+ f'num_layers({num_layers})'
+ self.num_layers = num_layers
+ self.num_upsample = num_upsample
+ self.conv = ModuleList()
+ for i in range(num_layers):
+ self.conv.append(
+ ConvModule(
+ in_channels,
+ inner_channels,
+ 3,
+ padding=1,
+ stride=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ **kwargs))
+ in_channels = inner_channels
+
+ def forward(self, x):
+ num_upsample = self.num_upsample
+ for i in range(self.num_layers):
+ x = self.conv[i](x)
+ if num_upsample > 0:
+ num_upsample -= 1
+ x = F.interpolate(
+ x, scale_factor=2, mode='bilinear', align_corners=False)
+ return x
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/csp_layer.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/csp_layer.py
new file mode 100644
index 0000000000000000000000000000000000000000..c8b547b8994862bfe14739033bb6b254ef886f29
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/csp_layer.py
@@ -0,0 +1,246 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule, DepthwiseSeparableConvModule
+from mmengine.model import BaseModule
+from torch import Tensor
+
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from .se_layer import ChannelAttention
+
+
+class DarknetBottleneck(BaseModule):
+ """The basic bottleneck block used in Darknet.
+
+ Each ResBlock consists of two ConvModules and the input is added to the
+ final output. Each ConvModule is composed of Conv, BN, and LeakyReLU.
+ The first convLayer has filter size of 1x1 and the second one has the
+ filter size of 3x3.
+
+ Args:
+ in_channels (int): The input channels of this Module.
+ out_channels (int): The output channels of this Module.
+ expansion (float): The kernel size of the convolution.
+ Defaults to 0.5.
+ add_identity (bool): Whether to add identity to the out.
+ Defaults to True.
+ use_depthwise (bool): Whether to use depthwise separable convolution.
+ Defaults to False.
+ conv_cfg (dict): Config dict for convolution layer. Defaults to None,
+ which means using conv2d.
+ norm_cfg (dict): Config dict for normalization layer.
+ Defaults to dict(type='BN').
+ act_cfg (dict): Config dict for activation layer.
+ Defaults to dict(type='Swish').
+ """
+
+ def __init__(self,
+ in_channels: int,
+ out_channels: int,
+ expansion: float = 0.5,
+ add_identity: bool = True,
+ use_depthwise: bool = False,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: ConfigType = dict(
+ type='BN', momentum=0.03, eps=0.001),
+ act_cfg: ConfigType = dict(type='Swish'),
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ hidden_channels = int(out_channels * expansion)
+ conv = DepthwiseSeparableConvModule if use_depthwise else ConvModule
+ self.conv1 = ConvModule(
+ in_channels,
+ hidden_channels,
+ 1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+ self.conv2 = conv(
+ hidden_channels,
+ out_channels,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+ self.add_identity = \
+ add_identity and in_channels == out_channels
+
+ def forward(self, x: Tensor) -> Tensor:
+ """Forward function."""
+ identity = x
+ out = self.conv1(x)
+ out = self.conv2(out)
+
+ if self.add_identity:
+ return out + identity
+ else:
+ return out
+
+
+class CSPNeXtBlock(BaseModule):
+ """The basic bottleneck block used in CSPNeXt.
+
+ Args:
+ in_channels (int): The input channels of this Module.
+ out_channels (int): The output channels of this Module.
+ expansion (float): Expand ratio of the hidden channel. Defaults to 0.5.
+ add_identity (bool): Whether to add identity to the out. Only works
+ when in_channels == out_channels. Defaults to True.
+ use_depthwise (bool): Whether to use depthwise separable convolution.
+ Defaults to False.
+ kernel_size (int): The kernel size of the second convolution layer.
+ Defaults to 5.
+ conv_cfg (dict): Config dict for convolution layer. Defaults to None,
+ which means using conv2d.
+ norm_cfg (dict): Config dict for normalization layer.
+ Defaults to dict(type='BN', momentum=0.03, eps=0.001).
+ act_cfg (dict): Config dict for activation layer.
+ Defaults to dict(type='SiLU').
+ init_cfg (:obj:`ConfigDict` or dict or list[dict] or
+ list[:obj:`ConfigDict`], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ in_channels: int,
+ out_channels: int,
+ expansion: float = 0.5,
+ add_identity: bool = True,
+ use_depthwise: bool = False,
+ kernel_size: int = 5,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: ConfigType = dict(
+ type='BN', momentum=0.03, eps=0.001),
+ act_cfg: ConfigType = dict(type='SiLU'),
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ hidden_channels = int(out_channels * expansion)
+ conv = DepthwiseSeparableConvModule if use_depthwise else ConvModule
+ self.conv1 = conv(
+ in_channels,
+ hidden_channels,
+ 3,
+ stride=1,
+ padding=1,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+ self.conv2 = DepthwiseSeparableConvModule(
+ hidden_channels,
+ out_channels,
+ kernel_size,
+ stride=1,
+ padding=kernel_size // 2,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+ self.add_identity = \
+ add_identity and in_channels == out_channels
+
+ def forward(self, x: Tensor) -> Tensor:
+ """Forward function."""
+ identity = x
+ out = self.conv1(x)
+ out = self.conv2(out)
+
+ if self.add_identity:
+ return out + identity
+ else:
+ return out
+
+
+class CSPLayer(BaseModule):
+ """Cross Stage Partial Layer.
+
+ Args:
+ in_channels (int): The input channels of the CSP layer.
+ out_channels (int): The output channels of the CSP layer.
+ expand_ratio (float): Ratio to adjust the number of channels of the
+ hidden layer. Defaults to 0.5.
+ num_blocks (int): Number of blocks. Defaults to 1.
+ add_identity (bool): Whether to add identity in blocks.
+ Defaults to True.
+ use_cspnext_block (bool): Whether to use CSPNeXt block.
+ Defaults to False.
+ use_depthwise (bool): Whether to use depthwise separable convolution in
+ blocks. Defaults to False.
+ channel_attention (bool): Whether to add channel attention in each
+ stage. Defaults to True.
+ conv_cfg (dict, optional): Config dict for convolution layer.
+ Defaults to None, which means using conv2d.
+ norm_cfg (dict): Config dict for normalization layer.
+ Defaults to dict(type='BN')
+ act_cfg (dict): Config dict for activation layer.
+ Defaults to dict(type='Swish')
+ init_cfg (:obj:`ConfigDict` or dict or list[dict] or
+ list[:obj:`ConfigDict`], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ in_channels: int,
+ out_channels: int,
+ expand_ratio: float = 0.5,
+ num_blocks: int = 1,
+ add_identity: bool = True,
+ use_depthwise: bool = False,
+ use_cspnext_block: bool = False,
+ channel_attention: bool = False,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: ConfigType = dict(
+ type='BN', momentum=0.03, eps=0.001),
+ act_cfg: ConfigType = dict(type='Swish'),
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ block = CSPNeXtBlock if use_cspnext_block else DarknetBottleneck
+ mid_channels = int(out_channels * expand_ratio)
+ self.channel_attention = channel_attention
+ self.main_conv = ConvModule(
+ in_channels,
+ mid_channels,
+ 1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+ self.short_conv = ConvModule(
+ in_channels,
+ mid_channels,
+ 1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+ self.final_conv = ConvModule(
+ 2 * mid_channels,
+ out_channels,
+ 1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+
+ self.blocks = nn.Sequential(*[
+ block(
+ mid_channels,
+ mid_channels,
+ 1.0,
+ add_identity,
+ use_depthwise,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg) for _ in range(num_blocks)
+ ])
+ if channel_attention:
+ self.attention = ChannelAttention(2 * mid_channels)
+
+ def forward(self, x: Tensor) -> Tensor:
+ """Forward function."""
+ x_short = self.short_conv(x)
+
+ x_main = self.main_conv(x)
+ x_main = self.blocks(x_main)
+
+ x_final = torch.cat((x_main, x_short), dim=1)
+
+ if self.channel_attention:
+ x_final = self.attention(x_final)
+ return self.final_conv(x_final)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/dropblock.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/dropblock.py
new file mode 100644
index 0000000000000000000000000000000000000000..7938199b761d637afdb1b2c62dbca01d1bf629eb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/dropblock.py
@@ -0,0 +1,86 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+
+from mmdet.registry import MODELS
+
+eps = 1e-6
+
+
+@MODELS.register_module()
+class DropBlock(nn.Module):
+ """Randomly drop some regions of feature maps.
+
+ Please refer to the method proposed in `DropBlock
+ `_ for details.
+
+ Args:
+ drop_prob (float): The probability of dropping each block.
+ block_size (int): The size of dropped blocks.
+ warmup_iters (int): The drop probability will linearly increase
+ from `0` to `drop_prob` during the first `warmup_iters` iterations.
+ Default: 2000.
+ """
+
+ def __init__(self, drop_prob, block_size, warmup_iters=2000, **kwargs):
+ super(DropBlock, self).__init__()
+ assert block_size % 2 == 1
+ assert 0 < drop_prob <= 1
+ assert warmup_iters >= 0
+ self.drop_prob = drop_prob
+ self.block_size = block_size
+ self.warmup_iters = warmup_iters
+ self.iter_cnt = 0
+
+ def forward(self, x):
+ """
+ Args:
+ x (Tensor): Input feature map on which some areas will be randomly
+ dropped.
+
+ Returns:
+ Tensor: The tensor after DropBlock layer.
+ """
+ if not self.training:
+ return x
+ self.iter_cnt += 1
+ N, C, H, W = list(x.shape)
+ gamma = self._compute_gamma((H, W))
+ mask_shape = (N, C, H - self.block_size + 1, W - self.block_size + 1)
+ mask = torch.bernoulli(torch.full(mask_shape, gamma, device=x.device))
+
+ mask = F.pad(mask, [self.block_size // 2] * 4, value=0)
+ mask = F.max_pool2d(
+ input=mask,
+ stride=(1, 1),
+ kernel_size=(self.block_size, self.block_size),
+ padding=self.block_size // 2)
+ mask = 1 - mask
+ x = x * mask * mask.numel() / (eps + mask.sum())
+ return x
+
+ def _compute_gamma(self, feat_size):
+ """Compute the value of gamma according to paper. gamma is the
+ parameter of bernoulli distribution, which controls the number of
+ features to drop.
+
+ gamma = (drop_prob * fm_area) / (drop_area * keep_area)
+
+ Args:
+ feat_size (tuple[int, int]): The height and width of feature map.
+
+ Returns:
+ float: The value of gamma.
+ """
+ gamma = (self.drop_prob * feat_size[0] * feat_size[1])
+ gamma /= ((feat_size[0] - self.block_size + 1) *
+ (feat_size[1] - self.block_size + 1))
+ gamma /= (self.block_size**2)
+ factor = (1.0 if self.iter_cnt > self.warmup_iters else self.iter_cnt /
+ self.warmup_iters)
+ return gamma * factor
+
+ def extra_repr(self):
+ return (f'drop_prob={self.drop_prob}, block_size={self.block_size}, '
+ f'warmup_iters={self.warmup_iters}')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/ema.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/ema.py
new file mode 100644
index 0000000000000000000000000000000000000000..73a0ca67c2888a0b17476e60b60eaf0b7eba4a6a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/ema.py
@@ -0,0 +1,66 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+from typing import Optional
+
+import torch
+import torch.nn as nn
+from mmengine.model import ExponentialMovingAverage
+from torch import Tensor
+
+from mmdet.registry import MODELS
+
+
+@MODELS.register_module()
+class ExpMomentumEMA(ExponentialMovingAverage):
+ """Exponential moving average (EMA) with exponential momentum strategy,
+ which is used in YOLOX.
+
+ Args:
+ model (nn.Module): The model to be averaged.
+ momentum (float): The momentum used for updating ema parameter.
+ Ema's parameter are updated with the formula:
+ `averaged_param = (1-momentum) * averaged_param + momentum *
+ source_param`. Defaults to 0.0002.
+ gamma (int): Use a larger momentum early in training and gradually
+ annealing to a smaller value to update the ema model smoothly. The
+ momentum is calculated as
+ `(1 - momentum) * exp(-(1 + steps) / gamma) + momentum`.
+ Defaults to 2000.
+ interval (int): Interval between two updates. Defaults to 1.
+ device (torch.device, optional): If provided, the averaged model will
+ be stored on the :attr:`device`. Defaults to None.
+ update_buffers (bool): if True, it will compute running averages for
+ both the parameters and the buffers of the model. Defaults to
+ False.
+ """
+
+ def __init__(self,
+ model: nn.Module,
+ momentum: float = 0.0002,
+ gamma: int = 2000,
+ interval=1,
+ device: Optional[torch.device] = None,
+ update_buffers: bool = False) -> None:
+ super().__init__(
+ model=model,
+ momentum=momentum,
+ interval=interval,
+ device=device,
+ update_buffers=update_buffers)
+ assert gamma > 0, f'gamma must be greater than 0, but got {gamma}'
+ self.gamma = gamma
+
+ def avg_func(self, averaged_param: Tensor, source_param: Tensor,
+ steps: int) -> None:
+ """Compute the moving average of the parameters using the exponential
+ momentum strategy.
+
+ Args:
+ averaged_param (Tensor): The averaged parameters.
+ source_param (Tensor): The source parameters.
+ steps (int): The number of times the parameters have been
+ updated.
+ """
+ momentum = (1 - self.momentum) * math.exp(
+ -float(1 + steps) / self.gamma) + self.momentum
+ averaged_param.lerp_(source_param, momentum)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/inverted_residual.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/inverted_residual.py
new file mode 100644
index 0000000000000000000000000000000000000000..a174ccc8835a1ee720f9cdaa7c5be210f5be8113
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/inverted_residual.py
@@ -0,0 +1,130 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch.nn as nn
+import torch.utils.checkpoint as cp
+from mmcv.cnn import ConvModule
+from mmcv.cnn.bricks import DropPath
+from mmengine.model import BaseModule
+
+from .se_layer import SELayer
+
+
+class InvertedResidual(BaseModule):
+ """Inverted Residual Block.
+
+ Args:
+ in_channels (int): The input channels of this Module.
+ out_channels (int): The output channels of this Module.
+ mid_channels (int): The input channels of the depthwise convolution.
+ kernel_size (int): The kernel size of the depthwise convolution.
+ Default: 3.
+ stride (int): The stride of the depthwise convolution. Default: 1.
+ se_cfg (dict): Config dict for se layer. Default: None, which means no
+ se layer.
+ with_expand_conv (bool): Use expand conv or not. If set False,
+ mid_channels must be the same with in_channels.
+ Default: True.
+ conv_cfg (dict): Config dict for convolution layer. Default: None,
+ which means using conv2d.
+ norm_cfg (dict): Config dict for normalization layer.
+ Default: dict(type='BN').
+ act_cfg (dict): Config dict for activation layer.
+ Default: dict(type='ReLU').
+ drop_path_rate (float): stochastic depth rate. Defaults to 0.
+ with_cp (bool): Use checkpoint or not. Using checkpoint will save some
+ memory while slowing down the training speed. Default: False.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None
+
+ Returns:
+ Tensor: The output tensor.
+ """
+
+ def __init__(self,
+ in_channels,
+ out_channels,
+ mid_channels,
+ kernel_size=3,
+ stride=1,
+ se_cfg=None,
+ with_expand_conv=True,
+ conv_cfg=None,
+ norm_cfg=dict(type='BN'),
+ act_cfg=dict(type='ReLU'),
+ drop_path_rate=0.,
+ with_cp=False,
+ init_cfg=None):
+ super(InvertedResidual, self).__init__(init_cfg)
+ self.with_res_shortcut = (stride == 1 and in_channels == out_channels)
+ assert stride in [1, 2], f'stride must in [1, 2]. ' \
+ f'But received {stride}.'
+ self.with_cp = with_cp
+ self.drop_path = DropPath(
+ drop_path_rate) if drop_path_rate > 0 else nn.Identity()
+ self.with_se = se_cfg is not None
+ self.with_expand_conv = with_expand_conv
+
+ if self.with_se:
+ assert isinstance(se_cfg, dict)
+ if not self.with_expand_conv:
+ assert mid_channels == in_channels
+
+ if self.with_expand_conv:
+ self.expand_conv = ConvModule(
+ in_channels=in_channels,
+ out_channels=mid_channels,
+ kernel_size=1,
+ stride=1,
+ padding=0,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+ self.depthwise_conv = ConvModule(
+ in_channels=mid_channels,
+ out_channels=mid_channels,
+ kernel_size=kernel_size,
+ stride=stride,
+ padding=kernel_size // 2,
+ groups=mid_channels,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+
+ if self.with_se:
+ self.se = SELayer(**se_cfg)
+
+ self.linear_conv = ConvModule(
+ in_channels=mid_channels,
+ out_channels=out_channels,
+ kernel_size=1,
+ stride=1,
+ padding=0,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=None)
+
+ def forward(self, x):
+
+ def _inner_forward(x):
+ out = x
+
+ if self.with_expand_conv:
+ out = self.expand_conv(out)
+
+ out = self.depthwise_conv(out)
+
+ if self.with_se:
+ out = self.se(out)
+
+ out = self.linear_conv(out)
+
+ if self.with_res_shortcut:
+ return x + self.drop_path(out)
+ else:
+ return out
+
+ if self.with_cp and x.requires_grad:
+ out = cp.checkpoint(_inner_forward, x)
+ else:
+ out = _inner_forward(x)
+
+ return out
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/matrix_nms.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/matrix_nms.py
new file mode 100644
index 0000000000000000000000000000000000000000..9dc8c4f74e28127fb69ccc684f0bdb2bd3943b20
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/matrix_nms.py
@@ -0,0 +1,121 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+
+
+def mask_matrix_nms(masks,
+ labels,
+ scores,
+ filter_thr=-1,
+ nms_pre=-1,
+ max_num=-1,
+ kernel='gaussian',
+ sigma=2.0,
+ mask_area=None):
+ """Matrix NMS for multi-class masks.
+
+ Args:
+ masks (Tensor): Has shape (num_instances, h, w)
+ labels (Tensor): Labels of corresponding masks,
+ has shape (num_instances,).
+ scores (Tensor): Mask scores of corresponding masks,
+ has shape (num_instances).
+ filter_thr (float): Score threshold to filter the masks
+ after matrix nms. Default: -1, which means do not
+ use filter_thr.
+ nms_pre (int): The max number of instances to do the matrix nms.
+ Default: -1, which means do not use nms_pre.
+ max_num (int, optional): If there are more than max_num masks after
+ matrix, only top max_num will be kept. Default: -1, which means
+ do not use max_num.
+ kernel (str): 'linear' or 'gaussian'.
+ sigma (float): std in gaussian method.
+ mask_area (Tensor): The sum of seg_masks.
+
+ Returns:
+ tuple(Tensor): Processed mask results.
+
+ - scores (Tensor): Updated scores, has shape (n,).
+ - labels (Tensor): Remained labels, has shape (n,).
+ - masks (Tensor): Remained masks, has shape (n, w, h).
+ - keep_inds (Tensor): The indices number of
+ the remaining mask in the input mask, has shape (n,).
+ """
+ assert len(labels) == len(masks) == len(scores)
+ if len(labels) == 0:
+ return scores.new_zeros(0), labels.new_zeros(0), masks.new_zeros(
+ 0, *masks.shape[-2:]), labels.new_zeros(0)
+ if mask_area is None:
+ mask_area = masks.sum((1, 2)).float()
+ else:
+ assert len(masks) == len(mask_area)
+
+ # sort and keep top nms_pre
+ scores, sort_inds = torch.sort(scores, descending=True)
+
+ keep_inds = sort_inds
+ if nms_pre > 0 and len(sort_inds) > nms_pre:
+ sort_inds = sort_inds[:nms_pre]
+ keep_inds = keep_inds[:nms_pre]
+ scores = scores[:nms_pre]
+ masks = masks[sort_inds]
+ mask_area = mask_area[sort_inds]
+ labels = labels[sort_inds]
+
+ num_masks = len(labels)
+ flatten_masks = masks.reshape(num_masks, -1).float()
+ # inter.
+ inter_matrix = torch.mm(flatten_masks, flatten_masks.transpose(1, 0))
+ expanded_mask_area = mask_area.expand(num_masks, num_masks)
+ # Upper triangle iou matrix.
+ iou_matrix = (inter_matrix /
+ (expanded_mask_area + expanded_mask_area.transpose(1, 0) -
+ inter_matrix)).triu(diagonal=1)
+ # label_specific matrix.
+ expanded_labels = labels.expand(num_masks, num_masks)
+ # Upper triangle label matrix.
+ label_matrix = (expanded_labels == expanded_labels.transpose(
+ 1, 0)).triu(diagonal=1)
+
+ # IoU compensation
+ compensate_iou, _ = (iou_matrix * label_matrix).max(0)
+ compensate_iou = compensate_iou.expand(num_masks,
+ num_masks).transpose(1, 0)
+
+ # IoU decay
+ decay_iou = iou_matrix * label_matrix
+
+ # Calculate the decay_coefficient
+ if kernel == 'gaussian':
+ decay_matrix = torch.exp(-1 * sigma * (decay_iou**2))
+ compensate_matrix = torch.exp(-1 * sigma * (compensate_iou**2))
+ decay_coefficient, _ = (decay_matrix / compensate_matrix).min(0)
+ elif kernel == 'linear':
+ decay_matrix = (1 - decay_iou) / (1 - compensate_iou)
+ decay_coefficient, _ = decay_matrix.min(0)
+ else:
+ raise NotImplementedError(
+ f'{kernel} kernel is not supported in matrix nms!')
+ # update the score.
+ scores = scores * decay_coefficient
+
+ if filter_thr > 0:
+ keep = scores >= filter_thr
+ keep_inds = keep_inds[keep]
+ if not keep.any():
+ return scores.new_zeros(0), labels.new_zeros(0), masks.new_zeros(
+ 0, *masks.shape[-2:]), labels.new_zeros(0)
+ masks = masks[keep]
+ scores = scores[keep]
+ labels = labels[keep]
+
+ # sort and keep top max_num
+ scores, sort_inds = torch.sort(scores, descending=True)
+ keep_inds = keep_inds[sort_inds]
+ if max_num > 0 and len(sort_inds) > max_num:
+ sort_inds = sort_inds[:max_num]
+ keep_inds = keep_inds[:max_num]
+ scores = scores[:max_num]
+ masks = masks[sort_inds]
+ labels = labels[sort_inds]
+
+ return scores, labels, masks, keep_inds
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/msdeformattn_pixel_decoder.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/msdeformattn_pixel_decoder.py
new file mode 100644
index 0000000000000000000000000000000000000000..a67dc3c4437f83ebe1c82d12b3ed91f429030ce7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/msdeformattn_pixel_decoder.py
@@ -0,0 +1,246 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple, Union
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import Conv2d, ConvModule
+from mmcv.cnn.bricks.transformer import MultiScaleDeformableAttention
+from mmengine.model import (BaseModule, ModuleList, caffe2_xavier_init,
+ normal_init, xavier_init)
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptMultiConfig
+from ..task_modules.prior_generators import MlvlPointGenerator
+from .positional_encoding import SinePositionalEncoding
+from .transformer import Mask2FormerTransformerEncoder
+
+
+@MODELS.register_module()
+class MSDeformAttnPixelDecoder(BaseModule):
+ """Pixel decoder with multi-scale deformable attention.
+
+ Args:
+ in_channels (list[int] | tuple[int]): Number of channels in the
+ input feature maps.
+ strides (list[int] | tuple[int]): Output strides of feature from
+ backbone.
+ feat_channels (int): Number of channels for feature.
+ out_channels (int): Number of channels for output.
+ num_outs (int): Number of output scales.
+ norm_cfg (:obj:`ConfigDict` or dict): Config for normalization.
+ Defaults to dict(type='GN', num_groups=32).
+ act_cfg (:obj:`ConfigDict` or dict): Config for activation.
+ Defaults to dict(type='ReLU').
+ encoder (:obj:`ConfigDict` or dict): Config for transformer
+ encoder. Defaults to None.
+ positional_encoding (:obj:`ConfigDict` or dict): Config for
+ transformer encoder position encoding. Defaults to
+ dict(num_feats=128, normalize=True).
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict], optional): Initialization config dict. Defaults to None.
+ """
+
+ def __init__(self,
+ in_channels: Union[List[int],
+ Tuple[int]] = [256, 512, 1024, 2048],
+ strides: Union[List[int], Tuple[int]] = [4, 8, 16, 32],
+ feat_channels: int = 256,
+ out_channels: int = 256,
+ num_outs: int = 3,
+ norm_cfg: ConfigType = dict(type='GN', num_groups=32),
+ act_cfg: ConfigType = dict(type='ReLU'),
+ encoder: ConfigType = None,
+ positional_encoding: ConfigType = dict(
+ num_feats=128, normalize=True),
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.strides = strides
+ self.num_input_levels = len(in_channels)
+ self.num_encoder_levels = \
+ encoder.layer_cfg.self_attn_cfg.num_levels
+ assert self.num_encoder_levels >= 1, \
+ 'num_levels in attn_cfgs must be at least one'
+ input_conv_list = []
+ # from top to down (low to high resolution)
+ for i in range(self.num_input_levels - 1,
+ self.num_input_levels - self.num_encoder_levels - 1,
+ -1):
+ input_conv = ConvModule(
+ in_channels[i],
+ feat_channels,
+ kernel_size=1,
+ norm_cfg=norm_cfg,
+ act_cfg=None,
+ bias=True)
+ input_conv_list.append(input_conv)
+ self.input_convs = ModuleList(input_conv_list)
+
+ self.encoder = Mask2FormerTransformerEncoder(**encoder)
+ self.postional_encoding = SinePositionalEncoding(**positional_encoding)
+ # high resolution to low resolution
+ self.level_encoding = nn.Embedding(self.num_encoder_levels,
+ feat_channels)
+
+ # fpn-like structure
+ self.lateral_convs = ModuleList()
+ self.output_convs = ModuleList()
+ self.use_bias = norm_cfg is None
+ # from top to down (low to high resolution)
+ # fpn for the rest features that didn't pass in encoder
+ for i in range(self.num_input_levels - self.num_encoder_levels - 1, -1,
+ -1):
+ lateral_conv = ConvModule(
+ in_channels[i],
+ feat_channels,
+ kernel_size=1,
+ bias=self.use_bias,
+ norm_cfg=norm_cfg,
+ act_cfg=None)
+ output_conv = ConvModule(
+ feat_channels,
+ feat_channels,
+ kernel_size=3,
+ stride=1,
+ padding=1,
+ bias=self.use_bias,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+ self.lateral_convs.append(lateral_conv)
+ self.output_convs.append(output_conv)
+
+ self.mask_feature = Conv2d(
+ feat_channels, out_channels, kernel_size=1, stride=1, padding=0)
+
+ self.num_outs = num_outs
+ self.point_generator = MlvlPointGenerator(strides)
+
+ def init_weights(self) -> None:
+ """Initialize weights."""
+ for i in range(0, self.num_encoder_levels):
+ xavier_init(
+ self.input_convs[i].conv,
+ gain=1,
+ bias=0,
+ distribution='uniform')
+
+ for i in range(0, self.num_input_levels - self.num_encoder_levels):
+ caffe2_xavier_init(self.lateral_convs[i].conv, bias=0)
+ caffe2_xavier_init(self.output_convs[i].conv, bias=0)
+
+ caffe2_xavier_init(self.mask_feature, bias=0)
+
+ normal_init(self.level_encoding, mean=0, std=1)
+ for p in self.encoder.parameters():
+ if p.dim() > 1:
+ nn.init.xavier_normal_(p)
+
+ # init_weights defined in MultiScaleDeformableAttention
+ for m in self.encoder.layers.modules():
+ if isinstance(m, MultiScaleDeformableAttention):
+ m.init_weights()
+
+ def forward(self, feats: List[Tensor]) -> Tuple[Tensor, Tensor]:
+ """
+ Args:
+ feats (list[Tensor]): Feature maps of each level. Each has
+ shape of (batch_size, c, h, w).
+
+ Returns:
+ tuple: A tuple containing the following:
+
+ - mask_feature (Tensor): shape (batch_size, c, h, w).
+ - multi_scale_features (list[Tensor]): Multi scale \
+ features, each in shape (batch_size, c, h, w).
+ """
+ # generate padding mask for each level, for each image
+ batch_size = feats[0].shape[0]
+ encoder_input_list = []
+ padding_mask_list = []
+ level_positional_encoding_list = []
+ spatial_shapes = []
+ reference_points_list = []
+ for i in range(self.num_encoder_levels):
+ level_idx = self.num_input_levels - i - 1
+ feat = feats[level_idx]
+ feat_projected = self.input_convs[i](feat)
+ feat_hw = torch._shape_as_tensor(feat)[2:].to(feat.device)
+
+ # no padding
+ padding_mask_resized = feat.new_zeros(
+ (batch_size, ) + feat.shape[-2:], dtype=torch.bool)
+ pos_embed = self.postional_encoding(padding_mask_resized)
+ level_embed = self.level_encoding.weight[i]
+ level_pos_embed = level_embed.view(1, -1, 1, 1) + pos_embed
+ # (h_i * w_i, 2)
+ reference_points = self.point_generator.single_level_grid_priors(
+ feat.shape[-2:], level_idx, device=feat.device)
+ # normalize
+ feat_wh = feat_hw.unsqueeze(0).flip(dims=[0, 1])
+ factor = feat_wh * self.strides[level_idx]
+ reference_points = reference_points / factor
+
+ # shape (batch_size, c, h_i, w_i) -> (h_i * w_i, batch_size, c)
+ feat_projected = feat_projected.flatten(2).permute(0, 2, 1)
+ level_pos_embed = level_pos_embed.flatten(2).permute(0, 2, 1)
+ padding_mask_resized = padding_mask_resized.flatten(1)
+
+ encoder_input_list.append(feat_projected)
+ padding_mask_list.append(padding_mask_resized)
+ level_positional_encoding_list.append(level_pos_embed)
+ spatial_shapes.append(feat_hw)
+ reference_points_list.append(reference_points)
+ # shape (batch_size, total_num_queries),
+ # total_num_queries=sum([., h_i * w_i,.])
+ padding_masks = torch.cat(padding_mask_list, dim=1)
+ # shape (total_num_queries, batch_size, c)
+ encoder_inputs = torch.cat(encoder_input_list, dim=1)
+ level_positional_encodings = torch.cat(
+ level_positional_encoding_list, dim=1)
+ # shape (num_encoder_levels, 2), from low
+ # resolution to high resolution
+ num_queries_per_level = [e[0] * e[1] for e in spatial_shapes]
+ spatial_shapes = torch.cat(spatial_shapes).view(-1, 2)
+ # shape (0, h_0*w_0, h_0*w_0+h_1*w_1, ...)
+ level_start_index = torch.cat((spatial_shapes.new_zeros(
+ (1, )), spatial_shapes.prod(1).cumsum(0)[:-1]))
+ reference_points = torch.cat(reference_points_list, dim=0)
+ reference_points = reference_points[None, :, None].repeat(
+ batch_size, 1, self.num_encoder_levels, 1)
+ valid_radios = reference_points.new_ones(
+ (batch_size, self.num_encoder_levels, 2))
+ # shape (num_total_queries, batch_size, c)
+ memory = self.encoder(
+ query=encoder_inputs,
+ query_pos=level_positional_encodings,
+ key_padding_mask=padding_masks,
+ spatial_shapes=spatial_shapes,
+ reference_points=reference_points,
+ level_start_index=level_start_index,
+ valid_ratios=valid_radios)
+ # (batch_size, c, num_total_queries)
+ memory = memory.permute(0, 2, 1)
+
+ # from low resolution to high resolution
+ outs = torch.split(memory, num_queries_per_level, dim=-1)
+ outs = [
+ x.reshape(batch_size, -1, spatial_shapes[i][0],
+ spatial_shapes[i][1]) for i, x in enumerate(outs)
+ ]
+
+ for i in range(self.num_input_levels - self.num_encoder_levels - 1, -1,
+ -1):
+ x = feats[i]
+ cur_feat = self.lateral_convs[i](x)
+ y = cur_feat + F.interpolate(
+ outs[-1],
+ size=cur_feat.shape[-2:],
+ mode='bilinear',
+ align_corners=False)
+ y = self.output_convs[i](y)
+ outs.append(y)
+ multi_scale_features = outs[:self.num_outs]
+
+ mask_feature = self.mask_feature(outs[-1])
+ return mask_feature, multi_scale_features
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/normed_predictor.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/normed_predictor.py
new file mode 100644
index 0000000000000000000000000000000000000000..592194b1dbbb8582f4c642bf29135573e1f8c3c8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/normed_predictor.py
@@ -0,0 +1,99 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmengine.utils import digit_version
+from torch import Tensor
+
+from mmdet.registry import MODELS
+
+MODELS.register_module('Linear', module=nn.Linear)
+
+
+@MODELS.register_module(name='NormedLinear')
+class NormedLinear(nn.Linear):
+ """Normalized Linear Layer.
+
+ Args:
+ tempeature (float, optional): Tempeature term. Defaults to 20.
+ power (int, optional): Power term. Defaults to 1.0.
+ eps (float, optional): The minimal value of divisor to
+ keep numerical stability. Defaults to 1e-6.
+ """
+
+ def __init__(self,
+ *args,
+ tempearture: float = 20,
+ power: int = 1.0,
+ eps: float = 1e-6,
+ **kwargs) -> None:
+ super().__init__(*args, **kwargs)
+ self.tempearture = tempearture
+ self.power = power
+ self.eps = eps
+ self.init_weights()
+
+ def init_weights(self) -> None:
+ """Initialize the weights."""
+ nn.init.normal_(self.weight, mean=0, std=0.01)
+ if self.bias is not None:
+ nn.init.constant_(self.bias, 0)
+
+ def forward(self, x: Tensor) -> Tensor:
+ """Forward function for `NormedLinear`."""
+ weight_ = self.weight / (
+ self.weight.norm(dim=1, keepdim=True).pow(self.power) + self.eps)
+ x_ = x / (x.norm(dim=1, keepdim=True).pow(self.power) + self.eps)
+ x_ = x_ * self.tempearture
+
+ return F.linear(x_, weight_, self.bias)
+
+
+@MODELS.register_module(name='NormedConv2d')
+class NormedConv2d(nn.Conv2d):
+ """Normalized Conv2d Layer.
+
+ Args:
+ tempeature (float, optional): Tempeature term. Defaults to 20.
+ power (int, optional): Power term. Defaults to 1.0.
+ eps (float, optional): The minimal value of divisor to
+ keep numerical stability. Defaults to 1e-6.
+ norm_over_kernel (bool, optional): Normalize over kernel.
+ Defaults to False.
+ """
+
+ def __init__(self,
+ *args,
+ tempearture: float = 20,
+ power: int = 1.0,
+ eps: float = 1e-6,
+ norm_over_kernel: bool = False,
+ **kwargs) -> None:
+ super().__init__(*args, **kwargs)
+ self.tempearture = tempearture
+ self.power = power
+ self.norm_over_kernel = norm_over_kernel
+ self.eps = eps
+
+ def forward(self, x: Tensor) -> Tensor:
+ """Forward function for `NormedConv2d`."""
+ if not self.norm_over_kernel:
+ weight_ = self.weight / (
+ self.weight.norm(dim=1, keepdim=True).pow(self.power) +
+ self.eps)
+ else:
+ weight_ = self.weight / (
+ self.weight.view(self.weight.size(0), -1).norm(
+ dim=1, keepdim=True).pow(self.power)[..., None, None] +
+ self.eps)
+ x_ = x / (x.norm(dim=1, keepdim=True).pow(self.power) + self.eps)
+ x_ = x_ * self.tempearture
+
+ if hasattr(self, 'conv2d_forward'):
+ x_ = self.conv2d_forward(x_, weight_)
+ else:
+ if digit_version(torch.__version__) >= digit_version('1.8'):
+ x_ = self._conv_forward(x_, weight_, self.bias)
+ else:
+ x_ = self._conv_forward(x_, weight_)
+ return x_
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/pixel_decoder.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/pixel_decoder.py
new file mode 100644
index 0000000000000000000000000000000000000000..fb61434045eb9996276518577800132e4a25eb3e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/pixel_decoder.py
@@ -0,0 +1,249 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple, Union
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import Conv2d, ConvModule
+from mmengine.model import BaseModule, ModuleList, caffe2_xavier_init
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptMultiConfig
+from .positional_encoding import SinePositionalEncoding
+from .transformer import DetrTransformerEncoder
+
+
+@MODELS.register_module()
+class PixelDecoder(BaseModule):
+ """Pixel decoder with a structure like fpn.
+
+ Args:
+ in_channels (list[int] | tuple[int]): Number of channels in the
+ input feature maps.
+ feat_channels (int): Number channels for feature.
+ out_channels (int): Number channels for output.
+ norm_cfg (:obj:`ConfigDict` or dict): Config for normalization.
+ Defaults to dict(type='GN', num_groups=32).
+ act_cfg (:obj:`ConfigDict` or dict): Config for activation.
+ Defaults to dict(type='ReLU').
+ encoder (:obj:`ConfigDict` or dict): Config for transorformer
+ encoder.Defaults to None.
+ positional_encoding (:obj:`ConfigDict` or dict): Config for
+ transformer encoder position encoding. Defaults to
+ dict(type='SinePositionalEncoding', num_feats=128,
+ normalize=True).
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict], optional): Initialization config dict. Defaults to None.
+ """
+
+ def __init__(self,
+ in_channels: Union[List[int], Tuple[int]],
+ feat_channels: int,
+ out_channels: int,
+ norm_cfg: ConfigType = dict(type='GN', num_groups=32),
+ act_cfg: ConfigType = dict(type='ReLU'),
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.in_channels = in_channels
+ self.num_inputs = len(in_channels)
+ self.lateral_convs = ModuleList()
+ self.output_convs = ModuleList()
+ self.use_bias = norm_cfg is None
+ for i in range(0, self.num_inputs - 1):
+ lateral_conv = ConvModule(
+ in_channels[i],
+ feat_channels,
+ kernel_size=1,
+ bias=self.use_bias,
+ norm_cfg=norm_cfg,
+ act_cfg=None)
+ output_conv = ConvModule(
+ feat_channels,
+ feat_channels,
+ kernel_size=3,
+ stride=1,
+ padding=1,
+ bias=self.use_bias,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+ self.lateral_convs.append(lateral_conv)
+ self.output_convs.append(output_conv)
+
+ self.last_feat_conv = ConvModule(
+ in_channels[-1],
+ feat_channels,
+ kernel_size=3,
+ padding=1,
+ stride=1,
+ bias=self.use_bias,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+ self.mask_feature = Conv2d(
+ feat_channels, out_channels, kernel_size=3, stride=1, padding=1)
+
+ def init_weights(self) -> None:
+ """Initialize weights."""
+ for i in range(0, self.num_inputs - 2):
+ caffe2_xavier_init(self.lateral_convs[i].conv, bias=0)
+ caffe2_xavier_init(self.output_convs[i].conv, bias=0)
+
+ caffe2_xavier_init(self.mask_feature, bias=0)
+ caffe2_xavier_init(self.last_feat_conv, bias=0)
+
+ def forward(self, feats: List[Tensor],
+ batch_img_metas: List[dict]) -> Tuple[Tensor, Tensor]:
+ """
+ Args:
+ feats (list[Tensor]): Feature maps of each level. Each has
+ shape of (batch_size, c, h, w).
+ batch_img_metas (list[dict]): List of image information.
+ Pass in for creating more accurate padding mask. Not
+ used here.
+
+ Returns:
+ tuple[Tensor, Tensor]: a tuple containing the following:
+
+ - mask_feature (Tensor): Shape (batch_size, c, h, w).
+ - memory (Tensor): Output of last stage of backbone.\
+ Shape (batch_size, c, h, w).
+ """
+ y = self.last_feat_conv(feats[-1])
+ for i in range(self.num_inputs - 2, -1, -1):
+ x = feats[i]
+ cur_feat = self.lateral_convs[i](x)
+ y = cur_feat + \
+ F.interpolate(y, size=cur_feat.shape[-2:], mode='nearest')
+ y = self.output_convs[i](y)
+
+ mask_feature = self.mask_feature(y)
+ memory = feats[-1]
+ return mask_feature, memory
+
+
+@MODELS.register_module()
+class TransformerEncoderPixelDecoder(PixelDecoder):
+ """Pixel decoder with transormer encoder inside.
+
+ Args:
+ in_channels (list[int] | tuple[int]): Number of channels in the
+ input feature maps.
+ feat_channels (int): Number channels for feature.
+ out_channels (int): Number channels for output.
+ norm_cfg (:obj:`ConfigDict` or dict): Config for normalization.
+ Defaults to dict(type='GN', num_groups=32).
+ act_cfg (:obj:`ConfigDict` or dict): Config for activation.
+ Defaults to dict(type='ReLU').
+ encoder (:obj:`ConfigDict` or dict): Config for transformer encoder.
+ Defaults to None.
+ positional_encoding (:obj:`ConfigDict` or dict): Config for
+ transformer encoder position encoding. Defaults to
+ dict(num_feats=128, normalize=True).
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict], optional): Initialization config dict. Defaults to None.
+ """
+
+ def __init__(self,
+ in_channels: Union[List[int], Tuple[int]],
+ feat_channels: int,
+ out_channels: int,
+ norm_cfg: ConfigType = dict(type='GN', num_groups=32),
+ act_cfg: ConfigType = dict(type='ReLU'),
+ encoder: ConfigType = None,
+ positional_encoding: ConfigType = dict(
+ num_feats=128, normalize=True),
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ in_channels=in_channels,
+ feat_channels=feat_channels,
+ out_channels=out_channels,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg,
+ init_cfg=init_cfg)
+ self.last_feat_conv = None
+
+ self.encoder = DetrTransformerEncoder(**encoder)
+ self.encoder_embed_dims = self.encoder.embed_dims
+ assert self.encoder_embed_dims == feat_channels, 'embed_dims({}) of ' \
+ 'tranformer encoder must equal to feat_channels({})'.format(
+ feat_channels, self.encoder_embed_dims)
+ self.positional_encoding = SinePositionalEncoding(
+ **positional_encoding)
+ self.encoder_in_proj = Conv2d(
+ in_channels[-1], feat_channels, kernel_size=1)
+ self.encoder_out_proj = ConvModule(
+ feat_channels,
+ feat_channels,
+ kernel_size=3,
+ stride=1,
+ padding=1,
+ bias=self.use_bias,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+
+ def init_weights(self) -> None:
+ """Initialize weights."""
+ for i in range(0, self.num_inputs - 2):
+ caffe2_xavier_init(self.lateral_convs[i].conv, bias=0)
+ caffe2_xavier_init(self.output_convs[i].conv, bias=0)
+
+ caffe2_xavier_init(self.mask_feature, bias=0)
+ caffe2_xavier_init(self.encoder_in_proj, bias=0)
+ caffe2_xavier_init(self.encoder_out_proj.conv, bias=0)
+
+ for p in self.encoder.parameters():
+ if p.dim() > 1:
+ nn.init.xavier_uniform_(p)
+
+ def forward(self, feats: List[Tensor],
+ batch_img_metas: List[dict]) -> Tuple[Tensor, Tensor]:
+ """
+ Args:
+ feats (list[Tensor]): Feature maps of each level. Each has
+ shape of (batch_size, c, h, w).
+ batch_img_metas (list[dict]): List of image information. Pass in
+ for creating more accurate padding mask.
+
+ Returns:
+ tuple: a tuple containing the following:
+
+ - mask_feature (Tensor): shape (batch_size, c, h, w).
+ - memory (Tensor): shape (batch_size, c, h, w).
+ """
+ feat_last = feats[-1]
+ bs, c, h, w = feat_last.shape
+ input_img_h, input_img_w = batch_img_metas[0]['batch_input_shape']
+ padding_mask = feat_last.new_ones((bs, input_img_h, input_img_w),
+ dtype=torch.float32)
+ for i in range(bs):
+ img_h, img_w = batch_img_metas[i]['img_shape']
+ padding_mask[i, :img_h, :img_w] = 0
+ padding_mask = F.interpolate(
+ padding_mask.unsqueeze(1),
+ size=feat_last.shape[-2:],
+ mode='nearest').to(torch.bool).squeeze(1)
+
+ pos_embed = self.positional_encoding(padding_mask)
+ feat_last = self.encoder_in_proj(feat_last)
+ # (batch_size, c, h, w) -> (batch_size, num_queries, c)
+ feat_last = feat_last.flatten(2).permute(0, 2, 1)
+ pos_embed = pos_embed.flatten(2).permute(0, 2, 1)
+ # (batch_size, h, w) -> (batch_size, h*w)
+ padding_mask = padding_mask.flatten(1)
+ memory = self.encoder(
+ query=feat_last,
+ query_pos=pos_embed,
+ key_padding_mask=padding_mask)
+ # (batch_size, num_queries, c) -> (batch_size, c, h, w)
+ memory = memory.permute(0, 2, 1).view(bs, self.encoder_embed_dims, h,
+ w)
+ y = self.encoder_out_proj(memory)
+ for i in range(self.num_inputs - 2, -1, -1):
+ x = feats[i]
+ cur_feat = self.lateral_convs[i](x)
+ y = cur_feat + \
+ F.interpolate(y, size=cur_feat.shape[-2:], mode='nearest')
+ y = self.output_convs[i](y)
+
+ mask_feature = self.mask_feature(y)
+ return mask_feature, memory
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/positional_encoding.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/positional_encoding.py
new file mode 100644
index 0000000000000000000000000000000000000000..87080d81a9f155839d453b8671103e5d51fbf88a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/positional_encoding.py
@@ -0,0 +1,269 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+from typing import Optional
+
+import torch
+import torch.nn as nn
+from mmengine.model import BaseModule
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import MultiConfig, OptMultiConfig
+
+
+@MODELS.register_module()
+class SinePositionalEncoding(BaseModule):
+ """Position encoding with sine and cosine functions.
+
+ See `End-to-End Object Detection with Transformers
+ `_ for details.
+
+ Args:
+ num_feats (int): The feature dimension for each position
+ along x-axis or y-axis. Note the final returned dimension
+ for each position is 2 times of this value.
+ temperature (int, optional): The temperature used for scaling
+ the position embedding. Defaults to 10000.
+ normalize (bool, optional): Whether to normalize the position
+ embedding. Defaults to False.
+ scale (float, optional): A scale factor that scales the position
+ embedding. The scale will be used only when `normalize` is True.
+ Defaults to 2*pi.
+ eps (float, optional): A value added to the denominator for
+ numerical stability. Defaults to 1e-6.
+ offset (float): offset add to embed when do the normalization.
+ Defaults to 0.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Defaults to None
+ """
+
+ def __init__(self,
+ num_feats: int,
+ temperature: int = 10000,
+ normalize: bool = False,
+ scale: float = 2 * math.pi,
+ eps: float = 1e-6,
+ offset: float = 0.,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ if normalize:
+ assert isinstance(scale, (float, int)), 'when normalize is set,' \
+ 'scale should be provided and in float or int type, ' \
+ f'found {type(scale)}'
+ self.num_feats = num_feats
+ self.temperature = temperature
+ self.normalize = normalize
+ self.scale = scale
+ self.eps = eps
+ self.offset = offset
+
+ def forward(self, mask: Tensor, input: Optional[Tensor] = None) -> Tensor:
+ """Forward function for `SinePositionalEncoding`.
+
+ Args:
+ mask (Tensor): ByteTensor mask. Non-zero values representing
+ ignored positions, while zero values means valid positions
+ for this image. Shape [bs, h, w].
+ input (Tensor, optional): Input image/feature Tensor.
+ Shape [bs, c, h, w]
+
+ Returns:
+ pos (Tensor): Returned position embedding with shape
+ [bs, num_feats*2, h, w].
+ """
+ assert not (mask is None and input is None)
+
+ if mask is not None:
+ B, H, W = mask.size()
+ device = mask.device
+ # For convenience of exporting to ONNX,
+ # it's required to convert
+ # `masks` from bool to int.
+ mask = mask.to(torch.int)
+ not_mask = 1 - mask # logical_not
+ y_embed = not_mask.cumsum(1, dtype=torch.float32)
+ x_embed = not_mask.cumsum(2, dtype=torch.float32)
+ else:
+ # single image or batch image with no padding
+ B, _, H, W = input.shape
+ device = input.device
+ x_embed = torch.arange(
+ 1, W + 1, dtype=torch.float32, device=device)
+ x_embed = x_embed.view(1, 1, -1).repeat(B, H, 1)
+ y_embed = torch.arange(
+ 1, H + 1, dtype=torch.float32, device=device)
+ y_embed = y_embed.view(1, -1, 1).repeat(B, 1, W)
+ if self.normalize:
+ y_embed = (y_embed + self.offset) / \
+ (y_embed[:, -1:, :] + self.eps) * self.scale
+ x_embed = (x_embed + self.offset) / \
+ (x_embed[:, :, -1:] + self.eps) * self.scale
+ dim_t = torch.arange(
+ self.num_feats, dtype=torch.float32, device=device)
+ dim_t = self.temperature**(2 * (dim_t // 2) / self.num_feats)
+ pos_x = x_embed[:, :, :, None] / dim_t
+ pos_y = y_embed[:, :, :, None] / dim_t
+ # use `view` instead of `flatten` for dynamically exporting to ONNX
+
+ pos_x = torch.stack(
+ (pos_x[:, :, :, 0::2].sin(), pos_x[:, :, :, 1::2].cos()),
+ dim=4).view(B, H, W, -1)
+ pos_y = torch.stack(
+ (pos_y[:, :, :, 0::2].sin(), pos_y[:, :, :, 1::2].cos()),
+ dim=4).view(B, H, W, -1)
+ pos = torch.cat((pos_y, pos_x), dim=3).permute(0, 3, 1, 2)
+ return pos
+
+ def __repr__(self) -> str:
+ """str: a string that describes the module"""
+ repr_str = self.__class__.__name__
+ repr_str += f'(num_feats={self.num_feats}, '
+ repr_str += f'temperature={self.temperature}, '
+ repr_str += f'normalize={self.normalize}, '
+ repr_str += f'scale={self.scale}, '
+ repr_str += f'eps={self.eps})'
+ return repr_str
+
+
+@MODELS.register_module()
+class LearnedPositionalEncoding(BaseModule):
+ """Position embedding with learnable embedding weights.
+
+ Args:
+ num_feats (int): The feature dimension for each position
+ along x-axis or y-axis. The final returned dimension for
+ each position is 2 times of this value.
+ row_num_embed (int, optional): The dictionary size of row embeddings.
+ Defaults to 50.
+ col_num_embed (int, optional): The dictionary size of col embeddings.
+ Defaults to 50.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ """
+
+ def __init__(
+ self,
+ num_feats: int,
+ row_num_embed: int = 50,
+ col_num_embed: int = 50,
+ init_cfg: MultiConfig = dict(type='Uniform', layer='Embedding')
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.row_embed = nn.Embedding(row_num_embed, num_feats)
+ self.col_embed = nn.Embedding(col_num_embed, num_feats)
+ self.num_feats = num_feats
+ self.row_num_embed = row_num_embed
+ self.col_num_embed = col_num_embed
+
+ def forward(self, mask: Tensor) -> Tensor:
+ """Forward function for `LearnedPositionalEncoding`.
+
+ Args:
+ mask (Tensor): ByteTensor mask. Non-zero values representing
+ ignored positions, while zero values means valid positions
+ for this image. Shape [bs, h, w].
+
+ Returns:
+ pos (Tensor): Returned position embedding with shape
+ [bs, num_feats*2, h, w].
+ """
+ h, w = mask.shape[-2:]
+ x = torch.arange(w, device=mask.device)
+ y = torch.arange(h, device=mask.device)
+ x_embed = self.col_embed(x)
+ y_embed = self.row_embed(y)
+ pos = torch.cat(
+ (x_embed.unsqueeze(0).repeat(h, 1, 1), y_embed.unsqueeze(1).repeat(
+ 1, w, 1)),
+ dim=-1).permute(2, 0,
+ 1).unsqueeze(0).repeat(mask.shape[0], 1, 1, 1)
+ return pos
+
+ def __repr__(self) -> str:
+ """str: a string that describes the module"""
+ repr_str = self.__class__.__name__
+ repr_str += f'(num_feats={self.num_feats}, '
+ repr_str += f'row_num_embed={self.row_num_embed}, '
+ repr_str += f'col_num_embed={self.col_num_embed})'
+ return repr_str
+
+
+@MODELS.register_module()
+class SinePositionalEncoding3D(SinePositionalEncoding):
+ """Position encoding with sine and cosine functions.
+
+ See `End-to-End Object Detection with Transformers
+ `_ for details.
+
+ Args:
+ num_feats (int): The feature dimension for each position
+ along x-axis or y-axis. Note the final returned dimension
+ for each position is 2 times of this value.
+ temperature (int, optional): The temperature used for scaling
+ the position embedding. Defaults to 10000.
+ normalize (bool, optional): Whether to normalize the position
+ embedding. Defaults to False.
+ scale (float, optional): A scale factor that scales the position
+ embedding. The scale will be used only when `normalize` is True.
+ Defaults to 2*pi.
+ eps (float, optional): A value added to the denominator for
+ numerical stability. Defaults to 1e-6.
+ offset (float): offset add to embed when do the normalization.
+ Defaults to 0.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def forward(self, mask: Tensor) -> Tensor:
+ """Forward function for `SinePositionalEncoding3D`.
+
+ Args:
+ mask (Tensor): ByteTensor mask. Non-zero values representing
+ ignored positions, while zero values means valid positions
+ for this image. Shape [bs, t, h, w].
+
+ Returns:
+ pos (Tensor): Returned position embedding with shape
+ [bs, num_feats*2, h, w].
+ """
+ assert mask.dim() == 4,\
+ f'{mask.shape} should be a 4-dimensional Tensor,' \
+ f' got {mask.dim()}-dimensional Tensor instead '
+ # For convenience of exporting to ONNX, it's required to convert
+ # `masks` from bool to int.
+ mask = mask.to(torch.int)
+ not_mask = 1 - mask # logical_not
+ z_embed = not_mask.cumsum(1, dtype=torch.float32)
+ y_embed = not_mask.cumsum(2, dtype=torch.float32)
+ x_embed = not_mask.cumsum(3, dtype=torch.float32)
+ if self.normalize:
+ z_embed = (z_embed + self.offset) / \
+ (z_embed[:, -1:, :, :] + self.eps) * self.scale
+ y_embed = (y_embed + self.offset) / \
+ (y_embed[:, :, -1:, :] + self.eps) * self.scale
+ x_embed = (x_embed + self.offset) / \
+ (x_embed[:, :, :, -1:] + self.eps) * self.scale
+ dim_t = torch.arange(
+ self.num_feats, dtype=torch.float32, device=mask.device)
+ dim_t = self.temperature**(2 * (dim_t // 2) / self.num_feats)
+
+ dim_t_z = torch.arange((self.num_feats * 2),
+ dtype=torch.float32,
+ device=mask.device)
+ dim_t_z = self.temperature**(2 * (dim_t_z // 2) / (self.num_feats * 2))
+
+ pos_x = x_embed[:, :, :, :, None] / dim_t
+ pos_y = y_embed[:, :, :, :, None] / dim_t
+ pos_z = z_embed[:, :, :, :, None] / dim_t_z
+ # use `view` instead of `flatten` for dynamically exporting to ONNX
+ B, T, H, W = mask.size()
+ pos_x = torch.stack(
+ (pos_x[:, :, :, :, 0::2].sin(), pos_x[:, :, :, :, 1::2].cos()),
+ dim=5).view(B, T, H, W, -1)
+ pos_y = torch.stack(
+ (pos_y[:, :, :, :, 0::2].sin(), pos_y[:, :, :, :, 1::2].cos()),
+ dim=5).view(B, T, H, W, -1)
+ pos_z = torch.stack(
+ (pos_z[:, :, :, :, 0::2].sin(), pos_z[:, :, :, :, 1::2].cos()),
+ dim=5).view(B, T, H, W, -1)
+ pos = (torch.cat((pos_y, pos_x), dim=4) + pos_z).permute(0, 1, 4, 2, 3)
+ return pos
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/res_layer.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/res_layer.py
new file mode 100644
index 0000000000000000000000000000000000000000..ff24d3e8562d1c3c724b35f7dc10cafe48e47650
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/res_layer.py
@@ -0,0 +1,195 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional
+
+from mmcv.cnn import build_conv_layer, build_norm_layer
+from mmengine.model import BaseModule, Sequential
+from torch import Tensor
+from torch import nn as nn
+
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+
+
+class ResLayer(Sequential):
+ """ResLayer to build ResNet style backbone.
+
+ Args:
+ block (nn.Module): block used to build ResLayer.
+ inplanes (int): inplanes of block.
+ planes (int): planes of block.
+ num_blocks (int): number of blocks.
+ stride (int): stride of the first block. Defaults to 1
+ avg_down (bool): Use AvgPool instead of stride conv when
+ downsampling in the bottleneck. Defaults to False
+ conv_cfg (dict): dictionary to construct and config conv layer.
+ Defaults to None
+ norm_cfg (dict): dictionary to construct and config norm layer.
+ Defaults to dict(type='BN')
+ downsample_first (bool): Downsample at the first block or last block.
+ False for Hourglass, True for ResNet. Defaults to True
+ """
+
+ def __init__(self,
+ block: BaseModule,
+ inplanes: int,
+ planes: int,
+ num_blocks: int,
+ stride: int = 1,
+ avg_down: bool = False,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: ConfigType = dict(type='BN'),
+ downsample_first: bool = True,
+ **kwargs) -> None:
+ self.block = block
+
+ downsample = None
+ if stride != 1 or inplanes != planes * block.expansion:
+ downsample = []
+ conv_stride = stride
+ if avg_down:
+ conv_stride = 1
+ downsample.append(
+ nn.AvgPool2d(
+ kernel_size=stride,
+ stride=stride,
+ ceil_mode=True,
+ count_include_pad=False))
+ downsample.extend([
+ build_conv_layer(
+ conv_cfg,
+ inplanes,
+ planes * block.expansion,
+ kernel_size=1,
+ stride=conv_stride,
+ bias=False),
+ build_norm_layer(norm_cfg, planes * block.expansion)[1]
+ ])
+ downsample = nn.Sequential(*downsample)
+
+ layers = []
+ if downsample_first:
+ layers.append(
+ block(
+ inplanes=inplanes,
+ planes=planes,
+ stride=stride,
+ downsample=downsample,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ **kwargs))
+ inplanes = planes * block.expansion
+ for _ in range(1, num_blocks):
+ layers.append(
+ block(
+ inplanes=inplanes,
+ planes=planes,
+ stride=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ **kwargs))
+
+ else: # downsample_first=False is for HourglassModule
+ for _ in range(num_blocks - 1):
+ layers.append(
+ block(
+ inplanes=inplanes,
+ planes=inplanes,
+ stride=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ **kwargs))
+ layers.append(
+ block(
+ inplanes=inplanes,
+ planes=planes,
+ stride=stride,
+ downsample=downsample,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ **kwargs))
+ super().__init__(*layers)
+
+
+class SimplifiedBasicBlock(BaseModule):
+ """Simplified version of original basic residual block. This is used in
+ `SCNet `_.
+
+ - Norm layer is now optional
+ - Last ReLU in forward function is removed
+ """
+ expansion = 1
+
+ def __init__(self,
+ inplanes: int,
+ planes: int,
+ stride: int = 1,
+ dilation: int = 1,
+ downsample: Optional[Sequential] = None,
+ style: ConfigType = 'pytorch',
+ with_cp: bool = False,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: ConfigType = dict(type='BN'),
+ dcn: OptConfigType = None,
+ plugins: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ assert dcn is None, 'Not implemented yet.'
+ assert plugins is None, 'Not implemented yet.'
+ assert not with_cp, 'Not implemented yet.'
+ self.with_norm = norm_cfg is not None
+ with_bias = True if norm_cfg is None else False
+ self.conv1 = build_conv_layer(
+ conv_cfg,
+ inplanes,
+ planes,
+ 3,
+ stride=stride,
+ padding=dilation,
+ dilation=dilation,
+ bias=with_bias)
+ if self.with_norm:
+ self.norm1_name, norm1 = build_norm_layer(
+ norm_cfg, planes, postfix=1)
+ self.add_module(self.norm1_name, norm1)
+ self.conv2 = build_conv_layer(
+ conv_cfg, planes, planes, 3, padding=1, bias=with_bias)
+ if self.with_norm:
+ self.norm2_name, norm2 = build_norm_layer(
+ norm_cfg, planes, postfix=2)
+ self.add_module(self.norm2_name, norm2)
+
+ self.relu = nn.ReLU(inplace=True)
+ self.downsample = downsample
+ self.stride = stride
+ self.dilation = dilation
+ self.with_cp = with_cp
+
+ @property
+ def norm1(self) -> Optional[BaseModule]:
+ """nn.Module: normalization layer after the first convolution layer"""
+ return getattr(self, self.norm1_name) if self.with_norm else None
+
+ @property
+ def norm2(self) -> Optional[BaseModule]:
+ """nn.Module: normalization layer after the second convolution layer"""
+ return getattr(self, self.norm2_name) if self.with_norm else None
+
+ def forward(self, x: Tensor) -> Tensor:
+ """Forward function for SimplifiedBasicBlock."""
+
+ identity = x
+
+ out = self.conv1(x)
+ if self.with_norm:
+ out = self.norm1(out)
+ out = self.relu(out)
+
+ out = self.conv2(out)
+ if self.with_norm:
+ out = self.norm2(out)
+
+ if self.downsample is not None:
+ identity = self.downsample(x)
+
+ out += identity
+
+ return out
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/se_layer.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/se_layer.py
new file mode 100644
index 0000000000000000000000000000000000000000..5598dabaf6f3b3a09f4348fcd65ff39897b7068f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/se_layer.py
@@ -0,0 +1,162 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from mmengine.model import BaseModule
+from mmengine.utils import digit_version, is_tuple_of
+from torch import Tensor
+
+from mmdet.utils import MultiConfig, OptConfigType, OptMultiConfig
+
+
+class SELayer(BaseModule):
+ """Squeeze-and-Excitation Module.
+
+ Args:
+ channels (int): The input (and output) channels of the SE layer.
+ ratio (int): Squeeze ratio in SELayer, the intermediate channel will be
+ ``int(channels/ratio)``. Defaults to 16.
+ conv_cfg (None or dict): Config dict for convolution layer.
+ Defaults to None, which means using conv2d.
+ act_cfg (dict or Sequence[dict]): Config dict for activation layer.
+ If act_cfg is a dict, two activation layers will be configurated
+ by this dict. If act_cfg is a sequence of dicts, the first
+ activation layer will be configurated by the first dict and the
+ second activation layer will be configurated by the second dict.
+ Defaults to (dict(type='ReLU'), dict(type='Sigmoid'))
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Defaults to None
+ """
+
+ def __init__(self,
+ channels: int,
+ ratio: int = 16,
+ conv_cfg: OptConfigType = None,
+ act_cfg: MultiConfig = (dict(type='ReLU'),
+ dict(type='Sigmoid')),
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ if isinstance(act_cfg, dict):
+ act_cfg = (act_cfg, act_cfg)
+ assert len(act_cfg) == 2
+ assert is_tuple_of(act_cfg, dict)
+ self.global_avgpool = nn.AdaptiveAvgPool2d(1)
+ self.conv1 = ConvModule(
+ in_channels=channels,
+ out_channels=int(channels / ratio),
+ kernel_size=1,
+ stride=1,
+ conv_cfg=conv_cfg,
+ act_cfg=act_cfg[0])
+ self.conv2 = ConvModule(
+ in_channels=int(channels / ratio),
+ out_channels=channels,
+ kernel_size=1,
+ stride=1,
+ conv_cfg=conv_cfg,
+ act_cfg=act_cfg[1])
+
+ def forward(self, x: Tensor) -> Tensor:
+ """Forward function for SELayer."""
+ out = self.global_avgpool(x)
+ out = self.conv1(out)
+ out = self.conv2(out)
+ return x * out
+
+
+class DyReLU(BaseModule):
+ """Dynamic ReLU (DyReLU) module.
+
+ See `Dynamic ReLU `_ for details.
+ Current implementation is specialized for task-aware attention in DyHead.
+ HSigmoid arguments in default act_cfg follow DyHead official code.
+ https://github.com/microsoft/DynamicHead/blob/master/dyhead/dyrelu.py
+
+ Args:
+ channels (int): The input (and output) channels of DyReLU module.
+ ratio (int): Squeeze ratio in Squeeze-and-Excitation-like module,
+ the intermediate channel will be ``int(channels/ratio)``.
+ Defaults to 4.
+ conv_cfg (None or dict): Config dict for convolution layer.
+ Defaults to None, which means using conv2d.
+ act_cfg (dict or Sequence[dict]): Config dict for activation layer.
+ If act_cfg is a dict, two activation layers will be configurated
+ by this dict. If act_cfg is a sequence of dicts, the first
+ activation layer will be configurated by the first dict and the
+ second activation layer will be configurated by the second dict.
+ Defaults to (dict(type='ReLU'), dict(type='HSigmoid', bias=3.0,
+ divisor=6.0))
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Defaults to None
+ """
+
+ def __init__(self,
+ channels: int,
+ ratio: int = 4,
+ conv_cfg: OptConfigType = None,
+ act_cfg: MultiConfig = (dict(type='ReLU'),
+ dict(
+ type='HSigmoid',
+ bias=3.0,
+ divisor=6.0)),
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ if isinstance(act_cfg, dict):
+ act_cfg = (act_cfg, act_cfg)
+ assert len(act_cfg) == 2
+ assert is_tuple_of(act_cfg, dict)
+ self.channels = channels
+ self.expansion = 4 # for a1, b1, a2, b2
+ self.global_avgpool = nn.AdaptiveAvgPool2d(1)
+ self.conv1 = ConvModule(
+ in_channels=channels,
+ out_channels=int(channels / ratio),
+ kernel_size=1,
+ stride=1,
+ conv_cfg=conv_cfg,
+ act_cfg=act_cfg[0])
+ self.conv2 = ConvModule(
+ in_channels=int(channels / ratio),
+ out_channels=channels * self.expansion,
+ kernel_size=1,
+ stride=1,
+ conv_cfg=conv_cfg,
+ act_cfg=act_cfg[1])
+
+ def forward(self, x: Tensor) -> Tensor:
+ """Forward function."""
+ coeffs = self.global_avgpool(x)
+ coeffs = self.conv1(coeffs)
+ coeffs = self.conv2(coeffs) - 0.5 # value range: [-0.5, 0.5]
+ a1, b1, a2, b2 = torch.split(coeffs, self.channels, dim=1)
+ a1 = a1 * 2.0 + 1.0 # [-1.0, 1.0] + 1.0
+ a2 = a2 * 2.0 # [-1.0, 1.0]
+ out = torch.max(x * a1 + b1, x * a2 + b2)
+ return out
+
+
+class ChannelAttention(BaseModule):
+ """Channel attention Module.
+
+ Args:
+ channels (int): The input (and output) channels of the attention layer.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Defaults to None
+ """
+
+ def __init__(self, channels: int, init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.global_avgpool = nn.AdaptiveAvgPool2d(1)
+ self.fc = nn.Conv2d(channels, channels, 1, 1, 0, bias=True)
+ if digit_version(torch.__version__) < (1, 7, 0):
+ self.act = nn.Hardsigmoid()
+ else:
+ self.act = nn.Hardsigmoid(inplace=True)
+
+ def forward(self, x: Tensor) -> Tensor:
+ """Forward function for ChannelAttention."""
+ with torch.cuda.amp.autocast(enabled=False):
+ out = self.global_avgpool(x)
+ out = self.fc(out)
+ out = self.act(out)
+ return x * out
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..839d936412673d765cd9f89a44a366a64976bb9c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/__init__.py
@@ -0,0 +1,41 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .conditional_detr_layers import (ConditionalDetrTransformerDecoder,
+ ConditionalDetrTransformerDecoderLayer)
+from .dab_detr_layers import (DABDetrTransformerDecoder,
+ DABDetrTransformerDecoderLayer,
+ DABDetrTransformerEncoder)
+from .ddq_detr_layers import DDQTransformerDecoder
+from .deformable_detr_layers import (DeformableDetrTransformerDecoder,
+ DeformableDetrTransformerDecoderLayer,
+ DeformableDetrTransformerEncoder,
+ DeformableDetrTransformerEncoderLayer)
+from .detr_layers import (DetrTransformerDecoder, DetrTransformerDecoderLayer,
+ DetrTransformerEncoder, DetrTransformerEncoderLayer)
+from .dino_layers import CdnQueryGenerator, DinoTransformerDecoder
+from .grounding_dino_layers import (GroundingDinoTransformerDecoder,
+ GroundingDinoTransformerDecoderLayer,
+ GroundingDinoTransformerEncoder)
+from .mask2former_layers import (Mask2FormerTransformerDecoder,
+ Mask2FormerTransformerDecoderLayer,
+ Mask2FormerTransformerEncoder)
+from .utils import (MLP, AdaptivePadding, ConditionalAttention, DynamicConv,
+ PatchEmbed, PatchMerging, coordinate_to_encoding,
+ inverse_sigmoid, nchw_to_nlc, nlc_to_nchw)
+
+__all__ = [
+ 'nlc_to_nchw', 'nchw_to_nlc', 'AdaptivePadding', 'PatchEmbed',
+ 'PatchMerging', 'inverse_sigmoid', 'DynamicConv', 'MLP',
+ 'DetrTransformerEncoder', 'DetrTransformerDecoder',
+ 'DetrTransformerEncoderLayer', 'DetrTransformerDecoderLayer',
+ 'DeformableDetrTransformerEncoder', 'DeformableDetrTransformerDecoder',
+ 'DeformableDetrTransformerEncoderLayer',
+ 'DeformableDetrTransformerDecoderLayer', 'coordinate_to_encoding',
+ 'ConditionalAttention', 'DABDetrTransformerDecoderLayer',
+ 'DABDetrTransformerDecoder', 'DABDetrTransformerEncoder',
+ 'DDQTransformerDecoder', 'ConditionalDetrTransformerDecoder',
+ 'ConditionalDetrTransformerDecoderLayer', 'DinoTransformerDecoder',
+ 'CdnQueryGenerator', 'Mask2FormerTransformerEncoder',
+ 'Mask2FormerTransformerDecoderLayer', 'Mask2FormerTransformerDecoder',
+ 'GroundingDinoTransformerDecoderLayer', 'GroundingDinoTransformerEncoder',
+ 'GroundingDinoTransformerDecoder'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/conditional_detr_layers.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/conditional_detr_layers.py
new file mode 100644
index 0000000000000000000000000000000000000000..6db12a1340c758996e8c0e96f0b21cbc6fa928c9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/conditional_detr_layers.py
@@ -0,0 +1,170 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+from mmcv.cnn import build_norm_layer
+from mmcv.cnn.bricks.transformer import FFN
+from torch import Tensor
+from torch.nn import ModuleList
+
+from .detr_layers import DetrTransformerDecoder, DetrTransformerDecoderLayer
+from .utils import MLP, ConditionalAttention, coordinate_to_encoding
+
+
+class ConditionalDetrTransformerDecoder(DetrTransformerDecoder):
+ """Decoder of Conditional DETR."""
+
+ def _init_layers(self) -> None:
+ """Initialize decoder layers and other layers."""
+ self.layers = ModuleList([
+ ConditionalDetrTransformerDecoderLayer(**self.layer_cfg)
+ for _ in range(self.num_layers)
+ ])
+ self.embed_dims = self.layers[0].embed_dims
+ self.post_norm = build_norm_layer(self.post_norm_cfg,
+ self.embed_dims)[1]
+ # conditional detr affline
+ self.query_scale = MLP(self.embed_dims, self.embed_dims,
+ self.embed_dims, 2)
+ self.ref_point_head = MLP(self.embed_dims, self.embed_dims, 2, 2)
+ # we have substitute 'qpos_proj' with 'qpos_sine_proj' except for
+ # the first decoder layer), so 'qpos_proj' should be deleted
+ # in other layers.
+ for layer_id in range(self.num_layers - 1):
+ self.layers[layer_id + 1].cross_attn.qpos_proj = None
+
+ def forward(self,
+ query: Tensor,
+ key: Tensor = None,
+ query_pos: Tensor = None,
+ key_pos: Tensor = None,
+ key_padding_mask: Tensor = None):
+ """Forward function of decoder.
+
+ Args:
+ query (Tensor): The input query with shape
+ (bs, num_queries, dim).
+ key (Tensor): The input key with shape (bs, num_keys, dim) If
+ `None`, the `query` will be used. Defaults to `None`.
+ query_pos (Tensor): The positional encoding for `query`, with the
+ same shape as `query`. If not `None`, it will be added to
+ `query` before forward function. Defaults to `None`.
+ key_pos (Tensor): The positional encoding for `key`, with the
+ same shape as `key`. If not `None`, it will be added to
+ `key` before forward function. If `None`, and `query_pos`
+ has the same shape as `key`, then `query_pos` will be used
+ as `key_pos`. Defaults to `None`.
+ key_padding_mask (Tensor): ByteTensor with shape (bs, num_keys).
+ Defaults to `None`.
+ Returns:
+ List[Tensor]: forwarded results with shape (num_decoder_layers,
+ bs, num_queries, dim) if `return_intermediate` is True, otherwise
+ with shape (1, bs, num_queries, dim). References with shape
+ (bs, num_queries, 2).
+ """
+ reference_unsigmoid = self.ref_point_head(
+ query_pos) # [bs, num_queries, 2]
+ reference = reference_unsigmoid.sigmoid()
+ reference_xy = reference[..., :2]
+ intermediate = []
+ for layer_id, layer in enumerate(self.layers):
+ if layer_id == 0:
+ pos_transformation = 1
+ else:
+ pos_transformation = self.query_scale(query)
+ # get sine embedding for the query reference
+ ref_sine_embed = coordinate_to_encoding(coord_tensor=reference_xy)
+ # apply transformation
+ ref_sine_embed = ref_sine_embed * pos_transformation
+ query = layer(
+ query,
+ key=key,
+ query_pos=query_pos,
+ key_pos=key_pos,
+ key_padding_mask=key_padding_mask,
+ ref_sine_embed=ref_sine_embed,
+ is_first=(layer_id == 0))
+ if self.return_intermediate:
+ intermediate.append(self.post_norm(query))
+
+ if self.return_intermediate:
+ return torch.stack(intermediate), reference
+
+ query = self.post_norm(query)
+ return query.unsqueeze(0), reference
+
+
+class ConditionalDetrTransformerDecoderLayer(DetrTransformerDecoderLayer):
+ """Implements decoder layer in Conditional DETR transformer."""
+
+ def _init_layers(self):
+ """Initialize self-attention, cross-attention, FFN, and
+ normalization."""
+ self.self_attn = ConditionalAttention(**self.self_attn_cfg)
+ self.cross_attn = ConditionalAttention(**self.cross_attn_cfg)
+ self.embed_dims = self.self_attn.embed_dims
+ self.ffn = FFN(**self.ffn_cfg)
+ norms_list = [
+ build_norm_layer(self.norm_cfg, self.embed_dims)[1]
+ for _ in range(3)
+ ]
+ self.norms = ModuleList(norms_list)
+
+ def forward(self,
+ query: Tensor,
+ key: Tensor = None,
+ query_pos: Tensor = None,
+ key_pos: Tensor = None,
+ self_attn_masks: Tensor = None,
+ cross_attn_masks: Tensor = None,
+ key_padding_mask: Tensor = None,
+ ref_sine_embed: Tensor = None,
+ is_first: bool = False):
+ """
+ Args:
+ query (Tensor): The input query, has shape (bs, num_queries, dim)
+ key (Tensor, optional): The input key, has shape (bs, num_keys,
+ dim). If `None`, the `query` will be used. Defaults to `None`.
+ query_pos (Tensor, optional): The positional encoding for `query`,
+ has the same shape as `query`. If not `None`, it will be
+ added to `query` before forward function. Defaults to `None`.
+ ref_sine_embed (Tensor): The positional encoding for query in
+ cross attention, with the same shape as `x`. Defaults to None.
+ key_pos (Tensor, optional): The positional encoding for `key`, has
+ the same shape as `key`. If not None, it will be added to
+ `key` before forward function. If None, and `query_pos` has
+ the same shape as `key`, then `query_pos` will be used for
+ `key_pos`. Defaults to None.
+ self_attn_masks (Tensor, optional): ByteTensor mask, has shape
+ (num_queries, num_keys), Same in `nn.MultiheadAttention.
+ forward`. Defaults to None.
+ cross_attn_masks (Tensor, optional): ByteTensor mask, has shape
+ (num_queries, num_keys), Same in `nn.MultiheadAttention.
+ forward`. Defaults to None.
+ key_padding_mask (Tensor, optional): ByteTensor, has shape
+ (bs, num_keys). Defaults to None.
+ is_first (bool): A indicator to tell whether the current layer
+ is the first layer of the decoder. Defaults to False.
+
+ Returns:
+ Tensor: Forwarded results, has shape (bs, num_queries, dim).
+ """
+ query = self.self_attn(
+ query=query,
+ key=query,
+ query_pos=query_pos,
+ key_pos=query_pos,
+ attn_mask=self_attn_masks)
+ query = self.norms[0](query)
+ query = self.cross_attn(
+ query=query,
+ key=key,
+ query_pos=query_pos,
+ key_pos=key_pos,
+ attn_mask=cross_attn_masks,
+ key_padding_mask=key_padding_mask,
+ ref_sine_embed=ref_sine_embed,
+ is_first=is_first)
+ query = self.norms[1](query)
+ query = self.ffn(query)
+ query = self.norms[2](query)
+
+ return query
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/dab_detr_layers.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/dab_detr_layers.py
new file mode 100644
index 0000000000000000000000000000000000000000..b8a6e7724a1b1ca18f26dd10455f3e3a4d696460
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/dab_detr_layers.py
@@ -0,0 +1,298 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import build_norm_layer
+from mmcv.cnn.bricks.transformer import FFN
+from mmengine.model import ModuleList
+from torch import Tensor
+
+from .detr_layers import (DetrTransformerDecoder, DetrTransformerDecoderLayer,
+ DetrTransformerEncoder, DetrTransformerEncoderLayer)
+from .utils import (MLP, ConditionalAttention, coordinate_to_encoding,
+ inverse_sigmoid)
+
+
+class DABDetrTransformerDecoderLayer(DetrTransformerDecoderLayer):
+ """Implements decoder layer in DAB-DETR transformer."""
+
+ def _init_layers(self):
+ """Initialize self-attention, cross-attention, FFN, normalization and
+ others."""
+ self.self_attn = ConditionalAttention(**self.self_attn_cfg)
+ self.cross_attn = ConditionalAttention(**self.cross_attn_cfg)
+ self.embed_dims = self.self_attn.embed_dims
+ self.ffn = FFN(**self.ffn_cfg)
+ norms_list = [
+ build_norm_layer(self.norm_cfg, self.embed_dims)[1]
+ for _ in range(3)
+ ]
+ self.norms = ModuleList(norms_list)
+ self.keep_query_pos = self.cross_attn.keep_query_pos
+
+ def forward(self,
+ query: Tensor,
+ key: Tensor,
+ query_pos: Tensor,
+ key_pos: Tensor,
+ ref_sine_embed: Tensor = None,
+ self_attn_masks: Tensor = None,
+ cross_attn_masks: Tensor = None,
+ key_padding_mask: Tensor = None,
+ is_first: bool = False,
+ **kwargs) -> Tensor:
+ """
+ Args:
+ query (Tensor): The input query with shape [bs, num_queries,
+ dim].
+ key (Tensor): The key tensor with shape [bs, num_keys,
+ dim].
+ query_pos (Tensor): The positional encoding for query in self
+ attention, with the same shape as `x`.
+ key_pos (Tensor): The positional encoding for `key`, with the
+ same shape as `key`.
+ ref_sine_embed (Tensor): The positional encoding for query in
+ cross attention, with the same shape as `x`.
+ Defaults to None.
+ self_attn_masks (Tensor): ByteTensor mask with shape [num_queries,
+ num_keys]. Same in `nn.MultiheadAttention.forward`.
+ Defaults to None.
+ cross_attn_masks (Tensor): ByteTensor mask with shape [num_queries,
+ num_keys]. Same in `nn.MultiheadAttention.forward`.
+ Defaults to None.
+ key_padding_mask (Tensor): ByteTensor with shape [bs, num_keys].
+ Defaults to None.
+ is_first (bool): A indicator to tell whether the current layer
+ is the first layer of the decoder.
+ Defaults to False.
+
+ Returns:
+ Tensor: forwarded results with shape
+ [bs, num_queries, dim].
+ """
+
+ query = self.self_attn(
+ query=query,
+ key=query,
+ query_pos=query_pos,
+ key_pos=query_pos,
+ attn_mask=self_attn_masks,
+ **kwargs)
+ query = self.norms[0](query)
+ query = self.cross_attn(
+ query=query,
+ key=key,
+ query_pos=query_pos,
+ key_pos=key_pos,
+ ref_sine_embed=ref_sine_embed,
+ attn_mask=cross_attn_masks,
+ key_padding_mask=key_padding_mask,
+ is_first=is_first,
+ **kwargs)
+ query = self.norms[1](query)
+ query = self.ffn(query)
+ query = self.norms[2](query)
+
+ return query
+
+
+class DABDetrTransformerDecoder(DetrTransformerDecoder):
+ """Decoder of DAB-DETR.
+
+ Args:
+ query_dim (int): The last dimension of query pos,
+ 4 for anchor format, 2 for point format.
+ Defaults to 4.
+ query_scale_type (str): Type of transformation applied
+ to content query. Defaults to `cond_elewise`.
+ with_modulated_hw_attn (bool): Whether to inject h&w info
+ during cross conditional attention. Defaults to True.
+ """
+
+ def __init__(self,
+ *args,
+ query_dim: int = 4,
+ query_scale_type: str = 'cond_elewise',
+ with_modulated_hw_attn: bool = True,
+ **kwargs):
+
+ self.query_dim = query_dim
+ self.query_scale_type = query_scale_type
+ self.with_modulated_hw_attn = with_modulated_hw_attn
+
+ super().__init__(*args, **kwargs)
+
+ def _init_layers(self):
+ """Initialize decoder layers and other layers."""
+ assert self.query_dim in [2, 4], \
+ f'{"dab-detr only supports anchor prior or reference point prior"}'
+ assert self.query_scale_type in [
+ 'cond_elewise', 'cond_scalar', 'fix_elewise'
+ ]
+
+ self.layers = ModuleList([
+ DABDetrTransformerDecoderLayer(**self.layer_cfg)
+ for _ in range(self.num_layers)
+ ])
+
+ embed_dims = self.layers[0].embed_dims
+ self.embed_dims = embed_dims
+
+ self.post_norm = build_norm_layer(self.post_norm_cfg, embed_dims)[1]
+ if self.query_scale_type == 'cond_elewise':
+ self.query_scale = MLP(embed_dims, embed_dims, embed_dims, 2)
+ elif self.query_scale_type == 'cond_scalar':
+ self.query_scale = MLP(embed_dims, embed_dims, 1, 2)
+ elif self.query_scale_type == 'fix_elewise':
+ self.query_scale = nn.Embedding(self.num_layers, embed_dims)
+ else:
+ raise NotImplementedError('Unknown query_scale_type: {}'.format(
+ self.query_scale_type))
+
+ self.ref_point_head = MLP(self.query_dim // 2 * embed_dims, embed_dims,
+ embed_dims, 2)
+
+ if self.with_modulated_hw_attn and self.query_dim == 4:
+ self.ref_anchor_head = MLP(embed_dims, embed_dims, 2, 2)
+
+ self.keep_query_pos = self.layers[0].keep_query_pos
+ if not self.keep_query_pos:
+ for layer_id in range(self.num_layers - 1):
+ self.layers[layer_id + 1].cross_attn.qpos_proj = None
+
+ def forward(self,
+ query: Tensor,
+ key: Tensor,
+ query_pos: Tensor,
+ key_pos: Tensor,
+ reg_branches: nn.Module,
+ key_padding_mask: Tensor = None,
+ **kwargs) -> List[Tensor]:
+ """Forward function of decoder.
+
+ Args:
+ query (Tensor): The input query with shape (bs, num_queries, dim).
+ key (Tensor): The input key with shape (bs, num_keys, dim).
+ query_pos (Tensor): The positional encoding for `query`, with the
+ same shape as `query`.
+ key_pos (Tensor): The positional encoding for `key`, with the
+ same shape as `key`.
+ reg_branches (nn.Module): The regression branch for dynamically
+ updating references in each layer.
+ key_padding_mask (Tensor): ByteTensor with shape (bs, num_keys).
+ Defaults to `None`.
+
+ Returns:
+ List[Tensor]: forwarded results with shape (num_decoder_layers,
+ bs, num_queries, dim) if `return_intermediate` is True, otherwise
+ with shape (1, bs, num_queries, dim). references with shape
+ (num_decoder_layers, bs, num_queries, 2/4).
+ """
+ output = query
+ unsigmoid_references = query_pos
+
+ reference_points = unsigmoid_references.sigmoid()
+ intermediate_reference_points = [reference_points]
+
+ intermediate = []
+ for layer_id, layer in enumerate(self.layers):
+ obj_center = reference_points[..., :self.query_dim]
+ ref_sine_embed = coordinate_to_encoding(
+ coord_tensor=obj_center, num_feats=self.embed_dims // 2)
+ query_pos = self.ref_point_head(
+ ref_sine_embed) # [bs, nq, 2c] -> [bs, nq, c]
+ # For the first decoder layer, do not apply transformation
+ if self.query_scale_type != 'fix_elewise':
+ if layer_id == 0:
+ pos_transformation = 1
+ else:
+ pos_transformation = self.query_scale(output)
+ else:
+ pos_transformation = self.query_scale.weight[layer_id]
+ # apply transformation
+ ref_sine_embed = ref_sine_embed[
+ ..., :self.embed_dims] * pos_transformation
+ # modulated height and weight attention
+ if self.with_modulated_hw_attn:
+ assert obj_center.size(-1) == 4
+ ref_hw = self.ref_anchor_head(output).sigmoid()
+ ref_sine_embed[..., self.embed_dims // 2:] *= \
+ (ref_hw[..., 0] / obj_center[..., 2]).unsqueeze(-1)
+ ref_sine_embed[..., : self.embed_dims // 2] *= \
+ (ref_hw[..., 1] / obj_center[..., 3]).unsqueeze(-1)
+
+ output = layer(
+ output,
+ key,
+ query_pos=query_pos,
+ ref_sine_embed=ref_sine_embed,
+ key_pos=key_pos,
+ key_padding_mask=key_padding_mask,
+ is_first=(layer_id == 0),
+ **kwargs)
+ # iter update
+ tmp_reg_preds = reg_branches(output)
+ tmp_reg_preds[..., :self.query_dim] += inverse_sigmoid(
+ reference_points)
+ new_reference_points = tmp_reg_preds[
+ ..., :self.query_dim].sigmoid()
+ if layer_id != self.num_layers - 1:
+ intermediate_reference_points.append(new_reference_points)
+ reference_points = new_reference_points.detach()
+
+ if self.return_intermediate:
+ intermediate.append(self.post_norm(output))
+
+ output = self.post_norm(output)
+
+ if self.return_intermediate:
+ return [
+ torch.stack(intermediate),
+ torch.stack(intermediate_reference_points),
+ ]
+ else:
+ return [
+ output.unsqueeze(0),
+ torch.stack(intermediate_reference_points)
+ ]
+
+
+class DABDetrTransformerEncoder(DetrTransformerEncoder):
+ """Encoder of DAB-DETR."""
+
+ def _init_layers(self):
+ """Initialize encoder layers."""
+ self.layers = ModuleList([
+ DetrTransformerEncoderLayer(**self.layer_cfg)
+ for _ in range(self.num_layers)
+ ])
+ embed_dims = self.layers[0].embed_dims
+ self.embed_dims = embed_dims
+ self.query_scale = MLP(embed_dims, embed_dims, embed_dims, 2)
+
+ def forward(self, query: Tensor, query_pos: Tensor,
+ key_padding_mask: Tensor, **kwargs):
+ """Forward function of encoder.
+
+ Args:
+ query (Tensor): Input queries of encoder, has shape
+ (bs, num_queries, dim).
+ query_pos (Tensor): The positional embeddings of the queries, has
+ shape (bs, num_feat_points, dim).
+ key_padding_mask (Tensor): ByteTensor, the key padding mask
+ of the queries, has shape (bs, num_feat_points).
+
+ Returns:
+ Tensor: With shape (num_queries, bs, dim).
+ """
+
+ for layer in self.layers:
+ pos_scales = self.query_scale(query)
+ query = layer(
+ query,
+ query_pos=query_pos * pos_scales,
+ key_padding_mask=key_padding_mask,
+ **kwargs)
+
+ return query
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/ddq_detr_layers.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/ddq_detr_layers.py
new file mode 100644
index 0000000000000000000000000000000000000000..57664c7ea2bdd17681ccdabe9140eb043a99e155
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/ddq_detr_layers.py
@@ -0,0 +1,223 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+
+import torch
+from mmcv.ops import batched_nms
+from torch import Tensor, nn
+
+from mmdet.structures.bbox import bbox_cxcywh_to_xyxy
+from .deformable_detr_layers import DeformableDetrTransformerDecoder
+from .utils import MLP, coordinate_to_encoding, inverse_sigmoid
+
+
+class DDQTransformerDecoder(DeformableDetrTransformerDecoder):
+ """Transformer decoder of DDQ."""
+
+ def _init_layers(self) -> None:
+ """Initialize encoder layers."""
+ super()._init_layers()
+ self.ref_point_head = MLP(self.embed_dims * 2, self.embed_dims,
+ self.embed_dims, 2)
+ self.norm = nn.LayerNorm(self.embed_dims)
+
+ def select_distinct_queries(self, reference_points: Tensor, query: Tensor,
+ self_attn_mask: Tensor, layer_index):
+ """Get updated `self_attn_mask` for distinct queries selection, it is
+ used in self attention layers of decoder.
+
+ Args:
+ reference_points (Tensor): The input reference of decoder,
+ has shape (bs, num_queries, 4) with the last dimension
+ arranged as (cx, cy, w, h).
+ query (Tensor): The input query of decoder, has shape
+ (bs, num_queries, dims).
+ self_attn_mask (Tensor): The input self attention mask of
+ last decoder layer, has shape (bs, num_queries_total,
+ num_queries_total).
+ layer_index (int): Last decoder layer index, used to get
+ classification score of last layer output, for
+ distinct queries selection.
+
+ Returns:
+ Tensor: `self_attn_mask` used in self attention layers
+ of decoder, has shape (bs, num_queries_total,
+ num_queries_total).
+ """
+ num_imgs = len(reference_points)
+ dis_start, num_dis = self.cache_dict['dis_query_info']
+ # shape of self_attn_mask
+ # (batch⋅num_heads, num_queries, embed_dims)
+ dis_mask = self_attn_mask[:, dis_start:dis_start + num_dis,
+ dis_start:dis_start + num_dis]
+ # cls_branches from DDQDETRHead
+ scores = self.cache_dict['cls_branches'][layer_index](
+ query[:, dis_start:dis_start + num_dis]).sigmoid().max(-1).values
+ proposals = reference_points[:, dis_start:dis_start + num_dis]
+ proposals = bbox_cxcywh_to_xyxy(proposals)
+
+ attn_mask_list = []
+ for img_id in range(num_imgs):
+ single_proposals = proposals[img_id]
+ single_scores = scores[img_id]
+ attn_mask = ~dis_mask[img_id * self.cache_dict['num_heads']][0]
+ # distinct query inds in this layer
+ ori_index = attn_mask.nonzero().view(-1)
+ _, keep_idxs = batched_nms(single_proposals[ori_index],
+ single_scores[ori_index],
+ torch.ones(len(ori_index)),
+ self.cache_dict['dqs_cfg'])
+
+ real_keep_index = ori_index[keep_idxs]
+
+ attn_mask = torch.ones_like(dis_mask[0]).bool()
+ # such a attn_mask give best result
+ # If it requires to keep index i, then all cells in row or column
+ # i should be kept in `attn_mask` . For example, if
+ # `real_keep_index` = [1, 4], and `attn_mask` size = [8, 8],
+ # then all cells at rows or columns [1, 4] should be kept, and
+ # all the other cells should be masked out. So the value of
+ # `attn_mask` should be:
+ #
+ # target\source 0 1 2 3 4 5 6 7
+ # 0 [ 0 1 0 0 1 0 0 0 ]
+ # 1 [ 1 1 1 1 1 1 1 1 ]
+ # 2 [ 0 1 0 0 1 0 0 0 ]
+ # 3 [ 0 1 0 0 1 0 0 0 ]
+ # 4 [ 1 1 1 1 1 1 1 1 ]
+ # 5 [ 0 1 0 0 1 0 0 0 ]
+ # 6 [ 0 1 0 0 1 0 0 0 ]
+ # 7 [ 0 1 0 0 1 0 0 0 ]
+ attn_mask[real_keep_index] = False
+ attn_mask[:, real_keep_index] = False
+
+ attn_mask = attn_mask[None].repeat(self.cache_dict['num_heads'], 1,
+ 1)
+ attn_mask_list.append(attn_mask)
+ attn_mask = torch.cat(attn_mask_list)
+ self_attn_mask = copy.deepcopy(self_attn_mask)
+ self_attn_mask[:, dis_start:dis_start + num_dis,
+ dis_start:dis_start + num_dis] = attn_mask
+ # will be used in loss and inference
+ self.cache_dict['distinct_query_mask'].append(~attn_mask)
+ return self_attn_mask
+
+ def forward(self, query: Tensor, value: Tensor, key_padding_mask: Tensor,
+ self_attn_mask: Tensor, reference_points: Tensor,
+ spatial_shapes: Tensor, level_start_index: Tensor,
+ valid_ratios: Tensor, reg_branches: nn.ModuleList,
+ **kwargs) -> Tensor:
+ """Forward function of Transformer decoder.
+
+ Args:
+ query (Tensor): The input query, has shape (bs, num_queries,
+ dims).
+ value (Tensor): The input values, has shape (bs, num_value, dim).
+ key_padding_mask (Tensor): The `key_padding_mask` of `cross_attn`
+ input. ByteTensor, has shape (bs, num_value).
+ self_attn_mask (Tensor): The attention mask to prevent information
+ leakage from different denoising groups, distinct queries and
+ dense queries, has shape (num_queries_total,
+ num_queries_total). It will be updated for distinct queries
+ selection in this forward function. It is `None` when
+ `self.training` is `False`.
+ reference_points (Tensor): The initial reference, has shape
+ (bs, num_queries, 4) with the last dimension arranged as
+ (cx, cy, w, h).
+ spatial_shapes (Tensor): Spatial shapes of features in all levels,
+ has shape (num_levels, 2), last dimension represents (h, w).
+ level_start_index (Tensor): The start index of each level.
+ A tensor has shape (num_levels, ) and can be represented
+ as [0, h_0*w_0, h_0*w_0+h_1*w_1, ...].
+ valid_ratios (Tensor): The ratios of the valid width and the valid
+ height relative to the width and the height of features in all
+ levels, has shape (bs, num_levels, 2).
+ reg_branches: (obj:`nn.ModuleList`): Used for refining the
+ regression results.
+
+ Returns:
+ tuple[Tensor]: Output queries and references of Transformer
+ decoder
+
+ - query (Tensor): Output embeddings of the last decoder, has
+ shape (bs, num_queries, embed_dims) when `return_intermediate`
+ is `False`. Otherwise, Intermediate output embeddings of all
+ decoder layers, has shape (num_decoder_layers, bs, num_queries,
+ embed_dims).
+ - reference_points (Tensor): The reference of the last decoder
+ layer, has shape (bs, num_queries, 4) when `return_intermediate`
+ is `False`. Otherwise, Intermediate references of all decoder
+ layers, has shape (1 + num_decoder_layers, bs, num_queries, 4).
+ The coordinates are arranged as (cx, cy, w, h).
+ """
+ intermediate = []
+ intermediate_reference_points = [reference_points]
+ self.cache_dict['distinct_query_mask'] = []
+ if self_attn_mask is None:
+ self_attn_mask = torch.zeros((query.size(1), query.size(1)),
+ device=query.device).bool()
+ # shape is (batch*number_heads, num_queries, num_queries)
+ self_attn_mask = self_attn_mask[None].repeat(
+ len(query) * self.cache_dict['num_heads'], 1, 1)
+ for layer_index, layer in enumerate(self.layers):
+ if reference_points.shape[-1] == 4:
+ reference_points_input = \
+ reference_points[:, :, None] * torch.cat(
+ [valid_ratios, valid_ratios], -1)[:, None]
+ else:
+ assert reference_points.shape[-1] == 2
+ reference_points_input = \
+ reference_points[:, :, None] * valid_ratios[:, None]
+
+ query_sine_embed = coordinate_to_encoding(
+ reference_points_input[:, :, 0, :],
+ num_feats=self.embed_dims // 2)
+ query_pos = self.ref_point_head(query_sine_embed)
+
+ query = layer(
+ query,
+ query_pos=query_pos,
+ value=value,
+ key_padding_mask=key_padding_mask,
+ self_attn_mask=self_attn_mask,
+ spatial_shapes=spatial_shapes,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios,
+ reference_points=reference_points_input,
+ **kwargs)
+
+ if not self.training:
+ tmp = reg_branches[layer_index](query)
+ assert reference_points.shape[-1] == 4
+ new_reference_points = tmp + inverse_sigmoid(
+ reference_points, eps=1e-3)
+ new_reference_points = new_reference_points.sigmoid()
+ reference_points = new_reference_points.detach()
+ if layer_index < (len(self.layers) - 1):
+ self_attn_mask = self.select_distinct_queries(
+ reference_points, query, self_attn_mask, layer_index)
+
+ else:
+ num_dense = self.cache_dict['num_dense_queries']
+ tmp = reg_branches[layer_index](query[:, :-num_dense])
+ tmp_dense = self.aux_reg_branches[layer_index](
+ query[:, -num_dense:])
+
+ tmp = torch.cat([tmp, tmp_dense], dim=1)
+ assert reference_points.shape[-1] == 4
+ new_reference_points = tmp + inverse_sigmoid(
+ reference_points, eps=1e-3)
+ new_reference_points = new_reference_points.sigmoid()
+ reference_points = new_reference_points.detach()
+ if layer_index < (len(self.layers) - 1):
+ self_attn_mask = self.select_distinct_queries(
+ reference_points, query, self_attn_mask, layer_index)
+
+ if self.return_intermediate:
+ intermediate.append(self.norm(query))
+ intermediate_reference_points.append(new_reference_points)
+
+ if self.return_intermediate:
+ return torch.stack(intermediate), torch.stack(
+ intermediate_reference_points)
+
+ return query, reference_points
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/deformable_detr_layers.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/deformable_detr_layers.py
new file mode 100644
index 0000000000000000000000000000000000000000..da6325d61270eb3546a39d5487587bc0610434d6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/deformable_detr_layers.py
@@ -0,0 +1,265 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Tuple, Union
+
+import torch
+from mmcv.cnn import build_norm_layer
+from mmcv.cnn.bricks.transformer import FFN, MultiheadAttention
+from mmcv.ops import MultiScaleDeformableAttention
+from mmengine.model import ModuleList
+from torch import Tensor, nn
+
+from .detr_layers import (DetrTransformerDecoder, DetrTransformerDecoderLayer,
+ DetrTransformerEncoder, DetrTransformerEncoderLayer)
+from .utils import inverse_sigmoid
+
+try:
+ from fairscale.nn.checkpoint import checkpoint_wrapper
+except Exception:
+ checkpoint_wrapper = None
+
+
+class DeformableDetrTransformerEncoder(DetrTransformerEncoder):
+ """Transformer encoder of Deformable DETR."""
+
+ def _init_layers(self) -> None:
+ """Initialize encoder layers."""
+ self.layers = ModuleList([
+ DeformableDetrTransformerEncoderLayer(**self.layer_cfg)
+ for _ in range(self.num_layers)
+ ])
+
+ if self.num_cp > 0:
+ if checkpoint_wrapper is None:
+ raise NotImplementedError(
+ 'If you want to reduce GPU memory usage, \
+ please install fairscale by executing the \
+ following command: pip install fairscale.')
+ for i in range(self.num_cp):
+ self.layers[i] = checkpoint_wrapper(self.layers[i])
+
+ self.embed_dims = self.layers[0].embed_dims
+
+ def forward(self, query: Tensor, query_pos: Tensor,
+ key_padding_mask: Tensor, spatial_shapes: Tensor,
+ level_start_index: Tensor, valid_ratios: Tensor,
+ **kwargs) -> Tensor:
+ """Forward function of Transformer encoder.
+
+ Args:
+ query (Tensor): The input query, has shape (bs, num_queries, dim).
+ query_pos (Tensor): The positional encoding for query, has shape
+ (bs, num_queries, dim).
+ key_padding_mask (Tensor): The `key_padding_mask` of `self_attn`
+ input. ByteTensor, has shape (bs, num_queries).
+ spatial_shapes (Tensor): Spatial shapes of features in all levels,
+ has shape (num_levels, 2), last dimension represents (h, w).
+ level_start_index (Tensor): The start index of each level.
+ A tensor has shape (num_levels, ) and can be represented
+ as [0, h_0*w_0, h_0*w_0+h_1*w_1, ...].
+ valid_ratios (Tensor): The ratios of the valid width and the valid
+ height relative to the width and the height of features in all
+ levels, has shape (bs, num_levels, 2).
+
+ Returns:
+ Tensor: Output queries of Transformer encoder, which is also
+ called 'encoder output embeddings' or 'memory', has shape
+ (bs, num_queries, dim)
+ """
+ reference_points = self.get_encoder_reference_points(
+ spatial_shapes, valid_ratios, device=query.device)
+ for layer in self.layers:
+ query = layer(
+ query=query,
+ query_pos=query_pos,
+ key_padding_mask=key_padding_mask,
+ spatial_shapes=spatial_shapes,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios,
+ reference_points=reference_points,
+ **kwargs)
+ return query
+
+ @staticmethod
+ def get_encoder_reference_points(
+ spatial_shapes: Tensor, valid_ratios: Tensor,
+ device: Union[torch.device, str]) -> Tensor:
+ """Get the reference points used in encoder.
+
+ Args:
+ spatial_shapes (Tensor): Spatial shapes of features in all levels,
+ has shape (num_levels, 2), last dimension represents (h, w).
+ valid_ratios (Tensor): The ratios of the valid width and the valid
+ height relative to the width and the height of features in all
+ levels, has shape (bs, num_levels, 2).
+ device (obj:`device` or str): The device acquired by the
+ `reference_points`.
+
+ Returns:
+ Tensor: Reference points used in decoder, has shape (bs, length,
+ num_levels, 2).
+ """
+
+ reference_points_list = []
+ for lvl, (H, W) in enumerate(spatial_shapes):
+ ref_y, ref_x = torch.meshgrid(
+ torch.linspace(
+ 0.5, H - 0.5, H, dtype=torch.float32, device=device),
+ torch.linspace(
+ 0.5, W - 0.5, W, dtype=torch.float32, device=device))
+ ref_y = ref_y.reshape(-1)[None] / (
+ valid_ratios[:, None, lvl, 1] * H)
+ ref_x = ref_x.reshape(-1)[None] / (
+ valid_ratios[:, None, lvl, 0] * W)
+ ref = torch.stack((ref_x, ref_y), -1)
+ reference_points_list.append(ref)
+ reference_points = torch.cat(reference_points_list, 1)
+ # [bs, sum(hw), num_level, 2]
+ reference_points = reference_points[:, :, None] * valid_ratios[:, None]
+ return reference_points
+
+
+class DeformableDetrTransformerDecoder(DetrTransformerDecoder):
+ """Transformer Decoder of Deformable DETR."""
+
+ def _init_layers(self) -> None:
+ """Initialize decoder layers."""
+ self.layers = ModuleList([
+ DeformableDetrTransformerDecoderLayer(**self.layer_cfg)
+ for _ in range(self.num_layers)
+ ])
+ self.embed_dims = self.layers[0].embed_dims
+ if self.post_norm_cfg is not None:
+ raise ValueError('There is not post_norm in '
+ f'{self._get_name()}')
+
+ def forward(self,
+ query: Tensor,
+ query_pos: Tensor,
+ value: Tensor,
+ key_padding_mask: Tensor,
+ reference_points: Tensor,
+ spatial_shapes: Tensor,
+ level_start_index: Tensor,
+ valid_ratios: Tensor,
+ reg_branches: Optional[nn.Module] = None,
+ **kwargs) -> Tuple[Tensor]:
+ """Forward function of Transformer decoder.
+
+ Args:
+ query (Tensor): The input queries, has shape (bs, num_queries,
+ dim).
+ query_pos (Tensor): The input positional query, has shape
+ (bs, num_queries, dim). It will be added to `query` before
+ forward function.
+ value (Tensor): The input values, has shape (bs, num_value, dim).
+ key_padding_mask (Tensor): The `key_padding_mask` of `cross_attn`
+ input. ByteTensor, has shape (bs, num_value).
+ reference_points (Tensor): The initial reference, has shape
+ (bs, num_queries, 4) with the last dimension arranged as
+ (cx, cy, w, h) when `as_two_stage` is `True`, otherwise has
+ shape (bs, num_queries, 2) with the last dimension arranged
+ as (cx, cy).
+ spatial_shapes (Tensor): Spatial shapes of features in all levels,
+ has shape (num_levels, 2), last dimension represents (h, w).
+ level_start_index (Tensor): The start index of each level.
+ A tensor has shape (num_levels, ) and can be represented
+ as [0, h_0*w_0, h_0*w_0+h_1*w_1, ...].
+ valid_ratios (Tensor): The ratios of the valid width and the valid
+ height relative to the width and the height of features in all
+ levels, has shape (bs, num_levels, 2).
+ reg_branches: (obj:`nn.ModuleList`, optional): Used for refining
+ the regression results. Only would be passed when
+ `with_box_refine` is `True`, otherwise would be `None`.
+
+ Returns:
+ tuple[Tensor]: Outputs of Deformable Transformer Decoder.
+
+ - output (Tensor): Output embeddings of the last decoder, has
+ shape (num_queries, bs, embed_dims) when `return_intermediate`
+ is `False`. Otherwise, Intermediate output embeddings of all
+ decoder layers, has shape (num_decoder_layers, num_queries, bs,
+ embed_dims).
+ - reference_points (Tensor): The reference of the last decoder
+ layer, has shape (bs, num_queries, 4) when `return_intermediate`
+ is `False`. Otherwise, Intermediate references of all decoder
+ layers, has shape (num_decoder_layers, bs, num_queries, 4). The
+ coordinates are arranged as (cx, cy, w, h)
+ """
+ output = query
+ intermediate = []
+ intermediate_reference_points = []
+ for layer_id, layer in enumerate(self.layers):
+ if reference_points.shape[-1] == 4:
+ reference_points_input = \
+ reference_points[:, :, None] * \
+ torch.cat([valid_ratios, valid_ratios], -1)[:, None]
+ else:
+ assert reference_points.shape[-1] == 2
+ reference_points_input = \
+ reference_points[:, :, None] * \
+ valid_ratios[:, None]
+ output = layer(
+ output,
+ query_pos=query_pos,
+ value=value,
+ key_padding_mask=key_padding_mask,
+ spatial_shapes=spatial_shapes,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios,
+ reference_points=reference_points_input,
+ **kwargs)
+
+ if reg_branches is not None:
+ tmp_reg_preds = reg_branches[layer_id](output)
+ if reference_points.shape[-1] == 4:
+ new_reference_points = tmp_reg_preds + inverse_sigmoid(
+ reference_points)
+ new_reference_points = new_reference_points.sigmoid()
+ else:
+ assert reference_points.shape[-1] == 2
+ new_reference_points = tmp_reg_preds
+ new_reference_points[..., :2] = tmp_reg_preds[
+ ..., :2] + inverse_sigmoid(reference_points)
+ new_reference_points = new_reference_points.sigmoid()
+ reference_points = new_reference_points.detach()
+
+ if self.return_intermediate:
+ intermediate.append(output)
+ intermediate_reference_points.append(reference_points)
+
+ if self.return_intermediate:
+ return torch.stack(intermediate), torch.stack(
+ intermediate_reference_points)
+
+ return output, reference_points
+
+
+class DeformableDetrTransformerEncoderLayer(DetrTransformerEncoderLayer):
+ """Encoder layer of Deformable DETR."""
+
+ def _init_layers(self) -> None:
+ """Initialize self_attn, ffn, and norms."""
+ self.self_attn = MultiScaleDeformableAttention(**self.self_attn_cfg)
+ self.embed_dims = self.self_attn.embed_dims
+ self.ffn = FFN(**self.ffn_cfg)
+ norms_list = [
+ build_norm_layer(self.norm_cfg, self.embed_dims)[1]
+ for _ in range(2)
+ ]
+ self.norms = ModuleList(norms_list)
+
+
+class DeformableDetrTransformerDecoderLayer(DetrTransformerDecoderLayer):
+ """Decoder layer of Deformable DETR."""
+
+ def _init_layers(self) -> None:
+ """Initialize self_attn, cross-attn, ffn, and norms."""
+ self.self_attn = MultiheadAttention(**self.self_attn_cfg)
+ self.cross_attn = MultiScaleDeformableAttention(**self.cross_attn_cfg)
+ self.embed_dims = self.self_attn.embed_dims
+ self.ffn = FFN(**self.ffn_cfg)
+ norms_list = [
+ build_norm_layer(self.norm_cfg, self.embed_dims)[1]
+ for _ in range(3)
+ ]
+ self.norms = ModuleList(norms_list)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/detr_layers.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/detr_layers.py
new file mode 100644
index 0000000000000000000000000000000000000000..6a83dd2faa660ed8f54bdd08271db1fcf6b53886
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/detr_layers.py
@@ -0,0 +1,374 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Union
+
+import torch
+from mmcv.cnn import build_norm_layer
+from mmcv.cnn.bricks.transformer import FFN, MultiheadAttention
+from mmengine import ConfigDict
+from mmengine.model import BaseModule, ModuleList
+from torch import Tensor
+
+from mmdet.utils import ConfigType, OptConfigType
+
+try:
+ from fairscale.nn.checkpoint import checkpoint_wrapper
+except Exception:
+ checkpoint_wrapper = None
+
+
+class DetrTransformerEncoder(BaseModule):
+ """Encoder of DETR.
+
+ Args:
+ num_layers (int): Number of encoder layers.
+ layer_cfg (:obj:`ConfigDict` or dict): the config of each encoder
+ layer. All the layers will share the same config.
+ num_cp (int): Number of checkpointing blocks in encoder layer.
+ Default to -1.
+ init_cfg (:obj:`ConfigDict` or dict, optional): the config to control
+ the initialization. Defaults to None.
+ """
+
+ def __init__(self,
+ num_layers: int,
+ layer_cfg: ConfigType,
+ num_cp: int = -1,
+ init_cfg: OptConfigType = None) -> None:
+
+ super().__init__(init_cfg=init_cfg)
+ self.num_layers = num_layers
+ self.layer_cfg = layer_cfg
+ self.num_cp = num_cp
+ assert self.num_cp <= self.num_layers
+ self._init_layers()
+
+ def _init_layers(self) -> None:
+ """Initialize encoder layers."""
+ self.layers = ModuleList([
+ DetrTransformerEncoderLayer(**self.layer_cfg)
+ for _ in range(self.num_layers)
+ ])
+
+ if self.num_cp > 0:
+ if checkpoint_wrapper is None:
+ raise NotImplementedError(
+ 'If you want to reduce GPU memory usage, \
+ please install fairscale by executing the \
+ following command: pip install fairscale.')
+ for i in range(self.num_cp):
+ self.layers[i] = checkpoint_wrapper(self.layers[i])
+
+ self.embed_dims = self.layers[0].embed_dims
+
+ def forward(self, query: Tensor, query_pos: Tensor,
+ key_padding_mask: Tensor, **kwargs) -> Tensor:
+ """Forward function of encoder.
+
+ Args:
+ query (Tensor): Input queries of encoder, has shape
+ (bs, num_queries, dim).
+ query_pos (Tensor): The positional embeddings of the queries, has
+ shape (bs, num_queries, dim).
+ key_padding_mask (Tensor): The `key_padding_mask` of `self_attn`
+ input. ByteTensor, has shape (bs, num_queries).
+
+ Returns:
+ Tensor: Has shape (bs, num_queries, dim) if `batch_first` is
+ `True`, otherwise (num_queries, bs, dim).
+ """
+ for layer in self.layers:
+ query = layer(query, query_pos, key_padding_mask, **kwargs)
+ return query
+
+
+class DetrTransformerDecoder(BaseModule):
+ """Decoder of DETR.
+
+ Args:
+ num_layers (int): Number of decoder layers.
+ layer_cfg (:obj:`ConfigDict` or dict): the config of each encoder
+ layer. All the layers will share the same config.
+ post_norm_cfg (:obj:`ConfigDict` or dict, optional): Config of the
+ post normalization layer. Defaults to `LN`.
+ return_intermediate (bool, optional): Whether to return outputs of
+ intermediate layers. Defaults to `True`,
+ init_cfg (:obj:`ConfigDict` or dict, optional): the config to control
+ the initialization. Defaults to None.
+ """
+
+ def __init__(self,
+ num_layers: int,
+ layer_cfg: ConfigType,
+ post_norm_cfg: OptConfigType = dict(type='LN'),
+ return_intermediate: bool = True,
+ init_cfg: Union[dict, ConfigDict] = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.layer_cfg = layer_cfg
+ self.num_layers = num_layers
+ self.post_norm_cfg = post_norm_cfg
+ self.return_intermediate = return_intermediate
+ self._init_layers()
+
+ def _init_layers(self) -> None:
+ """Initialize decoder layers."""
+ self.layers = ModuleList([
+ DetrTransformerDecoderLayer(**self.layer_cfg)
+ for _ in range(self.num_layers)
+ ])
+ self.embed_dims = self.layers[0].embed_dims
+ self.post_norm = build_norm_layer(self.post_norm_cfg,
+ self.embed_dims)[1]
+
+ def forward(self, query: Tensor, key: Tensor, value: Tensor,
+ query_pos: Tensor, key_pos: Tensor, key_padding_mask: Tensor,
+ **kwargs) -> Tensor:
+ """Forward function of decoder
+ Args:
+ query (Tensor): The input query, has shape (bs, num_queries, dim).
+ key (Tensor): The input key, has shape (bs, num_keys, dim).
+ value (Tensor): The input value with the same shape as `key`.
+ query_pos (Tensor): The positional encoding for `query`, with the
+ same shape as `query`.
+ key_pos (Tensor): The positional encoding for `key`, with the
+ same shape as `key`.
+ key_padding_mask (Tensor): The `key_padding_mask` of `cross_attn`
+ input. ByteTensor, has shape (bs, num_value).
+
+ Returns:
+ Tensor: The forwarded results will have shape
+ (num_decoder_layers, bs, num_queries, dim) if
+ `return_intermediate` is `True` else (1, bs, num_queries, dim).
+ """
+ intermediate = []
+ for layer in self.layers:
+ query = layer(
+ query,
+ key=key,
+ value=value,
+ query_pos=query_pos,
+ key_pos=key_pos,
+ key_padding_mask=key_padding_mask,
+ **kwargs)
+ if self.return_intermediate:
+ intermediate.append(self.post_norm(query))
+ query = self.post_norm(query)
+
+ if self.return_intermediate:
+ return torch.stack(intermediate)
+
+ return query.unsqueeze(0)
+
+
+class DetrTransformerEncoderLayer(BaseModule):
+ """Implements encoder layer in DETR transformer.
+
+ Args:
+ self_attn_cfg (:obj:`ConfigDict` or dict, optional): Config for self
+ attention.
+ ffn_cfg (:obj:`ConfigDict` or dict, optional): Config for FFN.
+ norm_cfg (:obj:`ConfigDict` or dict, optional): Config for
+ normalization layers. All the layers will share the same
+ config. Defaults to `LN`.
+ init_cfg (:obj:`ConfigDict` or dict, optional): Config to control
+ the initialization. Defaults to None.
+ """
+
+ def __init__(self,
+ self_attn_cfg: OptConfigType = dict(
+ embed_dims=256, num_heads=8, dropout=0.0),
+ ffn_cfg: OptConfigType = dict(
+ embed_dims=256,
+ feedforward_channels=1024,
+ num_fcs=2,
+ ffn_drop=0.,
+ act_cfg=dict(type='ReLU', inplace=True)),
+ norm_cfg: OptConfigType = dict(type='LN'),
+ init_cfg: OptConfigType = None) -> None:
+
+ super().__init__(init_cfg=init_cfg)
+
+ self.self_attn_cfg = self_attn_cfg
+ if 'batch_first' not in self.self_attn_cfg:
+ self.self_attn_cfg['batch_first'] = True
+ else:
+ assert self.self_attn_cfg['batch_first'] is True, 'First \
+ dimension of all DETRs in mmdet is `batch`, \
+ please set `batch_first` flag.'
+
+ self.ffn_cfg = ffn_cfg
+ self.norm_cfg = norm_cfg
+ self._init_layers()
+
+ def _init_layers(self) -> None:
+ """Initialize self-attention, FFN, and normalization."""
+ self.self_attn = MultiheadAttention(**self.self_attn_cfg)
+ self.embed_dims = self.self_attn.embed_dims
+ self.ffn = FFN(**self.ffn_cfg)
+ norms_list = [
+ build_norm_layer(self.norm_cfg, self.embed_dims)[1]
+ for _ in range(2)
+ ]
+ self.norms = ModuleList(norms_list)
+
+ def forward(self, query: Tensor, query_pos: Tensor,
+ key_padding_mask: Tensor, **kwargs) -> Tensor:
+ """Forward function of an encoder layer.
+
+ Args:
+ query (Tensor): The input query, has shape (bs, num_queries, dim).
+ query_pos (Tensor): The positional encoding for query, with
+ the same shape as `query`.
+ key_padding_mask (Tensor): The `key_padding_mask` of `self_attn`
+ input. ByteTensor. has shape (bs, num_queries).
+ Returns:
+ Tensor: forwarded results, has shape (bs, num_queries, dim).
+ """
+ query = self.self_attn(
+ query=query,
+ key=query,
+ value=query,
+ query_pos=query_pos,
+ key_pos=query_pos,
+ key_padding_mask=key_padding_mask,
+ **kwargs)
+ query = self.norms[0](query)
+ query = self.ffn(query)
+ query = self.norms[1](query)
+
+ return query
+
+
+class DetrTransformerDecoderLayer(BaseModule):
+ """Implements decoder layer in DETR transformer.
+
+ Args:
+ self_attn_cfg (:obj:`ConfigDict` or dict, optional): Config for self
+ attention.
+ cross_attn_cfg (:obj:`ConfigDict` or dict, optional): Config for cross
+ attention.
+ ffn_cfg (:obj:`ConfigDict` or dict, optional): Config for FFN.
+ norm_cfg (:obj:`ConfigDict` or dict, optional): Config for
+ normalization layers. All the layers will share the same
+ config. Defaults to `LN`.
+ init_cfg (:obj:`ConfigDict` or dict, optional): Config to control
+ the initialization. Defaults to None.
+ """
+
+ def __init__(self,
+ self_attn_cfg: OptConfigType = dict(
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.0,
+ batch_first=True),
+ cross_attn_cfg: OptConfigType = dict(
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.0,
+ batch_first=True),
+ ffn_cfg: OptConfigType = dict(
+ embed_dims=256,
+ feedforward_channels=1024,
+ num_fcs=2,
+ ffn_drop=0.,
+ act_cfg=dict(type='ReLU', inplace=True),
+ ),
+ norm_cfg: OptConfigType = dict(type='LN'),
+ init_cfg: OptConfigType = None) -> None:
+
+ super().__init__(init_cfg=init_cfg)
+
+ self.self_attn_cfg = self_attn_cfg
+ self.cross_attn_cfg = cross_attn_cfg
+ if 'batch_first' not in self.self_attn_cfg:
+ self.self_attn_cfg['batch_first'] = True
+ else:
+ assert self.self_attn_cfg['batch_first'] is True, 'First \
+ dimension of all DETRs in mmdet is `batch`, \
+ please set `batch_first` flag.'
+
+ if 'batch_first' not in self.cross_attn_cfg:
+ self.cross_attn_cfg['batch_first'] = True
+ else:
+ assert self.cross_attn_cfg['batch_first'] is True, 'First \
+ dimension of all DETRs in mmdet is `batch`, \
+ please set `batch_first` flag.'
+
+ self.ffn_cfg = ffn_cfg
+ self.norm_cfg = norm_cfg
+ self._init_layers()
+
+ def _init_layers(self) -> None:
+ """Initialize self-attention, FFN, and normalization."""
+ self.self_attn = MultiheadAttention(**self.self_attn_cfg)
+ self.cross_attn = MultiheadAttention(**self.cross_attn_cfg)
+ self.embed_dims = self.self_attn.embed_dims
+ self.ffn = FFN(**self.ffn_cfg)
+ norms_list = [
+ build_norm_layer(self.norm_cfg, self.embed_dims)[1]
+ for _ in range(3)
+ ]
+ self.norms = ModuleList(norms_list)
+
+ def forward(self,
+ query: Tensor,
+ key: Tensor = None,
+ value: Tensor = None,
+ query_pos: Tensor = None,
+ key_pos: Tensor = None,
+ self_attn_mask: Tensor = None,
+ cross_attn_mask: Tensor = None,
+ key_padding_mask: Tensor = None,
+ **kwargs) -> Tensor:
+ """
+ Args:
+ query (Tensor): The input query, has shape (bs, num_queries, dim).
+ key (Tensor, optional): The input key, has shape (bs, num_keys,
+ dim). If `None`, the `query` will be used. Defaults to `None`.
+ value (Tensor, optional): The input value, has the same shape as
+ `key`, as in `nn.MultiheadAttention.forward`. If `None`, the
+ `key` will be used. Defaults to `None`.
+ query_pos (Tensor, optional): The positional encoding for `query`,
+ has the same shape as `query`. If not `None`, it will be added
+ to `query` before forward function. Defaults to `None`.
+ key_pos (Tensor, optional): The positional encoding for `key`, has
+ the same shape as `key`. If not `None`, it will be added to
+ `key` before forward function. If None, and `query_pos` has the
+ same shape as `key`, then `query_pos` will be used for
+ `key_pos`. Defaults to None.
+ self_attn_mask (Tensor, optional): ByteTensor mask, has shape
+ (num_queries, num_keys), as in `nn.MultiheadAttention.forward`.
+ Defaults to None.
+ cross_attn_mask (Tensor, optional): ByteTensor mask, has shape
+ (num_queries, num_keys), as in `nn.MultiheadAttention.forward`.
+ Defaults to None.
+ key_padding_mask (Tensor, optional): The `key_padding_mask` of
+ `self_attn` input. ByteTensor, has shape (bs, num_value).
+ Defaults to None.
+
+ Returns:
+ Tensor: forwarded results, has shape (bs, num_queries, dim).
+ """
+
+ query = self.self_attn(
+ query=query,
+ key=query,
+ value=query,
+ query_pos=query_pos,
+ key_pos=query_pos,
+ attn_mask=self_attn_mask,
+ **kwargs)
+ query = self.norms[0](query)
+ query = self.cross_attn(
+ query=query,
+ key=key,
+ value=value,
+ query_pos=query_pos,
+ key_pos=key_pos,
+ attn_mask=cross_attn_mask,
+ key_padding_mask=key_padding_mask,
+ **kwargs)
+ query = self.norms[1](query)
+ query = self.ffn(query)
+ query = self.norms[2](query)
+
+ return query
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/dino_layers.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/dino_layers.py
new file mode 100644
index 0000000000000000000000000000000000000000..64610d0a7c0121a88f5e4279b6f854924230237e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/dino_layers.py
@@ -0,0 +1,562 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+from typing import Tuple, Union
+
+import torch
+from mmengine.model import BaseModule
+from torch import Tensor, nn
+
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import bbox_xyxy_to_cxcywh
+from mmdet.utils import OptConfigType
+from .deformable_detr_layers import DeformableDetrTransformerDecoder
+from .utils import MLP, coordinate_to_encoding, inverse_sigmoid
+
+
+class DinoTransformerDecoder(DeformableDetrTransformerDecoder):
+ """Transformer decoder of DINO."""
+
+ def _init_layers(self) -> None:
+ """Initialize decoder layers."""
+ super()._init_layers()
+ self.ref_point_head = MLP(self.embed_dims * 2, self.embed_dims,
+ self.embed_dims, 2)
+ self.norm = nn.LayerNorm(self.embed_dims)
+
+ def forward(self, query: Tensor, value: Tensor, key_padding_mask: Tensor,
+ self_attn_mask: Tensor, reference_points: Tensor,
+ spatial_shapes: Tensor, level_start_index: Tensor,
+ valid_ratios: Tensor, reg_branches: nn.ModuleList,
+ **kwargs) -> Tuple[Tensor]:
+ """Forward function of Transformer decoder.
+
+ Args:
+ query (Tensor): The input query, has shape (num_queries, bs, dim).
+ value (Tensor): The input values, has shape (num_value, bs, dim).
+ key_padding_mask (Tensor): The `key_padding_mask` of `self_attn`
+ input. ByteTensor, has shape (num_queries, bs).
+ self_attn_mask (Tensor): The attention mask to prevent information
+ leakage from different denoising groups and matching parts, has
+ shape (num_queries_total, num_queries_total). It is `None` when
+ `self.training` is `False`.
+ reference_points (Tensor): The initial reference, has shape
+ (bs, num_queries, 4) with the last dimension arranged as
+ (cx, cy, w, h).
+ spatial_shapes (Tensor): Spatial shapes of features in all levels,
+ has shape (num_levels, 2), last dimension represents (h, w).
+ level_start_index (Tensor): The start index of each level.
+ A tensor has shape (num_levels, ) and can be represented
+ as [0, h_0*w_0, h_0*w_0+h_1*w_1, ...].
+ valid_ratios (Tensor): The ratios of the valid width and the valid
+ height relative to the width and the height of features in all
+ levels, has shape (bs, num_levels, 2).
+ reg_branches: (obj:`nn.ModuleList`): Used for refining the
+ regression results.
+
+ Returns:
+ tuple[Tensor]: Output queries and references of Transformer
+ decoder
+
+ - query (Tensor): Output embeddings of the last decoder, has
+ shape (num_queries, bs, embed_dims) when `return_intermediate`
+ is `False`. Otherwise, Intermediate output embeddings of all
+ decoder layers, has shape (num_decoder_layers, num_queries, bs,
+ embed_dims).
+ - reference_points (Tensor): The reference of the last decoder
+ layer, has shape (bs, num_queries, 4) when `return_intermediate`
+ is `False`. Otherwise, Intermediate references of all decoder
+ layers, has shape (num_decoder_layers, bs, num_queries, 4). The
+ coordinates are arranged as (cx, cy, w, h)
+ """
+ intermediate = []
+ intermediate_reference_points = [reference_points]
+ for lid, layer in enumerate(self.layers):
+ if reference_points.shape[-1] == 4:
+ reference_points_input = \
+ reference_points[:, :, None] * torch.cat(
+ [valid_ratios, valid_ratios], -1)[:, None]
+ else:
+ assert reference_points.shape[-1] == 2
+ reference_points_input = \
+ reference_points[:, :, None] * valid_ratios[:, None]
+
+ query_sine_embed = coordinate_to_encoding(
+ reference_points_input[:, :, 0, :])
+ query_pos = self.ref_point_head(query_sine_embed)
+
+ query = layer(
+ query,
+ query_pos=query_pos,
+ value=value,
+ key_padding_mask=key_padding_mask,
+ self_attn_mask=self_attn_mask,
+ spatial_shapes=spatial_shapes,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios,
+ reference_points=reference_points_input,
+ **kwargs)
+
+ if reg_branches is not None:
+ tmp = reg_branches[lid](query)
+ assert reference_points.shape[-1] == 4
+ new_reference_points = tmp + inverse_sigmoid(
+ reference_points, eps=1e-3)
+ new_reference_points = new_reference_points.sigmoid()
+ reference_points = new_reference_points.detach()
+
+ if self.return_intermediate:
+ intermediate.append(self.norm(query))
+ intermediate_reference_points.append(new_reference_points)
+ # NOTE this is for the "Look Forward Twice" module,
+ # in the DeformDETR, reference_points was appended.
+
+ if self.return_intermediate:
+ return torch.stack(intermediate), torch.stack(
+ intermediate_reference_points)
+
+ return query, reference_points
+
+
+class CdnQueryGenerator(BaseModule):
+ """Implement query generator of the Contrastive denoising (CDN) proposed in
+ `DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object
+ Detection `_
+
+ Code is modified from the `official github repo
+ `_.
+
+ Args:
+ num_classes (int): Number of object classes.
+ embed_dims (int): The embedding dimensions of the generated queries.
+ num_matching_queries (int): The queries number of the matching part.
+ Used for generating dn_mask.
+ label_noise_scale (float): The scale of label noise, defaults to 0.5.
+ box_noise_scale (float): The scale of box noise, defaults to 1.0.
+ group_cfg (:obj:`ConfigDict` or dict, optional): The config of the
+ denoising queries grouping, includes `dynamic`, `num_dn_queries`,
+ and `num_groups`. Two grouping strategies, 'static dn groups' and
+ 'dynamic dn groups', are supported. When `dynamic` is `False`,
+ the `num_groups` should be set, and the number of denoising query
+ groups will always be `num_groups`. When `dynamic` is `True`, the
+ `num_dn_queries` should be set, and the group number will be
+ dynamic to ensure that the denoising queries number will not exceed
+ `num_dn_queries` to prevent large fluctuations of memory. Defaults
+ to `None`.
+ """
+
+ def __init__(self,
+ num_classes: int,
+ embed_dims: int,
+ num_matching_queries: int,
+ label_noise_scale: float = 0.5,
+ box_noise_scale: float = 1.0,
+ group_cfg: OptConfigType = None) -> None:
+ super().__init__()
+ self.num_classes = num_classes
+ self.embed_dims = embed_dims
+ self.num_matching_queries = num_matching_queries
+ self.label_noise_scale = label_noise_scale
+ self.box_noise_scale = box_noise_scale
+
+ # prepare grouping strategy
+ group_cfg = {} if group_cfg is None else group_cfg
+ self.dynamic_dn_groups = group_cfg.get('dynamic', True)
+ if self.dynamic_dn_groups:
+ if 'num_dn_queries' not in group_cfg:
+ warnings.warn("'num_dn_queries' should be set when using "
+ 'dynamic dn groups, use 100 as default.')
+ self.num_dn_queries = group_cfg.get('num_dn_queries', 100)
+ assert isinstance(self.num_dn_queries, int), \
+ f'Expected the num_dn_queries to have type int, but got ' \
+ f'{self.num_dn_queries}({type(self.num_dn_queries)}). '
+ else:
+ assert 'num_groups' in group_cfg, \
+ 'num_groups should be set when using static dn groups'
+ self.num_groups = group_cfg['num_groups']
+ assert isinstance(self.num_groups, int), \
+ f'Expected the num_groups to have type int, but got ' \
+ f'{self.num_groups}({type(self.num_groups)}). '
+
+ # NOTE The original repo of DINO set the num_embeddings 92 for coco,
+ # 91 (0~90) of which represents target classes and the 92 (91)
+ # indicates `Unknown` class. However, the embedding of `unknown` class
+ # is not used in the original DINO.
+ # TODO: num_classes + 1 or num_classes ?
+ self.label_embedding = nn.Embedding(self.num_classes, self.embed_dims)
+
+ def __call__(self, batch_data_samples: SampleList) -> tuple:
+ """Generate contrastive denoising (cdn) queries with ground truth.
+
+ Descriptions of the Number Values in code and comments:
+ - num_target_total: the total target number of the input batch
+ samples.
+ - max_num_target: the max target number of the input batch samples.
+ - num_noisy_targets: the total targets number after adding noise,
+ i.e., num_target_total * num_groups * 2.
+ - num_denoising_queries: the length of the output batched queries,
+ i.e., max_num_target * num_groups * 2.
+
+ NOTE The format of input bboxes in batch_data_samples is unnormalized
+ (x, y, x, y), and the output bbox queries are embedded by normalized
+ (cx, cy, w, h) format bboxes going through inverse_sigmoid.
+
+ Args:
+ batch_data_samples (list[:obj:`DetDataSample`]): List of the batch
+ data samples, each includes `gt_instance` which has attributes
+ `bboxes` and `labels`. The `bboxes` has unnormalized coordinate
+ format (x, y, x, y).
+
+ Returns:
+ tuple: The outputs of the dn query generator.
+
+ - dn_label_query (Tensor): The output content queries for denoising
+ part, has shape (bs, num_denoising_queries, dim), where
+ `num_denoising_queries = max_num_target * num_groups * 2`.
+ - dn_bbox_query (Tensor): The output reference bboxes as positions
+ of queries for denoising part, which are embedded by normalized
+ (cx, cy, w, h) format bboxes going through inverse_sigmoid, has
+ shape (bs, num_denoising_queries, 4) with the last dimension
+ arranged as (cx, cy, w, h).
+ - attn_mask (Tensor): The attention mask to prevent information
+ leakage from different denoising groups and matching parts,
+ will be used as `self_attn_mask` of the `decoder`, has shape
+ (num_queries_total, num_queries_total), where `num_queries_total`
+ is the sum of `num_denoising_queries` and `num_matching_queries`.
+ - dn_meta (Dict[str, int]): The dictionary saves information about
+ group collation, including 'num_denoising_queries' and
+ 'num_denoising_groups'. It will be used for split outputs of
+ denoising and matching parts and loss calculation.
+ """
+ # normalize bbox and collate ground truth (gt)
+ gt_labels_list = []
+ gt_bboxes_list = []
+ for sample in batch_data_samples:
+ img_h, img_w = sample.img_shape
+ bboxes = sample.gt_instances.bboxes
+ factor = bboxes.new_tensor([img_w, img_h, img_w,
+ img_h]).unsqueeze(0)
+ bboxes_normalized = bboxes / factor
+ gt_bboxes_list.append(bboxes_normalized)
+ gt_labels_list.append(sample.gt_instances.labels)
+ gt_labels = torch.cat(gt_labels_list) # (num_target_total, 4)
+ gt_bboxes = torch.cat(gt_bboxes_list)
+
+ num_target_list = [len(bboxes) for bboxes in gt_bboxes_list]
+ max_num_target = max(num_target_list)
+ num_groups = self.get_num_groups(max_num_target)
+
+ dn_label_query = self.generate_dn_label_query(gt_labels, num_groups)
+ dn_bbox_query = self.generate_dn_bbox_query(gt_bboxes, num_groups)
+
+ # The `batch_idx` saves the batch index of the corresponding sample
+ # for each target, has shape (num_target_total).
+ batch_idx = torch.cat([
+ torch.full_like(t.long(), i) for i, t in enumerate(gt_labels_list)
+ ])
+ dn_label_query, dn_bbox_query = self.collate_dn_queries(
+ dn_label_query, dn_bbox_query, batch_idx, len(batch_data_samples),
+ num_groups)
+
+ attn_mask = self.generate_dn_mask(
+ max_num_target, num_groups, device=dn_label_query.device)
+
+ dn_meta = dict(
+ num_denoising_queries=int(max_num_target * 2 * num_groups),
+ num_denoising_groups=num_groups)
+
+ return dn_label_query, dn_bbox_query, attn_mask, dn_meta
+
+ def get_num_groups(self, max_num_target: int = None) -> int:
+ """Calculate denoising query groups number.
+
+ Two grouping strategies, 'static dn groups' and 'dynamic dn groups',
+ are supported. When `self.dynamic_dn_groups` is `False`, the number
+ of denoising query groups will always be `self.num_groups`. When
+ `self.dynamic_dn_groups` is `True`, the group number will be dynamic,
+ ensuring the denoising queries number will not exceed
+ `self.num_dn_queries` to prevent large fluctuations of memory.
+
+ NOTE The `num_group` is shared for different samples in a batch. When
+ the target numbers in the samples varies, the denoising queries of the
+ samples containing fewer targets are padded to the max length.
+
+ Args:
+ max_num_target (int, optional): The max target number of the batch
+ samples. It will only be used when `self.dynamic_dn_groups` is
+ `True`. Defaults to `None`.
+
+ Returns:
+ int: The denoising group number of the current batch.
+ """
+ if self.dynamic_dn_groups:
+ assert max_num_target is not None, \
+ 'group_queries should be provided when using ' \
+ 'dynamic dn groups'
+ if max_num_target == 0:
+ num_groups = 1
+ else:
+ num_groups = self.num_dn_queries // max_num_target
+ else:
+ num_groups = self.num_groups
+ if num_groups < 1:
+ num_groups = 1
+ return int(num_groups)
+
+ def generate_dn_label_query(self, gt_labels: Tensor,
+ num_groups: int) -> Tensor:
+ """Generate noisy labels and their query embeddings.
+
+ The strategy for generating noisy labels is: Randomly choose labels of
+ `self.label_noise_scale * 0.5` proportion and override each of them
+ with a random object category label.
+
+ NOTE Not add noise to all labels. Besides, the `self.label_noise_scale
+ * 0.5` arg is the ratio of the chosen positions, which is higher than
+ the actual proportion of noisy labels, because the labels to override
+ may be correct. And the gap becomes larger as the number of target
+ categories decreases. The users should notice this and modify the scale
+ arg or the corresponding logic according to specific dataset.
+
+ Args:
+ gt_labels (Tensor): The concatenated gt labels of all samples
+ in the batch, has shape (num_target_total, ) where
+ `num_target_total = sum(num_target_list)`.
+ num_groups (int): The number of denoising query groups.
+
+ Returns:
+ Tensor: The query embeddings of noisy labels, has shape
+ (num_noisy_targets, embed_dims), where `num_noisy_targets =
+ num_target_total * num_groups * 2`.
+ """
+ assert self.label_noise_scale > 0
+ gt_labels_expand = gt_labels.repeat(2 * num_groups,
+ 1).view(-1) # Note `* 2` # noqa
+ p = torch.rand_like(gt_labels_expand.float())
+ chosen_indice = torch.nonzero(p < (self.label_noise_scale * 0.5)).view(
+ -1) # Note `* 0.5`
+ new_labels = torch.randint_like(chosen_indice, 0, self.num_classes)
+ noisy_labels_expand = gt_labels_expand.scatter(0, chosen_indice,
+ new_labels)
+ dn_label_query = self.label_embedding(noisy_labels_expand)
+ return dn_label_query
+
+ def generate_dn_bbox_query(self, gt_bboxes: Tensor,
+ num_groups: int) -> Tensor:
+ """Generate noisy bboxes and their query embeddings.
+
+ The strategy for generating noisy bboxes is as follow:
+
+ .. code:: text
+
+ +--------------------+
+ | negative |
+ | +----------+ |
+ | | positive | |
+ | | +-----|----+------------+
+ | | | | | |
+ | +----+-----+ | |
+ | | | |
+ +---------+----------+ |
+ | |
+ | gt bbox |
+ | |
+ | +---------+----------+
+ | | | |
+ | | +----+-----+ |
+ | | | | | |
+ +-------------|--- +----+ | |
+ | | positive | |
+ | +----------+ |
+ | negative |
+ +--------------------+
+
+ The random noise is added to the top-left and down-right point
+ positions, hence, normalized (x, y, x, y) format of bboxes are
+ required. The noisy bboxes of positive queries have the points
+ both within the inner square, while those of negative queries
+ have the points both between the inner and outer squares.
+
+ Besides, the length of outer square is twice as long as that of
+ the inner square, i.e., self.box_noise_scale * w_or_h / 2.
+ NOTE The noise is added to all the bboxes. Moreover, there is still
+ unconsidered case when one point is within the positive square and
+ the others is between the inner and outer squares.
+
+ Args:
+ gt_bboxes (Tensor): The concatenated gt bboxes of all samples
+ in the batch, has shape (num_target_total, 4) with the last
+ dimension arranged as (cx, cy, w, h) where
+ `num_target_total = sum(num_target_list)`.
+ num_groups (int): The number of denoising query groups.
+
+ Returns:
+ Tensor: The output noisy bboxes, which are embedded by normalized
+ (cx, cy, w, h) format bboxes going through inverse_sigmoid, has
+ shape (num_noisy_targets, 4) with the last dimension arranged as
+ (cx, cy, w, h), where
+ `num_noisy_targets = num_target_total * num_groups * 2`.
+ """
+ assert self.box_noise_scale > 0
+ device = gt_bboxes.device
+
+ # expand gt_bboxes as groups
+ gt_bboxes_expand = gt_bboxes.repeat(2 * num_groups, 1) # xyxy
+
+ # obtain index of negative queries in gt_bboxes_expand
+ positive_idx = torch.arange(
+ len(gt_bboxes), dtype=torch.long, device=device)
+ positive_idx = positive_idx.unsqueeze(0).repeat(num_groups, 1)
+ positive_idx += 2 * len(gt_bboxes) * torch.arange(
+ num_groups, dtype=torch.long, device=device)[:, None]
+ positive_idx = positive_idx.flatten()
+ negative_idx = positive_idx + len(gt_bboxes)
+
+ # determine the sign of each element in the random part of the added
+ # noise to be positive or negative randomly.
+ rand_sign = torch.randint_like(
+ gt_bboxes_expand, low=0, high=2,
+ dtype=torch.float32) * 2.0 - 1.0 # [low, high), 1 or -1, randomly
+
+ # calculate the random part of the added noise
+ rand_part = torch.rand_like(gt_bboxes_expand) # [0, 1)
+ rand_part[negative_idx] += 1.0 # pos: [0, 1); neg: [1, 2)
+ rand_part *= rand_sign # pos: (-1, 1); neg: (-2, -1] U [1, 2)
+
+ # add noise to the bboxes
+ bboxes_whwh = bbox_xyxy_to_cxcywh(gt_bboxes_expand)[:, 2:].repeat(1, 2)
+ noisy_bboxes_expand = gt_bboxes_expand + torch.mul(
+ rand_part, bboxes_whwh) * self.box_noise_scale / 2 # xyxy
+ noisy_bboxes_expand = noisy_bboxes_expand.clamp(min=0.0, max=1.0)
+ noisy_bboxes_expand = bbox_xyxy_to_cxcywh(noisy_bboxes_expand)
+
+ dn_bbox_query = inverse_sigmoid(noisy_bboxes_expand, eps=1e-3)
+ return dn_bbox_query
+
+ def collate_dn_queries(self, input_label_query: Tensor,
+ input_bbox_query: Tensor, batch_idx: Tensor,
+ batch_size: int, num_groups: int) -> Tuple[Tensor]:
+ """Collate generated queries to obtain batched dn queries.
+
+ The strategy for query collation is as follow:
+
+ .. code:: text
+
+ input_queries (num_target_total, query_dim)
+ P_A1 P_B1 P_B2 N_A1 N_B1 N_B2 P'A1 P'B1 P'B2 N'A1 N'B1 N'B2
+ |________ group1 ________| |________ group2 ________|
+ |
+ V
+ P_A1 Pad0 N_A1 Pad0 P'A1 Pad0 N'A1 Pad0
+ P_B1 P_B2 N_B1 N_B2 P'B1 P'B2 N'B1 N'B2
+ |____ group1 ____| |____ group2 ____|
+ batched_queries (batch_size, max_num_target, query_dim)
+
+ where query_dim is 4 for bbox and self.embed_dims for label.
+ Notation: _-group 1; '-group 2;
+ A-Sample1(has 1 target); B-sample2(has 2 targets)
+
+ Args:
+ input_label_query (Tensor): The generated label queries of all
+ targets, has shape (num_target_total, embed_dims) where
+ `num_target_total = sum(num_target_list)`.
+ input_bbox_query (Tensor): The generated bbox queries of all
+ targets, has shape (num_target_total, 4) with the last
+ dimension arranged as (cx, cy, w, h).
+ batch_idx (Tensor): The batch index of the corresponding sample
+ for each target, has shape (num_target_total).
+ batch_size (int): The size of the input batch.
+ num_groups (int): The number of denoising query groups.
+
+ Returns:
+ tuple[Tensor]: Output batched label and bbox queries.
+ - batched_label_query (Tensor): The output batched label queries,
+ has shape (batch_size, max_num_target, embed_dims).
+ - batched_bbox_query (Tensor): The output batched bbox queries,
+ has shape (batch_size, max_num_target, 4) with the last dimension
+ arranged as (cx, cy, w, h).
+ """
+ device = input_label_query.device
+ num_target_list = [
+ torch.sum(batch_idx == idx) for idx in range(batch_size)
+ ]
+ max_num_target = max(num_target_list)
+ num_denoising_queries = int(max_num_target * 2 * num_groups)
+
+ map_query_index = torch.cat([
+ torch.arange(num_target, device=device)
+ for num_target in num_target_list
+ ])
+ map_query_index = torch.cat([
+ map_query_index + max_num_target * i for i in range(2 * num_groups)
+ ]).long()
+ batch_idx_expand = batch_idx.repeat(2 * num_groups, 1).view(-1)
+ mapper = (batch_idx_expand, map_query_index)
+
+ batched_label_query = torch.zeros(
+ batch_size, num_denoising_queries, self.embed_dims, device=device)
+ batched_bbox_query = torch.zeros(
+ batch_size, num_denoising_queries, 4, device=device)
+
+ batched_label_query[mapper] = input_label_query
+ batched_bbox_query[mapper] = input_bbox_query
+ return batched_label_query, batched_bbox_query
+
+ def generate_dn_mask(self, max_num_target: int, num_groups: int,
+ device: Union[torch.device, str]) -> Tensor:
+ """Generate attention mask to prevent information leakage from
+ different denoising groups and matching parts.
+
+ .. code:: text
+
+ 0 0 0 0 1 1 1 1 0 0 0 0 0
+ 0 0 0 0 1 1 1 1 0 0 0 0 0
+ 0 0 0 0 1 1 1 1 0 0 0 0 0
+ 0 0 0 0 1 1 1 1 0 0 0 0 0
+ 1 1 1 1 0 0 0 0 0 0 0 0 0
+ 1 1 1 1 0 0 0 0 0 0 0 0 0
+ 1 1 1 1 0 0 0 0 0 0 0 0 0
+ 1 1 1 1 0 0 0 0 0 0 0 0 0
+ 1 1 1 1 1 1 1 1 0 0 0 0 0
+ 1 1 1 1 1 1 1 1 0 0 0 0 0
+ 1 1 1 1 1 1 1 1 0 0 0 0 0
+ 1 1 1 1 1 1 1 1 0 0 0 0 0
+ 1 1 1 1 1 1 1 1 0 0 0 0 0
+ max_num_target |_| |_________| num_matching_queries
+ |_____________| num_denoising_queries
+
+ 1 -> True (Masked), means 'can not see'.
+ 0 -> False (UnMasked), means 'can see'.
+
+ Args:
+ max_num_target (int): The max target number of the input batch
+ samples.
+ num_groups (int): The number of denoising query groups.
+ device (obj:`device` or str): The device of generated mask.
+
+ Returns:
+ Tensor: The attention mask to prevent information leakage from
+ different denoising groups and matching parts, will be used as
+ `self_attn_mask` of the `decoder`, has shape (num_queries_total,
+ num_queries_total), where `num_queries_total` is the sum of
+ `num_denoising_queries` and `num_matching_queries`.
+ """
+ num_denoising_queries = int(max_num_target * 2 * num_groups)
+ num_queries_total = num_denoising_queries + self.num_matching_queries
+ attn_mask = torch.zeros(
+ num_queries_total,
+ num_queries_total,
+ device=device,
+ dtype=torch.bool)
+ # Make the matching part cannot see the denoising groups
+ attn_mask[num_denoising_queries:, :num_denoising_queries] = True
+ # Make the denoising groups cannot see each other
+ for i in range(num_groups):
+ # Mask rows of one group per step.
+ row_scope = slice(max_num_target * 2 * i,
+ max_num_target * 2 * (i + 1))
+ left_scope = slice(max_num_target * 2 * i)
+ right_scope = slice(max_num_target * 2 * (i + 1),
+ num_denoising_queries)
+ attn_mask[row_scope, right_scope] = True
+ attn_mask[row_scope, left_scope] = True
+ return attn_mask
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/grounding_dino_layers.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/grounding_dino_layers.py
new file mode 100644
index 0000000000000000000000000000000000000000..3c285768f36af98075607b43e48e6f1018125ad1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/grounding_dino_layers.py
@@ -0,0 +1,270 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+import torch.nn as nn
+from mmcv.cnn import build_norm_layer
+from mmcv.cnn.bricks.transformer import FFN, MultiheadAttention
+from mmcv.ops import MultiScaleDeformableAttention
+from mmengine.model import ModuleList
+from torch import Tensor
+
+from mmdet.models.utils.vlfuse_helper import SingleScaleBiAttentionBlock
+from mmdet.utils import ConfigType, OptConfigType
+from .deformable_detr_layers import (DeformableDetrTransformerDecoderLayer,
+ DeformableDetrTransformerEncoder,
+ DeformableDetrTransformerEncoderLayer)
+from .detr_layers import DetrTransformerEncoderLayer
+from .dino_layers import DinoTransformerDecoder
+from .utils import MLP, get_text_sine_pos_embed
+
+try:
+ from fairscale.nn.checkpoint import checkpoint_wrapper
+except Exception:
+ checkpoint_wrapper = None
+
+
+class GroundingDinoTransformerDecoderLayer(
+ DeformableDetrTransformerDecoderLayer):
+
+ def __init__(self,
+ cross_attn_text_cfg: OptConfigType = dict(
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.0,
+ batch_first=True),
+ **kwargs) -> None:
+ """Decoder layer of Deformable DETR."""
+ self.cross_attn_text_cfg = cross_attn_text_cfg
+ if 'batch_first' not in self.cross_attn_text_cfg:
+ self.cross_attn_text_cfg['batch_first'] = True
+ super().__init__(**kwargs)
+
+ def _init_layers(self) -> None:
+ """Initialize self_attn, cross-attn, ffn, and norms."""
+ self.self_attn = MultiheadAttention(**self.self_attn_cfg)
+ self.cross_attn_text = MultiheadAttention(**self.cross_attn_text_cfg)
+ self.cross_attn = MultiScaleDeformableAttention(**self.cross_attn_cfg)
+ self.embed_dims = self.self_attn.embed_dims
+ self.ffn = FFN(**self.ffn_cfg)
+ norms_list = [
+ build_norm_layer(self.norm_cfg, self.embed_dims)[1]
+ for _ in range(4)
+ ]
+ self.norms = ModuleList(norms_list)
+
+ def forward(self,
+ query: Tensor,
+ key: Tensor = None,
+ value: Tensor = None,
+ query_pos: Tensor = None,
+ key_pos: Tensor = None,
+ self_attn_mask: Tensor = None,
+ cross_attn_mask: Tensor = None,
+ key_padding_mask: Tensor = None,
+ memory_text: Tensor = None,
+ text_attention_mask: Tensor = None,
+ **kwargs) -> Tensor:
+ """Implements decoder layer in Grounding DINO transformer.
+
+ Args:
+ query (Tensor): The input query, has shape (bs, num_queries, dim).
+ key (Tensor, optional): The input key, has shape (bs, num_keys,
+ dim). If `None`, the `query` will be used. Defaults to `None`.
+ value (Tensor, optional): The input value, has the same shape as
+ `key`, as in `nn.MultiheadAttention.forward`. If `None`, the
+ `key` will be used. Defaults to `None`.
+ query_pos (Tensor, optional): The positional encoding for `query`,
+ has the same shape as `query`. If not `None`, it will be added
+ to `query` before forward function. Defaults to `None`.
+ key_pos (Tensor, optional): The positional encoding for `key`, has
+ the same shape as `key`. If not `None`, it will be added to
+ `key` before forward function. If None, and `query_pos` has the
+ same shape as `key`, then `query_pos` will be used for
+ `key_pos`. Defaults to None.
+ self_attn_mask (Tensor, optional): ByteTensor mask, has shape
+ (num_queries, num_keys), as in `nn.MultiheadAttention.forward`.
+ Defaults to None.
+ cross_attn_mask (Tensor, optional): ByteTensor mask, has shape
+ (num_queries, num_keys), as in `nn.MultiheadAttention.forward`.
+ Defaults to None.
+ key_padding_mask (Tensor, optional): The `key_padding_mask` of
+ `self_attn` input. ByteTensor, has shape (bs, num_value).
+ Defaults to None.
+ memory_text (Tensor): Memory text. It has shape (bs, len_text,
+ text_embed_dims).
+ text_attention_mask (Tensor): Text token mask. It has shape (bs,
+ len_text).
+
+ Returns:
+ Tensor: forwarded results, has shape (bs, num_queries, dim).
+ """
+ # self attention
+ query = self.self_attn(
+ query=query,
+ key=query,
+ value=query,
+ query_pos=query_pos,
+ key_pos=query_pos,
+ attn_mask=self_attn_mask,
+ **kwargs)
+ query = self.norms[0](query)
+ # cross attention between query and text
+ query = self.cross_attn_text(
+ query=query,
+ query_pos=query_pos,
+ key=memory_text,
+ value=memory_text,
+ key_padding_mask=text_attention_mask)
+ query = self.norms[1](query)
+ # cross attention between query and image
+ query = self.cross_attn(
+ query=query,
+ key=key,
+ value=value,
+ query_pos=query_pos,
+ key_pos=key_pos,
+ attn_mask=cross_attn_mask,
+ key_padding_mask=key_padding_mask,
+ **kwargs)
+ query = self.norms[2](query)
+ query = self.ffn(query)
+ query = self.norms[3](query)
+
+ return query
+
+
+class GroundingDinoTransformerEncoder(DeformableDetrTransformerEncoder):
+
+ def __init__(self, text_layer_cfg: ConfigType,
+ fusion_layer_cfg: ConfigType, **kwargs) -> None:
+ self.text_layer_cfg = text_layer_cfg
+ self.fusion_layer_cfg = fusion_layer_cfg
+ super().__init__(**kwargs)
+
+ def _init_layers(self) -> None:
+ """Initialize encoder layers."""
+ self.layers = ModuleList([
+ DeformableDetrTransformerEncoderLayer(**self.layer_cfg)
+ for _ in range(self.num_layers)
+ ])
+ self.text_layers = ModuleList([
+ DetrTransformerEncoderLayer(**self.text_layer_cfg)
+ for _ in range(self.num_layers)
+ ])
+ self.fusion_layers = ModuleList([
+ SingleScaleBiAttentionBlock(**self.fusion_layer_cfg)
+ for _ in range(self.num_layers)
+ ])
+ self.embed_dims = self.layers[0].embed_dims
+ if self.num_cp > 0:
+ if checkpoint_wrapper is None:
+ raise NotImplementedError(
+ 'If you want to reduce GPU memory usage, \
+ please install fairscale by executing the \
+ following command: pip install fairscale.')
+ for i in range(self.num_cp):
+ self.layers[i] = checkpoint_wrapper(self.layers[i])
+ self.fusion_layers[i] = checkpoint_wrapper(
+ self.fusion_layers[i])
+
+ def forward(self,
+ query: Tensor,
+ query_pos: Tensor,
+ key_padding_mask: Tensor,
+ spatial_shapes: Tensor,
+ level_start_index: Tensor,
+ valid_ratios: Tensor,
+ memory_text: Tensor = None,
+ text_attention_mask: Tensor = None,
+ pos_text: Tensor = None,
+ text_self_attention_masks: Tensor = None,
+ position_ids: Tensor = None):
+ """Forward function of Transformer encoder.
+
+ Args:
+ query (Tensor): The input query, has shape (bs, num_queries, dim).
+ query_pos (Tensor): The positional encoding for query, has shape
+ (bs, num_queries, dim).
+ key_padding_mask (Tensor): The `key_padding_mask` of `self_attn`
+ input. ByteTensor, has shape (bs, num_queries).
+ spatial_shapes (Tensor): Spatial shapes of features in all levels,
+ has shape (num_levels, 2), last dimension represents (h, w).
+ level_start_index (Tensor): The start index of each level.
+ A tensor has shape (num_levels, ) and can be represented
+ as [0, h_0*w_0, h_0*w_0+h_1*w_1, ...].
+ valid_ratios (Tensor): The ratios of the valid width and the valid
+ height relative to the width and the height of features in all
+ levels, has shape (bs, num_levels, 2).
+ memory_text (Tensor, optional): Memory text. It has shape (bs,
+ len_text, text_embed_dims).
+ text_attention_mask (Tensor, optional): Text token mask. It has
+ shape (bs,len_text).
+ pos_text (Tensor, optional): The positional encoding for text.
+ Defaults to None.
+ text_self_attention_masks (Tensor, optional): Text self attention
+ mask. Defaults to None.
+ position_ids (Tensor, optional): Text position ids.
+ Defaults to None.
+ """
+ output = query
+ reference_points = self.get_encoder_reference_points(
+ spatial_shapes, valid_ratios, device=query.device)
+ if self.text_layers:
+ # generate pos_text
+ bs, n_text, _ = memory_text.shape
+ if pos_text is None and position_ids is None:
+ pos_text = (
+ torch.arange(n_text,
+ device=memory_text.device).float().unsqueeze(
+ 0).unsqueeze(-1).repeat(bs, 1, 1))
+ pos_text = get_text_sine_pos_embed(
+ pos_text, num_pos_feats=256, exchange_xy=False)
+ if position_ids is not None:
+ pos_text = get_text_sine_pos_embed(
+ position_ids[..., None],
+ num_pos_feats=256,
+ exchange_xy=False)
+
+ # main process
+ for layer_id, layer in enumerate(self.layers):
+ if self.fusion_layers:
+ output, memory_text = self.fusion_layers[layer_id](
+ visual_feature=output,
+ lang_feature=memory_text,
+ attention_mask_v=key_padding_mask,
+ attention_mask_l=text_attention_mask,
+ )
+ if self.text_layers:
+ text_num_heads = self.text_layers[
+ layer_id].self_attn_cfg.num_heads
+ memory_text = self.text_layers[layer_id](
+ query=memory_text,
+ query_pos=(pos_text if pos_text is not None else None),
+ attn_mask=~text_self_attention_masks.repeat(
+ text_num_heads, 1, 1), # note we use ~ for mask here
+ key_padding_mask=None,
+ )
+ output = layer(
+ query=output,
+ query_pos=query_pos,
+ reference_points=reference_points,
+ spatial_shapes=spatial_shapes,
+ level_start_index=level_start_index,
+ key_padding_mask=key_padding_mask)
+ return output, memory_text
+
+
+class GroundingDinoTransformerDecoder(DinoTransformerDecoder):
+
+ def _init_layers(self) -> None:
+ """Initialize decoder layers."""
+ self.layers = ModuleList([
+ GroundingDinoTransformerDecoderLayer(**self.layer_cfg)
+ for _ in range(self.num_layers)
+ ])
+ self.embed_dims = self.layers[0].embed_dims
+ if self.post_norm_cfg is not None:
+ raise ValueError('There is not post_norm in '
+ f'{self._get_name()}')
+ self.ref_point_head = MLP(self.embed_dims * 2, self.embed_dims,
+ self.embed_dims, 2)
+ self.norm = nn.LayerNorm(self.embed_dims)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/mask2former_layers.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/mask2former_layers.py
new file mode 100644
index 0000000000000000000000000000000000000000..dcc604e277d91151334ed520d78e6a5a8f388036
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/mask2former_layers.py
@@ -0,0 +1,135 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.cnn import build_norm_layer
+from mmengine.model import ModuleList
+from torch import Tensor
+
+from .deformable_detr_layers import DeformableDetrTransformerEncoder
+from .detr_layers import DetrTransformerDecoder, DetrTransformerDecoderLayer
+
+
+class Mask2FormerTransformerEncoder(DeformableDetrTransformerEncoder):
+ """Encoder in PixelDecoder of Mask2Former."""
+
+ def forward(self, query: Tensor, query_pos: Tensor,
+ key_padding_mask: Tensor, spatial_shapes: Tensor,
+ level_start_index: Tensor, valid_ratios: Tensor,
+ reference_points: Tensor, **kwargs) -> Tensor:
+ """Forward function of Transformer encoder.
+
+ Args:
+ query (Tensor): The input query, has shape (bs, num_queries, dim).
+ query_pos (Tensor): The positional encoding for query, has shape
+ (bs, num_queries, dim). If not None, it will be added to the
+ `query` before forward function. Defaults to None.
+ key_padding_mask (Tensor): The `key_padding_mask` of `self_attn`
+ input. ByteTensor, has shape (bs, num_queries).
+ spatial_shapes (Tensor): Spatial shapes of features in all levels,
+ has shape (num_levels, 2), last dimension represents (h, w).
+ level_start_index (Tensor): The start index of each level.
+ A tensor has shape (num_levels, ) and can be represented
+ as [0, h_0*w_0, h_0*w_0+h_1*w_1, ...].
+ valid_ratios (Tensor): The ratios of the valid width and the valid
+ height relative to the width and the height of features in all
+ levels, has shape (bs, num_levels, 2).
+ reference_points (Tensor): The initial reference, has shape
+ (bs, num_queries, 2) with the last dimension arranged
+ as (cx, cy).
+
+ Returns:
+ Tensor: Output queries of Transformer encoder, which is also
+ called 'encoder output embeddings' or 'memory', has shape
+ (bs, num_queries, dim)
+ """
+ for layer in self.layers:
+ query = layer(
+ query=query,
+ query_pos=query_pos,
+ key_padding_mask=key_padding_mask,
+ spatial_shapes=spatial_shapes,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios,
+ reference_points=reference_points,
+ **kwargs)
+ return query
+
+
+class Mask2FormerTransformerDecoder(DetrTransformerDecoder):
+ """Decoder of Mask2Former."""
+
+ def _init_layers(self) -> None:
+ """Initialize decoder layers."""
+ self.layers = ModuleList([
+ Mask2FormerTransformerDecoderLayer(**self.layer_cfg)
+ for _ in range(self.num_layers)
+ ])
+ self.embed_dims = self.layers[0].embed_dims
+ self.post_norm = build_norm_layer(self.post_norm_cfg,
+ self.embed_dims)[1]
+
+
+class Mask2FormerTransformerDecoderLayer(DetrTransformerDecoderLayer):
+ """Implements decoder layer in Mask2Former transformer."""
+
+ def forward(self,
+ query: Tensor,
+ key: Tensor = None,
+ value: Tensor = None,
+ query_pos: Tensor = None,
+ key_pos: Tensor = None,
+ self_attn_mask: Tensor = None,
+ cross_attn_mask: Tensor = None,
+ key_padding_mask: Tensor = None,
+ **kwargs) -> Tensor:
+ """
+ Args:
+ query (Tensor): The input query, has shape (bs, num_queries, dim).
+ key (Tensor, optional): The input key, has shape (bs, num_keys,
+ dim). If `None`, the `query` will be used. Defaults to `None`.
+ value (Tensor, optional): The input value, has the same shape as
+ `key`, as in `nn.MultiheadAttention.forward`. If `None`, the
+ `key` will be used. Defaults to `None`.
+ query_pos (Tensor, optional): The positional encoding for `query`,
+ has the same shape as `query`. If not `None`, it will be added
+ to `query` before forward function. Defaults to `None`.
+ key_pos (Tensor, optional): The positional encoding for `key`, has
+ the same shape as `key`. If not `None`, it will be added to
+ `key` before forward function. If None, and `query_pos` has the
+ same shape as `key`, then `query_pos` will be used for
+ `key_pos`. Defaults to None.
+ self_attn_mask (Tensor, optional): ByteTensor mask, has shape
+ (num_queries, num_keys), as in `nn.MultiheadAttention.forward`.
+ Defaults to None.
+ cross_attn_mask (Tensor, optional): ByteTensor mask, has shape
+ (num_queries, num_keys), as in `nn.MultiheadAttention.forward`.
+ Defaults to None.
+ key_padding_mask (Tensor, optional): The `key_padding_mask` of
+ `self_attn` input. ByteTensor, has shape (bs, num_value).
+ Defaults to None.
+
+ Returns:
+ Tensor: forwarded results, has shape (bs, num_queries, dim).
+ """
+
+ query = self.cross_attn(
+ query=query,
+ key=key,
+ value=value,
+ query_pos=query_pos,
+ key_pos=key_pos,
+ attn_mask=cross_attn_mask,
+ key_padding_mask=key_padding_mask,
+ **kwargs)
+ query = self.norms[0](query)
+ query = self.self_attn(
+ query=query,
+ key=query,
+ value=query,
+ query_pos=query_pos,
+ key_pos=query_pos,
+ attn_mask=self_attn_mask,
+ **kwargs)
+ query = self.norms[1](query)
+ query = self.ffn(query)
+ query = self.norms[2](query)
+
+ return query
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/utils.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/utils.py
new file mode 100644
index 0000000000000000000000000000000000000000..6e43a172ca7175b23c82f60894faf38ec6c437e3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/layers/transformer/utils.py
@@ -0,0 +1,915 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+import warnings
+from typing import Optional, Sequence, Tuple, Union
+
+import torch
+import torch.nn.functional as F
+from mmcv.cnn import (Linear, build_activation_layer, build_conv_layer,
+ build_norm_layer)
+from mmcv.cnn.bricks.drop import Dropout
+from mmengine.model import BaseModule, ModuleList
+from mmengine.utils import to_2tuple
+from torch import Tensor, nn
+
+from mmdet.registry import MODELS
+from mmdet.utils import OptConfigType, OptMultiConfig
+
+
+def nlc_to_nchw(x: Tensor, hw_shape: Sequence[int]) -> Tensor:
+ """Convert [N, L, C] shape tensor to [N, C, H, W] shape tensor.
+
+ Args:
+ x (Tensor): The input tensor of shape [N, L, C] before conversion.
+ hw_shape (Sequence[int]): The height and width of output feature map.
+
+ Returns:
+ Tensor: The output tensor of shape [N, C, H, W] after conversion.
+ """
+ H, W = hw_shape
+ assert len(x.shape) == 3
+ B, L, C = x.shape
+ assert L == H * W, 'The seq_len does not match H, W'
+ return x.transpose(1, 2).reshape(B, C, H, W).contiguous()
+
+
+def nchw_to_nlc(x):
+ """Flatten [N, C, H, W] shape tensor to [N, L, C] shape tensor.
+
+ Args:
+ x (Tensor): The input tensor of shape [N, C, H, W] before conversion.
+
+ Returns:
+ Tensor: The output tensor of shape [N, L, C] after conversion.
+ """
+ assert len(x.shape) == 4
+ return x.flatten(2).transpose(1, 2).contiguous()
+
+
+def coordinate_to_encoding(coord_tensor: Tensor,
+ num_feats: int = 128,
+ temperature: int = 10000,
+ scale: float = 2 * math.pi):
+ """Convert coordinate tensor to positional encoding.
+
+ Args:
+ coord_tensor (Tensor): Coordinate tensor to be converted to
+ positional encoding. With the last dimension as 2 or 4.
+ num_feats (int, optional): The feature dimension for each position
+ along x-axis or y-axis. Note the final returned dimension
+ for each position is 2 times of this value. Defaults to 128.
+ temperature (int, optional): The temperature used for scaling
+ the position embedding. Defaults to 10000.
+ scale (float, optional): A scale factor that scales the position
+ embedding. The scale will be used only when `normalize` is True.
+ Defaults to 2*pi.
+ Returns:
+ Tensor: Returned encoded positional tensor.
+ """
+ dim_t = torch.arange(
+ num_feats, dtype=torch.float32, device=coord_tensor.device)
+ dim_t = temperature**(2 * (dim_t // 2) / num_feats)
+ x_embed = coord_tensor[..., 0] * scale
+ y_embed = coord_tensor[..., 1] * scale
+ pos_x = x_embed[..., None] / dim_t
+ pos_y = y_embed[..., None] / dim_t
+ pos_x = torch.stack((pos_x[..., 0::2].sin(), pos_x[..., 1::2].cos()),
+ dim=-1).flatten(2)
+ pos_y = torch.stack((pos_y[..., 0::2].sin(), pos_y[..., 1::2].cos()),
+ dim=-1).flatten(2)
+ if coord_tensor.size(-1) == 2:
+ pos = torch.cat((pos_y, pos_x), dim=-1)
+ elif coord_tensor.size(-1) == 4:
+ w_embed = coord_tensor[..., 2] * scale
+ pos_w = w_embed[..., None] / dim_t
+ pos_w = torch.stack((pos_w[..., 0::2].sin(), pos_w[..., 1::2].cos()),
+ dim=-1).flatten(2)
+
+ h_embed = coord_tensor[..., 3] * scale
+ pos_h = h_embed[..., None] / dim_t
+ pos_h = torch.stack((pos_h[..., 0::2].sin(), pos_h[..., 1::2].cos()),
+ dim=-1).flatten(2)
+
+ pos = torch.cat((pos_y, pos_x, pos_w, pos_h), dim=-1)
+ else:
+ raise ValueError('Unknown pos_tensor shape(-1):{}'.format(
+ coord_tensor.size(-1)))
+ return pos
+
+
+def inverse_sigmoid(x: Tensor, eps: float = 1e-5) -> Tensor:
+ """Inverse function of sigmoid.
+
+ Args:
+ x (Tensor): The tensor to do the inverse.
+ eps (float): EPS avoid numerical overflow. Defaults 1e-5.
+ Returns:
+ Tensor: The x has passed the inverse function of sigmoid, has the same
+ shape with input.
+ """
+ x = x.clamp(min=0, max=1)
+ x1 = x.clamp(min=eps)
+ x2 = (1 - x).clamp(min=eps)
+ return torch.log(x1 / x2)
+
+
+class AdaptivePadding(nn.Module):
+ """Applies padding to input (if needed) so that input can get fully covered
+ by filter you specified. It support two modes "same" and "corner". The
+ "same" mode is same with "SAME" padding mode in TensorFlow, pad zero around
+ input. The "corner" mode would pad zero to bottom right.
+
+ Args:
+ kernel_size (int | tuple): Size of the kernel:
+ stride (int | tuple): Stride of the filter. Default: 1:
+ dilation (int | tuple): Spacing between kernel elements.
+ Default: 1
+ padding (str): Support "same" and "corner", "corner" mode
+ would pad zero to bottom right, and "same" mode would
+ pad zero around input. Default: "corner".
+ Example:
+ >>> kernel_size = 16
+ >>> stride = 16
+ >>> dilation = 1
+ >>> input = torch.rand(1, 1, 15, 17)
+ >>> adap_pad = AdaptivePadding(
+ >>> kernel_size=kernel_size,
+ >>> stride=stride,
+ >>> dilation=dilation,
+ >>> padding="corner")
+ >>> out = adap_pad(input)
+ >>> assert (out.shape[2], out.shape[3]) == (16, 32)
+ >>> input = torch.rand(1, 1, 16, 17)
+ >>> out = adap_pad(input)
+ >>> assert (out.shape[2], out.shape[3]) == (16, 32)
+ """
+
+ def __init__(self, kernel_size=1, stride=1, dilation=1, padding='corner'):
+
+ super(AdaptivePadding, self).__init__()
+
+ assert padding in ('same', 'corner')
+
+ kernel_size = to_2tuple(kernel_size)
+ stride = to_2tuple(stride)
+ padding = to_2tuple(padding)
+ dilation = to_2tuple(dilation)
+
+ self.padding = padding
+ self.kernel_size = kernel_size
+ self.stride = stride
+ self.dilation = dilation
+
+ def get_pad_shape(self, input_shape):
+ input_h, input_w = input_shape
+ kernel_h, kernel_w = self.kernel_size
+ stride_h, stride_w = self.stride
+ output_h = math.ceil(input_h / stride_h)
+ output_w = math.ceil(input_w / stride_w)
+ pad_h = max((output_h - 1) * stride_h +
+ (kernel_h - 1) * self.dilation[0] + 1 - input_h, 0)
+ pad_w = max((output_w - 1) * stride_w +
+ (kernel_w - 1) * self.dilation[1] + 1 - input_w, 0)
+ return pad_h, pad_w
+
+ def forward(self, x):
+ pad_h, pad_w = self.get_pad_shape(x.size()[-2:])
+ if pad_h > 0 or pad_w > 0:
+ if self.padding == 'corner':
+ x = F.pad(x, [0, pad_w, 0, pad_h])
+ elif self.padding == 'same':
+ x = F.pad(x, [
+ pad_w // 2, pad_w - pad_w // 2, pad_h // 2,
+ pad_h - pad_h // 2
+ ])
+ return x
+
+
+class PatchEmbed(BaseModule):
+ """Image to Patch Embedding.
+
+ We use a conv layer to implement PatchEmbed.
+
+ Args:
+ in_channels (int): The num of input channels. Default: 3
+ embed_dims (int): The dimensions of embedding. Default: 768
+ conv_type (str): The config dict for embedding
+ conv layer type selection. Default: "Conv2d.
+ kernel_size (int): The kernel_size of embedding conv. Default: 16.
+ stride (int): The slide stride of embedding conv.
+ Default: None (Would be set as `kernel_size`).
+ padding (int | tuple | string ): The padding length of
+ embedding conv. When it is a string, it means the mode
+ of adaptive padding, support "same" and "corner" now.
+ Default: "corner".
+ dilation (int): The dilation rate of embedding conv. Default: 1.
+ bias (bool): Bias of embed conv. Default: True.
+ norm_cfg (dict, optional): Config dict for normalization layer.
+ Default: None.
+ input_size (int | tuple | None): The size of input, which will be
+ used to calculate the out size. Only work when `dynamic_size`
+ is False. Default: None.
+ init_cfg (`mmengine.ConfigDict`, optional): The Config for
+ initialization. Default: None.
+ """
+
+ def __init__(self,
+ in_channels: int = 3,
+ embed_dims: int = 768,
+ conv_type: str = 'Conv2d',
+ kernel_size: int = 16,
+ stride: int = 16,
+ padding: Union[int, tuple, str] = 'corner',
+ dilation: int = 1,
+ bias: bool = True,
+ norm_cfg: OptConfigType = None,
+ input_size: Union[int, tuple] = None,
+ init_cfg: OptConfigType = None) -> None:
+ super(PatchEmbed, self).__init__(init_cfg=init_cfg)
+
+ self.embed_dims = embed_dims
+ if stride is None:
+ stride = kernel_size
+
+ kernel_size = to_2tuple(kernel_size)
+ stride = to_2tuple(stride)
+ dilation = to_2tuple(dilation)
+
+ if isinstance(padding, str):
+ self.adap_padding = AdaptivePadding(
+ kernel_size=kernel_size,
+ stride=stride,
+ dilation=dilation,
+ padding=padding)
+ # disable the padding of conv
+ padding = 0
+ else:
+ self.adap_padding = None
+ padding = to_2tuple(padding)
+
+ self.projection = build_conv_layer(
+ dict(type=conv_type),
+ in_channels=in_channels,
+ out_channels=embed_dims,
+ kernel_size=kernel_size,
+ stride=stride,
+ padding=padding,
+ dilation=dilation,
+ bias=bias)
+
+ if norm_cfg is not None:
+ self.norm = build_norm_layer(norm_cfg, embed_dims)[1]
+ else:
+ self.norm = None
+
+ if input_size:
+ input_size = to_2tuple(input_size)
+ # `init_out_size` would be used outside to
+ # calculate the num_patches
+ # when `use_abs_pos_embed` outside
+ self.init_input_size = input_size
+ if self.adap_padding:
+ pad_h, pad_w = self.adap_padding.get_pad_shape(input_size)
+ input_h, input_w = input_size
+ input_h = input_h + pad_h
+ input_w = input_w + pad_w
+ input_size = (input_h, input_w)
+
+ # https://pytorch.org/docs/stable/generated/torch.nn.Conv2d.html
+ h_out = (input_size[0] + 2 * padding[0] - dilation[0] *
+ (kernel_size[0] - 1) - 1) // stride[0] + 1
+ w_out = (input_size[1] + 2 * padding[1] - dilation[1] *
+ (kernel_size[1] - 1) - 1) // stride[1] + 1
+ self.init_out_size = (h_out, w_out)
+ else:
+ self.init_input_size = None
+ self.init_out_size = None
+
+ def forward(self, x: Tensor) -> Tuple[Tensor, Tuple[int]]:
+ """
+ Args:
+ x (Tensor): Has shape (B, C, H, W). In most case, C is 3.
+
+ Returns:
+ tuple: Contains merged results and its spatial shape.
+
+ - x (Tensor): Has shape (B, out_h * out_w, embed_dims)
+ - out_size (tuple[int]): Spatial shape of x, arrange as
+ (out_h, out_w).
+ """
+
+ if self.adap_padding:
+ x = self.adap_padding(x)
+
+ x = self.projection(x)
+ out_size = (x.shape[2], x.shape[3])
+ x = x.flatten(2).transpose(1, 2)
+ if self.norm is not None:
+ x = self.norm(x)
+ return x, out_size
+
+
+class PatchMerging(BaseModule):
+ """Merge patch feature map.
+
+ This layer groups feature map by kernel_size, and applies norm and linear
+ layers to the grouped feature map. Our implementation uses `nn.Unfold` to
+ merge patch, which is about 25% faster than original implementation.
+ Instead, we need to modify pretrained models for compatibility.
+
+ Args:
+ in_channels (int): The num of input channels.
+ to gets fully covered by filter and stride you specified..
+ Default: True.
+ out_channels (int): The num of output channels.
+ kernel_size (int | tuple, optional): the kernel size in the unfold
+ layer. Defaults to 2.
+ stride (int | tuple, optional): the stride of the sliding blocks in the
+ unfold layer. Default: None. (Would be set as `kernel_size`)
+ padding (int | tuple | string ): The padding length of
+ embedding conv. When it is a string, it means the mode
+ of adaptive padding, support "same" and "corner" now.
+ Default: "corner".
+ dilation (int | tuple, optional): dilation parameter in the unfold
+ layer. Default: 1.
+ bias (bool, optional): Whether to add bias in linear layer or not.
+ Defaults: False.
+ norm_cfg (dict, optional): Config dict for normalization layer.
+ Default: dict(type='LN').
+ init_cfg (dict, optional): The extra config for initialization.
+ Default: None.
+ """
+
+ def __init__(self,
+ in_channels: int,
+ out_channels: int,
+ kernel_size: Optional[Union[int, tuple]] = 2,
+ stride: Optional[Union[int, tuple]] = None,
+ padding: Union[int, tuple, str] = 'corner',
+ dilation: Optional[Union[int, tuple]] = 1,
+ bias: Optional[bool] = False,
+ norm_cfg: OptConfigType = dict(type='LN'),
+ init_cfg: OptConfigType = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.in_channels = in_channels
+ self.out_channels = out_channels
+ if stride:
+ stride = stride
+ else:
+ stride = kernel_size
+
+ kernel_size = to_2tuple(kernel_size)
+ stride = to_2tuple(stride)
+ dilation = to_2tuple(dilation)
+
+ if isinstance(padding, str):
+ self.adap_padding = AdaptivePadding(
+ kernel_size=kernel_size,
+ stride=stride,
+ dilation=dilation,
+ padding=padding)
+ # disable the padding of unfold
+ padding = 0
+ else:
+ self.adap_padding = None
+
+ padding = to_2tuple(padding)
+ self.sampler = nn.Unfold(
+ kernel_size=kernel_size,
+ dilation=dilation,
+ padding=padding,
+ stride=stride)
+
+ sample_dim = kernel_size[0] * kernel_size[1] * in_channels
+
+ if norm_cfg is not None:
+ self.norm = build_norm_layer(norm_cfg, sample_dim)[1]
+ else:
+ self.norm = None
+
+ self.reduction = nn.Linear(sample_dim, out_channels, bias=bias)
+
+ def forward(self, x: Tensor,
+ input_size: Tuple[int]) -> Tuple[Tensor, Tuple[int]]:
+ """
+ Args:
+ x (Tensor): Has shape (B, H*W, C_in).
+ input_size (tuple[int]): The spatial shape of x, arrange as (H, W).
+ Default: None.
+
+ Returns:
+ tuple: Contains merged results and its spatial shape.
+
+ - x (Tensor): Has shape (B, Merged_H * Merged_W, C_out)
+ - out_size (tuple[int]): Spatial shape of x, arrange as
+ (Merged_H, Merged_W).
+ """
+ B, L, C = x.shape
+ assert isinstance(input_size, Sequence), f'Expect ' \
+ f'input_size is ' \
+ f'`Sequence` ' \
+ f'but get {input_size}'
+
+ H, W = input_size
+ assert L == H * W, 'input feature has wrong size'
+
+ x = x.view(B, H, W, C).permute([0, 3, 1, 2]) # B, C, H, W
+ # Use nn.Unfold to merge patch. About 25% faster than original method,
+ # but need to modify pretrained model for compatibility
+
+ if self.adap_padding:
+ x = self.adap_padding(x)
+ H, W = x.shape[-2:]
+
+ x = self.sampler(x)
+ # if kernel_size=2 and stride=2, x should has shape (B, 4*C, H/2*W/2)
+
+ out_h = (H + 2 * self.sampler.padding[0] - self.sampler.dilation[0] *
+ (self.sampler.kernel_size[0] - 1) -
+ 1) // self.sampler.stride[0] + 1
+ out_w = (W + 2 * self.sampler.padding[1] - self.sampler.dilation[1] *
+ (self.sampler.kernel_size[1] - 1) -
+ 1) // self.sampler.stride[1] + 1
+
+ output_size = (out_h, out_w)
+ x = x.transpose(1, 2) # B, H/2*W/2, 4*C
+ x = self.norm(x) if self.norm else x
+ x = self.reduction(x)
+ return x, output_size
+
+
+class ConditionalAttention(BaseModule):
+ """A wrapper of conditional attention, dropout and residual connection.
+
+ Args:
+ embed_dims (int): The embedding dimension.
+ num_heads (int): Parallel attention heads.
+ attn_drop (float): A Dropout layer on attn_output_weights.
+ Default: 0.0.
+ proj_drop: A Dropout layer after `nn.MultiheadAttention`.
+ Default: 0.0.
+ cross_attn (bool): Whether the attention module is for cross attention.
+ Default: False
+ keep_query_pos (bool): Whether to transform query_pos before cross
+ attention.
+ Default: False.
+ batch_first (bool): When it is True, Key, Query and Value are shape of
+ (batch, n, embed_dim), otherwise (n, batch, embed_dim).
+ Default: True.
+ init_cfg (obj:`mmcv.ConfigDict`): The Config for initialization.
+ Default: None.
+ """
+
+ def __init__(self,
+ embed_dims: int,
+ num_heads: int,
+ attn_drop: float = 0.,
+ proj_drop: float = 0.,
+ cross_attn: bool = False,
+ keep_query_pos: bool = False,
+ batch_first: bool = True,
+ init_cfg: OptMultiConfig = None):
+ super().__init__(init_cfg=init_cfg)
+
+ assert batch_first is True, 'Set `batch_first`\
+ to False is NOT supported in ConditionalAttention. \
+ First dimension of all DETRs in mmdet is `batch`, \
+ please set `batch_first` to True.'
+
+ self.cross_attn = cross_attn
+ self.keep_query_pos = keep_query_pos
+ self.embed_dims = embed_dims
+ self.num_heads = num_heads
+ self.attn_drop = Dropout(attn_drop)
+ self.proj_drop = Dropout(proj_drop)
+
+ self._init_layers()
+
+ def _init_layers(self):
+ """Initialize layers for qkv projection."""
+ embed_dims = self.embed_dims
+ self.qcontent_proj = Linear(embed_dims, embed_dims)
+ self.qpos_proj = Linear(embed_dims, embed_dims)
+ self.kcontent_proj = Linear(embed_dims, embed_dims)
+ self.kpos_proj = Linear(embed_dims, embed_dims)
+ self.v_proj = Linear(embed_dims, embed_dims)
+ if self.cross_attn:
+ self.qpos_sine_proj = Linear(embed_dims, embed_dims)
+ self.out_proj = Linear(embed_dims, embed_dims)
+
+ nn.init.constant_(self.out_proj.bias, 0.)
+
+ def forward_attn(self,
+ query: Tensor,
+ key: Tensor,
+ value: Tensor,
+ attn_mask: Tensor = None,
+ key_padding_mask: Tensor = None) -> Tuple[Tensor]:
+ """Forward process for `ConditionalAttention`.
+
+ Args:
+ query (Tensor): The input query with shape [bs, num_queries,
+ embed_dims].
+ key (Tensor): The key tensor with shape [bs, num_keys,
+ embed_dims].
+ If None, the `query` will be used. Defaults to None.
+ value (Tensor): The value tensor with same shape as `key`.
+ Same in `nn.MultiheadAttention.forward`. Defaults to None.
+ If None, the `key` will be used.
+ attn_mask (Tensor): ByteTensor mask with shape [num_queries,
+ num_keys]. Same in `nn.MultiheadAttention.forward`.
+ Defaults to None.
+ key_padding_mask (Tensor): ByteTensor with shape [bs, num_keys].
+ Defaults to None.
+ Returns:
+ Tuple[Tensor]: Attention outputs of shape :math:`(N, L, E)`,
+ where :math:`N` is the batch size, :math:`L` is the target
+ sequence length , and :math:`E` is the embedding dimension
+ `embed_dim`. Attention weights per head of shape :math:`
+ (num_heads, L, S)`. where :math:`N` is batch size, :math:`L`
+ is target sequence length, and :math:`S` is the source sequence
+ length.
+ """
+ assert key.size(1) == value.size(1), \
+ f'{"key, value must have the same sequence length"}'
+ assert query.size(0) == key.size(0) == value.size(0), \
+ f'{"batch size must be equal for query, key, value"}'
+ assert query.size(2) == key.size(2), \
+ f'{"q_dims, k_dims must be equal"}'
+ assert value.size(2) == self.embed_dims, \
+ f'{"v_dims must be equal to embed_dims"}'
+
+ bs, tgt_len, hidden_dims = query.size()
+ _, src_len, _ = key.size()
+ head_dims = hidden_dims // self.num_heads
+ v_head_dims = self.embed_dims // self.num_heads
+ assert head_dims * self.num_heads == hidden_dims, \
+ f'{"hidden_dims must be divisible by num_heads"}'
+ scaling = float(head_dims)**-0.5
+
+ q = query * scaling
+ k = key
+ v = value
+
+ if attn_mask is not None:
+ assert attn_mask.dtype == torch.float32 or \
+ attn_mask.dtype == torch.float64 or \
+ attn_mask.dtype == torch.float16 or \
+ attn_mask.dtype == torch.uint8 or \
+ attn_mask.dtype == torch.bool, \
+ 'Only float, byte, and bool types are supported for \
+ attn_mask'
+
+ if attn_mask.dtype == torch.uint8:
+ warnings.warn('Byte tensor for attn_mask is deprecated.\
+ Use bool tensor instead.')
+ attn_mask = attn_mask.to(torch.bool)
+ if attn_mask.dim() == 2:
+ attn_mask = attn_mask.unsqueeze(0)
+ if list(attn_mask.size()) != [1, query.size(1), key.size(1)]:
+ raise RuntimeError(
+ 'The size of the 2D attn_mask is not correct.')
+ elif attn_mask.dim() == 3:
+ if list(attn_mask.size()) != [
+ bs * self.num_heads,
+ query.size(1),
+ key.size(1)
+ ]:
+ raise RuntimeError(
+ 'The size of the 3D attn_mask is not correct.')
+ else:
+ raise RuntimeError(
+ "attn_mask's dimension {} is not supported".format(
+ attn_mask.dim()))
+ # attn_mask's dim is 3 now.
+
+ if key_padding_mask is not None and key_padding_mask.dtype == int:
+ key_padding_mask = key_padding_mask.to(torch.bool)
+
+ q = q.contiguous().view(bs, tgt_len, self.num_heads,
+ head_dims).permute(0, 2, 1, 3).flatten(0, 1)
+ if k is not None:
+ k = k.contiguous().view(bs, src_len, self.num_heads,
+ head_dims).permute(0, 2, 1,
+ 3).flatten(0, 1)
+ if v is not None:
+ v = v.contiguous().view(bs, src_len, self.num_heads,
+ v_head_dims).permute(0, 2, 1,
+ 3).flatten(0, 1)
+
+ if key_padding_mask is not None:
+ assert key_padding_mask.size(0) == bs
+ assert key_padding_mask.size(1) == src_len
+
+ attn_output_weights = torch.bmm(q, k.transpose(1, 2))
+ assert list(attn_output_weights.size()) == [
+ bs * self.num_heads, tgt_len, src_len
+ ]
+
+ if attn_mask is not None:
+ if attn_mask.dtype == torch.bool:
+ attn_output_weights.masked_fill_(attn_mask, float('-inf'))
+ else:
+ attn_output_weights += attn_mask
+
+ if key_padding_mask is not None:
+ attn_output_weights = attn_output_weights.view(
+ bs, self.num_heads, tgt_len, src_len)
+ attn_output_weights = attn_output_weights.masked_fill(
+ key_padding_mask.unsqueeze(1).unsqueeze(2),
+ float('-inf'),
+ )
+ attn_output_weights = attn_output_weights.view(
+ bs * self.num_heads, tgt_len, src_len)
+
+ attn_output_weights = F.softmax(
+ attn_output_weights -
+ attn_output_weights.max(dim=-1, keepdim=True)[0],
+ dim=-1)
+ attn_output_weights = self.attn_drop(attn_output_weights)
+
+ attn_output = torch.bmm(attn_output_weights, v)
+ assert list(
+ attn_output.size()) == [bs * self.num_heads, tgt_len, v_head_dims]
+ attn_output = attn_output.view(bs, self.num_heads, tgt_len,
+ v_head_dims).permute(0, 2, 1,
+ 3).flatten(2)
+ attn_output = self.out_proj(attn_output)
+
+ # average attention weights over heads
+ attn_output_weights = attn_output_weights.view(bs, self.num_heads,
+ tgt_len, src_len)
+ return attn_output, attn_output_weights.sum(dim=1) / self.num_heads
+
+ def forward(self,
+ query: Tensor,
+ key: Tensor,
+ query_pos: Tensor = None,
+ ref_sine_embed: Tensor = None,
+ key_pos: Tensor = None,
+ attn_mask: Tensor = None,
+ key_padding_mask: Tensor = None,
+ is_first: bool = False) -> Tensor:
+ """Forward function for `ConditionalAttention`.
+ Args:
+ query (Tensor): The input query with shape [bs, num_queries,
+ embed_dims].
+ key (Tensor): The key tensor with shape [bs, num_keys,
+ embed_dims].
+ If None, the `query` will be used. Defaults to None.
+ query_pos (Tensor): The positional encoding for query in self
+ attention, with the same shape as `x`. If not None, it will
+ be added to `x` before forward function.
+ Defaults to None.
+ query_sine_embed (Tensor): The positional encoding for query in
+ cross attention, with the same shape as `x`. If not None, it
+ will be added to `x` before forward function.
+ Defaults to None.
+ key_pos (Tensor): The positional encoding for `key`, with the
+ same shape as `key`. Defaults to None. If not None, it will
+ be added to `key` before forward function. If None, and
+ `query_pos` has the same shape as `key`, then `query_pos`
+ will be used for `key_pos`. Defaults to None.
+ attn_mask (Tensor): ByteTensor mask with shape [num_queries,
+ num_keys]. Same in `nn.MultiheadAttention.forward`.
+ Defaults to None.
+ key_padding_mask (Tensor): ByteTensor with shape [bs, num_keys].
+ Defaults to None.
+ is_first (bool): A indicator to tell whether the current layer
+ is the first layer of the decoder.
+ Defaults to False.
+ Returns:
+ Tensor: forwarded results with shape
+ [bs, num_queries, embed_dims].
+ """
+
+ if self.cross_attn:
+ q_content = self.qcontent_proj(query)
+ k_content = self.kcontent_proj(key)
+ v = self.v_proj(key)
+
+ bs, nq, c = q_content.size()
+ _, hw, _ = k_content.size()
+
+ k_pos = self.kpos_proj(key_pos)
+ if is_first or self.keep_query_pos:
+ q_pos = self.qpos_proj(query_pos)
+ q = q_content + q_pos
+ k = k_content + k_pos
+ else:
+ q = q_content
+ k = k_content
+ q = q.view(bs, nq, self.num_heads, c // self.num_heads)
+ query_sine_embed = self.qpos_sine_proj(ref_sine_embed)
+ query_sine_embed = query_sine_embed.view(bs, nq, self.num_heads,
+ c // self.num_heads)
+ q = torch.cat([q, query_sine_embed], dim=3).view(bs, nq, 2 * c)
+ k = k.view(bs, hw, self.num_heads, c // self.num_heads)
+ k_pos = k_pos.view(bs, hw, self.num_heads, c // self.num_heads)
+ k = torch.cat([k, k_pos], dim=3).view(bs, hw, 2 * c)
+ ca_output = self.forward_attn(
+ query=q,
+ key=k,
+ value=v,
+ attn_mask=attn_mask,
+ key_padding_mask=key_padding_mask)[0]
+ query = query + self.proj_drop(ca_output)
+ else:
+ q_content = self.qcontent_proj(query)
+ q_pos = self.qpos_proj(query_pos)
+ k_content = self.kcontent_proj(query)
+ k_pos = self.kpos_proj(query_pos)
+ v = self.v_proj(query)
+ q = q_content if q_pos is None else q_content + q_pos
+ k = k_content if k_pos is None else k_content + k_pos
+ sa_output = self.forward_attn(
+ query=q,
+ key=k,
+ value=v,
+ attn_mask=attn_mask,
+ key_padding_mask=key_padding_mask)[0]
+ query = query + self.proj_drop(sa_output)
+
+ return query
+
+
+class MLP(BaseModule):
+ """Very simple multi-layer perceptron (also called FFN) with relu. Mostly
+ used in DETR series detectors.
+
+ Args:
+ input_dim (int): Feature dim of the input tensor.
+ hidden_dim (int): Feature dim of the hidden layer.
+ output_dim (int): Feature dim of the output tensor.
+ num_layers (int): Number of FFN layers. As the last
+ layer of MLP only contains FFN (Linear).
+ """
+
+ def __init__(self, input_dim: int, hidden_dim: int, output_dim: int,
+ num_layers: int) -> None:
+ super().__init__()
+ self.num_layers = num_layers
+ h = [hidden_dim] * (num_layers - 1)
+ self.layers = ModuleList(
+ Linear(n, k) for n, k in zip([input_dim] + h, h + [output_dim]))
+
+ def forward(self, x: Tensor) -> Tensor:
+ """Forward function of MLP.
+
+ Args:
+ x (Tensor): The input feature, has shape
+ (num_queries, bs, input_dim).
+ Returns:
+ Tensor: The output feature, has shape
+ (num_queries, bs, output_dim).
+ """
+ for i, layer in enumerate(self.layers):
+ x = F.relu(layer(x)) if i < self.num_layers - 1 else layer(x)
+ return x
+
+
+@MODELS.register_module()
+class DynamicConv(BaseModule):
+ """Implements Dynamic Convolution.
+
+ This module generate parameters for each sample and
+ use bmm to implement 1*1 convolution. Code is modified
+ from the `official github repo `_ .
+
+ Args:
+ in_channels (int): The input feature channel.
+ Defaults to 256.
+ feat_channels (int): The inner feature channel.
+ Defaults to 64.
+ out_channels (int, optional): The output feature channel.
+ When not specified, it will be set to `in_channels`
+ by default
+ input_feat_shape (int): The shape of input feature.
+ Defaults to 7.
+ with_proj (bool): Project two-dimentional feature to
+ one-dimentional feature. Default to True.
+ act_cfg (dict): The activation config for DynamicConv.
+ norm_cfg (dict): Config dict for normalization layer. Default
+ layer normalization.
+ init_cfg (obj:`mmengine.ConfigDict`): The Config for initialization.
+ Default: None.
+ """
+
+ def __init__(self,
+ in_channels: int = 256,
+ feat_channels: int = 64,
+ out_channels: Optional[int] = None,
+ input_feat_shape: int = 7,
+ with_proj: bool = True,
+ act_cfg: OptConfigType = dict(type='ReLU', inplace=True),
+ norm_cfg: OptConfigType = dict(type='LN'),
+ init_cfg: OptConfigType = None) -> None:
+ super(DynamicConv, self).__init__(init_cfg)
+ self.in_channels = in_channels
+ self.feat_channels = feat_channels
+ self.out_channels_raw = out_channels
+ self.input_feat_shape = input_feat_shape
+ self.with_proj = with_proj
+ self.act_cfg = act_cfg
+ self.norm_cfg = norm_cfg
+ self.out_channels = out_channels if out_channels else in_channels
+
+ self.num_params_in = self.in_channels * self.feat_channels
+ self.num_params_out = self.out_channels * self.feat_channels
+ self.dynamic_layer = nn.Linear(
+ self.in_channels, self.num_params_in + self.num_params_out)
+
+ self.norm_in = build_norm_layer(norm_cfg, self.feat_channels)[1]
+ self.norm_out = build_norm_layer(norm_cfg, self.out_channels)[1]
+
+ self.activation = build_activation_layer(act_cfg)
+
+ num_output = self.out_channels * input_feat_shape**2
+ if self.with_proj:
+ self.fc_layer = nn.Linear(num_output, self.out_channels)
+ self.fc_norm = build_norm_layer(norm_cfg, self.out_channels)[1]
+
+ def forward(self, param_feature: Tensor, input_feature: Tensor) -> Tensor:
+ """Forward function for `DynamicConv`.
+
+ Args:
+ param_feature (Tensor): The feature can be used
+ to generate the parameter, has shape
+ (num_all_proposals, in_channels).
+ input_feature (Tensor): Feature that
+ interact with parameters, has shape
+ (num_all_proposals, in_channels, H, W).
+
+ Returns:
+ Tensor: The output feature has shape
+ (num_all_proposals, out_channels).
+ """
+ input_feature = input_feature.flatten(2).permute(2, 0, 1)
+
+ input_feature = input_feature.permute(1, 0, 2)
+ parameters = self.dynamic_layer(param_feature)
+
+ param_in = parameters[:, :self.num_params_in].view(
+ -1, self.in_channels, self.feat_channels)
+ param_out = parameters[:, -self.num_params_out:].view(
+ -1, self.feat_channels, self.out_channels)
+
+ # input_feature has shape (num_all_proposals, H*W, in_channels)
+ # param_in has shape (num_all_proposals, in_channels, feat_channels)
+ # feature has shape (num_all_proposals, H*W, feat_channels)
+ features = torch.bmm(input_feature, param_in)
+ features = self.norm_in(features)
+ features = self.activation(features)
+
+ # param_out has shape (batch_size, feat_channels, out_channels)
+ features = torch.bmm(features, param_out)
+ features = self.norm_out(features)
+ features = self.activation(features)
+
+ if self.with_proj:
+ features = features.flatten(1)
+ features = self.fc_layer(features)
+ features = self.fc_norm(features)
+ features = self.activation(features)
+
+ return features
+
+
+def get_text_sine_pos_embed(
+ pos_tensor: torch.Tensor,
+ num_pos_feats: int = 128,
+ temperature: int = 10000,
+ exchange_xy: bool = True,
+):
+ """generate sine position embedding from a position tensor
+ Args:
+ pos_tensor (torch.Tensor): shape: [..., n].
+ num_pos_feats (int): projected shape for each float in the tensor.
+ temperature (int): temperature in the sine/cosine function.
+ exchange_xy (bool, optional): exchange pos x and pos y. For example,
+ input tensor is [x,y], the results will be [pos(y), pos(x)].
+ Defaults to True.
+ Returns:
+ pos_embed (torch.Tensor): shape: [..., n*num_pos_feats].
+ """
+ scale = 2 * math.pi
+ dim_t = torch.arange(
+ num_pos_feats, dtype=torch.float32, device=pos_tensor.device)
+ dim_t = temperature**(2 * torch.div(dim_t, 2, rounding_mode='floor') /
+ num_pos_feats)
+
+ def sine_func(x: torch.Tensor):
+ sin_x = x * scale / dim_t
+ sin_x = torch.stack((sin_x[..., 0::2].sin(), sin_x[..., 1::2].cos()),
+ dim=3).flatten(2)
+ return sin_x
+
+ pos_res = [
+ sine_func(x)
+ for x in pos_tensor.split([1] * pos_tensor.shape[-1], dim=-1)
+ ]
+ if exchange_xy:
+ pos_res[0], pos_res[1] = pos_res[1], pos_res[0]
+ pos_res = torch.cat(pos_res, dim=-1)
+ return pos_res
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..7c57a3a96879c6bd5eb61c300d316e2b4579b287
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/__init__.py
@@ -0,0 +1,42 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .accuracy import Accuracy, accuracy
+from .ae_loss import AssociativeEmbeddingLoss
+from .balanced_l1_loss import BalancedL1Loss, balanced_l1_loss
+from .cross_entropy_loss import (CrossEntropyCustomLoss, CrossEntropyLoss,
+ binary_cross_entropy, cross_entropy,
+ mask_cross_entropy)
+from .ddq_detr_aux_loss import DDQAuxLoss
+from .dice_loss import DiceLoss
+from .eqlv2_loss import EQLV2Loss
+from .focal_loss import FocalCustomLoss, FocalLoss, sigmoid_focal_loss
+from .gaussian_focal_loss import GaussianFocalLoss
+from .gfocal_loss import DistributionFocalLoss, QualityFocalLoss
+from .ghm_loss import GHMC, GHMR
+from .iou_loss import (BoundedIoULoss, CIoULoss, DIoULoss, EIoULoss, GIoULoss,
+ IoULoss, SIoULoss, bounded_iou_loss, iou_loss)
+from .kd_loss import KnowledgeDistillationKLDivLoss
+from .l2_loss import L2Loss
+from .margin_loss import MarginL2Loss
+from .mse_loss import MSELoss, mse_loss
+from .multipos_cross_entropy_loss import MultiPosCrossEntropyLoss
+from .pisa_loss import carl_loss, isr_p
+from .seesaw_loss import SeesawLoss
+from .smooth_l1_loss import L1Loss, SmoothL1Loss, l1_loss, smooth_l1_loss
+from .triplet_loss import TripletLoss
+from .utils import reduce_loss, weight_reduce_loss, weighted_loss
+from .varifocal_loss import VarifocalLoss
+
+__all__ = [
+ 'accuracy', 'Accuracy', 'cross_entropy', 'binary_cross_entropy',
+ 'mask_cross_entropy', 'CrossEntropyLoss', 'sigmoid_focal_loss',
+ 'FocalLoss', 'smooth_l1_loss', 'SmoothL1Loss', 'balanced_l1_loss',
+ 'BalancedL1Loss', 'mse_loss', 'MSELoss', 'iou_loss', 'bounded_iou_loss',
+ 'IoULoss', 'BoundedIoULoss', 'GIoULoss', 'DIoULoss', 'CIoULoss',
+ 'EIoULoss', 'SIoULoss', 'GHMC', 'GHMR', 'reduce_loss',
+ 'weight_reduce_loss', 'weighted_loss', 'L1Loss', 'l1_loss', 'isr_p',
+ 'carl_loss', 'AssociativeEmbeddingLoss', 'GaussianFocalLoss',
+ 'QualityFocalLoss', 'DistributionFocalLoss', 'VarifocalLoss',
+ 'KnowledgeDistillationKLDivLoss', 'SeesawLoss', 'DiceLoss', 'EQLV2Loss',
+ 'MarginL2Loss', 'MultiPosCrossEntropyLoss', 'L2Loss', 'TripletLoss',
+ 'DDQAuxLoss', 'CrossEntropyCustomLoss', 'FocalCustomLoss'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/accuracy.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/accuracy.py
new file mode 100644
index 0000000000000000000000000000000000000000..d68484e13965ced3bd6b104071d22657a9b3fde6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/accuracy.py
@@ -0,0 +1,77 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch.nn as nn
+
+
+def accuracy(pred, target, topk=1, thresh=None):
+ """Calculate accuracy according to the prediction and target.
+
+ Args:
+ pred (torch.Tensor): The model prediction, shape (N, num_class)
+ target (torch.Tensor): The target of each prediction, shape (N, )
+ topk (int | tuple[int], optional): If the predictions in ``topk``
+ matches the target, the predictions will be regarded as
+ correct ones. Defaults to 1.
+ thresh (float, optional): If not None, predictions with scores under
+ this threshold are considered incorrect. Default to None.
+
+ Returns:
+ float | tuple[float]: If the input ``topk`` is a single integer,
+ the function will return a single float as accuracy. If
+ ``topk`` is a tuple containing multiple integers, the
+ function will return a tuple containing accuracies of
+ each ``topk`` number.
+ """
+ assert isinstance(topk, (int, tuple))
+ if isinstance(topk, int):
+ topk = (topk, )
+ return_single = True
+ else:
+ return_single = False
+
+ maxk = max(topk)
+ if pred.size(0) == 0:
+ accu = [pred.new_tensor(0.) for i in range(len(topk))]
+ return accu[0] if return_single else accu
+ assert pred.ndim == 2 and target.ndim == 1
+ assert pred.size(0) == target.size(0)
+ assert maxk <= pred.size(1), \
+ f'maxk {maxk} exceeds pred dimension {pred.size(1)}'
+ pred_value, pred_label = pred.topk(maxk, dim=1)
+ pred_label = pred_label.t() # transpose to shape (maxk, N)
+ correct = pred_label.eq(target.view(1, -1).expand_as(pred_label))
+ if thresh is not None:
+ # Only prediction values larger than thresh are counted as correct
+ correct = correct & (pred_value > thresh).t()
+ res = []
+ for k in topk:
+ correct_k = correct[:k].reshape(-1).float().sum(0, keepdim=True)
+ res.append(correct_k.mul_(100.0 / pred.size(0)))
+ return res[0] if return_single else res
+
+
+class Accuracy(nn.Module):
+
+ def __init__(self, topk=(1, ), thresh=None):
+ """Module to calculate the accuracy.
+
+ Args:
+ topk (tuple, optional): The criterion used to calculate the
+ accuracy. Defaults to (1,).
+ thresh (float, optional): If not None, predictions with scores
+ under this threshold are considered incorrect. Default to None.
+ """
+ super().__init__()
+ self.topk = topk
+ self.thresh = thresh
+
+ def forward(self, pred, target):
+ """Forward function to calculate accuracy.
+
+ Args:
+ pred (torch.Tensor): Prediction of models.
+ target (torch.Tensor): Target for each prediction.
+
+ Returns:
+ tuple[float]: The accuracies under different topk criterions.
+ """
+ return accuracy(pred, target, self.topk, self.thresh)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/ae_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/ae_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..2aa7d696be4b937a2d45545a8309aaa936fe5f22
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/ae_loss.py
@@ -0,0 +1,101 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+
+from mmdet.registry import MODELS
+
+
+def ae_loss_per_image(tl_preds, br_preds, match):
+ """Associative Embedding Loss in one image.
+
+ Associative Embedding Loss including two parts: pull loss and push loss.
+ Pull loss makes embedding vectors from same object closer to each other.
+ Push loss distinguish embedding vector from different objects, and makes
+ the gap between them is large enough.
+
+ During computing, usually there are 3 cases:
+ - no object in image: both pull loss and push loss will be 0.
+ - one object in image: push loss will be 0 and pull loss is computed
+ by the two corner of the only object.
+ - more than one objects in image: pull loss is computed by corner pairs
+ from each object, push loss is computed by each object with all
+ other objects. We use confusion matrix with 0 in diagonal to
+ compute the push loss.
+
+ Args:
+ tl_preds (tensor): Embedding feature map of left-top corner.
+ br_preds (tensor): Embedding feature map of bottim-right corner.
+ match (list): Downsampled coordinates pair of each ground truth box.
+ """
+
+ tl_list, br_list, me_list = [], [], []
+ if len(match) == 0: # no object in image
+ pull_loss = tl_preds.sum() * 0.
+ push_loss = tl_preds.sum() * 0.
+ else:
+ for m in match:
+ [tl_y, tl_x], [br_y, br_x] = m
+ tl_e = tl_preds[:, tl_y, tl_x].view(-1, 1)
+ br_e = br_preds[:, br_y, br_x].view(-1, 1)
+ tl_list.append(tl_e)
+ br_list.append(br_e)
+ me_list.append((tl_e + br_e) / 2.0)
+
+ tl_list = torch.cat(tl_list)
+ br_list = torch.cat(br_list)
+ me_list = torch.cat(me_list)
+
+ assert tl_list.size() == br_list.size()
+
+ # N is object number in image, M is dimension of embedding vector
+ N, M = tl_list.size()
+
+ pull_loss = (tl_list - me_list).pow(2) + (br_list - me_list).pow(2)
+ pull_loss = pull_loss.sum() / N
+
+ margin = 1 # exp setting of CornerNet, details in section 3.3 of paper
+
+ # confusion matrix of push loss
+ conf_mat = me_list.expand((N, N, M)).permute(1, 0, 2) - me_list
+ conf_weight = 1 - torch.eye(N).type_as(me_list)
+ conf_mat = conf_weight * (margin - conf_mat.sum(-1).abs())
+
+ if N > 1: # more than one object in current image
+ push_loss = F.relu(conf_mat).sum() / (N * (N - 1))
+ else:
+ push_loss = tl_preds.sum() * 0.
+
+ return pull_loss, push_loss
+
+
+@MODELS.register_module()
+class AssociativeEmbeddingLoss(nn.Module):
+ """Associative Embedding Loss.
+
+ More details can be found in
+ `Associative Embedding `_ and
+ `CornerNet `_ .
+ Code is modified from `kp_utils.py `_ # noqa: E501
+
+ Args:
+ pull_weight (float): Loss weight for corners from same object.
+ push_weight (float): Loss weight for corners from different object.
+ """
+
+ def __init__(self, pull_weight=0.25, push_weight=0.25):
+ super(AssociativeEmbeddingLoss, self).__init__()
+ self.pull_weight = pull_weight
+ self.push_weight = push_weight
+
+ def forward(self, pred, target, match):
+ """Forward function."""
+ batch = pred.size(0)
+ pull_all, push_all = 0.0, 0.0
+ for i in range(batch):
+ pull, push = ae_loss_per_image(pred[i], target[i], match[i])
+
+ pull_all += self.pull_weight * pull
+ push_all += self.push_weight * push
+
+ return pull_all, push_all
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/balanced_l1_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/balanced_l1_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..25adaab2239e871476d9d4e3cbb1a238c3043041
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/balanced_l1_loss.py
@@ -0,0 +1,122 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import numpy as np
+import torch
+import torch.nn as nn
+
+from mmdet.registry import MODELS
+from .utils import weighted_loss
+
+
+@weighted_loss
+def balanced_l1_loss(pred,
+ target,
+ beta=1.0,
+ alpha=0.5,
+ gamma=1.5,
+ reduction='mean'):
+ """Calculate balanced L1 loss.
+
+ Please see the `Libra R-CNN `_
+
+ Args:
+ pred (torch.Tensor): The prediction with shape (N, 4).
+ target (torch.Tensor): The learning target of the prediction with
+ shape (N, 4).
+ beta (float): The loss is a piecewise function of prediction and target
+ and ``beta`` serves as a threshold for the difference between the
+ prediction and target. Defaults to 1.0.
+ alpha (float): The denominator ``alpha`` in the balanced L1 loss.
+ Defaults to 0.5.
+ gamma (float): The ``gamma`` in the balanced L1 loss.
+ Defaults to 1.5.
+ reduction (str, optional): The method that reduces the loss to a
+ scalar. Options are "none", "mean" and "sum".
+
+ Returns:
+ torch.Tensor: The calculated loss
+ """
+ assert beta > 0
+ if target.numel() == 0:
+ return pred.sum() * 0
+
+ assert pred.size() == target.size()
+
+ diff = torch.abs(pred - target)
+ b = np.e**(gamma / alpha) - 1
+ loss = torch.where(
+ diff < beta, alpha / b *
+ (b * diff + 1) * torch.log(b * diff / beta + 1) - alpha * diff,
+ gamma * diff + gamma / b - alpha * beta)
+
+ return loss
+
+
+@MODELS.register_module()
+class BalancedL1Loss(nn.Module):
+ """Balanced L1 Loss.
+
+ arXiv: https://arxiv.org/pdf/1904.02701.pdf (CVPR 2019)
+
+ Args:
+ alpha (float): The denominator ``alpha`` in the balanced L1 loss.
+ Defaults to 0.5.
+ gamma (float): The ``gamma`` in the balanced L1 loss. Defaults to 1.5.
+ beta (float, optional): The loss is a piecewise function of prediction
+ and target. ``beta`` serves as a threshold for the difference
+ between the prediction and target. Defaults to 1.0.
+ reduction (str, optional): The method that reduces the loss to a
+ scalar. Options are "none", "mean" and "sum".
+ loss_weight (float, optional): The weight of the loss. Defaults to 1.0
+ """
+
+ def __init__(self,
+ alpha=0.5,
+ gamma=1.5,
+ beta=1.0,
+ reduction='mean',
+ loss_weight=1.0):
+ super(BalancedL1Loss, self).__init__()
+ self.alpha = alpha
+ self.gamma = gamma
+ self.beta = beta
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+
+ def forward(self,
+ pred,
+ target,
+ weight=None,
+ avg_factor=None,
+ reduction_override=None,
+ **kwargs):
+ """Forward function of loss.
+
+ Args:
+ pred (torch.Tensor): The prediction with shape (N, 4).
+ target (torch.Tensor): The learning target of the prediction with
+ shape (N, 4).
+ weight (torch.Tensor, optional): Sample-wise loss weight with
+ shape (N, ).
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Options are "none", "mean" and "sum".
+
+ Returns:
+ torch.Tensor: The calculated loss
+ """
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ loss_bbox = self.loss_weight * balanced_l1_loss(
+ pred,
+ target,
+ weight,
+ alpha=self.alpha,
+ gamma=self.gamma,
+ beta=self.beta,
+ reduction=reduction,
+ avg_factor=avg_factor,
+ **kwargs)
+ return loss_bbox
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/cross_entropy_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/cross_entropy_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..49fac7743ceddd2454f44b76c63d514de43b5aef
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/cross_entropy_loss.py
@@ -0,0 +1,401 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+
+from mmdet.registry import MODELS
+from .accuracy import accuracy
+from .utils import weight_reduce_loss
+
+
+def cross_entropy(pred,
+ label,
+ weight=None,
+ reduction='mean',
+ avg_factor=None,
+ class_weight=None,
+ ignore_index=-100,
+ avg_non_ignore=False):
+ """Calculate the CrossEntropy loss.
+
+ Args:
+ pred (torch.Tensor): The prediction with shape (N, C), C is the number
+ of classes.
+ label (torch.Tensor): The learning label of the prediction.
+ weight (torch.Tensor, optional): Sample-wise loss weight.
+ reduction (str, optional): The method used to reduce the loss.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ class_weight (list[float], optional): The weight for each class.
+ ignore_index (int | None): The label index to be ignored.
+ If None, it will be set to default value. Default: -100.
+ avg_non_ignore (bool): The flag decides to whether the loss is
+ only averaged over non-ignored targets. Default: False.
+
+ Returns:
+ torch.Tensor: The calculated loss
+ """
+ # The default value of ignore_index is the same as F.cross_entropy
+ ignore_index = -100 if ignore_index is None else ignore_index
+ # element-wise losses
+ loss = F.cross_entropy(
+ pred,
+ label,
+ weight=class_weight,
+ reduction='none',
+ ignore_index=ignore_index)
+
+ # average loss over non-ignored elements
+ # pytorch's official cross_entropy average loss over non-ignored elements
+ # refer to https://github.com/pytorch/pytorch/blob/56b43f4fec1f76953f15a627694d4bba34588969/torch/nn/functional.py#L2660 # noqa
+ if (avg_factor is None) and avg_non_ignore and reduction == 'mean':
+ avg_factor = label.numel() - (label == ignore_index).sum().item()
+
+ # apply weights and do the reduction
+ if weight is not None:
+ weight = weight.float()
+ loss = weight_reduce_loss(
+ loss, weight=weight, reduction=reduction, avg_factor=avg_factor)
+
+ return loss
+
+
+def _expand_onehot_labels(labels, label_weights, label_channels, ignore_index):
+ """Expand onehot labels to match the size of prediction."""
+ bin_labels = labels.new_full((labels.size(0), label_channels), 0)
+ valid_mask = (labels >= 0) & (labels != ignore_index)
+ inds = torch.nonzero(
+ valid_mask & (labels < label_channels), as_tuple=False)
+
+ if inds.numel() > 0:
+ bin_labels[inds, labels[inds]] = 1
+
+ valid_mask = valid_mask.view(-1, 1).expand(labels.size(0),
+ label_channels).float()
+ if label_weights is None:
+ bin_label_weights = valid_mask
+ else:
+ bin_label_weights = label_weights.view(-1, 1).repeat(1, label_channels)
+ bin_label_weights *= valid_mask
+
+ return bin_labels, bin_label_weights, valid_mask
+
+
+def binary_cross_entropy(pred,
+ label,
+ weight=None,
+ reduction='mean',
+ avg_factor=None,
+ class_weight=None,
+ ignore_index=-100,
+ avg_non_ignore=False):
+ """Calculate the binary CrossEntropy loss.
+
+ Args:
+ pred (torch.Tensor): The prediction with shape (N, 1) or (N, ).
+ When the shape of pred is (N, 1), label will be expanded to
+ one-hot format, and when the shape of pred is (N, ), label
+ will not be expanded to one-hot format.
+ label (torch.Tensor): The learning label of the prediction,
+ with shape (N, ).
+ weight (torch.Tensor, optional): Sample-wise loss weight.
+ reduction (str, optional): The method used to reduce the loss.
+ Options are "none", "mean" and "sum".
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ class_weight (list[float], optional): The weight for each class.
+ ignore_index (int | None): The label index to be ignored.
+ If None, it will be set to default value. Default: -100.
+ avg_non_ignore (bool): The flag decides to whether the loss is
+ only averaged over non-ignored targets. Default: False.
+
+ Returns:
+ torch.Tensor: The calculated loss.
+ """
+ # The default value of ignore_index is the same as F.cross_entropy
+ ignore_index = -100 if ignore_index is None else ignore_index
+
+ if pred.dim() != label.dim():
+ label, weight, valid_mask = _expand_onehot_labels(
+ label, weight, pred.size(-1), ignore_index)
+ else:
+ # should mask out the ignored elements
+ valid_mask = ((label >= 0) & (label != ignore_index)).float()
+ if weight is not None:
+ # The inplace writing method will have a mismatched broadcast
+ # shape error if the weight and valid_mask dimensions
+ # are inconsistent such as (B,N,1) and (B,N,C).
+ weight = weight * valid_mask
+ else:
+ weight = valid_mask
+
+ # average loss over non-ignored elements
+ if (avg_factor is None) and avg_non_ignore and reduction == 'mean':
+ avg_factor = valid_mask.sum().item()
+
+ # weighted element-wise losses
+ weight = weight.float()
+ loss = F.binary_cross_entropy_with_logits(
+ pred, label.float(), pos_weight=class_weight, reduction='none')
+ # do the reduction for the weighted loss
+ loss = weight_reduce_loss(
+ loss, weight, reduction=reduction, avg_factor=avg_factor)
+
+ return loss
+
+
+def mask_cross_entropy(pred,
+ target,
+ label,
+ reduction='mean',
+ avg_factor=None,
+ class_weight=None,
+ ignore_index=None,
+ **kwargs):
+ """Calculate the CrossEntropy loss for masks.
+
+ Args:
+ pred (torch.Tensor): The prediction with shape (N, C, *), C is the
+ number of classes. The trailing * indicates arbitrary shape.
+ target (torch.Tensor): The learning label of the prediction.
+ label (torch.Tensor): ``label`` indicates the class label of the mask
+ corresponding object. This will be used to select the mask in the
+ of the class which the object belongs to when the mask prediction
+ if not class-agnostic.
+ reduction (str, optional): The method used to reduce the loss.
+ Options are "none", "mean" and "sum".
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ class_weight (list[float], optional): The weight for each class.
+ ignore_index (None): Placeholder, to be consistent with other loss.
+ Default: None.
+
+ Returns:
+ torch.Tensor: The calculated loss
+
+ Example:
+ >>> N, C = 3, 11
+ >>> H, W = 2, 2
+ >>> pred = torch.randn(N, C, H, W) * 1000
+ >>> target = torch.rand(N, H, W)
+ >>> label = torch.randint(0, C, size=(N,))
+ >>> reduction = 'mean'
+ >>> avg_factor = None
+ >>> class_weights = None
+ >>> loss = mask_cross_entropy(pred, target, label, reduction,
+ >>> avg_factor, class_weights)
+ >>> assert loss.shape == (1,)
+ """
+ assert ignore_index is None, 'BCE loss does not support ignore_index'
+ # TODO: handle these two reserved arguments
+ assert reduction == 'mean' and avg_factor is None
+ num_rois = pred.size()[0]
+ inds = torch.arange(0, num_rois, dtype=torch.long, device=pred.device)
+ pred_slice = pred[inds, label].squeeze(1)
+ return F.binary_cross_entropy_with_logits(
+ pred_slice, target, weight=class_weight, reduction='mean')[None]
+
+
+@MODELS.register_module()
+class CrossEntropyLoss(nn.Module):
+
+ def __init__(self,
+ use_sigmoid=False,
+ use_mask=False,
+ reduction='mean',
+ class_weight=None,
+ ignore_index=None,
+ loss_weight=1.0,
+ avg_non_ignore=False):
+ """CrossEntropyLoss.
+
+ Args:
+ use_sigmoid (bool, optional): Whether the prediction uses sigmoid
+ of softmax. Defaults to False.
+ use_mask (bool, optional): Whether to use mask cross entropy loss.
+ Defaults to False.
+ reduction (str, optional): . Defaults to 'mean'.
+ Options are "none", "mean" and "sum".
+ class_weight (list[float], optional): Weight of each class.
+ Defaults to None.
+ ignore_index (int | None): The label index to be ignored.
+ Defaults to None.
+ loss_weight (float, optional): Weight of the loss. Defaults to 1.0.
+ avg_non_ignore (bool): The flag decides to whether the loss is
+ only averaged over non-ignored targets. Default: False.
+ """
+ super(CrossEntropyLoss, self).__init__()
+ assert (use_sigmoid is False) or (use_mask is False)
+ self.use_sigmoid = use_sigmoid
+ self.use_mask = use_mask
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+ self.class_weight = class_weight
+ self.ignore_index = ignore_index
+ self.avg_non_ignore = avg_non_ignore
+ if ((ignore_index is not None) and not self.avg_non_ignore
+ and self.reduction == 'mean'):
+ warnings.warn(
+ 'Default ``avg_non_ignore`` is False, if you would like to '
+ 'ignore the certain label and average loss over non-ignore '
+ 'labels, which is the same with PyTorch official '
+ 'cross_entropy, set ``avg_non_ignore=True``.')
+
+ if self.use_sigmoid:
+ self.cls_criterion = binary_cross_entropy
+ elif self.use_mask:
+ self.cls_criterion = mask_cross_entropy
+ else:
+ self.cls_criterion = cross_entropy
+
+ def extra_repr(self):
+ """Extra repr."""
+ s = f'avg_non_ignore={self.avg_non_ignore}'
+ return s
+
+ def forward(self,
+ cls_score,
+ label,
+ weight=None,
+ avg_factor=None,
+ reduction_override=None,
+ ignore_index=None,
+ **kwargs):
+ """Forward function.
+
+ Args:
+ cls_score (torch.Tensor): The prediction.
+ label (torch.Tensor): The learning label of the prediction.
+ weight (torch.Tensor, optional): Sample-wise loss weight.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ reduction_override (str, optional): The method used to reduce the
+ loss. Options are "none", "mean" and "sum".
+ ignore_index (int | None): The label index to be ignored.
+ If not None, it will override the default value. Default: None.
+ Returns:
+ torch.Tensor: The calculated loss.
+ """
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ if ignore_index is None:
+ ignore_index = self.ignore_index
+
+ if self.class_weight is not None:
+ class_weight = cls_score.new_tensor(
+ self.class_weight, device=cls_score.device)
+ else:
+ class_weight = None
+ loss_cls = self.loss_weight * self.cls_criterion(
+ cls_score,
+ label,
+ weight,
+ class_weight=class_weight,
+ reduction=reduction,
+ avg_factor=avg_factor,
+ ignore_index=ignore_index,
+ avg_non_ignore=self.avg_non_ignore,
+ **kwargs)
+ return loss_cls
+
+
+@MODELS.register_module()
+class CrossEntropyCustomLoss(CrossEntropyLoss):
+
+ def __init__(self,
+ use_sigmoid=False,
+ use_mask=False,
+ reduction='mean',
+ num_classes=-1,
+ class_weight=None,
+ ignore_index=None,
+ loss_weight=1.0,
+ avg_non_ignore=False):
+ """CrossEntropyCustomLoss.
+
+ Args:
+ use_sigmoid (bool, optional): Whether the prediction uses sigmoid
+ of softmax. Defaults to False.
+ use_mask (bool, optional): Whether to use mask cross entropy loss.
+ Defaults to False.
+ reduction (str, optional): . Defaults to 'mean'.
+ Options are "none", "mean" and "sum".
+ num_classes (int): Number of classes to classify.
+ class_weight (list[float], optional): Weight of each class.
+ Defaults to None.
+ ignore_index (int | None): The label index to be ignored.
+ Defaults to None.
+ loss_weight (float, optional): Weight of the loss. Defaults to 1.0.
+ avg_non_ignore (bool): The flag decides to whether the loss is
+ only averaged over non-ignored targets. Default: False.
+ """
+ super(CrossEntropyCustomLoss, self).__init__()
+ assert (use_sigmoid is False) or (use_mask is False)
+ self.use_sigmoid = use_sigmoid
+ self.use_mask = use_mask
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+ self.class_weight = class_weight
+ self.ignore_index = ignore_index
+ self.avg_non_ignore = avg_non_ignore
+ if ((ignore_index is not None) and not self.avg_non_ignore
+ and self.reduction == 'mean'):
+ warnings.warn(
+ 'Default ``avg_non_ignore`` is False, if you would like to '
+ 'ignore the certain label and average loss over non-ignore '
+ 'labels, which is the same with PyTorch official '
+ 'cross_entropy, set ``avg_non_ignore=True``.')
+
+ if self.use_sigmoid:
+ self.cls_criterion = binary_cross_entropy
+ elif self.use_mask:
+ self.cls_criterion = mask_cross_entropy
+ else:
+ self.cls_criterion = cross_entropy
+
+ self.num_classes = num_classes
+
+ assert self.num_classes != -1
+
+ # custom output channels of the classifier
+ self.custom_cls_channels = True
+ # custom activation of cls_score
+ self.custom_activation = True
+ # custom accuracy of the classsifier
+ self.custom_accuracy = True
+
+ def get_cls_channels(self, num_classes):
+ assert num_classes == self.num_classes
+ if not self.use_sigmoid:
+ return num_classes + 1
+ else:
+ return num_classes
+
+ def get_activation(self, cls_score):
+
+ fine_cls_score = cls_score[:, :self.num_classes]
+
+ if not self.use_sigmoid:
+ bg_score = cls_score[:, [-1]]
+ new_score = torch.cat([fine_cls_score, bg_score], dim=-1)
+ scores = F.softmax(new_score, dim=-1)
+ else:
+ score_classes = fine_cls_score.sigmoid()
+ score_neg = 1 - score_classes.sum(dim=1, keepdim=True)
+ score_neg = score_neg.clamp(min=0, max=1)
+ scores = torch.cat([score_classes, score_neg], dim=1)
+
+ return scores
+
+ def get_accuracy(self, cls_score, labels):
+
+ fine_cls_score = cls_score[:, :self.num_classes]
+
+ pos_inds = labels < self.num_classes
+ acc_classes = accuracy(fine_cls_score[pos_inds], labels[pos_inds])
+ acc = dict()
+ acc['acc_classes'] = acc_classes
+ return acc
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/ddq_detr_aux_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/ddq_detr_aux_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..41f1c7166e6c7d05c5414cd04ad3eb3cd467f1b6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/ddq_detr_aux_loss.py
@@ -0,0 +1,303 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+import torch.nn as nn
+from mmengine.structures import BaseDataElement
+
+from mmdet.models.utils import multi_apply
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.utils import reduce_mean
+
+
+class DDQAuxLoss(nn.Module):
+ """DDQ auxiliary branches loss for dense queries.
+
+ Args:
+ loss_cls (dict):
+ Configuration of classification loss function.
+ loss_bbox (dict):
+ Configuration of bbox regression loss function.
+ train_cfg (dict):
+ Configuration of gt targets assigner for each predicted bbox.
+ """
+
+ def __init__(
+ self,
+ loss_cls=dict(
+ type='QualityFocalLoss',
+ use_sigmoid=True,
+ activated=True, # use probability instead of logit as input
+ beta=2.0,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=2.0),
+ train_cfg=dict(
+ assigner=dict(type='TopkHungarianAssigner', topk=8),
+ alpha=1,
+ beta=6),
+ ):
+ super(DDQAuxLoss, self).__init__()
+ self.train_cfg = train_cfg
+ self.loss_cls = MODELS.build(loss_cls)
+ self.loss_bbox = MODELS.build(loss_bbox)
+ self.assigner = TASK_UTILS.build(self.train_cfg['assigner'])
+
+ sampler_cfg = dict(type='PseudoSampler')
+ self.sampler = TASK_UTILS.build(sampler_cfg)
+
+ def loss_single(self, cls_score, bbox_pred, labels, label_weights,
+ bbox_targets, alignment_metrics):
+ """Calculate auxiliary branches loss for dense queries for one image.
+
+ Args:
+ cls_score (Tensor): Predicted normalized classification
+ scores for one image, has shape (num_dense_queries,
+ cls_out_channels).
+ bbox_pred (Tensor): Predicted unnormalized bbox coordinates
+ for one image, has shape (num_dense_queries, 4) with the
+ last dimension arranged as (x1, y1, x2, y2).
+ labels (Tensor): Labels for one image.
+ label_weights (Tensor): Label weights for one image.
+ bbox_targets (Tensor): Bbox targets for one image.
+ alignment_metrics (Tensor): Normalized alignment metrics for one
+ image.
+
+ Returns:
+ tuple: A tuple of loss components and loss weights.
+ """
+ bbox_targets = bbox_targets.reshape(-1, 4)
+ labels = labels.reshape(-1)
+ alignment_metrics = alignment_metrics.reshape(-1)
+ label_weights = label_weights.reshape(-1)
+ targets = (labels, alignment_metrics)
+ cls_loss_func = self.loss_cls
+
+ loss_cls = cls_loss_func(
+ cls_score, targets, label_weights, avg_factor=1.0)
+
+ # FG cat_id: [0, num_classes -1], BG cat_id: num_classes
+ bg_class_ind = cls_score.size(-1)
+ pos_inds = ((labels >= 0)
+ & (labels < bg_class_ind)).nonzero().squeeze(1)
+
+ if len(pos_inds) > 0:
+ pos_bbox_targets = bbox_targets[pos_inds]
+ pos_bbox_pred = bbox_pred[pos_inds]
+
+ pos_decode_bbox_pred = pos_bbox_pred
+ pos_decode_bbox_targets = pos_bbox_targets
+
+ # regression loss
+ pos_bbox_weight = alignment_metrics[pos_inds]
+
+ loss_bbox = self.loss_bbox(
+ pos_decode_bbox_pred,
+ pos_decode_bbox_targets,
+ weight=pos_bbox_weight,
+ avg_factor=1.0)
+ else:
+ loss_bbox = bbox_pred.sum() * 0
+ pos_bbox_weight = bbox_targets.new_tensor(0.)
+
+ return loss_cls, loss_bbox, alignment_metrics.sum(
+ ), pos_bbox_weight.sum()
+
+ def loss(self, cls_scores, bbox_preds, gt_bboxes, gt_labels, img_metas,
+ **kwargs):
+ """Calculate auxiliary branches loss for dense queries.
+
+ Args:
+ cls_scores (Tensor): Predicted normalized classification
+ scores, has shape (bs, num_dense_queries,
+ cls_out_channels).
+ bbox_preds (Tensor): Predicted unnormalized bbox coordinates,
+ has shape (bs, num_dense_queries, 4) with the last
+ dimension arranged as (x1, y1, x2, y2).
+ gt_bboxes (list[Tensor]): List of unnormalized ground truth
+ bboxes for each image, each has shape (num_gt, 4) with the
+ last dimension arranged as (x1, y1, x2, y2).
+ NOTE: num_gt is dynamic for each image.
+ gt_labels (list[Tensor]): List of ground truth classification
+ index for each image, each has shape (num_gt,).
+ NOTE: num_gt is dynamic for each image.
+ img_metas (list[dict]): Meta information for one image,
+ e.g., image size, scaling factor, etc.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ flatten_cls_scores = cls_scores
+ flatten_bbox_preds = bbox_preds
+
+ cls_reg_targets = self.get_targets(
+ flatten_cls_scores,
+ flatten_bbox_preds,
+ gt_bboxes,
+ img_metas,
+ gt_labels_list=gt_labels,
+ )
+ (labels_list, label_weights_list, bbox_targets_list,
+ alignment_metrics_list) = cls_reg_targets
+
+ losses_cls, losses_bbox, \
+ cls_avg_factors, bbox_avg_factors = multi_apply(
+ self.loss_single,
+ flatten_cls_scores,
+ flatten_bbox_preds,
+ labels_list,
+ label_weights_list,
+ bbox_targets_list,
+ alignment_metrics_list,
+ )
+
+ cls_avg_factor = reduce_mean(sum(cls_avg_factors)).clamp_(min=1).item()
+ losses_cls = list(map(lambda x: x / cls_avg_factor, losses_cls))
+
+ bbox_avg_factor = reduce_mean(
+ sum(bbox_avg_factors)).clamp_(min=1).item()
+ losses_bbox = list(map(lambda x: x / bbox_avg_factor, losses_bbox))
+ return dict(aux_loss_cls=losses_cls, aux_loss_bbox=losses_bbox)
+
+ def get_targets(self,
+ cls_scores,
+ bbox_preds,
+ gt_bboxes_list,
+ img_metas,
+ gt_labels_list=None,
+ **kwargs):
+ """Compute regression and classification targets for a batch images.
+
+ Args:
+ cls_scores (Tensor): Predicted normalized classification
+ scores, has shape (bs, num_dense_queries,
+ cls_out_channels).
+ bbox_preds (Tensor): Predicted unnormalized bbox coordinates,
+ has shape (bs, num_dense_queries, 4) with the last
+ dimension arranged as (x1, y1, x2, y2).
+ gt_bboxes_list (List[Tensor]): List of unnormalized ground truth
+ bboxes for each image, each has shape (num_gt, 4) with the
+ last dimension arranged as (x1, y1, x2, y2).
+ NOTE: num_gt is dynamic for each image.
+ img_metas (list[dict]): Meta information for one image,
+ e.g., image size, scaling factor, etc.
+ gt_labels_list (list[Tensor]): List of ground truth classification
+ index for each image, each has shape (num_gt,).
+ NOTE: num_gt is dynamic for each image.
+ Default: None.
+
+ Returns:
+ tuple: a tuple containing the following targets.
+
+ - all_labels (list[Tensor]): Labels for all images.
+ - all_label_weights (list[Tensor]): Label weights for all images.
+ - all_bbox_targets (list[Tensor]): Bbox targets for all images.
+ - all_assign_metrics (list[Tensor]): Normalized alignment metrics
+ for all images.
+ """
+ (all_labels, all_label_weights, all_bbox_targets,
+ all_assign_metrics) = multi_apply(self._get_target_single, cls_scores,
+ bbox_preds, gt_bboxes_list,
+ gt_labels_list, img_metas)
+
+ return (all_labels, all_label_weights, all_bbox_targets,
+ all_assign_metrics)
+
+ def _get_target_single(self, cls_scores, bbox_preds, gt_bboxes, gt_labels,
+ img_meta, **kwargs):
+ """Compute regression and classification targets for one image.
+
+ Args:
+ cls_scores (Tensor): Predicted normalized classification
+ scores for one image, has shape (num_dense_queries,
+ cls_out_channels).
+ bbox_preds (Tensor): Predicted unnormalized bbox coordinates
+ for one image, has shape (num_dense_queries, 4) with the
+ last dimension arranged as (x1, y1, x2, y2).
+ gt_bboxes (Tensor): Unnormalized ground truth
+ bboxes for one image, has shape (num_gt, 4) with the
+ last dimension arranged as (x1, y1, x2, y2).
+ NOTE: num_gt is dynamic for each image.
+ gt_labels (Tensor): Ground truth classification
+ index for the image, has shape (num_gt,).
+ NOTE: num_gt is dynamic for each image.
+ img_meta (dict): Meta information for one image.
+
+ Returns:
+ tuple[Tensor]: a tuple containing the following for one image.
+
+ - labels (Tensor): Labels for one image.
+ - label_weights (Tensor): Label weights for one image.
+ - bbox_targets (Tensor): Bbox targets for one image.
+ - norm_alignment_metrics (Tensor): Normalized alignment
+ metrics for one image.
+ """
+ if len(gt_labels) == 0:
+ num_valid_anchors = len(cls_scores)
+ bbox_targets = torch.zeros_like(bbox_preds)
+ labels = bbox_preds.new_full((num_valid_anchors, ),
+ cls_scores.size(-1),
+ dtype=torch.long)
+ label_weights = bbox_preds.new_zeros(
+ num_valid_anchors, dtype=torch.float)
+ norm_alignment_metrics = bbox_preds.new_zeros(
+ num_valid_anchors, dtype=torch.float)
+ return (labels, label_weights, bbox_targets,
+ norm_alignment_metrics)
+
+ assign_result = self.assigner.assign(cls_scores, bbox_preds, gt_bboxes,
+ gt_labels, img_meta)
+ assign_ious = assign_result.max_overlaps
+ assign_metrics = assign_result.assign_metrics
+
+ pred_instances = BaseDataElement()
+ gt_instances = BaseDataElement()
+
+ pred_instances.bboxes = bbox_preds
+ gt_instances.bboxes = gt_bboxes
+
+ pred_instances.priors = cls_scores
+ gt_instances.labels = gt_labels
+
+ sampling_result = self.sampler.sample(assign_result, pred_instances,
+ gt_instances)
+
+ num_valid_anchors = len(cls_scores)
+ bbox_targets = torch.zeros_like(bbox_preds)
+ labels = bbox_preds.new_full((num_valid_anchors, ),
+ cls_scores.size(-1),
+ dtype=torch.long)
+ label_weights = bbox_preds.new_zeros(
+ num_valid_anchors, dtype=torch.float)
+ norm_alignment_metrics = bbox_preds.new_zeros(
+ num_valid_anchors, dtype=torch.float)
+
+ pos_inds = sampling_result.pos_inds
+ neg_inds = sampling_result.neg_inds
+ if len(pos_inds) > 0:
+ # point-based
+ pos_bbox_targets = sampling_result.pos_gt_bboxes
+ bbox_targets[pos_inds, :] = pos_bbox_targets
+
+ if gt_labels is None:
+ # Only dense_heads gives gt_labels as None
+ # Foreground is the first class since v2.5.0
+ labels[pos_inds] = 0
+ else:
+ labels[pos_inds] = gt_labels[
+ sampling_result.pos_assigned_gt_inds]
+
+ label_weights[pos_inds] = 1.0
+
+ if len(neg_inds) > 0:
+ label_weights[neg_inds] = 1.0
+
+ class_assigned_gt_inds = torch.unique(
+ sampling_result.pos_assigned_gt_inds)
+ for gt_inds in class_assigned_gt_inds:
+ gt_class_inds = sampling_result.pos_assigned_gt_inds == gt_inds
+ pos_alignment_metrics = assign_metrics[gt_class_inds]
+ pos_ious = assign_ious[gt_class_inds]
+ pos_norm_alignment_metrics = pos_alignment_metrics / (
+ pos_alignment_metrics.max() + 10e-8) * pos_ious.max()
+ norm_alignment_metrics[
+ pos_inds[gt_class_inds]] = pos_norm_alignment_metrics
+
+ return (labels, label_weights, bbox_targets, norm_alignment_metrics)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/dice_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/dice_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..1d5cac1e9710a6a72fe0401db22b8b72cfe058f9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/dice_loss.py
@@ -0,0 +1,146 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+import torch.nn as nn
+
+from mmdet.registry import MODELS
+from .utils import weight_reduce_loss
+
+
+def dice_loss(pred,
+ target,
+ weight=None,
+ eps=1e-3,
+ reduction='mean',
+ naive_dice=False,
+ avg_factor=None):
+ """Calculate dice loss, there are two forms of dice loss is supported:
+
+ - the one proposed in `V-Net: Fully Convolutional Neural
+ Networks for Volumetric Medical Image Segmentation
+ `_.
+ - the dice loss in which the power of the number in the
+ denominator is the first power instead of the second
+ power.
+
+ Args:
+ pred (torch.Tensor): The prediction, has a shape (n, *)
+ target (torch.Tensor): The learning label of the prediction,
+ shape (n, *), same shape of pred.
+ weight (torch.Tensor, optional): The weight of loss for each
+ prediction, has a shape (n,). Defaults to None.
+ eps (float): Avoid dividing by zero. Default: 1e-3.
+ reduction (str, optional): The method used to reduce the loss into
+ a scalar. Defaults to 'mean'.
+ Options are "none", "mean" and "sum".
+ naive_dice (bool, optional): If false, use the dice
+ loss defined in the V-Net paper, otherwise, use the
+ naive dice loss in which the power of the number in the
+ denominator is the first power instead of the second
+ power.Defaults to False.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ """
+
+ input = pred.flatten(1)
+ target = target.flatten(1).float()
+
+ a = torch.sum(input * target, 1)
+ if naive_dice:
+ b = torch.sum(input, 1)
+ c = torch.sum(target, 1)
+ d = (2 * a + eps) / (b + c + eps)
+ else:
+ b = torch.sum(input * input, 1) + eps
+ c = torch.sum(target * target, 1) + eps
+ d = (2 * a) / (b + c)
+
+ loss = 1 - d
+ if weight is not None:
+ assert weight.ndim == loss.ndim
+ assert len(weight) == len(pred)
+ loss = weight_reduce_loss(loss, weight, reduction, avg_factor)
+ return loss
+
+
+@MODELS.register_module()
+class DiceLoss(nn.Module):
+
+ def __init__(self,
+ use_sigmoid=True,
+ activate=True,
+ reduction='mean',
+ naive_dice=False,
+ loss_weight=1.0,
+ eps=1e-3):
+ """Compute dice loss.
+
+ Args:
+ use_sigmoid (bool, optional): Whether to the prediction is
+ used for sigmoid or softmax. Defaults to True.
+ activate (bool): Whether to activate the predictions inside,
+ this will disable the inside sigmoid operation.
+ Defaults to True.
+ reduction (str, optional): The method used
+ to reduce the loss. Options are "none",
+ "mean" and "sum". Defaults to 'mean'.
+ naive_dice (bool, optional): If false, use the dice
+ loss defined in the V-Net paper, otherwise, use the
+ naive dice loss in which the power of the number in the
+ denominator is the first power instead of the second
+ power. Defaults to False.
+ loss_weight (float, optional): Weight of loss. Defaults to 1.0.
+ eps (float): Avoid dividing by zero. Defaults to 1e-3.
+ """
+
+ super(DiceLoss, self).__init__()
+ self.use_sigmoid = use_sigmoid
+ self.reduction = reduction
+ self.naive_dice = naive_dice
+ self.loss_weight = loss_weight
+ self.eps = eps
+ self.activate = activate
+
+ def forward(self,
+ pred,
+ target,
+ weight=None,
+ reduction_override=None,
+ avg_factor=None):
+ """Forward function.
+
+ Args:
+ pred (torch.Tensor): The prediction, has a shape (n, *).
+ target (torch.Tensor): The label of the prediction,
+ shape (n, *), same shape of pred.
+ weight (torch.Tensor, optional): The weight of loss for each
+ prediction, has a shape (n,). Defaults to None.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Options are "none", "mean" and "sum".
+
+ Returns:
+ torch.Tensor: The calculated loss
+ """
+
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+
+ if self.activate:
+ if self.use_sigmoid:
+ pred = pred.sigmoid()
+ else:
+ raise NotImplementedError
+
+ loss = self.loss_weight * dice_loss(
+ pred,
+ target,
+ weight,
+ eps=self.eps,
+ reduction=reduction,
+ naive_dice=self.naive_dice,
+ avg_factor=avg_factor)
+
+ return loss
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/eqlv2_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/eqlv2_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..ea1f4a9a8f7c71119c2bed743d714a34ab4db82c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/eqlv2_loss.py
@@ -0,0 +1,173 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import logging
+from functools import partial
+from typing import Optional
+
+import torch
+import torch.distributed as dist
+import torch.nn as nn
+import torch.nn.functional as F
+from mmengine.logging import print_log
+from torch import Tensor
+
+from mmdet.registry import MODELS
+
+
+@MODELS.register_module()
+class EQLV2Loss(nn.Module):
+
+ def __init__(self,
+ use_sigmoid: bool = True,
+ reduction: str = 'mean',
+ class_weight: Optional[Tensor] = None,
+ loss_weight: float = 1.0,
+ num_classes: int = 1203,
+ use_distributed: bool = False,
+ mu: float = 0.8,
+ alpha: float = 4.0,
+ gamma: int = 12,
+ vis_grad: bool = False,
+ test_with_obj: bool = True) -> None:
+ """`Equalization Loss v2 `_
+
+ Args:
+ use_sigmoid (bool): EQLv2 uses the sigmoid function to transform
+ the predicted logits to an estimated probability distribution.
+ reduction (str, optional): The method used to reduce the loss into
+ a scalar. Defaults to 'mean'.
+ class_weight (Tensor, optional): The weight of loss for each
+ prediction. Defaults to None.
+ loss_weight (float, optional): The weight of the total EQLv2 loss.
+ Defaults to 1.0.
+ num_classes (int): 1203 for lvis v1.0, 1230 for lvis v0.5.
+ use_distributed (bool, float): EQLv2 will calculate the gradients
+ on all GPUs if there is any. Change to True if you are using
+ distributed training. Default to False.
+ mu (float, optional): Defaults to 0.8
+ alpha (float, optional): A balance factor for the negative part of
+ EQLV2 Loss. Defaults to 4.0.
+ gamma (int, optional): The gamma for calculating the modulating
+ factor. Defaults to 12.
+ vis_grad (bool, optional): Default to False.
+ test_with_obj (bool, optional): Default to True.
+
+ Returns:
+ None.
+ """
+ super().__init__()
+ self.use_sigmoid = True
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+ self.class_weight = class_weight
+ self.num_classes = num_classes
+ self.group = True
+
+ # cfg for eqlv2
+ self.vis_grad = vis_grad
+ self.mu = mu
+ self.alpha = alpha
+ self.gamma = gamma
+ self.use_distributed = use_distributed
+
+ # initial variables
+ self.register_buffer('pos_grad', torch.zeros(self.num_classes))
+ self.register_buffer('neg_grad', torch.zeros(self.num_classes))
+ # At the beginning of training, we set a high value (eg. 100)
+ # for the initial gradient ratio so that the weight for pos
+ # gradients and neg gradients are 1.
+ self.register_buffer('pos_neg', torch.ones(self.num_classes) * 100)
+
+ self.test_with_obj = test_with_obj
+
+ def _func(x, gamma, mu):
+ return 1 / (1 + torch.exp(-gamma * (x - mu)))
+
+ self.map_func = partial(_func, gamma=self.gamma, mu=self.mu)
+
+ print_log(
+ f'build EQL v2, gamma: {gamma}, mu: {mu}, alpha: {alpha}',
+ logger='current',
+ level=logging.DEBUG)
+
+ def forward(self,
+ cls_score: Tensor,
+ label: Tensor,
+ weight: Optional[Tensor] = None,
+ avg_factor: Optional[int] = None,
+ reduction_override: Optional[Tensor] = None) -> Tensor:
+ """`Equalization Loss v2 `_
+
+ Args:
+ cls_score (Tensor): The prediction with shape (N, C), C is the
+ number of classes.
+ label (Tensor): The ground truth label of the predicted target with
+ shape (N, C), C is the number of classes.
+ weight (Tensor, optional): The weight of loss for each prediction.
+ Defaults to None.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Options are "none", "mean" and "sum".
+
+ Returns:
+ Tensor: The calculated loss
+ """
+ self.n_i, self.n_c = cls_score.size()
+ self.gt_classes = label
+ self.pred_class_logits = cls_score
+
+ def expand_label(pred, gt_classes):
+ target = pred.new_zeros(self.n_i, self.n_c)
+ target[torch.arange(self.n_i), gt_classes] = 1
+ return target
+
+ target = expand_label(cls_score, label)
+
+ pos_w, neg_w = self.get_weight(cls_score)
+
+ weight = pos_w * target + neg_w * (1 - target)
+
+ cls_loss = F.binary_cross_entropy_with_logits(
+ cls_score, target, reduction='none')
+ cls_loss = torch.sum(cls_loss * weight) / self.n_i
+
+ self.collect_grad(cls_score.detach(), target.detach(), weight.detach())
+
+ return self.loss_weight * cls_loss
+
+ def get_channel_num(self, num_classes):
+ num_channel = num_classes + 1
+ return num_channel
+
+ def get_activation(self, pred):
+ pred = torch.sigmoid(pred)
+ n_i, n_c = pred.size()
+ bg_score = pred[:, -1].view(n_i, 1)
+ if self.test_with_obj:
+ pred[:, :-1] *= (1 - bg_score)
+ return pred
+
+ def collect_grad(self, pred, target, weight):
+ prob = torch.sigmoid(pred)
+ grad = target * (prob - 1) + (1 - target) * prob
+ grad = torch.abs(grad)
+
+ # do not collect grad for objectiveness branch [:-1]
+ pos_grad = torch.sum(grad * target * weight, dim=0)[:-1]
+ neg_grad = torch.sum(grad * (1 - target) * weight, dim=0)[:-1]
+
+ if self.use_distributed:
+ dist.all_reduce(pos_grad)
+ dist.all_reduce(neg_grad)
+
+ self.pos_grad += pos_grad
+ self.neg_grad += neg_grad
+ self.pos_neg = self.pos_grad / (self.neg_grad + 1e-10)
+
+ def get_weight(self, pred):
+ neg_w = torch.cat([self.map_func(self.pos_neg), pred.new_ones(1)])
+ pos_w = 1 + self.alpha * (1 - neg_w)
+ neg_w = neg_w.view(1, -1).expand(self.n_i, self.n_c)
+ pos_w = pos_w.view(1, -1).expand(self.n_i, self.n_c)
+ return pos_w, neg_w
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/focal_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/focal_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..15bef293a591a7f4c099febdaa82abaf7fb4928a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/focal_loss.py
@@ -0,0 +1,371 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.ops import sigmoid_focal_loss as _sigmoid_focal_loss
+
+from mmdet.registry import MODELS
+from .accuracy import accuracy
+from .utils import weight_reduce_loss
+
+
+# This method is only for debugging
+def py_sigmoid_focal_loss(pred,
+ target,
+ weight=None,
+ gamma=2.0,
+ alpha=0.25,
+ reduction='mean',
+ avg_factor=None):
+ """PyTorch version of `Focal Loss `_.
+
+ Args:
+ pred (torch.Tensor): The prediction with shape (N, C), C is the
+ number of classes
+ target (torch.Tensor): The learning label of the prediction.
+ weight (torch.Tensor, optional): Sample-wise loss weight.
+ gamma (float, optional): The gamma for calculating the modulating
+ factor. Defaults to 2.0.
+ alpha (float, optional): A balanced form for Focal Loss.
+ Defaults to 0.25.
+ reduction (str, optional): The method used to reduce the loss into
+ a scalar. Defaults to 'mean'.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ """
+ pred_sigmoid = pred.sigmoid()
+ target = target.type_as(pred)
+ # Actually, pt here denotes (1 - pt) in the Focal Loss paper
+ pt = (1 - pred_sigmoid) * target + pred_sigmoid * (1 - target)
+ # Thus it's pt.pow(gamma) rather than (1 - pt).pow(gamma)
+ focal_weight = (alpha * target + (1 - alpha) *
+ (1 - target)) * pt.pow(gamma)
+ loss = F.binary_cross_entropy_with_logits(
+ pred, target, reduction='none') * focal_weight
+ if weight is not None:
+ if weight.shape != loss.shape:
+ if weight.size(0) == loss.size(0):
+ # For most cases, weight is of shape (num_priors, ),
+ # which means it does not have the second axis num_class
+ weight = weight.view(-1, 1)
+ else:
+ # Sometimes, weight per anchor per class is also needed. e.g.
+ # in FSAF. But it may be flattened of shape
+ # (num_priors x num_class, ), while loss is still of shape
+ # (num_priors, num_class).
+ assert weight.numel() == loss.numel()
+ weight = weight.view(loss.size(0), -1)
+ assert weight.ndim == loss.ndim
+ loss = weight_reduce_loss(loss, weight, reduction, avg_factor)
+ return loss
+
+
+def py_focal_loss_with_prob(pred,
+ target,
+ weight=None,
+ gamma=2.0,
+ alpha=0.25,
+ reduction='mean',
+ avg_factor=None):
+ """PyTorch version of `Focal Loss `_.
+ Different from `py_sigmoid_focal_loss`, this function accepts probability
+ as input.
+
+ Args:
+ pred (torch.Tensor): The prediction probability with shape (N, C),
+ C is the number of classes.
+ target (torch.Tensor): The learning label of the prediction.
+ The target shape support (N,C) or (N,), (N,C) means one-hot form.
+ weight (torch.Tensor, optional): Sample-wise loss weight.
+ gamma (float, optional): The gamma for calculating the modulating
+ factor. Defaults to 2.0.
+ alpha (float, optional): A balanced form for Focal Loss.
+ Defaults to 0.25.
+ reduction (str, optional): The method used to reduce the loss into
+ a scalar. Defaults to 'mean'.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ """
+ if pred.dim() != target.dim():
+ num_classes = pred.size(1)
+ target = F.one_hot(target, num_classes=num_classes + 1)
+ target = target[:, :num_classes]
+
+ target = target.type_as(pred)
+ pt = (1 - pred) * target + pred * (1 - target)
+ focal_weight = (alpha * target + (1 - alpha) *
+ (1 - target)) * pt.pow(gamma)
+ loss = F.binary_cross_entropy(
+ pred, target, reduction='none') * focal_weight
+ if weight is not None:
+ if weight.shape != loss.shape:
+ if weight.size(0) == loss.size(0):
+ # For most cases, weight is of shape (num_priors, ),
+ # which means it does not have the second axis num_class
+ weight = weight.view(-1, 1)
+ else:
+ # Sometimes, weight per anchor per class is also needed. e.g.
+ # in FSAF. But it may be flattened of shape
+ # (num_priors x num_class, ), while loss is still of shape
+ # (num_priors, num_class).
+ assert weight.numel() == loss.numel()
+ weight = weight.view(loss.size(0), -1)
+ assert weight.ndim == loss.ndim
+ loss = weight_reduce_loss(loss, weight, reduction, avg_factor)
+ return loss
+
+
+def sigmoid_focal_loss(pred,
+ target,
+ weight=None,
+ gamma=2.0,
+ alpha=0.25,
+ reduction='mean',
+ avg_factor=None):
+ r"""A wrapper of cuda version `Focal Loss
+ `_.
+
+ Args:
+ pred (torch.Tensor): The prediction with shape (N, C), C is the number
+ of classes.
+ target (torch.Tensor): The learning label of the prediction.
+ weight (torch.Tensor, optional): Sample-wise loss weight.
+ gamma (float, optional): The gamma for calculating the modulating
+ factor. Defaults to 2.0.
+ alpha (float, optional): A balanced form for Focal Loss.
+ Defaults to 0.25.
+ reduction (str, optional): The method used to reduce the loss into
+ a scalar. Defaults to 'mean'. Options are "none", "mean" and "sum".
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ """
+ # Function.apply does not accept keyword arguments, so the decorator
+ # "weighted_loss" is not applicable
+ loss = _sigmoid_focal_loss(pred.contiguous(), target.contiguous(), gamma,
+ alpha, None, 'none')
+ if weight is not None:
+ if weight.shape != loss.shape:
+ if weight.size(0) == loss.size(0):
+ # For most cases, weight is of shape (num_priors, ),
+ # which means it does not have the second axis num_class
+ weight = weight.view(-1, 1)
+ else:
+ # Sometimes, weight per anchor per class is also needed. e.g.
+ # in FSAF. But it may be flattened of shape
+ # (num_priors x num_class, ), while loss is still of shape
+ # (num_priors, num_class).
+ assert weight.numel() == loss.numel()
+ weight = weight.view(loss.size(0), -1)
+ assert weight.ndim == loss.ndim
+ loss = weight_reduce_loss(loss, weight, reduction, avg_factor)
+ return loss
+
+
+@MODELS.register_module()
+class FocalLoss(nn.Module):
+
+ def __init__(self,
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ reduction='mean',
+ loss_weight=1.0,
+ activated=False):
+ """`Focal Loss `_
+
+ Args:
+ use_sigmoid (bool, optional): Whether to the prediction is
+ used for sigmoid or softmax. Defaults to True.
+ gamma (float, optional): The gamma for calculating the modulating
+ factor. Defaults to 2.0.
+ alpha (float, optional): A balanced form for Focal Loss.
+ Defaults to 0.25.
+ reduction (str, optional): The method used to reduce the loss into
+ a scalar. Defaults to 'mean'. Options are "none", "mean" and
+ "sum".
+ loss_weight (float, optional): Weight of loss. Defaults to 1.0.
+ activated (bool, optional): Whether the input is activated.
+ If True, it means the input has been activated and can be
+ treated as probabilities. Else, it should be treated as logits.
+ Defaults to False.
+ """
+ super(FocalLoss, self).__init__()
+ assert use_sigmoid is True, 'Only sigmoid focal loss supported now.'
+ self.use_sigmoid = use_sigmoid
+ self.gamma = gamma
+ self.alpha = alpha
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+ self.activated = activated
+
+ def forward(self,
+ pred,
+ target,
+ weight=None,
+ avg_factor=None,
+ reduction_override=None):
+ """Forward function.
+
+ Args:
+ pred (torch.Tensor): The prediction.
+ target (torch.Tensor): The learning label of the prediction.
+ The target shape support (N,C) or (N,), (N,C) means
+ one-hot form.
+ weight (torch.Tensor, optional): The weight of loss for each
+ prediction. Defaults to None.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Options are "none", "mean" and "sum".
+
+ Returns:
+ torch.Tensor: The calculated loss
+ """
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ if self.use_sigmoid:
+ if self.activated:
+ calculate_loss_func = py_focal_loss_with_prob
+ else:
+ if pred.dim() == target.dim():
+ # this means that target is already in One-Hot form.
+ calculate_loss_func = py_sigmoid_focal_loss
+ elif torch.cuda.is_available() and pred.is_cuda:
+ calculate_loss_func = sigmoid_focal_loss
+ else:
+ num_classes = pred.size(1)
+ target = F.one_hot(target, num_classes=num_classes + 1)
+ target = target[:, :num_classes]
+ calculate_loss_func = py_sigmoid_focal_loss
+
+ loss_cls = self.loss_weight * calculate_loss_func(
+ pred,
+ target,
+ weight,
+ gamma=self.gamma,
+ alpha=self.alpha,
+ reduction=reduction,
+ avg_factor=avg_factor)
+
+ else:
+ raise NotImplementedError
+ return loss_cls
+
+
+@MODELS.register_module()
+class FocalCustomLoss(nn.Module):
+
+ def __init__(self,
+ use_sigmoid=True,
+ num_classes=-1,
+ gamma=2.0,
+ alpha=0.25,
+ reduction='mean',
+ loss_weight=1.0,
+ activated=False):
+ """`Focal Loss for V3Det `_
+
+ Args:
+ use_sigmoid (bool, optional): Whether to the prediction is
+ used for sigmoid or softmax. Defaults to True.
+ num_classes (int): Number of classes to classify.
+ gamma (float, optional): The gamma for calculating the modulating
+ factor. Defaults to 2.0.
+ alpha (float, optional): A balanced form for Focal Loss.
+ Defaults to 0.25.
+ reduction (str, optional): The method used to reduce the loss into
+ a scalar. Defaults to 'mean'. Options are "none", "mean" and
+ "sum".
+ loss_weight (float, optional): Weight of loss. Defaults to 1.0.
+ activated (bool, optional): Whether the input is activated.
+ If True, it means the input has been activated and can be
+ treated as probabilities. Else, it should be treated as logits.
+ Defaults to False.
+ """
+ super(FocalCustomLoss, self).__init__()
+ assert use_sigmoid is True, 'Only sigmoid focal loss supported now.'
+ self.use_sigmoid = use_sigmoid
+ self.num_classes = num_classes
+ self.gamma = gamma
+ self.alpha = alpha
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+ self.activated = activated
+
+ assert self.num_classes != -1
+
+ # custom output channels of the classifier
+ self.custom_cls_channels = True
+ # custom activation of cls_score
+ self.custom_activation = True
+ # custom accuracy of the classsifier
+ self.custom_accuracy = True
+
+ def get_cls_channels(self, num_classes):
+ assert num_classes == self.num_classes
+ return num_classes
+
+ def get_activation(self, cls_score):
+
+ fine_cls_score = cls_score[:, :self.num_classes]
+
+ score_classes = fine_cls_score.sigmoid()
+
+ return score_classes
+
+ def get_accuracy(self, cls_score, labels):
+
+ fine_cls_score = cls_score[:, :self.num_classes]
+
+ pos_inds = labels < self.num_classes
+ acc_classes = accuracy(fine_cls_score[pos_inds], labels[pos_inds])
+ acc = dict()
+ acc['acc_classes'] = acc_classes
+ return acc
+
+ def forward(self,
+ pred,
+ target,
+ weight=None,
+ avg_factor=None,
+ reduction_override=None):
+ """Forward function.
+
+ Args:
+ pred (torch.Tensor): The prediction.
+ target (torch.Tensor): The learning label of the prediction.
+ weight (torch.Tensor, optional): The weight of loss for each
+ prediction. Defaults to None.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Options are "none", "mean" and "sum".
+
+ Returns:
+ torch.Tensor: The calculated loss
+ """
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ if self.use_sigmoid:
+
+ num_classes = pred.size(1)
+ target = F.one_hot(target, num_classes=num_classes + 1)
+ target = target[:, :num_classes]
+ calculate_loss_func = py_sigmoid_focal_loss
+
+ loss_cls = self.loss_weight * calculate_loss_func(
+ pred,
+ target,
+ weight,
+ gamma=self.gamma,
+ alpha=self.alpha,
+ reduction=reduction,
+ avg_factor=avg_factor)
+
+ else:
+ raise NotImplementedError
+ return loss_cls
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/gaussian_focal_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/gaussian_focal_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..14fa8da462a5e7cabde2166878a1b9f2ccc16d62
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/gaussian_focal_loss.py
@@ -0,0 +1,186 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Union
+
+import torch.nn as nn
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from .utils import weight_reduce_loss, weighted_loss
+
+
+@weighted_loss
+def gaussian_focal_loss(pred: Tensor,
+ gaussian_target: Tensor,
+ alpha: float = 2.0,
+ gamma: float = 4.0,
+ pos_weight: float = 1.0,
+ neg_weight: float = 1.0) -> Tensor:
+ """`Focal Loss `_ for targets in gaussian
+ distribution.
+
+ Args:
+ pred (torch.Tensor): The prediction.
+ gaussian_target (torch.Tensor): The learning target of the prediction
+ in gaussian distribution.
+ alpha (float, optional): A balanced form for Focal Loss.
+ Defaults to 2.0.
+ gamma (float, optional): The gamma for calculating the modulating
+ factor. Defaults to 4.0.
+ pos_weight(float): Positive sample loss weight. Defaults to 1.0.
+ neg_weight(float): Negative sample loss weight. Defaults to 1.0.
+ """
+ eps = 1e-12
+ pos_weights = gaussian_target.eq(1)
+ neg_weights = (1 - gaussian_target).pow(gamma)
+ pos_loss = -(pred + eps).log() * (1 - pred).pow(alpha) * pos_weights
+ neg_loss = -(1 - pred + eps).log() * pred.pow(alpha) * neg_weights
+ return pos_weight * pos_loss + neg_weight * neg_loss
+
+
+def gaussian_focal_loss_with_pos_inds(
+ pred: Tensor,
+ gaussian_target: Tensor,
+ pos_inds: Tensor,
+ pos_labels: Tensor,
+ alpha: float = 2.0,
+ gamma: float = 4.0,
+ pos_weight: float = 1.0,
+ neg_weight: float = 1.0,
+ reduction: str = 'mean',
+ avg_factor: Optional[Union[int, float]] = None) -> Tensor:
+ """`Focal Loss `_ for targets in gaussian
+ distribution.
+
+ Note: The index with a value of 1 in ``gaussian_target`` in the
+ ``gaussian_focal_loss`` function is a positive sample, but in
+ ``gaussian_focal_loss_with_pos_inds`` the positive sample is passed
+ in through the ``pos_inds`` parameter.
+
+ Args:
+ pred (torch.Tensor): The prediction. The shape is (N, num_classes).
+ gaussian_target (torch.Tensor): The learning target of the prediction
+ in gaussian distribution. The shape is (N, num_classes).
+ pos_inds (torch.Tensor): The positive sample index.
+ The shape is (M, ).
+ pos_labels (torch.Tensor): The label corresponding to the positive
+ sample index. The shape is (M, ).
+ alpha (float, optional): A balanced form for Focal Loss.
+ Defaults to 2.0.
+ gamma (float, optional): The gamma for calculating the modulating
+ factor. Defaults to 4.0.
+ pos_weight(float): Positive sample loss weight. Defaults to 1.0.
+ neg_weight(float): Negative sample loss weight. Defaults to 1.0.
+ reduction (str): Options are "none", "mean" and "sum".
+ Defaults to 'mean`.
+ avg_factor (int, float, optional): Average factor that is used to
+ average the loss. Defaults to None.
+ """
+ eps = 1e-12
+ neg_weights = (1 - gaussian_target).pow(gamma)
+
+ pos_pred_pix = pred[pos_inds]
+ pos_pred = pos_pred_pix.gather(1, pos_labels.unsqueeze(1))
+ pos_loss = -(pos_pred + eps).log() * (1 - pos_pred).pow(alpha)
+ pos_loss = weight_reduce_loss(pos_loss, None, reduction, avg_factor)
+
+ neg_loss = -(1 - pred + eps).log() * pred.pow(alpha) * neg_weights
+ neg_loss = weight_reduce_loss(neg_loss, None, reduction, avg_factor)
+
+ return pos_weight * pos_loss + neg_weight * neg_loss
+
+
+@MODELS.register_module()
+class GaussianFocalLoss(nn.Module):
+ """GaussianFocalLoss is a variant of focal loss.
+
+ More details can be found in the `paper
+ `_
+ Code is modified from `kp_utils.py
+ `_ # noqa: E501
+ Please notice that the target in GaussianFocalLoss is a gaussian heatmap,
+ not 0/1 binary target.
+
+ Args:
+ alpha (float): Power of prediction.
+ gamma (float): Power of target for negative samples.
+ reduction (str): Options are "none", "mean" and "sum".
+ loss_weight (float): Loss weight of current loss.
+ pos_weight(float): Positive sample loss weight. Defaults to 1.0.
+ neg_weight(float): Negative sample loss weight. Defaults to 1.0.
+ """
+
+ def __init__(self,
+ alpha: float = 2.0,
+ gamma: float = 4.0,
+ reduction: str = 'mean',
+ loss_weight: float = 1.0,
+ pos_weight: float = 1.0,
+ neg_weight: float = 1.0) -> None:
+ super().__init__()
+ self.alpha = alpha
+ self.gamma = gamma
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+ self.pos_weight = pos_weight
+ self.neg_weight = neg_weight
+
+ def forward(self,
+ pred: Tensor,
+ target: Tensor,
+ pos_inds: Optional[Tensor] = None,
+ pos_labels: Optional[Tensor] = None,
+ weight: Optional[Tensor] = None,
+ avg_factor: Optional[Union[int, float]] = None,
+ reduction_override: Optional[str] = None) -> Tensor:
+ """Forward function.
+
+ If you want to manually determine which positions are
+ positive samples, you can set the pos_index and pos_label
+ parameter. Currently, only the CenterNet update version uses
+ the parameter.
+
+ Args:
+ pred (torch.Tensor): The prediction. The shape is (N, num_classes).
+ target (torch.Tensor): The learning target of the prediction
+ in gaussian distribution. The shape is (N, num_classes).
+ pos_inds (torch.Tensor): The positive sample index.
+ Defaults to None.
+ pos_labels (torch.Tensor): The label corresponding to the positive
+ sample index. Defaults to None.
+ weight (torch.Tensor, optional): The weight of loss for each
+ prediction. Defaults to None.
+ avg_factor (int, float, optional): Average factor that is used to
+ average the loss. Defaults to None.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Defaults to None.
+ """
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ if pos_inds is not None:
+ assert pos_labels is not None
+ # Only used by centernet update version
+ loss_reg = self.loss_weight * gaussian_focal_loss_with_pos_inds(
+ pred,
+ target,
+ pos_inds,
+ pos_labels,
+ alpha=self.alpha,
+ gamma=self.gamma,
+ pos_weight=self.pos_weight,
+ neg_weight=self.neg_weight,
+ reduction=reduction,
+ avg_factor=avg_factor)
+ else:
+ loss_reg = self.loss_weight * gaussian_focal_loss(
+ pred,
+ target,
+ weight,
+ alpha=self.alpha,
+ gamma=self.gamma,
+ pos_weight=self.pos_weight,
+ neg_weight=self.neg_weight,
+ reduction=reduction,
+ avg_factor=avg_factor)
+ return loss_reg
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/gfocal_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/gfocal_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..b3a1172207e859039ca5ed7e0604d8b787131c29
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/gfocal_loss.py
@@ -0,0 +1,295 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from functools import partial
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+
+from mmdet.models.losses.utils import weighted_loss
+from mmdet.registry import MODELS
+
+
+@weighted_loss
+def quality_focal_loss(pred, target, beta=2.0):
+ r"""Quality Focal Loss (QFL) is from `Generalized Focal Loss: Learning
+ Qualified and Distributed Bounding Boxes for Dense Object Detection
+ `_.
+
+ Args:
+ pred (torch.Tensor): Predicted joint representation of classification
+ and quality (IoU) estimation with shape (N, C), C is the number of
+ classes.
+ target (tuple([torch.Tensor])): Target category label with shape (N,)
+ and target quality label with shape (N,).
+ beta (float): The beta parameter for calculating the modulating factor.
+ Defaults to 2.0.
+
+ Returns:
+ torch.Tensor: Loss tensor with shape (N,).
+ """
+ assert len(target) == 2, """target for QFL must be a tuple of two elements,
+ including category label and quality label, respectively"""
+ # label denotes the category id, score denotes the quality score
+ label, score = target
+
+ # negatives are supervised by 0 quality score
+ pred_sigmoid = pred.sigmoid()
+ scale_factor = pred_sigmoid
+ zerolabel = scale_factor.new_zeros(pred.shape)
+ loss = F.binary_cross_entropy_with_logits(
+ pred, zerolabel, reduction='none') * scale_factor.pow(beta)
+
+ # FG cat_id: [0, num_classes -1], BG cat_id: num_classes
+ bg_class_ind = pred.size(1)
+ pos = ((label >= 0) & (label < bg_class_ind)).nonzero().squeeze(1)
+ pos_label = label[pos].long()
+ # positives are supervised by bbox quality (IoU) score
+ scale_factor = score[pos] - pred_sigmoid[pos, pos_label]
+ loss[pos, pos_label] = F.binary_cross_entropy_with_logits(
+ pred[pos, pos_label], score[pos],
+ reduction='none') * scale_factor.abs().pow(beta)
+
+ loss = loss.sum(dim=1, keepdim=False)
+ return loss
+
+
+@weighted_loss
+def quality_focal_loss_tensor_target(pred, target, beta=2.0, activated=False):
+ """`QualityFocal Loss `_
+ Args:
+ pred (torch.Tensor): The prediction with shape (N, C), C is the
+ number of classes
+ target (torch.Tensor): The learning target of the iou-aware
+ classification score with shape (N, C), C is the number of classes.
+ beta (float): The beta parameter for calculating the modulating factor.
+ Defaults to 2.0.
+ activated (bool): Whether the input is activated.
+ If True, it means the input has been activated and can be
+ treated as probabilities. Else, it should be treated as logits.
+ Defaults to False.
+ """
+ # pred and target should be of the same size
+ assert pred.size() == target.size()
+ if activated:
+ pred_sigmoid = pred
+ loss_function = F.binary_cross_entropy
+ else:
+ pred_sigmoid = pred.sigmoid()
+ loss_function = F.binary_cross_entropy_with_logits
+
+ scale_factor = pred_sigmoid
+ target = target.type_as(pred)
+
+ zerolabel = scale_factor.new_zeros(pred.shape)
+ loss = loss_function(
+ pred, zerolabel, reduction='none') * scale_factor.pow(beta)
+
+ pos = (target != 0)
+ scale_factor = target[pos] - pred_sigmoid[pos]
+ loss[pos] = loss_function(
+ pred[pos], target[pos],
+ reduction='none') * scale_factor.abs().pow(beta)
+
+ loss = loss.sum(dim=1, keepdim=False)
+ return loss
+
+
+@weighted_loss
+def quality_focal_loss_with_prob(pred, target, beta=2.0):
+ r"""Quality Focal Loss (QFL) is from `Generalized Focal Loss: Learning
+ Qualified and Distributed Bounding Boxes for Dense Object Detection
+ `_.
+ Different from `quality_focal_loss`, this function accepts probability
+ as input.
+
+ Args:
+ pred (torch.Tensor): Predicted joint representation of classification
+ and quality (IoU) estimation with shape (N, C), C is the number of
+ classes.
+ target (tuple([torch.Tensor])): Target category label with shape (N,)
+ and target quality label with shape (N,).
+ beta (float): The beta parameter for calculating the modulating factor.
+ Defaults to 2.0.
+
+ Returns:
+ torch.Tensor: Loss tensor with shape (N,).
+ """
+ assert len(target) == 2, """target for QFL must be a tuple of two elements,
+ including category label and quality label, respectively"""
+ # label denotes the category id, score denotes the quality score
+ label, score = target
+
+ # negatives are supervised by 0 quality score
+ pred_sigmoid = pred
+ scale_factor = pred_sigmoid
+ zerolabel = scale_factor.new_zeros(pred.shape)
+ loss = F.binary_cross_entropy(
+ pred, zerolabel, reduction='none') * scale_factor.pow(beta)
+
+ # FG cat_id: [0, num_classes -1], BG cat_id: num_classes
+ bg_class_ind = pred.size(1)
+ pos = ((label >= 0) & (label < bg_class_ind)).nonzero().squeeze(1)
+ pos_label = label[pos].long()
+ # positives are supervised by bbox quality (IoU) score
+ scale_factor = score[pos] - pred_sigmoid[pos, pos_label]
+ loss[pos, pos_label] = F.binary_cross_entropy(
+ pred[pos, pos_label], score[pos],
+ reduction='none') * scale_factor.abs().pow(beta)
+
+ loss = loss.sum(dim=1, keepdim=False)
+ return loss
+
+
+@weighted_loss
+def distribution_focal_loss(pred, label):
+ r"""Distribution Focal Loss (DFL) is from `Generalized Focal Loss: Learning
+ Qualified and Distributed Bounding Boxes for Dense Object Detection
+ `_.
+
+ Args:
+ pred (torch.Tensor): Predicted general distribution of bounding boxes
+ (before softmax) with shape (N, n+1), n is the max value of the
+ integral set `{0, ..., n}` in paper.
+ label (torch.Tensor): Target distance label for bounding boxes with
+ shape (N,).
+
+ Returns:
+ torch.Tensor: Loss tensor with shape (N,).
+ """
+ dis_left = label.long()
+ dis_right = dis_left + 1
+ weight_left = dis_right.float() - label
+ weight_right = label - dis_left.float()
+ loss = F.cross_entropy(pred, dis_left, reduction='none') * weight_left \
+ + F.cross_entropy(pred, dis_right, reduction='none') * weight_right
+ return loss
+
+
+@MODELS.register_module()
+class QualityFocalLoss(nn.Module):
+ r"""Quality Focal Loss (QFL) is a variant of `Generalized Focal Loss:
+ Learning Qualified and Distributed Bounding Boxes for Dense Object
+ Detection `_.
+
+ Args:
+ use_sigmoid (bool): Whether sigmoid operation is conducted in QFL.
+ Defaults to True.
+ beta (float): The beta parameter for calculating the modulating factor.
+ Defaults to 2.0.
+ reduction (str): Options are "none", "mean" and "sum".
+ loss_weight (float): Loss weight of current loss.
+ activated (bool, optional): Whether the input is activated.
+ If True, it means the input has been activated and can be
+ treated as probabilities. Else, it should be treated as logits.
+ Defaults to False.
+ """
+
+ def __init__(self,
+ use_sigmoid=True,
+ beta=2.0,
+ reduction='mean',
+ loss_weight=1.0,
+ activated=False):
+ super(QualityFocalLoss, self).__init__()
+ assert use_sigmoid is True, 'Only sigmoid in QFL supported now.'
+ self.use_sigmoid = use_sigmoid
+ self.beta = beta
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+ self.activated = activated
+
+ def forward(self,
+ pred,
+ target,
+ weight=None,
+ avg_factor=None,
+ reduction_override=None):
+ """Forward function.
+
+ Args:
+ pred (torch.Tensor): Predicted joint representation of
+ classification and quality (IoU) estimation with shape (N, C),
+ C is the number of classes.
+ target (Union(tuple([torch.Tensor]),Torch.Tensor)): The type is
+ tuple, it should be included Target category label with
+ shape (N,) and target quality label with shape (N,).The type
+ is torch.Tensor, the target should be one-hot form with
+ soft weights.
+ weight (torch.Tensor, optional): The weight of loss for each
+ prediction. Defaults to None.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Defaults to None.
+ """
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ if self.use_sigmoid:
+ if self.activated:
+ calculate_loss_func = quality_focal_loss_with_prob
+ else:
+ calculate_loss_func = quality_focal_loss
+ if isinstance(target, torch.Tensor):
+ # the target shape with (N,C) or (N,C,...), which means
+ # the target is one-hot form with soft weights.
+ calculate_loss_func = partial(
+ quality_focal_loss_tensor_target, activated=self.activated)
+
+ loss_cls = self.loss_weight * calculate_loss_func(
+ pred,
+ target,
+ weight,
+ beta=self.beta,
+ reduction=reduction,
+ avg_factor=avg_factor)
+ else:
+ raise NotImplementedError
+ return loss_cls
+
+
+@MODELS.register_module()
+class DistributionFocalLoss(nn.Module):
+ r"""Distribution Focal Loss (DFL) is a variant of `Generalized Focal Loss:
+ Learning Qualified and Distributed Bounding Boxes for Dense Object
+ Detection `_.
+
+ Args:
+ reduction (str): Options are `'none'`, `'mean'` and `'sum'`.
+ loss_weight (float): Loss weight of current loss.
+ """
+
+ def __init__(self, reduction='mean', loss_weight=1.0):
+ super(DistributionFocalLoss, self).__init__()
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+
+ def forward(self,
+ pred,
+ target,
+ weight=None,
+ avg_factor=None,
+ reduction_override=None):
+ """Forward function.
+
+ Args:
+ pred (torch.Tensor): Predicted general distribution of bounding
+ boxes (before softmax) with shape (N, n+1), n is the max value
+ of the integral set `{0, ..., n}` in paper.
+ target (torch.Tensor): Target distance label for bounding boxes
+ with shape (N,).
+ weight (torch.Tensor, optional): The weight of loss for each
+ prediction. Defaults to None.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Defaults to None.
+ """
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ loss_cls = self.loss_weight * distribution_focal_loss(
+ pred, target, weight, reduction=reduction, avg_factor=avg_factor)
+ return loss_cls
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/ghm_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/ghm_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..a874c0038cc4a77769705a3a06a95a56d3e8dd2d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/ghm_loss.py
@@ -0,0 +1,213 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+
+from mmdet.registry import MODELS
+from .utils import weight_reduce_loss
+
+
+def _expand_onehot_labels(labels, label_weights, label_channels):
+ bin_labels = labels.new_full((labels.size(0), label_channels), 0)
+ inds = torch.nonzero(
+ (labels >= 0) & (labels < label_channels), as_tuple=False).squeeze()
+ if inds.numel() > 0:
+ bin_labels[inds, labels[inds]] = 1
+ bin_label_weights = label_weights.view(-1, 1).expand(
+ label_weights.size(0), label_channels)
+ return bin_labels, bin_label_weights
+
+
+# TODO: code refactoring to make it consistent with other losses
+@MODELS.register_module()
+class GHMC(nn.Module):
+ """GHM Classification Loss.
+
+ Details of the theorem can be viewed in the paper
+ `Gradient Harmonized Single-stage Detector
+ `_.
+
+ Args:
+ bins (int): Number of the unit regions for distribution calculation.
+ momentum (float): The parameter for moving average.
+ use_sigmoid (bool): Can only be true for BCE based loss now.
+ loss_weight (float): The weight of the total GHM-C loss.
+ reduction (str): Options are "none", "mean" and "sum".
+ Defaults to "mean"
+ """
+
+ def __init__(self,
+ bins=10,
+ momentum=0,
+ use_sigmoid=True,
+ loss_weight=1.0,
+ reduction='mean'):
+ super(GHMC, self).__init__()
+ self.bins = bins
+ self.momentum = momentum
+ edges = torch.arange(bins + 1).float() / bins
+ self.register_buffer('edges', edges)
+ self.edges[-1] += 1e-6
+ if momentum > 0:
+ acc_sum = torch.zeros(bins)
+ self.register_buffer('acc_sum', acc_sum)
+ self.use_sigmoid = use_sigmoid
+ if not self.use_sigmoid:
+ raise NotImplementedError
+ self.loss_weight = loss_weight
+ self.reduction = reduction
+
+ def forward(self,
+ pred,
+ target,
+ label_weight,
+ reduction_override=None,
+ **kwargs):
+ """Calculate the GHM-C loss.
+
+ Args:
+ pred (float tensor of size [batch_num, class_num]):
+ The direct prediction of classification fc layer.
+ target (float tensor of size [batch_num, class_num]):
+ Binary class target for each sample.
+ label_weight (float tensor of size [batch_num, class_num]):
+ the value is 1 if the sample is valid and 0 if ignored.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Defaults to None.
+ Returns:
+ The gradient harmonized loss.
+ """
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ # the target should be binary class label
+ if pred.dim() != target.dim():
+ target, label_weight = _expand_onehot_labels(
+ target, label_weight, pred.size(-1))
+ target, label_weight = target.float(), label_weight.float()
+ edges = self.edges
+ mmt = self.momentum
+ weights = torch.zeros_like(pred)
+
+ # gradient length
+ g = torch.abs(pred.sigmoid().detach() - target)
+
+ valid = label_weight > 0
+ tot = max(valid.float().sum().item(), 1.0)
+ n = 0 # n valid bins
+ for i in range(self.bins):
+ inds = (g >= edges[i]) & (g < edges[i + 1]) & valid
+ num_in_bin = inds.sum().item()
+ if num_in_bin > 0:
+ if mmt > 0:
+ self.acc_sum[i] = mmt * self.acc_sum[i] \
+ + (1 - mmt) * num_in_bin
+ weights[inds] = tot / self.acc_sum[i]
+ else:
+ weights[inds] = tot / num_in_bin
+ n += 1
+ if n > 0:
+ weights = weights / n
+
+ loss = F.binary_cross_entropy_with_logits(
+ pred, target, reduction='none')
+ loss = weight_reduce_loss(
+ loss, weights, reduction=reduction, avg_factor=tot)
+ return loss * self.loss_weight
+
+
+# TODO: code refactoring to make it consistent with other losses
+@MODELS.register_module()
+class GHMR(nn.Module):
+ """GHM Regression Loss.
+
+ Details of the theorem can be viewed in the paper
+ `Gradient Harmonized Single-stage Detector
+ `_.
+
+ Args:
+ mu (float): The parameter for the Authentic Smooth L1 loss.
+ bins (int): Number of the unit regions for distribution calculation.
+ momentum (float): The parameter for moving average.
+ loss_weight (float): The weight of the total GHM-R loss.
+ reduction (str): Options are "none", "mean" and "sum".
+ Defaults to "mean"
+ """
+
+ def __init__(self,
+ mu=0.02,
+ bins=10,
+ momentum=0,
+ loss_weight=1.0,
+ reduction='mean'):
+ super(GHMR, self).__init__()
+ self.mu = mu
+ self.bins = bins
+ edges = torch.arange(bins + 1).float() / bins
+ self.register_buffer('edges', edges)
+ self.edges[-1] = 1e3
+ self.momentum = momentum
+ if momentum > 0:
+ acc_sum = torch.zeros(bins)
+ self.register_buffer('acc_sum', acc_sum)
+ self.loss_weight = loss_weight
+ self.reduction = reduction
+
+ # TODO: support reduction parameter
+ def forward(self,
+ pred,
+ target,
+ label_weight,
+ avg_factor=None,
+ reduction_override=None):
+ """Calculate the GHM-R loss.
+
+ Args:
+ pred (float tensor of size [batch_num, 4 (* class_num)]):
+ The prediction of box regression layer. Channel number can be 4
+ or 4 * class_num depending on whether it is class-agnostic.
+ target (float tensor of size [batch_num, 4 (* class_num)]):
+ The target regression values with the same size of pred.
+ label_weight (float tensor of size [batch_num, 4 (* class_num)]):
+ The weight of each sample, 0 if ignored.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Defaults to None.
+ Returns:
+ The gradient harmonized loss.
+ """
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ mu = self.mu
+ edges = self.edges
+ mmt = self.momentum
+
+ # ASL1 loss
+ diff = pred - target
+ loss = torch.sqrt(diff * diff + mu * mu) - mu
+
+ # gradient length
+ g = torch.abs(diff / torch.sqrt(mu * mu + diff * diff)).detach()
+ weights = torch.zeros_like(g)
+
+ valid = label_weight > 0
+ tot = max(label_weight.float().sum().item(), 1.0)
+ n = 0 # n: valid bins
+ for i in range(self.bins):
+ inds = (g >= edges[i]) & (g < edges[i + 1]) & valid
+ num_in_bin = inds.sum().item()
+ if num_in_bin > 0:
+ n += 1
+ if mmt > 0:
+ self.acc_sum[i] = mmt * self.acc_sum[i] \
+ + (1 - mmt) * num_in_bin
+ weights[inds] = tot / self.acc_sum[i]
+ else:
+ weights[inds] = tot / num_in_bin
+ if n > 0:
+ weights /= n
+ loss = weight_reduce_loss(
+ loss, weights, reduction=reduction, avg_factor=tot)
+ return loss * self.loss_weight
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/iou_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/iou_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..c8a2b977868cef6f4039b49277bfc853ffc720bd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/iou_loss.py
@@ -0,0 +1,926 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+import warnings
+from typing import Optional
+
+import torch
+import torch.nn as nn
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures.bbox import bbox_overlaps
+from .utils import weighted_loss
+
+
+@weighted_loss
+def iou_loss(pred: Tensor,
+ target: Tensor,
+ linear: bool = False,
+ mode: str = 'log',
+ eps: float = 1e-6) -> Tensor:
+ """IoU loss.
+
+ Computing the IoU loss between a set of predicted bboxes and target bboxes.
+ The loss is calculated as negative log of IoU.
+
+ Args:
+ pred (Tensor): Predicted bboxes of format (x1, y1, x2, y2),
+ shape (n, 4).
+ target (Tensor): Corresponding gt bboxes, shape (n, 4).
+ linear (bool, optional): If True, use linear scale of loss instead of
+ log scale. Default: False.
+ mode (str): Loss scaling mode, including "linear", "square", and "log".
+ Default: 'log'
+ eps (float): Epsilon to avoid log(0).
+
+ Return:
+ Tensor: Loss tensor.
+ """
+ assert mode in ['linear', 'square', 'log']
+ if linear:
+ mode = 'linear'
+ warnings.warn('DeprecationWarning: Setting "linear=True" in '
+ 'iou_loss is deprecated, please use "mode=`linear`" '
+ 'instead.')
+ # avoid fp16 overflow
+ if pred.dtype == torch.float16:
+ fp16 = True
+ pred = pred.to(torch.float32)
+ else:
+ fp16 = False
+
+ ious = bbox_overlaps(pred, target, is_aligned=True).clamp(min=eps)
+
+ if fp16:
+ ious = ious.to(torch.float16)
+
+ if mode == 'linear':
+ loss = 1 - ious
+ elif mode == 'square':
+ loss = 1 - ious**2
+ elif mode == 'log':
+ loss = -ious.log()
+ else:
+ raise NotImplementedError
+ return loss
+
+
+@weighted_loss
+def bounded_iou_loss(pred: Tensor,
+ target: Tensor,
+ beta: float = 0.2,
+ eps: float = 1e-3) -> Tensor:
+ """BIoULoss.
+
+ This is an implementation of paper
+ `Improving Object Localization with Fitness NMS and Bounded IoU Loss.
+ `_.
+
+ Args:
+ pred (Tensor): Predicted bboxes of format (x1, y1, x2, y2),
+ shape (n, 4).
+ target (Tensor): Corresponding gt bboxes, shape (n, 4).
+ beta (float, optional): Beta parameter in smoothl1.
+ eps (float, optional): Epsilon to avoid NaN values.
+
+ Return:
+ Tensor: Loss tensor.
+ """
+ pred_ctrx = (pred[:, 0] + pred[:, 2]) * 0.5
+ pred_ctry = (pred[:, 1] + pred[:, 3]) * 0.5
+ pred_w = pred[:, 2] - pred[:, 0]
+ pred_h = pred[:, 3] - pred[:, 1]
+ with torch.no_grad():
+ target_ctrx = (target[:, 0] + target[:, 2]) * 0.5
+ target_ctry = (target[:, 1] + target[:, 3]) * 0.5
+ target_w = target[:, 2] - target[:, 0]
+ target_h = target[:, 3] - target[:, 1]
+
+ dx = target_ctrx - pred_ctrx
+ dy = target_ctry - pred_ctry
+
+ loss_dx = 1 - torch.max(
+ (target_w - 2 * dx.abs()) /
+ (target_w + 2 * dx.abs() + eps), torch.zeros_like(dx))
+ loss_dy = 1 - torch.max(
+ (target_h - 2 * dy.abs()) /
+ (target_h + 2 * dy.abs() + eps), torch.zeros_like(dy))
+ loss_dw = 1 - torch.min(target_w / (pred_w + eps), pred_w /
+ (target_w + eps))
+ loss_dh = 1 - torch.min(target_h / (pred_h + eps), pred_h /
+ (target_h + eps))
+ # view(..., -1) does not work for empty tensor
+ loss_comb = torch.stack([loss_dx, loss_dy, loss_dw, loss_dh],
+ dim=-1).flatten(1)
+
+ loss = torch.where(loss_comb < beta, 0.5 * loss_comb * loss_comb / beta,
+ loss_comb - 0.5 * beta)
+ return loss
+
+
+@weighted_loss
+def giou_loss(pred: Tensor, target: Tensor, eps: float = 1e-7) -> Tensor:
+ r"""`Generalized Intersection over Union: A Metric and A Loss for Bounding
+ Box Regression `_.
+
+ Args:
+ pred (Tensor): Predicted bboxes of format (x1, y1, x2, y2),
+ shape (n, 4).
+ target (Tensor): Corresponding gt bboxes, shape (n, 4).
+ eps (float): Epsilon to avoid log(0).
+
+ Return:
+ Tensor: Loss tensor.
+ """
+ # avoid fp16 overflow
+ if pred.dtype == torch.float16:
+ fp16 = True
+ pred = pred.to(torch.float32)
+ else:
+ fp16 = False
+
+ gious = bbox_overlaps(pred, target, mode='giou', is_aligned=True, eps=eps)
+
+ if fp16:
+ gious = gious.to(torch.float16)
+
+ loss = 1 - gious
+ return loss
+
+
+@weighted_loss
+def diou_loss(pred: Tensor, target: Tensor, eps: float = 1e-7) -> Tensor:
+ r"""Implementation of `Distance-IoU Loss: Faster and Better
+ Learning for Bounding Box Regression https://arxiv.org/abs/1911.08287`_.
+
+ Code is modified from https://github.com/Zzh-tju/DIoU.
+
+ Args:
+ pred (Tensor): Predicted bboxes of format (x1, y1, x2, y2),
+ shape (n, 4).
+ target (Tensor): Corresponding gt bboxes, shape (n, 4).
+ eps (float): Epsilon to avoid log(0).
+
+ Return:
+ Tensor: Loss tensor.
+ """
+ # overlap
+ lt = torch.max(pred[:, :2], target[:, :2])
+ rb = torch.min(pred[:, 2:], target[:, 2:])
+ wh = (rb - lt).clamp(min=0)
+ overlap = wh[:, 0] * wh[:, 1]
+
+ # union
+ ap = (pred[:, 2] - pred[:, 0]) * (pred[:, 3] - pred[:, 1])
+ ag = (target[:, 2] - target[:, 0]) * (target[:, 3] - target[:, 1])
+ union = ap + ag - overlap + eps
+
+ # IoU
+ ious = overlap / union
+
+ # enclose area
+ enclose_x1y1 = torch.min(pred[:, :2], target[:, :2])
+ enclose_x2y2 = torch.max(pred[:, 2:], target[:, 2:])
+ enclose_wh = (enclose_x2y2 - enclose_x1y1).clamp(min=0)
+
+ cw = enclose_wh[:, 0]
+ ch = enclose_wh[:, 1]
+
+ c2 = cw**2 + ch**2 + eps
+
+ b1_x1, b1_y1 = pred[:, 0], pred[:, 1]
+ b1_x2, b1_y2 = pred[:, 2], pred[:, 3]
+ b2_x1, b2_y1 = target[:, 0], target[:, 1]
+ b2_x2, b2_y2 = target[:, 2], target[:, 3]
+
+ left = ((b2_x1 + b2_x2) - (b1_x1 + b1_x2))**2 / 4
+ right = ((b2_y1 + b2_y2) - (b1_y1 + b1_y2))**2 / 4
+ rho2 = left + right
+
+ # DIoU
+ dious = ious - rho2 / c2
+ loss = 1 - dious
+ return loss
+
+
+@weighted_loss
+def ciou_loss(pred: Tensor, target: Tensor, eps: float = 1e-7) -> Tensor:
+ r"""`Implementation of paper `Enhancing Geometric Factors into
+ Model Learning and Inference for Object Detection and Instance
+ Segmentation `_.
+
+ Code is modified from https://github.com/Zzh-tju/CIoU.
+
+ Args:
+ pred (Tensor): Predicted bboxes of format (x1, y1, x2, y2),
+ shape (n, 4).
+ target (Tensor): Corresponding gt bboxes, shape (n, 4).
+ eps (float): Epsilon to avoid log(0).
+
+ Return:
+ Tensor: Loss tensor.
+ """
+ # overlap
+ lt = torch.max(pred[:, :2], target[:, :2])
+ rb = torch.min(pred[:, 2:], target[:, 2:])
+ wh = (rb - lt).clamp(min=0)
+ overlap = wh[:, 0] * wh[:, 1]
+
+ # union
+ ap = (pred[:, 2] - pred[:, 0]) * (pred[:, 3] - pred[:, 1])
+ ag = (target[:, 2] - target[:, 0]) * (target[:, 3] - target[:, 1])
+ union = ap + ag - overlap + eps
+
+ # IoU
+ ious = overlap / union
+
+ # enclose area
+ enclose_x1y1 = torch.min(pred[:, :2], target[:, :2])
+ enclose_x2y2 = torch.max(pred[:, 2:], target[:, 2:])
+ enclose_wh = (enclose_x2y2 - enclose_x1y1).clamp(min=0)
+
+ cw = enclose_wh[:, 0]
+ ch = enclose_wh[:, 1]
+
+ c2 = cw**2 + ch**2 + eps
+
+ b1_x1, b1_y1 = pred[:, 0], pred[:, 1]
+ b1_x2, b1_y2 = pred[:, 2], pred[:, 3]
+ b2_x1, b2_y1 = target[:, 0], target[:, 1]
+ b2_x2, b2_y2 = target[:, 2], target[:, 3]
+
+ w1, h1 = b1_x2 - b1_x1, b1_y2 - b1_y1 + eps
+ w2, h2 = b2_x2 - b2_x1, b2_y2 - b2_y1 + eps
+
+ left = ((b2_x1 + b2_x2) - (b1_x1 + b1_x2))**2 / 4
+ right = ((b2_y1 + b2_y2) - (b1_y1 + b1_y2))**2 / 4
+ rho2 = left + right
+
+ factor = 4 / math.pi**2
+ v = factor * torch.pow(torch.atan(w2 / h2) - torch.atan(w1 / h1), 2)
+
+ with torch.no_grad():
+ alpha = (ious > 0.5).float() * v / (1 - ious + v)
+
+ # CIoU
+ cious = ious - (rho2 / c2 + alpha * v)
+ loss = 1 - cious.clamp(min=-1.0, max=1.0)
+ return loss
+
+
+@weighted_loss
+def eiou_loss(pred: Tensor,
+ target: Tensor,
+ smooth_point: float = 0.1,
+ eps: float = 1e-7) -> Tensor:
+ r"""Implementation of paper `Extended-IoU Loss: A Systematic
+ IoU-Related Method: Beyond Simplified Regression for Better
+ Localization `_
+
+ Code is modified from https://github.com//ShiqiYu/libfacedetection.train.
+
+ Args:
+ pred (Tensor): Predicted bboxes of format (x1, y1, x2, y2),
+ shape (n, 4).
+ target (Tensor): Corresponding gt bboxes, shape (n, 4).
+ smooth_point (float): hyperparameter, default is 0.1.
+ eps (float): Epsilon to avoid log(0).
+
+ Return:
+ Tensor: Loss tensor.
+ """
+ px1, py1, px2, py2 = pred[:, 0], pred[:, 1], pred[:, 2], pred[:, 3]
+ tx1, ty1, tx2, ty2 = target[:, 0], target[:, 1], target[:, 2], target[:, 3]
+
+ # extent top left
+ ex1 = torch.min(px1, tx1)
+ ey1 = torch.min(py1, ty1)
+
+ # intersection coordinates
+ ix1 = torch.max(px1, tx1)
+ iy1 = torch.max(py1, ty1)
+ ix2 = torch.min(px2, tx2)
+ iy2 = torch.min(py2, ty2)
+
+ # extra
+ xmin = torch.min(ix1, ix2)
+ ymin = torch.min(iy1, iy2)
+ xmax = torch.max(ix1, ix2)
+ ymax = torch.max(iy1, iy2)
+
+ # Intersection
+ intersection = (ix2 - ex1) * (iy2 - ey1) + (xmin - ex1) * (ymin - ey1) - (
+ ix1 - ex1) * (ymax - ey1) - (xmax - ex1) * (
+ iy1 - ey1)
+ # Union
+ union = (px2 - px1) * (py2 - py1) + (tx2 - tx1) * (
+ ty2 - ty1) - intersection + eps
+ # IoU
+ ious = 1 - (intersection / union)
+
+ # Smooth-EIoU
+ smooth_sign = (ious < smooth_point).detach().float()
+ loss = 0.5 * smooth_sign * (ious**2) / smooth_point + (1 - smooth_sign) * (
+ ious - 0.5 * smooth_point)
+ return loss
+
+
+@weighted_loss
+def siou_loss(pred, target, eps=1e-7, neg_gamma=False):
+ r"""`Implementation of paper `SIoU Loss: More Powerful Learning
+ for Bounding Box Regression `_.
+
+ Code is modified from https://github.com/meituan/YOLOv6.
+
+ Args:
+ pred (Tensor): Predicted bboxes of format (x1, y1, x2, y2),
+ shape (n, 4).
+ target (Tensor): Corresponding gt bboxes, shape (n, 4).
+ eps (float): Eps to avoid log(0).
+ neg_gamma (bool): `True` follows original implementation in paper.
+
+ Return:
+ Tensor: Loss tensor.
+ """
+ # overlap
+ lt = torch.max(pred[:, :2], target[:, :2])
+ rb = torch.min(pred[:, 2:], target[:, 2:])
+ wh = (rb - lt).clamp(min=0)
+ overlap = wh[:, 0] * wh[:, 1]
+
+ # union
+ ap = (pred[:, 2] - pred[:, 0]) * (pred[:, 3] - pred[:, 1])
+ ag = (target[:, 2] - target[:, 0]) * (target[:, 3] - target[:, 1])
+ union = ap + ag - overlap + eps
+
+ # IoU
+ ious = overlap / union
+
+ # enclose area
+ enclose_x1y1 = torch.min(pred[:, :2], target[:, :2])
+ enclose_x2y2 = torch.max(pred[:, 2:], target[:, 2:])
+ # modified clamp threshold zero to eps to avoid NaN
+ enclose_wh = (enclose_x2y2 - enclose_x1y1).clamp(min=eps)
+
+ cw = enclose_wh[:, 0]
+ ch = enclose_wh[:, 1]
+
+ b1_x1, b1_y1 = pred[:, 0], pred[:, 1]
+ b1_x2, b1_y2 = pred[:, 2], pred[:, 3]
+ b2_x1, b2_y1 = target[:, 0], target[:, 1]
+ b2_x2, b2_y2 = target[:, 2], target[:, 3]
+
+ w1, h1 = b1_x2 - b1_x1, b1_y2 - b1_y1 + eps
+ w2, h2 = b2_x2 - b2_x1, b2_y2 - b2_y1 + eps
+
+ # angle cost
+ s_cw = (b2_x1 + b2_x2 - b1_x1 - b1_x2) * 0.5 + eps
+ s_ch = (b2_y1 + b2_y2 - b1_y1 - b1_y2) * 0.5 + eps
+
+ sigma = torch.pow(s_cw**2 + s_ch**2, 0.5)
+
+ sin_alpha_1 = torch.abs(s_cw) / sigma
+ sin_alpha_2 = torch.abs(s_ch) / sigma
+ threshold = pow(2, 0.5) / 2
+ sin_alpha = torch.where(sin_alpha_1 > threshold, sin_alpha_2, sin_alpha_1)
+ angle_cost = torch.cos(torch.asin(sin_alpha) * 2 - math.pi / 2)
+
+ # distance cost
+ rho_x = (s_cw / cw)**2
+ rho_y = (s_ch / ch)**2
+
+ # `neg_gamma=True` follows original implementation in paper
+ # but setting `neg_gamma=False` makes training more stable.
+ gamma = angle_cost - 2 if neg_gamma else 2 - angle_cost
+ distance_cost = 2 - torch.exp(gamma * rho_x) - torch.exp(gamma * rho_y)
+
+ # shape cost
+ omiga_w = torch.abs(w1 - w2) / torch.max(w1, w2)
+ omiga_h = torch.abs(h1 - h2) / torch.max(h1, h2)
+ shape_cost = torch.pow(1 - torch.exp(-1 * omiga_w), 4) + torch.pow(
+ 1 - torch.exp(-1 * omiga_h), 4)
+
+ # SIoU
+ sious = ious - 0.5 * (distance_cost + shape_cost)
+ loss = 1 - sious.clamp(min=-1.0, max=1.0)
+ return loss
+
+
+@MODELS.register_module()
+class IoULoss(nn.Module):
+ """IoULoss.
+
+ Computing the IoU loss between a set of predicted bboxes and target bboxes.
+
+ Args:
+ linear (bool): If True, use linear scale of loss else determined
+ by mode. Default: False.
+ eps (float): Epsilon to avoid log(0).
+ reduction (str): Options are "none", "mean" and "sum".
+ loss_weight (float): Weight of loss.
+ mode (str): Loss scaling mode, including "linear", "square", and "log".
+ Default: 'log'
+ """
+
+ def __init__(self,
+ linear: bool = False,
+ eps: float = 1e-6,
+ reduction: str = 'mean',
+ loss_weight: float = 1.0,
+ mode: str = 'log') -> None:
+ super().__init__()
+ assert mode in ['linear', 'square', 'log']
+ if linear:
+ mode = 'linear'
+ warnings.warn('DeprecationWarning: Setting "linear=True" in '
+ 'IOULoss is deprecated, please use "mode=`linear`" '
+ 'instead.')
+ self.mode = mode
+ self.linear = linear
+ self.eps = eps
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+
+ def forward(self,
+ pred: Tensor,
+ target: Tensor,
+ weight: Optional[Tensor] = None,
+ avg_factor: Optional[int] = None,
+ reduction_override: Optional[str] = None,
+ **kwargs) -> Tensor:
+ """Forward function.
+
+ Args:
+ pred (Tensor): Predicted bboxes of format (x1, y1, x2, y2),
+ shape (n, 4).
+ target (Tensor): The learning target of the prediction,
+ shape (n, 4).
+ weight (Tensor, optional): The weight of loss for each
+ prediction. Defaults to None.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Defaults to None. Options are "none", "mean" and "sum".
+
+ Return:
+ Tensor: Loss tensor.
+ """
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ if (weight is not None) and (not torch.any(weight > 0)) and (
+ reduction != 'none'):
+ if pred.dim() == weight.dim() + 1:
+ weight = weight.unsqueeze(1)
+ return (pred * weight).sum() # 0
+ if weight is not None and weight.dim() > 1:
+ # TODO: remove this in the future
+ # reduce the weight of shape (n, 4) to (n,) to match the
+ # iou_loss of shape (n,)
+ assert weight.shape == pred.shape
+ weight = weight.mean(-1)
+ loss = self.loss_weight * iou_loss(
+ pred,
+ target,
+ weight,
+ mode=self.mode,
+ eps=self.eps,
+ reduction=reduction,
+ avg_factor=avg_factor,
+ **kwargs)
+ return loss
+
+
+@MODELS.register_module()
+class BoundedIoULoss(nn.Module):
+ """BIoULoss.
+
+ This is an implementation of paper
+ `Improving Object Localization with Fitness NMS and Bounded IoU Loss.
+ `_.
+
+ Args:
+ beta (float, optional): Beta parameter in smoothl1.
+ eps (float, optional): Epsilon to avoid NaN values.
+ reduction (str): Options are "none", "mean" and "sum".
+ loss_weight (float): Weight of loss.
+ """
+
+ def __init__(self,
+ beta: float = 0.2,
+ eps: float = 1e-3,
+ reduction: str = 'mean',
+ loss_weight: float = 1.0) -> None:
+ super().__init__()
+ self.beta = beta
+ self.eps = eps
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+
+ def forward(self,
+ pred: Tensor,
+ target: Tensor,
+ weight: Optional[Tensor] = None,
+ avg_factor: Optional[int] = None,
+ reduction_override: Optional[str] = None,
+ **kwargs) -> Tensor:
+ """Forward function.
+
+ Args:
+ pred (Tensor): Predicted bboxes of format (x1, y1, x2, y2),
+ shape (n, 4).
+ target (Tensor): The learning target of the prediction,
+ shape (n, 4).
+ weight (Optional[Tensor], optional): The weight of loss for each
+ prediction. Defaults to None.
+ avg_factor (Optional[int], optional): Average factor that is used
+ to average the loss. Defaults to None.
+ reduction_override (Optional[str], optional): The reduction method
+ used to override the original reduction method of the loss.
+ Defaults to None. Options are "none", "mean" and "sum".
+
+ Returns:
+ Tensor: Loss tensor.
+ """
+ if weight is not None and not torch.any(weight > 0):
+ if pred.dim() == weight.dim() + 1:
+ weight = weight.unsqueeze(1)
+ return (pred * weight).sum() # 0
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ loss = self.loss_weight * bounded_iou_loss(
+ pred,
+ target,
+ weight,
+ beta=self.beta,
+ eps=self.eps,
+ reduction=reduction,
+ avg_factor=avg_factor,
+ **kwargs)
+ return loss
+
+
+@MODELS.register_module()
+class GIoULoss(nn.Module):
+ r"""`Generalized Intersection over Union: A Metric and A Loss for Bounding
+ Box Regression `_.
+
+ Args:
+ eps (float): Epsilon to avoid log(0).
+ reduction (str): Options are "none", "mean" and "sum".
+ loss_weight (float): Weight of loss.
+ """
+
+ def __init__(self,
+ eps: float = 1e-6,
+ reduction: str = 'mean',
+ loss_weight: float = 1.0) -> None:
+ super().__init__()
+ self.eps = eps
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+
+ def forward(self,
+ pred: Tensor,
+ target: Tensor,
+ weight: Optional[Tensor] = None,
+ avg_factor: Optional[int] = None,
+ reduction_override: Optional[str] = None,
+ **kwargs) -> Tensor:
+ """Forward function.
+
+ Args:
+ pred (Tensor): Predicted bboxes of format (x1, y1, x2, y2),
+ shape (n, 4).
+ target (Tensor): The learning target of the prediction,
+ shape (n, 4).
+ weight (Optional[Tensor], optional): The weight of loss for each
+ prediction. Defaults to None.
+ avg_factor (Optional[int], optional): Average factor that is used
+ to average the loss. Defaults to None.
+ reduction_override (Optional[str], optional): The reduction method
+ used to override the original reduction method of the loss.
+ Defaults to None. Options are "none", "mean" and "sum".
+
+ Returns:
+ Tensor: Loss tensor.
+ """
+ if weight is not None and not torch.any(weight > 0):
+ if pred.dim() == weight.dim() + 1:
+ weight = weight.unsqueeze(1)
+ return (pred * weight).sum() # 0
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ if weight is not None and weight.dim() > 1:
+ # TODO: remove this in the future
+ # reduce the weight of shape (n, 4) to (n,) to match the
+ # giou_loss of shape (n,)
+ assert weight.shape == pred.shape
+ weight = weight.mean(-1)
+ loss = self.loss_weight * giou_loss(
+ pred,
+ target,
+ weight,
+ eps=self.eps,
+ reduction=reduction,
+ avg_factor=avg_factor,
+ **kwargs)
+ return loss
+
+
+@MODELS.register_module()
+class DIoULoss(nn.Module):
+ r"""Implementation of `Distance-IoU Loss: Faster and Better
+ Learning for Bounding Box Regression https://arxiv.org/abs/1911.08287`_.
+
+ Code is modified from https://github.com/Zzh-tju/DIoU.
+
+ Args:
+ eps (float): Epsilon to avoid log(0).
+ reduction (str): Options are "none", "mean" and "sum".
+ loss_weight (float): Weight of loss.
+ """
+
+ def __init__(self,
+ eps: float = 1e-6,
+ reduction: str = 'mean',
+ loss_weight: float = 1.0) -> None:
+ super().__init__()
+ self.eps = eps
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+
+ def forward(self,
+ pred: Tensor,
+ target: Tensor,
+ weight: Optional[Tensor] = None,
+ avg_factor: Optional[int] = None,
+ reduction_override: Optional[str] = None,
+ **kwargs) -> Tensor:
+ """Forward function.
+
+ Args:
+ pred (Tensor): Predicted bboxes of format (x1, y1, x2, y2),
+ shape (n, 4).
+ target (Tensor): The learning target of the prediction,
+ shape (n, 4).
+ weight (Optional[Tensor], optional): The weight of loss for each
+ prediction. Defaults to None.
+ avg_factor (Optional[int], optional): Average factor that is used
+ to average the loss. Defaults to None.
+ reduction_override (Optional[str], optional): The reduction method
+ used to override the original reduction method of the loss.
+ Defaults to None. Options are "none", "mean" and "sum".
+
+ Returns:
+ Tensor: Loss tensor.
+ """
+ if weight is not None and not torch.any(weight > 0):
+ if pred.dim() == weight.dim() + 1:
+ weight = weight.unsqueeze(1)
+ return (pred * weight).sum() # 0
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ if weight is not None and weight.dim() > 1:
+ # TODO: remove this in the future
+ # reduce the weight of shape (n, 4) to (n,) to match the
+ # giou_loss of shape (n,)
+ assert weight.shape == pred.shape
+ weight = weight.mean(-1)
+ loss = self.loss_weight * diou_loss(
+ pred,
+ target,
+ weight,
+ eps=self.eps,
+ reduction=reduction,
+ avg_factor=avg_factor,
+ **kwargs)
+ return loss
+
+
+@MODELS.register_module()
+class CIoULoss(nn.Module):
+ r"""`Implementation of paper `Enhancing Geometric Factors into
+ Model Learning and Inference for Object Detection and Instance
+ Segmentation `_.
+
+ Code is modified from https://github.com/Zzh-tju/CIoU.
+
+ Args:
+ eps (float): Epsilon to avoid log(0).
+ reduction (str): Options are "none", "mean" and "sum".
+ loss_weight (float): Weight of loss.
+ """
+
+ def __init__(self,
+ eps: float = 1e-6,
+ reduction: str = 'mean',
+ loss_weight: float = 1.0) -> None:
+ super().__init__()
+ self.eps = eps
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+
+ def forward(self,
+ pred: Tensor,
+ target: Tensor,
+ weight: Optional[Tensor] = None,
+ avg_factor: Optional[int] = None,
+ reduction_override: Optional[str] = None,
+ **kwargs) -> Tensor:
+ """Forward function.
+
+ Args:
+ pred (Tensor): Predicted bboxes of format (x1, y1, x2, y2),
+ shape (n, 4).
+ target (Tensor): The learning target of the prediction,
+ shape (n, 4).
+ weight (Optional[Tensor], optional): The weight of loss for each
+ prediction. Defaults to None.
+ avg_factor (Optional[int], optional): Average factor that is used
+ to average the loss. Defaults to None.
+ reduction_override (Optional[str], optional): The reduction method
+ used to override the original reduction method of the loss.
+ Defaults to None. Options are "none", "mean" and "sum".
+
+ Returns:
+ Tensor: Loss tensor.
+ """
+ if weight is not None and not torch.any(weight > 0):
+ if pred.dim() == weight.dim() + 1:
+ weight = weight.unsqueeze(1)
+ return (pred * weight).sum() # 0
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ if weight is not None and weight.dim() > 1:
+ # TODO: remove this in the future
+ # reduce the weight of shape (n, 4) to (n,) to match the
+ # giou_loss of shape (n,)
+ assert weight.shape == pred.shape
+ weight = weight.mean(-1)
+ loss = self.loss_weight * ciou_loss(
+ pred,
+ target,
+ weight,
+ eps=self.eps,
+ reduction=reduction,
+ avg_factor=avg_factor,
+ **kwargs)
+ return loss
+
+
+@MODELS.register_module()
+class EIoULoss(nn.Module):
+ r"""Implementation of paper `Extended-IoU Loss: A Systematic
+ IoU-Related Method: Beyond Simplified Regression for Better
+ Localization `_
+
+ Code is modified from https://github.com//ShiqiYu/libfacedetection.train.
+
+ Args:
+ eps (float): Epsilon to avoid log(0).
+ reduction (str): Options are "none", "mean" and "sum".
+ loss_weight (float): Weight of loss.
+ smooth_point (float): hyperparameter, default is 0.1.
+ """
+
+ def __init__(self,
+ eps: float = 1e-6,
+ reduction: str = 'mean',
+ loss_weight: float = 1.0,
+ smooth_point: float = 0.1) -> None:
+ super().__init__()
+ self.eps = eps
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+ self.smooth_point = smooth_point
+
+ def forward(self,
+ pred: Tensor,
+ target: Tensor,
+ weight: Optional[Tensor] = None,
+ avg_factor: Optional[int] = None,
+ reduction_override: Optional[str] = None,
+ **kwargs) -> Tensor:
+ """Forward function.
+
+ Args:
+ pred (Tensor): Predicted bboxes of format (x1, y1, x2, y2),
+ shape (n, 4).
+ target (Tensor): The learning target of the prediction,
+ shape (n, 4).
+ weight (Optional[Tensor], optional): The weight of loss for each
+ prediction. Defaults to None.
+ avg_factor (Optional[int], optional): Average factor that is used
+ to average the loss. Defaults to None.
+ reduction_override (Optional[str], optional): The reduction method
+ used to override the original reduction method of the loss.
+ Defaults to None. Options are "none", "mean" and "sum".
+
+ Returns:
+ Tensor: Loss tensor.
+ """
+ if weight is not None and not torch.any(weight > 0):
+ if pred.dim() == weight.dim() + 1:
+ weight = weight.unsqueeze(1)
+ return (pred * weight).sum() # 0
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ if weight is not None and weight.dim() > 1:
+ assert weight.shape == pred.shape
+ weight = weight.mean(-1)
+ loss = self.loss_weight * eiou_loss(
+ pred,
+ target,
+ weight,
+ smooth_point=self.smooth_point,
+ eps=self.eps,
+ reduction=reduction,
+ avg_factor=avg_factor,
+ **kwargs)
+ return loss
+
+
+@MODELS.register_module()
+class SIoULoss(nn.Module):
+ r"""`Implementation of paper `SIoU Loss: More Powerful Learning
+ for Bounding Box Regression `_.
+
+ Code is modified from https://github.com/meituan/YOLOv6.
+
+ Args:
+ pred (Tensor): Predicted bboxes of format (x1, y1, x2, y2),
+ shape (n, 4).
+ target (Tensor): Corresponding gt bboxes, shape (n, 4).
+ eps (float): Eps to avoid log(0).
+ neg_gamma (bool): `True` follows original implementation in paper.
+
+ Return:
+ Tensor: Loss tensor.
+ """
+
+ def __init__(self,
+ eps: float = 1e-6,
+ reduction: str = 'mean',
+ loss_weight: float = 1.0,
+ neg_gamma: bool = False) -> None:
+ super().__init__()
+ self.eps = eps
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+ self.neg_gamma = neg_gamma
+
+ def forward(self,
+ pred: Tensor,
+ target: Tensor,
+ weight: Optional[Tensor] = None,
+ avg_factor: Optional[int] = None,
+ reduction_override: Optional[str] = None,
+ **kwargs) -> Tensor:
+ """Forward function.
+
+ Args:
+ pred (Tensor): Predicted bboxes of format (x1, y1, x2, y2),
+ shape (n, 4).
+ target (Tensor): The learning target of the prediction,
+ shape (n, 4).
+ weight (Optional[Tensor], optional): The weight of loss for each
+ prediction. Defaults to None.
+ avg_factor (Optional[int], optional): Average factor that is used
+ to average the loss. Defaults to None.
+ reduction_override (Optional[str], optional): The reduction method
+ used to override the original reduction method of the loss.
+ Defaults to None. Options are "none", "mean" and "sum".
+
+ Returns:
+ Tensor: Loss tensor.
+ """
+ if weight is not None and not torch.any(weight > 0):
+ if pred.dim() == weight.dim() + 1:
+ weight = weight.unsqueeze(1)
+ return (pred * weight).sum() # 0
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ if weight is not None and weight.dim() > 1:
+ # TODO: remove this in the future
+ # reduce the weight of shape (n, 4) to (n,) to match the
+ # giou_loss of shape (n,)
+ assert weight.shape == pred.shape
+ weight = weight.mean(-1)
+ loss = self.loss_weight * siou_loss(
+ pred,
+ target,
+ weight,
+ eps=self.eps,
+ reduction=reduction,
+ avg_factor=avg_factor,
+ neg_gamma=self.neg_gamma,
+ **kwargs)
+ return loss
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/kd_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/kd_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..0a7d5ef24a0b0d7d7390a27c7cd9cbfdbe61d823
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/kd_loss.py
@@ -0,0 +1,95 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional
+
+import torch.nn as nn
+import torch.nn.functional as F
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from .utils import weighted_loss
+
+
+@weighted_loss
+def knowledge_distillation_kl_div_loss(pred: Tensor,
+ soft_label: Tensor,
+ T: int,
+ detach_target: bool = True) -> Tensor:
+ r"""Loss function for knowledge distilling using KL divergence.
+
+ Args:
+ pred (Tensor): Predicted logits with shape (N, n + 1).
+ soft_label (Tensor): Target logits with shape (N, N + 1).
+ T (int): Temperature for distillation.
+ detach_target (bool): Remove soft_label from automatic differentiation
+
+ Returns:
+ Tensor: Loss tensor with shape (N,).
+ """
+ assert pred.size() == soft_label.size()
+ target = F.softmax(soft_label / T, dim=1)
+ if detach_target:
+ target = target.detach()
+
+ kd_loss = F.kl_div(
+ F.log_softmax(pred / T, dim=1), target, reduction='none').mean(1) * (
+ T * T)
+
+ return kd_loss
+
+
+@MODELS.register_module()
+class KnowledgeDistillationKLDivLoss(nn.Module):
+ """Loss function for knowledge distilling using KL divergence.
+
+ Args:
+ reduction (str): Options are `'none'`, `'mean'` and `'sum'`.
+ loss_weight (float): Loss weight of current loss.
+ T (int): Temperature for distillation.
+ """
+
+ def __init__(self,
+ reduction: str = 'mean',
+ loss_weight: float = 1.0,
+ T: int = 10) -> None:
+ super().__init__()
+ assert T >= 1
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+ self.T = T
+
+ def forward(self,
+ pred: Tensor,
+ soft_label: Tensor,
+ weight: Optional[Tensor] = None,
+ avg_factor: Optional[int] = None,
+ reduction_override: Optional[str] = None) -> Tensor:
+ """Forward function.
+
+ Args:
+ pred (Tensor): Predicted logits with shape (N, n + 1).
+ soft_label (Tensor): Target logits with shape (N, N + 1).
+ weight (Tensor, optional): The weight of loss for each
+ prediction. Defaults to None.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Defaults to None.
+
+ Returns:
+ Tensor: Loss tensor.
+ """
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+
+ loss_kd = self.loss_weight * knowledge_distillation_kl_div_loss(
+ pred,
+ soft_label,
+ weight,
+ reduction=reduction,
+ avg_factor=avg_factor,
+ T=self.T)
+
+ return loss_kd
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/l2_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/l2_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..6210a3007b2c39540f022925cc93181c7328e42d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/l2_loss.py
@@ -0,0 +1,139 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Tuple, Union
+
+import numpy as np
+import torch
+from mmengine.model import BaseModule
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from .utils import weighted_loss
+
+
+@weighted_loss
+def l2_loss(pred: Tensor, target: Tensor) -> Tensor:
+ """L2 loss.
+
+ Args:
+ pred (torch.Tensor): The prediction.
+ target (torch.Tensor): The learning target of the prediction.
+
+ Returns:
+ torch.Tensor: Calculated loss
+ """
+ assert pred.size() == target.size()
+ loss = torch.abs(pred - target)**2
+ return loss
+
+
+@MODELS.register_module()
+class L2Loss(BaseModule):
+ """L2 loss.
+
+ Args:
+ reduction (str, optional): The method to reduce the loss.
+ Options are "none", "mean" and "sum".
+ loss_weight (float, optional): The weight of loss.
+ """
+
+ def __init__(self,
+ neg_pos_ub: int = -1,
+ pos_margin: float = -1,
+ neg_margin: float = -1,
+ hard_mining: bool = False,
+ reduction: str = 'mean',
+ loss_weight: float = 1.0):
+ super(L2Loss, self).__init__()
+ self.neg_pos_ub = neg_pos_ub
+ self.pos_margin = pos_margin
+ self.neg_margin = neg_margin
+ self.hard_mining = hard_mining
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+
+ def forward(self,
+ pred: Tensor,
+ target: Tensor,
+ weight: Optional[Tensor] = None,
+ avg_factor: Optional[float] = None,
+ reduction_override: Optional[str] = None) -> Tensor:
+ """Forward function.
+
+ Args:
+ pred (torch.Tensor): The prediction.
+ target (torch.Tensor): The learning target of the prediction.
+ weight (torch.Tensor, optional): The weight of loss for each
+ prediction. Defaults to None.
+ avg_factor (float, optional): Average factor that is used to
+ average the loss. Defaults to None.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Defaults to None.
+ """
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ pred, weight, avg_factor = self.update_weight(pred, target, weight,
+ avg_factor)
+ loss_bbox = self.loss_weight * l2_loss(
+ pred, target, weight, reduction=reduction, avg_factor=avg_factor)
+ return loss_bbox
+
+ def update_weight(self, pred: Tensor, target: Tensor, weight: Tensor,
+ avg_factor: float) -> Tuple[Tensor, Tensor, float]:
+ """Update the weight according to targets."""
+ if weight is None:
+ weight = target.new_ones(target.size())
+
+ invalid_inds = weight <= 0
+ target[invalid_inds] = -1
+ pos_inds = target == 1
+ neg_inds = target == 0
+
+ if self.pos_margin > 0:
+ pred[pos_inds] -= self.pos_margin
+ if self.neg_margin > 0:
+ pred[neg_inds] -= self.neg_margin
+ pred = torch.clamp(pred, min=0, max=1)
+
+ num_pos = int((target == 1).sum())
+ num_neg = int((target == 0).sum())
+ if self.neg_pos_ub > 0 and num_neg / (num_pos +
+ 1e-6) > self.neg_pos_ub:
+ num_neg = num_pos * self.neg_pos_ub
+ neg_idx = torch.nonzero(target == 0, as_tuple=False)
+
+ if self.hard_mining:
+ costs = l2_loss(
+ pred, target, reduction='none')[neg_idx[:, 0],
+ neg_idx[:, 1]].detach()
+ neg_idx = neg_idx[costs.topk(num_neg)[1], :]
+ else:
+ neg_idx = self.random_choice(neg_idx, num_neg)
+
+ new_neg_inds = neg_inds.new_zeros(neg_inds.size()).bool()
+ new_neg_inds[neg_idx[:, 0], neg_idx[:, 1]] = True
+
+ invalid_neg_inds = torch.logical_xor(neg_inds, new_neg_inds)
+ weight[invalid_neg_inds] = 0
+
+ avg_factor = (weight > 0).sum()
+ return pred, weight, avg_factor
+
+ @staticmethod
+ def random_choice(gallery: Union[list, np.ndarray, Tensor],
+ num: int) -> np.ndarray:
+ """Random select some elements from the gallery.
+
+ It seems that Pytorch's implementation is slower than numpy so we use
+ numpy to randperm the indices.
+ """
+ assert len(gallery) >= num
+ if isinstance(gallery, list):
+ gallery = np.array(gallery)
+ cands = np.arange(len(gallery))
+ np.random.shuffle(cands)
+ rand_inds = cands[:num]
+ if not isinstance(gallery, np.ndarray):
+ rand_inds = torch.from_numpy(rand_inds).long().to(gallery.device)
+ return gallery[rand_inds]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/margin_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/margin_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..0609e1db50edf89c8ae8b65709e8ab786f580366
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/margin_loss.py
@@ -0,0 +1,152 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Tuple, Union
+
+import numpy as np
+import torch
+from mmengine.model import BaseModule
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from .mse_loss import mse_loss
+
+
+@MODELS.register_module()
+class MarginL2Loss(BaseModule):
+ """L2 loss with margin.
+
+ Args:
+ neg_pos_ub (int, optional): The upper bound of negative to positive
+ samples in hard mining. Defaults to -1.
+ pos_margin (float, optional): The similarity margin for positive
+ samples in hard mining. Defaults to -1.
+ neg_margin (float, optional): The similarity margin for negative
+ samples in hard mining. Defaults to -1.
+ hard_mining (bool, optional): Whether to use hard mining. Defaults to
+ False.
+ reduction (str, optional): The method to reduce the loss.
+ Options are "none", "mean" and "sum". Defaults to "mean".
+ loss_weight (float, optional): The weight of loss. Defaults to 1.0.
+ """
+
+ def __init__(self,
+ neg_pos_ub: int = -1,
+ pos_margin: float = -1,
+ neg_margin: float = -1,
+ hard_mining: bool = False,
+ reduction: str = 'mean',
+ loss_weight: float = 1.0):
+ super(MarginL2Loss, self).__init__()
+ self.neg_pos_ub = neg_pos_ub
+ self.pos_margin = pos_margin
+ self.neg_margin = neg_margin
+ self.hard_mining = hard_mining
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+
+ def forward(self,
+ pred: Tensor,
+ target: Tensor,
+ weight: Optional[Tensor] = None,
+ avg_factor: Optional[float] = None,
+ reduction_override: Optional[str] = None) -> Tensor:
+ """Forward function.
+
+ Args:
+ pred (torch.Tensor): The prediction.
+ target (torch.Tensor): The learning target of the prediction.
+ weight (torch.Tensor, optional): The weight of loss for each
+ prediction. Defaults to None.
+ avg_factor (float, optional): Average factor that is used to
+ average the loss. Defaults to None.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Defaults to None.
+ """
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ pred, weight, avg_factor = self.update_weight(pred, target, weight,
+ avg_factor)
+ loss_bbox = self.loss_weight * mse_loss(
+ pred,
+ target.float(),
+ weight.float(),
+ reduction=reduction,
+ avg_factor=avg_factor)
+ return loss_bbox
+
+ def update_weight(self, pred: Tensor, target: Tensor, weight: Tensor,
+ avg_factor: float) -> Tuple[Tensor, Tensor, float]:
+ """Update the weight according to targets.
+
+ Args:
+ pred (torch.Tensor): The prediction.
+ target (torch.Tensor): The learning target of the prediction.
+ weight (torch.Tensor): The weight of loss for each prediction.
+ avg_factor (float): Average factor that is used to average the
+ loss.
+
+ Returns:
+ tuple[torch.Tensor]: The updated prediction, weight and average
+ factor.
+ """
+ if weight is None:
+ weight = target.new_ones(target.size())
+
+ invalid_inds = weight <= 0
+ target[invalid_inds] = -1
+ pos_inds = target == 1
+ neg_inds = target == 0
+
+ if self.pos_margin > 0:
+ pred[pos_inds] -= self.pos_margin
+ if self.neg_margin > 0:
+ pred[neg_inds] -= self.neg_margin
+ pred = torch.clamp(pred, min=0, max=1)
+
+ num_pos = int((target == 1).sum())
+ num_neg = int((target == 0).sum())
+ if self.neg_pos_ub > 0 and num_neg / (num_pos +
+ 1e-6) > self.neg_pos_ub:
+ num_neg = num_pos * self.neg_pos_ub
+ neg_idx = torch.nonzero(target == 0, as_tuple=False)
+
+ if self.hard_mining:
+ costs = mse_loss(
+ pred, target.float(),
+ reduction='none')[neg_idx[:, 0], neg_idx[:, 1]].detach()
+ neg_idx = neg_idx[costs.topk(num_neg)[1], :]
+ else:
+ neg_idx = self.random_choice(neg_idx, num_neg)
+
+ new_neg_inds = neg_inds.new_zeros(neg_inds.size()).bool()
+ new_neg_inds[neg_idx[:, 0], neg_idx[:, 1]] = True
+
+ invalid_neg_inds = torch.logical_xor(neg_inds, new_neg_inds)
+ weight[invalid_neg_inds] = 0
+
+ avg_factor = (weight > 0).sum()
+ return pred, weight, avg_factor
+
+ @staticmethod
+ def random_choice(gallery: Union[list, np.ndarray, Tensor],
+ num: int) -> np.ndarray:
+ """Random select some elements from the gallery.
+
+ It seems that Pytorch's implementation is slower than numpy so we use
+ numpy to randperm the indices.
+
+ Args:
+ gallery (list | np.ndarray | torch.Tensor): The gallery from
+ which to sample.
+ num (int): The number of elements to sample.
+ """
+ assert len(gallery) >= num
+ if isinstance(gallery, list):
+ gallery = np.array(gallery)
+ cands = np.arange(len(gallery))
+ np.random.shuffle(cands)
+ rand_inds = cands[:num]
+ if not isinstance(gallery, np.ndarray):
+ rand_inds = torch.from_numpy(rand_inds).long().to(gallery.device)
+ return gallery[rand_inds]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/mse_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/mse_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..6048218ad36a8105e7fa182f40fae93ef7c9268f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/mse_loss.py
@@ -0,0 +1,69 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional
+
+import torch.nn as nn
+import torch.nn.functional as F
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from .utils import weighted_loss
+
+
+@weighted_loss
+def mse_loss(pred: Tensor, target: Tensor) -> Tensor:
+ """A Wrapper of MSE loss.
+ Args:
+ pred (Tensor): The prediction.
+ target (Tensor): The learning target of the prediction.
+
+ Returns:
+ Tensor: loss Tensor
+ """
+ return F.mse_loss(pred, target, reduction='none')
+
+
+@MODELS.register_module()
+class MSELoss(nn.Module):
+ """MSELoss.
+
+ Args:
+ reduction (str, optional): The method that reduces the loss to a
+ scalar. Options are "none", "mean" and "sum".
+ loss_weight (float, optional): The weight of the loss. Defaults to 1.0
+ """
+
+ def __init__(self,
+ reduction: str = 'mean',
+ loss_weight: float = 1.0) -> None:
+ super().__init__()
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+
+ def forward(self,
+ pred: Tensor,
+ target: Tensor,
+ weight: Optional[Tensor] = None,
+ avg_factor: Optional[int] = None,
+ reduction_override: Optional[str] = None) -> Tensor:
+ """Forward function of loss.
+
+ Args:
+ pred (Tensor): The prediction.
+ target (Tensor): The learning target of the prediction.
+ weight (Tensor, optional): Weight of the loss for each
+ prediction. Defaults to None.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Defaults to None.
+
+ Returns:
+ Tensor: The calculated loss.
+ """
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ loss = self.loss_weight * mse_loss(
+ pred, target, weight, reduction=reduction, avg_factor=avg_factor)
+ return loss
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/multipos_cross_entropy_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/multipos_cross_entropy_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..a7d1561ed414b7c15412b5e746dff39ca0c53ba1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/multipos_cross_entropy_loss.py
@@ -0,0 +1,100 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional
+
+import torch
+from mmengine.model import BaseModule
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from .utils import weight_reduce_loss
+
+
+@MODELS.register_module()
+class MultiPosCrossEntropyLoss(BaseModule):
+ """multi-positive targets cross entropy loss.
+
+ Args:
+ reduction (str, optional): The method to reduce the loss.
+ Options are "none", "mean" and "sum". Defaults to "mean".
+ loss_weight (float, optional): The weight of loss. Defaults to 1.0.
+ """
+
+ def __init__(self, reduction: str = 'mean', loss_weight: float = 1.0):
+ super(MultiPosCrossEntropyLoss, self).__init__()
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+
+ def multi_pos_cross_entropy(self,
+ pred: Tensor,
+ label: Tensor,
+ weight: Optional[Tensor] = None,
+ reduction: str = 'mean',
+ avg_factor: Optional[float] = None) -> Tensor:
+ """Multi-positive targets cross entropy loss.
+
+ Args:
+ pred (torch.Tensor): The prediction.
+ label (torch.Tensor): The assigned label of the prediction.
+ weight (torch.Tensor): The element-wise weight.
+ reduction (str): Same as built-in losses of PyTorch.
+ avg_factor (float): Average factor when computing
+ the mean of losses.
+
+ Returns:
+ torch.Tensor: Calculated loss
+ """
+
+ pos_inds = (label >= 1)
+ neg_inds = (label == 0)
+ pred_pos = pred * pos_inds.float()
+ pred_neg = pred * neg_inds.float()
+ # use -inf to mask out unwanted elements.
+ pred_pos[neg_inds] = pred_pos[neg_inds] + float('inf')
+ pred_neg[pos_inds] = pred_neg[pos_inds] + float('-inf')
+
+ _pos_expand = torch.repeat_interleave(pred_pos, pred.shape[1], dim=1)
+ _neg_expand = pred_neg.repeat(1, pred.shape[1])
+
+ x = torch.nn.functional.pad((_neg_expand - _pos_expand), (0, 1),
+ 'constant', 0)
+ loss = torch.logsumexp(x, dim=1)
+
+ # apply weights and do the reduction
+ if weight is not None:
+ weight = weight.float()
+ loss = weight_reduce_loss(
+ loss, weight=weight, reduction=reduction, avg_factor=avg_factor)
+
+ return loss
+
+ def forward(self,
+ cls_score: Tensor,
+ label: Tensor,
+ weight: Optional[Tensor] = None,
+ avg_factor: Optional[float] = None,
+ reduction_override: Optional[str] = None,
+ **kwargs) -> Tensor:
+ """Forward function.
+
+ Args:
+ cls_score (torch.Tensor): The classification score.
+ label (torch.Tensor): The assigned label of the prediction.
+ weight (torch.Tensor): The element-wise weight.
+ avg_factor (float): Average factor when computing
+ the mean of losses.
+ reduction_override (str): Same as built-in losses of PyTorch.
+
+ Returns:
+ torch.Tensor: Calculated loss
+ """
+ assert cls_score.size() == label.size()
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ loss_cls = self.loss_weight * self.multi_pos_cross_entropy(
+ cls_score,
+ label,
+ weight,
+ reduction=reduction,
+ avg_factor=avg_factor)
+ return loss_cls
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/pisa_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/pisa_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..b192aa0dbc7eb554755eb2f242eab0ea7f1fc650
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/pisa_loss.py
@@ -0,0 +1,187 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple
+
+import torch
+import torch.nn as nn
+from torch import Tensor
+
+from mmdet.structures.bbox import bbox_overlaps
+from ..task_modules.coders import BaseBBoxCoder
+from ..task_modules.samplers import SamplingResult
+
+
+def isr_p(cls_score: Tensor,
+ bbox_pred: Tensor,
+ bbox_targets: Tuple[Tensor],
+ rois: Tensor,
+ sampling_results: List[SamplingResult],
+ loss_cls: nn.Module,
+ bbox_coder: BaseBBoxCoder,
+ k: float = 2,
+ bias: float = 0,
+ num_class: int = 80) -> tuple:
+ """Importance-based Sample Reweighting (ISR_P), positive part.
+
+ Args:
+ cls_score (Tensor): Predicted classification scores.
+ bbox_pred (Tensor): Predicted bbox deltas.
+ bbox_targets (tuple[Tensor]): A tuple of bbox targets, the are
+ labels, label_weights, bbox_targets, bbox_weights, respectively.
+ rois (Tensor): Anchors (single_stage) in shape (n, 4) or RoIs
+ (two_stage) in shape (n, 5).
+ sampling_results (:obj:`SamplingResult`): Sampling results.
+ loss_cls (:obj:`nn.Module`): Classification loss func of the head.
+ bbox_coder (:obj:`BaseBBoxCoder`): BBox coder of the head.
+ k (float): Power of the non-linear mapping. Defaults to 2.
+ bias (float): Shift of the non-linear mapping. Defaults to 0.
+ num_class (int): Number of classes, defaults to 80.
+
+ Return:
+ tuple([Tensor]): labels, imp_based_label_weights, bbox_targets,
+ bbox_target_weights
+ """
+
+ labels, label_weights, bbox_targets, bbox_weights = bbox_targets
+ pos_label_inds = ((labels >= 0) &
+ (labels < num_class)).nonzero().reshape(-1)
+ pos_labels = labels[pos_label_inds]
+
+ # if no positive samples, return the original targets
+ num_pos = float(pos_label_inds.size(0))
+ if num_pos == 0:
+ return labels, label_weights, bbox_targets, bbox_weights
+
+ # merge pos_assigned_gt_inds of per image to a single tensor
+ gts = list()
+ last_max_gt = 0
+ for i in range(len(sampling_results)):
+ gt_i = sampling_results[i].pos_assigned_gt_inds
+ gts.append(gt_i + last_max_gt)
+ if len(gt_i) != 0:
+ last_max_gt = gt_i.max() + 1
+ gts = torch.cat(gts)
+ assert len(gts) == num_pos
+
+ cls_score = cls_score.detach()
+ bbox_pred = bbox_pred.detach()
+
+ # For single stage detectors, rois here indicate anchors, in shape (N, 4)
+ # For two stage detectors, rois are in shape (N, 5)
+ if rois.size(-1) == 5:
+ pos_rois = rois[pos_label_inds][:, 1:]
+ else:
+ pos_rois = rois[pos_label_inds]
+
+ if bbox_pred.size(-1) > 4:
+ bbox_pred = bbox_pred.view(bbox_pred.size(0), -1, 4)
+ pos_delta_pred = bbox_pred[pos_label_inds, pos_labels].view(-1, 4)
+ else:
+ pos_delta_pred = bbox_pred[pos_label_inds].view(-1, 4)
+
+ # compute iou of the predicted bbox and the corresponding GT
+ pos_delta_target = bbox_targets[pos_label_inds].view(-1, 4)
+ pos_bbox_pred = bbox_coder.decode(pos_rois, pos_delta_pred)
+ target_bbox_pred = bbox_coder.decode(pos_rois, pos_delta_target)
+ ious = bbox_overlaps(pos_bbox_pred, target_bbox_pred, is_aligned=True)
+
+ pos_imp_weights = label_weights[pos_label_inds]
+ # Two steps to compute IoU-HLR. Samples are first sorted by IoU locally,
+ # then sorted again within the same-rank group
+ max_l_num = pos_labels.bincount().max()
+ for label in pos_labels.unique():
+ l_inds = (pos_labels == label).nonzero().view(-1)
+ l_gts = gts[l_inds]
+ for t in l_gts.unique():
+ t_inds = l_inds[l_gts == t]
+ t_ious = ious[t_inds]
+ _, t_iou_rank_idx = t_ious.sort(descending=True)
+ _, t_iou_rank = t_iou_rank_idx.sort()
+ ious[t_inds] += max_l_num - t_iou_rank.float()
+ l_ious = ious[l_inds]
+ _, l_iou_rank_idx = l_ious.sort(descending=True)
+ _, l_iou_rank = l_iou_rank_idx.sort() # IoU-HLR
+ # linearly map HLR to label weights
+ pos_imp_weights[l_inds] *= (max_l_num - l_iou_rank.float()) / max_l_num
+
+ pos_imp_weights = (bias + pos_imp_weights * (1 - bias)).pow(k)
+
+ # normalize to make the new weighted loss value equal to the original loss
+ pos_loss_cls = loss_cls(
+ cls_score[pos_label_inds], pos_labels, reduction_override='none')
+ if pos_loss_cls.dim() > 1:
+ ori_pos_loss_cls = pos_loss_cls * label_weights[pos_label_inds][:,
+ None]
+ new_pos_loss_cls = pos_loss_cls * pos_imp_weights[:, None]
+ else:
+ ori_pos_loss_cls = pos_loss_cls * label_weights[pos_label_inds]
+ new_pos_loss_cls = pos_loss_cls * pos_imp_weights
+ pos_loss_cls_ratio = ori_pos_loss_cls.sum() / new_pos_loss_cls.sum()
+ pos_imp_weights = pos_imp_weights * pos_loss_cls_ratio
+ label_weights[pos_label_inds] = pos_imp_weights
+
+ bbox_targets = labels, label_weights, bbox_targets, bbox_weights
+ return bbox_targets
+
+
+def carl_loss(cls_score: Tensor,
+ labels: Tensor,
+ bbox_pred: Tensor,
+ bbox_targets: Tensor,
+ loss_bbox: nn.Module,
+ k: float = 1,
+ bias: float = 0.2,
+ avg_factor: Optional[int] = None,
+ sigmoid: bool = False,
+ num_class: int = 80) -> dict:
+ """Classification-Aware Regression Loss (CARL).
+
+ Args:
+ cls_score (Tensor): Predicted classification scores.
+ labels (Tensor): Targets of classification.
+ bbox_pred (Tensor): Predicted bbox deltas.
+ bbox_targets (Tensor): Target of bbox regression.
+ loss_bbox (func): Regression loss func of the head.
+ bbox_coder (obj): BBox coder of the head.
+ k (float): Power of the non-linear mapping. Defaults to 1.
+ bias (float): Shift of the non-linear mapping. Defaults to 0.2.
+ avg_factor (int, optional): Average factor used in regression loss.
+ sigmoid (bool): Activation of the classification score.
+ num_class (int): Number of classes, defaults to 80.
+
+ Return:
+ dict: CARL loss dict.
+ """
+ pos_label_inds = ((labels >= 0) &
+ (labels < num_class)).nonzero().reshape(-1)
+ if pos_label_inds.numel() == 0:
+ return dict(loss_carl=cls_score.sum()[None] * 0.)
+ pos_labels = labels[pos_label_inds]
+
+ # multiply pos_cls_score with the corresponding bbox weight
+ # and remain gradient
+ if sigmoid:
+ pos_cls_score = cls_score.sigmoid()[pos_label_inds, pos_labels]
+ else:
+ pos_cls_score = cls_score.softmax(-1)[pos_label_inds, pos_labels]
+ carl_loss_weights = (bias + (1 - bias) * pos_cls_score).pow(k)
+
+ # normalize carl_loss_weight to make its sum equal to num positive
+ num_pos = float(pos_cls_score.size(0))
+ weight_ratio = num_pos / carl_loss_weights.sum()
+ carl_loss_weights *= weight_ratio
+
+ if avg_factor is None:
+ avg_factor = bbox_targets.size(0)
+ # if is class agnostic, bbox pred is in shape (N, 4)
+ # otherwise, bbox pred is in shape (N, #classes, 4)
+ if bbox_pred.size(-1) > 4:
+ bbox_pred = bbox_pred.view(bbox_pred.size(0), -1, 4)
+ pos_bbox_preds = bbox_pred[pos_label_inds, pos_labels]
+ else:
+ pos_bbox_preds = bbox_pred[pos_label_inds]
+ ori_loss_reg = loss_bbox(
+ pos_bbox_preds,
+ bbox_targets[pos_label_inds],
+ reduction_override='none') / avg_factor
+ loss_carl = (ori_loss_reg * carl_loss_weights[:, None]).sum()
+ return dict(loss_carl=loss_carl[None])
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/seesaw_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/seesaw_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..4dec62b0afdc01e848e0c7f53ba0b6b10b899ea4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/seesaw_loss.py
@@ -0,0 +1,278 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, Optional, Tuple, Union
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from .accuracy import accuracy
+from .cross_entropy_loss import cross_entropy
+from .utils import weight_reduce_loss
+
+
+def seesaw_ce_loss(cls_score: Tensor,
+ labels: Tensor,
+ label_weights: Tensor,
+ cum_samples: Tensor,
+ num_classes: int,
+ p: float,
+ q: float,
+ eps: float,
+ reduction: str = 'mean',
+ avg_factor: Optional[int] = None) -> Tensor:
+ """Calculate the Seesaw CrossEntropy loss.
+
+ Args:
+ cls_score (Tensor): The prediction with shape (N, C),
+ C is the number of classes.
+ labels (Tensor): The learning label of the prediction.
+ label_weights (Tensor): Sample-wise loss weight.
+ cum_samples (Tensor): Cumulative samples for each category.
+ num_classes (int): The number of classes.
+ p (float): The ``p`` in the mitigation factor.
+ q (float): The ``q`` in the compenstation factor.
+ eps (float): The minimal value of divisor to smooth
+ the computation of compensation factor
+ reduction (str, optional): The method used to reduce the loss.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+
+ Returns:
+ Tensor: The calculated loss
+ """
+ assert cls_score.size(-1) == num_classes
+ assert len(cum_samples) == num_classes
+
+ onehot_labels = F.one_hot(labels, num_classes)
+ seesaw_weights = cls_score.new_ones(onehot_labels.size())
+
+ # mitigation factor
+ if p > 0:
+ sample_ratio_matrix = cum_samples[None, :].clamp(
+ min=1) / cum_samples[:, None].clamp(min=1)
+ index = (sample_ratio_matrix < 1.0).float()
+ sample_weights = sample_ratio_matrix.pow(p) * index + (1 - index)
+ mitigation_factor = sample_weights[labels.long(), :]
+ seesaw_weights = seesaw_weights * mitigation_factor
+
+ # compensation factor
+ if q > 0:
+ scores = F.softmax(cls_score.detach(), dim=1)
+ self_scores = scores[
+ torch.arange(0, len(scores)).to(scores.device).long(),
+ labels.long()]
+ score_matrix = scores / self_scores[:, None].clamp(min=eps)
+ index = (score_matrix > 1.0).float()
+ compensation_factor = score_matrix.pow(q) * index + (1 - index)
+ seesaw_weights = seesaw_weights * compensation_factor
+
+ cls_score = cls_score + (seesaw_weights.log() * (1 - onehot_labels))
+
+ loss = F.cross_entropy(cls_score, labels, weight=None, reduction='none')
+
+ if label_weights is not None:
+ label_weights = label_weights.float()
+ loss = weight_reduce_loss(
+ loss, weight=label_weights, reduction=reduction, avg_factor=avg_factor)
+ return loss
+
+
+@MODELS.register_module()
+class SeesawLoss(nn.Module):
+ """
+ Seesaw Loss for Long-Tailed Instance Segmentation (CVPR 2021)
+ arXiv: https://arxiv.org/abs/2008.10032
+
+ Args:
+ use_sigmoid (bool, optional): Whether the prediction uses sigmoid
+ of softmax. Only False is supported.
+ p (float, optional): The ``p`` in the mitigation factor.
+ Defaults to 0.8.
+ q (float, optional): The ``q`` in the compenstation factor.
+ Defaults to 2.0.
+ num_classes (int, optional): The number of classes.
+ Default to 1203 for LVIS v1 dataset.
+ eps (float, optional): The minimal value of divisor to smooth
+ the computation of compensation factor
+ reduction (str, optional): The method that reduces the loss to a
+ scalar. Options are "none", "mean" and "sum".
+ loss_weight (float, optional): The weight of the loss. Defaults to 1.0
+ return_dict (bool, optional): Whether return the losses as a dict.
+ Default to True.
+ """
+
+ def __init__(self,
+ use_sigmoid: bool = False,
+ p: float = 0.8,
+ q: float = 2.0,
+ num_classes: int = 1203,
+ eps: float = 1e-2,
+ reduction: str = 'mean',
+ loss_weight: float = 1.0,
+ return_dict: bool = True) -> None:
+ super().__init__()
+ assert not use_sigmoid
+ self.use_sigmoid = False
+ self.p = p
+ self.q = q
+ self.num_classes = num_classes
+ self.eps = eps
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+ self.return_dict = return_dict
+
+ # 0 for pos, 1 for neg
+ self.cls_criterion = seesaw_ce_loss
+
+ # cumulative samples for each category
+ self.register_buffer(
+ 'cum_samples',
+ torch.zeros(self.num_classes + 1, dtype=torch.float))
+
+ # custom output channels of the classifier
+ self.custom_cls_channels = True
+ # custom activation of cls_score
+ self.custom_activation = True
+ # custom accuracy of the classsifier
+ self.custom_accuracy = True
+
+ def _split_cls_score(self, cls_score: Tensor) -> Tuple[Tensor, Tensor]:
+ """split cls_score.
+
+ Args:
+ cls_score (Tensor): The prediction with shape (N, C + 2).
+
+ Returns:
+ Tuple[Tensor, Tensor]: The score for classes and objectness,
+ respectively
+ """
+ # split cls_score to cls_score_classes and cls_score_objectness
+ assert cls_score.size(-1) == self.num_classes + 2
+ cls_score_classes = cls_score[..., :-2]
+ cls_score_objectness = cls_score[..., -2:]
+ return cls_score_classes, cls_score_objectness
+
+ def get_cls_channels(self, num_classes: int) -> int:
+ """Get custom classification channels.
+
+ Args:
+ num_classes (int): The number of classes.
+
+ Returns:
+ int: The custom classification channels.
+ """
+ assert num_classes == self.num_classes
+ return num_classes + 2
+
+ def get_activation(self, cls_score: Tensor) -> Tensor:
+ """Get custom activation of cls_score.
+
+ Args:
+ cls_score (Tensor): The prediction with shape (N, C + 2).
+
+ Returns:
+ Tensor: The custom activation of cls_score with shape
+ (N, C + 1).
+ """
+ cls_score_classes, cls_score_objectness = self._split_cls_score(
+ cls_score)
+ score_classes = F.softmax(cls_score_classes, dim=-1)
+ score_objectness = F.softmax(cls_score_objectness, dim=-1)
+ score_pos = score_objectness[..., [0]]
+ score_neg = score_objectness[..., [1]]
+ score_classes = score_classes * score_pos
+ scores = torch.cat([score_classes, score_neg], dim=-1)
+ return scores
+
+ def get_accuracy(self, cls_score: Tensor,
+ labels: Tensor) -> Dict[str, Tensor]:
+ """Get custom accuracy w.r.t. cls_score and labels.
+
+ Args:
+ cls_score (Tensor): The prediction with shape (N, C + 2).
+ labels (Tensor): The learning label of the prediction.
+
+ Returns:
+ Dict [str, Tensor]: The accuracy for objectness and classes,
+ respectively.
+ """
+ pos_inds = labels < self.num_classes
+ obj_labels = (labels == self.num_classes).long()
+ cls_score_classes, cls_score_objectness = self._split_cls_score(
+ cls_score)
+ acc_objectness = accuracy(cls_score_objectness, obj_labels)
+ acc_classes = accuracy(cls_score_classes[pos_inds], labels[pos_inds])
+ acc = dict()
+ acc['acc_objectness'] = acc_objectness
+ acc['acc_classes'] = acc_classes
+ return acc
+
+ def forward(
+ self,
+ cls_score: Tensor,
+ labels: Tensor,
+ label_weights: Optional[Tensor] = None,
+ avg_factor: Optional[int] = None,
+ reduction_override: Optional[str] = None
+ ) -> Union[Tensor, Dict[str, Tensor]]:
+ """Forward function.
+
+ Args:
+ cls_score (Tensor): The prediction with shape (N, C + 2).
+ labels (Tensor): The learning label of the prediction.
+ label_weights (Tensor, optional): Sample-wise loss weight.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ reduction (str, optional): The method used to reduce the loss.
+ Options are "none", "mean" and "sum".
+
+ Returns:
+ Tensor | Dict [str, Tensor]:
+ if return_dict == False: The calculated loss |
+ if return_dict == True: The dict of calculated losses
+ for objectness and classes, respectively.
+ """
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ assert cls_score.size(-1) == self.num_classes + 2
+ pos_inds = labels < self.num_classes
+ # 0 for pos, 1 for neg
+ obj_labels = (labels == self.num_classes).long()
+
+ # accumulate the samples for each category
+ unique_labels = labels.unique()
+ for u_l in unique_labels:
+ inds_ = labels == u_l.item()
+ self.cum_samples[u_l] += inds_.sum()
+
+ if label_weights is not None:
+ label_weights = label_weights.float()
+ else:
+ label_weights = labels.new_ones(labels.size(), dtype=torch.float)
+
+ cls_score_classes, cls_score_objectness = self._split_cls_score(
+ cls_score)
+ # calculate loss_cls_classes (only need pos samples)
+ if pos_inds.sum() > 0:
+ loss_cls_classes = self.loss_weight * self.cls_criterion(
+ cls_score_classes[pos_inds], labels[pos_inds],
+ label_weights[pos_inds], self.cum_samples[:self.num_classes],
+ self.num_classes, self.p, self.q, self.eps, reduction,
+ avg_factor)
+ else:
+ loss_cls_classes = cls_score_classes[pos_inds].sum()
+ # calculate loss_cls_objectness
+ loss_cls_objectness = self.loss_weight * cross_entropy(
+ cls_score_objectness, obj_labels, label_weights, reduction,
+ avg_factor)
+
+ if self.return_dict:
+ loss_cls = dict()
+ loss_cls['loss_cls_objectness'] = loss_cls_objectness
+ loss_cls['loss_cls_classes'] = loss_cls_classes
+ else:
+ loss_cls = loss_cls_classes + loss_cls_objectness
+ return loss_cls
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/smooth_l1_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/smooth_l1_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..102f9780706172a44ade2ebe1709c7a1e847db7c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/smooth_l1_loss.py
@@ -0,0 +1,165 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional
+
+import torch
+import torch.nn as nn
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from .utils import weighted_loss
+
+
+@weighted_loss
+def smooth_l1_loss(pred: Tensor, target: Tensor, beta: float = 1.0) -> Tensor:
+ """Smooth L1 loss.
+
+ Args:
+ pred (Tensor): The prediction.
+ target (Tensor): The learning target of the prediction.
+ beta (float, optional): The threshold in the piecewise function.
+ Defaults to 1.0.
+
+ Returns:
+ Tensor: Calculated loss
+ """
+ assert beta > 0
+ if target.numel() == 0:
+ return pred.sum() * 0
+
+ assert pred.size() == target.size()
+ diff = torch.abs(pred - target)
+ loss = torch.where(diff < beta, 0.5 * diff * diff / beta,
+ diff - 0.5 * beta)
+ return loss
+
+
+@weighted_loss
+def l1_loss(pred: Tensor, target: Tensor) -> Tensor:
+ """L1 loss.
+
+ Args:
+ pred (Tensor): The prediction.
+ target (Tensor): The learning target of the prediction.
+
+ Returns:
+ Tensor: Calculated loss
+ """
+ if target.numel() == 0:
+ return pred.sum() * 0
+
+ assert pred.size() == target.size()
+ loss = torch.abs(pred - target)
+ return loss
+
+
+@MODELS.register_module()
+class SmoothL1Loss(nn.Module):
+ """Smooth L1 loss.
+
+ Args:
+ beta (float, optional): The threshold in the piecewise function.
+ Defaults to 1.0.
+ reduction (str, optional): The method to reduce the loss.
+ Options are "none", "mean" and "sum". Defaults to "mean".
+ loss_weight (float, optional): The weight of loss.
+ """
+
+ def __init__(self,
+ beta: float = 1.0,
+ reduction: str = 'mean',
+ loss_weight: float = 1.0) -> None:
+ super().__init__()
+ self.beta = beta
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+
+ def forward(self,
+ pred: Tensor,
+ target: Tensor,
+ weight: Optional[Tensor] = None,
+ avg_factor: Optional[int] = None,
+ reduction_override: Optional[str] = None,
+ **kwargs) -> Tensor:
+ """Forward function.
+
+ Args:
+ pred (Tensor): The prediction.
+ target (Tensor): The learning target of the prediction.
+ weight (Tensor, optional): The weight of loss for each
+ prediction. Defaults to None.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Defaults to None.
+
+ Returns:
+ Tensor: Calculated loss
+ """
+ if weight is not None and not torch.any(weight > 0):
+ if pred.dim() == weight.dim() + 1:
+ weight = weight.unsqueeze(1)
+ return (pred * weight).sum()
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ loss_bbox = self.loss_weight * smooth_l1_loss(
+ pred,
+ target,
+ weight,
+ beta=self.beta,
+ reduction=reduction,
+ avg_factor=avg_factor,
+ **kwargs)
+ return loss_bbox
+
+
+@MODELS.register_module()
+class L1Loss(nn.Module):
+ """L1 loss.
+
+ Args:
+ reduction (str, optional): The method to reduce the loss.
+ Options are "none", "mean" and "sum".
+ loss_weight (float, optional): The weight of loss.
+ """
+
+ def __init__(self,
+ reduction: str = 'mean',
+ loss_weight: float = 1.0) -> None:
+ super().__init__()
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+
+ def forward(self,
+ pred: Tensor,
+ target: Tensor,
+ weight: Optional[Tensor] = None,
+ avg_factor: Optional[int] = None,
+ reduction_override: Optional[str] = None) -> Tensor:
+ """Forward function.
+
+ Args:
+ pred (Tensor): The prediction.
+ target (Tensor): The learning target of the prediction.
+ weight (Tensor, optional): The weight of loss for each
+ prediction. Defaults to None.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Defaults to None.
+
+ Returns:
+ Tensor: Calculated loss
+ """
+ if weight is not None and not torch.any(weight > 0):
+ if pred.dim() == weight.dim() + 1:
+ weight = weight.unsqueeze(1)
+ return (pred * weight).sum()
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ loss_bbox = self.loss_weight * l1_loss(
+ pred, target, weight, reduction=reduction, avg_factor=avg_factor)
+ return loss_bbox
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/triplet_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/triplet_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..4528239beb4bf122fa1a05ee2ce21cb1cb144bde
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/triplet_loss.py
@@ -0,0 +1,88 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+import torch.nn as nn
+from mmengine.model import BaseModule
+
+from mmdet.registry import MODELS
+
+
+@MODELS.register_module()
+class TripletLoss(BaseModule):
+ """Triplet loss with hard positive/negative mining.
+
+ Reference:
+ Hermans et al. In Defense of the Triplet Loss for
+ Person Re-Identification. arXiv:1703.07737.
+ Imported from ``_.
+ Args:
+ margin (float, optional): Margin for triplet loss. Defaults to 0.3.
+ loss_weight (float, optional): Weight of the loss. Defaults to 1.0.
+ hard_mining (bool, optional): Whether to perform hard mining.
+ Defaults to True.
+ """
+
+ def __init__(self,
+ margin: float = 0.3,
+ loss_weight: float = 1.0,
+ hard_mining=True):
+ super(TripletLoss, self).__init__()
+ self.margin = margin
+ self.ranking_loss = nn.MarginRankingLoss(margin=margin)
+ self.loss_weight = loss_weight
+ self.hard_mining = hard_mining
+
+ def hard_mining_triplet_loss_forward(
+ self, inputs: torch.Tensor,
+ targets: torch.LongTensor) -> torch.Tensor:
+ """
+ Args:
+ inputs (torch.Tensor): feature matrix with shape
+ (batch_size, feat_dim).
+ targets (torch.LongTensor): ground truth labels with shape
+ (batch_size).
+
+ Returns:
+ torch.Tensor: triplet loss with hard mining.
+ """
+
+ batch_size = inputs.size(0)
+
+ # Compute Euclidean distance
+ dist = torch.pow(inputs, 2).sum(
+ dim=1, keepdim=True).expand(batch_size, batch_size)
+ dist = dist + dist.t()
+ dist.addmm_(inputs, inputs.t(), beta=1, alpha=-2)
+ dist = dist.clamp(min=1e-12).sqrt() # for numerical stability
+
+ # For each anchor, find the furthest positive sample
+ # and nearest negative sample in the embedding space
+ mask = targets.expand(batch_size, batch_size).eq(
+ targets.expand(batch_size, batch_size).t())
+ dist_ap, dist_an = [], []
+ for i in range(batch_size):
+ dist_ap.append(dist[i][mask[i]].max().unsqueeze(0))
+ dist_an.append(dist[i][mask[i] == 0].min().unsqueeze(0))
+ dist_ap = torch.cat(dist_ap)
+ dist_an = torch.cat(dist_an)
+
+ # Compute ranking hinge loss
+ y = torch.ones_like(dist_an)
+ return self.loss_weight * self.ranking_loss(dist_an, dist_ap, y)
+
+ def forward(self, inputs: torch.Tensor,
+ targets: torch.LongTensor) -> torch.Tensor:
+ """
+ Args:
+ inputs (torch.Tensor): feature matrix with shape
+ (batch_size, feat_dim).
+ targets (torch.LongTensor): ground truth labels with shape
+ (num_classes).
+
+ Returns:
+ torch.Tensor: triplet loss.
+ """
+ if self.hard_mining:
+ return self.hard_mining_triplet_loss_forward(inputs, targets)
+ else:
+ raise NotImplementedError()
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/utils.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/utils.py
new file mode 100644
index 0000000000000000000000000000000000000000..5e6e7859f353f3e5456f0cfc1f66b4b0ad535427
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/utils.py
@@ -0,0 +1,125 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import functools
+from typing import Callable, Optional
+
+import torch
+import torch.nn.functional as F
+from torch import Tensor
+
+
+def reduce_loss(loss: Tensor, reduction: str) -> Tensor:
+ """Reduce loss as specified.
+
+ Args:
+ loss (Tensor): Elementwise loss tensor.
+ reduction (str): Options are "none", "mean" and "sum".
+
+ Return:
+ Tensor: Reduced loss tensor.
+ """
+ reduction_enum = F._Reduction.get_enum(reduction)
+ # none: 0, elementwise_mean:1, sum: 2
+ if reduction_enum == 0:
+ return loss
+ elif reduction_enum == 1:
+ return loss.mean()
+ elif reduction_enum == 2:
+ return loss.sum()
+
+
+def weight_reduce_loss(loss: Tensor,
+ weight: Optional[Tensor] = None,
+ reduction: str = 'mean',
+ avg_factor: Optional[float] = None) -> Tensor:
+ """Apply element-wise weight and reduce loss.
+
+ Args:
+ loss (Tensor): Element-wise loss.
+ weight (Optional[Tensor], optional): Element-wise weights.
+ Defaults to None.
+ reduction (str, optional): Same as built-in losses of PyTorch.
+ Defaults to 'mean'.
+ avg_factor (Optional[float], optional): Average factor when
+ computing the mean of losses. Defaults to None.
+
+ Returns:
+ Tensor: Processed loss values.
+ """
+ # if weight is specified, apply element-wise weight
+ if weight is not None:
+ loss = loss * weight
+
+ # if avg_factor is not specified, just reduce the loss
+ if avg_factor is None:
+ loss = reduce_loss(loss, reduction)
+ else:
+ # if reduction is mean, then average the loss by avg_factor
+ if reduction == 'mean':
+ # Avoid causing ZeroDivisionError when avg_factor is 0.0,
+ # i.e., all labels of an image belong to ignore index.
+ eps = torch.finfo(torch.float32).eps
+ loss = loss.sum() / (avg_factor + eps)
+ # if reduction is 'none', then do nothing, otherwise raise an error
+ elif reduction != 'none':
+ raise ValueError('avg_factor can not be used with reduction="sum"')
+ return loss
+
+
+def weighted_loss(loss_func: Callable) -> Callable:
+ """Create a weighted version of a given loss function.
+
+ To use this decorator, the loss function must have the signature like
+ `loss_func(pred, target, **kwargs)`. The function only needs to compute
+ element-wise loss without any reduction. This decorator will add weight
+ and reduction arguments to the function. The decorated function will have
+ the signature like `loss_func(pred, target, weight=None, reduction='mean',
+ avg_factor=None, **kwargs)`.
+
+ :Example:
+
+ >>> import torch
+ >>> @weighted_loss
+ >>> def l1_loss(pred, target):
+ >>> return (pred - target).abs()
+
+ >>> pred = torch.Tensor([0, 2, 3])
+ >>> target = torch.Tensor([1, 1, 1])
+ >>> weight = torch.Tensor([1, 0, 1])
+
+ >>> l1_loss(pred, target)
+ tensor(1.3333)
+ >>> l1_loss(pred, target, weight)
+ tensor(1.)
+ >>> l1_loss(pred, target, reduction='none')
+ tensor([1., 1., 2.])
+ >>> l1_loss(pred, target, weight, avg_factor=2)
+ tensor(1.5000)
+ """
+
+ @functools.wraps(loss_func)
+ def wrapper(pred: Tensor,
+ target: Tensor,
+ weight: Optional[Tensor] = None,
+ reduction: str = 'mean',
+ avg_factor: Optional[int] = None,
+ **kwargs) -> Tensor:
+ """
+ Args:
+ pred (Tensor): The prediction.
+ target (Tensor): Target bboxes.
+ weight (Optional[Tensor], optional): The weight of loss for each
+ prediction. Defaults to None.
+ reduction (str, optional): Options are "none", "mean" and "sum".
+ Defaults to 'mean'.
+ avg_factor (Optional[int], optional): Average factor that is used
+ to average the loss. Defaults to None.
+
+ Returns:
+ Tensor: Loss tensor.
+ """
+ # get element-wise loss
+ loss = loss_func(pred, target, **kwargs)
+ loss = weight_reduce_loss(loss, weight, reduction, avg_factor)
+ return loss
+
+ return wrapper
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/varifocal_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/varifocal_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..58ab167352e1ae32566f5e731339966d5fd10759
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/losses/varifocal_loss.py
@@ -0,0 +1,141 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional
+
+import torch.nn as nn
+import torch.nn.functional as F
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from .utils import weight_reduce_loss
+
+
+def varifocal_loss(pred: Tensor,
+ target: Tensor,
+ weight: Optional[Tensor] = None,
+ alpha: float = 0.75,
+ gamma: float = 2.0,
+ iou_weighted: bool = True,
+ reduction: str = 'mean',
+ avg_factor: Optional[int] = None) -> Tensor:
+ """`Varifocal Loss `_
+
+ Args:
+ pred (Tensor): The prediction with shape (N, C), C is the
+ number of classes.
+ target (Tensor): The learning target of the iou-aware
+ classification score with shape (N, C), C is the number of classes.
+ weight (Tensor, optional): The weight of loss for each
+ prediction. Defaults to None.
+ alpha (float, optional): A balance factor for the negative part of
+ Varifocal Loss, which is different from the alpha of Focal Loss.
+ Defaults to 0.75.
+ gamma (float, optional): The gamma for calculating the modulating
+ factor. Defaults to 2.0.
+ iou_weighted (bool, optional): Whether to weight the loss of the
+ positive example with the iou target. Defaults to True.
+ reduction (str, optional): The method used to reduce the loss into
+ a scalar. Defaults to 'mean'. Options are "none", "mean" and
+ "sum".
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+
+ Returns:
+ Tensor: Loss tensor.
+ """
+ # pred and target should be of the same size
+ assert pred.size() == target.size()
+ pred_sigmoid = pred.sigmoid()
+ target = target.type_as(pred)
+ if iou_weighted:
+ focal_weight = target * (target > 0.0).float() + \
+ alpha * (pred_sigmoid - target).abs().pow(gamma) * \
+ (target <= 0.0).float()
+ else:
+ focal_weight = (target > 0.0).float() + \
+ alpha * (pred_sigmoid - target).abs().pow(gamma) * \
+ (target <= 0.0).float()
+ loss = F.binary_cross_entropy_with_logits(
+ pred, target, reduction='none') * focal_weight
+ loss = weight_reduce_loss(loss, weight, reduction, avg_factor)
+ return loss
+
+
+@MODELS.register_module()
+class VarifocalLoss(nn.Module):
+
+ def __init__(self,
+ use_sigmoid: bool = True,
+ alpha: float = 0.75,
+ gamma: float = 2.0,
+ iou_weighted: bool = True,
+ reduction: str = 'mean',
+ loss_weight: float = 1.0) -> None:
+ """`Varifocal Loss `_
+
+ Args:
+ use_sigmoid (bool, optional): Whether the prediction is
+ used for sigmoid or softmax. Defaults to True.
+ alpha (float, optional): A balance factor for the negative part of
+ Varifocal Loss, which is different from the alpha of Focal
+ Loss. Defaults to 0.75.
+ gamma (float, optional): The gamma for calculating the modulating
+ factor. Defaults to 2.0.
+ iou_weighted (bool, optional): Whether to weight the loss of the
+ positive examples with the iou target. Defaults to True.
+ reduction (str, optional): The method used to reduce the loss into
+ a scalar. Defaults to 'mean'. Options are "none", "mean" and
+ "sum".
+ loss_weight (float, optional): Weight of loss. Defaults to 1.0.
+ """
+ super().__init__()
+ assert use_sigmoid is True, \
+ 'Only sigmoid varifocal loss supported now.'
+ assert alpha >= 0.0
+ self.use_sigmoid = use_sigmoid
+ self.alpha = alpha
+ self.gamma = gamma
+ self.iou_weighted = iou_weighted
+ self.reduction = reduction
+ self.loss_weight = loss_weight
+
+ def forward(self,
+ pred: Tensor,
+ target: Tensor,
+ weight: Optional[Tensor] = None,
+ avg_factor: Optional[int] = None,
+ reduction_override: Optional[str] = None) -> Tensor:
+ """Forward function.
+
+ Args:
+ pred (Tensor): The prediction with shape (N, C), C is the
+ number of classes.
+ target (Tensor): The learning target of the iou-aware
+ classification score with shape (N, C), C is
+ the number of classes.
+ weight (Tensor, optional): The weight of loss for each
+ prediction. Defaults to None.
+ avg_factor (int, optional): Average factor that is used to average
+ the loss. Defaults to None.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Options are "none", "mean" and "sum".
+
+ Returns:
+ Tensor: The calculated loss
+ """
+ assert reduction_override in (None, 'none', 'mean', 'sum')
+ reduction = (
+ reduction_override if reduction_override else self.reduction)
+ if self.use_sigmoid:
+ loss_cls = self.loss_weight * varifocal_loss(
+ pred,
+ target,
+ weight,
+ alpha=self.alpha,
+ gamma=self.gamma,
+ iou_weighted=self.iou_weighted,
+ reduction=reduction,
+ avg_factor=avg_factor)
+ else:
+ raise NotImplementedError
+ return loss_cls
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..1bd3c8d3ba53daad736e05b5d29a6abb377fd595
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/__init__.py
@@ -0,0 +1,11 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .base import BaseMOTModel
+from .bytetrack import ByteTrack
+from .deep_sort import DeepSORT
+from .ocsort import OCSORT
+from .qdtrack import QDTrack
+from .strongsort import StrongSORT
+
+__all__ = [
+ 'BaseMOTModel', 'ByteTrack', 'QDTrack', 'DeepSORT', 'StrongSORT', 'OCSORT'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/base.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/base.py
new file mode 100644
index 0000000000000000000000000000000000000000..9981417924af3970319b0cbe6a9cc8d8a1095451
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/base.py
@@ -0,0 +1,147 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from abc import ABCMeta, abstractmethod
+from typing import Dict, List, Tuple, Union
+
+from mmengine.model import BaseModel
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import OptTrackSampleList, TrackSampleList
+from mmdet.utils import OptConfigType, OptMultiConfig
+
+
+@MODELS.register_module()
+class BaseMOTModel(BaseModel, metaclass=ABCMeta):
+ """Base class for multiple object tracking.
+
+ Args:
+ data_preprocessor (dict or ConfigDict, optional): The pre-process
+ config of :class:`TrackDataPreprocessor`. it usually includes,
+ ``pad_size_divisor``, ``pad_value``, ``mean`` and ``std``.
+ init_cfg (dict or list[dict]): Initialization config dict.
+ """
+
+ def __init__(self,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ data_preprocessor=data_preprocessor, init_cfg=init_cfg)
+
+ def freeze_module(self, module: Union[List[str], Tuple[str], str]) -> None:
+ """Freeze module during training."""
+ if isinstance(module, str):
+ modules = [module]
+ else:
+ if not (isinstance(module, list) or isinstance(module, tuple)):
+ raise TypeError('module must be a str or a list.')
+ else:
+ modules = module
+ for module in modules:
+ m = getattr(self, module)
+ m.eval()
+ for param in m.parameters():
+ param.requires_grad = False
+
+ @property
+ def with_detector(self) -> bool:
+ """bool: whether the framework has a detector."""
+ return hasattr(self, 'detector') and self.detector is not None
+
+ @property
+ def with_reid(self) -> bool:
+ """bool: whether the framework has a reid model."""
+ return hasattr(self, 'reid') and self.reid is not None
+
+ @property
+ def with_motion(self) -> bool:
+ """bool: whether the framework has a motion model."""
+ return hasattr(self, 'motion') and self.motion is not None
+
+ @property
+ def with_track_head(self) -> bool:
+ """bool: whether the framework has a track_head."""
+ return hasattr(self, 'track_head') and self.track_head is not None
+
+ @property
+ def with_tracker(self) -> bool:
+ """bool: whether the framework has a tracker."""
+ return hasattr(self, 'tracker') and self.tracker is not None
+
+ def forward(self,
+ inputs: Dict[str, Tensor],
+ data_samples: OptTrackSampleList = None,
+ mode: str = 'predict',
+ **kwargs):
+ """The unified entry for a forward process in both training and test.
+
+ The method should accept three modes: "tensor", "predict" and "loss":
+
+ - "tensor": Forward the whole network and return tensor or tuple of
+ tensor without any post-processing, same as a common nn.Module.
+ - "predict": Forward and return the predictions, which are fully
+ processed to a list of :obj:`TrackDataSample`.
+ - "loss": Forward and return a dict of losses according to the given
+ inputs and data samples.
+
+ Note that this method doesn't handle neither back propagation nor
+ optimizer updating, which are done in the :meth:`train_step`.
+
+ Args:
+ inputs (Dict[str, Tensor]): of shape (N, T, C, H, W)
+ encoding input images. Typically these should be mean centered
+ and std scaled. The N denotes batch size. The T denotes the
+ number of key/reference frames.
+ - img (Tensor) : The key images.
+ - ref_img (Tensor): The reference images.
+ data_samples (list[:obj:`TrackDataSample`], optional): The
+ annotation data of every samples. Defaults to None.
+ mode (str): Return what kind of value. Defaults to 'predict'.
+
+ Returns:
+ The return type depends on ``mode``.
+
+ - If ``mode="tensor"``, return a tensor or a tuple of tensor.
+ - If ``mode="predict"``, return a list of :obj:`TrackDataSample`.
+ - If ``mode="loss"``, return a dict of tensor.
+ """
+ if mode == 'loss':
+ return self.loss(inputs, data_samples, **kwargs)
+ elif mode == 'predict':
+ return self.predict(inputs, data_samples, **kwargs)
+ elif mode == 'tensor':
+ return self._forward(inputs, data_samples, **kwargs)
+ else:
+ raise RuntimeError(f'Invalid mode "{mode}". '
+ 'Only supports loss, predict and tensor mode')
+
+ @abstractmethod
+ def loss(self, inputs: Dict[str, Tensor], data_samples: TrackSampleList,
+ **kwargs) -> Union[dict, tuple]:
+ """Calculate losses from a batch of inputs and data samples."""
+ pass
+
+ @abstractmethod
+ def predict(self, inputs: Dict[str, Tensor], data_samples: TrackSampleList,
+ **kwargs) -> TrackSampleList:
+ """Predict results from a batch of inputs and data samples with post-
+ processing."""
+ pass
+
+ def _forward(self,
+ inputs: Dict[str, Tensor],
+ data_samples: OptTrackSampleList = None,
+ **kwargs):
+ """Network forward process. Usually includes backbone, neck and head
+ forward without any post-processing.
+
+ Args:
+ inputs (Dict[str, Tensor]): of shape (N, T, C, H, W).
+ data_samples (List[:obj:`TrackDataSample`], optional): The
+ Data Samples. It usually includes information such as
+ `gt_instance`.
+
+ Returns:
+ tuple[list]: A tuple of features from ``head`` forward.
+ """
+ raise NotImplementedError(
+ "_forward function (namely 'tensor' mode) is not supported now")
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/bytetrack.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/bytetrack.py
new file mode 100644
index 0000000000000000000000000000000000000000..8a3bb867cb284aad9854de44b2942341a4a33be8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/bytetrack.py
@@ -0,0 +1,94 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, Optional
+
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList, TrackSampleList
+from mmdet.utils import OptConfigType, OptMultiConfig
+from .base import BaseMOTModel
+
+
+@MODELS.register_module()
+class ByteTrack(BaseMOTModel):
+ """ByteTrack: Multi-Object Tracking by Associating Every Detection Box.
+
+ This multi object tracker is the implementation of `ByteTrack
+ `_.
+
+ Args:
+ detector (dict): Configuration of detector. Defaults to None.
+ tracker (dict): Configuration of tracker. Defaults to None.
+ data_preprocessor (dict or ConfigDict, optional): The pre-process
+ config of :class:`TrackDataPreprocessor`. it usually includes,
+ ``pad_size_divisor``, ``pad_value``, ``mean`` and ``std``.
+ init_cfg (dict or list[dict]): Configuration of initialization.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ detector: Optional[dict] = None,
+ tracker: Optional[dict] = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None):
+ super().__init__(data_preprocessor, init_cfg)
+
+ if detector is not None:
+ self.detector = MODELS.build(detector)
+
+ if tracker is not None:
+ self.tracker = MODELS.build(tracker)
+
+ def loss(self, inputs: Tensor, data_samples: SampleList, **kwargs) -> dict:
+ """Calculate losses from a batch of inputs and data samples.
+
+ Args:
+ inputs (Tensor): of shape (N, C, H, W) encoding
+ input images. Typically these should be mean centered and std
+ scaled. The N denotes batch size
+ data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance`.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ return self.detector.loss(inputs, data_samples, **kwargs)
+
+ def predict(self, inputs: Dict[str, Tensor], data_samples: TrackSampleList,
+ **kwargs) -> TrackSampleList:
+ """Predict results from a video and data samples with post-processing.
+
+ Args:
+ inputs (Tensor): of shape (N, T, C, H, W) encoding
+ input images. The N denotes batch size.
+ The T denotes the number of frames in a video.
+ data_samples (list[:obj:`TrackDataSample`]): The batch
+ data samples. It usually includes information such
+ as `video_data_samples`.
+ Returns:
+ TrackSampleList: Tracking results of the inputs.
+ """
+ assert inputs.dim() == 5, 'The img must be 5D Tensor (N, T, C, H, W).'
+ assert inputs.size(0) == 1, \
+ 'Bytetrack inference only support ' \
+ '1 batch size per gpu for now.'
+
+ assert len(data_samples) == 1, \
+ 'Bytetrack inference only support 1 batch size per gpu for now.'
+
+ track_data_sample = data_samples[0]
+ video_len = len(track_data_sample)
+
+ for frame_id in range(video_len):
+ img_data_sample = track_data_sample[frame_id]
+ single_img = inputs[:, frame_id].contiguous()
+ # det_results List[DetDataSample]
+ det_results = self.detector.predict(single_img, [img_data_sample])
+ assert len(det_results) == 1, 'Batch inference is not supported.'
+
+ pred_track_instances = self.tracker.track(
+ data_sample=det_results[0], **kwargs)
+ img_data_sample.pred_track_instances = pred_track_instances
+
+ return [track_data_sample]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/deep_sort.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/deep_sort.py
new file mode 100644
index 0000000000000000000000000000000000000000..70b30c7b07b2211fd0ad70767f479e57b6cd33f6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/deep_sort.py
@@ -0,0 +1,110 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional
+
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import TrackSampleList
+from mmdet.utils import OptConfigType
+from .base import BaseMOTModel
+
+
+@MODELS.register_module()
+class DeepSORT(BaseMOTModel):
+ """Simple online and realtime tracking with a deep association metric.
+
+ Details can be found at `DeepSORT`_.
+
+ Args:
+ detector (dict): Configuration of detector. Defaults to None.
+ reid (dict): Configuration of reid. Defaults to None
+ tracker (dict): Configuration of tracker. Defaults to None.
+ data_preprocessor (dict or ConfigDict, optional): The pre-process
+ config of :class:`TrackDataPreprocessor`. it usually includes,
+ ``pad_size_divisor``, ``pad_value``, ``mean`` and ``std``.
+ init_cfg (dict or list[dict]): Configuration of initialization.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ detector: Optional[dict] = None,
+ reid: Optional[dict] = None,
+ tracker: Optional[dict] = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptConfigType = None):
+ super().__init__(data_preprocessor, init_cfg)
+
+ if detector is not None:
+ self.detector = MODELS.build(detector)
+
+ if reid is not None:
+ self.reid = MODELS.build(reid)
+
+ if tracker is not None:
+ self.tracker = MODELS.build(tracker)
+
+ self.preprocess_cfg = data_preprocessor
+
+ def loss(self, inputs: Tensor, data_samples: TrackSampleList,
+ **kwargs) -> dict:
+ """Calculate losses from a batch of inputs and data samples."""
+ raise NotImplementedError(
+ 'Please train `detector` and `reid` models firstly, then \
+ inference with SORT/DeepSORT.')
+
+ def predict(self,
+ inputs: Tensor,
+ data_samples: TrackSampleList,
+ rescale: bool = True,
+ **kwargs) -> TrackSampleList:
+ """Predict results from a video and data samples with post- processing.
+
+ Args:
+ inputs (Tensor): of shape (N, T, C, H, W) encoding
+ input images. The N denotes batch size.
+ The T denotes the number of key frames
+ and reference frames.
+ data_samples (list[:obj:`TrackDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance`.
+ rescale (bool, Optional): If False, then returned bboxes and masks
+ will fit the scale of img, otherwise, returned bboxes and masks
+ will fit the scale of original image shape. Defaults to True.
+
+ Returns:
+ TrackSampleList: List[TrackDataSample]
+ Tracking results of the input videos.
+ Each DetDataSample usually contains ``pred_track_instances``.
+ """
+ assert inputs.dim() == 5, 'The img must be 5D Tensor (N, T, C, H, W).'
+ assert inputs.size(0) == 1, \
+ 'SORT/DeepSORT inference only support ' \
+ '1 batch size per gpu for now.'
+
+ assert len(data_samples) == 1, \
+ 'SORT/DeepSORT inference only support ' \
+ '1 batch size per gpu for now.'
+
+ track_data_sample = data_samples[0]
+ video_len = len(track_data_sample)
+ if track_data_sample[0].frame_id == 0:
+ self.tracker.reset()
+
+ for frame_id in range(video_len):
+ img_data_sample = track_data_sample[frame_id]
+ single_img = inputs[:, frame_id].contiguous()
+ # det_results List[DetDataSample]
+ det_results = self.detector.predict(single_img, [img_data_sample])
+ assert len(det_results) == 1, 'Batch inference is not supported.'
+
+ pred_track_instances = self.tracker.track(
+ model=self,
+ img=single_img,
+ feats=None,
+ data_sample=det_results[0],
+ data_preprocessor=self.preprocess_cfg,
+ rescale=rescale,
+ **kwargs)
+ img_data_sample.pred_track_instances = pred_track_instances
+
+ return [track_data_sample]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/ocsort.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/ocsort.py
new file mode 100644
index 0000000000000000000000000000000000000000..abf4eb3b06e2b1b223fe948f30dac877248377e3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/ocsort.py
@@ -0,0 +1,82 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+from typing import Dict, Optional
+
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import TrackSampleList
+from mmdet.utils import OptConfigType, OptMultiConfig
+from .base import BaseMOTModel
+
+
+@MODELS.register_module()
+class OCSORT(BaseMOTModel):
+ """OCOSRT: Observation-Centric SORT: Rethinking SORT for Robust
+ Multi-Object Tracking
+
+ This multi object tracker is the implementation of `OC-SORT
+ `_.
+
+ Args:
+ detector (dict): Configuration of detector. Defaults to None.
+ tracker (dict): Configuration of tracker. Defaults to None.
+ motion (dict): Configuration of motion. Defaults to None.
+ init_cfg (dict): Configuration of initialization. Defaults to None.
+ """
+
+ def __init__(self,
+ detector: Optional[dict] = None,
+ tracker: Optional[dict] = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None):
+ super().__init__(data_preprocessor, init_cfg)
+
+ if detector is not None:
+ self.detector = MODELS.build(detector)
+
+ if tracker is not None:
+ self.tracker = MODELS.build(tracker)
+
+ def loss(self, inputs: Tensor, data_samples: TrackSampleList,
+ **kwargs) -> dict:
+ """Calculate losses from a batch of inputs and data samples."""
+ return self.detector.loss(inputs, data_samples, **kwargs)
+
+ def predict(self, inputs: Dict[str, Tensor], data_samples: TrackSampleList,
+ **kwargs) -> TrackSampleList:
+ """Predict results from a video and data samples with post-processing.
+
+ Args:
+ inputs (Tensor): of shape (N, T, C, H, W) encoding
+ input images. The N denotes batch size.
+ The T denotes the number of frames in a video.
+ data_samples (list[:obj:`TrackDataSample`]): The batch
+ data samples. It usually includes information such
+ as `video_data_samples`.
+ Returns:
+ TrackSampleList: Tracking results of the inputs.
+ """
+ assert inputs.dim() == 5, 'The img must be 5D Tensor (N, T, C, H, W).'
+ assert inputs.size(0) == 1, \
+ 'OCSORT inference only support ' \
+ '1 batch size per gpu for now.'
+
+ assert len(data_samples) == 1, \
+ 'OCSORT inference only support 1 batch size per gpu for now.'
+
+ track_data_sample = data_samples[0]
+ video_len = len(track_data_sample)
+
+ for frame_id in range(video_len):
+ img_data_sample = track_data_sample[frame_id]
+ single_img = inputs[:, frame_id].contiguous()
+ # det_results List[DetDataSample]
+ det_results = self.detector.predict(single_img, [img_data_sample])
+ assert len(det_results) == 1, 'Batch inference is not supported.'
+
+ pred_track_instances = self.tracker.track(
+ data_sample=det_results[0], **kwargs)
+ img_data_sample.pred_track_instances = pred_track_instances
+
+ return [track_data_sample]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/qdtrack.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/qdtrack.py
new file mode 100644
index 0000000000000000000000000000000000000000..43d5dd60b8af8a6200e21a196c47d00dd2812a46
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/qdtrack.py
@@ -0,0 +1,186 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Union
+
+import torch
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import TrackSampleList
+from mmdet.utils import OptConfigType, OptMultiConfig
+from .base import BaseMOTModel
+
+
+@MODELS.register_module()
+class QDTrack(BaseMOTModel):
+ """Quasi-Dense Similarity Learning for Multiple Object Tracking.
+
+ This multi object tracker is the implementation of `QDTrack
+ `_.
+
+ Args:
+ detector (dict): Configuration of detector. Defaults to None.
+ track_head (dict): Configuration of track head. Defaults to None.
+ tracker (dict): Configuration of tracker. Defaults to None.
+ freeze_detector (bool): If True, freeze the detector weights.
+ Defaults to False.
+ data_preprocessor (dict or ConfigDict, optional): The pre-process
+ config of :class:`TrackDataPreprocessor`. it usually includes,
+ ``pad_size_divisor``, ``pad_value``, ``mean`` and ``std``.
+ init_cfg (dict or list[dict]): Configuration of initialization.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ detector: Optional[dict] = None,
+ track_head: Optional[dict] = None,
+ tracker: Optional[dict] = None,
+ freeze_detector: bool = False,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None):
+ super().__init__(data_preprocessor, init_cfg)
+ if detector is not None:
+ self.detector = MODELS.build(detector)
+
+ if track_head is not None:
+ self.track_head = MODELS.build(track_head)
+
+ if tracker is not None:
+ self.tracker = MODELS.build(tracker)
+
+ self.freeze_detector = freeze_detector
+ if self.freeze_detector:
+ self.freeze_module('detector')
+
+ def predict(self,
+ inputs: Tensor,
+ data_samples: TrackSampleList,
+ rescale: bool = True,
+ **kwargs) -> TrackSampleList:
+ """Predict results from a video and data samples with post- processing.
+
+ Args:
+ inputs (Tensor): of shape (N, T, C, H, W) encoding
+ input images. The N denotes batch size.
+ The T denotes the number of frames in a video.
+ data_samples (list[:obj:`TrackDataSample`]): The batch
+ data samples. It usually includes information such
+ as `video_data_samples`.
+ rescale (bool, Optional): If False, then returned bboxes and masks
+ will fit the scale of img, otherwise, returned bboxes and masks
+ will fit the scale of original image shape. Defaults to True.
+
+ Returns:
+ TrackSampleList: Tracking results of the inputs.
+ """
+ assert inputs.dim() == 5, 'The img must be 5D Tensor (N, T, C, H, W).'
+ assert inputs.size(0) == 1, \
+ 'QDTrack inference only support 1 batch size per gpu for now.'
+
+ assert len(data_samples) == 1, \
+ 'QDTrack only support 1 batch size per gpu for now.'
+
+ track_data_sample = data_samples[0]
+ video_len = len(track_data_sample)
+ if track_data_sample[0].frame_id == 0:
+ self.tracker.reset()
+
+ for frame_id in range(video_len):
+ img_data_sample = track_data_sample[frame_id]
+ single_img = inputs[:, frame_id].contiguous()
+ x = self.detector.extract_feat(single_img)
+ rpn_results_list = self.detector.rpn_head.predict(
+ x, [img_data_sample])
+ # det_results List[InstanceData]
+ det_results = self.detector.roi_head.predict(
+ x, rpn_results_list, [img_data_sample], rescale=rescale)
+ assert len(det_results) == 1, 'Batch inference is not supported.'
+ img_data_sample.pred_instances = det_results[0]
+ frame_pred_track_instances = self.tracker.track(
+ model=self,
+ img=single_img,
+ feats=x,
+ data_sample=img_data_sample,
+ **kwargs)
+ img_data_sample.pred_track_instances = frame_pred_track_instances
+
+ return [track_data_sample]
+
+ def loss(self, inputs: Tensor, data_samples: TrackSampleList,
+ **kwargs) -> Union[dict, tuple]:
+ """Calculate losses from a batch of inputs and data samples.
+
+ Args:
+ inputs (Dict[str, Tensor]): of shape (N, T, C, H, W) encoding
+ input images. Typically these should be mean centered and std
+ scaled. The N denotes batch size. The T denotes the number of
+ frames.
+ data_samples (list[:obj:`TrackDataSample`]): The batch
+ data samples. It usually includes information such
+ as `video_data_samples`.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ # modify the inputs shape to fit mmdet
+ assert inputs.dim() == 5, 'The img must be 5D Tensor (N, T, C, H, W).'
+ assert inputs.size(1) == 2, \
+ 'QDTrack can only have 1 key frame and 1 reference frame.'
+
+ # split the data_samples into two aspects: key frames and reference
+ # frames
+ ref_data_samples, key_data_samples = [], []
+ key_frame_inds, ref_frame_inds = [], []
+ # set cat_id of gt_labels to 0 in RPN
+ for track_data_sample in data_samples:
+ key_frame_inds.append(track_data_sample.key_frames_inds[0])
+ ref_frame_inds.append(track_data_sample.ref_frames_inds[0])
+ key_data_sample = track_data_sample.get_key_frames()[0]
+ key_data_sample.gt_instances.labels = \
+ torch.zeros_like(key_data_sample.gt_instances.labels)
+ key_data_samples.append(key_data_sample)
+ ref_data_sample = track_data_sample.get_ref_frames()[0]
+ ref_data_samples.append(ref_data_sample)
+
+ key_frame_inds = torch.tensor(key_frame_inds, dtype=torch.int64)
+ ref_frame_inds = torch.tensor(ref_frame_inds, dtype=torch.int64)
+ batch_inds = torch.arange(len(inputs))
+ key_imgs = inputs[batch_inds, key_frame_inds].contiguous()
+ ref_imgs = inputs[batch_inds, ref_frame_inds].contiguous()
+
+ x = self.detector.extract_feat(key_imgs)
+ ref_x = self.detector.extract_feat(ref_imgs)
+
+ losses = dict()
+ # RPN head forward and loss
+ assert self.detector.with_rpn, \
+ 'QDTrack only support detector with RPN.'
+
+ proposal_cfg = self.detector.train_cfg.get('rpn_proposal',
+ self.detector.test_cfg.rpn)
+ rpn_losses, rpn_results_list = self.detector.rpn_head. \
+ loss_and_predict(x,
+ key_data_samples,
+ proposal_cfg=proposal_cfg,
+ **kwargs)
+ ref_rpn_results_list = self.detector.rpn_head.predict(
+ ref_x, ref_data_samples, **kwargs)
+
+ # avoid get same name with roi_head loss
+ keys = rpn_losses.keys()
+ for key in keys:
+ if 'loss' in key and 'rpn' not in key:
+ rpn_losses[f'rpn_{key}'] = rpn_losses.pop(key)
+ losses.update(rpn_losses)
+
+ # roi_head loss
+ losses_detect = self.detector.roi_head.loss(x, rpn_results_list,
+ key_data_samples, **kwargs)
+ losses.update(losses_detect)
+
+ # tracking head loss
+ losses_track = self.track_head.loss(x, ref_x, rpn_results_list,
+ ref_rpn_results_list, data_samples,
+ **kwargs)
+ losses.update(losses_track)
+
+ return losses
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/strongsort.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/strongsort.py
new file mode 100644
index 0000000000000000000000000000000000000000..6129bf49972233206b3c05daa2174f99723d1b9d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/mot/strongsort.py
@@ -0,0 +1,129 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional
+
+import numpy as np
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures import TrackSampleList
+from mmdet.utils import OptConfigType
+from .deep_sort import DeepSORT
+
+
+@MODELS.register_module()
+class StrongSORT(DeepSORT):
+ """StrongSORT: Make DeepSORT Great Again.
+
+ Details can be found at `StrongSORT`_.
+
+ Args:
+ detector (dict): Configuration of detector. Defaults to None.
+ reid (dict): Configuration of reid. Defaults to None
+ tracker (dict): Configuration of tracker. Defaults to None.
+ kalman (dict): Configuration of Kalman filter. Defaults to None.
+ cmc (dict): Configuration of camera model compensation.
+ Defaults to None.
+ data_preprocessor (dict or ConfigDict, optional): The pre-process
+ config of :class:`TrackDataPreprocessor`. it usually includes,
+ ``pad_size_divisor``, ``pad_value``, ``mean`` and ``std``.
+ init_cfg (dict or list[dict]): Configuration of initialization.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ detector: Optional[dict] = None,
+ reid: Optional[dict] = None,
+ cmc: Optional[dict] = None,
+ tracker: Optional[dict] = None,
+ postprocess_model: Optional[dict] = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptConfigType = None):
+ super().__init__(detector, reid, tracker, data_preprocessor, init_cfg)
+
+ if cmc is not None:
+ self.cmc = TASK_UTILS.build(cmc)
+
+ if postprocess_model is not None:
+ self.postprocess_model = TASK_UTILS.build(postprocess_model)
+
+ @property
+ def with_cmc(self):
+ """bool: whether the framework has a camera model compensation
+ model.
+ """
+ return hasattr(self, 'cmc') and self.cmc is not None
+
+ def predict(self,
+ inputs: Tensor,
+ data_samples: TrackSampleList,
+ rescale: bool = True,
+ **kwargs) -> TrackSampleList:
+ """Predict results from a video and data samples with post- processing.
+
+ Args:
+ inputs (Tensor): of shape (N, T, C, H, W) encoding
+ input images. The N denotes batch size.
+ The T denotes the number of key frames
+ and reference frames.
+ data_samples (list[:obj:`TrackDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance`.
+ rescale (bool, Optional): If False, then returned bboxes and masks
+ will fit the scale of img, otherwise, returned bboxes and masks
+ will fit the scale of original image shape. Defaults to True.
+
+ Returns:
+ TrackSampleList: List[TrackDataSample]
+ Tracking results of the input videos.
+ Each DetDataSample usually contains ``pred_track_instances``.
+ """
+ assert inputs.dim() == 5, 'The img must be 5D Tensor (N, T, C, H, W).'
+ assert inputs.size(0) == 1, \
+ 'SORT/DeepSORT inference only support ' \
+ '1 batch size per gpu for now.'
+
+ assert len(data_samples) == 1, \
+ 'SORT/DeepSORT inference only support ' \
+ '1 batch size per gpu for now.'
+
+ track_data_sample = data_samples[0]
+ video_len = len(track_data_sample)
+
+ video_track_instances = []
+ for frame_id in range(video_len):
+ img_data_sample = track_data_sample[frame_id]
+ single_img = inputs[:, frame_id].contiguous()
+ # det_results List[DetDataSample]
+ det_results = self.detector.predict(single_img, [img_data_sample])
+ assert len(det_results) == 1, 'Batch inference is not supported.'
+
+ pred_track_instances = self.tracker.track(
+ model=self,
+ img=single_img,
+ data_sample=det_results[0],
+ data_preprocessor=self.preprocess_cfg,
+ rescale=rescale,
+ **kwargs)
+ for i in range(len(pred_track_instances.instances_id)):
+ video_track_instances.append(
+ np.array([
+ frame_id + 1,
+ pred_track_instances.instances_id[i].cpu(),
+ pred_track_instances.bboxes[i][0].cpu(),
+ pred_track_instances.bboxes[i][1].cpu(),
+ (pred_track_instances.bboxes[i][2] -
+ pred_track_instances.bboxes[i][0]).cpu(),
+ (pred_track_instances.bboxes[i][3] -
+ pred_track_instances.bboxes[i][1]).cpu(),
+ pred_track_instances.scores[i].cpu()
+ ]))
+ video_track_instances = np.array(video_track_instances).reshape(-1, 7)
+ video_track_instances = self.postprocess_model.forward(
+ video_track_instances)
+ for frame_id in range(video_len):
+ track_data_sample[frame_id].pred_track_instances = \
+ InstanceData(bboxes=video_track_instances[
+ video_track_instances[:, 0] == frame_id + 1, :])
+
+ return [track_data_sample]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..343fbfefbd871d00e855d1c3cf4b531345e4dcf1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/__init__.py
@@ -0,0 +1,27 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .bfp import BFP
+from .channel_mapper import ChannelMapper
+from .cspnext_pafpn import CSPNeXtPAFPN
+from .ct_resnet_neck import CTResNetNeck
+from .dilated_encoder import DilatedEncoder
+from .dyhead import DyHead
+from .fpg import FPG
+from .fpn import FPN
+from .fpn_carafe import FPN_CARAFE
+from .fpn_dropblock import FPN_DropBlock
+from .hrfpn import HRFPN
+from .nas_fpn import NASFPN
+from .nasfcos_fpn import NASFCOS_FPN
+from .pafpn import PAFPN
+from .rfp import RFP
+from .ssd_neck import SSDNeck
+from .ssh import SSH
+from .yolo_neck import YOLOV3Neck
+from .yolox_pafpn import YOLOXPAFPN
+
+__all__ = [
+ 'FPN', 'BFP', 'ChannelMapper', 'HRFPN', 'NASFPN', 'FPN_CARAFE', 'PAFPN',
+ 'NASFCOS_FPN', 'RFP', 'YOLOV3Neck', 'FPG', 'DilatedEncoder',
+ 'CTResNetNeck', 'SSDNeck', 'YOLOXPAFPN', 'DyHead', 'CSPNeXtPAFPN', 'SSH',
+ 'FPN_DropBlock'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/bfp.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/bfp.py
new file mode 100644
index 0000000000000000000000000000000000000000..401cdb0f552b06c9e8eb185c3e8ae0ba7112a9d8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/bfp.py
@@ -0,0 +1,111 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Tuple
+
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule
+from mmcv.cnn.bricks import NonLocal2d
+from mmengine.model import BaseModule
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import OptConfigType, OptMultiConfig
+
+
+@MODELS.register_module()
+class BFP(BaseModule):
+ """BFP (Balanced Feature Pyramids)
+
+ BFP takes multi-level features as inputs and gather them into a single one,
+ then refine the gathered feature and scatter the refined results to
+ multi-level features. This module is used in Libra R-CNN (CVPR 2019), see
+ the paper `Libra R-CNN: Towards Balanced Learning for Object Detection
+ `_ for details.
+
+ Args:
+ in_channels (int): Number of input channels (feature maps of all levels
+ should have the same channels).
+ num_levels (int): Number of input feature levels.
+ refine_level (int): Index of integration and refine level of BSF in
+ multi-level features from bottom to top.
+ refine_type (str): Type of the refine op, currently support
+ [None, 'conv', 'non_local'].
+ conv_cfg (:obj:`ConfigDict` or dict, optional): The config dict for
+ convolution layers.
+ norm_cfg (:obj:`ConfigDict` or dict, optional): The config dict for
+ normalization layers.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or
+ dict], optional): Initialization config dict.
+ """
+
+ def __init__(
+ self,
+ in_channels: int,
+ num_levels: int,
+ refine_level: int = 2,
+ refine_type: str = None,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: OptConfigType = None,
+ init_cfg: OptMultiConfig = dict(
+ type='Xavier', layer='Conv2d', distribution='uniform')
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ assert refine_type in [None, 'conv', 'non_local']
+
+ self.in_channels = in_channels
+ self.num_levels = num_levels
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+
+ self.refine_level = refine_level
+ self.refine_type = refine_type
+ assert 0 <= self.refine_level < self.num_levels
+
+ if self.refine_type == 'conv':
+ self.refine = ConvModule(
+ self.in_channels,
+ self.in_channels,
+ 3,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg)
+ elif self.refine_type == 'non_local':
+ self.refine = NonLocal2d(
+ self.in_channels,
+ reduction=1,
+ use_scale=False,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg)
+
+ def forward(self, inputs: Tuple[Tensor]) -> Tuple[Tensor]:
+ """Forward function."""
+ assert len(inputs) == self.num_levels
+
+ # step 1: gather multi-level features by resize and average
+ feats = []
+ gather_size = inputs[self.refine_level].size()[2:]
+ for i in range(self.num_levels):
+ if i < self.refine_level:
+ gathered = F.adaptive_max_pool2d(
+ inputs[i], output_size=gather_size)
+ else:
+ gathered = F.interpolate(
+ inputs[i], size=gather_size, mode='nearest')
+ feats.append(gathered)
+
+ bsf = sum(feats) / len(feats)
+
+ # step 2: refine gathered features
+ if self.refine_type is not None:
+ bsf = self.refine(bsf)
+
+ # step 3: scatter refined features to multi-levels by a residual path
+ outs = []
+ for i in range(self.num_levels):
+ out_size = inputs[i].size()[2:]
+ if i < self.refine_level:
+ residual = F.interpolate(bsf, size=out_size, mode='nearest')
+ else:
+ residual = F.adaptive_max_pool2d(bsf, output_size=out_size)
+ outs.append(residual + inputs[i])
+
+ return tuple(outs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/channel_mapper.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/channel_mapper.py
new file mode 100644
index 0000000000000000000000000000000000000000..74293618f2b8a649328ae4a5a0571809de9991dd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/channel_mapper.py
@@ -0,0 +1,112 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple, Union
+
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from mmengine.model import BaseModule
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import OptConfigType, OptMultiConfig
+
+
+@MODELS.register_module()
+class ChannelMapper(BaseModule):
+ """Channel Mapper to reduce/increase channels of backbone features.
+
+ This is used to reduce/increase channels of backbone features.
+
+ Args:
+ in_channels (List[int]): Number of input channels per scale.
+ out_channels (int): Number of output channels (used at each scale).
+ kernel_size (int, optional): kernel_size for reducing channels (used
+ at each scale). Default: 3.
+ conv_cfg (:obj:`ConfigDict` or dict, optional): Config dict for
+ convolution layer. Default: None.
+ norm_cfg (:obj:`ConfigDict` or dict, optional): Config dict for
+ normalization layer. Default: None.
+ act_cfg (:obj:`ConfigDict` or dict, optional): Config dict for
+ activation layer in ConvModule. Default: dict(type='ReLU').
+ bias (bool | str): If specified as `auto`, it will be decided by the
+ norm_cfg. Bias will be set as True if `norm_cfg` is None, otherwise
+ False. Default: "auto".
+ num_outs (int, optional): Number of output feature maps. There would
+ be extra_convs when num_outs larger than the length of in_channels.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or dict],
+ optional): Initialization config dict.
+ Example:
+ >>> import torch
+ >>> in_channels = [2, 3, 5, 7]
+ >>> scales = [340, 170, 84, 43]
+ >>> inputs = [torch.rand(1, c, s, s)
+ ... for c, s in zip(in_channels, scales)]
+ >>> self = ChannelMapper(in_channels, 11, 3).eval()
+ >>> outputs = self.forward(inputs)
+ >>> for i in range(len(outputs)):
+ ... print(f'outputs[{i}].shape = {outputs[i].shape}')
+ outputs[0].shape = torch.Size([1, 11, 340, 340])
+ outputs[1].shape = torch.Size([1, 11, 170, 170])
+ outputs[2].shape = torch.Size([1, 11, 84, 84])
+ outputs[3].shape = torch.Size([1, 11, 43, 43])
+ """
+
+ def __init__(
+ self,
+ in_channels: List[int],
+ out_channels: int,
+ kernel_size: int = 3,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: OptConfigType = None,
+ act_cfg: OptConfigType = dict(type='ReLU'),
+ bias: Union[bool, str] = 'auto',
+ num_outs: int = None,
+ init_cfg: OptMultiConfig = dict(
+ type='Xavier', layer='Conv2d', distribution='uniform')
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ assert isinstance(in_channels, list)
+ self.extra_convs = None
+ if num_outs is None:
+ num_outs = len(in_channels)
+ self.convs = nn.ModuleList()
+ for in_channel in in_channels:
+ self.convs.append(
+ ConvModule(
+ in_channel,
+ out_channels,
+ kernel_size,
+ padding=(kernel_size - 1) // 2,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg,
+ bias=bias))
+ if num_outs > len(in_channels):
+ self.extra_convs = nn.ModuleList()
+ for i in range(len(in_channels), num_outs):
+ if i == len(in_channels):
+ in_channel = in_channels[-1]
+ else:
+ in_channel = out_channels
+ self.extra_convs.append(
+ ConvModule(
+ in_channel,
+ out_channels,
+ 3,
+ stride=2,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg,
+ bias=bias))
+
+ def forward(self, inputs: Tuple[Tensor]) -> Tuple[Tensor]:
+ """Forward function."""
+ assert len(inputs) == len(self.convs)
+ outs = [self.convs[i](inputs[i]) for i in range(len(inputs))]
+ if self.extra_convs:
+ for i in range(len(self.extra_convs)):
+ if i == 0:
+ outs.append(self.extra_convs[0](inputs[-1]))
+ else:
+ outs.append(self.extra_convs[i](outs[-1]))
+ return tuple(outs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/cspnext_pafpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/cspnext_pafpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..a52ba72d9b3e48c4866fb16507bc2118eb23010e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/cspnext_pafpn.py
@@ -0,0 +1,170 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+from typing import Sequence, Tuple
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule, DepthwiseSeparableConvModule
+from mmengine.model import BaseModule
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptMultiConfig
+from ..layers import CSPLayer
+
+
+@MODELS.register_module()
+class CSPNeXtPAFPN(BaseModule):
+ """Path Aggregation Network with CSPNeXt blocks.
+
+ Args:
+ in_channels (Sequence[int]): Number of input channels per scale.
+ out_channels (int): Number of output channels (used at each scale)
+ num_csp_blocks (int): Number of bottlenecks in CSPLayer.
+ Defaults to 3.
+ use_depthwise (bool): Whether to use depthwise separable convolution in
+ blocks. Defaults to False.
+ expand_ratio (float): Ratio to adjust the number of channels of the
+ hidden layer. Default: 0.5
+ upsample_cfg (dict): Config dict for interpolate layer.
+ Default: `dict(scale_factor=2, mode='nearest')`
+ conv_cfg (dict, optional): Config dict for convolution layer.
+ Default: None, which means using conv2d.
+ norm_cfg (dict): Config dict for normalization layer.
+ Default: dict(type='BN')
+ act_cfg (dict): Config dict for activation layer.
+ Default: dict(type='Swish')
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None.
+ """
+
+ def __init__(
+ self,
+ in_channels: Sequence[int],
+ out_channels: int,
+ num_csp_blocks: int = 3,
+ use_depthwise: bool = False,
+ expand_ratio: float = 0.5,
+ upsample_cfg: ConfigType = dict(scale_factor=2, mode='nearest'),
+ conv_cfg: bool = None,
+ norm_cfg: ConfigType = dict(type='BN', momentum=0.03, eps=0.001),
+ act_cfg: ConfigType = dict(type='Swish'),
+ init_cfg: OptMultiConfig = dict(
+ type='Kaiming',
+ layer='Conv2d',
+ a=math.sqrt(5),
+ distribution='uniform',
+ mode='fan_in',
+ nonlinearity='leaky_relu')
+ ) -> None:
+ super().__init__(init_cfg)
+ self.in_channels = in_channels
+ self.out_channels = out_channels
+
+ conv = DepthwiseSeparableConvModule if use_depthwise else ConvModule
+
+ # build top-down blocks
+ self.upsample = nn.Upsample(**upsample_cfg)
+ self.reduce_layers = nn.ModuleList()
+ self.top_down_blocks = nn.ModuleList()
+ for idx in range(len(in_channels) - 1, 0, -1):
+ self.reduce_layers.append(
+ ConvModule(
+ in_channels[idx],
+ in_channels[idx - 1],
+ 1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg))
+ self.top_down_blocks.append(
+ CSPLayer(
+ in_channels[idx - 1] * 2,
+ in_channels[idx - 1],
+ num_blocks=num_csp_blocks,
+ add_identity=False,
+ use_depthwise=use_depthwise,
+ use_cspnext_block=True,
+ expand_ratio=expand_ratio,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg))
+
+ # build bottom-up blocks
+ self.downsamples = nn.ModuleList()
+ self.bottom_up_blocks = nn.ModuleList()
+ for idx in range(len(in_channels) - 1):
+ self.downsamples.append(
+ conv(
+ in_channels[idx],
+ in_channels[idx],
+ 3,
+ stride=2,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg))
+ self.bottom_up_blocks.append(
+ CSPLayer(
+ in_channels[idx] * 2,
+ in_channels[idx + 1],
+ num_blocks=num_csp_blocks,
+ add_identity=False,
+ use_depthwise=use_depthwise,
+ use_cspnext_block=True,
+ expand_ratio=expand_ratio,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg))
+
+ self.out_convs = nn.ModuleList()
+ for i in range(len(in_channels)):
+ self.out_convs.append(
+ conv(
+ in_channels[i],
+ out_channels,
+ 3,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg))
+
+ def forward(self, inputs: Tuple[Tensor, ...]) -> Tuple[Tensor, ...]:
+ """
+ Args:
+ inputs (tuple[Tensor]): input features.
+
+ Returns:
+ tuple[Tensor]: YOLOXPAFPN features.
+ """
+ assert len(inputs) == len(self.in_channels)
+
+ # top-down path
+ inner_outs = [inputs[-1]]
+ for idx in range(len(self.in_channels) - 1, 0, -1):
+ feat_heigh = inner_outs[0]
+ feat_low = inputs[idx - 1]
+ feat_heigh = self.reduce_layers[len(self.in_channels) - 1 - idx](
+ feat_heigh)
+ inner_outs[0] = feat_heigh
+
+ upsample_feat = self.upsample(feat_heigh)
+
+ inner_out = self.top_down_blocks[len(self.in_channels) - 1 - idx](
+ torch.cat([upsample_feat, feat_low], 1))
+ inner_outs.insert(0, inner_out)
+
+ # bottom-up path
+ outs = [inner_outs[0]]
+ for idx in range(len(self.in_channels) - 1):
+ feat_low = outs[-1]
+ feat_height = inner_outs[idx + 1]
+ downsample_feat = self.downsamples[idx](feat_low)
+ out = self.bottom_up_blocks[idx](
+ torch.cat([downsample_feat, feat_height], 1))
+ outs.append(out)
+
+ # out convs
+ for idx, conv in enumerate(self.out_convs):
+ outs[idx] = conv(outs[idx])
+
+ return tuple(outs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/ct_resnet_neck.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/ct_resnet_neck.py
new file mode 100644
index 0000000000000000000000000000000000000000..9109fe79290fafecd954f223d5365ef619c0c301
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/ct_resnet_neck.py
@@ -0,0 +1,102 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+from typing import Sequence, Tuple
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from mmengine.model import BaseModule
+
+from mmdet.registry import MODELS
+from mmdet.utils import OptMultiConfig
+
+
+@MODELS.register_module()
+class CTResNetNeck(BaseModule):
+ """The neck used in `CenterNet `_ for
+ object classification and box regression.
+
+ Args:
+ in_channels (int): Number of input channels.
+ num_deconv_filters (tuple[int]): Number of filters per stage.
+ num_deconv_kernels (tuple[int]): Number of kernels per stage.
+ use_dcn (bool): If True, use DCNv2. Defaults to True.
+ init_cfg (:obj:`ConfigDict` or dict or list[dict] or
+ list[:obj:`ConfigDict`], optional): Initialization
+ config dict.
+ """
+
+ def __init__(self,
+ in_channels: int,
+ num_deconv_filters: Tuple[int, ...],
+ num_deconv_kernels: Tuple[int, ...],
+ use_dcn: bool = True,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ assert len(num_deconv_filters) == len(num_deconv_kernels)
+ self.fp16_enabled = False
+ self.use_dcn = use_dcn
+ self.in_channels = in_channels
+ self.deconv_layers = self._make_deconv_layer(num_deconv_filters,
+ num_deconv_kernels)
+
+ def _make_deconv_layer(
+ self, num_deconv_filters: Tuple[int, ...],
+ num_deconv_kernels: Tuple[int, ...]) -> nn.Sequential:
+ """use deconv layers to upsample backbone's output."""
+ layers = []
+ for i in range(len(num_deconv_filters)):
+ feat_channels = num_deconv_filters[i]
+ conv_module = ConvModule(
+ self.in_channels,
+ feat_channels,
+ 3,
+ padding=1,
+ conv_cfg=dict(type='DCNv2') if self.use_dcn else None,
+ norm_cfg=dict(type='BN'))
+ layers.append(conv_module)
+ upsample_module = ConvModule(
+ feat_channels,
+ feat_channels,
+ num_deconv_kernels[i],
+ stride=2,
+ padding=1,
+ conv_cfg=dict(type='deconv'),
+ norm_cfg=dict(type='BN'))
+ layers.append(upsample_module)
+ self.in_channels = feat_channels
+
+ return nn.Sequential(*layers)
+
+ def init_weights(self) -> None:
+ """Initialize the parameters."""
+ for m in self.modules():
+ if isinstance(m, nn.ConvTranspose2d):
+ # In order to be consistent with the source code,
+ # reset the ConvTranspose2d initialization parameters
+ m.reset_parameters()
+ # Simulated bilinear upsampling kernel
+ w = m.weight.data
+ f = math.ceil(w.size(2) / 2)
+ c = (2 * f - 1 - f % 2) / (2. * f)
+ for i in range(w.size(2)):
+ for j in range(w.size(3)):
+ w[0, 0, i, j] = \
+ (1 - math.fabs(i / f - c)) * (
+ 1 - math.fabs(j / f - c))
+ for c in range(1, w.size(0)):
+ w[c, 0, :, :] = w[0, 0, :, :]
+ elif isinstance(m, nn.BatchNorm2d):
+ nn.init.constant_(m.weight, 1)
+ nn.init.constant_(m.bias, 0)
+ # self.use_dcn is False
+ elif not self.use_dcn and isinstance(m, nn.Conv2d):
+ # In order to be consistent with the source code,
+ # reset the Conv2d initialization parameters
+ m.reset_parameters()
+
+ def forward(self, x: Sequence[torch.Tensor]) -> Tuple[torch.Tensor]:
+ """model forward."""
+ assert isinstance(x, (list, tuple))
+ outs = self.deconv_layers(x[-1])
+ return outs,
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/dilated_encoder.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/dilated_encoder.py
new file mode 100644
index 0000000000000000000000000000000000000000..e9beb3ea9b4289da8d0100ae7759927f045829bb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/dilated_encoder.py
@@ -0,0 +1,109 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch.nn as nn
+from mmcv.cnn import ConvModule, is_norm
+from mmengine.model import caffe2_xavier_init, constant_init, normal_init
+from torch.nn import BatchNorm2d
+
+from mmdet.registry import MODELS
+
+
+class Bottleneck(nn.Module):
+ """Bottleneck block for DilatedEncoder used in `YOLOF.
+
+ `.
+
+ The Bottleneck contains three ConvLayers and one residual connection.
+
+ Args:
+ in_channels (int): The number of input channels.
+ mid_channels (int): The number of middle output channels.
+ dilation (int): Dilation rate.
+ norm_cfg (dict): Dictionary to construct and config norm layer.
+ """
+
+ def __init__(self,
+ in_channels,
+ mid_channels,
+ dilation,
+ norm_cfg=dict(type='BN', requires_grad=True)):
+ super(Bottleneck, self).__init__()
+ self.conv1 = ConvModule(
+ in_channels, mid_channels, 1, norm_cfg=norm_cfg)
+ self.conv2 = ConvModule(
+ mid_channels,
+ mid_channels,
+ 3,
+ padding=dilation,
+ dilation=dilation,
+ norm_cfg=norm_cfg)
+ self.conv3 = ConvModule(
+ mid_channels, in_channels, 1, norm_cfg=norm_cfg)
+
+ def forward(self, x):
+ identity = x
+ out = self.conv1(x)
+ out = self.conv2(out)
+ out = self.conv3(out)
+ out = out + identity
+ return out
+
+
+@MODELS.register_module()
+class DilatedEncoder(nn.Module):
+ """Dilated Encoder for YOLOF `.
+
+ This module contains two types of components:
+ - the original FPN lateral convolution layer and fpn convolution layer,
+ which are 1x1 conv + 3x3 conv
+ - the dilated residual block
+
+ Args:
+ in_channels (int): The number of input channels.
+ out_channels (int): The number of output channels.
+ block_mid_channels (int): The number of middle block output channels
+ num_residual_blocks (int): The number of residual blocks.
+ block_dilations (list): The list of residual blocks dilation.
+ """
+
+ def __init__(self, in_channels, out_channels, block_mid_channels,
+ num_residual_blocks, block_dilations):
+ super(DilatedEncoder, self).__init__()
+ self.in_channels = in_channels
+ self.out_channels = out_channels
+ self.block_mid_channels = block_mid_channels
+ self.num_residual_blocks = num_residual_blocks
+ self.block_dilations = block_dilations
+ self._init_layers()
+
+ def _init_layers(self):
+ self.lateral_conv = nn.Conv2d(
+ self.in_channels, self.out_channels, kernel_size=1)
+ self.lateral_norm = BatchNorm2d(self.out_channels)
+ self.fpn_conv = nn.Conv2d(
+ self.out_channels, self.out_channels, kernel_size=3, padding=1)
+ self.fpn_norm = BatchNorm2d(self.out_channels)
+ encoder_blocks = []
+ for i in range(self.num_residual_blocks):
+ dilation = self.block_dilations[i]
+ encoder_blocks.append(
+ Bottleneck(
+ self.out_channels,
+ self.block_mid_channels,
+ dilation=dilation))
+ self.dilated_encoder_blocks = nn.Sequential(*encoder_blocks)
+
+ def init_weights(self):
+ caffe2_xavier_init(self.lateral_conv)
+ caffe2_xavier_init(self.fpn_conv)
+ for m in [self.lateral_norm, self.fpn_norm]:
+ constant_init(m, 1)
+ for m in self.dilated_encoder_blocks.modules():
+ if isinstance(m, nn.Conv2d):
+ normal_init(m, mean=0, std=0.01)
+ if is_norm(m):
+ constant_init(m, 1)
+
+ def forward(self, feature):
+ out = self.lateral_norm(self.lateral_conv(feature[-1]))
+ out = self.fpn_norm(self.fpn_conv(out))
+ return self.dilated_encoder_blocks(out),
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/dyhead.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/dyhead.py
new file mode 100644
index 0000000000000000000000000000000000000000..5f5ae0b285c20558a0c7bcc59cbb7b214684eab2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/dyhead.py
@@ -0,0 +1,173 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import build_activation_layer, build_norm_layer
+from mmcv.ops.modulated_deform_conv import ModulatedDeformConv2d
+from mmengine.model import BaseModule, constant_init, normal_init
+
+from mmdet.registry import MODELS
+from ..layers import DyReLU
+
+# Reference:
+# https://github.com/microsoft/DynamicHead
+# https://github.com/jshilong/SEPC
+
+
+class DyDCNv2(nn.Module):
+ """ModulatedDeformConv2d with normalization layer used in DyHead.
+
+ This module cannot be configured with `conv_cfg=dict(type='DCNv2')`
+ because DyHead calculates offset and mask from middle-level feature.
+
+ Args:
+ in_channels (int): Number of input channels.
+ out_channels (int): Number of output channels.
+ stride (int | tuple[int], optional): Stride of the convolution.
+ Default: 1.
+ norm_cfg (dict, optional): Config dict for normalization layer.
+ Default: dict(type='GN', num_groups=16, requires_grad=True).
+ """
+
+ def __init__(self,
+ in_channels,
+ out_channels,
+ stride=1,
+ norm_cfg=dict(type='GN', num_groups=16, requires_grad=True)):
+ super().__init__()
+ self.with_norm = norm_cfg is not None
+ bias = not self.with_norm
+ self.conv = ModulatedDeformConv2d(
+ in_channels, out_channels, 3, stride=stride, padding=1, bias=bias)
+ if self.with_norm:
+ self.norm = build_norm_layer(norm_cfg, out_channels)[1]
+
+ def forward(self, x, offset, mask):
+ """Forward function."""
+ x = self.conv(x.contiguous(), offset, mask)
+ if self.with_norm:
+ x = self.norm(x)
+ return x
+
+
+class DyHeadBlock(nn.Module):
+ """DyHead Block with three types of attention.
+
+ HSigmoid arguments in default act_cfg follow official code, not paper.
+ https://github.com/microsoft/DynamicHead/blob/master/dyhead/dyrelu.py
+
+ Args:
+ in_channels (int): Number of input channels.
+ out_channels (int): Number of output channels.
+ zero_init_offset (bool, optional): Whether to use zero init for
+ `spatial_conv_offset`. Default: True.
+ act_cfg (dict, optional): Config dict for the last activation layer of
+ scale-aware attention. Default: dict(type='HSigmoid', bias=3.0,
+ divisor=6.0).
+ """
+
+ def __init__(self,
+ in_channels,
+ out_channels,
+ zero_init_offset=True,
+ act_cfg=dict(type='HSigmoid', bias=3.0, divisor=6.0)):
+ super().__init__()
+ self.zero_init_offset = zero_init_offset
+ # (offset_x, offset_y, mask) * kernel_size_y * kernel_size_x
+ self.offset_and_mask_dim = 3 * 3 * 3
+ self.offset_dim = 2 * 3 * 3
+
+ self.spatial_conv_high = DyDCNv2(in_channels, out_channels)
+ self.spatial_conv_mid = DyDCNv2(in_channels, out_channels)
+ self.spatial_conv_low = DyDCNv2(in_channels, out_channels, stride=2)
+ self.spatial_conv_offset = nn.Conv2d(
+ in_channels, self.offset_and_mask_dim, 3, padding=1)
+ self.scale_attn_module = nn.Sequential(
+ nn.AdaptiveAvgPool2d(1), nn.Conv2d(out_channels, 1, 1),
+ nn.ReLU(inplace=True), build_activation_layer(act_cfg))
+ self.task_attn_module = DyReLU(out_channels)
+ self._init_weights()
+
+ def _init_weights(self):
+ for m in self.modules():
+ if isinstance(m, nn.Conv2d):
+ normal_init(m, 0, 0.01)
+ if self.zero_init_offset:
+ constant_init(self.spatial_conv_offset, 0)
+
+ def forward(self, x):
+ """Forward function."""
+ outs = []
+ for level in range(len(x)):
+ # calculate offset and mask of DCNv2 from middle-level feature
+ offset_and_mask = self.spatial_conv_offset(x[level])
+ offset = offset_and_mask[:, :self.offset_dim, :, :]
+ mask = offset_and_mask[:, self.offset_dim:, :, :].sigmoid()
+
+ mid_feat = self.spatial_conv_mid(x[level], offset, mask)
+ sum_feat = mid_feat * self.scale_attn_module(mid_feat)
+ summed_levels = 1
+ if level > 0:
+ low_feat = self.spatial_conv_low(x[level - 1], offset, mask)
+ sum_feat += low_feat * self.scale_attn_module(low_feat)
+ summed_levels += 1
+ if level < len(x) - 1:
+ # this upsample order is weird, but faster than natural order
+ # https://github.com/microsoft/DynamicHead/issues/25
+ high_feat = F.interpolate(
+ self.spatial_conv_high(x[level + 1], offset, mask),
+ size=x[level].shape[-2:],
+ mode='bilinear',
+ align_corners=True)
+ sum_feat += high_feat * self.scale_attn_module(high_feat)
+ summed_levels += 1
+ outs.append(self.task_attn_module(sum_feat / summed_levels))
+
+ return outs
+
+
+@MODELS.register_module()
+class DyHead(BaseModule):
+ """DyHead neck consisting of multiple DyHead Blocks.
+
+ See `Dynamic Head: Unifying Object Detection Heads with Attentions
+ `_ for details.
+
+ Args:
+ in_channels (int): Number of input channels.
+ out_channels (int): Number of output channels.
+ num_blocks (int, optional): Number of DyHead Blocks. Default: 6.
+ zero_init_offset (bool, optional): Whether to use zero init for
+ `spatial_conv_offset`. Default: True.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None.
+ """
+
+ def __init__(self,
+ in_channels,
+ out_channels,
+ num_blocks=6,
+ zero_init_offset=True,
+ init_cfg=None):
+ assert init_cfg is None, 'To prevent abnormal initialization ' \
+ 'behavior, init_cfg is not allowed to be set'
+ super().__init__(init_cfg=init_cfg)
+ self.in_channels = in_channels
+ self.out_channels = out_channels
+ self.num_blocks = num_blocks
+ self.zero_init_offset = zero_init_offset
+
+ dyhead_blocks = []
+ for i in range(num_blocks):
+ in_channels = self.in_channels if i == 0 else self.out_channels
+ dyhead_blocks.append(
+ DyHeadBlock(
+ in_channels,
+ self.out_channels,
+ zero_init_offset=zero_init_offset))
+ self.dyhead_blocks = nn.Sequential(*dyhead_blocks)
+
+ def forward(self, inputs):
+ """Forward function."""
+ assert isinstance(inputs, (tuple, list))
+ outs = self.dyhead_blocks(inputs)
+ return tuple(outs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/fpg.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/fpg.py
new file mode 100644
index 0000000000000000000000000000000000000000..73ee799bb83645ab2556fe871dcd8b1c5bbff89e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/fpg.py
@@ -0,0 +1,406 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule
+from mmengine.model import BaseModule
+
+from mmdet.registry import MODELS
+
+
+class Transition(BaseModule):
+ """Base class for transition.
+
+ Args:
+ in_channels (int): Number of input channels.
+ out_channels (int): Number of output channels.
+ """
+
+ def __init__(self, in_channels, out_channels, init_cfg=None):
+ super().__init__(init_cfg)
+ self.in_channels = in_channels
+ self.out_channels = out_channels
+
+ def forward(x):
+ pass
+
+
+class UpInterpolationConv(Transition):
+ """A transition used for up-sampling.
+
+ Up-sample the input by interpolation then refines the feature by
+ a convolution layer.
+
+ Args:
+ in_channels (int): Number of input channels.
+ out_channels (int): Number of output channels.
+ scale_factor (int): Up-sampling factor. Default: 2.
+ mode (int): Interpolation mode. Default: nearest.
+ align_corners (bool): Whether align corners when interpolation.
+ Default: None.
+ kernel_size (int): Kernel size for the conv. Default: 3.
+ """
+
+ def __init__(self,
+ in_channels,
+ out_channels,
+ scale_factor=2,
+ mode='nearest',
+ align_corners=None,
+ kernel_size=3,
+ init_cfg=None,
+ **kwargs):
+ super().__init__(in_channels, out_channels, init_cfg)
+ self.mode = mode
+ self.scale_factor = scale_factor
+ self.align_corners = align_corners
+ self.conv = ConvModule(
+ in_channels,
+ out_channels,
+ kernel_size,
+ padding=(kernel_size - 1) // 2,
+ **kwargs)
+
+ def forward(self, x):
+ x = F.interpolate(
+ x,
+ scale_factor=self.scale_factor,
+ mode=self.mode,
+ align_corners=self.align_corners)
+ x = self.conv(x)
+ return x
+
+
+class LastConv(Transition):
+ """A transition used for refining the output of the last stage.
+
+ Args:
+ in_channels (int): Number of input channels.
+ out_channels (int): Number of output channels.
+ num_inputs (int): Number of inputs of the FPN features.
+ kernel_size (int): Kernel size for the conv. Default: 3.
+ """
+
+ def __init__(self,
+ in_channels,
+ out_channels,
+ num_inputs,
+ kernel_size=3,
+ init_cfg=None,
+ **kwargs):
+ super().__init__(in_channels, out_channels, init_cfg)
+ self.num_inputs = num_inputs
+ self.conv_out = ConvModule(
+ in_channels,
+ out_channels,
+ kernel_size,
+ padding=(kernel_size - 1) // 2,
+ **kwargs)
+
+ def forward(self, inputs):
+ assert len(inputs) == self.num_inputs
+ return self.conv_out(inputs[-1])
+
+
+@MODELS.register_module()
+class FPG(BaseModule):
+ """FPG.
+
+ Implementation of `Feature Pyramid Grids (FPG)
+ `_.
+ This implementation only gives the basic structure stated in the paper.
+ But users can implement different type of transitions to fully explore the
+ the potential power of the structure of FPG.
+
+ Args:
+ in_channels (int): Number of input channels (feature maps of all levels
+ should have the same channels).
+ out_channels (int): Number of output channels (used at each scale)
+ num_outs (int): Number of output scales.
+ stack_times (int): The number of times the pyramid architecture will
+ be stacked.
+ paths (list[str]): Specify the path order of each stack level.
+ Each element in the list should be either 'bu' (bottom-up) or
+ 'td' (top-down).
+ inter_channels (int): Number of inter channels.
+ same_up_trans (dict): Transition that goes down at the same stage.
+ same_down_trans (dict): Transition that goes up at the same stage.
+ across_lateral_trans (dict): Across-pathway same-stage
+ across_down_trans (dict): Across-pathway bottom-up connection.
+ across_up_trans (dict): Across-pathway top-down connection.
+ across_skip_trans (dict): Across-pathway skip connection.
+ output_trans (dict): Transition that trans the output of the
+ last stage.
+ start_level (int): Index of the start input backbone level used to
+ build the feature pyramid. Default: 0.
+ end_level (int): Index of the end input backbone level (exclusive) to
+ build the feature pyramid. Default: -1, which means the last level.
+ add_extra_convs (bool): It decides whether to add conv
+ layers on top of the original feature maps. Default to False.
+ If True, its actual mode is specified by `extra_convs_on_inputs`.
+ norm_cfg (dict): Config dict for normalization layer. Default: None.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ """
+
+ transition_types = {
+ 'conv': ConvModule,
+ 'interpolation_conv': UpInterpolationConv,
+ 'last_conv': LastConv,
+ }
+
+ def __init__(self,
+ in_channels,
+ out_channels,
+ num_outs,
+ stack_times,
+ paths,
+ inter_channels=None,
+ same_down_trans=None,
+ same_up_trans=dict(
+ type='conv', kernel_size=3, stride=2, padding=1),
+ across_lateral_trans=dict(type='conv', kernel_size=1),
+ across_down_trans=dict(type='conv', kernel_size=3),
+ across_up_trans=None,
+ across_skip_trans=dict(type='identity'),
+ output_trans=dict(type='last_conv', kernel_size=3),
+ start_level=0,
+ end_level=-1,
+ add_extra_convs=False,
+ norm_cfg=None,
+ skip_inds=None,
+ init_cfg=[
+ dict(type='Caffe2Xavier', layer='Conv2d'),
+ dict(
+ type='Constant',
+ layer=[
+ '_BatchNorm', '_InstanceNorm', 'GroupNorm',
+ 'LayerNorm'
+ ],
+ val=1.0)
+ ]):
+ super(FPG, self).__init__(init_cfg)
+ assert isinstance(in_channels, list)
+ self.in_channels = in_channels
+ self.out_channels = out_channels
+ self.num_ins = len(in_channels)
+ self.num_outs = num_outs
+ if inter_channels is None:
+ self.inter_channels = [out_channels for _ in range(num_outs)]
+ elif isinstance(inter_channels, int):
+ self.inter_channels = [inter_channels for _ in range(num_outs)]
+ else:
+ assert isinstance(inter_channels, list)
+ assert len(inter_channels) == num_outs
+ self.inter_channels = inter_channels
+ self.stack_times = stack_times
+ self.paths = paths
+ assert isinstance(paths, list) and len(paths) == stack_times
+ for d in paths:
+ assert d in ('bu', 'td')
+
+ self.same_down_trans = same_down_trans
+ self.same_up_trans = same_up_trans
+ self.across_lateral_trans = across_lateral_trans
+ self.across_down_trans = across_down_trans
+ self.across_up_trans = across_up_trans
+ self.output_trans = output_trans
+ self.across_skip_trans = across_skip_trans
+
+ self.with_bias = norm_cfg is None
+ # skip inds must be specified if across skip trans is not None
+ if self.across_skip_trans is not None:
+ skip_inds is not None
+ self.skip_inds = skip_inds
+ assert len(self.skip_inds[0]) <= self.stack_times
+
+ if end_level == -1 or end_level == self.num_ins - 1:
+ self.backbone_end_level = self.num_ins
+ assert num_outs >= self.num_ins - start_level
+ else:
+ # if end_level is not the last level, no extra level is allowed
+ self.backbone_end_level = end_level + 1
+ assert end_level < self.num_ins
+ assert num_outs == end_level - start_level + 1
+ self.start_level = start_level
+ self.end_level = end_level
+ self.add_extra_convs = add_extra_convs
+
+ # build lateral 1x1 convs to reduce channels
+ self.lateral_convs = nn.ModuleList()
+ for i in range(self.start_level, self.backbone_end_level):
+ l_conv = nn.Conv2d(self.in_channels[i],
+ self.inter_channels[i - self.start_level], 1)
+ self.lateral_convs.append(l_conv)
+
+ extra_levels = num_outs - self.backbone_end_level + self.start_level
+ self.extra_downsamples = nn.ModuleList()
+ for i in range(extra_levels):
+ if self.add_extra_convs:
+ fpn_idx = self.backbone_end_level - self.start_level + i
+ extra_conv = nn.Conv2d(
+ self.inter_channels[fpn_idx - 1],
+ self.inter_channels[fpn_idx],
+ 3,
+ stride=2,
+ padding=1)
+ self.extra_downsamples.append(extra_conv)
+ else:
+ self.extra_downsamples.append(nn.MaxPool2d(1, stride=2))
+
+ self.fpn_transitions = nn.ModuleList() # stack times
+ for s in range(self.stack_times):
+ stage_trans = nn.ModuleList() # num of feature levels
+ for i in range(self.num_outs):
+ # same, across_lateral, across_down, across_up
+ trans = nn.ModuleDict()
+ if s in self.skip_inds[i]:
+ stage_trans.append(trans)
+ continue
+ # build same-stage down trans (used in bottom-up paths)
+ if i == 0 or self.same_up_trans is None:
+ same_up_trans = None
+ else:
+ same_up_trans = self.build_trans(
+ self.same_up_trans, self.inter_channels[i - 1],
+ self.inter_channels[i])
+ trans['same_up'] = same_up_trans
+ # build same-stage up trans (used in top-down paths)
+ if i == self.num_outs - 1 or self.same_down_trans is None:
+ same_down_trans = None
+ else:
+ same_down_trans = self.build_trans(
+ self.same_down_trans, self.inter_channels[i + 1],
+ self.inter_channels[i])
+ trans['same_down'] = same_down_trans
+ # build across lateral trans
+ across_lateral_trans = self.build_trans(
+ self.across_lateral_trans, self.inter_channels[i],
+ self.inter_channels[i])
+ trans['across_lateral'] = across_lateral_trans
+ # build across down trans
+ if i == self.num_outs - 1 or self.across_down_trans is None:
+ across_down_trans = None
+ else:
+ across_down_trans = self.build_trans(
+ self.across_down_trans, self.inter_channels[i + 1],
+ self.inter_channels[i])
+ trans['across_down'] = across_down_trans
+ # build across up trans
+ if i == 0 or self.across_up_trans is None:
+ across_up_trans = None
+ else:
+ across_up_trans = self.build_trans(
+ self.across_up_trans, self.inter_channels[i - 1],
+ self.inter_channels[i])
+ trans['across_up'] = across_up_trans
+ if self.across_skip_trans is None:
+ across_skip_trans = None
+ else:
+ across_skip_trans = self.build_trans(
+ self.across_skip_trans, self.inter_channels[i - 1],
+ self.inter_channels[i])
+ trans['across_skip'] = across_skip_trans
+ # build across_skip trans
+ stage_trans.append(trans)
+ self.fpn_transitions.append(stage_trans)
+
+ self.output_transition = nn.ModuleList() # output levels
+ for i in range(self.num_outs):
+ trans = self.build_trans(
+ self.output_trans,
+ self.inter_channels[i],
+ self.out_channels,
+ num_inputs=self.stack_times + 1)
+ self.output_transition.append(trans)
+
+ self.relu = nn.ReLU(inplace=True)
+
+ def build_trans(self, cfg, in_channels, out_channels, **extra_args):
+ cfg_ = cfg.copy()
+ trans_type = cfg_.pop('type')
+ trans_cls = self.transition_types[trans_type]
+ return trans_cls(in_channels, out_channels, **cfg_, **extra_args)
+
+ def fuse(self, fuse_dict):
+ out = None
+ for item in fuse_dict.values():
+ if item is not None:
+ if out is None:
+ out = item
+ else:
+ out = out + item
+ return out
+
+ def forward(self, inputs):
+ assert len(inputs) == len(self.in_channels)
+
+ # build all levels from original feature maps
+ feats = [
+ lateral_conv(inputs[i + self.start_level])
+ for i, lateral_conv in enumerate(self.lateral_convs)
+ ]
+ for downsample in self.extra_downsamples:
+ feats.append(downsample(feats[-1]))
+
+ outs = [feats]
+
+ for i in range(self.stack_times):
+ current_outs = outs[-1]
+ next_outs = []
+ direction = self.paths[i]
+ for j in range(self.num_outs):
+ if i in self.skip_inds[j]:
+ next_outs.append(outs[-1][j])
+ continue
+ # feature level
+ if direction == 'td':
+ lvl = self.num_outs - j - 1
+ else:
+ lvl = j
+ # get transitions
+ if direction == 'td':
+ same_trans = self.fpn_transitions[i][lvl]['same_down']
+ else:
+ same_trans = self.fpn_transitions[i][lvl]['same_up']
+ across_lateral_trans = self.fpn_transitions[i][lvl][
+ 'across_lateral']
+ across_down_trans = self.fpn_transitions[i][lvl]['across_down']
+ across_up_trans = self.fpn_transitions[i][lvl]['across_up']
+ across_skip_trans = self.fpn_transitions[i][lvl]['across_skip']
+ # init output
+ to_fuse = dict(
+ same=None, lateral=None, across_up=None, across_down=None)
+ # same downsample/upsample
+ if same_trans is not None:
+ to_fuse['same'] = same_trans(next_outs[-1])
+ # across lateral
+ if across_lateral_trans is not None:
+ to_fuse['lateral'] = across_lateral_trans(
+ current_outs[lvl])
+ # across downsample
+ if lvl > 0 and across_up_trans is not None:
+ to_fuse['across_up'] = across_up_trans(current_outs[lvl -
+ 1])
+ # across upsample
+ if (lvl < self.num_outs - 1 and across_down_trans is not None):
+ to_fuse['across_down'] = across_down_trans(
+ current_outs[lvl + 1])
+ if across_skip_trans is not None:
+ to_fuse['across_skip'] = across_skip_trans(outs[0][lvl])
+ x = self.fuse(to_fuse)
+ next_outs.append(x)
+
+ if direction == 'td':
+ outs.append(next_outs[::-1])
+ else:
+ outs.append(next_outs)
+
+ # output trans
+ final_outs = []
+ for i in range(self.num_outs):
+ lvl_out_list = []
+ for s in range(len(outs)):
+ lvl_out_list.append(outs[s][i])
+ lvl_out = self.output_transition[i](lvl_out_list)
+ final_outs.append(lvl_out)
+
+ return final_outs
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/fpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/fpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..67bd8879641f8539f329e6ffb94f88d25e417244
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/fpn.py
@@ -0,0 +1,221 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple, Union
+
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule
+from mmengine.model import BaseModule
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, MultiConfig, OptConfigType
+
+
+@MODELS.register_module()
+class FPN(BaseModule):
+ r"""Feature Pyramid Network.
+
+ This is an implementation of paper `Feature Pyramid Networks for Object
+ Detection `_.
+
+ Args:
+ in_channels (list[int]): Number of input channels per scale.
+ out_channels (int): Number of output channels (used at each scale).
+ num_outs (int): Number of output scales.
+ start_level (int): Index of the start input backbone level used to
+ build the feature pyramid. Defaults to 0.
+ end_level (int): Index of the end input backbone level (exclusive) to
+ build the feature pyramid. Defaults to -1, which means the
+ last level.
+ add_extra_convs (bool | str): If bool, it decides whether to add conv
+ layers on top of the original feature maps. Defaults to False.
+ If True, it is equivalent to `add_extra_convs='on_input'`.
+ If str, it specifies the source feature map of the extra convs.
+ Only the following options are allowed
+
+ - 'on_input': Last feat map of neck inputs (i.e. backbone feature).
+ - 'on_lateral': Last feature map after lateral convs.
+ - 'on_output': The last output feature map after fpn convs.
+ relu_before_extra_convs (bool): Whether to apply relu before the extra
+ conv. Defaults to False.
+ no_norm_on_lateral (bool): Whether to apply norm on lateral.
+ Defaults to False.
+ conv_cfg (:obj:`ConfigDict` or dict, optional): Config dict for
+ convolution layer. Defaults to None.
+ norm_cfg (:obj:`ConfigDict` or dict, optional): Config dict for
+ normalization layer. Defaults to None.
+ act_cfg (:obj:`ConfigDict` or dict, optional): Config dict for
+ activation layer in ConvModule. Defaults to None.
+ upsample_cfg (:obj:`ConfigDict` or dict, optional): Config dict
+ for interpolate layer. Defaults to dict(mode='nearest').
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict]): Initialization config dict.
+
+ Example:
+ >>> import torch
+ >>> in_channels = [2, 3, 5, 7]
+ >>> scales = [340, 170, 84, 43]
+ >>> inputs = [torch.rand(1, c, s, s)
+ ... for c, s in zip(in_channels, scales)]
+ >>> self = FPN(in_channels, 11, len(in_channels)).eval()
+ >>> outputs = self.forward(inputs)
+ >>> for i in range(len(outputs)):
+ ... print(f'outputs[{i}].shape = {outputs[i].shape}')
+ outputs[0].shape = torch.Size([1, 11, 340, 340])
+ outputs[1].shape = torch.Size([1, 11, 170, 170])
+ outputs[2].shape = torch.Size([1, 11, 84, 84])
+ outputs[3].shape = torch.Size([1, 11, 43, 43])
+ """
+
+ def __init__(
+ self,
+ in_channels: List[int],
+ out_channels: int,
+ num_outs: int,
+ start_level: int = 0,
+ end_level: int = -1,
+ add_extra_convs: Union[bool, str] = False,
+ relu_before_extra_convs: bool = False,
+ no_norm_on_lateral: bool = False,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: OptConfigType = None,
+ act_cfg: OptConfigType = None,
+ upsample_cfg: ConfigType = dict(mode='nearest'),
+ init_cfg: MultiConfig = dict(
+ type='Xavier', layer='Conv2d', distribution='uniform')
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ assert isinstance(in_channels, list)
+ self.in_channels = in_channels
+ self.out_channels = out_channels
+ self.num_ins = len(in_channels)
+ self.num_outs = num_outs
+ self.relu_before_extra_convs = relu_before_extra_convs
+ self.no_norm_on_lateral = no_norm_on_lateral
+ self.fp16_enabled = False
+ self.upsample_cfg = upsample_cfg.copy()
+
+ if end_level == -1 or end_level == self.num_ins - 1:
+ self.backbone_end_level = self.num_ins
+ assert num_outs >= self.num_ins - start_level
+ else:
+ # if end_level is not the last level, no extra level is allowed
+ self.backbone_end_level = end_level + 1
+ assert end_level < self.num_ins
+ assert num_outs == end_level - start_level + 1
+ self.start_level = start_level
+ self.end_level = end_level
+ self.add_extra_convs = add_extra_convs
+ assert isinstance(add_extra_convs, (str, bool))
+ if isinstance(add_extra_convs, str):
+ # Extra_convs_source choices: 'on_input', 'on_lateral', 'on_output'
+ assert add_extra_convs in ('on_input', 'on_lateral', 'on_output')
+ elif add_extra_convs: # True
+ self.add_extra_convs = 'on_input'
+
+ self.lateral_convs = nn.ModuleList()
+ self.fpn_convs = nn.ModuleList()
+
+ for i in range(self.start_level, self.backbone_end_level):
+ l_conv = ConvModule(
+ in_channels[i],
+ out_channels,
+ 1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg if not self.no_norm_on_lateral else None,
+ act_cfg=act_cfg,
+ inplace=False)
+ fpn_conv = ConvModule(
+ out_channels,
+ out_channels,
+ 3,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg,
+ inplace=False)
+
+ self.lateral_convs.append(l_conv)
+ self.fpn_convs.append(fpn_conv)
+
+ # add extra conv layers (e.g., RetinaNet)
+ extra_levels = num_outs - self.backbone_end_level + self.start_level
+ if self.add_extra_convs and extra_levels >= 1:
+ for i in range(extra_levels):
+ if i == 0 and self.add_extra_convs == 'on_input':
+ in_channels = self.in_channels[self.backbone_end_level - 1]
+ else:
+ in_channels = out_channels
+ extra_fpn_conv = ConvModule(
+ in_channels,
+ out_channels,
+ 3,
+ stride=2,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg,
+ inplace=False)
+ self.fpn_convs.append(extra_fpn_conv)
+
+ def forward(self, inputs: Tuple[Tensor]) -> tuple:
+ """Forward function.
+
+ Args:
+ inputs (tuple[Tensor]): Features from the upstream network, each
+ is a 4D-tensor.
+
+ Returns:
+ tuple: Feature maps, each is a 4D-tensor.
+ """
+ assert len(inputs) == len(self.in_channels)
+
+ # build laterals
+ laterals = [
+ lateral_conv(inputs[i + self.start_level])
+ for i, lateral_conv in enumerate(self.lateral_convs)
+ ]
+
+ # build top-down path
+ used_backbone_levels = len(laterals)
+ for i in range(used_backbone_levels - 1, 0, -1):
+ # In some cases, fixing `scale factor` (e.g. 2) is preferred, but
+ # it cannot co-exist with `size` in `F.interpolate`.
+ if 'scale_factor' in self.upsample_cfg:
+ # fix runtime error of "+=" inplace operation in PyTorch 1.10
+ laterals[i - 1] = laterals[i - 1] + F.interpolate(
+ laterals[i], **self.upsample_cfg)
+ else:
+ prev_shape = laterals[i - 1].shape[2:]
+ laterals[i - 1] = laterals[i - 1] + F.interpolate(
+ laterals[i], size=prev_shape, **self.upsample_cfg)
+
+ # build outputs
+ # part 1: from original levels
+ outs = [
+ self.fpn_convs[i](laterals[i]) for i in range(used_backbone_levels)
+ ]
+ # part 2: add extra levels
+ if self.num_outs > len(outs):
+ # use max pool to get more levels on top of outputs
+ # (e.g., Faster R-CNN, Mask R-CNN)
+ if not self.add_extra_convs:
+ for i in range(self.num_outs - used_backbone_levels):
+ outs.append(F.max_pool2d(outs[-1], 1, stride=2))
+ # add conv layers on top of original feature maps (RetinaNet)
+ else:
+ if self.add_extra_convs == 'on_input':
+ extra_source = inputs[self.backbone_end_level - 1]
+ elif self.add_extra_convs == 'on_lateral':
+ extra_source = laterals[-1]
+ elif self.add_extra_convs == 'on_output':
+ extra_source = outs[-1]
+ else:
+ raise NotImplementedError
+ outs.append(self.fpn_convs[used_backbone_levels](extra_source))
+ for i in range(used_backbone_levels + 1, self.num_outs):
+ if self.relu_before_extra_convs:
+ outs.append(self.fpn_convs[i](F.relu(outs[-1])))
+ else:
+ outs.append(self.fpn_convs[i](outs[-1]))
+ return tuple(outs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/fpn_carafe.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/fpn_carafe.py
new file mode 100644
index 0000000000000000000000000000000000000000..b393ff7c340c0c343fc4c91a4d87d341f66a3177
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/fpn_carafe.py
@@ -0,0 +1,275 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch.nn as nn
+from mmcv.cnn import ConvModule, build_upsample_layer
+from mmcv.ops.carafe import CARAFEPack
+from mmengine.model import BaseModule, ModuleList, xavier_init
+
+from mmdet.registry import MODELS
+
+
+@MODELS.register_module()
+class FPN_CARAFE(BaseModule):
+ """FPN_CARAFE is a more flexible implementation of FPN. It allows more
+ choice for upsample methods during the top-down pathway.
+
+ It can reproduce the performance of ICCV 2019 paper
+ CARAFE: Content-Aware ReAssembly of FEatures
+ Please refer to https://arxiv.org/abs/1905.02188 for more details.
+
+ Args:
+ in_channels (list[int]): Number of channels for each input feature map.
+ out_channels (int): Output channels of feature pyramids.
+ num_outs (int): Number of output stages.
+ start_level (int): Start level of feature pyramids.
+ (Default: 0)
+ end_level (int): End level of feature pyramids.
+ (Default: -1 indicates the last level).
+ norm_cfg (dict): Dictionary to construct and config norm layer.
+ activate (str): Type of activation function in ConvModule
+ (Default: None indicates w/o activation).
+ order (dict): Order of components in ConvModule.
+ upsample (str): Type of upsample layer.
+ upsample_cfg (dict): Dictionary to construct and config upsample layer.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None
+ """
+
+ def __init__(self,
+ in_channels,
+ out_channels,
+ num_outs,
+ start_level=0,
+ end_level=-1,
+ norm_cfg=None,
+ act_cfg=None,
+ order=('conv', 'norm', 'act'),
+ upsample_cfg=dict(
+ type='carafe',
+ up_kernel=5,
+ up_group=1,
+ encoder_kernel=3,
+ encoder_dilation=1),
+ init_cfg=None):
+ assert init_cfg is None, 'To prevent abnormal initialization ' \
+ 'behavior, init_cfg is not allowed to be set'
+ super(FPN_CARAFE, self).__init__(init_cfg)
+ assert isinstance(in_channels, list)
+ self.in_channels = in_channels
+ self.out_channels = out_channels
+ self.num_ins = len(in_channels)
+ self.num_outs = num_outs
+ self.norm_cfg = norm_cfg
+ self.act_cfg = act_cfg
+ self.with_bias = norm_cfg is None
+ self.upsample_cfg = upsample_cfg.copy()
+ self.upsample = self.upsample_cfg.get('type')
+ self.relu = nn.ReLU(inplace=False)
+
+ self.order = order
+ assert order in [('conv', 'norm', 'act'), ('act', 'conv', 'norm')]
+
+ assert self.upsample in [
+ 'nearest', 'bilinear', 'deconv', 'pixel_shuffle', 'carafe', None
+ ]
+ if self.upsample in ['deconv', 'pixel_shuffle']:
+ assert hasattr(
+ self.upsample_cfg,
+ 'upsample_kernel') and self.upsample_cfg.upsample_kernel > 0
+ self.upsample_kernel = self.upsample_cfg.pop('upsample_kernel')
+
+ if end_level == -1 or end_level == self.num_ins - 1:
+ self.backbone_end_level = self.num_ins
+ assert num_outs >= self.num_ins - start_level
+ else:
+ # if end_level is not the last level, no extra level is allowed
+ self.backbone_end_level = end_level + 1
+ assert end_level < self.num_ins
+ assert num_outs == end_level - start_level + 1
+ self.start_level = start_level
+ self.end_level = end_level
+
+ self.lateral_convs = ModuleList()
+ self.fpn_convs = ModuleList()
+ self.upsample_modules = ModuleList()
+
+ for i in range(self.start_level, self.backbone_end_level):
+ l_conv = ConvModule(
+ in_channels[i],
+ out_channels,
+ 1,
+ norm_cfg=norm_cfg,
+ bias=self.with_bias,
+ act_cfg=act_cfg,
+ inplace=False,
+ order=self.order)
+ fpn_conv = ConvModule(
+ out_channels,
+ out_channels,
+ 3,
+ padding=1,
+ norm_cfg=self.norm_cfg,
+ bias=self.with_bias,
+ act_cfg=act_cfg,
+ inplace=False,
+ order=self.order)
+ if i != self.backbone_end_level - 1:
+ upsample_cfg_ = self.upsample_cfg.copy()
+ if self.upsample == 'deconv':
+ upsample_cfg_.update(
+ in_channels=out_channels,
+ out_channels=out_channels,
+ kernel_size=self.upsample_kernel,
+ stride=2,
+ padding=(self.upsample_kernel - 1) // 2,
+ output_padding=(self.upsample_kernel - 1) // 2)
+ elif self.upsample == 'pixel_shuffle':
+ upsample_cfg_.update(
+ in_channels=out_channels,
+ out_channels=out_channels,
+ scale_factor=2,
+ upsample_kernel=self.upsample_kernel)
+ elif self.upsample == 'carafe':
+ upsample_cfg_.update(channels=out_channels, scale_factor=2)
+ else:
+ # suppress warnings
+ align_corners = (None
+ if self.upsample == 'nearest' else False)
+ upsample_cfg_.update(
+ scale_factor=2,
+ mode=self.upsample,
+ align_corners=align_corners)
+ upsample_module = build_upsample_layer(upsample_cfg_)
+ self.upsample_modules.append(upsample_module)
+ self.lateral_convs.append(l_conv)
+ self.fpn_convs.append(fpn_conv)
+
+ # add extra conv layers (e.g., RetinaNet)
+ extra_out_levels = (
+ num_outs - self.backbone_end_level + self.start_level)
+ if extra_out_levels >= 1:
+ for i in range(extra_out_levels):
+ in_channels = (
+ self.in_channels[self.backbone_end_level -
+ 1] if i == 0 else out_channels)
+ extra_l_conv = ConvModule(
+ in_channels,
+ out_channels,
+ 3,
+ stride=2,
+ padding=1,
+ norm_cfg=norm_cfg,
+ bias=self.with_bias,
+ act_cfg=act_cfg,
+ inplace=False,
+ order=self.order)
+ if self.upsample == 'deconv':
+ upsampler_cfg_ = dict(
+ in_channels=out_channels,
+ out_channels=out_channels,
+ kernel_size=self.upsample_kernel,
+ stride=2,
+ padding=(self.upsample_kernel - 1) // 2,
+ output_padding=(self.upsample_kernel - 1) // 2)
+ elif self.upsample == 'pixel_shuffle':
+ upsampler_cfg_ = dict(
+ in_channels=out_channels,
+ out_channels=out_channels,
+ scale_factor=2,
+ upsample_kernel=self.upsample_kernel)
+ elif self.upsample == 'carafe':
+ upsampler_cfg_ = dict(
+ channels=out_channels,
+ scale_factor=2,
+ **self.upsample_cfg)
+ else:
+ # suppress warnings
+ align_corners = (None
+ if self.upsample == 'nearest' else False)
+ upsampler_cfg_ = dict(
+ scale_factor=2,
+ mode=self.upsample,
+ align_corners=align_corners)
+ upsampler_cfg_['type'] = self.upsample
+ upsample_module = build_upsample_layer(upsampler_cfg_)
+ extra_fpn_conv = ConvModule(
+ out_channels,
+ out_channels,
+ 3,
+ padding=1,
+ norm_cfg=self.norm_cfg,
+ bias=self.with_bias,
+ act_cfg=act_cfg,
+ inplace=False,
+ order=self.order)
+ self.upsample_modules.append(upsample_module)
+ self.fpn_convs.append(extra_fpn_conv)
+ self.lateral_convs.append(extra_l_conv)
+
+ # default init_weights for conv(msra) and norm in ConvModule
+ def init_weights(self):
+ """Initialize the weights of module."""
+ super(FPN_CARAFE, self).init_weights()
+ for m in self.modules():
+ if isinstance(m, (nn.Conv2d, nn.ConvTranspose2d)):
+ xavier_init(m, distribution='uniform')
+ for m in self.modules():
+ if isinstance(m, CARAFEPack):
+ m.init_weights()
+
+ def slice_as(self, src, dst):
+ """Slice ``src`` as ``dst``
+
+ Note:
+ ``src`` should have the same or larger size than ``dst``.
+
+ Args:
+ src (torch.Tensor): Tensors to be sliced.
+ dst (torch.Tensor): ``src`` will be sliced to have the same
+ size as ``dst``.
+
+ Returns:
+ torch.Tensor: Sliced tensor.
+ """
+ assert (src.size(2) >= dst.size(2)) and (src.size(3) >= dst.size(3))
+ if src.size(2) == dst.size(2) and src.size(3) == dst.size(3):
+ return src
+ else:
+ return src[:, :, :dst.size(2), :dst.size(3)]
+
+ def tensor_add(self, a, b):
+ """Add tensors ``a`` and ``b`` that might have different sizes."""
+ if a.size() == b.size():
+ c = a + b
+ else:
+ c = a + self.slice_as(b, a)
+ return c
+
+ def forward(self, inputs):
+ """Forward function."""
+ assert len(inputs) == len(self.in_channels)
+
+ # build laterals
+ laterals = []
+ for i, lateral_conv in enumerate(self.lateral_convs):
+ if i <= self.backbone_end_level - self.start_level:
+ input = inputs[min(i + self.start_level, len(inputs) - 1)]
+ else:
+ input = laterals[-1]
+ lateral = lateral_conv(input)
+ laterals.append(lateral)
+
+ # build top-down path
+ for i in range(len(laterals) - 1, 0, -1):
+ if self.upsample is not None:
+ upsample_feat = self.upsample_modules[i - 1](laterals[i])
+ else:
+ upsample_feat = laterals[i]
+ laterals[i - 1] = self.tensor_add(laterals[i - 1], upsample_feat)
+
+ # build outputs
+ num_conv_outs = len(self.fpn_convs)
+ outs = []
+ for i in range(num_conv_outs):
+ out = self.fpn_convs[i](laterals[i])
+ outs.append(out)
+ return tuple(outs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/fpn_dropblock.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/fpn_dropblock.py
new file mode 100644
index 0000000000000000000000000000000000000000..473af924cdaaecf88aa4a0a6e1500511530b91a2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/fpn_dropblock.py
@@ -0,0 +1,90 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Tuple
+
+import torch.nn.functional as F
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from .fpn import FPN
+
+
+@MODELS.register_module()
+class FPN_DropBlock(FPN):
+
+ def __init__(self,
+ *args,
+ plugin: Optional[dict] = dict(
+ type='DropBlock',
+ drop_prob=0.3,
+ block_size=3,
+ warmup_iters=0),
+ **kwargs) -> None:
+ super().__init__(*args, **kwargs)
+ self.plugin = None
+ if plugin is not None:
+ self.plugin = MODELS.build(plugin)
+
+ def forward(self, inputs: Tuple[Tensor]) -> tuple:
+ """Forward function.
+
+ Args:
+ inputs (tuple[Tensor]): Features from the upstream network, each
+ is a 4D-tensor.
+
+ Returns:
+ tuple: Feature maps, each is a 4D-tensor.
+ """
+ assert len(inputs) == len(self.in_channels)
+
+ # build laterals
+ laterals = [
+ lateral_conv(inputs[i + self.start_level])
+ for i, lateral_conv in enumerate(self.lateral_convs)
+ ]
+
+ # build top-down path
+ used_backbone_levels = len(laterals)
+ for i in range(used_backbone_levels - 1, 0, -1):
+ # In some cases, fixing `scale factor` (e.g. 2) is preferred, but
+ # it cannot co-exist with `size` in `F.interpolate`.
+ if 'scale_factor' in self.upsample_cfg:
+ # fix runtime error of "+=" inplace operation in PyTorch 1.10
+ laterals[i - 1] = laterals[i - 1] + F.interpolate(
+ laterals[i], **self.upsample_cfg)
+ else:
+ prev_shape = laterals[i - 1].shape[2:]
+ laterals[i - 1] = laterals[i - 1] + F.interpolate(
+ laterals[i], size=prev_shape, **self.upsample_cfg)
+
+ if self.plugin is not None:
+ laterals[i - 1] = self.plugin(laterals[i - 1])
+
+ # build outputs
+ # part 1: from original levels
+ outs = [
+ self.fpn_convs[i](laterals[i]) for i in range(used_backbone_levels)
+ ]
+ # part 2: add extra levels
+ if self.num_outs > len(outs):
+ # use max pool to get more levels on top of outputs
+ # (e.g., Faster R-CNN, Mask R-CNN)
+ if not self.add_extra_convs:
+ for i in range(self.num_outs - used_backbone_levels):
+ outs.append(F.max_pool2d(outs[-1], 1, stride=2))
+ # add conv layers on top of original feature maps (RetinaNet)
+ else:
+ if self.add_extra_convs == 'on_input':
+ extra_source = inputs[self.backbone_end_level - 1]
+ elif self.add_extra_convs == 'on_lateral':
+ extra_source = laterals[-1]
+ elif self.add_extra_convs == 'on_output':
+ extra_source = outs[-1]
+ else:
+ raise NotImplementedError
+ outs.append(self.fpn_convs[used_backbone_levels](extra_source))
+ for i in range(used_backbone_levels + 1, self.num_outs):
+ if self.relu_before_extra_convs:
+ outs.append(self.fpn_convs[i](F.relu(outs[-1])))
+ else:
+ outs.append(self.fpn_convs[i](outs[-1]))
+ return tuple(outs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/hrfpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/hrfpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..d2627549b4cb8acc6833bc40425e459c28aa5c20
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/hrfpn.py
@@ -0,0 +1,100 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule
+from mmengine.model import BaseModule
+from torch.utils.checkpoint import checkpoint
+
+from mmdet.registry import MODELS
+
+
+@MODELS.register_module()
+class HRFPN(BaseModule):
+ """HRFPN (High Resolution Feature Pyramids)
+
+ paper: `High-Resolution Representations for Labeling Pixels and Regions
+ `_.
+
+ Args:
+ in_channels (list): number of channels for each branch.
+ out_channels (int): output channels of feature pyramids.
+ num_outs (int): number of output stages.
+ pooling_type (str): pooling for generating feature pyramids
+ from {MAX, AVG}.
+ conv_cfg (dict): dictionary to construct and config conv layer.
+ norm_cfg (dict): dictionary to construct and config norm layer.
+ with_cp (bool): Use checkpoint or not. Using checkpoint will save some
+ memory while slowing down the training speed.
+ stride (int): stride of 3x3 convolutional layers
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ """
+
+ def __init__(self,
+ in_channels,
+ out_channels,
+ num_outs=5,
+ pooling_type='AVG',
+ conv_cfg=None,
+ norm_cfg=None,
+ with_cp=False,
+ stride=1,
+ init_cfg=dict(type='Caffe2Xavier', layer='Conv2d')):
+ super(HRFPN, self).__init__(init_cfg)
+ assert isinstance(in_channels, list)
+ self.in_channels = in_channels
+ self.out_channels = out_channels
+ self.num_ins = len(in_channels)
+ self.num_outs = num_outs
+ self.with_cp = with_cp
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+
+ self.reduction_conv = ConvModule(
+ sum(in_channels),
+ out_channels,
+ kernel_size=1,
+ conv_cfg=self.conv_cfg,
+ act_cfg=None)
+
+ self.fpn_convs = nn.ModuleList()
+ for i in range(self.num_outs):
+ self.fpn_convs.append(
+ ConvModule(
+ out_channels,
+ out_channels,
+ kernel_size=3,
+ padding=1,
+ stride=stride,
+ conv_cfg=self.conv_cfg,
+ act_cfg=None))
+
+ if pooling_type == 'MAX':
+ self.pooling = F.max_pool2d
+ else:
+ self.pooling = F.avg_pool2d
+
+ def forward(self, inputs):
+ """Forward function."""
+ assert len(inputs) == self.num_ins
+ outs = [inputs[0]]
+ for i in range(1, self.num_ins):
+ outs.append(
+ F.interpolate(inputs[i], scale_factor=2**i, mode='bilinear'))
+ out = torch.cat(outs, dim=1)
+ if out.requires_grad and self.with_cp:
+ out = checkpoint(self.reduction_conv, out)
+ else:
+ out = self.reduction_conv(out)
+ outs = [out]
+ for i in range(1, self.num_outs):
+ outs.append(self.pooling(out, kernel_size=2**i, stride=2**i))
+ outputs = []
+
+ for i in range(self.num_outs):
+ if outs[i].requires_grad and self.with_cp:
+ tmp_out = checkpoint(self.fpn_convs[i], outs[i])
+ else:
+ tmp_out = self.fpn_convs[i](outs[i])
+ outputs.append(tmp_out)
+ return tuple(outputs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/nas_fpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/nas_fpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..8ec90cd6eed3aa65a3a192d332cbfd8c16d5bc36
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/nas_fpn.py
@@ -0,0 +1,171 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple
+
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from mmcv.ops.merge_cells import GlobalPoolingCell, SumCell
+from mmengine.model import BaseModule, ModuleList
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import MultiConfig, OptConfigType
+
+
+@MODELS.register_module()
+class NASFPN(BaseModule):
+ """NAS-FPN.
+
+ Implementation of `NAS-FPN: Learning Scalable Feature Pyramid Architecture
+ for Object Detection `_
+
+ Args:
+ in_channels (List[int]): Number of input channels per scale.
+ out_channels (int): Number of output channels (used at each scale)
+ num_outs (int): Number of output scales.
+ stack_times (int): The number of times the pyramid architecture will
+ be stacked.
+ start_level (int): Index of the start input backbone level used to
+ build the feature pyramid. Defaults to 0.
+ end_level (int): Index of the end input backbone level (exclusive) to
+ build the feature pyramid. Defaults to -1, which means the
+ last level.
+ norm_cfg (:obj:`ConfigDict` or dict, optional): Config dict for
+ normalization layer. Defaults to None.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict]): Initialization config dict.
+ """
+
+ def __init__(
+ self,
+ in_channels: List[int],
+ out_channels: int,
+ num_outs: int,
+ stack_times: int,
+ start_level: int = 0,
+ end_level: int = -1,
+ norm_cfg: OptConfigType = None,
+ init_cfg: MultiConfig = dict(type='Caffe2Xavier', layer='Conv2d')
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ assert isinstance(in_channels, list)
+ self.in_channels = in_channels
+ self.out_channels = out_channels
+ self.num_ins = len(in_channels) # num of input feature levels
+ self.num_outs = num_outs # num of output feature levels
+ self.stack_times = stack_times
+ self.norm_cfg = norm_cfg
+
+ if end_level == -1 or end_level == self.num_ins - 1:
+ self.backbone_end_level = self.num_ins
+ assert num_outs >= self.num_ins - start_level
+ else:
+ # if end_level is not the last level, no extra level is allowed
+ self.backbone_end_level = end_level + 1
+ assert end_level < self.num_ins
+ assert num_outs == end_level - start_level + 1
+ self.start_level = start_level
+ self.end_level = end_level
+
+ # add lateral connections
+ self.lateral_convs = nn.ModuleList()
+ for i in range(self.start_level, self.backbone_end_level):
+ l_conv = ConvModule(
+ in_channels[i],
+ out_channels,
+ 1,
+ norm_cfg=norm_cfg,
+ act_cfg=None)
+ self.lateral_convs.append(l_conv)
+
+ # add extra downsample layers (stride-2 pooling or conv)
+ extra_levels = num_outs - self.backbone_end_level + self.start_level
+ self.extra_downsamples = nn.ModuleList()
+ for i in range(extra_levels):
+ extra_conv = ConvModule(
+ out_channels, out_channels, 1, norm_cfg=norm_cfg, act_cfg=None)
+ self.extra_downsamples.append(
+ nn.Sequential(extra_conv, nn.MaxPool2d(2, 2)))
+
+ # add NAS FPN connections
+ self.fpn_stages = ModuleList()
+ for _ in range(self.stack_times):
+ stage = nn.ModuleDict()
+ # gp(p6, p4) -> p4_1
+ stage['gp_64_4'] = GlobalPoolingCell(
+ in_channels=out_channels,
+ out_channels=out_channels,
+ out_norm_cfg=norm_cfg)
+ # sum(p4_1, p4) -> p4_2
+ stage['sum_44_4'] = SumCell(
+ in_channels=out_channels,
+ out_channels=out_channels,
+ out_norm_cfg=norm_cfg)
+ # sum(p4_2, p3) -> p3_out
+ stage['sum_43_3'] = SumCell(
+ in_channels=out_channels,
+ out_channels=out_channels,
+ out_norm_cfg=norm_cfg)
+ # sum(p3_out, p4_2) -> p4_out
+ stage['sum_34_4'] = SumCell(
+ in_channels=out_channels,
+ out_channels=out_channels,
+ out_norm_cfg=norm_cfg)
+ # sum(p5, gp(p4_out, p3_out)) -> p5_out
+ stage['gp_43_5'] = GlobalPoolingCell(with_out_conv=False)
+ stage['sum_55_5'] = SumCell(
+ in_channels=out_channels,
+ out_channels=out_channels,
+ out_norm_cfg=norm_cfg)
+ # sum(p7, gp(p5_out, p4_2)) -> p7_out
+ stage['gp_54_7'] = GlobalPoolingCell(with_out_conv=False)
+ stage['sum_77_7'] = SumCell(
+ in_channels=out_channels,
+ out_channels=out_channels,
+ out_norm_cfg=norm_cfg)
+ # gp(p7_out, p5_out) -> p6_out
+ stage['gp_75_6'] = GlobalPoolingCell(
+ in_channels=out_channels,
+ out_channels=out_channels,
+ out_norm_cfg=norm_cfg)
+ self.fpn_stages.append(stage)
+
+ def forward(self, inputs: Tuple[Tensor]) -> tuple:
+ """Forward function.
+
+ Args:
+ inputs (tuple[Tensor]): Features from the upstream network, each
+ is a 4D-tensor.
+
+ Returns:
+ tuple: Feature maps, each is a 4D-tensor.
+ """
+ # build P3-P5
+ feats = [
+ lateral_conv(inputs[i + self.start_level])
+ for i, lateral_conv in enumerate(self.lateral_convs)
+ ]
+ # build P6-P7 on top of P5
+ for downsample in self.extra_downsamples:
+ feats.append(downsample(feats[-1]))
+
+ p3, p4, p5, p6, p7 = feats
+
+ for stage in self.fpn_stages:
+ # gp(p6, p4) -> p4_1
+ p4_1 = stage['gp_64_4'](p6, p4, out_size=p4.shape[-2:])
+ # sum(p4_1, p4) -> p4_2
+ p4_2 = stage['sum_44_4'](p4_1, p4, out_size=p4.shape[-2:])
+ # sum(p4_2, p3) -> p3_out
+ p3 = stage['sum_43_3'](p4_2, p3, out_size=p3.shape[-2:])
+ # sum(p3_out, p4_2) -> p4_out
+ p4 = stage['sum_34_4'](p3, p4_2, out_size=p4.shape[-2:])
+ # sum(p5, gp(p4_out, p3_out)) -> p5_out
+ p5_tmp = stage['gp_43_5'](p4, p3, out_size=p5.shape[-2:])
+ p5 = stage['sum_55_5'](p5, p5_tmp, out_size=p5.shape[-2:])
+ # sum(p7, gp(p5_out, p4_2)) -> p7_out
+ p7_tmp = stage['gp_54_7'](p5, p4_2, out_size=p7.shape[-2:])
+ p7 = stage['sum_77_7'](p7, p7_tmp, out_size=p7.shape[-2:])
+ # gp(p7_out, p5_out) -> p6_out
+ p6 = stage['gp_75_6'](p7, p5, out_size=p6.shape[-2:])
+
+ return p3, p4, p5, p6, p7
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/nasfcos_fpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/nasfcos_fpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..12d0848f7634bb0113e0b5a16b5b65ba8b7ebb9c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/nasfcos_fpn.py
@@ -0,0 +1,170 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule
+from mmcv.ops.merge_cells import ConcatCell
+from mmengine.model import BaseModule, caffe2_xavier_init
+
+from mmdet.registry import MODELS
+
+
+@MODELS.register_module()
+class NASFCOS_FPN(BaseModule):
+ """FPN structure in NASFPN.
+
+ Implementation of paper `NAS-FCOS: Fast Neural Architecture Search for
+ Object Detection `_
+
+ Args:
+ in_channels (List[int]): Number of input channels per scale.
+ out_channels (int): Number of output channels (used at each scale)
+ num_outs (int): Number of output scales.
+ start_level (int): Index of the start input backbone level used to
+ build the feature pyramid. Default: 0.
+ end_level (int): Index of the end input backbone level (exclusive) to
+ build the feature pyramid. Default: -1, which means the last level.
+ add_extra_convs (bool): It decides whether to add conv
+ layers on top of the original feature maps. Default to False.
+ If True, its actual mode is specified by `extra_convs_on_inputs`.
+ conv_cfg (dict): dictionary to construct and config conv layer.
+ norm_cfg (dict): dictionary to construct and config norm layer.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None
+ """
+
+ def __init__(self,
+ in_channels,
+ out_channels,
+ num_outs,
+ start_level=1,
+ end_level=-1,
+ add_extra_convs=False,
+ conv_cfg=None,
+ norm_cfg=None,
+ init_cfg=None):
+ assert init_cfg is None, 'To prevent abnormal initialization ' \
+ 'behavior, init_cfg is not allowed to be set'
+ super(NASFCOS_FPN, self).__init__(init_cfg)
+ assert isinstance(in_channels, list)
+ self.in_channels = in_channels
+ self.out_channels = out_channels
+ self.num_ins = len(in_channels)
+ self.num_outs = num_outs
+ self.norm_cfg = norm_cfg
+ self.conv_cfg = conv_cfg
+
+ if end_level == -1 or end_level == self.num_ins - 1:
+ self.backbone_end_level = self.num_ins
+ assert num_outs >= self.num_ins - start_level
+ else:
+ # if end_level is not the last level, no extra level is allowed
+ self.backbone_end_level = end_level + 1
+ assert end_level < self.num_ins
+ assert num_outs == end_level - start_level + 1
+ self.start_level = start_level
+ self.end_level = end_level
+ self.add_extra_convs = add_extra_convs
+
+ self.adapt_convs = nn.ModuleList()
+ for i in range(self.start_level, self.backbone_end_level):
+ adapt_conv = ConvModule(
+ in_channels[i],
+ out_channels,
+ 1,
+ stride=1,
+ padding=0,
+ bias=False,
+ norm_cfg=dict(type='BN'),
+ act_cfg=dict(type='ReLU', inplace=False))
+ self.adapt_convs.append(adapt_conv)
+
+ # C2 is omitted according to the paper
+ extra_levels = num_outs - self.backbone_end_level + self.start_level
+
+ def build_concat_cell(with_input1_conv, with_input2_conv):
+ cell_conv_cfg = dict(
+ kernel_size=1, padding=0, bias=False, groups=out_channels)
+ return ConcatCell(
+ in_channels=out_channels,
+ out_channels=out_channels,
+ with_out_conv=True,
+ out_conv_cfg=cell_conv_cfg,
+ out_norm_cfg=dict(type='BN'),
+ out_conv_order=('norm', 'act', 'conv'),
+ with_input1_conv=with_input1_conv,
+ with_input2_conv=with_input2_conv,
+ input_conv_cfg=conv_cfg,
+ input_norm_cfg=norm_cfg,
+ upsample_mode='nearest')
+
+ # Denote c3=f0, c4=f1, c5=f2 for convince
+ self.fpn = nn.ModuleDict()
+ self.fpn['c22_1'] = build_concat_cell(True, True)
+ self.fpn['c22_2'] = build_concat_cell(True, True)
+ self.fpn['c32'] = build_concat_cell(True, False)
+ self.fpn['c02'] = build_concat_cell(True, False)
+ self.fpn['c42'] = build_concat_cell(True, True)
+ self.fpn['c36'] = build_concat_cell(True, True)
+ self.fpn['c61'] = build_concat_cell(True, True) # f9
+ self.extra_downsamples = nn.ModuleList()
+ for i in range(extra_levels):
+ extra_act_cfg = None if i == 0 \
+ else dict(type='ReLU', inplace=False)
+ self.extra_downsamples.append(
+ ConvModule(
+ out_channels,
+ out_channels,
+ 3,
+ stride=2,
+ padding=1,
+ act_cfg=extra_act_cfg,
+ order=('act', 'norm', 'conv')))
+
+ def forward(self, inputs):
+ """Forward function."""
+ feats = [
+ adapt_conv(inputs[i + self.start_level])
+ for i, adapt_conv in enumerate(self.adapt_convs)
+ ]
+
+ for (i, module_name) in enumerate(self.fpn):
+ idx_1, idx_2 = int(module_name[1]), int(module_name[2])
+ res = self.fpn[module_name](feats[idx_1], feats[idx_2])
+ feats.append(res)
+
+ ret = []
+ for (idx, input_idx) in zip([9, 8, 7], [1, 2, 3]): # add P3, P4, P5
+ feats1, feats2 = feats[idx], feats[5]
+ feats2_resize = F.interpolate(
+ feats2,
+ size=feats1.size()[2:],
+ mode='bilinear',
+ align_corners=False)
+
+ feats_sum = feats1 + feats2_resize
+ ret.append(
+ F.interpolate(
+ feats_sum,
+ size=inputs[input_idx].size()[2:],
+ mode='bilinear',
+ align_corners=False))
+
+ for submodule in self.extra_downsamples:
+ ret.append(submodule(ret[-1]))
+
+ return tuple(ret)
+
+ def init_weights(self):
+ """Initialize the weights of module."""
+ super(NASFCOS_FPN, self).init_weights()
+ for module in self.fpn.values():
+ if hasattr(module, 'conv_out'):
+ caffe2_xavier_init(module.out_conv.conv)
+
+ for modules in [
+ self.adapt_convs.modules(),
+ self.extra_downsamples.modules()
+ ]:
+ for module in modules:
+ if isinstance(module, nn.Conv2d):
+ caffe2_xavier_init(module)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/pafpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/pafpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..557638f48a629691f780d3e1466e234bbe987518
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/pafpn.py
@@ -0,0 +1,157 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule
+
+from mmdet.registry import MODELS
+from .fpn import FPN
+
+
+@MODELS.register_module()
+class PAFPN(FPN):
+ """Path Aggregation Network for Instance Segmentation.
+
+ This is an implementation of the `PAFPN in Path Aggregation Network
+ `_.
+
+ Args:
+ in_channels (List[int]): Number of input channels per scale.
+ out_channels (int): Number of output channels (used at each scale)
+ num_outs (int): Number of output scales.
+ start_level (int): Index of the start input backbone level used to
+ build the feature pyramid. Default: 0.
+ end_level (int): Index of the end input backbone level (exclusive) to
+ build the feature pyramid. Default: -1, which means the last level.
+ add_extra_convs (bool | str): If bool, it decides whether to add conv
+ layers on top of the original feature maps. Default to False.
+ If True, it is equivalent to `add_extra_convs='on_input'`.
+ If str, it specifies the source feature map of the extra convs.
+ Only the following options are allowed
+
+ - 'on_input': Last feat map of neck inputs (i.e. backbone feature).
+ - 'on_lateral': Last feature map after lateral convs.
+ - 'on_output': The last output feature map after fpn convs.
+ relu_before_extra_convs (bool): Whether to apply relu before the extra
+ conv. Default: False.
+ no_norm_on_lateral (bool): Whether to apply norm on lateral.
+ Default: False.
+ conv_cfg (dict): Config dict for convolution layer. Default: None.
+ norm_cfg (dict): Config dict for normalization layer. Default: None.
+ act_cfg (str): Config dict for activation layer in ConvModule.
+ Default: None.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ """
+
+ def __init__(self,
+ in_channels,
+ out_channels,
+ num_outs,
+ start_level=0,
+ end_level=-1,
+ add_extra_convs=False,
+ relu_before_extra_convs=False,
+ no_norm_on_lateral=False,
+ conv_cfg=None,
+ norm_cfg=None,
+ act_cfg=None,
+ init_cfg=dict(
+ type='Xavier', layer='Conv2d', distribution='uniform')):
+ super(PAFPN, self).__init__(
+ in_channels,
+ out_channels,
+ num_outs,
+ start_level,
+ end_level,
+ add_extra_convs,
+ relu_before_extra_convs,
+ no_norm_on_lateral,
+ conv_cfg,
+ norm_cfg,
+ act_cfg,
+ init_cfg=init_cfg)
+ # add extra bottom up pathway
+ self.downsample_convs = nn.ModuleList()
+ self.pafpn_convs = nn.ModuleList()
+ for i in range(self.start_level + 1, self.backbone_end_level):
+ d_conv = ConvModule(
+ out_channels,
+ out_channels,
+ 3,
+ stride=2,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg,
+ inplace=False)
+ pafpn_conv = ConvModule(
+ out_channels,
+ out_channels,
+ 3,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg,
+ inplace=False)
+ self.downsample_convs.append(d_conv)
+ self.pafpn_convs.append(pafpn_conv)
+
+ def forward(self, inputs):
+ """Forward function."""
+ assert len(inputs) == len(self.in_channels)
+
+ # build laterals
+ laterals = [
+ lateral_conv(inputs[i + self.start_level])
+ for i, lateral_conv in enumerate(self.lateral_convs)
+ ]
+
+ # build top-down path
+ used_backbone_levels = len(laterals)
+ for i in range(used_backbone_levels - 1, 0, -1):
+ prev_shape = laterals[i - 1].shape[2:]
+ laterals[i - 1] = laterals[i - 1] + F.interpolate(
+ laterals[i], size=prev_shape, mode='nearest')
+
+ # build outputs
+ # part 1: from original levels
+ inter_outs = [
+ self.fpn_convs[i](laterals[i]) for i in range(used_backbone_levels)
+ ]
+
+ # part 2: add bottom-up path
+ for i in range(0, used_backbone_levels - 1):
+ inter_outs[i + 1] = inter_outs[i + 1] + \
+ self.downsample_convs[i](inter_outs[i])
+
+ outs = []
+ outs.append(inter_outs[0])
+ outs.extend([
+ self.pafpn_convs[i - 1](inter_outs[i])
+ for i in range(1, used_backbone_levels)
+ ])
+
+ # part 3: add extra levels
+ if self.num_outs > len(outs):
+ # use max pool to get more levels on top of outputs
+ # (e.g., Faster R-CNN, Mask R-CNN)
+ if not self.add_extra_convs:
+ for i in range(self.num_outs - used_backbone_levels):
+ outs.append(F.max_pool2d(outs[-1], 1, stride=2))
+ # add conv layers on top of original feature maps (RetinaNet)
+ else:
+ if self.add_extra_convs == 'on_input':
+ orig = inputs[self.backbone_end_level - 1]
+ outs.append(self.fpn_convs[used_backbone_levels](orig))
+ elif self.add_extra_convs == 'on_lateral':
+ outs.append(self.fpn_convs[used_backbone_levels](
+ laterals[-1]))
+ elif self.add_extra_convs == 'on_output':
+ outs.append(self.fpn_convs[used_backbone_levels](outs[-1]))
+ else:
+ raise NotImplementedError
+ for i in range(used_backbone_levels + 1, self.num_outs):
+ if self.relu_before_extra_convs:
+ outs.append(self.fpn_convs[i](F.relu(outs[-1])))
+ else:
+ outs.append(self.fpn_convs[i](outs[-1]))
+ return tuple(outs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/rfp.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/rfp.py
new file mode 100644
index 0000000000000000000000000000000000000000..7ec9b3753c5031bb12a2b4c88733f13bf27c44e2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/rfp.py
@@ -0,0 +1,134 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmengine.model import BaseModule, ModuleList, constant_init, xavier_init
+
+from mmdet.registry import MODELS
+from .fpn import FPN
+
+
+class ASPP(BaseModule):
+ """ASPP (Atrous Spatial Pyramid Pooling)
+
+ This is an implementation of the ASPP module used in DetectoRS
+ (https://arxiv.org/pdf/2006.02334.pdf)
+
+ Args:
+ in_channels (int): Number of input channels.
+ out_channels (int): Number of channels produced by this module
+ dilations (tuple[int]): Dilations of the four branches.
+ Default: (1, 3, 6, 1)
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ """
+
+ def __init__(self,
+ in_channels,
+ out_channels,
+ dilations=(1, 3, 6, 1),
+ init_cfg=dict(type='Kaiming', layer='Conv2d')):
+ super().__init__(init_cfg)
+ assert dilations[-1] == 1
+ self.aspp = nn.ModuleList()
+ for dilation in dilations:
+ kernel_size = 3 if dilation > 1 else 1
+ padding = dilation if dilation > 1 else 0
+ conv = nn.Conv2d(
+ in_channels,
+ out_channels,
+ kernel_size=kernel_size,
+ stride=1,
+ dilation=dilation,
+ padding=padding,
+ bias=True)
+ self.aspp.append(conv)
+ self.gap = nn.AdaptiveAvgPool2d(1)
+
+ def forward(self, x):
+ avg_x = self.gap(x)
+ out = []
+ for aspp_idx in range(len(self.aspp)):
+ inp = avg_x if (aspp_idx == len(self.aspp) - 1) else x
+ out.append(F.relu_(self.aspp[aspp_idx](inp)))
+ out[-1] = out[-1].expand_as(out[-2])
+ out = torch.cat(out, dim=1)
+ return out
+
+
+@MODELS.register_module()
+class RFP(FPN):
+ """RFP (Recursive Feature Pyramid)
+
+ This is an implementation of RFP in `DetectoRS
+ `_. Different from standard FPN, the
+ input of RFP should be multi level features along with origin input image
+ of backbone.
+
+ Args:
+ rfp_steps (int): Number of unrolled steps of RFP.
+ rfp_backbone (dict): Configuration of the backbone for RFP.
+ aspp_out_channels (int): Number of output channels of ASPP module.
+ aspp_dilations (tuple[int]): Dilation rates of four branches.
+ Default: (1, 3, 6, 1)
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None
+ """
+
+ def __init__(self,
+ rfp_steps,
+ rfp_backbone,
+ aspp_out_channels,
+ aspp_dilations=(1, 3, 6, 1),
+ init_cfg=None,
+ **kwargs):
+ assert init_cfg is None, 'To prevent abnormal initialization ' \
+ 'behavior, init_cfg is not allowed to be set'
+ super().__init__(init_cfg=init_cfg, **kwargs)
+ self.rfp_steps = rfp_steps
+ # Be careful! Pretrained weights cannot be loaded when use
+ # nn.ModuleList
+ self.rfp_modules = ModuleList()
+ for rfp_idx in range(1, rfp_steps):
+ rfp_module = MODELS.build(rfp_backbone)
+ self.rfp_modules.append(rfp_module)
+ self.rfp_aspp = ASPP(self.out_channels, aspp_out_channels,
+ aspp_dilations)
+ self.rfp_weight = nn.Conv2d(
+ self.out_channels,
+ 1,
+ kernel_size=1,
+ stride=1,
+ padding=0,
+ bias=True)
+
+ def init_weights(self):
+ # Avoid using super().init_weights(), which may alter the default
+ # initialization of the modules in self.rfp_modules that have missing
+ # keys in the pretrained checkpoint.
+ for convs in [self.lateral_convs, self.fpn_convs]:
+ for m in convs.modules():
+ if isinstance(m, nn.Conv2d):
+ xavier_init(m, distribution='uniform')
+ for rfp_idx in range(self.rfp_steps - 1):
+ self.rfp_modules[rfp_idx].init_weights()
+ constant_init(self.rfp_weight, 0)
+
+ def forward(self, inputs):
+ inputs = list(inputs)
+ assert len(inputs) == len(self.in_channels) + 1 # +1 for input image
+ img = inputs.pop(0)
+ # FPN forward
+ x = super().forward(tuple(inputs))
+ for rfp_idx in range(self.rfp_steps - 1):
+ rfp_feats = [x[0]] + list(
+ self.rfp_aspp(x[i]) for i in range(1, len(x)))
+ x_idx = self.rfp_modules[rfp_idx].rfp_forward(img, rfp_feats)
+ # FPN forward
+ x_idx = super().forward(x_idx)
+ x_new = []
+ for ft_idx in range(len(x_idx)):
+ add_weight = torch.sigmoid(self.rfp_weight(x_idx[ft_idx]))
+ x_new.append(add_weight * x_idx[ft_idx] +
+ (1 - add_weight) * x[ft_idx])
+ x = x_new
+ return x
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/ssd_neck.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/ssd_neck.py
new file mode 100644
index 0000000000000000000000000000000000000000..17ba319370b988b9c7e2d98c2f10607ff8f8b5c3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/ssd_neck.py
@@ -0,0 +1,129 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule, DepthwiseSeparableConvModule
+from mmengine.model import BaseModule
+
+from mmdet.registry import MODELS
+
+
+@MODELS.register_module()
+class SSDNeck(BaseModule):
+ """Extra layers of SSD backbone to generate multi-scale feature maps.
+
+ Args:
+ in_channels (Sequence[int]): Number of input channels per scale.
+ out_channels (Sequence[int]): Number of output channels per scale.
+ level_strides (Sequence[int]): Stride of 3x3 conv per level.
+ level_paddings (Sequence[int]): Padding size of 3x3 conv per level.
+ l2_norm_scale (float|None): L2 normalization layer init scale.
+ If None, not use L2 normalization on the first input feature.
+ last_kernel_size (int): Kernel size of the last conv layer.
+ Default: 3.
+ use_depthwise (bool): Whether to use DepthwiseSeparableConv.
+ Default: False.
+ conv_cfg (dict): Config dict for convolution layer. Default: None.
+ norm_cfg (dict): Dictionary to construct and config norm layer.
+ Default: None.
+ act_cfg (dict): Config dict for activation layer.
+ Default: dict(type='ReLU').
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ """
+
+ def __init__(self,
+ in_channels,
+ out_channels,
+ level_strides,
+ level_paddings,
+ l2_norm_scale=20.,
+ last_kernel_size=3,
+ use_depthwise=False,
+ conv_cfg=None,
+ norm_cfg=None,
+ act_cfg=dict(type='ReLU'),
+ init_cfg=[
+ dict(
+ type='Xavier', distribution='uniform',
+ layer='Conv2d'),
+ dict(type='Constant', val=1, layer='BatchNorm2d'),
+ ]):
+ super(SSDNeck, self).__init__(init_cfg)
+ assert len(out_channels) > len(in_channels)
+ assert len(out_channels) - len(in_channels) == len(level_strides)
+ assert len(level_strides) == len(level_paddings)
+ assert in_channels == out_channels[:len(in_channels)]
+
+ if l2_norm_scale:
+ self.l2_norm = L2Norm(in_channels[0], l2_norm_scale)
+ self.init_cfg += [
+ dict(
+ type='Constant',
+ val=self.l2_norm.scale,
+ override=dict(name='l2_norm'))
+ ]
+
+ self.extra_layers = nn.ModuleList()
+ extra_layer_channels = out_channels[len(in_channels):]
+ second_conv = DepthwiseSeparableConvModule if \
+ use_depthwise else ConvModule
+
+ for i, (out_channel, stride, padding) in enumerate(
+ zip(extra_layer_channels, level_strides, level_paddings)):
+ kernel_size = last_kernel_size \
+ if i == len(extra_layer_channels) - 1 else 3
+ per_lvl_convs = nn.Sequential(
+ ConvModule(
+ out_channels[len(in_channels) - 1 + i],
+ out_channel // 2,
+ 1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg),
+ second_conv(
+ out_channel // 2,
+ out_channel,
+ kernel_size,
+ stride=stride,
+ padding=padding,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg))
+ self.extra_layers.append(per_lvl_convs)
+
+ def forward(self, inputs):
+ """Forward function."""
+ outs = [feat for feat in inputs]
+ if hasattr(self, 'l2_norm'):
+ outs[0] = self.l2_norm(outs[0])
+
+ feat = outs[-1]
+ for layer in self.extra_layers:
+ feat = layer(feat)
+ outs.append(feat)
+ return tuple(outs)
+
+
+class L2Norm(nn.Module):
+
+ def __init__(self, n_dims, scale=20., eps=1e-10):
+ """L2 normalization layer.
+
+ Args:
+ n_dims (int): Number of dimensions to be normalized
+ scale (float, optional): Defaults to 20..
+ eps (float, optional): Used to avoid division by zero.
+ Defaults to 1e-10.
+ """
+ super(L2Norm, self).__init__()
+ self.n_dims = n_dims
+ self.weight = nn.Parameter(torch.Tensor(self.n_dims))
+ self.eps = eps
+ self.scale = scale
+
+ def forward(self, x):
+ """Forward function."""
+ # normalization layer convert to FP32 in FP16 training
+ x_float = x.float()
+ norm = x_float.pow(2).sum(1, keepdim=True).sqrt() + self.eps
+ return (self.weight[None, :, None, None].float().expand_as(x_float) *
+ x_float / norm).type_as(x)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/ssh.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/ssh.py
new file mode 100644
index 0000000000000000000000000000000000000000..75a6561489d8d3634fc34829dafe819bbf066ed4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/ssh.py
@@ -0,0 +1,216 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple
+
+import torch
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule
+from mmengine.model import BaseModule
+
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+
+
+class SSHContextModule(BaseModule):
+ """This is an implementation of `SSH context module` described in `SSH:
+ Single Stage Headless Face Detector.
+
+ `_.
+
+ Args:
+ in_channels (int): Number of input channels used at each scale.
+ out_channels (int): Number of output channels used at each scale.
+ conv_cfg (:obj:`ConfigDict` or dict, optional): Config dict for
+ convolution layer. Defaults to None.
+ norm_cfg (:obj:`ConfigDict` or dict): Config dict for normalization
+ layer. Defaults to dict(type='BN').
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or
+ list[dict], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ in_channels: int,
+ out_channels: int,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: ConfigType = dict(type='BN'),
+ init_cfg: OptMultiConfig = None):
+ super().__init__(init_cfg=init_cfg)
+ assert out_channels % 4 == 0
+
+ self.in_channels = in_channels
+ self.out_channels = out_channels
+
+ self.conv5x5_1 = ConvModule(
+ self.in_channels,
+ self.out_channels // 4,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ )
+
+ self.conv5x5_2 = ConvModule(
+ self.out_channels // 4,
+ self.out_channels // 4,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=None)
+
+ self.conv7x7_2 = ConvModule(
+ self.out_channels // 4,
+ self.out_channels // 4,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ )
+
+ self.conv7x7_3 = ConvModule(
+ self.out_channels // 4,
+ self.out_channels // 4,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=None,
+ )
+
+ def forward(self, x: torch.Tensor) -> tuple:
+ conv5x5_1 = self.conv5x5_1(x)
+ conv5x5 = self.conv5x5_2(conv5x5_1)
+ conv7x7_2 = self.conv7x7_2(conv5x5_1)
+ conv7x7 = self.conv7x7_3(conv7x7_2)
+
+ return (conv5x5, conv7x7)
+
+
+class SSHDetModule(BaseModule):
+ """This is an implementation of `SSH detection module` described in `SSH:
+ Single Stage Headless Face Detector.
+
+ `_.
+
+ Args:
+ in_channels (int): Number of input channels used at each scale.
+ out_channels (int): Number of output channels used at each scale.
+ conv_cfg (:obj:`ConfigDict` or dict, optional): Config dict for
+ convolution layer. Defaults to None.
+ norm_cfg (:obj:`ConfigDict` or dict): Config dict for normalization
+ layer. Defaults to dict(type='BN').
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or
+ list[dict], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ in_channels: int,
+ out_channels: int,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: ConfigType = dict(type='BN'),
+ init_cfg: OptMultiConfig = None):
+ super().__init__(init_cfg=init_cfg)
+ assert out_channels % 4 == 0
+
+ self.in_channels = in_channels
+ self.out_channels = out_channels
+
+ self.conv3x3 = ConvModule(
+ self.in_channels,
+ self.out_channels // 2,
+ 3,
+ stride=1,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=None)
+
+ self.context_module = SSHContextModule(
+ in_channels=self.in_channels,
+ out_channels=self.out_channels,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg)
+
+ def forward(self, x: torch.Tensor) -> torch.Tensor:
+ conv3x3 = self.conv3x3(x)
+ conv5x5, conv7x7 = self.context_module(x)
+ out = torch.cat([conv3x3, conv5x5, conv7x7], dim=1)
+ out = F.relu(out)
+
+ return out
+
+
+@MODELS.register_module()
+class SSH(BaseModule):
+ """`SSH Neck` used in `SSH: Single Stage Headless Face Detector.
+
+ `_.
+
+ Args:
+ num_scales (int): The number of scales / stages.
+ in_channels (list[int]): The number of input channels per scale.
+ out_channels (list[int]): The number of output channels per scale.
+ conv_cfg (:obj:`ConfigDict` or dict, optional): Config dict for
+ convolution layer. Defaults to None.
+ norm_cfg (:obj:`ConfigDict` or dict): Config dict for normalization
+ layer. Defaults to dict(type='BN').
+ init_cfg (:obj:`ConfigDict` or list[:obj:`ConfigDict`] or dict or
+ list[dict], optional): Initialization config dict.
+
+ Example:
+ >>> import torch
+ >>> in_channels = [8, 16, 32, 64]
+ >>> out_channels = [16, 32, 64, 128]
+ >>> scales = [340, 170, 84, 43]
+ >>> inputs = [torch.rand(1, c, s, s)
+ ... for c, s in zip(in_channels, scales)]
+ >>> self = SSH(num_scales=4, in_channels=in_channels,
+ ... out_channels=out_channels)
+ >>> outputs = self.forward(inputs)
+ >>> for i in range(len(outputs)):
+ ... print(f'outputs[{i}].shape = {outputs[i].shape}')
+ outputs[0].shape = torch.Size([1, 16, 340, 340])
+ outputs[1].shape = torch.Size([1, 32, 170, 170])
+ outputs[2].shape = torch.Size([1, 64, 84, 84])
+ outputs[3].shape = torch.Size([1, 128, 43, 43])
+ """
+
+ def __init__(self,
+ num_scales: int,
+ in_channels: List[int],
+ out_channels: List[int],
+ conv_cfg: OptConfigType = None,
+ norm_cfg: ConfigType = dict(type='BN'),
+ init_cfg: OptMultiConfig = dict(
+ type='Xavier', layer='Conv2d', distribution='uniform')):
+ super().__init__(init_cfg=init_cfg)
+ assert (num_scales == len(in_channels) == len(out_channels))
+ self.num_scales = num_scales
+ self.in_channels = in_channels
+ self.out_channels = out_channels
+
+ for idx in range(self.num_scales):
+ in_c, out_c = self.in_channels[idx], self.out_channels[idx]
+ self.add_module(
+ f'ssh_module{idx}',
+ SSHDetModule(
+ in_channels=in_c,
+ out_channels=out_c,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg))
+
+ def forward(self, inputs: Tuple[torch.Tensor]) -> tuple:
+ assert len(inputs) == self.num_scales
+
+ outs = []
+ for idx, x in enumerate(inputs):
+ ssh_module = getattr(self, f'ssh_module{idx}')
+ out = ssh_module(x)
+ outs.append(out)
+
+ return tuple(outs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/yolo_neck.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/yolo_neck.py
new file mode 100644
index 0000000000000000000000000000000000000000..48a6b1a4897c85083aa1e1e7d692263f66de67c3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/yolo_neck.py
@@ -0,0 +1,145 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+# Copyright (c) 2019 Western Digital Corporation or its affiliates.
+from typing import List, Tuple
+
+import torch
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule
+from mmengine.model import BaseModule
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+
+
+class DetectionBlock(BaseModule):
+ """Detection block in YOLO neck.
+
+ Let out_channels = n, the DetectionBlock contains:
+ Six ConvLayers, 1 Conv2D Layer and 1 YoloLayer.
+ The first 6 ConvLayers are formed the following way:
+ 1x1xn, 3x3x2n, 1x1xn, 3x3x2n, 1x1xn, 3x3x2n.
+ The Conv2D layer is 1x1x255.
+ Some block will have branch after the fifth ConvLayer.
+ The input channel is arbitrary (in_channels)
+
+ Args:
+ in_channels (int): The number of input channels.
+ out_channels (int): The number of output channels.
+ conv_cfg (dict): Config dict for convolution layer. Default: None.
+ norm_cfg (dict): Dictionary to construct and config norm layer.
+ Default: dict(type='BN', requires_grad=True)
+ act_cfg (dict): Config dict for activation layer.
+ Default: dict(type='LeakyReLU', negative_slope=0.1).
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None
+ """
+
+ def __init__(self,
+ in_channels: int,
+ out_channels: int,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: ConfigType = dict(type='BN', requires_grad=True),
+ act_cfg: ConfigType = dict(
+ type='LeakyReLU', negative_slope=0.1),
+ init_cfg: OptMultiConfig = None) -> None:
+ super(DetectionBlock, self).__init__(init_cfg)
+ double_out_channels = out_channels * 2
+
+ # shortcut
+ cfg = dict(conv_cfg=conv_cfg, norm_cfg=norm_cfg, act_cfg=act_cfg)
+ self.conv1 = ConvModule(in_channels, out_channels, 1, **cfg)
+ self.conv2 = ConvModule(
+ out_channels, double_out_channels, 3, padding=1, **cfg)
+ self.conv3 = ConvModule(double_out_channels, out_channels, 1, **cfg)
+ self.conv4 = ConvModule(
+ out_channels, double_out_channels, 3, padding=1, **cfg)
+ self.conv5 = ConvModule(double_out_channels, out_channels, 1, **cfg)
+
+ def forward(self, x: Tensor) -> Tensor:
+ tmp = self.conv1(x)
+ tmp = self.conv2(tmp)
+ tmp = self.conv3(tmp)
+ tmp = self.conv4(tmp)
+ out = self.conv5(tmp)
+ return out
+
+
+@MODELS.register_module()
+class YOLOV3Neck(BaseModule):
+ """The neck of YOLOV3.
+
+ It can be treated as a simplified version of FPN. It
+ will take the result from Darknet backbone and do some upsampling and
+ concatenation. It will finally output the detection result.
+
+ Note:
+ The input feats should be from top to bottom.
+ i.e., from high-lvl to low-lvl
+ But YOLOV3Neck will process them in reversed order.
+ i.e., from bottom (high-lvl) to top (low-lvl)
+
+ Args:
+ num_scales (int): The number of scales / stages.
+ in_channels (List[int]): The number of input channels per scale.
+ out_channels (List[int]): The number of output channels per scale.
+ conv_cfg (dict, optional): Config dict for convolution layer.
+ Default: None.
+ norm_cfg (dict, optional): Dictionary to construct and config norm
+ layer. Default: dict(type='BN', requires_grad=True)
+ act_cfg (dict, optional): Config dict for activation layer.
+ Default: dict(type='LeakyReLU', negative_slope=0.1).
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None
+ """
+
+ def __init__(self,
+ num_scales: int,
+ in_channels: List[int],
+ out_channels: List[int],
+ conv_cfg: OptConfigType = None,
+ norm_cfg: ConfigType = dict(type='BN', requires_grad=True),
+ act_cfg: ConfigType = dict(
+ type='LeakyReLU', negative_slope=0.1),
+ init_cfg: OptMultiConfig = None) -> None:
+ super(YOLOV3Neck, self).__init__(init_cfg)
+ assert (num_scales == len(in_channels) == len(out_channels))
+ self.num_scales = num_scales
+ self.in_channels = in_channels
+ self.out_channels = out_channels
+
+ # shortcut
+ cfg = dict(conv_cfg=conv_cfg, norm_cfg=norm_cfg, act_cfg=act_cfg)
+
+ # To support arbitrary scales, the code looks awful, but it works.
+ # Better solution is welcomed.
+ self.detect1 = DetectionBlock(in_channels[0], out_channels[0], **cfg)
+ for i in range(1, self.num_scales):
+ in_c, out_c = self.in_channels[i], self.out_channels[i]
+ inter_c = out_channels[i - 1]
+ self.add_module(f'conv{i}', ConvModule(inter_c, out_c, 1, **cfg))
+ # in_c + out_c : High-lvl feats will be cat with low-lvl feats
+ self.add_module(f'detect{i+1}',
+ DetectionBlock(in_c + out_c, out_c, **cfg))
+
+ def forward(self, feats=Tuple[Tensor]) -> Tuple[Tensor]:
+ assert len(feats) == self.num_scales
+
+ # processed from bottom (high-lvl) to top (low-lvl)
+ outs = []
+ out = self.detect1(feats[-1])
+ outs.append(out)
+
+ for i, x in enumerate(reversed(feats[:-1])):
+ conv = getattr(self, f'conv{i+1}')
+ tmp = conv(out)
+
+ # Cat with low-lvl feats
+ tmp = F.interpolate(tmp, scale_factor=2)
+ tmp = torch.cat((tmp, x), 1)
+
+ detect = getattr(self, f'detect{i+2}')
+ out = detect(tmp)
+ outs.append(out)
+
+ return tuple(outs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/yolox_pafpn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/yolox_pafpn.py
new file mode 100644
index 0000000000000000000000000000000000000000..8ec3d12bfde8158c1a817fbf223a8eea94798667
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/necks/yolox_pafpn.py
@@ -0,0 +1,156 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import math
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule, DepthwiseSeparableConvModule
+from mmengine.model import BaseModule
+
+from mmdet.registry import MODELS
+from ..layers import CSPLayer
+
+
+@MODELS.register_module()
+class YOLOXPAFPN(BaseModule):
+ """Path Aggregation Network used in YOLOX.
+
+ Args:
+ in_channels (List[int]): Number of input channels per scale.
+ out_channels (int): Number of output channels (used at each scale)
+ num_csp_blocks (int): Number of bottlenecks in CSPLayer. Default: 3
+ use_depthwise (bool): Whether to depthwise separable convolution in
+ blocks. Default: False
+ upsample_cfg (dict): Config dict for interpolate layer.
+ Default: `dict(scale_factor=2, mode='nearest')`
+ conv_cfg (dict, optional): Config dict for convolution layer.
+ Default: None, which means using conv2d.
+ norm_cfg (dict): Config dict for normalization layer.
+ Default: dict(type='BN')
+ act_cfg (dict): Config dict for activation layer.
+ Default: dict(type='Swish')
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Default: None.
+ """
+
+ def __init__(self,
+ in_channels,
+ out_channels,
+ num_csp_blocks=3,
+ use_depthwise=False,
+ upsample_cfg=dict(scale_factor=2, mode='nearest'),
+ conv_cfg=None,
+ norm_cfg=dict(type='BN', momentum=0.03, eps=0.001),
+ act_cfg=dict(type='Swish'),
+ init_cfg=dict(
+ type='Kaiming',
+ layer='Conv2d',
+ a=math.sqrt(5),
+ distribution='uniform',
+ mode='fan_in',
+ nonlinearity='leaky_relu')):
+ super(YOLOXPAFPN, self).__init__(init_cfg)
+ self.in_channels = in_channels
+ self.out_channels = out_channels
+
+ conv = DepthwiseSeparableConvModule if use_depthwise else ConvModule
+
+ # build top-down blocks
+ self.upsample = nn.Upsample(**upsample_cfg)
+ self.reduce_layers = nn.ModuleList()
+ self.top_down_blocks = nn.ModuleList()
+ for idx in range(len(in_channels) - 1, 0, -1):
+ self.reduce_layers.append(
+ ConvModule(
+ in_channels[idx],
+ in_channels[idx - 1],
+ 1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg))
+ self.top_down_blocks.append(
+ CSPLayer(
+ in_channels[idx - 1] * 2,
+ in_channels[idx - 1],
+ num_blocks=num_csp_blocks,
+ add_identity=False,
+ use_depthwise=use_depthwise,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg))
+
+ # build bottom-up blocks
+ self.downsamples = nn.ModuleList()
+ self.bottom_up_blocks = nn.ModuleList()
+ for idx in range(len(in_channels) - 1):
+ self.downsamples.append(
+ conv(
+ in_channels[idx],
+ in_channels[idx],
+ 3,
+ stride=2,
+ padding=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg))
+ self.bottom_up_blocks.append(
+ CSPLayer(
+ in_channels[idx] * 2,
+ in_channels[idx + 1],
+ num_blocks=num_csp_blocks,
+ add_identity=False,
+ use_depthwise=use_depthwise,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg))
+
+ self.out_convs = nn.ModuleList()
+ for i in range(len(in_channels)):
+ self.out_convs.append(
+ ConvModule(
+ in_channels[i],
+ out_channels,
+ 1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg))
+
+ def forward(self, inputs):
+ """
+ Args:
+ inputs (tuple[Tensor]): input features.
+
+ Returns:
+ tuple[Tensor]: YOLOXPAFPN features.
+ """
+ assert len(inputs) == len(self.in_channels)
+
+ # top-down path
+ inner_outs = [inputs[-1]]
+ for idx in range(len(self.in_channels) - 1, 0, -1):
+ feat_heigh = inner_outs[0]
+ feat_low = inputs[idx - 1]
+ feat_heigh = self.reduce_layers[len(self.in_channels) - 1 - idx](
+ feat_heigh)
+ inner_outs[0] = feat_heigh
+
+ upsample_feat = self.upsample(feat_heigh)
+
+ inner_out = self.top_down_blocks[len(self.in_channels) - 1 - idx](
+ torch.cat([upsample_feat, feat_low], 1))
+ inner_outs.insert(0, inner_out)
+
+ # bottom-up path
+ outs = [inner_outs[0]]
+ for idx in range(len(self.in_channels) - 1):
+ feat_low = outs[-1]
+ feat_height = inner_outs[idx + 1]
+ downsample_feat = self.downsamples[idx](feat_low)
+ out = self.bottom_up_blocks[idx](
+ torch.cat([downsample_feat, feat_height], 1))
+ outs.append(out)
+
+ # out convs
+ for idx, conv in enumerate(self.out_convs):
+ outs[idx] = conv(outs[idx])
+
+ return tuple(outs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/reid/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/reid/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..aca617f7dea0b8047891c666ddb684dbbd018c81
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/reid/__init__.py
@@ -0,0 +1,7 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .base_reid import BaseReID
+from .fc_module import FcModule
+from .gap import GlobalAveragePooling
+from .linear_reid_head import LinearReIDHead
+
+__all__ = ['BaseReID', 'GlobalAveragePooling', 'LinearReIDHead', 'FcModule']
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/reid/base_reid.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/reid/base_reid.py
new file mode 100644
index 0000000000000000000000000000000000000000..4c45964394aa1651f846f2a7e63da3ee70b78909
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/reid/base_reid.py
@@ -0,0 +1,65 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional
+
+import torch
+
+try:
+ import mmpretrain
+ from mmpretrain.models.classifiers import ImageClassifier
+except ImportError:
+ mmpretrain = None
+ ImageClassifier = object
+
+from mmdet.registry import MODELS
+from mmdet.structures import ReIDDataSample
+
+
+@MODELS.register_module()
+class BaseReID(ImageClassifier):
+ """Base model for re-identification."""
+
+ def __init__(self, *args, **kwargs):
+ if mmpretrain is None:
+ raise RuntimeError('Please run "pip install openmim" and '
+ 'run "mim install mmpretrain" to '
+ 'install mmpretrain first.')
+ super().__init__(*args, **kwargs)
+
+ def forward(self,
+ inputs: torch.Tensor,
+ data_samples: Optional[List[ReIDDataSample]] = None,
+ mode: str = 'tensor'):
+ """The unified entry for a forward process in both training and test.
+
+ The method should accept three modes: "tensor", "predict" and "loss":
+
+ - "tensor": Forward the whole network and return tensor or tuple of
+ tensor without any post-processing, same as a common nn.Module.
+ - "predict": Forward and return the predictions, which are fully
+ processed to a list of :obj:`ReIDDataSample`.
+ - "loss": Forward and return a dict of losses according to the given
+ inputs and data samples.
+
+ Note that this method doesn't handle neither back propagation nor
+ optimizer updating, which are done in the :meth:`train_step`.
+
+ Args:
+ inputs (torch.Tensor): The input tensor with shape
+ (N, C, H, W) or (N, T, C, H, W).
+ data_samples (List[ReIDDataSample], optional): The annotation
+ data of every sample. It's required if ``mode="loss"``.
+ Defaults to None.
+ mode (str): Return what kind of value. Defaults to 'tensor'.
+
+ Returns:
+ The return type depends on ``mode``.
+
+ - If ``mode="tensor"``, return a tensor or a tuple of tensor.
+ - If ``mode="predict"``, return a list of
+ :obj:`ReIDDataSample`.
+ - If ``mode="loss"``, return a dict of tensor.
+ """
+ if len(inputs.size()) == 5:
+ assert inputs.size(0) == 1
+ inputs = inputs[0]
+ return super().forward(inputs, data_samples, mode)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/reid/fc_module.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/reid/fc_module.py
new file mode 100644
index 0000000000000000000000000000000000000000..76e7efd66e300a242bb250cc6ba5cc68ed722034
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/reid/fc_module.py
@@ -0,0 +1,71 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch.nn as nn
+from mmcv.cnn import build_activation_layer, build_norm_layer
+from mmengine.model import BaseModule
+
+from mmdet.registry import MODELS
+
+
+@MODELS.register_module()
+class FcModule(BaseModule):
+ """Fully-connected layer module.
+
+ Args:
+ in_channels (int): Input channels.
+ out_channels (int): Ourput channels.
+ norm_cfg (dict, optional): Configuration of normlization method
+ after fc. Defaults to None.
+ act_cfg (dict, optional): Configuration of activation method after fc.
+ Defaults to dict(type='ReLU').
+ inplace (bool, optional): Whether inplace the activatation module.
+ Defaults to True.
+ init_cfg (dict, optional): Initialization config dict.
+ Defaults to dict(type='Kaiming', layer='Linear').
+ """
+
+ def __init__(self,
+ in_channels: int,
+ out_channels: int,
+ norm_cfg: dict = None,
+ act_cfg: dict = dict(type='ReLU'),
+ inplace: bool = True,
+ init_cfg=dict(type='Kaiming', layer='Linear')):
+ super(FcModule, self).__init__(init_cfg)
+ assert norm_cfg is None or isinstance(norm_cfg, dict)
+ assert act_cfg is None or isinstance(act_cfg, dict)
+ self.norm_cfg = norm_cfg
+ self.act_cfg = act_cfg
+ self.inplace = inplace
+
+ self.with_norm = norm_cfg is not None
+ self.with_activation = act_cfg is not None
+
+ self.fc = nn.Linear(in_channels, out_channels)
+ # build normalization layers
+ if self.with_norm:
+ self.norm_name, norm = build_norm_layer(norm_cfg, out_channels)
+ self.add_module(self.norm_name, norm)
+
+ # build activation layer
+ if self.with_activation:
+ act_cfg_ = act_cfg.copy()
+ # nn.Tanh has no 'inplace' argument
+ if act_cfg_['type'] not in [
+ 'Tanh', 'PReLU', 'Sigmoid', 'HSigmoid', 'Swish'
+ ]:
+ act_cfg_.setdefault('inplace', inplace)
+ self.activate = build_activation_layer(act_cfg_)
+
+ @property
+ def norm(self):
+ """Normalization."""
+ return getattr(self, self.norm_name)
+
+ def forward(self, x, activate=True, norm=True):
+ """Model forward."""
+ x = self.fc(x)
+ if norm and self.with_norm:
+ x = self.norm(x)
+ if activate and self.with_activation:
+ x = self.activate(x)
+ return x
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/reid/gap.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/reid/gap.py
new file mode 100644
index 0000000000000000000000000000000000000000..aadc25e7144f2ca9efb66b496bf8ffa5504619ff
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/reid/gap.py
@@ -0,0 +1,40 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+import torch.nn as nn
+from mmengine.model import BaseModule
+
+from mmdet.registry import MODELS
+
+
+@MODELS.register_module()
+class GlobalAveragePooling(BaseModule):
+ """Global Average Pooling neck.
+
+ Note that we use `view` to remove extra channel after pooling. We do not
+ use `squeeze` as it will also remove the batch dimension when the tensor
+ has a batch dimension of size 1, which can lead to unexpected errors.
+ """
+
+ def __init__(self, kernel_size=None, stride=None):
+ super(GlobalAveragePooling, self).__init__()
+ if kernel_size is None and stride is None:
+ self.gap = nn.AdaptiveAvgPool2d((1, 1))
+ else:
+ self.gap = nn.AvgPool2d(kernel_size, stride)
+
+ def forward(self, inputs):
+ if isinstance(inputs, tuple):
+ outs = tuple([self.gap(x) for x in inputs])
+ outs = tuple([
+ out.view(x.size(0),
+ torch.tensor(out.size()[1:]).prod())
+ for out, x in zip(outs, inputs)
+ ])
+ elif isinstance(inputs, torch.Tensor):
+ outs = self.gap(inputs)
+ outs = outs.view(
+ inputs.size(0),
+ torch.tensor(outs.size()[1:]).prod())
+ else:
+ raise TypeError('neck inputs should be tuple or torch.tensor')
+ return outs
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/reid/linear_reid_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/reid/linear_reid_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..f35aaf6c2fc57b60e36017268e2a632df60ed342
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/reid/linear_reid_head.py
@@ -0,0 +1,202 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+from typing import List, Optional, Tuple, Union
+
+import torch
+import torch.nn as nn
+
+try:
+ import mmpretrain
+ from mmpretrain.evaluation.metrics import Accuracy
+except ImportError:
+ mmpretrain = None
+
+from mmengine.model import BaseModule
+
+from mmdet.registry import MODELS
+from mmdet.structures import ReIDDataSample
+from .fc_module import FcModule
+
+
+@MODELS.register_module()
+class LinearReIDHead(BaseModule):
+ """Linear head for re-identification.
+
+ Args:
+ num_fcs (int): Number of fcs.
+ in_channels (int): Number of channels in the input.
+ fc_channels (int): Number of channels in the fcs.
+ out_channels (int): Number of channels in the output.
+ norm_cfg (dict, optional): Configuration of normlization method
+ after fc. Defaults to None.
+ act_cfg (dict, optional): Configuration of activation method after fc.
+ Defaults to None.
+ num_classes (int, optional): Number of the identities. Default to None.
+ loss_cls (dict, optional): Cross entropy loss to train the ReID module.
+ Defaults to None.
+ loss_triplet (dict, optional): Triplet loss to train the ReID module.
+ Defaults to None.
+ topk (int | Tuple[int]): Top-k accuracy. Defaults to ``(1, )``.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Defaults to dict(type='Normal',layer='Linear', mean=0, std=0.01,
+ bias=0).
+ """
+
+ def __init__(self,
+ num_fcs: int,
+ in_channels: int,
+ fc_channels: int,
+ out_channels: int,
+ norm_cfg: Optional[dict] = None,
+ act_cfg: Optional[dict] = None,
+ num_classes: Optional[int] = None,
+ loss_cls: Optional[dict] = None,
+ loss_triplet: Optional[dict] = None,
+ topk: Union[int, Tuple[int]] = (1, ),
+ init_cfg: Union[dict, List[dict]] = dict(
+ type='Normal', layer='Linear', mean=0, std=0.01, bias=0)):
+ if mmpretrain is None:
+ raise RuntimeError('Please run "pip install openmim" and '
+ 'run "mim install mmpretrain" to '
+ 'install mmpretrain first.')
+ super(LinearReIDHead, self).__init__(init_cfg=init_cfg)
+
+ assert isinstance(topk, (int, tuple))
+ if isinstance(topk, int):
+ topk = (topk, )
+ for _topk in topk:
+ assert _topk > 0, 'Top-k should be larger than 0'
+ self.topk = topk
+
+ if loss_cls is None:
+ if isinstance(num_classes, int):
+ warnings.warn('Since cross entropy is not set, '
+ 'the num_classes will be ignored.')
+ if loss_triplet is None:
+ raise ValueError('Please choose at least one loss in '
+ 'triplet loss and cross entropy loss.')
+ elif not isinstance(num_classes, int):
+ raise TypeError('The num_classes must be a current number, '
+ 'if there is cross entropy loss.')
+ self.loss_cls = MODELS.build(loss_cls) if loss_cls else None
+ self.loss_triplet = MODELS.build(loss_triplet) \
+ if loss_triplet else None
+
+ self.num_fcs = num_fcs
+ self.in_channels = in_channels
+ self.fc_channels = fc_channels
+ self.out_channels = out_channels
+ self.norm_cfg = norm_cfg
+ self.act_cfg = act_cfg
+ self.num_classes = num_classes
+
+ self._init_layers()
+
+ def _init_layers(self):
+ """Initialize fc layers."""
+ self.fcs = nn.ModuleList()
+ for i in range(self.num_fcs):
+ in_channels = self.in_channels if i == 0 else self.fc_channels
+ self.fcs.append(
+ FcModule(in_channels, self.fc_channels, self.norm_cfg,
+ self.act_cfg))
+ in_channels = self.in_channels if self.num_fcs == 0 else \
+ self.fc_channels
+ self.fc_out = nn.Linear(in_channels, self.out_channels)
+ if self.loss_cls:
+ self.bn = nn.BatchNorm1d(self.out_channels)
+ self.classifier = nn.Linear(self.out_channels, self.num_classes)
+
+ def forward(self, feats: Tuple[torch.Tensor]) -> torch.Tensor:
+ """The forward process."""
+ # Multiple stage inputs are acceptable
+ # but only the last stage will be used.
+ feats = feats[-1]
+
+ for m in self.fcs:
+ feats = m(feats)
+ feats = self.fc_out(feats)
+ return feats
+
+ def loss(self, feats: Tuple[torch.Tensor],
+ data_samples: List[ReIDDataSample]) -> dict:
+ """Calculate losses.
+
+ Args:
+ feats (tuple[Tensor]): The features extracted from the backbone.
+ data_samples (List[ReIDDataSample]): The annotation data of
+ every samples.
+
+ Returns:
+ dict: a dictionary of loss components
+ """
+ # The part can be traced by torch.fx
+ feats = self(feats)
+
+ # The part can not be traced by torch.fx
+ losses = self.loss_by_feat(feats, data_samples)
+ return losses
+
+ def loss_by_feat(self, feats: torch.Tensor,
+ data_samples: List[ReIDDataSample]) -> dict:
+ """Unpack data samples and compute loss."""
+ losses = dict()
+ gt_label = torch.cat([i.gt_label.label for i in data_samples])
+ gt_label = gt_label.to(feats.device)
+
+ if self.loss_triplet:
+ losses['triplet_loss'] = self.loss_triplet(feats, gt_label)
+
+ if self.loss_cls:
+ feats_bn = self.bn(feats)
+ cls_score = self.classifier(feats_bn)
+ losses['ce_loss'] = self.loss_cls(cls_score, gt_label)
+ acc = Accuracy.calculate(cls_score, gt_label, topk=self.topk)
+ losses.update(
+ {f'accuracy_top-{k}': a
+ for k, a in zip(self.topk, acc)})
+
+ return losses
+
+ def predict(
+ self,
+ feats: Tuple[torch.Tensor],
+ data_samples: List[ReIDDataSample] = None) -> List[ReIDDataSample]:
+ """Inference without augmentation.
+
+ Args:
+ feats (Tuple[Tensor]): The features extracted from the backbone.
+ Multiple stage inputs are acceptable but only the last stage
+ will be used.
+ data_samples (List[ReIDDataSample], optional): The annotation
+ data of every samples. If not None, set ``pred_label`` of
+ the input data samples. Defaults to None.
+
+ Returns:
+ List[ReIDDataSample]: A list of data samples which contains the
+ predicted results.
+ """
+ # The part can be traced by torch.fx
+ feats = self(feats)
+
+ # The part can not be traced by torch.fx
+ data_samples = self.predict_by_feat(feats, data_samples)
+
+ return data_samples
+
+ def predict_by_feat(
+ self,
+ feats: torch.Tensor,
+ data_samples: List[ReIDDataSample] = None) -> List[ReIDDataSample]:
+ """Add prediction features to data samples."""
+ if data_samples is not None:
+ for data_sample, feat in zip(data_samples, feats):
+ data_sample.pred_feature = feat
+ else:
+ data_samples = []
+ for feat in feats:
+ data_sample = ReIDDataSample()
+ data_sample.pred_feature = feat
+ data_samples.append(data_sample)
+
+ return data_samples
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..bba5664cc5ae5229ddebcb42f7583364ca9f77d8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/__init__.py
@@ -0,0 +1,38 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .base_roi_head import BaseRoIHead
+from .bbox_heads import (BBoxHead, ConvFCBBoxHead, DIIHead,
+ DoubleConvFCBBoxHead, SABLHead, SCNetBBoxHead,
+ Shared2FCBBoxHead, Shared4Conv1FCBBoxHead)
+from .cascade_roi_head import CascadeRoIHead
+from .double_roi_head import DoubleHeadRoIHead
+from .dynamic_roi_head import DynamicRoIHead
+from .grid_roi_head import GridRoIHead
+from .htc_roi_head import HybridTaskCascadeRoIHead
+from .mask_heads import (CoarseMaskHead, FCNMaskHead, FeatureRelayHead,
+ FusedSemanticHead, GlobalContextHead, GridHead,
+ HTCMaskHead, MaskIoUHead, MaskPointHead,
+ SCNetMaskHead, SCNetSemanticHead)
+from .mask_scoring_roi_head import MaskScoringRoIHead
+from .multi_instance_roi_head import MultiInstanceRoIHead
+from .pisa_roi_head import PISARoIHead
+from .point_rend_roi_head import PointRendRoIHead
+from .roi_extractors import (BaseRoIExtractor, GenericRoIExtractor,
+ SingleRoIExtractor)
+from .scnet_roi_head import SCNetRoIHead
+from .shared_heads import ResLayer
+from .sparse_roi_head import SparseRoIHead
+from .standard_roi_head import StandardRoIHead
+from .trident_roi_head import TridentRoIHead
+
+__all__ = [
+ 'BaseRoIHead', 'CascadeRoIHead', 'DoubleHeadRoIHead', 'MaskScoringRoIHead',
+ 'HybridTaskCascadeRoIHead', 'GridRoIHead', 'ResLayer', 'BBoxHead',
+ 'ConvFCBBoxHead', 'DIIHead', 'SABLHead', 'Shared2FCBBoxHead',
+ 'StandardRoIHead', 'Shared4Conv1FCBBoxHead', 'DoubleConvFCBBoxHead',
+ 'FCNMaskHead', 'HTCMaskHead', 'FusedSemanticHead', 'GridHead',
+ 'MaskIoUHead', 'BaseRoIExtractor', 'GenericRoIExtractor',
+ 'SingleRoIExtractor', 'PISARoIHead', 'PointRendRoIHead', 'MaskPointHead',
+ 'CoarseMaskHead', 'DynamicRoIHead', 'SparseRoIHead', 'TridentRoIHead',
+ 'SCNetRoIHead', 'SCNetMaskHead', 'SCNetSemanticHead', 'SCNetBBoxHead',
+ 'FeatureRelayHead', 'GlobalContextHead', 'MultiInstanceRoIHead'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/base_roi_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/base_roi_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..405f80a73ecc5db7343d81ca55518160fcbc2b63
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/base_roi_head.py
@@ -0,0 +1,129 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from abc import ABCMeta, abstractmethod
+from typing import Tuple
+
+from mmengine.model import BaseModule
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.utils import InstanceList, OptConfigType, OptMultiConfig
+
+
+class BaseRoIHead(BaseModule, metaclass=ABCMeta):
+ """Base class for RoIHeads."""
+
+ def __init__(self,
+ bbox_roi_extractor: OptMultiConfig = None,
+ bbox_head: OptMultiConfig = None,
+ mask_roi_extractor: OptMultiConfig = None,
+ mask_head: OptMultiConfig = None,
+ shared_head: OptConfigType = None,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+ if shared_head is not None:
+ self.shared_head = MODELS.build(shared_head)
+
+ if bbox_head is not None:
+ self.init_bbox_head(bbox_roi_extractor, bbox_head)
+
+ if mask_head is not None:
+ self.init_mask_head(mask_roi_extractor, mask_head)
+
+ self.init_assigner_sampler()
+
+ @property
+ def with_bbox(self) -> bool:
+ """bool: whether the RoI head contains a `bbox_head`"""
+ return hasattr(self, 'bbox_head') and self.bbox_head is not None
+
+ @property
+ def with_mask(self) -> bool:
+ """bool: whether the RoI head contains a `mask_head`"""
+ return hasattr(self, 'mask_head') and self.mask_head is not None
+
+ @property
+ def with_shared_head(self) -> bool:
+ """bool: whether the RoI head contains a `shared_head`"""
+ return hasattr(self, 'shared_head') and self.shared_head is not None
+
+ @abstractmethod
+ def init_bbox_head(self, *args, **kwargs):
+ """Initialize ``bbox_head``"""
+ pass
+
+ @abstractmethod
+ def init_mask_head(self, *args, **kwargs):
+ """Initialize ``mask_head``"""
+ pass
+
+ @abstractmethod
+ def init_assigner_sampler(self, *args, **kwargs):
+ """Initialize assigner and sampler."""
+ pass
+
+ @abstractmethod
+ def loss(self, x: Tuple[Tensor], rpn_results_list: InstanceList,
+ batch_data_samples: SampleList):
+ """Perform forward propagation and loss calculation of the roi head on
+ the features of the upstream network."""
+
+ def predict(self,
+ x: Tuple[Tensor],
+ rpn_results_list: InstanceList,
+ batch_data_samples: SampleList,
+ rescale: bool = False) -> InstanceList:
+ """Perform forward propagation of the roi head and predict detection
+ results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from upstream network. Each
+ has shape (N, C, H, W).
+ rpn_results_list (list[:obj:`InstanceData`]): list of region
+ proposals.
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool): Whether to rescale the results to
+ the original image. Defaults to True.
+
+ Returns:
+ list[obj:`InstanceData`]: Detection results of each image.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+ """
+ assert self.with_bbox, 'Bbox head must be implemented.'
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+
+ # TODO: nms_op in mmcv need be enhanced, the bbox result may get
+ # difference when not rescale in bbox_head
+
+ # If it has the mask branch, the bbox branch does not need
+ # to be scaled to the original image scale, because the mask
+ # branch will scale both bbox and mask at the same time.
+ bbox_rescale = rescale if not self.with_mask else False
+ results_list = self.predict_bbox(
+ x,
+ batch_img_metas,
+ rpn_results_list,
+ rcnn_test_cfg=self.test_cfg,
+ rescale=bbox_rescale)
+
+ if self.with_mask:
+ results_list = self.predict_mask(
+ x, batch_img_metas, results_list, rescale=rescale)
+
+ return results_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..d9e742abfecfc9dfe37b78822407fc92e9d64cc3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/__init__.py
@@ -0,0 +1,15 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .bbox_head import BBoxHead
+from .convfc_bbox_head import (ConvFCBBoxHead, Shared2FCBBoxHead,
+ Shared4Conv1FCBBoxHead)
+from .dii_head import DIIHead
+from .double_bbox_head import DoubleConvFCBBoxHead
+from .multi_instance_bbox_head import MultiInstanceBBoxHead
+from .sabl_head import SABLHead
+from .scnet_bbox_head import SCNetBBoxHead
+
+__all__ = [
+ 'BBoxHead', 'ConvFCBBoxHead', 'Shared2FCBBoxHead',
+ 'Shared4Conv1FCBBoxHead', 'DoubleConvFCBBoxHead', 'SABLHead', 'DIIHead',
+ 'SCNetBBoxHead', 'MultiInstanceBBoxHead'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/bbox_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/bbox_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..3b2e8aae0833ae0351b544099d79d296f082a76e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/bbox_head.py
@@ -0,0 +1,708 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple, Union
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmengine.config import ConfigDict
+from mmengine.model import BaseModule
+from mmengine.structures import InstanceData
+from torch import Tensor
+from torch.nn.modules.utils import _pair
+
+from mmdet.models.layers import multiclass_nms
+from mmdet.models.losses import accuracy
+from mmdet.models.task_modules.samplers import SamplingResult
+from mmdet.models.utils import empty_instances, multi_apply
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures.bbox import get_box_tensor, scale_boxes
+from mmdet.utils import ConfigType, InstanceList, OptMultiConfig
+
+
+@MODELS.register_module()
+class BBoxHead(BaseModule):
+ """Simplest RoI head, with only two fc layers for classification and
+ regression respectively."""
+
+ def __init__(self,
+ with_avg_pool: bool = False,
+ with_cls: bool = True,
+ with_reg: bool = True,
+ roi_feat_size: int = 7,
+ in_channels: int = 256,
+ num_classes: int = 80,
+ bbox_coder: ConfigType = dict(
+ type='DeltaXYWHBBoxCoder',
+ clip_border=True,
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ predict_box_type: str = 'hbox',
+ reg_class_agnostic: bool = False,
+ reg_decoded_bbox: bool = False,
+ reg_predictor_cfg: ConfigType = dict(type='Linear'),
+ cls_predictor_cfg: ConfigType = dict(type='Linear'),
+ loss_cls: ConfigType = dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox: ConfigType = dict(
+ type='SmoothL1Loss', beta=1.0, loss_weight=1.0),
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ assert with_cls or with_reg
+ self.with_avg_pool = with_avg_pool
+ self.with_cls = with_cls
+ self.with_reg = with_reg
+ self.roi_feat_size = _pair(roi_feat_size)
+ self.roi_feat_area = self.roi_feat_size[0] * self.roi_feat_size[1]
+ self.in_channels = in_channels
+ self.num_classes = num_classes
+ self.predict_box_type = predict_box_type
+ self.reg_class_agnostic = reg_class_agnostic
+ self.reg_decoded_bbox = reg_decoded_bbox
+ self.reg_predictor_cfg = reg_predictor_cfg
+ self.cls_predictor_cfg = cls_predictor_cfg
+
+ self.bbox_coder = TASK_UTILS.build(bbox_coder)
+ self.loss_cls = MODELS.build(loss_cls)
+ self.loss_bbox = MODELS.build(loss_bbox)
+
+ in_channels = self.in_channels
+ if self.with_avg_pool:
+ self.avg_pool = nn.AvgPool2d(self.roi_feat_size)
+ else:
+ in_channels *= self.roi_feat_area
+ if self.with_cls:
+ # need to add background class
+ if self.custom_cls_channels:
+ cls_channels = self.loss_cls.get_cls_channels(self.num_classes)
+ else:
+ cls_channels = num_classes + 1
+ cls_predictor_cfg_ = self.cls_predictor_cfg.copy()
+ cls_predictor_cfg_.update(
+ in_features=in_channels, out_features=cls_channels)
+ self.fc_cls = MODELS.build(cls_predictor_cfg_)
+ if self.with_reg:
+ box_dim = self.bbox_coder.encode_size
+ out_dim_reg = box_dim if reg_class_agnostic else \
+ box_dim * num_classes
+ reg_predictor_cfg_ = self.reg_predictor_cfg.copy()
+ if isinstance(reg_predictor_cfg_, (dict, ConfigDict)):
+ reg_predictor_cfg_.update(
+ in_features=in_channels, out_features=out_dim_reg)
+ self.fc_reg = MODELS.build(reg_predictor_cfg_)
+ self.debug_imgs = None
+ if init_cfg is None:
+ self.init_cfg = []
+ if self.with_cls:
+ self.init_cfg += [
+ dict(
+ type='Normal', std=0.01, override=dict(name='fc_cls'))
+ ]
+ if self.with_reg:
+ self.init_cfg += [
+ dict(
+ type='Normal', std=0.001, override=dict(name='fc_reg'))
+ ]
+
+ # TODO: Create a SeasawBBoxHead to simplified logic in BBoxHead
+ @property
+ def custom_cls_channels(self) -> bool:
+ """get custom_cls_channels from loss_cls."""
+ return getattr(self.loss_cls, 'custom_cls_channels', False)
+
+ # TODO: Create a SeasawBBoxHead to simplified logic in BBoxHead
+ @property
+ def custom_activation(self) -> bool:
+ """get custom_activation from loss_cls."""
+ return getattr(self.loss_cls, 'custom_activation', False)
+
+ # TODO: Create a SeasawBBoxHead to simplified logic in BBoxHead
+ @property
+ def custom_accuracy(self) -> bool:
+ """get custom_accuracy from loss_cls."""
+ return getattr(self.loss_cls, 'custom_accuracy', False)
+
+ def forward(self, x: Tuple[Tensor]) -> tuple:
+ """Forward features from the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: A tuple of classification scores and bbox prediction.
+
+ - cls_score (Tensor): Classification scores for all
+ scale levels, each is a 4D-tensor, the channels number
+ is num_base_priors * num_classes.
+ - bbox_pred (Tensor): Box energies / deltas for all
+ scale levels, each is a 4D-tensor, the channels number
+ is num_base_priors * 4.
+ """
+ if self.with_avg_pool:
+ if x.numel() > 0:
+ x = self.avg_pool(x)
+ x = x.view(x.size(0), -1)
+ else:
+ # avg_pool does not support empty tensor,
+ # so use torch.mean instead it
+ x = torch.mean(x, dim=(-1, -2))
+ cls_score = self.fc_cls(x) if self.with_cls else None
+ bbox_pred = self.fc_reg(x) if self.with_reg else None
+ return cls_score, bbox_pred
+
+ def _get_targets_single(self, pos_priors: Tensor, neg_priors: Tensor,
+ pos_gt_bboxes: Tensor, pos_gt_labels: Tensor,
+ cfg: ConfigDict) -> tuple:
+ """Calculate the ground truth for proposals in the single image
+ according to the sampling results.
+
+ Args:
+ pos_priors (Tensor): Contains all the positive boxes,
+ has shape (num_pos, 4), the last dimension 4
+ represents [tl_x, tl_y, br_x, br_y].
+ neg_priors (Tensor): Contains all the negative boxes,
+ has shape (num_neg, 4), the last dimension 4
+ represents [tl_x, tl_y, br_x, br_y].
+ pos_gt_bboxes (Tensor): Contains gt_boxes for
+ all positive samples, has shape (num_pos, 4),
+ the last dimension 4
+ represents [tl_x, tl_y, br_x, br_y].
+ pos_gt_labels (Tensor): Contains gt_labels for
+ all positive samples, has shape (num_pos, ).
+ cfg (obj:`ConfigDict`): `train_cfg` of R-CNN.
+
+ Returns:
+ Tuple[Tensor]: Ground truth for proposals
+ in a single image. Containing the following Tensors:
+
+ - labels(Tensor): Gt_labels for all proposals, has
+ shape (num_proposals,).
+ - label_weights(Tensor): Labels_weights for all
+ proposals, has shape (num_proposals,).
+ - bbox_targets(Tensor):Regression target for all
+ proposals, has shape (num_proposals, 4), the
+ last dimension 4 represents [tl_x, tl_y, br_x, br_y].
+ - bbox_weights(Tensor):Regression weights for all
+ proposals, has shape (num_proposals, 4).
+ """
+ num_pos = pos_priors.size(0)
+ num_neg = neg_priors.size(0)
+ num_samples = num_pos + num_neg
+
+ # original implementation uses new_zeros since BG are set to be 0
+ # now use empty & fill because BG cat_id = num_classes,
+ # FG cat_id = [0, num_classes-1]
+ labels = pos_priors.new_full((num_samples, ),
+ self.num_classes,
+ dtype=torch.long)
+ reg_dim = pos_gt_bboxes.size(-1) if self.reg_decoded_bbox \
+ else self.bbox_coder.encode_size
+ label_weights = pos_priors.new_zeros(num_samples)
+ bbox_targets = pos_priors.new_zeros(num_samples, reg_dim)
+ bbox_weights = pos_priors.new_zeros(num_samples, reg_dim)
+ if num_pos > 0:
+ labels[:num_pos] = pos_gt_labels
+ pos_weight = 1.0 if cfg.pos_weight <= 0 else cfg.pos_weight
+ label_weights[:num_pos] = pos_weight
+ if not self.reg_decoded_bbox:
+ pos_bbox_targets = self.bbox_coder.encode(
+ pos_priors, pos_gt_bboxes)
+ else:
+ # When the regression loss (e.g. `IouLoss`, `GIouLoss`)
+ # is applied directly on the decoded bounding boxes, both
+ # the predicted boxes and regression targets should be with
+ # absolute coordinate format.
+ pos_bbox_targets = get_box_tensor(pos_gt_bboxes)
+ bbox_targets[:num_pos, :] = pos_bbox_targets
+ bbox_weights[:num_pos, :] = 1
+ if num_neg > 0:
+ label_weights[-num_neg:] = 1.0
+
+ return labels, label_weights, bbox_targets, bbox_weights
+
+ def get_targets(self,
+ sampling_results: List[SamplingResult],
+ rcnn_train_cfg: ConfigDict,
+ concat: bool = True) -> tuple:
+ """Calculate the ground truth for all samples in a batch according to
+ the sampling_results.
+
+ Almost the same as the implementation in bbox_head, we passed
+ additional parameters pos_inds_list and neg_inds_list to
+ `_get_targets_single` function.
+
+ Args:
+ sampling_results (List[obj:SamplingResult]): Assign results of
+ all images in a batch after sampling.
+ rcnn_train_cfg (obj:ConfigDict): `train_cfg` of RCNN.
+ concat (bool): Whether to concatenate the results of all
+ the images in a single batch.
+
+ Returns:
+ Tuple[Tensor]: Ground truth for proposals in a single image.
+ Containing the following list of Tensors:
+
+ - labels (list[Tensor],Tensor): Gt_labels for all
+ proposals in a batch, each tensor in list has
+ shape (num_proposals,) when `concat=False`, otherwise
+ just a single tensor has shape (num_all_proposals,).
+ - label_weights (list[Tensor]): Labels_weights for
+ all proposals in a batch, each tensor in list has
+ shape (num_proposals,) when `concat=False`, otherwise
+ just a single tensor has shape (num_all_proposals,).
+ - bbox_targets (list[Tensor],Tensor): Regression target
+ for all proposals in a batch, each tensor in list
+ has shape (num_proposals, 4) when `concat=False`,
+ otherwise just a single tensor has shape
+ (num_all_proposals, 4), the last dimension 4 represents
+ [tl_x, tl_y, br_x, br_y].
+ - bbox_weights (list[tensor],Tensor): Regression weights for
+ all proposals in a batch, each tensor in list has shape
+ (num_proposals, 4) when `concat=False`, otherwise just a
+ single tensor has shape (num_all_proposals, 4).
+ """
+ pos_priors_list = [res.pos_priors for res in sampling_results]
+ neg_priors_list = [res.neg_priors for res in sampling_results]
+ pos_gt_bboxes_list = [res.pos_gt_bboxes for res in sampling_results]
+ pos_gt_labels_list = [res.pos_gt_labels for res in sampling_results]
+ labels, label_weights, bbox_targets, bbox_weights = multi_apply(
+ self._get_targets_single,
+ pos_priors_list,
+ neg_priors_list,
+ pos_gt_bboxes_list,
+ pos_gt_labels_list,
+ cfg=rcnn_train_cfg)
+
+ if concat:
+ labels = torch.cat(labels, 0)
+ label_weights = torch.cat(label_weights, 0)
+ bbox_targets = torch.cat(bbox_targets, 0)
+ bbox_weights = torch.cat(bbox_weights, 0)
+ return labels, label_weights, bbox_targets, bbox_weights
+
+ def loss_and_target(self,
+ cls_score: Tensor,
+ bbox_pred: Tensor,
+ rois: Tensor,
+ sampling_results: List[SamplingResult],
+ rcnn_train_cfg: ConfigDict,
+ concat: bool = True,
+ reduction_override: Optional[str] = None) -> dict:
+ """Calculate the loss based on the features extracted by the bbox head.
+
+ Args:
+ cls_score (Tensor): Classification prediction
+ results of all class, has shape
+ (batch_size * num_proposals_single_image, num_classes)
+ bbox_pred (Tensor): Regression prediction results,
+ has shape
+ (batch_size * num_proposals_single_image, 4), the last
+ dimension 4 represents [tl_x, tl_y, br_x, br_y].
+ rois (Tensor): RoIs with the shape
+ (batch_size * num_proposals_single_image, 5) where the first
+ column indicates batch id of each RoI.
+ sampling_results (List[obj:SamplingResult]): Assign results of
+ all images in a batch after sampling.
+ rcnn_train_cfg (obj:ConfigDict): `train_cfg` of RCNN.
+ concat (bool): Whether to concatenate the results of all
+ the images in a single batch. Defaults to True.
+ reduction_override (str, optional): The reduction
+ method used to override the original reduction
+ method of the loss. Options are "none",
+ "mean" and "sum". Defaults to None,
+
+ Returns:
+ dict: A dictionary of loss and targets components.
+ The targets are only used for cascade rcnn.
+ """
+
+ cls_reg_targets = self.get_targets(
+ sampling_results, rcnn_train_cfg, concat=concat)
+ losses = self.loss(
+ cls_score,
+ bbox_pred,
+ rois,
+ *cls_reg_targets,
+ reduction_override=reduction_override)
+
+ # cls_reg_targets is only for cascade rcnn
+ return dict(loss_bbox=losses, bbox_targets=cls_reg_targets)
+
+ def loss(self,
+ cls_score: Tensor,
+ bbox_pred: Tensor,
+ rois: Tensor,
+ labels: Tensor,
+ label_weights: Tensor,
+ bbox_targets: Tensor,
+ bbox_weights: Tensor,
+ reduction_override: Optional[str] = None) -> dict:
+ """Calculate the loss based on the network predictions and targets.
+
+ Args:
+ cls_score (Tensor): Classification prediction
+ results of all class, has shape
+ (batch_size * num_proposals_single_image, num_classes)
+ bbox_pred (Tensor): Regression prediction results,
+ has shape
+ (batch_size * num_proposals_single_image, 4), the last
+ dimension 4 represents [tl_x, tl_y, br_x, br_y].
+ rois (Tensor): RoIs with the shape
+ (batch_size * num_proposals_single_image, 5) where the first
+ column indicates batch id of each RoI.
+ labels (Tensor): Gt_labels for all proposals in a batch, has
+ shape (batch_size * num_proposals_single_image, ).
+ label_weights (Tensor): Labels_weights for all proposals in a
+ batch, has shape (batch_size * num_proposals_single_image, ).
+ bbox_targets (Tensor): Regression target for all proposals in a
+ batch, has shape (batch_size * num_proposals_single_image, 4),
+ the last dimension 4 represents [tl_x, tl_y, br_x, br_y].
+ bbox_weights (Tensor): Regression weights for all proposals in a
+ batch, has shape (batch_size * num_proposals_single_image, 4).
+ reduction_override (str, optional): The reduction
+ method used to override the original reduction
+ method of the loss. Options are "none",
+ "mean" and "sum". Defaults to None,
+
+ Returns:
+ dict: A dictionary of loss.
+ """
+
+ losses = dict()
+
+ if cls_score is not None:
+ avg_factor = max(torch.sum(label_weights > 0).float().item(), 1.)
+ if cls_score.numel() > 0:
+ loss_cls_ = self.loss_cls(
+ cls_score,
+ labels,
+ label_weights,
+ avg_factor=avg_factor,
+ reduction_override=reduction_override)
+ if isinstance(loss_cls_, dict):
+ losses.update(loss_cls_)
+ else:
+ losses['loss_cls'] = loss_cls_
+ if self.custom_activation:
+ acc_ = self.loss_cls.get_accuracy(cls_score, labels)
+ losses.update(acc_)
+ else:
+ losses['acc'] = accuracy(cls_score, labels)
+ if bbox_pred is not None:
+ bg_class_ind = self.num_classes
+ # 0~self.num_classes-1 are FG, self.num_classes is BG
+ pos_inds = (labels >= 0) & (labels < bg_class_ind)
+ # do not perform bounding box regression for BG anymore.
+ if pos_inds.any():
+ if self.reg_decoded_bbox:
+ # When the regression loss (e.g. `IouLoss`,
+ # `GIouLoss`, `DIouLoss`) is applied directly on
+ # the decoded bounding boxes, it decodes the
+ # already encoded coordinates to absolute format.
+ bbox_pred = self.bbox_coder.decode(rois[:, 1:], bbox_pred)
+ bbox_pred = get_box_tensor(bbox_pred)
+ if self.reg_class_agnostic:
+ pos_bbox_pred = bbox_pred.view(
+ bbox_pred.size(0), -1)[pos_inds.type(torch.bool)]
+ else:
+ pos_bbox_pred = bbox_pred.view(
+ bbox_pred.size(0), self.num_classes,
+ -1)[pos_inds.type(torch.bool),
+ labels[pos_inds.type(torch.bool)]]
+ losses['loss_bbox'] = self.loss_bbox(
+ pos_bbox_pred,
+ bbox_targets[pos_inds.type(torch.bool)],
+ bbox_weights[pos_inds.type(torch.bool)],
+ avg_factor=bbox_targets.size(0),
+ reduction_override=reduction_override)
+ else:
+ losses['loss_bbox'] = bbox_pred[pos_inds].sum()
+
+ return losses
+
+ def predict_by_feat(self,
+ rois: Tuple[Tensor],
+ cls_scores: Tuple[Tensor],
+ bbox_preds: Tuple[Tensor],
+ batch_img_metas: List[dict],
+ rcnn_test_cfg: Optional[ConfigDict] = None,
+ rescale: bool = False) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ bbox results.
+
+ Args:
+ rois (tuple[Tensor]): Tuple of boxes to be transformed.
+ Each has shape (num_boxes, 5). last dimension 5 arrange as
+ (batch_index, x1, y1, x2, y2).
+ cls_scores (tuple[Tensor]): Tuple of box scores, each has shape
+ (num_boxes, num_classes + 1).
+ bbox_preds (tuple[Tensor]): Tuple of box energies / deltas, each
+ has shape (num_boxes, num_classes * 4).
+ batch_img_metas (list[dict]): List of image information.
+ rcnn_test_cfg (obj:`ConfigDict`, optional): `test_cfg` of R-CNN.
+ Defaults to None.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Instance segmentation
+ results of each image after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ assert len(cls_scores) == len(bbox_preds)
+ result_list = []
+ for img_id in range(len(batch_img_metas)):
+ img_meta = batch_img_metas[img_id]
+ results = self._predict_by_feat_single(
+ roi=rois[img_id],
+ cls_score=cls_scores[img_id],
+ bbox_pred=bbox_preds[img_id],
+ img_meta=img_meta,
+ rescale=rescale,
+ rcnn_test_cfg=rcnn_test_cfg)
+ result_list.append(results)
+
+ return result_list
+
+ def _predict_by_feat_single(
+ self,
+ roi: Tensor,
+ cls_score: Tensor,
+ bbox_pred: Tensor,
+ img_meta: dict,
+ rescale: bool = False,
+ rcnn_test_cfg: Optional[ConfigDict] = None) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results.
+
+ Args:
+ roi (Tensor): Boxes to be transformed. Has shape (num_boxes, 5).
+ last dimension 5 arrange as (batch_index, x1, y1, x2, y2).
+ cls_score (Tensor): Box scores, has shape
+ (num_boxes, num_classes + 1).
+ bbox_pred (Tensor): Box energies / deltas.
+ has shape (num_boxes, num_classes * 4).
+ img_meta (dict): image information.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ rcnn_test_cfg (obj:`ConfigDict`): `test_cfg` of Bbox Head.
+ Defaults to None
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image\
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ results = InstanceData()
+ if roi.shape[0] == 0:
+ return empty_instances([img_meta],
+ roi.device,
+ task_type='bbox',
+ instance_results=[results],
+ box_type=self.predict_box_type,
+ use_box_type=False,
+ num_classes=self.num_classes,
+ score_per_cls=rcnn_test_cfg is None)[0]
+
+ # some loss (Seesaw loss..) may have custom activation
+ if self.custom_cls_channels:
+ scores = self.loss_cls.get_activation(cls_score)
+ else:
+ scores = F.softmax(
+ cls_score, dim=-1) if cls_score is not None else None
+
+ img_shape = img_meta['img_shape']
+ num_rois = roi.size(0)
+ # bbox_pred would be None in some detector when with_reg is False,
+ # e.g. Grid R-CNN.
+ if bbox_pred is not None:
+ num_classes = 1 if self.reg_class_agnostic else self.num_classes
+ roi = roi.repeat_interleave(num_classes, dim=0)
+ bbox_pred = bbox_pred.view(-1, self.bbox_coder.encode_size)
+ bboxes = self.bbox_coder.decode(
+ roi[..., 1:], bbox_pred, max_shape=img_shape)
+ else:
+ bboxes = roi[:, 1:].clone()
+ if img_shape is not None and bboxes.size(-1) == 4:
+ bboxes[:, [0, 2]].clamp_(min=0, max=img_shape[1])
+ bboxes[:, [1, 3]].clamp_(min=0, max=img_shape[0])
+
+ if rescale and bboxes.size(0) > 0:
+ assert img_meta.get('scale_factor') is not None
+ scale_factor = [1 / s for s in img_meta['scale_factor']]
+ bboxes = scale_boxes(bboxes, scale_factor)
+
+ # Get the inside tensor when `bboxes` is a box type
+ bboxes = get_box_tensor(bboxes)
+ box_dim = bboxes.size(-1)
+ bboxes = bboxes.view(num_rois, -1)
+
+ if rcnn_test_cfg is None:
+ # This means that it is aug test.
+ # It needs to return the raw results without nms.
+ results.bboxes = bboxes
+ results.scores = scores
+ else:
+ det_bboxes, det_labels = multiclass_nms(
+ bboxes,
+ scores,
+ rcnn_test_cfg.score_thr,
+ rcnn_test_cfg.nms,
+ rcnn_test_cfg.max_per_img,
+ box_dim=box_dim)
+ results.bboxes = det_bboxes[:, :-1]
+ results.scores = det_bboxes[:, -1]
+ results.labels = det_labels
+ return results
+
+ def refine_bboxes(self, sampling_results: Union[List[SamplingResult],
+ InstanceList],
+ bbox_results: dict,
+ batch_img_metas: List[dict]) -> InstanceList:
+ """Refine bboxes during training.
+
+ Args:
+ sampling_results (List[:obj:`SamplingResult`] or
+ List[:obj:`InstanceData`]): Sampling results.
+ :obj:`SamplingResult` is the real sampling results
+ calculate from bbox_head, while :obj:`InstanceData` is
+ fake sampling results, e.g., in Sparse R-CNN or QueryInst, etc.
+ bbox_results (dict): Usually is a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `rois` (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+ - `bbox_targets` (tuple): Ground truth for proposals in a
+ single image. Containing the following list of Tensors:
+ (labels, label_weights, bbox_targets, bbox_weights)
+ batch_img_metas (List[dict]): List of image information.
+
+ Returns:
+ list[:obj:`InstanceData`]: Refined bboxes of each image.
+
+ Example:
+ >>> # xdoctest: +REQUIRES(module:kwarray)
+ >>> import numpy as np
+ >>> from mmdet.models.task_modules.samplers.
+ ... sampling_result import random_boxes
+ >>> from mmdet.models.task_modules.samplers import SamplingResult
+ >>> self = BBoxHead(reg_class_agnostic=True)
+ >>> n_roi = 2
+ >>> n_img = 4
+ >>> scale = 512
+ >>> rng = np.random.RandomState(0)
+ ... batch_img_metas = [{'img_shape': (scale, scale)}
+ >>> for _ in range(n_img)]
+ >>> sampling_results = [SamplingResult.random(rng=10)
+ ... for _ in range(n_img)]
+ >>> # Create rois in the expected format
+ >>> roi_boxes = random_boxes(n_roi, scale=scale, rng=rng)
+ >>> img_ids = torch.randint(0, n_img, (n_roi,))
+ >>> img_ids = img_ids.float()
+ >>> rois = torch.cat([img_ids[:, None], roi_boxes], dim=1)
+ >>> # Create other args
+ >>> labels = torch.randint(0, 81, (scale,)).long()
+ >>> bbox_preds = random_boxes(n_roi, scale=scale, rng=rng)
+ >>> cls_score = torch.randn((scale, 81))
+ ... # For each image, pretend random positive boxes are gts
+ >>> bbox_targets = (labels, None, None, None)
+ ... bbox_results = dict(rois=rois, bbox_pred=bbox_preds,
+ ... cls_score=cls_score,
+ ... bbox_targets=bbox_targets)
+ >>> bboxes_list = self.refine_bboxes(sampling_results,
+ ... bbox_results,
+ ... batch_img_metas)
+ >>> print(bboxes_list)
+ """
+ pos_is_gts = [res.pos_is_gt for res in sampling_results]
+ # bbox_targets is a tuple
+ labels = bbox_results['bbox_targets'][0]
+ cls_scores = bbox_results['cls_score']
+ rois = bbox_results['rois']
+ bbox_preds = bbox_results['bbox_pred']
+ if self.custom_activation:
+ # TODO: Create a SeasawBBoxHead to simplified logic in BBoxHead
+ cls_scores = self.loss_cls.get_activation(cls_scores)
+ if cls_scores.numel() == 0:
+ return None
+ if cls_scores.shape[-1] == self.num_classes + 1:
+ # remove background class
+ cls_scores = cls_scores[:, :-1]
+ elif cls_scores.shape[-1] != self.num_classes:
+ raise ValueError('The last dim of `cls_scores` should equal to '
+ '`num_classes` or `num_classes + 1`,'
+ f'but got {cls_scores.shape[-1]}.')
+ labels = torch.where(labels == self.num_classes, cls_scores.argmax(1),
+ labels)
+
+ img_ids = rois[:, 0].long().unique(sorted=True)
+ assert img_ids.numel() <= len(batch_img_metas)
+
+ results_list = []
+ for i in range(len(batch_img_metas)):
+ inds = torch.nonzero(
+ rois[:, 0] == i, as_tuple=False).squeeze(dim=1)
+ num_rois = inds.numel()
+
+ bboxes_ = rois[inds, 1:]
+ label_ = labels[inds]
+ bbox_pred_ = bbox_preds[inds]
+ img_meta_ = batch_img_metas[i]
+ pos_is_gts_ = pos_is_gts[i]
+
+ bboxes = self.regress_by_class(bboxes_, label_, bbox_pred_,
+ img_meta_)
+ # filter gt bboxes
+ pos_keep = 1 - pos_is_gts_
+ keep_inds = pos_is_gts_.new_ones(num_rois)
+ keep_inds[:len(pos_is_gts_)] = pos_keep
+ results = InstanceData(bboxes=bboxes[keep_inds.type(torch.bool)])
+ results_list.append(results)
+
+ return results_list
+
+ def regress_by_class(self, priors: Tensor, label: Tensor,
+ bbox_pred: Tensor, img_meta: dict) -> Tensor:
+ """Regress the bbox for the predicted class. Used in Cascade R-CNN.
+
+ Args:
+ priors (Tensor): Priors from `rpn_head` or last stage
+ `bbox_head`, has shape (num_proposals, 4).
+ label (Tensor): Only used when `self.reg_class_agnostic`
+ is False, has shape (num_proposals, ).
+ bbox_pred (Tensor): Regression prediction of
+ current stage `bbox_head`. When `self.reg_class_agnostic`
+ is False, it has shape (n, num_classes * 4), otherwise
+ it has shape (n, 4).
+ img_meta (dict): Image meta info.
+
+ Returns:
+ Tensor: Regressed bboxes, the same shape as input rois.
+ """
+ reg_dim = self.bbox_coder.encode_size
+ if not self.reg_class_agnostic:
+ label = label * reg_dim
+ inds = torch.stack([label + i for i in range(reg_dim)], 1)
+ bbox_pred = torch.gather(bbox_pred, 1, inds)
+ assert bbox_pred.size()[1] == reg_dim
+
+ max_shape = img_meta['img_shape']
+ regressed_bboxes = self.bbox_coder.decode(
+ priors, bbox_pred, max_shape=max_shape)
+ return regressed_bboxes
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/convfc_bbox_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/convfc_bbox_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..cb6aadd86d34af3605d432492931442026432cc8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/convfc_bbox_head.py
@@ -0,0 +1,249 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Tuple, Union
+
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from mmengine.config import ConfigDict
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from .bbox_head import BBoxHead
+
+
+@MODELS.register_module()
+class ConvFCBBoxHead(BBoxHead):
+ r"""More general bbox head, with shared conv and fc layers and two optional
+ separated branches.
+
+ .. code-block:: none
+
+ /-> cls convs -> cls fcs -> cls
+ shared convs -> shared fcs
+ \-> reg convs -> reg fcs -> reg
+ """ # noqa: W605
+
+ def __init__(self,
+ num_shared_convs: int = 0,
+ num_shared_fcs: int = 0,
+ num_cls_convs: int = 0,
+ num_cls_fcs: int = 0,
+ num_reg_convs: int = 0,
+ num_reg_fcs: int = 0,
+ conv_out_channels: int = 256,
+ fc_out_channels: int = 1024,
+ conv_cfg: Optional[Union[dict, ConfigDict]] = None,
+ norm_cfg: Optional[Union[dict, ConfigDict]] = None,
+ init_cfg: Optional[Union[dict, ConfigDict]] = None,
+ *args,
+ **kwargs) -> None:
+ super().__init__(*args, init_cfg=init_cfg, **kwargs)
+ assert (num_shared_convs + num_shared_fcs + num_cls_convs +
+ num_cls_fcs + num_reg_convs + num_reg_fcs > 0)
+ if num_cls_convs > 0 or num_reg_convs > 0:
+ assert num_shared_fcs == 0
+ if not self.with_cls:
+ assert num_cls_convs == 0 and num_cls_fcs == 0
+ if not self.with_reg:
+ assert num_reg_convs == 0 and num_reg_fcs == 0
+ self.num_shared_convs = num_shared_convs
+ self.num_shared_fcs = num_shared_fcs
+ self.num_cls_convs = num_cls_convs
+ self.num_cls_fcs = num_cls_fcs
+ self.num_reg_convs = num_reg_convs
+ self.num_reg_fcs = num_reg_fcs
+ self.conv_out_channels = conv_out_channels
+ self.fc_out_channels = fc_out_channels
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+
+ # add shared convs and fcs
+ self.shared_convs, self.shared_fcs, last_layer_dim = \
+ self._add_conv_fc_branch(
+ self.num_shared_convs, self.num_shared_fcs, self.in_channels,
+ True)
+ self.shared_out_channels = last_layer_dim
+
+ # add cls specific branch
+ self.cls_convs, self.cls_fcs, self.cls_last_dim = \
+ self._add_conv_fc_branch(
+ self.num_cls_convs, self.num_cls_fcs, self.shared_out_channels)
+
+ # add reg specific branch
+ self.reg_convs, self.reg_fcs, self.reg_last_dim = \
+ self._add_conv_fc_branch(
+ self.num_reg_convs, self.num_reg_fcs, self.shared_out_channels)
+
+ if self.num_shared_fcs == 0 and not self.with_avg_pool:
+ if self.num_cls_fcs == 0:
+ self.cls_last_dim *= self.roi_feat_area
+ if self.num_reg_fcs == 0:
+ self.reg_last_dim *= self.roi_feat_area
+
+ self.relu = nn.ReLU(inplace=True)
+ # reconstruct fc_cls and fc_reg since input channels are changed
+ if self.with_cls:
+ if self.custom_cls_channels:
+ cls_channels = self.loss_cls.get_cls_channels(self.num_classes)
+ else:
+ cls_channels = self.num_classes + 1
+ cls_predictor_cfg_ = self.cls_predictor_cfg.copy()
+ cls_predictor_cfg_.update(
+ in_features=self.cls_last_dim, out_features=cls_channels)
+ self.fc_cls = MODELS.build(cls_predictor_cfg_)
+ if self.with_reg:
+ box_dim = self.bbox_coder.encode_size
+ out_dim_reg = box_dim if self.reg_class_agnostic else \
+ box_dim * self.num_classes
+ reg_predictor_cfg_ = self.reg_predictor_cfg.copy()
+ if isinstance(reg_predictor_cfg_, (dict, ConfigDict)):
+ reg_predictor_cfg_.update(
+ in_features=self.reg_last_dim, out_features=out_dim_reg)
+ self.fc_reg = MODELS.build(reg_predictor_cfg_)
+
+ if init_cfg is None:
+ # when init_cfg is None,
+ # It has been set to
+ # [[dict(type='Normal', std=0.01, override=dict(name='fc_cls'))],
+ # [dict(type='Normal', std=0.001, override=dict(name='fc_reg'))]
+ # after `super(ConvFCBBoxHead, self).__init__()`
+ # we only need to append additional configuration
+ # for `shared_fcs`, `cls_fcs` and `reg_fcs`
+ self.init_cfg += [
+ dict(
+ type='Xavier',
+ distribution='uniform',
+ override=[
+ dict(name='shared_fcs'),
+ dict(name='cls_fcs'),
+ dict(name='reg_fcs')
+ ])
+ ]
+
+ def _add_conv_fc_branch(self,
+ num_branch_convs: int,
+ num_branch_fcs: int,
+ in_channels: int,
+ is_shared: bool = False) -> tuple:
+ """Add shared or separable branch.
+
+ convs -> avg pool (optional) -> fcs
+ """
+ last_layer_dim = in_channels
+ # add branch specific conv layers
+ branch_convs = nn.ModuleList()
+ if num_branch_convs > 0:
+ for i in range(num_branch_convs):
+ conv_in_channels = (
+ last_layer_dim if i == 0 else self.conv_out_channels)
+ branch_convs.append(
+ ConvModule(
+ conv_in_channels,
+ self.conv_out_channels,
+ 3,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ last_layer_dim = self.conv_out_channels
+ # add branch specific fc layers
+ branch_fcs = nn.ModuleList()
+ if num_branch_fcs > 0:
+ # for shared branch, only consider self.with_avg_pool
+ # for separated branches, also consider self.num_shared_fcs
+ if (is_shared
+ or self.num_shared_fcs == 0) and not self.with_avg_pool:
+ last_layer_dim *= self.roi_feat_area
+ for i in range(num_branch_fcs):
+ fc_in_channels = (
+ last_layer_dim if i == 0 else self.fc_out_channels)
+ branch_fcs.append(
+ nn.Linear(fc_in_channels, self.fc_out_channels))
+ last_layer_dim = self.fc_out_channels
+ return branch_convs, branch_fcs, last_layer_dim
+
+ def forward(self, x: Tuple[Tensor]) -> tuple:
+ """Forward features from the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: A tuple of classification scores and bbox prediction.
+
+ - cls_score (Tensor): Classification scores for all \
+ scale levels, each is a 4D-tensor, the channels number \
+ is num_base_priors * num_classes.
+ - bbox_pred (Tensor): Box energies / deltas for all \
+ scale levels, each is a 4D-tensor, the channels number \
+ is num_base_priors * 4.
+ """
+ # shared part
+ if self.num_shared_convs > 0:
+ for conv in self.shared_convs:
+ x = conv(x)
+
+ if self.num_shared_fcs > 0:
+ if self.with_avg_pool:
+ x = self.avg_pool(x)
+
+ x = x.flatten(1)
+
+ for fc in self.shared_fcs:
+ x = self.relu(fc(x))
+ # separate branches
+ x_cls = x
+ x_reg = x
+
+ for conv in self.cls_convs:
+ x_cls = conv(x_cls)
+ if x_cls.dim() > 2:
+ if self.with_avg_pool:
+ x_cls = self.avg_pool(x_cls)
+ x_cls = x_cls.flatten(1)
+ for fc in self.cls_fcs:
+ x_cls = self.relu(fc(x_cls))
+
+ for conv in self.reg_convs:
+ x_reg = conv(x_reg)
+ if x_reg.dim() > 2:
+ if self.with_avg_pool:
+ x_reg = self.avg_pool(x_reg)
+ x_reg = x_reg.flatten(1)
+ for fc in self.reg_fcs:
+ x_reg = self.relu(fc(x_reg))
+
+ cls_score = self.fc_cls(x_cls) if self.with_cls else None
+ bbox_pred = self.fc_reg(x_reg) if self.with_reg else None
+ return cls_score, bbox_pred
+
+
+@MODELS.register_module()
+class Shared2FCBBoxHead(ConvFCBBoxHead):
+
+ def __init__(self, fc_out_channels: int = 1024, *args, **kwargs) -> None:
+ super().__init__(
+ num_shared_convs=0,
+ num_shared_fcs=2,
+ num_cls_convs=0,
+ num_cls_fcs=0,
+ num_reg_convs=0,
+ num_reg_fcs=0,
+ fc_out_channels=fc_out_channels,
+ *args,
+ **kwargs)
+
+
+@MODELS.register_module()
+class Shared4Conv1FCBBoxHead(ConvFCBBoxHead):
+
+ def __init__(self, fc_out_channels: int = 1024, *args, **kwargs) -> None:
+ super().__init__(
+ num_shared_convs=4,
+ num_shared_fcs=1,
+ num_cls_convs=0,
+ num_cls_fcs=0,
+ num_reg_convs=0,
+ num_reg_fcs=0,
+ fc_out_channels=fc_out_channels,
+ *args,
+ **kwargs)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/dii_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/dii_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..ae9a31bbeb2a8f1da62b457363fa05031d21925a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/dii_head.py
@@ -0,0 +1,422 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import build_activation_layer, build_norm_layer
+from mmcv.cnn.bricks.transformer import FFN, MultiheadAttention
+from mmengine.config import ConfigDict
+from mmengine.model import bias_init_with_prob
+from torch import Tensor
+
+from mmdet.models.losses import accuracy
+from mmdet.models.task_modules import SamplingResult
+from mmdet.models.utils import multi_apply
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptConfigType, reduce_mean
+from .bbox_head import BBoxHead
+
+
+@MODELS.register_module()
+class DIIHead(BBoxHead):
+ r"""Dynamic Instance Interactive Head for `Sparse R-CNN: End-to-End Object
+ Detection with Learnable Proposals `_
+
+ Args:
+ num_classes (int): Number of class in dataset.
+ Defaults to 80.
+ num_ffn_fcs (int): The number of fully-connected
+ layers in FFNs. Defaults to 2.
+ num_heads (int): The hidden dimension of FFNs.
+ Defaults to 8.
+ num_cls_fcs (int): The number of fully-connected
+ layers in classification subnet. Defaults to 1.
+ num_reg_fcs (int): The number of fully-connected
+ layers in regression subnet. Defaults to 3.
+ feedforward_channels (int): The hidden dimension
+ of FFNs. Defaults to 2048
+ in_channels (int): Hidden_channels of MultiheadAttention.
+ Defaults to 256.
+ dropout (float): Probability of drop the channel.
+ Defaults to 0.0
+ ffn_act_cfg (:obj:`ConfigDict` or dict): The activation config
+ for FFNs.
+ dynamic_conv_cfg (:obj:`ConfigDict` or dict): The convolution
+ config for DynamicConv.
+ loss_iou (:obj:`ConfigDict` or dict): The config for iou or
+ giou loss.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict]): Initialization config dict. Defaults to None.
+ """
+
+ def __init__(self,
+ num_classes: int = 80,
+ num_ffn_fcs: int = 2,
+ num_heads: int = 8,
+ num_cls_fcs: int = 1,
+ num_reg_fcs: int = 3,
+ feedforward_channels: int = 2048,
+ in_channels: int = 256,
+ dropout: float = 0.0,
+ ffn_act_cfg: ConfigType = dict(type='ReLU', inplace=True),
+ dynamic_conv_cfg: ConfigType = dict(
+ type='DynamicConv',
+ in_channels=256,
+ feat_channels=64,
+ out_channels=256,
+ input_feat_shape=7,
+ act_cfg=dict(type='ReLU', inplace=True),
+ norm_cfg=dict(type='LN')),
+ loss_iou: ConfigType = dict(type='GIoULoss', loss_weight=2.0),
+ init_cfg: OptConfigType = None,
+ **kwargs) -> None:
+ assert init_cfg is None, 'To prevent abnormal initialization ' \
+ 'behavior, init_cfg is not allowed to be set'
+ super().__init__(
+ num_classes=num_classes,
+ reg_decoded_bbox=True,
+ reg_class_agnostic=True,
+ init_cfg=init_cfg,
+ **kwargs)
+ self.loss_iou = MODELS.build(loss_iou)
+ self.in_channels = in_channels
+ self.fp16_enabled = False
+ self.attention = MultiheadAttention(in_channels, num_heads, dropout)
+ self.attention_norm = build_norm_layer(dict(type='LN'), in_channels)[1]
+
+ self.instance_interactive_conv = MODELS.build(dynamic_conv_cfg)
+ self.instance_interactive_conv_dropout = nn.Dropout(dropout)
+ self.instance_interactive_conv_norm = build_norm_layer(
+ dict(type='LN'), in_channels)[1]
+
+ self.ffn = FFN(
+ in_channels,
+ feedforward_channels,
+ num_ffn_fcs,
+ act_cfg=ffn_act_cfg,
+ dropout=dropout)
+ self.ffn_norm = build_norm_layer(dict(type='LN'), in_channels)[1]
+
+ self.cls_fcs = nn.ModuleList()
+ for _ in range(num_cls_fcs):
+ self.cls_fcs.append(
+ nn.Linear(in_channels, in_channels, bias=False))
+ self.cls_fcs.append(
+ build_norm_layer(dict(type='LN'), in_channels)[1])
+ self.cls_fcs.append(
+ build_activation_layer(dict(type='ReLU', inplace=True)))
+
+ # over load the self.fc_cls in BBoxHead
+ if self.loss_cls.use_sigmoid:
+ self.fc_cls = nn.Linear(in_channels, self.num_classes)
+ else:
+ self.fc_cls = nn.Linear(in_channels, self.num_classes + 1)
+
+ self.reg_fcs = nn.ModuleList()
+ for _ in range(num_reg_fcs):
+ self.reg_fcs.append(
+ nn.Linear(in_channels, in_channels, bias=False))
+ self.reg_fcs.append(
+ build_norm_layer(dict(type='LN'), in_channels)[1])
+ self.reg_fcs.append(
+ build_activation_layer(dict(type='ReLU', inplace=True)))
+ # over load the self.fc_cls in BBoxHead
+ self.fc_reg = nn.Linear(in_channels, 4)
+
+ assert self.reg_class_agnostic, 'DIIHead only ' \
+ 'suppport `reg_class_agnostic=True` '
+ assert self.reg_decoded_bbox, 'DIIHead only ' \
+ 'suppport `reg_decoded_bbox=True`'
+
+ def init_weights(self) -> None:
+ """Use xavier initialization for all weight parameter and set
+ classification head bias as a specific value when use focal loss."""
+ super().init_weights()
+ for p in self.parameters():
+ if p.dim() > 1:
+ nn.init.xavier_uniform_(p)
+ else:
+ # adopt the default initialization for
+ # the weight and bias of the layer norm
+ pass
+ if self.loss_cls.use_sigmoid:
+ bias_init = bias_init_with_prob(0.01)
+ nn.init.constant_(self.fc_cls.bias, bias_init)
+
+ def forward(self, roi_feat: Tensor, proposal_feat: Tensor) -> tuple:
+ """Forward function of Dynamic Instance Interactive Head.
+
+ Args:
+ roi_feat (Tensor): Roi-pooling features with shape
+ (batch_size*num_proposals, feature_dimensions,
+ pooling_h , pooling_w).
+ proposal_feat (Tensor): Intermediate feature get from
+ diihead in last stage, has shape
+ (batch_size, num_proposals, feature_dimensions)
+
+ Returns:
+ tuple[Tensor]: Usually a tuple of classification scores
+ and bbox prediction and a intermediate feature.
+
+ - cls_scores (Tensor): Classification scores for
+ all proposals, has shape
+ (batch_size, num_proposals, num_classes).
+ - bbox_preds (Tensor): Box energies / deltas for
+ all proposals, has shape
+ (batch_size, num_proposals, 4).
+ - obj_feat (Tensor): Object feature before classification
+ and regression subnet, has shape
+ (batch_size, num_proposal, feature_dimensions).
+ - attn_feats (Tensor): Intermediate feature.
+ """
+ N, num_proposals = proposal_feat.shape[:2]
+
+ # Self attention
+ proposal_feat = proposal_feat.permute(1, 0, 2)
+ proposal_feat = self.attention_norm(self.attention(proposal_feat))
+ attn_feats = proposal_feat.permute(1, 0, 2)
+
+ # instance interactive
+ proposal_feat = attn_feats.reshape(-1, self.in_channels)
+ proposal_feat_iic = self.instance_interactive_conv(
+ proposal_feat, roi_feat)
+ proposal_feat = proposal_feat + self.instance_interactive_conv_dropout(
+ proposal_feat_iic)
+ obj_feat = self.instance_interactive_conv_norm(proposal_feat)
+
+ # FFN
+ obj_feat = self.ffn_norm(self.ffn(obj_feat))
+
+ cls_feat = obj_feat
+ reg_feat = obj_feat
+
+ for cls_layer in self.cls_fcs:
+ cls_feat = cls_layer(cls_feat)
+ for reg_layer in self.reg_fcs:
+ reg_feat = reg_layer(reg_feat)
+
+ cls_score = self.fc_cls(cls_feat).view(
+ N, num_proposals, self.num_classes
+ if self.loss_cls.use_sigmoid else self.num_classes + 1)
+ bbox_delta = self.fc_reg(reg_feat).view(N, num_proposals, 4)
+
+ return cls_score, bbox_delta, obj_feat.view(
+ N, num_proposals, self.in_channels), attn_feats
+
+ def loss_and_target(self,
+ cls_score: Tensor,
+ bbox_pred: Tensor,
+ sampling_results: List[SamplingResult],
+ rcnn_train_cfg: ConfigType,
+ imgs_whwh: Tensor,
+ concat: bool = True,
+ reduction_override: str = None) -> dict:
+ """Calculate the loss based on the features extracted by the DIIHead.
+
+ Args:
+ cls_score (Tensor): Classification prediction
+ results of all class, has shape
+ (batch_size * num_proposals_single_image, num_classes)
+ bbox_pred (Tensor): Regression prediction results, has shape
+ (batch_size * num_proposals_single_image, 4), the last
+ dimension 4 represents [tl_x, tl_y, br_x, br_y].
+ sampling_results (List[obj:SamplingResult]): Assign results of
+ all images in a batch after sampling.
+ rcnn_train_cfg (obj:ConfigDict): `train_cfg` of RCNN.
+ imgs_whwh (Tensor): imgs_whwh (Tensor): Tensor with\
+ shape (batch_size, num_proposals, 4), the last
+ dimension means
+ [img_width,img_height, img_width, img_height].
+ concat (bool): Whether to concatenate the results of all
+ the images in a single batch. Defaults to True.
+ reduction_override (str, optional): The reduction
+ method used to override the original reduction
+ method of the loss. Options are "none",
+ "mean" and "sum". Defaults to None.
+
+ Returns:
+ dict: A dictionary of loss and targets components.
+ The targets are only used for cascade rcnn.
+ """
+ cls_reg_targets = self.get_targets(
+ sampling_results=sampling_results,
+ rcnn_train_cfg=rcnn_train_cfg,
+ concat=concat)
+ (labels, label_weights, bbox_targets, bbox_weights) = cls_reg_targets
+
+ losses = dict()
+ bg_class_ind = self.num_classes
+ # note in spare rcnn num_gt == num_pos
+ pos_inds = (labels >= 0) & (labels < bg_class_ind)
+ num_pos = pos_inds.sum().float()
+ avg_factor = reduce_mean(num_pos)
+ if cls_score is not None:
+ if cls_score.numel() > 0:
+ losses['loss_cls'] = self.loss_cls(
+ cls_score,
+ labels,
+ label_weights,
+ avg_factor=avg_factor,
+ reduction_override=reduction_override)
+ losses['pos_acc'] = accuracy(cls_score[pos_inds],
+ labels[pos_inds])
+ if bbox_pred is not None:
+ # 0~self.num_classes-1 are FG, self.num_classes is BG
+ # do not perform bounding box regression for BG anymore.
+ if pos_inds.any():
+ pos_bbox_pred = bbox_pred.reshape(bbox_pred.size(0),
+ 4)[pos_inds.type(torch.bool)]
+ imgs_whwh = imgs_whwh.reshape(bbox_pred.size(0),
+ 4)[pos_inds.type(torch.bool)]
+ losses['loss_bbox'] = self.loss_bbox(
+ pos_bbox_pred / imgs_whwh,
+ bbox_targets[pos_inds.type(torch.bool)] / imgs_whwh,
+ bbox_weights[pos_inds.type(torch.bool)],
+ avg_factor=avg_factor)
+ losses['loss_iou'] = self.loss_iou(
+ pos_bbox_pred,
+ bbox_targets[pos_inds.type(torch.bool)],
+ bbox_weights[pos_inds.type(torch.bool)],
+ avg_factor=avg_factor)
+ else:
+ losses['loss_bbox'] = bbox_pred.sum() * 0
+ losses['loss_iou'] = bbox_pred.sum() * 0
+ return dict(loss_bbox=losses, bbox_targets=cls_reg_targets)
+
+ def _get_targets_single(self, pos_inds: Tensor, neg_inds: Tensor,
+ pos_priors: Tensor, neg_priors: Tensor,
+ pos_gt_bboxes: Tensor, pos_gt_labels: Tensor,
+ cfg: ConfigDict) -> tuple:
+ """Calculate the ground truth for proposals in the single image
+ according to the sampling results.
+
+ Almost the same as the implementation in `bbox_head`,
+ we add pos_inds and neg_inds to select positive and
+ negative samples instead of selecting the first num_pos
+ as positive samples.
+
+ Args:
+ pos_inds (Tensor): The length is equal to the
+ positive sample numbers contain all index
+ of the positive sample in the origin proposal set.
+ neg_inds (Tensor): The length is equal to the
+ negative sample numbers contain all index
+ of the negative sample in the origin proposal set.
+ pos_priors (Tensor): Contains all the positive boxes,
+ has shape (num_pos, 4), the last dimension 4
+ represents [tl_x, tl_y, br_x, br_y].
+ neg_priors (Tensor): Contains all the negative boxes,
+ has shape (num_neg, 4), the last dimension 4
+ represents [tl_x, tl_y, br_x, br_y].
+ pos_gt_bboxes (Tensor): Contains gt_boxes for
+ all positive samples, has shape (num_pos, 4),
+ the last dimension 4
+ represents [tl_x, tl_y, br_x, br_y].
+ pos_gt_labels (Tensor): Contains gt_labels for
+ all positive samples, has shape (num_pos, ).
+ cfg (obj:`ConfigDict`): `train_cfg` of R-CNN.
+
+ Returns:
+ Tuple[Tensor]: Ground truth for proposals in a single image.
+ Containing the following Tensors:
+
+ - labels(Tensor): Gt_labels for all proposals, has
+ shape (num_proposals,).
+ - label_weights(Tensor): Labels_weights for all proposals, has
+ shape (num_proposals,).
+ - bbox_targets(Tensor):Regression target for all proposals, has
+ shape (num_proposals, 4), the last dimension 4
+ represents [tl_x, tl_y, br_x, br_y].
+ - bbox_weights(Tensor):Regression weights for all proposals,
+ has shape (num_proposals, 4).
+ """
+ num_pos = pos_priors.size(0)
+ num_neg = neg_priors.size(0)
+ num_samples = num_pos + num_neg
+
+ # original implementation uses new_zeros since BG are set to be 0
+ # now use empty & fill because BG cat_id = num_classes,
+ # FG cat_id = [0, num_classes-1]
+ labels = pos_priors.new_full((num_samples, ),
+ self.num_classes,
+ dtype=torch.long)
+ label_weights = pos_priors.new_zeros(num_samples)
+ bbox_targets = pos_priors.new_zeros(num_samples, 4)
+ bbox_weights = pos_priors.new_zeros(num_samples, 4)
+ if num_pos > 0:
+ labels[pos_inds] = pos_gt_labels
+ pos_weight = 1.0 if cfg.pos_weight <= 0 else cfg.pos_weight
+ label_weights[pos_inds] = pos_weight
+ if not self.reg_decoded_bbox:
+ pos_bbox_targets = self.bbox_coder.encode(
+ pos_priors, pos_gt_bboxes)
+ else:
+ pos_bbox_targets = pos_gt_bboxes
+ bbox_targets[pos_inds, :] = pos_bbox_targets
+ bbox_weights[pos_inds, :] = 1
+ if num_neg > 0:
+ label_weights[neg_inds] = 1.0
+
+ return labels, label_weights, bbox_targets, bbox_weights
+
+ def get_targets(self,
+ sampling_results: List[SamplingResult],
+ rcnn_train_cfg: ConfigDict,
+ concat: bool = True) -> tuple:
+ """Calculate the ground truth for all samples in a batch according to
+ the sampling_results.
+
+ Almost the same as the implementation in bbox_head, we passed
+ additional parameters pos_inds_list and neg_inds_list to
+ `_get_targets_single` function.
+
+ Args:
+ sampling_results (List[obj:SamplingResult]): Assign results of
+ all images in a batch after sampling.
+ rcnn_train_cfg (obj:ConfigDict): `train_cfg` of RCNN.
+ concat (bool): Whether to concatenate the results of all
+ the images in a single batch.
+
+ Returns:
+ Tuple[Tensor]: Ground truth for proposals in a single image.
+ Containing the following list of Tensors:
+
+ - labels (list[Tensor],Tensor): Gt_labels for all
+ proposals in a batch, each tensor in list has
+ shape (num_proposals,) when `concat=False`, otherwise just
+ a single tensor has shape (num_all_proposals,).
+ - label_weights (list[Tensor]): Labels_weights for
+ all proposals in a batch, each tensor in list has shape
+ (num_proposals,) when `concat=False`, otherwise just a
+ single tensor has shape (num_all_proposals,).
+ - bbox_targets (list[Tensor],Tensor): Regression target
+ for all proposals in a batch, each tensor in list has
+ shape (num_proposals, 4) when `concat=False`, otherwise
+ just a single tensor has shape (num_all_proposals, 4),
+ the last dimension 4 represents [tl_x, tl_y, br_x, br_y].
+ - bbox_weights (list[tensor],Tensor): Regression weights for
+ all proposals in a batch, each tensor in list has shape
+ (num_proposals, 4) when `concat=False`, otherwise just a
+ single tensor has shape (num_all_proposals, 4).
+ """
+ pos_inds_list = [res.pos_inds for res in sampling_results]
+ neg_inds_list = [res.neg_inds for res in sampling_results]
+ pos_priors_list = [res.pos_priors for res in sampling_results]
+ neg_priors_list = [res.neg_priors for res in sampling_results]
+ pos_gt_bboxes_list = [res.pos_gt_bboxes for res in sampling_results]
+ pos_gt_labels_list = [res.pos_gt_labels for res in sampling_results]
+ labels, label_weights, bbox_targets, bbox_weights = multi_apply(
+ self._get_targets_single,
+ pos_inds_list,
+ neg_inds_list,
+ pos_priors_list,
+ neg_priors_list,
+ pos_gt_bboxes_list,
+ pos_gt_labels_list,
+ cfg=rcnn_train_cfg)
+ if concat:
+ labels = torch.cat(labels, 0)
+ label_weights = torch.cat(label_weights, 0)
+ bbox_targets = torch.cat(bbox_targets, 0)
+ bbox_weights = torch.cat(bbox_weights, 0)
+ return labels, label_weights, bbox_targets, bbox_weights
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/double_bbox_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/double_bbox_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..076c35843375c7aef5e58786d55ebacd281d54a3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/double_bbox_head.py
@@ -0,0 +1,199 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Tuple
+
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from mmengine.model import BaseModule, ModuleList
+from torch import Tensor
+
+from mmdet.models.backbones.resnet import Bottleneck
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, MultiConfig, OptConfigType, OptMultiConfig
+from .bbox_head import BBoxHead
+
+
+class BasicResBlock(BaseModule):
+ """Basic residual block.
+
+ This block is a little different from the block in the ResNet backbone.
+ The kernel size of conv1 is 1 in this block while 3 in ResNet BasicBlock.
+
+ Args:
+ in_channels (int): Channels of the input feature map.
+ out_channels (int): Channels of the output feature map.
+ conv_cfg (:obj:`ConfigDict` or dict, optional): The config dict
+ for convolution layers.
+ norm_cfg (:obj:`ConfigDict` or dict): The config dict for
+ normalization layers.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict], optional): Initialization config dict. Defaults to None
+ """
+
+ def __init__(self,
+ in_channels: int,
+ out_channels: int,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: ConfigType = dict(type='BN'),
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+
+ # main path
+ self.conv1 = ConvModule(
+ in_channels,
+ in_channels,
+ kernel_size=3,
+ padding=1,
+ bias=False,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg)
+ self.conv2 = ConvModule(
+ in_channels,
+ out_channels,
+ kernel_size=1,
+ bias=False,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=None)
+
+ # identity path
+ self.conv_identity = ConvModule(
+ in_channels,
+ out_channels,
+ kernel_size=1,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=None)
+
+ self.relu = nn.ReLU(inplace=True)
+
+ def forward(self, x: Tensor) -> Tensor:
+ """Forward function."""
+ identity = x
+
+ x = self.conv1(x)
+ x = self.conv2(x)
+
+ identity = self.conv_identity(identity)
+ out = x + identity
+
+ out = self.relu(out)
+ return out
+
+
+@MODELS.register_module()
+class DoubleConvFCBBoxHead(BBoxHead):
+ r"""Bbox head used in Double-Head R-CNN
+
+ .. code-block:: none
+
+ /-> cls
+ /-> shared convs ->
+ \-> reg
+ roi features
+ /-> cls
+ \-> shared fc ->
+ \-> reg
+ """ # noqa: W605
+
+ def __init__(self,
+ num_convs: int = 0,
+ num_fcs: int = 0,
+ conv_out_channels: int = 1024,
+ fc_out_channels: int = 1024,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: ConfigType = dict(type='BN'),
+ init_cfg: MultiConfig = dict(
+ type='Normal',
+ override=[
+ dict(type='Normal', name='fc_cls', std=0.01),
+ dict(type='Normal', name='fc_reg', std=0.001),
+ dict(
+ type='Xavier',
+ name='fc_branch',
+ distribution='uniform')
+ ]),
+ **kwargs) -> None:
+ kwargs.setdefault('with_avg_pool', True)
+ super().__init__(init_cfg=init_cfg, **kwargs)
+ assert self.with_avg_pool
+ assert num_convs > 0
+ assert num_fcs > 0
+ self.num_convs = num_convs
+ self.num_fcs = num_fcs
+ self.conv_out_channels = conv_out_channels
+ self.fc_out_channels = fc_out_channels
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+
+ # increase the channel of input features
+ self.res_block = BasicResBlock(self.in_channels,
+ self.conv_out_channels)
+
+ # add conv heads
+ self.conv_branch = self._add_conv_branch()
+ # add fc heads
+ self.fc_branch = self._add_fc_branch()
+
+ out_dim_reg = 4 if self.reg_class_agnostic else 4 * self.num_classes
+ self.fc_reg = nn.Linear(self.conv_out_channels, out_dim_reg)
+
+ self.fc_cls = nn.Linear(self.fc_out_channels, self.num_classes + 1)
+ self.relu = nn.ReLU()
+
+ def _add_conv_branch(self) -> None:
+ """Add the fc branch which consists of a sequential of conv layers."""
+ branch_convs = ModuleList()
+ for i in range(self.num_convs):
+ branch_convs.append(
+ Bottleneck(
+ inplanes=self.conv_out_channels,
+ planes=self.conv_out_channels // 4,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ return branch_convs
+
+ def _add_fc_branch(self) -> None:
+ """Add the fc branch which consists of a sequential of fc layers."""
+ branch_fcs = ModuleList()
+ for i in range(self.num_fcs):
+ fc_in_channels = (
+ self.in_channels *
+ self.roi_feat_area if i == 0 else self.fc_out_channels)
+ branch_fcs.append(nn.Linear(fc_in_channels, self.fc_out_channels))
+ return branch_fcs
+
+ def forward(self, x_cls: Tensor, x_reg: Tensor) -> Tuple[Tensor]:
+ """Forward features from the upstream network.
+
+ Args:
+ x_cls (Tensor): Classification features of rois
+ x_reg (Tensor): Regression features from the upstream network.
+
+ Returns:
+ tuple: A tuple of classification scores and bbox prediction.
+
+ - cls_score (Tensor): Classification score predictions of rois.
+ each roi predicts num_classes + 1 channels.
+ - bbox_pred (Tensor): BBox deltas predictions of rois. each roi
+ predicts 4 * num_classes channels.
+ """
+ # conv head
+ x_conv = self.res_block(x_reg)
+
+ for conv in self.conv_branch:
+ x_conv = conv(x_conv)
+
+ if self.with_avg_pool:
+ x_conv = self.avg_pool(x_conv)
+
+ x_conv = x_conv.view(x_conv.size(0), -1)
+ bbox_pred = self.fc_reg(x_conv)
+
+ # fc head
+ x_fc = x_cls.view(x_cls.size(0), -1)
+ for fc in self.fc_branch:
+ x_fc = self.relu(fc(x_fc))
+
+ cls_score = self.fc_cls(x_fc)
+
+ return cls_score, bbox_pred
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/multi_instance_bbox_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/multi_instance_bbox_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..38e57d2eddd580b13256da63c9bd8723be98e764
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/multi_instance_bbox_head.py
@@ -0,0 +1,626 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple, Union
+
+import numpy as np
+import torch
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule
+from mmengine.config import ConfigDict
+from mmengine.structures import InstanceData
+from torch import Tensor, nn
+
+from mmdet.models.roi_heads.bbox_heads.bbox_head import BBoxHead
+from mmdet.models.task_modules.samplers import SamplingResult
+from mmdet.models.utils import empty_instances
+from mmdet.registry import MODELS
+from mmdet.structures.bbox import bbox_overlaps
+
+
+@MODELS.register_module()
+class MultiInstanceBBoxHead(BBoxHead):
+ r"""Bbox head used in CrowdDet.
+
+ .. code-block:: none
+
+ /-> cls convs_1 -> cls fcs_1 -> cls_1
+ |--
+ | \-> reg convs_1 -> reg fcs_1 -> reg_1
+ |
+ | /-> cls convs_2 -> cls fcs_2 -> cls_2
+ shared convs -> shared fcs |--
+ | \-> reg convs_2 -> reg fcs_2 -> reg_2
+ |
+ | ...
+ |
+ | /-> cls convs_k -> cls fcs_k -> cls_k
+ |--
+ \-> reg convs_k -> reg fcs_k -> reg_k
+
+
+ Args:
+ num_instance (int): The number of branches after shared fcs.
+ Defaults to 2.
+ with_refine (bool): Whether to use refine module. Defaults to False.
+ num_shared_convs (int): The number of shared convs. Defaults to 0.
+ num_shared_fcs (int): The number of shared fcs. Defaults to 2.
+ num_cls_convs (int): The number of cls convs. Defaults to 0.
+ num_cls_fcs (int): The number of cls fcs. Defaults to 0.
+ num_reg_convs (int): The number of reg convs. Defaults to 0.
+ num_reg_fcs (int): The number of reg fcs. Defaults to 0.
+ conv_out_channels (int): The number of conv out channels.
+ Defaults to 256.
+ fc_out_channels (int): The number of fc out channels. Defaults to 1024.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Defaults to None.
+ """ # noqa: W605
+
+ def __init__(self,
+ num_instance: int = 2,
+ with_refine: bool = False,
+ num_shared_convs: int = 0,
+ num_shared_fcs: int = 2,
+ num_cls_convs: int = 0,
+ num_cls_fcs: int = 0,
+ num_reg_convs: int = 0,
+ num_reg_fcs: int = 0,
+ conv_out_channels: int = 256,
+ fc_out_channels: int = 1024,
+ init_cfg: Optional[Union[dict, ConfigDict]] = None,
+ *args,
+ **kwargs) -> None:
+ super().__init__(*args, init_cfg=init_cfg, **kwargs)
+ assert (num_shared_convs + num_shared_fcs + num_cls_convs +
+ num_cls_fcs + num_reg_convs + num_reg_fcs > 0)
+ assert num_instance == 2, 'Currently only 2 instances are supported'
+ if num_cls_convs > 0 or num_reg_convs > 0:
+ assert num_shared_fcs == 0
+ if not self.with_cls:
+ assert num_cls_convs == 0 and num_cls_fcs == 0
+ if not self.with_reg:
+ assert num_reg_convs == 0 and num_reg_fcs == 0
+ self.num_instance = num_instance
+ self.num_shared_convs = num_shared_convs
+ self.num_shared_fcs = num_shared_fcs
+ self.num_cls_convs = num_cls_convs
+ self.num_cls_fcs = num_cls_fcs
+ self.num_reg_convs = num_reg_convs
+ self.num_reg_fcs = num_reg_fcs
+ self.conv_out_channels = conv_out_channels
+ self.fc_out_channels = fc_out_channels
+ self.with_refine = with_refine
+
+ # add shared convs and fcs
+ self.shared_convs, self.shared_fcs, last_layer_dim = \
+ self._add_conv_fc_branch(
+ self.num_shared_convs, self.num_shared_fcs, self.in_channels,
+ True)
+ self.shared_out_channels = last_layer_dim
+ self.relu = nn.ReLU(inplace=True)
+
+ if self.with_refine:
+ refine_model_cfg = {
+ 'type': 'Linear',
+ 'in_features': self.shared_out_channels + 20,
+ 'out_features': self.shared_out_channels
+ }
+ self.shared_fcs_ref = MODELS.build(refine_model_cfg)
+ self.fc_cls_ref = nn.ModuleList()
+ self.fc_reg_ref = nn.ModuleList()
+
+ self.cls_convs = nn.ModuleList()
+ self.cls_fcs = nn.ModuleList()
+ self.reg_convs = nn.ModuleList()
+ self.reg_fcs = nn.ModuleList()
+ self.cls_last_dim = list()
+ self.reg_last_dim = list()
+ self.fc_cls = nn.ModuleList()
+ self.fc_reg = nn.ModuleList()
+ for k in range(self.num_instance):
+ # add cls specific branch
+ cls_convs, cls_fcs, cls_last_dim = self._add_conv_fc_branch(
+ self.num_cls_convs, self.num_cls_fcs, self.shared_out_channels)
+ self.cls_convs.append(cls_convs)
+ self.cls_fcs.append(cls_fcs)
+ self.cls_last_dim.append(cls_last_dim)
+
+ # add reg specific branch
+ reg_convs, reg_fcs, reg_last_dim = self._add_conv_fc_branch(
+ self.num_reg_convs, self.num_reg_fcs, self.shared_out_channels)
+ self.reg_convs.append(reg_convs)
+ self.reg_fcs.append(reg_fcs)
+ self.reg_last_dim.append(reg_last_dim)
+
+ if self.num_shared_fcs == 0 and not self.with_avg_pool:
+ if self.num_cls_fcs == 0:
+ self.cls_last_dim *= self.roi_feat_area
+ if self.num_reg_fcs == 0:
+ self.reg_last_dim *= self.roi_feat_area
+
+ if self.with_cls:
+ if self.custom_cls_channels:
+ cls_channels = self.loss_cls.get_cls_channels(
+ self.num_classes)
+ else:
+ cls_channels = self.num_classes + 1
+ cls_predictor_cfg_ = self.cls_predictor_cfg.copy() # deepcopy
+ cls_predictor_cfg_.update(
+ in_features=self.cls_last_dim[k],
+ out_features=cls_channels)
+ self.fc_cls.append(MODELS.build(cls_predictor_cfg_))
+ if self.with_refine:
+ self.fc_cls_ref.append(MODELS.build(cls_predictor_cfg_))
+
+ if self.with_reg:
+ out_dim_reg = (4 if self.reg_class_agnostic else 4 *
+ self.num_classes)
+ reg_predictor_cfg_ = self.reg_predictor_cfg.copy()
+ reg_predictor_cfg_.update(
+ in_features=self.reg_last_dim[k], out_features=out_dim_reg)
+ self.fc_reg.append(MODELS.build(reg_predictor_cfg_))
+ if self.with_refine:
+ self.fc_reg_ref.append(MODELS.build(reg_predictor_cfg_))
+
+ if init_cfg is None:
+ # when init_cfg is None,
+ # It has been set to
+ # [[dict(type='Normal', std=0.01, override=dict(name='fc_cls'))],
+ # [dict(type='Normal', std=0.001, override=dict(name='fc_reg'))]
+ # after `super(ConvFCBBoxHead, self).__init__()`
+ # we only need to append additional configuration
+ # for `shared_fcs`, `cls_fcs` and `reg_fcs`
+ self.init_cfg += [
+ dict(
+ type='Xavier',
+ distribution='uniform',
+ override=[
+ dict(name='shared_fcs'),
+ dict(name='cls_fcs'),
+ dict(name='reg_fcs')
+ ])
+ ]
+
+ def _add_conv_fc_branch(self,
+ num_branch_convs: int,
+ num_branch_fcs: int,
+ in_channels: int,
+ is_shared: bool = False) -> tuple:
+ """Add shared or separable branch.
+
+ convs -> avg pool (optional) -> fcs
+ """
+ last_layer_dim = in_channels
+ # add branch specific conv layers
+ branch_convs = nn.ModuleList()
+ if num_branch_convs > 0:
+ for i in range(num_branch_convs):
+ conv_in_channels = (
+ last_layer_dim if i == 0 else self.conv_out_channels)
+ branch_convs.append(
+ ConvModule(
+ conv_in_channels, self.conv_out_channels, 3,
+ padding=1))
+ last_layer_dim = self.conv_out_channels
+ # add branch specific fc layers
+ branch_fcs = nn.ModuleList()
+ if num_branch_fcs > 0:
+ # for shared branch, only consider self.with_avg_pool
+ # for separated branches, also consider self.num_shared_fcs
+ if (is_shared
+ or self.num_shared_fcs == 0) and not self.with_avg_pool:
+ last_layer_dim *= self.roi_feat_area
+ for i in range(num_branch_fcs):
+ fc_in_channels = (
+ last_layer_dim if i == 0 else self.fc_out_channels)
+ branch_fcs.append(
+ nn.Linear(fc_in_channels, self.fc_out_channels))
+ last_layer_dim = self.fc_out_channels
+ return branch_convs, branch_fcs, last_layer_dim
+
+ def forward(self, x: Tuple[Tensor]) -> tuple:
+ """Forward features from the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: A tuple of classification scores and bbox prediction.
+
+ - cls_score (Tensor): Classification scores for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_base_priors * num_classes.
+ - bbox_pred (Tensor): Box energies / deltas for all scale
+ levels, each is a 4D-tensor, the channels number is
+ num_base_priors * 4.
+ - cls_score_ref (Tensor): The cls_score after refine model.
+ - bbox_pred_ref (Tensor): The bbox_pred after refine model.
+ """
+ # shared part
+ if self.num_shared_convs > 0:
+ for conv in self.shared_convs:
+ x = conv(x)
+
+ if self.num_shared_fcs > 0:
+ if self.with_avg_pool:
+ x = self.avg_pool(x)
+
+ x = x.flatten(1)
+ for fc in self.shared_fcs:
+ x = self.relu(fc(x))
+
+ x_cls = x
+ x_reg = x
+ # separate branches
+ cls_score = list()
+ bbox_pred = list()
+ for k in range(self.num_instance):
+ for conv in self.cls_convs[k]:
+ x_cls = conv(x_cls)
+ if x_cls.dim() > 2:
+ if self.with_avg_pool:
+ x_cls = self.avg_pool(x_cls)
+ x_cls = x_cls.flatten(1)
+ for fc in self.cls_fcs[k]:
+ x_cls = self.relu(fc(x_cls))
+
+ for conv in self.reg_convs[k]:
+ x_reg = conv(x_reg)
+ if x_reg.dim() > 2:
+ if self.with_avg_pool:
+ x_reg = self.avg_pool(x_reg)
+ x_reg = x_reg.flatten(1)
+ for fc in self.reg_fcs[k]:
+ x_reg = self.relu(fc(x_reg))
+
+ cls_score.append(self.fc_cls[k](x_cls) if self.with_cls else None)
+ bbox_pred.append(self.fc_reg[k](x_reg) if self.with_reg else None)
+
+ if self.with_refine:
+ x_ref = x
+ cls_score_ref = list()
+ bbox_pred_ref = list()
+ for k in range(self.num_instance):
+ feat_ref = cls_score[k].softmax(dim=-1)
+ feat_ref = torch.cat((bbox_pred[k], feat_ref[:, 1][:, None]),
+ dim=1).repeat(1, 4)
+ feat_ref = torch.cat((x_ref, feat_ref), dim=1)
+ feat_ref = F.relu_(self.shared_fcs_ref(feat_ref))
+
+ cls_score_ref.append(self.fc_cls_ref[k](feat_ref))
+ bbox_pred_ref.append(self.fc_reg_ref[k](feat_ref))
+
+ cls_score = torch.cat(cls_score, dim=1)
+ bbox_pred = torch.cat(bbox_pred, dim=1)
+ cls_score_ref = torch.cat(cls_score_ref, dim=1)
+ bbox_pred_ref = torch.cat(bbox_pred_ref, dim=1)
+ return cls_score, bbox_pred, cls_score_ref, bbox_pred_ref
+
+ cls_score = torch.cat(cls_score, dim=1)
+ bbox_pred = torch.cat(bbox_pred, dim=1)
+
+ return cls_score, bbox_pred
+
+ def get_targets(self,
+ sampling_results: List[SamplingResult],
+ rcnn_train_cfg: ConfigDict,
+ concat: bool = True) -> tuple:
+ """Calculate the ground truth for all samples in a batch according to
+ the sampling_results.
+
+ Almost the same as the implementation in bbox_head, we passed
+ additional parameters pos_inds_list and neg_inds_list to
+ `_get_targets_single` function.
+
+ Args:
+ sampling_results (List[obj:SamplingResult]): Assign results of
+ all images in a batch after sampling.
+ rcnn_train_cfg (obj:ConfigDict): `train_cfg` of RCNN.
+ concat (bool): Whether to concatenate the results of all
+ the images in a single batch.
+
+ Returns:
+ Tuple[Tensor]: Ground truth for proposals in a single image.
+ Containing the following list of Tensors:
+
+ - labels (list[Tensor],Tensor): Gt_labels for all proposals in a
+ batch, each tensor in list has shape (num_proposals,) when
+ `concat=False`, otherwise just a single tensor has shape
+ (num_all_proposals,).
+ - label_weights (list[Tensor]): Labels_weights for
+ all proposals in a batch, each tensor in list has shape
+ (num_proposals,) when `concat=False`, otherwise just a single
+ tensor has shape (num_all_proposals,).
+ - bbox_targets (list[Tensor],Tensor): Regression target for all
+ proposals in a batch, each tensor in list has shape
+ (num_proposals, 4) when `concat=False`, otherwise just a single
+ tensor has shape (num_all_proposals, 4), the last dimension 4
+ represents [tl_x, tl_y, br_x, br_y].
+ - bbox_weights (list[tensor],Tensor): Regression weights for
+ all proposals in a batch, each tensor in list has shape
+ (num_proposals, 4) when `concat=False`, otherwise just a
+ single tensor has shape (num_all_proposals, 4).
+ """
+ labels = []
+ bbox_targets = []
+ bbox_weights = []
+ label_weights = []
+ for i in range(len(sampling_results)):
+ sample_bboxes = torch.cat([
+ sampling_results[i].pos_gt_bboxes,
+ sampling_results[i].neg_gt_bboxes
+ ])
+ sample_priors = sampling_results[i].priors
+ sample_priors = sample_priors.repeat(1, self.num_instance).reshape(
+ -1, 4)
+ sample_bboxes = sample_bboxes.reshape(-1, 4)
+
+ if not self.reg_decoded_bbox:
+ _bbox_targets = self.bbox_coder.encode(sample_priors,
+ sample_bboxes)
+ else:
+ _bbox_targets = sample_priors
+ _bbox_targets = _bbox_targets.reshape(-1, self.num_instance * 4)
+ _bbox_weights = torch.ones(_bbox_targets.shape)
+ _labels = torch.cat([
+ sampling_results[i].pos_gt_labels,
+ sampling_results[i].neg_gt_labels
+ ])
+ _labels_weights = torch.ones(_labels.shape)
+
+ bbox_targets.append(_bbox_targets)
+ bbox_weights.append(_bbox_weights)
+ labels.append(_labels)
+ label_weights.append(_labels_weights)
+
+ if concat:
+ labels = torch.cat(labels, 0)
+ label_weights = torch.cat(label_weights, 0)
+ bbox_targets = torch.cat(bbox_targets, 0)
+ bbox_weights = torch.cat(bbox_weights, 0)
+ return labels, label_weights, bbox_targets, bbox_weights
+
+ def loss(self, cls_score: Tensor, bbox_pred: Tensor, rois: Tensor,
+ labels: Tensor, label_weights: Tensor, bbox_targets: Tensor,
+ bbox_weights: Tensor, **kwargs) -> dict:
+ """Calculate the loss based on the network predictions and targets.
+
+ Args:
+ cls_score (Tensor): Classification prediction results of all class,
+ has shape (batch_size * num_proposals_single_image,
+ (num_classes + 1) * k), k represents the number of prediction
+ boxes generated by each proposal box.
+ bbox_pred (Tensor): Regression prediction results, has shape
+ (batch_size * num_proposals_single_image, 4 * k), the last
+ dimension 4 represents [tl_x, tl_y, br_x, br_y].
+ rois (Tensor): RoIs with the shape
+ (batch_size * num_proposals_single_image, 5) where the first
+ column indicates batch id of each RoI.
+ labels (Tensor): Gt_labels for all proposals in a batch, has
+ shape (batch_size * num_proposals_single_image, k).
+ label_weights (Tensor): Labels_weights for all proposals in a
+ batch, has shape (batch_size * num_proposals_single_image, k).
+ bbox_targets (Tensor): Regression target for all proposals in a
+ batch, has shape (batch_size * num_proposals_single_image,
+ 4 * k), the last dimension 4 represents [tl_x, tl_y, br_x,
+ br_y].
+ bbox_weights (Tensor): Regression weights for all proposals in a
+ batch, has shape (batch_size * num_proposals_single_image,
+ 4 * k).
+
+ Returns:
+ dict: A dictionary of loss.
+ """
+ losses = dict()
+ if bbox_pred.numel():
+ loss_0 = self.emd_loss(bbox_pred[:, 0:4], cls_score[:, 0:2],
+ bbox_pred[:, 4:8], cls_score[:, 2:4],
+ bbox_targets, labels)
+ loss_1 = self.emd_loss(bbox_pred[:, 4:8], cls_score[:, 2:4],
+ bbox_pred[:, 0:4], cls_score[:, 0:2],
+ bbox_targets, labels)
+ loss = torch.cat([loss_0, loss_1], dim=1)
+ _, min_indices = loss.min(dim=1)
+ loss_emd = loss[torch.arange(loss.shape[0]), min_indices]
+ loss_emd = loss_emd.mean()
+ else:
+ loss_emd = bbox_pred.sum()
+ losses['loss_rcnn_emd'] = loss_emd
+ return losses
+
+ def emd_loss(self, bbox_pred_0: Tensor, cls_score_0: Tensor,
+ bbox_pred_1: Tensor, cls_score_1: Tensor, targets: Tensor,
+ labels: Tensor) -> Tensor:
+ """Calculate the emd loss.
+
+ Note:
+ This implementation is modified from https://github.com/Purkialo/
+ CrowdDet/blob/master/lib/det_oprs/loss_opr.py
+
+ Args:
+ bbox_pred_0 (Tensor): Part of regression prediction results, has
+ shape (batch_size * num_proposals_single_image, 4), the last
+ dimension 4 represents [tl_x, tl_y, br_x, br_y].
+ cls_score_0 (Tensor): Part of classification prediction results,
+ has shape (batch_size * num_proposals_single_image,
+ (num_classes + 1)), where 1 represents the background.
+ bbox_pred_1 (Tensor): The other part of regression prediction
+ results, has shape (batch_size*num_proposals_single_image, 4).
+ cls_score_1 (Tensor):The other part of classification prediction
+ results, has shape (batch_size * num_proposals_single_image,
+ (num_classes + 1)).
+ targets (Tensor):Regression target for all proposals in a
+ batch, has shape (batch_size * num_proposals_single_image,
+ 4 * k), the last dimension 4 represents [tl_x, tl_y, br_x,
+ br_y], k represents the number of prediction boxes generated
+ by each proposal box.
+ labels (Tensor): Gt_labels for all proposals in a batch, has
+ shape (batch_size * num_proposals_single_image, k).
+
+ Returns:
+ torch.Tensor: The calculated loss.
+ """
+
+ bbox_pred = torch.cat([bbox_pred_0, bbox_pred_1],
+ dim=1).reshape(-1, bbox_pred_0.shape[-1])
+ cls_score = torch.cat([cls_score_0, cls_score_1],
+ dim=1).reshape(-1, cls_score_0.shape[-1])
+ targets = targets.reshape(-1, 4)
+ labels = labels.long().flatten()
+
+ # masks
+ valid_masks = labels >= 0
+ fg_masks = labels > 0
+
+ # multiple class
+ bbox_pred = bbox_pred.reshape(-1, self.num_classes, 4)
+ fg_gt_classes = labels[fg_masks]
+ bbox_pred = bbox_pred[fg_masks, fg_gt_classes - 1, :]
+
+ # loss for regression
+ loss_bbox = self.loss_bbox(bbox_pred, targets[fg_masks])
+ loss_bbox = loss_bbox.sum(dim=1)
+
+ # loss for classification
+ labels = labels * valid_masks
+ loss_cls = self.loss_cls(cls_score, labels)
+
+ loss_cls[fg_masks] = loss_cls[fg_masks] + loss_bbox
+ loss = loss_cls.reshape(-1, 2).sum(dim=1)
+ return loss.reshape(-1, 1)
+
+ def _predict_by_feat_single(
+ self,
+ roi: Tensor,
+ cls_score: Tensor,
+ bbox_pred: Tensor,
+ img_meta: dict,
+ rescale: bool = False,
+ rcnn_test_cfg: Optional[ConfigDict] = None) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results.
+
+ Args:
+ roi (Tensor): Boxes to be transformed. Has shape (num_boxes, 5).
+ last dimension 5 arrange as (batch_index, x1, y1, x2, y2).
+ cls_score (Tensor): Box scores, has shape
+ (num_boxes, num_classes + 1).
+ bbox_pred (Tensor): Box energies / deltas. has shape
+ (num_boxes, num_classes * 4).
+ img_meta (dict): image information.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ rcnn_test_cfg (obj:`ConfigDict`): `test_cfg` of Bbox Head.
+ Defaults to None
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+
+ cls_score = cls_score.reshape(-1, self.num_classes + 1)
+ bbox_pred = bbox_pred.reshape(-1, 4)
+ roi = roi.repeat_interleave(self.num_instance, dim=0)
+
+ results = InstanceData()
+ if roi.shape[0] == 0:
+ return empty_instances([img_meta],
+ roi.device,
+ task_type='bbox',
+ instance_results=[results])[0]
+
+ scores = cls_score.softmax(dim=-1) if cls_score is not None else None
+ img_shape = img_meta['img_shape']
+ bboxes = self.bbox_coder.decode(
+ roi[..., 1:], bbox_pred, max_shape=img_shape)
+
+ if rescale and bboxes.size(0) > 0:
+ assert img_meta.get('scale_factor') is not None
+ scale_factor = bboxes.new_tensor(img_meta['scale_factor']).repeat(
+ (1, 2))
+ bboxes = (bboxes.view(bboxes.size(0), -1, 4) / scale_factor).view(
+ bboxes.size()[0], -1)
+
+ if rcnn_test_cfg is None:
+ # This means that it is aug test.
+ # It needs to return the raw results without nms.
+ results.bboxes = bboxes
+ results.scores = scores
+ else:
+ roi_idx = np.tile(
+ np.arange(bboxes.shape[0] / self.num_instance)[:, None],
+ (1, self.num_instance)).reshape(-1, 1)[:, 0]
+ roi_idx = torch.from_numpy(roi_idx).to(bboxes.device).reshape(
+ -1, 1)
+ bboxes = torch.cat([bboxes, roi_idx], dim=1)
+ det_bboxes, det_scores = self.set_nms(
+ bboxes, scores[:, 1], rcnn_test_cfg.score_thr,
+ rcnn_test_cfg.nms['iou_threshold'], rcnn_test_cfg.max_per_img)
+
+ results.bboxes = det_bboxes[:, :-1]
+ results.scores = det_scores
+ results.labels = torch.zeros_like(det_scores)
+
+ return results
+
+ @staticmethod
+ def set_nms(bboxes: Tensor,
+ scores: Tensor,
+ score_thr: float,
+ iou_threshold: float,
+ max_num: int = -1) -> Tuple[Tensor, Tensor]:
+ """NMS for multi-instance prediction. Please refer to
+ https://github.com/Purkialo/CrowdDet for more details.
+
+ Args:
+ bboxes (Tensor): predict bboxes.
+ scores (Tensor): The score of each predict bbox.
+ score_thr (float): bbox threshold, bboxes with scores lower than it
+ will not be considered.
+ iou_threshold (float): IoU threshold to be considered as
+ conflicted.
+ max_num (int, optional): if there are more than max_num bboxes
+ after NMS, only top max_num will be kept. Default to -1.
+
+ Returns:
+ Tuple[Tensor, Tensor]: (bboxes, scores).
+ """
+
+ bboxes = bboxes[scores > score_thr]
+ scores = scores[scores > score_thr]
+
+ ordered_scores, order = scores.sort(descending=True)
+ ordered_bboxes = bboxes[order]
+ roi_idx = ordered_bboxes[:, -1]
+
+ keep = torch.ones(len(ordered_bboxes)) == 1
+ ruler = torch.arange(len(ordered_bboxes))
+
+ keep = keep.to(bboxes.device)
+ ruler = ruler.to(bboxes.device)
+
+ while ruler.shape[0] > 0:
+ basement = ruler[0]
+ ruler = ruler[1:]
+ idx = roi_idx[basement]
+ # calculate the body overlap
+ basement_bbox = ordered_bboxes[:, :4][basement].reshape(-1, 4)
+ ruler_bbox = ordered_bboxes[:, :4][ruler].reshape(-1, 4)
+ overlap = bbox_overlaps(basement_bbox, ruler_bbox)
+ indices = torch.where(overlap > iou_threshold)[1]
+ loc = torch.where(roi_idx[ruler][indices] == idx)
+ # the mask won't change in the step
+ mask = keep[ruler[indices][loc]]
+ keep[ruler[indices]] = False
+ keep[ruler[indices][loc][mask]] = True
+ ruler[~keep[ruler]] = -1
+ ruler = ruler[ruler > 0]
+
+ keep = keep[order.sort()[1]]
+ return bboxes[keep][:max_num, :], scores[keep][:max_num]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/sabl_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/sabl_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..9a9ee6aba9669514ec8ce7218e8c97e026830f6c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/sabl_head.py
@@ -0,0 +1,684 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Sequence, Tuple
+
+import numpy as np
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule
+from mmengine.config import ConfigDict
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.models.layers import multiclass_nms
+from mmdet.models.losses import accuracy
+from mmdet.models.task_modules import SamplingResult
+from mmdet.models.utils import multi_apply
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.utils import ConfigType, InstanceList, OptConfigType, OptMultiConfig
+from .bbox_head import BBoxHead
+
+
+@MODELS.register_module()
+class SABLHead(BBoxHead):
+ """Side-Aware Boundary Localization (SABL) for RoI-Head.
+
+ Side-Aware features are extracted by conv layers
+ with an attention mechanism.
+ Boundary Localization with Bucketing and Bucketing Guided Rescoring
+ are implemented in BucketingBBoxCoder.
+
+ Please refer to https://arxiv.org/abs/1912.04260 for more details.
+
+ Args:
+ cls_in_channels (int): Input channels of cls RoI feature. \
+ Defaults to 256.
+ reg_in_channels (int): Input channels of reg RoI feature. \
+ Defaults to 256.
+ roi_feat_size (int): Size of RoI features. Defaults to 7.
+ reg_feat_up_ratio (int): Upsample ratio of reg features. \
+ Defaults to 2.
+ reg_pre_kernel (int): Kernel of 2D conv layers before \
+ attention pooling. Defaults to 3.
+ reg_post_kernel (int): Kernel of 1D conv layers after \
+ attention pooling. Defaults to 3.
+ reg_pre_num (int): Number of pre convs. Defaults to 2.
+ reg_post_num (int): Number of post convs. Defaults to 1.
+ num_classes (int): Number of classes in dataset. Defaults to 80.
+ cls_out_channels (int): Hidden channels in cls fcs. Defaults to 1024.
+ reg_offset_out_channels (int): Hidden and output channel \
+ of reg offset branch. Defaults to 256.
+ reg_cls_out_channels (int): Hidden and output channel \
+ of reg cls branch. Defaults to 256.
+ num_cls_fcs (int): Number of fcs for cls branch. Defaults to 1.
+ num_reg_fcs (int): Number of fcs for reg branch.. Defaults to 0.
+ reg_class_agnostic (bool): Class agnostic regression or not. \
+ Defaults to True.
+ norm_cfg (dict): Config of norm layers. Defaults to None.
+ bbox_coder (dict): Config of bbox coder. Defaults 'BucketingBBoxCoder'.
+ loss_cls (dict): Config of classification loss.
+ loss_bbox_cls (dict): Config of classification loss for bbox branch.
+ loss_bbox_reg (dict): Config of regression loss for bbox branch.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ num_classes: int,
+ cls_in_channels: int = 256,
+ reg_in_channels: int = 256,
+ roi_feat_size: int = 7,
+ reg_feat_up_ratio: int = 2,
+ reg_pre_kernel: int = 3,
+ reg_post_kernel: int = 3,
+ reg_pre_num: int = 2,
+ reg_post_num: int = 1,
+ cls_out_channels: int = 1024,
+ reg_offset_out_channels: int = 256,
+ reg_cls_out_channels: int = 256,
+ num_cls_fcs: int = 1,
+ num_reg_fcs: int = 0,
+ reg_class_agnostic: bool = True,
+ norm_cfg: OptConfigType = None,
+ bbox_coder: ConfigType = dict(
+ type='BucketingBBoxCoder',
+ num_buckets=14,
+ scale_factor=1.7),
+ loss_cls: ConfigType = dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ loss_bbox_cls: ConfigType = dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ loss_weight=1.0),
+ loss_bbox_reg: ConfigType = dict(
+ type='SmoothL1Loss', beta=0.1, loss_weight=1.0),
+ init_cfg: OptMultiConfig = None) -> None:
+ super(BBoxHead, self).__init__(init_cfg=init_cfg)
+ self.cls_in_channels = cls_in_channels
+ self.reg_in_channels = reg_in_channels
+ self.roi_feat_size = roi_feat_size
+ self.reg_feat_up_ratio = int(reg_feat_up_ratio)
+ self.num_buckets = bbox_coder['num_buckets']
+ assert self.reg_feat_up_ratio // 2 >= 1
+ self.up_reg_feat_size = roi_feat_size * self.reg_feat_up_ratio
+ assert self.up_reg_feat_size == bbox_coder['num_buckets']
+ self.reg_pre_kernel = reg_pre_kernel
+ self.reg_post_kernel = reg_post_kernel
+ self.reg_pre_num = reg_pre_num
+ self.reg_post_num = reg_post_num
+ self.num_classes = num_classes
+ self.cls_out_channels = cls_out_channels
+ self.reg_offset_out_channels = reg_offset_out_channels
+ self.reg_cls_out_channels = reg_cls_out_channels
+ self.num_cls_fcs = num_cls_fcs
+ self.num_reg_fcs = num_reg_fcs
+ self.reg_class_agnostic = reg_class_agnostic
+ assert self.reg_class_agnostic
+ self.norm_cfg = norm_cfg
+
+ self.bbox_coder = TASK_UTILS.build(bbox_coder)
+ self.loss_cls = MODELS.build(loss_cls)
+ self.loss_bbox_cls = MODELS.build(loss_bbox_cls)
+ self.loss_bbox_reg = MODELS.build(loss_bbox_reg)
+
+ self.cls_fcs = self._add_fc_branch(self.num_cls_fcs,
+ self.cls_in_channels,
+ self.roi_feat_size,
+ self.cls_out_channels)
+
+ self.side_num = int(np.ceil(self.num_buckets / 2))
+
+ if self.reg_feat_up_ratio > 1:
+ self.upsample_x = nn.ConvTranspose1d(
+ reg_in_channels,
+ reg_in_channels,
+ self.reg_feat_up_ratio,
+ stride=self.reg_feat_up_ratio)
+ self.upsample_y = nn.ConvTranspose1d(
+ reg_in_channels,
+ reg_in_channels,
+ self.reg_feat_up_ratio,
+ stride=self.reg_feat_up_ratio)
+
+ self.reg_pre_convs = nn.ModuleList()
+ for i in range(self.reg_pre_num):
+ reg_pre_conv = ConvModule(
+ reg_in_channels,
+ reg_in_channels,
+ kernel_size=reg_pre_kernel,
+ padding=reg_pre_kernel // 2,
+ norm_cfg=norm_cfg,
+ act_cfg=dict(type='ReLU'))
+ self.reg_pre_convs.append(reg_pre_conv)
+
+ self.reg_post_conv_xs = nn.ModuleList()
+ for i in range(self.reg_post_num):
+ reg_post_conv_x = ConvModule(
+ reg_in_channels,
+ reg_in_channels,
+ kernel_size=(1, reg_post_kernel),
+ padding=(0, reg_post_kernel // 2),
+ norm_cfg=norm_cfg,
+ act_cfg=dict(type='ReLU'))
+ self.reg_post_conv_xs.append(reg_post_conv_x)
+ self.reg_post_conv_ys = nn.ModuleList()
+ for i in range(self.reg_post_num):
+ reg_post_conv_y = ConvModule(
+ reg_in_channels,
+ reg_in_channels,
+ kernel_size=(reg_post_kernel, 1),
+ padding=(reg_post_kernel // 2, 0),
+ norm_cfg=norm_cfg,
+ act_cfg=dict(type='ReLU'))
+ self.reg_post_conv_ys.append(reg_post_conv_y)
+
+ self.reg_conv_att_x = nn.Conv2d(reg_in_channels, 1, 1)
+ self.reg_conv_att_y = nn.Conv2d(reg_in_channels, 1, 1)
+
+ self.fc_cls = nn.Linear(self.cls_out_channels, self.num_classes + 1)
+ self.relu = nn.ReLU(inplace=True)
+
+ self.reg_cls_fcs = self._add_fc_branch(self.num_reg_fcs,
+ self.reg_in_channels, 1,
+ self.reg_cls_out_channels)
+ self.reg_offset_fcs = self._add_fc_branch(self.num_reg_fcs,
+ self.reg_in_channels, 1,
+ self.reg_offset_out_channels)
+ self.fc_reg_cls = nn.Linear(self.reg_cls_out_channels, 1)
+ self.fc_reg_offset = nn.Linear(self.reg_offset_out_channels, 1)
+
+ if init_cfg is None:
+ self.init_cfg = [
+ dict(
+ type='Xavier',
+ layer='Linear',
+ distribution='uniform',
+ override=[
+ dict(type='Normal', name='reg_conv_att_x', std=0.01),
+ dict(type='Normal', name='reg_conv_att_y', std=0.01),
+ dict(type='Normal', name='fc_reg_cls', std=0.01),
+ dict(type='Normal', name='fc_cls', std=0.01),
+ dict(type='Normal', name='fc_reg_offset', std=0.001)
+ ])
+ ]
+ if self.reg_feat_up_ratio > 1:
+ self.init_cfg += [
+ dict(
+ type='Kaiming',
+ distribution='normal',
+ override=[
+ dict(name='upsample_x'),
+ dict(name='upsample_y')
+ ])
+ ]
+
+ def _add_fc_branch(self, num_branch_fcs: int, in_channels: int,
+ roi_feat_size: int,
+ fc_out_channels: int) -> nn.ModuleList:
+ """build fc layers."""
+ in_channels = in_channels * roi_feat_size * roi_feat_size
+ branch_fcs = nn.ModuleList()
+ for i in range(num_branch_fcs):
+ fc_in_channels = (in_channels if i == 0 else fc_out_channels)
+ branch_fcs.append(nn.Linear(fc_in_channels, fc_out_channels))
+ return branch_fcs
+
+ def cls_forward(self, cls_x: Tensor) -> Tensor:
+ """forward of classification fc layers."""
+ cls_x = cls_x.view(cls_x.size(0), -1)
+ for fc in self.cls_fcs:
+ cls_x = self.relu(fc(cls_x))
+ cls_score = self.fc_cls(cls_x)
+ return cls_score
+
+ def attention_pool(self, reg_x: Tensor) -> tuple:
+ """Extract direction-specific features fx and fy with attention
+ methanism."""
+ reg_fx = reg_x
+ reg_fy = reg_x
+ reg_fx_att = self.reg_conv_att_x(reg_fx).sigmoid()
+ reg_fy_att = self.reg_conv_att_y(reg_fy).sigmoid()
+ reg_fx_att = reg_fx_att / reg_fx_att.sum(dim=2).unsqueeze(2)
+ reg_fy_att = reg_fy_att / reg_fy_att.sum(dim=3).unsqueeze(3)
+ reg_fx = (reg_fx * reg_fx_att).sum(dim=2)
+ reg_fy = (reg_fy * reg_fy_att).sum(dim=3)
+ return reg_fx, reg_fy
+
+ def side_aware_feature_extractor(self, reg_x: Tensor) -> tuple:
+ """Refine and extract side-aware features without split them."""
+ for reg_pre_conv in self.reg_pre_convs:
+ reg_x = reg_pre_conv(reg_x)
+ reg_fx, reg_fy = self.attention_pool(reg_x)
+
+ if self.reg_post_num > 0:
+ reg_fx = reg_fx.unsqueeze(2)
+ reg_fy = reg_fy.unsqueeze(3)
+ for i in range(self.reg_post_num):
+ reg_fx = self.reg_post_conv_xs[i](reg_fx)
+ reg_fy = self.reg_post_conv_ys[i](reg_fy)
+ reg_fx = reg_fx.squeeze(2)
+ reg_fy = reg_fy.squeeze(3)
+ if self.reg_feat_up_ratio > 1:
+ reg_fx = self.relu(self.upsample_x(reg_fx))
+ reg_fy = self.relu(self.upsample_y(reg_fy))
+ reg_fx = torch.transpose(reg_fx, 1, 2)
+ reg_fy = torch.transpose(reg_fy, 1, 2)
+ return reg_fx.contiguous(), reg_fy.contiguous()
+
+ def reg_pred(self, x: Tensor, offset_fcs: nn.ModuleList,
+ cls_fcs: nn.ModuleList) -> tuple:
+ """Predict bucketing estimation (cls_pred) and fine regression (offset
+ pred) with side-aware features."""
+ x_offset = x.view(-1, self.reg_in_channels)
+ x_cls = x.view(-1, self.reg_in_channels)
+
+ for fc in offset_fcs:
+ x_offset = self.relu(fc(x_offset))
+ for fc in cls_fcs:
+ x_cls = self.relu(fc(x_cls))
+ offset_pred = self.fc_reg_offset(x_offset)
+ cls_pred = self.fc_reg_cls(x_cls)
+
+ offset_pred = offset_pred.view(x.size(0), -1)
+ cls_pred = cls_pred.view(x.size(0), -1)
+
+ return offset_pred, cls_pred
+
+ def side_aware_split(self, feat: Tensor) -> Tensor:
+ """Split side-aware features aligned with orders of bucketing
+ targets."""
+ l_end = int(np.ceil(self.up_reg_feat_size / 2))
+ r_start = int(np.floor(self.up_reg_feat_size / 2))
+ feat_fl = feat[:, :l_end]
+ feat_fr = feat[:, r_start:].flip(dims=(1, ))
+ feat_fl = feat_fl.contiguous()
+ feat_fr = feat_fr.contiguous()
+ feat = torch.cat([feat_fl, feat_fr], dim=-1)
+ return feat
+
+ def bbox_pred_split(self, bbox_pred: tuple,
+ num_proposals_per_img: Sequence[int]) -> tuple:
+ """Split batch bbox prediction back to each image."""
+ bucket_cls_preds, bucket_offset_preds = bbox_pred
+ bucket_cls_preds = bucket_cls_preds.split(num_proposals_per_img, 0)
+ bucket_offset_preds = bucket_offset_preds.split(
+ num_proposals_per_img, 0)
+ bbox_pred = tuple(zip(bucket_cls_preds, bucket_offset_preds))
+ return bbox_pred
+
+ def reg_forward(self, reg_x: Tensor) -> tuple:
+ """forward of regression branch."""
+ outs = self.side_aware_feature_extractor(reg_x)
+ edge_offset_preds = []
+ edge_cls_preds = []
+ reg_fx = outs[0]
+ reg_fy = outs[1]
+ offset_pred_x, cls_pred_x = self.reg_pred(reg_fx, self.reg_offset_fcs,
+ self.reg_cls_fcs)
+ offset_pred_y, cls_pred_y = self.reg_pred(reg_fy, self.reg_offset_fcs,
+ self.reg_cls_fcs)
+ offset_pred_x = self.side_aware_split(offset_pred_x)
+ offset_pred_y = self.side_aware_split(offset_pred_y)
+ cls_pred_x = self.side_aware_split(cls_pred_x)
+ cls_pred_y = self.side_aware_split(cls_pred_y)
+ edge_offset_preds = torch.cat([offset_pred_x, offset_pred_y], dim=-1)
+ edge_cls_preds = torch.cat([cls_pred_x, cls_pred_y], dim=-1)
+
+ return edge_cls_preds, edge_offset_preds
+
+ def forward(self, x: Tensor) -> tuple:
+ """Forward features from the upstream network."""
+ bbox_pred = self.reg_forward(x)
+ cls_score = self.cls_forward(x)
+
+ return cls_score, bbox_pred
+
+ def get_targets(self,
+ sampling_results: List[SamplingResult],
+ rcnn_train_cfg: ConfigDict,
+ concat: bool = True) -> tuple:
+ """Calculate the ground truth for all samples in a batch according to
+ the sampling_results."""
+ pos_proposals = [res.pos_bboxes for res in sampling_results]
+ neg_proposals = [res.neg_bboxes for res in sampling_results]
+ pos_gt_bboxes = [res.pos_gt_bboxes for res in sampling_results]
+ pos_gt_labels = [res.pos_gt_labels for res in sampling_results]
+ cls_reg_targets = self.bucket_target(
+ pos_proposals,
+ neg_proposals,
+ pos_gt_bboxes,
+ pos_gt_labels,
+ rcnn_train_cfg,
+ concat=concat)
+ (labels, label_weights, bucket_cls_targets, bucket_cls_weights,
+ bucket_offset_targets, bucket_offset_weights) = cls_reg_targets
+ return (labels, label_weights, (bucket_cls_targets,
+ bucket_offset_targets),
+ (bucket_cls_weights, bucket_offset_weights))
+
+ def bucket_target(self,
+ pos_proposals_list: list,
+ neg_proposals_list: list,
+ pos_gt_bboxes_list: list,
+ pos_gt_labels_list: list,
+ rcnn_train_cfg: ConfigDict,
+ concat: bool = True) -> tuple:
+ """Compute bucketing estimation targets and fine regression targets for
+ a batch of images."""
+ (labels, label_weights, bucket_cls_targets, bucket_cls_weights,
+ bucket_offset_targets, bucket_offset_weights) = multi_apply(
+ self._bucket_target_single,
+ pos_proposals_list,
+ neg_proposals_list,
+ pos_gt_bboxes_list,
+ pos_gt_labels_list,
+ cfg=rcnn_train_cfg)
+
+ if concat:
+ labels = torch.cat(labels, 0)
+ label_weights = torch.cat(label_weights, 0)
+ bucket_cls_targets = torch.cat(bucket_cls_targets, 0)
+ bucket_cls_weights = torch.cat(bucket_cls_weights, 0)
+ bucket_offset_targets = torch.cat(bucket_offset_targets, 0)
+ bucket_offset_weights = torch.cat(bucket_offset_weights, 0)
+ return (labels, label_weights, bucket_cls_targets, bucket_cls_weights,
+ bucket_offset_targets, bucket_offset_weights)
+
+ def _bucket_target_single(self, pos_proposals: Tensor,
+ neg_proposals: Tensor, pos_gt_bboxes: Tensor,
+ pos_gt_labels: Tensor, cfg: ConfigDict) -> tuple:
+ """Compute bucketing estimation targets and fine regression targets for
+ a single image.
+
+ Args:
+ pos_proposals (Tensor): positive proposals of a single image,
+ Shape (n_pos, 4)
+ neg_proposals (Tensor): negative proposals of a single image,
+ Shape (n_neg, 4).
+ pos_gt_bboxes (Tensor): gt bboxes assigned to positive proposals
+ of a single image, Shape (n_pos, 4).
+ pos_gt_labels (Tensor): gt labels assigned to positive proposals
+ of a single image, Shape (n_pos, ).
+ cfg (dict): Config of calculating targets
+
+ Returns:
+ tuple:
+
+ - labels (Tensor): Labels in a single image. Shape (n,).
+ - label_weights (Tensor): Label weights in a single image.
+ Shape (n,)
+ - bucket_cls_targets (Tensor): Bucket cls targets in
+ a single image. Shape (n, num_buckets*2).
+ - bucket_cls_weights (Tensor): Bucket cls weights in
+ a single image. Shape (n, num_buckets*2).
+ - bucket_offset_targets (Tensor): Bucket offset targets
+ in a single image. Shape (n, num_buckets*2).
+ - bucket_offset_targets (Tensor): Bucket offset weights
+ in a single image. Shape (n, num_buckets*2).
+ """
+ num_pos = pos_proposals.size(0)
+ num_neg = neg_proposals.size(0)
+ num_samples = num_pos + num_neg
+ labels = pos_gt_bboxes.new_full((num_samples, ),
+ self.num_classes,
+ dtype=torch.long)
+ label_weights = pos_proposals.new_zeros(num_samples)
+ bucket_cls_targets = pos_proposals.new_zeros(num_samples,
+ 4 * self.side_num)
+ bucket_cls_weights = pos_proposals.new_zeros(num_samples,
+ 4 * self.side_num)
+ bucket_offset_targets = pos_proposals.new_zeros(
+ num_samples, 4 * self.side_num)
+ bucket_offset_weights = pos_proposals.new_zeros(
+ num_samples, 4 * self.side_num)
+ if num_pos > 0:
+ labels[:num_pos] = pos_gt_labels
+ label_weights[:num_pos] = 1.0
+ (pos_bucket_offset_targets, pos_bucket_offset_weights,
+ pos_bucket_cls_targets,
+ pos_bucket_cls_weights) = self.bbox_coder.encode(
+ pos_proposals, pos_gt_bboxes)
+ bucket_cls_targets[:num_pos, :] = pos_bucket_cls_targets
+ bucket_cls_weights[:num_pos, :] = pos_bucket_cls_weights
+ bucket_offset_targets[:num_pos, :] = pos_bucket_offset_targets
+ bucket_offset_weights[:num_pos, :] = pos_bucket_offset_weights
+ if num_neg > 0:
+ label_weights[-num_neg:] = 1.0
+ return (labels, label_weights, bucket_cls_targets, bucket_cls_weights,
+ bucket_offset_targets, bucket_offset_weights)
+
+ def loss(self,
+ cls_score: Tensor,
+ bbox_pred: Tuple[Tensor, Tensor],
+ rois: Tensor,
+ labels: Tensor,
+ label_weights: Tensor,
+ bbox_targets: Tuple[Tensor, Tensor],
+ bbox_weights: Tuple[Tensor, Tensor],
+ reduction_override: Optional[str] = None) -> dict:
+ """Calculate the loss based on the network predictions and targets.
+
+ Args:
+ cls_score (Tensor): Classification prediction
+ results of all class, has shape
+ (batch_size * num_proposals_single_image, num_classes)
+ bbox_pred (Tensor): A tuple of regression prediction results
+ containing `bucket_cls_preds and` `bucket_offset_preds`.
+ rois (Tensor): RoIs with the shape
+ (batch_size * num_proposals_single_image, 5) where the first
+ column indicates batch id of each RoI.
+ labels (Tensor): Gt_labels for all proposals in a batch, has
+ shape (batch_size * num_proposals_single_image, ).
+ label_weights (Tensor): Labels_weights for all proposals in a
+ batch, has shape (batch_size * num_proposals_single_image, ).
+ bbox_targets (Tuple[Tensor, Tensor]): A tuple of regression target
+ containing `bucket_cls_targets` and `bucket_offset_targets`.
+ the last dimension 4 represents [tl_x, tl_y, br_x, br_y].
+ bbox_weights (Tuple[Tensor, Tensor]): A tuple of regression
+ weights containing `bucket_cls_weights` and
+ `bucket_offset_weights`.
+ reduction_override (str, optional): The reduction
+ method used to override the original reduction
+ method of the loss. Options are "none",
+ "mean" and "sum". Defaults to None,
+
+ Returns:
+ dict: A dictionary of loss.
+ """
+ losses = dict()
+ if cls_score is not None:
+ avg_factor = max(torch.sum(label_weights > 0).float().item(), 1.)
+ losses['loss_cls'] = self.loss_cls(
+ cls_score,
+ labels,
+ label_weights,
+ avg_factor=avg_factor,
+ reduction_override=reduction_override)
+ losses['acc'] = accuracy(cls_score, labels)
+
+ if bbox_pred is not None:
+ bucket_cls_preds, bucket_offset_preds = bbox_pred
+ bucket_cls_targets, bucket_offset_targets = bbox_targets
+ bucket_cls_weights, bucket_offset_weights = bbox_weights
+ # edge cls
+ bucket_cls_preds = bucket_cls_preds.view(-1, self.side_num)
+ bucket_cls_targets = bucket_cls_targets.view(-1, self.side_num)
+ bucket_cls_weights = bucket_cls_weights.view(-1, self.side_num)
+ losses['loss_bbox_cls'] = self.loss_bbox_cls(
+ bucket_cls_preds,
+ bucket_cls_targets,
+ bucket_cls_weights,
+ avg_factor=bucket_cls_targets.size(0),
+ reduction_override=reduction_override)
+
+ losses['loss_bbox_reg'] = self.loss_bbox_reg(
+ bucket_offset_preds,
+ bucket_offset_targets,
+ bucket_offset_weights,
+ avg_factor=bucket_offset_targets.size(0),
+ reduction_override=reduction_override)
+
+ return losses
+
+ def _predict_by_feat_single(
+ self,
+ roi: Tensor,
+ cls_score: Tensor,
+ bbox_pred: Tuple[Tensor, Tensor],
+ img_meta: dict,
+ rescale: bool = False,
+ rcnn_test_cfg: Optional[ConfigDict] = None) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results.
+
+ Args:
+ roi (Tensor): Boxes to be transformed. Has shape (num_boxes, 5).
+ last dimension 5 arrange as (batch_index, x1, y1, x2, y2).
+ cls_score (Tensor): Box scores, has shape
+ (num_boxes, num_classes + 1).
+ bbox_pred (Tuple[Tensor, Tensor]): Box cls preds and offset preds.
+ img_meta (dict): image information.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ rcnn_test_cfg (obj:`ConfigDict`): `test_cfg` of Bbox Head.
+ Defaults to None
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ results = InstanceData()
+ if isinstance(cls_score, list):
+ cls_score = sum(cls_score) / float(len(cls_score))
+ scores = F.softmax(cls_score, dim=1) if cls_score is not None else None
+ img_shape = img_meta['img_shape']
+ if bbox_pred is not None:
+ bboxes, confidences = self.bbox_coder.decode(
+ roi[:, 1:], bbox_pred, img_shape)
+ else:
+ bboxes = roi[:, 1:].clone()
+ confidences = None
+ if img_shape is not None:
+ bboxes[:, [0, 2]].clamp_(min=0, max=img_shape[1] - 1)
+ bboxes[:, [1, 3]].clamp_(min=0, max=img_shape[0] - 1)
+
+ if rescale and bboxes.size(0) > 0:
+ assert img_meta.get('scale_factor') is not None
+ scale_factor = bboxes.new_tensor(img_meta['scale_factor']).repeat(
+ (1, 2))
+ bboxes = (bboxes.view(bboxes.size(0), -1, 4) / scale_factor).view(
+ bboxes.size()[0], -1)
+
+ if rcnn_test_cfg is None:
+ results.bboxes = bboxes
+ results.scores = scores
+ else:
+ det_bboxes, det_labels = multiclass_nms(
+ bboxes,
+ scores,
+ rcnn_test_cfg.score_thr,
+ rcnn_test_cfg.nms,
+ rcnn_test_cfg.max_per_img,
+ score_factors=confidences)
+ results.bboxes = det_bboxes[:, :4]
+ results.scores = det_bboxes[:, -1]
+ results.labels = det_labels
+ return results
+
+ def refine_bboxes(self, sampling_results: List[SamplingResult],
+ bbox_results: dict,
+ batch_img_metas: List[dict]) -> InstanceList:
+ """Refine bboxes during training.
+
+ Args:
+ sampling_results (List[:obj:`SamplingResult`]): Sampling results.
+ bbox_results (dict): Usually is a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `rois` (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+ - `bbox_targets` (tuple): Ground truth for proposals in a
+ single image. Containing the following list of Tensors:
+ (labels, label_weights, bbox_targets, bbox_weights)
+ batch_img_metas (List[dict]): List of image information.
+
+ Returns:
+ list[:obj:`InstanceData`]: Refined bboxes of each image.
+ """
+ pos_is_gts = [res.pos_is_gt for res in sampling_results]
+ # bbox_targets is a tuple
+ labels = bbox_results['bbox_targets'][0]
+ cls_scores = bbox_results['cls_score']
+ rois = bbox_results['rois']
+ bbox_preds = bbox_results['bbox_pred']
+
+ if cls_scores.numel() == 0:
+ return None
+
+ labels = torch.where(labels == self.num_classes,
+ cls_scores[:, :-1].argmax(1), labels)
+
+ img_ids = rois[:, 0].long().unique(sorted=True)
+ assert img_ids.numel() <= len(batch_img_metas)
+
+ results_list = []
+ for i in range(len(batch_img_metas)):
+ inds = torch.nonzero(
+ rois[:, 0] == i, as_tuple=False).squeeze(dim=1)
+ num_rois = inds.numel()
+
+ bboxes_ = rois[inds, 1:]
+ label_ = labels[inds]
+ edge_cls_preds, edge_offset_preds = bbox_preds
+ edge_cls_preds_ = edge_cls_preds[inds]
+ edge_offset_preds_ = edge_offset_preds[inds]
+ bbox_pred_ = (edge_cls_preds_, edge_offset_preds_)
+ img_meta_ = batch_img_metas[i]
+ pos_is_gts_ = pos_is_gts[i]
+
+ bboxes = self.regress_by_class(bboxes_, label_, bbox_pred_,
+ img_meta_)
+ # filter gt bboxes
+ pos_keep = 1 - pos_is_gts_
+ keep_inds = pos_is_gts_.new_ones(num_rois)
+ keep_inds[:len(pos_is_gts_)] = pos_keep
+ results = InstanceData(bboxes=bboxes[keep_inds.type(torch.bool)])
+ results_list.append(results)
+
+ return results_list
+
+ def regress_by_class(self, rois: Tensor, label: Tensor, bbox_pred: tuple,
+ img_meta: dict) -> Tensor:
+ """Regress the bbox for the predicted class. Used in Cascade R-CNN.
+
+ Args:
+ rois (Tensor): shape (n, 4) or (n, 5)
+ label (Tensor): shape (n, )
+ bbox_pred (Tuple[Tensor]): shape [(n, num_buckets *2), \
+ (n, num_buckets *2)]
+ img_meta (dict): Image meta info.
+
+ Returns:
+ Tensor: Regressed bboxes, the same shape as input rois.
+ """
+ assert rois.size(1) == 4 or rois.size(1) == 5
+
+ if rois.size(1) == 4:
+ new_rois, _ = self.bbox_coder.decode(rois, bbox_pred,
+ img_meta['img_shape'])
+ else:
+ bboxes, _ = self.bbox_coder.decode(rois[:, 1:], bbox_pred,
+ img_meta['img_shape'])
+ new_rois = torch.cat((rois[:, [0]], bboxes), dim=1)
+
+ return new_rois
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/scnet_bbox_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/scnet_bbox_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..790b08fb207970927c7925cb8b3fb365bc183dc4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/bbox_heads/scnet_bbox_head.py
@@ -0,0 +1,101 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Tuple, Union
+
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from .convfc_bbox_head import ConvFCBBoxHead
+
+
+@MODELS.register_module()
+class SCNetBBoxHead(ConvFCBBoxHead):
+ """BBox head for `SCNet `_.
+
+ This inherits ``ConvFCBBoxHead`` with modified forward() function, allow us
+ to get intermediate shared feature.
+ """
+
+ def _forward_shared(self, x: Tensor) -> Tensor:
+ """Forward function for shared part.
+
+ Args:
+ x (Tensor): Input feature.
+
+ Returns:
+ Tensor: Shared feature.
+ """
+ if self.num_shared_convs > 0:
+ for conv in self.shared_convs:
+ x = conv(x)
+
+ if self.num_shared_fcs > 0:
+ if self.with_avg_pool:
+ x = self.avg_pool(x)
+
+ x = x.flatten(1)
+
+ for fc in self.shared_fcs:
+ x = self.relu(fc(x))
+
+ return x
+
+ def _forward_cls_reg(self, x: Tensor) -> Tuple[Tensor]:
+ """Forward function for classification and regression parts.
+
+ Args:
+ x (Tensor): Input feature.
+
+ Returns:
+ tuple[Tensor]:
+
+ - cls_score (Tensor): classification prediction.
+ - bbox_pred (Tensor): bbox prediction.
+ """
+ x_cls = x
+ x_reg = x
+
+ for conv in self.cls_convs:
+ x_cls = conv(x_cls)
+ if x_cls.dim() > 2:
+ if self.with_avg_pool:
+ x_cls = self.avg_pool(x_cls)
+ x_cls = x_cls.flatten(1)
+ for fc in self.cls_fcs:
+ x_cls = self.relu(fc(x_cls))
+
+ for conv in self.reg_convs:
+ x_reg = conv(x_reg)
+ if x_reg.dim() > 2:
+ if self.with_avg_pool:
+ x_reg = self.avg_pool(x_reg)
+ x_reg = x_reg.flatten(1)
+ for fc in self.reg_fcs:
+ x_reg = self.relu(fc(x_reg))
+
+ cls_score = self.fc_cls(x_cls) if self.with_cls else None
+ bbox_pred = self.fc_reg(x_reg) if self.with_reg else None
+
+ return cls_score, bbox_pred
+
+ def forward(
+ self,
+ x: Tensor,
+ return_shared_feat: bool = False) -> Union[Tensor, Tuple[Tensor]]:
+ """Forward function.
+
+ Args:
+ x (Tensor): input features
+ return_shared_feat (bool): If True, return cls-reg-shared feature.
+
+ Return:
+ out (tuple[Tensor]): contain ``cls_score`` and ``bbox_pred``,
+ if ``return_shared_feat`` is True, append ``x_shared`` to the
+ returned tuple.
+ """
+ x_shared = self._forward_shared(x)
+ out = self._forward_cls_reg(x_shared)
+
+ if return_shared_feat:
+ out += (x_shared, )
+
+ return out
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/cascade_roi_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/cascade_roi_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..81db671113a63beb7849abdc0e432a738ee46f5e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/cascade_roi_head.py
@@ -0,0 +1,568 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Sequence, Tuple, Union
+
+import torch
+import torch.nn as nn
+from mmengine.model import ModuleList
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.models.task_modules.samplers import SamplingResult
+from mmdet.models.test_time_augs import merge_aug_masks
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import bbox2roi, get_box_tensor
+from mmdet.utils import (ConfigType, InstanceList, MultiConfig, OptConfigType,
+ OptMultiConfig)
+from ..utils.misc import empty_instances, unpack_gt_instances
+from .base_roi_head import BaseRoIHead
+
+
+@MODELS.register_module()
+class CascadeRoIHead(BaseRoIHead):
+ """Cascade roi head including one bbox head and one mask head.
+
+ https://arxiv.org/abs/1712.00726
+ """
+
+ def __init__(self,
+ num_stages: int,
+ stage_loss_weights: Union[List[float], Tuple[float]],
+ bbox_roi_extractor: OptMultiConfig = None,
+ bbox_head: OptMultiConfig = None,
+ mask_roi_extractor: OptMultiConfig = None,
+ mask_head: OptMultiConfig = None,
+ shared_head: OptConfigType = None,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ init_cfg: OptMultiConfig = None) -> None:
+ assert bbox_roi_extractor is not None
+ assert bbox_head is not None
+ assert shared_head is None, \
+ 'Shared head is not supported in Cascade RCNN anymore'
+
+ self.num_stages = num_stages
+ self.stage_loss_weights = stage_loss_weights
+ super().__init__(
+ bbox_roi_extractor=bbox_roi_extractor,
+ bbox_head=bbox_head,
+ mask_roi_extractor=mask_roi_extractor,
+ mask_head=mask_head,
+ shared_head=shared_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ init_cfg=init_cfg)
+
+ def init_bbox_head(self, bbox_roi_extractor: MultiConfig,
+ bbox_head: MultiConfig) -> None:
+ """Initialize box head and box roi extractor.
+
+ Args:
+ bbox_roi_extractor (:obj:`ConfigDict`, dict or list):
+ Config of box roi extractor.
+ bbox_head (:obj:`ConfigDict`, dict or list): Config
+ of box in box head.
+ """
+ self.bbox_roi_extractor = ModuleList()
+ self.bbox_head = ModuleList()
+ if not isinstance(bbox_roi_extractor, list):
+ bbox_roi_extractor = [
+ bbox_roi_extractor for _ in range(self.num_stages)
+ ]
+ if not isinstance(bbox_head, list):
+ bbox_head = [bbox_head for _ in range(self.num_stages)]
+ assert len(bbox_roi_extractor) == len(bbox_head) == self.num_stages
+ for roi_extractor, head in zip(bbox_roi_extractor, bbox_head):
+ self.bbox_roi_extractor.append(MODELS.build(roi_extractor))
+ self.bbox_head.append(MODELS.build(head))
+
+ def init_mask_head(self, mask_roi_extractor: MultiConfig,
+ mask_head: MultiConfig) -> None:
+ """Initialize mask head and mask roi extractor.
+
+ Args:
+ mask_head (dict): Config of mask in mask head.
+ mask_roi_extractor (:obj:`ConfigDict`, dict or list):
+ Config of mask roi extractor.
+ """
+ self.mask_head = nn.ModuleList()
+ if not isinstance(mask_head, list):
+ mask_head = [mask_head for _ in range(self.num_stages)]
+ assert len(mask_head) == self.num_stages
+ for head in mask_head:
+ self.mask_head.append(MODELS.build(head))
+ if mask_roi_extractor is not None:
+ self.share_roi_extractor = False
+ self.mask_roi_extractor = ModuleList()
+ if not isinstance(mask_roi_extractor, list):
+ mask_roi_extractor = [
+ mask_roi_extractor for _ in range(self.num_stages)
+ ]
+ assert len(mask_roi_extractor) == self.num_stages
+ for roi_extractor in mask_roi_extractor:
+ self.mask_roi_extractor.append(MODELS.build(roi_extractor))
+ else:
+ self.share_roi_extractor = True
+ self.mask_roi_extractor = self.bbox_roi_extractor
+
+ def init_assigner_sampler(self) -> None:
+ """Initialize assigner and sampler for each stage."""
+ self.bbox_assigner = []
+ self.bbox_sampler = []
+ if self.train_cfg is not None:
+ for idx, rcnn_train_cfg in enumerate(self.train_cfg):
+ self.bbox_assigner.append(
+ TASK_UTILS.build(rcnn_train_cfg.assigner))
+ self.current_stage = idx
+ self.bbox_sampler.append(
+ TASK_UTILS.build(
+ rcnn_train_cfg.sampler,
+ default_args=dict(context=self)))
+
+ def _bbox_forward(self, stage: int, x: Tuple[Tensor],
+ rois: Tensor) -> dict:
+ """Box head forward function used in both training and testing.
+
+ Args:
+ stage (int): The current stage in Cascade RoI Head.
+ x (tuple[Tensor]): List of multi-level img features.
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+
+ Returns:
+ dict[str, Tensor]: Usually returns a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `bbox_feats` (Tensor): Extract bbox RoI features.
+ """
+ bbox_roi_extractor = self.bbox_roi_extractor[stage]
+ bbox_head = self.bbox_head[stage]
+ bbox_feats = bbox_roi_extractor(x[:bbox_roi_extractor.num_inputs],
+ rois)
+ # do not support caffe_c4 model anymore
+ cls_score, bbox_pred = bbox_head(bbox_feats)
+
+ bbox_results = dict(
+ cls_score=cls_score, bbox_pred=bbox_pred, bbox_feats=bbox_feats)
+ return bbox_results
+
+ def bbox_loss(self, stage: int, x: Tuple[Tensor],
+ sampling_results: List[SamplingResult]) -> dict:
+ """Run forward function and calculate loss for box head in training.
+
+ Args:
+ stage (int): The current stage in Cascade RoI Head.
+ x (tuple[Tensor]): List of multi-level img features.
+ sampling_results (list["obj:`SamplingResult`]): Sampling results.
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `bbox_feats` (Tensor): Extract bbox RoI features.
+ - `loss_bbox` (dict): A dictionary of bbox loss components.
+ - `rois` (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+ - `bbox_targets` (tuple): Ground truth for proposals in a
+ single image. Containing the following list of Tensors:
+ (labels, label_weights, bbox_targets, bbox_weights)
+ """
+ bbox_head = self.bbox_head[stage]
+ rois = bbox2roi([res.priors for res in sampling_results])
+ bbox_results = self._bbox_forward(stage, x, rois)
+ bbox_results.update(rois=rois)
+
+ bbox_loss_and_target = bbox_head.loss_and_target(
+ cls_score=bbox_results['cls_score'],
+ bbox_pred=bbox_results['bbox_pred'],
+ rois=rois,
+ sampling_results=sampling_results,
+ rcnn_train_cfg=self.train_cfg[stage])
+ bbox_results.update(bbox_loss_and_target)
+
+ return bbox_results
+
+ def _mask_forward(self, stage: int, x: Tuple[Tensor],
+ rois: Tensor) -> dict:
+ """Mask head forward function used in both training and testing.
+
+ Args:
+ stage (int): The current stage in Cascade RoI Head.
+ x (tuple[Tensor]): Tuple of multi-level img features.
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `mask_preds` (Tensor): Mask prediction.
+ """
+ mask_roi_extractor = self.mask_roi_extractor[stage]
+ mask_head = self.mask_head[stage]
+ mask_feats = mask_roi_extractor(x[:mask_roi_extractor.num_inputs],
+ rois)
+ # do not support caffe_c4 model anymore
+ mask_preds = mask_head(mask_feats)
+
+ mask_results = dict(mask_preds=mask_preds)
+ return mask_results
+
+ def mask_loss(self, stage: int, x: Tuple[Tensor],
+ sampling_results: List[SamplingResult],
+ batch_gt_instances: InstanceList) -> dict:
+ """Run forward function and calculate loss for mask head in training.
+
+ Args:
+ stage (int): The current stage in Cascade RoI Head.
+ x (tuple[Tensor]): Tuple of multi-level img features.
+ sampling_results (list["obj:`SamplingResult`]): Sampling results.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``labels``, and
+ ``masks`` attributes.
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `mask_preds` (Tensor): Mask prediction.
+ - `loss_mask` (dict): A dictionary of mask loss components.
+ """
+ pos_rois = bbox2roi([res.pos_priors for res in sampling_results])
+ mask_results = self._mask_forward(stage, x, pos_rois)
+
+ mask_head = self.mask_head[stage]
+
+ mask_loss_and_target = mask_head.loss_and_target(
+ mask_preds=mask_results['mask_preds'],
+ sampling_results=sampling_results,
+ batch_gt_instances=batch_gt_instances,
+ rcnn_train_cfg=self.train_cfg[stage])
+ mask_results.update(mask_loss_and_target)
+
+ return mask_results
+
+ def loss(self, x: Tuple[Tensor], rpn_results_list: InstanceList,
+ batch_data_samples: SampleList) -> dict:
+ """Perform forward propagation and loss calculation of the detection
+ roi on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components
+ """
+ # TODO: May add a new function in baseroihead
+ assert len(rpn_results_list) == len(batch_data_samples)
+ outputs = unpack_gt_instances(batch_data_samples)
+ batch_gt_instances, batch_gt_instances_ignore, batch_img_metas \
+ = outputs
+
+ num_imgs = len(batch_data_samples)
+ losses = dict()
+ results_list = rpn_results_list
+ for stage in range(self.num_stages):
+ self.current_stage = stage
+
+ stage_loss_weight = self.stage_loss_weights[stage]
+
+ # assign gts and sample proposals
+ sampling_results = []
+ if self.with_bbox or self.with_mask:
+ bbox_assigner = self.bbox_assigner[stage]
+ bbox_sampler = self.bbox_sampler[stage]
+
+ for i in range(num_imgs):
+ results = results_list[i]
+ # rename rpn_results.bboxes to rpn_results.priors
+ results.priors = results.pop('bboxes')
+
+ assign_result = bbox_assigner.assign(
+ results, batch_gt_instances[i],
+ batch_gt_instances_ignore[i])
+
+ sampling_result = bbox_sampler.sample(
+ assign_result,
+ results,
+ batch_gt_instances[i],
+ feats=[lvl_feat[i][None] for lvl_feat in x])
+ sampling_results.append(sampling_result)
+
+ # bbox head forward and loss
+ bbox_results = self.bbox_loss(stage, x, sampling_results)
+
+ for name, value in bbox_results['loss_bbox'].items():
+ losses[f's{stage}.{name}'] = (
+ value * stage_loss_weight if 'loss' in name else value)
+
+ # mask head forward and loss
+ if self.with_mask:
+ mask_results = self.mask_loss(stage, x, sampling_results,
+ batch_gt_instances)
+ for name, value in mask_results['loss_mask'].items():
+ losses[f's{stage}.{name}'] = (
+ value * stage_loss_weight if 'loss' in name else value)
+
+ # refine bboxes
+ if stage < self.num_stages - 1:
+ bbox_head = self.bbox_head[stage]
+ with torch.no_grad():
+ results_list = bbox_head.refine_bboxes(
+ sampling_results, bbox_results, batch_img_metas)
+ # Empty proposal
+ if results_list is None:
+ break
+ return losses
+
+ def predict_bbox(self,
+ x: Tuple[Tensor],
+ batch_img_metas: List[dict],
+ rpn_results_list: InstanceList,
+ rcnn_test_cfg: ConfigType,
+ rescale: bool = False,
+ **kwargs) -> InstanceList:
+ """Perform forward propagation of the bbox head and predict detection
+ results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Feature maps of all scale level.
+ batch_img_metas (list[dict]): List of image information.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ rcnn_test_cfg (obj:`ConfigDict`): `test_cfg` of R-CNN.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ proposals = [res.bboxes for res in rpn_results_list]
+ num_proposals_per_img = tuple(len(p) for p in proposals)
+ rois = bbox2roi(proposals)
+
+ if rois.shape[0] == 0:
+ return empty_instances(
+ batch_img_metas,
+ rois.device,
+ task_type='bbox',
+ box_type=self.bbox_head[-1].predict_box_type,
+ num_classes=self.bbox_head[-1].num_classes,
+ score_per_cls=rcnn_test_cfg is None)
+
+ rois, cls_scores, bbox_preds = self._refine_roi(
+ x=x,
+ rois=rois,
+ batch_img_metas=batch_img_metas,
+ num_proposals_per_img=num_proposals_per_img,
+ **kwargs)
+
+ results_list = self.bbox_head[-1].predict_by_feat(
+ rois=rois,
+ cls_scores=cls_scores,
+ bbox_preds=bbox_preds,
+ batch_img_metas=batch_img_metas,
+ rescale=rescale,
+ rcnn_test_cfg=rcnn_test_cfg)
+ return results_list
+
+ def predict_mask(self,
+ x: Tuple[Tensor],
+ batch_img_metas: List[dict],
+ results_list: List[InstanceData],
+ rescale: bool = False) -> List[InstanceData]:
+ """Perform forward propagation of the mask head and predict detection
+ results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Feature maps of all scale level.
+ batch_img_metas (list[dict]): List of image information.
+ results_list (list[:obj:`InstanceData`]): Detection results of
+ each image.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+ """
+ bboxes = [res.bboxes for res in results_list]
+ mask_rois = bbox2roi(bboxes)
+ if mask_rois.shape[0] == 0:
+ results_list = empty_instances(
+ batch_img_metas,
+ mask_rois.device,
+ task_type='mask',
+ instance_results=results_list,
+ mask_thr_binary=self.test_cfg.mask_thr_binary)
+ return results_list
+
+ num_mask_rois_per_img = [len(res) for res in results_list]
+ aug_masks = []
+ for stage in range(self.num_stages):
+ mask_results = self._mask_forward(stage, x, mask_rois)
+ mask_preds = mask_results['mask_preds']
+ # split batch mask prediction back to each image
+ mask_preds = mask_preds.split(num_mask_rois_per_img, 0)
+ aug_masks.append([m.sigmoid().detach() for m in mask_preds])
+
+ merged_masks = []
+ for i in range(len(batch_img_metas)):
+ aug_mask = [mask[i] for mask in aug_masks]
+ merged_mask = merge_aug_masks(aug_mask, batch_img_metas[i])
+ merged_masks.append(merged_mask)
+ results_list = self.mask_head[-1].predict_by_feat(
+ mask_preds=merged_masks,
+ results_list=results_list,
+ batch_img_metas=batch_img_metas,
+ rcnn_test_cfg=self.test_cfg,
+ rescale=rescale,
+ activate_map=True)
+ return results_list
+
+ def _refine_roi(self, x: Tuple[Tensor], rois: Tensor,
+ batch_img_metas: List[dict],
+ num_proposals_per_img: Sequence[int], **kwargs) -> tuple:
+ """Multi-stage refinement of RoI.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ rois (Tensor): shape (n, 5), [batch_ind, x1, y1, x2, y2]
+ batch_img_metas (list[dict]): List of image information.
+ num_proposals_per_img (sequence[int]): number of proposals
+ in each image.
+
+ Returns:
+ tuple:
+
+ - rois (Tensor): Refined RoI.
+ - cls_scores (list[Tensor]): Average predicted
+ cls score per image.
+ - bbox_preds (list[Tensor]): Bbox branch predictions
+ for the last stage of per image.
+ """
+ # "ms" in variable names means multi-stage
+ ms_scores = []
+ for stage in range(self.num_stages):
+ bbox_results = self._bbox_forward(
+ stage=stage, x=x, rois=rois, **kwargs)
+
+ # split batch bbox prediction back to each image
+ cls_scores = bbox_results['cls_score']
+ bbox_preds = bbox_results['bbox_pred']
+
+ rois = rois.split(num_proposals_per_img, 0)
+ cls_scores = cls_scores.split(num_proposals_per_img, 0)
+ ms_scores.append(cls_scores)
+
+ # some detector with_reg is False, bbox_preds will be None
+ if bbox_preds is not None:
+ # TODO move this to a sabl_roi_head
+ # the bbox prediction of some detectors like SABL is not Tensor
+ if isinstance(bbox_preds, torch.Tensor):
+ bbox_preds = bbox_preds.split(num_proposals_per_img, 0)
+ else:
+ bbox_preds = self.bbox_head[stage].bbox_pred_split(
+ bbox_preds, num_proposals_per_img)
+ else:
+ bbox_preds = (None, ) * len(batch_img_metas)
+
+ if stage < self.num_stages - 1:
+ bbox_head = self.bbox_head[stage]
+ if bbox_head.custom_activation:
+ cls_scores = [
+ bbox_head.loss_cls.get_activation(s)
+ for s in cls_scores
+ ]
+ refine_rois_list = []
+ for i in range(len(batch_img_metas)):
+ if rois[i].shape[0] > 0:
+ bbox_label = cls_scores[i][:, :-1].argmax(dim=1)
+ # Refactor `bbox_head.regress_by_class` to only accept
+ # box tensor without img_idx concatenated.
+ refined_bboxes = bbox_head.regress_by_class(
+ rois[i][:, 1:], bbox_label, bbox_preds[i],
+ batch_img_metas[i])
+ refined_bboxes = get_box_tensor(refined_bboxes)
+ refined_rois = torch.cat(
+ [rois[i][:, [0]], refined_bboxes], dim=1)
+ refine_rois_list.append(refined_rois)
+ rois = torch.cat(refine_rois_list)
+
+ # average scores of each image by stages
+ cls_scores = [
+ sum([score[i] for score in ms_scores]) / float(len(ms_scores))
+ for i in range(len(batch_img_metas))
+ ]
+ return rois, cls_scores, bbox_preds
+
+ def forward(self, x: Tuple[Tensor], rpn_results_list: InstanceList,
+ batch_data_samples: SampleList) -> tuple:
+ """Network forward process. Usually includes backbone, neck and head
+ forward without any post-processing.
+
+ Args:
+ x (List[Tensor]): Multi-level features that may have different
+ resolutions.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ batch_data_samples (list[:obj:`DetDataSample`]): Each item contains
+ the meta information of each image and corresponding
+ annotations.
+
+ Returns
+ tuple: A tuple of features from ``bbox_head`` and ``mask_head``
+ forward.
+ """
+ results = ()
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+ proposals = [rpn_results.bboxes for rpn_results in rpn_results_list]
+ num_proposals_per_img = tuple(len(p) for p in proposals)
+ rois = bbox2roi(proposals)
+ # bbox head
+ if self.with_bbox:
+ rois, cls_scores, bbox_preds = self._refine_roi(
+ x, rois, batch_img_metas, num_proposals_per_img)
+ results = results + (cls_scores, bbox_preds)
+ # mask head
+ if self.with_mask:
+ aug_masks = []
+ rois = torch.cat(rois)
+ for stage in range(self.num_stages):
+ mask_results = self._mask_forward(stage, x, rois)
+ mask_preds = mask_results['mask_preds']
+ mask_preds = mask_preds.split(num_proposals_per_img, 0)
+ aug_masks.append([m.sigmoid().detach() for m in mask_preds])
+
+ merged_masks = []
+ for i in range(len(batch_img_metas)):
+ aug_mask = [mask[i] for mask in aug_masks]
+ merged_mask = merge_aug_masks(aug_mask, batch_img_metas[i])
+ merged_masks.append(merged_mask)
+ results = results + (merged_masks, )
+ return results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/double_roi_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/double_roi_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..f9464ff55bafcca9f3545a3a72dde1eb3939cece
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/double_roi_head.py
@@ -0,0 +1,53 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Tuple
+
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from .standard_roi_head import StandardRoIHead
+
+
+@MODELS.register_module()
+class DoubleHeadRoIHead(StandardRoIHead):
+ """RoI head for `Double Head RCNN `_.
+
+ Args:
+ reg_roi_scale_factor (float): The scale factor to extend the rois
+ used to extract the regression features.
+ """
+
+ def __init__(self, reg_roi_scale_factor: float, **kwargs):
+ super().__init__(**kwargs)
+ self.reg_roi_scale_factor = reg_roi_scale_factor
+
+ def _bbox_forward(self, x: Tuple[Tensor], rois: Tensor) -> dict:
+ """Box head forward function used in both training and testing.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+
+ Returns:
+ dict[str, Tensor]: Usually returns a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `bbox_feats` (Tensor): Extract bbox RoI features.
+ """
+ bbox_cls_feats = self.bbox_roi_extractor(
+ x[:self.bbox_roi_extractor.num_inputs], rois)
+ bbox_reg_feats = self.bbox_roi_extractor(
+ x[:self.bbox_roi_extractor.num_inputs],
+ rois,
+ roi_scale_factor=self.reg_roi_scale_factor)
+ if self.with_shared_head:
+ bbox_cls_feats = self.shared_head(bbox_cls_feats)
+ bbox_reg_feats = self.shared_head(bbox_reg_feats)
+ cls_score, bbox_pred = self.bbox_head(bbox_cls_feats, bbox_reg_feats)
+
+ bbox_results = dict(
+ cls_score=cls_score,
+ bbox_pred=bbox_pred,
+ bbox_feats=bbox_cls_feats)
+ return bbox_results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/dynamic_roi_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/dynamic_roi_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..3c7f7bd2f68cab0fcdec725501f74b65274eb30e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/dynamic_roi_head.py
@@ -0,0 +1,163 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple
+
+import numpy as np
+import torch
+from torch import Tensor
+
+from mmdet.models.losses import SmoothL1Loss
+from mmdet.models.task_modules.samplers import SamplingResult
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import bbox2roi
+from mmdet.utils import InstanceList
+from ..utils.misc import unpack_gt_instances
+from .standard_roi_head import StandardRoIHead
+
+EPS = 1e-15
+
+
+@MODELS.register_module()
+class DynamicRoIHead(StandardRoIHead):
+ """RoI head for `Dynamic R-CNN `_."""
+
+ def __init__(self, **kwargs) -> None:
+ super().__init__(**kwargs)
+ assert isinstance(self.bbox_head.loss_bbox, SmoothL1Loss)
+ # the IoU history of the past `update_iter_interval` iterations
+ self.iou_history = []
+ # the beta history of the past `update_iter_interval` iterations
+ self.beta_history = []
+
+ def loss(self, x: Tuple[Tensor], rpn_results_list: InstanceList,
+ batch_data_samples: SampleList) -> dict:
+ """Forward function for training.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict[str, Tensor]: a dictionary of loss components
+ """
+ assert len(rpn_results_list) == len(batch_data_samples)
+ outputs = unpack_gt_instances(batch_data_samples)
+ batch_gt_instances, batch_gt_instances_ignore, _ = outputs
+
+ # assign gts and sample proposals
+ num_imgs = len(batch_data_samples)
+ sampling_results = []
+ cur_iou = []
+ for i in range(num_imgs):
+ # rename rpn_results.bboxes to rpn_results.priors
+ rpn_results = rpn_results_list[i]
+ rpn_results.priors = rpn_results.pop('bboxes')
+
+ assign_result = self.bbox_assigner.assign(
+ rpn_results, batch_gt_instances[i],
+ batch_gt_instances_ignore[i])
+ sampling_result = self.bbox_sampler.sample(
+ assign_result,
+ rpn_results,
+ batch_gt_instances[i],
+ feats=[lvl_feat[i][None] for lvl_feat in x])
+ # record the `iou_topk`-th largest IoU in an image
+ iou_topk = min(self.train_cfg.dynamic_rcnn.iou_topk,
+ len(assign_result.max_overlaps))
+ ious, _ = torch.topk(assign_result.max_overlaps, iou_topk)
+ cur_iou.append(ious[-1].item())
+ sampling_results.append(sampling_result)
+ # average the current IoUs over images
+ cur_iou = np.mean(cur_iou)
+ self.iou_history.append(cur_iou)
+
+ losses = dict()
+ # bbox head forward and loss
+ if self.with_bbox:
+ bbox_results = self.bbox_loss(x, sampling_results)
+ losses.update(bbox_results['loss_bbox'])
+
+ # mask head forward and loss
+ if self.with_mask:
+ mask_results = self.mask_loss(x, sampling_results,
+ bbox_results['bbox_feats'],
+ batch_gt_instances)
+ losses.update(mask_results['loss_mask'])
+
+ # update IoU threshold and SmoothL1 beta
+ update_iter_interval = self.train_cfg.dynamic_rcnn.update_iter_interval
+ if len(self.iou_history) % update_iter_interval == 0:
+ new_iou_thr, new_beta = self.update_hyperparameters()
+
+ return losses
+
+ def bbox_loss(self, x: Tuple[Tensor],
+ sampling_results: List[SamplingResult]) -> dict:
+ """Perform forward propagation and loss calculation of the bbox head on
+ the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ sampling_results (list["obj:`SamplingResult`]): Sampling results.
+
+ Returns:
+ dict[str, Tensor]: Usually returns a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `bbox_feats` (Tensor): Extract bbox RoI features.
+ - `loss_bbox` (dict): A dictionary of bbox loss components.
+ """
+ rois = bbox2roi([res.priors for res in sampling_results])
+ bbox_results = self._bbox_forward(x, rois)
+
+ bbox_loss_and_target = self.bbox_head.loss_and_target(
+ cls_score=bbox_results['cls_score'],
+ bbox_pred=bbox_results['bbox_pred'],
+ rois=rois,
+ sampling_results=sampling_results,
+ rcnn_train_cfg=self.train_cfg)
+ bbox_results.update(loss_bbox=bbox_loss_and_target['loss_bbox'])
+
+ # record the `beta_topk`-th smallest target
+ # `bbox_targets[2]` and `bbox_targets[3]` stand for bbox_targets
+ # and bbox_weights, respectively
+ bbox_targets = bbox_loss_and_target['bbox_targets']
+ pos_inds = bbox_targets[3][:, 0].nonzero().squeeze(1)
+ num_pos = len(pos_inds)
+ num_imgs = len(sampling_results)
+ if num_pos > 0:
+ cur_target = bbox_targets[2][pos_inds, :2].abs().mean(dim=1)
+ beta_topk = min(self.train_cfg.dynamic_rcnn.beta_topk * num_imgs,
+ num_pos)
+ cur_target = torch.kthvalue(cur_target, beta_topk)[0].item()
+ self.beta_history.append(cur_target)
+
+ return bbox_results
+
+ def update_hyperparameters(self):
+ """Update hyperparameters like IoU thresholds for assigner and beta for
+ SmoothL1 loss based on the training statistics.
+
+ Returns:
+ tuple[float]: the updated ``iou_thr`` and ``beta``.
+ """
+ new_iou_thr = max(self.train_cfg.dynamic_rcnn.initial_iou,
+ np.mean(self.iou_history))
+ self.iou_history = []
+ self.bbox_assigner.pos_iou_thr = new_iou_thr
+ self.bbox_assigner.neg_iou_thr = new_iou_thr
+ self.bbox_assigner.min_pos_iou = new_iou_thr
+ if (not self.beta_history) or (np.median(self.beta_history) < EPS):
+ # avoid 0 or too small value for new_beta
+ new_beta = self.bbox_head.loss_bbox.beta
+ else:
+ new_beta = min(self.train_cfg.dynamic_rcnn.initial_beta,
+ np.median(self.beta_history))
+ self.beta_history = []
+ self.bbox_head.loss_bbox.beta = new_beta
+ return new_iou_thr, new_beta
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/grid_roi_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/grid_roi_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..9eda7f01bcd4e44faca14b61ec4956ee2c372ad6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/grid_roi_head.py
@@ -0,0 +1,280 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple
+
+import torch
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import bbox2roi
+from mmdet.utils import ConfigType, InstanceList
+from ..task_modules.samplers import SamplingResult
+from ..utils.misc import unpack_gt_instances
+from .standard_roi_head import StandardRoIHead
+
+
+@MODELS.register_module()
+class GridRoIHead(StandardRoIHead):
+ """Implementation of `Grid RoI Head `_
+
+ Args:
+ grid_roi_extractor (:obj:`ConfigDict` or dict): Config of
+ roi extractor.
+ grid_head (:obj:`ConfigDict` or dict): Config of grid head
+ """
+
+ def __init__(self, grid_roi_extractor: ConfigType, grid_head: ConfigType,
+ **kwargs) -> None:
+ assert grid_head is not None
+ super().__init__(**kwargs)
+ if grid_roi_extractor is not None:
+ self.grid_roi_extractor = MODELS.build(grid_roi_extractor)
+ self.share_roi_extractor = False
+ else:
+ self.share_roi_extractor = True
+ self.grid_roi_extractor = self.bbox_roi_extractor
+ self.grid_head = MODELS.build(grid_head)
+
+ def _random_jitter(self,
+ sampling_results: List[SamplingResult],
+ batch_img_metas: List[dict],
+ amplitude: float = 0.15) -> List[SamplingResult]:
+ """Ramdom jitter positive proposals for training.
+
+ Args:
+ sampling_results (List[obj:SamplingResult]): Assign results of
+ all images in a batch after sampling.
+ batch_img_metas (list[dict]): List of image information.
+ amplitude (float): Amplitude of random offset. Defaults to 0.15.
+
+ Returns:
+ list[obj:SamplingResult]: SamplingResults after random jittering.
+ """
+ for sampling_result, img_meta in zip(sampling_results,
+ batch_img_metas):
+ bboxes = sampling_result.pos_priors
+ random_offsets = bboxes.new_empty(bboxes.shape[0], 4).uniform_(
+ -amplitude, amplitude)
+ # before jittering
+ cxcy = (bboxes[:, 2:4] + bboxes[:, :2]) / 2
+ wh = (bboxes[:, 2:4] - bboxes[:, :2]).abs()
+ # after jittering
+ new_cxcy = cxcy + wh * random_offsets[:, :2]
+ new_wh = wh * (1 + random_offsets[:, 2:])
+ # xywh to xyxy
+ new_x1y1 = (new_cxcy - new_wh / 2)
+ new_x2y2 = (new_cxcy + new_wh / 2)
+ new_bboxes = torch.cat([new_x1y1, new_x2y2], dim=1)
+ # clip bboxes
+ max_shape = img_meta['img_shape']
+ if max_shape is not None:
+ new_bboxes[:, 0::2].clamp_(min=0, max=max_shape[1] - 1)
+ new_bboxes[:, 1::2].clamp_(min=0, max=max_shape[0] - 1)
+
+ sampling_result.pos_priors = new_bboxes
+ return sampling_results
+
+ # TODO: Forward is incorrect and need to refactor.
+ def forward(self,
+ x: Tuple[Tensor],
+ rpn_results_list: InstanceList,
+ batch_data_samples: SampleList = None) -> tuple:
+ """Network forward process. Usually includes backbone, neck and head
+ forward without any post-processing.
+
+ Args:
+ x (Tuple[Tensor]): Multi-level features that may have different
+ resolutions.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ batch_data_samples (list[:obj:`DetDataSample`]): Each item contains
+ the meta information of each image and corresponding
+ annotations.
+
+ Returns
+ tuple: A tuple of features from ``bbox_head`` and ``mask_head``
+ forward.
+ """
+ results = ()
+ proposals = [rpn_results.bboxes for rpn_results in rpn_results_list]
+ rois = bbox2roi(proposals)
+ # bbox head
+ if self.with_bbox:
+ bbox_results = self._bbox_forward(x, rois)
+ results = results + (bbox_results['cls_score'], )
+ if self.bbox_head.with_reg:
+ results = results + (bbox_results['bbox_pred'], )
+
+ # grid head
+ grid_rois = rois[:100]
+ grid_feats = self.grid_roi_extractor(
+ x[:len(self.grid_roi_extractor.featmap_strides)], grid_rois)
+ if self.with_shared_head:
+ grid_feats = self.shared_head(grid_feats)
+ self.grid_head.test_mode = True
+ grid_preds = self.grid_head(grid_feats)
+ results = results + (grid_preds, )
+
+ # mask head
+ if self.with_mask:
+ mask_rois = rois[:100]
+ mask_results = self._mask_forward(x, mask_rois)
+ results = results + (mask_results['mask_preds'], )
+ return results
+
+ def loss(self, x: Tuple[Tensor], rpn_results_list: InstanceList,
+ batch_data_samples: SampleList, **kwargs) -> dict:
+ """Perform forward propagation and loss calculation of the detection
+ roi on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components
+ """
+ assert len(rpn_results_list) == len(batch_data_samples)
+ outputs = unpack_gt_instances(batch_data_samples)
+ (batch_gt_instances, batch_gt_instances_ignore,
+ batch_img_metas) = outputs
+
+ # assign gts and sample proposals
+ num_imgs = len(batch_data_samples)
+ sampling_results = []
+ for i in range(num_imgs):
+ # rename rpn_results.bboxes to rpn_results.priors
+ rpn_results = rpn_results_list[i]
+ rpn_results.priors = rpn_results.pop('bboxes')
+
+ assign_result = self.bbox_assigner.assign(
+ rpn_results, batch_gt_instances[i],
+ batch_gt_instances_ignore[i])
+ sampling_result = self.bbox_sampler.sample(
+ assign_result,
+ rpn_results,
+ batch_gt_instances[i],
+ feats=[lvl_feat[i][None] for lvl_feat in x])
+ sampling_results.append(sampling_result)
+
+ losses = dict()
+ # bbox head loss
+ if self.with_bbox:
+ bbox_results = self.bbox_loss(x, sampling_results, batch_img_metas)
+ losses.update(bbox_results['loss_bbox'])
+
+ # mask head forward and loss
+ if self.with_mask:
+ mask_results = self.mask_loss(x, sampling_results,
+ bbox_results['bbox_feats'],
+ batch_gt_instances)
+ losses.update(mask_results['loss_mask'])
+
+ return losses
+
+ def bbox_loss(self,
+ x: Tuple[Tensor],
+ sampling_results: List[SamplingResult],
+ batch_img_metas: Optional[List[dict]] = None) -> dict:
+ """Perform forward propagation and loss calculation of the bbox head on
+ the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ sampling_results (list[:obj:`SamplingResult`]): Sampling results.
+ batch_img_metas (list[dict], optional): Meta information of each
+ image, e.g., image size, scaling factor, etc.
+
+ Returns:
+ dict[str, Tensor]: Usually returns a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `bbox_feats` (Tensor): Extract bbox RoI features.
+ - `loss_bbox` (dict): A dictionary of bbox loss components.
+ """
+ assert batch_img_metas is not None
+ bbox_results = super().bbox_loss(x, sampling_results)
+
+ # Grid head forward and loss
+ sampling_results = self._random_jitter(sampling_results,
+ batch_img_metas)
+ pos_rois = bbox2roi([res.pos_bboxes for res in sampling_results])
+
+ # GN in head does not support zero shape input
+ if pos_rois.shape[0] == 0:
+ return bbox_results
+
+ grid_feats = self.grid_roi_extractor(
+ x[:self.grid_roi_extractor.num_inputs], pos_rois)
+ if self.with_shared_head:
+ grid_feats = self.shared_head(grid_feats)
+ # Accelerate training
+ max_sample_num_grid = self.train_cfg.get('max_num_grid', 192)
+ sample_idx = torch.randperm(
+ grid_feats.shape[0])[:min(grid_feats.shape[0], max_sample_num_grid
+ )]
+ grid_feats = grid_feats[sample_idx]
+ grid_pred = self.grid_head(grid_feats)
+
+ loss_grid = self.grid_head.loss(grid_pred, sample_idx,
+ sampling_results, self.train_cfg)
+
+ bbox_results['loss_bbox'].update(loss_grid)
+ return bbox_results
+
+ def predict_bbox(self,
+ x: Tuple[Tensor],
+ batch_img_metas: List[dict],
+ rpn_results_list: InstanceList,
+ rcnn_test_cfg: ConfigType,
+ rescale: bool = False) -> InstanceList:
+ """Perform forward propagation of the bbox head and predict detection
+ results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Feature maps of all scale level.
+ batch_img_metas (list[dict]): List of image information.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ rcnn_test_cfg (:obj:`ConfigDict`): `test_cfg` of R-CNN.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape \
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4), the last \
+ dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ results_list = super().predict_bbox(
+ x,
+ batch_img_metas=batch_img_metas,
+ rpn_results_list=rpn_results_list,
+ rcnn_test_cfg=rcnn_test_cfg,
+ rescale=False)
+
+ grid_rois = bbox2roi([res.bboxes for res in results_list])
+ if grid_rois.shape[0] != 0:
+ grid_feats = self.grid_roi_extractor(
+ x[:len(self.grid_roi_extractor.featmap_strides)], grid_rois)
+ if self.with_shared_head:
+ grid_feats = self.shared_head(grid_feats)
+ self.grid_head.test_mode = True
+ grid_preds = self.grid_head(grid_feats)
+ results_list = self.grid_head.predict_by_feat(
+ grid_preds=grid_preds,
+ results_list=results_list,
+ batch_img_metas=batch_img_metas,
+ rescale=rescale)
+
+ return results_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/htc_roi_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/htc_roi_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..0fdd99ddd5ce4d9d42345d1f1d14ecbcae658124
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/htc_roi_head.py
@@ -0,0 +1,581 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, List, Optional, Tuple
+
+import torch
+import torch.nn.functional as F
+from torch import Tensor
+
+from mmdet.models.test_time_augs import merge_aug_masks
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import bbox2roi
+from mmdet.utils import InstanceList, OptConfigType
+from ..layers import adaptive_avg_pool2d
+from ..task_modules.samplers import SamplingResult
+from ..utils import empty_instances, unpack_gt_instances
+from .cascade_roi_head import CascadeRoIHead
+
+
+@MODELS.register_module()
+class HybridTaskCascadeRoIHead(CascadeRoIHead):
+ """Hybrid task cascade roi head including one bbox head and one mask head.
+
+ https://arxiv.org/abs/1901.07518
+
+ Args:
+ num_stages (int): Number of cascade stages.
+ stage_loss_weights (list[float]): Loss weight for every stage.
+ semantic_roi_extractor (:obj:`ConfigDict` or dict, optional):
+ Config of semantic roi extractor. Defaults to None.
+ Semantic_head (:obj:`ConfigDict` or dict, optional):
+ Config of semantic head. Defaults to None.
+ interleaved (bool): Whether to interleaves the box branch and mask
+ branch. If True, the mask branch can take the refined bounding
+ box predictions. Defaults to True.
+ mask_info_flow (bool): Whether to turn on the mask information flow,
+ which means that feeding the mask features of the preceding stage
+ to the current stage. Defaults to True.
+ """
+
+ def __init__(self,
+ num_stages: int,
+ stage_loss_weights: List[float],
+ semantic_roi_extractor: OptConfigType = None,
+ semantic_head: OptConfigType = None,
+ semantic_fusion: Tuple[str] = ('bbox', 'mask'),
+ interleaved: bool = True,
+ mask_info_flow: bool = True,
+ **kwargs) -> None:
+ super().__init__(
+ num_stages=num_stages,
+ stage_loss_weights=stage_loss_weights,
+ **kwargs)
+ assert self.with_bbox
+ assert not self.with_shared_head # shared head is not supported
+
+ if semantic_head is not None:
+ self.semantic_roi_extractor = MODELS.build(semantic_roi_extractor)
+ self.semantic_head = MODELS.build(semantic_head)
+
+ self.semantic_fusion = semantic_fusion
+ self.interleaved = interleaved
+ self.mask_info_flow = mask_info_flow
+
+ # TODO move to base_roi_head later
+ @property
+ def with_semantic(self) -> bool:
+ """bool: whether the head has semantic head"""
+ return hasattr(self,
+ 'semantic_head') and self.semantic_head is not None
+
+ def _bbox_forward(
+ self,
+ stage: int,
+ x: Tuple[Tensor],
+ rois: Tensor,
+ semantic_feat: Optional[Tensor] = None) -> Dict[str, Tensor]:
+ """Box head forward function used in both training and testing.
+
+ Args:
+ stage (int): The current stage in Cascade RoI Head.
+ x (tuple[Tensor]): List of multi-level img features.
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+ semantic_feat (Tensor, optional): Semantic feature. Defaults to
+ None.
+
+ Returns:
+ dict[str, Tensor]: Usually returns a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `bbox_feats` (Tensor): Extract bbox RoI features.
+ """
+ bbox_roi_extractor = self.bbox_roi_extractor[stage]
+ bbox_head = self.bbox_head[stage]
+ bbox_feats = bbox_roi_extractor(x[:bbox_roi_extractor.num_inputs],
+ rois)
+ if self.with_semantic and 'bbox' in self.semantic_fusion:
+ bbox_semantic_feat = self.semantic_roi_extractor([semantic_feat],
+ rois)
+ if bbox_semantic_feat.shape[-2:] != bbox_feats.shape[-2:]:
+ bbox_semantic_feat = adaptive_avg_pool2d(
+ bbox_semantic_feat, bbox_feats.shape[-2:])
+ bbox_feats += bbox_semantic_feat
+ cls_score, bbox_pred = bbox_head(bbox_feats)
+
+ bbox_results = dict(cls_score=cls_score, bbox_pred=bbox_pred)
+ return bbox_results
+
+ def bbox_loss(self,
+ stage: int,
+ x: Tuple[Tensor],
+ sampling_results: List[SamplingResult],
+ semantic_feat: Optional[Tensor] = None) -> dict:
+ """Run forward function and calculate loss for box head in training.
+
+ Args:
+ stage (int): The current stage in Cascade RoI Head.
+ x (tuple[Tensor]): List of multi-level img features.
+ sampling_results (list["obj:`SamplingResult`]): Sampling results.
+ semantic_feat (Tensor, optional): Semantic feature. Defaults to
+ None.
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `bbox_feats` (Tensor): Extract bbox RoI features.
+ - `loss_bbox` (dict): A dictionary of bbox loss components.
+ - `rois` (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+ - `bbox_targets` (tuple): Ground truth for proposals in a
+ single image. Containing the following list of Tensors:
+ (labels, label_weights, bbox_targets, bbox_weights)
+ """
+ bbox_head = self.bbox_head[stage]
+ rois = bbox2roi([res.priors for res in sampling_results])
+ bbox_results = self._bbox_forward(
+ stage, x, rois, semantic_feat=semantic_feat)
+ bbox_results.update(rois=rois)
+
+ bbox_loss_and_target = bbox_head.loss_and_target(
+ cls_score=bbox_results['cls_score'],
+ bbox_pred=bbox_results['bbox_pred'],
+ rois=rois,
+ sampling_results=sampling_results,
+ rcnn_train_cfg=self.train_cfg[stage])
+ bbox_results.update(bbox_loss_and_target)
+ return bbox_results
+
+ def _mask_forward(self,
+ stage: int,
+ x: Tuple[Tensor],
+ rois: Tensor,
+ semantic_feat: Optional[Tensor] = None,
+ training: bool = True) -> Dict[str, Tensor]:
+ """Mask head forward function used only in training.
+
+ Args:
+ stage (int): The current stage in Cascade RoI Head.
+ x (tuple[Tensor]): Tuple of multi-level img features.
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+ semantic_feat (Tensor, optional): Semantic feature. Defaults to
+ None.
+ training (bool): Mask Forward is different between training and
+ testing. If True, use the mask forward in training.
+ Defaults to True.
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `mask_preds` (Tensor): Mask prediction.
+ """
+ mask_roi_extractor = self.mask_roi_extractor[stage]
+ mask_head = self.mask_head[stage]
+ mask_feats = mask_roi_extractor(x[:mask_roi_extractor.num_inputs],
+ rois)
+
+ # semantic feature fusion
+ # element-wise sum for original features and pooled semantic features
+ if self.with_semantic and 'mask' in self.semantic_fusion:
+ mask_semantic_feat = self.semantic_roi_extractor([semantic_feat],
+ rois)
+ if mask_semantic_feat.shape[-2:] != mask_feats.shape[-2:]:
+ mask_semantic_feat = F.adaptive_avg_pool2d(
+ mask_semantic_feat, mask_feats.shape[-2:])
+ mask_feats = mask_feats + mask_semantic_feat
+
+ # mask information flow
+ # forward all previous mask heads to obtain last_feat, and fuse it
+ # with the normal mask feature
+ if training:
+ if self.mask_info_flow:
+ last_feat = None
+ for i in range(stage):
+ last_feat = self.mask_head[i](
+ mask_feats, last_feat, return_logits=False)
+ mask_preds = mask_head(
+ mask_feats, last_feat, return_feat=False)
+ else:
+ mask_preds = mask_head(mask_feats, return_feat=False)
+
+ mask_results = dict(mask_preds=mask_preds)
+ else:
+ aug_masks = []
+ last_feat = None
+ for i in range(self.num_stages):
+ mask_head = self.mask_head[i]
+ if self.mask_info_flow:
+ mask_preds, last_feat = mask_head(mask_feats, last_feat)
+ else:
+ mask_preds = mask_head(mask_feats)
+ aug_masks.append(mask_preds)
+
+ mask_results = dict(mask_preds=aug_masks)
+
+ return mask_results
+
+ def mask_loss(self,
+ stage: int,
+ x: Tuple[Tensor],
+ sampling_results: List[SamplingResult],
+ batch_gt_instances: InstanceList,
+ semantic_feat: Optional[Tensor] = None) -> dict:
+ """Run forward function and calculate loss for mask head in training.
+
+ Args:
+ stage (int): The current stage in Cascade RoI Head.
+ x (tuple[Tensor]): Tuple of multi-level img features.
+ sampling_results (list["obj:`SamplingResult`]): Sampling results.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``labels``, and
+ ``masks`` attributes.
+ semantic_feat (Tensor, optional): Semantic feature. Defaults to
+ None.
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `mask_preds` (Tensor): Mask prediction.
+ - `loss_mask` (dict): A dictionary of mask loss components.
+ """
+ pos_rois = bbox2roi([res.pos_priors for res in sampling_results])
+ mask_results = self._mask_forward(
+ stage=stage,
+ x=x,
+ rois=pos_rois,
+ semantic_feat=semantic_feat,
+ training=True)
+
+ mask_head = self.mask_head[stage]
+ mask_loss_and_target = mask_head.loss_and_target(
+ mask_preds=mask_results['mask_preds'],
+ sampling_results=sampling_results,
+ batch_gt_instances=batch_gt_instances,
+ rcnn_train_cfg=self.train_cfg[stage])
+ mask_results.update(mask_loss_and_target)
+
+ return mask_results
+
+ def loss(self, x: Tuple[Tensor], rpn_results_list: InstanceList,
+ batch_data_samples: SampleList) -> dict:
+ """Perform forward propagation and loss calculation of the detection
+ roi on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components
+ """
+ assert len(rpn_results_list) == len(batch_data_samples)
+ outputs = unpack_gt_instances(batch_data_samples)
+ batch_gt_instances, batch_gt_instances_ignore, batch_img_metas \
+ = outputs
+
+ # semantic segmentation part
+ # 2 outputs: segmentation prediction and embedded features
+ losses = dict()
+ if self.with_semantic:
+ gt_semantic_segs = [
+ data_sample.gt_sem_seg.sem_seg
+ for data_sample in batch_data_samples
+ ]
+ gt_semantic_segs = torch.stack(gt_semantic_segs)
+ semantic_pred, semantic_feat = self.semantic_head(x)
+ loss_seg = self.semantic_head.loss(semantic_pred, gt_semantic_segs)
+ losses['loss_semantic_seg'] = loss_seg
+ else:
+ semantic_feat = None
+
+ results_list = rpn_results_list
+ num_imgs = len(batch_img_metas)
+ for stage in range(self.num_stages):
+ self.current_stage = stage
+
+ stage_loss_weight = self.stage_loss_weights[stage]
+
+ # assign gts and sample proposals
+ sampling_results = []
+ bbox_assigner = self.bbox_assigner[stage]
+ bbox_sampler = self.bbox_sampler[stage]
+ for i in range(num_imgs):
+ results = results_list[i]
+ # rename rpn_results.bboxes to rpn_results.priors
+ if 'bboxes' in results:
+ results.priors = results.pop('bboxes')
+
+ assign_result = bbox_assigner.assign(
+ results, batch_gt_instances[i],
+ batch_gt_instances_ignore[i])
+ sampling_result = bbox_sampler.sample(
+ assign_result,
+ results,
+ batch_gt_instances[i],
+ feats=[lvl_feat[i][None] for lvl_feat in x])
+ sampling_results.append(sampling_result)
+
+ # bbox head forward and loss
+ bbox_results = self.bbox_loss(
+ stage=stage,
+ x=x,
+ sampling_results=sampling_results,
+ semantic_feat=semantic_feat)
+
+ for name, value in bbox_results['loss_bbox'].items():
+ losses[f's{stage}.{name}'] = (
+ value * stage_loss_weight if 'loss' in name else value)
+
+ # mask head forward and loss
+ if self.with_mask:
+ # interleaved execution: use regressed bboxes by the box branch
+ # to train the mask branch
+ if self.interleaved:
+ bbox_head = self.bbox_head[stage]
+ with torch.no_grad():
+ results_list = bbox_head.refine_bboxes(
+ sampling_results, bbox_results, batch_img_metas)
+ # re-assign and sample 512 RoIs from 512 RoIs
+ sampling_results = []
+ for i in range(num_imgs):
+ results = results_list[i]
+ # rename rpn_results.bboxes to rpn_results.priors
+ results.priors = results.pop('bboxes')
+ assign_result = bbox_assigner.assign(
+ results, batch_gt_instances[i],
+ batch_gt_instances_ignore[i])
+ sampling_result = bbox_sampler.sample(
+ assign_result,
+ results,
+ batch_gt_instances[i],
+ feats=[lvl_feat[i][None] for lvl_feat in x])
+ sampling_results.append(sampling_result)
+ mask_results = self.mask_loss(
+ stage=stage,
+ x=x,
+ sampling_results=sampling_results,
+ batch_gt_instances=batch_gt_instances,
+ semantic_feat=semantic_feat)
+ for name, value in mask_results['loss_mask'].items():
+ losses[f's{stage}.{name}'] = (
+ value * stage_loss_weight if 'loss' in name else value)
+
+ # refine bboxes (same as Cascade R-CNN)
+ if stage < self.num_stages - 1 and not self.interleaved:
+ bbox_head = self.bbox_head[stage]
+ with torch.no_grad():
+ results_list = bbox_head.refine_bboxes(
+ sampling_results=sampling_results,
+ bbox_results=bbox_results,
+ batch_img_metas=batch_img_metas)
+
+ return losses
+
+ def predict(self,
+ x: Tuple[Tensor],
+ rpn_results_list: InstanceList,
+ batch_data_samples: SampleList,
+ rescale: bool = False) -> InstanceList:
+ """Perform forward propagation of the roi head and predict detection
+ results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from upstream network. Each
+ has shape (N, C, H, W).
+ rpn_results_list (list[:obj:`InstanceData`]): list of region
+ proposals.
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool): Whether to rescale the results to
+ the original image. Defaults to False.
+
+ Returns:
+ list[obj:`InstanceData`]: Detection results of each image.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+ """
+ assert self.with_bbox, 'Bbox head must be implemented.'
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+
+ if self.with_semantic:
+ _, semantic_feat = self.semantic_head(x)
+ else:
+ semantic_feat = None
+
+ # TODO: nms_op in mmcv need be enhanced, the bbox result may get
+ # difference when not rescale in bbox_head
+
+ # If it has the mask branch, the bbox branch does not need
+ # to be scaled to the original image scale, because the mask
+ # branch will scale both bbox and mask at the same time.
+ bbox_rescale = rescale if not self.with_mask else False
+ results_list = self.predict_bbox(
+ x=x,
+ semantic_feat=semantic_feat,
+ batch_img_metas=batch_img_metas,
+ rpn_results_list=rpn_results_list,
+ rcnn_test_cfg=self.test_cfg,
+ rescale=bbox_rescale)
+
+ if self.with_mask:
+ results_list = self.predict_mask(
+ x=x,
+ semantic_heat=semantic_feat,
+ batch_img_metas=batch_img_metas,
+ results_list=results_list,
+ rescale=rescale)
+
+ return results_list
+
+ def predict_mask(self,
+ x: Tuple[Tensor],
+ semantic_heat: Tensor,
+ batch_img_metas: List[dict],
+ results_list: InstanceList,
+ rescale: bool = False) -> InstanceList:
+ """Perform forward propagation of the mask head and predict detection
+ results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Feature maps of all scale level.
+ semantic_feat (Tensor): Semantic feature.
+ batch_img_metas (list[dict]): List of image information.
+ results_list (list[:obj:`InstanceData`]): Detection results of
+ each image.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+ """
+ num_imgs = len(batch_img_metas)
+ bboxes = [res.bboxes for res in results_list]
+ mask_rois = bbox2roi(bboxes)
+ if mask_rois.shape[0] == 0:
+ results_list = empty_instances(
+ batch_img_metas=batch_img_metas,
+ device=mask_rois.device,
+ task_type='mask',
+ instance_results=results_list,
+ mask_thr_binary=self.test_cfg.mask_thr_binary)
+ return results_list
+
+ num_mask_rois_per_img = [len(res) for res in results_list]
+ mask_results = self._mask_forward(
+ stage=-1,
+ x=x,
+ rois=mask_rois,
+ semantic_feat=semantic_heat,
+ training=False)
+ # split batch mask prediction back to each image
+ aug_masks = [[
+ mask.sigmoid().detach()
+ for mask in mask_preds.split(num_mask_rois_per_img, 0)
+ ] for mask_preds in mask_results['mask_preds']]
+
+ merged_masks = []
+ for i in range(num_imgs):
+ aug_mask = [mask[i] for mask in aug_masks]
+ merged_mask = merge_aug_masks(aug_mask, batch_img_metas[i])
+ merged_masks.append(merged_mask)
+
+ results_list = self.mask_head[-1].predict_by_feat(
+ mask_preds=merged_masks,
+ results_list=results_list,
+ batch_img_metas=batch_img_metas,
+ rcnn_test_cfg=self.test_cfg,
+ rescale=rescale,
+ activate_map=True)
+
+ return results_list
+
+ def forward(self, x: Tuple[Tensor], rpn_results_list: InstanceList,
+ batch_data_samples: SampleList) -> tuple:
+ """Network forward process. Usually includes backbone, neck and head
+ forward without any post-processing.
+
+ Args:
+ x (List[Tensor]): Multi-level features that may have different
+ resolutions.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ batch_data_samples (list[:obj:`DetDataSample`]): Each item contains
+ the meta information of each image and corresponding
+ annotations.
+
+ Returns
+ tuple: A tuple of features from ``bbox_head`` and ``mask_head``
+ forward.
+ """
+ results = ()
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+ num_imgs = len(batch_img_metas)
+
+ if self.with_semantic:
+ _, semantic_feat = self.semantic_head(x)
+ else:
+ semantic_feat = None
+
+ proposals = [rpn_results.bboxes for rpn_results in rpn_results_list]
+ num_proposals_per_img = tuple(len(p) for p in proposals)
+ rois = bbox2roi(proposals)
+ # bbox head
+ if self.with_bbox:
+ rois, cls_scores, bbox_preds = self._refine_roi(
+ x=x,
+ rois=rois,
+ semantic_feat=semantic_feat,
+ batch_img_metas=batch_img_metas,
+ num_proposals_per_img=num_proposals_per_img)
+ results = results + (cls_scores, bbox_preds)
+ # mask head
+ if self.with_mask:
+ rois = torch.cat(rois)
+ mask_results = self._mask_forward(
+ stage=-1,
+ x=x,
+ rois=rois,
+ semantic_feat=semantic_feat,
+ training=False)
+ aug_masks = [[
+ mask.sigmoid().detach()
+ for mask in mask_preds.split(num_proposals_per_img, 0)
+ ] for mask_preds in mask_results['mask_preds']]
+
+ merged_masks = []
+ for i in range(num_imgs):
+ aug_mask = [mask[i] for mask in aug_masks]
+ merged_mask = merge_aug_masks(aug_mask, batch_img_metas[i])
+ merged_masks.append(merged_mask)
+ results = results + (merged_masks, )
+ return results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..48a5d4227be41b8985403251e1803f78cf500636
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/__init__.py
@@ -0,0 +1,20 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .coarse_mask_head import CoarseMaskHead
+from .dynamic_mask_head import DynamicMaskHead
+from .fcn_mask_head import FCNMaskHead
+from .feature_relay_head import FeatureRelayHead
+from .fused_semantic_head import FusedSemanticHead
+from .global_context_head import GlobalContextHead
+from .grid_head import GridHead
+from .htc_mask_head import HTCMaskHead
+from .mask_point_head import MaskPointHead
+from .maskiou_head import MaskIoUHead
+from .scnet_mask_head import SCNetMaskHead
+from .scnet_semantic_head import SCNetSemanticHead
+
+__all__ = [
+ 'FCNMaskHead', 'HTCMaskHead', 'FusedSemanticHead', 'GridHead',
+ 'MaskIoUHead', 'CoarseMaskHead', 'MaskPointHead', 'SCNetMaskHead',
+ 'SCNetSemanticHead', 'GlobalContextHead', 'FeatureRelayHead',
+ 'DynamicMaskHead'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/coarse_mask_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/coarse_mask_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..1caa901228f2439492b82d1890eba468963eb28d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/coarse_mask_head.py
@@ -0,0 +1,110 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmcv.cnn import ConvModule, Linear
+from mmengine.model import ModuleList
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import MultiConfig
+from .fcn_mask_head import FCNMaskHead
+
+
+@MODELS.register_module()
+class CoarseMaskHead(FCNMaskHead):
+ """Coarse mask head used in PointRend.
+
+ Compared with standard ``FCNMaskHead``, ``CoarseMaskHead`` will downsample
+ the input feature map instead of upsample it.
+
+ Args:
+ num_convs (int): Number of conv layers in the head. Defaults to 0.
+ num_fcs (int): Number of fc layers in the head. Defaults to 2.
+ fc_out_channels (int): Number of output channels of fc layer.
+ Defaults to 1024.
+ downsample_factor (int): The factor that feature map is downsampled by.
+ Defaults to 2.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ """
+
+ def __init__(self,
+ num_convs: int = 0,
+ num_fcs: int = 2,
+ fc_out_channels: int = 1024,
+ downsample_factor: int = 2,
+ init_cfg: MultiConfig = dict(
+ type='Xavier',
+ override=[
+ dict(name='fcs'),
+ dict(type='Constant', val=0.001, name='fc_logits')
+ ]),
+ *arg,
+ **kwarg) -> None:
+ super().__init__(
+ *arg,
+ num_convs=num_convs,
+ upsample_cfg=dict(type=None),
+ init_cfg=None,
+ **kwarg)
+ self.init_cfg = init_cfg
+ self.num_fcs = num_fcs
+ assert self.num_fcs > 0
+ self.fc_out_channels = fc_out_channels
+ self.downsample_factor = downsample_factor
+ assert self.downsample_factor >= 1
+ # remove conv_logit
+ delattr(self, 'conv_logits')
+
+ if downsample_factor > 1:
+ downsample_in_channels = (
+ self.conv_out_channels
+ if self.num_convs > 0 else self.in_channels)
+ self.downsample_conv = ConvModule(
+ downsample_in_channels,
+ self.conv_out_channels,
+ kernel_size=downsample_factor,
+ stride=downsample_factor,
+ padding=0,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg)
+ else:
+ self.downsample_conv = None
+
+ self.output_size = (self.roi_feat_size[0] // downsample_factor,
+ self.roi_feat_size[1] // downsample_factor)
+ self.output_area = self.output_size[0] * self.output_size[1]
+
+ last_layer_dim = self.conv_out_channels * self.output_area
+
+ self.fcs = ModuleList()
+ for i in range(num_fcs):
+ fc_in_channels = (
+ last_layer_dim if i == 0 else self.fc_out_channels)
+ self.fcs.append(Linear(fc_in_channels, self.fc_out_channels))
+ last_layer_dim = self.fc_out_channels
+ output_channels = self.num_classes * self.output_area
+ self.fc_logits = Linear(last_layer_dim, output_channels)
+
+ def init_weights(self) -> None:
+ """Initialize weights."""
+ super(FCNMaskHead, self).init_weights()
+
+ def forward(self, x: Tensor) -> Tensor:
+ """Forward features from the upstream network.
+
+ Args:
+ x (Tensor): Extract mask RoI features.
+
+ Returns:
+ Tensor: Predicted foreground masks.
+ """
+ for conv in self.convs:
+ x = conv(x)
+
+ if self.downsample_conv is not None:
+ x = self.downsample_conv(x)
+
+ x = x.flatten(1)
+ for fc in self.fcs:
+ x = self.relu(fc(x))
+ mask_preds = self.fc_logits(x).view(
+ x.size(0), self.num_classes, *self.output_size)
+ return mask_preds
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/dynamic_mask_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/dynamic_mask_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..f33612b1b141668d0463435975c14a26fbe5a0cd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/dynamic_mask_head.py
@@ -0,0 +1,166 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List
+
+import torch
+import torch.nn as nn
+from mmengine.config import ConfigDict
+from torch import Tensor
+
+from mmdet.models.task_modules import SamplingResult
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, InstanceList, OptConfigType, reduce_mean
+from .fcn_mask_head import FCNMaskHead
+
+
+@MODELS.register_module()
+class DynamicMaskHead(FCNMaskHead):
+ r"""Dynamic Mask Head for
+ `Instances as Queries `_
+
+ Args:
+ num_convs (int): Number of convolution layer.
+ Defaults to 4.
+ roi_feat_size (int): The output size of RoI extractor,
+ Defaults to 14.
+ in_channels (int): Input feature channels.
+ Defaults to 256.
+ conv_kernel_size (int): Kernel size of convolution layers.
+ Defaults to 3.
+ conv_out_channels (int): Output channels of convolution layers.
+ Defaults to 256.
+ num_classes (int): Number of classes.
+ Defaults to 80
+ class_agnostic (int): Whether generate class agnostic prediction.
+ Defaults to False.
+ dropout (float): Probability of drop the channel.
+ Defaults to 0.0
+ upsample_cfg (:obj:`ConfigDict` or dict): The config for
+ upsample layer.
+ conv_cfg (:obj:`ConfigDict` or dict, optional): The convolution
+ layer config.
+ norm_cfg (:obj:`ConfigDict` or dict, optional): The norm layer config.
+ dynamic_conv_cfg (:obj:`ConfigDict` or dict): The dynamic convolution
+ layer config.
+ loss_mask (:obj:`ConfigDict` or dict): The config for mask loss.
+ """
+
+ def __init__(self,
+ num_convs: int = 4,
+ roi_feat_size: int = 14,
+ in_channels: int = 256,
+ conv_kernel_size: int = 3,
+ conv_out_channels: int = 256,
+ num_classes: int = 80,
+ class_agnostic: bool = False,
+ upsample_cfg: ConfigType = dict(
+ type='deconv', scale_factor=2),
+ conv_cfg: OptConfigType = None,
+ norm_cfg: OptConfigType = None,
+ dynamic_conv_cfg: ConfigType = dict(
+ type='DynamicConv',
+ in_channels=256,
+ feat_channels=64,
+ out_channels=256,
+ input_feat_shape=14,
+ with_proj=False,
+ act_cfg=dict(type='ReLU', inplace=True),
+ norm_cfg=dict(type='LN')),
+ loss_mask: ConfigType = dict(
+ type='DiceLoss', loss_weight=8.0),
+ **kwargs) -> None:
+ super().__init__(
+ num_convs=num_convs,
+ roi_feat_size=roi_feat_size,
+ in_channels=in_channels,
+ conv_kernel_size=conv_kernel_size,
+ conv_out_channels=conv_out_channels,
+ num_classes=num_classes,
+ class_agnostic=class_agnostic,
+ upsample_cfg=upsample_cfg,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ loss_mask=loss_mask,
+ **kwargs)
+ assert class_agnostic is False, \
+ 'DynamicMaskHead only support class_agnostic=False'
+ self.fp16_enabled = False
+
+ self.instance_interactive_conv = MODELS.build(dynamic_conv_cfg)
+
+ def init_weights(self) -> None:
+ """Use xavier initialization for all weight parameter and set
+ classification head bias as a specific value when use focal loss."""
+ for p in self.parameters():
+ if p.dim() > 1:
+ nn.init.xavier_uniform_(p)
+ nn.init.constant_(self.conv_logits.bias, 0.)
+
+ def forward(self, roi_feat: Tensor, proposal_feat: Tensor) -> Tensor:
+ """Forward function of DynamicMaskHead.
+
+ Args:
+ roi_feat (Tensor): Roi-pooling features with shape
+ (batch_size*num_proposals, feature_dimensions,
+ pooling_h , pooling_w).
+ proposal_feat (Tensor): Intermediate feature get from
+ diihead in last stage, has shape
+ (batch_size*num_proposals, feature_dimensions)
+
+ Returns:
+ mask_preds (Tensor): Predicted foreground masks with shape
+ (batch_size*num_proposals, num_classes, pooling_h*2, pooling_w*2).
+ """
+
+ proposal_feat = proposal_feat.reshape(-1, self.in_channels)
+ proposal_feat_iic = self.instance_interactive_conv(
+ proposal_feat, roi_feat)
+
+ x = proposal_feat_iic.permute(0, 2, 1).reshape(roi_feat.size())
+
+ for conv in self.convs:
+ x = conv(x)
+ if self.upsample is not None:
+ x = self.upsample(x)
+ if self.upsample_method == 'deconv':
+ x = self.relu(x)
+ mask_preds = self.conv_logits(x)
+ return mask_preds
+
+ def loss_and_target(self, mask_preds: Tensor,
+ sampling_results: List[SamplingResult],
+ batch_gt_instances: InstanceList,
+ rcnn_train_cfg: ConfigDict) -> dict:
+ """Calculate the loss based on the features extracted by the mask head.
+
+ Args:
+ mask_preds (Tensor): Predicted foreground masks, has shape
+ (num_pos, num_classes, h, w).
+ sampling_results (List[obj:SamplingResult]): Assign results of
+ all images in a batch after sampling.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``labels``, and
+ ``masks`` attributes.
+ rcnn_train_cfg (obj:ConfigDict): `train_cfg` of RCNN.
+
+ Returns:
+ dict: A dictionary of loss and targets components.
+ """
+ mask_targets = self.get_targets(
+ sampling_results=sampling_results,
+ batch_gt_instances=batch_gt_instances,
+ rcnn_train_cfg=rcnn_train_cfg)
+ pos_labels = torch.cat([res.pos_gt_labels for res in sampling_results])
+
+ num_pos = pos_labels.new_ones(pos_labels.size()).float().sum()
+ avg_factor = torch.clamp(reduce_mean(num_pos), min=1.).item()
+ loss = dict()
+ if mask_preds.size(0) == 0:
+ loss_mask = mask_preds.sum()
+ else:
+ loss_mask = self.loss_mask(
+ mask_preds[torch.arange(num_pos).long(), pos_labels,
+ ...].sigmoid(),
+ mask_targets,
+ avg_factor=avg_factor)
+ loss['loss_mask'] = loss_mask
+ return dict(loss_mask=loss, mask_targets=mask_targets)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/fcn_mask_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/fcn_mask_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..3a089dfafcb69784f2fc266f0945e6d56b0466d3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/fcn_mask_head.py
@@ -0,0 +1,474 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple
+
+import numpy as np
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule, build_conv_layer, build_upsample_layer
+from mmcv.ops.carafe import CARAFEPack
+from mmengine.config import ConfigDict
+from mmengine.model import BaseModule, ModuleList
+from mmengine.structures import InstanceData
+from torch import Tensor
+from torch.nn.modules.utils import _pair
+
+from mmdet.models.task_modules.samplers import SamplingResult
+from mmdet.models.utils import empty_instances
+from mmdet.registry import MODELS
+from mmdet.structures.mask import mask_target
+from mmdet.utils import ConfigType, InstanceList, OptConfigType, OptMultiConfig
+
+BYTES_PER_FLOAT = 4
+# TODO: This memory limit may be too much or too little. It would be better to
+# determine it based on available resources.
+GPU_MEM_LIMIT = 1024**3 # 1 GB memory limit
+
+
+@MODELS.register_module()
+class FCNMaskHead(BaseModule):
+
+ def __init__(self,
+ num_convs: int = 4,
+ roi_feat_size: int = 14,
+ in_channels: int = 256,
+ conv_kernel_size: int = 3,
+ conv_out_channels: int = 256,
+ num_classes: int = 80,
+ class_agnostic: int = False,
+ upsample_cfg: ConfigType = dict(
+ type='deconv', scale_factor=2),
+ conv_cfg: OptConfigType = None,
+ norm_cfg: OptConfigType = None,
+ predictor_cfg: ConfigType = dict(type='Conv'),
+ loss_mask: ConfigType = dict(
+ type='CrossEntropyLoss', use_mask=True, loss_weight=1.0),
+ init_cfg: OptMultiConfig = None) -> None:
+ assert init_cfg is None, 'To prevent abnormal initialization ' \
+ 'behavior, init_cfg is not allowed to be set'
+ super().__init__(init_cfg=init_cfg)
+ self.upsample_cfg = upsample_cfg.copy()
+ if self.upsample_cfg['type'] not in [
+ None, 'deconv', 'nearest', 'bilinear', 'carafe'
+ ]:
+ raise ValueError(
+ f'Invalid upsample method {self.upsample_cfg["type"]}, '
+ 'accepted methods are "deconv", "nearest", "bilinear", '
+ '"carafe"')
+ self.num_convs = num_convs
+ # WARN: roi_feat_size is reserved and not used
+ self.roi_feat_size = _pair(roi_feat_size)
+ self.in_channels = in_channels
+ self.conv_kernel_size = conv_kernel_size
+ self.conv_out_channels = conv_out_channels
+ self.upsample_method = self.upsample_cfg.get('type')
+ self.scale_factor = self.upsample_cfg.pop('scale_factor', None)
+ self.num_classes = num_classes
+ self.class_agnostic = class_agnostic
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self.predictor_cfg = predictor_cfg
+ self.loss_mask = MODELS.build(loss_mask)
+
+ self.convs = ModuleList()
+ for i in range(self.num_convs):
+ in_channels = (
+ self.in_channels if i == 0 else self.conv_out_channels)
+ padding = (self.conv_kernel_size - 1) // 2
+ self.convs.append(
+ ConvModule(
+ in_channels,
+ self.conv_out_channels,
+ self.conv_kernel_size,
+ padding=padding,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg))
+ upsample_in_channels = (
+ self.conv_out_channels if self.num_convs > 0 else in_channels)
+ upsample_cfg_ = self.upsample_cfg.copy()
+ if self.upsample_method is None:
+ self.upsample = None
+ elif self.upsample_method == 'deconv':
+ upsample_cfg_.update(
+ in_channels=upsample_in_channels,
+ out_channels=self.conv_out_channels,
+ kernel_size=self.scale_factor,
+ stride=self.scale_factor)
+ self.upsample = build_upsample_layer(upsample_cfg_)
+ elif self.upsample_method == 'carafe':
+ upsample_cfg_.update(
+ channels=upsample_in_channels, scale_factor=self.scale_factor)
+ self.upsample = build_upsample_layer(upsample_cfg_)
+ else:
+ # suppress warnings
+ align_corners = (None
+ if self.upsample_method == 'nearest' else False)
+ upsample_cfg_.update(
+ scale_factor=self.scale_factor,
+ mode=self.upsample_method,
+ align_corners=align_corners)
+ self.upsample = build_upsample_layer(upsample_cfg_)
+
+ out_channels = 1 if self.class_agnostic else self.num_classes
+ logits_in_channel = (
+ self.conv_out_channels
+ if self.upsample_method == 'deconv' else upsample_in_channels)
+ self.conv_logits = build_conv_layer(self.predictor_cfg,
+ logits_in_channel, out_channels, 1)
+ self.relu = nn.ReLU(inplace=True)
+ self.debug_imgs = None
+
+ def init_weights(self) -> None:
+ """Initialize the weights."""
+ super().init_weights()
+ for m in [self.upsample, self.conv_logits]:
+ if m is None:
+ continue
+ elif isinstance(m, CARAFEPack):
+ m.init_weights()
+ elif hasattr(m, 'weight') and hasattr(m, 'bias'):
+ nn.init.kaiming_normal_(
+ m.weight, mode='fan_out', nonlinearity='relu')
+ nn.init.constant_(m.bias, 0)
+
+ def forward(self, x: Tensor) -> Tensor:
+ """Forward features from the upstream network.
+
+ Args:
+ x (Tensor): Extract mask RoI features.
+
+ Returns:
+ Tensor: Predicted foreground masks.
+ """
+ for conv in self.convs:
+ x = conv(x)
+ if self.upsample is not None:
+ x = self.upsample(x)
+ if self.upsample_method == 'deconv':
+ x = self.relu(x)
+ mask_preds = self.conv_logits(x)
+ return mask_preds
+
+ def get_targets(self, sampling_results: List[SamplingResult],
+ batch_gt_instances: InstanceList,
+ rcnn_train_cfg: ConfigDict) -> Tensor:
+ """Calculate the ground truth for all samples in a batch according to
+ the sampling_results.
+
+ Args:
+ sampling_results (List[obj:SamplingResult]): Assign results of
+ all images in a batch after sampling.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``labels``, and
+ ``masks`` attributes.
+ rcnn_train_cfg (obj:ConfigDict): `train_cfg` of RCNN.
+
+ Returns:
+ Tensor: Mask target of each positive proposals in the image.
+ """
+ pos_proposals = [res.pos_priors for res in sampling_results]
+ pos_assigned_gt_inds = [
+ res.pos_assigned_gt_inds for res in sampling_results
+ ]
+ gt_masks = [res.masks for res in batch_gt_instances]
+ mask_targets = mask_target(pos_proposals, pos_assigned_gt_inds,
+ gt_masks, rcnn_train_cfg)
+ return mask_targets
+
+ def loss_and_target(self, mask_preds: Tensor,
+ sampling_results: List[SamplingResult],
+ batch_gt_instances: InstanceList,
+ rcnn_train_cfg: ConfigDict) -> dict:
+ """Calculate the loss based on the features extracted by the mask head.
+
+ Args:
+ mask_preds (Tensor): Predicted foreground masks, has shape
+ (num_pos, num_classes, h, w).
+ sampling_results (List[obj:SamplingResult]): Assign results of
+ all images in a batch after sampling.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``labels``, and
+ ``masks`` attributes.
+ rcnn_train_cfg (obj:ConfigDict): `train_cfg` of RCNN.
+
+ Returns:
+ dict: A dictionary of loss and targets components.
+ """
+ mask_targets = self.get_targets(
+ sampling_results=sampling_results,
+ batch_gt_instances=batch_gt_instances,
+ rcnn_train_cfg=rcnn_train_cfg)
+
+ pos_labels = torch.cat([res.pos_gt_labels for res in sampling_results])
+
+ loss = dict()
+ if mask_preds.size(0) == 0:
+ loss_mask = mask_preds.sum()
+ else:
+ if self.class_agnostic:
+ loss_mask = self.loss_mask(mask_preds, mask_targets,
+ torch.zeros_like(pos_labels))
+ else:
+ loss_mask = self.loss_mask(mask_preds, mask_targets,
+ pos_labels)
+ loss['loss_mask'] = loss_mask
+ # TODO: which algorithm requires mask_targets?
+ return dict(loss_mask=loss, mask_targets=mask_targets)
+
+ def predict_by_feat(self,
+ mask_preds: Tuple[Tensor],
+ results_list: List[InstanceData],
+ batch_img_metas: List[dict],
+ rcnn_test_cfg: ConfigDict,
+ rescale: bool = False,
+ activate_map: bool = False) -> InstanceList:
+ """Transform a batch of output features extracted from the head into
+ mask results.
+
+ Args:
+ mask_preds (tuple[Tensor]): Tuple of predicted foreground masks,
+ each has shape (n, num_classes, h, w).
+ results_list (list[:obj:`InstanceData`]): Detection results of
+ each image.
+ batch_img_metas (list[dict]): List of image information.
+ rcnn_test_cfg (obj:`ConfigDict`): `test_cfg` of Bbox Head.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ activate_map (book): Whether get results with augmentations test.
+ If True, the `mask_preds` will not process with sigmoid.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ after the post process. Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+ """
+ assert len(mask_preds) == len(results_list) == len(batch_img_metas)
+
+ for img_id in range(len(batch_img_metas)):
+ img_meta = batch_img_metas[img_id]
+ results = results_list[img_id]
+ bboxes = results.bboxes
+ if bboxes.shape[0] == 0:
+ results_list[img_id] = empty_instances(
+ [img_meta],
+ bboxes.device,
+ task_type='mask',
+ instance_results=[results],
+ mask_thr_binary=rcnn_test_cfg.mask_thr_binary)[0]
+ else:
+ im_mask = self._predict_by_feat_single(
+ mask_preds=mask_preds[img_id],
+ bboxes=bboxes,
+ labels=results.labels,
+ img_meta=img_meta,
+ rcnn_test_cfg=rcnn_test_cfg,
+ rescale=rescale,
+ activate_map=activate_map)
+ results.masks = im_mask
+ return results_list
+
+ def _predict_by_feat_single(self,
+ mask_preds: Tensor,
+ bboxes: Tensor,
+ labels: Tensor,
+ img_meta: dict,
+ rcnn_test_cfg: ConfigDict,
+ rescale: bool = False,
+ activate_map: bool = False) -> Tensor:
+ """Get segmentation masks from mask_preds and bboxes.
+
+ Args:
+ mask_preds (Tensor): Predicted foreground masks, has shape
+ (n, num_classes, h, w).
+ bboxes (Tensor): Predicted bboxes, has shape (n, 4)
+ labels (Tensor): Labels of bboxes, has shape (n, )
+ img_meta (dict): image information.
+ rcnn_test_cfg (obj:`ConfigDict`): `test_cfg` of Bbox Head.
+ Defaults to None.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ activate_map (book): Whether get results with augmentations test.
+ If True, the `mask_preds` will not process with sigmoid.
+ Defaults to False.
+
+ Returns:
+ Tensor: Encoded masks, has shape (n, img_w, img_h)
+
+ Example:
+ >>> from mmengine.config import Config
+ >>> from mmdet.models.roi_heads.mask_heads.fcn_mask_head import * # NOQA
+ >>> N = 7 # N = number of extracted ROIs
+ >>> C, H, W = 11, 32, 32
+ >>> # Create example instance of FCN Mask Head.
+ >>> self = FCNMaskHead(num_classes=C, num_convs=0)
+ >>> inputs = torch.rand(N, self.in_channels, H, W)
+ >>> mask_preds = self.forward(inputs)
+ >>> # Each input is associated with some bounding box
+ >>> bboxes = torch.Tensor([[1, 1, 42, 42 ]] * N)
+ >>> labels = torch.randint(0, C, size=(N,))
+ >>> rcnn_test_cfg = Config({'mask_thr_binary': 0, })
+ >>> ori_shape = (H * 4, W * 4)
+ >>> scale_factor = (1, 1)
+ >>> rescale = False
+ >>> img_meta = {'scale_factor': scale_factor,
+ ... 'ori_shape': ori_shape}
+ >>> # Encoded masks are a list for each category.
+ >>> encoded_masks = self._get_seg_masks_single(
+ ... mask_preds, bboxes, labels,
+ ... img_meta, rcnn_test_cfg, rescale)
+ >>> assert encoded_masks.size()[0] == N
+ >>> assert encoded_masks.size()[1:] == ori_shape
+ """
+ scale_factor = bboxes.new_tensor(img_meta['scale_factor']).repeat(
+ (1, 2))
+ img_h, img_w = img_meta['ori_shape'][:2]
+ device = bboxes.device
+
+ if not activate_map:
+ mask_preds = mask_preds.sigmoid()
+ else:
+ # In AugTest, has been activated before
+ mask_preds = bboxes.new_tensor(mask_preds)
+
+ if rescale: # in-placed rescale the bboxes
+ bboxes /= scale_factor
+ else:
+ w_scale, h_scale = scale_factor[0, 0], scale_factor[0, 1]
+ img_h = np.round(img_h * h_scale.item()).astype(np.int32)
+ img_w = np.round(img_w * w_scale.item()).astype(np.int32)
+
+ N = len(mask_preds)
+ # The actual implementation split the input into chunks,
+ # and paste them chunk by chunk.
+ if device.type == 'cpu':
+ # CPU is most efficient when they are pasted one by one with
+ # skip_empty=True, so that it performs minimal number of
+ # operations.
+ num_chunks = N
+ else:
+ # GPU benefits from parallelism for larger chunks,
+ # but may have memory issue
+ # the types of img_w and img_h are np.int32,
+ # when the image resolution is large,
+ # the calculation of num_chunks will overflow.
+ # so we need to change the types of img_w and img_h to int.
+ # See https://github.com/open-mmlab/mmdetection/pull/5191
+ num_chunks = int(
+ np.ceil(N * int(img_h) * int(img_w) * BYTES_PER_FLOAT /
+ GPU_MEM_LIMIT))
+ assert (num_chunks <=
+ N), 'Default GPU_MEM_LIMIT is too small; try increasing it'
+ chunks = torch.chunk(torch.arange(N, device=device), num_chunks)
+
+ threshold = rcnn_test_cfg.mask_thr_binary
+ im_mask = torch.zeros(
+ N,
+ img_h,
+ img_w,
+ device=device,
+ dtype=torch.bool if threshold >= 0 else torch.uint8)
+
+ if not self.class_agnostic:
+ mask_preds = mask_preds[range(N), labels][:, None]
+
+ for inds in chunks:
+ masks_chunk, spatial_inds = _do_paste_mask(
+ mask_preds[inds],
+ bboxes[inds],
+ img_h,
+ img_w,
+ skip_empty=device.type == 'cpu')
+
+ if threshold >= 0:
+ masks_chunk = (masks_chunk >= threshold).to(dtype=torch.bool)
+ else:
+ # for visualization and debugging
+ masks_chunk = (masks_chunk * 255).to(dtype=torch.uint8)
+
+ im_mask[(inds, ) + spatial_inds] = masks_chunk
+ return im_mask
+
+
+def _do_paste_mask(masks: Tensor,
+ boxes: Tensor,
+ img_h: int,
+ img_w: int,
+ skip_empty: bool = True) -> tuple:
+ """Paste instance masks according to boxes.
+
+ This implementation is modified from
+ https://github.com/facebookresearch/detectron2/
+
+ Args:
+ masks (Tensor): N, 1, H, W
+ boxes (Tensor): N, 4
+ img_h (int): Height of the image to be pasted.
+ img_w (int): Width of the image to be pasted.
+ skip_empty (bool): Only paste masks within the region that
+ tightly bound all boxes, and returns the results this region only.
+ An important optimization for CPU.
+
+ Returns:
+ tuple: (Tensor, tuple). The first item is mask tensor, the second one
+ is the slice object.
+
+ If skip_empty == False, the whole image will be pasted. It will
+ return a mask of shape (N, img_h, img_w) and an empty tuple.
+
+ If skip_empty == True, only area around the mask will be pasted.
+ A mask of shape (N, h', w') and its start and end coordinates
+ in the original image will be returned.
+ """
+ # On GPU, paste all masks together (up to chunk size)
+ # by using the entire image to sample the masks
+ # Compared to pasting them one by one,
+ # this has more operations but is faster on COCO-scale dataset.
+ device = masks.device
+ if skip_empty:
+ x0_int, y0_int = torch.clamp(
+ boxes.min(dim=0).values.floor()[:2] - 1,
+ min=0).to(dtype=torch.int32)
+ x1_int = torch.clamp(
+ boxes[:, 2].max().ceil() + 1, max=img_w).to(dtype=torch.int32)
+ y1_int = torch.clamp(
+ boxes[:, 3].max().ceil() + 1, max=img_h).to(dtype=torch.int32)
+ else:
+ x0_int, y0_int = 0, 0
+ x1_int, y1_int = img_w, img_h
+ x0, y0, x1, y1 = torch.split(boxes, 1, dim=1) # each is Nx1
+
+ N = masks.shape[0]
+
+ img_y = torch.arange(y0_int, y1_int, device=device).to(torch.float32) + 0.5
+ img_x = torch.arange(x0_int, x1_int, device=device).to(torch.float32) + 0.5
+ img_y = (img_y - y0) / (y1 - y0) * 2 - 1
+ img_x = (img_x - x0) / (x1 - x0) * 2 - 1
+ # img_x, img_y have shapes (N, w), (N, h)
+ # IsInf op is not supported with ONNX<=1.7.0
+ if not torch.onnx.is_in_onnx_export():
+ if torch.isinf(img_x).any():
+ inds = torch.where(torch.isinf(img_x))
+ img_x[inds] = 0
+ if torch.isinf(img_y).any():
+ inds = torch.where(torch.isinf(img_y))
+ img_y[inds] = 0
+
+ gx = img_x[:, None, :].expand(N, img_y.size(1), img_x.size(1))
+ gy = img_y[:, :, None].expand(N, img_y.size(1), img_x.size(1))
+ grid = torch.stack([gx, gy], dim=3)
+
+ img_masks = F.grid_sample(
+ masks.to(dtype=torch.float32), grid, align_corners=False)
+
+ if skip_empty:
+ return img_masks[:, 0], (slice(y0_int, y1_int), slice(x0_int, x1_int))
+ else:
+ return img_masks[:, 0], ()
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/feature_relay_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/feature_relay_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..0c34561fa5fd749329eda164465ce9787278d357
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/feature_relay_head.py
@@ -0,0 +1,68 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional
+
+import torch.nn as nn
+from mmengine.model import BaseModule
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import MultiConfig
+
+
+@MODELS.register_module()
+class FeatureRelayHead(BaseModule):
+ """Feature Relay Head used in `SCNet `_.
+
+ Args:
+ in_channels (int): number of input channels. Defaults to 256.
+ conv_out_channels (int): number of output channels before
+ classification layer. Defaults to 256.
+ roi_feat_size (int): roi feat size at box head. Default: 7.
+ scale_factor (int): scale factor to match roi feat size
+ at mask head. Defaults to 2.
+ init_cfg (:obj:`ConfigDict` or dict or list[dict] or
+ list[:obj:`ConfigDict`]): Initialization config dict. Defaults to
+ dict(type='Kaiming', layer='Linear').
+ """
+
+ def __init__(
+ self,
+ in_channels: int = 1024,
+ out_conv_channels: int = 256,
+ roi_feat_size: int = 7,
+ scale_factor: int = 2,
+ init_cfg: MultiConfig = dict(type='Kaiming', layer='Linear')
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ assert isinstance(roi_feat_size, int)
+
+ self.in_channels = in_channels
+ self.out_conv_channels = out_conv_channels
+ self.roi_feat_size = roi_feat_size
+ self.out_channels = (roi_feat_size**2) * out_conv_channels
+ self.scale_factor = scale_factor
+ self.fp16_enabled = False
+
+ self.fc = nn.Linear(self.in_channels, self.out_channels)
+ self.upsample = nn.Upsample(
+ scale_factor=scale_factor, mode='bilinear', align_corners=True)
+
+ def forward(self, x: Tensor) -> Optional[Tensor]:
+ """Forward function.
+
+ Args:
+ x (Tensor): Input feature.
+
+ Returns:
+ Optional[Tensor]: Output feature. When the first dim of input is
+ 0, None is returned.
+ """
+ N, _ = x.shape
+ if N > 0:
+ out_C = self.out_conv_channels
+ out_HW = self.roi_feat_size
+ x = self.fc(x)
+ x = x.reshape(N, out_C, out_HW, out_HW)
+ x = self.upsample(x)
+ return x
+ return None
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/fused_semantic_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/fused_semantic_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..d20beb2975a563f03e7b6b2afcef287cb41af05a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/fused_semantic_head.py
@@ -0,0 +1,144 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+from typing import Tuple
+
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule
+from mmengine.config import ConfigDict
+from mmengine.model import BaseModule
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import MultiConfig, OptConfigType
+
+
+@MODELS.register_module()
+class FusedSemanticHead(BaseModule):
+ r"""Multi-level fused semantic segmentation head.
+
+ .. code-block:: none
+
+ in_1 -> 1x1 conv ---
+ |
+ in_2 -> 1x1 conv -- |
+ ||
+ in_3 -> 1x1 conv - ||
+ ||| /-> 1x1 conv (mask prediction)
+ in_4 -> 1x1 conv -----> 3x3 convs (*4)
+ | \-> 1x1 conv (feature)
+ in_5 -> 1x1 conv ---
+ """ # noqa: W605
+
+ def __init__(
+ self,
+ num_ins: int,
+ fusion_level: int,
+ seg_scale_factor=1 / 8,
+ num_convs: int = 4,
+ in_channels: int = 256,
+ conv_out_channels: int = 256,
+ num_classes: int = 183,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: OptConfigType = None,
+ ignore_label: int = None,
+ loss_weight: float = None,
+ loss_seg: ConfigDict = dict(
+ type='CrossEntropyLoss', ignore_index=255, loss_weight=0.2),
+ init_cfg: MultiConfig = dict(
+ type='Kaiming', override=dict(name='conv_logits'))
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.num_ins = num_ins
+ self.fusion_level = fusion_level
+ self.seg_scale_factor = seg_scale_factor
+ self.num_convs = num_convs
+ self.in_channels = in_channels
+ self.conv_out_channels = conv_out_channels
+ self.num_classes = num_classes
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self.fp16_enabled = False
+
+ self.lateral_convs = nn.ModuleList()
+ for i in range(self.num_ins):
+ self.lateral_convs.append(
+ ConvModule(
+ self.in_channels,
+ self.in_channels,
+ 1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ inplace=False))
+
+ self.convs = nn.ModuleList()
+ for i in range(self.num_convs):
+ in_channels = self.in_channels if i == 0 else conv_out_channels
+ self.convs.append(
+ ConvModule(
+ in_channels,
+ conv_out_channels,
+ 3,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ self.conv_embedding = ConvModule(
+ conv_out_channels,
+ conv_out_channels,
+ 1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg)
+ self.conv_logits = nn.Conv2d(conv_out_channels, self.num_classes, 1)
+ if ignore_label:
+ loss_seg['ignore_index'] = ignore_label
+ if loss_weight:
+ loss_seg['loss_weight'] = loss_weight
+ if ignore_label or loss_weight:
+ warnings.warn('``ignore_label`` and ``loss_weight`` would be '
+ 'deprecated soon. Please set ``ingore_index`` and '
+ '``loss_weight`` in ``loss_seg`` instead.')
+ self.criterion = MODELS.build(loss_seg)
+
+ def forward(self, feats: Tuple[Tensor]) -> Tuple[Tensor]:
+ """Forward function.
+
+ Args:
+ feats (tuple[Tensor]): Multi scale feature maps.
+
+ Returns:
+ tuple[Tensor]:
+
+ - mask_preds (Tensor): Predicted mask logits.
+ - x (Tensor): Fused feature.
+ """
+ x = self.lateral_convs[self.fusion_level](feats[self.fusion_level])
+ fused_size = tuple(x.shape[-2:])
+ for i, feat in enumerate(feats):
+ if i != self.fusion_level:
+ feat = F.interpolate(
+ feat, size=fused_size, mode='bilinear', align_corners=True)
+ # fix runtime error of "+=" inplace operation in PyTorch 1.10
+ x = x + self.lateral_convs[i](feat)
+
+ for i in range(self.num_convs):
+ x = self.convs[i](x)
+
+ mask_preds = self.conv_logits(x)
+ x = self.conv_embedding(x)
+ return mask_preds, x
+
+ def loss(self, mask_preds: Tensor, labels: Tensor) -> Tensor:
+ """Loss function.
+
+ Args:
+ mask_preds (Tensor): Predicted mask logits.
+ labels (Tensor): Ground truth.
+
+ Returns:
+ Tensor: Semantic segmentation loss.
+ """
+ labels = F.interpolate(
+ labels.float(), scale_factor=self.seg_scale_factor, mode='nearest')
+ labels = labels.squeeze(1).long()
+ loss_semantic_seg = self.criterion(mask_preds, labels)
+ return loss_semantic_seg
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/global_context_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/global_context_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..cb947ea582227d2b74112cbb930e1a3f85b77ff5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/global_context_head.py
@@ -0,0 +1,127 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple
+
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from mmengine.model import BaseModule
+from torch import Tensor
+
+from mmdet.models.layers import ResLayer, SimplifiedBasicBlock
+from mmdet.registry import MODELS
+from mmdet.utils import MultiConfig, OptConfigType
+
+
+@MODELS.register_module()
+class GlobalContextHead(BaseModule):
+ """Global context head used in `SCNet `_.
+
+ Args:
+ num_convs (int, optional): number of convolutional layer in GlbCtxHead.
+ Defaults to 4.
+ in_channels (int, optional): number of input channels. Defaults to 256.
+ conv_out_channels (int, optional): number of output channels before
+ classification layer. Defaults to 256.
+ num_classes (int, optional): number of classes. Defaults to 80.
+ loss_weight (float, optional): global context loss weight.
+ Defaults to 1.
+ conv_cfg (dict, optional): config to init conv layer. Defaults to None.
+ norm_cfg (dict, optional): config to init norm layer. Defaults to None.
+ conv_to_res (bool, optional): if True, 2 convs will be grouped into
+ 1 `SimplifiedBasicBlock` using a skip connection.
+ Defaults to False.
+ init_cfg (:obj:`ConfigDict` or dict or list[dict] or
+ list[:obj:`ConfigDict`]): Initialization config dict. Defaults to
+ dict(type='Normal', std=0.01, override=dict(name='fc')).
+ """
+
+ def __init__(
+ self,
+ num_convs: int = 4,
+ in_channels: int = 256,
+ conv_out_channels: int = 256,
+ num_classes: int = 80,
+ loss_weight: float = 1.0,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: OptConfigType = None,
+ conv_to_res: bool = False,
+ init_cfg: MultiConfig = dict(
+ type='Normal', std=0.01, override=dict(name='fc'))
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.num_convs = num_convs
+ self.in_channels = in_channels
+ self.conv_out_channels = conv_out_channels
+ self.num_classes = num_classes
+ self.loss_weight = loss_weight
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self.conv_to_res = conv_to_res
+ self.fp16_enabled = False
+
+ if self.conv_to_res:
+ num_res_blocks = num_convs // 2
+ self.convs = ResLayer(
+ SimplifiedBasicBlock,
+ in_channels,
+ self.conv_out_channels,
+ num_res_blocks,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg)
+ self.num_convs = num_res_blocks
+ else:
+ self.convs = nn.ModuleList()
+ for i in range(self.num_convs):
+ in_channels = self.in_channels if i == 0 else conv_out_channels
+ self.convs.append(
+ ConvModule(
+ in_channels,
+ conv_out_channels,
+ 3,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+
+ self.pool = nn.AdaptiveAvgPool2d(1)
+ self.fc = nn.Linear(conv_out_channels, num_classes)
+
+ self.criterion = nn.BCEWithLogitsLoss()
+
+ def forward(self, feats: Tuple[Tensor]) -> Tuple[Tensor]:
+ """Forward function.
+
+ Args:
+ feats (Tuple[Tensor]): Multi-scale feature maps.
+
+ Returns:
+ Tuple[Tensor]:
+
+ - mc_pred (Tensor): Multi-class prediction.
+ - x (Tensor): Global context feature.
+ """
+ x = feats[-1]
+ for i in range(self.num_convs):
+ x = self.convs[i](x)
+ x = self.pool(x)
+
+ # multi-class prediction
+ mc_pred = x.reshape(x.size(0), -1)
+ mc_pred = self.fc(mc_pred)
+
+ return mc_pred, x
+
+ def loss(self, pred: Tensor, labels: List[Tensor]) -> Tensor:
+ """Loss function.
+
+ Args:
+ pred (Tensor): Logits.
+ labels (list[Tensor]): Grouth truths.
+
+ Returns:
+ Tensor: Loss.
+ """
+ labels = [lbl.unique() for lbl in labels]
+ targets = pred.new_zeros(pred.size())
+ for i, label in enumerate(labels):
+ targets[i, label] = 1.0
+ loss = self.loss_weight * self.criterion(pred, targets)
+ return loss
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/grid_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/grid_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..d9514ae7bcfc1b7d5613fa0107e9bd087e13dd46
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/grid_head.py
@@ -0,0 +1,490 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, List, Tuple
+
+import numpy as np
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import ConvModule
+from mmengine.config import ConfigDict
+from mmengine.model import BaseModule
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.models.task_modules.samplers import SamplingResult
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, InstanceList, MultiConfig, OptConfigType
+
+
+@MODELS.register_module()
+class GridHead(BaseModule):
+ """Implementation of `Grid Head `_
+
+ Args:
+ grid_points (int): The number of grid points. Defaults to 9.
+ num_convs (int): The number of convolution layers. Defaults to 8.
+ roi_feat_size (int): RoI feature size. Default to 14.
+ in_channels (int): The channel number of inputs features.
+ Defaults to 256.
+ conv_kernel_size (int): The kernel size of convolution layers.
+ Defaults to 3.
+ point_feat_channels (int): The number of channels of each point
+ features. Defaults to 64.
+ class_agnostic (bool): Whether use class agnostic classification.
+ If so, the output channels of logits will be 1. Defaults to False.
+ loss_grid (:obj:`ConfigDict` or dict): Config of grid loss.
+ conv_cfg (:obj:`ConfigDict` or dict, optional) dictionary to
+ construct and config conv layer.
+ norm_cfg (:obj:`ConfigDict` or dict): dictionary to construct and
+ config norm layer.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict]): Initialization config dict.
+ """
+
+ def __init__(
+ self,
+ grid_points: int = 9,
+ num_convs: int = 8,
+ roi_feat_size: int = 14,
+ in_channels: int = 256,
+ conv_kernel_size: int = 3,
+ point_feat_channels: int = 64,
+ deconv_kernel_size: int = 4,
+ class_agnostic: bool = False,
+ loss_grid: ConfigType = dict(
+ type='CrossEntropyLoss', use_sigmoid=True, loss_weight=15),
+ conv_cfg: OptConfigType = None,
+ norm_cfg: ConfigType = dict(type='GN', num_groups=36),
+ init_cfg: MultiConfig = [
+ dict(type='Kaiming', layer=['Conv2d', 'Linear']),
+ dict(
+ type='Normal',
+ layer='ConvTranspose2d',
+ std=0.001,
+ override=dict(
+ type='Normal',
+ name='deconv2',
+ std=0.001,
+ bias=-np.log(0.99 / 0.01)))
+ ]
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.grid_points = grid_points
+ self.num_convs = num_convs
+ self.roi_feat_size = roi_feat_size
+ self.in_channels = in_channels
+ self.conv_kernel_size = conv_kernel_size
+ self.point_feat_channels = point_feat_channels
+ self.conv_out_channels = self.point_feat_channels * self.grid_points
+ self.class_agnostic = class_agnostic
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ if isinstance(norm_cfg, dict) and norm_cfg['type'] == 'GN':
+ assert self.conv_out_channels % norm_cfg['num_groups'] == 0
+
+ assert self.grid_points >= 4
+ self.grid_size = int(np.sqrt(self.grid_points))
+ if self.grid_size * self.grid_size != self.grid_points:
+ raise ValueError('grid_points must be a square number')
+
+ # the predicted heatmap is half of whole_map_size
+ if not isinstance(self.roi_feat_size, int):
+ raise ValueError('Only square RoIs are supporeted in Grid R-CNN')
+ self.whole_map_size = self.roi_feat_size * 4
+
+ # compute point-wise sub-regions
+ self.sub_regions = self.calc_sub_regions()
+
+ self.convs = []
+ for i in range(self.num_convs):
+ in_channels = (
+ self.in_channels if i == 0 else self.conv_out_channels)
+ stride = 2 if i == 0 else 1
+ padding = (self.conv_kernel_size - 1) // 2
+ self.convs.append(
+ ConvModule(
+ in_channels,
+ self.conv_out_channels,
+ self.conv_kernel_size,
+ stride=stride,
+ padding=padding,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg,
+ bias=True))
+ self.convs = nn.Sequential(*self.convs)
+
+ self.deconv1 = nn.ConvTranspose2d(
+ self.conv_out_channels,
+ self.conv_out_channels,
+ kernel_size=deconv_kernel_size,
+ stride=2,
+ padding=(deconv_kernel_size - 2) // 2,
+ groups=grid_points)
+ self.norm1 = nn.GroupNorm(grid_points, self.conv_out_channels)
+ self.deconv2 = nn.ConvTranspose2d(
+ self.conv_out_channels,
+ grid_points,
+ kernel_size=deconv_kernel_size,
+ stride=2,
+ padding=(deconv_kernel_size - 2) // 2,
+ groups=grid_points)
+
+ # find the 4-neighbor of each grid point
+ self.neighbor_points = []
+ grid_size = self.grid_size
+ for i in range(grid_size): # i-th column
+ for j in range(grid_size): # j-th row
+ neighbors = []
+ if i > 0: # left: (i - 1, j)
+ neighbors.append((i - 1) * grid_size + j)
+ if j > 0: # up: (i, j - 1)
+ neighbors.append(i * grid_size + j - 1)
+ if j < grid_size - 1: # down: (i, j + 1)
+ neighbors.append(i * grid_size + j + 1)
+ if i < grid_size - 1: # right: (i + 1, j)
+ neighbors.append((i + 1) * grid_size + j)
+ self.neighbor_points.append(tuple(neighbors))
+ # total edges in the grid
+ self.num_edges = sum([len(p) for p in self.neighbor_points])
+
+ self.forder_trans = nn.ModuleList() # first-order feature transition
+ self.sorder_trans = nn.ModuleList() # second-order feature transition
+ for neighbors in self.neighbor_points:
+ fo_trans = nn.ModuleList()
+ so_trans = nn.ModuleList()
+ for _ in range(len(neighbors)):
+ # each transition module consists of a 5x5 depth-wise conv and
+ # 1x1 conv.
+ fo_trans.append(
+ nn.Sequential(
+ nn.Conv2d(
+ self.point_feat_channels,
+ self.point_feat_channels,
+ 5,
+ stride=1,
+ padding=2,
+ groups=self.point_feat_channels),
+ nn.Conv2d(self.point_feat_channels,
+ self.point_feat_channels, 1)))
+ so_trans.append(
+ nn.Sequential(
+ nn.Conv2d(
+ self.point_feat_channels,
+ self.point_feat_channels,
+ 5,
+ 1,
+ 2,
+ groups=self.point_feat_channels),
+ nn.Conv2d(self.point_feat_channels,
+ self.point_feat_channels, 1)))
+ self.forder_trans.append(fo_trans)
+ self.sorder_trans.append(so_trans)
+
+ self.loss_grid = MODELS.build(loss_grid)
+
+ def forward(self, x: Tensor) -> Dict[str, Tensor]:
+ """forward function of ``GridHead``.
+
+ Args:
+ x (Tensor): RoI features, has shape
+ (num_rois, num_channels, roi_feat_size, roi_feat_size).
+
+ Returns:
+ Dict[str, Tensor]: Return a dict including fused and unfused
+ heatmap.
+ """
+ assert x.shape[-1] == x.shape[-2] == self.roi_feat_size
+ # RoI feature transformation, downsample 2x
+ x = self.convs(x)
+
+ c = self.point_feat_channels
+ # first-order fusion
+ x_fo = [None for _ in range(self.grid_points)]
+ for i, points in enumerate(self.neighbor_points):
+ x_fo[i] = x[:, i * c:(i + 1) * c]
+ for j, point_idx in enumerate(points):
+ x_fo[i] = x_fo[i] + self.forder_trans[i][j](
+ x[:, point_idx * c:(point_idx + 1) * c])
+
+ # second-order fusion
+ x_so = [None for _ in range(self.grid_points)]
+ for i, points in enumerate(self.neighbor_points):
+ x_so[i] = x[:, i * c:(i + 1) * c]
+ for j, point_idx in enumerate(points):
+ x_so[i] = x_so[i] + self.sorder_trans[i][j](x_fo[point_idx])
+
+ # predicted heatmap with fused features
+ x2 = torch.cat(x_so, dim=1)
+ x2 = self.deconv1(x2)
+ x2 = F.relu(self.norm1(x2), inplace=True)
+ heatmap = self.deconv2(x2)
+
+ # predicted heatmap with original features (applicable during training)
+ if self.training:
+ x1 = x
+ x1 = self.deconv1(x1)
+ x1 = F.relu(self.norm1(x1), inplace=True)
+ heatmap_unfused = self.deconv2(x1)
+ else:
+ heatmap_unfused = heatmap
+
+ return dict(fused=heatmap, unfused=heatmap_unfused)
+
+ def calc_sub_regions(self) -> List[Tuple[float]]:
+ """Compute point specific representation regions.
+
+ See `Grid R-CNN Plus `_ for details.
+ """
+ # to make it consistent with the original implementation, half_size
+ # is computed as 2 * quarter_size, which is smaller
+ half_size = self.whole_map_size // 4 * 2
+ sub_regions = []
+ for i in range(self.grid_points):
+ x_idx = i // self.grid_size
+ y_idx = i % self.grid_size
+ if x_idx == 0:
+ sub_x1 = 0
+ elif x_idx == self.grid_size - 1:
+ sub_x1 = half_size
+ else:
+ ratio = x_idx / (self.grid_size - 1) - 0.25
+ sub_x1 = max(int(ratio * self.whole_map_size), 0)
+
+ if y_idx == 0:
+ sub_y1 = 0
+ elif y_idx == self.grid_size - 1:
+ sub_y1 = half_size
+ else:
+ ratio = y_idx / (self.grid_size - 1) - 0.25
+ sub_y1 = max(int(ratio * self.whole_map_size), 0)
+ sub_regions.append(
+ (sub_x1, sub_y1, sub_x1 + half_size, sub_y1 + half_size))
+ return sub_regions
+
+ def get_targets(self, sampling_results: List[SamplingResult],
+ rcnn_train_cfg: ConfigDict) -> Tensor:
+ """Calculate the ground truth for all samples in a batch according to
+ the sampling_results.".
+
+ Args:
+ sampling_results (List[:obj:`SamplingResult`]): Assign results of
+ all images in a batch after sampling.
+ rcnn_train_cfg (:obj:`ConfigDict`): `train_cfg` of RCNN.
+
+ Returns:
+ Tensor: Grid heatmap targets.
+ """
+ # mix all samples (across images) together.
+ pos_bboxes = torch.cat([res.pos_bboxes for res in sampling_results],
+ dim=0).cpu()
+ pos_gt_bboxes = torch.cat(
+ [res.pos_gt_bboxes for res in sampling_results], dim=0).cpu()
+ assert pos_bboxes.shape == pos_gt_bboxes.shape
+
+ # expand pos_bboxes to 2x of original size
+ x1 = pos_bboxes[:, 0] - (pos_bboxes[:, 2] - pos_bboxes[:, 0]) / 2
+ y1 = pos_bboxes[:, 1] - (pos_bboxes[:, 3] - pos_bboxes[:, 1]) / 2
+ x2 = pos_bboxes[:, 2] + (pos_bboxes[:, 2] - pos_bboxes[:, 0]) / 2
+ y2 = pos_bboxes[:, 3] + (pos_bboxes[:, 3] - pos_bboxes[:, 1]) / 2
+ pos_bboxes = torch.stack([x1, y1, x2, y2], dim=-1)
+ pos_bbox_ws = (pos_bboxes[:, 2] - pos_bboxes[:, 0]).unsqueeze(-1)
+ pos_bbox_hs = (pos_bboxes[:, 3] - pos_bboxes[:, 1]).unsqueeze(-1)
+
+ num_rois = pos_bboxes.shape[0]
+ map_size = self.whole_map_size
+ # this is not the final target shape
+ targets = torch.zeros((num_rois, self.grid_points, map_size, map_size),
+ dtype=torch.float)
+
+ # pre-compute interpolation factors for all grid points.
+ # the first item is the factor of x-dim, and the second is y-dim.
+ # for a 9-point grid, factors are like (1, 0), (0.5, 0.5), (0, 1)
+ factors = []
+ for j in range(self.grid_points):
+ x_idx = j // self.grid_size
+ y_idx = j % self.grid_size
+ factors.append((1 - x_idx / (self.grid_size - 1),
+ 1 - y_idx / (self.grid_size - 1)))
+
+ radius = rcnn_train_cfg.pos_radius
+ radius2 = radius**2
+ for i in range(num_rois):
+ # ignore small bboxes
+ if (pos_bbox_ws[i] <= self.grid_size
+ or pos_bbox_hs[i] <= self.grid_size):
+ continue
+ # for each grid point, mark a small circle as positive
+ for j in range(self.grid_points):
+ factor_x, factor_y = factors[j]
+ gridpoint_x = factor_x * pos_gt_bboxes[i, 0] + (
+ 1 - factor_x) * pos_gt_bboxes[i, 2]
+ gridpoint_y = factor_y * pos_gt_bboxes[i, 1] + (
+ 1 - factor_y) * pos_gt_bboxes[i, 3]
+
+ cx = int((gridpoint_x - pos_bboxes[i, 0]) / pos_bbox_ws[i] *
+ map_size)
+ cy = int((gridpoint_y - pos_bboxes[i, 1]) / pos_bbox_hs[i] *
+ map_size)
+
+ for x in range(cx - radius, cx + radius + 1):
+ for y in range(cy - radius, cy + radius + 1):
+ if x >= 0 and x < map_size and y >= 0 and y < map_size:
+ if (x - cx)**2 + (y - cy)**2 <= radius2:
+ targets[i, j, y, x] = 1
+ # reduce the target heatmap size by a half
+ # proposed in Grid R-CNN Plus (https://arxiv.org/abs/1906.05688).
+ sub_targets = []
+ for i in range(self.grid_points):
+ sub_x1, sub_y1, sub_x2, sub_y2 = self.sub_regions[i]
+ sub_targets.append(targets[:, [i], sub_y1:sub_y2, sub_x1:sub_x2])
+ sub_targets = torch.cat(sub_targets, dim=1)
+ sub_targets = sub_targets.to(sampling_results[0].pos_bboxes.device)
+ return sub_targets
+
+ def loss(self, grid_pred: Tensor, sample_idx: Tensor,
+ sampling_results: List[SamplingResult],
+ rcnn_train_cfg: ConfigDict) -> dict:
+ """Calculate the loss based on the features extracted by the grid head.
+
+ Args:
+ grid_pred (dict[str, Tensor]): Outputs of grid_head forward.
+ sample_idx (Tensor): The sampling index of ``grid_pred``.
+ sampling_results (List[obj:SamplingResult]): Assign results of
+ all images in a batch after sampling.
+ rcnn_train_cfg (obj:`ConfigDict`): `train_cfg` of RCNN.
+
+ Returns:
+ dict: A dictionary of loss and targets components.
+ """
+ grid_targets = self.get_targets(sampling_results, rcnn_train_cfg)
+ grid_targets = grid_targets[sample_idx]
+
+ loss_fused = self.loss_grid(grid_pred['fused'], grid_targets)
+ loss_unfused = self.loss_grid(grid_pred['unfused'], grid_targets)
+ loss_grid = loss_fused + loss_unfused
+ return dict(loss_grid=loss_grid)
+
+ def predict_by_feat(self,
+ grid_preds: Dict[str, Tensor],
+ results_list: List[InstanceData],
+ batch_img_metas: List[dict],
+ rescale: bool = False) -> InstanceList:
+ """Adjust the predicted bboxes from bbox head.
+
+ Args:
+ grid_preds (dict[str, Tensor]): dictionary outputted by forward
+ function.
+ results_list (list[:obj:`InstanceData`]): Detection results of
+ each image.
+ batch_img_metas (list[dict]): List of image information.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ after the post process. Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape \
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4), the last \
+ dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ num_roi_per_img = tuple(res.bboxes.size(0) for res in results_list)
+ grid_preds = {
+ k: v.split(num_roi_per_img, 0)
+ for k, v in grid_preds.items()
+ }
+
+ for i, results in enumerate(results_list):
+ if len(results) != 0:
+ bboxes = self._predict_by_feat_single(
+ grid_pred=grid_preds['fused'][i],
+ bboxes=results.bboxes,
+ img_meta=batch_img_metas[i],
+ rescale=rescale)
+ results.bboxes = bboxes
+ return results_list
+
+ def _predict_by_feat_single(self,
+ grid_pred: Tensor,
+ bboxes: Tensor,
+ img_meta: dict,
+ rescale: bool = False) -> Tensor:
+ """Adjust ``bboxes`` according to ``grid_pred``.
+
+ Args:
+ grid_pred (Tensor): Grid fused heatmap.
+ bboxes (Tensor): Predicted bboxes, has shape (n, 4)
+ img_meta (dict): image information.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ Tensor: adjusted bboxes.
+ """
+ assert bboxes.size(0) == grid_pred.size(0)
+ grid_pred = grid_pred.sigmoid()
+
+ R, c, h, w = grid_pred.shape
+ half_size = self.whole_map_size // 4 * 2
+ assert h == w == half_size
+ assert c == self.grid_points
+
+ # find the point with max scores in the half-sized heatmap
+ grid_pred = grid_pred.view(R * c, h * w)
+ pred_scores, pred_position = grid_pred.max(dim=1)
+ xs = pred_position % w
+ ys = pred_position // w
+
+ # get the position in the whole heatmap instead of half-sized heatmap
+ for i in range(self.grid_points):
+ xs[i::self.grid_points] += self.sub_regions[i][0]
+ ys[i::self.grid_points] += self.sub_regions[i][1]
+
+ # reshape to (num_rois, grid_points)
+ pred_scores, xs, ys = tuple(
+ map(lambda x: x.view(R, c), [pred_scores, xs, ys]))
+
+ # get expanded pos_bboxes
+ widths = (bboxes[:, 2] - bboxes[:, 0]).unsqueeze(-1)
+ heights = (bboxes[:, 3] - bboxes[:, 1]).unsqueeze(-1)
+ x1 = (bboxes[:, 0, None] - widths / 2)
+ y1 = (bboxes[:, 1, None] - heights / 2)
+ # map the grid point to the absolute coordinates
+ abs_xs = (xs.float() + 0.5) / w * widths + x1
+ abs_ys = (ys.float() + 0.5) / h * heights + y1
+
+ # get the grid points indices that fall on the bbox boundaries
+ x1_inds = [i for i in range(self.grid_size)]
+ y1_inds = [i * self.grid_size for i in range(self.grid_size)]
+ x2_inds = [
+ self.grid_points - self.grid_size + i
+ for i in range(self.grid_size)
+ ]
+ y2_inds = [(i + 1) * self.grid_size - 1 for i in range(self.grid_size)]
+
+ # voting of all grid points on some boundary
+ bboxes_x1 = (abs_xs[:, x1_inds] * pred_scores[:, x1_inds]).sum(
+ dim=1, keepdim=True) / (
+ pred_scores[:, x1_inds].sum(dim=1, keepdim=True))
+ bboxes_y1 = (abs_ys[:, y1_inds] * pred_scores[:, y1_inds]).sum(
+ dim=1, keepdim=True) / (
+ pred_scores[:, y1_inds].sum(dim=1, keepdim=True))
+ bboxes_x2 = (abs_xs[:, x2_inds] * pred_scores[:, x2_inds]).sum(
+ dim=1, keepdim=True) / (
+ pred_scores[:, x2_inds].sum(dim=1, keepdim=True))
+ bboxes_y2 = (abs_ys[:, y2_inds] * pred_scores[:, y2_inds]).sum(
+ dim=1, keepdim=True) / (
+ pred_scores[:, y2_inds].sum(dim=1, keepdim=True))
+
+ bboxes = torch.cat([bboxes_x1, bboxes_y1, bboxes_x2, bboxes_y2], dim=1)
+ bboxes[:, [0, 2]].clamp_(min=0, max=img_meta['img_shape'][1])
+ bboxes[:, [1, 3]].clamp_(min=0, max=img_meta['img_shape'][0])
+
+ if rescale:
+ assert img_meta.get('scale_factor') is not None
+ bboxes /= bboxes.new_tensor(img_meta['scale_factor']).repeat(
+ (1, 2))
+
+ return bboxes
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/htc_mask_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/htc_mask_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..73ac1e6e5f115927e1a2accdd693aae512cac753
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/htc_mask_head.py
@@ -0,0 +1,65 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Union
+
+from mmcv.cnn import ConvModule
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from .fcn_mask_head import FCNMaskHead
+
+
+@MODELS.register_module()
+class HTCMaskHead(FCNMaskHead):
+ """Mask head for HTC.
+
+ Args:
+ with_conv_res (bool): Whether add conv layer for ``res_feat``.
+ Defaults to True.
+ """
+
+ def __init__(self, with_conv_res: bool = True, *args, **kwargs) -> None:
+ super().__init__(*args, **kwargs)
+ self.with_conv_res = with_conv_res
+ if self.with_conv_res:
+ self.conv_res = ConvModule(
+ self.conv_out_channels,
+ self.conv_out_channels,
+ 1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg)
+
+ def forward(self,
+ x: Tensor,
+ res_feat: Optional[Tensor] = None,
+ return_logits: bool = True,
+ return_feat: bool = True) -> Union[Tensor, List[Tensor]]:
+ """
+ Args:
+ x (Tensor): Feature map.
+ res_feat (Tensor, optional): Feature for residual connection.
+ Defaults to None.
+ return_logits (bool): Whether return mask logits. Defaults to True.
+ return_feat (bool): Whether return feature map. Defaults to True.
+
+ Returns:
+ Union[Tensor, List[Tensor]]: The return result is one of three
+ results: res_feat, logits, or [logits, res_feat].
+ """
+ assert not (not return_logits and not return_feat)
+ if res_feat is not None:
+ assert self.with_conv_res
+ res_feat = self.conv_res(res_feat)
+ x = x + res_feat
+ for conv in self.convs:
+ x = conv(x)
+ res_feat = x
+ outs = []
+ if return_logits:
+ x = self.upsample(x)
+ if self.upsample_method == 'deconv':
+ x = self.relu(x)
+ mask_preds = self.conv_logits(x)
+ outs.append(mask_preds)
+ if return_feat:
+ outs.append(res_feat)
+ return outs if len(outs) > 1 else outs[0]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/mask_point_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/mask_point_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..2084f59f07b48bf2e5b05bb7af61172df8737478
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/mask_point_head.py
@@ -0,0 +1,284 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+# Modified from https://github.com/facebookresearch/detectron2/tree/master/projects/PointRend/point_head/point_head.py # noqa
+
+from typing import List, Tuple
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from mmcv.ops import point_sample, rel_roi_point_to_rel_img_point
+from mmengine.model import BaseModule
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.models.task_modules.samplers import SamplingResult
+from mmdet.models.utils import (get_uncertain_point_coords_with_randomness,
+ get_uncertainty)
+from mmdet.registry import MODELS
+from mmdet.structures.bbox import bbox2roi
+from mmdet.utils import ConfigType, InstanceList, MultiConfig, OptConfigType
+
+
+@MODELS.register_module()
+class MaskPointHead(BaseModule):
+ """A mask point head use in PointRend.
+
+ ``MaskPointHead`` use shared multi-layer perceptron (equivalent to
+ nn.Conv1d) to predict the logit of input points. The fine-grained feature
+ and coarse feature will be concatenate together for predication.
+
+ Args:
+ num_fcs (int): Number of fc layers in the head. Defaults to 3.
+ in_channels (int): Number of input channels. Defaults to 256.
+ fc_channels (int): Number of fc channels. Defaults to 256.
+ num_classes (int): Number of classes for logits. Defaults to 80.
+ class_agnostic (bool): Whether use class agnostic classification.
+ If so, the output channels of logits will be 1. Defaults to False.
+ coarse_pred_each_layer (bool): Whether concatenate coarse feature with
+ the output of each fc layer. Defaults to True.
+ conv_cfg (:obj:`ConfigDict` or dict): Dictionary to construct
+ and config conv layer. Defaults to dict(type='Conv1d')).
+ norm_cfg (:obj:`ConfigDict` or dict, optional): Dictionary to construct
+ and config norm layer. Defaults to None.
+ loss_point (:obj:`ConfigDict` or dict): Dictionary to construct and
+ config loss layer of point head. Defaults to
+ dict(type='CrossEntropyLoss', use_mask=True, loss_weight=1.0).
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict], optional): Initialization config dict.
+ """
+
+ def __init__(
+ self,
+ num_classes: int,
+ num_fcs: int = 3,
+ in_channels: int = 256,
+ fc_channels: int = 256,
+ class_agnostic: bool = False,
+ coarse_pred_each_layer: bool = True,
+ conv_cfg: ConfigType = dict(type='Conv1d'),
+ norm_cfg: OptConfigType = None,
+ act_cfg: ConfigType = dict(type='ReLU'),
+ loss_point: ConfigType = dict(
+ type='CrossEntropyLoss', use_mask=True, loss_weight=1.0),
+ init_cfg: MultiConfig = dict(
+ type='Normal', std=0.001, override=dict(name='fc_logits'))
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.num_fcs = num_fcs
+ self.in_channels = in_channels
+ self.fc_channels = fc_channels
+ self.num_classes = num_classes
+ self.class_agnostic = class_agnostic
+ self.coarse_pred_each_layer = coarse_pred_each_layer
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self.loss_point = MODELS.build(loss_point)
+
+ fc_in_channels = in_channels + num_classes
+ self.fcs = nn.ModuleList()
+ for _ in range(num_fcs):
+ fc = ConvModule(
+ fc_in_channels,
+ fc_channels,
+ kernel_size=1,
+ stride=1,
+ padding=0,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ act_cfg=act_cfg)
+ self.fcs.append(fc)
+ fc_in_channels = fc_channels
+ fc_in_channels += num_classes if self.coarse_pred_each_layer else 0
+
+ out_channels = 1 if self.class_agnostic else self.num_classes
+ self.fc_logits = nn.Conv1d(
+ fc_in_channels, out_channels, kernel_size=1, stride=1, padding=0)
+
+ def forward(self, fine_grained_feats: Tensor,
+ coarse_feats: Tensor) -> Tensor:
+ """Classify each point base on fine grained and coarse feats.
+
+ Args:
+ fine_grained_feats (Tensor): Fine grained feature sampled from FPN,
+ shape (num_rois, in_channels, num_points).
+ coarse_feats (Tensor): Coarse feature sampled from CoarseMaskHead,
+ shape (num_rois, num_classes, num_points).
+
+ Returns:
+ Tensor: Point classification results,
+ shape (num_rois, num_class, num_points).
+ """
+
+ x = torch.cat([fine_grained_feats, coarse_feats], dim=1)
+ for fc in self.fcs:
+ x = fc(x)
+ if self.coarse_pred_each_layer:
+ x = torch.cat((x, coarse_feats), dim=1)
+ return self.fc_logits(x)
+
+ def get_targets(self, rois: Tensor, rel_roi_points: Tensor,
+ sampling_results: List[SamplingResult],
+ batch_gt_instances: InstanceList,
+ cfg: ConfigType) -> Tensor:
+ """Get training targets of MaskPointHead for all images.
+
+ Args:
+ rois (Tensor): Region of Interest, shape (num_rois, 5).
+ rel_roi_points (Tensor): Points coordinates relative to RoI, shape
+ (num_rois, num_points, 2).
+ sampling_results (:obj:`SamplingResult`): Sampling result after
+ sampling and assignment.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``labels``, and
+ ``masks`` attributes.
+ cfg (obj:`ConfigDict` or dict): Training cfg.
+
+ Returns:
+ Tensor: Point target, shape (num_rois, num_points).
+ """
+
+ num_imgs = len(sampling_results)
+ rois_list = []
+ rel_roi_points_list = []
+ for batch_ind in range(num_imgs):
+ inds = (rois[:, 0] == batch_ind)
+ rois_list.append(rois[inds])
+ rel_roi_points_list.append(rel_roi_points[inds])
+ pos_assigned_gt_inds_list = [
+ res.pos_assigned_gt_inds for res in sampling_results
+ ]
+ cfg_list = [cfg for _ in range(num_imgs)]
+
+ point_targets = map(self._get_targets_single, rois_list,
+ rel_roi_points_list, pos_assigned_gt_inds_list,
+ batch_gt_instances, cfg_list)
+ point_targets = list(point_targets)
+
+ if len(point_targets) > 0:
+ point_targets = torch.cat(point_targets)
+
+ return point_targets
+
+ def _get_targets_single(self, rois: Tensor, rel_roi_points: Tensor,
+ pos_assigned_gt_inds: Tensor,
+ gt_instances: InstanceData,
+ cfg: ConfigType) -> Tensor:
+ """Get training target of MaskPointHead for each image."""
+ num_pos = rois.size(0)
+ num_points = cfg.num_points
+ if num_pos > 0:
+ gt_masks_th = (
+ gt_instances.masks.to_tensor(rois.dtype,
+ rois.device).index_select(
+ 0, pos_assigned_gt_inds))
+ gt_masks_th = gt_masks_th.unsqueeze(1)
+ rel_img_points = rel_roi_point_to_rel_img_point(
+ rois, rel_roi_points, gt_masks_th)
+ point_targets = point_sample(gt_masks_th,
+ rel_img_points).squeeze(1)
+ else:
+ point_targets = rois.new_zeros((0, num_points))
+ return point_targets
+
+ def loss_and_target(self, point_pred: Tensor, rel_roi_points: Tensor,
+ sampling_results: List[SamplingResult],
+ batch_gt_instances: InstanceList,
+ cfg: ConfigType) -> dict:
+ """Calculate loss for MaskPointHead.
+
+ Args:
+ point_pred (Tensor): Point predication result, shape
+ (num_rois, num_classes, num_points).
+ rel_roi_points (Tensor): Points coordinates relative to RoI, shape
+ (num_rois, num_points, 2).
+ sampling_results (:obj:`SamplingResult`): Sampling result after
+ sampling and assignment.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``labels``, and
+ ``masks`` attributes.
+ cfg (obj:`ConfigDict` or dict): Training cfg.
+
+ Returns:
+ dict: a dictionary of point loss and point target.
+ """
+ rois = bbox2roi([res.pos_bboxes for res in sampling_results])
+ pos_labels = torch.cat([res.pos_gt_labels for res in sampling_results])
+
+ point_target = self.get_targets(rois, rel_roi_points, sampling_results,
+ batch_gt_instances, cfg)
+ if self.class_agnostic:
+ loss_point = self.loss_point(point_pred, point_target,
+ torch.zeros_like(pos_labels))
+ else:
+ loss_point = self.loss_point(point_pred, point_target, pos_labels)
+
+ return dict(loss_point=loss_point, point_target=point_target)
+
+ def get_roi_rel_points_train(self, mask_preds: Tensor, labels: Tensor,
+ cfg: ConfigType) -> Tensor:
+ """Get ``num_points`` most uncertain points with random points during
+ train.
+
+ Sample points in [0, 1] x [0, 1] coordinate space based on their
+ uncertainty. The uncertainties are calculated for each point using
+ '_get_uncertainty()' function that takes point's logit prediction as
+ input.
+
+ Args:
+ mask_preds (Tensor): A tensor of shape (num_rois, num_classes,
+ mask_height, mask_width) for class-specific or class-agnostic
+ prediction.
+ labels (Tensor): The ground truth class for each instance.
+ cfg (:obj:`ConfigDict` or dict): Training config of point head.
+
+ Returns:
+ point_coords (Tensor): A tensor of shape (num_rois, num_points, 2)
+ that contains the coordinates sampled points.
+ """
+ point_coords = get_uncertain_point_coords_with_randomness(
+ mask_preds, labels, cfg.num_points, cfg.oversample_ratio,
+ cfg.importance_sample_ratio)
+ return point_coords
+
+ def get_roi_rel_points_test(self, mask_preds: Tensor, label_preds: Tensor,
+ cfg: ConfigType) -> Tuple[Tensor, Tensor]:
+ """Get ``num_points`` most uncertain points during test.
+
+ Args:
+ mask_preds (Tensor): A tensor of shape (num_rois, num_classes,
+ mask_height, mask_width) for class-specific or class-agnostic
+ prediction.
+ label_preds (Tensor): The predication class for each instance.
+ cfg (:obj:`ConfigDict` or dict): Testing config of point head.
+
+ Returns:
+ tuple:
+
+ - point_indices (Tensor): A tensor of shape (num_rois, num_points)
+ that contains indices from [0, mask_height x mask_width) of the
+ most uncertain points.
+ - point_coords (Tensor): A tensor of shape (num_rois, num_points,
+ 2) that contains [0, 1] x [0, 1] normalized coordinates of the
+ most uncertain points from the [mask_height, mask_width] grid.
+ """
+ num_points = cfg.subdivision_num_points
+ uncertainty_map = get_uncertainty(mask_preds, label_preds)
+ num_rois, _, mask_height, mask_width = uncertainty_map.shape
+
+ # During ONNX exporting, the type of each elements of 'shape' is
+ # `Tensor(float)`, while it is `float` during PyTorch inference.
+ if isinstance(mask_height, torch.Tensor):
+ h_step = 1.0 / mask_height.float()
+ w_step = 1.0 / mask_width.float()
+ else:
+ h_step = 1.0 / mask_height
+ w_step = 1.0 / mask_width
+ # cast to int to avoid dynamic K for TopK op in ONNX
+ mask_size = int(mask_height * mask_width)
+ uncertainty_map = uncertainty_map.view(num_rois, mask_size)
+ num_points = min(mask_size, num_points)
+ point_indices = uncertainty_map.topk(num_points, dim=1)[1]
+ xs = w_step / 2.0 + (point_indices % mask_width).float() * w_step
+ ys = h_step / 2.0 + (point_indices // mask_width).float() * h_step
+ point_coords = torch.stack([xs, ys], dim=2)
+ return point_indices, point_coords
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/maskiou_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/maskiou_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..8901871e754c491f7bc94eb68a27fa1b50e29148
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/maskiou_head.py
@@ -0,0 +1,277 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple
+
+import numpy as np
+import torch
+import torch.nn as nn
+from mmcv.cnn import Conv2d, Linear, MaxPool2d
+from mmengine.config import ConfigDict
+from mmengine.model import BaseModule
+from mmengine.structures import InstanceData
+from torch import Tensor
+from torch.nn.modules.utils import _pair
+
+from mmdet.models.task_modules.samplers import SamplingResult
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, InstanceList, OptMultiConfig
+
+
+@MODELS.register_module()
+class MaskIoUHead(BaseModule):
+ """Mask IoU Head.
+
+ This head predicts the IoU of predicted masks and corresponding gt masks.
+
+ Args:
+ num_convs (int): The number of convolution layers. Defaults to 4.
+ num_fcs (int): The number of fully connected layers. Defaults to 2.
+ roi_feat_size (int): RoI feature size. Default to 14.
+ in_channels (int): The channel number of inputs features.
+ Defaults to 256.
+ conv_out_channels (int): The feature channels of convolution layers.
+ Defaults to 256.
+ fc_out_channels (int): The feature channels of fully connected layers.
+ Defaults to 1024.
+ num_classes (int): Number of categories excluding the background
+ category. Defaults to 80.
+ loss_iou (:obj:`ConfigDict` or dict): IoU loss.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict], optional): Initialization config dict.
+ """
+
+ def __init__(
+ self,
+ num_convs: int = 4,
+ num_fcs: int = 2,
+ roi_feat_size: int = 14,
+ in_channels: int = 256,
+ conv_out_channels: int = 256,
+ fc_out_channels: int = 1024,
+ num_classes: int = 80,
+ loss_iou: ConfigType = dict(type='MSELoss', loss_weight=0.5),
+ init_cfg: OptMultiConfig = [
+ dict(type='Kaiming', override=dict(name='convs')),
+ dict(type='Caffe2Xavier', override=dict(name='fcs')),
+ dict(type='Normal', std=0.01, override=dict(name='fc_mask_iou'))
+ ]
+ ) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.in_channels = in_channels
+ self.conv_out_channels = conv_out_channels
+ self.fc_out_channels = fc_out_channels
+ self.num_classes = num_classes
+
+ self.convs = nn.ModuleList()
+ for i in range(num_convs):
+ if i == 0:
+ # concatenation of mask feature and mask prediction
+ in_channels = self.in_channels + 1
+ else:
+ in_channels = self.conv_out_channels
+ stride = 2 if i == num_convs - 1 else 1
+ self.convs.append(
+ Conv2d(
+ in_channels,
+ self.conv_out_channels,
+ 3,
+ stride=stride,
+ padding=1))
+
+ roi_feat_size = _pair(roi_feat_size)
+ pooled_area = (roi_feat_size[0] // 2) * (roi_feat_size[1] // 2)
+ self.fcs = nn.ModuleList()
+ for i in range(num_fcs):
+ in_channels = (
+ self.conv_out_channels *
+ pooled_area if i == 0 else self.fc_out_channels)
+ self.fcs.append(Linear(in_channels, self.fc_out_channels))
+
+ self.fc_mask_iou = Linear(self.fc_out_channels, self.num_classes)
+ self.relu = nn.ReLU()
+ self.max_pool = MaxPool2d(2, 2)
+ self.loss_iou = MODELS.build(loss_iou)
+
+ def forward(self, mask_feat: Tensor, mask_preds: Tensor) -> Tensor:
+ """Forward function.
+
+ Args:
+ mask_feat (Tensor): Mask features from upstream models.
+ mask_preds (Tensor): Mask predictions from mask head.
+
+ Returns:
+ Tensor: Mask IoU predictions.
+ """
+ mask_preds = mask_preds.sigmoid()
+ mask_pred_pooled = self.max_pool(mask_preds.unsqueeze(1))
+
+ x = torch.cat((mask_feat, mask_pred_pooled), 1)
+
+ for conv in self.convs:
+ x = self.relu(conv(x))
+ x = x.flatten(1)
+ for fc in self.fcs:
+ x = self.relu(fc(x))
+ mask_iou = self.fc_mask_iou(x)
+ return mask_iou
+
+ def loss_and_target(self, mask_iou_pred: Tensor, mask_preds: Tensor,
+ mask_targets: Tensor,
+ sampling_results: List[SamplingResult],
+ batch_gt_instances: InstanceList,
+ rcnn_train_cfg: ConfigDict) -> dict:
+ """Calculate the loss and targets of MaskIoUHead.
+
+ Args:
+ mask_iou_pred (Tensor): Mask IoU predictions results, has shape
+ (num_pos, num_classes)
+ mask_preds (Tensor): Mask predictions from mask head, has shape
+ (num_pos, mask_size, mask_size).
+ mask_targets (Tensor): The ground truth masks assigned with
+ predictions, has shape
+ (num_pos, mask_size, mask_size).
+ sampling_results (List[obj:SamplingResult]): Assign results of
+ all images in a batch after sampling.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It includes ``masks`` inside.
+ rcnn_train_cfg (obj:`ConfigDict`): `train_cfg` of RCNN.
+
+ Returns:
+ dict: A dictionary of loss and targets components.
+ The targets are only used for cascade rcnn.
+ """
+ mask_iou_targets = self.get_targets(
+ sampling_results=sampling_results,
+ batch_gt_instances=batch_gt_instances,
+ mask_preds=mask_preds,
+ mask_targets=mask_targets,
+ rcnn_train_cfg=rcnn_train_cfg)
+
+ pos_inds = mask_iou_targets > 0
+ if pos_inds.sum() > 0:
+ loss_mask_iou = self.loss_iou(mask_iou_pred[pos_inds],
+ mask_iou_targets[pos_inds])
+ else:
+ loss_mask_iou = mask_iou_pred.sum() * 0
+ return dict(loss_mask_iou=loss_mask_iou)
+
+ def get_targets(self, sampling_results: List[SamplingResult],
+ batch_gt_instances: InstanceList, mask_preds: Tensor,
+ mask_targets: Tensor,
+ rcnn_train_cfg: ConfigDict) -> Tensor:
+ """Compute target of mask IoU.
+
+ Mask IoU target is the IoU of the predicted mask (inside a bbox) and
+ the gt mask of corresponding gt mask (the whole instance).
+ The intersection area is computed inside the bbox, and the gt mask area
+ is computed with two steps, firstly we compute the gt area inside the
+ bbox, then divide it by the area ratio of gt area inside the bbox and
+ the gt area of the whole instance.
+
+ Args:
+ sampling_results (list[:obj:`SamplingResult`]): sampling results.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It includes ``masks`` inside.
+ mask_preds (Tensor): Predicted masks of each positive proposal,
+ shape (num_pos, h, w).
+ mask_targets (Tensor): Gt mask of each positive proposal,
+ binary map of the shape (num_pos, h, w).
+ rcnn_train_cfg (obj:`ConfigDict`): Training config for R-CNN part.
+
+ Returns:
+ Tensor: mask iou target (length == num positive).
+ """
+ pos_proposals = [res.pos_priors for res in sampling_results]
+ pos_assigned_gt_inds = [
+ res.pos_assigned_gt_inds for res in sampling_results
+ ]
+ gt_masks = [res.masks for res in batch_gt_instances]
+
+ # compute the area ratio of gt areas inside the proposals and
+ # the whole instance
+ area_ratios = map(self._get_area_ratio, pos_proposals,
+ pos_assigned_gt_inds, gt_masks)
+ area_ratios = torch.cat(list(area_ratios))
+ assert mask_targets.size(0) == area_ratios.size(0)
+
+ mask_preds = (mask_preds > rcnn_train_cfg.mask_thr_binary).float()
+ mask_pred_areas = mask_preds.sum((-1, -2))
+
+ # mask_preds and mask_targets are binary maps
+ overlap_areas = (mask_preds * mask_targets).sum((-1, -2))
+
+ # compute the mask area of the whole instance
+ gt_full_areas = mask_targets.sum((-1, -2)) / (area_ratios + 1e-7)
+
+ mask_iou_targets = overlap_areas / (
+ mask_pred_areas + gt_full_areas - overlap_areas)
+ return mask_iou_targets
+
+ def _get_area_ratio(self, pos_proposals: Tensor,
+ pos_assigned_gt_inds: Tensor,
+ gt_masks: InstanceData) -> Tensor:
+ """Compute area ratio of the gt mask inside the proposal and the gt
+ mask of the corresponding instance.
+
+ Args:
+ pos_proposals (Tensor): Positive proposals, has shape (num_pos, 4).
+ pos_assigned_gt_inds (Tensor): positive proposals assigned ground
+ truth index.
+ gt_masks (BitmapMask or PolygonMask): Gt masks (the whole instance)
+ of each image, with the same shape of the input image.
+
+ Returns:
+ Tensor: The area ratio of the gt mask inside the proposal and the
+ gt mask of the corresponding instance.
+ """
+ num_pos = pos_proposals.size(0)
+ if num_pos > 0:
+ area_ratios = []
+ proposals_np = pos_proposals.cpu().numpy()
+ pos_assigned_gt_inds = pos_assigned_gt_inds.cpu().numpy()
+ # compute mask areas of gt instances (batch processing for speedup)
+ gt_instance_mask_area = gt_masks.areas
+ for i in range(num_pos):
+ gt_mask = gt_masks[pos_assigned_gt_inds[i]]
+
+ # crop the gt mask inside the proposal
+ bbox = proposals_np[i, :].astype(np.int32)
+ gt_mask_in_proposal = gt_mask.crop(bbox)
+
+ ratio = gt_mask_in_proposal.areas[0] / (
+ gt_instance_mask_area[pos_assigned_gt_inds[i]] + 1e-7)
+ area_ratios.append(ratio)
+ area_ratios = torch.from_numpy(np.stack(area_ratios)).float().to(
+ pos_proposals.device)
+ else:
+ area_ratios = pos_proposals.new_zeros((0, ))
+ return area_ratios
+
+ def predict_by_feat(self, mask_iou_preds: Tuple[Tensor],
+ results_list: InstanceList) -> InstanceList:
+ """Predict the mask iou and calculate it into ``results.scores``.
+
+ Args:
+ mask_iou_preds (Tensor): Mask IoU predictions results, has shape
+ (num_proposals, num_classes)
+ results_list (list[:obj:`InstanceData`]): Detection results of
+ each image.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ after the post process. Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+ """
+ assert len(mask_iou_preds) == len(results_list)
+ for results, mask_iou_pred in zip(results_list, mask_iou_preds):
+ labels = results.labels
+ scores = results.scores
+ results.scores = scores * mask_iou_pred[range(labels.size(0)),
+ labels]
+ return results_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/scnet_mask_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/scnet_mask_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..ffd30c337c37f4e280980e459c126df177fe7efa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/scnet_mask_head.py
@@ -0,0 +1,28 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.models.layers import ResLayer, SimplifiedBasicBlock
+from mmdet.registry import MODELS
+from .fcn_mask_head import FCNMaskHead
+
+
+@MODELS.register_module()
+class SCNetMaskHead(FCNMaskHead):
+ """Mask head for `SCNet `_.
+
+ Args:
+ conv_to_res (bool, optional): if True, change the conv layers to
+ ``SimplifiedBasicBlock``.
+ """
+
+ def __init__(self, conv_to_res: bool = True, **kwargs) -> None:
+ super().__init__(**kwargs)
+ self.conv_to_res = conv_to_res
+ if conv_to_res:
+ assert self.conv_kernel_size == 3
+ self.num_res_blocks = self.num_convs // 2
+ self.convs = ResLayer(
+ SimplifiedBasicBlock,
+ self.in_channels,
+ self.conv_out_channels,
+ self.num_res_blocks,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/scnet_semantic_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/scnet_semantic_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..55c5c8e4fae7d4e941a770d985c7253fd70f2226
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_heads/scnet_semantic_head.py
@@ -0,0 +1,28 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.models.layers import ResLayer, SimplifiedBasicBlock
+from mmdet.registry import MODELS
+from .fused_semantic_head import FusedSemanticHead
+
+
+@MODELS.register_module()
+class SCNetSemanticHead(FusedSemanticHead):
+ """Mask head for `SCNet `_.
+
+ Args:
+ conv_to_res (bool, optional): if True, change the conv layers to
+ ``SimplifiedBasicBlock``.
+ """
+
+ def __init__(self, conv_to_res: bool = True, **kwargs) -> None:
+ super().__init__(**kwargs)
+ self.conv_to_res = conv_to_res
+ if self.conv_to_res:
+ num_res_blocks = self.num_convs // 2
+ self.convs = ResLayer(
+ SimplifiedBasicBlock,
+ self.in_channels,
+ self.conv_out_channels,
+ num_res_blocks,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg)
+ self.num_convs = num_res_blocks
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_scoring_roi_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_scoring_roi_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..6545c0ed41ee7ad17b5f1b841f8bc8d65a7b6391
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/mask_scoring_roi_head.py
@@ -0,0 +1,208 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple
+
+import torch
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import bbox2roi
+from mmdet.utils import ConfigType, InstanceList
+from ..task_modules.samplers import SamplingResult
+from ..utils.misc import empty_instances
+from .standard_roi_head import StandardRoIHead
+
+
+@MODELS.register_module()
+class MaskScoringRoIHead(StandardRoIHead):
+ """Mask Scoring RoIHead for `Mask Scoring RCNN.
+
+ `_.
+
+ Args:
+ mask_iou_head (:obj`ConfigDict`, dict): The config of mask_iou_head.
+ """
+
+ def __init__(self, mask_iou_head: ConfigType, **kwargs):
+ assert mask_iou_head is not None
+ super().__init__(**kwargs)
+ self.mask_iou_head = MODELS.build(mask_iou_head)
+
+ def forward(self,
+ x: Tuple[Tensor],
+ rpn_results_list: InstanceList,
+ batch_data_samples: SampleList = None) -> tuple:
+ """Network forward process. Usually includes backbone, neck and head
+ forward without any post-processing.
+
+ Args:
+ x (List[Tensor]): Multi-level features that may have different
+ resolutions.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ batch_data_samples (list[:obj:`DetDataSample`]): Each item contains
+ the meta information of each image and corresponding
+ annotations.
+
+ Returns
+ tuple: A tuple of features from ``bbox_head`` and ``mask_head``
+ forward.
+ """
+ results = ()
+ proposals = [rpn_results.bboxes for rpn_results in rpn_results_list]
+ rois = bbox2roi(proposals)
+ # bbox head
+ if self.with_bbox:
+ bbox_results = self._bbox_forward(x, rois)
+ results = results + (bbox_results['cls_score'],
+ bbox_results['bbox_pred'])
+ # mask head
+ if self.with_mask:
+ mask_rois = rois[:100]
+ mask_results = self._mask_forward(x, mask_rois)
+ results = results + (mask_results['mask_preds'], )
+
+ # mask iou head
+ cls_score = bbox_results['cls_score'][:100]
+ mask_preds = mask_results['mask_preds']
+ mask_feats = mask_results['mask_feats']
+ _, labels = cls_score[:, :self.bbox_head.num_classes].max(dim=1)
+ mask_iou_preds = self.mask_iou_head(
+ mask_feats, mask_preds[range(labels.size(0)), labels])
+ results = results + (mask_iou_preds, )
+
+ return results
+
+ def mask_loss(self, x: Tuple[Tensor],
+ sampling_results: List[SamplingResult], bbox_feats,
+ batch_gt_instances: InstanceList) -> dict:
+ """Perform forward propagation and loss calculation of the mask head on
+ the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Tuple of multi-level img features.
+ sampling_results (list["obj:`SamplingResult`]): Sampling results.
+ bbox_feats (Tensor): Extract bbox RoI features.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``labels``, and
+ ``masks`` attributes.
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `mask_preds` (Tensor): Mask prediction.
+ - `mask_feats` (Tensor): Extract mask RoI features.
+ - `mask_targets` (Tensor): Mask target of each positive\
+ proposals in the image.
+ - `loss_mask` (dict): A dictionary of mask loss components.
+ - `loss_mask_iou` (Tensor): mask iou loss.
+ """
+ if not self.share_roi_extractor:
+ pos_rois = bbox2roi([res.pos_priors for res in sampling_results])
+ mask_results = self._mask_forward(x, pos_rois)
+ else:
+ pos_inds = []
+ device = bbox_feats.device
+ for res in sampling_results:
+ pos_inds.append(
+ torch.ones(
+ res.pos_priors.shape[0],
+ device=device,
+ dtype=torch.uint8))
+ pos_inds.append(
+ torch.zeros(
+ res.neg_priors.shape[0],
+ device=device,
+ dtype=torch.uint8))
+ pos_inds = torch.cat(pos_inds)
+
+ mask_results = self._mask_forward(
+ x, pos_inds=pos_inds, bbox_feats=bbox_feats)
+
+ mask_loss_and_target = self.mask_head.loss_and_target(
+ mask_preds=mask_results['mask_preds'],
+ sampling_results=sampling_results,
+ batch_gt_instances=batch_gt_instances,
+ rcnn_train_cfg=self.train_cfg)
+ mask_targets = mask_loss_and_target['mask_targets']
+ mask_results.update(loss_mask=mask_loss_and_target['loss_mask'])
+ if mask_results['loss_mask'] is None:
+ return mask_results
+
+ # mask iou head forward and loss
+ pos_labels = torch.cat([res.pos_gt_labels for res in sampling_results])
+ pos_mask_pred = mask_results['mask_preds'][
+ range(mask_results['mask_preds'].size(0)), pos_labels]
+ mask_iou_pred = self.mask_iou_head(mask_results['mask_feats'],
+ pos_mask_pred)
+ pos_mask_iou_pred = mask_iou_pred[range(mask_iou_pred.size(0)),
+ pos_labels]
+
+ loss_mask_iou = self.mask_iou_head.loss_and_target(
+ pos_mask_iou_pred, pos_mask_pred, mask_targets, sampling_results,
+ batch_gt_instances, self.train_cfg)
+ mask_results['loss_mask'].update(loss_mask_iou)
+ return mask_results
+
+ def predict_mask(self,
+ x: Tensor,
+ batch_img_metas: List[dict],
+ results_list: InstanceList,
+ rescale: bool = False) -> InstanceList:
+ """Perform forward propagation of the mask head and predict detection
+ results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Feature maps of all scale level.
+ batch_img_metas (list[dict]): List of image information.
+ results_list (list[:obj:`InstanceData`]): Detection results of
+ each image.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+ """
+ bboxes = [res.bboxes for res in results_list]
+ mask_rois = bbox2roi(bboxes)
+ if mask_rois.shape[0] == 0:
+ results_list = empty_instances(
+ batch_img_metas,
+ mask_rois.device,
+ task_type='mask',
+ instance_results=results_list,
+ mask_thr_binary=self.test_cfg.mask_thr_binary)
+ return results_list
+
+ mask_results = self._mask_forward(x, mask_rois)
+ mask_preds = mask_results['mask_preds']
+ mask_feats = mask_results['mask_feats']
+ # get mask scores with mask iou head
+ labels = torch.cat([res.labels for res in results_list])
+ mask_iou_preds = self.mask_iou_head(
+ mask_feats, mask_preds[range(labels.size(0)), labels])
+ # split batch mask prediction back to each image
+ num_mask_rois_per_img = [len(res) for res in results_list]
+ mask_preds = mask_preds.split(num_mask_rois_per_img, 0)
+ mask_iou_preds = mask_iou_preds.split(num_mask_rois_per_img, 0)
+
+ # TODO: Handle the case where rescale is false
+ results_list = self.mask_head.predict_by_feat(
+ mask_preds=mask_preds,
+ results_list=results_list,
+ batch_img_metas=batch_img_metas,
+ rcnn_test_cfg=self.test_cfg,
+ rescale=rescale)
+ results_list = self.mask_iou_head.predict_by_feat(
+ mask_iou_preds=mask_iou_preds, results_list=results_list)
+ return results_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/multi_instance_roi_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/multi_instance_roi_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..fee55b0a5d341c03165649f59737fd34d85c207e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/multi_instance_roi_head.py
@@ -0,0 +1,226 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple
+
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import DetDataSample
+from mmdet.structures.bbox import bbox2roi
+from mmdet.utils import ConfigType, InstanceList
+from ..task_modules.samplers import SamplingResult
+from ..utils import empty_instances, unpack_gt_instances
+from .standard_roi_head import StandardRoIHead
+
+
+@MODELS.register_module()
+class MultiInstanceRoIHead(StandardRoIHead):
+ """The roi head for Multi-instance prediction."""
+
+ def __init__(self, num_instance: int = 2, *args, **kwargs) -> None:
+ self.num_instance = num_instance
+ super().__init__(*args, **kwargs)
+
+ def init_bbox_head(self, bbox_roi_extractor: ConfigType,
+ bbox_head: ConfigType) -> None:
+ """Initialize box head and box roi extractor.
+
+ Args:
+ bbox_roi_extractor (dict or ConfigDict): Config of box
+ roi extractor.
+ bbox_head (dict or ConfigDict): Config of box in box head.
+ """
+ self.bbox_roi_extractor = MODELS.build(bbox_roi_extractor)
+ self.bbox_head = MODELS.build(bbox_head)
+
+ def _bbox_forward(self, x: Tuple[Tensor], rois: Tensor) -> dict:
+ """Box head forward function used in both training and testing.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+
+ Returns:
+ dict[str, Tensor]: Usually returns a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `cls_score_ref` (Tensor): The cls_score after refine model.
+ - `bbox_pred_ref` (Tensor): The bbox_pred after refine model.
+ - `bbox_feats` (Tensor): Extract bbox RoI features.
+ """
+ # TODO: a more flexible way to decide which feature maps to use
+ bbox_feats = self.bbox_roi_extractor(
+ x[:self.bbox_roi_extractor.num_inputs], rois)
+ bbox_results = self.bbox_head(bbox_feats)
+
+ if self.bbox_head.with_refine:
+ bbox_results = dict(
+ cls_score=bbox_results[0],
+ bbox_pred=bbox_results[1],
+ cls_score_ref=bbox_results[2],
+ bbox_pred_ref=bbox_results[3],
+ bbox_feats=bbox_feats)
+ else:
+ bbox_results = dict(
+ cls_score=bbox_results[0],
+ bbox_pred=bbox_results[1],
+ bbox_feats=bbox_feats)
+
+ return bbox_results
+
+ def bbox_loss(self, x: Tuple[Tensor],
+ sampling_results: List[SamplingResult]) -> dict:
+ """Perform forward propagation and loss calculation of the bbox head on
+ the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ sampling_results (list["obj:`SamplingResult`]): Sampling results.
+
+ Returns:
+ dict[str, Tensor]: Usually returns a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `bbox_feats` (Tensor): Extract bbox RoI features.
+ - `loss_bbox` (dict): A dictionary of bbox loss components.
+ """
+ rois = bbox2roi([res.priors for res in sampling_results])
+ bbox_results = self._bbox_forward(x, rois)
+
+ # If there is a refining process, add refine loss.
+ if 'cls_score_ref' in bbox_results:
+ bbox_loss_and_target = self.bbox_head.loss_and_target(
+ cls_score=bbox_results['cls_score'],
+ bbox_pred=bbox_results['bbox_pred'],
+ rois=rois,
+ sampling_results=sampling_results,
+ rcnn_train_cfg=self.train_cfg)
+ bbox_results.update(loss_bbox=bbox_loss_and_target['loss_bbox'])
+ bbox_loss_and_target_ref = self.bbox_head.loss_and_target(
+ cls_score=bbox_results['cls_score_ref'],
+ bbox_pred=bbox_results['bbox_pred_ref'],
+ rois=rois,
+ sampling_results=sampling_results,
+ rcnn_train_cfg=self.train_cfg)
+ bbox_results['loss_bbox']['loss_rcnn_emd_ref'] = \
+ bbox_loss_and_target_ref['loss_bbox']['loss_rcnn_emd']
+ else:
+ bbox_loss_and_target = self.bbox_head.loss_and_target(
+ cls_score=bbox_results['cls_score'],
+ bbox_pred=bbox_results['bbox_pred'],
+ rois=rois,
+ sampling_results=sampling_results,
+ rcnn_train_cfg=self.train_cfg)
+ bbox_results.update(loss_bbox=bbox_loss_and_target['loss_bbox'])
+
+ return bbox_results
+
+ def loss(self, x: Tuple[Tensor], rpn_results_list: InstanceList,
+ batch_data_samples: List[DetDataSample]) -> dict:
+ """Perform forward propagation and loss calculation of the detection
+ roi on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components
+ """
+ assert len(rpn_results_list) == len(batch_data_samples)
+ outputs = unpack_gt_instances(batch_data_samples)
+ batch_gt_instances, batch_gt_instances_ignore, _ = outputs
+
+ sampling_results = []
+ for i in range(len(batch_data_samples)):
+ # rename rpn_results.bboxes to rpn_results.priors
+ rpn_results = rpn_results_list[i]
+ rpn_results.priors = rpn_results.pop('bboxes')
+
+ assign_result = self.bbox_assigner.assign(
+ rpn_results, batch_gt_instances[i],
+ batch_gt_instances_ignore[i])
+ sampling_result = self.bbox_sampler.sample(
+ assign_result,
+ rpn_results,
+ batch_gt_instances[i],
+ batch_gt_instances_ignore=batch_gt_instances_ignore[i])
+ sampling_results.append(sampling_result)
+
+ losses = dict()
+ # bbox head loss
+ if self.with_bbox:
+ bbox_results = self.bbox_loss(x, sampling_results)
+ losses.update(bbox_results['loss_bbox'])
+
+ return losses
+
+ def predict_bbox(self,
+ x: Tuple[Tensor],
+ batch_img_metas: List[dict],
+ rpn_results_list: InstanceList,
+ rcnn_test_cfg: ConfigType,
+ rescale: bool = False) -> InstanceList:
+ """Perform forward propagation of the bbox head and predict detection
+ results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Feature maps of all scale level.
+ batch_img_metas (list[dict]): List of image information.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ rcnn_test_cfg (obj:`ConfigDict`): `test_cfg` of R-CNN.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ proposals = [res.bboxes for res in rpn_results_list]
+ rois = bbox2roi(proposals)
+
+ if rois.shape[0] == 0:
+ return empty_instances(
+ batch_img_metas, rois.device, task_type='bbox')
+
+ bbox_results = self._bbox_forward(x, rois)
+
+ # split batch bbox prediction back to each image
+ if 'cls_score_ref' in bbox_results:
+ cls_scores = bbox_results['cls_score_ref']
+ bbox_preds = bbox_results['bbox_pred_ref']
+ else:
+ cls_scores = bbox_results['cls_score']
+ bbox_preds = bbox_results['bbox_pred']
+ num_proposals_per_img = tuple(len(p) for p in proposals)
+ rois = rois.split(num_proposals_per_img, 0)
+ cls_scores = cls_scores.split(num_proposals_per_img, 0)
+
+ if bbox_preds is not None:
+ bbox_preds = bbox_preds.split(num_proposals_per_img, 0)
+ else:
+ bbox_preds = (None, ) * len(proposals)
+
+ result_list = self.bbox_head.predict_by_feat(
+ rois=rois,
+ cls_scores=cls_scores,
+ bbox_preds=bbox_preds,
+ batch_img_metas=batch_img_metas,
+ rcnn_test_cfg=rcnn_test_cfg,
+ rescale=rescale)
+ return result_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/pisa_roi_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/pisa_roi_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..45d59879da73b48df790c55d40a4a88f1d099111
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/pisa_roi_head.py
@@ -0,0 +1,148 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple
+
+from torch import Tensor
+
+from mmdet.models.task_modules import SamplingResult
+from mmdet.registry import MODELS
+from mmdet.structures import DetDataSample
+from mmdet.structures.bbox import bbox2roi
+from mmdet.utils import InstanceList
+from ..losses.pisa_loss import carl_loss, isr_p
+from ..utils import unpack_gt_instances
+from .standard_roi_head import StandardRoIHead
+
+
+@MODELS.register_module()
+class PISARoIHead(StandardRoIHead):
+ r"""The RoI head for `Prime Sample Attention in Object Detection
+ `_."""
+
+ def loss(self, x: Tuple[Tensor], rpn_results_list: InstanceList,
+ batch_data_samples: List[DetDataSample]) -> dict:
+ """Perform forward propagation and loss calculation of the detection
+ roi on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components
+ """
+ assert len(rpn_results_list) == len(batch_data_samples)
+ outputs = unpack_gt_instances(batch_data_samples)
+ batch_gt_instances, batch_gt_instances_ignore, _ = outputs
+
+ # assign gts and sample proposals
+ num_imgs = len(batch_data_samples)
+ sampling_results = []
+ neg_label_weights = []
+ for i in range(num_imgs):
+ # rename rpn_results.bboxes to rpn_results.priors
+ rpn_results = rpn_results_list[i]
+ rpn_results.priors = rpn_results.pop('bboxes')
+
+ assign_result = self.bbox_assigner.assign(
+ rpn_results, batch_gt_instances[i],
+ batch_gt_instances_ignore[i])
+ sampling_result = self.bbox_sampler.sample(
+ assign_result,
+ rpn_results,
+ batch_gt_instances[i],
+ feats=[lvl_feat[i][None] for lvl_feat in x])
+ if isinstance(sampling_result, tuple):
+ sampling_result, neg_label_weight = sampling_result
+ sampling_results.append(sampling_result)
+ neg_label_weights.append(neg_label_weight)
+
+ losses = dict()
+ # bbox head forward and loss
+ if self.with_bbox:
+ bbox_results = self.bbox_loss(
+ x, sampling_results, neg_label_weights=neg_label_weights)
+ losses.update(bbox_results['loss_bbox'])
+
+ # mask head forward and loss
+ if self.with_mask:
+ mask_results = self.mask_loss(x, sampling_results,
+ bbox_results['bbox_feats'],
+ batch_gt_instances)
+ losses.update(mask_results['loss_mask'])
+
+ return losses
+
+ def bbox_loss(self,
+ x: Tuple[Tensor],
+ sampling_results: List[SamplingResult],
+ neg_label_weights: List[Tensor] = None) -> dict:
+ """Perform forward propagation and loss calculation of the bbox head on
+ the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ sampling_results (list["obj:`SamplingResult`]): Sampling results.
+
+ Returns:
+ dict[str, Tensor]: Usually returns a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `bbox_feats` (Tensor): Extract bbox RoI features.
+ - `loss_bbox` (dict): A dictionary of bbox loss components.
+ """
+ rois = bbox2roi([res.priors for res in sampling_results])
+ bbox_results = self._bbox_forward(x, rois)
+ bbox_targets = self.bbox_head.get_targets(sampling_results,
+ self.train_cfg)
+
+ # neg_label_weights obtained by sampler is image-wise, mapping back to
+ # the corresponding location in label weights
+ if neg_label_weights[0] is not None:
+ label_weights = bbox_targets[1]
+ cur_num_rois = 0
+ for i in range(len(sampling_results)):
+ num_pos = sampling_results[i].pos_inds.size(0)
+ num_neg = sampling_results[i].neg_inds.size(0)
+ label_weights[cur_num_rois + num_pos:cur_num_rois + num_pos +
+ num_neg] = neg_label_weights[i]
+ cur_num_rois += num_pos + num_neg
+
+ cls_score = bbox_results['cls_score']
+ bbox_pred = bbox_results['bbox_pred']
+
+ # Apply ISR-P
+ isr_cfg = self.train_cfg.get('isr', None)
+ if isr_cfg is not None:
+ bbox_targets = isr_p(
+ cls_score,
+ bbox_pred,
+ bbox_targets,
+ rois,
+ sampling_results,
+ self.bbox_head.loss_cls,
+ self.bbox_head.bbox_coder,
+ **isr_cfg,
+ num_class=self.bbox_head.num_classes)
+ loss_bbox = self.bbox_head.loss(cls_score, bbox_pred, rois,
+ *bbox_targets)
+
+ # Add CARL Loss
+ carl_cfg = self.train_cfg.get('carl', None)
+ if carl_cfg is not None:
+ loss_carl = carl_loss(
+ cls_score,
+ bbox_targets[0],
+ bbox_pred,
+ bbox_targets[2],
+ self.bbox_head.loss_bbox,
+ **carl_cfg,
+ num_class=self.bbox_head.num_classes)
+ loss_bbox.update(loss_carl)
+
+ bbox_results.update(loss_bbox=loss_bbox)
+ return bbox_results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/point_rend_roi_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/point_rend_roi_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..6a0641549631e243c3db25039b01fed64fb1e0d1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/point_rend_roi_head.py
@@ -0,0 +1,236 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+# Modified from https://github.com/facebookresearch/detectron2/tree/master/projects/PointRend # noqa
+from typing import List, Tuple
+
+import torch
+import torch.nn.functional as F
+from mmcv.ops import point_sample, rel_roi_point_to_rel_img_point
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures.bbox import bbox2roi
+from mmdet.utils import ConfigType, InstanceList
+from ..task_modules.samplers import SamplingResult
+from ..utils import empty_instances
+from .standard_roi_head import StandardRoIHead
+
+
+@MODELS.register_module()
+class PointRendRoIHead(StandardRoIHead):
+ """`PointRend `_."""
+
+ def __init__(self, point_head: ConfigType, *args, **kwargs) -> None:
+ super().__init__(*args, **kwargs)
+ assert self.with_bbox and self.with_mask
+ self.init_point_head(point_head)
+
+ def init_point_head(self, point_head: ConfigType) -> None:
+ """Initialize ``point_head``"""
+ self.point_head = MODELS.build(point_head)
+
+ def mask_loss(self, x: Tuple[Tensor],
+ sampling_results: List[SamplingResult], bbox_feats: Tensor,
+ batch_gt_instances: InstanceList) -> dict:
+ """Run forward function and calculate loss for mask head and point head
+ in training."""
+ mask_results = super().mask_loss(
+ x=x,
+ sampling_results=sampling_results,
+ bbox_feats=bbox_feats,
+ batch_gt_instances=batch_gt_instances)
+
+ mask_point_results = self._mask_point_loss(
+ x=x,
+ sampling_results=sampling_results,
+ mask_preds=mask_results['mask_preds'],
+ batch_gt_instances=batch_gt_instances)
+ mask_results['loss_mask'].update(
+ loss_point=mask_point_results['loss_point'])
+
+ return mask_results
+
+ def _mask_point_loss(self, x: Tuple[Tensor],
+ sampling_results: List[SamplingResult],
+ mask_preds: Tensor,
+ batch_gt_instances: InstanceList) -> dict:
+ """Run forward function and calculate loss for point head in
+ training."""
+ pos_labels = torch.cat([res.pos_gt_labels for res in sampling_results])
+ rel_roi_points = self.point_head.get_roi_rel_points_train(
+ mask_preds, pos_labels, cfg=self.train_cfg)
+ rois = bbox2roi([res.pos_bboxes for res in sampling_results])
+
+ fine_grained_point_feats = self._get_fine_grained_point_feats(
+ x, rois, rel_roi_points)
+ coarse_point_feats = point_sample(mask_preds, rel_roi_points)
+ mask_point_pred = self.point_head(fine_grained_point_feats,
+ coarse_point_feats)
+
+ loss_and_target = self.point_head.loss_and_target(
+ point_pred=mask_point_pred,
+ rel_roi_points=rel_roi_points,
+ sampling_results=sampling_results,
+ batch_gt_instances=batch_gt_instances,
+ cfg=self.train_cfg)
+
+ return loss_and_target
+
+ def _mask_point_forward_test(self, x: Tuple[Tensor], rois: Tensor,
+ label_preds: Tensor,
+ mask_preds: Tensor) -> Tensor:
+ """Mask refining process with point head in testing.
+
+ Args:
+ x (tuple[Tensor]): Feature maps of all scale level.
+ rois (Tensor): shape (num_rois, 5).
+ label_preds (Tensor): The predication class for each rois.
+ mask_preds (Tensor): The predication coarse masks of
+ shape (num_rois, num_classes, small_size, small_size).
+
+ Returns:
+ Tensor: The refined masks of shape (num_rois, num_classes,
+ large_size, large_size).
+ """
+ refined_mask_pred = mask_preds.clone()
+ for subdivision_step in range(self.test_cfg.subdivision_steps):
+ refined_mask_pred = F.interpolate(
+ refined_mask_pred,
+ scale_factor=self.test_cfg.scale_factor,
+ mode='bilinear',
+ align_corners=False)
+ # If `subdivision_num_points` is larger or equal to the
+ # resolution of the next step, then we can skip this step
+ num_rois, channels, mask_height, mask_width = \
+ refined_mask_pred.shape
+ if (self.test_cfg.subdivision_num_points >=
+ self.test_cfg.scale_factor**2 * mask_height * mask_width
+ and
+ subdivision_step < self.test_cfg.subdivision_steps - 1):
+ continue
+ point_indices, rel_roi_points = \
+ self.point_head.get_roi_rel_points_test(
+ refined_mask_pred, label_preds, cfg=self.test_cfg)
+
+ fine_grained_point_feats = self._get_fine_grained_point_feats(
+ x=x, rois=rois, rel_roi_points=rel_roi_points)
+ coarse_point_feats = point_sample(mask_preds, rel_roi_points)
+ mask_point_pred = self.point_head(fine_grained_point_feats,
+ coarse_point_feats)
+
+ point_indices = point_indices.unsqueeze(1).expand(-1, channels, -1)
+ refined_mask_pred = refined_mask_pred.reshape(
+ num_rois, channels, mask_height * mask_width)
+ refined_mask_pred = refined_mask_pred.scatter_(
+ 2, point_indices, mask_point_pred)
+ refined_mask_pred = refined_mask_pred.view(num_rois, channels,
+ mask_height, mask_width)
+
+ return refined_mask_pred
+
+ def _get_fine_grained_point_feats(self, x: Tuple[Tensor], rois: Tensor,
+ rel_roi_points: Tensor) -> Tensor:
+ """Sample fine grained feats from each level feature map and
+ concatenate them together.
+
+ Args:
+ x (tuple[Tensor]): Feature maps of all scale level.
+ rois (Tensor): shape (num_rois, 5).
+ rel_roi_points (Tensor): A tensor of shape (num_rois, num_points,
+ 2) that contains [0, 1] x [0, 1] normalized coordinates of the
+ most uncertain points from the [mask_height, mask_width] grid.
+
+ Returns:
+ Tensor: The fine grained features for each points,
+ has shape (num_rois, feats_channels, num_points).
+ """
+ assert rois.shape[0] > 0, 'RoI is a empty tensor.'
+ num_imgs = x[0].shape[0]
+ fine_grained_feats = []
+ for idx in range(self.mask_roi_extractor.num_inputs):
+ feats = x[idx]
+ spatial_scale = 1. / float(
+ self.mask_roi_extractor.featmap_strides[idx])
+ point_feats = []
+ for batch_ind in range(num_imgs):
+ # unravel batch dim
+ feat = feats[batch_ind].unsqueeze(0)
+ inds = (rois[:, 0].long() == batch_ind)
+ if inds.any():
+ rel_img_points = rel_roi_point_to_rel_img_point(
+ rois=rois[inds],
+ rel_roi_points=rel_roi_points[inds],
+ img=feat.shape[2:],
+ spatial_scale=spatial_scale).unsqueeze(0)
+ point_feat = point_sample(feat, rel_img_points)
+ point_feat = point_feat.squeeze(0).transpose(0, 1)
+ point_feats.append(point_feat)
+ fine_grained_feats.append(torch.cat(point_feats, dim=0))
+ return torch.cat(fine_grained_feats, dim=1)
+
+ def predict_mask(self,
+ x: Tuple[Tensor],
+ batch_img_metas: List[dict],
+ results_list: InstanceList,
+ rescale: bool = False) -> InstanceList:
+ """Perform forward propagation of the mask head and predict detection
+ results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Feature maps of all scale level.
+ batch_img_metas (list[dict]): List of image information.
+ results_list (list[:obj:`InstanceData`]): Detection results of
+ each image.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+ """
+ # don't need to consider aug_test.
+ bboxes = [res.bboxes for res in results_list]
+ mask_rois = bbox2roi(bboxes)
+ if mask_rois.shape[0] == 0:
+ results_list = empty_instances(
+ batch_img_metas,
+ mask_rois.device,
+ task_type='mask',
+ instance_results=results_list,
+ mask_thr_binary=self.test_cfg.mask_thr_binary)
+ return results_list
+
+ mask_results = self._mask_forward(x, mask_rois)
+ mask_preds = mask_results['mask_preds']
+ # split batch mask prediction back to each image
+ num_mask_rois_per_img = [len(res) for res in results_list]
+ mask_preds = mask_preds.split(num_mask_rois_per_img, 0)
+
+ # refine mask_preds
+ mask_rois = mask_rois.split(num_mask_rois_per_img, 0)
+ mask_preds_refined = []
+ for i in range(len(batch_img_metas)):
+ labels = results_list[i].labels
+ x_i = [xx[[i]] for xx in x]
+ mask_rois_i = mask_rois[i]
+ mask_rois_i[:, 0] = 0
+ mask_pred_i = self._mask_point_forward_test(
+ x_i, mask_rois_i, labels, mask_preds[i])
+ mask_preds_refined.append(mask_pred_i)
+
+ # TODO: Handle the case where rescale is false
+ results_list = self.mask_head.predict_by_feat(
+ mask_preds=mask_preds_refined,
+ results_list=results_list,
+ batch_img_metas=batch_img_metas,
+ rcnn_test_cfg=self.test_cfg,
+ rescale=rescale)
+ return results_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/roi_extractors/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/roi_extractors/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..0f60214991b0ed14cdbc3964aee15356c6aaf2aa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/roi_extractors/__init__.py
@@ -0,0 +1,6 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .base_roi_extractor import BaseRoIExtractor
+from .generic_roi_extractor import GenericRoIExtractor
+from .single_level_roi_extractor import SingleRoIExtractor
+
+__all__ = ['BaseRoIExtractor', 'SingleRoIExtractor', 'GenericRoIExtractor']
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/roi_extractors/base_roi_extractor.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/roi_extractors/base_roi_extractor.py
new file mode 100644
index 0000000000000000000000000000000000000000..a8de0518818aba8d9aac7b807e3215d0da6c9b99
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/roi_extractors/base_roi_extractor.py
@@ -0,0 +1,111 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from abc import ABCMeta, abstractmethod
+from typing import List, Optional, Tuple
+
+import torch
+import torch.nn as nn
+from mmcv import ops
+from mmengine.model import BaseModule
+from torch import Tensor
+
+from mmdet.utils import ConfigType, OptMultiConfig
+
+
+class BaseRoIExtractor(BaseModule, metaclass=ABCMeta):
+ """Base class for RoI extractor.
+
+ Args:
+ roi_layer (:obj:`ConfigDict` or dict): Specify RoI layer type and
+ arguments.
+ out_channels (int): Output channels of RoI layers.
+ featmap_strides (list[int]): Strides of input feature maps.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict], optional): Initialization config dict. Defaults to None.
+ """
+
+ def __init__(self,
+ roi_layer: ConfigType,
+ out_channels: int,
+ featmap_strides: List[int],
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.roi_layers = self.build_roi_layers(roi_layer, featmap_strides)
+ self.out_channels = out_channels
+ self.featmap_strides = featmap_strides
+
+ @property
+ def num_inputs(self) -> int:
+ """int: Number of input feature maps."""
+ return len(self.featmap_strides)
+
+ def build_roi_layers(self, layer_cfg: ConfigType,
+ featmap_strides: List[int]) -> nn.ModuleList:
+ """Build RoI operator to extract feature from each level feature map.
+
+ Args:
+ layer_cfg (:obj:`ConfigDict` or dict): Dictionary to construct and
+ config RoI layer operation. Options are modules under
+ ``mmcv/ops`` such as ``RoIAlign``.
+ featmap_strides (list[int]): The stride of input feature map w.r.t
+ to the original image size, which would be used to scale RoI
+ coordinate (original image coordinate system) to feature
+ coordinate system.
+
+ Returns:
+ :obj:`nn.ModuleList`: The RoI extractor modules for each level
+ feature map.
+ """
+
+ cfg = layer_cfg.copy()
+ layer_type = cfg.pop('type')
+ if isinstance(layer_type, str):
+ assert hasattr(ops, layer_type)
+ layer_cls = getattr(ops, layer_type)
+ else:
+ layer_cls = layer_type
+ roi_layers = nn.ModuleList(
+ [layer_cls(spatial_scale=1 / s, **cfg) for s in featmap_strides])
+ return roi_layers
+
+ def roi_rescale(self, rois: Tensor, scale_factor: float) -> Tensor:
+ """Scale RoI coordinates by scale factor.
+
+ Args:
+ rois (Tensor): RoI (Region of Interest), shape (n, 5)
+ scale_factor (float): Scale factor that RoI will be multiplied by.
+
+ Returns:
+ Tensor: Scaled RoI.
+ """
+
+ cx = (rois[:, 1] + rois[:, 3]) * 0.5
+ cy = (rois[:, 2] + rois[:, 4]) * 0.5
+ w = rois[:, 3] - rois[:, 1]
+ h = rois[:, 4] - rois[:, 2]
+ new_w = w * scale_factor
+ new_h = h * scale_factor
+ x1 = cx - new_w * 0.5
+ x2 = cx + new_w * 0.5
+ y1 = cy - new_h * 0.5
+ y2 = cy + new_h * 0.5
+ new_rois = torch.stack((rois[:, 0], x1, y1, x2, y2), dim=-1)
+ return new_rois
+
+ @abstractmethod
+ def forward(self,
+ feats: Tuple[Tensor],
+ rois: Tensor,
+ roi_scale_factor: Optional[float] = None) -> Tensor:
+ """Extractor ROI feats.
+
+ Args:
+ feats (Tuple[Tensor]): Multi-scale features.
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+ roi_scale_factor (Optional[float]): RoI scale factor.
+ Defaults to None.
+
+ Returns:
+ Tensor: RoI feature.
+ """
+ pass
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/roi_extractors/generic_roi_extractor.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/roi_extractors/generic_roi_extractor.py
new file mode 100644
index 0000000000000000000000000000000000000000..39d4c90135d853404d564391f029558841ac9cac
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/roi_extractors/generic_roi_extractor.py
@@ -0,0 +1,102 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Tuple
+
+from mmcv.cnn.bricks import build_plugin_layer
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import OptConfigType
+from .base_roi_extractor import BaseRoIExtractor
+
+
+@MODELS.register_module()
+class GenericRoIExtractor(BaseRoIExtractor):
+ """Extract RoI features from all level feature maps levels.
+
+ This is the implementation of `A novel Region of Interest Extraction Layer
+ for Instance Segmentation `_.
+
+ Args:
+ aggregation (str): The method to aggregate multiple feature maps.
+ Options are 'sum', 'concat'. Defaults to 'sum'.
+ pre_cfg (:obj:`ConfigDict` or dict): Specify pre-processing modules.
+ Defaults to None.
+ post_cfg (:obj:`ConfigDict` or dict): Specify post-processing modules.
+ Defaults to None.
+ kwargs (keyword arguments): Arguments that are the same
+ as :class:`BaseRoIExtractor`.
+ """
+
+ def __init__(self,
+ aggregation: str = 'sum',
+ pre_cfg: OptConfigType = None,
+ post_cfg: OptConfigType = None,
+ **kwargs) -> None:
+ super().__init__(**kwargs)
+
+ assert aggregation in ['sum', 'concat']
+
+ self.aggregation = aggregation
+ self.with_post = post_cfg is not None
+ self.with_pre = pre_cfg is not None
+ # build pre/post processing modules
+ if self.with_post:
+ self.post_module = build_plugin_layer(post_cfg, '_post_module')[1]
+ if self.with_pre:
+ self.pre_module = build_plugin_layer(pre_cfg, '_pre_module')[1]
+
+ def forward(self,
+ feats: Tuple[Tensor],
+ rois: Tensor,
+ roi_scale_factor: Optional[float] = None) -> Tensor:
+ """Extractor ROI feats.
+
+ Args:
+ feats (Tuple[Tensor]): Multi-scale features.
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+ roi_scale_factor (Optional[float]): RoI scale factor.
+ Defaults to None.
+
+ Returns:
+ Tensor: RoI feature.
+ """
+ out_size = self.roi_layers[0].output_size
+ num_levels = len(feats)
+ roi_feats = feats[0].new_zeros(
+ rois.size(0), self.out_channels, *out_size)
+
+ # some times rois is an empty tensor
+ if roi_feats.shape[0] == 0:
+ return roi_feats
+
+ if num_levels == 1:
+ return self.roi_layers[0](feats[0], rois)
+
+ if roi_scale_factor is not None:
+ rois = self.roi_rescale(rois, roi_scale_factor)
+
+ # mark the starting channels for concat mode
+ start_channels = 0
+ for i in range(num_levels):
+ roi_feats_t = self.roi_layers[i](feats[i], rois)
+ end_channels = start_channels + roi_feats_t.size(1)
+ if self.with_pre:
+ # apply pre-processing to a RoI extracted from each layer
+ roi_feats_t = self.pre_module(roi_feats_t)
+ if self.aggregation == 'sum':
+ # and sum them all
+ roi_feats += roi_feats_t
+ else:
+ # and concat them along channel dimension
+ roi_feats[:, start_channels:end_channels] = roi_feats_t
+ # update channels starting position
+ start_channels = end_channels
+ # check if concat channels match at the end
+ if self.aggregation == 'concat':
+ assert start_channels == self.out_channels
+
+ if self.with_post:
+ # apply post-processing before return the result
+ roi_feats = self.post_module(roi_feats)
+ return roi_feats
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/roi_extractors/single_level_roi_extractor.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/roi_extractors/single_level_roi_extractor.py
new file mode 100644
index 0000000000000000000000000000000000000000..59229e0b0b0a18dff81abca6f5c20cb50b0d542c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/roi_extractors/single_level_roi_extractor.py
@@ -0,0 +1,119 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple
+
+import torch
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.utils import ConfigType, OptMultiConfig
+from .base_roi_extractor import BaseRoIExtractor
+
+
+@MODELS.register_module()
+class SingleRoIExtractor(BaseRoIExtractor):
+ """Extract RoI features from a single level feature map.
+
+ If there are multiple input feature levels, each RoI is mapped to a level
+ according to its scale. The mapping rule is proposed in
+ `FPN `_.
+
+ Args:
+ roi_layer (:obj:`ConfigDict` or dict): Specify RoI layer type and
+ arguments.
+ out_channels (int): Output channels of RoI layers.
+ featmap_strides (List[int]): Strides of input feature maps.
+ finest_scale (int): Scale threshold of mapping to level 0.
+ Defaults to 56.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict], optional): Initialization config dict. Defaults to None.
+ """
+
+ def __init__(self,
+ roi_layer: ConfigType,
+ out_channels: int,
+ featmap_strides: List[int],
+ finest_scale: int = 56,
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(
+ roi_layer=roi_layer,
+ out_channels=out_channels,
+ featmap_strides=featmap_strides,
+ init_cfg=init_cfg)
+ self.finest_scale = finest_scale
+
+ def map_roi_levels(self, rois: Tensor, num_levels: int) -> Tensor:
+ """Map rois to corresponding feature levels by scales.
+
+ - scale < finest_scale * 2: level 0
+ - finest_scale * 2 <= scale < finest_scale * 4: level 1
+ - finest_scale * 4 <= scale < finest_scale * 8: level 2
+ - scale >= finest_scale * 8: level 3
+
+ Args:
+ rois (Tensor): Input RoIs, shape (k, 5).
+ num_levels (int): Total level number.
+
+ Returns:
+ Tensor: Level index (0-based) of each RoI, shape (k, )
+ """
+ scale = torch.sqrt(
+ (rois[:, 3] - rois[:, 1]) * (rois[:, 4] - rois[:, 2]))
+ target_lvls = torch.floor(torch.log2(scale / self.finest_scale + 1e-6))
+ target_lvls = target_lvls.clamp(min=0, max=num_levels - 1).long()
+ return target_lvls
+
+ def forward(self,
+ feats: Tuple[Tensor],
+ rois: Tensor,
+ roi_scale_factor: Optional[float] = None):
+ """Extractor ROI feats.
+
+ Args:
+ feats (Tuple[Tensor]): Multi-scale features.
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+ roi_scale_factor (Optional[float]): RoI scale factor.
+ Defaults to None.
+
+ Returns:
+ Tensor: RoI feature.
+ """
+ # convert fp32 to fp16 when amp is on
+ rois = rois.type_as(feats[0])
+ out_size = self.roi_layers[0].output_size
+ num_levels = len(feats)
+ roi_feats = feats[0].new_zeros(
+ rois.size(0), self.out_channels, *out_size)
+
+ # TODO: remove this when parrots supports
+ if torch.__version__ == 'parrots':
+ roi_feats.requires_grad = True
+
+ if num_levels == 1:
+ if len(rois) == 0:
+ return roi_feats
+ return self.roi_layers[0](feats[0], rois)
+
+ target_lvls = self.map_roi_levels(rois, num_levels)
+
+ if roi_scale_factor is not None:
+ rois = self.roi_rescale(rois, roi_scale_factor)
+
+ for i in range(num_levels):
+ mask = target_lvls == i
+ inds = mask.nonzero(as_tuple=False).squeeze(1)
+ if inds.numel() > 0:
+ rois_ = rois[inds]
+ roi_feats_t = self.roi_layers[i](feats[i], rois_)
+ roi_feats[inds] = roi_feats_t
+ else:
+ # Sometimes some pyramid levels will not be used for RoI
+ # feature extraction and this will cause an incomplete
+ # computation graph in one GPU, which is different from those
+ # in other GPUs and will cause a hanging error.
+ # Therefore, we add it to ensure each feature pyramid is
+ # included in the computation graph to avoid runtime bugs.
+ roi_feats += sum(
+ x.view(-1)[0]
+ for x in self.parameters()) * 0. + feats[i].sum() * 0.
+ return roi_feats
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/scnet_roi_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/scnet_roi_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..e6d2bc1915bae38011cc75a720e48ed53b51ddb5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/scnet_roi_head.py
@@ -0,0 +1,677 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple
+
+import torch
+import torch.nn.functional as F
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import bbox2roi
+from mmdet.utils import ConfigType, InstanceList, OptConfigType
+from ..layers import adaptive_avg_pool2d
+from ..task_modules.samplers import SamplingResult
+from ..utils import empty_instances, unpack_gt_instances
+from .cascade_roi_head import CascadeRoIHead
+
+
+@MODELS.register_module()
+class SCNetRoIHead(CascadeRoIHead):
+ """RoIHead for `SCNet `_.
+
+ Args:
+ num_stages (int): number of cascade stages.
+ stage_loss_weights (list): loss weight of cascade stages.
+ semantic_roi_extractor (dict): config to init semantic roi extractor.
+ semantic_head (dict): config to init semantic head.
+ feat_relay_head (dict): config to init feature_relay_head.
+ glbctx_head (dict): config to init global context head.
+ """
+
+ def __init__(self,
+ num_stages: int,
+ stage_loss_weights: List[float],
+ semantic_roi_extractor: OptConfigType = None,
+ semantic_head: OptConfigType = None,
+ feat_relay_head: OptConfigType = None,
+ glbctx_head: OptConfigType = None,
+ **kwargs) -> None:
+ super().__init__(
+ num_stages=num_stages,
+ stage_loss_weights=stage_loss_weights,
+ **kwargs)
+ assert self.with_bbox and self.with_mask
+ assert not self.with_shared_head # shared head is not supported
+
+ if semantic_head is not None:
+ self.semantic_roi_extractor = MODELS.build(semantic_roi_extractor)
+ self.semantic_head = MODELS.build(semantic_head)
+
+ if feat_relay_head is not None:
+ self.feat_relay_head = MODELS.build(feat_relay_head)
+
+ if glbctx_head is not None:
+ self.glbctx_head = MODELS.build(glbctx_head)
+
+ def init_mask_head(self, mask_roi_extractor: ConfigType,
+ mask_head: ConfigType) -> None:
+ """Initialize ``mask_head``"""
+ if mask_roi_extractor is not None:
+ self.mask_roi_extractor = MODELS.build(mask_roi_extractor)
+ self.mask_head = MODELS.build(mask_head)
+
+ # TODO move to base_roi_head later
+ @property
+ def with_semantic(self) -> bool:
+ """bool: whether the head has semantic head"""
+ return hasattr(self,
+ 'semantic_head') and self.semantic_head is not None
+
+ @property
+ def with_feat_relay(self) -> bool:
+ """bool: whether the head has feature relay head"""
+ return (hasattr(self, 'feat_relay_head')
+ and self.feat_relay_head is not None)
+
+ @property
+ def with_glbctx(self) -> bool:
+ """bool: whether the head has global context head"""
+ return hasattr(self, 'glbctx_head') and self.glbctx_head is not None
+
+ def _fuse_glbctx(self, roi_feats: Tensor, glbctx_feat: Tensor,
+ rois: Tensor) -> Tensor:
+ """Fuse global context feats with roi feats.
+
+ Args:
+ roi_feats (Tensor): RoI features.
+ glbctx_feat (Tensor): Global context feature..
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+
+ Returns:
+ Tensor: Fused feature.
+ """
+ assert roi_feats.size(0) == rois.size(0)
+ # RuntimeError: isDifferentiableType(variable.scalar_type())
+ # INTERNAL ASSERT FAILED if detach() is not used when calling
+ # roi_head.predict().
+ img_inds = torch.unique(rois[:, 0].detach().cpu(), sorted=True).long()
+ fused_feats = torch.zeros_like(roi_feats)
+ for img_id in img_inds:
+ inds = (rois[:, 0] == img_id.item())
+ fused_feats[inds] = roi_feats[inds] + glbctx_feat[img_id]
+ return fused_feats
+
+ def _slice_pos_feats(self, feats: Tensor,
+ sampling_results: List[SamplingResult]) -> Tensor:
+ """Get features from pos rois.
+
+ Args:
+ feats (Tensor): Input features.
+ sampling_results (list["obj:`SamplingResult`]): Sampling results.
+
+ Returns:
+ Tensor: Sliced features.
+ """
+ num_rois = [res.priors.size(0) for res in sampling_results]
+ num_pos_rois = [res.pos_priors.size(0) for res in sampling_results]
+ inds = torch.zeros(sum(num_rois), dtype=torch.bool)
+ start = 0
+ for i in range(len(num_rois)):
+ start = 0 if i == 0 else start + num_rois[i - 1]
+ stop = start + num_pos_rois[i]
+ inds[start:stop] = 1
+ sliced_feats = feats[inds]
+ return sliced_feats
+
+ def _bbox_forward(self,
+ stage: int,
+ x: Tuple[Tensor],
+ rois: Tensor,
+ semantic_feat: Optional[Tensor] = None,
+ glbctx_feat: Optional[Tensor] = None) -> dict:
+ """Box head forward function used in both training and testing.
+
+ Args:
+ stage (int): The current stage in Cascade RoI Head.
+ x (tuple[Tensor]): List of multi-level img features.
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+ semantic_feat (Tensor): Semantic feature. Defaults to None.
+ glbctx_feat (Tensor): Global context feature. Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: Usually returns a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `bbox_feats` (Tensor): Extract bbox RoI features.
+ """
+ bbox_roi_extractor = self.bbox_roi_extractor[stage]
+ bbox_head = self.bbox_head[stage]
+ bbox_feats = bbox_roi_extractor(x[:bbox_roi_extractor.num_inputs],
+ rois)
+ if self.with_semantic and semantic_feat is not None:
+ bbox_semantic_feat = self.semantic_roi_extractor([semantic_feat],
+ rois)
+ if bbox_semantic_feat.shape[-2:] != bbox_feats.shape[-2:]:
+ bbox_semantic_feat = adaptive_avg_pool2d(
+ bbox_semantic_feat, bbox_feats.shape[-2:])
+ bbox_feats += bbox_semantic_feat
+ if self.with_glbctx and glbctx_feat is not None:
+ bbox_feats = self._fuse_glbctx(bbox_feats, glbctx_feat, rois)
+ cls_score, bbox_pred, relayed_feat = bbox_head(
+ bbox_feats, return_shared_feat=True)
+
+ bbox_results = dict(
+ cls_score=cls_score,
+ bbox_pred=bbox_pred,
+ relayed_feat=relayed_feat)
+ return bbox_results
+
+ def _mask_forward(self,
+ x: Tuple[Tensor],
+ rois: Tensor,
+ semantic_feat: Optional[Tensor] = None,
+ glbctx_feat: Optional[Tensor] = None,
+ relayed_feat: Optional[Tensor] = None) -> dict:
+ """Mask head forward function used in both training and testing.
+
+ Args:
+ stage (int): The current stage in Cascade RoI Head.
+ x (tuple[Tensor]): Tuple of multi-level img features.
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+ semantic_feat (Tensor): Semantic feature. Defaults to None.
+ glbctx_feat (Tensor): Global context feature. Defaults to None.
+ relayed_feat (Tensor): Relayed feature. Defaults to None.
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `mask_preds` (Tensor): Mask prediction.
+ """
+ mask_feats = self.mask_roi_extractor(
+ x[:self.mask_roi_extractor.num_inputs], rois)
+ if self.with_semantic and semantic_feat is not None:
+ mask_semantic_feat = self.semantic_roi_extractor([semantic_feat],
+ rois)
+ if mask_semantic_feat.shape[-2:] != mask_feats.shape[-2:]:
+ mask_semantic_feat = F.adaptive_avg_pool2d(
+ mask_semantic_feat, mask_feats.shape[-2:])
+ mask_feats += mask_semantic_feat
+ if self.with_glbctx and glbctx_feat is not None:
+ mask_feats = self._fuse_glbctx(mask_feats, glbctx_feat, rois)
+ if self.with_feat_relay and relayed_feat is not None:
+ mask_feats = mask_feats + relayed_feat
+ mask_preds = self.mask_head(mask_feats)
+ mask_results = dict(mask_preds=mask_preds)
+
+ return mask_results
+
+ def bbox_loss(self,
+ stage: int,
+ x: Tuple[Tensor],
+ sampling_results: List[SamplingResult],
+ semantic_feat: Optional[Tensor] = None,
+ glbctx_feat: Optional[Tensor] = None) -> dict:
+ """Run forward function and calculate loss for box head in training.
+
+ Args:
+ stage (int): The current stage in Cascade RoI Head.
+ x (tuple[Tensor]): List of multi-level img features.
+ sampling_results (list["obj:`SamplingResult`]): Sampling results.
+ semantic_feat (Tensor): Semantic feature. Defaults to None.
+ glbctx_feat (Tensor): Global context feature. Defaults to None.
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `bbox_feats` (Tensor): Extract bbox RoI features.
+ - `loss_bbox` (dict): A dictionary of bbox loss components.
+ - `rois` (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+ - `bbox_targets` (tuple): Ground truth for proposals in a
+ single image. Containing the following list of Tensors:
+ (labels, label_weights, bbox_targets, bbox_weights)
+ """
+ bbox_head = self.bbox_head[stage]
+ rois = bbox2roi([res.priors for res in sampling_results])
+ bbox_results = self._bbox_forward(
+ stage,
+ x,
+ rois,
+ semantic_feat=semantic_feat,
+ glbctx_feat=glbctx_feat)
+ bbox_results.update(rois=rois)
+
+ bbox_loss_and_target = bbox_head.loss_and_target(
+ cls_score=bbox_results['cls_score'],
+ bbox_pred=bbox_results['bbox_pred'],
+ rois=rois,
+ sampling_results=sampling_results,
+ rcnn_train_cfg=self.train_cfg[stage])
+
+ bbox_results.update(bbox_loss_and_target)
+ return bbox_results
+
+ def mask_loss(self,
+ x: Tuple[Tensor],
+ sampling_results: List[SamplingResult],
+ batch_gt_instances: InstanceList,
+ semantic_feat: Optional[Tensor] = None,
+ glbctx_feat: Optional[Tensor] = None,
+ relayed_feat: Optional[Tensor] = None) -> dict:
+ """Run forward function and calculate loss for mask head in training.
+
+ Args:
+ x (tuple[Tensor]): Tuple of multi-level img features.
+ sampling_results (list["obj:`SamplingResult`]): Sampling results.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``labels``, and
+ ``masks`` attributes.
+ semantic_feat (Tensor): Semantic feature. Defaults to None.
+ glbctx_feat (Tensor): Global context feature. Defaults to None.
+ relayed_feat (Tensor): Relayed feature. Defaults to None.
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `mask_preds` (Tensor): Mask prediction.
+ - `loss_mask` (dict): A dictionary of mask loss components.
+ """
+ pos_rois = bbox2roi([res.pos_priors for res in sampling_results])
+ mask_results = self._mask_forward(
+ x,
+ pos_rois,
+ semantic_feat=semantic_feat,
+ glbctx_feat=glbctx_feat,
+ relayed_feat=relayed_feat)
+
+ mask_loss_and_target = self.mask_head.loss_and_target(
+ mask_preds=mask_results['mask_preds'],
+ sampling_results=sampling_results,
+ batch_gt_instances=batch_gt_instances,
+ rcnn_train_cfg=self.train_cfg[-1])
+ mask_results.update(mask_loss_and_target)
+
+ return mask_results
+
+ def semantic_loss(self, x: Tuple[Tensor],
+ batch_data_samples: SampleList) -> dict:
+ """Semantic segmentation loss.
+
+ Args:
+ x (Tuple[Tensor]): Tuple of multi-level img features.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `semantic_feat` (Tensor): Semantic feature.
+ - `loss_seg` (dict): Semantic segmentation loss.
+ """
+ gt_semantic_segs = [
+ data_sample.gt_sem_seg.sem_seg
+ for data_sample in batch_data_samples
+ ]
+ gt_semantic_segs = torch.stack(gt_semantic_segs)
+ semantic_pred, semantic_feat = self.semantic_head(x)
+ loss_seg = self.semantic_head.loss(semantic_pred, gt_semantic_segs)
+
+ semantic_results = dict(loss_seg=loss_seg, semantic_feat=semantic_feat)
+
+ return semantic_results
+
+ def global_context_loss(self, x: Tuple[Tensor],
+ batch_gt_instances: InstanceList) -> dict:
+ """Global context loss.
+
+ Args:
+ x (Tuple[Tensor]): Tuple of multi-level img features.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``labels``, and
+ ``masks`` attributes.
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `glbctx_feat` (Tensor): Global context feature.
+ - `loss_glbctx` (dict): Global context loss.
+ """
+ gt_labels = [
+ gt_instances.labels for gt_instances in batch_gt_instances
+ ]
+ mc_pred, glbctx_feat = self.glbctx_head(x)
+ loss_glbctx = self.glbctx_head.loss(mc_pred, gt_labels)
+ global_context_results = dict(
+ loss_glbctx=loss_glbctx, glbctx_feat=glbctx_feat)
+
+ return global_context_results
+
+ def loss(self, x: Tensor, rpn_results_list: InstanceList,
+ batch_data_samples: SampleList) -> dict:
+ """Perform forward propagation and loss calculation of the detection
+ roi on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components
+ """
+ assert len(rpn_results_list) == len(batch_data_samples)
+ outputs = unpack_gt_instances(batch_data_samples)
+ batch_gt_instances, batch_gt_instances_ignore, batch_img_metas \
+ = outputs
+
+ losses = dict()
+
+ # semantic segmentation branch
+ if self.with_semantic:
+ semantic_results = self.semantic_loss(
+ x=x, batch_data_samples=batch_data_samples)
+ losses['loss_semantic_seg'] = semantic_results['loss_seg']
+ semantic_feat = semantic_results['semantic_feat']
+ else:
+ semantic_feat = None
+
+ # global context branch
+ if self.with_glbctx:
+ global_context_results = self.global_context_loss(
+ x=x, batch_gt_instances=batch_gt_instances)
+ losses['loss_glbctx'] = global_context_results['loss_glbctx']
+ glbctx_feat = global_context_results['glbctx_feat']
+ else:
+ glbctx_feat = None
+
+ results_list = rpn_results_list
+ num_imgs = len(batch_img_metas)
+ for stage in range(self.num_stages):
+ stage_loss_weight = self.stage_loss_weights[stage]
+
+ # assign gts and sample proposals
+ sampling_results = []
+ bbox_assigner = self.bbox_assigner[stage]
+ bbox_sampler = self.bbox_sampler[stage]
+ for i in range(num_imgs):
+ results = results_list[i]
+ # rename rpn_results.bboxes to rpn_results.priors
+ results.priors = results.pop('bboxes')
+
+ assign_result = bbox_assigner.assign(
+ results, batch_gt_instances[i],
+ batch_gt_instances_ignore[i])
+ sampling_result = bbox_sampler.sample(
+ assign_result,
+ results,
+ batch_gt_instances[i],
+ feats=[lvl_feat[i][None] for lvl_feat in x])
+ sampling_results.append(sampling_result)
+
+ # bbox head forward and loss
+ bbox_results = self.bbox_loss(
+ stage=stage,
+ x=x,
+ sampling_results=sampling_results,
+ semantic_feat=semantic_feat,
+ glbctx_feat=glbctx_feat)
+
+ for name, value in bbox_results['loss_bbox'].items():
+ losses[f's{stage}.{name}'] = (
+ value * stage_loss_weight if 'loss' in name else value)
+
+ # refine bboxes
+ if stage < self.num_stages - 1:
+ bbox_head = self.bbox_head[stage]
+ with torch.no_grad():
+ results_list = bbox_head.refine_bboxes(
+ sampling_results=sampling_results,
+ bbox_results=bbox_results,
+ batch_img_metas=batch_img_metas)
+
+ if self.with_feat_relay:
+ relayed_feat = self._slice_pos_feats(bbox_results['relayed_feat'],
+ sampling_results)
+ relayed_feat = self.feat_relay_head(relayed_feat)
+ else:
+ relayed_feat = None
+
+ # mask head forward and loss
+ mask_results = self.mask_loss(
+ x=x,
+ sampling_results=sampling_results,
+ batch_gt_instances=batch_gt_instances,
+ semantic_feat=semantic_feat,
+ glbctx_feat=glbctx_feat,
+ relayed_feat=relayed_feat)
+ mask_stage_loss_weight = sum(self.stage_loss_weights)
+ losses['loss_mask'] = mask_stage_loss_weight * mask_results[
+ 'loss_mask']['loss_mask']
+
+ return losses
+
+ def predict(self,
+ x: Tuple[Tensor],
+ rpn_results_list: InstanceList,
+ batch_data_samples: SampleList,
+ rescale: bool = False) -> InstanceList:
+ """Perform forward propagation of the roi head and predict detection
+ results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from upstream network. Each
+ has shape (N, C, H, W).
+ rpn_results_list (list[:obj:`InstanceData`]): list of region
+ proposals.
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool): Whether to rescale the results to
+ the original image. Defaults to False.
+
+ Returns:
+ list[obj:`InstanceData`]: Detection results of each image.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+ """
+ assert self.with_bbox, 'Bbox head must be implemented.'
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+
+ if self.with_semantic:
+ _, semantic_feat = self.semantic_head(x)
+ else:
+ semantic_feat = None
+
+ if self.with_glbctx:
+ _, glbctx_feat = self.glbctx_head(x)
+ else:
+ glbctx_feat = None
+
+ # TODO: nms_op in mmcv need be enhanced, the bbox result may get
+ # difference when not rescale in bbox_head
+
+ # If it has the mask branch, the bbox branch does not need
+ # to be scaled to the original image scale, because the mask
+ # branch will scale both bbox and mask at the same time.
+ bbox_rescale = rescale if not self.with_mask else False
+ results_list = self.predict_bbox(
+ x=x,
+ semantic_feat=semantic_feat,
+ glbctx_feat=glbctx_feat,
+ batch_img_metas=batch_img_metas,
+ rpn_results_list=rpn_results_list,
+ rcnn_test_cfg=self.test_cfg,
+ rescale=bbox_rescale)
+
+ if self.with_mask:
+ results_list = self.predict_mask(
+ x=x,
+ semantic_heat=semantic_feat,
+ glbctx_feat=glbctx_feat,
+ batch_img_metas=batch_img_metas,
+ results_list=results_list,
+ rescale=rescale)
+
+ return results_list
+
+ def predict_mask(self,
+ x: Tuple[Tensor],
+ semantic_heat: Tensor,
+ glbctx_feat: Tensor,
+ batch_img_metas: List[dict],
+ results_list: List[InstanceData],
+ rescale: bool = False) -> List[InstanceData]:
+ """Perform forward propagation of the mask head and predict detection
+ results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Feature maps of all scale level.
+ semantic_feat (Tensor): Semantic feature.
+ glbctx_feat (Tensor): Global context feature.
+ batch_img_metas (list[dict]): List of image information.
+ results_list (list[:obj:`InstanceData`]): Detection results of
+ each image.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+ """
+ bboxes = [res.bboxes for res in results_list]
+ mask_rois = bbox2roi(bboxes)
+ if mask_rois.shape[0] == 0:
+ results_list = empty_instances(
+ batch_img_metas=batch_img_metas,
+ device=mask_rois.device,
+ task_type='mask',
+ instance_results=results_list,
+ mask_thr_binary=self.test_cfg.mask_thr_binary)
+ return results_list
+
+ bboxes_results = self._bbox_forward(
+ stage=-1,
+ x=x,
+ rois=mask_rois,
+ semantic_feat=semantic_heat,
+ glbctx_feat=glbctx_feat)
+ relayed_feat = bboxes_results['relayed_feat']
+ relayed_feat = self.feat_relay_head(relayed_feat)
+
+ mask_results = self._mask_forward(
+ x=x,
+ rois=mask_rois,
+ semantic_feat=semantic_heat,
+ glbctx_feat=glbctx_feat,
+ relayed_feat=relayed_feat)
+ mask_preds = mask_results['mask_preds']
+
+ # split batch mask prediction back to each image
+ num_bbox_per_img = tuple(len(_bbox) for _bbox in bboxes)
+ mask_preds = mask_preds.split(num_bbox_per_img, 0)
+
+ results_list = self.mask_head.predict_by_feat(
+ mask_preds=mask_preds,
+ results_list=results_list,
+ batch_img_metas=batch_img_metas,
+ rcnn_test_cfg=self.test_cfg,
+ rescale=rescale)
+
+ return results_list
+
+ def forward(self, x: Tuple[Tensor], rpn_results_list: InstanceList,
+ batch_data_samples: SampleList) -> tuple:
+ """Network forward process. Usually includes backbone, neck and head
+ forward without any post-processing.
+
+ Args:
+ x (List[Tensor]): Multi-level features that may have different
+ resolutions.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ batch_data_samples (list[:obj:`DetDataSample`]): Each item contains
+ the meta information of each image and corresponding
+ annotations.
+
+ Returns
+ tuple: A tuple of features from ``bbox_head`` and ``mask_head``
+ forward.
+ """
+ results = ()
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+
+ if self.with_semantic:
+ _, semantic_feat = self.semantic_head(x)
+ else:
+ semantic_feat = None
+
+ if self.with_glbctx:
+ _, glbctx_feat = self.glbctx_head(x)
+ else:
+ glbctx_feat = None
+
+ proposals = [rpn_results.bboxes for rpn_results in rpn_results_list]
+ num_proposals_per_img = tuple(len(p) for p in proposals)
+ rois = bbox2roi(proposals)
+ # bbox head
+ if self.with_bbox:
+ rois, cls_scores, bbox_preds = self._refine_roi(
+ x=x,
+ rois=rois,
+ semantic_feat=semantic_feat,
+ glbctx_feat=glbctx_feat,
+ batch_img_metas=batch_img_metas,
+ num_proposals_per_img=num_proposals_per_img)
+ results = results + (cls_scores, bbox_preds)
+ # mask head
+ if self.with_mask:
+ rois = torch.cat(rois)
+ bboxes_results = self._bbox_forward(
+ stage=-1,
+ x=x,
+ rois=rois,
+ semantic_feat=semantic_feat,
+ glbctx_feat=glbctx_feat)
+ relayed_feat = bboxes_results['relayed_feat']
+ relayed_feat = self.feat_relay_head(relayed_feat)
+ mask_results = self._mask_forward(
+ x=x,
+ rois=rois,
+ semantic_feat=semantic_feat,
+ glbctx_feat=glbctx_feat,
+ relayed_feat=relayed_feat)
+ mask_preds = mask_results['mask_preds']
+ mask_preds = mask_preds.split(num_proposals_per_img, 0)
+ results = results + (mask_preds, )
+ return results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/shared_heads/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/shared_heads/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..d56636ab34d1dd2592828238099bcdccf179d6d3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/shared_heads/__init__.py
@@ -0,0 +1,4 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .res_layer import ResLayer
+
+__all__ = ['ResLayer']
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/shared_heads/res_layer.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/shared_heads/res_layer.py
new file mode 100644
index 0000000000000000000000000000000000000000..d9210cb928fec92135a195d44d13a8588382b947
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/shared_heads/res_layer.py
@@ -0,0 +1,79 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+
+import torch.nn as nn
+from mmengine.model import BaseModule
+
+from mmdet.models.backbones import ResNet
+from mmdet.models.layers import ResLayer as _ResLayer
+from mmdet.registry import MODELS
+
+
+@MODELS.register_module()
+class ResLayer(BaseModule):
+
+ def __init__(self,
+ depth,
+ stage=3,
+ stride=2,
+ dilation=1,
+ style='pytorch',
+ norm_cfg=dict(type='BN', requires_grad=True),
+ norm_eval=True,
+ with_cp=False,
+ dcn=None,
+ pretrained=None,
+ init_cfg=None):
+ super(ResLayer, self).__init__(init_cfg)
+
+ self.norm_eval = norm_eval
+ self.norm_cfg = norm_cfg
+ self.stage = stage
+ self.fp16_enabled = False
+ block, stage_blocks = ResNet.arch_settings[depth]
+ stage_block = stage_blocks[stage]
+ planes = 64 * 2**stage
+ inplanes = 64 * 2**(stage - 1) * block.expansion
+
+ res_layer = _ResLayer(
+ block,
+ inplanes,
+ planes,
+ stage_block,
+ stride=stride,
+ dilation=dilation,
+ style=style,
+ with_cp=with_cp,
+ norm_cfg=self.norm_cfg,
+ dcn=dcn)
+ self.add_module(f'layer{stage + 1}', res_layer)
+
+ assert not (init_cfg and pretrained), \
+ 'init_cfg and pretrained cannot be specified at the same time'
+ if isinstance(pretrained, str):
+ warnings.warn('DeprecationWarning: pretrained is a deprecated, '
+ 'please use "init_cfg" instead')
+ self.init_cfg = dict(type='Pretrained', checkpoint=pretrained)
+ elif pretrained is None:
+ if init_cfg is None:
+ self.init_cfg = [
+ dict(type='Kaiming', layer='Conv2d'),
+ dict(
+ type='Constant',
+ val=1,
+ layer=['_BatchNorm', 'GroupNorm'])
+ ]
+ else:
+ raise TypeError('pretrained must be a str or None')
+
+ def forward(self, x):
+ res_layer = getattr(self, f'layer{self.stage + 1}')
+ out = res_layer(x)
+ return out
+
+ def train(self, mode=True):
+ super(ResLayer, self).train(mode)
+ if self.norm_eval:
+ for m in self.modules():
+ if isinstance(m, nn.BatchNorm2d):
+ m.eval()
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/sparse_roi_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/sparse_roi_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..19c3e1e335ca4e4a9d5befcbffcf4665b459cb5a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/sparse_roi_head.py
@@ -0,0 +1,601 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple
+
+import torch
+from mmengine.config import ConfigDict
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.models.task_modules.samplers import PseudoSampler
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import bbox2roi
+from mmdet.utils import ConfigType, InstanceList, OptConfigType
+from ..utils.misc import empty_instances, unpack_gt_instances
+from .cascade_roi_head import CascadeRoIHead
+
+
+@MODELS.register_module()
+class SparseRoIHead(CascadeRoIHead):
+ r"""The RoIHead for `Sparse R-CNN: End-to-End Object Detection with
+ Learnable Proposals `_
+ and `Instances as Queries `_
+
+ Args:
+ num_stages (int): Number of stage whole iterative process.
+ Defaults to 6.
+ stage_loss_weights (Tuple[float]): The loss
+ weight of each stage. By default all stages have
+ the same weight 1.
+ bbox_roi_extractor (:obj:`ConfigDict` or dict): Config of box
+ roi extractor.
+ mask_roi_extractor (:obj:`ConfigDict` or dict): Config of mask
+ roi extractor.
+ bbox_head (:obj:`ConfigDict` or dict): Config of box head.
+ mask_head (:obj:`ConfigDict` or dict): Config of mask head.
+ train_cfg (:obj:`ConfigDict` or dict, Optional): Configuration
+ information in train stage. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, Optional): Configuration
+ information in test stage. Defaults to None.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict]): Initialization config dict. Defaults to None.
+ """
+
+ def __init__(self,
+ num_stages: int = 6,
+ stage_loss_weights: Tuple[float] = (1, 1, 1, 1, 1, 1),
+ proposal_feature_channel: int = 256,
+ bbox_roi_extractor: ConfigType = dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(
+ type='RoIAlign', output_size=7, sampling_ratio=2),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32]),
+ mask_roi_extractor: OptConfigType = None,
+ bbox_head: ConfigType = dict(
+ type='DIIHead',
+ num_classes=80,
+ num_fcs=2,
+ num_heads=8,
+ num_cls_fcs=1,
+ num_reg_fcs=3,
+ feedforward_channels=2048,
+ hidden_channels=256,
+ dropout=0.0,
+ roi_feat_size=7,
+ ffn_act_cfg=dict(type='ReLU', inplace=True)),
+ mask_head: OptConfigType = None,
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ init_cfg: OptConfigType = None) -> None:
+ assert bbox_roi_extractor is not None
+ assert bbox_head is not None
+ assert len(stage_loss_weights) == num_stages
+ self.num_stages = num_stages
+ self.stage_loss_weights = stage_loss_weights
+ self.proposal_feature_channel = proposal_feature_channel
+ super().__init__(
+ num_stages=num_stages,
+ stage_loss_weights=stage_loss_weights,
+ bbox_roi_extractor=bbox_roi_extractor,
+ mask_roi_extractor=mask_roi_extractor,
+ bbox_head=bbox_head,
+ mask_head=mask_head,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ init_cfg=init_cfg)
+ # train_cfg would be None when run the test.py
+ if train_cfg is not None:
+ for stage in range(num_stages):
+ assert isinstance(self.bbox_sampler[stage], PseudoSampler), \
+ 'Sparse R-CNN and QueryInst only support `PseudoSampler`'
+
+ def bbox_loss(self, stage: int, x: Tuple[Tensor],
+ results_list: InstanceList, object_feats: Tensor,
+ batch_img_metas: List[dict],
+ batch_gt_instances: InstanceList) -> dict:
+ """Perform forward propagation and loss calculation of the bbox head on
+ the features of the upstream network.
+
+ Args:
+ stage (int): The current stage in iterative process.
+ x (tuple[Tensor]): List of multi-level img features.
+ results_list (List[:obj:`InstanceData`]) : List of region
+ proposals.
+ object_feats (Tensor): The object feature extracted from
+ the previous stage.
+ batch_img_metas (list[dict]): Meta information of each image.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``labels``, and
+ ``masks`` attributes.
+
+ Returns:
+ dict[str, Tensor]: Usually returns a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `bbox_feats` (Tensor): Extract bbox RoI features.
+ - `loss_bbox` (dict): A dictionary of bbox loss components.
+ """
+ proposal_list = [res.bboxes for res in results_list]
+ rois = bbox2roi(proposal_list)
+ bbox_results = self._bbox_forward(stage, x, rois, object_feats,
+ batch_img_metas)
+ imgs_whwh = torch.cat(
+ [res.imgs_whwh[None, ...] for res in results_list])
+ cls_pred_list = bbox_results['detached_cls_scores']
+ proposal_list = bbox_results['detached_proposals']
+
+ sampling_results = []
+ bbox_head = self.bbox_head[stage]
+ for i in range(len(batch_img_metas)):
+ pred_instances = InstanceData()
+ # TODO: Enhance the logic
+ pred_instances.bboxes = proposal_list[i] # for assinger
+ pred_instances.scores = cls_pred_list[i]
+ pred_instances.priors = proposal_list[i] # for sampler
+
+ assign_result = self.bbox_assigner[stage].assign(
+ pred_instances=pred_instances,
+ gt_instances=batch_gt_instances[i],
+ gt_instances_ignore=None,
+ img_meta=batch_img_metas[i])
+
+ sampling_result = self.bbox_sampler[stage].sample(
+ assign_result, pred_instances, batch_gt_instances[i])
+ sampling_results.append(sampling_result)
+
+ bbox_results.update(sampling_results=sampling_results)
+
+ cls_score = bbox_results['cls_score']
+ decoded_bboxes = bbox_results['decoded_bboxes']
+ cls_score = cls_score.view(-1, cls_score.size(-1))
+ decoded_bboxes = decoded_bboxes.view(-1, 4)
+ bbox_loss_and_target = bbox_head.loss_and_target(
+ cls_score,
+ decoded_bboxes,
+ sampling_results,
+ self.train_cfg[stage],
+ imgs_whwh=imgs_whwh,
+ concat=True)
+ bbox_results.update(bbox_loss_and_target)
+
+ # propose for the new proposal_list
+ proposal_list = []
+ for idx in range(len(batch_img_metas)):
+ results = InstanceData()
+ results.imgs_whwh = results_list[idx].imgs_whwh
+ results.bboxes = bbox_results['detached_proposals'][idx]
+ proposal_list.append(results)
+ bbox_results.update(results_list=proposal_list)
+ return bbox_results
+
+ def _bbox_forward(self, stage: int, x: Tuple[Tensor], rois: Tensor,
+ object_feats: Tensor,
+ batch_img_metas: List[dict]) -> dict:
+ """Box head forward function used in both training and testing. Returns
+ all regression, classification results and a intermediate feature.
+
+ Args:
+ stage (int): The current stage in iterative process.
+ x (tuple[Tensor]): List of multi-level img features.
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+ Each dimension means (img_index, x1, y1, x2, y2).
+ object_feats (Tensor): The object feature extracted from
+ the previous stage.
+ batch_img_metas (list[dict]): Meta information of each image.
+
+ Returns:
+ dict[str, Tensor]: a dictionary of bbox head outputs,
+ Containing the following results:
+
+ - cls_score (Tensor): The score of each class, has
+ shape (batch_size, num_proposals, num_classes)
+ when use focal loss or
+ (batch_size, num_proposals, num_classes+1)
+ otherwise.
+ - decoded_bboxes (Tensor): The regression results
+ with shape (batch_size, num_proposal, 4).
+ The last dimension 4 represents
+ [tl_x, tl_y, br_x, br_y].
+ - object_feats (Tensor): The object feature extracted
+ from current stage
+ - detached_cls_scores (list[Tensor]): The detached
+ classification results, length is batch_size, and
+ each tensor has shape (num_proposal, num_classes).
+ - detached_proposals (list[tensor]): The detached
+ regression results, length is batch_size, and each
+ tensor has shape (num_proposal, 4). The last
+ dimension 4 represents [tl_x, tl_y, br_x, br_y].
+ """
+ num_imgs = len(batch_img_metas)
+ bbox_roi_extractor = self.bbox_roi_extractor[stage]
+ bbox_head = self.bbox_head[stage]
+ bbox_feats = bbox_roi_extractor(x[:bbox_roi_extractor.num_inputs],
+ rois)
+ cls_score, bbox_pred, object_feats, attn_feats = bbox_head(
+ bbox_feats, object_feats)
+
+ fake_bbox_results = dict(
+ rois=rois,
+ bbox_targets=(rois.new_zeros(len(rois), dtype=torch.long), None),
+ bbox_pred=bbox_pred.view(-1, bbox_pred.size(-1)),
+ cls_score=cls_score.view(-1, cls_score.size(-1)))
+ fake_sampling_results = [
+ InstanceData(pos_is_gt=rois.new_zeros(object_feats.size(1)))
+ for _ in range(len(batch_img_metas))
+ ]
+
+ results_list = bbox_head.refine_bboxes(
+ sampling_results=fake_sampling_results,
+ bbox_results=fake_bbox_results,
+ batch_img_metas=batch_img_metas)
+ proposal_list = [res.bboxes for res in results_list]
+ bbox_results = dict(
+ cls_score=cls_score,
+ decoded_bboxes=torch.cat(proposal_list),
+ object_feats=object_feats,
+ attn_feats=attn_feats,
+ # detach then use it in label assign
+ detached_cls_scores=[
+ cls_score[i].detach() for i in range(num_imgs)
+ ],
+ detached_proposals=[item.detach() for item in proposal_list])
+
+ return bbox_results
+
+ def _mask_forward(self, stage: int, x: Tuple[Tensor], rois: Tensor,
+ attn_feats) -> dict:
+ """Mask head forward function used in both training and testing.
+
+ Args:
+ stage (int): The current stage in Cascade RoI Head.
+ x (tuple[Tensor]): Tuple of multi-level img features.
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+ attn_feats (Tensot): Intermediate feature get from the last
+ diihead, has shape
+ (batch_size*num_proposals, feature_dimensions)
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `mask_preds` (Tensor): Mask prediction.
+ """
+ mask_roi_extractor = self.mask_roi_extractor[stage]
+ mask_head = self.mask_head[stage]
+ mask_feats = mask_roi_extractor(x[:mask_roi_extractor.num_inputs],
+ rois)
+ # do not support caffe_c4 model anymore
+ mask_preds = mask_head(mask_feats, attn_feats)
+
+ mask_results = dict(mask_preds=mask_preds)
+ return mask_results
+
+ def mask_loss(self, stage: int, x: Tuple[Tensor], bbox_results: dict,
+ batch_gt_instances: InstanceList,
+ rcnn_train_cfg: ConfigDict) -> dict:
+ """Run forward function and calculate loss for mask head in training.
+
+ Args:
+ stage (int): The current stage in Cascade RoI Head.
+ x (tuple[Tensor]): Tuple of multi-level img features.
+ bbox_results (dict): Results obtained from `bbox_loss`.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``labels``, and
+ ``masks`` attributes.
+ rcnn_train_cfg (obj:ConfigDict): `train_cfg` of RCNN.
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `mask_preds` (Tensor): Mask prediction.
+ - `loss_mask` (dict): A dictionary of mask loss components.
+ """
+ attn_feats = bbox_results['attn_feats']
+ sampling_results = bbox_results['sampling_results']
+
+ pos_rois = bbox2roi([res.pos_priors for res in sampling_results])
+
+ attn_feats = torch.cat([
+ feats[res.pos_inds]
+ for (feats, res) in zip(attn_feats, sampling_results)
+ ])
+ mask_results = self._mask_forward(stage, x, pos_rois, attn_feats)
+
+ mask_loss_and_target = self.mask_head[stage].loss_and_target(
+ mask_preds=mask_results['mask_preds'],
+ sampling_results=sampling_results,
+ batch_gt_instances=batch_gt_instances,
+ rcnn_train_cfg=rcnn_train_cfg)
+ mask_results.update(mask_loss_and_target)
+
+ return mask_results
+
+ def loss(self, x: Tuple[Tensor], rpn_results_list: InstanceList,
+ batch_data_samples: SampleList) -> dict:
+ """Perform forward propagation and loss calculation of the detection
+ roi on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ rpn_results_list (List[:obj:`InstanceData`]): List of region
+ proposals.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict: a dictionary of loss components of all stage.
+ """
+ outputs = unpack_gt_instances(batch_data_samples)
+ batch_gt_instances, batch_gt_instances_ignore, batch_img_metas \
+ = outputs
+
+ object_feats = torch.cat(
+ [res.pop('features')[None, ...] for res in rpn_results_list])
+ results_list = rpn_results_list
+ losses = {}
+ for stage in range(self.num_stages):
+ stage_loss_weight = self.stage_loss_weights[stage]
+
+ # bbox head forward and loss
+ bbox_results = self.bbox_loss(
+ stage=stage,
+ x=x,
+ object_feats=object_feats,
+ results_list=results_list,
+ batch_img_metas=batch_img_metas,
+ batch_gt_instances=batch_gt_instances)
+
+ for name, value in bbox_results['loss_bbox'].items():
+ losses[f's{stage}.{name}'] = (
+ value * stage_loss_weight if 'loss' in name else value)
+
+ if self.with_mask:
+ mask_results = self.mask_loss(
+ stage=stage,
+ x=x,
+ bbox_results=bbox_results,
+ batch_gt_instances=batch_gt_instances,
+ rcnn_train_cfg=self.train_cfg[stage])
+
+ for name, value in mask_results['loss_mask'].items():
+ losses[f's{stage}.{name}'] = (
+ value * stage_loss_weight if 'loss' in name else value)
+
+ object_feats = bbox_results['object_feats']
+ results_list = bbox_results['results_list']
+ return losses
+
+ def predict_bbox(self,
+ x: Tuple[Tensor],
+ batch_img_metas: List[dict],
+ rpn_results_list: InstanceList,
+ rcnn_test_cfg: ConfigType,
+ rescale: bool = False) -> InstanceList:
+ """Perform forward propagation of the bbox head and predict detection
+ results on the features of the upstream network.
+
+ Args:
+ x(tuple[Tensor]): Feature maps of all scale level.
+ batch_img_metas (list[dict]): List of image information.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ rcnn_test_cfg (obj:`ConfigDict`): `test_cfg` of R-CNN.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ proposal_list = [res.bboxes for res in rpn_results_list]
+ object_feats = torch.cat(
+ [res.pop('features')[None, ...] for res in rpn_results_list])
+ if all([proposal.shape[0] == 0 for proposal in proposal_list]):
+ # There is no proposal in the whole batch
+ return empty_instances(
+ batch_img_metas, x[0].device, task_type='bbox')
+
+ for stage in range(self.num_stages):
+ rois = bbox2roi(proposal_list)
+ bbox_results = self._bbox_forward(stage, x, rois, object_feats,
+ batch_img_metas)
+ object_feats = bbox_results['object_feats']
+ cls_score = bbox_results['cls_score']
+ proposal_list = bbox_results['detached_proposals']
+
+ num_classes = self.bbox_head[-1].num_classes
+
+ if self.bbox_head[-1].loss_cls.use_sigmoid:
+ cls_score = cls_score.sigmoid()
+ else:
+ cls_score = cls_score.softmax(-1)[..., :-1]
+
+ topk_inds_list = []
+ results_list = []
+ for img_id in range(len(batch_img_metas)):
+ cls_score_per_img = cls_score[img_id]
+ scores_per_img, topk_inds = cls_score_per_img.flatten(0, 1).topk(
+ self.test_cfg.max_per_img, sorted=False)
+ labels_per_img = topk_inds % num_classes
+ bboxes_per_img = proposal_list[img_id][topk_inds // num_classes]
+ topk_inds_list.append(topk_inds)
+ if rescale and bboxes_per_img.size(0) > 0:
+ assert batch_img_metas[img_id].get('scale_factor') is not None
+ scale_factor = bboxes_per_img.new_tensor(
+ batch_img_metas[img_id]['scale_factor']).repeat((1, 2))
+ bboxes_per_img = (
+ bboxes_per_img.view(bboxes_per_img.size(0), -1, 4) /
+ scale_factor).view(bboxes_per_img.size()[0], -1)
+
+ results = InstanceData()
+ results.bboxes = bboxes_per_img
+ results.scores = scores_per_img
+ results.labels = labels_per_img
+ results_list.append(results)
+ if self.with_mask:
+ for img_id in range(len(batch_img_metas)):
+ # add positive information in InstanceData to predict
+ # mask results in `mask_head`.
+ proposals = bbox_results['detached_proposals'][img_id]
+ topk_inds = topk_inds_list[img_id]
+ attn_feats = bbox_results['attn_feats'][img_id]
+
+ results_list[img_id].proposals = proposals
+ results_list[img_id].topk_inds = topk_inds
+ results_list[img_id].attn_feats = attn_feats
+ return results_list
+
+ def predict_mask(self,
+ x: Tuple[Tensor],
+ batch_img_metas: List[dict],
+ results_list: InstanceList,
+ rescale: bool = False) -> InstanceList:
+ """Perform forward propagation of the mask head and predict detection
+ results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Feature maps of all scale level.
+ batch_img_metas (list[dict]): List of image information.
+ results_list (list[:obj:`InstanceData`]): Detection results of
+ each image. Each item usually contains following keys:
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - proposal (Tensor): Bboxes predicted from bbox_head,
+ has a shape (num_instances, 4).
+ - topk_inds (Tensor): Topk indices of each image, has
+ shape (num_instances, )
+ - attn_feats (Tensor): Intermediate feature get from the last
+ diihead, has shape (num_instances, feature_dimensions)
+
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+ """
+ proposal_list = [res.pop('proposals') for res in results_list]
+ topk_inds_list = [res.pop('topk_inds') for res in results_list]
+ attn_feats = torch.cat(
+ [res.pop('attn_feats')[None, ...] for res in results_list])
+
+ rois = bbox2roi(proposal_list)
+
+ if rois.shape[0] == 0:
+ results_list = empty_instances(
+ batch_img_metas,
+ rois.device,
+ task_type='mask',
+ instance_results=results_list,
+ mask_thr_binary=self.test_cfg.mask_thr_binary)
+ return results_list
+
+ last_stage = self.num_stages - 1
+ mask_results = self._mask_forward(last_stage, x, rois, attn_feats)
+
+ num_imgs = len(batch_img_metas)
+ mask_results['mask_preds'] = mask_results['mask_preds'].reshape(
+ num_imgs, -1, *mask_results['mask_preds'].size()[1:])
+ num_classes = self.bbox_head[-1].num_classes
+
+ mask_preds = []
+ for img_id in range(num_imgs):
+ topk_inds = topk_inds_list[img_id]
+ masks_per_img = mask_results['mask_preds'][img_id].flatten(
+ 0, 1)[topk_inds]
+ masks_per_img = masks_per_img[:, None,
+ ...].repeat(1, num_classes, 1, 1)
+ mask_preds.append(masks_per_img)
+ results_list = self.mask_head[-1].predict_by_feat(
+ mask_preds,
+ results_list,
+ batch_img_metas,
+ rcnn_test_cfg=self.test_cfg,
+ rescale=rescale)
+
+ return results_list
+
+ # TODO: Need to refactor later
+ def forward(self, x: Tuple[Tensor], rpn_results_list: InstanceList,
+ batch_data_samples: SampleList) -> tuple:
+ """Network forward process. Usually includes backbone, neck and head
+ forward without any post-processing.
+
+ Args:
+ x (List[Tensor]): Multi-level features that may have different
+ resolutions.
+ rpn_results_list (List[:obj:`InstanceData`]): List of region
+ proposals.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns
+ tuple: A tuple of features from ``bbox_head`` and ``mask_head``
+ forward.
+ """
+ outputs = unpack_gt_instances(batch_data_samples)
+ (batch_gt_instances, batch_gt_instances_ignore,
+ batch_img_metas) = outputs
+
+ all_stage_bbox_results = []
+ object_feats = torch.cat(
+ [res.pop('features')[None, ...] for res in rpn_results_list])
+ results_list = rpn_results_list
+ if self.with_bbox:
+ for stage in range(self.num_stages):
+ bbox_results = self.bbox_loss(
+ stage=stage,
+ x=x,
+ results_list=results_list,
+ object_feats=object_feats,
+ batch_img_metas=batch_img_metas,
+ batch_gt_instances=batch_gt_instances)
+ bbox_results.pop('loss_bbox')
+ # torch.jit does not support obj:SamplingResult
+ bbox_results.pop('results_list')
+ bbox_res = bbox_results.copy()
+ bbox_res.pop('sampling_results')
+ all_stage_bbox_results.append((bbox_res, ))
+
+ if self.with_mask:
+ attn_feats = bbox_results['attn_feats']
+ sampling_results = bbox_results['sampling_results']
+
+ pos_rois = bbox2roi(
+ [res.pos_priors for res in sampling_results])
+
+ attn_feats = torch.cat([
+ feats[res.pos_inds]
+ for (feats, res) in zip(attn_feats, sampling_results)
+ ])
+ mask_results = self._mask_forward(stage, x, pos_rois,
+ attn_feats)
+ all_stage_bbox_results[-1] += (mask_results, )
+ return tuple(all_stage_bbox_results)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/standard_roi_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/standard_roi_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..8d168eba0fb2ccf6aa89bde5c637160f10aea83a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/standard_roi_head.py
@@ -0,0 +1,419 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple
+
+import torch
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures import DetDataSample, SampleList
+from mmdet.structures.bbox import bbox2roi
+from mmdet.utils import ConfigType, InstanceList
+from ..task_modules.samplers import SamplingResult
+from ..utils import empty_instances, unpack_gt_instances
+from .base_roi_head import BaseRoIHead
+
+
+@MODELS.register_module()
+class StandardRoIHead(BaseRoIHead):
+ """Simplest base roi head including one bbox head and one mask head."""
+
+ def init_assigner_sampler(self) -> None:
+ """Initialize assigner and sampler."""
+ self.bbox_assigner = None
+ self.bbox_sampler = None
+ if self.train_cfg:
+ self.bbox_assigner = TASK_UTILS.build(self.train_cfg.assigner)
+ self.bbox_sampler = TASK_UTILS.build(
+ self.train_cfg.sampler, default_args=dict(context=self))
+
+ def init_bbox_head(self, bbox_roi_extractor: ConfigType,
+ bbox_head: ConfigType) -> None:
+ """Initialize box head and box roi extractor.
+
+ Args:
+ bbox_roi_extractor (dict or ConfigDict): Config of box
+ roi extractor.
+ bbox_head (dict or ConfigDict): Config of box in box head.
+ """
+ self.bbox_roi_extractor = MODELS.build(bbox_roi_extractor)
+ self.bbox_head = MODELS.build(bbox_head)
+
+ def init_mask_head(self, mask_roi_extractor: ConfigType,
+ mask_head: ConfigType) -> None:
+ """Initialize mask head and mask roi extractor.
+
+ Args:
+ mask_roi_extractor (dict or ConfigDict): Config of mask roi
+ extractor.
+ mask_head (dict or ConfigDict): Config of mask in mask head.
+ """
+ if mask_roi_extractor is not None:
+ self.mask_roi_extractor = MODELS.build(mask_roi_extractor)
+ self.share_roi_extractor = False
+ else:
+ self.share_roi_extractor = True
+ self.mask_roi_extractor = self.bbox_roi_extractor
+ self.mask_head = MODELS.build(mask_head)
+
+ # TODO: Need to refactor later
+ def forward(self,
+ x: Tuple[Tensor],
+ rpn_results_list: InstanceList,
+ batch_data_samples: SampleList = None) -> tuple:
+ """Network forward process. Usually includes backbone, neck and head
+ forward without any post-processing.
+
+ Args:
+ x (List[Tensor]): Multi-level features that may have different
+ resolutions.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ batch_data_samples (list[:obj:`DetDataSample`]): Each item contains
+ the meta information of each image and corresponding
+ annotations.
+
+ Returns
+ tuple: A tuple of features from ``bbox_head`` and ``mask_head``
+ forward.
+ """
+ results = ()
+ proposals = [rpn_results.bboxes for rpn_results in rpn_results_list]
+ rois = bbox2roi(proposals)
+ # bbox head
+ if self.with_bbox:
+ bbox_results = self._bbox_forward(x, rois)
+ results = results + (bbox_results['cls_score'],
+ bbox_results['bbox_pred'])
+ # mask head
+ if self.with_mask:
+ mask_rois = rois[:100]
+ mask_results = self._mask_forward(x, mask_rois)
+ results = results + (mask_results['mask_preds'], )
+ return results
+
+ def loss(self, x: Tuple[Tensor], rpn_results_list: InstanceList,
+ batch_data_samples: List[DetDataSample]) -> dict:
+ """Perform forward propagation and loss calculation of the detection
+ roi on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components
+ """
+ assert len(rpn_results_list) == len(batch_data_samples)
+ outputs = unpack_gt_instances(batch_data_samples)
+ batch_gt_instances, batch_gt_instances_ignore, _ = outputs
+
+ # assign gts and sample proposals
+ num_imgs = len(batch_data_samples)
+ sampling_results = []
+ for i in range(num_imgs):
+ # rename rpn_results.bboxes to rpn_results.priors
+ rpn_results = rpn_results_list[i]
+ rpn_results.priors = rpn_results.pop('bboxes')
+
+ assign_result = self.bbox_assigner.assign(
+ rpn_results, batch_gt_instances[i],
+ batch_gt_instances_ignore[i])
+ sampling_result = self.bbox_sampler.sample(
+ assign_result,
+ rpn_results,
+ batch_gt_instances[i],
+ feats=[lvl_feat[i][None] for lvl_feat in x])
+ sampling_results.append(sampling_result)
+
+ losses = dict()
+ # bbox head loss
+ if self.with_bbox:
+ bbox_results = self.bbox_loss(x, sampling_results)
+ losses.update(bbox_results['loss_bbox'])
+
+ # mask head forward and loss
+ if self.with_mask:
+ mask_results = self.mask_loss(x, sampling_results,
+ bbox_results['bbox_feats'],
+ batch_gt_instances)
+ losses.update(mask_results['loss_mask'])
+
+ return losses
+
+ def _bbox_forward(self, x: Tuple[Tensor], rois: Tensor) -> dict:
+ """Box head forward function used in both training and testing.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+
+ Returns:
+ dict[str, Tensor]: Usually returns a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `bbox_feats` (Tensor): Extract bbox RoI features.
+ """
+ # TODO: a more flexible way to decide which feature maps to use
+ bbox_feats = self.bbox_roi_extractor(
+ x[:self.bbox_roi_extractor.num_inputs], rois)
+ if self.with_shared_head:
+ bbox_feats = self.shared_head(bbox_feats)
+ cls_score, bbox_pred = self.bbox_head(bbox_feats)
+
+ bbox_results = dict(
+ cls_score=cls_score, bbox_pred=bbox_pred, bbox_feats=bbox_feats)
+ return bbox_results
+
+ def bbox_loss(self, x: Tuple[Tensor],
+ sampling_results: List[SamplingResult]) -> dict:
+ """Perform forward propagation and loss calculation of the bbox head on
+ the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ sampling_results (list["obj:`SamplingResult`]): Sampling results.
+
+ Returns:
+ dict[str, Tensor]: Usually returns a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `bbox_feats` (Tensor): Extract bbox RoI features.
+ - `loss_bbox` (dict): A dictionary of bbox loss components.
+ """
+ rois = bbox2roi([res.priors for res in sampling_results])
+ bbox_results = self._bbox_forward(x, rois)
+
+ bbox_loss_and_target = self.bbox_head.loss_and_target(
+ cls_score=bbox_results['cls_score'],
+ bbox_pred=bbox_results['bbox_pred'],
+ rois=rois,
+ sampling_results=sampling_results,
+ rcnn_train_cfg=self.train_cfg)
+
+ bbox_results.update(loss_bbox=bbox_loss_and_target['loss_bbox'])
+ return bbox_results
+
+ def mask_loss(self, x: Tuple[Tensor],
+ sampling_results: List[SamplingResult], bbox_feats: Tensor,
+ batch_gt_instances: InstanceList) -> dict:
+ """Perform forward propagation and loss calculation of the mask head on
+ the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Tuple of multi-level img features.
+ sampling_results (list["obj:`SamplingResult`]): Sampling results.
+ bbox_feats (Tensor): Extract bbox RoI features.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``labels``, and
+ ``masks`` attributes.
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `mask_preds` (Tensor): Mask prediction.
+ - `mask_feats` (Tensor): Extract mask RoI features.
+ - `mask_targets` (Tensor): Mask target of each positive\
+ proposals in the image.
+ - `loss_mask` (dict): A dictionary of mask loss components.
+ """
+ if not self.share_roi_extractor:
+ pos_rois = bbox2roi([res.pos_priors for res in sampling_results])
+ mask_results = self._mask_forward(x, pos_rois)
+ else:
+ pos_inds = []
+ device = bbox_feats.device
+ for res in sampling_results:
+ pos_inds.append(
+ torch.ones(
+ res.pos_priors.shape[0],
+ device=device,
+ dtype=torch.uint8))
+ pos_inds.append(
+ torch.zeros(
+ res.neg_priors.shape[0],
+ device=device,
+ dtype=torch.uint8))
+ pos_inds = torch.cat(pos_inds)
+
+ mask_results = self._mask_forward(
+ x, pos_inds=pos_inds, bbox_feats=bbox_feats)
+
+ mask_loss_and_target = self.mask_head.loss_and_target(
+ mask_preds=mask_results['mask_preds'],
+ sampling_results=sampling_results,
+ batch_gt_instances=batch_gt_instances,
+ rcnn_train_cfg=self.train_cfg)
+
+ mask_results.update(loss_mask=mask_loss_and_target['loss_mask'])
+ return mask_results
+
+ def _mask_forward(self,
+ x: Tuple[Tensor],
+ rois: Tensor = None,
+ pos_inds: Optional[Tensor] = None,
+ bbox_feats: Optional[Tensor] = None) -> dict:
+ """Mask head forward function used in both training and testing.
+
+ Args:
+ x (tuple[Tensor]): Tuple of multi-level img features.
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+ pos_inds (Tensor, optional): Indices of positive samples.
+ Defaults to None.
+ bbox_feats (Tensor): Extract bbox RoI features. Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: Usually returns a dictionary with keys:
+
+ - `mask_preds` (Tensor): Mask prediction.
+ - `mask_feats` (Tensor): Extract mask RoI features.
+ """
+ assert ((rois is not None) ^
+ (pos_inds is not None and bbox_feats is not None))
+ if rois is not None:
+ mask_feats = self.mask_roi_extractor(
+ x[:self.mask_roi_extractor.num_inputs], rois)
+ if self.with_shared_head:
+ mask_feats = self.shared_head(mask_feats)
+ else:
+ assert bbox_feats is not None
+ mask_feats = bbox_feats[pos_inds]
+
+ mask_preds = self.mask_head(mask_feats)
+ mask_results = dict(mask_preds=mask_preds, mask_feats=mask_feats)
+ return mask_results
+
+ def predict_bbox(self,
+ x: Tuple[Tensor],
+ batch_img_metas: List[dict],
+ rpn_results_list: InstanceList,
+ rcnn_test_cfg: ConfigType,
+ rescale: bool = False) -> InstanceList:
+ """Perform forward propagation of the bbox head and predict detection
+ results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Feature maps of all scale level.
+ batch_img_metas (list[dict]): List of image information.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ rcnn_test_cfg (obj:`ConfigDict`): `test_cfg` of R-CNN.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ proposals = [res.bboxes for res in rpn_results_list]
+ rois = bbox2roi(proposals)
+
+ if rois.shape[0] == 0:
+ return empty_instances(
+ batch_img_metas,
+ rois.device,
+ task_type='bbox',
+ box_type=self.bbox_head.predict_box_type,
+ num_classes=self.bbox_head.num_classes,
+ score_per_cls=rcnn_test_cfg is None)
+
+ bbox_results = self._bbox_forward(x, rois)
+
+ # split batch bbox prediction back to each image
+ cls_scores = bbox_results['cls_score']
+ bbox_preds = bbox_results['bbox_pred']
+ num_proposals_per_img = tuple(len(p) for p in proposals)
+ rois = rois.split(num_proposals_per_img, 0)
+ cls_scores = cls_scores.split(num_proposals_per_img, 0)
+
+ # some detector with_reg is False, bbox_preds will be None
+ if bbox_preds is not None:
+ # TODO move this to a sabl_roi_head
+ # the bbox prediction of some detectors like SABL is not Tensor
+ if isinstance(bbox_preds, torch.Tensor):
+ bbox_preds = bbox_preds.split(num_proposals_per_img, 0)
+ else:
+ bbox_preds = self.bbox_head.bbox_pred_split(
+ bbox_preds, num_proposals_per_img)
+ else:
+ bbox_preds = (None, ) * len(proposals)
+
+ result_list = self.bbox_head.predict_by_feat(
+ rois=rois,
+ cls_scores=cls_scores,
+ bbox_preds=bbox_preds,
+ batch_img_metas=batch_img_metas,
+ rcnn_test_cfg=rcnn_test_cfg,
+ rescale=rescale)
+ return result_list
+
+ def predict_mask(self,
+ x: Tuple[Tensor],
+ batch_img_metas: List[dict],
+ results_list: InstanceList,
+ rescale: bool = False) -> InstanceList:
+ """Perform forward propagation of the mask head and predict detection
+ results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Feature maps of all scale level.
+ batch_img_metas (list[dict]): List of image information.
+ results_list (list[:obj:`InstanceData`]): Detection results of
+ each image.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+ """
+ # don't need to consider aug_test.
+ bboxes = [res.bboxes for res in results_list]
+ mask_rois = bbox2roi(bboxes)
+ if mask_rois.shape[0] == 0:
+ results_list = empty_instances(
+ batch_img_metas,
+ mask_rois.device,
+ task_type='mask',
+ instance_results=results_list,
+ mask_thr_binary=self.test_cfg.mask_thr_binary)
+ return results_list
+
+ mask_results = self._mask_forward(x, mask_rois)
+ mask_preds = mask_results['mask_preds']
+ # split batch mask prediction back to each image
+ num_mask_rois_per_img = [len(res) for res in results_list]
+ mask_preds = mask_preds.split(num_mask_rois_per_img, 0)
+
+ # TODO: Handle the case where rescale is false
+ results_list = self.mask_head.predict_by_feat(
+ mask_preds=mask_preds,
+ results_list=results_list,
+ batch_img_metas=batch_img_metas,
+ rcnn_test_cfg=self.test_cfg,
+ rescale=rescale)
+ return results_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/test_mixins.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/test_mixins.py
new file mode 100644
index 0000000000000000000000000000000000000000..940490454d9cf1fde4d69c1f890c173b92d522a1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/test_mixins.py
@@ -0,0 +1,171 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+# TODO: delete this file after refactor
+import sys
+
+import torch
+
+from mmdet.models.layers import multiclass_nms
+from mmdet.models.test_time_augs import merge_aug_bboxes, merge_aug_masks
+from mmdet.structures.bbox import bbox2roi, bbox_mapping
+
+if sys.version_info >= (3, 7):
+ from mmdet.utils.contextmanagers import completed
+
+
+class BBoxTestMixin:
+
+ if sys.version_info >= (3, 7):
+ # TODO: Currently not supported
+ async def async_test_bboxes(self,
+ x,
+ img_metas,
+ proposals,
+ rcnn_test_cfg,
+ rescale=False,
+ **kwargs):
+ """Asynchronized test for box head without augmentation."""
+ rois = bbox2roi(proposals)
+ roi_feats = self.bbox_roi_extractor(
+ x[:len(self.bbox_roi_extractor.featmap_strides)], rois)
+ if self.with_shared_head:
+ roi_feats = self.shared_head(roi_feats)
+ sleep_interval = rcnn_test_cfg.get('async_sleep_interval', 0.017)
+
+ async with completed(
+ __name__, 'bbox_head_forward',
+ sleep_interval=sleep_interval):
+ cls_score, bbox_pred = self.bbox_head(roi_feats)
+
+ img_shape = img_metas[0]['img_shape']
+ scale_factor = img_metas[0]['scale_factor']
+ det_bboxes, det_labels = self.bbox_head.get_bboxes(
+ rois,
+ cls_score,
+ bbox_pred,
+ img_shape,
+ scale_factor,
+ rescale=rescale,
+ cfg=rcnn_test_cfg)
+ return det_bboxes, det_labels
+
+ # TODO: Currently not supported
+ def aug_test_bboxes(self, feats, img_metas, rpn_results_list,
+ rcnn_test_cfg):
+ """Test det bboxes with test time augmentation."""
+ aug_bboxes = []
+ aug_scores = []
+ for x, img_meta in zip(feats, img_metas):
+ # only one image in the batch
+ img_shape = img_meta[0]['img_shape']
+ scale_factor = img_meta[0]['scale_factor']
+ flip = img_meta[0]['flip']
+ flip_direction = img_meta[0]['flip_direction']
+ # TODO more flexible
+ proposals = bbox_mapping(rpn_results_list[0][:, :4], img_shape,
+ scale_factor, flip, flip_direction)
+ rois = bbox2roi([proposals])
+ bbox_results = self.bbox_forward(x, rois)
+ bboxes, scores = self.bbox_head.get_bboxes(
+ rois,
+ bbox_results['cls_score'],
+ bbox_results['bbox_pred'],
+ img_shape,
+ scale_factor,
+ rescale=False,
+ cfg=None)
+ aug_bboxes.append(bboxes)
+ aug_scores.append(scores)
+ # after merging, bboxes will be rescaled to the original image size
+ merged_bboxes, merged_scores = merge_aug_bboxes(
+ aug_bboxes, aug_scores, img_metas, rcnn_test_cfg)
+ if merged_bboxes.shape[0] == 0:
+ # There is no proposal in the single image
+ det_bboxes = merged_bboxes.new_zeros(0, 5)
+ det_labels = merged_bboxes.new_zeros((0, ), dtype=torch.long)
+ else:
+ det_bboxes, det_labels = multiclass_nms(merged_bboxes,
+ merged_scores,
+ rcnn_test_cfg.score_thr,
+ rcnn_test_cfg.nms,
+ rcnn_test_cfg.max_per_img)
+ return det_bboxes, det_labels
+
+
+class MaskTestMixin:
+
+ if sys.version_info >= (3, 7):
+ # TODO: Currently not supported
+ async def async_test_mask(self,
+ x,
+ img_metas,
+ det_bboxes,
+ det_labels,
+ rescale=False,
+ mask_test_cfg=None):
+ """Asynchronized test for mask head without augmentation."""
+ # image shape of the first image in the batch (only one)
+ ori_shape = img_metas[0]['ori_shape']
+ scale_factor = img_metas[0]['scale_factor']
+ if det_bboxes.shape[0] == 0:
+ segm_result = [[] for _ in range(self.mask_head.num_classes)]
+ else:
+ if rescale and not isinstance(scale_factor,
+ (float, torch.Tensor)):
+ scale_factor = det_bboxes.new_tensor(scale_factor)
+ _bboxes = (
+ det_bboxes[:, :4] *
+ scale_factor if rescale else det_bboxes)
+ mask_rois = bbox2roi([_bboxes])
+ mask_feats = self.mask_roi_extractor(
+ x[:len(self.mask_roi_extractor.featmap_strides)],
+ mask_rois)
+
+ if self.with_shared_head:
+ mask_feats = self.shared_head(mask_feats)
+ if mask_test_cfg and \
+ mask_test_cfg.get('async_sleep_interval'):
+ sleep_interval = mask_test_cfg['async_sleep_interval']
+ else:
+ sleep_interval = 0.035
+ async with completed(
+ __name__,
+ 'mask_head_forward',
+ sleep_interval=sleep_interval):
+ mask_pred = self.mask_head(mask_feats)
+ segm_result = self.mask_head.get_results(
+ mask_pred, _bboxes, det_labels, self.test_cfg, ori_shape,
+ scale_factor, rescale)
+ return segm_result
+
+ # TODO: Currently not supported
+ def aug_test_mask(self, feats, img_metas, det_bboxes, det_labels):
+ """Test for mask head with test time augmentation."""
+ if det_bboxes.shape[0] == 0:
+ segm_result = [[] for _ in range(self.mask_head.num_classes)]
+ else:
+ aug_masks = []
+ for x, img_meta in zip(feats, img_metas):
+ img_shape = img_meta[0]['img_shape']
+ scale_factor = img_meta[0]['scale_factor']
+ flip = img_meta[0]['flip']
+ flip_direction = img_meta[0]['flip_direction']
+ _bboxes = bbox_mapping(det_bboxes[:, :4], img_shape,
+ scale_factor, flip, flip_direction)
+ mask_rois = bbox2roi([_bboxes])
+ mask_results = self._mask_forward(x, mask_rois)
+ # convert to numpy array to save memory
+ aug_masks.append(
+ mask_results['mask_pred'].sigmoid().cpu().numpy())
+ merged_masks = merge_aug_masks(aug_masks, img_metas, self.test_cfg)
+
+ ori_shape = img_metas[0][0]['ori_shape']
+ scale_factor = det_bboxes.new_ones(4)
+ segm_result = self.mask_head.get_results(
+ merged_masks,
+ det_bboxes,
+ det_labels,
+ self.test_cfg,
+ ori_shape,
+ scale_factor=scale_factor,
+ rescale=False)
+ return segm_result
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/trident_roi_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/trident_roi_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..5215327296282a8e7ca502f3321aced8a4f840b7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/roi_heads/trident_roi_head.py
@@ -0,0 +1,112 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Tuple
+
+import torch
+from mmcv.ops import batched_nms
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.utils import InstanceList
+from .standard_roi_head import StandardRoIHead
+
+
+@MODELS.register_module()
+class TridentRoIHead(StandardRoIHead):
+ """Trident roi head.
+
+ Args:
+ num_branch (int): Number of branches in TridentNet.
+ test_branch_idx (int): In inference, all 3 branches will be used
+ if `test_branch_idx==-1`, otherwise only branch with index
+ `test_branch_idx` will be used.
+ """
+
+ def __init__(self, num_branch: int, test_branch_idx: int,
+ **kwargs) -> None:
+ self.num_branch = num_branch
+ self.test_branch_idx = test_branch_idx
+ super().__init__(**kwargs)
+
+ def merge_trident_bboxes(self,
+ trident_results: InstanceList) -> InstanceData:
+ """Merge bbox predictions of each branch.
+
+ Args:
+ trident_results (List[:obj:`InstanceData`]): A list of InstanceData
+ predicted from every branch.
+
+ Returns:
+ :obj:`InstanceData`: merged InstanceData.
+ """
+ bboxes = torch.cat([res.bboxes for res in trident_results])
+ scores = torch.cat([res.scores for res in trident_results])
+ labels = torch.cat([res.labels for res in trident_results])
+
+ nms_cfg = self.test_cfg['nms']
+ results = InstanceData()
+ if bboxes.numel() == 0:
+ results.bboxes = bboxes
+ results.scores = scores
+ results.labels = labels
+ else:
+ det_bboxes, keep = batched_nms(bboxes, scores, labels, nms_cfg)
+ results.bboxes = det_bboxes[:, :-1]
+ results.scores = det_bboxes[:, -1]
+ results.labels = labels[keep]
+
+ if self.test_cfg['max_per_img'] > 0:
+ results = results[:self.test_cfg['max_per_img']]
+ return results
+
+ def predict(self,
+ x: Tuple[Tensor],
+ rpn_results_list: InstanceList,
+ batch_data_samples: SampleList,
+ rescale: bool = False) -> InstanceList:
+ """Perform forward propagation of the roi head and predict detection
+ results on the features of the upstream network.
+
+ - Compute prediction bbox and label per branch.
+ - Merge predictions of each branch according to scores of
+ bboxes, i.e., bboxes with higher score are kept to give
+ top-k prediction.
+
+ Args:
+ x (tuple[Tensor]): Features from upstream network. Each
+ has shape (N, C, H, W).
+ rpn_results_list (list[:obj:`InstanceData`]): list of region
+ proposals.
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool): Whether to rescale the results to
+ the original image. Defaults to True.
+
+ Returns:
+ list[obj:`InstanceData`]: Detection results of each image.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ results_list = super().predict(
+ x=x,
+ rpn_results_list=rpn_results_list,
+ batch_data_samples=batch_data_samples,
+ rescale=rescale)
+
+ num_branch = self.num_branch \
+ if self.training or self.test_branch_idx == -1 else 1
+
+ merged_results_list = []
+ for i in range(len(batch_data_samples) // num_branch):
+ merged_results_list.append(
+ self.merge_trident_bboxes(results_list[i * num_branch:(i + 1) *
+ num_branch]))
+ return merged_results_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..b489a905b1e9b6cef2e8b9575600990563128e4e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/__init__.py
@@ -0,0 +1,3 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .panoptic_fpn_head import PanopticFPNHead # noqa: F401,F403
+from .panoptic_fusion_heads import * # noqa: F401,F403
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/base_semantic_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/base_semantic_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..1db71549d89766c45012517c20cef443f4760419
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/base_semantic_head.py
@@ -0,0 +1,113 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from abc import ABCMeta, abstractmethod
+from typing import Dict, List, Tuple, Union
+
+import torch.nn.functional as F
+from mmengine.model import BaseModule
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.utils import ConfigType, OptMultiConfig
+
+
+@MODELS.register_module()
+class BaseSemanticHead(BaseModule, metaclass=ABCMeta):
+ """Base module of Semantic Head.
+
+ Args:
+ num_classes (int): the number of classes.
+ seg_rescale_factor (float): the rescale factor for ``gt_sem_seg``,
+ which equals to ``1 / output_strides``. The output_strides is
+ for ``seg_preds``. Defaults to 1 / 4.
+ init_cfg (Optional[Union[:obj:`ConfigDict`, dict]]): the initialization
+ config.
+ loss_seg (Union[:obj:`ConfigDict`, dict]): the loss of the semantic
+ head.
+ """
+
+ def __init__(self,
+ num_classes: int,
+ seg_rescale_factor: float = 1 / 4.,
+ loss_seg: ConfigType = dict(
+ type='CrossEntropyLoss',
+ ignore_index=255,
+ loss_weight=1.0),
+ init_cfg: OptMultiConfig = None) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.loss_seg = MODELS.build(loss_seg)
+ self.num_classes = num_classes
+ self.seg_rescale_factor = seg_rescale_factor
+
+ @abstractmethod
+ def forward(self, x: Union[Tensor, Tuple[Tensor]]) -> Dict[str, Tensor]:
+ """Placeholder of forward function.
+
+ Args:
+ x (Tensor): Feature maps.
+
+ Returns:
+ Dict[str, Tensor]: A dictionary, including features
+ and predicted scores. Required keys: 'seg_preds'
+ and 'feats'.
+ """
+ pass
+
+ @abstractmethod
+ def loss(self, x: Union[Tensor, Tuple[Tensor]],
+ batch_data_samples: SampleList) -> Dict[str, Tensor]:
+ """
+ Args:
+ x (Union[Tensor, Tuple[Tensor]]): Feature maps.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Args:
+ x (Tensor): Feature maps.
+
+ Returns:
+ Dict[str, Tensor]: The loss of semantic head.
+ """
+ pass
+
+ def predict(self,
+ x: Union[Tensor, Tuple[Tensor]],
+ batch_img_metas: List[dict],
+ rescale: bool = False) -> List[Tensor]:
+ """Test without Augmentation.
+
+ Args:
+ x (Union[Tensor, Tuple[Tensor]]): Feature maps.
+ batch_img_metas (List[dict]): List of image information.
+ rescale (bool): Whether to rescale the results.
+ Defaults to False.
+
+ Returns:
+ list[Tensor]: semantic segmentation logits.
+ """
+ seg_preds = self.forward(x)['seg_preds']
+ seg_preds = F.interpolate(
+ seg_preds,
+ size=batch_img_metas[0]['batch_input_shape'],
+ mode='bilinear',
+ align_corners=False)
+ seg_preds = [seg_preds[i] for i in range(len(batch_img_metas))]
+
+ if rescale:
+ seg_pred_list = []
+ for i in range(len(batch_img_metas)):
+ h, w = batch_img_metas[i]['img_shape']
+ seg_pred = seg_preds[i][:, :h, :w]
+
+ h, w = batch_img_metas[i]['ori_shape']
+ seg_pred = F.interpolate(
+ seg_pred[None],
+ size=(h, w),
+ mode='bilinear',
+ align_corners=False)[0]
+ seg_pred_list.append(seg_pred)
+ else:
+ seg_pred_list = seg_preds
+
+ return seg_pred_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/panoptic_fpn_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/panoptic_fpn_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..8d8b901360922f6cdb9f8d15b60dac8d7514ee75
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/panoptic_fpn_head.py
@@ -0,0 +1,174 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, Tuple, Union
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmengine.model import ModuleList
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.utils import ConfigType, OptConfigType, OptMultiConfig
+from ..layers import ConvUpsample
+from ..utils import interpolate_as
+from .base_semantic_head import BaseSemanticHead
+
+
+@MODELS.register_module()
+class PanopticFPNHead(BaseSemanticHead):
+ """PanopticFPNHead used in Panoptic FPN.
+
+ In this head, the number of output channels is ``num_stuff_classes
+ + 1``, including all stuff classes and one thing class. The stuff
+ classes will be reset from ``0`` to ``num_stuff_classes - 1``, the
+ thing classes will be merged to ``num_stuff_classes``-th channel.
+
+ Arg:
+ num_things_classes (int): Number of thing classes. Default: 80.
+ num_stuff_classes (int): Number of stuff classes. Default: 53.
+ in_channels (int): Number of channels in the input feature
+ map.
+ inner_channels (int): Number of channels in inner features.
+ start_level (int): The start level of the input features
+ used in PanopticFPN.
+ end_level (int): The end level of the used features, the
+ ``end_level``-th layer will not be used.
+ conv_cfg (Optional[Union[ConfigDict, dict]]): Dictionary to construct
+ and config conv layer.
+ norm_cfg (Union[ConfigDict, dict]): Dictionary to construct and config
+ norm layer. Use ``GN`` by default.
+ init_cfg (Optional[Union[ConfigDict, dict]]): Initialization config
+ dict.
+ loss_seg (Union[ConfigDict, dict]): the loss of the semantic head.
+ """
+
+ def __init__(self,
+ num_things_classes: int = 80,
+ num_stuff_classes: int = 53,
+ in_channels: int = 256,
+ inner_channels: int = 128,
+ start_level: int = 0,
+ end_level: int = 4,
+ conv_cfg: OptConfigType = None,
+ norm_cfg: ConfigType = dict(
+ type='GN', num_groups=32, requires_grad=True),
+ loss_seg: ConfigType = dict(
+ type='CrossEntropyLoss', ignore_index=-1,
+ loss_weight=1.0),
+ init_cfg: OptMultiConfig = None) -> None:
+ seg_rescale_factor = 1 / 2**(start_level + 2)
+ super().__init__(
+ num_classes=num_stuff_classes + 1,
+ seg_rescale_factor=seg_rescale_factor,
+ loss_seg=loss_seg,
+ init_cfg=init_cfg)
+ self.num_things_classes = num_things_classes
+ self.num_stuff_classes = num_stuff_classes
+ # Used feature layers are [start_level, end_level)
+ self.start_level = start_level
+ self.end_level = end_level
+ self.num_stages = end_level - start_level
+ self.inner_channels = inner_channels
+
+ self.conv_upsample_layers = ModuleList()
+ for i in range(start_level, end_level):
+ self.conv_upsample_layers.append(
+ ConvUpsample(
+ in_channels,
+ inner_channels,
+ num_layers=i if i > 0 else 1,
+ num_upsample=i if i > 0 else 0,
+ conv_cfg=conv_cfg,
+ norm_cfg=norm_cfg,
+ ))
+ self.conv_logits = nn.Conv2d(inner_channels, self.num_classes, 1)
+
+ def _set_things_to_void(self, gt_semantic_seg: Tensor) -> Tensor:
+ """Merge thing classes to one class.
+
+ In PanopticFPN, the background labels will be reset from `0` to
+ `self.num_stuff_classes-1`, the foreground labels will be merged to
+ `self.num_stuff_classes`-th channel.
+ """
+ gt_semantic_seg = gt_semantic_seg.int()
+ fg_mask = gt_semantic_seg < self.num_things_classes
+ bg_mask = (gt_semantic_seg >= self.num_things_classes) * (
+ gt_semantic_seg < self.num_things_classes + self.num_stuff_classes)
+
+ new_gt_seg = torch.clone(gt_semantic_seg)
+ new_gt_seg = torch.where(bg_mask,
+ gt_semantic_seg - self.num_things_classes,
+ new_gt_seg)
+ new_gt_seg = torch.where(fg_mask,
+ fg_mask.int() * self.num_stuff_classes,
+ new_gt_seg)
+ return new_gt_seg
+
+ def loss(self, x: Union[Tensor, Tuple[Tensor]],
+ batch_data_samples: SampleList) -> Dict[str, Tensor]:
+ """
+ Args:
+ x (Union[Tensor, Tuple[Tensor]]): Feature maps.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ Dict[str, Tensor]: The loss of semantic head.
+ """
+ seg_preds = self(x)['seg_preds']
+ gt_semantic_segs = [
+ data_sample.gt_sem_seg.sem_seg
+ for data_sample in batch_data_samples
+ ]
+
+ gt_semantic_segs = torch.stack(gt_semantic_segs)
+ if self.seg_rescale_factor != 1.0:
+ gt_semantic_segs = F.interpolate(
+ gt_semantic_segs.float(),
+ scale_factor=self.seg_rescale_factor,
+ mode='nearest').squeeze(1)
+
+ # Things classes will be merged to one class in PanopticFPN.
+ gt_semantic_segs = self._set_things_to_void(gt_semantic_segs)
+
+ if seg_preds.shape[-2:] != gt_semantic_segs.shape[-2:]:
+ seg_preds = interpolate_as(seg_preds, gt_semantic_segs)
+ seg_preds = seg_preds.permute((0, 2, 3, 1))
+
+ loss_seg = self.loss_seg(
+ seg_preds.reshape(-1, self.num_classes), # => [NxHxW, C]
+ gt_semantic_segs.reshape(-1).long())
+
+ return dict(loss_seg=loss_seg)
+
+ def init_weights(self) -> None:
+ """Initialize weights."""
+ super().init_weights()
+ nn.init.normal_(self.conv_logits.weight.data, 0, 0.01)
+ self.conv_logits.bias.data.zero_()
+
+ def forward(self, x: Tuple[Tensor]) -> Dict[str, Tensor]:
+ """Forward.
+
+ Args:
+ x (Tuple[Tensor]): Multi scale Feature maps.
+
+ Returns:
+ dict[str, Tensor]: semantic segmentation predictions and
+ feature maps.
+ """
+ # the number of subnets must be not more than
+ # the length of features.
+ assert self.num_stages <= len(x)
+
+ feats = []
+ for i, layer in enumerate(self.conv_upsample_layers):
+ f = layer(x[self.start_level + i])
+ feats.append(f)
+
+ seg_feats = torch.sum(torch.stack(feats, dim=0), dim=0)
+ seg_preds = self.conv_logits(seg_feats)
+ out = dict(seg_preds=seg_preds, seg_feats=seg_feats)
+ return out
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/panoptic_fusion_heads/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/panoptic_fusion_heads/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..41625a61d6d1c38c633062c24b1e3455bd3ae2df
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/panoptic_fusion_heads/__init__.py
@@ -0,0 +1,5 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .base_panoptic_fusion_head import \
+ BasePanopticFusionHead # noqa: F401,F403
+from .heuristic_fusion_head import HeuristicFusionHead # noqa: F401,F403
+from .maskformer_fusion_head import MaskFormerFusionHead # noqa: F401,F403
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/panoptic_fusion_heads/base_panoptic_fusion_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/panoptic_fusion_heads/base_panoptic_fusion_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..f6b20e1cd144eaebd042b8017f143c0a643adde1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/panoptic_fusion_heads/base_panoptic_fusion_head.py
@@ -0,0 +1,43 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from abc import ABCMeta, abstractmethod
+
+from mmengine.model import BaseModule
+
+from mmdet.registry import MODELS
+from mmdet.utils import OptConfigType, OptMultiConfig
+
+
+@MODELS.register_module()
+class BasePanopticFusionHead(BaseModule, metaclass=ABCMeta):
+ """Base class for panoptic heads."""
+
+ def __init__(self,
+ num_things_classes: int = 80,
+ num_stuff_classes: int = 53,
+ test_cfg: OptConfigType = None,
+ loss_panoptic: OptConfigType = None,
+ init_cfg: OptMultiConfig = None,
+ **kwargs) -> None:
+ super().__init__(init_cfg=init_cfg)
+ self.num_things_classes = num_things_classes
+ self.num_stuff_classes = num_stuff_classes
+ self.num_classes = num_things_classes + num_stuff_classes
+ self.test_cfg = test_cfg
+
+ if loss_panoptic:
+ self.loss_panoptic = MODELS.build(loss_panoptic)
+ else:
+ self.loss_panoptic = None
+
+ @property
+ def with_loss(self) -> bool:
+ """bool: whether the panoptic head contains loss function."""
+ return self.loss_panoptic is not None
+
+ @abstractmethod
+ def loss(self, **kwargs):
+ """Loss function."""
+
+ @abstractmethod
+ def predict(self, **kwargs):
+ """Predict function."""
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/panoptic_fusion_heads/heuristic_fusion_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/panoptic_fusion_heads/heuristic_fusion_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..7a4a4200edd97f42e9a138e14a1d07328ad9b139
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/panoptic_fusion_heads/heuristic_fusion_head.py
@@ -0,0 +1,159 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List
+
+import torch
+from mmengine.structures import InstanceData, PixelData
+from torch import Tensor
+
+from mmdet.evaluation.functional import INSTANCE_OFFSET
+from mmdet.registry import MODELS
+from mmdet.utils import InstanceList, OptConfigType, OptMultiConfig, PixelList
+from .base_panoptic_fusion_head import BasePanopticFusionHead
+
+
+@MODELS.register_module()
+class HeuristicFusionHead(BasePanopticFusionHead):
+ """Fusion Head with Heuristic method."""
+
+ def __init__(self,
+ num_things_classes: int = 80,
+ num_stuff_classes: int = 53,
+ test_cfg: OptConfigType = None,
+ init_cfg: OptMultiConfig = None,
+ **kwargs) -> None:
+ super().__init__(
+ num_things_classes=num_things_classes,
+ num_stuff_classes=num_stuff_classes,
+ test_cfg=test_cfg,
+ loss_panoptic=None,
+ init_cfg=init_cfg,
+ **kwargs)
+
+ def loss(self, **kwargs) -> dict:
+ """HeuristicFusionHead has no training loss."""
+ return dict()
+
+ def _lay_masks(self,
+ mask_results: InstanceData,
+ overlap_thr: float = 0.5) -> Tensor:
+ """Lay instance masks to a result map.
+
+ Args:
+ mask_results (:obj:`InstanceData`): Instance segmentation results,
+ each contains ``bboxes``, ``labels``, ``scores`` and ``masks``.
+ overlap_thr (float): Threshold to determine whether two masks
+ overlap. default: 0.5.
+
+ Returns:
+ Tensor: The result map, (H, W).
+ """
+ bboxes = mask_results.bboxes
+ scores = mask_results.scores
+ labels = mask_results.labels
+ masks = mask_results.masks
+
+ num_insts = bboxes.shape[0]
+ id_map = torch.zeros(
+ masks.shape[-2:], device=bboxes.device, dtype=torch.long)
+ if num_insts == 0:
+ return id_map, labels
+
+ # Sort by score to use heuristic fusion
+ order = torch.argsort(-scores)
+ bboxes = bboxes[order]
+ labels = labels[order]
+ segm_masks = masks[order]
+
+ instance_id = 1
+ left_labels = []
+ for idx in range(bboxes.shape[0]):
+ _cls = labels[idx]
+ _mask = segm_masks[idx]
+ instance_id_map = torch.ones_like(
+ _mask, dtype=torch.long) * instance_id
+ area = _mask.sum()
+ if area == 0:
+ continue
+
+ pasted = id_map > 0
+ intersect = (_mask * pasted).sum()
+ if (intersect / (area + 1e-5)) > overlap_thr:
+ continue
+
+ _part = _mask * (~pasted)
+ id_map = torch.where(_part, instance_id_map, id_map)
+ left_labels.append(_cls)
+ instance_id += 1
+
+ if len(left_labels) > 0:
+ instance_labels = torch.stack(left_labels)
+ else:
+ instance_labels = bboxes.new_zeros((0, ), dtype=torch.long)
+ assert instance_id == (len(instance_labels) + 1)
+ return id_map, instance_labels
+
+ def _predict_single(self, mask_results: InstanceData, seg_preds: Tensor,
+ **kwargs) -> PixelData:
+ """Fuse the results of instance and semantic segmentations.
+
+ Args:
+ mask_results (:obj:`InstanceData`): Instance segmentation results,
+ each contains ``bboxes``, ``labels``, ``scores`` and ``masks``.
+ seg_preds (Tensor): The semantic segmentation results,
+ (num_stuff + 1, H, W).
+
+ Returns:
+ Tensor: The panoptic segmentation result, (H, W).
+ """
+ id_map, labels = self._lay_masks(mask_results,
+ self.test_cfg.mask_overlap)
+
+ seg_results = seg_preds.argmax(dim=0)
+ seg_results = seg_results + self.num_things_classes
+
+ pan_results = seg_results
+ instance_id = 1
+ for idx in range(len(mask_results)):
+ _mask = id_map == (idx + 1)
+ if _mask.sum() == 0:
+ continue
+ _cls = labels[idx]
+ # simply trust detection
+ segment_id = _cls + instance_id * INSTANCE_OFFSET
+ pan_results[_mask] = segment_id
+ instance_id += 1
+
+ ids, counts = torch.unique(
+ pan_results % INSTANCE_OFFSET, return_counts=True)
+ stuff_ids = ids[ids >= self.num_things_classes]
+ stuff_counts = counts[ids >= self.num_things_classes]
+ ignore_stuff_ids = stuff_ids[
+ stuff_counts < self.test_cfg.stuff_area_limit]
+
+ assert pan_results.ndim == 2
+ pan_results[(pan_results.unsqueeze(2) == ignore_stuff_ids.reshape(
+ 1, 1, -1)).any(dim=2)] = self.num_classes
+
+ pan_results = PixelData(sem_seg=pan_results[None].int())
+ return pan_results
+
+ def predict(self, mask_results_list: InstanceList,
+ seg_preds_list: List[Tensor], **kwargs) -> PixelList:
+ """Predict results by fusing the results of instance and semantic
+ segmentations.
+
+ Args:
+ mask_results_list (list[:obj:`InstanceData`]): Instance
+ segmentation results, each contains ``bboxes``, ``labels``,
+ ``scores`` and ``masks``.
+ seg_preds_list (Tensor): List of semantic segmentation results.
+
+ Returns:
+ List[PixelData]: Panoptic segmentation result.
+ """
+ results_list = [
+ self._predict_single(mask_results_list[i], seg_preds_list[i])
+ for i in range(len(mask_results_list))
+ ]
+
+ return results_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/panoptic_fusion_heads/maskformer_fusion_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/panoptic_fusion_heads/maskformer_fusion_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..1b76e6b45bb9be2584f8b3eca2e5e1c0809249fa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/seg_heads/panoptic_fusion_heads/maskformer_fusion_head.py
@@ -0,0 +1,266 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List
+
+import torch
+import torch.nn.functional as F
+from mmengine.structures import InstanceData, PixelData
+from torch import Tensor
+
+from mmdet.evaluation.functional import INSTANCE_OFFSET
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.structures.mask import mask2bbox
+from mmdet.utils import OptConfigType, OptMultiConfig
+from .base_panoptic_fusion_head import BasePanopticFusionHead
+
+
+@MODELS.register_module()
+class MaskFormerFusionHead(BasePanopticFusionHead):
+ """MaskFormer fusion head which postprocesses results for panoptic
+ segmentation, instance segmentation and semantic segmentation."""
+
+ def __init__(self,
+ num_things_classes: int = 80,
+ num_stuff_classes: int = 53,
+ test_cfg: OptConfigType = None,
+ loss_panoptic: OptConfigType = None,
+ init_cfg: OptMultiConfig = None,
+ **kwargs):
+ super().__init__(
+ num_things_classes=num_things_classes,
+ num_stuff_classes=num_stuff_classes,
+ test_cfg=test_cfg,
+ loss_panoptic=loss_panoptic,
+ init_cfg=init_cfg,
+ **kwargs)
+
+ def loss(self, **kwargs):
+ """MaskFormerFusionHead has no training loss."""
+ return dict()
+
+ def panoptic_postprocess(self, mask_cls: Tensor,
+ mask_pred: Tensor) -> PixelData:
+ """Panoptic segmengation inference.
+
+ Args:
+ mask_cls (Tensor): Classfication outputs of shape
+ (num_queries, cls_out_channels) for a image.
+ Note `cls_out_channels` should includes
+ background.
+ mask_pred (Tensor): Mask outputs of shape
+ (num_queries, h, w) for a image.
+
+ Returns:
+ :obj:`PixelData`: Panoptic segment result of shape \
+ (h, w), each element in Tensor means: \
+ ``segment_id = _cls + instance_id * INSTANCE_OFFSET``.
+ """
+ object_mask_thr = self.test_cfg.get('object_mask_thr', 0.8)
+ iou_thr = self.test_cfg.get('iou_thr', 0.8)
+ filter_low_score = self.test_cfg.get('filter_low_score', False)
+
+ scores, labels = F.softmax(mask_cls, dim=-1).max(-1)
+ mask_pred = mask_pred.sigmoid()
+
+ keep = labels.ne(self.num_classes) & (scores > object_mask_thr)
+ cur_scores = scores[keep]
+ cur_classes = labels[keep]
+ cur_masks = mask_pred[keep]
+
+ cur_prob_masks = cur_scores.view(-1, 1, 1) * cur_masks
+
+ h, w = cur_masks.shape[-2:]
+ panoptic_seg = torch.full((h, w),
+ self.num_classes,
+ dtype=torch.int32,
+ device=cur_masks.device)
+ if cur_masks.shape[0] == 0:
+ # We didn't detect any mask :(
+ pass
+ else:
+ cur_mask_ids = cur_prob_masks.argmax(0)
+ instance_id = 1
+ for k in range(cur_classes.shape[0]):
+ pred_class = int(cur_classes[k].item())
+ isthing = pred_class < self.num_things_classes
+ mask = cur_mask_ids == k
+ mask_area = mask.sum().item()
+ original_area = (cur_masks[k] >= 0.5).sum().item()
+
+ if filter_low_score:
+ mask = mask & (cur_masks[k] >= 0.5)
+
+ if mask_area > 0 and original_area > 0:
+ if mask_area / original_area < iou_thr:
+ continue
+
+ if not isthing:
+ # different stuff regions of same class will be
+ # merged here, and stuff share the instance_id 0.
+ panoptic_seg[mask] = pred_class
+ else:
+ panoptic_seg[mask] = (
+ pred_class + instance_id * INSTANCE_OFFSET)
+ instance_id += 1
+
+ return PixelData(sem_seg=panoptic_seg[None])
+
+ def semantic_postprocess(self, mask_cls: Tensor,
+ mask_pred: Tensor) -> PixelData:
+ """Semantic segmengation postprocess.
+
+ Args:
+ mask_cls (Tensor): Classfication outputs of shape
+ (num_queries, cls_out_channels) for a image.
+ Note `cls_out_channels` should includes
+ background.
+ mask_pred (Tensor): Mask outputs of shape
+ (num_queries, h, w) for a image.
+
+ Returns:
+ :obj:`PixelData`: Semantic segment result.
+ """
+ # TODO add semantic segmentation result
+ raise NotImplementedError
+
+ def instance_postprocess(self, mask_cls: Tensor,
+ mask_pred: Tensor) -> InstanceData:
+ """Instance segmengation postprocess.
+
+ Args:
+ mask_cls (Tensor): Classfication outputs of shape
+ (num_queries, cls_out_channels) for a image.
+ Note `cls_out_channels` should includes
+ background.
+ mask_pred (Tensor): Mask outputs of shape
+ (num_queries, h, w) for a image.
+
+ Returns:
+ :obj:`InstanceData`: Instance segmentation results.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+ """
+ max_per_image = self.test_cfg.get('max_per_image', 100)
+ num_queries = mask_cls.shape[0]
+ # shape (num_queries, num_class)
+ scores = F.softmax(mask_cls, dim=-1)[:, :-1]
+ # shape (num_queries * num_class, )
+ labels = torch.arange(self.num_classes, device=mask_cls.device).\
+ unsqueeze(0).repeat(num_queries, 1).flatten(0, 1)
+ scores_per_image, top_indices = scores.flatten(0, 1).topk(
+ max_per_image, sorted=False)
+ labels_per_image = labels[top_indices]
+
+ query_indices = top_indices // self.num_classes
+ mask_pred = mask_pred[query_indices]
+
+ # extract things
+ is_thing = labels_per_image < self.num_things_classes
+ scores_per_image = scores_per_image[is_thing]
+ labels_per_image = labels_per_image[is_thing]
+ mask_pred = mask_pred[is_thing]
+
+ mask_pred_binary = (mask_pred > 0).float()
+ mask_scores_per_image = (mask_pred.sigmoid() *
+ mask_pred_binary).flatten(1).sum(1) / (
+ mask_pred_binary.flatten(1).sum(1) + 1e-6)
+ det_scores = scores_per_image * mask_scores_per_image
+ mask_pred_binary = mask_pred_binary.bool()
+ bboxes = mask2bbox(mask_pred_binary)
+
+ results = InstanceData()
+ results.bboxes = bboxes
+ results.labels = labels_per_image
+ results.scores = det_scores
+ results.masks = mask_pred_binary
+ return results
+
+ def predict(self,
+ mask_cls_results: Tensor,
+ mask_pred_results: Tensor,
+ batch_data_samples: SampleList,
+ rescale: bool = False,
+ **kwargs) -> List[dict]:
+ """Test segment without test-time aumengtation.
+
+ Only the output of last decoder layers was used.
+
+ Args:
+ mask_cls_results (Tensor): Mask classification logits,
+ shape (batch_size, num_queries, cls_out_channels).
+ Note `cls_out_channels` should includes background.
+ mask_pred_results (Tensor): Mask logits, shape
+ (batch_size, num_queries, h, w).
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool): If True, return boxes in
+ original image space. Default False.
+
+ Returns:
+ list[dict]: Instance segmentation \
+ results and panoptic segmentation results for each \
+ image.
+
+ .. code-block:: none
+
+ [
+ {
+ 'pan_results': PixelData,
+ 'ins_results': InstanceData,
+ # semantic segmentation results are not supported yet
+ 'sem_results': PixelData
+ },
+ ...
+ ]
+ """
+ batch_img_metas = [
+ data_sample.metainfo for data_sample in batch_data_samples
+ ]
+ panoptic_on = self.test_cfg.get('panoptic_on', True)
+ semantic_on = self.test_cfg.get('semantic_on', False)
+ instance_on = self.test_cfg.get('instance_on', False)
+ assert not semantic_on, 'segmantic segmentation '\
+ 'results are not supported yet.'
+
+ results = []
+ for mask_cls_result, mask_pred_result, meta in zip(
+ mask_cls_results, mask_pred_results, batch_img_metas):
+ # remove padding
+ img_height, img_width = meta['img_shape'][:2]
+ mask_pred_result = mask_pred_result[:, :img_height, :img_width]
+
+ if rescale:
+ # return result in original resolution
+ ori_height, ori_width = meta['ori_shape'][:2]
+ mask_pred_result = F.interpolate(
+ mask_pred_result[:, None],
+ size=(ori_height, ori_width),
+ mode='bilinear',
+ align_corners=False)[:, 0]
+
+ result = dict()
+ if panoptic_on:
+ pan_results = self.panoptic_postprocess(
+ mask_cls_result, mask_pred_result)
+ result['pan_results'] = pan_results
+
+ if instance_on:
+ ins_results = self.instance_postprocess(
+ mask_cls_result, mask_pred_result)
+ result['ins_results'] = ins_results
+
+ if semantic_on:
+ sem_results = self.semantic_postprocess(
+ mask_cls_result, mask_pred_result)
+ result['sem_results'] = sem_results
+
+ results.append(result)
+
+ return results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..7bfd8f058ed656760e0b1a3fd6118f31a799cb11
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/__init__.py
@@ -0,0 +1,18 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .assigners import * # noqa: F401,F403
+from .builder import (ANCHOR_GENERATORS, BBOX_ASSIGNERS, BBOX_CODERS,
+ BBOX_SAMPLERS, IOU_CALCULATORS, MATCH_COSTS,
+ PRIOR_GENERATORS, build_anchor_generator, build_assigner,
+ build_bbox_coder, build_iou_calculator, build_match_cost,
+ build_prior_generator, build_sampler)
+from .coders import * # noqa: F401,F403
+from .prior_generators import * # noqa: F401,F403
+from .samplers import * # noqa: F401,F403
+from .tracking import * # noqa: F401,F403
+
+__all__ = [
+ 'ANCHOR_GENERATORS', 'PRIOR_GENERATORS', 'BBOX_ASSIGNERS', 'BBOX_SAMPLERS',
+ 'MATCH_COSTS', 'BBOX_CODERS', 'IOU_CALCULATORS', 'build_anchor_generator',
+ 'build_prior_generator', 'build_assigner', 'build_sampler',
+ 'build_iou_calculator', 'build_match_cost', 'build_bbox_coder'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..4e564f24c95b1cc6be8a35a1a309ebf10e582032
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/__init__.py
@@ -0,0 +1,32 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .approx_max_iou_assigner import ApproxMaxIoUAssigner
+from .assign_result import AssignResult
+from .atss_assigner import ATSSAssigner
+from .base_assigner import BaseAssigner
+from .center_region_assigner import CenterRegionAssigner
+from .dynamic_soft_label_assigner import DynamicSoftLabelAssigner
+from .grid_assigner import GridAssigner
+from .hungarian_assigner import HungarianAssigner
+from .iou2d_calculator import BboxOverlaps2D, BboxOverlaps2D_GLIP
+from .match_cost import (BBoxL1Cost, BinaryFocalLossCost, ClassificationCost,
+ CrossEntropyLossCost, DiceCost, FocalLossCost,
+ IoUCost)
+from .max_iou_assigner import MaxIoUAssigner
+from .multi_instance_assigner import MultiInstanceAssigner
+from .point_assigner import PointAssigner
+from .region_assigner import RegionAssigner
+from .sim_ota_assigner import SimOTAAssigner
+from .task_aligned_assigner import TaskAlignedAssigner
+from .topk_hungarian_assigner import TopkHungarianAssigner
+from .uniform_assigner import UniformAssigner
+
+__all__ = [
+ 'BaseAssigner', 'BinaryFocalLossCost', 'MaxIoUAssigner',
+ 'ApproxMaxIoUAssigner', 'AssignResult', 'PointAssigner', 'ATSSAssigner',
+ 'CenterRegionAssigner', 'GridAssigner', 'HungarianAssigner',
+ 'RegionAssigner', 'UniformAssigner', 'SimOTAAssigner',
+ 'TaskAlignedAssigner', 'TopkHungarianAssigner', 'BBoxL1Cost',
+ 'ClassificationCost', 'CrossEntropyLossCost', 'DiceCost', 'FocalLossCost',
+ 'IoUCost', 'BboxOverlaps2D', 'DynamicSoftLabelAssigner',
+ 'MultiInstanceAssigner', 'BboxOverlaps2D_GLIP'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/approx_max_iou_assigner.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/approx_max_iou_assigner.py
new file mode 100644
index 0000000000000000000000000000000000000000..471d54e578d640da242355b54cebe05658309ca2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/approx_max_iou_assigner.py
@@ -0,0 +1,162 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Union
+
+import torch
+from mmengine.config import ConfigDict
+from mmengine.structures import InstanceData
+
+from mmdet.registry import TASK_UTILS
+from .assign_result import AssignResult
+from .max_iou_assigner import MaxIoUAssigner
+
+
+@TASK_UTILS.register_module()
+class ApproxMaxIoUAssigner(MaxIoUAssigner):
+ """Assign a corresponding gt bbox or background to each bbox.
+
+ Each proposals will be assigned with an integer indicating the ground truth
+ index. (semi-positive index: gt label (0-based), -1: background)
+
+ - -1: negative sample, no assigned gt
+ - semi-positive integer: positive sample, index (0-based) of assigned gt
+
+ Args:
+ pos_iou_thr (float): IoU threshold for positive bboxes.
+ neg_iou_thr (float or tuple): IoU threshold for negative bboxes.
+ min_pos_iou (float): Minimum iou for a bbox to be considered as a
+ positive bbox. Positive samples can have smaller IoU than
+ pos_iou_thr due to the 4th step (assign max IoU sample to each gt).
+ gt_max_assign_all (bool): Whether to assign all bboxes with the same
+ highest overlap with some gt to that gt.
+ ignore_iof_thr (float): IoF threshold for ignoring bboxes (if
+ `gt_bboxes_ignore` is specified). Negative values mean not
+ ignoring any bboxes.
+ ignore_wrt_candidates (bool): Whether to compute the iof between
+ `bboxes` and `gt_bboxes_ignore`, or the contrary.
+ match_low_quality (bool): Whether to allow quality matches. This is
+ usually allowed for RPN and single stage detectors, but not allowed
+ in the second stage.
+ gpu_assign_thr (int): The upper bound of the number of GT for GPU
+ assign. When the number of gt is above this threshold, will assign
+ on CPU device. Negative values mean not assign on CPU.
+ iou_calculator (:obj:`ConfigDict` or dict): Config of overlaps
+ Calculator.
+ """
+
+ def __init__(
+ self,
+ pos_iou_thr: float,
+ neg_iou_thr: Union[float, tuple],
+ min_pos_iou: float = .0,
+ gt_max_assign_all: bool = True,
+ ignore_iof_thr: float = -1,
+ ignore_wrt_candidates: bool = True,
+ match_low_quality: bool = True,
+ gpu_assign_thr: int = -1,
+ iou_calculator: Union[ConfigDict, dict] = dict(type='BboxOverlaps2D')
+ ) -> None:
+ self.pos_iou_thr = pos_iou_thr
+ self.neg_iou_thr = neg_iou_thr
+ self.min_pos_iou = min_pos_iou
+ self.gt_max_assign_all = gt_max_assign_all
+ self.ignore_iof_thr = ignore_iof_thr
+ self.ignore_wrt_candidates = ignore_wrt_candidates
+ self.gpu_assign_thr = gpu_assign_thr
+ self.match_low_quality = match_low_quality
+ self.iou_calculator = TASK_UTILS.build(iou_calculator)
+
+ def assign(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ **kwargs) -> AssignResult:
+ """Assign gt to approxs.
+
+ This method assign a gt bbox to each group of approxs (bboxes),
+ each group of approxs is represent by a base approx (bbox) and
+ will be assigned with -1, or a semi-positive number.
+ background_label (-1) means negative sample,
+ semi-positive number is the index (0-based) of assigned gt.
+ The assignment is done in following steps, the order matters.
+
+ 1. assign every bbox to background_label (-1)
+ 2. use the max IoU of each group of approxs to assign
+ 2. assign proposals whose iou with all gts < neg_iou_thr to background
+ 3. for each bbox, if the iou with its nearest gt >= pos_iou_thr,
+ assign it to that bbox
+ 4. for each gt bbox, assign its nearest proposals (may be more than
+ one) to itself
+
+ Args:
+ pred_instances (:obj:`InstanceData`): Instances of model
+ predictions. It includes ``priors``, and the priors can
+ be anchors or points, or the bboxes predicted by the
+ previous stage, has shape (n, 4). ``approxs`` means the
+ group of approxs aligned with ``priors``, has shape
+ (n, num_approxs, 4).
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes``, with shape (k, 4),
+ and ``labels``, with shape (k, ).
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes``
+ attribute data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ :obj:`AssignResult`: The assign result.
+ """
+ squares = pred_instances.priors
+ approxs = pred_instances.approxs
+ gt_bboxes = gt_instances.bboxes
+ gt_labels = gt_instances.labels
+ gt_bboxes_ignore = None if gt_instances_ignore is None else \
+ gt_instances_ignore.get('bboxes', None)
+ approxs_per_octave = approxs.size(1)
+
+ num_squares = squares.size(0)
+ num_gts = gt_bboxes.size(0)
+
+ if num_squares == 0 or num_gts == 0:
+ # No predictions and/or truth, return empty assignment
+ overlaps = approxs.new(num_gts, num_squares)
+ assign_result = self.assign_wrt_overlaps(overlaps, gt_labels)
+ return assign_result
+
+ # re-organize anchors by approxs_per_octave x num_squares
+ approxs = torch.transpose(approxs, 0, 1).contiguous().view(-1, 4)
+ assign_on_cpu = True if (self.gpu_assign_thr > 0) and (
+ num_gts > self.gpu_assign_thr) else False
+ # compute overlap and assign gt on CPU when number of GT is large
+ if assign_on_cpu:
+ device = approxs.device
+ approxs = approxs.cpu()
+ gt_bboxes = gt_bboxes.cpu()
+ if gt_bboxes_ignore is not None:
+ gt_bboxes_ignore = gt_bboxes_ignore.cpu()
+ if gt_labels is not None:
+ gt_labels = gt_labels.cpu()
+ all_overlaps = self.iou_calculator(approxs, gt_bboxes)
+
+ overlaps, _ = all_overlaps.view(approxs_per_octave, num_squares,
+ num_gts).max(dim=0)
+ overlaps = torch.transpose(overlaps, 0, 1)
+
+ if (self.ignore_iof_thr > 0 and gt_bboxes_ignore is not None
+ and gt_bboxes_ignore.numel() > 0 and squares.numel() > 0):
+ if self.ignore_wrt_candidates:
+ ignore_overlaps = self.iou_calculator(
+ squares, gt_bboxes_ignore, mode='iof')
+ ignore_max_overlaps, _ = ignore_overlaps.max(dim=1)
+ else:
+ ignore_overlaps = self.iou_calculator(
+ gt_bboxes_ignore, squares, mode='iof')
+ ignore_max_overlaps, _ = ignore_overlaps.max(dim=0)
+ overlaps[:, ignore_max_overlaps > self.ignore_iof_thr] = -1
+
+ assign_result = self.assign_wrt_overlaps(overlaps, gt_labels)
+ if assign_on_cpu:
+ assign_result.gt_inds = assign_result.gt_inds.to(device)
+ assign_result.max_overlaps = assign_result.max_overlaps.to(device)
+ if assign_result.labels is not None:
+ assign_result.labels = assign_result.labels.to(device)
+ return assign_result
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/assign_result.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/assign_result.py
new file mode 100644
index 0000000000000000000000000000000000000000..56ca2c3c18fee94cc4a039b769e42521bd14907d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/assign_result.py
@@ -0,0 +1,198 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+from torch import Tensor
+
+from mmdet.utils import util_mixins
+
+
+class AssignResult(util_mixins.NiceRepr):
+ """Stores assignments between predicted and truth boxes.
+
+ Attributes:
+ num_gts (int): the number of truth boxes considered when computing this
+ assignment
+ gt_inds (Tensor): for each predicted box indicates the 1-based
+ index of the assigned truth box. 0 means unassigned and -1 means
+ ignore.
+ max_overlaps (Tensor): the iou between the predicted box and its
+ assigned truth box.
+ labels (Tensor): If specified, for each predicted box
+ indicates the category label of the assigned truth box.
+
+ Example:
+ >>> # An assign result between 4 predicted boxes and 9 true boxes
+ >>> # where only two boxes were assigned.
+ >>> num_gts = 9
+ >>> max_overlaps = torch.LongTensor([0, .5, .9, 0])
+ >>> gt_inds = torch.LongTensor([-1, 1, 2, 0])
+ >>> labels = torch.LongTensor([0, 3, 4, 0])
+ >>> self = AssignResult(num_gts, gt_inds, max_overlaps, labels)
+ >>> print(str(self)) # xdoctest: +IGNORE_WANT
+
+ >>> # Force addition of gt labels (when adding gt as proposals)
+ >>> new_labels = torch.LongTensor([3, 4, 5])
+ >>> self.add_gt_(new_labels)
+ >>> print(str(self)) # xdoctest: +IGNORE_WANT
+
+ """
+
+ def __init__(self, num_gts: int, gt_inds: Tensor, max_overlaps: Tensor,
+ labels: Tensor) -> None:
+ self.num_gts = num_gts
+ self.gt_inds = gt_inds
+ self.max_overlaps = max_overlaps
+ self.labels = labels
+ # Interface for possible user-defined properties
+ self._extra_properties = {}
+
+ @property
+ def num_preds(self):
+ """int: the number of predictions in this assignment"""
+ return len(self.gt_inds)
+
+ def set_extra_property(self, key, value):
+ """Set user-defined new property."""
+ assert key not in self.info
+ self._extra_properties[key] = value
+
+ def get_extra_property(self, key):
+ """Get user-defined property."""
+ return self._extra_properties.get(key, None)
+
+ @property
+ def info(self):
+ """dict: a dictionary of info about the object"""
+ basic_info = {
+ 'num_gts': self.num_gts,
+ 'num_preds': self.num_preds,
+ 'gt_inds': self.gt_inds,
+ 'max_overlaps': self.max_overlaps,
+ 'labels': self.labels,
+ }
+ basic_info.update(self._extra_properties)
+ return basic_info
+
+ def __nice__(self):
+ """str: a "nice" summary string describing this assign result"""
+ parts = []
+ parts.append(f'num_gts={self.num_gts!r}')
+ if self.gt_inds is None:
+ parts.append(f'gt_inds={self.gt_inds!r}')
+ else:
+ parts.append(f'gt_inds.shape={tuple(self.gt_inds.shape)!r}')
+ if self.max_overlaps is None:
+ parts.append(f'max_overlaps={self.max_overlaps!r}')
+ else:
+ parts.append('max_overlaps.shape='
+ f'{tuple(self.max_overlaps.shape)!r}')
+ if self.labels is None:
+ parts.append(f'labels={self.labels!r}')
+ else:
+ parts.append(f'labels.shape={tuple(self.labels.shape)!r}')
+ return ', '.join(parts)
+
+ @classmethod
+ def random(cls, **kwargs):
+ """Create random AssignResult for tests or debugging.
+
+ Args:
+ num_preds: number of predicted boxes
+ num_gts: number of true boxes
+ p_ignore (float): probability of a predicted box assigned to an
+ ignored truth
+ p_assigned (float): probability of a predicted box not being
+ assigned
+ p_use_label (float | bool): with labels or not
+ rng (None | int | numpy.random.RandomState): seed or state
+
+ Returns:
+ :obj:`AssignResult`: Randomly generated assign results.
+
+ Example:
+ >>> from mmdet.models.task_modules.assigners.assign_result import * # NOQA
+ >>> self = AssignResult.random()
+ >>> print(self.info)
+ """
+ from ..samplers.sampling_result import ensure_rng
+ rng = ensure_rng(kwargs.get('rng', None))
+
+ num_gts = kwargs.get('num_gts', None)
+ num_preds = kwargs.get('num_preds', None)
+ p_ignore = kwargs.get('p_ignore', 0.3)
+ p_assigned = kwargs.get('p_assigned', 0.7)
+ num_classes = kwargs.get('num_classes', 3)
+
+ if num_gts is None:
+ num_gts = rng.randint(0, 8)
+ if num_preds is None:
+ num_preds = rng.randint(0, 16)
+
+ if num_gts == 0:
+ max_overlaps = torch.zeros(num_preds, dtype=torch.float32)
+ gt_inds = torch.zeros(num_preds, dtype=torch.int64)
+ labels = torch.zeros(num_preds, dtype=torch.int64)
+
+ else:
+ import numpy as np
+
+ # Create an overlap for each predicted box
+ max_overlaps = torch.from_numpy(rng.rand(num_preds))
+
+ # Construct gt_inds for each predicted box
+ is_assigned = torch.from_numpy(rng.rand(num_preds) < p_assigned)
+ # maximum number of assignments constraints
+ n_assigned = min(num_preds, min(num_gts, is_assigned.sum()))
+
+ assigned_idxs = np.where(is_assigned)[0]
+ rng.shuffle(assigned_idxs)
+ assigned_idxs = assigned_idxs[0:n_assigned]
+ assigned_idxs.sort()
+
+ is_assigned[:] = 0
+ is_assigned[assigned_idxs] = True
+
+ is_ignore = torch.from_numpy(
+ rng.rand(num_preds) < p_ignore) & is_assigned
+
+ gt_inds = torch.zeros(num_preds, dtype=torch.int64)
+
+ true_idxs = np.arange(num_gts)
+ rng.shuffle(true_idxs)
+ true_idxs = torch.from_numpy(true_idxs)
+ gt_inds[is_assigned] = true_idxs[:n_assigned].long()
+
+ gt_inds = torch.from_numpy(
+ rng.randint(1, num_gts + 1, size=num_preds))
+ gt_inds[is_ignore] = -1
+ gt_inds[~is_assigned] = 0
+ max_overlaps[~is_assigned] = 0
+
+ if num_classes == 0:
+ labels = torch.zeros(num_preds, dtype=torch.int64)
+ else:
+ labels = torch.from_numpy(
+ # remind that we set FG labels to [0, num_class-1]
+ # since mmdet v2.0
+ # BG cat_id: num_class
+ rng.randint(0, num_classes, size=num_preds))
+ labels[~is_assigned] = 0
+
+ self = cls(num_gts, gt_inds, max_overlaps, labels)
+ return self
+
+ def add_gt_(self, gt_labels):
+ """Add ground truth as assigned results.
+
+ Args:
+ gt_labels (torch.Tensor): Labels of gt boxes
+ """
+ self_inds = torch.arange(
+ 1, len(gt_labels) + 1, dtype=torch.long, device=gt_labels.device)
+ self.gt_inds = torch.cat([self_inds, self.gt_inds])
+
+ self.max_overlaps = torch.cat(
+ [self.max_overlaps.new_ones(len(gt_labels)), self.max_overlaps])
+
+ self.labels = torch.cat([gt_labels, self.labels])
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/atss_assigner.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/atss_assigner.py
new file mode 100644
index 0000000000000000000000000000000000000000..2796b990c5ae4c56bcf314e1342671d950232ae6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/atss_assigner.py
@@ -0,0 +1,254 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+from typing import List, Optional
+
+import torch
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import TASK_UTILS
+from mmdet.utils import ConfigType
+from .assign_result import AssignResult
+from .base_assigner import BaseAssigner
+
+
+def bbox_center_distance(bboxes: Tensor, priors: Tensor) -> Tensor:
+ """Compute the center distance between bboxes and priors.
+
+ Args:
+ bboxes (Tensor): Shape (n, 4) for , "xyxy" format.
+ priors (Tensor): Shape (n, 4) for priors, "xyxy" format.
+
+ Returns:
+ Tensor: Center distances between bboxes and priors.
+ """
+ bbox_cx = (bboxes[:, 0] + bboxes[:, 2]) / 2.0
+ bbox_cy = (bboxes[:, 1] + bboxes[:, 3]) / 2.0
+ bbox_points = torch.stack((bbox_cx, bbox_cy), dim=1)
+
+ priors_cx = (priors[:, 0] + priors[:, 2]) / 2.0
+ priors_cy = (priors[:, 1] + priors[:, 3]) / 2.0
+ priors_points = torch.stack((priors_cx, priors_cy), dim=1)
+
+ distances = (priors_points[:, None, :] -
+ bbox_points[None, :, :]).pow(2).sum(-1).sqrt()
+
+ return distances
+
+
+@TASK_UTILS.register_module()
+class ATSSAssigner(BaseAssigner):
+ """Assign a corresponding gt bbox or background to each prior.
+
+ Each proposals will be assigned with `0` or a positive integer
+ indicating the ground truth index.
+
+ - 0: negative sample, no assigned gt
+ - positive integer: positive sample, index (1-based) of assigned gt
+
+ If ``alpha`` is not None, it means that the dynamic cost
+ ATSSAssigner is adopted, which is currently only used in the DDOD.
+
+ Args:
+ topk (int): number of priors selected in each level
+ alpha (float, optional): param of cost rate for each proposal only
+ in DDOD. Defaults to None.
+ iou_calculator (:obj:`ConfigDict` or dict): Config dict for iou
+ calculator. Defaults to ``dict(type='BboxOverlaps2D')``
+ ignore_iof_thr (float): IoF threshold for ignoring bboxes (if
+ `gt_bboxes_ignore` is specified). Negative values mean not
+ ignoring any bboxes. Defaults to -1.
+ """
+
+ def __init__(self,
+ topk: int,
+ alpha: Optional[float] = None,
+ iou_calculator: ConfigType = dict(type='BboxOverlaps2D'),
+ ignore_iof_thr: float = -1) -> None:
+ self.topk = topk
+ self.alpha = alpha
+ self.iou_calculator = TASK_UTILS.build(iou_calculator)
+ self.ignore_iof_thr = ignore_iof_thr
+
+ # https://github.com/sfzhang15/ATSS/blob/master/atss_core/modeling/rpn/atss/loss.py
+ def assign(
+ self,
+ pred_instances: InstanceData,
+ num_level_priors: List[int],
+ gt_instances: InstanceData,
+ gt_instances_ignore: Optional[InstanceData] = None
+ ) -> AssignResult:
+ """Assign gt to priors.
+
+ The assignment is done in following steps
+
+ 1. compute iou between all prior (prior of all pyramid levels) and gt
+ 2. compute center distance between all prior and gt
+ 3. on each pyramid level, for each gt, select k prior whose center
+ are closest to the gt center, so we total select k*l prior as
+ candidates for each gt
+ 4. get corresponding iou for the these candidates, and compute the
+ mean and std, set mean + std as the iou threshold
+ 5. select these candidates whose iou are greater than or equal to
+ the threshold as positive
+ 6. limit the positive sample's center in gt
+
+ If ``alpha`` is not None, and ``cls_scores`` and `bbox_preds`
+ are not None, the overlaps calculation in the first step
+ will also include dynamic cost, which is currently only used in
+ the DDOD.
+
+ Args:
+ pred_instances (:obj:`InstaceData`): Instances of model
+ predictions. It includes ``priors``, and the priors can
+ be anchors, points, or bboxes predicted by the model,
+ shape(n, 4).
+ num_level_priors (List): Number of bboxes in each level
+ gt_instances (:obj:`InstaceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ gt_instances_ignore (:obj:`InstaceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes``
+ attribute data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ :obj:`AssignResult`: The assign result.
+ """
+ gt_bboxes = gt_instances.bboxes
+ priors = pred_instances.priors
+ gt_labels = gt_instances.labels
+ if gt_instances_ignore is not None:
+ gt_bboxes_ignore = gt_instances_ignore.bboxes
+ else:
+ gt_bboxes_ignore = None
+
+ INF = 100000000
+ priors = priors[:, :4]
+ num_gt, num_priors = gt_bboxes.size(0), priors.size(0)
+
+ message = 'Invalid alpha parameter because cls_scores or ' \
+ 'bbox_preds are None. If you want to use the ' \
+ 'cost-based ATSSAssigner, please set cls_scores, ' \
+ 'bbox_preds and self.alpha at the same time. '
+
+ # compute iou between all bbox and gt
+ if self.alpha is None:
+ # ATSSAssigner
+ overlaps = self.iou_calculator(priors, gt_bboxes)
+ if ('scores' in pred_instances or 'bboxes' in pred_instances):
+ warnings.warn(message)
+
+ else:
+ # Dynamic cost ATSSAssigner in DDOD
+ assert ('scores' in pred_instances
+ and 'bboxes' in pred_instances), message
+ cls_scores = pred_instances.scores
+ bbox_preds = pred_instances.bboxes
+
+ # compute cls cost for bbox and GT
+ cls_cost = torch.sigmoid(cls_scores[:, gt_labels])
+
+ # compute iou between all bbox and gt
+ overlaps = self.iou_calculator(bbox_preds, gt_bboxes)
+
+ # make sure that we are in element-wise multiplication
+ assert cls_cost.shape == overlaps.shape
+
+ # overlaps is actually a cost matrix
+ overlaps = cls_cost**(1 - self.alpha) * overlaps**self.alpha
+
+ # assign 0 by default
+ assigned_gt_inds = overlaps.new_full((num_priors, ),
+ 0,
+ dtype=torch.long)
+
+ if num_gt == 0 or num_priors == 0:
+ # No ground truth or boxes, return empty assignment
+ max_overlaps = overlaps.new_zeros((num_priors, ))
+ if num_gt == 0:
+ # No truth, assign everything to background
+ assigned_gt_inds[:] = 0
+ assigned_labels = overlaps.new_full((num_priors, ),
+ -1,
+ dtype=torch.long)
+ return AssignResult(
+ num_gt, assigned_gt_inds, max_overlaps, labels=assigned_labels)
+
+ # compute center distance between all bbox and gt
+ distances = bbox_center_distance(gt_bboxes, priors)
+
+ if (self.ignore_iof_thr > 0 and gt_bboxes_ignore is not None
+ and gt_bboxes_ignore.numel() > 0 and priors.numel() > 0):
+ ignore_overlaps = self.iou_calculator(
+ priors, gt_bboxes_ignore, mode='iof')
+ ignore_max_overlaps, _ = ignore_overlaps.max(dim=1)
+ ignore_idxs = ignore_max_overlaps > self.ignore_iof_thr
+ distances[ignore_idxs, :] = INF
+ assigned_gt_inds[ignore_idxs] = -1
+
+ # Selecting candidates based on the center distance
+ candidate_idxs = []
+ start_idx = 0
+ for level, priors_per_level in enumerate(num_level_priors):
+ # on each pyramid level, for each gt,
+ # select k bbox whose center are closest to the gt center
+ end_idx = start_idx + priors_per_level
+ distances_per_level = distances[start_idx:end_idx, :]
+ selectable_k = min(self.topk, priors_per_level)
+ _, topk_idxs_per_level = distances_per_level.topk(
+ selectable_k, dim=0, largest=False)
+ candidate_idxs.append(topk_idxs_per_level + start_idx)
+ start_idx = end_idx
+ candidate_idxs = torch.cat(candidate_idxs, dim=0)
+
+ # get corresponding iou for the these candidates, and compute the
+ # mean and std, set mean + std as the iou threshold
+ candidate_overlaps = overlaps[candidate_idxs, torch.arange(num_gt)]
+ overlaps_mean_per_gt = candidate_overlaps.mean(0)
+ overlaps_std_per_gt = candidate_overlaps.std(0)
+ overlaps_thr_per_gt = overlaps_mean_per_gt + overlaps_std_per_gt
+
+ is_pos = candidate_overlaps >= overlaps_thr_per_gt[None, :]
+
+ # limit the positive sample's center in gt
+ for gt_idx in range(num_gt):
+ candidate_idxs[:, gt_idx] += gt_idx * num_priors
+ priors_cx = (priors[:, 0] + priors[:, 2]) / 2.0
+ priors_cy = (priors[:, 1] + priors[:, 3]) / 2.0
+ ep_priors_cx = priors_cx.view(1, -1).expand(
+ num_gt, num_priors).contiguous().view(-1)
+ ep_priors_cy = priors_cy.view(1, -1).expand(
+ num_gt, num_priors).contiguous().view(-1)
+ candidate_idxs = candidate_idxs.view(-1)
+
+ # calculate the left, top, right, bottom distance between positive
+ # prior center and gt side
+ l_ = ep_priors_cx[candidate_idxs].view(-1, num_gt) - gt_bboxes[:, 0]
+ t_ = ep_priors_cy[candidate_idxs].view(-1, num_gt) - gt_bboxes[:, 1]
+ r_ = gt_bboxes[:, 2] - ep_priors_cx[candidate_idxs].view(-1, num_gt)
+ b_ = gt_bboxes[:, 3] - ep_priors_cy[candidate_idxs].view(-1, num_gt)
+ is_in_gts = torch.stack([l_, t_, r_, b_], dim=1).min(dim=1)[0] > 0.01
+
+ is_pos = is_pos & is_in_gts
+
+ # if an anchor box is assigned to multiple gts,
+ # the one with the highest IoU will be selected.
+ overlaps_inf = torch.full_like(overlaps,
+ -INF).t().contiguous().view(-1)
+ index = candidate_idxs.view(-1)[is_pos.view(-1)]
+ overlaps_inf[index] = overlaps.t().contiguous().view(-1)[index]
+ overlaps_inf = overlaps_inf.view(num_gt, -1).t()
+
+ max_overlaps, argmax_overlaps = overlaps_inf.max(dim=1)
+ assigned_gt_inds[
+ max_overlaps != -INF] = argmax_overlaps[max_overlaps != -INF] + 1
+
+ assigned_labels = assigned_gt_inds.new_full((num_priors, ), -1)
+ pos_inds = torch.nonzero(
+ assigned_gt_inds > 0, as_tuple=False).squeeze()
+ if pos_inds.numel() > 0:
+ assigned_labels[pos_inds] = gt_labels[assigned_gt_inds[pos_inds] -
+ 1]
+ return AssignResult(
+ num_gt, assigned_gt_inds, max_overlaps, labels=assigned_labels)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/base_assigner.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/base_assigner.py
new file mode 100644
index 0000000000000000000000000000000000000000..b12280ad746c7557008313dd936a62a99e8c78d5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/base_assigner.py
@@ -0,0 +1,17 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from abc import ABCMeta, abstractmethod
+from typing import Optional
+
+from mmengine.structures import InstanceData
+
+
+class BaseAssigner(metaclass=ABCMeta):
+ """Base assigner that assigns boxes to ground truth boxes."""
+
+ @abstractmethod
+ def assign(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ **kwargs):
+ """Assign boxes to either a ground truth boxes or a negative boxes."""
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/center_region_assigner.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/center_region_assigner.py
new file mode 100644
index 0000000000000000000000000000000000000000..11c8055c67cdf46c1ae0f877e88192db33795581
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/center_region_assigner.py
@@ -0,0 +1,366 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Tuple
+
+import torch
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import TASK_UTILS
+from mmdet.utils import ConfigType
+from .assign_result import AssignResult
+from .base_assigner import BaseAssigner
+
+
+def scale_boxes(bboxes: Tensor, scale: float) -> Tensor:
+ """Expand an array of boxes by a given scale.
+
+ Args:
+ bboxes (Tensor): Shape (m, 4)
+ scale (float): The scale factor of bboxes
+
+ Returns:
+ Tensor: Shape (m, 4). Scaled bboxes
+ """
+ assert bboxes.size(1) == 4
+ w_half = (bboxes[:, 2] - bboxes[:, 0]) * .5
+ h_half = (bboxes[:, 3] - bboxes[:, 1]) * .5
+ x_c = (bboxes[:, 2] + bboxes[:, 0]) * .5
+ y_c = (bboxes[:, 3] + bboxes[:, 1]) * .5
+
+ w_half *= scale
+ h_half *= scale
+
+ boxes_scaled = torch.zeros_like(bboxes)
+ boxes_scaled[:, 0] = x_c - w_half
+ boxes_scaled[:, 2] = x_c + w_half
+ boxes_scaled[:, 1] = y_c - h_half
+ boxes_scaled[:, 3] = y_c + h_half
+ return boxes_scaled
+
+
+def is_located_in(points: Tensor, bboxes: Tensor) -> Tensor:
+ """Are points located in bboxes.
+
+ Args:
+ points (Tensor): Points, shape: (m, 2).
+ bboxes (Tensor): Bounding boxes, shape: (n, 4).
+
+ Return:
+ Tensor: Flags indicating if points are located in bboxes,
+ shape: (m, n).
+ """
+ assert points.size(1) == 2
+ assert bboxes.size(1) == 4
+ return (points[:, 0].unsqueeze(1) > bboxes[:, 0].unsqueeze(0)) & \
+ (points[:, 0].unsqueeze(1) < bboxes[:, 2].unsqueeze(0)) & \
+ (points[:, 1].unsqueeze(1) > bboxes[:, 1].unsqueeze(0)) & \
+ (points[:, 1].unsqueeze(1) < bboxes[:, 3].unsqueeze(0))
+
+
+def bboxes_area(bboxes: Tensor) -> Tensor:
+ """Compute the area of an array of bboxes.
+
+ Args:
+ bboxes (Tensor): The coordinates ox bboxes. Shape: (m, 4)
+
+ Returns:
+ Tensor: Area of the bboxes. Shape: (m, )
+ """
+ assert bboxes.size(1) == 4
+ w = (bboxes[:, 2] - bboxes[:, 0])
+ h = (bboxes[:, 3] - bboxes[:, 1])
+ areas = w * h
+ return areas
+
+
+@TASK_UTILS.register_module()
+class CenterRegionAssigner(BaseAssigner):
+ """Assign pixels at the center region of a bbox as positive.
+
+ Each proposals will be assigned with `-1`, `0`, or a positive integer
+ indicating the ground truth index.
+ - -1: negative samples
+ - semi-positive numbers: positive sample, index (0-based) of assigned gt
+
+ Args:
+ pos_scale (float): Threshold within which pixels are
+ labelled as positive.
+ neg_scale (float): Threshold above which pixels are
+ labelled as positive.
+ min_pos_iof (float): Minimum iof of a pixel with a gt to be
+ labelled as positive. Default: 1e-2
+ ignore_gt_scale (float): Threshold within which the pixels
+ are ignored when the gt is labelled as shadowed. Default: 0.5
+ foreground_dominate (bool): If True, the bbox will be assigned as
+ positive when a gt's kernel region overlaps with another's shadowed
+ (ignored) region, otherwise it is set as ignored. Default to False.
+ iou_calculator (:obj:`ConfigDict` or dict): Config of overlaps
+ Calculator.
+ """
+
+ def __init__(
+ self,
+ pos_scale: float,
+ neg_scale: float,
+ min_pos_iof: float = 1e-2,
+ ignore_gt_scale: float = 0.5,
+ foreground_dominate: bool = False,
+ iou_calculator: ConfigType = dict(type='BboxOverlaps2D')
+ ) -> None:
+ self.pos_scale = pos_scale
+ self.neg_scale = neg_scale
+ self.min_pos_iof = min_pos_iof
+ self.ignore_gt_scale = ignore_gt_scale
+ self.foreground_dominate = foreground_dominate
+ self.iou_calculator = TASK_UTILS.build(iou_calculator)
+
+ def get_gt_priorities(self, gt_bboxes: Tensor) -> Tensor:
+ """Get gt priorities according to their areas.
+
+ Smaller gt has higher priority.
+
+ Args:
+ gt_bboxes (Tensor): Ground truth boxes, shape (k, 4).
+
+ Returns:
+ Tensor: The priority of gts so that gts with larger priority is
+ more likely to be assigned. Shape (k, )
+ """
+ gt_areas = bboxes_area(gt_bboxes)
+ # Rank all gt bbox areas. Smaller objects has larger priority
+ _, sort_idx = gt_areas.sort(descending=True)
+ sort_idx = sort_idx.argsort()
+ return sort_idx
+
+ def assign(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ **kwargs) -> AssignResult:
+ """Assign gt to bboxes.
+
+ This method assigns gts to every prior (proposal/anchor), each prior
+ will be assigned with -1, or a semi-positive number. -1 means
+ negative sample, semi-positive number is the index (0-based) of
+ assigned gt.
+
+ Args:
+ pred_instances (:obj:`InstanceData`): Instances of model
+ predictions. It includes ``priors``, and the priors can
+ be anchors or points, or the bboxes predicted by the
+ previous stage, has shape (n, 4). The bboxes predicted by
+ the current model or stage will be named ``bboxes``,
+ ``labels``, and ``scores``, the same as the ``InstanceData``
+ in other places.
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes``, with shape (k, 4),
+ and ``labels``, with shape (k, ).
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes``
+ attribute data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ :obj:`AssignResult`: The assigned result. Note that shadowed_labels
+ of shape (N, 2) is also added as an `assign_result` attribute.
+ `shadowed_labels` is a tensor composed of N pairs of anchor_ind,
+ class_label], where N is the number of anchors that lie in the
+ outer region of a gt, anchor_ind is the shadowed anchor index
+ and class_label is the shadowed class label.
+
+ Example:
+ >>> from mmengine.structures import InstanceData
+ >>> self = CenterRegionAssigner(0.2, 0.2)
+ >>> pred_instances.priors = torch.Tensor([[0, 0, 10, 10],
+ ... [10, 10, 20, 20]])
+ >>> gt_instances = InstanceData()
+ >>> gt_instances.bboxes = torch.Tensor([[0, 0, 10, 10]])
+ >>> gt_instances.labels = torch.Tensor([0])
+ >>> assign_result = self.assign(pred_instances, gt_instances)
+ >>> expected_gt_inds = torch.LongTensor([1, 0])
+ >>> assert torch.all(assign_result.gt_inds == expected_gt_inds)
+ """
+ # There are in total 5 steps in the pixel assignment
+ # 1. Find core (the center region, say inner 0.2)
+ # and shadow (the relatively ourter part, say inner 0.2-0.5)
+ # regions of every gt.
+ # 2. Find all prior bboxes that lie in gt_core and gt_shadow regions
+ # 3. Assign prior bboxes in gt_core with a one-hot id of the gt in
+ # the image.
+ # 3.1. For overlapping objects, the prior bboxes in gt_core is
+ # assigned with the object with smallest area
+ # 4. Assign prior bboxes with class label according to its gt id.
+ # 4.1. Assign -1 to prior bboxes lying in shadowed gts
+ # 4.2. Assign positive prior boxes with the corresponding label
+ # 5. Find pixels lying in the shadow of an object and assign them with
+ # background label, but set the loss weight of its corresponding
+ # gt to zero.
+
+ # TODO not extract bboxes in assign.
+ gt_bboxes = gt_instances.bboxes
+ priors = pred_instances.priors
+ gt_labels = gt_instances.labels
+
+ assert priors.size(1) == 4, 'priors must have size of 4'
+ # 1. Find core positive and shadow region of every gt
+ gt_core = scale_boxes(gt_bboxes, self.pos_scale)
+ gt_shadow = scale_boxes(gt_bboxes, self.neg_scale)
+
+ # 2. Find prior bboxes that lie in gt_core and gt_shadow regions
+ prior_centers = (priors[:, 2:4] + priors[:, 0:2]) / 2
+ # The center points lie within the gt boxes
+ is_prior_in_gt = is_located_in(prior_centers, gt_bboxes)
+ # Only calculate prior and gt_core IoF. This enables small prior bboxes
+ # to match large gts
+ prior_and_gt_core_overlaps = self.iou_calculator(
+ priors, gt_core, mode='iof')
+ # The center point of effective priors should be within the gt box
+ is_prior_in_gt_core = is_prior_in_gt & (
+ prior_and_gt_core_overlaps > self.min_pos_iof) # shape (n, k)
+
+ is_prior_in_gt_shadow = (
+ self.iou_calculator(priors, gt_shadow, mode='iof') >
+ self.min_pos_iof)
+ # Rule out center effective positive pixels
+ is_prior_in_gt_shadow &= (~is_prior_in_gt_core)
+
+ num_gts, num_priors = gt_bboxes.size(0), priors.size(0)
+ if num_gts == 0 or num_priors == 0:
+ # If no gts exist, assign all pixels to negative
+ assigned_gt_ids = \
+ is_prior_in_gt_core.new_zeros((num_priors,),
+ dtype=torch.long)
+ pixels_in_gt_shadow = assigned_gt_ids.new_empty((0, 2))
+ else:
+ # Step 3: assign a one-hot gt id to each pixel, and smaller objects
+ # have high priority to assign the pixel.
+ sort_idx = self.get_gt_priorities(gt_bboxes)
+ assigned_gt_ids, pixels_in_gt_shadow = \
+ self.assign_one_hot_gt_indices(is_prior_in_gt_core,
+ is_prior_in_gt_shadow,
+ gt_priority=sort_idx)
+
+ if (gt_instances_ignore is not None
+ and gt_instances_ignore.bboxes.numel() > 0):
+ # No ground truth or boxes, return empty assignment
+ gt_bboxes_ignore = gt_instances_ignore.bboxes
+ gt_bboxes_ignore = scale_boxes(
+ gt_bboxes_ignore, scale=self.ignore_gt_scale)
+ is_prior_in_ignored_gts = is_located_in(prior_centers,
+ gt_bboxes_ignore)
+ is_prior_in_ignored_gts = is_prior_in_ignored_gts.any(dim=1)
+ assigned_gt_ids[is_prior_in_ignored_gts] = -1
+
+ # 4. Assign prior bboxes with class label according to its gt id.
+ # Default assigned label is the background (-1)
+ assigned_labels = assigned_gt_ids.new_full((num_priors, ), -1)
+ pos_inds = torch.nonzero(assigned_gt_ids > 0, as_tuple=False).squeeze()
+ if pos_inds.numel() > 0:
+ assigned_labels[pos_inds] = gt_labels[assigned_gt_ids[pos_inds] -
+ 1]
+ # 5. Find pixels lying in the shadow of an object
+ shadowed_pixel_labels = pixels_in_gt_shadow.clone()
+ if pixels_in_gt_shadow.numel() > 0:
+ pixel_idx, gt_idx =\
+ pixels_in_gt_shadow[:, 0], pixels_in_gt_shadow[:, 1]
+ assert (assigned_gt_ids[pixel_idx] != gt_idx).all(), \
+ 'Some pixels are dually assigned to ignore and gt!'
+ shadowed_pixel_labels[:, 1] = gt_labels[gt_idx - 1]
+ override = (
+ assigned_labels[pixel_idx] == shadowed_pixel_labels[:, 1])
+ if self.foreground_dominate:
+ # When a pixel is both positive and shadowed, set it as pos
+ shadowed_pixel_labels = shadowed_pixel_labels[~override]
+ else:
+ # When a pixel is both pos and shadowed, set it as shadowed
+ assigned_labels[pixel_idx[override]] = -1
+ assigned_gt_ids[pixel_idx[override]] = 0
+
+ assign_result = AssignResult(
+ num_gts, assigned_gt_ids, None, labels=assigned_labels)
+ # Add shadowed_labels as assign_result property. Shape: (num_shadow, 2)
+ assign_result.set_extra_property('shadowed_labels',
+ shadowed_pixel_labels)
+ return assign_result
+
+ def assign_one_hot_gt_indices(
+ self,
+ is_prior_in_gt_core: Tensor,
+ is_prior_in_gt_shadow: Tensor,
+ gt_priority: Optional[Tensor] = None) -> Tuple[Tensor, Tensor]:
+ """Assign only one gt index to each prior box.
+
+ Gts with large gt_priority are more likely to be assigned.
+
+ Args:
+ is_prior_in_gt_core (Tensor): Bool tensor indicating the prior
+ center is in the core area of a gt (e.g. 0-0.2).
+ Shape: (num_prior, num_gt).
+ is_prior_in_gt_shadow (Tensor): Bool tensor indicating the prior
+ center is in the shadowed area of a gt (e.g. 0.2-0.5).
+ Shape: (num_prior, num_gt).
+ gt_priority (Tensor): Priorities of gts. The gt with a higher
+ priority is more likely to be assigned to the bbox when the
+ bbox match with multiple gts. Shape: (num_gt, ).
+
+ Returns:
+ tuple: Returns (assigned_gt_inds, shadowed_gt_inds).
+
+ - assigned_gt_inds: The assigned gt index of each prior bbox \
+ (i.e. index from 1 to num_gts). Shape: (num_prior, ).
+ - shadowed_gt_inds: shadowed gt indices. It is a tensor of \
+ shape (num_ignore, 2) with first column being the shadowed prior \
+ bbox indices and the second column the shadowed gt \
+ indices (1-based).
+ """
+ num_bboxes, num_gts = is_prior_in_gt_core.shape
+
+ if gt_priority is None:
+ gt_priority = torch.arange(
+ num_gts, device=is_prior_in_gt_core.device)
+ assert gt_priority.size(0) == num_gts
+ # The bigger gt_priority, the more preferable to be assigned
+ # The assigned inds are by default 0 (background)
+ assigned_gt_inds = is_prior_in_gt_core.new_zeros((num_bboxes, ),
+ dtype=torch.long)
+ # Shadowed bboxes are assigned to be background. But the corresponding
+ # label is ignored during loss calculation, which is done through
+ # shadowed_gt_inds
+ shadowed_gt_inds = torch.nonzero(is_prior_in_gt_shadow, as_tuple=False)
+ if is_prior_in_gt_core.sum() == 0: # No gt match
+ shadowed_gt_inds[:, 1] += 1 # 1-based. For consistency issue
+ return assigned_gt_inds, shadowed_gt_inds
+
+ # The priority of each prior box and gt pair. If one prior box is
+ # matched bo multiple gts. Only the pair with the highest priority
+ # is saved
+ pair_priority = is_prior_in_gt_core.new_full((num_bboxes, num_gts),
+ -1,
+ dtype=torch.long)
+
+ # Each bbox could match with multiple gts.
+ # The following codes deal with this situation
+ # Matched bboxes (to any gt). Shape: (num_pos_anchor, )
+ inds_of_match = torch.any(is_prior_in_gt_core, dim=1)
+ # The matched gt index of each positive bbox. Length >= num_pos_anchor
+ # , since one bbox could match multiple gts
+ matched_bbox_gt_inds = torch.nonzero(
+ is_prior_in_gt_core, as_tuple=False)[:, 1]
+ # Assign priority to each bbox-gt pair.
+ pair_priority[is_prior_in_gt_core] = gt_priority[matched_bbox_gt_inds]
+ _, argmax_priority = pair_priority[inds_of_match].max(dim=1)
+ assigned_gt_inds[inds_of_match] = argmax_priority + 1 # 1-based
+ # Zero-out the assigned anchor box to filter the shadowed gt indices
+ is_prior_in_gt_core[inds_of_match, argmax_priority] = 0
+ # Concat the shadowed indices due to overlapping with that out side of
+ # effective scale. shape: (total_num_ignore, 2)
+ shadowed_gt_inds = torch.cat(
+ (shadowed_gt_inds,
+ torch.nonzero(is_prior_in_gt_core, as_tuple=False)),
+ dim=0)
+ # Change `is_prior_in_gt_core` back to keep arguments intact.
+ is_prior_in_gt_core[inds_of_match, argmax_priority] = 1
+ # 1-based shadowed gt indices, to be consistent with `assigned_gt_inds`
+ if shadowed_gt_inds.numel() > 0:
+ shadowed_gt_inds[:, 1] += 1
+ return assigned_gt_inds, shadowed_gt_inds
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/dynamic_soft_label_assigner.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/dynamic_soft_label_assigner.py
new file mode 100644
index 0000000000000000000000000000000000000000..3fc7af39b22cd6dc00248e330547176787c23963
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/dynamic_soft_label_assigner.py
@@ -0,0 +1,227 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Tuple
+
+import torch
+import torch.nn.functional as F
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import TASK_UTILS
+from mmdet.structures.bbox import BaseBoxes
+from mmdet.utils import ConfigType
+from .assign_result import AssignResult
+from .base_assigner import BaseAssigner
+
+INF = 100000000
+EPS = 1.0e-7
+
+
+def center_of_mass(masks: Tensor, eps: float = 1e-7) -> Tensor:
+ """Compute the masks center of mass.
+
+ Args:
+ masks: Mask tensor, has shape (num_masks, H, W).
+ eps: a small number to avoid normalizer to be zero.
+ Defaults to 1e-7.
+ Returns:
+ Tensor: The masks center of mass. Has shape (num_masks, 2).
+ """
+ n, h, w = masks.shape
+ grid_h = torch.arange(h, device=masks.device)[:, None]
+ grid_w = torch.arange(w, device=masks.device)
+ normalizer = masks.sum(dim=(1, 2)).float().clamp(min=eps)
+ center_y = (masks * grid_h).sum(dim=(1, 2)) / normalizer
+ center_x = (masks * grid_w).sum(dim=(1, 2)) / normalizer
+ center = torch.cat([center_x[:, None], center_y[:, None]], dim=1)
+ return center
+
+
+@TASK_UTILS.register_module()
+class DynamicSoftLabelAssigner(BaseAssigner):
+ """Computes matching between predictions and ground truth with dynamic soft
+ label assignment.
+
+ Args:
+ soft_center_radius (float): Radius of the soft center prior.
+ Defaults to 3.0.
+ topk (int): Select top-k predictions to calculate dynamic k
+ best matches for each gt. Defaults to 13.
+ iou_weight (float): The scale factor of iou cost. Defaults to 3.0.
+ iou_calculator (ConfigType): Config of overlaps Calculator.
+ Defaults to dict(type='BboxOverlaps2D').
+ """
+
+ def __init__(
+ self,
+ soft_center_radius: float = 3.0,
+ topk: int = 13,
+ iou_weight: float = 3.0,
+ iou_calculator: ConfigType = dict(type='BboxOverlaps2D')
+ ) -> None:
+ self.soft_center_radius = soft_center_radius
+ self.topk = topk
+ self.iou_weight = iou_weight
+ self.iou_calculator = TASK_UTILS.build(iou_calculator)
+
+ def assign(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ **kwargs) -> AssignResult:
+ """Assign gt to priors.
+
+ Args:
+ pred_instances (:obj:`InstanceData`): Instances of model
+ predictions. It includes ``priors``, and the priors can
+ be anchors or points, or the bboxes predicted by the
+ previous stage, has shape (n, 4). The bboxes predicted by
+ the current model or stage will be named ``bboxes``,
+ ``labels``, and ``scores``, the same as the ``InstanceData``
+ in other places.
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes``, with shape (k, 4),
+ and ``labels``, with shape (k, ).
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes``
+ attribute data that is ignored during training and testing.
+ Defaults to None.
+ Returns:
+ obj:`AssignResult`: The assigned result.
+ """
+ gt_bboxes = gt_instances.bboxes
+ gt_labels = gt_instances.labels
+ num_gt = gt_bboxes.size(0)
+
+ decoded_bboxes = pred_instances.bboxes
+ pred_scores = pred_instances.scores
+ priors = pred_instances.priors
+ num_bboxes = decoded_bboxes.size(0)
+
+ # assign 0 by default
+ assigned_gt_inds = decoded_bboxes.new_full((num_bboxes, ),
+ 0,
+ dtype=torch.long)
+ if num_gt == 0 or num_bboxes == 0:
+ # No ground truth or boxes, return empty assignment
+ max_overlaps = decoded_bboxes.new_zeros((num_bboxes, ))
+ if num_gt == 0:
+ # No truth, assign everything to background
+ assigned_gt_inds[:] = 0
+ assigned_labels = decoded_bboxes.new_full((num_bboxes, ),
+ -1,
+ dtype=torch.long)
+ return AssignResult(
+ num_gt, assigned_gt_inds, max_overlaps, labels=assigned_labels)
+
+ prior_center = priors[:, :2]
+ if isinstance(gt_bboxes, BaseBoxes):
+ is_in_gts = gt_bboxes.find_inside_points(prior_center)
+ else:
+ # Tensor boxes will be treated as horizontal boxes by defaults
+ lt_ = prior_center[:, None] - gt_bboxes[:, :2]
+ rb_ = gt_bboxes[:, 2:] - prior_center[:, None]
+
+ deltas = torch.cat([lt_, rb_], dim=-1)
+ is_in_gts = deltas.min(dim=-1).values > 0
+
+ valid_mask = is_in_gts.sum(dim=1) > 0
+
+ valid_decoded_bbox = decoded_bboxes[valid_mask]
+ valid_pred_scores = pred_scores[valid_mask]
+ num_valid = valid_decoded_bbox.size(0)
+
+ if num_valid == 0:
+ # No ground truth or boxes, return empty assignment
+ max_overlaps = decoded_bboxes.new_zeros((num_bboxes, ))
+ assigned_labels = decoded_bboxes.new_full((num_bboxes, ),
+ -1,
+ dtype=torch.long)
+ return AssignResult(
+ num_gt, assigned_gt_inds, max_overlaps, labels=assigned_labels)
+ if hasattr(gt_instances, 'masks'):
+ gt_center = center_of_mass(gt_instances.masks, eps=EPS)
+ elif isinstance(gt_bboxes, BaseBoxes):
+ gt_center = gt_bboxes.centers
+ else:
+ # Tensor boxes will be treated as horizontal boxes by defaults
+ gt_center = (gt_bboxes[:, :2] + gt_bboxes[:, 2:]) / 2.0
+ valid_prior = priors[valid_mask]
+ strides = valid_prior[:, 2]
+ distance = (valid_prior[:, None, :2] - gt_center[None, :, :]
+ ).pow(2).sum(-1).sqrt() / strides[:, None]
+ soft_center_prior = torch.pow(10, distance - self.soft_center_radius)
+
+ pairwise_ious = self.iou_calculator(valid_decoded_bbox, gt_bboxes)
+ iou_cost = -torch.log(pairwise_ious + EPS) * self.iou_weight
+
+ gt_onehot_label = (
+ F.one_hot(gt_labels.to(torch.int64),
+ pred_scores.shape[-1]).float().unsqueeze(0).repeat(
+ num_valid, 1, 1))
+ valid_pred_scores = valid_pred_scores.unsqueeze(1).repeat(1, num_gt, 1)
+
+ soft_label = gt_onehot_label * pairwise_ious[..., None]
+ scale_factor = soft_label - valid_pred_scores.sigmoid()
+ soft_cls_cost = F.binary_cross_entropy_with_logits(
+ valid_pred_scores, soft_label,
+ reduction='none') * scale_factor.abs().pow(2.0)
+ soft_cls_cost = soft_cls_cost.sum(dim=-1)
+
+ cost_matrix = soft_cls_cost + iou_cost + soft_center_prior
+
+ matched_pred_ious, matched_gt_inds = self.dynamic_k_matching(
+ cost_matrix, pairwise_ious, num_gt, valid_mask)
+
+ # convert to AssignResult format
+ assigned_gt_inds[valid_mask] = matched_gt_inds + 1
+ assigned_labels = assigned_gt_inds.new_full((num_bboxes, ), -1)
+ assigned_labels[valid_mask] = gt_labels[matched_gt_inds].long()
+ max_overlaps = assigned_gt_inds.new_full((num_bboxes, ),
+ -INF,
+ dtype=torch.float32)
+ max_overlaps[valid_mask] = matched_pred_ious
+ return AssignResult(
+ num_gt, assigned_gt_inds, max_overlaps, labels=assigned_labels)
+
+ def dynamic_k_matching(self, cost: Tensor, pairwise_ious: Tensor,
+ num_gt: int,
+ valid_mask: Tensor) -> Tuple[Tensor, Tensor]:
+ """Use IoU and matching cost to calculate the dynamic top-k positive
+ targets. Same as SimOTA.
+
+ Args:
+ cost (Tensor): Cost matrix.
+ pairwise_ious (Tensor): Pairwise iou matrix.
+ num_gt (int): Number of gt.
+ valid_mask (Tensor): Mask for valid bboxes.
+
+ Returns:
+ tuple: matched ious and gt indexes.
+ """
+ matching_matrix = torch.zeros_like(cost, dtype=torch.uint8)
+ # select candidate topk ious for dynamic-k calculation
+ candidate_topk = min(self.topk, pairwise_ious.size(0))
+ topk_ious, _ = torch.topk(pairwise_ious, candidate_topk, dim=0)
+ # calculate dynamic k for each gt
+ dynamic_ks = torch.clamp(topk_ious.sum(0).int(), min=1)
+ for gt_idx in range(num_gt):
+ _, pos_idx = torch.topk(
+ cost[:, gt_idx], k=dynamic_ks[gt_idx], largest=False)
+ matching_matrix[:, gt_idx][pos_idx] = 1
+
+ del topk_ious, dynamic_ks, pos_idx
+
+ prior_match_gt_mask = matching_matrix.sum(1) > 1
+ if prior_match_gt_mask.sum() > 0:
+ cost_min, cost_argmin = torch.min(
+ cost[prior_match_gt_mask, :], dim=1)
+ matching_matrix[prior_match_gt_mask, :] *= 0
+ matching_matrix[prior_match_gt_mask, cost_argmin] = 1
+ # get foreground mask inside box and center prior
+ fg_mask_inboxes = matching_matrix.sum(1) > 0
+ valid_mask[valid_mask.clone()] = fg_mask_inboxes
+
+ matched_gt_inds = matching_matrix[fg_mask_inboxes, :].argmax(1)
+ matched_pred_ious = (matching_matrix *
+ pairwise_ious).sum(1)[fg_mask_inboxes]
+ return matched_pred_ious, matched_gt_inds
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/grid_assigner.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/grid_assigner.py
new file mode 100644
index 0000000000000000000000000000000000000000..d8935d2df2937f90c71599e5b45ed9a3dff8cd7e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/grid_assigner.py
@@ -0,0 +1,177 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Tuple, Union
+
+import torch
+from mmengine.structures import InstanceData
+
+from mmdet.registry import TASK_UTILS
+from mmdet.utils import ConfigType
+from .assign_result import AssignResult
+from .base_assigner import BaseAssigner
+
+
+@TASK_UTILS.register_module()
+class GridAssigner(BaseAssigner):
+ """Assign a corresponding gt bbox or background to each bbox.
+
+ Each proposals will be assigned with `-1`, `0`, or a positive integer
+ indicating the ground truth index.
+
+ - -1: don't care
+ - 0: negative sample, no assigned gt
+ - positive integer: positive sample, index (1-based) of assigned gt
+
+ Args:
+ pos_iou_thr (float): IoU threshold for positive bboxes.
+ neg_iou_thr (float or tuple[float, float]): IoU threshold for negative
+ bboxes.
+ min_pos_iou (float): Minimum iou for a bbox to be considered as a
+ positive bbox. Positive samples can have smaller IoU than
+ pos_iou_thr due to the 4th step (assign max IoU sample to each gt).
+ Defaults to 0.
+ gt_max_assign_all (bool): Whether to assign all bboxes with the same
+ highest overlap with some gt to that gt.
+ iou_calculator (:obj:`ConfigDict` or dict): Config of overlaps
+ Calculator.
+ """
+
+ def __init__(
+ self,
+ pos_iou_thr: float,
+ neg_iou_thr: Union[float, Tuple[float, float]],
+ min_pos_iou: float = .0,
+ gt_max_assign_all: bool = True,
+ iou_calculator: ConfigType = dict(type='BboxOverlaps2D')
+ ) -> None:
+ self.pos_iou_thr = pos_iou_thr
+ self.neg_iou_thr = neg_iou_thr
+ self.min_pos_iou = min_pos_iou
+ self.gt_max_assign_all = gt_max_assign_all
+ self.iou_calculator = TASK_UTILS.build(iou_calculator)
+
+ def assign(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ **kwargs) -> AssignResult:
+ """Assign gt to bboxes. The process is very much like the max iou
+ assigner, except that positive samples are constrained within the cell
+ that the gt boxes fell in.
+
+ This method assign a gt bbox to every bbox (proposal/anchor), each bbox
+ will be assigned with -1, 0, or a positive number. -1 means don't care,
+ 0 means negative sample, positive number is the index (1-based) of
+ assigned gt.
+ The assignment is done in following steps, the order matters.
+
+ 1. assign every bbox to -1
+ 2. assign proposals whose iou with all gts <= neg_iou_thr to 0
+ 3. for each bbox within a cell, if the iou with its nearest gt >
+ pos_iou_thr and the center of that gt falls inside the cell,
+ assign it to that bbox
+ 4. for each gt bbox, assign its nearest proposals within the cell the
+ gt bbox falls in to itself.
+
+ Args:
+ pred_instances (:obj:`InstanceData`): Instances of model
+ predictions. It includes ``priors``, and the priors can
+ be anchors or points, or the bboxes predicted by the
+ previous stage, has shape (n, 4). The bboxes predicted by
+ the current model or stage will be named ``bboxes``,
+ ``labels``, and ``scores``, the same as the ``InstanceData``
+ in other places.
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes``, with shape (k, 4),
+ and ``labels``, with shape (k, ).
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes``
+ attribute data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ :obj:`AssignResult`: The assign result.
+ """
+ gt_bboxes = gt_instances.bboxes
+ gt_labels = gt_instances.labels
+
+ priors = pred_instances.priors
+ responsible_flags = pred_instances.responsible_flags
+
+ num_gts, num_priors = gt_bboxes.size(0), priors.size(0)
+
+ # compute iou between all gt and priors
+ overlaps = self.iou_calculator(gt_bboxes, priors)
+
+ # 1. assign -1 by default
+ assigned_gt_inds = overlaps.new_full((num_priors, ),
+ -1,
+ dtype=torch.long)
+
+ if num_gts == 0 or num_priors == 0:
+ # No ground truth or priors, return empty assignment
+ max_overlaps = overlaps.new_zeros((num_priors, ))
+ if num_gts == 0:
+ # No truth, assign everything to background
+ assigned_gt_inds[:] = 0
+ assigned_labels = overlaps.new_full((num_priors, ),
+ -1,
+ dtype=torch.long)
+ return AssignResult(
+ num_gts,
+ assigned_gt_inds,
+ max_overlaps,
+ labels=assigned_labels)
+
+ # 2. assign negative: below
+ # for each anchor, which gt best overlaps with it
+ # for each anchor, the max iou of all gts
+ # shape of max_overlaps == argmax_overlaps == num_priors
+ max_overlaps, argmax_overlaps = overlaps.max(dim=0)
+
+ if isinstance(self.neg_iou_thr, float):
+ assigned_gt_inds[(max_overlaps >= 0)
+ & (max_overlaps <= self.neg_iou_thr)] = 0
+ elif isinstance(self.neg_iou_thr, (tuple, list)):
+ assert len(self.neg_iou_thr) == 2
+ assigned_gt_inds[(max_overlaps > self.neg_iou_thr[0])
+ & (max_overlaps <= self.neg_iou_thr[1])] = 0
+
+ # 3. assign positive: falls into responsible cell and above
+ # positive IOU threshold, the order matters.
+ # the prior condition of comparison is to filter out all
+ # unrelated anchors, i.e. not responsible_flags
+ overlaps[:, ~responsible_flags.type(torch.bool)] = -1.
+
+ # calculate max_overlaps again, but this time we only consider IOUs
+ # for anchors responsible for prediction
+ max_overlaps, argmax_overlaps = overlaps.max(dim=0)
+
+ # for each gt, which anchor best overlaps with it
+ # for each gt, the max iou of all proposals
+ # shape of gt_max_overlaps == gt_argmax_overlaps == num_gts
+ gt_max_overlaps, gt_argmax_overlaps = overlaps.max(dim=1)
+
+ pos_inds = (max_overlaps > self.pos_iou_thr) & responsible_flags.type(
+ torch.bool)
+ assigned_gt_inds[pos_inds] = argmax_overlaps[pos_inds] + 1
+
+ # 4. assign positive to max overlapped anchors within responsible cell
+ for i in range(num_gts):
+ if gt_max_overlaps[i] > self.min_pos_iou:
+ if self.gt_max_assign_all:
+ max_iou_inds = (overlaps[i, :] == gt_max_overlaps[i]) & \
+ responsible_flags.type(torch.bool)
+ assigned_gt_inds[max_iou_inds] = i + 1
+ elif responsible_flags[gt_argmax_overlaps[i]]:
+ assigned_gt_inds[gt_argmax_overlaps[i]] = i + 1
+
+ # assign labels of positive anchors
+ assigned_labels = assigned_gt_inds.new_full((num_priors, ), -1)
+ pos_inds = torch.nonzero(
+ assigned_gt_inds > 0, as_tuple=False).squeeze()
+ if pos_inds.numel() > 0:
+ assigned_labels[pos_inds] = gt_labels[assigned_gt_inds[pos_inds] -
+ 1]
+
+ return AssignResult(
+ num_gts, assigned_gt_inds, max_overlaps, labels=assigned_labels)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/hungarian_assigner.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/hungarian_assigner.py
new file mode 100644
index 0000000000000000000000000000000000000000..a6745a36cdc713c74f801f62dae0d8fe3d03828f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/hungarian_assigner.py
@@ -0,0 +1,145 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Union
+
+import torch
+from mmengine import ConfigDict
+from mmengine.structures import InstanceData
+from scipy.optimize import linear_sum_assignment
+from torch import Tensor
+
+from mmdet.registry import TASK_UTILS
+from .assign_result import AssignResult
+from .base_assigner import BaseAssigner
+
+
+@TASK_UTILS.register_module()
+class HungarianAssigner(BaseAssigner):
+ """Computes one-to-one matching between predictions and ground truth.
+
+ This class computes an assignment between the targets and the predictions
+ based on the costs. The costs are weighted sum of some components.
+ For DETR the costs are weighted sum of classification cost, regression L1
+ cost and regression iou cost. The targets don't include the no_object, so
+ generally there are more predictions than targets. After the one-to-one
+ matching, the un-matched are treated as backgrounds. Thus each query
+ prediction will be assigned with `0` or a positive integer indicating the
+ ground truth index:
+
+ - 0: negative sample, no assigned gt
+ - positive integer: positive sample, index (1-based) of assigned gt
+
+ Args:
+ match_costs (:obj:`ConfigDict` or dict or \
+ List[Union[:obj:`ConfigDict`, dict]]): Match cost configs.
+ """
+
+ def __init__(
+ self, match_costs: Union[List[Union[dict, ConfigDict]], dict,
+ ConfigDict]
+ ) -> None:
+
+ if isinstance(match_costs, dict):
+ match_costs = [match_costs]
+ elif isinstance(match_costs, list):
+ assert len(match_costs) > 0, \
+ 'match_costs must not be a empty list.'
+
+ self.match_costs = [
+ TASK_UTILS.build(match_cost) for match_cost in match_costs
+ ]
+
+ def assign(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ img_meta: Optional[dict] = None,
+ **kwargs) -> AssignResult:
+ """Computes one-to-one matching based on the weighted costs.
+
+ This method assign each query prediction to a ground truth or
+ background. The `assigned_gt_inds` with -1 means don't care,
+ 0 means negative sample, and positive number is the index (1-based)
+ of assigned gt.
+ The assignment is done in the following steps, the order matters.
+
+ 1. assign every prediction to -1
+ 2. compute the weighted costs
+ 3. do Hungarian matching on CPU based on the costs
+ 4. assign all to 0 (background) first, then for each matched pair
+ between predictions and gts, treat this prediction as foreground
+ and assign the corresponding gt index (plus 1) to it.
+
+ Args:
+ pred_instances (:obj:`InstanceData`): Instances of model
+ predictions. It includes ``priors``, and the priors can
+ be anchors or points, or the bboxes predicted by the
+ previous stage, has shape (n, 4). The bboxes predicted by
+ the current model or stage will be named ``bboxes``,
+ ``labels``, and ``scores``, the same as the ``InstanceData``
+ in other places. It may includes ``masks``, with shape
+ (n, h, w) or (n, l).
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes``, with shape (k, 4),
+ ``labels``, with shape (k, ) and ``masks``, with shape
+ (k, h, w) or (k, l).
+ img_meta (dict): Image information.
+
+ Returns:
+ :obj:`AssignResult`: The assigned result.
+ """
+ assert isinstance(gt_instances.labels, Tensor)
+ num_gts, num_preds = len(gt_instances), len(pred_instances)
+ gt_labels = gt_instances.labels
+ device = gt_labels.device
+
+ # 1. assign -1 by default
+ assigned_gt_inds = torch.full((num_preds, ),
+ -1,
+ dtype=torch.long,
+ device=device)
+ assigned_labels = torch.full((num_preds, ),
+ -1,
+ dtype=torch.long,
+ device=device)
+
+ if num_gts == 0 or num_preds == 0:
+ # No ground truth or boxes, return empty assignment
+ if num_gts == 0:
+ # No ground truth, assign all to background
+ assigned_gt_inds[:] = 0
+ return AssignResult(
+ num_gts=num_gts,
+ gt_inds=assigned_gt_inds,
+ max_overlaps=None,
+ labels=assigned_labels)
+
+ # 2. compute weighted cost
+ cost_list = []
+ for match_cost in self.match_costs:
+ cost = match_cost(
+ pred_instances=pred_instances,
+ gt_instances=gt_instances,
+ img_meta=img_meta)
+ cost_list.append(cost)
+ cost = torch.stack(cost_list).sum(dim=0)
+
+ # 3. do Hungarian matching on CPU using linear_sum_assignment
+ cost = cost.detach().cpu()
+ if linear_sum_assignment is None:
+ raise ImportError('Please run "pip install scipy" '
+ 'to install scipy first.')
+
+ matched_row_inds, matched_col_inds = linear_sum_assignment(cost)
+ matched_row_inds = torch.from_numpy(matched_row_inds).to(device)
+ matched_col_inds = torch.from_numpy(matched_col_inds).to(device)
+
+ # 4. assign backgrounds and foregrounds
+ # assign all indices to backgrounds first
+ assigned_gt_inds[:] = 0
+ # assign foregrounds based on matching results
+ assigned_gt_inds[matched_row_inds] = matched_col_inds + 1
+ assigned_labels[matched_row_inds] = gt_labels[matched_col_inds]
+ return AssignResult(
+ num_gts=num_gts,
+ gt_inds=assigned_gt_inds,
+ max_overlaps=None,
+ labels=assigned_labels)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/iou2d_calculator.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/iou2d_calculator.py
new file mode 100644
index 0000000000000000000000000000000000000000..b6daa94feb46ac2f188df41c7be59ffdc3905e58
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/iou2d_calculator.py
@@ -0,0 +1,88 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+
+from mmdet.registry import TASK_UTILS
+from mmdet.structures.bbox import bbox_overlaps, get_box_tensor
+
+
+def cast_tensor_type(x, scale=1., dtype=None):
+ if dtype == 'fp16':
+ # scale is for preventing overflows
+ x = (x / scale).half()
+ return x
+
+
+@TASK_UTILS.register_module()
+class BboxOverlaps2D:
+ """2D Overlaps (e.g. IoUs, GIoUs) Calculator."""
+
+ def __init__(self, scale=1., dtype=None):
+ self.scale = scale
+ self.dtype = dtype
+
+ def __call__(self, bboxes1, bboxes2, mode='iou', is_aligned=False):
+ """Calculate IoU between 2D bboxes.
+
+ Args:
+ bboxes1 (Tensor or :obj:`BaseBoxes`): bboxes have shape (m, 4)
+ in format, or shape (m, 5) in format.
+ bboxes2 (Tensor or :obj:`BaseBoxes`): bboxes have shape (m, 4)
+ in format, shape (m, 5) in format, or be empty. If ``is_aligned `` is ``True``,
+ then m and n must be equal.
+ mode (str): "iou" (intersection over union), "iof" (intersection
+ over foreground), or "giou" (generalized intersection over
+ union).
+ is_aligned (bool, optional): If True, then m and n must be equal.
+ Default False.
+
+ Returns:
+ Tensor: shape (m, n) if ``is_aligned `` is False else shape (m,)
+ """
+ bboxes1 = get_box_tensor(bboxes1)
+ bboxes2 = get_box_tensor(bboxes2)
+ assert bboxes1.size(-1) in [0, 4, 5]
+ assert bboxes2.size(-1) in [0, 4, 5]
+ if bboxes2.size(-1) == 5:
+ bboxes2 = bboxes2[..., :4]
+ if bboxes1.size(-1) == 5:
+ bboxes1 = bboxes1[..., :4]
+
+ if self.dtype == 'fp16':
+ # change tensor type to save cpu and cuda memory and keep speed
+ bboxes1 = cast_tensor_type(bboxes1, self.scale, self.dtype)
+ bboxes2 = cast_tensor_type(bboxes2, self.scale, self.dtype)
+ overlaps = bbox_overlaps(bboxes1, bboxes2, mode, is_aligned)
+ if not overlaps.is_cuda and overlaps.dtype == torch.float16:
+ # resume cpu float32
+ overlaps = overlaps.float()
+ return overlaps
+
+ return bbox_overlaps(bboxes1, bboxes2, mode, is_aligned)
+
+ def __repr__(self):
+ """str: a string describing the module"""
+ repr_str = self.__class__.__name__ + f'(' \
+ f'scale={self.scale}, dtype={self.dtype})'
+ return repr_str
+
+
+@TASK_UTILS.register_module()
+class BboxOverlaps2D_GLIP(BboxOverlaps2D):
+
+ def __call__(self, bboxes1, bboxes2, mode='iou', is_aligned=False):
+ TO_REMOVE = 1
+ area1 = (bboxes1[:, 2] - bboxes1[:, 0] + TO_REMOVE) * (
+ bboxes1[:, 3] - bboxes1[:, 1] + TO_REMOVE)
+ area2 = (bboxes2[:, 2] - bboxes2[:, 0] + TO_REMOVE) * (
+ bboxes2[:, 3] - bboxes2[:, 1] + TO_REMOVE)
+
+ lt = torch.max(bboxes1[:, None, :2], bboxes2[:, :2]) # [N,M,2]
+ rb = torch.min(bboxes1[:, None, 2:], bboxes2[:, 2:]) # [N,M,2]
+
+ wh = (rb - lt + TO_REMOVE).clamp(min=0) # [N,M,2]
+ inter = wh[:, :, 0] * wh[:, :, 1] # [N,M]
+
+ iou = inter / (area1[:, None] + area2 - inter)
+ return iou
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/match_cost.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/match_cost.py
new file mode 100644
index 0000000000000000000000000000000000000000..5fc62f01f29138cba31ef2b41254f497351fe0d0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/match_cost.py
@@ -0,0 +1,525 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from abc import abstractmethod
+from typing import Optional, Union
+
+import torch
+import torch.nn.functional as F
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import TASK_UTILS
+from mmdet.structures.bbox import bbox_overlaps, bbox_xyxy_to_cxcywh
+
+
+class BaseMatchCost:
+ """Base match cost class.
+
+ Args:
+ weight (Union[float, int]): Cost weight. Defaults to 1.
+ """
+
+ def __init__(self, weight: Union[float, int] = 1.) -> None:
+ self.weight = weight
+
+ @abstractmethod
+ def __call__(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ img_meta: Optional[dict] = None,
+ **kwargs) -> Tensor:
+ """Compute match cost.
+
+ Args:
+ pred_instances (:obj:`InstanceData`): Instances of model
+ predictions. It includes ``priors``, and the priors can
+ be anchors or points, or the bboxes predicted by the
+ previous stage, has shape (n, 4). The bboxes predicted by
+ the current model or stage will be named ``bboxes``,
+ ``labels``, and ``scores``, the same as the ``InstanceData``
+ in other places.
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes``, with shape (k, 4),
+ and ``labels``, with shape (k, ).
+ img_meta (dict, optional): Image information.
+
+ Returns:
+ Tensor: Match Cost matrix of shape (num_preds, num_gts).
+ """
+ pass
+
+
+@TASK_UTILS.register_module()
+class BBoxL1Cost(BaseMatchCost):
+ """BBoxL1Cost.
+
+ Note: ``bboxes`` in ``InstanceData`` passed in is of format 'xyxy'
+ and its coordinates are unnormalized.
+
+ Args:
+ box_format (str, optional): 'xyxy' for DETR, 'xywh' for Sparse_RCNN.
+ Defaults to 'xyxy'.
+ weight (Union[float, int]): Cost weight. Defaults to 1.
+
+ Examples:
+ >>> from mmdet.models.task_modules.assigners.
+ ... match_costs.match_cost import BBoxL1Cost
+ >>> import torch
+ >>> self = BBoxL1Cost()
+ >>> bbox_pred = torch.rand(1, 4)
+ >>> gt_bboxes= torch.FloatTensor([[0, 0, 2, 4], [1, 2, 3, 4]])
+ >>> factor = torch.tensor([10, 8, 10, 8])
+ >>> self(bbox_pred, gt_bboxes, factor)
+ tensor([[1.6172, 1.6422]])
+ """
+
+ def __init__(self,
+ box_format: str = 'xyxy',
+ weight: Union[float, int] = 1.) -> None:
+ super().__init__(weight=weight)
+ assert box_format in ['xyxy', 'xywh']
+ self.box_format = box_format
+
+ def __call__(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ img_meta: Optional[dict] = None,
+ **kwargs) -> Tensor:
+ """Compute match cost.
+
+ Args:
+ pred_instances (:obj:`InstanceData`): ``bboxes`` inside is
+ predicted boxes with unnormalized coordinate
+ (x, y, x, y).
+ gt_instances (:obj:`InstanceData`): ``bboxes`` inside is gt
+ bboxes with unnormalized coordinate (x, y, x, y).
+ img_meta (Optional[dict]): Image information. Defaults to None.
+
+ Returns:
+ Tensor: Match Cost matrix of shape (num_preds, num_gts).
+ """
+ pred_bboxes = pred_instances.bboxes
+ gt_bboxes = gt_instances.bboxes
+
+ # convert box format
+ if self.box_format == 'xywh':
+ gt_bboxes = bbox_xyxy_to_cxcywh(gt_bboxes)
+ pred_bboxes = bbox_xyxy_to_cxcywh(pred_bboxes)
+
+ # normalized
+ img_h, img_w = img_meta['img_shape']
+ factor = gt_bboxes.new_tensor([img_w, img_h, img_w,
+ img_h]).unsqueeze(0)
+ gt_bboxes = gt_bboxes / factor
+ pred_bboxes = pred_bboxes / factor
+
+ bbox_cost = torch.cdist(pred_bboxes, gt_bboxes, p=1)
+ return bbox_cost * self.weight
+
+
+@TASK_UTILS.register_module()
+class IoUCost(BaseMatchCost):
+ """IoUCost.
+
+ Note: ``bboxes`` in ``InstanceData`` passed in is of format 'xyxy'
+ and its coordinates are unnormalized.
+
+ Args:
+ iou_mode (str): iou mode such as 'iou', 'giou'. Defaults to 'giou'.
+ weight (Union[float, int]): Cost weight. Defaults to 1.
+
+ Examples:
+ >>> from mmdet.models.task_modules.assigners.
+ ... match_costs.match_cost import IoUCost
+ >>> import torch
+ >>> self = IoUCost()
+ >>> bboxes = torch.FloatTensor([[1,1, 2, 2], [2, 2, 3, 4]])
+ >>> gt_bboxes = torch.FloatTensor([[0, 0, 2, 4], [1, 2, 3, 4]])
+ >>> self(bboxes, gt_bboxes)
+ tensor([[-0.1250, 0.1667],
+ [ 0.1667, -0.5000]])
+ """
+
+ def __init__(self, iou_mode: str = 'giou', weight: Union[float, int] = 1.):
+ super().__init__(weight=weight)
+ self.iou_mode = iou_mode
+
+ def __call__(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ img_meta: Optional[dict] = None,
+ **kwargs):
+ """Compute match cost.
+
+ Args:
+ pred_instances (:obj:`InstanceData`): ``bboxes`` inside is
+ predicted boxes with unnormalized coordinate
+ (x, y, x, y).
+ gt_instances (:obj:`InstanceData`): ``bboxes`` inside is gt
+ bboxes with unnormalized coordinate (x, y, x, y).
+ img_meta (Optional[dict]): Image information. Defaults to None.
+
+ Returns:
+ Tensor: Match Cost matrix of shape (num_preds, num_gts).
+ """
+ pred_bboxes = pred_instances.bboxes
+ gt_bboxes = gt_instances.bboxes
+
+ # avoid fp16 overflow
+ if pred_bboxes.dtype == torch.float16:
+ fp16 = True
+ pred_bboxes = pred_bboxes.to(torch.float32)
+ else:
+ fp16 = False
+
+ overlaps = bbox_overlaps(
+ pred_bboxes, gt_bboxes, mode=self.iou_mode, is_aligned=False)
+
+ if fp16:
+ overlaps = overlaps.to(torch.float16)
+
+ # The 1 is a constant that doesn't change the matching, so omitted.
+ iou_cost = -overlaps
+ return iou_cost * self.weight
+
+
+@TASK_UTILS.register_module()
+class ClassificationCost(BaseMatchCost):
+ """ClsSoftmaxCost.
+
+ Args:
+ weight (Union[float, int]): Cost weight. Defaults to 1.
+
+ Examples:
+ >>> from mmdet.models.task_modules.assigners.
+ ... match_costs.match_cost import ClassificationCost
+ >>> import torch
+ >>> self = ClassificationCost()
+ >>> cls_pred = torch.rand(4, 3)
+ >>> gt_labels = torch.tensor([0, 1, 2])
+ >>> factor = torch.tensor([10, 8, 10, 8])
+ >>> self(cls_pred, gt_labels)
+ tensor([[-0.3430, -0.3525, -0.3045],
+ [-0.3077, -0.2931, -0.3992],
+ [-0.3664, -0.3455, -0.2881],
+ [-0.3343, -0.2701, -0.3956]])
+ """
+
+ def __init__(self, weight: Union[float, int] = 1) -> None:
+ super().__init__(weight=weight)
+
+ def __call__(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ img_meta: Optional[dict] = None,
+ **kwargs) -> Tensor:
+ """Compute match cost.
+
+ Args:
+ pred_instances (:obj:`InstanceData`): ``scores`` inside is
+ predicted classification logits, of shape
+ (num_queries, num_class).
+ gt_instances (:obj:`InstanceData`): ``labels`` inside should have
+ shape (num_gt, ).
+ img_meta (Optional[dict]): _description_. Defaults to None.
+
+ Returns:
+ Tensor: Match Cost matrix of shape (num_preds, num_gts).
+ """
+ pred_scores = pred_instances.scores
+ gt_labels = gt_instances.labels
+
+ pred_scores = pred_scores.softmax(-1)
+ cls_cost = -pred_scores[:, gt_labels]
+
+ return cls_cost * self.weight
+
+
+@TASK_UTILS.register_module()
+class FocalLossCost(BaseMatchCost):
+ """FocalLossCost.
+
+ Args:
+ alpha (Union[float, int]): focal_loss alpha. Defaults to 0.25.
+ gamma (Union[float, int]): focal_loss gamma. Defaults to 2.
+ eps (float): Defaults to 1e-12.
+ binary_input (bool): Whether the input is binary. Currently,
+ binary_input = True is for masks input, binary_input = False
+ is for label input. Defaults to False.
+ weight (Union[float, int]): Cost weight. Defaults to 1.
+ """
+
+ def __init__(self,
+ alpha: Union[float, int] = 0.25,
+ gamma: Union[float, int] = 2,
+ eps: float = 1e-12,
+ binary_input: bool = False,
+ weight: Union[float, int] = 1.) -> None:
+ super().__init__(weight=weight)
+ self.alpha = alpha
+ self.gamma = gamma
+ self.eps = eps
+ self.binary_input = binary_input
+
+ def _focal_loss_cost(self, cls_pred: Tensor, gt_labels: Tensor) -> Tensor:
+ """
+ Args:
+ cls_pred (Tensor): Predicted classification logits, shape
+ (num_queries, num_class).
+ gt_labels (Tensor): Label of `gt_bboxes`, shape (num_gt,).
+
+ Returns:
+ torch.Tensor: cls_cost value with weight
+ """
+ cls_pred = cls_pred.sigmoid()
+ neg_cost = -(1 - cls_pred + self.eps).log() * (
+ 1 - self.alpha) * cls_pred.pow(self.gamma)
+ pos_cost = -(cls_pred + self.eps).log() * self.alpha * (
+ 1 - cls_pred).pow(self.gamma)
+
+ cls_cost = pos_cost[:, gt_labels] - neg_cost[:, gt_labels]
+ return cls_cost * self.weight
+
+ def _mask_focal_loss_cost(self, cls_pred, gt_labels) -> Tensor:
+ """
+ Args:
+ cls_pred (Tensor): Predicted classification logits.
+ in shape (num_queries, d1, ..., dn), dtype=torch.float32.
+ gt_labels (Tensor): Ground truth in shape (num_gt, d1, ..., dn),
+ dtype=torch.long. Labels should be binary.
+
+ Returns:
+ Tensor: Focal cost matrix with weight in shape\
+ (num_queries, num_gt).
+ """
+ cls_pred = cls_pred.flatten(1)
+ gt_labels = gt_labels.flatten(1).float()
+ n = cls_pred.shape[1]
+ cls_pred = cls_pred.sigmoid()
+ neg_cost = -(1 - cls_pred + self.eps).log() * (
+ 1 - self.alpha) * cls_pred.pow(self.gamma)
+ pos_cost = -(cls_pred + self.eps).log() * self.alpha * (
+ 1 - cls_pred).pow(self.gamma)
+
+ cls_cost = torch.einsum('nc,mc->nm', pos_cost, gt_labels) + \
+ torch.einsum('nc,mc->nm', neg_cost, (1 - gt_labels))
+ return cls_cost / n * self.weight
+
+ def __call__(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ img_meta: Optional[dict] = None,
+ **kwargs) -> Tensor:
+ """Compute match cost.
+
+ Args:
+ pred_instances (:obj:`InstanceData`): Predicted instances which
+ must contain ``scores`` or ``masks``.
+ gt_instances (:obj:`InstanceData`): Ground truth which must contain
+ ``labels`` or ``mask``.
+ img_meta (Optional[dict]): Image information. Defaults to None.
+
+ Returns:
+ Tensor: Match Cost matrix of shape (num_preds, num_gts).
+ """
+ if self.binary_input:
+ pred_masks = pred_instances.masks
+ gt_masks = gt_instances.masks
+ return self._mask_focal_loss_cost(pred_masks, gt_masks)
+ else:
+ pred_scores = pred_instances.scores
+ gt_labels = gt_instances.labels
+ return self._focal_loss_cost(pred_scores, gt_labels)
+
+
+@TASK_UTILS.register_module()
+class BinaryFocalLossCost(FocalLossCost):
+
+ def _focal_loss_cost(self, cls_pred: Tensor, gt_labels: Tensor) -> Tensor:
+ """
+ Args:
+ cls_pred (Tensor): Predicted classification logits, shape
+ (num_queries, num_class).
+ gt_labels (Tensor): Label of `gt_bboxes`, shape (num_gt,).
+
+ Returns:
+ torch.Tensor: cls_cost value with weight
+ """
+ cls_pred = cls_pred.flatten(1)
+ gt_labels = gt_labels.flatten(1).float()
+ cls_pred = cls_pred.sigmoid()
+ neg_cost = -(1 - cls_pred + self.eps).log() * (
+ 1 - self.alpha) * cls_pred.pow(self.gamma)
+ pos_cost = -(cls_pred + self.eps).log() * self.alpha * (
+ 1 - cls_pred).pow(self.gamma)
+
+ cls_cost = torch.einsum('nc,mc->nm', pos_cost, gt_labels) + \
+ torch.einsum('nc,mc->nm', neg_cost, (1 - gt_labels))
+ return cls_cost * self.weight
+
+ def __call__(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ img_meta: Optional[dict] = None,
+ **kwargs) -> Tensor:
+ """Compute match cost.
+
+ Args:
+ pred_instances (:obj:`InstanceData`): Predicted instances which
+ must contain ``scores`` or ``masks``.
+ gt_instances (:obj:`InstanceData`): Ground truth which must contain
+ ``labels`` or ``mask``.
+ img_meta (Optional[dict]): Image information. Defaults to None.
+
+ Returns:
+ Tensor: Match Cost matrix of shape (num_preds, num_gts).
+ """
+ # gt_instances.text_token_mask is a repeated tensor of the same length
+ # of instances. Only gt_instances.text_token_mask[0] is useful
+ text_token_mask = torch.nonzero(
+ gt_instances.text_token_mask[0]).squeeze(-1)
+ pred_scores = pred_instances.scores[:, text_token_mask]
+ gt_labels = gt_instances.positive_maps[:, text_token_mask]
+ return self._focal_loss_cost(pred_scores, gt_labels)
+
+
+@TASK_UTILS.register_module()
+class DiceCost(BaseMatchCost):
+ """Cost of mask assignments based on dice losses.
+
+ Args:
+ pred_act (bool): Whether to apply sigmoid to mask_pred.
+ Defaults to False.
+ eps (float): Defaults to 1e-3.
+ naive_dice (bool): If True, use the naive dice loss
+ in which the power of the number in the denominator is
+ the first power. If False, use the second power that
+ is adopted by K-Net and SOLO. Defaults to True.
+ weight (Union[float, int]): Cost weight. Defaults to 1.
+ """
+
+ def __init__(self,
+ pred_act: bool = False,
+ eps: float = 1e-3,
+ naive_dice: bool = True,
+ weight: Union[float, int] = 1.) -> None:
+ super().__init__(weight=weight)
+ self.pred_act = pred_act
+ self.eps = eps
+ self.naive_dice = naive_dice
+
+ def _binary_mask_dice_loss(self, mask_preds: Tensor,
+ gt_masks: Tensor) -> Tensor:
+ """
+ Args:
+ mask_preds (Tensor): Mask prediction in shape (num_queries, *).
+ gt_masks (Tensor): Ground truth in shape (num_gt, *)
+ store 0 or 1, 0 for negative class and 1 for
+ positive class.
+
+ Returns:
+ Tensor: Dice cost matrix in shape (num_queries, num_gt).
+ """
+ mask_preds = mask_preds.flatten(1)
+ gt_masks = gt_masks.flatten(1).float()
+ numerator = 2 * torch.einsum('nc,mc->nm', mask_preds, gt_masks)
+ if self.naive_dice:
+ denominator = mask_preds.sum(-1)[:, None] + \
+ gt_masks.sum(-1)[None, :]
+ else:
+ denominator = mask_preds.pow(2).sum(1)[:, None] + \
+ gt_masks.pow(2).sum(1)[None, :]
+ loss = 1 - (numerator + self.eps) / (denominator + self.eps)
+ return loss
+
+ def __call__(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ img_meta: Optional[dict] = None,
+ **kwargs) -> Tensor:
+ """Compute match cost.
+
+ Args:
+ pred_instances (:obj:`InstanceData`): Predicted instances which
+ must contain ``masks``.
+ gt_instances (:obj:`InstanceData`): Ground truth which must contain
+ ``mask``.
+ img_meta (Optional[dict]): Image information. Defaults to None.
+
+ Returns:
+ Tensor: Match Cost matrix of shape (num_preds, num_gts).
+ """
+ pred_masks = pred_instances.masks
+ gt_masks = gt_instances.masks
+
+ if self.pred_act:
+ pred_masks = pred_masks.sigmoid()
+ dice_cost = self._binary_mask_dice_loss(pred_masks, gt_masks)
+ return dice_cost * self.weight
+
+
+@TASK_UTILS.register_module()
+class CrossEntropyLossCost(BaseMatchCost):
+ """CrossEntropyLossCost.
+
+ Args:
+ use_sigmoid (bool): Whether the prediction uses sigmoid
+ of softmax. Defaults to True.
+ weight (Union[float, int]): Cost weight. Defaults to 1.
+ """
+
+ def __init__(self,
+ use_sigmoid: bool = True,
+ weight: Union[float, int] = 1.) -> None:
+ super().__init__(weight=weight)
+ self.use_sigmoid = use_sigmoid
+
+ def _binary_cross_entropy(self, cls_pred: Tensor,
+ gt_labels: Tensor) -> Tensor:
+ """
+ Args:
+ cls_pred (Tensor): The prediction with shape (num_queries, 1, *) or
+ (num_queries, *).
+ gt_labels (Tensor): The learning label of prediction with
+ shape (num_gt, *).
+
+ Returns:
+ Tensor: Cross entropy cost matrix in shape (num_queries, num_gt).
+ """
+ cls_pred = cls_pred.flatten(1).float()
+ gt_labels = gt_labels.flatten(1).float()
+ n = cls_pred.shape[1]
+ pos = F.binary_cross_entropy_with_logits(
+ cls_pred, torch.ones_like(cls_pred), reduction='none')
+ neg = F.binary_cross_entropy_with_logits(
+ cls_pred, torch.zeros_like(cls_pred), reduction='none')
+ cls_cost = torch.einsum('nc,mc->nm', pos, gt_labels) + \
+ torch.einsum('nc,mc->nm', neg, 1 - gt_labels)
+ cls_cost = cls_cost / n
+
+ return cls_cost
+
+ def __call__(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ img_meta: Optional[dict] = None,
+ **kwargs) -> Tensor:
+ """Compute match cost.
+
+ Args:
+ pred_instances (:obj:`InstanceData`): Predicted instances which
+ must contain ``scores`` or ``masks``.
+ gt_instances (:obj:`InstanceData`): Ground truth which must contain
+ ``labels`` or ``masks``.
+ img_meta (Optional[dict]): Image information. Defaults to None.
+
+ Returns:
+ Tensor: Match Cost matrix of shape (num_preds, num_gts).
+ """
+ pred_masks = pred_instances.masks
+ gt_masks = gt_instances.masks
+ if self.use_sigmoid:
+ cls_cost = self._binary_cross_entropy(pred_masks, gt_masks)
+ else:
+ raise NotImplementedError
+
+ return cls_cost * self.weight
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/max_iou_assigner.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/max_iou_assigner.py
new file mode 100644
index 0000000000000000000000000000000000000000..71da54429ae0526bf52277bc3b1d24630acceaed
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/max_iou_assigner.py
@@ -0,0 +1,325 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from typing import Optional, Union
+
+import torch
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import TASK_UTILS
+from .assign_result import AssignResult
+from .base_assigner import BaseAssigner
+
+
+def _perm_box(bboxes,
+ iou_calculator,
+ iou_thr=0.97,
+ perm_range=0.01,
+ counter=0,
+ max_iter=5):
+ """Compute the permuted bboxes.
+
+ Args:
+ bboxes (Tensor): Shape (n, 4) for , "xyxy" format.
+ iou_calculator (obj): Overlaps Calculator.
+ iou_thr (float): The permuted bboxes should have IoU > iou_thr.
+ perm_range (float): The scale of permutation.
+ counter (int): Counter of permutation iteration.
+ max_iter (int): The max iterations of permutation.
+ Returns:
+ Tensor: The permuted bboxes.
+ """
+ ori_bboxes = copy.deepcopy(bboxes)
+ is_valid = True
+ N = bboxes.size(0)
+ perm_factor = bboxes.new_empty(N, 4).uniform_(1 - perm_range,
+ 1 + perm_range)
+ bboxes *= perm_factor
+ new_wh = bboxes[:, 2:] - bboxes[:, :2]
+ if (new_wh <= 0).any():
+ is_valid = False
+ iou = iou_calculator(ori_bboxes.unique(dim=0), bboxes)
+ if (iou < iou_thr).any():
+ is_valid = False
+ if not is_valid and counter < max_iter:
+ return _perm_box(
+ ori_bboxes,
+ iou_calculator,
+ perm_range=max(perm_range - counter * 0.001, 1e-3),
+ counter=counter + 1)
+ return bboxes
+
+
+def perm_repeat_bboxes(bboxes, iou_calculator=None, perm_repeat_cfg=None):
+ """Permute the repeated bboxes.
+
+ Args:
+ bboxes (Tensor): Shape (n, 4) for , "xyxy" format.
+ iou_calculator (obj): Overlaps Calculator.
+ perm_repeat_cfg (Dict): Config of permutation.
+ Returns:
+ Tensor: Bboxes after permuted repeated bboxes.
+ """
+ assert isinstance(bboxes, torch.Tensor)
+ if iou_calculator is None:
+ import torchvision
+ iou_calculator = torchvision.ops.box_iou
+ bboxes = copy.deepcopy(bboxes)
+ unique_bboxes = bboxes.unique(dim=0)
+ iou_thr = perm_repeat_cfg.get('iou_thr', 0.97)
+ perm_range = perm_repeat_cfg.get('perm_range', 0.01)
+ for box in unique_bboxes:
+ inds = (bboxes == box).sum(-1).float() == 4
+ if inds.float().sum().item() == 1:
+ continue
+ bboxes[inds] = _perm_box(
+ bboxes[inds],
+ iou_calculator,
+ iou_thr=iou_thr,
+ perm_range=perm_range,
+ counter=0)
+ return bboxes
+
+
+@TASK_UTILS.register_module()
+class MaxIoUAssigner(BaseAssigner):
+ """Assign a corresponding gt bbox or background to each bbox.
+
+ Each proposals will be assigned with `-1`, or a semi-positive integer
+ indicating the ground truth index.
+
+ - -1: negative sample, no assigned gt
+ - semi-positive integer: positive sample, index (0-based) of assigned gt
+
+ Args:
+ pos_iou_thr (float): IoU threshold for positive bboxes.
+ neg_iou_thr (float or tuple): IoU threshold for negative bboxes.
+ min_pos_iou (float): Minimum iou for a bbox to be considered as a
+ positive bbox. Positive samples can have smaller IoU than
+ pos_iou_thr due to the 4th step (assign max IoU sample to each gt).
+ `min_pos_iou` is set to avoid assigning bboxes that have extremely
+ small iou with GT as positive samples. It brings about 0.3 mAP
+ improvements in 1x schedule but does not affect the performance of
+ 3x schedule. More comparisons can be found in
+ `PR #7464 `_.
+ gt_max_assign_all (bool): Whether to assign all bboxes with the same
+ highest overlap with some gt to that gt.
+ ignore_iof_thr (float): IoF threshold for ignoring bboxes (if
+ `gt_bboxes_ignore` is specified). Negative values mean not
+ ignoring any bboxes.
+ ignore_wrt_candidates (bool): Whether to compute the iof between
+ `bboxes` and `gt_bboxes_ignore`, or the contrary.
+ match_low_quality (bool): Whether to allow low quality matches. This is
+ usually allowed for RPN and single stage detectors, but not allowed
+ in the second stage. Details are demonstrated in Step 4.
+ gpu_assign_thr (int): The upper bound of the number of GT for GPU
+ assign. When the number of gt is above this threshold, will assign
+ on CPU device. Negative values mean not assign on CPU.
+ iou_calculator (dict): Config of overlaps Calculator.
+ perm_repeat_gt_cfg (dict): Config of permute repeated gt bboxes.
+ """
+
+ def __init__(self,
+ pos_iou_thr: float,
+ neg_iou_thr: Union[float, tuple],
+ min_pos_iou: float = .0,
+ gt_max_assign_all: bool = True,
+ ignore_iof_thr: float = -1,
+ ignore_wrt_candidates: bool = True,
+ match_low_quality: bool = True,
+ gpu_assign_thr: float = -1,
+ iou_calculator: dict = dict(type='BboxOverlaps2D'),
+ perm_repeat_gt_cfg=None):
+ self.pos_iou_thr = pos_iou_thr
+ self.neg_iou_thr = neg_iou_thr
+ self.min_pos_iou = min_pos_iou
+ self.gt_max_assign_all = gt_max_assign_all
+ self.ignore_iof_thr = ignore_iof_thr
+ self.ignore_wrt_candidates = ignore_wrt_candidates
+ self.gpu_assign_thr = gpu_assign_thr
+ self.match_low_quality = match_low_quality
+ self.iou_calculator = TASK_UTILS.build(iou_calculator)
+ self.perm_repeat_gt_cfg = perm_repeat_gt_cfg
+
+ def assign(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ **kwargs) -> AssignResult:
+ """Assign gt to bboxes.
+
+ This method assign a gt bbox to every bbox (proposal/anchor), each bbox
+ will be assigned with -1, or a semi-positive number. -1 means negative
+ sample, semi-positive number is the index (0-based) of assigned gt.
+ The assignment is done in following steps, the order matters.
+
+ 1. assign every bbox to the background
+ 2. assign proposals whose iou with all gts < neg_iou_thr to 0
+ 3. for each bbox, if the iou with its nearest gt >= pos_iou_thr,
+ assign it to that bbox
+ 4. for each gt bbox, assign its nearest proposals (may be more than
+ one) to itself
+
+ Args:
+ pred_instances (:obj:`InstanceData`): Instances of model
+ predictions. It includes ``priors``, and the priors can
+ be anchors or points, or the bboxes predicted by the
+ previous stage, has shape (n, 4). The bboxes predicted by
+ the current model or stage will be named ``bboxes``,
+ ``labels``, and ``scores``, the same as the ``InstanceData``
+ in other places.
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes``, with shape (k, 4),
+ and ``labels``, with shape (k, ).
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes``
+ attribute data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ :obj:`AssignResult`: The assign result.
+
+ Example:
+ >>> from mmengine.structures import InstanceData
+ >>> self = MaxIoUAssigner(0.5, 0.5)
+ >>> pred_instances = InstanceData()
+ >>> pred_instances.priors = torch.Tensor([[0, 0, 10, 10],
+ ... [10, 10, 20, 20]])
+ >>> gt_instances = InstanceData()
+ >>> gt_instances.bboxes = torch.Tensor([[0, 0, 10, 9]])
+ >>> gt_instances.labels = torch.Tensor([0])
+ >>> assign_result = self.assign(pred_instances, gt_instances)
+ >>> expected_gt_inds = torch.LongTensor([1, 0])
+ >>> assert torch.all(assign_result.gt_inds == expected_gt_inds)
+ """
+ gt_bboxes = gt_instances.bboxes
+ priors = pred_instances.priors
+ gt_labels = gt_instances.labels
+ if gt_instances_ignore is not None:
+ gt_bboxes_ignore = gt_instances_ignore.bboxes
+ else:
+ gt_bboxes_ignore = None
+
+ assign_on_cpu = True if (self.gpu_assign_thr > 0) and (
+ gt_bboxes.shape[0] > self.gpu_assign_thr) else False
+ # compute overlap and assign gt on CPU when number of GT is large
+ if assign_on_cpu:
+ device = priors.device
+ priors = priors.cpu()
+ gt_bboxes = gt_bboxes.cpu()
+ gt_labels = gt_labels.cpu()
+ if gt_bboxes_ignore is not None:
+ gt_bboxes_ignore = gt_bboxes_ignore.cpu()
+
+ if self.perm_repeat_gt_cfg is not None and priors.numel() > 0:
+ gt_bboxes_unique = perm_repeat_bboxes(gt_bboxes,
+ self.iou_calculator,
+ self.perm_repeat_gt_cfg)
+ else:
+ gt_bboxes_unique = gt_bboxes
+ overlaps = self.iou_calculator(gt_bboxes_unique, priors)
+
+ if (self.ignore_iof_thr > 0 and gt_bboxes_ignore is not None
+ and gt_bboxes_ignore.numel() > 0 and priors.numel() > 0):
+ if self.ignore_wrt_candidates:
+ ignore_overlaps = self.iou_calculator(
+ priors, gt_bboxes_ignore, mode='iof')
+ ignore_max_overlaps, _ = ignore_overlaps.max(dim=1)
+ else:
+ ignore_overlaps = self.iou_calculator(
+ gt_bboxes_ignore, priors, mode='iof')
+ ignore_max_overlaps, _ = ignore_overlaps.max(dim=0)
+ overlaps[:, ignore_max_overlaps > self.ignore_iof_thr] = -1
+
+ assign_result = self.assign_wrt_overlaps(overlaps, gt_labels)
+ if assign_on_cpu:
+ assign_result.gt_inds = assign_result.gt_inds.to(device)
+ assign_result.max_overlaps = assign_result.max_overlaps.to(device)
+ if assign_result.labels is not None:
+ assign_result.labels = assign_result.labels.to(device)
+ return assign_result
+
+ def assign_wrt_overlaps(self, overlaps: Tensor,
+ gt_labels: Tensor) -> AssignResult:
+ """Assign w.r.t. the overlaps of priors with gts.
+
+ Args:
+ overlaps (Tensor): Overlaps between k gt_bboxes and n bboxes,
+ shape(k, n).
+ gt_labels (Tensor): Labels of k gt_bboxes, shape (k, ).
+
+ Returns:
+ :obj:`AssignResult`: The assign result.
+ """
+ num_gts, num_bboxes = overlaps.size(0), overlaps.size(1)
+
+ # 1. assign -1 by default
+ assigned_gt_inds = overlaps.new_full((num_bboxes, ),
+ -1,
+ dtype=torch.long)
+
+ if num_gts == 0 or num_bboxes == 0:
+ # No ground truth or boxes, return empty assignment
+ max_overlaps = overlaps.new_zeros((num_bboxes, ))
+ assigned_labels = overlaps.new_full((num_bboxes, ),
+ -1,
+ dtype=torch.long)
+ if num_gts == 0:
+ # No truth, assign everything to background
+ assigned_gt_inds[:] = 0
+ return AssignResult(
+ num_gts=num_gts,
+ gt_inds=assigned_gt_inds,
+ max_overlaps=max_overlaps,
+ labels=assigned_labels)
+
+ # for each anchor, which gt best overlaps with it
+ # for each anchor, the max iou of all gts
+ max_overlaps, argmax_overlaps = overlaps.max(dim=0)
+ # for each gt, which anchor best overlaps with it
+ # for each gt, the max iou of all proposals
+ gt_max_overlaps, gt_argmax_overlaps = overlaps.max(dim=1)
+
+ # 2. assign negative: below
+ # the negative inds are set to be 0
+ if isinstance(self.neg_iou_thr, float):
+ assigned_gt_inds[(max_overlaps >= 0)
+ & (max_overlaps < self.neg_iou_thr)] = 0
+ elif isinstance(self.neg_iou_thr, tuple):
+ assert len(self.neg_iou_thr) == 2
+ assigned_gt_inds[(max_overlaps >= self.neg_iou_thr[0])
+ & (max_overlaps < self.neg_iou_thr[1])] = 0
+
+ # 3. assign positive: above positive IoU threshold
+ pos_inds = max_overlaps >= self.pos_iou_thr
+ assigned_gt_inds[pos_inds] = argmax_overlaps[pos_inds] + 1
+
+ if self.match_low_quality:
+ # Low-quality matching will overwrite the assigned_gt_inds assigned
+ # in Step 3. Thus, the assigned gt might not be the best one for
+ # prediction.
+ # For example, if bbox A has 0.9 and 0.8 iou with GT bbox 1 & 2,
+ # bbox 1 will be assigned as the best target for bbox A in step 3.
+ # However, if GT bbox 2's gt_argmax_overlaps = A, bbox A's
+ # assigned_gt_inds will be overwritten to be bbox 2.
+ # This might be the reason that it is not used in ROI Heads.
+ for i in range(num_gts):
+ if gt_max_overlaps[i] >= self.min_pos_iou:
+ if self.gt_max_assign_all:
+ max_iou_inds = overlaps[i, :] == gt_max_overlaps[i]
+ assigned_gt_inds[max_iou_inds] = i + 1
+ else:
+ assigned_gt_inds[gt_argmax_overlaps[i]] = i + 1
+
+ assigned_labels = assigned_gt_inds.new_full((num_bboxes, ), -1)
+ pos_inds = torch.nonzero(
+ assigned_gt_inds > 0, as_tuple=False).squeeze()
+ if pos_inds.numel() > 0:
+ assigned_labels[pos_inds] = gt_labels[assigned_gt_inds[pos_inds] -
+ 1]
+
+ return AssignResult(
+ num_gts=num_gts,
+ gt_inds=assigned_gt_inds,
+ max_overlaps=max_overlaps,
+ labels=assigned_labels)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/multi_instance_assigner.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/multi_instance_assigner.py
new file mode 100644
index 0000000000000000000000000000000000000000..1ba32afe856b3c2ad03ed89562d080f15b6ccf30
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/multi_instance_assigner.py
@@ -0,0 +1,140 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional
+
+import torch
+from mmengine.structures import InstanceData
+
+from mmdet.registry import TASK_UTILS
+from .assign_result import AssignResult
+from .max_iou_assigner import MaxIoUAssigner
+
+
+@TASK_UTILS.register_module()
+class MultiInstanceAssigner(MaxIoUAssigner):
+ """Assign a corresponding gt bbox or background to each proposal bbox. If
+ we need to use a proposal box to generate multiple predict boxes,
+ `MultiInstanceAssigner` can assign multiple gt to each proposal box.
+
+ Args:
+ num_instance (int): How many bboxes are predicted by each proposal box.
+ """
+
+ def __init__(self, num_instance: int = 2, **kwargs):
+ super().__init__(**kwargs)
+ self.num_instance = num_instance
+
+ def assign(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ **kwargs) -> AssignResult:
+ """Assign gt to bboxes.
+
+ This method assign gt bboxes to every bbox (proposal/anchor), each bbox
+ is assigned a set of gts, and the number of gts in this set is defined
+ by `self.num_instance`.
+
+ Args:
+ pred_instances (:obj:`InstanceData`): Instances of model
+ predictions. It includes ``priors``, and the priors can
+ be anchors or points, or the bboxes predicted by the
+ previous stage, has shape (n, 4). The bboxes predicted by
+ the current model or stage will be named ``bboxes``,
+ ``labels``, and ``scores``, the same as the ``InstanceData``
+ in other places.
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes``, with shape (k, 4),
+ and ``labels``, with shape (k, ).
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes``
+ attribute data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ :obj:`AssignResult`: The assign result.
+ """
+ gt_bboxes = gt_instances.bboxes
+ priors = pred_instances.priors
+ # Set the FG label to 1 and add ignored annotations
+ gt_labels = gt_instances.labels + 1
+ if gt_instances_ignore is not None:
+ gt_bboxes_ignore = gt_instances_ignore.bboxes
+ if hasattr(gt_instances_ignore, 'labels'):
+ gt_labels_ignore = gt_instances_ignore.labels
+ else:
+ gt_labels_ignore = torch.ones_like(gt_bboxes_ignore)[:, 0] * -1
+ else:
+ gt_bboxes_ignore = None
+ gt_labels_ignore = None
+
+ assign_on_cpu = True if (self.gpu_assign_thr > 0) and (
+ gt_bboxes.shape[0] > self.gpu_assign_thr) else False
+ # compute overlap and assign gt on CPU when number of GT is large
+ if assign_on_cpu:
+ device = priors.device
+ priors = priors.cpu()
+ gt_bboxes = gt_bboxes.cpu()
+ gt_labels = gt_labels.cpu()
+ if gt_bboxes_ignore is not None:
+ gt_bboxes_ignore = gt_bboxes_ignore.cpu()
+ gt_labels_ignore = gt_labels_ignore.cpu()
+
+ if gt_bboxes_ignore is not None:
+ all_bboxes = torch.cat([gt_bboxes, gt_bboxes_ignore], dim=0)
+ all_labels = torch.cat([gt_labels, gt_labels_ignore], dim=0)
+ else:
+ all_bboxes = gt_bboxes
+ all_labels = gt_labels
+ all_priors = torch.cat([priors, all_bboxes], dim=0)
+
+ overlaps_normal = self.iou_calculator(
+ all_priors, all_bboxes, mode='iou')
+ overlaps_ignore = self.iou_calculator(
+ all_priors, all_bboxes, mode='iof')
+ gt_ignore_mask = all_labels.eq(-1).repeat(all_priors.shape[0], 1)
+ overlaps_normal = overlaps_normal * ~gt_ignore_mask
+ overlaps_ignore = overlaps_ignore * gt_ignore_mask
+
+ overlaps_normal, overlaps_normal_indices = overlaps_normal.sort(
+ descending=True, dim=1)
+ overlaps_ignore, overlaps_ignore_indices = overlaps_ignore.sort(
+ descending=True, dim=1)
+
+ # select the roi with the higher score
+ max_overlaps_normal = overlaps_normal[:, :self.num_instance].flatten()
+ gt_assignment_normal = overlaps_normal_indices[:, :self.
+ num_instance].flatten()
+ max_overlaps_ignore = overlaps_ignore[:, :self.num_instance].flatten()
+ gt_assignment_ignore = overlaps_ignore_indices[:, :self.
+ num_instance].flatten()
+
+ # ignore or not
+ ignore_assign_mask = (max_overlaps_normal < self.pos_iou_thr) * (
+ max_overlaps_ignore > max_overlaps_normal)
+ overlaps = (max_overlaps_normal * ~ignore_assign_mask) + (
+ max_overlaps_ignore * ignore_assign_mask)
+ gt_assignment = (gt_assignment_normal * ~ignore_assign_mask) + (
+ gt_assignment_ignore * ignore_assign_mask)
+
+ assigned_labels = all_labels[gt_assignment]
+ fg_mask = (overlaps >= self.pos_iou_thr) * (assigned_labels != -1)
+ bg_mask = (overlaps < self.neg_iou_thr) * (overlaps >= 0)
+ assigned_labels[fg_mask] = 1
+ assigned_labels[bg_mask] = 0
+
+ overlaps = overlaps.reshape(-1, self.num_instance)
+ gt_assignment = gt_assignment.reshape(-1, self.num_instance)
+ assigned_labels = assigned_labels.reshape(-1, self.num_instance)
+
+ assign_result = AssignResult(
+ num_gts=all_bboxes.size(0),
+ gt_inds=gt_assignment,
+ max_overlaps=overlaps,
+ labels=assigned_labels)
+
+ if assign_on_cpu:
+ assign_result.gt_inds = assign_result.gt_inds.to(device)
+ assign_result.max_overlaps = assign_result.max_overlaps.to(device)
+ if assign_result.labels is not None:
+ assign_result.labels = assign_result.labels.to(device)
+ return assign_result
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/point_assigner.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/point_assigner.py
new file mode 100644
index 0000000000000000000000000000000000000000..4da60a490b0022ac76c46db8a34f814bc9da8e2e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/point_assigner.py
@@ -0,0 +1,155 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional
+
+import torch
+from mmengine.structures import InstanceData
+
+from mmdet.registry import TASK_UTILS
+from .assign_result import AssignResult
+from .base_assigner import BaseAssigner
+
+
+@TASK_UTILS.register_module()
+class PointAssigner(BaseAssigner):
+ """Assign a corresponding gt bbox or background to each point.
+
+ Each proposals will be assigned with `0`, or a positive integer
+ indicating the ground truth index.
+
+ - 0: negative sample, no assigned gt
+ - positive integer: positive sample, index (1-based) of assigned gt
+ """
+
+ def __init__(self, scale: int = 4, pos_num: int = 3) -> None:
+ self.scale = scale
+ self.pos_num = pos_num
+
+ def assign(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ **kwargs) -> AssignResult:
+ """Assign gt to points.
+
+ This method assign a gt bbox to every points set, each points set
+ will be assigned with the background_label (-1), or a label number.
+ -1 is background, and semi-positive number is the index (0-based) of
+ assigned gt.
+ The assignment is done in following steps, the order matters.
+
+ 1. assign every points to the background_label (-1)
+ 2. A point is assigned to some gt bbox if
+ (i) the point is within the k closest points to the gt bbox
+ (ii) the distance between this point and the gt is smaller than
+ other gt bboxes
+
+ Args:
+ pred_instances (:obj:`InstanceData`): Instances of model
+ predictions. It includes ``priors``, and the priors can
+ be anchors or points, or the bboxes predicted by the
+ previous stage, has shape (n, 4). The bboxes predicted by
+ the current model or stage will be named ``bboxes``,
+ ``labels``, and ``scores``, the same as the ``InstanceData``
+ in other places.
+
+
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes``, with shape (k, 4),
+ and ``labels``, with shape (k, ).
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes``
+ attribute data that is ignored during training and testing.
+ Defaults to None.
+ Returns:
+ :obj:`AssignResult`: The assign result.
+ """
+ gt_bboxes = gt_instances.bboxes
+ gt_labels = gt_instances.labels
+ # points to be assigned, shape(n, 3) while last
+ # dimension stands for (x, y, stride).
+ points = pred_instances.priors
+
+ num_points = points.shape[0]
+ num_gts = gt_bboxes.shape[0]
+
+ if num_gts == 0 or num_points == 0:
+ # If no truth assign everything to the background
+ assigned_gt_inds = points.new_full((num_points, ),
+ 0,
+ dtype=torch.long)
+ assigned_labels = points.new_full((num_points, ),
+ -1,
+ dtype=torch.long)
+ return AssignResult(
+ num_gts=num_gts,
+ gt_inds=assigned_gt_inds,
+ max_overlaps=None,
+ labels=assigned_labels)
+
+ points_xy = points[:, :2]
+ points_stride = points[:, 2]
+ points_lvl = torch.log2(
+ points_stride).int() # [3...,4...,5...,6...,7...]
+ lvl_min, lvl_max = points_lvl.min(), points_lvl.max()
+
+ # assign gt box
+ gt_bboxes_xy = (gt_bboxes[:, :2] + gt_bboxes[:, 2:]) / 2
+ gt_bboxes_wh = (gt_bboxes[:, 2:] - gt_bboxes[:, :2]).clamp(min=1e-6)
+ scale = self.scale
+ gt_bboxes_lvl = ((torch.log2(gt_bboxes_wh[:, 0] / scale) +
+ torch.log2(gt_bboxes_wh[:, 1] / scale)) / 2).int()
+ gt_bboxes_lvl = torch.clamp(gt_bboxes_lvl, min=lvl_min, max=lvl_max)
+
+ # stores the assigned gt index of each point
+ assigned_gt_inds = points.new_zeros((num_points, ), dtype=torch.long)
+ # stores the assigned gt dist (to this point) of each point
+ assigned_gt_dist = points.new_full((num_points, ), float('inf'))
+ points_range = torch.arange(points.shape[0])
+
+ for idx in range(num_gts):
+ gt_lvl = gt_bboxes_lvl[idx]
+ # get the index of points in this level
+ lvl_idx = gt_lvl == points_lvl
+ points_index = points_range[lvl_idx]
+ # get the points in this level
+ lvl_points = points_xy[lvl_idx, :]
+ # get the center point of gt
+ gt_point = gt_bboxes_xy[[idx], :]
+ # get width and height of gt
+ gt_wh = gt_bboxes_wh[[idx], :]
+ # compute the distance between gt center and
+ # all points in this level
+ points_gt_dist = ((lvl_points - gt_point) / gt_wh).norm(dim=1)
+ # find the nearest k points to gt center in this level
+ min_dist, min_dist_index = torch.topk(
+ points_gt_dist, self.pos_num, largest=False)
+ # the index of nearest k points to gt center in this level
+ min_dist_points_index = points_index[min_dist_index]
+ # The less_than_recorded_index stores the index
+ # of min_dist that is less then the assigned_gt_dist. Where
+ # assigned_gt_dist stores the dist from previous assigned gt
+ # (if exist) to each point.
+ less_than_recorded_index = min_dist < assigned_gt_dist[
+ min_dist_points_index]
+ # The min_dist_points_index stores the index of points satisfy:
+ # (1) it is k nearest to current gt center in this level.
+ # (2) it is closer to current gt center than other gt center.
+ min_dist_points_index = min_dist_points_index[
+ less_than_recorded_index]
+ # assign the result
+ assigned_gt_inds[min_dist_points_index] = idx + 1
+ assigned_gt_dist[min_dist_points_index] = min_dist[
+ less_than_recorded_index]
+
+ assigned_labels = assigned_gt_inds.new_full((num_points, ), -1)
+ pos_inds = torch.nonzero(
+ assigned_gt_inds > 0, as_tuple=False).squeeze()
+ if pos_inds.numel() > 0:
+ assigned_labels[pos_inds] = gt_labels[assigned_gt_inds[pos_inds] -
+ 1]
+
+ return AssignResult(
+ num_gts=num_gts,
+ gt_inds=assigned_gt_inds,
+ max_overlaps=None,
+ labels=assigned_labels)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/region_assigner.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/region_assigner.py
new file mode 100644
index 0000000000000000000000000000000000000000..df549143086c1195efaf12a2f3e81259da0e6c97
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/region_assigner.py
@@ -0,0 +1,239 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple
+
+import torch
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import TASK_UTILS
+from ..prior_generators import anchor_inside_flags
+from .assign_result import AssignResult
+from .base_assigner import BaseAssigner
+
+
+def calc_region(
+ bbox: Tensor,
+ ratio: float,
+ stride: int,
+ featmap_size: Optional[Tuple[int, int]] = None) -> Tuple[Tensor]:
+ """Calculate region of the box defined by the ratio, the ratio is from the
+ center of the box to every edge."""
+ # project bbox on the feature
+ f_bbox = bbox / stride
+ x1 = torch.round((1 - ratio) * f_bbox[0] + ratio * f_bbox[2])
+ y1 = torch.round((1 - ratio) * f_bbox[1] + ratio * f_bbox[3])
+ x2 = torch.round(ratio * f_bbox[0] + (1 - ratio) * f_bbox[2])
+ y2 = torch.round(ratio * f_bbox[1] + (1 - ratio) * f_bbox[3])
+ if featmap_size is not None:
+ x1 = x1.clamp(min=0, max=featmap_size[1])
+ y1 = y1.clamp(min=0, max=featmap_size[0])
+ x2 = x2.clamp(min=0, max=featmap_size[1])
+ y2 = y2.clamp(min=0, max=featmap_size[0])
+ return (x1, y1, x2, y2)
+
+
+def anchor_ctr_inside_region_flags(anchors: Tensor, stride: int,
+ region: Tuple[Tensor]) -> Tensor:
+ """Get the flag indicate whether anchor centers are inside regions."""
+ x1, y1, x2, y2 = region
+ f_anchors = anchors / stride
+ x = (f_anchors[:, 0] + f_anchors[:, 2]) * 0.5
+ y = (f_anchors[:, 1] + f_anchors[:, 3]) * 0.5
+ flags = (x >= x1) & (x <= x2) & (y >= y1) & (y <= y2)
+ return flags
+
+
+@TASK_UTILS.register_module()
+class RegionAssigner(BaseAssigner):
+ """Assign a corresponding gt bbox or background to each bbox.
+
+ Each proposals will be assigned with `-1`, `0`, or a positive integer
+ indicating the ground truth index.
+
+ - -1: don't care
+ - 0: negative sample, no assigned gt
+ - positive integer: positive sample, index (1-based) of assigned gt
+
+ Args:
+ center_ratio (float): ratio of the region in the center of the bbox to
+ define positive sample.
+ ignore_ratio (float): ratio of the region to define ignore samples.
+ """
+
+ def __init__(self,
+ center_ratio: float = 0.2,
+ ignore_ratio: float = 0.5) -> None:
+ self.center_ratio = center_ratio
+ self.ignore_ratio = ignore_ratio
+
+ def assign(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ img_meta: dict,
+ featmap_sizes: List[Tuple[int, int]],
+ num_level_anchors: List[int],
+ anchor_scale: int,
+ anchor_strides: List[int],
+ gt_instances_ignore: Optional[InstanceData] = None,
+ allowed_border: int = 0) -> AssignResult:
+ """Assign gt to anchors.
+
+ This method assign a gt bbox to every bbox (proposal/anchor), each bbox
+ will be assigned with -1, 0, or a positive number. -1 means don't care,
+ 0 means negative sample, positive number is the index (1-based) of
+ assigned gt.
+
+ The assignment is done in following steps, and the order matters.
+
+ 1. Assign every anchor to 0 (negative)
+ 2. (For each gt_bboxes) Compute ignore flags based on ignore_region
+ then assign -1 to anchors w.r.t. ignore flags
+ 3. (For each gt_bboxes) Compute pos flags based on center_region then
+ assign gt_bboxes to anchors w.r.t. pos flags
+ 4. (For each gt_bboxes) Compute ignore flags based on adjacent anchor
+ level then assign -1 to anchors w.r.t. ignore flags
+ 5. Assign anchor outside of image to -1
+
+ Args:
+ pred_instances (:obj:`InstanceData`): Instances of model
+ predictions. It includes ``priors``, and the priors can
+ be anchors or points, or the bboxes predicted by the
+ previous stage, has shape (n, 4). The bboxes predicted by
+ the current model or stage will be named ``bboxes``,
+ ``labels``, and ``scores``, the same as the ``InstanceData``
+ in other places.
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes``, with shape (k, 4),
+ and ``labels``, with shape (k, ).
+ img_meta (dict): Meta info of image.
+ featmap_sizes (list[tuple[int, int]]): Feature map size each level.
+ num_level_anchors (list[int]): The number of anchors in each level.
+ anchor_scale (int): Scale of the anchor.
+ anchor_strides (list[int]): Stride of the anchor.
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes``
+ attribute data that is ignored during training and testing.
+ Defaults to None.
+ allowed_border (int, optional): The border to allow the valid
+ anchor. Defaults to 0.
+
+ Returns:
+ :obj:`AssignResult`: The assign result.
+ """
+ if gt_instances_ignore is not None:
+ raise NotImplementedError
+
+ num_gts = len(gt_instances)
+ num_bboxes = len(pred_instances)
+
+ gt_bboxes = gt_instances.bboxes
+ gt_labels = gt_instances.labels
+ flat_anchors = pred_instances.priors
+ flat_valid_flags = pred_instances.valid_flags
+ mlvl_anchors = torch.split(flat_anchors, num_level_anchors)
+
+ if num_gts == 0 or num_bboxes == 0:
+ # No ground truth or boxes, return empty assignment
+ max_overlaps = gt_bboxes.new_zeros((num_bboxes, ))
+ assigned_gt_inds = gt_bboxes.new_zeros((num_bboxes, ),
+ dtype=torch.long)
+ assigned_labels = gt_bboxes.new_full((num_bboxes, ),
+ -1,
+ dtype=torch.long)
+ return AssignResult(
+ num_gts=num_gts,
+ gt_inds=assigned_gt_inds,
+ max_overlaps=max_overlaps,
+ labels=assigned_labels)
+
+ num_lvls = len(mlvl_anchors)
+ r1 = (1 - self.center_ratio) / 2
+ r2 = (1 - self.ignore_ratio) / 2
+
+ scale = torch.sqrt((gt_bboxes[:, 2] - gt_bboxes[:, 0]) *
+ (gt_bboxes[:, 3] - gt_bboxes[:, 1]))
+ min_anchor_size = scale.new_full(
+ (1, ), float(anchor_scale * anchor_strides[0]))
+ target_lvls = torch.floor(
+ torch.log2(scale) - torch.log2(min_anchor_size) + 0.5)
+ target_lvls = target_lvls.clamp(min=0, max=num_lvls - 1).long()
+
+ # 1. assign 0 (negative) by default
+ mlvl_assigned_gt_inds = []
+ mlvl_ignore_flags = []
+ for lvl in range(num_lvls):
+ assigned_gt_inds = gt_bboxes.new_full((num_level_anchors[lvl], ),
+ 0,
+ dtype=torch.long)
+ ignore_flags = torch.zeros_like(assigned_gt_inds)
+ mlvl_assigned_gt_inds.append(assigned_gt_inds)
+ mlvl_ignore_flags.append(ignore_flags)
+
+ for gt_id in range(num_gts):
+ lvl = target_lvls[gt_id].item()
+ featmap_size = featmap_sizes[lvl]
+ stride = anchor_strides[lvl]
+ anchors = mlvl_anchors[lvl]
+ gt_bbox = gt_bboxes[gt_id, :4]
+
+ # Compute regions
+ ignore_region = calc_region(gt_bbox, r2, stride, featmap_size)
+ ctr_region = calc_region(gt_bbox, r1, stride, featmap_size)
+
+ # 2. Assign -1 to ignore flags
+ ignore_flags = anchor_ctr_inside_region_flags(
+ anchors, stride, ignore_region)
+ mlvl_assigned_gt_inds[lvl][ignore_flags] = -1
+
+ # 3. Assign gt_bboxes to pos flags
+ pos_flags = anchor_ctr_inside_region_flags(anchors, stride,
+ ctr_region)
+ mlvl_assigned_gt_inds[lvl][pos_flags] = gt_id + 1
+
+ # 4. Assign -1 to ignore adjacent lvl
+ if lvl > 0:
+ d_lvl = lvl - 1
+ d_anchors = mlvl_anchors[d_lvl]
+ d_featmap_size = featmap_sizes[d_lvl]
+ d_stride = anchor_strides[d_lvl]
+ d_ignore_region = calc_region(gt_bbox, r2, d_stride,
+ d_featmap_size)
+ ignore_flags = anchor_ctr_inside_region_flags(
+ d_anchors, d_stride, d_ignore_region)
+ mlvl_ignore_flags[d_lvl][ignore_flags] = 1
+ if lvl < num_lvls - 1:
+ u_lvl = lvl + 1
+ u_anchors = mlvl_anchors[u_lvl]
+ u_featmap_size = featmap_sizes[u_lvl]
+ u_stride = anchor_strides[u_lvl]
+ u_ignore_region = calc_region(gt_bbox, r2, u_stride,
+ u_featmap_size)
+ ignore_flags = anchor_ctr_inside_region_flags(
+ u_anchors, u_stride, u_ignore_region)
+ mlvl_ignore_flags[u_lvl][ignore_flags] = 1
+
+ # 4. (cont.) Assign -1 to ignore adjacent lvl
+ for lvl in range(num_lvls):
+ ignore_flags = mlvl_ignore_flags[lvl]
+ mlvl_assigned_gt_inds[lvl][ignore_flags == 1] = -1
+
+ # 5. Assign -1 to anchor outside of image
+ flat_assigned_gt_inds = torch.cat(mlvl_assigned_gt_inds)
+ assert (flat_assigned_gt_inds.shape[0] == flat_anchors.shape[0] ==
+ flat_valid_flags.shape[0])
+ inside_flags = anchor_inside_flags(flat_anchors, flat_valid_flags,
+ img_meta['img_shape'],
+ allowed_border)
+ outside_flags = ~inside_flags
+ flat_assigned_gt_inds[outside_flags] = -1
+
+ assigned_labels = torch.zeros_like(flat_assigned_gt_inds)
+ pos_flags = flat_assigned_gt_inds > 0
+ assigned_labels[pos_flags] = gt_labels[flat_assigned_gt_inds[pos_flags]
+ - 1]
+
+ return AssignResult(
+ num_gts=num_gts,
+ gt_inds=flat_assigned_gt_inds,
+ max_overlaps=None,
+ labels=assigned_labels)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/sim_ota_assigner.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/sim_ota_assigner.py
new file mode 100644
index 0000000000000000000000000000000000000000..d54a8b91d132d9bf661267de666bfed7e915a65a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/sim_ota_assigner.py
@@ -0,0 +1,223 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Tuple
+
+import torch
+import torch.nn.functional as F
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import TASK_UTILS
+from mmdet.utils import ConfigType
+from .assign_result import AssignResult
+from .base_assigner import BaseAssigner
+
+INF = 100000.0
+EPS = 1.0e-7
+
+
+@TASK_UTILS.register_module()
+class SimOTAAssigner(BaseAssigner):
+ """Computes matching between predictions and ground truth.
+
+ Args:
+ center_radius (float): Ground truth center size
+ to judge whether a prior is in center. Defaults to 2.5.
+ candidate_topk (int): The candidate top-k which used to
+ get top-k ious to calculate dynamic-k. Defaults to 10.
+ iou_weight (float): The scale factor for regression
+ iou cost. Defaults to 3.0.
+ cls_weight (float): The scale factor for classification
+ cost. Defaults to 1.0.
+ iou_calculator (ConfigType): Config of overlaps Calculator.
+ Defaults to dict(type='BboxOverlaps2D').
+ """
+
+ def __init__(self,
+ center_radius: float = 2.5,
+ candidate_topk: int = 10,
+ iou_weight: float = 3.0,
+ cls_weight: float = 1.0,
+ iou_calculator: ConfigType = dict(type='BboxOverlaps2D')):
+ self.center_radius = center_radius
+ self.candidate_topk = candidate_topk
+ self.iou_weight = iou_weight
+ self.cls_weight = cls_weight
+ self.iou_calculator = TASK_UTILS.build(iou_calculator)
+
+ def assign(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ **kwargs) -> AssignResult:
+ """Assign gt to priors using SimOTA.
+
+ Args:
+ pred_instances (:obj:`InstanceData`): Instances of model
+ predictions. It includes ``priors``, and the priors can
+ be anchors or points, or the bboxes predicted by the
+ previous stage, has shape (n, 4). The bboxes predicted by
+ the current model or stage will be named ``bboxes``,
+ ``labels``, and ``scores``, the same as the ``InstanceData``
+ in other places.
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes``, with shape (k, 4),
+ and ``labels``, with shape (k, ).
+ gt_instances_ignore (:obj:`InstanceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes``
+ attribute data that is ignored during training and testing.
+ Defaults to None.
+ Returns:
+ obj:`AssignResult`: The assigned result.
+ """
+ gt_bboxes = gt_instances.bboxes
+ gt_labels = gt_instances.labels
+ num_gt = gt_bboxes.size(0)
+
+ decoded_bboxes = pred_instances.bboxes
+ pred_scores = pred_instances.scores
+ priors = pred_instances.priors
+ num_bboxes = decoded_bboxes.size(0)
+
+ # assign 0 by default
+ assigned_gt_inds = decoded_bboxes.new_full((num_bboxes, ),
+ 0,
+ dtype=torch.long)
+ if num_gt == 0 or num_bboxes == 0:
+ # No ground truth or boxes, return empty assignment
+ max_overlaps = decoded_bboxes.new_zeros((num_bboxes, ))
+ assigned_labels = decoded_bboxes.new_full((num_bboxes, ),
+ -1,
+ dtype=torch.long)
+ return AssignResult(
+ num_gt, assigned_gt_inds, max_overlaps, labels=assigned_labels)
+
+ valid_mask, is_in_boxes_and_center = self.get_in_gt_and_in_center_info(
+ priors, gt_bboxes)
+ valid_decoded_bbox = decoded_bboxes[valid_mask]
+ valid_pred_scores = pred_scores[valid_mask]
+ num_valid = valid_decoded_bbox.size(0)
+ if num_valid == 0:
+ # No valid bboxes, return empty assignment
+ max_overlaps = decoded_bboxes.new_zeros((num_bboxes, ))
+ assigned_labels = decoded_bboxes.new_full((num_bboxes, ),
+ -1,
+ dtype=torch.long)
+ return AssignResult(
+ num_gt, assigned_gt_inds, max_overlaps, labels=assigned_labels)
+
+ pairwise_ious = self.iou_calculator(valid_decoded_bbox, gt_bboxes)
+ iou_cost = -torch.log(pairwise_ious + EPS)
+
+ gt_onehot_label = (
+ F.one_hot(gt_labels.to(torch.int64),
+ pred_scores.shape[-1]).float().unsqueeze(0).repeat(
+ num_valid, 1, 1))
+
+ valid_pred_scores = valid_pred_scores.unsqueeze(1).repeat(1, num_gt, 1)
+ # disable AMP autocast and calculate BCE with FP32 to avoid overflow
+ with torch.cuda.amp.autocast(enabled=False):
+ cls_cost = (
+ F.binary_cross_entropy(
+ valid_pred_scores.to(dtype=torch.float32),
+ gt_onehot_label,
+ reduction='none',
+ ).sum(-1).to(dtype=valid_pred_scores.dtype))
+
+ cost_matrix = (
+ cls_cost * self.cls_weight + iou_cost * self.iou_weight +
+ (~is_in_boxes_and_center) * INF)
+
+ matched_pred_ious, matched_gt_inds = \
+ self.dynamic_k_matching(
+ cost_matrix, pairwise_ious, num_gt, valid_mask)
+
+ # convert to AssignResult format
+ assigned_gt_inds[valid_mask] = matched_gt_inds + 1
+ assigned_labels = assigned_gt_inds.new_full((num_bboxes, ), -1)
+ assigned_labels[valid_mask] = gt_labels[matched_gt_inds].long()
+ max_overlaps = assigned_gt_inds.new_full((num_bboxes, ),
+ -INF,
+ dtype=torch.float32)
+ max_overlaps[valid_mask] = matched_pred_ious
+ return AssignResult(
+ num_gt, assigned_gt_inds, max_overlaps, labels=assigned_labels)
+
+ def get_in_gt_and_in_center_info(
+ self, priors: Tensor, gt_bboxes: Tensor) -> Tuple[Tensor, Tensor]:
+ """Get the information of which prior is in gt bboxes and gt center
+ priors."""
+ num_gt = gt_bboxes.size(0)
+
+ repeated_x = priors[:, 0].unsqueeze(1).repeat(1, num_gt)
+ repeated_y = priors[:, 1].unsqueeze(1).repeat(1, num_gt)
+ repeated_stride_x = priors[:, 2].unsqueeze(1).repeat(1, num_gt)
+ repeated_stride_y = priors[:, 3].unsqueeze(1).repeat(1, num_gt)
+
+ # is prior centers in gt bboxes, shape: [n_prior, n_gt]
+ l_ = repeated_x - gt_bboxes[:, 0]
+ t_ = repeated_y - gt_bboxes[:, 1]
+ r_ = gt_bboxes[:, 2] - repeated_x
+ b_ = gt_bboxes[:, 3] - repeated_y
+
+ deltas = torch.stack([l_, t_, r_, b_], dim=1)
+ is_in_gts = deltas.min(dim=1).values > 0
+ is_in_gts_all = is_in_gts.sum(dim=1) > 0
+
+ # is prior centers in gt centers
+ gt_cxs = (gt_bboxes[:, 0] + gt_bboxes[:, 2]) / 2.0
+ gt_cys = (gt_bboxes[:, 1] + gt_bboxes[:, 3]) / 2.0
+ ct_box_l = gt_cxs - self.center_radius * repeated_stride_x
+ ct_box_t = gt_cys - self.center_radius * repeated_stride_y
+ ct_box_r = gt_cxs + self.center_radius * repeated_stride_x
+ ct_box_b = gt_cys + self.center_radius * repeated_stride_y
+
+ cl_ = repeated_x - ct_box_l
+ ct_ = repeated_y - ct_box_t
+ cr_ = ct_box_r - repeated_x
+ cb_ = ct_box_b - repeated_y
+
+ ct_deltas = torch.stack([cl_, ct_, cr_, cb_], dim=1)
+ is_in_cts = ct_deltas.min(dim=1).values > 0
+ is_in_cts_all = is_in_cts.sum(dim=1) > 0
+
+ # in boxes or in centers, shape: [num_priors]
+ is_in_gts_or_centers = is_in_gts_all | is_in_cts_all
+
+ # both in boxes and centers, shape: [num_fg, num_gt]
+ is_in_boxes_and_centers = (
+ is_in_gts[is_in_gts_or_centers, :]
+ & is_in_cts[is_in_gts_or_centers, :])
+ return is_in_gts_or_centers, is_in_boxes_and_centers
+
+ def dynamic_k_matching(self, cost: Tensor, pairwise_ious: Tensor,
+ num_gt: int,
+ valid_mask: Tensor) -> Tuple[Tensor, Tensor]:
+ """Use IoU and matching cost to calculate the dynamic top-k positive
+ targets."""
+ matching_matrix = torch.zeros_like(cost, dtype=torch.uint8)
+ # select candidate topk ious for dynamic-k calculation
+ candidate_topk = min(self.candidate_topk, pairwise_ious.size(0))
+ topk_ious, _ = torch.topk(pairwise_ious, candidate_topk, dim=0)
+ # calculate dynamic k for each gt
+ dynamic_ks = torch.clamp(topk_ious.sum(0).int(), min=1)
+ for gt_idx in range(num_gt):
+ _, pos_idx = torch.topk(
+ cost[:, gt_idx], k=dynamic_ks[gt_idx], largest=False)
+ matching_matrix[:, gt_idx][pos_idx] = 1
+
+ del topk_ious, dynamic_ks, pos_idx
+
+ prior_match_gt_mask = matching_matrix.sum(1) > 1
+ if prior_match_gt_mask.sum() > 0:
+ cost_min, cost_argmin = torch.min(
+ cost[prior_match_gt_mask, :], dim=1)
+ matching_matrix[prior_match_gt_mask, :] *= 0
+ matching_matrix[prior_match_gt_mask, cost_argmin] = 1
+ # get foreground mask inside box and center prior
+ fg_mask_inboxes = matching_matrix.sum(1) > 0
+ valid_mask[valid_mask.clone()] = fg_mask_inboxes
+
+ matched_gt_inds = matching_matrix[fg_mask_inboxes, :].argmax(1)
+ matched_pred_ious = (matching_matrix *
+ pairwise_ious).sum(1)[fg_mask_inboxes]
+ return matched_pred_ious, matched_gt_inds
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/task_aligned_assigner.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/task_aligned_assigner.py
new file mode 100644
index 0000000000000000000000000000000000000000..220ea8485933ab3243f6c1e205dbf1b973df08d7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/task_aligned_assigner.py
@@ -0,0 +1,158 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional
+
+import torch
+from mmengine.structures import InstanceData
+
+from mmdet.registry import TASK_UTILS
+from mmdet.utils import ConfigType
+from .assign_result import AssignResult
+from .base_assigner import BaseAssigner
+
+INF = 100000000
+
+
+@TASK_UTILS.register_module()
+class TaskAlignedAssigner(BaseAssigner):
+ """Task aligned assigner used in the paper:
+ `TOOD: Task-aligned One-stage Object Detection.
+ `_.
+
+ Assign a corresponding gt bbox or background to each predicted bbox.
+ Each bbox will be assigned with `0` or a positive integer
+ indicating the ground truth index.
+
+ - 0: negative sample, no assigned gt
+ - positive integer: positive sample, index (1-based) of assigned gt
+
+ Args:
+ topk (int): number of bbox selected in each level
+ iou_calculator (:obj:`ConfigDict` or dict): Config dict for iou
+ calculator. Defaults to ``dict(type='BboxOverlaps2D')``
+ """
+
+ def __init__(self,
+ topk: int,
+ iou_calculator: ConfigType = dict(type='BboxOverlaps2D')):
+ assert topk >= 1
+ self.topk = topk
+ self.iou_calculator = TASK_UTILS.build(iou_calculator)
+
+ def assign(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ gt_instances_ignore: Optional[InstanceData] = None,
+ alpha: int = 1,
+ beta: int = 6) -> AssignResult:
+ """Assign gt to bboxes.
+
+ The assignment is done in following steps
+
+ 1. compute alignment metric between all bbox (bbox of all pyramid
+ levels) and gt
+ 2. select top-k bbox as candidates for each gt
+ 3. limit the positive sample's center in gt (because the anchor-free
+ detector only can predict positive distance)
+
+
+ Args:
+ pred_instances (:obj:`InstaceData`): Instances of model
+ predictions. It includes ``priors``, and the priors can
+ be anchors, points, or bboxes predicted by the model,
+ shape(n, 4).
+ gt_instances (:obj:`InstaceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ gt_instances_ignore (:obj:`InstaceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes``
+ attribute data that is ignored during training and testing.
+ Defaults to None.
+ alpha (int): Hyper-parameters related to alignment_metrics.
+ Defaults to 1.
+ beta (int): Hyper-parameters related to alignment_metrics.
+ Defaults to 6.
+
+ Returns:
+ :obj:`TaskAlignedAssignResult`: The assign result.
+ """
+ priors = pred_instances.priors
+ decode_bboxes = pred_instances.bboxes
+ pred_scores = pred_instances.scores
+ gt_bboxes = gt_instances.bboxes
+ gt_labels = gt_instances.labels
+
+ priors = priors[:, :4]
+ num_gt, num_bboxes = gt_bboxes.size(0), priors.size(0)
+ # compute alignment metric between all bbox and gt
+ overlaps = self.iou_calculator(decode_bboxes, gt_bboxes).detach()
+ bbox_scores = pred_scores[:, gt_labels].detach()
+ # assign 0 by default
+ assigned_gt_inds = priors.new_full((num_bboxes, ), 0, dtype=torch.long)
+ assign_metrics = priors.new_zeros((num_bboxes, ))
+
+ if num_gt == 0 or num_bboxes == 0:
+ # No ground truth or boxes, return empty assignment
+ max_overlaps = priors.new_zeros((num_bboxes, ))
+ if num_gt == 0:
+ # No gt boxes, assign everything to background
+ assigned_gt_inds[:] = 0
+ assigned_labels = priors.new_full((num_bboxes, ),
+ -1,
+ dtype=torch.long)
+ assign_result = AssignResult(
+ num_gt, assigned_gt_inds, max_overlaps, labels=assigned_labels)
+ assign_result.assign_metrics = assign_metrics
+ return assign_result
+
+ # select top-k bboxes as candidates for each gt
+ alignment_metrics = bbox_scores**alpha * overlaps**beta
+ topk = min(self.topk, alignment_metrics.size(0))
+ _, candidate_idxs = alignment_metrics.topk(topk, dim=0, largest=True)
+ candidate_metrics = alignment_metrics[candidate_idxs,
+ torch.arange(num_gt)]
+ is_pos = candidate_metrics > 0
+
+ # limit the positive sample's center in gt
+ priors_cx = (priors[:, 0] + priors[:, 2]) / 2.0
+ priors_cy = (priors[:, 1] + priors[:, 3]) / 2.0
+ for gt_idx in range(num_gt):
+ candidate_idxs[:, gt_idx] += gt_idx * num_bboxes
+ ep_priors_cx = priors_cx.view(1, -1).expand(
+ num_gt, num_bboxes).contiguous().view(-1)
+ ep_priors_cy = priors_cy.view(1, -1).expand(
+ num_gt, num_bboxes).contiguous().view(-1)
+ candidate_idxs = candidate_idxs.view(-1)
+
+ # calculate the left, top, right, bottom distance between positive
+ # bbox center and gt side
+ l_ = ep_priors_cx[candidate_idxs].view(-1, num_gt) - gt_bboxes[:, 0]
+ t_ = ep_priors_cy[candidate_idxs].view(-1, num_gt) - gt_bboxes[:, 1]
+ r_ = gt_bboxes[:, 2] - ep_priors_cx[candidate_idxs].view(-1, num_gt)
+ b_ = gt_bboxes[:, 3] - ep_priors_cy[candidate_idxs].view(-1, num_gt)
+ is_in_gts = torch.stack([l_, t_, r_, b_], dim=1).min(dim=1)[0] > 0.01
+ is_pos = is_pos & is_in_gts
+
+ # if an anchor box is assigned to multiple gts,
+ # the one with the highest iou will be selected.
+ overlaps_inf = torch.full_like(overlaps,
+ -INF).t().contiguous().view(-1)
+ index = candidate_idxs.view(-1)[is_pos.view(-1)]
+ overlaps_inf[index] = overlaps.t().contiguous().view(-1)[index]
+ overlaps_inf = overlaps_inf.view(num_gt, -1).t()
+
+ max_overlaps, argmax_overlaps = overlaps_inf.max(dim=1)
+ assigned_gt_inds[
+ max_overlaps != -INF] = argmax_overlaps[max_overlaps != -INF] + 1
+ assign_metrics[max_overlaps != -INF] = alignment_metrics[
+ max_overlaps != -INF, argmax_overlaps[max_overlaps != -INF]]
+
+ assigned_labels = assigned_gt_inds.new_full((num_bboxes, ), -1)
+ pos_inds = torch.nonzero(
+ assigned_gt_inds > 0, as_tuple=False).squeeze()
+ if pos_inds.numel() > 0:
+ assigned_labels[pos_inds] = gt_labels[assigned_gt_inds[pos_inds] -
+ 1]
+ assign_result = AssignResult(
+ num_gt, assigned_gt_inds, max_overlaps, labels=assigned_labels)
+ assign_result.assign_metrics = assign_metrics
+ return assign_result
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/topk_hungarian_assigner.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/topk_hungarian_assigner.py
new file mode 100644
index 0000000000000000000000000000000000000000..e48f092ac1ae99eadfdf7502b591b57c782e6354
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/topk_hungarian_assigner.py
@@ -0,0 +1,182 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+from mmengine.structures import BaseDataElement
+from scipy.optimize import linear_sum_assignment
+
+from mmdet.registry import TASK_UTILS
+from .assign_result import AssignResult
+from .task_aligned_assigner import TaskAlignedAssigner
+
+
+@TASK_UTILS.register_module()
+class TopkHungarianAssigner(TaskAlignedAssigner):
+ """Computes 1-to-k matching between ground truth and predictions.
+
+ This class computes an assignment between the targets and the predictions
+ based on the costs. The costs are weighted sum of some components.
+ For DETR the costs are weighted sum of classification cost, regression L1
+ cost and regression iou cost. The targets don't include the no_object, so
+ generally there are more predictions than targets. After the 1-to-k
+ gt-pred matching, the un-matched are treated as backgrounds. Thus each
+ query prediction will be assigned with `0` or a positive integer
+ indicating the ground truth index:
+
+ - 0: negative sample, no assigned gt
+ - positive integer: positive sample, index (1-based) of assigned gt
+
+ Args:
+ cls_cost (dict): Classification cost configuration.
+ reg_cost (dict): Regression L1 cost configuration.
+ iou_cost (dict): Regression iou cost configuration.
+ """
+
+ def __init__(self,
+ *args,
+ cls_cost=dict(type='FocalLossCost', weight=2.0),
+ reg_cost=dict(type='BBoxL1Cost', weight=5.0),
+ iou_cost=dict(type='IoUCost', iou_mode='giou', weight=2.0),
+ **kwargs):
+ super(TopkHungarianAssigner, self).__init__(*args, **kwargs)
+
+ self.cls_cost = TASK_UTILS.build(cls_cost)
+ self.reg_cost = TASK_UTILS.build(reg_cost)
+ self.iou_cost = TASK_UTILS.build(iou_cost)
+
+ def assign(self,
+ pred_scores,
+ decode_bboxes,
+ gt_bboxes,
+ gt_labels,
+ img_meta,
+ alpha=1,
+ beta=6,
+ **kwargs):
+ """Computes 1-to-k gt-pred matching based on the weighted costs.
+
+ This method assign each query prediction to a ground truth or
+ background. The `assigned_gt_inds` with -1 means don't care,
+ 0 means negative sample, and positive number is the index (1-based)
+ of assigned gt.
+ The assignment is done in the following steps, the order matters.
+
+ 1. Assign every prediction to -1.
+ 2. Compute the weighted costs, each cost has shape (num_pred, num_gt).
+ 3. Update topk to be min(topk, int(num_pred / num_gt)), then repeat
+ costs topk times to shape: (num_pred, num_gt * topk), so that each
+ gt will match topk predictions.
+ 3. Do Hungarian matching on CPU based on the costs.
+ 4. Assign all to 0 (background) first, then for each matched pair
+ between predictions and gts, treat this prediction as foreground
+ and assign the corresponding gt index (plus 1) to it.
+ 5. Calculate alignment metrics and overlaps of each matched pred-gt
+ pair.
+
+ Args:
+ pred_scores (Tensor): Predicted normalized classification
+ scores for one image, has shape (num_dense_queries,
+ cls_out_channels).
+ decode_bboxes (Tensor): Predicted unnormalized bbox coordinates
+ for one image, has shape (num_dense_queries, 4) with the
+ last dimension arranged as (x1, y1, x2, y2).
+ gt_bboxes (Tensor): Unnormalized ground truth
+ bboxes for one image, has shape (num_gt, 4) with the
+ last dimension arranged as (x1, y1, x2, y2).
+ NOTE: num_gt is dynamic for each image.
+ gt_labels (Tensor): Ground truth classification
+ index for the image, has shape (num_gt,).
+ NOTE: num_gt is dynamic for each image.
+ img_meta (dict): Meta information for one image.
+ alpha (int): Hyper-parameters related to alignment_metrics.
+ Defaults to 1.
+ beta (int): Hyper-parameters related to alignment_metrics.
+ Defaults to 6.
+
+ Returns:
+ :obj:`AssignResult`: The assigned result.
+ """
+ pred_scores = pred_scores.detach()
+ decode_bboxes = decode_bboxes.detach()
+ temp_overlaps = self.iou_calculator(decode_bboxes, gt_bboxes).detach()
+ bbox_scores = pred_scores[:, gt_labels].detach()
+ alignment_metrics = bbox_scores**alpha * temp_overlaps**beta
+
+ pred_instances = BaseDataElement()
+ gt_instances = BaseDataElement()
+
+ pred_instances.bboxes = decode_bboxes
+ gt_instances.bboxes = gt_bboxes
+
+ pred_instances.scores = pred_scores
+ gt_instances.labels = gt_labels
+
+ reg_cost = self.reg_cost(pred_instances, gt_instances, img_meta)
+ iou_cost = self.iou_cost(pred_instances, gt_instances, img_meta)
+ cls_cost = self.cls_cost(pred_instances, gt_instances, img_meta)
+ all_cost = cls_cost + reg_cost + iou_cost
+
+ num_gt, num_bboxes = gt_bboxes.size(0), pred_scores.size(0)
+ if num_gt > 0:
+ # assign 0 by default
+ assigned_gt_inds = pred_scores.new_full((num_bboxes, ),
+ 0,
+ dtype=torch.long)
+ select_cost = all_cost
+
+ topk = min(self.topk, int(len(select_cost) / num_gt))
+
+ # Repeat the ground truth `topk` times to perform 1-to-k gt-pred
+ # matching. For example, if `num_pred` = 900, `num_gt` = 3, then
+ # there are only 3 gt-pred pairs in sum for 1-1 matching.
+ # However, for 1-k gt-pred matching, if `topk` = 4, then each
+ # gt is assigned 4 unique predictions, so there would be 12
+ # gt-pred pairs in sum.
+ repeat_select_cost = select_cost[...,
+ None].repeat(1, 1, topk).view(
+ select_cost.size(0), -1)
+ # anchor index and gt index
+ matched_row_inds, matched_col_inds = linear_sum_assignment(
+ repeat_select_cost.detach().cpu().numpy())
+ matched_row_inds = torch.from_numpy(matched_row_inds).to(
+ pred_scores.device)
+ matched_col_inds = torch.from_numpy(matched_col_inds).to(
+ pred_scores.device)
+
+ match_gt_ids = matched_col_inds // topk
+ candidate_idxs = matched_row_inds
+
+ assigned_labels = assigned_gt_inds.new_full((num_bboxes, ), -1)
+
+ if candidate_idxs.numel() > 0:
+ assigned_labels[candidate_idxs] = gt_labels[match_gt_ids]
+ else:
+ assigned_labels = None
+
+ assigned_gt_inds[candidate_idxs] = match_gt_ids + 1
+
+ overlaps = self.iou_calculator(
+ decode_bboxes[candidate_idxs],
+ gt_bboxes[match_gt_ids],
+ is_aligned=True).detach()
+
+ temp_pos_alignment_metrics = alignment_metrics[candidate_idxs]
+ pos_alignment_metrics = torch.gather(temp_pos_alignment_metrics, 1,
+ match_gt_ids[:,
+ None]).view(-1)
+ assign_result = AssignResult(
+ num_gt, assigned_gt_inds, overlaps, labels=assigned_labels)
+
+ assign_result.assign_metrics = pos_alignment_metrics
+ return assign_result
+ else:
+
+ assigned_gt_inds = pred_scores.new_full((num_bboxes, ),
+ -1,
+ dtype=torch.long)
+
+ assigned_labels = pred_scores.new_full((num_bboxes, ),
+ -1,
+ dtype=torch.long)
+
+ assigned_gt_inds[:] = 0
+ return AssignResult(
+ 0, assigned_gt_inds, None, labels=assigned_labels)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/uniform_assigner.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/uniform_assigner.py
new file mode 100644
index 0000000000000000000000000000000000000000..9a83bfd0b46a3690dce9cf0adf2c1e676f304d06
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/assigners/uniform_assigner.py
@@ -0,0 +1,173 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional
+
+import torch
+from mmengine.structures import InstanceData
+
+from mmdet.registry import TASK_UTILS
+from mmdet.structures.bbox import bbox_xyxy_to_cxcywh
+from mmdet.utils import ConfigType
+from .assign_result import AssignResult
+from .base_assigner import BaseAssigner
+
+
+@TASK_UTILS.register_module()
+class UniformAssigner(BaseAssigner):
+ """Uniform Matching between the priors and gt boxes, which can achieve
+ balance in positive priors, and gt_bboxes_ignore was not considered for
+ now.
+
+ Args:
+ pos_ignore_thr (float): the threshold to ignore positive priors
+ neg_ignore_thr (float): the threshold to ignore negative priors
+ match_times(int): Number of positive priors for each gt box.
+ Defaults to 4.
+ iou_calculator (:obj:`ConfigDict` or dict): Config dict for iou
+ calculator. Defaults to ``dict(type='BboxOverlaps2D')``
+ """
+
+ def __init__(self,
+ pos_ignore_thr: float,
+ neg_ignore_thr: float,
+ match_times: int = 4,
+ iou_calculator: ConfigType = dict(type='BboxOverlaps2D')):
+ self.match_times = match_times
+ self.pos_ignore_thr = pos_ignore_thr
+ self.neg_ignore_thr = neg_ignore_thr
+ self.iou_calculator = TASK_UTILS.build(iou_calculator)
+
+ def assign(
+ self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ gt_instances_ignore: Optional[InstanceData] = None
+ ) -> AssignResult:
+ """Assign gt to priors.
+
+ The assignment is done in following steps
+
+ 1. assign -1 by default
+ 2. compute the L1 cost between boxes. Note that we use priors and
+ predict boxes both
+ 3. compute the ignore indexes use gt_bboxes and predict boxes
+ 4. compute the ignore indexes of positive sample use priors and
+ predict boxes
+
+
+ Args:
+ pred_instances (:obj:`InstaceData`): Instances of model
+ predictions. It includes ``priors``, and the priors can
+ be priors, points, or bboxes predicted by the model,
+ shape(n, 4).
+ gt_instances (:obj:`InstaceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ gt_instances_ignore (:obj:`InstaceData`, optional): Instances
+ to be ignored during training. It includes ``bboxes``
+ attribute data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ :obj:`AssignResult`: The assign result.
+ """
+
+ gt_bboxes = gt_instances.bboxes
+ gt_labels = gt_instances.labels
+ priors = pred_instances.priors
+ bbox_pred = pred_instances.decoder_priors
+
+ num_gts, num_bboxes = gt_bboxes.size(0), bbox_pred.size(0)
+
+ # 1. assign -1 by default
+ assigned_gt_inds = bbox_pred.new_full((num_bboxes, ),
+ 0,
+ dtype=torch.long)
+ assigned_labels = bbox_pred.new_full((num_bboxes, ),
+ -1,
+ dtype=torch.long)
+ if num_gts == 0 or num_bboxes == 0:
+ # No ground truth or boxes, return empty assignment
+ if num_gts == 0:
+ # No ground truth, assign all to background
+ assigned_gt_inds[:] = 0
+ assign_result = AssignResult(
+ num_gts, assigned_gt_inds, None, labels=assigned_labels)
+ assign_result.set_extra_property(
+ 'pos_idx', bbox_pred.new_empty(0, dtype=torch.bool))
+ assign_result.set_extra_property('pos_predicted_boxes',
+ bbox_pred.new_empty((0, 4)))
+ assign_result.set_extra_property('target_boxes',
+ bbox_pred.new_empty((0, 4)))
+ return assign_result
+
+ # 2. Compute the L1 cost between boxes
+ # Note that we use priors and predict boxes both
+ cost_bbox = torch.cdist(
+ bbox_xyxy_to_cxcywh(bbox_pred),
+ bbox_xyxy_to_cxcywh(gt_bboxes),
+ p=1)
+ cost_bbox_priors = torch.cdist(
+ bbox_xyxy_to_cxcywh(priors), bbox_xyxy_to_cxcywh(gt_bboxes), p=1)
+
+ # We found that topk function has different results in cpu and
+ # cuda mode. In order to ensure consistency with the source code,
+ # we also use cpu mode.
+ # TODO: Check whether the performance of cpu and cuda are the same.
+ C = cost_bbox.cpu()
+ C1 = cost_bbox_priors.cpu()
+
+ # self.match_times x n
+ index = torch.topk(
+ C, # c=b,n,x c[i]=n,x
+ k=self.match_times,
+ dim=0,
+ largest=False)[1]
+
+ # self.match_times x n
+ index1 = torch.topk(C1, k=self.match_times, dim=0, largest=False)[1]
+ # (self.match_times*2) x n
+ indexes = torch.cat((index, index1),
+ dim=1).reshape(-1).to(bbox_pred.device)
+
+ pred_overlaps = self.iou_calculator(bbox_pred, gt_bboxes)
+ anchor_overlaps = self.iou_calculator(priors, gt_bboxes)
+ pred_max_overlaps, _ = pred_overlaps.max(dim=1)
+ anchor_max_overlaps, _ = anchor_overlaps.max(dim=0)
+
+ # 3. Compute the ignore indexes use gt_bboxes and predict boxes
+ ignore_idx = pred_max_overlaps > self.neg_ignore_thr
+ assigned_gt_inds[ignore_idx] = -1
+
+ # 4. Compute the ignore indexes of positive sample use priors
+ # and predict boxes
+ pos_gt_index = torch.arange(
+ 0, C1.size(1),
+ device=bbox_pred.device).repeat(self.match_times * 2)
+ pos_ious = anchor_overlaps[indexes, pos_gt_index]
+ pos_ignore_idx = pos_ious < self.pos_ignore_thr
+
+ pos_gt_index_with_ignore = pos_gt_index + 1
+ pos_gt_index_with_ignore[pos_ignore_idx] = -1
+ assigned_gt_inds[indexes] = pos_gt_index_with_ignore
+
+ if gt_labels is not None:
+ assigned_labels = assigned_gt_inds.new_full((num_bboxes, ), -1)
+ pos_inds = torch.nonzero(
+ assigned_gt_inds > 0, as_tuple=False).squeeze()
+ if pos_inds.numel() > 0:
+ assigned_labels[pos_inds] = gt_labels[
+ assigned_gt_inds[pos_inds] - 1]
+ else:
+ assigned_labels = None
+
+ assign_result = AssignResult(
+ num_gts,
+ assigned_gt_inds,
+ anchor_max_overlaps,
+ labels=assigned_labels)
+ assign_result.set_extra_property('pos_idx', ~pos_ignore_idx)
+ assign_result.set_extra_property('pos_predicted_boxes',
+ bbox_pred[indexes])
+ assign_result.set_extra_property('target_boxes',
+ gt_bboxes[pos_gt_index])
+ return assign_result
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/builder.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/builder.py
new file mode 100644
index 0000000000000000000000000000000000000000..6736049fef688e0d663d6195c79ec9688dc4c5d7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/builder.py
@@ -0,0 +1,62 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+
+from mmdet.registry import TASK_UTILS
+
+PRIOR_GENERATORS = TASK_UTILS
+ANCHOR_GENERATORS = TASK_UTILS
+BBOX_ASSIGNERS = TASK_UTILS
+BBOX_SAMPLERS = TASK_UTILS
+BBOX_CODERS = TASK_UTILS
+MATCH_COSTS = TASK_UTILS
+IOU_CALCULATORS = TASK_UTILS
+
+
+def build_bbox_coder(cfg, **default_args):
+ """Builder of box coder."""
+ warnings.warn('``build_sampler`` would be deprecated soon, please use '
+ '``mmdet.registry.TASK_UTILS.build()`` ')
+ return TASK_UTILS.build(cfg, default_args=default_args)
+
+
+def build_iou_calculator(cfg, default_args=None):
+ """Builder of IoU calculator."""
+ warnings.warn(
+ '``build_iou_calculator`` would be deprecated soon, please use '
+ '``mmdet.registry.TASK_UTILS.build()`` ')
+ return TASK_UTILS.build(cfg, default_args=default_args)
+
+
+def build_match_cost(cfg, default_args=None):
+ """Builder of IoU calculator."""
+ warnings.warn('``build_match_cost`` would be deprecated soon, please use '
+ '``mmdet.registry.TASK_UTILS.build()`` ')
+ return TASK_UTILS.build(cfg, default_args=default_args)
+
+
+def build_assigner(cfg, **default_args):
+ """Builder of box assigner."""
+ warnings.warn('``build_assigner`` would be deprecated soon, please use '
+ '``mmdet.registry.TASK_UTILS.build()`` ')
+ return TASK_UTILS.build(cfg, default_args=default_args)
+
+
+def build_sampler(cfg, **default_args):
+ """Builder of box sampler."""
+ warnings.warn('``build_sampler`` would be deprecated soon, please use '
+ '``mmdet.registry.TASK_UTILS.build()`` ')
+ return TASK_UTILS.build(cfg, default_args=default_args)
+
+
+def build_prior_generator(cfg, default_args=None):
+ warnings.warn(
+ '``build_prior_generator`` would be deprecated soon, please use '
+ '``mmdet.registry.TASK_UTILS.build()`` ')
+ return TASK_UTILS.build(cfg, default_args=default_args)
+
+
+def build_anchor_generator(cfg, default_args=None):
+ warnings.warn(
+ '``build_anchor_generator`` would be deprecated soon, please use '
+ '``mmdet.registry.TASK_UTILS.build()`` ')
+ return TASK_UTILS.build(cfg, default_args=default_args)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..97c3982140021958dabdd03f8040519f946250ff
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/__init__.py
@@ -0,0 +1,16 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .base_bbox_coder import BaseBBoxCoder
+from .bucketing_bbox_coder import BucketingBBoxCoder
+from .delta_xywh_bbox_coder import (DeltaXYWHBBoxCoder,
+ DeltaXYWHBBoxCoderForGLIP)
+from .distance_point_bbox_coder import DistancePointBBoxCoder
+from .legacy_delta_xywh_bbox_coder import LegacyDeltaXYWHBBoxCoder
+from .pseudo_bbox_coder import PseudoBBoxCoder
+from .tblr_bbox_coder import TBLRBBoxCoder
+from .yolo_bbox_coder import YOLOBBoxCoder
+
+__all__ = [
+ 'BaseBBoxCoder', 'PseudoBBoxCoder', 'DeltaXYWHBBoxCoder',
+ 'LegacyDeltaXYWHBBoxCoder', 'TBLRBBoxCoder', 'YOLOBBoxCoder',
+ 'BucketingBBoxCoder', 'DistancePointBBoxCoder', 'DeltaXYWHBBoxCoderForGLIP'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/base_bbox_coder.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/base_bbox_coder.py
new file mode 100644
index 0000000000000000000000000000000000000000..806d2651869e02173578c9eb331758743a068dd9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/base_bbox_coder.py
@@ -0,0 +1,26 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from abc import ABCMeta, abstractmethod
+
+
+class BaseBBoxCoder(metaclass=ABCMeta):
+ """Base bounding box coder.
+
+ Args:
+ use_box_type (bool): Whether to warp decoded boxes with the
+ box type data structure. Defaults to False.
+ """
+
+ # The size of the last of dimension of the encoded tensor.
+ encode_size = 4
+
+ def __init__(self, use_box_type: bool = False, **kwargs):
+ self.use_box_type = use_box_type
+
+ @abstractmethod
+ def encode(self, bboxes, gt_bboxes):
+ """Encode deltas between bboxes and ground truth boxes."""
+
+ @abstractmethod
+ def decode(self, bboxes, bboxes_pred):
+ """Decode the predicted bboxes according to prediction and base
+ boxes."""
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/bucketing_bbox_coder.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/bucketing_bbox_coder.py
new file mode 100644
index 0000000000000000000000000000000000000000..4044e1cd91d619521606f3c03032a40a9fc27130
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/bucketing_bbox_coder.py
@@ -0,0 +1,366 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Sequence, Tuple, Union
+
+import numpy as np
+import torch
+import torch.nn.functional as F
+from torch import Tensor
+
+from mmdet.registry import TASK_UTILS
+from mmdet.structures.bbox import (BaseBoxes, HorizontalBoxes, bbox_rescale,
+ get_box_tensor)
+from .base_bbox_coder import BaseBBoxCoder
+
+
+@TASK_UTILS.register_module()
+class BucketingBBoxCoder(BaseBBoxCoder):
+ """Bucketing BBox Coder for Side-Aware Boundary Localization (SABL).
+
+ Boundary Localization with Bucketing and Bucketing Guided Rescoring
+ are implemented here.
+
+ Please refer to https://arxiv.org/abs/1912.04260 for more details.
+
+ Args:
+ num_buckets (int): Number of buckets.
+ scale_factor (int): Scale factor of proposals to generate buckets.
+ offset_topk (int): Topk buckets are used to generate
+ bucket fine regression targets. Defaults to 2.
+ offset_upperbound (float): Offset upperbound to generate
+ bucket fine regression targets.
+ To avoid too large offset displacements. Defaults to 1.0.
+ cls_ignore_neighbor (bool): Ignore second nearest bucket or Not.
+ Defaults to True.
+ clip_border (bool, optional): Whether clip the objects outside the
+ border of the image. Defaults to True.
+ """
+
+ def __init__(self,
+ num_buckets: int,
+ scale_factor: int,
+ offset_topk: int = 2,
+ offset_upperbound: float = 1.0,
+ cls_ignore_neighbor: bool = True,
+ clip_border: bool = True,
+ **kwargs) -> None:
+ super().__init__(**kwargs)
+ self.num_buckets = num_buckets
+ self.scale_factor = scale_factor
+ self.offset_topk = offset_topk
+ self.offset_upperbound = offset_upperbound
+ self.cls_ignore_neighbor = cls_ignore_neighbor
+ self.clip_border = clip_border
+
+ def encode(self, bboxes: Union[Tensor, BaseBoxes],
+ gt_bboxes: Union[Tensor, BaseBoxes]) -> Tuple[Tensor]:
+ """Get bucketing estimation and fine regression targets during
+ training.
+
+ Args:
+ bboxes (torch.Tensor or :obj:`BaseBoxes`): source boxes,
+ e.g., object proposals.
+ gt_bboxes (torch.Tensor or :obj:`BaseBoxes`): target of the
+ transformation, e.g., ground truth boxes.
+
+ Returns:
+ encoded_bboxes(tuple[Tensor]): bucketing estimation
+ and fine regression targets and weights
+ """
+ bboxes = get_box_tensor(bboxes)
+ gt_bboxes = get_box_tensor(gt_bboxes)
+ assert bboxes.size(0) == gt_bboxes.size(0)
+ assert bboxes.size(-1) == gt_bboxes.size(-1) == 4
+ encoded_bboxes = bbox2bucket(bboxes, gt_bboxes, self.num_buckets,
+ self.scale_factor, self.offset_topk,
+ self.offset_upperbound,
+ self.cls_ignore_neighbor)
+ return encoded_bboxes
+
+ def decode(
+ self,
+ bboxes: Union[Tensor, BaseBoxes],
+ pred_bboxes: Tensor,
+ max_shape: Optional[Tuple[int]] = None
+ ) -> Tuple[Union[Tensor, BaseBoxes], Tensor]:
+ """Apply transformation `pred_bboxes` to `boxes`.
+ Args:
+ boxes (torch.Tensor or :obj:`BaseBoxes`): Basic boxes.
+ pred_bboxes (torch.Tensor): Predictions for bucketing estimation
+ and fine regression
+ max_shape (tuple[int], optional): Maximum shape of boxes.
+ Defaults to None.
+
+ Returns:
+ Union[torch.Tensor, :obj:`BaseBoxes`]: Decoded boxes.
+ """
+ bboxes = get_box_tensor(bboxes)
+ assert len(pred_bboxes) == 2
+ cls_preds, offset_preds = pred_bboxes
+ assert cls_preds.size(0) == bboxes.size(0) and offset_preds.size(
+ 0) == bboxes.size(0)
+ bboxes, loc_confidence = bucket2bbox(bboxes, cls_preds, offset_preds,
+ self.num_buckets,
+ self.scale_factor, max_shape,
+ self.clip_border)
+ if self.use_box_type:
+ bboxes = HorizontalBoxes(bboxes, clone=False)
+ return bboxes, loc_confidence
+
+
+def generat_buckets(proposals: Tensor,
+ num_buckets: int,
+ scale_factor: float = 1.0) -> Tuple[Tensor]:
+ """Generate buckets w.r.t bucket number and scale factor of proposals.
+
+ Args:
+ proposals (Tensor): Shape (n, 4)
+ num_buckets (int): Number of buckets.
+ scale_factor (float): Scale factor to rescale proposals.
+
+ Returns:
+ tuple[Tensor]: (bucket_w, bucket_h, l_buckets, r_buckets,
+ t_buckets, d_buckets)
+
+ - bucket_w: Width of buckets on x-axis. Shape (n, ).
+ - bucket_h: Height of buckets on y-axis. Shape (n, ).
+ - l_buckets: Left buckets. Shape (n, ceil(side_num/2)).
+ - r_buckets: Right buckets. Shape (n, ceil(side_num/2)).
+ - t_buckets: Top buckets. Shape (n, ceil(side_num/2)).
+ - d_buckets: Down buckets. Shape (n, ceil(side_num/2)).
+ """
+ proposals = bbox_rescale(proposals, scale_factor)
+
+ # number of buckets in each side
+ side_num = int(np.ceil(num_buckets / 2.0))
+ pw = proposals[..., 2] - proposals[..., 0]
+ ph = proposals[..., 3] - proposals[..., 1]
+ px1 = proposals[..., 0]
+ py1 = proposals[..., 1]
+ px2 = proposals[..., 2]
+ py2 = proposals[..., 3]
+
+ bucket_w = pw / num_buckets
+ bucket_h = ph / num_buckets
+
+ # left buckets
+ l_buckets = px1[:, None] + (0.5 + torch.arange(
+ 0, side_num).to(proposals).float())[None, :] * bucket_w[:, None]
+ # right buckets
+ r_buckets = px2[:, None] - (0.5 + torch.arange(
+ 0, side_num).to(proposals).float())[None, :] * bucket_w[:, None]
+ # top buckets
+ t_buckets = py1[:, None] + (0.5 + torch.arange(
+ 0, side_num).to(proposals).float())[None, :] * bucket_h[:, None]
+ # down buckets
+ d_buckets = py2[:, None] - (0.5 + torch.arange(
+ 0, side_num).to(proposals).float())[None, :] * bucket_h[:, None]
+ return bucket_w, bucket_h, l_buckets, r_buckets, t_buckets, d_buckets
+
+
+def bbox2bucket(proposals: Tensor,
+ gt: Tensor,
+ num_buckets: int,
+ scale_factor: float,
+ offset_topk: int = 2,
+ offset_upperbound: float = 1.0,
+ cls_ignore_neighbor: bool = True) -> Tuple[Tensor]:
+ """Generate buckets estimation and fine regression targets.
+
+ Args:
+ proposals (Tensor): Shape (n, 4)
+ gt (Tensor): Shape (n, 4)
+ num_buckets (int): Number of buckets.
+ scale_factor (float): Scale factor to rescale proposals.
+ offset_topk (int): Topk buckets are used to generate
+ bucket fine regression targets. Defaults to 2.
+ offset_upperbound (float): Offset allowance to generate
+ bucket fine regression targets.
+ To avoid too large offset displacements. Defaults to 1.0.
+ cls_ignore_neighbor (bool): Ignore second nearest bucket or Not.
+ Defaults to True.
+
+ Returns:
+ tuple[Tensor]: (offsets, offsets_weights, bucket_labels, cls_weights).
+
+ - offsets: Fine regression targets. \
+ Shape (n, num_buckets*2).
+ - offsets_weights: Fine regression weights. \
+ Shape (n, num_buckets*2).
+ - bucket_labels: Bucketing estimation labels. \
+ Shape (n, num_buckets*2).
+ - cls_weights: Bucketing estimation weights. \
+ Shape (n, num_buckets*2).
+ """
+ assert proposals.size() == gt.size()
+
+ # generate buckets
+ proposals = proposals.float()
+ gt = gt.float()
+ (bucket_w, bucket_h, l_buckets, r_buckets, t_buckets,
+ d_buckets) = generat_buckets(proposals, num_buckets, scale_factor)
+
+ gx1 = gt[..., 0]
+ gy1 = gt[..., 1]
+ gx2 = gt[..., 2]
+ gy2 = gt[..., 3]
+
+ # generate offset targets and weights
+ # offsets from buckets to gts
+ l_offsets = (l_buckets - gx1[:, None]) / bucket_w[:, None]
+ r_offsets = (r_buckets - gx2[:, None]) / bucket_w[:, None]
+ t_offsets = (t_buckets - gy1[:, None]) / bucket_h[:, None]
+ d_offsets = (d_buckets - gy2[:, None]) / bucket_h[:, None]
+
+ # select top-k nearest buckets
+ l_topk, l_label = l_offsets.abs().topk(
+ offset_topk, dim=1, largest=False, sorted=True)
+ r_topk, r_label = r_offsets.abs().topk(
+ offset_topk, dim=1, largest=False, sorted=True)
+ t_topk, t_label = t_offsets.abs().topk(
+ offset_topk, dim=1, largest=False, sorted=True)
+ d_topk, d_label = d_offsets.abs().topk(
+ offset_topk, dim=1, largest=False, sorted=True)
+
+ offset_l_weights = l_offsets.new_zeros(l_offsets.size())
+ offset_r_weights = r_offsets.new_zeros(r_offsets.size())
+ offset_t_weights = t_offsets.new_zeros(t_offsets.size())
+ offset_d_weights = d_offsets.new_zeros(d_offsets.size())
+ inds = torch.arange(0, proposals.size(0)).to(proposals).long()
+
+ # generate offset weights of top-k nearest buckets
+ for k in range(offset_topk):
+ if k >= 1:
+ offset_l_weights[inds, l_label[:,
+ k]] = (l_topk[:, k] <
+ offset_upperbound).float()
+ offset_r_weights[inds, r_label[:,
+ k]] = (r_topk[:, k] <
+ offset_upperbound).float()
+ offset_t_weights[inds, t_label[:,
+ k]] = (t_topk[:, k] <
+ offset_upperbound).float()
+ offset_d_weights[inds, d_label[:,
+ k]] = (d_topk[:, k] <
+ offset_upperbound).float()
+ else:
+ offset_l_weights[inds, l_label[:, k]] = 1.0
+ offset_r_weights[inds, r_label[:, k]] = 1.0
+ offset_t_weights[inds, t_label[:, k]] = 1.0
+ offset_d_weights[inds, d_label[:, k]] = 1.0
+
+ offsets = torch.cat([l_offsets, r_offsets, t_offsets, d_offsets], dim=-1)
+ offsets_weights = torch.cat([
+ offset_l_weights, offset_r_weights, offset_t_weights, offset_d_weights
+ ],
+ dim=-1)
+
+ # generate bucket labels and weight
+ side_num = int(np.ceil(num_buckets / 2.0))
+ labels = torch.stack(
+ [l_label[:, 0], r_label[:, 0], t_label[:, 0], d_label[:, 0]], dim=-1)
+
+ batch_size = labels.size(0)
+ bucket_labels = F.one_hot(labels.view(-1), side_num).view(batch_size,
+ -1).float()
+ bucket_cls_l_weights = (l_offsets.abs() < 1).float()
+ bucket_cls_r_weights = (r_offsets.abs() < 1).float()
+ bucket_cls_t_weights = (t_offsets.abs() < 1).float()
+ bucket_cls_d_weights = (d_offsets.abs() < 1).float()
+ bucket_cls_weights = torch.cat([
+ bucket_cls_l_weights, bucket_cls_r_weights, bucket_cls_t_weights,
+ bucket_cls_d_weights
+ ],
+ dim=-1)
+ # ignore second nearest buckets for cls if necessary
+ if cls_ignore_neighbor:
+ bucket_cls_weights = (~((bucket_cls_weights == 1) &
+ (bucket_labels == 0))).float()
+ else:
+ bucket_cls_weights[:] = 1.0
+ return offsets, offsets_weights, bucket_labels, bucket_cls_weights
+
+
+def bucket2bbox(proposals: Tensor,
+ cls_preds: Tensor,
+ offset_preds: Tensor,
+ num_buckets: int,
+ scale_factor: float = 1.0,
+ max_shape: Optional[Union[Sequence[int], Tensor,
+ Sequence[Sequence[int]]]] = None,
+ clip_border: bool = True) -> Tuple[Tensor]:
+ """Apply bucketing estimation (cls preds) and fine regression (offset
+ preds) to generate det bboxes.
+
+ Args:
+ proposals (Tensor): Boxes to be transformed. Shape (n, 4)
+ cls_preds (Tensor): bucketing estimation. Shape (n, num_buckets*2).
+ offset_preds (Tensor): fine regression. Shape (n, num_buckets*2).
+ num_buckets (int): Number of buckets.
+ scale_factor (float): Scale factor to rescale proposals.
+ max_shape (tuple[int, int]): Maximum bounds for boxes. specifies (H, W)
+ clip_border (bool, optional): Whether clip the objects outside the
+ border of the image. Defaults to True.
+
+ Returns:
+ tuple[Tensor]: (bboxes, loc_confidence).
+
+ - bboxes: predicted bboxes. Shape (n, 4)
+ - loc_confidence: localization confidence of predicted bboxes.
+ Shape (n,).
+ """
+
+ side_num = int(np.ceil(num_buckets / 2.0))
+ cls_preds = cls_preds.view(-1, side_num)
+ offset_preds = offset_preds.view(-1, side_num)
+
+ scores = F.softmax(cls_preds, dim=1)
+ score_topk, score_label = scores.topk(2, dim=1, largest=True, sorted=True)
+
+ rescaled_proposals = bbox_rescale(proposals, scale_factor)
+
+ pw = rescaled_proposals[..., 2] - rescaled_proposals[..., 0]
+ ph = rescaled_proposals[..., 3] - rescaled_proposals[..., 1]
+ px1 = rescaled_proposals[..., 0]
+ py1 = rescaled_proposals[..., 1]
+ px2 = rescaled_proposals[..., 2]
+ py2 = rescaled_proposals[..., 3]
+
+ bucket_w = pw / num_buckets
+ bucket_h = ph / num_buckets
+
+ score_inds_l = score_label[0::4, 0]
+ score_inds_r = score_label[1::4, 0]
+ score_inds_t = score_label[2::4, 0]
+ score_inds_d = score_label[3::4, 0]
+ l_buckets = px1 + (0.5 + score_inds_l.float()) * bucket_w
+ r_buckets = px2 - (0.5 + score_inds_r.float()) * bucket_w
+ t_buckets = py1 + (0.5 + score_inds_t.float()) * bucket_h
+ d_buckets = py2 - (0.5 + score_inds_d.float()) * bucket_h
+
+ offsets = offset_preds.view(-1, 4, side_num)
+ inds = torch.arange(proposals.size(0)).to(proposals).long()
+ l_offsets = offsets[:, 0, :][inds, score_inds_l]
+ r_offsets = offsets[:, 1, :][inds, score_inds_r]
+ t_offsets = offsets[:, 2, :][inds, score_inds_t]
+ d_offsets = offsets[:, 3, :][inds, score_inds_d]
+
+ x1 = l_buckets - l_offsets * bucket_w
+ x2 = r_buckets - r_offsets * bucket_w
+ y1 = t_buckets - t_offsets * bucket_h
+ y2 = d_buckets - d_offsets * bucket_h
+
+ if clip_border and max_shape is not None:
+ x1 = x1.clamp(min=0, max=max_shape[1] - 1)
+ y1 = y1.clamp(min=0, max=max_shape[0] - 1)
+ x2 = x2.clamp(min=0, max=max_shape[1] - 1)
+ y2 = y2.clamp(min=0, max=max_shape[0] - 1)
+ bboxes = torch.cat([x1[:, None], y1[:, None], x2[:, None], y2[:, None]],
+ dim=-1)
+
+ # bucketing guided rescoring
+ loc_confidence = score_topk[:, 0]
+ top2_neighbor_inds = (score_label[:, 0] - score_label[:, 1]).abs() == 1
+ loc_confidence += score_topk[:, 1] * top2_neighbor_inds.float()
+ loc_confidence = loc_confidence.view(-1, 4).mean(dim=1)
+
+ return bboxes, loc_confidence
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/delta_xywh_bbox_coder.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/delta_xywh_bbox_coder.py
new file mode 100644
index 0000000000000000000000000000000000000000..c2b60b5ee791e05ce4f5f8d8e1876f7f61e964ed
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/delta_xywh_bbox_coder.py
@@ -0,0 +1,579 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+from typing import Optional, Sequence, Union
+
+import numpy as np
+import torch
+from torch import Tensor
+
+from mmdet.registry import TASK_UTILS
+from mmdet.structures.bbox import BaseBoxes, HorizontalBoxes, get_box_tensor
+from .base_bbox_coder import BaseBBoxCoder
+
+
+@TASK_UTILS.register_module()
+class DeltaXYWHBBoxCoder(BaseBBoxCoder):
+ """Delta XYWH BBox coder.
+
+ Following the practice in `R-CNN `_,
+ this coder encodes bbox (x1, y1, x2, y2) into delta (dx, dy, dw, dh) and
+ decodes delta (dx, dy, dw, dh) back to original bbox (x1, y1, x2, y2).
+
+ Args:
+ target_means (Sequence[float]): Denormalizing means of target for
+ delta coordinates
+ target_stds (Sequence[float]): Denormalizing standard deviation of
+ target for delta coordinates
+ clip_border (bool, optional): Whether clip the objects outside the
+ border of the image. Defaults to True.
+ add_ctr_clamp (bool): Whether to add center clamp, when added, the
+ predicted box is clamped is its center is too far away from
+ the original anchor's center. Only used by YOLOF. Default False.
+ ctr_clamp (int): the maximum pixel shift to clamp. Only used by YOLOF.
+ Default 32.
+ """
+
+ def __init__(self,
+ target_means: Sequence[float] = (0., 0., 0., 0.),
+ target_stds: Sequence[float] = (1., 1., 1., 1.),
+ clip_border: bool = True,
+ add_ctr_clamp: bool = False,
+ ctr_clamp: int = 32,
+ **kwargs) -> None:
+ super().__init__(**kwargs)
+ self.means = target_means
+ self.stds = target_stds
+ self.clip_border = clip_border
+ self.add_ctr_clamp = add_ctr_clamp
+ self.ctr_clamp = ctr_clamp
+
+ def encode(self, bboxes: Union[Tensor, BaseBoxes],
+ gt_bboxes: Union[Tensor, BaseBoxes]) -> Tensor:
+ """Get box regression transformation deltas that can be used to
+ transform the ``bboxes`` into the ``gt_bboxes``.
+
+ Args:
+ bboxes (torch.Tensor or :obj:`BaseBoxes`): Source boxes,
+ e.g., object proposals.
+ gt_bboxes (torch.Tensor or :obj:`BaseBoxes`): Target of the
+ transformation, e.g., ground-truth boxes.
+
+ Returns:
+ torch.Tensor: Box transformation deltas
+ """
+ bboxes = get_box_tensor(bboxes)
+ gt_bboxes = get_box_tensor(gt_bboxes)
+ assert bboxes.size(0) == gt_bboxes.size(0)
+ assert bboxes.size(-1) == gt_bboxes.size(-1) == 4
+ encoded_bboxes = bbox2delta(bboxes, gt_bboxes, self.means, self.stds)
+ return encoded_bboxes
+
+ def decode(
+ self,
+ bboxes: Union[Tensor, BaseBoxes],
+ pred_bboxes: Tensor,
+ max_shape: Optional[Union[Sequence[int], Tensor,
+ Sequence[Sequence[int]]]] = None,
+ wh_ratio_clip: Optional[float] = 16 / 1000
+ ) -> Union[Tensor, BaseBoxes]:
+ """Apply transformation `pred_bboxes` to `boxes`.
+
+ Args:
+ bboxes (torch.Tensor or :obj:`BaseBoxes`): Basic boxes. Shape
+ (B, N, 4) or (N, 4)
+ pred_bboxes (Tensor): Encoded offsets with respect to each roi.
+ Has shape (B, N, num_classes * 4) or (B, N, 4) or
+ (N, num_classes * 4) or (N, 4). Note N = num_anchors * W * H
+ when rois is a grid of anchors.Offset encoding follows [1]_.
+ max_shape (Sequence[int] or torch.Tensor or Sequence[
+ Sequence[int]],optional): Maximum bounds for boxes, specifies
+ (H, W, C) or (H, W). If bboxes shape is (B, N, 4), then
+ the max_shape should be a Sequence[Sequence[int]]
+ and the length of max_shape should also be B.
+ wh_ratio_clip (float, optional): The allowed ratio between
+ width and height.
+
+ Returns:
+ Union[torch.Tensor, :obj:`BaseBoxes`]: Decoded boxes.
+ """
+ bboxes = get_box_tensor(bboxes)
+ assert pred_bboxes.size(0) == bboxes.size(0)
+ if pred_bboxes.ndim == 3:
+ assert pred_bboxes.size(1) == bboxes.size(1)
+
+ if pred_bboxes.ndim == 2 and not torch.onnx.is_in_onnx_export():
+ # single image decode
+ decoded_bboxes = delta2bbox(bboxes, pred_bboxes, self.means,
+ self.stds, max_shape, wh_ratio_clip,
+ self.clip_border, self.add_ctr_clamp,
+ self.ctr_clamp)
+ else:
+ if pred_bboxes.ndim == 3 and not torch.onnx.is_in_onnx_export():
+ warnings.warn(
+ 'DeprecationWarning: onnx_delta2bbox is deprecated '
+ 'in the case of batch decoding and non-ONNX, '
+ 'please use “delta2bbox” instead. In order to improve '
+ 'the decoding speed, the batch function will no '
+ 'longer be supported. ')
+ decoded_bboxes = onnx_delta2bbox(bboxes, pred_bboxes, self.means,
+ self.stds, max_shape,
+ wh_ratio_clip, self.clip_border,
+ self.add_ctr_clamp,
+ self.ctr_clamp)
+
+ if self.use_box_type:
+ assert decoded_bboxes.size(-1) == 4, \
+ ('Cannot warp decoded boxes with box type when decoded boxes'
+ 'have shape of (N, num_classes * 4)')
+ decoded_bboxes = HorizontalBoxes(decoded_bboxes)
+ return decoded_bboxes
+
+
+@TASK_UTILS.register_module()
+class DeltaXYWHBBoxCoderForGLIP(DeltaXYWHBBoxCoder):
+ """This is designed specifically for the GLIP algorithm.
+
+ In order to completely match the official performance, we need to perform
+ special calculations in the encoding and decoding processes, such as
+ additional +1 and -1 calculations. However, this is not a user-friendly
+ design.
+ """
+
+ def encode(self, bboxes: Union[Tensor, BaseBoxes],
+ gt_bboxes: Union[Tensor, BaseBoxes]) -> Tensor:
+ """Get box regression transformation deltas that can be used to
+ transform the ``bboxes`` into the ``gt_bboxes``.
+
+ Args:
+ bboxes (torch.Tensor or :obj:`BaseBoxes`): Source boxes,
+ e.g., object proposals.
+ gt_bboxes (torch.Tensor or :obj:`BaseBoxes`): Target of the
+ transformation, e.g., ground-truth boxes.
+
+ Returns:
+ torch.Tensor: Box transformation deltas
+ """
+ bboxes = get_box_tensor(bboxes)
+ gt_bboxes = get_box_tensor(gt_bboxes)
+ assert bboxes.size(0) == gt_bboxes.size(0)
+ assert bboxes.size(-1) == gt_bboxes.size(-1) == 4
+ encoded_bboxes = bbox2delta(bboxes, gt_bboxes, self.means, self.stds)
+ return encoded_bboxes
+
+ def decode(
+ self,
+ bboxes: Union[Tensor, BaseBoxes],
+ pred_bboxes: Tensor,
+ max_shape: Optional[Union[Sequence[int], Tensor,
+ Sequence[Sequence[int]]]] = None,
+ wh_ratio_clip: Optional[float] = 16 / 1000
+ ) -> Union[Tensor, BaseBoxes]:
+ """Apply transformation `pred_bboxes` to `boxes`.
+
+ Args:
+ bboxes (torch.Tensor or :obj:`BaseBoxes`): Basic boxes. Shape
+ (B, N, 4) or (N, 4)
+ pred_bboxes (Tensor): Encoded offsets with respect to each roi.
+ Has shape (B, N, num_classes * 4) or (B, N, 4) or
+ (N, num_classes * 4) or (N, 4). Note N = num_anchors * W * H
+ when rois is a grid of anchors.Offset encoding follows [1]_.
+ max_shape (Sequence[int] or torch.Tensor or Sequence[
+ Sequence[int]],optional): Maximum bounds for boxes, specifies
+ (H, W, C) or (H, W). If bboxes shape is (B, N, 4), then
+ the max_shape should be a Sequence[Sequence[int]]
+ and the length of max_shape should also be B.
+ wh_ratio_clip (float, optional): The allowed ratio between
+ width and height.
+
+ Returns:
+ Union[torch.Tensor, :obj:`BaseBoxes`]: Decoded boxes.
+ """
+ bboxes = get_box_tensor(bboxes)
+ assert pred_bboxes.size(0) == bboxes.size(0)
+ if pred_bboxes.ndim == 3:
+ assert pred_bboxes.size(1) == bboxes.size(1)
+
+ if pred_bboxes.ndim == 2 and not torch.onnx.is_in_onnx_export():
+ # single image decode
+ decoded_bboxes = delta2bbox_glip(bboxes, pred_bboxes, self.means,
+ self.stds, max_shape,
+ wh_ratio_clip, self.clip_border,
+ self.add_ctr_clamp,
+ self.ctr_clamp)
+ else:
+ raise NotImplementedError()
+
+ if self.use_box_type:
+ assert decoded_bboxes.size(-1) == 4, \
+ ('Cannot warp decoded boxes with box type when decoded boxes'
+ 'have shape of (N, num_classes * 4)')
+ decoded_bboxes = HorizontalBoxes(decoded_bboxes)
+ return decoded_bboxes
+
+
+def bbox2delta(
+ proposals: Tensor,
+ gt: Tensor,
+ means: Sequence[float] = (0., 0., 0., 0.),
+ stds: Sequence[float] = (1., 1., 1., 1.)
+) -> Tensor:
+ """Compute deltas of proposals w.r.t. gt.
+
+ We usually compute the deltas of x, y, w, h of proposals w.r.t ground
+ truth bboxes to get regression target.
+ This is the inverse function of :func:`delta2bbox`.
+
+ Args:
+ proposals (Tensor): Boxes to be transformed, shape (N, ..., 4)
+ gt (Tensor): Gt bboxes to be used as base, shape (N, ..., 4)
+ means (Sequence[float]): Denormalizing means for delta coordinates
+ stds (Sequence[float]): Denormalizing standard deviation for delta
+ coordinates
+
+ Returns:
+ Tensor: deltas with shape (N, 4), where columns represent dx, dy,
+ dw, dh.
+ """
+ assert proposals.size() == gt.size()
+
+ proposals = proposals.float()
+ gt = gt.float()
+ px = (proposals[..., 0] + proposals[..., 2]) * 0.5
+ py = (proposals[..., 1] + proposals[..., 3]) * 0.5
+ pw = proposals[..., 2] - proposals[..., 0]
+ ph = proposals[..., 3] - proposals[..., 1]
+
+ gx = (gt[..., 0] + gt[..., 2]) * 0.5
+ gy = (gt[..., 1] + gt[..., 3]) * 0.5
+ gw = gt[..., 2] - gt[..., 0]
+ gh = gt[..., 3] - gt[..., 1]
+
+ dx = (gx - px) / pw
+ dy = (gy - py) / ph
+ dw = torch.log(gw / pw)
+ dh = torch.log(gh / ph)
+ deltas = torch.stack([dx, dy, dw, dh], dim=-1)
+
+ means = deltas.new_tensor(means).unsqueeze(0)
+ stds = deltas.new_tensor(stds).unsqueeze(0)
+ deltas = deltas.sub_(means).div_(stds)
+
+ return deltas
+
+
+def delta2bbox(rois: Tensor,
+ deltas: Tensor,
+ means: Sequence[float] = (0., 0., 0., 0.),
+ stds: Sequence[float] = (1., 1., 1., 1.),
+ max_shape: Optional[Union[Sequence[int], Tensor,
+ Sequence[Sequence[int]]]] = None,
+ wh_ratio_clip: float = 16 / 1000,
+ clip_border: bool = True,
+ add_ctr_clamp: bool = False,
+ ctr_clamp: int = 32) -> Tensor:
+ """Apply deltas to shift/scale base boxes.
+
+ Typically the rois are anchor or proposed bounding boxes and the deltas are
+ network outputs used to shift/scale those boxes.
+ This is the inverse function of :func:`bbox2delta`.
+
+ Args:
+ rois (Tensor): Boxes to be transformed. Has shape (N, 4).
+ deltas (Tensor): Encoded offsets relative to each roi.
+ Has shape (N, num_classes * 4) or (N, 4). Note
+ N = num_base_anchors * W * H, when rois is a grid of
+ anchors. Offset encoding follows [1]_.
+ means (Sequence[float]): Denormalizing means for delta coordinates.
+ Default (0., 0., 0., 0.).
+ stds (Sequence[float]): Denormalizing standard deviation for delta
+ coordinates. Default (1., 1., 1., 1.).
+ max_shape (tuple[int, int]): Maximum bounds for boxes, specifies
+ (H, W). Default None.
+ wh_ratio_clip (float): Maximum aspect ratio for boxes. Default
+ 16 / 1000.
+ clip_border (bool, optional): Whether clip the objects outside the
+ border of the image. Default True.
+ add_ctr_clamp (bool): Whether to add center clamp. When set to True,
+ the center of the prediction bounding box will be clamped to
+ avoid being too far away from the center of the anchor.
+ Only used by YOLOF. Default False.
+ ctr_clamp (int): the maximum pixel shift to clamp. Only used by YOLOF.
+ Default 32.
+
+ Returns:
+ Tensor: Boxes with shape (N, num_classes * 4) or (N, 4), where 4
+ represent tl_x, tl_y, br_x, br_y.
+
+ References:
+ .. [1] https://arxiv.org/abs/1311.2524
+
+ Example:
+ >>> rois = torch.Tensor([[ 0., 0., 1., 1.],
+ >>> [ 0., 0., 1., 1.],
+ >>> [ 0., 0., 1., 1.],
+ >>> [ 5., 5., 5., 5.]])
+ >>> deltas = torch.Tensor([[ 0., 0., 0., 0.],
+ >>> [ 1., 1., 1., 1.],
+ >>> [ 0., 0., 2., -1.],
+ >>> [ 0.7, -1.9, -0.5, 0.3]])
+ >>> delta2bbox(rois, deltas, max_shape=(32, 32, 3))
+ tensor([[0.0000, 0.0000, 1.0000, 1.0000],
+ [0.1409, 0.1409, 2.8591, 2.8591],
+ [0.0000, 0.3161, 4.1945, 0.6839],
+ [5.0000, 5.0000, 5.0000, 5.0000]])
+ """
+ num_bboxes, num_classes = deltas.size(0), deltas.size(1) // 4
+ if num_bboxes == 0:
+ return deltas
+
+ deltas = deltas.reshape(-1, 4)
+
+ means = deltas.new_tensor(means).view(1, -1)
+ stds = deltas.new_tensor(stds).view(1, -1)
+ denorm_deltas = deltas * stds + means
+
+ dxy = denorm_deltas[:, :2]
+ dwh = denorm_deltas[:, 2:]
+
+ # Compute width/height of each roi
+ rois_ = rois.repeat(1, num_classes).reshape(-1, 4)
+ pxy = ((rois_[:, :2] + rois_[:, 2:]) * 0.5)
+ pwh = (rois_[:, 2:] - rois_[:, :2])
+
+ dxy_wh = pwh * dxy
+
+ max_ratio = np.abs(np.log(wh_ratio_clip))
+ if add_ctr_clamp:
+ dxy_wh = torch.clamp(dxy_wh, max=ctr_clamp, min=-ctr_clamp)
+ dwh = torch.clamp(dwh, max=max_ratio)
+ else:
+ dwh = dwh.clamp(min=-max_ratio, max=max_ratio)
+
+ gxy = pxy + dxy_wh
+ gwh = pwh * dwh.exp()
+ x1y1 = gxy - (gwh * 0.5)
+ x2y2 = gxy + (gwh * 0.5)
+ bboxes = torch.cat([x1y1, x2y2], dim=-1)
+ if clip_border and max_shape is not None:
+ bboxes[..., 0::2].clamp_(min=0, max=max_shape[1])
+ bboxes[..., 1::2].clamp_(min=0, max=max_shape[0])
+ bboxes = bboxes.reshape(num_bboxes, -1)
+ return bboxes
+
+
+def onnx_delta2bbox(rois: Tensor,
+ deltas: Tensor,
+ means: Sequence[float] = (0., 0., 0., 0.),
+ stds: Sequence[float] = (1., 1., 1., 1.),
+ max_shape: Optional[Union[Sequence[int], Tensor,
+ Sequence[Sequence[int]]]] = None,
+ wh_ratio_clip: float = 16 / 1000,
+ clip_border: Optional[bool] = True,
+ add_ctr_clamp: bool = False,
+ ctr_clamp: int = 32) -> Tensor:
+ """Apply deltas to shift/scale base boxes.
+
+ Typically the rois are anchor or proposed bounding boxes and the deltas are
+ network outputs used to shift/scale those boxes.
+ This is the inverse function of :func:`bbox2delta`.
+
+ Args:
+ rois (Tensor): Boxes to be transformed. Has shape (N, 4) or (B, N, 4)
+ deltas (Tensor): Encoded offsets with respect to each roi.
+ Has shape (B, N, num_classes * 4) or (B, N, 4) or
+ (N, num_classes * 4) or (N, 4). Note N = num_anchors * W * H
+ when rois is a grid of anchors.Offset encoding follows [1]_.
+ means (Sequence[float]): Denormalizing means for delta coordinates.
+ Default (0., 0., 0., 0.).
+ stds (Sequence[float]): Denormalizing standard deviation for delta
+ coordinates. Default (1., 1., 1., 1.).
+ max_shape (Sequence[int] or torch.Tensor or Sequence[
+ Sequence[int]],optional): Maximum bounds for boxes, specifies
+ (H, W, C) or (H, W). If rois shape is (B, N, 4), then
+ the max_shape should be a Sequence[Sequence[int]]
+ and the length of max_shape should also be B. Default None.
+ wh_ratio_clip (float): Maximum aspect ratio for boxes.
+ Default 16 / 1000.
+ clip_border (bool, optional): Whether clip the objects outside the
+ border of the image. Default True.
+ add_ctr_clamp (bool): Whether to add center clamp, when added, the
+ predicted box is clamped is its center is too far away from
+ the original anchor's center. Only used by YOLOF. Default False.
+ ctr_clamp (int): the maximum pixel shift to clamp. Only used by YOLOF.
+ Default 32.
+
+ Returns:
+ Tensor: Boxes with shape (B, N, num_classes * 4) or (B, N, 4) or
+ (N, num_classes * 4) or (N, 4), where 4 represent
+ tl_x, tl_y, br_x, br_y.
+
+ References:
+ .. [1] https://arxiv.org/abs/1311.2524
+
+ Example:
+ >>> rois = torch.Tensor([[ 0., 0., 1., 1.],
+ >>> [ 0., 0., 1., 1.],
+ >>> [ 0., 0., 1., 1.],
+ >>> [ 5., 5., 5., 5.]])
+ >>> deltas = torch.Tensor([[ 0., 0., 0., 0.],
+ >>> [ 1., 1., 1., 1.],
+ >>> [ 0., 0., 2., -1.],
+ >>> [ 0.7, -1.9, -0.5, 0.3]])
+ >>> delta2bbox(rois, deltas, max_shape=(32, 32, 3))
+ tensor([[0.0000, 0.0000, 1.0000, 1.0000],
+ [0.1409, 0.1409, 2.8591, 2.8591],
+ [0.0000, 0.3161, 4.1945, 0.6839],
+ [5.0000, 5.0000, 5.0000, 5.0000]])
+ """
+ means = deltas.new_tensor(means).view(1,
+ -1).repeat(1,
+ deltas.size(-1) // 4)
+ stds = deltas.new_tensor(stds).view(1, -1).repeat(1, deltas.size(-1) // 4)
+ denorm_deltas = deltas * stds + means
+ dx = denorm_deltas[..., 0::4]
+ dy = denorm_deltas[..., 1::4]
+ dw = denorm_deltas[..., 2::4]
+ dh = denorm_deltas[..., 3::4]
+
+ x1, y1 = rois[..., 0], rois[..., 1]
+ x2, y2 = rois[..., 2], rois[..., 3]
+ # Compute center of each roi
+ px = ((x1 + x2) * 0.5).unsqueeze(-1).expand_as(dx)
+ py = ((y1 + y2) * 0.5).unsqueeze(-1).expand_as(dy)
+ # Compute width/height of each roi
+ pw = (x2 - x1).unsqueeze(-1).expand_as(dw)
+ ph = (y2 - y1).unsqueeze(-1).expand_as(dh)
+
+ dx_width = pw * dx
+ dy_height = ph * dy
+
+ max_ratio = np.abs(np.log(wh_ratio_clip))
+ if add_ctr_clamp:
+ dx_width = torch.clamp(dx_width, max=ctr_clamp, min=-ctr_clamp)
+ dy_height = torch.clamp(dy_height, max=ctr_clamp, min=-ctr_clamp)
+ dw = torch.clamp(dw, max=max_ratio)
+ dh = torch.clamp(dh, max=max_ratio)
+ else:
+ dw = dw.clamp(min=-max_ratio, max=max_ratio)
+ dh = dh.clamp(min=-max_ratio, max=max_ratio)
+ # Use exp(network energy) to enlarge/shrink each roi
+ gw = pw * dw.exp()
+ gh = ph * dh.exp()
+ # Use network energy to shift the center of each roi
+ gx = px + dx_width
+ gy = py + dy_height
+ # Convert center-xy/width/height to top-left, bottom-right
+ x1 = gx - gw * 0.5
+ y1 = gy - gh * 0.5
+ x2 = gx + gw * 0.5
+ y2 = gy + gh * 0.5
+
+ bboxes = torch.stack([x1, y1, x2, y2], dim=-1).view(deltas.size())
+
+ if clip_border and max_shape is not None:
+ # clip bboxes with dynamic `min` and `max` for onnx
+ if torch.onnx.is_in_onnx_export():
+ from mmdet.core.export import dynamic_clip_for_onnx
+ x1, y1, x2, y2 = dynamic_clip_for_onnx(x1, y1, x2, y2, max_shape)
+ bboxes = torch.stack([x1, y1, x2, y2], dim=-1).view(deltas.size())
+ return bboxes
+ if not isinstance(max_shape, torch.Tensor):
+ max_shape = x1.new_tensor(max_shape)
+ max_shape = max_shape[..., :2].type_as(x1)
+ if max_shape.ndim == 2:
+ assert bboxes.ndim == 3
+ assert max_shape.size(0) == bboxes.size(0)
+
+ min_xy = x1.new_tensor(0)
+ max_xy = torch.cat(
+ [max_shape] * (deltas.size(-1) // 2),
+ dim=-1).flip(-1).unsqueeze(-2)
+ bboxes = torch.where(bboxes < min_xy, min_xy, bboxes)
+ bboxes = torch.where(bboxes > max_xy, max_xy, bboxes)
+
+ return bboxes
+
+
+def delta2bbox_glip(rois: Tensor,
+ deltas: Tensor,
+ means: Sequence[float] = (0., 0., 0., 0.),
+ stds: Sequence[float] = (1., 1., 1., 1.),
+ max_shape: Optional[Union[Sequence[int], Tensor,
+ Sequence[Sequence[int]]]] = None,
+ wh_ratio_clip: float = 16 / 1000,
+ clip_border: bool = True,
+ add_ctr_clamp: bool = False,
+ ctr_clamp: int = 32) -> Tensor:
+ """Apply deltas to shift/scale base boxes.
+
+ Typically the rois are anchor or proposed bounding boxes and the deltas are
+ network outputs used to shift/scale those boxes.
+ This is the inverse function of :func:`bbox2delta`.
+
+ Args:
+ rois (Tensor): Boxes to be transformed. Has shape (N, 4).
+ deltas (Tensor): Encoded offsets relative to each roi.
+ Has shape (N, num_classes * 4) or (N, 4). Note
+ N = num_base_anchors * W * H, when rois is a grid of
+ anchors. Offset encoding follows [1]_.
+ means (Sequence[float]): Denormalizing means for delta coordinates.
+ Default (0., 0., 0., 0.).
+ stds (Sequence[float]): Denormalizing standard deviation for delta
+ coordinates. Default (1., 1., 1., 1.).
+ max_shape (tuple[int, int]): Maximum bounds for boxes, specifies
+ (H, W). Default None.
+ wh_ratio_clip (float): Maximum aspect ratio for boxes. Default
+ 16 / 1000.
+ clip_border (bool, optional): Whether clip the objects outside the
+ border of the image. Default True.
+ add_ctr_clamp (bool): Whether to add center clamp. When set to True,
+ the center of the prediction bounding box will be clamped to
+ avoid being too far away from the center of the anchor.
+ Only used by YOLOF. Default False.
+ ctr_clamp (int): the maximum pixel shift to clamp. Only used by YOLOF.
+ Default 32.
+
+ Returns:
+ Tensor: Boxes with shape (N, num_classes * 4) or (N, 4), where 4
+ represent tl_x, tl_y, br_x, br_y.
+ """
+ num_bboxes, num_classes = deltas.size(0), deltas.size(1) // 4
+ if num_bboxes == 0:
+ return deltas
+
+ deltas = deltas.reshape(-1, 4)
+
+ means = deltas.new_tensor(means).view(1, -1)
+ stds = deltas.new_tensor(stds).view(1, -1)
+ denorm_deltas = deltas * stds + means
+
+ dxy = denorm_deltas[:, :2]
+ dwh = denorm_deltas[:, 2:]
+
+ # Compute width/height of each roi
+ rois_ = rois.repeat(1, num_classes).reshape(-1, 4)
+ pxy = ((rois_[:, :2] + rois_[:, 2:] - 1) * 0.5) # note
+ pwh = (rois_[:, 2:] - rois_[:, :2])
+
+ dxy_wh = pwh * dxy
+
+ max_ratio = np.abs(np.log(wh_ratio_clip))
+ if add_ctr_clamp:
+ dxy_wh = torch.clamp(dxy_wh, max=ctr_clamp, min=-ctr_clamp)
+ dwh = torch.clamp(dwh, max=max_ratio)
+ else:
+ dwh = dwh.clamp(min=-max_ratio, max=max_ratio)
+
+ gxy = pxy + dxy_wh
+ gwh = pwh * dwh.exp()
+
+ x1y1 = gxy - (gwh - 1) * 0.5 # Note
+ x2y2 = gxy + (gwh - 1) * 0.5 # Note
+
+ bboxes = torch.cat([x1y1, x2y2], dim=-1)
+
+ if clip_border and max_shape is not None:
+ bboxes[..., 0::2].clamp_(min=0, max=max_shape[1] - 1) # Note
+ bboxes[..., 1::2].clamp_(min=0, max=max_shape[0] - 1) # Note
+ bboxes = bboxes.reshape(num_bboxes, -1)
+ return bboxes
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/distance_point_bbox_coder.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/distance_point_bbox_coder.py
new file mode 100644
index 0000000000000000000000000000000000000000..ab26bf4b96c48df689da3722c23aa65e646348db
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/distance_point_bbox_coder.py
@@ -0,0 +1,85 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Sequence, Union
+
+from torch import Tensor
+
+from mmdet.registry import TASK_UTILS
+from mmdet.structures.bbox import (BaseBoxes, HorizontalBoxes, bbox2distance,
+ distance2bbox, get_box_tensor)
+from .base_bbox_coder import BaseBBoxCoder
+
+
+@TASK_UTILS.register_module()
+class DistancePointBBoxCoder(BaseBBoxCoder):
+ """Distance Point BBox coder.
+
+ This coder encodes gt bboxes (x1, y1, x2, y2) into (top, bottom, left,
+ right) and decode it back to the original.
+
+ Args:
+ clip_border (bool, optional): Whether clip the objects outside the
+ border of the image. Defaults to True.
+ """
+
+ def __init__(self, clip_border: Optional[bool] = True, **kwargs) -> None:
+ super().__init__(**kwargs)
+ self.clip_border = clip_border
+
+ def encode(self,
+ points: Tensor,
+ gt_bboxes: Union[Tensor, BaseBoxes],
+ max_dis: Optional[float] = None,
+ eps: float = 0.1) -> Tensor:
+ """Encode bounding box to distances.
+
+ Args:
+ points (Tensor): Shape (N, 2), The format is [x, y].
+ gt_bboxes (Tensor or :obj:`BaseBoxes`): Shape (N, 4), The format
+ is "xyxy"
+ max_dis (float): Upper bound of the distance. Default None.
+ eps (float): a small value to ensure target < max_dis, instead <=.
+ Default 0.1.
+
+ Returns:
+ Tensor: Box transformation deltas. The shape is (N, 4).
+ """
+ gt_bboxes = get_box_tensor(gt_bboxes)
+ assert points.size(0) == gt_bboxes.size(0)
+ assert points.size(-1) == 2
+ assert gt_bboxes.size(-1) == 4
+ return bbox2distance(points, gt_bboxes, max_dis, eps)
+
+ def decode(
+ self,
+ points: Tensor,
+ pred_bboxes: Tensor,
+ max_shape: Optional[Union[Sequence[int], Tensor,
+ Sequence[Sequence[int]]]] = None
+ ) -> Union[Tensor, BaseBoxes]:
+ """Decode distance prediction to bounding box.
+
+ Args:
+ points (Tensor): Shape (B, N, 2) or (N, 2).
+ pred_bboxes (Tensor): Distance from the given point to 4
+ boundaries (left, top, right, bottom). Shape (B, N, 4)
+ or (N, 4)
+ max_shape (Sequence[int] or torch.Tensor or Sequence[
+ Sequence[int]],optional): Maximum bounds for boxes, specifies
+ (H, W, C) or (H, W). If priors shape is (B, N, 4), then
+ the max_shape should be a Sequence[Sequence[int]],
+ and the length of max_shape should also be B.
+ Default None.
+ Returns:
+ Union[Tensor, :obj:`BaseBoxes`]: Boxes with shape (N, 4) or
+ (B, N, 4)
+ """
+ assert points.size(0) == pred_bboxes.size(0)
+ assert points.size(-1) == 2
+ assert pred_bboxes.size(-1) == 4
+ if self.clip_border is False:
+ max_shape = None
+ bboxes = distance2bbox(points, pred_bboxes, max_shape)
+
+ if self.use_box_type:
+ bboxes = HorizontalBoxes(bboxes)
+ return bboxes
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/legacy_delta_xywh_bbox_coder.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/legacy_delta_xywh_bbox_coder.py
new file mode 100644
index 0000000000000000000000000000000000000000..9eb1bedb3fbe19433c8bdb37f80891efa2cb72fc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/legacy_delta_xywh_bbox_coder.py
@@ -0,0 +1,235 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Sequence, Union
+
+import numpy as np
+import torch
+from torch import Tensor
+
+from mmdet.registry import TASK_UTILS
+from mmdet.structures.bbox import BaseBoxes, HorizontalBoxes, get_box_tensor
+from .base_bbox_coder import BaseBBoxCoder
+
+
+@TASK_UTILS.register_module()
+class LegacyDeltaXYWHBBoxCoder(BaseBBoxCoder):
+ """Legacy Delta XYWH BBox coder used in MMDet V1.x.
+
+ Following the practice in R-CNN [1]_, this coder encodes bbox (x1, y1, x2,
+ y2) into delta (dx, dy, dw, dh) and decodes delta (dx, dy, dw, dh)
+ back to original bbox (x1, y1, x2, y2).
+
+ Note:
+ The main difference between :class`LegacyDeltaXYWHBBoxCoder` and
+ :class:`DeltaXYWHBBoxCoder` is whether ``+ 1`` is used during width and
+ height calculation. We suggest to only use this coder when testing with
+ MMDet V1.x models.
+
+ References:
+ .. [1] https://arxiv.org/abs/1311.2524
+
+ Args:
+ target_means (Sequence[float]): denormalizing means of target for
+ delta coordinates
+ target_stds (Sequence[float]): denormalizing standard deviation of
+ target for delta coordinates
+ """
+
+ def __init__(self,
+ target_means: Sequence[float] = (0., 0., 0., 0.),
+ target_stds: Sequence[float] = (1., 1., 1., 1.),
+ **kwargs) -> None:
+ super().__init__(**kwargs)
+ self.means = target_means
+ self.stds = target_stds
+
+ def encode(self, bboxes: Union[Tensor, BaseBoxes],
+ gt_bboxes: Union[Tensor, BaseBoxes]) -> Tensor:
+ """Get box regression transformation deltas that can be used to
+ transform the ``bboxes`` into the ``gt_bboxes``.
+
+ Args:
+ bboxes (torch.Tensor or :obj:`BaseBoxes`): source boxes,
+ e.g., object proposals.
+ gt_bboxes (torch.Tensor or :obj:`BaseBoxes`): target of the
+ transformation, e.g., ground-truth boxes.
+
+ Returns:
+ torch.Tensor: Box transformation deltas
+ """
+ bboxes = get_box_tensor(bboxes)
+ gt_bboxes = get_box_tensor(gt_bboxes)
+ assert bboxes.size(0) == gt_bboxes.size(0)
+ assert bboxes.size(-1) == gt_bboxes.size(-1) == 4
+ encoded_bboxes = legacy_bbox2delta(bboxes, gt_bboxes, self.means,
+ self.stds)
+ return encoded_bboxes
+
+ def decode(
+ self,
+ bboxes: Union[Tensor, BaseBoxes],
+ pred_bboxes: Tensor,
+ max_shape: Optional[Union[Sequence[int], Tensor,
+ Sequence[Sequence[int]]]] = None,
+ wh_ratio_clip: Optional[float] = 16 / 1000
+ ) -> Union[Tensor, BaseBoxes]:
+ """Apply transformation `pred_bboxes` to `boxes`.
+
+ Args:
+ boxes (torch.Tensor or :obj:`BaseBoxes`): Basic boxes.
+ pred_bboxes (torch.Tensor): Encoded boxes with shape
+ max_shape (tuple[int], optional): Maximum shape of boxes.
+ Defaults to None.
+ wh_ratio_clip (float, optional): The allowed ratio between
+ width and height.
+
+ Returns:
+ Union[torch.Tensor, :obj:`BaseBoxes`]: Decoded boxes.
+ """
+ bboxes = get_box_tensor(bboxes)
+ assert pred_bboxes.size(0) == bboxes.size(0)
+ decoded_bboxes = legacy_delta2bbox(bboxes, pred_bboxes, self.means,
+ self.stds, max_shape, wh_ratio_clip)
+
+ if self.use_box_type:
+ assert decoded_bboxes.size(-1) == 4, \
+ ('Cannot warp decoded boxes with box type when decoded boxes'
+ 'have shape of (N, num_classes * 4)')
+ decoded_bboxes = HorizontalBoxes(decoded_bboxes)
+ return decoded_bboxes
+
+
+def legacy_bbox2delta(
+ proposals: Tensor,
+ gt: Tensor,
+ means: Sequence[float] = (0., 0., 0., 0.),
+ stds: Sequence[float] = (1., 1., 1., 1.)
+) -> Tensor:
+ """Compute deltas of proposals w.r.t. gt in the MMDet V1.x manner.
+
+ We usually compute the deltas of x, y, w, h of proposals w.r.t ground
+ truth bboxes to get regression target.
+ This is the inverse function of `delta2bbox()`
+
+ Args:
+ proposals (Tensor): Boxes to be transformed, shape (N, ..., 4)
+ gt (Tensor): Gt bboxes to be used as base, shape (N, ..., 4)
+ means (Sequence[float]): Denormalizing means for delta coordinates
+ stds (Sequence[float]): Denormalizing standard deviation for delta
+ coordinates
+
+ Returns:
+ Tensor: deltas with shape (N, 4), where columns represent dx, dy,
+ dw, dh.
+ """
+ assert proposals.size() == gt.size()
+
+ proposals = proposals.float()
+ gt = gt.float()
+ px = (proposals[..., 0] + proposals[..., 2]) * 0.5
+ py = (proposals[..., 1] + proposals[..., 3]) * 0.5
+ pw = proposals[..., 2] - proposals[..., 0] + 1.0
+ ph = proposals[..., 3] - proposals[..., 1] + 1.0
+
+ gx = (gt[..., 0] + gt[..., 2]) * 0.5
+ gy = (gt[..., 1] + gt[..., 3]) * 0.5
+ gw = gt[..., 2] - gt[..., 0] + 1.0
+ gh = gt[..., 3] - gt[..., 1] + 1.0
+
+ dx = (gx - px) / pw
+ dy = (gy - py) / ph
+ dw = torch.log(gw / pw)
+ dh = torch.log(gh / ph)
+ deltas = torch.stack([dx, dy, dw, dh], dim=-1)
+
+ means = deltas.new_tensor(means).unsqueeze(0)
+ stds = deltas.new_tensor(stds).unsqueeze(0)
+ deltas = deltas.sub_(means).div_(stds)
+
+ return deltas
+
+
+def legacy_delta2bbox(rois: Tensor,
+ deltas: Tensor,
+ means: Sequence[float] = (0., 0., 0., 0.),
+ stds: Sequence[float] = (1., 1., 1., 1.),
+ max_shape: Optional[
+ Union[Sequence[int], Tensor,
+ Sequence[Sequence[int]]]] = None,
+ wh_ratio_clip: float = 16 / 1000) -> Tensor:
+ """Apply deltas to shift/scale base boxes in the MMDet V1.x manner.
+
+ Typically the rois are anchor or proposed bounding boxes and the deltas are
+ network outputs used to shift/scale those boxes.
+ This is the inverse function of `bbox2delta()`
+
+ Args:
+ rois (Tensor): Boxes to be transformed. Has shape (N, 4)
+ deltas (Tensor): Encoded offsets with respect to each roi.
+ Has shape (N, 4 * num_classes). Note N = num_anchors * W * H when
+ rois is a grid of anchors. Offset encoding follows [1]_.
+ means (Sequence[float]): Denormalizing means for delta coordinates
+ stds (Sequence[float]): Denormalizing standard deviation for delta
+ coordinates
+ max_shape (tuple[int, int]): Maximum bounds for boxes. specifies (H, W)
+ wh_ratio_clip (float): Maximum aspect ratio for boxes.
+
+ Returns:
+ Tensor: Boxes with shape (N, 4), where columns represent
+ tl_x, tl_y, br_x, br_y.
+
+ References:
+ .. [1] https://arxiv.org/abs/1311.2524
+
+ Example:
+ >>> rois = torch.Tensor([[ 0., 0., 1., 1.],
+ >>> [ 0., 0., 1., 1.],
+ >>> [ 0., 0., 1., 1.],
+ >>> [ 5., 5., 5., 5.]])
+ >>> deltas = torch.Tensor([[ 0., 0., 0., 0.],
+ >>> [ 1., 1., 1., 1.],
+ >>> [ 0., 0., 2., -1.],
+ >>> [ 0.7, -1.9, -0.5, 0.3]])
+ >>> legacy_delta2bbox(rois, deltas, max_shape=(32, 32))
+ tensor([[0.0000, 0.0000, 1.5000, 1.5000],
+ [0.0000, 0.0000, 5.2183, 5.2183],
+ [0.0000, 0.1321, 7.8891, 0.8679],
+ [5.3967, 2.4251, 6.0033, 3.7749]])
+ """
+ means = deltas.new_tensor(means).repeat(1, deltas.size(1) // 4)
+ stds = deltas.new_tensor(stds).repeat(1, deltas.size(1) // 4)
+ denorm_deltas = deltas * stds + means
+ dx = denorm_deltas[:, 0::4]
+ dy = denorm_deltas[:, 1::4]
+ dw = denorm_deltas[:, 2::4]
+ dh = denorm_deltas[:, 3::4]
+ max_ratio = np.abs(np.log(wh_ratio_clip))
+ dw = dw.clamp(min=-max_ratio, max=max_ratio)
+ dh = dh.clamp(min=-max_ratio, max=max_ratio)
+ # Compute center of each roi
+ px = ((rois[:, 0] + rois[:, 2]) * 0.5).unsqueeze(1).expand_as(dx)
+ py = ((rois[:, 1] + rois[:, 3]) * 0.5).unsqueeze(1).expand_as(dy)
+ # Compute width/height of each roi
+ pw = (rois[:, 2] - rois[:, 0] + 1.0).unsqueeze(1).expand_as(dw)
+ ph = (rois[:, 3] - rois[:, 1] + 1.0).unsqueeze(1).expand_as(dh)
+ # Use exp(network energy) to enlarge/shrink each roi
+ gw = pw * dw.exp()
+ gh = ph * dh.exp()
+ # Use network energy to shift the center of each roi
+ gx = px + pw * dx
+ gy = py + ph * dy
+ # Convert center-xy/width/height to top-left, bottom-right
+
+ # The true legacy box coder should +- 0.5 here.
+ # However, current implementation improves the performance when testing
+ # the models trained in MMDetection 1.X (~0.5 bbox AP, 0.2 mask AP)
+ x1 = gx - gw * 0.5
+ y1 = gy - gh * 0.5
+ x2 = gx + gw * 0.5
+ y2 = gy + gh * 0.5
+ if max_shape is not None:
+ x1 = x1.clamp(min=0, max=max_shape[1] - 1)
+ y1 = y1.clamp(min=0, max=max_shape[0] - 1)
+ x2 = x2.clamp(min=0, max=max_shape[1] - 1)
+ y2 = y2.clamp(min=0, max=max_shape[0] - 1)
+ bboxes = torch.stack([x1, y1, x2, y2], dim=-1).view_as(deltas)
+ return bboxes
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/pseudo_bbox_coder.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/pseudo_bbox_coder.py
new file mode 100644
index 0000000000000000000000000000000000000000..9ee74311f6d12bde49d0c678edb60540a8c95c8b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/pseudo_bbox_coder.py
@@ -0,0 +1,29 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Union
+
+from torch import Tensor
+
+from mmdet.registry import TASK_UTILS
+from mmdet.structures.bbox import BaseBoxes, HorizontalBoxes, get_box_tensor
+from .base_bbox_coder import BaseBBoxCoder
+
+
+@TASK_UTILS.register_module()
+class PseudoBBoxCoder(BaseBBoxCoder):
+ """Pseudo bounding box coder."""
+
+ def __init__(self, **kwargs):
+ super().__init__(**kwargs)
+
+ def encode(self, bboxes: Tensor, gt_bboxes: Union[Tensor,
+ BaseBoxes]) -> Tensor:
+ """torch.Tensor: return the given ``bboxes``"""
+ gt_bboxes = get_box_tensor(gt_bboxes)
+ return gt_bboxes
+
+ def decode(self, bboxes: Tensor, pred_bboxes: Union[Tensor,
+ BaseBoxes]) -> Tensor:
+ """torch.Tensor: return the given ``pred_bboxes``"""
+ if self.use_box_type:
+ pred_bboxes = HorizontalBoxes(pred_bboxes)
+ return pred_bboxes
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/tblr_bbox_coder.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/tblr_bbox_coder.py
new file mode 100644
index 0000000000000000000000000000000000000000..74b388f7bad6ebc1911cee5b0b7d73bbd04de17a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/tblr_bbox_coder.py
@@ -0,0 +1,228 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Sequence, Union
+
+import torch
+from torch import Tensor
+
+from mmdet.registry import TASK_UTILS
+from mmdet.structures.bbox import BaseBoxes, HorizontalBoxes, get_box_tensor
+from .base_bbox_coder import BaseBBoxCoder
+
+
+@TASK_UTILS.register_module()
+class TBLRBBoxCoder(BaseBBoxCoder):
+ """TBLR BBox coder.
+
+ Following the practice in `FSAF `_,
+ this coder encodes gt bboxes (x1, y1, x2, y2) into (top, bottom, left,
+ right) and decode it back to the original.
+
+ Args:
+ normalizer (list | float): Normalization factor to be
+ divided with when coding the coordinates. If it is a list, it should
+ have length of 4 indicating normalization factor in tblr dims.
+ Otherwise it is a unified float factor for all dims. Default: 4.0
+ clip_border (bool, optional): Whether clip the objects outside the
+ border of the image. Defaults to True.
+ """
+
+ def __init__(self,
+ normalizer: Union[Sequence[float], float] = 4.0,
+ clip_border: bool = True,
+ **kwargs) -> None:
+ super().__init__(**kwargs)
+ self.normalizer = normalizer
+ self.clip_border = clip_border
+
+ def encode(self, bboxes: Union[Tensor, BaseBoxes],
+ gt_bboxes: Union[Tensor, BaseBoxes]) -> Tensor:
+ """Get box regression transformation deltas that can be used to
+ transform the ``bboxes`` into the ``gt_bboxes`` in the (top, left,
+ bottom, right) order.
+
+ Args:
+ bboxes (torch.Tensor or :obj:`BaseBoxes`): source boxes,
+ e.g., object proposals.
+ gt_bboxes (torch.Tensor or :obj:`BaseBoxes`): target of the
+ transformation, e.g., ground truth boxes.
+
+ Returns:
+ torch.Tensor: Box transformation deltas
+ """
+ bboxes = get_box_tensor(bboxes)
+ gt_bboxes = get_box_tensor(gt_bboxes)
+ assert bboxes.size(0) == gt_bboxes.size(0)
+ assert bboxes.size(-1) == gt_bboxes.size(-1) == 4
+ encoded_bboxes = bboxes2tblr(
+ bboxes, gt_bboxes, normalizer=self.normalizer)
+ return encoded_bboxes
+
+ def decode(
+ self,
+ bboxes: Union[Tensor, BaseBoxes],
+ pred_bboxes: Tensor,
+ max_shape: Optional[Union[Sequence[int], Tensor,
+ Sequence[Sequence[int]]]] = None
+ ) -> Union[Tensor, BaseBoxes]:
+ """Apply transformation `pred_bboxes` to `boxes`.
+
+ Args:
+ bboxes (torch.Tensor or :obj:`BaseBoxes`): Basic boxes.Shape
+ (B, N, 4) or (N, 4)
+ pred_bboxes (torch.Tensor): Encoded boxes with shape
+ (B, N, 4) or (N, 4)
+ max_shape (Sequence[int] or torch.Tensor or Sequence[
+ Sequence[int]],optional): Maximum bounds for boxes, specifies
+ (H, W, C) or (H, W). If bboxes shape is (B, N, 4), then
+ the max_shape should be a Sequence[Sequence[int]]
+ and the length of max_shape should also be B.
+
+ Returns:
+ Union[torch.Tensor, :obj:`BaseBoxes`]: Decoded boxes.
+ """
+ bboxes = get_box_tensor(bboxes)
+ decoded_bboxes = tblr2bboxes(
+ bboxes,
+ pred_bboxes,
+ normalizer=self.normalizer,
+ max_shape=max_shape,
+ clip_border=self.clip_border)
+
+ if self.use_box_type:
+ decoded_bboxes = HorizontalBoxes(decoded_bboxes)
+ return decoded_bboxes
+
+
+def bboxes2tblr(priors: Tensor,
+ gts: Tensor,
+ normalizer: Union[Sequence[float], float] = 4.0,
+ normalize_by_wh: bool = True) -> Tensor:
+ """Encode ground truth boxes to tblr coordinate.
+
+ It first convert the gt coordinate to tblr format,
+ (top, bottom, left, right), relative to prior box centers.
+ The tblr coordinate may be normalized by the side length of prior bboxes
+ if `normalize_by_wh` is specified as True, and it is then normalized by
+ the `normalizer` factor.
+
+ Args:
+ priors (Tensor): Prior boxes in point form
+ Shape: (num_proposals,4).
+ gts (Tensor): Coords of ground truth for each prior in point-form
+ Shape: (num_proposals, 4).
+ normalizer (Sequence[float] | float): normalization parameter of
+ encoded boxes. If it is a list, it has to have length = 4.
+ Default: 4.0
+ normalize_by_wh (bool): Whether to normalize tblr coordinate by the
+ side length (wh) of prior bboxes.
+
+ Return:
+ encoded boxes (Tensor), Shape: (num_proposals, 4)
+ """
+
+ # dist b/t match center and prior's center
+ if not isinstance(normalizer, float):
+ normalizer = torch.tensor(normalizer, device=priors.device)
+ assert len(normalizer) == 4, 'Normalizer must have length = 4'
+ assert priors.size(0) == gts.size(0)
+ prior_centers = (priors[:, 0:2] + priors[:, 2:4]) / 2
+ xmin, ymin, xmax, ymax = gts.split(1, dim=1)
+ top = prior_centers[:, 1].unsqueeze(1) - ymin
+ bottom = ymax - prior_centers[:, 1].unsqueeze(1)
+ left = prior_centers[:, 0].unsqueeze(1) - xmin
+ right = xmax - prior_centers[:, 0].unsqueeze(1)
+ loc = torch.cat((top, bottom, left, right), dim=1)
+ if normalize_by_wh:
+ # Normalize tblr by anchor width and height
+ wh = priors[:, 2:4] - priors[:, 0:2]
+ w, h = torch.split(wh, 1, dim=1)
+ loc[:, :2] /= h # tb is normalized by h
+ loc[:, 2:] /= w # lr is normalized by w
+ # Normalize tblr by the given normalization factor
+ return loc / normalizer
+
+
+def tblr2bboxes(priors: Tensor,
+ tblr: Tensor,
+ normalizer: Union[Sequence[float], float] = 4.0,
+ normalize_by_wh: bool = True,
+ max_shape: Optional[Union[Sequence[int], Tensor,
+ Sequence[Sequence[int]]]] = None,
+ clip_border: bool = True) -> Tensor:
+ """Decode tblr outputs to prediction boxes.
+
+ The process includes 3 steps: 1) De-normalize tblr coordinates by
+ multiplying it with `normalizer`; 2) De-normalize tblr coordinates by the
+ prior bbox width and height if `normalize_by_wh` is `True`; 3) Convert
+ tblr (top, bottom, left, right) pair relative to the center of priors back
+ to (xmin, ymin, xmax, ymax) coordinate.
+
+ Args:
+ priors (Tensor): Prior boxes in point form (x0, y0, x1, y1)
+ Shape: (N,4) or (B, N, 4).
+ tblr (Tensor): Coords of network output in tblr form
+ Shape: (N, 4) or (B, N, 4).
+ normalizer (Sequence[float] | float): Normalization parameter of
+ encoded boxes. By list, it represents the normalization factors at
+ tblr dims. By float, it is the unified normalization factor at all
+ dims. Default: 4.0
+ normalize_by_wh (bool): Whether the tblr coordinates have been
+ normalized by the side length (wh) of prior bboxes.
+ max_shape (Sequence[int] or torch.Tensor or Sequence[
+ Sequence[int]],optional): Maximum bounds for boxes, specifies
+ (H, W, C) or (H, W). If priors shape is (B, N, 4), then
+ the max_shape should be a Sequence[Sequence[int]]
+ and the length of max_shape should also be B.
+ clip_border (bool, optional): Whether clip the objects outside the
+ border of the image. Defaults to True.
+
+ Return:
+ encoded boxes (Tensor): Boxes with shape (N, 4) or (B, N, 4)
+ """
+ if not isinstance(normalizer, float):
+ normalizer = torch.tensor(normalizer, device=priors.device)
+ assert len(normalizer) == 4, 'Normalizer must have length = 4'
+ assert priors.size(0) == tblr.size(0)
+ if priors.ndim == 3:
+ assert priors.size(1) == tblr.size(1)
+
+ loc_decode = tblr * normalizer
+ prior_centers = (priors[..., 0:2] + priors[..., 2:4]) / 2
+ if normalize_by_wh:
+ wh = priors[..., 2:4] - priors[..., 0:2]
+ w, h = torch.split(wh, 1, dim=-1)
+ # Inplace operation with slice would failed for exporting to ONNX
+ th = h * loc_decode[..., :2] # tb
+ tw = w * loc_decode[..., 2:] # lr
+ loc_decode = torch.cat([th, tw], dim=-1)
+ # Cannot be exported using onnx when loc_decode.split(1, dim=-1)
+ top, bottom, left, right = loc_decode.split((1, 1, 1, 1), dim=-1)
+ xmin = prior_centers[..., 0].unsqueeze(-1) - left
+ xmax = prior_centers[..., 0].unsqueeze(-1) + right
+ ymin = prior_centers[..., 1].unsqueeze(-1) - top
+ ymax = prior_centers[..., 1].unsqueeze(-1) + bottom
+
+ bboxes = torch.cat((xmin, ymin, xmax, ymax), dim=-1)
+
+ if clip_border and max_shape is not None:
+ # clip bboxes with dynamic `min` and `max` for onnx
+ if torch.onnx.is_in_onnx_export():
+ from mmdet.core.export import dynamic_clip_for_onnx
+ xmin, ymin, xmax, ymax = dynamic_clip_for_onnx(
+ xmin, ymin, xmax, ymax, max_shape)
+ bboxes = torch.cat([xmin, ymin, xmax, ymax], dim=-1)
+ return bboxes
+ if not isinstance(max_shape, torch.Tensor):
+ max_shape = priors.new_tensor(max_shape)
+ max_shape = max_shape[..., :2].type_as(priors)
+ if max_shape.ndim == 2:
+ assert bboxes.ndim == 3
+ assert max_shape.size(0) == bboxes.size(0)
+
+ min_xy = priors.new_tensor(0)
+ max_xy = torch.cat([max_shape, max_shape],
+ dim=-1).flip(-1).unsqueeze(-2)
+ bboxes = torch.where(bboxes < min_xy, min_xy, bboxes)
+ bboxes = torch.where(bboxes > max_xy, max_xy, bboxes)
+
+ return bboxes
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/yolo_bbox_coder.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/yolo_bbox_coder.py
new file mode 100644
index 0000000000000000000000000000000000000000..2e1c766789bec844ff359e225435bc3b2f5dd736
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/coders/yolo_bbox_coder.py
@@ -0,0 +1,94 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Union
+
+import torch
+from torch import Tensor
+
+from mmdet.registry import TASK_UTILS
+from mmdet.structures.bbox import BaseBoxes, HorizontalBoxes, get_box_tensor
+from .base_bbox_coder import BaseBBoxCoder
+
+
+@TASK_UTILS.register_module()
+class YOLOBBoxCoder(BaseBBoxCoder):
+ """YOLO BBox coder.
+
+ Following `YOLO `_, this coder divide
+ image into grids, and encode bbox (x1, y1, x2, y2) into (cx, cy, dw, dh).
+ cx, cy in [0., 1.], denotes relative center position w.r.t the center of
+ bboxes. dw, dh are the same as :obj:`DeltaXYWHBBoxCoder`.
+
+ Args:
+ eps (float): Min value of cx, cy when encoding.
+ """
+
+ def __init__(self, eps: float = 1e-6, **kwargs):
+ super().__init__(**kwargs)
+ self.eps = eps
+
+ def encode(self, bboxes: Union[Tensor, BaseBoxes],
+ gt_bboxes: Union[Tensor, BaseBoxes],
+ stride: Union[Tensor, int]) -> Tensor:
+ """Get box regression transformation deltas that can be used to
+ transform the ``bboxes`` into the ``gt_bboxes``.
+
+ Args:
+ bboxes (torch.Tensor or :obj:`BaseBoxes`): Source boxes,
+ e.g., anchors.
+ gt_bboxes (torch.Tensor or :obj:`BaseBoxes`): Target of the
+ transformation, e.g., ground-truth boxes.
+ stride (torch.Tensor | int): Stride of bboxes.
+
+ Returns:
+ torch.Tensor: Box transformation deltas
+ """
+ bboxes = get_box_tensor(bboxes)
+ gt_bboxes = get_box_tensor(gt_bboxes)
+ assert bboxes.size(0) == gt_bboxes.size(0)
+ assert bboxes.size(-1) == gt_bboxes.size(-1) == 4
+ x_center_gt = (gt_bboxes[..., 0] + gt_bboxes[..., 2]) * 0.5
+ y_center_gt = (gt_bboxes[..., 1] + gt_bboxes[..., 3]) * 0.5
+ w_gt = gt_bboxes[..., 2] - gt_bboxes[..., 0]
+ h_gt = gt_bboxes[..., 3] - gt_bboxes[..., 1]
+ x_center = (bboxes[..., 0] + bboxes[..., 2]) * 0.5
+ y_center = (bboxes[..., 1] + bboxes[..., 3]) * 0.5
+ w = bboxes[..., 2] - bboxes[..., 0]
+ h = bboxes[..., 3] - bboxes[..., 1]
+ w_target = torch.log((w_gt / w).clamp(min=self.eps))
+ h_target = torch.log((h_gt / h).clamp(min=self.eps))
+ x_center_target = ((x_center_gt - x_center) / stride + 0.5).clamp(
+ self.eps, 1 - self.eps)
+ y_center_target = ((y_center_gt - y_center) / stride + 0.5).clamp(
+ self.eps, 1 - self.eps)
+ encoded_bboxes = torch.stack(
+ [x_center_target, y_center_target, w_target, h_target], dim=-1)
+ return encoded_bboxes
+
+ def decode(self, bboxes: Union[Tensor, BaseBoxes], pred_bboxes: Tensor,
+ stride: Union[Tensor, int]) -> Union[Tensor, BaseBoxes]:
+ """Apply transformation `pred_bboxes` to `boxes`.
+
+ Args:
+ boxes (torch.Tensor or :obj:`BaseBoxes`): Basic boxes,
+ e.g. anchors.
+ pred_bboxes (torch.Tensor): Encoded boxes with shape
+ stride (torch.Tensor | int): Strides of bboxes.
+
+ Returns:
+ Union[torch.Tensor, :obj:`BaseBoxes`]: Decoded boxes.
+ """
+ bboxes = get_box_tensor(bboxes)
+ assert pred_bboxes.size(-1) == bboxes.size(-1) == 4
+ xy_centers = (bboxes[..., :2] + bboxes[..., 2:]) * 0.5 + (
+ pred_bboxes[..., :2] - 0.5) * stride
+ whs = (bboxes[..., 2:] -
+ bboxes[..., :2]) * 0.5 * pred_bboxes[..., 2:].exp()
+ decoded_bboxes = torch.stack(
+ (xy_centers[..., 0] - whs[..., 0], xy_centers[..., 1] -
+ whs[..., 1], xy_centers[..., 0] + whs[..., 0],
+ xy_centers[..., 1] + whs[..., 1]),
+ dim=-1)
+
+ if self.use_box_type:
+ decoded_bboxes = HorizontalBoxes(decoded_bboxes)
+ return decoded_bboxes
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/prior_generators/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/prior_generators/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..7795e98ca77bb5ffc77ff1da848130717d8f85a6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/prior_generators/__init__.py
@@ -0,0 +1,11 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .anchor_generator import (AnchorGenerator, LegacyAnchorGenerator,
+ SSDAnchorGenerator, YOLOAnchorGenerator)
+from .point_generator import MlvlPointGenerator, PointGenerator
+from .utils import anchor_inside_flags, calc_region
+
+__all__ = [
+ 'AnchorGenerator', 'LegacyAnchorGenerator', 'anchor_inside_flags',
+ 'PointGenerator', 'calc_region', 'YOLOAnchorGenerator',
+ 'MlvlPointGenerator', 'SSDAnchorGenerator'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/prior_generators/anchor_generator.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/prior_generators/anchor_generator.py
new file mode 100644
index 0000000000000000000000000000000000000000..2757697ce2283ec8b46ba89325e63fad0be4a7e8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/prior_generators/anchor_generator.py
@@ -0,0 +1,848 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+from typing import List, Optional, Tuple, Union
+
+import numpy as np
+import torch
+from mmengine.utils import is_tuple_of
+from torch import Tensor
+from torch.nn.modules.utils import _pair
+
+from mmdet.registry import TASK_UTILS
+from mmdet.structures.bbox import HorizontalBoxes
+
+DeviceType = Union[str, torch.device]
+
+
+@TASK_UTILS.register_module()
+class AnchorGenerator:
+ """Standard anchor generator for 2D anchor-based detectors.
+
+ Args:
+ strides (list[int] | list[tuple[int, int]]): Strides of anchors
+ in multiple feature levels in order (w, h).
+ ratios (list[float]): The list of ratios between the height and width
+ of anchors in a single level.
+ scales (list[int], Optional): Anchor scales for anchors
+ in a single level. It cannot be set at the same time
+ if `octave_base_scale` and `scales_per_octave` are set.
+ base_sizes (list[int], Optional): The basic sizes
+ of anchors in multiple levels.
+ If None is given, strides will be used as base_sizes.
+ (If strides are non square, the shortest stride is taken.)
+ scale_major (bool): Whether to multiply scales first when generating
+ base anchors. If true, the anchors in the same row will have the
+ same scales. By default it is True in V2.0
+ octave_base_scale (int, Optional): The base scale of octave.
+ scales_per_octave (int, Optional): Number of scales for each octave.
+ `octave_base_scale` and `scales_per_octave` are usually used in
+ retinanet and the `scales` should be None when they are set.
+ centers (list[tuple[float]], Optional): The centers of the anchor
+ relative to the feature grid center in multiple feature levels.
+ By default it is set to be None and not used. If a list of tuple of
+ float is given, they will be used to shift the centers of anchors.
+ center_offset (float): The offset of center in proportion to anchors'
+ width and height. By default it is 0 in V2.0.
+ use_box_type (bool): Whether to warp anchors with the box type data
+ structure. Defaults to False.
+
+ Examples:
+ >>> from mmdet.models.task_modules.
+ ... prior_generators import AnchorGenerator
+ >>> self = AnchorGenerator([16], [1.], [1.], [9])
+ >>> all_anchors = self.grid_priors([(2, 2)], device='cpu')
+ >>> print(all_anchors)
+ [tensor([[-4.5000, -4.5000, 4.5000, 4.5000],
+ [11.5000, -4.5000, 20.5000, 4.5000],
+ [-4.5000, 11.5000, 4.5000, 20.5000],
+ [11.5000, 11.5000, 20.5000, 20.5000]])]
+ >>> self = AnchorGenerator([16, 32], [1.], [1.], [9, 18])
+ >>> all_anchors = self.grid_priors([(2, 2), (1, 1)], device='cpu')
+ >>> print(all_anchors)
+ [tensor([[-4.5000, -4.5000, 4.5000, 4.5000],
+ [11.5000, -4.5000, 20.5000, 4.5000],
+ [-4.5000, 11.5000, 4.5000, 20.5000],
+ [11.5000, 11.5000, 20.5000, 20.5000]]), \
+ tensor([[-9., -9., 9., 9.]])]
+ """
+
+ def __init__(self,
+ strides: Union[List[int], List[Tuple[int, int]]],
+ ratios: List[float],
+ scales: Optional[List[int]] = None,
+ base_sizes: Optional[List[int]] = None,
+ scale_major: bool = True,
+ octave_base_scale: Optional[int] = None,
+ scales_per_octave: Optional[int] = None,
+ centers: Optional[List[Tuple[float, float]]] = None,
+ center_offset: float = 0.,
+ use_box_type: bool = False) -> None:
+ # check center and center_offset
+ if center_offset != 0:
+ assert centers is None, 'center cannot be set when center_offset' \
+ f'!=0, {centers} is given.'
+ if not (0 <= center_offset <= 1):
+ raise ValueError('center_offset should be in range [0, 1], '
+ f'{center_offset} is given.')
+ if centers is not None:
+ assert len(centers) == len(strides), \
+ 'The number of strides should be the same as centers, got ' \
+ f'{strides} and {centers}'
+
+ # calculate base sizes of anchors
+ self.strides = [_pair(stride) for stride in strides]
+ self.base_sizes = [min(stride) for stride in self.strides
+ ] if base_sizes is None else base_sizes
+ assert len(self.base_sizes) == len(self.strides), \
+ 'The number of strides should be the same as base sizes, got ' \
+ f'{self.strides} and {self.base_sizes}'
+
+ # calculate scales of anchors
+ assert ((octave_base_scale is not None
+ and scales_per_octave is not None) ^ (scales is not None)), \
+ 'scales and octave_base_scale with scales_per_octave cannot' \
+ ' be set at the same time'
+ if scales is not None:
+ self.scales = torch.Tensor(scales)
+ elif octave_base_scale is not None and scales_per_octave is not None:
+ octave_scales = np.array(
+ [2**(i / scales_per_octave) for i in range(scales_per_octave)])
+ scales = octave_scales * octave_base_scale
+ self.scales = torch.Tensor(scales)
+ else:
+ raise ValueError('Either scales or octave_base_scale with '
+ 'scales_per_octave should be set')
+
+ self.octave_base_scale = octave_base_scale
+ self.scales_per_octave = scales_per_octave
+ self.ratios = torch.Tensor(ratios)
+ self.scale_major = scale_major
+ self.centers = centers
+ self.center_offset = center_offset
+ self.base_anchors = self.gen_base_anchors()
+ self.use_box_type = use_box_type
+
+ @property
+ def num_base_anchors(self) -> List[int]:
+ """list[int]: total number of base anchors in a feature grid"""
+ return self.num_base_priors
+
+ @property
+ def num_base_priors(self) -> List[int]:
+ """list[int]: The number of priors (anchors) at a point
+ on the feature grid"""
+ return [base_anchors.size(0) for base_anchors in self.base_anchors]
+
+ @property
+ def num_levels(self) -> int:
+ """int: number of feature levels that the generator will be applied"""
+ return len(self.strides)
+
+ def gen_base_anchors(self) -> List[Tensor]:
+ """Generate base anchors.
+
+ Returns:
+ list(torch.Tensor): Base anchors of a feature grid in multiple \
+ feature levels.
+ """
+ multi_level_base_anchors = []
+ for i, base_size in enumerate(self.base_sizes):
+ center = None
+ if self.centers is not None:
+ center = self.centers[i]
+ multi_level_base_anchors.append(
+ self.gen_single_level_base_anchors(
+ base_size,
+ scales=self.scales,
+ ratios=self.ratios,
+ center=center))
+ return multi_level_base_anchors
+
+ def gen_single_level_base_anchors(self,
+ base_size: Union[int, float],
+ scales: Tensor,
+ ratios: Tensor,
+ center: Optional[Tuple[float]] = None) \
+ -> Tensor:
+ """Generate base anchors of a single level.
+
+ Args:
+ base_size (int | float): Basic size of an anchor.
+ scales (torch.Tensor): Scales of the anchor.
+ ratios (torch.Tensor): The ratio between the height
+ and width of anchors in a single level.
+ center (tuple[float], optional): The center of the base anchor
+ related to a single feature grid. Defaults to None.
+
+ Returns:
+ torch.Tensor: Anchors in a single-level feature maps.
+ """
+ w = base_size
+ h = base_size
+ if center is None:
+ x_center = self.center_offset * w
+ y_center = self.center_offset * h
+ else:
+ x_center, y_center = center
+
+ h_ratios = torch.sqrt(ratios)
+ w_ratios = 1 / h_ratios
+ if self.scale_major:
+ ws = (w * w_ratios[:, None] * scales[None, :]).view(-1)
+ hs = (h * h_ratios[:, None] * scales[None, :]).view(-1)
+ else:
+ ws = (w * scales[:, None] * w_ratios[None, :]).view(-1)
+ hs = (h * scales[:, None] * h_ratios[None, :]).view(-1)
+
+ # use float anchor and the anchor's center is aligned with the
+ # pixel center
+ base_anchors = [
+ x_center - 0.5 * ws, y_center - 0.5 * hs, x_center + 0.5 * ws,
+ y_center + 0.5 * hs
+ ]
+ base_anchors = torch.stack(base_anchors, dim=-1)
+
+ return base_anchors
+
+ def _meshgrid(self,
+ x: Tensor,
+ y: Tensor,
+ row_major: bool = True) -> Tuple[Tensor]:
+ """Generate mesh grid of x and y.
+
+ Args:
+ x (torch.Tensor): Grids of x dimension.
+ y (torch.Tensor): Grids of y dimension.
+ row_major (bool): Whether to return y grids first.
+ Defaults to True.
+
+ Returns:
+ tuple[torch.Tensor]: The mesh grids of x and y.
+ """
+ # use shape instead of len to keep tracing while exporting to onnx
+ xx = x.repeat(y.shape[0])
+ yy = y.view(-1, 1).repeat(1, x.shape[0]).view(-1)
+ if row_major:
+ return xx, yy
+ else:
+ return yy, xx
+
+ def grid_priors(self,
+ featmap_sizes: List[Tuple],
+ dtype: torch.dtype = torch.float32,
+ device: DeviceType = 'cuda') -> List[Tensor]:
+ """Generate grid anchors in multiple feature levels.
+
+ Args:
+ featmap_sizes (list[tuple]): List of feature map sizes in
+ multiple feature levels.
+ dtype (:obj:`torch.dtype`): Dtype of priors.
+ Defaults to torch.float32.
+ device (str | torch.device): The device where the anchors
+ will be put on.
+
+ Return:
+ list[torch.Tensor]: Anchors in multiple feature levels. \
+ The sizes of each tensor should be [N, 4], where \
+ N = width * height * num_base_anchors, width and height \
+ are the sizes of the corresponding feature level, \
+ num_base_anchors is the number of anchors for that level.
+ """
+ assert self.num_levels == len(featmap_sizes)
+ multi_level_anchors = []
+ for i in range(self.num_levels):
+ anchors = self.single_level_grid_priors(
+ featmap_sizes[i], level_idx=i, dtype=dtype, device=device)
+ multi_level_anchors.append(anchors)
+ return multi_level_anchors
+
+ def single_level_grid_priors(self,
+ featmap_size: Tuple[int, int],
+ level_idx: int,
+ dtype: torch.dtype = torch.float32,
+ device: DeviceType = 'cuda') -> Tensor:
+ """Generate grid anchors of a single level.
+
+ Note:
+ This function is usually called by method ``self.grid_priors``.
+
+ Args:
+ featmap_size (tuple[int, int]): Size of the feature maps.
+ level_idx (int): The index of corresponding feature map level.
+ dtype (obj:`torch.dtype`): Date type of points.Defaults to
+ ``torch.float32``.
+ device (str | torch.device): The device the tensor will be put on.
+ Defaults to 'cuda'.
+
+ Returns:
+ torch.Tensor: Anchors in the overall feature maps.
+ """
+
+ base_anchors = self.base_anchors[level_idx].to(device).to(dtype)
+ feat_h, feat_w = featmap_size
+ stride_w, stride_h = self.strides[level_idx]
+ # First create Range with the default dtype, than convert to
+ # target `dtype` for onnx exporting.
+ shift_x = torch.arange(0, feat_w, device=device).to(dtype) * stride_w
+ shift_y = torch.arange(0, feat_h, device=device).to(dtype) * stride_h
+
+ shift_xx, shift_yy = self._meshgrid(shift_x, shift_y)
+ shifts = torch.stack([shift_xx, shift_yy, shift_xx, shift_yy], dim=-1)
+ # first feat_w elements correspond to the first row of shifts
+ # add A anchors (1, A, 4) to K shifts (K, 1, 4) to get
+ # shifted anchors (K, A, 4), reshape to (K*A, 4)
+
+ all_anchors = base_anchors[None, :, :] + shifts[:, None, :]
+ all_anchors = all_anchors.view(-1, 4)
+ # first A rows correspond to A anchors of (0, 0) in feature map,
+ # then (0, 1), (0, 2), ...
+ if self.use_box_type:
+ all_anchors = HorizontalBoxes(all_anchors)
+ return all_anchors
+
+ def sparse_priors(self,
+ prior_idxs: Tensor,
+ featmap_size: Tuple[int, int],
+ level_idx: int,
+ dtype: torch.dtype = torch.float32,
+ device: DeviceType = 'cuda') -> Tensor:
+ """Generate sparse anchors according to the ``prior_idxs``.
+
+ Args:
+ prior_idxs (Tensor): The index of corresponding anchors
+ in the feature map.
+ featmap_size (tuple[int, int]): feature map size arrange as (h, w).
+ level_idx (int): The level index of corresponding feature
+ map.
+ dtype (obj:`torch.dtype`): Date type of points.Defaults to
+ ``torch.float32``.
+ device (str | torch.device): The device where the points is
+ located.
+ Returns:
+ Tensor: Anchor with shape (N, 4), N should be equal to
+ the length of ``prior_idxs``.
+ """
+
+ height, width = featmap_size
+ num_base_anchors = self.num_base_anchors[level_idx]
+ base_anchor_id = prior_idxs % num_base_anchors
+ x = (prior_idxs //
+ num_base_anchors) % width * self.strides[level_idx][0]
+ y = (prior_idxs // width //
+ num_base_anchors) % height * self.strides[level_idx][1]
+ priors = torch.stack([x, y, x, y], 1).to(dtype).to(device) + \
+ self.base_anchors[level_idx][base_anchor_id, :].to(device)
+
+ return priors
+
+ def grid_anchors(self,
+ featmap_sizes: List[Tuple],
+ device: DeviceType = 'cuda') -> List[Tensor]:
+ """Generate grid anchors in multiple feature levels.
+
+ Args:
+ featmap_sizes (list[tuple]): List of feature map sizes in
+ multiple feature levels.
+ device (str | torch.device): Device where the anchors will be
+ put on.
+
+ Return:
+ list[torch.Tensor]: Anchors in multiple feature levels. \
+ The sizes of each tensor should be [N, 4], where \
+ N = width * height * num_base_anchors, width and height \
+ are the sizes of the corresponding feature level, \
+ num_base_anchors is the number of anchors for that level.
+ """
+ warnings.warn('``grid_anchors`` would be deprecated soon. '
+ 'Please use ``grid_priors`` ')
+
+ assert self.num_levels == len(featmap_sizes)
+ multi_level_anchors = []
+ for i in range(self.num_levels):
+ anchors = self.single_level_grid_anchors(
+ self.base_anchors[i].to(device),
+ featmap_sizes[i],
+ self.strides[i],
+ device=device)
+ multi_level_anchors.append(anchors)
+ return multi_level_anchors
+
+ def single_level_grid_anchors(self,
+ base_anchors: Tensor,
+ featmap_size: Tuple[int, int],
+ stride: Tuple[int, int] = (16, 16),
+ device: DeviceType = 'cuda') -> Tensor:
+ """Generate grid anchors of a single level.
+
+ Note:
+ This function is usually called by method ``self.grid_anchors``.
+
+ Args:
+ base_anchors (torch.Tensor): The base anchors of a feature grid.
+ featmap_size (tuple[int]): Size of the feature maps.
+ stride (tuple[int, int]): Stride of the feature map in order
+ (w, h). Defaults to (16, 16).
+ device (str | torch.device): Device the tensor will be put on.
+ Defaults to 'cuda'.
+
+ Returns:
+ torch.Tensor: Anchors in the overall feature maps.
+ """
+
+ warnings.warn(
+ '``single_level_grid_anchors`` would be deprecated soon. '
+ 'Please use ``single_level_grid_priors`` ')
+
+ # keep featmap_size as Tensor instead of int, so that we
+ # can convert to ONNX correctly
+ feat_h, feat_w = featmap_size
+ shift_x = torch.arange(0, feat_w, device=device) * stride[0]
+ shift_y = torch.arange(0, feat_h, device=device) * stride[1]
+
+ shift_xx, shift_yy = self._meshgrid(shift_x, shift_y)
+ shifts = torch.stack([shift_xx, shift_yy, shift_xx, shift_yy], dim=-1)
+ shifts = shifts.type_as(base_anchors)
+ # first feat_w elements correspond to the first row of shifts
+ # add A anchors (1, A, 4) to K shifts (K, 1, 4) to get
+ # shifted anchors (K, A, 4), reshape to (K*A, 4)
+
+ all_anchors = base_anchors[None, :, :] + shifts[:, None, :]
+ all_anchors = all_anchors.view(-1, 4)
+ # first A rows correspond to A anchors of (0, 0) in feature map,
+ # then (0, 1), (0, 2), ...
+ return all_anchors
+
+ def valid_flags(self,
+ featmap_sizes: List[Tuple[int, int]],
+ pad_shape: Tuple,
+ device: DeviceType = 'cuda') -> List[Tensor]:
+ """Generate valid flags of anchors in multiple feature levels.
+
+ Args:
+ featmap_sizes (list(tuple[int, int])): List of feature map sizes in
+ multiple feature levels.
+ pad_shape (tuple): The padded shape of the image.
+ device (str | torch.device): Device where the anchors will be
+ put on.
+
+ Return:
+ list(torch.Tensor): Valid flags of anchors in multiple levels.
+ """
+ assert self.num_levels == len(featmap_sizes)
+ multi_level_flags = []
+ for i in range(self.num_levels):
+ anchor_stride = self.strides[i]
+ feat_h, feat_w = featmap_sizes[i]
+ h, w = pad_shape[:2]
+ valid_feat_h = min(int(np.ceil(h / anchor_stride[1])), feat_h)
+ valid_feat_w = min(int(np.ceil(w / anchor_stride[0])), feat_w)
+ flags = self.single_level_valid_flags((feat_h, feat_w),
+ (valid_feat_h, valid_feat_w),
+ self.num_base_anchors[i],
+ device=device)
+ multi_level_flags.append(flags)
+ return multi_level_flags
+
+ def single_level_valid_flags(self,
+ featmap_size: Tuple[int, int],
+ valid_size: Tuple[int, int],
+ num_base_anchors: int,
+ device: DeviceType = 'cuda') -> Tensor:
+ """Generate the valid flags of anchor in a single feature map.
+
+ Args:
+ featmap_size (tuple[int]): The size of feature maps, arrange
+ as (h, w).
+ valid_size (tuple[int]): The valid size of the feature maps.
+ num_base_anchors (int): The number of base anchors.
+ device (str | torch.device): Device where the flags will be put on.
+ Defaults to 'cuda'.
+
+ Returns:
+ torch.Tensor: The valid flags of each anchor in a single level \
+ feature map.
+ """
+ feat_h, feat_w = featmap_size
+ valid_h, valid_w = valid_size
+ assert valid_h <= feat_h and valid_w <= feat_w
+ valid_x = torch.zeros(feat_w, dtype=torch.bool, device=device)
+ valid_y = torch.zeros(feat_h, dtype=torch.bool, device=device)
+ valid_x[:valid_w] = 1
+ valid_y[:valid_h] = 1
+ valid_xx, valid_yy = self._meshgrid(valid_x, valid_y)
+ valid = valid_xx & valid_yy
+ valid = valid[:, None].expand(valid.size(0),
+ num_base_anchors).contiguous().view(-1)
+ return valid
+
+ def __repr__(self) -> str:
+ """str: a string that describes the module"""
+ indent_str = ' '
+ repr_str = self.__class__.__name__ + '(\n'
+ repr_str += f'{indent_str}strides={self.strides},\n'
+ repr_str += f'{indent_str}ratios={self.ratios},\n'
+ repr_str += f'{indent_str}scales={self.scales},\n'
+ repr_str += f'{indent_str}base_sizes={self.base_sizes},\n'
+ repr_str += f'{indent_str}scale_major={self.scale_major},\n'
+ repr_str += f'{indent_str}octave_base_scale='
+ repr_str += f'{self.octave_base_scale},\n'
+ repr_str += f'{indent_str}scales_per_octave='
+ repr_str += f'{self.scales_per_octave},\n'
+ repr_str += f'{indent_str}num_levels={self.num_levels}\n'
+ repr_str += f'{indent_str}centers={self.centers},\n'
+ repr_str += f'{indent_str}center_offset={self.center_offset})'
+ return repr_str
+
+
+@TASK_UTILS.register_module()
+class SSDAnchorGenerator(AnchorGenerator):
+ """Anchor generator for SSD.
+
+ Args:
+ strides (list[int] | list[tuple[int, int]]): Strides of anchors
+ in multiple feature levels.
+ ratios (list[float]): The list of ratios between the height and width
+ of anchors in a single level.
+ min_sizes (list[float]): The list of minimum anchor sizes on each
+ level.
+ max_sizes (list[float]): The list of maximum anchor sizes on each
+ level.
+ basesize_ratio_range (tuple(float)): Ratio range of anchors. Being
+ used when not setting min_sizes and max_sizes.
+ input_size (int): Size of feature map, 300 for SSD300, 512 for
+ SSD512. Being used when not setting min_sizes and max_sizes.
+ scale_major (bool): Whether to multiply scales first when generating
+ base anchors. If true, the anchors in the same row will have the
+ same scales. It is always set to be False in SSD.
+ use_box_type (bool): Whether to warp anchors with the box type data
+ structure. Defaults to False.
+ """
+
+ def __init__(self,
+ strides: Union[List[int], List[Tuple[int, int]]],
+ ratios: List[float],
+ min_sizes: Optional[List[float]] = None,
+ max_sizes: Optional[List[float]] = None,
+ basesize_ratio_range: Tuple[float] = (0.15, 0.9),
+ input_size: int = 300,
+ scale_major: bool = True,
+ use_box_type: bool = False) -> None:
+ assert len(strides) == len(ratios)
+ assert not (min_sizes is None) ^ (max_sizes is None)
+ self.strides = [_pair(stride) for stride in strides]
+ self.centers = [(stride[0] / 2., stride[1] / 2.)
+ for stride in self.strides]
+
+ if min_sizes is None and max_sizes is None:
+ # use hard code to generate SSD anchors
+ self.input_size = input_size
+ assert is_tuple_of(basesize_ratio_range, float)
+ self.basesize_ratio_range = basesize_ratio_range
+ # calculate anchor ratios and sizes
+ min_ratio, max_ratio = basesize_ratio_range
+ min_ratio = int(min_ratio * 100)
+ max_ratio = int(max_ratio * 100)
+ step = int(np.floor(max_ratio - min_ratio) / (self.num_levels - 2))
+ min_sizes = []
+ max_sizes = []
+ for ratio in range(int(min_ratio), int(max_ratio) + 1, step):
+ min_sizes.append(int(self.input_size * ratio / 100))
+ max_sizes.append(int(self.input_size * (ratio + step) / 100))
+ if self.input_size == 300:
+ if basesize_ratio_range[0] == 0.15: # SSD300 COCO
+ min_sizes.insert(0, int(self.input_size * 7 / 100))
+ max_sizes.insert(0, int(self.input_size * 15 / 100))
+ elif basesize_ratio_range[0] == 0.2: # SSD300 VOC
+ min_sizes.insert(0, int(self.input_size * 10 / 100))
+ max_sizes.insert(0, int(self.input_size * 20 / 100))
+ else:
+ raise ValueError(
+ 'basesize_ratio_range[0] should be either 0.15'
+ 'or 0.2 when input_size is 300, got '
+ f'{basesize_ratio_range[0]}.')
+ elif self.input_size == 512:
+ if basesize_ratio_range[0] == 0.1: # SSD512 COCO
+ min_sizes.insert(0, int(self.input_size * 4 / 100))
+ max_sizes.insert(0, int(self.input_size * 10 / 100))
+ elif basesize_ratio_range[0] == 0.15: # SSD512 VOC
+ min_sizes.insert(0, int(self.input_size * 7 / 100))
+ max_sizes.insert(0, int(self.input_size * 15 / 100))
+ else:
+ raise ValueError(
+ 'When not setting min_sizes and max_sizes,'
+ 'basesize_ratio_range[0] should be either 0.1'
+ 'or 0.15 when input_size is 512, got'
+ f' {basesize_ratio_range[0]}.')
+ else:
+ raise ValueError(
+ 'Only support 300 or 512 in SSDAnchorGenerator when '
+ 'not setting min_sizes and max_sizes, '
+ f'got {self.input_size}.')
+
+ assert len(min_sizes) == len(max_sizes) == len(strides)
+
+ anchor_ratios = []
+ anchor_scales = []
+ for k in range(len(self.strides)):
+ scales = [1., np.sqrt(max_sizes[k] / min_sizes[k])]
+ anchor_ratio = [1.]
+ for r in ratios[k]:
+ anchor_ratio += [1 / r, r] # 4 or 6 ratio
+ anchor_ratios.append(torch.Tensor(anchor_ratio))
+ anchor_scales.append(torch.Tensor(scales))
+
+ self.base_sizes = min_sizes
+ self.scales = anchor_scales
+ self.ratios = anchor_ratios
+ self.scale_major = scale_major
+ self.center_offset = 0
+ self.base_anchors = self.gen_base_anchors()
+ self.use_box_type = use_box_type
+
+ def gen_base_anchors(self) -> List[Tensor]:
+ """Generate base anchors.
+
+ Returns:
+ list(torch.Tensor): Base anchors of a feature grid in multiple \
+ feature levels.
+ """
+ multi_level_base_anchors = []
+ for i, base_size in enumerate(self.base_sizes):
+ base_anchors = self.gen_single_level_base_anchors(
+ base_size,
+ scales=self.scales[i],
+ ratios=self.ratios[i],
+ center=self.centers[i])
+ indices = list(range(len(self.ratios[i])))
+ indices.insert(1, len(indices))
+ base_anchors = torch.index_select(base_anchors, 0,
+ torch.LongTensor(indices))
+ multi_level_base_anchors.append(base_anchors)
+ return multi_level_base_anchors
+
+ def __repr__(self) -> str:
+ """str: a string that describes the module"""
+ indent_str = ' '
+ repr_str = self.__class__.__name__ + '(\n'
+ repr_str += f'{indent_str}strides={self.strides},\n'
+ repr_str += f'{indent_str}scales={self.scales},\n'
+ repr_str += f'{indent_str}scale_major={self.scale_major},\n'
+ repr_str += f'{indent_str}input_size={self.input_size},\n'
+ repr_str += f'{indent_str}scales={self.scales},\n'
+ repr_str += f'{indent_str}ratios={self.ratios},\n'
+ repr_str += f'{indent_str}num_levels={self.num_levels},\n'
+ repr_str += f'{indent_str}base_sizes={self.base_sizes},\n'
+ repr_str += f'{indent_str}basesize_ratio_range='
+ repr_str += f'{self.basesize_ratio_range})'
+ return repr_str
+
+
+@TASK_UTILS.register_module()
+class LegacyAnchorGenerator(AnchorGenerator):
+ """Legacy anchor generator used in MMDetection V1.x.
+
+ Note:
+ Difference to the V2.0 anchor generator:
+
+ 1. The center offset of V1.x anchors are set to be 0.5 rather than 0.
+ 2. The width/height are minused by 1 when calculating the anchors' \
+ centers and corners to meet the V1.x coordinate system.
+ 3. The anchors' corners are quantized.
+
+ Args:
+ strides (list[int] | list[tuple[int]]): Strides of anchors
+ in multiple feature levels.
+ ratios (list[float]): The list of ratios between the height and width
+ of anchors in a single level.
+ scales (list[int] | None): Anchor scales for anchors in a single level.
+ It cannot be set at the same time if `octave_base_scale` and
+ `scales_per_octave` are set.
+ base_sizes (list[int]): The basic sizes of anchors in multiple levels.
+ If None is given, strides will be used to generate base_sizes.
+ scale_major (bool): Whether to multiply scales first when generating
+ base anchors. If true, the anchors in the same row will have the
+ same scales. By default it is True in V2.0
+ octave_base_scale (int): The base scale of octave.
+ scales_per_octave (int): Number of scales for each octave.
+ `octave_base_scale` and `scales_per_octave` are usually used in
+ retinanet and the `scales` should be None when they are set.
+ centers (list[tuple[float, float]] | None): The centers of the anchor
+ relative to the feature grid center in multiple feature levels.
+ By default it is set to be None and not used. It a list of float
+ is given, this list will be used to shift the centers of anchors.
+ center_offset (float): The offset of center in proportion to anchors'
+ width and height. By default it is 0.5 in V2.0 but it should be 0.5
+ in v1.x models.
+ use_box_type (bool): Whether to warp anchors with the box type data
+ structure. Defaults to False.
+
+ Examples:
+ >>> from mmdet.models.task_modules.
+ ... prior_generators import LegacyAnchorGenerator
+ >>> self = LegacyAnchorGenerator(
+ >>> [16], [1.], [1.], [9], center_offset=0.5)
+ >>> all_anchors = self.grid_anchors(((2, 2),), device='cpu')
+ >>> print(all_anchors)
+ [tensor([[ 0., 0., 8., 8.],
+ [16., 0., 24., 8.],
+ [ 0., 16., 8., 24.],
+ [16., 16., 24., 24.]])]
+ """
+
+ def gen_single_level_base_anchors(self,
+ base_size: Union[int, float],
+ scales: Tensor,
+ ratios: Tensor,
+ center: Optional[Tuple[float]] = None) \
+ -> Tensor:
+ """Generate base anchors of a single level.
+
+ Note:
+ The width/height of anchors are minused by 1 when calculating \
+ the centers and corners to meet the V1.x coordinate system.
+
+ Args:
+ base_size (int | float): Basic size of an anchor.
+ scales (torch.Tensor): Scales of the anchor.
+ ratios (torch.Tensor): The ratio between the height.
+ and width of anchors in a single level.
+ center (tuple[float], optional): The center of the base anchor
+ related to a single feature grid. Defaults to None.
+
+ Returns:
+ torch.Tensor: Anchors in a single-level feature map.
+ """
+ w = base_size
+ h = base_size
+ if center is None:
+ x_center = self.center_offset * (w - 1)
+ y_center = self.center_offset * (h - 1)
+ else:
+ x_center, y_center = center
+
+ h_ratios = torch.sqrt(ratios)
+ w_ratios = 1 / h_ratios
+ if self.scale_major:
+ ws = (w * w_ratios[:, None] * scales[None, :]).view(-1)
+ hs = (h * h_ratios[:, None] * scales[None, :]).view(-1)
+ else:
+ ws = (w * scales[:, None] * w_ratios[None, :]).view(-1)
+ hs = (h * scales[:, None] * h_ratios[None, :]).view(-1)
+
+ # use float anchor and the anchor's center is aligned with the
+ # pixel center
+ base_anchors = [
+ x_center - 0.5 * (ws - 1), y_center - 0.5 * (hs - 1),
+ x_center + 0.5 * (ws - 1), y_center + 0.5 * (hs - 1)
+ ]
+ base_anchors = torch.stack(base_anchors, dim=-1).round()
+
+ return base_anchors
+
+
+@TASK_UTILS.register_module()
+class LegacySSDAnchorGenerator(SSDAnchorGenerator, LegacyAnchorGenerator):
+ """Legacy anchor generator used in MMDetection V1.x.
+
+ The difference between `LegacySSDAnchorGenerator` and `SSDAnchorGenerator`
+ can be found in `LegacyAnchorGenerator`.
+ """
+
+ def __init__(self,
+ strides: Union[List[int], List[Tuple[int, int]]],
+ ratios: List[float],
+ basesize_ratio_range: Tuple[float],
+ input_size: int = 300,
+ scale_major: bool = True,
+ use_box_type: bool = False) -> None:
+ super(LegacySSDAnchorGenerator, self).__init__(
+ strides=strides,
+ ratios=ratios,
+ basesize_ratio_range=basesize_ratio_range,
+ input_size=input_size,
+ scale_major=scale_major,
+ use_box_type=use_box_type)
+ self.centers = [((stride - 1) / 2., (stride - 1) / 2.)
+ for stride in strides]
+ self.base_anchors = self.gen_base_anchors()
+
+
+@TASK_UTILS.register_module()
+class YOLOAnchorGenerator(AnchorGenerator):
+ """Anchor generator for YOLO.
+
+ Args:
+ strides (list[int] | list[tuple[int, int]]): Strides of anchors
+ in multiple feature levels.
+ base_sizes (list[list[tuple[int, int]]]): The basic sizes
+ of anchors in multiple levels.
+ """
+
+ def __init__(self,
+ strides: Union[List[int], List[Tuple[int, int]]],
+ base_sizes: List[List[Tuple[int, int]]],
+ use_box_type: bool = False) -> None:
+ self.strides = [_pair(stride) for stride in strides]
+ self.centers = [(stride[0] / 2., stride[1] / 2.)
+ for stride in self.strides]
+ self.base_sizes = []
+ num_anchor_per_level = len(base_sizes[0])
+ for base_sizes_per_level in base_sizes:
+ assert num_anchor_per_level == len(base_sizes_per_level)
+ self.base_sizes.append(
+ [_pair(base_size) for base_size in base_sizes_per_level])
+ self.base_anchors = self.gen_base_anchors()
+ self.use_box_type = use_box_type
+
+ @property
+ def num_levels(self) -> int:
+ """int: number of feature levels that the generator will be applied"""
+ return len(self.base_sizes)
+
+ def gen_base_anchors(self) -> List[Tensor]:
+ """Generate base anchors.
+
+ Returns:
+ list(torch.Tensor): Base anchors of a feature grid in multiple \
+ feature levels.
+ """
+ multi_level_base_anchors = []
+ for i, base_sizes_per_level in enumerate(self.base_sizes):
+ center = None
+ if self.centers is not None:
+ center = self.centers[i]
+ multi_level_base_anchors.append(
+ self.gen_single_level_base_anchors(base_sizes_per_level,
+ center))
+ return multi_level_base_anchors
+
+ def gen_single_level_base_anchors(self,
+ base_sizes_per_level: List[Tuple[int]],
+ center: Optional[Tuple[float]] = None) \
+ -> Tensor:
+ """Generate base anchors of a single level.
+
+ Args:
+ base_sizes_per_level (list[tuple[int]]): Basic sizes of
+ anchors.
+ center (tuple[float], optional): The center of the base anchor
+ related to a single feature grid. Defaults to None.
+
+ Returns:
+ torch.Tensor: Anchors in a single-level feature maps.
+ """
+ x_center, y_center = center
+ base_anchors = []
+ for base_size in base_sizes_per_level:
+ w, h = base_size
+
+ # use float anchor and the anchor's center is aligned with the
+ # pixel center
+ base_anchor = torch.Tensor([
+ x_center - 0.5 * w, y_center - 0.5 * h, x_center + 0.5 * w,
+ y_center + 0.5 * h
+ ])
+ base_anchors.append(base_anchor)
+ base_anchors = torch.stack(base_anchors, dim=0)
+
+ return base_anchors
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/prior_generators/point_generator.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/prior_generators/point_generator.py
new file mode 100644
index 0000000000000000000000000000000000000000..c87ad656c61cb251bfdfcbd23b1cc5263c68bf5f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/prior_generators/point_generator.py
@@ -0,0 +1,321 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple, Union
+
+import numpy as np
+import torch
+from torch import Tensor
+from torch.nn.modules.utils import _pair
+
+from mmdet.registry import TASK_UTILS
+
+DeviceType = Union[str, torch.device]
+
+
+@TASK_UTILS.register_module()
+class PointGenerator:
+
+ def _meshgrid(self,
+ x: Tensor,
+ y: Tensor,
+ row_major: bool = True) -> Tuple[Tensor, Tensor]:
+ """Generate mesh grid of x and y.
+
+ Args:
+ x (torch.Tensor): Grids of x dimension.
+ y (torch.Tensor): Grids of y dimension.
+ row_major (bool): Whether to return y grids first.
+ Defaults to True.
+
+ Returns:
+ tuple[torch.Tensor]: The mesh grids of x and y.
+ """
+ xx = x.repeat(len(y))
+ yy = y.view(-1, 1).repeat(1, len(x)).view(-1)
+ if row_major:
+ return xx, yy
+ else:
+ return yy, xx
+
+ def grid_points(self,
+ featmap_size: Tuple[int, int],
+ stride=16,
+ device: DeviceType = 'cuda') -> Tensor:
+ """Generate grid points of a single level.
+
+ Args:
+ featmap_size (tuple[int, int]): Size of the feature maps.
+ stride (int): The stride of corresponding feature map.
+ device (str | torch.device): The device the tensor will be put on.
+ Defaults to 'cuda'.
+
+ Returns:
+ torch.Tensor: grid point in a feature map.
+ """
+ feat_h, feat_w = featmap_size
+ shift_x = torch.arange(0., feat_w, device=device) * stride
+ shift_y = torch.arange(0., feat_h, device=device) * stride
+ shift_xx, shift_yy = self._meshgrid(shift_x, shift_y)
+ stride = shift_x.new_full((shift_xx.shape[0], ), stride)
+ shifts = torch.stack([shift_xx, shift_yy, stride], dim=-1)
+ all_points = shifts.to(device)
+ return all_points
+
+ def valid_flags(self,
+ featmap_size: Tuple[int, int],
+ valid_size: Tuple[int, int],
+ device: DeviceType = 'cuda') -> Tensor:
+ """Generate valid flags of anchors in a feature map.
+
+ Args:
+ featmap_sizes (list(tuple[int, int])): List of feature map sizes in
+ multiple feature levels.
+ valid_shape (tuple[int, int]): The valid shape of the image.
+ device (str | torch.device): Device where the anchors will be
+ put on.
+
+ Return:
+ torch.Tensor: Valid flags of anchors in a level.
+ """
+ feat_h, feat_w = featmap_size
+ valid_h, valid_w = valid_size
+ assert valid_h <= feat_h and valid_w <= feat_w
+ valid_x = torch.zeros(feat_w, dtype=torch.bool, device=device)
+ valid_y = torch.zeros(feat_h, dtype=torch.bool, device=device)
+ valid_x[:valid_w] = 1
+ valid_y[:valid_h] = 1
+ valid_xx, valid_yy = self._meshgrid(valid_x, valid_y)
+ valid = valid_xx & valid_yy
+ return valid
+
+
+@TASK_UTILS.register_module()
+class MlvlPointGenerator:
+ """Standard points generator for multi-level (Mlvl) feature maps in 2D
+ points-based detectors.
+
+ Args:
+ strides (list[int] | list[tuple[int, int]]): Strides of anchors
+ in multiple feature levels in order (w, h).
+ offset (float): The offset of points, the value is normalized with
+ corresponding stride. Defaults to 0.5.
+ """
+
+ def __init__(self,
+ strides: Union[List[int], List[Tuple[int, int]]],
+ offset: float = 0.5) -> None:
+ self.strides = [_pair(stride) for stride in strides]
+ self.offset = offset
+
+ @property
+ def num_levels(self) -> int:
+ """int: number of feature levels that the generator will be applied"""
+ return len(self.strides)
+
+ @property
+ def num_base_priors(self) -> List[int]:
+ """list[int]: The number of priors (points) at a point
+ on the feature grid"""
+ return [1 for _ in range(len(self.strides))]
+
+ def _meshgrid(self,
+ x: Tensor,
+ y: Tensor,
+ row_major: bool = True) -> Tuple[Tensor, Tensor]:
+ yy, xx = torch.meshgrid(y, x)
+ if row_major:
+ # warning .flatten() would cause error in ONNX exporting
+ # have to use reshape here
+ return xx.reshape(-1), yy.reshape(-1)
+
+ else:
+ return yy.reshape(-1), xx.reshape(-1)
+
+ def grid_priors(self,
+ featmap_sizes: List[Tuple],
+ dtype: torch.dtype = torch.float32,
+ device: DeviceType = 'cuda',
+ with_stride: bool = False) -> List[Tensor]:
+ """Generate grid points of multiple feature levels.
+
+ Args:
+ featmap_sizes (list[tuple]): List of feature map sizes in
+ multiple feature levels, each size arrange as
+ as (h, w).
+ dtype (:obj:`dtype`): Dtype of priors. Defaults to torch.float32.
+ device (str | torch.device): The device where the anchors will be
+ put on.
+ with_stride (bool): Whether to concatenate the stride to
+ the last dimension of points.
+
+ Return:
+ list[torch.Tensor]: Points of multiple feature levels.
+ The sizes of each tensor should be (N, 2) when with stride is
+ ``False``, where N = width * height, width and height
+ are the sizes of the corresponding feature level,
+ and the last dimension 2 represent (coord_x, coord_y),
+ otherwise the shape should be (N, 4),
+ and the last dimension 4 represent
+ (coord_x, coord_y, stride_w, stride_h).
+ """
+
+ assert self.num_levels == len(featmap_sizes)
+ multi_level_priors = []
+ for i in range(self.num_levels):
+ priors = self.single_level_grid_priors(
+ featmap_sizes[i],
+ level_idx=i,
+ dtype=dtype,
+ device=device,
+ with_stride=with_stride)
+ multi_level_priors.append(priors)
+ return multi_level_priors
+
+ def single_level_grid_priors(self,
+ featmap_size: Tuple[int],
+ level_idx: int,
+ dtype: torch.dtype = torch.float32,
+ device: DeviceType = 'cuda',
+ with_stride: bool = False) -> Tensor:
+ """Generate grid Points of a single level.
+
+ Note:
+ This function is usually called by method ``self.grid_priors``.
+
+ Args:
+ featmap_size (tuple[int]): Size of the feature maps, arrange as
+ (h, w).
+ level_idx (int): The index of corresponding feature map level.
+ dtype (:obj:`dtype`): Dtype of priors. Defaults to torch.float32.
+ device (str | torch.device): The device the tensor will be put on.
+ Defaults to 'cuda'.
+ with_stride (bool): Concatenate the stride to the last dimension
+ of points.
+
+ Return:
+ Tensor: Points of single feature levels.
+ The shape of tensor should be (N, 2) when with stride is
+ ``False``, where N = width * height, width and height
+ are the sizes of the corresponding feature level,
+ and the last dimension 2 represent (coord_x, coord_y),
+ otherwise the shape should be (N, 4),
+ and the last dimension 4 represent
+ (coord_x, coord_y, stride_w, stride_h).
+ """
+ feat_h, feat_w = featmap_size
+ stride_w, stride_h = self.strides[level_idx]
+ shift_x = (torch.arange(0, feat_w, device=device) +
+ self.offset) * stride_w
+ # keep featmap_size as Tensor instead of int, so that we
+ # can convert to ONNX correctly
+ shift_x = shift_x.to(dtype)
+
+ shift_y = (torch.arange(0, feat_h, device=device) +
+ self.offset) * stride_h
+ # keep featmap_size as Tensor instead of int, so that we
+ # can convert to ONNX correctly
+ shift_y = shift_y.to(dtype)
+ shift_xx, shift_yy = self._meshgrid(shift_x, shift_y)
+ if not with_stride:
+ shifts = torch.stack([shift_xx, shift_yy], dim=-1)
+ else:
+ # use `shape[0]` instead of `len(shift_xx)` for ONNX export
+ stride_w = shift_xx.new_full((shift_xx.shape[0], ),
+ stride_w).to(dtype)
+ stride_h = shift_xx.new_full((shift_yy.shape[0], ),
+ stride_h).to(dtype)
+ shifts = torch.stack([shift_xx, shift_yy, stride_w, stride_h],
+ dim=-1)
+ all_points = shifts.to(device)
+ return all_points
+
+ def valid_flags(self,
+ featmap_sizes: List[Tuple[int, int]],
+ pad_shape: Tuple[int],
+ device: DeviceType = 'cuda') -> List[Tensor]:
+ """Generate valid flags of points of multiple feature levels.
+
+ Args:
+ featmap_sizes (list(tuple)): List of feature map sizes in
+ multiple feature levels, each size arrange as
+ as (h, w).
+ pad_shape (tuple(int)): The padded shape of the image,
+ arrange as (h, w).
+ device (str | torch.device): The device where the anchors will be
+ put on.
+
+ Return:
+ list(torch.Tensor): Valid flags of points of multiple levels.
+ """
+ assert self.num_levels == len(featmap_sizes)
+ multi_level_flags = []
+ for i in range(self.num_levels):
+ point_stride = self.strides[i]
+ feat_h, feat_w = featmap_sizes[i]
+ h, w = pad_shape[:2]
+ valid_feat_h = min(int(np.ceil(h / point_stride[1])), feat_h)
+ valid_feat_w = min(int(np.ceil(w / point_stride[0])), feat_w)
+ flags = self.single_level_valid_flags((feat_h, feat_w),
+ (valid_feat_h, valid_feat_w),
+ device=device)
+ multi_level_flags.append(flags)
+ return multi_level_flags
+
+ def single_level_valid_flags(self,
+ featmap_size: Tuple[int, int],
+ valid_size: Tuple[int, int],
+ device: DeviceType = 'cuda') -> Tensor:
+ """Generate the valid flags of points of a single feature map.
+
+ Args:
+ featmap_size (tuple[int]): The size of feature maps, arrange as
+ as (h, w).
+ valid_size (tuple[int]): The valid size of the feature maps.
+ The size arrange as as (h, w).
+ device (str | torch.device): The device where the flags will be
+ put on. Defaults to 'cuda'.
+
+ Returns:
+ torch.Tensor: The valid flags of each points in a single level \
+ feature map.
+ """
+ feat_h, feat_w = featmap_size
+ valid_h, valid_w = valid_size
+ assert valid_h <= feat_h and valid_w <= feat_w
+ valid_x = torch.zeros(feat_w, dtype=torch.bool, device=device)
+ valid_y = torch.zeros(feat_h, dtype=torch.bool, device=device)
+ valid_x[:valid_w] = 1
+ valid_y[:valid_h] = 1
+ valid_xx, valid_yy = self._meshgrid(valid_x, valid_y)
+ valid = valid_xx & valid_yy
+ return valid
+
+ def sparse_priors(self,
+ prior_idxs: Tensor,
+ featmap_size: Tuple[int],
+ level_idx: int,
+ dtype: torch.dtype = torch.float32,
+ device: DeviceType = 'cuda') -> Tensor:
+ """Generate sparse points according to the ``prior_idxs``.
+
+ Args:
+ prior_idxs (Tensor): The index of corresponding anchors
+ in the feature map.
+ featmap_size (tuple[int]): feature map size arrange as (w, h).
+ level_idx (int): The level index of corresponding feature
+ map.
+ dtype (obj:`torch.dtype`): Date type of points. Defaults to
+ ``torch.float32``.
+ device (str | torch.device): The device where the points is
+ located.
+ Returns:
+ Tensor: Anchor with shape (N, 2), N should be equal to
+ the length of ``prior_idxs``. And last dimension
+ 2 represent (coord_x, coord_y).
+ """
+ height, width = featmap_size
+ x = (prior_idxs % width + self.offset) * self.strides[level_idx][0]
+ y = ((prior_idxs // width) % height +
+ self.offset) * self.strides[level_idx][1]
+ prioris = torch.stack([x, y], 1).to(dtype)
+ prioris = prioris.to(device)
+ return prioris
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/prior_generators/utils.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/prior_generators/utils.py
new file mode 100644
index 0000000000000000000000000000000000000000..3aa2dfd49669ba931d20ad9482cb841698cceb8a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/prior_generators/utils.py
@@ -0,0 +1,70 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Tuple
+
+import torch
+from torch import Tensor
+
+from mmdet.structures.bbox import BaseBoxes
+
+
+def anchor_inside_flags(flat_anchors: Tensor,
+ valid_flags: Tensor,
+ img_shape: Tuple[int],
+ allowed_border: int = 0) -> Tensor:
+ """Check whether the anchors are inside the border.
+
+ Args:
+ flat_anchors (torch.Tensor): Flatten anchors, shape (n, 4).
+ valid_flags (torch.Tensor): An existing valid flags of anchors.
+ img_shape (tuple(int)): Shape of current image.
+ allowed_border (int): The border to allow the valid anchor.
+ Defaults to 0.
+
+ Returns:
+ torch.Tensor: Flags indicating whether the anchors are inside a \
+ valid range.
+ """
+ img_h, img_w = img_shape[:2]
+ if allowed_border >= 0:
+ if isinstance(flat_anchors, BaseBoxes):
+ inside_flags = valid_flags & \
+ flat_anchors.is_inside([img_h, img_w],
+ all_inside=True,
+ allowed_border=allowed_border)
+ else:
+ inside_flags = valid_flags & \
+ (flat_anchors[:, 0] >= -allowed_border) & \
+ (flat_anchors[:, 1] >= -allowed_border) & \
+ (flat_anchors[:, 2] < img_w + allowed_border) & \
+ (flat_anchors[:, 3] < img_h + allowed_border)
+ else:
+ inside_flags = valid_flags
+ return inside_flags
+
+
+def calc_region(bbox: Tensor,
+ ratio: float,
+ featmap_size: Optional[Tuple] = None) -> Tuple[int]:
+ """Calculate a proportional bbox region.
+
+ The bbox center are fixed and the new h' and w' is h * ratio and w * ratio.
+
+ Args:
+ bbox (Tensor): Bboxes to calculate regions, shape (n, 4).
+ ratio (float): Ratio of the output region.
+ featmap_size (tuple, Optional): Feature map size in (height, width)
+ order used for clipping the boundary. Defaults to None.
+
+ Returns:
+ tuple: x1, y1, x2, y2
+ """
+ x1 = torch.round((1 - ratio) * bbox[0] + ratio * bbox[2]).long()
+ y1 = torch.round((1 - ratio) * bbox[1] + ratio * bbox[3]).long()
+ x2 = torch.round(ratio * bbox[0] + (1 - ratio) * bbox[2]).long()
+ y2 = torch.round(ratio * bbox[1] + (1 - ratio) * bbox[3]).long()
+ if featmap_size is not None:
+ x1 = x1.clamp(min=0, max=featmap_size[1])
+ y1 = y1.clamp(min=0, max=featmap_size[0])
+ x2 = x2.clamp(min=0, max=featmap_size[1])
+ y2 = y2.clamp(min=0, max=featmap_size[0])
+ return (x1, y1, x2, y2)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..3782eb898cf8acace63b4f16204cae6c07eb6e30
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/__init__.py
@@ -0,0 +1,22 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .base_sampler import BaseSampler
+from .combined_sampler import CombinedSampler
+from .instance_balanced_pos_sampler import InstanceBalancedPosSampler
+from .iou_balanced_neg_sampler import IoUBalancedNegSampler
+from .mask_pseudo_sampler import MaskPseudoSampler
+from .mask_sampling_result import MaskSamplingResult
+from .multi_instance_random_sampler import MultiInsRandomSampler
+from .multi_instance_sampling_result import MultiInstanceSamplingResult
+from .ohem_sampler import OHEMSampler
+from .pseudo_sampler import PseudoSampler
+from .random_sampler import RandomSampler
+from .sampling_result import SamplingResult
+from .score_hlr_sampler import ScoreHLRSampler
+
+__all__ = [
+ 'BaseSampler', 'PseudoSampler', 'RandomSampler',
+ 'InstanceBalancedPosSampler', 'IoUBalancedNegSampler', 'CombinedSampler',
+ 'OHEMSampler', 'SamplingResult', 'ScoreHLRSampler', 'MaskPseudoSampler',
+ 'MaskSamplingResult', 'MultiInstanceSamplingResult',
+ 'MultiInsRandomSampler'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/base_sampler.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/base_sampler.py
new file mode 100644
index 0000000000000000000000000000000000000000..be8a9a5ee3ec4e70b19aeea21b7998cf2b131d59
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/base_sampler.py
@@ -0,0 +1,136 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from abc import ABCMeta, abstractmethod
+
+import torch
+from mmengine.structures import InstanceData
+
+from mmdet.structures.bbox import BaseBoxes, cat_boxes
+from ..assigners import AssignResult
+from .sampling_result import SamplingResult
+
+
+class BaseSampler(metaclass=ABCMeta):
+ """Base class of samplers.
+
+ Args:
+ num (int): Number of samples
+ pos_fraction (float): Fraction of positive samples
+ neg_pos_up (int): Upper bound number of negative and
+ positive samples. Defaults to -1.
+ add_gt_as_proposals (bool): Whether to add ground truth
+ boxes as proposals. Defaults to True.
+ """
+
+ def __init__(self,
+ num: int,
+ pos_fraction: float,
+ neg_pos_ub: int = -1,
+ add_gt_as_proposals: bool = True,
+ **kwargs) -> None:
+ self.num = num
+ self.pos_fraction = pos_fraction
+ self.neg_pos_ub = neg_pos_ub
+ self.add_gt_as_proposals = add_gt_as_proposals
+ self.pos_sampler = self
+ self.neg_sampler = self
+
+ @abstractmethod
+ def _sample_pos(self, assign_result: AssignResult, num_expected: int,
+ **kwargs):
+ """Sample positive samples."""
+ pass
+
+ @abstractmethod
+ def _sample_neg(self, assign_result: AssignResult, num_expected: int,
+ **kwargs):
+ """Sample negative samples."""
+ pass
+
+ def sample(self, assign_result: AssignResult, pred_instances: InstanceData,
+ gt_instances: InstanceData, **kwargs) -> SamplingResult:
+ """Sample positive and negative bboxes.
+
+ This is a simple implementation of bbox sampling given candidates,
+ assigning results and ground truth bboxes.
+
+ Args:
+ assign_result (:obj:`AssignResult`): Assigning results.
+ pred_instances (:obj:`InstanceData`): Instances of model
+ predictions. It includes ``priors``, and the priors can
+ be anchors or points, or the bboxes predicted by the
+ previous stage, has shape (n, 4). The bboxes predicted by
+ the current model or stage will be named ``bboxes``,
+ ``labels``, and ``scores``, the same as the ``InstanceData``
+ in other places.
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes``, with shape (k, 4),
+ and ``labels``, with shape (k, ).
+
+ Returns:
+ :obj:`SamplingResult`: Sampling result.
+
+ Example:
+ >>> from mmengine.structures import InstanceData
+ >>> from mmdet.models.task_modules.samplers import RandomSampler,
+ >>> from mmdet.models.task_modules.assigners import AssignResult
+ >>> from mmdet.models.task_modules.samplers.
+ ... sampling_result import ensure_rng, random_boxes
+ >>> rng = ensure_rng(None)
+ >>> assign_result = AssignResult.random(rng=rng)
+ >>> pred_instances = InstanceData()
+ >>> pred_instances.priors = random_boxes(assign_result.num_preds,
+ ... rng=rng)
+ >>> gt_instances = InstanceData()
+ >>> gt_instances.bboxes = random_boxes(assign_result.num_gts,
+ ... rng=rng)
+ >>> gt_instances.labels = torch.randint(
+ ... 0, 5, (assign_result.num_gts,), dtype=torch.long)
+ >>> self = RandomSampler(num=32, pos_fraction=0.5, neg_pos_ub=-1,
+ >>> add_gt_as_proposals=False)
+ >>> self = self.sample(assign_result, pred_instances, gt_instances)
+ """
+ gt_bboxes = gt_instances.bboxes
+ priors = pred_instances.priors
+ gt_labels = gt_instances.labels
+ if len(priors.shape) < 2:
+ priors = priors[None, :]
+
+ gt_flags = priors.new_zeros((priors.shape[0], ), dtype=torch.uint8)
+ if self.add_gt_as_proposals and len(gt_bboxes) > 0:
+ # When `gt_bboxes` and `priors` are all box type, convert
+ # `gt_bboxes` type to `priors` type.
+ if (isinstance(gt_bboxes, BaseBoxes)
+ and isinstance(priors, BaseBoxes)):
+ gt_bboxes_ = gt_bboxes.convert_to(type(priors))
+ else:
+ gt_bboxes_ = gt_bboxes
+ priors = cat_boxes([gt_bboxes_, priors], dim=0)
+ assign_result.add_gt_(gt_labels)
+ gt_ones = priors.new_ones(gt_bboxes_.shape[0], dtype=torch.uint8)
+ gt_flags = torch.cat([gt_ones, gt_flags])
+
+ num_expected_pos = int(self.num * self.pos_fraction)
+ pos_inds = self.pos_sampler._sample_pos(
+ assign_result, num_expected_pos, bboxes=priors, **kwargs)
+ # We found that sampled indices have duplicated items occasionally.
+ # (may be a bug of PyTorch)
+ pos_inds = pos_inds.unique()
+ num_sampled_pos = pos_inds.numel()
+ num_expected_neg = self.num - num_sampled_pos
+ if self.neg_pos_ub >= 0:
+ _pos = max(1, num_sampled_pos)
+ neg_upper_bound = int(self.neg_pos_ub * _pos)
+ if num_expected_neg > neg_upper_bound:
+ num_expected_neg = neg_upper_bound
+ neg_inds = self.neg_sampler._sample_neg(
+ assign_result, num_expected_neg, bboxes=priors, **kwargs)
+ neg_inds = neg_inds.unique()
+
+ sampling_result = SamplingResult(
+ pos_inds=pos_inds,
+ neg_inds=neg_inds,
+ priors=priors,
+ gt_bboxes=gt_bboxes,
+ assign_result=assign_result,
+ gt_flags=gt_flags)
+ return sampling_result
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/combined_sampler.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/combined_sampler.py
new file mode 100644
index 0000000000000000000000000000000000000000..8e0560e372efffe865fa32028d823280a8bd5d87
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/combined_sampler.py
@@ -0,0 +1,21 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmdet.registry import TASK_UTILS
+from .base_sampler import BaseSampler
+
+
+@TASK_UTILS.register_module()
+class CombinedSampler(BaseSampler):
+ """A sampler that combines positive sampler and negative sampler."""
+
+ def __init__(self, pos_sampler, neg_sampler, **kwargs):
+ super(CombinedSampler, self).__init__(**kwargs)
+ self.pos_sampler = TASK_UTILS.build(pos_sampler, default_args=kwargs)
+ self.neg_sampler = TASK_UTILS.build(neg_sampler, default_args=kwargs)
+
+ def _sample_pos(self, **kwargs):
+ """Sample positive samples."""
+ raise NotImplementedError
+
+ def _sample_neg(self, **kwargs):
+ """Sample negative samples."""
+ raise NotImplementedError
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/instance_balanced_pos_sampler.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/instance_balanced_pos_sampler.py
new file mode 100644
index 0000000000000000000000000000000000000000..e48d8e9158e8dabf0bb4072b8e421de9b6410d00
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/instance_balanced_pos_sampler.py
@@ -0,0 +1,56 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import numpy as np
+import torch
+
+from mmdet.registry import TASK_UTILS
+from .random_sampler import RandomSampler
+
+
+@TASK_UTILS.register_module()
+class InstanceBalancedPosSampler(RandomSampler):
+ """Instance balanced sampler that samples equal number of positive samples
+ for each instance."""
+
+ def _sample_pos(self, assign_result, num_expected, **kwargs):
+ """Sample positive boxes.
+
+ Args:
+ assign_result (:obj:`AssignResult`): The assigned results of boxes.
+ num_expected (int): The number of expected positive samples
+
+ Returns:
+ Tensor or ndarray: sampled indices.
+ """
+ pos_inds = torch.nonzero(assign_result.gt_inds > 0, as_tuple=False)
+ if pos_inds.numel() != 0:
+ pos_inds = pos_inds.squeeze(1)
+ if pos_inds.numel() <= num_expected:
+ return pos_inds
+ else:
+ unique_gt_inds = assign_result.gt_inds[pos_inds].unique()
+ num_gts = len(unique_gt_inds)
+ num_per_gt = int(round(num_expected / float(num_gts)) + 1)
+ sampled_inds = []
+ for i in unique_gt_inds:
+ inds = torch.nonzero(
+ assign_result.gt_inds == i.item(), as_tuple=False)
+ if inds.numel() != 0:
+ inds = inds.squeeze(1)
+ else:
+ continue
+ if len(inds) > num_per_gt:
+ inds = self.random_choice(inds, num_per_gt)
+ sampled_inds.append(inds)
+ sampled_inds = torch.cat(sampled_inds)
+ if len(sampled_inds) < num_expected:
+ num_extra = num_expected - len(sampled_inds)
+ extra_inds = np.array(
+ list(set(pos_inds.cpu()) - set(sampled_inds.cpu())))
+ if len(extra_inds) > num_extra:
+ extra_inds = self.random_choice(extra_inds, num_extra)
+ extra_inds = torch.from_numpy(extra_inds).to(
+ assign_result.gt_inds.device).long()
+ sampled_inds = torch.cat([sampled_inds, extra_inds])
+ elif len(sampled_inds) > num_expected:
+ sampled_inds = self.random_choice(sampled_inds, num_expected)
+ return sampled_inds
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/iou_balanced_neg_sampler.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/iou_balanced_neg_sampler.py
new file mode 100644
index 0000000000000000000000000000000000000000..dc1f46413c99d115f31ef190b4fb198b588a156e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/iou_balanced_neg_sampler.py
@@ -0,0 +1,158 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import numpy as np
+import torch
+
+from mmdet.registry import TASK_UTILS
+from .random_sampler import RandomSampler
+
+
+@TASK_UTILS.register_module()
+class IoUBalancedNegSampler(RandomSampler):
+ """IoU Balanced Sampling.
+
+ arXiv: https://arxiv.org/pdf/1904.02701.pdf (CVPR 2019)
+
+ Sampling proposals according to their IoU. `floor_fraction` of needed RoIs
+ are sampled from proposals whose IoU are lower than `floor_thr` randomly.
+ The others are sampled from proposals whose IoU are higher than
+ `floor_thr`. These proposals are sampled from some bins evenly, which are
+ split by `num_bins` via IoU evenly.
+
+ Args:
+ num (int): number of proposals.
+ pos_fraction (float): fraction of positive proposals.
+ floor_thr (float): threshold (minimum) IoU for IoU balanced sampling,
+ set to -1 if all using IoU balanced sampling.
+ floor_fraction (float): sampling fraction of proposals under floor_thr.
+ num_bins (int): number of bins in IoU balanced sampling.
+ """
+
+ def __init__(self,
+ num,
+ pos_fraction,
+ floor_thr=-1,
+ floor_fraction=0,
+ num_bins=3,
+ **kwargs):
+ super(IoUBalancedNegSampler, self).__init__(num, pos_fraction,
+ **kwargs)
+ assert floor_thr >= 0 or floor_thr == -1
+ assert 0 <= floor_fraction <= 1
+ assert num_bins >= 1
+
+ self.floor_thr = floor_thr
+ self.floor_fraction = floor_fraction
+ self.num_bins = num_bins
+
+ def sample_via_interval(self, max_overlaps, full_set, num_expected):
+ """Sample according to the iou interval.
+
+ Args:
+ max_overlaps (torch.Tensor): IoU between bounding boxes and ground
+ truth boxes.
+ full_set (set(int)): A full set of indices of boxes。
+ num_expected (int): Number of expected samples。
+
+ Returns:
+ np.ndarray: Indices of samples
+ """
+ max_iou = max_overlaps.max()
+ iou_interval = (max_iou - self.floor_thr) / self.num_bins
+ per_num_expected = int(num_expected / self.num_bins)
+
+ sampled_inds = []
+ for i in range(self.num_bins):
+ start_iou = self.floor_thr + i * iou_interval
+ end_iou = self.floor_thr + (i + 1) * iou_interval
+ tmp_set = set(
+ np.where(
+ np.logical_and(max_overlaps >= start_iou,
+ max_overlaps < end_iou))[0])
+ tmp_inds = list(tmp_set & full_set)
+ if len(tmp_inds) > per_num_expected:
+ tmp_sampled_set = self.random_choice(tmp_inds,
+ per_num_expected)
+ else:
+ tmp_sampled_set = np.array(tmp_inds, dtype=np.int64)
+ sampled_inds.append(tmp_sampled_set)
+
+ sampled_inds = np.concatenate(sampled_inds)
+ if len(sampled_inds) < num_expected:
+ num_extra = num_expected - len(sampled_inds)
+ extra_inds = np.array(list(full_set - set(sampled_inds)))
+ if len(extra_inds) > num_extra:
+ extra_inds = self.random_choice(extra_inds, num_extra)
+ sampled_inds = np.concatenate([sampled_inds, extra_inds])
+
+ return sampled_inds
+
+ def _sample_neg(self, assign_result, num_expected, **kwargs):
+ """Sample negative boxes.
+
+ Args:
+ assign_result (:obj:`AssignResult`): The assigned results of boxes.
+ num_expected (int): The number of expected negative samples
+
+ Returns:
+ Tensor or ndarray: sampled indices.
+ """
+ neg_inds = torch.nonzero(assign_result.gt_inds == 0, as_tuple=False)
+ if neg_inds.numel() != 0:
+ neg_inds = neg_inds.squeeze(1)
+ if len(neg_inds) <= num_expected:
+ return neg_inds
+ else:
+ max_overlaps = assign_result.max_overlaps.cpu().numpy()
+ # balance sampling for negative samples
+ neg_set = set(neg_inds.cpu().numpy())
+
+ if self.floor_thr > 0:
+ floor_set = set(
+ np.where(
+ np.logical_and(max_overlaps >= 0,
+ max_overlaps < self.floor_thr))[0])
+ iou_sampling_set = set(
+ np.where(max_overlaps >= self.floor_thr)[0])
+ elif self.floor_thr == 0:
+ floor_set = set(np.where(max_overlaps == 0)[0])
+ iou_sampling_set = set(
+ np.where(max_overlaps > self.floor_thr)[0])
+ else:
+ floor_set = set()
+ iou_sampling_set = set(
+ np.where(max_overlaps > self.floor_thr)[0])
+ # for sampling interval calculation
+ self.floor_thr = 0
+
+ floor_neg_inds = list(floor_set & neg_set)
+ iou_sampling_neg_inds = list(iou_sampling_set & neg_set)
+ num_expected_iou_sampling = int(num_expected *
+ (1 - self.floor_fraction))
+ if len(iou_sampling_neg_inds) > num_expected_iou_sampling:
+ if self.num_bins >= 2:
+ iou_sampled_inds = self.sample_via_interval(
+ max_overlaps, set(iou_sampling_neg_inds),
+ num_expected_iou_sampling)
+ else:
+ iou_sampled_inds = self.random_choice(
+ iou_sampling_neg_inds, num_expected_iou_sampling)
+ else:
+ iou_sampled_inds = np.array(
+ iou_sampling_neg_inds, dtype=np.int64)
+ num_expected_floor = num_expected - len(iou_sampled_inds)
+ if len(floor_neg_inds) > num_expected_floor:
+ sampled_floor_inds = self.random_choice(
+ floor_neg_inds, num_expected_floor)
+ else:
+ sampled_floor_inds = np.array(floor_neg_inds, dtype=np.int64)
+ sampled_inds = np.concatenate(
+ (sampled_floor_inds, iou_sampled_inds))
+ if len(sampled_inds) < num_expected:
+ num_extra = num_expected - len(sampled_inds)
+ extra_inds = np.array(list(neg_set - set(sampled_inds)))
+ if len(extra_inds) > num_extra:
+ extra_inds = self.random_choice(extra_inds, num_extra)
+ sampled_inds = np.concatenate((sampled_inds, extra_inds))
+ sampled_inds = torch.from_numpy(sampled_inds).long().to(
+ assign_result.gt_inds.device)
+ return sampled_inds
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/mask_pseudo_sampler.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/mask_pseudo_sampler.py
new file mode 100644
index 0000000000000000000000000000000000000000..307dd5d15c962b97dc60b899e60170d0bfed90a7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/mask_pseudo_sampler.py
@@ -0,0 +1,60 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+"""copy from
+https://github.com/ZwwWayne/K-Net/blob/main/knet/det/mask_pseudo_sampler.py."""
+
+import torch
+from mmengine.structures import InstanceData
+
+from mmdet.registry import TASK_UTILS
+from ..assigners import AssignResult
+from .base_sampler import BaseSampler
+from .mask_sampling_result import MaskSamplingResult
+
+
+@TASK_UTILS.register_module()
+class MaskPseudoSampler(BaseSampler):
+ """A pseudo sampler that does not do sampling actually."""
+
+ def __init__(self, **kwargs):
+ pass
+
+ def _sample_pos(self, **kwargs):
+ """Sample positive samples."""
+ raise NotImplementedError
+
+ def _sample_neg(self, **kwargs):
+ """Sample negative samples."""
+ raise NotImplementedError
+
+ def sample(self, assign_result: AssignResult, pred_instances: InstanceData,
+ gt_instances: InstanceData, *args, **kwargs):
+ """Directly returns the positive and negative indices of samples.
+
+ Args:
+ assign_result (:obj:`AssignResult`): Mask assigning results.
+ pred_instances (:obj:`InstanceData`): Instances of model
+ predictions. It includes ``scores`` and ``masks`` predicted
+ by the model.
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``labels`` and ``masks``
+ attributes.
+
+ Returns:
+ :obj:`SamplingResult`: sampler results
+ """
+ pred_masks = pred_instances.masks
+ gt_masks = gt_instances.masks
+ pos_inds = torch.nonzero(
+ assign_result.gt_inds > 0, as_tuple=False).squeeze(-1).unique()
+ neg_inds = torch.nonzero(
+ assign_result.gt_inds == 0, as_tuple=False).squeeze(-1).unique()
+ gt_flags = pred_masks.new_zeros(pred_masks.shape[0], dtype=torch.uint8)
+ sampling_result = MaskSamplingResult(
+ pos_inds=pos_inds,
+ neg_inds=neg_inds,
+ masks=pred_masks,
+ gt_masks=gt_masks,
+ assign_result=assign_result,
+ gt_flags=gt_flags,
+ avg_factor_with_neg=False)
+ return sampling_result
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/mask_sampling_result.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/mask_sampling_result.py
new file mode 100644
index 0000000000000000000000000000000000000000..adaa62e8a0af28bb004a34b961f672ec03988d2c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/mask_sampling_result.py
@@ -0,0 +1,68 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+"""copy from
+https://github.com/ZwwWayne/K-Net/blob/main/knet/det/mask_pseudo_sampler.py."""
+
+import torch
+from torch import Tensor
+
+from ..assigners import AssignResult
+from .sampling_result import SamplingResult
+
+
+class MaskSamplingResult(SamplingResult):
+ """Mask sampling result."""
+
+ def __init__(self,
+ pos_inds: Tensor,
+ neg_inds: Tensor,
+ masks: Tensor,
+ gt_masks: Tensor,
+ assign_result: AssignResult,
+ gt_flags: Tensor,
+ avg_factor_with_neg: bool = True) -> None:
+ self.pos_inds = pos_inds
+ self.neg_inds = neg_inds
+ self.num_pos = max(pos_inds.numel(), 1)
+ self.num_neg = max(neg_inds.numel(), 1)
+ self.avg_factor = self.num_pos + self.num_neg \
+ if avg_factor_with_neg else self.num_pos
+
+ self.pos_masks = masks[pos_inds]
+ self.neg_masks = masks[neg_inds]
+ self.pos_is_gt = gt_flags[pos_inds]
+
+ self.num_gts = gt_masks.shape[0]
+ self.pos_assigned_gt_inds = assign_result.gt_inds[pos_inds] - 1
+
+ if gt_masks.numel() == 0:
+ # hack for index error case
+ assert self.pos_assigned_gt_inds.numel() == 0
+ self.pos_gt_masks = torch.empty_like(gt_masks)
+ else:
+ self.pos_gt_masks = gt_masks[self.pos_assigned_gt_inds, :]
+
+ @property
+ def masks(self) -> Tensor:
+ """torch.Tensor: concatenated positive and negative masks."""
+ return torch.cat([self.pos_masks, self.neg_masks])
+
+ def __nice__(self) -> str:
+ data = self.info.copy()
+ data['pos_masks'] = data.pop('pos_masks').shape
+ data['neg_masks'] = data.pop('neg_masks').shape
+ parts = [f"'{k}': {v!r}" for k, v in sorted(data.items())]
+ body = ' ' + ',\n '.join(parts)
+ return '{\n' + body + '\n}'
+
+ @property
+ def info(self) -> dict:
+ """Returns a dictionary of info about the object."""
+ return {
+ 'pos_inds': self.pos_inds,
+ 'neg_inds': self.neg_inds,
+ 'pos_masks': self.pos_masks,
+ 'neg_masks': self.neg_masks,
+ 'pos_is_gt': self.pos_is_gt,
+ 'num_gts': self.num_gts,
+ 'pos_assigned_gt_inds': self.pos_assigned_gt_inds,
+ }
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/multi_instance_random_sampler.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/multi_instance_random_sampler.py
new file mode 100644
index 0000000000000000000000000000000000000000..8b74054e3a11ed6025e98e90bd0addb131a1dc02
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/multi_instance_random_sampler.py
@@ -0,0 +1,130 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Union
+
+import torch
+from mmengine.structures import InstanceData
+from numpy import ndarray
+from torch import Tensor
+
+from mmdet.registry import TASK_UTILS
+from ..assigners import AssignResult
+from .multi_instance_sampling_result import MultiInstanceSamplingResult
+from .random_sampler import RandomSampler
+
+
+@TASK_UTILS.register_module()
+class MultiInsRandomSampler(RandomSampler):
+ """Random sampler for multi instance.
+
+ Note:
+ Multi-instance means to predict multiple detection boxes with
+ one proposal box. `AssignResult` may assign multiple gt boxes
+ to each proposal box, in this case `RandomSampler` should be
+ replaced by `MultiInsRandomSampler`
+ """
+
+ def _sample_pos(self, assign_result: AssignResult, num_expected: int,
+ **kwargs) -> Union[Tensor, ndarray]:
+ """Randomly sample some positive samples.
+
+ Args:
+ assign_result (:obj:`AssignResult`): Bbox assigning results.
+ num_expected (int): The number of expected positive samples
+
+ Returns:
+ Tensor or ndarray: sampled indices.
+ """
+ pos_inds = torch.nonzero(
+ assign_result.labels[:, 0] > 0, as_tuple=False)
+ if pos_inds.numel() != 0:
+ pos_inds = pos_inds.squeeze(1)
+ if pos_inds.numel() <= num_expected:
+ return pos_inds
+ else:
+ return self.random_choice(pos_inds, num_expected)
+
+ def _sample_neg(self, assign_result: AssignResult, num_expected: int,
+ **kwargs) -> Union[Tensor, ndarray]:
+ """Randomly sample some negative samples.
+
+ Args:
+ assign_result (:obj:`AssignResult`): Bbox assigning results.
+ num_expected (int): The number of expected positive samples
+
+ Returns:
+ Tensor or ndarray: sampled indices.
+ """
+ neg_inds = torch.nonzero(
+ assign_result.labels[:, 0] == 0, as_tuple=False)
+ if neg_inds.numel() != 0:
+ neg_inds = neg_inds.squeeze(1)
+ if len(neg_inds) <= num_expected:
+ return neg_inds
+ else:
+ return self.random_choice(neg_inds, num_expected)
+
+ def sample(self, assign_result: AssignResult, pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ **kwargs) -> MultiInstanceSamplingResult:
+ """Sample positive and negative bboxes.
+
+ Args:
+ assign_result (:obj:`AssignResult`): Assigning results from
+ MultiInstanceAssigner.
+ pred_instances (:obj:`InstanceData`): Instances of model
+ predictions. It includes ``priors``, and the priors can
+ be anchors or points, or the bboxes predicted by the
+ previous stage, has shape (n, 4). The bboxes predicted by
+ the current model or stage will be named ``bboxes``,
+ ``labels``, and ``scores``, the same as the ``InstanceData``
+ in other places.
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes``, with shape (k, 4),
+ and ``labels``, with shape (k, ).
+
+ Returns:
+ :obj:`MultiInstanceSamplingResult`: Sampling result.
+ """
+
+ assert 'batch_gt_instances_ignore' in kwargs, \
+ 'batch_gt_instances_ignore is necessary for MultiInsRandomSampler'
+
+ gt_bboxes = gt_instances.bboxes
+ ignore_bboxes = kwargs['batch_gt_instances_ignore'].bboxes
+ gt_and_ignore_bboxes = torch.cat([gt_bboxes, ignore_bboxes], dim=0)
+ priors = pred_instances.priors
+ if len(priors.shape) < 2:
+ priors = priors[None, :]
+ priors = priors[:, :4]
+
+ gt_flags = priors.new_zeros((priors.shape[0], ), dtype=torch.uint8)
+ priors = torch.cat([priors, gt_and_ignore_bboxes], dim=0)
+ gt_ones = priors.new_ones(
+ gt_and_ignore_bboxes.shape[0], dtype=torch.uint8)
+ gt_flags = torch.cat([gt_flags, gt_ones])
+
+ num_expected_pos = int(self.num * self.pos_fraction)
+ pos_inds = self.pos_sampler._sample_pos(assign_result,
+ num_expected_pos)
+ # We found that sampled indices have duplicated items occasionally.
+ # (may be a bug of PyTorch)
+ pos_inds = pos_inds.unique()
+ num_sampled_pos = pos_inds.numel()
+ num_expected_neg = self.num - num_sampled_pos
+ if self.neg_pos_ub >= 0:
+ _pos = max(1, num_sampled_pos)
+ neg_upper_bound = int(self.neg_pos_ub * _pos)
+ if num_expected_neg > neg_upper_bound:
+ num_expected_neg = neg_upper_bound
+ neg_inds = self.neg_sampler._sample_neg(assign_result,
+ num_expected_neg)
+ neg_inds = neg_inds.unique()
+
+ sampling_result = MultiInstanceSamplingResult(
+ pos_inds=pos_inds,
+ neg_inds=neg_inds,
+ priors=priors,
+ gt_and_ignore_bboxes=gt_and_ignore_bboxes,
+ assign_result=assign_result,
+ gt_flags=gt_flags)
+ return sampling_result
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/multi_instance_sampling_result.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/multi_instance_sampling_result.py
new file mode 100644
index 0000000000000000000000000000000000000000..438a0aa91c0cc8904f6d8bba7139408dd99b98cf
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/multi_instance_sampling_result.py
@@ -0,0 +1,56 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+from torch import Tensor
+
+from ..assigners import AssignResult
+from .sampling_result import SamplingResult
+
+
+class MultiInstanceSamplingResult(SamplingResult):
+ """Bbox sampling result. Further encapsulation of SamplingResult. Three
+ attributes neg_assigned_gt_inds, neg_gt_labels, and neg_gt_bboxes have been
+ added for SamplingResult.
+
+ Args:
+ pos_inds (Tensor): Indices of positive samples.
+ neg_inds (Tensor): Indices of negative samples.
+ priors (Tensor): The priors can be anchors or points,
+ or the bboxes predicted by the previous stage.
+ gt_and_ignore_bboxes (Tensor): Ground truth and ignore bboxes.
+ assign_result (:obj:`AssignResult`): Assigning results.
+ gt_flags (Tensor): The Ground truth flags.
+ avg_factor_with_neg (bool): If True, ``avg_factor`` equal to
+ the number of total priors; Otherwise, it is the number of
+ positive priors. Defaults to True.
+ """
+
+ def __init__(self,
+ pos_inds: Tensor,
+ neg_inds: Tensor,
+ priors: Tensor,
+ gt_and_ignore_bboxes: Tensor,
+ assign_result: AssignResult,
+ gt_flags: Tensor,
+ avg_factor_with_neg: bool = True) -> None:
+ self.neg_assigned_gt_inds = assign_result.gt_inds[neg_inds]
+ self.neg_gt_labels = assign_result.labels[neg_inds]
+
+ if gt_and_ignore_bboxes.numel() == 0:
+ self.neg_gt_bboxes = torch.empty_like(gt_and_ignore_bboxes).view(
+ -1, 4)
+ else:
+ if len(gt_and_ignore_bboxes.shape) < 2:
+ gt_and_ignore_bboxes = gt_and_ignore_bboxes.view(-1, 4)
+ self.neg_gt_bboxes = gt_and_ignore_bboxes[
+ self.neg_assigned_gt_inds.long(), :]
+
+ # To resist the minus 1 operation in `SamplingResult.init()`.
+ assign_result.gt_inds += 1
+ super().__init__(
+ pos_inds=pos_inds,
+ neg_inds=neg_inds,
+ priors=priors,
+ gt_bboxes=gt_and_ignore_bboxes,
+ assign_result=assign_result,
+ gt_flags=gt_flags,
+ avg_factor_with_neg=avg_factor_with_neg)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/ohem_sampler.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/ohem_sampler.py
new file mode 100644
index 0000000000000000000000000000000000000000..f478a448cde00d64caeba1d0ba613d2497a7fb12
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/ohem_sampler.py
@@ -0,0 +1,111 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+
+from mmdet.registry import TASK_UTILS
+from mmdet.structures.bbox import bbox2roi
+from .base_sampler import BaseSampler
+
+
+@TASK_UTILS.register_module()
+class OHEMSampler(BaseSampler):
+ r"""Online Hard Example Mining Sampler described in `Training Region-based
+ Object Detectors with Online Hard Example Mining
+ `_.
+ """
+
+ def __init__(self,
+ num,
+ pos_fraction,
+ context,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True,
+ loss_key='loss_cls',
+ **kwargs):
+ super(OHEMSampler, self).__init__(num, pos_fraction, neg_pos_ub,
+ add_gt_as_proposals)
+ self.context = context
+ if not hasattr(self.context, 'num_stages'):
+ self.bbox_head = self.context.bbox_head
+ else:
+ self.bbox_head = self.context.bbox_head[self.context.current_stage]
+
+ self.loss_key = loss_key
+
+ def hard_mining(self, inds, num_expected, bboxes, labels, feats):
+ with torch.no_grad():
+ rois = bbox2roi([bboxes])
+ if not hasattr(self.context, 'num_stages'):
+ bbox_results = self.context._bbox_forward(feats, rois)
+ else:
+ bbox_results = self.context._bbox_forward(
+ self.context.current_stage, feats, rois)
+ cls_score = bbox_results['cls_score']
+ loss = self.bbox_head.loss(
+ cls_score=cls_score,
+ bbox_pred=None,
+ rois=rois,
+ labels=labels,
+ label_weights=cls_score.new_ones(cls_score.size(0)),
+ bbox_targets=None,
+ bbox_weights=None,
+ reduction_override='none')[self.loss_key]
+ _, topk_loss_inds = loss.topk(num_expected)
+ return inds[topk_loss_inds]
+
+ def _sample_pos(self,
+ assign_result,
+ num_expected,
+ bboxes=None,
+ feats=None,
+ **kwargs):
+ """Sample positive boxes.
+
+ Args:
+ assign_result (:obj:`AssignResult`): Assigned results
+ num_expected (int): Number of expected positive samples
+ bboxes (torch.Tensor, optional): Boxes. Defaults to None.
+ feats (list[torch.Tensor], optional): Multi-level features.
+ Defaults to None.
+
+ Returns:
+ torch.Tensor: Indices of positive samples
+ """
+ # Sample some hard positive samples
+ pos_inds = torch.nonzero(assign_result.gt_inds > 0, as_tuple=False)
+ if pos_inds.numel() != 0:
+ pos_inds = pos_inds.squeeze(1)
+ if pos_inds.numel() <= num_expected:
+ return pos_inds
+ else:
+ return self.hard_mining(pos_inds, num_expected, bboxes[pos_inds],
+ assign_result.labels[pos_inds], feats)
+
+ def _sample_neg(self,
+ assign_result,
+ num_expected,
+ bboxes=None,
+ feats=None,
+ **kwargs):
+ """Sample negative boxes.
+
+ Args:
+ assign_result (:obj:`AssignResult`): Assigned results
+ num_expected (int): Number of expected negative samples
+ bboxes (torch.Tensor, optional): Boxes. Defaults to None.
+ feats (list[torch.Tensor], optional): Multi-level features.
+ Defaults to None.
+
+ Returns:
+ torch.Tensor: Indices of negative samples
+ """
+ # Sample some hard negative samples
+ neg_inds = torch.nonzero(assign_result.gt_inds == 0, as_tuple=False)
+ if neg_inds.numel() != 0:
+ neg_inds = neg_inds.squeeze(1)
+ if len(neg_inds) <= num_expected:
+ return neg_inds
+ else:
+ neg_labels = assign_result.labels.new_empty(
+ neg_inds.size(0)).fill_(self.bbox_head.num_classes)
+ return self.hard_mining(neg_inds, num_expected, bboxes[neg_inds],
+ neg_labels, feats)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/pseudo_sampler.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/pseudo_sampler.py
new file mode 100644
index 0000000000000000000000000000000000000000..a8186cc3364516f34abe1c293017db6e2042d92a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/pseudo_sampler.py
@@ -0,0 +1,60 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+from mmengine.structures import InstanceData
+
+from mmdet.registry import TASK_UTILS
+from ..assigners import AssignResult
+from .base_sampler import BaseSampler
+from .sampling_result import SamplingResult
+
+
+@TASK_UTILS.register_module()
+class PseudoSampler(BaseSampler):
+ """A pseudo sampler that does not do sampling actually."""
+
+ def __init__(self, **kwargs):
+ pass
+
+ def _sample_pos(self, **kwargs):
+ """Sample positive samples."""
+ raise NotImplementedError
+
+ def _sample_neg(self, **kwargs):
+ """Sample negative samples."""
+ raise NotImplementedError
+
+ def sample(self, assign_result: AssignResult, pred_instances: InstanceData,
+ gt_instances: InstanceData, *args, **kwargs):
+ """Directly returns the positive and negative indices of samples.
+
+ Args:
+ assign_result (:obj:`AssignResult`): Bbox assigning results.
+ pred_instances (:obj:`InstanceData`): Instances of model
+ predictions. It includes ``priors``, and the priors can
+ be anchors, points, or bboxes predicted by the model,
+ shape(n, 4).
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes`` and ``labels``
+ attributes.
+
+ Returns:
+ :obj:`SamplingResult`: sampler results
+ """
+ gt_bboxes = gt_instances.bboxes
+ priors = pred_instances.priors
+
+ pos_inds = torch.nonzero(
+ assign_result.gt_inds > 0, as_tuple=False).squeeze(-1).unique()
+ neg_inds = torch.nonzero(
+ assign_result.gt_inds == 0, as_tuple=False).squeeze(-1).unique()
+
+ gt_flags = priors.new_zeros(priors.shape[0], dtype=torch.uint8)
+ sampling_result = SamplingResult(
+ pos_inds=pos_inds,
+ neg_inds=neg_inds,
+ priors=priors,
+ gt_bboxes=gt_bboxes,
+ assign_result=assign_result,
+ gt_flags=gt_flags,
+ avg_factor_with_neg=False)
+ return sampling_result
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/random_sampler.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/random_sampler.py
new file mode 100644
index 0000000000000000000000000000000000000000..fa03665fc36cc6a0084431324b16727b2dc8993e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/random_sampler.py
@@ -0,0 +1,109 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Union
+
+import torch
+from numpy import ndarray
+from torch import Tensor
+
+from mmdet.registry import TASK_UTILS
+from ..assigners import AssignResult
+from .base_sampler import BaseSampler
+
+
+@TASK_UTILS.register_module()
+class RandomSampler(BaseSampler):
+ """Random sampler.
+
+ Args:
+ num (int): Number of samples
+ pos_fraction (float): Fraction of positive samples
+ neg_pos_up (int): Upper bound number of negative and
+ positive samples. Defaults to -1.
+ add_gt_as_proposals (bool): Whether to add ground truth
+ boxes as proposals. Defaults to True.
+ """
+
+ def __init__(self,
+ num: int,
+ pos_fraction: float,
+ neg_pos_ub: int = -1,
+ add_gt_as_proposals: bool = True,
+ **kwargs):
+ from .sampling_result import ensure_rng
+ super().__init__(
+ num=num,
+ pos_fraction=pos_fraction,
+ neg_pos_ub=neg_pos_ub,
+ add_gt_as_proposals=add_gt_as_proposals)
+ self.rng = ensure_rng(kwargs.get('rng', None))
+
+ def random_choice(self, gallery: Union[Tensor, ndarray, list],
+ num: int) -> Union[Tensor, ndarray]:
+ """Random select some elements from the gallery.
+
+ If `gallery` is a Tensor, the returned indices will be a Tensor;
+ If `gallery` is a ndarray or list, the returned indices will be a
+ ndarray.
+
+ Args:
+ gallery (Tensor | ndarray | list): indices pool.
+ num (int): expected sample num.
+
+ Returns:
+ Tensor or ndarray: sampled indices.
+ """
+ assert len(gallery) >= num
+
+ is_tensor = isinstance(gallery, torch.Tensor)
+ if not is_tensor:
+ if torch.cuda.is_available():
+ device = torch.cuda.current_device()
+ else:
+ device = 'cpu'
+ gallery = torch.tensor(gallery, dtype=torch.long, device=device)
+ # This is a temporary fix. We can revert the following code
+ # when PyTorch fixes the abnormal return of torch.randperm.
+ # See: https://github.com/open-mmlab/mmdetection/pull/5014
+ perm = torch.randperm(gallery.numel())[:num].to(device=gallery.device)
+ rand_inds = gallery[perm]
+ if not is_tensor:
+ rand_inds = rand_inds.cpu().numpy()
+ return rand_inds
+
+ def _sample_pos(self, assign_result: AssignResult, num_expected: int,
+ **kwargs) -> Union[Tensor, ndarray]:
+ """Randomly sample some positive samples.
+
+ Args:
+ assign_result (:obj:`AssignResult`): Bbox assigning results.
+ num_expected (int): The number of expected positive samples
+
+ Returns:
+ Tensor or ndarray: sampled indices.
+ """
+ pos_inds = torch.nonzero(assign_result.gt_inds > 0, as_tuple=False)
+ if pos_inds.numel() != 0:
+ pos_inds = pos_inds.squeeze(1)
+ if pos_inds.numel() <= num_expected:
+ return pos_inds
+ else:
+ return self.random_choice(pos_inds, num_expected)
+
+ def _sample_neg(self, assign_result: AssignResult, num_expected: int,
+ **kwargs) -> Union[Tensor, ndarray]:
+ """Randomly sample some negative samples.
+
+ Args:
+ assign_result (:obj:`AssignResult`): Bbox assigning results.
+ num_expected (int): The number of expected positive samples
+
+ Returns:
+ Tensor or ndarray: sampled indices.
+ """
+ neg_inds = torch.nonzero(assign_result.gt_inds == 0, as_tuple=False)
+ if neg_inds.numel() != 0:
+ neg_inds = neg_inds.squeeze(1)
+ if len(neg_inds) <= num_expected:
+ return neg_inds
+ else:
+ return self.random_choice(neg_inds, num_expected)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/sampling_result.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/sampling_result.py
new file mode 100644
index 0000000000000000000000000000000000000000..cb510ee68f24b8c444b6ed447016bfc785b825c2
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/sampling_result.py
@@ -0,0 +1,240 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+
+import numpy as np
+import torch
+from torch import Tensor
+
+from mmdet.structures.bbox import BaseBoxes, cat_boxes
+from mmdet.utils import util_mixins
+from mmdet.utils.util_random import ensure_rng
+from ..assigners import AssignResult
+
+
+def random_boxes(num=1, scale=1, rng=None):
+ """Simple version of ``kwimage.Boxes.random``
+
+ Returns:
+ Tensor: shape (n, 4) in x1, y1, x2, y2 format.
+
+ References:
+ https://gitlab.kitware.com/computer-vision/kwimage/blob/master/kwimage/structs/boxes.py#L1390
+
+ Example:
+ >>> num = 3
+ >>> scale = 512
+ >>> rng = 0
+ >>> boxes = random_boxes(num, scale, rng)
+ >>> print(boxes)
+ tensor([[280.9925, 278.9802, 308.6148, 366.1769],
+ [216.9113, 330.6978, 224.0446, 456.5878],
+ [405.3632, 196.3221, 493.3953, 270.7942]])
+ """
+ rng = ensure_rng(rng)
+
+ tlbr = rng.rand(num, 4).astype(np.float32)
+
+ tl_x = np.minimum(tlbr[:, 0], tlbr[:, 2])
+ tl_y = np.minimum(tlbr[:, 1], tlbr[:, 3])
+ br_x = np.maximum(tlbr[:, 0], tlbr[:, 2])
+ br_y = np.maximum(tlbr[:, 1], tlbr[:, 3])
+
+ tlbr[:, 0] = tl_x * scale
+ tlbr[:, 1] = tl_y * scale
+ tlbr[:, 2] = br_x * scale
+ tlbr[:, 3] = br_y * scale
+
+ boxes = torch.from_numpy(tlbr)
+ return boxes
+
+
+class SamplingResult(util_mixins.NiceRepr):
+ """Bbox sampling result.
+
+ Args:
+ pos_inds (Tensor): Indices of positive samples.
+ neg_inds (Tensor): Indices of negative samples.
+ priors (Tensor): The priors can be anchors or points,
+ or the bboxes predicted by the previous stage.
+ gt_bboxes (Tensor): Ground truth of bboxes.
+ assign_result (:obj:`AssignResult`): Assigning results.
+ gt_flags (Tensor): The Ground truth flags.
+ avg_factor_with_neg (bool): If True, ``avg_factor`` equal to
+ the number of total priors; Otherwise, it is the number of
+ positive priors. Defaults to True.
+
+ Example:
+ >>> # xdoctest: +IGNORE_WANT
+ >>> from mmdet.models.task_modules.samplers.sampling_result import * # NOQA
+ >>> self = SamplingResult.random(rng=10)
+ >>> print(f'self = {self}')
+ self =
+ """
+
+ def __init__(self,
+ pos_inds: Tensor,
+ neg_inds: Tensor,
+ priors: Tensor,
+ gt_bboxes: Tensor,
+ assign_result: AssignResult,
+ gt_flags: Tensor,
+ avg_factor_with_neg: bool = True) -> None:
+ self.pos_inds = pos_inds
+ self.neg_inds = neg_inds
+ self.num_pos = max(pos_inds.numel(), 1)
+ self.num_neg = max(neg_inds.numel(), 1)
+ self.avg_factor_with_neg = avg_factor_with_neg
+ self.avg_factor = self.num_pos + self.num_neg \
+ if avg_factor_with_neg else self.num_pos
+ self.pos_priors = priors[pos_inds]
+ self.neg_priors = priors[neg_inds]
+ self.pos_is_gt = gt_flags[pos_inds]
+
+ self.num_gts = gt_bboxes.shape[0]
+ self.pos_assigned_gt_inds = assign_result.gt_inds[pos_inds] - 1
+ self.pos_gt_labels = assign_result.labels[pos_inds]
+ box_dim = gt_bboxes.box_dim if isinstance(gt_bboxes, BaseBoxes) else 4
+ if gt_bboxes.numel() == 0:
+ # hack for index error case
+ assert self.pos_assigned_gt_inds.numel() == 0
+ self.pos_gt_bboxes = gt_bboxes.view(-1, box_dim)
+ else:
+ if len(gt_bboxes.shape) < 2:
+ gt_bboxes = gt_bboxes.view(-1, box_dim)
+ self.pos_gt_bboxes = gt_bboxes[self.pos_assigned_gt_inds.long()]
+
+ @property
+ def priors(self):
+ """torch.Tensor: concatenated positive and negative priors"""
+ return cat_boxes([self.pos_priors, self.neg_priors])
+
+ @property
+ def bboxes(self):
+ """torch.Tensor: concatenated positive and negative boxes"""
+ warnings.warn('DeprecationWarning: bboxes is deprecated, '
+ 'please use "priors" instead')
+ return self.priors
+
+ @property
+ def pos_bboxes(self):
+ warnings.warn('DeprecationWarning: pos_bboxes is deprecated, '
+ 'please use "pos_priors" instead')
+ return self.pos_priors
+
+ @property
+ def neg_bboxes(self):
+ warnings.warn('DeprecationWarning: neg_bboxes is deprecated, '
+ 'please use "neg_priors" instead')
+ return self.neg_priors
+
+ def to(self, device):
+ """Change the device of the data inplace.
+
+ Example:
+ >>> self = SamplingResult.random()
+ >>> print(f'self = {self.to(None)}')
+ >>> # xdoctest: +REQUIRES(--gpu)
+ >>> print(f'self = {self.to(0)}')
+ """
+ _dict = self.__dict__
+ for key, value in _dict.items():
+ if isinstance(value, (torch.Tensor, BaseBoxes)):
+ _dict[key] = value.to(device)
+ return self
+
+ def __nice__(self):
+ data = self.info.copy()
+ data['pos_priors'] = data.pop('pos_priors').shape
+ data['neg_priors'] = data.pop('neg_priors').shape
+ parts = [f"'{k}': {v!r}" for k, v in sorted(data.items())]
+ body = ' ' + ',\n '.join(parts)
+ return '{\n' + body + '\n}'
+
+ @property
+ def info(self):
+ """Returns a dictionary of info about the object."""
+ return {
+ 'pos_inds': self.pos_inds,
+ 'neg_inds': self.neg_inds,
+ 'pos_priors': self.pos_priors,
+ 'neg_priors': self.neg_priors,
+ 'pos_is_gt': self.pos_is_gt,
+ 'num_gts': self.num_gts,
+ 'pos_assigned_gt_inds': self.pos_assigned_gt_inds,
+ 'num_pos': self.num_pos,
+ 'num_neg': self.num_neg,
+ 'avg_factor': self.avg_factor
+ }
+
+ @classmethod
+ def random(cls, rng=None, **kwargs):
+ """
+ Args:
+ rng (None | int | numpy.random.RandomState): seed or state.
+ kwargs (keyword arguments):
+ - num_preds: Number of predicted boxes.
+ - num_gts: Number of true boxes.
+ - p_ignore (float): Probability of a predicted box assigned to
+ an ignored truth.
+ - p_assigned (float): probability of a predicted box not being
+ assigned.
+
+ Returns:
+ :obj:`SamplingResult`: Randomly generated sampling result.
+
+ Example:
+ >>> from mmdet.models.task_modules.samplers.sampling_result import * # NOQA
+ >>> self = SamplingResult.random()
+ >>> print(self.__dict__)
+ """
+ from mmengine.structures import InstanceData
+
+ from mmdet.models.task_modules.assigners import AssignResult
+ from mmdet.models.task_modules.samplers import RandomSampler
+ rng = ensure_rng(rng)
+
+ # make probabilistic?
+ num = 32
+ pos_fraction = 0.5
+ neg_pos_ub = -1
+
+ assign_result = AssignResult.random(rng=rng, **kwargs)
+
+ # Note we could just compute an assignment
+ priors = random_boxes(assign_result.num_preds, rng=rng)
+ gt_bboxes = random_boxes(assign_result.num_gts, rng=rng)
+ gt_labels = torch.randint(
+ 0, 5, (assign_result.num_gts, ), dtype=torch.long)
+
+ pred_instances = InstanceData()
+ pred_instances.priors = priors
+
+ gt_instances = InstanceData()
+ gt_instances.bboxes = gt_bboxes
+ gt_instances.labels = gt_labels
+
+ add_gt_as_proposals = True
+
+ sampler = RandomSampler(
+ num,
+ pos_fraction,
+ neg_pos_ub=neg_pos_ub,
+ add_gt_as_proposals=add_gt_as_proposals,
+ rng=rng)
+ self = sampler.sample(
+ assign_result=assign_result,
+ pred_instances=pred_instances,
+ gt_instances=gt_instances)
+ return self
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/score_hlr_sampler.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/score_hlr_sampler.py
new file mode 100644
index 0000000000000000000000000000000000000000..0227585b92329625d053f1e9f8c161fd02af8aef
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/samplers/score_hlr_sampler.py
@@ -0,0 +1,290 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Union
+
+import torch
+from mmcv.ops import nms_match
+from mmengine.structures import InstanceData
+from numpy import ndarray
+from torch import Tensor
+
+from mmdet.registry import TASK_UTILS
+from mmdet.structures.bbox import bbox2roi
+from ..assigners import AssignResult
+from .base_sampler import BaseSampler
+from .sampling_result import SamplingResult
+
+
+@TASK_UTILS.register_module()
+class ScoreHLRSampler(BaseSampler):
+ r"""Importance-based Sample Reweighting (ISR_N), described in `Prime Sample
+ Attention in Object Detection `_.
+
+ Score hierarchical local rank (HLR) differentiates with RandomSampler in
+ negative part. It firstly computes Score-HLR in a two-step way,
+ then linearly maps score hlr to the loss weights.
+
+ Args:
+ num (int): Total number of sampled RoIs.
+ pos_fraction (float): Fraction of positive samples.
+ context (:obj:`BaseRoIHead`): RoI head that the sampler belongs to.
+ neg_pos_ub (int): Upper bound of the ratio of num negative to num
+ positive, -1 means no upper bound. Defaults to -1.
+ add_gt_as_proposals (bool): Whether to add ground truth as proposals.
+ Defaults to True.
+ k (float): Power of the non-linear mapping. Defaults to 0.5
+ bias (float): Shift of the non-linear mapping. Defaults to 0.
+ score_thr (float): Minimum score that a negative sample is to be
+ considered as valid bbox. Defaults to 0.05.
+ iou_thr (float): IoU threshold for NMS match. Defaults to 0.5.
+ """
+
+ def __init__(self,
+ num: int,
+ pos_fraction: float,
+ context,
+ neg_pos_ub: int = -1,
+ add_gt_as_proposals: bool = True,
+ k: float = 0.5,
+ bias: float = 0,
+ score_thr: float = 0.05,
+ iou_thr: float = 0.5,
+ **kwargs) -> None:
+ super().__init__(
+ num=num,
+ pos_fraction=pos_fraction,
+ neg_pos_ub=neg_pos_ub,
+ add_gt_as_proposals=add_gt_as_proposals)
+ self.k = k
+ self.bias = bias
+ self.score_thr = score_thr
+ self.iou_thr = iou_thr
+ self.context = context
+ # context of cascade detectors is a list, so distinguish them here.
+ if not hasattr(context, 'num_stages'):
+ self.bbox_roi_extractor = context.bbox_roi_extractor
+ self.bbox_head = context.bbox_head
+ self.with_shared_head = context.with_shared_head
+ if self.with_shared_head:
+ self.shared_head = context.shared_head
+ else:
+ self.bbox_roi_extractor = context.bbox_roi_extractor[
+ context.current_stage]
+ self.bbox_head = context.bbox_head[context.current_stage]
+
+ @staticmethod
+ def random_choice(gallery: Union[Tensor, ndarray, list],
+ num: int) -> Union[Tensor, ndarray]:
+ """Randomly select some elements from the gallery.
+
+ If `gallery` is a Tensor, the returned indices will be a Tensor;
+ If `gallery` is a ndarray or list, the returned indices will be a
+ ndarray.
+
+ Args:
+ gallery (Tensor or ndarray or list): indices pool.
+ num (int): expected sample num.
+
+ Returns:
+ Tensor or ndarray: sampled indices.
+ """
+ assert len(gallery) >= num
+
+ is_tensor = isinstance(gallery, torch.Tensor)
+ if not is_tensor:
+ if torch.cuda.is_available():
+ device = torch.cuda.current_device()
+ else:
+ device = 'cpu'
+ gallery = torch.tensor(gallery, dtype=torch.long, device=device)
+ perm = torch.randperm(gallery.numel(), device=gallery.device)[:num]
+ rand_inds = gallery[perm]
+ if not is_tensor:
+ rand_inds = rand_inds.cpu().numpy()
+ return rand_inds
+
+ def _sample_pos(self, assign_result: AssignResult, num_expected: int,
+ **kwargs) -> Union[Tensor, ndarray]:
+ """Randomly sample some positive samples.
+
+ Args:
+ assign_result (:obj:`AssignResult`): Bbox assigning results.
+ num_expected (int): The number of expected positive samples
+
+ Returns:
+ Tensor or ndarray: sampled indices.
+ """
+ pos_inds = torch.nonzero(assign_result.gt_inds > 0).flatten()
+ if pos_inds.numel() <= num_expected:
+ return pos_inds
+ else:
+ return self.random_choice(pos_inds, num_expected)
+
+ def _sample_neg(self, assign_result: AssignResult, num_expected: int,
+ bboxes: Tensor, feats: Tensor,
+ **kwargs) -> Union[Tensor, ndarray]:
+ """Sample negative samples.
+
+ Score-HLR sampler is done in the following steps:
+ 1. Take the maximum positive score prediction of each negative samples
+ as s_i.
+ 2. Filter out negative samples whose s_i <= score_thr, the left samples
+ are called valid samples.
+ 3. Use NMS-Match to divide valid samples into different groups,
+ samples in the same group will greatly overlap with each other
+ 4. Rank the matched samples in two-steps to get Score-HLR.
+ (1) In the same group, rank samples with their scores.
+ (2) In the same score rank across different groups,
+ rank samples with their scores again.
+ 5. Linearly map Score-HLR to the final label weights.
+
+ Args:
+ assign_result (:obj:`AssignResult`): result of assigner.
+ num_expected (int): Expected number of samples.
+ bboxes (Tensor): bbox to be sampled.
+ feats (Tensor): Features come from FPN.
+
+ Returns:
+ Tensor or ndarray: sampled indices.
+ """
+ neg_inds = torch.nonzero(assign_result.gt_inds == 0).flatten()
+ num_neg = neg_inds.size(0)
+ if num_neg == 0:
+ return neg_inds, None
+ with torch.no_grad():
+ neg_bboxes = bboxes[neg_inds]
+ neg_rois = bbox2roi([neg_bboxes])
+ bbox_result = self.context._bbox_forward(feats, neg_rois)
+ cls_score, bbox_pred = bbox_result['cls_score'], bbox_result[
+ 'bbox_pred']
+
+ ori_loss = self.bbox_head.loss(
+ cls_score=cls_score,
+ bbox_pred=None,
+ rois=None,
+ labels=neg_inds.new_full((num_neg, ),
+ self.bbox_head.num_classes),
+ label_weights=cls_score.new_ones(num_neg),
+ bbox_targets=None,
+ bbox_weights=None,
+ reduction_override='none')['loss_cls']
+
+ # filter out samples with the max score lower than score_thr
+ max_score, argmax_score = cls_score.softmax(-1)[:, :-1].max(-1)
+ valid_inds = (max_score > self.score_thr).nonzero().view(-1)
+ invalid_inds = (max_score <= self.score_thr).nonzero().view(-1)
+ num_valid = valid_inds.size(0)
+ num_invalid = invalid_inds.size(0)
+
+ num_expected = min(num_neg, num_expected)
+ num_hlr = min(num_valid, num_expected)
+ num_rand = num_expected - num_hlr
+ if num_valid > 0:
+ valid_rois = neg_rois[valid_inds]
+ valid_max_score = max_score[valid_inds]
+ valid_argmax_score = argmax_score[valid_inds]
+ valid_bbox_pred = bbox_pred[valid_inds]
+
+ # valid_bbox_pred shape: [num_valid, #num_classes, 4]
+ valid_bbox_pred = valid_bbox_pred.view(
+ valid_bbox_pred.size(0), -1, 4)
+ selected_bbox_pred = valid_bbox_pred[range(num_valid),
+ valid_argmax_score]
+ pred_bboxes = self.bbox_head.bbox_coder.decode(
+ valid_rois[:, 1:], selected_bbox_pred)
+ pred_bboxes_with_score = torch.cat(
+ [pred_bboxes, valid_max_score[:, None]], -1)
+ group = nms_match(pred_bboxes_with_score, self.iou_thr)
+
+ # imp: importance
+ imp = cls_score.new_zeros(num_valid)
+ for g in group:
+ g_score = valid_max_score[g]
+ # g_score has already sorted
+ rank = g_score.new_tensor(range(g_score.size(0)))
+ imp[g] = num_valid - rank + g_score
+ _, imp_rank_inds = imp.sort(descending=True)
+ _, imp_rank = imp_rank_inds.sort()
+ hlr_inds = imp_rank_inds[:num_expected]
+
+ if num_rand > 0:
+ rand_inds = torch.randperm(num_invalid)[:num_rand]
+ select_inds = torch.cat(
+ [valid_inds[hlr_inds], invalid_inds[rand_inds]])
+ else:
+ select_inds = valid_inds[hlr_inds]
+
+ neg_label_weights = cls_score.new_ones(num_expected)
+
+ up_bound = max(num_expected, num_valid)
+ imp_weights = (up_bound -
+ imp_rank[hlr_inds].float()) / up_bound
+ neg_label_weights[:num_hlr] = imp_weights
+ neg_label_weights[num_hlr:] = imp_weights.min()
+ neg_label_weights = (self.bias +
+ (1 - self.bias) * neg_label_weights).pow(
+ self.k)
+ ori_selected_loss = ori_loss[select_inds]
+ new_loss = ori_selected_loss * neg_label_weights
+ norm_ratio = ori_selected_loss.sum() / new_loss.sum()
+ neg_label_weights *= norm_ratio
+ else:
+ neg_label_weights = cls_score.new_ones(num_expected)
+ select_inds = torch.randperm(num_neg)[:num_expected]
+
+ return neg_inds[select_inds], neg_label_weights
+
+ def sample(self, assign_result: AssignResult, pred_instances: InstanceData,
+ gt_instances: InstanceData, **kwargs) -> SamplingResult:
+ """Sample positive and negative bboxes.
+
+ This is a simple implementation of bbox sampling given candidates,
+ assigning results and ground truth bboxes.
+
+ Args:
+ assign_result (:obj:`AssignResult`): Assigning results.
+ pred_instances (:obj:`InstanceData`): Instances of model
+ predictions. It includes ``priors``, and the priors can
+ be anchors or points, or the bboxes predicted by the
+ previous stage, has shape (n, 4). The bboxes predicted by
+ the current model or stage will be named ``bboxes``,
+ ``labels``, and ``scores``, the same as the ``InstanceData``
+ in other places.
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes``, with shape (k, 4),
+ and ``labels``, with shape (k, ).
+
+ Returns:
+ :obj:`SamplingResult`: Sampling result.
+ """
+ gt_bboxes = gt_instances.bboxes
+ priors = pred_instances.priors
+ gt_labels = gt_instances.labels
+
+ gt_flags = priors.new_zeros((priors.shape[0], ), dtype=torch.uint8)
+ if self.add_gt_as_proposals and len(gt_bboxes) > 0:
+ priors = torch.cat([gt_bboxes, priors], dim=0)
+ assign_result.add_gt_(gt_labels)
+ gt_ones = priors.new_ones(gt_bboxes.shape[0], dtype=torch.uint8)
+ gt_flags = torch.cat([gt_ones, gt_flags])
+
+ num_expected_pos = int(self.num * self.pos_fraction)
+ pos_inds = self.pos_sampler._sample_pos(
+ assign_result, num_expected_pos, bboxes=priors, **kwargs)
+ num_sampled_pos = pos_inds.numel()
+ num_expected_neg = self.num - num_sampled_pos
+ if self.neg_pos_ub >= 0:
+ _pos = max(1, num_sampled_pos)
+ neg_upper_bound = int(self.neg_pos_ub * _pos)
+ if num_expected_neg > neg_upper_bound:
+ num_expected_neg = neg_upper_bound
+ neg_inds, neg_label_weights = self.neg_sampler._sample_neg(
+ assign_result, num_expected_neg, bboxes=priors, **kwargs)
+
+ sampling_result = SamplingResult(
+ pos_inds=pos_inds,
+ neg_inds=neg_inds,
+ priors=priors,
+ gt_bboxes=gt_bboxes,
+ assign_result=assign_result,
+ gt_flags=gt_flags)
+ return sampling_result, neg_label_weights
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/tracking/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/tracking/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..57a86d739d586e47e007d26de4542d6bdeced755
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/tracking/__init__.py
@@ -0,0 +1,11 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .aflink import AppearanceFreeLink
+from .camera_motion_compensation import CameraMotionCompensation
+from .interpolation import InterpolateTracklets
+from .kalman_filter import KalmanFilter
+from .similarity import embed_similarity
+
+__all__ = [
+ 'KalmanFilter', 'InterpolateTracklets', 'embed_similarity',
+ 'AppearanceFreeLink', 'CameraMotionCompensation'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/tracking/aflink.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/tracking/aflink.py
new file mode 100644
index 0000000000000000000000000000000000000000..52461067e372b30bbd28325ead00f5381c546326
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/tracking/aflink.py
@@ -0,0 +1,281 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from collections import defaultdict
+from typing import Tuple
+
+import numpy as np
+import torch
+from mmengine.model import BaseModule
+from mmengine.runner.checkpoint import load_checkpoint
+from scipy.optimize import linear_sum_assignment
+from torch import Tensor, nn
+
+from mmdet.registry import TASK_UTILS
+
+INFINITY = 1e5
+
+
+class TemporalBlock(BaseModule):
+ """The temporal block of AFLink model.
+
+ Args:
+ in_channel (int): the dimension of the input channels.
+ out_channel (int): the dimension of the output channels.
+ """
+
+ def __init__(self,
+ in_channel: int,
+ out_channel: int,
+ kernel_size: tuple = (7, 1)):
+ super(TemporalBlock, self).__init__()
+ self.conv = nn.Conv2d(in_channel, out_channel, kernel_size, bias=False)
+ self.relu = nn.ReLU(inplace=True)
+ self.bnf = nn.BatchNorm1d(out_channel)
+ self.bnx = nn.BatchNorm1d(out_channel)
+ self.bny = nn.BatchNorm1d(out_channel)
+
+ def bn(self, x: Tensor) -> Tensor:
+ x[:, :, :, 0] = self.bnf(x[:, :, :, 0])
+ x[:, :, :, 1] = self.bnx(x[:, :, :, 1])
+ x[:, :, :, 2] = self.bny(x[:, :, :, 2])
+ return x
+
+ def forward(self, x: Tensor) -> Tensor:
+ x = self.conv(x)
+ x = self.bn(x)
+ x = self.relu(x)
+ return x
+
+
+class FusionBlock(BaseModule):
+ """The fusion block of AFLink model.
+
+ Args:
+ in_channel (int): the dimension of the input channels.
+ out_channel (int): the dimension of the output channels.
+ """
+
+ def __init__(self, in_channel: int, out_channel: int):
+ super(FusionBlock, self).__init__()
+ self.conv = nn.Conv2d(in_channel, out_channel, (1, 3), bias=False)
+ self.bn = nn.BatchNorm2d(out_channel)
+ self.relu = nn.ReLU(inplace=True)
+
+ def forward(self, x: Tensor) -> Tensor:
+ x = self.conv(x)
+ x = self.bn(x)
+ x = self.relu(x)
+ return x
+
+
+class Classifier(BaseModule):
+ """The classifier of AFLink model.
+
+ Args:
+ in_channel (int): the dimension of the input channels.
+ """
+
+ def __init__(self, in_channel: int, out_channel: int):
+ super(Classifier, self).__init__()
+ self.fc1 = nn.Linear(in_channel * 2, in_channel // 2)
+ self.relu = nn.ReLU(inplace=True)
+ self.fc2 = nn.Linear(in_channel // 2, out_channel)
+
+ def forward(self, x1: Tensor, x2: Tensor) -> Tensor:
+ x = torch.cat((x1, x2), dim=1)
+ x = self.fc1(x)
+ x = self.relu(x)
+ x = self.fc2(x)
+ return x
+
+
+class AFLinkModel(BaseModule):
+ """Appearance-Free Link Model."""
+
+ def __init__(self,
+ temporal_module_channels: list = [1, 32, 64, 128, 256],
+ fusion_module_channels: list = [256, 256],
+ classifier_channels: list = [256, 2]):
+ super(AFLinkModel, self).__init__()
+ self.TemporalModule_1 = nn.Sequential(*[
+ TemporalBlock(temporal_module_channels[i],
+ temporal_module_channels[i + 1])
+ for i in range(len(temporal_module_channels) - 1)
+ ])
+
+ self.TemporalModule_2 = nn.Sequential(*[
+ TemporalBlock(temporal_module_channels[i],
+ temporal_module_channels[i + 1])
+ for i in range(len(temporal_module_channels) - 1)
+ ])
+
+ self.FusionBlock_1 = FusionBlock(*fusion_module_channels)
+ self.FusionBlock_2 = FusionBlock(*fusion_module_channels)
+
+ self.pooling = nn.AdaptiveAvgPool2d((1, 1))
+ self.classifier = Classifier(*classifier_channels)
+
+ def forward(self, x1: Tensor, x2: Tensor) -> Tensor:
+ assert not self.training, 'Only testing is supported for AFLink.'
+ x1 = x1[:, :, :, :3]
+ x2 = x2[:, :, :, :3]
+ x1 = self.TemporalModule_1(x1) # [B,1,30,3] -> [B,256,6,3]
+ x2 = self.TemporalModule_2(x2)
+ x1 = self.FusionBlock_1(x1)
+ x2 = self.FusionBlock_2(x2)
+ x1 = self.pooling(x1).squeeze(-1).squeeze(-1)
+ x2 = self.pooling(x2).squeeze(-1).squeeze(-1)
+ y = self.classifier(x1, x2)
+ y = torch.softmax(y, dim=1)[0, 1]
+ return y
+
+
+@TASK_UTILS.register_module()
+class AppearanceFreeLink(BaseModule):
+ """Appearance-Free Link method.
+
+ This method is proposed in
+ "StrongSORT: Make DeepSORT Great Again"
+ `StrongSORT`_.
+
+ Args:
+ checkpoint (str): Checkpoint path.
+ temporal_threshold (tuple, optional): The temporal constraint
+ for tracklets association. Defaults to (0, 30).
+ spatial_threshold (int, optional): The spatial constraint for
+ tracklets association. Defaults to 75.
+ confidence_threshold (float, optional): The minimum confidence
+ threshold for tracklets association. Defaults to 0.95.
+ """
+
+ def __init__(self,
+ checkpoint: str,
+ temporal_threshold: tuple = (0, 30),
+ spatial_threshold: int = 75,
+ confidence_threshold: float = 0.95):
+ super(AppearanceFreeLink, self).__init__()
+ self.temporal_threshold = temporal_threshold
+ self.spatial_threshold = spatial_threshold
+ self.confidence_threshold = confidence_threshold
+
+ self.model = AFLinkModel()
+ if checkpoint:
+ load_checkpoint(self.model, checkpoint)
+ if torch.cuda.is_available():
+ self.model.cuda()
+ self.model.eval()
+
+ self.device = next(self.model.parameters()).device
+ self.fn_l2 = lambda x, y: np.sqrt(x**2 + y**2)
+
+ def data_transform(self,
+ track1: np.ndarray,
+ track2: np.ndarray,
+ length: int = 30) -> Tuple[np.ndarray]:
+ """Data Transformation. This is used to standardize the length of
+ tracks to a unified length. Then perform min-max normalization to the
+ motion embeddings.
+
+ Args:
+ track1 (ndarray): the first track with shape (N,C).
+ track2 (ndarray): the second track with shape (M,C).
+ length (int): the unified length of tracks. Defaults to 30.
+
+ Returns:
+ Tuple[ndarray]: the transformed track1 and track2.
+ """
+ # fill or cut track1
+ length_1 = track1.shape[0]
+ track1 = track1[-length:] if length_1 >= length else \
+ np.pad(track1, ((length - length_1, 0), (0, 0)))
+
+ # fill or cut track1
+ length_2 = track2.shape[0]
+ track2 = track2[:length] if length_2 >= length else \
+ np.pad(track2, ((0, length - length_2), (0, 0)))
+
+ # min-max normalization
+ min_ = np.concatenate((track1, track2), axis=0).min(axis=0)
+ max_ = np.concatenate((track1, track2), axis=0).max(axis=0)
+ subtractor = (max_ + min_) / 2
+ divisor = (max_ - min_) / 2 + 1e-5
+ track1 = (track1 - subtractor) / divisor
+ track2 = (track2 - subtractor) / divisor
+
+ return track1, track2
+
+ def forward(self, pred_tracks: np.ndarray) -> np.ndarray:
+ """Forward function.
+
+ pred_tracks (ndarray): With shape (N, 7). Each row denotes
+ (frame_id, track_id, x1, y1, x2, y2, score).
+
+ Returns:
+ ndarray: The linked tracks with shape (N, 7). Each row denotes
+ (frame_id, track_id, x1, y1, x2, y2, score)
+ """
+ # sort tracks by the frame id
+ pred_tracks = pred_tracks[np.argsort(pred_tracks[:, 0])]
+
+ # gather tracks information
+ id2info = defaultdict(list)
+ for row in pred_tracks:
+ frame_id, track_id, x1, y1, x2, y2 = row[:6]
+ id2info[track_id].append([frame_id, x1, y1, x2 - x1, y2 - y1])
+ id2info = {k: np.array(v) for k, v in id2info.items()}
+ num_track = len(id2info)
+ track_ids = np.array(list(id2info))
+ cost_matrix = np.full((num_track, num_track), INFINITY)
+
+ # compute the cost matrix
+ for i, id_i in enumerate(track_ids):
+ for j, id_j in enumerate(track_ids):
+ if id_i == id_j:
+ continue
+ info_i, info_j = id2info[id_i], id2info[id_j]
+ frame_i, box_i = info_i[-1][0], info_i[-1][1:3]
+ frame_j, box_j = info_j[0][0], info_j[0][1:3]
+ # temporal constraint
+ if not self.temporal_threshold[0] <= \
+ frame_j - frame_i <= self.temporal_threshold[1]:
+ continue
+ # spatial constraint
+ if self.fn_l2(box_i[0] - box_j[0], box_i[1] - box_j[1]) \
+ > self.spatial_threshold:
+ continue
+ # confidence constraint
+ track_i, track_j = self.data_transform(info_i, info_j)
+
+ # numpy to torch
+ track_i = torch.tensor(
+ track_i, dtype=torch.float).to(self.device)
+ track_j = torch.tensor(
+ track_j, dtype=torch.float).to(self.device)
+ track_i = track_i.unsqueeze(0).unsqueeze(0)
+ track_j = track_j.unsqueeze(0).unsqueeze(0)
+
+ confidence = self.model(track_i,
+ track_j).detach().cpu().numpy()
+ if confidence >= self.confidence_threshold:
+ cost_matrix[i, j] = 1 - confidence
+
+ # linear assignment
+ indices = linear_sum_assignment(cost_matrix)
+ _id2id = dict() # the temporary assignment results
+ id2id = dict() # the final assignment results
+ for i, j in zip(indices[0], indices[1]):
+ if cost_matrix[i, j] < INFINITY:
+ _id2id[i] = j
+ for k, v in _id2id.items():
+ if k in id2id:
+ id2id[v] = id2id[k]
+ else:
+ id2id[v] = k
+
+ # link
+ for k, v in id2id.items():
+ pred_tracks[pred_tracks[:, 1] == k, 1] = v
+
+ # deduplicate
+ _, index = np.unique(pred_tracks[:, :2], return_index=True, axis=0)
+
+ return pred_tracks[index]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/tracking/camera_motion_compensation.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/tracking/camera_motion_compensation.py
new file mode 100644
index 0000000000000000000000000000000000000000..1a6298494fd1c24e0e7bba457dd50864725f98c8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/tracking/camera_motion_compensation.py
@@ -0,0 +1,104 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import cv2
+import numpy as np
+import torch
+from torch import Tensor
+
+from mmdet.registry import TASK_UTILS
+from mmdet.structures.bbox import bbox_cxcyah_to_xyxy, bbox_xyxy_to_cxcyah
+
+
+@TASK_UTILS.register_module()
+class CameraMotionCompensation:
+ """Camera motion compensation.
+
+ Args:
+ warp_mode (str): Warp mode in opencv.
+ Defaults to 'cv2.MOTION_EUCLIDEAN'.
+ num_iters (int): Number of the iterations. Defaults to 50.
+ stop_eps (float): Terminate threshold. Defaults to 0.001.
+ """
+
+ def __init__(self,
+ warp_mode: str = 'cv2.MOTION_EUCLIDEAN',
+ num_iters: int = 50,
+ stop_eps: float = 0.001):
+ self.warp_mode = eval(warp_mode)
+ self.num_iters = num_iters
+ self.stop_eps = stop_eps
+
+ def get_warp_matrix(self, img: np.ndarray, ref_img: np.ndarray) -> Tensor:
+ """Calculate warping matrix between two images."""
+ img = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY)
+ ref_img = cv2.cvtColor(ref_img, cv2.COLOR_BGR2GRAY)
+
+ warp_matrix = np.eye(2, 3, dtype=np.float32)
+ criteria = (cv2.TERM_CRITERIA_EPS | cv2.TERM_CRITERIA_COUNT,
+ self.num_iters, self.stop_eps)
+ cc, warp_matrix = cv2.findTransformECC(img, ref_img, warp_matrix,
+ self.warp_mode, criteria, None,
+ 1)
+ warp_matrix = torch.from_numpy(warp_matrix)
+ return warp_matrix
+
+ def warp_bboxes(self, bboxes: Tensor, warp_matrix: Tensor) -> Tensor:
+ """Warp bounding boxes according to the warping matrix."""
+ tl, br = bboxes[:, :2], bboxes[:, 2:]
+ tl = torch.cat((tl, torch.ones(tl.shape[0], 1).to(bboxes.device)),
+ dim=1)
+ br = torch.cat((br, torch.ones(tl.shape[0], 1).to(bboxes.device)),
+ dim=1)
+ trans_tl = torch.mm(warp_matrix, tl.t()).t()
+ trans_br = torch.mm(warp_matrix, br.t()).t()
+ trans_bboxes = torch.cat((trans_tl, trans_br), dim=1)
+ return trans_bboxes.to(bboxes.device)
+
+ def warp_means(self, means: np.ndarray, warp_matrix: Tensor) -> np.ndarray:
+ """Warp track.mean according to the warping matrix."""
+ cxcyah = torch.from_numpy(means[:, :4]).float()
+ xyxy = bbox_cxcyah_to_xyxy(cxcyah)
+ warped_xyxy = self.warp_bboxes(xyxy, warp_matrix)
+ warped_cxcyah = bbox_xyxy_to_cxcyah(warped_xyxy).numpy()
+ means[:, :4] = warped_cxcyah
+ return means
+
+ def track(self, img: Tensor, ref_img: Tensor, tracks: dict,
+ num_samples: int, frame_id: int, metainfo: dict) -> dict:
+ """Tracking forward."""
+ img = img.squeeze(0).cpu().numpy().transpose((1, 2, 0))
+ ref_img = ref_img.squeeze(0).cpu().numpy().transpose((1, 2, 0))
+ warp_matrix = self.get_warp_matrix(img, ref_img)
+
+ # rescale the warp_matrix due to the `resize` in pipeline
+ scale_factor_h, scale_factor_w = metainfo['scale_factor']
+ warp_matrix[0, 2] = warp_matrix[0, 2] / scale_factor_w
+ warp_matrix[1, 2] = warp_matrix[1, 2] / scale_factor_h
+
+ bboxes = []
+ num_bboxes = []
+ means = []
+ for k, v in tracks.items():
+ if int(v['frame_ids'][-1]) < frame_id - 1:
+ _num = 1
+ else:
+ _num = min(num_samples, len(v.bboxes))
+ num_bboxes.append(_num)
+ bboxes.extend(v.bboxes[-_num:])
+ if len(v.mean) > 0:
+ means.append(v.mean)
+ bboxes = torch.cat(bboxes, dim=0)
+ warped_bboxes = self.warp_bboxes(bboxes, warp_matrix.to(bboxes.device))
+
+ warped_bboxes = torch.split(warped_bboxes, num_bboxes)
+ for b, (k, v) in zip(warped_bboxes, tracks.items()):
+ _num = b.shape[0]
+ b = torch.split(b, [1] * _num)
+ tracks[k].bboxes[-_num:] = b
+
+ if means:
+ means = np.asarray(means)
+ warped_means = self.warp_means(means, warp_matrix)
+ for m, (k, v) in zip(warped_means, tracks.items()):
+ tracks[k].mean = m
+
+ return tracks
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/tracking/interpolation.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/tracking/interpolation.py
new file mode 100644
index 0000000000000000000000000000000000000000..fb6a25af4f253e3ec6b9781831ff43c6bafe50e1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/tracking/interpolation.py
@@ -0,0 +1,168 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import numpy as np
+
+try:
+ from sklearn.gaussian_process import GaussianProcessRegressor as GPR
+ from sklearn.gaussian_process.kernels import RBF
+ HAS_SKIKIT_LEARN = True
+except ImportError:
+ HAS_SKIKIT_LEARN = False
+
+from mmdet.registry import TASK_UTILS
+
+
+@TASK_UTILS.register_module()
+class InterpolateTracklets:
+ """Interpolate tracks to make tracks more complete.
+
+ Args:
+ min_num_frames (int, optional): The minimum length of a track that will
+ be interpolated. Defaults to 5.
+ max_num_frames (int, optional): The maximum disconnected length in
+ a track. Defaults to 20.
+ use_gsi (bool, optional): Whether to use the GSI (Gaussian-smoothed
+ interpolation) method. Defaults to False.
+ smooth_tau (int, optional): smoothing parameter in GSI. Defaults to 10.
+ """
+
+ def __init__(self,
+ min_num_frames: int = 5,
+ max_num_frames: int = 20,
+ use_gsi: bool = False,
+ smooth_tau: int = 10):
+ if not HAS_SKIKIT_LEARN:
+ raise RuntimeError('sscikit-learn is not installed,\
+ please install it by: pip install scikit-learn')
+ self.min_num_frames = min_num_frames
+ self.max_num_frames = max_num_frames
+ self.use_gsi = use_gsi
+ self.smooth_tau = smooth_tau
+
+ def _interpolate_track(self,
+ track: np.ndarray,
+ track_id: int,
+ max_num_frames: int = 20) -> np.ndarray:
+ """Interpolate a track linearly to make the track more complete.
+
+ This function is proposed in
+ "ByteTrack: Multi-Object Tracking by Associating Every Detection Box."
+ `ByteTrack`_.
+
+ Args:
+ track (ndarray): With shape (N, 7). Each row denotes
+ (frame_id, track_id, x1, y1, x2, y2, score).
+ max_num_frames (int, optional): The maximum disconnected length in
+ the track. Defaults to 20.
+
+ Returns:
+ ndarray: The interpolated track with shape (N, 7). Each row denotes
+ (frame_id, track_id, x1, y1, x2, y2, score)
+ """
+ assert (track[:, 1] == track_id).all(), \
+ 'The track id should not changed when interpolate a track.'
+
+ frame_ids = track[:, 0]
+ interpolated_track = np.zeros((0, 7))
+ # perform interpolation for the disconnected frames in the track.
+ for i in np.where(np.diff(frame_ids) > 1)[0]:
+ left_frame_id = frame_ids[i]
+ right_frame_id = frame_ids[i + 1]
+ num_disconnected_frames = int(right_frame_id - left_frame_id)
+
+ if 1 < num_disconnected_frames < max_num_frames:
+ left_bbox = track[i, 2:6]
+ right_bbox = track[i + 1, 2:6]
+
+ # perform interpolation for two adjacent tracklets.
+ for j in range(1, num_disconnected_frames):
+ cur_bbox = j / (num_disconnected_frames) * (
+ right_bbox - left_bbox) + left_bbox
+ cur_result = np.ones((7, ))
+ cur_result[0] = j + left_frame_id
+ cur_result[1] = track_id
+ cur_result[2:6] = cur_bbox
+
+ interpolated_track = np.concatenate(
+ (interpolated_track, cur_result[None]), axis=0)
+
+ interpolated_track = np.concatenate((track, interpolated_track),
+ axis=0)
+ return interpolated_track
+
+ def gaussian_smoothed_interpolation(self,
+ track: np.ndarray,
+ smooth_tau: int = 10) -> np.ndarray:
+ """Gaussian-Smoothed Interpolation.
+
+ This function is proposed in
+ "StrongSORT: Make DeepSORT Great Again"
+ `StrongSORT`_.
+
+ Args:
+ track (ndarray): With shape (N, 7). Each row denotes
+ (frame_id, track_id, x1, y1, x2, y2, score).
+ smooth_tau (int, optional): smoothing parameter in GSI.
+ Defaults to 10.
+
+ Returns:
+ ndarray: The interpolated tracks with shape (N, 7). Each row
+ denotes (frame_id, track_id, x1, y1, x2, y2, score)
+ """
+ len_scale = np.clip(smooth_tau * np.log(smooth_tau**3 / len(track)),
+ smooth_tau**-1, smooth_tau**2)
+ gpr = GPR(RBF(len_scale, 'fixed'))
+ t = track[:, 0].reshape(-1, 1)
+ x1 = track[:, 2].reshape(-1, 1)
+ y1 = track[:, 3].reshape(-1, 1)
+ x2 = track[:, 4].reshape(-1, 1)
+ y2 = track[:, 5].reshape(-1, 1)
+ gpr.fit(t, x1)
+ x1_gpr = gpr.predict(t)
+ gpr.fit(t, y1)
+ y1_gpr = gpr.predict(t)
+ gpr.fit(t, x2)
+ x2_gpr = gpr.predict(t)
+ gpr.fit(t, y2)
+ y2_gpr = gpr.predict(t)
+ gsi_track = [[
+ t[i, 0], track[i, 1], x1_gpr[i], y1_gpr[i], x2_gpr[i], y2_gpr[i],
+ track[i, 6]
+ ] for i in range(len(t))]
+ return np.array(gsi_track)
+
+ def forward(self, pred_tracks: np.ndarray) -> np.ndarray:
+ """Forward function.
+
+ pred_tracks (ndarray): With shape (N, 7). Each row denotes
+ (frame_id, track_id, x1, y1, x2, y2, score).
+
+ Returns:
+ ndarray: The interpolated tracks with shape (N, 7). Each row
+ denotes (frame_id, track_id, x1, y1, x2, y2, score).
+ """
+ max_track_id = int(np.max(pred_tracks[:, 1]))
+ min_track_id = int(np.min(pred_tracks[:, 1]))
+
+ # perform interpolation for each track
+ interpolated_tracks = []
+ for track_id in range(min_track_id, max_track_id + 1):
+ inds = pred_tracks[:, 1] == track_id
+ track = pred_tracks[inds]
+ num_frames = len(track)
+ if num_frames <= 2:
+ continue
+
+ if num_frames > self.min_num_frames:
+ interpolated_track = self._interpolate_track(
+ track, track_id, self.max_num_frames)
+ else:
+ interpolated_track = track
+
+ if self.use_gsi:
+ interpolated_track = self.gaussian_smoothed_interpolation(
+ interpolated_track, self.smooth_tau)
+
+ interpolated_tracks.append(interpolated_track)
+
+ interpolated_tracks = np.concatenate(interpolated_tracks)
+ return interpolated_tracks[interpolated_tracks[:, 0].argsort()]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/tracking/kalman_filter.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/tracking/kalman_filter.py
new file mode 100644
index 0000000000000000000000000000000000000000..a8ae1416af69bce17fd20dd5231eba2f12f7ed64
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/tracking/kalman_filter.py
@@ -0,0 +1,267 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Tuple
+
+import numpy as np
+import torch
+
+try:
+ import scipy.linalg
+ HAS_SCIPY = True
+except ImportError:
+ HAS_SCIPY = False
+
+from mmdet.registry import TASK_UTILS
+
+
+@TASK_UTILS.register_module()
+class KalmanFilter:
+ """A simple Kalman filter for tracking bounding boxes in image space.
+
+ The implementation is referred to https://github.com/nwojke/deep_sort.
+
+ Args:
+ center_only (bool): If True, distance computation is done with
+ respect to the bounding box center position only.
+ Defaults to False.
+ use_nsa (bool): Whether to use the NSA (Noise Scale Adaptive) Kalman
+ Filter, which adaptively modulates the noise scale according to
+ the quality of detections. More details in
+ https://arxiv.org/abs/2202.11983. Defaults to False.
+ """
+ chi2inv95 = {
+ 1: 3.8415,
+ 2: 5.9915,
+ 3: 7.8147,
+ 4: 9.4877,
+ 5: 11.070,
+ 6: 12.592,
+ 7: 14.067,
+ 8: 15.507,
+ 9: 16.919
+ }
+
+ def __init__(self, center_only: bool = False, use_nsa: bool = False):
+ if not HAS_SCIPY:
+ raise RuntimeError('sscikit-learn is not installed,\
+ please install it by: pip install scikit-learn')
+ self.center_only = center_only
+ if self.center_only:
+ self.gating_threshold = self.chi2inv95[2]
+ else:
+ self.gating_threshold = self.chi2inv95[4]
+
+ self.use_nsa = use_nsa
+ ndim, dt = 4, 1.
+
+ # Create Kalman filter model matrices.
+ self._motion_mat = np.eye(2 * ndim, 2 * ndim)
+ for i in range(ndim):
+ self._motion_mat[i, ndim + i] = dt
+ self._update_mat = np.eye(ndim, 2 * ndim)
+
+ # Motion and observation uncertainty are chosen relative to the current
+ # state estimate. These weights control the amount of uncertainty in
+ # the model. This is a bit hacky.
+ self._std_weight_position = 1. / 20
+ self._std_weight_velocity = 1. / 160
+
+ def initiate(self, measurement: np.array) -> Tuple[np.array, np.array]:
+ """Create track from unassociated measurement.
+
+ Args:
+ measurement (ndarray): Bounding box coordinates (x, y, a, h) with
+ center position (x, y), aspect ratio a, and height h.
+
+ Returns:
+ (ndarray, ndarray): Returns the mean vector (8 dimensional) and
+ covariance matrix (8x8 dimensional) of the new track.
+ Unobserved velocities are initialized to 0 mean.
+ """
+ mean_pos = measurement
+ mean_vel = np.zeros_like(mean_pos)
+ mean = np.r_[mean_pos, mean_vel]
+
+ std = [
+ 2 * self._std_weight_position * measurement[3],
+ 2 * self._std_weight_position * measurement[3], 1e-2,
+ 2 * self._std_weight_position * measurement[3],
+ 10 * self._std_weight_velocity * measurement[3],
+ 10 * self._std_weight_velocity * measurement[3], 1e-5,
+ 10 * self._std_weight_velocity * measurement[3]
+ ]
+ covariance = np.diag(np.square(std))
+ return mean, covariance
+
+ def predict(self, mean: np.array,
+ covariance: np.array) -> Tuple[np.array, np.array]:
+ """Run Kalman filter prediction step.
+
+ Args:
+ mean (ndarray): The 8 dimensional mean vector of the object
+ state at the previous time step.
+
+ covariance (ndarray): The 8x8 dimensional covariance matrix
+ of the object state at the previous time step.
+
+ Returns:
+ (ndarray, ndarray): Returns the mean vector and covariance
+ matrix of the predicted state. Unobserved velocities are
+ initialized to 0 mean.
+ """
+ std_pos = [
+ self._std_weight_position * mean[3],
+ self._std_weight_position * mean[3], 1e-2,
+ self._std_weight_position * mean[3]
+ ]
+ std_vel = [
+ self._std_weight_velocity * mean[3],
+ self._std_weight_velocity * mean[3], 1e-5,
+ self._std_weight_velocity * mean[3]
+ ]
+ motion_cov = np.diag(np.square(np.r_[std_pos, std_vel]))
+
+ mean = np.dot(self._motion_mat, mean)
+ covariance = np.linalg.multi_dot(
+ (self._motion_mat, covariance, self._motion_mat.T)) + motion_cov
+
+ return mean, covariance
+
+ def project(self,
+ mean: np.array,
+ covariance: np.array,
+ bbox_score: float = 0.) -> Tuple[np.array, np.array]:
+ """Project state distribution to measurement space.
+
+ Args:
+ mean (ndarray): The state's mean vector (8 dimensional array).
+ covariance (ndarray): The state's covariance matrix (8x8
+ dimensional).
+ bbox_score (float): The confidence score of the bbox.
+ Defaults to 0.
+
+ Returns:
+ (ndarray, ndarray): Returns the projected mean and covariance
+ matrix of the given state estimate.
+ """
+ std = [
+ self._std_weight_position * mean[3],
+ self._std_weight_position * mean[3], 1e-1,
+ self._std_weight_position * mean[3]
+ ]
+
+ if self.use_nsa:
+ std = [(1 - bbox_score) * x for x in std]
+
+ innovation_cov = np.diag(np.square(std))
+
+ mean = np.dot(self._update_mat, mean)
+ covariance = np.linalg.multi_dot(
+ (self._update_mat, covariance, self._update_mat.T))
+ return mean, covariance + innovation_cov
+
+ def update(self,
+ mean: np.array,
+ covariance: np.array,
+ measurement: np.array,
+ bbox_score: float = 0.) -> Tuple[np.array, np.array]:
+ """Run Kalman filter correction step.
+
+ Args:
+ mean (ndarray): The predicted state's mean vector (8 dimensional).
+ covariance (ndarray): The state's covariance matrix (8x8
+ dimensional).
+ measurement (ndarray): The 4 dimensional measurement vector
+ (x, y, a, h), where (x, y) is the center position, a the
+ aspect ratio, and h the height of the bounding box.
+ bbox_score (float): The confidence score of the bbox.
+ Defaults to 0.
+
+ Returns:
+ (ndarray, ndarray): Returns the measurement-corrected state
+ distribution.
+ """
+ projected_mean, projected_cov = \
+ self.project(mean, covariance, bbox_score)
+
+ chol_factor, lower = scipy.linalg.cho_factor(
+ projected_cov, lower=True, check_finite=False)
+ kalman_gain = scipy.linalg.cho_solve((chol_factor, lower),
+ np.dot(covariance,
+ self._update_mat.T).T,
+ check_finite=False).T
+ innovation = measurement - projected_mean
+
+ new_mean = mean + np.dot(innovation, kalman_gain.T)
+ new_covariance = covariance - np.linalg.multi_dot(
+ (kalman_gain, projected_cov, kalman_gain.T))
+ return new_mean, new_covariance
+
+ def gating_distance(self,
+ mean: np.array,
+ covariance: np.array,
+ measurements: np.array,
+ only_position: bool = False) -> np.array:
+ """Compute gating distance between state distribution and measurements.
+
+ A suitable distance threshold can be obtained from `chi2inv95`. If
+ `only_position` is False, the chi-square distribution has 4 degrees of
+ freedom, otherwise 2.
+
+ Args:
+ mean (ndarray): Mean vector over the state distribution (8
+ dimensional).
+ covariance (ndarray): Covariance of the state distribution (8x8
+ dimensional).
+ measurements (ndarray): An Nx4 dimensional matrix of N
+ measurements, each in format (x, y, a, h) where (x, y) is the
+ bounding box center position, a the aspect ratio, and h the
+ height.
+ only_position (bool, optional): If True, distance computation is
+ done with respect to the bounding box center position only.
+ Defaults to False.
+
+ Returns:
+ ndarray: Returns an array of length N, where the i-th element
+ contains the squared Mahalanobis distance between
+ (mean, covariance) and `measurements[i]`.
+ """
+ mean, covariance = self.project(mean, covariance)
+ if only_position:
+ mean, covariance = mean[:2], covariance[:2, :2]
+ measurements = measurements[:, :2]
+
+ cholesky_factor = np.linalg.cholesky(covariance)
+ d = measurements - mean
+ z = scipy.linalg.solve_triangular(
+ cholesky_factor,
+ d.T,
+ lower=True,
+ check_finite=False,
+ overwrite_b=True)
+ squared_maha = np.sum(z * z, axis=0)
+ return squared_maha
+
+ def track(self, tracks: dict,
+ bboxes: torch.Tensor) -> Tuple[dict, np.array]:
+ """Track forward.
+
+ Args:
+ tracks (dict[int:dict]): Track buffer.
+ bboxes (Tensor): Detected bounding boxes.
+
+ Returns:
+ (dict[int:dict], ndarray): Updated tracks and bboxes.
+ """
+ costs = []
+ for id, track in tracks.items():
+ track.mean, track.covariance = self.predict(
+ track.mean, track.covariance)
+ gating_distance = self.gating_distance(track.mean,
+ track.covariance,
+ bboxes.cpu().numpy(),
+ self.center_only)
+ costs.append(gating_distance)
+
+ costs = np.stack(costs, 0)
+ costs[costs > self.gating_threshold] = np.nan
+ return tracks, costs
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/tracking/similarity.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/tracking/similarity.py
new file mode 100644
index 0000000000000000000000000000000000000000..730e43b86214ae92ffdcab8ae39e6f9261075caa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/task_modules/tracking/similarity.py
@@ -0,0 +1,34 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+import torch.nn.functional as F
+from torch import Tensor
+
+
+def embed_similarity(key_embeds: Tensor,
+ ref_embeds: Tensor,
+ method: str = 'dot_product',
+ temperature: int = -1) -> Tensor:
+ """Calculate feature similarity from embeddings.
+
+ Args:
+ key_embeds (Tensor): Shape (N1, C).
+ ref_embeds (Tensor): Shape (N2, C).
+ method (str, optional): Method to calculate the similarity,
+ options are 'dot_product' and 'cosine'. Defaults to
+ 'dot_product'.
+ temperature (int, optional): Softmax temperature. Defaults to -1.
+
+ Returns:
+ Tensor: Similarity matrix of shape (N1, N2).
+ """
+ assert method in ['dot_product', 'cosine']
+
+ if method == 'cosine':
+ key_embeds = F.normalize(key_embeds, p=2, dim=1)
+ ref_embeds = F.normalize(ref_embeds, p=2, dim=1)
+
+ similarity = torch.mm(key_embeds, ref_embeds.T)
+
+ if temperature > 0:
+ similarity /= float(temperature)
+ return similarity
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/test_time_augs/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/test_time_augs/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..f5e4926efb011b45b3ab7d3d303fb2d105aaa192
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/test_time_augs/__init__.py
@@ -0,0 +1,10 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .det_tta import DetTTAModel
+from .merge_augs import (merge_aug_bboxes, merge_aug_masks,
+ merge_aug_proposals, merge_aug_results,
+ merge_aug_scores)
+
+__all__ = [
+ 'merge_aug_bboxes', 'merge_aug_masks', 'merge_aug_proposals',
+ 'merge_aug_scores', 'merge_aug_results', 'DetTTAModel'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/test_time_augs/det_tta.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/test_time_augs/det_tta.py
new file mode 100644
index 0000000000000000000000000000000000000000..95f91db9e1250358db0e1a572cf4c37cc7fe6e6f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/test_time_augs/det_tta.py
@@ -0,0 +1,144 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple
+
+import torch
+from mmcv.ops import batched_nms
+from mmengine.model import BaseTTAModel
+from mmengine.registry import MODELS
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.structures import DetDataSample
+from mmdet.structures.bbox import bbox_flip
+
+
+@MODELS.register_module()
+class DetTTAModel(BaseTTAModel):
+ """Merge augmented detection results, only bboxes corresponding score under
+ flipping and multi-scale resizing can be processed now.
+
+ Examples:
+ >>> tta_model = dict(
+ >>> type='DetTTAModel',
+ >>> tta_cfg=dict(nms=dict(
+ >>> type='nms',
+ >>> iou_threshold=0.5),
+ >>> max_per_img=100))
+ >>>
+ >>> tta_pipeline = [
+ >>> dict(type='LoadImageFromFile',
+ >>> backend_args=None),
+ >>> dict(
+ >>> type='TestTimeAug',
+ >>> transforms=[[
+ >>> dict(type='Resize',
+ >>> scale=(1333, 800),
+ >>> keep_ratio=True),
+ >>> ], [
+ >>> dict(type='RandomFlip', prob=1.),
+ >>> dict(type='RandomFlip', prob=0.)
+ >>> ], [
+ >>> dict(
+ >>> type='PackDetInputs',
+ >>> meta_keys=('img_id', 'img_path', 'ori_shape',
+ >>> 'img_shape', 'scale_factor', 'flip',
+ >>> 'flip_direction'))
+ >>> ]])]
+ """
+
+ def __init__(self, tta_cfg=None, **kwargs):
+ super().__init__(**kwargs)
+ self.tta_cfg = tta_cfg
+
+ def merge_aug_bboxes(self, aug_bboxes: List[Tensor],
+ aug_scores: List[Tensor],
+ img_metas: List[str]) -> Tuple[Tensor, Tensor]:
+ """Merge augmented detection bboxes and scores.
+
+ Args:
+ aug_bboxes (list[Tensor]): shape (n, 4*#class)
+ aug_scores (list[Tensor] or None): shape (n, #class)
+ Returns:
+ tuple[Tensor]: ``bboxes`` with shape (n,4), where
+ 4 represent (tl_x, tl_y, br_x, br_y)
+ and ``scores`` with shape (n,).
+ """
+ recovered_bboxes = []
+ for bboxes, img_info in zip(aug_bboxes, img_metas):
+ ori_shape = img_info['ori_shape']
+ flip = img_info['flip']
+ flip_direction = img_info['flip_direction']
+ if flip:
+ bboxes = bbox_flip(
+ bboxes=bboxes,
+ img_shape=ori_shape,
+ direction=flip_direction)
+ recovered_bboxes.append(bboxes)
+ bboxes = torch.cat(recovered_bboxes, dim=0)
+ if aug_scores is None:
+ return bboxes
+ else:
+ scores = torch.cat(aug_scores, dim=0)
+ return bboxes, scores
+
+ def merge_preds(self, data_samples_list: List[List[DetDataSample]]):
+ """Merge batch predictions of enhanced data.
+
+ Args:
+ data_samples_list (List[List[DetDataSample]]): List of predictions
+ of all enhanced data. The outer list indicates images, and the
+ inner list corresponds to the different views of one image.
+ Each element of the inner list is a ``DetDataSample``.
+ Returns:
+ List[DetDataSample]: Merged batch prediction.
+ """
+ merged_data_samples = []
+ for data_samples in data_samples_list:
+ merged_data_samples.append(self._merge_single_sample(data_samples))
+ return merged_data_samples
+
+ def _merge_single_sample(
+ self, data_samples: List[DetDataSample]) -> DetDataSample:
+ """Merge predictions which come form the different views of one image
+ to one prediction.
+
+ Args:
+ data_samples (List[DetDataSample]): List of predictions
+ of enhanced data which come form one image.
+ Returns:
+ List[DetDataSample]: Merged prediction.
+ """
+ aug_bboxes = []
+ aug_scores = []
+ aug_labels = []
+ img_metas = []
+ # TODO: support instance segmentation TTA
+ assert data_samples[0].pred_instances.get('masks', None) is None, \
+ 'TTA of instance segmentation does not support now.'
+ for data_sample in data_samples:
+ aug_bboxes.append(data_sample.pred_instances.bboxes)
+ aug_scores.append(data_sample.pred_instances.scores)
+ aug_labels.append(data_sample.pred_instances.labels)
+ img_metas.append(data_sample.metainfo)
+
+ merged_bboxes, merged_scores = self.merge_aug_bboxes(
+ aug_bboxes, aug_scores, img_metas)
+ merged_labels = torch.cat(aug_labels, dim=0)
+
+ if merged_bboxes.numel() == 0:
+ return data_samples[0]
+
+ det_bboxes, keep_idxs = batched_nms(merged_bboxes, merged_scores,
+ merged_labels, self.tta_cfg.nms)
+
+ det_bboxes = det_bboxes[:self.tta_cfg.max_per_img]
+ det_labels = merged_labels[keep_idxs][:self.tta_cfg.max_per_img]
+
+ results = InstanceData()
+ _det_bboxes = det_bboxes.clone()
+ results.bboxes = _det_bboxes[:, :-1]
+ results.scores = _det_bboxes[:, -1]
+ results.labels = det_labels
+ det_results = data_samples[0]
+ det_results.pred_instances = results
+ return det_results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/test_time_augs/merge_augs.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/test_time_augs/merge_augs.py
new file mode 100644
index 0000000000000000000000000000000000000000..5935a8614c39d70253a09a339f51c144661c64fb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/test_time_augs/merge_augs.py
@@ -0,0 +1,219 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import warnings
+from typing import List, Optional, Union
+
+import numpy as np
+import torch
+from mmcv.ops import nms
+from mmengine.config import ConfigDict
+from torch import Tensor
+
+from mmdet.structures.bbox import bbox_mapping_back
+
+
+# TODO remove this, never be used in mmdet
+def merge_aug_proposals(aug_proposals, img_metas, cfg):
+ """Merge augmented proposals (multiscale, flip, etc.)
+
+ Args:
+ aug_proposals (list[Tensor]): proposals from different testing
+ schemes, shape (n, 5). Note that they are not rescaled to the
+ original image size.
+
+ img_metas (list[dict]): list of image info dict where each dict has:
+ 'img_shape', 'scale_factor', 'flip', and may also contain
+ 'filename', 'ori_shape', 'pad_shape', and 'img_norm_cfg'.
+ For details on the values of these keys see
+ `mmdet/datasets/pipelines/formatting.py:Collect`.
+
+ cfg (dict): rpn test config.
+
+ Returns:
+ Tensor: shape (n, 4), proposals corresponding to original image scale.
+ """
+
+ cfg = copy.deepcopy(cfg)
+
+ # deprecate arguments warning
+ if 'nms' not in cfg or 'max_num' in cfg or 'nms_thr' in cfg:
+ warnings.warn(
+ 'In rpn_proposal or test_cfg, '
+ 'nms_thr has been moved to a dict named nms as '
+ 'iou_threshold, max_num has been renamed as max_per_img, '
+ 'name of original arguments and the way to specify '
+ 'iou_threshold of NMS will be deprecated.')
+ if 'nms' not in cfg:
+ cfg.nms = ConfigDict(dict(type='nms', iou_threshold=cfg.nms_thr))
+ if 'max_num' in cfg:
+ if 'max_per_img' in cfg:
+ assert cfg.max_num == cfg.max_per_img, f'You set max_num and ' \
+ f'max_per_img at the same time, but get {cfg.max_num} ' \
+ f'and {cfg.max_per_img} respectively' \
+ f'Please delete max_num which will be deprecated.'
+ else:
+ cfg.max_per_img = cfg.max_num
+ if 'nms_thr' in cfg:
+ assert cfg.nms.iou_threshold == cfg.nms_thr, f'You set ' \
+ f'iou_threshold in nms and ' \
+ f'nms_thr at the same time, but get ' \
+ f'{cfg.nms.iou_threshold} and {cfg.nms_thr}' \
+ f' respectively. Please delete the nms_thr ' \
+ f'which will be deprecated.'
+
+ recovered_proposals = []
+ for proposals, img_info in zip(aug_proposals, img_metas):
+ img_shape = img_info['img_shape']
+ scale_factor = img_info['scale_factor']
+ flip = img_info['flip']
+ flip_direction = img_info['flip_direction']
+ _proposals = proposals.clone()
+ _proposals[:, :4] = bbox_mapping_back(_proposals[:, :4], img_shape,
+ scale_factor, flip,
+ flip_direction)
+ recovered_proposals.append(_proposals)
+ aug_proposals = torch.cat(recovered_proposals, dim=0)
+ merged_proposals, _ = nms(aug_proposals[:, :4].contiguous(),
+ aug_proposals[:, -1].contiguous(),
+ cfg.nms.iou_threshold)
+ scores = merged_proposals[:, 4]
+ _, order = scores.sort(0, descending=True)
+ num = min(cfg.max_per_img, merged_proposals.shape[0])
+ order = order[:num]
+ merged_proposals = merged_proposals[order, :]
+ return merged_proposals
+
+
+# TODO remove this, never be used in mmdet
+def merge_aug_bboxes(aug_bboxes, aug_scores, img_metas, rcnn_test_cfg):
+ """Merge augmented detection bboxes and scores.
+
+ Args:
+ aug_bboxes (list[Tensor]): shape (n, 4*#class)
+ aug_scores (list[Tensor] or None): shape (n, #class)
+ img_shapes (list[Tensor]): shape (3, ).
+ rcnn_test_cfg (dict): rcnn test config.
+
+ Returns:
+ tuple: (bboxes, scores)
+ """
+ recovered_bboxes = []
+ for bboxes, img_info in zip(aug_bboxes, img_metas):
+ img_shape = img_info[0]['img_shape']
+ scale_factor = img_info[0]['scale_factor']
+ flip = img_info[0]['flip']
+ flip_direction = img_info[0]['flip_direction']
+ bboxes = bbox_mapping_back(bboxes, img_shape, scale_factor, flip,
+ flip_direction)
+ recovered_bboxes.append(bboxes)
+ bboxes = torch.stack(recovered_bboxes).mean(dim=0)
+ if aug_scores is None:
+ return bboxes
+ else:
+ scores = torch.stack(aug_scores).mean(dim=0)
+ return bboxes, scores
+
+
+def merge_aug_results(aug_batch_results, aug_batch_img_metas):
+ """Merge augmented detection results, only bboxes corresponding score under
+ flipping and multi-scale resizing can be processed now.
+
+ Args:
+ aug_batch_results (list[list[[obj:`InstanceData`]]):
+ Detection results of multiple images with
+ different augmentations.
+ The outer list indicate the augmentation . The inter
+ list indicate the batch dimension.
+ Each item usually contains the following keys.
+
+ - scores (Tensor): Classification scores, in shape
+ (num_instance,)
+ - labels (Tensor): Labels of bboxes, in shape
+ (num_instances,).
+ - bboxes (Tensor): In shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ aug_batch_img_metas (list[list[dict]]): The outer list
+ indicates test-time augs (multiscale, flip, etc.)
+ and the inner list indicates
+ images in a batch. Each dict in the list contains
+ information of an image in the batch.
+
+ Returns:
+ batch_results (list[obj:`InstanceData`]): Same with
+ the input `aug_results` except that all bboxes have
+ been mapped to the original scale.
+ """
+ num_augs = len(aug_batch_results)
+ num_imgs = len(aug_batch_results[0])
+
+ batch_results = []
+ aug_batch_results = copy.deepcopy(aug_batch_results)
+ for img_id in range(num_imgs):
+ aug_results = []
+ for aug_id in range(num_augs):
+ img_metas = aug_batch_img_metas[aug_id][img_id]
+ results = aug_batch_results[aug_id][img_id]
+
+ img_shape = img_metas['img_shape']
+ scale_factor = img_metas['scale_factor']
+ flip = img_metas['flip']
+ flip_direction = img_metas['flip_direction']
+ bboxes = bbox_mapping_back(results.bboxes, img_shape, scale_factor,
+ flip, flip_direction)
+ results.bboxes = bboxes
+ aug_results.append(results)
+ merged_aug_results = results.cat(aug_results)
+ batch_results.append(merged_aug_results)
+
+ return batch_results
+
+
+def merge_aug_scores(aug_scores):
+ """Merge augmented bbox scores."""
+ if isinstance(aug_scores[0], torch.Tensor):
+ return torch.mean(torch.stack(aug_scores), dim=0)
+ else:
+ return np.mean(aug_scores, axis=0)
+
+
+def merge_aug_masks(aug_masks: List[Tensor],
+ img_metas: dict,
+ weights: Optional[Union[list, Tensor]] = None) -> Tensor:
+ """Merge augmented mask prediction.
+
+ Args:
+ aug_masks (list[Tensor]): each has shape
+ (n, c, h, w).
+ img_metas (dict): Image information.
+ weights (list or Tensor): Weight of each aug_masks,
+ the length should be n.
+
+ Returns:
+ Tensor: has shape (n, c, h, w)
+ """
+ recovered_masks = []
+ for i, mask in enumerate(aug_masks):
+ if weights is not None:
+ assert len(weights) == len(aug_masks)
+ weight = weights[i]
+ else:
+ weight = 1
+ flip = img_metas.get('flip', False)
+ if flip:
+ flip_direction = img_metas['flip_direction']
+ if flip_direction == 'horizontal':
+ mask = mask[:, :, :, ::-1]
+ elif flip_direction == 'vertical':
+ mask = mask[:, :, ::-1, :]
+ elif flip_direction == 'diagonal':
+ mask = mask[:, :, :, ::-1]
+ mask = mask[:, :, ::-1, :]
+ else:
+ raise ValueError(
+ f"Invalid flipping direction '{flip_direction}'")
+ recovered_masks.append(mask[None, :] * weight)
+
+ merged_masks = torch.cat(recovered_masks, 0).mean(dim=0)
+ if weights is not None:
+ merged_masks = merged_masks * len(weights) / sum(weights)
+ return merged_masks
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..00284bb7b40dd007c28b6cc9175ac26a52c6c528
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/__init__.py
@@ -0,0 +1,13 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .base_tracker import BaseTracker
+from .byte_tracker import ByteTracker
+from .masktrack_rcnn_tracker import MaskTrackRCNNTracker
+from .ocsort_tracker import OCSORTTracker
+from .quasi_dense_tracker import QuasiDenseTracker
+from .sort_tracker import SORTTracker
+from .strongsort_tracker import StrongSORTTracker
+
+__all__ = [
+ 'BaseTracker', 'ByteTracker', 'QuasiDenseTracker', 'SORTTracker',
+ 'StrongSORTTracker', 'OCSORTTracker', 'MaskTrackRCNNTracker'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/base_tracker.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/base_tracker.py
new file mode 100644
index 0000000000000000000000000000000000000000..0cf188653cd9adda59decd45f65fc4ede63fe3a7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/base_tracker.py
@@ -0,0 +1,240 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from abc import ABCMeta, abstractmethod
+from typing import List, Optional, Tuple
+
+import torch
+import torch.nn.functional as F
+from addict import Dict
+
+
+class BaseTracker(metaclass=ABCMeta):
+ """Base tracker model.
+
+ Args:
+ momentums (dict[str:float], optional): Momentums to update the buffers.
+ The `str` indicates the name of the buffer while the `float`
+ indicates the momentum. Defaults to None.
+ num_frames_retain (int, optional). If a track is disappeared more than
+ `num_frames_retain` frames, it will be deleted in the memo.
+ Defaults to 10.
+ """
+
+ def __init__(self,
+ momentums: Optional[dict] = None,
+ num_frames_retain: int = 10) -> None:
+ super().__init__()
+ if momentums is not None:
+ assert isinstance(momentums, dict), 'momentums must be a dict'
+ self.momentums = momentums
+ self.num_frames_retain = num_frames_retain
+
+ self.reset()
+
+ def reset(self) -> None:
+ """Reset the buffer of the tracker."""
+ self.num_tracks = 0
+ self.tracks = dict()
+
+ @property
+ def empty(self) -> bool:
+ """Whether the buffer is empty or not."""
+ return False if self.tracks else True
+
+ @property
+ def ids(self) -> List[dict]:
+ """All ids in the tracker."""
+ return list(self.tracks.keys())
+
+ @property
+ def with_reid(self) -> bool:
+ """bool: whether the framework has a reid model"""
+ return hasattr(self, 'reid') and self.reid is not None
+
+ def update(self, **kwargs) -> None:
+ """Update the tracker.
+
+ Args:
+ kwargs (dict[str: Tensor | int]): The `str` indicates the
+ name of the input variable. `ids` and `frame_ids` are
+ obligatory in the keys.
+ """
+ memo_items = [k for k, v in kwargs.items() if v is not None]
+ rm_items = [k for k in kwargs.keys() if k not in memo_items]
+ for item in rm_items:
+ kwargs.pop(item)
+ if not hasattr(self, 'memo_items'):
+ self.memo_items = memo_items
+ else:
+ assert memo_items == self.memo_items
+
+ assert 'ids' in memo_items
+ num_objs = len(kwargs['ids'])
+ id_indice = memo_items.index('ids')
+ assert 'frame_ids' in memo_items
+ frame_id = int(kwargs['frame_ids'])
+ if isinstance(kwargs['frame_ids'], int):
+ kwargs['frame_ids'] = torch.tensor([kwargs['frame_ids']] *
+ num_objs)
+ # cur_frame_id = int(kwargs['frame_ids'][0])
+ for k, v in kwargs.items():
+ if len(v) != num_objs:
+ raise ValueError('kwargs value must both equal')
+
+ for obj in zip(*kwargs.values()):
+ id = int(obj[id_indice])
+ if id in self.tracks:
+ self.update_track(id, obj)
+ else:
+ self.init_track(id, obj)
+
+ self.pop_invalid_tracks(frame_id)
+
+ def pop_invalid_tracks(self, frame_id: int) -> None:
+ """Pop out invalid tracks."""
+ invalid_ids = []
+ for k, v in self.tracks.items():
+ if frame_id - v['frame_ids'][-1] >= self.num_frames_retain:
+ invalid_ids.append(k)
+ for invalid_id in invalid_ids:
+ self.tracks.pop(invalid_id)
+
+ def update_track(self, id: int, obj: Tuple[torch.Tensor]):
+ """Update a track."""
+ for k, v in zip(self.memo_items, obj):
+ v = v[None]
+ if self.momentums is not None and k in self.momentums:
+ m = self.momentums[k]
+ self.tracks[id][k] = (1 - m) * self.tracks[id][k] + m * v
+ else:
+ self.tracks[id][k].append(v)
+
+ def init_track(self, id: int, obj: Tuple[torch.Tensor]):
+ """Initialize a track."""
+ self.tracks[id] = Dict()
+ for k, v in zip(self.memo_items, obj):
+ v = v[None]
+ if self.momentums is not None and k in self.momentums:
+ self.tracks[id][k] = v
+ else:
+ self.tracks[id][k] = [v]
+
+ @property
+ def memo(self) -> dict:
+ """Return all buffers in the tracker."""
+ outs = Dict()
+ for k in self.memo_items:
+ outs[k] = []
+
+ for id, objs in self.tracks.items():
+ for k, v in objs.items():
+ if k not in outs:
+ continue
+ if self.momentums is not None and k in self.momentums:
+ v = v
+ else:
+ v = v[-1]
+ outs[k].append(v)
+
+ for k, v in outs.items():
+ outs[k] = torch.cat(v, dim=0)
+ return outs
+
+ def get(self,
+ item: str,
+ ids: Optional[list] = None,
+ num_samples: Optional[int] = None,
+ behavior: Optional[str] = None) -> torch.Tensor:
+ """Get the buffer of a specific item.
+
+ Args:
+ item (str): The demanded item.
+ ids (list[int], optional): The demanded ids. Defaults to None.
+ num_samples (int, optional): Number of samples to calculate the
+ results. Defaults to None.
+ behavior (str, optional): Behavior to calculate the results.
+ Options are `mean` | None. Defaults to None.
+
+ Returns:
+ Tensor: The results of the demanded item.
+ """
+ if ids is None:
+ ids = self.ids
+
+ outs = []
+ for id in ids:
+ out = self.tracks[id][item]
+ if isinstance(out, list):
+ if num_samples is not None:
+ out = out[-num_samples:]
+ out = torch.cat(out, dim=0)
+ if behavior == 'mean':
+ out = out.mean(dim=0, keepdim=True)
+ elif behavior is None:
+ out = out[None]
+ else:
+ raise NotImplementedError()
+ else:
+ out = out[-1]
+ outs.append(out)
+ return torch.cat(outs, dim=0)
+
+ @abstractmethod
+ def track(self, *args, **kwargs):
+ """Tracking forward function."""
+ pass
+
+ def crop_imgs(self,
+ img: torch.Tensor,
+ meta_info: dict,
+ bboxes: torch.Tensor,
+ rescale: bool = False) -> torch.Tensor:
+ """Crop the images according to some bounding boxes. Typically for re-
+ identification sub-module.
+
+ Args:
+ img (Tensor): of shape (T, C, H, W) encoding input image.
+ Typically these should be mean centered and std scaled.
+ meta_info (dict): image information dict where each dict
+ has: 'img_shape', 'scale_factor', 'flip', and may also contain
+ 'filename', 'ori_shape', 'pad_shape', and 'img_norm_cfg'.
+ bboxes (Tensor): of shape (N, 4) or (N, 5).
+ rescale (bool, optional): If True, the bounding boxes should be
+ rescaled to fit the scale of the image. Defaults to False.
+
+ Returns:
+ Tensor: Image tensor of shape (T, C, H, W).
+ """
+ h, w = meta_info['img_shape']
+ img = img[:, :, :h, :w]
+ if rescale:
+ factor_x, factor_y = meta_info['scale_factor']
+ bboxes[:, :4] *= torch.tensor(
+ [factor_x, factor_y, factor_x, factor_y]).to(bboxes.device)
+ bboxes[:, 0] = torch.clamp(bboxes[:, 0], min=0, max=w - 1)
+ bboxes[:, 1] = torch.clamp(bboxes[:, 1], min=0, max=h - 1)
+ bboxes[:, 2] = torch.clamp(bboxes[:, 2], min=1, max=w)
+ bboxes[:, 3] = torch.clamp(bboxes[:, 3], min=1, max=h)
+
+ crop_imgs = []
+ for bbox in bboxes:
+ x1, y1, x2, y2 = map(int, bbox)
+ if x2 <= x1:
+ x2 = x1 + 1
+ if y2 <= y1:
+ y2 = y1 + 1
+ crop_img = img[:, :, y1:y2, x1:x2]
+ if self.reid.get('img_scale', False):
+ crop_img = F.interpolate(
+ crop_img,
+ size=self.reid['img_scale'],
+ mode='bilinear',
+ align_corners=False)
+ crop_imgs.append(crop_img)
+
+ if len(crop_imgs) > 0:
+ return torch.cat(crop_imgs, dim=0)
+ elif self.reid.get('img_scale', False):
+ _h, _w = self.reid['img_scale']
+ return img.new_zeros((0, 3, _h, _w))
+ else:
+ return img.new_zeros((0, 3, h, w))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/byte_tracker.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/byte_tracker.py
new file mode 100644
index 0000000000000000000000000000000000000000..11f3adc53c58339f6289cbfa77aed738259fc98c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/byte_tracker.py
@@ -0,0 +1,334 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple
+
+try:
+ import lap
+except ImportError:
+ lap = None
+import numpy as np
+import torch
+from mmengine.structures import InstanceData
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures import DetDataSample
+from mmdet.structures.bbox import (bbox_cxcyah_to_xyxy, bbox_overlaps,
+ bbox_xyxy_to_cxcyah)
+from .base_tracker import BaseTracker
+
+
+@MODELS.register_module()
+class ByteTracker(BaseTracker):
+ """Tracker for ByteTrack.
+
+ Args:
+ motion (dict): Configuration of motion. Defaults to None.
+ obj_score_thrs (dict): Detection score threshold for matching objects.
+ - high (float): Threshold of the first matching. Defaults to 0.6.
+ - low (float): Threshold of the second matching. Defaults to 0.1.
+ init_track_thr (float): Detection score threshold for initializing a
+ new tracklet. Defaults to 0.7.
+ weight_iou_with_det_scores (bool): Whether using detection scores to
+ weight IOU which is used for matching. Defaults to True.
+ match_iou_thrs (dict): IOU distance threshold for matching between two
+ frames.
+ - high (float): Threshold of the first matching. Defaults to 0.1.
+ - low (float): Threshold of the second matching. Defaults to 0.5.
+ - tentative (float): Threshold of the matching for tentative
+ tracklets. Defaults to 0.3.
+ num_tentatives (int, optional): Number of continuous frames to confirm
+ a track. Defaults to 3.
+ """
+
+ def __init__(self,
+ motion: Optional[dict] = None,
+ obj_score_thrs: dict = dict(high=0.6, low=0.1),
+ init_track_thr: float = 0.7,
+ weight_iou_with_det_scores: bool = True,
+ match_iou_thrs: dict = dict(high=0.1, low=0.5, tentative=0.3),
+ num_tentatives: int = 3,
+ **kwargs):
+ super().__init__(**kwargs)
+
+ if lap is None:
+ raise RuntimeError('lap is not installed,\
+ please install it by: pip install lap')
+ if motion is not None:
+ self.motion = TASK_UTILS.build(motion)
+
+ self.obj_score_thrs = obj_score_thrs
+ self.init_track_thr = init_track_thr
+
+ self.weight_iou_with_det_scores = weight_iou_with_det_scores
+ self.match_iou_thrs = match_iou_thrs
+
+ self.num_tentatives = num_tentatives
+
+ @property
+ def confirmed_ids(self) -> List:
+ """Confirmed ids in the tracker."""
+ ids = [id for id, track in self.tracks.items() if not track.tentative]
+ return ids
+
+ @property
+ def unconfirmed_ids(self) -> List:
+ """Unconfirmed ids in the tracker."""
+ ids = [id for id, track in self.tracks.items() if track.tentative]
+ return ids
+
+ def init_track(self, id: int, obj: Tuple[torch.Tensor]) -> None:
+ """Initialize a track."""
+ super().init_track(id, obj)
+ if self.tracks[id].frame_ids[-1] == 0:
+ self.tracks[id].tentative = False
+ else:
+ self.tracks[id].tentative = True
+ bbox = bbox_xyxy_to_cxcyah(self.tracks[id].bboxes[-1]) # size = (1, 4)
+ assert bbox.ndim == 2 and bbox.shape[0] == 1
+ bbox = bbox.squeeze(0).cpu().numpy()
+ self.tracks[id].mean, self.tracks[id].covariance = self.kf.initiate(
+ bbox)
+
+ def update_track(self, id: int, obj: Tuple[torch.Tensor]) -> None:
+ """Update a track."""
+ super().update_track(id, obj)
+ if self.tracks[id].tentative:
+ if len(self.tracks[id]['bboxes']) >= self.num_tentatives:
+ self.tracks[id].tentative = False
+ bbox = bbox_xyxy_to_cxcyah(self.tracks[id].bboxes[-1]) # size = (1, 4)
+ assert bbox.ndim == 2 and bbox.shape[0] == 1
+ bbox = bbox.squeeze(0).cpu().numpy()
+ track_label = self.tracks[id]['labels'][-1]
+ label_idx = self.memo_items.index('labels')
+ obj_label = obj[label_idx]
+ assert obj_label == track_label
+ self.tracks[id].mean, self.tracks[id].covariance = self.kf.update(
+ self.tracks[id].mean, self.tracks[id].covariance, bbox)
+
+ def pop_invalid_tracks(self, frame_id: int) -> None:
+ """Pop out invalid tracks."""
+ invalid_ids = []
+ for k, v in self.tracks.items():
+ # case1: disappeared frames >= self.num_frames_retrain
+ case1 = frame_id - v['frame_ids'][-1] >= self.num_frames_retain
+ # case2: tentative tracks but not matched in this frame
+ case2 = v.tentative and v['frame_ids'][-1] != frame_id
+ if case1 or case2:
+ invalid_ids.append(k)
+ for invalid_id in invalid_ids:
+ self.tracks.pop(invalid_id)
+
+ def assign_ids(
+ self,
+ ids: List[int],
+ det_bboxes: torch.Tensor,
+ det_labels: torch.Tensor,
+ det_scores: torch.Tensor,
+ weight_iou_with_det_scores: Optional[bool] = False,
+ match_iou_thr: Optional[float] = 0.5
+ ) -> Tuple[np.ndarray, np.ndarray]:
+ """Assign ids.
+
+ Args:
+ ids (list[int]): Tracking ids.
+ det_bboxes (Tensor): of shape (N, 4)
+ det_labels (Tensor): of shape (N,)
+ det_scores (Tensor): of shape (N,)
+ weight_iou_with_det_scores (bool, optional): Whether using
+ detection scores to weight IOU which is used for matching.
+ Defaults to False.
+ match_iou_thr (float, optional): Matching threshold.
+ Defaults to 0.5.
+
+ Returns:
+ tuple(np.ndarray, np.ndarray): The assigning ids.
+ """
+ # get track_bboxes
+ track_bboxes = np.zeros((0, 4))
+ for id in ids:
+ track_bboxes = np.concatenate(
+ (track_bboxes, self.tracks[id].mean[:4][None]), axis=0)
+ track_bboxes = torch.from_numpy(track_bboxes).to(det_bboxes)
+ track_bboxes = bbox_cxcyah_to_xyxy(track_bboxes)
+
+ # compute distance
+ ious = bbox_overlaps(track_bboxes, det_bboxes)
+ if weight_iou_with_det_scores:
+ ious *= det_scores
+ # support multi-class association
+ track_labels = torch.tensor([
+ self.tracks[id]['labels'][-1] for id in ids
+ ]).to(det_bboxes.device)
+
+ cate_match = det_labels[None, :] == track_labels[:, None]
+ # to avoid det and track of different categories are matched
+ cate_cost = (1 - cate_match.int()) * 1e6
+
+ dists = (1 - ious + cate_cost).cpu().numpy()
+
+ # bipartite match
+ if dists.size > 0:
+ cost, row, col = lap.lapjv(
+ dists, extend_cost=True, cost_limit=1 - match_iou_thr)
+ else:
+ row = np.zeros(len(ids)).astype(np.int32) - 1
+ col = np.zeros(len(det_bboxes)).astype(np.int32) - 1
+ return row, col
+
+ def track(self, data_sample: DetDataSample, **kwargs) -> InstanceData:
+ """Tracking forward function.
+
+ Args:
+ data_sample (:obj:`DetDataSample`): The data sample.
+ It includes information such as `pred_instances`.
+
+ Returns:
+ :obj:`InstanceData`: Tracking results of the input images.
+ Each InstanceData usually contains ``bboxes``, ``labels``,
+ ``scores`` and ``instances_id``.
+ """
+ metainfo = data_sample.metainfo
+ bboxes = data_sample.pred_instances.bboxes
+ labels = data_sample.pred_instances.labels
+ scores = data_sample.pred_instances.scores
+
+ frame_id = metainfo.get('frame_id', -1)
+ if frame_id == 0:
+ self.reset()
+ if not hasattr(self, 'kf'):
+ self.kf = self.motion
+
+ if self.empty or bboxes.size(0) == 0:
+ valid_inds = scores > self.init_track_thr
+ scores = scores[valid_inds]
+ bboxes = bboxes[valid_inds]
+ labels = labels[valid_inds]
+ num_new_tracks = bboxes.size(0)
+ ids = torch.arange(self.num_tracks,
+ self.num_tracks + num_new_tracks).to(labels)
+ self.num_tracks += num_new_tracks
+
+ else:
+ # 0. init
+ ids = torch.full((bboxes.size(0), ),
+ -1,
+ dtype=labels.dtype,
+ device=labels.device)
+
+ # get the detection bboxes for the first association
+ first_det_inds = scores > self.obj_score_thrs['high']
+ first_det_bboxes = bboxes[first_det_inds]
+ first_det_labels = labels[first_det_inds]
+ first_det_scores = scores[first_det_inds]
+ first_det_ids = ids[first_det_inds]
+
+ # get the detection bboxes for the second association
+ second_det_inds = (~first_det_inds) & (
+ scores > self.obj_score_thrs['low'])
+ second_det_bboxes = bboxes[second_det_inds]
+ second_det_labels = labels[second_det_inds]
+ second_det_scores = scores[second_det_inds]
+ second_det_ids = ids[second_det_inds]
+
+ # 1. use Kalman Filter to predict current location
+ for id in self.confirmed_ids:
+ # track is lost in previous frame
+ if self.tracks[id].frame_ids[-1] != frame_id - 1:
+ self.tracks[id].mean[7] = 0
+ (self.tracks[id].mean,
+ self.tracks[id].covariance) = self.kf.predict(
+ self.tracks[id].mean, self.tracks[id].covariance)
+
+ # 2. first match
+ first_match_track_inds, first_match_det_inds = self.assign_ids(
+ self.confirmed_ids, first_det_bboxes, first_det_labels,
+ first_det_scores, self.weight_iou_with_det_scores,
+ self.match_iou_thrs['high'])
+ # '-1' mean a detection box is not matched with tracklets in
+ # previous frame
+ valid = first_match_det_inds > -1
+ first_det_ids[valid] = torch.tensor(
+ self.confirmed_ids)[first_match_det_inds[valid]].to(labels)
+
+ first_match_det_bboxes = first_det_bboxes[valid]
+ first_match_det_labels = first_det_labels[valid]
+ first_match_det_scores = first_det_scores[valid]
+ first_match_det_ids = first_det_ids[valid]
+ assert (first_match_det_ids > -1).all()
+
+ first_unmatch_det_bboxes = first_det_bboxes[~valid]
+ first_unmatch_det_labels = first_det_labels[~valid]
+ first_unmatch_det_scores = first_det_scores[~valid]
+ first_unmatch_det_ids = first_det_ids[~valid]
+ assert (first_unmatch_det_ids == -1).all()
+
+ # 3. use unmatched detection bboxes from the first match to match
+ # the unconfirmed tracks
+ (tentative_match_track_inds,
+ tentative_match_det_inds) = self.assign_ids(
+ self.unconfirmed_ids, first_unmatch_det_bboxes,
+ first_unmatch_det_labels, first_unmatch_det_scores,
+ self.weight_iou_with_det_scores,
+ self.match_iou_thrs['tentative'])
+ valid = tentative_match_det_inds > -1
+ first_unmatch_det_ids[valid] = torch.tensor(self.unconfirmed_ids)[
+ tentative_match_det_inds[valid]].to(labels)
+
+ # 4. second match for unmatched tracks from the first match
+ first_unmatch_track_ids = []
+ for i, id in enumerate(self.confirmed_ids):
+ # tracklet is not matched in the first match
+ case_1 = first_match_track_inds[i] == -1
+ # tracklet is not lost in the previous frame
+ case_2 = self.tracks[id].frame_ids[-1] == frame_id - 1
+ if case_1 and case_2:
+ first_unmatch_track_ids.append(id)
+
+ second_match_track_inds, second_match_det_inds = self.assign_ids(
+ first_unmatch_track_ids, second_det_bboxes, second_det_labels,
+ second_det_scores, False, self.match_iou_thrs['low'])
+ valid = second_match_det_inds > -1
+ second_det_ids[valid] = torch.tensor(first_unmatch_track_ids)[
+ second_match_det_inds[valid]].to(ids)
+
+ # 5. gather all matched detection bboxes from step 2-4
+ # we only keep matched detection bboxes in second match, which
+ # means the id != -1
+ valid = second_det_ids > -1
+ bboxes = torch.cat(
+ (first_match_det_bboxes, first_unmatch_det_bboxes), dim=0)
+ bboxes = torch.cat((bboxes, second_det_bboxes[valid]), dim=0)
+
+ labels = torch.cat(
+ (first_match_det_labels, first_unmatch_det_labels), dim=0)
+ labels = torch.cat((labels, second_det_labels[valid]), dim=0)
+
+ scores = torch.cat(
+ (first_match_det_scores, first_unmatch_det_scores), dim=0)
+ scores = torch.cat((scores, second_det_scores[valid]), dim=0)
+
+ ids = torch.cat((first_match_det_ids, first_unmatch_det_ids),
+ dim=0)
+ ids = torch.cat((ids, second_det_ids[valid]), dim=0)
+
+ # 6. assign new ids
+ new_track_inds = ids == -1
+ ids[new_track_inds] = torch.arange(
+ self.num_tracks,
+ self.num_tracks + new_track_inds.sum()).to(labels)
+ self.num_tracks += new_track_inds.sum()
+
+ self.update(
+ ids=ids,
+ bboxes=bboxes,
+ scores=scores,
+ labels=labels,
+ frame_ids=frame_id)
+
+ # update pred_track_instances
+ pred_track_instances = InstanceData()
+ pred_track_instances.bboxes = bboxes
+ pred_track_instances.labels = labels
+ pred_track_instances.scores = scores
+ pred_track_instances.instances_id = ids
+
+ return pred_track_instances
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/masktrack_rcnn_tracker.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/masktrack_rcnn_tracker.py
new file mode 100644
index 0000000000000000000000000000000000000000..cc167786b8b412629885a4f134a1bf79f3dfaa93
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/masktrack_rcnn_tracker.py
@@ -0,0 +1,189 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List
+
+import torch
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import DetDataSample
+from mmdet.structures.bbox import bbox_overlaps
+from .base_tracker import BaseTracker
+
+
+@MODELS.register_module()
+class MaskTrackRCNNTracker(BaseTracker):
+ """Tracker for MaskTrack R-CNN.
+
+ Args:
+ match_weights (dict[str : float]): The Weighting factor when computing
+ the match score. It contains keys as follows:
+
+ - det_score (float): The coefficient of `det_score` when computing
+ match score.
+ - iou (float): The coefficient of `ious` when computing match
+ score.
+ - det_label (float): The coefficient of `label_deltas` when
+ computing match score.
+ """
+
+ def __init__(self,
+ match_weights: dict = dict(
+ det_score=1.0, iou=2.0, det_label=10.0),
+ **kwargs):
+ super().__init__(**kwargs)
+ self.match_weights = match_weights
+
+ def get_match_score(self, bboxes: Tensor, labels: Tensor, scores: Tensor,
+ prev_bboxes: Tensor, prev_labels: Tensor,
+ similarity_logits: Tensor) -> Tensor:
+ """Get the match score.
+
+ Args:
+ bboxes (torch.Tensor): of shape (num_current_bboxes, 4) in
+ [tl_x, tl_y, br_x, br_y] format. Denoting the detection
+ bboxes of current frame.
+ labels (torch.Tensor): of shape (num_current_bboxes, )
+ scores (torch.Tensor): of shape (num_current_bboxes, )
+ prev_bboxes (torch.Tensor): of shape (num_previous_bboxes, 4) in
+ [tl_x, tl_y, br_x, br_y] format. Denoting the detection bboxes
+ of previous frame.
+ prev_labels (torch.Tensor): of shape (num_previous_bboxes, )
+ similarity_logits (torch.Tensor): of shape (num_current_bboxes,
+ num_previous_bboxes + 1). Denoting the similarity logits from
+ track head.
+
+ Returns:
+ torch.Tensor: The matching score of shape (num_current_bboxes,
+ num_previous_bboxes + 1)
+ """
+ similarity_scores = similarity_logits.softmax(dim=1)
+
+ ious = bbox_overlaps(bboxes, prev_bboxes)
+ iou_dummy = ious.new_zeros(ious.shape[0], 1)
+ ious = torch.cat((iou_dummy, ious), dim=1)
+
+ label_deltas = (labels.view(-1, 1) == prev_labels).float()
+ label_deltas_dummy = label_deltas.new_ones(label_deltas.shape[0], 1)
+ label_deltas = torch.cat((label_deltas_dummy, label_deltas), dim=1)
+
+ match_score = similarity_scores.log()
+ match_score += self.match_weights['det_score'] * \
+ scores.view(-1, 1).log()
+ match_score += self.match_weights['iou'] * ious
+ match_score += self.match_weights['det_label'] * label_deltas
+
+ return match_score
+
+ def assign_ids(self, match_scores: Tensor):
+ num_prev_bboxes = match_scores.shape[1] - 1
+ _, match_ids = match_scores.max(dim=1)
+
+ ids = match_ids.new_zeros(match_ids.shape[0]) - 1
+ best_match_scores = match_scores.new_zeros(num_prev_bboxes) - 1e6
+ for idx, match_id in enumerate(match_ids):
+ if match_id == 0:
+ ids[idx] = self.num_tracks
+ self.num_tracks += 1
+ else:
+ match_score = match_scores[idx, match_id]
+ # TODO: fix the bug where multiple candidate might match
+ # with the same previous object.
+ if match_score > best_match_scores[match_id - 1]:
+ ids[idx] = self.ids[match_id - 1]
+ best_match_scores[match_id - 1] = match_score
+ return ids, best_match_scores
+
+ def track(self,
+ model: torch.nn.Module,
+ feats: List[torch.Tensor],
+ data_sample: DetDataSample,
+ rescale=True,
+ **kwargs) -> InstanceData:
+ """Tracking forward function.
+
+ Args:
+ model (nn.Module): VIS model.
+ img (Tensor): of shape (T, C, H, W) encoding input image.
+ Typically these should be mean centered and std scaled.
+ The T denotes the number of key images and usually is 1 in
+ MaskTrackRCNN method.
+ feats (list[Tensor]): Multi level feature maps of `img`.
+ data_sample (:obj:`TrackDataSample`): The data sample.
+ It includes information such as `pred_det_instances`.
+ rescale (bool, optional): If True, the bounding boxes should be
+ rescaled to fit the original scale of the image. Defaults to
+ True.
+
+ Returns:
+ :obj:`InstanceData`: Tracking results of the input images.
+ Each InstanceData usually contains ``bboxes``, ``labels``,
+ ``scores`` and ``instances_id``.
+ """
+ metainfo = data_sample.metainfo
+ bboxes = data_sample.pred_instances.bboxes
+ masks = data_sample.pred_instances.masks
+ labels = data_sample.pred_instances.labels
+ scores = data_sample.pred_instances.scores
+
+ frame_id = metainfo.get('frame_id', -1)
+ # create pred_track_instances
+ pred_track_instances = InstanceData()
+
+ if bboxes.shape[0] == 0:
+ ids = torch.zeros_like(labels)
+ pred_track_instances = data_sample.pred_instances.clone()
+ pred_track_instances.instances_id = ids
+ return pred_track_instances
+
+ rescaled_bboxes = bboxes.clone()
+ if rescale:
+ scale_factor = rescaled_bboxes.new_tensor(
+ metainfo['scale_factor']).repeat((1, 2))
+ rescaled_bboxes = rescaled_bboxes * scale_factor
+ roi_feats, _ = model.track_head.extract_roi_feats(
+ feats, [rescaled_bboxes])
+
+ if self.empty:
+ num_new_tracks = bboxes.size(0)
+ ids = torch.arange(
+ self.num_tracks,
+ self.num_tracks + num_new_tracks,
+ dtype=torch.long)
+ self.num_tracks += num_new_tracks
+ else:
+ prev_bboxes = self.get('bboxes')
+ prev_labels = self.get('labels')
+ prev_roi_feats = self.get('roi_feats')
+
+ similarity_logits = model.track_head.predict(
+ roi_feats, prev_roi_feats)
+ match_scores = self.get_match_score(bboxes, labels, scores,
+ prev_bboxes, prev_labels,
+ similarity_logits)
+ ids, _ = self.assign_ids(match_scores)
+
+ valid_inds = ids > -1
+ ids = ids[valid_inds]
+ bboxes = bboxes[valid_inds]
+ labels = labels[valid_inds]
+ scores = scores[valid_inds]
+ masks = masks[valid_inds]
+ roi_feats = roi_feats[valid_inds]
+
+ self.update(
+ ids=ids,
+ bboxes=bboxes,
+ labels=labels,
+ scores=scores,
+ masks=masks,
+ roi_feats=roi_feats,
+ frame_ids=frame_id)
+ # update pred_track_instances
+ pred_track_instances.bboxes = bboxes
+ pred_track_instances.masks = masks
+ pred_track_instances.labels = labels
+ pred_track_instances.scores = scores
+ pred_track_instances.instances_id = ids
+
+ return pred_track_instances
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/ocsort_tracker.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/ocsort_tracker.py
new file mode 100644
index 0000000000000000000000000000000000000000..4e09990c603aee8ced3bf3a65ceb530142e6e873
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/ocsort_tracker.py
@@ -0,0 +1,531 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple
+
+try:
+ import lap
+except ImportError:
+ lap = None
+import numpy as np
+import torch
+from addict import Dict
+from mmengine.structures import InstanceData
+
+from mmdet.registry import MODELS
+from mmdet.structures import DetDataSample
+from mmdet.structures.bbox import (bbox_cxcyah_to_xyxy, bbox_overlaps,
+ bbox_xyxy_to_cxcyah)
+from .sort_tracker import SORTTracker
+
+
+@MODELS.register_module()
+class OCSORTTracker(SORTTracker):
+ """Tracker for OC-SORT.
+
+ Args:
+ motion (dict): Configuration of motion. Defaults to None.
+ obj_score_thrs (float): Detection score threshold for matching objects.
+ Defaults to 0.3.
+ init_track_thr (float): Detection score threshold for initializing a
+ new tracklet. Defaults to 0.7.
+ weight_iou_with_det_scores (bool): Whether using detection scores to
+ weight IOU which is used for matching. Defaults to True.
+ match_iou_thr (float): IOU distance threshold for matching between two
+ frames. Defaults to 0.3.
+ num_tentatives (int, optional): Number of continuous frames to confirm
+ a track. Defaults to 3.
+ vel_consist_weight (float): Weight of the velocity consistency term in
+ association (OCM term in the paper).
+ vel_delta_t (int): The difference of time step for calculating of the
+ velocity direction of tracklets.
+ init_cfg (dict or list[dict], optional): Initialization config dict.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ motion: Optional[dict] = None,
+ obj_score_thr: float = 0.3,
+ init_track_thr: float = 0.7,
+ weight_iou_with_det_scores: bool = True,
+ match_iou_thr: float = 0.3,
+ num_tentatives: int = 3,
+ vel_consist_weight: float = 0.2,
+ vel_delta_t: int = 3,
+ **kwargs):
+ if lap is None:
+ raise RuntimeError('lap is not installed,\
+ please install it by: pip install lap')
+ super().__init__(motion=motion, **kwargs)
+ self.obj_score_thr = obj_score_thr
+ self.init_track_thr = init_track_thr
+
+ self.weight_iou_with_det_scores = weight_iou_with_det_scores
+ self.match_iou_thr = match_iou_thr
+ self.vel_consist_weight = vel_consist_weight
+ self.vel_delta_t = vel_delta_t
+
+ self.num_tentatives = num_tentatives
+
+ @property
+ def unconfirmed_ids(self):
+ """Unconfirmed ids in the tracker."""
+ ids = [id for id, track in self.tracks.items() if track.tentative]
+ return ids
+
+ def init_track(self, id: int, obj: Tuple[torch.Tensor]):
+ """Initialize a track."""
+ super().init_track(id, obj)
+ if self.tracks[id].frame_ids[-1] == 0:
+ self.tracks[id].tentative = False
+ else:
+ self.tracks[id].tentative = True
+ bbox = bbox_xyxy_to_cxcyah(self.tracks[id].bboxes[-1]) # size = (1, 4)
+ assert bbox.ndim == 2 and bbox.shape[0] == 1
+ bbox = bbox.squeeze(0).cpu().numpy()
+ self.tracks[id].mean, self.tracks[id].covariance = self.kf.initiate(
+ bbox)
+ # track.obs maintains the history associated detections to this track
+ self.tracks[id].obs = []
+ bbox_id = self.memo_items.index('bboxes')
+ self.tracks[id].obs.append(obj[bbox_id])
+ # a placefolder to save mean/covariance before losing tracking it
+ # parameters to save: mean, covariance, measurement
+ self.tracks[id].tracked = True
+ self.tracks[id].saved_attr = Dict()
+ self.tracks[id].velocity = torch.tensor(
+ (-1, -1)).to(obj[bbox_id].device) # placeholder
+
+ def update_track(self, id: int, obj: Tuple[torch.Tensor]):
+ """Update a track."""
+ super().update_track(id, obj)
+ if self.tracks[id].tentative:
+ if len(self.tracks[id]['bboxes']) >= self.num_tentatives:
+ self.tracks[id].tentative = False
+ bbox = bbox_xyxy_to_cxcyah(self.tracks[id].bboxes[-1]) # size = (1, 4)
+ assert bbox.ndim == 2 and bbox.shape[0] == 1
+ bbox = bbox.squeeze(0).cpu().numpy()
+ self.tracks[id].mean, self.tracks[id].covariance = self.kf.update(
+ self.tracks[id].mean, self.tracks[id].covariance, bbox)
+ self.tracks[id].tracked = True
+ bbox_id = self.memo_items.index('bboxes')
+ self.tracks[id].obs.append(obj[bbox_id])
+
+ bbox1 = self.k_step_observation(self.tracks[id])
+ bbox2 = obj[bbox_id]
+ self.tracks[id].velocity = self.vel_direction(bbox1, bbox2).to(
+ obj[bbox_id].device)
+
+ def vel_direction(self, bbox1: torch.Tensor, bbox2: torch.Tensor):
+ """Estimate the direction vector between two boxes."""
+ if bbox1.sum() < 0 or bbox2.sum() < 0:
+ return torch.tensor((-1, -1))
+ cx1, cy1 = (bbox1[0] + bbox1[2]) / 2.0, (bbox1[1] + bbox1[3]) / 2.0
+ cx2, cy2 = (bbox2[0] + bbox2[2]) / 2.0, (bbox2[1] + bbox2[3]) / 2.0
+ speed = torch.tensor([cy2 - cy1, cx2 - cx1])
+ norm = torch.sqrt((speed[0])**2 + (speed[1])**2) + 1e-6
+ return speed / norm
+
+ def vel_direction_batch(self, bboxes1: torch.Tensor,
+ bboxes2: torch.Tensor):
+ """Estimate the direction vector given two batches of boxes."""
+ cx1, cy1 = (bboxes1[:, 0] + bboxes1[:, 2]) / 2.0, (bboxes1[:, 1] +
+ bboxes1[:, 3]) / 2.0
+ cx2, cy2 = (bboxes2[:, 0] + bboxes2[:, 2]) / 2.0, (bboxes2[:, 1] +
+ bboxes2[:, 3]) / 2.0
+ speed_diff_y = cy2[None, :] - cy1[:, None]
+ speed_diff_x = cx2[None, :] - cx1[:, None]
+ speed = torch.cat((speed_diff_y[..., None], speed_diff_x[..., None]),
+ dim=-1)
+ norm = torch.sqrt((speed[:, :, 0])**2 + (speed[:, :, 1])**2) + 1e-6
+ speed[:, :, 0] /= norm
+ speed[:, :, 1] /= norm
+ return speed
+
+ def k_step_observation(self, track: Dict):
+ """return the observation k step away before."""
+ obs_seqs = track.obs
+ num_obs = len(obs_seqs)
+ if num_obs == 0:
+ return torch.tensor((-1, -1, -1, -1)).to(track.obs[0].device)
+ elif num_obs > self.vel_delta_t:
+ if obs_seqs[num_obs - 1 - self.vel_delta_t] is not None:
+ return obs_seqs[num_obs - 1 - self.vel_delta_t]
+ else:
+ return self.last_obs(track)
+ else:
+ return self.last_obs(track)
+
+ def ocm_assign_ids(self,
+ ids: List[int],
+ det_bboxes: torch.Tensor,
+ det_labels: torch.Tensor,
+ det_scores: torch.Tensor,
+ weight_iou_with_det_scores: Optional[bool] = False,
+ match_iou_thr: Optional[float] = 0.5):
+ """Apply Observation-Centric Momentum (OCM) to assign ids.
+
+ OCM adds movement direction consistency into the association cost
+ matrix. This term requires no additional assumption but from the
+ same linear motion assumption as the canonical Kalman Filter in SORT.
+
+ Args:
+ ids (list[int]): Tracking ids.
+ det_bboxes (Tensor): of shape (N, 4)
+ det_labels (Tensor): of shape (N,)
+ det_scores (Tensor): of shape (N,)
+ weight_iou_with_det_scores (bool, optional): Whether using
+ detection scores to weight IOU which is used for matching.
+ Defaults to False.
+ match_iou_thr (float, optional): Matching threshold.
+ Defaults to 0.5.
+
+ Returns:
+ tuple(int): The assigning ids.
+
+ OC-SORT uses velocity consistency besides IoU for association
+ """
+ # get track_bboxes
+ track_bboxes = np.zeros((0, 4))
+ for id in ids:
+ track_bboxes = np.concatenate(
+ (track_bboxes, self.tracks[id].mean[:4][None]), axis=0)
+ track_bboxes = torch.from_numpy(track_bboxes).to(det_bboxes)
+ track_bboxes = bbox_cxcyah_to_xyxy(track_bboxes)
+
+ # compute distance
+ ious = bbox_overlaps(track_bboxes, det_bboxes)
+ if weight_iou_with_det_scores:
+ ious *= det_scores
+
+ # support multi-class association
+ track_labels = torch.tensor([
+ self.tracks[id]['labels'][-1] for id in ids
+ ]).to(det_bboxes.device)
+ cate_match = det_labels[None, :] == track_labels[:, None]
+ # to avoid det and track of different categories are matched
+ cate_cost = (1 - cate_match.int()) * 1e6
+
+ dists = (1 - ious + cate_cost).cpu().numpy()
+
+ if len(ids) > 0 and len(det_bboxes) > 0:
+ track_velocities = torch.stack(
+ [self.tracks[id].velocity for id in ids]).to(det_bboxes.device)
+ k_step_observations = torch.stack([
+ self.k_step_observation(self.tracks[id]) for id in ids
+ ]).to(det_bboxes.device)
+ # valid1: if the track has previous observations to estimate speed
+ # valid2: if the associated observation k steps ago is a detection
+ valid1 = track_velocities.sum(dim=1) != -2
+ valid2 = k_step_observations.sum(dim=1) != -4
+ valid = valid1 & valid2
+
+ vel_to_match = self.vel_direction_batch(k_step_observations,
+ det_bboxes)
+ track_velocities = track_velocities[:, None, :].repeat(
+ 1, det_bboxes.shape[0], 1)
+
+ angle_cos = (vel_to_match * track_velocities).sum(dim=-1)
+ angle_cos = torch.clamp(angle_cos, min=-1, max=1)
+ angle = torch.acos(angle_cos) # [0, pi]
+ norm_angle = (angle - np.pi / 2.) / np.pi # [-0.5, 0.5]
+ valid_matrix = valid[:, None].int().repeat(1, det_bboxes.shape[0])
+ # set non-valid entries 0
+ valid_norm_angle = norm_angle * valid_matrix
+
+ dists += valid_norm_angle.cpu().numpy() * self.vel_consist_weight
+
+ # bipartite match
+ if dists.size > 0:
+ cost, row, col = lap.lapjv(
+ dists, extend_cost=True, cost_limit=1 - match_iou_thr)
+ else:
+ row = np.zeros(len(ids)).astype(np.int32) - 1
+ col = np.zeros(len(det_bboxes)).astype(np.int32) - 1
+ return row, col
+
+ def last_obs(self, track: Dict):
+ """extract the last associated observation."""
+ for bbox in track.obs[::-1]:
+ if bbox is not None:
+ return bbox
+
+ def ocr_assign_ids(self,
+ track_obs: torch.Tensor,
+ last_track_labels: torch.Tensor,
+ det_bboxes: torch.Tensor,
+ det_labels: torch.Tensor,
+ det_scores: torch.Tensor,
+ weight_iou_with_det_scores: Optional[bool] = False,
+ match_iou_thr: Optional[float] = 0.5):
+ """association for Observation-Centric Recovery.
+
+ As try to recover tracks from being lost whose estimated velocity is
+ out- to-date, we use IoU-only matching strategy.
+
+ Args:
+ track_obs (Tensor): the list of historical associated
+ detections of tracks
+ det_bboxes (Tensor): of shape (N, 5), unmatched detections
+ det_labels (Tensor): of shape (N,)
+ det_scores (Tensor): of shape (N,)
+ weight_iou_with_det_scores (bool, optional): Whether using
+ detection scores to weight IOU which is used for matching.
+ Defaults to False.
+ match_iou_thr (float, optional): Matching threshold.
+ Defaults to 0.5.
+
+ Returns:
+ tuple(int): The assigning ids.
+ """
+ # compute distance
+ ious = bbox_overlaps(track_obs, det_bboxes)
+ if weight_iou_with_det_scores:
+ ious *= det_scores
+
+ # support multi-class association
+ cate_match = det_labels[None, :] == last_track_labels[:, None]
+ # to avoid det and track of different categories are matched
+ cate_cost = (1 - cate_match.int()) * 1e6
+
+ dists = (1 - ious + cate_cost).cpu().numpy()
+
+ # bipartite match
+ if dists.size > 0:
+ cost, row, col = lap.lapjv(
+ dists, extend_cost=True, cost_limit=1 - match_iou_thr)
+ else:
+ row = np.zeros(len(track_obs)).astype(np.int32) - 1
+ col = np.zeros(len(det_bboxes)).astype(np.int32) - 1
+ return row, col
+
+ def online_smooth(self, track: Dict, obj: torch.Tensor):
+ """Once a track is recovered from being lost, online smooth its
+ parameters to fix the error accumulated during being lost.
+
+ NOTE: you can use different virtual trajectory generation
+ strategies, we adopt the naive linear interpolation as default
+ """
+ last_match_bbox = self.last_obs(track)
+ new_match_bbox = obj
+ unmatch_len = 0
+ for bbox in track.obs[::-1]:
+ if bbox is None:
+ unmatch_len += 1
+ else:
+ break
+ bbox_shift_per_step = (new_match_bbox - last_match_bbox) / (
+ unmatch_len + 1)
+ track.mean = track.saved_attr.mean
+ track.covariance = track.saved_attr.covariance
+ for i in range(unmatch_len):
+ virtual_bbox = last_match_bbox + (i + 1) * bbox_shift_per_step
+ virtual_bbox = bbox_xyxy_to_cxcyah(virtual_bbox[None, :])
+ virtual_bbox = virtual_bbox.squeeze(0).cpu().numpy()
+ track.mean, track.covariance = self.kf.update(
+ track.mean, track.covariance, virtual_bbox)
+
+ def track(self, data_sample: DetDataSample, **kwargs) -> InstanceData:
+ """Tracking forward function.
+ NOTE: this implementation is slightly different from the original
+ OC-SORT implementation (https://github.com/noahcao/OC_SORT)that we
+ do association between detections and tentative/non-tentative tracks
+ independently while the original implementation combines them together.
+
+ Args:
+ data_sample (:obj:`DetDataSample`): The data sample.
+ It includes information such as `pred_instances`.
+
+ Returns:
+ :obj:`InstanceData`: Tracking results of the input images.
+ Each InstanceData usually contains ``bboxes``, ``labels``,
+ ``scores`` and ``instances_id``.
+ """
+ metainfo = data_sample.metainfo
+ bboxes = data_sample.pred_instances.bboxes
+ labels = data_sample.pred_instances.labels
+ scores = data_sample.pred_instances.scores
+ frame_id = metainfo.get('frame_id', -1)
+ if frame_id == 0:
+ self.reset()
+ if not hasattr(self, 'kf'):
+ self.kf = self.motion
+
+ if self.empty or bboxes.size(0) == 0:
+ valid_inds = scores > self.init_track_thr
+ scores = scores[valid_inds]
+ bboxes = bboxes[valid_inds]
+ labels = labels[valid_inds]
+ num_new_tracks = bboxes.size(0)
+ ids = torch.arange(self.num_tracks,
+ self.num_tracks + num_new_tracks).to(labels)
+ self.num_tracks += num_new_tracks
+ else:
+ # 0. init
+ ids = torch.full((bboxes.size(0), ),
+ -1,
+ dtype=labels.dtype,
+ device=labels.device)
+
+ # get the detection bboxes for the first association
+ det_inds = scores > self.obj_score_thr
+ det_bboxes = bboxes[det_inds]
+ det_labels = labels[det_inds]
+ det_scores = scores[det_inds]
+ det_ids = ids[det_inds]
+
+ # 1. predict by Kalman Filter
+ for id in self.confirmed_ids:
+ # track is lost in previous frame
+ if self.tracks[id].frame_ids[-1] != frame_id - 1:
+ self.tracks[id].mean[7] = 0
+ if self.tracks[id].tracked:
+ self.tracks[id].saved_attr.mean = self.tracks[id].mean
+ self.tracks[id].saved_attr.covariance = self.tracks[
+ id].covariance
+ (self.tracks[id].mean,
+ self.tracks[id].covariance) = self.kf.predict(
+ self.tracks[id].mean, self.tracks[id].covariance)
+
+ # 2. match detections and tracks' predicted locations
+ match_track_inds, raw_match_det_inds = self.ocm_assign_ids(
+ self.confirmed_ids, det_bboxes, det_labels, det_scores,
+ self.weight_iou_with_det_scores, self.match_iou_thr)
+ # '-1' mean a detection box is not matched with tracklets in
+ # previous frame
+ valid = raw_match_det_inds > -1
+ det_ids[valid] = torch.tensor(
+ self.confirmed_ids)[raw_match_det_inds[valid]].to(labels)
+
+ match_det_bboxes = det_bboxes[valid]
+ match_det_labels = det_labels[valid]
+ match_det_scores = det_scores[valid]
+ match_det_ids = det_ids[valid]
+ assert (match_det_ids > -1).all()
+
+ # unmatched tracks and detections
+ unmatch_det_bboxes = det_bboxes[~valid]
+ unmatch_det_labels = det_labels[~valid]
+ unmatch_det_scores = det_scores[~valid]
+ unmatch_det_ids = det_ids[~valid]
+ assert (unmatch_det_ids == -1).all()
+
+ # 3. use unmatched detection bboxes from the first match to match
+ # the unconfirmed tracks
+ (tentative_match_track_inds,
+ tentative_match_det_inds) = self.ocm_assign_ids(
+ self.unconfirmed_ids, unmatch_det_bboxes, unmatch_det_labels,
+ unmatch_det_scores, self.weight_iou_with_det_scores,
+ self.match_iou_thr)
+ valid = tentative_match_det_inds > -1
+ unmatch_det_ids[valid] = torch.tensor(self.unconfirmed_ids)[
+ tentative_match_det_inds[valid]].to(labels)
+
+ match_det_bboxes = torch.cat(
+ (match_det_bboxes, unmatch_det_bboxes[valid]), dim=0)
+ match_det_labels = torch.cat(
+ (match_det_labels, unmatch_det_labels[valid]), dim=0)
+ match_det_scores = torch.cat(
+ (match_det_scores, unmatch_det_scores[valid]), dim=0)
+ match_det_ids = torch.cat((match_det_ids, unmatch_det_ids[valid]),
+ dim=0)
+ assert (match_det_ids > -1).all()
+
+ unmatch_det_bboxes = unmatch_det_bboxes[~valid]
+ unmatch_det_labels = unmatch_det_labels[~valid]
+ unmatch_det_scores = unmatch_det_scores[~valid]
+ unmatch_det_ids = unmatch_det_ids[~valid]
+ assert (unmatch_det_ids == -1).all()
+
+ all_track_ids = [id for id, _ in self.tracks.items()]
+ unmatched_track_inds = torch.tensor(
+ [ind for ind in all_track_ids if ind not in match_det_ids])
+
+ if len(unmatched_track_inds) > 0:
+ # 4. still some tracks not associated yet, perform OCR
+ last_observations = []
+ for id in unmatched_track_inds:
+ last_box = self.last_obs(self.tracks[id.item()])
+ last_observations.append(last_box)
+ last_observations = torch.stack(last_observations)
+ last_track_labels = torch.tensor([
+ self.tracks[id.item()]['labels'][-1]
+ for id in unmatched_track_inds
+ ]).to(det_bboxes.device)
+
+ remain_det_ids = torch.full((unmatch_det_bboxes.size(0), ),
+ -1,
+ dtype=labels.dtype,
+ device=labels.device)
+
+ _, ocr_match_det_inds = self.ocr_assign_ids(
+ last_observations, last_track_labels, unmatch_det_bboxes,
+ unmatch_det_labels, unmatch_det_scores,
+ self.weight_iou_with_det_scores, self.match_iou_thr)
+
+ valid = ocr_match_det_inds > -1
+ remain_det_ids[valid] = unmatched_track_inds.clone()[
+ ocr_match_det_inds[valid]].to(labels)
+
+ ocr_match_det_bboxes = unmatch_det_bboxes[valid]
+ ocr_match_det_labels = unmatch_det_labels[valid]
+ ocr_match_det_scores = unmatch_det_scores[valid]
+ ocr_match_det_ids = remain_det_ids[valid]
+ assert (ocr_match_det_ids > -1).all()
+
+ ocr_unmatch_det_bboxes = unmatch_det_bboxes[~valid]
+ ocr_unmatch_det_labels = unmatch_det_labels[~valid]
+ ocr_unmatch_det_scores = unmatch_det_scores[~valid]
+ ocr_unmatch_det_ids = remain_det_ids[~valid]
+ assert (ocr_unmatch_det_ids == -1).all()
+
+ unmatch_det_bboxes = ocr_unmatch_det_bboxes
+ unmatch_det_labels = ocr_unmatch_det_labels
+ unmatch_det_scores = ocr_unmatch_det_scores
+ unmatch_det_ids = ocr_unmatch_det_ids
+ match_det_bboxes = torch.cat(
+ (match_det_bboxes, ocr_match_det_bboxes), dim=0)
+ match_det_labels = torch.cat(
+ (match_det_labels, ocr_match_det_labels), dim=0)
+ match_det_scores = torch.cat(
+ (match_det_scores, ocr_match_det_scores), dim=0)
+ match_det_ids = torch.cat((match_det_ids, ocr_match_det_ids),
+ dim=0)
+
+ # 5. summarize the track results
+ for i in range(len(match_det_ids)):
+ det_bbox = match_det_bboxes[i]
+ track_id = match_det_ids[i].item()
+ if not self.tracks[track_id].tracked:
+ # the track is lost before this step
+ self.online_smooth(self.tracks[track_id], det_bbox)
+
+ for track_id in all_track_ids:
+ if track_id not in match_det_ids:
+ self.tracks[track_id].tracked = False
+ self.tracks[track_id].obs.append(None)
+
+ bboxes = torch.cat((match_det_bboxes, unmatch_det_bboxes), dim=0)
+ labels = torch.cat((match_det_labels, unmatch_det_labels), dim=0)
+ scores = torch.cat((match_det_scores, unmatch_det_scores), dim=0)
+ ids = torch.cat((match_det_ids, unmatch_det_ids), dim=0)
+ # 6. assign new ids
+ new_track_inds = ids == -1
+
+ ids[new_track_inds] = torch.arange(
+ self.num_tracks,
+ self.num_tracks + new_track_inds.sum()).to(labels)
+ self.num_tracks += new_track_inds.sum()
+
+ self.update(
+ ids=ids,
+ bboxes=bboxes,
+ labels=labels,
+ scores=scores,
+ frame_ids=frame_id)
+
+ # update pred_track_instances
+ pred_track_instances = InstanceData()
+ pred_track_instances.bboxes = bboxes
+ pred_track_instances.labels = labels
+ pred_track_instances.scores = scores
+ pred_track_instances.instances_id = ids
+ return pred_track_instances
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/quasi_dense_tracker.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/quasi_dense_tracker.py
new file mode 100644
index 0000000000000000000000000000000000000000..c93c3c4c3bd5c8939e77195f30a7eb2f0314e225
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/quasi_dense_tracker.py
@@ -0,0 +1,316 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple
+
+import torch
+import torch.nn.functional as F
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.registry import MODELS
+from mmdet.structures import TrackDataSample
+from mmdet.structures.bbox import bbox_overlaps
+from .base_tracker import BaseTracker
+
+
+@MODELS.register_module()
+class QuasiDenseTracker(BaseTracker):
+ """Tracker for Quasi-Dense Tracking.
+
+ Args:
+ init_score_thr (float): The cls_score threshold to
+ initialize a new tracklet. Defaults to 0.8.
+ obj_score_thr (float): The cls_score threshold to
+ update a tracked tracklet. Defaults to 0.5.
+ match_score_thr (float): The match threshold. Defaults to 0.5.
+ memo_tracklet_frames (int): The most frames in a tracklet memory.
+ Defaults to 10.
+ memo_backdrop_frames (int): The most frames in the backdrops.
+ Defaults to 1.
+ memo_momentum (float): The momentum value for embeds updating.
+ Defaults to 0.8.
+ nms_conf_thr (float): The nms threshold for confidence.
+ Defaults to 0.5.
+ nms_backdrop_iou_thr (float): The nms threshold for backdrop IoU.
+ Defaults to 0.3.
+ nms_class_iou_thr (float): The nms threshold for class IoU.
+ Defaults to 0.7.
+ with_cats (bool): Whether to track with the same category.
+ Defaults to True.
+ match_metric (str): The match metric. Defaults to 'bisoftmax'.
+ """
+
+ def __init__(self,
+ init_score_thr: float = 0.8,
+ obj_score_thr: float = 0.5,
+ match_score_thr: float = 0.5,
+ memo_tracklet_frames: int = 10,
+ memo_backdrop_frames: int = 1,
+ memo_momentum: float = 0.8,
+ nms_conf_thr: float = 0.5,
+ nms_backdrop_iou_thr: float = 0.3,
+ nms_class_iou_thr: float = 0.7,
+ with_cats: bool = True,
+ match_metric: str = 'bisoftmax',
+ **kwargs):
+ super().__init__(**kwargs)
+ assert 0 <= memo_momentum <= 1.0
+ assert memo_tracklet_frames >= 0
+ assert memo_backdrop_frames >= 0
+ self.init_score_thr = init_score_thr
+ self.obj_score_thr = obj_score_thr
+ self.match_score_thr = match_score_thr
+ self.memo_tracklet_frames = memo_tracklet_frames
+ self.memo_backdrop_frames = memo_backdrop_frames
+ self.memo_momentum = memo_momentum
+ self.nms_conf_thr = nms_conf_thr
+ self.nms_backdrop_iou_thr = nms_backdrop_iou_thr
+ self.nms_class_iou_thr = nms_class_iou_thr
+ self.with_cats = with_cats
+ assert match_metric in ['bisoftmax', 'softmax', 'cosine']
+ self.match_metric = match_metric
+
+ self.num_tracks = 0
+ self.tracks = dict()
+ self.backdrops = []
+
+ def reset(self):
+ """Reset the buffer of the tracker."""
+ self.num_tracks = 0
+ self.tracks = dict()
+ self.backdrops = []
+
+ def update(self, ids: Tensor, bboxes: Tensor, embeds: Tensor,
+ labels: Tensor, scores: Tensor, frame_id: int) -> None:
+ """Tracking forward function.
+
+ Args:
+ ids (Tensor): of shape(N, ).
+ bboxes (Tensor): of shape (N, 5).
+ embeds (Tensor): of shape (N, 256).
+ labels (Tensor): of shape (N, ).
+ scores (Tensor): of shape (N, ).
+ frame_id (int): The id of current frame, 0-index.
+ """
+ tracklet_inds = ids > -1
+
+ for id, bbox, embed, label, score in zip(ids[tracklet_inds],
+ bboxes[tracklet_inds],
+ embeds[tracklet_inds],
+ labels[tracklet_inds],
+ scores[tracklet_inds]):
+ id = int(id)
+ # update the tracked ones and initialize new tracks
+ if id in self.tracks.keys():
+ velocity = (bbox - self.tracks[id]['bbox']) / (
+ frame_id - self.tracks[id]['last_frame'])
+ self.tracks[id]['bbox'] = bbox
+ self.tracks[id]['embed'] = (
+ 1 - self.memo_momentum
+ ) * self.tracks[id]['embed'] + self.memo_momentum * embed
+ self.tracks[id]['last_frame'] = frame_id
+ self.tracks[id]['label'] = label
+ self.tracks[id]['score'] = score
+ self.tracks[id]['velocity'] = (
+ self.tracks[id]['velocity'] * self.tracks[id]['acc_frame']
+ + velocity) / (
+ self.tracks[id]['acc_frame'] + 1)
+ self.tracks[id]['acc_frame'] += 1
+ else:
+ self.tracks[id] = dict(
+ bbox=bbox,
+ embed=embed,
+ label=label,
+ score=score,
+ last_frame=frame_id,
+ velocity=torch.zeros_like(bbox),
+ acc_frame=0)
+ # backdrop update according to IoU
+ backdrop_inds = torch.nonzero(ids == -1, as_tuple=False).squeeze(1)
+ ious = bbox_overlaps(bboxes[backdrop_inds], bboxes)
+ for i, ind in enumerate(backdrop_inds):
+ if (ious[i, :ind] > self.nms_backdrop_iou_thr).any():
+ backdrop_inds[i] = -1
+ backdrop_inds = backdrop_inds[backdrop_inds > -1]
+ # old backdrops would be removed at first
+ self.backdrops.insert(
+ 0,
+ dict(
+ bboxes=bboxes[backdrop_inds],
+ embeds=embeds[backdrop_inds],
+ labels=labels[backdrop_inds]))
+
+ # pop memo
+ invalid_ids = []
+ for k, v in self.tracks.items():
+ if frame_id - v['last_frame'] >= self.memo_tracklet_frames:
+ invalid_ids.append(k)
+ for invalid_id in invalid_ids:
+ self.tracks.pop(invalid_id)
+
+ if len(self.backdrops) > self.memo_backdrop_frames:
+ self.backdrops.pop()
+
+ @property
+ def memo(self) -> Tuple[Tensor, ...]:
+ """Get tracks memory."""
+ memo_embeds = []
+ memo_ids = []
+ memo_bboxes = []
+ memo_labels = []
+ # velocity of tracks
+ memo_vs = []
+ # get tracks
+ for k, v in self.tracks.items():
+ memo_bboxes.append(v['bbox'][None, :])
+ memo_embeds.append(v['embed'][None, :])
+ memo_ids.append(k)
+ memo_labels.append(v['label'].view(1, 1))
+ memo_vs.append(v['velocity'][None, :])
+ memo_ids = torch.tensor(memo_ids, dtype=torch.long).view(1, -1)
+ # get backdrops
+ for backdrop in self.backdrops:
+ backdrop_ids = torch.full((1, backdrop['embeds'].size(0)),
+ -1,
+ dtype=torch.long)
+ backdrop_vs = torch.zeros_like(backdrop['bboxes'])
+ memo_bboxes.append(backdrop['bboxes'])
+ memo_embeds.append(backdrop['embeds'])
+ memo_ids = torch.cat([memo_ids, backdrop_ids], dim=1)
+ memo_labels.append(backdrop['labels'][:, None])
+ memo_vs.append(backdrop_vs)
+
+ memo_bboxes = torch.cat(memo_bboxes, dim=0)
+ memo_embeds = torch.cat(memo_embeds, dim=0)
+ memo_labels = torch.cat(memo_labels, dim=0).squeeze(1)
+ memo_vs = torch.cat(memo_vs, dim=0)
+ return memo_bboxes, memo_labels, memo_embeds, memo_ids.squeeze(
+ 0), memo_vs
+
+ def track(self,
+ model: torch.nn.Module,
+ img: torch.Tensor,
+ feats: List[torch.Tensor],
+ data_sample: TrackDataSample,
+ rescale=True,
+ **kwargs) -> InstanceData:
+ """Tracking forward function.
+
+ Args:
+ model (nn.Module): MOT model.
+ img (Tensor): of shape (T, C, H, W) encoding input image.
+ Typically these should be mean centered and std scaled.
+ The T denotes the number of key images and usually is 1 in
+ QDTrack method.
+ feats (list[Tensor]): Multi level feature maps of `img`.
+ data_sample (:obj:`TrackDataSample`): The data sample.
+ It includes information such as `pred_instances`.
+ rescale (bool, optional): If True, the bounding boxes should be
+ rescaled to fit the original scale of the image. Defaults to
+ True.
+
+ Returns:
+ :obj:`InstanceData`: Tracking results of the input images.
+ Each InstanceData usually contains ``bboxes``, ``labels``,
+ ``scores`` and ``instances_id``.
+ """
+ metainfo = data_sample.metainfo
+ bboxes = data_sample.pred_instances.bboxes
+ labels = data_sample.pred_instances.labels
+ scores = data_sample.pred_instances.scores
+
+ frame_id = metainfo.get('frame_id', -1)
+ # create pred_track_instances
+ pred_track_instances = InstanceData()
+
+ # return zero bboxes if there is no track targets
+ if bboxes.shape[0] == 0:
+ ids = torch.zeros_like(labels)
+ pred_track_instances = data_sample.pred_instances.clone()
+ pred_track_instances.instances_id = ids
+ return pred_track_instances
+
+ # get track feats
+ rescaled_bboxes = bboxes.clone()
+ if rescale:
+ scale_factor = rescaled_bboxes.new_tensor(
+ metainfo['scale_factor']).repeat((1, 2))
+ rescaled_bboxes = rescaled_bboxes * scale_factor
+ track_feats = model.track_head.predict(feats, [rescaled_bboxes])
+ # sort according to the object_score
+ _, inds = scores.sort(descending=True)
+ bboxes = bboxes[inds]
+ scores = scores[inds]
+ labels = labels[inds]
+ embeds = track_feats[inds, :]
+
+ # duplicate removal for potential backdrops and cross classes
+ valids = bboxes.new_ones((bboxes.size(0)))
+ ious = bbox_overlaps(bboxes, bboxes)
+ for i in range(1, bboxes.size(0)):
+ thr = self.nms_backdrop_iou_thr if scores[
+ i] < self.obj_score_thr else self.nms_class_iou_thr
+ if (ious[i, :i] > thr).any():
+ valids[i] = 0
+ valids = valids == 1
+ bboxes = bboxes[valids]
+ scores = scores[valids]
+ labels = labels[valids]
+ embeds = embeds[valids, :]
+
+ # init ids container
+ ids = torch.full((bboxes.size(0), ), -1, dtype=torch.long)
+
+ # match if buffer is not empty
+ if bboxes.size(0) > 0 and not self.empty:
+ (memo_bboxes, memo_labels, memo_embeds, memo_ids,
+ memo_vs) = self.memo
+
+ if self.match_metric == 'bisoftmax':
+ feats = torch.mm(embeds, memo_embeds.t())
+ d2t_scores = feats.softmax(dim=1)
+ t2d_scores = feats.softmax(dim=0)
+ match_scores = (d2t_scores + t2d_scores) / 2
+ elif self.match_metric == 'softmax':
+ feats = torch.mm(embeds, memo_embeds.t())
+ match_scores = feats.softmax(dim=1)
+ elif self.match_metric == 'cosine':
+ match_scores = torch.mm(
+ F.normalize(embeds, p=2, dim=1),
+ F.normalize(memo_embeds, p=2, dim=1).t())
+ else:
+ raise NotImplementedError
+ # track with the same category
+ if self.with_cats:
+ cat_same = labels.view(-1, 1) == memo_labels.view(1, -1)
+ match_scores *= cat_same.float().to(match_scores.device)
+ # track according to match_scores
+ for i in range(bboxes.size(0)):
+ conf, memo_ind = torch.max(match_scores[i, :], dim=0)
+ id = memo_ids[memo_ind]
+ if conf > self.match_score_thr:
+ if id > -1:
+ # keep bboxes with high object score
+ # and remove background bboxes
+ if scores[i] > self.obj_score_thr:
+ ids[i] = id
+ match_scores[:i, memo_ind] = 0
+ match_scores[i + 1:, memo_ind] = 0
+ else:
+ if conf > self.nms_conf_thr:
+ ids[i] = -2
+ # initialize new tracks
+ new_inds = (ids == -1) & (scores > self.init_score_thr).cpu()
+ num_news = new_inds.sum()
+ ids[new_inds] = torch.arange(
+ self.num_tracks, self.num_tracks + num_news, dtype=torch.long)
+ self.num_tracks += num_news
+
+ self.update(ids, bboxes, embeds, labels, scores, frame_id)
+ tracklet_inds = ids > -1
+ # update pred_track_instances
+ pred_track_instances.bboxes = bboxes[tracklet_inds]
+ pred_track_instances.labels = labels[tracklet_inds]
+ pred_track_instances.scores = scores[tracklet_inds]
+ pred_track_instances.instances_id = ids[tracklet_inds]
+
+ return pred_track_instances
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/sort_tracker.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/sort_tracker.py
new file mode 100644
index 0000000000000000000000000000000000000000..c4a4fed92702f7d1ea66917a7157fcf5d0773a30
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/sort_tracker.py
@@ -0,0 +1,268 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple
+
+import numpy as np
+import torch
+from mmengine.structures import InstanceData
+
+try:
+ import motmetrics
+ from motmetrics.lap import linear_sum_assignment
+except ImportError:
+ motmetrics = None
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures import DetDataSample
+from mmdet.structures.bbox import bbox_overlaps, bbox_xyxy_to_cxcyah
+from mmdet.utils import OptConfigType
+from ..utils import imrenormalize
+from .base_tracker import BaseTracker
+
+
+@MODELS.register_module()
+class SORTTracker(BaseTracker):
+ """Tracker for SORT/DeepSORT.
+
+ Args:
+ obj_score_thr (float, optional): Threshold to filter the objects.
+ Defaults to 0.3.
+ motion (dict): Configuration of motion. Defaults to None.
+ reid (dict, optional): Configuration for the ReID model.
+ - num_samples (int, optional): Number of samples to calculate the
+ feature embeddings of a track. Default to 10.
+ - image_scale (tuple, optional): Input scale of the ReID model.
+ Default to (256, 128).
+ - img_norm_cfg (dict, optional): Configuration to normalize the
+ input. Default to None.
+ - match_score_thr (float, optional): Similarity threshold for the
+ matching process. Default to 2.0.
+ match_iou_thr (float, optional): Threshold of the IoU matching process.
+ Defaults to 0.7.
+ num_tentatives (int, optional): Number of continuous frames to confirm
+ a track. Defaults to 3.
+ """
+
+ def __init__(self,
+ motion: Optional[dict] = None,
+ obj_score_thr: float = 0.3,
+ reid: dict = dict(
+ num_samples=10,
+ img_scale=(256, 128),
+ img_norm_cfg=None,
+ match_score_thr=2.0),
+ match_iou_thr: float = 0.7,
+ num_tentatives: int = 3,
+ **kwargs):
+ if motmetrics is None:
+ raise RuntimeError('motmetrics is not installed,\
+ please install it by: pip install motmetrics')
+ super().__init__(**kwargs)
+ if motion is not None:
+ self.motion = TASK_UTILS.build(motion)
+ assert self.motion is not None, 'SORT/Deep SORT need KalmanFilter'
+ self.obj_score_thr = obj_score_thr
+ self.reid = reid
+ self.match_iou_thr = match_iou_thr
+ self.num_tentatives = num_tentatives
+
+ @property
+ def confirmed_ids(self) -> List:
+ """Confirmed ids in the tracker."""
+ ids = [id for id, track in self.tracks.items() if not track.tentative]
+ return ids
+
+ def init_track(self, id: int, obj: Tuple[Tensor]) -> None:
+ """Initialize a track."""
+ super().init_track(id, obj)
+ self.tracks[id].tentative = True
+ bbox = bbox_xyxy_to_cxcyah(self.tracks[id].bboxes[-1]) # size = (1, 4)
+ assert bbox.ndim == 2 and bbox.shape[0] == 1
+ bbox = bbox.squeeze(0).cpu().numpy()
+ self.tracks[id].mean, self.tracks[id].covariance = self.kf.initiate(
+ bbox)
+
+ def update_track(self, id: int, obj: Tuple[Tensor]) -> None:
+ """Update a track."""
+ super().update_track(id, obj)
+ if self.tracks[id].tentative:
+ if len(self.tracks[id]['bboxes']) >= self.num_tentatives:
+ self.tracks[id].tentative = False
+ bbox = bbox_xyxy_to_cxcyah(self.tracks[id].bboxes[-1]) # size = (1, 4)
+ assert bbox.ndim == 2 and bbox.shape[0] == 1
+ bbox = bbox.squeeze(0).cpu().numpy()
+ self.tracks[id].mean, self.tracks[id].covariance = self.kf.update(
+ self.tracks[id].mean, self.tracks[id].covariance, bbox)
+
+ def pop_invalid_tracks(self, frame_id: int) -> None:
+ """Pop out invalid tracks."""
+ invalid_ids = []
+ for k, v in self.tracks.items():
+ # case1: disappeared frames >= self.num_frames_retrain
+ case1 = frame_id - v['frame_ids'][-1] >= self.num_frames_retain
+ # case2: tentative tracks but not matched in this frame
+ case2 = v.tentative and v['frame_ids'][-1] != frame_id
+ if case1 or case2:
+ invalid_ids.append(k)
+ for invalid_id in invalid_ids:
+ self.tracks.pop(invalid_id)
+
+ def track(self,
+ model: torch.nn.Module,
+ img: Tensor,
+ data_sample: DetDataSample,
+ data_preprocessor: OptConfigType = None,
+ rescale: bool = False,
+ **kwargs) -> InstanceData:
+ """Tracking forward function.
+
+ Args:
+ model (nn.Module): MOT model.
+ img (Tensor): of shape (T, C, H, W) encoding input image.
+ Typically these should be mean centered and std scaled.
+ The T denotes the number of key images and usually is 1 in
+ SORT method.
+ data_sample (:obj:`TrackDataSample`): The data sample.
+ It includes information such as `pred_det_instances`.
+ data_preprocessor (dict or ConfigDict, optional): The pre-process
+ config of :class:`TrackDataPreprocessor`. it usually includes,
+ ``pad_size_divisor``, ``pad_value``, ``mean`` and ``std``.
+ rescale (bool, optional): If True, the bounding boxes should be
+ rescaled to fit the original scale of the image. Defaults to
+ False.
+
+ Returns:
+ :obj:`InstanceData`: Tracking results of the input images.
+ Each InstanceData usually contains ``bboxes``, ``labels``,
+ ``scores`` and ``instances_id``.
+ """
+ metainfo = data_sample.metainfo
+ bboxes = data_sample.pred_instances.bboxes
+ labels = data_sample.pred_instances.labels
+ scores = data_sample.pred_instances.scores
+
+ frame_id = metainfo.get('frame_id', -1)
+ if frame_id == 0:
+ self.reset()
+ if not hasattr(self, 'kf'):
+ self.kf = self.motion
+
+ if self.with_reid:
+ if self.reid.get('img_norm_cfg', False):
+ img_norm_cfg = dict(
+ mean=data_preprocessor['mean'],
+ std=data_preprocessor['std'],
+ to_bgr=data_preprocessor['rgb_to_bgr'])
+ reid_img = imrenormalize(img, img_norm_cfg,
+ self.reid['img_norm_cfg'])
+ else:
+ reid_img = img.clone()
+
+ valid_inds = scores > self.obj_score_thr
+ bboxes = bboxes[valid_inds]
+ labels = labels[valid_inds]
+ scores = scores[valid_inds]
+
+ if self.empty or bboxes.size(0) == 0:
+ num_new_tracks = bboxes.size(0)
+ ids = torch.arange(
+ self.num_tracks,
+ self.num_tracks + num_new_tracks,
+ dtype=torch.long).to(bboxes.device)
+ self.num_tracks += num_new_tracks
+ if self.with_reid:
+ crops = self.crop_imgs(reid_img, metainfo, bboxes.clone(),
+ rescale)
+ if crops.size(0) > 0:
+ embeds = model.reid(crops, mode='tensor')
+ else:
+ embeds = crops.new_zeros((0, model.reid.head.out_channels))
+ else:
+ ids = torch.full((bboxes.size(0), ), -1,
+ dtype=torch.long).to(bboxes.device)
+
+ # motion
+ self.tracks, costs = self.motion.track(self.tracks,
+ bbox_xyxy_to_cxcyah(bboxes))
+
+ active_ids = self.confirmed_ids
+ if self.with_reid:
+ crops = self.crop_imgs(reid_img, metainfo, bboxes.clone(),
+ rescale)
+ embeds = model.reid(crops, mode='tensor')
+
+ # reid
+ if len(active_ids) > 0:
+ track_embeds = self.get(
+ 'embeds',
+ active_ids,
+ self.reid.get('num_samples', None),
+ behavior='mean')
+ reid_dists = torch.cdist(track_embeds, embeds)
+
+ # support multi-class association
+ track_labels = torch.tensor([
+ self.tracks[id]['labels'][-1] for id in active_ids
+ ]).to(bboxes.device)
+ cate_match = labels[None, :] == track_labels[:, None]
+ cate_cost = (1 - cate_match.int()) * 1e6
+ reid_dists = (reid_dists + cate_cost).cpu().numpy()
+
+ valid_inds = [list(self.ids).index(_) for _ in active_ids]
+ reid_dists[~np.isfinite(costs[valid_inds, :])] = np.nan
+
+ row, col = linear_sum_assignment(reid_dists)
+ for r, c in zip(row, col):
+ dist = reid_dists[r, c]
+ if not np.isfinite(dist):
+ continue
+ if dist <= self.reid['match_score_thr']:
+ ids[c] = active_ids[r]
+
+ active_ids = [
+ id for id in self.ids if id not in ids
+ and self.tracks[id].frame_ids[-1] == frame_id - 1
+ ]
+ if len(active_ids) > 0:
+ active_dets = torch.nonzero(ids == -1).squeeze(1)
+ track_bboxes = self.get('bboxes', active_ids)
+ ious = bbox_overlaps(track_bboxes, bboxes[active_dets])
+
+ # support multi-class association
+ track_labels = torch.tensor([
+ self.tracks[id]['labels'][-1] for id in active_ids
+ ]).to(bboxes.device)
+ cate_match = labels[None, active_dets] == track_labels[:, None]
+ cate_cost = (1 - cate_match.int()) * 1e6
+
+ dists = (1 - ious + cate_cost).cpu().numpy()
+
+ row, col = linear_sum_assignment(dists)
+ for r, c in zip(row, col):
+ dist = dists[r, c]
+ if dist < 1 - self.match_iou_thr:
+ ids[active_dets[c]] = active_ids[r]
+
+ new_track_inds = ids == -1
+ ids[new_track_inds] = torch.arange(
+ self.num_tracks,
+ self.num_tracks + new_track_inds.sum(),
+ dtype=torch.long).to(bboxes.device)
+ self.num_tracks += new_track_inds.sum()
+
+ self.update(
+ ids=ids,
+ bboxes=bboxes,
+ scores=scores,
+ labels=labels,
+ embeds=embeds if self.with_reid else None,
+ frame_ids=frame_id)
+
+ # update pred_track_instances
+ pred_track_instances = InstanceData()
+ pred_track_instances.bboxes = bboxes
+ pred_track_instances.labels = labels
+ pred_track_instances.scores = scores
+ pred_track_instances.instances_id = ids
+
+ return pred_track_instances
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/strongsort_tracker.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/strongsort_tracker.py
new file mode 100644
index 0000000000000000000000000000000000000000..9d7075701bc3205b9ea30f03790cfa1c42a97822
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/trackers/strongsort_tracker.py
@@ -0,0 +1,273 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Tuple
+
+import numpy as np
+import torch
+from mmengine.structures import InstanceData
+
+try:
+ import motmetrics
+ from motmetrics.lap import linear_sum_assignment
+except ImportError:
+ motmetrics = None
+from torch import Tensor
+
+from mmdet.models.utils import imrenormalize
+from mmdet.registry import MODELS
+from mmdet.structures import TrackDataSample
+from mmdet.structures.bbox import bbox_overlaps, bbox_xyxy_to_cxcyah
+from mmdet.utils import OptConfigType
+from .sort_tracker import SORTTracker
+
+
+def cosine_distance(x: Tensor, y: Tensor) -> np.ndarray:
+ """compute the cosine distance.
+
+ Args:
+ x (Tensor): embeddings with shape (N,C).
+ y (Tensor): embeddings with shape (M,C).
+
+ Returns:
+ ndarray: cosine distance with shape (N,M).
+ """
+ x = x.cpu().numpy()
+ y = y.cpu().numpy()
+ x = x / np.linalg.norm(x, axis=1, keepdims=True)
+ y = y / np.linalg.norm(y, axis=1, keepdims=True)
+ dists = 1. - np.dot(x, y.T)
+ return dists
+
+
+@MODELS.register_module()
+class StrongSORTTracker(SORTTracker):
+ """Tracker for StrongSORT.
+
+ Args:
+ obj_score_thr (float, optional): Threshold to filter the objects.
+ Defaults to 0.6.
+ motion (dict): Configuration of motion. Defaults to None.
+ reid (dict, optional): Configuration for the ReID model.
+ - num_samples (int, optional): Number of samples to calculate the
+ feature embeddings of a track. Default to None.
+ - image_scale (tuple, optional): Input scale of the ReID model.
+ Default to (256, 128).
+ - img_norm_cfg (dict, optional): Configuration to normalize the
+ input. Default to None.
+ - match_score_thr (float, optional): Similarity threshold for the
+ matching process. Default to 0.3.
+ - motion_weight (float, optional): the weight of the motion cost.
+ Defaults to 0.02.
+ match_iou_thr (float, optional): Threshold of the IoU matching process.
+ Defaults to 0.7.
+ num_tentatives (int, optional): Number of continuous frames to confirm
+ a track. Defaults to 2.
+ """
+
+ def __init__(self,
+ motion: Optional[dict] = None,
+ obj_score_thr: float = 0.6,
+ reid: dict = dict(
+ num_samples=None,
+ img_scale=(256, 128),
+ img_norm_cfg=None,
+ match_score_thr=0.3,
+ motion_weight=0.02),
+ match_iou_thr: float = 0.7,
+ num_tentatives: int = 2,
+ **kwargs):
+ if motmetrics is None:
+ raise RuntimeError('motmetrics is not installed,\
+ please install it by: pip install motmetrics')
+ super().__init__(motion, obj_score_thr, reid, match_iou_thr,
+ num_tentatives, **kwargs)
+
+ def update_track(self, id: int, obj: Tuple[Tensor]) -> None:
+ """Update a track."""
+ for k, v in zip(self.memo_items, obj):
+ v = v[None]
+ if self.momentums is not None and k in self.momentums:
+ m = self.momentums[k]
+ self.tracks[id][k] = (1 - m) * self.tracks[id][k] + m * v
+ else:
+ self.tracks[id][k].append(v)
+
+ if self.tracks[id].tentative:
+ if len(self.tracks[id]['bboxes']) >= self.num_tentatives:
+ self.tracks[id].tentative = False
+ bbox = bbox_xyxy_to_cxcyah(self.tracks[id].bboxes[-1]) # size = (1, 4)
+ assert bbox.ndim == 2 and bbox.shape[0] == 1
+ bbox = bbox.squeeze(0).cpu().numpy()
+ score = float(self.tracks[id].scores[-1].cpu())
+ self.tracks[id].mean, self.tracks[id].covariance = self.kf.update(
+ self.tracks[id].mean, self.tracks[id].covariance, bbox, score)
+
+ def track(self,
+ model: torch.nn.Module,
+ img: Tensor,
+ data_sample: TrackDataSample,
+ data_preprocessor: OptConfigType = None,
+ rescale: bool = False,
+ **kwargs) -> InstanceData:
+ """Tracking forward function.
+
+ Args:
+ model (nn.Module): MOT model.
+ img (Tensor): of shape (T, C, H, W) encoding input image.
+ Typically these should be mean centered and std scaled.
+ The T denotes the number of key images and usually is 1 in
+ SORT method.
+ feats (list[Tensor]): Multi level feature maps of `img`.
+ data_sample (:obj:`TrackDataSample`): The data sample.
+ It includes information such as `pred_det_instances`.
+ data_preprocessor (dict or ConfigDict, optional): The pre-process
+ config of :class:`TrackDataPreprocessor`. it usually includes,
+ ``pad_size_divisor``, ``pad_value``, ``mean`` and ``std``.
+ rescale (bool, optional): If True, the bounding boxes should be
+ rescaled to fit the original scale of the image. Defaults to
+ False.
+
+ Returns:
+ :obj:`InstanceData`: Tracking results of the input images.
+ Each InstanceData usually contains ``bboxes``, ``labels``,
+ ``scores`` and ``instances_id``.
+ """
+ metainfo = data_sample.metainfo
+ bboxes = data_sample.pred_instances.bboxes
+ labels = data_sample.pred_instances.labels
+ scores = data_sample.pred_instances.scores
+
+ frame_id = metainfo.get('frame_id', -1)
+ if frame_id == 0:
+ self.reset()
+ if not hasattr(self, 'kf'):
+ self.kf = self.motion
+
+ if self.with_reid:
+ if self.reid.get('img_norm_cfg', False):
+ img_norm_cfg = dict(
+ mean=data_preprocessor.get('mean', [0, 0, 0]),
+ std=data_preprocessor.get('std', [1, 1, 1]),
+ to_bgr=data_preprocessor.get('rgb_to_bgr', False))
+ reid_img = imrenormalize(img, img_norm_cfg,
+ self.reid['img_norm_cfg'])
+ else:
+ reid_img = img.clone()
+
+ valid_inds = scores > self.obj_score_thr
+ bboxes = bboxes[valid_inds]
+ labels = labels[valid_inds]
+ scores = scores[valid_inds]
+
+ if self.empty or bboxes.size(0) == 0:
+ num_new_tracks = bboxes.size(0)
+ ids = torch.arange(
+ self.num_tracks,
+ self.num_tracks + num_new_tracks,
+ dtype=torch.long).to(bboxes.device)
+ self.num_tracks += num_new_tracks
+ if self.with_reid:
+ crops = self.crop_imgs(reid_img, metainfo, bboxes.clone(),
+ rescale)
+ if crops.size(0) > 0:
+ embeds = model.reid(crops, mode='tensor')
+ else:
+ embeds = crops.new_zeros((0, model.reid.head.out_channels))
+ else:
+ ids = torch.full((bboxes.size(0), ), -1,
+ dtype=torch.long).to(bboxes.device)
+
+ # motion
+ if model.with_cmc:
+ num_samples = 1
+ self.tracks = model.cmc.track(self.last_img, img, self.tracks,
+ num_samples, frame_id, metainfo)
+
+ self.tracks, motion_dists = self.motion.track(
+ self.tracks, bbox_xyxy_to_cxcyah(bboxes))
+
+ active_ids = self.confirmed_ids
+ if self.with_reid:
+ crops = self.crop_imgs(reid_img, metainfo, bboxes.clone(),
+ rescale)
+ embeds = model.reid(crops, mode='tensor')
+
+ # reid
+ if len(active_ids) > 0:
+ track_embeds = self.get(
+ 'embeds',
+ active_ids,
+ self.reid.get('num_samples', None),
+ behavior='mean')
+ reid_dists = cosine_distance(track_embeds, embeds)
+ valid_inds = [list(self.ids).index(_) for _ in active_ids]
+ reid_dists[~np.isfinite(motion_dists[
+ valid_inds, :])] = np.nan
+
+ weight_motion = self.reid.get('motion_weight')
+ match_dists = (1 - weight_motion) * reid_dists + \
+ weight_motion * motion_dists[valid_inds]
+
+ # support multi-class association
+ track_labels = torch.tensor([
+ self.tracks[id]['labels'][-1] for id in active_ids
+ ]).to(bboxes.device)
+ cate_match = labels[None, :] == track_labels[:, None]
+ cate_cost = ((1 - cate_match.int()) * 1e6).cpu().numpy()
+ match_dists = match_dists + cate_cost
+
+ row, col = linear_sum_assignment(match_dists)
+ for r, c in zip(row, col):
+ dist = match_dists[r, c]
+ if not np.isfinite(dist):
+ continue
+ if dist <= self.reid['match_score_thr']:
+ ids[c] = active_ids[r]
+
+ active_ids = [
+ id for id in self.ids if id not in ids
+ and self.tracks[id].frame_ids[-1] == frame_id - 1
+ ]
+ if len(active_ids) > 0:
+ active_dets = torch.nonzero(ids == -1).squeeze(1)
+ track_bboxes = self.get('bboxes', active_ids)
+ ious = bbox_overlaps(track_bboxes, bboxes[active_dets])
+
+ # support multi-class association
+ track_labels = torch.tensor([
+ self.tracks[id]['labels'][-1] for id in active_ids
+ ]).to(bboxes.device)
+ cate_match = labels[None, active_dets] == track_labels[:, None]
+ cate_cost = (1 - cate_match.int()) * 1e6
+
+ dists = (1 - ious + cate_cost).cpu().numpy()
+
+ row, col = linear_sum_assignment(dists)
+ for r, c in zip(row, col):
+ dist = dists[r, c]
+ if dist < 1 - self.match_iou_thr:
+ ids[active_dets[c]] = active_ids[r]
+
+ new_track_inds = ids == -1
+ ids[new_track_inds] = torch.arange(
+ self.num_tracks,
+ self.num_tracks + new_track_inds.sum(),
+ dtype=torch.long).to(bboxes.device)
+ self.num_tracks += new_track_inds.sum()
+
+ self.update(
+ ids=ids,
+ bboxes=bboxes,
+ scores=scores,
+ labels=labels,
+ embeds=embeds if self.with_reid else None,
+ frame_ids=frame_id)
+ self.last_img = img
+
+ # update pred_track_instances
+ pred_track_instances = InstanceData()
+ pred_track_instances.bboxes = bboxes
+ pred_track_instances.labels = labels
+ pred_track_instances.scores = scores
+ pred_track_instances.instances_id = ids
+
+ return pred_track_instances
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/tracking_heads/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/tracking_heads/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..bd1f0561cc076f2a603a64eb479cc6de0372a438
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/tracking_heads/__init__.py
@@ -0,0 +1,11 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .mask2former_track_head import Mask2FormerTrackHead
+from .quasi_dense_embed_head import QuasiDenseEmbedHead
+from .quasi_dense_track_head import QuasiDenseTrackHead
+from .roi_embed_head import RoIEmbedHead
+from .roi_track_head import RoITrackHead
+
+__all__ = [
+ 'QuasiDenseEmbedHead', 'QuasiDenseTrackHead', 'Mask2FormerTrackHead',
+ 'RoIEmbedHead', 'RoITrackHead'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/tracking_heads/mask2former_track_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/tracking_heads/mask2former_track_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..0877241bc33fcd1ef8f7ed154d503d9dbd8ab938
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/tracking_heads/mask2former_track_head.py
@@ -0,0 +1,729 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from collections import defaultdict
+from typing import Dict, List, Tuple
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import Conv2d
+from mmcv.ops import point_sample
+from mmengine.model import ModuleList
+from mmengine.model.weight_init import caffe2_xavier_init
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.models.dense_heads import AnchorFreeHead, MaskFormerHead
+from mmdet.models.utils import get_uncertain_point_coords_with_randomness
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures import TrackDataSample, TrackSampleList
+from mmdet.structures.mask import mask2bbox
+from mmdet.utils import (ConfigType, InstanceList, OptConfigType,
+ OptMultiConfig, reduce_mean)
+from ..layers import Mask2FormerTransformerDecoder
+
+
+@MODELS.register_module()
+class Mask2FormerTrackHead(MaskFormerHead):
+ """Implements the Mask2Former head.
+
+ See `Masked-attention Mask Transformer for Universal Image
+ Segmentation `_ for details.
+
+ Args:
+ in_channels (list[int]): Number of channels in the input feature map.
+ feat_channels (int): Number of channels for features.
+ out_channels (int): Number of channels for output.
+ num_classes (int): Number of VIS classes.
+ num_queries (int): Number of query in Transformer decoder.
+ Defaults to 100.
+ num_transformer_feat_level (int): Number of feats levels.
+ Defaults to 3.
+ pixel_decoder (:obj:`ConfigDict` or dict): Config for pixel
+ decoder.
+ enforce_decoder_input_project (bool, optional): Whether to add
+ a layer to change the embed_dim of transformer encoder in
+ pixel decoder to the embed_dim of transformer decoder.
+ Defaults to False.
+ transformer_decoder (:obj:`ConfigDict` or dict): Config for
+ transformer decoder.
+ positional_encoding (:obj:`ConfigDict` or dict): Config for
+ transformer decoder position encoding.
+ Defaults to `SinePositionalEncoding3D`.
+ loss_cls (:obj:`ConfigDict` or dict): Config of the classification
+ loss. Defaults to `CrossEntropyLoss`.
+ loss_mask (:obj:`ConfigDict` or dict): Config of the mask loss.
+ Defaults to 'CrossEntropyLoss'.
+ loss_dice (:obj:`ConfigDict` or dict): Config of the dice loss.
+ Defaults to 'DiceLoss'.
+ train_cfg (:obj:`ConfigDict` or dict, optional): Training config of
+ Mask2Former head. Defaults to None.
+ test_cfg (:obj:`ConfigDict` or dict, optional): Testing config of
+ Mask2Former head. Defaults to None.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict], optional): Initialization config dict. Defaults to None.
+ """
+
+ def __init__(self,
+ in_channels: List[int],
+ feat_channels: int,
+ out_channels: int,
+ num_classes: int,
+ num_frames: int = 2,
+ num_queries: int = 100,
+ num_transformer_feat_level: int = 3,
+ pixel_decoder: ConfigType = ...,
+ enforce_decoder_input_project: bool = False,
+ transformer_decoder: ConfigType = ...,
+ positional_encoding: ConfigType = dict(
+ num_feats=128, normalize=True),
+ loss_cls: ConfigType = dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=2.0,
+ reduction='mean',
+ class_weight=[1.0] * 133 + [0.1]),
+ loss_mask: ConfigType = dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ reduction='mean',
+ loss_weight=5.0),
+ loss_dice: ConfigType = dict(
+ type='DiceLoss',
+ use_sigmoid=True,
+ activate=True,
+ reduction='mean',
+ naive_dice=True,
+ eps=1.0,
+ loss_weight=5.0),
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ init_cfg: OptMultiConfig = None,
+ **kwargs) -> None:
+ super(AnchorFreeHead, self).__init__(init_cfg=init_cfg)
+ self.num_classes = num_classes
+ self.num_frames = num_frames
+ self.num_queries = num_queries
+ self.num_transformer_feat_level = num_transformer_feat_level
+ self.num_transformer_feat_level = num_transformer_feat_level
+ self.num_heads = transformer_decoder.layer_cfg.cross_attn_cfg.num_heads
+ self.num_transformer_decoder_layers = transformer_decoder.num_layers
+ assert pixel_decoder.encoder.layer_cfg. \
+ self_attn_cfg.num_levels == num_transformer_feat_level
+ pixel_decoder_ = copy.deepcopy(pixel_decoder)
+ pixel_decoder_.update(
+ in_channels=in_channels,
+ feat_channels=feat_channels,
+ out_channels=out_channels)
+ self.pixel_decoder = MODELS.build(pixel_decoder_)
+ self.transformer_decoder = Mask2FormerTransformerDecoder(
+ **transformer_decoder)
+ self.decoder_embed_dims = self.transformer_decoder.embed_dims
+
+ self.decoder_input_projs = ModuleList()
+ # from low resolution to high resolution
+ for _ in range(num_transformer_feat_level):
+ if (self.decoder_embed_dims != feat_channels
+ or enforce_decoder_input_project):
+ self.decoder_input_projs.append(
+ Conv2d(
+ feat_channels, self.decoder_embed_dims, kernel_size=1))
+ else:
+ self.decoder_input_projs.append(nn.Identity())
+ self.decoder_positional_encoding = MODELS.build(positional_encoding)
+ self.query_embed = nn.Embedding(self.num_queries, feat_channels)
+ self.query_feat = nn.Embedding(self.num_queries, feat_channels)
+ # from low resolution to high resolution
+ self.level_embed = nn.Embedding(self.num_transformer_feat_level,
+ feat_channels)
+
+ self.cls_embed = nn.Linear(feat_channels, self.num_classes + 1)
+ self.mask_embed = nn.Sequential(
+ nn.Linear(feat_channels, feat_channels), nn.ReLU(inplace=True),
+ nn.Linear(feat_channels, feat_channels), nn.ReLU(inplace=True),
+ nn.Linear(feat_channels, out_channels))
+
+ self.test_cfg = test_cfg
+ self.train_cfg = train_cfg
+ if train_cfg:
+ self.assigner = TASK_UTILS.build(self.train_cfg.assigner)
+ self.sampler = TASK_UTILS.build(
+ # self.train_cfg.sampler, default_args=dict(context=self))
+ self.train_cfg['sampler'],
+ default_args=dict(context=self))
+ self.num_points = self.train_cfg.get('num_points', 12544)
+ self.oversample_ratio = self.train_cfg.get('oversample_ratio', 3.0)
+ self.importance_sample_ratio = self.train_cfg.get(
+ 'importance_sample_ratio', 0.75)
+
+ self.class_weight = loss_cls.class_weight
+ self.loss_cls = MODELS.build(loss_cls)
+ self.loss_mask = MODELS.build(loss_mask)
+ self.loss_dice = MODELS.build(loss_dice)
+
+ def init_weights(self) -> None:
+ for m in self.decoder_input_projs:
+ if isinstance(m, Conv2d):
+ caffe2_xavier_init(m, bias=0)
+
+ self.pixel_decoder.init_weights()
+
+ for p in self.transformer_decoder.parameters():
+ if p.dim() > 1:
+ nn.init.xavier_normal_(p)
+
+ def preprocess_gt(self, batch_gt_instances: InstanceList) -> InstanceList:
+ """Preprocess the ground truth for all images.
+
+ It aims to reorganize the `gt`. For example, in the
+ `batch_data_sample.gt_instances.mask`, its shape is
+ `(all_num_gts, h, w)`, but we don't know each gt belongs to which `img`
+ (assume `num_frames` is 2). So, this func used to reshape the `gt_mask`
+ to `(num_gts_per_img, num_frames, h, w)`. In addition, we can't
+ guarantee that the number of instances in these two images is equal,
+ so `-1` refers to nonexistent instances.
+
+ Args:
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``labels``, each is
+ ground truth labels of each bbox, with shape (num_gts, )
+ and ``masks``, each is ground truth masks of each instances
+ of an image, shape (num_gts, h, w).
+
+ Returns:
+ list[obj:`InstanceData`]: each contains the following keys
+
+ - labels (Tensor): Ground truth class indices\
+ for an image, with shape (n, ), n is the sum of\
+ number of stuff type and number of instance in an image.
+ - masks (Tensor): Ground truth mask for a\
+ image, with shape (n, t, h, w).
+ """
+ final_batch_gt_instances = []
+ batch_size = len(batch_gt_instances) // self.num_frames
+ for batch_idx in range(batch_size):
+ pair_gt_insatences = batch_gt_instances[batch_idx *
+ self.num_frames:batch_idx *
+ self.num_frames +
+ self.num_frames]
+
+ assert len(
+ pair_gt_insatences
+ ) > 1, f'mask2former for vis need multi frames to train, \
+ but you only use {len(pair_gt_insatences)} frames'
+
+ _device = pair_gt_insatences[0].labels.device
+
+ for gt_instances in pair_gt_insatences:
+ gt_instances.masks = gt_instances.masks.to_tensor(
+ dtype=torch.bool, device=_device)
+ all_ins_id = torch.cat([
+ gt_instances.instances_ids
+ for gt_instances in pair_gt_insatences
+ ])
+ all_ins_id = all_ins_id.unique().tolist()
+ map_ins_id = dict()
+ for i, ins_id in enumerate(all_ins_id):
+ map_ins_id[ins_id] = i
+
+ num_instances = len(all_ins_id)
+ mask_shape = [
+ num_instances, self.num_frames,
+ pair_gt_insatences[0].masks.shape[1],
+ pair_gt_insatences[0].masks.shape[2]
+ ]
+ gt_masks_per_video = torch.zeros(
+ mask_shape, dtype=torch.bool, device=_device)
+ gt_ids_per_video = torch.full((num_instances, self.num_frames),
+ -1,
+ dtype=torch.long,
+ device=_device)
+ gt_labels_per_video = torch.full((num_instances, ),
+ -1,
+ dtype=torch.long,
+ device=_device)
+
+ for frame_id in range(self.num_frames):
+ cur_frame_gts = pair_gt_insatences[frame_id]
+ ins_ids = cur_frame_gts.instances_ids.tolist()
+ for i, id in enumerate(ins_ids):
+ gt_masks_per_video[map_ins_id[id],
+ frame_id, :, :] = cur_frame_gts.masks[i]
+ gt_ids_per_video[map_ins_id[id],
+ frame_id] = cur_frame_gts.instances_ids[i]
+ gt_labels_per_video[
+ map_ins_id[id]] = cur_frame_gts.labels[i]
+
+ tmp_instances = InstanceData(
+ labels=gt_labels_per_video,
+ masks=gt_masks_per_video.long(),
+ instances_id=gt_ids_per_video)
+ final_batch_gt_instances.append(tmp_instances)
+
+ return final_batch_gt_instances
+
+ def _get_targets_single(self, cls_score: Tensor, mask_pred: Tensor,
+ gt_instances: InstanceData,
+ img_meta: dict) -> Tuple[Tensor]:
+ """Compute classification and mask targets for one image.
+
+ Args:
+ cls_score (Tensor): Mask score logits from a single decoder layer
+ for one image. Shape (num_queries, cls_out_channels).
+ mask_pred (Tensor): Mask logits for a single decoder layer for one
+ image. Shape (num_queries, num_frames, h, w).
+ gt_instances (:obj:`InstanceData`): It contains ``labels`` and
+ ``masks``.
+ img_meta (dict): Image informtation.
+
+ Returns:
+ tuple[Tensor]: A tuple containing the following for one image.
+
+ - labels (Tensor): Labels of each image. \
+ shape (num_queries, ).
+ - label_weights (Tensor): Label weights of each image. \
+ shape (num_queries, ).
+ - mask_targets (Tensor): Mask targets of each image. \
+ shape (num_queries, num_frames, h, w).
+ - mask_weights (Tensor): Mask weights of each image. \
+ shape (num_queries, ).
+ - pos_inds (Tensor): Sampled positive indices for each \
+ image.
+ - neg_inds (Tensor): Sampled negative indices for each \
+ image.
+ - sampling_result (:obj:`SamplingResult`): Sampling results.
+ """
+ # (num_gts, )
+ gt_labels = gt_instances.labels
+ # (num_gts, num_frames, h, w)
+ gt_masks = gt_instances.masks
+ # sample points
+ num_queries = cls_score.shape[0]
+ num_gts = gt_labels.shape[0]
+
+ point_coords = torch.rand((1, self.num_points, 2),
+ device=cls_score.device)
+
+ # shape (num_queries, num_points)
+ mask_points_pred = point_sample(mask_pred,
+ point_coords.repeat(num_queries, 1,
+ 1)).flatten(1)
+ # shape (num_gts, num_points)
+ gt_points_masks = point_sample(gt_masks.float(),
+ point_coords.repeat(num_gts, 1,
+ 1)).flatten(1)
+
+ sampled_gt_instances = InstanceData(
+ labels=gt_labels, masks=gt_points_masks)
+ sampled_pred_instances = InstanceData(
+ scores=cls_score, masks=mask_points_pred)
+ # assign and sample
+ assign_result = self.assigner.assign(
+ pred_instances=sampled_pred_instances,
+ gt_instances=sampled_gt_instances,
+ img_meta=img_meta)
+ pred_instances = InstanceData(scores=cls_score, masks=mask_pred)
+ sampling_result = self.sampler.sample(
+ assign_result=assign_result,
+ pred_instances=pred_instances,
+ gt_instances=gt_instances)
+ pos_inds = sampling_result.pos_inds
+ neg_inds = sampling_result.neg_inds
+
+ # label target
+ labels = gt_labels.new_full((self.num_queries, ),
+ self.num_classes,
+ dtype=torch.long)
+ labels[pos_inds] = gt_labels[sampling_result.pos_assigned_gt_inds]
+ label_weights = gt_labels.new_ones((self.num_queries, ))
+
+ # mask target
+ mask_targets = gt_masks[sampling_result.pos_assigned_gt_inds]
+ mask_weights = mask_pred.new_zeros((self.num_queries, ))
+ mask_weights[pos_inds] = 1.0
+
+ return (labels, label_weights, mask_targets, mask_weights, pos_inds,
+ neg_inds, sampling_result)
+
+ def _loss_by_feat_single(self, cls_scores: Tensor, mask_preds: Tensor,
+ batch_gt_instances: List[InstanceData],
+ batch_img_metas: List[dict]) -> Tuple[Tensor]:
+ """Loss function for outputs from a single decoder layer.
+
+ Args:
+ cls_scores (Tensor): Mask score logits from a single decoder layer
+ for all images. Shape (batch_size, num_queries,
+ cls_out_channels). Note `cls_out_channels` should include
+ background.
+ mask_preds (Tensor): Mask logits for a pixel decoder for all
+ images. Shape (batch_size, num_queries, num_frames,h, w).
+ batch_gt_instances (list[obj:`InstanceData`]): each contains
+ ``labels`` and ``masks``.
+ batch_img_metas (list[dict]): List of image meta information.
+
+ Returns:
+ tuple[Tensor]: Loss components for outputs from a single \
+ decoder layer.
+ """
+ num_imgs = cls_scores.size(0)
+ cls_scores_list = [cls_scores[i] for i in range(num_imgs)]
+ mask_preds_list = [mask_preds[i] for i in range(num_imgs)]
+ (labels_list, label_weights_list, mask_targets_list, mask_weights_list,
+ avg_factor) = self.get_targets(cls_scores_list, mask_preds_list,
+ batch_gt_instances, batch_img_metas)
+ # shape (batch_size, num_queries)
+ labels = torch.stack(labels_list, dim=0)
+ # shape (batch_size, num_queries)
+ label_weights = torch.stack(label_weights_list, dim=0)
+ # shape (num_total_gts, num_frames, h, w)
+ mask_targets = torch.cat(mask_targets_list, dim=0)
+ # shape (batch_size, num_queries)
+ mask_weights = torch.stack(mask_weights_list, dim=0)
+
+ # classfication loss
+ # shape (batch_size * num_queries, )
+ cls_scores = cls_scores.flatten(0, 1)
+ labels = labels.flatten(0, 1)
+ label_weights = label_weights.flatten(0, 1)
+
+ class_weight = cls_scores.new_tensor(self.class_weight)
+ loss_cls = self.loss_cls(
+ cls_scores,
+ labels,
+ label_weights,
+ avg_factor=class_weight[labels].sum())
+
+ num_total_masks = reduce_mean(cls_scores.new_tensor([avg_factor]))
+ num_total_masks = max(num_total_masks, 1)
+
+ # extract positive ones
+ # shape (batch_size, num_queries, num_frames, h, w)
+ # -> (num_total_gts, num_frames, h, w)
+ mask_preds = mask_preds[mask_weights > 0]
+
+ if mask_targets.shape[0] == 0:
+ # zero match
+ loss_dice = mask_preds.sum()
+ loss_mask = mask_preds.sum()
+ return loss_cls, loss_mask, loss_dice
+
+ with torch.no_grad():
+ points_coords = get_uncertain_point_coords_with_randomness(
+ mask_preds.flatten(0, 1).unsqueeze(1), None, self.num_points,
+ self.oversample_ratio, self.importance_sample_ratio)
+ # shape (num_total_gts * num_frames, h, w) ->
+ # (num_total_gts, num_points)
+ mask_point_targets = point_sample(
+ mask_targets.flatten(0, 1).unsqueeze(1).float(),
+ points_coords).squeeze(1)
+ # shape (num_total_gts * num_frames, num_points)
+ mask_point_preds = point_sample(
+ mask_preds.flatten(0, 1).unsqueeze(1), points_coords).squeeze(1)
+
+ # dice loss
+ loss_dice = self.loss_dice(
+ mask_point_preds, mask_point_targets, avg_factor=num_total_masks)
+
+ # mask loss
+ # shape (num_total_gts * num_frames, num_points) ->
+ # (num_total_gts * num_frames * num_points, )
+ mask_point_preds = mask_point_preds.reshape(-1)
+ # shape (num_total_gts, num_points) -> (num_total_gts * num_points, )
+ mask_point_targets = mask_point_targets.reshape(-1)
+ loss_mask = self.loss_mask(
+ mask_point_preds,
+ mask_point_targets,
+ avg_factor=num_total_masks * self.num_points / self.num_frames)
+
+ return loss_cls, loss_mask, loss_dice
+
+ def _forward_head(
+ self, decoder_out: Tensor, mask_feature: Tensor,
+ attn_mask_target_size: Tuple[int,
+ int]) -> Tuple[Tensor, Tensor, Tensor]:
+ """Forward for head part which is called after every decoder layer.
+
+ Args:
+ decoder_out (Tensor): in shape (num_queries, batch_size, c).
+ mask_feature (Tensor): in shape (batch_size, t, c, h, w).
+ attn_mask_target_size (tuple[int, int]): target attention
+ mask size.
+
+ Returns:
+ tuple: A tuple contain three elements.
+
+ - cls_pred (Tensor): Classification scores in shape \
+ (batch_size, num_queries, cls_out_channels). \
+ Note `cls_out_channels` should include background.
+ - mask_pred (Tensor): Mask scores in shape \
+ (batch_size, num_queries,h, w).
+ - attn_mask (Tensor): Attention mask in shape \
+ (batch_size * num_heads, num_queries, h, w).
+ """
+ decoder_out = self.transformer_decoder.post_norm(decoder_out)
+ cls_pred = self.cls_embed(decoder_out)
+ mask_embed = self.mask_embed(decoder_out)
+
+ # shape (batch_size, num_queries, t, h, w)
+ mask_pred = torch.einsum('bqc,btchw->bqthw', mask_embed, mask_feature)
+ b, q, t, _, _ = mask_pred.shape
+
+ attn_mask = F.interpolate(
+ mask_pred.flatten(0, 1),
+ attn_mask_target_size,
+ mode='bilinear',
+ align_corners=False).view(b, q, t, attn_mask_target_size[0],
+ attn_mask_target_size[1])
+
+ # shape (batch_size, num_queries, t, h, w) ->
+ # (batch_size, num_queries, t*h*w) ->
+ # (batch_size, num_head, num_queries, t*h*w) ->
+ # (batch_size*num_head, num_queries, t*h*w)
+ attn_mask = attn_mask.flatten(2).unsqueeze(1).repeat(
+ (1, self.num_heads, 1, 1)).flatten(0, 1)
+ attn_mask = attn_mask.sigmoid() < 0.5
+ attn_mask = attn_mask.detach()
+
+ return cls_pred, mask_pred, attn_mask
+
+ def forward(
+ self, x: List[Tensor], data_samples: TrackDataSample
+ ) -> Tuple[List[Tensor], List[Tensor]]:
+ """Forward function.
+
+ Args:
+ x (list[Tensor]): Multi scale Features from the
+ upstream network, each is a 4D-tensor.
+ data_samples (List[:obj:`TrackDataSample`]): The Data
+ Samples. It usually includes information such as `gt_instance`.
+
+ Returns:
+ tuple[list[Tensor]]: A tuple contains two elements.
+
+ - cls_pred_list (list[Tensor)]: Classification logits \
+ for each decoder layer. Each is a 3D-tensor with shape \
+ (batch_size, num_queries, cls_out_channels). \
+ Note `cls_out_channels` should include background.
+ - mask_pred_list (list[Tensor]): Mask logits for each \
+ decoder layer. Each with shape (batch_size, num_queries, \
+ h, w).
+ """
+ mask_features, multi_scale_memorys = self.pixel_decoder(x)
+ bt, c_m, h_m, w_m = mask_features.shape
+ batch_size = bt // self.num_frames if self.training else 1
+ t = bt // batch_size
+ mask_features = mask_features.view(batch_size, t, c_m, h_m, w_m)
+ # multi_scale_memorys (from low resolution to high resolution)
+ decoder_inputs = []
+ decoder_positional_encodings = []
+ for i in range(self.num_transformer_feat_level):
+ decoder_input = self.decoder_input_projs[i](multi_scale_memorys[i])
+ decoder_input = decoder_input.flatten(2)
+ level_embed = self.level_embed.weight[i][None, :, None]
+ decoder_input = decoder_input + level_embed
+ _, c, hw = decoder_input.shape
+ # shape (batch_size*t, c, h, w) ->
+ # (batch_size, t, c, hw) ->
+ # (batch_size, t*h*w, c)
+ decoder_input = decoder_input.view(batch_size, t, c,
+ hw).permute(0, 1, 3,
+ 2).flatten(1, 2)
+ # shape (batch_size, c, h, w) -> (h*w, batch_size, c)
+ mask = decoder_input.new_zeros(
+ (batch_size, t) + multi_scale_memorys[i].shape[-2:],
+ dtype=torch.bool)
+ decoder_positional_encoding = self.decoder_positional_encoding(
+ mask)
+ decoder_positional_encoding = decoder_positional_encoding.flatten(
+ 3).permute(0, 1, 3, 2).flatten(1, 2)
+ decoder_inputs.append(decoder_input)
+ decoder_positional_encodings.append(decoder_positional_encoding)
+ # shape (num_queries, c) -> (batch_size, num_queries, c)
+ query_feat = self.query_feat.weight.unsqueeze(0).repeat(
+ (batch_size, 1, 1))
+ query_embed = self.query_embed.weight.unsqueeze(0).repeat(
+ (batch_size, 1, 1))
+
+ cls_pred_list = []
+ mask_pred_list = []
+ cls_pred, mask_pred, attn_mask = self._forward_head(
+ query_feat, mask_features, multi_scale_memorys[0].shape[-2:])
+ cls_pred_list.append(cls_pred)
+ mask_pred_list.append(mask_pred)
+
+ for i in range(self.num_transformer_decoder_layers):
+ level_idx = i % self.num_transformer_feat_level
+ # if a mask is all True(all background), then set it all False.
+ attn_mask[torch.where(
+ attn_mask.sum(-1) == attn_mask.shape[-1])] = False
+
+ # cross_attn + self_attn
+ layer = self.transformer_decoder.layers[i]
+ query_feat = layer(
+ query=query_feat,
+ key=decoder_inputs[level_idx],
+ value=decoder_inputs[level_idx],
+ query_pos=query_embed,
+ key_pos=decoder_positional_encodings[level_idx],
+ cross_attn_mask=attn_mask,
+ query_key_padding_mask=None,
+ # here we do not apply masking on padded region
+ key_padding_mask=None)
+ cls_pred, mask_pred, attn_mask = self._forward_head(
+ query_feat, mask_features, multi_scale_memorys[
+ (i + 1) % self.num_transformer_feat_level].shape[-2:])
+
+ cls_pred_list.append(cls_pred)
+ mask_pred_list.append(mask_pred)
+
+ return cls_pred_list, mask_pred_list
+
+ def loss(
+ self,
+ x: Tuple[Tensor],
+ data_samples: TrackSampleList,
+ ) -> Dict[str, Tensor]:
+ """Perform forward propagation and loss calculation of the track head
+ on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Multi-level features from the upstream
+ network, each is a 4D-tensor.
+ data_samples (List[:obj:`TrackDataSample`]): The Data
+ Samples. It usually includes information such as `gt_instance`.
+
+ Returns:
+ dict[str, Tensor]: a dictionary of loss components
+ """
+ batch_img_metas = []
+ batch_gt_instances = []
+
+ for data_sample in data_samples:
+ video_img_metas = defaultdict(list)
+ for image_idx in range(len(data_sample)):
+ batch_gt_instances.append(data_sample[image_idx].gt_instances)
+ for key, value in data_sample[image_idx].metainfo.items():
+ video_img_metas[key].append(value)
+ batch_img_metas.append(video_img_metas)
+
+ # forward
+ all_cls_scores, all_mask_preds = self(x, data_samples)
+
+ # preprocess ground truth
+ batch_gt_instances = self.preprocess_gt(batch_gt_instances)
+ # loss
+ losses = self.loss_by_feat(all_cls_scores, all_mask_preds,
+ batch_gt_instances, batch_img_metas)
+
+ return losses
+
+ def predict(self,
+ x: Tuple[Tensor],
+ data_samples: TrackDataSample,
+ rescale: bool = True) -> InstanceList:
+ """Test without augmentation.
+
+ Args:
+ x (tuple[Tensor]): Multi-level features from the
+ upstream network, each is a 4D-tensor.
+ data_samples (List[:obj:`TrackDataSample`]): The Data
+ Samples. It usually includes information such as `gt_instance`.
+ rescale (bool, Optional): If False, then returned bboxes and masks
+ will fit the scale of img, otherwise, returned bboxes and masks
+ will fit the scale of original image shape. Defaults to True.
+
+ Returns:
+ list[obj:`InstanceData`]: each contains the following keys
+ - labels (Tensor): Prediction class indices\
+ for an image, with shape (n, ), n is the sum of\
+ number of stuff type and number of instance in an image.
+ - masks (Tensor): Prediction mask for a\
+ image, with shape (n, t, h, w).
+ """
+
+ batch_img_metas = [
+ data_samples[img_idx].metainfo
+ for img_idx in range(len(data_samples))
+ ]
+ all_cls_scores, all_mask_preds = self(x, data_samples)
+ mask_cls_results = all_cls_scores[-1]
+ mask_pred_results = all_mask_preds[-1]
+
+ mask_cls_results = mask_cls_results[0]
+ # upsample masks
+ img_shape = batch_img_metas[0]['batch_input_shape']
+ mask_pred_results = F.interpolate(
+ mask_pred_results[0],
+ size=(img_shape[0], img_shape[1]),
+ mode='bilinear',
+ align_corners=False)
+
+ results = self.predict_by_feat(mask_cls_results, mask_pred_results,
+ batch_img_metas)
+ return results
+
+ def predict_by_feat(self,
+ mask_cls_results: List[Tensor],
+ mask_pred_results: List[Tensor],
+ batch_img_metas: List[dict],
+ rescale: bool = True) -> InstanceList:
+ """Get top-10 predictions.
+
+ Args:
+ mask_cls_results (Tensor): Mask classification logits,\
+ shape (batch_size, num_queries, cls_out_channels).
+ Note `cls_out_channels` should include background.
+ mask_pred_results (Tensor): Mask logits, shape \
+ (batch_size, num_queries, h, w).
+ batch_img_metas (list[dict]): List of image meta information.
+ rescale (bool, Optional): If False, then returned bboxes and masks
+ will fit the scale of img, otherwise, returned bboxes and masks
+ will fit the scale of original image shape. Defaults to True.
+
+ Returns:
+ list[obj:`InstanceData`]: each contains the following keys
+ - labels (Tensor): Prediction class indices\
+ for an image, with shape (n, ), n is the sum of\
+ number of stuff type and number of instance in an image.
+ - masks (Tensor): Prediction mask for a\
+ image, with shape (n, t, h, w).
+ """
+ results = []
+ if len(mask_cls_results) > 0:
+ scores = F.softmax(mask_cls_results, dim=-1)[:, :-1]
+ labels = torch.arange(self.num_classes).unsqueeze(0).repeat(
+ self.num_queries, 1).flatten(0, 1).to(scores.device)
+ # keep top-10 predictions
+ scores_per_image, topk_indices = scores.flatten(0, 1).topk(
+ 10, sorted=False)
+ labels_per_image = labels[topk_indices]
+ topk_indices = topk_indices // self.num_classes
+ mask_pred_results = mask_pred_results[topk_indices]
+
+ img_shape = batch_img_metas[0]['img_shape']
+ mask_pred_results = \
+ mask_pred_results[:, :, :img_shape[0], :img_shape[1]]
+ if rescale:
+ # return result in original resolution
+ ori_height, ori_width = batch_img_metas[0]['ori_shape'][:2]
+ mask_pred_results = F.interpolate(
+ mask_pred_results,
+ size=(ori_height, ori_width),
+ mode='bilinear',
+ align_corners=False)
+
+ masks = mask_pred_results > 0.
+
+ # format top-10 predictions
+ for img_idx in range(len(batch_img_metas)):
+ pred_track_instances = InstanceData()
+
+ pred_track_instances.masks = masks[:, img_idx]
+ pred_track_instances.bboxes = mask2bbox(masks[:, img_idx])
+ pred_track_instances.labels = labels_per_image
+ pred_track_instances.scores = scores_per_image
+ pred_track_instances.instances_id = torch.arange(10)
+
+ results.append(pred_track_instances)
+
+ return results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/tracking_heads/quasi_dense_embed_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/tracking_heads/quasi_dense_embed_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..55e3c05b7aba188608f7dd2fdda54e0759cee03c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/tracking_heads/quasi_dense_embed_head.py
@@ -0,0 +1,347 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Tuple
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from mmengine.model import BaseModule
+from torch import Tensor
+from torch.nn.modules.utils import _pair
+
+from mmdet.models.task_modules import SamplingResult
+from mmdet.registry import MODELS
+from ..task_modules.tracking import embed_similarity
+
+
+@MODELS.register_module()
+class QuasiDenseEmbedHead(BaseModule):
+ """The quasi-dense roi embed head.
+
+ Args:
+ embed_channels (int): The input channel of embed features.
+ Defaults to 256.
+ softmax_temp (int): Softmax temperature. Defaults to -1.
+ loss_track (dict): The loss function for tracking. Defaults to
+ MultiPosCrossEntropyLoss.
+ loss_track_aux (dict): The auxiliary loss function for tracking.
+ Defaults to MarginL2Loss.
+ init_cfg (:obj:`ConfigDict` or dict or list[:obj:`ConfigDict` or \
+ dict]): Initialization config dict.
+ """
+
+ def __init__(self,
+ num_convs: int = 0,
+ num_fcs: int = 0,
+ roi_feat_size: int = 7,
+ in_channels: int = 256,
+ conv_out_channels: int = 256,
+ with_avg_pool: bool = False,
+ fc_out_channels: int = 1024,
+ conv_cfg: Optional[dict] = None,
+ norm_cfg: Optional[dict] = None,
+ embed_channels: int = 256,
+ softmax_temp: int = -1,
+ loss_track: Optional[dict] = None,
+ loss_track_aux: dict = dict(
+ type='MarginL2Loss',
+ sample_ratio=3,
+ margin=0.3,
+ loss_weight=1.0,
+ hard_mining=True),
+ init_cfg: dict = dict(
+ type='Xavier',
+ layer='Linear',
+ distribution='uniform',
+ bias=0,
+ override=dict(
+ type='Normal',
+ name='fc_embed',
+ mean=0,
+ std=0.01,
+ bias=0))):
+ super(QuasiDenseEmbedHead, self).__init__(init_cfg=init_cfg)
+ self.num_convs = num_convs
+ self.num_fcs = num_fcs
+ self.roi_feat_size = _pair(roi_feat_size)
+ self.roi_feat_area = self.roi_feat_size[0] * self.roi_feat_size[1]
+ self.in_channels = in_channels
+ self.conv_out_channels = conv_out_channels
+ self.with_avg_pool = with_avg_pool
+ self.fc_out_channels = fc_out_channels
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+
+ if self.with_avg_pool:
+ self.avg_pool = nn.AvgPool2d(self.roi_feat_size)
+ # add convs and fcs
+ self.convs, self.fcs, self.last_layer_dim = self._add_conv_fc_branch(
+ self.num_convs, self.num_fcs, self.in_channels)
+ self.relu = nn.ReLU(inplace=True)
+
+ if loss_track is None:
+ loss_track = dict(
+ type='MultiPosCrossEntropyLoss', loss_weight=0.25)
+
+ self.fc_embed = nn.Linear(self.last_layer_dim, embed_channels)
+ self.softmax_temp = softmax_temp
+ self.loss_track = MODELS.build(loss_track)
+ if loss_track_aux is not None:
+ self.loss_track_aux = MODELS.build(loss_track_aux)
+ else:
+ self.loss_track_aux = None
+
+ def _add_conv_fc_branch(
+ self, num_branch_convs: int, num_branch_fcs: int,
+ in_channels: int) -> Tuple[nn.ModuleList, nn.ModuleList, int]:
+ """Add shared or separable branch. convs -> avg pool (optional) -> fcs.
+
+ Args:
+ num_branch_convs (int): The number of convoluational layers.
+ num_branch_fcs (int): The number of fully connection layers.
+ in_channels (int): The input channel of roi features.
+
+ Returns:
+ Tuple[nn.ModuleList, nn.ModuleList, int]: The convs, fcs and the
+ last layer dimension.
+ """
+ last_layer_dim = in_channels
+ # add branch specific conv layers
+ branch_convs = nn.ModuleList()
+ if num_branch_convs > 0:
+ for i in range(num_branch_convs):
+ conv_in_channels = (
+ last_layer_dim if i == 0 else self.conv_out_channels)
+ branch_convs.append(
+ ConvModule(
+ conv_in_channels,
+ self.conv_out_channels,
+ 3,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ last_layer_dim = self.conv_out_channels
+
+ # add branch specific fc layers
+ branch_fcs = nn.ModuleList()
+ if num_branch_fcs > 0:
+ if not self.with_avg_pool:
+ last_layer_dim *= self.roi_feat_area
+ for i in range(num_branch_fcs):
+ fc_in_channels = (
+ last_layer_dim if i == 0 else self.fc_out_channels)
+ branch_fcs.append(
+ nn.Linear(fc_in_channels, self.fc_out_channels))
+ last_layer_dim = self.fc_out_channels
+
+ return branch_convs, branch_fcs, last_layer_dim
+
+ def forward(self, x: Tensor) -> Tensor:
+ """Forward function.
+
+ Args:
+ x (Tensor): The input features from ROI head.
+
+ Returns:
+ Tensor: The embedding feature map.
+ """
+
+ if self.num_convs > 0:
+ for conv in self.convs:
+ x = conv(x)
+ x = x.flatten(1)
+ if self.num_fcs > 0:
+ for fc in self.fcs:
+ x = self.relu(fc(x))
+ x = self.fc_embed(x)
+ return x
+
+ def get_targets(
+ self, gt_match_indices: List[Tensor],
+ key_sampling_results: List[SamplingResult],
+ ref_sampling_results: List[SamplingResult]) -> Tuple[List, List]:
+ """Calculate the track targets and track weights for all samples in a
+ batch according to the sampling_results.
+
+ Args:
+ gt_match_indices (list(Tensor)): Mapping from gt_instance_ids to
+ ref_gt_instance_ids of the same tracklet in a pair of images.
+ key_sampling_results (List[obj:SamplingResult]): Assign results of
+ all images in a batch after sampling.
+ ref_sampling_results (List[obj:SamplingResult]): Assign results of
+ all reference images in a batch after sampling.
+
+ Returns:
+ Tuple[list[Tensor]]: Association results.
+ Containing the following list of Tensors:
+
+ - track_targets (list[Tensor]): The mapping instance ids from
+ all positive proposals in the key image to all proposals
+ in the reference image, each tensor in list has
+ shape (len(key_pos_bboxes), len(ref_bboxes)).
+ - track_weights (list[Tensor]): Loss weights for all positive
+ proposals in a batch, each tensor in list has
+ shape (len(key_pos_bboxes),).
+ """
+
+ track_targets = []
+ track_weights = []
+ for _gt_match_indices, key_res, ref_res in zip(gt_match_indices,
+ key_sampling_results,
+ ref_sampling_results):
+ targets = _gt_match_indices.new_zeros(
+ (key_res.pos_bboxes.size(0), ref_res.bboxes.size(0)),
+ dtype=torch.int)
+ _match_indices = _gt_match_indices[key_res.pos_assigned_gt_inds]
+ pos2pos = (_match_indices.view(
+ -1, 1) == ref_res.pos_assigned_gt_inds.view(1, -1)).int()
+ targets[:, :pos2pos.size(1)] = pos2pos
+ weights = (targets.sum(dim=1) > 0).float()
+ track_targets.append(targets)
+ track_weights.append(weights)
+ return track_targets, track_weights
+
+ def match(
+ self, key_embeds: Tensor, ref_embeds: Tensor,
+ key_sampling_results: List[SamplingResult],
+ ref_sampling_results: List[SamplingResult]
+ ) -> Tuple[List[Tensor], List[Tensor]]:
+ """Calculate the dist matrixes for loss measurement.
+
+ Args:
+ key_embeds (Tensor): Embeds of positive bboxes in sampling results
+ of key image.
+ ref_embeds (Tensor): Embeds of all bboxes in sampling results
+ of the reference image.
+ key_sampling_results (List[obj:SamplingResults]): Assign results of
+ all images in a batch after sampling.
+ ref_sampling_results (List[obj:SamplingResults]): Assign results of
+ all reference images in a batch after sampling.
+
+ Returns:
+ Tuple[list[Tensor]]: Calculation results.
+ Containing the following list of Tensors:
+
+ - dists (list[Tensor]): Dot-product dists between
+ key_embeds and ref_embeds, each tensor in list has
+ shape (len(key_pos_bboxes), len(ref_bboxes)).
+ - cos_dists (list[Tensor]): Cosine dists between
+ key_embeds and ref_embeds, each tensor in list has
+ shape (len(key_pos_bboxes), len(ref_bboxes)).
+ """
+
+ num_key_rois = [res.pos_bboxes.size(0) for res in key_sampling_results]
+ key_embeds = torch.split(key_embeds, num_key_rois)
+ num_ref_rois = [res.bboxes.size(0) for res in ref_sampling_results]
+ ref_embeds = torch.split(ref_embeds, num_ref_rois)
+
+ dists, cos_dists = [], []
+ for key_embed, ref_embed in zip(key_embeds, ref_embeds):
+ dist = embed_similarity(
+ key_embed,
+ ref_embed,
+ method='dot_product',
+ temperature=self.softmax_temp)
+ dists.append(dist)
+ if self.loss_track_aux is not None:
+ cos_dist = embed_similarity(
+ key_embed, ref_embed, method='cosine')
+ cos_dists.append(cos_dist)
+ else:
+ cos_dists.append(None)
+ return dists, cos_dists
+
+ def loss(self, key_roi_feats: Tensor, ref_roi_feats: Tensor,
+ key_sampling_results: List[SamplingResult],
+ ref_sampling_results: List[SamplingResult],
+ gt_match_indices_list: List[Tensor]) -> dict:
+ """Calculate the track loss and the auxiliary track loss.
+
+ Args:
+ key_roi_feats (Tensor): Embeds of positive bboxes in sampling
+ results of key image.
+ ref_roi_feats (Tensor): Embeds of all bboxes in sampling results
+ of the reference image.
+ key_sampling_results (List[obj:SamplingResults]): Assign results of
+ all images in a batch after sampling.
+ ref_sampling_results (List[obj:SamplingResults]): Assign results of
+ all reference images in a batch after sampling.
+ gt_match_indices_list (list(Tensor)): Mapping from gt_instances_ids
+ to ref_gt_instances_ids of the same tracklet in a pair of
+ images.
+
+ Returns:
+ Dict [str: Tensor]: Calculation results.
+ Containing the following list of Tensors:
+
+ - loss_track (Tensor): Results of loss_track function.
+ - loss_track_aux (Tensor): Results of loss_track_aux function.
+ """
+ key_track_feats = self(key_roi_feats)
+ ref_track_feats = self(ref_roi_feats)
+
+ losses = self.loss_by_feat(key_track_feats, ref_track_feats,
+ key_sampling_results, ref_sampling_results,
+ gt_match_indices_list)
+ return losses
+
+ def loss_by_feat(self, key_track_feats: Tensor, ref_track_feats: Tensor,
+ key_sampling_results: List[SamplingResult],
+ ref_sampling_results: List[SamplingResult],
+ gt_match_indices_list: List[Tensor]) -> dict:
+ """Calculate the track loss and the auxiliary track loss.
+
+ Args:
+ key_track_feats (Tensor): Embeds of positive bboxes in sampling
+ results of key image.
+ ref_track_feats (Tensor): Embeds of all bboxes in sampling results
+ of the reference image.
+ key_sampling_results (List[obj:SamplingResults]): Assign results of
+ all images in a batch after sampling.
+ ref_sampling_results (List[obj:SamplingResults]): Assign results of
+ all reference images in a batch after sampling.
+ gt_match_indices_list (list(Tensor)): Mapping from instances_ids
+ from key image to reference image of the same tracklet in a
+ pair of images.
+
+ Returns:
+ Dict [str: Tensor]: Calculation results.
+ Containing the following list of Tensors:
+
+ - loss_track (Tensor): Results of loss_track function.
+ - loss_track_aux (Tensor): Results of loss_track_aux function.
+ """
+ dists, cos_dists = self.match(key_track_feats, ref_track_feats,
+ key_sampling_results,
+ ref_sampling_results)
+ targets, weights = self.get_targets(gt_match_indices_list,
+ key_sampling_results,
+ ref_sampling_results)
+ losses = dict()
+
+ loss_track = 0.
+ loss_track_aux = 0.
+ for _dists, _cos_dists, _targets, _weights in zip(
+ dists, cos_dists, targets, weights):
+ loss_track += self.loss_track(
+ _dists, _targets, _weights, avg_factor=_weights.sum())
+ if self.loss_track_aux is not None:
+ loss_track_aux += self.loss_track_aux(_cos_dists, _targets)
+ losses['loss_track'] = loss_track / len(dists)
+
+ if self.loss_track_aux is not None:
+ losses['loss_track_aux'] = loss_track_aux / len(dists)
+
+ return losses
+
+ def predict(self, bbox_feats: Tensor) -> Tensor:
+ """Perform forward propagation of the tracking head and predict
+ tracking results on the features of the upstream network.
+
+ Args:
+ bbox_feats: The extracted roi features.
+
+ Returns:
+ Tensor: The extracted track features.
+ """
+ track_feats = self(bbox_feats)
+ return track_feats
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/tracking_heads/quasi_dense_track_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/tracking_heads/quasi_dense_track_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..bd078dac827e35c7514330870cf884001985156b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/tracking_heads/quasi_dense_track_head.py
@@ -0,0 +1,178 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional
+
+from mmengine.model import BaseModule
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures import TrackSampleList
+from mmdet.structures.bbox import bbox2roi
+from mmdet.utils import InstanceList
+
+
+@MODELS.register_module()
+class QuasiDenseTrackHead(BaseModule):
+ """The quasi-dense track head."""
+
+ def __init__(self,
+ roi_extractor: Optional[dict] = None,
+ embed_head: Optional[dict] = None,
+ regress_head: Optional[dict] = None,
+ train_cfg: Optional[dict] = None,
+ test_cfg: Optional[dict] = None,
+ init_cfg: Optional[dict] = None,
+ **kwargs):
+ super().__init__(init_cfg=init_cfg)
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+
+ if embed_head is not None:
+ self.init_embed_head(roi_extractor, embed_head)
+
+ if regress_head is not None:
+ raise NotImplementedError('Regression head is not supported yet.')
+
+ self.init_assigner_sampler()
+
+ def init_embed_head(self, roi_extractor, embed_head) -> None:
+ """Initialize ``embed_head``
+
+ Args:
+ roi_extractor (dict, optional): Configuration of roi extractor.
+ Defaults to None.
+ embed_head (dict, optional): Configuration of embed head. Defaults
+ to None.
+ """
+ self.roi_extractor = MODELS.build(roi_extractor)
+ self.embed_head = MODELS.build(embed_head)
+
+ def init_assigner_sampler(self) -> None:
+ """Initialize assigner and sampler."""
+ self.bbox_assigner = None
+ self.bbox_sampler = None
+ if self.train_cfg:
+ self.bbox_assigner = TASK_UTILS.build(self.train_cfg.assigner)
+ self.bbox_sampler = TASK_UTILS.build(
+ self.train_cfg.sampler, default_args=dict(context=self))
+
+ @property
+ def with_track(self) -> bool:
+ """bool: whether the multi-object tracker has an embed head"""
+ return hasattr(self, 'embed_head') and self.embed_head is not None
+
+ def extract_roi_feats(self, feats: List[Tensor],
+ bboxes: List[Tensor]) -> Tensor:
+ """Extract roi features.
+
+ Args:
+ feats (list[Tensor]): list of multi-level image features.
+ bboxes (list[Tensor]): list of bboxes in sampling result.
+
+ Returns:
+ Tensor: The extracted roi features.
+ """
+ rois = bbox2roi(bboxes)
+ bbox_feats = self.roi_extractor(feats[:self.roi_extractor.num_inputs],
+ rois)
+ return bbox_feats
+
+ def loss(self, key_feats: List[Tensor], ref_feats: List[Tensor],
+ rpn_results_list: InstanceList,
+ ref_rpn_results_list: InstanceList, data_samples: TrackSampleList,
+ **kwargs) -> dict:
+ """Calculate losses from a batch of inputs and data samples.
+
+ Args:
+ key_feats (list[Tensor]): list of multi-level image features.
+ ref_feats (list[Tensor]): list of multi-level ref_img features.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals of key img.
+ ref_rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals of ref img.
+ data_samples (list[:obj:`TrackDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance`.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ assert self.with_track
+ num_imgs = len(data_samples)
+ batch_gt_instances = []
+ ref_batch_gt_instances = []
+ batch_gt_instances_ignore = []
+ gt_match_indices_list = []
+ for track_data_sample in data_samples:
+ key_data_sample = track_data_sample.get_key_frames()[0]
+ ref_data_sample = track_data_sample.get_ref_frames()[0]
+ batch_gt_instances.append(key_data_sample.gt_instances)
+ ref_batch_gt_instances.append(ref_data_sample.gt_instances)
+ if 'ignored_instances' in key_data_sample:
+ batch_gt_instances_ignore.append(
+ key_data_sample.ignored_instances)
+ else:
+ batch_gt_instances_ignore.append(None)
+ # get gt_match_indices
+ ins_ids = key_data_sample.gt_instances.instances_ids.tolist()
+ ref_ins_ids = ref_data_sample.gt_instances.instances_ids.tolist()
+ match_indices = Tensor([
+ ref_ins_ids.index(i) if (i in ref_ins_ids and i > 0) else -1
+ for i in ins_ids
+ ]).to(key_feats[0].device)
+ gt_match_indices_list.append(match_indices)
+
+ key_sampling_results, ref_sampling_results = [], []
+ for i in range(num_imgs):
+ rpn_results = rpn_results_list[i]
+ ref_rpn_results = ref_rpn_results_list[i]
+ # rename ref_rpn_results.bboxes to ref_rpn_results.priors
+ ref_rpn_results.priors = ref_rpn_results.pop('bboxes')
+
+ assign_result = self.bbox_assigner.assign(
+ rpn_results, batch_gt_instances[i],
+ batch_gt_instances_ignore[i])
+ sampling_result = self.bbox_sampler.sample(
+ assign_result,
+ rpn_results,
+ batch_gt_instances[i],
+ feats=[lvl_feat[i][None] for lvl_feat in key_feats])
+ key_sampling_results.append(sampling_result)
+
+ ref_assign_result = self.bbox_assigner.assign(
+ ref_rpn_results, ref_batch_gt_instances[i],
+ batch_gt_instances_ignore[i])
+ ref_sampling_result = self.bbox_sampler.sample(
+ ref_assign_result,
+ ref_rpn_results,
+ ref_batch_gt_instances[i],
+ feats=[lvl_feat[i][None] for lvl_feat in ref_feats])
+ ref_sampling_results.append(ref_sampling_result)
+
+ key_bboxes = [res.pos_bboxes for res in key_sampling_results]
+ key_roi_feats = self.extract_roi_feats(key_feats, key_bboxes)
+ ref_bboxes = [res.bboxes for res in ref_sampling_results]
+ ref_roi_feats = self.extract_roi_feats(ref_feats, ref_bboxes)
+
+ loss_track = self.embed_head.loss(key_roi_feats, ref_roi_feats,
+ key_sampling_results,
+ ref_sampling_results,
+ gt_match_indices_list)
+
+ return loss_track
+
+ def predict(self, feats: List[Tensor],
+ rescaled_bboxes: List[Tensor]) -> Tensor:
+ """Perform forward propagation of the tracking head and predict
+ tracking results on the features of the upstream network.
+
+ Args:
+ feats (list[Tensor]): Multi level feature maps of `img`.
+ rescaled_bboxes (list[Tensor]): list of rescaled bboxes in sampling
+ result.
+
+ Returns:
+ Tensor: The extracted track features.
+ """
+ bbox_feats = self.extract_roi_feats(feats, rescaled_bboxes)
+ track_feats = self.embed_head.predict(bbox_feats)
+ return track_feats
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/tracking_heads/roi_embed_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/tracking_heads/roi_embed_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..e18b81fbe52e109e7afb3e6d5e8e6624ef48242f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/tracking_heads/roi_embed_head.py
@@ -0,0 +1,391 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from collections import defaultdict
+from typing import List, Optional, Tuple
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import ConvModule
+from mmengine.model import BaseModule
+from torch import Tensor
+from torch.nn.modules.utils import _pair
+
+from mmdet.models.losses import accuracy
+from mmdet.models.task_modules import SamplingResult
+from mmdet.models.task_modules.tracking import embed_similarity
+from mmdet.registry import MODELS
+
+
+@MODELS.register_module()
+class RoIEmbedHead(BaseModule):
+ """The roi embed head.
+
+ This module is used in multi-object tracking methods, such as MaskTrack
+ R-CNN.
+
+ Args:
+ num_convs (int): The number of convoluational layers to embed roi
+ features. Defaults to 0.
+ num_fcs (int): The number of fully connection layers to embed roi
+ features. Defaults to 0.
+ roi_feat_size (int|tuple(int)): The spatial size of roi features.
+ Defaults to 7.
+ in_channels (int): The input channel of roi features. Defaults to 256.
+ conv_out_channels (int): The output channel of roi features after
+ forwarding convoluational layers. Defaults to 256.
+ with_avg_pool (bool): Whether use average pooling before passing roi
+ features into fully connection layers. Defaults to False.
+ fc_out_channels (int): The output channel of roi features after
+ forwarding fully connection layers. Defaults to 1024.
+ conv_cfg (dict): Config dict for convolution layer. Defaults to None,
+ which means using conv2d.
+ norm_cfg (dict): Config dict for normalization layer. Defaults to None.
+ loss_match (dict): The loss function. Defaults to
+ dict(type='CrossEntropyLoss', use_sigmoid=False, loss_weight=1.0)
+ init_cfg (dict): Configuration of initialization. Defaults to None.
+ """
+
+ def __init__(self,
+ num_convs: int = 0,
+ num_fcs: int = 0,
+ roi_feat_size: int = 7,
+ in_channels: int = 256,
+ conv_out_channels: int = 256,
+ with_avg_pool: bool = False,
+ fc_out_channels: int = 1024,
+ conv_cfg: Optional[dict] = None,
+ norm_cfg: Optional[dict] = None,
+ loss_match: dict = dict(
+ type='mmdet.CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0),
+ init_cfg: Optional[dict] = None,
+ **kwargs):
+ super(RoIEmbedHead, self).__init__(init_cfg=init_cfg)
+ self.num_convs = num_convs
+ self.num_fcs = num_fcs
+ self.roi_feat_size = _pair(roi_feat_size)
+ self.roi_feat_area = self.roi_feat_size[0] * self.roi_feat_size[1]
+ self.in_channels = in_channels
+ self.conv_out_channels = conv_out_channels
+ self.with_avg_pool = with_avg_pool
+ self.fc_out_channels = fc_out_channels
+ self.conv_cfg = conv_cfg
+ self.norm_cfg = norm_cfg
+ self.loss_match = MODELS.build(loss_match)
+ self.fp16_enabled = False
+
+ if self.with_avg_pool:
+ self.avg_pool = nn.AvgPool2d(self.roi_feat_size)
+ # add convs and fcs
+ self.convs, self.fcs, self.last_layer_dim = self._add_conv_fc_branch(
+ self.num_convs, self.num_fcs, self.in_channels)
+ self.relu = nn.ReLU(inplace=True)
+
+ def _add_conv_fc_branch(
+ self, num_branch_convs: int, num_branch_fcs: int,
+ in_channels: int) -> Tuple[nn.ModuleList, nn.ModuleList, int]:
+ """Add shared or separable branch.
+
+ convs -> avg pool (optional) -> fcs
+ """
+ last_layer_dim = in_channels
+ # add branch specific conv layers
+ branch_convs = nn.ModuleList()
+ if num_branch_convs > 0:
+ for i in range(num_branch_convs):
+ conv_in_channels = (
+ last_layer_dim if i == 0 else self.conv_out_channels)
+ branch_convs.append(
+ ConvModule(
+ conv_in_channels,
+ self.conv_out_channels,
+ 3,
+ padding=1,
+ conv_cfg=self.conv_cfg,
+ norm_cfg=self.norm_cfg))
+ last_layer_dim = self.conv_out_channels
+
+ # add branch specific fc layers
+ branch_fcs = nn.ModuleList()
+ if num_branch_fcs > 0:
+ if not self.with_avg_pool:
+ last_layer_dim *= self.roi_feat_area
+ for i in range(num_branch_fcs):
+ fc_in_channels = (
+ last_layer_dim if i == 0 else self.fc_out_channels)
+ branch_fcs.append(
+ nn.Linear(fc_in_channels, self.fc_out_channels))
+ last_layer_dim = self.fc_out_channels
+
+ return branch_convs, branch_fcs, last_layer_dim
+
+ @property
+ def custom_activation(self):
+ return getattr(self.loss_match, 'custom_activation', False)
+
+ def extract_feat(self, x: Tensor,
+ num_x_per_img: List[int]) -> Tuple[Tensor]:
+ """Extract feature from the input `x`, and split the output to a list.
+
+ Args:
+ x (Tensor): of shape [N, C, H, W]. N is the number of proposals.
+ num_x_per_img (list[int]): The `x` contains proposals of
+ multi-images. `num_x_per_img` denotes the number of proposals
+ for each image.
+
+ Returns:
+ list[Tensor]: Each Tensor denotes the embed features belonging to
+ an image in a batch.
+ """
+ if self.num_convs > 0:
+ for conv in self.convs:
+ x = conv(x)
+
+ if self.num_fcs > 0:
+ if self.with_avg_pool:
+ x = self.avg_pool(x)
+ x = x.flatten(1)
+ for fc in self.fcs:
+ x = self.relu(fc(x))
+ else:
+ x = x.flatten(1)
+
+ x_split = torch.split(x, num_x_per_img, dim=0)
+ return x_split
+
+ def forward(
+ self, x: Tensor, ref_x: Tensor, num_x_per_img: List[int],
+ num_x_per_ref_img: List[int]
+ ) -> Tuple[Tuple[Tensor], Tuple[Tensor]]:
+ """Computing the similarity scores between `x` and `ref_x`.
+
+ Args:
+ x (Tensor): of shape [N, C, H, W]. N is the number of key frame
+ proposals.
+ ref_x (Tensor): of shape [M, C, H, W]. M is the number of reference
+ frame proposals.
+ num_x_per_img (list[int]): The `x` contains proposals of
+ multi-images. `num_x_per_img` denotes the number of proposals
+ for each key image.
+ num_x_per_ref_img (list[int]): The `ref_x` contains proposals of
+ multi-images. `num_x_per_ref_img` denotes the number of
+ proposals for each reference image.
+
+ Returns:
+ tuple[tuple[Tensor], tuple[Tensor]]: Each tuple of tensor denotes
+ the embed features belonging to an image in a batch.
+ """
+ x_split = self.extract_feat(x, num_x_per_img)
+ ref_x_split = self.extract_feat(ref_x, num_x_per_ref_img)
+
+ return x_split, ref_x_split
+
+ def get_targets(self, sampling_results: List[SamplingResult],
+ gt_instance_ids: List[Tensor],
+ ref_gt_instance_ids: List[Tensor]) -> Tuple[List, List]:
+ """Calculate the ground truth for all samples in a batch according to
+ the sampling_results.
+
+ Args:
+ sampling_results (List[obj:SamplingResult]): Assign results of
+ all images in a batch after sampling.
+ gt_instance_ids (list[Tensor]): The instance ids of gt_bboxes of
+ all images in a batch, each tensor has shape (num_gt, ).
+ ref_gt_instance_ids (list[Tensor]): The instance ids of gt_bboxes
+ of all reference images in a batch, each tensor has shape
+ (num_gt, ).
+
+ Returns:
+ Tuple[list[Tensor]]: Ground truth for proposals in a batch.
+ Containing the following list of Tensors:
+
+ - track_id_targets (list[Tensor]): The instance ids of
+ Gt_labels for all proposals in a batch, each tensor in list
+ has shape (num_proposals,).
+ - track_id_weights (list[Tensor]): Labels_weights for
+ all proposals in a batch, each tensor in list has
+ shape (num_proposals,).
+ """
+ track_id_targets = []
+ track_id_weights = []
+
+ for res, gt_instance_id, ref_gt_instance_id in zip(
+ sampling_results, gt_instance_ids, ref_gt_instance_ids):
+ pos_instance_ids = gt_instance_id[res.pos_assigned_gt_inds]
+ pos_match_id = gt_instance_id.new_zeros(len(pos_instance_ids))
+ for i, id in enumerate(pos_instance_ids):
+ if id in ref_gt_instance_id:
+ pos_match_id[i] = ref_gt_instance_id.tolist().index(id) + 1
+
+ track_id_target = gt_instance_id.new_zeros(
+ len(res.bboxes), dtype=torch.int64)
+ track_id_target[:len(res.pos_bboxes)] = pos_match_id
+ track_id_weight = res.bboxes.new_zeros(len(res.bboxes))
+ track_id_weight[:len(res.pos_bboxes)] = 1.0
+
+ track_id_targets.append(track_id_target)
+ track_id_weights.append(track_id_weight)
+
+ return track_id_targets, track_id_weights
+
+ def loss(
+ self,
+ bbox_feats: Tensor,
+ ref_bbox_feats: Tensor,
+ num_bbox_per_img: int,
+ num_bbox_per_ref_img: int,
+ sampling_results: List[SamplingResult],
+ gt_instance_ids: List[Tensor],
+ ref_gt_instance_ids: List[Tensor],
+ reduction_override: Optional[str] = None,
+ ) -> dict:
+ """Calculate the loss in a batch.
+
+ Args:
+ bbox_feats (Tensor): of shape [N, C, H, W]. N is the number of
+ bboxes.
+ ref_bbox_feats (Tensor): of shape [M, C, H, W]. M is the number of
+ reference bboxes.
+ num_bbox_per_img (list[int]): The `bbox_feats` contains proposals
+ of multi-images. `num_bbox_per_img` denotes the number of
+ proposals for each key image.
+ num_bbox_per_ref_img (list[int]): The `ref_bbox_feats` contains
+ proposals of multi-images. `num_bbox_per_ref_img` denotes the
+ number of proposals for each reference image.
+ sampling_results (List[obj:SamplingResult]): Assign results of
+ all images in a batch after sampling.
+ gt_instance_ids (list[Tensor]): The instance ids of gt_bboxes of
+ all images in a batch, each tensor has shape (num_gt, ).
+ ref_gt_instance_ids (list[Tensor]): The instance ids of gt_bboxes
+ of all reference images in a batch, each tensor has shape
+ (num_gt, ).
+ reduction_override (str, optional): The method used to reduce the
+ loss. Options are "none", "mean" and "sum".
+
+ Returns:
+ dict[str, Tensor]: a dictionary of loss components.
+ """
+ x_split, ref_x_split = self(bbox_feats, ref_bbox_feats,
+ num_bbox_per_img, num_bbox_per_ref_img)
+
+ losses = self.loss_by_feat(x_split, ref_x_split, sampling_results,
+ gt_instance_ids, ref_gt_instance_ids,
+ reduction_override)
+ return losses
+
+ def loss_by_feat(self,
+ x_split: Tuple[Tensor],
+ ref_x_split: Tuple[Tensor],
+ sampling_results: List[SamplingResult],
+ gt_instance_ids: List[Tensor],
+ ref_gt_instance_ids: List[Tensor],
+ reduction_override: Optional[str] = None) -> dict:
+ """Calculate losses.
+
+ Args:
+ x_split (Tensor): The embed features belonging to key image.
+ ref_x_split (Tensor): The embed features belonging to ref image.
+ sampling_results (List[obj:SamplingResult]): Assign results of
+ all images in a batch after sampling.
+ gt_instance_ids (list[Tensor]): The instance ids of gt_bboxes of
+ all images in a batch, each tensor has shape (num_gt, ).
+ ref_gt_instance_ids (list[Tensor]): The instance ids of gt_bboxes
+ of all reference images in a batch, each tensor has shape
+ (num_gt, ).
+ reduction_override (str, optional): The method used to reduce the
+ loss. Options are "none", "mean" and "sum".
+
+ Returns:
+ dict[str, Tensor]: a dictionary of loss components.
+ """
+ track_id_targets, track_id_weights = self.get_targets(
+ sampling_results, gt_instance_ids, ref_gt_instance_ids)
+ assert isinstance(track_id_targets, list)
+ assert isinstance(track_id_weights, list)
+ assert len(track_id_weights) == len(track_id_targets)
+
+ losses = defaultdict(list)
+ similarity_logits = []
+ for one_x, one_ref_x in zip(x_split, ref_x_split):
+ similarity_logit = embed_similarity(
+ one_x, one_ref_x, method='dot_product')
+ dummy = similarity_logit.new_zeros(one_x.shape[0], 1)
+ similarity_logit = torch.cat((dummy, similarity_logit), dim=1)
+ similarity_logits.append(similarity_logit)
+ assert isinstance(similarity_logits, list)
+ assert len(similarity_logits) == len(track_id_targets)
+
+ for similarity_logit, track_id_target, track_id_weight in zip(
+ similarity_logits, track_id_targets, track_id_weights):
+ avg_factor = max(torch.sum(track_id_target > 0).float().item(), 1.)
+ if similarity_logit.numel() > 0:
+ loss_match = self.loss_match(
+ similarity_logit,
+ track_id_target,
+ track_id_weight,
+ avg_factor=avg_factor,
+ reduction_override=reduction_override)
+ if isinstance(loss_match, dict):
+ for key, value in loss_match.items():
+ losses[key].append(value)
+ else:
+ losses['loss_match'].append(loss_match)
+
+ valid_index = track_id_weight > 0
+ valid_similarity_logit = similarity_logit[valid_index]
+ valid_track_id_target = track_id_target[valid_index]
+ if self.custom_activation:
+ match_accuracy = self.loss_match.get_accuracy(
+ valid_similarity_logit, valid_track_id_target)
+ for key, value in match_accuracy.items():
+ losses[key].append(value)
+ else:
+ losses['match_accuracy'].append(
+ accuracy(valid_similarity_logit,
+ valid_track_id_target))
+
+ for key, value in losses.items():
+ losses[key] = sum(losses[key]) / len(similarity_logits)
+ return losses
+
+ def predict(self, roi_feats: Tensor,
+ prev_roi_feats: Tensor) -> List[Tensor]:
+ """Perform forward propagation of the tracking head and predict
+ tracking results on the features of the upstream network.
+
+ Args:
+ roi_feats (Tensor): Feature map of current images rois.
+ prev_roi_feats (Tensor): Feature map of previous images rois.
+
+ Returns:
+ list[Tensor]: The predicted similarity_logits of each pair of key
+ image and reference image.
+ """
+ x_split, ref_x_split = self(roi_feats, prev_roi_feats,
+ [roi_feats.shape[0]],
+ [prev_roi_feats.shape[0]])
+
+ similarity_logits = self.predict_by_feat(x_split, ref_x_split)
+
+ return similarity_logits
+
+ def predict_by_feat(self, x_split: Tuple[Tensor],
+ ref_x_split: Tuple[Tensor]) -> List[Tensor]:
+ """Get similarity_logits.
+
+ Args:
+ x_split (Tensor): The embed features belonging to key image.
+ ref_x_split (Tensor): The embed features belonging to ref image.
+
+ Returns:
+ list[Tensor]: The predicted similarity_logits of each pair of key
+ image and reference image.
+ """
+ similarity_logits = []
+ for one_x, one_ref_x in zip(x_split, ref_x_split):
+ similarity_logit = embed_similarity(
+ one_x, one_ref_x, method='dot_product')
+ dummy = similarity_logit.new_zeros(one_x.shape[0], 1)
+ similarity_logit = torch.cat((dummy, similarity_logit), dim=1)
+ similarity_logits.append(similarity_logit)
+ return similarity_logits
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/tracking_heads/roi_track_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/tracking_heads/roi_track_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..c51c810022cc856411e1de83278e38fdc2b670c8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/tracking_heads/roi_track_head.py
@@ -0,0 +1,178 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from abc import ABCMeta
+from typing import List, Optional, Tuple
+
+from mmengine.model import BaseModule
+from torch import Tensor
+
+from mmdet.registry import MODELS, TASK_UTILS
+from mmdet.structures import TrackSampleList
+from mmdet.structures.bbox import bbox2roi
+from mmdet.utils import InstanceList
+
+
+@MODELS.register_module()
+class RoITrackHead(BaseModule, metaclass=ABCMeta):
+ """The roi track head.
+
+ This module is used in multi-object tracking methods, such as MaskTrack
+ R-CNN.
+
+ Args:
+ roi_extractor (dict): Configuration of roi extractor. Defaults to None.
+ embed_head (dict): Configuration of embed head. Defaults to None.
+ train_cfg (dict): Configuration when training. Defaults to None.
+ test_cfg (dict): Configuration when testing. Defaults to None.
+ init_cfg (dict): Configuration of initialization. Defaults to None.
+ """
+
+ def __init__(self,
+ roi_extractor: Optional[dict] = None,
+ embed_head: Optional[dict] = None,
+ regress_head: Optional[dict] = None,
+ train_cfg: Optional[dict] = None,
+ test_cfg: Optional[dict] = None,
+ init_cfg: Optional[dict] = None,
+ *args,
+ **kwargs):
+ super().__init__(init_cfg=init_cfg)
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+
+ if embed_head is not None:
+ self.init_embed_head(roi_extractor, embed_head)
+
+ if regress_head is not None:
+ raise NotImplementedError('Regression head is not supported yet.')
+
+ self.init_assigner_sampler()
+
+ def init_embed_head(self, roi_extractor, embed_head) -> None:
+ """Initialize ``embed_head``"""
+ self.roi_extractor = MODELS.build(roi_extractor)
+ self.embed_head = MODELS.build(embed_head)
+
+ def init_assigner_sampler(self) -> None:
+ """Initialize assigner and sampler."""
+ self.bbox_assigner = None
+ self.bbox_sampler = None
+ if self.train_cfg:
+ self.bbox_assigner = TASK_UTILS.build(self.train_cfg.assigner)
+ self.bbox_sampler = TASK_UTILS.build(
+ self.train_cfg.sampler, default_args=dict(context=self))
+
+ @property
+ def with_track(self) -> bool:
+ """bool: whether the multi-object tracker has an embed head"""
+ return hasattr(self, 'embed_head') and self.embed_head is not None
+
+ def extract_roi_feats(
+ self, feats: List[Tensor],
+ bboxes: List[Tensor]) -> Tuple[Tuple[Tensor], List[int]]:
+ """Extract roi features.
+
+ Args:
+ feats (list[Tensor]): list of multi-level image features.
+ bboxes (list[Tensor]): list of bboxes in sampling result.
+
+ Returns:
+ tuple[tuple[Tensor], list[int]]: The extracted roi features and
+ the number of bboxes in each image.
+ """
+ rois = bbox2roi(bboxes)
+ bbox_feats = self.roi_extractor(feats[:self.roi_extractor.num_inputs],
+ rois)
+ num_bbox_per_img = [len(bbox) for bbox in bboxes]
+ return bbox_feats, num_bbox_per_img
+
+ def loss(self, key_feats: List[Tensor], ref_feats: List[Tensor],
+ rpn_results_list: InstanceList, data_samples: TrackSampleList,
+ **kwargs) -> dict:
+ """Calculate losses from a batch of inputs and data samples.
+
+ Args:
+ key_feats (list[Tensor]): list of multi-level image features.
+ ref_feats (list[Tensor]): list of multi-level ref_img features.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ data_samples (list[:obj:`TrackDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance`.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+ assert self.with_track
+ batch_gt_instances = []
+ ref_batch_gt_instances = []
+ batch_gt_instances_ignore = []
+ gt_instance_ids = []
+ ref_gt_instance_ids = []
+ for track_data_sample in data_samples:
+ key_data_sample = track_data_sample.get_key_frames()[0]
+ ref_data_sample = track_data_sample.get_ref_frames()[0]
+ batch_gt_instances.append(key_data_sample.gt_instances)
+ ref_batch_gt_instances.append(ref_data_sample.gt_instances)
+ if 'ignored_instances' in key_data_sample:
+ batch_gt_instances_ignore.append(
+ key_data_sample.ignored_instances)
+ else:
+ batch_gt_instances_ignore.append(None)
+
+ gt_instance_ids.append(key_data_sample.gt_instances.instances_ids)
+ ref_gt_instance_ids.append(
+ ref_data_sample.gt_instances.instances_ids)
+
+ losses = dict()
+ num_imgs = len(data_samples)
+ if batch_gt_instances_ignore is None:
+ batch_gt_instances_ignore = [None] * num_imgs
+ sampling_results = []
+ for i in range(num_imgs):
+ rpn_results = rpn_results_list[i]
+
+ assign_result = self.bbox_assigner.assign(
+ rpn_results, batch_gt_instances[i],
+ batch_gt_instances_ignore[i])
+ sampling_result = self.bbox_sampler.sample(
+ assign_result,
+ rpn_results,
+ batch_gt_instances[i],
+ feats=[lvl_feat[i][None] for lvl_feat in key_feats])
+ sampling_results.append(sampling_result)
+
+ bboxes = [res.bboxes for res in sampling_results]
+ bbox_feats, num_bbox_per_img = self.extract_roi_feats(
+ key_feats, bboxes)
+
+ # batch_size is 1
+ ref_gt_bboxes = [
+ ref_batch_gt_instance.bboxes
+ for ref_batch_gt_instance in ref_batch_gt_instances
+ ]
+ ref_bbox_feats, num_bbox_per_ref_img = self.extract_roi_feats(
+ ref_feats, ref_gt_bboxes)
+
+ loss_track = self.embed_head.loss(bbox_feats, ref_bbox_feats,
+ num_bbox_per_img,
+ num_bbox_per_ref_img,
+ sampling_results, gt_instance_ids,
+ ref_gt_instance_ids)
+ losses.update(loss_track)
+
+ return losses
+
+ def predict(self, roi_feats: Tensor,
+ prev_roi_feats: Tensor) -> List[Tensor]:
+ """Perform forward propagation of the tracking head and predict
+ tracking results on the features of the upstream network.
+
+ Args:
+ roi_feats (Tensor): Feature map of current images rois.
+ prev_roi_feats (Tensor): Feature map of previous images rois.
+
+ Returns:
+ list[Tensor]: The predicted similarity_logits of each pair of key
+ image and reference image.
+ """
+ return self.embed_head.predict(roi_feats, prev_roi_feats)[0]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..a00d9a37f33169dc1c523c68db55f823dd0424fa
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/__init__.py
@@ -0,0 +1,37 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .gaussian_target import (gather_feat, gaussian_radius,
+ gen_gaussian_target, get_local_maximum,
+ get_topk_from_heatmap, transpose_and_gather_feat)
+from .image import imrenormalize
+from .make_divisible import make_divisible
+# Disable yapf because it conflicts with isort.
+# yapf: disable
+from .misc import (align_tensor, aligned_bilinear, center_of_mass,
+ empty_instances, filter_gt_instances,
+ filter_scores_and_topk, flip_tensor, generate_coordinate,
+ images_to_levels, interpolate_as, levels_to_images,
+ mask2ndarray, multi_apply, relative_coordinate_maps,
+ rename_loss_dict, reweight_loss_dict,
+ samplelist_boxtype2tensor, select_single_mlvl,
+ sigmoid_geometric_mean, unfold_wo_center, unmap,
+ unpack_gt_instances)
+from .panoptic_gt_processing import preprocess_panoptic_gt
+from .point_sample import (get_uncertain_point_coords_with_randomness,
+ get_uncertainty)
+from .vlfuse_helper import BertEncoderLayer, VLFuse, permute_and_flatten
+from .wbf import weighted_boxes_fusion
+
+__all__ = [
+ 'gaussian_radius', 'gen_gaussian_target', 'make_divisible',
+ 'get_local_maximum', 'get_topk_from_heatmap', 'transpose_and_gather_feat',
+ 'interpolate_as', 'sigmoid_geometric_mean', 'gather_feat',
+ 'preprocess_panoptic_gt', 'get_uncertain_point_coords_with_randomness',
+ 'get_uncertainty', 'unpack_gt_instances', 'empty_instances',
+ 'center_of_mass', 'filter_scores_and_topk', 'flip_tensor',
+ 'generate_coordinate', 'levels_to_images', 'mask2ndarray', 'multi_apply',
+ 'select_single_mlvl', 'unmap', 'images_to_levels',
+ 'samplelist_boxtype2tensor', 'filter_gt_instances', 'rename_loss_dict',
+ 'reweight_loss_dict', 'relative_coordinate_maps', 'aligned_bilinear',
+ 'unfold_wo_center', 'imrenormalize', 'VLFuse', 'permute_and_flatten',
+ 'BertEncoderLayer', 'align_tensor', 'weighted_boxes_fusion'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/gaussian_target.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/gaussian_target.py
new file mode 100644
index 0000000000000000000000000000000000000000..5bf4d558ce05c4f953e1c3fcf75016e5874afce1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/gaussian_target.py
@@ -0,0 +1,268 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from math import sqrt
+
+import torch
+import torch.nn.functional as F
+
+
+def gaussian2D(radius, sigma=1, dtype=torch.float32, device='cpu'):
+ """Generate 2D gaussian kernel.
+
+ Args:
+ radius (int): Radius of gaussian kernel.
+ sigma (int): Sigma of gaussian function. Default: 1.
+ dtype (torch.dtype): Dtype of gaussian tensor. Default: torch.float32.
+ device (str): Device of gaussian tensor. Default: 'cpu'.
+
+ Returns:
+ h (Tensor): Gaussian kernel with a
+ ``(2 * radius + 1) * (2 * radius + 1)`` shape.
+ """
+ x = torch.arange(
+ -radius, radius + 1, dtype=dtype, device=device).view(1, -1)
+ y = torch.arange(
+ -radius, radius + 1, dtype=dtype, device=device).view(-1, 1)
+
+ h = (-(x * x + y * y) / (2 * sigma * sigma)).exp()
+
+ h[h < torch.finfo(h.dtype).eps * h.max()] = 0
+ return h
+
+
+def gen_gaussian_target(heatmap, center, radius, k=1):
+ """Generate 2D gaussian heatmap.
+
+ Args:
+ heatmap (Tensor): Input heatmap, the gaussian kernel will cover on
+ it and maintain the max value.
+ center (list[int]): Coord of gaussian kernel's center.
+ radius (int): Radius of gaussian kernel.
+ k (int): Coefficient of gaussian kernel. Default: 1.
+
+ Returns:
+ out_heatmap (Tensor): Updated heatmap covered by gaussian kernel.
+ """
+ diameter = 2 * radius + 1
+ gaussian_kernel = gaussian2D(
+ radius, sigma=diameter / 6, dtype=heatmap.dtype, device=heatmap.device)
+
+ x, y = center
+
+ height, width = heatmap.shape[:2]
+
+ left, right = min(x, radius), min(width - x, radius + 1)
+ top, bottom = min(y, radius), min(height - y, radius + 1)
+
+ masked_heatmap = heatmap[y - top:y + bottom, x - left:x + right]
+ masked_gaussian = gaussian_kernel[radius - top:radius + bottom,
+ radius - left:radius + right]
+ out_heatmap = heatmap
+ torch.max(
+ masked_heatmap,
+ masked_gaussian * k,
+ out=out_heatmap[y - top:y + bottom, x - left:x + right])
+
+ return out_heatmap
+
+
+def gaussian_radius(det_size, min_overlap):
+ r"""Generate 2D gaussian radius.
+
+ This function is modified from the `official github repo
+ `_.
+
+ Given ``min_overlap``, radius could computed by a quadratic equation
+ according to Vieta's formulas.
+
+ There are 3 cases for computing gaussian radius, details are following:
+
+ - Explanation of figure: ``lt`` and ``br`` indicates the left-top and
+ bottom-right corner of ground truth box. ``x`` indicates the
+ generated corner at the limited position when ``radius=r``.
+
+ - Case1: one corner is inside the gt box and the other is outside.
+
+ .. code:: text
+
+ |< width >|
+
+ lt-+----------+ -
+ | | | ^
+ +--x----------+--+
+ | | | |
+ | | | | height
+ | | overlap | |
+ | | | |
+ | | | | v
+ +--+---------br--+ -
+ | | |
+ +----------+--x
+
+ To ensure IoU of generated box and gt box is larger than ``min_overlap``:
+
+ .. math::
+ \cfrac{(w-r)*(h-r)}{w*h+(w+h)r-r^2} \ge {iou} \quad\Rightarrow\quad
+ {r^2-(w+h)r+\cfrac{1-iou}{1+iou}*w*h} \ge 0 \\
+ {a} = 1,\quad{b} = {-(w+h)},\quad{c} = {\cfrac{1-iou}{1+iou}*w*h}
+ {r} \le \cfrac{-b-\sqrt{b^2-4*a*c}}{2*a}
+
+ - Case2: both two corners are inside the gt box.
+
+ .. code:: text
+
+ |< width >|
+
+ lt-+----------+ -
+ | | | ^
+ +--x-------+ |
+ | | | |
+ | |overlap| | height
+ | | | |
+ | +-------x--+
+ | | | v
+ +----------+-br -
+
+ To ensure IoU of generated box and gt box is larger than ``min_overlap``:
+
+ .. math::
+ \cfrac{(w-2*r)*(h-2*r)}{w*h} \ge {iou} \quad\Rightarrow\quad
+ {4r^2-2(w+h)r+(1-iou)*w*h} \ge 0 \\
+ {a} = 4,\quad {b} = {-2(w+h)},\quad {c} = {(1-iou)*w*h}
+ {r} \le \cfrac{-b-\sqrt{b^2-4*a*c}}{2*a}
+
+ - Case3: both two corners are outside the gt box.
+
+ .. code:: text
+
+ |< width >|
+
+ x--+----------------+
+ | | |
+ +-lt-------------+ | -
+ | | | | ^
+ | | | |
+ | | overlap | | height
+ | | | |
+ | | | | v
+ | +------------br--+ -
+ | | |
+ +----------------+--x
+
+ To ensure IoU of generated box and gt box is larger than ``min_overlap``:
+
+ .. math::
+ \cfrac{w*h}{(w+2*r)*(h+2*r)} \ge {iou} \quad\Rightarrow\quad
+ {4*iou*r^2+2*iou*(w+h)r+(iou-1)*w*h} \le 0 \\
+ {a} = {4*iou},\quad {b} = {2*iou*(w+h)},\quad {c} = {(iou-1)*w*h} \\
+ {r} \le \cfrac{-b+\sqrt{b^2-4*a*c}}{2*a}
+
+ Args:
+ det_size (list[int]): Shape of object.
+ min_overlap (float): Min IoU with ground truth for boxes generated by
+ keypoints inside the gaussian kernel.
+
+ Returns:
+ radius (int): Radius of gaussian kernel.
+ """
+ height, width = det_size
+
+ a1 = 1
+ b1 = (height + width)
+ c1 = width * height * (1 - min_overlap) / (1 + min_overlap)
+ sq1 = sqrt(b1**2 - 4 * a1 * c1)
+ r1 = (b1 - sq1) / (2 * a1)
+
+ a2 = 4
+ b2 = 2 * (height + width)
+ c2 = (1 - min_overlap) * width * height
+ sq2 = sqrt(b2**2 - 4 * a2 * c2)
+ r2 = (b2 - sq2) / (2 * a2)
+
+ a3 = 4 * min_overlap
+ b3 = -2 * min_overlap * (height + width)
+ c3 = (min_overlap - 1) * width * height
+ sq3 = sqrt(b3**2 - 4 * a3 * c3)
+ r3 = (b3 + sq3) / (2 * a3)
+ return min(r1, r2, r3)
+
+
+def get_local_maximum(heat, kernel=3):
+ """Extract local maximum pixel with given kernel.
+
+ Args:
+ heat (Tensor): Target heatmap.
+ kernel (int): Kernel size of max pooling. Default: 3.
+
+ Returns:
+ heat (Tensor): A heatmap where local maximum pixels maintain its
+ own value and other positions are 0.
+ """
+ pad = (kernel - 1) // 2
+ hmax = F.max_pool2d(heat, kernel, stride=1, padding=pad)
+ keep = (hmax == heat).float()
+ return heat * keep
+
+
+def get_topk_from_heatmap(scores, k=20):
+ """Get top k positions from heatmap.
+
+ Args:
+ scores (Tensor): Target heatmap with shape
+ [batch, num_classes, height, width].
+ k (int): Target number. Default: 20.
+
+ Returns:
+ tuple[torch.Tensor]: Scores, indexes, categories and coords of
+ topk keypoint. Containing following Tensors:
+
+ - topk_scores (Tensor): Max scores of each topk keypoint.
+ - topk_inds (Tensor): Indexes of each topk keypoint.
+ - topk_clses (Tensor): Categories of each topk keypoint.
+ - topk_ys (Tensor): Y-coord of each topk keypoint.
+ - topk_xs (Tensor): X-coord of each topk keypoint.
+ """
+ batch, _, height, width = scores.size()
+ topk_scores, topk_inds = torch.topk(scores.view(batch, -1), k)
+ topk_clses = topk_inds // (height * width)
+ topk_inds = topk_inds % (height * width)
+ topk_ys = topk_inds // width
+ topk_xs = (topk_inds % width).int().float()
+ return topk_scores, topk_inds, topk_clses, topk_ys, topk_xs
+
+
+def gather_feat(feat, ind, mask=None):
+ """Gather feature according to index.
+
+ Args:
+ feat (Tensor): Target feature map.
+ ind (Tensor): Target coord index.
+ mask (Tensor | None): Mask of feature map. Default: None.
+
+ Returns:
+ feat (Tensor): Gathered feature.
+ """
+ dim = feat.size(2)
+ ind = ind.unsqueeze(2).repeat(1, 1, dim)
+ feat = feat.gather(1, ind)
+ if mask is not None:
+ mask = mask.unsqueeze(2).expand_as(feat)
+ feat = feat[mask]
+ feat = feat.view(-1, dim)
+ return feat
+
+
+def transpose_and_gather_feat(feat, ind):
+ """Transpose and gather feature according to index.
+
+ Args:
+ feat (Tensor): Target feature map.
+ ind (Tensor): Target coord index.
+
+ Returns:
+ feat (Tensor): Transposed and gathered feature.
+ """
+ feat = feat.permute(0, 2, 3, 1).contiguous()
+ feat = feat.view(feat.size(0), -1, feat.size(3))
+ feat = gather_feat(feat, ind)
+ return feat
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/image.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/image.py
new file mode 100644
index 0000000000000000000000000000000000000000..16b5787a78232e46f47585c99526ca2b4ca9d1a1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/image.py
@@ -0,0 +1,52 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Union
+
+import mmcv
+import numpy as np
+import torch
+from torch import Tensor
+
+
+def imrenormalize(img: Union[Tensor, np.ndarray], img_norm_cfg: dict,
+ new_img_norm_cfg: dict) -> Union[Tensor, np.ndarray]:
+ """Re-normalize the image.
+
+ Args:
+ img (Tensor | ndarray): Input image. If the input is a Tensor, the
+ shape is (1, C, H, W). If the input is a ndarray, the shape
+ is (H, W, C).
+ img_norm_cfg (dict): Original configuration for the normalization.
+ new_img_norm_cfg (dict): New configuration for the normalization.
+
+ Returns:
+ Tensor | ndarray: Output image with the same type and shape of
+ the input.
+ """
+ if isinstance(img, torch.Tensor):
+ assert img.ndim == 4 and img.shape[0] == 1
+ new_img = img.squeeze(0).cpu().numpy().transpose(1, 2, 0)
+ new_img = _imrenormalize(new_img, img_norm_cfg, new_img_norm_cfg)
+ new_img = new_img.transpose(2, 0, 1)[None]
+ return torch.from_numpy(new_img).to(img)
+ else:
+ return _imrenormalize(img, img_norm_cfg, new_img_norm_cfg)
+
+
+def _imrenormalize(img: Union[Tensor, np.ndarray], img_norm_cfg: dict,
+ new_img_norm_cfg: dict) -> Union[Tensor, np.ndarray]:
+ """Re-normalize the image."""
+ img_norm_cfg = img_norm_cfg.copy()
+ new_img_norm_cfg = new_img_norm_cfg.copy()
+ for k, v in img_norm_cfg.items():
+ if (k == 'mean' or k == 'std') and not isinstance(v, np.ndarray):
+ img_norm_cfg[k] = np.array(v, dtype=img.dtype)
+ # reverse cfg
+ if 'bgr_to_rgb' in img_norm_cfg:
+ img_norm_cfg['rgb_to_bgr'] = img_norm_cfg['bgr_to_rgb']
+ img_norm_cfg.pop('bgr_to_rgb')
+ for k, v in new_img_norm_cfg.items():
+ if (k == 'mean' or k == 'std') and not isinstance(v, np.ndarray):
+ new_img_norm_cfg[k] = np.array(v, dtype=img.dtype)
+ img = mmcv.imdenormalize(img, **img_norm_cfg)
+ img = mmcv.imnormalize(img, **new_img_norm_cfg)
+ return img
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/make_divisible.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/make_divisible.py
new file mode 100644
index 0000000000000000000000000000000000000000..ed42c2eeea2a6aed03a0be5516b8d1ef1139e486
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/make_divisible.py
@@ -0,0 +1,28 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+def make_divisible(value, divisor, min_value=None, min_ratio=0.9):
+ """Make divisible function.
+
+ This function rounds the channel number to the nearest value that can be
+ divisible by the divisor. It is taken from the original tf repo. It ensures
+ that all layers have a channel number that is divisible by divisor. It can
+ be seen here: https://github.com/tensorflow/models/blob/master/research/slim/nets/mobilenet/mobilenet.py # noqa
+
+ Args:
+ value (int): The original channel number.
+ divisor (int): The divisor to fully divide the channel number.
+ min_value (int): The minimum value of the output channel.
+ Default: None, means that the minimum value equal to the divisor.
+ min_ratio (float): The minimum ratio of the rounded channel number to
+ the original channel number. Default: 0.9.
+
+ Returns:
+ int: The modified output channel number.
+ """
+
+ if min_value is None:
+ min_value = divisor
+ new_value = max(min_value, int(value + divisor / 2) // divisor * divisor)
+ # Make sure that round down does not go down by more than (1-min_ratio).
+ if new_value < min_ratio * value:
+ new_value += divisor
+ return new_value
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/misc.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/misc.py
new file mode 100644
index 0000000000000000000000000000000000000000..2cf429153ba7e0be025396b069aef8212144e34d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/misc.py
@@ -0,0 +1,697 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from functools import partial
+from typing import List, Optional, Sequence, Tuple, Union
+
+import numpy as np
+import torch
+from mmengine.structures import InstanceData
+from mmengine.utils import digit_version
+from six.moves import map, zip
+from torch import Tensor
+from torch.autograd import Function
+from torch.nn import functional as F
+
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import BaseBoxes, get_box_type, stack_boxes
+from mmdet.structures.mask import BitmapMasks, PolygonMasks
+from mmdet.utils import OptInstanceList
+
+
+class SigmoidGeometricMean(Function):
+ """Forward and backward function of geometric mean of two sigmoid
+ functions.
+
+ This implementation with analytical gradient function substitutes
+ the autograd function of (x.sigmoid() * y.sigmoid()).sqrt(). The
+ original implementation incurs none during gradient backprapagation
+ if both x and y are very small values.
+ """
+
+ @staticmethod
+ def forward(ctx, x, y):
+ x_sigmoid = x.sigmoid()
+ y_sigmoid = y.sigmoid()
+ z = (x_sigmoid * y_sigmoid).sqrt()
+ ctx.save_for_backward(x_sigmoid, y_sigmoid, z)
+ return z
+
+ @staticmethod
+ def backward(ctx, grad_output):
+ x_sigmoid, y_sigmoid, z = ctx.saved_tensors
+ grad_x = grad_output * z * (1 - x_sigmoid) / 2
+ grad_y = grad_output * z * (1 - y_sigmoid) / 2
+ return grad_x, grad_y
+
+
+sigmoid_geometric_mean = SigmoidGeometricMean.apply
+
+
+def interpolate_as(source, target, mode='bilinear', align_corners=False):
+ """Interpolate the `source` to the shape of the `target`.
+
+ The `source` must be a Tensor, but the `target` can be a Tensor or a
+ np.ndarray with the shape (..., target_h, target_w).
+
+ Args:
+ source (Tensor): A 3D/4D Tensor with the shape (N, H, W) or
+ (N, C, H, W).
+ target (Tensor | np.ndarray): The interpolation target with the shape
+ (..., target_h, target_w).
+ mode (str): Algorithm used for interpolation. The options are the
+ same as those in F.interpolate(). Default: ``'bilinear'``.
+ align_corners (bool): The same as the argument in F.interpolate().
+
+ Returns:
+ Tensor: The interpolated source Tensor.
+ """
+ assert len(target.shape) >= 2
+
+ def _interpolate_as(source, target, mode='bilinear', align_corners=False):
+ """Interpolate the `source` (4D) to the shape of the `target`."""
+ target_h, target_w = target.shape[-2:]
+ source_h, source_w = source.shape[-2:]
+ if target_h != source_h or target_w != source_w:
+ source = F.interpolate(
+ source,
+ size=(target_h, target_w),
+ mode=mode,
+ align_corners=align_corners)
+ return source
+
+ if len(source.shape) == 3:
+ source = source[:, None, :, :]
+ source = _interpolate_as(source, target, mode, align_corners)
+ return source[:, 0, :, :]
+ else:
+ return _interpolate_as(source, target, mode, align_corners)
+
+
+def unpack_gt_instances(batch_data_samples: SampleList) -> tuple:
+ """Unpack ``gt_instances``, ``gt_instances_ignore`` and ``img_metas`` based
+ on ``batch_data_samples``
+
+ Args:
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+
+ Returns:
+ tuple:
+
+ - batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ - batch_gt_instances_ignore (list[:obj:`InstanceData`]):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+ - batch_img_metas (list[dict]): Meta information of each image,
+ e.g., image size, scaling factor, etc.
+ """
+ batch_gt_instances = []
+ batch_gt_instances_ignore = []
+ batch_img_metas = []
+ for data_sample in batch_data_samples:
+ batch_img_metas.append(data_sample.metainfo)
+ batch_gt_instances.append(data_sample.gt_instances)
+ if 'ignored_instances' in data_sample:
+ batch_gt_instances_ignore.append(data_sample.ignored_instances)
+ else:
+ batch_gt_instances_ignore.append(None)
+
+ return batch_gt_instances, batch_gt_instances_ignore, batch_img_metas
+
+
+def empty_instances(batch_img_metas: List[dict],
+ device: torch.device,
+ task_type: str,
+ instance_results: OptInstanceList = None,
+ mask_thr_binary: Union[int, float] = 0,
+ box_type: Union[str, type] = 'hbox',
+ use_box_type: bool = False,
+ num_classes: int = 80,
+ score_per_cls: bool = False) -> List[InstanceData]:
+ """Handle predicted instances when RoI is empty.
+
+ Note: If ``instance_results`` is not None, it will be modified
+ in place internally, and then return ``instance_results``
+
+ Args:
+ batch_img_metas (list[dict]): List of image information.
+ device (torch.device): Device of tensor.
+ task_type (str): Expected returned task type. it currently
+ supports bbox and mask.
+ instance_results (list[:obj:`InstanceData`]): List of instance
+ results.
+ mask_thr_binary (int, float): mask binarization threshold.
+ Defaults to 0.
+ box_type (str or type): The empty box type. Defaults to `hbox`.
+ use_box_type (bool): Whether to warp boxes with the box type.
+ Defaults to False.
+ num_classes (int): num_classes of bbox_head. Defaults to 80.
+ score_per_cls (bool): Whether to generate classwise score for
+ the empty instance. ``score_per_cls`` will be True when the model
+ needs to produce raw results without nms. Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ """
+ assert task_type in ('bbox', 'mask'), 'Only support bbox and mask,' \
+ f' but got {task_type}'
+
+ if instance_results is not None:
+ assert len(instance_results) == len(batch_img_metas)
+
+ results_list = []
+ for img_id in range(len(batch_img_metas)):
+ if instance_results is not None:
+ results = instance_results[img_id]
+ assert isinstance(results, InstanceData)
+ else:
+ results = InstanceData()
+
+ if task_type == 'bbox':
+ _, box_type = get_box_type(box_type)
+ bboxes = torch.zeros(0, box_type.box_dim, device=device)
+ if use_box_type:
+ bboxes = box_type(bboxes, clone=False)
+ results.bboxes = bboxes
+ score_shape = (0, num_classes + 1) if score_per_cls else (0, )
+ results.scores = torch.zeros(score_shape, device=device)
+ results.labels = torch.zeros((0, ),
+ device=device,
+ dtype=torch.long)
+ else:
+ # TODO: Handle the case where rescale is false
+ img_h, img_w = batch_img_metas[img_id]['ori_shape'][:2]
+ # the type of `im_mask` will be torch.bool or torch.uint8,
+ # where uint8 if for visualization and debugging.
+ im_mask = torch.zeros(
+ 0,
+ img_h,
+ img_w,
+ device=device,
+ dtype=torch.bool if mask_thr_binary >= 0 else torch.uint8)
+ results.masks = im_mask
+ results_list.append(results)
+ return results_list
+
+
+def multi_apply(func, *args, **kwargs):
+ """Apply function to a list of arguments.
+
+ Note:
+ This function applies the ``func`` to multiple inputs and
+ map the multiple outputs of the ``func`` into different
+ list. Each list contains the same type of outputs corresponding
+ to different inputs.
+
+ Args:
+ func (Function): A function that will be applied to a list of
+ arguments
+
+ Returns:
+ tuple(list): A tuple containing multiple list, each list contains \
+ a kind of returned results by the function
+ """
+ pfunc = partial(func, **kwargs) if kwargs else func
+ map_results = map(pfunc, *args)
+ return tuple(map(list, zip(*map_results)))
+
+
+def unmap(data, count, inds, fill=0):
+ """Unmap a subset of item (data) back to the original set of items (of size
+ count)"""
+ if data.dim() == 1:
+ ret = data.new_full((count, ), fill)
+ ret[inds.type(torch.bool)] = data
+ else:
+ new_size = (count, ) + data.size()[1:]
+ ret = data.new_full(new_size, fill)
+ ret[inds.type(torch.bool), :] = data
+ return ret
+
+
+def mask2ndarray(mask):
+ """Convert Mask to ndarray..
+
+ Args:
+ mask (:obj:`BitmapMasks` or :obj:`PolygonMasks` or
+ torch.Tensor or np.ndarray): The mask to be converted.
+
+ Returns:
+ np.ndarray: Ndarray mask of shape (n, h, w) that has been converted
+ """
+ if isinstance(mask, (BitmapMasks, PolygonMasks)):
+ mask = mask.to_ndarray()
+ elif isinstance(mask, torch.Tensor):
+ mask = mask.detach().cpu().numpy()
+ elif not isinstance(mask, np.ndarray):
+ raise TypeError(f'Unsupported {type(mask)} data type')
+ return mask
+
+
+def flip_tensor(src_tensor, flip_direction):
+ """flip tensor base on flip_direction.
+
+ Args:
+ src_tensor (Tensor): input feature map, shape (B, C, H, W).
+ flip_direction (str): The flipping direction. Options are
+ 'horizontal', 'vertical', 'diagonal'.
+
+ Returns:
+ out_tensor (Tensor): Flipped tensor.
+ """
+ assert src_tensor.ndim == 4
+ valid_directions = ['horizontal', 'vertical', 'diagonal']
+ assert flip_direction in valid_directions
+ if flip_direction == 'horizontal':
+ out_tensor = torch.flip(src_tensor, [3])
+ elif flip_direction == 'vertical':
+ out_tensor = torch.flip(src_tensor, [2])
+ else:
+ out_tensor = torch.flip(src_tensor, [2, 3])
+ return out_tensor
+
+
+def select_single_mlvl(mlvl_tensors, batch_id, detach=True):
+ """Extract a multi-scale single image tensor from a multi-scale batch
+ tensor based on batch index.
+
+ Note: The default value of detach is True, because the proposal gradient
+ needs to be detached during the training of the two-stage model. E.g
+ Cascade Mask R-CNN.
+
+ Args:
+ mlvl_tensors (list[Tensor]): Batch tensor for all scale levels,
+ each is a 4D-tensor.
+ batch_id (int): Batch index.
+ detach (bool): Whether detach gradient. Default True.
+
+ Returns:
+ list[Tensor]: Multi-scale single image tensor.
+ """
+ assert isinstance(mlvl_tensors, (list, tuple))
+ num_levels = len(mlvl_tensors)
+
+ if detach:
+ mlvl_tensor_list = [
+ mlvl_tensors[i][batch_id].detach() for i in range(num_levels)
+ ]
+ else:
+ mlvl_tensor_list = [
+ mlvl_tensors[i][batch_id] for i in range(num_levels)
+ ]
+ return mlvl_tensor_list
+
+
+def filter_scores_and_topk(scores, score_thr, topk, results=None):
+ """Filter results using score threshold and topk candidates.
+
+ Args:
+ scores (Tensor): The scores, shape (num_bboxes, K).
+ score_thr (float): The score filter threshold.
+ topk (int): The number of topk candidates.
+ results (dict or list or Tensor, Optional): The results to
+ which the filtering rule is to be applied. The shape
+ of each item is (num_bboxes, N).
+
+ Returns:
+ tuple: Filtered results
+
+ - scores (Tensor): The scores after being filtered, \
+ shape (num_bboxes_filtered, ).
+ - labels (Tensor): The class labels, shape \
+ (num_bboxes_filtered, ).
+ - anchor_idxs (Tensor): The anchor indexes, shape \
+ (num_bboxes_filtered, ).
+ - filtered_results (dict or list or Tensor, Optional): \
+ The filtered results. The shape of each item is \
+ (num_bboxes_filtered, N).
+ """
+ valid_mask = scores > score_thr
+ scores = scores[valid_mask]
+ valid_idxs = torch.nonzero(valid_mask)
+
+ num_topk = min(topk, valid_idxs.size(0))
+ # torch.sort is actually faster than .topk (at least on GPUs)
+ scores, idxs = scores.sort(descending=True)
+ scores = scores[:num_topk]
+ topk_idxs = valid_idxs[idxs[:num_topk]]
+ keep_idxs, labels = topk_idxs.unbind(dim=1)
+
+ filtered_results = None
+ if results is not None:
+ if isinstance(results, dict):
+ filtered_results = {k: v[keep_idxs] for k, v in results.items()}
+ elif isinstance(results, list):
+ filtered_results = [result[keep_idxs] for result in results]
+ elif isinstance(results, torch.Tensor):
+ filtered_results = results[keep_idxs]
+ else:
+ raise NotImplementedError(f'Only supports dict or list or Tensor, '
+ f'but get {type(results)}.')
+ return scores, labels, keep_idxs, filtered_results
+
+
+def center_of_mass(mask, esp=1e-6):
+ """Calculate the centroid coordinates of the mask.
+
+ Args:
+ mask (Tensor): The mask to be calculated, shape (h, w).
+ esp (float): Avoid dividing by zero. Default: 1e-6.
+
+ Returns:
+ tuple[Tensor]: the coordinates of the center point of the mask.
+
+ - center_h (Tensor): the center point of the height.
+ - center_w (Tensor): the center point of the width.
+ """
+ h, w = mask.shape
+ grid_h = torch.arange(h, device=mask.device)[:, None]
+ grid_w = torch.arange(w, device=mask.device)
+ normalizer = mask.sum().float().clamp(min=esp)
+ center_h = (mask * grid_h).sum() / normalizer
+ center_w = (mask * grid_w).sum() / normalizer
+ return center_h, center_w
+
+
+def generate_coordinate(featmap_sizes, device='cuda'):
+ """Generate the coordinate.
+
+ Args:
+ featmap_sizes (tuple): The feature to be calculated,
+ of shape (N, C, W, H).
+ device (str): The device where the feature will be put on.
+ Returns:
+ coord_feat (Tensor): The coordinate feature, of shape (N, 2, W, H).
+ """
+
+ x_range = torch.linspace(-1, 1, featmap_sizes[-1], device=device)
+ y_range = torch.linspace(-1, 1, featmap_sizes[-2], device=device)
+ y, x = torch.meshgrid(y_range, x_range)
+ y = y.expand([featmap_sizes[0], 1, -1, -1])
+ x = x.expand([featmap_sizes[0], 1, -1, -1])
+ coord_feat = torch.cat([x, y], 1)
+
+ return coord_feat
+
+
+def levels_to_images(mlvl_tensor: List[torch.Tensor]) -> List[torch.Tensor]:
+ """Concat multi-level feature maps by image.
+
+ [feature_level0, feature_level1...] -> [feature_image0, feature_image1...]
+ Convert the shape of each element in mlvl_tensor from (N, C, H, W) to
+ (N, H*W , C), then split the element to N elements with shape (H*W, C), and
+ concat elements in same image of all level along first dimension.
+
+ Args:
+ mlvl_tensor (list[Tensor]): list of Tensor which collect from
+ corresponding level. Each element is of shape (N, C, H, W)
+
+ Returns:
+ list[Tensor]: A list that contains N tensors and each tensor is
+ of shape (num_elements, C)
+ """
+ batch_size = mlvl_tensor[0].size(0)
+ batch_list = [[] for _ in range(batch_size)]
+ channels = mlvl_tensor[0].size(1)
+ for t in mlvl_tensor:
+ t = t.permute(0, 2, 3, 1)
+ t = t.view(batch_size, -1, channels).contiguous()
+ for img in range(batch_size):
+ batch_list[img].append(t[img])
+ return [torch.cat(item, 0) for item in batch_list]
+
+
+def images_to_levels(target, num_levels):
+ """Convert targets by image to targets by feature level.
+
+ [target_img0, target_img1] -> [target_level0, target_level1, ...]
+ """
+ target = stack_boxes(target, 0)
+ level_targets = []
+ start = 0
+ for n in num_levels:
+ end = start + n
+ # level_targets.append(target[:, start:end].squeeze(0))
+ level_targets.append(target[:, start:end])
+ start = end
+ return level_targets
+
+
+def samplelist_boxtype2tensor(batch_data_samples: SampleList) -> SampleList:
+ for data_samples in batch_data_samples:
+ if 'gt_instances' in data_samples:
+ bboxes = data_samples.gt_instances.get('bboxes', None)
+ if isinstance(bboxes, BaseBoxes):
+ data_samples.gt_instances.bboxes = bboxes.tensor
+ if 'pred_instances' in data_samples:
+ bboxes = data_samples.pred_instances.get('bboxes', None)
+ if isinstance(bboxes, BaseBoxes):
+ data_samples.pred_instances.bboxes = bboxes.tensor
+ if 'ignored_instances' in data_samples:
+ bboxes = data_samples.ignored_instances.get('bboxes', None)
+ if isinstance(bboxes, BaseBoxes):
+ data_samples.ignored_instances.bboxes = bboxes.tensor
+
+
+_torch_version_div_indexing = (
+ 'parrots' not in torch.__version__
+ and digit_version(torch.__version__) >= digit_version('1.8'))
+
+
+def floordiv(dividend, divisor, rounding_mode='trunc'):
+ if _torch_version_div_indexing:
+ return torch.div(dividend, divisor, rounding_mode=rounding_mode)
+ else:
+ return dividend // divisor
+
+
+def _filter_gt_instances_by_score(batch_data_samples: SampleList,
+ score_thr: float) -> SampleList:
+ """Filter ground truth (GT) instances by score.
+
+ Args:
+ batch_data_samples (SampleList): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ score_thr (float): The score filter threshold.
+
+ Returns:
+ SampleList: The Data Samples filtered by score.
+ """
+ for data_samples in batch_data_samples:
+ assert 'scores' in data_samples.gt_instances, \
+ 'there does not exit scores in instances'
+ if data_samples.gt_instances.bboxes.shape[0] > 0:
+ data_samples.gt_instances = data_samples.gt_instances[
+ data_samples.gt_instances.scores > score_thr]
+ return batch_data_samples
+
+
+def _filter_gt_instances_by_size(batch_data_samples: SampleList,
+ wh_thr: tuple) -> SampleList:
+ """Filter ground truth (GT) instances by size.
+
+ Args:
+ batch_data_samples (SampleList): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ wh_thr (tuple): Minimum width and height of bbox.
+
+ Returns:
+ SampleList: The Data Samples filtered by score.
+ """
+ for data_samples in batch_data_samples:
+ bboxes = data_samples.gt_instances.bboxes
+ if bboxes.shape[0] > 0:
+ w = bboxes[:, 2] - bboxes[:, 0]
+ h = bboxes[:, 3] - bboxes[:, 1]
+ data_samples.gt_instances = data_samples.gt_instances[
+ (w > wh_thr[0]) & (h > wh_thr[1])]
+ return batch_data_samples
+
+
+def filter_gt_instances(batch_data_samples: SampleList,
+ score_thr: float = None,
+ wh_thr: tuple = None):
+ """Filter ground truth (GT) instances by score and/or size.
+
+ Args:
+ batch_data_samples (SampleList): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ score_thr (float): The score filter threshold.
+ wh_thr (tuple): Minimum width and height of bbox.
+
+ Returns:
+ SampleList: The Data Samples filtered by score and/or size.
+ """
+
+ if score_thr is not None:
+ batch_data_samples = _filter_gt_instances_by_score(
+ batch_data_samples, score_thr)
+ if wh_thr is not None:
+ batch_data_samples = _filter_gt_instances_by_size(
+ batch_data_samples, wh_thr)
+ return batch_data_samples
+
+
+def rename_loss_dict(prefix: str, losses: dict) -> dict:
+ """Rename the key names in loss dict by adding a prefix.
+
+ Args:
+ prefix (str): The prefix for loss components.
+ losses (dict): A dictionary of loss components.
+
+ Returns:
+ dict: A dictionary of loss components with prefix.
+ """
+ return {prefix + k: v for k, v in losses.items()}
+
+
+def reweight_loss_dict(losses: dict, weight: float) -> dict:
+ """Reweight losses in the dict by weight.
+
+ Args:
+ losses (dict): A dictionary of loss components.
+ weight (float): Weight for loss components.
+
+ Returns:
+ dict: A dictionary of weighted loss components.
+ """
+ for name, loss in losses.items():
+ if 'loss' in name:
+ if isinstance(loss, Sequence):
+ losses[name] = [item * weight for item in loss]
+ else:
+ losses[name] = loss * weight
+ return losses
+
+
+def relative_coordinate_maps(
+ locations: Tensor,
+ centers: Tensor,
+ strides: Tensor,
+ size_of_interest: int,
+ feat_sizes: Tuple[int],
+) -> Tensor:
+ """Generate the relative coordinate maps with feat_stride.
+
+ Args:
+ locations (Tensor): The prior location of mask feature map.
+ It has shape (num_priors, 2).
+ centers (Tensor): The prior points of a object in
+ all feature pyramid. It has shape (num_pos, 2)
+ strides (Tensor): The prior strides of a object in
+ all feature pyramid. It has shape (num_pos, 1)
+ size_of_interest (int): The size of the region used in rel coord.
+ feat_sizes (Tuple[int]): The feature size H and W, which has 2 dims.
+ Returns:
+ rel_coord_feat (Tensor): The coordinate feature
+ of shape (num_pos, 2, H, W).
+ """
+
+ H, W = feat_sizes
+ rel_coordinates = centers.reshape(-1, 1, 2) - locations.reshape(1, -1, 2)
+ rel_coordinates = rel_coordinates.permute(0, 2, 1).float()
+ rel_coordinates = rel_coordinates / (
+ strides[:, None, None] * size_of_interest)
+ return rel_coordinates.reshape(-1, 2, H, W)
+
+
+def aligned_bilinear(tensor: Tensor, factor: int) -> Tensor:
+ """aligned bilinear, used in original implement in CondInst:
+
+ https://github.com/aim-uofa/AdelaiDet/blob/\
+ c0b2092ce72442b0f40972f7c6dda8bb52c46d16/adet/utils/comm.py#L23
+ """
+
+ assert tensor.dim() == 4
+ assert factor >= 1
+ assert int(factor) == factor
+
+ if factor == 1:
+ return tensor
+
+ h, w = tensor.size()[2:]
+ tensor = F.pad(tensor, pad=(0, 1, 0, 1), mode='replicate')
+ oh = factor * h + 1
+ ow = factor * w + 1
+ tensor = F.interpolate(
+ tensor, size=(oh, ow), mode='bilinear', align_corners=True)
+ tensor = F.pad(
+ tensor, pad=(factor // 2, 0, factor // 2, 0), mode='replicate')
+
+ return tensor[:, :, :oh - 1, :ow - 1]
+
+
+def unfold_wo_center(x, kernel_size: int, dilation: int) -> Tensor:
+ """unfold_wo_center, used in original implement in BoxInst:
+
+ https://github.com/aim-uofa/AdelaiDet/blob/\
+ 4a3a1f7372c35b48ebf5f6adc59f135a0fa28d60/\
+ adet/modeling/condinst/condinst.py#L53
+ """
+ assert x.dim() == 4
+ assert kernel_size % 2 == 1
+
+ # using SAME padding
+ padding = (kernel_size + (dilation - 1) * (kernel_size - 1)) // 2
+ unfolded_x = F.unfold(
+ x, kernel_size=kernel_size, padding=padding, dilation=dilation)
+ unfolded_x = unfolded_x.reshape(
+ x.size(0), x.size(1), -1, x.size(2), x.size(3))
+ # remove the center pixels
+ size = kernel_size**2
+ unfolded_x = torch.cat(
+ (unfolded_x[:, :, :size // 2], unfolded_x[:, :, size // 2 + 1:]),
+ dim=2)
+
+ return unfolded_x
+
+
+def padding_to(input_tensor: Tensor, max_len: int = 300) -> Tensor:
+ """Pad the first dimension of `input_tensor` to `max_len`.
+
+ Args:
+ input_tensor (Tensor): The tensor to be padded,
+ max_len (int): Padding target size in the first dimension.
+ Default: 300
+ https://github.com/jshilong/DDQ/blob/ddq_detr/projects/models/utils.py#L19
+ Returns:
+ Tensor: The tensor padded with the first dimension size `max_len`.
+ """
+ if max_len is None:
+ return input_tensor
+ num_padding = max_len - len(input_tensor)
+ if input_tensor.dim() > 1:
+ padding = input_tensor.new_zeros(
+ num_padding, *input_tensor.size()[1:], dtype=input_tensor.dtype)
+ else:
+ padding = input_tensor.new_zeros(num_padding, dtype=input_tensor.dtype)
+ output_tensor = torch.cat([input_tensor, padding], dim=0)
+ return output_tensor
+
+
+def align_tensor(inputs: List[Tensor],
+ max_len: Optional[int] = None) -> Tensor:
+ """Pad each input to `max_len`, then stack them. If `max_len` is None, then
+ it is the max size of the first dimension of each input.
+
+ https://github.com/jshilong/DDQ/blob/ddq_detr/projects/models/\
+ utils.py#L12
+
+ Args:
+ inputs (list[Tensor]): The tensors to be padded,
+ Each input should have the same shape except the first dimension.
+ max_len (int): Padding target size in the first dimension.
+ Default: None
+ Returns:
+ Tensor: Stacked inputs after padding in the first dimension.
+ """
+ if max_len is None:
+ max_len = max([len(item) for item in inputs])
+
+ return torch.stack([padding_to(item, max_len) for item in inputs])
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/panoptic_gt_processing.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/panoptic_gt_processing.py
new file mode 100644
index 0000000000000000000000000000000000000000..7a3bc95fc04040b4a2a13fa63f2d02f092f725e6
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/panoptic_gt_processing.py
@@ -0,0 +1,70 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Tuple
+
+import torch
+from torch import Tensor
+
+
+def preprocess_panoptic_gt(gt_labels: Tensor, gt_masks: Tensor,
+ gt_semantic_seg: Tensor, num_things: int,
+ num_stuff: int) -> Tuple[Tensor, Tensor]:
+ """Preprocess the ground truth for a image.
+
+ Args:
+ gt_labels (Tensor): Ground truth labels of each bbox,
+ with shape (num_gts, ).
+ gt_masks (BitmapMasks): Ground truth masks of each instances
+ of a image, shape (num_gts, h, w).
+ gt_semantic_seg (Tensor | None): Ground truth of semantic
+ segmentation with the shape (1, h, w).
+ [0, num_thing_class - 1] means things,
+ [num_thing_class, num_class-1] means stuff,
+ 255 means VOID. It's None when training instance segmentation.
+
+ Returns:
+ tuple[Tensor, Tensor]: a tuple containing the following targets.
+
+ - labels (Tensor): Ground truth class indices for a
+ image, with shape (n, ), n is the sum of number
+ of stuff type and number of instance in a image.
+ - masks (Tensor): Ground truth mask for a image, with
+ shape (n, h, w). Contains stuff and things when training
+ panoptic segmentation, and things only when training
+ instance segmentation.
+ """
+ num_classes = num_things + num_stuff
+ things_masks = gt_masks.to_tensor(
+ dtype=torch.bool, device=gt_labels.device)
+
+ if gt_semantic_seg is None:
+ masks = things_masks.long()
+ return gt_labels, masks
+
+ things_labels = gt_labels
+ gt_semantic_seg = gt_semantic_seg.squeeze(0)
+
+ semantic_labels = torch.unique(
+ gt_semantic_seg,
+ sorted=False,
+ return_inverse=False,
+ return_counts=False)
+ stuff_masks_list = []
+ stuff_labels_list = []
+ for label in semantic_labels:
+ if label < num_things or label >= num_classes:
+ continue
+ stuff_mask = gt_semantic_seg == label
+ stuff_masks_list.append(stuff_mask)
+ stuff_labels_list.append(label)
+
+ if len(stuff_masks_list) > 0:
+ stuff_masks = torch.stack(stuff_masks_list, dim=0)
+ stuff_labels = torch.stack(stuff_labels_list, dim=0)
+ labels = torch.cat([things_labels, stuff_labels], dim=0)
+ masks = torch.cat([things_masks, stuff_masks], dim=0)
+ else:
+ labels = things_labels
+ masks = things_masks
+
+ masks = masks.long()
+ return labels, masks
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/point_sample.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/point_sample.py
new file mode 100644
index 0000000000000000000000000000000000000000..1afc957f3da7d1dc030c21d40311c768c6952ea4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/point_sample.py
@@ -0,0 +1,88 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+from mmcv.ops import point_sample
+from torch import Tensor
+
+
+def get_uncertainty(mask_preds: Tensor, labels: Tensor) -> Tensor:
+ """Estimate uncertainty based on pred logits.
+
+ We estimate uncertainty as L1 distance between 0.0 and the logits
+ prediction in 'mask_preds' for the foreground class in `classes`.
+
+ Args:
+ mask_preds (Tensor): mask predication logits, shape (num_rois,
+ num_classes, mask_height, mask_width).
+
+ labels (Tensor): Either predicted or ground truth label for
+ each predicted mask, of length num_rois.
+
+ Returns:
+ scores (Tensor): Uncertainty scores with the most uncertain
+ locations having the highest uncertainty score,
+ shape (num_rois, 1, mask_height, mask_width)
+ """
+ if mask_preds.shape[1] == 1:
+ gt_class_logits = mask_preds.clone()
+ else:
+ inds = torch.arange(mask_preds.shape[0], device=mask_preds.device)
+ gt_class_logits = mask_preds[inds, labels].unsqueeze(1)
+ return -torch.abs(gt_class_logits)
+
+
+def get_uncertain_point_coords_with_randomness(
+ mask_preds: Tensor, labels: Tensor, num_points: int,
+ oversample_ratio: float, importance_sample_ratio: float) -> Tensor:
+ """Get ``num_points`` most uncertain points with random points during
+ train.
+
+ Sample points in [0, 1] x [0, 1] coordinate space based on their
+ uncertainty. The uncertainties are calculated for each point using
+ 'get_uncertainty()' function that takes point's logit prediction as
+ input.
+
+ Args:
+ mask_preds (Tensor): A tensor of shape (num_rois, num_classes,
+ mask_height, mask_width) for class-specific or class-agnostic
+ prediction.
+ labels (Tensor): The ground truth class for each instance.
+ num_points (int): The number of points to sample.
+ oversample_ratio (float): Oversampling parameter.
+ importance_sample_ratio (float): Ratio of points that are sampled
+ via importnace sampling.
+
+ Returns:
+ point_coords (Tensor): A tensor of shape (num_rois, num_points, 2)
+ that contains the coordinates sampled points.
+ """
+ assert oversample_ratio >= 1
+ assert 0 <= importance_sample_ratio <= 1
+ batch_size = mask_preds.shape[0]
+ num_sampled = int(num_points * oversample_ratio)
+ point_coords = torch.rand(
+ batch_size, num_sampled, 2, device=mask_preds.device)
+ point_logits = point_sample(mask_preds, point_coords)
+ # It is crucial to calculate uncertainty based on the sampled
+ # prediction value for the points. Calculating uncertainties of the
+ # coarse predictions first and sampling them for points leads to
+ # incorrect results. To illustrate this: assume uncertainty func(
+ # logits)=-abs(logits), a sampled point between two coarse
+ # predictions with -1 and 1 logits has 0 logits, and therefore 0
+ # uncertainty value. However, if we calculate uncertainties for the
+ # coarse predictions first, both will have -1 uncertainty,
+ # and sampled point will get -1 uncertainty.
+ point_uncertainties = get_uncertainty(point_logits, labels)
+ num_uncertain_points = int(importance_sample_ratio * num_points)
+ num_random_points = num_points - num_uncertain_points
+ idx = torch.topk(
+ point_uncertainties[:, 0, :], k=num_uncertain_points, dim=1)[1]
+ shift = num_sampled * torch.arange(
+ batch_size, dtype=torch.long, device=mask_preds.device)
+ idx += shift[:, None]
+ point_coords = point_coords.view(-1, 2)[idx.view(-1), :].view(
+ batch_size, num_uncertain_points, 2)
+ if num_random_points > 0:
+ rand_roi_coords = torch.rand(
+ batch_size, num_random_points, 2, device=mask_preds.device)
+ point_coords = torch.cat((point_coords, rand_roi_coords), dim=1)
+ return point_coords
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/vlfuse_helper.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/vlfuse_helper.py
new file mode 100644
index 0000000000000000000000000000000000000000..76b54de317c1f24d7cb40573954f988fd94fef42
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/vlfuse_helper.py
@@ -0,0 +1,773 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+# Modified from https://github.com/microsoft/GLIP/blob/main/maskrcnn_benchmark/utils/fuse_helper.py # noqa
+# and https://github.com/microsoft/GLIP/blob/main/maskrcnn_benchmark/modeling/rpn/modeling_bert.py # noqa
+import math
+from typing import Dict, Optional, Tuple
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+import torch.utils.checkpoint as checkpoint
+from mmcv.cnn.bricks import DropPath
+from torch import Tensor
+
+try:
+ from transformers import BertConfig, BertPreTrainedModel
+ from transformers.modeling_utils import apply_chunking_to_forward
+ from transformers.models.bert.modeling_bert import \
+ BertAttention as HFBertAttention
+ from transformers.models.bert.modeling_bert import \
+ BertIntermediate as HFBertIntermediate
+ from transformers.models.bert.modeling_bert import \
+ BertOutput as HFBertOutput
+except ImportError:
+ BertConfig = None
+ BertPreTrainedModel = object
+ apply_chunking_to_forward = None
+ HFBertAttention = object
+ HFBertIntermediate = object
+ HFBertOutput = object
+
+MAX_CLAMP_VALUE = 50000
+
+
+def permute_and_flatten(layer: Tensor, N: int, A: int, C: int, H: int,
+ W: int) -> Tensor:
+ """Permute and then flatten a tensor,
+
+ from size (N, A, C, H, W) to (N, H * W * A, C).
+
+ Args:
+ layer (Tensor): Tensor of shape (N, C, H, W).
+ N (int): Batch size.
+ A (int): Number of attention heads.
+ C (int): Number of channels.
+ H (int): Height of feature map.
+ W (int): Width of feature map.
+
+ Returns:
+ Tensor: A Tensor of shape (N, H * W * A, C).
+ """
+ layer = layer.view(N, A, C, H, W)
+ layer = layer.permute(0, 3, 4, 1, 2)
+ layer = layer.reshape(N, -1, C)
+ return layer
+
+
+def clamp_values(vector: Tensor) -> Tensor:
+ """Clamp the values of a vector to the range [-MAX_CLAMP_VALUE,
+ MAX_CLAMP_VALUE].
+
+ Args:
+ vector (Tensor): Tensor of shape (N, C, H, W).
+
+ Returns:
+ Tensor: A Tensor of shape (N, C, H, W) with clamped values.
+ """
+ vector = torch.clamp(vector, min=-MAX_CLAMP_VALUE, max=MAX_CLAMP_VALUE)
+ return vector
+
+
+class BiMultiHeadAttention(nn.Module):
+ """Bidirectional fusion Multi-Head Attention layer.
+
+ Args:
+ v_dim (int): The dimension of the vision input.
+ l_dim (int): The dimension of the language input.
+ embed_dim (int): The embedding dimension for the attention operation.
+ num_heads (int): The number of attention heads.
+ dropout (float, optional): The dropout probability. Defaults to 0.1.
+ """
+
+ def __init__(self,
+ v_dim: int,
+ l_dim: int,
+ embed_dim: int,
+ num_heads: int,
+ dropout: float = 0.1):
+ super(BiMultiHeadAttention, self).__init__()
+
+ self.embed_dim = embed_dim
+ self.num_heads = num_heads
+ self.head_dim = embed_dim // num_heads
+ self.v_dim = v_dim
+ self.l_dim = l_dim
+
+ assert (
+ self.head_dim * self.num_heads == self.embed_dim
+ ), 'embed_dim must be divisible by num_heads ' \
+ f'(got `embed_dim`: {self.embed_dim} ' \
+ f'and `num_heads`: {self.num_heads}).'
+ self.scale = self.head_dim**(-0.5)
+ self.dropout = dropout
+
+ self.v_proj = nn.Linear(self.v_dim, self.embed_dim)
+ self.l_proj = nn.Linear(self.l_dim, self.embed_dim)
+ self.values_v_proj = nn.Linear(self.v_dim, self.embed_dim)
+ self.values_l_proj = nn.Linear(self.l_dim, self.embed_dim)
+
+ self.out_v_proj = nn.Linear(self.embed_dim, self.v_dim)
+ self.out_l_proj = nn.Linear(self.embed_dim, self.l_dim)
+
+ self.stable_softmax_2d = False
+ self.clamp_min_for_underflow = True
+ self.clamp_max_for_overflow = True
+
+ self._reset_parameters()
+
+ def _shape(self, tensor: Tensor, seq_len: int, bsz: int):
+ return tensor.view(bsz, seq_len, self.num_heads,
+ self.head_dim).transpose(1, 2).contiguous()
+
+ def _reset_parameters(self):
+ nn.init.xavier_uniform_(self.v_proj.weight)
+ self.v_proj.bias.data.fill_(0)
+ nn.init.xavier_uniform_(self.l_proj.weight)
+ self.l_proj.bias.data.fill_(0)
+ nn.init.xavier_uniform_(self.values_v_proj.weight)
+ self.values_v_proj.bias.data.fill_(0)
+ nn.init.xavier_uniform_(self.values_l_proj.weight)
+ self.values_l_proj.bias.data.fill_(0)
+ nn.init.xavier_uniform_(self.out_v_proj.weight)
+ self.out_v_proj.bias.data.fill_(0)
+ nn.init.xavier_uniform_(self.out_l_proj.weight)
+ self.out_l_proj.bias.data.fill_(0)
+
+ def forward(
+ self,
+ vision: Tensor,
+ lang: Tensor,
+ attention_mask_v: Optional[Tensor] = None,
+ attention_mask_l: Optional[Tensor] = None,
+ ) -> Tuple[Tensor, Tensor]:
+ bsz, tgt_len, _ = vision.size()
+
+ query_states = self.v_proj(vision) * self.scale
+ key_states = self._shape(self.l_proj(lang), -1, bsz)
+ value_v_states = self._shape(self.values_v_proj(vision), -1, bsz)
+ value_l_states = self._shape(self.values_l_proj(lang), -1, bsz)
+
+ proj_shape = (bsz * self.num_heads, -1, self.head_dim)
+ query_states = self._shape(query_states, tgt_len,
+ bsz).view(*proj_shape)
+ key_states = key_states.view(*proj_shape)
+ value_v_states = value_v_states.view(*proj_shape)
+ value_l_states = value_l_states.view(*proj_shape)
+
+ src_len = key_states.size(1)
+ attn_weights = torch.bmm(query_states, key_states.transpose(1, 2))
+
+ if attn_weights.size() != (bsz * self.num_heads, tgt_len, src_len):
+ raise ValueError(
+ f'Attention weights should be of '
+ f'size {(bsz * self.num_heads, tgt_len, src_len)}, '
+ f'but is {attn_weights.size()}')
+
+ if self.stable_softmax_2d:
+ attn_weights = attn_weights - attn_weights.max()
+
+ if self.clamp_min_for_underflow:
+ # Do not increase -50000, data type half has quite limited range
+ attn_weights = torch.clamp(attn_weights, min=-MAX_CLAMP_VALUE)
+ if self.clamp_max_for_overflow:
+ # Do not increase 50000, data type half has quite limited range
+ attn_weights = torch.clamp(attn_weights, max=MAX_CLAMP_VALUE)
+
+ attn_weights_T = attn_weights.transpose(1, 2)
+ attn_weights_l = (
+ attn_weights_T -
+ torch.max(attn_weights_T, dim=-1, keepdim=True)[0])
+ if self.clamp_min_for_underflow:
+ # Do not increase -50000, data type half has quite limited range
+ attn_weights_l = torch.clamp(attn_weights_l, min=-MAX_CLAMP_VALUE)
+ if self.clamp_max_for_overflow:
+ # Do not increase 50000, data type half has quite limited range
+ attn_weights_l = torch.clamp(attn_weights_l, max=MAX_CLAMP_VALUE)
+
+ if attention_mask_v is not None:
+ attention_mask_v = (
+ attention_mask_v[:, None,
+ None, :].repeat(1, self.num_heads, 1,
+ 1).flatten(0, 1))
+ attn_weights_l.masked_fill_(attention_mask_v, float('-inf'))
+
+ attn_weights_l = attn_weights_l.softmax(dim=-1)
+
+ if attention_mask_l is not None:
+ assert (attention_mask_l.dim() == 2)
+ attention_mask = attention_mask_l.unsqueeze(1).unsqueeze(1)
+ attention_mask = attention_mask.expand(bsz, 1, tgt_len, src_len)
+ attention_mask = attention_mask.masked_fill(
+ attention_mask == 0, -9e15)
+
+ if attention_mask.size() != (bsz, 1, tgt_len, src_len):
+ raise ValueError('Attention mask should be of '
+ f'size {(bsz, 1, tgt_len, src_len)}')
+ attn_weights = attn_weights.view(bsz, self.num_heads, tgt_len,
+ src_len) + attention_mask
+ attn_weights = attn_weights.view(bsz * self.num_heads, tgt_len,
+ src_len)
+
+ attn_weights_v = nn.functional.softmax(attn_weights, dim=-1)
+
+ attn_probs_v = F.dropout(
+ attn_weights_v, p=self.dropout, training=self.training)
+ attn_probs_l = F.dropout(
+ attn_weights_l, p=self.dropout, training=self.training)
+
+ attn_output_v = torch.bmm(attn_probs_v, value_l_states)
+ attn_output_l = torch.bmm(attn_probs_l, value_v_states)
+
+ if attn_output_v.size() != (bsz * self.num_heads, tgt_len,
+ self.head_dim):
+ raise ValueError(
+ '`attn_output_v` should be of '
+ f'size {(bsz, self.num_heads, tgt_len, self.head_dim)}, '
+ f'but is {attn_output_v.size()}')
+
+ if attn_output_l.size() != (bsz * self.num_heads, src_len,
+ self.head_dim):
+ raise ValueError(
+ '`attn_output_l` should be of size '
+ f'{(bsz, self.num_heads, src_len, self.head_dim)}, '
+ f'but is {attn_output_l.size()}')
+
+ attn_output_v = attn_output_v.view(bsz, self.num_heads, tgt_len,
+ self.head_dim)
+ attn_output_v = attn_output_v.transpose(1, 2)
+ attn_output_v = attn_output_v.reshape(bsz, tgt_len, self.embed_dim)
+
+ attn_output_l = attn_output_l.view(bsz, self.num_heads, src_len,
+ self.head_dim)
+ attn_output_l = attn_output_l.transpose(1, 2)
+ attn_output_l = attn_output_l.reshape(bsz, src_len, self.embed_dim)
+
+ attn_output_v = self.out_v_proj(attn_output_v)
+ attn_output_l = self.out_l_proj(attn_output_l)
+
+ return attn_output_v, attn_output_l
+
+
+class BiAttentionBlock(nn.Module):
+ """BiAttentionBlock Module:
+
+ First, multi-level visual features are concat; Then the concat visual
+ feature and lang feature are fused by attention; Finally the newly visual
+ feature are split into multi levels.
+
+ Args:
+ v_dim (int): The dimension of the visual features.
+ l_dim (int): The dimension of the language feature.
+ embed_dim (int): The embedding dimension for the attention operation.
+ num_heads (int): The number of attention heads.
+ dropout (float, optional): The dropout probability. Defaults to 0.1.
+ drop_path (float, optional): The drop path probability.
+ Defaults to 0.0.
+ init_values (float, optional):
+ The initial value for the scaling parameter.
+ Defaults to 1e-4.
+ """
+
+ def __init__(self,
+ v_dim: int,
+ l_dim: int,
+ embed_dim: int,
+ num_heads: int,
+ dropout: float = 0.1,
+ drop_path: float = .0,
+ init_values: float = 1e-4):
+ super().__init__()
+
+ # pre layer norm
+ self.layer_norm_v = nn.LayerNorm(v_dim)
+ self.layer_norm_l = nn.LayerNorm(l_dim)
+ self.attn = BiMultiHeadAttention(
+ v_dim=v_dim,
+ l_dim=l_dim,
+ embed_dim=embed_dim,
+ num_heads=num_heads,
+ dropout=dropout)
+
+ # add layer scale for training stability
+ self.drop_path = DropPath(
+ drop_path) if drop_path > 0. else nn.Identity()
+ self.gamma_v = nn.Parameter(
+ init_values * torch.ones(v_dim), requires_grad=True)
+ self.gamma_l = nn.Parameter(
+ init_values * torch.ones(l_dim), requires_grad=True)
+
+ def forward(self,
+ vf0: Tensor,
+ vf1: Tensor,
+ vf2: Tensor,
+ vf3: Tensor,
+ vf4: Tensor,
+ lang_feature: Tensor,
+ attention_mask_l=None):
+ visual_features = [vf0, vf1, vf2, vf3, vf4]
+ size_per_level, visual_features_flatten = [], []
+ for i, feat_per_level in enumerate(visual_features):
+ bs, c, h, w = feat_per_level.shape
+ size_per_level.append([h, w])
+ feat = permute_and_flatten(feat_per_level, bs, -1, c, h, w)
+ visual_features_flatten.append(feat)
+ visual_features_flatten = torch.cat(visual_features_flatten, dim=1)
+ new_v, new_lang_feature = self.single_attention_call(
+ visual_features_flatten,
+ lang_feature,
+ attention_mask_l=attention_mask_l)
+ # [bs, N, C] -> [bs, C, N]
+ new_v = new_v.transpose(1, 2).contiguous()
+
+ start = 0
+ # fvfs is mean fusion_visual_features
+ fvfs = []
+ for (h, w) in size_per_level:
+ new_v_per_level = new_v[:, :,
+ start:start + h * w].view(bs, -1, h,
+ w).contiguous()
+ fvfs.append(new_v_per_level)
+ start += h * w
+
+ return fvfs[0], fvfs[1], fvfs[2], fvfs[3], fvfs[4], new_lang_feature
+
+ def single_attention_call(
+ self,
+ visual: Tensor,
+ lang: Tensor,
+ attention_mask_v: Optional[Tensor] = None,
+ attention_mask_l: Optional[Tensor] = None,
+ ) -> Tuple[Tensor, Tensor]:
+ """Perform a single attention call between the visual and language
+ inputs.
+
+ Args:
+ visual (Tensor): The visual input tensor.
+ lang (Tensor): The language input tensor.
+ attention_mask_v (Optional[Tensor]):
+ An optional attention mask tensor for the visual input.
+ attention_mask_l (Optional[Tensor]):
+ An optional attention mask tensor for the language input.
+
+ Returns:
+ Tuple[Tensor, Tensor]: A tuple containing the updated
+ visual and language tensors after the attention call.
+ """
+ visual = self.layer_norm_v(visual)
+ lang = self.layer_norm_l(lang)
+ delta_v, delta_l = self.attn(
+ visual,
+ lang,
+ attention_mask_v=attention_mask_v,
+ attention_mask_l=attention_mask_l)
+ # visual, lang = visual + delta_v, l + delta_l
+ visual = visual + self.drop_path(self.gamma_v * delta_v)
+ lang = lang + self.drop_path(self.gamma_l * delta_l)
+ return visual, lang
+
+
+class SingleScaleBiAttentionBlock(BiAttentionBlock):
+ """This is a single-scale implementation of `BiAttentionBlock`.
+
+ The only differenece between it and `BiAttentionBlock` is that the
+ `forward` function of `SingleScaleBiAttentionBlock` only accepts a single
+ flatten visual feature map, while the `forward` function in
+ `BiAttentionBlock` accepts multiple visual feature maps.
+ """
+
+ def forward(self,
+ visual_feature: Tensor,
+ lang_feature: Tensor,
+ attention_mask_v=None,
+ attention_mask_l=None):
+ """Single-scale forward pass.
+
+ Args:
+ visual_feature (Tensor): The visual input tensor. Tensor of
+ shape (bs, patch_len, ch).
+ lang_feature (Tensor): The language input tensor. Tensor of
+ shape (bs, text_len, ch).
+ attention_mask_v (_type_, optional): Visual feature attention
+ mask. Defaults to None.
+ attention_mask_l (_type_, optional): Language feature attention
+ mask.Defaults to None.
+ """
+ new_v, new_lang_feature = self.single_attention_call(
+ visual_feature,
+ lang_feature,
+ attention_mask_v=attention_mask_v,
+ attention_mask_l=attention_mask_l)
+ return new_v, new_lang_feature
+
+
+class VLFuse(nn.Module):
+ """Early Fusion Module.
+
+ Args:
+ v_dim (int): Dimension of visual features.
+ l_dim (int): Dimension of language features.
+ embed_dim (int): The embedding dimension for the attention operation.
+ num_heads (int): Number of attention heads.
+ dropout (float): Dropout probability.
+ drop_path (float): Drop path probability.
+ use_checkpoint (bool): Whether to use PyTorch's checkpoint function.
+ """
+
+ def __init__(self,
+ v_dim: int = 256,
+ l_dim: int = 768,
+ embed_dim: int = 2048,
+ num_heads: int = 8,
+ dropout: float = 0.1,
+ drop_path: float = 0.0,
+ use_checkpoint: bool = False):
+ super().__init__()
+ self.use_checkpoint = use_checkpoint
+ self.b_attn = BiAttentionBlock(
+ v_dim=v_dim,
+ l_dim=l_dim,
+ embed_dim=embed_dim,
+ num_heads=num_heads,
+ dropout=dropout,
+ drop_path=drop_path,
+ init_values=1.0 / 6.0)
+
+ def forward(self, x: dict) -> dict:
+ """Forward pass of the VLFuse module."""
+ visual_features = x['visual']
+ language_dict_features = x['lang']
+
+ if self.use_checkpoint:
+ # vf is mean visual_features
+ # checkpoint does not allow complex data structures as input,
+ # such as list, so we must split them.
+ vf0, vf1, vf2, vf3, vf4, language_features = checkpoint.checkpoint(
+ self.b_attn, *visual_features,
+ language_dict_features['hidden'],
+ language_dict_features['masks'])
+ else:
+ vf0, vf1, vf2, vf3, vf4, language_features = self.b_attn(
+ *visual_features, language_dict_features['hidden'],
+ language_dict_features['masks'])
+
+ language_dict_features['hidden'] = language_features
+ fused_language_dict_features = language_dict_features
+
+ features_dict = {
+ 'visual': [vf0, vf1, vf2, vf3, vf4],
+ 'lang': fused_language_dict_features
+ }
+
+ return features_dict
+
+
+class BertEncoderLayer(BertPreTrainedModel):
+ """A modified version of the `BertLayer` class from the
+ `transformers.models.bert.modeling_bert` module.
+
+ Args:
+ config (:class:`~transformers.BertConfig`):
+ The configuration object that
+ contains various parameters for the model.
+ clamp_min_for_underflow (bool, optional):
+ Whether to clamp the minimum value of the hidden states
+ to prevent underflow. Defaults to `False`.
+ clamp_max_for_overflow (bool, optional):
+ Whether to clamp the maximum value of the hidden states
+ to prevent overflow. Defaults to `False`.
+ """
+
+ def __init__(self,
+ config: BertConfig,
+ clamp_min_for_underflow: bool = False,
+ clamp_max_for_overflow: bool = False):
+ super().__init__(config)
+ self.config = config
+ self.chunk_size_feed_forward = config.chunk_size_feed_forward
+ self.seq_len_dim = 1
+
+ self.attention = BertAttention(config, clamp_min_for_underflow,
+ clamp_max_for_overflow)
+ self.intermediate = BertIntermediate(config)
+ self.output = BertOutput(config)
+
+ def forward(
+ self, inputs: Dict[str, Dict[str, torch.Tensor]]
+ ) -> Dict[str, Dict[str, torch.Tensor]]:
+ """Applies the BertEncoderLayer to the input features."""
+ language_dict_features = inputs['lang']
+ hidden_states = language_dict_features['hidden']
+ attention_mask = language_dict_features['masks']
+
+ device = hidden_states.device
+ input_shape = hidden_states.size()[:-1]
+ extended_attention_mask = self.get_extended_attention_mask(
+ attention_mask, input_shape, device)
+
+ self_attention_outputs = self.attention(
+ hidden_states,
+ extended_attention_mask,
+ None,
+ output_attentions=False,
+ past_key_value=None)
+ attention_output = self_attention_outputs[0]
+ outputs = self_attention_outputs[1:]
+ layer_output = apply_chunking_to_forward(self.feed_forward_chunk,
+ self.chunk_size_feed_forward,
+ self.seq_len_dim,
+ attention_output)
+ outputs = (layer_output, ) + outputs
+ hidden_states = outputs[0]
+
+ language_dict_features['hidden'] = hidden_states
+
+ features_dict = {
+ 'visual': inputs['visual'],
+ 'lang': language_dict_features
+ }
+
+ return features_dict
+
+ def feed_forward_chunk(self, attention_output: Tensor) -> Tensor:
+ """Applies the intermediate and output layers of the BertEncoderLayer
+ to a chunk of the input sequence."""
+ intermediate_output = self.intermediate(attention_output)
+ layer_output = self.output(intermediate_output, attention_output)
+ return layer_output
+
+
+# The following code is the same as the Huggingface code,
+# with the only difference being the additional clamp operation.
+class BertSelfAttention(nn.Module):
+ """BERT self-attention layer from Huggingface transformers.
+
+ Compared to the BertSelfAttention of Huggingface, only add the clamp.
+
+ Args:
+ config (:class:`~transformers.BertConfig`):
+ The configuration object that
+ contains various parameters for the model.
+ clamp_min_for_underflow (bool, optional):
+ Whether to clamp the minimum value of the hidden states
+ to prevent underflow. Defaults to `False`.
+ clamp_max_for_overflow (bool, optional):
+ Whether to clamp the maximum value of the hidden states
+ to prevent overflow. Defaults to `False`.
+ """
+
+ def __init__(self,
+ config: BertConfig,
+ clamp_min_for_underflow: bool = False,
+ clamp_max_for_overflow: bool = False):
+ super().__init__()
+ if config.hidden_size % config.num_attention_heads != 0 and \
+ not hasattr(config, 'embedding_size'):
+ raise ValueError(f'The hidden size ({config.hidden_size}) is '
+ 'not a multiple of the number of attention '
+ f'heads ({config.num_attention_heads})')
+
+ self.num_attention_heads = config.num_attention_heads
+ self.attention_head_size = int(config.hidden_size /
+ config.num_attention_heads)
+ self.all_head_size = self.num_attention_heads * \
+ self.attention_head_size
+
+ self.query = nn.Linear(config.hidden_size, self.all_head_size)
+ self.key = nn.Linear(config.hidden_size, self.all_head_size)
+ self.value = nn.Linear(config.hidden_size, self.all_head_size)
+
+ self.dropout = nn.Dropout(config.attention_probs_dropout_prob)
+ self.position_embedding_type = getattr(config,
+ 'position_embedding_type',
+ 'absolute')
+ if self.position_embedding_type == 'relative_key' or \
+ self.position_embedding_type == 'relative_key_query':
+ self.max_position_embeddings = config.max_position_embeddings
+ self.distance_embedding = nn.Embedding(
+ 2 * config.max_position_embeddings - 1,
+ self.attention_head_size)
+ self.clamp_min_for_underflow = clamp_min_for_underflow
+ self.clamp_max_for_overflow = clamp_max_for_overflow
+
+ self.is_decoder = config.is_decoder
+
+ def transpose_for_scores(self, x: Tensor) -> Tensor:
+ """Transpose the dimensions of `x`."""
+ new_x_shape = x.size()[:-1] + (self.num_attention_heads,
+ self.attention_head_size)
+ x = x.view(*new_x_shape)
+ return x.permute(0, 2, 1, 3)
+
+ def forward(
+ self,
+ hidden_states: Tensor,
+ attention_mask: Optional[Tensor] = None,
+ head_mask: Optional[Tensor] = None,
+ encoder_hidden_states: Optional[Tensor] = None,
+ encoder_attention_mask: Optional[Tensor] = None,
+ past_key_value: Optional[Tuple[Tensor, Tensor]] = None,
+ output_attentions: bool = False,
+ ) -> Tuple[Tensor, ...]:
+ """Perform a forward pass through the BERT self-attention layer."""
+
+ mixed_query_layer = self.query(hidden_states)
+
+ # If this is instantiated as a cross-attention module, the keys
+ # and values come from an encoder; the attention mask needs to be
+ # such that the encoder's padding tokens are not attended to.
+ is_cross_attention = encoder_hidden_states is not None
+
+ if is_cross_attention and past_key_value is not None:
+ # reuse k,v, cross_attentions
+ key_layer = past_key_value[0]
+ value_layer = past_key_value[1]
+ attention_mask = encoder_attention_mask
+ elif is_cross_attention:
+ key_layer = self.transpose_for_scores(
+ self.key(encoder_hidden_states))
+ value_layer = self.transpose_for_scores(
+ self.value(encoder_hidden_states))
+ attention_mask = encoder_attention_mask
+ elif past_key_value is not None:
+ key_layer = self.transpose_for_scores(self.key(hidden_states))
+ value_layer = self.transpose_for_scores(self.value(hidden_states))
+ key_layer = torch.cat([past_key_value[0], key_layer], dim=2)
+ value_layer = torch.cat([past_key_value[1], value_layer], dim=2)
+ else:
+ key_layer = self.transpose_for_scores(self.key(hidden_states))
+ value_layer = self.transpose_for_scores(self.value(hidden_states))
+
+ query_layer = self.transpose_for_scores(mixed_query_layer)
+
+ if self.is_decoder:
+ past_key_value = (key_layer, value_layer)
+
+ # Take the dot product between "query" and "key"
+ # to get the raw attention scores.
+ attention_scores = torch.matmul(query_layer,
+ key_layer.transpose(-1, -2))
+
+ if self.position_embedding_type == 'relative_key' or \
+ self.position_embedding_type == 'relative_key_query':
+ seq_length = hidden_states.size()[1]
+ position_ids_l = torch.arange(
+ seq_length, dtype=torch.long,
+ device=hidden_states.device).view(-1, 1)
+ position_ids_r = torch.arange(
+ seq_length, dtype=torch.long,
+ device=hidden_states.device).view(1, -1)
+ distance = position_ids_l - position_ids_r
+ positional_embedding = self.distance_embedding(
+ distance + self.max_position_embeddings - 1)
+ positional_embedding = positional_embedding.to(
+ dtype=query_layer.dtype) # fp16 compatibility
+
+ if self.position_embedding_type == 'relative_key':
+ relative_position_scores = torch.einsum(
+ 'bhld,lrd->bhlr', query_layer, positional_embedding)
+ attention_scores = attention_scores + relative_position_scores
+ elif self.position_embedding_type == 'relative_key_query':
+ relative_position_scores_query = torch.einsum(
+ 'bhld,lrd->bhlr', query_layer, positional_embedding)
+ relative_position_scores_key = torch.einsum(
+ 'bhrd,lrd->bhlr', key_layer, positional_embedding)
+ attention_scores = attention_scores + \
+ relative_position_scores_query + \
+ relative_position_scores_key
+
+ attention_scores = attention_scores / math.sqrt(
+ self.attention_head_size)
+
+ if self.clamp_min_for_underflow:
+ attention_scores = torch.clamp(
+ attention_scores, min=-MAX_CLAMP_VALUE
+ ) # Do not increase -50000, data type half has quite limited range
+ if self.clamp_max_for_overflow:
+ attention_scores = torch.clamp(
+ attention_scores, max=MAX_CLAMP_VALUE
+ ) # Do not increase 50000, data type half has quite limited range
+
+ if attention_mask is not None:
+ # Apply the attention mask is
+ # (precomputed for all layers in BertModel forward() function)
+ attention_scores = attention_scores + attention_mask
+
+ # Normalize the attention scores to probabilities.
+ attention_probs = nn.Softmax(dim=-1)(attention_scores)
+
+ # This is actually dropping out entire tokens to attend to, which might
+ # seem a bit unusual, but is taken from the original Transformer paper.
+ attention_probs = self.dropout(attention_probs)
+
+ # Mask heads if we want to
+ if head_mask is not None:
+ attention_probs = attention_probs * head_mask
+
+ context_layer = torch.matmul(attention_probs, value_layer)
+
+ context_layer = context_layer.permute(0, 2, 1, 3).contiguous()
+ new_context_layer_shape = context_layer.size()[:-2] + (
+ self.all_head_size, )
+ context_layer = context_layer.view(*new_context_layer_shape)
+
+ outputs = (context_layer,
+ attention_probs) if output_attentions else (context_layer, )
+
+ if self.is_decoder:
+ outputs = outputs + (past_key_value, )
+ return outputs
+
+
+class BertAttention(HFBertAttention):
+ """BertAttention is made up of self-attention and intermediate+output.
+
+ Compared to the BertAttention of Huggingface, only add the clamp.
+
+ Args:
+ config (:class:`~transformers.BertConfig`):
+ The configuration object that
+ contains various parameters for the model.
+ clamp_min_for_underflow (bool, optional):
+ Whether to clamp the minimum value of the hidden states
+ to prevent underflow. Defaults to `False`.
+ clamp_max_for_overflow (bool, optional):
+ Whether to clamp the maximum value of the hidden states
+ to prevent overflow. Defaults to `False`.
+ """
+
+ def __init__(self,
+ config: BertConfig,
+ clamp_min_for_underflow: bool = False,
+ clamp_max_for_overflow: bool = False):
+ super().__init__(config)
+ self.self = BertSelfAttention(config, clamp_min_for_underflow,
+ clamp_max_for_overflow)
+
+
+class BertIntermediate(HFBertIntermediate):
+ """Modified from transformers.models.bert.modeling_bert.BertIntermediate.
+
+ Compared to the BertIntermediate of Huggingface, only add the clamp.
+ """
+
+ def forward(self, hidden_states: Tensor) -> Tensor:
+ hidden_states = self.dense(hidden_states)
+ hidden_states = clamp_values(hidden_states)
+ hidden_states = self.intermediate_act_fn(hidden_states)
+ hidden_states = clamp_values(hidden_states)
+ return hidden_states
+
+
+class BertOutput(HFBertOutput):
+ """Modified from transformers.models.bert.modeling_bert.BertOutput.
+
+ Compared to the BertOutput of Huggingface, only add the clamp.
+ """
+
+ def forward(self, hidden_states: Tensor, input_tensor: Tensor) -> Tensor:
+ hidden_states = self.dense(hidden_states)
+ hidden_states = self.dropout(hidden_states)
+ hidden_states = clamp_values(hidden_states)
+ hidden_states = self.LayerNorm(hidden_states + input_tensor)
+ hidden_states = clamp_values(hidden_states)
+ return hidden_states
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/wbf.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/wbf.py
new file mode 100644
index 0000000000000000000000000000000000000000..b26a2c669a520467c6fcf52d0eec53a69834a16a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/utils/wbf.py
@@ -0,0 +1,250 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+import warnings
+from typing import Tuple
+
+import numpy as np
+import torch
+from torch import Tensor
+
+
+# References: https://github.com/ZFTurbo/Weighted-Boxes-Fusion
+def weighted_boxes_fusion(
+ bboxes_list: list,
+ scores_list: list,
+ labels_list: list,
+ weights: list = None,
+ iou_thr: float = 0.55,
+ skip_box_thr: float = 0.0,
+ conf_type: str = 'avg',
+ allows_overflow: bool = False) -> Tuple[Tensor, Tensor, Tensor]:
+ """weighted boxes fusion is a method for
+ fusing predictions from different object detection models, which utilizes
+ confidence scores of all proposed bounding boxes to construct averaged
+ boxes.
+
+ Args:
+ bboxes_list(list): list of boxes predictions from each model,
+ each box is 4 numbers.
+ scores_list(list): list of scores for each model
+ labels_list(list): list of labels for each model
+ weights: list of weights for each model.
+ Default: None, which means weight == 1 for each model
+ iou_thr: IoU value for boxes to be a match
+ skip_box_thr: exclude boxes with score lower than this variable.
+ conf_type: how to calculate confidence in weighted boxes.
+ 'avg': average value,
+ 'max': maximum value,
+ 'box_and_model_avg': box and model wise hybrid weighted average,
+ 'absent_model_aware_avg': weighted average that takes into
+ account the absent model.
+ allows_overflow: false if we want confidence score not exceed 1.0.
+
+ Returns:
+ bboxes(Tensor): boxes coordinates (Order of boxes: x1, y1, x2, y2).
+ scores(Tensor): confidence scores
+ labels(Tensor): boxes labels
+ """
+
+ if weights is None:
+ weights = np.ones(len(bboxes_list))
+ if len(weights) != len(bboxes_list):
+ print('Warning: incorrect number of weights {}. Must be: '
+ '{}. Set weights equal to 1.'.format(
+ len(weights), len(bboxes_list)))
+ weights = np.ones(len(bboxes_list))
+ weights = np.array(weights)
+
+ if conf_type not in [
+ 'avg', 'max', 'box_and_model_avg', 'absent_model_aware_avg'
+ ]:
+ print('Unknown conf_type: {}. Must be "avg", '
+ '"max" or "box_and_model_avg", '
+ 'or "absent_model_aware_avg"'.format(conf_type))
+ exit()
+
+ filtered_boxes = prefilter_boxes(bboxes_list, scores_list, labels_list,
+ weights, skip_box_thr)
+ if len(filtered_boxes) == 0:
+ return torch.Tensor(), torch.Tensor(), torch.Tensor()
+
+ overall_boxes = []
+
+ for label in filtered_boxes:
+ boxes = filtered_boxes[label]
+ new_boxes = []
+ weighted_boxes = np.empty((0, 8))
+
+ # Clusterize boxes
+ for j in range(0, len(boxes)):
+ index, best_iou = find_matching_box_fast(weighted_boxes, boxes[j],
+ iou_thr)
+
+ if index != -1:
+ new_boxes[index].append(boxes[j])
+ weighted_boxes[index] = get_weighted_box(
+ new_boxes[index], conf_type)
+ else:
+ new_boxes.append([boxes[j].copy()])
+ weighted_boxes = np.vstack((weighted_boxes, boxes[j].copy()))
+
+ # Rescale confidence based on number of models and boxes
+ for i in range(len(new_boxes)):
+ clustered_boxes = new_boxes[i]
+ if conf_type == 'box_and_model_avg':
+ clustered_boxes = np.array(clustered_boxes)
+ # weighted average for boxes
+ weighted_boxes[i, 1] = weighted_boxes[i, 1] * len(
+ clustered_boxes) / weighted_boxes[i, 2]
+ # identify unique model index by model index column
+ _, idx = np.unique(clustered_boxes[:, 3], return_index=True)
+ # rescale by unique model weights
+ weighted_boxes[i, 1] = weighted_boxes[i, 1] * clustered_boxes[
+ idx, 2].sum() / weights.sum()
+ elif conf_type == 'absent_model_aware_avg':
+ clustered_boxes = np.array(clustered_boxes)
+ # get unique model index in the cluster
+ models = np.unique(clustered_boxes[:, 3]).astype(int)
+ # create a mask to get unused model weights
+ mask = np.ones(len(weights), dtype=bool)
+ mask[models] = False
+ # absent model aware weighted average
+ weighted_boxes[
+ i, 1] = weighted_boxes[i, 1] * len(clustered_boxes) / (
+ weighted_boxes[i, 2] + weights[mask].sum())
+ elif conf_type == 'max':
+ weighted_boxes[i, 1] = weighted_boxes[i, 1] / weights.max()
+ elif not allows_overflow:
+ weighted_boxes[i, 1] = weighted_boxes[i, 1] * min(
+ len(weights), len(clustered_boxes)) / weights.sum()
+ else:
+ weighted_boxes[i, 1] = weighted_boxes[i, 1] * len(
+ clustered_boxes) / weights.sum()
+ overall_boxes.append(weighted_boxes)
+ overall_boxes = np.concatenate(overall_boxes, axis=0)
+ overall_boxes = overall_boxes[overall_boxes[:, 1].argsort()[::-1]]
+
+ bboxes = torch.Tensor(overall_boxes[:, 4:])
+ scores = torch.Tensor(overall_boxes[:, 1])
+ labels = torch.Tensor(overall_boxes[:, 0]).int()
+
+ return bboxes, scores, labels
+
+
+def prefilter_boxes(boxes, scores, labels, weights, thr):
+
+ new_boxes = dict()
+
+ for t in range(len(boxes)):
+
+ if len(boxes[t]) != len(scores[t]):
+ print('Error. Length of boxes arrays not equal to '
+ 'length of scores array: {} != {}'.format(
+ len(boxes[t]), len(scores[t])))
+ exit()
+
+ if len(boxes[t]) != len(labels[t]):
+ print('Error. Length of boxes arrays not equal to '
+ 'length of labels array: {} != {}'.format(
+ len(boxes[t]), len(labels[t])))
+ exit()
+
+ for j in range(len(boxes[t])):
+ score = scores[t][j]
+ if score < thr:
+ continue
+ label = int(labels[t][j])
+ box_part = boxes[t][j]
+ x1 = float(box_part[0])
+ y1 = float(box_part[1])
+ x2 = float(box_part[2])
+ y2 = float(box_part[3])
+
+ # Box data checks
+ if x2 < x1:
+ warnings.warn('X2 < X1 value in box. Swap them.')
+ x1, x2 = x2, x1
+ if y2 < y1:
+ warnings.warn('Y2 < Y1 value in box. Swap them.')
+ y1, y2 = y2, y1
+ if (x2 - x1) * (y2 - y1) == 0.0:
+ warnings.warn('Zero area box skipped: {}.'.format(box_part))
+ continue
+
+ # [label, score, weight, model index, x1, y1, x2, y2]
+ b = [
+ int(label),
+ float(score) * weights[t], weights[t], t, x1, y1, x2, y2
+ ]
+
+ if label not in new_boxes:
+ new_boxes[label] = []
+ new_boxes[label].append(b)
+
+ # Sort each list in dict by score and transform it to numpy array
+ for k in new_boxes:
+ current_boxes = np.array(new_boxes[k])
+ new_boxes[k] = current_boxes[current_boxes[:, 1].argsort()[::-1]]
+
+ return new_boxes
+
+
+def get_weighted_box(boxes, conf_type='avg'):
+
+ box = np.zeros(8, dtype=np.float32)
+ conf = 0
+ conf_list = []
+ w = 0
+ for b in boxes:
+ box[4:] += (b[1] * b[4:])
+ conf += b[1]
+ conf_list.append(b[1])
+ w += b[2]
+ box[0] = boxes[0][0]
+ if conf_type in ('avg', 'box_and_model_avg', 'absent_model_aware_avg'):
+ box[1] = conf / len(boxes)
+ elif conf_type == 'max':
+ box[1] = np.array(conf_list).max()
+ box[2] = w
+ box[3] = -1
+ box[4:] /= conf
+
+ return box
+
+
+def find_matching_box_fast(boxes_list, new_box, match_iou):
+
+ def bb_iou_array(boxes, new_box):
+ # bb intersection over union
+ xA = np.maximum(boxes[:, 0], new_box[0])
+ yA = np.maximum(boxes[:, 1], new_box[1])
+ xB = np.minimum(boxes[:, 2], new_box[2])
+ yB = np.minimum(boxes[:, 3], new_box[3])
+
+ interArea = np.maximum(xB - xA, 0) * np.maximum(yB - yA, 0)
+
+ # compute the area of both the prediction and ground-truth rectangles
+ boxAArea = (boxes[:, 2] - boxes[:, 0]) * (boxes[:, 3] - boxes[:, 1])
+ boxBArea = (new_box[2] - new_box[0]) * (new_box[3] - new_box[1])
+
+ iou = interArea / (boxAArea + boxBArea - interArea)
+
+ return iou
+
+ if boxes_list.shape[0] == 0:
+ return -1, match_iou
+
+ boxes = boxes_list
+
+ ious = bb_iou_array(boxes[:, 4:], new_box[4:])
+
+ ious[boxes[:, 0] != new_box[0]] = -1
+
+ best_idx = np.argmax(ious)
+ best_iou = ious[best_idx]
+
+ if best_iou <= match_iou:
+ best_iou = match_iou
+ best_idx = -1
+
+ return best_idx, best_iou
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/vis/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/vis/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..ab63a9066bcf6cd25d7c9063cc66d9b0390b3d42
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/vis/__init__.py
@@ -0,0 +1,5 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .mask2former_vis import Mask2FormerVideo
+from .masktrack_rcnn import MaskTrackRCNN
+
+__all__ = ['Mask2FormerVideo', 'MaskTrackRCNN']
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/vis/mask2former_vis.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/vis/mask2former_vis.py
new file mode 100644
index 0000000000000000000000000000000000000000..6ab04296e120622f4b5e28739f4c3323d253f7d5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/vis/mask2former_vis.py
@@ -0,0 +1,120 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Union
+
+from torch import Tensor
+
+from mmdet.models.mot import BaseMOTModel
+from mmdet.registry import MODELS
+from mmdet.structures import TrackDataSample, TrackSampleList
+from mmdet.utils import OptConfigType, OptMultiConfig
+
+
+@MODELS.register_module()
+class Mask2FormerVideo(BaseMOTModel):
+ r"""Implementation of `Masked-attention Mask
+ Transformer for Universal Image Segmentation
+ `_.
+
+ Args:
+ backbone (dict): Configuration of backbone. Defaults to None.
+ track_head (dict): Configuration of track head. Defaults to None.
+ data_preprocessor (dict or ConfigDict, optional): The pre-process
+ config of :class:`TrackDataPreprocessor`. it usually includes,
+ ``pad_size_divisor``, ``pad_value``, ``mean`` and ``std``.
+ Defaults to None.
+ init_cfg (dict or list[dict]): Configuration of initialization.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ backbone: Optional[dict] = None,
+ track_head: Optional[dict] = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None):
+ super(BaseMOTModel, self).__init__(
+ data_preprocessor=data_preprocessor, init_cfg=init_cfg)
+
+ if backbone is not None:
+ self.backbone = MODELS.build(backbone)
+
+ if track_head is not None:
+ self.track_head = MODELS.build(track_head)
+
+ self.num_classes = self.track_head.num_classes
+
+ def _load_from_state_dict(self, state_dict, prefix, local_metadata, strict,
+ missing_keys, unexpected_keys, error_msgs):
+ """Overload in order to load mmdet pretrained ckpt."""
+ for key in list(state_dict):
+ if key.startswith('panoptic_head'):
+ state_dict[key.replace('panoptic',
+ 'track')] = state_dict.pop(key)
+
+ super()._load_from_state_dict(state_dict, prefix, local_metadata,
+ strict, missing_keys, unexpected_keys,
+ error_msgs)
+
+ def loss(self, inputs: Tensor, data_samples: TrackSampleList,
+ **kwargs) -> Union[dict, tuple]:
+ """
+ Args:
+ inputs (Tensor): Input images of shape (N, T, C, H, W).
+ These should usually be mean centered and std scaled.
+ data_samples (list[:obj:`TrackDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance`.
+
+ Returns:
+ dict[str, Tensor]: a dictionary of loss components
+ """
+ assert inputs.dim() == 5, 'The img must be 5D Tensor (N, T, C, H, W).'
+ # shape (N * T, C, H, W)
+ img = inputs.flatten(0, 1)
+
+ x = self.backbone(img)
+ losses = self.track_head.loss(x, data_samples)
+
+ return losses
+
+ def predict(self,
+ inputs: Tensor,
+ data_samples: TrackSampleList,
+ rescale: bool = True) -> TrackSampleList:
+ """Predict results from a batch of inputs and data samples with
+ postprocessing.
+
+ Args:
+ inputs (Tensor): of shape (N, T, C, H, W) encoding
+ input images. The N denotes batch size.
+ The T denotes the number of frames in a video.
+ data_samples (list[:obj:`TrackDataSample`]): The batch
+ data samples. It usually includes information such
+ as `video_data_samples`.
+ rescale (bool, Optional): If False, then returned bboxes and masks
+ will fit the scale of img, otherwise, returned bboxes and masks
+ will fit the scale of original image shape. Defaults to True.
+
+ Returns:
+ TrackSampleList: Tracking results of the inputs.
+ """
+ assert inputs.dim() == 5, 'The img must be 5D Tensor (N, T, C, H, W).'
+
+ assert len(data_samples) == 1, \
+ 'Mask2former only support 1 batch size per gpu for now.'
+
+ # [T, C, H, W]
+ img = inputs[0]
+ track_data_sample = data_samples[0]
+ feats = self.backbone(img)
+ pred_track_ins_list = self.track_head.predict(feats, track_data_sample,
+ rescale)
+
+ det_data_samples_list = []
+ for idx, pred_track_ins in enumerate(pred_track_ins_list):
+ img_data_sample = track_data_sample[idx]
+ img_data_sample.pred_track_instances = pred_track_ins
+ det_data_samples_list.append(img_data_sample)
+
+ results = TrackDataSample()
+ results.video_data_samples = det_data_samples_list
+ return [results]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/vis/masktrack_rcnn.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/vis/masktrack_rcnn.py
new file mode 100644
index 0000000000000000000000000000000000000000..9c28e7b8529d3d53d5a59ecff0ea46662d035f23
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/models/vis/masktrack_rcnn.py
@@ -0,0 +1,181 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional
+
+import torch
+from torch import Tensor
+
+from mmdet.models.mot import BaseMOTModel
+from mmdet.registry import MODELS
+from mmdet.structures import TrackSampleList
+from mmdet.utils import OptConfigType, OptMultiConfig
+
+
+@MODELS.register_module()
+class MaskTrackRCNN(BaseMOTModel):
+ """Video Instance Segmentation.
+
+ This video instance segmentor is the implementation of`MaskTrack R-CNN
+ `_.
+
+ Args:
+ detector (dict): Configuration of detector. Defaults to None.
+ track_head (dict): Configuration of track head. Defaults to None.
+ tracker (dict): Configuration of tracker. Defaults to None.
+ data_preprocessor (dict or ConfigDict, optional): The pre-process
+ config of :class:`TrackDataPreprocessor`. it usually includes,
+ ``pad_size_divisor``, ``pad_value``, ``mean`` and ``std``.
+ init_cfg (dict or list[dict]): Configuration of initialization.
+ Defaults to None.
+ """
+
+ def __init__(self,
+ detector: Optional[dict] = None,
+ track_head: Optional[dict] = None,
+ tracker: Optional[dict] = None,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None):
+ super().__init__(data_preprocessor, init_cfg)
+
+ if detector is not None:
+ self.detector = MODELS.build(detector)
+ assert hasattr(self.detector, 'roi_head'), \
+ 'MaskTrack R-CNN only supports two stage detectors.'
+
+ if track_head is not None:
+ self.track_head = MODELS.build(track_head)
+ if tracker is not None:
+ self.tracker = MODELS.build(tracker)
+
+ def loss(self, inputs: Tensor, data_samples: TrackSampleList,
+ **kwargs) -> dict:
+ """Calculate losses from a batch of inputs and data samples.
+
+ Args:
+ inputs (Dict[str, Tensor]): of shape (N, T, C, H, W) encoding
+ input images. Typically these should be mean centered and std
+ scaled. The N denotes batch size. The T denotes the number of
+ frames.
+ data_samples (list[:obj:`TrackDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance`.
+
+ Returns:
+ dict: A dictionary of loss components.
+ """
+
+ assert inputs.dim() == 5, 'The img must be 5D Tensor (N, T, C, H, W).'
+ assert inputs.size(1) == 2, \
+ 'MaskTrackRCNN can only have 1 key frame and 1 reference frame.'
+
+ # split the data_samples into two aspects: key frames and reference
+ # frames
+ ref_data_samples, key_data_samples = [], []
+ key_frame_inds, ref_frame_inds = [], []
+
+ # set cat_id of gt_labels to 0 in RPN
+ for track_data_sample in data_samples:
+ key_data_sample = track_data_sample.get_key_frames()[0]
+ key_data_samples.append(key_data_sample)
+ ref_data_sample = track_data_sample.get_ref_frames()[0]
+ ref_data_samples.append(ref_data_sample)
+ key_frame_inds.append(track_data_sample.key_frames_inds[0])
+ ref_frame_inds.append(track_data_sample.ref_frames_inds[0])
+
+ key_frame_inds = torch.tensor(key_frame_inds, dtype=torch.int64)
+ ref_frame_inds = torch.tensor(ref_frame_inds, dtype=torch.int64)
+ batch_inds = torch.arange(len(inputs))
+ key_imgs = inputs[batch_inds, key_frame_inds].contiguous()
+ ref_imgs = inputs[batch_inds, ref_frame_inds].contiguous()
+
+ x = self.detector.extract_feat(key_imgs)
+ ref_x = self.detector.extract_feat(ref_imgs)
+
+ losses = dict()
+
+ # RPN forward and loss
+ if self.detector.with_rpn:
+ proposal_cfg = self.detector.train_cfg.get(
+ 'rpn_proposal', self.detector.test_cfg.rpn)
+
+ rpn_losses, rpn_results_list = self.detector.rpn_head. \
+ loss_and_predict(x,
+ key_data_samples,
+ proposal_cfg=proposal_cfg,
+ **kwargs)
+
+ # avoid get same name with roi_head loss
+ keys = rpn_losses.keys()
+ for key in keys:
+ if 'loss' in key and 'rpn' not in key:
+ rpn_losses[f'rpn_{key}'] = rpn_losses.pop(key)
+ losses.update(rpn_losses)
+ else:
+ # TODO: Not support currently, should have a check at Fast R-CNN
+ assert key_data_samples[0].get('proposals', None) is not None
+ # use pre-defined proposals in InstanceData for the second stage
+ # to extract ROI features.
+ rpn_results_list = [
+ key_data_sample.proposals
+ for key_data_sample in key_data_samples
+ ]
+
+ losses_detect = self.detector.roi_head.loss(x, rpn_results_list,
+ key_data_samples, **kwargs)
+ losses.update(losses_detect)
+
+ losses_track = self.track_head.loss(x, ref_x, rpn_results_list,
+ data_samples, **kwargs)
+ losses.update(losses_track)
+
+ return losses
+
+ def predict(self,
+ inputs: Tensor,
+ data_samples: TrackSampleList,
+ rescale: bool = True,
+ **kwargs) -> TrackSampleList:
+ """Test without augmentation.
+
+ Args:
+ inputs (Tensor): of shape (N, T, C, H, W) encoding
+ input images. The N denotes batch size.
+ The T denotes the number of frames in a video.
+ data_samples (list[:obj:`TrackDataSample`]): The batch
+ data samples. It usually includes information such
+ as `video_data_samples`.
+ rescale (bool, Optional): If False, then returned bboxes and masks
+ will fit the scale of img, otherwise, returned bboxes and masks
+ will fit the scale of original image shape. Defaults to True.
+
+ Returns:
+ TrackSampleList: Tracking results of the inputs.
+ """
+ assert inputs.dim() == 5, 'The img must be 5D Tensor (N, T, C, H, W).'
+
+ assert len(data_samples) == 1, \
+ 'MaskTrackRCNN only support 1 batch size per gpu for now.'
+
+ track_data_sample = data_samples[0]
+ video_len = len(track_data_sample)
+ if track_data_sample[0].frame_id == 0:
+ self.tracker.reset()
+
+ for frame_id in range(video_len):
+ img_data_sample = track_data_sample[frame_id]
+ single_img = inputs[:, frame_id].contiguous()
+ x = self.detector.extract_feat(single_img)
+
+ rpn_results_list = self.detector.rpn_head.predict(
+ x, [img_data_sample])
+ # det_results List[InstanceData]
+ det_results = self.detector.roi_head.predict(
+ x, rpn_results_list, [img_data_sample], rescale=rescale)
+ assert len(det_results) == 1, 'Batch inference is not supported.'
+ assert 'masks' in det_results[0], 'There are no mask results.'
+
+ img_data_sample.pred_instances = det_results[0]
+ frame_pred_track_instances = self.tracker.track(
+ model=self, feats=x, data_sample=img_data_sample, **kwargs)
+ img_data_sample.pred_track_instances = frame_pred_track_instances
+
+ return [track_data_sample]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/registry.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/registry.py
new file mode 100644
index 0000000000000000000000000000000000000000..3a5b2b28a4f80a488994b48a99043a20c604e55e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/registry.py
@@ -0,0 +1,121 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+"""MMDetection provides 17 registry nodes to support using modules across
+projects. Each node is a child of the root registry in MMEngine.
+
+More details can be found at
+https://mmengine.readthedocs.io/en/latest/tutorials/registry.html.
+"""
+
+from mmengine.registry import DATA_SAMPLERS as MMENGINE_DATA_SAMPLERS
+from mmengine.registry import DATASETS as MMENGINE_DATASETS
+from mmengine.registry import EVALUATOR as MMENGINE_EVALUATOR
+from mmengine.registry import HOOKS as MMENGINE_HOOKS
+from mmengine.registry import LOG_PROCESSORS as MMENGINE_LOG_PROCESSORS
+from mmengine.registry import LOOPS as MMENGINE_LOOPS
+from mmengine.registry import METRICS as MMENGINE_METRICS
+from mmengine.registry import MODEL_WRAPPERS as MMENGINE_MODEL_WRAPPERS
+from mmengine.registry import MODELS as MMENGINE_MODELS
+from mmengine.registry import \
+ OPTIM_WRAPPER_CONSTRUCTORS as MMENGINE_OPTIM_WRAPPER_CONSTRUCTORS
+from mmengine.registry import OPTIM_WRAPPERS as MMENGINE_OPTIM_WRAPPERS
+from mmengine.registry import OPTIMIZERS as MMENGINE_OPTIMIZERS
+from mmengine.registry import PARAM_SCHEDULERS as MMENGINE_PARAM_SCHEDULERS
+from mmengine.registry import \
+ RUNNER_CONSTRUCTORS as MMENGINE_RUNNER_CONSTRUCTORS
+from mmengine.registry import RUNNERS as MMENGINE_RUNNERS
+from mmengine.registry import TASK_UTILS as MMENGINE_TASK_UTILS
+from mmengine.registry import TRANSFORMS as MMENGINE_TRANSFORMS
+from mmengine.registry import VISBACKENDS as MMENGINE_VISBACKENDS
+from mmengine.registry import VISUALIZERS as MMENGINE_VISUALIZERS
+from mmengine.registry import \
+ WEIGHT_INITIALIZERS as MMENGINE_WEIGHT_INITIALIZERS
+from mmengine.registry import Registry
+
+# manage all kinds of runners like `EpochBasedRunner` and `IterBasedRunner`
+RUNNERS = Registry(
+ 'runner', parent=MMENGINE_RUNNERS, locations=['mmdet.engine.runner'])
+# manage runner constructors that define how to initialize runners
+RUNNER_CONSTRUCTORS = Registry(
+ 'runner constructor',
+ parent=MMENGINE_RUNNER_CONSTRUCTORS,
+ locations=['mmdet.engine.runner'])
+# manage all kinds of loops like `EpochBasedTrainLoop`
+LOOPS = Registry(
+ 'loop', parent=MMENGINE_LOOPS, locations=['mmdet.engine.runner'])
+# manage all kinds of hooks like `CheckpointHook`
+HOOKS = Registry(
+ 'hook', parent=MMENGINE_HOOKS, locations=['mmdet.engine.hooks'])
+
+# manage data-related modules
+DATASETS = Registry(
+ 'dataset', parent=MMENGINE_DATASETS, locations=['mmdet.datasets'])
+DATA_SAMPLERS = Registry(
+ 'data sampler',
+ parent=MMENGINE_DATA_SAMPLERS,
+ locations=['mmdet.datasets.samplers'])
+TRANSFORMS = Registry(
+ 'transform',
+ parent=MMENGINE_TRANSFORMS,
+ locations=['mmdet.datasets.transforms'])
+
+# manage all kinds of modules inheriting `nn.Module`
+MODELS = Registry('model', parent=MMENGINE_MODELS, locations=['mmdet.models'])
+# manage all kinds of model wrappers like 'MMDistributedDataParallel'
+MODEL_WRAPPERS = Registry(
+ 'model_wrapper',
+ parent=MMENGINE_MODEL_WRAPPERS,
+ locations=['mmdet.models'])
+# manage all kinds of weight initialization modules like `Uniform`
+WEIGHT_INITIALIZERS = Registry(
+ 'weight initializer',
+ parent=MMENGINE_WEIGHT_INITIALIZERS,
+ locations=['mmdet.models'])
+
+# manage all kinds of optimizers like `SGD` and `Adam`
+OPTIMIZERS = Registry(
+ 'optimizer',
+ parent=MMENGINE_OPTIMIZERS,
+ locations=['mmdet.engine.optimizers'])
+# manage optimizer wrapper
+OPTIM_WRAPPERS = Registry(
+ 'optim_wrapper',
+ parent=MMENGINE_OPTIM_WRAPPERS,
+ locations=['mmdet.engine.optimizers'])
+# manage constructors that customize the optimization hyperparameters.
+OPTIM_WRAPPER_CONSTRUCTORS = Registry(
+ 'optimizer constructor',
+ parent=MMENGINE_OPTIM_WRAPPER_CONSTRUCTORS,
+ locations=['mmdet.engine.optimizers'])
+# manage all kinds of parameter schedulers like `MultiStepLR`
+PARAM_SCHEDULERS = Registry(
+ 'parameter scheduler',
+ parent=MMENGINE_PARAM_SCHEDULERS,
+ locations=['mmdet.engine.schedulers'])
+# manage all kinds of metrics
+METRICS = Registry(
+ 'metric', parent=MMENGINE_METRICS, locations=['mmdet.evaluation'])
+# manage evaluator
+EVALUATOR = Registry(
+ 'evaluator', parent=MMENGINE_EVALUATOR, locations=['mmdet.evaluation'])
+
+# manage task-specific modules like anchor generators and box coders
+TASK_UTILS = Registry(
+ 'task util', parent=MMENGINE_TASK_UTILS, locations=['mmdet.models'])
+
+# manage visualizer
+VISUALIZERS = Registry(
+ 'visualizer',
+ parent=MMENGINE_VISUALIZERS,
+ locations=['mmdet.visualization'])
+# manage visualizer backend
+VISBACKENDS = Registry(
+ 'vis_backend',
+ parent=MMENGINE_VISBACKENDS,
+ locations=['mmdet.visualization'])
+
+# manage logprocessor
+LOG_PROCESSORS = Registry(
+ 'log_processor',
+ parent=MMENGINE_LOG_PROCESSORS,
+ # TODO: update the location when mmdet has its own log processor
+ locations=['mmdet.engine'])
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..381c6a4f4549c2c4395d994cbd860a3e52eb9994
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/__init__.py
@@ -0,0 +1,10 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .det_data_sample import DetDataSample, OptSampleList, SampleList
+from .reid_data_sample import ReIDDataSample
+from .track_data_sample import (OptTrackSampleList, TrackDataSample,
+ TrackSampleList)
+
+__all__ = [
+ 'DetDataSample', 'SampleList', 'OptSampleList', 'TrackDataSample',
+ 'TrackSampleList', 'OptTrackSampleList', 'ReIDDataSample'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/bbox/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/bbox/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..4d531986509ad1b2141118449aab39343bbde82c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/bbox/__init__.py
@@ -0,0 +1,25 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .base_boxes import BaseBoxes
+from .bbox_overlaps import bbox_overlaps
+from .box_type import (autocast_box_type, convert_box_type, get_box_type,
+ register_box, register_box_converter)
+from .horizontal_boxes import HorizontalBoxes
+from .transforms import bbox_cxcyah_to_xyxy # noqa: E501
+from .transforms import (bbox2corner, bbox2distance, bbox2result, bbox2roi,
+ bbox_cxcywh_to_xyxy, bbox_flip, bbox_mapping,
+ bbox_mapping_back, bbox_project, bbox_rescale,
+ bbox_xyxy_to_cxcyah, bbox_xyxy_to_cxcywh, cat_boxes,
+ corner2bbox, distance2bbox, empty_box_as,
+ find_inside_bboxes, get_box_tensor, get_box_wh,
+ roi2bbox, scale_boxes, stack_boxes)
+
+__all__ = [
+ 'bbox_overlaps', 'bbox_flip', 'bbox_mapping', 'bbox_mapping_back',
+ 'bbox2roi', 'roi2bbox', 'bbox2result', 'distance2bbox', 'bbox2distance',
+ 'bbox_rescale', 'bbox_cxcywh_to_xyxy', 'bbox_xyxy_to_cxcywh',
+ 'find_inside_bboxes', 'bbox2corner', 'corner2bbox', 'bbox_project',
+ 'BaseBoxes', 'convert_box_type', 'get_box_type', 'register_box',
+ 'register_box_converter', 'HorizontalBoxes', 'autocast_box_type',
+ 'cat_boxes', 'stack_boxes', 'scale_boxes', 'get_box_wh', 'get_box_tensor',
+ 'empty_box_as', 'bbox_xyxy_to_cxcyah', 'bbox_cxcyah_to_xyxy'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/bbox/base_boxes.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/bbox/base_boxes.py
new file mode 100644
index 0000000000000000000000000000000000000000..0ed667664a8a57a1b9b7e422af03d41274882747
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/bbox/base_boxes.py
@@ -0,0 +1,549 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from abc import ABCMeta, abstractmethod, abstractproperty, abstractstaticmethod
+from typing import List, Optional, Sequence, Tuple, Type, TypeVar, Union
+
+import numpy as np
+import torch
+from torch import BoolTensor, Tensor
+
+from mmdet.structures.mask.structures import BitmapMasks, PolygonMasks
+
+T = TypeVar('T')
+DeviceType = Union[str, torch.device]
+IndexType = Union[slice, int, list, torch.LongTensor, torch.cuda.LongTensor,
+ torch.BoolTensor, torch.cuda.BoolTensor, np.ndarray]
+MaskType = Union[BitmapMasks, PolygonMasks]
+
+
+class BaseBoxes(metaclass=ABCMeta):
+ """The base class for 2D box types.
+
+ The functions of ``BaseBoxes`` lie in three fields:
+
+ - Verify the boxes shape.
+ - Support tensor-like operations.
+ - Define abstract functions for 2D boxes.
+
+ In ``__init__`` , ``BaseBoxes`` verifies the validity of the data shape
+ w.r.t ``box_dim``. The tensor with the dimension >= 2 and the length
+ of the last dimension being ``box_dim`` will be regarded as valid.
+ ``BaseBoxes`` will restore them at the field ``tensor``. It's necessary
+ to override ``box_dim`` in subclass to guarantee the data shape is
+ correct.
+
+ There are many basic tensor-like functions implemented in ``BaseBoxes``.
+ In most cases, users can operate ``BaseBoxes`` instance like a normal
+ tensor. To protect the validity of data shape, All tensor-like functions
+ cannot modify the last dimension of ``self.tensor``.
+
+ When creating a new box type, users need to inherit from ``BaseBoxes``
+ and override abstract methods and specify the ``box_dim``. Then, register
+ the new box type by using the decorator ``register_box_type``.
+
+ Args:
+ data (Tensor or np.ndarray or Sequence): The box data with shape
+ (..., box_dim).
+ dtype (torch.dtype, Optional): data type of boxes. Defaults to None.
+ device (str or torch.device, Optional): device of boxes.
+ Default to None.
+ clone (bool): Whether clone ``boxes`` or not. Defaults to True.
+ """
+
+ # Used to verify the last dimension length
+ # Should override it in subclass.
+ box_dim: int = 0
+
+ def __init__(self,
+ data: Union[Tensor, np.ndarray, Sequence],
+ dtype: Optional[torch.dtype] = None,
+ device: Optional[DeviceType] = None,
+ clone: bool = True) -> None:
+ if isinstance(data, (np.ndarray, Tensor, Sequence)):
+ data = torch.as_tensor(data)
+ else:
+ raise TypeError('boxes should be Tensor, ndarray, or Sequence, ',
+ f'but got {type(data)}')
+
+ if device is not None or dtype is not None:
+ data = data.to(dtype=dtype, device=device)
+ # Clone the data to avoid potential bugs
+ if clone:
+ data = data.clone()
+ # handle the empty input like []
+ if data.numel() == 0:
+ data = data.reshape((-1, self.box_dim))
+
+ assert data.dim() >= 2 and data.size(-1) == self.box_dim, \
+ ('The boxes dimension must >= 2 and the length of the last '
+ f'dimension must be {self.box_dim}, but got boxes with '
+ f'shape {data.shape}.')
+ self.tensor = data
+
+ def convert_to(self, dst_type: Union[str, type]) -> 'BaseBoxes':
+ """Convert self to another box type.
+
+ Args:
+ dst_type (str or type): destination box type.
+
+ Returns:
+ :obj:`BaseBoxes`: destination box type object .
+ """
+ from .box_type import convert_box_type
+ return convert_box_type(self, dst_type=dst_type)
+
+ def empty_boxes(self: T,
+ dtype: Optional[torch.dtype] = None,
+ device: Optional[DeviceType] = None) -> T:
+ """Create empty box.
+
+ Args:
+ dtype (torch.dtype, Optional): data type of boxes.
+ device (str or torch.device, Optional): device of boxes.
+
+ Returns:
+ T: empty boxes with shape of (0, box_dim).
+ """
+ empty_box = self.tensor.new_zeros(
+ 0, self.box_dim, dtype=dtype, device=device)
+ return type(self)(empty_box, clone=False)
+
+ def fake_boxes(self: T,
+ sizes: Tuple[int],
+ fill: float = 0,
+ dtype: Optional[torch.dtype] = None,
+ device: Optional[DeviceType] = None) -> T:
+ """Create fake boxes with specific sizes and fill values.
+
+ Args:
+ sizes (Tuple[int]): The size of fake boxes. The last value must
+ be equal with ``self.box_dim``.
+ fill (float): filling value. Defaults to 0.
+ dtype (torch.dtype, Optional): data type of boxes.
+ device (str or torch.device, Optional): device of boxes.
+
+ Returns:
+ T: Fake boxes with shape of ``sizes``.
+ """
+ fake_boxes = self.tensor.new_full(
+ sizes, fill, dtype=dtype, device=device)
+ return type(self)(fake_boxes, clone=False)
+
+ def __getitem__(self: T, index: IndexType) -> T:
+ """Rewrite getitem to protect the last dimension shape."""
+ boxes = self.tensor
+ if isinstance(index, np.ndarray):
+ index = torch.as_tensor(index, device=self.device)
+ if isinstance(index, Tensor) and index.dtype == torch.bool:
+ assert index.dim() < boxes.dim()
+ elif isinstance(index, tuple):
+ assert len(index) < boxes.dim()
+ # `Ellipsis`(...) is commonly used in index like [None, ...].
+ # When `Ellipsis` is in index, it must be the last item.
+ if Ellipsis in index:
+ assert index[-1] is Ellipsis
+
+ boxes = boxes[index]
+ if boxes.dim() == 1:
+ boxes = boxes.reshape(1, -1)
+ return type(self)(boxes, clone=False)
+
+ def __setitem__(self: T, index: IndexType, values: Union[Tensor, T]) -> T:
+ """Rewrite setitem to protect the last dimension shape."""
+ assert type(values) is type(self), \
+ 'The value to be set must be the same box type as self'
+ values = values.tensor
+
+ if isinstance(index, np.ndarray):
+ index = torch.as_tensor(index, device=self.device)
+ if isinstance(index, Tensor) and index.dtype == torch.bool:
+ assert index.dim() < self.tensor.dim()
+ elif isinstance(index, tuple):
+ assert len(index) < self.tensor.dim()
+ # `Ellipsis`(...) is commonly used in index like [None, ...].
+ # When `Ellipsis` is in index, it must be the last item.
+ if Ellipsis in index:
+ assert index[-1] is Ellipsis
+
+ self.tensor[index] = values
+
+ def __len__(self) -> int:
+ """Return the length of self.tensor first dimension."""
+ return self.tensor.size(0)
+
+ def __deepcopy__(self, memo):
+ """Only clone the ``self.tensor`` when applying deepcopy."""
+ cls = self.__class__
+ other = cls.__new__(cls)
+ memo[id(self)] = other
+ other.tensor = self.tensor.clone()
+ return other
+
+ def __repr__(self) -> str:
+ """Return a strings that describes the object."""
+ return self.__class__.__name__ + '(\n' + str(self.tensor) + ')'
+
+ def new_tensor(self, *args, **kwargs) -> Tensor:
+ """Reload ``new_tensor`` from self.tensor."""
+ return self.tensor.new_tensor(*args, **kwargs)
+
+ def new_full(self, *args, **kwargs) -> Tensor:
+ """Reload ``new_full`` from self.tensor."""
+ return self.tensor.new_full(*args, **kwargs)
+
+ def new_empty(self, *args, **kwargs) -> Tensor:
+ """Reload ``new_empty`` from self.tensor."""
+ return self.tensor.new_empty(*args, **kwargs)
+
+ def new_ones(self, *args, **kwargs) -> Tensor:
+ """Reload ``new_ones`` from self.tensor."""
+ return self.tensor.new_ones(*args, **kwargs)
+
+ def new_zeros(self, *args, **kwargs) -> Tensor:
+ """Reload ``new_zeros`` from self.tensor."""
+ return self.tensor.new_zeros(*args, **kwargs)
+
+ def size(self, dim: Optional[int] = None) -> Union[int, torch.Size]:
+ """Reload new_zeros from self.tensor."""
+ # self.tensor.size(dim) cannot work when dim=None.
+ return self.tensor.size() if dim is None else self.tensor.size(dim)
+
+ def dim(self) -> int:
+ """Reload ``dim`` from self.tensor."""
+ return self.tensor.dim()
+
+ @property
+ def device(self) -> torch.device:
+ """Reload ``device`` from self.tensor."""
+ return self.tensor.device
+
+ @property
+ def dtype(self) -> torch.dtype:
+ """Reload ``dtype`` from self.tensor."""
+ return self.tensor.dtype
+
+ @property
+ def shape(self) -> torch.Size:
+ return self.tensor.shape
+
+ def numel(self) -> int:
+ """Reload ``numel`` from self.tensor."""
+ return self.tensor.numel()
+
+ def numpy(self) -> np.ndarray:
+ """Reload ``numpy`` from self.tensor."""
+ return self.tensor.numpy()
+
+ def to(self: T, *args, **kwargs) -> T:
+ """Reload ``to`` from self.tensor."""
+ return type(self)(self.tensor.to(*args, **kwargs), clone=False)
+
+ def cpu(self: T) -> T:
+ """Reload ``cpu`` from self.tensor."""
+ return type(self)(self.tensor.cpu(), clone=False)
+
+ def cuda(self: T, *args, **kwargs) -> T:
+ """Reload ``cuda`` from self.tensor."""
+ return type(self)(self.tensor.cuda(*args, **kwargs), clone=False)
+
+ def clone(self: T) -> T:
+ """Reload ``clone`` from self.tensor."""
+ return type(self)(self.tensor)
+
+ def detach(self: T) -> T:
+ """Reload ``detach`` from self.tensor."""
+ return type(self)(self.tensor.detach(), clone=False)
+
+ def view(self: T, *shape: Tuple[int]) -> T:
+ """Reload ``view`` from self.tensor."""
+ return type(self)(self.tensor.view(shape), clone=False)
+
+ def reshape(self: T, *shape: Tuple[int]) -> T:
+ """Reload ``reshape`` from self.tensor."""
+ return type(self)(self.tensor.reshape(shape), clone=False)
+
+ def expand(self: T, *sizes: Tuple[int]) -> T:
+ """Reload ``expand`` from self.tensor."""
+ return type(self)(self.tensor.expand(sizes), clone=False)
+
+ def repeat(self: T, *sizes: Tuple[int]) -> T:
+ """Reload ``repeat`` from self.tensor."""
+ return type(self)(self.tensor.repeat(sizes), clone=False)
+
+ def transpose(self: T, dim0: int, dim1: int) -> T:
+ """Reload ``transpose`` from self.tensor."""
+ ndim = self.tensor.dim()
+ assert dim0 != -1 and dim0 != ndim - 1
+ assert dim1 != -1 and dim1 != ndim - 1
+ return type(self)(self.tensor.transpose(dim0, dim1), clone=False)
+
+ def permute(self: T, *dims: Tuple[int]) -> T:
+ """Reload ``permute`` from self.tensor."""
+ assert dims[-1] == -1 or dims[-1] == self.tensor.dim() - 1
+ return type(self)(self.tensor.permute(dims), clone=False)
+
+ def split(self: T,
+ split_size_or_sections: Union[int, Sequence[int]],
+ dim: int = 0) -> List[T]:
+ """Reload ``split`` from self.tensor."""
+ assert dim != -1 and dim != self.tensor.dim() - 1
+ boxes_list = self.tensor.split(split_size_or_sections, dim=dim)
+ return [type(self)(boxes, clone=False) for boxes in boxes_list]
+
+ def chunk(self: T, chunks: int, dim: int = 0) -> List[T]:
+ """Reload ``chunk`` from self.tensor."""
+ assert dim != -1 and dim != self.tensor.dim() - 1
+ boxes_list = self.tensor.chunk(chunks, dim=dim)
+ return [type(self)(boxes, clone=False) for boxes in boxes_list]
+
+ def unbind(self: T, dim: int = 0) -> T:
+ """Reload ``unbind`` from self.tensor."""
+ assert dim != -1 and dim != self.tensor.dim() - 1
+ boxes_list = self.tensor.unbind(dim=dim)
+ return [type(self)(boxes, clone=False) for boxes in boxes_list]
+
+ def flatten(self: T, start_dim: int = 0, end_dim: int = -2) -> T:
+ """Reload ``flatten`` from self.tensor."""
+ assert end_dim != -1 and end_dim != self.tensor.dim() - 1
+ return type(self)(self.tensor.flatten(start_dim, end_dim), clone=False)
+
+ def squeeze(self: T, dim: Optional[int] = None) -> T:
+ """Reload ``squeeze`` from self.tensor."""
+ boxes = self.tensor.squeeze() if dim is None else \
+ self.tensor.squeeze(dim)
+ return type(self)(boxes, clone=False)
+
+ def unsqueeze(self: T, dim: int) -> T:
+ """Reload ``unsqueeze`` from self.tensor."""
+ assert dim != -1 and dim != self.tensor.dim()
+ return type(self)(self.tensor.unsqueeze(dim), clone=False)
+
+ @classmethod
+ def cat(cls: Type[T], box_list: Sequence[T], dim: int = 0) -> T:
+ """Cancatenates a box instance list into one single box instance.
+ Similar to ``torch.cat``.
+
+ Args:
+ box_list (Sequence[T]): A sequence of box instances.
+ dim (int): The dimension over which the box are concatenated.
+ Defaults to 0.
+
+ Returns:
+ T: Concatenated box instance.
+ """
+ assert isinstance(box_list, Sequence)
+ if len(box_list) == 0:
+ raise ValueError('box_list should not be a empty list.')
+
+ assert dim != -1 and dim != box_list[0].dim() - 1
+ assert all(isinstance(boxes, cls) for boxes in box_list)
+
+ th_box_list = [boxes.tensor for boxes in box_list]
+ return cls(torch.cat(th_box_list, dim=dim), clone=False)
+
+ @classmethod
+ def stack(cls: Type[T], box_list: Sequence[T], dim: int = 0) -> T:
+ """Concatenates a sequence of tensors along a new dimension. Similar to
+ ``torch.stack``.
+
+ Args:
+ box_list (Sequence[T]): A sequence of box instances.
+ dim (int): Dimension to insert. Defaults to 0.
+
+ Returns:
+ T: Concatenated box instance.
+ """
+ assert isinstance(box_list, Sequence)
+ if len(box_list) == 0:
+ raise ValueError('box_list should not be a empty list.')
+
+ assert dim != -1 and dim != box_list[0].dim()
+ assert all(isinstance(boxes, cls) for boxes in box_list)
+
+ th_box_list = [boxes.tensor for boxes in box_list]
+ return cls(torch.stack(th_box_list, dim=dim), clone=False)
+
+ @abstractproperty
+ def centers(self) -> Tensor:
+ """Return a tensor representing the centers of boxes."""
+ pass
+
+ @abstractproperty
+ def areas(self) -> Tensor:
+ """Return a tensor representing the areas of boxes."""
+ pass
+
+ @abstractproperty
+ def widths(self) -> Tensor:
+ """Return a tensor representing the widths of boxes."""
+ pass
+
+ @abstractproperty
+ def heights(self) -> Tensor:
+ """Return a tensor representing the heights of boxes."""
+ pass
+
+ @abstractmethod
+ def flip_(self,
+ img_shape: Tuple[int, int],
+ direction: str = 'horizontal') -> None:
+ """Flip boxes horizontally or vertically in-place.
+
+ Args:
+ img_shape (Tuple[int, int]): A tuple of image height and width.
+ direction (str): Flip direction, options are "horizontal",
+ "vertical" and "diagonal". Defaults to "horizontal"
+ """
+ pass
+
+ @abstractmethod
+ def translate_(self, distances: Tuple[float, float]) -> None:
+ """Translate boxes in-place.
+
+ Args:
+ distances (Tuple[float, float]): translate distances. The first
+ is horizontal distance and the second is vertical distance.
+ """
+ pass
+
+ @abstractmethod
+ def clip_(self, img_shape: Tuple[int, int]) -> None:
+ """Clip boxes according to the image shape in-place.
+
+ Args:
+ img_shape (Tuple[int, int]): A tuple of image height and width.
+ """
+ pass
+
+ @abstractmethod
+ def rotate_(self, center: Tuple[float, float], angle: float) -> None:
+ """Rotate all boxes in-place.
+
+ Args:
+ center (Tuple[float, float]): Rotation origin.
+ angle (float): Rotation angle represented in degrees. Positive
+ values mean clockwise rotation.
+ """
+ pass
+
+ @abstractmethod
+ def project_(self, homography_matrix: Union[Tensor, np.ndarray]) -> None:
+ """Geometric transformat boxes in-place.
+
+ Args:
+ homography_matrix (Tensor or np.ndarray]):
+ Shape (3, 3) for geometric transformation.
+ """
+ pass
+
+ @abstractmethod
+ def rescale_(self, scale_factor: Tuple[float, float]) -> None:
+ """Rescale boxes w.r.t. rescale_factor in-place.
+
+ Note:
+ Both ``rescale_`` and ``resize_`` will enlarge or shrink boxes
+ w.r.t ``scale_facotr``. The difference is that ``resize_`` only
+ changes the width and the height of boxes, but ``rescale_`` also
+ rescales the box centers simultaneously.
+
+ Args:
+ scale_factor (Tuple[float, float]): factors for scaling boxes.
+ The length should be 2.
+ """
+ pass
+
+ @abstractmethod
+ def resize_(self, scale_factor: Tuple[float, float]) -> None:
+ """Resize the box width and height w.r.t scale_factor in-place.
+
+ Note:
+ Both ``rescale_`` and ``resize_`` will enlarge or shrink boxes
+ w.r.t ``scale_facotr``. The difference is that ``resize_`` only
+ changes the width and the height of boxes, but ``rescale_`` also
+ rescales the box centers simultaneously.
+
+ Args:
+ scale_factor (Tuple[float, float]): factors for scaling box
+ shapes. The length should be 2.
+ """
+ pass
+
+ @abstractmethod
+ def is_inside(self,
+ img_shape: Tuple[int, int],
+ all_inside: bool = False,
+ allowed_border: int = 0) -> BoolTensor:
+ """Find boxes inside the image.
+
+ Args:
+ img_shape (Tuple[int, int]): A tuple of image height and width.
+ all_inside (bool): Whether the boxes are all inside the image or
+ part inside the image. Defaults to False.
+ allowed_border (int): Boxes that extend beyond the image shape
+ boundary by more than ``allowed_border`` are considered
+ "outside" Defaults to 0.
+ Returns:
+ BoolTensor: A BoolTensor indicating whether the box is inside
+ the image. Assuming the original boxes have shape (m, n, box_dim),
+ the output has shape (m, n).
+ """
+ pass
+
+ @abstractmethod
+ def find_inside_points(self,
+ points: Tensor,
+ is_aligned: bool = False) -> BoolTensor:
+ """Find inside box points. Boxes dimension must be 2.
+
+ Args:
+ points (Tensor): Points coordinates. Has shape of (m, 2).
+ is_aligned (bool): Whether ``points`` has been aligned with boxes
+ or not. If True, the length of boxes and ``points`` should be
+ the same. Defaults to False.
+
+ Returns:
+ BoolTensor: A BoolTensor indicating whether a point is inside
+ boxes. Assuming the boxes has shape of (n, box_dim), if
+ ``is_aligned`` is False. The index has shape of (m, n). If
+ ``is_aligned`` is True, m should be equal to n and the index has
+ shape of (m, ).
+ """
+ pass
+
+ @abstractstaticmethod
+ def overlaps(boxes1: 'BaseBoxes',
+ boxes2: 'BaseBoxes',
+ mode: str = 'iou',
+ is_aligned: bool = False,
+ eps: float = 1e-6) -> Tensor:
+ """Calculate overlap between two set of boxes with their types
+ converted to the present box type.
+
+ Args:
+ boxes1 (:obj:`BaseBoxes`): BaseBoxes with shape of (m, box_dim)
+ or empty.
+ boxes2 (:obj:`BaseBoxes`): BaseBoxes with shape of (n, box_dim)
+ or empty.
+ mode (str): "iou" (intersection over union), "iof" (intersection
+ over foreground). Defaults to "iou".
+ is_aligned (bool): If True, then m and n must be equal. Defaults
+ to False.
+ eps (float): A value added to the denominator for numerical
+ stability. Defaults to 1e-6.
+
+ Returns:
+ Tensor: shape (m, n) if ``is_aligned`` is False else shape (m,)
+ """
+ pass
+
+ @abstractstaticmethod
+ def from_instance_masks(masks: MaskType) -> 'BaseBoxes':
+ """Create boxes from instance masks.
+
+ Args:
+ masks (:obj:`BitmapMasks` or :obj:`PolygonMasks`): BitmapMasks or
+ PolygonMasks instance with length of n.
+
+ Returns:
+ :obj:`BaseBoxes`: Converted boxes with shape of (n, box_dim).
+ """
+ pass
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/bbox/bbox_overlaps.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/bbox/bbox_overlaps.py
new file mode 100644
index 0000000000000000000000000000000000000000..8e3435d28b38a5479a6c791f52a76d8ba293a6eb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/bbox/bbox_overlaps.py
@@ -0,0 +1,199 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+
+
+def fp16_clamp(x, min=None, max=None):
+ if not x.is_cuda and x.dtype == torch.float16:
+ # clamp for cpu float16, tensor fp16 has no clamp implementation
+ return x.float().clamp(min, max).half()
+
+ return x.clamp(min, max)
+
+
+def bbox_overlaps(bboxes1, bboxes2, mode='iou', is_aligned=False, eps=1e-6):
+ """Calculate overlap between two set of bboxes.
+
+ FP16 Contributed by https://github.com/open-mmlab/mmdetection/pull/4889
+ Note:
+ Assume bboxes1 is M x 4, bboxes2 is N x 4, when mode is 'iou',
+ there are some new generated variable when calculating IOU
+ using bbox_overlaps function:
+
+ 1) is_aligned is False
+ area1: M x 1
+ area2: N x 1
+ lt: M x N x 2
+ rb: M x N x 2
+ wh: M x N x 2
+ overlap: M x N x 1
+ union: M x N x 1
+ ious: M x N x 1
+
+ Total memory:
+ S = (9 x N x M + N + M) * 4 Byte,
+
+ When using FP16, we can reduce:
+ R = (9 x N x M + N + M) * 4 / 2 Byte
+ R large than (N + M) * 4 * 2 is always true when N and M >= 1.
+ Obviously, N + M <= N * M < 3 * N * M, when N >=2 and M >=2,
+ N + 1 < 3 * N, when N or M is 1.
+
+ Given M = 40 (ground truth), N = 400000 (three anchor boxes
+ in per grid, FPN, R-CNNs),
+ R = 275 MB (one times)
+
+ A special case (dense detection), M = 512 (ground truth),
+ R = 3516 MB = 3.43 GB
+
+ When the batch size is B, reduce:
+ B x R
+
+ Therefore, CUDA memory runs out frequently.
+
+ Experiments on GeForce RTX 2080Ti (11019 MiB):
+
+ | dtype | M | N | Use | Real | Ideal |
+ |:----:|:----:|:----:|:----:|:----:|:----:|
+ | FP32 | 512 | 400000 | 8020 MiB | -- | -- |
+ | FP16 | 512 | 400000 | 4504 MiB | 3516 MiB | 3516 MiB |
+ | FP32 | 40 | 400000 | 1540 MiB | -- | -- |
+ | FP16 | 40 | 400000 | 1264 MiB | 276MiB | 275 MiB |
+
+ 2) is_aligned is True
+ area1: N x 1
+ area2: N x 1
+ lt: N x 2
+ rb: N x 2
+ wh: N x 2
+ overlap: N x 1
+ union: N x 1
+ ious: N x 1
+
+ Total memory:
+ S = 11 x N * 4 Byte
+
+ When using FP16, we can reduce:
+ R = 11 x N * 4 / 2 Byte
+
+ So do the 'giou' (large than 'iou').
+
+ Time-wise, FP16 is generally faster than FP32.
+
+ When gpu_assign_thr is not -1, it takes more time on cpu
+ but not reduce memory.
+ There, we can reduce half the memory and keep the speed.
+
+ If ``is_aligned`` is ``False``, then calculate the overlaps between each
+ bbox of bboxes1 and bboxes2, otherwise the overlaps between each aligned
+ pair of bboxes1 and bboxes2.
+
+ Args:
+ bboxes1 (Tensor): shape (B, m, 4) in format or empty.
+ bboxes2 (Tensor): shape (B, n, 4) in format or empty.
+ B indicates the batch dim, in shape (B1, B2, ..., Bn).
+ If ``is_aligned`` is ``True``, then m and n must be equal.
+ mode (str): "iou" (intersection over union), "iof" (intersection over
+ foreground) or "giou" (generalized intersection over union).
+ Default "iou".
+ is_aligned (bool, optional): If True, then m and n must be equal.
+ Default False.
+ eps (float, optional): A value added to the denominator for numerical
+ stability. Default 1e-6.
+
+ Returns:
+ Tensor: shape (m, n) if ``is_aligned`` is False else shape (m,)
+
+ Example:
+ >>> bboxes1 = torch.FloatTensor([
+ >>> [0, 0, 10, 10],
+ >>> [10, 10, 20, 20],
+ >>> [32, 32, 38, 42],
+ >>> ])
+ >>> bboxes2 = torch.FloatTensor([
+ >>> [0, 0, 10, 20],
+ >>> [0, 10, 10, 19],
+ >>> [10, 10, 20, 20],
+ >>> ])
+ >>> overlaps = bbox_overlaps(bboxes1, bboxes2)
+ >>> assert overlaps.shape == (3, 3)
+ >>> overlaps = bbox_overlaps(bboxes1, bboxes2, is_aligned=True)
+ >>> assert overlaps.shape == (3, )
+
+ Example:
+ >>> empty = torch.empty(0, 4)
+ >>> nonempty = torch.FloatTensor([[0, 0, 10, 9]])
+ >>> assert tuple(bbox_overlaps(empty, nonempty).shape) == (0, 1)
+ >>> assert tuple(bbox_overlaps(nonempty, empty).shape) == (1, 0)
+ >>> assert tuple(bbox_overlaps(empty, empty).shape) == (0, 0)
+ """
+
+ assert mode in ['iou', 'iof', 'giou'], f'Unsupported mode {mode}'
+ # Either the boxes are empty or the length of boxes' last dimension is 4
+ assert (bboxes1.size(-1) == 4 or bboxes1.size(0) == 0)
+ assert (bboxes2.size(-1) == 4 or bboxes2.size(0) == 0)
+
+ # Batch dim must be the same
+ # Batch dim: (B1, B2, ... Bn)
+ assert bboxes1.shape[:-2] == bboxes2.shape[:-2]
+ batch_shape = bboxes1.shape[:-2]
+
+ rows = bboxes1.size(-2)
+ cols = bboxes2.size(-2)
+ if is_aligned:
+ assert rows == cols
+
+ if rows * cols == 0:
+ if is_aligned:
+ return bboxes1.new(batch_shape + (rows, ))
+ else:
+ return bboxes1.new(batch_shape + (rows, cols))
+
+ area1 = (bboxes1[..., 2] - bboxes1[..., 0]) * (
+ bboxes1[..., 3] - bboxes1[..., 1])
+ area2 = (bboxes2[..., 2] - bboxes2[..., 0]) * (
+ bboxes2[..., 3] - bboxes2[..., 1])
+
+ if is_aligned:
+ lt = torch.max(bboxes1[..., :2], bboxes2[..., :2]) # [B, rows, 2]
+ rb = torch.min(bboxes1[..., 2:], bboxes2[..., 2:]) # [B, rows, 2]
+
+ wh = fp16_clamp(rb - lt, min=0)
+ overlap = wh[..., 0] * wh[..., 1]
+
+ if mode in ['iou', 'giou']:
+ union = area1 + area2 - overlap
+ else:
+ union = area1
+ if mode == 'giou':
+ enclosed_lt = torch.min(bboxes1[..., :2], bboxes2[..., :2])
+ enclosed_rb = torch.max(bboxes1[..., 2:], bboxes2[..., 2:])
+ else:
+ lt = torch.max(bboxes1[..., :, None, :2],
+ bboxes2[..., None, :, :2]) # [B, rows, cols, 2]
+ rb = torch.min(bboxes1[..., :, None, 2:],
+ bboxes2[..., None, :, 2:]) # [B, rows, cols, 2]
+
+ wh = fp16_clamp(rb - lt, min=0)
+ overlap = wh[..., 0] * wh[..., 1]
+
+ if mode in ['iou', 'giou']:
+ union = area1[..., None] + area2[..., None, :] - overlap
+ else:
+ union = area1[..., None]
+ if mode == 'giou':
+ enclosed_lt = torch.min(bboxes1[..., :, None, :2],
+ bboxes2[..., None, :, :2])
+ enclosed_rb = torch.max(bboxes1[..., :, None, 2:],
+ bboxes2[..., None, :, 2:])
+
+ eps = union.new_tensor([eps])
+ union = torch.max(union, eps)
+ ious = overlap / union
+ if mode in ['iou', 'iof']:
+ return ious
+ # calculate gious
+ enclose_wh = fp16_clamp(enclosed_rb - enclosed_lt, min=0)
+ enclose_area = enclose_wh[..., 0] * enclose_wh[..., 1]
+ enclose_area = torch.max(enclose_area, eps)
+ gious = ious - (enclose_area - union) / enclose_area
+ return gious
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/bbox/box_type.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/bbox/box_type.py
new file mode 100644
index 0000000000000000000000000000000000000000..c7eb5494c36c8efcbb414897f7c2532a6d3a1ddb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/bbox/box_type.py
@@ -0,0 +1,296 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Callable, Optional, Tuple, Type, Union
+
+import numpy as np
+import torch
+from torch import Tensor
+
+from .base_boxes import BaseBoxes
+
+BoxType = Union[np.ndarray, Tensor, BaseBoxes]
+
+box_types: dict = {}
+_box_type_to_name: dict = {}
+box_converters: dict = {}
+
+
+def _register_box(name: str, box_type: Type, force: bool = False) -> None:
+ """Register a box type.
+
+ Args:
+ name (str): The name of box type.
+ box_type (type): Box mode class to be registered.
+ force (bool): Whether to override an existing class with the same
+ name. Defaults to False.
+ """
+ assert issubclass(box_type, BaseBoxes)
+ name = name.lower()
+
+ if not force and (name in box_types or box_type in _box_type_to_name):
+ raise KeyError(f'box type {name} has been registered')
+ elif name in box_types:
+ _box_type = box_types.pop(name)
+ _box_type_to_name.pop(_box_type)
+ elif box_type in _box_type_to_name:
+ _name = _box_type_to_name.pop(box_type)
+ box_types.pop(_name)
+
+ box_types[name] = box_type
+ _box_type_to_name[box_type] = name
+
+
+def register_box(name: str,
+ box_type: Type = None,
+ force: bool = False) -> Union[Type, Callable]:
+ """Register a box type.
+
+ A record will be added to ``bbox_types``, whose key is the box type name
+ and value is the box type itself. Simultaneously, a reverse dictionary
+ ``_box_type_to_name`` will be updated. It can be used as a decorator or
+ a normal function.
+
+ Args:
+ name (str): The name of box type.
+ bbox_type (type, Optional): Box type class to be registered.
+ Defaults to None.
+ force (bool): Whether to override the existing box type with the same
+ name. Defaults to False.
+
+ Examples:
+ >>> from mmdet.structures.bbox import register_box
+ >>> from mmdet.structures.bbox import BaseBoxes
+
+ >>> # as a decorator
+ >>> @register_box('hbox')
+ >>> class HorizontalBoxes(BaseBoxes):
+ >>> pass
+
+ >>> # as a normal function
+ >>> class RotatedBoxes(BaseBoxes):
+ >>> pass
+ >>> register_box('rbox', RotatedBoxes)
+ """
+ if not isinstance(force, bool):
+ raise TypeError(f'force must be a boolean, but got {type(force)}')
+
+ # use it as a normal method: register_box(name, box_type=BoxCls)
+ if box_type is not None:
+ _register_box(name=name, box_type=box_type, force=force)
+ return box_type
+
+ # use it as a decorator: @register_box(name)
+ def _register(cls):
+ _register_box(name=name, box_type=cls, force=force)
+ return cls
+
+ return _register
+
+
+def _register_box_converter(src_type: Union[str, type],
+ dst_type: Union[str, type],
+ converter: Callable,
+ force: bool = False) -> None:
+ """Register a box converter.
+
+ Args:
+ src_type (str or type): source box type name or class.
+ dst_type (str or type): destination box type name or class.
+ converter (Callable): Convert function.
+ force (bool): Whether to override the existing box type with the same
+ name. Defaults to False.
+ """
+ assert callable(converter)
+ src_type_name, _ = get_box_type(src_type)
+ dst_type_name, _ = get_box_type(dst_type)
+
+ converter_name = src_type_name + '2' + dst_type_name
+ if not force and converter_name in box_converters:
+ raise KeyError(f'The box converter from {src_type_name} to '
+ f'{dst_type_name} has been registered.')
+
+ box_converters[converter_name] = converter
+
+
+def register_box_converter(src_type: Union[str, type],
+ dst_type: Union[str, type],
+ converter: Optional[Callable] = None,
+ force: bool = False) -> Callable:
+ """Register a box converter.
+
+ A record will be added to ``box_converter``, whose key is
+ '{src_type_name}2{dst_type_name}' and value is the convert function.
+ It can be used as a decorator or a normal function.
+
+ Args:
+ src_type (str or type): source box type name or class.
+ dst_type (str or type): destination box type name or class.
+ converter (Callable): Convert function. Defaults to None.
+ force (bool): Whether to override the existing box type with the same
+ name. Defaults to False.
+
+ Examples:
+ >>> from mmdet.structures.bbox import register_box_converter
+ >>> # as a decorator
+ >>> @register_box_converter('hbox', 'rbox')
+ >>> def converter_A(boxes):
+ >>> pass
+
+ >>> # as a normal function
+ >>> def converter_B(boxes):
+ >>> pass
+ >>> register_box_converter('rbox', 'hbox', converter_B)
+ """
+ if not isinstance(force, bool):
+ raise TypeError(f'force must be a boolean, but got {type(force)}')
+
+ # use it as a normal method:
+ # register_box_converter(src_type, dst_type, converter=Func)
+ if converter is not None:
+ _register_box_converter(
+ src_type=src_type,
+ dst_type=dst_type,
+ converter=converter,
+ force=force)
+ return converter
+
+ # use it as a decorator: @register_box_converter(name)
+ def _register(func):
+ _register_box_converter(
+ src_type=src_type, dst_type=dst_type, converter=func, force=force)
+ return func
+
+ return _register
+
+
+def get_box_type(box_type: Union[str, type]) -> Tuple[str, type]:
+ """get both box type name and class.
+
+ Args:
+ box_type (str or type): Single box type name or class.
+
+ Returns:
+ Tuple[str, type]: A tuple of box type name and class.
+ """
+ if isinstance(box_type, str):
+ type_name = box_type.lower()
+ assert type_name in box_types, \
+ f"Box type {type_name} hasn't been registered in box_types."
+ type_cls = box_types[type_name]
+ elif issubclass(box_type, BaseBoxes):
+ assert box_type in _box_type_to_name, \
+ f"Box type {box_type} hasn't been registered in box_types."
+ type_name = _box_type_to_name[box_type]
+ type_cls = box_type
+ else:
+ raise KeyError('box_type must be a str or class inheriting from '
+ f'BaseBoxes, but got {type(box_type)}.')
+ return type_name, type_cls
+
+
+def convert_box_type(boxes: BoxType,
+ *,
+ src_type: Union[str, type] = None,
+ dst_type: Union[str, type] = None) -> BoxType:
+ """Convert boxes from source type to destination type.
+
+ If ``boxes`` is a instance of BaseBoxes, the ``src_type`` will be set
+ as the type of ``boxes``.
+
+ Args:
+ boxes (np.ndarray or Tensor or :obj:`BaseBoxes`): boxes need to
+ convert.
+ src_type (str or type, Optional): source box type. Defaults to None.
+ dst_type (str or type, Optional): destination box type. Defaults to
+ None.
+
+ Returns:
+ Union[np.ndarray, Tensor, :obj:`BaseBoxes`]: Converted boxes. It's type
+ is consistent with the input's type.
+ """
+ assert dst_type is not None
+ dst_type_name, dst_type_cls = get_box_type(dst_type)
+
+ is_box_cls = False
+ is_numpy = False
+ if isinstance(boxes, BaseBoxes):
+ src_type_name, _ = get_box_type(type(boxes))
+ is_box_cls = True
+ elif isinstance(boxes, (Tensor, np.ndarray)):
+ assert src_type is not None
+ src_type_name, _ = get_box_type(src_type)
+ if isinstance(boxes, np.ndarray):
+ is_numpy = True
+ else:
+ raise TypeError('boxes must be a instance of BaseBoxes, Tensor or '
+ f'ndarray, but get {type(boxes)}.')
+
+ if src_type_name == dst_type_name:
+ return boxes
+
+ converter_name = src_type_name + '2' + dst_type_name
+ assert converter_name in box_converters, \
+ "Convert function hasn't been registered in box_converters."
+ converter = box_converters[converter_name]
+
+ if is_box_cls:
+ boxes = converter(boxes.tensor)
+ return dst_type_cls(boxes)
+ elif is_numpy:
+ boxes = converter(torch.from_numpy(boxes))
+ return boxes.numpy()
+ else:
+ return converter(boxes)
+
+
+def autocast_box_type(dst_box_type='hbox') -> Callable:
+ """A decorator which automatically casts results['gt_bboxes'] to the
+ destination box type.
+
+ It commenly used in mmdet.datasets.transforms to make the transforms up-
+ compatible with the np.ndarray type of results['gt_bboxes'].
+
+ The speed of processing of np.ndarray and BaseBoxes data are the same:
+
+ - np.ndarray: 0.0509 img/s
+ - BaseBoxes: 0.0551 img/s
+
+ Args:
+ dst_box_type (str): Destination box type.
+ """
+ _, box_type_cls = get_box_type(dst_box_type)
+
+ def decorator(func: Callable) -> Callable:
+
+ def wrapper(self, results: dict, *args, **kwargs) -> dict:
+ if ('gt_bboxes' not in results
+ or isinstance(results['gt_bboxes'], BaseBoxes)):
+ return func(self, results)
+ elif isinstance(results['gt_bboxes'], np.ndarray):
+ results['gt_bboxes'] = box_type_cls(
+ results['gt_bboxes'], clone=False)
+ if 'mix_results' in results:
+ for res in results['mix_results']:
+ if isinstance(res['gt_bboxes'], np.ndarray):
+ res['gt_bboxes'] = box_type_cls(
+ res['gt_bboxes'], clone=False)
+
+ _results = func(self, results, *args, **kwargs)
+
+ # In some cases, the function will process gt_bboxes in-place
+ # Simultaneously convert inputting and outputting gt_bboxes
+ # back to np.ndarray
+ if isinstance(_results, dict) and 'gt_bboxes' in _results:
+ if isinstance(_results['gt_bboxes'], BaseBoxes):
+ _results['gt_bboxes'] = _results['gt_bboxes'].numpy()
+ if isinstance(results['gt_bboxes'], BaseBoxes):
+ results['gt_bboxes'] = results['gt_bboxes'].numpy()
+ return _results
+ else:
+ raise TypeError(
+ "auto_box_type requires results['gt_bboxes'] to "
+ 'be BaseBoxes or np.ndarray, but got '
+ f"{type(results['gt_bboxes'])}")
+
+ return wrapper
+
+ return decorator
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/bbox/horizontal_boxes.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/bbox/horizontal_boxes.py
new file mode 100644
index 0000000000000000000000000000000000000000..b3a78518105fda02cef2d3a2bcaceb410759165c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/bbox/horizontal_boxes.py
@@ -0,0 +1,432 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Tuple, TypeVar, Union
+
+import cv2
+import numpy as np
+import torch
+from torch import BoolTensor, Tensor
+
+from mmdet.structures.mask.structures import BitmapMasks, PolygonMasks
+from .base_boxes import BaseBoxes
+from .bbox_overlaps import bbox_overlaps
+from .box_type import register_box
+
+T = TypeVar('T')
+DeviceType = Union[str, torch.device]
+MaskType = Union[BitmapMasks, PolygonMasks]
+
+
+@register_box(name='hbox')
+class HorizontalBoxes(BaseBoxes):
+ """The horizontal box class used in MMDetection by default.
+
+ The ``box_dim`` of ``HorizontalBoxes`` is 4, which means the length of
+ the last dimension of the data should be 4. Two modes of box data are
+ supported in ``HorizontalBoxes``:
+
+ - 'xyxy': Each row of data indicates (x1, y1, x2, y2), which are the
+ coordinates of the left-top and right-bottom points.
+ - 'cxcywh': Each row of data indicates (x, y, w, h), where (x, y) are the
+ coordinates of the box centers and (w, h) are the width and height.
+
+ ``HorizontalBoxes`` only restores 'xyxy' mode of data. If the the data is
+ in 'cxcywh' mode, users need to input ``in_mode='cxcywh'`` and The code
+ will convert the 'cxcywh' data to 'xyxy' automatically.
+
+ Args:
+ data (Tensor or np.ndarray or Sequence): The box data with shape of
+ (..., 4).
+ dtype (torch.dtype, Optional): data type of boxes. Defaults to None.
+ device (str or torch.device, Optional): device of boxes.
+ Default to None.
+ clone (bool): Whether clone ``boxes`` or not. Defaults to True.
+ mode (str, Optional): the mode of boxes. If it is 'cxcywh', the
+ `data` will be converted to 'xyxy' mode. Defaults to None.
+ """
+
+ box_dim: int = 4
+
+ def __init__(self,
+ data: Union[Tensor, np.ndarray],
+ dtype: torch.dtype = None,
+ device: DeviceType = None,
+ clone: bool = True,
+ in_mode: Optional[str] = None) -> None:
+ super().__init__(data=data, dtype=dtype, device=device, clone=clone)
+ if isinstance(in_mode, str):
+ if in_mode not in ('xyxy', 'cxcywh'):
+ raise ValueError(f'Get invalid mode {in_mode}.')
+ if in_mode == 'cxcywh':
+ self.tensor = self.cxcywh_to_xyxy(self.tensor)
+
+ @staticmethod
+ def cxcywh_to_xyxy(boxes: Tensor) -> Tensor:
+ """Convert box coordinates from (cx, cy, w, h) to (x1, y1, x2, y2).
+
+ Args:
+ boxes (Tensor): cxcywh boxes tensor with shape of (..., 4).
+
+ Returns:
+ Tensor: xyxy boxes tensor with shape of (..., 4).
+ """
+ ctr, wh = boxes.split((2, 2), dim=-1)
+ return torch.cat([(ctr - wh / 2), (ctr + wh / 2)], dim=-1)
+
+ @staticmethod
+ def xyxy_to_cxcywh(boxes: Tensor) -> Tensor:
+ """Convert box coordinates from (x1, y1, x2, y2) to (cx, cy, w, h).
+
+ Args:
+ boxes (Tensor): xyxy boxes tensor with shape of (..., 4).
+
+ Returns:
+ Tensor: cxcywh boxes tensor with shape of (..., 4).
+ """
+ xy1, xy2 = boxes.split((2, 2), dim=-1)
+ return torch.cat([(xy2 + xy1) / 2, (xy2 - xy1)], dim=-1)
+
+ @property
+ def cxcywh(self) -> Tensor:
+ """Return a tensor representing the cxcywh boxes."""
+ return self.xyxy_to_cxcywh(self.tensor)
+
+ @property
+ def centers(self) -> Tensor:
+ """Return a tensor representing the centers of boxes."""
+ boxes = self.tensor
+ return (boxes[..., :2] + boxes[..., 2:]) / 2
+
+ @property
+ def areas(self) -> Tensor:
+ """Return a tensor representing the areas of boxes."""
+ boxes = self.tensor
+ return (boxes[..., 2] - boxes[..., 0]) * (
+ boxes[..., 3] - boxes[..., 1])
+
+ @property
+ def widths(self) -> Tensor:
+ """Return a tensor representing the widths of boxes."""
+ boxes = self.tensor
+ return boxes[..., 2] - boxes[..., 0]
+
+ @property
+ def heights(self) -> Tensor:
+ """Return a tensor representing the heights of boxes."""
+ boxes = self.tensor
+ return boxes[..., 3] - boxes[..., 1]
+
+ def flip_(self,
+ img_shape: Tuple[int, int],
+ direction: str = 'horizontal') -> None:
+ """Flip boxes horizontally or vertically in-place.
+
+ Args:
+ img_shape (Tuple[int, int]): A tuple of image height and width.
+ direction (str): Flip direction, options are "horizontal",
+ "vertical" and "diagonal". Defaults to "horizontal"
+ """
+ assert direction in ['horizontal', 'vertical', 'diagonal']
+ flipped = self.tensor
+ boxes = flipped.clone()
+ if direction == 'horizontal':
+ flipped[..., 0] = img_shape[1] - boxes[..., 2]
+ flipped[..., 2] = img_shape[1] - boxes[..., 0]
+ elif direction == 'vertical':
+ flipped[..., 1] = img_shape[0] - boxes[..., 3]
+ flipped[..., 3] = img_shape[0] - boxes[..., 1]
+ else:
+ flipped[..., 0] = img_shape[1] - boxes[..., 2]
+ flipped[..., 1] = img_shape[0] - boxes[..., 3]
+ flipped[..., 2] = img_shape[1] - boxes[..., 0]
+ flipped[..., 3] = img_shape[0] - boxes[..., 1]
+
+ def translate_(self, distances: Tuple[float, float]) -> None:
+ """Translate boxes in-place.
+
+ Args:
+ distances (Tuple[float, float]): translate distances. The first
+ is horizontal distance and the second is vertical distance.
+ """
+ boxes = self.tensor
+ assert len(distances) == 2
+ self.tensor = boxes + boxes.new_tensor(distances).repeat(2)
+
+ def clip_(self, img_shape: Tuple[int, int]) -> None:
+ """Clip boxes according to the image shape in-place.
+
+ Args:
+ img_shape (Tuple[int, int]): A tuple of image height and width.
+ """
+ boxes = self.tensor
+ boxes[..., 0::2] = boxes[..., 0::2].clamp(0, img_shape[1])
+ boxes[..., 1::2] = boxes[..., 1::2].clamp(0, img_shape[0])
+
+ def rotate_(self, center: Tuple[float, float], angle: float) -> None:
+ """Rotate all boxes in-place.
+
+ Args:
+ center (Tuple[float, float]): Rotation origin.
+ angle (float): Rotation angle represented in degrees. Positive
+ values mean clockwise rotation.
+ """
+ boxes = self.tensor
+ rotation_matrix = boxes.new_tensor(
+ cv2.getRotationMatrix2D(center, -angle, 1))
+
+ corners = self.hbox2corner(boxes)
+ corners = torch.cat(
+ [corners, corners.new_ones(*corners.shape[:-1], 1)], dim=-1)
+ corners_T = torch.transpose(corners, -1, -2)
+ corners_T = torch.matmul(rotation_matrix, corners_T)
+ corners = torch.transpose(corners_T, -1, -2)
+ self.tensor = self.corner2hbox(corners)
+
+ def project_(self, homography_matrix: Union[Tensor, np.ndarray]) -> None:
+ """Geometric transformat boxes in-place.
+
+ Args:
+ homography_matrix (Tensor or np.ndarray]):
+ Shape (3, 3) for geometric transformation.
+ """
+ boxes = self.tensor
+ if isinstance(homography_matrix, np.ndarray):
+ homography_matrix = boxes.new_tensor(homography_matrix)
+ corners = self.hbox2corner(boxes)
+ corners = torch.cat(
+ [corners, corners.new_ones(*corners.shape[:-1], 1)], dim=-1)
+ corners_T = torch.transpose(corners, -1, -2)
+ corners_T = torch.matmul(homography_matrix, corners_T)
+ corners = torch.transpose(corners_T, -1, -2)
+ # Convert to homogeneous coordinates by normalization
+ corners = corners[..., :2] / corners[..., 2:3]
+ self.tensor = self.corner2hbox(corners)
+
+ @staticmethod
+ def hbox2corner(boxes: Tensor) -> Tensor:
+ """Convert box coordinates from (x1, y1, x2, y2) to corners ((x1, y1),
+ (x2, y1), (x1, y2), (x2, y2)).
+
+ Args:
+ boxes (Tensor): Horizontal box tensor with shape of (..., 4).
+
+ Returns:
+ Tensor: Corner tensor with shape of (..., 4, 2).
+ """
+ x1, y1, x2, y2 = torch.split(boxes, 1, dim=-1)
+ corners = torch.cat([x1, y1, x2, y1, x1, y2, x2, y2], dim=-1)
+ return corners.reshape(*corners.shape[:-1], 4, 2)
+
+ @staticmethod
+ def corner2hbox(corners: Tensor) -> Tensor:
+ """Convert box coordinates from corners ((x1, y1), (x2, y1), (x1, y2),
+ (x2, y2)) to (x1, y1, x2, y2).
+
+ Args:
+ corners (Tensor): Corner tensor with shape of (..., 4, 2).
+
+ Returns:
+ Tensor: Horizontal box tensor with shape of (..., 4).
+ """
+ if corners.numel() == 0:
+ return corners.new_zeros((0, 4))
+ min_xy = corners.min(dim=-2)[0]
+ max_xy = corners.max(dim=-2)[0]
+ return torch.cat([min_xy, max_xy], dim=-1)
+
+ def rescale_(self, scale_factor: Tuple[float, float]) -> None:
+ """Rescale boxes w.r.t. rescale_factor in-place.
+
+ Note:
+ Both ``rescale_`` and ``resize_`` will enlarge or shrink boxes
+ w.r.t ``scale_facotr``. The difference is that ``resize_`` only
+ changes the width and the height of boxes, but ``rescale_`` also
+ rescales the box centers simultaneously.
+
+ Args:
+ scale_factor (Tuple[float, float]): factors for scaling boxes.
+ The length should be 2.
+ """
+ boxes = self.tensor
+ assert len(scale_factor) == 2
+ scale_factor = boxes.new_tensor(scale_factor).repeat(2)
+ self.tensor = boxes * scale_factor
+
+ def resize_(self, scale_factor: Tuple[float, float]) -> None:
+ """Resize the box width and height w.r.t scale_factor in-place.
+
+ Note:
+ Both ``rescale_`` and ``resize_`` will enlarge or shrink boxes
+ w.r.t ``scale_facotr``. The difference is that ``resize_`` only
+ changes the width and the height of boxes, but ``rescale_`` also
+ rescales the box centers simultaneously.
+
+ Args:
+ scale_factor (Tuple[float, float]): factors for scaling box
+ shapes. The length should be 2.
+ """
+ boxes = self.tensor
+ assert len(scale_factor) == 2
+ ctrs = (boxes[..., 2:] + boxes[..., :2]) / 2
+ wh = boxes[..., 2:] - boxes[..., :2]
+ scale_factor = boxes.new_tensor(scale_factor)
+ wh = wh * scale_factor
+ xy1 = ctrs - 0.5 * wh
+ xy2 = ctrs + 0.5 * wh
+ self.tensor = torch.cat([xy1, xy2], dim=-1)
+
+ def is_inside(self,
+ img_shape: Tuple[int, int],
+ all_inside: bool = False,
+ allowed_border: int = 0) -> BoolTensor:
+ """Find boxes inside the image.
+
+ Args:
+ img_shape (Tuple[int, int]): A tuple of image height and width.
+ all_inside (bool): Whether the boxes are all inside the image or
+ part inside the image. Defaults to False.
+ allowed_border (int): Boxes that extend beyond the image shape
+ boundary by more than ``allowed_border`` are considered
+ "outside" Defaults to 0.
+ Returns:
+ BoolTensor: A BoolTensor indicating whether the box is inside
+ the image. Assuming the original boxes have shape (m, n, 4),
+ the output has shape (m, n).
+ """
+ img_h, img_w = img_shape
+ boxes = self.tensor
+ if all_inside:
+ return (boxes[:, 0] >= -allowed_border) & \
+ (boxes[:, 1] >= -allowed_border) & \
+ (boxes[:, 2] < img_w + allowed_border) & \
+ (boxes[:, 3] < img_h + allowed_border)
+ else:
+ return (boxes[..., 0] < img_w + allowed_border) & \
+ (boxes[..., 1] < img_h + allowed_border) & \
+ (boxes[..., 2] > -allowed_border) & \
+ (boxes[..., 3] > -allowed_border)
+
+ def find_inside_points(self,
+ points: Tensor,
+ is_aligned: bool = False) -> BoolTensor:
+ """Find inside box points. Boxes dimension must be 2.
+
+ Args:
+ points (Tensor): Points coordinates. Has shape of (m, 2).
+ is_aligned (bool): Whether ``points`` has been aligned with boxes
+ or not. If True, the length of boxes and ``points`` should be
+ the same. Defaults to False.
+
+ Returns:
+ BoolTensor: A BoolTensor indicating whether a point is inside
+ boxes. Assuming the boxes has shape of (n, 4), if ``is_aligned``
+ is False. The index has shape of (m, n). If ``is_aligned`` is
+ True, m should be equal to n and the index has shape of (m, ).
+ """
+ boxes = self.tensor
+ assert boxes.dim() == 2, 'boxes dimension must be 2.'
+
+ if not is_aligned:
+ boxes = boxes[None, :, :]
+ points = points[:, None, :]
+ else:
+ assert boxes.size(0) == points.size(0)
+
+ x_min, y_min, x_max, y_max = boxes.unbind(dim=-1)
+ return (points[..., 0] >= x_min) & (points[..., 0] <= x_max) & \
+ (points[..., 1] >= y_min) & (points[..., 1] <= y_max)
+
+ def create_masks(self, img_shape: Tuple[int, int]) -> BitmapMasks:
+ """
+ Args:
+ img_shape (Tuple[int, int]): A tuple of image height and width.
+
+ Returns:
+ :obj:`BitmapMasks`: Converted masks
+ """
+ img_h, img_w = img_shape
+ boxes = self.tensor
+
+ xmin, ymin = boxes[:, 0:1], boxes[:, 1:2]
+ xmax, ymax = boxes[:, 2:3], boxes[:, 3:4]
+ gt_masks = np.zeros((len(boxes), img_h, img_w), dtype=np.uint8)
+ for i in range(len(boxes)):
+ gt_masks[i,
+ int(ymin[i]):int(ymax[i]),
+ int(xmin[i]):int(xmax[i])] = 1
+ return BitmapMasks(gt_masks, img_h, img_w)
+
+ @staticmethod
+ def overlaps(boxes1: BaseBoxes,
+ boxes2: BaseBoxes,
+ mode: str = 'iou',
+ is_aligned: bool = False,
+ eps: float = 1e-6) -> Tensor:
+ """Calculate overlap between two set of boxes with their types
+ converted to ``HorizontalBoxes``.
+
+ Args:
+ boxes1 (:obj:`BaseBoxes`): BaseBoxes with shape of (m, box_dim)
+ or empty.
+ boxes2 (:obj:`BaseBoxes`): BaseBoxes with shape of (n, box_dim)
+ or empty.
+ mode (str): "iou" (intersection over union), "iof" (intersection
+ over foreground). Defaults to "iou".
+ is_aligned (bool): If True, then m and n must be equal. Defaults
+ to False.
+ eps (float): A value added to the denominator for numerical
+ stability. Defaults to 1e-6.
+
+ Returns:
+ Tensor: shape (m, n) if ``is_aligned`` is False else shape (m,)
+ """
+ boxes1 = boxes1.convert_to('hbox')
+ boxes2 = boxes2.convert_to('hbox')
+ return bbox_overlaps(
+ boxes1.tensor,
+ boxes2.tensor,
+ mode=mode,
+ is_aligned=is_aligned,
+ eps=eps)
+
+ @staticmethod
+ def from_instance_masks(masks: MaskType) -> 'HorizontalBoxes':
+ """Create horizontal boxes from instance masks.
+
+ Args:
+ masks (:obj:`BitmapMasks` or :obj:`PolygonMasks`): BitmapMasks or
+ PolygonMasks instance with length of n.
+
+ Returns:
+ :obj:`HorizontalBoxes`: Converted boxes with shape of (n, 4).
+ """
+ num_masks = len(masks)
+ boxes = np.zeros((num_masks, 4), dtype=np.float32)
+ if isinstance(masks, BitmapMasks):
+ x_any = masks.masks.any(axis=1)
+ y_any = masks.masks.any(axis=2)
+ for idx in range(num_masks):
+ x = np.where(x_any[idx, :])[0]
+ y = np.where(y_any[idx, :])[0]
+ if len(x) > 0 and len(y) > 0:
+ # use +1 for x_max and y_max so that the right and bottom
+ # boundary of instance masks are fully included by the box
+ boxes[idx, :] = np.array(
+ [x[0], y[0], x[-1] + 1, y[-1] + 1], dtype=np.float32)
+ elif isinstance(masks, PolygonMasks):
+ for idx, poly_per_obj in enumerate(masks.masks):
+ # simply use a number that is big enough for comparison with
+ # coordinates
+ xy_min = np.array([masks.width * 2, masks.height * 2],
+ dtype=np.float32)
+ xy_max = np.zeros(2, dtype=np.float32)
+ for p in poly_per_obj:
+ xy = np.array(p).reshape(-1, 2).astype(np.float32)
+ xy_min = np.minimum(xy_min, np.min(xy, axis=0))
+ xy_max = np.maximum(xy_max, np.max(xy, axis=0))
+ boxes[idx, :2] = xy_min
+ boxes[idx, 2:] = xy_max
+ else:
+ raise TypeError(
+ '`masks` must be `BitmapMasks` or `PolygonMasks`, '
+ f'but got {type(masks)}.')
+ return HorizontalBoxes(boxes)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/bbox/transforms.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/bbox/transforms.py
new file mode 100644
index 0000000000000000000000000000000000000000..287e6aa6fcaeaf09a8a2838a04a97157cd02a00c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/bbox/transforms.py
@@ -0,0 +1,498 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Sequence, Tuple, Union
+
+import numpy as np
+import torch
+from torch import Tensor
+
+from mmdet.structures.bbox import BaseBoxes
+
+
+def find_inside_bboxes(bboxes: Tensor, img_h: int, img_w: int) -> Tensor:
+ """Find bboxes as long as a part of bboxes is inside the image.
+
+ Args:
+ bboxes (Tensor): Shape (N, 4).
+ img_h (int): Image height.
+ img_w (int): Image width.
+
+ Returns:
+ Tensor: Index of the remaining bboxes.
+ """
+ inside_inds = (bboxes[:, 0] < img_w) & (bboxes[:, 2] > 0) \
+ & (bboxes[:, 1] < img_h) & (bboxes[:, 3] > 0)
+ return inside_inds
+
+
+def bbox_flip(bboxes: Tensor,
+ img_shape: Tuple[int],
+ direction: str = 'horizontal') -> Tensor:
+ """Flip bboxes horizontally or vertically.
+
+ Args:
+ bboxes (Tensor): Shape (..., 4*k)
+ img_shape (Tuple[int]): Image shape.
+ direction (str): Flip direction, options are "horizontal", "vertical",
+ "diagonal". Default: "horizontal"
+
+ Returns:
+ Tensor: Flipped bboxes.
+ """
+ assert bboxes.shape[-1] % 4 == 0
+ assert direction in ['horizontal', 'vertical', 'diagonal']
+ flipped = bboxes.clone()
+ if direction == 'horizontal':
+ flipped[..., 0::4] = img_shape[1] - bboxes[..., 2::4]
+ flipped[..., 2::4] = img_shape[1] - bboxes[..., 0::4]
+ elif direction == 'vertical':
+ flipped[..., 1::4] = img_shape[0] - bboxes[..., 3::4]
+ flipped[..., 3::4] = img_shape[0] - bboxes[..., 1::4]
+ else:
+ flipped[..., 0::4] = img_shape[1] - bboxes[..., 2::4]
+ flipped[..., 1::4] = img_shape[0] - bboxes[..., 3::4]
+ flipped[..., 2::4] = img_shape[1] - bboxes[..., 0::4]
+ flipped[..., 3::4] = img_shape[0] - bboxes[..., 1::4]
+ return flipped
+
+
+def bbox_mapping(bboxes: Tensor,
+ img_shape: Tuple[int],
+ scale_factor: Union[float, Tuple[float]],
+ flip: bool,
+ flip_direction: str = 'horizontal') -> Tensor:
+ """Map bboxes from the original image scale to testing scale."""
+ new_bboxes = bboxes * bboxes.new_tensor(scale_factor)
+ if flip:
+ new_bboxes = bbox_flip(new_bboxes, img_shape, flip_direction)
+ return new_bboxes
+
+
+def bbox_mapping_back(bboxes: Tensor,
+ img_shape: Tuple[int],
+ scale_factor: Union[float, Tuple[float]],
+ flip: bool,
+ flip_direction: str = 'horizontal') -> Tensor:
+ """Map bboxes from testing scale to original image scale."""
+ new_bboxes = bbox_flip(bboxes, img_shape,
+ flip_direction) if flip else bboxes
+ new_bboxes = new_bboxes.view(-1, 4) / new_bboxes.new_tensor(scale_factor)
+ return new_bboxes.view(bboxes.shape)
+
+
+def bbox2roi(bbox_list: List[Union[Tensor, BaseBoxes]]) -> Tensor:
+ """Convert a list of bboxes to roi format.
+
+ Args:
+ bbox_list (List[Union[Tensor, :obj:`BaseBoxes`]): a list of bboxes
+ corresponding to a batch of images.
+
+ Returns:
+ Tensor: shape (n, box_dim + 1), where ``box_dim`` depends on the
+ different box types. For example, If the box type in ``bbox_list``
+ is HorizontalBoxes, the output shape is (n, 5). Each row of data
+ indicates [batch_ind, x1, y1, x2, y2].
+ """
+ rois_list = []
+ for img_id, bboxes in enumerate(bbox_list):
+ bboxes = get_box_tensor(bboxes)
+ img_inds = bboxes.new_full((bboxes.size(0), 1), img_id)
+ rois = torch.cat([img_inds, bboxes], dim=-1)
+ rois_list.append(rois)
+ rois = torch.cat(rois_list, 0)
+ return rois
+
+
+def roi2bbox(rois: Tensor) -> List[Tensor]:
+ """Convert rois to bounding box format.
+
+ Args:
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+
+ Returns:
+ List[Tensor]: Converted boxes of corresponding rois.
+ """
+ bbox_list = []
+ img_ids = torch.unique(rois[:, 0].cpu(), sorted=True)
+ for img_id in img_ids:
+ inds = (rois[:, 0] == img_id.item())
+ bbox = rois[inds, 1:]
+ bbox_list.append(bbox)
+ return bbox_list
+
+
+# TODO remove later
+def bbox2result(bboxes: Union[Tensor, np.ndarray], labels: Union[Tensor,
+ np.ndarray],
+ num_classes: int) -> List[np.ndarray]:
+ """Convert detection results to a list of numpy arrays.
+
+ Args:
+ bboxes (Tensor | np.ndarray): shape (n, 5)
+ labels (Tensor | np.ndarray): shape (n, )
+ num_classes (int): class number, including background class
+
+ Returns:
+ List(np.ndarray]): bbox results of each class
+ """
+ if bboxes.shape[0] == 0:
+ return [np.zeros((0, 5), dtype=np.float32) for i in range(num_classes)]
+ else:
+ if isinstance(bboxes, torch.Tensor):
+ bboxes = bboxes.detach().cpu().numpy()
+ labels = labels.detach().cpu().numpy()
+ return [bboxes[labels == i, :] for i in range(num_classes)]
+
+
+def distance2bbox(
+ points: Tensor,
+ distance: Tensor,
+ max_shape: Optional[Union[Sequence[int], Tensor,
+ Sequence[Sequence[int]]]] = None
+) -> Tensor:
+ """Decode distance prediction to bounding box.
+
+ Args:
+ points (Tensor): Shape (B, N, 2) or (N, 2).
+ distance (Tensor): Distance from the given point to 4
+ boundaries (left, top, right, bottom). Shape (B, N, 4) or (N, 4)
+ max_shape (Union[Sequence[int], Tensor, Sequence[Sequence[int]]],
+ optional): Maximum bounds for boxes, specifies
+ (H, W, C) or (H, W). If priors shape is (B, N, 4), then
+ the max_shape should be a Sequence[Sequence[int]]
+ and the length of max_shape should also be B.
+
+ Returns:
+ Tensor: Boxes with shape (N, 4) or (B, N, 4)
+ """
+
+ x1 = points[..., 0] - distance[..., 0]
+ y1 = points[..., 1] - distance[..., 1]
+ x2 = points[..., 0] + distance[..., 2]
+ y2 = points[..., 1] + distance[..., 3]
+
+ bboxes = torch.stack([x1, y1, x2, y2], -1)
+
+ if max_shape is not None:
+ if bboxes.dim() == 2 and not torch.onnx.is_in_onnx_export():
+ # speed up
+ bboxes[:, 0::2].clamp_(min=0, max=max_shape[1])
+ bboxes[:, 1::2].clamp_(min=0, max=max_shape[0])
+ return bboxes
+
+ # clip bboxes with dynamic `min` and `max` for onnx
+ if torch.onnx.is_in_onnx_export():
+ # TODO: delete
+ from mmdet.core.export import dynamic_clip_for_onnx
+ x1, y1, x2, y2 = dynamic_clip_for_onnx(x1, y1, x2, y2, max_shape)
+ bboxes = torch.stack([x1, y1, x2, y2], dim=-1)
+ return bboxes
+ if not isinstance(max_shape, torch.Tensor):
+ max_shape = x1.new_tensor(max_shape)
+ max_shape = max_shape[..., :2].type_as(x1)
+ if max_shape.ndim == 2:
+ assert bboxes.ndim == 3
+ assert max_shape.size(0) == bboxes.size(0)
+
+ min_xy = x1.new_tensor(0)
+ max_xy = torch.cat([max_shape, max_shape],
+ dim=-1).flip(-1).unsqueeze(-2)
+ bboxes = torch.where(bboxes < min_xy, min_xy, bboxes)
+ bboxes = torch.where(bboxes > max_xy, max_xy, bboxes)
+
+ return bboxes
+
+
+def bbox2distance(points: Tensor,
+ bbox: Tensor,
+ max_dis: Optional[float] = None,
+ eps: float = 0.1) -> Tensor:
+ """Decode bounding box based on distances.
+
+ Args:
+ points (Tensor): Shape (n, 2) or (b, n, 2), [x, y].
+ bbox (Tensor): Shape (n, 4) or (b, n, 4), "xyxy" format
+ max_dis (float, optional): Upper bound of the distance.
+ eps (float): a small value to ensure target < max_dis, instead <=
+
+ Returns:
+ Tensor: Decoded distances.
+ """
+ left = points[..., 0] - bbox[..., 0]
+ top = points[..., 1] - bbox[..., 1]
+ right = bbox[..., 2] - points[..., 0]
+ bottom = bbox[..., 3] - points[..., 1]
+ if max_dis is not None:
+ left = left.clamp(min=0, max=max_dis - eps)
+ top = top.clamp(min=0, max=max_dis - eps)
+ right = right.clamp(min=0, max=max_dis - eps)
+ bottom = bottom.clamp(min=0, max=max_dis - eps)
+ return torch.stack([left, top, right, bottom], -1)
+
+
+def bbox_rescale(bboxes: Tensor, scale_factor: float = 1.0) -> Tensor:
+ """Rescale bounding box w.r.t. scale_factor.
+
+ Args:
+ bboxes (Tensor): Shape (n, 4) for bboxes or (n, 5) for rois
+ scale_factor (float): rescale factor
+
+ Returns:
+ Tensor: Rescaled bboxes.
+ """
+ if bboxes.size(1) == 5:
+ bboxes_ = bboxes[:, 1:]
+ inds_ = bboxes[:, 0]
+ else:
+ bboxes_ = bboxes
+ cx = (bboxes_[:, 0] + bboxes_[:, 2]) * 0.5
+ cy = (bboxes_[:, 1] + bboxes_[:, 3]) * 0.5
+ w = bboxes_[:, 2] - bboxes_[:, 0]
+ h = bboxes_[:, 3] - bboxes_[:, 1]
+ w = w * scale_factor
+ h = h * scale_factor
+ x1 = cx - 0.5 * w
+ x2 = cx + 0.5 * w
+ y1 = cy - 0.5 * h
+ y2 = cy + 0.5 * h
+ if bboxes.size(1) == 5:
+ rescaled_bboxes = torch.stack([inds_, x1, y1, x2, y2], dim=-1)
+ else:
+ rescaled_bboxes = torch.stack([x1, y1, x2, y2], dim=-1)
+ return rescaled_bboxes
+
+
+def bbox_cxcywh_to_xyxy(bbox: Tensor) -> Tensor:
+ """Convert bbox coordinates from (cx, cy, w, h) to (x1, y1, x2, y2).
+
+ Args:
+ bbox (Tensor): Shape (n, 4) for bboxes.
+
+ Returns:
+ Tensor: Converted bboxes.
+ """
+ cx, cy, w, h = bbox.split((1, 1, 1, 1), dim=-1)
+ bbox_new = [(cx - 0.5 * w), (cy - 0.5 * h), (cx + 0.5 * w), (cy + 0.5 * h)]
+ return torch.cat(bbox_new, dim=-1)
+
+
+def bbox_xyxy_to_cxcywh(bbox: Tensor) -> Tensor:
+ """Convert bbox coordinates from (x1, y1, x2, y2) to (cx, cy, w, h).
+
+ Args:
+ bbox (Tensor): Shape (n, 4) for bboxes.
+
+ Returns:
+ Tensor: Converted bboxes.
+ """
+ x1, y1, x2, y2 = bbox.split((1, 1, 1, 1), dim=-1)
+ bbox_new = [(x1 + x2) / 2, (y1 + y2) / 2, (x2 - x1), (y2 - y1)]
+ return torch.cat(bbox_new, dim=-1)
+
+
+def bbox2corner(bboxes: torch.Tensor) -> torch.Tensor:
+ """Convert bbox coordinates from (x1, y1, x2, y2) to corners ((x1, y1),
+ (x2, y1), (x1, y2), (x2, y2)).
+
+ Args:
+ bboxes (Tensor): Shape (n, 4) for bboxes.
+ Returns:
+ Tensor: Shape (n*4, 2) for corners.
+ """
+ x1, y1, x2, y2 = torch.split(bboxes, 1, dim=1)
+ return torch.cat([x1, y1, x2, y1, x1, y2, x2, y2], dim=1).reshape(-1, 2)
+
+
+def corner2bbox(corners: torch.Tensor) -> torch.Tensor:
+ """Convert bbox coordinates from corners ((x1, y1), (x2, y1), (x1, y2),
+ (x2, y2)) to (x1, y1, x2, y2).
+
+ Args:
+ corners (Tensor): Shape (n*4, 2) for corners.
+ Returns:
+ Tensor: Shape (n, 4) for bboxes.
+ """
+ corners = corners.reshape(-1, 4, 2)
+ min_xy = corners.min(dim=1)[0]
+ max_xy = corners.max(dim=1)[0]
+ return torch.cat([min_xy, max_xy], dim=1)
+
+
+def bbox_project(
+ bboxes: Union[torch.Tensor, np.ndarray],
+ homography_matrix: Union[torch.Tensor, np.ndarray],
+ img_shape: Optional[Tuple[int, int]] = None
+) -> Union[torch.Tensor, np.ndarray]:
+ """Geometric transformation for bbox.
+
+ Args:
+ bboxes (Union[torch.Tensor, np.ndarray]): Shape (n, 4) for bboxes.
+ homography_matrix (Union[torch.Tensor, np.ndarray]):
+ Shape (3, 3) for geometric transformation.
+ img_shape (Tuple[int, int], optional): Image shape. Defaults to None.
+ Returns:
+ Union[torch.Tensor, np.ndarray]: Converted bboxes.
+ """
+ bboxes_type = type(bboxes)
+ if bboxes_type is np.ndarray:
+ bboxes = torch.from_numpy(bboxes)
+ if isinstance(homography_matrix, np.ndarray):
+ homography_matrix = torch.from_numpy(homography_matrix)
+ corners = bbox2corner(bboxes)
+ corners = torch.cat(
+ [corners, corners.new_ones(corners.shape[0], 1)], dim=1)
+ corners = torch.matmul(homography_matrix, corners.t()).t()
+ # Convert to homogeneous coordinates by normalization
+ corners = corners[:, :2] / corners[:, 2:3]
+ bboxes = corner2bbox(corners)
+ if img_shape is not None:
+ bboxes[:, 0::2] = bboxes[:, 0::2].clamp(0, img_shape[1])
+ bboxes[:, 1::2] = bboxes[:, 1::2].clamp(0, img_shape[0])
+ if bboxes_type is np.ndarray:
+ bboxes = bboxes.numpy()
+ return bboxes
+
+
+def cat_boxes(data_list: List[Union[Tensor, BaseBoxes]],
+ dim: int = 0) -> Union[Tensor, BaseBoxes]:
+ """Concatenate boxes with type of tensor or box type.
+
+ Args:
+ data_list (List[Union[Tensor, :obj:`BaseBoxes`]]): A list of tensors
+ or box types need to be concatenated.
+ dim (int): The dimension over which the box are concatenated.
+ Defaults to 0.
+
+ Returns:
+ Union[Tensor, :obj`BaseBoxes`]: Concatenated results.
+ """
+ if data_list and isinstance(data_list[0], BaseBoxes):
+ return data_list[0].cat(data_list, dim=dim)
+ else:
+ return torch.cat(data_list, dim=dim)
+
+
+def stack_boxes(data_list: List[Union[Tensor, BaseBoxes]],
+ dim: int = 0) -> Union[Tensor, BaseBoxes]:
+ """Stack boxes with type of tensor or box type.
+
+ Args:
+ data_list (List[Union[Tensor, :obj:`BaseBoxes`]]): A list of tensors
+ or box types need to be stacked.
+ dim (int): The dimension over which the box are stacked.
+ Defaults to 0.
+
+ Returns:
+ Union[Tensor, :obj`BaseBoxes`]: Stacked results.
+ """
+ if data_list and isinstance(data_list[0], BaseBoxes):
+ return data_list[0].stack(data_list, dim=dim)
+ else:
+ return torch.stack(data_list, dim=dim)
+
+
+def scale_boxes(boxes: Union[Tensor, BaseBoxes],
+ scale_factor: Tuple[float, float]) -> Union[Tensor, BaseBoxes]:
+ """Scale boxes with type of tensor or box type.
+
+ Args:
+ boxes (Tensor or :obj:`BaseBoxes`): boxes need to be scaled. Its type
+ can be a tensor or a box type.
+ scale_factor (Tuple[float, float]): factors for scaling boxes.
+ The length should be 2.
+
+ Returns:
+ Union[Tensor, :obj:`BaseBoxes`]: Scaled boxes.
+ """
+ if isinstance(boxes, BaseBoxes):
+ boxes.rescale_(scale_factor)
+ return boxes
+ else:
+ # Tensor boxes will be treated as horizontal boxes
+ repeat_num = int(boxes.size(-1) / 2)
+ scale_factor = boxes.new_tensor(scale_factor).repeat((1, repeat_num))
+ return boxes * scale_factor
+
+
+def get_box_wh(boxes: Union[Tensor, BaseBoxes]) -> Tuple[Tensor, Tensor]:
+ """Get the width and height of boxes with type of tensor or box type.
+
+ Args:
+ boxes (Tensor or :obj:`BaseBoxes`): boxes with type of tensor
+ or box type.
+
+ Returns:
+ Tuple[Tensor, Tensor]: the width and height of boxes.
+ """
+ if isinstance(boxes, BaseBoxes):
+ w = boxes.widths
+ h = boxes.heights
+ else:
+ # Tensor boxes will be treated as horizontal boxes by defaults
+ w = boxes[:, 2] - boxes[:, 0]
+ h = boxes[:, 3] - boxes[:, 1]
+ return w, h
+
+
+def get_box_tensor(boxes: Union[Tensor, BaseBoxes]) -> Tensor:
+ """Get tensor data from box type boxes.
+
+ Args:
+ boxes (Tensor or BaseBoxes): boxes with type of tensor or box type.
+ If its type is a tensor, the boxes will be directly returned.
+ If its type is a box type, the `boxes.tensor` will be returned.
+
+ Returns:
+ Tensor: boxes tensor.
+ """
+ if isinstance(boxes, BaseBoxes):
+ boxes = boxes.tensor
+ return boxes
+
+
+def empty_box_as(boxes: Union[Tensor, BaseBoxes]) -> Union[Tensor, BaseBoxes]:
+ """Generate empty box according to input ``boxes` type and device.
+
+ Args:
+ boxes (Tensor or :obj:`BaseBoxes`): boxes with type of tensor
+ or box type.
+
+ Returns:
+ Union[Tensor, BaseBoxes]: Generated empty box.
+ """
+ if isinstance(boxes, BaseBoxes):
+ return boxes.empty_boxes()
+ else:
+ # Tensor boxes will be treated as horizontal boxes by defaults
+ return boxes.new_zeros(0, 4)
+
+
+def bbox_xyxy_to_cxcyah(bboxes: torch.Tensor) -> torch.Tensor:
+ """Convert bbox coordinates from (x1, y1, x2, y2) to (cx, cy, ratio, h).
+
+ Args:
+ bbox (Tensor): Shape (n, 4) for bboxes.
+
+ Returns:
+ Tensor: Converted bboxes.
+ """
+ cx = (bboxes[:, 2] + bboxes[:, 0]) / 2
+ cy = (bboxes[:, 3] + bboxes[:, 1]) / 2
+ w = bboxes[:, 2] - bboxes[:, 0]
+ h = bboxes[:, 3] - bboxes[:, 1]
+ xyah = torch.stack([cx, cy, w / h, h], -1)
+ return xyah
+
+
+def bbox_cxcyah_to_xyxy(bboxes: torch.Tensor) -> torch.Tensor:
+ """Convert bbox coordinates from (cx, cy, ratio, h) to (x1, y1, x2, y2).
+
+ Args:
+ bbox (Tensor): Shape (n, 4) for bboxes.
+ Returns:
+ Tensor: Converted bboxes.
+ """
+ cx, cy, ratio, h = bboxes.split((1, 1, 1, 1), dim=-1)
+ w = ratio * h
+ x1y1x2y2 = [cx - w / 2.0, cy - h / 2.0, cx + w / 2.0, cy + h / 2.0]
+ return torch.cat(x1y1x2y2, dim=-1)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/det_data_sample.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/det_data_sample.py
new file mode 100644
index 0000000000000000000000000000000000000000..37dd74725ed2ff5eb8a088c9d23a9ac5469b07a3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/det_data_sample.py
@@ -0,0 +1,237 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional
+
+from mmengine.structures import BaseDataElement, InstanceData, PixelData
+
+
+class DetDataSample(BaseDataElement):
+ """A data structure interface of MMDetection. They are used as interfaces
+ between different components.
+
+ The attributes in ``DetDataSample`` are divided into several parts:
+
+ - ``proposals``(InstanceData): Region proposals used in two-stage
+ detectors.
+ - ``gt_instances``(InstanceData): Ground truth of instance annotations.
+ - ``pred_instances``(InstanceData): Instances of detection predictions.
+ - ``pred_track_instances``(InstanceData): Instances of tracking
+ predictions.
+ - ``ignored_instances``(InstanceData): Instances to be ignored during
+ training/testing.
+ - ``gt_panoptic_seg``(PixelData): Ground truth of panoptic
+ segmentation.
+ - ``pred_panoptic_seg``(PixelData): Prediction of panoptic
+ segmentation.
+ - ``gt_sem_seg``(PixelData): Ground truth of semantic segmentation.
+ - ``pred_sem_seg``(PixelData): Prediction of semantic segmentation.
+
+ Examples:
+ >>> import torch
+ >>> import numpy as np
+ >>> from mmengine.structures import InstanceData
+ >>> from mmdet.structures import DetDataSample
+
+ >>> data_sample = DetDataSample()
+ >>> img_meta = dict(img_shape=(800, 1196),
+ ... pad_shape=(800, 1216))
+ >>> gt_instances = InstanceData(metainfo=img_meta)
+ >>> gt_instances.bboxes = torch.rand((5, 4))
+ >>> gt_instances.labels = torch.rand((5,))
+ >>> data_sample.gt_instances = gt_instances
+ >>> assert 'img_shape' in data_sample.gt_instances.metainfo_keys()
+ >>> len(data_sample.gt_instances)
+ 5
+ >>> print(data_sample)
+
+ ) at 0x7f21fb1b9880>
+ >>> pred_instances = InstanceData(metainfo=img_meta)
+ >>> pred_instances.bboxes = torch.rand((5, 4))
+ >>> pred_instances.scores = torch.rand((5,))
+ >>> data_sample = DetDataSample(pred_instances=pred_instances)
+ >>> assert 'pred_instances' in data_sample
+
+ >>> pred_track_instances = InstanceData(metainfo=img_meta)
+ >>> pred_track_instances.bboxes = torch.rand((5, 4))
+ >>> pred_track_instances.scores = torch.rand((5,))
+ >>> data_sample = DetDataSample(
+ ... pred_track_instances=pred_track_instances)
+ >>> assert 'pred_track_instances' in data_sample
+
+ >>> data_sample = DetDataSample()
+ >>> gt_instances_data = dict(
+ ... bboxes=torch.rand(2, 4),
+ ... labels=torch.rand(2),
+ ... masks=np.random.rand(2, 2, 2))
+ >>> gt_instances = InstanceData(**gt_instances_data)
+ >>> data_sample.gt_instances = gt_instances
+ >>> assert 'gt_instances' in data_sample
+ >>> assert 'masks' in data_sample.gt_instances
+
+ >>> data_sample = DetDataSample()
+ >>> gt_panoptic_seg_data = dict(panoptic_seg=torch.rand(2, 4))
+ >>> gt_panoptic_seg = PixelData(**gt_panoptic_seg_data)
+ >>> data_sample.gt_panoptic_seg = gt_panoptic_seg
+ >>> print(data_sample)
+
+ gt_panoptic_seg:
+ ) at 0x7f66c2bb7280>
+ >>> data_sample = DetDataSample()
+ >>> gt_segm_seg_data = dict(segm_seg=torch.rand(2, 2, 2))
+ >>> gt_segm_seg = PixelData(**gt_segm_seg_data)
+ >>> data_sample.gt_segm_seg = gt_segm_seg
+ >>> assert 'gt_segm_seg' in data_sample
+ >>> assert 'segm_seg' in data_sample.gt_segm_seg
+ """
+
+ @property
+ def proposals(self) -> InstanceData:
+ return self._proposals
+
+ @proposals.setter
+ def proposals(self, value: InstanceData):
+ self.set_field(value, '_proposals', dtype=InstanceData)
+
+ @proposals.deleter
+ def proposals(self):
+ del self._proposals
+
+ @property
+ def gt_instances(self) -> InstanceData:
+ return self._gt_instances
+
+ @gt_instances.setter
+ def gt_instances(self, value: InstanceData):
+ self.set_field(value, '_gt_instances', dtype=InstanceData)
+
+ @gt_instances.deleter
+ def gt_instances(self):
+ del self._gt_instances
+
+ @property
+ def pred_instances(self) -> InstanceData:
+ return self._pred_instances
+
+ @pred_instances.setter
+ def pred_instances(self, value: InstanceData):
+ self.set_field(value, '_pred_instances', dtype=InstanceData)
+
+ @pred_instances.deleter
+ def pred_instances(self):
+ del self._pred_instances
+
+ # directly add ``pred_track_instances`` in ``DetDataSample``
+ # so that the ``TrackDataSample`` does not bother to access the
+ # instance-level information.
+ @property
+ def pred_track_instances(self) -> InstanceData:
+ return self._pred_track_instances
+
+ @pred_track_instances.setter
+ def pred_track_instances(self, value: InstanceData):
+ self.set_field(value, '_pred_track_instances', dtype=InstanceData)
+
+ @pred_track_instances.deleter
+ def pred_track_instances(self):
+ del self._pred_track_instances
+
+ @property
+ def ignored_instances(self) -> InstanceData:
+ return self._ignored_instances
+
+ @ignored_instances.setter
+ def ignored_instances(self, value: InstanceData):
+ self.set_field(value, '_ignored_instances', dtype=InstanceData)
+
+ @ignored_instances.deleter
+ def ignored_instances(self):
+ del self._ignored_instances
+
+ @property
+ def gt_panoptic_seg(self) -> PixelData:
+ return self._gt_panoptic_seg
+
+ @gt_panoptic_seg.setter
+ def gt_panoptic_seg(self, value: PixelData):
+ self.set_field(value, '_gt_panoptic_seg', dtype=PixelData)
+
+ @gt_panoptic_seg.deleter
+ def gt_panoptic_seg(self):
+ del self._gt_panoptic_seg
+
+ @property
+ def pred_panoptic_seg(self) -> PixelData:
+ return self._pred_panoptic_seg
+
+ @pred_panoptic_seg.setter
+ def pred_panoptic_seg(self, value: PixelData):
+ self.set_field(value, '_pred_panoptic_seg', dtype=PixelData)
+
+ @pred_panoptic_seg.deleter
+ def pred_panoptic_seg(self):
+ del self._pred_panoptic_seg
+
+ @property
+ def gt_sem_seg(self) -> PixelData:
+ return self._gt_sem_seg
+
+ @gt_sem_seg.setter
+ def gt_sem_seg(self, value: PixelData):
+ self.set_field(value, '_gt_sem_seg', dtype=PixelData)
+
+ @gt_sem_seg.deleter
+ def gt_sem_seg(self):
+ del self._gt_sem_seg
+
+ @property
+ def pred_sem_seg(self) -> PixelData:
+ return self._pred_sem_seg
+
+ @pred_sem_seg.setter
+ def pred_sem_seg(self, value: PixelData):
+ self.set_field(value, '_pred_sem_seg', dtype=PixelData)
+
+ @pred_sem_seg.deleter
+ def pred_sem_seg(self):
+ del self._pred_sem_seg
+
+
+SampleList = List[DetDataSample]
+OptSampleList = Optional[SampleList]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/mask/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/mask/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..f78394701df1b493259c4c23a79aea5c5cb8be95
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/mask/__init__.py
@@ -0,0 +1,11 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .mask_target import mask_target
+from .structures import (BaseInstanceMasks, BitmapMasks, PolygonMasks,
+ bitmap_to_polygon, polygon_to_bitmap)
+from .utils import encode_mask_results, mask2bbox, split_combined_polys
+
+__all__ = [
+ 'split_combined_polys', 'mask_target', 'BaseInstanceMasks', 'BitmapMasks',
+ 'PolygonMasks', 'encode_mask_results', 'mask2bbox', 'polygon_to_bitmap',
+ 'bitmap_to_polygon'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/mask/mask_target.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/mask/mask_target.py
new file mode 100644
index 0000000000000000000000000000000000000000..b2fc5f1878300446b114c9f57c6a885fea8c927c
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/mask/mask_target.py
@@ -0,0 +1,127 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import numpy as np
+import torch
+from torch.nn.modules.utils import _pair
+
+
+def mask_target(pos_proposals_list, pos_assigned_gt_inds_list, gt_masks_list,
+ cfg):
+ """Compute mask target for positive proposals in multiple images.
+
+ Args:
+ pos_proposals_list (list[Tensor]): Positive proposals in multiple
+ images, each has shape (num_pos, 4).
+ pos_assigned_gt_inds_list (list[Tensor]): Assigned GT indices for each
+ positive proposals, each has shape (num_pos,).
+ gt_masks_list (list[:obj:`BaseInstanceMasks`]): Ground truth masks of
+ each image.
+ cfg (dict): Config dict that specifies the mask size.
+
+ Returns:
+ Tensor: Mask target of each image, has shape (num_pos, w, h).
+
+ Example:
+ >>> from mmengine.config import Config
+ >>> import mmdet
+ >>> from mmdet.data_elements.mask import BitmapMasks
+ >>> from mmdet.data_elements.mask.mask_target import *
+ >>> H, W = 17, 18
+ >>> cfg = Config({'mask_size': (13, 14)})
+ >>> rng = np.random.RandomState(0)
+ >>> # Positive proposals (tl_x, tl_y, br_x, br_y) for each image
+ >>> pos_proposals_list = [
+ >>> torch.Tensor([
+ >>> [ 7.2425, 5.5929, 13.9414, 14.9541],
+ >>> [ 7.3241, 3.6170, 16.3850, 15.3102],
+ >>> ]),
+ >>> torch.Tensor([
+ >>> [ 4.8448, 6.4010, 7.0314, 9.7681],
+ >>> [ 5.9790, 2.6989, 7.4416, 4.8580],
+ >>> [ 0.0000, 0.0000, 0.1398, 9.8232],
+ >>> ]),
+ >>> ]
+ >>> # Corresponding class index for each proposal for each image
+ >>> pos_assigned_gt_inds_list = [
+ >>> torch.LongTensor([7, 0]),
+ >>> torch.LongTensor([5, 4, 1]),
+ >>> ]
+ >>> # Ground truth mask for each true object for each image
+ >>> gt_masks_list = [
+ >>> BitmapMasks(rng.rand(8, H, W), height=H, width=W),
+ >>> BitmapMasks(rng.rand(6, H, W), height=H, width=W),
+ >>> ]
+ >>> mask_targets = mask_target(
+ >>> pos_proposals_list, pos_assigned_gt_inds_list,
+ >>> gt_masks_list, cfg)
+ >>> assert mask_targets.shape == (5,) + cfg['mask_size']
+ """
+ cfg_list = [cfg for _ in range(len(pos_proposals_list))]
+ mask_targets = map(mask_target_single, pos_proposals_list,
+ pos_assigned_gt_inds_list, gt_masks_list, cfg_list)
+ mask_targets = list(mask_targets)
+ if len(mask_targets) > 0:
+ mask_targets = torch.cat(mask_targets)
+ return mask_targets
+
+
+def mask_target_single(pos_proposals, pos_assigned_gt_inds, gt_masks, cfg):
+ """Compute mask target for each positive proposal in the image.
+
+ Args:
+ pos_proposals (Tensor): Positive proposals.
+ pos_assigned_gt_inds (Tensor): Assigned GT inds of positive proposals.
+ gt_masks (:obj:`BaseInstanceMasks`): GT masks in the format of Bitmap
+ or Polygon.
+ cfg (dict): Config dict that indicate the mask size.
+
+ Returns:
+ Tensor: Mask target of each positive proposals in the image.
+
+ Example:
+ >>> from mmengine.config import Config
+ >>> import mmdet
+ >>> from mmdet.data_elements.mask import BitmapMasks
+ >>> from mmdet.data_elements.mask.mask_target import * # NOQA
+ >>> H, W = 32, 32
+ >>> cfg = Config({'mask_size': (7, 11)})
+ >>> rng = np.random.RandomState(0)
+ >>> # Masks for each ground truth box (relative to the image)
+ >>> gt_masks_data = rng.rand(3, H, W)
+ >>> gt_masks = BitmapMasks(gt_masks_data, height=H, width=W)
+ >>> # Predicted positive boxes in one image
+ >>> pos_proposals = torch.FloatTensor([
+ >>> [ 16.2, 5.5, 19.9, 20.9],
+ >>> [ 17.3, 13.6, 19.3, 19.3],
+ >>> [ 14.8, 16.4, 17.0, 23.7],
+ >>> [ 0.0, 0.0, 16.0, 16.0],
+ >>> [ 4.0, 0.0, 20.0, 16.0],
+ >>> ])
+ >>> # For each predicted proposal, its assignment to a gt mask
+ >>> pos_assigned_gt_inds = torch.LongTensor([0, 1, 2, 1, 1])
+ >>> mask_targets = mask_target_single(
+ >>> pos_proposals, pos_assigned_gt_inds, gt_masks, cfg)
+ >>> assert mask_targets.shape == (5,) + cfg['mask_size']
+ """
+ device = pos_proposals.device
+ mask_size = _pair(cfg.mask_size)
+ binarize = not cfg.get('soft_mask_target', False)
+ num_pos = pos_proposals.size(0)
+ if num_pos > 0:
+ proposals_np = pos_proposals.cpu().numpy()
+ maxh, maxw = gt_masks.height, gt_masks.width
+ proposals_np[:, [0, 2]] = np.clip(proposals_np[:, [0, 2]], 0, maxw)
+ proposals_np[:, [1, 3]] = np.clip(proposals_np[:, [1, 3]], 0, maxh)
+ pos_assigned_gt_inds = pos_assigned_gt_inds.cpu().numpy()
+
+ mask_targets = gt_masks.crop_and_resize(
+ proposals_np,
+ mask_size,
+ device=device,
+ inds=pos_assigned_gt_inds,
+ binarize=binarize).to_ndarray()
+
+ mask_targets = torch.from_numpy(mask_targets).float().to(device)
+ else:
+ mask_targets = pos_proposals.new_zeros((0, ) + mask_size)
+
+ return mask_targets
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/mask/structures.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/mask/structures.py
new file mode 100644
index 0000000000000000000000000000000000000000..b4fdd27570b0d11d92eba4e8f854e153750135a4
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/mask/structures.py
@@ -0,0 +1,1193 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import itertools
+from abc import ABCMeta, abstractmethod
+from typing import Sequence, Type, TypeVar
+
+import cv2
+import mmcv
+import numpy as np
+import pycocotools.mask as maskUtils
+import shapely.geometry as geometry
+import torch
+from mmcv.ops.roi_align import roi_align
+
+T = TypeVar('T')
+
+
+class BaseInstanceMasks(metaclass=ABCMeta):
+ """Base class for instance masks."""
+
+ @abstractmethod
+ def rescale(self, scale, interpolation='nearest'):
+ """Rescale masks as large as possible while keeping the aspect ratio.
+ For details can refer to `mmcv.imrescale`.
+
+ Args:
+ scale (tuple[int]): The maximum size (h, w) of rescaled mask.
+ interpolation (str): Same as :func:`mmcv.imrescale`.
+
+ Returns:
+ BaseInstanceMasks: The rescaled masks.
+ """
+
+ @abstractmethod
+ def resize(self, out_shape, interpolation='nearest'):
+ """Resize masks to the given out_shape.
+
+ Args:
+ out_shape: Target (h, w) of resized mask.
+ interpolation (str): See :func:`mmcv.imresize`.
+
+ Returns:
+ BaseInstanceMasks: The resized masks.
+ """
+
+ @abstractmethod
+ def flip(self, flip_direction='horizontal'):
+ """Flip masks alone the given direction.
+
+ Args:
+ flip_direction (str): Either 'horizontal' or 'vertical'.
+
+ Returns:
+ BaseInstanceMasks: The flipped masks.
+ """
+
+ @abstractmethod
+ def pad(self, out_shape, pad_val):
+ """Pad masks to the given size of (h, w).
+
+ Args:
+ out_shape (tuple[int]): Target (h, w) of padded mask.
+ pad_val (int): The padded value.
+
+ Returns:
+ BaseInstanceMasks: The padded masks.
+ """
+
+ @abstractmethod
+ def crop(self, bbox):
+ """Crop each mask by the given bbox.
+
+ Args:
+ bbox (ndarray): Bbox in format [x1, y1, x2, y2], shape (4, ).
+
+ Return:
+ BaseInstanceMasks: The cropped masks.
+ """
+
+ @abstractmethod
+ def crop_and_resize(self,
+ bboxes,
+ out_shape,
+ inds,
+ device,
+ interpolation='bilinear',
+ binarize=True):
+ """Crop and resize masks by the given bboxes.
+
+ This function is mainly used in mask targets computation.
+ It firstly align mask to bboxes by assigned_inds, then crop mask by the
+ assigned bbox and resize to the size of (mask_h, mask_w)
+
+ Args:
+ bboxes (Tensor): Bboxes in format [x1, y1, x2, y2], shape (N, 4)
+ out_shape (tuple[int]): Target (h, w) of resized mask
+ inds (ndarray): Indexes to assign masks to each bbox,
+ shape (N,) and values should be between [0, num_masks - 1].
+ device (str): Device of bboxes
+ interpolation (str): See `mmcv.imresize`
+ binarize (bool): if True fractional values are rounded to 0 or 1
+ after the resize operation. if False and unsupported an error
+ will be raised. Defaults to True.
+
+ Return:
+ BaseInstanceMasks: the cropped and resized masks.
+ """
+
+ @abstractmethod
+ def expand(self, expanded_h, expanded_w, top, left):
+ """see :class:`Expand`."""
+
+ @property
+ @abstractmethod
+ def areas(self):
+ """ndarray: areas of each instance."""
+
+ @abstractmethod
+ def to_ndarray(self):
+ """Convert masks to the format of ndarray.
+
+ Return:
+ ndarray: Converted masks in the format of ndarray.
+ """
+
+ @abstractmethod
+ def to_tensor(self, dtype, device):
+ """Convert masks to the format of Tensor.
+
+ Args:
+ dtype (str): Dtype of converted mask.
+ device (torch.device): Device of converted masks.
+
+ Returns:
+ Tensor: Converted masks in the format of Tensor.
+ """
+
+ @abstractmethod
+ def translate(self,
+ out_shape,
+ offset,
+ direction='horizontal',
+ border_value=0,
+ interpolation='bilinear'):
+ """Translate the masks.
+
+ Args:
+ out_shape (tuple[int]): Shape for output mask, format (h, w).
+ offset (int | float): The offset for translate.
+ direction (str): The translate direction, either "horizontal"
+ or "vertical".
+ border_value (int | float): Border value. Default 0.
+ interpolation (str): Same as :func:`mmcv.imtranslate`.
+
+ Returns:
+ Translated masks.
+ """
+
+ def shear(self,
+ out_shape,
+ magnitude,
+ direction='horizontal',
+ border_value=0,
+ interpolation='bilinear'):
+ """Shear the masks.
+
+ Args:
+ out_shape (tuple[int]): Shape for output mask, format (h, w).
+ magnitude (int | float): The magnitude used for shear.
+ direction (str): The shear direction, either "horizontal"
+ or "vertical".
+ border_value (int | tuple[int]): Value used in case of a
+ constant border. Default 0.
+ interpolation (str): Same as in :func:`mmcv.imshear`.
+
+ Returns:
+ ndarray: Sheared masks.
+ """
+
+ @abstractmethod
+ def rotate(self, out_shape, angle, center=None, scale=1.0, border_value=0):
+ """Rotate the masks.
+
+ Args:
+ out_shape (tuple[int]): Shape for output mask, format (h, w).
+ angle (int | float): Rotation angle in degrees. Positive values
+ mean counter-clockwise rotation.
+ center (tuple[float], optional): Center point (w, h) of the
+ rotation in source image. If not specified, the center of
+ the image will be used.
+ scale (int | float): Isotropic scale factor.
+ border_value (int | float): Border value. Default 0 for masks.
+
+ Returns:
+ Rotated masks.
+ """
+
+ def get_bboxes(self, dst_type='hbb'):
+ """Get the certain type boxes from masks.
+
+ Please refer to ``mmdet.structures.bbox.box_type`` for more details of
+ the box type.
+
+ Args:
+ dst_type: Destination box type.
+
+ Returns:
+ :obj:`BaseBoxes`: Certain type boxes.
+ """
+ from ..bbox import get_box_type
+ _, box_type_cls = get_box_type(dst_type)
+ return box_type_cls.from_instance_masks(self)
+
+ @classmethod
+ @abstractmethod
+ def cat(cls: Type[T], masks: Sequence[T]) -> T:
+ """Concatenate a sequence of masks into one single mask instance.
+
+ Args:
+ masks (Sequence[T]): A sequence of mask instances.
+
+ Returns:
+ T: Concatenated mask instance.
+ """
+
+
+class BitmapMasks(BaseInstanceMasks):
+ """This class represents masks in the form of bitmaps.
+
+ Args:
+ masks (ndarray): ndarray of masks in shape (N, H, W), where N is
+ the number of objects.
+ height (int): height of masks
+ width (int): width of masks
+
+ Example:
+ >>> from mmdet.data_elements.mask.structures import * # NOQA
+ >>> num_masks, H, W = 3, 32, 32
+ >>> rng = np.random.RandomState(0)
+ >>> masks = (rng.rand(num_masks, H, W) > 0.1).astype(np.int64)
+ >>> self = BitmapMasks(masks, height=H, width=W)
+
+ >>> # demo crop_and_resize
+ >>> num_boxes = 5
+ >>> bboxes = np.array([[0, 0, 30, 10.0]] * num_boxes)
+ >>> out_shape = (14, 14)
+ >>> inds = torch.randint(0, len(self), size=(num_boxes,))
+ >>> device = 'cpu'
+ >>> interpolation = 'bilinear'
+ >>> new = self.crop_and_resize(
+ ... bboxes, out_shape, inds, device, interpolation)
+ >>> assert len(new) == num_boxes
+ >>> assert new.height, new.width == out_shape
+ """
+
+ def __init__(self, masks, height, width):
+ self.height = height
+ self.width = width
+ if len(masks) == 0:
+ self.masks = np.empty((0, self.height, self.width), dtype=np.uint8)
+ else:
+ assert isinstance(masks, (list, np.ndarray))
+ if isinstance(masks, list):
+ assert isinstance(masks[0], np.ndarray)
+ assert masks[0].ndim == 2 # (H, W)
+ else:
+ assert masks.ndim == 3 # (N, H, W)
+
+ self.masks = np.stack(masks).reshape(-1, height, width)
+ assert self.masks.shape[1] == self.height
+ assert self.masks.shape[2] == self.width
+
+ def __getitem__(self, index):
+ """Index the BitmapMask.
+
+ Args:
+ index (int | ndarray): Indices in the format of integer or ndarray.
+
+ Returns:
+ :obj:`BitmapMasks`: Indexed bitmap masks.
+ """
+ masks = self.masks[index].reshape(-1, self.height, self.width)
+ return BitmapMasks(masks, self.height, self.width)
+
+ def __iter__(self):
+ return iter(self.masks)
+
+ def __repr__(self):
+ s = self.__class__.__name__ + '('
+ s += f'num_masks={len(self.masks)}, '
+ s += f'height={self.height}, '
+ s += f'width={self.width})'
+ return s
+
+ def __len__(self):
+ """Number of masks."""
+ return len(self.masks)
+
+ def rescale(self, scale, interpolation='nearest'):
+ """See :func:`BaseInstanceMasks.rescale`."""
+ if len(self.masks) == 0:
+ new_w, new_h = mmcv.rescale_size((self.width, self.height), scale)
+ rescaled_masks = np.empty((0, new_h, new_w), dtype=np.uint8)
+ else:
+ rescaled_masks = np.stack([
+ mmcv.imrescale(mask, scale, interpolation=interpolation)
+ for mask in self.masks
+ ])
+ height, width = rescaled_masks.shape[1:]
+ return BitmapMasks(rescaled_masks, height, width)
+
+ def resize(self, out_shape, interpolation='nearest'):
+ """See :func:`BaseInstanceMasks.resize`."""
+ if len(self.masks) == 0:
+ resized_masks = np.empty((0, *out_shape), dtype=np.uint8)
+ else:
+ resized_masks = np.stack([
+ mmcv.imresize(
+ mask, out_shape[::-1], interpolation=interpolation)
+ for mask in self.masks
+ ])
+ return BitmapMasks(resized_masks, *out_shape)
+
+ def flip(self, flip_direction='horizontal'):
+ """See :func:`BaseInstanceMasks.flip`."""
+ assert flip_direction in ('horizontal', 'vertical', 'diagonal')
+
+ if len(self.masks) == 0:
+ flipped_masks = self.masks
+ else:
+ flipped_masks = np.stack([
+ mmcv.imflip(mask, direction=flip_direction)
+ for mask in self.masks
+ ])
+ return BitmapMasks(flipped_masks, self.height, self.width)
+
+ def pad(self, out_shape, pad_val=0):
+ """See :func:`BaseInstanceMasks.pad`."""
+ if len(self.masks) == 0:
+ padded_masks = np.empty((0, *out_shape), dtype=np.uint8)
+ else:
+ padded_masks = np.stack([
+ mmcv.impad(mask, shape=out_shape, pad_val=pad_val)
+ for mask in self.masks
+ ])
+ return BitmapMasks(padded_masks, *out_shape)
+
+ def crop(self, bbox):
+ """See :func:`BaseInstanceMasks.crop`."""
+ assert isinstance(bbox, np.ndarray)
+ assert bbox.ndim == 1
+
+ # clip the boundary
+ bbox = bbox.copy()
+ bbox[0::2] = np.clip(bbox[0::2], 0, self.width)
+ bbox[1::2] = np.clip(bbox[1::2], 0, self.height)
+ x1, y1, x2, y2 = bbox
+ w = np.maximum(x2 - x1, 1)
+ h = np.maximum(y2 - y1, 1)
+
+ if len(self.masks) == 0:
+ cropped_masks = np.empty((0, h, w), dtype=np.uint8)
+ else:
+ cropped_masks = self.masks[:, y1:y1 + h, x1:x1 + w]
+ return BitmapMasks(cropped_masks, h, w)
+
+ def crop_and_resize(self,
+ bboxes,
+ out_shape,
+ inds,
+ device='cpu',
+ interpolation='bilinear',
+ binarize=True):
+ """See :func:`BaseInstanceMasks.crop_and_resize`."""
+ if len(self.masks) == 0:
+ empty_masks = np.empty((0, *out_shape), dtype=np.uint8)
+ return BitmapMasks(empty_masks, *out_shape)
+
+ # convert bboxes to tensor
+ if isinstance(bboxes, np.ndarray):
+ bboxes = torch.from_numpy(bboxes).to(device=device)
+ if isinstance(inds, np.ndarray):
+ inds = torch.from_numpy(inds).to(device=device)
+
+ num_bbox = bboxes.shape[0]
+ fake_inds = torch.arange(
+ num_bbox, device=device).to(dtype=bboxes.dtype)[:, None]
+ rois = torch.cat([fake_inds, bboxes], dim=1) # Nx5
+ rois = rois.to(device=device)
+ if num_bbox > 0:
+ gt_masks_th = torch.from_numpy(self.masks).to(device).index_select(
+ 0, inds).to(dtype=rois.dtype)
+ targets = roi_align(gt_masks_th[:, None, :, :], rois, out_shape,
+ 1.0, 0, 'avg', True).squeeze(1)
+ if binarize:
+ resized_masks = (targets >= 0.5).cpu().numpy()
+ else:
+ resized_masks = targets.cpu().numpy()
+ else:
+ resized_masks = []
+ return BitmapMasks(resized_masks, *out_shape)
+
+ def expand(self, expanded_h, expanded_w, top, left):
+ """See :func:`BaseInstanceMasks.expand`."""
+ if len(self.masks) == 0:
+ expanded_mask = np.empty((0, expanded_h, expanded_w),
+ dtype=np.uint8)
+ else:
+ expanded_mask = np.zeros((len(self), expanded_h, expanded_w),
+ dtype=np.uint8)
+ expanded_mask[:, top:top + self.height,
+ left:left + self.width] = self.masks
+ return BitmapMasks(expanded_mask, expanded_h, expanded_w)
+
+ def translate(self,
+ out_shape,
+ offset,
+ direction='horizontal',
+ border_value=0,
+ interpolation='bilinear'):
+ """Translate the BitmapMasks.
+
+ Args:
+ out_shape (tuple[int]): Shape for output mask, format (h, w).
+ offset (int | float): The offset for translate.
+ direction (str): The translate direction, either "horizontal"
+ or "vertical".
+ border_value (int | float): Border value. Default 0 for masks.
+ interpolation (str): Same as :func:`mmcv.imtranslate`.
+
+ Returns:
+ BitmapMasks: Translated BitmapMasks.
+
+ Example:
+ >>> from mmdet.data_elements.mask.structures import BitmapMasks
+ >>> self = BitmapMasks.random(dtype=np.uint8)
+ >>> out_shape = (32, 32)
+ >>> offset = 4
+ >>> direction = 'horizontal'
+ >>> border_value = 0
+ >>> interpolation = 'bilinear'
+ >>> # Note, There seem to be issues when:
+ >>> # * the mask dtype is not supported by cv2.AffineWarp
+ >>> new = self.translate(out_shape, offset, direction,
+ >>> border_value, interpolation)
+ >>> assert len(new) == len(self)
+ >>> assert new.height, new.width == out_shape
+ """
+ if len(self.masks) == 0:
+ translated_masks = np.empty((0, *out_shape), dtype=np.uint8)
+ else:
+ masks = self.masks
+ if masks.shape[-2:] != out_shape:
+ empty_masks = np.zeros((masks.shape[0], *out_shape),
+ dtype=masks.dtype)
+ min_h = min(out_shape[0], masks.shape[1])
+ min_w = min(out_shape[1], masks.shape[2])
+ empty_masks[:, :min_h, :min_w] = masks[:, :min_h, :min_w]
+ masks = empty_masks
+ translated_masks = mmcv.imtranslate(
+ masks.transpose((1, 2, 0)),
+ offset,
+ direction,
+ border_value=border_value,
+ interpolation=interpolation)
+ if translated_masks.ndim == 2:
+ translated_masks = translated_masks[:, :, None]
+ translated_masks = translated_masks.transpose(
+ (2, 0, 1)).astype(self.masks.dtype)
+ return BitmapMasks(translated_masks, *out_shape)
+
+ def shear(self,
+ out_shape,
+ magnitude,
+ direction='horizontal',
+ border_value=0,
+ interpolation='bilinear'):
+ """Shear the BitmapMasks.
+
+ Args:
+ out_shape (tuple[int]): Shape for output mask, format (h, w).
+ magnitude (int | float): The magnitude used for shear.
+ direction (str): The shear direction, either "horizontal"
+ or "vertical".
+ border_value (int | tuple[int]): Value used in case of a
+ constant border.
+ interpolation (str): Same as in :func:`mmcv.imshear`.
+
+ Returns:
+ BitmapMasks: The sheared masks.
+ """
+ if len(self.masks) == 0:
+ sheared_masks = np.empty((0, *out_shape), dtype=np.uint8)
+ else:
+ sheared_masks = mmcv.imshear(
+ self.masks.transpose((1, 2, 0)),
+ magnitude,
+ direction,
+ border_value=border_value,
+ interpolation=interpolation)
+ if sheared_masks.ndim == 2:
+ sheared_masks = sheared_masks[:, :, None]
+ sheared_masks = sheared_masks.transpose(
+ (2, 0, 1)).astype(self.masks.dtype)
+ return BitmapMasks(sheared_masks, *out_shape)
+
+ def rotate(self,
+ out_shape,
+ angle,
+ center=None,
+ scale=1.0,
+ border_value=0,
+ interpolation='bilinear'):
+ """Rotate the BitmapMasks.
+
+ Args:
+ out_shape (tuple[int]): Shape for output mask, format (h, w).
+ angle (int | float): Rotation angle in degrees. Positive values
+ mean counter-clockwise rotation.
+ center (tuple[float], optional): Center point (w, h) of the
+ rotation in source image. If not specified, the center of
+ the image will be used.
+ scale (int | float): Isotropic scale factor.
+ border_value (int | float): Border value. Default 0 for masks.
+ interpolation (str): Same as in :func:`mmcv.imrotate`.
+
+ Returns:
+ BitmapMasks: Rotated BitmapMasks.
+ """
+ if len(self.masks) == 0:
+ rotated_masks = np.empty((0, *out_shape), dtype=self.masks.dtype)
+ else:
+ rotated_masks = mmcv.imrotate(
+ self.masks.transpose((1, 2, 0)),
+ angle,
+ center=center,
+ scale=scale,
+ border_value=border_value,
+ interpolation=interpolation)
+ if rotated_masks.ndim == 2:
+ # case when only one mask, (h, w)
+ rotated_masks = rotated_masks[:, :, None] # (h, w, 1)
+ rotated_masks = rotated_masks.transpose(
+ (2, 0, 1)).astype(self.masks.dtype)
+ return BitmapMasks(rotated_masks, *out_shape)
+
+ @property
+ def areas(self):
+ """See :py:attr:`BaseInstanceMasks.areas`."""
+ return self.masks.sum((1, 2))
+
+ def to_ndarray(self):
+ """See :func:`BaseInstanceMasks.to_ndarray`."""
+ return self.masks
+
+ def to_tensor(self, dtype, device):
+ """See :func:`BaseInstanceMasks.to_tensor`."""
+ return torch.tensor(self.masks, dtype=dtype, device=device)
+
+ @classmethod
+ def random(cls,
+ num_masks=3,
+ height=32,
+ width=32,
+ dtype=np.uint8,
+ rng=None):
+ """Generate random bitmap masks for demo / testing purposes.
+
+ Example:
+ >>> from mmdet.data_elements.mask.structures import BitmapMasks
+ >>> self = BitmapMasks.random()
+ >>> print('self = {}'.format(self))
+ self = BitmapMasks(num_masks=3, height=32, width=32)
+ """
+ from mmdet.utils.util_random import ensure_rng
+ rng = ensure_rng(rng)
+ masks = (rng.rand(num_masks, height, width) > 0.1).astype(dtype)
+ self = cls(masks, height=height, width=width)
+ return self
+
+ @classmethod
+ def cat(cls: Type[T], masks: Sequence[T]) -> T:
+ """Concatenate a sequence of masks into one single mask instance.
+
+ Args:
+ masks (Sequence[BitmapMasks]): A sequence of mask instances.
+
+ Returns:
+ BitmapMasks: Concatenated mask instance.
+ """
+ assert isinstance(masks, Sequence)
+ if len(masks) == 0:
+ raise ValueError('masks should not be an empty list.')
+ assert all(isinstance(m, cls) for m in masks)
+
+ mask_array = np.concatenate([m.masks for m in masks], axis=0)
+ return cls(mask_array, *mask_array.shape[1:])
+
+
+class PolygonMasks(BaseInstanceMasks):
+ """This class represents masks in the form of polygons.
+
+ Polygons is a list of three levels. The first level of the list
+ corresponds to objects, the second level to the polys that compose the
+ object, the third level to the poly coordinates
+
+ Args:
+ masks (list[list[ndarray]]): The first level of the list
+ corresponds to objects, the second level to the polys that
+ compose the object, the third level to the poly coordinates
+ height (int): height of masks
+ width (int): width of masks
+
+ Example:
+ >>> from mmdet.data_elements.mask.structures import * # NOQA
+ >>> masks = [
+ >>> [ np.array([0, 0, 10, 0, 10, 10., 0, 10, 0, 0]) ]
+ >>> ]
+ >>> height, width = 16, 16
+ >>> self = PolygonMasks(masks, height, width)
+
+ >>> # demo translate
+ >>> new = self.translate((16, 16), 4., direction='horizontal')
+ >>> assert np.all(new.masks[0][0][1::2] == masks[0][0][1::2])
+ >>> assert np.all(new.masks[0][0][0::2] == masks[0][0][0::2] + 4)
+
+ >>> # demo crop_and_resize
+ >>> num_boxes = 3
+ >>> bboxes = np.array([[0, 0, 30, 10.0]] * num_boxes)
+ >>> out_shape = (16, 16)
+ >>> inds = torch.randint(0, len(self), size=(num_boxes,))
+ >>> device = 'cpu'
+ >>> interpolation = 'bilinear'
+ >>> new = self.crop_and_resize(
+ ... bboxes, out_shape, inds, device, interpolation)
+ >>> assert len(new) == num_boxes
+ >>> assert new.height, new.width == out_shape
+ """
+
+ def __init__(self, masks, height, width):
+ assert isinstance(masks, list)
+ if len(masks) > 0:
+ assert isinstance(masks[0], list)
+ assert isinstance(masks[0][0], np.ndarray)
+
+ self.height = height
+ self.width = width
+ self.masks = masks
+
+ def __getitem__(self, index):
+ """Index the polygon masks.
+
+ Args:
+ index (ndarray | List): The indices.
+
+ Returns:
+ :obj:`PolygonMasks`: The indexed polygon masks.
+ """
+ if isinstance(index, np.ndarray):
+ if index.dtype == bool:
+ index = np.where(index)[0].tolist()
+ else:
+ index = index.tolist()
+ if isinstance(index, list):
+ masks = [self.masks[i] for i in index]
+ else:
+ try:
+ masks = self.masks[index]
+ except Exception:
+ raise ValueError(
+ f'Unsupported input of type {type(index)} for indexing!')
+ if len(masks) and isinstance(masks[0], np.ndarray):
+ masks = [masks] # ensure a list of three levels
+ return PolygonMasks(masks, self.height, self.width)
+
+ def __iter__(self):
+ return iter(self.masks)
+
+ def __repr__(self):
+ s = self.__class__.__name__ + '('
+ s += f'num_masks={len(self.masks)}, '
+ s += f'height={self.height}, '
+ s += f'width={self.width})'
+ return s
+
+ def __len__(self):
+ """Number of masks."""
+ return len(self.masks)
+
+ def rescale(self, scale, interpolation=None):
+ """see :func:`BaseInstanceMasks.rescale`"""
+ new_w, new_h = mmcv.rescale_size((self.width, self.height), scale)
+ if len(self.masks) == 0:
+ rescaled_masks = PolygonMasks([], new_h, new_w)
+ else:
+ rescaled_masks = self.resize((new_h, new_w))
+ return rescaled_masks
+
+ def resize(self, out_shape, interpolation=None):
+ """see :func:`BaseInstanceMasks.resize`"""
+ if len(self.masks) == 0:
+ resized_masks = PolygonMasks([], *out_shape)
+ else:
+ h_scale = out_shape[0] / self.height
+ w_scale = out_shape[1] / self.width
+ resized_masks = []
+ for poly_per_obj in self.masks:
+ resized_poly = []
+ for p in poly_per_obj:
+ p = p.copy()
+ p[0::2] = p[0::2] * w_scale
+ p[1::2] = p[1::2] * h_scale
+ resized_poly.append(p)
+ resized_masks.append(resized_poly)
+ resized_masks = PolygonMasks(resized_masks, *out_shape)
+ return resized_masks
+
+ def flip(self, flip_direction='horizontal'):
+ """see :func:`BaseInstanceMasks.flip`"""
+ assert flip_direction in ('horizontal', 'vertical', 'diagonal')
+ if len(self.masks) == 0:
+ flipped_masks = PolygonMasks([], self.height, self.width)
+ else:
+ flipped_masks = []
+ for poly_per_obj in self.masks:
+ flipped_poly_per_obj = []
+ for p in poly_per_obj:
+ p = p.copy()
+ if flip_direction == 'horizontal':
+ p[0::2] = self.width - p[0::2]
+ elif flip_direction == 'vertical':
+ p[1::2] = self.height - p[1::2]
+ else:
+ p[0::2] = self.width - p[0::2]
+ p[1::2] = self.height - p[1::2]
+ flipped_poly_per_obj.append(p)
+ flipped_masks.append(flipped_poly_per_obj)
+ flipped_masks = PolygonMasks(flipped_masks, self.height,
+ self.width)
+ return flipped_masks
+
+ def crop(self, bbox):
+ """see :func:`BaseInstanceMasks.crop`"""
+ assert isinstance(bbox, np.ndarray)
+ assert bbox.ndim == 1
+
+ # clip the boundary
+ bbox = bbox.copy()
+ bbox[0::2] = np.clip(bbox[0::2], 0, self.width)
+ bbox[1::2] = np.clip(bbox[1::2], 0, self.height)
+ x1, y1, x2, y2 = bbox
+ w = np.maximum(x2 - x1, 1)
+ h = np.maximum(y2 - y1, 1)
+
+ if len(self.masks) == 0:
+ cropped_masks = PolygonMasks([], h, w)
+ else:
+ # reference: https://github.com/facebookresearch/fvcore/blob/main/fvcore/transforms/transform.py # noqa
+ crop_box = geometry.box(x1, y1, x2, y2).buffer(0.0)
+ cropped_masks = []
+ # suppress shapely warnings util it incorporates GEOS>=3.11.2
+ # reference: https://github.com/shapely/shapely/issues/1345
+ initial_settings = np.seterr()
+ np.seterr(invalid='ignore')
+ for poly_per_obj in self.masks:
+ cropped_poly_per_obj = []
+ for p in poly_per_obj:
+ p = p.copy()
+ p = geometry.Polygon(p.reshape(-1, 2)).buffer(0.0)
+ # polygon must be valid to perform intersection.
+ if not p.is_valid:
+ continue
+ cropped = p.intersection(crop_box)
+ if cropped.is_empty:
+ continue
+ if isinstance(cropped,
+ geometry.collection.BaseMultipartGeometry):
+ cropped = cropped.geoms
+ else:
+ cropped = [cropped]
+ # one polygon may be cropped to multiple ones
+ for poly in cropped:
+ # ignore lines or points
+ if not isinstance(
+ poly, geometry.Polygon) or not poly.is_valid:
+ continue
+ coords = np.asarray(poly.exterior.coords)
+ # remove an extra identical vertex at the end
+ coords = coords[:-1]
+ coords[:, 0] -= x1
+ coords[:, 1] -= y1
+ cropped_poly_per_obj.append(coords.reshape(-1))
+ # a dummy polygon to avoid misalignment between masks and boxes
+ if len(cropped_poly_per_obj) == 0:
+ cropped_poly_per_obj = [np.array([0, 0, 0, 0, 0, 0])]
+ cropped_masks.append(cropped_poly_per_obj)
+ np.seterr(**initial_settings)
+ cropped_masks = PolygonMasks(cropped_masks, h, w)
+ return cropped_masks
+
+ def pad(self, out_shape, pad_val=0):
+ """padding has no effect on polygons`"""
+ return PolygonMasks(self.masks, *out_shape)
+
+ def expand(self, *args, **kwargs):
+ """TODO: Add expand for polygon"""
+ raise NotImplementedError
+
+ def crop_and_resize(self,
+ bboxes,
+ out_shape,
+ inds,
+ device='cpu',
+ interpolation='bilinear',
+ binarize=True):
+ """see :func:`BaseInstanceMasks.crop_and_resize`"""
+ out_h, out_w = out_shape
+ if len(self.masks) == 0:
+ return PolygonMasks([], out_h, out_w)
+
+ if not binarize:
+ raise ValueError('Polygons are always binary, '
+ 'setting binarize=False is unsupported')
+
+ resized_masks = []
+ for i in range(len(bboxes)):
+ mask = self.masks[inds[i]]
+ bbox = bboxes[i, :]
+ x1, y1, x2, y2 = bbox
+ w = np.maximum(x2 - x1, 1)
+ h = np.maximum(y2 - y1, 1)
+ h_scale = out_h / max(h, 0.1) # avoid too large scale
+ w_scale = out_w / max(w, 0.1)
+
+ resized_mask = []
+ for p in mask:
+ p = p.copy()
+ # crop
+ # pycocotools will clip the boundary
+ p[0::2] = p[0::2] - bbox[0]
+ p[1::2] = p[1::2] - bbox[1]
+
+ # resize
+ p[0::2] = p[0::2] * w_scale
+ p[1::2] = p[1::2] * h_scale
+ resized_mask.append(p)
+ resized_masks.append(resized_mask)
+ return PolygonMasks(resized_masks, *out_shape)
+
+ def translate(self,
+ out_shape,
+ offset,
+ direction='horizontal',
+ border_value=None,
+ interpolation=None):
+ """Translate the PolygonMasks.
+
+ Example:
+ >>> self = PolygonMasks.random(dtype=np.int64)
+ >>> out_shape = (self.height, self.width)
+ >>> new = self.translate(out_shape, 4., direction='horizontal')
+ >>> assert np.all(new.masks[0][0][1::2] == self.masks[0][0][1::2])
+ >>> assert np.all(new.masks[0][0][0::2] == self.masks[0][0][0::2] + 4) # noqa: E501
+ """
+ assert border_value is None or border_value == 0, \
+ 'Here border_value is not '\
+ f'used, and defaultly should be None or 0. got {border_value}.'
+ if len(self.masks) == 0:
+ translated_masks = PolygonMasks([], *out_shape)
+ else:
+ translated_masks = []
+ for poly_per_obj in self.masks:
+ translated_poly_per_obj = []
+ for p in poly_per_obj:
+ p = p.copy()
+ if direction == 'horizontal':
+ p[0::2] = np.clip(p[0::2] + offset, 0, out_shape[1])
+ elif direction == 'vertical':
+ p[1::2] = np.clip(p[1::2] + offset, 0, out_shape[0])
+ translated_poly_per_obj.append(p)
+ translated_masks.append(translated_poly_per_obj)
+ translated_masks = PolygonMasks(translated_masks, *out_shape)
+ return translated_masks
+
+ def shear(self,
+ out_shape,
+ magnitude,
+ direction='horizontal',
+ border_value=0,
+ interpolation='bilinear'):
+ """See :func:`BaseInstanceMasks.shear`."""
+ if len(self.masks) == 0:
+ sheared_masks = PolygonMasks([], *out_shape)
+ else:
+ sheared_masks = []
+ if direction == 'horizontal':
+ shear_matrix = np.stack([[1, magnitude],
+ [0, 1]]).astype(np.float32)
+ elif direction == 'vertical':
+ shear_matrix = np.stack([[1, 0], [magnitude,
+ 1]]).astype(np.float32)
+ for poly_per_obj in self.masks:
+ sheared_poly = []
+ for p in poly_per_obj:
+ p = np.stack([p[0::2], p[1::2]], axis=0) # [2, n]
+ new_coords = np.matmul(shear_matrix, p) # [2, n]
+ new_coords[0, :] = np.clip(new_coords[0, :], 0,
+ out_shape[1])
+ new_coords[1, :] = np.clip(new_coords[1, :], 0,
+ out_shape[0])
+ sheared_poly.append(
+ new_coords.transpose((1, 0)).reshape(-1))
+ sheared_masks.append(sheared_poly)
+ sheared_masks = PolygonMasks(sheared_masks, *out_shape)
+ return sheared_masks
+
+ def rotate(self,
+ out_shape,
+ angle,
+ center=None,
+ scale=1.0,
+ border_value=0,
+ interpolation='bilinear'):
+ """See :func:`BaseInstanceMasks.rotate`."""
+ if len(self.masks) == 0:
+ rotated_masks = PolygonMasks([], *out_shape)
+ else:
+ rotated_masks = []
+ rotate_matrix = cv2.getRotationMatrix2D(center, -angle, scale)
+ for poly_per_obj in self.masks:
+ rotated_poly = []
+ for p in poly_per_obj:
+ p = p.copy()
+ coords = np.stack([p[0::2], p[1::2]], axis=1) # [n, 2]
+ # pad 1 to convert from format [x, y] to homogeneous
+ # coordinates format [x, y, 1]
+ coords = np.concatenate(
+ (coords, np.ones((coords.shape[0], 1), coords.dtype)),
+ axis=1) # [n, 3]
+ rotated_coords = np.matmul(
+ rotate_matrix[None, :, :],
+ coords[:, :, None])[..., 0] # [n, 2, 1] -> [n, 2]
+ rotated_coords[:, 0] = np.clip(rotated_coords[:, 0], 0,
+ out_shape[1])
+ rotated_coords[:, 1] = np.clip(rotated_coords[:, 1], 0,
+ out_shape[0])
+ rotated_poly.append(rotated_coords.reshape(-1))
+ rotated_masks.append(rotated_poly)
+ rotated_masks = PolygonMasks(rotated_masks, *out_shape)
+ return rotated_masks
+
+ def to_bitmap(self):
+ """convert polygon masks to bitmap masks."""
+ bitmap_masks = self.to_ndarray()
+ return BitmapMasks(bitmap_masks, self.height, self.width)
+
+ @property
+ def areas(self):
+ """Compute areas of masks.
+
+ This func is modified from `detectron2
+ `_.
+ The function only works with Polygons using the shoelace formula.
+
+ Return:
+ ndarray: areas of each instance
+ """ # noqa: W501
+ area = []
+ for polygons_per_obj in self.masks:
+ area_per_obj = 0
+ for p in polygons_per_obj:
+ area_per_obj += self._polygon_area(p[0::2], p[1::2])
+ area.append(area_per_obj)
+ return np.asarray(area)
+
+ def _polygon_area(self, x, y):
+ """Compute the area of a component of a polygon.
+
+ Using the shoelace formula:
+ https://stackoverflow.com/questions/24467972/calculate-area-of-polygon-given-x-y-coordinates
+
+ Args:
+ x (ndarray): x coordinates of the component
+ y (ndarray): y coordinates of the component
+
+ Return:
+ float: the are of the component
+ """ # noqa: 501
+ return 0.5 * np.abs(
+ np.dot(x, np.roll(y, 1)) - np.dot(y, np.roll(x, 1)))
+
+ def to_ndarray(self):
+ """Convert masks to the format of ndarray."""
+ if len(self.masks) == 0:
+ return np.empty((0, self.height, self.width), dtype=np.uint8)
+ bitmap_masks = []
+ for poly_per_obj in self.masks:
+ bitmap_masks.append(
+ polygon_to_bitmap(poly_per_obj, self.height, self.width))
+ return np.stack(bitmap_masks)
+
+ def to_tensor(self, dtype, device):
+ """See :func:`BaseInstanceMasks.to_tensor`."""
+ if len(self.masks) == 0:
+ return torch.empty((0, self.height, self.width),
+ dtype=dtype,
+ device=device)
+ ndarray_masks = self.to_ndarray()
+ return torch.tensor(ndarray_masks, dtype=dtype, device=device)
+
+ @classmethod
+ def random(cls,
+ num_masks=3,
+ height=32,
+ width=32,
+ n_verts=5,
+ dtype=np.float32,
+ rng=None):
+ """Generate random polygon masks for demo / testing purposes.
+
+ Adapted from [1]_
+
+ References:
+ .. [1] https://gitlab.kitware.com/computer-vision/kwimage/-/blob/928cae35ca8/kwimage/structs/polygon.py#L379 # noqa: E501
+
+ Example:
+ >>> from mmdet.data_elements.mask.structures import PolygonMasks
+ >>> self = PolygonMasks.random()
+ >>> print('self = {}'.format(self))
+ """
+ from mmdet.utils.util_random import ensure_rng
+ rng = ensure_rng(rng)
+
+ def _gen_polygon(n, irregularity, spikeyness):
+ """Creates the polygon by sampling points on a circle around the
+ centre. Random noise is added by varying the angular spacing
+ between sequential points, and by varying the radial distance of
+ each point from the centre.
+
+ Based on original code by Mike Ounsworth
+
+ Args:
+ n (int): number of vertices
+ irregularity (float): [0,1] indicating how much variance there
+ is in the angular spacing of vertices. [0,1] will map to
+ [0, 2pi/numberOfVerts]
+ spikeyness (float): [0,1] indicating how much variance there is
+ in each vertex from the circle of radius aveRadius. [0,1]
+ will map to [0, aveRadius]
+
+ Returns:
+ a list of vertices, in CCW order.
+ """
+ from scipy.stats import truncnorm
+
+ # Generate around the unit circle
+ cx, cy = (0.0, 0.0)
+ radius = 1
+
+ tau = np.pi * 2
+
+ irregularity = np.clip(irregularity, 0, 1) * 2 * np.pi / n
+ spikeyness = np.clip(spikeyness, 1e-9, 1)
+
+ # generate n angle steps
+ lower = (tau / n) - irregularity
+ upper = (tau / n) + irregularity
+ angle_steps = rng.uniform(lower, upper, n)
+
+ # normalize the steps so that point 0 and point n+1 are the same
+ k = angle_steps.sum() / (2 * np.pi)
+ angles = (angle_steps / k).cumsum() + rng.uniform(0, tau)
+
+ # Convert high and low values to be wrt the standard normal range
+ # https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.truncnorm.html
+ low = 0
+ high = 2 * radius
+ mean = radius
+ std = spikeyness
+ a = (low - mean) / std
+ b = (high - mean) / std
+ tnorm = truncnorm(a=a, b=b, loc=mean, scale=std)
+
+ # now generate the points
+ radii = tnorm.rvs(n, random_state=rng)
+ x_pts = cx + radii * np.cos(angles)
+ y_pts = cy + radii * np.sin(angles)
+
+ points = np.hstack([x_pts[:, None], y_pts[:, None]])
+
+ # Scale to 0-1 space
+ points = points - points.min(axis=0)
+ points = points / points.max(axis=0)
+
+ # Randomly place within 0-1 space
+ points = points * (rng.rand() * .8 + .2)
+ min_pt = points.min(axis=0)
+ max_pt = points.max(axis=0)
+
+ high = (1 - max_pt)
+ low = (0 - min_pt)
+ offset = (rng.rand(2) * (high - low)) + low
+ points = points + offset
+ return points
+
+ def _order_vertices(verts):
+ """
+ References:
+ https://stackoverflow.com/questions/1709283/how-can-i-sort-a-coordinate-list-for-a-rectangle-counterclockwise
+ """
+ mlat = verts.T[0].sum() / len(verts)
+ mlng = verts.T[1].sum() / len(verts)
+
+ tau = np.pi * 2
+ angle = (np.arctan2(mlat - verts.T[0], verts.T[1] - mlng) +
+ tau) % tau
+ sortx = angle.argsort()
+ verts = verts.take(sortx, axis=0)
+ return verts
+
+ # Generate a random exterior for each requested mask
+ masks = []
+ for _ in range(num_masks):
+ exterior = _order_vertices(_gen_polygon(n_verts, 0.9, 0.9))
+ exterior = (exterior * [(width, height)]).astype(dtype)
+ masks.append([exterior.ravel()])
+
+ self = cls(masks, height, width)
+ return self
+
+ @classmethod
+ def cat(cls: Type[T], masks: Sequence[T]) -> T:
+ """Concatenate a sequence of masks into one single mask instance.
+
+ Args:
+ masks (Sequence[PolygonMasks]): A sequence of mask instances.
+
+ Returns:
+ PolygonMasks: Concatenated mask instance.
+ """
+ assert isinstance(masks, Sequence)
+ if len(masks) == 0:
+ raise ValueError('masks should not be an empty list.')
+ assert all(isinstance(m, cls) for m in masks)
+
+ mask_list = list(itertools.chain(*[m.masks for m in masks]))
+ return cls(mask_list, masks[0].height, masks[0].width)
+
+
+def polygon_to_bitmap(polygons, height, width):
+ """Convert masks from the form of polygons to bitmaps.
+
+ Args:
+ polygons (list[ndarray]): masks in polygon representation
+ height (int): mask height
+ width (int): mask width
+
+ Return:
+ ndarray: the converted masks in bitmap representation
+ """
+ rles = maskUtils.frPyObjects(polygons, height, width)
+ rle = maskUtils.merge(rles)
+ bitmap_mask = maskUtils.decode(rle).astype(bool)
+ return bitmap_mask
+
+
+def bitmap_to_polygon(bitmap):
+ """Convert masks from the form of bitmaps to polygons.
+
+ Args:
+ bitmap (ndarray): masks in bitmap representation.
+
+ Return:
+ list[ndarray]: the converted mask in polygon representation.
+ bool: whether the mask has holes.
+ """
+ bitmap = np.ascontiguousarray(bitmap).astype(np.uint8)
+ # cv2.RETR_CCOMP: retrieves all of the contours and organizes them
+ # into a two-level hierarchy. At the top level, there are external
+ # boundaries of the components. At the second level, there are
+ # boundaries of the holes. If there is another contour inside a hole
+ # of a connected component, it is still put at the top level.
+ # cv2.CHAIN_APPROX_NONE: stores absolutely all the contour points.
+ outs = cv2.findContours(bitmap, cv2.RETR_CCOMP, cv2.CHAIN_APPROX_NONE)
+ contours = outs[-2]
+ hierarchy = outs[-1]
+ if hierarchy is None:
+ return [], False
+ # hierarchy[i]: 4 elements, for the indexes of next, previous,
+ # parent, or nested contours. If there is no corresponding contour,
+ # it will be -1.
+ with_hole = (hierarchy.reshape(-1, 4)[:, 3] >= 0).any()
+ contours = [c.reshape(-1, 2) for c in contours]
+ return contours, with_hole
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/mask/utils.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/mask/utils.py
new file mode 100644
index 0000000000000000000000000000000000000000..6bd445e4fce1a312949f222d54d230a1a622d726
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/mask/utils.py
@@ -0,0 +1,77 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import numpy as np
+import pycocotools.mask as mask_util
+import torch
+from mmengine.utils import slice_list
+
+
+def split_combined_polys(polys, poly_lens, polys_per_mask):
+ """Split the combined 1-D polys into masks.
+
+ A mask is represented as a list of polys, and a poly is represented as
+ a 1-D array. In dataset, all masks are concatenated into a single 1-D
+ tensor. Here we need to split the tensor into original representations.
+
+ Args:
+ polys (list): a list (length = image num) of 1-D tensors
+ poly_lens (list): a list (length = image num) of poly length
+ polys_per_mask (list): a list (length = image num) of poly number
+ of each mask
+
+ Returns:
+ list: a list (length = image num) of list (length = mask num) of \
+ list (length = poly num) of numpy array.
+ """
+ mask_polys_list = []
+ for img_id in range(len(polys)):
+ polys_single = polys[img_id]
+ polys_lens_single = poly_lens[img_id].tolist()
+ polys_per_mask_single = polys_per_mask[img_id].tolist()
+
+ split_polys = slice_list(polys_single, polys_lens_single)
+ mask_polys = slice_list(split_polys, polys_per_mask_single)
+ mask_polys_list.append(mask_polys)
+ return mask_polys_list
+
+
+# TODO: move this function to more proper place
+def encode_mask_results(mask_results):
+ """Encode bitmap mask to RLE code.
+
+ Args:
+ mask_results (list): bitmap mask results.
+
+ Returns:
+ list | tuple: RLE encoded mask.
+ """
+ encoded_mask_results = []
+ for mask in mask_results:
+ encoded_mask_results.append(
+ mask_util.encode(
+ np.array(mask[:, :, np.newaxis], order='F',
+ dtype='uint8'))[0]) # encoded with RLE
+ return encoded_mask_results
+
+
+def mask2bbox(masks):
+ """Obtain tight bounding boxes of binary masks.
+
+ Args:
+ masks (Tensor): Binary mask of shape (n, h, w).
+
+ Returns:
+ Tensor: Bboxe with shape (n, 4) of \
+ positive region in binary mask.
+ """
+ N = masks.shape[0]
+ bboxes = masks.new_zeros((N, 4), dtype=torch.float32)
+ x_any = torch.any(masks, dim=1)
+ y_any = torch.any(masks, dim=2)
+ for i in range(N):
+ x = torch.where(x_any[i, :])[0]
+ y = torch.where(y_any[i, :])[0]
+ if len(x) > 0 and len(y) > 0:
+ bboxes[i, :] = bboxes.new_tensor(
+ [x[0], y[0], x[-1] + 1, y[-1] + 1])
+
+ return bboxes
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/reid_data_sample.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/reid_data_sample.py
new file mode 100644
index 0000000000000000000000000000000000000000..69958eece3671c9040c1f5561e724ca2d5f8e155
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/reid_data_sample.py
@@ -0,0 +1,123 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from numbers import Number
+from typing import Sequence, Union
+
+import mmengine
+import numpy as np
+import torch
+from mmengine.structures import BaseDataElement, LabelData
+
+
+def format_label(value: Union[torch.Tensor, np.ndarray, Sequence, int],
+ num_classes: int = None) -> LabelData:
+ """Convert label of various python types to :obj:`mmengine.LabelData`.
+
+ Supported types are: :class:`numpy.ndarray`, :class:`torch.Tensor`,
+ :class:`Sequence`, :class:`int`.
+
+ Args:
+ value (torch.Tensor | numpy.ndarray | Sequence | int): Label value.
+ num_classes (int, optional): The number of classes. If not None, set
+ it to the metainfo. Defaults to None.
+
+ Returns:
+ :obj:`mmengine.LabelData`: The foramtted label data.
+ """
+
+ # Handle single number
+ if isinstance(value, (torch.Tensor, np.ndarray)) and value.ndim == 0:
+ value = int(value.item())
+
+ if isinstance(value, np.ndarray):
+ value = torch.from_numpy(value)
+ elif isinstance(value, Sequence) and not mmengine.utils.is_str(value):
+ value = torch.tensor(value)
+ elif isinstance(value, int):
+ value = torch.LongTensor([value])
+ elif not isinstance(value, torch.Tensor):
+ raise TypeError(f'Type {type(value)} is not an available label type.')
+
+ metainfo = {}
+ if num_classes is not None:
+ metainfo['num_classes'] = num_classes
+ if value.max() >= num_classes:
+ raise ValueError(f'The label data ({value}) should not '
+ f'exceed num_classes ({num_classes}).')
+ label = LabelData(label=value, metainfo=metainfo)
+ return label
+
+
+class ReIDDataSample(BaseDataElement):
+ """A data structure interface of ReID task.
+
+ It's used as interfaces between different components.
+
+ Meta field:
+ img_shape (Tuple): The shape of the corresponding input image.
+ Used for visualization.
+ ori_shape (Tuple): The original shape of the corresponding image.
+ Used for visualization.
+ num_classes (int): The number of all categories.
+ Used for label format conversion.
+
+ Data field:
+ gt_label (LabelData): The ground truth label.
+ pred_label (LabelData): The predicted label.
+ scores (torch.Tensor): The outputs of model.
+ """
+
+ @property
+ def gt_label(self):
+ return self._gt_label
+
+ @gt_label.setter
+ def gt_label(self, value: LabelData):
+ self.set_field(value, '_gt_label', dtype=LabelData)
+
+ @gt_label.deleter
+ def gt_label(self):
+ del self._gt_label
+
+ def set_gt_label(
+ self, value: Union[np.ndarray, torch.Tensor, Sequence[Number], Number]
+ ) -> 'ReIDDataSample':
+ """Set label of ``gt_label``."""
+ label = format_label(value, self.get('num_classes'))
+ if 'gt_label' in self: # setting for the second time
+ self.gt_label.label = label.label
+ else: # setting for the first time
+ self.gt_label = label
+ return self
+
+ def set_gt_score(self, value: torch.Tensor) -> 'ReIDDataSample':
+ """Set score of ``gt_label``."""
+ assert isinstance(value, torch.Tensor), \
+ f'The value should be a torch.Tensor but got {type(value)}.'
+ assert value.ndim == 1, \
+ f'The dims of value should be 1, but got {value.ndim}.'
+
+ if 'num_classes' in self:
+ assert value.size(0) == self.num_classes, \
+ f"The length of value ({value.size(0)}) doesn't "\
+ f'match the num_classes ({self.num_classes}).'
+ metainfo = {'num_classes': self.num_classes}
+ else:
+ metainfo = {'num_classes': value.size(0)}
+
+ if 'gt_label' in self: # setting for the second time
+ self.gt_label.score = value
+ else: # setting for the first time
+ self.gt_label = LabelData(score=value, metainfo=metainfo)
+ return self
+
+ @property
+ def pred_feature(self):
+ return self._pred_feature
+
+ @pred_feature.setter
+ def pred_feature(self, value: torch.Tensor):
+ self.set_field(value, '_pred_feature', dtype=torch.Tensor)
+
+ @pred_feature.deleter
+ def pred_feature(self):
+ del self._pred_feature
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/track_data_sample.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/track_data_sample.py
new file mode 100644
index 0000000000000000000000000000000000000000..d005a5a42f57682d0b76d60d3dae463c4b4dc727
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/structures/track_data_sample.py
@@ -0,0 +1,273 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Sequence
+
+import numpy as np
+import torch
+from mmengine.structures import BaseDataElement
+
+from .det_data_sample import DetDataSample
+
+
+class TrackDataSample(BaseDataElement):
+ """A data structure interface of tracking task in MMDetection. It is used
+ as interfaces between different components.
+
+ This data structure can be viewd as a wrapper of multiple DetDataSample to
+ some extent. Specifically, it only contains a property:
+ ``video_data_samples`` which is a list of DetDataSample, each of which
+ corresponds to a single frame. If you want to get the property of a single
+ frame, you must first get the corresponding ``DetDataSample`` by indexing
+ and then get the property of the frame, such as ``gt_instances``,
+ ``pred_instances`` and so on. As for metainfo, it differs from
+ ``DetDataSample`` in that each value corresponds to the metainfo key is a
+ list where each element corresponds to information of a single frame.
+
+ Examples:
+ >>> import torch
+ >>> from mmengine.structures import InstanceData
+ >>> from mmdet.structures import DetDataSample, TrackDataSample
+ >>> track_data_sample = TrackDataSample()
+ >>> # set the 1st frame
+ >>> frame1_data_sample = DetDataSample(metainfo=dict(
+ ... img_shape=(100, 100), frame_id=0))
+ >>> frame1_gt_instances = InstanceData()
+ >>> frame1_gt_instances.bbox = torch.zeros([2, 4])
+ >>> frame1_data_sample.gt_instances = frame1_gt_instances
+ >>> # set the 2nd frame
+ >>> frame2_data_sample = DetDataSample(metainfo=dict(
+ ... img_shape=(100, 100), frame_id=1))
+ >>> frame2_gt_instances = InstanceData()
+ >>> frame2_gt_instances.bbox = torch.ones([3, 4])
+ >>> frame2_data_sample.gt_instances = frame2_gt_instances
+ >>> track_data_sample.video_data_samples = [frame1_data_sample,
+ ... frame2_data_sample]
+ >>> # set metainfo for track_data_sample
+ >>> track_data_sample.set_metainfo(dict(key_frames_inds=[0]))
+ >>> track_data_sample.set_metainfo(dict(ref_frames_inds=[1]))
+ >>> print(track_data_sample)
+
+ ) at 0x7f64bd223340>,
+ ) at 0x7f64bd1346d0>]
+ ) at 0x7f64bd2237f0>
+ >>> print(len(track_data_sample))
+ 2
+ >>> key_data_sample = track_data_sample.get_key_frames()
+ >>> print(key_data_sample[0].frame_id)
+ 0
+ >>> ref_data_sample = track_data_sample.get_ref_frames()
+ >>> print(ref_data_sample[0].frame_id)
+ 1
+ >>> frame1_data_sample = track_data_sample[0]
+ >>> print(frame1_data_sample.gt_instances.bbox)
+ tensor([[0., 0., 0., 0.],
+ [0., 0., 0., 0.]])
+ >>> # Tensor-like methods
+ >>> cuda_track_data_sample = track_data_sample.to('cuda')
+ >>> cuda_track_data_sample = track_data_sample.cuda()
+ >>> cpu_track_data_sample = track_data_sample.cpu()
+ >>> cpu_track_data_sample = track_data_sample.to('cpu')
+ >>> fp16_instances = cuda_track_data_sample.to(
+ ... device=None, dtype=torch.float16, non_blocking=False,
+ ... copy=False, memory_format=torch.preserve_format)
+ """
+
+ @property
+ def video_data_samples(self) -> List[DetDataSample]:
+ return self._video_data_samples
+
+ @video_data_samples.setter
+ def video_data_samples(self, value: List[DetDataSample]):
+ if isinstance(value, DetDataSample):
+ value = [value]
+ assert isinstance(value, list), 'video_data_samples must be a list'
+ assert isinstance(
+ value[0], DetDataSample
+ ), 'video_data_samples must be a list of DetDataSample, but got '
+ f'{value[0]}'
+ self.set_field(value, '_video_data_samples', dtype=list)
+
+ @video_data_samples.deleter
+ def video_data_samples(self):
+ del self._video_data_samples
+
+ def __getitem__(self, index):
+ assert hasattr(self,
+ '_video_data_samples'), 'video_data_samples not set'
+ return self._video_data_samples[index]
+
+ def get_key_frames(self):
+ assert hasattr(self, 'key_frames_inds'), \
+ 'key_frames_inds not set'
+ assert isinstance(self.key_frames_inds, Sequence)
+ key_frames_info = []
+ for index in self.key_frames_inds:
+ key_frames_info.append(self[index])
+ return key_frames_info
+
+ def get_ref_frames(self):
+ assert hasattr(self, 'ref_frames_inds'), \
+ 'ref_frames_inds not set'
+ ref_frames_info = []
+ assert isinstance(self.ref_frames_inds, Sequence)
+ for index in self.ref_frames_inds:
+ ref_frames_info.append(self[index])
+ return ref_frames_info
+
+ def __len__(self):
+ return len(self._video_data_samples) if hasattr(
+ self, '_video_data_samples') else 0
+
+ # TODO: add UT for this Tensor-like method
+ # Tensor-like methods
+ def to(self, *args, **kwargs) -> 'BaseDataElement':
+ """Apply same name function to all tensors in data_fields."""
+ new_data = self.new()
+ for k, v_list in self.items():
+ data_list = []
+ for v in v_list:
+ if hasattr(v, 'to'):
+ v = v.to(*args, **kwargs)
+ data_list.append(v)
+ if len(data_list) > 0:
+ new_data.set_data({f'{k}': data_list})
+ return new_data
+
+ # Tensor-like methods
+ def cpu(self) -> 'BaseDataElement':
+ """Convert all tensors to CPU in data."""
+ new_data = self.new()
+ for k, v_list in self.items():
+ data_list = []
+ for v in v_list:
+ if isinstance(v, (torch.Tensor, BaseDataElement)):
+ v = v.cpu()
+ data_list.append(v)
+ if len(data_list) > 0:
+ new_data.set_data({f'{k}': data_list})
+ return new_data
+
+ # Tensor-like methods
+ def cuda(self) -> 'BaseDataElement':
+ """Convert all tensors to GPU in data."""
+ new_data = self.new()
+ for k, v_list in self.items():
+ data_list = []
+ for v in v_list:
+ if isinstance(v, (torch.Tensor, BaseDataElement)):
+ v = v.cuda()
+ data_list.append(v)
+ if len(data_list) > 0:
+ new_data.set_data({f'{k}': data_list})
+ return new_data
+
+ # Tensor-like methods
+ def npu(self) -> 'BaseDataElement':
+ """Convert all tensors to NPU in data."""
+ new_data = self.new()
+ for k, v_list in self.items():
+ data_list = []
+ for v in v_list:
+ if isinstance(v, (torch.Tensor, BaseDataElement)):
+ v = v.npu()
+ data_list.append(v)
+ if len(data_list) > 0:
+ new_data.set_data({f'{k}': data_list})
+ return new_data
+
+ # Tensor-like methods
+ def detach(self) -> 'BaseDataElement':
+ """Detach all tensors in data."""
+ new_data = self.new()
+ for k, v_list in self.items():
+ data_list = []
+ for v in v_list:
+ if isinstance(v, (torch.Tensor, BaseDataElement)):
+ v = v.detach()
+ data_list.append(v)
+ if len(data_list) > 0:
+ new_data.set_data({f'{k}': data_list})
+ return new_data
+
+ # Tensor-like methods
+ def numpy(self) -> 'BaseDataElement':
+ """Convert all tensors to np.ndarray in data."""
+ new_data = self.new()
+ for k, v_list in self.items():
+ data_list = []
+ for v in v_list:
+ if isinstance(v, (torch.Tensor, BaseDataElement)):
+ v = v.detach().cpu().numpy()
+ data_list.append(v)
+ if len(data_list) > 0:
+ new_data.set_data({f'{k}': data_list})
+ return new_data
+
+ def to_tensor(self) -> 'BaseDataElement':
+ """Convert all np.ndarray to tensor in data."""
+ new_data = self.new()
+ for k, v_list in self.items():
+ data_list = []
+ for v in v_list:
+ if isinstance(v, np.ndarray):
+ v = torch.from_numpy(v)
+ elif isinstance(v, BaseDataElement):
+ v = v.to_tensor()
+ data_list.append(v)
+ if len(data_list) > 0:
+ new_data.set_data({f'{k}': data_list})
+ return new_data
+
+ # Tensor-like methods
+ def clone(self) -> 'BaseDataElement':
+ """Deep copy the current data element.
+
+ Returns:
+ BaseDataElement: The copy of current data element.
+ """
+ clone_data = self.__class__()
+ clone_data.set_metainfo(dict(self.metainfo_items()))
+
+ for k, v_list in self.items():
+ clone_item_list = []
+ for v in v_list:
+ clone_item_list.append(v.clone())
+ clone_data.set_data({k: clone_item_list})
+ return clone_data
+
+
+TrackSampleList = List[TrackDataSample]
+OptTrackSampleList = Optional[TrackSampleList]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/testing/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/testing/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..766fb471022ee6f2e4e1ff13a52040ae57772e53
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/testing/__init__.py
@@ -0,0 +1,12 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from ._fast_stop_training_hook import FastStopTrainingHook # noqa: F401,F403
+from ._utils import (demo_mm_inputs, demo_mm_proposals,
+ demo_mm_sampling_results, demo_track_inputs,
+ get_detector_cfg, get_roi_head_cfg, random_boxes,
+ replace_to_ceph)
+
+__all__ = [
+ 'demo_mm_inputs', 'get_detector_cfg', 'get_roi_head_cfg',
+ 'demo_mm_proposals', 'demo_mm_sampling_results', 'replace_to_ceph',
+ 'demo_track_inputs', 'VideoDataSampleFeeder', 'random_boxes'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/testing/_fast_stop_training_hook.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/testing/_fast_stop_training_hook.py
new file mode 100644
index 0000000000000000000000000000000000000000..f8e3d11439f875d2c9a6ce6b8a0b33acc832c2c5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/testing/_fast_stop_training_hook.py
@@ -0,0 +1,27 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.hooks import Hook
+
+from mmdet.registry import HOOKS
+
+
+@HOOKS.register_module()
+class FastStopTrainingHook(Hook):
+ """Set runner's epoch information to the model."""
+
+ def __init__(self, by_epoch, save_ckpt=False, stop_iter_or_epoch=5):
+ self.by_epoch = by_epoch
+ self.save_ckpt = save_ckpt
+ self.stop_iter_or_epoch = stop_iter_or_epoch
+
+ def after_train_iter(self, runner, batch_idx: int, data_batch: None,
+ outputs: None) -> None:
+ if self.save_ckpt and self.by_epoch:
+ # If it is epoch-based and want to save weights,
+ # we must run at least 1 epoch.
+ return
+ if runner.iter >= self.stop_iter_or_epoch:
+ raise RuntimeError('quick exit')
+
+ def after_train_epoch(self, runner) -> None:
+ if runner.epoch >= self.stop_iter_or_epoch - 1:
+ raise RuntimeError('quick exit')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/testing/_utils.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/testing/_utils.py
new file mode 100644
index 0000000000000000000000000000000000000000..c4d3a86deab17e9c5acd1b1fe7f42e0bfa78943d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/testing/_utils.py
@@ -0,0 +1,469 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from os.path import dirname, exists, join
+
+import numpy as np
+import torch
+from mmengine.config import Config
+from mmengine.dataset import pseudo_collate
+from mmengine.structures import InstanceData, PixelData
+
+from mmdet.utils.util_random import ensure_rng
+from ..registry import TASK_UTILS
+from ..structures import DetDataSample, TrackDataSample
+from ..structures.bbox import HorizontalBoxes
+
+
+def _get_config_directory():
+ """Find the predefined detector config directory."""
+ try:
+ # Assume we are running in the source mmdetection repo
+ repo_dpath = dirname(dirname(dirname(__file__)))
+ except NameError:
+ # For IPython development when this __file__ is not defined
+ import mmdet
+ repo_dpath = dirname(dirname(mmdet.__file__))
+ config_dpath = join(repo_dpath, 'configs')
+ if not exists(config_dpath):
+ raise Exception('Cannot find config path')
+ return config_dpath
+
+
+def _get_config_module(fname):
+ """Load a configuration as a python module."""
+ config_dpath = _get_config_directory()
+ config_fpath = join(config_dpath, fname)
+ config_mod = Config.fromfile(config_fpath)
+ return config_mod
+
+
+def get_detector_cfg(fname):
+ """Grab configs necessary to create a detector.
+
+ These are deep copied to allow for safe modification of parameters without
+ influencing other tests.
+ """
+ config = _get_config_module(fname)
+ model = copy.deepcopy(config.model)
+ return model
+
+
+def get_roi_head_cfg(fname):
+ """Grab configs necessary to create a roi_head.
+
+ These are deep copied to allow for safe modification of parameters without
+ influencing other tests.
+ """
+ config = _get_config_module(fname)
+ model = copy.deepcopy(config.model)
+
+ roi_head = model.roi_head
+ train_cfg = None if model.train_cfg is None else model.train_cfg.rcnn
+ test_cfg = None if model.test_cfg is None else model.test_cfg.rcnn
+ roi_head.update(dict(train_cfg=train_cfg, test_cfg=test_cfg))
+ return roi_head
+
+
+def _rand_bboxes(rng, num_boxes, w, h):
+ cx, cy, bw, bh = rng.rand(num_boxes, 4).T
+
+ tl_x = ((cx * w) - (w * bw / 2)).clip(0, w)
+ tl_y = ((cy * h) - (h * bh / 2)).clip(0, h)
+ br_x = ((cx * w) + (w * bw / 2)).clip(0, w)
+ br_y = ((cy * h) + (h * bh / 2)).clip(0, h)
+
+ bboxes = np.vstack([tl_x, tl_y, br_x, br_y]).T
+ return bboxes
+
+
+def _rand_masks(rng, num_boxes, bboxes, img_w, img_h):
+ from mmdet.structures.mask import BitmapMasks
+ masks = np.zeros((num_boxes, img_h, img_w))
+ for i, bbox in enumerate(bboxes):
+ bbox = bbox.astype(np.int32)
+ mask = (rng.rand(1, bbox[3] - bbox[1], bbox[2] - bbox[0]) >
+ 0.3).astype(np.int64)
+ masks[i:i + 1, bbox[1]:bbox[3], bbox[0]:bbox[2]] = mask
+ return BitmapMasks(masks, height=img_h, width=img_w)
+
+
+def demo_mm_inputs(batch_size=2,
+ image_shapes=(3, 128, 128),
+ num_items=None,
+ num_classes=10,
+ sem_seg_output_strides=1,
+ with_mask=False,
+ with_semantic=False,
+ use_box_type=False,
+ device='cpu',
+ texts=None,
+ custom_entities=False):
+ """Create a superset of inputs needed to run test or train batches.
+
+ Args:
+ batch_size (int): batch size. Defaults to 2.
+ image_shapes (List[tuple], Optional): image shape.
+ Defaults to (3, 128, 128)
+ num_items (None | List[int]): specifies the number
+ of boxes in each batch item. Default to None.
+ num_classes (int): number of different labels a
+ box might have. Defaults to 10.
+ with_mask (bool): Whether to return mask annotation.
+ Defaults to False.
+ with_semantic (bool): whether to return semantic.
+ Defaults to False.
+ device (str): Destination device type. Defaults to cpu.
+ """
+ rng = np.random.RandomState(0)
+
+ if isinstance(image_shapes, list):
+ assert len(image_shapes) == batch_size
+ else:
+ image_shapes = [image_shapes] * batch_size
+
+ if isinstance(num_items, list):
+ assert len(num_items) == batch_size
+
+ if texts is not None:
+ assert batch_size == len(texts)
+
+ packed_inputs = []
+ for idx in range(batch_size):
+ image_shape = image_shapes[idx]
+ c, h, w = image_shape
+
+ image = rng.randint(0, 255, size=image_shape, dtype=np.uint8)
+
+ mm_inputs = dict()
+ mm_inputs['inputs'] = torch.from_numpy(image).to(device)
+
+ img_meta = {
+ 'img_id': idx,
+ 'img_shape': image_shape[1:],
+ 'ori_shape': image_shape[1:],
+ 'filename': '.png',
+ 'scale_factor': np.array([1.1, 1.2]),
+ 'flip': False,
+ 'flip_direction': None,
+ 'border': [1, 1, 1, 1] # Only used by CenterNet
+ }
+
+ if texts:
+ img_meta['text'] = texts[idx]
+ img_meta['custom_entities'] = custom_entities
+
+ data_sample = DetDataSample()
+ data_sample.set_metainfo(img_meta)
+
+ # gt_instances
+ gt_instances = InstanceData()
+ if num_items is None:
+ num_boxes = rng.randint(1, 10)
+ else:
+ num_boxes = num_items[idx]
+
+ bboxes = _rand_bboxes(rng, num_boxes, w, h)
+ labels = rng.randint(1, num_classes, size=num_boxes)
+ # TODO: remove this part when all model adapted with BaseBoxes
+ if use_box_type:
+ gt_instances.bboxes = HorizontalBoxes(bboxes, dtype=torch.float32)
+ else:
+ gt_instances.bboxes = torch.FloatTensor(bboxes)
+ gt_instances.labels = torch.LongTensor(labels)
+
+ if with_mask:
+ masks = _rand_masks(rng, num_boxes, bboxes, w, h)
+ gt_instances.masks = masks
+
+ # TODO: waiting for ci to be fixed
+ # masks = np.random.randint(0, 2, (len(bboxes), h, w), dtype=np.uint8)
+ # gt_instances.mask = BitmapMasks(masks, h, w)
+
+ data_sample.gt_instances = gt_instances
+
+ # ignore_instances
+ ignore_instances = InstanceData()
+ bboxes = _rand_bboxes(rng, num_boxes, w, h)
+ if use_box_type:
+ ignore_instances.bboxes = HorizontalBoxes(
+ bboxes, dtype=torch.float32)
+ else:
+ ignore_instances.bboxes = torch.FloatTensor(bboxes)
+ data_sample.ignored_instances = ignore_instances
+
+ # gt_sem_seg
+ if with_semantic:
+ # assume gt_semantic_seg using scale 1/8 of the img
+ gt_semantic_seg = torch.from_numpy(
+ np.random.randint(
+ 0,
+ num_classes, (1, h // sem_seg_output_strides,
+ w // sem_seg_output_strides),
+ dtype=np.uint8))
+ gt_sem_seg_data = dict(sem_seg=gt_semantic_seg)
+ data_sample.gt_sem_seg = PixelData(**gt_sem_seg_data)
+
+ mm_inputs['data_samples'] = data_sample.to(device)
+
+ # TODO: gt_ignore
+
+ packed_inputs.append(mm_inputs)
+ data = pseudo_collate(packed_inputs)
+ return data
+
+
+def demo_mm_proposals(image_shapes, num_proposals, device='cpu'):
+ """Create a list of fake porposals.
+
+ Args:
+ image_shapes (list[tuple[int]]): Batch image shapes.
+ num_proposals (int): The number of fake proposals.
+ """
+ rng = np.random.RandomState(0)
+
+ results = []
+ for img_shape in image_shapes:
+ result = InstanceData()
+ w, h = img_shape[1:]
+ proposals = _rand_bboxes(rng, num_proposals, w, h)
+ result.bboxes = torch.from_numpy(proposals).float()
+ result.scores = torch.from_numpy(rng.rand(num_proposals)).float()
+ result.labels = torch.zeros(num_proposals).long()
+ results.append(result.to(device))
+ return results
+
+
+def demo_mm_sampling_results(proposals_list,
+ batch_gt_instances,
+ batch_gt_instances_ignore=None,
+ assigner_cfg=None,
+ sampler_cfg=None,
+ feats=None):
+ """Create sample results that can be passed to BBoxHead.get_targets."""
+ assert len(proposals_list) == len(batch_gt_instances)
+ if batch_gt_instances_ignore is None:
+ batch_gt_instances_ignore = [None for _ in batch_gt_instances]
+ else:
+ assert len(batch_gt_instances_ignore) == len(batch_gt_instances)
+
+ default_assigner_cfg = dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ ignore_iof_thr=-1)
+ assigner_cfg = assigner_cfg if assigner_cfg is not None \
+ else default_assigner_cfg
+ default_sampler_cfg = dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True)
+ sampler_cfg = sampler_cfg if sampler_cfg is not None \
+ else default_sampler_cfg
+ bbox_assigner = TASK_UTILS.build(assigner_cfg)
+ bbox_sampler = TASK_UTILS.build(sampler_cfg)
+
+ sampling_results = []
+ for i in range(len(batch_gt_instances)):
+ if feats is not None:
+ feats = [lvl_feat[i][None] for lvl_feat in feats]
+ # rename proposals.bboxes to proposals.priors
+ proposals = proposals_list[i]
+ proposals.priors = proposals.pop('bboxes')
+
+ assign_result = bbox_assigner.assign(proposals, batch_gt_instances[i],
+ batch_gt_instances_ignore[i])
+ sampling_result = bbox_sampler.sample(
+ assign_result, proposals, batch_gt_instances[i], feats=feats)
+ sampling_results.append(sampling_result)
+
+ return sampling_results
+
+
+def demo_track_inputs(batch_size=1,
+ num_frames=2,
+ key_frames_inds=None,
+ image_shapes=(3, 128, 128),
+ num_items=None,
+ num_classes=1,
+ with_mask=False,
+ with_semantic=False):
+ """Create a superset of inputs needed to run test or train batches.
+
+ Args:
+ batch_size (int): batch size. Default to 1.
+ num_frames (int): The number of frames.
+ key_frames_inds (List): The indices of key frames.
+ image_shapes (List[tuple], Optional): image shape.
+ Default to (3, 128, 128)
+ num_items (None | List[int]): specifies the number
+ of boxes in each batch item. Default to None.
+ num_classes (int): number of different labels a
+ box might have. Default to 1.
+ with_mask (bool): Whether to return mask annotation.
+ Defaults to False.
+ with_semantic (bool): whether to return semantic.
+ Default to False.
+ """
+ rng = np.random.RandomState(0)
+
+ # Make sure the length of image_shapes is equal to ``batch_size``
+ if isinstance(image_shapes, list):
+ assert len(image_shapes) == batch_size
+ else:
+ image_shapes = [image_shapes] * batch_size
+
+ packed_inputs = []
+ for idx in range(batch_size):
+ mm_inputs = dict(inputs=dict())
+ _, h, w = image_shapes[idx]
+
+ imgs = rng.randint(
+ 0, 255, size=(num_frames, *image_shapes[idx]), dtype=np.uint8)
+ mm_inputs['inputs'] = torch.from_numpy(imgs)
+
+ img_meta = {
+ 'img_id': idx,
+ 'img_shape': image_shapes[idx][-2:],
+ 'ori_shape': image_shapes[idx][-2:],
+ 'filename': '.png',
+ 'scale_factor': np.array([1.1, 1.2]),
+ 'flip': False,
+ 'flip_direction': None,
+ 'is_video_data': True,
+ }
+
+ video_data_samples = []
+ for i in range(num_frames):
+ data_sample = DetDataSample()
+ img_meta['frame_id'] = i
+ data_sample.set_metainfo(img_meta)
+
+ # gt_instances
+ gt_instances = InstanceData()
+ if num_items is None:
+ num_boxes = rng.randint(1, 10)
+ else:
+ num_boxes = num_items[idx]
+
+ bboxes = _rand_bboxes(rng, num_boxes, w, h)
+ labels = rng.randint(0, num_classes, size=num_boxes)
+ instances_id = rng.randint(100, num_classes + 100, size=num_boxes)
+ gt_instances.bboxes = torch.FloatTensor(bboxes)
+ gt_instances.labels = torch.LongTensor(labels)
+ gt_instances.instances_ids = torch.LongTensor(instances_id)
+
+ if with_mask:
+ masks = _rand_masks(rng, num_boxes, bboxes, w, h)
+ gt_instances.masks = masks
+
+ data_sample.gt_instances = gt_instances
+ # ignore_instances
+ ignore_instances = InstanceData()
+ bboxes = _rand_bboxes(rng, num_boxes, w, h)
+ ignore_instances.bboxes = bboxes
+ data_sample.ignored_instances = ignore_instances
+
+ video_data_samples.append(data_sample)
+
+ track_data_sample = TrackDataSample()
+ track_data_sample.video_data_samples = video_data_samples
+ if key_frames_inds is not None:
+ assert isinstance(
+ key_frames_inds,
+ list) and len(key_frames_inds) < num_frames and max(
+ key_frames_inds) < num_frames
+ ref_frames_inds = [
+ i for i in range(num_frames) if i not in key_frames_inds
+ ]
+ track_data_sample.set_metainfo(
+ dict(key_frames_inds=key_frames_inds))
+ track_data_sample.set_metainfo(
+ dict(ref_frames_inds=ref_frames_inds))
+ mm_inputs['data_samples'] = track_data_sample
+
+ # TODO: gt_ignore
+ packed_inputs.append(mm_inputs)
+ data = pseudo_collate(packed_inputs)
+ return data
+
+
+def random_boxes(num=1, scale=1, rng=None):
+ """Simple version of ``kwimage.Boxes.random``
+ Returns:
+ Tensor: shape (n, 4) in x1, y1, x2, y2 format.
+ References:
+ https://gitlab.kitware.com/computer-vision/kwimage/blob/master/kwimage/structs/boxes.py#L1390 # noqa: E501
+ Example:
+ >>> num = 3
+ >>> scale = 512
+ >>> rng = 0
+ >>> boxes = random_boxes(num, scale, rng)
+ >>> print(boxes)
+ tensor([[280.9925, 278.9802, 308.6148, 366.1769],
+ [216.9113, 330.6978, 224.0446, 456.5878],
+ [405.3632, 196.3221, 493.3953, 270.7942]])
+ """
+ rng = ensure_rng(rng)
+
+ tlbr = rng.rand(num, 4).astype(np.float32)
+
+ tl_x = np.minimum(tlbr[:, 0], tlbr[:, 2])
+ tl_y = np.minimum(tlbr[:, 1], tlbr[:, 3])
+ br_x = np.maximum(tlbr[:, 0], tlbr[:, 2])
+ br_y = np.maximum(tlbr[:, 1], tlbr[:, 3])
+
+ tlbr[:, 0] = tl_x * scale
+ tlbr[:, 1] = tl_y * scale
+ tlbr[:, 2] = br_x * scale
+ tlbr[:, 3] = br_y * scale
+
+ boxes = torch.from_numpy(tlbr)
+ return boxes
+
+
+# TODO: Support full ceph
+def replace_to_ceph(cfg):
+ backend_args = dict(
+ backend='petrel',
+ path_mapping=dict({
+ './data/': 's3://openmmlab/datasets/detection/',
+ 'data/': 's3://openmmlab/datasets/detection/'
+ }))
+
+ # TODO: name is a reserved interface, which will be used later.
+ def _process_pipeline(dataset, name):
+
+ def replace_img(pipeline):
+ if pipeline['type'] == 'LoadImageFromFile':
+ pipeline['backend_args'] = backend_args
+
+ def replace_ann(pipeline):
+ if pipeline['type'] == 'LoadAnnotations' or pipeline[
+ 'type'] == 'LoadPanopticAnnotations':
+ pipeline['backend_args'] = backend_args
+
+ if 'pipeline' in dataset:
+ replace_img(dataset.pipeline[0])
+ replace_ann(dataset.pipeline[1])
+ if 'dataset' in dataset:
+ # dataset wrapper
+ replace_img(dataset.dataset.pipeline[0])
+ replace_ann(dataset.dataset.pipeline[1])
+ else:
+ # dataset wrapper
+ replace_img(dataset.dataset.pipeline[0])
+ replace_ann(dataset.dataset.pipeline[1])
+
+ def _process_evaluator(evaluator, name):
+ if evaluator['type'] == 'CocoPanopticMetric':
+ evaluator['backend_args'] = backend_args
+
+ # half ceph
+ _process_pipeline(cfg.train_dataloader.dataset, cfg.filename)
+ _process_pipeline(cfg.val_dataloader.dataset, cfg.filename)
+ _process_pipeline(cfg.test_dataloader.dataset, cfg.filename)
+ _process_evaluator(cfg.val_evaluator, cfg.filename)
+ _process_evaluator(cfg.test_evaluator, cfg.filename)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..449a890bac411f84790eb3d014175e3a48757847
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/__init__.py
@@ -0,0 +1,28 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .collect_env import collect_env
+from .compat_config import compat_cfg
+from .dist_utils import (all_reduce_dict, allreduce_grads, reduce_mean,
+ sync_random_seed)
+from .logger import get_caller_name, log_img_scale
+from .memory import AvoidCUDAOOM, AvoidOOM
+from .misc import (find_latest_checkpoint, get_test_pipeline_cfg,
+ update_data_root)
+from .mot_error_visualize import imshow_mot_errors
+from .replace_cfg_vals import replace_cfg_vals
+from .setup_env import (register_all_modules, setup_cache_size_limit_of_dynamo,
+ setup_multi_processes)
+from .split_batch import split_batch
+from .typing_utils import (ConfigType, InstanceList, MultiConfig,
+ OptConfigType, OptInstanceList, OptMultiConfig,
+ OptPixelList, PixelList, RangeType)
+
+__all__ = [
+ 'collect_env', 'find_latest_checkpoint', 'update_data_root',
+ 'setup_multi_processes', 'get_caller_name', 'log_img_scale', 'compat_cfg',
+ 'split_batch', 'register_all_modules', 'replace_cfg_vals', 'AvoidOOM',
+ 'AvoidCUDAOOM', 'all_reduce_dict', 'allreduce_grads', 'reduce_mean',
+ 'sync_random_seed', 'ConfigType', 'InstanceList', 'MultiConfig',
+ 'OptConfigType', 'OptInstanceList', 'OptMultiConfig', 'OptPixelList',
+ 'PixelList', 'RangeType', 'get_test_pipeline_cfg',
+ 'setup_cache_size_limit_of_dynamo', 'imshow_mot_errors'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/benchmark.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/benchmark.py
new file mode 100644
index 0000000000000000000000000000000000000000..5419b2d175e3c48c063a39ae28758b386f9ab597
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/benchmark.py
@@ -0,0 +1,529 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import time
+from functools import partial
+from typing import List, Optional, Union
+
+import numpy as np
+import torch
+import torch.nn as nn
+from mmcv.cnn import fuse_conv_bn
+# TODO need update
+# from mmcv.runner import wrap_fp16_model
+from mmengine import MMLogger
+from mmengine.config import Config
+from mmengine.device import get_max_cuda_memory
+from mmengine.dist import get_world_size
+from mmengine.runner import Runner, load_checkpoint
+from mmengine.utils.dl_utils import set_multi_processing
+from torch.nn.parallel import DistributedDataParallel
+
+from mmdet.registry import DATASETS, MODELS
+
+try:
+ import psutil
+except ImportError:
+ psutil = None
+
+
+def custom_round(value: Union[int, float],
+ factor: Union[int, float],
+ precision: int = 2) -> float:
+ """Custom round function."""
+ return round(value / factor, precision)
+
+
+gb_round = partial(custom_round, factor=1024**3)
+
+
+def print_log(msg: str, logger: Optional[MMLogger] = None) -> None:
+ """Print a log message."""
+ if logger is None:
+ print(msg, flush=True)
+ else:
+ logger.info(msg)
+
+
+def print_process_memory(p: psutil.Process,
+ logger: Optional[MMLogger] = None) -> None:
+ """print process memory info."""
+ mem_used = gb_round(psutil.virtual_memory().used)
+ memory_full_info = p.memory_full_info()
+ uss_mem = gb_round(memory_full_info.uss)
+ if hasattr(memory_full_info, 'pss'):
+ pss_mem = gb_round(memory_full_info.pss)
+
+ for children in p.children():
+ child_mem_info = children.memory_full_info()
+ uss_mem += gb_round(child_mem_info.uss)
+ if hasattr(child_mem_info, 'pss'):
+ pss_mem += gb_round(child_mem_info.pss)
+
+ process_count = 1 + len(p.children())
+
+ log_msg = f'(GB) mem_used: {mem_used:.2f} | uss: {uss_mem:.2f} | '
+ if hasattr(memory_full_info, 'pss'):
+ log_msg += f'pss: {pss_mem:.2f} | '
+ log_msg += f'total_proc: {process_count}'
+ print_log(log_msg, logger)
+
+
+class BaseBenchmark:
+ """The benchmark base class.
+
+ The ``run`` method is an external calling interface, and it will
+ call the ``run_once`` method ``repeat_num`` times for benchmarking.
+ Finally, call the ``average_multiple_runs`` method to further process
+ the results of multiple runs.
+
+ Args:
+ max_iter (int): maximum iterations of benchmark.
+ log_interval (int): interval of logging.
+ num_warmup (int): Number of Warmup.
+ logger (MMLogger, optional): Formatted logger used to record messages.
+ """
+
+ def __init__(self,
+ max_iter: int,
+ log_interval: int,
+ num_warmup: int,
+ logger: Optional[MMLogger] = None):
+ self.max_iter = max_iter
+ self.log_interval = log_interval
+ self.num_warmup = num_warmup
+ self.logger = logger
+
+ def run(self, repeat_num: int = 1) -> dict:
+ """benchmark entry method.
+
+ Args:
+ repeat_num (int): Number of repeat benchmark.
+ Defaults to 1.
+ """
+ assert repeat_num >= 1
+
+ results = []
+ for _ in range(repeat_num):
+ results.append(self.run_once())
+
+ results = self.average_multiple_runs(results)
+ return results
+
+ def run_once(self) -> dict:
+ """Executes the benchmark once."""
+ raise NotImplementedError()
+
+ def average_multiple_runs(self, results: List[dict]) -> dict:
+ """Average the results of multiple runs."""
+ raise NotImplementedError()
+
+
+class InferenceBenchmark(BaseBenchmark):
+ """The inference benchmark class. It will be statistical inference FPS,
+ CUDA memory and CPU memory information.
+
+ Args:
+ cfg (mmengine.Config): config.
+ checkpoint (str): Accept local filepath, URL, ``torchvision://xxx``,
+ ``open-mmlab://xxx``.
+ distributed (bool): distributed testing flag.
+ is_fuse_conv_bn (bool): Whether to fuse conv and bn, this will
+ slightly increase the inference speed.
+ max_iter (int): maximum iterations of benchmark. Defaults to 2000.
+ log_interval (int): interval of logging. Defaults to 50.
+ num_warmup (int): Number of Warmup. Defaults to 5.
+ logger (MMLogger, optional): Formatted logger used to record messages.
+ """
+
+ def __init__(self,
+ cfg: Config,
+ checkpoint: str,
+ distributed: bool,
+ is_fuse_conv_bn: bool,
+ max_iter: int = 2000,
+ log_interval: int = 50,
+ num_warmup: int = 5,
+ logger: Optional[MMLogger] = None):
+ super().__init__(max_iter, log_interval, num_warmup, logger)
+
+ assert get_world_size(
+ ) == 1, 'Inference benchmark does not allow distributed multi-GPU'
+
+ self.cfg = copy.deepcopy(cfg)
+ self.distributed = distributed
+
+ if psutil is None:
+ raise ImportError('psutil is not installed, please install it by: '
+ 'pip install psutil')
+
+ self._process = psutil.Process()
+ env_cfg = self.cfg.get('env_cfg')
+ if env_cfg.get('cudnn_benchmark'):
+ torch.backends.cudnn.benchmark = True
+
+ mp_cfg: dict = env_cfg.get('mp_cfg', {})
+ set_multi_processing(**mp_cfg, distributed=self.distributed)
+
+ print_log('before build: ', self.logger)
+ print_process_memory(self._process, self.logger)
+
+ self.model = self._init_model(checkpoint, is_fuse_conv_bn)
+
+ # Because multiple processes will occupy additional CPU resources,
+ # FPS statistics will be more unstable when num_workers is not 0.
+ # It is reasonable to set num_workers to 0.
+ dataloader_cfg = cfg.test_dataloader
+ dataloader_cfg['num_workers'] = 0
+ dataloader_cfg['batch_size'] = 1
+ dataloader_cfg['persistent_workers'] = False
+ self.data_loader = Runner.build_dataloader(dataloader_cfg)
+
+ print_log('after build: ', self.logger)
+ print_process_memory(self._process, self.logger)
+
+ def _init_model(self, checkpoint: str, is_fuse_conv_bn: bool) -> nn.Module:
+ """Initialize the model."""
+ model = MODELS.build(self.cfg.model)
+ # TODO need update
+ # fp16_cfg = self.cfg.get('fp16', None)
+ # if fp16_cfg is not None:
+ # wrap_fp16_model(model)
+
+ load_checkpoint(model, checkpoint, map_location='cpu')
+ if is_fuse_conv_bn:
+ model = fuse_conv_bn(model)
+
+ model = model.cuda()
+
+ if self.distributed:
+ model = DistributedDataParallel(
+ model,
+ device_ids=[torch.cuda.current_device()],
+ broadcast_buffers=False,
+ find_unused_parameters=False)
+
+ model.eval()
+ return model
+
+ def run_once(self) -> dict:
+ """Executes the benchmark once."""
+ pure_inf_time = 0
+ fps = 0
+
+ for i, data in enumerate(self.data_loader):
+
+ if (i + 1) % self.log_interval == 0:
+ print_log('==================================', self.logger)
+
+ torch.cuda.synchronize()
+ start_time = time.perf_counter()
+
+ with torch.no_grad():
+ self.model.test_step(data)
+
+ torch.cuda.synchronize()
+ elapsed = time.perf_counter() - start_time
+
+ if i >= self.num_warmup:
+ pure_inf_time += elapsed
+ if (i + 1) % self.log_interval == 0:
+ fps = (i + 1 - self.num_warmup) / pure_inf_time
+ cuda_memory = get_max_cuda_memory()
+
+ print_log(
+ f'Done image [{i + 1:<3}/{self.max_iter}], '
+ f'fps: {fps:.1f} img/s, '
+ f'times per image: {1000 / fps:.1f} ms/img, '
+ f'cuda memory: {cuda_memory} MB', self.logger)
+ print_process_memory(self._process, self.logger)
+
+ if (i + 1) == self.max_iter:
+ fps = (i + 1 - self.num_warmup) / pure_inf_time
+ break
+
+ return {'fps': fps}
+
+ def average_multiple_runs(self, results: List[dict]) -> dict:
+ """Average the results of multiple runs."""
+ print_log('============== Done ==================', self.logger)
+
+ fps_list_ = [round(result['fps'], 1) for result in results]
+ avg_fps_ = sum(fps_list_) / len(fps_list_)
+ outputs = {'avg_fps': avg_fps_, 'fps_list': fps_list_}
+
+ if len(fps_list_) > 1:
+ times_pre_image_list_ = [
+ round(1000 / result['fps'], 1) for result in results
+ ]
+ avg_times_pre_image_ = sum(times_pre_image_list_) / len(
+ times_pre_image_list_)
+
+ print_log(
+ f'Overall fps: {fps_list_}[{avg_fps_:.1f}] img/s, '
+ 'times per image: '
+ f'{times_pre_image_list_}[{avg_times_pre_image_:.1f}] '
+ 'ms/img', self.logger)
+ else:
+ print_log(
+ f'Overall fps: {fps_list_[0]:.1f} img/s, '
+ f'times per image: {1000 / fps_list_[0]:.1f} ms/img',
+ self.logger)
+
+ print_log(f'cuda memory: {get_max_cuda_memory()} MB', self.logger)
+ print_process_memory(self._process, self.logger)
+
+ return outputs
+
+
+class DataLoaderBenchmark(BaseBenchmark):
+ """The dataloader benchmark class. It will be statistical inference FPS and
+ CPU memory information.
+
+ Args:
+ cfg (mmengine.Config): config.
+ distributed (bool): distributed testing flag.
+ dataset_type (str): benchmark data type, only supports ``train``,
+ ``val`` and ``test``.
+ max_iter (int): maximum iterations of benchmark. Defaults to 2000.
+ log_interval (int): interval of logging. Defaults to 50.
+ num_warmup (int): Number of Warmup. Defaults to 5.
+ logger (MMLogger, optional): Formatted logger used to record messages.
+ """
+
+ def __init__(self,
+ cfg: Config,
+ distributed: bool,
+ dataset_type: str,
+ max_iter: int = 2000,
+ log_interval: int = 50,
+ num_warmup: int = 5,
+ logger: Optional[MMLogger] = None):
+ super().__init__(max_iter, log_interval, num_warmup, logger)
+
+ assert dataset_type in ['train', 'val', 'test'], \
+ 'dataset_type only supports train,' \
+ f' val and test, but got {dataset_type}'
+ assert get_world_size(
+ ) == 1, 'Dataloader benchmark does not allow distributed multi-GPU'
+
+ self.cfg = copy.deepcopy(cfg)
+ self.distributed = distributed
+
+ if psutil is None:
+ raise ImportError('psutil is not installed, please install it by: '
+ 'pip install psutil')
+ self._process = psutil.Process()
+
+ mp_cfg = self.cfg.get('env_cfg', {}).get('mp_cfg')
+ if mp_cfg is not None:
+ set_multi_processing(distributed=self.distributed, **mp_cfg)
+ else:
+ set_multi_processing(distributed=self.distributed)
+
+ print_log('before build: ', self.logger)
+ print_process_memory(self._process, self.logger)
+
+ if dataset_type == 'train':
+ self.data_loader = Runner.build_dataloader(cfg.train_dataloader)
+ elif dataset_type == 'test':
+ self.data_loader = Runner.build_dataloader(cfg.test_dataloader)
+ else:
+ self.data_loader = Runner.build_dataloader(cfg.val_dataloader)
+
+ self.batch_size = self.data_loader.batch_size
+ self.num_workers = self.data_loader.num_workers
+
+ print_log('after build: ', self.logger)
+ print_process_memory(self._process, self.logger)
+
+ def run_once(self) -> dict:
+ """Executes the benchmark once."""
+ pure_inf_time = 0
+ fps = 0
+
+ # benchmark with 2000 image and take the average
+ start_time = time.perf_counter()
+ for i, data in enumerate(self.data_loader):
+ elapsed = time.perf_counter() - start_time
+
+ if (i + 1) % self.log_interval == 0:
+ print_log('==================================', self.logger)
+
+ if i >= self.num_warmup:
+ pure_inf_time += elapsed
+ if (i + 1) % self.log_interval == 0:
+ fps = (i + 1 - self.num_warmup) / pure_inf_time
+
+ print_log(
+ f'Done batch [{i + 1:<3}/{self.max_iter}], '
+ f'fps: {fps:.1f} batch/s, '
+ f'times per batch: {1000 / fps:.1f} ms/batch, '
+ f'batch size: {self.batch_size}, num_workers: '
+ f'{self.num_workers}', self.logger)
+ print_process_memory(self._process, self.logger)
+
+ if (i + 1) == self.max_iter:
+ fps = (i + 1 - self.num_warmup) / pure_inf_time
+ break
+
+ start_time = time.perf_counter()
+
+ return {'fps': fps}
+
+ def average_multiple_runs(self, results: List[dict]) -> dict:
+ """Average the results of multiple runs."""
+ print_log('============== Done ==================', self.logger)
+
+ fps_list_ = [round(result['fps'], 1) for result in results]
+ avg_fps_ = sum(fps_list_) / len(fps_list_)
+ outputs = {'avg_fps': avg_fps_, 'fps_list': fps_list_}
+
+ if len(fps_list_) > 1:
+ times_pre_image_list_ = [
+ round(1000 / result['fps'], 1) for result in results
+ ]
+ avg_times_pre_image_ = sum(times_pre_image_list_) / len(
+ times_pre_image_list_)
+
+ print_log(
+ f'Overall fps: {fps_list_}[{avg_fps_:.1f}] img/s, '
+ 'times per batch: '
+ f'{times_pre_image_list_}[{avg_times_pre_image_:.1f}] '
+ f'ms/batch, batch size: {self.batch_size}, num_workers: '
+ f'{self.num_workers}', self.logger)
+ else:
+ print_log(
+ f'Overall fps: {fps_list_[0]:.1f} batch/s, '
+ f'times per batch: {1000 / fps_list_[0]:.1f} ms/batch, '
+ f'batch size: {self.batch_size}, num_workers: '
+ f'{self.num_workers}', self.logger)
+
+ print_process_memory(self._process, self.logger)
+
+ return outputs
+
+
+class DatasetBenchmark(BaseBenchmark):
+ """The dataset benchmark class. It will be statistical inference FPS, FPS
+ pre transform and CPU memory information.
+
+ Args:
+ cfg (mmengine.Config): config.
+ dataset_type (str): benchmark data type, only supports ``train``,
+ ``val`` and ``test``.
+ max_iter (int): maximum iterations of benchmark. Defaults to 2000.
+ log_interval (int): interval of logging. Defaults to 50.
+ num_warmup (int): Number of Warmup. Defaults to 5.
+ logger (MMLogger, optional): Formatted logger used to record messages.
+ """
+
+ def __init__(self,
+ cfg: Config,
+ dataset_type: str,
+ max_iter: int = 2000,
+ log_interval: int = 50,
+ num_warmup: int = 5,
+ logger: Optional[MMLogger] = None):
+ super().__init__(max_iter, log_interval, num_warmup, logger)
+ assert dataset_type in ['train', 'val', 'test'], \
+ 'dataset_type only supports train,' \
+ f' val and test, but got {dataset_type}'
+ assert get_world_size(
+ ) == 1, 'Dataset benchmark does not allow distributed multi-GPU'
+ self.cfg = copy.deepcopy(cfg)
+
+ if dataset_type == 'train':
+ dataloader_cfg = copy.deepcopy(cfg.train_dataloader)
+ elif dataset_type == 'test':
+ dataloader_cfg = copy.deepcopy(cfg.test_dataloader)
+ else:
+ dataloader_cfg = copy.deepcopy(cfg.val_dataloader)
+
+ dataset_cfg = dataloader_cfg.pop('dataset')
+ dataset = DATASETS.build(dataset_cfg)
+ if hasattr(dataset, 'full_init'):
+ dataset.full_init()
+ self.dataset = dataset
+
+ def run_once(self) -> dict:
+ """Executes the benchmark once."""
+ pure_inf_time = 0
+ fps = 0
+
+ total_index = list(range(len(self.dataset)))
+ np.random.shuffle(total_index)
+
+ start_time = time.perf_counter()
+ for i, idx in enumerate(total_index):
+ if (i + 1) % self.log_interval == 0:
+ print_log('==================================', self.logger)
+
+ get_data_info_start_time = time.perf_counter()
+ data_info = self.dataset.get_data_info(idx)
+ get_data_info_elapsed = time.perf_counter(
+ ) - get_data_info_start_time
+
+ if (i + 1) % self.log_interval == 0:
+ print_log(f'get_data_info - {get_data_info_elapsed * 1000} ms',
+ self.logger)
+
+ for t in self.dataset.pipeline.transforms:
+ transform_start_time = time.perf_counter()
+ data_info = t(data_info)
+ transform_elapsed = time.perf_counter() - transform_start_time
+
+ if (i + 1) % self.log_interval == 0:
+ print_log(
+ f'{t.__class__.__name__} - '
+ f'{transform_elapsed * 1000} ms', self.logger)
+
+ if data_info is None:
+ break
+
+ elapsed = time.perf_counter() - start_time
+
+ if i >= self.num_warmup:
+ pure_inf_time += elapsed
+ if (i + 1) % self.log_interval == 0:
+ fps = (i + 1 - self.num_warmup) / pure_inf_time
+
+ print_log(
+ f'Done img [{i + 1:<3}/{self.max_iter}], '
+ f'fps: {fps:.1f} img/s, '
+ f'times per img: {1000 / fps:.1f} ms/img', self.logger)
+
+ if (i + 1) == self.max_iter:
+ fps = (i + 1 - self.num_warmup) / pure_inf_time
+ break
+
+ start_time = time.perf_counter()
+
+ return {'fps': fps}
+
+ def average_multiple_runs(self, results: List[dict]) -> dict:
+ """Average the results of multiple runs."""
+ print_log('============== Done ==================', self.logger)
+
+ fps_list_ = [round(result['fps'], 1) for result in results]
+ avg_fps_ = sum(fps_list_) / len(fps_list_)
+ outputs = {'avg_fps': avg_fps_, 'fps_list': fps_list_}
+
+ if len(fps_list_) > 1:
+ times_pre_image_list_ = [
+ round(1000 / result['fps'], 1) for result in results
+ ]
+ avg_times_pre_image_ = sum(times_pre_image_list_) / len(
+ times_pre_image_list_)
+
+ print_log(
+ f'Overall fps: {fps_list_}[{avg_fps_:.1f}] img/s, '
+ 'times per img: '
+ f'{times_pre_image_list_}[{avg_times_pre_image_:.1f}] '
+ 'ms/img', self.logger)
+ else:
+ print_log(
+ f'Overall fps: {fps_list_[0]:.1f} img/s, '
+ f'times per img: {1000 / fps_list_[0]:.1f} ms/img',
+ self.logger)
+
+ return outputs
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/collect_env.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/collect_env.py
new file mode 100644
index 0000000000000000000000000000000000000000..b0eed80fe2e4630b78ea3b13fde6046914e47e8b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/collect_env.py
@@ -0,0 +1,17 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from mmengine.utils import get_git_hash
+from mmengine.utils.dl_utils import collect_env as collect_base_env
+
+import mmdet
+
+
+def collect_env():
+ """Collect the information of the running environments."""
+ env_info = collect_base_env()
+ env_info['MMDetection'] = mmdet.__version__ + '+' + get_git_hash()[:7]
+ return env_info
+
+
+if __name__ == '__main__':
+ for name, val in collect_env().items():
+ print(f'{name}: {val}')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/compat_config.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/compat_config.py
new file mode 100644
index 0000000000000000000000000000000000000000..133adb65c2276401eca947e223e5b7c1760de418
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/compat_config.py
@@ -0,0 +1,139 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+import warnings
+
+from mmengine.config import ConfigDict
+
+
+def compat_cfg(cfg):
+ """This function would modify some filed to keep the compatibility of
+ config.
+
+ For example, it will move some args which will be deprecated to the correct
+ fields.
+ """
+ cfg = copy.deepcopy(cfg)
+ cfg = compat_imgs_per_gpu(cfg)
+ cfg = compat_loader_args(cfg)
+ cfg = compat_runner_args(cfg)
+ return cfg
+
+
+def compat_runner_args(cfg):
+ if 'runner' not in cfg:
+ cfg.runner = ConfigDict({
+ 'type': 'EpochBasedRunner',
+ 'max_epochs': cfg.total_epochs
+ })
+ warnings.warn(
+ 'config is now expected to have a `runner` section, '
+ 'please set `runner` in your config.', UserWarning)
+ else:
+ if 'total_epochs' in cfg:
+ assert cfg.total_epochs == cfg.runner.max_epochs
+ return cfg
+
+
+def compat_imgs_per_gpu(cfg):
+ cfg = copy.deepcopy(cfg)
+ if 'imgs_per_gpu' in cfg.data:
+ warnings.warn('"imgs_per_gpu" is deprecated in MMDet V2.0. '
+ 'Please use "samples_per_gpu" instead')
+ if 'samples_per_gpu' in cfg.data:
+ warnings.warn(
+ f'Got "imgs_per_gpu"={cfg.data.imgs_per_gpu} and '
+ f'"samples_per_gpu"={cfg.data.samples_per_gpu}, "imgs_per_gpu"'
+ f'={cfg.data.imgs_per_gpu} is used in this experiments')
+ else:
+ warnings.warn('Automatically set "samples_per_gpu"="imgs_per_gpu"='
+ f'{cfg.data.imgs_per_gpu} in this experiments')
+ cfg.data.samples_per_gpu = cfg.data.imgs_per_gpu
+ return cfg
+
+
+def compat_loader_args(cfg):
+ """Deprecated sample_per_gpu in cfg.data."""
+
+ cfg = copy.deepcopy(cfg)
+ if 'train_dataloader' not in cfg.data:
+ cfg.data['train_dataloader'] = ConfigDict()
+ if 'val_dataloader' not in cfg.data:
+ cfg.data['val_dataloader'] = ConfigDict()
+ if 'test_dataloader' not in cfg.data:
+ cfg.data['test_dataloader'] = ConfigDict()
+
+ # special process for train_dataloader
+ if 'samples_per_gpu' in cfg.data:
+
+ samples_per_gpu = cfg.data.pop('samples_per_gpu')
+ assert 'samples_per_gpu' not in \
+ cfg.data.train_dataloader, ('`samples_per_gpu` are set '
+ 'in `data` field and ` '
+ 'data.train_dataloader` '
+ 'at the same time. '
+ 'Please only set it in '
+ '`data.train_dataloader`. ')
+ cfg.data.train_dataloader['samples_per_gpu'] = samples_per_gpu
+
+ if 'persistent_workers' in cfg.data:
+
+ persistent_workers = cfg.data.pop('persistent_workers')
+ assert 'persistent_workers' not in \
+ cfg.data.train_dataloader, ('`persistent_workers` are set '
+ 'in `data` field and ` '
+ 'data.train_dataloader` '
+ 'at the same time. '
+ 'Please only set it in '
+ '`data.train_dataloader`. ')
+ cfg.data.train_dataloader['persistent_workers'] = persistent_workers
+
+ if 'workers_per_gpu' in cfg.data:
+
+ workers_per_gpu = cfg.data.pop('workers_per_gpu')
+ cfg.data.train_dataloader['workers_per_gpu'] = workers_per_gpu
+ cfg.data.val_dataloader['workers_per_gpu'] = workers_per_gpu
+ cfg.data.test_dataloader['workers_per_gpu'] = workers_per_gpu
+
+ # special process for val_dataloader
+ if 'samples_per_gpu' in cfg.data.val:
+ # keep default value of `sample_per_gpu` is 1
+ assert 'samples_per_gpu' not in \
+ cfg.data.val_dataloader, ('`samples_per_gpu` are set '
+ 'in `data.val` field and ` '
+ 'data.val_dataloader` at '
+ 'the same time. '
+ 'Please only set it in '
+ '`data.val_dataloader`. ')
+ cfg.data.val_dataloader['samples_per_gpu'] = \
+ cfg.data.val.pop('samples_per_gpu')
+ # special process for val_dataloader
+
+ # in case the test dataset is concatenated
+ if isinstance(cfg.data.test, dict):
+ if 'samples_per_gpu' in cfg.data.test:
+ assert 'samples_per_gpu' not in \
+ cfg.data.test_dataloader, ('`samples_per_gpu` are set '
+ 'in `data.test` field and ` '
+ 'data.test_dataloader` '
+ 'at the same time. '
+ 'Please only set it in '
+ '`data.test_dataloader`. ')
+
+ cfg.data.test_dataloader['samples_per_gpu'] = \
+ cfg.data.test.pop('samples_per_gpu')
+
+ elif isinstance(cfg.data.test, list):
+ for ds_cfg in cfg.data.test:
+ if 'samples_per_gpu' in ds_cfg:
+ assert 'samples_per_gpu' not in \
+ cfg.data.test_dataloader, ('`samples_per_gpu` are set '
+ 'in `data.test` field and ` '
+ 'data.test_dataloader` at'
+ ' the same time. '
+ 'Please only set it in '
+ '`data.test_dataloader`. ')
+ samples_per_gpu = max(
+ [ds_cfg.pop('samples_per_gpu', 1) for ds_cfg in cfg.data.test])
+ cfg.data.test_dataloader['samples_per_gpu'] = samples_per_gpu
+
+ return cfg
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/contextmanagers.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/contextmanagers.py
new file mode 100644
index 0000000000000000000000000000000000000000..fa12bfcaff1e781b0a8cc7d7c8b839c2f2955a05
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/contextmanagers.py
@@ -0,0 +1,122 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import asyncio
+import contextlib
+import logging
+import os
+import time
+from typing import List
+
+import torch
+
+logger = logging.getLogger(__name__)
+
+DEBUG_COMPLETED_TIME = bool(os.environ.get('DEBUG_COMPLETED_TIME', False))
+
+
+@contextlib.asynccontextmanager
+async def completed(trace_name='',
+ name='',
+ sleep_interval=0.05,
+ streams: List[torch.cuda.Stream] = None):
+ """Async context manager that waits for work to complete on given CUDA
+ streams."""
+ if not torch.cuda.is_available():
+ yield
+ return
+
+ stream_before_context_switch = torch.cuda.current_stream()
+ if not streams:
+ streams = [stream_before_context_switch]
+ else:
+ streams = [s if s else stream_before_context_switch for s in streams]
+
+ end_events = [
+ torch.cuda.Event(enable_timing=DEBUG_COMPLETED_TIME) for _ in streams
+ ]
+
+ if DEBUG_COMPLETED_TIME:
+ start = torch.cuda.Event(enable_timing=True)
+ stream_before_context_switch.record_event(start)
+
+ cpu_start = time.monotonic()
+ logger.debug('%s %s starting, streams: %s', trace_name, name, streams)
+ grad_enabled_before = torch.is_grad_enabled()
+ try:
+ yield
+ finally:
+ current_stream = torch.cuda.current_stream()
+ assert current_stream == stream_before_context_switch
+
+ if DEBUG_COMPLETED_TIME:
+ cpu_end = time.monotonic()
+ for i, stream in enumerate(streams):
+ event = end_events[i]
+ stream.record_event(event)
+
+ grad_enabled_after = torch.is_grad_enabled()
+
+ # observed change of torch.is_grad_enabled() during concurrent run of
+ # async_test_bboxes code
+ assert (grad_enabled_before == grad_enabled_after
+ ), 'Unexpected is_grad_enabled() value change'
+
+ are_done = [e.query() for e in end_events]
+ logger.debug('%s %s completed: %s streams: %s', trace_name, name,
+ are_done, streams)
+ with torch.cuda.stream(stream_before_context_switch):
+ while not all(are_done):
+ await asyncio.sleep(sleep_interval)
+ are_done = [e.query() for e in end_events]
+ logger.debug(
+ '%s %s completed: %s streams: %s',
+ trace_name,
+ name,
+ are_done,
+ streams,
+ )
+
+ current_stream = torch.cuda.current_stream()
+ assert current_stream == stream_before_context_switch
+
+ if DEBUG_COMPLETED_TIME:
+ cpu_time = (cpu_end - cpu_start) * 1000
+ stream_times_ms = ''
+ for i, stream in enumerate(streams):
+ elapsed_time = start.elapsed_time(end_events[i])
+ stream_times_ms += f' {stream} {elapsed_time:.2f} ms'
+ logger.info('%s %s %.2f ms %s', trace_name, name, cpu_time,
+ stream_times_ms)
+
+
+@contextlib.asynccontextmanager
+async def concurrent(streamqueue: asyncio.Queue,
+ trace_name='concurrent',
+ name='stream'):
+ """Run code concurrently in different streams.
+
+ :param streamqueue: asyncio.Queue instance.
+
+ Queue tasks define the pool of streams used for concurrent execution.
+ """
+ if not torch.cuda.is_available():
+ yield
+ return
+
+ initial_stream = torch.cuda.current_stream()
+
+ with torch.cuda.stream(initial_stream):
+ stream = await streamqueue.get()
+ assert isinstance(stream, torch.cuda.Stream)
+
+ try:
+ with torch.cuda.stream(stream):
+ logger.debug('%s %s is starting, stream: %s', trace_name, name,
+ stream)
+ yield
+ current = torch.cuda.current_stream()
+ assert current == stream
+ logger.debug('%s %s has finished, stream: %s', trace_name,
+ name, stream)
+ finally:
+ streamqueue.task_done()
+ streamqueue.put_nowait(stream)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/dist_utils.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/dist_utils.py
new file mode 100644
index 0000000000000000000000000000000000000000..2f2c8614a181ec0594ba157002a2760737e2c6e3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/dist_utils.py
@@ -0,0 +1,184 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import functools
+import pickle
+import warnings
+from collections import OrderedDict
+
+import numpy as np
+import torch
+import torch.distributed as dist
+from mmengine.dist import get_dist_info
+from torch._utils import (_flatten_dense_tensors, _take_tensors,
+ _unflatten_dense_tensors)
+
+
+def _allreduce_coalesced(tensors, world_size, bucket_size_mb=-1):
+ if bucket_size_mb > 0:
+ bucket_size_bytes = bucket_size_mb * 1024 * 1024
+ buckets = _take_tensors(tensors, bucket_size_bytes)
+ else:
+ buckets = OrderedDict()
+ for tensor in tensors:
+ tp = tensor.type()
+ if tp not in buckets:
+ buckets[tp] = []
+ buckets[tp].append(tensor)
+ buckets = buckets.values()
+
+ for bucket in buckets:
+ flat_tensors = _flatten_dense_tensors(bucket)
+ dist.all_reduce(flat_tensors)
+ flat_tensors.div_(world_size)
+ for tensor, synced in zip(
+ bucket, _unflatten_dense_tensors(flat_tensors, bucket)):
+ tensor.copy_(synced)
+
+
+def allreduce_grads(params, coalesce=True, bucket_size_mb=-1):
+ """Allreduce gradients.
+
+ Args:
+ params (list[torch.Parameters]): List of parameters of a model
+ coalesce (bool, optional): Whether allreduce parameters as a whole.
+ Defaults to True.
+ bucket_size_mb (int, optional): Size of bucket, the unit is MB.
+ Defaults to -1.
+ """
+ grads = [
+ param.grad.data for param in params
+ if param.requires_grad and param.grad is not None
+ ]
+ world_size = dist.get_world_size()
+ if coalesce:
+ _allreduce_coalesced(grads, world_size, bucket_size_mb)
+ else:
+ for tensor in grads:
+ dist.all_reduce(tensor.div_(world_size))
+
+
+def reduce_mean(tensor):
+ """"Obtain the mean of tensor on different GPUs."""
+ if not (dist.is_available() and dist.is_initialized()):
+ return tensor
+ tensor = tensor.clone()
+ dist.all_reduce(tensor.div_(dist.get_world_size()), op=dist.ReduceOp.SUM)
+ return tensor
+
+
+def obj2tensor(pyobj, device='cuda'):
+ """Serialize picklable python object to tensor."""
+ storage = torch.ByteStorage.from_buffer(pickle.dumps(pyobj))
+ return torch.ByteTensor(storage).to(device=device)
+
+
+def tensor2obj(tensor):
+ """Deserialize tensor to picklable python object."""
+ return pickle.loads(tensor.cpu().numpy().tobytes())
+
+
+@functools.lru_cache()
+def _get_global_gloo_group():
+ """Return a process group based on gloo backend, containing all the ranks
+ The result is cached."""
+ if dist.get_backend() == 'nccl':
+ return dist.new_group(backend='gloo')
+ else:
+ return dist.group.WORLD
+
+
+def all_reduce_dict(py_dict, op='sum', group=None, to_float=True):
+ """Apply all reduce function for python dict object.
+
+ The code is modified from https://github.com/Megvii-
+ BaseDetection/YOLOX/blob/main/yolox/utils/allreduce_norm.py.
+
+ NOTE: make sure that py_dict in different ranks has the same keys and
+ the values should be in the same shape. Currently only supports
+ nccl backend.
+
+ Args:
+ py_dict (dict): Dict to be applied all reduce op.
+ op (str): Operator, could be 'sum' or 'mean'. Default: 'sum'
+ group (:obj:`torch.distributed.group`, optional): Distributed group,
+ Default: None.
+ to_float (bool): Whether to convert all values of dict to float.
+ Default: True.
+
+ Returns:
+ OrderedDict: reduced python dict object.
+ """
+ warnings.warn(
+ 'group` is deprecated. Currently only supports NCCL backend.')
+ _, world_size = get_dist_info()
+ if world_size == 1:
+ return py_dict
+
+ # all reduce logic across different devices.
+ py_key = list(py_dict.keys())
+ if not isinstance(py_dict, OrderedDict):
+ py_key_tensor = obj2tensor(py_key)
+ dist.broadcast(py_key_tensor, src=0)
+ py_key = tensor2obj(py_key_tensor)
+
+ tensor_shapes = [py_dict[k].shape for k in py_key]
+ tensor_numels = [py_dict[k].numel() for k in py_key]
+
+ if to_float:
+ warnings.warn('Note: the "to_float" is True, you need to '
+ 'ensure that the behavior is reasonable.')
+ flatten_tensor = torch.cat(
+ [py_dict[k].flatten().float() for k in py_key])
+ else:
+ flatten_tensor = torch.cat([py_dict[k].flatten() for k in py_key])
+
+ dist.all_reduce(flatten_tensor, op=dist.ReduceOp.SUM)
+ if op == 'mean':
+ flatten_tensor /= world_size
+
+ split_tensors = [
+ x.reshape(shape) for x, shape in zip(
+ torch.split(flatten_tensor, tensor_numels), tensor_shapes)
+ ]
+ out_dict = {k: v for k, v in zip(py_key, split_tensors)}
+ if isinstance(py_dict, OrderedDict):
+ out_dict = OrderedDict(out_dict)
+ return out_dict
+
+
+def sync_random_seed(seed=None, device='cuda'):
+ """Make sure different ranks share the same seed.
+
+ All workers must call this function, otherwise it will deadlock.
+ This method is generally used in `DistributedSampler`,
+ because the seed should be identical across all processes
+ in the distributed group.
+
+ In distributed sampling, different ranks should sample non-overlapped
+ data in the dataset. Therefore, this function is used to make sure that
+ each rank shuffles the data indices in the same order based
+ on the same seed. Then different ranks could use different indices
+ to select non-overlapped data from the same data list.
+
+ Args:
+ seed (int, Optional): The seed. Default to None.
+ device (str): The device where the seed will be put on.
+ Default to 'cuda'.
+
+ Returns:
+ int: Seed to be used.
+ """
+ if seed is None:
+ seed = np.random.randint(2**31)
+ assert isinstance(seed, int)
+
+ rank, world_size = get_dist_info()
+
+ if world_size == 1:
+ return seed
+
+ if rank == 0:
+ random_num = torch.tensor(seed, dtype=torch.int32, device=device)
+ else:
+ random_num = torch.tensor(0, dtype=torch.int32, device=device)
+ dist.broadcast(random_num, src=0)
+ return random_num.item()
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/large_image.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/large_image.py
new file mode 100644
index 0000000000000000000000000000000000000000..f1f07c2bdc6958f2b3bdd69da0a639276252a91e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/large_image.py
@@ -0,0 +1,104 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Sequence, Tuple
+
+import torch
+from mmcv.ops import batched_nms
+from mmengine.structures import InstanceData
+
+from mmdet.structures import DetDataSample, SampleList
+
+
+def shift_rbboxes(bboxes: torch.Tensor, offset: Sequence[int]):
+ """Shift rotated bboxes with offset.
+
+ Args:
+ bboxes (Tensor): The rotated bboxes need to be translated.
+ With shape (n, 5), which means (x, y, w, h, a).
+ offset (Sequence[int]): The translation offsets with shape of (2, ).
+ Returns:
+ Tensor: Shifted rotated bboxes.
+ """
+ offset_tensor = bboxes.new_tensor(offset)
+ shifted_bboxes = bboxes.clone()
+ shifted_bboxes[:, 0:2] = shifted_bboxes[:, 0:2] + offset_tensor
+ return shifted_bboxes
+
+
+def shift_predictions(det_data_samples: SampleList,
+ offsets: Sequence[Tuple[int, int]],
+ src_image_shape: Tuple[int, int]) -> SampleList:
+ """Shift predictions to the original image.
+
+ Args:
+ det_data_samples (List[:obj:`DetDataSample`]): A list of patch results.
+ offsets (Sequence[Tuple[int, int]]): Positions of the left top points
+ of patches.
+ src_image_shape (Tuple[int, int]): A (height, width) tuple of the large
+ image's width and height.
+ Returns:
+ (List[:obj:`DetDataSample`]): shifted results.
+ """
+ try:
+ from sahi.slicing import shift_bboxes, shift_masks
+ except ImportError:
+ raise ImportError('Please run "pip install -U sahi" '
+ 'to install sahi first for large image inference.')
+
+ assert len(det_data_samples) == len(
+ offsets), 'The `results` should has the ' 'same length with `offsets`.'
+ shifted_predictions = []
+ for det_data_sample, offset in zip(det_data_samples, offsets):
+ pred_inst = det_data_sample.pred_instances.clone()
+
+ # Check bbox type
+ if pred_inst.bboxes.size(-1) == 4:
+ # Horizontal bboxes
+ shifted_bboxes = shift_bboxes(pred_inst.bboxes, offset)
+ elif pred_inst.bboxes.size(-1) == 5:
+ # Rotated bboxes
+ shifted_bboxes = shift_rbboxes(pred_inst.bboxes, offset)
+ else:
+ raise NotImplementedError
+
+ # shift bboxes and masks
+ pred_inst.bboxes = shifted_bboxes
+ if 'masks' in det_data_sample:
+ pred_inst.masks = shift_masks(pred_inst.masks, offset,
+ src_image_shape)
+
+ shifted_predictions.append(pred_inst.clone())
+
+ shifted_predictions = InstanceData.cat(shifted_predictions)
+
+ return shifted_predictions
+
+
+def merge_results_by_nms(results: SampleList, offsets: Sequence[Tuple[int,
+ int]],
+ src_image_shape: Tuple[int, int],
+ nms_cfg: dict) -> DetDataSample:
+ """Merge patch results by nms.
+
+ Args:
+ results (List[:obj:`DetDataSample`]): A list of patch results.
+ offsets (Sequence[Tuple[int, int]]): Positions of the left top points
+ of patches.
+ src_image_shape (Tuple[int, int]): A (height, width) tuple of the large
+ image's width and height.
+ nms_cfg (dict): it should specify nms type and other parameters
+ like `iou_threshold`.
+ Returns:
+ :obj:`DetDataSample`: merged results.
+ """
+ shifted_instances = shift_predictions(results, offsets, src_image_shape)
+
+ _, keeps = batched_nms(
+ boxes=shifted_instances.bboxes,
+ scores=shifted_instances.scores,
+ idxs=shifted_instances.labels,
+ nms_cfg=nms_cfg)
+ merged_instances = shifted_instances[keeps]
+
+ merged_result = results[0].clone()
+ merged_result.pred_instances = merged_instances
+ return merged_result
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/logger.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/logger.py
new file mode 100644
index 0000000000000000000000000000000000000000..9fec08bbad5517c9169eedb15b4768e7d88d39c7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/logger.py
@@ -0,0 +1,49 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import inspect
+
+from mmengine.logging import print_log
+
+
+def get_caller_name():
+ """Get name of caller method."""
+ # this_func_frame = inspect.stack()[0][0] # i.e., get_caller_name
+ # callee_frame = inspect.stack()[1][0] # e.g., log_img_scale
+ caller_frame = inspect.stack()[2][0] # e.g., caller of log_img_scale
+ caller_method = caller_frame.f_code.co_name
+ try:
+ caller_class = caller_frame.f_locals['self'].__class__.__name__
+ return f'{caller_class}.{caller_method}'
+ except KeyError: # caller is a function
+ return caller_method
+
+
+def log_img_scale(img_scale, shape_order='hw', skip_square=False):
+ """Log image size.
+
+ Args:
+ img_scale (tuple): Image size to be logged.
+ shape_order (str, optional): The order of image shape.
+ 'hw' for (height, width) and 'wh' for (width, height).
+ Defaults to 'hw'.
+ skip_square (bool, optional): Whether to skip logging for square
+ img_scale. Defaults to False.
+
+ Returns:
+ bool: Whether to have done logging.
+ """
+ if shape_order == 'hw':
+ height, width = img_scale
+ elif shape_order == 'wh':
+ width, height = img_scale
+ else:
+ raise ValueError(f'Invalid shape_order {shape_order}.')
+
+ if skip_square and (height == width):
+ return False
+
+ caller = get_caller_name()
+ print_log(
+ f'image shape: height={height}, width={width} in {caller}',
+ logger='current')
+
+ return True
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/memory.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/memory.py
new file mode 100644
index 0000000000000000000000000000000000000000..b6f9cbc7f9e5f54e2cc429e5e655b2a27d38d61f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/memory.py
@@ -0,0 +1,212 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import warnings
+from collections import abc
+from contextlib import contextmanager
+from functools import wraps
+
+import torch
+from mmengine.logging import MMLogger
+
+
+def cast_tensor_type(inputs, src_type=None, dst_type=None):
+ """Recursively convert Tensor in inputs from ``src_type`` to ``dst_type``.
+
+ Args:
+ inputs: Inputs that to be casted.
+ src_type (torch.dtype | torch.device): Source type.
+ src_type (torch.dtype | torch.device): Destination type.
+
+ Returns:
+ The same type with inputs, but all contained Tensors have been cast.
+ """
+ assert dst_type is not None
+ if isinstance(inputs, torch.Tensor):
+ if isinstance(dst_type, torch.device):
+ # convert Tensor to dst_device
+ if hasattr(inputs, 'to') and \
+ hasattr(inputs, 'device') and \
+ (inputs.device == src_type or src_type is None):
+ return inputs.to(dst_type)
+ else:
+ return inputs
+ else:
+ # convert Tensor to dst_dtype
+ if hasattr(inputs, 'to') and \
+ hasattr(inputs, 'dtype') and \
+ (inputs.dtype == src_type or src_type is None):
+ return inputs.to(dst_type)
+ else:
+ return inputs
+ # we need to ensure that the type of inputs to be casted are the same
+ # as the argument `src_type`.
+ elif isinstance(inputs, abc.Mapping):
+ return type(inputs)({
+ k: cast_tensor_type(v, src_type=src_type, dst_type=dst_type)
+ for k, v in inputs.items()
+ })
+ elif isinstance(inputs, abc.Iterable):
+ return type(inputs)(
+ cast_tensor_type(item, src_type=src_type, dst_type=dst_type)
+ for item in inputs)
+ # TODO: Currently not supported
+ # elif isinstance(inputs, InstanceData):
+ # for key, value in inputs.items():
+ # inputs[key] = cast_tensor_type(
+ # value, src_type=src_type, dst_type=dst_type)
+ # return inputs
+ else:
+ return inputs
+
+
+@contextmanager
+def _ignore_torch_cuda_oom():
+ """A context which ignores CUDA OOM exception from pytorch.
+
+ Code is modified from
+ # noqa: E501
+ """
+ try:
+ yield
+ except RuntimeError as e:
+ # NOTE: the string may change?
+ if 'CUDA out of memory. ' in str(e):
+ pass
+ else:
+ raise
+
+
+class AvoidOOM:
+ """Try to convert inputs to FP16 and CPU if got a PyTorch's CUDA Out of
+ Memory error. It will do the following steps:
+
+ 1. First retry after calling `torch.cuda.empty_cache()`.
+ 2. If that still fails, it will then retry by converting inputs
+ to FP16.
+ 3. If that still fails trying to convert inputs to CPUs.
+ In this case, it expects the function to dispatch to
+ CPU implementation.
+
+ Args:
+ to_cpu (bool): Whether to convert outputs to CPU if get an OOM
+ error. This will slow down the code significantly.
+ Defaults to True.
+ test (bool): Skip `_ignore_torch_cuda_oom` operate that can use
+ lightweight data in unit test, only used in
+ test unit. Defaults to False.
+
+ Examples:
+ >>> from mmdet.utils.memory import AvoidOOM
+ >>> AvoidCUDAOOM = AvoidOOM()
+ >>> output = AvoidOOM.retry_if_cuda_oom(
+ >>> some_torch_function)(input1, input2)
+ >>> # To use as a decorator
+ >>> # from mmdet.utils import AvoidCUDAOOM
+ >>> @AvoidCUDAOOM.retry_if_cuda_oom
+ >>> def function(*args, **kwargs):
+ >>> return None
+ ```
+
+ Note:
+ 1. The output may be on CPU even if inputs are on GPU. Processing
+ on CPU will slow down the code significantly.
+ 2. When converting inputs to CPU, it will only look at each argument
+ and check if it has `.device` and `.to` for conversion. Nested
+ structures of tensors are not supported.
+ 3. Since the function might be called more than once, it has to be
+ stateless.
+ """
+
+ def __init__(self, to_cpu=True, test=False):
+ self.to_cpu = to_cpu
+ self.test = test
+
+ def retry_if_cuda_oom(self, func):
+ """Makes a function retry itself after encountering pytorch's CUDA OOM
+ error.
+
+ The implementation logic is referred to
+ https://github.com/facebookresearch/detectron2/blob/main/detectron2/utils/memory.py
+
+ Args:
+ func: a stateless callable that takes tensor-like objects
+ as arguments.
+ Returns:
+ func: a callable which retries `func` if OOM is encountered.
+ """ # noqa: W605
+
+ @wraps(func)
+ def wrapped(*args, **kwargs):
+
+ # raw function
+ if not self.test:
+ with _ignore_torch_cuda_oom():
+ return func(*args, **kwargs)
+
+ # Clear cache and retry
+ torch.cuda.empty_cache()
+ with _ignore_torch_cuda_oom():
+ return func(*args, **kwargs)
+
+ # get the type and device of first tensor
+ dtype, device = None, None
+ values = args + tuple(kwargs.values())
+ for value in values:
+ if isinstance(value, torch.Tensor):
+ dtype = value.dtype
+ device = value.device
+ break
+ if dtype is None or device is None:
+ raise ValueError('There is no tensor in the inputs, '
+ 'cannot get dtype and device.')
+
+ # Convert to FP16
+ fp16_args = cast_tensor_type(args, dst_type=torch.half)
+ fp16_kwargs = cast_tensor_type(kwargs, dst_type=torch.half)
+ logger = MMLogger.get_current_instance()
+ logger.warning(f'Attempting to copy inputs of {str(func)} '
+ 'to FP16 due to CUDA OOM')
+
+ # get input tensor type, the output type will same as
+ # the first parameter type.
+ with _ignore_torch_cuda_oom():
+ output = func(*fp16_args, **fp16_kwargs)
+ output = cast_tensor_type(
+ output, src_type=torch.half, dst_type=dtype)
+ if not self.test:
+ return output
+ logger.warning('Using FP16 still meet CUDA OOM')
+
+ # Try on CPU. This will slow down the code significantly,
+ # therefore print a notice.
+ if self.to_cpu:
+ logger.warning(f'Attempting to copy inputs of {str(func)} '
+ 'to CPU due to CUDA OOM')
+ cpu_device = torch.empty(0).device
+ cpu_args = cast_tensor_type(args, dst_type=cpu_device)
+ cpu_kwargs = cast_tensor_type(kwargs, dst_type=cpu_device)
+
+ # convert outputs to GPU
+ with _ignore_torch_cuda_oom():
+ logger.warning(f'Convert outputs to GPU (device={device})')
+ output = func(*cpu_args, **cpu_kwargs)
+ output = cast_tensor_type(
+ output, src_type=cpu_device, dst_type=device)
+ return output
+
+ warnings.warn('Cannot convert output to GPU due to CUDA OOM, '
+ 'the output is now on CPU, which might cause '
+ 'errors if the output need to interact with GPU '
+ 'data in subsequent operations')
+ logger.warning('Cannot convert output to GPU due to '
+ 'CUDA OOM, the output is on CPU now.')
+
+ return func(*cpu_args, **cpu_kwargs)
+ else:
+ # may still get CUDA OOM error
+ return func(*args, **kwargs)
+
+ return wrapped
+
+
+# To use AvoidOOM as a decorator
+AvoidCUDAOOM = AvoidOOM()
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/misc.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/misc.py
new file mode 100644
index 0000000000000000000000000000000000000000..8dfb394465196cbd1e60c96f5be3aaee416d59cf
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/misc.py
@@ -0,0 +1,149 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import glob
+import os
+import os.path as osp
+import urllib
+import warnings
+from typing import Union
+
+import torch
+from mmengine.config import Config, ConfigDict
+from mmengine.logging import print_log
+from mmengine.utils import scandir
+
+IMG_EXTENSIONS = ('.jpg', '.jpeg', '.png', '.ppm', '.bmp', '.pgm', '.tif',
+ '.tiff', '.webp')
+
+
+def find_latest_checkpoint(path, suffix='pth'):
+ """Find the latest checkpoint from the working directory.
+
+ Args:
+ path(str): The path to find checkpoints.
+ suffix(str): File extension.
+ Defaults to pth.
+
+ Returns:
+ latest_path(str | None): File path of the latest checkpoint.
+ References:
+ .. [1] https://github.com/microsoft/SoftTeacher
+ /blob/main/ssod/utils/patch.py
+ """
+ if not osp.exists(path):
+ warnings.warn('The path of checkpoints does not exist.')
+ return None
+ if osp.exists(osp.join(path, f'latest.{suffix}')):
+ return osp.join(path, f'latest.{suffix}')
+
+ checkpoints = glob.glob(osp.join(path, f'*.{suffix}'))
+ if len(checkpoints) == 0:
+ warnings.warn('There are no checkpoints in the path.')
+ return None
+ latest = -1
+ latest_path = None
+ for checkpoint in checkpoints:
+ count = int(osp.basename(checkpoint).split('_')[-1].split('.')[0])
+ if count > latest:
+ latest = count
+ latest_path = checkpoint
+ return latest_path
+
+
+def update_data_root(cfg, logger=None):
+ """Update data root according to env MMDET_DATASETS.
+
+ If set env MMDET_DATASETS, update cfg.data_root according to
+ MMDET_DATASETS. Otherwise, using cfg.data_root as default.
+
+ Args:
+ cfg (:obj:`Config`): The model config need to modify
+ logger (logging.Logger | str | None): the way to print msg
+ """
+ assert isinstance(cfg, Config), \
+ f'cfg got wrong type: {type(cfg)}, expected mmengine.Config'
+
+ if 'MMDET_DATASETS' in os.environ:
+ dst_root = os.environ['MMDET_DATASETS']
+ print_log(f'MMDET_DATASETS has been set to be {dst_root}.'
+ f'Using {dst_root} as data root.')
+ else:
+ return
+
+ assert isinstance(cfg, Config), \
+ f'cfg got wrong type: {type(cfg)}, expected mmengine.Config'
+
+ def update(cfg, src_str, dst_str):
+ for k, v in cfg.items():
+ if isinstance(v, ConfigDict):
+ update(cfg[k], src_str, dst_str)
+ if isinstance(v, str) and src_str in v:
+ cfg[k] = v.replace(src_str, dst_str)
+
+ update(cfg.data, cfg.data_root, dst_root)
+ cfg.data_root = dst_root
+
+
+def get_test_pipeline_cfg(cfg: Union[str, ConfigDict]) -> ConfigDict:
+ """Get the test dataset pipeline from entire config.
+
+ Args:
+ cfg (str or :obj:`ConfigDict`): the entire config. Can be a config
+ file or a ``ConfigDict``.
+
+ Returns:
+ :obj:`ConfigDict`: the config of test dataset.
+ """
+ if isinstance(cfg, str):
+ cfg = Config.fromfile(cfg)
+
+ def _get_test_pipeline_cfg(dataset_cfg):
+ if 'pipeline' in dataset_cfg:
+ return dataset_cfg.pipeline
+ # handle dataset wrapper
+ elif 'dataset' in dataset_cfg:
+ return _get_test_pipeline_cfg(dataset_cfg.dataset)
+ # handle dataset wrappers like ConcatDataset
+ elif 'datasets' in dataset_cfg:
+ return _get_test_pipeline_cfg(dataset_cfg.datasets[0])
+
+ raise RuntimeError('Cannot find `pipeline` in `test_dataloader`')
+
+ return _get_test_pipeline_cfg(cfg.test_dataloader.dataset)
+
+
+def get_file_list(source_root: str) -> [list, dict]:
+ """Get file list.
+
+ Args:
+ source_root (str): image or video source path
+
+ Return:
+ source_file_path_list (list): A list for all source file.
+ source_type (dict): Source type: file or url or dir.
+ """
+ is_dir = os.path.isdir(source_root)
+ is_url = source_root.startswith(('http:/', 'https:/'))
+ is_file = os.path.splitext(source_root)[-1].lower() in IMG_EXTENSIONS
+
+ source_file_path_list = []
+ if is_dir:
+ # when input source is dir
+ for file in scandir(source_root, IMG_EXTENSIONS, recursive=True):
+ source_file_path_list.append(os.path.join(source_root, file))
+ elif is_url:
+ # when input source is url
+ filename = os.path.basename(
+ urllib.parse.unquote(source_root).split('?')[0])
+ file_save_path = os.path.join(os.getcwd(), filename)
+ print(f'Downloading source file to {file_save_path}')
+ torch.hub.download_url_to_file(source_root, file_save_path)
+ source_file_path_list = [file_save_path]
+ elif is_file:
+ # when input source is single image
+ source_file_path_list = [source_root]
+ else:
+ print('Cannot find image file.')
+
+ source_type = dict(is_dir=is_dir, is_url=is_url, is_file=is_file)
+
+ return source_file_path_list, source_type
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/mot_error_visualize.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/mot_error_visualize.py
new file mode 100644
index 0000000000000000000000000000000000000000..01bf8645d340aa1f5ab8251211a719f2de9845b1
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/mot_error_visualize.py
@@ -0,0 +1,273 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import os.path as osp
+from typing import Union
+
+try:
+ import seaborn as sns
+except ImportError:
+ sns = None
+import cv2
+import matplotlib.pyplot as plt
+import mmcv
+import numpy as np
+from matplotlib.patches import Rectangle
+from mmengine.utils import mkdir_or_exist
+
+
+def imshow_mot_errors(*args, backend: str = 'cv2', **kwargs):
+ """Show the wrong tracks on the input image.
+
+ Args:
+ backend (str, optional): Backend of visualization.
+ Defaults to 'cv2'.
+ """
+ if backend == 'cv2':
+ return _cv2_show_wrong_tracks(*args, **kwargs)
+ elif backend == 'plt':
+ return _plt_show_wrong_tracks(*args, **kwargs)
+ else:
+ raise NotImplementedError()
+
+
+def _cv2_show_wrong_tracks(img: Union[str, np.ndarray],
+ bboxes: np.ndarray,
+ ids: np.ndarray,
+ error_types: np.ndarray,
+ thickness: int = 2,
+ font_scale: float = 0.4,
+ text_width: int = 10,
+ text_height: int = 15,
+ show: bool = False,
+ wait_time: int = 100,
+ out_file: str = None) -> np.ndarray:
+ """Show the wrong tracks with opencv.
+
+ Args:
+ img (str or ndarray): The image to be displayed.
+ bboxes (ndarray): A ndarray of shape (k, 5).
+ ids (ndarray): A ndarray of shape (k, ).
+ error_types (ndarray): A ndarray of shape (k, ), where 0 denotes
+ false positives, 1 denotes false negative and 2 denotes ID switch.
+ thickness (int, optional): Thickness of lines.
+ Defaults to 2.
+ font_scale (float, optional): Font scale to draw id and score.
+ Defaults to 0.4.
+ text_width (int, optional): Width to draw id and score.
+ Defaults to 10.
+ text_height (int, optional): Height to draw id and score.
+ Defaults to 15.
+ show (bool, optional): Whether to show the image on the fly.
+ Defaults to False.
+ wait_time (int, optional): Value of waitKey param.
+ Defaults to 100.
+ out_file (str, optional): The filename to write the image.
+ Defaults to None.
+
+ Returns:
+ ndarray: Visualized image.
+ """
+ if sns is None:
+ raise ImportError('please run pip install seaborn')
+ assert bboxes.ndim == 2, \
+ f' bboxes ndim should be 2, but its ndim is {bboxes.ndim}.'
+ assert ids.ndim == 1, \
+ f' ids ndim should be 1, but its ndim is {ids.ndim}.'
+ assert error_types.ndim == 1, \
+ f' error_types ndim should be 1, but its ndim is {error_types.ndim}.'
+ assert bboxes.shape[0] == ids.shape[0], \
+ 'bboxes.shape[0] and ids.shape[0] should have the same length.'
+ assert bboxes.shape[1] == 5, \
+ f' bboxes.shape[1] should be 5, but its {bboxes.shape[1]}.'
+
+ bbox_colors = sns.color_palette()
+ # red, yellow, blue
+ bbox_colors = [bbox_colors[3], bbox_colors[1], bbox_colors[0]]
+ bbox_colors = [[int(255 * _c) for _c in bbox_color][::-1]
+ for bbox_color in bbox_colors]
+
+ if isinstance(img, str):
+ img = mmcv.imread(img)
+ else:
+ assert img.ndim == 3
+
+ img_shape = img.shape
+ bboxes[:, 0::2] = np.clip(bboxes[:, 0::2], 0, img_shape[1])
+ bboxes[:, 1::2] = np.clip(bboxes[:, 1::2], 0, img_shape[0])
+
+ for bbox, error_type, id in zip(bboxes, error_types, ids):
+ x1, y1, x2, y2 = bbox[:4].astype(np.int32)
+ score = float(bbox[-1])
+
+ # bbox
+ bbox_color = bbox_colors[error_type]
+ cv2.rectangle(img, (x1, y1), (x2, y2), bbox_color, thickness=thickness)
+
+ # FN does not have id and score
+ if error_type == 1:
+ continue
+
+ # score
+ text = '{:.02f}'.format(score)
+ width = (len(text) - 1) * text_width
+ img[y1:y1 + text_height, x1:x1 + width, :] = bbox_color
+ cv2.putText(
+ img,
+ text, (x1, y1 + text_height - 2),
+ cv2.FONT_HERSHEY_COMPLEX,
+ font_scale,
+ color=(0, 0, 0))
+
+ # id
+ text = str(id)
+ width = len(text) * text_width
+ img[y1 + text_height:y1 + text_height * 2,
+ x1:x1 + width, :] = bbox_color
+ cv2.putText(
+ img,
+ str(id), (x1, y1 + text_height * 2 - 2),
+ cv2.FONT_HERSHEY_COMPLEX,
+ font_scale,
+ color=(0, 0, 0))
+
+ if show:
+ mmcv.imshow(img, wait_time=wait_time)
+ if out_file is not None:
+ mmcv.imwrite(img, out_file)
+
+ return img
+
+
+def _plt_show_wrong_tracks(img: Union[str, np.ndarray],
+ bboxes: np.ndarray,
+ ids: np.ndarray,
+ error_types: np.ndarray,
+ thickness: float = 0.1,
+ font_scale: float = 3.0,
+ text_width: int = 8,
+ text_height: int = 13,
+ show: bool = False,
+ wait_time: int = 100,
+ out_file: str = None) -> np.ndarray:
+ """Show the wrong tracks with matplotlib.
+
+ Args:
+ img (str or ndarray): The image to be displayed.
+ bboxes (ndarray): A ndarray of shape (k, 5).
+ ids (ndarray): A ndarray of shape (k, ).
+ error_types (ndarray): A ndarray of shape (k, ), where 0 denotes
+ false positives, 1 denotes false negative and 2 denotes ID switch.
+ thickness (float, optional): Thickness of lines.
+ Defaults to 0.1.
+ font_scale (float, optional): Font scale to draw id and score.
+ Defaults to 3.0.
+ text_width (int, optional): Width to draw id and score.
+ Defaults to 8.
+ text_height (int, optional): Height to draw id and score.
+ Defaults to 13.
+ show (bool, optional): Whether to show the image on the fly.
+ Defaults to False.
+ wait_time (int, optional): Value of waitKey param.
+ Defaults to 100.
+ out_file (str, optional): The filename to write the image.
+ Defaults to None.
+
+ Returns:
+ ndarray: Original image.
+ """
+ assert bboxes.ndim == 2, \
+ f' bboxes ndim should be 2, but its ndim is {bboxes.ndim}.'
+ assert ids.ndim == 1, \
+ f' ids ndim should be 1, but its ndim is {ids.ndim}.'
+ assert error_types.ndim == 1, \
+ f' error_types ndim should be 1, but its ndim is {error_types.ndim}.'
+ assert bboxes.shape[0] == ids.shape[0], \
+ 'bboxes.shape[0] and ids.shape[0] should have the same length.'
+ assert bboxes.shape[1] == 5, \
+ f' bboxes.shape[1] should be 5, but its {bboxes.shape[1]}.'
+
+ bbox_colors = sns.color_palette()
+ # red, yellow, blue
+ bbox_colors = [bbox_colors[3], bbox_colors[1], bbox_colors[0]]
+
+ if isinstance(img, str):
+ img = plt.imread(img)
+ else:
+ assert img.ndim == 3
+ img = mmcv.bgr2rgb(img)
+
+ img_shape = img.shape
+ bboxes[:, 0::2] = np.clip(bboxes[:, 0::2], 0, img_shape[1])
+ bboxes[:, 1::2] = np.clip(bboxes[:, 1::2], 0, img_shape[0])
+
+ plt.imshow(img)
+ plt.gca().set_axis_off()
+ plt.autoscale(False)
+ plt.subplots_adjust(
+ top=1, bottom=0, right=1, left=0, hspace=None, wspace=None)
+ plt.margins(0, 0)
+ plt.gca().xaxis.set_major_locator(plt.NullLocator())
+ plt.gca().yaxis.set_major_locator(plt.NullLocator())
+ plt.rcParams['figure.figsize'] = img_shape[1], img_shape[0]
+
+ for bbox, error_type, id in zip(bboxes, error_types, ids):
+ x1, y1, x2, y2, score = bbox
+ w, h = int(x2 - x1), int(y2 - y1)
+ left_top = (int(x1), int(y1))
+
+ # bbox
+ plt.gca().add_patch(
+ Rectangle(
+ left_top,
+ w,
+ h,
+ thickness,
+ edgecolor=bbox_colors[error_type],
+ facecolor='none'))
+
+ # FN does not have id and score
+ if error_type == 1:
+ continue
+
+ # score
+ text = '{:.02f}'.format(score)
+ width = len(text) * text_width
+ plt.gca().add_patch(
+ Rectangle((left_top[0], left_top[1]),
+ width,
+ text_height,
+ thickness,
+ edgecolor=bbox_colors[error_type],
+ facecolor=bbox_colors[error_type]))
+
+ plt.text(
+ left_top[0],
+ left_top[1] + text_height + 2,
+ text,
+ fontsize=font_scale)
+
+ # id
+ text = str(id)
+ width = len(text) * text_width
+ plt.gca().add_patch(
+ Rectangle((left_top[0], left_top[1] + text_height + 1),
+ width,
+ text_height,
+ thickness,
+ edgecolor=bbox_colors[error_type],
+ facecolor=bbox_colors[error_type]))
+ plt.text(
+ left_top[0],
+ left_top[1] + 2 * (text_height + 1),
+ text,
+ fontsize=font_scale)
+
+ if out_file is not None:
+ mkdir_or_exist(osp.abspath(osp.dirname(out_file)))
+ plt.savefig(out_file, dpi=300, bbox_inches='tight', pad_inches=0.0)
+
+ if show:
+ plt.draw()
+ plt.pause(wait_time / 1000.)
+
+ plt.clf()
+ return img
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/profiling.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/profiling.py
new file mode 100644
index 0000000000000000000000000000000000000000..2f53f456c72db57bfa69a8d022c92d153580209e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/profiling.py
@@ -0,0 +1,40 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import contextlib
+import sys
+import time
+
+import torch
+
+if sys.version_info >= (3, 7):
+
+ @contextlib.contextmanager
+ def profile_time(trace_name,
+ name,
+ enabled=True,
+ stream=None,
+ end_stream=None):
+ """Print time spent by CPU and GPU.
+
+ Useful as a temporary context manager to find sweet spots of code
+ suitable for async implementation.
+ """
+ if (not enabled) or not torch.cuda.is_available():
+ yield
+ return
+ stream = stream if stream else torch.cuda.current_stream()
+ end_stream = end_stream if end_stream else stream
+ start = torch.cuda.Event(enable_timing=True)
+ end = torch.cuda.Event(enable_timing=True)
+ stream.record_event(start)
+ try:
+ cpu_start = time.monotonic()
+ yield
+ finally:
+ cpu_end = time.monotonic()
+ end_stream.record_event(end)
+ end.synchronize()
+ cpu_time = (cpu_end - cpu_start) * 1000
+ gpu_time = start.elapsed_time(end)
+ msg = f'{trace_name} {name} cpu_time {cpu_time:.2f} ms '
+ msg += f'gpu_time {gpu_time:.2f} ms stream {stream}'
+ print(msg, end_stream)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/replace_cfg_vals.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/replace_cfg_vals.py
new file mode 100644
index 0000000000000000000000000000000000000000..a3331a36ce5a22fcc4d4a955d757f5e8b6bfc6bb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/replace_cfg_vals.py
@@ -0,0 +1,70 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import re
+
+from mmengine.config import Config
+
+
+def replace_cfg_vals(ori_cfg):
+ """Replace the string "${key}" with the corresponding value.
+
+ Replace the "${key}" with the value of ori_cfg.key in the config. And
+ support replacing the chained ${key}. Such as, replace "${key0.key1}"
+ with the value of cfg.key0.key1. Code is modified from `vars.py
+ < https://github.com/microsoft/SoftTeacher/blob/main/ssod/utils/vars.py>`_ # noqa: E501
+
+ Args:
+ ori_cfg (mmengine.config.Config):
+ The origin config with "${key}" generated from a file.
+
+ Returns:
+ updated_cfg [mmengine.config.Config]:
+ The config with "${key}" replaced by the corresponding value.
+ """
+
+ def get_value(cfg, key):
+ for k in key.split('.'):
+ cfg = cfg[k]
+ return cfg
+
+ def replace_value(cfg):
+ if isinstance(cfg, dict):
+ return {key: replace_value(value) for key, value in cfg.items()}
+ elif isinstance(cfg, list):
+ return [replace_value(item) for item in cfg]
+ elif isinstance(cfg, tuple):
+ return tuple([replace_value(item) for item in cfg])
+ elif isinstance(cfg, str):
+ # the format of string cfg may be:
+ # 1) "${key}", which will be replaced with cfg.key directly
+ # 2) "xxx${key}xxx" or "xxx${key1}xxx${key2}xxx",
+ # which will be replaced with the string of the cfg.key
+ keys = pattern_key.findall(cfg)
+ values = [get_value(ori_cfg, key[2:-1]) for key in keys]
+ if len(keys) == 1 and keys[0] == cfg:
+ # the format of string cfg is "${key}"
+ cfg = values[0]
+ else:
+ for key, value in zip(keys, values):
+ # the format of string cfg is
+ # "xxx${key}xxx" or "xxx${key1}xxx${key2}xxx"
+ assert not isinstance(value, (dict, list, tuple)), \
+ f'for the format of string cfg is ' \
+ f"'xxxxx${key}xxxxx' or 'xxx${key}xxx${key}xxx', " \
+ f"the type of the value of '${key}' " \
+ f'can not be dict, list, or tuple' \
+ f'but you input {type(value)} in {cfg}'
+ cfg = cfg.replace(key, str(value))
+ return cfg
+ else:
+ return cfg
+
+ # the pattern of string "${key}"
+ pattern_key = re.compile(r'\$\{[a-zA-Z\d_.]*\}')
+ # the type of ori_cfg._cfg_dict is mmengine.config.ConfigDict
+ updated_cfg = Config(
+ replace_value(ori_cfg._cfg_dict), filename=ori_cfg.filename)
+ # replace the model with model_wrapper
+ if updated_cfg.get('model_wrapper', None) is not None:
+ updated_cfg.model = updated_cfg.model_wrapper
+ updated_cfg.pop('model_wrapper')
+ return updated_cfg
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/setup_env.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/setup_env.py
new file mode 100644
index 0000000000000000000000000000000000000000..a7b37845a883752a1659fabf62c7404cff971191
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/setup_env.py
@@ -0,0 +1,118 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import datetime
+import logging
+import os
+import platform
+import warnings
+
+import cv2
+import torch.multiprocessing as mp
+from mmengine import DefaultScope
+from mmengine.logging import print_log
+from mmengine.utils import digit_version
+
+
+def setup_cache_size_limit_of_dynamo():
+ """Setup cache size limit of dynamo.
+
+ Note: Due to the dynamic shape of the loss calculation and
+ post-processing parts in the object detection algorithm, these
+ functions must be compiled every time they are run.
+ Setting a large value for torch._dynamo.config.cache_size_limit
+ may result in repeated compilation, which can slow down training
+ and testing speed. Therefore, we need to set the default value of
+ cache_size_limit smaller. An empirical value is 4.
+ """
+
+ import torch
+ if digit_version(torch.__version__) >= digit_version('2.0.0'):
+ if 'DYNAMO_CACHE_SIZE_LIMIT' in os.environ:
+ import torch._dynamo
+ cache_size_limit = int(os.environ['DYNAMO_CACHE_SIZE_LIMIT'])
+ torch._dynamo.config.cache_size_limit = cache_size_limit
+ print_log(
+ f'torch._dynamo.config.cache_size_limit is force '
+ f'set to {cache_size_limit}.',
+ logger='current',
+ level=logging.WARNING)
+
+
+def setup_multi_processes(cfg):
+ """Setup multi-processing environment variables."""
+ # set multi-process start method as `fork` to speed up the training
+ if platform.system() != 'Windows':
+ mp_start_method = cfg.get('mp_start_method', 'fork')
+ current_method = mp.get_start_method(allow_none=True)
+ if current_method is not None and current_method != mp_start_method:
+ warnings.warn(
+ f'Multi-processing start method `{mp_start_method}` is '
+ f'different from the previous setting `{current_method}`.'
+ f'It will be force set to `{mp_start_method}`. You can change '
+ f'this behavior by changing `mp_start_method` in your config.')
+ mp.set_start_method(mp_start_method, force=True)
+
+ # disable opencv multithreading to avoid system being overloaded
+ opencv_num_threads = cfg.get('opencv_num_threads', 0)
+ cv2.setNumThreads(opencv_num_threads)
+
+ # setup OMP threads
+ # This code is referred from https://github.com/pytorch/pytorch/blob/master/torch/distributed/run.py # noqa
+ workers_per_gpu = cfg.data.get('workers_per_gpu', 1)
+ if 'train_dataloader' in cfg.data:
+ workers_per_gpu = \
+ max(cfg.data.train_dataloader.get('workers_per_gpu', 1),
+ workers_per_gpu)
+
+ if 'OMP_NUM_THREADS' not in os.environ and workers_per_gpu > 1:
+ omp_num_threads = 1
+ warnings.warn(
+ f'Setting OMP_NUM_THREADS environment variable for each process '
+ f'to be {omp_num_threads} in default, to avoid your system being '
+ f'overloaded, please further tune the variable for optimal '
+ f'performance in your application as needed.')
+ os.environ['OMP_NUM_THREADS'] = str(omp_num_threads)
+
+ # setup MKL threads
+ if 'MKL_NUM_THREADS' not in os.environ and workers_per_gpu > 1:
+ mkl_num_threads = 1
+ warnings.warn(
+ f'Setting MKL_NUM_THREADS environment variable for each process '
+ f'to be {mkl_num_threads} in default, to avoid your system being '
+ f'overloaded, please further tune the variable for optimal '
+ f'performance in your application as needed.')
+ os.environ['MKL_NUM_THREADS'] = str(mkl_num_threads)
+
+
+def register_all_modules(init_default_scope: bool = True) -> None:
+ """Register all modules in mmdet into the registries.
+
+ Args:
+ init_default_scope (bool): Whether initialize the mmdet default scope.
+ When `init_default_scope=True`, the global default scope will be
+ set to `mmdet`, and all registries will build modules from mmdet's
+ registry node. To understand more about the registry, please refer
+ to https://github.com/open-mmlab/mmengine/blob/main/docs/en/tutorials/registry.md
+ Defaults to True.
+ """ # noqa
+ import mmdet.datasets # noqa: F401,F403
+ import mmdet.engine # noqa: F401,F403
+ import mmdet.evaluation # noqa: F401,F403
+ import mmdet.models # noqa: F401,F403
+ import mmdet.visualization # noqa: F401,F403
+
+ if init_default_scope:
+ never_created = DefaultScope.get_current_instance() is None \
+ or not DefaultScope.check_instance_created('mmdet')
+ if never_created:
+ DefaultScope.get_instance('mmdet', scope_name='mmdet')
+ return
+ current_scope = DefaultScope.get_current_instance()
+ if current_scope.scope_name != 'mmdet':
+ warnings.warn('The current default scope '
+ f'"{current_scope.scope_name}" is not "mmdet", '
+ '`register_all_modules` will force the current'
+ 'default scope to be "mmdet". If this is not '
+ 'expected, please set `init_default_scope=False`.')
+ # avoid name conflict
+ new_instance_name = f'mmdet-{datetime.datetime.now()}'
+ DefaultScope.get_instance(new_instance_name, scope_name='mmdet')
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/split_batch.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/split_batch.py
new file mode 100644
index 0000000000000000000000000000000000000000..0276fb331f23c1a7f7451faf2a8f768e616d45fd
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/split_batch.py
@@ -0,0 +1,45 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import torch
+
+
+def split_batch(img, img_metas, kwargs):
+ """Split data_batch by tags.
+
+ Code is modified from
+ # noqa: E501
+
+ Args:
+ img (Tensor): of shape (N, C, H, W) encoding input images.
+ Typically these should be mean centered and std scaled.
+ img_metas (list[dict]): List of image info dict where each dict
+ has: 'img_shape', 'scale_factor', 'flip', and may also contain
+ 'filename', 'ori_shape', 'pad_shape', and 'img_norm_cfg'.
+ For details on the values of these keys, see
+ :class:`mmdet.datasets.pipelines.Collect`.
+ kwargs (dict): Specific to concrete implementation.
+
+ Returns:
+ data_groups (dict): a dict that data_batch splited by tags,
+ such as 'sup', 'unsup_teacher', and 'unsup_student'.
+ """
+
+ # only stack img in the batch
+ def fuse_list(obj_list, obj):
+ return torch.stack(obj_list) if isinstance(obj,
+ torch.Tensor) else obj_list
+
+ # select data with tag from data_batch
+ def select_group(data_batch, current_tag):
+ group_flag = [tag == current_tag for tag in data_batch['tag']]
+ return {
+ k: fuse_list([vv for vv, gf in zip(v, group_flag) if gf], v)
+ for k, v in data_batch.items()
+ }
+
+ kwargs.update({'img': img, 'img_metas': img_metas})
+ kwargs.update({'tag': [meta['tag'] for meta in img_metas]})
+ tags = list(set(kwargs['tag']))
+ data_groups = {tag: select_group(kwargs, tag) for tag in tags}
+ for tag, group in data_groups.items():
+ group.pop('tag')
+ return data_groups
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/typing_utils.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/typing_utils.py
new file mode 100644
index 0000000000000000000000000000000000000000..6caf6de53274594e139dbe7c1973c747229bf010
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/typing_utils.py
@@ -0,0 +1,22 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+"""Collecting some commonly used type hint in mmdetection."""
+from typing import List, Optional, Sequence, Tuple, Union
+
+from mmengine.config import ConfigDict
+from mmengine.structures import InstanceData, PixelData
+
+# TODO: Need to avoid circular import with assigner and sampler
+# Type hint of config data
+ConfigType = Union[ConfigDict, dict]
+OptConfigType = Optional[ConfigType]
+# Type hint of one or more config data
+MultiConfig = Union[ConfigType, List[ConfigType]]
+OptMultiConfig = Optional[MultiConfig]
+
+InstanceList = List[InstanceData]
+OptInstanceList = Optional[InstanceList]
+
+PixelList = List[PixelData]
+OptPixelList = Optional[PixelList]
+
+RangeType = Sequence[Tuple[int, int]]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/util_mixins.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/util_mixins.py
new file mode 100644
index 0000000000000000000000000000000000000000..b83b6617f5e4a202067e1659bf448962a2a2bc72
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/util_mixins.py
@@ -0,0 +1,105 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+"""This module defines the :class:`NiceRepr` mixin class, which defines a
+``__repr__`` and ``__str__`` method that only depend on a custom ``__nice__``
+method, which you must define. This means you only have to overload one
+function instead of two. Furthermore, if the object defines a ``__len__``
+method, then the ``__nice__`` method defaults to something sensible, otherwise
+it is treated as abstract and raises ``NotImplementedError``.
+
+To use simply have your object inherit from :class:`NiceRepr`
+(multi-inheritance should be ok).
+
+This code was copied from the ubelt library: https://github.com/Erotemic/ubelt
+
+Example:
+ >>> # Objects that define __nice__ have a default __str__ and __repr__
+ >>> class Student(NiceRepr):
+ ... def __init__(self, name):
+ ... self.name = name
+ ... def __nice__(self):
+ ... return self.name
+ >>> s1 = Student('Alice')
+ >>> s2 = Student('Bob')
+ >>> print(f's1 = {s1}')
+ >>> print(f's2 = {s2}')
+ s1 =
+ s2 =
+
+Example:
+ >>> # Objects that define __len__ have a default __nice__
+ >>> class Group(NiceRepr):
+ ... def __init__(self, data):
+ ... self.data = data
+ ... def __len__(self):
+ ... return len(self.data)
+ >>> g = Group([1, 2, 3])
+ >>> print(f'g = {g}')
+ g =
+"""
+import warnings
+
+
+class NiceRepr:
+ """Inherit from this class and define ``__nice__`` to "nicely" print your
+ objects.
+
+ Defines ``__str__`` and ``__repr__`` in terms of ``__nice__`` function
+ Classes that inherit from :class:`NiceRepr` should redefine ``__nice__``.
+ If the inheriting class has a ``__len__``, method then the default
+ ``__nice__`` method will return its length.
+
+ Example:
+ >>> class Foo(NiceRepr):
+ ... def __nice__(self):
+ ... return 'info'
+ >>> foo = Foo()
+ >>> assert str(foo) == ''
+ >>> assert repr(foo).startswith('>> class Bar(NiceRepr):
+ ... pass
+ >>> bar = Bar()
+ >>> import pytest
+ >>> with pytest.warns(None) as record:
+ >>> assert 'object at' in str(bar)
+ >>> assert 'object at' in repr(bar)
+
+ Example:
+ >>> class Baz(NiceRepr):
+ ... def __len__(self):
+ ... return 5
+ >>> baz = Baz()
+ >>> assert str(baz) == ''
+ """
+
+ def __nice__(self):
+ """str: a "nice" summary string describing this module"""
+ if hasattr(self, '__len__'):
+ # It is a common pattern for objects to use __len__ in __nice__
+ # As a convenience we define a default __nice__ for these objects
+ return str(len(self))
+ else:
+ # In all other cases force the subclass to overload __nice__
+ raise NotImplementedError(
+ f'Define the __nice__ method for {self.__class__!r}')
+
+ def __repr__(self):
+ """str: the string of the module"""
+ try:
+ nice = self.__nice__()
+ classname = self.__class__.__name__
+ return f'<{classname}({nice}) at {hex(id(self))}>'
+ except NotImplementedError as ex:
+ warnings.warn(str(ex), category=RuntimeWarning)
+ return object.__repr__(self)
+
+ def __str__(self):
+ """str: the string of the module"""
+ try:
+ classname = self.__class__.__name__
+ nice = self.__nice__()
+ return f'<{classname}({nice})>'
+ except NotImplementedError as ex:
+ warnings.warn(str(ex), category=RuntimeWarning)
+ return object.__repr__(self)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/util_random.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/util_random.py
new file mode 100644
index 0000000000000000000000000000000000000000..dc1ecb6c03b026156c9947cb6d356a822448be0f
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/utils/util_random.py
@@ -0,0 +1,34 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+"""Helpers for random number generators."""
+import numpy as np
+
+
+def ensure_rng(rng=None):
+ """Coerces input into a random number generator.
+
+ If the input is None, then a global random state is returned.
+
+ If the input is a numeric value, then that is used as a seed to construct a
+ random state. Otherwise the input is returned as-is.
+
+ Adapted from [1]_.
+
+ Args:
+ rng (int | numpy.random.RandomState | None):
+ if None, then defaults to the global rng. Otherwise this can be an
+ integer or a RandomState class
+ Returns:
+ (numpy.random.RandomState) : rng -
+ a numpy random number generator
+
+ References:
+ .. [1] https://gitlab.kitware.com/computer-vision/kwarray/blob/master/kwarray/util_random.py#L270 # noqa: E501
+ """
+
+ if rng is None:
+ rng = np.random.mtrand._rand
+ elif isinstance(rng, int):
+ rng = np.random.RandomState(rng)
+ else:
+ rng = rng
+ return rng
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/version.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/version.py
new file mode 100644
index 0000000000000000000000000000000000000000..47989fc0a31f8d8eaa3adff72ab83db61b25b529
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/version.py
@@ -0,0 +1,27 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+
+__version__ = '3.3.0'
+short_version = __version__
+
+
+def parse_version_info(version_str):
+ """Parse a version string into a tuple.
+
+ Args:
+ version_str (str): The version string.
+ Returns:
+ tuple[int | str]: The version info, e.g., "1.3.0" is parsed into
+ (1, 3, 0), and "2.0.0rc1" is parsed into (2, 0, 0, 'rc1').
+ """
+ version_info = []
+ for x in version_str.split('.'):
+ if x.isdigit():
+ version_info.append(int(x))
+ elif x.find('rc') != -1:
+ patch_version = x.split('rc')
+ version_info.append(int(patch_version[0]))
+ version_info.append(f'rc{patch_version[1]}')
+ return tuple(version_info)
+
+
+version_info = parse_version_info(__version__)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/visualization/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/visualization/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..a7edaed9d8701b1be72ff2f7ca646b865007e2eb
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/visualization/__init__.py
@@ -0,0 +1,8 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .local_visualizer import DetLocalVisualizer, TrackLocalVisualizer
+from .palette import get_palette, jitter_color, palette_val
+
+__all__ = [
+ 'palette_val', 'get_palette', 'DetLocalVisualizer', 'jitter_color',
+ 'TrackLocalVisualizer'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/visualization/local_visualizer.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/visualization/local_visualizer.py
new file mode 100644
index 0000000000000000000000000000000000000000..c1b413191e71255cfd4cc324612e069655801ba9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/visualization/local_visualizer.py
@@ -0,0 +1,710 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, List, Optional, Tuple, Union
+
+import cv2
+import mmcv
+import numpy as np
+
+try:
+ import seaborn as sns
+except ImportError:
+ sns = None
+import torch
+from mmengine.dist import master_only
+from mmengine.structures import InstanceData, PixelData
+from mmengine.visualization import Visualizer
+
+from ..evaluation import INSTANCE_OFFSET
+from ..registry import VISUALIZERS
+from ..structures import DetDataSample
+from ..structures.mask import BitmapMasks, PolygonMasks, bitmap_to_polygon
+from .palette import _get_adaptive_scales, get_palette, jitter_color
+
+
+@VISUALIZERS.register_module()
+class DetLocalVisualizer(Visualizer):
+ """MMDetection Local Visualizer.
+
+ Args:
+ name (str): Name of the instance. Defaults to 'visualizer'.
+ image (np.ndarray, optional): the origin image to draw. The format
+ should be RGB. Defaults to None.
+ vis_backends (list, optional): Visual backend config list.
+ Defaults to None.
+ save_dir (str, optional): Save file dir for all storage backends.
+ If it is None, the backend storage will not save any data.
+ bbox_color (str, tuple(int), optional): Color of bbox lines.
+ The tuple of color should be in BGR order. Defaults to None.
+ text_color (str, tuple(int), optional): Color of texts.
+ The tuple of color should be in BGR order.
+ Defaults to (200, 200, 200).
+ mask_color (str, tuple(int), optional): Color of masks.
+ The tuple of color should be in BGR order.
+ Defaults to None.
+ line_width (int, float): The linewidth of lines.
+ Defaults to 3.
+ alpha (int, float): The transparency of bboxes or mask.
+ Defaults to 0.8.
+
+ Examples:
+ >>> import numpy as np
+ >>> import torch
+ >>> from mmengine.structures import InstanceData
+ >>> from mmdet.structures import DetDataSample
+ >>> from mmdet.visualization import DetLocalVisualizer
+
+ >>> det_local_visualizer = DetLocalVisualizer()
+ >>> image = np.random.randint(0, 256,
+ ... size=(10, 12, 3)).astype('uint8')
+ >>> gt_instances = InstanceData()
+ >>> gt_instances.bboxes = torch.Tensor([[1, 2, 2, 5]])
+ >>> gt_instances.labels = torch.randint(0, 2, (1,))
+ >>> gt_det_data_sample = DetDataSample()
+ >>> gt_det_data_sample.gt_instances = gt_instances
+ >>> det_local_visualizer.add_datasample('image', image,
+ ... gt_det_data_sample)
+ >>> det_local_visualizer.add_datasample(
+ ... 'image', image, gt_det_data_sample,
+ ... out_file='out_file.jpg')
+ >>> det_local_visualizer.add_datasample(
+ ... 'image', image, gt_det_data_sample,
+ ... show=True)
+ >>> pred_instances = InstanceData()
+ >>> pred_instances.bboxes = torch.Tensor([[2, 4, 4, 8]])
+ >>> pred_instances.labels = torch.randint(0, 2, (1,))
+ >>> pred_det_data_sample = DetDataSample()
+ >>> pred_det_data_sample.pred_instances = pred_instances
+ >>> det_local_visualizer.add_datasample('image', image,
+ ... gt_det_data_sample,
+ ... pred_det_data_sample)
+ """
+
+ def __init__(self,
+ name: str = 'visualizer',
+ image: Optional[np.ndarray] = None,
+ vis_backends: Optional[Dict] = None,
+ save_dir: Optional[str] = None,
+ bbox_color: Optional[Union[str, Tuple[int]]] = None,
+ text_color: Optional[Union[str,
+ Tuple[int]]] = (200, 200, 200),
+ mask_color: Optional[Union[str, Tuple[int]]] = None,
+ line_width: Union[int, float] = 3,
+ alpha: float = 0.8) -> None:
+ super().__init__(
+ name=name,
+ image=image,
+ vis_backends=vis_backends,
+ save_dir=save_dir)
+ self.bbox_color = bbox_color
+ self.text_color = text_color
+ self.mask_color = mask_color
+ self.line_width = line_width
+ self.alpha = alpha
+ # Set default value. When calling
+ # `DetLocalVisualizer().dataset_meta=xxx`,
+ # it will override the default value.
+ self.dataset_meta = {}
+
+ def _draw_instances(self, image: np.ndarray, instances: ['InstanceData'],
+ classes: Optional[List[str]],
+ palette: Optional[List[tuple]]) -> np.ndarray:
+ """Draw instances of GT or prediction.
+
+ Args:
+ image (np.ndarray): The image to draw.
+ instances (:obj:`InstanceData`): Data structure for
+ instance-level annotations or predictions.
+ classes (List[str], optional): Category information.
+ palette (List[tuple], optional): Palette information
+ corresponding to the category.
+
+ Returns:
+ np.ndarray: the drawn image which channel is RGB.
+ """
+ self.set_image(image)
+
+ if 'bboxes' in instances and instances.bboxes.sum() > 0:
+ bboxes = instances.bboxes
+ labels = instances.labels
+
+ max_label = int(max(labels) if len(labels) > 0 else 0)
+ text_palette = get_palette(self.text_color, max_label + 1)
+ text_colors = [text_palette[label] for label in labels]
+
+ bbox_color = palette if self.bbox_color is None \
+ else self.bbox_color
+ bbox_palette = get_palette(bbox_color, max_label + 1)
+ colors = [bbox_palette[label] for label in labels]
+ self.draw_bboxes(
+ bboxes,
+ edge_colors=colors,
+ alpha=self.alpha,
+ line_widths=self.line_width)
+
+ positions = bboxes[:, :2] + self.line_width
+ areas = (bboxes[:, 3] - bboxes[:, 1]) * (
+ bboxes[:, 2] - bboxes[:, 0])
+ scales = _get_adaptive_scales(areas)
+
+ for i, (pos, label) in enumerate(zip(positions, labels)):
+ if 'label_names' in instances:
+ label_text = instances.label_names[i]
+ else:
+ label_text = classes[
+ label] if classes is not None else f'class {label}'
+ if 'scores' in instances:
+ score = round(float(instances.scores[i]) * 100, 1)
+ label_text += f': {score}'
+
+ import matplotlib, os
+ PROJECT_DIR = os.getenv('DSP_PROJECT_DIR', '/path/to/DSP_PROJECT_DIR') # Set this manually if the environment variable is unavailable
+ font_properties = matplotlib.font_manager.FontProperties(fname=os.path.join(PROJECT_DIR, 'fonts/GILI____.TTF'))
+ self.draw_texts(
+ label_text,
+ pos,
+ colors=text_colors[i],
+ vertical_alignments='bottom',
+ # vertical_alignments='top',
+ horizontal_alignments='left',
+ # font_sizes=int(13 * scales[i]),
+ # font_sizes=int(24 * scales[i]),
+ font_sizes=32,
+ # font_families='monospace',
+ font_properties=font_properties,
+ bboxes=[{
+ 'facecolor': 'black',
+ 'alpha': 0.8,
+ 'pad': 0.7,
+ 'edgecolor': 'none'
+ }])
+
+ if 'masks' in instances:
+ labels = instances.labels
+ masks = instances.masks
+ if isinstance(masks, torch.Tensor):
+ masks = masks.numpy()
+ elif isinstance(masks, (PolygonMasks, BitmapMasks)):
+ masks = masks.to_ndarray()
+
+ masks = masks.astype(bool)
+
+ max_label = int(max(labels) if len(labels) > 0 else 0)
+ mask_color = palette if self.mask_color is None \
+ else self.mask_color
+ mask_palette = get_palette(mask_color, max_label + 1)
+ colors = [jitter_color(mask_palette[label]) for label in labels]
+ text_palette = get_palette(self.text_color, max_label + 1)
+ text_colors = [text_palette[label] for label in labels]
+
+ polygons = []
+ for i, mask in enumerate(masks):
+ contours, _ = bitmap_to_polygon(mask)
+ polygons.extend(contours)
+ self.draw_polygons(polygons, edge_colors='w', alpha=self.alpha)
+ self.draw_binary_masks(masks, colors=colors, alphas=self.alpha)
+
+ if len(labels) > 0 and \
+ ('bboxes' not in instances or
+ instances.bboxes.sum() == 0):
+ # instances.bboxes.sum()==0 represent dummy bboxes.
+ # A typical example of SOLO does not exist bbox branch.
+ areas = []
+ positions = []
+ for mask in masks:
+ _, _, stats, centroids = cv2.connectedComponentsWithStats(
+ mask.astype(np.uint8), connectivity=8)
+ if stats.shape[0] > 1:
+ largest_id = np.argmax(stats[1:, -1]) + 1
+ positions.append(centroids[largest_id])
+ areas.append(stats[largest_id, -1])
+ areas = np.stack(areas, axis=0)
+ scales = _get_adaptive_scales(areas)
+
+ for i, (pos, label) in enumerate(zip(positions, labels)):
+ if 'label_names' in instances:
+ label_text = instances.label_names[i]
+ else:
+ label_text = classes[
+ label] if classes is not None else f'class {label}'
+ if 'scores' in instances:
+ score = round(float(instances.scores[i]) * 100, 1)
+ label_text += f': {score}'
+
+ self.draw_texts(
+ label_text,
+ pos,
+ colors=text_colors[i],
+ # font_sizes=int(13 * scales[i]),
+ font_sizes=24,
+ horizontal_alignments='center',
+ bboxes=[{
+ 'facecolor': 'black',
+ 'alpha': 0.8,
+ 'pad': 0.7,
+ 'edgecolor': 'none'
+ }])
+ return self.get_image()
+
+ def _draw_panoptic_seg(self, image: np.ndarray,
+ panoptic_seg: ['PixelData'],
+ classes: Optional[List[str]],
+ palette: Optional[List]) -> np.ndarray:
+ """Draw panoptic seg of GT or prediction.
+
+ Args:
+ image (np.ndarray): The image to draw.
+ panoptic_seg (:obj:`PixelData`): Data structure for
+ pixel-level annotations or predictions.
+ classes (List[str], optional): Category information.
+
+ Returns:
+ np.ndarray: the drawn image which channel is RGB.
+ """
+ # TODO: Is there a way to bypass?
+ num_classes = len(classes)
+
+ panoptic_seg_data = panoptic_seg.sem_seg[0]
+
+ ids = np.unique(panoptic_seg_data)[::-1]
+
+ if 'label_names' in panoptic_seg:
+ # open set panoptic segmentation
+ classes = panoptic_seg.metainfo['label_names']
+ ignore_index = panoptic_seg.metainfo.get('ignore_index',
+ len(classes))
+ ids = ids[ids != ignore_index]
+ else:
+ # for VOID label
+ ids = ids[ids != num_classes]
+
+ labels = np.array([id % INSTANCE_OFFSET for id in ids], dtype=np.int64)
+ segms = (panoptic_seg_data[None] == ids[:, None, None])
+
+ max_label = int(max(labels) if len(labels) > 0 else 0)
+
+ mask_color = palette if self.mask_color is None \
+ else self.mask_color
+ mask_palette = get_palette(mask_color, max_label + 1)
+ colors = [mask_palette[label] for label in labels]
+
+ self.set_image(image)
+
+ # draw segm
+ polygons = []
+ for i, mask in enumerate(segms):
+ contours, _ = bitmap_to_polygon(mask)
+ polygons.extend(contours)
+ self.draw_polygons(polygons, edge_colors='w', alpha=self.alpha)
+ self.draw_binary_masks(segms, colors=colors, alphas=self.alpha)
+
+ # draw label
+ areas = []
+ positions = []
+ for mask in segms:
+ _, _, stats, centroids = cv2.connectedComponentsWithStats(
+ mask.astype(np.uint8), connectivity=8)
+ max_id = np.argmax(stats[1:, -1]) + 1
+ positions.append(centroids[max_id])
+ areas.append(stats[max_id, -1])
+ areas = np.stack(areas, axis=0)
+ scales = _get_adaptive_scales(areas)
+
+ text_palette = get_palette(self.text_color, max_label + 1)
+ text_colors = [text_palette[label] for label in labels]
+
+ for i, (pos, label) in enumerate(zip(positions, labels)):
+ label_text = classes[label]
+
+ self.draw_texts(
+ label_text,
+ pos,
+ colors=text_colors[i],
+ font_sizes=int(13 * scales[i]),
+ bboxes=[{
+ 'facecolor': 'black',
+ 'alpha': 0.8,
+ 'pad': 0.7,
+ 'edgecolor': 'none'
+ }],
+ horizontal_alignments='center')
+ return self.get_image()
+
+ def _draw_sem_seg(self, image: np.ndarray, sem_seg: PixelData,
+ classes: Optional[List],
+ palette: Optional[List]) -> np.ndarray:
+ """Draw semantic seg of GT or prediction.
+
+ Args:
+ image (np.ndarray): The image to draw.
+ sem_seg (:obj:`PixelData`): Data structure for pixel-level
+ annotations or predictions.
+ classes (list, optional): Input classes for result rendering, as
+ the prediction of segmentation model is a segment map with
+ label indices, `classes` is a list which includes items
+ responding to the label indices. If classes is not defined,
+ visualizer will take `cityscapes` classes by default.
+ Defaults to None.
+ palette (list, optional): Input palette for result rendering, which
+ is a list of color palette responding to the classes.
+ Defaults to None.
+
+ Returns:
+ np.ndarray: the drawn image which channel is RGB.
+ """
+ sem_seg_data = sem_seg.sem_seg
+ if isinstance(sem_seg_data, torch.Tensor):
+ sem_seg_data = sem_seg_data.numpy()
+
+ # 0 ~ num_class, the value 0 means background
+ ids = np.unique(sem_seg_data)
+ ignore_index = sem_seg.metainfo.get('ignore_index', 255)
+ ids = ids[ids != ignore_index]
+
+ if 'label_names' in sem_seg:
+ # open set semseg
+ label_names = sem_seg.metainfo['label_names']
+ else:
+ label_names = classes
+
+ labels = np.array(ids, dtype=np.int64)
+ colors = [palette[label] for label in labels]
+
+ self.set_image(image)
+
+ # draw semantic masks
+ for i, (label, color) in enumerate(zip(labels, colors)):
+ masks = sem_seg_data == label
+ self.draw_binary_masks(masks, colors=[color], alphas=self.alpha)
+ label_text = label_names[label]
+ _, _, stats, centroids = cv2.connectedComponentsWithStats(
+ masks[0].astype(np.uint8), connectivity=8)
+ if stats.shape[0] > 1:
+ largest_id = np.argmax(stats[1:, -1]) + 1
+ centroids = centroids[largest_id]
+
+ areas = stats[largest_id, -1]
+ scales = _get_adaptive_scales(areas)
+
+ self.draw_texts(
+ label_text,
+ centroids,
+ colors=(255, 255, 255),
+ font_sizes=int(13 * scales),
+ horizontal_alignments='center',
+ bboxes=[{
+ 'facecolor': 'black',
+ 'alpha': 0.8,
+ 'pad': 0.7,
+ 'edgecolor': 'none'
+ }])
+
+ return self.get_image()
+
+ @master_only
+ def add_datasample(
+ self,
+ name: str,
+ image: np.ndarray,
+ data_sample: Optional['DetDataSample'] = None,
+ draw_gt: bool = True,
+ draw_pred: bool = True,
+ show: bool = False,
+ wait_time: float = 0,
+ # TODO: Supported in mmengine's Viusalizer.
+ out_file: Optional[str] = None,
+ pred_score_thr: float = 0.3,
+ step: int = 0) -> None:
+ """Draw datasample and save to all backends.
+
+ - If GT and prediction are plotted at the same time, they are
+ displayed in a stitched image where the left image is the
+ ground truth and the right image is the prediction.
+ - If ``show`` is True, all storage backends are ignored, and
+ the images will be displayed in a local window.
+ - If ``out_file`` is specified, the drawn image will be
+ saved to ``out_file``. t is usually used when the display
+ is not available.
+
+ Args:
+ name (str): The image identifier.
+ image (np.ndarray): The image to draw.
+ data_sample (:obj:`DetDataSample`, optional): A data
+ sample that contain annotations and predictions.
+ Defaults to None.
+ draw_gt (bool): Whether to draw GT DetDataSample. Default to True.
+ draw_pred (bool): Whether to draw Prediction DetDataSample.
+ Defaults to True.
+ show (bool): Whether to display the drawn image. Default to False.
+ wait_time (float): The interval of show (s). Defaults to 0.
+ out_file (str): Path to output file. Defaults to None.
+ pred_score_thr (float): The threshold to visualize the bboxes
+ and masks. Defaults to 0.3.
+ step (int): Global step value to record. Defaults to 0.
+ """
+ image = image.clip(0, 255).astype(np.uint8)
+ classes = self.dataset_meta.get('classes', None)
+ palette = self.dataset_meta.get('palette', None)
+
+ gt_img_data = None
+ pred_img_data = None
+
+ if data_sample is not None:
+ data_sample = data_sample.cpu()
+
+ if draw_gt and data_sample is not None:
+ gt_img_data = image
+ if 'gt_instances' in data_sample:
+ gt_img_data = self._draw_instances(image,
+ data_sample.gt_instances,
+ classes, palette)
+ if 'gt_sem_seg' in data_sample:
+ gt_img_data = self._draw_sem_seg(gt_img_data,
+ data_sample.gt_sem_seg,
+ classes, palette)
+
+ if 'gt_panoptic_seg' in data_sample:
+ assert classes is not None, 'class information is ' \
+ 'not provided when ' \
+ 'visualizing panoptic ' \
+ 'segmentation results.'
+ gt_img_data = self._draw_panoptic_seg(
+ gt_img_data, data_sample.gt_panoptic_seg, classes, palette)
+
+ if draw_pred and data_sample is not None:
+ pred_img_data = image
+ if 'pred_instances' in data_sample:
+ pred_instances = data_sample.pred_instances
+ pred_instances = pred_instances[
+ pred_instances.scores > pred_score_thr]
+ pred_img_data = self._draw_instances(image, pred_instances,
+ classes, palette)
+
+ if 'pred_sem_seg' in data_sample:
+ pred_img_data = self._draw_sem_seg(pred_img_data,
+ data_sample.pred_sem_seg,
+ classes, palette)
+
+ if 'pred_panoptic_seg' in data_sample:
+ assert classes is not None, 'class information is ' \
+ 'not provided when ' \
+ 'visualizing panoptic ' \
+ 'segmentation results.'
+ pred_img_data = self._draw_panoptic_seg(
+ pred_img_data, data_sample.pred_panoptic_seg.numpy(),
+ classes, palette)
+
+ if gt_img_data is not None and pred_img_data is not None:
+ drawn_img = np.concatenate((gt_img_data, pred_img_data), axis=1)
+ elif gt_img_data is not None:
+ drawn_img = gt_img_data
+ elif pred_img_data is not None:
+ drawn_img = pred_img_data
+ else:
+ # Display the original image directly if nothing is drawn.
+ drawn_img = image
+
+ # It is convenient for users to obtain the drawn image.
+ # For example, the user wants to obtain the drawn image and
+ # save it as a video during video inference.
+ self.set_image(drawn_img)
+
+ if show:
+ self.show(drawn_img, win_name=name, wait_time=wait_time)
+
+ if out_file is not None:
+ mmcv.imwrite(drawn_img[..., ::-1], out_file)
+ else:
+ self.add_image(name, drawn_img, step)
+
+
+def random_color(seed):
+ """Random a color according to the input seed."""
+ if sns is None:
+ raise RuntimeError('motmetrics is not installed,\
+ please install it by: pip install seaborn')
+ np.random.seed(seed)
+ colors = sns.color_palette()
+ color = colors[np.random.choice(range(len(colors)))]
+ color = tuple([int(255 * c) for c in color])
+ return color
+
+
+@VISUALIZERS.register_module()
+class TrackLocalVisualizer(Visualizer):
+ """Tracking Local Visualizer for the MOT, VIS tasks.
+
+ Args:
+ name (str): Name of the instance. Defaults to 'visualizer'.
+ image (np.ndarray, optional): the origin image to draw. The format
+ should be RGB. Defaults to None.
+ vis_backends (list, optional): Visual backend config list.
+ Defaults to None.
+ save_dir (str, optional): Save file dir for all storage backends.
+ If it is None, the backend storage will not save any data.
+ line_width (int, float): The linewidth of lines.
+ Defaults to 3.
+ alpha (int, float): The transparency of bboxes or mask.
+ Defaults to 0.8.
+ """
+
+ def __init__(self,
+ name: str = 'visualizer',
+ image: Optional[np.ndarray] = None,
+ vis_backends: Optional[Dict] = None,
+ save_dir: Optional[str] = None,
+ line_width: Union[int, float] = 3,
+ alpha: float = 0.8) -> None:
+ super().__init__(name, image, vis_backends, save_dir)
+ self.line_width = line_width
+ self.alpha = alpha
+ # Set default value. When calling
+ # `TrackLocalVisualizer().dataset_meta=xxx`,
+ # it will override the default value.
+ self.dataset_meta = {}
+
+ def _draw_instances(self, image: np.ndarray,
+ instances: InstanceData) -> np.ndarray:
+ """Draw instances of GT or prediction.
+
+ Args:
+ image (np.ndarray): The image to draw.
+ instances (:obj:`InstanceData`): Data structure for
+ instance-level annotations or predictions.
+ Returns:
+ np.ndarray: the drawn image which channel is RGB.
+ """
+ self.set_image(image)
+ classes = self.dataset_meta.get('classes', None)
+
+ # get colors and texts
+ # for the MOT and VIS tasks
+ colors = [random_color(_id) for _id in instances.instances_id]
+ categories = [
+ classes[label] if classes is not None else f'cls{label}'
+ for label in instances.labels
+ ]
+ if 'scores' in instances:
+ texts = [
+ f'{category_name}\n{instance_id} | {score:.2f}'
+ for category_name, instance_id, score in zip(
+ categories, instances.instances_id, instances.scores)
+ ]
+ else:
+ texts = [
+ f'{category_name}\n{instance_id}' for category_name,
+ instance_id in zip(categories, instances.instances_id)
+ ]
+
+ # draw bboxes and texts
+ if 'bboxes' in instances:
+ # draw bboxes
+ bboxes = instances.bboxes.clone()
+ self.draw_bboxes(
+ bboxes,
+ edge_colors=colors,
+ alpha=self.alpha,
+ line_widths=self.line_width)
+ # draw texts
+ if texts is not None:
+ positions = bboxes[:, :2] + self.line_width
+ areas = (bboxes[:, 3] - bboxes[:, 1]) * (
+ bboxes[:, 2] - bboxes[:, 0])
+ scales = _get_adaptive_scales(areas.cpu().numpy())
+ for i, pos in enumerate(positions):
+ self.draw_texts(
+ texts[i],
+ pos,
+ colors='black',
+ font_sizes=int(13 * scales[i]),
+ bboxes=[{
+ 'facecolor': [c / 255 for c in colors[i]],
+ 'alpha': 0.8,
+ 'pad': 0.7,
+ 'edgecolor': 'none'
+ }])
+
+ # draw masks
+ if 'masks' in instances:
+ masks = instances.masks
+ polygons = []
+ for i, mask in enumerate(masks):
+ contours, _ = bitmap_to_polygon(mask)
+ polygons.extend(contours)
+ self.draw_polygons(polygons, edge_colors='w', alpha=self.alpha)
+ self.draw_binary_masks(masks, colors=colors, alphas=self.alpha)
+
+ return self.get_image()
+
+ @master_only
+ def add_datasample(
+ self,
+ name: str,
+ image: np.ndarray,
+ data_sample: DetDataSample = None,
+ draw_gt: bool = True,
+ draw_pred: bool = True,
+ show: bool = False,
+ wait_time: int = 0,
+ # TODO: Supported in mmengine's Viusalizer.
+ out_file: Optional[str] = None,
+ pred_score_thr: float = 0.3,
+ step: int = 0) -> None:
+ """Draw datasample and save to all backends.
+
+ - If GT and prediction are plotted at the same time, they are
+ displayed in a stitched image where the left image is the
+ ground truth and the right image is the prediction.
+ - If ``show`` is True, all storage backends are ignored, and
+ the images will be displayed in a local window.
+ - If ``out_file`` is specified, the drawn image will be
+ saved to ``out_file``. t is usually used when the display
+ is not available.
+ Args:
+ name (str): The image identifier.
+ image (np.ndarray): The image to draw.
+ data_sample (OptTrackSampleList): A data
+ sample that contain annotations and predictions.
+ Defaults to None.
+ draw_gt (bool): Whether to draw GT TrackDataSample.
+ Default to True.
+ draw_pred (bool): Whether to draw Prediction TrackDataSample.
+ Defaults to True.
+ show (bool): Whether to display the drawn image. Default to False.
+ wait_time (int): The interval of show (s). Defaults to 0.
+ out_file (str): Path to output file. Defaults to None.
+ pred_score_thr (float): The threshold to visualize the bboxes
+ and masks. Defaults to 0.3.
+ step (int): Global step value to record. Defaults to 0.
+ """
+ gt_img_data = None
+ pred_img_data = None
+
+ if data_sample is not None:
+ data_sample = data_sample.cpu()
+
+ if draw_gt and data_sample is not None:
+ assert 'gt_instances' in data_sample
+ gt_img_data = self._draw_instances(image, data_sample.gt_instances)
+
+ if draw_pred and data_sample is not None:
+ assert 'pred_track_instances' in data_sample
+ pred_instances = data_sample.pred_track_instances
+ if 'scores' in pred_instances:
+ pred_instances = pred_instances[
+ pred_instances.scores > pred_score_thr].cpu()
+ pred_img_data = self._draw_instances(image, pred_instances)
+
+ if gt_img_data is not None and pred_img_data is not None:
+ drawn_img = np.concatenate((gt_img_data, pred_img_data), axis=1)
+ elif gt_img_data is not None:
+ drawn_img = gt_img_data
+ else:
+ drawn_img = pred_img_data
+
+ if show:
+ self.show(drawn_img, win_name=name, wait_time=wait_time)
+
+ if out_file is not None:
+ mmcv.imwrite(drawn_img[..., ::-1], out_file)
+ else:
+ self.add_image(name, drawn_img, step)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/visualization/local_visualizer.py.bak b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/visualization/local_visualizer.py.bak
new file mode 100644
index 0000000000000000000000000000000000000000..cc6521c56eb167c2c94a3f058594d9e832fb15ad
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/visualization/local_visualizer.py.bak
@@ -0,0 +1,699 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Dict, List, Optional, Tuple, Union
+
+import cv2
+import mmcv
+import numpy as np
+
+try:
+ import seaborn as sns
+except ImportError:
+ sns = None
+import torch
+from mmengine.dist import master_only
+from mmengine.structures import InstanceData, PixelData
+from mmengine.visualization import Visualizer
+
+from ..evaluation import INSTANCE_OFFSET
+from ..registry import VISUALIZERS
+from ..structures import DetDataSample
+from ..structures.mask import BitmapMasks, PolygonMasks, bitmap_to_polygon
+from .palette import _get_adaptive_scales, get_palette, jitter_color
+
+
+@VISUALIZERS.register_module()
+class DetLocalVisualizer(Visualizer):
+ """MMDetection Local Visualizer.
+
+ Args:
+ name (str): Name of the instance. Defaults to 'visualizer'.
+ image (np.ndarray, optional): the origin image to draw. The format
+ should be RGB. Defaults to None.
+ vis_backends (list, optional): Visual backend config list.
+ Defaults to None.
+ save_dir (str, optional): Save file dir for all storage backends.
+ If it is None, the backend storage will not save any data.
+ bbox_color (str, tuple(int), optional): Color of bbox lines.
+ The tuple of color should be in BGR order. Defaults to None.
+ text_color (str, tuple(int), optional): Color of texts.
+ The tuple of color should be in BGR order.
+ Defaults to (200, 200, 200).
+ mask_color (str, tuple(int), optional): Color of masks.
+ The tuple of color should be in BGR order.
+ Defaults to None.
+ line_width (int, float): The linewidth of lines.
+ Defaults to 3.
+ alpha (int, float): The transparency of bboxes or mask.
+ Defaults to 0.8.
+
+ Examples:
+ >>> import numpy as np
+ >>> import torch
+ >>> from mmengine.structures import InstanceData
+ >>> from mmdet.structures import DetDataSample
+ >>> from mmdet.visualization import DetLocalVisualizer
+
+ >>> det_local_visualizer = DetLocalVisualizer()
+ >>> image = np.random.randint(0, 256,
+ ... size=(10, 12, 3)).astype('uint8')
+ >>> gt_instances = InstanceData()
+ >>> gt_instances.bboxes = torch.Tensor([[1, 2, 2, 5]])
+ >>> gt_instances.labels = torch.randint(0, 2, (1,))
+ >>> gt_det_data_sample = DetDataSample()
+ >>> gt_det_data_sample.gt_instances = gt_instances
+ >>> det_local_visualizer.add_datasample('image', image,
+ ... gt_det_data_sample)
+ >>> det_local_visualizer.add_datasample(
+ ... 'image', image, gt_det_data_sample,
+ ... out_file='out_file.jpg')
+ >>> det_local_visualizer.add_datasample(
+ ... 'image', image, gt_det_data_sample,
+ ... show=True)
+ >>> pred_instances = InstanceData()
+ >>> pred_instances.bboxes = torch.Tensor([[2, 4, 4, 8]])
+ >>> pred_instances.labels = torch.randint(0, 2, (1,))
+ >>> pred_det_data_sample = DetDataSample()
+ >>> pred_det_data_sample.pred_instances = pred_instances
+ >>> det_local_visualizer.add_datasample('image', image,
+ ... gt_det_data_sample,
+ ... pred_det_data_sample)
+ """
+
+ def __init__(self,
+ name: str = 'visualizer',
+ image: Optional[np.ndarray] = None,
+ vis_backends: Optional[Dict] = None,
+ save_dir: Optional[str] = None,
+ bbox_color: Optional[Union[str, Tuple[int]]] = None,
+ text_color: Optional[Union[str,
+ Tuple[int]]] = (200, 200, 200),
+ mask_color: Optional[Union[str, Tuple[int]]] = None,
+ line_width: Union[int, float] = 3,
+ alpha: float = 0.8) -> None:
+ super().__init__(
+ name=name,
+ image=image,
+ vis_backends=vis_backends,
+ save_dir=save_dir)
+ self.bbox_color = bbox_color
+ self.text_color = text_color
+ self.mask_color = mask_color
+ self.line_width = line_width
+ self.alpha = alpha
+ # Set default value. When calling
+ # `DetLocalVisualizer().dataset_meta=xxx`,
+ # it will override the default value.
+ self.dataset_meta = {}
+
+ def _draw_instances(self, image: np.ndarray, instances: ['InstanceData'],
+ classes: Optional[List[str]],
+ palette: Optional[List[tuple]]) -> np.ndarray:
+ """Draw instances of GT or prediction.
+
+ Args:
+ image (np.ndarray): The image to draw.
+ instances (:obj:`InstanceData`): Data structure for
+ instance-level annotations or predictions.
+ classes (List[str], optional): Category information.
+ palette (List[tuple], optional): Palette information
+ corresponding to the category.
+
+ Returns:
+ np.ndarray: the drawn image which channel is RGB.
+ """
+ self.set_image(image)
+
+ if 'bboxes' in instances and instances.bboxes.sum() > 0:
+ bboxes = instances.bboxes
+ labels = instances.labels
+
+ max_label = int(max(labels) if len(labels) > 0 else 0)
+ text_palette = get_palette(self.text_color, max_label + 1)
+ text_colors = [text_palette[label] for label in labels]
+
+ bbox_color = palette if self.bbox_color is None \
+ else self.bbox_color
+ bbox_palette = get_palette(bbox_color, max_label + 1)
+ colors = [bbox_palette[label] for label in labels]
+ self.draw_bboxes(
+ bboxes,
+ edge_colors=colors,
+ alpha=self.alpha,
+ line_widths=self.line_width)
+
+ positions = bboxes[:, :2] + self.line_width
+ areas = (bboxes[:, 3] - bboxes[:, 1]) * (
+ bboxes[:, 2] - bboxes[:, 0])
+ scales = _get_adaptive_scales(areas)
+
+ for i, (pos, label) in enumerate(zip(positions, labels)):
+ if 'label_names' in instances:
+ label_text = instances.label_names[i]
+ else:
+ label_text = classes[
+ label] if classes is not None else f'class {label}'
+ if 'scores' in instances:
+ score = round(float(instances.scores[i]) * 100, 1)
+ label_text += f': {score}'
+
+ self.draw_texts(
+ label_text,
+ pos,
+ colors=text_colors[i],
+ font_sizes=int(13 * scales[i]),
+ bboxes=[{
+ 'facecolor': 'black',
+ 'alpha': 0.8,
+ 'pad': 0.7,
+ 'edgecolor': 'none'
+ }])
+
+ if 'masks' in instances:
+ labels = instances.labels
+ masks = instances.masks
+ if isinstance(masks, torch.Tensor):
+ masks = masks.numpy()
+ elif isinstance(masks, (PolygonMasks, BitmapMasks)):
+ masks = masks.to_ndarray()
+
+ masks = masks.astype(bool)
+
+ max_label = int(max(labels) if len(labels) > 0 else 0)
+ mask_color = palette if self.mask_color is None \
+ else self.mask_color
+ mask_palette = get_palette(mask_color, max_label + 1)
+ colors = [jitter_color(mask_palette[label]) for label in labels]
+ text_palette = get_palette(self.text_color, max_label + 1)
+ text_colors = [text_palette[label] for label in labels]
+
+ polygons = []
+ for i, mask in enumerate(masks):
+ contours, _ = bitmap_to_polygon(mask)
+ polygons.extend(contours)
+ self.draw_polygons(polygons, edge_colors='w', alpha=self.alpha)
+ self.draw_binary_masks(masks, colors=colors, alphas=self.alpha)
+
+ if len(labels) > 0 and \
+ ('bboxes' not in instances or
+ instances.bboxes.sum() == 0):
+ # instances.bboxes.sum()==0 represent dummy bboxes.
+ # A typical example of SOLO does not exist bbox branch.
+ areas = []
+ positions = []
+ for mask in masks:
+ _, _, stats, centroids = cv2.connectedComponentsWithStats(
+ mask.astype(np.uint8), connectivity=8)
+ if stats.shape[0] > 1:
+ largest_id = np.argmax(stats[1:, -1]) + 1
+ positions.append(centroids[largest_id])
+ areas.append(stats[largest_id, -1])
+ areas = np.stack(areas, axis=0)
+ scales = _get_adaptive_scales(areas)
+
+ for i, (pos, label) in enumerate(zip(positions, labels)):
+ if 'label_names' in instances:
+ label_text = instances.label_names[i]
+ else:
+ label_text = classes[
+ label] if classes is not None else f'class {label}'
+ if 'scores' in instances:
+ score = round(float(instances.scores[i]) * 100, 1)
+ label_text += f': {score}'
+
+ self.draw_texts(
+ label_text,
+ pos,
+ colors=text_colors[i],
+ font_sizes=int(13 * scales[i]),
+ horizontal_alignments='center',
+ bboxes=[{
+ 'facecolor': 'black',
+ 'alpha': 0.8,
+ 'pad': 0.7,
+ 'edgecolor': 'none'
+ }])
+ return self.get_image()
+
+ def _draw_panoptic_seg(self, image: np.ndarray,
+ panoptic_seg: ['PixelData'],
+ classes: Optional[List[str]],
+ palette: Optional[List]) -> np.ndarray:
+ """Draw panoptic seg of GT or prediction.
+
+ Args:
+ image (np.ndarray): The image to draw.
+ panoptic_seg (:obj:`PixelData`): Data structure for
+ pixel-level annotations or predictions.
+ classes (List[str], optional): Category information.
+
+ Returns:
+ np.ndarray: the drawn image which channel is RGB.
+ """
+ # TODO: Is there a way to bypass?
+ num_classes = len(classes)
+
+ panoptic_seg_data = panoptic_seg.sem_seg[0]
+
+ ids = np.unique(panoptic_seg_data)[::-1]
+
+ if 'label_names' in panoptic_seg:
+ # open set panoptic segmentation
+ classes = panoptic_seg.metainfo['label_names']
+ ignore_index = panoptic_seg.metainfo.get('ignore_index',
+ len(classes))
+ ids = ids[ids != ignore_index]
+ else:
+ # for VOID label
+ ids = ids[ids != num_classes]
+
+ labels = np.array([id % INSTANCE_OFFSET for id in ids], dtype=np.int64)
+ segms = (panoptic_seg_data[None] == ids[:, None, None])
+
+ max_label = int(max(labels) if len(labels) > 0 else 0)
+
+ mask_color = palette if self.mask_color is None \
+ else self.mask_color
+ mask_palette = get_palette(mask_color, max_label + 1)
+ colors = [mask_palette[label] for label in labels]
+
+ self.set_image(image)
+
+ # draw segm
+ polygons = []
+ for i, mask in enumerate(segms):
+ contours, _ = bitmap_to_polygon(mask)
+ polygons.extend(contours)
+ self.draw_polygons(polygons, edge_colors='w', alpha=self.alpha)
+ self.draw_binary_masks(segms, colors=colors, alphas=self.alpha)
+
+ # draw label
+ areas = []
+ positions = []
+ for mask in segms:
+ _, _, stats, centroids = cv2.connectedComponentsWithStats(
+ mask.astype(np.uint8), connectivity=8)
+ max_id = np.argmax(stats[1:, -1]) + 1
+ positions.append(centroids[max_id])
+ areas.append(stats[max_id, -1])
+ areas = np.stack(areas, axis=0)
+ scales = _get_adaptive_scales(areas)
+
+ text_palette = get_palette(self.text_color, max_label + 1)
+ text_colors = [text_palette[label] for label in labels]
+
+ for i, (pos, label) in enumerate(zip(positions, labels)):
+ label_text = classes[label]
+
+ self.draw_texts(
+ label_text,
+ pos,
+ colors=text_colors[i],
+ font_sizes=int(13 * scales[i]),
+ bboxes=[{
+ 'facecolor': 'black',
+ 'alpha': 0.8,
+ 'pad': 0.7,
+ 'edgecolor': 'none'
+ }],
+ horizontal_alignments='center')
+ return self.get_image()
+
+ def _draw_sem_seg(self, image: np.ndarray, sem_seg: PixelData,
+ classes: Optional[List],
+ palette: Optional[List]) -> np.ndarray:
+ """Draw semantic seg of GT or prediction.
+
+ Args:
+ image (np.ndarray): The image to draw.
+ sem_seg (:obj:`PixelData`): Data structure for pixel-level
+ annotations or predictions.
+ classes (list, optional): Input classes for result rendering, as
+ the prediction of segmentation model is a segment map with
+ label indices, `classes` is a list which includes items
+ responding to the label indices. If classes is not defined,
+ visualizer will take `cityscapes` classes by default.
+ Defaults to None.
+ palette (list, optional): Input palette for result rendering, which
+ is a list of color palette responding to the classes.
+ Defaults to None.
+
+ Returns:
+ np.ndarray: the drawn image which channel is RGB.
+ """
+ sem_seg_data = sem_seg.sem_seg
+ if isinstance(sem_seg_data, torch.Tensor):
+ sem_seg_data = sem_seg_data.numpy()
+
+ # 0 ~ num_class, the value 0 means background
+ ids = np.unique(sem_seg_data)
+ ignore_index = sem_seg.metainfo.get('ignore_index', 255)
+ ids = ids[ids != ignore_index]
+
+ if 'label_names' in sem_seg:
+ # open set semseg
+ label_names = sem_seg.metainfo['label_names']
+ else:
+ label_names = classes
+
+ labels = np.array(ids, dtype=np.int64)
+ colors = [palette[label] for label in labels]
+
+ self.set_image(image)
+
+ # draw semantic masks
+ for i, (label, color) in enumerate(zip(labels, colors)):
+ masks = sem_seg_data == label
+ self.draw_binary_masks(masks, colors=[color], alphas=self.alpha)
+ label_text = label_names[label]
+ _, _, stats, centroids = cv2.connectedComponentsWithStats(
+ masks[0].astype(np.uint8), connectivity=8)
+ if stats.shape[0] > 1:
+ largest_id = np.argmax(stats[1:, -1]) + 1
+ centroids = centroids[largest_id]
+
+ areas = stats[largest_id, -1]
+ scales = _get_adaptive_scales(areas)
+
+ self.draw_texts(
+ label_text,
+ centroids,
+ colors=(255, 255, 255),
+ font_sizes=int(13 * scales),
+ horizontal_alignments='center',
+ bboxes=[{
+ 'facecolor': 'black',
+ 'alpha': 0.8,
+ 'pad': 0.7,
+ 'edgecolor': 'none'
+ }])
+
+ return self.get_image()
+
+ @master_only
+ def add_datasample(
+ self,
+ name: str,
+ image: np.ndarray,
+ data_sample: Optional['DetDataSample'] = None,
+ draw_gt: bool = True,
+ draw_pred: bool = True,
+ show: bool = False,
+ wait_time: float = 0,
+ # TODO: Supported in mmengine's Viusalizer.
+ out_file: Optional[str] = None,
+ pred_score_thr: float = 0.3,
+ step: int = 0) -> None:
+ """Draw datasample and save to all backends.
+
+ - If GT and prediction are plotted at the same time, they are
+ displayed in a stitched image where the left image is the
+ ground truth and the right image is the prediction.
+ - If ``show`` is True, all storage backends are ignored, and
+ the images will be displayed in a local window.
+ - If ``out_file`` is specified, the drawn image will be
+ saved to ``out_file``. t is usually used when the display
+ is not available.
+
+ Args:
+ name (str): The image identifier.
+ image (np.ndarray): The image to draw.
+ data_sample (:obj:`DetDataSample`, optional): A data
+ sample that contain annotations and predictions.
+ Defaults to None.
+ draw_gt (bool): Whether to draw GT DetDataSample. Default to True.
+ draw_pred (bool): Whether to draw Prediction DetDataSample.
+ Defaults to True.
+ show (bool): Whether to display the drawn image. Default to False.
+ wait_time (float): The interval of show (s). Defaults to 0.
+ out_file (str): Path to output file. Defaults to None.
+ pred_score_thr (float): The threshold to visualize the bboxes
+ and masks. Defaults to 0.3.
+ step (int): Global step value to record. Defaults to 0.
+ """
+ image = image.clip(0, 255).astype(np.uint8)
+ classes = self.dataset_meta.get('classes', None)
+ palette = self.dataset_meta.get('palette', None)
+
+ gt_img_data = None
+ pred_img_data = None
+
+ if data_sample is not None:
+ data_sample = data_sample.cpu()
+
+ if draw_gt and data_sample is not None:
+ gt_img_data = image
+ if 'gt_instances' in data_sample:
+ gt_img_data = self._draw_instances(image,
+ data_sample.gt_instances,
+ classes, palette)
+ if 'gt_sem_seg' in data_sample:
+ gt_img_data = self._draw_sem_seg(gt_img_data,
+ data_sample.gt_sem_seg,
+ classes, palette)
+
+ if 'gt_panoptic_seg' in data_sample:
+ assert classes is not None, 'class information is ' \
+ 'not provided when ' \
+ 'visualizing panoptic ' \
+ 'segmentation results.'
+ gt_img_data = self._draw_panoptic_seg(
+ gt_img_data, data_sample.gt_panoptic_seg, classes, palette)
+
+ if draw_pred and data_sample is not None:
+ pred_img_data = image
+ if 'pred_instances' in data_sample:
+ pred_instances = data_sample.pred_instances
+ pred_instances = pred_instances[
+ pred_instances.scores > pred_score_thr]
+ pred_img_data = self._draw_instances(image, pred_instances,
+ classes, palette)
+
+ if 'pred_sem_seg' in data_sample:
+ pred_img_data = self._draw_sem_seg(pred_img_data,
+ data_sample.pred_sem_seg,
+ classes, palette)
+
+ if 'pred_panoptic_seg' in data_sample:
+ assert classes is not None, 'class information is ' \
+ 'not provided when ' \
+ 'visualizing panoptic ' \
+ 'segmentation results.'
+ pred_img_data = self._draw_panoptic_seg(
+ pred_img_data, data_sample.pred_panoptic_seg.numpy(),
+ classes, palette)
+
+ if gt_img_data is not None and pred_img_data is not None:
+ drawn_img = np.concatenate((gt_img_data, pred_img_data), axis=1)
+ elif gt_img_data is not None:
+ drawn_img = gt_img_data
+ elif pred_img_data is not None:
+ drawn_img = pred_img_data
+ else:
+ # Display the original image directly if nothing is drawn.
+ drawn_img = image
+
+ # It is convenient for users to obtain the drawn image.
+ # For example, the user wants to obtain the drawn image and
+ # save it as a video during video inference.
+ self.set_image(drawn_img)
+
+ if show:
+ self.show(drawn_img, win_name=name, wait_time=wait_time)
+
+ if out_file is not None:
+ mmcv.imwrite(drawn_img[..., ::-1], out_file)
+ else:
+ self.add_image(name, drawn_img, step)
+
+
+def random_color(seed):
+ """Random a color according to the input seed."""
+ if sns is None:
+ raise RuntimeError('motmetrics is not installed,\
+ please install it by: pip install seaborn')
+ np.random.seed(seed)
+ colors = sns.color_palette()
+ color = colors[np.random.choice(range(len(colors)))]
+ color = tuple([int(255 * c) for c in color])
+ return color
+
+
+@VISUALIZERS.register_module()
+class TrackLocalVisualizer(Visualizer):
+ """Tracking Local Visualizer for the MOT, VIS tasks.
+
+ Args:
+ name (str): Name of the instance. Defaults to 'visualizer'.
+ image (np.ndarray, optional): the origin image to draw. The format
+ should be RGB. Defaults to None.
+ vis_backends (list, optional): Visual backend config list.
+ Defaults to None.
+ save_dir (str, optional): Save file dir for all storage backends.
+ If it is None, the backend storage will not save any data.
+ line_width (int, float): The linewidth of lines.
+ Defaults to 3.
+ alpha (int, float): The transparency of bboxes or mask.
+ Defaults to 0.8.
+ """
+
+ def __init__(self,
+ name: str = 'visualizer',
+ image: Optional[np.ndarray] = None,
+ vis_backends: Optional[Dict] = None,
+ save_dir: Optional[str] = None,
+ line_width: Union[int, float] = 3,
+ alpha: float = 0.8) -> None:
+ super().__init__(name, image, vis_backends, save_dir)
+ self.line_width = line_width
+ self.alpha = alpha
+ # Set default value. When calling
+ # `TrackLocalVisualizer().dataset_meta=xxx`,
+ # it will override the default value.
+ self.dataset_meta = {}
+
+ def _draw_instances(self, image: np.ndarray,
+ instances: InstanceData) -> np.ndarray:
+ """Draw instances of GT or prediction.
+
+ Args:
+ image (np.ndarray): The image to draw.
+ instances (:obj:`InstanceData`): Data structure for
+ instance-level annotations or predictions.
+ Returns:
+ np.ndarray: the drawn image which channel is RGB.
+ """
+ self.set_image(image)
+ classes = self.dataset_meta.get('classes', None)
+
+ # get colors and texts
+ # for the MOT and VIS tasks
+ colors = [random_color(_id) for _id in instances.instances_id]
+ categories = [
+ classes[label] if classes is not None else f'cls{label}'
+ for label in instances.labels
+ ]
+ if 'scores' in instances:
+ texts = [
+ f'{category_name}\n{instance_id} | {score:.2f}'
+ for category_name, instance_id, score in zip(
+ categories, instances.instances_id, instances.scores)
+ ]
+ else:
+ texts = [
+ f'{category_name}\n{instance_id}' for category_name,
+ instance_id in zip(categories, instances.instances_id)
+ ]
+
+ # draw bboxes and texts
+ if 'bboxes' in instances:
+ # draw bboxes
+ bboxes = instances.bboxes.clone()
+ self.draw_bboxes(
+ bboxes,
+ edge_colors=colors,
+ alpha=self.alpha,
+ line_widths=self.line_width)
+ # draw texts
+ if texts is not None:
+ positions = bboxes[:, :2] + self.line_width
+ areas = (bboxes[:, 3] - bboxes[:, 1]) * (
+ bboxes[:, 2] - bboxes[:, 0])
+ scales = _get_adaptive_scales(areas.cpu().numpy())
+ for i, pos in enumerate(positions):
+ self.draw_texts(
+ texts[i],
+ pos,
+ colors='black',
+ font_sizes=int(13 * scales[i]),
+ bboxes=[{
+ 'facecolor': [c / 255 for c in colors[i]],
+ 'alpha': 0.8,
+ 'pad': 0.7,
+ 'edgecolor': 'none'
+ }])
+
+ # draw masks
+ if 'masks' in instances:
+ masks = instances.masks
+ polygons = []
+ for i, mask in enumerate(masks):
+ contours, _ = bitmap_to_polygon(mask)
+ polygons.extend(contours)
+ self.draw_polygons(polygons, edge_colors='w', alpha=self.alpha)
+ self.draw_binary_masks(masks, colors=colors, alphas=self.alpha)
+
+ return self.get_image()
+
+ @master_only
+ def add_datasample(
+ self,
+ name: str,
+ image: np.ndarray,
+ data_sample: DetDataSample = None,
+ draw_gt: bool = True,
+ draw_pred: bool = True,
+ show: bool = False,
+ wait_time: int = 0,
+ # TODO: Supported in mmengine's Viusalizer.
+ out_file: Optional[str] = None,
+ pred_score_thr: float = 0.3,
+ step: int = 0) -> None:
+ """Draw datasample and save to all backends.
+
+ - If GT and prediction are plotted at the same time, they are
+ displayed in a stitched image where the left image is the
+ ground truth and the right image is the prediction.
+ - If ``show`` is True, all storage backends are ignored, and
+ the images will be displayed in a local window.
+ - If ``out_file`` is specified, the drawn image will be
+ saved to ``out_file``. t is usually used when the display
+ is not available.
+ Args:
+ name (str): The image identifier.
+ image (np.ndarray): The image to draw.
+ data_sample (OptTrackSampleList): A data
+ sample that contain annotations and predictions.
+ Defaults to None.
+ draw_gt (bool): Whether to draw GT TrackDataSample.
+ Default to True.
+ draw_pred (bool): Whether to draw Prediction TrackDataSample.
+ Defaults to True.
+ show (bool): Whether to display the drawn image. Default to False.
+ wait_time (int): The interval of show (s). Defaults to 0.
+ out_file (str): Path to output file. Defaults to None.
+ pred_score_thr (float): The threshold to visualize the bboxes
+ and masks. Defaults to 0.3.
+ step (int): Global step value to record. Defaults to 0.
+ """
+ gt_img_data = None
+ pred_img_data = None
+
+ if data_sample is not None:
+ data_sample = data_sample.cpu()
+
+ if draw_gt and data_sample is not None:
+ assert 'gt_instances' in data_sample
+ gt_img_data = self._draw_instances(image, data_sample.gt_instances)
+
+ if draw_pred and data_sample is not None:
+ assert 'pred_track_instances' in data_sample
+ pred_instances = data_sample.pred_track_instances
+ if 'scores' in pred_instances:
+ pred_instances = pred_instances[
+ pred_instances.scores > pred_score_thr].cpu()
+ pred_img_data = self._draw_instances(image, pred_instances)
+
+ if gt_img_data is not None and pred_img_data is not None:
+ drawn_img = np.concatenate((gt_img_data, pred_img_data), axis=1)
+ elif gt_img_data is not None:
+ drawn_img = gt_img_data
+ else:
+ drawn_img = pred_img_data
+
+ if show:
+ self.show(drawn_img, win_name=name, wait_time=wait_time)
+
+ if out_file is not None:
+ mmcv.imwrite(drawn_img[..., ::-1], out_file)
+ else:
+ self.add_image(name, drawn_img, step)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/visualization/palette.py b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/visualization/palette.py
new file mode 100644
index 0000000000000000000000000000000000000000..3c402c08823a60759c984093ba7f05f1e310dbd9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/mmdet/visualization/palette.py
@@ -0,0 +1,108 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Tuple, Union
+
+import mmcv
+import numpy as np
+from mmengine.utils import is_str
+
+
+def palette_val(palette: List[tuple]) -> List[tuple]:
+ """Convert palette to matplotlib palette.
+
+ Args:
+ palette (List[tuple]): A list of color tuples.
+
+ Returns:
+ List[tuple[float]]: A list of RGB matplotlib color tuples.
+ """
+ new_palette = []
+ for color in palette:
+ color = [c / 255 for c in color]
+ new_palette.append(tuple(color))
+ return new_palette
+
+
+def get_palette(palette: Union[List[tuple], str, tuple],
+ num_classes: int) -> List[Tuple[int]]:
+ """Get palette from various inputs.
+
+ Args:
+ palette (list[tuple] | str | tuple): palette inputs.
+ num_classes (int): the number of classes.
+
+ Returns:
+ list[tuple[int]]: A list of color tuples.
+ """
+ assert isinstance(num_classes, int)
+
+ if isinstance(palette, list):
+ dataset_palette = palette
+ elif isinstance(palette, tuple):
+ dataset_palette = [palette] * num_classes
+ elif palette == 'random' or palette is None:
+ state = np.random.get_state()
+ # random color
+ np.random.seed(42)
+ palette = np.random.randint(0, 256, size=(num_classes, 3))
+ np.random.set_state(state)
+ dataset_palette = [tuple(c) for c in palette]
+ elif palette == 'coco':
+ from mmdet.datasets import CocoDataset, CocoPanopticDataset
+ dataset_palette = CocoDataset.METAINFO['palette']
+ if len(dataset_palette) < num_classes:
+ dataset_palette = CocoPanopticDataset.METAINFO['palette']
+ elif palette == 'citys':
+ from mmdet.datasets import CityscapesDataset
+ dataset_palette = CityscapesDataset.METAINFO['palette']
+ elif palette == 'voc':
+ from mmdet.datasets import VOCDataset
+ dataset_palette = VOCDataset.METAINFO['palette']
+ elif is_str(palette):
+ dataset_palette = [mmcv.color_val(palette)[::-1]] * num_classes
+ else:
+ raise TypeError(f'Invalid type for palette: {type(palette)}')
+
+ assert len(dataset_palette) >= num_classes, \
+ 'The length of palette should not be less than `num_classes`.'
+ return dataset_palette
+
+
+def _get_adaptive_scales(areas: np.ndarray,
+ min_area: int = 800,
+ max_area: int = 30000) -> np.ndarray:
+ """Get adaptive scales according to areas.
+
+ The scale range is [0.5, 1.0]. When the area is less than
+ ``min_area``, the scale is 0.5 while the area is larger than
+ ``max_area``, the scale is 1.0.
+
+ Args:
+ areas (ndarray): The areas of bboxes or masks with the
+ shape of (n, ).
+ min_area (int): Lower bound areas for adaptive scales.
+ Defaults to 800.
+ max_area (int): Upper bound areas for adaptive scales.
+ Defaults to 30000.
+
+ Returns:
+ ndarray: The adaotive scales with the shape of (n, ).
+ """
+ scales = 0.5 + (areas - min_area) // (max_area - min_area)
+ scales = np.clip(scales, 0.5, 1.0)
+ return scales
+
+
+def jitter_color(color: tuple) -> tuple:
+ """Randomly jitter the given color in order to better distinguish instances
+ with the same class.
+
+ Args:
+ color (tuple): The RGB color tuple. Each value is between [0, 255].
+
+ Returns:
+ tuple: The jittered color tuple.
+ """
+ jitter = np.random.rand(3)
+ jitter = (jitter / np.linalg.norm(jitter) - 0.5) * 0.5 * 255
+ color = np.clip(jitter + color, 0, 255).astype(np.uint8)
+ return tuple(color)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..33690fe0c432f877f406e30ae10ba64e3fc54906
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/README.md
@@ -0,0 +1,33 @@
+# AlignDETR
+
+> [Align-DETR: Improving DETR with Simple IoU-aware BCE loss](https://arxiv.org/abs/2304.07527)
+
+
+
+## Abstract
+
+DETR has set up a simple end-to-end pipeline for object detection by formulating this task as a set prediction problem, showing promising potential. However, despite the significant progress in improving DETR, this paper identifies a problem of misalignment in the output distribution, which prevents the best-regressed samples from being assigned with high confidence, hindering the model's accuracy. We propose a metric, recall of best-regressed samples, to quantitively evaluate the misalignment problem. Observing its importance, we propose a novel Align-DETR that incorporates a localization precision-aware classification loss in optimization. The proposed loss, IA-BCE, guides the training of DETR to build a strong correlation between classification score and localization precision. We also adopt the mixed-matching strategy, to facilitate DETR-based detectors with faster training convergence while keeping an end-to-end scheme. Moreover, to overcome the dramatic decrease in sample quality induced by the sparsity of queries, we introduce a prime sample weighting mechanism to suppress the interference of unimportant samples. Extensive experiments are conducted with very competitive results reported. In particular, it delivers a 46 (+3.8)% AP on the DAB-DETR baseline with the ResNet-50 backbone and reaches a new SOTA performance of 50.2% AP in the 1x setting on the COCO validation set when employing the strong baseline DINO.
+
+
+
+## Results and Models
+
+| Backbone | Model | Lr schd | box AP | Config | Download |
+| :------: | :---------: | :-----: | :----: | :------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| R-50 | DINO-4scale | 12e | 50.5 | [config](./align_detr-4scale_r50_8xb2-12e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/align_detr/align_detr-4scale_r50_8xb2-12e_coco/align_detr-4scale_r50_8xb2-12e_coco_20230914_095734-61f921af.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/align_detr/align_detr-4scale_r50_8xb2-12e_coco/align_detr-4scale_r50_8xb2-12e_coco_20230914_095734.log.json) |
+| R-50 | DINO-4scale | 24e | 51.4 | [config](./align_detr-4scale_r50_8xb2-24e_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/align_detr/align_detr-4scale_r50_8xb2-24e_coco/align_detr-4scale_r50_8xb2-24e_coco_20230919_152414-f4b6cf76.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/align_detr/align_detr-4scale_r50_8xb2-24e_coco/align_detr-4scale_r50_8xb2-24e_coco_20230919_152414.log.json) |
+
+## Citation
+
+We provide the config files for AlignDETR: [Align-DETR: Improving DETR with Simple IoU-aware BCE loss](https://arxiv.org/abs/2304.07527).
+
+```latex
+@misc{cai2023aligndetr,
+ title={Align-DETR: Improving DETR with Simple IoU-aware BCE loss},
+ author={Zhi Cai and Songtao Liu and Guodong Wang and Zheng Ge and Xiangyu Zhang and Di Huang},
+ year={2023},
+ eprint={2304.07527},
+ archivePrefix={arXiv},
+ primaryClass={cs.CV}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/align_detr/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/align_detr/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..26a49b52476dfa72bde2923db73662ec25f50f33
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/align_detr/__init__.py
@@ -0,0 +1,5 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .align_detr_head import AlignDETRHead
+from .mixed_hungarian_assigner import MixedHungarianAssigner
+
+__all__ = ['AlignDETRHead', 'MixedHungarianAssigner']
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/align_detr/align_detr_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/align_detr/align_detr_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..c06d1bd404c40d3a66340f52211e760e468a47a8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/align_detr/align_detr_head.py
@@ -0,0 +1,508 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Any, Dict, List, Tuple, Union
+
+import torch
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.models.dense_heads import DINOHead
+from mmdet.registry import MODELS
+from mmdet.structures.bbox import (bbox_cxcywh_to_xyxy, bbox_overlaps,
+ bbox_xyxy_to_cxcywh)
+from mmdet.utils import InstanceList
+from .utils import KeysRecorder
+
+
+@MODELS.register_module()
+class AlignDETRHead(DINOHead):
+ r"""Head of the Align-DETR: Improving DETR with Simple IoU-aware BCE loss
+
+ Code is modified from the `official github repo
+ `_.
+
+ More details can be found in the `paper
+ `_ .
+
+ Args:
+ all_layers_num_gt_repeat List[int]: Number to repeat gt for 1-to-k
+ matching between ground truth and predictions of each decoder
+ layer. Only used for matching queries, not for denoising queries.
+ Element count is `num_pred_layer`. If `as_two_stage` is True, then
+ the last element is for encoder output, and the others for
+ decoder layers. Otherwise, all elements are for decoder layers.
+ Defaults to a list of `1` for the last decoder layer and `2` for
+ the others.
+ alpha (float): Hyper-parameter of classification loss that controls
+ the proportion of each item to calculate `t`, the weighted
+ geometric average of the confident score and the IoU score, to
+ align classification and regression scores. Defaults to `0.25`.
+ gamma (float): Hyper-parameter of classification loss to do the hard
+ negative mining. Defaults to `2.0`.
+ tau (float): Hyper-parameter of classification and regression losses,
+ it is the temperature controlling the sharpness of the function
+ to calculate positive sample weight. Defaults to `1.5`.
+ """
+
+ def __init__(self,
+ *args,
+ all_layers_num_gt_repeat: List[int] = None,
+ alpha: float = 0.25,
+ gamma: float = 2.0,
+ tau: float = 1.5,
+ **kwargs) -> None:
+ self.all_layers_num_gt_repeat = all_layers_num_gt_repeat
+ self.alpha = alpha
+ self.gamma = gamma
+ self.tau = tau
+ self.weight_table = torch.zeros(
+ len(all_layers_num_gt_repeat), max(all_layers_num_gt_repeat))
+ for layer_index, num_gt_repeat in enumerate(all_layers_num_gt_repeat):
+ self.weight_table[layer_index][:num_gt_repeat] = torch.exp(
+ -torch.arange(num_gt_repeat) / tau)
+
+ super().__init__(*args, **kwargs)
+ assert len(self.all_layers_num_gt_repeat) == self.num_pred_layer
+
+ def loss_by_feat(self, all_layers_cls_scores: Tensor, *args,
+ **kwargs) -> Any:
+ """Loss function.
+ AlignDETR: This method is based on `DINOHead.loss_by_feat`.
+
+ Args:
+ all_layers_cls_scores (Tensor): Classification scores of all
+ decoder layers, has shape (num_decoder_layers, bs,
+ num_queries_total, cls_out_channels), where
+ `num_queries_total` is the sum of `num_denoising_queries`
+ and `num_matching_queries`.
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ # Wrap `all_layers_cls_scores` with KeysRecorder to record its
+ # `__getitem__` keys and get decoder layer index.
+ all_layers_cls_scores = KeysRecorder(all_layers_cls_scores)
+ result = super(AlignDETRHead,
+ self).loss_by_feat(all_layers_cls_scores, *args,
+ **kwargs)
+ return result
+
+ def loss_by_feat_single(self, cls_scores: Union[KeysRecorder, Tensor],
+ bbox_preds: Tensor,
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict]) -> Tuple[Tensor]:
+ """Loss function for outputs from a single decoder layer of a single
+ feature level.
+ AlignDETR: This method is based on `DINOHead.loss_by_feat_single`.
+
+ Args:
+ cls_scores (Union[KeysRecorder, Tensor]): Box score logits from a
+ single decoder layer for all images, has shape (bs,
+ num_queries, cls_out_channels).
+ bbox_preds (Tensor): Sigmoid outputs from a single decoder layer
+ for all images, with normalized coordinate (cx, cy, w, h) and
+ shape (bs, num_queries, 4).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+
+ Returns:
+ Tuple[Tensor]: A tuple including `loss_cls`, `loss_box` and
+ `loss_iou`.
+ """
+ # AlignDETR: Get layer_index.
+ if isinstance(cls_scores, KeysRecorder):
+ # Outputs are from decoder layer. Get layer_index from
+ # `__getitem__` keys history.
+ keys = [key for key in cls_scores.keys if isinstance(key, int)]
+ assert len(keys) == 1, \
+ 'Failed to extract key from cls_scores.keys: {}'.format(keys)
+ layer_index = keys[0]
+ # Get dn_cls_scores tensor.
+ cls_scores = cls_scores.obj
+ else:
+ # Outputs are from encoder layer.
+ layer_index = self.num_pred_layer - 1
+
+ for img_meta in batch_img_metas:
+ img_meta['layer_index'] = layer_index
+
+ results = super(AlignDETRHead, self).loss_by_feat_single(
+ cls_scores,
+ bbox_preds,
+ batch_gt_instances=batch_gt_instances,
+ batch_img_metas=batch_img_metas)
+ return results
+
+ def get_targets(self, cls_scores_list: List[Tensor],
+ bbox_preds_list: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict]) -> tuple:
+ """Compute regression and classification targets for a batch image.
+
+ Outputs from a single decoder layer of a single feature level are used.
+ AlignDETR: This method is based on `DETRHead.get_targets`.
+
+ Args:
+ cls_scores_list (list[Tensor]): Box score logits from a single
+ decoder layer for each image, has shape [num_queries,
+ cls_out_channels].
+ bbox_preds_list (list[Tensor]): Sigmoid outputs from a single
+ decoder layer for each image, with normalized coordinate
+ (cx, cy, w, h) and shape [num_queries, 4].
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+
+ Returns:
+ tuple: a tuple containing the following targets.
+
+ - labels_list (list[Tensor]): Labels for all images.
+ - label_weights_list (list[Tensor]): Label weights for all images.
+ - bbox_targets_list (list[Tensor]): BBox targets for all images.
+ - bbox_weights_list (list[Tensor]): BBox weights for all images.
+ - num_total_pos (int): Number of positive samples in all images.
+ - num_total_neg (int): Number of negative samples in all images.
+ """
+ results = super(AlignDETRHead,
+ self).get_targets(cls_scores_list, bbox_preds_list,
+ batch_gt_instances, batch_img_metas)
+
+ # AlignDETR: `num_total_pos` for matching queries is the number of
+ # unique gt bboxes in the batch. Refer to AlignDETR official code:
+ # https://github.com/FelixCaae/AlignDETR/blob/8c2b1806026e1b33fe1c282577de1647e352d7f0/aligndetr/criterions/base_criterion.py#L195C15-L195C15 # noqa: E501
+ num_total_pos = sum(
+ len(gt_instances) for gt_instances in batch_gt_instances)
+
+ results = list(results)
+ results[-2] = num_total_pos
+ return tuple(results)
+
+ def _get_targets_single(self, cls_score: Tensor, bbox_pred: Tensor,
+ gt_instances: InstanceData,
+ img_meta: dict) -> tuple:
+ """Compute regression and classification targets for one image.
+
+ Outputs from a single decoder layer of a single feature level are used.
+ AlignDETR: This method is based on `DETRHead._get_targets_single`.
+
+ Args:
+ cls_score (Tensor): Box score logits from a single decoder layer
+ for one image. Shape [num_queries, cls_out_channels].
+ bbox_pred (Tensor): Sigmoid outputs from a single decoder layer
+ for one image, with normalized coordinate (cx, cy, w, h) and
+ shape [num_queries, 4].
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes`` and ``labels``
+ attributes.
+ img_meta (dict): Meta information for one image.
+ layer_index (int): Decoder layer index for the outputs. Defaults
+ to `-1`.
+
+ Returns:
+ tuple[Tensor]: a tuple containing the following for one image.
+
+ - labels (Tensor): Labels of each image.
+ - label_weights (Tensor]): Label weights of each image.
+ - bbox_targets (Tensor): BBox targets of each image.
+ - bbox_weights (Tensor): BBox weights of each image.
+ - pos_inds (Tensor): Sampled positive indices for each image.
+ - neg_inds (Tensor): Sampled negative indices for each image.
+ """
+ img_h, img_w = img_meta['img_shape']
+ factor = bbox_pred.new_tensor([img_w, img_h, img_w,
+ img_h]).unsqueeze(0)
+ # convert bbox_pred from xywh, normalized to xyxy, unnormalized
+ bbox_pred = bbox_cxcywh_to_xyxy(bbox_pred)
+ bbox_pred = bbox_pred * factor
+
+ pred_instances = InstanceData(scores=cls_score, bboxes=bbox_pred)
+
+ # assigner and sampler
+ # AlignDETR: Get `k` of current layer.
+ layer_index = img_meta['layer_index']
+ num_gt_repeat = self.all_layers_num_gt_repeat[layer_index]
+ assign_result = self.assigner.assign(
+ pred_instances=pred_instances,
+ gt_instances=gt_instances,
+ img_meta=img_meta,
+ k=num_gt_repeat)
+
+ gt_bboxes = gt_instances.bboxes
+ gt_labels = gt_instances.labels
+ pos_inds = torch.nonzero(
+ assign_result.gt_inds > 0, as_tuple=False).squeeze(-1).unique()
+ neg_inds = torch.nonzero(
+ assign_result.gt_inds == 0, as_tuple=False).squeeze(-1).unique()
+ pos_assigned_gt_inds = assign_result.gt_inds[pos_inds] - 1
+ pos_gt_bboxes = gt_bboxes[pos_assigned_gt_inds.long(), :]
+
+ # AlignDETR: Get label targets, label weights, and bbox weights.
+ target_results = self._get_align_detr_targets_single(
+ cls_score,
+ bbox_pred,
+ gt_labels,
+ pos_gt_bboxes,
+ pos_inds,
+ pos_assigned_gt_inds,
+ layer_index,
+ is_matching_queries=True)
+
+ label_targets, label_weights, bbox_weights = target_results
+
+ # bbox targets
+ bbox_targets = torch.zeros_like(bbox_pred, dtype=gt_bboxes.dtype)
+
+ # DETR regress the relative position of boxes (cxcywh) in the image.
+ # Thus the learning target should be normalized by the image size, also
+ # the box format should be converted from defaultly x1y1x2y2 to cxcywh.
+ pos_gt_bboxes_normalized = pos_gt_bboxes / factor
+ pos_gt_bboxes_targets = bbox_xyxy_to_cxcywh(pos_gt_bboxes_normalized)
+ bbox_targets[pos_inds] = pos_gt_bboxes_targets
+ return (label_targets, label_weights, bbox_targets, bbox_weights,
+ pos_inds, neg_inds)
+
+ def _loss_dn_single(self, dn_cls_scores: KeysRecorder,
+ dn_bbox_preds: Tensor,
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ dn_meta: Dict[str, int]) -> Tuple[Tensor]:
+ """Denoising loss for outputs from a single decoder layer.
+ AlignDETR: This method is based on `DINOHead._loss_dn_single`.
+
+ Args:
+ dn_cls_scores (KeysRecorder): Classification scores of a single
+ decoder layer in denoising part, has shape (bs,
+ num_denoising_queries, cls_out_channels).
+ dn_bbox_preds (Tensor): Regression outputs of a single decoder
+ layer in denoising part. Each is a 4D-tensor with normalized
+ coordinate format (cx, cy, w, h) and has shape
+ (bs, num_denoising_queries, 4).
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ dn_meta (Dict[str, int]): The dictionary saves information about
+ group collation, including 'num_denoising_queries' and
+ 'num_denoising_groups'. It will be used for split outputs of
+ denoising and matching parts and loss calculation.
+
+ Returns:
+ Tuple[Tensor]: A tuple including `loss_cls`, `loss_box` and
+ `loss_iou`.
+ """
+ # AlignDETR: Get dn_cls_scores tensor.
+ dn_cls_scores = dn_cls_scores.obj
+
+ # AlignDETR: Add layer outputs to meta info because they are not
+ # variables of method `_get_dn_targets_single`.
+ for image_index, img_meta in enumerate(batch_img_metas):
+ img_meta['dn_cls_score'] = dn_cls_scores[image_index]
+ img_meta['dn_bbox_pred'] = dn_bbox_preds[image_index]
+
+ results = super()._loss_dn_single(dn_cls_scores, dn_bbox_preds,
+ batch_gt_instances, batch_img_metas,
+ dn_meta)
+ return results
+
+ def _get_dn_targets_single(self, gt_instances: InstanceData,
+ img_meta: dict, dn_meta: Dict[str,
+ int]) -> tuple:
+ """Get targets in denoising part for one image.
+ AlignDETR: This method is based on
+ `DINOHead._get_dn_targets_single`.
+ and 1) Added passing `dn_cls_score`, `dn_bbox_pred` to this
+ method; 2) Modified the way to get targets.
+ Args:
+ dn_cls_score (Tensor): Box score logits from a single decoder
+ layer in denoising part for one image, has shape
+ [num_denoising_queries, cls_out_channels].
+ dn_bbox_pred (Tensor): Sigmoid outputs from a single decoder
+ layer in denoising part for one image, with
+ normalized coordinate (cx, cy, w, h) and shape
+ [num_denoising_queries, 4].
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It should includes ``bboxes`` and ``labels``
+ attributes.
+ img_meta (dict): Meta information for one image.
+ dn_meta (Dict[str, int]): The dictionary saves information about
+ group collation, including 'num_denoising_queries' and
+ 'num_denoising_groups'. It will be used for split outputs of
+ denoising and matching parts and loss calculation.
+
+ Returns:
+ tuple[Tensor]: a tuple containing the following for one image.
+
+ - labels (Tensor): Labels of each image.
+ - label_weights (Tensor]): Label weights of each image.
+ - bbox_targets (Tensor): BBox targets of each image.
+ - bbox_weights (Tensor): BBox weights of each image.
+ - pos_inds (Tensor): Sampled positive indices for each image.
+ - neg_inds (Tensor): Sampled negative indices for each image.
+ """
+ gt_bboxes = gt_instances.bboxes
+ gt_labels = gt_instances.labels
+ num_groups = dn_meta['num_denoising_groups']
+ num_denoising_queries = dn_meta['num_denoising_queries']
+ num_queries_each_group = int(num_denoising_queries / num_groups)
+ device = gt_bboxes.device
+
+ if len(gt_labels) > 0:
+ t = torch.arange(len(gt_labels), dtype=torch.long, device=device)
+ t = t.unsqueeze(0).repeat(num_groups, 1)
+ pos_assigned_gt_inds = t.flatten()
+ pos_inds = torch.arange(
+ num_groups, dtype=torch.long, device=device)
+ pos_inds = pos_inds.unsqueeze(1) * num_queries_each_group + t
+ pos_inds = pos_inds.flatten()
+ else:
+ pos_inds = pos_assigned_gt_inds = \
+ gt_bboxes.new_tensor([], dtype=torch.long)
+
+ neg_inds = pos_inds + num_queries_each_group // 2
+
+ # AlignDETR: Get meta info and layer outputs.
+ img_h, img_w = img_meta['img_shape']
+ dn_cls_score = img_meta['dn_cls_score']
+ dn_bbox_pred = img_meta['dn_bbox_pred']
+ factor = dn_bbox_pred.new_tensor([img_w, img_h, img_w,
+ img_h]).unsqueeze(0)
+
+ # AlignDETR: Convert dn_bbox_pred from xywh, normalized to xyxy,
+ # unnormalized.
+ dn_bbox_pred = bbox_cxcywh_to_xyxy(dn_bbox_pred)
+ dn_bbox_pred = dn_bbox_pred * factor
+
+ # AlignDETR: Get label targets, label weights, and bbox weights.
+ target_results = self._get_align_detr_targets_single(
+ dn_cls_score, dn_bbox_pred, gt_labels,
+ gt_bboxes.repeat([num_groups, 1]), pos_inds, pos_assigned_gt_inds)
+
+ label_targets, label_weights, bbox_weights = target_results
+
+ # bbox targets
+ bbox_targets = torch.zeros(num_denoising_queries, 4, device=device)
+
+ # DETR regress the relative position of boxes (cxcywh) in the image.
+ # Thus the learning target should be normalized by the image size, also
+ # the box format should be converted from defaultly x1y1x2y2 to cxcywh.
+ gt_bboxes_normalized = gt_bboxes / factor
+ gt_bboxes_targets = bbox_xyxy_to_cxcywh(gt_bboxes_normalized)
+ bbox_targets[pos_inds] = gt_bboxes_targets.repeat([num_groups, 1])
+
+ return (label_targets, label_weights, bbox_targets, bbox_weights,
+ pos_inds, neg_inds)
+
+ def _get_align_detr_targets_single(self,
+ cls_score: Tensor,
+ bbox_pred: Tensor,
+ gt_labels: Tensor,
+ pos_gt_bboxes: Tensor,
+ pos_inds: Tensor,
+ pos_assigned_gt_inds: Tensor,
+ layer_index: int = -1,
+ is_matching_queries: bool = False):
+ '''AlignDETR: Get label targets, label weights, and bbox weights based
+ on `t`, the weighted geometric average of the confident score and
+ the IoU score, to align classification and regression scores.
+
+ Args:
+ cls_score (Tensor): Box score logits from the last encoder layer
+ or a single decoder layer for one image. Shape
+ [num_queries or num_denoising_queries, cls_out_channels].
+ bbox_pred (Tensor): Sigmoid outputs from the last encoder layer
+ or a single decoder layer for one image, with unnormalized
+ coordinate (x, y, x, y) and shape
+ [num_queries or num_denoising_queries, 4].
+ gt_labels (Tensor): Ground truth classification labels for one
+ image, has shape [num_gt].
+ pos_gt_bboxes (Tensor): Positive ground truth bboxes for one
+ image, with unnormalized coordinate (x, y, x, y) and shape
+ [num_positive, 4].
+ pos_inds (Tensor): Positive prediction box indices, has shape
+ [num_positive].
+ pos_assigned_gt_inds Tensor: Positive ground truth box indices,
+ has shape [num_positive].
+ layer_index (int): decoder layer index for the outputs. Defaults
+ to `-1`.
+ is_matching_queries (bool): The outputs are from matching
+ queries or denoising queries. Defaults to `False`.
+
+ Returns:
+ tuple[Tensor]: a tuple containing the following for one image.
+
+ - label_targets (Tensor): Labels of one image. Shape
+ [num_queries or num_denoising_queries, cls_out_channels].
+ - label_weights (Tensor): Label weights of one image. Shape
+ [num_queries or num_denoising_queries, cls_out_channels].
+ - bbox_weights (Tensor): BBox weights of one image. Shape
+ [num_queries or num_denoising_queries, 4].
+ '''
+
+ # Classification loss
+ # = 1 * BCE(prob, t * rank_weights) for positive sample;
+ # = prob**gamma * BCE(prob, 0) for negative sample.
+ # That is,
+ # label_targets = 0 for negative sample;
+ # = t * rank_weights for positive sample.
+ # label_weights = pred**gamma for negative sample;
+ # = 1 for positive sample.
+ cls_prob = cls_score.sigmoid()
+ label_targets = torch.zeros_like(
+ cls_score, device=pos_gt_bboxes.device)
+ label_weights = cls_prob**self.gamma
+
+ bbox_weights = torch.zeros_like(bbox_pred, dtype=pos_gt_bboxes.dtype)
+
+ if len(pos_inds) == 0:
+ return label_targets, label_weights, bbox_weights
+
+ pos_cls_score_inds = (pos_inds, gt_labels[pos_assigned_gt_inds])
+ iou_scores = bbox_overlaps(
+ bbox_pred[pos_inds], pos_gt_bboxes, is_aligned=True)
+
+ # t (Tensor): The weighted geometric average of the confident score
+ # and the IoU score, to align classification and regression scores.
+ # Shape [num_positive].
+ t = (
+ cls_prob[pos_cls_score_inds]**self.alpha *
+ iou_scores**(1 - self.alpha))
+ t = torch.clamp(t, 0.01).detach()
+
+ # Calculate rank_weights for matching queries.
+ if is_matching_queries:
+ # rank_weights (Tensor): Weights of each group of predictions
+ # assigned to the same positive gt bbox. Shape [num_positive].
+ rank_weights = torch.zeros_like(t, dtype=self.weight_table.dtype)
+
+ assert 0 <= layer_index < len(self.weight_table), layer_index
+ rank_to_weight = self.weight_table[layer_index].to(
+ rank_weights.device)
+ unique_gt_inds = torch.unique(pos_assigned_gt_inds)
+
+ # For each positive gt bbox, get all predictions assigned to it,
+ # then calculate rank weights for this group of predictions.
+ for gt_index in unique_gt_inds:
+ pred_group_cond = pos_assigned_gt_inds == gt_index
+ # Weights are based on their rank sorted by t in the group.
+ pred_group = t[pred_group_cond]
+ indices = pred_group.sort(descending=True)[1]
+ group_weights = torch.zeros_like(
+ indices, dtype=self.weight_table.dtype)
+ group_weights[indices] = rank_to_weight[:len(indices)]
+ rank_weights[pred_group_cond] = group_weights
+
+ t = t * rank_weights
+ pos_bbox_weights = rank_weights.unsqueeze(-1).repeat(
+ 1, bbox_pred.size(-1))
+ bbox_weights[pos_inds] = pos_bbox_weights
+ else:
+ bbox_weights[pos_inds] = 1.0
+
+ label_targets[pos_cls_score_inds] = t
+ label_weights[pos_cls_score_inds] = 1.0
+
+ return label_targets, label_weights, bbox_weights
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/align_detr/mixed_hungarian_assigner.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/align_detr/mixed_hungarian_assigner.py
new file mode 100644
index 0000000000000000000000000000000000000000..cc31b5e6aa67e546d8655c5fd513419e7cbe437e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/align_detr/mixed_hungarian_assigner.py
@@ -0,0 +1,162 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Optional, Union
+
+import torch
+from mmengine import ConfigDict
+from mmengine.structures import InstanceData
+from scipy.optimize import linear_sum_assignment
+from torch import Tensor
+
+from mmdet.models.task_modules import AssignResult, BaseAssigner
+from mmdet.registry import TASK_UTILS
+
+
+@TASK_UTILS.register_module()
+class MixedHungarianAssigner(BaseAssigner):
+ """Computes 1-to-k matching between ground truth and predictions.
+
+ This class computes an assignment between the targets and the predictions
+ based on the costs. The costs are weighted sum of some components.
+ For DETR the costs are weighted sum of classification cost, regression L1
+ cost and regression iou cost. The targets don't include the no_object, so
+ generally there are more predictions than targets. After the 1-to-k
+ gt-pred matching, the un-matched are treated as backgrounds. Thus
+ each query prediction will be assigned with `0` or a positive integer
+ indicating the ground truth index:
+
+ - 0: negative sample, no assigned gt
+ - positive integer: positive sample, index (1-based) of assigned gt
+
+ Args:
+ match_costs (:obj:`ConfigDict` or dict or \
+ List[Union[:obj:`ConfigDict`, dict]]): Match cost configs.
+ """
+
+ def __init__(
+ self, match_costs: Union[List[Union[dict, ConfigDict]], dict,
+ ConfigDict]
+ ) -> None:
+
+ if isinstance(match_costs, dict):
+ match_costs = [match_costs]
+ elif isinstance(match_costs, list):
+ assert len(match_costs) > 0, \
+ 'match_costs must not be a empty list.'
+
+ self.match_costs = [
+ TASK_UTILS.build(match_cost) for match_cost in match_costs
+ ]
+
+ def assign(self,
+ pred_instances: InstanceData,
+ gt_instances: InstanceData,
+ img_meta: Optional[dict] = None,
+ k: int = 1,
+ **kwargs) -> AssignResult:
+ """Computes 1-to-k gt-pred matching based on the weighted costs.
+
+ This method assign each query prediction to a ground truth or
+ background. The `assigned_gt_inds` with -1 means don't care,
+ 0 means negative sample, and positive number is the index (1-based)
+ of assigned gt.
+ The assignment is done in the following steps, the order matters.
+
+ 1. Assign every prediction to -1.
+ 2. Compute the weighted costs, each cost has shape
+ (num_preds, num_gts).
+ 3. Update k according to num_preds and num_gts, then repeat
+ costs k times to shape: (num_preds, k * num_gts), so that each
+ gt will match k predictions.
+ 4. Do Hungarian matching on CPU based on the costs.
+ 5. Assign all to 0 (background) first, then for each matched pair
+ between predictions and gts, treat this prediction as foreground
+ and assign the corresponding gt index (plus 1) to it.
+
+ Args:
+ pred_instances (:obj:`InstanceData`): Instances of model
+ predictions. It includes ``priors``, and the priors can
+ be anchors or points, or the bboxes predicted by the
+ previous stage, has shape (n, 4). The bboxes predicted by
+ the current model or stage will be named ``bboxes``,
+ ``labels``, and ``scores``, the same as the ``InstanceData``
+ in other places. It may includes ``masks``, with shape
+ (n, h, w) or (n, l).
+ gt_instances (:obj:`InstanceData`): Ground truth of instance
+ annotations. It usually includes ``bboxes``, with shape (k, 4),
+ ``labels``, with shape (k, ) and ``masks``, with shape
+ (k, h, w) or (k, l).
+ img_meta (dict): Image information for one image.
+
+ Returns:
+ :obj:`AssignResult`: The assigned result.
+ """
+ assert isinstance(gt_instances.labels, Tensor)
+ num_gts, num_preds = len(gt_instances), len(pred_instances)
+ gt_labels = gt_instances.labels
+ device = gt_labels.device
+
+ # 1. Assign -1 by default.
+ assigned_gt_inds = torch.full((num_preds, ),
+ -1,
+ dtype=torch.long,
+ device=device)
+ assigned_labels = torch.full((num_preds, ),
+ -1,
+ dtype=torch.long,
+ device=device)
+
+ if num_gts == 0 or num_preds == 0:
+ # No ground truth or boxes, return empty assignment.
+ if num_gts == 0:
+ # No ground truth, assign all to background.
+ assigned_gt_inds[:] = 0
+ return AssignResult(
+ num_gts=num_gts,
+ gt_inds=assigned_gt_inds,
+ max_overlaps=None,
+ labels=assigned_labels)
+
+ # 2. Compute weighted costs.
+ cost_list = []
+ for match_cost in self.match_costs:
+ cost = match_cost(
+ pred_instances=pred_instances,
+ gt_instances=gt_instances,
+ img_meta=img_meta)
+ cost_list.append(cost)
+ cost = torch.stack(cost_list).sum(dim=0)
+
+ # 3. Update k according to num_preds and num_gts, then
+ # repeat the ground truth k times to perform 1-to-k gt-pred
+ # matching. For example, if num_preds = 900, num_gts = 3, then
+ # there are only 3 gt-pred pairs in sum for 1-1 matching.
+ # However, for 1-k gt-pred matching, if k = 4, then each
+ # gt is assigned 4 unique predictions, so there would be 12
+ # gt-pred pairs in sum.
+ k = max(1, min(k, num_preds // num_gts))
+ cost = cost.repeat(1, k)
+
+ # 4. Do Hungarian matching on CPU using linear_sum_assignment.
+ cost = cost.detach().cpu()
+ if linear_sum_assignment is None:
+ raise ImportError('Please run "pip install scipy" '
+ 'to install scipy first.')
+
+ matched_row_inds, matched_col_inds = linear_sum_assignment(cost)
+ matched_row_inds = torch.from_numpy(matched_row_inds).to(device)
+ matched_col_inds = torch.from_numpy(matched_col_inds).to(device)
+
+ matched_col_inds = matched_col_inds % num_gts
+ # 5. Assign backgrounds and foregrounds.
+ # Assign all indices to backgrounds first.
+ assigned_gt_inds[:] = 0
+ # Assign foregrounds based on matching results.
+ assigned_gt_inds[matched_row_inds] = matched_col_inds + 1
+ assigned_labels[matched_row_inds] = gt_labels[matched_col_inds]
+ assign_result = AssignResult(
+ num_gts=k * num_gts,
+ gt_inds=assigned_gt_inds,
+ max_overlaps=None,
+ labels=assigned_labels)
+
+ return assign_result
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/align_detr/utils.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/align_detr/utils.py
new file mode 100644
index 0000000000000000000000000000000000000000..5a3c17ec5dac3b036613c04a70b3d385a4230ed9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/align_detr/utils.py
@@ -0,0 +1,34 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Any, List, Optional
+
+
+class KeysRecorder:
+ """Wrap object to record its `__getitem__` keys in the history.
+
+ Args:
+ obj (object): Any object that supports `__getitem__`.
+ keys (List): List of keys already recorded. Default to None.
+ """
+
+ def __init__(self, obj: Any, keys: Optional[List[Any]] = None) -> None:
+ self.obj = obj
+
+ if keys is None:
+ keys = []
+ self.keys = keys
+
+ def __getitem__(self, key: Any) -> 'KeysRecorder':
+ """Wrap method `__getitem__` to record its keys.
+
+ Args:
+ key: Key that is passed to the object.
+
+ Returns:
+ result (KeysRecorder): KeysRecorder instance that wraps sub_obj.
+ """
+ sub_obj = self.obj.__getitem__(key)
+ keys = self.keys.copy()
+ keys.append(key)
+ # Create a KeysRecorder instance from the sub_obj.
+ result = KeysRecorder(sub_obj, keys)
+ return result
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/configs/align_detr-4scale_r50_8xb2-12e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/configs/align_detr-4scale_r50_8xb2-12e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..0fe069905e0ccab6d387fe95c050cdfe92fc6da3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/configs/align_detr-4scale_r50_8xb2-12e_coco.py
@@ -0,0 +1,185 @@
+_base_ = [
+ '../../../configs/_base_/datasets/coco_detection.py',
+ '../../../configs/_base_/default_runtime.py'
+]
+custom_imports = dict(
+ imports=['projects.AlignDETR.align_detr'], allow_failed_imports=False)
+
+model = dict(
+ type='DINO',
+ num_queries=900, # num_matching_queries
+ with_box_refine=True,
+ as_two_stage=True,
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=1),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(1, 2, 3),
+ # AlignDETR: Only freeze stem.
+ frozen_stages=0,
+ norm_cfg=dict(type='FrozenBN', requires_grad=False),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='ChannelMapper',
+ in_channels=[512, 1024, 2048],
+ kernel_size=1,
+ out_channels=256,
+ # AlignDETR: Add conv bias.
+ bias=True,
+ act_cfg=None,
+ norm_cfg=dict(type='GN', num_groups=32),
+ num_outs=4),
+ encoder=dict(
+ num_layers=6,
+ layer_cfg=dict(
+ self_attn_cfg=dict(embed_dims=256, num_levels=4,
+ dropout=0.0), # 0.1 for DeformDETR
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048, # 1024 for DeformDETR
+ ffn_drop=0.0))), # 0.1 for DeformDETR
+ decoder=dict(
+ num_layers=6,
+ return_intermediate=True,
+ layer_cfg=dict(
+ self_attn_cfg=dict(embed_dims=256, num_heads=8,
+ dropout=0.0), # 0.1 for DeformDETR
+ cross_attn_cfg=dict(embed_dims=256, num_levels=4,
+ dropout=0.0), # 0.1 for DeformDETR
+ ffn_cfg=dict(
+ embed_dims=256,
+ feedforward_channels=2048, # 1024 for DeformDETR
+ ffn_drop=0.0)), # 0.1 for DeformDETR
+ post_norm_cfg=None),
+ positional_encoding=dict(
+ num_feats=128,
+ normalize=True,
+ # AlignDETR: Set offset and temperature the same as DeformDETR.
+ offset=-0.5, # -0.5 for DeformDETR
+ temperature=10000), # 10000 for DeformDETR
+ bbox_head=dict(
+ type='AlignDETRHead',
+ # AlignDETR: First 6 elements of `all_layers_num_gt_repeat` are for
+ # decoder layers' outputs. The last element is for encoder layer.
+ all_layers_num_gt_repeat=[2, 2, 2, 2, 2, 1, 2],
+ alpha=0.25,
+ gamma=2.0,
+ tau=1.5,
+ num_classes=80,
+ sync_cls_avg_factor=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True,
+ loss_weight=1.0), # 2.0 in DeformDETR
+ loss_bbox=dict(type='L1Loss', loss_weight=5.0),
+ loss_iou=dict(type='GIoULoss', loss_weight=2.0)),
+ dn_cfg=dict( # TODO: Move to model.train_cfg ?
+ label_noise_scale=0.5,
+ box_noise_scale=1.0, # 0.4 for DN-DETR
+ group_cfg=dict(dynamic=True, num_groups=None,
+ num_dn_queries=100)), # TODO: half num_dn_queries
+ # training and testing settings
+ train_cfg=dict(
+ assigner=dict(
+ type='MixedHungarianAssigner',
+ match_costs=[
+ dict(type='FocalLossCost', weight=2.0),
+ dict(type='BBoxL1Cost', weight=5.0, box_format='xywh'),
+ dict(type='IoUCost', iou_mode='giou', weight=2.0)
+ ])),
+ test_cfg=dict(max_per_img=300)) # 100 for DeformDETR
+
+# train_pipeline, NOTE the img_scale and the Pad's size_divisor is different
+# from the default setting in mmdet.
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='PackDetInputs')
+]
+train_dataloader = dict(
+ dataset=dict(
+ # AlignDETR: Filter empty gt.
+ filter_cfg=dict(filter_empty_gt=True),
+ pipeline=train_pipeline))
+
+# optimizer
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(
+ type='AdamW',
+ lr=0.0001, # 0.0002 for DeformDETR
+ weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(
+ custom_keys={'backbone': dict(lr_mult=0.1)},
+ # AlignDETR: No norm decay.
+ norm_decay_mult=0.0)
+) # custom_keys contains sampling_offsets and reference_points in DeformDETR # noqa
+
+# learning policy
+max_epochs = 12
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=0.0001,
+ by_epoch=False,
+ begin=0,
+ end=2000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[11],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/configs/align_detr-4scale_r50_8xb2-24e_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/configs/align_detr-4scale_r50_8xb2-24e_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..f62114ce0a8d4d1755603064b9f173beeff12ce5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/AlignDETR/configs/align_detr-4scale_r50_8xb2-24e_coco.py
@@ -0,0 +1,19 @@
+_base_ = './align_detr-4scale_r50_8xb2-12e_coco.py'
+max_epochs = 24
+train_cfg = dict(
+ type='EpochBasedTrainLoop', max_epochs=max_epochs, val_interval=1)
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=0.0001,
+ by_epoch=False,
+ begin=0,
+ end=2000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[20],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..787592ade508337d341342058a25d471603d93fe
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/README.md
@@ -0,0 +1,32 @@
+# CO-DETR
+
+> [DETRs with Collaborative Hybrid Assignments Training](https://arxiv.org/abs/2211.12860)
+
+
+
+## Abstract
+
+In this paper, we provide the observation that too few queries assigned as positive samples in DETR with one-to-one set matching leads to sparse supervision on the encoder's output which considerably hurt the discriminative feature learning of the encoder and vice visa for attention learning in the decoder. To alleviate this, we present a novel collaborative hybrid assignments training scheme, namely Co-DETR, to learn more efficient and effective DETR-based detectors from versatile label assignment manners. This new training scheme can easily enhance the encoder's learning ability in end-to-end detectors by training the multiple parallel auxiliary heads supervised by one-to-many label assignments such as ATSS and Faster RCNN. In addition, we conduct extra customized positive queries by extracting the positive coordinates from these auxiliary heads to improve the training efficiency of positive samples in the decoder. In inference, these auxiliary heads are discarded and thus our method introduces no additional parameters and computational cost to the original detector while requiring no hand-crafted non-maximum suppression (NMS). We conduct extensive experiments to evaluate the effectiveness of the proposed approach on DETR variants, including DAB-DETR, Deformable-DETR, and DINO-Deformable-DETR. The state-of-the-art DINO-Deformable-DETR with Swin-L can be improved from 58.5% to 59.5% AP on COCO val. Surprisingly, incorporated with ViT-L backbone, we achieve 66.0% AP on COCO test-dev and 67.9% AP on LVIS val, outperforming previous methods by clear margins with much fewer model sizes.
+
+
+

+
+
+## Results and Models
+
+| Model | Backbone | Epochs | Aug | Dataset | box AP | Config | Download |
+| :-------: | :------: | :----: | :--: | :---------------------------: | :----: | :--------------------------------------------------------------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Co-DINO | R50 | 12 | LSJ | COCO | 52.0 | [config](configs/codino/co_dino_5scale_r50_lsj_8xb2_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/codetr/co_dino_5scale_r50_lsj_8xb2_1x_coco/co_dino_5scale_r50_lsj_8xb2_1x_coco-69a72d67.pth)\\ [log](https://download.openmmlab.com/mmdetection/v3.0/codetr/co_dino_5scale_r50_lsj_8xb2_1x_coco/co_dino_5scale_r50_lsj_8xb2_1x_coco_20230818_150457.json) |
+| Co-DINO\* | R50 | 12 | DETR | COCO | 52.1 | [config](configs/codino/co_dino_5scale_r50_8xb2_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/codetr/co_dino_5scale_r50_1x_coco-7481f903.pth) |
+| Co-DINO\* | R50 | 36 | LSJ | COCO | 54.8 | [config](configs/codino/co_dino_5scale_r50_lsj_8xb2_3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/codetr/co_dino_5scale_lsj_r50_3x_coco-fe5a6829.pth) |
+| Co-DINO\* | Swin-L | 12 | DETR | COCO | 58.9 | [config](configs/codino/co_dino_5scale_swin_l_16xb1_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/codetr/co_dino_5scale_swin_large_1x_coco-27c13da4.pth) |
+| Co-DINO\* | Swin-L | 12 | LSJ | COCO | 59.3 | [config](configs/codino/co_dino_5scale_swin_l_lsj_16xb1_1x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/codetr/co_dino_5scale_lsj_swin_large_1x_coco-3af73af2.pth) |
+| Co-DINO\* | Swin-L | 36 | DETR | COCO | 60.0 | [config](configs/codino/co_dino_5scale_swin_l_16xb1_3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/codetr/co_dino_5scale_swin_large_3x_coco-d7a6d8af.pth) |
+| Co-DINO\* | Swin-L | 36 | LSJ | COCO | 60.7 | [config](configs/codino/co_dino_5scale_swin_l_lsj_16xb1_3x_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/codetr/co_dino_5scale_lsj_swin_large_1x_coco-3af73af2.pth) |
+| Co-DINO\* | Swin-L | 16 | DETR | Objects365 pre-trained + COCO | 64.1 | [config](configs/codino/co_dino_5scale_swin_l_16xb1_16e_o365tococo.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/codetr/co_dino_5scale_swin_large_16e_o365tococo-614254c9.pth) |
+
+Note
+
+- Models labeled * are not trained by us, but from [CO-DETR](https://github.com/Sense-X/Co-DETR) official website.
+- We find that the performance is unstable and may fluctuate by about 0.3 mAP.
+- If you want to save GPU memory by enabling checkpointing, please use the `pip install fairscale` command.
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/codetr/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/codetr/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..2ca4c02d9f7b71643b3b63ef4df254b87d4f9661
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/codetr/__init__.py
@@ -0,0 +1,13 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .co_atss_head import CoATSSHead
+from .co_dino_head import CoDINOHead
+from .co_roi_head import CoStandardRoIHead
+from .codetr import CoDETR
+from .transformer import (CoDinoTransformer, DetrTransformerDecoderLayer,
+ DetrTransformerEncoder, DinoTransformerDecoder)
+
+__all__ = [
+ 'CoDETR', 'CoDinoTransformer', 'DinoTransformerDecoder', 'CoDINOHead',
+ 'CoATSSHead', 'CoStandardRoIHead', 'DetrTransformerEncoder',
+ 'DetrTransformerDecoderLayer'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/codetr/co_atss_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/codetr/co_atss_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..c6ae0180da7be292b67a5bb83c1ad34b848ff17a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/codetr/co_atss_head.py
@@ -0,0 +1,153 @@
+from typing import List
+
+import torch
+from torch import Tensor
+
+from mmdet.models.dense_heads import ATSSHead
+from mmdet.models.utils import images_to_levels, multi_apply
+from mmdet.registry import MODELS
+from mmdet.utils import InstanceList, OptInstanceList, reduce_mean
+
+
+@MODELS.register_module()
+class CoATSSHead(ATSSHead):
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ centernesses: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None) -> dict:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level
+ Has shape (N, num_anchors * num_classes, H, W)
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level with shape (N, num_anchors * 4, H, W)
+ centernesses (list[Tensor]): Centerness for each scale
+ level with shape (N, num_anchors * 1, H, W)
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], Optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ assert len(featmap_sizes) == self.prior_generator.num_levels
+
+ device = cls_scores[0].device
+ anchor_list, valid_flag_list = self.get_anchors(
+ featmap_sizes, batch_img_metas, device=device)
+
+ cls_reg_targets = self.get_targets(
+ anchor_list,
+ valid_flag_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore=batch_gt_instances_ignore)
+
+ (anchor_list, labels_list, label_weights_list, bbox_targets_list,
+ bbox_weights_list, avg_factor, ori_anchors, ori_labels,
+ ori_bbox_targets) = cls_reg_targets
+
+ avg_factor = reduce_mean(
+ torch.tensor(avg_factor, dtype=torch.float, device=device)).item()
+
+ losses_cls, losses_bbox, loss_centerness, \
+ bbox_avg_factor = multi_apply(
+ self.loss_by_feat_single,
+ anchor_list,
+ cls_scores,
+ bbox_preds,
+ centernesses,
+ labels_list,
+ label_weights_list,
+ bbox_targets_list,
+ avg_factor=avg_factor)
+
+ bbox_avg_factor = sum(bbox_avg_factor)
+ bbox_avg_factor = reduce_mean(bbox_avg_factor).clamp_(min=1).item()
+ losses_bbox = list(map(lambda x: x / bbox_avg_factor, losses_bbox))
+
+ # diff
+ pos_coords = (ori_anchors, ori_labels, ori_bbox_targets, 'atss')
+ return dict(
+ loss_cls=losses_cls,
+ loss_bbox=losses_bbox,
+ loss_centerness=loss_centerness,
+ pos_coords=pos_coords)
+
+ def get_targets(self,
+ anchor_list: List[List[Tensor]],
+ valid_flag_list: List[List[Tensor]],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None,
+ unmap_outputs: bool = True) -> tuple:
+ """Get targets for ATSS head.
+
+ This method is almost the same as `AnchorHead.get_targets()`. Besides
+ returning the targets as the parent method does, it also returns the
+ anchors as the first element of the returned tuple.
+ """
+ num_imgs = len(batch_img_metas)
+ assert len(anchor_list) == len(valid_flag_list) == num_imgs
+
+ # anchor number of multi levels
+ num_level_anchors = [anchors.size(0) for anchors in anchor_list[0]]
+ num_level_anchors_list = [num_level_anchors] * num_imgs
+
+ # concat all level anchors and flags to a single tensor
+ for i in range(num_imgs):
+ assert len(anchor_list[i]) == len(valid_flag_list[i])
+ anchor_list[i] = torch.cat(anchor_list[i])
+ valid_flag_list[i] = torch.cat(valid_flag_list[i])
+
+ # compute targets for each image
+ if batch_gt_instances_ignore is None:
+ batch_gt_instances_ignore = [None] * num_imgs
+ (all_anchors, all_labels, all_label_weights, all_bbox_targets,
+ all_bbox_weights, pos_inds_list, neg_inds_list,
+ sampling_results_list) = multi_apply(
+ self._get_targets_single,
+ anchor_list,
+ valid_flag_list,
+ num_level_anchors_list,
+ batch_gt_instances,
+ batch_img_metas,
+ batch_gt_instances_ignore,
+ unmap_outputs=unmap_outputs)
+ # Get `avg_factor` of all images, which calculate in `SamplingResult`.
+ # When using sampling method, avg_factor is usually the sum of
+ # positive and negative priors. When using `PseudoSampler`,
+ # `avg_factor` is usually equal to the number of positive priors.
+ avg_factor = sum(
+ [results.avg_factor for results in sampling_results_list])
+ # split targets to a list w.r.t. multiple levels
+ anchors_list = images_to_levels(all_anchors, num_level_anchors)
+ labels_list = images_to_levels(all_labels, num_level_anchors)
+ label_weights_list = images_to_levels(all_label_weights,
+ num_level_anchors)
+ bbox_targets_list = images_to_levels(all_bbox_targets,
+ num_level_anchors)
+ bbox_weights_list = images_to_levels(all_bbox_weights,
+ num_level_anchors)
+
+ # diff
+ ori_anchors = all_anchors
+ ori_labels = all_labels
+ ori_bbox_targets = all_bbox_targets
+ return (anchors_list, labels_list, label_weights_list,
+ bbox_targets_list, bbox_weights_list, avg_factor, ori_anchors,
+ ori_labels, ori_bbox_targets)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/codetr/co_dino_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/codetr/co_dino_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..192acf97d86c5d24b623608a46d564a8753b5b7b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/codetr/co_dino_head.py
@@ -0,0 +1,677 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from typing import List
+
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmcv.cnn import Linear
+from mmcv.ops import batched_nms
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.models import DINOHead
+from mmdet.models.layers import CdnQueryGenerator
+from mmdet.models.layers.transformer import inverse_sigmoid
+from mmdet.models.utils import multi_apply
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import (bbox_cxcywh_to_xyxy, bbox_overlaps,
+ bbox_xyxy_to_cxcywh)
+from mmdet.utils import InstanceList, reduce_mean
+
+
+@MODELS.register_module()
+class CoDINOHead(DINOHead):
+
+ def __init__(self,
+ *args,
+ num_query=900,
+ transformer=None,
+ in_channels=2048,
+ max_pos_coords=300,
+ dn_cfg=None,
+ use_zero_padding=False,
+ positional_encoding=dict(
+ type='SinePositionalEncoding',
+ num_feats=128,
+ normalize=True),
+ **kwargs):
+ self.with_box_refine = True
+ self.mixed_selection = True
+ self.in_channels = in_channels
+ self.max_pos_coords = max_pos_coords
+ self.positional_encoding = positional_encoding
+ self.num_query = num_query
+ self.use_zero_padding = use_zero_padding
+
+ if 'two_stage_num_proposals' in transformer:
+ assert transformer['two_stage_num_proposals'] == num_query, \
+ 'two_stage_num_proposals must be equal to num_query for DINO'
+ else:
+ transformer['two_stage_num_proposals'] = num_query
+ transformer['as_two_stage'] = True
+ if self.mixed_selection:
+ transformer['mixed_selection'] = self.mixed_selection
+ self.transformer = transformer
+ self.act_cfg = transformer.get('act_cfg',
+ dict(type='ReLU', inplace=True))
+
+ super().__init__(*args, **kwargs)
+
+ self.activate = MODELS.build(self.act_cfg)
+ self.positional_encoding = MODELS.build(self.positional_encoding)
+ self.init_denoising(dn_cfg)
+
+ def _init_layers(self):
+ self.transformer = MODELS.build(self.transformer)
+ self.embed_dims = self.transformer.embed_dims
+ assert hasattr(self.positional_encoding, 'num_feats')
+ num_feats = self.positional_encoding.num_feats
+ assert num_feats * 2 == self.embed_dims, 'embed_dims should' \
+ f' be exactly 2 times of num_feats. Found {self.embed_dims}' \
+ f' and {num_feats}.'
+ """Initialize classification branch and regression branch of head."""
+ fc_cls = Linear(self.embed_dims, self.cls_out_channels)
+ reg_branch = []
+ for _ in range(self.num_reg_fcs):
+ reg_branch.append(Linear(self.embed_dims, self.embed_dims))
+ reg_branch.append(nn.ReLU())
+ reg_branch.append(Linear(self.embed_dims, 4))
+ reg_branch = nn.Sequential(*reg_branch)
+
+ def _get_clones(module, N):
+ return nn.ModuleList([copy.deepcopy(module) for i in range(N)])
+
+ # last reg_branch is used to generate proposal from
+ # encode feature map when as_two_stage is True.
+ num_pred = (self.transformer.decoder.num_layers + 1) if \
+ self.as_two_stage else self.transformer.decoder.num_layers
+
+ self.cls_branches = _get_clones(fc_cls, num_pred)
+ self.reg_branches = _get_clones(reg_branch, num_pred)
+
+ self.downsample = nn.Sequential(
+ nn.Conv2d(
+ self.embed_dims,
+ self.embed_dims,
+ kernel_size=3,
+ stride=2,
+ padding=1), nn.GroupNorm(32, self.embed_dims))
+
+ def init_denoising(self, dn_cfg):
+ if dn_cfg is not None:
+ dn_cfg['num_classes'] = self.num_classes
+ dn_cfg['num_matching_queries'] = self.num_query
+ dn_cfg['embed_dims'] = self.embed_dims
+ self.dn_generator = CdnQueryGenerator(**dn_cfg)
+
+ def forward(self,
+ mlvl_feats,
+ img_metas,
+ dn_label_query=None,
+ dn_bbox_query=None,
+ attn_mask=None):
+ batch_size = mlvl_feats[0].size(0)
+ input_img_h, input_img_w = img_metas[0]['batch_input_shape']
+ img_masks = mlvl_feats[0].new_ones(
+ (batch_size, input_img_h, input_img_w))
+ for img_id in range(batch_size):
+ img_h, img_w = img_metas[img_id]['img_shape']
+ img_masks[img_id, :img_h, :img_w] = 0
+
+ mlvl_masks = []
+ mlvl_positional_encodings = []
+ for feat in mlvl_feats:
+ mlvl_masks.append(
+ F.interpolate(img_masks[None],
+ size=feat.shape[-2:]).to(torch.bool).squeeze(0))
+ mlvl_positional_encodings.append(
+ self.positional_encoding(mlvl_masks[-1]))
+
+ query_embeds = None
+ hs, inter_references, topk_score, topk_anchor, enc_outputs = \
+ self.transformer(
+ mlvl_feats,
+ mlvl_masks,
+ query_embeds,
+ mlvl_positional_encodings,
+ dn_label_query,
+ dn_bbox_query,
+ attn_mask,
+ reg_branches=self.reg_branches if self.with_box_refine else None, # noqa:E501
+ cls_branches=self.cls_branches if self.as_two_stage else None # noqa:E501
+ )
+ outs = []
+ num_level = len(mlvl_feats)
+ start = 0
+ for lvl in range(num_level):
+ bs, c, h, w = mlvl_feats[lvl].shape
+ end = start + h * w
+ feat = enc_outputs[start:end].permute(1, 2, 0).contiguous()
+ start = end
+ outs.append(feat.reshape(bs, c, h, w))
+ outs.append(self.downsample(outs[-1]))
+
+ hs = hs.permute(0, 2, 1, 3)
+
+ if dn_label_query is not None and dn_label_query.size(1) == 0:
+ # NOTE: If there is no target in the image, the parameters of
+ # label_embedding won't be used in producing loss, which raises
+ # RuntimeError when using distributed mode.
+ hs[0] += self.dn_generator.label_embedding.weight[0, 0] * 0.0
+
+ outputs_classes = []
+ outputs_coords = []
+
+ for lvl in range(hs.shape[0]):
+ reference = inter_references[lvl]
+ reference = inverse_sigmoid(reference, eps=1e-3)
+ outputs_class = self.cls_branches[lvl](hs[lvl])
+ tmp = self.reg_branches[lvl](hs[lvl])
+ if reference.shape[-1] == 4:
+ tmp += reference
+ else:
+ assert reference.shape[-1] == 2
+ tmp[..., :2] += reference
+ outputs_coord = tmp.sigmoid()
+ outputs_classes.append(outputs_class)
+ outputs_coords.append(outputs_coord)
+
+ outputs_classes = torch.stack(outputs_classes)
+ outputs_coords = torch.stack(outputs_coords)
+
+ return outputs_classes, outputs_coords, topk_score, topk_anchor, outs
+
+ def predict(self,
+ feats: List[Tensor],
+ batch_data_samples: SampleList,
+ rescale: bool = True) -> InstanceList:
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+ outs = self.forward(feats, batch_img_metas)
+
+ predictions = self.predict_by_feat(
+ *outs, batch_img_metas=batch_img_metas, rescale=rescale)
+
+ return predictions
+
+ def predict_by_feat(self,
+ all_cls_scores,
+ all_bbox_preds,
+ enc_cls_scores,
+ enc_bbox_preds,
+ enc_outputs,
+ batch_img_metas,
+ rescale=True):
+
+ cls_scores = all_cls_scores[-1]
+ bbox_preds = all_bbox_preds[-1]
+
+ result_list = []
+ for img_id in range(len(batch_img_metas)):
+ cls_score = cls_scores[img_id]
+ bbox_pred = bbox_preds[img_id]
+ img_meta = batch_img_metas[img_id]
+ results = self._predict_by_feat_single(cls_score, bbox_pred,
+ img_meta, rescale)
+ result_list.append(results)
+ return result_list
+
+ def _predict_by_feat_single(self,
+ cls_score: Tensor,
+ bbox_pred: Tensor,
+ img_meta: dict,
+ rescale: bool = True) -> InstanceData:
+ """Transform outputs from the last decoder layer into bbox predictions
+ for each image.
+
+ Args:
+ cls_score (Tensor): Box score logits from the last decoder layer
+ for each image. Shape [num_queries, cls_out_channels].
+ bbox_pred (Tensor): Sigmoid outputs from the last decoder layer
+ for each image, with coordinate format (cx, cy, w, h) and
+ shape [num_queries, 4].
+ img_meta (dict): Image meta info.
+ rescale (bool): If True, return boxes in original image
+ space. Default True.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ assert len(cls_score) == len(bbox_pred) # num_queries
+ max_per_img = self.test_cfg.get('max_per_img', self.num_query)
+ score_thr = self.test_cfg.get('score_thr', 0)
+ with_nms = self.test_cfg.get('nms', None)
+
+ img_shape = img_meta['img_shape']
+ # exclude background
+ if self.loss_cls.use_sigmoid:
+ cls_score = cls_score.sigmoid()
+ scores, indexes = cls_score.view(-1).topk(max_per_img)
+ det_labels = indexes % self.num_classes
+ bbox_index = indexes // self.num_classes
+ bbox_pred = bbox_pred[bbox_index]
+ else:
+ scores, det_labels = F.softmax(cls_score, dim=-1)[..., :-1].max(-1)
+ scores, bbox_index = scores.topk(max_per_img)
+ bbox_pred = bbox_pred[bbox_index]
+ det_labels = det_labels[bbox_index]
+
+ if score_thr > 0:
+ valid_mask = scores > score_thr
+ scores = scores[valid_mask]
+ bbox_pred = bbox_pred[valid_mask]
+ det_labels = det_labels[valid_mask]
+
+ det_bboxes = bbox_cxcywh_to_xyxy(bbox_pred)
+ det_bboxes[:, 0::2] = det_bboxes[:, 0::2] * img_shape[1]
+ det_bboxes[:, 1::2] = det_bboxes[:, 1::2] * img_shape[0]
+ det_bboxes[:, 0::2].clamp_(min=0, max=img_shape[1])
+ det_bboxes[:, 1::2].clamp_(min=0, max=img_shape[0])
+ if rescale:
+ assert img_meta.get('scale_factor') is not None
+ det_bboxes /= det_bboxes.new_tensor(
+ img_meta['scale_factor']).repeat((1, 2))
+
+ results = InstanceData()
+ results.bboxes = det_bboxes
+ results.scores = scores
+ results.labels = det_labels
+
+ if with_nms and results.bboxes.numel() > 0:
+ det_bboxes, keep_idxs = batched_nms(results.bboxes, results.scores,
+ results.labels,
+ self.test_cfg.nms)
+ results = results[keep_idxs]
+ results.scores = det_bboxes[:, -1]
+ results = results[:max_per_img]
+
+ return results
+
+ def loss(self, x, batch_data_samples):
+ assert self.dn_generator is not None, '"dn_cfg" must be set'
+
+ batch_gt_instances = []
+ batch_img_metas = []
+ for data_sample in batch_data_samples:
+ batch_img_metas.append(data_sample.metainfo)
+ batch_gt_instances.append(data_sample.gt_instances)
+
+ dn_label_query, dn_bbox_query, attn_mask, dn_meta = \
+ self.dn_generator(batch_data_samples)
+
+ outs = self(x, batch_img_metas, dn_label_query, dn_bbox_query,
+ attn_mask)
+
+ loss_inputs = outs[:-1] + (batch_gt_instances, batch_img_metas,
+ dn_meta)
+ losses = self.loss_by_feat(*loss_inputs)
+ enc_outputs = outs[-1]
+ return losses, enc_outputs
+
+ def forward_aux(self, mlvl_feats, img_metas, aux_targets, head_idx):
+ """Forward function.
+
+ Args:
+ mlvl_feats (tuple[Tensor]): Features from the upstream
+ network, each is a 4D-tensor with shape
+ (N, C, H, W).
+ img_metas (list[dict]): List of image information.
+
+ Returns:
+ all_cls_scores (Tensor): Outputs from the classification head, \
+ shape [nb_dec, bs, num_query, cls_out_channels]. Note \
+ cls_out_channels should includes background.
+ all_bbox_preds (Tensor): Sigmoid outputs from the regression \
+ head with normalized coordinate format (cx, cy, w, h). \
+ Shape [nb_dec, bs, num_query, 4].
+ enc_outputs_class (Tensor): The score of each point on encode \
+ feature map, has shape (N, h*w, num_class). Only when \
+ as_two_stage is True it would be returned, otherwise \
+ `None` would be returned.
+ enc_outputs_coord (Tensor): The proposal generate from the \
+ encode feature map, has shape (N, h*w, 4). Only when \
+ as_two_stage is True it would be returned, otherwise \
+ `None` would be returned.
+ """
+ aux_coords, aux_labels, aux_targets, aux_label_weights, \
+ aux_bbox_weights, aux_feats, attn_masks = aux_targets
+ batch_size = mlvl_feats[0].size(0)
+ input_img_h, input_img_w = img_metas[0]['batch_input_shape']
+ img_masks = mlvl_feats[0].new_ones(
+ (batch_size, input_img_h, input_img_w))
+ for img_id in range(batch_size):
+ img_h, img_w = img_metas[img_id]['img_shape']
+ img_masks[img_id, :img_h, :img_w] = 0
+
+ mlvl_masks = []
+ mlvl_positional_encodings = []
+ for feat in mlvl_feats:
+ mlvl_masks.append(
+ F.interpolate(img_masks[None],
+ size=feat.shape[-2:]).to(torch.bool).squeeze(0))
+ mlvl_positional_encodings.append(
+ self.positional_encoding(mlvl_masks[-1]))
+
+ query_embeds = None
+ hs, inter_references = self.transformer.forward_aux(
+ mlvl_feats,
+ mlvl_masks,
+ query_embeds,
+ mlvl_positional_encodings,
+ aux_coords,
+ pos_feats=aux_feats,
+ reg_branches=self.reg_branches if self.with_box_refine else None,
+ cls_branches=self.cls_branches if self.as_two_stage else None,
+ return_encoder_output=True,
+ attn_masks=attn_masks,
+ head_idx=head_idx)
+
+ hs = hs.permute(0, 2, 1, 3)
+ outputs_classes = []
+ outputs_coords = []
+
+ for lvl in range(hs.shape[0]):
+ reference = inter_references[lvl]
+ reference = inverse_sigmoid(reference, eps=1e-3)
+ outputs_class = self.cls_branches[lvl](hs[lvl])
+ tmp = self.reg_branches[lvl](hs[lvl])
+ if reference.shape[-1] == 4:
+ tmp += reference
+ else:
+ assert reference.shape[-1] == 2
+ tmp[..., :2] += reference
+ outputs_coord = tmp.sigmoid()
+ outputs_classes.append(outputs_class)
+ outputs_coords.append(outputs_coord)
+
+ outputs_classes = torch.stack(outputs_classes)
+ outputs_coords = torch.stack(outputs_coords)
+
+ return outputs_classes, outputs_coords, None, None
+
+ def loss_aux(self,
+ x,
+ pos_coords=None,
+ head_idx=0,
+ batch_data_samples=None):
+ batch_gt_instances = []
+ batch_img_metas = []
+ for data_sample in batch_data_samples:
+ batch_img_metas.append(data_sample.metainfo)
+ batch_gt_instances.append(data_sample.gt_instances)
+
+ gt_bboxes = [b.bboxes for b in batch_gt_instances]
+ gt_labels = [b.labels for b in batch_gt_instances]
+
+ aux_targets = self.get_aux_targets(pos_coords, batch_img_metas, x,
+ head_idx)
+ outs = self.forward_aux(x[:-1], batch_img_metas, aux_targets, head_idx)
+ outs = outs + aux_targets
+ if gt_labels is None:
+ loss_inputs = outs + (gt_bboxes, batch_img_metas)
+ else:
+ loss_inputs = outs + (gt_bboxes, gt_labels, batch_img_metas)
+ losses = self.loss_aux_by_feat(*loss_inputs)
+ return losses
+
+ def get_aux_targets(self, pos_coords, img_metas, mlvl_feats, head_idx):
+ coords, labels, targets = pos_coords[:3]
+ head_name = pos_coords[-1]
+ bs, c = len(coords), mlvl_feats[0].shape[1]
+ max_num_coords = 0
+ all_feats = []
+ for i in range(bs):
+ label = labels[i]
+ feats = [
+ feat[i].reshape(c, -1).transpose(1, 0) for feat in mlvl_feats
+ ]
+ feats = torch.cat(feats, dim=0)
+ bg_class_ind = self.num_classes
+ pos_inds = ((label >= 0)
+ & (label < bg_class_ind)).nonzero().squeeze(1)
+ max_num_coords = max(max_num_coords, len(pos_inds))
+ all_feats.append(feats)
+ max_num_coords = min(self.max_pos_coords, max_num_coords)
+ max_num_coords = max(9, max_num_coords)
+
+ if self.use_zero_padding:
+ attn_masks = []
+ label_weights = coords[0].new_zeros([bs, max_num_coords])
+ else:
+ attn_masks = None
+ label_weights = coords[0].new_ones([bs, max_num_coords])
+ bbox_weights = coords[0].new_zeros([bs, max_num_coords, 4])
+
+ aux_coords, aux_labels, aux_targets, aux_feats = [], [], [], []
+
+ for i in range(bs):
+ coord, label, target = coords[i], labels[i], targets[i]
+ feats = all_feats[i]
+ if 'rcnn' in head_name:
+ feats = pos_coords[-2][i]
+ num_coords_per_point = 1
+ else:
+ num_coords_per_point = coord.shape[0] // feats.shape[0]
+ feats = feats.unsqueeze(1).repeat(1, num_coords_per_point, 1)
+ feats = feats.reshape(feats.shape[0] * num_coords_per_point,
+ feats.shape[-1])
+ img_meta = img_metas[i]
+ img_h, img_w = img_meta['img_shape']
+ factor = coord.new_tensor([img_w, img_h, img_w,
+ img_h]).unsqueeze(0)
+ bg_class_ind = self.num_classes
+ pos_inds = ((label >= 0)
+ & (label < bg_class_ind)).nonzero().squeeze(1)
+ neg_inds = (label == bg_class_ind).nonzero().squeeze(1)
+ if pos_inds.shape[0] > max_num_coords:
+ indices = torch.randperm(
+ pos_inds.shape[0])[:max_num_coords].cuda()
+ pos_inds = pos_inds[indices]
+
+ coord = bbox_xyxy_to_cxcywh(coord[pos_inds] / factor)
+ label = label[pos_inds]
+ target = bbox_xyxy_to_cxcywh(target[pos_inds] / factor)
+ feat = feats[pos_inds]
+
+ if self.use_zero_padding:
+ label_weights[i][:len(label)] = 1
+ bbox_weights[i][:len(label)] = 1
+ attn_mask = torch.zeros([
+ max_num_coords,
+ max_num_coords,
+ ]).bool().to(coord.device)
+ else:
+ bbox_weights[i][:len(label)] = 1
+
+ if coord.shape[0] < max_num_coords:
+ padding_shape = max_num_coords - coord.shape[0]
+ if self.use_zero_padding:
+ padding_coord = coord.new_zeros([padding_shape, 4])
+ padding_label = label.new_ones([padding_shape
+ ]) * self.num_classes
+ padding_target = target.new_zeros([padding_shape, 4])
+ padding_feat = feat.new_zeros([padding_shape, c])
+ attn_mask[coord.shape[0]:, 0:coord.shape[0], ] = True
+ attn_mask[:, coord.shape[0]:, ] = True
+ else:
+ indices = torch.randperm(
+ neg_inds.shape[0])[:padding_shape].cuda()
+ neg_inds = neg_inds[indices]
+ padding_coord = bbox_xyxy_to_cxcywh(coords[i][neg_inds] /
+ factor)
+ padding_label = labels[i][neg_inds]
+ padding_target = bbox_xyxy_to_cxcywh(targets[i][neg_inds] /
+ factor)
+ padding_feat = feats[neg_inds]
+ coord = torch.cat((coord, padding_coord), dim=0)
+ label = torch.cat((label, padding_label), dim=0)
+ target = torch.cat((target, padding_target), dim=0)
+ feat = torch.cat((feat, padding_feat), dim=0)
+ if self.use_zero_padding:
+ attn_masks.append(attn_mask.unsqueeze(0))
+ aux_coords.append(coord.unsqueeze(0))
+ aux_labels.append(label.unsqueeze(0))
+ aux_targets.append(target.unsqueeze(0))
+ aux_feats.append(feat.unsqueeze(0))
+
+ if self.use_zero_padding:
+ attn_masks = torch.cat(
+ attn_masks, dim=0).unsqueeze(1).repeat(1, 8, 1, 1)
+ attn_masks = attn_masks.reshape(bs * 8, max_num_coords,
+ max_num_coords)
+ else:
+ attn_masks = None
+
+ aux_coords = torch.cat(aux_coords, dim=0)
+ aux_labels = torch.cat(aux_labels, dim=0)
+ aux_targets = torch.cat(aux_targets, dim=0)
+ aux_feats = torch.cat(aux_feats, dim=0)
+ aux_label_weights = label_weights
+ aux_bbox_weights = bbox_weights
+ return (aux_coords, aux_labels, aux_targets, aux_label_weights,
+ aux_bbox_weights, aux_feats, attn_masks)
+
+ def loss_aux_by_feat(self,
+ all_cls_scores,
+ all_bbox_preds,
+ enc_cls_scores,
+ enc_bbox_preds,
+ aux_coords,
+ aux_labels,
+ aux_targets,
+ aux_label_weights,
+ aux_bbox_weights,
+ aux_feats,
+ attn_masks,
+ gt_bboxes_list,
+ gt_labels_list,
+ img_metas,
+ gt_bboxes_ignore=None):
+ num_dec_layers = len(all_cls_scores)
+ all_labels = [aux_labels for _ in range(num_dec_layers)]
+ all_label_weights = [aux_label_weights for _ in range(num_dec_layers)]
+ all_bbox_targets = [aux_targets for _ in range(num_dec_layers)]
+ all_bbox_weights = [aux_bbox_weights for _ in range(num_dec_layers)]
+ img_metas_list = [img_metas for _ in range(num_dec_layers)]
+ all_gt_bboxes_ignore_list = [
+ gt_bboxes_ignore for _ in range(num_dec_layers)
+ ]
+
+ losses_cls, losses_bbox, losses_iou = multi_apply(
+ self._loss_aux_by_feat_single, all_cls_scores, all_bbox_preds,
+ all_labels, all_label_weights, all_bbox_targets, all_bbox_weights,
+ img_metas_list, all_gt_bboxes_ignore_list)
+
+ loss_dict = dict()
+ # loss of proposal generated from encode feature map.
+
+ # loss from the last decoder layer
+ loss_dict['loss_cls_aux'] = losses_cls[-1]
+ loss_dict['loss_bbox_aux'] = losses_bbox[-1]
+ loss_dict['loss_iou_aux'] = losses_iou[-1]
+ # loss from other decoder layers
+ num_dec_layer = 0
+ for loss_cls_i, loss_bbox_i, loss_iou_i in zip(losses_cls[:-1],
+ losses_bbox[:-1],
+ losses_iou[:-1]):
+ loss_dict[f'd{num_dec_layer}.loss_cls_aux'] = loss_cls_i
+ loss_dict[f'd{num_dec_layer}.loss_bbox_aux'] = loss_bbox_i
+ loss_dict[f'd{num_dec_layer}.loss_iou_aux'] = loss_iou_i
+ num_dec_layer += 1
+ return loss_dict
+
+ def _loss_aux_by_feat_single(self,
+ cls_scores,
+ bbox_preds,
+ labels,
+ label_weights,
+ bbox_targets,
+ bbox_weights,
+ img_metas,
+ gt_bboxes_ignore_list=None):
+ num_imgs = cls_scores.size(0)
+ num_q = cls_scores.size(1)
+
+ try:
+ labels = labels.reshape(num_imgs * num_q)
+ label_weights = label_weights.reshape(num_imgs * num_q)
+ bbox_targets = bbox_targets.reshape(num_imgs * num_q, 4)
+ bbox_weights = bbox_weights.reshape(num_imgs * num_q, 4)
+ except Exception:
+ return cls_scores.mean() * 0, cls_scores.mean(
+ ) * 0, cls_scores.mean() * 0
+
+ bg_class_ind = self.num_classes
+ num_total_pos = len(
+ ((labels >= 0) & (labels < bg_class_ind)).nonzero().squeeze(1))
+ num_total_neg = num_imgs * num_q - num_total_pos
+
+ # classification loss
+ cls_scores = cls_scores.reshape(-1, self.cls_out_channels)
+ # construct weighted avg_factor to match with the official DETR repo
+ cls_avg_factor = num_total_pos * 1.0 + \
+ num_total_neg * self.bg_cls_weight
+ if self.sync_cls_avg_factor:
+ cls_avg_factor = reduce_mean(
+ cls_scores.new_tensor([cls_avg_factor]))
+ cls_avg_factor = max(cls_avg_factor, 1)
+
+ bg_class_ind = self.num_classes
+ pos_inds = ((labels >= 0)
+ & (labels < bg_class_ind)).nonzero().squeeze(1)
+ scores = label_weights.new_zeros(labels.shape)
+ pos_bbox_targets = bbox_targets[pos_inds]
+ pos_decode_bbox_targets = bbox_cxcywh_to_xyxy(pos_bbox_targets)
+ pos_bbox_pred = bbox_preds.reshape(-1, 4)[pos_inds]
+ pos_decode_bbox_pred = bbox_cxcywh_to_xyxy(pos_bbox_pred)
+ scores[pos_inds] = bbox_overlaps(
+ pos_decode_bbox_pred.detach(),
+ pos_decode_bbox_targets,
+ is_aligned=True)
+ loss_cls = self.loss_cls(
+ cls_scores, (labels, scores),
+ weight=label_weights,
+ avg_factor=cls_avg_factor)
+
+ # Compute the average number of gt boxes across all gpus, for
+ # normalization purposes
+ num_total_pos = loss_cls.new_tensor([num_total_pos])
+ num_total_pos = torch.clamp(reduce_mean(num_total_pos), min=1).item()
+
+ # construct factors used for rescale bboxes
+ factors = []
+ for img_meta, bbox_pred in zip(img_metas, bbox_preds):
+ img_h, img_w = img_meta['img_shape']
+ factor = bbox_pred.new_tensor([img_w, img_h, img_w,
+ img_h]).unsqueeze(0).repeat(
+ bbox_pred.size(0), 1)
+ factors.append(factor)
+ factors = torch.cat(factors, 0)
+
+ # DETR regress the relative position of boxes (cxcywh) in the image,
+ # thus the learning target is normalized by the image size. So here
+ # we need to re-scale them for calculating IoU loss
+ bbox_preds = bbox_preds.reshape(-1, 4)
+ bboxes = bbox_cxcywh_to_xyxy(bbox_preds) * factors
+ bboxes_gt = bbox_cxcywh_to_xyxy(bbox_targets) * factors
+
+ # regression IoU loss, defaultly GIoU loss
+ loss_iou = self.loss_iou(
+ bboxes, bboxes_gt, bbox_weights, avg_factor=num_total_pos)
+
+ # regression L1 loss
+ loss_bbox = self.loss_bbox(
+ bbox_preds, bbox_targets, bbox_weights, avg_factor=num_total_pos)
+ return loss_cls, loss_bbox, loss_iou
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/codetr/co_roi_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/codetr/co_roi_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..9aafb53beddf07428e59d83e9de832ff5102821a
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/codetr/co_roi_head.py
@@ -0,0 +1,108 @@
+from typing import List, Tuple
+
+import torch
+from torch import Tensor
+
+from mmdet.models.roi_heads import StandardRoIHead
+from mmdet.models.task_modules.samplers import SamplingResult
+from mmdet.models.utils import unpack_gt_instances
+from mmdet.registry import MODELS
+from mmdet.structures import DetDataSample
+from mmdet.structures.bbox import bbox2roi
+from mmdet.utils import InstanceList
+
+
+@MODELS.register_module()
+class CoStandardRoIHead(StandardRoIHead):
+
+ def loss(self, x: Tuple[Tensor], rpn_results_list: InstanceList,
+ batch_data_samples: List[DetDataSample]) -> dict:
+ max_proposal = 2000
+
+ assert len(rpn_results_list) == len(batch_data_samples)
+ outputs = unpack_gt_instances(batch_data_samples)
+ batch_gt_instances, batch_gt_instances_ignore, _ = outputs
+
+ # assign gts and sample proposals
+ num_imgs = len(batch_data_samples)
+ sampling_results = []
+ for i in range(num_imgs):
+ # rename rpn_results.bboxes to rpn_results.priors
+ rpn_results = rpn_results_list[i]
+ rpn_results.priors = rpn_results.pop('bboxes')
+
+ assign_result = self.bbox_assigner.assign(
+ rpn_results, batch_gt_instances[i],
+ batch_gt_instances_ignore[i])
+ sampling_result = self.bbox_sampler.sample(
+ assign_result,
+ rpn_results,
+ batch_gt_instances[i],
+ feats=[lvl_feat[i][None] for lvl_feat in x])
+ sampling_results.append(sampling_result)
+
+ losses = dict()
+ # bbox head forward and loss
+ if self.with_bbox:
+ bbox_results = self.bbox_loss(x, sampling_results)
+ losses.update(bbox_results['loss_bbox'])
+
+ bbox_targets = bbox_results['bbox_targets']
+ for res in sampling_results:
+ max_proposal = min(max_proposal, res.bboxes.shape[0])
+ ori_coords = bbox2roi([res.bboxes for res in sampling_results])
+ ori_proposals, ori_labels, \
+ ori_bbox_targets, ori_bbox_feats = [], [], [], []
+ for i in range(num_imgs):
+ idx = (ori_coords[:, 0] == i).nonzero().squeeze(1)
+ idx = idx[:max_proposal]
+ ori_proposal = ori_coords[idx][:, 1:].unsqueeze(0)
+ ori_label = bbox_targets[0][idx].unsqueeze(0)
+ ori_bbox_target = bbox_targets[2][idx].unsqueeze(0)
+ ori_bbox_feat = bbox_results['bbox_feats'].mean(-1).mean(-1)
+ ori_bbox_feat = ori_bbox_feat[idx].unsqueeze(0)
+ ori_proposals.append(ori_proposal)
+ ori_labels.append(ori_label)
+ ori_bbox_targets.append(ori_bbox_target)
+ ori_bbox_feats.append(ori_bbox_feat)
+ ori_coords = torch.cat(ori_proposals, dim=0)
+ ori_labels = torch.cat(ori_labels, dim=0)
+ ori_bbox_targets = torch.cat(ori_bbox_targets, dim=0)
+ ori_bbox_feats = torch.cat(ori_bbox_feats, dim=0)
+ pos_coords = (ori_coords, ori_labels, ori_bbox_targets,
+ ori_bbox_feats, 'rcnn')
+ losses.update(pos_coords=pos_coords)
+
+ return losses
+
+ def bbox_loss(self, x: Tuple[Tensor],
+ sampling_results: List[SamplingResult]) -> dict:
+ """Perform forward propagation and loss calculation of the bbox head on
+ the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ sampling_results (list["obj:`SamplingResult`]): Sampling results.
+
+ Returns:
+ dict[str, Tensor]: Usually returns a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `bbox_feats` (Tensor): Extract bbox RoI features.
+ - `loss_bbox` (dict): A dictionary of bbox loss components.
+ """
+ rois = bbox2roi([res.priors for res in sampling_results])
+ bbox_results = self._bbox_forward(x, rois)
+
+ bbox_loss_and_target = self.bbox_head.loss_and_target(
+ cls_score=bbox_results['cls_score'],
+ bbox_pred=bbox_results['bbox_pred'],
+ rois=rois,
+ sampling_results=sampling_results,
+ rcnn_train_cfg=self.train_cfg)
+
+ bbox_results.update(loss_bbox=bbox_loss_and_target['loss_bbox'])
+ # diff
+ bbox_results.update(bbox_targets=bbox_loss_and_target['bbox_targets'])
+ return bbox_results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/codetr/codetr.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/codetr/codetr.py
new file mode 100644
index 0000000000000000000000000000000000000000..82826f641075c0af7eebd322b6b36b53390cc648
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/codetr/codetr.py
@@ -0,0 +1,320 @@
+import copy
+from typing import Tuple, Union
+
+import torch
+import torch.nn as nn
+from torch import Tensor
+
+from mmdet.models.detectors.base import BaseDetector
+from mmdet.registry import MODELS
+from mmdet.structures import OptSampleList, SampleList
+from mmdet.utils import InstanceList, OptConfigType, OptMultiConfig
+
+
+@MODELS.register_module()
+class CoDETR(BaseDetector):
+
+ def __init__(
+ self,
+ backbone,
+ neck=None,
+ query_head=None, # detr head
+ rpn_head=None, # two-stage rpn
+ roi_head=[None], # two-stage
+ bbox_head=[None], # one-stage
+ train_cfg=[None, None],
+ test_cfg=[None, None],
+ # Control whether to consider positive samples
+ # from the auxiliary head as additional positive queries.
+ with_pos_coord=True,
+ use_lsj=True,
+ eval_module='detr',
+ # Evaluate the Nth head.
+ eval_index=0,
+ data_preprocessor: OptConfigType = None,
+ init_cfg: OptMultiConfig = None):
+ super(CoDETR, self).__init__(
+ data_preprocessor=data_preprocessor, init_cfg=init_cfg)
+ self.with_pos_coord = with_pos_coord
+ self.use_lsj = use_lsj
+
+ assert eval_module in ['detr', 'one-stage', 'two-stage']
+ self.eval_module = eval_module
+
+ self.backbone = MODELS.build(backbone)
+ if neck is not None:
+ self.neck = MODELS.build(neck)
+ # Module index for evaluation
+ self.eval_index = eval_index
+ head_idx = 0
+ if query_head is not None:
+ query_head.update(train_cfg=train_cfg[head_idx] if (
+ train_cfg is not None and train_cfg[head_idx] is not None
+ ) else None)
+ query_head.update(test_cfg=test_cfg[head_idx])
+ self.query_head = MODELS.build(query_head)
+ self.query_head.init_weights()
+ head_idx += 1
+
+ if rpn_head is not None:
+ rpn_train_cfg = train_cfg[head_idx].rpn if (
+ train_cfg is not None
+ and train_cfg[head_idx] is not None) else None
+ rpn_head_ = rpn_head.copy()
+ rpn_head_.update(
+ train_cfg=rpn_train_cfg, test_cfg=test_cfg[head_idx].rpn)
+ self.rpn_head = MODELS.build(rpn_head_)
+ self.rpn_head.init_weights()
+
+ self.roi_head = nn.ModuleList()
+ for i in range(len(roi_head)):
+ if roi_head[i]:
+ rcnn_train_cfg = train_cfg[i + head_idx].rcnn if (
+ train_cfg
+ and train_cfg[i + head_idx] is not None) else None
+ roi_head[i].update(train_cfg=rcnn_train_cfg)
+ roi_head[i].update(test_cfg=test_cfg[i + head_idx].rcnn)
+ self.roi_head.append(MODELS.build(roi_head[i]))
+ self.roi_head[-1].init_weights()
+
+ self.bbox_head = nn.ModuleList()
+ for i in range(len(bbox_head)):
+ if bbox_head[i]:
+ bbox_head[i].update(
+ train_cfg=train_cfg[i + head_idx + len(self.roi_head)] if (
+ train_cfg and train_cfg[i + head_idx +
+ len(self.roi_head)] is not None
+ ) else None)
+ bbox_head[i].update(test_cfg=test_cfg[i + head_idx +
+ len(self.roi_head)])
+ self.bbox_head.append(MODELS.build(bbox_head[i]))
+ self.bbox_head[-1].init_weights()
+
+ self.head_idx = head_idx
+ self.train_cfg = train_cfg
+ self.test_cfg = test_cfg
+
+ @property
+ def with_rpn(self):
+ """bool: whether the detector has RPN"""
+ return hasattr(self, 'rpn_head') and self.rpn_head is not None
+
+ @property
+ def with_query_head(self):
+ """bool: whether the detector has a RoI head"""
+ return hasattr(self, 'query_head') and self.query_head is not None
+
+ @property
+ def with_roi_head(self):
+ """bool: whether the detector has a RoI head"""
+ return hasattr(self, 'roi_head') and self.roi_head is not None and len(
+ self.roi_head) > 0
+
+ @property
+ def with_shared_head(self):
+ """bool: whether the detector has a shared head in the RoI Head"""
+ return hasattr(self, 'roi_head') and self.roi_head[0].with_shared_head
+
+ @property
+ def with_bbox(self):
+ """bool: whether the detector has a bbox head"""
+ return ((hasattr(self, 'roi_head') and self.roi_head is not None
+ and len(self.roi_head) > 0)
+ or (hasattr(self, 'bbox_head') and self.bbox_head is not None
+ and len(self.bbox_head) > 0))
+
+ def extract_feat(self, batch_inputs: Tensor) -> Tuple[Tensor]:
+ """Extract features.
+
+ Args:
+ batch_inputs (Tensor): Image tensor, has shape (bs, dim, H, W).
+
+ Returns:
+ tuple[Tensor]: Tuple of feature maps from neck. Each feature map
+ has shape (bs, dim, H, W).
+ """
+ x = self.backbone(batch_inputs)
+ if self.with_neck:
+ x = self.neck(x)
+ return x
+
+ def _forward(self,
+ batch_inputs: Tensor,
+ batch_data_samples: OptSampleList = None):
+ pass
+
+ def loss(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> Union[dict, list]:
+ batch_input_shape = batch_data_samples[0].batch_input_shape
+ if self.use_lsj:
+ for data_samples in batch_data_samples:
+ img_metas = data_samples.metainfo
+ input_img_h, input_img_w = batch_input_shape
+ img_metas['img_shape'] = [input_img_h, input_img_w]
+
+ x = self.extract_feat(batch_inputs)
+
+ losses = dict()
+
+ def upd_loss(losses, idx, weight=1):
+ new_losses = dict()
+ for k, v in losses.items():
+ new_k = '{}{}'.format(k, idx)
+ if isinstance(v, list) or isinstance(v, tuple):
+ new_losses[new_k] = [i * weight for i in v]
+ else:
+ new_losses[new_k] = v * weight
+ return new_losses
+
+ # DETR encoder and decoder forward
+ if self.with_query_head:
+ bbox_losses, x = self.query_head.loss(x, batch_data_samples)
+ losses.update(bbox_losses)
+
+ # RPN forward and loss
+ if self.with_rpn:
+ proposal_cfg = self.train_cfg[self.head_idx].get(
+ 'rpn_proposal', self.test_cfg[self.head_idx].rpn)
+
+ rpn_data_samples = copy.deepcopy(batch_data_samples)
+ # set cat_id of gt_labels to 0 in RPN
+ for data_sample in rpn_data_samples:
+ data_sample.gt_instances.labels = \
+ torch.zeros_like(data_sample.gt_instances.labels)
+
+ rpn_losses, proposal_list = self.rpn_head.loss_and_predict(
+ x, rpn_data_samples, proposal_cfg=proposal_cfg)
+
+ # avoid get same name with roi_head loss
+ keys = rpn_losses.keys()
+ for key in list(keys):
+ if 'loss' in key and 'rpn' not in key:
+ rpn_losses[f'rpn_{key}'] = rpn_losses.pop(key)
+
+ losses.update(rpn_losses)
+ else:
+ assert batch_data_samples[0].get('proposals', None) is not None
+ # use pre-defined proposals in InstanceData for the second stage
+ # to extract ROI features.
+ proposal_list = [
+ data_sample.proposals for data_sample in batch_data_samples
+ ]
+
+ positive_coords = []
+ for i in range(len(self.roi_head)):
+ roi_losses = self.roi_head[i].loss(x, proposal_list,
+ batch_data_samples)
+ if self.with_pos_coord:
+ positive_coords.append(roi_losses.pop('pos_coords'))
+ else:
+ if 'pos_coords' in roi_losses.keys():
+ roi_losses.pop('pos_coords')
+ roi_losses = upd_loss(roi_losses, idx=i)
+ losses.update(roi_losses)
+
+ for i in range(len(self.bbox_head)):
+ bbox_losses = self.bbox_head[i].loss(x, batch_data_samples)
+ if self.with_pos_coord:
+ pos_coords = bbox_losses.pop('pos_coords')
+ positive_coords.append(pos_coords)
+ else:
+ if 'pos_coords' in bbox_losses.keys():
+ bbox_losses.pop('pos_coords')
+ bbox_losses = upd_loss(bbox_losses, idx=i + len(self.roi_head))
+ losses.update(bbox_losses)
+
+ if self.with_pos_coord and len(positive_coords) > 0:
+ for i in range(len(positive_coords)):
+ bbox_losses = self.query_head.loss_aux(x, positive_coords[i],
+ i, batch_data_samples)
+ bbox_losses = upd_loss(bbox_losses, idx=i)
+ losses.update(bbox_losses)
+
+ return losses
+
+ def predict(self,
+ batch_inputs: Tensor,
+ batch_data_samples: SampleList,
+ rescale: bool = True) -> SampleList:
+ """Predict results from a batch of inputs and data samples with post-
+ processing.
+
+ Args:
+ batch_inputs (Tensor): Inputs, has shape (bs, dim, H, W).
+ batch_data_samples (List[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+ rescale (bool): Whether to rescale the results.
+ Defaults to True.
+
+ Returns:
+ list[:obj:`DetDataSample`]: Detection results of the input images.
+ Each DetDataSample usually contain 'pred_instances'. And the
+ `pred_instances` usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ assert self.eval_module in ['detr', 'one-stage', 'two-stage']
+
+ if self.use_lsj:
+ for data_samples in batch_data_samples:
+ img_metas = data_samples.metainfo
+ input_img_h, input_img_w = img_metas['batch_input_shape']
+ img_metas['img_shape'] = [input_img_h, input_img_w]
+
+ img_feats = self.extract_feat(batch_inputs)
+ if self.with_bbox and self.eval_module == 'one-stage':
+ results_list = self.predict_bbox_head(
+ img_feats, batch_data_samples, rescale=rescale)
+ elif self.with_roi_head and self.eval_module == 'two-stage':
+ results_list = self.predict_roi_head(
+ img_feats, batch_data_samples, rescale=rescale)
+ else:
+ results_list = self.predict_query_head(
+ img_feats, batch_data_samples, rescale=rescale)
+
+ batch_data_samples = self.add_pred_to_datasample(
+ batch_data_samples, results_list)
+ return batch_data_samples
+
+ def predict_query_head(self,
+ mlvl_feats: Tuple[Tensor],
+ batch_data_samples: SampleList,
+ rescale: bool = True) -> InstanceList:
+ return self.query_head.predict(
+ mlvl_feats, batch_data_samples=batch_data_samples, rescale=rescale)
+
+ def predict_roi_head(self,
+ mlvl_feats: Tuple[Tensor],
+ batch_data_samples: SampleList,
+ rescale: bool = True) -> InstanceList:
+ assert self.with_bbox, 'Bbox head must be implemented.'
+ if self.with_query_head:
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+ results = self.query_head.forward(mlvl_feats, batch_img_metas)
+ mlvl_feats = results[-1]
+ rpn_results_list = self.rpn_head.predict(
+ mlvl_feats, batch_data_samples, rescale=False)
+ return self.roi_head[self.eval_index].predict(
+ mlvl_feats, rpn_results_list, batch_data_samples, rescale=rescale)
+
+ def predict_bbox_head(self,
+ mlvl_feats: Tuple[Tensor],
+ batch_data_samples: SampleList,
+ rescale: bool = True) -> InstanceList:
+ assert self.with_bbox, 'Bbox head must be implemented.'
+ if self.with_query_head:
+ batch_img_metas = [
+ data_samples.metainfo for data_samples in batch_data_samples
+ ]
+ results = self.query_head.forward(mlvl_feats, batch_img_metas)
+ mlvl_feats = results[-1]
+ return self.bbox_head[self.eval_index].predict(
+ mlvl_feats, batch_data_samples, rescale=rescale)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/codetr/transformer.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/codetr/transformer.py
new file mode 100644
index 0000000000000000000000000000000000000000..009f94a8bcc88c584b336bab272a48b4960202de
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/codetr/transformer.py
@@ -0,0 +1,1376 @@
+import math
+import warnings
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import build_norm_layer
+from mmcv.cnn.bricks.transformer import (BaseTransformerLayer,
+ TransformerLayerSequence,
+ build_transformer_layer_sequence)
+from mmcv.ops import MultiScaleDeformableAttention
+from mmengine.model import BaseModule
+from mmengine.model.weight_init import xavier_init
+from torch.nn.init import normal_
+
+from mmdet.models.layers.transformer import inverse_sigmoid
+from mmdet.registry import MODELS
+
+try:
+ from fairscale.nn.checkpoint import checkpoint_wrapper
+except Exception:
+ checkpoint_wrapper = None
+
+# In order to save the cost and effort of reproduction,
+# I did not refactor it into the style of mmdet 3.x DETR.
+
+
+class Transformer(BaseModule):
+ """Implements the DETR transformer.
+
+ Following the official DETR implementation, this module copy-paste
+ from torch.nn.Transformer with modifications:
+
+ * positional encodings are passed in MultiheadAttention
+ * extra LN at the end of encoder is removed
+ * decoder returns a stack of activations from all decoding layers
+
+ See `paper: End-to-End Object Detection with Transformers
+ `_ for details.
+
+ Args:
+ encoder (`mmcv.ConfigDict` | Dict): Config of
+ TransformerEncoder. Defaults to None.
+ decoder ((`mmcv.ConfigDict` | Dict)): Config of
+ TransformerDecoder. Defaults to None
+ init_cfg (obj:`mmcv.ConfigDict`): The Config for initialization.
+ Defaults to None.
+ """
+
+ def __init__(self, encoder=None, decoder=None, init_cfg=None):
+ super(Transformer, self).__init__(init_cfg=init_cfg)
+ self.encoder = build_transformer_layer_sequence(encoder)
+ self.decoder = build_transformer_layer_sequence(decoder)
+ self.embed_dims = self.encoder.embed_dims
+
+ def init_weights(self):
+ # follow the official DETR to init parameters
+ for m in self.modules():
+ if hasattr(m, 'weight') and m.weight.dim() > 1:
+ xavier_init(m, distribution='uniform')
+ self._is_init = True
+
+ def forward(self, x, mask, query_embed, pos_embed):
+ """Forward function for `Transformer`.
+
+ Args:
+ x (Tensor): Input query with shape [bs, c, h, w] where
+ c = embed_dims.
+ mask (Tensor): The key_padding_mask used for encoder and decoder,
+ with shape [bs, h, w].
+ query_embed (Tensor): The query embedding for decoder, with shape
+ [num_query, c].
+ pos_embed (Tensor): The positional encoding for encoder and
+ decoder, with the same shape as `x`.
+
+ Returns:
+ tuple[Tensor]: results of decoder containing the following tensor.
+
+ - out_dec: Output from decoder. If return_intermediate_dec \
+ is True output has shape [num_dec_layers, bs,
+ num_query, embed_dims], else has shape [1, bs, \
+ num_query, embed_dims].
+ - memory: Output results from encoder, with shape \
+ [bs, embed_dims, h, w].
+ """
+ bs, c, h, w = x.shape
+ # use `view` instead of `flatten` for dynamically exporting to ONNX
+ x = x.view(bs, c, -1).permute(2, 0, 1) # [bs, c, h, w] -> [h*w, bs, c]
+ pos_embed = pos_embed.view(bs, c, -1).permute(2, 0, 1)
+ query_embed = query_embed.unsqueeze(1).repeat(
+ 1, bs, 1) # [num_query, dim] -> [num_query, bs, dim]
+ mask = mask.view(bs, -1) # [bs, h, w] -> [bs, h*w]
+ memory = self.encoder(
+ query=x,
+ key=None,
+ value=None,
+ query_pos=pos_embed,
+ query_key_padding_mask=mask)
+ target = torch.zeros_like(query_embed)
+ # out_dec: [num_layers, num_query, bs, dim]
+ out_dec = self.decoder(
+ query=target,
+ key=memory,
+ value=memory,
+ key_pos=pos_embed,
+ query_pos=query_embed,
+ key_padding_mask=mask)
+ out_dec = out_dec.transpose(1, 2)
+ memory = memory.permute(1, 2, 0).reshape(bs, c, h, w)
+ return out_dec, memory
+
+
+@MODELS.register_module(force=True)
+class DeformableDetrTransformerDecoder(TransformerLayerSequence):
+ """Implements the decoder in DETR transformer.
+
+ Args:
+ return_intermediate (bool): Whether to return intermediate outputs.
+ coder_norm_cfg (dict): Config of last normalization layer. Default:
+ `LN`.
+ """
+
+ def __init__(self, *args, return_intermediate=False, **kwargs):
+
+ super(DeformableDetrTransformerDecoder, self).__init__(*args, **kwargs)
+ self.return_intermediate = return_intermediate
+
+ def forward(self,
+ query,
+ *args,
+ reference_points=None,
+ valid_ratios=None,
+ reg_branches=None,
+ **kwargs):
+ """Forward function for `TransformerDecoder`.
+
+ Args:
+ query (Tensor): Input query with shape
+ `(num_query, bs, embed_dims)`.
+ reference_points (Tensor): The reference
+ points of offset. has shape
+ (bs, num_query, 4) when as_two_stage,
+ otherwise has shape ((bs, num_query, 2).
+ valid_ratios (Tensor): The radios of valid
+ points on the feature map, has shape
+ (bs, num_levels, 2)
+ reg_branch: (obj:`nn.ModuleList`): Used for
+ refining the regression results. Only would
+ be passed when with_box_refine is True,
+ otherwise would be passed a `None`.
+
+ Returns:
+ Tensor: Results with shape [1, num_query, bs, embed_dims] when
+ return_intermediate is `False`, otherwise it has shape
+ [num_layers, num_query, bs, embed_dims].
+ """
+ output = query
+ intermediate = []
+ intermediate_reference_points = []
+ for lid, layer in enumerate(self.layers):
+ if reference_points.shape[-1] == 4:
+ reference_points_input = reference_points[:, :, None] * \
+ torch.cat([valid_ratios, valid_ratios], -1)[:, None]
+ else:
+ assert reference_points.shape[-1] == 2
+ reference_points_input = reference_points[:, :, None] * \
+ valid_ratios[:, None]
+ output = layer(
+ output,
+ *args,
+ reference_points=reference_points_input,
+ **kwargs)
+ output = output.permute(1, 0, 2)
+
+ if reg_branches is not None:
+ tmp = reg_branches[lid](output)
+ if reference_points.shape[-1] == 4:
+ new_reference_points = tmp + inverse_sigmoid(
+ reference_points)
+ new_reference_points = new_reference_points.sigmoid()
+ else:
+ assert reference_points.shape[-1] == 2
+ new_reference_points = tmp
+ new_reference_points[..., :2] = tmp[
+ ..., :2] + inverse_sigmoid(reference_points)
+ new_reference_points = new_reference_points.sigmoid()
+ reference_points = new_reference_points.detach()
+
+ output = output.permute(1, 0, 2)
+ if self.return_intermediate:
+ intermediate.append(output)
+ intermediate_reference_points.append(reference_points)
+
+ if self.return_intermediate:
+ return torch.stack(intermediate), torch.stack(
+ intermediate_reference_points)
+
+ return output, reference_points
+
+
+@MODELS.register_module(force=True)
+class DeformableDetrTransformer(Transformer):
+ """Implements the DeformableDETR transformer.
+
+ Args:
+ as_two_stage (bool): Generate query from encoder features.
+ Default: False.
+ num_feature_levels (int): Number of feature maps from FPN:
+ Default: 4.
+ two_stage_num_proposals (int): Number of proposals when set
+ `as_two_stage` as True. Default: 300.
+ """
+
+ def __init__(self,
+ as_two_stage=False,
+ num_feature_levels=4,
+ two_stage_num_proposals=300,
+ **kwargs):
+ super(DeformableDetrTransformer, self).__init__(**kwargs)
+ self.as_two_stage = as_two_stage
+ self.num_feature_levels = num_feature_levels
+ self.two_stage_num_proposals = two_stage_num_proposals
+ self.embed_dims = self.encoder.embed_dims
+ self.init_layers()
+
+ def init_layers(self):
+ """Initialize layers of the DeformableDetrTransformer."""
+ self.level_embeds = nn.Parameter(
+ torch.Tensor(self.num_feature_levels, self.embed_dims))
+
+ if self.as_two_stage:
+ self.enc_output = nn.Linear(self.embed_dims, self.embed_dims)
+ self.enc_output_norm = nn.LayerNorm(self.embed_dims)
+ self.pos_trans = nn.Linear(self.embed_dims * 2,
+ self.embed_dims * 2)
+ self.pos_trans_norm = nn.LayerNorm(self.embed_dims * 2)
+ else:
+ self.reference_points = nn.Linear(self.embed_dims, 2)
+
+ def init_weights(self):
+ """Initialize the transformer weights."""
+ for p in self.parameters():
+ if p.dim() > 1:
+ nn.init.xavier_uniform_(p)
+ for m in self.modules():
+ if isinstance(m, MultiScaleDeformableAttention):
+ m.init_weights()
+ if not self.as_two_stage:
+ xavier_init(self.reference_points, distribution='uniform', bias=0.)
+ normal_(self.level_embeds)
+
+ def gen_encoder_output_proposals(self, memory, memory_padding_mask,
+ spatial_shapes):
+ """Generate proposals from encoded memory.
+
+ Args:
+ memory (Tensor) : The output of encoder,
+ has shape (bs, num_key, embed_dim). num_key is
+ equal the number of points on feature map from
+ all level.
+ memory_padding_mask (Tensor): Padding mask for memory.
+ has shape (bs, num_key).
+ spatial_shapes (Tensor): The shape of all feature maps.
+ has shape (num_level, 2).
+
+ Returns:
+ tuple: A tuple of feature map and bbox prediction.
+
+ - output_memory (Tensor): The input of decoder, \
+ has shape (bs, num_key, embed_dim). num_key is \
+ equal the number of points on feature map from \
+ all levels.
+ - output_proposals (Tensor): The normalized proposal \
+ after a inverse sigmoid, has shape \
+ (bs, num_keys, 4).
+ """
+
+ N, S, C = memory.shape
+ proposals = []
+ _cur = 0
+ for lvl, (H, W) in enumerate(spatial_shapes):
+ mask_flatten_ = memory_padding_mask[:, _cur:(_cur + H * W)].view(
+ N, H, W, 1)
+ valid_H = torch.sum(~mask_flatten_[:, :, 0, 0], 1)
+ valid_W = torch.sum(~mask_flatten_[:, 0, :, 0], 1)
+
+ grid_y, grid_x = torch.meshgrid(
+ torch.linspace(
+ 0, H - 1, H, dtype=torch.float32, device=memory.device),
+ torch.linspace(
+ 0, W - 1, W, dtype=torch.float32, device=memory.device))
+ grid = torch.cat([grid_x.unsqueeze(-1), grid_y.unsqueeze(-1)], -1)
+
+ scale = torch.cat([valid_W.unsqueeze(-1),
+ valid_H.unsqueeze(-1)], 1).view(N, 1, 1, 2)
+ grid = (grid.unsqueeze(0).expand(N, -1, -1, -1) + 0.5) / scale
+ wh = torch.ones_like(grid) * 0.05 * (2.0**lvl)
+ proposal = torch.cat((grid, wh), -1).view(N, -1, 4)
+ proposals.append(proposal)
+ _cur += (H * W)
+ output_proposals = torch.cat(proposals, 1)
+ output_proposals_valid = ((output_proposals > 0.01) &
+ (output_proposals < 0.99)).all(
+ -1, keepdim=True)
+ output_proposals = torch.log(output_proposals / (1 - output_proposals))
+ output_proposals = output_proposals.masked_fill(
+ memory_padding_mask.unsqueeze(-1), float('inf'))
+ output_proposals = output_proposals.masked_fill(
+ ~output_proposals_valid, float('inf'))
+
+ output_memory = memory
+ output_memory = output_memory.masked_fill(
+ memory_padding_mask.unsqueeze(-1), float(0))
+ output_memory = output_memory.masked_fill(~output_proposals_valid,
+ float(0))
+ output_memory = self.enc_output_norm(self.enc_output(output_memory))
+ return output_memory, output_proposals
+
+ @staticmethod
+ def get_reference_points(spatial_shapes, valid_ratios, device):
+ """Get the reference points used in decoder.
+
+ Args:
+ spatial_shapes (Tensor): The shape of all
+ feature maps, has shape (num_level, 2).
+ valid_ratios (Tensor): The radios of valid
+ points on the feature map, has shape
+ (bs, num_levels, 2)
+ device (obj:`device`): The device where
+ reference_points should be.
+
+ Returns:
+ Tensor: reference points used in decoder, has \
+ shape (bs, num_keys, num_levels, 2).
+ """
+ reference_points_list = []
+ for lvl, (H, W) in enumerate(spatial_shapes):
+ ref_y, ref_x = torch.meshgrid(
+ torch.linspace(
+ 0.5, H - 0.5, H, dtype=torch.float32, device=device),
+ torch.linspace(
+ 0.5, W - 0.5, W, dtype=torch.float32, device=device))
+ ref_y = ref_y.reshape(-1)[None] / (
+ valid_ratios[:, None, lvl, 1] * H)
+ ref_x = ref_x.reshape(-1)[None] / (
+ valid_ratios[:, None, lvl, 0] * W)
+ ref = torch.stack((ref_x, ref_y), -1)
+ reference_points_list.append(ref)
+ reference_points = torch.cat(reference_points_list, 1)
+ reference_points = reference_points[:, :, None] * valid_ratios[:, None]
+ return reference_points
+
+ def get_valid_ratio(self, mask):
+ """Get the valid radios of feature maps of all level."""
+ _, H, W = mask.shape
+ valid_H = torch.sum(~mask[:, :, 0], 1)
+ valid_W = torch.sum(~mask[:, 0, :], 1)
+ valid_ratio_h = valid_H.float() / H
+ valid_ratio_w = valid_W.float() / W
+ valid_ratio = torch.stack([valid_ratio_w, valid_ratio_h], -1)
+ return valid_ratio
+
+ def get_proposal_pos_embed(self,
+ proposals,
+ num_pos_feats=128,
+ temperature=10000):
+ """Get the position embedding of proposal."""
+ scale = 2 * math.pi
+ dim_t = torch.arange(
+ num_pos_feats, dtype=torch.float32, device=proposals.device)
+ dim_t = temperature**(2 * (dim_t // 2) / num_pos_feats)
+ # N, L, 4
+ proposals = proposals.sigmoid() * scale
+ # N, L, 4, 128
+ pos = proposals[:, :, :, None] / dim_t
+ # N, L, 4, 64, 2
+ pos = torch.stack((pos[:, :, :, 0::2].sin(), pos[:, :, :, 1::2].cos()),
+ dim=4).flatten(2)
+ return pos
+
+ def forward(self,
+ mlvl_feats,
+ mlvl_masks,
+ query_embed,
+ mlvl_pos_embeds,
+ reg_branches=None,
+ cls_branches=None,
+ **kwargs):
+ """Forward function for `Transformer`.
+
+ Args:
+ mlvl_feats (list(Tensor)): Input queries from
+ different level. Each element has shape
+ [bs, embed_dims, h, w].
+ mlvl_masks (list(Tensor)): The key_padding_mask from
+ different level used for encoder and decoder,
+ each element has shape [bs, h, w].
+ query_embed (Tensor): The query embedding for decoder,
+ with shape [num_query, c].
+ mlvl_pos_embeds (list(Tensor)): The positional encoding
+ of feats from different level, has the shape
+ [bs, embed_dims, h, w].
+ reg_branches (obj:`nn.ModuleList`): Regression heads for
+ feature maps from each decoder layer. Only would
+ be passed when
+ `with_box_refine` is True. Default to None.
+ cls_branches (obj:`nn.ModuleList`): Classification heads
+ for feature maps from each decoder layer. Only would
+ be passed when `as_two_stage`
+ is True. Default to None.
+
+
+ Returns:
+ tuple[Tensor]: results of decoder containing the following tensor.
+
+ - inter_states: Outputs from decoder. If
+ return_intermediate_dec is True output has shape \
+ (num_dec_layers, bs, num_query, embed_dims), else has \
+ shape (1, bs, num_query, embed_dims).
+ - init_reference_out: The initial value of reference \
+ points, has shape (bs, num_queries, 4).
+ - inter_references_out: The internal value of reference \
+ points in decoder, has shape \
+ (num_dec_layers, bs,num_query, embed_dims)
+ - enc_outputs_class: The classification score of \
+ proposals generated from \
+ encoder's feature maps, has shape \
+ (batch, h*w, num_classes). \
+ Only would be returned when `as_two_stage` is True, \
+ otherwise None.
+ - enc_outputs_coord_unact: The regression results \
+ generated from encoder's feature maps., has shape \
+ (batch, h*w, 4). Only would \
+ be returned when `as_two_stage` is True, \
+ otherwise None.
+ """
+ assert self.as_two_stage or query_embed is not None
+
+ feat_flatten = []
+ mask_flatten = []
+ lvl_pos_embed_flatten = []
+ spatial_shapes = []
+ for lvl, (feat, mask, pos_embed) in enumerate(
+ zip(mlvl_feats, mlvl_masks, mlvl_pos_embeds)):
+ bs, c, h, w = feat.shape
+ spatial_shape = (h, w)
+ spatial_shapes.append(spatial_shape)
+ feat = feat.flatten(2).transpose(1, 2)
+ mask = mask.flatten(1)
+ pos_embed = pos_embed.flatten(2).transpose(1, 2)
+ lvl_pos_embed = pos_embed + self.level_embeds[lvl].view(1, 1, -1)
+ lvl_pos_embed_flatten.append(lvl_pos_embed)
+ feat_flatten.append(feat)
+ mask_flatten.append(mask)
+ feat_flatten = torch.cat(feat_flatten, 1)
+ mask_flatten = torch.cat(mask_flatten, 1)
+ lvl_pos_embed_flatten = torch.cat(lvl_pos_embed_flatten, 1)
+ spatial_shapes = torch.as_tensor(
+ spatial_shapes, dtype=torch.long, device=feat_flatten.device)
+ level_start_index = torch.cat((spatial_shapes.new_zeros(
+ (1, )), spatial_shapes.prod(1).cumsum(0)[:-1]))
+ valid_ratios = torch.stack(
+ [self.get_valid_ratio(m) for m in mlvl_masks], 1)
+
+ reference_points = \
+ self.get_reference_points(spatial_shapes,
+ valid_ratios,
+ device=feat.device)
+
+ feat_flatten = feat_flatten.permute(1, 0, 2) # (H*W, bs, embed_dims)
+ lvl_pos_embed_flatten = lvl_pos_embed_flatten.permute(
+ 1, 0, 2) # (H*W, bs, embed_dims)
+ memory = self.encoder(
+ query=feat_flatten,
+ key=None,
+ value=None,
+ query_pos=lvl_pos_embed_flatten,
+ query_key_padding_mask=mask_flatten,
+ spatial_shapes=spatial_shapes,
+ reference_points=reference_points,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios,
+ **kwargs)
+
+ memory = memory.permute(1, 0, 2)
+ bs, _, c = memory.shape
+ if self.as_two_stage:
+ output_memory, output_proposals = \
+ self.gen_encoder_output_proposals(
+ memory, mask_flatten, spatial_shapes)
+ enc_outputs_class = cls_branches[self.decoder.num_layers](
+ output_memory)
+ enc_outputs_coord_unact = \
+ reg_branches[
+ self.decoder.num_layers](output_memory) + output_proposals
+
+ topk = self.two_stage_num_proposals
+ # We only use the first channel in enc_outputs_class as foreground,
+ # the other (num_classes - 1) channels are actually not used.
+ # Its targets are set to be 0s, which indicates the first
+ # class (foreground) because we use [0, num_classes - 1] to
+ # indicate class labels, background class is indicated by
+ # num_classes (similar convention in RPN).
+ # See https://github.com/open-mmlab/mmdetection/blob/master/mmdet/models/dense_heads/deformable_detr_head.py#L241 # noqa
+ # This follows the official implementation of Deformable DETR.
+ topk_proposals = torch.topk(
+ enc_outputs_class[..., 0], topk, dim=1)[1]
+ topk_coords_unact = torch.gather(
+ enc_outputs_coord_unact, 1,
+ topk_proposals.unsqueeze(-1).repeat(1, 1, 4))
+ topk_coords_unact = topk_coords_unact.detach()
+ reference_points = topk_coords_unact.sigmoid()
+ init_reference_out = reference_points
+ pos_trans_out = self.pos_trans_norm(
+ self.pos_trans(self.get_proposal_pos_embed(topk_coords_unact)))
+ query_pos, query = torch.split(pos_trans_out, c, dim=2)
+ else:
+ query_pos, query = torch.split(query_embed, c, dim=1)
+ query_pos = query_pos.unsqueeze(0).expand(bs, -1, -1)
+ query = query.unsqueeze(0).expand(bs, -1, -1)
+ reference_points = self.reference_points(query_pos).sigmoid()
+ init_reference_out = reference_points
+
+ # decoder
+ query = query.permute(1, 0, 2)
+ memory = memory.permute(1, 0, 2)
+ query_pos = query_pos.permute(1, 0, 2)
+ inter_states, inter_references = self.decoder(
+ query=query,
+ key=None,
+ value=memory,
+ query_pos=query_pos,
+ key_padding_mask=mask_flatten,
+ reference_points=reference_points,
+ spatial_shapes=spatial_shapes,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios,
+ reg_branches=reg_branches,
+ **kwargs)
+
+ inter_references_out = inter_references
+ if self.as_two_stage:
+ return inter_states, init_reference_out,\
+ inter_references_out, enc_outputs_class,\
+ enc_outputs_coord_unact
+ return inter_states, init_reference_out, \
+ inter_references_out, None, None
+
+
+@MODELS.register_module()
+class CoDeformableDetrTransformerDecoder(TransformerLayerSequence):
+ """Implements the decoder in DETR transformer.
+
+ Args:
+ return_intermediate (bool): Whether to return intermediate outputs.
+ coder_norm_cfg (dict): Config of last normalization layer. Default:
+ `LN`.
+ """
+
+ def __init__(self,
+ *args,
+ return_intermediate=False,
+ look_forward_twice=False,
+ **kwargs):
+
+ super(CoDeformableDetrTransformerDecoder,
+ self).__init__(*args, **kwargs)
+ self.return_intermediate = return_intermediate
+ self.look_forward_twice = look_forward_twice
+
+ def forward(self,
+ query,
+ *args,
+ reference_points=None,
+ valid_ratios=None,
+ reg_branches=None,
+ **kwargs):
+ """Forward function for `TransformerDecoder`.
+
+ Args:
+ query (Tensor): Input query with shape
+ `(num_query, bs, embed_dims)`.
+ reference_points (Tensor): The reference
+ points of offset. has shape
+ (bs, num_query, 4) when as_two_stage,
+ otherwise has shape ((bs, num_query, 2).
+ valid_ratios (Tensor): The radios of valid
+ points on the feature map, has shape
+ (bs, num_levels, 2)
+ reg_branch: (obj:`nn.ModuleList`): Used for
+ refining the regression results. Only would
+ be passed when with_box_refine is True,
+ otherwise would be passed a `None`.
+
+ Returns:
+ Tensor: Results with shape [1, num_query, bs, embed_dims] when
+ return_intermediate is `False`, otherwise it has shape
+ [num_layers, num_query, bs, embed_dims].
+ """
+ output = query
+ intermediate = []
+ intermediate_reference_points = []
+ for lid, layer in enumerate(self.layers):
+ if reference_points.shape[-1] == 4:
+ reference_points_input = reference_points[:, :, None] * \
+ torch.cat([valid_ratios, valid_ratios], -1)[:, None]
+ else:
+ assert reference_points.shape[-1] == 2
+ reference_points_input = reference_points[:, :, None] * \
+ valid_ratios[:, None]
+ output = layer(
+ output,
+ *args,
+ reference_points=reference_points_input,
+ **kwargs)
+ output = output.permute(1, 0, 2)
+
+ if reg_branches is not None:
+ tmp = reg_branches[lid](output)
+ if reference_points.shape[-1] == 4:
+ new_reference_points = tmp + inverse_sigmoid(
+ reference_points)
+ new_reference_points = new_reference_points.sigmoid()
+ else:
+ assert reference_points.shape[-1] == 2
+ new_reference_points = tmp
+ new_reference_points[..., :2] = tmp[
+ ..., :2] + inverse_sigmoid(reference_points)
+ new_reference_points = new_reference_points.sigmoid()
+ reference_points = new_reference_points.detach()
+
+ output = output.permute(1, 0, 2)
+ if self.return_intermediate:
+ intermediate.append(output)
+ intermediate_reference_points.append(
+ new_reference_points if self.
+ look_forward_twice else reference_points)
+ if self.return_intermediate:
+ return torch.stack(intermediate), torch.stack(
+ intermediate_reference_points)
+
+ return output, reference_points
+
+
+@MODELS.register_module()
+class CoDeformableDetrTransformer(DeformableDetrTransformer):
+
+ def __init__(self,
+ mixed_selection=True,
+ with_pos_coord=True,
+ with_coord_feat=True,
+ num_co_heads=1,
+ **kwargs):
+ self.mixed_selection = mixed_selection
+ self.with_pos_coord = with_pos_coord
+ self.with_coord_feat = with_coord_feat
+ self.num_co_heads = num_co_heads
+ super(CoDeformableDetrTransformer, self).__init__(**kwargs)
+ self._init_layers()
+
+ def _init_layers(self):
+ """Initialize layers of the CoDeformableDetrTransformer."""
+ if self.with_pos_coord:
+ if self.num_co_heads > 0:
+ # bug: this code should be 'self.head_pos_embed =
+ # nn.Embedding(self.num_co_heads, self.embed_dims)',
+ # we keep this bug for reproducing our results with ResNet-50.
+ # You can fix this bug when reproducing results with
+ # swin transformer.
+ self.head_pos_embed = nn.Embedding(self.num_co_heads, 1, 1,
+ self.embed_dims)
+ self.aux_pos_trans = nn.ModuleList()
+ self.aux_pos_trans_norm = nn.ModuleList()
+ self.pos_feats_trans = nn.ModuleList()
+ self.pos_feats_norm = nn.ModuleList()
+ for i in range(self.num_co_heads):
+ self.aux_pos_trans.append(
+ nn.Linear(self.embed_dims * 2, self.embed_dims * 2))
+ self.aux_pos_trans_norm.append(
+ nn.LayerNorm(self.embed_dims * 2))
+ if self.with_coord_feat:
+ self.pos_feats_trans.append(
+ nn.Linear(self.embed_dims, self.embed_dims))
+ self.pos_feats_norm.append(
+ nn.LayerNorm(self.embed_dims))
+
+ def get_proposal_pos_embed(self,
+ proposals,
+ num_pos_feats=128,
+ temperature=10000):
+ """Get the position embedding of proposal."""
+ num_pos_feats = self.embed_dims // 2
+ scale = 2 * math.pi
+ dim_t = torch.arange(
+ num_pos_feats, dtype=torch.float32, device=proposals.device)
+ dim_t = temperature**(2 * (dim_t // 2) / num_pos_feats)
+ # N, L, 4
+ proposals = proposals.sigmoid() * scale
+ # N, L, 4, 128
+ pos = proposals[:, :, :, None] / dim_t
+ # N, L, 4, 64, 2
+ pos = torch.stack((pos[:, :, :, 0::2].sin(), pos[:, :, :, 1::2].cos()),
+ dim=4).flatten(2)
+ return pos
+
+ def forward(self,
+ mlvl_feats,
+ mlvl_masks,
+ query_embed,
+ mlvl_pos_embeds,
+ reg_branches=None,
+ cls_branches=None,
+ return_encoder_output=False,
+ attn_masks=None,
+ **kwargs):
+ """Forward function for `Transformer`.
+
+ Args:
+ mlvl_feats (list(Tensor)): Input queries from
+ different level. Each element has shape
+ [bs, embed_dims, h, w].
+ mlvl_masks (list(Tensor)): The key_padding_mask from
+ different level used for encoder and decoder,
+ each element has shape [bs, h, w].
+ query_embed (Tensor): The query embedding for decoder,
+ with shape [num_query, c].
+ mlvl_pos_embeds (list(Tensor)): The positional encoding
+ of feats from different level, has the shape
+ [bs, embed_dims, h, w].
+ reg_branches (obj:`nn.ModuleList`): Regression heads for
+ feature maps from each decoder layer. Only would
+ be passed when
+ `with_box_refine` is True. Default to None.
+ cls_branches (obj:`nn.ModuleList`): Classification heads
+ for feature maps from each decoder layer. Only would
+ be passed when `as_two_stage`
+ is True. Default to None.
+
+
+ Returns:
+ tuple[Tensor]: results of decoder containing the following tensor.
+
+ - inter_states: Outputs from decoder. If
+ return_intermediate_dec is True output has shape \
+ (num_dec_layers, bs, num_query, embed_dims), else has \
+ shape (1, bs, num_query, embed_dims).
+ - init_reference_out: The initial value of reference \
+ points, has shape (bs, num_queries, 4).
+ - inter_references_out: The internal value of reference \
+ points in decoder, has shape \
+ (num_dec_layers, bs,num_query, embed_dims)
+ - enc_outputs_class: The classification score of \
+ proposals generated from \
+ encoder's feature maps, has shape \
+ (batch, h*w, num_classes). \
+ Only would be returned when `as_two_stage` is True, \
+ otherwise None.
+ - enc_outputs_coord_unact: The regression results \
+ generated from encoder's feature maps., has shape \
+ (batch, h*w, 4). Only would \
+ be returned when `as_two_stage` is True, \
+ otherwise None.
+ """
+ assert self.as_two_stage or query_embed is not None
+
+ feat_flatten = []
+ mask_flatten = []
+ lvl_pos_embed_flatten = []
+ spatial_shapes = []
+ for lvl, (feat, mask, pos_embed) in enumerate(
+ zip(mlvl_feats, mlvl_masks, mlvl_pos_embeds)):
+ bs, c, h, w = feat.shape
+ spatial_shape = (h, w)
+ spatial_shapes.append(spatial_shape)
+ feat = feat.flatten(2).transpose(1, 2)
+ mask = mask.flatten(1)
+ pos_embed = pos_embed.flatten(2).transpose(1, 2)
+ lvl_pos_embed = pos_embed + self.level_embeds[lvl].view(1, 1, -1)
+ lvl_pos_embed_flatten.append(lvl_pos_embed)
+ feat_flatten.append(feat)
+ mask_flatten.append(mask)
+ feat_flatten = torch.cat(feat_flatten, 1)
+ mask_flatten = torch.cat(mask_flatten, 1)
+ lvl_pos_embed_flatten = torch.cat(lvl_pos_embed_flatten, 1)
+ spatial_shapes = torch.as_tensor(
+ spatial_shapes, dtype=torch.long, device=feat_flatten.device)
+ level_start_index = torch.cat((spatial_shapes.new_zeros(
+ (1, )), spatial_shapes.prod(1).cumsum(0)[:-1]))
+ valid_ratios = torch.stack(
+ [self.get_valid_ratio(m) for m in mlvl_masks], 1)
+
+ reference_points = \
+ self.get_reference_points(spatial_shapes,
+ valid_ratios,
+ device=feat.device)
+
+ feat_flatten = feat_flatten.permute(1, 0, 2) # (H*W, bs, embed_dims)
+ lvl_pos_embed_flatten = lvl_pos_embed_flatten.permute(
+ 1, 0, 2) # (H*W, bs, embed_dims)
+ memory = self.encoder(
+ query=feat_flatten,
+ key=None,
+ value=None,
+ query_pos=lvl_pos_embed_flatten,
+ query_key_padding_mask=mask_flatten,
+ spatial_shapes=spatial_shapes,
+ reference_points=reference_points,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios,
+ **kwargs)
+
+ memory = memory.permute(1, 0, 2)
+ bs, _, c = memory.shape
+ if self.as_two_stage:
+ output_memory, output_proposals = \
+ self.gen_encoder_output_proposals(
+ memory, mask_flatten, spatial_shapes)
+ enc_outputs_class = cls_branches[self.decoder.num_layers](
+ output_memory)
+ enc_outputs_coord_unact = \
+ reg_branches[
+ self.decoder.num_layers](output_memory) + output_proposals
+
+ topk = self.two_stage_num_proposals
+ topk = query_embed.shape[0]
+ topk_proposals = torch.topk(
+ enc_outputs_class[..., 0], topk, dim=1)[1]
+ topk_coords_unact = torch.gather(
+ enc_outputs_coord_unact, 1,
+ topk_proposals.unsqueeze(-1).repeat(1, 1, 4))
+ topk_coords_unact = topk_coords_unact.detach()
+ reference_points = topk_coords_unact.sigmoid()
+ init_reference_out = reference_points
+ pos_trans_out = self.pos_trans_norm(
+ self.pos_trans(self.get_proposal_pos_embed(topk_coords_unact)))
+
+ if not self.mixed_selection:
+ query_pos, query = torch.split(pos_trans_out, c, dim=2)
+ else:
+ # query_embed here is the content embed for deformable DETR
+ query = query_embed.unsqueeze(0).expand(bs, -1, -1)
+ query_pos, _ = torch.split(pos_trans_out, c, dim=2)
+ else:
+ query_pos, query = torch.split(query_embed, c, dim=1)
+ query_pos = query_pos.unsqueeze(0).expand(bs, -1, -1)
+ query = query.unsqueeze(0).expand(bs, -1, -1)
+ reference_points = self.reference_points(query_pos).sigmoid()
+ init_reference_out = reference_points
+
+ # decoder
+ query = query.permute(1, 0, 2)
+ memory = memory.permute(1, 0, 2)
+ query_pos = query_pos.permute(1, 0, 2)
+ inter_states, inter_references = self.decoder(
+ query=query,
+ key=None,
+ value=memory,
+ query_pos=query_pos,
+ key_padding_mask=mask_flatten,
+ reference_points=reference_points,
+ spatial_shapes=spatial_shapes,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios,
+ reg_branches=reg_branches,
+ attn_masks=attn_masks,
+ **kwargs)
+
+ inter_references_out = inter_references
+ if self.as_two_stage:
+ if return_encoder_output:
+ return inter_states, init_reference_out,\
+ inter_references_out, enc_outputs_class,\
+ enc_outputs_coord_unact, memory
+ return inter_states, init_reference_out,\
+ inter_references_out, enc_outputs_class,\
+ enc_outputs_coord_unact
+ if return_encoder_output:
+ return inter_states, init_reference_out, \
+ inter_references_out, None, None, memory
+ return inter_states, init_reference_out, \
+ inter_references_out, None, None
+
+ def forward_aux(self,
+ mlvl_feats,
+ mlvl_masks,
+ query_embed,
+ mlvl_pos_embeds,
+ pos_anchors,
+ pos_feats=None,
+ reg_branches=None,
+ cls_branches=None,
+ return_encoder_output=False,
+ attn_masks=None,
+ head_idx=0,
+ **kwargs):
+ feat_flatten = []
+ mask_flatten = []
+ spatial_shapes = []
+ for lvl, (feat, mask, pos_embed) in enumerate(
+ zip(mlvl_feats, mlvl_masks, mlvl_pos_embeds)):
+ bs, c, h, w = feat.shape
+ spatial_shape = (h, w)
+ spatial_shapes.append(spatial_shape)
+ feat = feat.flatten(2).transpose(1, 2)
+ mask = mask.flatten(1)
+ feat_flatten.append(feat)
+ mask_flatten.append(mask)
+ feat_flatten = torch.cat(feat_flatten, 1)
+ mask_flatten = torch.cat(mask_flatten, 1)
+ spatial_shapes = torch.as_tensor(
+ spatial_shapes, dtype=torch.long, device=feat_flatten.device)
+ level_start_index = torch.cat((spatial_shapes.new_zeros(
+ (1, )), spatial_shapes.prod(1).cumsum(0)[:-1]))
+ valid_ratios = torch.stack(
+ [self.get_valid_ratio(m) for m in mlvl_masks], 1)
+
+ feat_flatten = feat_flatten.permute(1, 0, 2) # (H*W, bs, embed_dims)
+
+ memory = feat_flatten
+ memory = memory.permute(1, 0, 2)
+ bs, _, c = memory.shape
+
+ topk_coords_unact = inverse_sigmoid(pos_anchors)
+ reference_points = pos_anchors
+ init_reference_out = reference_points
+ if self.num_co_heads > 0:
+ pos_trans_out = self.aux_pos_trans_norm[head_idx](
+ self.aux_pos_trans[head_idx](
+ self.get_proposal_pos_embed(topk_coords_unact)))
+ query_pos, query = torch.split(pos_trans_out, c, dim=2)
+ if self.with_coord_feat:
+ query = query + self.pos_feats_norm[head_idx](
+ self.pos_feats_trans[head_idx](pos_feats))
+ query_pos = query_pos + self.head_pos_embed.weight[head_idx]
+
+ # decoder
+ query = query.permute(1, 0, 2)
+ memory = memory.permute(1, 0, 2)
+ query_pos = query_pos.permute(1, 0, 2)
+ inter_states, inter_references = self.decoder(
+ query=query,
+ key=None,
+ value=memory,
+ query_pos=query_pos,
+ key_padding_mask=mask_flatten,
+ reference_points=reference_points,
+ spatial_shapes=spatial_shapes,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios,
+ reg_branches=reg_branches,
+ attn_masks=attn_masks,
+ **kwargs)
+
+ inter_references_out = inter_references
+ return inter_states, init_reference_out, \
+ inter_references_out
+
+
+def build_MLP(input_dim, hidden_dim, output_dim, num_layers):
+ assert num_layers > 1, \
+ f'num_layers should be greater than 1 but got {num_layers}'
+ h = [hidden_dim] * (num_layers - 1)
+ layers = list()
+ for n, k in zip([input_dim] + h[:-1], h):
+ layers.extend((nn.Linear(n, k), nn.ReLU()))
+ # Note that the relu func of MLP in original DETR repo is set
+ # 'inplace=False', however the ReLU cfg of FFN in mmdet is set
+ # 'inplace=True' by default.
+ layers.append(nn.Linear(hidden_dim, output_dim))
+ return nn.Sequential(*layers)
+
+
+@MODELS.register_module()
+class DinoTransformerDecoder(DeformableDetrTransformerDecoder):
+
+ def __init__(self, *args, **kwargs):
+ super(DinoTransformerDecoder, self).__init__(*args, **kwargs)
+ self._init_layers()
+
+ def _init_layers(self):
+ self.ref_point_head = build_MLP(self.embed_dims * 2, self.embed_dims,
+ self.embed_dims, 2)
+ self.norm = nn.LayerNorm(self.embed_dims)
+
+ @staticmethod
+ def gen_sineembed_for_position(pos_tensor, pos_feat):
+ # n_query, bs, _ = pos_tensor.size()
+ # sineembed_tensor = torch.zeros(n_query, bs, 256)
+ scale = 2 * math.pi
+ dim_t = torch.arange(
+ pos_feat, dtype=torch.float32, device=pos_tensor.device)
+ dim_t = 10000**(2 * (dim_t // 2) / pos_feat)
+ x_embed = pos_tensor[:, :, 0] * scale
+ y_embed = pos_tensor[:, :, 1] * scale
+ pos_x = x_embed[:, :, None] / dim_t
+ pos_y = y_embed[:, :, None] / dim_t
+ pos_x = torch.stack((pos_x[:, :, 0::2].sin(), pos_x[:, :, 1::2].cos()),
+ dim=3).flatten(2)
+ pos_y = torch.stack((pos_y[:, :, 0::2].sin(), pos_y[:, :, 1::2].cos()),
+ dim=3).flatten(2)
+ if pos_tensor.size(-1) == 2:
+ pos = torch.cat((pos_y, pos_x), dim=2)
+ elif pos_tensor.size(-1) == 4:
+ w_embed = pos_tensor[:, :, 2] * scale
+ pos_w = w_embed[:, :, None] / dim_t
+ pos_w = torch.stack(
+ (pos_w[:, :, 0::2].sin(), pos_w[:, :, 1::2].cos()),
+ dim=3).flatten(2)
+
+ h_embed = pos_tensor[:, :, 3] * scale
+ pos_h = h_embed[:, :, None] / dim_t
+ pos_h = torch.stack(
+ (pos_h[:, :, 0::2].sin(), pos_h[:, :, 1::2].cos()),
+ dim=3).flatten(2)
+
+ pos = torch.cat((pos_y, pos_x, pos_w, pos_h), dim=2)
+ else:
+ raise ValueError('Unknown pos_tensor shape(-1):{}'.format(
+ pos_tensor.size(-1)))
+ return pos
+
+ def forward(self,
+ query,
+ *args,
+ reference_points=None,
+ valid_ratios=None,
+ reg_branches=None,
+ **kwargs):
+ output = query
+ intermediate = []
+ intermediate_reference_points = [reference_points]
+ for lid, layer in enumerate(self.layers):
+ if reference_points.shape[-1] == 4:
+ reference_points_input = \
+ reference_points[:, :, None] * torch.cat(
+ [valid_ratios, valid_ratios], -1)[:, None]
+ else:
+ assert reference_points.shape[-1] == 2
+ reference_points_input = \
+ reference_points[:, :, None] * valid_ratios[:, None]
+
+ query_sine_embed = self.gen_sineembed_for_position(
+ reference_points_input[:, :, 0, :], self.embed_dims // 2)
+ query_pos = self.ref_point_head(query_sine_embed)
+
+ query_pos = query_pos.permute(1, 0, 2)
+ output = layer(
+ output,
+ *args,
+ query_pos=query_pos,
+ reference_points=reference_points_input,
+ **kwargs)
+ output = output.permute(1, 0, 2)
+
+ if reg_branches is not None:
+ tmp = reg_branches[lid](output)
+ assert reference_points.shape[-1] == 4
+ new_reference_points = tmp + inverse_sigmoid(
+ reference_points, eps=1e-3)
+ new_reference_points = new_reference_points.sigmoid()
+ reference_points = new_reference_points.detach()
+
+ output = output.permute(1, 0, 2)
+ if self.return_intermediate:
+ intermediate.append(self.norm(output))
+ intermediate_reference_points.append(new_reference_points)
+ # NOTE this is for the "Look Forward Twice" module,
+ # in the DeformDETR, reference_points was appended.
+
+ if self.return_intermediate:
+ return torch.stack(intermediate), torch.stack(
+ intermediate_reference_points)
+
+ return output, reference_points
+
+
+@MODELS.register_module()
+class CoDinoTransformer(CoDeformableDetrTransformer):
+
+ def __init__(self, *args, **kwargs):
+ super(CoDinoTransformer, self).__init__(*args, **kwargs)
+
+ def init_layers(self):
+ """Initialize layers of the DinoTransformer."""
+ self.level_embeds = nn.Parameter(
+ torch.Tensor(self.num_feature_levels, self.embed_dims))
+ self.enc_output = nn.Linear(self.embed_dims, self.embed_dims)
+ self.enc_output_norm = nn.LayerNorm(self.embed_dims)
+ self.query_embed = nn.Embedding(self.two_stage_num_proposals,
+ self.embed_dims)
+
+ def _init_layers(self):
+ if self.with_pos_coord:
+ if self.num_co_heads > 0:
+ self.aux_pos_trans = nn.ModuleList()
+ self.aux_pos_trans_norm = nn.ModuleList()
+ self.pos_feats_trans = nn.ModuleList()
+ self.pos_feats_norm = nn.ModuleList()
+ for i in range(self.num_co_heads):
+ self.aux_pos_trans.append(
+ nn.Linear(self.embed_dims * 2, self.embed_dims))
+ self.aux_pos_trans_norm.append(
+ nn.LayerNorm(self.embed_dims))
+ if self.with_coord_feat:
+ self.pos_feats_trans.append(
+ nn.Linear(self.embed_dims, self.embed_dims))
+ self.pos_feats_norm.append(
+ nn.LayerNorm(self.embed_dims))
+
+ def init_weights(self):
+ super().init_weights()
+ nn.init.normal_(self.query_embed.weight.data)
+
+ def forward(self,
+ mlvl_feats,
+ mlvl_masks,
+ query_embed,
+ mlvl_pos_embeds,
+ dn_label_query,
+ dn_bbox_query,
+ attn_mask,
+ reg_branches=None,
+ cls_branches=None,
+ **kwargs):
+ assert self.as_two_stage and query_embed is None, \
+ 'as_two_stage must be True for DINO'
+
+ feat_flatten = []
+ mask_flatten = []
+ lvl_pos_embed_flatten = []
+ spatial_shapes = []
+ for lvl, (feat, mask, pos_embed) in enumerate(
+ zip(mlvl_feats, mlvl_masks, mlvl_pos_embeds)):
+ bs, c, h, w = feat.shape
+ spatial_shape = (h, w)
+ spatial_shapes.append(spatial_shape)
+ feat = feat.flatten(2).transpose(1, 2)
+ mask = mask.flatten(1)
+ pos_embed = pos_embed.flatten(2).transpose(1, 2)
+ lvl_pos_embed = pos_embed + self.level_embeds[lvl].view(1, 1, -1)
+ lvl_pos_embed_flatten.append(lvl_pos_embed)
+ feat_flatten.append(feat)
+ mask_flatten.append(mask)
+ feat_flatten = torch.cat(feat_flatten, 1)
+ mask_flatten = torch.cat(mask_flatten, 1)
+ lvl_pos_embed_flatten = torch.cat(lvl_pos_embed_flatten, 1)
+ spatial_shapes = torch.as_tensor(
+ spatial_shapes, dtype=torch.long, device=feat_flatten.device)
+ level_start_index = torch.cat((spatial_shapes.new_zeros(
+ (1, )), spatial_shapes.prod(1).cumsum(0)[:-1]))
+ valid_ratios = torch.stack(
+ [self.get_valid_ratio(m) for m in mlvl_masks], 1)
+
+ reference_points = self.get_reference_points(
+ spatial_shapes, valid_ratios, device=feat.device)
+
+ feat_flatten = feat_flatten.permute(1, 0, 2) # (H*W, bs, embed_dims)
+ lvl_pos_embed_flatten = lvl_pos_embed_flatten.permute(
+ 1, 0, 2) # (H*W, bs, embed_dims)
+ memory = self.encoder(
+ query=feat_flatten,
+ key=None,
+ value=None,
+ query_pos=lvl_pos_embed_flatten,
+ query_key_padding_mask=mask_flatten,
+ spatial_shapes=spatial_shapes,
+ reference_points=reference_points,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios,
+ **kwargs)
+ memory = memory.permute(1, 0, 2)
+ bs, _, c = memory.shape
+
+ output_memory, output_proposals = self.gen_encoder_output_proposals(
+ memory, mask_flatten, spatial_shapes)
+ enc_outputs_class = cls_branches[self.decoder.num_layers](
+ output_memory)
+ enc_outputs_coord_unact = reg_branches[self.decoder.num_layers](
+ output_memory) + output_proposals
+ cls_out_features = cls_branches[self.decoder.num_layers].out_features
+ topk = self.two_stage_num_proposals
+ # NOTE In DeformDETR, enc_outputs_class[..., 0] is used for topk
+ topk_indices = torch.topk(enc_outputs_class.max(-1)[0], topk, dim=1)[1]
+
+ topk_score = torch.gather(
+ enc_outputs_class, 1,
+ topk_indices.unsqueeze(-1).repeat(1, 1, cls_out_features))
+ topk_coords_unact = torch.gather(
+ enc_outputs_coord_unact, 1,
+ topk_indices.unsqueeze(-1).repeat(1, 1, 4))
+ topk_anchor = topk_coords_unact.sigmoid()
+ topk_coords_unact = topk_coords_unact.detach()
+
+ query = self.query_embed.weight[:, None, :].repeat(1, bs,
+ 1).transpose(0, 1)
+ # NOTE the query_embed here is not spatial query as in DETR.
+ # It is actually content query, which is named tgt in other
+ # DETR-like models
+ if dn_label_query is not None:
+ query = torch.cat([dn_label_query, query], dim=1)
+ if dn_bbox_query is not None:
+ reference_points = torch.cat([dn_bbox_query, topk_coords_unact],
+ dim=1)
+ else:
+ reference_points = topk_coords_unact
+ reference_points = reference_points.sigmoid()
+ # decoder
+ query = query.permute(1, 0, 2)
+ memory = memory.permute(1, 0, 2)
+ inter_states, inter_references = self.decoder(
+ query=query,
+ key=None,
+ value=memory,
+ attn_masks=attn_mask,
+ key_padding_mask=mask_flatten,
+ reference_points=reference_points,
+ spatial_shapes=spatial_shapes,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios,
+ reg_branches=reg_branches,
+ **kwargs)
+
+ inter_references_out = inter_references
+
+ return inter_states, inter_references_out, \
+ topk_score, topk_anchor, memory
+
+ def forward_aux(self,
+ mlvl_feats,
+ mlvl_masks,
+ query_embed,
+ mlvl_pos_embeds,
+ pos_anchors,
+ pos_feats=None,
+ reg_branches=None,
+ cls_branches=None,
+ return_encoder_output=False,
+ attn_masks=None,
+ head_idx=0,
+ **kwargs):
+ feat_flatten = []
+ mask_flatten = []
+ spatial_shapes = []
+ for lvl, (feat, mask, pos_embed) in enumerate(
+ zip(mlvl_feats, mlvl_masks, mlvl_pos_embeds)):
+ bs, c, h, w = feat.shape
+ spatial_shape = (h, w)
+ spatial_shapes.append(spatial_shape)
+ feat = feat.flatten(2).transpose(1, 2)
+ mask = mask.flatten(1)
+ feat_flatten.append(feat)
+ mask_flatten.append(mask)
+ feat_flatten = torch.cat(feat_flatten, 1)
+ mask_flatten = torch.cat(mask_flatten, 1)
+ spatial_shapes = torch.as_tensor(
+ spatial_shapes, dtype=torch.long, device=feat_flatten.device)
+ level_start_index = torch.cat((spatial_shapes.new_zeros(
+ (1, )), spatial_shapes.prod(1).cumsum(0)[:-1]))
+ valid_ratios = torch.stack(
+ [self.get_valid_ratio(m) for m in mlvl_masks], 1)
+
+ feat_flatten = feat_flatten.permute(1, 0, 2) # (H*W, bs, embed_dims)
+
+ memory = feat_flatten
+ memory = memory.permute(1, 0, 2)
+ bs, _, c = memory.shape
+
+ topk_coords_unact = inverse_sigmoid(pos_anchors)
+ reference_points = pos_anchors
+ if self.num_co_heads > 0:
+ pos_trans_out = self.aux_pos_trans_norm[head_idx](
+ self.aux_pos_trans[head_idx](
+ self.get_proposal_pos_embed(topk_coords_unact)))
+ query = pos_trans_out
+ if self.with_coord_feat:
+ query = query + self.pos_feats_norm[head_idx](
+ self.pos_feats_trans[head_idx](pos_feats))
+
+ # decoder
+ query = query.permute(1, 0, 2)
+ memory = memory.permute(1, 0, 2)
+ inter_states, inter_references = self.decoder(
+ query=query,
+ key=None,
+ value=memory,
+ attn_masks=None,
+ key_padding_mask=mask_flatten,
+ reference_points=reference_points,
+ spatial_shapes=spatial_shapes,
+ level_start_index=level_start_index,
+ valid_ratios=valid_ratios,
+ reg_branches=reg_branches,
+ **kwargs)
+
+ inter_references_out = inter_references
+
+ return inter_states, inter_references_out
+
+
+@MODELS.register_module()
+class DetrTransformerEncoder(TransformerLayerSequence):
+ """TransformerEncoder of DETR.
+
+ Args:
+ post_norm_cfg (dict): Config of last normalization layer. Default:
+ `LN`. Only used when `self.pre_norm` is `True`
+ """
+
+ def __init__(self,
+ *args,
+ post_norm_cfg=dict(type='LN'),
+ with_cp=-1,
+ **kwargs):
+ super(DetrTransformerEncoder, self).__init__(*args, **kwargs)
+ if post_norm_cfg is not None:
+ self.post_norm = build_norm_layer(
+ post_norm_cfg, self.embed_dims)[1] if self.pre_norm else None
+ else:
+ assert not self.pre_norm, f'Use prenorm in ' \
+ f'{self.__class__.__name__},' \
+ f'Please specify post_norm_cfg'
+ self.post_norm = None
+ self.with_cp = with_cp
+ if self.with_cp > 0:
+ if checkpoint_wrapper is None:
+ warnings.warn('If you want to reduce GPU memory usage, \
+ please install fairscale by executing the \
+ following command: pip install fairscale.')
+ return
+ for i in range(self.with_cp):
+ self.layers[i] = checkpoint_wrapper(self.layers[i])
+
+
+@MODELS.register_module()
+class DetrTransformerDecoderLayer(BaseTransformerLayer):
+ """Implements decoder layer in DETR transformer.
+
+ Args:
+ attn_cfgs (list[`mmcv.ConfigDict`] | list[dict] | dict )):
+ Configs for self_attention or cross_attention, the order
+ should be consistent with it in `operation_order`. If it is
+ a dict, it would be expand to the number of attention in
+ `operation_order`.
+ feedforward_channels (int): The hidden dimension for FFNs.
+ ffn_dropout (float): Probability of an element to be zeroed
+ in ffn. Default 0.0.
+ operation_order (tuple[str]): The execution order of operation
+ in transformer. Such as ('self_attn', 'norm', 'ffn', 'norm').
+ Default:None
+ act_cfg (dict): The activation config for FFNs. Default: `LN`
+ norm_cfg (dict): Config dict for normalization layer.
+ Default: `LN`.
+ ffn_num_fcs (int): The number of fully-connected layers in FFNs.
+ Default:2.
+ """
+
+ def __init__(self,
+ attn_cfgs,
+ feedforward_channels,
+ ffn_dropout=0.0,
+ operation_order=None,
+ act_cfg=dict(type='ReLU', inplace=True),
+ norm_cfg=dict(type='LN'),
+ ffn_num_fcs=2,
+ **kwargs):
+ super(DetrTransformerDecoderLayer, self).__init__(
+ attn_cfgs=attn_cfgs,
+ feedforward_channels=feedforward_channels,
+ ffn_dropout=ffn_dropout,
+ operation_order=operation_order,
+ act_cfg=act_cfg,
+ norm_cfg=norm_cfg,
+ ffn_num_fcs=ffn_num_fcs,
+ **kwargs)
+ assert len(operation_order) == 6
+ assert set(operation_order) == set(
+ ['self_attn', 'norm', 'cross_attn', 'ffn'])
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_r50_8xb2_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_r50_8xb2_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..1a4130437666428213eb3250f8eee9d2a4d1442b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_r50_8xb2_1x_coco.py
@@ -0,0 +1,68 @@
+_base_ = './co_dino_5scale_r50_lsj_8xb2_1x_coco.py'
+
+model = dict(
+ use_lsj=False, data_preprocessor=dict(pad_mask=False, batch_augments=None))
+
+# train_pipeline, NOTE the img_scale and the Pad's size_divisor is different
+# from the default setting in mmdet.
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 1333), (512, 1333), (544, 1333), (576, 1333),
+ (608, 1333), (640, 1333), (672, 1333), (704, 1333),
+ (736, 1333), (768, 1333), (800, 1333)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ _delete_=True,
+ type=_base_.dataset_type,
+ data_root=_base_.data_root,
+ ann_file='annotations/instances_train2017.json',
+ data_prefix=dict(img='train2017/'),
+ filter_cfg=dict(filter_empty_gt=False, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=_base_.backend_args))
+
+test_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_r50_lsj_8xb2_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_r50_lsj_8xb2_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..876b90f89c8795186d830689c9bdb420b0cfbb18
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_r50_lsj_8xb2_1x_coco.py
@@ -0,0 +1,359 @@
+_base_ = 'mmdet::common/ssj_scp_270k_coco-instance.py'
+
+custom_imports = dict(
+ imports=['projects.CO-DETR.codetr'], allow_failed_imports=False)
+
+# model settings
+num_dec_layer = 6
+loss_lambda = 2.0
+num_classes = 80
+
+image_size = (1024, 1024)
+batch_augments = [
+ dict(type='BatchFixedSizePad', size=image_size, pad_mask=True)
+]
+model = dict(
+ type='CoDETR',
+ # If using the lsj augmentation,
+ # it is recommended to set it to True.
+ use_lsj=True,
+ # detr: 52.1
+ # one-stage: 49.4
+ # two-stage: 47.9
+ eval_module='detr', # in ['detr', 'one-stage', 'two-stage']
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_mask=True,
+ batch_augments=batch_augments),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(0, 1, 2, 3),
+ frozen_stages=1,
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ style='pytorch',
+ init_cfg=dict(type='Pretrained', checkpoint='torchvision://resnet50')),
+ neck=dict(
+ type='ChannelMapper',
+ in_channels=[256, 512, 1024, 2048],
+ kernel_size=1,
+ out_channels=256,
+ act_cfg=None,
+ norm_cfg=dict(type='GN', num_groups=32),
+ num_outs=5),
+ query_head=dict(
+ type='CoDINOHead',
+ num_query=900,
+ num_classes=num_classes,
+ in_channels=2048,
+ as_two_stage=True,
+ dn_cfg=dict(
+ label_noise_scale=0.5,
+ box_noise_scale=1.0,
+ group_cfg=dict(dynamic=True, num_groups=None, num_dn_queries=100)),
+ transformer=dict(
+ type='CoDinoTransformer',
+ with_coord_feat=False,
+ num_co_heads=2, # ATSS Aux Head + Faster RCNN Aux Head
+ num_feature_levels=5,
+ encoder=dict(
+ type='DetrTransformerEncoder',
+ num_layers=6,
+ # number of layers that use checkpoint.
+ # The maximum value for the setting is num_layers.
+ # FairScale must be installed for it to work.
+ with_cp=4,
+ transformerlayers=dict(
+ type='BaseTransformerLayer',
+ attn_cfgs=dict(
+ type='MultiScaleDeformableAttention',
+ embed_dims=256,
+ num_levels=5,
+ dropout=0.0),
+ feedforward_channels=2048,
+ ffn_dropout=0.0,
+ operation_order=('self_attn', 'norm', 'ffn', 'norm'))),
+ decoder=dict(
+ type='DinoTransformerDecoder',
+ num_layers=6,
+ return_intermediate=True,
+ transformerlayers=dict(
+ type='DetrTransformerDecoderLayer',
+ attn_cfgs=[
+ dict(
+ type='MultiheadAttention',
+ embed_dims=256,
+ num_heads=8,
+ dropout=0.0),
+ dict(
+ type='MultiScaleDeformableAttention',
+ embed_dims=256,
+ num_levels=5,
+ dropout=0.0),
+ ],
+ feedforward_channels=2048,
+ ffn_dropout=0.0,
+ operation_order=('self_attn', 'norm', 'cross_attn', 'norm',
+ 'ffn', 'norm')))),
+ positional_encoding=dict(
+ type='SinePositionalEncoding',
+ num_feats=128,
+ temperature=20,
+ normalize=True),
+ loss_cls=dict( # Different from the DINO
+ type='QualityFocalLoss',
+ use_sigmoid=True,
+ beta=2.0,
+ loss_weight=1.0),
+ loss_bbox=dict(type='L1Loss', loss_weight=5.0),
+ loss_iou=dict(type='GIoULoss', loss_weight=2.0)),
+ rpn_head=dict(
+ type='RPNHead',
+ in_channels=256,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ octave_base_scale=4,
+ scales_per_octave=3,
+ ratios=[0.5, 1.0, 2.0],
+ strides=[4, 8, 16, 32, 64, 128]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[1.0, 1.0, 1.0, 1.0]),
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ loss_weight=1.0 * num_dec_layer * loss_lambda),
+ loss_bbox=dict(
+ type='L1Loss', loss_weight=1.0 * num_dec_layer * loss_lambda)),
+ roi_head=[
+ dict(
+ type='CoStandardRoIHead',
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(
+ type='RoIAlign', output_size=7, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[4, 8, 16, 32, 64],
+ finest_scale=56),
+ bbox_head=dict(
+ type='Shared2FCBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=num_classes,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=False,
+ reg_decoded_bbox=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=False,
+ loss_weight=1.0 * num_dec_layer * loss_lambda),
+ loss_bbox=dict(
+ type='GIoULoss',
+ loss_weight=10.0 * num_dec_layer * loss_lambda)))
+ ],
+ bbox_head=[
+ dict(
+ type='CoATSSHead',
+ num_classes=num_classes,
+ in_channels=256,
+ stacked_convs=1,
+ feat_channels=256,
+ anchor_generator=dict(
+ type='AnchorGenerator',
+ ratios=[1.0],
+ octave_base_scale=8,
+ scales_per_octave=1,
+ strides=[4, 8, 16, 32, 64, 128]),
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[.0, .0, .0, .0],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ loss_cls=dict(
+ type='FocalLoss',
+ use_sigmoid=True,
+ gamma=2.0,
+ alpha=0.25,
+ loss_weight=1.0 * num_dec_layer * loss_lambda),
+ loss_bbox=dict(
+ type='GIoULoss',
+ loss_weight=2.0 * num_dec_layer * loss_lambda),
+ loss_centerness=dict(
+ type='CrossEntropyLoss',
+ use_sigmoid=True,
+ loss_weight=1.0 * num_dec_layer * loss_lambda)),
+ ],
+ # model training and testing settings
+ train_cfg=[
+ dict(
+ assigner=dict(
+ type='HungarianAssigner',
+ match_costs=[
+ dict(type='FocalLossCost', weight=2.0),
+ dict(type='BBoxL1Cost', weight=5.0, box_format='xywh'),
+ dict(type='IoUCost', iou_mode='giou', weight=2.0)
+ ])),
+ dict(
+ rpn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ match_low_quality=True,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False),
+ rpn_proposal=dict(
+ nms_pre=4000,
+ max_per_img=1000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.5,
+ neg_iou_thr=0.5,
+ min_pos_iou=0.5,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ pos_weight=-1,
+ debug=False)),
+ dict(
+ assigner=dict(type='ATSSAssigner', topk=9),
+ allowed_border=-1,
+ pos_weight=-1,
+ debug=False)
+ ],
+ test_cfg=[
+ # Deferent from the DINO, we use the NMS.
+ dict(
+ max_per_img=300,
+ # NMS can improve the mAP by 0.2.
+ nms=dict(type='soft_nms', iou_threshold=0.8)),
+ dict(
+ rpn=dict(
+ nms_pre=1000,
+ max_per_img=1000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=dict(
+ score_thr=0.0,
+ nms=dict(type='nms', iou_threshold=0.5),
+ max_per_img=100)),
+ dict(
+ # atss bbox head:
+ nms_pre=1000,
+ min_bbox_size=0,
+ score_thr=0.0,
+ nms=dict(type='nms', iou_threshold=0.6),
+ max_per_img=100),
+ # soft-nms is also supported for rcnn testing
+ # e.g., nms=dict(type='soft_nms', iou_threshold=0.5, min_score=0.05)
+ ])
+
+# LSJ + CopyPaste
+load_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomResize',
+ scale=image_size,
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size,
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size=image_size, pad_val=dict(img=(114, 114, 114))),
+]
+
+train_pipeline = [
+ dict(type='CopyPaste', max_num_pasted=100),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ dataset=dict(
+ pipeline=train_pipeline,
+ dataset=dict(
+ filter_cfg=dict(filter_empty_gt=False), pipeline=load_pipeline)))
+
+# follow ViTDet
+test_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='Resize', scale=image_size, keep_ratio=True), # diff
+ dict(type='Pad', size=image_size, pad_val=dict(img=(114, 114, 114))),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+optim_wrapper = dict(
+ _delete_=True,
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=2e-4, weight_decay=0.0001),
+ clip_grad=dict(max_norm=0.1, norm_type=2),
+ paramwise_cfg=dict(custom_keys={'backbone': dict(lr_mult=0.1)}))
+
+val_evaluator = dict(metric='bbox')
+test_evaluator = val_evaluator
+
+max_epochs = 12
+train_cfg = dict(
+ _delete_=True,
+ type='EpochBasedTrainLoop',
+ max_epochs=max_epochs,
+ val_interval=1)
+
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[11],
+ gamma=0.1)
+]
+
+default_hooks = dict(
+ checkpoint=dict(by_epoch=True, interval=1, max_keep_ckpts=3))
+log_processor = dict(by_epoch=True)
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (2 samples per GPU)
+auto_scale_lr = dict(base_batch_size=16)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_r50_lsj_8xb2_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_r50_lsj_8xb2_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..9a9fc34f680a3de3f96a548817f3d4e37983fee7
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_r50_lsj_8xb2_3x_coco.py
@@ -0,0 +1,4 @@
+_base_ = ['co_dino_5scale_r50_lsj_8xb2_1x_coco.py']
+
+param_scheduler = [dict(milestones=[30])]
+train_cfg = dict(max_epochs=36)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_swin_l_16xb1_16e_o365tococo.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_swin_l_16xb1_16e_o365tococo.py
new file mode 100644
index 0000000000000000000000000000000000000000..77821c380f3407c2288377dc78232fd12205fc76
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_swin_l_16xb1_16e_o365tococo.py
@@ -0,0 +1,115 @@
+_base_ = ['co_dino_5scale_r50_8xb2_1x_coco.py']
+
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_large_patch4_window12_384_22k.pth' # noqa
+load_from = 'https://download.openmmlab.com/mmdetection/v3.0/codetr/co_dino_5scale_swin_large_16e_o365tococo-614254c9.pth' # noqa
+
+# model settings
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ pretrain_img_size=384,
+ embed_dims=192,
+ depths=[2, 2, 18, 2],
+ num_heads=[6, 12, 24, 48],
+ window_size=12,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(0, 1, 2, 3),
+ # Please only add indices that would be used
+ # in FPN, otherwise some parameter will not be used
+ with_cp=True,
+ convert_weights=True,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ neck=dict(in_channels=[192, 384, 768, 1536]),
+ query_head=dict(
+ dn_cfg=dict(box_noise_scale=0.4, group_cfg=dict(num_dn_queries=500)),
+ transformer=dict(encoder=dict(with_cp=6))))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(
+ type='RandomChoice',
+ transforms=[
+ [
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 2048), (512, 2048), (544, 2048), (576, 2048),
+ (608, 2048), (640, 2048), (672, 2048), (704, 2048),
+ (736, 2048), (768, 2048), (800, 2048), (832, 2048),
+ (864, 2048), (896, 2048), (928, 2048), (960, 2048),
+ (992, 2048), (1024, 2048), (1056, 2048),
+ (1088, 2048), (1120, 2048), (1152, 2048),
+ (1184, 2048), (1216, 2048), (1248, 2048),
+ (1280, 2048), (1312, 2048), (1344, 2048),
+ (1376, 2048), (1408, 2048), (1440, 2048),
+ (1472, 2048), (1504, 2048), (1536, 2048)],
+ keep_ratio=True)
+ ],
+ [
+ dict(
+ type='RandomChoiceResize',
+ # The radio of all image in train dataset < 7
+ # follow the original implement
+ scales=[(400, 4200), (500, 4200), (600, 4200)],
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(384, 600),
+ allow_negative_crop=True),
+ dict(
+ type='RandomChoiceResize',
+ scales=[(480, 2048), (512, 2048), (544, 2048), (576, 2048),
+ (608, 2048), (640, 2048), (672, 2048), (704, 2048),
+ (736, 2048), (768, 2048), (800, 2048), (832, 2048),
+ (864, 2048), (896, 2048), (928, 2048), (960, 2048),
+ (992, 2048), (1024, 2048), (1056, 2048),
+ (1088, 2048), (1120, 2048), (1152, 2048),
+ (1184, 2048), (1216, 2048), (1248, 2048),
+ (1280, 2048), (1312, 2048), (1344, 2048),
+ (1376, 2048), (1408, 2048), (1440, 2048),
+ (1472, 2048), (1504, 2048), (1536, 2048)],
+ keep_ratio=True)
+ ]
+ ]),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(
+ batch_size=1, num_workers=1, dataset=dict(pipeline=train_pipeline))
+
+test_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='Resize', scale=(2048, 1280), keep_ratio=True),
+ dict(type='LoadAnnotations', with_bbox=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+
+optim_wrapper = dict(optimizer=dict(lr=1e-4))
+
+max_epochs = 16
+train_cfg = dict(max_epochs=max_epochs)
+
+param_scheduler = [
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[8],
+ gamma=0.1)
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_swin_l_16xb1_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_swin_l_16xb1_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..d4a873464d422334a42d72543bbccc3b344aa97e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_swin_l_16xb1_1x_coco.py
@@ -0,0 +1,31 @@
+_base_ = ['co_dino_5scale_r50_8xb2_1x_coco.py']
+
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_large_patch4_window12_384_22k.pth' # noqa
+
+# model settings
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ pretrain_img_size=384,
+ embed_dims=192,
+ depths=[2, 2, 18, 2],
+ num_heads=[6, 12, 24, 48],
+ window_size=12,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(0, 1, 2, 3),
+ # Please only add indices that would be used
+ # in FPN, otherwise some parameter will not be used
+ with_cp=False,
+ convert_weights=True,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ neck=dict(in_channels=[192, 384, 768, 1536]),
+ query_head=dict(transformer=dict(encoder=dict(with_cp=6))))
+
+train_dataloader = dict(batch_size=1, num_workers=1)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_swin_l_16xb1_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_swin_l_16xb1_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..c2fce29b98b5ffe7e51396b8b88b289fc4c8ffbc
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_swin_l_16xb1_3x_coco.py
@@ -0,0 +1,6 @@
+_base_ = ['co_dino_5scale_swin_l_16xb1_1x_coco.py']
+# model settings
+model = dict(backbone=dict(drop_path_rate=0.6))
+
+param_scheduler = [dict(milestones=[30])]
+train_cfg = dict(max_epochs=36)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_swin_l_lsj_16xb1_1x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_swin_l_lsj_16xb1_1x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..4a9b3688b8ebf6525f4d96526dd543576ae6253b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_swin_l_lsj_16xb1_1x_coco.py
@@ -0,0 +1,72 @@
+_base_ = ['co_dino_5scale_r50_lsj_8xb2_1x_coco.py']
+
+image_size = (1280, 1280)
+batch_augments = [
+ dict(type='BatchFixedSizePad', size=image_size, pad_mask=True)
+]
+pretrained = 'https://github.com/SwinTransformer/storage/releases/download/v1.0.0/swin_large_patch4_window12_384_22k.pth' # noqa
+
+# model settings
+model = dict(
+ data_preprocessor=dict(batch_augments=batch_augments),
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ pretrain_img_size=384,
+ embed_dims=192,
+ depths=[2, 2, 18, 2],
+ num_heads=[6, 12, 24, 48],
+ window_size=12,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(0, 1, 2, 3),
+ # Please only add indices that would be used
+ # in FPN, otherwise some parameter will not be used
+ with_cp=False,
+ convert_weights=True,
+ init_cfg=dict(type='Pretrained', checkpoint=pretrained)),
+ neck=dict(in_channels=[192, 384, 768, 1536]),
+ query_head=dict(transformer=dict(encoder=dict(with_cp=6))))
+
+load_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomResize',
+ scale=image_size,
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size,
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='Pad', size=image_size, pad_val=dict(img=(114, 114, 114))),
+]
+
+train_dataloader = dict(
+ batch_size=1,
+ num_workers=1,
+ dataset=dict(dataset=dict(pipeline=load_pipeline)))
+
+test_pipeline = [
+ dict(type='LoadImageFromFile'),
+ dict(type='Resize', scale=image_size, keep_ratio=True),
+ dict(type='Pad', size=image_size, pad_val=dict(img=(114, 114, 114))),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_swin_l_lsj_16xb1_3x_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_swin_l_lsj_16xb1_3x_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..bf9cd4f439287d7174f9b773b7177ade179cd536
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/CO-DETR/configs/codino/co_dino_5scale_swin_l_lsj_16xb1_3x_coco.py
@@ -0,0 +1,7 @@
+_base_ = ['co_dino_5scale_swin_l_lsj_16xb1_1x_coco.py']
+
+model = dict(backbone=dict(drop_path_rate=0.5))
+
+param_scheduler = [dict(type='MultiStepLR', milestones=[30])]
+
+train_cfg = dict(max_epochs=36)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/ConvNeXt-V2/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/projects/ConvNeXt-V2/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..7a9f56cd24777130a4301573d44bcfd655f721b8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/ConvNeXt-V2/README.md
@@ -0,0 +1,37 @@
+# ConvNeXt-V2
+
+> [ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders](http://arxiv.org/abs/2301.00808)
+
+## Abstract
+
+Driven by improved architectures and better representation learning frameworks, the field of visual recognition has enjoyed rapid modernization and performance boost in the early 2020s. For example, modern ConvNets, represented by ConvNeXt \[52\], have demonstrated strong performance in various scenarios. While these models were originally designed for supervised learning with ImageNet labels, they can also potentially benefit from self-supervised learning techniques such as masked autoencoders (MAE) . However, we found that simply combining these two approaches leads to subpar performance. In this paper, we propose a fully convolutional masked autoencoder framework and a new Global Response Normalization (GRN) layer that can be added to the ConvNeXt architecture to enhance inter-channel feature competition. This co-design of self-supervised learning techniques and architectural improvement results in a new model family called ConvNeXt V2, which significantly improves the performance of pure ConvNets on various recognition benchmarks, including ImageNet classification, COCO detection, and ADE20K segmentation. We also provide pre-trained ConvNeXt V2 models of various sizes, ranging from an efficient 3.7Mparameter Atto model with 76.7% top-1 accuracy on Im-ageNet, to a 650M Huge model that achieves a state-of-theart 88.9% accuracy using only public training data.
+
+
+

+
+
+## Results and models
+
+| Method | Backbone | Pretrain | Lr schd | Augmentation | Mem (GB) | box AP | mask AP | Config | Download |
+| :--------: | :-----------: | :------: | :-----: | :----------: | :------: | :----: | :-----: | :----------------------------------------------------------: | :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Mask R-CNN | ConvNeXt-V2-B | FCMAE | 3x | LSJ | 22.5 | 52.9 | 46.4 | [config](./mask-rcnn_convnext-v2-b_fpn_lsj-3x-fcmae_coco.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/convnextv2/mask-rcnn_convnext-v2-b_fpn_lsj-3x-fcmae_coco/mask-rcnn_convnext-v2-b_fpn_lsj-3x-fcmae_coco_20230113_110947-757ee2dd.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/convnextv2/mask-rcnn_convnext-v2-b_fpn_lsj-3x-fcmae_coco/mask-rcnn_convnext-v2-b_fpn_lsj-3x-fcmae_coco_20230113_110947.log.json) |
+
+**Note**:
+
+- This is a pre-release version of ConvNeXt-V2 object detection. The official finetuning setting of ConvNeXt-V2 has not been released yet.
+- ConvNeXt backbone needs to install [MMPretrain](https://github.com/open-mmlab/mmpretrain/) first, which has abundant backbones for downstream tasks.
+
+```shell
+pip install mmpretrain
+```
+
+## Citation
+
+```bibtex
+@article{Woo2023ConvNeXtV2,
+ title={ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders},
+ author={Sanghyun Woo, Shoubhik Debnath, Ronghang Hu, Xinlei Chen, Zhuang Liu, In So Kweon and Saining Xie},
+ year={2023},
+ journal={arXiv preprint arXiv:2301.00808},
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/ConvNeXt-V2/configs/mask-rcnn_convnext-v2-b_fpn_lsj-3x-fcmae_coco.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/ConvNeXt-V2/configs/mask-rcnn_convnext-v2-b_fpn_lsj-3x-fcmae_coco.py
new file mode 100644
index 0000000000000000000000000000000000000000..59e89550459c18d49185a20de9df140adab54b05
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/ConvNeXt-V2/configs/mask-rcnn_convnext-v2-b_fpn_lsj-3x-fcmae_coco.py
@@ -0,0 +1,92 @@
+_base_ = [
+ 'mmdet::_base_/models/mask-rcnn_r50_fpn.py',
+ 'mmdet::_base_/datasets/coco_instance.py',
+ 'mmdet::_base_/schedules/schedule_1x.py',
+ 'mmdet::_base_/default_runtime.py'
+]
+
+# please install the mmpretrain
+# import mmpretrain.models to trigger register_module in mmpretrain
+custom_imports = dict(
+ imports=['mmpretrain.models'], allow_failed_imports=False)
+checkpoint_file = 'https://download.openmmlab.com/mmclassification/v0/convnext-v2/convnext-v2-base_3rdparty-fcmae_in1k_20230104-8a798eaf.pth' # noqa
+image_size = (1024, 1024)
+
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='mmpretrain.ConvNeXt',
+ arch='base',
+ out_indices=[0, 1, 2, 3],
+ # TODO: verify stochastic depth rate {0.1, 0.2, 0.3, 0.4}
+ drop_path_rate=0.4,
+ layer_scale_init_value=0., # disable layer scale when using GRN
+ gap_before_final_norm=False,
+ use_grn=True, # V2 uses GRN
+ init_cfg=dict(
+ type='Pretrained', checkpoint=checkpoint_file,
+ prefix='backbone.')),
+ neck=dict(in_channels=[128, 256, 512, 1024]),
+ test_cfg=dict(
+ rpn=dict(nms=dict(type='nms')), # TODO: does RPN use soft_nms?
+ rcnn=dict(nms=dict(type='soft_nms'))))
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=_base_.backend_args),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomResize',
+ scale=image_size,
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size,
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(
+ batch_size=4, # total_batch_size 32 = 8 GPUS x 4 images
+ num_workers=8,
+ dataset=dict(pipeline=train_pipeline))
+
+max_epochs = 36
+train_cfg = dict(max_epochs=max_epochs)
+
+# learning rate
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0,
+ end=1000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=max_epochs,
+ by_epoch=True,
+ milestones=[27, 33],
+ gamma=0.1)
+]
+
+# Enable automatic-mixed-precision training with AmpOptimWrapper.
+optim_wrapper = dict(
+ type='AmpOptimWrapper',
+ constructor='LearningRateDecayOptimizerConstructor',
+ paramwise_cfg={
+ 'decay_rate': 0.95,
+ 'decay_type': 'layer_wise', # TODO: sweep layer-wise lr decay?
+ 'num_layers': 12
+ },
+ optimizer=dict(
+ _delete_=True,
+ type='AdamW',
+ lr=0.0001,
+ betas=(0.9, 0.999),
+ weight_decay=0.05,
+ ))
+
+default_hooks = dict(checkpoint=dict(max_keep_ckpts=1))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..98cd705b040e507c4513dd60f77da2904962bbec
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/README.md
@@ -0,0 +1,156 @@
+# Note: This project has been deprecated, please use [Detic_new](../Detic_new).
+
+# Detecting Twenty-thousand Classes using Image-level Supervision
+
+## Description
+
+**Detic**: A **Det**ector with **i**mage **c**lasses that can use image-level labels to easily train detectors.
+
+
+
+> [**Detecting Twenty-thousand Classes using Image-level Supervision**](http://arxiv.org/abs/2201.02605),
+> Xingyi Zhou, Rohit Girdhar, Armand Joulin, Philipp Krähenbühl, Ishan Misra,
+> *ECCV 2022 ([arXiv 2201.02605](http://arxiv.org/abs/2201.02605))*
+
+## Usage
+
+
+
+## Installation
+
+Detic requires to install CLIP.
+
+```shell
+pip install git+https://github.com/openai/CLIP.git
+```
+
+### Demo
+
+#### Inference with existing dataset vocabulary embeddings
+
+First, go to the Detic project folder.
+
+```shell
+cd projects/Detic
+```
+
+Then, download the pre-computed CLIP embeddings from [dataset metainfo](https://github.com/facebookresearch/Detic/tree/main/datasets/metadata) to the `datasets/metadata` folder.
+The CLIP embeddings will be loaded to the zero-shot classifier during inference.
+For example, you can download LVIS's class name embeddings with the following command:
+
+```shell
+wget -P datasets/metadata https://raw.githubusercontent.com/facebookresearch/Detic/main/datasets/metadata/lvis_v1_clip_a%2Bcname.npy
+```
+
+You can run demo like this:
+
+```shell
+python demo.py \
+ ${IMAGE_PATH} \
+ ${CONFIG_PATH} \
+ ${MODEL_PATH} \
+ --show \
+ --score-thr 0.5 \
+ --dataset lvis
+```
+
+
+
+### Inference with custom vocabularies
+
+- Detic can detects any class given class names by using CLIP.
+
+You can detect custom classes with `--class-name` command:
+
+```
+python demo.py \
+ ${IMAGE_PATH} \
+ ${CONFIG_PATH} \
+ ${MODEL_PATH} \
+ --show \
+ --score-thr 0.3 \
+ --class-name headphone webcam paper coffe
+```
+
+
+
+Note that `headphone`, `paper` and `coffe` (typo intended) are not LVIS classes. Despite the misspelled class name, Detic can produce a reasonable detection for `coffe`.
+
+## Results
+
+Here we only provide the Detic Swin-B model for the open vocabulary demo. Multi-dataset training and open-vocabulary testing will be supported in the future.
+
+To find more variants, please visit the [official model zoo](https://github.com/facebookresearch/Detic/blob/main/docs/MODEL_ZOO.md).
+
+| Backbone | Training data | Config | Download |
+| :------: | :------------------------: | :-------------------------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Swin-B | ImageNet-21K & LVIS & COCO | [config](./configs/detic_centernet2_swin-b_fpn_4x_lvis-coco-in21k.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/detic/detic_centernet2_swin-b_fpn_4x_lvis-coco-in21k/detic_centernet2_swin-b_fpn_4x_lvis-coco-in21k_20230120-0d301978.pth) |
+
+## Citation
+
+If you find Detic is useful in your research or applications, please consider giving a star 🌟 to the [official repository](https://github.com/facebookresearch/Detic) and citing Detic by the following BibTeX entry.
+
+```BibTeX
+@inproceedings{zhou2022detecting,
+ title={Detecting Twenty-thousand Classes using Image-level Supervision},
+ author={Zhou, Xingyi and Girdhar, Rohit and Joulin, Armand and Kr{\"a}henb{\"u}hl, Philipp and Misra, Ishan},
+ booktitle={ECCV},
+ year={2022}
+}
+
+```
+
+## Checklist
+
+
+
+- [x] Milestone 1: PR-ready, and acceptable to be one of the `projects/`.
+
+ - [x] Finish the code
+
+
+
+ - [x] Basic docstrings & proper citation
+
+
+
+ - [x] Test-time correctness
+
+
+
+ - [x] A full README
+
+
+
+- [ ] Milestone 2: Indicates a successful model implementation.
+
+ - [ ] Training-time correctness
+
+
+
+- [ ] Milestone 3: Good to be a part of our core package!
+
+ - [ ] Type hints and docstrings
+
+
+
+ - [ ] Unit tests
+
+
+
+ - [ ] Code polishing
+
+
+
+ - [ ] Metafile.yml
+
+
+
+- [ ] Move your modules into the core package following the codebase's file hierarchy structure.
+
+
+
+- [ ] Refactor your modules into the core package following the codebase's file hierarchy structure.
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/configs/detic_centernet2_swin-b_fpn_4x_lvis-coco-in21k.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/configs/detic_centernet2_swin-b_fpn_4x_lvis-coco-in21k.py
new file mode 100644
index 0000000000000000000000000000000000000000..d554c40ec20bb77632711c357179d402cc978809
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/configs/detic_centernet2_swin-b_fpn_4x_lvis-coco-in21k.py
@@ -0,0 +1,298 @@
+_base_ = 'mmdet::common/lsj-200e_coco-detection.py'
+
+custom_imports = dict(
+ imports=['projects.Detic.detic'], allow_failed_imports=False)
+
+image_size = (1024, 1024)
+batch_augments = [dict(type='BatchFixedSizePad', size=image_size)]
+
+cls_layer = dict(
+ type='ZeroShotClassifier',
+ zs_weight_path='rand',
+ zs_weight_dim=512,
+ use_bias=0.0,
+ norm_weight=True,
+ norm_temperature=50.0)
+reg_layer = [
+ dict(type='Linear', in_features=1024, out_features=1024),
+ dict(type='ReLU', inplace=True),
+ dict(type='Linear', in_features=1024, out_features=4)
+]
+
+num_classes = 22047
+
+model = dict(
+ type='CascadeRCNN',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32,
+ batch_augments=batch_augments),
+ backbone=dict(
+ type='SwinTransformer',
+ embed_dims=128,
+ depths=[2, 2, 18, 2],
+ num_heads=[4, 8, 16, 32],
+ window_size=7,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(1, 2, 3),
+ with_cp=False),
+ neck=dict(
+ type='FPN',
+ in_channels=[256, 512, 1024],
+ out_channels=256,
+ start_level=0,
+ add_extra_convs='on_output',
+ num_outs=5,
+ init_cfg=dict(type='Caffe2Xavier', layer='Conv2d'),
+ relu_before_extra_convs=True),
+ rpn_head=dict(
+ type='CenterNetRPNHead',
+ num_classes=1,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ strides=[8, 16, 32, 64, 128],
+ conv_bias=True,
+ norm_cfg=dict(type='GN', num_groups=32, requires_grad=True),
+ loss_cls=dict(
+ type='GaussianFocalLoss',
+ pos_weight=0.25,
+ neg_weight=0.75,
+ loss_weight=1.0),
+ loss_bbox=dict(type='GIoULoss', loss_weight=2.0),
+ ),
+ roi_head=dict(
+ type='DeticRoIHead',
+ num_stages=3,
+ stage_loss_weights=[1, 0.5, 0.25],
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(
+ type='RoIAlign',
+ output_size=7,
+ sampling_ratio=0,
+ use_torchvision=True),
+ out_channels=256,
+ featmap_strides=[8, 16, 32],
+ # approximately equal to
+ # canonical_box_size=224, canonical_level=4 in D2
+ finest_scale=112),
+ bbox_head=[
+ dict(
+ type='DeticBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=num_classes,
+ cls_predictor_cfg=cls_layer,
+ reg_predictor_cfg=reg_layer,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0,
+ loss_weight=1.0)),
+ dict(
+ type='DeticBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=num_classes,
+ cls_predictor_cfg=cls_layer,
+ reg_predictor_cfg=reg_layer,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.05, 0.05, 0.1, 0.1]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0,
+ loss_weight=1.0)),
+ dict(
+ type='DeticBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=num_classes,
+ cls_predictor_cfg=cls_layer,
+ reg_predictor_cfg=reg_layer,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.033, 0.033, 0.067, 0.067]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=1.0, loss_weight=1.0))
+ ],
+ mask_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=14, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[8, 16, 32],
+ # approximately equal to
+ # canonical_box_size=224, canonical_level=4 in D2
+ finest_scale=112),
+ mask_head=dict(
+ type='FCNMaskHead',
+ num_convs=4,
+ in_channels=256,
+ conv_out_channels=256,
+ class_agnostic=True,
+ num_classes=num_classes,
+ loss_mask=dict(
+ type='CrossEntropyLoss', use_mask=True, loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ match_low_quality=True,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=0,
+ pos_weight=-1,
+ debug=False),
+ rpn_proposal=dict(
+ nms_pre=2000,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.7),
+ min_bbox_size=0),
+ rcnn=[
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.6,
+ neg_iou_thr=0.6,
+ min_pos_iou=0.6,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ mask_size=28,
+ pos_weight=-1,
+ debug=False),
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.7,
+ min_pos_iou=0.7,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ mask_size=28,
+ pos_weight=-1,
+ debug=False),
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.8,
+ neg_iou_thr=0.8,
+ min_pos_iou=0.8,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ mask_size=28,
+ pos_weight=-1,
+ debug=False)
+ ]),
+ test_cfg=dict(
+ rpn=dict(
+ score_thr=0.0001,
+ nms_pre=1000,
+ max_per_img=256,
+ nms=dict(type='nms', iou_threshold=0.9),
+ min_bbox_size=0),
+ rcnn=dict(
+ score_thr=0.02,
+ nms=dict(type='nms', iou_threshold=0.5),
+ max_per_img=300,
+ mask_thr_binary=0.5)))
+
+backend = 'pillow'
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ backend_args=_base_.backend_args,
+ imdecode_backend=backend),
+ dict(type='Resize', scale=(1333, 800), keep_ratio=True, backend=backend),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ poly2mask=False),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(batch_size=8, num_workers=4)
+val_dataloader = dict(dataset=dict(pipeline=test_pipeline))
+test_dataloader = val_dataloader
+# Enable automatic-mixed-precision training with AmpOptimWrapper.
+optim_wrapper = dict(
+ type='AmpOptimWrapper',
+ optimizer=dict(
+ type='SGD', lr=0.01 * 4, momentum=0.9, weight_decay=0.00004),
+ paramwise_cfg=dict(norm_decay_mult=0.))
+
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=0.00025,
+ by_epoch=False,
+ begin=0,
+ end=4000),
+ dict(
+ type='MultiStepLR',
+ begin=0,
+ end=25,
+ by_epoch=True,
+ milestones=[22, 24],
+ gamma=0.1)
+]
+
+# NOTE: `auto_scale_lr` is for automatically scaling LR,
+# USER SHOULD NOT CHANGE ITS VALUES.
+# base_batch_size = (8 GPUs) x (8 samples per GPU)
+auto_scale_lr = dict(base_batch_size=64)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/demo.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/demo.py
new file mode 100644
index 0000000000000000000000000000000000000000..d5c80c9aa5fcce38003d8f105c50c7316cdd2e51
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/demo.py
@@ -0,0 +1,142 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import os
+import urllib
+from argparse import ArgumentParser
+
+import mmcv
+import torch
+from mmengine.logging import print_log
+from mmengine.utils import ProgressBar, scandir
+
+from mmdet.apis import inference_detector, init_detector
+from mmdet.registry import VISUALIZERS
+from mmdet.utils import register_all_modules
+
+IMG_EXTENSIONS = ('.jpg', '.jpeg', '.png', '.ppm', '.bmp', '.pgm', '.tif',
+ '.tiff', '.webp')
+
+
+def get_file_list(source_root: str) -> [list, dict]:
+ """Get file list.
+
+ Args:
+ source_root (str): image or video source path
+
+ Return:
+ source_file_path_list (list): A list for all source file.
+ source_type (dict): Source type: file or url or dir.
+ """
+ is_dir = os.path.isdir(source_root)
+ is_url = source_root.startswith(('http:/', 'https:/'))
+ is_file = os.path.splitext(source_root)[-1].lower() in IMG_EXTENSIONS
+
+ source_file_path_list = []
+ if is_dir:
+ # when input source is dir
+ for file in scandir(source_root, IMG_EXTENSIONS, recursive=True):
+ source_file_path_list.append(os.path.join(source_root, file))
+ elif is_url:
+ # when input source is url
+ filename = os.path.basename(
+ urllib.parse.unquote(source_root).split('?')[0])
+ file_save_path = os.path.join(os.getcwd(), filename)
+ print(f'Downloading source file to {file_save_path}')
+ torch.hub.download_url_to_file(source_root, file_save_path)
+ source_file_path_list = [file_save_path]
+ elif is_file:
+ # when input source is single image
+ source_file_path_list = [source_root]
+ else:
+ print('Cannot find image file.')
+
+ source_type = dict(is_dir=is_dir, is_url=is_url, is_file=is_file)
+
+ return source_file_path_list, source_type
+
+
+def parse_args():
+ parser = ArgumentParser()
+ parser.add_argument(
+ 'img', help='Image path, include image file, dir and URL.')
+ parser.add_argument('config', help='Config file')
+ parser.add_argument('checkpoint', help='Checkpoint file')
+ parser.add_argument(
+ '--out-dir', default='./output', help='Path to output file')
+ parser.add_argument(
+ '--device', default='cuda:0', help='Device used for inference')
+ parser.add_argument(
+ '--show', action='store_true', help='Show the detection results')
+ parser.add_argument(
+ '--score-thr', type=float, default=0.3, help='Bbox score threshold')
+ parser.add_argument(
+ '--dataset', type=str, help='dataset name to load the text embedding')
+ parser.add_argument(
+ '--class-name', nargs='+', type=str, help='custom class names')
+ args = parser.parse_args()
+ return args
+
+
+def main():
+ args = parse_args()
+
+ # register all modules in mmdet into the registries
+ register_all_modules()
+
+ # build the model from a config file and a checkpoint file
+ model = init_detector(args.config, args.checkpoint, device=args.device)
+
+ if not os.path.exists(args.out_dir) and not args.show:
+ os.mkdir(args.out_dir)
+
+ # init visualizer
+ visualizer = VISUALIZERS.build(model.cfg.visualizer)
+ visualizer.dataset_meta = model.dataset_meta
+
+ # get file list
+ files, source_type = get_file_list(args.img)
+ from detic.utils import (get_class_names, get_text_embeddings,
+ reset_cls_layer_weight)
+
+ # class name embeddings
+ if args.class_name:
+ dataset_classes = args.class_name
+ elif args.dataset:
+ dataset_classes = get_class_names(args.dataset)
+ embedding = get_text_embeddings(
+ dataset=args.dataset, custom_vocabulary=args.class_name)
+ visualizer.dataset_meta['classes'] = dataset_classes
+ reset_cls_layer_weight(model, embedding)
+
+ # start detector inference
+ progress_bar = ProgressBar(len(files))
+ for file in files:
+ result = inference_detector(model, file)
+
+ img = mmcv.imread(file)
+ img = mmcv.imconvert(img, 'bgr', 'rgb')
+
+ if source_type['is_dir']:
+ filename = os.path.relpath(file, args.img).replace('/', '_')
+ else:
+ filename = os.path.basename(file)
+ out_file = None if args.show else os.path.join(args.out_dir, filename)
+
+ progress_bar.update()
+
+ visualizer.add_datasample(
+ filename,
+ img,
+ data_sample=result,
+ draw_gt=False,
+ show=args.show,
+ wait_time=0,
+ out_file=out_file,
+ pred_score_thr=args.score_thr)
+
+ if not args.show:
+ print_log(
+ f'\nResults have been saved at {os.path.abspath(args.out_dir)}')
+
+
+if __name__ == '__main__':
+ main()
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..d0ad070259aa1dfe3cd037b67d721bb0babd53b0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/__init__.py
@@ -0,0 +1,9 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .centernet_rpn_head import CenterNetRPNHead
+from .detic_bbox_head import DeticBBoxHead
+from .detic_roi_head import DeticRoIHead
+from .zero_shot_classifier import ZeroShotClassifier
+
+__all__ = [
+ 'CenterNetRPNHead', 'DeticBBoxHead', 'DeticRoIHead', 'ZeroShotClassifier'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/centernet_rpn_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/centernet_rpn_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..765d6dfb2b6425bf66dff71e3ac1cd5cd6e28707
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/centernet_rpn_head.py
@@ -0,0 +1,196 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from typing import List, Sequence, Tuple
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import Scale
+from mmengine import ConfigDict
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.models.dense_heads import CenterNetUpdateHead
+from mmdet.models.utils import multi_apply
+from mmdet.registry import MODELS
+
+INF = 1000000000
+RangeType = Sequence[Tuple[int, int]]
+
+
+@MODELS.register_module(force=True) # avoid bug
+class CenterNetRPNHead(CenterNetUpdateHead):
+ """CenterNetUpdateHead is an improved version of CenterNet in CenterNet2.
+
+ Paper link ``_.
+ """
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self._init_reg_convs()
+ self._init_predictor()
+
+ def _init_predictor(self) -> None:
+ """Initialize predictor layers of the head."""
+ self.conv_cls = nn.Conv2d(
+ self.feat_channels, self.num_classes, 3, padding=1)
+ self.conv_reg = nn.Conv2d(self.feat_channels, 4, 3, padding=1)
+
+ def forward(self, x: Tuple[Tensor]) -> Tuple[List[Tensor], List[Tensor]]:
+ """Forward features from the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Features from the upstream network, each is
+ a 4D-tensor.
+
+ Returns:
+ tuple: A tuple of each level outputs.
+
+ - cls_scores (list[Tensor]): Box scores for each scale level, \
+ each is a 4D-tensor, the channel number is num_classes.
+ - bbox_preds (list[Tensor]): Box energies / deltas for each \
+ scale level, each is a 4D-tensor, the channel number is 4.
+ """
+ res = multi_apply(self.forward_single, x, self.scales, self.strides)
+ return res
+
+ def forward_single(self, x: Tensor, scale: Scale,
+ stride: int) -> Tuple[Tensor, Tensor]:
+ """Forward features of a single scale level.
+
+ Args:
+ x (Tensor): FPN feature maps of the specified stride.
+ scale (:obj:`mmcv.cnn.Scale`): Learnable scale module to resize
+ the bbox prediction.
+ stride (int): The corresponding stride for feature maps.
+
+ Returns:
+ tuple: scores for each class, bbox predictions of
+ input feature maps.
+ """
+ for m in self.reg_convs:
+ x = m(x)
+ cls_score = self.conv_cls(x)
+ bbox_pred = self.conv_reg(x)
+ # scale the bbox_pred of different level
+ # float to avoid overflow when enabling FP16
+ bbox_pred = scale(bbox_pred).float()
+ # bbox_pred needed for gradient computation has been modified
+ # by F.relu(bbox_pred) when run with PyTorch 1.10. So replace
+ # F.relu(bbox_pred) with bbox_pred.clamp(min=0)
+ bbox_pred = bbox_pred.clamp(min=0)
+ if not self.training:
+ bbox_pred *= stride
+ return cls_score, bbox_pred # score aligned, box larger
+
+ def _predict_by_feat_single(self,
+ cls_score_list: List[Tensor],
+ bbox_pred_list: List[Tensor],
+ score_factor_list: List[Tensor],
+ mlvl_priors: List[Tensor],
+ img_meta: dict,
+ cfg: ConfigDict,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results.
+
+ Args:
+ cls_score_list (list[Tensor]): Box scores from all scale
+ levels of a single image, each item has shape
+ (num_priors * num_classes, H, W).
+ bbox_pred_list (list[Tensor]): Box energies / deltas from
+ all scale levels of a single image, each item has shape
+ (num_priors * 4, H, W).
+ score_factor_list (list[Tensor]): Score factor from all scale
+ levels of a single image, each item has shape
+ (num_priors * 1, H, W).
+ mlvl_priors (list[Tensor]): Each element in the list is
+ the priors of a single level in feature pyramid. In all
+ anchor-based methods, it has shape (num_priors, 4). In
+ all anchor-free methods, it has shape (num_priors, 2)
+ when `with_stride=True`, otherwise it still has shape
+ (num_priors, 4).
+ img_meta (dict): Image meta info.
+ cfg (mmengine.Config): Test / postprocessing configuration,
+ if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+
+ cfg = self.test_cfg if cfg is None else cfg
+ cfg = copy.deepcopy(cfg)
+ nms_pre = cfg.get('nms_pre', -1)
+
+ mlvl_bbox_preds = []
+ mlvl_valid_priors = []
+ mlvl_scores = []
+ mlvl_labels = []
+
+ for level_idx, (cls_score, bbox_pred, score_factor, priors) in \
+ enumerate(zip(cls_score_list, bbox_pred_list,
+ score_factor_list, mlvl_priors)):
+
+ assert cls_score.size()[-2:] == bbox_pred.size()[-2:]
+
+ dim = self.bbox_coder.encode_size
+ bbox_pred = bbox_pred.permute(1, 2, 0).reshape(-1, dim)
+ cls_score = cls_score.permute(1, 2,
+ 0).reshape(-1, self.cls_out_channels)
+ heatmap = cls_score.sigmoid()
+ score_thr = cfg.get('score_thr', 0)
+
+ candidate_inds = heatmap > score_thr # 0.05
+ pre_nms_top_n = candidate_inds.sum() # N
+ pre_nms_top_n = pre_nms_top_n.clamp(max=nms_pre) # N
+
+ heatmap = heatmap[candidate_inds] # n
+
+ candidate_nonzeros = candidate_inds.nonzero() # n
+ box_loc = candidate_nonzeros[:, 0] # n
+ labels = candidate_nonzeros[:, 1] # n
+
+ bbox_pred = bbox_pred[box_loc] # n x 4
+ per_grids = priors[box_loc] # n x 2
+
+ if candidate_inds.sum().item() > pre_nms_top_n.item():
+ heatmap, top_k_indices = \
+ heatmap.topk(pre_nms_top_n, sorted=False)
+ labels = labels[top_k_indices]
+ bbox_pred = bbox_pred[top_k_indices]
+ per_grids = per_grids[top_k_indices]
+
+ bboxes = self.bbox_coder.decode(per_grids, bbox_pred)
+ # avoid invalid boxes in RoI heads
+ bboxes[:, 2] = torch.max(bboxes[:, 2], bboxes[:, 0] + 0.01)
+ bboxes[:, 3] = torch.max(bboxes[:, 3], bboxes[:, 1] + 0.01)
+
+ mlvl_bbox_preds.append(bboxes)
+ mlvl_valid_priors.append(priors)
+ mlvl_scores.append(torch.sqrt(heatmap))
+ mlvl_labels.append(labels)
+
+ results = InstanceData()
+ results.bboxes = torch.cat(mlvl_bbox_preds)
+ results.scores = torch.cat(mlvl_scores)
+ results.labels = torch.cat(mlvl_labels)
+
+ return self._bbox_post_process(
+ results=results,
+ cfg=cfg,
+ rescale=rescale,
+ with_nms=with_nms,
+ img_meta=img_meta)
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/detic_bbox_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/detic_bbox_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..9408cbe04fd94e9d2490b4c9589b22d043f1e5b9
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/detic_bbox_head.py
@@ -0,0 +1,112 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Union
+
+from mmengine.config import ConfigDict
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.models.layers import multiclass_nms
+from mmdet.models.roi_heads.bbox_heads import Shared2FCBBoxHead
+from mmdet.models.utils import empty_instances
+from mmdet.registry import MODELS
+from mmdet.structures.bbox import get_box_tensor, scale_boxes
+
+
+@MODELS.register_module(force=True) # avoid bug
+class DeticBBoxHead(Shared2FCBBoxHead):
+
+ def __init__(self,
+ *args,
+ init_cfg: Optional[Union[dict, ConfigDict]] = None,
+ **kwargs) -> None:
+ super().__init__(*args, init_cfg=init_cfg, **kwargs)
+ # reconstruct fc_cls and fc_reg since input channels are changed
+ assert self.with_cls
+ cls_channels = self.num_classes
+ cls_predictor_cfg_ = self.cls_predictor_cfg.copy()
+ cls_predictor_cfg_.update(
+ in_features=self.cls_last_dim, out_features=cls_channels)
+ self.fc_cls = MODELS.build(cls_predictor_cfg_)
+
+ def _predict_by_feat_single(
+ self,
+ roi: Tensor,
+ cls_score: Tensor,
+ bbox_pred: Tensor,
+ img_meta: dict,
+ rescale: bool = False,
+ rcnn_test_cfg: Optional[ConfigDict] = None) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results.
+
+ Args:
+ roi (Tensor): Boxes to be transformed. Has shape (num_boxes, 5).
+ last dimension 5 arrange as (batch_index, x1, y1, x2, y2).
+ cls_score (Tensor): Box scores, has shape
+ (num_boxes, num_classes + 1).
+ bbox_pred (Tensor): Box energies / deltas.
+ has shape (num_boxes, num_classes * 4).
+ img_meta (dict): image information.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ rcnn_test_cfg (obj:`ConfigDict`): `test_cfg` of Bbox Head.
+ Defaults to None
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image\
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ results = InstanceData()
+ if roi.shape[0] == 0:
+ return empty_instances([img_meta],
+ roi.device,
+ task_type='bbox',
+ instance_results=[results],
+ box_type=self.predict_box_type,
+ use_box_type=False,
+ num_classes=self.num_classes,
+ score_per_cls=rcnn_test_cfg is None)[0]
+ scores = cls_score
+ img_shape = img_meta['img_shape']
+ num_rois = roi.size(0)
+
+ num_classes = 1 if self.reg_class_agnostic else self.num_classes
+ roi = roi.repeat_interleave(num_classes, dim=0)
+ bbox_pred = bbox_pred.view(-1, self.bbox_coder.encode_size)
+ bboxes = self.bbox_coder.decode(
+ roi[..., 1:], bbox_pred, max_shape=img_shape)
+
+ if rescale and bboxes.size(0) > 0:
+ assert img_meta.get('scale_factor') is not None
+ scale_factor = [1 / s for s in img_meta['scale_factor']]
+ bboxes = scale_boxes(bboxes, scale_factor)
+
+ # Get the inside tensor when `bboxes` is a box type
+ bboxes = get_box_tensor(bboxes)
+ box_dim = bboxes.size(-1)
+ bboxes = bboxes.view(num_rois, -1)
+
+ if rcnn_test_cfg is None:
+ # This means that it is aug test.
+ # It needs to return the raw results without nms.
+ results.bboxes = bboxes
+ results.scores = scores
+ else:
+ det_bboxes, det_labels = multiclass_nms(
+ bboxes,
+ scores,
+ rcnn_test_cfg.score_thr,
+ rcnn_test_cfg.nms,
+ rcnn_test_cfg.max_per_img,
+ box_dim=box_dim)
+ results.bboxes = det_bboxes[:, :-1]
+ results.scores = det_bboxes[:, -1]
+ results.labels = det_labels
+ return results
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/detic_roi_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/detic_roi_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..a09c11c6e698b94e89a81e9562fbcdc666a62d95
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/detic_roi_head.py
@@ -0,0 +1,326 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Sequence, Tuple
+
+import torch
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.models.roi_heads import CascadeRoIHead
+from mmdet.models.task_modules.samplers import SamplingResult
+from mmdet.models.test_time_augs import merge_aug_masks
+from mmdet.models.utils.misc import empty_instances
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import bbox2roi, get_box_tensor
+from mmdet.utils import ConfigType, InstanceList, MultiConfig
+
+
+@MODELS.register_module(force=True) # avoid bug
+class DeticRoIHead(CascadeRoIHead):
+
+ def init_mask_head(self, mask_roi_extractor: MultiConfig,
+ mask_head: MultiConfig) -> None:
+ """Initialize mask head and mask roi extractor.
+
+ Args:
+ mask_head (dict): Config of mask in mask head.
+ mask_roi_extractor (:obj:`ConfigDict`, dict or list):
+ Config of mask roi extractor.
+ """
+ self.mask_head = MODELS.build(mask_head)
+
+ if mask_roi_extractor is not None:
+ self.share_roi_extractor = False
+ self.mask_roi_extractor = MODELS.build(mask_roi_extractor)
+ else:
+ self.share_roi_extractor = True
+ self.mask_roi_extractor = self.bbox_roi_extractor
+
+ def _refine_roi(self, x: Tuple[Tensor], rois: Tensor,
+ batch_img_metas: List[dict],
+ num_proposals_per_img: Sequence[int], **kwargs) -> tuple:
+ """Multi-stage refinement of RoI.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ rois (Tensor): shape (n, 5), [batch_ind, x1, y1, x2, y2]
+ batch_img_metas (list[dict]): List of image information.
+ num_proposals_per_img (sequence[int]): number of proposals
+ in each image.
+
+ Returns:
+ tuple:
+
+ - rois (Tensor): Refined RoI.
+ - cls_scores (list[Tensor]): Average predicted
+ cls score per image.
+ - bbox_preds (list[Tensor]): Bbox branch predictions
+ for the last stage of per image.
+ """
+ # "ms" in variable names means multi-stage
+ ms_scores = []
+ for stage in range(self.num_stages):
+ bbox_results = self._bbox_forward(
+ stage=stage, x=x, rois=rois, **kwargs)
+
+ # split batch bbox prediction back to each image
+ cls_scores = bbox_results['cls_score'].sigmoid()
+ bbox_preds = bbox_results['bbox_pred']
+
+ rois = rois.split(num_proposals_per_img, 0)
+ cls_scores = cls_scores.split(num_proposals_per_img, 0)
+ ms_scores.append(cls_scores)
+ bbox_preds = bbox_preds.split(num_proposals_per_img, 0)
+
+ if stage < self.num_stages - 1:
+ bbox_head = self.bbox_head[stage]
+ refine_rois_list = []
+ for i in range(len(batch_img_metas)):
+ if rois[i].shape[0] > 0:
+ bbox_label = cls_scores[i][:, :-1].argmax(dim=1)
+ # Refactor `bbox_head.regress_by_class` to only accept
+ # box tensor without img_idx concatenated.
+ refined_bboxes = bbox_head.regress_by_class(
+ rois[i][:, 1:], bbox_label, bbox_preds[i],
+ batch_img_metas[i])
+ refined_bboxes = get_box_tensor(refined_bboxes)
+ refined_rois = torch.cat(
+ [rois[i][:, [0]], refined_bboxes], dim=1)
+ refine_rois_list.append(refined_rois)
+ rois = torch.cat(refine_rois_list)
+ # ms_scores aligned
+ # average scores of each image by stages
+ cls_scores = [
+ sum([score[i] for score in ms_scores]) / float(len(ms_scores))
+ for i in range(len(batch_img_metas))
+ ] # aligned
+ return rois, cls_scores, bbox_preds
+
+ def _bbox_forward(self, stage: int, x: Tuple[Tensor],
+ rois: Tensor) -> dict:
+ """Box head forward function used in both training and testing.
+
+ Args:
+ stage (int): The current stage in Cascade RoI Head.
+ x (tuple[Tensor]): List of multi-level img features.
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+
+ Returns:
+ dict[str, Tensor]: Usually returns a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `bbox_feats` (Tensor): Extract bbox RoI features.
+ """
+ bbox_roi_extractor = self.bbox_roi_extractor[stage]
+ bbox_head = self.bbox_head[stage]
+ bbox_feats = bbox_roi_extractor(x[:bbox_roi_extractor.num_inputs],
+ rois)
+ # do not support caffe_c4 model anymore
+ cls_score, bbox_pred = bbox_head(bbox_feats)
+
+ bbox_results = dict(
+ cls_score=cls_score, bbox_pred=bbox_pred, bbox_feats=bbox_feats)
+ return bbox_results
+
+ def predict_bbox(self,
+ x: Tuple[Tensor],
+ batch_img_metas: List[dict],
+ rpn_results_list: InstanceList,
+ rcnn_test_cfg: ConfigType,
+ rescale: bool = False,
+ **kwargs) -> InstanceList:
+ """Perform forward propagation of the bbox head and predict detection
+ results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Feature maps of all scale level.
+ batch_img_metas (list[dict]): List of image information.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ rcnn_test_cfg (obj:`ConfigDict`): `test_cfg` of R-CNN.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ proposals = [res.bboxes for res in rpn_results_list]
+ proposal_scores = [res.scores for res in rpn_results_list]
+ num_proposals_per_img = tuple(len(p) for p in proposals)
+ rois = bbox2roi(proposals)
+
+ if rois.shape[0] == 0:
+ return empty_instances(
+ batch_img_metas,
+ rois.device,
+ task_type='bbox',
+ box_type=self.bbox_head[-1].predict_box_type,
+ num_classes=self.bbox_head[-1].num_classes,
+ score_per_cls=rcnn_test_cfg is None)
+ # rois aligned
+ rois, cls_scores, bbox_preds = self._refine_roi(
+ x=x,
+ rois=rois,
+ batch_img_metas=batch_img_metas,
+ num_proposals_per_img=num_proposals_per_img,
+ **kwargs)
+
+ # score reweighting in centernet2
+ cls_scores = [(s * ps[:, None])**0.5
+ for s, ps in zip(cls_scores, proposal_scores)]
+ cls_scores = [
+ s * (s == s[:, :-1].max(dim=1)[0][:, None]).float()
+ for s in cls_scores
+ ]
+
+ # fast_rcnn_inference
+ results_list = self.bbox_head[-1].predict_by_feat(
+ rois=rois,
+ cls_scores=cls_scores,
+ bbox_preds=bbox_preds,
+ batch_img_metas=batch_img_metas,
+ rescale=rescale,
+ rcnn_test_cfg=rcnn_test_cfg)
+ return results_list
+
+ def _mask_forward(self, x: Tuple[Tensor], rois: Tensor) -> dict:
+ """Mask head forward function used in both training and testing.
+
+ Args:
+ stage (int): The current stage in Cascade RoI Head.
+ x (tuple[Tensor]): Tuple of multi-level img features.
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `mask_preds` (Tensor): Mask prediction.
+ """
+ mask_feats = self.mask_roi_extractor(
+ x[:self.mask_roi_extractor.num_inputs], rois)
+ # do not support caffe_c4 model anymore
+ mask_preds = self.mask_head(mask_feats)
+
+ mask_results = dict(mask_preds=mask_preds)
+ return mask_results
+
+ def mask_loss(self, x, sampling_results: List[SamplingResult],
+ batch_gt_instances: InstanceList) -> dict:
+ """Run forward function and calculate loss for mask head in training.
+
+ Args:
+ x (tuple[Tensor]): Tuple of multi-level img features.
+ sampling_results (list["obj:`SamplingResult`]): Sampling results.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``labels``, and
+ ``masks`` attributes.
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `mask_preds` (Tensor): Mask prediction.
+ - `loss_mask` (dict): A dictionary of mask loss components.
+ """
+ pos_rois = bbox2roi([res.pos_priors for res in sampling_results])
+ mask_results = self._mask_forward(x, pos_rois)
+
+ mask_loss_and_target = self.mask_head.loss_and_target(
+ mask_preds=mask_results['mask_preds'],
+ sampling_results=sampling_results,
+ batch_gt_instances=batch_gt_instances,
+ rcnn_train_cfg=self.train_cfg[-1])
+ mask_results.update(mask_loss_and_target)
+
+ return mask_results
+
+ def loss(self, x: Tuple[Tensor], rpn_results_list: InstanceList,
+ batch_data_samples: SampleList) -> dict:
+ """Perform forward propagation and loss calculation of the detection
+ roi on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components
+ """
+ raise NotImplementedError
+
+ def predict_mask(self,
+ x: Tuple[Tensor],
+ batch_img_metas: List[dict],
+ results_list: List[InstanceData],
+ rescale: bool = False) -> List[InstanceData]:
+ """Perform forward propagation of the mask head and predict detection
+ results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Feature maps of all scale level.
+ batch_img_metas (list[dict]): List of image information.
+ results_list (list[:obj:`InstanceData`]): Detection results of
+ each image.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+ """
+ bboxes = [res.bboxes for res in results_list]
+ mask_rois = bbox2roi(bboxes)
+ if mask_rois.shape[0] == 0:
+ results_list = empty_instances(
+ batch_img_metas,
+ mask_rois.device,
+ task_type='mask',
+ instance_results=results_list,
+ mask_thr_binary=self.test_cfg.mask_thr_binary)
+ return results_list
+
+ num_mask_rois_per_img = [len(res) for res in results_list]
+ aug_masks = []
+ mask_results = self._mask_forward(x, mask_rois)
+ mask_preds = mask_results['mask_preds']
+ # split batch mask prediction back to each image
+ mask_preds = mask_preds.split(num_mask_rois_per_img, 0)
+ aug_masks.append([m.sigmoid().detach() for m in mask_preds])
+
+ merged_masks = []
+ for i in range(len(batch_img_metas)):
+ aug_mask = [mask[i] for mask in aug_masks]
+ merged_mask = merge_aug_masks(aug_mask, batch_img_metas[i])
+ merged_masks.append(merged_mask)
+ results_list = self.mask_head.predict_by_feat(
+ mask_preds=merged_masks,
+ results_list=results_list,
+ batch_img_metas=batch_img_metas,
+ rcnn_test_cfg=self.test_cfg,
+ rescale=rescale,
+ activate_map=True)
+ return results_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/text_encoder.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/text_encoder.py
new file mode 100644
index 0000000000000000000000000000000000000000..f0024efaf3025bbdb3ae1cfd2e805dbadae61674
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/text_encoder.py
@@ -0,0 +1,50 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Union
+
+import torch
+import torch.nn as nn
+
+
+class CLIPTextEncoder(nn.Module):
+
+ def __init__(self, model_name='ViT-B/32'):
+ super().__init__()
+ import clip
+ from clip.simple_tokenizer import SimpleTokenizer
+ self.tokenizer = SimpleTokenizer()
+ pretrained_model, _ = clip.load(model_name, device='cpu')
+ self.clip = pretrained_model
+
+ @property
+ def device(self):
+ return self.clip.device
+
+ @property
+ def dtype(self):
+ return self.clip.dtype
+
+ def tokenize(self,
+ texts: Union[str, List[str]],
+ context_length: int = 77) -> torch.LongTensor:
+ if isinstance(texts, str):
+ texts = [texts]
+
+ sot_token = self.tokenizer.encoder['<|startoftext|>']
+ eot_token = self.tokenizer.encoder['<|endoftext|>']
+ all_tokens = [[sot_token] + self.tokenizer.encode(text) + [eot_token]
+ for text in texts]
+ result = torch.zeros(len(all_tokens), context_length, dtype=torch.long)
+
+ for i, tokens in enumerate(all_tokens):
+ if len(tokens) > context_length:
+ st = torch.randint(len(tokens) - context_length + 1,
+ (1, ))[0].item()
+ tokens = tokens[st:st + context_length]
+ result[i, :len(tokens)] = torch.tensor(tokens)
+
+ return result
+
+ def forward(self, text):
+ text = self.tokenize(text)
+ text_features = self.clip.encode_text(text)
+ return text_features
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/utils.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/utils.py
new file mode 100644
index 0000000000000000000000000000000000000000..56d4fd429d72910bb486630fb2e38edcd986cab0
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/utils.py
@@ -0,0 +1,78 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import numpy as np
+import torch
+import torch.nn.functional as F
+from mmengine.logging import print_log
+
+from .text_encoder import CLIPTextEncoder
+
+# download from
+# https://github.com/facebookresearch/Detic/tree/main/datasets/metadata
+DATASET_EMBEDDINGS = {
+ 'lvis': 'datasets/metadata/lvis_v1_clip_a+cname.npy',
+ 'objects365': 'datasets/metadata/o365_clip_a+cnamefix.npy',
+ 'openimages': 'datasets/metadata/oid_clip_a+cname.npy',
+ 'coco': 'datasets/metadata/coco_clip_a+cname.npy',
+}
+
+
+def get_text_embeddings(dataset=None,
+ custom_vocabulary=None,
+ prompt_prefix='a '):
+ assert (dataset is None) ^ (custom_vocabulary is None), \
+ 'Either `dataset` or `custom_vocabulary` should be specified.'
+ if dataset:
+ if dataset in DATASET_EMBEDDINGS:
+ return DATASET_EMBEDDINGS[dataset]
+ else:
+ custom_vocabulary = get_class_names(dataset)
+
+ text_encoder = CLIPTextEncoder()
+ text_encoder.eval()
+ texts = [prompt_prefix + x for x in custom_vocabulary]
+ print_log(
+ f'Computing text embeddings for {len(custom_vocabulary)} classes.')
+ embeddings = text_encoder(texts).detach().permute(1, 0).contiguous().cpu()
+ return embeddings
+
+
+def get_class_names(dataset):
+ if dataset == 'coco':
+ from mmdet.datasets import CocoDataset
+ class_names = CocoDataset.METAINFO['classes']
+ elif dataset == 'cityscapes':
+ from mmdet.datasets import CityscapesDataset
+ class_names = CityscapesDataset.METAINFO['classes']
+ elif dataset == 'voc':
+ from mmdet.datasets import VOCDataset
+ class_names = VOCDataset.METAINFO['classes']
+ elif dataset == 'openimages':
+ from mmdet.datasets import OpenImagesDataset
+ class_names = OpenImagesDataset.METAINFO['classes']
+ elif dataset == 'lvis':
+ from mmdet.datasets import LVISV1Dataset
+ class_names = LVISV1Dataset.METAINFO['classes']
+ else:
+ raise TypeError(f'Invalid type for dataset name: {type(dataset)}')
+ return class_names
+
+
+def reset_cls_layer_weight(model, weight):
+ if type(weight) == str:
+ print_log(f'Resetting cls_layer_weight from file: {weight}')
+ zs_weight = torch.tensor(
+ np.load(weight),
+ dtype=torch.float32).permute(1, 0).contiguous() # D x C
+ else:
+ zs_weight = weight
+ zs_weight = torch.cat(
+ [zs_weight, zs_weight.new_zeros(
+ (zs_weight.shape[0], 1))], dim=1) # D x (C + 1)
+ zs_weight = F.normalize(zs_weight, p=2, dim=0)
+ zs_weight = zs_weight.to('cuda')
+ num_classes = zs_weight.shape[-1]
+
+ for bbox_head in model.roi_head.bbox_head:
+ bbox_head.num_classes = num_classes
+ del bbox_head.fc_cls.zs_weight
+ bbox_head.fc_cls.zs_weight = zs_weight
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/zero_shot_classifier.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/zero_shot_classifier.py
new file mode 100644
index 0000000000000000000000000000000000000000..35c9e49285ca1fe82f027be8ebe7f197cf772988
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic/detic/zero_shot_classifier.py
@@ -0,0 +1,73 @@
+# Copyright (c) Facebook, Inc. and its affiliates.
+import numpy as np
+import torch
+from torch import nn
+from torch.nn import functional as F
+
+from mmdet.registry import MODELS
+
+
+@MODELS.register_module(force=True) # avoid bug
+class ZeroShotClassifier(nn.Module):
+
+ def __init__(
+ self,
+ in_features: int,
+ out_features: int, # num_classes
+ zs_weight_path: str,
+ zs_weight_dim: int = 512,
+ use_bias: float = 0.0,
+ norm_weight: bool = True,
+ norm_temperature: float = 50.0,
+ ):
+ super().__init__()
+ num_classes = out_features
+ self.norm_weight = norm_weight
+ self.norm_temperature = norm_temperature
+
+ self.use_bias = use_bias < 0
+ if self.use_bias:
+ self.cls_bias = nn.Parameter(torch.ones(1) * use_bias)
+
+ self.linear = nn.Linear(in_features, zs_weight_dim)
+
+ if zs_weight_path == 'rand':
+ zs_weight = torch.randn((zs_weight_dim, num_classes))
+ nn.init.normal_(zs_weight, std=0.01)
+ else:
+ zs_weight = torch.tensor(
+ np.load(zs_weight_path),
+ dtype=torch.float32).permute(1, 0).contiguous() # D x C
+ zs_weight = torch.cat(
+ [zs_weight, zs_weight.new_zeros(
+ (zs_weight_dim, 1))], dim=1) # D x (C + 1)
+
+ if self.norm_weight:
+ zs_weight = F.normalize(zs_weight, p=2, dim=0)
+
+ if zs_weight_path == 'rand':
+ self.zs_weight = nn.Parameter(zs_weight)
+ else:
+ self.register_buffer('zs_weight', zs_weight)
+
+ assert self.zs_weight.shape[1] == num_classes + 1, self.zs_weight.shape
+
+ def forward(self, x, classifier=None):
+ '''
+ Inputs:
+ x: B x D'
+ classifier_info: (C', C' x D)
+ '''
+ x = self.linear(x)
+ if classifier is not None:
+ zs_weight = classifier.permute(1, 0).contiguous() # D x C'
+ zs_weight = F.normalize(zs_weight, p=2, dim=0) \
+ if self.norm_weight else zs_weight
+ else:
+ zs_weight = self.zs_weight
+ if self.norm_weight:
+ x = self.norm_temperature * F.normalize(x, p=2, dim=1)
+ x = torch.mm(x, zs_weight)
+ if self.use_bias:
+ x = x + self.cls_bias
+ return x
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..3c7714c36a957bcd9786772791894356ab1cae71
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/README.md
@@ -0,0 +1,248 @@
+# Detecting Twenty-thousand Classes using Image-level Supervision
+
+## Description
+
+**Detic**: A **Det**ector with **i**mage **c**lasses that can use image-level labels to easily train detectors.
+
+
+
+> [**Detecting Twenty-thousand Classes using Image-level Supervision**](http://arxiv.org/abs/2201.02605),
+> Xingyi Zhou, Rohit Girdhar, Armand Joulin, Philipp Krähenbühl, Ishan Misra,
+> *ECCV 2022 ([arXiv 2201.02605](http://arxiv.org/abs/2201.02605))*
+
+## Usage
+
+
+
+## Installation
+
+Detic requires to install CLIP.
+
+```shell
+pip install git+https://github.com/openai/CLIP.git
+```
+
+## Prepare Datasets
+
+It is recommended to download and extract the dataset somewhere outside the project directory and symlink the dataset root to `$MMDETECTION/data` as below. If your folder structure is different, you may need to change the corresponding paths in config files.
+
+### LVIS
+
+LVIS dataset is adopted as box-labeled data, [LVIS](https://www.lvisdataset.org/) is available from official website or mirror. You need to generate `lvis_v1_train_norare.json` according to the [official prepare datasets](https://github.com/facebookresearch/Detic/blob/main/datasets/README.md#coco-and-lvis) for open-vocabulary LVIS, which removes the labels of 337 rare-class from training. You can also download [lvis_v1_train_norare.json](https://download.openmmlab.com/mmdetection/v3.0/detic/data/lvis/annotations/lvis_v1_train_norare.json) from our backup. The directory should be like this.
+
+```shell
+mmdetection
+├── data
+│ ├── lvis
+│ │ ├── annotations
+│ │ | ├── lvis_v1_train.json
+│ │ | ├── lvis_v1_val.json
+│ │ | ├── lvis_v1_train_norare.json
+│ │ ├── train2017
+│ │ ├── val2017
+```
+
+### ImageNet-LVIS
+
+ImageNet-LVIS is adopted as image-labeled data. You can download [ImageNet-21K](https://www.image-net.org/download.php) dataset from the official website. Then you need to unzip the overlapping classes of LVIS and convert them into LVIS annotation format according to the [official prepare datasets](https://github.com/facebookresearch/Detic/blob/main/datasets/README.md#imagenet-21k). The directory should be like this.
+
+```shell
+mmdetection
+├── data
+│ ├── imagenet
+│ │ ├── annotations
+│ │ | ├── imagenet_lvis_image_info.json
+│ │ ├── ImageNet-21K
+│ │ | ├── n00007846
+│ │ | ├── n01318894
+│ │ | ├── ...
+```
+
+### Metadata
+
+`data/metadata/` is the preprocessed meta-data (included in the repo). Please follow the [official instruction](https://github.com/facebookresearch/Detic/blob/main/datasets/README.md#metadata) to pre-process the LVIS dataset. You will generate `lvis_v1_train_cat_info.json` for Federated loss, which contains the frequency of each category of training set of LVIS. In addition, `lvis_v1_clip_a+cname.npy` is the pre-computed CLIP embeddings for each category of LVIS. You can also choose to directly download [lvis_v1_train_cat_info](https://download.openmmlab.com/mmdetection/v3.0/detic/data/metadata/lvis_v1_train_cat_info.json) and [lvis_v1_clip_a+cname.npy](https://download.openmmlab.com/mmdetection/v3.0/detic/data/metadata/lvis_v1_clip_a%2Bcname.npy) form our backup. The directory should be like this.
+
+```shell
+mmdetection
+├── data
+│ ├── metadata
+│ │ ├── lvis_v1_train_cat_info.json
+│ │ ├── lvis_v1_clip_a+cname.npy
+```
+
+## Demo
+
+Here we provide the Detic model for the open vocabulary demo. This model is trained on combined LVIS-COCO and ImageNet-21K for better demo purposes. LVIS models do not detect persons well due to its federated annotation protocol. LVIS+COCO models give better visual results.
+
+| Backbone | Training data | Config | Download |
+| :------: | :----------------------------: | :-------------------------------------------------------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| Swin-B | LVIS & COCO & ImageNet-21K | [config](./configs/detic_centernet2_swin-b_fpn_4x_lvis_coco_in21k.py) | [model](https://download.openmmlab.com/mmdetection/v3.0/detic/detic_centernet2_swin-b_fpn_4x_lvis-coco-in21k/detic_centernet2_swin-b_fpn_4x_lvis-coco-in21k_20230120-0d301978.pth) |
+
+You can also download other models from [official model zoo](https://github.com/facebookresearch/Detic/blob/main/docs/MODEL_ZOO.md), and convert the format by run
+
+```shell
+python tools/model_converters/detic_to_mmdet.py --src /path/to/detic_weight.pth --dst /path/to/mmdet_weight.pth
+```
+
+### Inference with existing dataset vocabulary
+
+You can detect classes of existing dataset with `--texts` command:
+
+```shell
+python demo/image_demo.py \
+ ${IMAGE_PATH} \
+ ${CONFIG_PATH} \
+ ${MODEL_PATH} \
+ --texts lvis \
+ --pred-score-thr 0.5 \
+ --palette 'random'
+```
+
+
+
+### Inference with custom vocabularies
+
+Detic can detects any class given class names by using CLIP. You can detect customized classes with `--texts` command:
+
+```shell
+python demo/image_demo.py \
+ ${IMAGE_PATH} \
+ ${CONFIG_PATH} \
+ ${MODEL_PATH} \
+ --texts 'headphone . webcam . paper . coffe.' \
+ --pred-score-thr 0.3 \
+ --palette 'random'
+```
+
+
+
+Note that `headphone`, `paper` and `coffe` (typo intended) are not LVIS classes. Despite the misspelled class name, Detic can produce a reasonable detection for `coffe`.
+
+## Models and Results
+
+### Training
+
+There are two stages in the whole training process. The first stage is to train a model using images with box labels as the baseline. The second stage is to finetune from the baseline model and leverage image-labeled data.
+
+#### First stage
+
+To train the baseline with box-supervised, run
+
+```shell
+bash ./tools/dist_train.sh projects/Detic_new/detic_centernet2_r50_fpn_4x_lvis_boxsup.py 8
+```
+
+| Model (Config) | mask mAP | mask mAP(official) | mask mAP_rare | mask mAP_rare(officical) |
+| :---------------------------------------------------------------------------------------------: | :------: | :----------------: | :-----------: | :----------------------: |
+| [detic_centernet2_r50_fpn_4x_lvis_boxsup](./configs/detic_centernet2_r50_fpn_4x_lvis_boxsup.py) | 31.6 | 31.5 | 26.6 | 25.6 |
+
+#### Second stage
+
+The second stage uses both object detection and image classification datasets.
+
+##### Multi-Datasets Config
+
+We provide improved dataset_wrapper `ConcatDataset` to concatenate multiple datasets, all datasets could have different annotation types and different pipelines (e.g., image_size). You can also obtain the index of `dataset_source` for each sample through ` get_dataset_source` . We provide sampler `MultiDataSampler` to custom the ratios of different datasets. Beside, we provide batch_sampler `MultiDataAspectRatioBatchSampler` to enable different datasets to have different batchsizes. The config of multiple datasets is as follows:
+
+```python
+dataset_det = dict(
+ type='ClassBalancedDataset',
+ oversample_thr=1e-3,
+ dataset=dict(
+ type='LVISV1Dataset',
+ data_root='data/lvis/',
+ ann_file='annotations/lvis_v1_train.json',
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline_det,
+ backend_args=backend_args))
+
+dataset_cls = dict(
+ type='ImageNetLVISV1Dataset',
+ data_root='data/imagenet',
+ ann_file='annotations/imagenet_lvis_image_info.json',
+ data_prefix=dict(img='ImageNet-LVIS/'),
+ pipeline=train_pipeline_cls,
+ backend_args=backend_args)
+
+train_dataloader = dict(
+ batch_size=[8, 32],
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(
+ type='MultiDataSampler',
+ dataset_ratio=[1, 4]),
+ batch_sampler=dict(
+ type='MultiDataAspectRatioBatchSampler',
+ num_datasets=2),
+ dataset=dict(
+ type='ConcatDataset',
+ datasets=[dataset_det, dataset_cls]))
+```
+
+###### Note:
+
+- If the one of the multiple datasets is `ConcatDataset` , it is still considered as a dataset for `num_datasets` in `MultiDataAspectRatioBatchSampler`.
+
+To finetune the baseline model with image-labeled data, run:
+
+```shell
+bash ./tools/dist_train.sh projects/Detic_new/detic_centernet2_r50_fpn_4x_lvis_in21k-lvis.py 8
+```
+
+| Model (Config) | mask mAP | mask mAP(official) | mask mAP_rare | mask mAP_rare(officical) |
+| :-----------------------------------------------------------------------------------------------------: | :------: | :----------------: | :-----------: | :----------------------: |
+| [detic_centernet2_r50_fpn_4x_lvis_in21k-lvis](./configs/detic_centernet2_r50_fpn_4x_lvis_in21k-lvis.py) | 32.9 | 33.2 | 30.9 | 29.7 |
+
+#### Standard LVIS Results
+
+| Model (Config) | mask mAP | mask mAP(official) | mask mAP_rare | mask mAP_rare(officical) | Download |
+| :-----------------------------------------------------------------------------------------------------------: | :------: | :----------------: | :-----------: | :----------------------: | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| [detic_centernet2_r50_fpn_4x_lvis_boxsup](./configs/detic_centernet2_r50_fpn_4x_lvis_boxsup.py) | 31.6 | 31.5 | 26.6 | 25.6 | [model](https://download.openmmlab.com/mmdetection/v3.0/detic/detic_centernet2_r50_fpn_4x_lvis_boxsup/detic_centernet2_r50_fpn_4x_lvis_boxsup_20230911_233514-54116677.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/detic/detic_centernet2_r50_fpn_4x_lvis_boxsup/detic_centernet2_r50_fpn_4x_lvis_boxsup_20230911_233514.log.json) |
+| [detic_centernet2_r50_fpn_4x_lvis_in21k-lvis](./configs/detic_centernet2_r50_fpn_4x_lvis_in21k-lvis.py) | 32.9 | 33.2 | 30.9 | 29.7 | [model](https://download.openmmlab.com/mmdetection/v3.0/detic/detic_centernet2_r50_fpn_4x_lvis_in21k-lvis/detic_centernet2_r50_fpn_4x_lvis_in21k-lvis_20230912_040619-9e7a3258.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/detic/detic_centernet2_r50_fpn_4x_lvis_in21k-lvis/detic_centernet2_r50_fpn_4x_lvis_in21k-lvis_20230912_040619.log.json) |
+| [detic_centernet2_swin-b_fpn_4x_lvis_boxsup](./configs/detic_centernet2_swin-b_fpn_4x_lvis_boxsup.py) | 40.7 | 40.7 | 38.0 | 35.9 | [model](https://download.openmmlab.com/mmdetection/v3.0/detic/detic_centernet2_swin-b_fpn_4x_lvis_boxsup/detic_centernet2_swin-b_fpn_4x_lvis_boxsup_20230825_061737-328e85f9.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/detic/detic_centernet2_swin-b_fpn_4x_lvis_boxsup/detic_centernet2_swin-b_fpn_4x_lvis_boxsup_20230825_061737.log.json) |
+| [detic_centernet2_swin-b_fpn_4x_lvis_in21k-lvis](./configs/detic_centernet2_swin-b_fpn_4x_lvis_in21k-lvis.py) | 41.7 | 41.7 | 41.7 | 41.7 | [model](https://download.openmmlab.com/mmdetection/v3.0/detic/detic_centernet2_swin-b_fpn_4x_lvis_in21k-lvis/detic_centernet2_swin-b_fpn_4x_lvis_in21k-lvis_20230926_235410-0c152391.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/detic/detic_centernet2_swin-b_fpn_4x_lvis_in21k-lvis/detic_centernet2_swin-b_fpn_4x_lvis_in21k-lvis_20230926_235410.log.json) |
+
+#### Open-vocabulary LVIS Results
+
+| Model (Config) | mask mAP | mask mAP(official) | mask mAP_rare | mask mAP_rare(officical) | Download |
+| :---------------------------------------------------------------------------------------------------------------: | :------: | :----------------: | :-----------: | :----------------------: | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| [detic_centernet2_r50_fpn_4x_lvis-base_boxsup](./configs/detic_centernet2_r50_fpn_4x_lvis-base_boxsup.py) | 30.4 | 30.2 | 16.2 | 16.4 | [model](https://download.openmmlab.com/mmdetection/v3.0/detic/detic_centernet2_r50_fpn_4x_lvis-base_boxsup/detic_centernet2_r50_fpn_4x_lvis-base_boxsup_20230921_180638-c1685ee2.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/detic/detic_centernet2_r50_fpn_4x_lvis-base_boxsup/detic_centernet2_r50_fpn_4x_lvis-base_boxsup_20230921_180638.log.json) |
+| [detic_centernet2_r50_fpn_4x_lvis-base_in21k-lvis](./configs/detic_centernet2_r50_fpn_4x_lvis-base_in21k-lvis.py) | 32.6 | 32.4 | 27.4 | 24.9 | [model](https://download.openmmlab.com/mmdetection/v3.0/detic/detic_centernet2_r50_fpn_4x_lvis-base_in21k-lvis/detic_centernet2_r50_fpn_4x_lvis-base_in21k-lvis_20230925_014315-2d2cc8b7.pth) \| [log](https://download.openmmlab.com/mmdetection/v3.0/detic/detic_centernet2_r50_fpn_4x_lvis-base_in21k-lvis/detic_centernet2_r50_fpn_4x_lvis-base_in21k-lvis_20230925_014315.log.json) |
+
+### Testing
+
+#### Test Command
+
+To evaluate a model with a trained model, run
+
+```shell
+python ./tools/test.py ${CONFIG_FILE} ${CHECKPOINT_FILE}
+```
+
+#### Open-vocabulary LVIS Results
+
+The models are converted from the official model zoo.
+
+| Model (Config) | mask mAP | mask mAP_novel | Download |
+| :---------------------------------------------------------------------------------------------------------------------: | :------: | :------------: | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: |
+| [detic_centernet2_swin-b_fpn_4x_lvis-base_boxsup](./configs/detic_centernet2_swin-b_fpn_4x_lvis-base_boxsup.py) | 38.4 | 21.9 | [model](https://download.openmmlab.com/mmdetection/v3.0/detic/detic_centernet2_swin-b_fpn_4x_lvis-base_boxsup/detic_centernet2_swin-b_fpn_4x_lvis-base_boxsup-481281c8.pth) |
+| [detic_centernet2_swin-b_fpn_4x_lvis-base_in21k-lvis](./configs/detic_centernet2_swin-b_fpn_4x_lvis-base_in21k-lvis.py) | 40.7 | 34.0 | [model](https://download.openmmlab.com/mmdetection/v3.0/detic/detic_centernet2_swin-b_fpn_4x_lvis-base_in21k-lvis/detic_centernet2_swin-b_fpn_4x_lvis-base_in21k-lvis-ec91245d.pth) |
+
+###### Note:
+
+- The open-vocabulary LVIS setup is LVIS without rare class annotations in training, termed `lvisbase`. We evaluate rare classes as novel classes in testing.
+- ` in21k-lvis` denotes that the model use the overlap classes between ImageNet-21K and LVIS as image-labeled data.
+
+## Citation
+
+If you find Detic is useful in your research or applications, please consider giving a star 🌟 to the [official repository](https://github.com/facebookresearch/Detic) and citing Detic by the following BibTeX entry.
+
+```BibTeX
+@inproceedings{zhou2022detecting,
+ title={Detecting Twenty-thousand Classes using Image-level Supervision},
+ author={Zhou, Xingyi and Girdhar, Rohit and Joulin, Armand and Kr{\"a}henb{\"u}hl, Philipp and Misra, Ishan},
+ booktitle={ECCV},
+ year={2022}
+}
+```
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_r50_fpn_4x_lvis-base_boxsup.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_r50_fpn_4x_lvis-base_boxsup.py
new file mode 100644
index 0000000000000000000000000000000000000000..8ca57b77d7f3bb50e926e3fdb1f85cb29e53777e
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_r50_fpn_4x_lvis-base_boxsup.py
@@ -0,0 +1,9 @@
+_base_ = './detic_centernet2_r50_fpn_4x_lvis_boxsup.py'
+
+# 'lvis_v1_train_norare.json' is the annotations of lvis_v1
+# removing the labels of 337 rare-class
+train_dataloader = dict(
+ dataset=dict(
+ type='ClassBalancedDataset',
+ oversample_thr=1e-3,
+ dataset=dict(ann_file='annotations/lvis_v1_train_norare.json')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_r50_fpn_4x_lvis-base_in21k-lvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_r50_fpn_4x_lvis-base_in21k-lvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..034acb6ebc49a05c459f950a01053aed326109c8
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_r50_fpn_4x_lvis-base_in21k-lvis.py
@@ -0,0 +1,93 @@
+_base_ = './detic_centernet2_r50_fpn_4x_lvis_boxsup.py'
+dataset_type = ['LVISV1Dataset', 'ImageNetLVISV1Dataset']
+image_size_det = (640, 640)
+image_size_cls = (320, 320)
+
+# backend = 'pillow'
+backend_args = None
+
+train_pipeline_det = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomResize',
+ scale=image_size_det,
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size_det,
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_pipeline_cls = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=False, with_label=True),
+ dict(
+ type='RandomResize',
+ scale=image_size_cls,
+ ratio_range=(0.5, 1.5),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size_cls,
+ recompute_bbox=False,
+ bbox_clip_border=False,
+ allow_negative_crop=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+# 'lvis_v1_train_norare.json' is the annotations of lvis_v1
+# removing the labels of 337 rare-class
+dataset_det = dict(
+ type='ClassBalancedDataset',
+ oversample_thr=1e-3,
+ dataset=dict(
+ type='LVISV1Dataset',
+ data_root='data/lvis/',
+ ann_file='annotations/lvis_v1_train_norare.json',
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline_det,
+ backend_args=backend_args))
+
+dataset_cls = dict(
+ type='ImageNetLVISV1Dataset',
+ data_root='data/imagenet',
+ ann_file='annotations/imagenet_lvis_image_info.json',
+ data_prefix=dict(img='ImageNet-LVIS/'),
+ pipeline=train_pipeline_cls,
+ backend_args=backend_args)
+
+train_dataloader = dict(
+ _delete_=True,
+ batch_size=[8, 32],
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='MultiDataSampler', dataset_ratio=[1, 4]),
+ batch_sampler=dict(
+ type='MultiDataAspectRatioBatchSampler', num_datasets=2),
+ dataset=dict(type='ConcatDataset', datasets=[dataset_det, dataset_cls]))
+
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0,
+ end=1000),
+ dict(
+ type='CosineAnnealingLR',
+ begin=0,
+ by_epoch=False,
+ T_max=90000,
+ )
+]
+
+load_from = './first_stage/detic_centernet2_r50_fpn_4x_lvis-base_boxsup.pth'
+
+find_unused_parameters = True
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_r50_fpn_4x_lvis_boxsup.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_r50_fpn_4x_lvis_boxsup.py
new file mode 100644
index 0000000000000000000000000000000000000000..a11be374cc785903ca65e253e7ae3c7e6ff87429
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_r50_fpn_4x_lvis_boxsup.py
@@ -0,0 +1,410 @@
+_base_ = 'mmdet::_base_/default_runtime.py'
+dataset_type = 'LVISV1Dataset'
+custom_imports = dict(
+ imports=['projects.Detic_new.detic'], allow_failed_imports=False)
+
+num_classes = 1203
+lvis_cat_frequency_info = 'data/metadata/lvis_v1_train_cat_info.json'
+
+# 'data/metadata/lvis_v1_clip_a+cname.npy' is pre-computed
+# CLIP embeddings for each category
+cls_layer = dict(
+ type='ZeroShotClassifier',
+ zs_weight_path='data/metadata/lvis_v1_clip_a+cname.npy',
+ zs_weight_dim=512,
+ use_bias=0.0,
+ norm_weight=True,
+ norm_temperature=50.0)
+reg_layer = [
+ dict(type='Linear', in_features=1024, out_features=1024),
+ dict(type='ReLU', inplace=True),
+ dict(type='Linear', in_features=1024, out_features=4)
+]
+
+model = dict(
+ type='Detic',
+ data_preprocessor=dict(
+ type='DetDataPreprocessor',
+ mean=[123.675, 116.28, 103.53],
+ std=[58.395, 57.12, 57.375],
+ bgr_to_rgb=True,
+ pad_size_divisor=32),
+ backbone=dict(
+ type='ResNet',
+ depth=50,
+ num_stages=4,
+ out_indices=(1, 2, 3),
+ norm_cfg=dict(type='BN', requires_grad=False),
+ norm_eval=True,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='https://miil-public-eu.oss-eu-central-1.aliyuncs.com/'
+ 'model-zoo/ImageNet_21K_P/models/resnet50_miil_21k.pth')),
+ neck=dict(
+ type='FPN',
+ in_channels=[512, 1024, 2048],
+ out_channels=256,
+ start_level=0,
+ add_extra_convs='on_output',
+ num_outs=5,
+ init_cfg=dict(type='Caffe2Xavier', layer='Conv2d'),
+ relu_before_extra_convs=True),
+ rpn_head=dict(
+ type='CenterNetRPNHead',
+ num_classes=1,
+ in_channels=256,
+ stacked_convs=4,
+ feat_channels=256,
+ strides=[8, 16, 32, 64, 128],
+ conv_bias=True,
+ norm_cfg=dict(type='GN', num_groups=32, requires_grad=True),
+ loss_cls=dict(
+ type='HeatmapFocalLoss',
+ alpha=0.25,
+ beta=4.0,
+ gamma=2.0,
+ pos_weight=0.5,
+ neg_weight=0.5,
+ loss_weight=1.0,
+ ignore_high_fp=0.85,
+ ),
+ loss_bbox=dict(type='GIoULoss', eps=1e-6, loss_weight=1.0),
+ ),
+ roi_head=dict(
+ type='DeticRoIHead',
+ num_stages=3,
+ stage_loss_weights=[1.0, 1.0, 1.0],
+ bbox_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(
+ type='RoIAlign',
+ output_size=7,
+ sampling_ratio=0,
+ use_torchvision=True),
+ out_channels=256,
+ featmap_strides=[8, 16, 32],
+ # approximately equal to
+ # canonical_box_size=224, canonical_level=4 in D2
+ finest_scale=112),
+ bbox_head=[
+ dict(
+ type='DeticBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=num_classes,
+ cls_predictor_cfg=cls_layer,
+ reg_predictor_cfg=reg_layer,
+ use_fed_loss=True,
+ cat_freq_path=lvis_cat_frequency_info,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.1, 0.1, 0.2, 0.2]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=0.1,
+ loss_weight=1.0)),
+ dict(
+ type='DeticBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=num_classes,
+ cls_predictor_cfg=cls_layer,
+ reg_predictor_cfg=reg_layer,
+ use_fed_loss=True,
+ cat_freq_path=lvis_cat_frequency_info,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.05, 0.05, 0.1, 0.1]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=0.1,
+ loss_weight=1.0)),
+ dict(
+ type='DeticBBoxHead',
+ in_channels=256,
+ fc_out_channels=1024,
+ roi_feat_size=7,
+ num_classes=num_classes,
+ cls_predictor_cfg=cls_layer,
+ reg_predictor_cfg=reg_layer,
+ use_fed_loss=True,
+ cat_freq_path=lvis_cat_frequency_info,
+ bbox_coder=dict(
+ type='DeltaXYWHBBoxCoder',
+ target_means=[0., 0., 0., 0.],
+ target_stds=[0.033, 0.033, 0.067, 0.067]),
+ reg_class_agnostic=True,
+ loss_cls=dict(
+ type='CrossEntropyLoss', use_sigmoid=True,
+ loss_weight=1.0),
+ loss_bbox=dict(type='SmoothL1Loss', beta=0.1, loss_weight=1.0))
+ ],
+ mask_roi_extractor=dict(
+ type='SingleRoIExtractor',
+ roi_layer=dict(type='RoIAlign', output_size=14, sampling_ratio=0),
+ out_channels=256,
+ featmap_strides=[8, 16, 32],
+ # approximately equal to
+ # canonical_box_size=224, canonical_level=4 in D2
+ finest_scale=112),
+ mask_head=dict(
+ type='FCNMaskHead',
+ num_convs=4,
+ in_channels=256,
+ conv_out_channels=256,
+ class_agnostic=True,
+ num_classes=num_classes,
+ loss_mask=dict(
+ type='CrossEntropyLoss', use_mask=True, loss_weight=1.0))),
+ # model training and testing settings
+ train_cfg=dict(
+ rpn=dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.3,
+ min_pos_iou=0.3,
+ match_low_quality=True,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=256,
+ pos_fraction=0.5,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ allowed_border=0,
+ pos_weight=-1,
+ debug=False),
+ rpn_proposal=dict(
+ score_thr=0.0001,
+ nms_pre=4000,
+ max_per_img=2000,
+ nms=dict(type='nms', iou_threshold=0.9),
+ min_bbox_size=0),
+ rcnn=[
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.6,
+ neg_iou_thr=0.6,
+ min_pos_iou=0.6,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=True),
+ mask_size=28,
+ pos_weight=-1,
+ debug=False),
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.7,
+ neg_iou_thr=0.7,
+ min_pos_iou=0.7,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ mask_size=28,
+ pos_weight=-1,
+ debug=False),
+ dict(
+ assigner=dict(
+ type='MaxIoUAssigner',
+ pos_iou_thr=0.8,
+ neg_iou_thr=0.8,
+ min_pos_iou=0.8,
+ match_low_quality=False,
+ ignore_iof_thr=-1),
+ sampler=dict(
+ type='RandomSampler',
+ num=512,
+ pos_fraction=0.25,
+ neg_pos_ub=-1,
+ add_gt_as_proposals=False),
+ mask_size=28,
+ pos_weight=-1,
+ debug=False)
+ ]),
+ test_cfg=dict(
+ rpn=dict(
+ score_thr=0.0001,
+ nms_pre=1000,
+ max_per_img=256,
+ nms=dict(type='nms', iou_threshold=0.9),
+ min_bbox_size=0),
+ rcnn=dict(
+ score_thr=0.02,
+ nms=dict(type='nms', iou_threshold=0.5),
+ max_per_img=300,
+ mask_thr_binary=0.5)))
+
+# backend = 'pillow'
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomResize',
+ scale=(640, 640),
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(640, 640),
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+test_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ backend_args=backend_args,
+ imdecode_backend=backend_args),
+ dict(
+ type='Resize',
+ scale=(1333, 800),
+ keep_ratio=True,
+ backend=backend_args),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ poly2mask=False),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor', 'text', 'custom_entities'))
+]
+
+val_pipeline = [
+ dict(
+ type='LoadImageFromFile',
+ backend_args=backend_args,
+ imdecode_backend=backend_args),
+ dict(
+ type='Resize',
+ scale=(1333, 800),
+ keep_ratio=True,
+ backend=backend_args),
+ dict(
+ type='LoadAnnotations',
+ with_bbox=True,
+ with_mask=True,
+ poly2mask=False),
+ dict(
+ type='PackDetInputs',
+ meta_keys=('img_id', 'img_path', 'ori_shape', 'img_shape',
+ 'scale_factor'))
+]
+
+train_dataloader = dict(
+ batch_size=8,
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='DefaultSampler', shuffle=True),
+ batch_sampler=dict(type='AspectRatioBatchSampler'),
+ dataset=dict(
+ type='ClassBalancedDataset',
+ oversample_thr=1e-3,
+ dataset=dict(
+ type='LVISV1Dataset',
+ data_root='data/lvis/',
+ ann_file='annotations/lvis_v1_train_norare.json',
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline,
+ backend_args=backend_args)))
+
+val_dataloader = dict(
+ batch_size=8,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ pin_memory=True,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type='LVISV1Dataset',
+ data_root='data/lvis/',
+ ann_file='annotations/lvis_v1_val.json',
+ data_prefix=dict(img=''),
+ pipeline=val_pipeline,
+ return_classes=False))
+
+test_dataloader = dict(
+ batch_size=8,
+ num_workers=2,
+ persistent_workers=True,
+ drop_last=False,
+ pin_memory=True,
+ sampler=dict(type='DefaultSampler', shuffle=False),
+ dataset=dict(
+ type='LVISV1Dataset',
+ data_root='data/lvis/',
+ ann_file='annotations/lvis_v1_val.json',
+ data_prefix=dict(img=''),
+ pipeline=test_pipeline,
+ return_classes=True))
+
+val_evaluator = dict(
+ type='LVISMetric',
+ ann_file='data/lvis/annotations/lvis_v1_val.json',
+ metric=['bbox', 'segm'])
+test_evaluator = val_evaluator
+
+# training schedule for 90k with batch_size of 64
+# with total batch_size of 16, 90k iters is equivalent to '1x' (12 epochs)
+# with total batch_size of 64, 90k iters is equivalent to '4x'
+max_iter = 90000
+train_cfg = dict(
+ type='IterBasedTrainLoop', max_iters=max_iter, val_interval=90000)
+val_cfg = dict(type='ValLoop')
+test_cfg = dict(type='TestLoop')
+
+# Enable automatic-mixed-precision training with AmpOptimWrapper.
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0002, weight_decay=0.0001),
+ paramwise_cfg=dict(norm_decay_mult=0.),
+ clip_grad=dict(max_norm=1.0, norm_type=2))
+
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=0.0001,
+ by_epoch=False,
+ begin=0,
+ end=10000),
+ dict(
+ type='CosineAnnealingLR',
+ begin=0,
+ by_epoch=False,
+ T_max=max_iter,
+ )
+]
+
+# only keep latest 5 checkpoints
+default_hooks = dict(
+ checkpoint=dict(by_epoch=False, interval=30000, max_keep_ckpts=5),
+ logger=dict(type='LoggerHook', interval=50))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_r50_fpn_4x_lvis_in21k-lvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_r50_fpn_4x_lvis_in21k-lvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..ce97ed6d589aad27e47d6a5e836f933a67eba42d
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_r50_fpn_4x_lvis_in21k-lvis.py
@@ -0,0 +1,91 @@
+_base_ = './detic_centernet2_r50_fpn_4x_lvis_boxsup.py'
+dataset_type = ['LVISV1Dataset', 'ImageNetLVISV1Dataset']
+image_size_det = (640, 640)
+image_size_cls = (320, 320)
+
+# backend = 'pillow'
+backend_args = None
+
+train_pipeline_det = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomResize',
+ scale=image_size_det,
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size_det,
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_pipeline_cls = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=False, with_label=True),
+ dict(
+ type='RandomResize',
+ scale=image_size_cls,
+ ratio_range=(0.5, 1.5),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size_cls,
+ recompute_bbox=False,
+ bbox_clip_border=False,
+ allow_negative_crop=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+dataset_det = dict(
+ type='ClassBalancedDataset',
+ oversample_thr=1e-3,
+ dataset=dict(
+ type='LVISV1Dataset',
+ data_root='data/lvis/',
+ ann_file='annotations/lvis_v1_train.json',
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline_det,
+ backend_args=backend_args))
+
+dataset_cls = dict(
+ type='ImageNetLVISV1Dataset',
+ data_root='data/imagenet',
+ ann_file='annotations/imagenet_lvis_image_info.json',
+ data_prefix=dict(img='ImageNet-LVIS/'),
+ pipeline=train_pipeline_cls,
+ backend_args=backend_args)
+
+train_dataloader = dict(
+ _delete_=True,
+ batch_size=[8, 32],
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='MultiDataSampler', dataset_ratio=[1, 4]),
+ batch_sampler=dict(
+ type='MultiDataAspectRatioBatchSampler', num_datasets=2),
+ dataset=dict(type='ConcatDataset', datasets=[dataset_det, dataset_cls]))
+
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0,
+ end=1000),
+ dict(
+ type='CosineAnnealingLR',
+ begin=0,
+ by_epoch=False,
+ T_max=90000,
+ )
+]
+
+load_from = './first_stage/detic_centernet2_r50_fpn_4x_lvis_boxsup.pth'
+
+find_unused_parameters = True
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_swin-b_fpn_4x_lvis-base_boxsup.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_swin-b_fpn_4x_lvis-base_boxsup.py
new file mode 100644
index 0000000000000000000000000000000000000000..efedd111e2f796bc50ec56aabfe359fc32dbd4d3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_swin-b_fpn_4x_lvis-base_boxsup.py
@@ -0,0 +1,9 @@
+_base_ = './detic_centernet2_swin-b_fpn_4x_lvis_boxsup.py'
+
+# 'lvis_v1_train_norare.json' is the annotations of lvis_v1
+# removing the labels of 337 rare-class
+train_dataloader = dict(
+ dataset=dict(
+ type='ClassBalancedDataset',
+ oversample_thr=1e-3,
+ dataset=dict(ann_file='annotations/lvis_v1_train_norare.json')))
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_swin-b_fpn_4x_lvis-base_in21k-lvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_swin-b_fpn_4x_lvis-base_in21k-lvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..1df70970e2d97cd1c794e6db7eb390799264117b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_swin-b_fpn_4x_lvis-base_in21k-lvis.py
@@ -0,0 +1,118 @@
+_base_ = './detic_centernet2_r50_fpn_4x_lvis_in21k-lvis.py'
+
+image_size_det = (896, 896)
+image_size_cls = (448, 448)
+
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ embed_dims=128,
+ depths=[2, 2, 18, 2],
+ num_heads=[4, 8, 16, 32],
+ window_size=7,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(1, 2, 3),
+ with_cp=False),
+ neck=dict(in_channels=[256, 512, 1024]))
+
+backend_args = None
+train_pipeline_det = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomResize',
+ scale=image_size_det,
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size_det,
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_pipeline_cls = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=False, with_label=True),
+ dict(
+ type='RandomResize',
+ scale=image_size_cls,
+ ratio_range=(0.5, 1.5),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size_cls,
+ recompute_bbox=False,
+ bbox_clip_border=False,
+ allow_negative_crop=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+# 'lvis_v1_train_norare.json' is the annotations of lvis_v1
+# removing the labels of 337 rare-class
+dataset_det = dict(
+ type='ClassBalancedDataset',
+ oversample_thr=1e-3,
+ dataset=dict(
+ type='LVISV1Dataset',
+ data_root='data/lvis/',
+ ann_file='annotations/lvis_v1_train_norare.json',
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline_det,
+ backend_args=backend_args))
+
+dataset_cls = dict(
+ type='ImageNetLVISV1Dataset',
+ data_root='data/imagenet',
+ ann_file='annotations/imagenet_lvis_image_info.json',
+ data_prefix=dict(img='ImageNet-LVIS/'),
+ pipeline=train_pipeline_cls,
+ backend_args=backend_args)
+
+train_dataloader = dict(
+ _delete_=True,
+ batch_size=[4, 16],
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='MultiDataSampler', dataset_ratio=[1, 4]),
+ batch_sampler=dict(
+ type='MultiDataAspectRatioBatchSampler', num_datasets=2),
+ dataset=dict(type='ConcatDataset', datasets=[dataset_det, dataset_cls]))
+
+# training schedule for 180k
+max_iter = 180000
+train_cfg = dict(
+ type='IterBasedTrainLoop', max_iters=max_iter, val_interval=180000)
+
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0001, weight_decay=0.0001))
+
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0,
+ end=1000),
+ dict(
+ type='CosineAnnealingLR',
+ begin=0,
+ by_epoch=False,
+ T_max=max_iter,
+ )
+]
+
+load_from = './first_stage/detic_centernet2_swin-b_fpn_4x_lvis-base_boxsup.pth'
+find_unused_parameters = True
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_swin-b_fpn_4x_lvis_boxsup.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_swin-b_fpn_4x_lvis_boxsup.py
new file mode 100644
index 0000000000000000000000000000000000000000..ce04a815facba0a1f2753e2351a6f1b903824d32
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_swin-b_fpn_4x_lvis_boxsup.py
@@ -0,0 +1,78 @@
+_base_ = './detic_centernet2_r50_fpn_4x_lvis_boxsup.py'
+
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ embed_dims=128,
+ depths=[2, 2, 18, 2],
+ num_heads=[4, 8, 16, 32],
+ window_size=7,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(1, 2, 3),
+ with_cp=False,
+ convert_weights=True,
+ init_cfg=dict(
+ type='Pretrained',
+ checkpoint='https://github.com/SwinTransformer/storage/releases/'
+ 'download/v1.0.0/swin_base_patch4_window7_224_22k.pth')),
+ neck=dict(in_channels=[256, 512, 1024]))
+
+# backend = 'pillow'
+backend_args = None
+
+train_pipeline = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomResize',
+ scale=(896, 896),
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=(896, 896),
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_dataloader = dict(
+ dataset=dict(
+ type='ClassBalancedDataset',
+ oversample_thr=1e-3,
+ dataset=dict(pipeline=train_pipeline)))
+
+# training schedule for 180k
+max_iter = 180000
+train_cfg = dict(
+ type='IterBasedTrainLoop', max_iters=max_iter, val_interval=180000)
+
+# Enable automatic-mixed-precision training with AmpOptimWrapper.
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0001, weight_decay=0.0001))
+
+param_scheduler = [
+ dict(
+ type='LinearLR',
+ start_factor=0.0001,
+ by_epoch=False,
+ begin=0,
+ end=10000),
+ dict(
+ type='CosineAnnealingLR',
+ begin=0,
+ by_epoch=False,
+ T_max=max_iter,
+ )
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_swin-b_fpn_4x_lvis_coco_in21k.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_swin-b_fpn_4x_lvis_coco_in21k.py
new file mode 100644
index 0000000000000000000000000000000000000000..a9ab2c69adaa7eaa94082a9fd8e960c00f4925f3
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_swin-b_fpn_4x_lvis_coco_in21k.py
@@ -0,0 +1,2 @@
+# not support training, only for testing
+_base_ = './detic_centernet2_swin-b_fpn_4x_lvis_in21k-lvis.py'
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_swin-b_fpn_4x_lvis_in21k-lvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_swin-b_fpn_4x_lvis_in21k-lvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..de358ac3460934738ceb3fae39a37b8ab0216069
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/configs/detic_centernet2_swin-b_fpn_4x_lvis_in21k-lvis.py
@@ -0,0 +1,116 @@
+_base_ = './detic_centernet2_r50_fpn_4x_lvis_in21k-lvis.py'
+
+image_size_det = (896, 896)
+image_size_cls = (448, 448)
+
+model = dict(
+ backbone=dict(
+ _delete_=True,
+ type='SwinTransformer',
+ embed_dims=128,
+ depths=[2, 2, 18, 2],
+ num_heads=[4, 8, 16, 32],
+ window_size=7,
+ mlp_ratio=4,
+ qkv_bias=True,
+ qk_scale=None,
+ drop_rate=0.,
+ attn_drop_rate=0.,
+ drop_path_rate=0.3,
+ patch_norm=True,
+ out_indices=(1, 2, 3),
+ with_cp=False),
+ neck=dict(in_channels=[256, 512, 1024]))
+
+backend_args = None
+train_pipeline_det = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=True, with_mask=True),
+ dict(
+ type='RandomResize',
+ scale=image_size_det,
+ ratio_range=(0.1, 2.0),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size_det,
+ recompute_bbox=True,
+ allow_negative_crop=True),
+ dict(type='FilterAnnotations', min_gt_bbox_wh=(1e-2, 1e-2)),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+train_pipeline_cls = [
+ dict(type='LoadImageFromFile', backend_args=backend_args),
+ dict(type='LoadAnnotations', with_bbox=False, with_label=True),
+ dict(
+ type='RandomResize',
+ scale=image_size_cls,
+ ratio_range=(0.5, 1.5),
+ keep_ratio=True),
+ dict(
+ type='RandomCrop',
+ crop_type='absolute_range',
+ crop_size=image_size_cls,
+ recompute_bbox=False,
+ bbox_clip_border=False,
+ allow_negative_crop=True),
+ dict(type='RandomFlip', prob=0.5),
+ dict(type='PackDetInputs')
+]
+
+dataset_det = dict(
+ type='ClassBalancedDataset',
+ oversample_thr=1e-3,
+ dataset=dict(
+ type='LVISV1Dataset',
+ data_root='data/lvis/',
+ ann_file='annotations/lvis_v1_train.json',
+ data_prefix=dict(img=''),
+ filter_cfg=dict(filter_empty_gt=True, min_size=32),
+ pipeline=train_pipeline_det,
+ backend_args=backend_args))
+
+dataset_cls = dict(
+ type='ImageNetLVISV1Dataset',
+ data_root='data/imagenet',
+ ann_file='annotations/imagenet_lvis_image_info.json',
+ data_prefix=dict(img='ImageNet-LVIS/'),
+ pipeline=train_pipeline_cls,
+ backend_args=backend_args)
+
+train_dataloader = dict(
+ _delete_=True,
+ batch_size=[4, 16],
+ num_workers=2,
+ persistent_workers=True,
+ sampler=dict(type='MultiDataSampler', dataset_ratio=[1, 4]),
+ batch_sampler=dict(
+ type='MultiDataAspectRatioBatchSampler', num_datasets=2),
+ dataset=dict(type='ConcatDataset', datasets=[dataset_det, dataset_cls]))
+
+# training schedule for 180k
+max_iter = 180000
+train_cfg = dict(
+ type='IterBasedTrainLoop', max_iters=max_iter, val_interval=180000)
+
+optim_wrapper = dict(
+ type='OptimWrapper',
+ optimizer=dict(type='AdamW', lr=0.0001, weight_decay=0.0001))
+
+param_scheduler = [
+ dict(
+ type='LinearLR', start_factor=0.001, by_epoch=False, begin=0,
+ end=1000),
+ dict(
+ type='CosineAnnealingLR',
+ begin=0,
+ by_epoch=False,
+ T_max=max_iter,
+ )
+]
+
+load_from = './first_stage/detic_centernet2_swin-b_fpn_4x_lvis_boxsup.pth'
+find_unused_parameters = True
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/__init__.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/__init__.py
new file mode 100644
index 0000000000000000000000000000000000000000..e4b0d7bb8c883849df66d5183f096c64ab7c119b
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/__init__.py
@@ -0,0 +1,13 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from .centernet_rpn_head import CenterNetRPNHead
+from .detic import Detic
+from .detic_bbox_head import DeticBBoxHead
+from .detic_roi_head import DeticRoIHead
+from .heatmap_focal_loss import HeatmapFocalLoss
+from .imagenet_lvis import ImageNetLVISV1Dataset
+from .zero_shot_classifier import ZeroShotClassifier
+
+__all__ = [
+ 'CenterNetRPNHead', 'Detic', 'DeticBBoxHead', 'DeticRoIHead',
+ 'ZeroShotClassifier', 'HeatmapFocalLoss', 'ImageNetLVISV1Dataset'
+]
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/centernet_rpn_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/centernet_rpn_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..629872824d7c8f23542e3414f7faed71ea4efa93
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/centernet_rpn_head.py
@@ -0,0 +1,573 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from typing import Dict, List, Optional, Sequence, Tuple
+
+import torch
+import torch.nn as nn
+from mmcv.cnn import Scale
+from mmengine import ConfigDict
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.models.dense_heads import CenterNetUpdateHead
+from mmdet.models.utils import unpack_gt_instances
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import bbox2distance
+from mmdet.utils import (ConfigType, InstanceList, OptConfigType,
+ OptInstanceList, reduce_mean)
+from .iou_loss import IOULoss
+
+# from .heatmap_focal_loss import binary_heatmap_focal_loss_jit
+INF = 1000000000
+RangeType = Sequence[Tuple[int, int]]
+
+
+@MODELS.register_module()
+class CenterNetRPNHead(CenterNetUpdateHead):
+ """CenterNetUpdateHead is an improved version of CenterNet in CenterNet2.
+
+ Paper link ``_.
+ Args:
+ num_classes (int): Number of categories excluding the background
+ category.
+ in_channels (int): Number of channel in the input feature map.
+ regress_ranges (Sequence[Tuple[int, int]]): Regress range of multiple
+ level points.
+ hm_min_radius (int): Heatmap target minimum radius of cls branch.
+ Defaults to 4.
+ hm_min_overlap (float): Heatmap target minimum overlap of cls branch.
+ Defaults to 0.8.
+ more_pos_thresh (float): The filtering threshold when the cls branch
+ adds more positive samples. Defaults to 0.2.
+ more_pos_topk (int): The maximum number of additional positive samples
+ added to each gt. Defaults to 9.
+ soft_weight_on_reg (bool): Whether to use the soft target of the
+ cls branch as the soft weight of the bbox branch.
+ Defaults to False.
+ loss_cls (:obj:`ConfigDict` or dict): Config of cls loss. Defaults to
+ dict(type='GaussianFocalLoss', loss_weight=1.0)
+ loss_bbox (:obj:`ConfigDict` or dict): Config of bbox loss. Defaults to
+ dict(type='GIoULoss', loss_weight=2.0).
+ norm_cfg (:obj:`ConfigDict` or dict, optional): dictionary to construct
+ and config norm layer. Defaults to
+ ``norm_cfg=dict(type='GN', num_groups=32, requires_grad=True)``.
+ train_cfg (:obj:`ConfigDict` or dict, optional): Training config.
+ Unused in CenterNet. Reserved for compatibility with
+ SingleStageDetector.
+ test_cfg (:obj:`ConfigDict` or dict, optional): Testing config
+ of CenterNet.
+ """
+
+ def __init__(self,
+ num_classes: int,
+ in_channels: int,
+ regress_ranges: RangeType = ((0, 80), (64, 160), (128, 320),
+ (256, 640), (512, INF)),
+ hm_min_radius: int = 4,
+ hm_min_overlap: float = 0.8,
+ more_pos: bool = False,
+ more_pos_thresh: float = 0.2,
+ more_pos_topk: int = 9,
+ soft_weight_on_reg: bool = False,
+ not_clamp_box: bool = False,
+ loss_cls: ConfigType = dict(
+ type='HeatmapFocalLoss',
+ alpha=0.25,
+ beta=4.0,
+ gamma=2.0,
+ pos_weight=1.0,
+ neg_weight=1.0,
+ sigmoid_clamp=1e-4,
+ ignore_high_fp=-1.0,
+ loss_weight=1.0,
+ ),
+ loss_bbox: ConfigType = dict(
+ type='GIoULoss', loss_weight=2.0),
+ norm_cfg: OptConfigType = dict(
+ type='GN', num_groups=32, requires_grad=True),
+ train_cfg: OptConfigType = None,
+ test_cfg: OptConfigType = None,
+ **kwargs) -> None:
+ super().__init__(
+ num_classes=num_classes,
+ in_channels=in_channels,
+ # loss_bbox=loss_bbox,
+ loss_cls=loss_cls,
+ norm_cfg=norm_cfg,
+ train_cfg=train_cfg,
+ test_cfg=test_cfg,
+ **kwargs)
+ self.soft_weight_on_reg = soft_weight_on_reg
+ self.hm_min_radius = hm_min_radius
+ self.more_pos_thresh = more_pos_thresh
+ self.more_pos_topk = more_pos_topk
+ self.more_pos = more_pos
+ self.not_clamp_box = not_clamp_box
+ self.delta = (1 - hm_min_overlap) / (1 + hm_min_overlap)
+ self.loss_bbox = IOULoss('giou')
+
+ # GaussianFocalLoss must be sigmoid mode
+ self.use_sigmoid_cls = True
+ self.cls_out_channels = num_classes
+
+ self.regress_ranges = regress_ranges
+ self.scales = nn.ModuleList([Scale(1.0) for _ in self.strides])
+
+ def _init_layers(self) -> None:
+ """Initialize layers of the head."""
+ self._init_reg_convs()
+ self._init_predictor()
+
+ def forward_single(self, x: Tensor, scale: Scale,
+ stride: int) -> Tuple[Tensor, Tensor]:
+ """Forward features of a single scale level.
+
+ Args:
+ x (Tensor): FPN feature maps of the specified stride.
+ scale (:obj:`mmcv.cnn.Scale`): Learnable scale module to resize
+ the bbox prediction.
+ stride (int): The corresponding stride for feature maps.
+
+ Returns:
+ tuple: scores for each class, bbox predictions of
+ input feature maps.
+ """
+ for m in self.reg_convs:
+ x = m(x)
+ cls_score = self.conv_cls(x)
+ bbox_pred = self.conv_reg(x)
+ # scale the bbox_pred of different level
+ # float to avoid overflow when enabling FP16
+ bbox_pred = scale(bbox_pred).float()
+ # bbox_pred needed for gradient computation has been modified
+ # by F.relu(bbox_pred) when run with PyTorch 1.10. So replace
+ # F.relu(bbox_pred) with bbox_pred.clamp(min=0)
+ bbox_pred = bbox_pred.clamp(min=0)
+ return cls_score, bbox_pred # score aligned, box larger
+
+ def loss_by_feat(
+ self,
+ cls_scores: List[Tensor],
+ bbox_preds: List[Tensor],
+ batch_gt_instances: InstanceList,
+ batch_img_metas: List[dict],
+ batch_gt_instances_ignore: OptInstanceList = None
+ ) -> Dict[str, Tensor]:
+ """Calculate the loss based on the features extracted by the detection
+ head.
+
+ Args:
+ cls_scores (list[Tensor]): Box scores for each scale level,
+ each is a 4D-tensor, the channel number is num_classes.
+ bbox_preds (list[Tensor]): Box energies / deltas for each scale
+ level, each is a 4D-tensor, the channel number is 4.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes`` and ``labels``
+ attributes.
+ batch_img_metas (list[dict]): Meta information of each image, e.g.,
+ image size, scaling factor, etc.
+ batch_gt_instances_ignore (list[:obj:`InstanceData`], optional):
+ Batch of gt_instances_ignore. It includes ``bboxes`` attribute
+ data that is ignored during training and testing.
+ Defaults to None.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components.
+ """
+
+ num_imgs = cls_scores[0].size(0)
+ assert len(cls_scores) == len(bbox_preds)
+ featmap_sizes = [featmap.size()[-2:] for featmap in cls_scores]
+ all_level_points = self.prior_generator.grid_priors(
+ featmap_sizes,
+ dtype=bbox_preds[0].dtype,
+ device=bbox_preds[0].device)
+
+ # 1 flatten outputs
+ flatten_cls_scores = [
+ cls_score.permute(0, 2, 3, 1).reshape(-1, self.cls_out_channels)
+ for cls_score in cls_scores
+ ]
+ flatten_bbox_preds = [
+ bbox_pred.permute(0, 2, 3, 1).reshape(-1, 4)
+ for bbox_pred in bbox_preds
+ ]
+ flatten_cls_scores = torch.cat(flatten_cls_scores)
+ flatten_bbox_preds = torch.cat(flatten_bbox_preds)
+
+ # repeat points to align with bbox_preds
+ flatten_points = torch.cat(
+ [points.repeat(num_imgs, 1) for points in all_level_points])
+
+ assert (torch.isfinite(flatten_bbox_preds).all().item())
+
+ # 2 calc reg and cls branch targets
+ cls_targets, bbox_targets = self.get_targets(all_level_points,
+ batch_gt_instances)
+
+ # 3 pos index for cls branch
+ featmap_sizes = flatten_points.new_tensor(featmap_sizes)
+
+ if self.more_pos:
+ pos_inds, cls_labels = self.add_cls_pos_inds(
+ flatten_points, flatten_bbox_preds, featmap_sizes,
+ batch_gt_instances)
+ else:
+ pos_inds = self._get_label_inds(batch_gt_instances,
+ batch_img_metas, featmap_sizes)
+
+ # 4 calc cls loss
+ if pos_inds is None:
+ # num_gts=0
+ num_pos_cls = bbox_preds[0].new_tensor(0, dtype=torch.float)
+ else:
+ num_pos_cls = bbox_preds[0].new_tensor(
+ len(pos_inds), dtype=torch.float)
+ num_pos_cls = max(reduce_mean(num_pos_cls), 1.0)
+
+ cat_agn_cls_targets = cls_targets.max(dim=1)[0] # M
+
+ cls_pos_loss, cls_neg_loss = self.loss_cls(
+ flatten_cls_scores.squeeze(1), cat_agn_cls_targets, pos_inds,
+ num_pos_cls)
+
+ # 5 calc reg loss
+ pos_bbox_inds = torch.nonzero(
+ bbox_targets.max(dim=1)[0] >= 0).squeeze(1)
+ pos_bbox_preds = flatten_bbox_preds[pos_bbox_inds]
+ pos_bbox_targets = bbox_targets[pos_bbox_inds]
+
+ bbox_weight_map = cls_targets.max(dim=1)[0]
+ bbox_weight_map = bbox_weight_map[pos_bbox_inds]
+ bbox_weight_map = bbox_weight_map if self.soft_weight_on_reg \
+ else torch.ones_like(bbox_weight_map)
+
+ num_pos_bbox = max(reduce_mean(bbox_weight_map.sum()), 1.0)
+
+ if len(pos_bbox_inds) > 0:
+ bbox_loss = self.loss_bbox(
+ pos_bbox_preds,
+ pos_bbox_targets,
+ bbox_weight_map,
+ reduction='sum') / num_pos_bbox
+ else:
+ bbox_loss = flatten_bbox_preds.sum() * 0
+
+ return dict(
+ loss_bbox=bbox_loss,
+ loss_cls_pos=cls_pos_loss,
+ loss_cls_neg=cls_neg_loss)
+
+ def loss_and_predict(
+ self,
+ x: Tuple[Tensor],
+ batch_data_samples: SampleList,
+ proposal_cfg: Optional[ConfigDict] = None
+ ) -> Tuple[dict, InstanceList]:
+ """Perform forward propagation of the head, then calculate loss and
+ predictions from the features and data samples.
+
+ Args:
+ x (tuple[Tensor]): Features from FPN.
+ batch_data_samples (list[:obj:`DetDataSample`]): Each item contains
+ the meta information of each image and corresponding
+ annotations.
+ proposal_cfg (ConfigDict, optional): Test / postprocessing
+ configuration, if None, test_cfg would be used.
+ Defaults to None.
+
+ Returns:
+ tuple: the return value is a tuple contains:
+
+ - losses: (dict[str, Tensor]): A dictionary of loss components.
+ - predictions (list[:obj:`InstanceData`]): Detection
+ results of each image after the post process.
+ """
+ outputs = unpack_gt_instances(batch_data_samples)
+ (batch_gt_instances, batch_gt_instances_ignore,
+ batch_img_metas) = outputs
+
+ outs = self(x)
+
+ loss_inputs = outs + (batch_gt_instances, batch_img_metas,
+ batch_gt_instances_ignore)
+ losses = self.loss_by_feat(*loss_inputs)
+ predictions = self.predict_by_feat(
+ *outs, batch_img_metas=batch_img_metas, cfg=proposal_cfg)
+ return losses, predictions
+
+ def _predict_by_feat_single(self,
+ cls_score_list: List[Tensor],
+ bbox_pred_list: List[Tensor],
+ score_factor_list: List[Tensor],
+ mlvl_priors: List[Tensor],
+ img_meta: dict,
+ cfg: ConfigDict,
+ rescale: bool = False,
+ with_nms: bool = True) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results.
+
+ Args:
+ cls_score_list (list[Tensor]): Box scores from all scale
+ levels of a single image, each item has shape
+ (num_priors * num_classes, H, W).
+ bbox_pred_list (list[Tensor]): Box energies / deltas from
+ all scale levels of a single image, each item has shape
+ (num_priors * 4, H, W).
+ score_factor_list (list[Tensor]): Score factor from all scale
+ levels of a single image, each item has shape
+ (num_priors * 1, H, W).
+ mlvl_priors (list[Tensor]): Each element in the list is
+ the priors of a single level in feature pyramid. In all
+ anchor-based methods, it has shape (num_priors, 4). In
+ all anchor-free methods, it has shape (num_priors, 2)
+ when `with_stride=True`, otherwise it still has shape
+ (num_priors, 4).
+ img_meta (dict): Image meta info.
+ cfg (mmengine.Config): Test / postprocessing configuration,
+ if None, test_cfg would be used.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ with_nms (bool): If True, do nms before return boxes.
+ Defaults to True.
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+
+ cfg = self.test_cfg if cfg is None else cfg
+ cfg = copy.deepcopy(cfg)
+ nms_pre = cfg.get('nms_pre', -1)
+
+ mlvl_bbox_preds = []
+ mlvl_valid_priors = []
+ mlvl_scores = []
+ mlvl_labels = []
+
+ for level_idx, (cls_score, bbox_pred, score_factor, priors) in \
+ enumerate(zip(cls_score_list, bbox_pred_list,
+ score_factor_list, mlvl_priors)):
+
+ assert cls_score.size()[-2:] == bbox_pred.size()[-2:]
+
+ bbox_pred = bbox_pred * self.strides[level_idx]
+
+ dim = self.bbox_coder.encode_size
+ bbox_pred = bbox_pred.permute(1, 2, 0).reshape(-1, dim)
+ cls_score = cls_score.permute(1, 2,
+ 0).reshape(-1, self.cls_out_channels)
+ heatmap = cls_score.sigmoid()
+ score_thr = cfg.get('score_thr', 0)
+
+ candidate_inds = heatmap > score_thr # 0.05
+ pre_nms_top_n = candidate_inds.sum() # N
+ pre_nms_top_n = pre_nms_top_n.clamp(max=nms_pre) # N
+
+ heatmap = heatmap[candidate_inds] # n
+
+ candidate_nonzeros = candidate_inds.nonzero() # n
+ box_loc = candidate_nonzeros[:, 0] # n
+ labels = candidate_nonzeros[:, 1] # n
+
+ bbox_pred = bbox_pred[box_loc] # n x 4
+ per_grids = priors[box_loc] # n x 2
+
+ if candidate_inds.sum().item() > pre_nms_top_n.item():
+ heatmap, top_k_indices = \
+ heatmap.topk(pre_nms_top_n, sorted=False)
+ labels = labels[top_k_indices]
+ bbox_pred = bbox_pred[top_k_indices]
+ per_grids = per_grids[top_k_indices]
+
+ bboxes = torch.stack([
+ per_grids[:, 0] - bbox_pred[:, 0],
+ per_grids[:, 1] - bbox_pred[:, 1],
+ per_grids[:, 0] + bbox_pred[:, 2],
+ per_grids[:, 1] + bbox_pred[:, 3],
+ ],
+ dim=1) # n x 4
+
+ # avoid invalid boxes in RoI heads
+ bboxes[:, 2] = torch.max(bboxes[:, 2], bboxes[:, 0] + 0.01)
+ bboxes[:, 3] = torch.max(bboxes[:, 3], bboxes[:, 1] + 0.01)
+
+ # bboxes = self.bbox_coder.decode(per_grids, bbox_pred)
+ # # avoid invalid boxes in RoI heads
+ # bboxes[:, 2] = torch.max(bboxes[:, 2], bboxes[:, 0] + 0.01)
+ # bboxes[:, 3] = torch.max(bboxes[:, 3], bboxes[:, 1] + 0.01)
+
+ mlvl_bbox_preds.append(bboxes)
+ mlvl_valid_priors.append(priors)
+ mlvl_scores.append(torch.sqrt(heatmap))
+ mlvl_labels.append(labels)
+
+ results = InstanceData()
+ results.bboxes = torch.cat(mlvl_bbox_preds)
+ results.scores = torch.cat(mlvl_scores)
+ results.labels = torch.cat(mlvl_labels)
+
+ return self._bbox_post_process(
+ results=results,
+ cfg=cfg,
+ rescale=rescale,
+ with_nms=with_nms,
+ img_meta=img_meta)
+
+ def _get_label_inds(self, batch_gt_instances, batch_img_metas,
+ shapes_per_level):
+ '''
+ Inputs:
+ batch_gt_instances: [n_i], sum n_i = N
+ shapes_per_level: L x 2 [(h_l, w_l)]_L
+ Returns:
+ pos_inds: N'
+ labels: N'
+ '''
+ pos_inds = []
+ L = len(self.strides)
+ B = len(batch_gt_instances)
+ shapes_per_level = shapes_per_level.long()
+ loc_per_level = (shapes_per_level[:, 0] *
+ shapes_per_level[:, 1]).long() # L
+ level_bases = []
+ s = 0
+ for i in range(L):
+ level_bases.append(s)
+ s = s + B * loc_per_level[i]
+ level_bases = shapes_per_level.new_tensor(level_bases).long() # L
+ strides_default = shapes_per_level.new_tensor(
+ self.strides).float() # L
+ for im_i in range(B):
+ targets_per_im = batch_gt_instances[im_i]
+ if hasattr(targets_per_im, 'bboxes'):
+ bboxes = targets_per_im.bboxes # n x 4
+ else:
+ bboxes = targets_per_im.labels.new_tensor(
+ [], dtype=torch.float).reshape(-1, 4)
+ n = bboxes.shape[0]
+ centers = ((bboxes[:, [0, 1]] + bboxes[:, [2, 3]]) / 2) # n x 2
+ centers = centers.view(n, 1, 2).expand(n, L, 2).contiguous()
+ if self.not_clamp_box:
+ h, w = batch_img_metas[im_i]._image_size
+ centers[:, :, 0].clamp_(min=0).clamp_(max=w - 1)
+ centers[:, :, 1].clamp_(min=0).clamp_(max=h - 1)
+ strides = strides_default.view(1, L, 1).expand(n, L, 2)
+ centers_inds = (centers / strides).long() # n x L x 2
+ Ws = shapes_per_level[:, 1].view(1, L).expand(n, L)
+ pos_ind = level_bases.view(1, L).expand(n, L) \
+ + im_i * loc_per_level.view(1, L).expand(n, L) \
+ + centers_inds[:, :, 1] * Ws + centers_inds[:, :, 0] # n x L
+ is_cared_in_the_level = self.assign_fpn_level(bboxes)
+ pos_ind = pos_ind[is_cared_in_the_level].view(-1)
+
+ pos_inds.append(pos_ind) # n'
+ pos_inds = torch.cat(pos_inds, dim=0).long()
+ return pos_inds # N, N
+
+ def assign_fpn_level(self, boxes):
+ '''
+ Inputs:
+ boxes: n x 4
+ size_ranges: L x 2
+ Return:
+ is_cared_in_the_level: n x L
+ '''
+ size_ranges = boxes.new_tensor(self.regress_ranges).view(
+ len(self.regress_ranges), 2) # L x 2
+ crit = ((boxes[:, 2:] - boxes[:, :2])**2).sum(dim=1)**0.5 / 2 # n
+ n, L = crit.shape[0], size_ranges.shape[0]
+ crit = crit.view(n, 1).expand(n, L)
+ size_ranges_expand = size_ranges.view(1, L, 2).expand(n, L, 2)
+ is_cared_in_the_level = (crit >= size_ranges_expand[:, :, 0]) & \
+ (crit <= size_ranges_expand[:, :, 1])
+ return is_cared_in_the_level
+
+ def _get_targets_single(self, gt_instances: InstanceData, points: Tensor,
+ regress_ranges: Tensor,
+ strides: Tensor) -> Tuple[Tensor, Tensor]:
+ """Compute classification and bbox targets for a single image."""
+ num_points = points.size(0)
+ num_gts = len(gt_instances)
+ gt_labels = gt_instances.labels
+
+ if not hasattr(gt_instances, 'bboxes'):
+ gt_bboxes = gt_labels.new_tensor([], dtype=torch.float)
+ else:
+ gt_bboxes = gt_instances.bboxes
+
+ if not hasattr(gt_instances, 'bboxes') or num_gts == 0:
+ return gt_labels.new_full((num_points,
+ self.num_classes),
+ self.num_classes,
+ dtype=torch.float), \
+ gt_bboxes.new_full((num_points, 4), -1)
+
+ # Calculate the regression tblr target corresponding to all points
+ points = points[:, None].expand(num_points, num_gts, 2)
+ gt_bboxes = gt_bboxes[None].expand(num_points, num_gts, 4)
+ strides = strides[:, None, None].expand(num_points, num_gts, 2)
+
+ bbox_target = bbox2distance(points, gt_bboxes) # M x N x 4
+
+ # condition1: inside a gt bbox
+ inside_gt_bbox_mask = bbox_target.min(dim=2)[0] > 0 # M x N
+
+ # condition2: Calculate the nearest points from
+ # the upper, lower, left and right ranges from
+ # the center of the gt bbox
+ centers = ((gt_bboxes[..., [0, 1]] + gt_bboxes[..., [2, 3]]) / 2)
+ centers_discret = ((centers / strides).int() * strides).float() + \
+ strides / 2
+
+ centers_discret_dist = points - centers_discret
+ dist_x = centers_discret_dist[..., 0].abs()
+ dist_y = centers_discret_dist[..., 1].abs()
+ inside_gt_center3x3_mask = (dist_x <= strides[..., 0]) & \
+ (dist_y <= strides[..., 0])
+
+ # condition3: limit the regression range for each location
+ bbox_target_wh = bbox_target[..., :2] + bbox_target[..., 2:]
+ crit = (bbox_target_wh**2).sum(dim=2)**0.5 / 2
+ inside_fpn_level_mask = (crit >= regress_ranges[:, [0]]) & \
+ (crit <= regress_ranges[:, [1]])
+ bbox_target_mask = inside_gt_bbox_mask & \
+ inside_gt_center3x3_mask & \
+ inside_fpn_level_mask
+
+ # Calculate the distance weight map
+ gt_center_peak_mask = ((centers_discret_dist**2).sum(dim=2) == 0)
+ weighted_dist = ((points - centers)**2).sum(dim=2) # M x N
+ weighted_dist[gt_center_peak_mask] = 0
+
+ areas = (gt_bboxes[..., 2] - gt_bboxes[..., 0]) * (
+ gt_bboxes[..., 3] - gt_bboxes[..., 1])
+ radius = self.delta**2 * 2 * areas
+ radius = torch.clamp(radius, min=self.hm_min_radius**2)
+ weighted_dist = weighted_dist / radius
+
+ # Calculate bbox_target
+ bbox_weighted_dist = weighted_dist.clone()
+ bbox_weighted_dist[bbox_target_mask == 0] = INF * 1.0
+ min_dist, min_inds = bbox_weighted_dist.min(dim=1)
+ bbox_target = bbox_target[range(len(bbox_target)),
+ min_inds] # M x N x 4 --> M x 4
+ bbox_target[min_dist == INF] = -INF
+
+ # Convert to feature map scale
+ bbox_target /= strides[:, 0, :].repeat(1, 2)
+
+ # Calculate cls_target
+ cls_target = self._create_heatmaps_from_dist(weighted_dist, gt_labels)
+
+ return cls_target, bbox_target
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/detic.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/detic.py
new file mode 100644
index 0000000000000000000000000000000000000000..7028690ace9b1fa902bfe2b4ea2536b7afdc26be
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/detic.py
@@ -0,0 +1,274 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import copy
+from typing import List, Union
+
+import numpy as np
+import torch
+import torch.nn as nn
+import torch.nn.functional as F
+from mmengine.logging import print_log
+from torch import Tensor
+
+from mmdet.datasets import LVISV1Dataset
+from mmdet.models.detectors.cascade_rcnn import CascadeRCNN
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+
+
+class CLIPTextEncoder(nn.Module):
+
+ def __init__(self, model_name='ViT-B/32'):
+ super().__init__()
+ import clip
+ from clip.simple_tokenizer import SimpleTokenizer
+ self.tokenizer = SimpleTokenizer()
+ pretrained_model, _ = clip.load(model_name, device='cpu')
+ self.clip = pretrained_model
+
+ @property
+ def device(self):
+ return self.clip.device
+
+ @property
+ def dtype(self):
+ return self.clip.dtype
+
+ def tokenize(self,
+ texts: Union[str, List[str]],
+ context_length: int = 77) -> torch.LongTensor:
+ if isinstance(texts, str):
+ texts = [texts]
+
+ sot_token = self.tokenizer.encoder['<|startoftext|>']
+ eot_token = self.tokenizer.encoder['<|endoftext|>']
+ all_tokens = [[sot_token] + self.tokenizer.encode(text) + [eot_token]
+ for text in texts]
+ result = torch.zeros(len(all_tokens), context_length, dtype=torch.long)
+
+ for i, tokens in enumerate(all_tokens):
+ if len(tokens) > context_length:
+ st = torch.randint(len(tokens) - context_length + 1,
+ (1, ))[0].item()
+ tokens = tokens[st:st + context_length]
+ result[i, :len(tokens)] = torch.tensor(tokens)
+
+ return result
+
+ def forward(self, text):
+ text = self.tokenize(text)
+ text_features = self.clip.encode_text(text)
+ return text_features
+
+
+def get_class_weight(original_caption, prompt_prefix='a '):
+ if isinstance(original_caption, str):
+ if original_caption == 'coco':
+ from mmdet.datasets import CocoDataset
+ class_names = CocoDataset.METAINFO['classes']
+ elif original_caption == 'cityscapes':
+ from mmdet.datasets import CityscapesDataset
+ class_names = CityscapesDataset.METAINFO['classes']
+ elif original_caption == 'voc':
+ from mmdet.datasets import VOCDataset
+ class_names = VOCDataset.METAINFO['classes']
+ elif original_caption == 'openimages':
+ from mmdet.datasets import OpenImagesDataset
+ class_names = OpenImagesDataset.METAINFO['classes']
+ elif original_caption == 'lvis':
+ from mmdet.datasets import LVISV1Dataset
+ class_names = LVISV1Dataset.METAINFO['classes']
+ else:
+ if not original_caption.endswith('.'):
+ original_caption = original_caption + ' . '
+ original_caption = original_caption.split(' . ')
+ class_names = list(filter(lambda x: len(x) > 0, original_caption))
+
+ # for test.py
+ else:
+ class_names = list(original_caption)
+
+ text_encoder = CLIPTextEncoder()
+ text_encoder.eval()
+ texts = [prompt_prefix + x for x in class_names]
+ print_log(f'Computing text embeddings for {len(class_names)} classes.')
+ embeddings = text_encoder(texts).detach().permute(1, 0).contiguous().cpu()
+ return class_names, embeddings
+
+
+def reset_cls_layer_weight(roi_head, weight):
+ if type(weight) == str:
+ print_log(f'Resetting cls_layer_weight from file: {weight}')
+ zs_weight = torch.tensor(
+ np.load(weight),
+ dtype=torch.float32).permute(1, 0).contiguous() # D x C
+ else:
+ zs_weight = weight
+ zs_weight = torch.cat(
+ [zs_weight, zs_weight.new_zeros(
+ (zs_weight.shape[0], 1))], dim=1) # D x (C + 1)
+ zs_weight = F.normalize(zs_weight, p=2, dim=0)
+ zs_weight = zs_weight.to('cuda')
+ num_classes = zs_weight.shape[-1]
+
+ for bbox_head in roi_head.bbox_head:
+ bbox_head.num_classes = num_classes
+ del bbox_head.fc_cls.zs_weight
+ bbox_head.fc_cls.zs_weight = zs_weight
+
+
+@MODELS.register_module()
+class Detic(CascadeRCNN):
+
+ def __init__(self,
+ with_image_labels: bool = False,
+ sync_caption_batch: bool = False,
+ fp16: bool = False,
+ roi_head_name: str = '',
+ cap_batch_ratio: int = 4,
+ with_caption: bool = False,
+ dynamic_classifier: bool = False,
+ **kwargs) -> None:
+ super().__init__(**kwargs)
+
+ self._entities = LVISV1Dataset.METAINFO['classes']
+ self._text_prompts = None
+ # Turn on co-training with classification data
+ self.with_image_labels = with_image_labels
+ # Caption losses
+ self.with_caption = with_caption
+ # synchronize across GPUs to enlarge # "classes"
+ self.sync_caption_batch = sync_caption_batch
+ # Ratio between detection data and caption data
+ self.cap_batch_ratio = cap_batch_ratio
+ self.fp16 = fp16
+ self.roi_head_name = roi_head_name
+ # dynamic class sampling when training with 21K classes,
+ # Federated loss is enabled when DYNAMIC_CLASSIFIER is on
+ self.dynamic_classifier = dynamic_classifier
+ self.return_proposal = False
+ if self.dynamic_classifier:
+ self.freq_weight = kwargs.pop('freq_weight')
+ self.num_classes = kwargs.pop('num_classes')
+ self.num_sample_cats = kwargs.pop('num_sample_cats')
+
+ def loss(self, batch_inputs: Tensor,
+ batch_data_samples: SampleList) -> dict:
+ """Calculate losses from a batch of inputs and data samples.
+
+ Args:
+ batch_inputs (Tensor): Input images of shape (N, C, H, W).
+ These should usually be mean centered and std scaled.
+ batch_data_samples (List[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict: A dictionary of loss components
+ """
+
+ x = self.extract_feat(batch_inputs)
+ losses = dict()
+
+ # RPN forward and loss
+ if self.with_rpn:
+ proposal_cfg = self.train_cfg.get('rpn_proposal',
+ self.test_cfg.rpn)
+ rpn_data_samples = copy.deepcopy(batch_data_samples)
+ # set cat_id of gt_labels to 0 in RPN
+ for data_sample in rpn_data_samples:
+ data_sample.gt_instances.labels = \
+ torch.zeros_like(data_sample.gt_instances.labels)
+
+ rpn_losses, rpn_results_list = self.rpn_head.loss_and_predict(
+ x, rpn_data_samples, proposal_cfg=proposal_cfg)
+
+ # avoid get same name with roi_head loss
+ keys = rpn_losses.keys()
+ for key in list(keys):
+ if 'loss' in key and 'rpn' not in key:
+ rpn_losses[f'rpn_{key}'] = rpn_losses.pop(key)
+ losses.update(rpn_losses)
+ # if not hasattr(batch_data_samples[0].gt_instances, 'bboxes'):
+ # losses.update({k: v * 0 for k, v in rpn_losses.items()})
+ # else:
+ # losses.update(rpn_losses)
+ else:
+ assert batch_data_samples[0].get('proposals', None) is not None
+ # use pre-defined proposals in InstanceData for the second stage
+ # to extract ROI features.
+ rpn_results_list = [
+ data_sample.proposals for data_sample in batch_data_samples
+ ]
+
+ roi_losses = self.roi_head.loss(x, rpn_results_list,
+ batch_data_samples)
+
+ losses.update(roi_losses)
+
+ return losses
+
+ def predict(self,
+ batch_inputs: Tensor,
+ batch_data_samples: SampleList,
+ rescale: bool = True) -> SampleList:
+ """Predict results from a batch of inputs and data samples with post-
+ processing.
+
+ Args:
+ batch_inputs (Tensor): Inputs with shape (N, C, H, W).
+ batch_data_samples (List[:obj:`DetDataSample`]): The Data
+ Samples. It usually includes information such as
+ `gt_instance`, `gt_panoptic_seg` and `gt_sem_seg`.
+ rescale (bool): Whether to rescale the results.
+ Defaults to True.
+
+ Returns:
+ list[:obj:`DetDataSample`]: Return the detection results of the
+ input images. The returns value is DetDataSample,
+ which usually contain 'pred_instances'. And the
+ ``pred_instances`` usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+ """
+ # For single image inference
+ if 'custom_entities' in batch_data_samples[0]:
+ text_prompts = batch_data_samples[0].text
+ if text_prompts != self._text_prompts:
+ self._text_prompts = text_prompts
+ class_names, zs_weight = get_class_weight(text_prompts)
+ self._entities = class_names
+ reset_cls_layer_weight(self.roi_head, zs_weight)
+
+ assert self.with_bbox, 'Bbox head must be implemented.'
+
+ x = self.extract_feat(batch_inputs)
+
+ # If there are no pre-defined proposals, use RPN to get proposals
+ if batch_data_samples[0].get('proposals', None) is None:
+ rpn_results_list = self.rpn_head.predict(
+ x, batch_data_samples, rescale=False)
+ else:
+ rpn_results_list = [
+ data_sample.proposals for data_sample in batch_data_samples
+ ]
+
+ results_list = self.roi_head.predict(
+ x, rpn_results_list, batch_data_samples, rescale=rescale)
+
+ for data_sample, pred_instances in zip(batch_data_samples,
+ results_list):
+ if len(pred_instances) > 0:
+ label_names = []
+ for labels in pred_instances.labels:
+ label_names.append(self._entities[labels])
+ # for visualization
+ pred_instances.label_names = label_names
+ data_sample.pred_instances = pred_instances
+
+ return batch_data_samples
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/detic_bbox_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/detic_bbox_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..8779494ba137e03267b6f3d191da46587334c6e5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/detic_bbox_head.py
@@ -0,0 +1,434 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+import json
+from typing import List, Optional
+
+import torch
+from mmengine.config import ConfigDict
+from mmengine.structures import InstanceData
+from torch import Tensor
+from torch.nn import functional as F
+
+from mmdet.models.layers import multiclass_nms
+from mmdet.models.losses import accuracy
+from mmdet.models.roi_heads.bbox_heads import Shared2FCBBoxHead
+from mmdet.models.utils import empty_instances
+from mmdet.registry import MODELS
+from mmdet.structures.bbox import get_box_tensor, scale_boxes
+from mmdet.utils import ConfigType, InstanceList
+
+
+def load_class_freq(path='datasets/metadata/lvis_v1_train_cat_info.json',
+ freq_weight=0.5):
+ cat_info = json.load(open(path, 'r'))
+ cat_info = torch.tensor(
+ [c['image_count'] for c in sorted(cat_info, key=lambda x: x['id'])])
+ freq_weight = cat_info.float()**freq_weight
+ return freq_weight
+
+
+def get_fed_loss_inds(labels, num_sample_cats, C, weight=None):
+
+ appeared = torch.unique(labels) # C'
+ prob = appeared.new_ones(C + 1).float()
+ prob[-1] = 0
+ if len(appeared) < num_sample_cats:
+ if weight is not None:
+ prob[:C] = weight.float().clone()
+ prob[appeared] = 0
+ more_appeared = torch.multinomial(
+ prob, num_sample_cats - len(appeared), replacement=False)
+ appeared = torch.cat([appeared, more_appeared])
+ return appeared
+
+
+@MODELS.register_module()
+class DeticBBoxHead(Shared2FCBBoxHead):
+
+ def __init__(self,
+ image_loss_weight: float = 0.1,
+ use_fed_loss: bool = False,
+ cat_freq_path: str = '',
+ fed_loss_freq_weight: float = 0.5,
+ fed_loss_num_cat: int = 50,
+ cls_predictor_cfg: ConfigType = dict(
+ type='ZeroShotClassifier'),
+ *args,
+ **kwargs) -> None:
+ super().__init__(*args, **kwargs)
+ # reconstruct fc_cls and fc_reg since input channels are changed
+ assert self.with_cls
+
+ self.cls_predictor_cfg = cls_predictor_cfg
+ cls_channels = self.num_classes
+ self.cls_predictor_cfg.update(
+ in_features=self.cls_last_dim, out_features=cls_channels)
+ self.fc_cls = MODELS.build(self.cls_predictor_cfg)
+
+ self.init_cfg += [
+ dict(type='Caffe2Xavier', override=dict(name='reg_fcs'))
+ ]
+
+ self.image_loss_weight = image_loss_weight
+ self.use_fed_loss = use_fed_loss
+ self.cat_freq_path = cat_freq_path
+ self.fed_loss_freq_weight = fed_loss_freq_weight
+ self.fed_loss_num_cat = fed_loss_num_cat
+
+ if self.use_fed_loss:
+ freq_weight = load_class_freq(cat_freq_path, fed_loss_freq_weight)
+ self.register_buffer('freq_weight', freq_weight)
+ else:
+ self.freq_weight = None
+
+ def _predict_by_feat_single(
+ self,
+ roi: Tensor,
+ cls_score: Tensor,
+ bbox_pred: Tensor,
+ img_meta: dict,
+ rescale: bool = False,
+ rcnn_test_cfg: Optional[ConfigDict] = None) -> InstanceData:
+ """Transform a single image's features extracted from the head into
+ bbox results.
+
+ Args:
+ roi (Tensor): Boxes to be transformed. Has shape (num_boxes, 5).
+ last dimension 5 arrange as (batch_index, x1, y1, x2, y2).
+ cls_score (Tensor): Box scores, has shape
+ (num_boxes, num_classes + 1).
+ bbox_pred (Tensor): Box energies / deltas.
+ has shape (num_boxes, num_classes * 4).
+ img_meta (dict): image information.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+ rcnn_test_cfg (obj:`ConfigDict`): `test_cfg` of Bbox Head.
+ Defaults to None
+
+ Returns:
+ :obj:`InstanceData`: Detection results of each image\
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ results = InstanceData()
+ if roi.shape[0] == 0:
+ return empty_instances([img_meta],
+ roi.device,
+ task_type='bbox',
+ instance_results=[results],
+ box_type=self.predict_box_type,
+ use_box_type=False,
+ num_classes=self.num_classes,
+ score_per_cls=rcnn_test_cfg is None)[0]
+ scores = cls_score
+ img_shape = img_meta['img_shape']
+ num_rois = roi.size(0)
+
+ num_classes = 1 if self.reg_class_agnostic else self.num_classes
+ roi = roi.repeat_interleave(num_classes, dim=0)
+ bbox_pred = bbox_pred.view(-1, self.bbox_coder.encode_size)
+ bboxes = self.bbox_coder.decode(
+ roi[..., 1:], bbox_pred, max_shape=img_shape)
+
+ if rescale and bboxes.size(0) > 0:
+ assert img_meta.get('scale_factor') is not None
+ scale_factor = [1 / s for s in img_meta['scale_factor']]
+ bboxes = scale_boxes(bboxes, scale_factor)
+
+ # Get the inside tensor when `bboxes` is a box type
+ bboxes = get_box_tensor(bboxes)
+ box_dim = bboxes.size(-1)
+ bboxes = bboxes.view(num_rois, -1)
+
+ if rcnn_test_cfg is None:
+ # This means that it is aug test.
+ # It needs to return the raw results without nms.
+ results.bboxes = bboxes
+ results.scores = scores
+ else:
+ det_bboxes, det_labels = multiclass_nms(
+ bboxes,
+ scores,
+ rcnn_test_cfg.score_thr,
+ rcnn_test_cfg.nms,
+ rcnn_test_cfg.max_per_img,
+ box_dim=box_dim)
+ results.bboxes = det_bboxes[:, :-1]
+ results.scores = det_bboxes[:, -1]
+ results.labels = det_labels
+ return results
+
+ def loss(self,
+ cls_score: Tensor,
+ bbox_pred: Tensor,
+ rois: Tensor,
+ labels: Tensor,
+ label_weights: Tensor,
+ bbox_targets: Tensor,
+ bbox_weights: Tensor,
+ reduction_override: Optional[str] = None) -> dict:
+ """Calculate the loss based on the network predictions and targets.
+
+ Args:
+ cls_score (Tensor): Classification prediction
+ results of all class, has shape
+ (batch_size * num_proposals_single_image, num_classes)
+ bbox_pred (Tensor): Regression prediction results,
+ has shape
+ (batch_size * num_proposals_single_image, 4), the last
+ dimension 4 represents [tl_x, tl_y, br_x, br_y].
+ rois (Tensor): RoIs with the shape
+ (batch_size * num_proposals_single_image, 5) where the first
+ column indicates batch id of each RoI.
+ labels (Tensor): Gt_labels for all proposals in a batch, has
+ shape (batch_size * num_proposals_single_image, ).
+ label_weights (Tensor): Labels_weights for all proposals in a
+ batch, has shape (batch_size * num_proposals_single_image, ).
+ bbox_targets (Tensor): Regression target for all proposals in a
+ batch, has shape (batch_size * num_proposals_single_image, 4),
+ the last dimension 4 represents [tl_x, tl_y, br_x, br_y].
+ bbox_weights (Tensor): Regression weights for all proposals in a
+ batch, has shape (batch_size * num_proposals_single_image, 4).
+ reduction_override (str, optional): The reduction
+ method used to override the original reduction
+ method of the loss. Options are "none",
+ "mean" and "sum". Defaults to None,
+
+ Returns:
+ dict: A dictionary of loss.
+ """
+
+ losses = dict()
+
+ if cls_score is not None:
+
+ if cls_score.numel() > 0:
+ loss_cls_ = self.sigmoid_cross_entropy_loss(cls_score, labels)
+ if isinstance(loss_cls_, dict):
+ losses.update(loss_cls_)
+ else:
+ losses['loss_cls'] = loss_cls_
+ if self.custom_activation:
+ acc_ = self.loss_cls.get_accuracy(cls_score, labels)
+ losses.update(acc_)
+ else:
+ losses['acc'] = accuracy(cls_score, labels)
+ if bbox_pred is not None:
+ bg_class_ind = self.num_classes
+ # 0~self.num_classes-1 are FG, self.num_classes is BG
+ pos_inds = (labels >= 0) & (labels < bg_class_ind)
+ # do not perform bounding box regression for BG anymore.
+ if pos_inds.any():
+ if self.reg_decoded_bbox:
+ # When the regression loss (e.g. `IouLoss`,
+ # `GIouLoss`, `DIouLoss`) is applied directly on
+ # the decoded bounding boxes, it decodes the
+ # already encoded coordinates to absolute format.
+ bbox_pred = self.bbox_coder.decode(rois[:, 1:], bbox_pred)
+ bbox_pred = get_box_tensor(bbox_pred)
+ if self.reg_class_agnostic:
+ pos_bbox_pred = bbox_pred.view(
+ bbox_pred.size(0), -1)[pos_inds.type(torch.bool)]
+ else:
+ pos_bbox_pred = bbox_pred.view(
+ bbox_pred.size(0), self.num_classes,
+ -1)[pos_inds.type(torch.bool),
+ labels[pos_inds.type(torch.bool)]]
+
+ losses['loss_bbox'] = self.loss_bbox(
+ pos_bbox_pred,
+ bbox_targets[pos_inds.type(torch.bool)],
+ bbox_weights[pos_inds.type(torch.bool)],
+ avg_factor=bbox_targets.size(0),
+ reduction_override=reduction_override)
+ else:
+ losses['loss_bbox'] = bbox_pred[pos_inds].sum()
+ return losses
+
+ def sigmoid_cross_entropy_loss(self, cls_score, labels):
+ if cls_score.numel() == 0:
+ return cls_score.new_zeros(
+ [1])[0] # This is more robust than .sum() * 0.
+ B = cls_score.shape[0]
+ C = cls_score.shape[1] - 1
+
+ target = cls_score.new_zeros(B, C + 1)
+ target[range(len(labels)), labels] = 1 # B x (C + 1)
+ target = target[:, :C] # B x C
+
+ weight = 1
+ if self.use_fed_loss and (self.freq_weight is not None): # fedloss
+ appeared = get_fed_loss_inds(
+ labels,
+ num_sample_cats=self.fed_loss_num_cat,
+ C=C,
+ weight=self.freq_weight)
+ appeared_mask = appeared.new_zeros(C + 1)
+ appeared_mask[appeared] = 1 # C + 1
+ appeared_mask = appeared_mask[:C]
+ fed_w = appeared_mask.view(1, C).expand(B, C)
+ weight = weight * fed_w.float()
+ # if self.ignore_zero_cats and (self.freq_weight is not None):
+ # w = (self.freq_weight.view(-1) > 1e-4).float()
+ # weight = weight * w.view(1, C).expand(B, C)
+ # # import pdb; pdb.set_trace()
+
+ cls_loss = F.binary_cross_entropy_with_logits(
+ cls_score[:, :-1], target, reduction='none') # B x C
+ loss = torch.sum(cls_loss * weight) / B
+ return loss
+
+ def image_label_losses(self, cls_score, sampling_results, image_labels):
+ '''
+ Inputs:
+ cls_score: N x (C + 1)
+ image_labels B x 1
+ '''
+ num_inst_per_image = [
+ len(pred_instances) for pred_instances in sampling_results
+ ]
+ cls_score = cls_score.split(
+ num_inst_per_image, dim=0) # B x n x (C + 1)
+ B = len(cls_score)
+ loss = cls_score[0].new_zeros([1])[0]
+ for (score, labels, pred_instances) in zip(cls_score, image_labels,
+ sampling_results):
+ if score.shape[0] == 0:
+ loss += score.new_zeros([1])[0]
+ continue
+ # find out max-size idx
+ bboxes = pred_instances.bboxes
+ areas = (bboxes[:, 2] - bboxes[:, 0]) * (
+ bboxes[:, 3] - bboxes[:, 1])
+ idx = areas[:-1].argmax().item() if len(areas) > 1 else 0
+
+ for label in labels:
+ target = score.new_zeros(score.shape[1])
+ target[label] = 1
+ loss_i = F.binary_cross_entropy_with_logits(
+ score[idx], target, reduction='sum')
+ loss += loss_i / len(labels)
+ loss = loss / B
+
+ return loss * self.image_loss_weight
+
+ def refine_bboxes(self, bbox_results: dict,
+ batch_img_metas: List[dict]) -> InstanceList:
+ """Refine bboxes during training.
+
+ Args:
+ bbox_results (dict): Usually is a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `rois` (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+ - `bbox_targets` (tuple): Ground truth for proposals in a
+ single image. Containing the following list of Tensors:
+ (labels, label_weights, bbox_targets, bbox_weights)
+ batch_img_metas (List[dict]): List of image information.
+
+ Returns:
+ list[:obj:`InstanceData`]: Refined bboxes of each image.
+
+ Example:
+ >>> # xdoctest: +REQUIRES(module:kwarray)
+ >>> import numpy as np
+ >>> from mmdet.models.task_modules.samplers.
+ ... sampling_result import random_boxes
+ >>> from mmdet.models.task_modules.samplers import SamplingResult
+ >>> self = BBoxHead(reg_class_agnostic=True)
+ >>> n_roi = 2
+ >>> n_img = 4
+ >>> scale = 512
+ >>> rng = np.random.RandomState(0)
+ ... batch_img_metas = [{'img_shape': (scale, scale)}
+ >>> for _ in range(n_img)]
+ >>> sampling_results = [SamplingResult.random(rng=10)
+ ... for _ in range(n_img)]
+ >>> # Create rois in the expected format
+ >>> roi_boxes = random_boxes(n_roi, scale=scale, rng=rng)
+ >>> img_ids = torch.randint(0, n_img, (n_roi,))
+ >>> img_ids = img_ids.float()
+ >>> rois = torch.cat([img_ids[:, None], roi_boxes], dim=1)
+ >>> # Create other args
+ >>> labels = torch.randint(0, 81, (scale,)).long()
+ >>> bbox_preds = random_boxes(n_roi, scale=scale, rng=rng)
+ >>> cls_score = torch.randn((scale, 81))
+ ... # For each image, pretend random positive boxes are gts
+ >>> bbox_targets = (labels, None, None, None)
+ ... bbox_results = dict(rois=rois, bbox_pred=bbox_preds,
+ ... cls_score=cls_score,
+ ... bbox_targets=bbox_targets)
+ >>> bboxes_list = self.refine_bboxes(sampling_results,
+ ... bbox_results,
+ ... batch_img_metas)
+ >>> print(bboxes_list)
+ """
+ # bbox_targets is a tuple
+ cls_scores = bbox_results['cls_score']
+ rois = bbox_results['rois']
+ bbox_preds = bbox_results['bbox_pred']
+ if self.custom_activation:
+ # TODO: Create a SeasawBBoxHead to simplified logic in BBoxHead
+ cls_scores = self.loss_cls.get_activation(cls_scores)
+ if cls_scores.numel() == 0:
+ return None
+ if cls_scores.shape[-1] == self.num_classes + 1:
+ # remove background class
+ cls_scores = cls_scores[:, :-1]
+ elif cls_scores.shape[-1] != self.num_classes:
+ raise ValueError('The last dim of `cls_scores` should equal to '
+ '`num_classes` or `num_classes + 1`,'
+ f'but got {cls_scores.shape[-1]}.')
+
+ img_ids = rois[:, 0].long().unique(sorted=True)
+ assert img_ids.numel() <= len(batch_img_metas)
+
+ results_list = []
+ for i in range(len(batch_img_metas)):
+ inds = torch.nonzero(
+ rois[:, 0] == i, as_tuple=False).squeeze(dim=1)
+
+ bboxes_ = rois[inds, 1:]
+ bbox_pred_ = bbox_preds[inds]
+ img_meta_ = batch_img_metas[i]
+
+ bboxes = self.regress(bboxes_, bbox_pred_, img_meta_)
+
+ # don't filter gt bboxes like D2
+ results = InstanceData(bboxes=bboxes)
+ results_list.append(results)
+
+ return results_list
+
+ def regress(self, priors: Tensor, bbox_pred: Tensor,
+ img_meta: dict) -> Tensor:
+ """Regress the bbox for the predicted class. Used in Cascade R-CNN.
+
+ Args:
+ priors (Tensor): Priors from `rpn_head` or last stage
+ `bbox_head`, has shape (num_proposals, 4).
+ label (Tensor): Only used when `self.reg_class_agnostic`
+ is False, has shape (num_proposals, ).
+ bbox_pred (Tensor): Regression prediction of
+ current stage `bbox_head`. When `self.reg_class_agnostic`
+ is False, it has shape (n, num_classes * 4), otherwise
+ it has shape (n, 4).
+ img_meta (dict): Image meta info.
+
+ Returns:
+ Tensor: Regressed bboxes, the same shape as input rois.
+ """
+ reg_dim = self.bbox_coder.encode_size
+ assert bbox_pred.size()[1] == reg_dim
+
+ max_shape = img_meta['img_shape']
+ regressed_bboxes = self.bbox_coder.decode(
+ priors, bbox_pred, max_shape=max_shape)
+ return regressed_bboxes
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/detic_roi_head.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/detic_roi_head.py
new file mode 100644
index 0000000000000000000000000000000000000000..35785cda7434ed95063a96f4a3dca8a1ddfd3b25
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/detic_roi_head.py
@@ -0,0 +1,440 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import List, Sequence, Tuple
+
+import torch
+from mmengine.structures import InstanceData
+from torch import Tensor
+
+from mmdet.models.roi_heads import CascadeRoIHead
+from mmdet.models.task_modules.samplers import SamplingResult
+from mmdet.models.test_time_augs import merge_aug_masks
+from mmdet.models.utils import empty_instances, unpack_gt_instances
+from mmdet.registry import MODELS
+from mmdet.structures import SampleList
+from mmdet.structures.bbox import bbox2roi, get_box_tensor
+from mmdet.utils import ConfigType, InstanceList, MultiConfig
+
+
+@MODELS.register_module()
+class DeticRoIHead(CascadeRoIHead):
+
+ def __init__(
+ self,
+ *,
+ mult_proposal_score: bool = False,
+ with_image_labels: bool = False,
+ add_image_box: bool = False,
+ image_box_size: float = 1.0,
+ ws_num_props: int = 128,
+ add_feature_to_prop: bool = False,
+ mask_weight: float = 1.0,
+ one_class_per_proposal: bool = False,
+ **kwargs,
+ ):
+ super().__init__(**kwargs)
+ self.mult_proposal_score = mult_proposal_score
+ self.with_image_labels = with_image_labels
+ self.add_image_box = add_image_box
+ self.image_box_size = image_box_size
+ self.ws_num_props = ws_num_props
+ self.add_feature_to_prop = add_feature_to_prop
+ self.mask_weight = mask_weight
+ self.one_class_per_proposal = one_class_per_proposal
+
+ def init_mask_head(self, mask_roi_extractor: MultiConfig,
+ mask_head: MultiConfig) -> None:
+ """Initialize mask head and mask roi extractor.
+
+ Args:
+ mask_head (dict): Config of mask in mask head.
+ mask_roi_extractor (:obj:`ConfigDict`, dict or list):
+ Config of mask roi extractor.
+ """
+ self.mask_head = MODELS.build(mask_head)
+
+ if mask_roi_extractor is not None:
+ self.share_roi_extractor = False
+ self.mask_roi_extractor = MODELS.build(mask_roi_extractor)
+ else:
+ self.share_roi_extractor = True
+ self.mask_roi_extractor = self.bbox_roi_extractor
+
+ def _refine_roi(self, x: Tuple[Tensor], rois: Tensor,
+ batch_img_metas: List[dict],
+ num_proposals_per_img: Sequence[int], **kwargs) -> tuple:
+ """Multi-stage refinement of RoI.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ rois (Tensor): shape (n, 5), [batch_ind, x1, y1, x2, y2]
+ batch_img_metas (list[dict]): List of image information.
+ num_proposals_per_img (sequence[int]): number of proposals
+ in each image.
+
+ Returns:
+ tuple:
+
+ - rois (Tensor): Refined RoI.
+ - cls_scores (list[Tensor]): Average predicted
+ cls score per image.
+ - bbox_preds (list[Tensor]): Bbox branch predictions
+ for the last stage of per image.
+ """
+ # "ms" in variable names means multi-stage
+ ms_scores = []
+ for stage in range(self.num_stages):
+ bbox_results = self._bbox_forward(
+ stage=stage, x=x, rois=rois, **kwargs)
+
+ # split batch bbox prediction back to each image
+ cls_scores = bbox_results['cls_score'].sigmoid()
+ bbox_preds = bbox_results['bbox_pred']
+
+ rois = rois.split(num_proposals_per_img, 0)
+ cls_scores = cls_scores.split(num_proposals_per_img, 0)
+ ms_scores.append(cls_scores)
+ bbox_preds = bbox_preds.split(num_proposals_per_img, 0)
+
+ if stage < self.num_stages - 1:
+ bbox_head = self.bbox_head[stage]
+ refine_rois_list = []
+ for i in range(len(batch_img_metas)):
+ if rois[i].shape[0] > 0:
+ bbox_label = cls_scores[i][:, :-1].argmax(dim=1)
+ # Refactor `bbox_head.regress_by_class` to only accept
+ # box tensor without img_idx concatenated.
+ refined_bboxes = bbox_head.regress_by_class(
+ rois[i][:, 1:], bbox_label, bbox_preds[i],
+ batch_img_metas[i])
+ refined_bboxes = get_box_tensor(refined_bboxes)
+ refined_rois = torch.cat(
+ [rois[i][:, [0]], refined_bboxes], dim=1)
+ refine_rois_list.append(refined_rois)
+ rois = torch.cat(refine_rois_list)
+ # ms_scores aligned
+ # average scores of each image by stages
+ cls_scores = [
+ sum([score[i] for score in ms_scores]) / float(len(ms_scores))
+ for i in range(len(batch_img_metas))
+ ] # aligned
+ return rois, cls_scores, bbox_preds
+
+ def predict_bbox(self,
+ x: Tuple[Tensor],
+ batch_img_metas: List[dict],
+ rpn_results_list: InstanceList,
+ rcnn_test_cfg: ConfigType,
+ rescale: bool = False,
+ **kwargs) -> InstanceList:
+ """Perform forward propagation of the bbox head and predict detection
+ results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Feature maps of all scale level.
+ batch_img_metas (list[dict]): List of image information.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ rcnn_test_cfg (obj:`ConfigDict`): `test_cfg` of R-CNN.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ """
+ proposals = [res.bboxes for res in rpn_results_list]
+ proposal_scores = [res.scores for res in rpn_results_list]
+ num_proposals_per_img = tuple(len(p) for p in proposals)
+ rois = bbox2roi(proposals)
+
+ if rois.shape[0] == 0:
+ return empty_instances(
+ batch_img_metas,
+ rois.device,
+ task_type='bbox',
+ box_type=self.bbox_head[-1].predict_box_type,
+ num_classes=self.bbox_head[-1].num_classes,
+ score_per_cls=rcnn_test_cfg is None)
+ # rois aligned
+ rois, cls_scores, bbox_preds = self._refine_roi(
+ x=x,
+ rois=rois,
+ batch_img_metas=batch_img_metas,
+ num_proposals_per_img=num_proposals_per_img,
+ **kwargs)
+
+ # score reweighting in centernet2
+ cls_scores = [(s * ps[:, None])**0.5
+ for s, ps in zip(cls_scores, proposal_scores)]
+ # # for demo
+ # cls_scores = [
+ # s * (s == s[:, :-1].max(dim=1)[0][:, None]).float()
+ # for s in cls_scores
+ # ]
+
+ # fast_rcnn_inference
+ results_list = self.bbox_head[-1].predict_by_feat(
+ rois=rois,
+ cls_scores=cls_scores,
+ bbox_preds=bbox_preds,
+ batch_img_metas=batch_img_metas,
+ rescale=rescale,
+ rcnn_test_cfg=rcnn_test_cfg)
+ return results_list
+
+ def _mask_forward(self, x: Tuple[Tensor], rois: Tensor) -> dict:
+ """Mask head forward function used in both training and testing.
+
+ Args:
+ stage (int): The current stage in Cascade RoI Head.
+ x (tuple[Tensor]): Tuple of multi-level img features.
+ rois (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `mask_preds` (Tensor): Mask prediction.
+ """
+ mask_feats = self.mask_roi_extractor(
+ x[:self.mask_roi_extractor.num_inputs], rois)
+ # do not support caffe_c4 model anymore
+ mask_preds = self.mask_head(mask_feats)
+
+ mask_results = dict(mask_preds=mask_preds)
+ return mask_results
+
+ def mask_loss(self, x, sampling_results: List[SamplingResult],
+ batch_gt_instances: InstanceList) -> dict:
+ """Run forward function and calculate loss for mask head in training.
+
+ Args:
+ x (tuple[Tensor]): Tuple of multi-level img features.
+ sampling_results (list["obj:`SamplingResult`]): Sampling results.
+ batch_gt_instances (list[:obj:`InstanceData`]): Batch of
+ gt_instance. It usually includes ``bboxes``, ``labels``, and
+ ``masks`` attributes.
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `mask_preds` (Tensor): Mask prediction.
+ - `loss_mask` (dict): A dictionary of mask loss components.
+ """
+ pos_rois = bbox2roi([res.pos_priors for res in sampling_results])
+ mask_results = self._mask_forward(x, pos_rois)
+
+ mask_loss_and_target = self.mask_head.loss_and_target(
+ mask_preds=mask_results['mask_preds'],
+ sampling_results=sampling_results,
+ batch_gt_instances=batch_gt_instances,
+ rcnn_train_cfg=self.train_cfg[-1])
+ mask_results.update(mask_loss_and_target)
+
+ return mask_results
+
+ def loss(self, x: Tuple[Tensor], rpn_results_list: InstanceList,
+ batch_data_samples: SampleList) -> dict:
+ """Perform forward propagation and loss calculation of the detection
+ roi on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): List of multi-level img features.
+ rpn_results_list (list[:obj:`InstanceData`]): List of region
+ proposals.
+ batch_data_samples (list[:obj:`DetDataSample`]): The batch
+ data samples. It usually includes information such
+ as `gt_instance` or `gt_panoptic_seg` or `gt_sem_seg`.
+
+ Returns:
+ dict[str, Tensor]: A dictionary of loss components
+ """
+ assert len(rpn_results_list) == len(batch_data_samples)
+ outputs = unpack_gt_instances(batch_data_samples)
+ batch_gt_instances, batch_gt_instances_ignore, batch_img_metas \
+ = outputs
+
+ num_imgs = len(batch_data_samples)
+ image_labels = [x.gt_instances.labels for x in batch_data_samples]
+ losses = dict()
+ results_list = rpn_results_list
+
+ for stage in range(self.num_stages):
+ self.current_stage = stage
+ stage_loss_weight = self.stage_loss_weights[stage]
+ if hasattr(batch_gt_instances[0], 'bboxes'):
+ # assign gts and sample proposals
+ sampling_results = []
+ if self.with_bbox or self.with_mask:
+ bbox_assigner = self.bbox_assigner[stage]
+ bbox_sampler = self.bbox_sampler[stage]
+
+ for i in range(num_imgs):
+ results = results_list[i]
+ # rename rpn_results.bboxes to rpn_results.priors
+ results.priors = results.pop('bboxes')
+
+ assign_result = bbox_assigner.assign(
+ results, batch_gt_instances[i],
+ batch_gt_instances_ignore[i])
+
+ sampling_result = bbox_sampler.sample(
+ assign_result,
+ results,
+ batch_gt_instances[i],
+ feats=[lvl_feat[i][None] for lvl_feat in x])
+
+ sampling_results.append(sampling_result)
+
+ # bbox head forward and loss
+ bbox_results = self.bbox_loss(stage, x, sampling_results)
+
+ for name, value in bbox_results['loss_bbox'].items():
+ losses[f's{stage}.{name}'] = (
+ value * stage_loss_weight if 'loss' in name else value)
+ losses[f's{stage}.image_loss'] = x[0].new_zeros([1])[0]
+
+ # mask head forward and loss
+ # D2 only forward stage.0
+ if self.with_mask and stage == 0:
+ mask_results = self.mask_loss(x, sampling_results,
+ batch_gt_instances)
+ for name, value in mask_results['loss_mask'].items():
+ losses[name] = (
+ value *
+ stage_loss_weight if 'loss' in name else value)
+
+ else:
+ # get ws_num_props pred_instances for each image
+ sampling_results = [
+ pred_instances[:self.ws_num_props]
+ for pred_instances in results_list
+ ]
+ for i, pred_instances in enumerate(sampling_results):
+ pred_instances.bboxes = pred_instances.bboxes.detach()
+ bbox_results = self.image_loss(stage, x, sampling_results,
+ image_labels)
+ losses[f's{stage}.image_loss'] = bbox_results['image_loss']
+
+ for name in ['loss_cls', 'loss_bbox']:
+ losses[f's{stage}.{name}'] = x[0].new_zeros([1])[0]
+ if stage == 0:
+ losses['loss_mask'] = x[0].new_zeros([1])[0]
+
+ # refine bboxes
+ if stage < self.num_stages - 1:
+ bbox_head = self.bbox_head[stage]
+ with torch.no_grad():
+ results_list = bbox_head.refine_bboxes(
+ bbox_results, batch_img_metas)
+ # Empty proposal
+ if results_list is None:
+ break
+
+ return losses
+
+ def image_loss(self, stage: int, x: Tuple[Tensor],
+ sampling_results: List[SamplingResult],
+ image_labels) -> dict:
+ """Run forward function and calculate loss for box head in training.
+
+ Args:
+ stage (int): The current stage in Cascade RoI Head.
+ x (tuple[Tensor]): List of multi-level img features.
+ sampling_results (list["obj:`SamplingResult`]): Sampling results.
+
+ Returns:
+ dict: Usually returns a dictionary with keys:
+
+ - `cls_score` (Tensor): Classification scores.
+ - `bbox_pred` (Tensor): Box energies / deltas.
+ - `bbox_feats` (Tensor): Extract bbox RoI features.
+ - `loss_bbox` (dict): A dictionary of bbox loss components.
+ - `rois` (Tensor): RoIs with the shape (n, 5) where the first
+ column indicates batch id of each RoI.
+ - `bbox_targets` (tuple): Ground truth for proposals in a
+ single image. Containing the following list of Tensors:
+ (labels, label_weights, bbox_targets, bbox_weights)
+ """
+ bbox_head = self.bbox_head[stage]
+ rois = bbox2roi([res.bboxes for res in sampling_results])
+ bbox_results = self._bbox_forward(stage, x, rois)
+ bbox_results.update(rois=rois)
+
+ image_loss = bbox_head.image_label_losses(
+ cls_score=bbox_results['cls_score'],
+ sampling_results=sampling_results,
+ image_labels=image_labels)
+ bbox_results.update(dict(image_loss=image_loss))
+
+ return bbox_results
+
+ def predict_mask(self,
+ x: Tuple[Tensor],
+ batch_img_metas: List[dict],
+ results_list: List[InstanceData],
+ rescale: bool = False) -> List[InstanceData]:
+ """Perform forward propagation of the mask head and predict detection
+ results on the features of the upstream network.
+
+ Args:
+ x (tuple[Tensor]): Feature maps of all scale level.
+ batch_img_metas (list[dict]): List of image information.
+ results_list (list[:obj:`InstanceData`]): Detection results of
+ each image.
+ rescale (bool): If True, return boxes in original image space.
+ Defaults to False.
+
+ Returns:
+ list[:obj:`InstanceData`]: Detection results of each image
+ after the post process.
+ Each item usually contains following keys.
+
+ - scores (Tensor): Classification scores, has a shape
+ (num_instance, )
+ - labels (Tensor): Labels of bboxes, has a shape
+ (num_instances, ).
+ - bboxes (Tensor): Has a shape (num_instances, 4),
+ the last dimension 4 arrange as (x1, y1, x2, y2).
+ - masks (Tensor): Has a shape (num_instances, H, W).
+ """
+ bboxes = [res.bboxes for res in results_list]
+ mask_rois = bbox2roi(bboxes)
+ if mask_rois.shape[0] == 0:
+ results_list = empty_instances(
+ batch_img_metas,
+ mask_rois.device,
+ task_type='mask',
+ instance_results=results_list,
+ mask_thr_binary=self.test_cfg.mask_thr_binary)
+ return results_list
+
+ num_mask_rois_per_img = [len(res) for res in results_list]
+ aug_masks = []
+ mask_results = self._mask_forward(x, mask_rois)
+ mask_preds = mask_results['mask_preds']
+ # split batch mask prediction back to each image
+ mask_preds = mask_preds.split(num_mask_rois_per_img, 0)
+ aug_masks.append([m.sigmoid().detach() for m in mask_preds])
+
+ merged_masks = []
+ for i in range(len(batch_img_metas)):
+ aug_mask = [mask[i] for mask in aug_masks]
+ merged_mask = merge_aug_masks(aug_mask, batch_img_metas[i])
+ merged_masks.append(merged_mask)
+ results_list = self.mask_head.predict_by_feat(
+ mask_preds=merged_masks,
+ results_list=results_list,
+ batch_img_metas=batch_img_metas,
+ rcnn_test_cfg=self.test_cfg,
+ rescale=rescale,
+ activate_map=True)
+ return results_list
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/heatmap_focal_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/heatmap_focal_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..021a5b22d91e3bb35b951a135f839245551c9b99
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/heatmap_focal_loss.py
@@ -0,0 +1,131 @@
+# Copyright (c) OpenMMLab. All rights reserved.
+from typing import Optional, Union
+
+import torch
+import torch.nn as nn
+from torch import Tensor
+
+from mmdet.registry import MODELS
+
+
+# support class-agnostic heatmap_focal_loss
+def heatmap_focal_loss_with_pos_inds(
+ pred: Tensor,
+ targets: Tensor,
+ pos_inds: Tensor,
+ alpha: float = 2.0,
+ beta: float = 4.0,
+ gamma: float = 4.0,
+ sigmoid_clamp: float = 1e-4,
+ ignore_high_fp: float = -1.0,
+ pos_weight: float = 1.0,
+ neg_weight: float = 1.0,
+ avg_factor: Optional[Union[int, float]] = None) -> Tensor:
+
+ pred = torch.clamp(
+ pred.sigmoid_(), min=sigmoid_clamp, max=1 - sigmoid_clamp)
+
+ neg_weights = torch.pow(1 - targets, beta)
+
+ pos_pred = pred[pos_inds]
+ pos_loss = torch.log(pos_pred) * torch.pow(1 - pos_pred, gamma)
+ neg_loss = torch.log(1 - pred) * torch.pow(pred, gamma) * neg_weights
+ if ignore_high_fp > 0:
+ not_high_fp = (pred < ignore_high_fp).float()
+ neg_loss = not_high_fp * neg_loss
+
+ pos_loss = -pos_loss.sum()
+ neg_loss = -neg_loss.sum()
+ if alpha >= 0:
+ pos_loss = alpha * pos_loss
+ neg_loss = (1 - alpha) * neg_loss
+
+ pos_loss = pos_weight * pos_loss / avg_factor
+ neg_loss = neg_weight * neg_loss / avg_factor
+
+ return pos_loss, neg_loss
+
+
+@MODELS.register_module()
+class HeatmapFocalLoss(nn.Module):
+ """GaussianFocalLoss is a variant of focal loss.
+
+ More details can be found in the `paper
+ `_
+ Code is modified from `kp_utils.py
+ `_ # noqa: E501
+ Please notice that the target in GaussianFocalLoss is a gaussian heatmap,
+ not 0/1 binary target.
+
+ Args:
+ alpha (float): Power of prediction.
+ gamma (float): Power of target for negative samples.
+ reduction (str): Options are "none", "mean" and "sum".
+ loss_weight (float): Loss weight of current loss.
+ pos_weight(float): Positive sample loss weight. Defaults to 1.0.
+ neg_weight(float): Negative sample loss weight. Defaults to 1.0.
+ """
+
+ def __init__(
+ self,
+ alpha: float = 2.0,
+ beta: float = 4.0,
+ gamma: float = 4.0,
+ sigmoid_clamp: float = 1e-4,
+ ignore_high_fp: float = -1.0,
+ loss_weight: float = 1.0,
+ pos_weight: float = 1.0,
+ neg_weight: float = 1.0,
+ ) -> None:
+ super().__init__()
+ self.alpha = alpha
+ self.beta = beta
+ self.gamma = gamma
+ self.sigmoid_clamp = sigmoid_clamp
+ self.ignore_high_fp = ignore_high_fp
+ self.loss_weight = loss_weight
+ self.pos_weight = pos_weight
+ self.neg_weight = neg_weight
+
+ def forward(self,
+ pred: Tensor,
+ target: Tensor,
+ pos_inds: Optional[Tensor] = None,
+ avg_factor: Optional[Union[int, float]] = None) -> Tensor:
+ """Forward function.
+
+ If you want to manually determine which positions are
+ positive samples, you can set the pos_index and pos_label
+ parameter. Currently, only the CenterNet update version uses
+ the parameter.
+
+ Args:
+ pred (torch.Tensor): The prediction. The shape is (N, num_classes).
+ target (torch.Tensor): The learning target of the prediction
+ in gaussian distribution. The shape is (N, num_classes).
+ pos_inds (torch.Tensor): The positive sample index.
+ Defaults to None.
+ pos_labels (torch.Tensor): The label corresponding to the positive
+ sample index. Defaults to None.
+ weight (torch.Tensor, optional): The weight of loss for each
+ prediction. Defaults to None.
+ avg_factor (int, float, optional): Average factor that is used to
+ average the loss. Defaults to None.
+ reduction_override (str, optional): The reduction method used to
+ override the original reduction method of the loss.
+ Defaults to None.
+ """
+
+ pos_loss, neg_loss = heatmap_focal_loss_with_pos_inds(
+ pred,
+ target,
+ pos_inds,
+ alpha=self.alpha,
+ beta=self.beta,
+ gamma=self.gamma,
+ sigmoid_clamp=self.sigmoid_clamp,
+ ignore_high_fp=self.ignore_high_fp,
+ pos_weight=self.pos_weight,
+ neg_weight=self.neg_weight,
+ avg_factor=avg_factor)
+ return pos_loss, neg_loss
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/imagenet_lvis.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/imagenet_lvis.py
new file mode 100644
index 0000000000000000000000000000000000000000..3375a086682908ef8e68c5959ac5226eb6564ca5
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/imagenet_lvis.py
@@ -0,0 +1,395 @@
+# Copyright (c) OpenMMLab. All rights reserved.METAINFO
+import copy
+import os.path as osp
+import pickle
+import warnings
+from typing import List, Union
+
+from mmengine.fileio import get_local_path
+
+from mmdet.datasets import LVISV1Dataset
+from mmdet.registry import DATASETS
+
+
+@DATASETS.register_module()
+class ImageNetLVISV1Dataset(LVISV1Dataset):
+ """LVIS v1 dataset for detection."""
+
+ METAINFO = {
+ 'classes':
+ ('aerosol_can', 'air_conditioner', 'airplane', 'alarm_clock',
+ 'alcohol', 'alligator', 'almond', 'ambulance', 'amplifier', 'anklet',
+ 'antenna', 'apple', 'applesauce', 'apricot', 'apron', 'aquarium',
+ 'arctic_(type_of_shoe)', 'armband', 'armchair', 'armoire', 'armor',
+ 'artichoke', 'trash_can', 'ashtray', 'asparagus', 'atomizer',
+ 'avocado', 'award', 'awning', 'ax', 'baboon', 'baby_buggy',
+ 'basketball_backboard', 'backpack', 'handbag', 'suitcase', 'bagel',
+ 'bagpipe', 'baguet', 'bait', 'ball', 'ballet_skirt', 'balloon',
+ 'bamboo', 'banana', 'Band_Aid', 'bandage', 'bandanna', 'banjo',
+ 'banner', 'barbell', 'barge', 'barrel', 'barrette', 'barrow',
+ 'baseball_base', 'baseball', 'baseball_bat', 'baseball_cap',
+ 'baseball_glove', 'basket', 'basketball', 'bass_horn', 'bat_(animal)',
+ 'bath_mat', 'bath_towel', 'bathrobe', 'bathtub', 'batter_(food)',
+ 'battery', 'beachball', 'bead', 'bean_curd', 'beanbag', 'beanie',
+ 'bear', 'bed', 'bedpan', 'bedspread', 'cow', 'beef_(food)', 'beeper',
+ 'beer_bottle', 'beer_can', 'beetle', 'bell', 'bell_pepper', 'belt',
+ 'belt_buckle', 'bench', 'beret', 'bib', 'Bible', 'bicycle', 'visor',
+ 'billboard', 'binder', 'binoculars', 'bird', 'birdfeeder', 'birdbath',
+ 'birdcage', 'birdhouse', 'birthday_cake', 'birthday_card',
+ 'pirate_flag', 'black_sheep', 'blackberry', 'blackboard', 'blanket',
+ 'blazer', 'blender', 'blimp', 'blinker', 'blouse', 'blueberry',
+ 'gameboard', 'boat', 'bob', 'bobbin', 'bobby_pin', 'boiled_egg',
+ 'bolo_tie', 'deadbolt', 'bolt', 'bonnet', 'book', 'bookcase',
+ 'booklet', 'bookmark', 'boom_microphone', 'boot', 'bottle',
+ 'bottle_opener', 'bouquet', 'bow_(weapon)',
+ 'bow_(decorative_ribbons)', 'bow-tie', 'bowl', 'pipe_bowl',
+ 'bowler_hat', 'bowling_ball', 'box', 'boxing_glove', 'suspenders',
+ 'bracelet', 'brass_plaque', 'brassiere', 'bread-bin', 'bread',
+ 'breechcloth', 'bridal_gown', 'briefcase', 'broccoli', 'broach',
+ 'broom', 'brownie', 'brussels_sprouts', 'bubble_gum', 'bucket',
+ 'horse_buggy', 'bull', 'bulldog', 'bulldozer', 'bullet_train',
+ 'bulletin_board', 'bulletproof_vest', 'bullhorn', 'bun', 'bunk_bed',
+ 'buoy', 'burrito', 'bus_(vehicle)', 'business_card', 'butter',
+ 'butterfly', 'button', 'cab_(taxi)', 'cabana', 'cabin_car', 'cabinet',
+ 'locker', 'cake', 'calculator', 'calendar', 'calf', 'camcorder',
+ 'camel', 'camera', 'camera_lens', 'camper_(vehicle)', 'can',
+ 'can_opener', 'candle', 'candle_holder', 'candy_bar', 'candy_cane',
+ 'walking_cane', 'canister', 'canoe', 'cantaloup', 'canteen',
+ 'cap_(headwear)', 'bottle_cap', 'cape', 'cappuccino',
+ 'car_(automobile)', 'railcar_(part_of_a_train)', 'elevator_car',
+ 'car_battery', 'identity_card', 'card', 'cardigan', 'cargo_ship',
+ 'carnation', 'horse_carriage', 'carrot', 'tote_bag', 'cart', 'carton',
+ 'cash_register', 'casserole', 'cassette', 'cast', 'cat',
+ 'cauliflower', 'cayenne_(spice)', 'CD_player', 'celery',
+ 'cellular_telephone', 'chain_mail', 'chair', 'chaise_longue',
+ 'chalice', 'chandelier', 'chap', 'checkbook', 'checkerboard',
+ 'cherry', 'chessboard', 'chicken_(animal)', 'chickpea',
+ 'chili_(vegetable)', 'chime', 'chinaware', 'crisp_(potato_chip)',
+ 'poker_chip', 'chocolate_bar', 'chocolate_cake', 'chocolate_milk',
+ 'chocolate_mousse', 'choker', 'chopping_board', 'chopstick',
+ 'Christmas_tree', 'slide', 'cider', 'cigar_box', 'cigarette',
+ 'cigarette_case', 'cistern', 'clarinet', 'clasp', 'cleansing_agent',
+ 'cleat_(for_securing_rope)', 'clementine', 'clip', 'clipboard',
+ 'clippers_(for_plants)', 'cloak', 'clock', 'clock_tower',
+ 'clothes_hamper', 'clothespin', 'clutch_bag', 'coaster', 'coat',
+ 'coat_hanger', 'coatrack', 'cock', 'cockroach', 'cocoa_(beverage)',
+ 'coconut', 'coffee_maker', 'coffee_table', 'coffeepot', 'coil',
+ 'coin', 'colander', 'coleslaw', 'coloring_material',
+ 'combination_lock', 'pacifier', 'comic_book', 'compass',
+ 'computer_keyboard', 'condiment', 'cone', 'control',
+ 'convertible_(automobile)', 'sofa_bed', 'cooker', 'cookie',
+ 'cooking_utensil', 'cooler_(for_food)', 'cork_(bottle_plug)',
+ 'corkboard', 'corkscrew', 'edible_corn', 'cornbread', 'cornet',
+ 'cornice', 'cornmeal', 'corset', 'costume', 'cougar', 'coverall',
+ 'cowbell', 'cowboy_hat', 'crab_(animal)', 'crabmeat', 'cracker',
+ 'crape', 'crate', 'crayon', 'cream_pitcher', 'crescent_roll', 'crib',
+ 'crock_pot', 'crossbar', 'crouton', 'crow', 'crowbar', 'crown',
+ 'crucifix', 'cruise_ship', 'police_cruiser', 'crumb', 'crutch',
+ 'cub_(animal)', 'cube', 'cucumber', 'cufflink', 'cup', 'trophy_cup',
+ 'cupboard', 'cupcake', 'hair_curler', 'curling_iron', 'curtain',
+ 'cushion', 'cylinder', 'cymbal', 'dagger', 'dalmatian', 'dartboard',
+ 'date_(fruit)', 'deck_chair', 'deer', 'dental_floss', 'desk',
+ 'detergent', 'diaper', 'diary', 'die', 'dinghy', 'dining_table',
+ 'tux', 'dish', 'dish_antenna', 'dishrag', 'dishtowel', 'dishwasher',
+ 'dishwasher_detergent', 'dispenser', 'diving_board', 'Dixie_cup',
+ 'dog', 'dog_collar', 'doll', 'dollar', 'dollhouse', 'dolphin',
+ 'domestic_ass', 'doorknob', 'doormat', 'doughnut', 'dove',
+ 'dragonfly', 'drawer', 'underdrawers', 'dress', 'dress_hat',
+ 'dress_suit', 'dresser', 'drill', 'drone', 'dropper',
+ 'drum_(musical_instrument)', 'drumstick', 'duck', 'duckling',
+ 'duct_tape', 'duffel_bag', 'dumbbell', 'dumpster', 'dustpan', 'eagle',
+ 'earphone', 'earplug', 'earring', 'easel', 'eclair', 'eel', 'egg',
+ 'egg_roll', 'egg_yolk', 'eggbeater', 'eggplant', 'electric_chair',
+ 'refrigerator', 'elephant', 'elk', 'envelope', 'eraser', 'escargot',
+ 'eyepatch', 'falcon', 'fan', 'faucet', 'fedora', 'ferret',
+ 'Ferris_wheel', 'ferry', 'fig_(fruit)', 'fighter_jet', 'figurine',
+ 'file_cabinet', 'file_(tool)', 'fire_alarm', 'fire_engine',
+ 'fire_extinguisher', 'fire_hose', 'fireplace', 'fireplug',
+ 'first-aid_kit', 'fish', 'fish_(food)', 'fishbowl', 'fishing_rod',
+ 'flag', 'flagpole', 'flamingo', 'flannel', 'flap', 'flash',
+ 'flashlight', 'fleece', 'flip-flop_(sandal)', 'flipper_(footwear)',
+ 'flower_arrangement', 'flute_glass', 'foal', 'folding_chair',
+ 'food_processor', 'football_(American)', 'football_helmet',
+ 'footstool', 'fork', 'forklift', 'freight_car', 'French_toast',
+ 'freshener', 'frisbee', 'frog', 'fruit_juice', 'frying_pan', 'fudge',
+ 'funnel', 'futon', 'gag', 'garbage', 'garbage_truck', 'garden_hose',
+ 'gargle', 'gargoyle', 'garlic', 'gasmask', 'gazelle', 'gelatin',
+ 'gemstone', 'generator', 'giant_panda', 'gift_wrap', 'ginger',
+ 'giraffe', 'cincture', 'glass_(drink_container)', 'globe', 'glove',
+ 'goat', 'goggles', 'goldfish', 'golf_club', 'golfcart',
+ 'gondola_(boat)', 'goose', 'gorilla', 'gourd', 'grape', 'grater',
+ 'gravestone', 'gravy_boat', 'green_bean', 'green_onion', 'griddle',
+ 'grill', 'grits', 'grizzly', 'grocery_bag', 'guitar', 'gull', 'gun',
+ 'hairbrush', 'hairnet', 'hairpin', 'halter_top', 'ham', 'hamburger',
+ 'hammer', 'hammock', 'hamper', 'hamster', 'hair_dryer', 'hand_glass',
+ 'hand_towel', 'handcart', 'handcuff', 'handkerchief', 'handle',
+ 'handsaw', 'hardback_book', 'harmonium', 'hat', 'hatbox', 'veil',
+ 'headband', 'headboard', 'headlight', 'headscarf', 'headset',
+ 'headstall_(for_horses)', 'heart', 'heater', 'helicopter', 'helmet',
+ 'heron', 'highchair', 'hinge', 'hippopotamus', 'hockey_stick', 'hog',
+ 'home_plate_(baseball)', 'honey', 'fume_hood', 'hook', 'hookah',
+ 'hornet', 'horse', 'hose', 'hot-air_balloon', 'hotplate', 'hot_sauce',
+ 'hourglass', 'houseboat', 'hummingbird', 'hummus', 'polar_bear',
+ 'icecream', 'popsicle', 'ice_maker', 'ice_pack', 'ice_skate',
+ 'igniter', 'inhaler', 'iPod', 'iron_(for_clothing)', 'ironing_board',
+ 'jacket', 'jam', 'jar', 'jean', 'jeep', 'jelly_bean', 'jersey',
+ 'jet_plane', 'jewel', 'jewelry', 'joystick', 'jumpsuit', 'kayak',
+ 'keg', 'kennel', 'kettle', 'key', 'keycard', 'kilt', 'kimono',
+ 'kitchen_sink', 'kitchen_table', 'kite', 'kitten', 'kiwi_fruit',
+ 'knee_pad', 'knife', 'knitting_needle', 'knob', 'knocker_(on_a_door)',
+ 'koala', 'lab_coat', 'ladder', 'ladle', 'ladybug', 'lamb_(animal)',
+ 'lamb-chop', 'lamp', 'lamppost', 'lampshade', 'lantern', 'lanyard',
+ 'laptop_computer', 'lasagna', 'latch', 'lawn_mower', 'leather',
+ 'legging_(clothing)', 'Lego', 'legume', 'lemon', 'lemonade',
+ 'lettuce', 'license_plate', 'life_buoy', 'life_jacket', 'lightbulb',
+ 'lightning_rod', 'lime', 'limousine', 'lion', 'lip_balm', 'liquor',
+ 'lizard', 'log', 'lollipop', 'speaker_(stereo_equipment)', 'loveseat',
+ 'machine_gun', 'magazine', 'magnet', 'mail_slot', 'mailbox_(at_home)',
+ 'mallard', 'mallet', 'mammoth', 'manatee', 'mandarin_orange',
+ 'manger', 'manhole', 'map', 'marker', 'martini', 'mascot',
+ 'mashed_potato', 'masher', 'mask', 'mast', 'mat_(gym_equipment)',
+ 'matchbox', 'mattress', 'measuring_cup', 'measuring_stick',
+ 'meatball', 'medicine', 'melon', 'microphone', 'microscope',
+ 'microwave_oven', 'milestone', 'milk', 'milk_can', 'milkshake',
+ 'minivan', 'mint_candy', 'mirror', 'mitten', 'mixer_(kitchen_tool)',
+ 'money', 'monitor_(computer_equipment) computer_monitor', 'monkey',
+ 'motor', 'motor_scooter', 'motor_vehicle', 'motorcycle',
+ 'mound_(baseball)', 'mouse_(computer_equipment)', 'mousepad',
+ 'muffin', 'mug', 'mushroom', 'music_stool', 'musical_instrument',
+ 'nailfile', 'napkin', 'neckerchief', 'necklace', 'necktie', 'needle',
+ 'nest', 'newspaper', 'newsstand', 'nightshirt',
+ 'nosebag_(for_animals)', 'noseband_(for_animals)', 'notebook',
+ 'notepad', 'nut', 'nutcracker', 'oar', 'octopus_(food)',
+ 'octopus_(animal)', 'oil_lamp', 'olive_oil', 'omelet', 'onion',
+ 'orange_(fruit)', 'orange_juice', 'ostrich', 'ottoman', 'oven',
+ 'overalls_(clothing)', 'owl', 'packet', 'inkpad', 'pad', 'paddle',
+ 'padlock', 'paintbrush', 'painting', 'pajamas', 'palette',
+ 'pan_(for_cooking)', 'pan_(metal_container)', 'pancake', 'pantyhose',
+ 'papaya', 'paper_plate', 'paper_towel', 'paperback_book',
+ 'paperweight', 'parachute', 'parakeet', 'parasail_(sports)',
+ 'parasol', 'parchment', 'parka', 'parking_meter', 'parrot',
+ 'passenger_car_(part_of_a_train)', 'passenger_ship', 'passport',
+ 'pastry', 'patty_(food)', 'pea_(food)', 'peach', 'peanut_butter',
+ 'pear', 'peeler_(tool_for_fruit_and_vegetables)', 'wooden_leg',
+ 'pegboard', 'pelican', 'pen', 'pencil', 'pencil_box',
+ 'pencil_sharpener', 'pendulum', 'penguin', 'pennant', 'penny_(coin)',
+ 'pepper', 'pepper_mill', 'perfume', 'persimmon', 'person', 'pet',
+ 'pew_(church_bench)', 'phonebook', 'phonograph_record', 'piano',
+ 'pickle', 'pickup_truck', 'pie', 'pigeon', 'piggy_bank', 'pillow',
+ 'pin_(non_jewelry)', 'pineapple', 'pinecone', 'ping-pong_ball',
+ 'pinwheel', 'tobacco_pipe', 'pipe', 'pistol', 'pita_(bread)',
+ 'pitcher_(vessel_for_liquid)', 'pitchfork', 'pizza', 'place_mat',
+ 'plate', 'platter', 'playpen', 'pliers', 'plow_(farm_equipment)',
+ 'plume', 'pocket_watch', 'pocketknife', 'poker_(fire_stirring_tool)',
+ 'pole', 'polo_shirt', 'poncho', 'pony', 'pool_table', 'pop_(soda)',
+ 'postbox_(public)', 'postcard', 'poster', 'pot', 'flowerpot',
+ 'potato', 'potholder', 'pottery', 'pouch', 'power_shovel', 'prawn',
+ 'pretzel', 'printer', 'projectile_(weapon)', 'projector', 'propeller',
+ 'prune', 'pudding', 'puffer_(fish)', 'puffin', 'pug-dog', 'pumpkin',
+ 'puncher', 'puppet', 'puppy', 'quesadilla', 'quiche', 'quilt',
+ 'rabbit', 'race_car', 'racket', 'radar', 'radiator', 'radio_receiver',
+ 'radish', 'raft', 'rag_doll', 'raincoat', 'ram_(animal)', 'raspberry',
+ 'rat', 'razorblade', 'reamer_(juicer)', 'rearview_mirror', 'receipt',
+ 'recliner', 'record_player', 'reflector', 'remote_control',
+ 'rhinoceros', 'rib_(food)', 'rifle', 'ring', 'river_boat', 'road_map',
+ 'robe', 'rocking_chair', 'rodent', 'roller_skate', 'Rollerblade',
+ 'rolling_pin', 'root_beer', 'router_(computer_equipment)',
+ 'rubber_band', 'runner_(carpet)', 'plastic_bag',
+ 'saddle_(on_an_animal)', 'saddle_blanket', 'saddlebag', 'safety_pin',
+ 'sail', 'salad', 'salad_plate', 'salami', 'salmon_(fish)',
+ 'salmon_(food)', 'salsa', 'saltshaker', 'sandal_(type_of_shoe)',
+ 'sandwich', 'satchel', 'saucepan', 'saucer', 'sausage', 'sawhorse',
+ 'saxophone', 'scale_(measuring_instrument)', 'scarecrow', 'scarf',
+ 'school_bus', 'scissors', 'scoreboard', 'scraper', 'screwdriver',
+ 'scrubbing_brush', 'sculpture', 'seabird', 'seahorse', 'seaplane',
+ 'seashell', 'sewing_machine', 'shaker', 'shampoo', 'shark',
+ 'sharpener', 'Sharpie', 'shaver_(electric)', 'shaving_cream', 'shawl',
+ 'shears', 'sheep', 'shepherd_dog', 'sherbert', 'shield', 'shirt',
+ 'shoe', 'shopping_bag', 'shopping_cart', 'short_pants', 'shot_glass',
+ 'shoulder_bag', 'shovel', 'shower_head', 'shower_cap',
+ 'shower_curtain', 'shredder_(for_paper)', 'signboard', 'silo', 'sink',
+ 'skateboard', 'skewer', 'ski', 'ski_boot', 'ski_parka', 'ski_pole',
+ 'skirt', 'skullcap', 'sled', 'sleeping_bag', 'sling_(bandage)',
+ 'slipper_(footwear)', 'smoothie', 'snake', 'snowboard', 'snowman',
+ 'snowmobile', 'soap', 'soccer_ball', 'sock', 'sofa', 'softball',
+ 'solar_array', 'sombrero', 'soup', 'soup_bowl', 'soupspoon',
+ 'sour_cream', 'soya_milk', 'space_shuttle', 'sparkler_(fireworks)',
+ 'spatula', 'spear', 'spectacles', 'spice_rack', 'spider', 'crawfish',
+ 'sponge', 'spoon', 'sportswear', 'spotlight', 'squid_(food)',
+ 'squirrel', 'stagecoach', 'stapler_(stapling_machine)', 'starfish',
+ 'statue_(sculpture)', 'steak_(food)', 'steak_knife', 'steering_wheel',
+ 'stepladder', 'step_stool', 'stereo_(sound_system)', 'stew',
+ 'stirrer', 'stirrup', 'stool', 'stop_sign', 'brake_light', 'stove',
+ 'strainer', 'strap', 'straw_(for_drinking)', 'strawberry',
+ 'street_sign', 'streetlight', 'string_cheese', 'stylus', 'subwoofer',
+ 'sugar_bowl', 'sugarcane_(plant)', 'suit_(clothing)', 'sunflower',
+ 'sunglasses', 'sunhat', 'surfboard', 'sushi', 'mop', 'sweat_pants',
+ 'sweatband', 'sweater', 'sweatshirt', 'sweet_potato', 'swimsuit',
+ 'sword', 'syringe', 'Tabasco_sauce', 'table-tennis_table', 'table',
+ 'table_lamp', 'tablecloth', 'tachometer', 'taco', 'tag', 'taillight',
+ 'tambourine', 'army_tank', 'tank_(storage_vessel)',
+ 'tank_top_(clothing)', 'tape_(sticky_cloth_or_paper)', 'tape_measure',
+ 'tapestry', 'tarp', 'tartan', 'tassel', 'tea_bag', 'teacup',
+ 'teakettle', 'teapot', 'teddy_bear', 'telephone', 'telephone_booth',
+ 'telephone_pole', 'telephoto_lens', 'television_camera',
+ 'television_set', 'tennis_ball', 'tennis_racket', 'tequila',
+ 'thermometer', 'thermos_bottle', 'thermostat', 'thimble', 'thread',
+ 'thumbtack', 'tiara', 'tiger', 'tights_(clothing)', 'timer',
+ 'tinfoil', 'tinsel', 'tissue_paper', 'toast_(food)', 'toaster',
+ 'toaster_oven', 'toilet', 'toilet_tissue', 'tomato', 'tongs',
+ 'toolbox', 'toothbrush', 'toothpaste', 'toothpick', 'cover',
+ 'tortilla', 'tow_truck', 'towel', 'towel_rack', 'toy',
+ 'tractor_(farm_equipment)', 'traffic_light', 'dirt_bike',
+ 'trailer_truck', 'train_(railroad_vehicle)', 'trampoline', 'tray',
+ 'trench_coat', 'triangle_(musical_instrument)', 'tricycle', 'tripod',
+ 'trousers', 'truck', 'truffle_(chocolate)', 'trunk', 'vat', 'turban',
+ 'turkey_(food)', 'turnip', 'turtle', 'turtleneck_(clothing)',
+ 'typewriter', 'umbrella', 'underwear', 'unicycle', 'urinal', 'urn',
+ 'vacuum_cleaner', 'vase', 'vending_machine', 'vent', 'vest',
+ 'videotape', 'vinegar', 'violin', 'vodka', 'volleyball', 'vulture',
+ 'waffle', 'waffle_iron', 'wagon', 'wagon_wheel', 'walking_stick',
+ 'wall_clock', 'wall_socket', 'wallet', 'walrus', 'wardrobe',
+ 'washbasin', 'automatic_washer', 'watch', 'water_bottle',
+ 'water_cooler', 'water_faucet', 'water_heater', 'water_jug',
+ 'water_gun', 'water_scooter', 'water_ski', 'water_tower',
+ 'watering_can', 'watermelon', 'weathervane', 'webcam', 'wedding_cake',
+ 'wedding_ring', 'wet_suit', 'wheel', 'wheelchair', 'whipped_cream',
+ 'whistle', 'wig', 'wind_chime', 'windmill', 'window_box_(for_plants)',
+ 'windshield_wiper', 'windsock', 'wine_bottle', 'wine_bucket',
+ 'wineglass', 'blinder_(for_horses)', 'wok', 'wolf', 'wooden_spoon',
+ 'wreath', 'wrench', 'wristband', 'wristlet', 'yacht', 'yogurt',
+ 'yoke_(animal_equipment)', 'zebra', 'zucchini'),
+ 'palette':
+ None
+ }
+
+ def get_data_info(self, idx: int) -> dict:
+ """Get annotation by index and automatically call ``full_init`` if the
+ dataset has not been fully initialized.
+
+ Args:
+ idx (int): The index of data.
+
+ Returns:
+ dict: The idx-th annotation of the dataset.
+ """
+ if self.serialize_data:
+ start_addr = 0 if idx == 0 else self.data_address[idx - 1].item()
+ end_addr = self.data_address[idx].item()
+ bytes = memoryview(
+ self.data_bytes[start_addr:end_addr]) # type: ignore
+ data_info = pickle.loads(bytes) # type: ignore
+ else:
+ data_info = copy.deepcopy(self.data_list[idx])
+
+ # Some codebase needs `sample_idx` of data information. Here we convert
+ # the idx to a positive number and save it in data information.
+ if idx >= 0:
+ data_info['sample_idx'] = idx
+ else:
+ data_info['sample_idx'] = len(self) + idx
+
+ return data_info
+
+ def load_data_list(self) -> List[dict]:
+ """Load annotations from an annotation file named as ``self.ann_file``
+
+ Returns:
+ List[dict]: A list of annotation.
+ """ # noqa: E501
+ try:
+ import lvis
+ if getattr(lvis, '__version__', '0') >= '10.5.3':
+ warnings.warn(
+ 'mmlvis is deprecated, please install official lvis-api by "pip install git+https://github.com/lvis-dataset/lvis-api.git"', # noqa: E501
+ UserWarning)
+ from lvis import LVIS
+ except ImportError:
+ raise ImportError(
+ 'Package lvis is not installed. Please run "pip install git+https://github.com/lvis-dataset/lvis-api.git".' # noqa: E501
+ )
+ with get_local_path(
+ self.ann_file, backend_args=self.backend_args) as local_path:
+ self.lvis = LVIS(local_path)
+ self.cat_ids = self.lvis.get_cat_ids()
+ self.cat2label = {cat_id: i for i, cat_id in enumerate(self.cat_ids)}
+ self.cat_img_map = copy.deepcopy(self.lvis.cat_img_map)
+ img_ids = self.lvis.get_img_ids()
+ data_list = []
+ total_ann_ids = []
+ for img_id in img_ids:
+ raw_img_info = self.lvis.load_imgs([img_id])[0]
+ raw_img_info['img_id'] = img_id
+
+ ann_ids = self.lvis.get_ann_ids(img_ids=[img_id])
+ total_ann_ids.extend(ann_ids)
+ parsed_data_info = self.parse_data_info(
+ {'raw_img_info': raw_img_info})
+ data_list.append(parsed_data_info)
+ if self.ANN_ID_UNIQUE:
+ assert len(set(total_ann_ids)) == len(
+ total_ann_ids
+ ), f"Annotation ids in '{self.ann_file}' are not unique!"
+
+ del self.lvis
+ # print(data_list)
+ return data_list
+
+ def parse_data_info(self, raw_data_info: dict) -> Union[dict, List[dict]]:
+ """Parse raw annotation to target format.
+
+ Args:
+ raw_data_info (dict): Raw data information load from ``ann_file``
+
+ Returns:
+ Union[dict, List[dict]]: Parsed annotation.
+ """
+ img_info = raw_data_info['raw_img_info']
+
+ data_info = {}
+
+ # TODO: need to change data_prefix['img'] to data_prefix['img_path']
+ img_path = osp.join(self.data_prefix['img'], img_info['file_name'])
+ if self.data_prefix.get('seg', None):
+ seg_map_path = osp.join(
+ self.data_prefix['seg'],
+ img_info['file_name'].rsplit('.', 1)[0] + self.seg_map_suffix)
+ else:
+ seg_map_path = None
+ data_info['img_path'] = img_path
+ data_info['img_id'] = img_info['img_id']
+ data_info['seg_map_path'] = seg_map_path
+ data_info['height'] = img_info['height']
+ data_info['width'] = img_info['width']
+
+ if self.return_classes:
+ data_info['text'] = self.metainfo['classes']
+ data_info['custom_entities'] = True
+
+ instances = []
+ image_labels = [
+ self.cat2label[x] for x in img_info['pos_category_ids']
+ ]
+ for image_label in image_labels:
+ instance = {}
+ instance['bbox_label'] = image_label
+ instances.append(instance)
+ data_info['instances'] = instances
+
+ return data_info
+
+ def get_cat_ids(self, idx: int) -> List[int]:
+ """Get COCO category ids by index.
+
+ Args:
+ idx (int): Index of data.
+
+ Returns:
+ List[int]: All categories in the image of specified index.
+ """
+ data_info = self.get_data_info(idx)
+ image_labels = []
+ for instance in data_info['instances']:
+ image_labels.append(instance['bbox_label'])
+
+ return image_labels
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/iou_loss.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/iou_loss.py
new file mode 100644
index 0000000000000000000000000000000000000000..349545cf54d379c905047f9ac6d635bd5f9217be
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/iou_loss.py
@@ -0,0 +1,125 @@
+import torch
+from torch import nn
+
+
+# support calculate IOULoss with box_pred
+class IOULoss(nn.Module):
+
+ def __init__(self, loc_loss_type='iou'):
+ super(IOULoss, self).__init__()
+ self.loc_loss_type = loc_loss_type
+
+ def forward(self, pred, target, weight=None, reduction='sum'):
+ pred_left = pred[:, 0]
+ pred_top = pred[:, 1]
+ pred_right = pred[:, 2]
+ pred_bottom = pred[:, 3]
+
+ target_left = target[:, 0]
+ target_top = target[:, 1]
+ target_right = target[:, 2]
+ target_bottom = target[:, 3]
+
+ target_aera = (target_left + target_right) * (
+ target_top + target_bottom)
+ pred_aera = (pred_left + pred_right) * (pred_top + pred_bottom)
+
+ w_intersect = torch.min(pred_left, target_left) + torch.min(
+ pred_right, target_right)
+ h_intersect = torch.min(pred_bottom, target_bottom) + torch.min(
+ pred_top, target_top)
+
+ g_w_intersect = torch.max(pred_left, target_left) + torch.max(
+ pred_right, target_right)
+ g_h_intersect = torch.max(pred_bottom, target_bottom) + torch.max(
+ pred_top, target_top)
+ ac_uion = g_w_intersect * g_h_intersect
+
+ area_intersect = w_intersect * h_intersect
+ area_union = target_aera + pred_aera - area_intersect
+
+ ious = (area_intersect + 1.0) / (area_union + 1.0)
+ gious = ious - (ac_uion - area_union) / ac_uion
+ if self.loc_loss_type == 'iou':
+ losses = -torch.log(ious)
+ elif self.loc_loss_type == 'linear_iou':
+ losses = 1 - ious
+ elif self.loc_loss_type == 'giou':
+ losses = 1 - gious
+ else:
+ raise NotImplementedError
+
+ if weight is not None:
+ losses = losses * weight
+ else:
+ losses = losses
+
+ if reduction == 'sum':
+ return losses.sum()
+ elif reduction == 'batch':
+ return losses.sum(dim=[1])
+ elif reduction == 'none':
+ return losses
+ else:
+ raise NotImplementedError
+
+
+def giou_loss(
+ boxes1: torch.Tensor,
+ boxes2: torch.Tensor,
+ reduction: str = 'none',
+ eps: float = 1e-7,
+) -> torch.Tensor:
+ """Generalized Intersection over Union Loss (Hamid Rezatofighi et.
+
+ al)
+ https://arxiv.org/abs/1902.09630
+ Gradient-friendly IoU loss with an additional penalty that is
+ non-zero when the boxes do not overlap and scales with the size
+ of their smallest enclosing box. This loss is symmetric, so the
+ boxes1 and boxes2 arguments are interchangeable.
+ Args:
+ boxes1, boxes2 (Tensor): box locations in XYXY format, shape
+ (N, 4) or (4,).
+ reduction: 'none' | 'mean' | 'sum'
+ 'none': No reduction will be applied to the output.
+ 'mean': The output will be averaged.
+ 'sum': The output will be summed.
+ eps (float): small number to prevent division by zero
+ """
+
+ x1, y1, x2, y2 = boxes1.unbind(dim=-1)
+ x1g, y1g, x2g, y2g = boxes2.unbind(dim=-1)
+
+ assert (x2 >= x1).all(), 'bad box: x1 larger than x2'
+ assert (y2 >= y1).all(), 'bad box: y1 larger than y2'
+
+ # Intersection keypoints
+ xkis1 = torch.max(x1, x1g)
+ ykis1 = torch.max(y1, y1g)
+ xkis2 = torch.min(x2, x2g)
+ ykis2 = torch.min(y2, y2g)
+
+ intsctk = torch.zeros_like(x1)
+ mask = (ykis2 > ykis1) & (xkis2 > xkis1)
+ intsctk[mask] = (xkis2[mask] - xkis1[mask]) * (ykis2[mask] - ykis1[mask])
+ unionk = (x2 - x1) * (y2 - y1) + (x2g - x1g) * (y2g - y1g) - intsctk
+ iouk = intsctk / (unionk + eps)
+
+ # smallest enclosing box
+ xc1 = torch.min(x1, x1g)
+ yc1 = torch.min(y1, y1g)
+ xc2 = torch.max(x2, x2g)
+ yc2 = torch.max(y2, y2g)
+
+ area_c = (xc2 - xc1) * (yc2 - yc1)
+ miouk = iouk - ((area_c - unionk) / (area_c + eps))
+
+ loss = 1 - miouk
+
+ if reduction == 'mean':
+ loss = loss.mean()
+ elif reduction == 'sum':
+ loss = loss.sum()
+
+ return loss
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/zero_shot_classifier.py b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/zero_shot_classifier.py
new file mode 100644
index 0000000000000000000000000000000000000000..cb9946d582501b2b75f844193c71ff338c384825
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/Detic_new/detic/zero_shot_classifier.py
@@ -0,0 +1,73 @@
+# Copyright (c) Facebook, Inc. and its affiliates.
+import numpy as np
+import torch
+from torch import nn
+from torch.nn import functional as F
+
+from mmdet.registry import MODELS
+
+
+@MODELS.register_module()
+class ZeroShotClassifier(nn.Module):
+
+ def __init__(
+ self,
+ in_features: int,
+ out_features: int, # num_classes
+ zs_weight_path: str,
+ zs_weight_dim: int = 512,
+ use_bias: float = 0.0,
+ norm_weight: bool = True,
+ norm_temperature: float = 50.0,
+ ):
+ super().__init__()
+ num_classes = out_features
+ self.norm_weight = norm_weight
+ self.norm_temperature = norm_temperature
+
+ self.use_bias = use_bias < 0
+ if self.use_bias:
+ self.cls_bias = nn.Parameter(torch.ones(1) * use_bias)
+
+ self.linear = nn.Linear(in_features, zs_weight_dim)
+
+ if zs_weight_path == 'rand':
+ zs_weight = torch.randn((zs_weight_dim, num_classes))
+ nn.init.normal_(zs_weight, std=0.01)
+ else:
+ zs_weight = torch.tensor(
+ np.load(zs_weight_path),
+ dtype=torch.float32).permute(1, 0).contiguous() # D x C
+ zs_weight = torch.cat(
+ [zs_weight, zs_weight.new_zeros(
+ (zs_weight_dim, 1))], dim=1) # D x (C + 1)
+
+ if self.norm_weight:
+ zs_weight = F.normalize(zs_weight, p=2, dim=0)
+
+ if zs_weight_path == 'rand':
+ self.zs_weight = nn.Parameter(zs_weight)
+ else:
+ self.register_buffer('zs_weight', zs_weight)
+
+ assert self.zs_weight.shape[1] == num_classes + 1, self.zs_weight.shape
+
+ def forward(self, x, classifier=None):
+ '''
+ Inputs:
+ x: B x D'
+ classifier_info: (C', C' x D)
+ '''
+ x = self.linear(x)
+ if classifier is not None:
+ zs_weight = classifier.permute(1, 0).contiguous() # D x C'
+ zs_weight = F.normalize(zs_weight, p=2, dim=0) \
+ if self.norm_weight else zs_weight
+ else:
+ zs_weight = self.zs_weight
+ if self.norm_weight:
+ x = self.norm_temperature * F.normalize(x, p=2, dim=1)
+ x = torch.mm(x, zs_weight)
+ if self.use_bias:
+ x = x + self.cls_bias
+ return x
diff --git a/scripts/evaluation/FasterRCNN_score-mmdet/projects/DiffusionDet/README.md b/scripts/evaluation/FasterRCNN_score-mmdet/projects/DiffusionDet/README.md
new file mode 100644
index 0000000000000000000000000000000000000000..5542d9a59a041c2983b3bb7a9fc77f599148ce72
--- /dev/null
+++ b/scripts/evaluation/FasterRCNN_score-mmdet/projects/DiffusionDet/README.md
@@ -0,0 +1,172 @@
+## Description
+
+This is an implementation of [DiffusionDet](https://github.com/ShoufaChen/DiffusionDet) based on [MMDetection](https://github.com/open-mmlab/mmdetection/tree/main), [MMCV](https://github.com/open-mmlab/mmcv), and [MMEngine](https://github.com/open-mmlab/mmengine).
+
+
+